e:/uni_stuff/papers/routinisation-in-dialogue-d&d'10/gargett-main.dvi dialogue and discourse 2(1) (2011) 171–197 doi: 10.5087/dad.2011.108 incrementality and the dynamics of routines in dialogue andrew gargett andrew.gargett@uaeu.ac.ae department of linguistics u.a.e. university al ain, u.a.e. editor: hannes reiser and david schlangen abstract we propose a novel dual processing model of linguistic routinisation, specifically formulaic expressions (from relatively fixed idioms, all the way throughto looser collocational phenomena). this model is formalised using the dynamic syntax (ds) formal account of language processing, whereby we make a specific extension to the core ds lexical architecture to capture the dynamics of linguistic routinisation. this extension is inspired by work within cognitive science more broadly. ds has a range of attractive modelling features, such as fullincrementality, as well as recent accounts of using resources of the core grammar for modelling arange of dialogue phenomena, all of which we deploy in our account. this leads to not only a fully incremental model of formulaic language, but further, this straightforwardly extends to routinised dialogue phenomena. we consider this approach to be a proof of concept of how interdisciplinary work within cognitive science holds out the promise of meeting challenges faced by modellers of dialogue and discourse. c©2011 andrew gargett submitted 1/2010; accepted 11/2010; published online 5/2011 gargett 1. orientation in this paper we propose a unified approach to the relation between linguisticknowledge and linguistic experience, specifically, we present a new approach to modelling linguistic routinisation which, among other things, offers a way of capturing within the same framework the use of both formulaic and non-formulaic language in dialogue. as we discuss below, these phenomena have numerous distinct properties, yet they also share features directly relevant for formal modelling (see nunberg et al. (1994)). focusing on actual language use, we formally model the relative incrementality of both formulaic and non-formulaic language. our theoretical framework is inspired by work on dual process models ofcognitive phenomena, in particular interaction within dialogue (barr and keysar (2006) provide inpart a recent overview). specifically, we model the interaction between rule-based and memory-based processing of natural language, formally implementing this within dynamic syntax (ds, kempson et al. (2001), cann et al. (2005)). while ds has typically focused on rule-based processing, we seek to extend this by arguing that actual patterns of linguistic phenomena, as found for examplein dialogue, emerge out of the interaction between these distinct processes. for us, processingformulaic language is more likely to involve retrieval of items stored as wholes than computed online (details about this below), compared to the rule-driven processing underlying non-formulaic language. for this formalisation of dual processing, we extend the lexical architecture of ds, so that lexical entries incorporate both the usual ds lexical actions, but also include the output semantic structures which result from employing such lexical actions. this sets up two competing processes for updating the representation being constructed for a speaker’s utterance,a slower one based on lexical actions, and a faster one based on stored semantic structure (see gargett (2010) for details). importantly, in making this extension, we are concerned with retaining attractiveproperties of the ds account. in particular, we aim to preserve the dynamics of the model, retaining fully incremental processing within the context of interaction. we are able to effectively account for the dynamics of routinised dialogue phenomena via modifications to the ds model at the lexical level alone. this might be viewed as one in a line of recent proposals for lexicalist modelling ofdialogue phenomena (e.g. kecskes (2008)). 172 incrementality and routines in dialogue 2. motivations our approach is distinct from previous dual process models of dialoguein that we focus on core grammatical resources to model interaction, in line with previous dynamic syntaxwork (purver et al. (2006), gargett et al. (2009)), aiming to investigate the extent to which dialogue can be modelled using mechanisms specified by the grammar.1 this mechanistic approach to modelling formulaic language, ormulti-word expressions(mwes), makes a break with the orthodoxy of property-list approaches which define mwes via a list of linguistic properties, such as relative compositionality, idiosyncrasy of meaning, selectivity for discontinuous formulaic expressions, etc (e.g. nunberg et al. (1994)). we avoid modelling directly in terms of properties, following thereasoning in rawson (2004), regarding drawbacks of this. instead, rawson suggests mechanistic models (specifically those of logan (1988) and anderson (1992)), which we find fit remarkably well to our adopted approach to incremental processing. for our purposes, linguistic routinisation involves long-term storing and reuse of context, with context taken to be the previous words plus their mode of construal (detailsto be made precise). interlocutors may routinise the grammatical or semantic aspects of words or phrases, within a single turn or across multiple turns. such routinisation is highly sensitive to specific features of the context of an interaction, such features typically triggering the routine. a note: defining the time periods relevant for the emergence of routines is somewhat problematic (given thefuzziness of notions of language, language use and context); here we will define short-term reuse as reuse of words and their construal immediately following the initial use, medium term as reuse later in theinteraction, and long-term as reuse on some subsequent occasion of interaction. 2.1 dynamics of incrementality milward (1994) proposes that incrementality involves: (i) as much information being extracted as soon as possible (ii) carried out in small steps approximately as each word is encountered in a way that we will make more precise below, we can think of incremental processing of an incoming string of words as a kind ofstepping throughthe string of items while constructing the unfolding representation. then the processing of formulaic language can be seen asskipping over chunks of items rather than stepping through every possible individual item.consider how a hearer might process the information provided by an utterance of: (1) ‘bob left’ natural language utterances, like all kinds of natural phenomena, unfold in time, such temporal dimension of processing being a feature ofdynamical systems(ward (2002)). such systems are inherently incremental, with input updating one state to the next. first, the occurrence of the name ‘bob’: (2) state1 bob → state2 1. to a first approximation, and for ease of exposition, grammatical mechanisms are here simply taken to be the rules and representations of an adopted formal grammar model, in this case ds (analogous to the second representational level in marr (1982), that of the algorithms and data structures of some formal model). 173 gargett is the initial term in the unfolding utterance, and also licenses the hearer to expect an upcoming predicate, among other terms, in the construal of that utterance. further,the hearer could even try to make a guess as to which bob is being referred to (assuming a place-holdermodel of names). such considerations suggest this kind of processing is goal-driven. presuming an intuitive, evidence-driven model of communication,2 interlocutors are presumed to operate in a goal-driven way, guided by expectations about what is coming up (determined via current evidence from the ongoing interaction). the expectation of a predicate is satisfied by the occurrence of ‘left’, triggering a transition to the next state: (3) state2 left → state3 however, note that prior to the actual occurrence of this item, and the particular information it imparts (e.g. that no other items are required for a saturated proposition), any of a number of other ways of completing the utterance are possible (e.g. a transitive predicate such aslikes). another possibility is that of completing the utterance with the idiomatic ‘kicked himself’, processed as a chunk: (4) state2 kicked himself −→ state3 further, we might also wonder whether alternatives such as ‘kicked herself’, ‘kicked’ (non-idiomatic), etc, are available, and at what point (see discussion of implementation issues below). dynamical modelling involves observing changes to some phenomenon over timeat discrete time-steps (ward (2002)), whereby complex systems can be modelled via “snapshots” of successive stages in the system’s progress, these being idealised as successive states of the system, plus transitions between these states (milward (1994)). as a dynamic process, parsing can be characterised in terms of observed states and transitions between such states. some parsingresearch suggests such transitions are madeeagerlyrather than delayed (e.g. until more information is available), with problems arising if a chosen search path turns out to be wrong with respect to upcoming linguistic information (e.g. the so-called garden-pathing of expressions like ‘the horse raced past the barn fell’) (sturt and lombardo (2005)). so in this more abstract view of incrementality as stepping through stages in the development of an output process of some system, key implementation issues include (e.g. crocker (2010)): (i) the degree of eagerness, (ii) whether parallel or serial parsing strategies are employed, (iii) whether/not the process is monotonic. this suggests classifying incremental approaches depending on whether/not they exhibit such features. to this end, we aim for a unified account of formulaic and non-formulaic language with a parallel flavour (see section (2.3) below for details of how we go about this). 2.2 empirical details 2.2.1 idioms in dialogue let’s start with some clearly context-dependent set of phenomena provided by the following elliptical dialogue fragments ((a) – (e)): 2. which is to say, interlocutors communicate by cycling through stages of presenting evidence of the information they wish to impart (in the case of speakers), and evidence of the informationthey have accepted (in the case of hearers). 174 incrementality and routines in dialogue (5) a: have you seen mary? (a) b: mary? (b) b: no, i haven’t. (c) b: she’s out i think. (d) b: i have, bill. (e) b: no, nor bob. the interpretation of such fragments requires at least linguistic context (a’s question), as well as non-linguistic context. indeed, context either directly provides linguistic evidence for completion, or else this is more indirect. the examples in (5) above are all examples wherethe context provides direct linguistic input for processing the fragments, for example, a’s original utterance forming the context for b’s utterance in (a), providing predicate and subject information. a complication arises with a more indirect mode of construal occurring in so-called “sloppy” ellipsis. consider the following examples:3 (6) john took his clients to the cleaners, and so did bill “bill took john’s clients to the cleaners” (strict) “bill took bill’s clients to the cleaners” (sloppy) so in (6), we have a strict interpretation in which some content is taken directlyfrom context, but in addition, a sloppy reading, via reusing some aspect of meaning from context (the referent resolving the anaphoric reference of ‘his’) but nevertheless with distinct content.4 now, placing the idiom in example (6) within an elliptical context illustrates how idiomsare compositional and sensitive to items they combine with (see nunberg et al. (1994)): (7) john took his clients to the cleaners, but never his shirts in example (7), substituting an inanimate for an animate nominal (i.e. ‘shirts’ for‘clients’) removes the cues which trigger the idiom, resulting in a switch to non-formulaic processing, despitethe very same representationbeing reused for the second clause (see the analysis of ellipsis in cann etal. (2007)). example (7) demonstrates one kind of incrementality of formulaic language, and here is another in the context of dialogue: (8) a: bob took his clients to... b: the cleaners. a: actually, i was going to say to his uncle’s restaurant. splitting idioms across dialogue turns suggests their processing is as incremental as non-formulaic language. note, this involves not only the processing of strings, but alsothe construction of representations.this perhaps would not come as a surprise to those accounts which for the last decade have been arguing for the compositionality of formulaic language (e.g. nunberg et al. (1994)). 3. the meaning of ‘take np to the cleaners’ being “to take a significant quantity of np’s money or valuables, through gambling, unfavorable investing, fraud, litigation, etc.’ (http://en.wiktionary.org/wiki/taketo the cleaners). 4. these two readings are given a formally unitary account through theuse of abstraction and higher order unification by dalrymple et al. (dalrymple et al. (1991)), but the ds account proposed by purver et al. (2006) is that in which the actions used in building up interpretation are themselves stored in context and, in the sloppy forms of interpretation, re-used to yield distinct denotational content. 175 gargett 2.2.2 pre-fabs in maze descriptions a key focus here is the emergence within dialogue of looser forms of routinised language, such as pre-fab (i.e. conventionalised collocation, see bybee (2006), more details in section (2.3) below). pickering and garrod (2005) present an influential model of linguistic routinisation, based in large part on patterns discovered in maze task experiments. now, a robust observation about these experiments is that participants tend to produce a restricted number of so-calleddescription types, from more concrete (e.g.figural “the sticking out bit at the bottom”,line “at the end of bottom row”) to more abstract (e.g.path“two across, one up”,matrix “2,9”) ways of describing/conceptualising the maze. however, whether these categorisations in fact capture distinct, independenttypesof language is not entirely clear, given the complexity of the various constraints under which the linguistic system operates during dialogue.5 for example, figurative language might be employed in conditions of greater uncertainty (as proposed in bavelas (2009)). consider thisin terms of minimal effort vs. maximal effect: a more concrete yet elaborate description (e.g. “the left indicator bit on the right top corner”) could be harder to produce yet more likely understood by someone, whereas a more truncated and specialised form (e.g. “two across, one up”) will be easier to produce but understanding it may require task-relevant experience. an experimental task that naturally leads to routinisation is the maze location description task reported in healey (1997). here interlocutors communicate with each other inorder to identify a set of twenty maze locations. given the repetitive nature of this task, over the course of interaction, they tend to routinise these descriptions in predictable ways. consider the following example of one such maze location description: (9) (i) a: ummm, it’s the top right hand corner two down, (ii) b: two down right so [it’s the] (iii) a: [so it’s] kind of three down really but it’s only two down, (iv) b: okay (v) a: if there isn’t one in the top left hand corner, (vi) b: right, (vii) a: right? here the meaning of ‘down’ is negotiated quite explicitly, with respect to the specific context, namely the shape of the particular maze a is describing to b. this kind of dialogue routinisation involves linguistically encoding some aspect of non-linguistic context which the interlocutors have made mutually salient through their interaction (using specific linguistic resources useful for picking out bits of the world). importantly, routinisation adds to this the possibilitythat future similar interaction will reflect/reuse aspects of this particular interaction. yet, it has been observed that this trend is not unidirectional, and can in fact reverse, especially when problems arise (e.g. healey (2008)). for example, in subsequent interaction, ‘some numberdown’ is more likely to refer to a point from the top of the actual maze than, say, the putative top of the smallest square the entire maze fits into. interestingly, a and b might employ ‘down’ in this way in subsequent interactions with other interlocutors, and some times this may work, particularly if they are doing this against a background 5. at the very least, there is likely interaction between linguistic and conceptual systems, involving not only preferred ways of talking in interaction, but also preferred ways of conceptualisingmazes. 176 incrementality and routines in dialogue community-wide set of interactions. however, if in egocentric fashion they wrongly presume the same routine should work with any subsequent interlocutor,6 they may well run into difficulty that causes such routines to be ineffective. we are here interested in the question of what might they do next. one answer is that they will typically switch to other, more computational(rather than memory-based) kinds of linguistic processing, which we discuss elsewhere (e.g. gargett (2010)). 2.3 simulating routinisation let’s bring out some issues more distinctly by considering the life-cycle of a linguistic routine, adapting the mechanistic approaches in logan (1988), anderson (1992), rawson (2004). imagine a (perhaps unusual) adult speaker of english encountering the following for the first time: (10) bob pulled some strings at work. given the capacity to construct a representation of this string, complete with asuitable construal (e.g. determining that bob does not physically handle string for a living), withrepeated exposure to this phrase our adult speaker can store this representation which she can later re-use. for this speaker, ‘pull strings’ has become a pre-fab (i.e. conventionalised collocation, see bybee (2006)) with the sense of (broadly) “manipulate”. in line with our dual processing perspective, we further assume that updating proceeds via either the immediate use of lexical actions for constructing representations from scratch, or else the reuse of stored representationsoutputted from previous use of lexical actions. following logan, such interaction between competing optionscan be modelled as a race between them to effect update. at the beginning of the cycle of routinisation, rules may be favoured in suchraces due to their generality (cf. logan (1988), rawson (2004)). but over time, richercontexts accumulate for the output representations (via the surrounding words, etc) and within whichprocessing takes place, so that specialisation of these representations to specific contexts in which processingtypicallyoccurs, leads to their being favoured over rules.7 the intermediate term is marked by a period of shifting between one kind of processing and another, and over-specialisation of semantic structures can lead to these failing to respond in novel contexts. importantly,computational processing is still available when such failure occurs, and may in fact reappear in such cases.8 however, what about the longer term? it is here that rules make a resurgence, in the form of routines, or complexes of actions. idioms are a good example of this, and various earlier accounts have looked at the relative compositionality of these (e.g. nunberg et al. (1994)). an implication of logan’s model seldom taken up is that computational processing neveractually abandons the race, and might even somehow gain a competitive edge (even after some period of dominance by memory-based processing).9 we would suggest that various ways of effecting computational efficiency, such as production compilation (see taatgen et al. (2008)) could provide computational 6. or even if it is more mechanistic than this, say, triggered by cues within memory. 7. we are of course talking about a probabilistic phenomenon, and a probabilistic version of the ds parser is currently underway. 8. our approach to coordinating the two processes in this way, albeit indirectly, attempts to extend the model in logan (1988) for linguistic purposes, and is in line with suggestions in the linguistics literature (e.g. wray (2002) on holophrastic processing operating in tandem with more analytical processing, rawson and middleton (2009) on novelty and automaticity in text comprehension). 9. recently, rawson and middleton (2009) has discussed logan’s theory in similar terms (interestingly, by way of proposing an account of the response of automatic processing to novelty). 177 gargett processing with the necessary boost to eventually win out over memory-based processing, thus linking linguistic routinisation with models of automaticity in cognitive psychology. production compilation essentially involves linking together otherwise sequential, separateproductions (each with their own triggering conditions and output effects) into a complex unit with asingle triggering condition and combined output effects.10 2.4 previous accounts there is extensive work on routinisation (the process leading to the formation of routines) throughout the cognitive sciences. common to many such theories is the idea that routines arise from practice, becoming points of stability in the face of contextual exigencies, typically expressed as a list of various properties, such as their being: (i) rigid in form, (ii) somewhat truncated, (iii) relatively non-compositional, (iv) yet highly specialised to the context (ruh etal. (2005), chernova and arkin (2007)). suchrepetition effectsare familiar enough from everyday experience, and this list only hints at the complexity involved with routinisation. within cognitive psychology, routinisation surfaces chiefly in models of automaticity (e.g. logan (1988), anderson (1992), bargh (1992)), ranging from the more common property listing accounts, to the recently emerging mechanistic approaches (recently critiqued in rawson (2004)). within linguistics, routinisation surfaces chiefly in the vast literature on formulaic language (e.g. bolinger (1976), jackendoff (1997), erman and warren (2000), sag et al. (2002)), much of which employs property-listing of one sort or another as a key component of their modelling strategy. despite detailed accounts of routinisation phenomena within dialogue (e.g. kuiper (1996), aijmer (1996)), these also involve extensive listing of properties, with little formal and computational work in this area. pickering and garrod (2005) present an explicit, testable model of linguistic routinisation.11 however, their account lacks the formal details we require, despite their suggestions for adapting ideas from jackendoff (2003), the latter being a form of the property-list approach to routinisation, which we are trying to avoid. rawson (2004) presents a non-property-listing alternative, demonstrating the usefulness of mechanistic approaches to linguistic modelling (see also logan (1997)), comparing the rule-based account of anderson (1992) to memory-based models of automaticity, such asthat of logan (1988), in a series of reading experiments. while the results were complex, with memory-based processing clearly driving the bulk of speed-up effects associated with practice (sothat there were typically reduced reading times for the same texts, but elevated times for novel texts),she did find evidence for some involvement of rule-based processing, particularly in response to novel items. in our account, rather than devising our own models and possibly re-inventing several wheels, we look to such models from psychology to provide the basis for our account of linguistic routinisation. to capture earlier observations in section (2.2) regarding the relative incrementality of linguistic routines, as well as proposals for how such routines might emerge in section (2.3), we suggest that the dynamics of routinisation stems from the race between memory-based vs. rule-based processes. thus, the shift from the mid-term where memory-based processing holds sway, to the longer-term 10. schematically, compiling two productions triggered bycondition1 andcondition2, and which lead to effects update1 andupdate2, respectively, may lead to a rule with compound effects, but which does not requirecondition2 as a trigger, something like: if [condition1] then [update1 ∧ update2] 11. within dynamic syntax, bouzouita (2008) proposes a way of formallymodelling aspects of their account. 178 incrementality and routines in dialogue where rule-based processing, via production compilation, becomes more competitive, is externally verified by proposals from specific cognitive psychology accounts (chiefly anderson (1992), logan (1988), rawson (2004)). in this way, we formulate a novel account of the emergence of formulaic language, in terms of the dynamics of the linguistic system, by extending the framework of dynamic syntax (ds, e.g. kempson et al. (2001), cann et al. (2005)), especially as this has been applied to dialogue (e.g. purver et al. (2006)). we show how this yields not only a new explanation of the continuum from the context specialisation of pre-fabs to selectivity of discontinuous idioms, but further, we can employ the ds model of language processing as applied to dialogue phenomena, to potentially model all manner of formulaic dialogue phenomena. 179 gargett 3. formal modelling 3.1 formal details of the model informally, the dynamic syntax (ds, kempson et al. (2001), cann et al. (2005)) account of how contextual information can be incorporatedas it ariseswith linguistic information during dialogue, has three main characteristics: fully incremental processing, modelling update as cycles of the enrichment of underspecified representations, plus parity of parsing and production (formal details for each below). ds provides a fully incremental parsing model, with update modelled as transitions between succeeding parse states, essentially, enrichment of partial tree structures. parsing is then the sequence of pairings of natural language strings of termss with the logical formular representing the semantic structure of those terms: (11) {〈s(i), ri〉 , 〈s(i+ 1), ri+1〉 , . . .} thus,ri results from parsings(i). more generally, these successive parse states are modelled as triples (12) 〈pt, fs, fa〉 of (partial) tree structurespt , functionfs mapping partial tree structures to items of the formal language, and functionfa mapping actions (from sets of actions,a) for transition between trees to pairs of partial trees. structured logical formulae representing (predicate-argument) contentare mapped to decorated (binary) finite partial trees. thus, parsing the string ‘bob left’ results in the unreduced lambda term((λxleft(x)), bob), represented by the following decorated finite partial tree which includes the decorations for the topmost,ty(t) mother node (both formulafo and typing informationty included):12 [0fo(left′(bob′)), t y(t)[0fo(bob′), t y(e)][1fo(λxleft′(x)), t y(e > t)]] dominance relations between nodes specify tree structure, from more local argument-daughter(≺0) andfunction-daughter(≺1) relations between immediate neighbour nodes, to more global relations holding over collections of nested sub-trees (neighbours of neighbours of nodes). another crucial component of the framework is a so-called link mechanism for constructing pairs of trees, effectively conjoining the information contained in trees so linked (more below).13 transitions from onept to the next (in the sequence of updates) are effected through three main kinds of actions: lexical actions, a finite set of incremental actions associated with every word in the language, the occurrence of this word effectively triggering this instruction set, computational actions, a finite set of actions of a more general nature for building linguistic structure, which are triggered independently of the occurrence of individual words, and finally 12. following the bracketed format of meyer-viol (2001) signifying predicate-argument tree structure, where outermost [0...] is the top level, subsequent[0...] enclose argument daughters within the tree, and any[1...] enclose function daughters. 13. linking symbolised via an〈l〉 modality. 180 incrementality and routines in dialogue pragmatic actions, a finite set of actions which operate to connect contextual information with the currentpt under construction. each of these “macro”-level actions are actually composed of lower-level constructional actions for creating new nodes, decorating nodes, substitution at nodes, or else pointer movement.14 regarding this latter, thepointer device (symbolised by♦) is central to modelling transitions, singling out the specific node in the tree which is the current focus of update, pairing this node with the partial tree currently under construction (kempson et al. 2001, p. 275),15 this pairing then being an element of the set ofpointed partial trees (ppts). an important feature of the eventual dialogue account (detailed below) is that, while parsing may begin in the simplest cases with the so-called axiom,?ty(t), an initial state expressing the requirement to simply build a propositional object, modelling sequences of contributions by succeeding interlocutors in dialogue may potentially require starting from any point along the partial order of states (more below). indeed, parsing could then start from a context provided by the previous speaker’s utterance (as we will see). summarising the presentation so far, the basic units of the framework (kempson et al. 2001, p. 269) are decoratedppts, described using a language pairing elements from the setf of labels or features, likety (“type”), fo (“formula”) or tn (“tree node”), with elements from the setd of formulae (“decorations”), likee > t, λxleft′(x) or 01, the latter effectively values for these features (cf. attribute value matrices), but also including a setmv of metavariables (details of these latter below).16 the enrichment of underspecifiedpptsis crucial to the account of goal-directed information growth for any dimension of tree structures and decorations (formula andtype information). three kinds of underspecification are involved, structural, formula, and type value, the goal-directedness of enrichment/update modelled explicitly in terms ofrequirements, so that for any labelx, adding a requirement?x imposes a goal to establishx. all aspects of underspecification have an associated requirement for update, so that requirements may take the form?ty(t), ?ty(e), ?ty(e > t), ?〈↓〉ty(e > t), ?∃xfo(x), ?∃xtn(x), etc. pronouns illustrate formula underspecification, for example, the pronounhe licensing projection of a metavariablefo(umale′(u)) of ty(e) with requirement?∃xfo(x) (the latter a requirement for a fully specified formula). such metavariables are replaced by asubstitutionprocess from a term available in context. names too are modelled as projecting a metavariable, so that the occurrence of “bill” projects a metavariable annotated as fo(ubill′(u)), with instruction to construct a link transition to a linked tree of topnodety(t) decorated with the formula valuebill′(u), characterising the predicatebeing named bill, this constituting a constraint on the logical constant to be assigned as a construal of the use of that name in the particular context. all such metavariable-based terms are then enrichedvia the pragmatic action of substitution, and which may itself be suitably constrained (e.g. see discussion in purver et al. (2006)). 14. such basic actions can combine into complex ones. suppressing some formal details (for which see (meyer-viol 2001, p. 171), (kempson et al. 2001, p. 308)), basic actions includeones likemake, put andgo, which are very low level actions for constructing and decorating nodes, and moving between sub-trees under construction. macros of actions can be assembled from basic actions, chained together with the dynamic logic concatenation operator;. for example, the procedure for moving to a particular node to decorate it with some fact or requirement is specified as: make(r); go(r); put(φ); go(r−1). 15. thus,tn ≺i t ′n′ if t = t ′ andn ≺i n ′ (for somei belonging to the set of dominance relations presented above). 16. this is somewhat simplified, for details see kempson et al. (2001). 181 gargett 3.2 parity, dialogue and bi-directional grammars in ds, interactional dynamics during dialogue can be captured directly in terms of the core resources of the grammar, whereby the transition between speakers is modelled as the transition between parsing states. ds models generation as driven by the same underlying processing mechanisms as parsing, in common with many bi-directional grammars (e.g. shieber (1988)),thereby ensuring that generation is as incremental and goal-directed as parsing. in ds, a hearer builds a succession of partial parse trees representing components of the speaker’s utterance. speaking is modelled in the same way, with the addition of agoal treerepresenting what the speaker wishes to say.17 thus, a hearer can switch to speaking immediately and work from the same representation for both. further, parsing and generation can start from any point (axiom being only one possibility), so that interlocutors may in fact work off the immediately preceding context, or else from some store of structures, both strategies being crucial for modelling routinisation in dialogue. we can then model the switching back-and-forth between parsing and generation in a fully incremental way, this providing a mechanism for the emergence ofparity during dialogue.18 such tight coupling of goal-directed parsing and generation captures how interlocutors can make micro-adjustments in and through the interaction itself, thus directly modelling how the emergent dialogue isshaped on the fly through such fine-grained interaction between interlocutors. more formally, purver et al. (2006) define a parse statep as the triple〈t,w,a〉, with t a (possibly partial) tree,w the associated sequence of words, anda the associated sequence of lexical and computational actions. at any point in the parsing process, the context c for a particular partial treet in the setp can be taken to consist of the results of previous parses{. . . , 〈ti,wi, ai〉, . . .}. later we draw on this for our model of tripartite minimal exchanges, with the context consisting of a parse statep ′ resulting from some utterance initiating the exchange, any partial trees established in subsequent parsed fragments associated with some clarification or extension of aspects ofp ′, and finally partial trees established for the response of the initiating speaker. the generation model consists of incremental (word-by-word) parsing, and lexicon search for words which provide appropriate tree-update relative to agoal tree(what it is the speaker wishes to say), and through this process speakers produce the natural language string associated with the goal tree (purver et al. (2006)). a generator stateg is thus a pair〈tg, x〉 consisting of: (i) a goal treetg, and (ii) a setx of pairs〈s, p 〉, s a candidate partial string,p an associated parser state.19 the context-dependence of generation comes to the fore where lexical search can include context wherever possible, the effect being to reduce the production task. such reuse of context drives coordination between speakers via generation as well as parsing, the dynamics of this process arising indirectly out of the interaction. the following example demonstrates reuse for simple questionanswer: 17. being a so-called tactical generation model only, the question of how this is arrived at is put on hold for now. 18. within cognitive science, parity involves sharing representations across processes within the same individual, such as where representations for both speaking and hearing are built usingthe same underlying mechanism (pickering and garrod (2004)). bi-directionality is then an important ingredient withinthe present account of how parity across understanding and generation systems is achieved, and is central to ourmodelling of dialogue phenomena based on the ds grammar model. 19. as defined in section (3.1) above. 182 incrementality and routines in dialogue (13) a: who left? b: bill did. from a ds perspective (e.g. cann et al. (2007)), for their answer, breuses the context provided by interpreting a’s question (details below). note that for a to understand b’squestion, a needs to understand it against the context of their own initial utterance. moreover, the combined incremental, goal-directed framework enables processing linguistic stringssub-propositionally, making the ds dialogue model radically different to established approaches to dialogue modelling (e.g. traum (1994), asher and lascarides (2003)) which retain commitment to rather coarse-grained units of analysis, typically propositions.20 ds models the processing of fragments in contexts as steps toward the construction of a fully propositional term, so that grounding (the process whereby interlocutors arrive at mutualunderstanding of their utterances) is driven sub-propositionally. there is ample evidence that during dialogue people quite happily interact sub-propositionally, such as:21 (14) trains91, dialogue 1.2, lines 12.1 to 18.1 s: okay : so we’ll say m: send s: e2 : i guess : ... from elmira m: tshh : yeah s: and send them ... m: to corning s: to corning here, m and s switch turns, seemingly together constructing the eventual proposition. of course, each must understand the whole and where their own contribution fits, so each must separately entertain some proposition commensurate with that expressed by the final utterance (or if not, then they would know they were mistaken). for ds, such sub-propositional units are grammatical,22 and interlocutors may be not only working toward constructing a propositional term (i.e. the output of a complete tree), they may also be engaged in more partial interpretive work and processing material at sub-propositional levels below this. our modelling strategy for dealing with the complex of dialogue phenomena is to recast this in terms of a minimal exchange model, focusing directly on theinitiative which dialogue agents display in the following manner: 20. although others have recently claimed to also be modelling sub-propositionally, e.g. matheson et al. (2000), poesio and rieser (2010). 21. note that the format of the following dialogue from the trains91 dialogue corpus (http://www.cs.rochester.edu/research/speech/trains.html) is simplifiedfor expository purposes. . . . represents an extended pause of noticeable duration. 22. in the sense that grammaticality is relative to context, if when combined insuitable ways with some information from context the result is a grammatical unit (for explicit definition of grammatical, outside our scope here, see cann et al. (2007). 183 gargett (i) a context tree: thestart statethis is the tree the initiator is starting their parse from (ii) a goal tree: thefinal state the initiator will end up with a tree matching this, if grounding is successful (iii) a construction tree: an intermediate state bridging (i) and (ii), and which replaces the context tree after every update step this covers a range of core dialogue phenomena we are interested in modelling, such as clarifications, reformulations, and corrections, and we have demonstrated (e.g. gargett et al. (2009)) that our approach accounts for such core dialogue phenomena. however, it is also important to point out that we are not currently explicitly modelling the content of such exchanges, and we are not claiming to have (as yet) a complete dialogue model in this sense. let’s consider example (13) using a more schematic presentation of these three, inter-related kinds of trees. in table (1), the context for step 1 is the following alreadyformed tree structure resulting from b’s parse of a’s question, including both the subject nodedecorated withwh, and predicate node decorated withleft′: (15) ty(t), fo(left′(wh)),♦ fo(wh), t y(e) fo(leave′) ty(e > t) in our example, the subject node is updated with information licensed by occurrence of ‘bob’, this reflecting b’s analysis of a’s question (plus relevant wider knowledge that bob is the correct answer in this case). next, occurrence of ‘did’ licenses update in accordance with the following lexical actions (see purver et al. (2006) for details):23 (16) ‘do’ if ?ty(e > t) then put(fo(u)); put(ty(e > t)) put(?∃x.fo(x)) else abort having uttered ‘did’, the next update step 2 in table (1) requires specifying the metavariablefo(u) decorating the predicate node, by enabling reuse of the formula decorating thety(e > t) node of the tree in (15). the dynamics of the process stem from how construction trees iteratively become in turn context trees for subsequent construction trees, with a progression of construction trees recycled as context. these cycles of contribution-response-contribution enable narrowingof focus to a specific point in the representation under construction, providing interlocutors the opportunity for quite fine-grained adjustments of understanding.24 23. for convenience, we are ignoring tense information. 24. so that such continual switching of speakers does not necessarilyindicate misunderstanding and breakdown of communication. 184 incrementality and routines in dialogue contextb: constructionb: goalb: step 1 ?ty(t) bill′, t y(e) ty(e > t),♦ ?ty(t) bill′, t y(e) fo(u), t y(e > t), ?∃x.fo(x), ♦ step 2 ?ty(t) bill′, t y(e) fo(u), t y(e > t), ?∃x.fo(x), ♦ ?ty(t) bill′, t y(e) fo(u), t y(e > t), ?∃x.fo(x), ♦ ⇑ fo(left′) ty(t), left′(bill′) bill′, t y(e) fo(left′) ty(e > t) step 3 ?ty(t) bill′, t y(e) fo(u), t y(e > t), ?∃x.fo(x), ♦ ⇑ fo(left′) ty(t), left′(bill′),♦ bill′, t y(e) fo(left′), t y(e > t) table 1: tripartite model of minimal exchanges: simple question and answer. note that the predicate informationfo(left′) substituted into the construction tree in step 2, is provided by b’s own representation of a’s previous utterance of ‘who left?’ (see discussion of (15) for details). 185 gargett 4. from prefabs to idioms in language use this section attempts a formal demonstration of the approach to routinisation suggested in sections (2.3) and (2.4), whereby pre-fabs emerge within focused dialogueover the intermediate term, with increasing efficiency of lexical actions over the longer term. another source of contact with our hybrid modelling approach in the literature are the results from the psycholinguistic investigation of idioms by sprenger et al. (2006). their model is hybrid between idioms being unitary, at a conceptual level, and compositional, at a lexemic level. thus, the particular lemma ‘bucket’ will be activated for both the literal and idiomatic meanings of ‘kicked the bucket’, but the source of activation of the lemma is different in each case. however, the formal model in sprenger et al. (2006) focuses on linking distinct syntactic and semantic/conceptual levels, whereas we focus on the additional aspects of the dynamics of such links in contexts of interaction. in what follows, we provide details of our extension to the core ds framework, incorporating hybrid rule-based and memory-based modelling, and formally detailing a dualprocess model of the emergence of formulaic language. we finish by drawing out the other thread pursued in this paper, that of incremental context-dependence of formulaic language, demonstrating this in a dialogue setting. note that while the analyses in what follows involve constructed data,we are currently extending our analyses to actual dialogue data (such as that reported in healey (1997), discussed in section (2.2.2) above). 4.1 extending the core ds account our aim here is to extend the ds framework, by modelling lexical entries as nodes within a network of such entries, consisting of tuples〈w, t,a〉 of phonological informationw, semantic structuret and lexical actionsa. these nodes are accessed primarily through recognising/producing sequences of phonological stringswiwjwk . . ., so that nodes may themselves consist in part of string sets such as those for ‘kick’, ‘kick himself’, ‘kick herself’, ‘kick themselves’, etc (more details on the structure of lexical entries below). this leads to transitions between states being effected either via lexical actions, triggering building of tree structure by basic actions (asnoted in section (3.1) above), or else by directly contributing (previously stored) structure atthe appropriate place in the unfolding tree structure. in what follows, we first consider the formal modelling issues which formulaic language presents, then suggest how fixed idiomatic forms may emerge from relatively less fixed pre-fabs. finally, we will show how relevant lexical entries can be extended, before finishing this subsection with a schematic presentation of our formal proposal for extending ds lexical entries to lexical nodes. 4.1.1 basic lexical entries as mentioned, nodes may consist of sets of phonological strings, together with some associated lexical actions, such as (in all following examples we ignore tense for convenience):25 25. note the use of relational operators for various purposes, including locating the current node within the larger tree structure, such as〈↓1〉ty(e > t), which specifies that at the function daughter below the current node is one of ty(e > t), or even for pointing out a direction, as ingo(〈↑0〉), which is to say, “go up the argument (’0’) branch from here”. 186 incrementality and routines in dialogue (17) kick if ?ty(e > t) then make(〈↓1〉), go(〈↓1〉), put(fo(λxy.kick′(xy)), t y(e > (e > t))) go(〈↑1〉),make(〈↓0〉), go(〈↓0〉), put(?ty(e)) else abort (18) xself if ?ty(e > t) then put(fo(uanaph) ∧ ty(e)) else abort note that the operation of the reflexive also relies on a special local version of the pragmatic action substitution, which enriches the metavariable decorating the object node with the formula information decorating the subject node (following the analysis in cann et al. (2005)). now, applying the lexical entries in (17) and (18) capture non-idiomatic meaning only but consider the occurrence of idiomatic ‘bob kicked himself’(=“bob reproached himself”). recall from section (2.1) the following schematic sequence of transitions for this idiomatic expression, repeating here the previous analysis that parsing this expression involves stepping through one less state than the non-idiomatic expression: (19) s1 bob → s2 kicked himself −→ s3 upon completing a parse of this sequence we might arrive at the following final state (the formulation of names here is simplified, but see gargett (2010) for details): (20) ty(t), fo(reproach′(bob′)(bob′))♦ ty(e) bob′ ty(e > t) fo(λy.reproach′(bob′)(y)) ty(e) bob′ ty(e > (e > t)) fo(λxy.reproach′(x)(y)) we need to show how our model captures the entire sequence of updates leading to (20) by providing lexical actions for the idiomatic expression. rather than simply stipulating these directly, the advantage of the dual process account we are taking is that we are able to model the process underlying the emergence of these entries (specifically, the memory-basedprocessing of these in terms of their access and retrieval). 4.1.2 the emergence and storage of semantic structures we propose that the sequence of the verbkick plus reflexivehimselfbecomes routinised over time, with storage of the semantic structure outputted at the associated parse statess2 ands3, essentially the tree in example (20). over time, there will be an accumulation of many instances of such output, for example, both of the following: 187 gargett (21) {tn(0), fo(reproach′(mary′)(mary′)), t y(t)} {〈↑0〉tn(0), fo(mary′), t y(e)} {〈↑1〉tn(0), fo(λx.reproach′(mary′)(x)), t y(e > t)} {〈↑0↑1〉tn(0), fo(mary), t y(e)} {〈↑1↑1〉tn(0), fo(λxy.reproach′(y)(x)), t y(e > (e > t))} (22) {tn(0), fo(reproach′(bob′)(bob′)), t y(t)} {〈↑0〉tn(0), fo(bob′), t y(e)} {〈↑1〉tn(0), fo(λx.reproach′(bob′)(x)), t y(e > t)} {〈↑0↑1〉tn(0), fo(bob), t y(e)} {〈↑1↑1〉tn(0), fo(λxy.reproach′(y)(x)), t y(e > (e > t))} by virtue of being stored locally, these largely similar semantic structures may berelated via a process oftree abstraction. here we adapt a proposal by wilfried meyer-viol to formalise how such abstraction might proceed.26 recall that update via the transition function involves moving the unfolding tree structure along the partial order≤ from less to more specified states. the basic idea of abstraction involves movingbackwardsalong≤, effectively unwinding the complete tree to an earlier point at whichever nodes the information of the source trees differs, replacing any formulae at these nodes with metavariables and requirements for anfo. thus, the structures in examples (21) and (22) can be abstracted as follows: (23) {tn(0), ?ty(t)} {〈↑0〉tn(0), fo(u), ?∃yfo(y), ?ty(e)} {〈↑1〉tn(0), ?ty(e > t)} {〈↑0↑1〉tn(0), fo(u), ?∃xfo(x), ?ty(e)} {〈↑1↑1〉tn(0), fo(λxy.reproach′(y)(x)), t y(e > (e > t))} the resulting abstracted tree in (23) is not an expected output of parsingsome utterance in english: although the verb node is fixed and decorated with a fully specified formula,the subject node is fixed yet also decorated by an underspecified formula expression.27 such abstraction essentially pinpoints the similarities in these structures, and might be expected from structures being stored locally within some network of such structures, as a memory-based effect (discussed further below). the metavariables here represent kinds of abstractions over the placesthey are the holder for. these places were originally occupied by items which had some similarity with respect to each other, in most general terms (employing featural definition of categories) this involved [+animacy], and more specifically, it involved [+human]. note that metavariables at both subject and object nodes in (23) are identical, and this captures the identity of the formulae at these nodes resulting from occurrence of thereflexive (see (18) above). yet, as it stands, our analysis is incomplete, since we need to derive a structure which can be employed incrementally at the appropriate point in the parse. recall the schematic of this sequence in (19): after the occurrence of ‘bob’, the parse state is asfollows: 26. personal communication, ms. 27. a node can of course be unfixed and decorated with underspecified formula, as occurs in the case of left-dislocation, see kempson et al. (2001) for details. 188 incrementality and routines in dialogue (24) ?ty(t) ty(e) bob′ ?ty(e > t),♦ proceeding to the next update step via stored semantic structure requires retrieval of a sub-tree with topmost node of type?ty(e > t). however, the tree in example (23) contains a subject node, which is superfluous since we require only information stored for the combined predicate and object (i.e. in ds terms, the sub-tree with topmost node of typety(e > t)). at this point, we do not have a detailed theoretical account of how such superfluous information might bediscarded.28 for now, we simply prune structure above thety(e > t) node, with the result as follows: (25) {〈↑1〉tn(0), ?ty(e > t)} {〈↑0↑1〉tn(0), fo(fo(uanaph) ∧ ty(e))} {〈↑1↑1〉tn(0), fo(λxy.reproach′(y)(x)), t y(e > (e > t))} the final version in (25) models the structure that would provide update ofthe partial tree in (24). note that this additional step extends the original abstraction operation by decorating the object node with a reflexive anaphor (to which can be applied the special local version of substitution advocated by cann et al. (2005) for reflexives), this being triggeredduring the pruning process by identical metavariables occurring on subject and object nodes of (23). 4.1.3 extending lexical entries crucial to our proposed account of the dynamics of the emergence of formulaic language (as discussed in section (2.3) above) is the competition between lexical actions and semantic structures to update the tree currently under construction, in response to occurrences of the phonological string. thus, the occurrence of ‘kick himself’ sets in train a race between the processes underlying both lexical actions and semantic structure, to produce the material which updatesthe unfolding tree structure through the sequence of transitions represented above in (19). depending on the outcome of this race, it may be the abstracted semantic structure in (25) which updatesthe unfolding tree, or it may be the lexical actions triggered by ‘kick’ followed by those triggered by ‘himself’. we consider this account to be essentially a linguistic recasting of logan’s model of automaticity (see section (2.4) for details). we saw in section (4.1.2) how semantic structures might emerge and provide structure for updating the tree currently under construction. we propose that a semantic structure suitably optimised (such as after undergoing the pruning process described in section (4.1.2)), would win the race to provide update. indeed, the degree of specialisation of the semantic structure for the particular context, leads to it taking over update of the parse state. now, keeping with our linguistically inspired extension of logan’s model, the only way that the lexical action can become competitive again, and thus take over processing in response to the occurrence of, say, idiomatic ‘kick himself’ or ‘kick herself’, is if somehow there is a reduction in processing time. an obvious mechanism for this is the use of procedural compilation in various models of working memory (e.g. act-r, taatgen et al. (2008), discussedfurther below). 28. for example, some way of restricting extraction to that information from context which is “relevant” is likely needed, although we do not currently have a way of operationalising such a notion of relevance of information. 189 gargett we propose then that lexical rules may undergo their own form of optimisationin order to become once again competitive in this process, in particular through a process of procedural compilation, whereby lexical actions are chained together to provide complexesof such actions. the resulting complex lexical action for idiomatic ‘kick oneself’ is as follows: (26) kick oneself if ?ty(e > t) then make(〈↓1〉), go(〈↓1〉), put(fo(λxy.reproach′(x)(y)), t y(e > (e > t))), go(〈↑1〉),make(〈↓0〉), go(〈↓0〉), put(fo(uanaph) ∧ ty(e)) else abort note that the lexical entry in (26) for ‘kick oneself’ contains additional procedures for decorating the tree node with the010 address (the object node),29 these additional procedures being deployed by extending the formalism for lexical actions with theand structuring device for bundling together procedures dealing with both ‘kick’ as well as ‘oneself’ into a more complexlexical action for ‘kick oneself’. 4.1.4 interim summary in summary, the mechanism we are proposing for the emergence of linguistic routines is quite indirect, driven by thecompetitionbetween semantic structure and lexical actions.30 it should be emphasised that of course ds provides the possibility that either actions or structures can be used as possible updates for the unfolding tree structure, so of course we couldwell have represented (23) in terms of actions rather than structure. however, what we are seeking here is a way of using these formal mechanisms to model processes underlying the patterns we see in dialogue in cognitive terms. to this end, these structures are employed to suggest a memory-based account for how outputs of the parser may be stored and reused on subsequent occasions (perhaps even over the much longer term in the case of stable forms of formulaic language). further, we have shown that, despite idioms involving skipping over some state/s, rather than stepping through each and every possible individual state (recall discussion in section (2.1)), the process is still incremental, just that there are overall fewer actual transitions between states, this being the effect of more the complex lexical actions for idioms.31 4.1.5 architecture of lexical nodes now, while our model integrates rules and stored structures, both are essentially computational.32 the final component in our dual process account is to fully implement retrieval of structures in 29. reading the address from right to left, and thereby “back up” the tree: argument daughter (0) of function daughter (1) of the root node (0). 30. note how close we are to the account in kecskes (2008) of the dynamics of linguistic meaning arising from the interplay between current knowledge and growing experience. 31. in fact, as pointed out by an astute reviewer, as this entry stands, it may suggest that no concept of ‘kick’ is accessed when invoking the entry in (26). this is not intended by our account, although at this stage we are unable to directly address this issue. as noted immediately below, we intend revisiting these andother issues, in particular in light of a recent model proposed by sprenger et al. (2006). 32. via lexical procedures, on the one hand, and via tree abstraction, on the other. 190 incrementality and routines in dialogue a properly memory-based fashion, in order to simulate competition between update by either rulebased computation or via retrieval of structure. in gargett (2010), we propose modelling this competition in terms of retrieval of semantic structures conditioned through the manipulation of activation weights, and also in terms of utilities assigned to productions that govern how these operate (e.g. their speed).33 these aspects of our model are currently being implemented, and details of these are not included here. our dual processing account of lexical architecture models lexical nodes as consisting not only of strings and lexical actions for computing representations, but also stored semantic structure. figure (1) presents a schematic model of lexical nodes: each lexical node being bundles of phonological strings, lexical actions and semantic structures. these semantic structuresare essentially the structures outputted from previous uses of the associated lexical actions. accessing the information in these nodes is essentially via phonological strings (noted above), ranging from more compositional (like ‘kick’), to more formulaic (like ‘kick himself’, ‘kick herself’, ‘kick themselves’). by making both rules and structures available via some string set, our proposal aimsto reflect the hybrid compositional/formulaic nature of the lexicon (e.g. sprenger et al. (2006)).34 lexical node phon lexacts semstr figure 1: representation of lexical nodes (phon= phonological material,lexacts= lexical actions, semstr= stored semantic structures) in summary, we have shown with our series of examples employing tree abstraction, that this enables modelling the interaction between the processes that give rise to semantic structures as well as the process whereby these structures may be stored and subsequently retrieved. the result is a 33. such weights degrade over time, thereby modelling recency effects, so that both storing and retrieval of semantic structures, say, boosts their activation levels, and analogously, for theutility levels of productions when successfully firing. 34. indeed, a crucial element of sprenger et al.’s account is the relationship between the lexical and conceptual levels, which is beyond the scope of our proposal here. we are not here proposing a full-fledged lexical architecture, and our highly schematic model presented here does not detail the links between specific phonological forms and specific actions and/or semantic representations, required to model how uttering ‘kick himself’ increases the likelihood of use of the semantic representation for this specific string. indeed, in order to develop a more detailed account, we will aim for a model along the lines of the proposal in sprenger et al. (2006), but bringing this in line with our more procedural approach. 191 gargett unified story of formulaic and non-formulaic language,35 extending the ds lexicon to consist of a network of nodes, each consisting of strings (as the locus of lexical entries), with their associated lexical actions (core linguistic knowledge) and semantic structures (accumulating via “linguistic” experiences). 4.2 idioms in context finally, we demonstrate how our approach draws on aspects of the ds grammar model, to model idioms in dialogue. consider the following examples of the idiom ‘kick oneself’: (27) a: bob kicked himself b: and so did jill? by way of demonstrating how dialogue phenomena are modelled in ds via core grammatical resources, table (2) displays an analysis of example (27). first, on our extended model, given the idiomatic reading of a’s utterance, there are two possible routes to constructing a representation of a’s utterance: (i) use of the complex lexical actions for idiomatic reading of a’s utterance (as set out in (26)), (ii) retrieval of long-term stored structure for the idiomatic reading (displayed in (25)). in accordance with our dual process account, each of these are viable alternatives which compete against one another to provide update. second, b’s response in (27)cannot mean that jill kicked bob, so that if immediate context is reused here, this cannot consist of the output structure, complete with its value for the metavariable, since this would wrongly allow the meaning jill kicked bob. however, depending on which was initially used, complex actions or else stored semantic structure, this would be available for reuse in this case. this analysis reveals where we need to focus future work. for b’s response to a, a ds analysis of “do” posits a metavariable ofty(e > t), enabling substitution of predicate information from context, with an obvious candidate being the stored semantic structure retrieved for parsing a’s utterance (see (25)). another candidate may in fact be the structure immediately following application of the complex lexical action triggered by the idiom (see (26)). for the analysis in table (2), both alternatives, reuse of complex actions and retrieval of semantic structure, are theoretically possible given our dual process account (see section (4.1)). determining the strategy actually selected in this competition is an implementation issue; in future work we aim to implement the proposal in section (4.1.5), modelling the competition between lexical actions and semantic structure in these terms. 35. while we focus processing at the level of lexical nodes, others have suggested distributing this across multiple lexica (contra becker (1975), wray (2002), among others). 192 incrementality and routines in dialogue context tree: construction tree: update 1 “bob kicked himself” ?ty(t) ty(e) bob′ ?ty(e > t) ♦ ?ty(t) ty(e) bob′ ?ty(e > t) ty(e) ∧ uanaph,♦ ty(e > (e > t)) reproach′ update 2 “and so did jill” ?ty(t) ty(e) jill′ ?ty(e > t), fo(v),♦ ?ty(t) ty(e) jill′ ?ty(e > t) ⇑ ty(e) ∧ uanaph,♦ ty(e > (e > t)) reproach′ table 2: b’s context and construction trees for example (27). updates tothe context tree licensed by idiomatic reading of the string are symbolised by the blue dashed line (althoughthese dashed lines are for expository purposes, and have no formal significance).note that for update 2, the separate steps of update from context and then resolving reference via a local form of substitution to ensure identical formulae on subject and object nodes, are placed together on the same tree diagram for convenience, but they are in fact separate steps. the grayed section in update 2 highlights the subtree drawn from context (in fact thety(e > t) subtree from the construction tree in update 1). 193 gargett 5. conclusions: modelling context and routinisation in dialogue we have provided a novel dual process account of modelling the dynamics of formulaic language, as an alternative to the more common property listing approach. by setting our account firmly within the ds framework, we retain features of this framework useful for modelling. taking this approach, we can straightforwardly extend the framework to modelroutinedialogue phenomena. we extend the ds lexical architecture in order to model the emergence of formulaic language indirectly, within a model of language that focuses on replicating the processes underlying patterns of usage. our focus on the relationship between idioms and prefabs demonstrates the ds model of language as a system flexible enough to provide multiple strategies for a single form. an additional novel aspect of our approach to linguistic routinisation, is that we providea unified account of processing, focusing this at the level of lexical nodes, rather than distributing this across multiple lexica as has been proposed elsewhere. in sum, our contributions are fourfold: (1) a unified approach to formulaic and non-formulaic language, (2) a novel dual process account of formulaic language,(3) an extension of ds, in particular with respect to lexical architecture, and (4) a model of the dynamics of linguistic routinisation in dialogue. as an added bonus, our account turns out to be a linguistic implementation of the model for routinisation originally proposed by logan from within cognitive psychology, directly demonstrating how the complexity of dialogue can be tackled by integrating insights across disciplines within cognitive science. acknowledgments. i would like to acknowledge ruth kempson, eleni gregoromichelaki and matt purver for insightful discussions which helped in clarifying the ideaspresented here (matt also kindly commented on a very early draft). also many thanks to david schlangen and hannes rieser for their expert editorial guidance, as well as three anonymous reviewers who generously provided acutely critical comments which vastly improved the final version (normal disclaimers apply). 194 incrementality and routines in dialogue references k. aijmer. conversational routines in english: convention and creativity. longman, 1996. j.r. anderson. automaticity and the act* theory.american journal of psychology, 105:165–180, 1992. n. asher and a. lascarides.logics of conversation. cambridge university press, 2003. j.a. bargh. the ecology of automaticity: toward establishing the conditions needed to produce automatic processing effects.american journal of psychology, 105:181–199, 1992. d.j. barr and b. keysar. perspective-taking and the coordination of meaning in language use. in m.j. traxler and m.a. gernsbacher, editors,handbook of psycholinguistics: second edition, pages 901–938. amsterdam: elsevier, 2006. j.b. bavelas. what’s unique about dialogue? hand gestures, figurative language, facial displays, and direct quotation. inproceedings of sigdial 2009: the 10th annual meeting of the special interest group in discourse and dialogue, 2009. j.d. becker. the phrasal lexicon. inproceedings of the 1975 workshop on theoretical issues in natural language processing, 1975. d. bolinger. meaning and memory.forum linguisticum, 1:1–14, 1976. m. bouzouita. at the syntax-pragmatics interface: clitics in the history of spanish. in r. cooper and r. kempson, editors,language evolution and change. cambridge university press, 2008. j.l. bybee. from usage to grammar: the mind’s response to repetition.language, 82:711–733, 2006. r. cann, r. kempson, and l. marten.the dynamics of language: an introduction. elsevier, 2005. r. cann, r. kempson, and m. purver. context and wellformedness: thedynamics of ellipsis. research on language and computation, 5, 2007. s. chernova and r.c. arkin. from deliberative to routine behaviors: a cognitively-inspired action selection mechanism for routine behavior capture.adaptive behavior journal, 15(2):199–216, 2007. m. crocker. computational psycholinguistics. in a. clark, c. fox, and s. lappin, editors,the handbook of computational linguistics and natural language processing. wiley-blackwell, 2010. m. dalrymple, s.m. shieber, and f.c.n. pereira. ellipsis and higher-orderunification. linguistics and philosophy, 14(4):399–452, 1991. b. erman and b. warren. the idiom principle and the open-choice principle. text, 20:29–62, 2000. a. gargett.context and routinisation in dialogue. ph.d., kings college london, 2010. 195 gargett a. gargett, e. gregoromichelaki, r. kempson, m. purver, and y. sato. grammar resources for modelling dialogue dynamically.cognitive neurodynamics, 3(4):347–63, 2009. p. healey. expertise or expert-ese: the emergence of task-oriented sub-languages. in m.d. shafto and p. langley, editors,proceedings of the 19th annual conference of the cognitive science society, pages 301–306, stanford, california, august 1997. stanford university. p. healey. interactive misalignment: the role of repair in the development ofgroxford university press sub-languages. in r. cooper and r. kempson, editors,language in flux: dialogue coordination, language variation, change and evolution. college publications, 2008. r. jackendoff.the architecture of the language faculty. mit press, 1997. r. jackendoff. foundations of language: brain, meaning, grammar, and evolution. oxford university press, 2003. i. kecskes. dueling context: a dynamic model of meaning.journal of pragmatics, 2008. r. kempson, w. meyer-viol, and d. gabbay.dynamic syntax: the flow of language understanding. blackwell, 2001. k. kuiper. smooth talkers: the linguistic performance of auctioneers and sportscasters. lawrence erlbaum associates, 1996. g.d. logan. toward an instance theory of automatization.psychological review, 95:492–527, 1988. g.d. logan. automaticity and reading: perspectives from the instance theory of automatization. reading and writing quarterly, 13:123–146, 1997. d. marr. vision. w.h. freeman, 1982. c. matheson, m. poesio, and d. traum. modeling grounding and discourse obligations using update rules. inproceedings of the 1st meeting of the north american chapter of the association for computational linguistics (naacl), seattle, washington, april 2000. w. meyer-viol. sequential construction of logical forms.lnai 2014, pages 159–178, 2001. d. milward. dynamic dependency grammar.linguistics and philosophy, 17:561–605, 1994. g. nunberg, i.a. sag, and t. wasow. idioms.language, 70:491–538, 1994. m.j. pickering and s. garrod. toward a mechanistic psychology of dialogue. behavioral and brain sciences, 27:169–225, 2004. m.j. pickering and s. garrod. establishing and using routines during dialogue: implications for psychology and linguistics. in a. cutler, editor,twenty-first century psycholinguistics: four cornerstones, pages 85–101. lawrence erlbaum associates, 2005. m. poesio and h. rieser. completions, coordination, and alignment in dialogue. dialogue & discourse, 1(1), 2010. 196 incrementality and routines in dialogue m. purver, r. cann, and r. kempson. grammars as parsers: meeting thedialogue challenge. research on language and computation, 4(2-3):289–326, 2006. k.a. rawson. exploring automaticity in text processing: syntactic ambiguity as atest case.cognitive psychology, 49:333–369, 2004. k.a. rawson and e.l. middleton. memory-based processing as a mechanism of automaticity in text comprehension.journal of experimental psychology: learning, memory, and cognition, 35 (2):353–370, 2009. n. ruh, r.p. cooper, and d. mareschal. a reinforcement model of sequential routine action. in proceedings of international and interdisciplinary conference on adaptive knowledge representation and reasoning, pages 65–70, 2005. i.a. sag, t. baldwin, f. bond, a.a. copestake, and d. flickinger. multiword expressions: a pain in the neck for nlp. inproceedings of the third international conference on computational linguistics and intelligent text processing, pages 1–15, 2002. s. m. shieber. a uniform architecture for parsing and generation. inproceedings of the 12th international conference on computational linguistics, pages 614–619, 1988. s.a. sprenger, w.j.m. levelt, and g. kempen. lexical access during theproduction of idiomatic phrases.journal of memory and language, 54(2):161–184, 2006. p. sturt and v. lombardo. processing coordinated structures: incrementality and connectedness. cognitive science, 29(2):291–305, 2005. n.a. taatgen, d. huss, d. dickison, and j.r. anderson. the acquisitionof robust and flexible cognitive skills.journal of experimental psychology: general, 137(3):548–565, 2008. d. traum.a computational theory of grounding in natural language conversation. phd thesis, university of rochester, 1994. l. m. ward. dynamical cognitive science. mit press, 2002. a. wray. formulaic language and the lexicon. cambridge university press, 2002. 197 dialogue and discourse 3(2) (2012) 125–146 doi: 0.5087/dad.2012.206 creating conversational characters using question generation tools xuchen yao (now at johns hopkins university) xuchen@cs.jhu.edu emma tosch (now at the university of massachusetts amherst) etosch@gmail.com grace chen (now at california state university, long beach) grace.chen@student.csulb.edu elnaz nouri nouri@ict.usc.edu ron artstein artstein@ict.usc.edu anton leuski leuski@ict.usc.edu kenji sagae sagae@ict.usc.edu david traum traum@ict.usc.edu university of southern california institute for creative technologies 12015 waterfront drive, playa vista ca 90094-2536, usa editors: paul piwek and kristy elizabeth boyer abstract this article describes a new tool for extracting question-answer pairs from text articles, and reports three experiments that investigate how suitable this technique is for supplying knowledge to conversational characters. experiment 1 demonstrates the feasibility of our method by creating characters for 14 distinct topics and evaluating them using hand-authored questions. experiment 2 evaluates three of these characters using questions collected from naive participants, showing that the generated characters provide full or partial answers to about half of the questions asked. experiment 3 adds automatically extracted knowledge to an existing, hand-authored character, demonstrating that augmented characters can answer questions about new topics but with some degradation of the ability to answer questions about topics that the original character was trained to answer. overall, the results show that question generation is a promising method for creating or augmenting a question answering conversational character using an existing text. 1. introduction virtual question-answering characters (leuski et al., 2006a) are useful for many purposes, such as communication skills training (traum et al., 2007), informal education (swartout et al., 2010), and entertainment (hartholt et al., 2009). question answering ability can also form the basis for characters with capabilities for longer, sustained dialogue (roque and traum, 2007; artstein et al., 2009). in order to answer questions from users, a character needs to know the required answers; this has proved to be a major bottleneck in creating and deploying such characters because the requisite knowledge needs to be authored manually. there is research into ways of generating characters from corpora (gandhe and traum, 2008, 2010), but this requires large amounts of relevant conversational data, which are not readily available. on the other hand, plenty of information is available on many topics in text form, and text has been successfully transformed into dialogue that is acted out by conversational agents (piwek et al., 2007; hernault et al., 2008). we would like to be able to leverage textual resources in order to create a conversational character that responds to users – think c⃝2012 yao, tosch, chen, nouri, artstein, leuski, sagae, traum submitted 3/11; accepted 1/12; published online 3/12 yao, tosch, chen, nouri, artstein, leuski, sagae, traum of it as a dialogue system which can be given an article or other piece of text, and is then able to answer questions based on that text. there are essentially two ways to use textual resources to provide knowledge to a virtual character: the resources can be precompiled into a knowledge base that the character can use, or the character can consult the text directly on an as-needed basis, using generic software for answering questions based on textual material. automatic question answering has been studied extensively in recent years, retrieving answers from information databases (katz, 1988) as well as unstructured text collections, as in the question-answering track at the text retrieval conference (trec) (voorhees, 2004); interactive question answering uses a dialogue system as a front-end for adapting and refining user questions (quarteroni and manandhar, 2007). we are aware of one effort of integrating automatic question answering with a conversational character (mehta and corradini, 2008): when the character encounters a question for which it does not know an answer, it uses a question-answering system to query the web and find an appropriate answer. our approach follows the alternative route, of taking a textual resource and compiling it into a format that the character can use. this approach is extremely practical, as the virtual characters are created using an existing, publicly available toolkit – the ict virtual human toolkit1 – which uses a knowledge base in the form of linked question-answer pairs (leuski and traum, 2010). instead of authoring the knowledge base by hand, we populate it with question-answer pairs derived from a text through the use of question generation tools – question transducer (heilman and smith, 2009) and our own reimplementation, called openaryhpe.2 such question generation algorithms were initially developed for the purpose of reading comprehension assessment and practice, but turn out to be general enough that we can use them to create a character knowledge base. the extracted knowledge base can be used as a stand-alone conversational character or added to an existing one. the approach has the additional advantage that all the knowledge of the resulting character resides in a single, unified database, which can be further tweaked manually for improved performance. this article describes a set of experiments which demonstrate the viability of this approach. we start with a sample of encyclopedia texts on a variety of topics, which we transform into sets of question-answer pairs using the two question generation tools mentioned above. these questionanswer pairs form knowledge bases for a series of virtual question answering characters. experiment 1 tests 14 such characters using a small, hand-crafted set of test questions, demonstrating that the method is workable. experiment 2 uses a large test set of questions and answers collected from a pool of naive participants in order to provide a more thorough evaluation of three of the characters from the previous experiment. overall, the characters work reasonably well when asked questions for which the source text has an answer, giving a full or partial answer more than half the time. experiment 3 adds some of the automatically extracted knowledge bases to an existing, hand-authored character, and tests it on questions about topics it was originally designed to answer as well as from the newly added topics. performance on the new topics is similar to that of the corresponding stand-alone characters, though there is some degradation of the character’s ability to respond to questions about its original topics. the remainder of the article gives an overview of question generation and the tools we use, describes in detail the setup and results for the three experiments, and concludes with broader implications and directions for further study. 1. http://vhtoolkit.ict.usc.edu 2. http://code.google.com/p/openaryhpe 126 conversational characters 2. question generation and openaryhpe 2.1 background question generation is the task of generating reasonable questions and corresponding answers from a text. the source text can be a single sentence, a paragraph, an article and even a fact table or knowledge base. current approaches use the source text to provide answers to the generated questions. for instance, given the sentence jack went to a chinese restaurant yesterday, a question generator might produce where did jack go yesterday? but it would not ask did he give tips? this behavior of asking for known information makes question generation applicable in areas such as intelligent tutoring systems (to help learners check their accomplishment) and closed-domain question answering systems (to assemble question-answer pairs automatically). a taxonomy of questions used in the question answering track of the text retrieval conference divides questions into three types: factoid, list, and definition (voorhees, 2004). factoid questions have a single correct answer, for example how tall is the eiffel tower? list questions ask for a set of answer terms, for example what are army rankings from top to bottom? finally, definition questions are more open-ended, for example what is mold? in later editions of the question answering track, definition questions were replaced by “other” questions, specifically asking for other information on a topic beyond what was asked in previous questions. in terms of target complexity, question generation can be divided into shallow and deep (graesser et al., 2009). fact-based question generation produces typical shallow questions such as who, when, where, what, how many, how much and yes/no questions. automatic question generators usually achieve this by employing various named entity recognizers, or retrieving corresponding entity annotations from ontologies and semantic web resources; the generators then formulate questions which ask for these named entities. the types of questions that are generated depend on the named entities that the generator is able to recognize in the input sentence. a special type of shallow question involves an interrogative pronoun restricted by a head noun, for example what stone, which instrument or how many rivers. to generate such questions, generators typically use information about hierarchical semantic relations, gleaned from lexical resources such as wordnet. for example, given a sentence kim plays violin, a question generator may use the information that instrument is a hypernym of violin in order to come up with a question such as what instrument does kim play? this type of question is particularly prone to overgeneration if a lexical item is ambiguous and the generator cannot determine the appropriate sense. for example, given the source sentence a sword is a hand-held weapon made of metal, and knowing that metal is a genre of music, one of our generators formulated the question what music type is a sword a hand-held weapon made? (see table 1 below). deep questions, such as why, how, why not and what if questions, are considered a more difficult task than shallow questions because they typically involve inference that goes beyond reordering and substitution of words in the surface text. such inference is beyond the capabilities of most current question generators. one exception is mostow and chen (2009), which generates what, why and how questions about mental states following clues of a set of pre-defined modal verbs. 2.2 question transducer current approaches to question generation may be based on templates (mostow and chen, 2009), syntax (wyse and piwek, 2009; heilman and smith, 2009), or semantics (schwartz et al., 2004). 127 yao, tosch, chen, nouri, artstein, leuski, sagae, traum our experiments are based on a tool called question transducer (heilman and smith, 2009). the core idea of question transducer is to transduce a syntactic tree of a declarative sentence into that of an interrogative following a set of hand-crafted rules. the algorithm uses tregex (a tree query language) and tsurgeon (tree manipulation language) to manipulate sentences into questions in three stages: selecting and forming declarative sentences using paraphrasing rules and word substitutions, syntactically transforming the declarative sentences into questions, and scoring and ranking the resulting questions. question transducer employs a set of named entity recognizers (ners) to identify the target term and determine the type of questions to be asked. thus correctness of the question types relies on the ners, and the grammaticality of the generated sentences depends on the transformation rules. the idea of syntactic transformation from declarative sentences into questions is shared by some work from the question answering community. given a question, question answering attempts to find the most probable answer, or the sentence that contains the answer. thus, a measure of how a question is related to a sentence containing the answer is defined over syntactic transformation from the sentence to the question. this transformation is done, for instance, in the noisy channel approach of echihabi and marcu (2003) via a distortion model (specifically, ibm model 4 in brown et al. 1993) and in wang et al. (2007) via the quasi-synchronous grammar formalism of smith and eisner (2006). 2.3 openaryhpe in addition to using the question transducer tool, we also developed a partial reimplementation in the framework of openephyra (schlaefer et al., 2006) which we call openaryhpe. the difference between the two tools is that while question transducer only uses the stanford named entity recognizer (finkel et al., 2005), openaryhpe also includes ners based on ontologies (to expand to new domains) and regular expressions (to recognize time, distance, measurement more precisely). thus, openaryhpe is able to produce more questions by recognizing more terms. also, openaryhpe is able to ask more specific questions by utilizing the hypernyms provided in the ontology list. for instance, openaryhpe includes 90 lists of common newspapers, films, animals, authors, etc. openaryhpe is able to ask what newspaper or what film questions, providing more hints to the user about the particular type of answer the system is seeking. however, these hypernyms are not disambiguated, which leads to overgeneration as discussed above. moreover, openaryhpe does not implement the question ranking module that question transducer uses to output scored questions (heilman and smith, 2010). this is not a limitation for the purpose of the experiments reported in this article, because the experiments do not use the ranking of question-answer pairs. openaryhpe and question transducer work on individual sentences of the original text, so they only generate question-answer pairs that are contained in a single sentence. table 1 gives some examples of question-answer pairs extracted from a single source sentence (the list is not exhaustive – many more pairs were extracted from the same sentence). this sample shows that the two question generation tools are not equivalent, as each tool gives different pairs (though there is some overlap). looking at all the questions and answers, we see that there is a many-to-many mapping – a single extracted question is paired with multiple answers, and a single answer is paired with multiple questions. the question generation tools also add knowledge that is not available in the source – question transducer has figured out that a sword is a kind of weapon, and openaryhpe knows that metal is a genre of music (though it is not aware that this is not the relevant meaning 128 conversational characters source: a sword is a hand-held weapon made of metal. oaa what is a hand-held weapon made of metal? — a sword. oa qtb what is a sword? — a sword is a hand-held weapon made of metal. oa qt what is a sword? — a hand-held weapon made of metal. oa what music type is a sword a hand-held weapon made? — a hand-held weapon made of metal. oa what music type is a sword a hand-held weapon made? — of metal. qt what kind of weapon is a sword? — a sword is a hand-held weapon made of metal. qt what kind of weapon is a sword? — a hand-held weapon made of metal. aopenaryhpe bquestion transducer table 1: sample question-answer pairs extracted from a single source in this case). the music question also shows us that the questions are not always grammatically well-formed; the same holds for the answers, though it is not apparent from this sample. we did not perform a direct evaluation on the quality of the questions generated by the two systems, as in our experiments these questions are actually hidden from the humans talking to the conversational characters. however, we do evaluate how these generated questions affect the performance of the generated conversational characters (see experiment 1 and table 4). 3. experiment 1: hand-authored questions 3.1 method our first experiment was intended to investigate whether a character knowledge base in questionanswer format, created automatically from a source text using a question generator, could provide good answers to questions about the source text. the experiment involved the following steps. 1. select texts to serve as the raw source for the character knowledge base, and create a set of questions and answers based on the texts to serve as a “gold standard” test set. 2. create a character knowledge base in question-answer format, using question generation tools to extract question-answer pairs from the source text. 3. present the test questions to the character, and evaluate the quality of resulting responses using the answers in the test set as a reference. the above three steps are presented in the following sections. 3.1.1 materials as raw material for generating the character knowledge bases we selected 14 text excerpts from simple english wikipedia3 (table 2). these texts were chosen to represent a variety of topics. we chose the simple english version because it contains a smaller vocabulary and simpler sentence structure than the regular english version, and therefore lends itself better for processing by the 3. http://simple.wikipedia.org 129 yao, tosch, chen, nouri, artstein, leuski, sagae, traum question-answer pairs source article length (words) test extracted oaa qtb tot.c albert einstein 385 12 162 295 405 australia 567 9 240 308 500 beer 299 10 146 231 315 chicago blackhawks 338 8 134 164 291 excalibur 342 10 146 256 379 greenhouse gas 378 7 164 206 345 ludvig van beethoven 765 14 338 635 889 movie 543 12 162 246 393 river 368 13 108 262 339 roman empire 466 12 270 310 550 rugby football 478 14 238 260 473 scientific theory 408 9 134 234 333 sword 363 12 168 268 407 united states 426 10 194 270 436 aopenaryhpe bquestion transducer cthe total is less than the sum of oa and qt due to overlap. table 2: wikipedia text excerpts and question-answer pairs question generation tools (limitations of current question generation technology, and specifically the tools we used, mean that source texts for creating virtual characters need to be chosen with care; such restrictions are likely to be relaxed as general question generation technology matures and improves). articles were retrieved on june 10, 2010, and text excerpts were manually copied and pasted from a web browser into a text editor. the lengths of the individual texts ranged from 299 to 765 words (mean 438, standard deviation 122). for each text the third author constructed a set of questions with answers that can be found in the text, to serve as a test set for the generated characters. the number of question-answer pairs per text ranged from 7 to 14, for a total of 152 (mean 10.9, standard deviation 2.2). the test set concentrated on the types of questions that the characters were expected to handle, with an overwhelming majority of what questions (table 3) (our what category includes also what and which restricted by a head noun or preposition phrase, for example what instrument or which of the generals). 3.1.2 question generation we created question-answer pairs from the texts using the two question generation tools described in section 2: question transducer (heilman and smith, 2009) and openaryhpe. table 2 shows the number of question-answer pairs extracted from each text by each tool. the number of extracted pairs is large because both tools are biased towards overgeneration, creating multiple questions and answers for each identified keyword. question transducer overgenerates because the generation step is followed by ranking the question-answer pairs to find the best ones. we did not use the rank130 conversational characters question type total n % a lb er t e in st ei n a us tr al ia b ee r c hi ca go b la ck ha w ks e xc al ib ur g re en ho us e ga s l ud vi g va n b ee th ov en m ov ie r iv er r om an e m pi re r ug by fo ot ba ll sc ie nt ifi c th eo ry sw or d u ni te d st at es what 95 62 10 6 8 1 3 3 6 11 13 5 9 6 11 3 who 23 15 . 1 . 4 4 . 5 1 . 6 . 1 . 1 how much 8 5 . . . 1 . 1 . . . . 1 . . 5 how 7 5 . . 1 . . 2 . . . 1 1 1 . 1 when 6 4 . . 1 1 . . 2 . . . 1 . 1 . yes/no 5 3 2 . . . . 1 1 . . . . 1 . . where 5 3 . 1 . 1 2 . . . . . 1 . . . why 3 2 . 1 . . 1 . . . . . 1 . . . table 3: test set question types ing in our experiments (and openaryhpe did not even implement the ranking model), but the results in section 3.2 suggest that overgeneration is also useful for our method of creating conversational characters, because overgeneration provides many options from which the engine that drives the characters is able to identify the appropriate answers. the generated question-answer pairs were imported as a character knowledge base into npceditor (leuski and traum, 2010), a text classification system that drives virtual characters and is available for download as part of the ict virtual human toolkit (see footnote 1). npceditor is trained on a knowledge base of linked question-answer pairs, and is able to answer novel questions by selecting the most appropriate response from the available answers in the knowledge base. for each new input question, npceditor computes a language model for the ideal answer using the linked training data; it then compares the language model of the ideal answer to those of all of the answers in the knowledge base, and selects the closest available answer based on a similarity metric between language models. the use of language models allows npceditor to overcome some variation in the phrasing of questions, and retrieve appropriate responses for questions it has not seen in the training data. the training questions are only an avenue for selecting the answer and are never seen by the user interacting with the character; it is therefore not crucial that they be grammatically well-formed, only that they provide useful information for selecting an appropriate answer. for each source text we used npceditor to train 3 characters: one was trained on the questionanswer pairs extracted by openaryhpe, another on those extracted by question transducer, and a third character was trained on the pooled set of question-answer pairs. 3.1.3 evaluation we evaluated the character knowledge bases by presenting each of the test questions to the appropriate character in npceditor and rating the resulting response against the predetermined correct answer. this was done separately for the three sets of training data – the question-answer pairs 131 yao, tosch, chen, nouri, artstein, leuski, sagae, traum q-a generator mean n distribution 0 0.5 1 1.5 2 openaryhpe 1.27 152 48 3 9 2 90 question transducer 1.33 152 44 3 9 1 95 combined 1.49 152 30 2 14 0 106 table 4: rating of character answers (scale 0–2, two annotators) extracted by openaryhpe, those extracted by question transducer, and the pooled set. two raters (the second and third authors) rated all of the responses on the following three-point scale. 0 incorrect answer. 1 partly correct answer. 2 fully correct answer. agreement between the annotators was very high: α = 0.985,4 indicating that the rating is fairly straightforward: of the 456 responses rated, the annotators disagreed only on 11, and in all of those instances the magnitude of the disagreement was just 1. we therefore proceeded with the analysis using the mean of the two raters as a single score. 3.2 results table 4 shows the distribution of ratings for the character answers, broken down by the source of the question-answer pairs. the success rate is comparable for openaryhpe and question transducer, and somewhat better when the character is trained on the combined output of the two tools. the difference, however, does not appear to be significant: for the distribution data in table 4, χ2(8) = 9.6, p = 0.30. we do find a marginally significant effect of question generation tool when we run an anova modeling the individual mean ratings as an effect of tool and source text (a 3× 14 design): f(2,414) = 2.68, p = 0.07. to the extent that the difference is meaningful, we conjecture that the pooled knowledge bases provide the best responses because they contain the largest number of question-answer pairs in the training data, and thus offer the most choices for selecting an appropriate answer. the analysis does not show a significant interaction between question-generation tool and source text (f(26,414) = 0.59, p > 0.9), but a highly significant main effect of source text (f(13,414) = 3.25, p < 0.001), indicating that the source text and the questions asked during testing have a profound effect on the success of the character. the mean answer rating for characters trained on the pooled output from the two question generation tools ranged from 1.00 for the australia topic to 2.00 for the chicago blackhawks topic (table 5). 4. krippendorff’s α (krippendorff, 1980) is a chance-corrected agreement coefficient, similar to the more familiar k statistic (siegel and castellan, 1988). like k, α ranges from −1 to 1, where 1 signifies perfect agreement, 0 obtains when agreement is at chance level, and negative values show systematic disagreement. we chose to use α because it allows a variety of distance metrics between the judgments; here we used the interval metric as an approximation of the notion that partly correct answers fall somewhere between incorrect and fully correct answers. 132 conversational characters source text mean n distribution 0 0.5 1 1.5 2 albert einstein 1.33 12 3 . 2 . 7 australia 1.00 9 4 . 1 . 4 beer 1.50 10 2 . 1 . 7 chicago blackhawks 2.00 8 . . . . 8 excalibur 1.40 10 2 . 2 . 6 greenhouse gas 1.71 7 . . 2 . 5 ludvig van beethoven 1.43 14 4 . . . 10 movie 1.25 12 4 . 1 . 7 river 1.77 13 1 . 1 . 11 roman empire 1.67 12 2 . . . 10 rugby football 1.32 14 4 1 . . 9 scientific theory 1.50 9 1 1 1 . 6 sword 1.75 12 1 . 1 . 10 united states 1.40 10 2 . 2 . 6 table 5: rating of answers by characters trained on the pooled output from the question generation tools since the success of a character depends both on the training data and the questions asked, we designed the next experiment to look at a wider array of questions, which would better represent what a conversational character might encounter in a live interaction with people. 4. experiment 2: questions collected from users 4.1 method our second experiment was intended to investigate whether a character knowledge base, created automatically as in the previous experiment, could provide good answers to typical questions asked by users who are not familiar with the system. the experiment involved the following steps.5 1. select texts and create a character knowledge base as in the previous experiment. 2. collect questions and answers from naive participants to serve as a test set. 3. present the test questions to the character, and evaluate the quality of resulting responses using the answers in the test set as a reference. 4.1.1 materials since this experiment concentrated on evaluating character responses to collected user questions, we used a small selection of the source texts from the previous experiment. we expected performance to drop due to the more varied questions, so we chose three of the five top-performing source texts: 5. this experiment was reported in abbreviated form in chen et al. (2011). 133 yao, tosch, chen, nouri, artstein, leuski, sagae, traum sword, river and roman empire (see table 2). we only trained the characters based on the pooled output of the two question generation tools, since previously that resulted in the best performance. 4.1.2 test set since our method is intended to create a virtual character that can answer questions by human users, we need to test our character knowledge base against typical questions that a person may ask; at the same time, the test set should take into account the fact that the character can only respond with information that is present in the source texts. we therefore collected test questions from participants both before reading the source text and after having read it. the procedure for collecting the test data was as follows. 1. the participant wrote five questions about a topic (swords, rivers, or the roman empire), without having read any text materials about it. 2. the participant read the source text about the topic. 3. the participant wrote five additional questions about the topic, based on the source text. 4. the participant provided answers to all of their questions, where each answer was a contiguous segment of text from the source. if the participant felt that the text did not contain an answer, they marked the answer as “n/a”. the data were collected using the qualtrics on-line survey tool.6 the procedure was repeated three times, once for each topic, so in total each participant provided 30 questions and answers – 5 questions for each topic before reading the text and 5 after reading. we had 22 participants for a total of 660 collected questions and corresponding answers (220 for each topic). topics were presented to all the participants in the same order – first swords, then rivers, and finally the roman empire. 4.1.3 evaluation we evaluated the character knowledge bases by presenting each of the test questions to npceditor and rating the resulting answer against the user-provided answer. two raters (the second and third authors) rated all of the responses independently on four dimensions – two binary dimensions, and two dimensions on a three-point scale. question quality 1 comprehensible, 0 incomprehensible. system response 2 full answer, 1 partial answer, 0 not an answer. is the system response good enough 1 yes, 0 no. is the system response the best in the text 1 yes, 0 no, −1 there is no good response in the text. the first dimension, question quality, was rated merely to make sure that the questions created by the participants were understandable; the vast majority of questions (634 of 660) were rated as comprehensible by both coders, so we did not analyze this further. the main dimension for evaluation was the second, answer quality, and it forms the basis for the analysis in the next section; a few examples of this annotation are shown in table 6. the third and fourth dimensions were 6. http://qualtrics.com 134 conversational characters question user answer system answer ratings what is a sword made of? metal a sword is a hand-held weapon made of metal 2 2 what are swords made out of? a sword is a hand-held weapon made of metal larger swords such as longswords or claymores are used with two hands 0 0 what is a katana sword? a katana has one sharp edge and a small guard near the handle. a sword is a hand-held weapon made of metal 0 0 what is a bayonet? n/a a sword is a hand-held weapon made of metal 0 0 how fast is a river? it sometimes flows very slowly a wide slow river is called an “old river” 1 0 table 6: sample rating of system answers by 2 coders intended to give insight into nuances that are not captured by a single numerical rating for answer quality: whether the answer provided by the system was good enough as a response for a virtual character engaged in dialogue, and whether it was the best possible answer that could be found in the text. the latter question also allowed the raters to judge whether an answer to a question was available, independently of the judgment of the participant who provided the question. agreement between the annotators was high for all the rating dimensions: α = 0.865 for answer quality, α = 0.804 for whether an answer was good enough, and α = 0.784 for whether an answer was the best available.7 while the three questions were intended to capture distinct aspects of an answer’s suitability, in practice there was not much difference in the annotators’ responses, and the ratings display very high correlations (r = 0.92 for one coder and r = 0.95 for the other coder between answer quality and good enough, and r = 0.81 and r = 0.82 between answer quality and best available). we therefore proceed with the analysis using only the results from the answer quality annotation. 4.2 results 4.2.1 question distribution the questions collected from experiment participants show us what human users think they may want to ask a virtual character (since these questions were asked off-line rather than in conversation, they are less indicative as to what users actually do ask). the participants produced questions under two conditions – before reading the source text and after having read it – and provided their own judgments as to whether an answer was available in the text. table 7 shows the differences between the two conditions, broken down by question type. the most common question type was what, 7. for the answer quality question we used α with the interval metric as explained in footnote 4, whereas for the best available question the judgments are categorical so we used the nominal metric; for binary distinctions like the “good enough” question, the two metrics are equivalent. 135 yao, tosch, chen, nouri, artstein, leuski, sagae, traum before reading after reading question type total avail n/a avail n/a n % n % n % n % n % what 363 55 90 51 66 43 196 63 11 61 yes/no 59 9 10 6 16 10 28 9 5 28 who 50 8 17 10 10 7 23 7 0 0 when 46 7 24 14 5 3 17 5 0 0 where 46 7 12 7 15 10 17 5 2 11 how much 45 7 13 7 24 16 8 3 0 0 how 39 6 6 3 13 8 20 6 0 0 why 9 1 5 3 3 2 1 0 0 0 other 3 0 0 0 1 1 2 1 0 0 total 660 100 177 100 153 100 312 100 18 100 table 7: test question types question type total river roman sword n % n % n % n % what 363 55 143 65 99 45 121 55 yes/no 59 9 19 9 13 6 27 12 who 50 8 0 0 41 19 9 4 when 46 7 0 0 30 14 16 7 where 46 7 26 12 6 3 14 6 how much 45 7 15 7 15 7 15 7 how 39 6 15 7 10 5 14 6 why 9 1 1 0 5 2 3 1 other 3 0 1 0 1 0 1 0 total 660 100 220 100 220 100 220 100 table 8: test questions by topic constituting 55% of all questions (including what and which restricted by a noun or preposition phrase). the distribution is different for questions produced before and after reading the source text (χ2(8) = 37, p < 0.001): participants produced more what and yes/no questions after reading the source text – they had asked more varied question types before. of the questions asked before reading the text, 46% had no answers available in the text; of those asked after, only 5% were without answers. question types also differed by topic as shown in table 8, and again the difference was significant (χ2(16) = 115, p < 0.001). the topic of rivers received no who or when questions, which together constituted almost a third of the questions about the roman empire. 136 conversational characters distribution authoring time answer mean n 0 0.5 1 1.5 2 before reading available 0.53 177 103 19 21 8 26 n/a 0.12 153 131 12 7 1 2 after reading available 0.99 312 127 26 21 4 134 n/a 0.36 18 13 2 0 1 2 total 0.65 660 374 59 49 14 164 table 9: quality rating of character responses (0–2 scale) 4.2.2 answer quality the quality of the answers provided by the character to the user questions varied depending on the question authoring time (before or after reading the text) and the availability of an answer in the source text. we used the mean of the scores given by the two raters, so each answer received a score between 0 and 2 in half-point increments. table 9 shows the mean rating in each group as well as the distribution of scores; recall that the availability of an answer was judged by the participants whereas the quality of an answer was judged by the raters, which explains why a small number of answers are considered good even though the original participant thought there was no good answer in the text. the vast majority of good answers come from questions that were authored after the participant had read the source text, whereas for questions written before having read the text, npceditor has a much harder time finding an appropriate answer even when one is available. a likely explanation for this is that questions written after having read the text are more likely to use vocabulary and constructions found in the text – that is, the texts cause some form of lexical priming (levelt and kelter, 1982) or syntactic priming (bock, 1986). looking only at user questions with an available answer, questions asked after having read the text have an out-of-vocabulary word token rate of 42%, compared to 52% for questions asked before having read the text. npceditor is built on cross-language information retrieval techniques (leuski and traum, 2010) and thus it does not require the questions and answers to share a vocabulary, but it does require that the test questions be reasonably similar to the training questions. since the questions in the training data are derived from the source text, a better alignment of user questions with the source text should make it easier to map the questions to appropriate answers. finally, we note that answers provided by the characters to questions that do not have an answer in the source text are typically very poor (note however that npceditor is able to tell when its confidence in an answer is low, see section 5.2 below). the quality of the answer is also affected by the question type. table 10 gives the mean ratings for each question type, broken down by authoring time, both for all questions of the type as well as just those questions that have an answer available in the text. we see substantial differences between the question types – who questions do particularly well, whereas yes/no questions do rather poorly. this may be a byproduct of the question generation tools, which are able to identify some types of information better than others. for instance, the named entity recognizers have been trained on annotated data to recognize people names, and thus who questions are usually linked to a correct target answer in the training data. in contrast, why questions are generated from lexical matching of causal clue words, such as the reason or due to; this matching is not discriminate enough in 137 yao, tosch, chen, nouri, artstein, leuski, sagae, traum question type all questions before reading after reading all avail all avail what 0.70 0.31 0.46 1.00 1.02 yes/no 0.27 0.15 0.25 0.36 0.41 who 1.15 0.67 0.94 1.72 1.72 when 0.79 0.62 0.67 1.09 1.09 where 0.54 0.22 0.50 1.00 1.12 how much 0.53 0.39 0.65 1.19 1.19 how 0.22 0.16 0.50 0.28 0.28 why 0.11 0.12 0.20 0.00 0.00 other 1.33 0.00 — 2.00 2.00 table 10: mean ratings by question type (0–2 scale) finding out the reason and result of events, which may be a reason that generated why questions are generally of low quality. interestingly enough, we did not find an effect of source text on the quality of the responses. we conducted a 4-way anova looking at source text, availability of a response, question authoring time, and question type. the three factors discussed above came out as highly significant main effects: response availability (f(1,592) = 106, p < 0.001), authoring time (f(1,592) = 53, p < 0.001) and question type (f(8,592) = 6.5, p < 0.001); there was also a significant interaction between source text and question type (f(14,592) = 3.6, p < 0.001). but the main effect of source text was not significant (f(2,592) = 2.8, p = 0.06), and there was only one additional marginally significant interaction, between answer availability and question type (f(7,592) = 2.1, p = 0.04). 5. experiment 3: character augmentation 5.1 method the first two experiments showed that question generation can be used for creating question answering virtual characters from a text. our third experiment set out to investigate how the addition of an automatically generated question-answer knowledge base affects the performance of an existing, hand-authored character.8 ideally, such an augmented character would be able to answer the same questions as the original character without a substantial performance loss, and also be able to answer some questions covered by the added knowledge base. we chose to experiment with an existing character for which we already had an extensive test set of questions with known correct responses; to this character we added successive knowledge bases generated by the question-answering tools. 5.1.1 materials the base character for the experiment was the twins ada and grace, a pair of virtual characters situated in the museum of science in boston where they serve as virtual guides (swartout et al., 2010). the twins answer questions from visitors about exhibits in the museum and about science 8. this experiment was reported in nouri et al. (2011). 138 conversational characters source text questions answers q-a pairs twins 406 148 483 twins + australia 652 342 999 twins + beer 559 268 807 twins + beethoven 849 421 1412 twins + australia + beer 804 462 1323 twins + australia + beer + beethoven 1245 735 2252 table 11: training data for the augmented characters in general; these topics will be referred to as the original topics, because these are what the original knowledge base was designed for. all the training data for the twins were authored by hand. the base character was successively augmented by adding three of the automatically generated knowledge bases from experiment 1 (the choice of knowledge bases was arbitrary). these will be referred to as the new topics. we trained a total of five augmented characters in addition to the baseline; table 11 shows the number of questions, answers and links in each of the sets of training data. 5.1.2 test set original topics. to test performance of the augmented characters on questions from the twins’ original topics we use an extensive test set collected during the initial stages of the twins’ deployment at the museum of science, when visitor interaction was done primarily through trained handlers (relying on handlers allowed us to deploy the characters prior to collecting the required amount of visitor speech, mostly from children, necessary to train acoustic models for speech recognition). the handlers relay the visitors’ questions through a microphone to be processed by a speech recognizer; they also tend to reformulate user questions to better match the questions in the twins’ knowledge base, and many of their utterances are a precise word for word match of utterances in the twins’ training data. such utterances are a good test case for the classifier because the intended correct responses are known, but actual performance varies due to speech recognition errors; they thus test the ability of the classifier to overcome a noisy input. the same test set was used in wang et al. (2011) to compare various methods of handling speech recognizer output; here we use it to compare different character knowledge bases. the speech recognizer output remains constant in the different test runs – all characters are tested on exactly the same utterance texts. to a classifier for the original topics, question-answer pairs from the new topics can be considered as training noise; what the different characters test, then, is how the addition of knowledge bases for the new topics affects the performance of the original, handauthored part of the character. the test set consists of 7690 utterances. these utterances were collected on 56 individual days so they represent several hundred visitors; the majority of the utterances (almost 6000) come from two handlers. each utterance contains the original speech recognizer output retrieved from the system logs (speech recognition was performed using the sonic toolkit, pellom and hacıoğlu, 2001/2005). some of the utterances are identical – there is a total of 2264 utterance types (speech recognizer output), corresponding to 265 transcribed utterance types (transcriptions were performed manually). the median word error rate for the utterances is 20% (mean 29%, standard deviation 139 yao, tosch, chen, nouri, artstein, leuski, sagae, traum 36%). this level of word error rate is acceptable for this application – as we will see below, the original character fails to understand only 10% of the input utterances, and this error rate declines rapidly when the character is allowed to identify its own non-understanding (figure 1). new topics. we also tested performance of the augmented characters on questions relating to the new topics. since we do not have an extensive set of spoken utterances as for the twins’ original topics, we used the same test sets constructed for experiment 1. 5.1.3 evaluation original topics. to evaluate performance on questions from the original topics, we ran our test set through each of the characters in table 11. for each utterance we sent the text of the speech recognizer output to npceditor, and compared the response to the answers linked to the corresponding manual transcription. a response was scored as correct if it matched one of the linked answers, otherwise it was scored as incorrect. we also collected the confidence scores reported by npceditor in order to enable the analysis in figure 1 below (the confidence score is the inverse of the kullback-leibler divergence between the language models of the ideal response and the actual response; see leuski and traum, 2010). new topics. performance on questions relating to the added knowledge bases was evaluated as in the previous experiments, by sending the text of the question to npceditor and manually comparing the response to the predetermined answer key. since experiments 1 and 2 have already established that this procedure is highly reliable, the rating was performed by just one person (the fourth author). 5.2 results 5.2.1 performance on the original topics just counting the correct and incorrect responses is not sufficient for evaluating character performance, because npceditor employs dialogue management logic designed to avoid the worst outputs. during training, npceditor calculates a response threshold based on the classifier’s confidence in the appropriateness of selected responses: this threshold finds an optimal balance between false positives (inappropriate responses above threshold) and false negatives (appropriate responses below threshold) on the training data. at runtime, if the confidence for a selected response falls below the predetermined threshold, that response is replaced with an “off-topic” utterance that asks the user to repeat the question or takes initiative and changes the topic (leuski et al., 2006b); such failure to return a response (also called non-understanding, bohus and rudnicky, 2005) is usually preferred over returning an inappropriate one (misunderstanding). the capability to not return a response is crucial in keeping conversational characters coherent, but it is not captured by standard classifier evaluation methods such as accuracy, recall (proportion of correct responses that were retrieved), or precision (proportion of retrieved responses that are correct). we cannot use the default threshold calculated by npceditor during training, because these default thresholds yield different return rates for different characters. we therefore use a visual evaluation method that looks at the full trade-off between return levels and error rates (artstein, 2011). for each test utterance we logged the top-ranked response together with its confidence score, and then we plotted the rate of off-topics against errors at each possible threshold; this was done 140 conversational characters figure 1: trade-off between errors and non-returns on the original topics separately for each character (since confidence scores are based on parameters learned during training, they are not comparable across characters). figure 1 shows the curves for the baseline character and the five augmented characters: non-returns are plotted on the horizontal axis and corresponding error rates on the vertical axis; at the extreme right, where no responses are returned, error rates are necessarily zero for all characters. lower curves indicate better performance. the best performer on the test set for the original topics is the original twins character, with a 10% error rate when all responses are returned, and virtually no errors with a non-return rate above 20%. performance degrades somewhat with the successive addition of automatically generated questions from the new topics, though the degradation is mitigated to some extent when higher nonreturn rates are acceptable. in exchange for an increased error rate on questions from the original topics, the augmented characters can now answer questions pertaining to the new topics. 5.2.2 performance on the new topics the added knowledge bases are identical to those in experiment 1, so that represents the ceiling we can expect for performance of the augmented characters on the new topics. we found that performance was only slightly degraded. we tested each question set on those characters that included the relevant knowledge base. the number of correct or partially correct answers is shown in table 12; in each case, the correct answers are a subset of the correct answers from experiment 1. we also tested the question sets on the original twins character – as expected, none of the returned responses was a correct answer. 6. discussion the experiments demonstrate that our approach is viable – using question generation tools to populate character knowledge bases in question-answer format results in virtual characters that can give appropriate answers to user questions at least some of the time. some types of questions do better than others, and who questions do particularly well. the differences between the question types probably have to do with the question generation tools and the kinds of question-answer pairs they 141 yao, tosch, chen, nouri, artstein, leuski, sagae, traum test set character australia beer beethoven n = 9 10 14 experiment 1 5 8 9a twins + australia 5 twins + beer 7 twins + beethoven 9 twins + australia + beer 5 7 twins + australia + beer + beethoven 5 6 9 twins 0 0 0 athis number differs from that in table 5 because one answer which was marked as correct in experiment 1 was marked as incorrect in experiment 3. table 12: correct answers from the augmented characters extract. it is no surprise that characters rarely give an appropriate answer to questions without an answer in the source text; the solution to this problem is twofold – find source texts that contain the information users want to ask about, and enable a mechanism for the character to recognize questions that cannot be answered, together with strategies for appropriate responses that are not answers (patel et al., 2006; artstein et al., 2009). however, there remain many user questions with an answer in the source text that the character is not able to find, and this is where there is substantial room for improvement. a key factor for any question answering character is getting a good match between actual questions the users want to ask and the answers the character is able to provide. a study of user questions can guide the creators of a character towards appropriate texts that contain answers to the common questions. the questions collected in our study show that people ask different kinds of questions for the various topics presented to them; this can serve as the beginning of a systematic study of question patterns that depend on the topic. there is also a need to bridge the gap between the vocabulary of user questions and that of questions extracted from the source texts, through improvements to the question generation process and the use of lexical resources. the current work suggests several directions for future research. our experiments always rated the response that got the highest ranking from npceditor. however, npceditor is more nuanced than that, and it can use the confidence scores that rank the responses to tell to some degree whether the chosen response is likely to be correct or whether it is more likely that an appropriate response is not available. this functionality may allow the character itself to judge the quality of its answers. additionally, our main test set of questions and answers was collected ahead of time in a questionnaire format. it constitutes a broad test set that can be used to compare different question generation or classification mechanisms. ultimately, however, the purpose of this research is to create conversational virtual characters, so it would be appropriate to also test the characters in conversation. 142 conversational characters acknowledgments we wish to thank michael heilman, author of question transducer, for giving us access to his code and allowing us to use it in our experiments. the project or effort described here has been sponsored by the u.s. army research, development, and engineering command (rdecom). statements and opinions expressed do not necessarily reflect the position or the policy of the united states government, and no official endorsement should be inferred. references ron artstein. error return plots. in proceedings of the sigdial 2011 conference, pages 319–324, portland, oregon, june 2011. association for computational linguistics. url http://www.aclweb.org/anthology/w/w11/w11-2037. ron artstein, sudeep gandhe, jillian gerten, anton leuski, and david traum. semi-formal evaluation of conversational characters. in orna grumberg, michael kaminski, shmuel katz, and shuly wintner, editors, languages: from formal to natural. essays dedicated to nissim francez on the occasion of his 65th birthday, volume 5533 of lecture notes in computer science, pages 22–35. springer, heidelberg, may 2009. j. kathryn bock. syntactic persistence in language production. cognitive psychology, 18(3):355– 387, 1986. dan bohus and alexander i. rudnicky. sorry, i didn’t catch that! – an investigation of non-understanding errors and recovery strategies. in proceedings of the 6th sigdial workshop on discourse and dialogue, pages 128–143, lisbon, portugal, september 2005. peter f. brown, stephen a. della pietra, vincent j. della pietra, and robert l. mercer. the mathematics of statistical machine translation: parameter estimation. computational linguistics, 19 (2):263–311, 1993. url http://www.aclweb.org/anthology/j/j93/j93-2003.pdf. grace chen, emma tosch, ron artstein, anton leuski, and david traum. evaluating conversational characters created through question generation. in proceedings of the twenty-fourth international florida artificial intelligence research society conference, pages 343–344, palm beach, florida, may 2011. aaai press. abdessamad echihabi and daniel marcu. a noisy-channel approach to question answering. in proceedings of the 41st annual meeting of the association for computational linguistics, pages 16–23, sapporo, japan, july 2003. association for computational linguistics. url http://www.aclweb.org/anthology/p/p03/p03-1003.pdf. jenny rose finkel, trond grenager, and christopher manning. incorporating non-local information into information extraction systems by gibbs sampling. in acl ’05: proceedings of the 43rd annual meeting on association for computational linguistics, pages 363–370, morristown, nj, usa, 2005. doi: 10.3115/1219840.1219885. sudeep gandhe and david traum. evaluation understudy for dialogue coherence models. in 9th sigdial workshop on discourse and dialogue, columbus, ohio, 2008. 143 yao, tosch, chen, nouri, artstein, leuski, sagae, traum sudeep gandhe and david traum. i’ve said it before, and i’ll say it again: an empirical investigation of the upper bound of the selection approach to dialogue. in proceedings of the sigdial 2010 conference, tokyo, 2010. a. graesser, j. otero, a. corbett, d. flickinger, a. joshi, and l. vanderwende. guidelines for question generation shared task and evaluation campaigns. in v. rus and a. graesser, editors, the question generation shared task and evaluation challenge workshop report. the university of memphis, 2009. arno hartholt, jonathan gratch, lori weiss, and the gunslinger team. at the virtual frontier: introducing gunslinger, a multi-character, mixed-reality, story-driven experience. in zsófia ruttkay, michael kipp, anton nijholt, and hannes högni vilhjálmsson, editors, intelligent virtual agents: 9th international conference, iva 2009, amsterdam, the netherlands, september 14– 16, 2009 proceedings, volume 5773 of lecture notes in artificial intelligence, pages 500–501, heidelberg, september 2009. springer. doi: 0.1007/978-3-642-04380-2 62. michael heilman and noah a. smith. question generation via overgenerating transformations and ranking. technical report cmu-lti-09-013, carnegie mellon university language technologies institute, 2009. michael heilman and noah a. smith. good question! statistical ranking for question generation. in proc. of naacl/hlt, 2010. hugo hernault, paul piwek, helmut prendinger, and mitsuru ishizuka. generating dialogues for virtual agents using nested textual coherence relations. in helmut prendinger, james lester, and mitsuru ishizuka, editors, intelligent virtual agents: 8th international conference, iva 2008, tokyo, japan, september 1–3, 2008 proceedings, volume 5208 of lecture notes in artificial intelligence, pages 139–145, heidelberg, september 2008. springer. doi: 10.1007/978-3-54085483-8 14. boris katz. using english for indexing and retrieving. in proceedings of the 1st riao conference on user-oriented content-based text and image handling (riao ’88), 1988. klaus krippendorff. content analysis: an introduction to its methodology, chapter 12, pages 129– 154. sage, beverly hills, california, 1980. anton leuski and david traum. practical language processing for virtual humans. in twentysecond annual conference on innovative applications of artificial intelligence (iaai-10), 2010. anton leuski, brandon kennedy, ronakkumar patel, and david traum. asking questions to limited domain virtual characters: how good does speech recognition have to be? in proceedings of the 25th army science conference, 2006a. anton leuski, ronakkumar patel, david traum, and brandon kennedy. building effective question answering characters. in proceedings of the 7th sigdial workshop on discourse and dialogue, sydney, australia, july 2006b. willem j. m. levelt and stephanie kelter. surface form and memory in question answering. cognitive psychology, 14(1):78–106, 1982. 144 conversational characters manish mehta and andrea corradini. handling out of domain topics by a conversational character. in proceedings of the 3rd international conference on digital interactive media in entertainment and arts, dimea ’08, pages 273–280, new york, ny, usa, 2008. acm. doi: 10.1145/1413634.1413686. jack mostow and wei chen. generating instruction automatically for the reading strategy of selfquestioning. in vania dimitrova, riichiro mizoguchi, benedict du boulay, and art graesser, editors, artificial intelligence in education – building learning systems that care: from knowledge representation to affective modelling, volume 200 of frontiers in artificial intelligence and applications, pages 465–472, amsterdam, 2009. ios press. elnaz nouri, ron artstein, anton leuski, and david traum. augmenting conversational characters with generated question-answer pairs. in question generation: papers from the aaai fall symposium, pages 49–52, arlington, virginia, november 2011. aaai press. ronakkumar patel, anton leuski, and david traum. dealing with out of domain questions in virtual characters. in jonathan gratch, michael young, ruth aylett, daniel ballin, and patrick olivier, editors, intelligent virtual agents: 6th international conference, iva 2006, marina del rey, ca, usa, august 21–23, 2006 proceedings, volume 4133 of lecture notes in artificial intelligence, pages 121–131, heidelberg, august 2006. springer. doi: 10.1007/11821830 10. bryan pellom and kadri hacıoğlu. sonic: the university of colorado continuous speech recognizer. technical report tr-cslr-2001-01, university of colorado, boulder, 2001/2005. url http://www.bltek.com/images/research/virtual-teachers/sonic/pellom-tr-cslr-2001-01.pdf. paul piwek, hugo hernault, helmut prendinger, and mitsuru ishizuka. t2d: generating dialogues between virtual agents automatically from text. in catherine pelachaud, jean-claude martin, elisabeth andré, gérard chollet, kostas karpouzis, and danielle pelé, editors, intelligent virtual agents: 7th international conference, iva 2007, paris, france, september 17–19, 2007 proceedings, volume 4722 of lecture notes in artificial intelligence, pages 161–174, heidelberg, september 2007. springer. doi: 10.1007/978-3-540-74997-4 16. silvia quarteroni and suresh manandhar. a chatbot-based interactive question answering system. in ron artstein and laure vieu, editors, decalog 2007: proceedings of the 11th workshop on the semantics and pragmatics of dialogue, pages 83–90, rovereto, italy, may 2007. antonio roque and david traum. a model of compliance and emotion for potentially adversarial dialogue agents. in proceedings of the 8th sigdial workshop on discourse and dialogue, pages 35–38, antwerp, belgium, september 2007. nico schlaefer, petra gieselman, and guido sautter. the ephyra qa system at trec 2006. in the fifteenth text retrieval conference proceedings, gaithersburg, md, november 2006. lee schwartz, takako aikawa, and michel pahud. dynamic language learning tools. in proceedings of the 2004 instil/icall symposium, 2004. sidney siegel and n. john castellan, jr. nonparametric statistics for the behavioral sciences, chapter 9.8, pages 284–291. mcgraw-hill, new york, second edition, 1988. 145 yao, tosch, chen, nouri, artstein, leuski, sagae, traum david a. smith and jason eisner. quasi-synchronous grammars: alignment by soft projection of syntactic dependencies. in proceedings of the workshop on statistical machine translation, pages 23–30, new york, june 2006. association for computational linguistics. url http://www.aclweb.org/anthology/w/w06/w06-3104.pdf. william swartout, david traum, ron artstein, dan noren, paul debevec, kerry bronnenkant, josh williams, anton leuski, shrikanth narayanan, diane piepol, chad lane, jacquelyn morie, priti aggarwal, matt liewer, jen-yuan chiang, jillian gerten, selina chu, and kyle white. ada and grace: toward realistic and engaging virtual museum guides. in jan allbeck, norman badler, timothy bickmore, and alla pelachaud, catherine safonova, editors, intelligent virtual agents: 10th international conference, iva 2010, philadelphia, pa, usa, september 20–22, 2010 proceedings, volume 6356 of lecture notes in artificial intelligence, pages 286–300, heidelberg, september 2010. springer. david traum, antonio roque, anton leuski, panayiotis georgiou, jillian gerten, bilyana martinovski, shrikanth narayanan, susan robinson, and ashish vaswani. hassan: a virtual human for tactical questioning. in proceedings of the 8th sigdial workshop on discourse and dialogue, pages 71–74, antwerp, belgium, september 2007. ellen m. voorhees. overview of the trec 2003 question answering track. in proceedings of the twelfth text retrieval conference (trec 2003), pages 54–68, 2004. url http://trec.nist.gov/pubs/trec12/papers/qa.overview.pdf. mengqiu wang, noah a. smith, and teruko mitamura. what is the jeopardy model? a quasisynchronous grammar for qa. in proceedings of the 2007 joint conference on empirical methods in natural language processing and computational natural language learning (emnlpconll), pages 22–32, prague, czech republic, june 2007. association for computational linguistics. url http://www.aclweb.org/anthology/d/d07/d07-1003.pdf. william yang wang, ron artstein, anton leuski, and david traum. improving spoken dialogue understanding using phonetic mixture models. in proceedings of the twenty-fourth international florida artificial intelligence research society conference, pages 329–334, palm beach, florida, may 2011. aaai press. brendan wyse and paul piwek. generating questions from openlearn study units. in scotty d. craig and darina dicheva, editors, aied 2009: 14th international conference on artificial intelligence in education, workshops proceedings: volume 1, the 2nd workshop on question generation, pages 66–73, brighton, uk, july 2009. 146 ibanezetal-edit-ms-190320 dialogue & discourse 11(1) doi: 10.5087/dad.2020.102 ©2020 romualdo ibáñez, fernando moncada, benjamín cárcamo and valentina marín this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). signaling of causal coherence relations in spanish: variety, functionality, and specificity in academic contexts romualdo ibáñez romualdo.ibanez@pucv.cl pontificia universidad católica de valparaíso el bosque 1290, viña del mar, chile fernando moncada fmoncada@ubiobio.cl universidad del bío bío av. brasil 1180, chillán, chile benjamín cárcamo benjamin.carcamo@pucv.cl pontificia universidad católica de valparaíso el bosque 1290, viña del mar, chile valentina marín valentina.marin.b @mail.pucv.cl pontificia universidad católica de valparaíso el bosque 1290, viña del mar, chile editor: manfred stede submitted 01/2019; accepted 01/2020; published online 03/2020 abstract while recent studies on spanish have shown that some causal connectives specialize in expressing certain types of causal relations, others have revealed that causal relations may be signaled by a variety of linguistic devices. given that we were interested not only in specificity and variety, but also in the functionality of causal connective expressions, our objective in the present study was threefold. first, to identify the variety of connective expressions (connectives and cue phrases) used to signal causal relations in spanish. second, to determine whether a relationship of specificity exists between connective expressions and particular types of causal relations. third, to describe the functionality of those connective expressions. we analyzed a corpus of 2,514 causal coherence relations previously annotated and identified in a corpus of academic texts. 41 different linguistic devices used to signal causal relations were identified. these devices were grouped into two main functional classes: connectives and cue phrases. regarding the functionality of the signals, we found that 8 of the most frequent connective expressions were used to signal different relations. as for specificity, in terms of syntactic categories, it was observed that various conjunctions and conjunctive adverbs specialize in signaling specific relations. keywords: causal coherence relations, connectives, cue phrases, spanish signaling of causal coherence relations in spanish 41 1 introduction based on the assumption that there is no one-to-one mapping between connective expressions and coherence relations, scholars have devoted considerable attention to the description of coherence relation signaling. this interest has given rise to studies that explore patterns of signaling from different theoretical and methodological approaches. from an approach that assumes coherence relations as explicit or implicit, studies have shown that some causal connectives specialize in expressing certain types of causal relations. this specialization is directly related to the notion of specificity proposed by spooren (1997), who claims that a coherence relation is said to be specified, when it is marked by a connective that is prototypically used to encode its meaning, and underspecified when is marked by one that is not. the notion of specificity has also been explored by scholars interested in causal connectives, who refer to them as specified causal connectives or underspecified causal connectives, depending on the degree of subjectivity a connective is involved with (li et al., 2017). for example, the dutch connective dus (‘so’) is frequently used in epistemic relations (stukker et al., 2009) while omdat (‘because’) is predominant in volitional content ones (sanders et al., 2012); according to this perspective, both can be considered specified, as they specify the degree of subjectivity the relations are involved with. this pattern of use has also been observed in other languages. the french car (‘for’) and puisque (‘since’) are mostly used in epistemic relations while parce que, in content ones (anscombre & ducrot, 1983; degand & pander maat, 2003; zufferey, 2010). in german, weil (‘because’) is predominant in content relations whereas denn (‘because’) is typical in epistemic ones (stukker & sanders, 2012). in spanish, porque (‘because’), ya que (‘since’) and puesto que (‘given that’) are typically used to signal causal coherence relations (montolío, 2001); and puesto que (‘given that’) is preferred over porque (‘because’) to express subjective relations (santana et al., 2018). a pattern of use different from specificity is that of polyfunctionality, which means that a connective can be used to signal different types of coherence relations (fischer, 2006; blackwell, 2016). in spanish, connectives porque (‘because’) and debido a (‘due to’) may be considered polyfunctional, as they may signal either subjective or objective causal coherence relations (santana et al., 2017; santana et al., 2018; cárcamo, 2019). from a different approach, other researchers question the existence of implicit coherence relations (taboada, 2009; taboada & das, 2013; das, 2014; das & taboada, 2017). so, they focus on the variety of linguistic devices that may signal coherence relations, beyond connectives (or discourse markers). studies that take this approach have shown that, in spanish, the same coherence relation (concession) can be signaled by 18 different types of signals, including connectives (pero/‘but’); other connective expressions, such as prepositional phrases (para+ np, con+np); and a variety of linguistic devices, such as gerunds (siendo/‘being’), and impersonal clauses (bien es cierto que/‘it is well known that’) (taboada & gómez-gonzález, 2012). with regard to causal coherence relations in spanish, it has been observed, not only that most of them (97%) are signaled, but also, that they may be signaled by linguistic devices other than connective expressions, such as lexical items (psychological verbs, such as enojarse/‘to get angry’), nonfinite verbs (miguel fue castigado por llegar tarde a casa/‘miguel was punished for arriving home late’), and genre structure (objective or purpose coherence relations in the third rhetorical move of an abstract) (duque, 2014). studies from these two approaches have shed light on important issues regarding the signaling of causal coherence relations in spanish; however, there are still other aspects that have been understudied. one of them is the degree of specificity of the relation between particular types of coherence relations and connective expressions other than connectives (or discourse markers). another aspect is the phenomenon of functionality of causal connective expressions, i.e. the possibility that a connective expression may signal different types of causal relations (fischer, 2006; redeker & gruber, 2014). finally, it also seems necessary to gather information regarding the interaction between connective expressions and coherence relations in different ibáñez, moncada, cárcamo and marín 42 contexts of use. to go further into the description of these phenomena would not only widen our understanding of the way coherence relations are signaled in spanish, it would also contribute to the delineation of the way connective expressions and causal coherence relations interact across languages. in order to account not only for specificity and variety, but also for functionality in the signaling of causal coherence relations in spanish, we take an integrative approach. hence, our objective is threefold. first, to identify the variety of connective expressions (connectives and cue phrases) used to signal causal relations in spanish. second, to determine whether a relationship of specificity between connective expressions and particular types of causal relations exists. third, to describe the functionality of those connective expressions. to achieve our objectives, we analyzed a corpus of 2,514 causal coherence relations previously identified by ibáñez et al. (2015). we added signaling information in what could be understood as a process of annotation upon annotation (taboada & das, 2013). given that the analysis was carried out in causal coherence relations identified in texts belonging to different academic genres used in university programs of biology and law, this study provides information regarding the signaling of causal coherence relations in academic contexts. the article is organized as follows. first, we provide an introduction to the concept of causal coherence relations. second, we present our conception of the relation between causal coherence relations and different connective expressions. third, we present the methods. fourth, we present and discuss the results. finally, we discuss the implications of those results and provide the conclusion. 2 causal coherence relations in the last four decades, several proposals have been developed for the description and classification of the different existent types of coherence relations (mann & thompson, 1988; carlson & marcu, 2001; sanders et al., 1992; hovy et al., 1992; polanyi et al., 2004, among many others). in spite of the fact that such proposals vary considerably in terms of the types and number of relations or the specificity of the groupings, all of them include the group of causal relations. the predominance of such relations can be explained because of the importance that causality has not only in human cognition (salmon, 1997; sanders, 2005; noordman & de blijzer, 2000), but also in text organization (meyer, 2000). in fact, it is claimed that all languages have specific means to express causality (pander maat & sanders, 2000). for the spanish language, extensive studies have been conducted to describe those means, paying special attention, among other aspects, to the encoding of causality in syntactic structures (galán, 1999; gutiérrez ordoñez, 2002), the use of explicative causal conjunctions (goethals, 2002), the relation between causativity, agentivity and transitivity (gozalo gómez, 2004) and the interaction between punctuation markers and connectives in causal constructions (figueras, 2000) (for an excellent overview, see arroyo, 2017). important attempts have been made to describe and classify the coherence relations that encode causality and to explain their system and use. one approach to describing coherence relations, particularly causal ones, is the cognitive approach to coherence relations (ccr) (sanders, et al., 1992). in ccr, causality is defined as the implicational meaning that can be inferred between consecutive discourse segments. therefore, causality is not restricted exclusively to those causal relations that connect two events in the physical world (fact 1 leads to fact 2), but it is also present in cases where one event leads to a conclusion (based on fact 1, someone concludes x), where one event occurs given certain circumstances (if x, then y), when someone performs an action to reach a purpose (action x is done to achieve goal y), and even in cases where a final state/result is not expected (x, however y). therefore, relations typically signaling of causal coherence relations in spanish 43 referred to as condition, purpose, or concession are considered causal in ccr and other frameworks based on it. in its original version, ccr presents a set of four basic cognitive primitives that can be used to organize different types of coherence relations (mostly causal) that language users infer between two or more segments in a text: basic operation, source of coherence, order of segments and polarity. basic operation indicates the strength of the semantic link between segment 1 (s1) and segment 2 (s2), and it has two values: causal and additive. a relation is causal if there is an implication relation (p → q) between the two segments, in which p is antecedent and q is consequent. on the contrary, in additive relations the only relation that can be inferred between the segments is conjunction (p & q). source of coherence refers to the nature of the link established between the segments. it is semantic, if the link is established at the level of the propositional content; or pragmatic, if it is established at the level of illocutionary meaning. in further developments these values have been reformulated. semantic relations have been referred to as content relations while pragmatic relations have been subdivided into speech acts (links motivated by illocutionary force) and epistemic ones (connections that involve logical reasoning and inferences) (spooren & sanders, 2008). order of segments accounts for the correspondence between each discourse segment and its role as either antecedent (p) or consequent (q) in the coherence relation. thus, there is basic order when p corresponds to the s1 and q, to the s2. conversely, the order is nonbasic when the opposite sequence is present. provided this criterion reflects two different orders in which p and q can be presented in the connected discourse segments, it distinguishes causeconsequence relations (e.g. s1. as a result, s2) from consequence-cause relations (s1 because s2) or claim-argument (s1 since s2) from argument-claim (s1 therefore s2) relations. finally, polarity distinguishes between negative and positive relations. a negative relation holds when the relation between s1 and s2 involves the negation of the propositional content of one of the segments, while a positive relation holds when there is no such a negation. negative relations also involve the violation of the expectations generated by p, whereas in positive relations, q is in line with what can be expected according to p. for instance, in peter studied hard during the semester and he passed the course, polarity is positive since passing a course (q) is something one can expect for someone who has studied hard (p). on the contrary, in peter studied hard during the semester but he failed the course, polarity is negative because failing a course (q) is not a typical (or expected) result for someone who has studied hard (p). since some negative relations involve an implication operation (as described above), they are considered as causal in ccr, different from cases like peter studied hard during the semester but i did not (where there is no implication). let us consider examples (1), (2), and (3) to illustrate the classification of causal relations according to the ccr approach. (1) [en la regulación alostérica, el efector se combina con la enzima en un lugar diferente del centro activo, denominado centro alostérico.] s1 por ello, [ocurre una modificación en el centro activo de la enzima.] s2 [‘in allosteric regulation, the effector combines with the enzyme in a different place of the active center, named allosteric center.] s1 because of that, [a modification in the enzyme active center occurs’]. s2 (2) [los neurotransmisores regulan la transmisión de impulsos nerviosos.]s1 por lo tanto, [desempeñan un papel fundamental en el funcionamiento del sistema nervioso.] s2 [‘neurotransmitters regulate the transmission of nervous impulses.] s1 therefore, [they play a crucial role in the functioning of the nervous system.’] s2 ibáñez, moncada, cárcamo and marín 44 (3) [los cloroplastos transfirieron menos dna, en comparación con las mitocondrias.] s1 por lo tanto, [es posible que la endobiosis de las mitocondrias haya ocurrido antes de la endobiosis que originó los cloroplastos.] s2 [‘chloroplasts transferred less dna compared to mitocondria.] s1 therefore, [it is possible that the endobiosis of mitochondria had occurred before the endobiosis that created chloroplasts.’] s2 in (1), (2) and (3), the basic operation is causal since a relation of implication can be deduced between s1 and s2. in all cases, the polarity is positive since s2 does not imply a negation of s1. the order of segments is basic, given that s1 operates as the antecedent (p) and segment 2, as the consequent (q). one may conclude, then, that these examples reflect the same type of causality and, therefore, that they could be classified under the same label. however, the difference between these examples can be identified by taking into account their source of coherence. according to spooren and sanders (2008), in (1) a content relation holds since the states of affairs described in s1 and s2 occur in the physical world. on the contrary, in (2) a speech act relation holds given that s2 is a claim by the author based on evidence presented in s1. in (3), the relation is epistemic since s2 is an inference made by the author based on the evidence provided in s1. therefore, according to the ccr approach, (1) is a case of nonvolitional cause, while (2) is evaluation, and (3) interpretation. ccr has become an important discourse annotation scheme. it has been used in several studies covering languages such as german (pit, 2003), dutch (stukker, 2005; spooren & sanders, 2008), mandarin chinese (li et al., 2013; wei, 2018) and recently, spanish (santana et al., 2018). since the original 1992 proposal, several modifications and updates have been proposed for the interpretation and operationalization of the original primitives (see hoek, 2018; hoek et al., 2019 for a state-of-the-art in ccr). for instance, a further distinction has been made in source of coherence, regarding the degree of subjectivity causal relations may encode. the notion of subjectivity refers to the degree to which a reasoning entity or subject of consciousness (soc) is involved in the construal of the coherence relation (degand & pander maat, 2003; pander maat & degand, 2001; pander maat & sanders, 2000, 2001). the more present an soc is in the construction of the relation, the more subjective the relation is (epistemic relations). hence, if the soc is absent, a relation will be objective (non-volitional content relations). following those principles, (1) can be classified as an objective relation because the causal link is established between two events of the physical world, without the intervention of an soc. on the other hand, (2) and (3) are considered subjective since the causal link is mediated by the participation of an soc. ccr’s original primitives and later reformulations have greatly influenced other approaches to coherence relations. such is the case of the top-down bottom-up approach developed by ibáñez et al. (2015) and ibáñez et al. (2019) for the spanish language. this approach has been used to annotate corpora of academic genres and school textbooks of different disciplines, allowing the identification of types of relations that vary in their frequency depending on the discipline. for instance, when analyzing a corpus of academic genres used in law and biology, ibáñez, et al. (2015) identified that causal relations that involve obligation (as (4) below) were highly frequent in legal texts but were not present in the texts from biology. these relations, labeled as condition-obligation, differ from the other conditional relations of the taxonomy in that the consequent (q) is not an action that an agent performs voluntarily (condition-action) nor is a state that results when a condition is met (condition-event) but is an action that an agent is required to do (see appendix). signaling of causal coherence relations in spanish 45 (4) [si el ciudadano no respeta la norma,] s1 [debe someterse a juicio.] s2 [‘if the citizen shall not respect the law,’] s1 [‘he or she must be judged.’] s2 as described in section 3.1, the study reported in this paper was based on the taxonomy of ibáñez et al. (2015), and it was conducted upon their corpus of causal coherence relations extracted from academic genres. 2.1 causal coherence relations and connective expressions despite the considerable amount of research on coherence relations and their signals, there is still no general consensus on some fundamental issues. the first has to do with the nature of the signal. to some researchers, coherence relations are explicit, when they are signaled by a connective expression that indicates the link between the discourse segments, and implicit, when there is no connective expression involved (knott & dale, 1994; meyer & webber, 2013; fraser, 2009; van der vliet & redeker, 2014). others claim that there are no implicit coherence relations, given that coherence relations may also be signaled by a range of linguistic devices whose primary function is not to mark coherence relations, such as non-finite verbs, genre-structure, punctuation, lexical items, or even sentence mood (taboada, 2009; taboada & das, 2013; das & taboada, 2017; das, 2014). another issue in which there seems to be no consensus has to do with the names given to connective expressions. in the literature, they are referred to as cue phrases (knott & dale, 1994; knott, 1996; knott & sanders, 1998), discourse markers (schiffrin, 1987; portolés, 1998; zorraquino & portolés, 1999; taboada, 2006; redeker & gruber, 2014), pragmatic markers (fraser, 1999, 2009), discourse operators (redeker, 1990, 1991), discourse particles (fischer, 2006; aijmer, 2002; briz et al., 2008), connectives (pons, 1998; degand & pander maat, 2003; van der vliet & redeker 2014; evers-vermeul et al., 2017), among other names (fraser, 2009). there is also lack of consensus regarding the way connective expressions are conceived. to some researchers (redeker & gruber, 2014), connectives are the same as discourse markers, while to others (portolés, 1993, 1998; zorraquino & portolés, 1999; duque, 2014), the former correspond to a subgroup of the latter. simultaneously, some other researchers (fraser, 1996, 2009) understand discourse markers as part of pragmatic markers, while others (knott & dale, 1994; knott, 1996) use cue phrase as a covering term for all kinds of expressions that signal coherence relations, including discourse markers. despite the terminological and conceptual heterogeneity (das 2014), there is consensus on the fact that connective expressions comprise a functional class that signals the coherence relation that holds between two discourse segments. though, there is no one-to-one link between connective expressions and the type of coherence relation they signal. in fact, sometimes, a type of coherence relation may be signaled by different connective expressions (condition relation may be signalled by if, unless, when, and since), showing lack of specificity (spooren, 1997), while some other times, the same connective expression may signal different coherence relations – a phenomenon identified as ambiguity by some researchers (stede, 2014) and as a polyfunctional profile of use by others (fischer, 2006; redeker & gruber, 2014) (connective but may signal relations of contrast, concession and antithesis). as a functional class, connective expressions are drawn from different syntactic classes, such as subordinating and coordinating conjunctions, complex prepositions, and conjunctive adverbs (fraser, 1999, 2009; redeker & gruber, 2014). they can be one word (such as adverbs) or multiword expressions (such as complex prepositions) (stede, 2014; crible & cuenca, 2017). they may be fixed expressions, such as conjunctive adverbs, or less frozen expressions, such as prepositional phrases or other unsystematic constructions (das, 2014; duque, 2014). descriptions of connective expressions in spanish treat multi-word fixed expressions as locutions (pons, 1998; zorraquino & portolés, 1999; galán, 1999; real academia española, 2019). hence, and ibáñez, moncada, cárcamo and marín 46 depending on the syntactic category the multi-word fixed expression is associated to, it may be classified as prepositional locution (a pesar de/‘in spite of’), adverbial locution (sin embargo/‘however’), or conjunctive locution (puesto que/‘given that’) (pavón lucero, 1999). in the present study, two main types of causal connective expressions are distinguished: causal connectives (cc), conceived as one-word or multi-word invariable expressions, whose main function is to signal the causal coherence relation that holds between discourse segments; and cue phrases (cp), which comprise those less frozen expressions that signal the causal coherence relation that holds between discourse segments, and may allow for syntactic modification. while the former are drawn from syntactic categories, such as conjunction (porque/‘because’) – including conjunctive locution (ya que/‘given that’) –, conjunctive adverb (sin embargo/‘however’), and complex preposition (a pesar de/‘in spite of’), the latter are usually expressed by prepositional phrases (por esa razón/‘for that reason’) and other unsystematic constructions (preposition + infinitive). coherence relation frequency 1 cause effect 228 2 effect causes 33 3 action reason 69 4 reason action 59 5 act purpose 112 6 purpose act 83 7 claim argument 446 8 argument claim 311 9 condition event 374 10 event condition 264 11 condition obligation 128 12 obligation condition 7 13 basic contrast 281 14 non basic contrast 70 15 evidence deduction 47 16 deduction evidence 2 total 2,514 table 1. types and instances of causal coherence relations of the corpus. signaling of causal coherence relations in spanish 47 3 methods in order to account for variety, functionality, and specificity of causal connective expressions, we followed the approach adopted by taboada and das (2013). we added a new layer of information (causal connective expressions) to a corpus already annotated for coherence relations (ibáñez et al., 2015). the original corpus consisted of 27 complete exemplars (762,737 words) of academic genres (textbook, disciplinary text, and research article), written in spanish and used in undergraduate programs of law and biology. 3.1 corpus the corpus of the present study corresponds to the 2,514 causal coherence relations identified in the original corpus (ibáñez et al., 2015), which are classified into sixteen types, as shown above in table 1. 3.2 the annotation of connective expressions the annotation of connective expressions was carried out following two sequential steps: identification and classification. identification: this procedure started by distinguishing explicit from implicit relations based on whether or not they were marked, following some of taboada and das (2013) and das (2014) conditions for considering an expression to be a signal: 1. the scope of the function of a signal is a single discourse sequence comprising adjacent discourse segments in a relation. 2. signals mark relations that hold between two discourse segments. 3. signals constitute a functional class of lexical expressions drawn from different syntactic classes. as shown in (5), the two discourse segments (in square brackets) are connected by a claim-argument relation. the relation is explicit since it is signaled by the connective ya que. (5) [tales juicios no son apropiados] s1, ya que [no resuelven el asunto de manera definitiva.]s2 [‘such trials are not appropriate,’] s1 since [‘they do not definitively solve the issue.’] s2 classification: in order to classify the connective expressions identified in our corpus, an extensive bibliographical review of different proposals of spanish causal connective expressions was carried out (pavón lucero, 1999; domínguez garcía, 2007; martí, 2008; martínez, 1997; montolío, 2001; portolés, 1998; pons, 1998; zorraquino & portolés, 1999; galán, 1999; real academia española, 2019). most of these proposals classify connectives as functional categories, according to the type of coherence relation they signal (concessive marker/connective, contrastive connective/marker, additive connective/marker, etc.). given that this study is driven by the assumption that there is no one-to-one correspondence between signals and types of relations, and that, as functional expressions, they are drawn from different syntactic classes, causal connective expressions were also classified according to syntactic classes (pons, 1998; real academia española, 2019). hence, for the functional class of causal connective, we used the categories of conjunction, conjunctive locution, conjunctive adverb, and complex preposition, while for the functional class of cue phrase, we used the categories of prepositional phrases and constructions (preposition+infinitive). ibáñez, moncada, cárcamo and marín 48 3.3 procedure due to the number and variety of forms and, mainly, to their use in combination, classification of causal signals is neither simple nor objective. for that reason, in order to ensure inter rater agreement, a two-coders-discuss strategy was used (spooren & degand, 2010). after a training period of one month, two pairs of coders (all postgraduate students in linguistics) were assigned the 2,514 causal relations. each pair was asked to identify and classify the causal connective expressions in the whole corpus, according to the categories selected for the study. first, causal connective expressions were classified independently by the two coders in each pair. after that, in case of disagreement, differences were discussed until agreement was reached by the pair. subsequently, the results of both pairs were compared. the agreement index was k= 0.78, which can be interpreted in this context as substantial (landis & koch, 1977; spooren & degand, 2010). 4 results in this section, and based on the analysis performed to our corpus of academic texts, we first present the variety of connective expressions used to signal causal coherence relations in spanish. then, we show the functionality of the most frequent connective expressions, and, finally, the degree of specificity of the relation between those signals and particular causal relations. 4.1 variety of connective expressions the first objective of the present study was to identify the variety of connective expressions used to signal causal relations in spanish. out of the 2,514 causal relations analyzed, 2,284 (92%) were signaled, while only 230 were not (8%), which shows the predominance of explicit causal relations in the corpus. this is in line with prior research that has established that spanish coherence relations tend to be signaled in some manner most of the time (taboada & das, 2013; duque, 2014). the analysis also revealed the existence of 41 different linguistic devices, which were grouped into two functional classes (causal connectives and cue phrases) and six syntactic categories, as shown in table 2. table 2 shows the wide variety of linguistic devices used to signal causal relations in the corpus of study. it can be observed that out of the two functional classes, connective encompasses most cases of causal coherence relations signaling. by far, the most frequent syntactic class observed is conjunction, which constitutes on its own 51.72% of all the linguistic devices identified in the corpus. this is significant considering conjunction only includes 9 of the devices used. the second highest frequency can be attributed to conjunctive adverb, which constitutes 14.93%. as for cue phrase, the linguistic devices are very similarly distributed between prepositional phrases and preposition + infinitive construction, which constitute 7.97% and 9.16% respectively. based on this data, it can be suggested that causal coherence relations are prototypically signaled by conjunction in spanish written texts, in the academic contexts of biology and law. these findings contribute to the study of signaling devices in spanish, considering most of the research conducted on this matter is diachronic rather than synchronic (garcía-cervigón, 2006; herrero, 1999; zagona, 2002) or purely theoretical (cid, 2002) rather than empirical. signaling of causal coherence relations in spanish 49 functional class syntactic category occurrence percentage linguistic devices connective conjunction 1174 51.72% si, aunque, porque, como, pero, pues, cuando, mientras, así conjunctive locution 322 14.19% ya que, para que, puesto que, siempre que, de modo que, a menos que, aun cuando, salvo que, dado que, con tal que conjunctive adverb 339 14.93% sin embargo, por lo tanto, si bien, por consiguiente, en consecuencia, por cuanto, no obstante, una vez, entonces complex preposition 46 2.03% a pesar de, a fin de, en caso de cue phrase prepositional phrase 181 7.97% por eso, por lo que, por ello, en este caso, por lo mismo, por lo cual, por esa razón preposition + infinitive constructions 208 9.16% para + infinitivo, por + infinitivo, al + infinitive table 2. connective expressions, classification and occurrence. 4.2 functionality of connective expressions in order to describe the functionality of the causal connective expressions, the 10 most frequent linguistic devices identified in our corpus (which represent 68.3% of the signaled relations) were examined. the results are shown in table 3. as it can be observed, 9 of the 10 most frequent linguistic devices belong to the category of connective and one to cue phrase. our data confirms that most linguistic devices are polyfunctional, since they signal various types of causal relations, ranging from two (aunque/‘although’) to five (porque/‘because’), even though there are others that only signal one relation such as pero (‘but’) and sin embargo (‘however’). table 3 shows that the conjunction si (‘if’), the most frequent in the corpus, allows the signaling of four types of relations that involve conditionality. over half of the cases in which si is used correspond to condition-event (53%), a relation that follows a basic order (antecedent-consequent) and whose consequent depicts a ibáñez, moncada, cárcamo and marín 50 connective expressions functional class frequency % coherence relations signalled si (‘if’) connective 434 19 event-condition (25%) condition-event (53%) condition-obligation (21%) obligation-condition (1%) porque (‘because’) connective 200 8.75 action-reason (14%) claim-argument (74%) argument-claim (2%) deduction-evidence (1%) effect-cause (9%) para + infinitivo (‘to + infinitive’) cue phrase 171 7.48 act-purpose (49%) purpose-act (44%) event-condition (7%) cuando (‘when’) connective 137 5.99 event-condition (40%) condition-obligation (6%) condition-event (54%) pues (‘since’) connective 133 5.82 action-reason (8%) claim-argument (77%) effect-cause (2%) argument-claim (13%) ya que (‘since’) connective 125 5.47 action-reason (8%) argument-claim (3%) claim-argument (86%) effect-cause (3%) sin embargo (‘however’) connective 102 4.46 basic contrast (100%) pero (‘but’) connective 90 3.94 basic contrast (100%) aunque (‘although’) connective 85 3.72 basic contrast (51%) non basic contrast (49%) por lo tanto (‘therefore’) connective 83 3.63 argument-claim (72%) cause-effect (18%) evidence-deduction (6%) reason-action (4%) 1,560 68.3 table 3. 10 most frequent causal connective expressions and the relations they signal. signaling of causal coherence relations in spanish 51 situation occurring in the physical world without intentionality. the same conjunction is also used in event-condition (25%), a relation that differs from the previous one only in terms of the order in which the segments are presented (consequent-antecedent). this conjunction frequently also signals relations whose consequent represents an obligatory situation (condition-obligation, 21%). another interesting profile of use is observed in the conjunction porque (‘because’), since it is a device that signals the widest variety of relations (5). among them, it signals relations mediated by intentions (action-reason, 14%), relations that involve speakers’ stance (claimargument (74%) and argument-claim (2%)), and relations that express causality in the physical world (effect-cause, 9%). 4.3 specificity of connective expressions in order to determine whether a relation of specificity between connective expressions and types of causal relations exists, the chi-square goodness of fit test (χ2) was used. this statistical test assesses the distribution of categorical data and compares it against a specific proportion by which observed and expected frequencies are examined. a χ2 was run for each linguistic device based on the coherence relations they signaled. in all instances an even hypothesized distribution was assumed for each test. this decision was methodologically important since it allowed us to differentiate the expectations of each of the devices. when a connective was found to signal different coherence relations, it was assumed as the null hypothesis for the chi-square goodness of fit test that it would not show any preference toward a particular coherence relation. in those cases where the null hypothesis was rejected, adjusted standardized residuals were used to complement the significance of the statistical test. those cases where the residuals were higher than 3.0 were considered to represent specificity, since this number has been considered to reflect whether the number of observations is significantly larger than expected (agresti, 2018). following this procedure, several connective expressions were found to signal specific types of coherence relations, as shown in table 4. it is worth mentioning that there are two connectives (pero/‘but’ and sin embargo/‘however’) that are used exclusively in basic contrast relations. it should be noted that in (ibáñez et al., 2015), and similar to ccr (although with another label), this type of relation is regarded as causal negative because a relation of implication holds between the connected discourse segments. such implication, in turn, involves the generation of expectations, which in negative relations are not fulfilled. in other words, in negative relations, like basic contrast, the consequent (q) is not in line with the expectations triggered by the antecedent (p), as illustrated in (6) and (7) below. given that pero (‘but’) and sin embargo (‘however’) are used only in basic contrast relations, both connectives could be regarded as having a profile of specificity towards basic contrast. (6) en general, las células de los cultivos primarios mueren después de un cierto número de mitosis (50 a 100 mitosis). pero a veces algunas células experimentan mutación y se hacen inmortales. ‘in general, cells of primary cultures die after a certain number of mitosis (50 to 100 mitosis). but sometimes some cells experience mutation and become immortal.’ (basic contrast/biology/handbook) (7) el sistema nos parece mucho más adelantado que el de la república argentina y el de brasil. sin embargo, su fundamento es inadmisible y utilitario. ‘the system seems to us to be much more advanced than that of argentina’s and brazil’s. however, the argument is inadmissible and utilitarian.’ (basic contrast/law/handbook) ibáñez, moncada, cárcamo and marín 52 connective expressions observed frequency expected frequency chi-square goodness of fit test specific coherence relation cuando (‘when’) 71 55 34 34 (χ2(3) = 100.635, p = <.05) condition-event event-condition para + inf. (‘to+ infinitive’) 84 75 43 43 (χ2(3) = 128.485, p = <.05) act-purpose purpose-act por lo tanto (‘therefore’) 59 17 (χ2(4)= 142.361, p = <.05) argument-claim porque (‘because’) 147 40 (χ2(4) = 369.350, p = <.05) claim-argument pues (‘since’) 102 27 (χ2(4) = 273.729, p = <.05) claim-argument si (‘if’) 229 109 χ2 (3) = 237.318, p = <.05 condition-event ya que (‘since’) 108 31 (χ2(3) = 252.248, p = <.05) claim-argument table 4. observed and expected frequency of the most frequent connective expressions across the coherence relations they signal.1 among those connectives used to signal more than one relation, a clear pattern of specificity was found for si (‘if’). this conjunction signals various types of conditional relations, but, as shown in table 4, it is predominantly used in condition-event. this relation is established when the (non) occurrence of one or more events determine the (non) occurrence of others, as illustrated in (8) and (9). 1 sin embargo (‘however’) and pero (‘but’) have not been included in table 4 because all observed instances are basic contrast. aunque (‘although’) has not been included because it shows an even distribution between basic contrast (n=43) and non-basic contrast (n=42) (χ2(1) = .012, p = 0.914). signaling of causal coherence relations in spanish 53 (8) si un aparato mitótico aislado se trata con estas sustancias, su estructura se pierde rápidamente. ‘if an isolated mitotic apparatus is treated with these substances, its structure is lost quickly.’ (condition-event/biology/handbook) (9) si los tribunales nacionales carecen de ella, la ley fija su falta de competencia. ‘if national courts do not have it, the law fixes its lack of competence.’ (condition-event/law/handbook) another clear case of specificity was found for porque (‘because’) and pues (‘since’), and ya que (‘since’). as mentioned above, these connectives signal a wide repertoire of causal relations: porque signals 5 types of relations, and pues and ya que, 4 (see table 3). interestingly, table 4 shows that in spite of such polyfunctionality, these connectives show specificity for the same relation: claim-argument. this causal relation, as illustrated in (10), (11) and (12), involves the statement of an opinion or judgment, which is expressed in the first segment of the causal relation and functions as the consequent (q). the evidence or arguments for such a claim is presented in the second segment of the causal relation and functions as the antecedent (p). (10) será necesaria la prueba indirecta, porque el hecho no está presente o ha dejado de existir. ‘the indirect evidence will be required because the fact is not present or has ceased to exist.’ (claim-argument/law/disciplinary text) (11) no es sólo una simple mezcla de estas sustancias, pues el protoplasma tiene una organización muy compleja. ‘it is not just a simple mixture of substances, since the protoplasm has a very complex organization.’ (claim-argument/biology/disciplinary text) (12) la glicolisis es un proceso poco eficiente, ya que de las 690 kcal mol presentes en la glucosa, apenas 20 son aprovechadas. ‘glycolysis is not an efficient process since out of the 690 kcal present in glucose, only 20 are used.’ (claim-argument/biology/disciplinary text) another profile of specificity was found for cuando (‘when’). this connective, usually used as a signal of temporality in spanish, is also frequently used for causal relations, particularly in those that involve conditionality. although cuando is used to signal 3 types of relations (see table 3), it shows a preference for two: condition-event and event-condition. these relations are used to present a state (p), whose occurrence results in the occurrence of another (q), as shown in (13), or to present a state (q), whose occurrence results from the occurrence of another (p), as shown in (14). ibáñez, moncada, cárcamo and marín 54 (13) la presión predatoria tiende a aumentar cuando crece la población de presas. ‘predatory pressure tends to increase when prey population grows.’ (event-condition/biology/disciplinary text) (14) cuando son colocadas en una solución hipertónica, las células disminuyen de volumen. ‘when placed in a hypertonic solution, cells decrease their volume’. (condition-event/biology/disciplinary text) in addition, while conducting the analysis, it was noticed that the connective expressions also show patterns of specificity towards subjective or objective causal relations when considering the nature of the coherence relations they predominantly signal. for example, following ibáñez et al. (2015), (15) is a case of condition-event, a relation whose source of coherence is non-volitional. in cases like this, the link between the antecedent (p) and the consequent (q) is established in the physical world and there is no volition or obligation involved in the state of affairs expressed in q. therefore, it can be regarded as an objective relation. on the contrary, in (16) an argumentclaim relation holds, where the source of coherence is speech act. in cases like this, q corresponds to a claim used by the speaker with the support of the argument presented in p. therefore, the speaker is responsible for the construal of the causal relation. thus, (16) is considered a subjective relation. (15) cuando las células musculares o hepáticas son expuestas a la hormona adrenalina, hay un aumento en el contenido intracelular de camp. ‘when muscular or hepatic cells are exposed to the adrenaline hormone, there is an increase in the intracellular content of camp.’ (condition-event/biology/handbook) (16) ya que gobernar y ser gobernado es algo diferente, hemos de suponer que la moderación de los gobernantes no es idéntica a la moderación de los gobernados. ‘since to govern is different from being governed, we have to expect that the moderation of the leaders is not identical to the moderation of those being governed.’ (argument-claim/law/disciplinary text) following this line of reasoning, cuando (‘when’), pero (‘but’), sin embargo (‘however’) and aunque (‘although’) can be regarded as showing a preference for objective relations since they predominantly signal objective relations such as condition-event and basic contrast. on the other hand, por lo tanto (‘therefore’), porque (‘because’), pues (‘as’), and ya que (‘since’) may be considered as subjective since they predominantly signal subjective relations such as argumentclaim and claim-argument (see table 4). the results reported up to this point reveal that regardless of the fact that some connectives show a polyfunctional profile (i.e., they signal different types of relations), they feature a clear pattern of specificity. these findings lead us to suggest that, at least in two academic contexts (biology and law), some causal connectives are highly specific for particular relations. from another perspective, it could be argued that certain causal relations are predominantly signaled by specific connective expressions. therefore, further studies could focus on specific sets or subsets of causal relations to determine whether such relations show a preference for certain connectives. signaling of causal coherence relations in spanish 55 5 discussion and conclusions the relation between coherence relations and connectives has been understudied in spanish. in fact, most research has traditionally been aimed at providing comprehensive descriptions and grammar-oriented classifications of connectives, without considering coherence relations (gili gaya, 1961; portolés, 1993, 1998; zorraquino & portolés 1999, fuentes 1987, 1996; iglesias recuero, 2000; caravedo, 2003; arroyo, 2017). more recent discourse-oriented studies have switched the focus, paying special attention to the relation between coherence relations and connectives. in this scenario, it is possible to identify, at least, two types of studies on explicit causal coherence relations: those interested in describing the prototypical use of causal connectives (cao et al., 2016; santana et al., 2017; santana et al., 2018) and those intended to account for the variety of linguistic devices that may signal causal relations (taboada & gómezgonzález, 2012; duque, 2014). to the best of our knowledge, there have not been previous studies combining both approaches. therefore, the current corpus-based study aimed not only to describe the variety of linguistic devices used to signal causal relations in the spanish language, but also to explore their functionality and specificity. in line with taboada and das’ suggestions (2013), instead of starting from scratch, we re-used an already annotated corpus to which we added an extra layer (their signaling). we used a corpus of 2,514 causal coherence relations previously annotated (ibáñez et al., 2015), which were identified in academic written texts in spanish. the manual analysis carried out enabled us, in the first place, to identify 41 different connective expressions used to signal causal relations. these devices were grouped into two main classes: connectives and cue phrases. within the class of connectives, four types were distinguished: conjunctions (si/‘if’, porque/‘because’, pues/‘since’, pero/‘but’), conjunctive locutions (puesto que/‘given that’, ya que/‘since’), conjunctive adverbs (sin embargo/‘however’) and complex prepositions (a pesar de/‘in spite of’). in the case of cue phrases, two types were identified: prepositional phrases (por esa razón/‘for that reason’) and preposition + infinitive constructions (para conseguir/‘to achieve’). an interesting finding is that out of the 10 most frequent connective expressions identified, 9 are connectives. among these, the most frequent one is the conjunction si (‘if’), which is coherent with the predominance of conditional relations in the corpus. this reinforces the idea that connectives are the prototypical means of marking coherence relations in spanish. furthermore, in line with previous studies in spanish (montolío, 2001; domínguez garcía, 2007), our data shows that porque (‘because’) is one of the most typical ones. actually, in our corpus it has the second highest frequency, which supports the idea that this conjunction is one of the most commonly used to express causality in spanish (dominguez garcía, 2007). some of the other frequent conjunctions in our data, such as ya que (‘since’), have been characterized as typical signals of causality in spanish in previous research (pit et al., 1996; goethals, 2002; duque, 2016; santana et al., 2018). therefore, our results contribute to provide new empirical evidence on the most typical signals used to mark causality in the spanish language. regarding the functionality of the connective expressions, we analyzed the 10 most frequent ones and found that 8 of them signal more than one relation. an interesting case was observed in the conjunction porque (‘because’). as claimed in previous studies (montolío, 2001; domínguez garcía, 2007), this conjunction is one of the most frequently used because it allows to express different types of causal relations. our data provides evidence not only for its high frequency in language use, but also for its polyfunctionality. specifically, our study shows that porque signals 5 different types of relations, ranging from those that involve volition (action-reason) to those mediated by the speaker's stance (claim-argument, argument-claim) to those belonging to the content domain (effect-cause). a similar case was observed in the profile of use of cuando (‘when’), a conjunction which, in the spanish language, is often used as a signal for temporality (cuando era joven, solía viajar en motocicleta, ‘when i was young, i used to ride a motorbike’). ibáñez, moncada, cárcamo and marín 56 our data shows that cuando has the fourth highest frequency, which proves that its use in causal relations is not peripheral. in fact, cuando allows the signaling of 3 different types of conditional relations, having the same meaning as si (‘if’). cases like these, when the semantics of the connective used to explicit a relation does not fully match the semantics of the relation that is intended by the speaker/writer, is what spooren (1997) calls underspecification. previous psycholinguistic research (spooren, 1997; li et al., 2017) has shown that underspecified coherence relations impose different cognitive efforts compared to those relations signaled by a typical (specific) connective. therefore, further research is needed to investigate whether such differences are also found in the spanish language. regarding specificity, and based on previous literature, we expected that certain connective expressions would show a preference for specific causal relations. our results were in line with our expectations since it was observed that various conjunctions and conjunctive adverbs specialize in signaling specific relations. for instance, pero (‘but’) and sin embargo (‘however’) are used to signal exclusively negative causal relations (basic contrast). other cases are the conjunctions si (‘if’), which is used mostly in conditional relations (condition-event), and pues (‘since’), which is used frequently in relations involving a speaker’s stance (claim-argument). in addition, based on the causal relation they frequently signal, our data showed that certain connective expressions show a preference for subjective meanings. por lo tanto (‘therefore’), porque (‘because’), pues (‘as’), and ya que (‘since’) are mostly used in subjective relations, where the speaker is involved in the construal of the relation, for instance through the statement of a conclusion or an opinion. these findings differ from previous corpus-based studies. santana et al (2018), for instance, found that porque (‘because’), and ya que (‘since’) were not associated with subjective features, which led them to conclude that spanish does not have connectives that have a clear subjectivity profile. they argued that spanish porque, similar to english because, is used to express both subjective and objective relations (sweetser, 1990; knott & sanders, 1998). a possible explanation for this difference is how subjectivity was operationalized. santana et al. (2018) used an analytical model that decomposes the notion of subjectivity in a series of components (domain, modality, presence of the soc and identity of the soc), whereas in our study the distinction between objective and subjective relations was based only on sweetser’s (1990) distinction for the source of coherence of the relation. besides, since our analysis was performed on academic genres only, it could be possible that the subjective pattern shown by porque and ya que may be due to the predominantly argumentative nature of the genres they were extracted from and of the disciplines (law). it could be that in other (non-academic) genres, with other discourse organization modes (such as narrative or descriptive), these connectives may display more objective profiles. in consequence, further research is required to extend our understanding of subjectivity in spanish and the devices that express it. complementary, psycholinguistic studies could provide us with new insights about the processing of connectives depending on their degree of specificity towards subjectivity. it could be that some differences may be observed in the processing and comprehension of subjective relations signaled by connectives that vary their degrees of specificity. specifically, it could be expected, for instance, that a claim-argument relation is processed faster when it is signaled by ya que than by porque. the main contribution of the current study is the approach adopted to account for the signaling of causal relations in spanish. different from previous studies that have analyzed either how a particular relation is marked or whether a specific marker is associated to specific relations, we combined both perspectives. in addition, given that we did not construct a taxonomy of connective expressions to analyze their profiles of use, our results provide a richer picture on the ways causal relations are actually signaled and on how those connective expressions are used, at least in academic written texts. further research, focused on spontaneous conversations, for instance, may shed new light on usage pattern of causal connective expressions in conversations, which may differ from the patterns reported here. studies such as the ones conducted by günthner (1993) and keller (1995) have demonstrated that a connective like german weil signaling of causal coherence relations in spanish 57 (‘because’) usually expresses epistemic relations in conversations, while in written texts it is used for content domain relations. therefore, it could be possible that a typical connective like porque may present different usage patterns depending on modality. acknowledgments the first author’s work was enabled by a grant awarded by fondecyt (project 1160094) from the national commission for scientific and technological research (conicyt). the second author was funded by postdoctoral grant fondecyt 3180779 from the national commission for scientific and technological research (conicyt). the third author was funded by doctoral grant conicyt-pfcha/doctorado nacional/2017-21170031. appendix. causal coherence relations used by ibáñez et al. (2015) causal coherence order of events polarity source of coherence content speech act epistemic basic positive causeeffect reason action condition obligation argumentclaim evidence deduction non basic positive effectcause actionreason obligation -condition claim argument basic negative basic contrast non basic negative non basic contrast basic positive conditionevent condition– action non basic positive eventcondition basic positive purposeact non basic positive actpurpose references alan agresti (2018). an introduction to categorical data analysis. new york, wiley & sons. karin aijmer (2002). english discourse particles. evidence from a corpus. studies in corpus linguistics. amsterdam, john benjamins. jean-claude anscombre and oswald ducrot (1983). l'argumentation dans la langue. brussels, editions mardaga. ignacio arroyo (2017). la expresión de causa en español. madrid, visor libros. sarah blackwell (2016). porque in spanish oral narratives: semantic porque, (meta) pragmatic or both? perspectives in pragmatics, philosophy & psychology, 4: 615-651. antonio briz, salvador pons and josé portolés (2008). diccionario de partículas discursivas del español. in el diccionario como puente entre las lenguas y culturas del mundo. actas del ii congreso internacional de lexicografía hispánica. alicante, biblioteca virtual cervantes: 217-227. shuyuan cao, iría da cunha and nuria bel (2016). an analysis of the concession relation based on the discourse marker aunque in spanish chinese parallel corpus. procesamiento del lenguaje natural, 56: 81-88. ibáñez, moncada, cárcamo and marín 58 rocío caravedo (2003). principios del cambio lingüístico. una contribución sincrónica a la lingüística histórica. revista de filología española 83(1): 39-62. benjamín cárcamo (2019). subjectivity in spanish causal connectives: differentiating porque, ya que and debido a que. spanish in context 16(1): 51-76. lynn carlson and daniel marcu (2001). discourse tagging reference manual. isi technical report isi-tr-545: 54-56. manuel cid (2002). las conjunciones coordinantes del español actual desde el punto de vista funcional. boletín de lingüística 18: 49-70. ludivine crible and maría josé cuenca (2017). discourse markers in speech: characteristics and challenges for corpus annotation. dialogue & discourse 8(2): 149-166. debopam das (2014). signaling of coherence relations in discourse. phd thesis, simon fraser university, burnaby. debopam das and maite taboada (2017). rst signalling corpus: a corpus of signals of coherence relations. language resources & evaluation 52(1):149-184. liesbeth degand and henk pander maat (2003). a contrastive study of dutch and french causal connectives on the speaker involvement scale. in a. verhagen & j. van de weijer (eds.), usage based approaches to dutch: 75-199. utrecht: lot. maría domínguez garcía (2007). conectores discursivos en textos argumentativos breves. madrid, arco/libros. eladio duque (2014). signaling causal coherence relations. discourse studies 16(1): 25-46. eladio duque (2016). las relaciones del discurso. madrid, arco/libros. jacqueline evers-vermeul, jet hoek and merel scholman (2017). on temporality in discourse annotation: theoretical and practical considerations. dialogue & discourse 8(2): 1-20. carolina figueras solanilla (2000). puntuación e interpretación de las expresiones causales en el texto escrito. in j. de bustos tovar et al (eds.), lengua, discurso, texto: i simposio internacional de análisis del discurso: 281-295. madrid, visor libros. kerstin fischer (ed.). (2006). approaches to discourse particles. amsterdam, elsevier. bruce fraser (1996). pragmatic markers. pragmatics. quarterly publication of the international pragmatics association (ipra) 6(2): 167-190. bruce fraser (1999). what are discourse markers? journal of pragmatics 31(7): 931-952. bruce fraser (2009). an account of discourse markers. international review of pragmatics 1(2): 293-320. catalina fuentes (1996). aproximación a la estructura del texto. málaga, ágora. catalina fuentes (1987). enlaces extraoracionales. sevilla, universida de sevilla. carmen galán (1999). la subordinación causal y final. in v. demonte and i. bosque, gramática descriptiva la lengua española: 3597-3642. madrid, espasa calpe. alberto garcía-cervigón (2006). el grupo del nombre en la analogía de la grae, 1771-1917. madrid, editorial complutense. samuel gili gaya (1961). curso superior de sintaxis española. barcelona, bibliografs/a. patrick goethals (2002). las conjunciones causales explicativas en castellano. un estudio semiótico-lingüísta. leeuven, peeters. paula gozalo gómez (2004). la expresión de la causa en castellano. madrid, universidad autónoma de madrid. susanne günthner (1993). weil – man kann es ja wissenschaftlich untersuchen: diskurspragmatische aspekte der wortstellung in weil-sätzen. linguistische berichte 143: 37-55. salvador gutiérrez ordóñez (2002). «causales», boletín de la real academia española. in s. gutiérrez ordóñez (ed.). forma y sentido en sintaxis: 100-208. madrid, arco libros. francisco javier herrero (1999). sobre la evolución de las oraciones y conjunciones adversativas. revista de filología española 79(3): 292-328. signaling of causal coherence relations in spanish 59 jet hoek (2018). making sense of discourse: on discourse segmentation and the linguistic marking of coherence relations. phd thesis, utrecht university. utrecht: lot. jet hoek, jacqueline evers-vermeul and ted sanders (2019). using the cognitive approach to coherence relations for discourse annotation. dialogue & discourse 10(2): 1-33. eduard hovy, julia lavid, ema maier, vibhu mittal and cecile paris (1992). employing knowledge resources in a new text planner architecture. in aspects of automated natural language generation, pages 57-72. springer, berlin, heidelberg. romualdo ibáñez, fernando moncada and benjamín cárcamo (2019). coherence relations in primary school textbooks: variation across school subjects. discourse processes, 56(8), 764-785. romualdo ibáñez, fernando moncada and andrea santana (2015). variación disciplinar en el discurso académico de la biología y del derecho: un estudio a partir de las relaciones de coherencia. onomázein (32): 101-131. silvia iglesias recuero (2000). la evolución del “pues” como marcador discursivo hasta el siglo xv. boletín de la real academia española 280(80): 209-308. rudi, keller (1995). the epistemic weil. in d. stein and s. wright (eds.), subjectivity and subjectivisation: linguistic perspectives. 16-30. cambridge: cambridge university press. alistair knott (1996). a data-driven methodology for motivating a set of coherence relations. phd thesis, university of edinburgh. alistair knott and robert dale (1994). using linguistic phenomena to motivate a set of coherence relations. discourse processes 18(1): 35-62. alistair knott and ted sanders (1998). the classification of coherence relations and their linguistic markers: an exploration of two languages. journal of pragmatics 30(2): 135 175. richard landis and gary koch (1977). an application of hierarchical kappa-type statistics in the assessment of majority agreement among multiple observers. biometrics 33(2): 363-374. fang li, jacqueline evers-vermeul and ted sanders (2013). subjectivity and result marking in mandarin. chinese language & discourse 4(1): 74–119. fang li, willem mak, jacqueline evers-vermeul and ted sanders (2017). on the online effects of subjectivity encoded in causal connectives. review of cognitive linguistics 15(1): 34 57. william mann and sandra thompson (1988). rhetorical structure theory: toward a functional theory of text organization. text interdisciplinary journal for the study of discourse 8(3): 243-281. manuel martí sánchez (2008). los marcadores en español l/e: conectores discursivos y operadores pragmátics. madrid, arco/libros. roser martínez (1997). conectando texto: guía para el uso efectivo de elementos conectores en castellano. barcelona, octaedro. paul meyer (2000). the relevance of causality. in e. couper-kuhlen & b. kortmann (eds.), cause, condition, concession and contrast: cognitive and discourse perspectives: 9-34. berlin, mouton de gruyter. thomas meyer and bonnie webber (2013). implicitation of discourse connectives in (machine) translation. in proceedings of the workshop on discourse in machine translation, pages 19-26, sofìa. estrella montolío (2001). conectores de la lengua escrita. contraargumentativos, consecutivos, aditivos y organizadores de la información. barcelona, ariel. leo noordman and femke de blijzer (2000). on the processing of causal relations. in e. couper-kuhlen & b. kortmann (eds.), cause, condition, concession and contrast: cognitive and discourse perspectives: 35-56. berlin, mouton de gruyter. ibáñez, moncada, cárcamo and marín 60 henk pander maat and liesbeth degand (2001). scaling causal relations and connectives in terms of speaker involvement. cognitive linguistics 12(3): 211-246. henk pander maat and ted sanders (2000). domains of use or subjectivity? the distribution of three dutch causal connectives explained. in e. couper-kuhlen and b. kortmann (eds.), cause, condition, concession, contrast: cognitive and discourse perspectives: 57-82. berlin, mouton de gruyter. henk pander maat and ted sanders (2001). subjectivity in causal connectives: an empirical study of language in use. cognitive linguistics 12(3): 247–274. maría victoria pavón lucero (1999). clases de partículas. in v. demonte and i. bosque, gramática descriptiva de la lengua española: 565-656. madrid, espasa calpe. mirna pit (2003). how to express yourself with causal connective: subjectivity and causal connectives in dutch, german and french. new york, rodopi. mirna pit, jacqueline hulst and paander mat (1996). subjectiviteit en de spaanse connectieven porque, ya que en puesto que. gramma/ttt 5(3): 221-240. livia polanyi, chris culy, martin van den berg, gian lorenzo thione and david ahn (2004). a rule-based approach to discourse parsing. in proceedings of the 5th sigdial workshop on discourse and dialogue at hlt-naacl 2004, pages 108-117. cambridge, massachusetts, usa salvador pons (1998). conexión y conectores. estudio de su relación en el registro informal de la lengua. cuadernos de filología, anexo xxvii. valencia, universitat de valència. josé portóles (1993). la distinción entre los conectores y otros marcadores del discurso en español. verba: anuario galego de filoloxia 20: 141-170. josé portóles (1998). marcadores del discurso. barcelona, ariel. real academia española: diccionario de la lengua española, 23.ª ed., [on line version 23.3]. [october, 2019]. gisela redeker (1990). ideational and pragmatic markers of discourse structure. journal of pragmatics 14(3): 367-381. gisela redeker (1991). linguistic markers of discourse structure. linguistics 29(6): 1139-1172. gisela redeker and helmut gruber (2014). introduction the pragmatics of discourse coherence. in h. gruber and g. redeker (eds.), the pragmatics of discourse coherence: theories and applications: 1-22. amsterdam, john benjamins. william salmon (1997). causality and explanations. new york, oxford university press. josé sanders, ted sanders and eve sweetser. (2012). responsible subjects and discourse causality. how mental spaces and perspective help identifying subjectivity in dutch backward causal connectives. journal of pragmatics 44(2): 191-213. ted sanders (2005). coherence, causality and cognitive complexity in discourse. in m. aurnague & m. bras (eds.), proceedings of the first international symposium on the exploration and modelling of meaning, pages 31-46. université de toulouse-le-mirail, france. ted sanders, wilbert spooren and leo noordman. (1992). toward a taxonomy of coherence relations. discourse processes 15(1): 1-35. andrea santana, dorien nieuwenhuijsen, wilbert spooren and ted sanders (2017). causality and subjectivity in spanish connectives: exploring the use of automatic subjectivity analyses in various text types. discours. (20). andrea santana, wilbert spooren, dorien nieuwenhuijsen and ted sanders (2018). subjectivity in spanish discourse: explicit and implicit causal relations in different text types. dialogue & discourse 9(1): 163-191. deborah schiffrin (1987). discourse markers. studies in interactional sociolinguistics. cambridge: cambridge university press. wilbert spooren (1997). the processing of underspecified coherence relations. discourse processes 24(1):149-168. signaling of causal coherence relations in spanish 61 wilbert spooren and liesbeth degand (2010). coding coherence relations: reliability and validity. corpus linguistics and linguistic theory 6(2): 241–266. wilbert spooren and ted sanders (2008). the acquisition order of coherence relations: on cognitive complexity in discourse. journal of pragmatics 40(12): 2003-2026. manfred stede (2014). resolving connective ambiguity: a prerequisite for discourse parsing. in h. gruber and g. redeker (eds.), the pragmatics of discourse coherence: 121-141. amsterdam, john benjamins. ninke stukker (2005). causality marking across levels of language structure. a cognitive semantic analysis of causal connectives in dutch. phd thesis. utrecht university. ninke stukker and ted sanders (2012). subjectivity and prototype structure in causal connectives: a cross-linguistic perspective. journal of pragmatics 44(2): 169-190. ninke stukker, ted sanders and arie verhagen (2009). categories of subjectivity in dutch causal connectives: a usage-based analysis. in t. sanders and e. sweetser (eds.), causal categories in discourse and cognition: 119-171. berlin, de gruyter mouton. eve sweetser (1990). from etymology to pragmatics: metaphorical and cultural aspects of semantic structure. cambridge university press. maite taboada (2006). discourse markers as signals (or not) of rhetorical relations. journal of pragmatics 38(4): 567-592. maite taboada (2009). implicit and explicit coherence relations. in jan renkema (ed.), discourse of course: 127–140. amsterdam, john benjamins. maite taboada and debopam das (2013). annotation upon annotation: adding signalling information to a corpus of discourse relations. dialogue & discourse 4(2): 249-281. maite taboada and maría gómez-gonzález (2012). discourse markers and coherence relations: comparison across markers, languages and modalities. linguistics and the human sciences 6(1-3): 17-41. nynke van der vliet and gisela redeker (2014). explicit and implicit coherence relations in dutch texts. in h. gruber and g. redeker (eds.), the pragmatics of discourse coherence: 23-52. amsterdam, john benjamins. yipu wei (2018). causal connectives and perspective markers in chinese: the encoding and processing of subjectivity in discourse. phd thesis. utrecht university. karen zagona (2002). the syntax of spanish. cambridge university press. martín zorraquino and antonia portolés (1999). los marcadores del discurso. in v. demonte and i. bosque, gramática descriptiva la lengua española: 4051-4213. madrid, espasa calpe. sandrine zufferey (2010). lexical pragmatics and theory of mind: the acquisition of connectives. amsterdam, john benjamins. dialogue & discourse 8(1) (2017) 1–30 doi: 10.5087/dad.2017.101 non-native differences in prosodic-construction use nigel g. ward nigelward@acm.org department of computer science university of texas at el paso and kyoto university paola gallardo pgallardo@miners.utep.edu department of computer science university of texas at el paso editor: dr. amanda stent submitted 10/2015; accepted 12/2016; published online 01/2017 abstract many language learners never acquire truly native-sounding prosody, and often are weak on the dialog-related uses of prosody. previous work has suggested this may involve deficits with specific prosodic constructions, but this has not been systematically investigated. we developed semiautomatic analysis methods able to identify and characterize such differences. starting with two sets of dialog data, one of native speakers and one of non-natives, we applied principal components analysis, and then identified differences in distributions and in the constructions themselves. applied to recordings of six advanced-level native-spanish learners conversing in english, these methods revealed differences in their uses of speaking rate and pitch in turn-taking, and infrequent and variant use of the english prosodic constructions for showing involvement and for explaining. keywords: conversation, dialog, interaction, pragmatics, language learning, non-native prosody, comparison, second language, l2, american english, mexican-spanish speakers, l1 transfer, turntaking, principal components analysis, semi-automatic methods 1. introduction non-native speakers often have saliently non-native prosody, even when their other language skills are good (zimmerer et al., 2014; mennen, 2015). among the various functions of prosody, it has been suggested that the dialog-related aspects might be most important for language learners, since incomplete command of the prosodic forms used for pragmatic functions can impact interactional competence and the achievement of communicative goals (barraja-rohan, 2011). further, nonnatives may show an “over-use of a limited variety of intonation patterns in the l2” and an “underuse” of others (ramirez verdugo, 2003, 2006). this paper reports a corpus-based exploration of these issues. the contributions are 1) semi-automated methods for discovering prosodic differences from data, 2) an inventory of some important prosodic constructions of english, 3) descriptions of the prosodic forms and pragmatic functions of some of these constructions, and 4) the identification of three ways in which the prosody of native-spanish learners of english differs from that of native english speakers. c©2017 nigel ward and paola gallardo this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). ward and gallardo in this article our interest is in “dialog prosody,” by which we mean uses of prosody that help coordinate the interaction and convey attitude, pragmatics, and related functions. we are not here interested in the more-commonly studied aspects of prosody — prosody as it relates to words, syntax, semantics, emotion, and paralinguistics — although of course all uses of prosody are complexly interrelated (couper-kuhlen and selting, 1996b; hirschberg, 2002; szczepek reed, 2012; wichmann, 2014). this paper is organized as follows. first we discuss previous approaches and their limitations (section 2). choosing to describe prosodic skills in terms of prosodic constructions, we applied principal components analysis (pca) (section 3) to native-native dialog data, resulting in the identification of 32 common prosodic constructions of english dialog (section 4). no suitable nonnative corpora existing, we collected 90 minutes of data from six advanced native-spanish speakers in english conversations with native english speakers (section 5). simple statistical measures revealed some differences in prosodic construction use (section 6). model-based comparison revealed others: for this we first identified the non-native prosodic constructions, and then compared these to the native speakers’ constructions (section 7). in section 8 we summarize and note questions for future research. 2. previous methods for characterizing non-native prosodic differences it is not easy to accurately characterize how non-native prosody differs from native-speaker prosody, especially for the dialog-related uses. this section overviews four commonly used approaches and their limitations. the first approach starts with a pragmatic function. the classic example is gumperz’s discussion of cafeteria servers who offered side dishes using a falling accent, a function which english native speakers perform with a rising accent (gumperz, 1982). other work in this vein has examined english question, focus and list-final intonation (ramirez verdugo, 2003; swerts and zerbian, 2010; kainada and lengeris, 2015), italian contrast (turco et al., 2015), spanish turn keeping and information seeking (aronsson and fant, 2014), and backchanneling in german and vietnamese (ha et al., 2016). such analyses rely on the existence of a clear intended function, and thus they cannot give us the whole story. one reason is that speakers in dialog commonly pursue multiple goals with each utterance (bunt, 2011; heritage, 2012). in such cases it is impossible to say definitively what the appropriate prosodic form would be, as native speakers in the same situation might differ in which of the pragmatic functions they choose to prosodically highlight. another reason is that such studies can only discover differences for previously-identified functions, and there is no guarantee that the pragmatic functions identified to date are exhaustive, or even cover those most important in actual conversation. a second approach starts with a syntactic form or sentence type and looks at differences between learners’ and native speakers’ prosodic realizations of it. while often informative, this approach is complicated by the fact that a given syntactic structure may serve different discourse functions, depending on the context and the prosody. indeed, prosody is often determined more by the pragmatic function than the syntactic form, especially in dialog (lai, 2012; hedberg et al., 2014), so findings obtained from such production tasks may not fully reflect actual behavior in dialog. nevertheless most detailed studies of learners’ prosody have used this approach. a third approach characterizes non-native prosody with reference to a model of the appropriate prosodic forms. for example, toivanen’s examination of the distribution of tone types (fall, rise-fall, 2 non-native prosody fall-plus-rise, etc.) showed that learners used rising tones less than the native speakers (toivanen, 2003). however model-based analysis also has its limitations. for models of prosody that are symbolic rather than phonetic, labor-intensive segmentation and/or hand labeling is required before they can be applied to data. more generally, such approaches only work for the aspects of prosody that a model handles, and these are always limited. for example, the most popular current models of prosodic forms are based on monolog data, and they mostly handle only pitch (intonation), leaving out speaking rate, timing, and intensity, although these aspects are also important in modeling learners’ skills (trouvain and gut, 2007; romero-trillo, 2012). a fourth approach uses raw statistics on prosodic usage over corpora. for example, zimmerer’s measurements showed that non-native speakers use much less pitch variation (zimmerer et al., 2014). this method can exploit large amounts of data and is entirely objective. it is also robust: while in any given utterance a non-native speaker may have good reason to use compressed pitch range — for example when losing interest in a topic and preparing to close it out — consistently limited pitch variation across a corpus is good evidence for a real difference. other work in this vein has shown that values of speaking rate and pitch range correlate with assessments of comprehensibility and accentedness (kang, 2010). however raw statistics, being context-independent, cannot pinpoint the locations of the differences, nor their communicative significance. for example, they cannot tell in which specific contexts a wider pitch range would have been appropriate. these methods have provided insights regarding many specific differences in prosody (mennen, 2015). nevertheless each has its weaknesses, and thus we wish to explore a new way to investigate non-native prosody. our approach, like the latter two above, starts with forms, rather than with functions, and is corpus-based. because we are centrally interested in what people actually do in conversation, we wish to discover from the data itself what functions are expressed with prosody, and how non-native speakers differ. 3. prosodic constructions and their automatic discovery describing prosodic behavior is difficult, and there is currently no consensus on how to represent prosodic knowledge. rather there are many different approaches, with different assumptions, methods, and descriptive vocabularies. several good surveys exist (cutler and ladd, 1983; couperkuhlen, 1986; ladd, 1996; wells, 2006; szczepek reed, 2006; van santen et al., 2008; arvaniti, 2011; xu, 2011; prieto, 2015). prosody as used for dialog purposes is especially problematic (kalathottukaren et al., 2015). for this study we chose to use an approach based on an inventory of constructions, for two reasons: it supports semi-automated analysis, and it directly represents the prosodic forms associated with pragmatic functions. this section explains the notion of prosodic construction and how we discover the constructions of a language from a collection of dialog recordings. 3.1 prosodic constructions recently a shared notion of prosodic construction has emerged from work in several research traditions, including conversation analysis, experimental phonetics, autosegmental-metrical intonation modeling, and big-data analysis (ogden, 2007, 2012; petrone and niebuhr, 2013; niebuhr, 2014; hedberg et al., 2014; ward, 2014). prosodic constructions are recurring temporal patterns of prosodic activity that express specific meanings and functions. they typically involve not only 3 ward and gallardo pitch contours but also energy, rate, timing and articulation properties, and may involve synchronized contributions by two participants. for example, in the upgraded assessment construction, as described by ogden (ogden, 2012; ward, 2014), a listener expresses agreement with an assessment by producing an upgraded version, for example when one speaker (a) observes it’s pretty and the other (b) follows with absolutely gorgeous. the upgraded assessment is generally produced with increased intensity, pitch height, and pitch range, and with a ‘tighter’ articulation. often this upgraded assessment follows a bid for some kind of empathy or affiliation. prosodically this involves a speaking loudly for a bit but then trailing off, where the trailing-off is in a lower pitch, and then falling silent for a moment. b’s upgraded assessment in turn is often followed by resumed speech by a that is again louder and tends to last for a few seconds. thus this construction involves interleaved prosodic behaviors by two participants, with specific sequencing and timing. table 1 roughly shows the prototypical temporal configuration of this construction. in jointly performing this construction the participants each express specific attitudes, and together establish a shared assessment and joint interest. time speaker a speaker b –3000 to –300ms speaking quiet –2500 to –1300 speaking louder quiet –800 to 0 tapering in loudness to silence –300 to +300 quiet or silent loud and fast +300 to +600 resuming speaking slowing and fading out +600 to +1300 speaking loudly quiet +1300 to +3200 speaking quiet table 1: major components of a prototypical rendition of the upgraded assessment construction. times are in milliseconds relative to the end of speaker a’s assessment. prosodic constructions share much with the classical notion of intonation contour (liberman and sag, 1974; ladd, 1978). they describe a recurring sequence of prosodic elements in a specific temporal configuration, with some dialog function. these functions often affect the future course of the dialog or the unfolding relationship between the participants. they may in addition have meanings or expressive values, although these are often abstract and highly context-dependent. prosodic constructions extend intonation contours in three ways (niebuhr, 2014; ward, 2014). first, they describe not only patterns of pitch but also include other prosodic features, such as intensity, rate, and timing: thus they are multistream models. second, prosodic constructions are not limited to the behavior of a single speaker, but often describe coordinated actions by two parties. third, they are not necessarily linked to sentences or utterances, but instead can cover arbitrary regions of time. prosodic constructions resemble grammatical constructions (goldberg, 2013) — form-function pairings where the form is a syntactic template and the function is some conventionalized semantic or pragmatic content — in particular in being composable. constructions can be modeled in various ways: qualitatively, symbolically (hedberg et al., 2014), or quantitatively (lai, 2012). for this paper we use quantitative descriptions, as they have two useful properties. first they are superimposable, which suits the fact that any specific time 4 non-native prosody in a dialog may involve multiple prosodic behavior sequences expressing simultaneously-present pragmatic functions. this seems descriptively necessary, and is an essential part of many modern models of prosody (van santen et al., 2004; chen et al., 2004; xu, 2005). second, their presence is graded, meaning that a construction is not simply present or absent, rather it can be present to varying degrees, to the extent that more of the component features are more strongly present and their temporal configuration more closely matches the prototype. for example, a weak version of the upgraded assessment construction might function as a somewhat perfunctory acknowledgment. despite its limitations (ward, 2014), this approach to prosody has the important advantage of enabling the automatic detection of prosodic constructions in unlabeled data. this enables the computation of statistics on construction use and measurements of construction differences. 3.2 prosodic construction discovery by principal components analysis in order to examine non-native uses of dialog prosody as comprehensively as possible, we need a large inventory of constructions. constructions can be discovered in many ways, including inductively by conversation analysis (ogden, 2012), statistically over realizations of known pragmatic functions (hedberg et al., 2003; niebuhr, 2014; hedberg et al., 2014), and semi-automatically (ward, 2014). given similar data, it appears that these methods can all give similar results, so here we used the fastest and easiest: a semi-automated one. several automated and semi-automated discovery methods for intonation contours and other prosodic elements have recently been developed, including some based on clustering, functional data analysis (fda), and pca (itahashi and tanaka, 1993; chen et al., 2005; gubian et al., 2011; parrell et al., 2013; jokisch et al., 2014; reichel, 2014). in this paper we use pca, because it is relatively simple and because it works for raw dialog data, without needing preliminary segmentation or annotation. pca can be described in several ways, but it is convenient to view it as an iterative analysis process. in each stage, pca finds the factor that explains as much as possible of the observed variation, across many datapoints and many variables. it then subtracts out what that factor explains, finds another factor to explain much of the remaining variation, and iterates. for example, if we have statistics on children, including height, weight, running speed, arm strength, lung capacity, stamina, and so on, the first underlying factor may be age, the second something like skinny-chubby, the third socioeconomic status, and so on. the observed variable values for any datapoint (child) are modeled as linear combinations of these underlying factors. conversely, given the observed values for a datapoint, it is trivial to compute the values of the underlying factors, by a simple matrix multiplication. prosodic constructions as we model them — being graded and superimposable — perfectly suit the assumptions of pca: they can serve as the underlying factors that explain the surface, observed, prosody. that is, the observed prosody over any short region of a dialog can be explained as the superimposed effects of multiple, simultaneously-active constructions. thus, our method is to apply pca to datapoints, each of which is a point in time, each described by various observed prosodic features. the output is then a set of dimensions, which are configurations of features that frequently occur together. in this particular study, the datapoints are taken every 10 milliseconds throughout the conversations. this means that, rather than considering prosody only at turn-ends, or only when computed over utterances, as is done in many approaches, the method considers prosody as it appears every5 ward and gallardo where in the conversations. (indeed, datapoints are taken even during silent regions. this makes sense, because silence is also a dialog phenomenon, and can be part of larger patterns of behavior, as will be seen.) 3.3 prosodic features used this subsection documents the prosodic features used as input to pca. since our interest is in dialog prosody, we wanted to use features relevant to the prosodic forms involved in expressing dialog-related functions. while there is no sharp distinction between the prosodic forms involved in in different functions — for example, the same feature can convey either lexical identity or pragmatic function, depending on the language or the speaker — there are tendencies. in particular, it seems that prosodic features which are anchored to or aligned with other linguistic units – syllables, words, sentences — tend to relate more to lexical and syntactic functions. we therefore use unaligned features. this does not mean that other prosodic functions will be entirely excluded from our models, but it does reduce their effects. the features were computed at every timepoint, without relying on any segmentation of the input. while prosodic analysis is often done subsequent to a segmentation of the input, for example into turns, here we do without such preprocessing. one reason is that some prosodic patterns are not turn-aligned, so if we restricted attention to turn-aligned features, the analysis would likely not find such patterns. another reason is that, although the notion of “turn” seems straightforward, in spontaneous conversations turns are difficult to identify reliably, even by humans annotators following strict guidelines, so by avoiding this step, we simplify the process. since we need features that can support the discovery of temporal patterns, for each datapoint we used a number of features at different offsets to broadly represent the local prosodic context. for example, in addition to the intensity over the past 50 milliseconds, we also used the intensity over a 50 millisecond window centered 75 ms in the past, over a 100 ms window centered 150 ms in the past, and so on, for both past and future windows, spanning about 6 seconds centered around the point of interest. including such offset features enables the use of pca for time-series analysis. following previous work, we chose windows of various sizes so as to give greater temporal resolution near the time of interest, that is the timepoint at the center of all the features. since our interest is in prosody, not just intonation, we included not only pitch features but also features for speaking rate, intensity, and creaky voice. while there are many more features that could be included, this set was designed to capture most of the prosodic information that has been found most useful for many tasks (schuller, 2011; shriberg and stolcke, 2004; ward et al., 2011). we followed previous work in having more windows and finer resolution for the features that usually are most informative, notably intensity. in total we used 176 features, as listed in figure 1. these were computed using our open-source toolkit (ward, 2015). this includes built-in normalizations to make the features fairly speakerindependent. for loudness we used log energy normalized per track to correct for different recording conditions and different speakers. for the pitch-height and pitch-range features we used percentiles in the distribution of pitch seen for that track, thus again normalizing for speaker. for speaking rate, we used a simple frame-by-frame energy-difference measure. to avoid the problems associated with interpolating pitch over nonvoiced regions, we used features representing the strength of evidence for the pitch being low (respectively, high, narrow, wide) over a region, using evidence computed only over valid pitch points. for example, our narrow-pitch feature counts the number of pairs of 6 non-native prosody pitch points within a window between which the pitch varies less than 2%. this feature, like the others, was designed to be robust: to roughly match perceptions over a great variety of voices, dialog activities and noise levels. none are entirely reliable, and in particular the speaking-rate proxy, although intended to detect fast speech versus lengthening, also responds to precise articulation (enunciation) versus phonetic reduction, and to creaky versus modal voice. this feature set was designed and refined based largely on experience with various prediction tasks (ward and vega, 2012; ward et al., 2011); experience also shows that minor changes to the feature windows or feature implementations have little effect on the dimensions that pca finds, doubtless due to their overall robustness and to the size of the dataset used. amplitude low pitch, high pitch, creakiness narrow pitch, wide pitch speaking rate (16 per speaker) (14 each, per speaker) (10 each, per speaker) (10 per speaker) -3200 – -1600 -1600 – -800 -1600 – -800 -1600 – -800 -1600 – -800 -800 – -400 -800 – -400 -800 – -400 -800 – -400 -400 – -300 -400 – -300 -400 – -300 -400 – -200 -300 – -200 -300 – -200 -300 – -200 -200 – -100 -200 – -100 -200 – 0 -200 – -100 -100 – -50 -100 – -50 -100 – 0 -50 – 0 -50 – 0 0 – 50 0 – 50 50 – 100 50 – 100 0 – 100 100 – 200 100 – 200 0 – 200 100 – 200 200 – 300 200 – 300 200 – 300 300 – 400 300 – 400 300 – 400 200 – 400 400 – 800 400 – 800 400 – 800 400 – 800 800 – 1600 800 – 1600 800 – 1600 800 – 1600 1600 – 3200 figure 1: the prosodic feature inventory. start and end times for each window in milliseconds offset from the point of interest. these features are computed for these windows for both left and right speakers, giving 176 in total. 3.4 from dimensions to constructions the workflow is summarized in figure 2. for each timepoint in the corpus the prosodic features are computed. pca digests all this data and outputs dimensions. each dimension has a weight on each of the features. for example, on the data set discussed below, dimension 1 (principal component 1) has a high negative weight on speaker-a-amplitudeover-0-50-milliseconds, a high negative weight on speaker-a-amplitude-over-50-100-milliseconds, a positive weight on speaker-b-energy-amplitude-over-0-50-milliseconds, and so on. thus, at times when speaker a is speaking and b silent, the value on dimension 1 will be negative, and for the opposite configuration it will be positive. every dimension codes for two patterns in this way: one when it is present positively, and one when it is present negatively. 7 ward and gallardo dimensions usually have, moreover, temporal variation in the loadings. that is, a certain feature, like high pitch, may be indicative of a pattern being present when it occurs at one time, but not at another. for example, for dimension 1 the loadings on the high-pitch features are high for early windows but then fall, to the extent that, by the 800-1600 millisecond window, the low-pitch loading is greater than the high-pitch loading. this particular pattern of loadings is easy to understand: it corresponds to the well-known prosodic phenomenon of declination. the fact that a simple mathematical operation, pca, can find such a pattern, even though we were not looking for it, illustrates its power. because all patterns observed so far involve extensive temporal variation in loadings, and thus represent temporal configurations of features, it is appropriate to call them constructions, as we will henceforth. feature-vector description feature-vector description feature-vector description feature-vector descriptions … principal components analysis principal components … … … … … … … interpretation feature extractors six seconds of context 720,000 samples from native-speaker dialogs 32 prosodic constructions … figure 2: principal components analysis workflow. 3.5 inferring construction meanings we are interested not only in prosodic patterns, but also in what they mean. while some prosodic patterns, like declination, may just be facts of language (or of a specific language), many have meaning or serve a function. to fully understand how learners’ prosody differs from that of native speakers, we first need to understand these meanings. this is, however not easy: identifying meanings for prosodic constructions is rife with methodological pitfalls (arvaniti, 2011; prieto, 2015). therefore our identifications of meanings for constructions must be regarded as tentative. this section starts with illustrations, and then describes the interpretation process systematically. as mentioned above, one outcome of pca is a prosodic description of each construction: its weightings for each of the features. for example, dimension 1 had a loading of –0.08 on the speaker a amplitude-over-800-to-1600-ms feature. this information being overwhelming with 176 features, it is convenient to use visualizations. for example, figure 3 shows the loadings for dimension 3, showing that the loading for the speaker-a-log-energy (“volume”) feature from -1600 to -800 ms is positive, and so on. examining the other loadings, it is clear that this dimension involves the a speaker (top) speaking and then falling silent, and the b speaker, conversely, being silent and then speaking. thus it encompasses a turn-yielding construction (“dimension 3 lo”) and a turn-taking construction (“dimension 3 hi”). from the figure it is easy to also see some of the prosodic correlates typical of turn-yielding pattern in english: notably increases in intensity, speaking rate, and creakiness, followed by a further increases on the latter two and a simultaneous drop in pitch. table 8 non-native prosody 2 gives a simplified summary of this construction. for reasons of space, this paper only discusses aspects of the loadings that are relevant to non-native speakers’ differences, but all are available at http://www.cs.utep.edu/nigel/l2english/, both numerically and as visualizations. figure 3: loadings of dimension 3. purple solid lines are for the a speaker; green dashed lines for b. time is in milliseconds. the dotted lines are zeros, with points above them indicating positively loaded features and points below negative. the “pitch height” line shows the difference between the loadings of the high-pitch and low-pitch features; similarly “pitch width” is the difference of the wide and narrow features. while this figure shows the strengths of factor loadings, rather than average values for pitch height etc., in practice, instances in the dialogs where this dimension is strongly present do tend to have feature values varying over time as the figure suggests. while the intensity features extend out to 3200 ms before and after the point of interest, to save space we show only 4 secondsworth of feature loadings. of course, no actual instance of turn yield will exactly match this pattern, due to the simultaneous presence of other, superimposed, constructions. however there are cases that match quite well: an example is seen in figure 4, transcribed as example 1. (audio for the examples also is available at the above url.) example 1 (soc008@165.1s) a: i just need to get that lab done, and i’m done with that lab. 9 ward and gallardo time speaker a speaker b –2000 ∼ –1600 ms loud, fast, creaky –1600 ∼ –800 ms louder, faster, creakier –800 ∼ 400 ms pitch drops, quieter –400 ∼ 0 ms quieter falling to silent 0 ∼ 400 ms silent or a quiet, tentative start 400 ∼ 800 ms loud, creaky, fast, high in pitch 800 ∼ 1200 ms feature values revert to typical table 2: major components of a prototypical rendition of the basic turn hand-off construction of english. times are in milliseconds relative to the point halfway between the original speaker’s end and the new speaker’s start. b: what, what about, where, where did you guys get in the homework? we note in passing that most of the properties of the turn hand-off construction, as found here, are also present in other descriptions. for example gravano and hirschberg (2011) found similar turn-yield tendencies involving rate, intensity, non-modal voice, and final pitch. it is also interesting to note that in dimension 3 the loadings of features for the two speakers are nearly symmetric: past-future mirror images across the point of interest (0 milliseconds). we do not ascribe any deep significance to this: pca often results in dimensions with some form of symmetry, and this tendency is stronger here because the features computed for the two sides are identical, and because the two sides are slightly correlated, due to a small amount of cross-track bleeding. dimension 3 was thus easy to understand, but this was not true for all dimensions. for most dimensions we couldn’t infer the meaning from the loadings alone, so we relied more on examination of places in the corpus where a construction was strongly present. specifically, we examined ten to twenty exemplars, places where the value on a given dimension was highest or lowest. for each of these we noted aspects of the context and dialog activity, and the pragmatic functions that were being expressed. these we inferred primarily from information in the dialog itself, including the words being said and the behaviors of both participants in the immediate context, a standard conversation-analysis technique (sidnell, 2011). we also occasionally engaged our own intuitions about what was being conveyed by the observed prosodic form, sometimes by considering the contrast to other forms that the speaker might have used. we then used qualitative-inductive methods to find commonalities among the noted functions and meanings. for some dimensions the commonalities were obvious from just a few examples; for others the commonality did not become clear until we had examined many. in every case, after forming an initial hypothesized meaning, we examined more examples, either finding confirmation or, less often, discovering that we needed to refine or change it. this interpretation process generally went smoothly. however, some of the exemplars were hard to relate to a general hypothesized meaning. one reason is that the pragmatic force of any individual construction depends on the local context, including other constructions simultaneously present. for example, the swift turn exchange in example 1 is not only high on dimension 3, but also fairly high on dimension 2, since there is a lot of talk by both speakers, and low on dimension 10 non-native prosody figure 4: a swift turn exchange, high on dimension 3. pitch is shown at the bottom of each track. 16, primarily since the last syllable of the bottom speaker is short and creaky, indicating his attitude towards the lab. it is also likely that our working assumption, that the meanings contributed by the different constructions are compositional, is not entirely accurate. a second thing that complicated interpretation was individual differences in behavior and uses. for example, for some of the highside exemplars for dimension 10, it seemed that the speaker was being provocative, not just simply disagreeing or diverging, the general functions. quite often, the constructions appeared polysemous. as another example, for some of the low-side dimension 10 exemplars, one speaker was producing a short check questions, or saying something they expected the other to agree with, and the other usually did. in the table we generalize over these and related meanings and refer to the function as “agreeing or aligning.” other analysts could make other choices for such summary phrases. a third complication for the analysis was the presence of creative and deliberate uses of prosody, including for non-literal meanings and for reported speech (rao, 2013b; estelles-arguedas, 2015), many of which did not follow the general tendencies. in a sufficiently advanced model, that properly accounted for all these factors, perhaps the prosody-meaning mappings could be seen to be functioning as exceptionless rules. as it is, our policy was to ascribe meanings to forms if those meanings were clearly present in most of the exem11 ward and gallardo plars. despite this imprecision, we see that pca was effective in identifying factors (dimensions, patterns) that not only explain the observations, but also are meaningful. of course, the demonstrated ability of pca to do this for various types of data is why we chose it, so its effectiveness also for prosody is not a surprise (ward and vega, 2012; ward, 2014; ward et al., 2016). 4. some prosodic constructions of english to quantify the prosodic behavior patterns of the non-native speakers, we need to compare them to some standard. the language norm that our non-native speakers were most familiar with is general american english, especially as spoken by young people in the southwest, so we decided to use the social speech collection (ward and werner, 2013). like the primary data set, described below, this consists of unconstrained conversations among university students studying computer science, although this was recorded for a different purpose, recorded two years earlier, and recorded with different microphones. we took 6 native-native conversations from this corpus, lasting about 10 minutes each, and computed the prosodic features described above for 720,000 data points, taken every 10 milliseconds for both speakers. we then applied the methods described above. table 3 summarizes the results. the second column shows the percentage of variance accounted for by each dimension. it shows, for example, that describing prosodic behavior just with one value, the value on dimension 1, explains 17% of the variance across all 176 features. together the top 16 dimensions account for 55% percent of the variation in these dialogs, suggesting that examination of the top 32 constructions can cover most of the dialog-prosody skillset that learners need. (the number 16 has no special significance; we chose to stop at 16 due simply to limitations of time and space.) our interpretations are summarized in the third column. ideally we would like a full and final listing of the prosodic constructions of english before going on to examine non-native speakers’ differences. of course this listing falls short. in addition to the issues noted above, the details of the feature set we chose are somewhat arbitrary, and with different feature sets the resulting dimensions vary slightly (ward and vega, 2012; ward, 2014). the descriptions are only suggestive, due to reasons of space, although each construction really deserves a paper in itself, to treat its form and function in detail, and to relate it to alternate possible descriptions. we give detail, below, only for constructions that turn out to be used differently by non-native speakers. we also note that our method did not identify exclusively dialog-related aspects of prosody, nor, certainly, all dialog-relevant constructions, not least because our features only cover six-second spans. thus we do not propose this listing as a universally valid or verified list of the pragmatic functions of english prosody. nevertheless all of these functions have been previously identified in the literature as important for dialog (wells, 2006; international standards organization, 2012; riggenbach, 1991; couperkuhlen and selting, 1996a; sidnell, 2011; clark, 1996; szczepek reed, 2010), and pca-based studies of other corpora have revealed similar constructions and functions (ward and vega, 2012; ward, 2014), so this list does seem likely to be of some generality. 5. non-native dialog data for this study we chose to work with advanced non-native speakers, inspired by reports of those who, despite years of immersion, still have weak prosodic skills (zimmerer et al., 2014), and from personal observation of friends and family members for which this is the case. in this we diverge 12 non-native prosody 1 17% lo: speaker a speaking, speaker b silent §3.4, §6, §7.2, §8 hi: speaker b speaking, speaker a silent §3.4, §6, §7.2, §8 2 8% lo: both speakers silent hi: both speakers talking together or laughing together §3.5, §6, §7.2 3 4% lo: b yields the turn and a takes the turn §3.5, §6, §7.2 hi: a yields the turn and b takes the turn §4, §6, §7.2 4 3% lo: b makes a small contribution during a’s turn §6 hi: a makes a small contribution during b’s turn §6 5 3% lo: high involvement §7.1, §8 hi: low involvement §7.1, §8 6 3% lo: pivot point of a rhetorical structure, etc. hi: pausing while thinking how to continue §6 7 3% lo: bidding for empathy, inviting an inference §7.1, §8 hi: giving factual information, explaining something or some actions §7.1, §8 8 2% lo: a confident, speaking with authority or based on personal experience hi: b confident, speaking with authority or based on personal experience 9 2% lo: a disfluent, hesitant, or silent; b silent or fluent hi: b disfluent, hesitant, or silent; a silent or fluent 10 2% lo: speakers agreeing or aligning §3.5, §6, §8 hi: speakers disagreeing or diverging §4 11 2% lo: b yields floor to a hi: a yields floor to b 12 1% lo: b interpolates a short comment §7.2 hi: a interpolates a short comment §7.2 13 1% lo: lack of new information hi: knowledge asymmetry between speakers 14 1% lo: personal-situation comments hi: complaints about third parties 15 1% lo: being positive about one’s own prospects or a past experience hi: negative feeling about something/someone distant 16 1% lo: displeasure, annoyance §3.5 hi: amusement, positive evaluation of something/someone 18 1% lo: a reveals downside or b reveals silver lining §6 hi: b reveals downside or a reveals silver lining §6 21 1% lo: memory recall §6 hi: rushed turn grab or hold §6 table 3: the top sixteen prosodic dimensions in the reference corpus, plus two more. the second field is the amount of variance explained by the dimension. the third field summarizes our interpretations of the dimension when negatively or positively present, that is, the “lo-side” and “hi-side” constructions. the fourth field indexes further discussion. 13 ward and gallardo from the common practice of studying non-native prosody using data from learners still in language classes (van engen et al., 2010). this section summarizes some of the important properties of our data sets; the details appear elsewhere (ward and gallardo, 2015). we chose learners whose native language was spanish, based on the ease of recruiting them. the segmental, lexical, and syntactic aspects of spanish prosody are known to differ from english, as are the expressions of some pragmatic functions, including questions, back-channeling, complaining, and expressing probability and usuality (bowen, 1956; farias, 2013; hualde, 2005; berry, 1994; ramirez verdugo, 2005; rivera and ward, 2006; rao, 2013a; santiago and delais-roussarie, 2015; de la mota et al., 2010). spanish also expresses some pragmatic functions less with prosody than with word order, discourse particles, or gesture (borras-comes et al., 2014; ortega-llebaria and colantoni, 2014). accordingly it seemed likely that there would be differences in dialog prosody also. we recruited among friends and acquaintances in our department; as a result, all participants had completed at least one semester of college in the united states. participants gave informed consent, and we compensated them with $15 for participating. each non-native speaker was recorded in dialog with an english-monolingual native-speaker partner. later, after listening to the recordings, we decided to exclude four speakers who seemed to have almost-native conversation skills. thus we obtained 9 conversations, including 6 different non-native speakers. these speakers all had strong vocabulary and good fluency, but all were noticeably non-native in pronunciation. all had grown up in northern mexico. when recording we did not ask the speakers to do anything more specific than talk to each other. their conversations were spontaneous and varied widely in topic. while there are advantages to using conversation data based on scripted or role-play interactions, spontaneous conversations may more closely approximate real-world interaction. while producing appropriate prosody in monolog or scripted dialog is, in essence, “merely” a question of choosing the appropriate form and applying it to a sentence, producing appropriate prosody in dialog is a much greater challenge. realization of each construction requires using multiple prosodic features in specific temporal configurations, multiple constructions must often be simultaneously realized, and all this must be done under the time pressure of choosing words and listening to and coordinating with the dialog partner. as noted above, we selected the non-native speakers for the corpus based on our perceptions of awkwardness with english, without explicit consideration of prosodic behaviors. however we did a post-hoc examination to see whether their prosody also appeared non-native. casual listening showed that it was, most saliently in having: a tendency to syllable-timing rather than stress timing, unusual patterns of utterance-final lengthening or lack of lengthening, and misplaced stresses and accents. there seemed to be other differences but we did not attempt to categorize them, preferring to move directly to the model-based analyses. we must here note two potential issues with this data. one is that, since each pair of speakers includes a native speaker, and since each pattern involves behavior by both speakers, it is possible that some observed differences could be due to the native speakers behaving differently when interacting with non-native speakers, rather than to differences in the behavior of the non-native speakers themselves. however we saw only rare evidence for this, and only for one speaker pair, so this is probably not a major problem. another potential issue is that, statistically, a pattern may be detected as often used, when in fact this may be mostly due to times when the native speaker perfectly executes his side of the pattern, with little or no support from the non-native speaker. thus our method may understate the non-native differences. 14 non-native prosody for comparisons we used three other data sets. two we recorded ourselves: one of monolingual native english speakers talking with other native speakers, and one of spanish speakers speaking together in spanish. both of these collections included many speakers from the primary collection. all were recorded in the same environment with the same equipment. we also used native speakers from the well-known switchboard corpus (godfrey et al., 1992), specifically, 7 randomly-selected dialogs (14 speakers, 35 minutes total). although these are also dyadic conversations in american english, in these conversations the participants were strangers, they were generally much older, they spoke by telephone, and they started with suggested topics, such as crime and childcare, although most of the conversations rapidly moved on to other topics. table 4 summarizes the five data sets used. use language speakers used social speech building the model, reference english english-native 60 min. switchboard exploring distributions english english-native 35 min. utep non-native english finding skill deficits english spanish-native 90 min. utep native english comparing difference magnitudes english english-native 70 min. utep spanish exploring l1 influences spanish spanish-native 90 min. table 4: summary of the data sets 6. distribution differences we expected that the dialog-prosody deficits of non-native speakers could be associated with specific constructions. logically, following mennen’s (2015) categories, these deficits could be of four kinds: not knowing a construction, using it too often or not enough, using it for the wrong functions, or not producing it accurately. in our first approach we looked for differences of the first two kinds by comparing distributions on the various dimensions. if non-native speakers are using a construction with the same frequency as natives, the distributions on the associated dimension should be the same; conversely, if the distributions differ their prosody may be significantly different. figure 5 overviews the workflow. the first stages are fully automatic: given the dimensions, the prosody in the immediate context of every point can be represented as the sum of the contributions of all the dimensions active at that time. thus we applied the loadings discovered by pca to samples taken every 10 milliseconds throughout the data, simply by taking the dot product. to test this method, we first applied it to another set of english data to see whether it would find differences that made sense. specifically we used the switchboard data. on dimension 2 there was a large distribution difference relative to the reference data, as seen in figure 6: the switchboard speakers exhibit fewer high values on this dimension. referring to the interpretation in table 3, this indicates that in this data less often were both participants simultaneously talking or laughing together. by listening to some conversations we readily confirmed that this was in fact the case. this is unsurprising, given that turn-taking is generally more formal in telephone conversations and in conversations between strangers. for reasons of space we do not show distributions for the other dimensions. instead summary statistics are given in table 5, where columns 2 and 3 show the means and standard deviations of 15 ward and gallardo feature-vector description feature-vector description feature extractors six seconds of context 1,080,000 non-native audio samples feature-vector description feature-vector description … conversion (dot product) principal components … … … … … … … principalcomponents description principalcomponents description principalcomponents description principalcomponents description … aggregation statistics [from figure 1] re-represented samples interpretation 32 prosodic constructions … figure 5: workflow for comparing distributions. figure 6: distribution of values on dimension 2 for the reference data and the switchboard data. switchboard speakers’ uses of the top 16 dimensions, plus 2 more. the mean for the reference set is zero on each dimension, due to normalization. both the means and the standard deviations shown have been normalized by (divided by) by the standard deviation of the same dimension in the reference set. thus in columns 2, 4, and 6 the units for the means are standard deviations, with negative values where the switchboard speakers tended to be lower on that dimension and positive values when higher. in columns 3, 5, and 7, for the standard deviations, values less than 1 mean 16 non-native prosody switchboard native non-native dimension mean std. dev. mean std. dev. mean std. dev. 1 -0.00 1.22 -0.00 1.05 -0.11 0.97 2 -0.31 0.53 -0.33 0.76 -0.45 0.75 3 0.00 0.91 0.00 0.91 0.01 0.86 4 0.00 0.87 -0.00 0.85 -0.05 0.83 5 -0.12 0.74 -0.01 0.94 0.01 0.91 6 -0.17 0.65 0.10 0.86 0.14 0.84 7 -0.06 0.86 -0.06 0.97 -0.03 0.97 8 -0.00 0.81 0.00 0.86 0.03 0.86 9 0.00 0.85 0.00 0.82 -0.00 0.81 10 -0.43 0.97 -0.10 1.01 -0.21 0.99 11 -0.00 0.94 -0.00 0.91 -0.01 0.93 12 -0.00 0.92 -0.00 0.89 -0.00 0.89 13 0.11 0.71 0.03 0.91 0.03 0.90 14 -0.33 0.79 -0.07 0.94 -0.09 0.95 15 0.37 0.87 0.07 0.97 0.07 0.98 16 0.00 0.88 -0.00 0.92 -0.02 0.92 18 -0.00 1.02 -0.00 0.97 0.11 0.91 21 0.00 0.95 0.00 0.88 -0.13 0.87 table 5: statistics on dimension use in three data collections, relative to the reference set. for the reference set, not shown, the means are all 0 and the standard deviations all 1, due to normalization. bolding in columns 2 and 3 marks the largest differences between switchboard and the reference set, and in columns 6 and 7 the largest differences between the non-native and the native datasets. the switchboard speakers had narrower distributions than the reference speakers, and greater values wider distributions. among the differences, the largest was for dimension 10, for which the switchboard speakers tended to the negative side. prosodically, the 10-lo construction involves quieter-than-average utterances, with gradually increasing pitch and a moment of creakier, faster speech with expanded pitch range. pragmatically this is associated with alignment and agreement, as seen in example 2. in this fragment g is agreeing that the class r plans to take is interesting, starting quietly and then building, culminating with an enthusiastic huh. example 2 (eng001@419.2s) r: but i still really want to do it, because it sounds interesting g: yeah, that does sound interesting, r: and (trails off) g: huh! that’d be interesting ... thus the distribution of values on dimension 10 suggests that switchboard conversations exhibit more alignment and agreement than the reference dialogs. this was again easily confirmed by some 17 ward and gallardo listening. it is also easy to understand: strangers who have no desired outcomes beyond having a pleasant conversation tend to find things they can agree on. thus it seems that comparing distributions can reveal differences in prosodic behavior that reflect real differences in dialog activities and interaction styles. having verified the approach, we next applied the same method to the non-native data, and, for comparison, to the monolingual native data. table 5 shows the results. columns 4 through 7 are the means and standard deviations on each dimension for both sets. most relevant are dimensions where the non-native behaviors differ not only from the reference data, but also from our native-speaker comparison data. while we expect variation between any two random sets of speakers, if the non-natives differ from both the native data and the reference data, that indicates a real difference. however, such differences were slight. indeed, contrary to expectation, the non-native means were mostly closer to the native means than the switchboard speakers’ means were. we examined the differences statistically, using unmatched, two-tailed heteroskedastic t-tests with bonferroni corrections, taking as independent samples the means of each speaker’s values, and found no statistically significant differences. nevertheless there appear to be tendencies; the rest of this section discusses the dimensions with larger differences. for dimension 1 the averages were noticeably different, indicating that the non-natives speakers talk rather more than the native speakers. this is easy to confirm, by listening, and easy to understand: often language learners take more words and more attempts to convey what they want to say. for dimension 2 also, the averages were different, indicating that the non-native speakers tended to have less overlapped speech; something that was again easy to confirm by listening. for dimension 3, although there was only a tiny difference in means, the non-native data exhibited narrower variation. this dearth of extreme values on this dimension indicates that the non-native speakers had fewer prototypical turn takes and turn yields. from this statistic alone we cannot tell exactly how they differed, but when we listened to the data, we did notice a tendency to have longer gaps between speaker changes. for dimension 4 the non-native speakers averaged slightly lower. dimension 4-hi involved the speaker making an interleaved short contribution in the midst of speech from the interlocutor. these short contributions were usually a backchannel, a short question, laughter, or a suggestion of a word that the other was looking for. (the interlocutor’s prosody in the vicinity involved peaks in intensity, creakiness, and enunciation or speaking rate about two seconds apart, often with a gap in between.) this fact that the non-native average was lower indicates that the non-native speakers less commonly produced small utterances precisely interleaved in the other’s turn. again, listening confirmed this tendency. for dimension 6 the non-native data averaged slightly higher. the prosody of dimension 6-hi involves a region of low intensity, generally a pause, surrounded by two regions of high intensity, wide pitch range, and creakiness; thus this indicates a tendency for the non-native speakers to pause more often to think. for dimension 10 the non-native speakers averaged lower. as noted above, this indicates a greater tendency to align or agree with the other person. example 3 is an exemplar of this prosodic pattern: the words squeeze you are fast, creaky and in wide pitch range, and they lead into a highpitched laugh. this example requires some explanation. a and b have been talking about pets, and b has related a time when he tried to use a dog as a pillow. he re-enacts the situation, speaking as he might have addressed the dog in apology. a shows empathy by herself acting out he might have felt, in the last clause herself addressing an imagined dog. 18 non-native prosody example 3 (nn007@126.8s) b: i’m sorry but you’re so fluffy a: [laughs] i just want to use it as a pillow, and squeeze you [laughs] while our focus has been on the top 16 dimensions, we ran the statistics down further, and noted large differences for some others. space permits discussion of only two. for dimension 18 the non-natives speakers had fewer low values, indicating fewer positive-tonegative perspective shifts. examples of these included (my favorite class is) programming languages, because it’s the only hope i have (to get an a), with the last clause wry in tone, and the material’s really easy, so a lot of people, like stop paying attention to the class, and that’s what i did (and that’s why i failed it last time). prosodically the 18-low construction involves a region of high pitch and high pitch range, followed directly by a region of low pitch. (it is not that the non-native speakers were unfamiliar with wry humor; in fact, when we applied pca to the spanish data wry humor showed up quite high, in dimension 10. however the spanish form is very different, involving a few seconds of creaky voice and a fast speaking rate, including a short period of increased pitch range in the vicinity of a short pause.) for dimension 21 the non-native speakers averaged lower. dimension 21-lo was associated with filler production while recalling something from memory, where the filler was flat in pitch and initially creaky. 21-hi was associated with a rushed start to grab or hold a turn, with wide pitch range, and often followed by a reformulation. fillers were indeed common in the non-native utterances, and aggressive turn starts rare. thus this method led us to find interesting differences, of various types. some appear to reflect actual prosodic deficits (dimension 18 and, as discussed below, dimension 3). others appear to reflect processing limitations (1, 3, 4, 6, 21), which is plausible since learners may be slow to comprehend and/or need more time to create fluent utterances (wiberg, 2003). yet others seem more to involve cultural factors (2, 10, 21), which are known to often transcend considerations of what language is currently being spoken (tannen, 1989). for example, the greater use of construction 10lo (alignment) can be related to well-attested norms of mexican culture (condon, 1985). it seems likely that the non-native speakers were behaving as they thought appropriate, rather than trying to behave like native speakers but failing due to a prosodic skill deficit. similarly, the reduced use of dimension 21-high (aggressive/rushed turn holding), could be explained as a choice not to use (or not to acquire) a behavior that seems rude in many contexts. however distributions cannot tell the whole story: there are gross difference in prosodic behavior that they do not show. this point was driven home for us when we looked at the distributions of our spanish data on the same reference dimensions. to our surprise, there were only minor differences in the distributions. we speculate that this is because the space of possible prosodic variation is limited, and so each language tends to use the entire space, although for different purposes. be that as it may, it was clear that we needed another way to look at the data. 7. dimension differences mennen’s (2015) categories include two other kinds of differences, relating to realization and to semantics. we therefore set out to look for differences in the details of how non-native speakers produce specific prosodic constructions and of how they map pragmatic functions to constructions. to do this we developed a second method, that of comparing the patterns of non-native behavior 19 ward and gallardo to the native patterns, rather than directly comparing non-native data to the native patterns. thus, in this approach the first step is to characterize the prosodic behaviors of the non-native speakers in their own terms. this is done, again, using pca. figure 7 shows the concept. we assume that, if the non-native behavior is similar to that of the native behavior in some respect, then the relevant pattern of native behavior will be well matched by some non-native pattern. conversely, we assume that native patterns that lack a counterpart will be behaviors that the non-native speakers have not mastered. to reduce the extent of “muddying” due to the behaviors of the native-speaker partners, for the pca we ensured that the non-natives were always in the a track. similarity comparisons reference dimensions (constructions) nonnative data non-native dimensions (constructions) interpretation [from figure 1] … … figure 7: workflow for comparing construction inventories. since we use patterns based on pca-derived dimensions, finding counterparts is easy. each dimension is defined by its loadings on the raw features, so two dimensions are similar to the extent that their loadings are similar. we use a simple operator for this, the cosine: by this metric, two dimensions are similar if, for every feature, when one dimension’s loading on that feature is strongly positive then the other is also, and conversely if negative. table 6 shows the result: the cosines between the top five reference dimensions and the top five non-native dimensions. it is clear that the top four dimensions do have counterparts, but for reference dimension 5 there is no strongly-similar non-native dimension. to identify other reference dimensions without a clear non-native counterpart, we computed table 7, showing the cosines of the best matching dimensions from both the non-native data and the native-data comparison set. as might be expected, this second population of native speakers does not show exactly the same prosodic behavior as the reference population: column two is never 1.00. at the same time, also unsurprisingly, the natives are almost always closer to the reference than are the non-natives. we first discuss major differences then minor differences, as they seem to reflect different kinds of issues. 20 non-native prosody 7.1 differences in use table 7 indicates that the non-native speakers lack patterns corresponding to the reference dimensions 5 and 7. in this subsection we describe the corresponding patterns in native-speaker dialogs, then examine how the non-native speakers differed. the low-side pattern of dimension 5 involves about a half-second of increased intensity, starting with high pitch and ending creaky. at times in the native data when dimension 5 was most strongly negative we frequently saw discourse markers, such as yeah, ah, ooh, and but, being used assertively. in general, when dimension 5 is negative the speaker is showing involvement. in the non-native data, we found that only some speakers used this prosodic pattern for this function. (as always, we inferred their intended functions from their words and their behavior in the wider context.) one appeared not to use it at all; that is, there were no times where her speech was highly negative on this dimension. another used this prosody frequently on question-initial so, making her questions sound incongruously aggressive. the high-side pattern of dimension 5 involves low pitch over several seconds, and within that a short region that is creaky and even more strongly low in pitch. this frequently occurs with words like and, um, like, and you know, for example when a speaker is musing about his future plans. in general, at times when dimension 5 was high, the speaker had low involvement in the topic and/or the dialog itself. in the non-native data, while some speakers used this prosodic pattern for low involvement, they also used it in other contexts, for example in offering help, in marking disfluencies, and in greeting. as a second way to help understand what was going on, we examined the functions of the non-native dimensions which were (somewhat) similar to this dimension, namely 4, 5, and 6. their functions included the co-construction of utterances, floor holding, backchanneling, and marking the point of a story, but not involvement. for dimension 7, the low-side pattern involves pitch strongly high and with a region with a fairly slow drop in intensity, rate and creakiness over about 1.5 seconds. native speakers frequently use this trailing off or “intensity-fade” pattern to solicit empathy, and sometimes also when leaving something unsaid and inviting the listener to infer it. there were many cases where the non-native speakers were using essentially this same pattern for essentially the same function: soliciting empathy, understanding, or an inference. thus there was no apparent deficit on the negative side. the high-side pattern of dimension 7 involves strongly low pitch over about 3 seconds. when native speakers used this they were generally explaining something, usually something factual, such as a software project’s architecture, or how a study group had arranged to turn in a joint assignment. nn 1 nn 2 nn 3 nn 4 nn 5 ref 1 .97 –.07 –.00 –.01 .07 ref 2 .11 .90 –.00 .09 –.05 ref 3 –.01 .01 –.95 –.10 –.09 ref 4 .04 .03 .08 –.43 –.80 ref 5 .03 –.14 .05 .61 –.44 table 6: similarities between reference dimensions and non-native dimensions. in each row the highest value is bolded. 21 ward and gallardo cosine of the most reference strongly-related dimension dimension native non-native ratio 1 .99 .97 .98 2 .90 .90 1.00 3 .92 .95 1.03 4 .92 .80 .87 5 .80 .61 .76 6 .64 .63 .98 7 .83 .57 .69 8 .73 .63 .86 9 .74 .69 .93 10 .67 .64 .96 11 .87 .83 .95 12 .62 .79 1.27 13 .74 .72 .95 14 .75 .64 .85 15 .74 .68 .92 16 .60 .56 .93 table 7: for each of the reference dimensions, the cosine of the best-matching dimension found for the other data sets. listening to places in the non-native data high on dimension 7, we found no cases where a nonnative speaker used a long region of low pitch in the course of explaining things. it is not that they never explained technical things; rather they tended to do so in an interactive style, including lots of pitch variation, for example on interleaved questions to check that the listener was following. some non-native speakers didn’t use the long low-pitch region pattern at all; others did, however not for explaining but rather when talking about something personal, such as family background, likes and dislikes, habits, or intentions. overall, this suggests that the lack of a non-native dimension corresponding to a reference dimension really does indicate a weakness with their prosodic expression of the corresponding functions 7.2 differences in realization the method is effective not only for identifying gaps in learners’ skills, but also for detecting where they are using essentially the same constructions in the same ways, but with small differences. although small — after all, as table 7 shows, for many dimensions the non-native differences are minor, reflecting their advanced level — the differences are revealing. in this subsection we consider just the three top dimensions. non-native dimension 1 is very similar in loadings to reference dimension 1, as seen by the 0.97 cosine, and in the recordings they obviously serve the same function: positive when the left 22 non-native prosody speaker has the floor, and negative when the right speaker does. however the loadings are not identical: in particular for some speaking rate features the loadings are higher for the native speakers. this suggests that native speakers talk faster when they have the floor, but the non-native speakers have no such tendency. (to investigate whether this might be due to transfer from spanish, we also ran pca on the spanish data: the corresponding dimension there indeed lacked a tendency for the person holding the floor to speak faster than average.) reference dimension 2 and non-native dimension 2 also differ slightly in loadings, indicating that on the high side, during regions of overlapped speech by the two speakers, native speakers tend to have a fast speaking rate, whereas non-native speakers tend to have a higher pitch. for reference dimension 3 and non-native dimension 3, turn hand-offs, the major difference is that the native speakers tend to speak faster at turn starts, but for non-native speakers this tendency is much weaker. the reality of these differences, suggested by the loadings, was readily confirmed by listening to some of the non-native data. thus this method also seems valid: the differences that it uncovers are real ones. however we must note that this method does not appear to be reliable for less-frequent constructions. looking again at table 7, it is clear that the lower-ranked reference dimensions tend to align less well to the dimensions of the other data sets. this likely reflects a lack of robustness to extraneous sources of variability. this can be seen in the results for reference dimension 12. this pattern involves one speaker interleaving a short comment during a brief pause by the other, often showing alignment, appreciation, or empathy. while the non-native speakers appear to be doing this fairly successfully (a cosine of .79), the comparison native speakers appear to lack this pattern (highest cosine of .62). looking at the data, we believe that this reflects not differences in prosodic competence but rather the fact that the comparison set of native speakers tended to talk more about technical topics than about personal ones, giving them fewer opportunities to be supportive. 8. summary and prospects we have presented new methods for finding prosodic patterns that are under-used or variant in form in non-native speech. these methods work automatically from data, once an initial analysis of the native-speaker patterns has been done. software supporting this workflow is available as opensource (ward, 2015). in these methods we use pca to reduce the high-dimensional space of all prosodic features to a lower-dimensional space in which the patterns of the two speaker populations can be meaningfully compared. this improves on previous methods in going beyond mere featurelevel differences, while avoiding some of the biases inherent in top-down theory-driven analyses. we applied these methods to discover some important prosodic constructions of english dialog, and used these to discover ways in which spanish-native learners differed, including in the prosody of turn-taking, showing involvement, and explaining. it is interesting to speculate about how these prosodic differences may relate to perceived cultural differences. american businessmen often perceive mexicans, it has been said, as being leisurely and disinclined to rush, and as tending to bring personal and emotional considerations into business discussions, rather than rationally sticking to facts (condon, 1985). we earlier noted that the non-native speakers did not tend to pick up the pace of speaking even when they have the floor (dimension 1), do not tend to mark involvement (5), tend to agree more (10), and may not mark factual, explanatory information as different from personally-relevant information (7). thus, 23 ward and gallardo while there may be real cultural differences, these cross-cultural perceptions may also reflect mere prosodic-behavior differences. we note that these results must be interpreted cautiously. here our aim was to explore new methods and previously unremarked phenomena, not to definitively establish any fact or settle any issue. one limitation of this study was that the speakers were not matched across the conditions, so some of the differences found may reflect different personality types or other uncontrolled differences across the datasets. another limitation is that the only verifications done involved looking at the data to check the reasonableness of what the automatic methods suggested. a more methodologically-sound procedure would be to first obtain an independently-created and validated listing of non-native dialog-prosody deficits, something that unfortunately does not exist at present. a third limitation is that, in places where the method required subjective judgments, we used only our own. judgments from more observers would be required to have full confidence in the interpretations given. there are many open scientific questions relating to our methods and our findings. one is how to find a better feature set, that is, one that leads to results that accurately reflect what is perceptually most salient and communicatively most important. another set of questions involves the details of each construction. here we relied heavily on automated methods, seeking a big-picture inventory and a broad-brushstroke understanding. like many other big-data methods (swanson and charniak, 2014) this was efficient, but has its limits. further examination using more sensitive methods could better tie these construction-based descriptions to those developed within other theoretical frameworks and improve the accuracy of the descriptions. another scientific question is which differences in prosody actually matter. on the one hand, differences may “make the speaker sound strange, typical of their origin, boring or annoying . . . [but] . . . not cause much of an actual breakdown in communication” (wells, 2014). on the other hand, such differences may affect perceptions and dialog outcomes (tannen, 2005; curhan and pentland, 2007). identifying which differences matter will require further study involving broader consideration of interpersonal and social factors. despite these limitations and open questions, this work may be useful in various ways. these findings may be useful for teachers. second language teaching, when it treats prosody at all, usually focuses on just a few well-understood aspects (diepenbroek and derwing, 2013; busa, 2012). here we have identified some frequently-used dialog-related aspects of prosody that are almost certainly of value to learners, but are never taught, nor acquired naturally even after years of immersion. future work should refine our findings into descriptions that will be useful for learners and teachers. future work should also develop better techniques for teaching the prosody related to interaction patterns in dialog, as this involves special challenges (ward et al., 2007; betz and huth, 2014). the methods developed here may be useful for rating speakers. assessment is important for gatekeeping purposes — including placing learners into the appropriate level of instruction – and recently there is growing interest in ways to rate not only language knowledge but also interaction skills and communicative effectiveness. this is true not only for non-native speakers (mitchell et al., 2014; litman et al., 2016), but also for native speakers who wish to become more effective in interviews or other unfamiliar situations (hoque et al., 2013). finally, beyond assessment of overall competence, these methods may be useful for pinpointing the prosodic deficits of individual speakers. among other challenges, this will require ways to obtain reliable results with less data. 24 non-native prosody acknowledgments we thank the participants who let us record their conversations, and richard ogden, david novick, david devault, juergen trouvain, francisco torreira, and shizuka nakamura for discussion. this work was supported in part by the us national science foundation by means of research experience for undergraduates supplements to iis 191-4868 and iis 144-9093, and in part by the fulbright program. references berit aronsson and lars fant. boundary tones in non-native speech: the transfer of pragmatics strategies from l1 swedish into l2 spanish. intercultural pragmatics, 11:159–198, 2014. amalia arvaniti. the representation of intonation. in marc van oostendorp, colin j. ewen, elizabeth v. hume, and keren rice, editors, the blackwell companion to phonology. wiley, 2011. anne-marie barraja-rohan. using conversation analysis in the second language classroom to teach interactional competence. language teaching research, 15:479–507, 2011. anne berry. spanish and american turn-taking styles: a comparative study. in l. f. boulton, editor, pragmatics and language learning monograph series, volume 5, pages 180–190. university of illinois, urbana-champaign: division of english as an international language, 1994. emma m. betz and thorsten huth. beyond grammar: teaching interaction in the german language classroom. die unterrichtspraxis/teaching german, 47(2):140–163, 2014. joan borras-comes, constantijn kaland, pilar prieto, and marc swerts. audiovisual correlates of interrogativity: a comparative analysis of catalan and dutch. journal of nonverbal behavior, 38:53–66, 2014. j. donald bowen. a comparison of the intonation patterns of english and spanish. hispania, 39: 30–35, 1956. harry bunt. multifunctionality in dialogue. computer speech and language, 25:222–245, 2011. maria grazia busa. the role of prosody in pronunciation teaching: a growing appreciation. in maria grazia busa and antonio stella, editors, methodological perspectives on second language prosody, pages 101–105, 2012. gao-peng chen, gérard bailly, qing-feng liu, and ren-hua wang. a superposed prosodic model for chinese text-to-speech synthesis. in international symposium on chinese spoken language processing, pages 177–180. ieee, 2004. zi-he chen, yuan-fu liao, and yau-tarng juang. prosody modeling and eigen-prosody analysis for robust speaker recognition. icassp, 2005. herbert h. clark. using language. cambridge university press, 1996. john c. condon. good neighbors: communicating with the mexicans. intercultural press, 1985. 25 ward and gallardo elizabeth couper-kuhlen. an introduction to english prosody. edward arnold, 1986. elizabeth couper-kuhlen and margret selting. prosody in conversation: interactional studies. cambridge university press, 1996a. elizabeth couper-kuhlen and margret selting. towards and interactional perspective on prosody and a prosodic perspective on interaction. in elizabeth couper-kuhlen and margret selting, editors, prosody in conversation: interactional studies. cambridge university press, 1996b. jared r. curhan and alex pentland. thin slices of negotiation: predicting outcomes from conversational dynamics within the first 5 minutes. journal of applied psychology, 92:802, 2007. anne cutler and d. robert ladd. models and measurements in the study of prosody. in anne cutler and d. robert ladd, editors, prosody: models and measurements. springer, 1983. carme de la mota, pedro martin butragueno, and pilar prieto. mexican spanish intonation. in p. prieto and p. roseano, editors, transcription of intonation of the spanish language, pages 319–350. lincom europa, 2010. lori g. diepenbroek and tracey m. derwing. to what extent do popular esl textbooks incorporate oral fluency and pragmatic development? tesl canada journal, 30:1–20, 2013. maria estelles-arguedas. expressing evidentiality through prosody? prosodic voicing in reported speech in spanish colloquial conversations. journal of pragmatics, 85:138–154, 2015. maria gabriela valenzuela farias. a comparative analysis of intonation between spanish and english speakers in tag questions, wh-questions, inverted questions, and repetition questions. revista brasileira de linguı́stica aplicada, 13:1061–1083, 2013. john j. godfrey, edward c. holliman, and jane mcdaniel. switchboard: telephone speech corpus for research and development. in proceedings of icassp, pages 517–520, 1992. adele e. goldberg. constructionist approaches. in thomas hoffman and graeme trousdale, editors, the oxford handbook of construction grammar, pages 15–31. oxford university press, 2013. agustin gravano and julia hirschberg. turn-taking cues in task-oriented dialogue. computer speech and language, 25:601–634, 2011. michele gubian, lou boves, and francesco cangemi. joint analysis of f0 and speech rate with functional data analysis. in icassp, pages 4972–4975, 2011. john j. gumperz. discourse strategies. cambridge university press, 1982. kieu-phuong ha, samuel ebner, and martine grice. speech prosody and possible misunderstandings in intercultural talk: a study of listener behaviour in standard vietnamese and german dialogues. in speech prosody, pages 801–805, 2016. nancy hedberg, juan m. sosa, and lorna fadden. the intonation of contradictions in american english. in prosody and pragmatics conference, 2003. 26 non-native prosody nancy hedberg, juan m. sosa, and emrah gorgulu. the meaning of intonation in yes-no questions in american english. corpus linguistics and linguistic theory, to appear, 2014. john heritage. epistemics in action: action formation and territories of knowledge. research on language & social interaction, 45(1):1–29, 2012. julia hirschberg. communication and prosody: functional aspects of prosody. speech communication, 36:31–43, 2002. mohammed ehsan hoque, matthieu courgeon, jean-claude martin, bilge mutlu, and rosalind w. picard. mach: my automated conversation coach. in proceedings of the 2013 acm international joint conference on pervasive and ubiquitous computing, pages 697–706, 2013. jose ignacio hualde. the sounds of spanish, chapter 14: intonation. cambridge university press, 2005. international standards organization. language resource management – semantic annotation framework (semaf) – part 2: dialogue acts. iso 24618-2:2012, 2012. shuichi itahashi and kimihito tanaka. a method of classification among japanese dialects. in eurospeech, 1993. oliver jokisch, tristan langenberg, and gabor pinter. intonation-based classification of language proficiency using fda. in speech prosody, 2014. evia kainada and angelos lengeris. native language influences on the production of secondlanguage prosody. journal of the international phonetic association, 45:269–287, 2015. rose thomas kalathottukaren, suzanne c. purdy, and elaine ballard. behavioral measures to evaluate prosodic skills: a review of assessment tools for children and adults. contemporary issues in communication science and disorders, 42:138–154, 2015. okim kang. relative salience of suprasegmental features on judgments of l2 comprehensibility and accentedness. system, 38:301–315, 2010. d. robert ladd. stylized intonation. language, pages 517–540, 1978. d. robert ladd. intonational phonology. cambridge university press, 1996. catherine lai. response types and the prosody of declaratives. in speech prosody, 2012. mark liberman and ivan sag. prosodic form and discourse function. in papers from tenth regional meeting, chicago linguistic society, pages 402–427, 1974. diane litman, steve young, mark gales, kate knill, karen ottewell, rogier van dalen, and david vandyke. towards using conversations with spoken dialogue systems in the automated assessment of non-native speakers of english. in 17th annual meeting of the special interest group on discourse and dialogue, pages 270–275, 2016. ineke mennen. beyond segments: towards a l2 intonation learning theory. in elisabeth delaisroussairie, mathieu avanzi, and sophie herment, editors, prosody and language in contact, pages 171–188. springer, 2015. 27 ward and gallardo christopher m. mitchell, keelan evanini, and klaus zechner. a trialogue-based spoken dialogue system for assessment of english language learners. in proceedings of the international workshop on spoken dialogue systems (iwsds), 2014. oliver niebuhr. resistance is futile: the intonation between continuation rise and calling contour in german. in interspeech, pages 132–136, 2014. richard ogden. prosodies in conversation. in oliver niebuhr, editor, understanding prosody: the role of context, function, and communication, pages 201–217. de gruyter, 2012. richard a. ogden. linguistic resources for complaints in conversation. in international congress of the phonetic sciences, pages 1321–1324, 2007. marta ortega-llebaria and laura colantoni. l2 english intonation: relations between formmeaning associations, access to meaning, and l1 transfer. studies in second language acquisition, 36:331–353, 2014. benjamin parrell, sungbok lee, and dani byrd. evaluation of prosodic juncture strength using functional data analysis. journal of phonetics, 41(6):442–452, 2013. caterina petrone and oliver niebuhr. on the intonation of german intonation questions: the role of the prenuclear region. language and speech, 57:108–146, 2013. pilar prieto. intonational meaning. wiley interdisciplinary reviews: cognitive science, 6:371–381, 2015. dolores ramirez verdugo. the nature and patterning of native and non-native intonation in the expression of certainty and uncertainty: pragmatic effects. journal of pragmatics, 37:2086– 2115, 2005. dolores ramirez verdugo. a study of intonation awareness and learning in non-native speakers of english. language awareness, 15:141–159, 2006. ma. dolores ramirez verdugo. non-native interlanguage intonation systems: a study based on a computerized corpus of spanish learners of english. icame journal, 26:115–132, 2003. rajiv rao. intonational variation in third party complaints in spanish. journal of speech sciences, 3:141–168, 2013a. rajiv rao. prosodic consequences of sarcasm versus sincerity in mexican spanish. concentric: studies in linguistics, 39(2):33–59, 2013b. uwe d. reichel. linking bottom-up intonation stylization to discourse structure. computer speech and language, 28:1340–1365, 2014. heidi riggenbach. toward an understanding of fluency: a microanalysis of nonnative speaker conversations. discourse processes, 14:423–441, 1991. anais g. rivera and nigel ward. prosodic cues that lead to back-channel feedback in northern mexican spanish. hdls-7 conference, high desert linguistics society, university of new mexico, 2006. 28 non-native prosody jesus romero-trillo. pragmatics and prosody in english language teaching. springer science & business media, 2012. fabian santiago and elisabeth delais-roussarie. the acquisition of question intonation by mexican spanish learners of french. in prosody and language in contact, pages 243–270. springer, 2015. bjorn schuller. voice and speech analysis in search of states and traits. in albert ali salah and theo gevers, editors, computer analysis of human behavior, pages 227–253. springer, 2011. elizabeth e. shriberg and andreas stolcke. direct modeling of prosody: an overview of applications in automatic speech processing. in proceedings of the international conference on speech prosody, pages 575–582, 2004. jack sidnell. conversation analysis: an introduction. john wiley & sons, 2011. ben swanson and eugene charniak. data driven language transfer hypothesis. in 14th conference of the european chapter of the association for computational linguistics, pages 169–173, 2014. marc swerts and s. zerbian. intonational differences between l1 and l2 english in south africa. phonetica, 67:127–146, 2010. beatrice szczepek reed. prosodic orientation in english conversation. palgrave, 2006. beatrice szczepek reed. analysing conversation: an introduction to prosody. palgrave macmillan, 2010. beatrice szczepek reed. prosody in conversation: implications for teaching english pronunciation. in j. romero-trillo, editor, pragmatics and prosody in english language teaching, pages 147– 168. springer, 2012. deborah tannen. that’s not what i meant! how conversational style makes or breaks relationships. ballantine, 1989. deborah tannen. interactional sociolinguistics as a resource for intercultural pragmatics. intercultural pragmatics, 2:205–208, 2005. juhani toivanen. tone choice in the english intonation of proficient non-native speakers. in in proceedings of fonetik, phonum 9, pages 165–168, 2003. jürgen trouvain and ulrike gut. non-native prosody: phonetic description and teaching practice. walter de gruyter, 2007. giuseppina turco, christine dimroth, and bettina braun. prosodic and lexical marking of contrast in l2 italian. second language research, 2015. kristin j. van engen, melissa baese-berk, rachel e. baker, arim choi, midam kim, and ann r. bradlow. the wildcat corpus of native-and foreign-accented english: communicative efficiency across conversational dyads with varying language alignment profiles. language and speech, 53: 510–540, 2010. 29 ward and gallardo jan van santen, taniya mishra, and esther klabbers. prosodic processing. in springer handbook of speech processing, pages 471–488. springer, 2008. jan p.h. van santen, taniya mishra, and esther klabbers. estimating phrase curves in the general superpositional intonation model. in fifth isca workshop on speech synthesis, pages 61–66, 2004. nigel g. ward. automatic discovery of simply-composable prosodic elements. in speech prosody, pages 915–919, 2014. nigel g. ward. midlevel prosodic features toolkit. https://github.com/nigelgward/midlevel, 2015. nigel g. ward and paola gallardo. a corpus for investigating english-language learners’ dialog behaviors. technical report utep-cs-15-33, university of texas at el paso, department of computer science, 2015. nigel g. ward and alejandro vega. a bottom-up exploration of the dimensions of dialog state in spoken interaction. in 13th annual sigdial meeting on discourse and dialogue, 2012. nigel g. ward and steven d. werner. data collection for the similar segments in social speech task. technical report utep-cs-13-58, university of texas at el paso, 2013. nigel g. ward, rafael escalante, yaffa al bayyari, and thamar solorio. learning to show you’re listening. computer assisted language learning, 20:385–407, 2007. nigel g. ward, alejandro vega, and timo baumann. prosodic and temporal features for language modeling for dialog. speech communication, 54:161–174, 2011. nigel g. ward, yuanchao li, tianyu zhao, and tatsuya kawahara. interactional and pragmaticsrelated prosodic patterns in mandarin dialog. in speech prosody, 2016. j. c. wells. english intonation: an introduction. cambridge, 2006. john c. wells. sounds interesting. cambridge, 2014. eva wiberg. interactional context in l2 dialogues. journal of pragmatics, 35:389–407, 2003. anne wichmann. discourse intonation. covenant journal of language studies, 2(1), 2014. yi xu. speech melody as articulatorily implemented communicative functions. speech communication, 46:220–251, 2005. yi xu. speech prosody: a methodological review. journal of speech sciences, 1:85–115, 2011. frank zimmerer, jeanin jugler, bistra andreeva, bernd mobius, and jürgen trouvain. too cautious to vary more? a comparison of pitch variation in native and non-native productions of french and german speakers. in speech prosody conference, 2014. 30 dialogue and discourse 1(3) (2010) 1-33 doi: 10.5087/dad.2010.003 hilda: a discourse parser using support vector machine classification∗ hugo hernault hugo@mi.ci.i.u-tokyo.ac.jp graduate school of information science & technology the university of tokyo 7-3-1 hongo, bunkyo-ku, tokyo 113-8656, japan helmut prendinger helmut@nii.ac.jp national institute of informatics 2-1-2 hitotsubashi, chiyoda-ku, tokyo 101-8430, japan david a. duverle dave@kuicr.kyoto-u.ac.jp bioinformatics center, institute for chemical research kyoto university gokasho, uji, kyoto 611-0011, japan mitsuru ishizuka ishizuka@i.u-tokyo.ac.jp graduate school of information science & technology the university of tokyo 7-3-1 hongo, bunkyo-ku, tokyo 113-8656, japan editor: tim paek abstract discourse structures have a central role in several computational tasks, such as question–answering or dialogue generation. in particular, the framework of the rhetorical structure theory (rst) offers a sound formalism for hierarchical text organization. in this article, we present hilda, an implemented discourse parser based on rst and support vector machine (svm) classification. svm classifiers are trained and applied to discourse segmentation and relation labeling. by combining labeling with a greedy bottom-up tree building approach, we are able to create accurate discourse trees in linear time complexity with respect to the length of the input text. importantly, our parser can parse entire texts, whereas the publicly available parser spade (soricut and marcu, 2003) is limited to sentence level analysis. hilda outperforms other discourse parsers for tree structure construction and discourse relation labeling. for the discourse parsing task, our system reaches 78.3% of the performance level of human annotators. compared to a state-of-the-art rule-based discourse parser, our system achieves an performance increase of 11.6%. keywords: discourse parser, rhetorical structure theory, support vector machines 1. introduction in the last twenty years, the study of discourse has received continuous attention from the natural language processing community (moore and wiemer-hastings, 2003). discourse structure is fundamental to many text-based applications, such as question–answering (chai and jin, 2004) or ∗. this article is a significantly improved and extended version of duverle and prendinger (2009). c©2010 hugo hernault, helmut prendinger, david a. duverle, and mitsuru ishizuka submitted 2/10; accepted 11/10; published online 12/10 hernault, prendinger, duverle, and ishizuka dialogue generation (prendinger et al., 2007). those applications require the availability of robust and efficient discourse parsers. several attempts have been made to create discourse parsers in the framework of the rhetorical structure theory (mann and thompson, 1988), which is one of the most widely used theories of text organization. in rst, a text is first divided into non-overlapping text chunks, called elementary discourse units (abbreviated: edus). for instance, the following sentence, taken from the rst discourse treebank corpus (carlson et al., 2001): farm lending was enacted to correct this problem by providing a reliable flow of lendable funds. can be segmented into edus as shown in figure 1. [farm lending was enacted]1a [to correct this problem]1b [by providing a reliable flow of lendable funds.]1c (wsj1131) figure 1: segmentation of a sentence into edus next, consecutive edus are put in relation with each other, using a pre-defined set of rhetorical, or discourse, relations. the final goal of the discourse parser is to produce a tree structure as a representation of how all units of the text relate to each other. figure 2 shows two types of discourse relations in rst: hypotactic (‘mononuclear’) and paratactic (‘multi-nuclear’). the left-hand tree is an example of a mononuclear relation. here the two discourse units connected by the relation have different status, which is indicated by the direction of the arrow. the endpoint of the arrow denotes the ‘nucleus’ of the relation, whereas the unit at the other end is referred to as the ‘satellite’. the nucleus is considered as more prominent than the satellite. for example, condition is defined in carlson et al. (2001) as: “in a condition relation, the truth of the proposition associated with the nucleus is a consequence of the fulfillment of the condition in the satellite. the satellite presents a situation that is not realized.” other mononuclear discourse relations described in (mann and thompson, 1988; carlson et al., 2001) include: background, circumstance, elaboration and purpose. paratactic (‘multi-nuclear’) relations, on the other hand, have no distinguished nucleus and the relation can connect an arbitrary number of discourse units. the relation to the right in figure 2 is multi-nuclear. for example, list is defined as “[...] a multinuclear relation whose elements can be listed, but which are not in a comparison [...]” (carlson et al., 2001). the following discourse relations are also multi-nuclear (see mann and thompson (1988); carlson et al. (2001)): contrast, disjunction, sequence, and topic-comment. r relationname satellite nucleus relationname nucleus nucleus figure 2: the two relation types in rst. left: mononuclear; right: multi-nuclear figure 3 shows the tree representation of the preceding example (see figure 1). 2 hilda: a discourse parser using support vector machine classification enablement farm lending was enacted manner-means to correct this problem by providing a reliable flow of lendable funds. figure 3: a simple discourse tree (wsj1131) the tasks of discourse segmentation and relation labeling have been previously modeled using rule-based and statistical methods. in rule-based systems, rules manipulating syntactic and lexical information are defined manually. when applied to a text, these rules enable to detect the presence of edu boundaries or select the discourse relations holding between edus. however, given the heterogeneity of all possible texts, a rule-based model requires the creation of a large number of rules. supervised learning techniques offer an interesting alternative. with those techniques, labeled instances are extracted from a large body of text. classifiers are then trained to differentiate between edu boundary and non-boundary words, or assign discourse relation labels to edu pairs. the machine learning based approach for discourse parsing was greatly facilitated by the release of the rst discourse treebank (rst-dt) (carlson et al., 2001), a corpus of rst-annotated snippets from the wall street journal. previous supervised approaches were aimed at producing sentence level analysis (soricut and marcu, 2003) or at describing partially-implemented systems (reitter, 2003b). by contrast, our work targets discourse structure at text level. specifically, we created a fully-implemented, extensivelyevaluated system. in this article, we will present hilda (high-level discourse analyzer), a text level discourse parser based on support vector machine classification. in section 2, we will report on available approaches using rule-based and statistical discourse segmentation and parsing. in section 3, we will explain the two core tasks for an automated discourse parser (segmentation and relation labeling) and describe our working assumptions. section 3.4 is dedicated to the set of lexical, syntactic and structural features used in our model. in section 4, we will provide (1) a detailed evaluation of the discourse segmentation and relation labeling modules, (2) a full system evaluation, and (3) a comparison to other discourse parsers. finally, in section 5, we will conclude the article and discuss future work. 2. related work since marcu’s first attempt at developing a rule-based discourse parser (marcu, 2000), several algorithms for discourse parsing have been proposed, both statistical and rule-based. each of them extracts a different set of features from the input, and demonstrates different run-time complexity. in this section we present the most notable ones, which also proved instrumental in identifying 3 hernault, prendinger, duverle, and ishizuka useful features for our own algorithm. we start with discussing statistical approaches to discourse segmenting and parsing. spade (soricut and marcu, 2003) is a sentence level discourse parser. two probabilistic models are employed that use syntactic and lexical information to segment and to parse text. issues of algorithmic complexity that made their original algorithm impractical on large instances (marcu, 2000) are overcome by a dynamic programming approach. although spade does not support full text discourse parsing, it demonstrates the correlation between syntactic and discourse information and their capability to identifying structure and relations empirically, particularly when no signaling cue-word is present (between 60% and 70% of the total (taboada, 2006)). the authors’ exploration of ‘dominance sets’ in lexicalized syntax trees has provided the basis for a set of lexico-syntactic features in our own algorithm. a key difference to our approach is we are able to process full text input, rather than individual sentences. to the best of our knowledge, reitter (2003b) presents the only method based exclusively on feature-rich supervised learning to produce text level discourse parse trees. his algorithm relies on training a set of support vector machine classifiers to score and to label relations between spans. although the author’s suggestion of a bottom-up tree building algorithm using chart parsing style techniques has not been implemented yet, the results for raw instance classification provides a good intermediate benchmark for the evaluation of our own instance classifier. more recently, baldridge and lascarides (2005) successfully implemented a probabilistic parser that uses headed trees to label discourse relations. the authors employed the more specific framework of segmented discourse representation theory (asher and lascarides, 2003) rather than rst, which they applied to texts in dialogue form. in sagae (2009), a discourse parser based on shift-reduced parsing is presented. the parser is trained on the rst-dt, employs transition algorithms for dependency and constituent trees, and uses its own syntax and dependency tagger. it brings noticeable improvements in accuracy and speed against marcu’s original chart parsing approach (marcu, 2000). the parser also includes a discourse segmenter based on a binary classifier trained on lexico-syntactic features. the discourse segmenter performs better than the one presented in reitter (2003b) and is significantly faster. importantly, compared to soricut and marcu (2003), the proposed system is able to create discourse structures for full texts, not only sentences. for discourse segmentation, the author reports an f-score of 86.7% and an f-score of 44.5% for creating text level discourse trees. subba and di eugenio (2009) propose a discourse parser based on inductive logic programming, a first-order logic approach for learning discourse relations. in their system, besides traditional syntactic and discourse cue information, rules are learnt based on semantic information from wordnet (fellbaum, 1998), similarity measures, structural properties between text segments, as well as compositional semantics. then, a shift-reduce parsing model is employed to produce discourse structures. the authors employ a custom corpus of instructional texts, manually segmented into discourse units, and annotated with a custom set of 26 discourse relations. for sentence-level relation labeling, their system reaches an f-score of 63%, while the f-score for text level discourse trees is 35.44%. a statistical discourse segmenter based on artificial neural networks is presented in subba and di eugenio (2007). the system employs a multi-layer perceptron model and is trained on the rst-dt using syntactic and lexical information, in particular discourse cues. bagging (breiman, 1996) is applied to reduce over-fitting. the performance of this segmenter is comparable to the one 4 hilda: a discourse parser using support vector machine classification employed in soricut and marcu (2003), with an f-score of 84.4% (86% when using perfect parse trees). next we turn to rule-based approaches to segmentation and discourse parsing. le thanh et al. (2004b) propose a multi-step algorithm to segment text and organize its spans into trees for each successive level of text organization: sentence level, paragraph level, and text level. within each level, the authors explore a search space consisting of all valid trees fitting the ‘sequentiality principle’ (i.e., two spans connected in the tree are adjacent in the original text, see section 3.2). the multi-level nature of their algorithm mitigates the combinatorial explosion effect. however, at the text level, the algorithm has to score a large number of trees to extract the best candidate—despite the use of beam search as a method to explore the solution space. hence, this approach is impractical for large input text. furthermore, the assumption of strong tree-consistency (the adjacency hypothesis) at each organizational level is not supported in the rst corpus, where a majority of documents contain organizational blocks that do not map to valid subtrees. tofiloski et al. (2009) argue that using rules has certain advantages over automatic learning methods and present a rule-based discourse segmenter. their proposed model does not depend on a specific training corpus and supports high-precision segmentation by inserting fewer but ‘quality’ boundaries. however, segmentation is done in a manner different from the segmentation guidelines used in the rst-dt corpus (carlson et al., 2001). first, the authors only create edus that contain verbs. they try to capture specific relations (e.g., condition, purpose) and avoid less informative relations, such as elaboration. further, complement clauses are not placed in independent units in order not to break np constituency, and avoid same-unit relations. for instance, “he said that” is not considered an independent edu. in this article, we will not enter the discussion of what constitutes the best segmentation guidelines, and simply take the rst-dt corpus as the basis for segmentation by supervised methods. 3. approach 3.1 the rst discourse treebank as previously mentioned, rst-dt is a large corpus of documents annotated with edu segmentation and full text level rhetorical structure (carlson et al., 2001). the corpus was released in 2002 and contains 385 articles transcribed from the wall street journal. these articles are a subset of the penn treebank corpus of english (marcus et al., 1993). the corpus allows us to both train and evaluate the performance of our algorithm on a large number of documents of varied lengths. although the original rst set is composed of 24 discourse relations (mann and thompson, 1988), this list has evolved to fit the type of application and expressive power required by each researcher. in the rst-dt, relation annotation is done using a set of 78 rhetorical relations, which enables a high level of expressivity. these relations are divided in two categories: 53 mononuclear relations and 25 multi-nuclear relations. 3.2 working assumptions an essential characteristic of rst is that “a left-to-right reading of the terminal frontier of a discourse tree associated with a text corresponds to the span of text analyzed, in the same linear order.” (marcu, 2000). this ‘sequentiality principle’ guarantees that only consecutive spans of a text can be put into relation, which dramatically reduces the size of the solution space (see section 3.3). 5 hernault, prendinger, duverle, and ishizuka when building discourse parsers, researchers have opted to use small relation sets, which makes the problem of assigning relations easier. for instance, marcu (2000) used 15 relations, while le thanh et al. (2004b) used 14 relations. although the rich 78-relations set of the rst-dt enables a high level of expressivity, which is useful in the case of linguistic studies, it has unnecessary precision for most text analysis applications. furthermore, from a machine learning perspective, working with a smaller set of relations improves the computational properties of the problem. a problem with fewer classes guarantees a better separability between the different classes, as it abstracts from the ambiguities inherent to fine-grained relations. we adopt a similar strategy and work with the 18 relations defined in carlson et al. (2001), and previously used by soricut and marcu (2003). in this reduced set, the original rst-dt relations are partitioned into 16 categories, which correspond to (merged) relations (called “classes” in carlson et al. (2001)), depending on their rhetorical similarity.1 for instance, the semantically similar problem-solution, question-answer, statement-response, topic-comment and comment-topic relations, are merged to topiccomment. two extra relations are used for helping the structuring of the text, textual-organization and same-unit. in rst, multi-nuclear relations can take an arbitrary number of arguments. for instance, the list relation can connect any number of spans, leading to n-ary tree representations. in order to maintain compatibility with our svm classification approach, we have to convert n-ary trees to binary trees. this conversion can be done trivially, by consecutively nesting the arguments of multi-nuclear relations, in the fashion shown in figure 4. list [1] [2] [3] [4] 1 list [1] list [2] list [3] [4] 1 figure 4: binarization of multi-nuclear relations it is interesting to note that the reverse transformation is possible most of the time. however, a loss of reversibility occurs when two similar multi-nuclear relations are nested consecutively (for instance, two lists). however, these cases occur rarely in practice—only 15 times in the 380 documents of our corpus. 3.3 outline of discourse segmentation and discourse relation labeling figure 5 shows the basic workflow of the hilda parser (details of the workflow are provided in appendix a). 1. for a complete list of the relations used, see appendix b. 6 hilda: a discourse parser using support vector machine classification • a text is first segmented into edus. • then the relation labeling step evaluates which relations are likely to hold between consecutive edus. the two edus which are most probably connected by a rhetorical relation are merged into a rhetorical structure tree of two edus. • next, we go back to the labeling step to re-evaluate which relations are the most likely to hold between spans (rhetorical structure trees of any size, including atomical edus). • this procedure is repeated until all spans are merged. the outlined method enables us to find a good approximation of ‘perfect’ discourse trees (i.e., the ones created by human annotators) in linear time complexity with respect to the length of the input text. input text bdiscourse segmentation relation labeling tree construction discourse tree figure 5: basic parser workflow the tasks of discourse segmentation and relation labeling are modeled as supervised classification tasks. we chose to use support vector machines (svm) (vapnik, 1995). svms are a set of maximum-margin classifiers, i.e., they minimize the classification error and maximize the geometric margin. they have been used extensively in domains ranging from bio-informatics to natural language processing, and are considered as a state-of-the-art classifier for many practical tasks. in discourse segmentation, the problem is to assign each word of the input text an observation category c ∈ {0, 1}, where 0 indicates that a word is not a boundary, and 1 indicates that it is a boundary. for example, in the snippet from figure 1, the words said, quarter and yesterday are assigned 1, whereas all other words belong to category 0. hence, we train a binary classifier seg : w → {0, 1}, wherew is the set of all words occurring in the input text, denoted by inputtext. after the segmentation step, we obtain a list of edus, e = 〈e1, e2, . . .〉. the concatenation of these edus, from left to right, gives back inputtext. for relation labeling and tree construction, we apply the following method. given two consecutive spans (edus or subtrees), we determine (1) the likelihood of a direct structural relation, (2) the probabilities for the label of the relation with nuclearity of the spans. a full rhetorical structure tree is produced by applying this mechanism repeatedly in a straightforward bottom-up fashion. definition 1 (discourse relation set) a discourse relation setr is defined as a set r = {r1, r2, . . . ri, . . . , rn}, such that ∀i, ri = 〈rri, lefti, righti〉, whereby: • rri ∈ {attribution, cause, elaboration, list, . . . } (see appendix b for a full list of relations considered in hilda); • lefti, righti ∈ {nucleus, satellite}; 7 hernault, prendinger, duverle, and ishizuka • all ri triplets are valid, i.e., given a rhetorical relation rri, the left and right nuclearities lefti and righti are allowed nuclearity options. our working rhetorical relation set and allowed nuclearity options are defined in appendix b. we employ 18 rhetorical relations which, when nuclearized, give a total of 41 possible ri triplets (the cardinal number ofr). definition 2 (valid rs-tree) a rs-tree t is valid if it satisfies the following properties: 1. all leaf nodes of t are edus;2 2. all non-leaf nodes of t are tagged with a discourse relation ri ∈ r. t = {t1, t2, . . . ti, . . .} denotes the set of all valid rs-trees. in our algorithm, rs-trees are implemented using binary tree structures. to obtain a good classification accuracy, we train two separate classifiers: • a binary classifier for structure labeling, i.e., existence of a rhetorical relation between two valid rs-trees: struct : t × t → {0, 1}. • a multi-class classifier for rhetorical relation and nuclearity labeling: label : t × t → r. label returns predicted relation label and nuclearities. the detailed algorithm for tree construction is provided in figure 6. we first create a list containing all edus of the input text, in left-to-right reading order. the binary classifier struct is then applied to all pairs of consecutive elements from this list, in order to determine the probability of a structural relation between consecutive edus. next we select the pair with the highest probability and apply label to it, in order to determine the relation’s label and nuclearity. we can now remove the two selected edus from the list, and replace them by a newly created rs-tree labeled by the estimated relation label and nuclearity, whose left and right children are the selected edus. we update the probabilities of structural relations holding between the newly created tree and adjacent elements in the list, using struct. subsequently, we repeat the process until the list contains only one element, which is our final discourse tree, built bottom-to-top. in doing so, we have adopted a ‘greedy’ approach, i.e., at each iteration the two spans most likely connected by a relation have been merged. hence, our algorithm, which makes locally-optimal choices, performs in time complexity of o(n). 3.4 features 3.4.1 discourse segmentation the first step of the parser is to segment the text into units. for this task, we use a combination of syntactic and lexical features: words, pos tags, and lexical heads. in particular, we use the lexicosyntactic features of soricut and marcu (2003), which were found good indicators for the presence of edu boundaries. figure 9 shows the parse tree of the sentence “farm lending was enacted to correct this problem by providing a reliable flow of lendable funds.” (see also figure 1). here, lexical heads are generated using projection rules from magerman (1995) and indicated between brackets. for a word w, we look at its highest ancestor in the parse tree with a lexical head equal to w, and with a right-sibling. 2. note that a single edu also constitutes a valid rs-tree. 8 hilda: a discourse parser using support vector machine classification require: e = 〈e1, e2 . . . 〉, list of all the text’s edu ensure: finaltree is a valid rs-tree for inputtext l← e for all (li, li+1) in l do scores[i]← struct(li, li+1) a end for while |l| > 1 do i← argmax(scores) newlabel← label(li, li+1) newsubtree← createtree(li, li+1, newlabel)b scores[i− 1]← struct(li−1, newsubtree) scores[i+ 2]← struct(newsubtree, li+2) delete(scores[i]) delete(scores[i+ 1]) l← [l0, . . . , li−1, newsubtree, li+2, . . . ] end while finaltree← l0 return finaltree a. li denotes the i-th element of list l. b. createtree is a function that takes two rs-trees t and t ′, a relation r ∈ r, and returns the rs-tree whose left child is t , whose right child is t ′, and which is tagged with relation r. figure 6: bottom-up tree construction this highest-ancestor node is called nw. then, we call its parent np, and its right-sibling nr. for instance, when applying this process to the word “enacted”, we obtain nw = vbn(enacted), np = vp(enacted), nr = s(correct). we can now define the contextual features of the word at position i in the text. it is the set composed of the word wi, its pos, as well as the pos and lexical heads of nwi, npi, and nri. next, the features for position i in the text are created by concatenating the contextual features at positions i − 2, i − 1, and i. those positions were determined empirically to give the maximum contextual coverage. 3.4.2 relation labeling for relation labeling, several shallow lexical and syntactic features are taken into account, including features of textual organization, lexical features, ‘dominance sets’ (soricut and marcu, 2003), and structural features. the first type of features we incorporate is related to textual organization. a number of previous efforts (marcu, 1996; soricut and marcu, 2003) have shown that there is a strong correlation between different organizational levels of textual units and subtrees of the rs-tree, both at the sentence level and at the paragraph level. hence, the following features provide valuable high-level cues for the task of scoring span relation priority (classifier struct). as hypothesized by reitter (2003b), there is a correlation between span length and type of rhetorical relations. for instance, the satellite of a contrast relation tends to be shorter than its nucleus. 9 hernault, prendinger, duverle, and ishizuka the organizational features taken into account in our system are presented in table 1. in this table, all subtree-specific features, which are symmetrically extracted from both left and right candidate spans, are suffixed by “s[pan]”. the other features, calculated as a function of the two subtrees considered as a pair, are indicated by “f[ull]”. table 1: features encoding textual organization feature name scope belong to same sentence f belong to same paragraph f number of paragraph boundaries s number of sentence boundaries s length in tokens s length in edus s distance to beginning of sentence in tokens s size of span over sentence in edus s size of span over sentence in tokens s size of both spans over sentence in tokens f distance to beginning of sentence in edus s distance to beginning of text in tokens s distance to end of sentence in tokens s discourse cues, such as because, however, etc are another type of feature that frequently indicate the presence of discourse relations. instead of working with a pre-defined dictionary of discourse cues, we chose to build an empirical n-gram dictionary from the training corpus, ranked by frequency, in order to keep it to a reasonable size. this method has the advantage of capturing non-lexical signals, such as punctuation and paragraph boundaries. we encode the prefix and suffix of each span as ordered 3-grams, which were found to give the best signal-to-noise ratio (see manning and schütze, 1999). encoding the beginning and end of a span is motivated by the hypothesis that the most meaningful rhetorical signals are found at the edges of the span (schilder, 2002). for instance, for the span in figure 3 consisting of units 1b and 1c connected by a manner-means relation, the beginning and end-of-span 3-grams we encode in our features (to, correct, this) and (of, lendable, funds), respectively. in total, our 3-gram dictionary is of size 12,000, and encoding prefix and suffix raise the encoding size to 22 × |3-gram dictionary| = 48, 000. this approach was validated by comparing it to results obtained from using a fixed dictionary of 300 cue-phrases instead (oberlander and moore, 1999). performance was found to be lower when using only the cue-phrase dictionary, and marginally lower when using both together. to improve discourse cue detection and to make it less dependent on lexical information, we complement it with shallow syntactic information, by encoding pos tags for the prefix and suffix of each span. the penn treebank contains npos = 384 tags, hence it requires 2 × 3 × npos = 2304 additional dimensions. to reflect the way relations attach in rs-trees, we use the notion of dominance sets as defined in soricut and marcu (2003). a dominance set is formed by consecutive edus of a sentence that are connected in the lexicalized syntax tree by a head node (the dominating node) and an attachment node (the dominated node). for instance, the sentence of figure 1 contains three edus, which 10 hilda: a discourse parser using support vector machine classification are connected by two rhetorical relations, enablement and manner-means. this leads to two possible tree constructions (see figure 8), in which edu 1b is connected either to 1a or to 1c. in the left-hand construction, 1a and 1b are connected first, then the subtree they span is connected to 1c. in the right-hand construction, 1b and 1c are connected first, then the subtree they span is connected to 1a. it is possible to infer the correct logical nesting order by studying the associated syntax tree, and observing the sub-trees spanned by each edu. in figure 9, looking at the frontiers between dominating nodes (diamond-shaped) and dominated nodes (oval-shaped) indicates the correct dominance order, 1a > 1b > 1c. thus, 1b should be attached to 1c, and the correct tree in figure 8 is the right-hand one. soricut and marcu (2003) also note that only taking the pos tags of the dominance nodes into account is too general. then, in order to augment information about the lexico-syntactic context, lexical heads of those nodes are also taken into account. for approximating soricut and marcu’s definition of dominance sets, we incorporate the lexicosyntactic and structural features presented in table 2 into our model. it is worth noting that most of these features apply only when both spans belong to the same sentence. otherwise, the parse trees of spans are disjunct, and the notion of dominance does not apply. in this case, the features receive default values during feature extraction. table 2: features encoding dominance sets feature name scope distance to root of the syntax tree s distance to common ancestor in the syntax tree s delta of distances to common ancestor f dominating node’s lexical head in span s common ancestor’s pos tag f common ancestor’s lexical head f dominating node’s pos tag f dominating node’s lexical head f dominated node’s pos tag f dominated node’s lexical head f dominated node’s sibling’s pos tag f dominated node’s sibling’s lexical head f relative position of lexical head in sentence s the strong compositionality criterion was found by marcu through empirical analysis of large rst trees: “[. . . ] whenever two large text spans are connected through a rhetorical relation, that rhetorical relation holds also between the most important parts (i.e. the nuclei) of the constituent spans.” (see marcu (1996, 2)). this criterion, he argues, provides a good indicator of validity in the case of text level trees where relations can connect large spans of text beyond the sentence level. the idea of strong compositionality is implemented by replicating the set of lexical and syntactic features of table 2, which are extracted from the ‘most important’ edu, called main constituent, of each span. the main constituent of the left span is found by recursively selecting the right-most nucleus, starting from the top relation. for the right span, the left-most nucleus is selected instead. we chose the leftmost and rightmost nuclei, respectively, so that the edus considered are as close 11 hernault, prendinger, duverle, and ishizuka as possible to each other. marcu (1996) recommends selecting the union of all nuclei in each span. however, the restriction we impose is necessary to accommodate the fixed length nature of our feature vectors. finally, to guide decisions regarding high-level relation classification on large spans, we encode the rhetorical structure of a span into the feature vector. we (informally) observed some level of correlation between relations at different levels of the tree throughout the corpus. this is trivially the case for n-ary relations such list which have been binarized in our representation; the presence of several list relations in right-most nodes of a subtree greatly increases the chance that the parent relation might be a list itself. we encode this structure as a breadth-first, flat list representation of the binary tree, whereby each element of the list is split over one of the 41 potential class labels. as the feature vectors have a fixed length, we set an arbitrary limitation on the depth encoded. this increases the encoding size by 2×|r|×2d = 41×2d+1, where d is the maximal depth considered. experiments have shown optimal performances for 2 ≤ d ≤ 3 (we selected d = 3) with sharp decreases for values of d > 4. this result can be explained by the excessively noisy data brought about by an exponentially growing set of features. 4. experiments 4.1 discourse segmentation we first measure the performance of the seg classifier. after feature extraction from the 341 training documents and 38 test documents of the rst-dt, we obtain 177,633 training vectors and 21,667 test vectors for this task. because a sentence typically contains few edu boundaries, our dataset is skewed, with approximately 89% negative training examples and 11% positive examples. our segmenter is evaluated with respect to two types of competing systems. first, we measure the segmentation result when using parse trees from the penn treebank (marcus et al., 1993) as our gold standard. second, as a practical evaluation, we compare the performance when using parse trees generated respectively by the stanford parser3 (klein and manning, 2003) and by the charniak parser4 (charniak, 2000). evaluation is done on the test subset of the rst-dt. in our experiments, we use the same measure as soricut and marcu (2003) and tofiloski et al. (2009), i.e., we only measure the score on boundaries inside sentences. this avoids artificially boosting the performance by including obvious end-of-sentence or end-of-paragraph boundaries. the radial basis function (rbf) (vapnik, 1995) kernel is selected, and parameter estimation is done using grid search with 5-fold cross validation (staelin, 2003). in practice, we did not observe any problem related to the class imbalance in the training set with these parameters. the performance of our and other available segmenters is shown in table 3. nnds refers to the neural-networks discourse segmenter (subba and di eugenio, 2007), spade-seg is the segmentation module of spade, described in soricut and marcu (2003), and sag refers to the segmenter included in the discourse parser of sagae (2009). table 3 (top row) shows that hilda-seg significantly outperforms other discourse segmenters. with gold standard trees, spade-seg and nnds yield f-scores of 84.7% and 86%, respectively, versus 95% for hilda-seg. the measure of the human annotators’ agreement for the segmentation task has been 3. http://nlp.stanford.edu/software/lex-parser.shtml 4. ftp://ftp.cs.brown.edu/pub/nlparser/ 12 hilda: a discourse parser using support vector machine classification table 3: performance comparison with other segmenters system trees precision recall f-score spade-seg penn 84.1 85.4 84.7 nnds penn 85.5 86.6 86.0 hilda-seg penn 95.5 94.5 95.0 spade-seg charniak 83.5 82.7 83.1 nnds charniak 83.9 84.8 84.4 hilda-seg charniak 94.7 93.4 94.0 sag sag 87.4 86.0 86.7 hilda-seg stanford 94.5 93.1 93.8 human agreement – 98.5 98.2 98.3 calculated in soricut and marcu (2003), with a f-score of 98.3%. using perfect parse trees, hildaseg reaches 96.6% of the annotators’ level. when the stanford parse trees are selected, 95.4% of this level is reached. the current state-of-the-art discourse segmenter, sag, in which a custom syntax parser is employed, has an f-score of 86.7% for this task. when using stanford parse trees, our segmenter thus provides a 8.2% performance increase compared to sag (9.6% when using perfect trees). in table 3, we decided not to include the results of the rule-based segmenters of le thanh et al. (2004a) and tofiloski et al. (2009), for several reasons. first, le thanh et al. (2004a) report their results using a ‘softer’ metric, in which end-of-sentence boundaries are taken into account. the authors used penn treebank parse trees, and after evaluation on 8 texts of the rst-dt, obtain a precision value of 81.4% and a recall value of 79.2%. finally, the results of tofiloski et al. (2009) cannot be compared directly to ours, as different segmentation guidelines were used. the authors report a precision value of 89% and a recall value of 86% when using charniak parse trees, and a precision value of 82% and a recall value of 86% when using stanford trees. moreover, these scores were measured on three texts of the rst-dt only, which makes comparison even more difficult. to investigate which features contributed to the performance of the discourse segmenter, we run several experiments that measure the segmentation performance when training with different combinations of features, taken at different relative positions. for each experiment, we use parse trees from the stanford parser. the results are indicated in table 4. position (0) and (−1, 0) indicate that, when extracting features for the word at position i in the text, we only used contextual features from position i, and the concatenation of contextual features from positions i−1 and i, respectively. we observe that, for position (0), using as features the current word w and its part of speech pos(w) only, we obtain a precision of 82.7%, but a comparatively lower recall of 70.8%. however, when using as features the three nodes nw, np, nr of the parse tree only (see 3.4.1), the precision is slightly lower (82%) but the recall higher than the previous case, at 78.4%. when using the combination of all these features, we obtain a precision of 81.9% and a recall of 79.4%. with this set of relevant features defined, we now increase the coverage by also including features extracted from the previous word in our feature set. with these features, now taken from positions 13 hernault, prendinger, duverle, and ishizuka (−1, 0), we notice a sharp f-score increase of 14.5%. in particular, precision reaches 93.3% and recall 91.5%. finally, we add the contextual features for the word at position i − 2 to the feature set. the f-score increase in this case is much smaller (1.5%). we now attain a precision of 94.5% and a recall of 93.1%. from this point forward, adding features from anterior words did not improve performance, although it significantly increased training time. finally, we did not notice any significant performance increase when including contextual features for words located ahead of the current position. table 4: segmentation performance with different sets of features positions features p r f (0) w 80.7 69.4 74.6 (0) w, pos(w) 82.7 70.8 76.3 (0) nw, np, nr 82.0 78.4 80.2 (0) w, pos(w), nw, np, nr 81.9 79.4 80.7 (−1, 0) w, pos(w), nw, np, nr 93.3 91.5 92.4 (−2,−1, 0) w, pos(w), nw, np, nr 94.5 93.1 93.8 4.2 relation labeling although our final goal is to achieve good performance on the entire tree-building task, a useful intermediate evaluation of our system can be conducted by measuring the raw performance of our svm classifiers. the binary classifier struct is trained on 52,683 instances (split approximately between 1/3 positive and 2/3 negative examples), extracted from 341 documents, and tested on 8,558 instances extracted from 38 documents. the feature space has a dimension of 136,987. the other classifier, label, is trained on 17,742 instances (labeled across 41 classes) and tested on 2,887 instances, the same feature vector dimension as struct. the optimal training parameters for each kernel function are obtained through automated grid search with 5-fold cross-validation (staelin, 2003). table 5 shows the training time and accuracy for these experiments, using various software and selected different kernels, for both struct and label, as well as a comparison to the results of reitter (2003a). for struct we used the regression mode of the svm software to obtain probabilistic output. note that in table 5, the results for struct and label are independent, as they indicate performance for cross-validation. 14 hilda: a discourse parser using support vector machine classification table 5: svm classifiers performances – ll: liblineara, svml: svm lightb, svmm: svm multiclassc, ls: libsvmd classifier struct label reitter kernel linear polyn. rbf linear rbf rbf software ll svml svml svmm ls svml training time 21.4s 5m53s 12m 15m 23m 216m accuracy 82.2 85.0 82.9 65.8 66.8 61.0 a http://www.csie.ntu.edu.tw/˜cjlin/liblinear/ b http://svmlight.joachims.org/ c http://svmlight.joachims.org/svm multiclass.html d http://www.csie.ntu.edu.tw/˜cjlin/libsvm/ despite their simplicity, the noticeably good performances of linear kernel methods in the results presented in table 5, as compared to the more complex polynomial and rbf kernels, indicate that our data separates fairly well linearly. this is a commonly observed effect of high-dimensional input such as ours (over 100,000 features) (chen et al., 2007). reitter (2003a) provides a baseline for absolute comparison on the multi-label classification task for a similar classifier that assumes perfect segmentation of the input, as ours does. reitter’s accuracy results of 61% for a smaller set of training instances (7976 instances from 240 documents compared to 17,742 instances in our case) with considerably less classes (16 rhetorical relation labels with no nuclearity, as opposed to our 41 nuclearized relation classes) seems to indicate that this sub-component of our system with an accuracy of 66.8% performs well. in hilda, we select a linear kernel for struct and an optimally parameterized rbf kernel for lab, considering performance and runtime complexity. we use modified versions of the liblinear (optimized for large-instances) and libsvm packages. all further evaluation results reported in this article were conducted using these kernels. to give a more precise idea of the classifiers’ performance, we look into the results for each class, when trained on rst-dt’s standard training set and evaluated on the standard test set. table 6 indicates the results for struct, while table 7 indicates the results for label. in table 6, we notice in particular that the performance for the class ‘+1’, corresponding to the presence of a structural relation between to rs-trees is lower (f-score of 73.3%) than for the absence of a relation (f-score of 86.7%). because rst-dt’s test set is relatively small (38 documents), and as certain discourse relations occur rarely in natural language texts, several of label’s classes are not present in the test set and thus not presented in table 7. we observe that performance varies widely across classes. for instance, attribution is very well detected, with f-scores close to 95%, whereas cause[n][s] has a very low f-score of 3.9%. here, the average f-score is 47.7%. as we employed a linear kernel in struct, it is possible to identify which features were useful, by looking at the weight assigned by the classifier to each feature in the svm model file. table 8 contains the 25 features with the highest weights in absolute value. these features are thus useful for detecting the presence of a discourse relation between two rs trees. in particular, we observe 15 hernault, prendinger, duverle, and ishizuka table 6: class-specific performance for struct, evaluated on rst-dt’s test set svm class precision recall f-score +1 74.0 72.7 73.3 −1 86.3 87.1 86.7 table 7: class-specific performance for label, evaluated on rst-dt’s test set svm class precision recall f-score attribution[n][s] 93.6 96.2 94.9 attribution[s][n] 95.7 93.7 94.7 background[n][s] 47.8 41.5 44.4 background[s][n] 38.7 20.7 27.0 cause[n][s] 33.3 2.1 3.9 comparison[n][s] 50.0 5.9 10.5 condition[n][s] 100.0 47.8 64.7 condition[s][n] 85.7 72.0 78.3 contrast[n][n] 31.1 21.9 25.7 contrast[n][s] 50.0 20.8 29.4 contrast[s][n] 51.1 39.7 44.7 elaboration[n][s] 58.1 94.5 72.0 enablement[n][s] 61.9 59.1 60.5 enablement[s][n] 50.0 50.0 50.0 explanation[n][s] 66.7 16.8 26.9 joint[n][n] 62.2 55.2 58.5 manner-means[n][s] 53.8 29.2 37.8 same-unit[n][n] 78.8 96.9 86.9 summary[n][s] 100.0 31.2 47.6 temporal[n][n] 83.3 11.9 20.8 temporal[n][s] 85.7 24.0 37.5 textual-organization[n][n] 33.3 22.2 26.7 topic-change[n][n] 83.3 38.5 52.6 that the most heavily-weighted feature is the binary feature indicating if the two input spans belong to the same sentence. among these features, we notice several measures of distance, such as the distance of the left span to the sentence beginning, the relative size of the left and right span’s edus over the sentence length, or the number of paragraphs in the right span. we also find several features related to dominance sets, such as features encoding the dominated and dominating node’s pos tags and lexical heads. 16 hilda: a discourse parser using support vector machine classification table 8: top 25 svm weights of the linear kernel for struct. feature weight both spans belong to the same sentence 4.118836 size of span over sentence in edus 3.582545 distance of the left span to beginning of sentence in edus -3.437157 common ancestor’s pos tag is ‘prn’ -2.911269 dominating node’s lexical head is ‘which’ -2.668148 pos tag of the right span’s last token is ‘.’ 2.636921 size of left span over sentence in tokens -2.341654 size of both spans over sentence in tokens -2.222655 left and right span belong to the same sentence -2.217709 pos tag of the left span’s last token is ‘.’ 2.170483 number of paragraphs in the right span 2.123776 relative position of lexical head in sentence 2.084602 pos tag of the left span’s last token is ‘”’ 2.051767 dominated node’s sibling’s pos tag is ‘vbg’ -1.875610 pos tag of the right span’s first token is ‘to’ -1.869705 distance of the right span from the sentence beginning -1.839178 pos tag of the right span’s last token is ‘,’ -1.814842 top lexical head of the left span is ‘which’ 1.777975 top lexical head of the sentence belongs to the right span -1.774898 top lexical head of the left span is ‘of’ -1.706626 dominating node’s lexical head is ‘set’ -1.670822 dominated node’s lexical head is ‘ideas’ -1.669270 dominating node’s pos tag is ‘nn’ -1.650378 dominated node’s sibling’s pos tag is ‘vbn’ 1.628170 dominating node’s pos tag is ‘vbn’ -1.617607 4.3 full system performance a measure of the performance of our full system is realized by comparing the structure and labeling of the rs-tree produced by our algorithm to that obtained through manual annotation (our gold standard). here, we use the parseval metrics (black et al., 1991) that defines the performance metrics (precision, recall, f-score) of a parsing system. for comparability, the tree constituents used in the method are enumerated using marcu’s formalism (see marcu, 2000, 143–144), which assigns relation labels to the child nodes of the relation instead of the parent node, thus making it easier to represent n-ary relations. in practice, this formalism yields lower scores than our own, likely because of the greater emphasis placed on tree structure accuracy over relation labeling. hilda is evaluated based on three experiments. 1. perfectly segmented input taken from the rst-dt (indicated as ‘manual’ in the following tables). in this case, we use parse trees from the penn treebank in all steps of the parsing. 17 hernault, prendinger, duverle, and ishizuka 2. the system is run using the output of spade’s segmenter (spade-seg), in which case charniak’s parser is used for generating parse trees. 3. the system is run with our own segmenter (hilda-seg), presented in section 4.1. in this case, we use parse trees from the stanford parser. the first measure gives us a good estimate of our system’s optimal performance (given optimal input), while the others provide a concrete performance evaluation under practical conditions. our measurements are taken based on the standard test subset of 41 files provided by the rst-dt corpus. for each experiment, parse trees are evaluated on four successive, increasingly complex criteria. first, on the blank tree structure (‘s’), then on the tree structure with nuclearity indication (‘n’), then the tree structure with rhetorical relation indication but no nuclearity indication (‘r’), and finally on the fully-labeled tree structure with both nuclearity and rhetorical relation labels (‘f’). for each criteria, precision is calculated by taking the ratio of the number of identical tree constituents found in both the generated rs-tree and the gold-standard tree, against the total number of constituents in the generated discourse tree. recall is calculated by taking the ratio of the number of identical tree constituents found in both the generated rs-tree and the gold-standard tree, against the total number of constituents in the gold-standard discourse tree. the results are summarized in figure 9. note that when using perfect segmentation, precision and recall are identical since both trees have same number of constituents. table 9: full parser evaluation, using the standard test subset seg. manual spade-seg hilda-seg s n r f s n r f s n r f precision 83.0 68.4 55.3 54.8 69.5 56.1 44.9 44.4 73.0 59.7 48.2 47.7 recall 83.0 68.4 55.3 54.8 69.2 55.8 44.7 44.2 71.7 58.6 47.4 46.9 f-score 83.0 68.4 55.3 54.8 69.3 56.0 44.8 44.3 72.3 59.1 47.8 47.3 as expected, we observe the highest performance when using manual segmentation, with an f-score of 83% for structure annotation, and 54.8% for the fully-annotated tree. using spade yields lower results, with an f-score of 69.3% for structure annotation, and 44.3% for the full discourse tree. on the other hand, when using the discourse segmenter presented in section 4.1, structure annotation produces f-scores of 72.3% for structure annotation, and 47.3% for the full tree. these differences indicate that, although these segmenters score high on the segmentation task (above 80%), errors occurring at the beginning of the pipeline are amplified and weight heavily on the performance of the output. furthermore, the parse trees of spade-seg and hilda-seg are generated automatically, which also explains the lower results. in order to compare our results to the gold standard (defined as manual agreement between human annotators) more accurately, we also evaluated performances using another test set (see table 10) of 52 files annotated by two labelers of the rst-dt. in each case, the remaining 340-350 files are used for training. the f-score for human agreement is 65.3%. when using manual segmentation, we obtain an f-score of 55.1% on this test set, which is 84% of human performance level. when using our own segmenter, we obtain an f-score of 48.8%, which is 74.7% of human performance level. 18 hilda: a discourse parser using support vector machine classification table 10: evaluation on a doubly-annotated subset and comparison to human agreement system performance human agreement seg. manual hilda-seg – s n r f s n r f s n r f precision 84.1 70.6 55.6 55.1 74.0 61.7 49.4 48.9 88.0 77.5 66.0 65.2 recall 84.1 70.6 55.6 55.1 73.7 61.5 49.1 48.7 88.1 77.6 66.1 65.3 f-score 84.1 70.6 55.6 55.1 73.8 61.6 49.2 48.8 88.1 77.5 66.0 65.3 4.4 comparison to other systems in this section, we attempt to compare our system to other text level discourse parsers, namely marcu’s decision-tree based parser (marcu, 2000), the multi-level rule-based system built by le thanh et al. (2004b), and the shift-reduce parser of sagae (2009). for comparison with marcu and lethanh’s systems, we face several difficulties: marcu (2000) reports his results on unspecified documents and we do not have access to the parser software. therefore we could not evaluate his parser on our test set. then, in the case of lethanh’s system, we do have access to the documents on which the system was evaluated, but we do not have access to the parser software either. another issue comes from the number of discourse relations employed: in marcu’s case, 15 relations were employed, while lethanh used 14, and we used 18. for these two systems, we have no choice but to reproduce the results provided by each author, in separate tables for fairness. we define a measure, falgo/fhuman, computed for each system, as an informal indicator of how each parser is estimated to perform compared to the human agreement. as each author found a different value for human agreement f-score, we scale the score of each system by the human agreement f-score found by its respective author. for instance, le thanh et al. (2004b) report a surprisingly low 72.7% f-score for structure, when marcu (2000) notes 81%, and we found 87%. this suggests that, despite our best efforts, evaluation metrics used by each author might differ. again, this measure is an informal indicator of performance. table 11 indicates the results given by marcu (2000), table 12 the results given by le thanh et al. (2004b), and table 13 the results for our system. note that in table 13, in order to make our experimental setting as close as possible to lethanh’s, we used lethanh’s set of 21 documents as testing subset, and the rest of the corpus for training. using our calculated human-agreement f-scores, we estimate that our system reaches 86.4% of the human performance level for structure labeling, while for nuclearity and relation labeling, we obtain 79.8% and 78.3%, respectively, of the human f-score. marcu (2000) calculated reaching 25.7% of the human performance level for relation labeling, while le thanh et al. (2004b) calculated reaching 39.9% of this level. direct comparison to our system is only possible in the case of sagae (2009), as sagae employed the same 18 standard rst-dt relations, and evaluation was done on the rst-dt’s test subset. a performance comparison of both systems is presented in table 14. the scores are for full discourse tree creation. we observe that the performance of the two parsers is relatively similar. in particular, recall is 46.2% in sag, while we obtained 46.9% in hilda. precision for hilda is however 11% higher than sag, at 47.7%, against 42.9% for sagae’s system. 19 hernault, prendinger, duverle, and ishizuka table 11: performance for the parser of marcu (2000), using 15 discourse relations, evaluated on unspecified documents structure nuclearity relations precision 65.8 54.0 34.3 recall 34.0 21.6 13.0 f-score 44.8 30.9 18.8 falgo/fhuman 56.0 42.9 25.7 table 12: performance for the parser of le thanh et al. (2004b), using 14 discourse relations, evaluated on 21 documents from the rst-dt structure nuclearity relations precision 54.5 47.8 40.5 recall 52.9 46.4 39.3 f-score 53.7 47.1 39.9 falgo/fhuman 73.9 71.8 70.1 4.5 time performance to complete the discussion of the parser’s actual performance, we measure the time taken to parse standard texts from the test section of the rst-dt. the computer used in the tests has a 2.53 ghz processor and 4 gb of ram. on average, loading the stanford syntax parser and svm models altogether takes four seconds. this is done only once. we use modified command-line versions of libsvm and liblinear to accept data from stdin, so that loading of models is performed only once. after this, we can perform fast classification decisions, taking a few tenths of a second each. table 15 shows the performance of the parser on various texts. tseg is the time taken to segment the text, and tbuild the time taken to build the tree. as expected, we empirically observe that tseg and tbuild follow a o(n) time complexity. however, in the system, the only process not under our control is the syntax parser. it is important to ensure that the time taken to generate parse trees is reasonable. on average, on our test computer, the stanford parser takes 10 seconds to parse 300 words. of the three processes (syntax parsing, discourse segmentation, tree building), syntax parsing is always the slowest task, and thus the bottleneck of the system. still, this level of performance makes the parser usable as part of an interactive system. 20 hilda: a discourse parser using support vector machine classification table 13: performance for hilda, using 18 discourse relations, evaluated on 21 documents from the rst-dt (lethanh’s) structure nuclearity relations precision 76.0 61.4 51.2 recall 75.6 61.2 50.6 f-score 75.8 61.3 50.9 falgo/fhuman 86.4 79.8 78.3 table 14: performance comparison between the proposed parser (hilda) and the parser of sagae (2009) (sag), for full discourse tree creation, using 18 discourse relations, evaluated on the rst-dt’s test set sag hilda precision 42.9 47.7 recall 46.2 46.9 f-score 44.5 47.3 4.6 limitations in this section, we present some limitations of the parser and discuss possible improvements and directions for future work. when evaluating the proposed discourse segmenter on the rst-dt’s test set, we observe that all segmentation errors are due to over-segmentation, i.e. words which are not edu boundaries are mistaken for boundaries. table 16 shows the ten most frequent words on which segmentation errors occur, as well as how prevalent these errors are among all segmentation errors. several of these mistakes are linked to punctuation, particularly to quotes and dash marks. also, several mistakes seem related to the segmenter creating excessively small clause-like units. typical improper seg21 hernault, prendinger, duverle, and ishizuka table 15: time taken to parse various texts (in seconds) text sentences words tseg (s) tbuild (s) wsj0616 36 838 5.02 11.30 wsj1113 5 102 0.40 0.51 wsj1129 2 45 0.26 0.40 wsj2336 10 262 1.71 3.72 wsj2385 50 516 3.05 7.17 mentation decisions (marked ↑) related to these words are seen in the following sentences, where the true segmentation decisions (when present) are marked ⇑: ‘‘ [...] the banco exterior group has a lot ↑ to offer a potential suitor.’’ ‘‘ [...] when the market does recover, ↑ the damage is done [...]’’ ‘‘ [...] cut costs, increase capital ↑ and build new areas of business [...]’’ ‘‘ [...] and perhaps even destroy ↑ -⇑ a $2.38 settlement fund [...]’’ ‘‘ in august, ↑ mr. lewis pleaded guilty to three felony counts.’’ ‘‘ when it misses one month ↑ it tends to miss the next month.’’ table 16: words on which erroneous segmentation decisions are most frequent, and their proportion among all segmentation errors on the rst-dt’s test set word pos prevalence among errors (%) to to 7.7 the dt 6.6 and cc 4.6 “ “ 3.4 – : 2.3 mr. nnp 2.3 it prp 2.1 for in 2.0 if in 1.8 said vbd 1.6 22 hilda: a discourse parser using support vector machine classification although a quantitative evaluation of the discourse trees produced by our parser is made possible by the usage of metrics traditionally employed for evaluating the performance of syntax parsers, we found that a qualitative evaluation of the generated discourse trees was difficult. in particular, we hypothesize that finding patterns of errors commonly present in the generated discourse trees is made difficult by the greedy approach adopted. because of the proposed algorithm and the way the classifiers are cascaded, classification errors occurring anywhere when building the discourse tree will affect both sibling nodes and higher level relations of the tree, making the real cause of errors difficult to identify. indeed, the presented greedy algorithm is time-efficient, but provides only locally optimal solutions. using generic global probabilistic optimization meta-algorithms such as simulated annealing (kirkpatrick et al., 1983), we hope to address the problem of local optimality while maintaining reasonable time complexity. secondly, we plan to investigate the use of sequential labeling methods (lafferty et al., 2001) for finding optimal sequences of discourse relations. another related shortcoming of the current algorithm, as presented in figure 6, occurs when two pairs of edus have an equal probability of being connected by a discourse relation. in the present system, the default behavior in case of conflict is to select the first pair of units in text order. however, ideally, all candidates should be considered, with their respective trees built independently, somehow scored, and the best candidate being finally returned. we hope to address this point by employing the aforementioned optimization methods. in section 4.2, we noticed important discrepancies in f-scores (see table 7) for the different classes of label. we hypothesize that this issue is mainly related to the nature of discourse relations. indeed, a notoriously challenging task for discourse parsers is to detect ‘implicit relations’, i.e. relations not signaled by discourse cues, such as ‘thus’, ‘however’, ‘but’, etc. certain relations, such as attribution, are invariably realized with a discourse cue, indicated by the verb ‘say’, such as ‘x said that y’. this relation is easy to detect, and we empirically observe high performance on attribution’s classes (see table 7). on the other hand, an implicit contrast relation holds between the two following discourse units: [mr. roman comes across as a low-key executive;]2a [mr. phillips has a flashier personality.]2b in this example, the contrast is captured by the antonymy between the adjectives low-key and flashy. as there is no cue, it is harder to detect, and we observe in practice a lower performance for contrast relations. taboada (2006) estimates that 60% to 70% of naturally-occurring discourse relations are implicit. in general, while explicit discourse relations are well detected (pitler et al., 2008), implicit relations are challenging to detect. for instance, we fed the proposed parser ten short sentences containing an implicit contrast relation. all relations were wrongly classified, either as elaboration in six cases or as joint in four cases. similar errors were seen when trying to classify implicit comparison and temporal relations. this indicates that our set of features is not fully-adapted for implicit relations. recent evidence (webber, 2009) even suggests that the features present when an explicit discourse relation is realized differ significantly from those present in the case of an implicit relation. a possible solution is to train a separate discourse relation classifier for implicit discourse relations. a different set of features will be required. for this task, several researchers (marcu and echihabi, 2002; pitler et al., 2009; lin et al., 2009) have noted relatively good performance of word pair features taken between the relation’s two argument 23 hernault, prendinger, duverle, and ishizuka spans. for instance, using the same example, we would encode lemmatized word pairs such as (mr., mr.), (mr., phillips), . . . , (low-key, flashy), . . . , (executive, personality), etc. however, word pair features alone are not sufficient to obtain high performance rates. we plan to investigate implicit discourse relation classification by employing features based on measures of semantic similarity between edus, using for instance wordnet (fellbaum, 1998). finally, we plan to apply feature selection in order to effectively reduce the size of the feature set. this method has the promise of shorter training and testing times. another interesting aspect of feature selection is the potential reduction of noise, which may lead to an improvement in accuracy. 5. conclusions we presented hilda, an automated discourse parser that analyzes text in the framework of the rhetorical structure theory. the system is composed of two modules, (1) a discourse segmenter, and (2) a relation classifier and tree builder. both tasks are based on support vector machine classification. the discourse segmenter performs at around 96% of the human level, while the relation classification and tree-building module reaches 84% of human level. our system combines both modules for creating a fully automatic text level parser, reaching 78.3% of the human performance level. importantly, the tree-building process performs in linear time, which allows us to use the system in the context of a (near) real-time or interactive system, such as dialogue generation or questionanswering. for instance, the proposed parser can be integrated in a dialogue generation system such as text-to-dialogue (prendinger et al., 2007; hernault et al., 2008). in this system, rst structures are used as the backbone for generating dialogues between two agents representing a layperson and an expert. question-answer pairs between the two agents are created by applying relation-specific mapping rules manipulating the leaves of the rst tree. if the rst tree utilized contains errors, such as improperly-segmented edus or erroneous discourse relations, the generated dialogue will be semantically incorrect. hence, a sound discourse parser is important. more recently, the coda corpus (stoyanchev and piwek, 2010) was released. this corpus contains 700 expository dialogues labeled with dialogue acts, paired with human-authored, equivalent monologues annotated with rst discourse relations. using this corpus, one can learn how to generate an expository dialogue from an rst-annotated monologue. an ambitious next step is to automatize the generation process, i.e. creating a meaningful expository dialogue given any plain, non-annotated monologue. this is a promising perspective for creating systems that are able to present information automatically, such as tutoring applications. a prerequisite is the accurate detection of the discourse relations in the input monologues, and hence, the role of the discourse parser is once again central. with these applications in mind, we plan to make the proposed parser more efficient, by addressing the issues presented in section 4.6. in particular, we intend to improve the tree-building method, in order to create globally optimal discourse structures. finally, we plan to train a separate classifier for detecting implicit discourse relations as well. 24 hilda: a discourse parser using support vector machine classification appendix a. system workflow test document syntax parsing (stanford parser) syntax trees lexicalization lexicalized syntax trees segmentation feature extraction svm classification tokenized edus svm classification relation labeling feature extraction scored rs subtrees rhetorical structure tree bottom-up tree construction rst discourse treebank penn treebank edus lexicalization lexicalized syntax trees syntax trees alignment segmentation feature extraction svm training tokenization tokenized edus relation labeling feature extraction svm training segmentation modelrelation labeling model alignment 25 hernault, prendinger, duverle, and ishizuka appendix b. relation list relation nuclearity options attribution nucleus satellite satellite nucleus background nucleus satellite satellite nucleus cause nucleus nucleus nucleus satellite satellite nucleus comparison nucleus nucleus nucleus satellite satellite nucleus condition nucleus nucleus nucleus satellite satellite nucleus contrast nucleus nucleus nucleus satellite satellite nucleus elaboration nucleus satellite satellite nucleus enablement nucleus satellite satellite nucleus evaluation nucleus nucleus nucleus satellite satellite nucleus explanation nucleus nucleus nucleus satellite satellite nucleus joint nucleus nucleus manner-means nucleus satellite satellite nucleus summary nucleus satellite satellite nucleus temporal nucleus nucleus nucleus satellite satellite nucleus topic-change nucleus nucleus nucleus satellite topic-comment nucleus nucleus nucleus satellite satellite nucleus same-unit nucleus nucleus textual-organization nucleus nucleus 26 hilda: a discourse parser using support vector machine classification references nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. jason baldridge and alex lascarides. probabilistic head-driven parsing for discourse structure. in proceedings of the ninth conference on computational natural language learning (conll2005), pages 96–103, ann arbor, michigan, june 2005. association for computational linguistics. ezra black, steven p. abney, d. flickenger, claudia gdaniec, ralph grishman, p. harrison, donald hindle, robert ingria, frederick jelinek, judith l. klavans, mark liberman, mitchell p. marcus, salim roukos, beatrice santorini, and tomek strzalkowski. a procedure for quantitatively comparing the syntactic coverage of english grammars. in proceedings of workshop on speech and natural language, pages 306–311. association for computational linguistics morristown, nj, usa, 1991. leo breiman. bagging predictors. machine learning, 24(2):123–140, 1996. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. proceedings of second sigdial workshop on discourse and dialogue-volume 16, pages 1–10, 2001. joyce y. chai and rong jin. discourse structure for context question answering. in sanda harabagiu and finley lacatusu, editors, hlt-naacl 2004: workshop on pragmatics of question answering, pages 23–30, boston, massachusetts, usa, may 2 may 7 2004. association for computational linguistics. eugene charniak. a maximum-entropy-inspired parser. in proceedings of the 1st north american chapter of the association for computational linguistics conference, pages 132–139, san francisco, ca, usa, 2000. morgan kaufmann publishers inc. degang chen, qiang he, and xizhao wang. on linear separability of data sets in feature space. neurocomputing, 70(13-15):2441–2448, 2007. david duverle and helmut prendinger. a novel discourse parser based on support vector machine classification. in proceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural language processing of the afnlp, pages 665–673, suntec, singapore, august 2009. association for computational linguistics. christiane fellbaum, editor. wordnet: an electronic lexical database. mit press, 1998. hugo hernault, paul piwek, helmut prendinger, and mitsuru ishizuka. generating dialogues for virtual agents using nested textual coherence relations. in iva ’08: proceedings of the 8th international conference on intelligent virtual agents, pages 139–145, berlin, heidelberg, 2008. springer-verlag. scott kirkpatrick, c. daniel gelatt, and mario p. vecchi. optimization by simulated annealing. science, 220(4598):671–680, 1983. 27 hernault, prendinger, duverle, and ishizuka dan klein and christopher d. manning. fast exact inference with a factored model for natural language parsing. in advances in neural information processing systems, volume 15, pages 3–10. mit press, 2003. john d. lafferty, andrew mccallum, and fernando c. n. pereira. conditional random fields: probabilistic models for segmenting and labeling sequence data. in icml’01: proceedings of the eighteenth international conference on machine learning, pages 282–289, san francisco, ca, usa, 2001. morgan kaufmann publishers inc. huong le thanh, geetha abeysinghe, and christian huyck. automated discourse segmentation by syntactic information and cue phrases. in proceedings of aia’04, innsbruck, austria, february 16–18 2004a. huong le thanh, geetha abeysinghe, and christian huyck. generating discourse structures for written texts. in proceedings of the 20th international conference on computational linguistics, pages 329–335, geneva, switzerland, aug 23–aug 27 2004b. coling. ziheng lin, min-yen kan, and hwee tou ng. recognizing implicit discourse relations in the penn discourse treebank. in proceedings of the 2009 conference on empirical methods in natural language processing, pages 343–351, singapore, august 2009. association for computational linguistics. david m. magerman. statistical decision-tree models for parsing. proceedings of the 33rd annual meeting on association for computational linguistics, pages 276–283, 1995. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text, 8(3):243–281, 1988. christopher d. manning and hinrich schütze. foundations of statistical natural language processing. mit press, 1999. daniel marcu. building up rhetorical structure trees. proceedings of the national conference on artificial intelligence, pages 1069–1074, 1996. daniel marcu. the theory and practice of discourse parsing and summarization. mit press, 2000. daniel marcu and abdessamad echihabi. an unsupervised approach to recognizing discourse relations. in proceedings of 40th annual meeting of the association for computational linguistics, pages 368–375, philadelphia, pennsylvania, usa, july 2002. association for computational linguistics. doi: 10.3115/1073083.1073145. mitchell p. marcus, mary ann marcinkiewicz, and beatrice santorini. building a large annotated corpus of english: the penn treebank. computational linguistics, 19(2):313–330, 1993. johanna d. moore and peter wiemer-hastings. discourse in computational linguistics and artificial intelligence. in a. graesser, m. gernsbacher, and s. goldman, editors, handbook of discourse processes, pages 439–486. erlbaum, mahwah, nj, 2003. 28 hilda: a discourse parser using support vector machine classification jon oberlander and johanna d. moore. cue phrases in discourse: further evidence for the core:contributor distinction. in workshop on levels of representation in discourse, pages 87–93, edinburgh, uk, 1999. emily pitler, mridhula raghupathy, hena mehta, ani nenkova, alan lee, and aravind joshi. easily identifiable discourse relations. in coling 2008: companion volume: posters, pages 87–90, manchester, uk, august 2008. coling 2008 organizing committee. emily pitler, annie louis, and ani nenkova. automatic sense prediction for implicit discourse relations in text. in proceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural language processing of the afnlp, pages 683–691, suntec, singapore, august 2009. association for computational linguistics. helmut prendinger, paul piwek, and mitsuru ishizuka. a novel method for automatically generating multi-modal dialogue from text. international journal of semantic computing, 1(3):319–334, 2007. david reitter. rhetorical analysis with rich-feature support vector models. unpublished master’s thesis, university of potsdam, potsdam, germany, 2003a. david reitter. simple signals for complex rhetorics: on rhetorical analysis with rich-feature support vector models. ldv forum, 18(1/2):38–52, 2003b. kenji sagae. analysis of discourse structure with syntactic dependencies and data-driven shiftreduce parsing. in proceedings of the 11th international conference on parsing technologies (iwpt’09), pages 81–84, paris, france, october 2009. association for computational linguistics. frank schilder. robust discourse parsing via discourse markers, topicality and position. natural language engineering, 8(2-3):235–255, 2002. radu soricut and daniel marcu. sentence level discourse parsing using syntactic and lexical information. proceedings of the 2003 conference of the north american chapter of the association for computational linguistics on human language technology, 1:149–156, 2003. carl staelin. parameter selection for support vector machines. hewlett-packard company, tech. rep. hpl-2002-354r1, 2003. svetlana stoyanchev and paul piwek. constructing the coda corpus: a parallel corpus of monologues and expository dialogues. in proceedings of the seventh conference on international language resources and evaluation (lrec’10), valletta, malta, may 2010. european language resources association. rajen subba and barbara di eugenio. automatic discourse segmentation using neural networks. in proceedings of 11th workshop on the semantics and pragmatics of dialogue, pages 189–190, trento, italy, 2007. rajen subba and barbara di eugenio. an effective discourse parser that uses rich linguistic information. in proceedings of human language technologies: the 2009 annual conference of 29 hernault, prendinger, duverle, and ishizuka the north american chapter of the association for computational linguistics, pages 566–574, boulder, colorado, june 2009. association for computational linguistics. maite taboada. discourse markers as signals (or not) of rhetorical relations. journal of pragmatics, 38(4):567–592, 2006. milan tofiloski, julian brooke, and maite taboada. a syntactic and lexical-based discourse segmenter. in acl ’09, pages 77–80, suntec, singapore, august 2009. association for computational linguistics. vladimir n. vapnik. the nature of statistical learning theory. springer-verlag new york, inc., new york, ny, usa, 1995. bonnie webber. genre distinctions for discourse in the penn treebank. in proceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural language processing of the afnlp, pages 674–682, suntec, singapore, august 2009. association for computational linguistics. 30 hilda: a discourse parser using support vector machine classification s np nn fa rm nn le nd in g vp vb d w as vp vb n en ac te d s vp to to vp vb co rre ct np dt th is nn pr ob le m pp in by s vp vb g pr ov id in g np np dt a jj re lia bl e nn flo w pp in of np jj le nd ab le nn s fu nd s . . (e na ct ed ) (e na ct ed ) (c or re ct ) (w as ) (w as ) (p ro bl em ) (c or re ct ) (b y) fi gu re 7: pa rt ia lly -l ex ic al iz ed sy nt ax tr ee (l ex ic al he ad s ar e in di ca te d be tw ee n pa re nt he se s) 31 hernault, prendinger, duverle, and ishizuka manner-means enablement 1a 1b 1c enablement 1a manner-means 1b 1c figure 8: two possible attachments in the rs-tree of the text in figure 1. 32 hilda: a discourse parser using support vector machine classification s np nn fa rm nn le nd in g vp vb d w as vp vb n en ac te d s vp to to vp vb co rre ct np dt th is nn pr ob le m pp in by s vp vb g pr ov id in g np np dt a jj re lia bl e nn flo w pp in of np jj le nd ab le nn s fu nd s . . (e na ct ed ) (c or re ct ) (w as ) (w as ) (c or re ct ) (b y) 1a 1b 1c fi gu re 9: pa rt ia lly -l ex ic al iz ed sy nt ax tr ee ,s ho w in g do m in an ce se ts 33 rusetal_editorproof-02072012 dialogue and discourse 3(2) (2012) 177–204 doi: 0.5087/dad.2012.208 ©2012 vasile rus, brendan wyse, et al. submitted 3/11; accepted 1/12; published online 3/12 a detailed account of the first question generation shared task evaluation challenge vasile rus vrus@memphis.edu department of computer science the university of memphis usa brendan wyse bjwyse@gmail.com aol, dublin ireland paul piwek p.piwek@open.ac.uk centre for research in computing open university, uk mihai lintean mclinten@memphis.edu department of computer science the university of memphis memphis, tn 38152 usa svetlana stoyanchev s.stoyanchev@open.ac.uk centre for research in computing open university, uk cristian moldovan cmldovan@memphis.edu department of computer science the university of memphis memphis, tn 38152 usa editor:kristy elizabeth boyer abstract the paper provides a detailed account of the first shared task evaluation challenge on question generation that took place in 2010. the campaign included two tasks that take text as input and produce text, i.e. questions, as output: task a – question generation from paragraphs and task b – question generation from sentences. motivation, data sets, evaluation criteria, guidelines for judges, and results are presented for the two tasks. lessons learned and advice for future question generation shared task evaluation challenges (qg-stec) are also offered. keywords: question generation, shared task evaluation campaign. rus, wyse, et al. 178 1 introduction question generation is an essential component of learning environments, help systems, information seeking systems, multi-modal conversations between virtual agents, and a myriad of other applications (lauer, peacock, and graesser, 1992; piwek et al, 2007). question generation has been recently defined as the task of automatically generating questions from some form of input (rus & graesser, 2009). the input could vary from information in a database to a deep semantic representation to raw text. question generation is viewed as a three-step process (rus & graesser, 2009): content selection, selection of question type (who, why, yes/no, etc.), and question construction. content selection is about deciding what the question should be about given the various inputs available to the system and the context in which the question is being asked. in other words, this step focuses on deciding what is important to ask about in a particular context. the step of question type selection decides the most appropriate type of the question given the selected content and context. that is, this step focuses on how to ask the question. for instance, a how or why question type may be selected or even a yes/no type of question. the last step, question construction, is about realizing the question given the selected content, question type, and the context in which the question is being asked. this step too focuses on how to ask the question. a fourth step has been proposed by mostow and chen (2009) which refers to the decision of when to ask a question in a particular context. a question has its most desired effect when asked at the right moment. while finding the right moment seems to be an open research challenge at this moment, preliminary work has noted that a question should not be asked “too soon after the prior one” in the context of reading comprehension assessment or scaffolding (beck, mostow, & bey, 2004; mostow et al., 2004). question generation (qg) is primarily a dialogue and discourse task. it draws on and overlaps with research in both natural language understanding (nlu) and natural language generation (nlg). for example, qg from raw text, e.g. a textbook or the dialogue history between a tutor and tutee, will involve some sort of analysis, i.e., nlu, of the input to select content, detect patterns, etc. the construction of the output question will typically involve nlg: a representation of the input is mapped to a sentence. qg from non-linguistics input knowledge representations (e.g., semantic web ontology languages), can be viewed directly as a typical nlg task that takes as input non-linguistic data and produces output in a human language (reiter & dale, 1997; reiter & dale, 2000). qg as a focused area of research is relatively young. it basically started in 2008 with the national science foundation-sponsored workshop on the question generation shared task and evaluation challenge1 (rus & graesser, 2009). sporadic research on question generation did happen before (wolfe, 1976; kunichika et al., 2001; mitkov, ha, & karamanis, 2006; rus, cai, & graesser, 2007a). the qg research community has decided to offer shared task evaluation campaigns or challenges (stecs) as a way to stimulate research and bring the community closer. the first question generation shared task evaluation challenge (qg-stec) follows a long tradition of stecs in natural language processing: see various tracks at the text retrieval conference2 (trec), e.g. the question answering track (voorhees & tice, 2000), the semantic evaluation challenges under the senseval3 umbrella (edmonds, 2002), or the annual tasks run by the conference on natural language learning4 (conll). in particular, the idea of a qgstec was inspired by the recent activity in the natural language generation (nlg) community to offer shared task evaluation campaigns as a potential avenue to provide a focus for research in 1 www.questiongeneration.org 2 http://trec.nist.gov 3 www.senseval.org 4 http://www.cnts.ua.ac.be/conll/ the first question generation shared task evaluation challenge 179 nlg and to increase the visibility of nlg in the wider natural language processing (nlp) community (dale and white, 2007). the nlg community has offered under the umbrella of generation challenges5 several shared tasks on language generation including generation of natural-language instructions to aid human task-solving in a virtual environment (the give-2 challenge6; koller et al., 2010) and post-processing referring expressions in extractive summaries (grec’10; belz et al., 2008; belz & kow, 2010). in designing the first qg-stec, we had to balance conceptual and pragmatic issues. two core aspects of a question are the goal of the question and its importance. it is difficult to determine whether a particular question is good without knowing the context in which it is posed; ideally one would like to have information about what counts as important and what the goals are in the current context. this suggests that a stec on qg should be tied to a particular application, e.g. tutoring systems. however, an application-specific stec would limit the pool of potential participants to those interested in the target application. therefore, the challenge was to find a framework in which the goal and importance are intrinsic to the source of questions and less tied to a particular context/application. one possibility was to have the general goal of asking questions about salient items in a source of information, e.g. core ideas in a paragraph of text. our task a (described later) has been defined with this concept in mind. this basic idea does not really hold for generation of questions from single sentences, which is our task b (defined later). for task b, we have basically decided to ignore the importance of the question and chose to accept all questions as long as they fit some other minimal criteria for a good question (fluency, ambiguity, relevance), which are application independent. ignoring the importance of the questions for task b seemed appropriate given that the research on this type of systems is still in a very early phase. task b focused more on evaluating the capacity of systems to construct questions rather than select content for asking questions about. adopting the basic principle of application-independence has the advantage of escaping the problem of a limited pool of participants (to those interested in a particular application had that application been chosen as the target for a qg-stec). besides the advantage of a larger pool of potential participants, an application-independent qg-stec would provide a more fair ground for comparison as teams already working on a certain application would not be advantaged as would be the case if the application had been the focus of a stec. it should be noted that the idea of an application-independent stec is not new. an example of an application-independent stec would be generic summaries (as opposed to query-specific summaries) in summarization. another decision aimed at attracting as many participants as possible and promoting a more fair comparison environment concerned the input for the qg tasks. a particular semantic representation would have provided an advantage to groups already working with it and at the same time it would have raised the barrier-to-entry for newcomers. instead, we have adopted a second guiding principle for the first qg-stec tasks: no representational commitment. that is, we wanted to have as generic an input as possible. therefore, the input to both task a and b in the first qg-stec was raw text. that is, the first qg-stec falls in the wider category of textto-questions tasks identified at the first workshop on question generation (rus & graesser, 2009). also, it can be viewed as a text-to-text tasks as proposed by rus and colleagues (2007b). task a and b offered in the first qg-stec fall in the text-to-question category of qg tasks identified by the first workshop on question generation (www.questiongeneration.org). the first workshop identified four categories of qg tasks (rus & graesser, 2009): text-to-question, tutorial dialogue, assessment, and query-to-question. using another categorization, tasks a 5 http://www.itri.brighton.ac.uk/research/genchal10/ 6 http://www.give-challenge.org/ rus, wyse, et al. 180 and b are part of the text-to-text natural language generation task category identified by the natural language generation community (rus et al., 2007b). there was overlap between tasks a and b in the first qg-stec. this was intentional with the aim of encouraging people preferring one task to also participate in the other. in particular, generating questions from individual sentences (task b) was also included as one goal among several others in task a (where the input consisted of paragraphs). we opted for human-based evaluation for the first qg-stec. evaluation in question generation is closest to evaluations in natural language generation, machine translation, and summarization, as the outputs in all these areas correspond to texts in a human language, similar to the question generated shared tasks in the first qg-stec. in machine translation and summarization, the input is also text, as for the tasks in the first qg-stec. while in machine translation and summarization automatic scoring procedures are nowadays acceptable as they have been shown to correlate with human ratings (papineni et al., 2002; lin, 2004; lin and och, 2004), in natural language generation human-based evaluations are the current norm (walker, owen, & rogati, 2002). for the time being, we have opted for the conservative option of relying on human judgment for both conceptual and pragmatic reasons. as in natural language generation, for a given input text there is in theory a large number of questions that can be asked about. in other words, there is no uniquely and clearly defined correct answer for question generation. additionally, given the early stage of research in question generation, the safe approach to evaluation would be the human-based approach. automatic scoring is an interesting topic of future research for which the first qg-stec can be used as a testbed. for certain types of questions, automatic scoring is possible based on extrinsic criteria. for instance, the quality of the automatically generated multiple-choice questions by mitkov, ha, and karamanis (2006) were evaluated using item test theory. this scoring resembles the task-based (extrinsic) form of evaluation used in natural language generations in some cases. furthermore, at the time of planning the first qg-stec there was only one previous work, to the best of our knowledge, that evaluated questions of the type we focused on, i.e. non multiple choice questions, generated from raw text (rus, cai, & graesser, 2007a). in line with most existing evaluation schemes, for the qg-stec the focus was on measuring the quality of the output (i.e., the generated questions) along several dimensions. such an absolute measure does, however, not take into account the quality of the input. ideally, especially for criteria such as fluency, one may want to measure to what extent the score for the output question is an increase or decrease relative to the score for the input text, e.g., is the fluency of the output question better or worse than that of the input text and how much better or worse? such an approach measuring relative quality of outputs has been proposed and applied by piwek & stoyanchev (2011) to the task of generating short fragments of dialogue (e.g., question-answer pairs) from expository monologue. overall, we had one submission for task a and four submissions for task b. the submissions were evaluated through a peer-review system for task b. task a was evaluated by two external judges as there was only one submission and the peer-review mechanism could not be applied. there are several explanations for the relatively small number of participants. first, there are no well-established research programs in this area which would allow researchers to dedicate their time to develop competitive qg systems and participate in qg-stecs. this argument is further supported by the discrepancy between the number of research groups who expressed interest in participating and those who really submitted results for evaluation. in particular for task a, the discrepancy was quite significant: five teams expressed explicit interest to participate (as required by the qg-stec organizers, but only one team submitted results for evaluation. in a way, task a is more difficult than task b (it requires discourse level processing), which explains the lower number of submitted teams, although the number of interested teams was comparable. second, being the first qg-stec ever offered, the effort necessary to competitively participate was substantial. in subsequent qg-stecs, participating teams will be able to build on the tools and insight generated either by themselves or other teams that the first question generation shared task evaluation challenge 181 participated in the first qg-stec. third, in future qg-stecs publicity for the stec could be improved, perhaps by aligning with the generation challenges initiative or another mainstream shared task initiative, and by allowing for more time between release of the stec instructions/development data and the release of test data. all the data together with detailed descriptions of the two tasks and the papers describing the approaches proposed by the participants are accessible from the main question generation website (see footnote 1). it is important to say that the two tasks offered in the first qg-stec were selected among five candidate tasks by the members of the qg community. a preference poll was conducted and the most preferred tasks, question generation from paragraphs (task a) and question generation from sentences (task b), were chosen to be offered in the first qg-stec. the other three candidate tasks were: ranking automatically generated questions (michael heilman and noah smith), concept identification and ordering (rodney nielsen and lee becker), and question type identification (vasile rus and arthur graesser). the involvement of the qg community at large contributed to quality of tasks offered and ultimately the success of the first qg-stec. 2 task a: question generation from paragraphs in this section, we present the details of task a, question generation from paragraphs, together with the results of the participating system. all the data and guidelines are accessible from the question generation wiki7. 2.1 task definition the question generation from paragraphs (qgp) task challenged participants to generate a list of questions from a given input paragraph. the questions should be at three specificity/scope levels: general/broad (triggered by the entire input paragraph), medium (one or more clauses or sentences), and specific (phrase level or less). participants were asked to generate one general question, two medium questions, and three specific questions for a total of six questions per input paragraph. the specificity/scope was defined by the portion of the paragraph that answered the question. if multiple questions could be generated at one level, only the specified number should be submitted. for example, for a paragraph that answers two broad questions, only one question at that level of specificity should be submitted. for the qgp task, questions are considered important and interesting if they ask about the core idea(s) in the paragraph and an average person reading the paragraph would consider them so, based on a quick analysis of the contents of the paragraph. simple, trivial questions such as what is x? or generic questions such as what is the paragraph about? were to be avoided. in addition, implied questions (whose answer is not explicitly stated in the paragraph) were not allowed as the emphasis was on questions triggered and answered by the paragraph. furthermore, questions could not be compounded as in what is … and who … ? questions had to be grammatically and semantically correct and related to the topic of the given input paragraph. question types (who/what/why/…) generated for each paragraph should be diverse, i.e. different question types are preferred in the set of 6 required questions. 2.2 guidelines for human judges we next show an example paragraph together with six interesting, application-independent questions that could be generated. the paragraph is about two-handed backhands in tennis and was collected from wikipedia. table 1 shows the paragraph while table 2 shows six questions triggered by this paragraph. each question is at a different level of specificity. the first question 7 http://www.questiongeneration.org/mediawiki/index.php/qg-stec_2010 rus, wyse, et al. 182 is general/broad because its answer is the entire paragraph. one may argue that the first sentence in the paragraph, which usually is the topic sentence summarizing the paragraph, also forms a valid answer to the question. this is true in general. for this reason, we instructed the judges to consider the widest scope possible when judging a question and also to judge the question within the scope indicated by participants. in other words, if the participants selected the entire paragraph as opposed to just the topic sentence as the content that triggered the question then judges should rate the question scope within the indicated scope: does the question fit the participant-selected scope? table 1. example of input paragraph (from http://en.wikipedia.org/wiki/backhand) together with the answers to the questions shown in table 2. input paragraph two-handed backhands have some important advantages over one-handed backhands. two-handed backhands are generally more accurate because by having two hands on the racquet, this makes it easier to inflict topspintopspin on the ball allowing for more control of the shot. two-handed backhands are easier to hit for most high ballshigh balls . two-handed backhands can be hit with an open stancean open stance , whereas one-handers usually have to have a closed stance, which adds further steps (which is a problem at higher levels of play). table 2. examples of questions and scores for the paragraph in table 1. the fragment in between [] is optional. questions scope what are the advantages of two-handed backhands in tennis? general why are two-handed backhands more accurate [when compared to one-handers]? medium what is one consequence of inflicting topspin on a tennis ball? medium what kind of spin does a two-handed backhand inflict on the ball? specific what stance is needed to hit a two-handed backhand? specific what types of balls are easier to hit with a two-handed backhand? specific the next two questions in table 2 are medium-scope questions as their answers are entire sentences (see the underlined and bold sentences, respectively, in the input paragraph shown in table 1 which form the answers to the two questions). the bottom three questions in table 2 are specific questions whose answers are a phrase or less in length. the phrase answers to the example specific questions are shown in bold face in the input paragraph in table 1. the three specificity levels proposed for task a were consciously chosen by the proposing team as explained next. the specific questions were inspired by the factoid questions used for shared tasks by the question answering community (voorhees & tice, 2000). the answer to these factoid questions were usually short snippets of text in the form of a phrase or even less, e.g. the name of a person such as president obama. examples of factoids questions used by the question answering stecs are who is the voice of miss piggy?, whose answer is “frank oz”, and how much could you rent a volkswagen bug for in 1966? whose answer is “$1/day”. by focusing on specific facts in the input paragraph, qg-stec participants could be challenged to generate specific questions. additionally, these questions corresponded closely to the questions the first question generation shared task evaluation challenge 183 required for task b (on qg from sentences). the medium-scope questions were inspired by discourse-level relations which usually hold among larger chunks of texts in the input paragraph (see piwek et al., 2007 for automatic generation of such questions in expository dialogue). as an example, we use the cause-effect relation between the underlined versus the simple italic text portion of the sentence “two-handed backhands are generally more accurate because by having two hands on the racquet, this makes it easier to inflict topspin on the ball allowing for more control of the shot.” (from the input paragraph in table 1). the discourse relations could be used to generate questions where one fragment in the discourse relation forms the body of the question while the other fragment forms the answer to the question, i.e. the target content triggering the question. the cause-effect relation example provided above can be used to trigger a why question whose body is the effect portion of the relation (why are two-handed backhands more accurate?) while the answer is the cause portion of the relation: “because by having two hands on the racquet, this makes it easier to inflict topspin on the ball allowing for more control of the shot.” the general/broad scope questions could be generated by doing a global analysis of the information content of the paragraph based on which the more salient concepts could be used to trigger such general/broad questions whose answer is the entire paragraph. we now turn to describing the evaluation criteria used to judge submitted questions. these criteria were used by human raters to judge the questions. we will use the paragraph and questions in tables 1 and 2 to illustrate the judging criteria. a set of five scores, one for each of the following criteria was generated for each question. • specificity • syntax • semantics • question type correctness • diversity additionally, we anticipated creating a composite score based on the scores for the individual criteria. this score would range from 1 (first/top ranked, best) to 4 (lowest rank), with 1 meaning that the best possible score was achieved on each of the individual criteria. however, deciding on a valid combination of the individual scores is extremely difficult (e.g., are the criteria all of the same weight?), whilst it would force a single ranking on the performance of participating systems. for these reasons, we decided not to calculate such a composite score, but keep this on the agenda as an issue for further discussion in the run-up to the next qg-stec for task a. the specificity scores are assigned primarily based on the answer span in the input paragraph. the broadest question is the one whose answer spans the entire paragraph. the most specific question is the one whose answer is less than a sentence: a clause, phrase, word, or collocation. scores were assigned based on the following rubric: 1 – input paragraph, 2 – multiple sentences, 3 – a clause or less, 4 – trivial/generic, implied, no question (empty question), or undecided, e.g. a semantically wrong question which may not be understood well enough to judge its scope. as we expected six questions as output, if one level is missed we encouraged participants to generate questions of a narrower scope. for instance, if a broad-scope question could not be generated then a medium or specific question should have been submitted. this assured that participants submitted as many as six questions for each input paragraph. the best question specificity scores for six questions would be 1, 2, 2, 3, 3, 3. this best configuration of scores would be possible for paragraphs that could trigger the required number of questions at each scope level, which may not always be the case. in preparing the data for the qg-stec, we aimed at selecting paragraphs that made it possible to obtain perfect scope scores. while the initial plan was for the judges to look at the question itself and select themselves the portion of the paragraph that may have triggered the question, we opted instead, for practical reasons, to allow the judges to see the span of text submitted by participants for each question and decide based on the participant-submitted span the specificity of the question. the advantage of rus, wyse, et al. 184 the initial plan is that the judges’ selected text span could have been automatically compared to the span submitted by participants for an automated scoring process of the scope criterion. however, this initial plan proved to be more challenging because if judges were allowed to select their own answer span without seeing the one provided by participants then that may have undesired effects on human-judging of the questions on the other criteria, e.g. question type correctness or semantic correctness. for instance, it may be the case that a judge selects a slightly different span for the question than the one targeted by the participant, which may imply a question type different from the one chosen by the participants. another problem occurs for questions that have semantic issues. for such questions, guessing the fragment in the input paragraph that triggered them can be challenging for the human judges. having the fragment already available as indicated by participants would allow the judges to better evaluate the question along all criteria. the semantic correctness was judged using the following scores: 1 – semantically correct and idiomatic/natural, 2 – semantically correct and close to the text or other questions, 3 – some semantic issues, 4 – semantically unacceptable. table 3 shows examples of real questions submitted by the university of pennsylvania team and the corresponding semantic correctness scores. scores of 3 and 4 were relatively easy to assign. sometimes, it was harder to differentiate between scores of 1 and 2 because while the question seems natural and almost impossible to formulate in a more natural way, it was too close or identical to the text. table 3. examples of questions corresponding to different levels of semantic correctness. questions semantic score what is the porosity of an aquifer? 1 how are diamonds brought close to the earth surface by a magma, which cools into igneous rocks known as kimberlites and lamproites? 2 who is ibn sahl credited with? 3 what might she to spend like his other childless concubines as? 4 the syntactic correctness was judged using the following scores: 1 – grammatically correct and idiomatic/natural, 2 – grammatically correct, 3 – some grammar problems, 4 – grammatically unacceptable. table 4 shows examples of real questions submitted by participants and the corresponding syntactic correctness scores. scores of 3 and 4 were easy to assign. differentiating between scores of 1 and 2 was at times difficult because the naturalness of the question is more difficult to judge when there is no obvious more natural way to ask a question given a particular target fragment in the input paragraph. correctness of question type means the specified type by a participant is agreed upon by the human judge. this is a binary dimension: 1 – means the judge agrees with the specified answer type, 0 – means the judge disagrees. an example of a wrong question type is provided in the following question, who do we see in the next scene?, whose correct type was what. question type correctness depends on the target content that triggered the question, as well as the body of the question. again, we assumed the target content submitted by the participants as being fixed/correct and judged the question type correctness with this assumption in mind. there are several cases that had to be considered. first, the question type was deemed incorrect when it did not match the target content and the question body. second, when the question body is semantically unacceptable, the question type can either be considered incorrect by default or can be judged correct if the target content does imply the selected question type. we chose to judge the correctness of question type with respect to the target content, ignoring the question body the first question generation shared task evaluation challenge 185 when semantically unacceptable. the reason for this choice was to avoid penalizing participants whose question body construction module was less developed. as a final remark on body when table 4. examples of questions corresponding to different levels of syntactic correctness. questions syntactic score what was the immediate cause of the first crusade? 1 how are diamonds brought close to the earth surface by a magma, which cools into igneous rocks known as kimberlites and lamproites ? 2 what do different materials have a different albedo so reflect a different amount of solar energy? 3 what does seek we to apply to life the understanding that separate parts of the ecosystem function as a whole? 4 semantically unacceptable. the reason for this choice was to avoid penalizing participants whose question body construction module was less developed. as a final remark on question type correctness, we would like to note that question type is a binary judgment (0-1) as opposed to the other dimensions where we used finer grain ratings (1-4), allowing for some gray ratings instead of just crisp distinctions. this difference in granularity of ratings may explain the differences in score values compared to the scores on the other dimensions. diversity of question types was also evaluated. at each scope level, ideally, each question had a different question type. a question type is loosely defined as being formed by the question word (e.g., wh-word or auxiliary) and by the head of the immediately following phrase. for instance, in what u.s. researcher …? the head of the phrase u.s. researcher that follows the question word what indicates a person, which means the question is actually a who question and not a what question. however, preference was given to diversity of question words. full question types, i.e. including the head of the phrase following the question word, were considered in special cases when the use of diverse question words was constrained by the input paragraph; that is, when different question words are hard to employ in order to generate different question types. for instance, some paragraphs may facilitate the generation of true what questions, i.e. what question types, but not when questions. for diversity ratings, we will use the following rubric: 1 – diverse in terms of question type and main body, 2 – diverse in terms of main body, 3 – paraphrase of a previous question, and 4 – similar-to-identical to a previous question. while full diversity would be ideal, it can be quite challenging for some input paragraphs. for pragmatic reasons, we relaxed the diversity criteria. we assigned the highest score of diversity if at least 50% of the question types in the whole 6-question set were different and the distribution of types was balanced, e.g. 2-who, 2-what, and 2-where would be scored higher than 4-who, 1what, and 1-where. we also defined overall average scores for each criterion. the overall average scores were defined as the average of individual scores. an individual score summarizes the scores along a dimension, e.g. syntactic correctness, and is computed by taking the average of individual scores shown by the formula below where q is the number of questions. || _ || q scoreindividual scoreoverallsyntactic q ∑ =−− the overall average score for specificity is more challenging to define. the goal would be to have a summative score with values from 1 to 4. 1 should be assigned to perfect system that generates 6 questions for each paragraph with the required distribution of specificity levels: 1 general, 2 medium, and 3 specific. for instance, if there are 4 specific questions in a set of 6 questions, then when judging the fourth specific question it will be penalized because the number rus, wyse, et al. 186 of expected specific questions (3) have been exhausted and another question at general or medium scope has not been generated. as of this writing, we are refining our overall average score for specificity. 2.3 data sources and annotation the primary source of input paragraphs used for the first qg-stec were: wikipedia, openlearn,8 and yahoo! answers. the three sources have their own peculiarities. wikipedia and openlearn are collections of well-written texts with wikipedia being developed in a more ad-hoc manner by volunteer contributors while openlearn, a repository of learning materials from open university courses that has been released for free access by the general public, has been created by professionals, i.e. academics and editors at open university. yahoo! answers is a communitybased question answering online service in which web surfers ask questions which are then answered by other surfers. yahoo! answers contains texts that are less edited resulting in texts that quite often contain ungrammatical sentences. the difference in quality between wikipedia and yahoo! answers texts can be explained by the differences in the way users contribute text and this text is subsequently dealt with. in wikipedia, once a text is drafted it can be polished by others over iterations until a stable version is reached. in yahoo! answers, a first contribution, i.e. answer to a question, is almost never re-edited by another contributor. we chose to use yahoo! answers as a source due to the large pool of question-answer pairs available for almost any type of questions (rus et al., 2009). eventually, the question-answer pairs can be used to train a system to generate the question given the answer. we collected (almost) paragraphs from each of these three sources. the paragraphs cover randomly selected topics of general interest. half of the paragraphs from each source were allocated to a development data set (65 paragraphs) and a test data set (60 paragraphs), respectively. for the development data set, we manually generated and scored 6 questions per paragraph for a total of 6 x 65 = 390 questions. paragraphs were selected such that they were self-contained (no need for previous context to be interpreted, e.g. will have no unresolved pronouns) and have around 5-7 sentences for a total of 100-200 tokens (excluding punctuation). in addition, we aimed for paragraphs that can facilitate the generation of 1 broad question, 2 medium questions, and 3 specific questions. examples of paragraphs from each source are given in table 5 below. we decided to provide minimal annotation for input in order to allow individual participants to choose their own preprocessing tools. we did not offer annotations for lemmas, pos tags, syntactic information, or propbank-style predicate-argument structures. this linguistic information can be obtained with acceptable levels of accuracy from open source tools. this approach favors comparison of full systems in a black-box manner as opposed to more specific components. we did provide discourse relations based on hilda, a freely available automatic discourse parser (duverle & prendinger, 2009). 2.4 submission format for each input paragraph, participants were asked to submit six lines of output. each line had to contain four items that are tab separated as shown below: rank index-set question-type question where rank is the rank or identifier of the question starting with 1, index-set is a set of indices in the input paragraph that indicate the target content in terms of span of tokens (words 8 openlearn gives free access to learning materials from the open university (http://open.ac.uk/openlearn) the first question generation shared task evaluation challenge 187 and punctuation) that triggered the question (or can form the answer to the question), question-type is the type of question, and question is the generated question. the index-set should contain pairs of start-end token indices delimited by commas: <0-10, 2035>. the index-set itself is delimited by ‘<’ and ‘>’. question-type can be one of the following (who, where, when, which, how, how many/long, generic, yes/no; see the questionfrom-sentences task description for details regarding these question types) or a new one submitted by participants. examples of lines in the above format are shown below. 1 <0-145> who who is abraham lincoln? 2 <98-126> what what major measures did president lincoln introduce? 3 <66-66> when when was abraham lincoln elected president? table 5. examples of input paragraphs from the three sources: wikipedia, openlearn, and yahoo! answers. source paragraph wikipedia enzymes are mainly proteins, that catalyze (i.e., increase the rates of) chemical reactions. an important function of enzymes is in the digestive systems of animals. enzymes such as amylases and proteases break down large molecules (starch or proteins, respectively) into smaller ones, so they can be absorbed by the intestines. starch molecules, for example, are too large to be absorbed from the intestine, but enzymes hydrolyse the starch chains into smaller molecules such as maltose and eventually glucose, which can then be absorbed. different enzymes digest different food substances. in ruminants which have herbivorous diets, microorganisms in the gut produce another enzyme, cellulase to break down the cellulose cell walls of plant fiber. openlearn there are two distinct zones containing water beneath the ground surface. the unsaturated zone has mainly air-filled pores, with water held by surface tension in a film around the soil or rock particles. water moves downwards by gravity through this zone, into the saturated zone beneath, in which all the pores are filled with water. the boundary surface between the unsaturated zone and the saturated zone is the water table, which is the level of water in a well (strictly, in a well that just penetrates to the water table). water below the water table, in the saturated zone, is groundwater. just above the water table is a zone called the capillary fringe, in which water has not yet reached the water table, because it has been held up by capillary retention. in this process water tends to cling to the walls of narrow openings. the width of the capillary fringe depends on the size of the pore spaces and the number of interconnected pores. it is generally greater for small pore spaces than for larger ones. yahoo! answers hi. depending on the type of duck, some can be trained. keep them in a dog pen if they are indoors, but with lots of "outside" time. keeping them out of food and water means getting special dishes just for this. any good farm supply store will have special dishes which have holes for the beak, and a sloped top so that ducks and poultry can't get into the dish. i'm including a link for the actual "care" of the duckling. ducks get big and male ducks, at maturity, can get a little bit mean. if you want a pet duck try to get a female. make sure to feed appropriate food. most ducks eat poultry food. muscovites eat game bird food. babies need crumbles, and adults get pellets. you can supplement food with scratch grain (but not too much, and not until they are a couple weeks old). bread should be kept as a once in a while thing. although, you may find that young ducks have a love of "milk sop". milk sop is when you soak stale bread in milk so that it's mushy. it's got calcium and protein and while it shouldn't be fed regularly it makes a great treat. 2.5 results and discussion for task a, there was one submission out of five registered participants. the participating team was from university of pennsylvania (mannem, prasad, & joshi, 2010). this section presents their approach and results. the approach proposed by the university of pennsylvania team (mannem, prasad, and joshi, 2010) for the task of generating questions from paragraphs (task a) is an over-generation approach in which many questions are first generated from a myriad of potential content items in the input paragraph followed by a ranking phase. the importance of a question is determined in rus, wyse, et al. 188 two separate steps. first, they use predicate argument structures along with semantic roles to identify important aspects of paragraphs. required and optional arguments of predicates are considered as good content candidates for asking questions about. predicates with less than two arguments are excluded and copular verbs are treated differently because of the limitations of the semantic role labeler. it should be noted that the type of arguments is also used for question type selection. this step has limitations when it comes to generating questions that should rely on cross-sentence information, e.g. a cause-effect discourse relation between two sentences. such cross-sentence information is important for medium and general questions in task a, as mentioned earlier. second, in the ranking phase mannem, prasad, and joshi prefer questions generated from main clauses and questions that do not include pronouns, because their approach did not include any coreference resolution. the question formulation step uses the selected content and corresponding verb complex in a sentence together with a set of reformulation rules to generate the actual questions. the initial plan in terms of evaluating the submissions for task a was to do peer-reviewing. however, because we only received one submission peer-reviewing was not possible. instead, we adopted an independent-judges approach in which two independents human raters judged the submitted questions using the interface depicted in figure 3. table 6. summary of results for university of pennsylvania. score results/inter-rater agreement specificity general= 90%; medium=121%; specific=80%; other = 1.39%/68.76% syntactic correctness 1.82/87.64% semantic correctness 1.97/78.73% question diversity 1.85/100% question type correctness 83.62%/78.22% table 7. summary of results by source of input paragraphs for university of pennsylvania. score wikipedia openlearn yahoo! answers specificity 19/46/50/2 19/45/47/1 16/54/47/2 syntactic correctness 1.63 1.86 1.98 semantic correctness 1.90 1.96 2.08 question diversity 1.55 2 1.95 question type correctness 91.73% 87.6% 71.53% for the 60 input paragraphs used for testing, participants were supposed to submit 60 general questions, 120 medium questions, and 180 specific questions, for a total of 360 questions. the university of pennsylvania team submitted 349 questions of which 54 were rated general based on the span of the input paragraph fragment that triggered the question (index-set field in the submission format), 145 as medium, 144 as specific, and 5 as other. the table indicates the percentages of each type of question that were submitted out of what they were supposed to submit. in the other category, we report the percentage of questions that could not be classified as either general, medium, or specific, out of the total number of questions that were supposed to be submitted, i.e. 360. examples of questions that were classified in the other category are those in which participants submitted an empty span for the input fragment based on which the question was generated, e.g. the index-set field has a value of 285-285, or the question was so hard to understand that it was impossible to evaluate its scope. the inter-rater agreement, i.e. the percentage of annotated items on which judges agree, on the specificity criterion was 68.76%. the results indicate that participants tended to overgenerate medium questions (they submitted the first question generation shared task evaluation challenge 189 more questions than asked for) and undergenerate general and specific questions (participants submitted less questions than required). in terms of syntax, the majority of the submitted questions were deemed correct (score of 1.97 ~ a rate of 2 means grammatically correct) but not necessarily idiomatic or natural. a good majority of the questions were generated by re-using entire chunks of text from the input paragraph, i.e. a selection-based generation approach for the question construction phase. semantics-wise, the submitted questions were deemed semantically correct and close to the text and/or other questions (average score of 1.97 ~ 2) but not idiomatic or natural. the diversity of the submitted questions was rated at 1.85 level, which is close to a score of 2. a score of 2 indicates diversity in terms of the main body of the questions. the types of question and their distribution is shown in the table 8 and charts in figures 8 and 9 below. it should be noted that the types were obtained by only considering the question word, e.g. wh-word, and without any finer-grain analysis, e.g. distinguishing definition questions such as what is …? from other what questions. we would like to point out that all what questions that were submitted are true what questions, as opposed to disguised other types as in what researchers discovered dna? which is semantically a who question (asking for a person rather than a thing). that is, all the submitted what questions followed the pattern what auxiliary-verb …? the question type correctness was 83.62% overall with inter-rater agreement of 78.22%. table 8. distribution of question types (overall and by input paragraph source). question type overall wikipedia openlearn yahoo! answers how 40 16 12 12 what 258 80 87 91 when 22 5 5 12 where 12 7 4 1 who 9 5 1 3 why 8 4 3 1 figure 1. distribution of question types. rus, wyse, et al. 190 figure 2. distribution of question types by input paragraph source. figure 3. a screenshot of the question generation from paragraphs rating tool. 3 task b: question generation from sentences 3.1 task definition this task had four participants: jadavpur university, university of lethbridge, saarland university, and university of wolverhampton. participants were given a set of inputs, with each input consisting of: § a single sentence and § a specific target question type (e.g., who?, why?, how?, when?; see below for the complete list of types used in the challenge). the first question generation shared task evaluation challenge 191 for each input, the task was to generate 2 questions of the specified target question type. input sentences, 90 in total, were selected from openlearn, wikipedia and yahoo! answers (30 inputs from each source). extremely short (<5 words) or long sentences (>35 words) were not included. prior to receiving the actual test data, participants were provided with a development/example data set consisting of sentences from the aforementioned sources and, for one or more target question types, examples of questions. these questions were manually authored and crosschecked by the team organizing task b. the three examples in table 9 are taken from the development data set, one each from openlearn, wikipedia and yahoo! answers. note that input sentences were provided as raw text. annotations were not provided. there are a variety of nlp open source tools available to potential participants and the choice of tools and how these tools are used was considered a fundamental part of the challenge. table 9. examples of sentences and corresponding questions from three sources. source openlearn sentence the poet rudyard kipling lost his only son in the trenches in 1915. target question type who who lost his only son in the trenches in 1915? when when did rudyard kipling lose his son? how many how many sons did rudyard kipling have? source wikipedia sentence two important variables used for the classification of igneous rocks are particle size, which largely depends upon the cooling history, and the mineral composition of the rock. target question type which which two important variables are used for the classification of igneous rocks? source yahoo!answers sentence in australia you no longer can buy the ordinary incandescent globes, as you probably already know. target question type where where can you no longer buy the ordinary incandescent globes? yes/no can you buy the ordinary incandescent globes in australia? what what can you no longer buy in australia? participants were also provided with the following list specifying the target question types: § who?: the answer to the generated question is a person (e.g. abraham lincoln) or group of people (e.g. the american people) named in the input sentence. § where?: the answer to the generated question is a placename (e.g. dublin, mars) or location (north-west, to the left of) which is contained in or can be derived from the input sentence. § when?: the answer is a specific date (e.g. 3rd july 1973, 4th july), time (e.g. 2:35, 10 seconds ago), era or other representation of time. § which?: the answer will be a member of a category (e.g. invertebrate or vertebrate) or group (e.g. colours, race) or a choice of entities (e.g. union or confederacy) given in the input sentence. rus, wyse, et al. 192 § what?: the question might describe a specific entity mentioned in the input sentence and ask what it is. the question may also ask the purpose, attributes or relations of an entity as described in the input sentence. § why?: the question asks the reasoning behind some statement made in the input sentence § how many/long?: the answer will be a duration of time or range of values (e.g. 2 days) or a specific count of entities (e.g. 32 counties) within the input sentence. § yes/no: the generated question should ask whether a fact contained in the input sentence is either true or false (e.g. are mathematical co-ordinate grids used in graphs?). 3.2 guidelines for human judges the evaluation criteria fulfilled two roles. firstly, they were provided to the participants as a specification of the kind of questions that their systems should aim to generate. secondly, they also played the role of guidelines for the judges of system outputs in the evaluation exercise. for this task, five criteria were identified: § relevance § question type § syntactic correctness and fluency § ambiguity § variety all criteria are associated with a scale from 1 to n (where n is 2, 3 or 4), with 1 being the best score and n the worst score. the criteria are defined as follows. 3.2.1 relevance questions should be relevant to the input sentence. this criterion measures how well the question can be answered based on what the input sentence says. table 10. scoring rubric for relevance. rank description 1 the question is completely relevant to the input sentence. 2 the question relates mostly to the input sentence. 3 the question is only slightly related to the input sentence. 4 the question is totally unrelated to the input sentence. 3.2.2 question type questions should be of the specified target question type. table 11. scoring rubric for question type. rank description 1 the question is of the target question type. 2 the type of the generated question and the target question type are different. 3.2.3 syntactic correctness and fluency the syntactic correctness is rated to ensure systems can generate grammatical output. in addition, those questions which read fluently are ranked higher. the first question generation shared task evaluation challenge 193 table 12. scoring rubric for syntactic correctness and fluency. rank description example 1 the question is grammatically correct and idiomatic/natural. in which type of animals are phagocytes highly developed? 2 the question is grammatically correct but does not read as fluently as we would like. in which type of animals are phagocytes, which are important throughout the animal kingdom, highly developed? 3 there are some grammatical errors in the question. in which type of animals is phagocytes, which are important throughout the animal kingdom, highly developed? 4 the question is grammatically unacceptable. on which type of animals is phagocytes, which are important throughout the animal kingdom, developed? 3.2.4 ambiguity the question should make sense when asked more or less out of the blue. typically, an unambiguous question will have one very clear answer. table 13. scoring rubric for ambiguity. rank description example 1 the question is unambiguous. who was nominated in 1997 to the u.s. court of appeals for the second circuit? 2 the question could provide more information. who was nominated in 1997? 3 the question is clearly ambiguous when asked out of the blue. who was nominated? 3.2.5 variety pairs of questions in answer to a single input (i.e., with the same target question type) are evaluated on how different they are from each other. this rewards those systems which are capable of generating a range of different questions for the same input. table 14. scoring rubric for variety. rank description example 1 the two questions are different in content. where was x born?, where did x work? 2 both ask the same question, but there are grammatical and/or lexical differences. what is x for?, what purpose does x serve? 3 the two questions are identical. each of the criteria is applied independently of the other criteria to each of the generated questions. scoring on the criteria was done by two judges for each generated question (with the judges not knowing which system generated the question). for each question, one of the judges was a member of the qg-stec team and the other a member of one of the participating teams rus, wyse, et al. 194 (though they never got to rate questions from their own system). in short, we had a mix of independent judges and peer-review. the scores of the two judges for each item were averaged, with the average score being used in the further calculations reported below. 3.3 results & discussion this section contains a presentation and analysis of the results of the qg from sentences task of qg-stec 2010.9 the results concern a variety of criteria in order to determine possible improvements which might be made to both the qg from sentences task and to future qg systems. the first of these criteria is the total number of questions generated for a data instance by all qg systems. the results for the rating criteria used by the human evaluators were also analysed both in relation to the different sources and systems. the marking system employed used lower scores for better questions. the best score for all criteria was 1. thus, in the tables and charts used in this report the lower values indicate a better score. 3.3.1 number of questions generated by resource the qg from sentences task used 30 sentences from each of three online resources, openlearn, wikipedia and yahoo! answers, giving a total of 90 sentences. for each sentence one or more target question types was specified resulting in 180 possible generated questions. in order to measure the ability of a qg system to generate a variety of questions, participants were asked to generate two questions for each target sentence and type. the maximum number of submitted questions possible was then 360 times the number of participants (4) resulting in a total of 1440. no. of sentences average no. of question types two questions per question type possible # of questions per participant 90 x 2 x 2 = 360 analysing the actual number of questions generated with regard to the online resource can provide some indication as to which resources are more difficult to work with from a qg perspective. table 15 below shows the actual questions generated by participants for input sentences generated from each of the online resources. table 15. percentage of questions generated per data resource. resource max achievable questions per resource actual generated questions per resource percentage generated openlearn 480 312 65% wikipedia 480 309 64% yahoo!answers 480 275 57% all systems followed the guidelines returning up to two questions for each question type. the number of generated questions varied across the systems resulting in 62% of all requested questions. for sentences extracted from either the openlearn or wikipedia resource it was found that systems generated approximately 65% of the maximum number of questions possible. the yahoo! answers resource appears to have been more challenging. for sentences extracted from this resource the systems only generated 57% of the maximum possible. the resource does contain many spelling errors and other language imperfections which make it difficult for nlp and this was also reflected in the stec results. 9 the raw data with the results can be found at http://computing.open.ac.uk/coda/data.html the first question generation shared task evaluation challenge 195 3.3.2 number of questions generated by question type the target question types were specified for each sentence with the intention of discovering those types which provided more of a challenge for current qg systems. analysing the results for actual questions generated per question type provides an indication of this. table 16 below shows the percentage of actual generated questions per question type. table 16. percentage of questions generated per question type. max achievable questions per type actual generated questions per type percentage generated when 144 107 74 % who 120 84 70 % where 120 74 62 % what 464 285 61 % which 176 108 61 % why 120 72 60 % how many 184 106 58 % yes/no 112 60 54 % assuming that the number of actually generated questions provides some indication of the ability to generate questions for a specific type, the data suggest that questions of some types are easier to generate than others. participating systems generated 70% and above of the maximum possible sentences for ‘when’ and ‘who’ question types. for ‘yes/no’ question types it was found that only 54% of the maximum possible sentences were generated. this could indicate that ‘yes/no’ questions are more difficult to generate. this particular figure is, however, skewed because one of the participating systems had no rules at all for generating yes-no questions and therefore never generated this type of question. 3.3.3 average scores by resource questions were rated by our judges using five criteria. these criteria were then used to measure the performance of systems with regard to the criteria and to determine areas where qg systems might be improved. we first present and discuss the scores for the questions organised by resource. analysing the scores on the criteria with regard to the resource should again indicate any difficulty a particular resource might present, in this case according to a specific criterion. we also again analyse the results according to the target question type to discover any relation between target question types and the evaluation criteria. table 17 and figure 4 below show the average scores for each criterion for each resource used in the stec. lower scores are better. table 17. percentage of questions generated per question type. relevance question type correct ness ambiguity variety openlearn 1.54 1.07 2.14 1.54 1.65 wikipedia 1.61 1.12 2.24 1.61 1.89 yahoo! answers 1.57 1.12 2.23 1.72 2.04 rus, wyse, et al. 196 average relevance scores per resource 1.5 1.52 1.54 1.56 1.58 1.6 1.62 1.64 openlearn yahoo! answers wikipedia average question type scores per resource 1 1.02 1.04 1.06 1.08 1.1 1.12 1.14 1.16 1.18 1.2 openlearn wikipedia yahoo! answers average correctness scores per resource 2.1 2.12 2.14 2.16 2.18 2.2 2.22 2.24 2.26 2.28 2.3 openlearn yahoo! answers wikipedia average ambiguity scores per resource 1.45 1.5 1.55 1.6 1.65 1.7 1.75 1.8 openlearn wikipedia yahoo! answers average variety scores per resource 1.55 1.65 1.75 1.85 1.95 2.05 openlearn wikipedia yahoo! answers figure 4. scores for each question data source. the first question generation shared task evaluation challenge 197 the data show that the sentences from the openlearn resource achieved as good as or better generation results than the other two resources on all criteria. one possible reason for this is that openlearn study units are rigorously checked before being published. question type, which ensures that generated questions are of the type requested by the test instance, is close to 1 (the optimal value) for all resources. in general, most generated questions were of the type specified. input sentences derived from yahoo! answers reflect the nature of that resource. text from yahoo! answers contains a lot of vague, badly spelled, ambiguous and generally imperfect language. 3.3.4 average scores by question type the evaluation criteria can also be used together with the various target question types in order to determine whether particular question types present difficulties for generation systems. table 18 and figure 5 below show the average scores for the five criteria according to the target question type for the data instance. table 18. scores for each question type. relevance question type correctness ambiguity variety how many 1.41 1.1 2.44 1.51 1.98 who 1.45 1.06 1.88 1.48 1.57 when 1.54 1.08 2.37 1.72 1.87 yes/no 1.57 1.05 2.35 1.63 2.02 what 1.58 1.08 1.98 1.63 1.68 which 1.64 1.22 2.36 1.51 1.99 where 1.66 1.05 2.36 1.69 1.99 why 1.74 1.15 2.31 1.83 2.16 average relevance scores per question type 1.2 1.3 1.4 1.5 1.6 1.7 1.8 how m any who when yes/n o what whic h where why average question type scores per question type 1 1.05 1.1 1.15 1.2 1.25 1.3 yes/n o where who when what how m any why whic h rus, wyse, et al. 198 average correctness scores per question type 1.8 1.9 2 2.1 2.2 2.3 2.4 2.5 who what why yes/n o where whic h when how m any average ambiguity scores per question type 1.4 1.45 1.5 1.55 1.6 1.65 1.7 1.75 1.8 1.85 1.9 who whic h how m any yes/n o what where when why average variety scores per question type 1.5 1.6 1.7 1.8 1.9 2 2.1 2.2 who what when how m any where whic h yes/n o why figure 5. scores for each question type. there is variation in the results for the different criteria when grouped by target question type. the type ‘how many’ achieved the best results for matching question type and also the worst result for ‘correctness’. ‘why’ question types were amongst the poorest rated generated questions for all criteria. ‘who’ question types achieved the best results in most criteria. 3.3.5 average scores by system five systems entered this task: mrsqg saarland (yao & zhang, 2010), wlv wolverhampton (varga & ha, 2010), jugg jadavpur (pal et al., 2010) and lethbridge (ali et al., 2010). the averaged results for the systems on each of the evaluation criteria are depicted in figure 6, with lower values indicating better scores. wlv scores best on all criteria except for “variety”. the picture changes when systems are penalized for missing questions (figure 7), i.e., when a missing question is given the lowest possible score on each of the criteria. now, mrsqg outperforms the other systems on all criteria. the first question generation shared task evaluation challenge 199 0.00 0.50 1.00 1.50 2.00 2.50 3.00 mrsqg saarland wlv wolverhampton jugg jadavpur lethbridge system a ve ra ge s co re relevance question type correctness ambiguity variety figure 6. results for qg from sentences (without penalty for missing questions) results with penalty missing questions 0.00 1.00 2.00 3.00 4.00 mrsqg saarland wlv wolverhampton jugg jadavpur lethbridge systems a ve ra ge s co re s relevance question type correctness ambiguity variety figure 7. results for qg from sentence (with penalty for missing questions) the lethbridge system uses syntactic parsing, part of speech (pos) tagging and a named entity analyzer to decorate the input sentence with annotations. based on the syntactic parse, the sentence is simplified, dividing it into short sentences. these are then matched with rules. a rule consists essentially of a template that is matched against the input and corresponding question patterns. if the template is successfully matched with the structure of an input sentence, the result is one or more instantiated question patterns, i.e, questions. a similar approach is followed by juqgg, though it uses a slightly different set of tools for decorating the input text with structure (e.g., it includes also semantic role labelling). there is also no sentence simplification step. there is, however, as with the lethbridge system, a stage where the processed input is matched with templates and then mapped to question patterns an approach also followed previously by, for example, wyse & piwek (2009). the wlv system also uses this approach. it is, however, set apart from the two other systems by not having any rules for yes-no questions and including a checking stage where tests are performed on the input (does it contain negation, complex structures, etc.) which determine whether a question of sufficient quality can be generated. finally, mrsqg forms a class of its own. it uses natural language understanding technology to produce a language-independent deep semantic representation of the input sentence. the representation of the declarative content of the input sentence is transformed, using rules, to a representation of the content of a question. it then takes this question representation and uses an existing natural language generation system to construct the surface realization of the question. rus, wyse, et al. 200 when we penalize systems for missing questions, the semantics-based deep approach used by mrsqg, on the qg-stec data, outperforms the rule-based shallow approach of the other three systems. interestingly, if we compute the scores using only the values for the actually generated questions, the wlv system comes out best. here it seems that the wlv system strategy of checking the input prior to generation (and not generating an output if the input doesn’t satisfy certain criteria) pays off. 4 lessons learned the first qg-stec was definitely a success by many measures including number of participants, results, and resources created, given the lack of funded projects on the topic. as already mentioned, there was a big discrepancy between number of researchers who expressed their interest to participate and the number of researchers who actually participated. in the absence of research funding targeting qg research efforts, researchers could not afford spending substantial effort on developing competitive algorithms to participate in qg-stecs. funding of systematic research on qg is desperately needed to advance the field and keep the momentum going. after reflecting on the process and outcomes, we have identified several aspects that have helped making the task a success and some that could be improved in the next round of shared evaluations of tasks a and b or that could be used in planning other qg tasks. one key aspect of process to put a qg-stec together was the involvement of the qg community. the community was asked to submit candidate qg tasks for a first qg-stec. then, the community at large was polled regarding which tasks to choose for the first qg-stec. we believed that more than two tasks would have diluted the success of a first qg-stec. by polling the community, we indirectly assured maximum participation. one downside of the community polling process was the introduction of some delays regarding the official announcement of the qg tasks to be offered in the first qg-stec. we recommend the next qg-stec should be announced at least one year in advance of the expected publication of the results. this will give enough lead time to advertise the event and also give participants sufficient time to get ready for the challenge. in terms of planning, we recommend that the test period be scheduled at least 4 months before the publication of the results such that a peer-evaluation process can be effectively implemented. one other positive aspect of the first qg-stec was the post-challenge group discussions at the 3rd workshop on question generation where the results of the stec were made publicly available. the participants had a chance to provide suggestions for improving the next round of the stec. for instance, the participants recommended that for task a, question generation from paragraphs, we eliminate yahoo! answers as a source of texts due to language challenges posed by unedited yahoo! answers paragraphs. indeed, we compared means of the five quality criteria with source of input paragraphs as the factor (n=117 for wikipedia; n=112 for openlearn; n=120 for yahoo! answers). a significant difference was found for the syntactic complexity criterion (p=0.015<0.05). a post-hoc tukey test revealed significant differences between wikipedia and yahoo! answers groups (p=0.013<0.05). as replacement for yahoo! answers, there was a suggestion to consider texts from domains that are of high-interest, for example, in biomedical domains. this may attract research groups working on biomedical texts, which could give a boost to the qg-stec and qg research efforts in general. another suggestion was to not only ask participants to generate questions and indicate the fragment in the input paragraph that triggered the question, but also to ask participants to generate the answers themselves. this latter suggestion needs more thought before it is implemented, as it seems to add another level of complexity to the qg-stec. one last general suggestion is to consider replacing rating scales, which we used in evaluating the submissions of the first qg-stec, with preference judgments as the former seem to pose some challenges. they may be unintuitive for raters and the inter-rater agreement tends to the first question generation shared task evaluation challenge 201 be low when using rating scales (belz & kow, 2010a). a rating scale makes raters judge questions in absolute terms. in contrast, preference judgments would allow them to judge which of two items they prefer for a particular evaluation criterion, e.g. syntactic correctness. if preference judgments are implemented in a future qg-stec, raters will be shown several questions at a time, submitted by different participants, and asked to rank them based on each of the criteria. belz and kow (2010b) showed that using preference judgments for language generation tasks leads to better inter-annotator agreement and also explains a larger proportion of variation accounted for by system differences. we also have some specific recommendations for running the next round tasks of a and b. a spin-off of task a could be a task focusing on content selection as this step seems to be the more challenging for this task as it involves discourse level processing. a user model can be added to such a task in order to make it more interesting. furthermore, a task on content selection can target a specific application such that there is a clear set of goals that could be used to quantify the importance of various content items. other spin-off tasks can focus on question type selection and question construction as well. for task a, there was on approach proposed in which many questions are first generated from a myriad of potential content items in the input paragraph followed by a ranking phase. the importance of a question was determined in two separate steps. first, predicate argument structures along with semantic roles are used to identify important aspects of paragraphs. this step has limitations when it comes to generating questions that should rely on cross-sentence information, e.g. a cause-effect discourse relation between two sentences. second, the approach preferred questions generated from main clauses and questions that do not include pronouns because no coreference resolution was used. the question formulation step used the selected content and corresponding verb complex in a sentence together with a set of reformulation rules to generate the actual questions. they avoided using discourse parsers in their approach due to the modest current performance of these parsers. we will have to provide discourse annotations for the input paragraphs in a future offering of task a, either by manually improving the output of available discourse parsers, such as hilda, or by manually annotating the input paragraphs. while this may seem inconsistent with our principle of using raw input for the task, we feel this minimal annotation of input for discourse relations is necessary given the young stage of development of discourse parsers relative to the maturity of other standard language processing steps such as part of speech tagging or syntactic parsing. for task b, there were essentially two approaches: a ‘shallow’ and a ‘deep’ one. whereas the shallow approach relied on a combination of syntactic and semantic annotations of the input to guide rules for mapping declarative text to questions, the deep approach relied on full computation of a language-independent semantic representation of the input which is mapped to a semantic representation of the output, which is then mapped in turn to a question using language generation technology. the deep approach (as used by the mrsqg system) performed best, provided that missing questions were penalized (i.e., questions that the system was asked to generate but did not generate). one of the ‘shallow’ systems, wlv, did however come out best when missing questions were not taken into account. this particular system distinguished itself from the other systems by including a checking stage, where, based on an analysis of the input, the system decided whether to generate an output. we should also add that the wlv system could have performed even better if it had included rules for yes-no questions (these were completely missing from this system). there are several other ways to improve future runs of the two tasks offered in the first qgstec. a suggestion is to increase the number of input data items to at least 60 items per sources such that statistical significance tests could be performed to compare results across the various sources of input text. another suggestion is to add a specificity field to the submission requirements. the specificity field will require participants to indicate what the scope is for each of the questions rus, wyse, et al. 202 (general, medium, or specific). in the first qg-stec, the scope was somehow inferred from the answer location range by the rater. we also need to develop a smart scheme to combine the individual scores into an overall score for tasks a and b. we believe a second run of the tasks may give us enough evidence to develop an appropriate overall score by weighting the various criteria according to some principles. 5 conclusions the first qg-stec was a success in terms of participants and resources created (detailed results are not discussed due to space constraints, but see the links to the raw results data in the sections above for task a and b). we now have several datasets of text-question pairs together with human judgments using various criteria. also, we developed a software tool that can be used by humans to rate questions (see figure 1). participants have proposed and implemented several approaches that can be used as a starting point for future qg-stecs and sources of inspiration by newcomers. while these resources are valuable, there is still need for larger datasets, linguistic annotations for the input data, and more tools and improved evaluation methodologies. the short term goal should be scaling up the qg-stec for tasks a (question generation from paragraphs) and b (question generation from sentences) such that more data is available to enable more solid results and more participants. in the long term, the qg-stec should scale up in terms of tasks, e.g. more targeted tasks that focus on specific aspects of the question generation process such as question type selection. the survival and future successes of the qg-stec movement and of the qg research community at large is highly dependent on the availability of funding resources directed specifically at question generation research projects. acknowledgments we are grateful to a number of people who contributed to the success of the first shared task evaluation challenge on question generation: rodney nielsen, amanda stent, arthur graesser, jose otero, and james lester. also, we would like to thank the national science foundation who partially supported this work through grants ri-0836259 and ri-0938239 (awarded to vasile rus), the institute for education sciences who partially funded this work through grant r305a100875 (awarded to vasile rus), and the engineering and physical sciences research council who partially supported the effort on task b through grant ep/g020981/1 (awarded to paul piwek). the views expressed in this paper are solely the authors’. references ali, h, chali, y., hasan, s. (2010). automation of question generation from sentences, in: boyer & piwek (2010), pp. 58-67. beck, j.e., mostow, j., and bey, j. (2004). can automated questions scaffold children's reading comprehension? proceedings of the 7th international conference on intelligent tutoring systems, 478490. 2004. maceio, brazil. belz, a., kow, e., viethen, j. and gatt, a. (2008) the grec challenge 2008: overview and evaluation results. in proceedings of the 5th international natural language generation conference (inlg'08), pp. 183-191. belz, a. and kow, e. (2010a). the grec challenges 2010: overview and evaluation results. in proceedings of the 6th international natural language generation conference (inlg'10), pp. 219229. the first question generation shared task evaluation challenge 203 belz, a. and kow, e. (2010b). comparing rating scales and preference judgements in language evaluation. in proceedings of the 6th international natural language generation conference (inlg'10), pp. 7-15. boyer, k. and piwek, p. (2010) (eds). qg2010: the third workshop on question generation, carnegie mellon university, pittsburgh, pa, usa. dale, r. & m. white (2007) (eds.). position papers of the workshop on shared tasks and comparative evaluation in natural language generation. duverle, d. and prendinger, h. (2009). a novel discourse parser based on support vector machines. proc 47th annual meeting of the association for computational linguistics and the 4th int'l joint conf on natural language processing of the asian federation of natural language processing (aclijcnlp'09), singapore, aug 2009 (acl and afnlp), pp 665-673. edmonds, p. (2002). introduction to senseval. elra newsletter, october 2002. judita preiss and david yarowsky (2001). editors. the proceedings of senseval-2: second international workshop on evaluating word sense disambiguation systems. koller, a., striegnitz, k., gargett, a., byron, d., cassell, j., dale, r., moore, j., and oberlander, j. (2010). report on the second nlg challenge on generating instructions in virtual environments (give-2). in proceedings of the 6th international natural language generation conference (inlg), dublin, ireland. kunichika, h., katayama, t., hirashima, t., & takeuchi, a. (2001). automated question generation methods for intelligent english learning systems and its evaluation, proc. of icce01. lauer, t., peacock, e., & graesser, a. c. (1992) (eds.). questions and information systems. hillsdale, nj: erlbaum. lin, c.y. (2004). rouge: a package for automatic evaluation of summaries. in proceedings of the workshop on text summarization branches out, post-conference workshop of acl 2004, barcelona, spain. lin, c. and och, f. (2004). automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics, in proceedings of the 42nd annual meeting of the association of computational linguistics. mannem, p., prasad, r., and joshi, a. (2010). question generation from paragraphs at upenn: qgstec system description, proceedings of the third workshop on question generation (qg 2010), pittsburgh, pa, june 2010. mitkov, r., ha, l. a. and karamanis, n. (2006): a computer-aided environment for generating multiplechoice test items. natural language engineering 12(2). 177-194. cambridge university press. mostow, j. and chen, w. (2009). generating instruction automatically for the reading strategy of selfquestioning. proceedings of the 14th international conference on artificial intelligence in education (aied2009), brighton, uk, 465-472. mostow, j., beck, j., bey, j., cuneo, a., sison, j., tobin, b., and valeri, j. (2004). using automated questions to assess reading comprehension, vocabulary, and effects of tutorial interventions. technology, instruction, cognition and learning, 2004. 2: p. 97-134. pal, s., mondal, t., pakray, p., das, d., bandyopadhyay, s. (2010). qgstec system description: juqgg: a rule based approach. in: boyer & piwek (2010), pp.76-79. papineni, k., roukos, s., ward, t., and zhu, w. j. (2002). bleu: a method for automatic evaluation of machine translation, in acl-2002: 40th annual meeting of the association for computational linguistics pp. 311–318. piwek, p., h. hernault, h. prendinger, m. ishizuka (2007). t2d: generating dialogues between virtual agents automatically from text. in: intelligent virtual agents: proceedings of iva07, lnai 4722, september 17-19, 2007, paris, france, (springer-verlag, berlin heidelberg) pp.161-174. piwek, p. and s. stoyanchev (2011). data-oriented monologue-to-dialogue generation. proceedings of the 49th annual meeting of the association for computational linguistics:shortpapers, pages 242-247, portland, oregon, june 19-24, 2011. reiter, e. & dale, r. (1997). building applied natural-language generation systems. journal of naturallanguage engineering, 3:57-87. reiter, e. & dale, r. (2000). building applied natural-language generation systems, oxford university press, 2000. rus, v., cai, z., graesser, a.c. (2007a). experiments on generating questions about facts. alexander f. gelbukh (ed.): computational linguistics and intelligent text processing, 8th international conference, cicling 2007, mexico city, mexico, february 18-24, 2007 rus, wyse, et al. 204 rus, v., graesser, a.c., stent, a., walker, m., and white, m. (2007b). text-to-text generation, in shared tasks and comparative evaluation in natural language generation by robert dale and michael white, november, 2007, pages 33-46. rus, v. and graesser, a.c. (2009). workshop report: the question generation task and evaluation challenge, institute for intelligent systems, memphis, tn, isbn: 978-0-615-27428-7. varga, a. and ha, l.a. (2010). wlv: a question generation system for the qgstec 2010 task b. in: boyer & piwek (2010), pp. 80-83. voorhees, e. m. and tice, d.m. (2000). the trec-8 question answering track evaluation. in e.m. voorhees and d.k. harman, editors, proceedings of the eighth text retrieval conference (trec-8), pages 83-105, 2000. nist special publication 500-246. walker, m., rambow, o., & rogati, m. (2002). training a sentence planner for spoken dialogue using boosting, computer speech and language special issue on spoken language generation , july 2002. wolfe, j.h. (1976). "automatic question generation from text an aid to independent study." sigcue outlook 10(si): 104--112. wyse, b. and p. piwek (2009). generating questions from openlearn study units. in: v. rus and j. lester (eds.), proceedings of the 2nd workshop on question generation, aied 2009 workshop proceedings, pp. 66-73. yao, x and zhang, y. (2010). question generation with minimal recursion semantics. in: boyer & piwek (2010), pp.68-75. dialogue and discourse 3(2) (2012) 75–99 doi: 0.5087/dad.2012.204 question generation from concept maps andrew m. olney aolney@memphis.edu arthur c. graesser a-graesser@memphis.edu institute for intelligent systems university of memphis 365 innovation drive memphis, tn 38152 natalie k. person person@rhodes.edu department of psychology rhodes college 2000 n. parkway, memphis, tn 38112-1690 editors: paul piwek and kristy elizabeth boyer abstract in this paper we present a question generation approach suitable for tutorial dialogues. the approach is based on previous psychological theories that hypothesize questions are generated from a knowledge representation modeled as a concept map. our model semi-automatically extracts concept maps from a textbook and uses them to generate questions. the purpose of the study is to generate and evaluate pedagogically-appropriate questions at varying levels of specificity across one or more sentences. the evaluation metrics include scales from the question generation shared task and evaluation challenge and a new scale specific to the pedagogical nature of questions in tutoring 1. introduction a large body of research exists on question-related phenomena. much of this research derives from the tradition of question answering rather than question asking. whereas question answering has received extensive attention in computer science for several decades, with increasing interest over the last 10 years (voorhees and dang 2005, winograd 1972), question asking has received attention primarily in educational/psychological circles (beck et al. 1997, bransford et al. 1991, brown 1988, collins 1988, dillon 1988, edelson et al. 1999, palinscar and brown 1984, piaget 1952, pressley and forrest-pressley 1985, scardamalia and bereiter 1985, schank 1999, zimmerman 1989) until the recent surge of interest in computational accounts (rus and graesser 2009). the distinction between question answering and question asking in some ways parallels the distinction between natural language understanding and natural language generation. natural language understanding takes a piece of text and maps it to one of many possible representations, while natural language generation maps from many possible representations to one piece of text (dale et al. 1998). question answering, after all, requires some understanding of what the question is about in order for the correct answer to be recognized when found. likewise question asking seems logically tied to the idea of selection, since only one question can be asked at a given moment. though the above description appeals to representations, the use of representations in question answering has been largely avoided in state of the art systems (ferrucci et al. 2010, moldovan et al. 2007). given the success of representation-free approaches to question answering, a pertinent question is whether the same can be true for question generation (yao 2010). the simplest instantiation of representation-free approach would be to syntactically transform a sentence into a question using operations like wh-fronting and wh-inversion, e.g. “john will eat sushi” becomes “what will john c©2012 andrew olney, arthur graesser and natalie person submitted 3/11; accepted 1/12; published online 3/12 olney, graesser and person eat?” and indeed, this basic approach has been used extensively in computational models of question generation (ali et al. 2010, heilman and smith 2009, 2010, kalady et al. 2010, varga and ha 2010, wyse and piwek 2009). these approaches are knowledge-poor in that they do not take into account properties of words other than their syntactic function in a sentence. one significant problem with knowledge poor approaches is determining the question type. in the example above, it is necessary to determine if “sushi” is a person, place, or thing and use “who,” “where,” or “what,” respectively. so it appears that some degree of knowledge representation is required to generate questions with this higher degree of precision. the literature on question generation has approached this problem in primarily two different ways. the first is by using named entity recognizers to classify phrases into useful categories like place, person, organization, quantity, etc. (ali et al. 2010, kalady et al. 2010, mannem et al. 2010, varga and ha 2010, yao and zhang 2010). the second approach to determining question type is semantic role labeling (chen et al. 2009, mannem et al. 2010, pal et al. 2010, sag and flickinger 2008, yao and zhang 2010). semantic role labelers assign argument structures to parse trees, specifying the subject, patient, instrument, and other roles associated with predicates (palmer et al. 2005). role labels thus make it easier to deal with phenomena like passivization which invert the order of arguments. role labels also provide useful adjunct roles specifying causal, manner, temporal, locative, and other relationships (carreras and màrquez 2004). previous work on question generation has demonstrated that named entities and adjunct roles may both be mapped to question type categories in a straightforward way. for example, the sentence “charles darwin was impressed enough with earthworms that he devoted years-and an entire book-to their study.” may be transformed into the question “who was impressed enough with earthworms that he devoted years-and an entire bookto their study?” by mapping the person named entity “charles darwin” to the question type “who.” likewise the sentence “because fermentation does not require oxygen, fermentation is said to be anaerobic.” may be transformed into the question “why is fermentation said to be anaerobic?” by mapping the causal adjunct am-cau to the question type “why.” clearly semantic role labelers and named entity recognizers bring significant value to the question generation process by adding knowledge. one might speculate that if a little knowledge is good, perhaps a lot of knowledge is even better. but what sort of knowledge? recent work by chen and colleagues has explored the role that knowledge may play in generating questions (chen et al. 2009, chen 2009, mostow and chen 2009). rather than generating questions one sentence at a time, their system builds a situation model of the text and then generates questions from that model. in an early version of the system, which works exclusively with narrative text, chen et al. (2009) employ the clever trick of maintaining a situation model only of characters’ mental states. because mental states of characters are usually fundamental in explaining the evolution of a narrative, questions generated from the model appear to satisfy a nontrivial challenge for question generation: generating important questions that span more than one sentence (vanderwende 2008). although the early version of the system only generates simple yes/no questions, later versions can generate “why” and “how” questions for both narrative and informational text. an example of a mental state “why” question for narrative text is give in figure 1. 76 question generation from concept maps once upon a time a town mouse, on a trip to the country, met a country mouse. they soon became friends. the country mouse took his new friend into the meadows and forest. to thank his friend for the lovely time, he invited the country mouse to visit him in the town. and when the country mouse saw the cheese, cake, honey, jam and other goodies at the house, he was pleasantly surprised. right now the question i’m thinking about is, why was the country mouse surprised? figure 1: an example passage and question from mostow and chen (2009) all versions of the system use parsing and semantic role labeling to extract argument structures, with particular emphasis on modal verbs, e.g. “believe,” “fear,” etc, which are then mapped to the situation model, a semantic network. the mapping process involves a semantic decomposition step, in which modal verbs like “pretend” are represented as “x is not present in reality and person p1’s belief, but p1 wants person p2 to believe it” (chen 2009). thus the network supports inference based on lexical knowledge, but it does not encode general world knowledge1. questions are generated from this knowledge structure by filling templates (chen et al. 2009, mostow and chen 2009): • why/how did ? • what happens ? • why was/were ? • why ? the work of chen et al. is much more closely aligned with psychological theories of question generation than work that generates questions from single sentences (ali et al. 2010, heilman and smith 2009, 2010, kalady et al. 2010, mannem et al. 2010, pal et al. 2010, sag and flickinger 2008, varga and ha 2010, wyse and piwek 2009, yao and zhang 2010) in at least two ways. by attempting to comprehend the text before asking questions, the work of chen et al. acknowledges that question asking and comprehension are inextricably linked (collins and gentner 1980, graesser and person 1994, hilton 1990, olson et al. 1985). additionally, chen et al.’s work underscores the pedagogical nature of questions, namely that question asking can involve a tutor on behalf of a student as well as the student alone. the primary goal of the research reported here is to continue the progress made by chen et al. in connecting the computational work on question generation with previous research on question asking in the psychological literature. this study attempts to close the psychological/computational gap in two ways. first, we review work from the psychological literature that is relevant to a computational account of question generation. secondly, we present a model of tutorial question generation derived from previous psychological models of question asking. thirdly, we present an evaluation of this model using a version of the question generation shared task and evaluation challenge metrics (rus et al. 2010a; rus et al., this volume) augmented for the pedagogical nature of this task. 2. psychological framework in tutorial contexts, question generation by both human instructors (tutors) and students has been observed by graesser, person, and colleagues (graesser and person 1994, graesser et al. 1995, person 1. world knowledge is often included in situation models (mcnamara and magliano 2009). 77 olney, graesser and person table 1: graesser, person, and huber (1992)’s question categories question category abstract specification 1. verification is x true or false? did an event occur? 2. disjunctive is x, y, or z the case? 3. concept completion who? what? when? where? 4. feature specification what qualitative properties does x have? 5. quantification how much? how many? 6. definition questions what does x mean? 7. example questions what is an example of a category? 8. comparison how is x similar to or different from y? 9. interpretation what can be inferred from given data? 10. causal antecedent what state causally led to another state? 11. causal consequence what are the consequences of a state? 12. goal orientation what are the goals behind an agent action? 13. procedural what process allows an agent to reach a goal? 14. enablement what resource allows an agent to reach a goal? 15. expectation why did some expected event not occur? 16. judgmental what value does the answerer give to an idea? 17. assertion a declarative statement that indicates the speaker does not understand an idea. 18. request/directive the questioner wants the listener to perform some action. et al. 1994). intuitively, though questions are being asked by both student and tutor, the goal behind each question depends greatly upon who is asking it. for example, human tutor questions, unlike student questions, do not signify knowledge deficits and are instead attempts to facilitate student learning. graesser et al. (1992) present an analysis of questions that occur during tutoring sessions that decouples the surface form of the question, the content of the question, the mechanism that generated the question, and the specificity of the question. each of these are independent dimensions along which questions may vary. although deep theoretical issues are addressed in the graesser et al. (1992) analysis, in what follows we present the analysis as a descriptive or taxonomic framework to organize further discussion of human question generation. 2.1 content a question taxonomy can focus on surface features, such as the question stem used, or alternatively can focus on the conceptual content behind the question. as discussed by graesser et al. (1992), there are several reasons to prefer a conceptual organization. one reason is that question stems underspecify the nature of the question. for example, “what happened” requires a causal description of events while “what is that” requires only a label or definition, even though both use the same stem, “what.” likewise questions can be explicitly marked with a question mark, but they can also be pragmatically implied, e.g. “i don’t understand gravity.” these distinctions motivate table 1, which draws on previous research and has been validated in multiple studies (graesser and person 1994, graesser et al. 1995, person et al. 1994). in particular, question types 10 through 15 are highly correlated with the deeper levels of cognition in bloom’s taxonomy of educational objectives in the cognitive domain (bloom 1956, graesser and person 1994). thus one key finding along this dimension of analysis is that generation of optimal questions for student learning should be sensitive to conceptual content rather than merely surface form. 2.2 mechanism graesser et al. (1992) specify four major question generation mechanisms. the first of these is knowledge deficit, which, being information seeking, generates learner questions more often than tutor questions (graesser and person 1994). the second mechanism is common ground. questions 78 question generation from concept maps table 2: graesser, person, and huber (1992)’s question generation mechanisms correction of knowledge deficit 1. obstacle in plan or problem solving 2. deciding among alternatives that are equally attractive 3. gap in knowledge that is needed for comprehension 4. glitch in explanation of an event 5. contradiction monitoring common ground 6. estimating or establishing common ground 7. confirmation of a belief 8. accumulating additional knowledge about a topic 9. comprehension gauging 10. questioner’s assessment of answerer’s knowledge 11. questioner’s attempt to have answerer generate an inference social coordination of action 12. indirect request 13. indirect advice 14. asking permission 15. offer 16. bargaining control of conversation and attention 17. greeting 18. reply to summons 19. change in speaker 20. focus on agent’s actions 21. rhetorical question 22. gripe generated by this mechanism seek to maintain mutual knowledge and understanding between tutor and student, e.g. “have you covered factorial designs?”. the third mechanism coordinates social actions including requests, permission, and negotiation. tutors ask these kinds of questions to engage the student in activities with pedagogical significance. the fourth and final question generation mechanism is conversation-control, by which the tutor and student affect the flow of the conversation and each other’s attention, e.g. greetings, gripes, and rhetorical questions. specific examples of the four question generation mechanisms are given in table 2. arguably, the most important question generation mechanism for tutors is the common ground mechanism, accounting for more than 90% of tutor questions in two different domains, with the vast majority of these being student assessment questions (graesser and person 1994). student assessment questions probe student understanding to confirm that it agrees with the tutor’s understanding. since the dimension of question content is independent of question mechanism, any of the questions in table 1 can be generated as a student assessment question. an interesting question for future research is the extent to which a question generated by a specific mechanism can have pragmatic effects consistent with other mechanisms, or alternatively, whether multiple mechanisms can be involved in the generation of a question. intuitively, a tutor question can have multiple pragmatic effects on the student such as focusing attention, highlighting the importance of a topic, stimulating student self-assessment on that topic, assessing the student on that topic, creating an opportunity for student knowledge construction, or creating a context for further instruction (answer feedback, remediation, etc.). some of these effects can be manifested by the surface realization of the question, which we turn to next. 2.3 specificity when the information sought by a question is explicitly marked, then that question is said to have a high degree of specificity. for example, “what is the first element of the list a, b, c?” is highly 79 olney, graesser and person table 3: question specificity question type specificity example pumps low can you say more? hints medium what’s going on with friction here? prompts high what’s the force resisting the sliding motion of surfaces? explicit and requires only knowledge of how to apply a first operator to a list. contrastingly, “what’s the first letter of the alphabet?” presupposes that the listener has the world knowledge and dialogue context to correctly identify the implied list, e.g. the latin alphabet rather than the cyrillic. questions can be even less specified “what is the first letter?” or “what is it?” requiring greater degrees of common ground between the tutor and student. previous research has characterized specificity as being low, medium, or high to allow reliable coding for discriminative analyses (graesser et al. 1992, graesser and person 1994, person et al. 1994). however, it appears that low, medium, and high specificity questions map onto the kinds of questions that tutors ask in naturalistic settings known as pumps, hints, and prompts (graesser et al. 1995, hume et al. 1996) as is shown in table 3. questions at the beginning of the table provide less information to the student than questions towards the end of the table. while questions generated via the student assessment mechanism described in section 2.2 all have the consequence of creating an opportunity to assess student knowledge, there are several other possible effects. first, a tutorial strategy that asks more specific questions only when a student is floundering promotes active construction of knowledge (graesser et al. 1995, chi et al. 2001). secondly, a very specific question, like a prompt, focuses attention on the word prompted for and highlights its importance (d’mello et al. 2010). thirdly, a less specific question can lead to an extended discussion, creating a context for further instruction (chi et al. 2008). this list of effects is not meant to be exhaustive, but rather it highlights that there are many desirable effects that can be obtained by varying the specificity of tutorial questions. 3. psychological models the psychological framework outlined in section 2 has led to detailed psychological models of question asking. however, question asking in students has received more attention than in tutors because deep student questions are positively correlated with their test scores (graesser et al. 1995). thus there is considerable interest in scaffolding students to generate deep questions (beck et al. 1997, bransford et al. 1991, brown 1988, collins 1988, dillon 1988, edelson et al. 1999, palinscar and brown 1984, piaget 1952, pressley and forrest-pressley 1985, scardamalia and bereiter 1985, schank 1999, zimmerman 1989). indeed there is an extensive literature investigating the improvements in the comprehension, learning, and memory of technical material that can be achieved by training students to ask questions during comprehension (ciardiello 1998, craig et al. 2006, davey and mcbride 1986, foos 1994, gavelek and raphael 1985, king 1989, 1992, 1994, odafe 1998, palinscar and brown 1984, rosenshine et al. 1996, singer and donlan 1982, wong 1985). rosenshine et al. (1996) present a meta-analysis of 26 empirical studies investigating question generation learning effects. a cognitive computational model of question asking has been developed by graesser and colleagues (graesser and olde 2003, otero and graesser 2001). the model is called preg, which is a root morpheme for “question” in the spanish language. according to the preg model, cognitive disequilibrium drives the asking of questions (berlyne 1960, chinn and brewer 1993, collins 1988, festinger 1962, flammer 1981, graesser et al. 1996, graesser and mcmahen 1993, schank 1999). thus the preg model primarily focuses on the knowledge deficit mechanisms of table 2, and in its current form is most applicable to student-generated questions. 80 question generation from concept maps part abdomenarthropod posterior has-part is-a has-property figure 2: a concept map fragment. key terms have black nodes. preg has two primary components. the first is a set of production rules that specify the categories of questions that are asked under particular conditions (i.e., content features of text and knowledge states of individuals). the second component is a conceptual graph, which is a particular instantiation of a semantic network (graesser and clark 1985, sowa 1992). in this formulation of conceptual graphs, nodes themselves can be propositions, e.g. “a girl wants to play with a doll,” and relations are (as much as possible) limited to a generic set of propositions for a given domain. for example, one such categorization consists of 21 relations including is-a, has-property, has-consequence, reason, implies, outcome, and means (gordon et al. 1993). a particular advantage of limiting relations to these categories is that the categories can then be set into correspondence with the question types described in table 1 for both the purposes of answering questions (graesser and franklin 1990) as well as generating them (gordon et al. 1993). in one study, the predictions of preg were compared to the questions generated by middle and high school students reading science texts (otero and graesser 2001). the preg model was able to not only account for nearly all of the questions the students asked, but it was also able to account for the questions the students didn’t ask. the empirical evidence suggests that conceptual graph structures provide sufficient analytical detail to capture the systematic mechanisms of question asking. 4. computational model the previous discussion in sections 2 and 3 offers substantial guidance in the design of a computational model of tutorial question generation. section 2 describes a conceptual question taxonomy in which deeper questions are correlated with deeper student learning, question generation mechanisms behind tutorial questions, and the role of specificity in question asking. section 3 presents a psychological model, preg, that builds on this framework. although the preg model is not completely aligned with the present objective of generating tutorial questions, it provides a roadmap for a twostep modeling approach. in the first step, conceptual graphs are created for a particular domain. automatic extraction from text is desirable because hand-authoring knowledge representations for intelligent tutoring systems is extremely labor intensive (murray 1998, corbett 2002, aleven et al. 2006). secondly, conceptual graphs are used to generate student-assessment questions that span the question types of table 1 and the specificity levels of table 3. we have previously presented an approach for extracting concept maps from a biology textbook (olney 2010, olney et al. 2010, 2011) that we will briefly review. concept maps are similar to conceptual graph structures, but they are generally less structured (fisher et al. 2000, mintzes et al. 2005). our concept map representation has two significant structural elements. the first is key terms, shown as black nodes in figure 2. these are terms in our domain that are pedagogically 81 olney, graesser and person significant. only key terms can be the start of a triple, e.g. abdomen is-a part. end nodes can contain key terms, other words, or complete propositions. the second central aspect of our representation is labeled edges, shown as boxes in figure 2. as noted by fisher et al. (2000), a small set of edges can account for a large percentage of relationships in a domain. in comparison, the conceptual graph structures of gordon et al. (1993) contain nodes that can be concepts or statements, and each node is categorized as a state, event, style, goal, or action. their nodes may be connected with directional or bi-directional arcs from a fixed set of relations. however, conceptual graphs and our concept maps share a prescribed set of edges, and as will be discussed below, this feature facilitates linkages between the graph representations and question asking/answering (graesser and clark 1985, graesser and franklin 1990, gordon et al. 1993). in the next few sections we provide an overview of the concept map extraction process. 4.1 key terms as discussed by vanderwende (2008), one goal of question generation systems should be to ask important questions. in order to do so reliably, one must identify the key ideas in the domain of interest. note that at this level, importance is defined as generally important, rather than important with respect to a particular student. defining important questions relative to a particular model of student understanding is outside the scope of this study. though general purpose key term extraction procedures have been proposed (medelyan et al. 2009), they are less relevant in a pedagogical context where key terms are often already provided, whether in glossaries (navigli and velardi 2008), or textbook indices (larrañaga et al. 2004). to develop our key terms, we used the glossary and index from a textbook in the domain of biology (miller and levine 2002) as well as the keywords given in a test-prep study guide (cypress curriculum services 2008). this process yielded approximately 3,000 distinct terms, with 2,500 coming from the index and the remainder coming from the glossary and test-prep guide. thus we can skip the keyword extraction step of previous work on concept map extraction (valerio and leake 2008, zouaq and nkambou 2009). 4.2 edge relations edge relations used in conceptual graphs typically depict abstract and domain-independent relationships (graesser and clark 1985, gordon et al. 1993). however, previous work suggests that while a large percentage of relations in a domain are from a small set, new content can drive new additions to that set (fisher et al. 2000). in order to verify the completeness of our edge relations, we undertook an analysis of concept maps from biology. we manually analyzed 4,371 biology triples available on the internet2. these triples span the two topics of molecules & cells and population biology. because these two topics represent the extremes of levels of description in biology, we presume that their relations will mostly generalize to the levels between them. a frequency analysis of these triples revealed that 50% of all relations are is-a, has-part, or has-property. in the set of 4,371 biology triples, only 252 distinct relation types were present. we then manually clustered the 252 relation types into 20 relations. the reduction in relation types lost little information because the original data set had many subclassing relationships, e.g. part-of had the subclasses composed of, has organelle, organelle of, component in, subcellular structure of, and has subcellular structure. since this subclassing is often recoverable by knowing the node type connected, e.g. an organelle, the mapping of the 252 relation types to 20 relations loses less information than might be expected. for example, weighted by frequency, 75% of relations in the 4,371 triples can be accounted for by the 20 relations by considering subclasses as above. if one additionally allows verbal relations to be subclassed by has-consequence, e.g. eats becomes has-consequence eats x where the subclassed verbal relation is merged with the end node 2. http://www.biologylessons.sdsu.edu 82 question generation from concept maps x, the 20 relations account for 91% of triples by frequency. likewise applying part-of relationships to processes, as illustrated by interphase part-of cell cycle raises coverage to 95%. the remaining relations that don’t fit well into the 20 clusters tend to be highly specific composites such as often co-evolves with or gains phosphate to form; these are subclassed under has-property in the same way as verbal relations are subclassed under has-consequence. of the 20 relations, 8 overlap with domain-independent relationships described in psychological research (graesser and clark 1985, gordon et al. 1993) and an additional 4 are domain-general adjunct roles such as not and direction. the remaining 8 are biology and education specific, like convert and definition. these 20 relations were augmented with 10 additional relations from gordon et al. and adjunct roles, including after, contrast, enable, has-consequence, lack, produce, before, convert, example, has-part, location, purpose, combine, definition, extent, has-property, manner, reciprocal, connect, direction, follow, implies, not, require, contain, during, function, isa, possibility, and same-as, in order to maximize coverage on unseen biology texts. a detailed analysis of the overlap between the clustered relations, relations from gordon et al. (1993), and adjunct relations is presented in olney et al. (2011). 4.3 automatic extraction the automatic extraction process creates triples matching the representational scheme described in the previous sections. since each triple begins with a key term, multiple triples starting with that term can be indexed according to that term. alternatively, one can consider each key term to be the center of a radial graph of triples. triples containing key terms as start and end nodes bridge these radial graphs. input texts for automatic extraction for the present study consisted of the textbook and glossary described in section 4.1 as well as a set of 118 mini-lectures written by a biology teacher. each of these mini-lectures is approximately 300-500 words long and is written in an informal style. the large number of anaphora in the mini-lectures necessitated using an anaphora resolution algorithm. otherwise, triples that start with an anaphor, e.g. “it is contributing to the destruction of the rainforest,” will be discarded because they do not begin with a key term. thus the first step in automatic extraction for the mini-lectures is anaphora resolution. empronoun is a state of the art pronoun anaphora resolution utility that has an accuracy of 68% on the penn treebank, roughly 10% better performance than previous methods implemented in javarap, open-nlp, bart and guitar (charniak and elsner 2009). using the empronoun algorithm, each lecture was rewritten by replacing anaphora with their respective reference, e.g. replacing “it” with “logging” in the previous example. the second step is semantic parsing. the lth srl3 parser is a semantic role labeling parser that outputs a dependency parse annotated with propbank and nombank predicate/argument structures (johansson and nugues 2008, meyers et al. 2004, palmer et al. 2005). for each word token in a parse, the parser returns information about the word token’s part of speech, lemma, head, and relation to the head. moreover, it uses propbank and nombank to identify predicates in the parse, either verbal predicates ( propbank) or nominal predicates (nombank), and their associated arguments. for example, consider the sentence, “many athletes now use a dietary supplement called creatine to enhance their performance.” the lth srl parser outputs five predicates for this sentence: use (athletes/a0) (now/am-tmp) (supplements/a1) (to/a2) supplement (dietary/a1) (supplement/a2) called (supplement/a1) (creatine/a2) enhance (supplement/a0) (performance/a1) 3. the swedish “lunds tekniska högskola” translates as “faculty of engineering.” 83 olney, graesser and person table 4: example predicate maps predicate pos edge relation frequency start end have.03 v has property 1,210 a0 span use.01 v use 1,101 a0 span produce.01 v produce 825 a0 span call.01 v has definition 663 a1 a2 performance (their/a0) three of these predicates are verbal predicates: “use,” “called,” and “enhance.” verbal predicates often, but not always, have agent roles specified by a0 and patient roles specified by a1. however, consistent generalizations in role labeling are often lacking, particularly for arguments beyond a0 and a1 (palmer et al. 2005). moreover, the presence of a0 and a1 can’t be guaranteed: although “use” and “enhance” both have a0 and a1, “called” has no a0 because there’s no clear agent doing the calling in this situation. passive verbs frequently have no a0. finally, verbal adjuncts can specify adverbial properties such as time, here specified by am-tmp. two of the predicates are nominal predicates: “supplement” and “performance.” nominal predicates have much more complex argument role assignment rules than verbal predicates (meyers et al. 2004). predicates that are nominalizations of verbs, such as “performance,” have roles more closely corresponding to verbal predicate roles like a0 and a1. for other kinds of nominal predicates, overt agents taking a0 roles are less common. as a result, many nominal predicates have only an a1 filling a theme role, as “dietary” is the theme of supplement. the third step in concept map extraction is the actual extraction step. we have defined four extractor algorithms that target specific syntactic/semantic features of the parse, is-a, adjectives, prepositions, and predicates. each extractor begins with an attempt to identify a key term as a possible start node. since key terms can be phrases, e.g. “homologous structure,” the search for key terms greedily follows the syntactic dependents of a potential key term while applying morphological rules. for example, in the sentence “homologous structures have the same structure,” “structures” is the subject of the verb “have” and a target for a key term. the search process follows syntactic dependents of the subject to map “homologous structures” to the known key term “homologous structure.” in many cases, no key term will be found, so the prospective triple is discarded. several edge relations are handled purely syntactically. is-a relations are indicated when the root verb of the sentence is “be,” but not a helping verb. is-a relations can create a special context for processing additional relations. for example, in the sentence, “an abdomen is a posterior part of an arthropod’s body,” “posterior” modifies “part,” but the desired triple is abdomen has-property posterior. this is an example of the adjective extraction algorithm running in the context of an is-a relation. prepositions can create a variety of edge relations. for example, if the preposition is in and has a loc dependency relation to its head (a locative relation), then the appropriate relation is location, e.g. “by migrating whales in the pacific ocean” becomes whales location in the pacific ocean. relations from propbank and nombank require a slightly more sophisticated approach. as illustrated in some of the preceding examples, not all predicates have an a0. likewise not all predicates have patient/instrument roles like a1 and a2. the variability in predicate arguments makes simple mapping, e.g. a0 is the start node, predicate is the edge relation, and a1 is the end node, unrealistic. therefore we created a manual mapping between predicates, arguments, and edge relations for every predicate that occurred more than 40 times in the corpus (358 predicates total). table 4 lists the four most common predicates and their mappings. the label “span” in the last column indicates that the end node of the triple should be the text dominated by the predicate. for example, “carbohydrates give cells structure.” has ao “carbohydrates” and a1 “structure.” however, it is more desirable to extract the triple carbohydrates 84 question generation from concept maps has-property give cells structure than carbohydrates has-property structure. end nodes based on predicate spans tend to contain more words and therefore have closer fidelity to the original sentence. finally, we apply some filters to remove triples that are either not particularly useful for question generation or appear to be from mis-parsed sentences. we apply filters on the back end rather than during concept map extraction because some of the filtered triples are useful for other purposes besides question generation, e.g. student modeling and assessment. three of the filters are straightforward and require little explanation: the repetition filter, the adjective filter, and the nominal filter. the repetition filter considers the number of words in common between the start and end nodes. if the number of shared words is more than half the words in the end node, the triple is filtered. this helps alleviate redundant triples such as cell has-property cell. the adjective filter removes any triple whose key term is an adjective. these triples violate the assumption by the question generator that all key terms are nouns. for example, ’reproducing’ might be a key term, but in a particular triple it might function as an adjective rather than a noun. these cases are often caused by mis-parsed sentences. finally, the nominal filter removes all nombank predicates except has-part predicates, because these often have span end nodes and so contain themselves, e.g. light has-property the energy of sunlight. the most sophisticated filter is the likelihood ratio filter. this filter measures the association between the start and end node using likelihood ratios (dunning 1993) and a chi-square significance criterion to remove triples with insignificant association. words from the end node that have low log entropy are removed prior to calculation, and the remaining words from start and end nodes are pooled. the significance criterion for the chi-square test is .0001. this filter weeds out start and end nodes that do not have a strong association. in the document set used for the present study, 43,833 of the originally extracted triples were filtered to a set of 19,143 triples. the filtered triples were distributed around 1,165 start nodes out of approximately 3,000 possible key terms. the five most connected key terms in the filtered set are cell, plant, organism, species, and animal, which collectively account for 18% of the total connections. of the possible 30 edge relations, only 18 were present in the filtered triples, excluding before, convert, direction, during, extent, follow, function, implies, manner, possibility, reciprocal, and after. the top five edge relations extracted were has-property, has-consequence, has-part, location, and is-a, making up roughly 82% of the total relations. by themselves, the relations has-property, is-a, and has-part make up 56% of all edge relations, which is consistent with human concept maps for biology domains (fisher et al. 2000). 4.4 question generation our question generation approach uses the concept map described in previous sections to generate questions. questions may either be generated from individual triples or by combinations of triples. question generation from individual triples is very similar to generating questions from individual sentences. both cases ignore how information is structured across the domain. question generation from combinations of triples, in contrast, introduces a limited form of reasoning over the concept map knowledge representation. we consider each of these in turn. 4.4.1 generation from individual triples question generation using individual triples can generate questions of varying specificity as described in section 2.3. in general, the less specific a question is, the easier it is to generate the question with a template. for example, pumps, the least specific question type, can be generated with a fixed bag of expressions like, “what else?” or “can you say more?” although this level of specificity is trivial, the generation of hints and prompts warrants some discussion. hints have an intermediate level of specificity. our approach to hint generation makes use of parametrized question templates that consider the start node and edge relation of a triple. some example hint templates of the 14 that were used are given in table 5. question templates are 85 olney, graesser and person table 5: example hint question templates edge relation question template ? and what do we know about kt? has consequence what do kt do? has definition kt, what is that? has part what do kt have? isa so kt be? selected based on the edge relation of the source triple, or a wildcard template (?) can be used. templates may have placeholders for the key term (kt) as well as verb lemmas (do, be). question templates for hints are populated in two steps. first, a determiner is added to the key term. the algorithm decides what determiner to add based on whether a determiner modified the key term in the source sentence, the key term is a mass noun, or the key term is plural. the second step involves matching the agreement features of the key term with verb lemma (if present). agreement features are derived using the specialist lexicon (browne et al. 2000), which has been previously used in the natural language generation community (gatt and reiter 2009). prompts have the highest level of specificity, because by definition prompts seek only a word or phrase for an answer (graesser et al. 1995). additionally, prompts are often generated as incomplete assertions (see table 1), e.g. “the first letter of the alphabet is...?” by making use of this assertionoriented structure as well as the single relations encoded in triples, our prompt generator attempts to avoid extremely complex questions that can be created by syntactic transformation, e.g. “where was penicillin discovered by alexander fleming in 1928?” prompt generation is a two step process. in the first step, the triple is rendered as an assertion. rendering an assertion requires assembling the start node, edge relation, and end node into a coherent sentence. constructing a well formed declarative sentence from these fragments requires managing a large number of possible cases involving tense, determiners, modifiers, passivization, and conjunction, amongst others. the assertion generation process can be considered as a black box for the purposes of the current discussion. if the edge relation is not, then the declarative sentence is transformed into a verification question. otherwise the declarative sentence is scanned for the word with the highest log entropy weight (dumais 1991). this word is substituted by “what” to make the final prompt. by using log entropy as a criterion for querying, we are attempting to maximize the probability that the deleted item is relevant in the domain. log entropy also gives a principled way of selecting amongst multiple key terms in a sentence, though in our implementation any word was a target. 4.4.2 generation from multiple triples one significant advantage to building a knowledge representation of text prior to generating questions is that knowledge may be integrated across the text. in our model, integration naturally falls out of the restriction that all triples must start with a key term. as long as triples begin with the same key term, they may be considered as relating to the same concept even if they are many pages apart. the additional structure this provides allows for some interesting questioning strategies. we briefly define three questioning strategies that we call contextual verification, forced choice, and causal chain, respectively. both contextual verification questions and forced choice questions compare and contrast features of two related concepts. empirical evidence suggests that this kind of discriminative learning, in which salient features are compared and contrasted across categories, is an important part of concept learning (tennyson and park 1980). there are some similarities between this strategy and the teaching with analogies philosophy (glynn 2008). in both cases, the student is reminded of some86 question generation from concept maps thing they know in order to learn something new. the context primes existing student knowledge and helps the student associate that knowledge with the situation presented in the question. contextual verification questions present a context and then ask a verification question. an example contextual verification question from our system is “an ecosystem is a community. is that true for a forest?” in order to generate contextual verification questions, we index all of the key nodes that are subtypes (is-a) of a common node. for example, cat and dog are both subtypes of animal. each key node has an associated set of triples that can be intersected with the other sets to discover common and unique properties. for example, cats and dogs both have tails and four legs, but only cats chase mice. common properties can be used to generate contextual verification questions that should be answered positively, e.g. “cats have a tail. is that true for dogs?” while unique properties can be used to generate questions that should be answered negatively “cats chase mice. is that true for dogs?” the context component is generated as an assertion as described previously for prompts. the verification question itself is easily generated using a question template as described previously for hints. forced choice questions in contrast can only apply to the disjunctive case in which a property is not held in common by two subtypes of a common node. an example generated by our system is “what resides in skin, melanin or phytochrome?” the triples used to generate forced choice questions are selected using the same method as for contextual verification questions, and likewise may be generated using a question template. causal chain questions are based on causal relationships that connect multiple key terms. previous research has investigated how students can learn an underlying causal concept map and use it to solve problems and construct explanations (mills et al. 2004). qualitatively speaking, causal relations have a +/valence, indicating a direct or inverse relationship of one concept/variable to another. for example, carrots may be directly causally related (+) to rabbits, such that an increase in the number of carrots leads to an increase in the number of rabbits. causal chain questions are constructed by joining two triples from the concept map, such that the end of the first triple is the same as the beginning of the second triple. an example causal chain question produced by our system is “how do bacteria produce energy?” which bridges the key terms of bacteria, food, and energy. since causal chaining requires the edge relations be transitive, we restricted causal chaining to the edge relations produce and has-consequence. 5. evaluation 5.1 method we conducted an evaluation of the question generation system described in section 4. three judges who were experts in questions and pedagogy evaluated questions generated by the system. each question was rated on the following five dimensions: whether the sentence was of the target type (question type), the relevance of the question to the source sentence (relevance), the syntactic fluency of the question (fluency), the ambiguity of the question (ambiguity), and the pedagogical value of the question (pedagogy). the first four of these dimensions were derived from the question generation shared task evaluation challenge (qgstec) (rus et al. 2010b; rus et al., this volume). a consistent four item scale was used for all dimensions except question type, which was binary. an example of the four item scale is shown in table 6. the full set of scales is included in the appendix. the evaluation set rated by the judges was constructed using the three text sources described in section 4.3: the textbook, glossary, and mini-lectures. each judge blindly rated the same 30 hint and prompt questions from each of these sources, for a total of 60 questions from each source. in addition, each judge rated approximately 30 contextual verification, forced choice, and causal chain questions. since these generation methods are hard hit by sparse data, all three text sources were used to generate these three question types. each hint or prompt was preceded by its source sentence so that comparative judgments, like relevance, could be made: 87 olney, graesser and person table 6: rating scale for relevance score criteria 1 the question is completely relevant to the input sentence. 2 the question relates mostly to the input sentence. 3 the question is only slightly related to the input sentence. 4 the question is totally unrelated to the input sentence. table 7: inter-rater reliability scale judge pair cronbach’s α question type ab .43 relevance ab .82 fluency ab .79 ambiguity ab .74 pedagogy bc .80 an antheridium is a male reproductive structure in some algae and plants. tell me about an antheridium. (hint) an antheridium is what? (prompt) the contextual verification, forced choice, and causal chain questions were preceded by the two source sentences for their respective triples: enzymes are proteins that act as biological catalysts. hemoglobin is the protein in red blood cells that carries oxygen. enzymes are proteins. is that true for hemoglobin? inter-rater reliability was calculated on each of the five measures, using a two-way random effect model to calculate average measure intra-class correlation. cronbach’s α for each measure is presented in table 7. the pair of judges with the highest alpha for each measure were used to calculate composite scores for each question in later analyses. most of the reliability scores in table 7 are close to .80, which is considered satisfactory reliability. however, reliability for question type was poor at α = .43. this smaller value is attributable to the fact that these judgments were binary and to the conservative test statistic: proportion agreement for question type was .80 for hints and .75 for prompts. 5.2 results our two guiding hypotheses in the evaluation were that question source and question type would affect ratings scores. question source is likely to affect ratings because sentences from the texts vary in terms of their syntactic complexity. for example, sentences from the glossary have the prototype structure “an x is a y that ...” while sentences from the other sources have no such restrictions. additionally, anaphora resolution was used on mini-lectures but not on the other texts. this could affect questions generated from mini-lectures negatively by introducing errors in anaphora resolution or positively by removing ambiguous anaphora from questions. likewise question categories vary considerably in the complexity of processes used to generate them. hints are more template based than prompts, and both of these question types are simpler than contextual verification, forced choice, and causal chain questions that construct questions over multiple sentences. in the following sections we present statistical tests of significance between these conditions. 88 question generation from concept maps table 8: mean ratings across question sources textbook mini-lectures glossary scale mean sd mean sd mean sd question type 1.41 0.36 1.44 0.38 1.27 0.42 relevance 2.72 1.02 2.54 0.92 1.77 1.08 fluency 2.29 1.01 2.17 0.98 2.25 1.14 ambiguity 3.35 0.89 3.19 0.88 2.79 0.82 pedagogy 2.88 1.07 3.09 1.03 2.25 1.18 5.2.1 results of question source a one-way anova was used to test for relevance differences among the three question sources. relevance differed significantly across the three sources, f(2,177) = 14.80, p = .0001. the effect size, calculated using cohen’s f2, was .17. scheffé post-hoc comparisons of the three sources indicate that the glossary questions (m = 1.78, sd = 1.08) had significantly better relevance than the textbook questions (m = 2.72, sd = 1.02), p = .0001 and significantly better relevance than the mini-lecture questions (m = 2.54, sd = 0.92), p = .0001. this pattern was repeated for ambiguity, f(2,177) = 6.65, p = .002. the effect size, calculated using cohen’s f2, was .08. scheffé post-hoc comparisons indicate that the glossary questions (m = 2.79, sd = 0.82) had significantly lower ambiguity than the textbook questions (m = 3.35, sd = 0.89), p = .002 and the mini-lecture questions (m = 3.19, sd = 0.88), p = .043. the same pattern held for pedagogy, f(2,177) = 9.60, p = .0001. the effect size, calculated using cohen’s f2, was .11. scheffé post-hoc comparisons indicate that the glossary questions (m = 2.25, sd = 1.18) had significantly better pedagogy than the textbook questions (m = 2.88, sd = 1.08), p = .008 and the mini-lecture questions (m = 3.09, sd = 1.03), p = .0001. one-way anovas were used to test for question type and fluency differences among the three question sources, but neither question type nor fluency significantly differed across the three sources of questions, p = .05. mean ratings across question sources are presented in table 8. 5.2.2 results of question category a one-way anova was used to test for question type differences across the five question categories. question type differed significantly across the five categories, f(4,263) = 12.00, p = .0001. the effect size, calculated using cohen’s f2, was .18. scheffé post-hoc comparisons of the five question categories indicate that the prompt questions (m = 1.55, sd = 0.42) were significantly less likely to be of the appropriate type than the hint questions (m = 1.2, sd = 0.26), p = .0001 and significantly less likely to be of the appropriate type than the forced choice questions (m = 1.27, sd = 0.37), p = .009. a one-way anova was used to test for fluency differences across the five categories. fluency differed significantly across the five categories, f(4,263) = 29.40, p = .0001. the effect size, calculated using cohen’s f2, was .45. scheffé post-hoc comparisons of the five question categories indicate that the hint questions (m = 1.56, sd = 0.74) were significantly more fluent than prompts (m = 2.92, sd = 0.83), p = .0001, forced choice questions (m = 2.64, sd = 1.04), p = .0001, contextual verification questions (m = 2.25, sd = 1.02), p = .007, and causal chain questions ( m = 2.40, sd = 0.98), p = .0001. post-hoc comparisons further indicated that contextual verification questions were significantly more fluent that prompts, p = .01. a one-way anova was used to test for pedagogy differences across the five question categories. pedagogy differed significantly across the five categories, f(4,263) = 8.10, p = .0001. the effect size, calculated using cohen’s f2, was .12. scheffé post-hoc comparisons of the five question categories indicate that the hint questions (m = 2.39, sd = 1.15) were significantly more pedagogic than prompts (m = 3.09, sd = 1.03), p = .001, forced choice questions (m = 3.21, sd = 1.07), p = .015, 89 olney, graesser and person table 9: mean ratings for single triple question types prompt hint scale mean sd mean sd question type 1.55 0.42 1.20 0.26 relevance 2.45 1.12 2.24 1.04 fluency 2.92 0.83 1.56 0.74 ambiguity 3.12 0.97 3.10 0.81 pedagogy 3.09 1.03 2.39 1.15 table 10: mean ratings for multiple triple question types forced contextual causal choice verification chain scale mean sd mean sd mean sd question type 1.27 0.37 1.42 0.37 1.37 0.29 relevance 2.52 0.88 2.88 0.95 2.13 0.63 fluency 2.64 1.04 2.25 1.02 2.40 0.98 ambiguity 3.07 0.89 3.13 0.83 2.85 0.77 pedagogy 3.21 1.07 3.33 1.02 3.18 1.05 contextual verification questions (m = 3.33, sd = 1.02), p = .002, and causal chain questions ( m = 3.18, sd = 1.05), p = .018. one-way anovas were used to test for relevance and ambiguity differences among the five question categories, but neither relevance nor ambiguity significantly differed across the five question categories, p = .05. mean ratings across question categories are presented in table 9 and table 10. 6. discussion ideally, the results in section 5.2 would be interpreted relative to previous results in the literature. however, since no official results from the 2010 question generation shared task evaluation challenge (qgstec) appear to have been released, comparison to previous results is rather limited. the only work that reports informal results using the qgstec scales appears to be the wlv system (varga and ha 2010). in this study, inter-annotator agreement for relevance and fluency in cohen’s κ is .21 and .22, for average ratings of 2.65 and 2.98 respectively. in comparison, our agreement on these scales in κ was .82 and .86. as no other inter-annotator agreement appears on the qgstec scales appears to have been reported elsewhere, one contribution of our study is to show that high reliability can be achieved on these five scales. although our range of relevance, 2.13-2.88, and our range of fluency, 1.56-2.92, appear to compare favorably to those of varga and ha (2010), it’s not clear which of our question categories (if any) map onto theirs. therefore it is more informative to discuss the comparisons of question source and question category and reflect on how these comparisons might inform our psychologically-based model of question generation. overall, most means in tables 8, 9, and 10 fall in the indeterminate range between slightly correct and slightly incorrect. notable exceptions to this trend are pedagogy for multiple triple categories and prompts, which tend towards slightly unreasonable, and ambiguity scores across tables that tend to slightly ambiguous. encouragingly, these few means are closer in score indicating they are slightly unreasonable rather than completely unreasonable. one sample t-tests confirm that all these means are statistically different from completely unreasonable, p = .05. this is particularly encouraging for the question categories that make use of multiple triples, contextual verification, forced choice, and causal chain, because they have more opportunities to make errors. 90 question generation from concept maps the comparisons between question sources in section 5.2.1 support the conclusion that our model converts glossary sentences into better questions than sentences from the textbook or mini-lectures, although the effect sizes are small, with f2 ranging from .08-.17. this result supports the hypothesis that glossary sentences, having less syntactic variability than sentences from the other texts, don’t require as complex algorithms to generate questions. glossary sentences produce acceptably relevant questions, with a mean score of 1.77. on the other dimensions, glossary sentences are roughly in the middle of the scale between somewhat correct and somewhat incorrect. it is somewhat surprising that there were no differences between the textbook and the minilectures, because the mini-lectures used the empronoun anaphora resolution algorithm (charniak and elsner 2009). anaphora resolution errors have been previously reported as significant detractors in question generation systems that use knowledge representations (chen 2009, chen et al. 2009, mostow and chen 2009). however, these previous studies did not use empronoun, which has roughly 10% better performance than other state of the art methods (charniak and elsner 2009). since no significant differences were found in our evaluation, we tentatively conclude that the empronoun algorithm, by resolving out-of-context pronouns, boosted scores as much as it hurt them. the comparison between question categories in section 5.2.2 is perhaps the richest and most interesting comparison in our evaluation. the first main finding of the question category comparison is that prompt questions are less likely to be of the correct type than hints or forced choice. the similarity in ambiguity ratings between prompts (m = 3.12) and hints (m = 3.10) suggests that prompts are not very specific and might be more like hints than prompts. the second major finding in section 5.2.2 is that hints are more fluent and have better pedagogy than all the other categories. the success of hints in this regard is probably due to their template-based generation, which is the simplest question generation method of the five methods evaluated. however, it’s also worth noting what we didn’t find, which are additional differences between the questions generated from multiple triples and those generated from a single triple. all three multiple triple question categories in table 10 have neutral fluency and relevance scores, suggesting that they are on average acceptable questions. additionally, no differences were found between single triple and multiple triple question categories for the measures of relevance and ambiguity. on the other hand, pedagogy scores for multiple triple question categories are less reasonable, indicating that more work needs to be done before these questions can be used in an e-learning environment. 7. conclusion the major goal of this study was to bridge the gap between psychological theories of question asking and computational models of question generation. the model we presented in section 4 is heavily grounded in several decades of psychological research on question asking and answering. our objective is not to build linkages to psychological theory merely for the sake of it. on the contrary, we believe that the psychological mechanisms that have been proposed for question asking hold great promise for future computational models of question generation. by generating plausible questions from concept maps, our model lends some support to this claim. the computational model we presented in this study is one step towards the larger goal of understanding what question to generate, dynamically, for an individual student. however, many additional questions need to be answered before this goal can be realized. we must track what the student knows and use pedagogical knowledge to trigger the right questions at the right time. solving these research challenges will deepen our understanding and advance the art and science of question generation. 91 olney, graesser and person appendix a. modified qgstec rating scales relevance 1 the question is completely relevant to the input sentence. 2 the question relates mostly to the input sentence. 3 the question is only slightly related to the input sentence. 4 the question is totally unrelated to the input sentence. question type 1 the question is of the target question type. 2 the type of the generated question and the target question type are different. fluency 1 the question is grammatically correct and idiomatic/natural. 2 the question is grammatically correct but does not read as fluently as we would like. 3 there are some grammatical errors in the question. 4 the question is grammatically unacceptable. ambiguity 1 the question is completely unambiguous. 2 the question is mostly unambiguous. 3 the question is slightly ambiguous. 4 the question is totally ambiguous when asked out of the blue. pedagogy 1 very reasonable 2 somewhat reasonable 3 somewhat unreasonable 4 very unreasonable references vincent aleven, bruce m. mclaren, jonathan sewall, and kenneth r. koedinger. the cognitive tutor authoring tools (ctat): preliminary evaluation of efficiency gains. in intelligent tutoring systems, pages 61–70, 2006. husam ali, yllias chali, and sadid a. hasan. automation of question generation from sentences. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 58–67, pittsburgh, june 2010. questiongeneration.org. url http: //oro.open.ac.uk/22343/. 92 question generation from concept maps isabel l. beck, margaret g. mckeown, rebecca l. hamilton, and linda kucan. questioning the author: an approach for enhancing student engagement with text. international reading association, 1997. daniel ellis berlyne. conflict, arousal, and curiosity. mcgraw-hill new york, 1960. benjamin s. bloom, editor. taxonomy of educational objectives. handbook i: cognitive domain. mckay, new york, 1956. john d. bransford, susan r. goldman, and nancy j. vye. making a difference in people’s ability to think: reflections on a decade of work and some hopes for the future. in r. j. sternberg and l. okagaki, editors, influences on children, pages 147–180. erlbaum, hillsdale, nj, 1991. ann l. brown. motivation to learn and understand: on taking charge of one’s own learning. cognition and instruction, 5:311–321, 1988. allen c. browne, alexa t. mccray, and suresh srinivasan. the specialist lexicon. technical report, national library of medicine, bethesda, maryland, june 2000. xavier carreras and llúıs màrquez. introduction to the conll-2004 shared task: semantic role labeling. in hwee tou ng and ellen riloff, editors, hlt-naacl 2004 workshop: eighth conference on computational natural language learning (conll-2004), pages 89–97, boston, massachusetts, usa, may 6 may 7 2004. association for computational linguistics. eugene charniak and micha elsner. em works for pronoun anaphora resolution. in proceedings of the 12th conference of the european chapter of the acl (eacl 2009), pages 148–156, athens, greece, march 2009. association for computational linguistics. url http://www.aclweb.org/ anthology/e09-1018. wei chen. understanding mental states in natural language. in proceedings of the eighth international conference on computational semantics, iwcs-8 ’09, pages 61–72, stroudsburg, pa, usa, 2009. association for computational linguistics. isbn 978-90-74029-34-6. url http://portal.acm.org/citation.cfm?id=1693756.1693766. wei chen, gregory aist, and jack mostow. generating questions automatically from informational text. in scotty d. craig and darina dicheva, editors, the 2nd workshop on question generation, volume 1 of aied 2009 workshops proceedings, pages 17–24, brighton, uk, july 2009. international artificial intelligence in education society. michelene t. h. chi, stephanie a. siler, heisawn jeong, takashi yamauchi, and robert g. hausmann. learning from tutoring. cognitive science, 25:471–533, 2001. michelene t. h. chi, marguerite roy, and robert g. m. hausmann. observing tutorial dialogues collaboratively: insights about human tutoring effectiveness from vicarious learning. cognitive science, 32(2):301–341, 2008. clark a. chinn and william f. brewer. the role of anomalous data in knowledge acquisition: a theoretical framework and implications for science instruction. review of educational research, 63(1):1–49, 1993. angelo v. ciardiello. did you ask a good question today? alternative cognitive and metacognitive strategies. journal of adolescent & adult literacy, 42(3):210–219, 1998. alan collins. different goals of inquiry teaching. questioning exchange, 2:39–45, 1988. 93 olney, graesser and person alan collins and dedre gentner. a framework for a cognitive theory of writing. in l. gregg and e. r. steinberg, editors, cognitive processes in writing, pages 51–72. lawrence erlbaum associates, hillsdale, new jersey, 1980. albert corbett. cognitive tutor algebra i: adaptive student modeling in widespread classroom use. in technology and assessment: thinking ahead. proceedings from a workshop, pages 50–62, washington, d.c., november 2002. national academy press. scotty d. craig, jeremiah sullins, amy witherspoon, and barry gholson. the deep-level-reasoningquestion effect: the role of dialogue and deep-level-reasoning questions during vicarious learning. cognition and instruction, 24(4):565–591, 2006. llc cypress curriculum services. tennessee gateway coach, biology. triumph learning, new york, ny, 2008. robert dale, barbara di eugenio, and donia scott. introduction to the special issue on natural language generation. computational linguistics, 24(3):345–354, september 1998. beth davey and susan mcbride. effects of question-generation training on reading comprehension. journal of educational psychology, 78(4):256–262, 1986. james thomas dillon. questioning and teaching. a manual of practice. teachers college press, new york, 1988. sidney d’mello, patrick hays, claire williams, whitney cade, jennifer brown, and andrew m. olney. collaborative lecturing by human and computer tutors. in intelligent tutoring systems, lecture notes in computer science, pages 178–187, berlin, 2010. springer. susan dumais. improving the retrieval of information from external sources. behavior research methods, instruments and computers, 23(2):229–236, 1991. ted dunning. accurate methods for the statistics of surprise and coincidence. computational linguistics, 19:61–74, march 1993. issn 0891-2017. url http://portal.acm.org/citation. cfm?id=972450.972454. daniel c. edelson, douglas n. gordin, and roy d. pea. addressing the challenges of inquiry-based learning through technology and curriculum design. journal of the learning sciences, 8(3 & 4): 391–450, 1999. david ferrucci, eric brown, jennifer chu-carroll, james fan, david gondek, aditya a. kalyanpur, adam lally, j. william murdock, eric nyberg, john prager, nico schlaefer, and chris welty. building watson: an overview of the deepqa project. ai magazine, 31(3), 2010. url http: //www.aaai.org.proxy.lib.sfu.ca/ojs/index.php/aimagazine/article/view/2303. leon festinger. a theory of cognitive dissonance. tavistock publications, 1962. kathleen m. fisher, james h. wandersee, and david e. moody. mapping biology knowledge. kluwer academic pub, 2000. august flammer. towards a theory of question asking. psychological research, 43(4):407–420, 1981. paul w. foos. student study techniques and the generation effect. journal of educational psychology, 86(4):567–76, 1994. albert gatt and ehud reiter. simplenlg: a realisation engine for practical applications. in enlg ’09: proceedings of the 12th european workshop on natural language generation, pages 90–93, morristown, nj, usa, 2009. association for computational linguistics. 94 question generation from concept maps james r. gavelek and taffy e. raphael. metacognition, instruction, and the role of questioning activities. in d.l. forrest-pressley, g.e. mackinnon, and g.t. waller, editors, metacognition, cognition, and human performance: instructional practices, volume 2, pages 103–136. academic press, orlando, fl, 1985. shawn m. glynn. making science concepts meaningful to students: teaching with analogies. in s. mikelskis-seifert, u. ringelband, and m. bruckmann, editors, four decades of research in science education: from curriculum development to quality improvement, pages 113–125. waxmann, mnster, germany, 2008. sallie e. gordon, kimberly a. schmierer, and richard t. gill. conceptual graph analysis: knowledge acquisition for instructional system design. human factors: the journal of the human factors and ergonomics society, 35(3):459–481, 1993. arthur c. graesser and leslie c. clark. structures and procedures of implicit knowledge. ablex, norwood, nj, 1985. arthur c. graesser and stanley p. franklin. quest: a cognitive model of question answering. discourse processes, 13:279–303, 1990. arthur c. graesser and cathy l. mcmahen. anomalous information triggers questions when adults solve problems and comprehend stories. journal of educational psychology, 85:136–151, 1993. arthur c. graesser and brent a. olde. how does one know whether a person understands a device? the quality of the questions the person asks when the device breaks down. journal of educational psychology, 95(3):524–536, 2003. arthur c. graesser and natalie k. person. question asking during tutoring. american educational research journal, 31(1):104–137, 1994. issn 0002-8312. arthur c. graesser, sallie e. gordon, and lawrence e. brainerd. quest: a model of question answering. computers and mathematics with applications, 23:733–745, 1992. arthur c. graesser, natalie k. person, and joseph p. magliano. collaborative dialogue patterns in naturalistic one-to-one tutoring. applied cognitive psychology, 9:1–28, 1995. arthur c. graesser, william baggett, and kent williams. question-driven explanatory reasoning. applied cognitive psychology, 10:s17–s32, 1996. michael heilman and noah a. smith. ranking automatically generated questions as a shared task. in scotty d. craig and darina dicheva, editors, the 2nd workshop on question generation, volume 1 of aied 2009 workshops proceedings, pages 30–37, brighton, uk, july 2009. international artificial intelligence in education society. michael heilman and noah a. smith. extracting simplified statements for factual question generation. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 11–20, pittsburgh, june 2010. questiongeneration.org. url http://oro.open.ac.uk/22343/. denis j. hilton. conversational processes and causal explanation. psychological bulletin, 107(1): 65–81, 1990. gregory hume, joel michael, allen rovick, and martha evens. hinting as a tactic in one-on-one tutoring. the journal of the learning sciences, 5:23–47, 1996. 95 olney, graesser and person richard johansson and pierre nugues. dependency-based syntactic-semantic analysis with propbank and nombank. in conll ’08: proceedings of the twelfth conference on computational natural language learning, pages 183–187, morristown, nj, usa, 2008. association for computational linguistics. isbn 978-1-905593-48-4. saidalavi kalady, ajeesh elikkottil, and rajarshi das. natural language question generation using syntax and keywords. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 1–10, pittsburgh, june 2010. questiongeneration.org. url http://oro.open.ac.uk/22343/. alison king. effects of self-questioning training on college students’ comprehension of lectures. contemporary educational psychology, 14(4):1–16, 1989. alison king. comparison of self-questioning, summarizing, and notetaking-review as strategies for learning from lectures. american educational research journal, 29(2):303–323, 1992. alison king. guiding knowledge construction in the classroom: effects of teaching children how to question and how to explain. american educational research journal, 31(2):338–368, 1994. mikel larrañaga, urko rueda, jon a. elorriaga, and ana arruarte lasa. acquisition of the domain structure from document indexes using heuristic reasoning. in intelligent tutoring systems, pages 175–186, 2004. prashanth mannem, rashmi prasad, and aravind joshi. question generation from paragraphs at upenn: qgstec system description. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 84–91, pittsburgh, june 2010. questiongeneration.org. url http://oro.open.ac.uk/22343/. danielle s. mcnamara and joseph p. magliano. toward a comprehensive model of comprehension. in brian h. ross, editor, the psychology of learning and motivation, volume 51, chapter 9, pages 297–384. academic press, new york, 2009. isbn 978-0-12-374489-0. olena medelyan, eibe frank, and ian h. witten. human-competitive tagging using automatic keyphrase extraction. in proceedings of the 2009 conference on empirical methods in natural language processing, pages 1318–1327, singapore, august 2009. association for computational linguistics. url http://www.aclweb.org/anthology/d/d09/d09-1137. adam meyers, ruth reeves, catherine macleod, rachel szekely, veronika zielinska, brian young, and ralph grishman. the nombank project: an interim report. in a. meyers, editor, hltnaacl 2004 workshop: frontiers in corpus annotation, pages 24–31, boston, massachusetts, usa, may 2 may 7 2004. association for computational linguistics. kenneth r. miller and joseph s. levine. prentice hall biology. pearson education, new jersey, 2002. bruce mills, martha evens, and reva freedman. implementing directed lines of reasoning in an intelligent tutoring system using the atlas planning environment. in proceedings of the international conference on information technology, coding and computing, volume 1, pages 729–733, las vegas, nevada, april 2004. joel j. mintzes, james h. wandersee, and joseph d. novak. assessing science understanding: a human constructivist view. academic press, 2005. dan i. moldovan, christine clark, and moldovan bowden. lymba’s poweranswer 4 in trec 2007. in trec, 2007. 96 question generation from concept maps jack mostow and wei chen. generating instruction automatically for the reading strategy of selfquestioning. in vania dimitrova, riichiro mizoguchi, benedict du boulay, and art graesser, editors, proceeding of the 2009 conference on artificial intelligence in education: building learning systems that care: from knowledge representation to affective modelling, pages 465–472, amsterdam, the netherlands, 2009. ios press. tom murray. authoring knowledge-based tutors: tools for content, instructional strategy, student model, and interface design. journal of the learning sciences, 7(1):5, 1998. issn 1050-8406. roberto navigli and paola velardi. from glossaries to ontologies: extracting semantic structure from textual definitions. in proceeding of the 2008 conference on ontology learning and population: bridging the gap between text and knowledge, pages 71–87, amsterdam, the netherlands, the netherlands, 2008. ios press. isbn 978-1-58603-818-2. victor u. odafe. students generating test items: a teaching and assessment strategy. mathematics teacher, 91(3):198–203, 1998. andrew olney, whitney cade, and claire williams. generating concept map exercises from textbooks. in proceedings of the sixth workshop on innovative use of nlp for building educational applications, pages 111–119, portland, oregon, june 2011. association for computational linguistics. url http://www.aclweb.org/anthology/w11-1414. andrew m. olney. extraction of concept maps from textbooks for domain modeling. in vincent aleven, judy kay, and jack mostow, editors, intelligent tutoring systems, volume 6095 of lecture notes in computer science, pages 390–392. springer berlin / heidelberg, 2010. url http: //dx.doi.org/10.1007/978-3-642-13437-1\ 80. 10.1007/978-3-642-13437-1 80. andrew m. olney, arthur c. graesser, and natalie k. person. tutorial dialog in natural language. in r. nkambou, j. bourdeau, and r. mizoguchi, editors, advances in intelligent tutoring systems, volume 308 of studies in computational intelligence, pages 181–206. springer-verlag, berlin, 2010. gary m. olson, susan a. duffy, and robert l. mack. question-asking as a component of text comprehension. in arthur c. graesser and j. b. black, editors, the psychology of questions, pages 219–226. lawrence earlbaum, hillsdale, n.j., 1985. jose otero and arthur c. graesser. preg: elements of a model of question asking. cognition and instruction, 19(2):143–175, 2001. santanu pal, tapabrata mondal, partha pakray, dipankar das, and sivaji bandyopadhyay. qgstec system description juqgg: a rule based approach. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 76–79, pittsburgh, june 2010. questiongeneration.org. url http://oro.open.ac.uk/22343/. annemarie s. palinscar and ann l. brown. reciprocal teaching of comprehension-fostering and comprehension-monitoring activities. cognition and instruction, 1:117–175, 1984. martha palmer, daniel gildea, and paul kingsbury. the proposition bank: an annotated corpus of semantic roles. comput. linguist., 31(1):71–106, 2005. issn 0891-2017. doi: http://dx.doi. org/10.1162/0891201053630264. natalie k. person, arthur c. graesser, joseph p. magliano, and roger j. kreuz. inferring what the student knows in one-to-one tutoring: the role of student questions and answers. learning and individual differences, 6(2):205–229, 1994. jean piaget. the origins of intelligence. international university press, new york, 1952. 97 olney, graesser and person michael pressley and donna forrest-pressley. questions and children’s cognitive processing. in a.c. graesser and b. john, editors, the psychology of questions, pages 277–296. lawrence erlbaum associates, hillsdale, nj, 1985. barak rosenshine, carla meister, and saul chapman. teaching students to generate questions: a review of the intervention studies. review of educational research, 66(2):181–221, 1996. vasile rus and arthur c. graesser. the question generation shared task and evaluation challenge. technical report, university of memphis, 2009. isbn:978-0-615-27428-7. vasile rus, brendan wyse, paul piwek, mihai lintean, svetlana stoyanchev, and cristian moldovan. overview of the first question generation shared task evaluation challenge. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 45–57, pittsburgh, june 2010a. questiongeneration.org. url http: //oro.open.ac.uk/22343/. vasile rus, brendan wyse, paul piwek, mihai lintean, stoyanchev svetlana, and cristian moldovan. the first question generation shared task evaluation challenge. in proceedings of the 6th international natural language generation conference (inlg 2010), dublin, ireland, 2010b. ivan a. sag and dan flickinger. generating questions with deep reversible grammars. in workshop on the question generation shared task and evaluation challenge, arlington, va, september 2008. marlene scardamalia and carl bereiter. fostering the development of self-regulation in children’s knowledge processing. in susan f. chipman, judity w. segal, and robert glaser, editors, thinking and learning skills, volume 2, pages 563–577. erlbaum, hillsdale, nj, 1985. roger c. schank. dynamic memory revisited. cambridge university press, 1999. isbn 0521633982, 9780521633987. harry singer and dan donlan. active comprehension: problem-solving schema with question generation for comprehension of complex short stories. reading research quarterly, 17(2):166–186, 1982. john f. sowa. semantic networks. in stuart c shapiro, editor, encyclopedia of artificial intelligence. wiley, 1992. robert d. tennyson and ok-choon park. the teaching of concepts: a review of instructional design research literature. review of educational research, 50(1):55 –70, 1980. doi: 10.3102/ 00346543050001055. url http://rer.sagepub.com/content/50/1/55.abstract. alejandro valerio and david b. leake. associating documents to concept maps in context. in a. j. canas, p. reiska, m. ahlberg, and j. d. novak, editors, proceedings of the third international conference on concept mapping, 2008. lucy vanderwende. the importance of being important: question generation. in vasile rus and art graesser, editors, proceedings of the workshop on the question generation shared task and evaluation challenge, september 25-26 2008. andrea varga and le an ha. wlv: a question generation system for the qgstec 2010 task b. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 80–83, pittsburgh, june 2010. questiongeneration.org. url http: //oro.open.ac.uk/22343/. 98 question generation from concept maps ellen m. voorhees and hoa trang dang. overview of the trec 2005 question answering track. in nist special publication 500-266: the fourteenth text retrieval conference proceedings (trec 2005), 2005. url citeseer.ist.psu.edu/article/voorhees02overview.html. terry winograd. understanding natural language. academic press, new york, 1972. bernice y. l. wong. self-questioning instructional research: a review. review of educational research, 55(2):227–268, 1985. brendan wyse and paul piwek. generating questions from openlearn study units. in scotty d. craig and darina dicheva, editors, the 2nd workshop on question generation, volume 1 of aied 2009 workshops proceedings, pages 66–73, brighton, uk, july 2009. international artificial intelligence in education society. xuchen yao. question generation with minimal recursion semantics. master’s thesis, saarland university & university of groningen, 2010. url http://cs.jhu.edu/∼xuchen/paper/ yao2010master.pdf. xuchen yao and yi zhang. question generation with minimal recursion semantics. in kristy elizabeth boyer and paul piwek, editors, proceedings of qg2010: the third workshop on question generation, pages 68–75, pittsburgh, june 2010. questiongeneration.org. url http: //oro.open.ac.uk/22343/. barry j. zimmerman. a social cognitive view of self-regulated academic learning. journal of educational psychology, 81(3):329–339, 1989. amal zouaq and roger nkambou. evaluating the generation of domain ontologies in the knowledge puzzle project. ieee trans. on knowl. and data eng., 21(11):1559–1572, 2009. issn 1041-4347. doi: http://dx.doi.org/10.1109/tkde.2009.25. 99 dialogue and discourse 3(2) (2012) 11–42 doi: 0.5087/dad.2012.202 semantics-based question generation and implementation xuchen yao xuchen@cs.jhu.edu department of computer science, johns hopkins university, 3400 n. charles street, baltimore, usa gosse bouma g.bouma@rug.nl information science, university of groningen, po box 716, 9700 as groningen, the netherlands yi zhang yzhang@coli.uni-sb.de lt-lab, german research center for artificial intelligence (dfki gmbh) department of computational linguistics, saarland university, 66123 saarbrücken, germany editors: paul piwek and kristy elizabeth boyer abstract this paper presents a question generation system based on the approach of semantic rewriting. state-of-the-art deep linguistic parsing and generation tools are employed to map natural language sentences into their meaning representations in the form of minimal recursion semantics (mrs) and vice versa. by carefully operating on the semantic structures, we obtain a principled way of generating questions which avoids ad-hoc manipulation of syntactic structures. based on the (partial) understanding of the sentence meaning, the system generates questions that are semantically grounded and purposeful. as the generator uses a deep linguistic grammar, the grammaticality of the generation results is licensed by the grammar. with a specialized ranking model, the linguistic realizations from the general purpose generation model are further refined for the question generation task. the evaluation results from qgstec2010 show promising results for the proposed approach. 1. introduction question generation (qg) is the task of generating reasonable questions from an input, which can be structured (e.g. a database) or unstructured (e.g. a text). in this paper, we narrow the task of qg down to taking a natural language text as input (thus textual qg), as it is a more interesting challenge that involves a joint effort between natural language understanding (nlu) and natural language generation (nlg). simply put, if natural language understanding maps text to symbols and natural language generation maps symbols to text, then question generation maps text to text, through an inner mapping from symbols for declarative sentences to symbols for interrogative sentences, as shown in figure 1. here we use symbols as an organized data form that can represent the semantics of natural languages and that can be processed by a machinery, artificial or otherwise. the task of question generation contains multiple subareas. usually, the approach taken for qg depends on the purpose of the qg application. generally speaking, a qg system can be helpful in the following areas: • intelligent tutoring systems. qg can ask questions based on learning materials in order to check learners’ accomplishment or help them focus on the keystones in study. qg can also help tutors to prepare questions intended for learners or prepare for potential questions from learners. c©2012 xuchen yao, gosse bouma and yi zhang submitted 1/11; accepted 1/12; published online 3/12 yao, bouma and zhang • closed-domain question answering (qa) systems. some closed-domain qa systems use predefined (sometimes hand-written) question-answer pairs to provide qa services. by employing a qg approach such systems could be ported to other domains with little or no effort. • natural language summarization/generation systems. qg can help to generate, for instance, frequently asked questions from the provided information source in order to provide a list of faq candidates. in terms of target complexity, qg can be divided into deep qg and shallow qg (graesser et al., 2009). deep qg generates deep questions that involve more logical thinking (such as why, why not, what-if, what-if-not and how questions) whereas shallow qg generates shallow questions that focus more on facts (such as who, what, when, where, which, how many/much and yes/no questions). given the current state of qg, most of the applications listed above have limited themselves to shallow qg. we describe a semantics-based system, mrsqg1, that generates questions from a given text, specifically, a text that contains only a single declarative sentence. this restriction was motivated by task b of the question generation shared task and evaluation challenge (qgstec2010; rus et al., 2010; rus et al., this volume), which provides participants with a single sentence and asks them to generate questions according to a required target question type. by concentrating on single sentences, systems can focus on generating well-formed and purposeful questions, without having to deal (initially) with text analysis at the discourse level. the basic intuition of mrsqg can be explained by the following example. think of generating a few simple questions from the sentence “john plays football.”: example 1 john plays football. (a) who plays football? (b) what does john play? natural language text natural language questions transformationsymbolic representation for text symbolic representation for questions nlu nlg question generation figure 1: the relation between question generation and its two components: natural language understanding (nlu) and natural language generation (nlg). 1. http://code.google.com/p/mrsqg/ 12 http://code.google.com/p/mrsqg/ semantics-based question generation and implementation when people perform question generation, a transformation from declarative sentences to interrogatives happens. this transformation can be described at different levels of abstraction. an intuitive one is provided by predicate logic: example 2 play(john, football) ⇐ john plays football. (a) play(who, football) ⇐ who plays football? (b) play(john, what) ⇐ what does john play? if the above abstraction can be described and obtained in a formal language and transformation can be done according to some well-formed mechanism, then the task of question generation has a solution. we propose a semantics-based method of transforming the minimal recursion semantics (mrs, copestake et al., 2005) representation of declarative sentences to that of interrogative sentences. the mrs analysis is obtained from pet (callmeier, 2000) running with the english resource grammar (erg, flickinger, 2000) as the core linguistic component, while the generation function is delivered by the linguistic knowledge builder (lkb, copestake, 2002). the advantage of this approach is that the mapping from declarative to interrogative sentence is done on the semantic representations. in this way, we are able to use an independently developed parser and generator for the analysis and generation stage. the generator will usually propose several different surface realizations of a given input due to its extensive grammatical coverage. this means that the system is able to produce more diverse questions, but also that ranking of various surface forms becomes an issue. additional advantages of the semantic approach are that it is to a large extent language independent, and that it provides a principled level of representation for incorporating lexical semantic resources. for instance, given that “sport” is a hypernym of “football”, we can have the following transformation: example 3 play(john, football) & hypernym(sport, football) ⇒ play(john, which sport) the hypernym relation between “sport” and “football” can either be obtained from ontologies, such as a list of different sports, or semantic networks, such as wordnet (fellbaum, 1998). this paper is organized as follows. section 2 reviews related work in question generation and explains why a semantics-based approach is proposed. it also identifies major challenges in the task of question generation. after a brief background introduction in section 3, section 4 provides corresponding solutions and describes our implemented system mrsqg. specifically, section 4.1 and 4.2 present the key idea in this paper. evaluation results are presented in section 5 and several aspects of the proposed method are discussed in section 6. finally, section 7 concludes and addresses future work. 2. related work generally speaking, the following issues have been addressed in question generation: 1. question transformation. as shown in figure 1, this requires a theoretically-sound and practically-feasible algorithm to build a mapping from symbolic representation of declarative sentences to interrogative sentences. 2. sentence simplification. complex and long sentences widely occur in written languages. but questions are rarely very long. on the one hand, complex input sentences are hard to match against pre-defined patterns. on the other hand, most current question generation approaches transform the input sentence into questions, thus it is better to keep the input short and succinct in order to avoid lengthy and awkward questions. thus sentence simplification is usually performed as a pre-processing step. 13 yao, bouma and zhang natural language text natural language questions symbolic representation for text symbolic representation for questions nlu nlg question generation simplification ranking transformation figure 2: three major problems (sentence simplification, question transformation and question ranking) in the process of question generation, as shown in the framed boxes. 3. question ranking. in the case of over generation, a ranking algorithm to grade the grammaticality and naturalness of questions must be developed. furthermore, a good ranking algorithm should also select relevant and appropriate questions according to the content requirements. figure 2 shows these three problems in an overview of a complete question generation framework. the following section reviews research on these three issues in more detail. 2.1 question transformation there are generally three approaches to question transformation and generation: template-based, syntax-based and semantics-based. mostow and chen (2009) reported on a template-based system under a self-questioning strategy to help children generate questions from narrative fiction. three question templates are used to produce what/why/how questions. of 769 questions evaluated, 71.3% were rated acceptable. this work was further expanded by chen et al. (2009) with 4 more templates to generate what-wouldhappen-if, when-would-x-happen, what-would-happen-when and why-x questions from informational text questions. template-based approaches are mostly suitable for applications with a special purpose, which sometimes come with a closed-domain. the trade-off between coverage and cost is hard to balance because human labor is required to produce high-quality templates. thus, it is generally considered unsuitable for open-domain general purpose applications. syntax-based approaches include wyse and piwek (2009) and heilman and smith (2009), both of which use a very similar method for manipulating syntactic trees. the core idea is to transform a syntactic tree of a declarative sentence into that of an interrogative. specific matching and transformation rules are defined by experienced linguists and operate on tree structures. all operations 14 semantics-based question generation and implementation are straightforward from a syntactic point of view. heilman and smith (2009) reported 43.3% acceptability for the top 10 ranked questions and produced an average of 6.8 acceptable questions per 250 words on wikipedia texts. semantics-based approaches are less explored in previous research. schwartz et al. (2004) introduce a content question generator that uses the logical form to represent the semantic relationships of the arguments within a sentence and generate wh-questions. however, the paper only introduces the result of this generator whereas the inner mechanism is not presented. sag and flickinger (2008) discuss the possibility and feasibility of using the english resource grammar for generation under a head-driven phrase structure grammar (hpsg, pollard and sag, 1994) framework. minimal recursion semantics is the input to erg and linguistic realizations come from the linguistic knowledge builder system. successful applications are listed in support of their arguments for generating through erg and lkb. in this paper, we present the implementation of a question generation system following their proposed approach. 2.2 sentence simplification sentence simplification reduces the average length and syntactic complexity of a sentence, which is usually marked by a reduction in reading time and an increase in comprehension. chandrasekar et al. (1996) reported two rule-based methods to perform text simplification. they take simplification as a two stage process. the first stage gives a structural representation of a sentence and the second stage transforms this representation into a simpler one, using handcrafted rules. the two methods differ in that the structural representation is different and thus rules also change accordingly (one uses chunk parsing and the other supertagging (bangalore and joshi, 1999)). in natural language processing, the methods are also used as a pre-processing technique to alleviate the overload of the parser, the information retrieval engine, etc. thus the analysis of sentences is no deeper than a syntactic tree. in the context of question generation, the analyzing power is not confined to this level. for instance, dorr et al. (2003) and heilman and smith (2010b) use a syntactic parser to obtain a tree analysis of a whole sentence and define heuristics over tree structures. cohn and lapata (2009) compress sentences by rewriting tree structures over synchronous tree substitution grammar (stsg, eisner, 2003). 2.3 question ranking question ranking falls into the topic of realization ranking, which can be taken as the final step of natural language generation. in the context of generating questions from semantics, there can be multiple projections from a single semantic representation to surface realizations. velldal and oepen (2006) compared different statistical models to discriminate between competing surface realizations. the performance of a language model, a maximum entropy (maxent) model and a support vector machine (svm) ranker is investigated. the language model is a trigram model trained on the british national corpus (bnc) with 100 million words. sentences with higher probability are ranked better. the maxent model and svm ranker use features defined over derivation trees as well as lexical trigram models. their result shows that maxent is slightly better than svm, while both models significantly outperform the language model. this result is mostly based on declarative sentences. our proposed question ranking module combines the maxent model from velldal and oepen (2006) with a language model trained on questions. details are in section 4.3. heilman and smith (2010a) worked directly on ranking questions. they employed an overgenerateand-rank approach. the overgenerated questions were ranked by a logistic regression model trained on a human-annotated corpus. the features used by the ranker covered various aspects of the questions, including length, n -gram language model, wh-words, grammatical features, etc. while 27.3% of all test set questions were acceptable, 52.3% of the top 20% ranked questions were acceptable. 15 yao, bouma and zhang 2.4 perception of related work among the three subtasks of question generation, most of the work related to sentence simplification and question transformation is heavily syntax-based. the internal symbolic representation of languages is encoded with syntax and transformation rules are defined over syntax. depending on the depth of processing, these syntactic structures can be either flat (with only pos or chunk information) or structured (with parse trees). however, these approaches also introduce syntaxspecific limitations, including language dependencies, ad-hoc transformations and complex syntactic formalisms. semantics-based approaches are to some extent amenable to these problems. they have more expressive power by utilizing lexical semantics resources and employing more flexible generators. specifically, semantics-based methods involve both parsing to semantics and generating from semantics, as well as transforming via semantics. question generation happens to be one application that requires all of these operations. 3. background: theory, grammar and tools the proposed system consists of multiple syntactic and semantic processing components based on the delph-in tool-chain (e.g. erg/lkb/pet), while the theoretical support comes from minimal recursion semantics. this section introduces all the components involved, but concentrates on the parts that are actually used in later sections. figure 3 gives an overview of the functionalities of different components. minimal recursion semantics is a theory of semantic representation of natural language sentences with a focus on the underspecification of scope ambiguities. pet is used as a parser to interpret a natural language sentence into mrs with the guidance of a hand-written grammar. in turn, the generation component of lkb takes in an mrs structure and produces various realizations as natural language sentences. both directions of processing are guided by the english resource grammar, which includes a large scale hand-crafted lexicon and sophisticated grammar rules that cover most essential syntactic constructions of english, and connects natural language sentences with their meaning representations in mrs. the tools are developed in the context of the delph-in2 collaboration, and are freely available as an open source repository. 3.1 minimal recursion semantics and english resource grammar minimal recursion semantics is a meta-level language for describing semantic structures in some underlying object language (copestake et al., 2005). while the representation can be combined freely with various linguistic frameworks, mrs is particularly convenient to be encoded in typed feature structures. as it turns out to be also practically suitable for encoding semantic compositions in grammar engineering, it is no surprise to see mrs being adopted as the semantic representation for most of the delph-in hpsg grammars. an mrs structure is composed of the following components: ltop, index, rels, hcons, as shown in figure 4. ltop is the topmost label of this mrs. the index usually contains an event variable “e”, which is co-indexed with the arg0 property of the main predicate ( like v rel in this case), i.e. its bound variable. rels is a bag of elementary predications, or eps, in which a single ep means a single relation with its arguments, such as like v rel(e2, x5, x9). any ep with rstr and body features corresponds to a generalized quantifier. it takes a form of rel(arg0, rstr, body) where arg0 refers to the bound variable and rstr puts a scopal restriction on some other relation by the “qeq” relation specified in hcons (handle constraints). here in figure 4 the relation proper q rel(x5, h4, 2. deep linguistic processing with hpsg: http://www.delph-in.net/ 16 http://www.delph-in.net/ semantics-based question generation and implementation index: e2 rels: < [ proper_q_rel<0:4> lbl: h3 arg0: x6 rstr: h5 body: h4 ] [ _like_v_1_rel<5:10> lbl: h8 arg0: e2 [ e sf: prop tense: pres ] arg1: x6 arg2: x9 [ proper_q_rel<11:17> lbl: h10 arg0: x9 rstr: h12 body: h11 ] > hcons: < h5 qeq h7 h12 qeq h13 > [ named_rel<0:4> lbl: h7 arg0: x6 (pers: 3 num: sg) carg: "john" ] [ named_rel<11:17> lbl: h13 arg0: x9 (pers: 3 num: sg) carg: "mary" ] john likes mary. like(john, mary) parsing with pet generation with lkb john likes mary. minimal recursion semantics english resource grammar figure 3: different components (pet/lkb/mrs/erg) of the hpsg-centralized delph-in community. the predicate logic form like(john, mary) can be further represented by minimal recursion semantics, a more powerful and complex semantic formalism, which is adopted in the english resource grammar.  mrs ltop h1 index e2 rels 〈  proper q lbl h3 arg0 x5 rstr h4 body h6 , named lbl h7 arg0 x5 carg ”mary” ,  like v lbl h8 arg0 e2 arg1 x5 arg2 x9 ,  udef q lbl h10 arg0 x9 rstr h11 body h12 ,  red a lbl h13 arg0 e14 arg1 x9 ,[ rose n lbl h13 arg0 x9 ] 〉 hcons 〈 [ qeq harg h4 larg h7 ] , [ qeq harg h11 larg h13 ] 〉  figure 4: mrs representation of “mary likes red roses” with some linguistic attributes omitted. types are in bold. features (or attributes) are in capital. values are in �. values in 〈〉 are list values. note that some features are token-identical. the suffix rel in relation names is dropped to save space. 17 yao, bouma and zhang like v named(”mary”) proper q rstr/h arg1/neq rose n udef q rstr/h red a arg1/eq arg2/neq figure 5: the dependency mrs structure of “mary likes red roses”. relations between eps are indicated by the labels of arrowed arcs. refer back to figure 4 for the original mrs. h6) with a “h4 qeq h7 ” constraint means that proper q rel outscopes the named rel, which has a label h7. mrs to some extent preserves the ambiguities that are inherent to natural language, and represents them in a compact way using underspecification. for instance, for the sentence “every man loves a woman”, there are at least two major readings: every man loves the same woman or every man loves a different woman. this sentence has only one mrs structure but two scopal readings: it can either be that the ep for “a” outscopes the ep for “every” (every man loves the same woman) or the other way around. choosing mrs as our semantic formalism has gained a bonus for the present application: semantic transformation operates on mrss that are not resolved to unambiguous logical representations, and generation also operates on such structures. thus, the application can be realized without being forced to do complete ambiguity resolution for the input sentence. dependency mrs (dmrs, copestake, 2008) serves as an interface between flat and non-flat structures. recall that in figure 4, the value of rels is a bag of eps. this flat structure is verbose and does not easily show how the eps in an mrs are connected with each other. dmrs is designed to be more direct in representing the direct relations between eps, but still preserves all of the original semantic information. a dmrs is a connected acyclic graph. figure 5 shows the dmrs of “mary likes red roses.”, originally from figure 4. the directional arcs represent regular semantic dependencies (i.e. the semantic head points to its children) with the labels of arcs describing the relations in detail. a label (e.g. arg1/neq) has two parts: the part before the slash is inherited from the feature name of the original mrs; the part after the slash indicates the type of a scopal relation. possible values are: h (qeq relationship), eq (label equality), neq (label non-equality), heq (one ep’s argument is the other ep’s label) and null (underspecified label relationships). the semantic composition rules are encoded in the english resource grammar, together with the thorough modeling of the syntax. the grammar is based on the hpsg framework, and through over 15 years of continuous development, has achieved broad coverage while maintaining high linguistic accuracy. it consists of a large set of lexical entries under a detailed hierarchy of lexical types, with a modest set of lexical rules for production. the erg uses a davidsonian representation in which all verbs introduce events. this explains why the like v rel relation in figures 3 and 4 has e2 as its arg0. 3.2 linguistic knowledge builder we use the generation component of lkb for sentence realization. the linguistic knowledge builder is a grammar engineering platform for developing linguistic grammars in typed-feature structures 18 semantics-based question generation and implementation and unification-based formalisms. it can examine the competence and performance of a grammar by means of parsing and generation. the generation component takes as input a valid mrs structure, and tries to find all possible realizations of the same meaning representation in natural language sentences according to the grammar. the generator uses a chart-based algorithm as described in kay (1996), carroll et al. (1999), carroll and oepen (2005). the latter two also tackle efficiency problems. in case multiple realizations are available, the generator ranks them with a statistical disambiguation model trained with large scale treebanks. 3.3 parsing with pet although erg has a large lexicon, there are always unknown words in real text. thus, a robust processing strategy is needed as fallback. pet is a platform for doing experiments with efficient processing of unification-based grammars (callmeier, 2000). it employs a two-stage parsing model. firstly, hpsg rules are used in parsing. since these rules are sometimes too restrictive and only produce a partial chart, a second stage with a permissive pcfg backbone predicts a full-spanning pseudo-derivation, and the fragmented mrs structures are extracted from these partial analyses. (zhang et al., 2007). the fragmented mrs does not have full feature structures for all constituents because of rule conflict, otherwise it could have been parsed in the first stage. 4. proposed method recall in section 2 we addressed three major problems in question generation: question transformation, sentence simplification and question ranking. a semantics-based system, mrsqg, is developed to tackle these problems by corresponding solutions: mrs transformation for simple sentences, mrs decomposition for complex sentences and automatic generation with rankings. also, practical issues such as robust generation with fallbacks are addressed. figure 6 shows the processing pipelines of mrsqg. the following is a brief description of each step. 1. term extraction. the stanford named entity recognizer (finkel et al., 2005), a regex ne tagger, an ontology ne tagger and wordnet (fellbaum, 1998) are used to extract terms. 2. fsc construction. the feature structure chart (fsc) format3 is an xml-based format that introduces tokenization and external annotation to the erg grammar and pet parser. using fsc makes the terms annotated by named entity recognizers known to the parser. 3. parsing with pet. the chart mapping (adolphs et al., 2008) functionality of pet accepts fsc input and outputs mrs structures. 4. mrs decomposition. complex sentences first need to be broken into shorter ones with valid and meaningful semantic representation. also, it must be possible to generate a natural language sentence from this shorter semantic representation. this is the key point in our semantics-based approach. details in section 4.2. 5. mrs transformation. given a valid mrs structure of a sentence, we replace eps for terms with eps for (wh) question words. section 4.1 gives a detailed description with examples. 6. generating with lkb. given an mrs for a question, the generator of lkb produces possible surface realizations. 7. output selection. given a well-formed mrs structure, lkb might give multiple outputs. depending on how erg does generation, some output strings might not sound fluent or even 3. http://wiki.delph-in.net/moin/petinputfsc 19 http://wiki.delph-in.net/moin/petinputfsc yao, bouma and zhang mrs xml plain text term extraction fsc construction parsing with pet 2 3 1 generation with lkb output selection 5 6 7 mrs transformationmrs decomposition 8 output to console/xml fsc xml apposition decomposer coordination decomposer subclause decomposer subordinate decomposer why decomposer mrs xml 4 figure 6: pipelines of an mrs transformation based question generation system. grammatical. thus there must be ranking algorithms to select the best one. details are in section 4.3. 8. output to console/xml. mrsqg can output its results either to a console for user interaction, or (to a file) in xml format for formal evaluation. by way of illustration, we now present some actual questions generated by mrsqg: example 4 jackson was born on august 29, 1958 in gary, indiana. generated who questions: (a) who was born in gary , indiana on august 29 , 1958? (b) who was born on august 29 , 1958 in gary , indiana? generated where questions: (c) where was jackson born on august 29 , 1958? generated when questions: (d) when was jackson born in gary , indiana? generated yes/no questions: (e) jackson was born on august 29 , 1958 in gary , indiana? (f) jackson was born in gary , indiana on august 29 , 1958? (g) was jackson born on august 29 , 1958 in gary , indiana? (h) was jackson born in gary , indiana on august 29 , 1958? (i) in gary , indiana was jackson born on august 29 , 1958? (j) in gary , indiana, was jackson born on august 29 , 1958? (k) on august 29 , 1958 was jackson born in gary , indiana? (l) on august 29 , 1958, was jackson born in gary , indiana? the examples show that the system generates various question types. different word orders are allowed by the general erg used in generation. not all of these sound equally natural. section 4.3 addresses this issue. the following sections present more details on steps that need further elaboration. 20 semantics-based question generation and implementation like v 1 named(”john”) proper q rstr/h arg1/neq named(”mary”) proper q rstr/h arg2/neq like v 1 person which q rstr/h arg1/neq named(”mary”) proper q rstr/h arg2/neq (a) “john likes mary” → “who likes mary?” sing v 1 named(”mary”) proper q rstr/h arg1/neq on p named(”broadway”) proper q rstr/h arg2/neq arg1/eq sing v 1 named(”mary”) proper q rstr/h arg1/neq loc nonsp place n which q rstr/h arg2/neq arg1/eq (b) “mary sings on broadway.” → “where does mary sing?” sing v 1 named(”mary”) proper q rstr/h arg1/neq at p temp numbered hour(”10”) def implicit q rstr/h arg2/neq arg1/eq sing v 1 named(”mary”) proper q rstr/h arg1/neq loc nonsp time which q rstr/h arg2/neq arg1/eq (c) “mary sings at 10.” → “when does mary sing?” figure 7: (continued in the next page) 21 yao, bouma and zhang fight v 1 named(”john”) proper q rstr/h arg1/neq for p named(”mary”) proper q rstr/h arg2/neq arg1/eq fight v 1 named(”john”) proper q rstr/h arg1/neq for p reason q which q rstr/h arg2/neq arg1/eq (d) “john fights for mary.” → “why does john fight?” figure 7: mrs transformation from declarative sentences to wh questions in a form of dependency graph. 4.1 mrs transformation for simple sentences the transformation from declarative sentences into interrogatives follows a mapping between elementary predications (eps) of relations. figure 7 shows this mapping. many terms in preprocessing are tagged as proper nouns (nnp or nnps). thus the eps of a term turns out to consist of two eps: proper q rel (a quantification relation) and named rel (a naming relation), with proper q rel outscoping and governing named_rel. the eps of wh-question words have a similar structure. for instance, the eps of “who” consist of two relations: which q rel and person rel, with which q rel outscoping and governing person rel. changing the eps of terms to eps of wh-question words naturally results in an mrs for wh-questions. similarly, in where/when/why questions, the eps for the wh question word are which q rel and place rel/time rel/reason rel. special attention must be paid to the preposition word that usually comes before location/time. in a dependency tree, a preposition word governs the head of the phrase with a post-slash eq relation (as shown in figure 7(bcd)). the ep of the preposition must be changed to a loc nonsp rel ep (an implicit locative which does not specify a preposition) which takes the wh word relation as an argument in both cases of when/where. this ep avoids generating non-grammatical phrases such as “in where” and “on when”. generally, there is one rule per question type. the choice for a transformation rule is triggered by the named entity output. most questions regarding person/location/time can be asked in two ways: either by who/where/when questions or by which person/location/time questions. we encoded both types of these rules in the system. for instance, “on august 29, 1958” can be transformed to either “when” or “on which date”. both questions are generated at this point, and we leave it to the question ranker to select the best one based on properties of the whole question. as we also use an overgenerate-and-rank approach, we always try to generate as many questions as possible. changing the sf (sentence force) attribute of the main event variable from prop to ques generates yes/no questions. this is the simplest case in question generation. 22 semantics-based question generation and implementation english sentence structure complex dependent clause + independent clause subordinate clause causal | non-causal relative clause compound coordination of sentences simple independent & simple clause coordination of phrases apposition others coordinationsubclausesubordinatewhy apposition decomposer pool decomposed sentence figure 8: the structure of english sentences and corresponding decomposers (in red boxes). sentence decomposition does not know the type of sentence beforehand thus all sentences will go through the pool of decomposers. the dashed arrows just indicate which decomposer works on which type of sentence. 4.2 mrs decomposition for complex sentences 4.2.1 overview the mrs mapping between declarative and interrogative sentences only works for simple sentences. it generates lengthy questions from complex sentences, as illustrated in example 4. this is not a desirable result, as too much unnecessary information is provided in the questions. thus methods must be developed to obtain partial but intact semantic representations from complex sentences, so we can generate from simpler mrs representations. note, however, that the production of lengthy sentences is not a weakness of the generator, but in essence a problem of sentence simplification. this is what our mrs decomposition rules intend to tackle from the semantic level. mrsqg employs four decomposers for apposition, coordination, subclause and subordinate clauses. an extra why decomposer splits a causal sentence into two parts, reason and result, by extracting the arguments of the causal conjunction word, such as “because”, “the reason”, etc. the distribution of these decomposers is not random but depends on the structure of english sentences. english sentences are generally categorized into three types: simple, compound and complex depending on the type of clauses they contain. simple sentences do not contain dependent clauses. compound sentences are composed of at least two independent clauses. complex sentences must have at least one dependent clause and one independent clause. dependent clauses can be subordinate or relative clauses. subordinate clauses are typically introduced by a subordinating conjunction, while relatives are typically introduced by relative pronouns. 23 yao, bouma and zhang other phenomena that call for sentence simplification are coordination and apposition. coordination can be either clausal or phrasal. each type of sentence or grammar construction has a corresponding decomposer, shown in figure 8: • apposition is formed by two adjacent nouns describing the same reference in a sentence. an apposition decomposer simplifies the sentence “search giant google was found guilty” to “search giant was found guilty” and “google was found guilty” in order to prevent an ungrammatical replacement “*search giant who was found guilty?”. note that it is not usually easy to directly replace the compound “search giant google” with a question word since a named entity recognizer does not work well on compounds or noun phrase chunks. • coordination is formed by two or more elements connected with coordinators such as “and”, “or”. a coordination decomposer simplifies the sentence coordination “john likes cats and mary likes dogs” to “john likes cats” and “mary likes dogs” then generates from each simpler sentence. however, it avoids splitting coordination of nouns. for instance, “the cat and the dog live in the same place” is not split into “the cat live in the same place”, which is nonsense and ungrammatical. • a subordinate decomposer works on sentences containing dependent clauses. a dependent clause “depends” on the main clause, or an independent clause. thus it cannot stand alone as a complete sentence. it starts with a subordinate conjunction and also contains a subject and a predicate. for instance, the sentence “given that bart chases dogs, bart is a brave cat”is decomposed into two parts: “bart chases dogs”from the subordinate clause and“bart is a brave cat” from the independent clause. note that in some cases the proposition of the subordinate clause might be changed after decomposition. for instance, extracting “bart chases dogs” from “if bart chases dogs, bart is a brave cat” makes the proposition “bart chases dogs” true, while its true value is undetermined in the original sentence. the following subsection takes the subclause decomposer as an example and illustrates how it works. the order of applying these decomposers is not considered since in question generation from single sentences text cohesion is not important. 4.2.2 subclause decomposer a subclause decomposer works on sentences that contain relative clauses, such as this one. a relative clause is mainly indicated by relative pronouns, i.e., who, whom, whose, which, whomever, whatever, and that. extracting relative clauses from a sentence helps to ask better questions. for instance, given the following sentence: example 5 (a) bart is the cat that chases the dog. extracted relative clause after decomposition: (b) the cat chases the dog. parsing the above clause into logic form: (c) chase(cat, dog) replacing named entities with question words: (d) chase(which animal, dog) and chase(cat, which animal) generated questions from the transformed logic form: (e) which animal chases the dog? (f) which animal does the cat chase? it is impossible to ask a short question such as (e) and (f) directly from the original sentence (a) without dropping the main clause. a subclause decomposer serves to change this situation. 24 semantics-based question generation and implementation be v id named(”bart”) proper q rstr/h arg1/neq cat n 1 the q rstr/h chase v 1 dog n 1 the q rstr/h arg2/neqarg1/eq arg2/neq be v id named(”bart”) proper q rstr/h arg1/neq cat n 1 the q rstr/h arg2/neq be v id cat n 1 the q rstr/h chase v 1 dog n 1 the q rstr/h arg2/neqarg1/neq arg2/neq (a): bart is the cat that chases the dog. (b): bart is the cat. (c): the cat chases the dog. decompose({ chase v 1},{}) decompose({ be v id},{},keepeq = 0) figure 9: a subclause decomposer extracts relative clauses by finding all non-verb eps that are directly related to a verb ep. it also relaxes scope constraint from a relative clause. a strict eq relation in the top graph represents an np “the cat that chases the dog”. relaxing it to neq generates a sentence “the cat chases the dog.” in the bottom right graph. 25 yao, bouma and zhang be v id named(”mary”) proper q rstr/h arg1/neq girl n 1 the q rstr/h with p live v 1 named(”john”) proper q rstr/h arg1/neq arg1/eqarg2/eq arg2/neq be v id named(”mary”) proper q rstr/h arg1/neq girl n 1 the q rstr/h arg2/neq be v id girl n 1 the q rstr/h with p live v 1 proper q(named(”john”)) arg1/neq arg1/eqarg2/neq (a): mary is the girl john lives with. (b): mary is the girl. (c): john lives with the girl. decompose({ with p},{}) decompose({ be v id},{},keepeq = 0) figure 10: subclause decomposition for relative clauses with a preposition. the preposition ep with p connects the subclause with the main clause. after decomposition, the dependency relation between with p and girl n 1 is relaxed from arg2/eq in (a) to arg2/neq in (c). 26 semantics-based question generation and implementation though relative pronouns indicate relative clauses, in an mrs structure, these relative pronouns are not explicitly represented. for instance, in figure 9(a), there is no ep for the relative pronoun ”that”. however, the verb ep chase v 1 governs its subject by a post-slash eq relation. this indicates that chase v 1 and cat n 1 share the same label and have the same scope. after decomposing the sentence, this constraint of the same scope should be relaxed. thus in the mrs of “the cat chases the dog.”, chase v 1 and cat n 1 have different scopes, indicated by a post-slash neq relation. a generic decomposition algorithm should have a scope-relaxing step at the final stage. it is also possible to form a relative clause by leaving out the object of a preposition. although there is usually a relative pronoun in such cases (example 6a), it can also be left out (example 6b): example 6 (a) mary is the girl with whom john lives. (with a relative pronoun) (b) mary is the girl john lives with. (zero relative pronoun) the erg produces an identical mrs for such cases. the preposition is the word that connects the relative clause to the main clause. thus the subclause decomposer first starts from the preposition rather than the verb in the relative clause, as shown in figure 10. 4.2.3 general algorithm in this subsection we describe the general algorithm for sentence decomposition. this generic algorithm is the key step for decomposing complex mrs structures. before describing it, we first sum up the dependencies between eps and their intricacies. note that the following is only a list of examples and is incomplete. • arg*/eq: a special case that indicates scope identity. – adjectives govern nouns, such as brave a 1 → cat n 1 from “a brave cat”. – adverbs govern verbs, such as very+much a 1→ like v 1 from“like ... very much”. – prepositions govern and attach to verbs, such as with p → live v 1 from “live with ...”. – passive verbs govern and modify phrases, such as stir v 1→ martini n 1 from“a stirred martini”. – in a relative clause: ∗ verbs govern and modify phrases that are represented by relative pronouns, such as chase v 1 → cat n 1 from “the cat that chases ...” in figure 9. ∗ prepositions govern phrases that are represented by relative pronouns, such as with p → girl n 1 from “the girl whom john lives with”. • arg*/neq: the most general case that underspecifies scopes. – verbs govern nouns (subjects, objects, indirect objects, etc), such as chase v 1→ named(“bart”) and chase v 1 → dog n 1 from “bart chases dogs”. – prepositions govern phrases after them, such as with p → girl n 1 from “lives with the girl” in figure 10(c). – some grammatical construction relations govern their arguments, such as compound name → named(“google”) from “search giant google”. • arg*/null: a special case where an argument is empty. – passive verbs govern and modify phrases, e.g. in chase v 1→ dog n 1 from“the chased dog” and “the dog that was chased”, chase v 1 has no subject. 27 yao, bouma and zhang • arg*/h: a rare case where an argument qeq an ep. – subordinating conjunctions govern their arguments, such as given+that x subord → be v id from “given that arg2, arg1.”. – verbs govern verbs, such as tell v 1 rel→ like v 1 rel from“john told peter he likes mary.”. • null/eq: a rare case where two eps share the same label but preserve no equalities. • rstr/h: a common case for every noun phrase. – a noun is governed by its quantifier through a qeq relation, such as proper q→ named(“bart”) from “bart”. • l|r-index/neq: a common case for every coordinating conjunction. – coordinating conjunctions govern their arguments, such as and c → like v 1 from “lindex and r-index”. • l|r-hndl/heq: a special case for coordination of verb phrases. coordination of noun phrases do not have these relations. – the governor’s argument is its dependent’s label, such as but c→ like v 1 from “l-hndl but r-hndl”. we first give a formal definition of a connected dmrs graph, then a generic algorithm for dmrs decomposition. connected dmrs graph a connected dmrs graph is a tuple g = (n,e,l, spre, spost) of: a set n , whose elements are called nodes; a set e of connected pairs of vertices, called edges; a function l that returns the associated label for edges in e; a set spre of pre-slash labels and a set spost of post-slash labels. specifically, n is the set of all elementary predications (eps) defined in a grammar; spre contains all pre-slash labels, namely {arg*, rstr, l-index, r-index, l-hndl, r-hndl, null}; spost contains all post-slash labels, namely {eq, neq, h, heq, null}; l is defined as: l(x, y) = [pre/post, . . .]. for every node x, y ∈ n , l returns a list of pairs pre/post that pre ∈ spre, post ∈ spost. if pre 6= null, then the edge between (x, y) is directed: x is the governor, y is the dependant; otherwise the edge between x and y is not directed. if post = null, then y = null, x has no dependant by a pre relation. algorithm 1 shows how to find all related eps for some target eps. it is a graph traversal algorithm starting from a set of target eps for which we want to find related eps. it also accepts a set of exception eps that we always want to exclude. finally it returns all related eps to the initial set of target eps. there are two optional parameters for the algorithm that define different behaviors of graph traversal. relaxeq controls whether we want to relax the scope constraints from /eq in a subclause to /neq in a main clause. it works on both verbs and prepositions that head a relative clause. examples of this option taking effect can be found in figure 9(c) and 10(c). keepeq controls whether we want to keep a subclause introduced by a verb or preposition. it is set to true by 28 semantics-based question generation and implementation algorithm 1 a generic decomposing algorithm for connected dmrs graphs. function decompose(reps, eeps, relaxeq = 1, keepeq = 1) parameters: reps: a set of eps for which we want to find related eps. eeps: a set of exception eps. relaxeq: a boolean value of whether to relax the post-slash value from eq to neq for verbs and prepositions (optional, default:1). keepeq: a boolean value of whether to keep verbs and prepositions with a post-slash eq value (optional, default:1). returns: a set of eps that are related to reps ;; assuming concurrent modification of a set is permitted in a for loop aeps ← the set of all eps in the dmrs graph reteps ← ∅ ;; initialize an empty set for tep ∈ reps and tep /∈ eeps do for ep ∈ aeps and ep /∈ eeps and ep /∈ reps do pre/post← l(tep, ep) ;; ep is the dependant of tep if pre 6= null then ;; ep exists if relaxeq and post = eq and (tep is a verb ep or (tep is a preposition ep and pre = arg2)) then assign ep a new label and change its qeq relation accordingly end if reteps.add(ep) , aeps.remove(ep) end if pre/post← l(ep, tep ) ;; ep is the governor of tep if pre 6= null then ;; ep exists if keepeq = 0 and ep is a (verb ep or preposition ep) and post = eq and ep has no empty arg* then continue ;; continue the loop without going further below end if if not (ep is a verb ep and post = neq or post = h) then reteps.add(ep) , aeps.remove(ep) end if end if end for end for if reteps 6= ∅ then return decompose(reps ∪ reteps, eeps, relaxeq = 0) ;; the union of two else return reps end if 29 yao, bouma and zhang default. if set to false, the relative clause will be removed from the main clause. examples can be found in figure 9(b) and 10(b). all decomposers except the apposition decomposer employ this generic algorithm. figures 9 to 10 are marked with corresponding function calls of this algorithm that decompose one complex mrs structure to two simpler ones. 4.3 automatic generation with rankings mrsqg usually produces too much output. the reasons are twofold: firstly, pet and erg “overparse”. sometimes even a simple sentence has tens of parsed mrs structures due to subtle distinctions in the grammar. secondly, lkb and erg overgenerate. sometimes even a simple mrs structure has tens of realizations, due to word order freedom, lexical choice, etc. thus ranking is needed to select the best output. generation from lkb has already incorporated the maxent model by velldal and oepen (2006), which works best for declarative sentences as its training corpus does not include questions. to rank questions, we could have trained the maxent model on questions. but this requires a syntactically annotated (in hpsg treebank style) question corpus, which is time consuming and expensive to create. thus we took a shortcut to simply use language models trained on questions only. to construct a corpus consisting of only questions, data from the following sources was collected: • the question answering track (1999-2007)4 of the text retrieval conference (trec, voorhees, 2001) • the multilingual question answering campaign (2003-2009)5 from the cross-language evaluation forum (clef, braschler and peters, 2004) • a question classification (qc) dataset (li and roth, 2002, hovy et al., 2001)6 • a collection of yahoo!answers (liu and agichtein, 2008)7 the whole language model training procedure follows a standard protocol of“build-adapt-prune-test”. firstly we build a small language model for questions only. then this language model is adapted with the whole english wikipedia8 to increase lexicon coverage. finally we prune it for rapid access and test it for effectiveness. the irst language modeling toolkit (federico and cettolo, 2007) was used. to combine the scores from both the maxent model and the language model, we first project the log-based maxent score to linear space and then we use a weighted formula for linear scores to combine the linear maxent score with the sentence probability from the language model. suppose for a single mrs representation there are n different realizations r = [r1, r2, . . . rn ]. p (ri|me) is the normalized maxent score for the i-th realization under a maximum entropy model me. p (ri|lm) is the probability of each of the n different realizations given the language model lm. with two rankings p (ri|me) and p (ri|lm) we can borrow the idea of f-measure from information retrieval and combine them: r(ri) = fβ = (1 + β2) p (ri|me)p (ri|lm) β2p (ri|me) + p (ri|lm) (1) with β = 1 the ranking is unbiased. with β = 2 the ranking weights p (ri|lm) twice as much as p (ri|me). 4. http://trec.nist.gov/data/qamain.html 5. http://celct.isti.cnr.it/respubliqa/index.php?page=pages/pastcampaigns.php 6. http://l2r.cs.uiuc.edu/~cogcomp/data/qa/qc/ 7. http://ir.mathcs.emory.edu/shared/ 8. http://meta.wikimedia.org/wiki/wikipedia_machine_translation_project 30 http://trec.nist.gov/data/qamain.html http://celct.isti.cnr.it/respubliqa/index.php?page=pages/pastcampaigns.php http://l2r.cs.uiuc.edu/~cogcomp/data/qa/qc/ http://ir.mathcs.emory.edu/shared/ http://meta.wikimedia.org/wiki/wikipedia_machine_translation_project semantics-based question generation and implementation interrogatives declaratives trec clef qc yahooans all wikipedia sentence count 4,950 13,682 5,452 581,348 605,432 30,716,301 word count (k) 36.6 107.3 61.1 5,773.8 5,978.9 722,845.6 words/sentence 7.41 7.94 11.20 9.93 9.88 23.53 table 1: statistics of the data sources. “all” is the sum of the first four datasets. wikipedia is about 120 times larger than “all” in terms of words. figure 11 illustrates how question ranking works. since the maxent model is only trained on declarative sentences, it prefers sentences with a declarative structure. however, the language model trained with questions prefers auxiliary fronting. the weighted score can help to select the best interrogative sentences. 5. evaluation the evaluation of question generation with semantics was conducted as part of the question generation shared task and evaluation challenge (qgstec2010; rus et al., 2010; rus et al., this volume). participants are given a set of inputs consisting of an input sentence + question type and their system should generate two questions for each type. question types include yes/no, which, what, when, how many, where, why and who. input sources are wikipedia, openlearn9 and yahoo! answers. each source contributes 30 input sentences. there will be 360 questions generated in total. evaluation was conducted by independent human raters (but not necessarily native speakers). they follow the following criteria: 1. relevance. questions should be relevant to the input sentence. best/worse score: 1/4. 2. question type. questions should be of the specified target question type. best/worse score: 1/2. 3. syntactic correctness and fluency. the syntactic correctness is rated to ensure systems can generate sensible output. best/worse score: 1/4. 4. ambiguity. the question should make sense when asked more or less out of the blue. best/worse score: 1/3. 5. variety. pairs of questions in answer to a single input are evaluated on how different they are from each other. best/worse score: 1/3. note that all participants were asked to generate two questions of the same type. if only one question was generated, then the variety ranking of this question receives the lowest score and the missing question receives the lowest scores for all criteria. 5.1 evaluation results four systems participated in task b of qgstec2010: • lethbridge (ali et al., 2010), university of lethbridge, canada 9. http://openlearn.open.ac.uk 31 http://openlearn.open.ac.uk yao, bouma and zhang unranked realizations from lkb for “will the wedding be held next monday?” next monday the wedding will be held? next monday will the wedding be held? next monday, the wedding will be held? next monday, will the wedding be held? the wedding will be held next monday? will the wedding be held next monday? maxent scores 4.31 the wedding will be held next monday? 1.63 will the wedding be held next monday? 1.35 next monday the wedding will be held? 1.14 will the wedding be held next monday? 0.77 next monday, the wedding will be held? 0.51 next monday will the wedding be held? 0.29 next monday, will the wedding be held? language model scores 1.97 next monday will the wedding be held? 1.97 will the wedding be held next monday? 1.97 will the wedding be held next monday? 1.38 next monday, will the wedding be held? 1.01 the wedding will be held next monday? 0.95 next monday the wedding will be held? 0.75 next monday, the wedding will be held? ranked f1 scores 1.78 will the wedding be held next monday? 1.64 the wedding will be held next monday? 1.44 will the wedding be held next monday? 1.11 next monday the wedding will be held? 0.81 next monday will the wedding be held? 0.76 next monday, the wedding will be held? 0.48 next monday, will the wedding be held? figure 11: combined ranking using scores from a maxent model and a language model. the final combined scores come from equation 1 with β = 1. all scores are multiplied by 10 for better reading. note that the question “will the wedding be held next monday?” appears twice and has different scores with the maxent model. this is because the internal generation chart is different and thus it is regarded as different, even though the lexical wording is the same. 32 semantics-based question generation and implementation input sentences output questions count %(/90) mean std count %(/360) mean std mrsqg 89 98.9 19.27 6.94 354 98.3 12.36 7.40 wlv 76 84.4 19.36 7.21 165 45.8 13.75 7.22 juqgg 85 94.4 19.61 6.91 209 58.1 13.32 7.34 lethbridge 50 55.6 20.18 5.75 168 46.7 8.26 3.85 (a) coverage page 1 mrsqg wlv juqgg lethbridge 0.00% 10.00% 20.00% 30.00% 40.00% 50.00% 60.00% 70.00% 80.00% 90.00% 100.00% coverages on input and output (generating 360 questions from 90 sentences) sentences questions (b) table 2: generation coverage of four participants. each was supposed to generate 360 questions in all from 90 sentences. “mean” and “std” indicate the average and standard deviation of the input sentence length and output question length in words. • mrsqg, saarland university, germany • juqgg (pal et al., 2010), jadavpur university, india • wlv (varga and ha, 2010), university of wolverhampton, uk in this section we describe the evaluation result of mrsqg in comparison with other systems. 5.1.1 generation coverage the low coverage number in most cases of table 2 shows that generation from all input sentences and from all required question types was not an easy task. none of the systems reached 100%. among them, mrsqg, wlv and juqgg generated from more than 80% of the 90 sentences but lethbridge only managed to generate from 55.6% of them. as for the required 360 questions, wlv, juqgg and lethbridge only generated around 40% ˜ 60% of them. most systems can respond to the input sentences but only mrsqg has a good coverage on required output questions. figure 2b in table 2 illustrates this. good coverage on required questions usually depends on whether the employed named entity recognizer is able to identify corresponding terms for a question type and whether the reproduction 33 yao, bouma and zhang relevance question type correctness ambiguity variety mrsqg 1.61 1.13 2.06 1.52 1.78 wlv 1.17 1.06 1.75 1.30 2.08 juqgg 1.68 1.19 2.44 1.76 1.86 lethbridge 1.74 1.05 2.64 1.96 1.76 best/worst 1/4 1/2 1/4 1/3 1/3 agreement 63% 88% 46% 55% 58% (a) relevance question type correctness ambiguity variety 0.00 0.50 1.00 1.50 2.00 2.50 3.00 3.50 4.00 results per criterion without penalty on missing questions wlv mrsqg juqgg lethbridge worst (b) table 3: results per participant without penalty for missing questions. lower grades are better. agreement closer to 100% indicates better inter-rater reliability. ”worst” indicates the worst possible scores per evaluation category. rule covers enough structural variants. it also depends on whether the systems have been tuned on the development set and the strategy they employed. for instance, in order to guarantee high coverage, mrsqg can choose to sacrifice some performance in sentence correctness. some systems, such as wlv, seemed to focus on performance rather than coverage. the next subsection shows this point. 5.1.2 overall evaluation grades two human raters gave grades to each question according to the established criteria. then all grades were averaged and inter-rater agreement was calculated. due to the fact that most systems do not have a good coverage on required questions (c.f. table 2), the final grades were calculated with and without penalty on missing questions. table 3 presents the results without penalty on missing questions. grades were calculated based only on generated questions from each system. out of all grading criteria, syntactic correctness and fluency appears to be the hardest for all systems. apparently, syntactic or semantic transformations do not guarantee grammaticality and fluency of the generated question. the best scores are obtained for question type and relevance. this is not surprising, as the question type is given and the questions are generated on the basis of a single sentence. 34 semantics-based question generation and implementation relevance question type correctness ambiguity variety mrsqg 1.65 1.15 2.09 1.54 1.80 wlv 2.70 1.57 2.97 2.22 2.58 juqgg 2.65 1.53 3.10 2.28 2.34 lethbridge 2.95 1.56 3.36 2.52 2.42 (a) relevance question type correctness ambiguity variety 0.00 0.50 1.00 1.50 2.00 2.50 3.00 3.50 4.00 results per criterion with penalty on missing questions mrsqg wlv juqgg lethbridge worst (b) table 4: results per participant with penalty for missing questions. lower grades are better. ”worst” indicates the worst possible scores per evaluation category. table 4 shows the results if all questions that should be generated are taken into account. if a system failed to generate a question, it is assumed this system generates a question with the worst scores in all criteria. since wlv, juqgg and lethbridge have relatively low coverage (between 40% ˜ 60%), their scores deteriorate considerably. the scores for mrsqg are not affected that much, as it has a coverage of 99%. the inter-rater agreement for the scores shown by table 3 is not satisfactory. a score of over 80% usually indicates good agreement but only question type has achieved this standard. landis and koch (1977) have argued values between 0–.20 as slight, .21–.40 as fair, .41–.60 as moderate, .61–.80 as substantial, and .81–1 as almost perfect agreement. according to this standard, all the agreement numbers for other criteria only show a moderate or weakly substantial agreement between raters. these values reflect the fact that the rating criteria were ambiguously defined. 5.1.3 evaluation grades per question type qgstec2010 requires eight types of questions: yes/no, which, what, when, how many, where, why and who. figure 12 shows the individual scores of mrsqg on these questions. the quality of these questions mostly depends on the term extraction component. for instance, if a named entity recognizer fails or the ontology cannot provide a hypernym for a term, a which question cannot be properly generated. in general, mrsqg did the worst for which questions and the best for who questions. this reveals the strong and weak points of the term extraction component employed in 35 yao, bouma and zhang relevance question type correctness ambiguity variety 0 0.5 1 1.5 2 2.5 3 3.5 4 performance of mrsqg per question type yes/no(28) which(42) what(116) when(36) how many(44) where(28) why(30) who(30) worst(354) figure 12: performance of mrsqg per question type. numbers in () indicate how many questions were generated regarding each question type. the number in worst() is the sum of all questions. ”worst” indicates the worst possible scores per evaluation category. mrsqg. also, why questions involve more reasoning than others. since mrsqg does not have a reasoning component, it received the worst score in “ambiguity” on why questions. 6. discussion this section discusses various aspects in question generation with semantics. some of the discussion is based on a theoretical perspective while some involves implementation issues. the merits of deep processing with semantics are emphasized but their disadvantages are also addressed. 6.1 generation with semantics can produce better sentences with fewer rules one semantic representation can lead to several surface realizations. with a good ranking mechanism the best realization can be selected, which makes the generated question even more natural than the input sentence in some cases. take the english active and passive voice as an example: example 7 (a) the dog was chased by bart. question generated from syntactic transformation: (b) by whom was the dog chased? extra question generated from semantic transformation: (c) who chased the dog? by replacing the term with question words and fronting auxiliary verbs and question words, a syntaxbased system will normally only generate questions as good as (b). a semantics-based system, generating from chase(who, the dog), can produce both the passive form in (b) and the active form in (c). in most contexts, (c) would be preferred. (c) is shorter than (b) and sounds more natural. one might argue that with special treatment and carefully tested transformation rules, a syntaxbased system is capable of generating questions both in active and passive voices. but then again, similar problems arise in the case of di-transitive verbs: example 8 (a) john gave the waitress a one-hundred-dollar tip. 36 semantics-based question generation and implementation asking a question on the tip giver: (b) who gave the waitress a one-hundred-dollar tip? extra question generated from semantics-based system: (c) who gave a one-hundred-dollar tip to the waitress? a syntax-based system should have no problems producing (b). however, an additional rule is required to generate (c). for a semantics-based system, both (b) and (c) can be generated from the logical form give(who, one-hundred-dollar tip, waitress) and no special rules are required. generally in a semantics-based system, the generator is responsible for realizing a sentence in all the ways a language’s grammar permits. the question transformer only needs to address the semantic divergence between different questions. thus, a semantics-based system in general requires fewer rules to generate more questions (i.e. more word order variation) than a syntax-based system. 6.2 interface to lexical semantics resources although one semantic representation can produce multiple syntactic realizations, this procedure is handled internally through chart generation with erg. also, different syntactic realizations of the same meaning can be symbolized into the same semantic representation through parsing with erg. thus mrsqg is relieved of the burden to secure grammaticality and is able to focus on the semantics of languages. this in turn makes mrsqg capable of incorporating other lexical semantics resources seamlessly. for instance, given the predicate logic form have(john, dog) from the sentence “john has a dog” and applying ontologies to recognize that a dog is an animal, a predicate logic form have(john, what animal) for the question “what animal does john have” can be naturally derived. apart from using hypernym-hyponym10 relations to produce which questions, other lexical semantic relations can be employed as well. holonym-meronym11 relations can help ask more specific questions. for instance, given that a clutch is a part of a transmission system and a sentence “a clutch helps to change gears”, a more specific question “what part of a transmission system helps to change gears?” instead of merely “what helps to change gears?” can be asked. also, given the synonym of the cue word the task of lexical substitution can be performed to produce more lexical variations in a question. using lexical semantics resources needs the attention of a word sense disambiguation component. a misidentification of the sense of word might lead to nonsense questions. so expanding the range of questions with lexical semantics resources should be used with caution. note that there is no distinct difference between semantics-based system and syntax-based system in the ability to incorporate lexical semantics resources. but in terms of expressive simplicity and ease of use, a semantics-based system seems to us a more natural way to utilize lexical semantics resources. 6.3 language independence and domain adaptability in theory, mrsqg is language-neutral as it is based on semantic transformations. as long as there is a grammar12 conforming with the hpsg structure and lkb, adapting it to other languages should require little or no modification. however, the experience in multi-lingual grammar engineering has shown that although mrs offers a higher level of abstraction than syntax, it is difficult to guarantee absolute language independence. as a syntax-semantics interface, part of the mrs representation will inevitably carry some language specificity. as a consequence, the mrs transfer rules need to be adapted for the specific grammars, similar to the situation in mrs-based machine translation (oepen et al., 2004). the domain adaptability is confined to the following parts: 10. y is a hypernym of x if every x is a (kind of) y; y is a hyponym of x if every y is a (kind of) x. 11. y is a holonym of x if x is a part of y; y is a meronym of x if y is a part of x. 12. for a list of available grammars, check http://wiki.delph-in.net/moin/matrixtop 37 http://wiki.delph-in.net/moin/matrixtop yao, bouma and zhang 1. named entity recognizers. for a different domain, the recognizers must be re-trained. mrsqg also uses an ontology-based named entity recognizer. thus collections of domain-specific named entities can be easily plugged-in to mrsqg. 2. hpsg parser. the pet parser needs to be re-trained on a new domain with an hpsg treebank. however, since the underlying hpsg grammars are mainly hand-written, they normally generalize well and have a steady performance on different domains. 6.4 limitations of proposed method a semantics-based question generation system is theoretically sound and intuitive. but the implementation is limited to tools that are currently available, such as the grammar, the parser, the generator and the preprocessors. the central theme of mrsqg is mrs, which is an abstract syntacticsemantic interface employed by erg. the parser and generator cannot work without the erg either. thus the erg is indeed the backbone of the whole system. the heavy machinery employed by a deep precision grammar decreases both the parsing and generation speed and requires large memory footprints. thus mrsqg needs more resources and time to process the same sentence than a syntax-based system does. on the parsing side, mrsqg is not robust against ungrammatical sentences, due to the fact that erg is a rule-based grammar and only accepts grammatical sentences. the grammar coverage also decides the system performance in terms of recall value. but since erg has been developed for over ten years with great effort, robustness against ungrammaticality and rare grammatical constructions is only a minor limitation to mrsqg. on the generation side, all parsed mrs structures in theory should generate. but there exists the problem of overgeneration. as shown by velldal and oepen (2006), the maxent model achieves a 64.28% accuracy on the best sentence and 83.60% accuracy on the top-5 best sentences. thus the question given by mrsqg might not be the best one in some cases. on the preprocessing side, the types of questions that can be asked depend on the named entity recognizers and ontologies. if the named entity recognizer fails to recognize a term, mrsqg is only able to generate a yes/no question that requires no term extraction. even worse, if the named entity recognizer mistakenly recognizes a term, mrsqg generates wrong questions that might be confusing to people. on the side of sentence simplification, the decomposed sentences are not necessarily grammatical. even grammatical sentences might not make sense due to lack of information. for instance, given a restrictive relative clause, “bart is the cat that chases the dog”, one does not ask “who is the cat?” but rather “who is the cat that chases the dog?”. in this case, sentence simplification produces questions that are too vague and that cannot be understood without a context. our proposed system can only rely on the question ranking module to select a best one. in some other cases, such as the hypothetical clause “if bart chases dogs, bart is a brave cat”, the subordinate decomposer might extract a false statement “bart chases dogs” or “bart is a brave cat”, then a question that cannot find its answer from the input would be generated. ruling out these types of questions require either some extra rules, or a textual entailment module to judge whether the extracted simple clause can be inferred from the original sentence. from the theoretical point of view, the theory underlying mrsqg is dependency minimal recursion semantics, a variant of mrs. dmrs provides a connecting interface between a semantic language and a dependency structure. although still under development, it has been successfully applied to question generation via mrsqg . however, there are still redundancies in dmrs (such as the nondirectional null/eq “dependence”). some of the redundancies are even crucial: a mis-manipulation of the mrs structure inevitably leads to generation failure and thus makes the mrs transformation process fragile. fixing this issue is not trivial however, which requires the joint effort from the mrs theory, the erg grammar and the lkb generator. 38 semantics-based question generation and implementation from the application point of view, mrsqg limits itself to only generating from single sentences. expanding the input range to paragraphs or even articles is more interesting for applications but needs more sophisticated processing. thus mrsqg currently only serves as a starting point for the full task of question generation. 7. conclusion and future work this paper introduces the task of question generation and proposes a semantics-based method to perform this task. generating questions from the semantics of languages is intuitive but also has its difficulties, namely sentence simplification, question transformation and question ranking. this paper proposes three methods to address these issues: mrs decomposition for complex sentences to simplify sentences, mrs transformation for simple sentences to convert the semantic form of declarative sentences into that of interrogative sentences, and hybrid ranking to select the best questions. the underlying theoretical support comes from a dependency semantic representation (dmrs) while the backbone is an english deep precision grammar (erg) based on the hpsg framework. the core technology used in this paper is mrs decomposition and transfer. a generic decomposition algorithm is developed to perform sentence simplification, which boils down to solving a graph traversal problem following the labels of edges that encode linguistic properties. evaluation results reveal some of the fine points of this semantics-based method and also challenges that indicate future work. the proposed method works better than most other syntax/rulebased systems in terms of the correctness and variety of generated questions, mainly benefiting from the underlying precision grammar. however, the quality of yes/no questions is not the best among other question types. this indicates that the sentence decomposer does not work very well. the main reason is that it does not take text cohesion and context into account. thus sometimes the simplified sentences are ambiguous or even not related to the original sentences. enhancing the sentence decomposer to simplify complex sentences but still preserving enough information and semantic integrity is one of the future works. the low inter-rater agreement shows that the evaluation criteria are not well defined. also, the evaluation focuses on standalone questions without putting the task of question generation into an application scenario. this disconnects question generation from the requirements of actual applications. for instance, an intelligent tutoring system might prefer precise questions (achieving high precision by sacrificing recall) whilst a closed-domain question answering system might need as many questions as possible (achieving high recall by sacrificing precision). since the research of question generation has just started, efforts and results are still in a preliminary stage. combining question generation with specific application requirements has been put into the long-term schedule. to sum up, the method proposed by this paper produces so far the first open-source semanticsbased question generation system. it properly employs a series of deep processing steps which in turn lead to better results than most other shallow methods. theoretically, it describes the linguistic structure of a sentence as a dependency semantic graph and performs sentence simplification algorithmically, which has its potential usage in other fields of natural language processing. practically, the developed system is open-source and provides an automatic application framework that combines various tools for preprocessing, parsing, generation and mrs manipulation. the usage of this method will be further tested in specific application scenarios, such as intelligent tutoring systems and closed-domain question answering systems. acknowledgments this article is based on the first author’s master thesis (yao, 2010). 39 yao, bouma and zhang references peter adolphs, stephan oepen, ulrich callmeier, berthold crysmann, dan flickinger, and bernd kiefer. some fine points of hybrid natural language parsing. in proceedings of the sixth international language resources and evaluation (lrec’08), marrakech, morocco, may 2008. european language resources association (elra). husam ali, yllias chali, and sadid a. hasan. automation of question generation from sentences. in kristy elizabeth boyer and paul piwek, editors, proceedings of the third workshop on question generation, pittsburgh, pennsylvania, usa, 2010. s. bangalore and a.k. joshi. supertagging: an approach to almost parsing. computational linguistics, 25(2):237–265, 1999. m. braschler and c. peters. cross-language evaluation forum: objectives, results, achievements. information retrieval, 7(1):7–31, 2004. ulrich callmeier. pet – a platform for experimentation with efficient hpsg processing techniques. nat. lang. eng., 6(1):99–107, 2000. issn 1351-3249. j. carroll and s. oepen. high efficiency realization for a wide-coverage unification grammar. lecture notes in computer science, 3651:165, 2005. john carroll, ann copestake, dan flickinger, and victor poznanski. an efficient chart generator for (semi-)lexicalist grammars. in proceedings of the 7th european workshop on natural language generation (ewnlg’99), pages 86–95, toulouse, france, 1999. r. chandrasekar, c. doran, and b. srinivas. motivations and methods for text simplification. in proceedings of the 16th conference on computational linguistics, volume 2, pages 1041–1044, 1996. wei chen, gregory aist, and jack mostow. generating questions automatically from informational text. in proceedings of the 2nd workshop on question generation in craig, s.d. & dicheva, s. (eds.) (2009) aied 2009: 14th international conference on artificial intelligence in education: workshops proceedings, 2009. trevor cohn and mirella lapata. sentence compression as tree transduction. j. artif. intell. res. (jair), 34:637–674, 2009. a. copestake, d. flickinger, c. pollard, and i.a. sag. minimal recursion semantics: an introduction. research on language & computation, 3(4):281–332, 2005. ann copestake. implementing typed feature structure grammars. csli: stanford, 2002. ann copestake. dependency and (r)mrs. http://www.cl.cam.ac.uk/~aac10/papers/dmrs.pdf, 2008. b. dorr, d. zajic, and r. schwartz. hedge trimmer: a parse-and-trim approach to headline generation. in proceedings of the hlt-naacl 03 on text summarization workshop, volume 5, pages 1–8, 2003. jason eisner. learning non-isomorphic tree mappings for machine translation. in proceedings of the 41st annual meeting on association for computational linguistics volume 2, acl ’03, pages 205–208, stroudsburg, pa, usa, 2003. association for computational linguistics. marcello federico and mauro cettolo. efficient handling of n-gram language models for statistical machine translation. in statmt ’07: proceedings of the second workshop on statistical machine translation, pages 88–95, morristown, nj, usa, 2007. 40 http://www.cl.cam.ac.uk/~aac10/papers/dmrs.pdf semantics-based question generation and implementation christiane fellbaum, editor. wordnet: an electronic lexical database. mit press cambridge, ma, 1998. jenny rose finkel, trond grenager, and christopher manning. incorporating non-local information into information extraction systems by gibbs sampling. in acl ’05: proceedings of the 43rd annual meeting on association for computational linguistics, pages 363–370, morristown, nj, usa, 2005. dan flickinger. on building a more efficient grammar by exploiting types. natural language engineering, 6(1):15–28, 2000. issn 1351-3249. a. graesser, j. otero, a. corbett, d. flickinger, a. joshi, and l. vanderwende. guidelines for question generation shared task and evaluation campaigns. in v. rus and a. graesser, editors, the question generation shared task and evaluation challenge workshop report. the university of memphis, 2009. m. heilman and n. a. smith. question generation via overgenerating transformations and ranking. technical report, language technologies institute, carnegie mellon university technical report cmu-lti-09-013, 2009. m. heilman and n. a. smith. good question! statistical ranking for question generation. in proc. of naacl/hlt, 2010a. michael heilman and noah a. smith. extracting simplified statements for factual question generation. in proceedings of the 3rd workshop on question generation., 2010b. eduard hovy, laurie gerber, ulf hermjakob, chin-yew lin, and deepak ravichandran. toward semantics-based answer pinpointing. in hlt ’01: proceedings of the first international conference on human language technology research, pages 1–7, morristown, nj, usa, 2001. martin kay. chart generation. in proceedings of the 34th annual meeting on association for computational linguistics, pages 200–204, morristown, nj, usa, 1996. j. r. landis and g. g. koch. the measurement of observer agreement for categorical data. biometrics, 33(1):159–174, march 1977. xin li and dan roth. learning question classifiers. in proceedings of the 19th international conference on computational linguistics, pages 1–7, morristown, nj, usa, 2002. yandong liu and eugene agichtein. you’ve got answers: towards personalized models for predicting success in community question answering. in proceedings of annual meeting of the association of computational linguistics (acl), 2008. jack mostow and wei chen. generating instruction automatically for the reading strategy of selfquestioning. in proceeding of the 2009 conference on artificial intelligence in education, pages 465–472, amsterdam, the netherlands, 2009. ios press. stephan oepen, helge dyvik, jan tore lønning, erik velldal, dorothee beermann, john carroll, dan flickinger, lars hellan, janne bondi johannessen, paul meurer, torbjørn nordg̊ard, and victoria rosén. som å kapp-ete med trollet? towards mrs-based norwegian-english machine translation. in proceedings of the 10th international conference on theoretical and methodological issues in machine translation, baltimore, md, october 2004. santanu pal, tapabrata mondal, partha pakray, dipankar das, and sivaji bandyopadhyay. qgstec system description – juqgg: a rule based approach. in kristy elizabeth boyer and paul piwek, editors, proceedings of the third workshop on question generation, pittsburgh, pennsylvania, usa, 2010. 41 yao, bouma and zhang c.j. pollard and i.a. sag. head-driven phrase structure grammar. university of chicago press, 1994. v. rus, b. wyse, p. piwek, m. lintean, stoyanchev s., and c. moldovan. the first question generation shared task evaluation challenge. in proceedings of the 6th international natural language generation conference (inlg 2010), dublin, ireland, 2010. ivan a. sag and dan flickinger. generating questions with deep reversible grammars. in proceedings of the first workshop on the question generation shared task and evaluation challenge, arlington, va: nsf, 2008. lee schwartz, takako aikawa, and michel pahud. dynamic language learning tools. in proceedings of the 2004 instil/icall symposium, 2004. andrea varga and le an ha. wlv: a question generation system for the qgstec 2010 task b. in kristy elizabeth boyer and paul piwek, editors, proceedings of the third workshop on question generation, pittsburgh, pennsylvania, usa, 2010. e. velldal and s. oepen. statistical ranking in tactical generation. in proceedings of the 2006 conference on empirical methods in natural language processing, pages 517–525, 2006. ellen m. voorhees. the trec question answering track. nat. lang. eng., 7(4):361–378, 2001. brendan wyse and paul piwek. generating questions from openlearn study units. in proceedings of the 2nd workshop on question generation in craig, s.d. & dicheva, s. (eds.) (2009) aied 2009: 14th international conference on artificial intelligence in education: workshops proceedings, 2009. xuchen yao. question generation with minimal recursion semantics. master’s thesis, saarland university & university of groningen, 2010. yi zhang, valia kordoni, and erin fitzgerald. partial parse selection for robust deep processing. in proceedings of acl 2007 workshop on deep linguistic processing, pages 128–135, 2007. 42 1 introduction 2 related work 2.1 question transformation 2.2 sentence simplification 2.3 question ranking 2.4 perception of related work 3 background: theory, grammar and tools 3.1 minimal recursion semantics and english resource grammar 3.2 linguistic knowledge builder 3.3 parsing with pet 4 proposed method 4.1 mrs transformation for simple sentences 4.2 mrs decomposition for complex sentences 4.2.1 overview 4.2.2 subclause decomposer 4.2.3 general algorithm 4.3 automatic generation with rankings 5 evaluation 5.1 evaluation results 5.1.1 generation coverage 5.1.2 overall evaluation grades 5.1.3 evaluation grades per question type 6 discussion 6.1 generation with semantics can produce better sentences with fewer rules 6.2 interface to lexical semantics resources 6.3 language independence and domain adaptability 6.4 limitations of proposed method 7 conclusion and future work agmo_final_style.dvi dialogue and discourse 2(1) (2011) 83–111 doi: 10.5087/dad.2011.105 a general, abstract model of incremental dialogue processing david schlangen david .schlangen@uni-bielefeld.de faculty of linguistics and literature bielefeld university bielefeld, germany gabriel skantze gabriel@speech.kth .se department of speech, music and hearing kth stockholm, sweden editor: hannes rieser abstract we present a general model and conceptual framework for specifying architectures for incremental processing in dialogue systems, in particular with respectto the topology of the network of modules that make up the system, the way information flows through this network, how information increments are ‘packaged’, and how these increments are processed by the modules. this model enables the precise specification of incremental systems and hence facilitates detailed comparisons between systems, as well as giving guidance on designing newsystems. in particular, the model can serve as a framework for specifying module communication in such systems, as we illustrate with some examples. keywords: incremental processing, dialogue systems, software architecture 1. introduction dialogue processing is, by its very nature,incremental. no dialogue agent (artificial or natural) processes whole dialogues, if only for the simple reason that dialogues are createdincrementally, by participants taking turns at speaking (or typing, in computer-mediated, chat-like interaction). at this level, most current implemented dialogue systems are incremental: they process user utterances as a whole and produce their response utterances as a whole. incremental processing, as the term is commonly used, means more than this, however, namely that processing starts before the input is complete (e.g., (kilger and finkler, 1995)). incremental systems hence are those where “[e]ach processing component will be triggered into activity by a minimal amount of its characteristic input” (levelt, 1989). if we assume that the characteristic input of a dialogue system is the utterance (see (traum and heeman, 1997) for an attempt to define this unit), we wouldexpect an incremental system to work on units smaller than utterances. doing this then brings into the reach of computational modelling a whole range of behaviours that cannot otherwise be captured, like concurrent feedback (“uh-huh”, “yeah”), fast turn-taking, and collaborative utterance construction. our aim in the work presented here is to describe and give names to the options available to designers of incremental systems. the model that we specify isgeneralin the sense that it describes elements that, as we believe, are essential to incremental processing, and hence can be expected to c©2011 david schlangen and gabriel skantze submitted 1/2010; accepted 3/2011; published online 5/2011 schlangen and skantze play a role in most, if not even all, systems performing such processing, andabstractin the sense that it only describes properties of these elements and relations between them, but not concrete ways to instantiate them in computer implementations.1 specifically, we define some abstract data types, some abstract methods that are applicable to them, and a range of possible constraints on processing modules. the notions introduced here allow the (abstract) specification of a wide range of different systems, from non-incremental pipelines to fully incremental, asynchronous, parallel, predictive systems, thus making it possible to be explicit about similarities and differences between systems. we believe that this will be of great use in the future development of such systems, in that it makes clear the choices and trade-offs one can make. while we discuss some issues that arise when directly basing a module communication infrastructure for incremental systems on the model, our main focus here is on the conceptual framework, not on implementation details. what we are alsonot doing here is to argue for one particular ‘best architecture’—what this is depends on the particular aims of an implementation/model and on more low-level technical considerations (e.g., availability of processing modules). as we are not trying to proveproperties of the specified systems here, the formalisations we give are not supported by a formal semantics. in the next section, we first motivate the use of incremental processing in dialogue systems, and then give some examples of the possible differences in system architectures that we want to capture. in section 3, we present the abstract model that underlies the system andmodule specifications, of which we give some examples in section 4.2 in section 5, we discuss how the model can be used when designing the communication infrastructure of new incremental dialoguesystems. we close with a brief discussion of related work. 2. motivation 2.1 why model incremental processing? before we go into the details of our model, we take a step back and discuss possible reasons for using incremental processing in models of dialogue processing. this discussionwill be brief, however, as our main interest in this paper is not to convince anyone of choosing this styleof processing, but is rather to describe some of the options that are available once that choice has been made.3 depending on the goals one has when building a system, incremental processing may offer certain advantages over non-incremental processing (i.e., processingwhere only full utterances, or even longer units like turns, are considered by the system). for example, ifone tries to optimize thereactivityof the system, then incrementality may be attractive for the straightforward engineering reason that starting to process an input before it is complete can lead to the processing being finished faster compared to starting processing only after the input is complete. this can be true even if the incremental processing takes somewhat longer than the non-incremental processing, if that difference is not too high, as illustrated in figure 1. (of course, the algorithms used in incre1. we do not claim, of course, that the list of issues discussed here is exhaustive. moreover, even within the space that we sketch here, there are many detail questions that will still need much work to be answered fully. 2. these sections are based on schlangen and skantze (2009), but are revised and significantly extended. 3. see e.g. allen et al. (2001); aist et al. (2007); skantze and schlangen (2009); buß et al. (2010) for more comprehensive discussions and some empirical results regarding the useof incremental processing in applied dialogue systems. regarding models of human dialogue processing, there is an even longer thread of literature in psycholinguistics, amassing evidence that the human language processor worksincrementally; see eg. marslen-wilson (1973); altmann and steedman (1988); tanenhaus et al. (1995); van berkumet al. (2007). 84 an abstract model of incremental processing mental processing should not have a systematically worse complexity than theirnon-incremental counterparts.) figure 1: input processed by incremental and sequential processor. even a much slower incremental processor can finish before a sequential processor. in a similar vein, incremental processing can potentially promise betterquality of processing, if the system has provisions for interactions between modules, where the results of higher-level processing of earlier bits of input can influence lower-level processing of later bits of input. (e.g., if the syntactic structure computed for the recognised speech so far constrains how subsequent speech is recognised—standard practice in speech recognition.) if one takes, more generally,naturalnessas a parameter to optimisation, then incremental processing can bring a range of phenomena into reach that cannot otherwise be modelled. examples of such phenomena are, as mentioned above, concurrent feedback (“uh-huh”, “yeah”), fast turn-taking, collaborative utterance construction (buß and schlangen, 2010), andgeneration of hedges and selfcorrections (skantze and hjalmarsson, 2010). (see edlund et al. (2008) for a general discussion of the goal of naturalness for dialogue systems.) lastly, if one goes even further and choosesrealismas a modelling goal, treating the dialogue system as a computational model of human cognition (schlangen, 2009), then one would be well advised to make the model work incrementally, given the wealth of evidence that the human language processor works incrementally (see references above). 2.2 invariants and variants in incremental processing 2.2.1 modularity implemented dialogue systems are typically modular systems, where separable processing tasks are encapsulated in separate software modules. encapsulation of processing tasks can also be found in cognitively motivated theories (e.g., levelt’s model of generation, levelt (1989)).4 our first goal in the present work is to characterise such modular structure. this involvesa) identifying the modules in a system (abstractly, without saying yet too much about what exactly theydo) and b) specifying the connections between them. figure 2 shows three examples ofmodule networks, representations of systems in terms of their component modules and the connections between them. modules are represented by boxes, and connections by arrows indicating the path along which information flows between modules, the direction of this flow, and, through the use of parallel connections, the“bandwidth” of the 4. as an aside, note that modularity does not preclude extensive interaction between information sources, as postulated in current constraint-based theories of language processing (macdonald, 1994), see discussion in altmann and steedman (1988). 85 schlangen and skantze connection (see below). arrows not coming from or going to modules represent the global input(s) and output(s) to and from the system. figure 2: module network topologies our goal now is to facilitate exact and concise description of the differences between module networks such as those in the example. informally, the network on the left can be described as a simple pipeline with no parallel paths, the one in the middle as a pipeline enhanced with a parallel path, and the one on the right as a star-architecture; we want to be able to describe exactly the constraints that define each type of network. an important element of such a specification is to state how information flows in thesystem and between the modules, again in an abstract way, without saying much about the information itself (as the nature of the information depends on details of the actual modules). as mentioned above, the figure 2 indicates the direction of information flow (i.e., whose output is whoseinput) through the direction of edges and the “bandwidth” of a connection between modules. aconnection between two modules may be single-directional, if information can only flow in one way (which typically will be from a ‘lower-level’ to a ‘higher-level’ module, e.g. from one that iscloser to the input source to one that is further away, or from one that is closer to the generation source to one that is closer to the output channel). in a bi-directional setting, higher-level information may influence lower-level processing, for example through expressing expectations that the lower-level module should work towards. the ‘bandwidth’ of the connection is represented via the number of edgesbetween two nodes (the more edges, the higher the bandwidth), and this refers to whether parallel information can be sent or not. one possible source for such parallelism in an incrementaldialogue system is illustrated in figure 3 (right), where for some stretches of an input signal (a sound wave), alternative hypotheses are entertained (note that the boxes here donot represent modules, but rather bits of incremental information, namely words; flattened into strings, one alternativeconsists of four words, the other of three). we can view these alternative hypotheses about the same original signal as being parallel to each other (with respect to the input they are grounded in). 2.2.2 granularity even systems that modularise the dialogue processing task in roughly the sameway can still differ along another dimension of incremental processing, namely in how they divide up larger units into increments. as an example, one system may decide on words as the minimal units—we will call such units simply theincremental units(iu) of a system—whereas another system may package incremental results triggered by time, e.g. every 500ms, and not by linguistic units. more formally, 86 an abstract model of incremental processing figure 3: parallel information streams (left) and alternative hypotheses (right) figure 4: incremental input mapped to (less) incremental output we will express this difference as a difference in relation of incremental units to the larger unit of which they are increments. 2.2.3 input/output relations we also want to be able to specify how incremental bits of input of a processing module can relate to incremental bits of its output. figure 4 shows one possible configuration, where over time incremental bits of input (shown in the left column) accumulate before one bit of output (in the right column) is produced. (as for example in a parser that waits until it can compute a major phrase out of the words that are its input.) describing the range of possible module behaviours with respect to such input/output relations is another important element of the abstract modelpresented here. 2.2.4 revisability it is in the nature of incremental processing, where output is generated on the basis of incomplete input, that such output may have to be revised once more information becomesavailable (see (baumann et al., 2009) for a detailed investigation of the behaviour of asr in this respect). figure 5 illustrates such a case. at time-stept1, the available frames of acoustic features (shown in the left column) lead the processor, an automatic speech recogniser, to hypothesise that the word “four” has been spoken. this hypothesis is passed on from the internal state of the processor (middle column) to the output (right column). however, at time-pointt2, as additional acoustic frames have come in, it becomes clear that “forty” is a better hypothesis about the previous frames together with the new ones. it is now not enough to just output the new hypothesis: it is possiblethat later modules have 87 schlangen and skantze figure 5: example of hypothesis revision already started to work with the hypothesis “four”, so the changed status of this hypothesis has to be communicated as well. this is shown at time-stept3. note that the situation just described is different from, and comes additionalto, the uncertainty management in dialogue systems in general. it is often seen desirable to be able to represent a module’sconfidencein its hypotheses; indeed the increasingly popular approaches to dialogue management that make use of probablistic models for representing uncertaindialogue states (see e.g. williams and young (2007)) rely on such information. we want to distinguish provisions for representing such ‘staticconfidence’ (confidence in a hypothesis at a given time) from the dynamic aspects of potentially having to revise this confidence in the light of later information and having to percolate this revision through the module network. it is the latter that figure5 illustrates, and it is another aspect of incremental systems that may be handled differently indifferent systems. capturing these differences is the final aim of our model. 2.3 related work the model described here is inspired partially by young et al. (1989)’s token passing architecture; our model can be seen as a (substantial) generalisation of the idea of passing smaller information bits around, out of the domain of asr and into the system as a whole. some of the characterisations of the behaviour of incremental modules are inspired by (kilger and finkler, 1995), but adapted and extended from incremental generation to the general case of incrementalprocessing. while there recently have been a number of papers about incremental systems (e.g., (devault and stone, 2003; aist et al., 2006; brick and scheutz, 2007)), none of those offer general considerations about architectures,5 which is what we are trying to offer in the following. 3. the model 3.1 overview we model a dialogue processing system in an abstract way as a collection ofconnected processing modules, where information is passed between the modules along these connections. the third component beside the modules and their connections is the basic unit of information that is communicated between the modules, which we call theincremental unit(iu). we will only characterise those properties of ius that are needed for our purpose of specifying different system types and ba5. despite its title, (aist et al., 2006) also only describes one particular setup. 88 an abstract model of incremental processing sic operations needed for incremental processing; we will not say anything about the actual, module specificpayloads of these units. the processing module itself is modelled as consisting of aleft buffer (lb), the processor proper, and aright buffer(rb). when talking about operations of the processor, we will sometimes useleft buffer-incremental unit(lb-iu) for units in lb andright buffer-incremental unit(rbiu) for units in rb. this setup is illustrated in figure 5 above. ius in lb (here, acoustic frames as input to an asr) are consumedby the processor (i.e., are processed) which creates an internal result; in the case shown here, this internal result ispostedas an rb-iu only after a series of lb-ius have accumulated. in reality, such processing will obviously take some time to complete. for the purposes of our abstract description here, however, we abstract away from such processing times and describe processors as relations between (sets of) lbs and rbs. we begin our description of the model with the specification of network topologies. 3.2 network topology connections between modules are expressed throughconnectedness axiomswhich simply state that ius in one module’s right buffer are also in another buffer’s left buffer. (again, in an implemented system communication between modules will take time, but we abstract away fromthis here.) this connection can also be partial or filtered. for example,∀x(x ∈ rb1 ∧ np (x) ↔ x ∈ lb2) expresses that all and only nps in module one’s right buffer appear in module two’s left buffer. if desired, a given rb can be connected to more than one lb, and more than one rb can feed into the same lb (see the middle example in figure 2). together, the set of these axiomsdefines the network topology of a concrete system. different topology types can then be defined through constraints on module sets and their connections. i.e., a pipeline system is one in which it cannot happen that an iu is in more than one right buffer and more than one left buffer. note that for now we are assuming token identity and not for example copyingof data structures. that is, we assume that it indeed is thesameiu that is in the left and right buffers of connected modules, and hence any changes made to an iu are immediately known to all modules that contain it in a buffer. this allows us to abstract away from the actual “transportation” of ius between modules (but see section 5 below). 3.3 incremental units so far, all we have said about ius is that they are holding a ‘minimal amount of characteristic input’ (or, of course, a minimal amount of characteristicoutput, which is to become some other module’s input). communicating just these minimal information bits is enough only for the simplest kind of system that we consider, a pipeline with only a single stream of information and no revision. if more advanced features are desired, there needs to be more structure tothe ius. in this section we describe the kinds of information that the most capable systems need to represent, and that make possible operations like hypothesis revision, prediction, and parallel hypothesis processing. (these operations will be explained in the next section.) if in a particular system someof these operations aren’t required, some of the structure on ius can be simplified. informally, the representational desiderata are as follows. first, one needs to be able to express certain relations between ius, which record how ius within a module and between modules belong together. (recall that ius are meant to be incremental, minimal bits of what will at some point 89 schlangen and skantze amount to one larger unit; e.g. words that form an utterance, or chunks of semantics that will eventually form the interpretation of an utterance, etc.) these relations come in two basic types; we discuss them in detail in the next two subsections. apart from the relations,we want ius to carry three other types of information as well: a confidence score representingthe confidence its producer had in it being accurate; information about whether revisions of the iu arestill to be expected or not; and information about whether the iu has already been processed by consumers, and if so, by whom. a full specification of all types of information about ius one might want to track is given below in section 3.3.4. 3.3.1 horizontal relations between ius the first type of relation expresses what could be described as ‘horizontal’ relationships between ius; horizontal, because these relationships hold only between ius of thesame level (produced by the same module), and typically reflect in some way the temporal order in which the ius were created. the most important instance of this type of relation is thesuccessorrelation, which links ius in such a way that from a chain of ius linked with this relation, the (possibly still partial) larger unit these ius are increments of can be read off. to make this description concrete: this relation will link ius representing words (for example, in the output of an asr) in such a way that the (possibly still partial) utterance hypothesis can be produced by following the links. moreformally, if two asr word hypothesis-ius iu1 and iu2 stand in this relation, this expresses that the asr takes iu2 to be the continuation of the utterance whose previous element is iu1. if another iu, iu3, were linked to iu1 as well, this would express that the sequences iu1-iu2 and iu1-iu3 form alternatives. (see figure 3, right; this relation is also illustrated in figure 6, discussed below.) this is the most basic kind of horizontal relation, which most incremental systems will have to realise at least implicitly (since it keeps track of how increments form larger units of information). for other purposes, one may also want to allow other relations of this type. for example, dependency relations between concepts would be of this type as well. figure 6 illustrates thisfor a possible semantic representation for the two recognition alternatives “move the ball twofields” / “move the ball to field z”. here we have one iu holding the concept of a move action, and linked to it, as further specifying it, a concept representing the patient of the action (ball -5) and, as alternatives corresponding to the alternative recognition results, two different directional concepts (distance:2 / “two fields” andtarget:z / “field z”). to recognise at this level that these two directional concepts are alternatives and do not both hold, one must follow thesuccessorrelation, where they are on different branches. this illustrates that thespecifiesrelation as used here (and represented in the figure by dashed arrows) cannot be subsumed by the successor relation (solid arrows in the figure). 3.3.2 hierarchical relations the other kind of relation orders ius into an informational hierarchy, linkingius to those bits of information—other ius—on which they depend informationally, or, as we will call it, in which they aregrounded. examples for this are the links between ius representing segments of audio material and ius representing the corresponding word hypotheses (i.e.,lb-ius and rb-ius of a speech recogniser); between word hypotheses and semantic representations (i.e., lb-ius and rbius of a semantic parser; in both these examples the links cut across module boundaries); between larger phrases in a parser and their constituents (e.g., an iu representing a sentence-level parse and 90 an abstract model of incremental processing figure 6: an example network, built by the relationssuccessor(solid arrows),specifies(wide dashed arrows),grounded in(narrow dashed arrows) ius representing an np and a vp, in which case the link holds between ius from the same level); or between states of the dialogue manager and plans for system utterances. figure 6 shows such grounded inlinks between word-ius and semantic concept ius, as narrow dashed lines. this relation effectively tracks the flow of information through the system; following its transitive closure one can go back from the highest level iu, which is output by the system, to the input iu or set of input ius on which it is ultimately grounded.6 going in the other direction, this is also the path along which doubts of a module about the quality of an iu, or ultimately its revocation, can percolate through the module network. for example, should an asr decide toremove support for a word hypothesis, then all ius that are linked to this iu via this relation will be thrown into doubt as well. (see the purge operation described below.) 3.3.3 meta-information we also assume that it will be useful in some systems to be able to represent and communicate additional information about ius. for example, in certain setups it may be useful for a module to learn whether its consumers have found a use for a hypothesis or not. this could be a parser that gets told whether an np it has built has a denotation in the current dialogue context or not, and that could then base a decision on whether to extend the parse or not on that information. (see below in section 4.1 the discussion of ‘early interaction systems’.) 6. of course, the output of a dialogue system is not fully determined just by its input. other elements of the state figure as well; e.g., some parts of an answer to a query will depend on a database and not just the query, however, the intention to respond with a statement will still be grounded in the input signal. 91 schlangen and skantze the other type of information concerns the confidence of a module in its hypotheses. as noted above, this confidence can change during the processing of further input; in extreme cases, such changes in confidence may lead to revocation of hypotheses on which possibly already later hypotheses were built. finally, in building systems ourselves, we have found that it can be usefulto add a binary notion of commitmentthat signals that a iu will not (or shall not) be modified again. a module can declare such commitment on its own ius in cases where for technical reasons (e.g., tokeep internal state to a manageable size) it decides that it will not touch them anymore. it can alsodemand commitment from the producers of its input ius, if it has based unrevisable actions on them; should the producing module still have a need to revise them, then this would lead to an exceptional situation that should be handled specifically. (e.g., by performing an explicit self-correction.)note that we intend a rather technical, house-keeping use for this. there is a more general sense in which modules are ‘committed’ to their ius, but this we think is covered by providing confidence information and by agreeing on a system-internal strategy for putting out ius in the first place(e.g., by either optimizing for speed, sending out even possibly very transient hypotheses, orby optimizing for quality, possibly holding back ius until internal confidence has reached a certain threshold). 3.3.4 formal specification we define formally the universeu in which incremental units live as follows. it consists of a set of iu objects,iu (which includes a special iu⊤), a set of module labelsm, and a collection of functions and relations defined on these: • i is an identification function, which, for ease of reference, maps each iuobject onto a unique id (e.g., a natural number). • a function c that maps each iu to a module label, its creator. (from this follows that each iu can only have one creator.) • a family of relations between ius created by the same module (i.e., iusα, β, wherei(α) 6= i(β) andc(α) = c(β)); we will call these relationssame level links. one instance of this type is thesuccessorrelation discussed above, which defines a (partial) order on ius. we specify that by default, ius are successor of a special iu⊤. this guarantees that ius linked via the successorrelation form a connected graph rooted in the⊤ element. the type of the graph will depend on the purposes of the sending and consuming module(s). for a one-best output of an asr it might be enough for the graph to be a chain (andsuccessorhence be a total order), whereas an n-best output might be better represented as a tree (with all first words linked to ⊤) or a lattice (as in figure 3, right). there can be other kinds of such same level relations, as described above. • g is thegrounded inrelation, connecting an iu to one or more ius out of which it was built. for example, an iu holding a (partial) parse might be grounded in a set of word hypothesis ius, and these in turn might be grounded in sets of ius holding acoustic features. while thesame level linkalways points to ius on the same level,grounded inlinks can hold both between ius from the same level as well as between ius of connectedmodules; in both cases it expresses that one bit of information depends on other bits. thetransitive closure of this relation hence links system output ius to a set of system input ius. forconvenience, we may define a relationsupports(x,y)for cases wherey is grounded inx; and hence the closure of this relation links input-ius to the output that is (eventually) built on them. 92 an abstract model of incremental processing this is also the hook for the mechanism that realises the revision process described above with figure 5: if a module decides to revoke one of its hypotheses, it sets its confidence value (see below) to 0; on noticing this event, all consuming modules can then check whether they have produced rb-ius that link to this lb-iu, and do the same for them. in this way, information about revision will automatically percolate through the module network. finally, we can also define a special value (e.g.,n.n.) for this relation and use it to trigger prediction: if an rb-iu is grounded in-related ton.n. , this can be understood as a directive to the processor to find evidence for this iu (i.e., to prove it), using the information in its left buffer. • t is theconfidence(or trust) function, which maps each iu to a numerical value (depending on the module’s needs, either fromn or r) through which the generating processor can pass on its confidence in its hypothesis. this then can have an influence ondecisions of the consuming processor. for example, if there are parallel hypothesesof different quality (confidence), a processor may decide to process (and produce output for) the best first. a special value (e.g., 0, or -1) can be defined to flag hypotheses that are being revoked by a processor, as described above. • c is thecommittedproperty, which holds when a producing module has committed to the iu, i.e., it guarantees that it will never revoke the iu. see below for a discussion of how such a decision may be made, and how it travels through the module network. • s is theseenrelation, relating ius and module labels. using this relation, processors can record whether they have “looked at”—that is, attempted to process—the iu. in the simplest case, the positive fact can be represented simply by adding the processor id to the list (assuming here that all processors in a system are identified by a unique id); in more complicated setups one may want to offer status information like “is being processed by module id” or “no use has been found for iu by module id”, or “i assign this iu a probability of n” via additional relations. this allows processors both to keep track of which lb-ius they have already looked at (and hence, to more easily identify new material that may have entered their lb) and to recognise which of its rb-ius have been of use to later modules,information which can then be used for example to make decisions on which hypothesis to expand next. • p finally relates ius to linguistic objects (like words or parses) which are their actualpayload, i.e., the module-specific unit of ‘characteristic input’ (or output). systems can differ in which of these elements they realise, and even modules within a system can differ along this dimension. it seems plausible to assume that most incremental systems will have concepts that can be mapped to what may be called the core set of iu properties,(i, c, same level link, grounded in link, t ,p), while more sophisticated processing (e.g., using prediction and a high degree of interaction between modules) will make use of the other properties. 3.4 modules as explained above, modules consist of left and right buffers, and processors with internal state that operate on input ius to create output ius. normally, the direction of thisupdate operation is from the left to the right (meaning that lb-ius are ‘consumed’ to producenew rb-ius that are grounded in them), but in the case of expectation-guided modules can also be from right to left (i.e., attempting to find evidence for some predicted output). a central concept when looking at modules is that of the update step, which has three stages: 1) the left buffer of themodule is updated from 93 schlangen and skantze figure 7: a schematic view of the update process state lbt to state lbt′ ; 2) the current module recognises what has changed, and performs itsown update first of its internal state (ist to is′ t) and then 3) of its own right buffer (rbt to state rb′t; this in turn will be the first update step for all consuming modules). figure 7 illustrates this flow of information through a module, showing only those parts of the module that are relevant at each timestep. we use these labels for the update stages when we now first discuss some more properties of buffers, then list operations that processors have to perform, and finally describe different possible module behaviours. 3.4.1 buffers for the purposes of the abstract model, we can conceptualise buffers simply as sets of ius, for which the constraints hold that are specified by the connectedness axioms as described above. during the execution of a system, buffers change over time (through the updates described above). to denote the state of a buffer at a given timet, we define a functionstate from time labels to iu sets. the changes made to a buffer from timet to t′ (i.e., to get fromstate(t) to state(t′)) we can then denote by∆t,t′ . we will assume for now that the processor only receives this delta, and bases its computations on this. we discuss alternative set-ups below in section 5. we also define a function are which returns the currentactive right edgeof a buffer; this is the set of ius that are a) currently active (we assume that ius that belong to input that has been fully processed are marked as inactive); b) not revoked (see below); and c) maximal with respect to thesuccessorrelation (recall that this defines a partial order rooted in⊤). this is the right edge insofar as this is where new increments will attach (e.g., where a new word will be added as more audio material comes in). 94 an abstract model of incremental processing 3.4.2 processoroperations at the most abstract level, the job of processors is to react to changes in their buffers (the∆t,t′) by performing appropriate updates. the elementary task here is to react tothe appearance of new lb-ius and eventually build new rb-ius out of them (or, in the case of a predictive system, to react to new expectations being entered as rb-ius and to evaluate subsequentlb-ius as to whether they provide evidence for these expectations). how exactly this is done is specific to the individual tasks of the module (e.g., asr, parser, dialogue manager, etc.), and we won’t have anything to say about this here; what we will describe here are the differenttypesof updates that can be implemented in an incremental processing module. we also leave openhow processors are triggered into action; we simply assume that on receiving new lb-ius or rb-ius or noticing changes to already known ius—more generally, on being given∆t,t′ of the appropriate buffer—they will eventually perform these operations. again, we describe here the complete set of operations; systems may differ in which subset of the functions they implement, or even whether they implement these operations as recognisably separate steps at all. purge lb-ius that are revoked by their producer must be purged from the internal state of the processor (so that they will not be used in future updates) and all rb-ius grounded in them must be revoked as well. some reasons for revoking hypotheses have already been mentioned. for example, a speech recogniser might decide that a previously output word hypothesis is not valid anymore (i.e., is not anymore among the n-best that are passed on). or, a parser might decide in the light of new evidence that a certain structure it has built is a dead end, and withdraw support for it. in all these cases,all ‘later’ hypotheses that build on this iu (i.e., all hypotheses that are in the transitive closure of this iu’s supportrelation) must be purged. if all modules implement the purge operation, this revision information will be guaranteed to travel through the network. new iu update new lb-ius have to be integrated into the internal state, and eventually new rbius are built based on them (not necessarily in the same frequency as new lb-ius are received; see figure 4 above, and discussion below). the new rb-ius have to be related appropriately to other ius (e.g., viasame level links, grounded in points, etc.). as mentioned above, this is the most basic operation of a processor, and can be expected to be implemented in all systems. processors can takesupportsinformation into account when deciding on the order in which they update. a processor might for example decide to first try to use the newinformation (in its lb) to extend structures that have already proven useful to later modules(that is, that support new ius). for example, a parser might decide to follow an interpretation path thatis deemed more likely by a contextual processing module (which has grounded hypotheses in the partial path). this may result in better use of resources—the downside of such a strategy of course is that modules can be garden-pathed.7 update may also work towards a goal. as mentioned above, putting ungrounded ius in a module’s rb can be understood as a request to the module to try to find evidencefor it. for example, the dialogue manager might decide based on the dialogue context that a certain type of dialogue act is likely to follow. by requesting the dialogue act recognition module to find evidence for this hypothesis, it can direct processing resources towards this task. (the dialogue recognition module then can in turn decide on which evidence it would like to see, and ask lower modules to prove this. ideally, 7. it depends on the goals behind building the model whether this is considered a downside or desired behaviour. 95 schlangen and skantze this could filter down to the interface module, the asr, and guide its hypothesis forming. technically, something like this is probably easier to realise by other means; see (schuler et al., 2009), briefly discussed below, for an example of an integrated approach, where semantics and reference resolution can directly bear on the speech recognition process.) we finally note that in certain setups it may be necessary to consume different types of ius in one module. as explained above, we allow more than one module to feed into another module lb. an example where something like this could be useful is in the processing of multi-modal information, where information about both words spoken and gestures performed may be needed to compute an interpretation. commit there are three ways in which a processor may have to deal with commits. first, it can decide for itself to commit rb-ius. for example, a parser may decide to commit to apreviously built structure if it failed to integrate into it a certain number of new words, thusassuming that the previous structure is complete. second, a processor may notice that a previous module has committed to ius in its lb. this might be used by the processor to remove internal state kept for potential revisions. eventually, this commitment of previous modules might lead theprocessor to also commit to its output, thus triggering a chain of commitments. interestingly, it can also make sense to let commits flow from right to left, as briefly discussed above. for example, if the system has committed to a certain interpretation by making a publicly observable action (e.g., an utterance, or an action in another modality), this can be represented as a commit on ius. this information would then travel down the processing network; leading to the potential for a clash between a revoke message coming from the left and thecommit directive from the right. in such a case, where the justification for an action is revoked when the action has already been performed, self-correction behaviours can be executed.8 other updates finally, in some settings it may also be desirable to let modules change the confidence score of ius after having put them into the rb (and so after they have already potentially been consumed by later modules); the consuming modules then might need to react to this change, perhaps by updating their internal state, by changing their future update strategy, or by changing their own confidence in something they have passed on into their own right buffer. it may also be useful in certain settings to allow other aspects of ius to be changed later aswell, such assame level links; again, this would be something that consuming modules need to notice and react to. 3.4.3 characterising module behaviour modules can also be characterised through a description of the changesthat updates yield to buffers and internal states, and the relations between changes to left buffers and those to right buffers. we list several dimensions along which such a characterisation can be made. updates to iu sequences using the notion of aright edgefrom section 3.4.1, we can transfer some terms from (wiŕen, 1992): a module isleft-to-right incrementalif it only producesextensions to the current right edge; within what wirén (1992) callsfully incremental, we can make further distinctions, namely according to whether only revisions or also insertions and deletions are allowed. (revisions are covered by the revoke operation described above; insertions and deletions can be expressed in our model by allowingsuccessorlinks to be changed appropriately.) when we want 8. in future work, we will explore if and how (e.g. through the implementation of a self-monitoring cycle with commits and revokes) the various types of dysfluency described by levelt (1989) and others can be modeled. 96 an abstract model of incremental processing to make this further distinction, we call the formerright edge revision-incrementaland the latter insertion/deletion incremental. processor-internal incrementality we can also distinguish modules according to how they update their internal state. we call modules that keep their internal state betweenupdate steps and only enrich it according to the current∆ of their input bufferinternally incremental(and the algorithms they use for doing sofully incremental algorithms). while this perhaps conforms best with an intuitive understanding of what incremental processing is, one can also imagine a different strategy (which has indeed been realised (devault et al., 2009), as briefly reviewed below). in this strategy, all internal state is thrown away between updates, and output is always computed from scratch using the full currently available input and not just the newest increments of it; wewill call such modules restart incremental. this strategy can be used when one has available more conventional processing modules which happen to be robust against partial input, but are not built to handle incremental changes to their input. update frequency this dimension concerns how the update frequency of lb-ius relates to that of (connected) rb-ius. we write f:in=out for modules that guarantee that every new lb-iu will lead to a new rb-iu (that is grounded in the lb-iu). in such a setup, the consuming module lagsbehind the sending module only for exactly the time it needs to process the input. following nivre (2004), we can call this strict incrementality. f:in≥outdescribes modules that potentially collect a certain amount of lb-ius before producing an rb-iu based on them. this situation has been depicted in figure 4 above. f:out≥in characterises modules that update rbmoreoften than their lb is updated. this could happen in modules that produce endogenous information like clock signals,or that produce continuously improving hypotheses over the same input (see below), or modules that ‘expand’ their input, like a tts that produces audio frames. connectedness we may also want to distinguish between modules that produce ‘island’ hypotheses that are, at least when initially posted, not connected to previously generated output ius via a common element that dominates them throughgrounded inlinks, and those that guarantee that this is not the case. for example, to achieve anf:in=out behaviour, a parser may output hypotheses that are not connected to previous hypotheses, in which case we may call the hypotheses ‘unconnected’. conversely, to guarantee connectedness, a parsing module might need toaccumulate input, resulting in an f:in≥out behaviour, or may need to speculate on continuations, possibly resulting inf:in≤out behaviour.9 completeness we define thecompletenessof a set of ius which are connected via the successor relation informally as the relation of the sequence they form (e.g., a sequence of words understood as a prefix of an utterance) to (the type of) what would count as a maximal sequence. for example, for an asr module, such a maximal sequence may be the transcription of a whole utterance and 9. the notion ofconnectednessis adapted from (sturt and lombardo, 2005), who provide evidence that the human parser strives for connectedness. 97 schlangen and skantze not just a prefix of one; for the parser maximal output may be a parse of type sentence (as opposed to one of type np, for example), etc.10 building on this notion, we can characterise modules according to completeness of their lb and rb. in ac:in=out-type module, the most complete set of rb-ius is only as complete as the most complete set of lb-ius. that is, the module does not speculate about completions, nor does it lag behind. (this may technically be difficult to realise, and practically not veryrelevant.) more interesting is the difference between the following types: in ac:in≥out-type module, the most complete set of rb-ius potentially lags behind the most complete set of lb-ius. this will typically be the case inf:in≥out modules.c:out≥in-type modules finally potentially produce output that ismore complete than their input, i.e., theypredict continuations. an extreme case would be a module that always predicts complete output, given partial input. such a module may be useful in cases where modules have to be used later in the processing chain that can only handle complete input (that is, are non-incremental); we may call such a systemprefix-based predictive, semi-incremental. (again, (devault et al., 2009) is an example of such a module; as is (schlangen et al., 2009).) with these categories in hand, we can make further distinctions within what dean and boddy (1988) call anytime algorithms. such algorithms are defined as a) producing output at any time, which however b) improves in quality as the algorithm is given more time. incremental modulesby definition implement a reduced form of a): they may not produce an output at any time,but they do produce output at more times than non-incremental modules. this output then also improves over time, fulfilling condition b), since more input becomes available and either the guessesthe module made (if it is a c:out≥in module) will improve or the completeness in general increases (as more complete rb-ius are produced). processing modules, however, can also be anytime algorithms in a more restricted sense, namely if they continuously produce new and improved output even for a constant set of lb-ius, i.e. without changes on the input side. (which would bring themtowards thef:out≥in behaviour.) as a final note, we can now see that a non-incremental system can be characterised as a special case of an incremental system, namely one where ius are always maximally complete (with c:in=out) and where all modules update in one go (f:in=out). (typically, in such systems ius will also always be committed, but this need not necessarily be the case for asystem to be nonincremental.) 3.5 system specification combining all these elements, we can finally define a system specification as thefollowing: • a list of modules that are part of the system. • for each of those a description in terms of which operations from section 3.4.2 the module implements, and a characterisation of its behaviour in the terms of section 3.4.3. • a set of axioms describing the connections between module buffers (and hence the network topology), as explained in section 3.2. 10. this definition is only used here for abstractly classifying modules. practically, it is of course rarely possible to know how complete or incomplete an ongoing input is. investigating how a dialoguesystem can better predict completion of an utterance is in fact one of the aims of the project in which this framework was developed. 98 an abstract model of incremental processing figure 8: the numbers system architecture (ca = communicative act) • specifications of the format of the ius that are produced by each module, interms of the definition of slots in section 3.3. for technical reasons, one may also want to specify which information given about ius may be changed later, and which can be considered immutable. 4. some example specifications 4.1 ‘early interaction’ systems most previous work on incremental processing has focused on one ofthe many possible advantages of this processing style, namely on making available ‘higher-level’ informationto ‘lower-level’ processes, where this information can then be used to guide processing.this typically has taken the form of letting a parser interact with extra-syntactic knowledge. (despite considerable differences in the way this effect is achieved, (devault and stone, 2003; stoness etal., 2005; aist et al., 2006, 2007; brick et al., 2007; brick and scheutz, 2007) can all be subsumedunder this description.) the general approach can be described in iu terms as follows: the parser posts certain constituents (nps and vps) as rb-ius, connected modules filter out the phrase type they care about and evaluate them in the domain (e.g., a domain ontology checks which operations are possible in the domain, and what likely frames are that express a certain action; or a module tests whether nps have a denotation in the domain). this evaluation is attached to the iu (via theseenrelation, for example); the parser then must be capable of noticing updates to an rb-iu and act accordingly (e.g., modify the chart that forms its internal state). the cited papers all focus on this interaction and do not say much about thesystems in which this interaction is realised, so we cannot give full system specifications here. 4.2 the numbers system the numbers system (skantze and schlangen, 2009) has a special status here because it can not just be usefully described in the terms explained here, it actually directly instantiates some of the concepts and methods described in this paper. the module network topology of the system is shown in figure 8. this is pretty much a standard dialogue system layout, with the exception that prosodic analysis is done in the asr and that 99 schlangen and skantze dialogue management is divided into a discourse modelling module and an action manager. as can be seen in the figure, there is also a self-monitoring feedback loop—the system’s actions are sent from the tts (text-to-speech synthesizer) to the discourse modeller. thesystem has two modules that interface with the environment (i.e., are system boundaries): the asr and the tts. a single hypothesis chain connects the modules (that is, no two same level linkspoint to the same iu). modules pass messages between them that can be seen as xml-encodings of iu-tokens. information strictly flows from lb to rb. all iu slots except seen (s) are realised. the purge and commit operations are fully implemented. in the asr, revision occurs as already described above with figure 4, and word-hypothesis ius are committed (and the speech recognition search space is cleared) after 2 seconds of silence are detected. (note that later moduleswork with all ius from the moment that they are sent, and do not have to wait for them being committed.) the parser may revoke its hypotheses if the asr revokes the words it produces, but also if it recovers from a “garden path”, having built and closed off a larger structure too early. as a heuristic, the parser waits until a syntactic construct is followed by three words that are not part of it untilit commits. for each new discourse model increment, the action manager may produce new communicative acts (cas), and possibly revoke previous ones that have become obsolete. when the system has spoken a ca, this ca becomes committed, which is recorded by the discourse modeller. no hypothesis testing is done (that is, no un-grounded information is put onrbs). all modules have af:in≥out; c:in≥outcharacteristic; that is, they may collect information in the form of lb-ius before they generate rb-ius and hence potentially lag behind somewhat. the system achieves a very high degree of responsiveness—by using incremental asr and prosodic analysis for turn-taking decisions, it can react in around 200ms when suitable places for backchannels are detected, which should be compared to a typical minimum latency of 750ms in common systems where only a simple silence threshold is used.11 4.3 a prefix-based, predictive system the module described in (devault et al., 2009) has already been mentioneda couple of times above. the module is an nlu component that outputs full semantic frames, even if its input is only a partial utterance. in our terms, it isprefix-based predictive, with c:out≤in. the module is also ‘internalevent based’ in that updates are triggered by events of an internal clock (the asr is polled every 200ms) and not by the event of receiving a new lb-iu (more on this distinction in the next section). no internal state is kept between update steps, so output is always computed on the basis of the latest, possibly still partial, full input and not on the newest increments only;the module is only restart-incremental. 4.4 incremental generation in the deal system skantze and hjalmarsson (2010) describe an approach to incremental speech generation in dialogue systems that is based on the model presented here. the approach allows adialogue system to incrementally interpret spoken input while simultaneously planning, realising and self-monitoring the system response. if the system detects that the user has stopped speaking and it is appropriate for the system to take the turn, the system may start to speak, even if it does not yet have a complete plan of what to say, or if the input ius are not yet committed. as the input is processed, the action 11. a video showing an example interaction with the system can be found athttp://www.purl.org/ net/numbers-sds-video. 100 an abstract model of incremental processing figure 9: the right buffer of an asr (top) and the speech plan that at isincrementally produced (bottom). vertex s1 is associated with w1, s3 with w3, etc. optional speech segments are marked with dashed outline. manager in the system builds a tentative speech plan, which may then later be revised. if output ius need to be revised, and they have already been spoken, the system will automatically perform an overt self-repair, using editing terms such as “sorry, i mean”. in order to facilitate incremental speech generation, system utterances are made up of smaller ius. the action manager incrementally produces aspeech plan, which is a graph that represents different ways of realising a message. each edge in this graph is associated withspeech segment. the speech output module may then traverse this graph in different ways,depending on a number of constraints, such as timing. each speech segment is also made up of smallerspeech units. these mark locations where an utterance may be aborted, or where self-corrections may occur. figure 9 illustrates how a speech plan may be incrementally produced, as words are recognized by the speech recognizer. the grounded-in links may then be followed all the way back from the speech units to the speech segments, to the speech plan, to the communicative acts in the discourse model that it was a response to, and finally to the phrases and the words in the user utterance. thereby, a revision in the speech recogniser may trigger a revision in the spoken output. this system has been evaluated in a wizard-of-oz setting, and achieved faster reaction times than a version of the system with incremental generation disabled, and was judged more polite, more effective, and better at indicating when to speak. 5. using the iu model as a middle layer in incremental dialogue systems as we said in the introduction, the model as laid out in the previous sections is meant to describe the design space for incremental systems, and one of its applications is the uniform description (and consequently, easier comparison) of extant models of incremental processing (including different models of human incremental language processing). however, the model has also proved useful for us in the design of new systems (as described above; other systems are currently under development). in this section, we discuss some additional conceptual issues that need to be addressed when basing a system on this model.12 we present this in the form of a list of questions that a system designer must answer. again, we do not give recommendations for one particular solution but rather 12. we will remain mostly on the conceptual level here. a lower level description of implementational problems and links to a collection of reference implementations of the framework can be found in schlangen et al. (2010). 101 schlangen and skantze figure 10: ius as structured objects. the fields are: id, producing module, target ofsame-levellink, target ofgrounded-in-link, confidence value, committed, seen-by, payload. describe the available options, and the circumstances under which certain choices are advisable. in subsection 5.2 we then present a more detailed description of two differentways to derive a working system infrastructure from these concepts. 5.1 design issues how are ius represented? the first question to decide is how to represent ius. as we said above, besides the payload of an iu, which holds the actual incremental ‘chunk’ of information that is to be exchanged, there are a number of other properties and relations of ius that one might need to keep track of in a system. a straightforward way of doing so is to realise ius as a data structure with several fields that hold values (for which most programming languages have basic built-in datatypes), with booleans for properties and values for relations;figure 10 illustrates this approach (for two asr-ius that are linked viasuccessor, of which the first is committed, and both have been used by an nlu component). however, a less direct approach may also be appropriate, where properties of ius are indirectly represented via properties of thebuffers.13 section 5.2 will give an example of such an approach, where the properties of being revoked or being committed are represented indirectly and must be inferred from the state of the buffer. how are buffers synchronised between modules?the second, and more interesting challenge is to implement the flow of information—i.e., the flow of ius—through the system. in the abstract model explained above, we have treated buffers as sets of entities, and have represented connections between modules (via their buffers) by relations between such sets. doing this is enough from the point of view ofanalysisof a system, as there it is enough to know that information in one buffer is guaranteed to appear in another buffer as well. when designing a system, however, one has to actually make this happen. what is to be achieved, then, is that ius in a producer’s right buffer must appear in all of its consumers’ left buffers, and all changes of properties of the ius must be synchronised; in other words, it must be guaranteed that after an update to a rb (step rbt to rbt′ in figure 7) all connected lbs must be updated as well (must perform their step lbt to lbt′), so that the connectedness axioms hold for rbt′ and lbt′ . (or, respectively, for updates to lbs in case of right-to-left information flow.) there are two interrelated aspects to this: one is how the path along which the information travels is realised, the other is what actually travels along this path. figure 11sketches some of the available options. 13. this is one reason why we have taken care in section 3.3.4 above to only specify that there are certain properties that need to be represented, and have not said anything about the concrete data structures to be used for representing them. 102 an abstract model of incremental processing figure 11: options for realising paths and messages: shared memory (left); direct communication with either full copies of buffers or edit messages (middle); mediated communication via blackboard, again full copies or edit messages (right) first, the path: in a set-up where the modules in question operate in a sharedmemory space (figure 11, left), strictly speaking there is no path along which the informationhas to travel. in such a case ius can simply be data objects in shared memory, and these objects andall changes made to them are immediately accessible to all modules; in this sense, this realisations corresponds most closely to the abstract conceptualisation explained above. (if modules work asynchronously—see discussion below—the usual precautions have to be taken to avoid inconsistencies when modifying shared memory.) in set-ups where the modules need to be more independent from each other, perhaps running on different machines and/or being implemented in different programming languages, a ‘virtual shared memory’ can be created via message passing. figure 11 in the middle shows a set-up where producing module and consuming module exchange instructions on how to achieve synchronisation, and on the right it shows a variant of this where the ‘virtual shared memory’ is managed by a dedicated module that implements a ‘blackboard’ (a common set-up in ai systems,provided for example by the open agent architecture, (cheyer and martin, 2001)). second, the information that travels: the question here is whether the updated buffer is communicated as a whole (rbt′ is sent, to replace lbt and yield lbt′) or whether only the changes that were made when updating it are communicated (as message like “this is how you can turn lbt into lbt′”). in the former case, it is left to the consuming module to compute∆t,t′ , whereas in the latter case this information is provided by the producer; in most cases this may be the more efficient solution. all these options realise the same functionality, namely that buffers are synchronised. which option is the most appropriate in a given situation depends on other considerations, e.g. on how tightly the modules can be integrated, on whether existing resources have to be re-used, etc. how are updates to modules triggered? another question then is how updates in a module are triggered. there are two basic modes of operation here: one, updates can be triggered by the receptionof new information. this requires that new information is automaticallypushedto the consuming module; updates are hence driven by the external event of aproducing module pushing new information into the buffer. (more accurately: this event triggers acheckof whether an update 103 schlangen and skantze to internal state is required; there may be cases where new information doesnot actually have an effect on the consuming module.) the other option is to let a consuming module query the producing modules it is interested in for new information at self-determined intervals (or triggered by other, endogenous, internal events). this could then be called a ‘pull model’. mixtures of these approaches are of course possible, where some information is pushed and other pulled on demand (for example, new ius are pushed, but later changes to the confidence slot may not be pushed but only be queried when that information is relevant). how is control distributed between modules? on creation of an iu, are all further processing steps performed in sequence in all later modules, before control returnsand the next iu can be created by this module? or are modules running asynchronously (in threads, processes, agents)? the former case is more similar to normal dialogue systems, with the only change being that smaller increments travel through the network. the latter case requires more changes, to deal with possible concurrency problems. (of course, asynchronicity is not only possible in incremental systems, see e.g. (boye et al., 2000) for an early example of an asynchronous non-incremental system.) what is preserved between updates? is internal state cleared at the beginning of each update, or not? if it is cleared, then module needs to process the complete buffer containing all ius that span the input so far (and in fact doesn’t really work fully incrementally;see above the definition of restart incremental.) this is an appropriate choice if a module is used that can handle partial input, but cannot incrementally update its internal state. what is the relation between buffers and internal states? lastly, one needs to think about how deeply the modules are encapsulated and protected from each other. theway we have described it so far, the buffers are distinct from the internal state of a module, and a module may not need to take everything from its left buffer into its internal state, and not every processing step that leads to changes in its internal state needs to lead to changes of its right buffer. however, one can also imagine cases where there is a more direct connection between modules, andoutput of one module is written directly into the internal state of another module (e.g., words from an asr are put directly into a parse chart) or something is read directly out of an internal state. 5.2 two examples of iu-architectures we now discuss some more concrete details of how to plan the communication infrastructure of an incremental system. we do this in two variants, where different system goals and preconditions are taken into account. these descriptions are loosely based on our current work on two different systems.14 5.2.1 modules asagents; bi-directional information flow assume that we are in the following situation: we need the modules of the system we’re building to run on different machines (with different computer architectures), because we have to use legacy components (e.g., a vision system that only runs under microsoft windowsr© whereas other components run onunix machines). this means that we cannot rely on shared memory for implementing 14. note that while the choices illustrated in these specifications do cluster together naturally, they by no means are completely dependent—it is possible to combine features of the model in other ways, e.g. have a distributed system with “right edge” encoding of changes (see below). 104 an abstract model of incremental processing the module buffers. we also want the system to be predictive, expectation-driven; that is, we want ‘right-to-left’ information flow. what we sketch here as a solution to these requirements is a pretty straightforward implementation of the model as detailed above. first, we decide to implement ius as objects,with information fields as illustrated above in figure 10; as mentioned there, support for such objects is present in many different programming languages, so we’re not (very) restrictedon this. since we can’t rely on (process specific) shared memory, we have to achieve synchronisation of buffers through explicit communication between modules. details of how to set up such communication (registration of modules, routing and buffering of messages etc.) are beyond the scope of this paper, we just note that there is a variety of off-the-shelf solutions for this tasks (e.g., corba,15 ice,16 oaa (cheyer and martin, 2001), and many more). what’s more interesting is which messages we want to send. the following is a list of (informally specified) messages that realise the full spectrum of capabilities modelled in the iu approach. keep in mind that the purpose of sending these messages is twofold: to achieve synchronisation between buffers, and at the same time to encapsulate what has changed; that is, the messages represent the∆ of the update. first, from producer to consumer(s): • add this iu (full iu, all fields); this is, as noted above, the most basic type of information exchange, presenting a new increment for consumption. if parallel hypotheses are allowed in the system, the consumer has to look up to which of the active hypothesis chains the new iu is to be added. it can do this by checking thesuccessorlink on the new iu. • revoke this iu (iu id); indicates that a producer withdraws support for a previously posted hypothesis. • reinstate this iu (iu id); makes a previously revoked iu active again. • confidence change (iu id, new value); notifying consumer(s) about a change of the confidence slot. (note that if revocation is signalled via special values for confidence, then the previous two messages are only syntactic sugar for certain settings of this message.) • commit (iu id); which is a guarantee to the consumer(s) that no update or revocation will be performed anymore on this iu. • no support (iu id); this communicates to the consumer that an iu that the consumer wanted to be proven (that is, prediction; see below) could not be validated, and that the module has given up. (it is in some sense the ‘negative’ of a commitment.) consumers can send ‘back’ to producers the following messages: • used (iu id); indicating that the consumer could make use of a given iu, and hence the larger hypothesis it is part of may be promising. • useless (iu id); the consumer can show that no use can be made of the hypothesis path ending in this iu in the current context. 15. see e.g.http://www.omg.org/gettingstarted/corbafaq.htm. 16.http://www.zeroc.com/ 105 schlangen and skantze • commit (iu id); the consumer has based an irreversible decision on this iu, e.g. by making an observable action, and so an exception should be raised if the producer wants to revoke it. • support (full iu); indicating that the consumer expects this to be the producer’s output. while it would also be possible for consumers to query producers for newmessages, the more straightforward solution is to let the producers send such messages whenever they have made changes to their right buffers (i.e., on making the update step rbt to rb′ t from figure 7). updates to a module can then be triggered by reception of these messages; i.e., the module is event-driven. from the requirement stated at the beginning of this section that modules needto run on different machines it follows that they need to run in different processes. again thestraightforward solution then is to let the modules communicate asynchronously and run concurrently,with all the possible pitfalls that concurrent programming entails (see e.g. (van roy and haridi,2004)).17 5.2.2 close coupling ofmodules; shared memory; efficient encoding of updates imagine a different scenario. we know that we are going to build all modules of the incremental system from scratch, and we know that we can do this within one programminglanguage. hence, we can make use of shared memory to realise buffers, the option sketched infigure 11 (left). we also do not anticipate a need for prediction. in such a setting, we can use acompact alternative representation of iu networks, which makes communication more efficient. in a shared memory setup, synchronisation of buffers is not an issue, as right buffers and left buffers are in fact the same memory space. only the task of notifying consuming modules about what exactly changed remains, of what they have to look for when they access the shared memory. the idea here is now to reify positional information by superimposing a network of position nodes over the iu network, with the ius being associated with edges in that network. these positional nodes then give us names for certain update stages, and so revisions can be efficiently encoded by reference to these nodes. an example can make this clearer. figure 12 shows five update steps in the right buffer of an incremental asr module. by reference to positional nodes, we can communicate easily a) what the newest (rightmost; as explained above in section 3.4.1) committed iu is (we will call this nc fornewest committed; indicated in the figure as a shaded node) and b) what the newest non-revoked or active iu is (i.e., the right edge (re); indicated in the figure as a node with a dashed line). so, the change between the state at timet1 andt2 is signalled by re taking on a different value. this value (w3) has not been seen before, and so the consumingmodule can infer that the network has been extended; it can find out which ius have been addedby going back from the new re to the last previously seen position (in this case, w2; note that this would withno changes work for parallel hypothesis threads, if the positional network were a lattice). at t3, a retraction of a hypothesis is signalled by a return to a previous state, w2. all consuming modules have to do now is to return to an internal state linked to this previous input state (more on this in a second). finally, t5 illustrates a commitment, where nc changes, and all ius on the path from the newnc to the last committed iu now count as committed. this representation style can be considered more parsimonious, as what inthe approach from the previous section were three different messages (add, revoke, commit) ishere implicitly expressed in 17. it would be interesting to formally analyse the potential for parallelisationin dialogue processing, with petri-nets or other modelling tools. we leave this to future work. 106 an abstract model of incremental processing ����� ����� � ���� ������ ������������ �� ��� ���� � ��� �� �� � � ����� ��� � ��� ����� ��� ����� � ��� �� �� ��� ���������� ��� � ��� ����� ��� ����� ��� � ��� ���� �� ��� � �� ��� � � �� �� �� ��� � � �� �� �� ��� � � �� �� � �� ������ �� �� ��� � � �� �� � �� ������ �� figure 12: the right buffer of an asr module, and update messages compactly representing revokes and commits ����� ����� � ���� ������ ������� ������ ��� ������ � � ��� ������� �� � � ��� ������ � � � ����� � �� �� �� � ����� � � � � �� �� �� ���� � ����� � � � � �� �� �� ���� figure 13: connecting input states and output states of a buffer 107 schlangen and skantze changes to the pair (nc, re). the status of an iu, and changes to it, are represented in the network and this pair, and so ius themselves can be immutable (unchangeable) objectsin this approach. this avoids the pitfalls of having concurrent modules altering and accessing thestate of the ius. message size is always constant in this approach, even if comprehensive changes are made to a buffer. e.g., in t4 in figure 12, two ius are added, but this still requires only one message, an update to what the label of the current right edge is. also, coupling between modules can betighter in this approach, as states of the consumer can be directly tied to states of the producer. this is illustrated in figure 13. states of a parser (which is, for ease of presentation, not doing a veryinteresting job here, simply mapping number words to their logical representations as numbers) can be linked to the updates that created them, so that revision on the input state just requires going back to aprevious output state. in the example, this happens int3, where the return of re to w2 results in a return of the previously computed consumer state p2. 6. conclusions and future work we have presented a general, abstract model of incremental dialogue processing. the model is general in the sense that it describes elements that are essential to incremental processing, and hence can be expected to play a role in most, if not even all, systems performing such processing. it is abstractin the sense that it only describes properties of these elements and relationsbetween them, but not concrete ways to instantiate them in computer implementations. we illustrated the notions developed here through the description of a number of existing systems in these terms, and we discussed some questions that arise when trying to build such systems. in future work, we will attempt to describe more existing systems (such as (devault and stone, 2003; aist et al., 2006; brick and scheutz, 2007)) in the terms developedhere, to more thoroughly investigate the coverage of our concepts. we are also currently exploring how more cognitively motivated models such as the model of speech generation by (levelt, 1989)can be specified in our framework. a further direction for extension is the implementation of modalityfusion as iuprocessing. lastly, we are now starting to work on connecting the model for incremental processing and grounding of interpretations in previous processing results described here with models of dialogue-level grounding in the information-state update tradition (larssonand traum, 2000). the first point of contact here will be the investigation of self-corrections, as a phenomenon that connects sub-utterance processing and discourse-level processing (ginzburg et al., 2007). acknowledgements this work was supported by a grant from dfg in the emmy noether programme. thanks to timo baumann and okko buß for discussions of the ideas reported here. we would also like to thank the anonymous reviewers of the precursor to this paper, (schlangen and skantze, 2009), and especially of this revised and extended version, for their very detailed and helpfulcomments. references gregory aist, james allen, ellen campana, lucian galescu, carlos a. gomezgallo, scott stoness, mary swift, and michael k. tanenhaus. software architectures for incremental understanding of human speech. inproceedings of the international conference on spoken language processing 108 an abstract model of incremental processing (icslp), pittsburgh, pa, usa, september 2006. gregory aist, james allen, ellen campana, carlos gomez gallo, scott stoness, mary swift, and michael k. tanenhaus. incremental understanding in human-computer dialogue and experimental evidence for advantages over nonincremental methods. inproceedings of decalog 2007, the 11th international workshop on the semantics and pragmatics of dialogue, trento, italy, 2007. james allen, george ferguson, and amanda stent. an architecture for morerealistic conversational systems. inproceedings of the conference on intelligent user interfaces, santa fe, usa, june 2001. gerry altmann and mark steedman. interaction with context during human sentence processing. cognition, 30:191–238, 1988. timo baumann, michaela atterer, and david schlangen. assessing and improving the performance of speech recognition for incremental systems. inproceedings of the north american chapter of the association for computational linguistics human language technologies (naacl hlt) 2009 conference, boulder, colorado, usa, may 2009. johan boye, beth ann hockey, and manny rayner. asynchronous dialogue management: two case-studies. inproceedings of the 4th workshop on semantics and pragmatics of dialogue (götalog2000), pages 51–55, gothenburg, sweden, 2000. timothy brick and matthias scheutz. incremental natural language processing for hri. inproceedings of the second acm ieee international conference on human-robot interaction, pages 263–270, washington, dc, usa, 2007. timothy brick, paul schermerhorn, and matthias scheutz. speech and action: integration of action and language for mobile robots. inproceedings of the 2007 ieee/rsj international conference on intelligent robots and systems (iros ’07), san diego, ca, usa, october/november 2007. okko buß and david schlangen. modelling sub-utterance phenomena in spoken dialogue systems. in proceedings of the 14th international workshop on the semantics and pragmatics of dialogue (pozdial 2010), pages 33–41, poznan, poland, june 2010. okko buß, timo baumann, and david schlangen. collaborating on utterances with a spoken dialogue system using an isu-based approach to incremental dialogue management. inproceedings of the sigdial 2010 conference, pages 233–236, tokyo, japan, september 2010. adam cheyer and david martin. the open agent architecture.journal of autonomous agents and multi-agent systems, 4(1):143–148, march 2001. oaa. thomas dean and mark boddy. an analysis of time-dependent planning. in proceedings of aaai88, pages 49–54. aaai, 1988. david devault and matthew stone. domain inference in incremental interpretation. inproceedings of icos 4: workshop on inference in computational semantics, nancy, france, september 2003. inria lorraine. 109 schlangen and skantze david devault, kenji sagae, and david traum. can i finish? learning whento respond to incremental interpretation results in interactive dialogue. inproceedings of the 10th annual sigdial meeting on discourse and dialogue (sigdial’09), london, uk, september 2009. jens edlund, joakim gustafson, mattias heldner, and anna hjalmarsson.towards human-like spoken dialogue systems.speech communication, 50:630–645, 2008. jonathan ginzburg, raquel fernández, and david schlangen. unifying selfand other-repair. in proceeding of decalog, the 11th international workshop on the semantics and pragmatics of dialogue (semdial07), trento, italy, june 2007. anne kilger and wolfgang finkler. incremental generation for real-time applications. technical report rr-95-11, dfki, saarbrücken, germany, 1995. staffan larsson and david traum. information state and dialogue management in the trindi dialogue move engine toolkit.natural language engineering, pages 323–340, 2000. willem j.m. levelt.speaking. mit press, cambridge, usa, 1989. maryellen c. macdonald. probabilistic constraints and syntactic ambiguity resolution. language and cognitive processes, 9(2):157–201, 1994. william d. marslen-wilson. linguistic structure and speech shadowing at very short latencies. nature, 244:522–523, august 1973. joakim nivre. incrementality in deterministic dependency parsing. pages 50–57, barcelona, spain, july 2004. david schlangen. what we can learn from dialogue systems that don’t work: on dialogue systems as cognitive models. inproceedings of diaholmia, the 13th international workshop on the semantics and pragmatics of dialogue (semdial 2009), pages 51–58, stockholm, sweden, june 2009. david schlangen and gabriel skantze. a general, abstract model of incremental dialogue processing. inproceedings of the 12th conference of the european chapter of the association for computational linguistics (eacl 2009), pages 710–718, athens, greece, march 2009. david schlangen, timo baumann, and michaela atterer. incremental reference resolution: the task, metrics for evaluation, and a bayesian filtering model that is sensitive todisfluencies. in proceedings of sigdial 2009, the 10th annual sigdial meeting on discourse and dialogue, london, uk, september 2009. david schlangen, timo baumann, hendrik buschmeier, okko buß, stefan kopp, gabriel skantze, and ramin yaghoubzadeh. middleware for incremental processing in conversational agents. in proceedings of the sigdial 2010 conference, pages 51–54, tokyo, japan, september 2010. william schuler, stephen wu, and lane schwartz. a framework for fast incremental interpretation during speech decoding.computational linguistics, 35(3), 2009. gabriel skantze and anna hjalmarsson. towards incremental speech generation in dialogue systems. inproceedings of the sigdial 2010 conference, pages 1–8, tokyo, japan, september 2010. 110 an abstract model of incremental processing gabriel skantze and david schlangen. incremental dialogue processing in a micro-domain. inproceedings of the 12th conference of the european chapter of the association for computational linguistics (eacl 2009), pages 745–753, athens, greece, march 2009. scott c. stoness, james allen, greg aist, and mary swift. using real-worldreference to improve spoken language understanding. inproceedings of workshop on spoken language understanding at aaai05, pittsburgh, pa, usa, 2005. patrick sturt and vincenzo lombardo. processing coordinated structures: incrementality and connectedness.cognitive science, 29:291–305, 2005. michael k. tanenhaus, michael j. spivey-knowlton, kathleen m. eberhard, and julie c. sedivy. intergration of visual and linguistic information in spoken language comprehension. science, 268, 1995. d. traum and p. heeman. utterance units in spoken dialogue. in e. maier,m. mast, and s. luperfoy, editors,dialogue processing in spoken language systems, lecture notes in artificial intelligence. springer-verlag, 1997. jos j.a. van berkum, arnout w. koornneef, marte otten, and mante s. nieuwland. establishing reference in language comprehension: an electrophysiological perspective. brain research, 1146:158–171, 2007. peter van roy and seif haridi.concepts, techniques, and models of computer programming. mit press, cambridge, massachusetts, usa, 2004. jason williams and steve young. partially observable markov decision processes for spoken dialog systems.computer speech and language, 21(2):231–422, 2007. mats wiŕen. studies in incremental natural language analysis. phd thesis, link̈oping university, linköping, sweden, 1992. s.j. young, n.h. russell, and j.h.s. thornton. token passing: a conceptual model for connected speech recognition systems. technical report cued/finfeng/tr 38, cambridge university engineering department, 1989. 111 microsoft word submission_dd_discan_v3.docx dialogue & discourse 7(2) (2016) 1-28 doi: 10.5087/dad.2016.201 ©2016 merel scholman, jacqueline evers-vermeul, and ted sanders this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). a step-wise approach to discourse annotation: towards a reliable categorization of coherence relations merel c.j. scholman m.c.j.scholman@coli.uni-saarland.de computational linguistics and phonetics, saarland university campus c7.4, 66123, saarbrücken, germany jacqueline evers-vermeul j.evers@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands ted j.m. sanders t.j.m.sanders@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands editor: barbara di eugenio submitted 11/2014; accepted 12/2015; published online 02/2016 abstract over the last decade, annotating coherence relations has gained increasing interest of the linguistics research community. often, trained linguists are employed for discourse annotation tasks. in this article, we investigate whether non-trained, non-expert annotators are capable of annotating coherence relations. for this goal, substitution and paraphrase tests are introduced that guide annotators during the process, and a systematic, step-wise annotation scheme is proposed. this annotation scheme is based on the cognitive approach to coherence relations (sanders et al., 1992, 1993), which consists of a taxonomy of coherence relations in terms of four cognitive primitives. the reliability of this annotation scheme is tested in an annotation experiment with 40 non-trained, non-expert annotators. the results show that two of the four primitives, polarity and order of the segments, can be applied reliably by non-trained annotators. the other two primitives, basic operation and source of coherence, are more problematic. participants using an explicit instruction with substitution and paraphrase tests show higher agreement on the primitives than participants using an implicit instruction without such tests. we identify categories on which the annotators disagree and propose adaptations to the manual and instructions for future studies. it is concluded that non-trained, non-expert annotators can be employed for discourse annotation, that a step-wise approach to coherence relations based on cognitively plausible principles is a promising method for annotating discourse, and that text-linguistic tests can guide annotators during the annotation process. keywords: discourse annotation, corpora, coherence relations, interrater reliability scholman, evers-vermeul & sanders 2 1 the complexity of discourse annotation the advent of linguistic corpora has had a large impact on the field of linguistics. by gathering and annotating large-scale collections of texts, researchers have gained new possibilities for analyzing language. corpora can be used to, for example, investigate characteristics associated with the use of a language feature, examine the realizations of a particular function of language, characterize a variety of languages, and map occurrences of a feature through entire texts (conrad, 2002). the focus area of corpora has mainly been on lexical, syntactic and semantic characteristics of language. existing corpora often lack annotations on the discourse level (carlson, marcu & okurowski, 2003; versley & gastel, 2012). however, the notion of “discourse”, and more specifically the coherence relations between parts of discourse such as cause-consequence and claim-argument, has become increasingly important in linguistics. this has led to the international tendency in the last decade to create discourse-annotated corpora. leading examples are the penn discourse treebank (prasad et al., 2008), the rhetorical structure theory (rst) treebank (carlson et al., 2003), the segmented discourse representation theory (sdrt; asher & lascarides, 2003) and the potsdam commentary corpus (stede, 2004). while discourse annotation guidelines generally agree on the idea of relations between discourse segments, they differ in other important aspects, such as which features of a relation are analyzed and the types of relations that are distinguished. some proposals present sets of approximately 20 relations, such as the one developed by mann and thompson (1988) and the set of core discourse relations in the iso project (prasad & bunt, 2015), others of only two relations (grosz & sidner, 1986). the pdtb contains a three-tiered hierarchical classification of 43 sense tags (prasad et al., 2008), and the annotation scheme used for the rst treebank distinguishes 78 relations that can be partitioned in 16 classes (carlson et al., 2003). the relational discourse analysis (rda) corpus (moser, moore & glendening, 1996), which is based on rst and the theory proposed by grosz and sidner (1986), distinguishes 29 relations on an intentional and an informational level. hence, it is not clear which and how many categories or classes of relations (for example, contingency, causal, or informational) and end labels (for example, result, volitional cause, and cause-consequence are all labels for causal relations) are needed to adequately describe and distinguish coherence relations. one thing that is clear is that annotation has proven to be a difficult task, which is reflected in low inter-annotator agreement scores (poesio & artstein, 2005). in current proposals, the developers often make use of two solutions to strive for sufficient agreement scores: (1) employing ‘experts’, namely professional linguists or annotators who have received extensive training, or (2) providing the annotators with large manuals. for example, for the rst treebank, professional language analysts with prior experience in other types of data annotation were employed. they also underwent extensive hands-on training (carlson et al., 2003). similarly, two linguistics graduate students were employed for the creation of the rda corpus. these annotators were required to do readings on discourse structure, and they received multiple training sessions (moser & moore, 1996). linguists have more knowledge about language and linguistic phenomena, and are therefore more sensitive to certain linguistic structures. likewise, annotators who have received extensive training have detailed knowledge about the phenomena that they are annotating. often, trained annotators have had the opportunity to discuss their annotations and check them with those of other annotators, which benefits the annotation quality. another solution while striving for sufficient agreement scores is to provide annotators with large manuals that describe the annotation process in great detail. for example, the manual for the pdtb corpus consists of 97 pages (pdtb research group, 2007), and the manual for the rst treebank consists of 87 pages (carlson & marcu, 2001), the latter including more detailed information about segmentation. these manuals contain necessary information for categories of coherence relations in discourse annotation 3 the annotators to be able to analyze texts reliably, but considering the length it can be assumed that annotators need time to work through them. in order to expand the field of discourse annotation and annotate discourse relations on a larger scale, it would be easier if non-trained, non-expert annotators, such as undergraduate students in the humanities, could be employed, and if smaller manuals could be used. working with non-trained, non-expert annotators has the practical advantage that they are easier to come by, and it is therefore also easier to employ a larger number of annotators. using non-trained, non-expert (also referred to as naive) annotators is not new; naive annotators have for example already been employed in anaphoric annotation (poesio & artstein, 2005). the benefit of employing naive annotators has also been recognized in other fields. alonso and mizzaro (2012) describe the advantages of using crowdsourcing for conducting different kinds of relevance evaluations: the outsourcing of tasks to a large group of people makes it possible to conduct information retrieval experiments extremely fast, with good results, and at a low cost. moreover, nowak and rüger (2010) found in a multi-label image annotation experiment that the annotations of non-expert annotators were of a comparable quality to the annotations of experts. although these studies investigated annotations in different fields than discourse analysis, their results do provide insight into the usability and reliability of non-expert annotators for discourse annotation. however, working with less-trained annotators should not affect the quality of the annotations. it therefore needs to be investigated what type of instructions naive annotators need in order to annotate reliably. the annotators in the current study are not as naive as the annotators in crowdsourcing tasks, since the annotators in this study are freshman and senior undergraduate students in the humanities. these students usually have affinity with language, and they were at least trained in some sort of analysis of language, varying from grammatical analyses to literary analyses. we chose this type of annotators because the discourse annotation task is likely to be too complex for people who are not used to work with languages in such a conscious manner. however, the annotators in this study can be considered significantly less expert than linguistics graduate students and professional linguists. the current study sets out to investigate whether non-trained, non-expert annotators can be employed to annotate discourse relations reliably. these types of annotators might benefit from a different annotation process. more specifically, the annotation task could be less complex for them if they could make use of a step-wise approach, in which they annotate characteristics of coherence relations one at a time (for example, deciding whether the relation is causal or additive and whether the relation is subjective or objective). in many of the current annotation proposals, annotators are required to define the coherence relation at hand by assigning an end label to it. this end label is the type of coherence relation, such as a result, claim-argument, contrast, or exception. we believe that the annotation task might become less complex if the process of defining a relation is broken up into several steps. this is explained in more detail in the next section. additionally, it is hypothesized that several (text-)linguistic tests could help non-trained, non-expert annotators during the annotation process. instructions containing tests that make use of connective properties and paraphrase tests could guide annotators during the interpretation of the coherence relation at hand. this is further explained in section 3. 2 a step-wise approach to coherence relations the discourse annotation task might become less complex if annotators can make use of a stepwise annotation approach. in many of the current proposals, annotators are required to define the relations in terms of end labels. for example, in rst annotators can choose the end label ‘cause’, which is used to describe a causal relation such as (1). (1) (in addition,) its machines are typically easier to operate, so customers require less assistance from software. (penn discourse treebank, fragment 1887) scholman, evers-vermeul & sanders 4 although it is not explicitly acknowledged by rst, the classification process can be broken up into several smaller steps: the coherence relation is a causal relation (rather than a temporal or additive relation), the polarity of the relation is positive (rather than negative, such as in contrastive relations), and it is an objective relation (rather than a subjective relation). these types of relations are explained in more detail in section 4. for this relation, the fact that it is causal is quite clear, and it is therefore not expected that its classification leads to many disagreements. however, other types of relations are more difficult. especially for these types of relations, breaking up the classification into several smaller types might be beneficial. this is illustrated with example (2), which is an ‘anti-thesis’ relation according to the rst manual. the end label ‘anti-thesis’ is described as a specific kind of contrast in which one cannot have a positive regard for both of the situations described (carlson & marcu, 2001: 45). (2) although the legality of these sales is still an open question, the disclosure couldn’t be better timed to support the position of export-control hawks in the pentagon and the intelligence community. (penn discourse treebank, fragment 2326) the classification process of this relation can also be broken up into several smaller steps: the coherence relation is causal (rather than temporal or additive), it is negative (rather than positive), and it involves the speaker’s reasoning and is therefore subjective (rather than objective). an annotation scheme that breaks up the classification of such a coherence relation into more and smaller steps might help the annotators during the process. rather than deciding on the end label of the relation at hand, they can decide on separate aspects of relations, which will eventually lead to an end label. a step-wise approach therefore might reduce the need for intensive training and still lead to high agreement between annotators. another advantage of a step-wise approach is that it makes use of similarities as well as differences between coherence relations, and therefore shows links between conceptually related relations. end labels carry the risk of dividing these related relations into separate classes. this is illustrated with examples (3) and (4). (3) operating revenue rose 69% to $8.48 billion from $5.01 billion. but the net interest has jumped 85% to $687.7 million from $371.1 million. (pdtb research group, 2007: 33) (4) (the biotechnology concern said) spanish authorities must still clear the price for the treatment but that it expects to receive such approval by year-end. (pitler & nenkova, 2009: 16) both relations in (3) and (4) are expressed by the connective but, but they fall in different classes according to the pdtb tagset. fragment (3), taken from the pdtb manual, is an example of a typical contrastive relation belonging to the class comparison. the relation in (4) is coded in the pdtb as belonging to the class expansion (pitler & nenkova, 2009), even though it is actually also a contrastive relation. although the pdtb does justice to the fact that these relations differ from each other (for example, (3) is additive and (4) is causal), it disregards the fact that the relations are both negative and contrastive, and that they are therefore conceptually related. by assigning the relations end labels, the conceptual relationship between the two coherence relations is not acknowledged. in contrast, an approach that classifies relations based on a combination of characteristics does account for this: such an approach does not only show differences between relations, it also shows similarities between different relations. in order to create an annotation scheme in which coherence relations are broken up into several characteristics, a classification of coherence relations is necessary that supports this stepcategories of coherence relations in discourse annotation 5 wise process. the cognitive approach to coherence relations (ccr), proposed by sanders, spooren and noordman (1992, 1993) is exactly such a theory in which the coherence relations are defined by their characteristics. the theory is built on the assumption that coherence relations are cognitive, psychological constructs that language users make use of when interpreting text, and not just descriptive constructs that are created by linguists. sanders et al. (1992, 1993) believe that understanding discourse means constructing a coherent representation of that discourse. since coherence relations play a crucial role in this representation, different relations over the same discourse will result in different representations. in line with hobbs’ (1979, 1985) and kehler’s (2002) work on coherence relations as cognitive elements of the discourse representation, sanders et al. (1992, 1993) set out to describe the link between the structure of a discourse as a linguistic object and its cognitive representation. sanders et al. (1992, 1993) distinguish four cognitive primitives that they claim to be relevant for every coherence relation. what distinguishes these primitives from other, possibly relevant characteristics or primitives is that they all concern the additional meaning provided by the relations, namely they concern the informational surplus that the coherence relation adds to the interpretation of the discourse segments in isolation. the four cognitive primitives are: polarity (relations are positive or negative), basic operation (causal or additive), source of coherence (objective or subjective), and order of the segments (basic or non-basic order).1 a detailed explanation of the primitives is given in section 5. besides the fact that ccr allows for a step-wise approach, there is another argument for the applicability of ccr for discourse annotation: several studies have shown that these basic primitives and the categories they define are cognitively relevant. for example, acquisition studies have shown that positive relations are acquired before negative relations (bloom et al., 1980, spooren & sanders, 2008), and that additive relations are acquired before causal relations (bloom et al., 1980; evers-vermeul & sanders, 2009). processing studies show that once causal relations are acquired, they are processed faster and generate better recall compared to additive and temporal relations (noordman & vonk, 1998; sanders & noordman, 2000). furthermore, objective causal relations are processed faster than subjective causal relations (canestrelli, mak & sanders, 2013; traxler, bybee & pickering, 1997; traxler, sanford, aked & moxey, 1997). and finally, studies have shown that coherence relations with a basic order of the segments are easier to process than coherence relations with a non-basic order (noordman & de blijzer, 2000; noordman & vonk, 1998). these studies indicate that the primitives and their categories affect language acquisition and processing, and are therefore cognitively relevant. the four primitives are hypothesized to be useful for discourse annotation because they allow for a step-wise annotation process. they can be visualized in a flowchart, leading to a compressed annotation scheme that can be used to make systematical decisions. this can be beneficial to trained annotators, but perhaps non-trained, non-expert annotators are also capable of applying the cognitive categories method in discourse annotation. although there is evidence for the relevance of the basic primitives and their categories, it has not been investigated how reliably they can be used to annotate coherence relations in everyday corpora of language use. the present study aims to explore this in an annotation experiment for which a large number of naive annotators analyze a sample corpus. 3 instructions guiding the annotation process the aim of the current study is to investigate whether non-trained, non-expert annotators can annotate coherence relations reliably. it also investigates whether the reliability increases when these annotators can make use of linguistic tests during the annotation process. there is a lot of 1 originally, the terms objective and subjective were defined in the literature as semantic and pragmatic, respectively (see pander maat & sanders, 2000 for a discussion of this transition). scholman, evers-vermeul & sanders 6 variation in the types of instructions that manuals of different proposals contain. for example, the manual for the rda corpus contains an instruction for a diagnostic test, for which the annotator has to imagine the context in which the relation occurs. the manual also explicitly mentions that annotators should not use discourse cues as a basis for deciding what relation occurs between the two segments (moser, moore & glendening, 1996). in contrast, the pdtb manual encourages annotators to take the discourse cue into account, and supplies the annotators with information on which relations a certain connective can signal (pdtb research group, 2007). the rst manual also mentions several typical discourse cues that often occur in certain types of relations (carlson & marcu, 2001). however, the pdtb and rst manuals do not explicitly provide the annotators with systematic tests that can be used as a diagnostic tool during the process. in one of the conditions in the current study, two types of tests are used to guide the annotator during the annotation process, namely a substitution test and a paraphrase test. both tests will be explained consecutively. the substitution test is based on characteristics of connectives. according to ccr, the cognitive primitives and their categories can be distinguished by the connectives they co-occur with. in other words, certain connectives signal certain types of relations, and readers or listeners can therefore use these connectives as processing instructions on how to relate the incoming information to the previous discourse segment. the idea of connectives as processing instructions was already suggested several decades ago by ducrot (1980) and lang (1984). ever since halliday’s and hasan’s (1976) seminal work, it has been argued that connectives differ in the type of relation they signal. for instance, because signals a positive causal relation; meanwhile signals a positive temporal relation; and but signals a negative relation (knott & dale, 1994). restrictions on the use of connectives can also be more subtle, because they can also hold within the same class of relations (pander maat & sanders, 2000). for example, within the class of negative relations, the connectives although and whereas signal different types, namely negative causal and negative non-causal relations, respectively. given that connectives indicate how two segments are related to each other, they can be used by annotators to guide them while analyzing the relation at hand. this can be done by employing substitution tests, which is a method for testing semantic intuitions (knott & dale, 1994; knott & sanders, 1998; pander maat & sanders, 2000). in a substitution test, the original connective is (mentally) substituted by another connective known for signaling a certain type of relation, while the meaning of the original relation is preserved. if there is no original connective present, the proposed connective is merely mentally inserted. for example, an annotator can ask himself for any given relation: can these two segments be connected by but? or by because? substitution tests therefore rely on the properties of the connectives, such as the polarity and degree of subjectivity they signal (pander maat & sanders, 2000). if two connectives are inter-substitutable in a coherence relation, they should be classified in the same category of coherence relations (knott & dale, 1994). substitution tests are not the only type of tests that annotators can apply; paraphrase tests can also facilitate the interpretation process (sanders, 1997; knott & sanders, 1998). in a paraphrase test, the annotator is instructed to choose one of two given paraphrases that best suits the coherence relation expressed in the text. the paraphrases both restate the two segments of the relation to give the meaning of the relation in another form. for example, in order to determine the order of the segments, the annotator can ask himself for a given objective causal relation: can the two segments be paraphrased as ‘segment 1 presents the cause, and segment 2 presents the consequence’ or ‘segment 1 present the consequence, and segment 2 presents the cause’? substitution tests and paraphrase tests have been used widely in studies on connectives in language use in various languages and across genres and media (see among others, degand, 2001; degand & pander maat, 2003; evers-vermeul, 2005; knott & dale, 1994; knott & sanders, 1998; li, evers-vermeul & sanders, 2013; pander maat & degand, 2001; pit, 2007; sanders, 1997; sanders & spooren, 2015; stukker & sanders, 2012; stukker, sanders & verhagen, 2008; categories of coherence relations in discourse annotation 7 zufferey, 2012), as well as in studies of connective acquisition (evers-vermeul & sanders, 2009, 2011; spooren & sanders, 2008). in all these studies, the tests have been applied successfully by expert annotators. whether such tests will also guide non-expert, non-trained annotators while analyzing real-life texts, is not clear yet. in the remainder of this paper, an annotation experiment is presented that set out to investigate this. 4 method in this experiment, 40 non-expert, non-trained subjects were asked to annotate a sample corpus making use of a step-wise approach based on ccr. the ccr approach allows for paraphrase and substitution tests to be used to determine the correct value for a primitive. these tests facilitate the decision making process, and are therefore expected to benefit the reliability of the method. in order to test whether this is true, two versions of the instruction were created: an implicit instruction and an explicit instruction. the implicit instruction relies only on the annotator’s knowledge of the categories obtained from the manual. the explicit instruction relies on this knowledge, as well as on paraphrase and substitution tests. this is explained in more detail below, but first the four primitives and their categories are explained. 4.1 material the material for the experiment consisted of a manual, a flowchart, two versions of an instruction and a sample corpus of 36 coherence relations. 4.1.1 manual and flowchart each subject received a nine-page manual for the cognitive approach to coherence relations, and a flowchart presenting all annotation choices. participants received no additional training besides this manual. in the manual, discourse annotation and segmentation is explained, followed by an explanation of every value of each primitive. after this explanation, examples are given for every possible combination, thereby illustrating the categories. a description of the cognitive primitives and their categories, similar to the description given in the manual, can be found below. polarity the first primitive in the taxonomy is polarity. this refers to the positive or negative character of a segment. a relation is positive if the propositions p and q, expressed in the two discourse segments s1 and s2, are linked directly, without a negation of one of these propositions. a relation with a positive polarity is typically connected by connectives such as and or because. (5) is an example of a relation with a positive polarity. 2 the brackets indicate where the first segment (s1) and the second segment (s2) start and end. (5) [the stocks can decrease tremendously in value]s1 and [thereby result in a loss for the investor.]s2 in example (5), the second segment has a direct link to the first segment. the second segment is an expected consequence and there is no negation of the entire segment present. a relation is negative if the negative counterpart of either p or q functions in the relation. a relation with a negative polarity is typically connected by connectives such as but and although, as is illustrated in (6). (6) [the biofuel is more expensive to produce,]s1 but [by reducing the excise-tax the government makes it possible to sell the fuel for the same price.]s2 2 all examples are (translations of fragments) taken from the dutch discan corpus. scholman, evers-vermeul & sanders 8 in (6), a logical positive second segment would be that the biofuel costs more, as a consequence of the higher production costs. however, the second segment presents a denial of this expectation: the fuel is not sold at a higher price due to a reduced excise-tax. the second segment expresses not-q, that is, the negation of the consequent of the relation. this negation causes the relation to have a negative polarity. basic operation the second primitive that sanders et al. (1992, 1993) distinguish is the basic operation. this primitive concerns the operation that has to be carried out on the two discourse segments. three types of basic operation underlie coherence relations: the causal, additive and temporal basic operation.3 these operations were proposed because they justify the basic intuition that discourse segments are either strongly connected (causality) or weakly connected (addition and temporality). for negative relations, the additive and temporal relations have been taken together as ‘non-causal’ relations. a relation is causal if an implicit relation (p " q) can be deduced between the two discourse segments, as in (7). (7) [the athletics union was forced to emigrate to belgium,]s1 because [there was no accommodation available in the netherlands.]s2 in (7), the consequence is presented in s1, and the cause in s2: a lack of accommodation has led to the emigration of the athletics union. the class of causal relations can be further divided in non-conditional (causal) and conditional relations. an example of a conditional causal relation can be seen in (8). (8) if [you don’t answer,] s1 [i will arrest you.]s2   in (8), the speaker confronts the listener with a condition. if the listener does not answer, there will be a consequence: he will be arrested. a relation is additive if the segments are connected by a logical conjunction (p & q), as in (9). (9) [the quality of this fuel with bio component is completely similar to shell’s regular euro 95] s1 and [the price at the pump is the same as well.]s2 the relation in (9) consists of two segments that both describe a fact about fuel with a bio component. the segments are in an equal relation to each other: there is no cause, consequence, condition or contrast present. a relation is temporal if the two segments are linked by their occurrence in the world or worlds being described. temporal relations have an additive nature, but differ in that the segments contain two events that are ordered in time. (10) is an example of a temporal relation. (10) [next thursday a second meeting will follow.]s1 [the unsatisfied ret-employees will decide after this meeting if they deem it necessary to continue protesting.]s2 3 the original proposal did not distinguish temporality as a basic operation, but included temporal relations in the category of positive additive relations. this value was now added at the basic level to improve descriptive adequacy, and because temporal relations have been shown to be relevant in the order of acquisition (evers-vermeul & sanders, 2009). still, there is some discussion on how basic temporality is. categories of coherence relations in discourse annotation 9 example (10) consists of two sequential (future) events. the events have an order in time: s2 follows s1. source of coherence the third primitive is the source of coherence, which can be divided into two categories: objective and subjective. a relation is objective if the discourse segments are connected by their propositional content. in other words, both segments describe situations in the real world or the world being described, as in (11). the speaker merely reports these facts, and is not actively involved in the construction of the relation. (11) [the plaintiff received his car,]s1 because [the advertisement was formulated ambiguously.]s2 relations are subjective if speakers or authors are actively engaged in the construction of these relations, either because they are reasoning, or because they perform a speech act in one or both segments. subjective relations, such as (12), usually express the speaker’s opinion, argument, claim or conclusion. (12) [drugs destroy people’s lives,]s1 so [drugs have to be battled judicially.]s2 in (12), the statement in the first segment is not the cause for the second segment, but a reason that is given to support the claim in the second segment. order the fourth primitive is the order of the segments. two segments in a causal relation can be connected in a basic or a non-basic order. the order of the segments is not applicable for additive relations, as they are logically symmetrical. a relation with a basic order has an antecedent as s1, followed by a consequent in s2, as in (13). the antecedent is the cause or the argument, and the consequent is the consequence or the claim. in a relation with a non-basic order, such as (14), the consequent precedes the antecedent. (13) sometimes children tease me. [but i don’t reply,]s1 that’s why [they don’t do it anymore.]s2 (14) [universities supposedly cancel subscriptions to scientific journals more often]s1 because [there is more information available through the internet.]s2 flowchart the four primitives can be represented in a flowchart, which can be used for annotating discourse and allows for a systematical, step-wise decision-making process. the entire flowchart can be seen in figure 1. this flowchart will be explained step by step. starting with a discourse relation, the first step in the annotation process is determining the polarity. the category of negatives differs greatly from that of positives; therefore this step is the first one in the flowchart. second, the basic operation has to be decided upon. for positive relations, this can be causal, causal-conditional, additive or temporal. for negative relations, this basic operation can be divided into the categories causal and non-causal (any negative relation that is not causal). this step is taken as the second step because the remaining two steps are not applicable to every relation. the third step is determining the source of coherence, which consists of the same two categories for all relations (objective and subjective), except for temporal and non-causal scholman, evers-vermeul & sanders 10 relations. because temporal relations are made up of a description of two events that are ordered in time, this type of relation is always objective. the final step concerns the order, which can be basic or non-basic. the order is not applicable for additive and non-causal relations, since the two segments in such relations are logically symmetric, and for temporal relations in which the segments describe events that occur simultaneously. figure 1. flowchart of the step-wise annotation instruction. 4.1.2 instructions annotators analyzed fragments using an instruction. two experimental conditions were created in this study: one group annotated according to an implicit instruction (see appendix a), and one according to an explicit instruction (see appendix b), which included text-linguistic tests. each possible answer for the steps in the instruction was preceded by a box, which participants could tick if they thought that this value was the correct answer. the implicit instruction consists of four steps; one step for each cognitive primitive. the instruction is straightforward and relies on the annotator’s knowledge of the categories. the annotator is instructed to determine the value and is reminded of any anomalies. take, for example, step 3 of the implicit instruction (originally, this instruction is in dutch): box 1. fragment of the implicit instruction the explicit instruction consists of five multileveled steps and contains two types of tests: paraphrase and substitution tests (see section 3). decisions for source of coherence and order are based on knowledge of the categories and paraphrase tests (sanders, 1997; knott & sanders, 1998). an example of paraphrase tests for order can be seen in box 2. 3. determine the source of coherence: is the relation objective or subjective? this does not apply to temporal or non-causal negative relations, because they do not differ in source of coherence. therefore, for these relations tick not applicable. objective subjective not applicable categories of coherence relations in discourse annotation 11 box 2. fragment of the explicit instruction in step 2a in box 2, the annotator is given two paraphrases that can be used to determine the source of coherence of a relation. a paraphrase test was also used to determine the order of subjective relations; in that case claim and reason were used instead of cause and consequence. in a substitution test, the annotator is first instructed to mentally take out the connective (if present in the relation), and then to replace it with different connectives. substitution tests are used because they rely on the connective properties. in the current study, substitution tests are used in the explicit instruction for determining the polarity and the basic operation. box 3 provides an example of a substitution test. box 3. fragment of the explicit instruction in step 1, the explicit instruction guides the annotator in his choice for polarity. in this case, the annotator is instructed to substitute the original connective with the connective but. this type of substitution test was also used for causal relations (“can you use because to connect the segments?”), conditional relations (“can you use if to connect the segments?”), additive relations (“can you use and to connect the segments?”) and temporal relations (“can you use then to connect the segments?”). 4.1.3 sample corpus the sample corpus consists of 36 dutch coherence relations with context, taken from the discan corpus. the discan corpus is a dutch corpus with annotated discourse relations, which was developed using an annotation scheme based on ccr (sanders, vis & broeder, 2012). this corpus currently consists of approximately 1500 fragments and includes seven subcorpora used in previous corpus-based research (see, for example, degand, 2001; sanders & spooren, 2009; stukker, 2010). these subcorpora mainly consist of newspaper articles, but also contain fragments from novels, spoken discourse, and chat fragments. the annotations that are included in the discan corpus were taken from the original annotations of the seven subcorpora and supplemented if any primitives were missing. currently, discan only contains explicit relations, although several additional subcorpora containing implicit relations have been prepared for inclusion in the discan corpus. 2a can you paraphrase the relation between s1 and s2 as in option a or rather option b below? a. the situation / fact / event in one segment causes the situation / fact / event in the other segment. or b. one segment describes the reason for the claim or conclusion given in the other segment. paraphrase a, then the source of coherence is objective. proceed to question 2b. paraphrase b, then the source of coherence is subjective. proceed to question 2c. 1. can you use but to connect the segments? yes, then the polarity is negative. proceed to 1a. no, then the polarity is positive. proceed to 2. scholman, evers-vermeul & sanders 12 for the current experiment, both spoken and written texts are incorporated in the corpus, as well as chat fragments. the fragments were included in their original formulation, to ensure that the task resembles a real-life annotation task. the fragments were presented with the segment boundaries indicated. this was done to limit effects of segmentation. 4.2 annotators 40 non-trained, non-expert subjects took part in this experiment and were paid for their participation. 20 subjects were freshman students and 20 subjects were senior students. all participants were students of the faculty of humanities at utrecht university. none had experience with discourse analysis. to ensure that participants in this experiment had an affinity with language and text, participants were recruited from undergraduate studies in modern languages, linguistics and communication sciences. these participants were expected to have basic meta-linguistic skills. a comparison is made between freshman and senior undergraduate students in order to investigate whether the amount of formal education in a field in humanities had an influence on the extent to which annotators can apply a classification scheme to coherence relations. 4.3 procedure all materials were presented on paper. the annotators were asked to meticulously read the manual and ask questions if anything was unclear. they were also ensured that they could consult the manual and ask questions throughout the entire experiment. questions could only concern the interpretation of a value; not the interpretation of a fragment. after the participants had read the manual and instruction, they could start annotating the sample corpus. each fragment of the sample corpus was followed by the instructions, in which annotators could tick their choices. they were instructed to follow the steps presented in the instruction. they were allowed to annotate at their own pace, take breaks and divide the workload into two sessions. all coders annotated independently. participants took approximately an hour and a half to read the manual and annotate all fragments. 4.4 processing the data consistency is a challenge for each discourse annotation project. although inter-annotator agreement is an important issue in the field of discourse analysis, the reliability and validity of coding is still a concern (spooren & degand, 2010). to deal with this problem, different statistics were calculated in the current study, namely the percentages of agreement, kappa (κ) scores, and recall, precision and f-scores. percentages of agreement are often reported in similar studies (artstein & poesio, 2008). it is the simplest measure of agreement, but it does not correct for chance agreement. this measure is therefore biased in favor of dimensions with a small number of categories (scott, 1955). kappa scores do correct for chance agreement, and therefore show a less biased picture of the data (carletta, 1996). when there is total agreement, κ is one. when there is no agreement besides chance agreement, κ is zero. concerning the acceptability of a kappa score, there is no clear definition of what passes as an acceptable agreement score (artstein & poesio, 2008). for the current study, it was decided to follow the conventions proposed by krippendorf (1980: 147), as reported by carletta (1996: 252): a category with almost perfect agreement (κ > 0.81) indicates a reliable method; a category with substantial agreement (0.61 < κ < 0.81) allows for tentative conclusions to be drawn; and everything below substantial agreement (κ < 0.61) indicates that the method is not reliable enough. finally, recall, precision, and f-scores are included to calculate the agreement with the original annotations per value of each primitive. these measures calculate the number of true and false positives and negatives. to illustrate this, consider table 1 (which is based on ting, 2010). categories of coherence relations in discourse annotation 13 assigned class by non-trained annotators positive negative assigned class by expert annotator positive true positive (tp) false negative (fn) negative false positive (fp) true negative (tn) table 1. the outcomes of classification into positive and negative classes in table 1, the values positive and negative can be considered to represent the actual values positive and negative for the primitive polarity, for example. true positives and true negatives are correct answers; namely when the subject agrees with the expert annotator. a false positive occurs when the subject assigns a positive polarity to an item that actually has a negative polarity. similarly, a false negative occurs when a subject assigns a negative polarity to a coherence relation that actually has a positive polarity. based on these outcomes, precision and recall can be calculated as follows: recall = true positives / total number of actual positives assigned by the expert annotator = tp / (tp + fn) precision = true positives / total number of positives assigned by the subject = tp / (tp + fp) in other words, recall represents the number of times the annotators assigned a value correctly, out of all the times that the expert annotators had assigned the value. precision shows the number of times the annotators assigned a value correctly, divided by all the times they assigned the value. instead of two measures, these scores are often combined to provide a single, harmonic measure of agreement called the f-measure (brants, 2000): f-measure = (2 * recall * precision) / (recall + precision) all three scores are reported in the current study. it is not determined what score is acceptable or unacceptable; rather, these scores are used to identify problems with specific categories of primitives, which are further discussed in section 5.4. 5 results agreement statistics were calculated for each primitive separately. first, the kappa statistics for agreement between annotators are presented. then the agreement with the original annotations in the discan corpus is shown in kappa statistics, followed by the agreement per type of instruction in percentages. the section is concluded with a more detailed analysis of agreement on separate categories in recall, precision and f-scores. 5.1 agreement between annotators table 2 shows agreement between annotators for each condition separately. scholman, evers-vermeul & sanders 14 primitive overall first year third year implicit instruction explicit instruction implicit instruction explicit instruction polarity .73 .68 .84 .64 .78 basic operation .42 .33 .47 .50 .47 source of coherence .31 .28 .27 .39 .32 order .47 .34 .49 .49 .66 table 2. kappa statistics for each primitive in general and per condition. table 2 shows that the non-expert, non-trained annotators agree substantially on the categories of polarity (κ = .73). agreement is moderate for the primitives basic operation (κ = .42) and order (κ = .47). agreement on source of coherence is fair (κ = .31). hence, of the four primitives, polarity yields the highest agreement and source of coherence is least agreed on. these results are in line with earlier results (sanders et al., 1992, 1993). when analyzed per year, most conditions show a kappa similar to the overall kappa scores. agreement on polarity is substantial in most conditions (.64 < κ < .78), but almost perfect in the first year explicit condition (κ = .84). agreement on basic operation is moderate in most conditions (.45 < κ < .50), but it is fair in the first year implicit condition (κ = .33). for source of coherence, agreement is fair in all conditions (.27 < κ < .39). agreement for the primitive order is fair for the first year implicit condition (κ = .34) and substantial in the third year explicit condition (κ = .66), whereas it’s moderate in the other conditions (κ = .49). note that it is possible that annotators show agreement on categories that are not the correct ones according to the original annotations. in other words, they can agree on the wrong categories. the kappa statistics for agreement with original annotations in section 5.2 will show whether participants annotated the correct categories. this will provide more insight into the quality of the instructions: how well do the instructions convey the information that annotators are supposed to know? 5.2 agreement with original annotations table 3 shows the agreement with the original annotations for each condition separately. primitive overall first year third year implicit instruction explicit instruction implicit instruction explicit instruction polarity .86 .84 .91 .79 .88 basic operation .49 .41 .52 .48 .52 source of coherence .31 .31 .25 .36 .31 order .61 .50 .62 .59 .69 table 3. agreement with original annotations in kappa statistics overall and per condition. the annotators showed almost perfect agreement with the original annotations on the primitive polarity (κ = .86). agreement with order was substantial (κ = .61). agreement of the annotators with the original annotations on basic operation was moderate (κ = .49) and agreement on source of coherence was fair (κ = .31). again, the results show that polarity yields highest agreement with the original annotations and source of coherence the lowest. when the conditions are analyzed separately, most scores of the primitives remain in the same range. all conditions show moderate agreement for the primitive basic operation (.41 < κ < .52) and fair agreement for the primitive source of coherence (.25 < κ < .36). for polarity, most conditions show almost perfect agreement (.84 < κ < 91), except for the third year implicit categories of coherence relations in discourse annotation 15 condition, which shows substantial agreement (κ = .79). for the primitive order, the two explicit conditions show substantial agreement (.62 < κ < .69), but only moderate agreement is found in the first year implicit condition (κ = .50) and third year implicit condition (κ = .59). to determine whether these differences in agreement with original annotations between conditions were significant, a univariate anova was performed. results indicated a significant main effect of type of instruction on agreement with the original annotations (f (1, 5717) = 12.28; p < .001). a significant main effect was also found for primitive (f (3, 5717) = 228.00; p < .001). no main effect of undergraduate year was found (f (1, 5717) = 3.76; p = .052), nor any interaction effects of undergraduate year and type of instruction or primitive. therefore, the distinction between first and third year students was not taken into account in further analyses. 5.3 agreement per type of instruction table 4 shows the percentages of agreement with the original annotations per type of instruction. primitive implicit instruction explicit instruction polarity 94 (.24) 96 (.19) basic operation 63 (.48) 71 (.46) source of coherence 57 (.50) 54 (.50) order 70 (.46) 78 (.42) table 4. percentages of agreement (and standard deviations) with the original annotations per type of instruction. an interaction effect was found for type of instruction and primitive (f (3, 5717) = 5.50; p = .001). participants using the explicit instruction showed more agreement with the original annotations on certain primitives than participants using the implicit instruction. further analyses showed a significant difference in agreement with original annotations between the implicit and explicit instructions for the primitives polarity (t (1356.83) = -2.19; p = .03), basic operation (t (1427.10) = -3.33; p = .001) and order (t (1418.45) = 3.32; p = .001). the annotators using the explicit instruction showed more agreement with original annotations for these three primitives than annotators using the implicit instruction. there was no significant difference in agreement with original annotations for the primitive source of coherence between participants using the implicit instruction and participants using the explicit instruction (t (1425.57) = 1.10; p = .27). 5.4 agreement on separate values per primitive in order to examine which values of a primitive were annotated better or worse than others, recall, precision and f-scores were calculated per value, primitive and instruction. as described in section 4.4, recall represents the number of times the annotators assigned a value correctly, divided by all the times that the expert annotator assigned the value. precision represents the number of times the annotators assigned a value correctly, divided by all the times they assigned the value (both correctly and incorrectly). the f-score is the harmonic mean of both. together, these three scores will provide more insight into which categories cause annotation problems, and which categories are often confused with each other. table 5 shows the recall, precision and fscores for the primitive polarity. according to the original annotations, there were 28 positive and eight negative relations in the sample corpus. scholman, evers-vermeul & sanders 16 value implicit instruction explicit instruction recall precision f-score recall precision f-score positive .95 .97 .96 .99 .97 .98 negative .89 .83 .85 .89 .95 .92 table 5. recall, precision and f-scores for the primitive polarity. the f-scores reported in table 5 show that the value negative was annotated correctly more often in the explicit instruction than in the implicit instruction. hence, in the implicit condition, subjects annotated positive relations as having a negative polarity more often. this is reflected in the precision and f-scores, and suggests that the substitution test used for the explicit instruction led to higher agreement on the category negative for polarity. overall, the value polarity was annotated well. this was already indicated by the percentages of agreement in table 4. table 6 shows the recall, precision and f-scores for the primitive basic operation.4 value implicit instruction explicit instruction recall precision f-score recall precision f-score causal .91 .63 .75 .92 .79 .85 conditional .29 .75 .42 .52 .75 .61 additive .47 .73 .58 .48 .76 .59 temporal .61 .43 .51 .55 .24 .34 non-causal .25 .65 .36 .41 .83 .55 table 6. recall, precision and f-scores for the primitive basic operation. as was indicated in section 5.3, the substitution and paraphrase tests for basic operation were helpful for the participants in the explicit condition. this finding is confirmed by the f-scores on the categories causal, conditional, and non-causal, which are higher in the explicit than in the implicit condition. overall, however, the values temporal and non-causal are problematic. the value temporal is often mistaken for the value additive, especially in the explicit condition. this leads to low precision scores for the value temporal, and low recall scores for the value additive. these results indicate that the substitution test for the temporal relations (“can you use then or when to connect the segments?”) was misleading: annotators did not think they could use then, leading them to the next substitution test: “can you use and to connect the segments?” table 6 also shows that the value non-causal was used for relations that were actually not noncausal, especially in the implicit condition. this led to lower recall scores for the value noncausal, and lower precision scores for the value causal. it should be noted here that these outcomes were based on only two coherence relations with a non-causal basic operation. finally, it appears that the value conditional was also applied to causal relations when it should not have been, especially in the implicit condition. again, this should be interpreted with care, since there was only one fragment with a conditional relation in the sample corpus. table 7 presents the scores for the primitive source of coherence.5 4 the sample contained 22 causal, one conditional, six additive, five temporal, and two non-causal relations. 5 the sample contained twelve objective and seventeen subjective relations, and seven relations to which source of coherence did not apply. categories of coherence relations in discourse annotation 17 value implicit instruction explicit instruction recall precision f-score recall precision f-score objective .47 .67 .55 .41 .57 .48 subjective .79 .54 .64 .72 .55 .62 not applicable .44 .46 .45 .48 .44 .46 table 7. recall, precision and f-scores for the primitive source of coherence. table 7 indicates that every value of source of coherence is problematic, but especially the values objective and not applicable. the subjects often annotated subjective relations as objective relations in both the explicit and implicit conditions, as shown by the low precision score for subjective relations, and low recall score for objective relations. also, the subjects often annotated relations for which the source of coherence does not apply as objective relations. this is reflected in the low precision score for the not applicable relations, and the low recall score for the objective relations. these results may be attributed to the step-wise aspect of the approach, which will be discussed in more detail after the recall, precision and f-scores for the primitive order are presented.6 value implicit instruction explicit instruction recall precision f-score recall precision f-score basic .55 .66 .60 .71 .55 .62 non-basic .80 .61 .69 .91 .78 .84 not applicable .76 .80 .78 .74 .93 .82 table 8. recall, precision and f-scores for the primitive order of the segments. for the primitive order, the subjects in the implicit condition often coded relations with a nonbasic order as relations with a basic order. this is reflected in the low recall score for the basic order, and the lower precision score for the non-basic order. apparently, the substitution and paraphrase tests were helpful in this area, as the participants in the explicit condition obtained higher recall scores for the basic as well as the non-basic order, and higher precision scores for the non-basic category. however, in the explicit condition, the subjects still coded basic relations as having an order that is not applicable, which is reflected in their lower precision score for the basic order. similar to the source of coherence, the low scores for the values of order could be due to the step-wise aspect of the taxonomy. recall that the annotators were instructed to first annotate the basic operation, and then the source of coherence and the order. if they made a mistake in the basic operation, for example they annotated additive relations as temporal relations, or vice versa, they would automatically annotate the wrong category of source of coherence and order. this is because temporal relations do not differ in their source of coherence, whereas additive relations do; and temporal relations can have different segment orders, whereas additive relations cannot. since the results indicated that the subjects did indeed often annotate temporal relations as additive relations, it is likely that this influenced the results. in order to determine the influence of the step-wise approach on the results, the percentages of correct annotations based on the correct annotations of the previous step was calculated. in other words, the percentages of correct annotations were calculated only for those relations in which the previous step was also annotated correctly. table 9 shows the results. 6 according to the original annotations, there were ten basic and eleven non-basic relations, and fifteen fragments for which order of the segments was not applicable. scholman, evers-vermeul & sanders 18 primitive implicit instruction explicit instruction 1. polarity 94 (n=720) 96 (n=720) 2. basic operation (based on correct annotation of step 1) 65 (n=673) 72 (n=693) 3. source of coherence (based on correct annotation of steps 1 and 2) 73 (n=440) 66 (n=500) 4. order (based on correct annotation of steps 1-3) 80 (n=322) 91 (n=331) correct annotation of all steps 36 (n=720) 42 (n=720) table 9. percentages (and actual numbers) of correct annotations for each step, based on correct annotations of previous steps (maximum n = 20 annotators per type of instruction×36 relations = 720 annotations). the results in table 9 indicate that the step-wise nature of this approach has a large influence. first, it results in a relatively low number of relations that were annotated correctly for all four primitives: 36% of the implicit annotations and 42% of the explicit annotations were entirely correct. second, the step-wise approach had a negative impact on the reliability of certain primitives. more specifically, the results indicate that the primitive source of coherence might not be as problematic as the previous results suggested. looking at the annotations of source of coherence irrespective of the correctness of previous annotation steps, the subjects annotate this primitive correct in 57% of the relations in the implicit condition, and 54% of the relations in the explicit condition, as shown in table 4. but when the relations in which the basic operation was incorrectly annotated are excluded, the percentages of correct annotations rise to 73% in the implicit condition and 66% in the explicit condition. in other words, 73% of the relations that were annotated correctly for their basic operation were also annotated correctly for their source of coherence in the implicit condition. these results indicate that the annotations for the different primitives are related: if the subjects annotate the basic operation incorrectly, they also annotate the source of coherence incorrectly more often than when they annotate the basic operation correctly. a similar conclusion can be drawn for the primitive order: without taking the step-wise process into account, the subjects annotated the order correctly in 70% of the relation in the implicit condition, and 78% of the relations in the explicit condition, as shown in table 4. however, when the step-wise approach is taken into account, the percentages of agreement rise to 80% in the implicit condition, and 91% in the explicit condition. 6 discussion and conclusion the research question that was formulated for this study was: are non-expert, non-trained annotators capable of annotating coherence relations by using a step-wise approach that is based on cognitively plausible primitives, and do substitution and paraphrase tests improve the quality of their annotations? in the following subsections we will address the merits and drawbacks of a step-wise approach (section 6.1), the usefulness of substitution and paraphrase tests (6.2), and the generalizability of our approach to other annotation systems (6.3). 6.1 step-wise approach at a first glance, when looking at the percentages of correct annotations of all steps taken together, the step-wise approach may not seem very promising. if the naive annotators in our study would have had to come up with an end label on the basis of their choices on all four primitives, the participants in the implicit condition would have chosen a correct end-label in 36% of the relations, and the participants in the explicit condition in 42% of the relations (see table 9). these scores are lower than the results reported in previous studies with expert annotators, who received intensive training before and during the annotation process. a study for categories of coherence relations in discourse annotation 19 agreement using rst showed a kappa ranging from .6 – 1.0 (carlson et al., 2003) and a study using the penn discourse treebank annotation scheme resulted in percentages of agreement ranging from 59.6% – 95.7% (miltsakaki et al., 2004). al-saif and markert (2010) report a kappa value of .57 for their pdtb-inspired scheme for arabic and a study on the dutch rst corpus resulted in a kappa score of .57 as well (van der vliet et al., 2011). however, if we look at the outcomes of the individual primitives, the step-wise approach does show potential, as the percentages of agreement ranged from 54% (for source of coherence) to 96% (for polarity), and the kappa statistic for the explicit condition averaged over the four primitives is .59. these results are comparable to the scores in the aforementioned studies with expert annotators. however, it should be noted that this averaged kappa does not include the correct sequence of annotations for all four primitives. still, given that the annotators were not trained in discourse annotation and only received a nine-page manual and instructions varying from one to three pages, this amount of agreement is promising. for polarity, the reliability was satisfactory: the annotators frequently agreed on this value, with each other as well as with the original annotations. on the basis of the agreement with the original annotations, we can also draw tentative conclusions for the primitive order of the segments, although the agreement among the forty annotators was moderate. for the other two primitives, basic operation and source of coherence, there is room for improvement, as we did not find adequate agreement among annotators nor between the naive annotators and the original annotations. we will discuss these two primitives in turn. the primitive basic operation only yielded moderate agreement. in particular, the results showed low agreement with original annotations on the categories temporal and non-causal. the category temporal was often mistaken for the category additive, especially in the explicit condition. this indicates that the substitution test for temporal relations (“can you use then/when to connect the segments?”) was misleading (see section 6.2 for a more extensive discussion of this issue). however, since the agreement on the category temporal in the implicit condition was also not acceptable, it can be concluded that the manual did not provide enough information to clarify this concept. after completing the experiment, several subjects declared that the distinction between additive and temporal was not entirely clear. if more annotators experienced this, they might have employed different definitions for basic operation and the categories additive and temporal, leading to different annotations. a similar study with clearer instructions and different substitution tests could shed light on the specific issue of temporal relations. regarding non-causal relations, the results showed that the annotators frequently analyzed causal relations as non-causal or conditional, especially in the implicit condition. the confusion with non-causal relations implies that annotators especially ran into problems with negative causal relations, since the non-causal category only occurs in relations with a negative polarity. negative causal relations are known to be more complex than, for example, positive causal and negative additive relations (evers-vermeul & sanders, 2009). this suggests that naive annotators might need some additional instruction on the interpretation of causal relations with a negative polarity. as the discussion in section 6.2 will show, substitution tests can be part of this additional instruction, as these reduce the number of mistakes with the (negative) causal category. the second primitive for which agreement was not high enough was the source of coherence: the recall, precision, and f-scores showed that every category of this primitive seems to be problematic. however, an investigation of the influence of the step-wise nature of the approach indicated that the relatively low reliability of the primitive source of coherence is at least partially related to problems with the primitive basic operation. when subjects annotated the basic operation correctly, they also showed greater reliability for the source of coherence, with percentages of correct annotations rising to 73% and 66% for the implicit respectively explicit condition (see table 9). it is therefore likely that the reliability of source of coherence will increase if the annotators have a better understanding of the basic operation. scholman, evers-vermeul & sanders 20 our findings suggest that a step-wise approach can be applied by naive annotators, but that the reliability of this approach still can and should be improved. in the current study, we implemented a hierarchical version of the step-wise approach: participants had to first code polarity, and then basic operation, source of coherence and order of the segments respectively. we thought this would help these naive annotators: as the flow chart in figure 1 illustrates, specific options are ruled out once annotators have made certain decisions. for example, if annotators selected the value negative for the primitive polarity, they would only have a choice between causal and non-causal relations, and could not select the categories additive and temporal anymore, as these were grouped together in the non-causal category. similarly, if they wrongly marked a relation as temporal, they did not have the option anymore to indicate whether the relation was objective or subjective. however, results showed that wrong choices on earlier primitives (also) negatively influenced choices on the following primitive(s), as reliability scores went up if only relations were taken into account that were annotated correctly during previous steps. it is worthwhile to explore whether the step-wise approach can be applied presenting the primitives independently of each other, that is without organizing the steps in a hierarchical way. 6.2 substitution and paraphrase tests the current experiment also tested the potential benefits of substitution and paraphrase tests. it was expected that annotators using the explicit instruction – with such tests – would show more agreement than annotators using the implicit instruction without such tests. the results confirmed this hypothesis for the primitives polarity, basic operation and order. no significant differences between the two conditions were found for the primitive source of coherence. these results indicate that the paraphrase and substitution tests indeed guide the annotators in interpreting the relation, except for the paraphrase tests used for source of coherence. the substitution test used for polarity (“can you use but to connect the segments?”) increased the kappa score from .81 to .89, and resulted in higher precision and f-scores for the negative relations. the substitution tests for basic operation (“can you use because / although / whereas / if / then / and to connect the segments?”) increased its kappa score from .45 to .53. more precisely, the substitution tests resulted in higher f-scores for causal, conditional, and non-causal relations: the results in section 5.4 indicated that participants in the explicit condition less frequently classified causal relations as non-causal than participants in the implicit condition. this indicates that the substitution test for this distinction (“can you use although or whereas to connect the segments?”) led to higher agreement. a similar result was found for conditional relations: there was more agreement on this value in the explicit condition than in the implicit condition, although it should be noted that both conditional and non-causal relations were relatively infrequent in the sample corpus. the recall, precision and f-scores showed that the substitution test for temporal relations led to more disagreement. in hindsight, this test (“can you use then/when to connect the segments?”) might indeed have been problematic when applied to specific relations participants had to annotate. several of the temporal relations already included another temporal marker, which made it harder to use then or when for signalling the temporal relation between the segments. for example, the coherence relation in (15) contains the temporal markers het komend jaar ‘next year’ in s1, and vervolgens ‘subsequently’ and tot 2010 ‘till 2010’ in s2. here, the annotators should have removed vervolgens ‘subsequently’ in order to be able to apply the substitution test. this indeed opens up the possibility to insert dan ‘then’, but if the naive annotators in the current study did not recognize vervolgens as a connective, they might have failed to do so. categories of coherence relations in discourse annotation 21 (15) de intercity zoals we die nu kennen wordt afgeschaft; de intercity nieuwe stijl stopt op meer stations en lijkt op de huidige sneltrein. daardoor kan de reistijd langer worden. [er zullen het komend jaar zeven stations bijkomen.]s1 [vervolgens worden tot 2010 in totaal 15 nieuwe stations geopend.]s2 ‘the intercity as we know it now will be abolished; the intercity new style stops at more stations and looks like the current express train. as a result travel time will increase. [next year seven stations will be added.]s1 [subsequently till 2010 a total of 15 new stations will be opened.’]s2 the paraphrase test used for order of the segments (“are s1 and s2 ordered as ‘s1 is the cause, s2 is the consequence’ or ‘s2 is the cause, s1 is the consequence’?”) worked better: it increased the kappa score from .54 to .65, and especially improved the recall, precision, and f-scores of non-basic relations. only the paraphrase test used for source of coherence (“can you paraphrase the relation between s1 and s2 as ‘the fact in one segment causes the fact in the other segment’ or ‘one segment expresses the reason for claiming something in the other segment’?”) did not significantly improve the amount of agreement. taken together, these results indicate that substitution and paraphrase tests can be beneficial, especially to help annotators identify negative from positive relations and causal from additive relations. it is likely that a step-wise approach will yield more agreement if the explanation of certain concepts, such as temporality and subjectivity, when the manual is adapted, and the substitution and paraphrase tests for these concepts are adjusted. 6.3 generalizability of the approach at this point we would like to emphasize again that this study was a first investigation into the viability of a step-wise approach and employing naive annotators for discourse annotation. many participants were not even acquainted with the notions of discourse and coherence, although all of them were undergraduate students in the humanities. their coding of the primitives polarity and order yielded considerable amounts of agreement, but source of coherence and basic operation were shown to be problematic. the step-wise approach was designed to test relatively naive annotators’ potential for performing a discourse annotation task. this does not mean, however, that this approach is restricted to this type of annotators or to the ccr. the question arises whether a step-wise approach might be useful for other types of annotators and applicable to other annotation systems as well. our answer to this question is ‘yes’, given that the step-wise approach yielded satisfactory amounts of agreement for polarity and order, and given that the problematic scores for the primitives source of coherence and basic operation were at least partially due to the hierarchical implementation of the step-wise approach and to problems with specific substitution tests. if naive annotators achieve agreement scores on individual primitives that are similar to results from expert annotations of end labels, this makes us wonder what expert annotators like linguists would do if they were provided with the same materials. a follow-up experiment with expert annotators using a step-wise approach might give insight into the specific facets of discourse annotation that give rise to low interrater reliability scores. additionally, it could be tested what a step-wise approach yields if it is applied to other annotation systems. future research might also reveal how much training exactly is needed for various aspects of discourse annotation. note that more experience in the field of humanities was not helpful for the annotators in our study, given that there were no significant differences in the agreement scores between first and third year undergraduate students. the current study showed relatively low agreement scores for source of coherence, a primitive that is known for being difficult to determine in everyday texts, even for trained annotators. previous studies have discussed this already (see stukker & sanders, 2012, for a recent overview). hence, it is possible that this primitive is too complicated to be annotated reliably by non-expert, non-trained annotators. scholman, evers-vermeul & sanders 22 however, it should be noted that the participants used in this study still managed to reach fair agreement on this primitive without any form of training. the only source of information that was available to the non-trained, non-expert annotators employed in this experiment was two paragraphs in the manual and a paraphrase test in the explicit condition. it needs to be investigated whether a slightly more extensive training of the annotators would improve the reliability of source of coherence, or whether this primitive is better left to experts in the field. this future study could follow suggestions by spooren and degand (2010) by investigating whether non-expert annotators can reach higher agreement if they are able to see and possibly discuss the correct annotations of – for example – the first fifteen fragments after they have annotated them. alternatively, naive annotators might be given just one of the four primitives, which they would have to apply to more relations. this might make these annotators more experienced as they continue to code more coherence relations. similarly, the use of substitution and paraphrase tests need not be restricted to the cognitive approach to coherence relations advocated by sanders et al. (1992, 1993), and applied here. other annotation systems, such as pdtb and rst, might be supplemented with substitution and paraphrase tests as well. given that current interrater reliabilities still leave room for improvement, it seems an attractive option to streamline the annotation schemas for these other approaches, and test whether this leads to similar conclusions. for now, a first investigation into the usability of non-trained, non-expert annotators in discourse annotation has shown that they can yield considerable amounts of agreement in discourse annotation tasks. analyzing coherence relations is a difficult task, even with extensive training and experience. yet non-trained, non-expert annotators using a step-wise approach based on cognitively plausible primitives manage to reach moderate to substantial agreement with little instructions. this indicates that a systematic, step-wise annotation process can decrease the complexity of the annotation task. moreover, it has been shown that an explicit instruction that includes substitution and paraphrase tests benefits annotator agreement. more extensive studies should be conducted to be able to further investigate the extent to which various annotators are able to reliably annotate coherence relations, but the results from the current study can be taken as a clue to the viability of such an approach. acknowledgements this research was partly funded by a clarin-nl grant awarded to ted sanders for the discanproject, and is also based on merel scholman’s ba-thesis. we are grateful to sandrine zufferey, kirsten vis and daan broeder for their valuable input. we also thank philippine geelhand for her help with collecting the data. earlier versions of this paper were presented at the textlink meeting in louvain-la-neuve (2015), clarin-nl meetings in utrecht and soesterberg (2013, 2014), the viot conference in leuven (2014), and the ipra conference in antwerp (2015). we thank the colleagues present at these meetings for the inspiring discussions we had. finally, we thank three anonymous reviewers and the associate editor of this journal for valuable feedback. categories of coherence relations in discourse annotation 23 appendix a : implicit instruction (translated to english, originally in dutch) fragment 1: the amount of biocomponent in shell’s euro 95 is in accordance with the guidelines of secretary of state van geel as announced on prinsjesdag. (vrom) for logistical reasons the biocomponent is mixed on one of shell’s depositories. this means that the percentage of biocomponent in the euro 95 per district can differ. [the euro 95 with biocomponent has the same high quality as the regular euro 95 from shell] [and clients can alternate between the two without doubt.] worldwide shell is active on multiple fronts in the area of biofuels. 1. determine the polarity: is the relation positive or negative? positive negative 2. determine the basic operation: is the relation causal, additive, temporal or, in the case of negative relations, non-causal? if the relation is causal, is it formulated conditionally? causal additive temporal non-causal causal-conditional 3. determine the source of coherence: is the relation objective or subjective? this does not apply to temporal or non-causal negative relations, because they do not differ in source of coherence. therefore, for these relations tick not applicable. objective subjective not applicable 4. determine the order: is the order of the segments basic of non-basic? this does not apply for additive and negative relations, because they do not differ in order. therefore, for these relations tick not applicable. basic non-basic not applicable scholman, evers-vermeul & sanders 24 appendix b: explicit instruction (translated to english, originally in dutch) fragment 1: the amount of biocomponent in shell’s euro 95 is in accordance with the guidelines of secretary of state van geel as announced on prinsjesdag. (vrom) for logistical reasons the biocomponent is mixed on one of shell’s depositories. this means that the percentage of biocomponent in the euro 95 per district can differ. [the euro 95 with biocomponent has the same high quality as the regular euro 95 from shell] [and clients can alternate between the two without doubt.] worldwide shell is active on multiple fronts in the area of biofuels. 0 if the relation contains a connective, take this out of the relation (mentally). do take the original connective into account during your interpretation, so that the meaning of the relation does not change. if a relation contains multiple connectives, such as ‘[i am tired, and therefore i am going to bed early]’, then take both connectives out of the relation. 1 can you use but to connect the segments? yes, then the polarity is negative and the relation belongs to the class of negatives. proceed to question 1a. no, then the polarity is positive. continue to question 2. 1a which of the two connectives best expresses the relation: although or whereas? although, then the basic operation is causal. this relation does not have an order. proceed to question 1b. whereas, then the basic operation is non-causal. this relation does not have a source of coherence or an order. you’ve finished analyzing this relation. 1b can you paraphrase the relation between s1 and s2 as option a or option b? a. a. one segment describes a situation / fact / event which occurred despite the situation / fact / event in the other segment. or b. b. one segment describes a conclusion / claim, despite the situation / fact / event that is described in the other segment. paraphrase a, then the source of coherence is objective. you’ve finished analyzing this relation. paraphrase b, then the source of coherence is subjective. you’ve finished analyzing this relation. 2 can you use because or if to connect the segments? because, then the basic operation is causal. proceed to question 2a. if, then the basic operation is conditional. proceed to question 2a. if neither can be used, not applicable, proceed to question 3. 2a can you paraphrase the relation between s1 and s2 as in option a or rather option b below? c. a. the situation / fact / event in one segment causes the situation / fact / event in the other segment. or d. b. one segment describes the reason for the claim or conclusion given in the other segment. paraphrase a, then the source of coherence is objective. proceed to question 2b. paraphrase b, then the source of coherence is subjective. proceed to question 2c. categories of coherence relations in discourse annotation 25 2b can the order of the segments be described as option a or option b? a. a. s1 is the cause, s2 is the consequence. or b. b. s1 is the consequence, s2 is the cause. paraphrase a, then the relation has a basic order. you’ve finished analyzing this relation. paraphrase b, then the relation has a non-basic order. you’ve finished analyzing this relation. 2c can the order of the segments be described as option a or option b? a. a. s1 describes the reason / argument, s2 describes the claim / conclusion. or b. b. s1 describes the claim / conclusion, s2 describes the reason / argument. paraphrase a, then the relation has a basic order. you’ve finished analyzing this relation. paraphrase b, then the relation has a non-basic order. you’ve finished analyzing this relation. 3 can you use then or when to connect the segments? yes, then the basic operation is temporal. these relations do not have a source of coherence. proceed to question 3a. no, then proceed to question 4. 3a are s1 and s2 chronologically order in time, anti-chronologically or do they happen simultaneously? chronologically, then the order is basic. you’ve finished analyzing this relation. anti-chronologically, then the order is non-basic. you’ve finished analyzing this relation. simultaneously, then the order is not applicable. you’ve finished analyzing this relation. 4 can you use and to connect the segments? yes, then the basic operation is additive. these relations do not differ in order. proceed to question 4a. no, then start again from question 1 and choose the most fitting connective. 4a can you paraphrase the relation between s1 and s2 as option a or option b? a. a. both segments describe a situation / fact / event. or b. b. one or both segments describe an opinion / claim / conclusion. paraphrase a, then the source of coherence is objective. you’ve finished analyzing this relation. paraphrase b, then the source of coherence is subjective. you’ve finished analyzing this relation. scholman, evers-vermeul & sanders 26 references omar alonso and stefano mizzaro (2012). using crowdsourcing for trec relevance assessment. information processing and management, 48(6): 1053-1066. amal al-saif and katja markert (2010). the leeds arabic discourse treebank: annotating discourse connectives for arabic. in n. calzolari, k. choukri, b. maegaard, j. mariani, j. odijk, s. piperidis, m. rosner & d. tapias (eds.), proceedings of 7th international conference on language resources and evaluation (lrec 2010): 2046-2053, malta.ron artstein and massimo poesio (2008). inter-coder agreement for computational linguistics. computational linguistics, 34(4): 555-596. nicholas asher and alex lascarides (2003). logics of conversation. cambridge: cambridge university press. lois bloom, margaret lahey, lois hood, karin lifter and kathleen fiess (1980). complex sentences: acquisition of syntactic connectives and the semantic relations they encode. journal of child language, 7(2): 235-261. thorsten brants (2000). inter-annotator agreement for a german newspaper corpus. proceedings of the sixth conference on applied natural language processing (lrec). seattle, wa. anneloes r. canestrelli, willem m. mak and ted j.m. sanders (2013). causal connectives in discourse processing: how differences in subjectivity are reflected in eye-movements. language and cognitive processes, 28(9): 1394-1413. jean carletta (1996). assessing agreement on classification tasks: the kappa statistic. computational linguistics, 22(2): 249-254. lynn carlson and daniel marcu (2001). discourse tagging reference manual. available online via http://www.isi.edu/~marcu/discourse/tagging-ref-manual.pdf. lynn carlson, daniel marcu and mary e. okurowski (2003). building a discourse-tagged corpus in the framework of rhetorical structure theory. in j. van kuppevelt and r. smith (eds.), current directions in discourse and dialogue: 85-112. dordrecht: kluwer academic publishers. susan conrad (2002). corpus linguistic approaches for discourse analysis. annual review of applied linguistics, 22: 75-95. liesbeth degand and henk pander maat (2003). a contrastive study of dutch and french causal connectives on the speaker involvement scale. in a. verhagen and j. van de weijer (eds.), usage-based approaches to dutch: 175-199. utrecht: lot. liesbeth degand (2001). form and function of causation: a theoretical and empirical investigation of causal constructions in dutch. leuven: peeters. oswald ducrot (1980). essai d’application: mais les allusions à l’énonciation – délocutifs, performatifs, discours indirect. in: h. parret (ed.), le langage en context: etudes philosophiques et linguistiques de pragmatique : 487-575. amsterdam: john benjamins. jacqueline evers-vermeul (2005). the development of dutch connectives; change and acquisition as windows on form-function relations. ph.d. dissertation. utrecht: lot. available online via http://www.lotpublications.nl/documents/110_fulltext.pdf. jacqueline evers-vermeul and ted j.m. sanders (2009). the emergence of dutch connectives: how cumulative cognitive complexity explains the order of acquisition. journal of child language, 36(4): 829-854. barbara j. grosz and candace l. sidner (1986). attention, intentions and the structure of discourse. computational linguistics, 12(3): 175-204. michael a.k. halliday and ruqaiya hasan (1976). cohesion in english. london: longman. jerry r. hobbs (1979). coherence and coreference. cognitive science, 3(1): 67-90. jerry r. hobbs (1985). on the coherence and structure of discourse. csli center for the study of language and information, stanford university. categories of coherence relations in discourse annotation 27 andrew kehler (2002). coherence, reference, and the theory of grammar. stanford, ca: csli publications. alistair knott and robert dale (1994). using linguistic phenomena to motivate a set of coherence relations. discourse processes, 18: 35-62. alistair knott and ted j.m. sanders (1998). the classification of coherence relations and their linguistic markers: an exploration of two languages. journal of pragmatics, 30: 135-175. klaus krippendorff (1980). content analysis: an introduction to its methodology. beverly hills, ca: sage. ewald lang (1984). the semantics of coordination. amsterdam: john benjamins. fang li, jacqueline evers-vermeul and ted j.m. sanders (2013). subjectivity and result marking in mandarin: a corpus-based investigation. chinese language and discourse, 4(1): 74119. william c. mann and sandra a. thompson (1988). rhetorical structure theory: towards a functional theory of text organization. text, 8(3): 243-281. eleni miltsakaki, rashmi prasad, aravind joshi and bonnie webber (2004). annotating discourse connectives and their arguments. proceedings of the frontiers in corpus annotation 2004 naacl/hlt conference workshop, boston. megan moser and johanna d. moore (1996). on the correlation of cues with discourse structure: results from a corpus study. university of pittsburgh, learning research and development center. available online via homepages.inf.ed.ac.uk/jmoore/papers/rda.ps. megan moser, johanna d. moore and erin glendening (1996). instructions for coding explanations: identifying segments, relations and minimal units. university of pittsburgh, learning research and development center. available online via http://homepages.inf.ed.ac.uk/jmoore/papers/rda-instr.ps. leo g.m. noordman and femke de blijzer (2000). on the processing of causal relations. in: e. couper kuhlen & b. kortmann (eds.), cause, condition, concession, contrast: cognitive and discourse perspectives. berlin, new york: mouton de gruyter. leo g.m. noordman and wietske vonk (1998): memory-­‐based processing in understanding causal information, discourse processes, 26(2-3): 191-212. stefanie nowak and stefan rüger (2010). how reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation. proceedings of the international conference on multimedia information retrieval (mir), philadelphia, usa. henk pander maat and ted j.m. sanders (2000). domains of use or subjectivity? the distribution of three dutch causal connectives explained. topics in english linguistics, 33: 57-82. pdtb research group (2007). the penn discourse treebank 2.0 annotation manual. available online via http://www.seas.upenn.edu/~pdtb/pdtbapi/pdtb-annotation-manual.pdf. mirna pit (2007). cross-linguistic analyses of backward causal connectives in dutch, german and french. languages in contrast, 7(1): 53-82. emily pitler and ani nenkova (2009). using syntax to disambiguate explicit discourse connectives in text. proceedings of the acl-ijcnlp 2009 conference: 13-16, singapore. massimo poesio and ron artstein (2005). the reliability of anaphoric annotation, reconsidered: taking ambiguity into account. proceedings of the workshop on frontiers in corpus annotations ii: pie in the sky: 76-83.rashmi prasad and harry bunt (2015). semantic relations in discourse: the current state of iso 24617-8. in h. bunt (ed.), proceedings of the 11th joint acl iso workshop on interoperable semantic annotation (isa-11): 8091. tilburg: ticc, tilburg university. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi and bonnie webber (2008). the penn discourse treebank 2.0. proceedings of the 6th international conference of language resources and evaluation (lrec 2008), marrakech. scholman, evers-vermeul & sanders 28 ted j.m. sanders (1997). semantic and pragmatic sources of coherence: on the categorization of coherence relations in context. discourse processes, 24: 119-147. ted j.m. sanders and leo g.m. noordman (2000). the role of coherence relations and their linguistic markers in text processing. discourse processes, 29: 37-60. ted j.m. sanders and wilbert p.m.s. spooren (2009). causal categories in discourse: converging evidence from language use. in: t.j.m. sanders and e. sweetser (eds.), causal categories in discourse and cognition. berlin: walter de gruyter. ted j.m. sanders and wilbert p.m.s. spooren (2015). causality and subjectivity in discourse: the meaning and use of causal connectives in spontaneous conversation, chat interactions and written text. linguistics, 53(1): 53-92. ted j.m. sanders, wilbert p.m.s. spooren and leo g.m. noordman (1992). toward a taxonomy of coherence relations. discourse processes, 15: 1-35. ted j.m. sanders, wilbert p.m.s. spooren and leo g.m. noordman (1993). coherence relations in a cognitive theory of discourse representation. cognitive linguistics, 4(2): 93-133. ted j.m. sanders, kirsten vis and daan broeder (2012). project notes of clarin project discan: towards a discourse annotation system for dutch language corpora. project notes. utrecht: utrecht university. william a. scott (1955). reliability of content analysis: the case of nominal scale coding. public opinion quarterly, 19(3): 321-325. wilbert p.m.s. spooren and liesbeth degand (2010). coding coherence relations: reliability and validity. corpus linguistics and linguistic theory, 6(2): 241-266. wilbert p.m.s. spooren and ted j.m. sanders (2008). the acquisition order of coherence relations: on cognitive complexity in discourse. journal of pragmatics, 40: 2003-2026. manfred stede (2004). the potsdam commentary corpus. proceedings acl workshop on discourse annotation. pennsylvania: acl. ninke m. stukker and ted j.m. sanders (2012). subjectivity and prototype structure in causal connectives. a cross-linguistic perspective. journal of pragmatics, 44(2): 169-190. ninke stukker, ted j.m. sanders and arie verhagen (2008). causality in verbs and in discourse connectives: converging evidence of cross-level parallels in dutch linguistic categorization. journal of pragmatics, 40:1296-1322. kai-ming ting (2010). precision and recall. in: c. sammut and g.i. webb (eds.), encyclopedia of machine learning. new york: springer us. matthew j. traxler, michael d. bybee and martin j. pickering (1997). influences of connectives on language comprehension: eye-tracking evidence for incremental interpretation. quarterly journal of experimental psychology, 50(3): 481-497. matthew j. traxler, anthony j. sanford, joy p. aked and linda m. moxey (1997). processing causal and diagnostic statements in discourse. journal of experimental psychology: learning, memory, and cognition, 23(1): 88-101. nynke van der vliet, ildikó berzlanovich, gosse bouma, markus egg and gisela redeker (2011). building a discourse-annotated dutch text corpus. in: s. dipper and h. zinsmeister (eds.), proceedings beyond semantics (dgfs workshop). bochumer linguistische arbeitsberichte 3: 157–171. yannick versley and anna gastel (2012). linguistic tests for discourse relations in the tüba-d/z corpus of written german. dialogue and discourse, 4(2): 142-173. sandrine zufferey (2012). “car, parce que, puisque” revisited: three empirical studies on french causal connectives. journal of pragmatics, 44(2): 138-153. reasoning between the lines: a logic of relational propositions dialogue & discourse 9(2) 80-110 doi: 10.5087/dad.2018.203 ©2019 andrew potter this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). reasoning between the lines: a logic of relational propositions andrew potter apotter1@una.edu computer science & information systems university of north alabama editor: manfred stede submitted 08/2017; accepted 12/2018; published online 01/2019 abstract this paper describes how rhetorical structure theory (rst) and relational propositions can be used to define a method for rendering and analyzing texts as expressions in propositional logic. relational propositions, the implicit assertions that correspond to rst relations, are defined using standard logical operators and rules of inference. the resulting logical forms are used to construct logical expressions that map to rst tree structures. the resulting expressions show that inference is pervasive within coherent texts. to support reasoning over these expressions, a set of rules for negation is defined. the logical forms and their negation rules can be used to examine the flow of reasoning and the effects of incoherence. because there is a correspondence between logical coherence and the functional relationships of rst, an rst analysis that cannot pass the test of logic is indicative either of a problematic analysis or of an incoherent text. the result is a method for analyzing the logic implicit within discursive reasoning. keywords: rhetorical structure theory, rst, relational propositions, propositional logic 1 introduction an important contribution of rhetorical structure theory (rst) has been its usefulness in describing the organization of text. an analysis created using rst presents the text structurally, identifying the functional relationships between the parts of the text and positioning them hierarchically to reflect their relative status within the overall textual organization. this provides, among other things, a basis for examination of the nature of the coherence of the text, and it has led to diverse studies and applications in discourse analysis, argumentation, theoretical linguistics, psycholinguistics, and computational linguistics (taboada & mann, 2006a). an area of research that has received little attention has been the interrelationship between rst and propositional logic. although the relationship between logic and language has been of longstanding interest among philosophers, logicians, linguists, and psychologists, studies specific to rst and propositional logic are not to be found. and yet the characteristics of rst and propositional logic suggest there may be untapped synergies. for both, the object of analysis is a well-defined domain: in rst, the domain is the text under analysis (mann & thompson, 1987b), and in propositional logic, the domain is the universe of discourse, an explicit or implicit boundary within which the analysis is delimited (boole, 1854). both involve analysis of relationships among discourse units at the sentential level, and both can be used to discover the sentential organization of discourse. both are concerned with discursive coherence, and both are useful in the analysis of discursive reasoning. the purpose of this paper is to examine the use of rst for rendering and analyzing texts as expressions in propositional logic. the primary claim to be developed is that any text analyzable potter 81 using rst can be restated as a boolean expression. that is, for any rst structure, there is a corresponding boolean expression consisting of a set of variables representing the discourse elements and connected using logical operators. the logical relations within such a boolean expression are inferentially interdependent. a set of rules of negation is formulated, and when applied to a text, they facilitate analysis of logical flows from one part to another. this is useful in analyzing rhetorical structures and their logical consequences for the interpretation of text. the result is a method for analyzing the logic implicit within discursive reasoning. rst is a functional theory of textual coherence. it defines coherence in terms of the relations that occur between parts of the text. the parts of the text may be elementary discourse units, or they may be structures of multiple related units. while there has been considerable debate as to the appropriate set of relations to be used in such an analysis (e.g., carlson & marcu, 2001; grosz & sidner, 1986; hovy, 1990; hovy & maier, 1993; knott, oberlander, o'donnell, & mellish, 2001; mann & thompson, 1987b; marcu, 2000; sanders, spooren, & noordman, 1992; stede, taboada, & das, 2017), there seems to be general agreement that there are identifiable relations and these may be used in describing the nature of coherence in a text. rst further defines coherence using a set of constraints that define the structural characteristics of the text as a whole. these constraints are identified as completeness, connectedness, uniqueness, and adjacency (mann & thompson, 1987b). the application of these constraints enables the analyst to identify the underlying hierarchical structure of the text. logic defines coherence as consistency (kolodny, 2007; thagard, 1989). to be considered coherent, the universe of discourse must be free of contradiction. the universe of discourse specifies the assertions constituting the universe and their logical interdependencies. the propositions constituting a coherent universe are true, either as elementary propositions asserted as axiomatic, or as inferred propositions, having been deduced either from elementary propositions or from other inferred propositions in accordance with the rules of inference. propositional logic includes a set of rules of inference, the logical functions that determine which conclusions may be drawn from a given set of premises (boole, 1854; russell, 1906). thus, a coherent universe of discourse may be viewed as an inferential structure of logically related propositions. these two concepts of coherence, one based on rst and the other on propositional logic, though distinct, have an important similarity: rst represents a text as discourse units connected by relations, and logic represents the universe of discourse as propositions connected by logical operators. if it were the case that rst relations corresponded one to one with logical operators and units with propositions, the task of restating rst analyses as boolean expressions would not be interesting. however, the task is not so simple. rst relations do not correspond one to one with logical operators. as will be developed in section 3, teasing out the logical expressions corresponding to rst relations involves a close analysis of the rst definitions as well as examination of examples of rst analyses. some of the resulting expressions are complex. of additional concern is the ostensible assumption of parity between elementary discourse units and propositions. in logic, a proposition is defined as being either true or false, but not both, a definition normally associated with declarative sentences. but the discourse units of rst need not be declarative. they can be any sentence type whatsoever, and in some cases, the units are sentence fragments or dependent clauses. while this issue may be addressed in part through the recognition that many non-declarative units may be restated as propositions, this does not fully speak to the issue. to fully address this issue, it is first necessary to consider propositional logic within the larger context of boolean algebra. propositional logic and the boolean algebra are interchangeable conceptualizations. essentially, propositional logic is an application of the boolean algebra. the logical operations of the boolean algebra correspond to the theorems of propositional logic, and boolean variables correspond to elementary propositions. boolean logic is more general and lends reasoning between the lines 82 itself to a variety of applications, including the design of switching circuits, mathematics, set theory, digital logic, and database query languages. although when discussing propositional logic we tend to speak in terms of truth values and truth functions, these semantics are logically arbitrary, and we could just as well speak of + and -, 1 and 0, yes and no, satisfiability and unsatisfiability, or any other bivalent conceptualization, such as belief and non-belief, positive and negative regard, desire and indifference, interest and disinterest, understanding and misunderstanding, or ability and inability. to the extent that the primitives of rst can be understood in terms of bivalent values, they are amenable to logical treatment. any discourse unit that can be either accepted or rejected can be treated as a proposition and subjected to the rules of inference. furthermore, in this paper, discourse units are of less concern as isolated units, but rather are of interest primarily with respect to the role they perform within a relational structure. this is perhaps most easily understood in terms of relational propositions, the alter ego of rst relations (mann & thompson, 1986a, 1986b, 2000a). relational propositions will be explained in detail shortly, but for now, suffice it to say that they are implicit assertions that arise between clauses within a text and are essential to the effective functioning of the text. they are rst relations reformulated as propositions. a relational proposition consists of a predicate and two variables. the predicate corresponds to an rst relation, and the variables map to the units or spans the relation references, one to the nucleus, the other to the satellite. the predicate defines the relationship between these variables. these propositions provide the basis for the logical interpretation of rst analyses. while it is convenient to discuss relational propositions in terms of rst relations, as propositions, these relations may also be treated as truth-functional assertions. thus, the perspective taken in this paper is that units can be considered not only within a broader boolean conceptualization, but that their interpretation is reliant on the role they play as arguments within their relational propositions. the value of this approach should become apparent through the analysis of specific relations and structures in sections 3-5. an additional potential concern is with text spans. in rst, there are, in addition to units and relations, text spans that represent uninterrupted linear intervals of text, consisting of multiple units and relations. the mann and thompson (2000b) analysis shown in figure 1 contains several such spans, such as the span represented as 2-4, representing the relations occurring among units 2, 3, and 4. such multiunit text spans are not explicitly defined in rst, raising the possibility of difficulties in logical interpretation. the elaboration relation shown here is between one unit and an evidence span of two others. mann and thompson (2000a) claimed the corresponding relational proposition would be a complex and framework-dependent assertion. however, the only framework upon which the corresponding relational proposition is dependent is rhetorical structure theory itself. the use of multiunit text spans is a diagrammatic shorthand for the structures encapsulated below them, and from the logical perspective, they function much the same as parentheses in a symbolic expression. this will be discussed in detail in section 2.2. the organization of the remainder of this paper is as follows. section 2 presents brief overviews of the theoretical concepts used in the paper. these include rst, relational propositions, and propositional logic. these overviews are followed in section 3 with an analysis of rst relations, providing a specification of the logical form of each. section 4 presents a set of rules for negating relational propositions. these rules are based on potential points of failure within the logical forms. in section 5, the logical forms and the rules of negation are used to examine the logical coherence of two texts. finally, section 6 provides a perspective on this research, locates it within the larger context of studies in the relationship between coherence relations and logic, summarizes the results, and suggests some areas for future application and research. potter 83 figure 1. rst analysis containing multiunit spans 2 theoretical background this section provides brief overviews of the major theoretical frameworks that support this paper. these include rhetorical structure theory (rst), relational propositions, and propositional logic. rst provides the foundation for describing the sentential organization of a text. it also provides the framework for using relational propositions. the theory of relational propositions is an alternative conceptualization of rst, such that the relationships between text spans are viewed as implicit assertions that occur between clauses in a text. propositional logic can then be used to define these assertions as logical expressions constructed of propositions and logical operators. 2.1 rhetorical structure theory rhetorical structure theory (rst) is a descriptive theory of text organization. it is a tool for describing and characterizing texts in terms of the relations that hold among the clauses comprising the text. the elements of rst are relations, schemas, schema applications, and structures. an rst relation consists of three parts: the satellite, the nucleus, and the relation. the satellite and nucleus are text spans. the distinction between satellite and nucleus arises as a result of the asymmetry of the relations. within a relation, the nucleus is more salient than the satellite. relations are defined in terms of four fields: constraints on the nucleus, constraints on the satellite, constraints on the combination of nucleus and satellite, and the intended effect. for example, in the definition of the evidence relation, the constraint on the nucleus is that the reader might not believe the information presented in the nucleus, the constraint on the satellite is that the reader must believe it or at least find it credible, and the constraint on the combination of the nucleus and satellite is that by comprehending the satellite, the reader will have an increased belief in the nucleus. the effect of evidence is that the reader’s belief in the nucleus is increased. all of these fields must be satisfied for an application of the evidence relation. mann and thompson (1987b) proposed that relations could be grouped into two general categories: presentational and subject matter. the distinction is based on intended effect. as defined by mann and thompson (p. 18), presentational relations are relations “whose intended effect is to increase some inclination in the reader, such as the desire to act or the degree of positive regard for, belief in, or acceptance of the nucleus.” an instance of a presentational relation is, when successful, an act of persuasion. subject matter relations are relations whose intended effect is that reasoning between the lines 84 the reader recognize the relation in question. unlike the presentational relations, they involve no act of persuasion. the significance of this distinction for determining logical form will become clear in section 3. schemas are abstract patterns used to specify the relations that may hold among a small set of text spans. mann and thompson (1987b) defined five schema types: 1) a relation with a nucleus and satellite, 2) multiple relations with a single nucleus and a satellite for each relation, 3) contrastive multinuclear relations, 4) joint multinuclear relations, and 5) sequential multinuclear relations. the first schema, a relation with a nucleus and a satellite, is employed for most rst relations and is the schema most often found in texts. schema applications define the conventions for schema instantiation. by convention, the application of schemas places no constraint on the order of spans (i.e. the nucleus may precede or follow the satellite), multi-relational schemas do not restrict which relations are used in their application, and any relation may be repeated multiple times in the application of a schema. the rst structure of a text refers to its overall composition in terms of schema applications. for the text to be judged coherent, it must conform to four constraints: completeness, connectedness, uniqueness, and adjacency. to be complete, the text must be analyzable as a unified structure. to meet the connectedness constraint, all text spans must be either minimal units or a constituent within a schema application. uniqueness requires that each schema application consist of a unique set of text spans. the adjacency constraint states that the text spans comprising any schema application constitute one contiguous text span. a consequence of these constraints is that an analysis performed using rst results in a tree structure. rst has been used for a variety of purposes. it has been used as a way of describing the relations between clauses of a text, irrespective of explicit grammatical or lexical signals. it has been used for analysis of a wide range of text types, such as expository prose, narrative discourse, and news broadcasts. it has been used in studies in contrastive rhetoric. of importance in this paper, rst has been used as a framework for investigating relational propositions, which are the topic of the next section. 2.2 relational propositions relational propositions are implicit coherence-producing assertions that serve to bind together the explicit parts of a text and are essential to the effective functioning of the text. this example (updated for the modern reader) from mann and thompson (1986b) is illuminating: i love to collect classic automobiles. my favorite car is my 1899 duryea. i love to collect classic automobiles. my favorite car is my 2008 toyota camry. even without knowing much about classic cars, most readers would agree that the first text is coherent and the second is not. in the first text, there is an implicit relational proposition that binds the two parts of the text together. the relational proposition asserts that the second part of the text is an elaboration of the first. the reader does not need to be explicitly informed of this relationship. the parts go together. in the second text, this is not the case. the second part of the text does not follow from the first. so essential are relational propositions that their failure may destroy the coherence of a text. these propositions are pervasive throughout the text, and form the basis for various kinds of inference. while lying outside mainstream rst research (taboada & mann, 2006b), the existence and nature of these propositions has been generally acknowledged (mann & thompson, 1986a, 2000a; nicholas, 1994; taboada & mann, 2006b). for each relation in an rst analysis, there is a corresponding relational proposition. rst and relational propositions provide parallel accounts of coherence. while rst accounts for coherence relations among the spans within a text, they do not identify the implicit relational acts that account for how the text functions (mann & thompson, 1986b). that is the unique contribution of potter 85 relational propositions. because relational propositions may be derived from the definitions of rst relations, they need not be sought directly in the text itself (mann & thompson, 1986a, 2000a). this means that relational propositions may be used as adjunct to rst relations. a relational proposition consists of a predicate and a pair of arguments. the predicate corresponds to the rst relation, and the arguments correspond to the satellite and nucleus. as presented by mann and thompson (2000a), the ordering of arguments in a relational proposition need not reflect their original ordering in the text, nor their canonical orderings in rst. the satellite precedes the nucleus, such that (s)predicate(n). for example, the relational proposition for the duryea example above would be (my favorite car is my 1899 duryea) is an elaboration of (i love to collect classic automobiles) these expressions can be made more concise by using a conventional predicate notation with the initial argument as the satellite and the second argument as the nucleus; for example the relational proposition corresponding to the elaboration relation would be elaboration(s, n). for relational propositions to serve as truth-functional representations of rst structures, they not only need to be capable of expressing elementary relational assertions, they must be applicable to complete rst structures. this is an area that has not been previously addressed. as noted above, mann and thompson (2000a) claimed that representing multiunit text spans would be a complex framework-dependent undertaking, and did not address it. they confined their research to simple relational propositions. however, rst structures contain multiunit text spans. for example the structure shown above in figure 1, from (mann & thompson, 2000b), contains several multiunit spans (1-6, 2-6, 2-4, 3-4, and 5-6). as noted earlier, such text spans are a diagrammatic shorthand for the structures appearing below them. they provide the same organizational function as parentheses in a symbolic notation. 3-4 is one group of units, 5-6 is another. there is a nesting of relational propositions that corresponds to the structure of the rst tree. the evidence relation between units 3 and 4 is the satellite of the elaboration relation. the concession relation between units 5 and 6 is the satellite of the interpretation relation, of which the elaboration relation is the nucleus. the resulting structure is the nucleus of the preparation relation: preparation(1,interpretation(concession(5,6), elaboration(evidence(4,3),2))) thus nested relational propositions enable the identification, not only of individual rst relations as assertions, but the expression of an entire rst structure as a relational proposition. the evaluation of such an expression begins with the innermost relational propositions, and works outward through the nested propositions. this suggests that there is a logical dependence that flows from the leaf nodes to the root. such an evaluation requires that each relation have a logical specification. supplying these specifications is the topic of section 3 of this paper. 2.3 propositional logic in propositional logic, propositions are treated as primitive units that may be joined using logical operators to form more complex logical expressions. these logical operators are used in this paper: negation   p conjunction  p  q disjunction  p  q exclusive disjunction  p  q material implication → p → q material equivalence  p  q reasoning between the lines 86 when propositions are combined using these operators, the resulting expressions depend on the constituent propositions and the semantics of the operators for their truth value. if p is true then  p is false. the conjunction of two propositions (p  q) is true if both p and q are true. disjunction can be defined in terms of negation and conjunction, as  ( p   q)), i.e., it is not the case that both p and q are false; (p  q) is true only if either p or q or both are true. exclusive disjunction (p  q) is true if p is true or q is true but not both p and q are true. material implication states that (p → q) is true if it is the case that  (p   q). that is, it will never be the case that p is true when q is false. material equivalence (or biconditionality) is a relation of mutual implication. if ((p → q)  (q→ p)), then (p  q). for logical definitions of rst relations, discourse units within a satellite-nucleus relation are identified using the symbols s and n. for example, the definition of antithesis, discussed below in section 3.1.1, is ((s  n)  s) → n for structures involving more than two variables, the traditional symbols from propositional logic are used. thus, a structure including three discourse units identifies the symbols p, q, and r, as in (((p → q)  p) → q)  (((r → q)  r) → q) and when providing a logical interpretation of a specific rst analysis, the units will be identified numerically, reflecting the numbering of units in the analysis, such as (((1 → 2)  1) → 2)  (((3 → 2)  3) → 2) the application of logic to rst, as with any application of logic to the real world, involves the use of judgment under conditions of uncertainty. mann and thompson (1987b) specify that when specifying an rst relation between text spans, the analyst must make a judgment with respect to each of these spans, and that owing to the nature of text analysis, these judgments must be made in terms of plausibility rather than certainty. additional uncertainty will be introduced with the logical interpretations of the relations. the rst relations, although specific, are applicable over a range of situations, and some will conform more closely to the interpretations than others. subject to these limitations, the following interpretation of rst relations will identify plausible inferential characteristics of rhetorical relations and show how they can be modeled using elementary operations of formal logic. 3 a logical interpretation of rhetorical relations for rst to be useful in rendering texts as expressions in propositional logic, it is necessary to establish that for each rst relation there is a corresponding logical form. in some cases this amounts to a simple one-to-one correspondence between the relation and a standard rule of inference, and in others the interpretation results in more complex inferential expressions. for each rst relation it is necessary to examine the nature of the asymmetry between the satellite and nucleus. in some cases the relationship is clearly implicative. in others, discovery of the inferential quality of the relation requires more probing. this section presents a discussion of each of the rst relations in the extended mann and thompson relation set, as identified on the rst website,1 1 http://www.sfu.ca/rst/index.html potter 87 providing a logical interpretation and rationale for each. the relations are grouped according to their traditional categories, starting with presentational relations, followed by subject matter relations and then multinuclear relations. 3.1 presentational relations presentational relations are relations “whose intended effect is to increase some inclination in the reader, such as the desire to act or the degree of positive regard for, belief in, or acceptance of the nucleus” (mann & thompson, 1987b, p. 18). an instance of a presentational relation is, when successful, an act of persuasion. in each of the presentational relations, the satellite serves as the impetus for the reader’s acceptance of the intended effect. in this way the writer seeks to influence the reader’s reasoning processes. increases in belief, desire, acceptance, regard, understanding, ability, and interest are the results of inferences specified by the relational propositions associated with these relations. consequently, these relations are logically argumentative; that is, the logical form of each relation presents a valid argument, consisting of a set of premises and a conclusion which follows from the premises. a consequence of this is that the logical forms of the presentational relations are tautologies. that these relations are tautologous arises due to the relationship between valid arguments and tautologies. for any valid argument, there is a corresponding conditional statement whose antecedent is the conjunction of the premises and whose consequent is the conclusion, and this corresponding statement will be a tautology (copi, 1967). thus any relation that is logically argumentative will be tautologous. the implications of this will be explored in detail in sections 4 and 5. 3.1.1 antithesis with the antithesis relation, the intended effect is to increase the reader’s positive regard for the situation presented in the nucleus. the satellite is incompatible with the nucleus, such that the reader cannot have positive regard for both the nucleus and the satellite. this incompatibility increases the reader’s positive regard for the nucleus. in this example of antithesis, from thompson and mann (1987), s: rather than winning them with our arms, n: we'd win them by our example, and their desire to follow it. two alternatives are presented, with the nucleus clearly preferred over the satellite. the logic is simple, based on disjunctive syllogism: (((s  n)   s) → n). to paraphrase, s or n but not s, therefore n. 3.1.2 concession with the concession relation, the writer concedes the situation presented in the satellite and asserts that, though there might seem to be a potential incompatibility between the satellite and the nucleus, the satellite and nucleus are indeed compatible. the writer holds the nucleus in positive regard, and by indicating a lack of incompatibility with the satellite, the writer seeks to increase the reader’s positive regard for the nucleus. to understand the logic of concession, it is necessary to consider the situation in which it is presented. in their paper on concessive relations, thompson and mann (1986, p. 441) observed: only in terms of its discourse context can we understand how concession is a 'conceding' of something: it concedes the potential incompatibility of two situations in order to forestall an objection that could interfere with the reader's belief of the point the writer wants to make. reasoning between the lines 88 the objection having been forestalled, its potential interference with the reader’s belief has been removed. an example of concession occurs in the common cause letter, as analyzed by mann and thompson in several of their papers (mann & thompson, 1986a, 1986b; thompson & mann, 1987): s: tempting as it may be, n: we shouldn’t embrace every popular issue that comes along. the satellite and nucleus are not incompatible: one may be tempted, yet still resist embracing every popular issue. as such, it is not the case that the satellite provides grounds for rejecting the nucleus:  (s →  n) and upon neutralizing this objection, the writer further invites the reader to infer from this the claim presented by the nucleus. if the satellite does not imply the negation of the nucleus, then the nucleus holds. the reasoning thus becomes an instance of modus ponens: (( (s → n) → n)   (s →  n)) → n antithesis and concession both increase positive regard for the nucleus at the expense of the satellite. antithesis negates the satellite, and concession dismisses it. in analyzing texts, it can sometimes be difficult to determine which relation is applicable. in the analysis shown above in figure 1, mann and thompson coded the relation for text span 5-6 as concessive: although there is no recession, the economy is not doing all that well. the antithetical reading is that the writer presents a choice between outright recession and protracted sluggishness, and ruling out recession leads to an inference of protracted sluggishness. the writer is negating the satellite while confirming the nucleus. that fits the logical form of antithesis precisely. in this manner, the logical forms can be useful in analyzing rhetorical structures. this will be developed further in section 5. 3.1.3 evidence with the evidence relation, the satellite provides evidence in support of the nucleus. for the relation to achieve its intended effect, it is necessary that the reader accept the satellite, either axiomatically as an assumption, or inferentially derived from subordinate relations in the text structure. the reader must recognize the implicative relationship between the satellite and the nucleus, with the result that the reader’s belief in the nucleus is increased. in argumentative terms, the satellite is the ground and the nucleus is the claim. the implicative relationship is subject to two constraints. evidence must be defined to specify that if the antecedent is believable, the consequent will also be believable (material implication) and that the antecedent is believable (modus ponens). material implication specifies a conditional relationship between the antecedent and the consequent, such that if the antecedent is believable, the consequent will also be believable: p → q. by definition, material implication does not specify whether the antecedent is true or not. as such, material implication is hypothetical. the evidence relation is not hypothetical. to achieve its effect, evidence requires that the antecedent (i.e. the satellite) be asserted. consequently the implicative relationship is, in addition to material implication, modus ponens: ((s → n)  s) → n potter 89 although the writer anticipates that the reader will accept the satellite, no such assumption is made about the nucleus. if such an assumption could be made, there would be no need for the evidence relation: the relation could be more properly coded as elaboration. in this example of the evidence relation, n: the program as published for calendar year 1980 really works. s: in only a few minutes, i entered all the figures from my 1980 tax return and got a result which agreed with my hand calculations to the penny. the writer is praising an income tax preparation software application; efficiency and consistency between programmatic and hand calculations are presented as evidence that the tax program really works (mann & thompson, 1987b). the satellite is asserted and is in an implicative relationship with the nucleus. as discussed below, this use of the modus ponens is applicable to the remaining presentational relations, including justify, enablement, motivation, background, preparation, summary, and restatement. 3.1.4 justify the satellite of justify is used to increase the reader’s readiness to accept the writer’s right to present what follows in the nucleus. this right to present may come from a variety of sources for use in a variety of situations. for example, the writer may wish to assert a position of authority, make known educational qualifications or professional credentials, announce a leadership role within an organization, or cite qualifying experience within the area under discussion. in this example from ai magazine, ken forbus (2016) makes reference to his experience and expertise in artificial intelligence: s: indeed, in my experience, n: today’s general-purpose ai systems tend to skate a very narrow line between catatonia and attention deficit disorder. his expertise is invoked to provide authority to the claim. without this authority, the comparison of ai systems to catatonia and attention deficit disorder would seem unworthy of attention. writers may also use justify when seeking to set aside authority and speak inclusively, so as to obtain acceptance of what follows. a good example of this is the famous quotation from us president john kennedy, in which he used justify to support an antithesis: s: and so, my fellow americans: n: ask not what your country can do for you; ask what you can do for your country. assuming an authoritative stance when asking the audience to give up their selfish desires in favor of altruistic contributions would be to risk accusation of self-serving hypocrisy. justify is also used when the writer is in a subordinate position, seeking permission to speak, or as a (perhaps ritualistic) show of respect, as when a lawyer in a u.s. courtroom begins an oral argument with the phrase, may it please the court…. stede et al. (2017) observe that with justify the satellite may be used to present a basic attitude of the writer, making it easier for the reader to accept the nucleus. an example of this can be found in the common cause analysis, as shown in figure 2. the writer seeks to dispel any perception of conflict between his personal views and his strategic view regarding the nuclear freeze proposal. justify achieves credibility through advantageously positioning the writer with respect to the claim made in the nucleus. reasoning between the lines 90 figure 2. common cause example of justify 3.1.5 motivation and enablement mann and thompson (1987b) treat motivation and enablement as a subgroup within the presentational relations. in both of these relations the nucleus presents an action. the action may be presented as an imperative or as an implicit call to action. motivation increases the reader's desire to perform the action, and enablement informs the reader how to do it. these two relations are often found together, their satellites linked to a common nucleus. this structural arrangement has sometimes been referred to as a motivation-enablement schema (mann & thompson, 1987b). the relations work well together, where one increases a desire and the other shows how to satisfy it. motivation and enablement are implicative and follow the same logical form as evidence. if the motivation satellite fails to motivate or if the enablement satellite fails to increase the reader’s ability to perform the action, the action presented in the nucleus is less likely to occur. when multiple satellites of a presentational relation share the same nucleus, they combine conjunctively to support the nucleus. in the motivation-enablement schema, desire and ability join forces to persuade the reader to take the indicated action. figure 3. linked argument using motivation and enablement in the analysis shown in figure 3, from mann and thompson (1987a), the action to be taken is to attend the 1983 annual staff breakfast. information about when and where the event will be held enables the reader to attend, and the opportunity to meet other staff members, staff assembly potter 91 representatives, and the university president provides reason for attending. the absence of either satellite would not necessarily lead to failure to take action. but given the presence of both satellites, if either of the relational propositions were to fail, i.e. if the intended motivation were demotivating, or the enablement contained incorrect or irrelevant information, then the motivation-enablement schema would fail. taken together they present a linked argument, sharing a common locus of effect: (((2 → 1)  2) → 1)  (((3 → 1)  3) → 1) 3.1.6 background background is used to assure comprehensibility of the text. a nucleus that would otherwise be incomprehensible is rendered understandable by virtue of its satellite. in this example from mann and thompson (1987a), the satellite provides background for the nucleus. without the satellite, the nucleus would be difficult and perhaps impossible to understand: s: i am having my car repaired in santa monica (1522 lincoln blvd.) this thursday 19th. n: would anyone be able to bring me to isi from there in the morning or drop me back there by 5 pm please? background has the same logical form as evidence, though the implication is one of comprehension rather than belief or acceptance. 3.1.7 preparation stede et al. (2017) specified that preparation be used when the satellite serves no stronger purpose than setting the topic for the nucleus. but setting the topic can be decisive in determining whether the writer achieves the intended effect of the text. preparation may determine whether the reader bothers to read the text. in the analysis shown in figure 4, the title serves to announce and delimit the topic of the text. the satellite precedes the nucleus and tends to make the reader more ready, interested or oriented for reading the nucleus. used in this way, preparation is intended to increase an inclination in the reader, namely an inclination to read the article. thus, preparation is presentational and subject to the same logical form as evidence. figure 4. the preparation relation reasoning between the lines 92 3.1.8 restatement and summary the status of restatement and summary as presentational relations is less than obvious. they were originally defined as subject matter relations (mann & thompson, 1987b, 1988). subsequently, on the rst website, they were re-categorized as presentational, but their definitions were not updated to reflect this change. with the restatement relation, the satellite restates the nucleus, and the satellite and nucleus are of comparable length. the summary relation is similar. the satellite restates the nucleus, but unlike restatement, the satellite is shorter than the nucleus. the defined effect of both of these relations is that the reader recognizes that the satellite restates the nucleus. if this were all there was to say on the matter, it would be difficult to see the intended effect as presentational – there would be no identifiable inclination in the reader that these relations increase. stede et al. (2017) argued that these relations should be treated as textual, and as with preparation, there is clearly a textual dimension to their use. however, going beyond the definitions, examining examples indicates that they are presentational. restatement and summary are used to reinforce a point the writer seeks to make. in the following example of restatement, from mann and thompson (1987b), the writer seeks to emphasize the claim that a clean car makes a statement about the character of the owner, and implicitly, that an unkempt car does too: n: a well-groomed car reflects its owner s: the car you drive says a lot about you. this is not simply a case of organizing the text. the restatement drives home the point. summary can be used in the same way. by recapping a complex text with a short summary, the writer can reinforce the claim. used in this way, restatement and summary support reader comprehension, belief, or action in a manner similar to background and preparation, and thus they adopt the same logical form. 3.2 subject matter relations subject matter relations are relations whose intended effect is that the reader will recognize the relation in question (mann & thompson, 1987b). unlike the presentational relations, they involve no act of persuasion. the locus of effect resides in both the nucleus and the satellite. because the acceptability of the nucleus and the satellite is not in question, the minimal logical form for a subject matter relation could be considered to be the conjunction of the satellite and nucleus, s  n. but the asymmetry of the subject matter relations indicate logical information beyond conjunction. the intended effect is a recognition of some feature of the organization of the subject matter (mann & matthiessen, 1990), and the logical forms corresponding to the subject matter relations should capture such features. for example, in a causal relation, the dependency between cause and effect may be represented as an implicative relationship. however, unlike the presentational relations, the point here is not that the implication is needed to establish the acceptability of the consequent. the antecedent and the consequent are both presumed to hold, and the relation indicates that they are in an implicative relationship. such a relation would take the form, s  n  (s → n). from a purely logical perspective, the implicative portion of this form is redundant: s → n will follow, by definition, from s  n. but the implicative characterizes the inferential processes implicit within the intended effect. the relational proposition corresponding to the logical form of an implicative subject matter relation is true only when both s and n are true and the implication holds. denial of the implication would be an indication of a failure to achieve the writer’s intended effect. in other words, the text would be logically incoherent. this is explored in detail in sections 4 and 5. the following sections provide definitions of the logical forms for each of the subject matter relations. potter 93 3.2.1 causal relations the term cause can be defined broadly as any situation or activity that brings about a change in some other situation or activity. in that sense, a causal relation is a relation that holds between parts of text, in which one part presents a cause and another part the effect. from this perspective, there are several subject matter relations that describe cause and effect relationships. mann and thompson (1987b) originally identified the causal relations as consisting of volitional cause, non-volitional cause, volitional result, non-volitional result, and purpose. they used the concept of volitionality as a means for distinguishing causes that involve the action of an agent, typically but not necessarily an action of a person, from causes that involve consequentiality without a chosen outcome. some analysts have elected to disregard this distinction between volitionality and non-volitionality. nicholas (1994) argued that since volitionality modifies only the nucleus or the satellite, but not the relationship between the two, it is rhetorically irrelevant. in their rst annotation guidelines, stede et al. (2017) ignored volitionality altogether and subsumed the mann and thompson relations under the more general categories of cause and result. because volitionality has no bearing on the logical analysis, i have followed their practice in this discussion of causal relations. stede et al. (2017) also specified that the rst relations condition, otherwise, and unless relations should also be treated as causal. these relations, along with purpose, apply to hypothetical situations. with condition, the satellite presents a hypothetical, future, or otherwise unrealized situation upon which the nucleus depends. with otherwise, realization of the nuclear situation prevents realization of the satellite situation. with the unless relation, realization of the nucleus occurs provided that realization of the satellite does not. with purpose, the nucleus is an activity that will lead to the situation identified in the satellite. two additional relations that can be considered as causal are means and unconditional. means specifies that the satellite presents a method or instrument which tends to make realization of the activity presented in the nucleus more likely. in this regard it is akin to enablement, except the presented activity is not intended to increase an inclination in the reader. the satellite presents a situation that brings about the activity presented in the nucleus. the unconditional relation asserts that the nucleus has no dependency on the satellite. it is a relation of negative causality – that is, unconditional negates any causal relation between the satellite and the nucleus. the following provides logical interpretations for each of these causal relations. 3.2.1.1 cause and result cause specifies that the satellite caused the nucleus, and result specifies that the nucleus caused the satellite. they are both subject matter implicative. some activity or situation arises as a result of some other activity or situation. their relational propositions assert that there is a causal relationship between the satellite and nucleus. here is an example of cause excerpted from the salvage and archaeology analysis (available on the rst website): s: since the objects in a [ship]wreck represent a single moment in time, n: they provide better chronological information than even the most carefully excavated terrestrial site. if the analyst had viewed this relation as argumentative, if there were some doubt that objects in a shipwreck provide better chronological information, then the relation would have been coded as evidence. but in the analyst’s judgment, this is not the case. the intended effect is that the reader recognizes that the satellite is the cause for the nucleus. implication is needed to establish not the truth of the consequent, but that the antecedent and the consequent are in an implicative relationship. the logical form of the cause relation documents this implicative relationship: reasoning between the lines 94 s  n  (s → n) the result relation is similar, except now the nucleus is the cause: s  n  (n → s). in this example from taboada and renkema (2008), the speaker’s difficulty in smiling at jesus arises as a result of the demands he places on her. hence the satellite follows from the nucleus: s: … i find it hard sometimes to smile at jesus. n: he can be very demanding. 3.2.1.2 means the means relation specifies that the satellite presents a method or instrument which tends to make realization of the activity presented in the nucleus more likely. in this analysis from taboada (2006), means represents features of a method reflected in the design of the human body: n: ... the visual system resolves confusion s: by applying some tricks that reflect a built-in knowledge of properties of the physical world. thus, means is a special case of cause, and it has the same subject matter implicative form. 3.2.1.3 condition with condition, realization of the satellite results in realization of the nucleus. because the realization of the nucleus depends upon realization of the satellite, without realization of the satellite there will be no realization of the nucleus. in this example from mann and thompson (1987b), the satellite specifies the conditions under which employees are urged to complete new beneficiary designation forms: n: employees are urged to complete new beneficiary designation forms for retirement or life insurance benefits s: whenever there is a change in marital or family status. a change in marital or family status is the condition under which employees are urged to complete new beneficiary designation forms. the reader recognizes that the realization of the nucleus depends on the realization of satellite. as long as the satellite remains unrealized, so will the nucleus. thus the relation is one of material equivalence (s  n). 3.2.1.4 purpose purpose is logically similar to condition. with the purpose relation the situation presented in the satellite is to be realized through the activity presented in the nucleus. the satellite is an as yet unrealized result of the nucleus. in this example from mann and thompson (1987b), s: to see which syncom diskette will replace the ones you're using now n: send for our free "flexi-finder" selection guide and the name of the supplier nearest you. the purpose of sending for the selection guide is to determine the appropriate syncom diskette. without the activity identified in the nucleus, the situation identified in the satellite is without potter 95 purpose and will not be realized. the logic of purpose is not only n → s, but also  n →  s. so the logic is material equivalence, n  s. 3.2.1.5 unconditional unconditional has received little attention in the literature, and some studies suggest that it occurs infrequently (e.g. iruskieta, 2014; rubin & vashchilko, 2012). this may be due in part to difficulties in differentiating it from concession. like concession, although the satellite could conceivably affect the realization of the nucleus, the nucleus is not dependent on the satellite. figure 5 shows an example from an article on pre-performance anxiety (brooks, 2013). the nucleus holds, regardless of the value of the satellite: (s   s) → n. figure 5. unconditional relation 3.2.2 otherwise and unless the otherwise and unless relations are similar to one another. with otherwise, realization of the nuclear situation prevents realization of the satellite situation. in preventing realization of the satellite situation, the nucleus effectively negates it. in this example from the rst website, n: project leaders should submit their entries for the revised brochure immediately. s: otherwise the existing entry will be used. only by submitting their new entries can project leaders prevent the use of existing entries, and failure to submit will affirm the satellite. consequently, otherwise indicates an exclusive disjunction, i.e., (n  s), or ((n  s)   (n  s)). either n will hold, or s will hold, but not both. unless is similar to otherwise, except now the realization of the satellite would prevent realization of the nucleus, rather than the nucleus preventing the satellite. 3.2.3 solutionhood with solutionhood, the satellite poses a problem, for which the nucleus presents a solution. the satellite may be presented as a question, request, problem, or other expressed need. the logic of solutionhood is not to be found in the potentially interrogative nature of the satellite. the logic can be derived from its rst definition and corresponding relational proposition. to explain this, it is useful to examine an analogy between interrogatives and propositional functions, proposed by hintikka (2007). a propositional function is an “expression containing an undetermined constituent,” as defined by whitehead and russell (1910, p. 92). because of this undetermined constituent (or free variable), a propositional function is not itself a proposition, so it has no truth value and cannot serve as an element within an argument. a propositional function defines the form of a proposition without asserting the proposition itself. the propositional function x reasoning between the lines 96 specifies some predicate  of x without indicating whether there might be any x. in this sense, according to hintikka, a propositional function is a specification of a query. but on the basis of such a query alone, it cannot be known whether any result will be forthcoming. the query could return a null set. the query can be made more specific by binding the variable x. this binding, or existential quantification, is expressed as (x)x, which asserts that for at least one instance of x, x is the case, although the identity of x is unspecified. as such it is a proposition. but to specify that there is an identifiable x for x, the variable x must be replaced with a constant, v. since we know that (x)x, we can infer that there is some constant that will support this. this is accomplished using existential instantiation: (x)x  v from the query (x)x, the answer, v is inferred. with solutionhood, the writer has committed to providing a solution. the problem is not simply posed and left for the reader to resolve. the problem presented by the satellite is an indispensable part of a relation that must be satisfied by the solution presented as the nucleus. the presentment of the problem, whether interrogative or otherwise, obliges the writer to provide an answer. otherwise there is no solutionhood. the questions, requests, or problems posed in a solutionhood relation are not to be taken out of context. the writer is providing a problem-solution matching set. the satellite is the specification of a query. since the writer is committed to providing a solution, the satellite is existentially quantified. the proposition (x)x represents the satellite of the solutionhood relation, where  is the problem, and x is the undetermined solution. that is, it is asserted that there is a solution to the problem, but as yet the solution has not been identified. the solution is identified in the nucleus, where x becomes v, thanks to existential instantiation. therefore, solutionhood is an inferential relation between the satellite and the nucleus. since (x)x and v are propositions, the inference may be restated propositionally as s  (s → n) given the assertion of s, n. this is not to claim that the answer to a question can be deduced from the question. rather that, in a coherent text, the posing of a problem can be understood to anticipate the solution, much as a database query uses a subset of first order logic to implicate the result set (sowa, 2000). an example will make this more concrete. this example is from mann and thompson (1985): s: i'm hungry. n: let's go to the fuji gardens. the problem is that the speaker is hungry, and the solution is that going to fuji gardens (presumably a restaurant) solves the hunger problem. this relation makes some presuppositions about hunger and the means by which the discomfort that attends it may be relieved. it is the desire to relieve the discomfort associated with hunger that gives the first statement its status as a problem, and it is the relief from discomfort resulting from the act of eating that gives the second statement its status as the solution to the problem. it is through recognition of these presuppositions that the statements achieve coherence, and it is on this level that the inferential nature of the solutionhood relational proposition becomes apparent. expressed as a propositional function, the problem to be solved is some as yet unidentified source of relief (x). posed as the satellite to a potter 97 solutionhood relation, the problem is existentially quantified, expressed as (x)relief(x). the nucleus, using existential instantiation then substitutes fuji gardens (v) as the source of relief: (x)relief(x) → relief(fujigardens) this is not the only effort to define solutionhood in terms amenable to inferential interpretation. abelen, redeker, and thompson (1993) treat solutionhood as an interpersonal relation, which they define as a relation used for persuasion, and they categorized it alongside presentational relations such as evidence, justify, and motivation. similarly, azar (1999) cited the persuasive power of solutionhood in his study of the rhetorical structure of argumentative texts. when leveraging the satellite to accept the nucleus, there would be an inferential dimension in solutionhood. and while persuasiveness may not be a factor in the fuji gardens example, it can be found in other cases, such as the syncom example used by mann and thompson (1987b), as well in azar’s analysis. but persuasiveness is not an essential feature of solutionhood, and for the purpose of discovering its underlying logic, it is unnecessary, as has been demonstrated here. 3.2.4 elaboration with elaboration, the satellite presents additional detail about the situation or some element of subject matter which is presented or inferentially accessible in the nucleus. mann and thompson (1987b) identified six subtypes of the relation, as listed in table 1. as noted by taboada and mann (2006b), any one of these subtypes is sufficient for the elaboration relation to hold, and in theory it could be possible to treat them as separate relations. and while carlson and marcu (2001) did just that, for the purpose of logical interpretation this is unnecessary. the common characteristic among the subtypes is that, as defined, the satellite presents additional detail about the nucleus. nucleus satellite set member abstraction instance whole part process step object attribute generalization specific table 1. elaboration subtypes whether through set membership, instantiation, mereology, process decomposition, object attribution, or specification, the satellite is logically implicit in the nucleus. were this subsumptive relationship inaccessible to the reader, the relevance of the satellite to the nucleus would be lost: the inference is necessary for the reader to recognize that the situation presented in the satellite provides additional information for the nucleus. in the example shown in figure 6, in segment 1, the writer identifies an ongoing activity. as analyzed by mann and thompson (1987a), segment 2 elaborates this by identifying a type of activity (set-member), and unit 3 further details this activity (generalization-specific). each of the nuclei progressively circumscribe the satellites of their relations. thus elaboration presents a subject matter implication, with the nucleus as antecedent and the satellite as consequent: s  n  (n → s). reasoning between the lines 98 figure 6. elaboration nuclei circumscribe their respective satellites 3.2.5 evaluation and interpretation evaluation and interpretation involve assessment of the situation presented in the nucleus. evaluation assesses the situation presented by the nucleus with respect to the writer’s positive regard. interpretation relates the nuclear situation to a framework of ideas not involved in the nucleus and not concerned with the writer’s positive regard (mann & thompson, 1987b). in both relations, the assessment presented in the satellite is a plausible inference made on the basis of the information provided in the nucleus. to the extent that the inference offered by the satellite is plausible, the relation is implicative: ((n → s)  n) → s. the implicative nature of these relations has been hinted at by several other researchers. both stent (2000) and stede (2008) note that evaluation involves the subjective judgment of the writer, and earlier abelen et al. (1993) categorized both evaluation and interpretation as interpersonal relations. deriving their definition of the term interpersonal from halliday and hasan (1976), abelen et al. view these relations as being notable for their persuasive power. what distinguishes evaluation from the presentational relations is that, while the writer presents an expression of positive regard, the effect is not explicitly directed towards any prospective inclination in the reader. the rst definition of the relation obscures this characteristic, because the effect occurs not in the nucleus but in the satellite. in an excerpt from mann and thompson’s (1987b) syncom analysis, shown in figure 7, the assessment presented in the satellite follows as a result of the information presented in the nucleus. an evaluation that did not follow from its nucleus would be incoherent. although interpretation does not deal in positive regard, the logical interpretation is identical to evaluation. unless the interpretation follows from its basis, it will be incoherent. just as with evaluation, the assessment presented in the satellite must be a plausible inference made on the basis of the information provided in the nucleus. 3.2.6 circumstance the satellite of the circumstance relation sets the framework within which the reader is intended to interpret the nucleus. in this respect, circumstance is similar to background, as noted by stede et al. (2017). the framework imposed by the satellite functions as a constraint on the nucleus, delimiting it with respect to temporal, spatial, or other considerations. figure 8 shows an analysis with two circumstance relations, one spatial, the other temporal. sometimes an instance of circumstance includes both spatial and temporal constraints, as in the example from the rst website shown in figure 9. this serves to illustrate the constraining effect of the relation. potter 99 figure 7. the evaluation relation figure 8. two circumstance relations, one spatial the other temporal figure 9. one circumstance relation with both temporal and spatial characteristics although some researchers have found circumstance to be non-causal (e.g., carlson & marcu, 2001), there are instances that combine temporal or spatial constraints with causality. the following examples use causal circumstance, the first from mann and thompson (1987b) and the second from stede et al. (2017). in the visitors fever example, the writer attributes the cause of the ailment to a visit with relatives in the midwest. in the second example, deregulation of the electricity market provides the reason for squeezing suppliers: reasoning between the lines 100 n: probably the most extreme case of visitors fever i have ever witnessed was a few summers ago s: when i visited relatives in the midwest. s: when veag came under pressure because of the deregulation of the electricity market, n: they compensated for this by squeezing their suppliers. while circumstance is similar to elaboration with respect to the satellite providing additional information about the nucleus, the nature of the additional information is quite different. rather than the nucleus subsuming the satellite, the effect of the additional information provided by the satellite circumscribes the scope of the nucleus. the nucleus holds within the domain specified by the satellite. the satellite specifies a condition under which the nucleus holds. the logical form of circumstance is subject matter implicative: s  n  (s → n). 3.3 multinuclear relations most of the multinuclear relations are logical conjunctions, with exceptions being disjunction and multinuclear-restatement. disjunction, as the name suggests, is disjunctive. the elements of multinuclear-restatement are materially equivalent assertions. while the conjunctive nature of conjunction, contrast, joint, and list should seem readily apparent, sequence might seem serially inferential. with sequence, a succession relationship between the situations is presented. mann and thompson suggest recipes as good examples. there is a serial dependency among the nuclei. the dependency could be temporal, process, or any other series. to the extent that this is the case, failure of any nucleus may render the successors unreachable. but this means that for a sequence to hold, all of its members must hold. hence sequence may be defined as the logical conjunction of its members. mann and thompson (1985) claimed that relational propositions are always present in multisentence texts. and yet none of their work on the topic mentions multinuclear relational propositions. i propose that they would take the form of list(n1, n2, n3…), which would transform into the logical expression n1  n2  n3… and disjunction would follow the same pattern, using the disjunction logical connector. multinuclear-restatement seems to occur as pairs of biconditional nuclei, although in principle there could be larger sets. 4 negating relational propositions the logical forms defined in section 3 can be used to construct logical interpretations of a text. this is accomplished by restating the rst structure of the text as a nested relational proposition and mapping the predicates comprising the proposition to their respective logical forms. this suggests the possibility of analyzing texts as truth functional expressions. if so, we would expect that these expressions would be true or false, and that the determinants of their truth value would be indicative of the logical coherence of the discourse. this presupposes the ability to negate logical forms. and yet, based on the definitions presented in section 3, it would seem that all relational propositions must be true. presentational relational propositions are true because they are tautologies. subject matter relational propositions are true because their satellites and nuclei are non-controversial. and while this observation—that relational propositions are by definition true—follows from the logical forms, and while these forms are in turn derived from their respective rst definitions, this clearly conflicts with both intuition and experience. rst analyzability is no guarantee of truth or logical coherence. as mann and thompson (1986b) and marcu (1996) have shown, it is possible to assign plausible rst relations for incoherent discourse. the apparent difficulty arises as a result of a disparity between the writer’s intended effect and the actual effect on the reader. the rst relations are defined to indicate the writer’s intended potter 101 effect. this is consistent with theories of rhetoric in general (heuboeck, 2009), but when reduced to its logical essence manifests itself as self-affirming reasoning, since it assumes that the effect has been achieved. thus it is necessary to specify not only the forms of intended effect but their truth conditions as viewed from a skeptical perspective as well. the perspective of a skeptical reader can be used to identify potential points of failure within the logical forms. the potential points of failure can be used in constructing rules of negation for use in reasoning over logical forms. the fundamental differences between presentational and subject matter relations require separate approaches for accomplishing this. the intended effect of presentational relations, to increase some inclination in the reader, is identified in the nucleus. if the writer is successful, i.e., if the reader accepts the effect, the reader accepts the situation as presented in the nucleus. the logical forms of the presentational relations build-in this acceptance. this is evident in the implicative structure of these relations. presentational relational propositions are logically valid arguments. the premises of the argument are located in the left hand side (lhs) of the logical form, to the left of the outermost implication. from these premises, the right hand side (rhs) or conclusion, consisting of the nucleus, is deductively inferred. thus the tautological structure of these relations presumes success for the writer’s intended effect. while presenting the logical forms from the writer’s perspective is consistent with the rst definitions, for the skeptical reader this amounts to begging the question. for the skeptical reader, it is the soundness of the lhs that is of interest. although negation of the lhs will not necessarily negate the rhs, an lhs once negated provides no functional support for the intended effect. since the burden of persuasion is on the lhs, its negation is sufficient for rejection of the rhs. were this not the case, the relation would be subject matter rather than presentational. the lhs of a presentational relation consists of a major premise and a minor premise, as identified in table 2. denial of the lhs of the form can consist of the negation of either or both of the premises, or of the rhs itself. relation lhs rhs major premise minor premise conclusion antithesis (s  n)  s n concession ( (s →  n) → n)  (s →  n) n evidence, justify, motivation, enablement, background, preparation, restatement, summary (s → n) s n table 2. presentational relations summary the lhs of the antithesis relation is ((s  n)   s), which implies the rhs, n, such that ((s  n)   s) → n. the lhs includes several options for disputation by the skeptical reader. the negation of the lhs,  ((s  n)   s) follows whenever s holds (the minor premise is negated), n does not hold, or when the major premise is negated: (s   n   (s  n)) →  ((s  n)   s) for concession to succeed it is necessary that the satellite be compatible with the nucleus. the concession lhs is negated whenever the major or minor premise is negated, or when n does not hold. negating the minor premise challenges the compatibility between the satellite and the nucleus: reasoning between the lines 102 ( ( (s →  n) → n)  (s →  n)   n) →  (( (s →  n) → n)   (s → n)) following the same approach, evidence is negated when the major or minor premise is negated, or when the nucleus is negated. negating the major premise denies the implicative relationship between the satellite and the nucleus. negating the minor premise (s) presents a valid argument for negation of the lhs. these same rules of negation also apply to the justify, motivation, enablement, background, preparation, restatement, and summary relations:  (s → n)   s   n →  ((s → n)  s) as defined in section 3, several of the subject matter relations are subject matter implicatives. for some of these, the implication flows from satellite to nucleus (cause, means, circumstance, solutionhood), and for others the flow is from nucleus to satellite (result, elaboration, evaluation, interpretation), but they all share the same general logical form, consisting of a conjunction of the satellite and nucleus and an implicative relation between the two. the negation rule for the s → n group is ( s   n   (s → n)) →  (s  n  (s → n)) and for the n → s group, the rule is ( s   n   (n → s)) →  (s  n  (n → s)) since the logical forms of condition and purpose are material equivalence (s  n), negation requires that either the satellite be false while the nucleus is true, or the nucleus be false while the satellite is true: ((s   n)  ( s  n)) →  (s  n) unconditional asserts that the nucleus is not dependent on the satellite, in the sense that n, regardless of the value of s, or (s   s) → n. negation of n suffices to negate the relation:  n →  ((s   s) → n) otherwise and unless, both being exclusive disjunctions, are negated when either both the satellite and the nucleus are true or when both are false: ((n  s)   (n  s)) →  ((n  s)   (n  s)) 5 the logic of relational propositions the logical forms and rules of negation can be used in exploring the coherence of texts. more precisely, these forms and rules can be used to explore textual reasoning as reflected in an rst analysis. any rst analysis can be restated as a nested relational proposition, and this nested relational proposition can be restated as a logical expression. as a simple example, the music day analysis from mann and thompson (1987b), shown in figure 10, can be restated as this relational proposition: justify(concession(2,3), 1). the concession relation is the satellite of the justify relation. because the logical form of justify is modus ponens (((s → n)  s) → n), the nested potter 103 concession(2,3) appears twice, first as the conditional part of the major premise and again as the minor premise, as shown in figure 11. figure 10. the music day analysis figure 11. music day logic since both concession and justify are presentational relations, for the purpose of this analysis, only the lhs of the forms are of interest. here is the nested relational proposition reduced to the lhs of each relation: ((((( (2 →  3) → 3)   (2 →  3)) → 3) → 1)  ((( (2 →  3) → 3)   (2 →  3)) → 3)) to this we can selectively apply rules of negation. for example, if the reader challenges the concession relation by denying the compatibility between its nucleus and satellite, the result is negation of the entire expression: (2 →  3) →  ((((( (2 →  3) → 3)   (2 →  3)) → 3) → 1)  ((( (2 →  3) → 3)   (2 →  3)) → 3)) the negation is systemic because the negated concession is the minor premise of the justify modus ponens argument. within an rst analysis, when a rule of negation can be plausibly applied to a relational proposition, this is an indication of a potential misalignment either between the logic and the rhetorical coherence of the text or between the logic and the rst analysis. thus the effects of faulty reasoning on the part of the writer and problematic rst encodings on the part of the analyst can be examined using the logic of relational propositions. the following presents a logical analysis of a more complex example. reasoning between the lines 104 the rst analysis to be considered is “bouquets in a basket,” shown in figure 12. according to mann and thompson (1987a), the text is expository. the rst analysis consists primarily of subject matter relations. the text begins with the statement, there is a gardening revolution going on (1). units 2-3 elaborate on the revolution mentioned in 1. the structure consisting of 1-3 is the satellite of a background relation. the nucleus of the background consists of a purpose relation. mann and thompson mention that unit 4, the satellite of a purpose relation, presents a possible goal, of creating "your own" bouquet, with segments 5-8 providing the method. figure 12. rst analysis of the bouquets in a basket text the logic of the text, as implicit in this rst analysis, may be summarized as follows: as the nucleus of an elaboration relation, the statement that there is a gardening revolution going on (1) is the antecedent of an implicative relation with units 2-3. by definition of elaboration, that people are planting flower baskets with living plants, and mixing many types in one container for a full summer of floral beauty follows from the assertion that there is a gardening revolution going on: (1  (2  3  (2 → 3))  (1 → (2  3  (2 → 3)))) by definition of background, from 1-3, it follows that the reader may create a bouquet of flowers (4) using the method presented in 5-8. and by definition of purpose, the method presented in text span 5-8 is materially equivalent with 4: (((5  6  (5 → 6))  (5  (7  8  (7 → 8))  (5 → (7  8  (7 → 8)))))  4) since background is a presentational relation, and it is an overarching relation for this text, the logical expression representing the text is this tautology: ((((1  (2  3  (2 → 3))  (1 → (2  3  (2 → 3)))) → (((5  (7  8  (7 → 8))  (5 → (7  8  (7 → 8))))  (5  6  (5 → 6)))  4))  (1  (2  3  (2 → 3))  (1 → (2  3  (2 → 3))))) → (((5  (7  8  (7 → 8))  (5 → (7  8  (7 → 8))))  (5  6  (5 → 6)))  4)) thus, according to the rst analysis, the text is logically coherent. this would be expected of any rst analysis, since the logical forms arise from the rst definitions. the logic becomes more potter 105 interesting when the rules of negation are applied. there are a few inferences within this rst analysis that may seem problematic to the skeptical reader. first, as a subject matter relation, elaboration assumes that the reader will accept both the satellite and nucleus. if either is questionable, some other relation must be applied. taken on its own, there is a gardening revolution going on (1) will seem puzzling to readers if the hyperbolic use of the term revolution is not obvious. in that case, substantiation would be needed. if for lack of substantiation the statement is denied, the consequences are systemic, and the lhs of the background relation is negated, resulting in this tautology: (1 →  (((1  (2  3  (2 → 3))  (1 → (2  3  (2 → 3)))) → (((5  (7  8  (7 → 8))  (5 → (7  8  (7 → 8))))  (5  6  (5 → 6)))  4))  (1  (2  3  (2 → 3))  (1 → (2  3  (2 → 3)))))) however, the lack of substantiation is not within the text, but emerges only when accessing the logic of the text through the intermediation of the rst analysis. that is, the logic of the elaboration relation is inconsistent with the logical coherence of the text. the satellite, rather than following from the nucleus, is needed to give credence to the claim asserted by the nucleus. this can be satisfied using evidence rather than elaboration for the relation between 1 and 2. a second problematic inference is found in the purpose relation. it is significant that the satellite of this relation, segment 4, addresses the reader directly: …create your own "victorian" bouquet of flowers. without some indication of why and how the reader should carry out such an activity, the background relation provides no support other than comprehension. moreover, the nucleus of the background is not 4, but the purpose relation itself, containing the structure of nested elaboration relations (5-8) for which 4 is satellite. the purpose of 5-8, to create your own "victorian" bouquet of flowers, must stand on its own. this places 4 at risk of negation. if the reader is disinclined to create a bouquet and denies 4, the purpose relation fails. negation of the nucleus implies negation of the satellite, and this results in negation of the background relation. concerns such as these are opportunities for assessing the structure of the text. an alternative analysis, one that would align the logic of the text and its rst structure, would be to code segment 4 as the nucleus of a motivation/enablement schema, as shown in figure 13. segments 1-3 are intended not merely to impart comprehension of what follows, but to increase the reader’s desire to perform the action, i.e., to create a bouquet. by offering the reader the prospect of a full summer of floral beauty, segments 1-3 are intended to increase the reader’s desire to create a basket, and segments 5-8 increase the reader’s ability to do so, by providing a method involving choices of shapes, sizes, forms, and colors of the flowers, height of the plants, leaf textures, and leaf colors. with this revised analysis, segment 4 emerges as the locus of effect for the text, and as the logical consequent as well. without alignment between rst structure with the logic of relational propositions, the inferences indicated by the original rst analysis lack plausibility and are subject to negation. by aligning the rst structure with its logical inferences, the organization of the text comes into focus: segment 1 follows from text span 2-3, and segment 4 follows from 1-3, representing the motivation part of the text. conjunctively, 4 also follows from 5-6, representing the enablement part of the text. here is the motivation conjunct: (((((((2  3  (2 → 3)) → 1)  (2  3  (2 → 3))) → 1) → 4)  ((((2  3  (2 → 3)) → 1)  (2  3  (2 → 3))) → 1)) → 4) and this is the enablement conjunct: reasoning between the lines 106 (((((5  6  (5 → 6))  (5  (7  8  (7 → 8))  (5 → (7  8  (7 → 8))))) → 4)  ((5  6  (5 → 6))  (5  (7  8  (7 → 8))  (5 → (7  8  (7 → 8)))))) → 4) in their commentary on this analysis, mann and thompson seemed to have recognized this, as indicated by their observation that 4 presents the possible goal, and that segments 5-8 provide the method. but they did not choose to code the text as such. an rst analysis that cannot pass the test of logic is an analysis in need of attention. as demonstrated here, the revisions indicated by the logic of relational propositions need not be merely a matter of tweaking a few relations. it may provide an alternative perspective on the text. in this example, the difference is between expository and argumentative. that the logic of relational propositions can be used to assess an existing rst analysis shows that it can also be used in the development of new analyses as well. the decision process would be the same: for any structure under consideration, identify the corresponding logical expressions and their consequences. if these lead to inconsistencies, either the text is problematic or the structure is in need of review. figure 13. alternative of analysis of bouquets in a basket 6 conclusion the relationship between coherence relations and logic has been an active area of research at least since hobbs (1979, 1985) used logic to define a set of relations for use in an automated inference system. using this approach, relations could be represented in a manner akin to (but not identical with) the predicate calculus (hobbs, 1979). there have been numerous other studies in the logic of coherence relations (e.g., danlos, 2008; gonzález & ribas, 2008; groenendijk, 2009; marcu, 2000; sanders et al., 1992; wong, 1986). much of the recent research in this area has been influenced by the work of asher and lascarides (2003), with their development of segmented discourse representation theory (sdrt). sdrt is based in part on discourse representation theory (drt), a theory of dynamic semantics that attempts to account for the context dependence of meaning, based on the observation that how a sentence is interpreted depends on what has occurred previously within the discourse (kamp & reyle, 1993; kamp, van genabith, & reyle, 2011). that is, the discourse context is dynamically modified with each successive utterance, and this updated context informs the interpretation of whatever comes next. building on drt, sdrt uses dynamic semantics to generate interpretations of discourse. updates to the interpretation are made using a glue logic which identifies appropriate rhetorical relations for linking new information to old, assigning a compositional and dynamic semantic interpretation for the discourse. the aims of this paper have been more modest than those of sdrt. in this paper i have demonstrated the use of relational propositions as a means for investigating texts as expressions in propositional logic. while sdrt defines a set of rules for deriving rhetorical structures from potter 107 logical forms, the logic of relational propositions provides a set of rules for deriving logical forms from rhetorical structures. while sdrt is a theory of dynamic discourse interpretation, the logic of relational propositions is based on a static logical interpretation of intended effect. thus, the two theories are each motivated by rather different goals, with rather different results. this paper has shown that any text analyzable using rst can be restated as an expression in propositional logic. relational propositions play a key role in this. by providing a theoretical basis for treating rst structures as assertions, relational propositions provide a bridge between rst analyses and propositional logic. in order to apply the logic of relational propositions to complete texts, rather than individual relations, it was necessary to extend the theory, so that any multiunit rst span can be restated as a nested relational proposition. for each of the rst relations there is a defined logical form reflecting the logical relationship between the satellite and the nucleus. within a nested relational proposition, the forms may be substituted for each of the constituent relations. mapping the forms into a nested relational proposition yields a logical expression of the text. the reduction of texts to logical forms reveals that logical inference is pervasive within coherent text. the paper has also defined a set of rules for negation of logical forms. the logical forms and rules of negation can be used to examine the logical coherence of texts. the status of the logical coherence of a text can be used to support rst analysis. because there is a correspondence between logical coherence and the functional relationships of rst, an rst analysis that cannot pass the test of logic is indicative either of a problematic analysis or textual incoherence. there are several potential applications of this research. if integrated with computational methods for generating rst analyses (e.g., corston-oliver, 1998; hernault, prendinger, duverle, & ishizuka, 2010; pardo, nunes, & rino, 2004; soricut & marcu, 2003), the method presented here could lead to useful tools for scalable analysis of large text collections. some other potential uses include contributions to knowledge representation, automated reasoning, controlled natural languages, cross-document analysis (cardoso, jorge, & pardo, 2015; radev, 2000), and integration with research in semantic equivalence, entailment, and knowledge extraction (androutsopoulos & malakasiotis, 2010; gangemi, 2013; zhang & patrick, 2005). some areas for future study include development of a computational framework for the study of the logic of rst analyses, studies in the interrelationship between logic and rhetoric, a more in-depth look at multinuclear relations, and investigation of various genres of discourse using the methods described in this paper. references abelen, e., redeker, g., & thompson, s. (1993). the rhetorical structure of us-american and dutch fund-raising letters. text, 3, 323-350. androutsopoulos, i., & malakasiotis, p. (2010). a survey of paraphrasing and textual entailment methods. journal of artificial intelligence research, 38, 135-187. asher, n., & lascarides, a. (2003). logics of conversation. cambridge, uk: cambridge university press. azar, m. (1999). argumentative text as rhetorical structure: an application of rhetorical structure theory. argumentation, 13(1), 97-114. boole, g. (1854). an investigation of the laws of thought on which are founded the mathematical theories of logic and probabilities (dover 1958 reprint ed.). new york: macmillan. brooks, a. w. (2013). get excited: reappraising pre-performance anxiety as excitement. journal of experimental psychology, 143(3), 1144-1158. cardoso, p. c. f., jorge, m. l. r. c., & pardo, t. a. s. (2015). exploring the rhetorical structure theory for multi-document summarization congreso de la sociedad española para el procesamiento del lenguaje natural (vol. xxxi). alicante, spain: sociedad española para el procesamiento del lenguaje natural. reasoning between the lines 108 carlson, l., & marcu, d. (2001, september). discourse tagging reference manual. retrieved from ftp://ftp.isi.edu/isi-pubs/tr-545.pdf copi, i. m. (1967). symbolic logic. new york: macmillan. corston-oliver, s. h. (1998). computing representations of the structure of written discourse. (dissertation), university of california, santa barbara, ca. danlos, l. (2008). strong generative capacity of rst, sdrt and discourse dependency dagss. in a. benz & p. kühnlein (eds.), constraints in discourse (pp. 69–95). amsterdam: benjamins. forbus, k. d. (2016). software social organisms: implications for measuring ai progress. ai magazine, 37(1), 85-90. gangemi, a. (2013). a comparison of knowledge extraction tools for the semantic web. in p. cimiano, o. corcho, v. presutti, l. hollink, & s. rudolph (eds.), the semantic web: semantics and big data (pp. 351-366). berlin, heidelberg: springer. gonzález, m., & ribas, m. (2008). the construction of epistemic space via causal connectives. in i. kecskes & j. mey (eds.), intention, common ground and the egocentric speaker-hearer (pp. 127-149). berlin: de gruyter. groenendijk, j. (2009). inquisitive semantics: two possibilities for disjunction. in p. bosch, d. gabelaia, & j. lang (eds.), logic, language, and computation (pp. 80-94). berlin, heidelberg: springer berlin heidelberg. grosz, b., & sidner, c. (1986). attention, intentions, and the structure of discourse. computational linguistics, 12(3), 175-204. halliday, m. a. k., & hasan, r. (1976). cohesion in english. london: longman. hernault, h., prendinger, h., duverle, d. a., & ishizuka, m. (2010). hilda: a discourse parser using support vector machine classification. dialogue and discourse, 1(3), 1-33. heuboeck, a. (2009). some aspects of coherence, genre and rhetorical structure – and their integration in a generic model of text. language studies working papers, 1, 35-45. hintikka, j. (2007). socratic epistemology: explorations of knowledge-seeking by questioning. new york: cambridge university press. hobbs, j. r. (1979). coherence and coreference. cognitive science, 3, 67-90. hobbs, j. r. (1985). on the coherence and structure of discourse (csli-85-37). stanford, ca: center for the study of language and information, stanford university. retrieved from http://www.isi.edu/~hobbs/ocsd.pdf hovy, e. h. (1990). parsimonious and profligate approaches to the question of discourse structure relations. proceedings of the fifth international workshop on natural language generation. pittsburgh, pa: association for computational linguistics. hovy, e. h., & maier, e. (1993). parsimonious or profligate: how many and which discourse relations? marina del rey, ca: information sciences institute, university of southern california. iruskieta, m. (2014). a description of pragmatics rhetorical structure and its evaluation in computational linguistics. paper presented at the programa de pós-graduação em letras, maringa, brazil. kamp, h., & reyle, u. (1993). from discourse to logic: introduction to model-theoretic semantics of natural language, formal logic and discourse representation theory. dordrecht: kluwer. kamp, h., van genabith, j., & reyle, u. (2011). discourse representation theory. in d. gabbay & f. guenthner (eds.), handbook of philosophical logic, volume 15 (2nd ed., pp. 125-294). dordrecht: springer. knott, a., oberlander, j., o'donnell, m., & mellish, c. (2001). beyond elaboration: the interaction of relations and focus in coherent text. in t. sanders, j. schilperoord, & w. spooren (eds.), ftp://ftp.isi.edu/isi-pubs/tr-545.pdf http://www.isi.edu/~hobbs/ocsd.pdf potter 109 text representation: linguistic and psycholinguistic aspects (pp. 181-196). amsterdam: john benjamins. kolodny, n. (2007). how does coherence matter? proceedings of the aristotelian society (vol. 107, pp. 229-263). oxford: oxford university press. mann, w. c., & matthiessen, c. m. i. m. (1990). functions of language in two frameworks. marina del rey, ca: information sciences institute. mann, w. c., & thompson, s. a. (1985). assertions from discourse structure (technical report no. isi/rs-85-155). marina del rey, california: information sciences institute. mann, w. c., & thompson, s. a. (1986a). assertions from discourse structure. hlt '86: proceedings of the workshop on strategic computing natural language (pp. 257-270). morristown, nj: association for computational linguistics. mann, w. c., & thompson, s. a. (1986b). relational propositions in discourse. discourse processes, 9(1), 57-90. mann, w. c., & thompson, s. a. (1987a). rhetorical structure theory: a framework for the analysis of texts. ipra papers in pragmatics, 1, 1-21. mann, w. c., & thompson, s. a. (1987b). rhetorical structure theory: a theory of text organization (isi/rs-87-190). marina del rey, ca: university of southern california, information sciences institute (isi). mann, w. c., & thompson, s. a. (1988). rhetorical structure theory: towards a functional theory of text organization. text, 8(3), 243-281. mann, w. c., & thompson, s. a. (2000a). toward a theory of reading between the lines: an exploration in discourse structure and implicit communication. paper presented at the seventh international pragmatics conference, budapest, hungary. mann, w. c., & thompson, s. a. (2000b). two views of rhetorical structure theory. proceedings of the 10th annual meeting of the society for text and discourse. lyon, france. marcu, d. (1996). distinguishing between coherent and incoherent texts. the proceedings of the student conference on computational linguistics in montreal (pp. 136–143). marcu, d. (2000). the theory and practice of discourse parsing and summarization. cambridge, ma: mit press. nicholas, n. (1994). problems in the application of rhetorical structure theory to text generation. (masters thesis), university of melbourne, melbourne, australia. pardo, t. a. s., nunes, m. d. g. v., & rino, l. h. m. (2004). dizer: an automatic discourse analyzer for brazilian portuguese. advances in artificial intelligence – sbia 2004 17th brazilian symposium on artificial intelligence, sao luis, maranhao, brazil, september 29ocotber 1, 2004. proceedings. berlin: springer. radev, d. r. (2000). a common theory of information fusion from multiple text sources step one: cross-document structure proceedings of the 1st sigdial workshop on discourse and dialogue volume 10 (pp. 74-83). hong kong: association for computational linguistics. rubin, v. l., & vashchilko, t. (2012). identification of truth and deception in text: application of vector space model to rhetorical structure theory proceedings of the workshop on computational approaches to deception detection (pp. 97-106). avignon, france: association for computational linguistics. russell, b. (1906). the theory of implication. american journal of mathematics, 28(2), 159-202 sanders, t. j. m., spooren, w. p. m., & noordman, l. g. m. (1992). toward a taxonomy of coherence relations. discourse processes, 15, 1-35. soricut, r., & marcu, d. (2003). sentence level discourse parsing using syntactic and lexical information. proceedings of the 2003 conference of the north american chapter of the association for computational linguistics on human language technology volume 1 (pp. 149-156). edmonton, canada: association for computational linguistics. reasoning between the lines 110 sowa, j. f. (2000). knowledge representation: logical, philosophical, and computational foundations. pacific grove, ca: brooks/cole. stede, m. (2008). disambiguating rhetorical structure. research on language and computation, 6(3), 311-332. stede, m., taboada, m., & das, d. (2017). annotation guidelines for rhetorical structure. university of potsdam and simon fraser university. retrieved from http://www.sfu.ca/~mtaboada/docs/research/rst_annotation_guidelines.pdf stent, a. j. (2000). rhetorical structure in dialog. paper presented at the 2nd international natural language generation conference (inlg'2000), patras, greece. taboada, m. (2006). discourse markers as signals (or not) of rhetorical relations. journal of pragmatics, 38(4), 567-592. taboada, m., & mann, w. c. (2006a). applications of rhetorical structure theory. discourse studies, 8(4), 567-588. taboada, m., & mann, w. c. (2006b). rhetorical structure theory: looking back and moving ahead. discourse studies, 8(3), 423-459. taboada, m., & renkema, j. (2008). discourse relations reference corpus: rst analyses of 65 texts. burnaby, tilburg: simon fraser university and tilburg university. retrieved from http://www.sfu.ca/rst/06tools/discourse_relations_corpus.html thagard, p. (1989). extending explanatory coherence. behavioral and brain sciences, 12, 435502. thompson, s. a., & mann, w. c. (1986). a discourse view of concession in written english. proceedings of the second annual meeting of the pacific linguistics conference (pp. 435447). eugene, oregon: university of oregon, eugene. thompson, s. a., & mann, w. c. (1987). antithesis: a study in clause combining and discourse structure. in r. steele & t. threadgold (eds.), language topics: essays in honour of michael halliday, volume ii (pp. 359-381). amsterdam: john benjamins. whitehead, a. n., & russell, b. (1910). principia mathematica (vol. 1). cambridge, england: cambridge university press. wong, w.-k. c. (1986). a theory of argument coherence (tr86-29). austin, texas: artificial intelligence laboratory. zhang, y., & patrick, j. (2005). paraphrase identification by text canonicalization. proceedings of the australasian language technology workshop (pp. 160-166). sydney, australia. http://www.sfu.ca/~mtaboada/docs/research/rst_annotation_guidelines.pdf http://www.sfu.ca/rst/06tools/discourse_relations_corpus.html reasoning between the lines: a logic of relational propositions 1 introduction 2 theoretical background 2.1 rhetorical structure theory 2.2 relational propositions 2.3 propositional logic 3 a logical interpretation of rhetorical relations 3.1 presentational relations 3.1.1 antithesis 3.1.2 concession 3.1.3 evidence 3.1.4 justify 3.1.5 motivation and enablement 3.1.6 background 3.1.7 preparation 3.1.8 restatement and summary 3.2 subject matter relations 3.2.1 causal relations 3.2.1.1 cause and result 3.2.1.2 means 3.2.1.3 condition 3.2.1.4 purpose 3.2.1.5 unconditional 3.2.2 otherwise and unless 3.2.3 solutionhood 3.2.4 elaboration 3.2.5 evaluation and interpretation 3.2.6 circumstance 3.3 multinuclear relations 4 negating relational propositions 5 the logic of relational propositions 6 conclusion references evaluation and optimisation of incremental processors dialogue and discourse 2(1) (2011) 113-141 doi: 10.5087/dad.2011.106 evaluation and optimisation of incremental processors timo baumann timo@ling.uni-potsdam.de okko buß okko@ling.uni-potsdam.de department for linguistics university of potsdam germany david schlangen david.schlangen@uni-bielefeld.de faculty of linguistics and literature bielefeld university germany editor: hannes rieser abstract incremental spoken dialogue systems, which process user input as it unfolds, pose additional engineering challenges compared to more standard non-incremental systems: their processing components must be able to accept partial, and possibly subsequently revised input, and must produce output that is at the same time as accurate as possible and delivered with as little delay as possible. in this article, we define metrics that measure how well a given processor meets these challenges, and we identify types of gold standards for evaluation. we exemplify these metrics in the evaluation of several incremental processors that we have developed. we also present generic means to optimise some of the measures, if certain trade-offs are accepted. we believe that this work will help enable principled comparison of components for incremental dialogue systems and portability of results. keywords: evaluation, incrementality, spoken dialogue systems, incremental processing, timing 1. introduction recent work (aist et al. 2007, skantze and schlangen 2009, skantze and hjalmarsson 2010) has shown that incremental dialogue systems, that is, dialogue systems which process user input while it is still ongoing, offer the potential to increase user satisfaction through more natural behaviour. research on how to build such systems, however, is in its infancy, and there are so far no agreed upon standards for how to develop, nor how to evaluate such systems or—as dialogue systems are typically built in a modular fashion—their components. we address the latter question, of how to specify and evaluate such components, which we call incremental processors, in this article. typical processing components in dialogue systems are speech recognition, parsers (often with grammars that are more semantically than syntactically motivated), dialogue act recognition, dialogue management, language generation and text-to-speech synthesis (allen et al. 2000, larsson and traum 2000). of these, all except dialogue managers are used in, and are often most actively developed for, other computational linguistics tasks. even in non-incremental dialogue systems, adaptation and combination of processing components developed for other tasks is not trivial. turning such c©2011 timo baumann, okko buß and david schlangen submitted 1/10; accepted 3/11; published online 5/11 baumann, buss and schlangen processors into incremental processors, i. e. enabling them to generate partial results given partial input that may come from other incremental components in the system, poses additional challenges: how can the desired behaviour with respect to the abilities of accepting partial, possibly uncertain and subsequently revised input, and of producing incrementally growing and improving output, be described, and actual behaviour be evaluated? incrementality adds a temporal dimension to the evaluation problem: there is no longer just one final output of a processor that needs to be evaluated; potentially, there is a sequence of such outputs, whose appropriateness given the itself partial input must be judged, as well as its “stability” over time and the timeliness of its production. in this article, we approach these challenges ‘from the outside’, as it were, by abstractly describing requirements for incremental processors, and deriving from this description metrics for evaluating their performance. we bring together in a unified framework our own previous work on evaluating incremental processors we have developed for our own purposes, and show how this framework can be used to compare this work to attempts by other groups. we also demonstrate how our conceptualisation of the desired outputs allows for the development of generic optimisation techniques that can improve incremental processors along certain parameters. it is our hope that this specification of an evaluation framework will be a valuable contribution to the developing field of incremental dialogue systems, by offering metrics that make work comparable and hence may help achieve the goal of building a toolbox of plug-and-play components. the remainder of the article is structured as follows: we discuss previous work addressing the evaluation of incremental processing components and similar work in the following section. in section 3, we explain our notion of incremental processing and some of the challenges involved in devising incremental processors. sections 4 and 5 are the main part of this article where we discuss our evaluation methodology (in section 4) and how it can be put to use to characterise processing components, namely several processors that we have built (in section 5). in particular, we discuss two different kinds of gold-standards in section 4.1 (those including or not including incrementality information) that can be used in the evaluation of incremental processors. we then discuss how these different kinds of gold standards can be used in evaluations that target different aspects of incremental processing: similarity metrics (section 4.2.1), i. e. measures of equality with or similarity to the gold standard, timing metrics (section 4.2.2), i. e. measures of the timing of relevant phenomena w. r. t. the gold standard, and diachronic metrics (section 4.2.3), which measure the change of incremental hypotheses over time. we give a summary of the metrics (section 4.4) which also contains a tabular overview (table 1). additionally, we exemplify all the metrics presented in the evaluations (section 5). using the considerations about the characteristics of incremental processing from the previous two sections, we present in section 6 some generic ways to optimise incremental processors with respect to some of the metrics. we close with a discussion of possible future work in section 7. 2. previous work accounts of incrementality and incremental processing have been given before (e. g. (amtrup 1999, guhe 2007)). however, most of that work does not systematically deal with the fact that incremental processing requires that a component be able to “change its mind”, i. e. it requires the possibility for an incremental processor to revert previously output hypotheses in the light of additional input. this requirement is described by schlangen and skantze (2009), who introduce a model of incremental processing which supports changes and revocation of previously output hypotheses of an individual 114 evaluation and optimisation of incremental processors processor. this allows a processor to output an hypothesis as soon as possible, while keeping the option of changing or revoking this hypothesis in favour of a different, better one, later on. significant changes and extensions in evaluation methodology become necessary for incremental processing following this approach. many problems in spoken language processing have traditionally been tackled incrementally (e. g. young et al. (1989)), partially to keep memory requirements low. there are even examples of early incremental speech understanding systems (reddy et al. 1976). however, in the system cited above, incrementality was only seen as a means to improve non-incremental system performance, and no evaluation of incremental aspects took place. “consciously incremental” processors for many different tasks have also been described in the literature (e. g. wachsmuth et al. (1998), sagae et al. (2009), and others mentioned below), and these descriptions usually also include evaluations. an often-used method here is to use standard metrics (such as word error rate (wer; hunt (1990)) for speech recognizers) for comparing a non-incremental process and an incremental processor which never change their own previously output hypotheses. as this kind of incremental processing is limited in right context, its results will be worse than non-incremental processing, and the trade-off between incrementality and degraded results is assessed (e. g. wachsmuth et al. (1998)). a notable exception from the tradition of using standard metrics for the evaluation of incremental processors is (kato et al. 2004) in the field of incremental parsing, which deals with the evaluation of incrementally generated partial parse trees. kato et al. define a partial parse’s validity given partial input and implement a method for withholding output hypotheses for a given length of time (similarly to our methods to be presented in section 6), measuring the imposed average delay of this method as well as loss in precision compared to non-incremental processing. while this evaluation method goes some way towards capturing peculiarities of incremental processing, it still cannot account for the full flexibility of the incremental model by schlangen and skantze (2009), which allows for incremental processors to revise previously output hypotheses so that they may eventually produce final results that are as good as those generated by their non-incremental counterparts. such an incremental processor may trade timeliness of incremental results for quality of final results. as we will argue here, for this kind of incremental processing evaluation metrics are needed that capture specifically the incremental aspects of the processors and the evolution over time of incremental results. what needs to be measured is what happens when, as simply comparing final results of incremental and non-incremental settings is not enough. in our own previous work on incremental processors, we have developed metrics—out of a need to do justice to the complexity of the problem of incremental processing, which wasn’t captured by standard metrics—for evaluating incremental asr (baumann et al. 2009a) and for balancing incremental quality and responsiveness; for evaluating incremental reference resolution (schlangen et al. 2009), where the focus was on measuring the “stability” of hypotheses; for evaluating incremental natural language understanding more generally (atterer et al. 2009), looking at the distribution of certain important events (correct hypothesis first found, and final decision reached); and finally for n-best processing (baumann et al. 2009b). it is our goal in the present article to generalise these metrics; we return to this previous work in section 5, where we revisit the results from these papers to illustrate our generalised metrics. there are similarities between the problems posed by incremental processing and those posed by what is often called “anytime processing” (dean and boddy 1988). particularly relevant from that field is zilberstein (1996), who discusses the evaluation of anytime processing. anytime processing is concerned with algorithms that have some result ready at every instant, even when all processing 115 baumann, buss and schlangen figure 1: schematic view of relation between input increments and growing output in incremental processing possibilities have not yet been exhausted. as noted by schlangen and skantze (2009), incremental processing can be seen as a generalisation of anytime processing, where responses are required also for partial input. in anytime processing, there exists a trade-off between the processor’s deliberation time (when to stop a heuristic search) and the quality of results. in incremental processing, the deliberation time can be seen as the amount of input that the processor awaits before returning a partial result. likewise to us arguing that incremental evaluation can not be condensed to one single performance metric, zilberstein (1996) notes for anytime processing that “[t]he binary notion of correctness is replaced with a multivalued quality measure associated with each answer.” he also introduces quality maps and performance profiles as ways to characterise the “expected output quality with execution time t” for anytime processors, a notion which we will adjust to relate output quality to amount of context. 3. our notion of incremental processing incremental processing is concerned with the processing of minimal amounts of input (we call these input increments) in a piece-meal fashion as soon as they become available (guhe 2007). additionally to consuming input incrementally, we also see production of output while the input is still ongoing as an integral part of incremental processing (kilger and finkler 1995)—without this condition, any incrementality a processor might possess internally would be unobservable to the outside world. in this article, we call a processor that consumes input incrementally and generates output incrementally an incremental processor. we are solely concerned with the evaluation of individual incremental processors and do not consider the combination of processors to systems. as a consequence, we do not have to consider the inter-play of loops when processors feed back information to predecessors. in the sense of wirén (1992, as cited by guhe (2007)), our processing is ordered left-to-right. we will detail the input and output of an incremental processor and data structures suitable for this purpose in the following subsections. 116 evaluation and optimisation of incremental processors 3.1 incremental processors a schematic example of input and output of an incremental processor is given in figure 1. the input, consisting of a sequence of eight input increments, is shown in the top row and the output after the processing of each input increment is shown in the subsequent rows. the time spanned by the input increments is shown on the horizontal axis, while the times at which the outputs have been produced are indicated on the vertical axis. we here assume that the output at some instant t is based on all input up to t and represent this correspondence also through the width (and alignment) of the bars representing output in the figure. as an example, output out5 is the output generated at time t5, after all input increments up to and including in5 have been consumed. it is incidental in the figure that input increment times are equidistant: in many cases, there will be irregular intervals between increments. in the discussions in this article, we ignore processing delays imposed by the actual consumption of input and generation of output by the processor. (in reality, tout5 , the time at which output out5 has been created would be slightly after tin5 , the time at which input in5 was consumed.) in the terminology of guhe (2007), we assume processing times to be bounded, i. e. constant per input increment, and assume that it is negligible compared to the time spanned by each increment. notice that not always is new output being generated after consuming an input increment (there is no new output after consuming in1, in4, and in7). this is called moderate (as opposed to massive) incrementality by hildebrandt et al. (1999). for anytime processing, zilberstein (1996) calls the property of having a result ready at all times interruptibility and notes that any non-interruptible processor can be made interruptible by caching the most recent result. this holds analogously for incremental processing. the figure illustrates that there are two temporal dimensions that are relevant when talking about incremental processing: one is the time when the output was made, the other is which segment of the input that output is based on. when evaluating, accuracy of the output should only be determined relative to the input available, and timeliness of the output should be determined relative to the relevant input. 3.2 representing incremental data in figure 1 we deliberately left out the question of how output should be structured. finkler (1997) introduced the distinction between quantitative and qualitative incrementality. in the former, all output is repeated after every input increment and hence the question of output structure is irrelevant for evaluation methodology. at the same time, quantitative incrementality does not facilitate building modular incremental systems, as output from one component cannot easily be fed to a subsequent component in a piece-meal fashion, as the link between individual bits of information is lost. contrastingly, in qualitative incrementality output increments build on the output of the preceding processing step and the relation between successive outputs can be traced: does a later output simply extend an earlier one, or does it (partially) contradict it? or is there no difference and hence no update by successive processors is necessary? for this we need to structure the output of an incremental processor into output increments. furthermore, we may be interested in tracing which parts of the input have lead to a certain output increment in order to make more informed decisions in a later module. 117 baumann, buss and schlangen figure 2: asr hypotheses during incremental recognition. raw asr hypotheses are shown on the left, the corresponding iu network on the right, and step-wise edits in the center. we use the notion of incremental units (ius), introduced in (schlangen and skantze 2009)1 to arrive at a parsimonious representation for both input and output. ius are the smallest ‘chunks’ of information that can trigger a processor into action. ius typically are part of larger units, for example as individual words are parts of an utterance. this relation of being part of the same larger unit is recorded through same level links; the information that was used in creating a given iu is linked to it via grounded in links.2 as ius are connected with these links, an iu network emerges that represents the current output of the processor at the given instant. during incremental processing, the processor incrementally builds up the network of ius that will eventually form its final output, once all input has been consumed. figure 2 shows (stylised) output of an incremental speech recognition (asr) component. in the figure, asr is recognising the utterance “nimm bitte das kreuz [take please the cross]” with intermittent misrecognitions substituting “bis” (“until”) for “bitte” (“please”) and “es” (“it”) for “das” (“the”). the representation on the left of the figure is similar to that of figure 1 above, but here the actual content and structure of the output hypotheses is shown. the rightmost column shows the iu network as it emerges after each processing step. each iu contains a word hypothesis and ius are connected via same level links, eventually forming an utterance. the middle column lists the changes that occur between consecutive states of the output iu network. the current hypothesis (or set of hypotheses), i. e. the current ‘total output’ of an incremental processor during incremental processing can be read off the output iu network by following the same level links from the newest iu(s) backwards. figure 3 shows iu networks for asrs in different configurations; for one-best output, there is one final node, for n-best lists there are n final nodes 1. for an extended version see also (schlangen and skantze 2011) in this volume. 2. the sequential relationship expressed by same level links is represented implicitly in the input sequence in figure 1 simply by placing the increments in sequence; figures 2 (right side), 3 and later figures make them explicit with diamond-headed arrows between ius. informational dependency between input and output is represented in figures 1 and 2 (left side) through horizontal alignment of inputs and outputs while the input has been left out in figure 3 and no dependencies are shown. in later figures, grounded-in links for informational dependency will be expressed by arrows with regular heads. 118 evaluation and optimisation of incremental processors figure 3: one-best (a), n-best (b), or lattice (c) iu output of an asr in the iu framework. with unconnected networks and for lattice-/tree-like structures paths are connected in a common root, ending possibly in several final nodes. the network representation enables a parsimonious representation of the changes to the current hypothesis from one time step to the next (i. e., the real output increments): extension of a previous output is represented as addition of linked ius to a network, while (partial) contradiction is represented as the un-linking of ius from a network and the linking of other ius in their place. finally, when looking at the evolution of the iu network during incremental processing, we are able to notice when certain ius are added or removed. using this representation, both the current full or ‘quantitative’ outputs as well as the changes between consecutive outputs can be traced. this representation along with how it changes over time allows us to define metrics that describe (a) the quality of the results encoded in the network, (b) when the individual contributions that form this result were created by the processor, and (c) what else has been hypothesized by the processor along the way. we will describe how to evaluate these three aspects of incremental processing in the following section, after briefly discussing the types of gold standard information that are required. 4. evaluating incremental processors in the sense that will concern us in this article, evaluation is the comparison of a processor’s actual output to some ideal output in order to assess the processor’s performance.3 such ideal output (often called gold standard) can be manually constructed or automatically generated using additional knowledge that is not available to the processor that is being evaluated. the ideal output of a process will look differently, depending on the task, and the specific goals set for a processor. take as an example the disfluent utterance “pick up the fork and the knife, i mean, the spoon.” even though “knife” is shown to be uttered by mistake with the interregnum “i mean” and later corrected by the reparans “spoon”, it should form part of the output of an incremental speech recognizer as these words have undeniably been spoken. similarly, a simple semantic analyzer should probably generate output for both “knife” and “i mean”, though the latter should be marked as being an interregnum. however, a pragmatic interpreter should not output take(knife), not even intermittently, as this meaning is not what the speaker intends. such an interpreter’s prudence will come at the cost of timeliness, because output can only be generated when it is reasonably certain. these considerations should be taken into account when designing the gold standard for a specific application. 3. another paradigm for the evaluation of dialogue systems is to subjectively evaluate the end-to-end performance of whole systems interacting with users. we will not deal with that methodology in this article, but see section 7 for some remarks. 119 baumann, buss and schlangen comparing ideal to actual output is easy if one is only interested in exact correspondence of the two: either they are equal, or they are not. evaluation is much more difficult if some kind of mismatch between actual and ideal output is allowed and different kinds of mismatches should be rated differently, that is, if what is being measured is the similarity between actual output and gold standard. what counts as similarity between two different possible outputs is highly task-dependent and many different similarity metrics have been developed for different tasks. (to pick two examples, see (nist website 2003) for metrics for evaluating utterance segmentation, and (papineni et al. 2002) for a metric for evaluating the output of machine translators.) incremental evaluation concerns the evaluation of incremental processors, that is, the comparison of actual outputs of such processors to ideal outputs. as explained in the previous section, an incremental processor produces a sequence of partial outputs (one per input increment). hence, there is more to do than just comparing one actual output to one ideal output: we want to compare the content of incremental outputs, the timing of contributions to the result, and the evolution of the output over the course of processing the input sequence. before defining metrics for these three aspects, we will first have to turn to the targets for comparison: the gold standard. 4.1 two types of gold standards for evaluation we have said above that evaluation (of the kind we are aiming at in this article) is the comparison of actual output of a given processor to output that we would ideally like to see. in the previous section, we have shown how we can represent output of an incremental processor in a way that makes all dimensions of incremental information (content, timing, and evolution of incremental results) easily accessible. the question now is how we can get ideal output data that is also in this format. sometimes, this is easy to achieve; in other cases, some information used in this representation cannot be recovered from typical evaluation resources, and an approximation has to be found. we deal with the former case first, and then with the approximation case. 4.1.1 evaluation with incremental gold standards figure 2 from the previous section showed the state of an incremental processor’s output iu network after each input increment consumed. this is the format in which we would like a gold standard for evaluation to be. luckily, the information required for this is often available in existing asr evaluation resources: for the content, we need the sequence of output ius, which in this case is a sequence of words which is provided by the transliteration of the input (either done manually, or automatically with an asr). we also need the link between input increments and output increments which in this case means that we need an alignment between words and audio signal. again, this is often provided by asr evaluation resources, and if not, can be produced automatically via forced alignment. in figure 2 (left side), an aligned gold standard sequence is shown in the top row (labelled “gold”). we can see this as the final state which our incremental processor should ideally reach; but what about the intermediate stages? these can be created from the final state by going backwards through the input ius and removing the current rightmost output iu whenever we go past the input iu that marks its beginning (e. g. at time 10 we would remove “kreuz”, at time 8 “das”, and so on.) following this method, the resulting gold standard demands that an output increment be created as soon as the first corresponding input increment has been consumed; e.g., a word-iu should be produced by an asr as soon as the first audio frame that is part of the word in the gold standard is 120 evaluation and optimisation of incremental processors figure 4: four subsequent outputs for an incremental semantics component as input words are being processed. received. while this will often be impossible to achieve, it provides us with a well-defined upper boundary of the performance that can be expected from an incremental processor. we call the resulting intermediate stages the current gold standard relative to a given input increment. this method is directly transferable to other kinds of input and output. figure 4 shows the incremental growth of a network representing a frame-semantics4; this time the input increments (in this case words, not bits of audio) are shown in the bottom row, and the grounded-in links which relate output to input are represented by arrows. (as the networks in this example are more complex, steps are drawn next to each other and not in rows as in the previous figures.) if we have available a corpus of utterances annotated with their final semantics together with information about which words are responsible for which bits of that final semantics, we can use the same method to go backwards through the input ius and create the full corresponding set of iu states for partial inputs. however, such resources are rare, as making the link between what should be known based on partial input may not even be easy for human annotators (but see (gallo et al. 2007) for an effort to create such a resource). typically, only the final correct semantics is available, with no indication of how to create it (see e. g. the atis corpus as used in (he and young 2005)). in such a case, the intermediate outputs must be approximated from the final state; we will explain how in the next section. 4.1.2 evaluation with non-incremental gold standards figure 5 shows a situation in which a fine-grained link between input and output increments cannot be recovered from the available evaluation resource. we then simply assume that all output ius are grounded in all input ius, which is the equivalent of saying that every input increment contributed to every output increment. the figure only shows the final state, we again derive the incremental steps from this by going backwards through the input ius, as above. however, no desired output will disappear from the gold standard because every output is already grounded in the very first input increment (as we don’t know what input increment some output increment logically depends on). viewed in the direction of time this means that the gold standard is demanding that all output increments be known from the beginning; this is clearly an unreasonable assumption, but as it is kept constant, it allows to measure the gradual approach towards this ideal. such a representation then of course gives us less information, and hence an evaluation based on it can only give a coarse-grained insight into the processor’s performance. if we assume that in reality not all output information is available immediately, the best a processor can do against such 4. in the example given, the slot filling for “modus” depends on “bitte” and all following words. this is because the modus could easily turn out to be e. g. “sarcastic” if some other word had been added later on. 121 baumann, buss and schlangen figure 5: a semantic frame represented in the iu framework without specific dependencies between slots and words. a gold standard is that it fares better and better as more input comes in, and as more and more of what will be the final representation is recovered. likewise, we loose the ability to make fine-grained statements about the timeliness of each output increment. there is a third common case that can be subsumed under this one. sometimes one may want to build a processor that is only incremental on the input side, producing for each input increment an output of the same type as it would for a complete, non-incremental input. an example for this would be a processor that predicts a ‘complete’ utterance meaning based on utterance prefixes. (this has recently been explored by sagae et al. (2009), schlangen et al. (2009) and heintze et al. (2010).) in iu-terms, the output iu is grounded in all input ius and hence such a processor can be evaluated against non-incremental gold standards without loss of information. 4.2 metrics for evaluation of incremental processors we now discuss metrics that quantify differences between actual and ideal output (i. e. the gold standard).5 we identify three categories of metrics: overall similarity metrics (measures of equality with or similarity to a gold standard), timing metrics (measures of the timing of relevant phenomena w. r. t. the gold standard) and diachronic metrics (measuring change of the incremental hypotheses over time), which we will look at in turn. these metrics illuminate the different aspects of incremental performance of a processor, but they are not completely independent of each other (e. g. timing can only be measured if something is correct, absolute correctness entails perfect timing and evolution, etc.). interrelations of metrics will be further discussed in section 4.3. 4.2.1 similarity metrics similarity metrics compare what should ideally be known at some point in time to what is known at that point. the only difference that incremental evaluation brings with it is that the comparison is not done only once, for the final output given complete input, but also for all stages that lead to this final output. an incremental similarity evaluation hence will result in a sequence of results per full input token (e. g. per utterance), where non-incremental similarity evaluation yields only one. to be able 5. when describing our metrics in the following subsections, we will not give fully formalised definitions. first, we believe that the chosen level of abstraction communicates our ideas in a flexible, yet precise manner that allows reproduction and transfer to different situations. secondly, a full formalisation of our metrics would first require the development of a formalism that would need to be capable of expressing subtle details about incremental processing (including revocation of previously output hypotheses, concurrency of processing components and possible delays in message passing). this would be, and hopefully will be, the topic of another paper. 122 evaluation and optimisation of incremental processors figure 6: an asr’s actual incremental output (left), the incremental gold standard (right) and a basic similarity judgement (center): correct (c), prefix-correct (p), or incorrect (!). to evaluate after every input increment, we need a gold standard that covers the ideal outputs after every input increment, as explained in the previous subsection. figure 6 shows such an incremental gold standard and the iu network (for the same utterance as in figure 2) produced by an incremental asr. the most basic measure of similarity is correctness: we simply count how often the output iu network is identical to the current gold standard and divide this by the total number of increments. in figure 6, the output is correct four times, resulting in a correctness of 40 % (ignoring the empty, trivially correct hypotheses 1 and 2 in the calculation). incremental processors often lag behind in producing output for recent input. if this delay (∆) is known in advance, we can take it into account, defining a delay-discounted correctness which, if measured at time t only expects a correctness relative to the gold standard at time t − ∆. however, the processor’s lag will often vary with the input, which a fixed ∆ cannot account for. in this case, we propose counting the number of times that the processor’s output is a prefix of the ideal output at that instant and call the corresponding metric p-correctness. in figure 6, p-correctness is 80 %. correctness has a shortcoming, however, namely that many processors generate output that is often not correct: even the final output—that is, the output when all input increments have been considered—may contain errors compared to the gold standard. (imagine that the asr in the example from figure 6 had generated “nehme” instead of “nimm”; this would have rendered all output increments incorrect in this strict understanding.) in such a case, a processor’s non-incremental deficiencies (i. e. deficiencies that also show in non-incremental processing) block the incremental evaluation. secondly, comparing correctness relative to an independent gold standard conflates the overall accuracy of the processor with the quality aspects of incremental processing. there are two possible remedies for the problem: relaxing the equality condition to some task-dependent similarity measure (that is, changing what counts as correct), or changing the gold standard in such a way that it is guaranteed that final results are regarded as correct. we discuss the latter approach first as it allows for a clean separation between incremental and non-incremental performance. if we want to ensure that the final output of a processor counts as correct, we can simply use this final output as the gold standard and derive from it the set of current gold standards. to distinguish this from correctness as compared to an independent gold standard, we call the measure r-correctness (for relatively correct, relative to the processor’s final output). this separates the evaluation of incremental and non-incremental quality aspects: to learn about the overall (non-incremental) 123 baumann, buss and schlangen quality of the processor, we use standard metrics to compare its output given a complete input with a “true” (externally generated) gold standard; to focus on the incremental quality, we use r-correctness. as a general rule of thumb, incremental performance will improve with improvements to non-incremental performance. the alternative approach of relaxing the equality condition leads us to using task dependent evaluation techniques that make it possible to measure the similarity (and not just identity) of iu networks. which non-incremental measure should be used as the basis for such incremental performance measure depends on the task at hand: for example, incremental speech recognition can be evaluated with wer or cer (boros et al. 1996), syntactic parsing with measures for tree comparison (carroll et al. 1998), semantics by calculating f-measure between frame pairs, or, as a more complex example, specific metrics for natural language generation (reiter and belz 2009). we can ‘incrementalize’ such metrics simply by making comparisons at each input time step, comparing actual output and current gold standard, as explained above. typically, we will be interested in the average of this metric (how similar in general is the output, no matter how partial the input?), but we may also want to explore results at different grades of completeness (does the processor perform worse on smaller prefixes than on longer ones?), or the performance development over time. also, we may want to allow for certain output differences over time, something that is unique to incremental processing. for example, an asr may produce in sequence the hypotheses “grün”, “grüne”, and finally “grünes” as it successively consumes more input. in certain settings, already the first hypothesis may be close enough so that the consuming processor can start to work, and the subsequent hypotheses would not count as contradictions. in such a case, we can set up the similarity metric so that it would allow all variants in the comparison with the gold standard, creating a kind of incremental concept error metric which weighs “sensible” mistakes as less serious than non-sensible mistakes. we close with a discussion of two more similarity metrics. f-score, the harmonic mean of precision and recall, is a useful metric for similarity when the output representation is a ‘bag’ of concepts and no timing information is available in the gold standard. in such a setting, we can expect the score to be very low at the beginning (hardly any slot, compared to the correct representation for the complete input, will have been filled in the beginning) and to rise over time, as more slots are being filled (hopefully correctly). in this case, the shape of f-score curves plotted over different grades of input completeness can become an informative derived measure. finally, we can value mistakes differently at different stages through the input. this is especially appropriate for processors that predict complete outputs (see discussion above at the end of section 4.1.2). defining a time-adjusted error, we can value certain events (e. g. a processor deciding on the special class “undecided”) differently, depending on how much of the input has been seen: indecision is understandable at the beginning of input, but the more one has seen, the more not making a decision becomes just like making a wrong decision.6 4.2.2 timing metrics timing metrics measure when some notable event happens in incremental output relative to some reference time from the gold standard. all other things being equal, we prefer incremental processors 6. there is a moral for life here. 124 evaluation and optimisation of incremental processors that give good and reliable results as early as possible; our timing metrics allow us to give a precise meaning to this. as mentioned above, a characteristic of incremental processing is that hypotheses may have to be revised in the light of subsequent input. this revision means that there are two events specifically that are informative about the performance of the processor: when information becomes available, and when it is not revised any more. we capture this with the following two metrics: • timing of the first occurrence of an output increment (fo), • timing of the final decision for an output increment (fd). the statistics of these events over a test corpus convey useful information: the average fo tells us when (on average) consuming components can usefully start to work on their input, the average fd tells us when a decision becomes stable and can be relied upon. in principle, we might be specifically interested in these metrics for certain tokens in our corpus, or would like to take a close look at the distributions of fo and fd in the corpus. however, we restrict the analyses in section 5 to reporting means, medians and standard deviations to describe the distributions. to determine timing measures, we use the time difference between the occurrence of the iu in the output relative to the gold standard iu’s timing. there are two issues with this that have to be resolved: one is that the gold standard by definition only contains timing information for correct output, and hence, we can only compute these timing measures for increments that we find in the gold standard; the other is that for those correct outputs, we need to decide on how to anchor the timing comparison. for the problem of (possibly) missing references in the gold standard for incorrect increments we can use an automatic alignment between output and gold standard to find references for non-matching increments. alternatively, the final (possibly incorrect) output of the processor may be used to derive the incremental gold standard (as we did for similarity metrics). again, this separates the question of how well the processor performs compared to an external standard from the question of how well it performs incrementally; here, the measures will tell us how fast and how stable the processor is in making the decisions it will ultimately make anyway. regarding the anchoring of the metrics, we measure fo from the beginning of the gold iu; this again encourages speculation to happen as soon as some of the cues relevant for the iu become available. the ‘faster’ a processor is, the lower its average fo will be—some processors may even achieve a negative average fo, if they often predict output before any (direct) evidence for it was available. conversely, for fd (the time until an iu is not withdrawn anymore) it makes sense to anchor in the end of the gold iu, which is the moment when all information regarding this iu has been seen. the more reliable a processor is, the lower its average fd will be. (fd can only be measured after processing has finished; during processing it is unknown whether an iu will be withdrawn later on.) there will be detailed examples of the application of these metrics in the following section, but as a first illustration, for “bitte” in figure 6 above, the first occurrence is at step 7, while it should have appeared at time 6, hence fo(bitte) is 1. as “bitte” is temporarily replaced by “bis”, the final decision is at step 9. looking back at the time-alignment in figure 2 we see that the alignment of “bitte” actually changes between step 8 and 9 (and is only correct at step 9). depending on whether we include this difference in our evaluation fd(bitte) is 1 or 2. whether timing metrics should be measured on an absolute or relative scale will depend on the task at hand and the granularity of the available gold standard. for example, we evaluate timing of asr output in milliseconds, but timing of reference resolution as utterance-percentages in section 5. 125 baumann, buss and schlangen 4.2.3 diachronic metrics in some sense, the types of metrics presented so far cast a static look on the processing results, with the similarity metrics telling us whether a result (at a given time step) is correct or not, and the timing measures telling us when correct results become available, given the whole set of outputs. the metrics we discuss now round out the set of metrics by telling us what happens over the course of processing the input—the diachronic evolution of the final results. we measure this diachronically in the processor’s incremental output by counting the edits to the iu network that are made over time, and in particular, counting the proportion of unnecessary edits. as discussed above, from one processing step to the next, there may be additions, revocations, and substitutions of ius in the network. (this was illustrated in the middle column in figure 2 above.) to build a final result comprised of n increments, it takes at least n addition operations. if an incremental processor ever ‘changes its mind’ about previous output, it will need more edit operations to revoke or substitute a previously output increment. we call the proportion of unnecessary or even harmful edits among the overall edits the processor’s edit overhead (eo).7 (in our running example from figure 2 above, there are 10 edit operations for a total of 4 words, resulting in an eo of 60 %.) why does edit overhead matter, and why do we need another metric to cover this? remember that in incremental systems, the idea is that subsequent components start to work immediately with incremental outputs of their predecessors. hence, unnecessary edits mean unnecessary work for all consumers (and, possibly, for their consumers). a processor that frequently changes previously output hypotheses (we call such changes jitter) may still go unpunished in similarity metrics (as it may produce the same amount of correct intermediate results, but in a “harmful” order) and while there are influences on timing metrics (see next subsection), edit overhead allows for a direct quantification. another kind of diachronic metric can be derived from timing metrics: we call the (average) difference between fo and fd, i. e. the time it takes for an increment to be first hypothesized and to have settled, correction time. we can base an external confidence measure on this statistic, combined with the ‘age’ of an increment (the time that it has survived without being recalled): for example, if a processor is known to have a correction time of less than 500 ms for, say, 90 % of its output ius, then we can be certain to a degree of 90 % that some iu will not change any more once it has been around for 500 ms. while this does not tell us whether an iu is final or not, it helps us to make probabilistic judgements. 4.3 interrelations between metrics in general, one wants a processor to perform as good as possible in all the metrics presented above. in practice, this may often be impossible. to begin with, there are some simple interrelations: perfect correctness is equivalent to perfect fo and fd. both entail zero edit overhead, however the reverse does not hold as the ‘good’ edits may have happened at incorrect processing steps. a processor that is not always correct will always be faced with trade-offs: improving timeliness (fo) by making it hazard guesses earlier will mean that it is more likely to get something wrong (hurting both correctness and eo). reducing eo, while being one of our main objectives in section 6, usually also delays decisions, hurting fo. contrastingly, fd may improve when effort is put into 7. depending on the operational costs for the consumers of the incremental output, we may want to assign different costs to addition, revocation and substitution. for example, in the example evaluations discussed in the next section, we assign a cost of 2 to the substitution operation. 126 evaluation and optimisation of incremental processors table 1: summary of metrics for incremental evaluations name characterisation see also similarity metrics sec. 4.2.1 correctness proportion of ideal intermediate results sec. 5.2 r-correctness incrementally correct relative to the non-incremental result sec. 5.1 discounted (r-)correctness results at time t are compared to gold standard at time t− ∆ sec. 6.1 p-correctness results that are a prefix of their gold standard sec. 5.1 time-adjusted error weighs indecision depending on how much has been seen sec. 5.4 f-score curves average development of f-score over time sec. 5.3 wer/cer curves development of task-specific metrics over time sec. 5.4 timing metrics sec. 4.2.2 fo first occurrence of an output increment fd final decision for an output increment – measured in milliseconds sec. 5.1 – measured as utterance percentage sec. 5.2 diachronic metrics sec. 4.2.3 eo edit overhead, proportion of unnecessary edits sec. 5.1, 5.2 correction time fd – fo, time it takes for an increment to become final sec. 5.1, 6.2 reducing spurious edits. high correctness is not a suitable target per se—a clairvoyant processor outputting the correct final result right from the beginning would even be incorrect against an incremental gold standard up until the very end. however, similarity is a good indicator of whether processing is running as expected and the development of a measure over time can give insights in both processing and properties of the corpus. as a general rule of thumb, improving a processor’s non-incremental performance (that is, the result for the complete input sequence) will also improve the incremental properties, as non-incremental performance is an upper bound for all incremental processing. finally, good performance in some metric may be more important than in another for a given application. this has to be taken into account when comparing different processors (or post-processing techniques as described in section 6) for a given task. 4.4 summary of metrics table 1 gives a short characterisation for the metrics described in this section. we first described correctness, as base-measure for similarity and explained how it can be extended to full-blown, task-dependent similarity metrics based on traditional, non-incremental similarity metrics, with f-score being the most generic, and that it is helpful to plot curves of these metrics over time. also, we noted that using the processor’s final output as gold standard enables the separation of the evaluation of incremental performance from non-incremental performance (r-correctness). we also proposed a time-adjusted error if a fully incremental gold standard is not available. we then described how the timing of a processor’s output can be evaluated by measuring when increments first occur (fo) and when the final decision for them happens (fd). we also discussed anchoring points and temporal scales for these timings. finally, we discussed the diachronic dimension of incremental processing and used edit overhead (possibly with differently weighted edit operations to account for consumer’s needs) and correction time as our diachronic metrics. 127 baumann, buss and schlangen figure 7: the twelve pentomino pieces individually (left) and arranged to form an elephant. we will instantiate these metrics and give exemplary evaluations in the following section before showing how to make trade-offs by exploiting the interrelations between metrics in section 6. as a quick guide, we have also noted the most relevant subsections for every metric in table 1. 5. example evaluations to give concrete examples of how the metrics developed in the previous section can be used to make statements about the performance of incremental processing modules, we now present evaluation experiments we have performed on incremental processors that we have developed.8 the presentation is in order of increasing complexity of the iu networks that are output by the processors. we start with examples of our incremental speech recognizer (which produces sequences of ius) and proceed to our semantic components, first a reference resolver (which outputs just one current iu) and then a semantic chunker (which creates hierarchically ordered sets of ius). the last part of this section is dedicated to a look at n-best processing, which includes results from both asr as well as the semantic chunker. 5.1 evaluation of incremental asr in this section we report our evaluation effort of incremental speech recognition. we first show how to capture performance with our metrics, and then investigate how the metrics are affected by variations to the quality of the underlying speech recognition. 5.1.1 capturing the performance of incremental speech recognisers we use the large-vocabulary continuous-speech recognition framework sphinx-4 (walker et al. 2004) for our experiments, using the built-in lextree decoder, extended by us to provide incremental results. we trained acoustic models for german and a trigram language model that is based on 8. these data have been published before (for section 5.1 in (baumann et al. 2009a), for section 5.2 in (schlangen et al. 2009), for section 5.3 in (atterer et al. 2009), and finally for section 5.4 in (baumann et al. 2009b)) but we reformulate the results in the light of the unified presentation of the metrics outlined above in section 4. we describe the processor’s details and experimental setups only as is necessary for this article’s topic; for details see the aforementioned papers. additionally, please note that all our experiments were conducted with german data. at first blush, this may be seen as problematic for a claim of general applicability. however, we do not see any principled reason why the methodology should not transfer to other languages. concrete results (e. g. timing tendencies) may be different across different languages, for example due to differences in canonical word order, but the evaluative function of the metrics when applied within one language should be preserved, as higher measures in a metric still signal better performance. 128 evaluation and optimisation of incremental processors table 2: base measurements for our incremental asr component ser (non-incremental) 68.2 % wer (non-incremental) 18.8 % r-correctness 30.9 % p-correctness 53.1 % edit overhead 90.5 % fo: mean, stddev, median 0.276 s, 0.186 s, 0.230 s fd: mean, stddev, median 0.004 s, 0.268 s, –0.06 s mean word duration 0.378 s immediately correct 58.6 % spontaneous instructions in a puzzle building domain. the test data is also from that domain. in our domain, which is also the basis for the other evaluations reported below, twelve geometric puzzle pieces (pentominos) are to be described, moved, rotated or otherwise manipulated by the experiment participants. in some of the experiments, puzzle pieces are coloured; in dialogue and wizard-of-oz experiments, an instructor tells a follower (or the wizard) how to place the pieces, most often to form an elephant as in figure 7. for this evaluation, we tried to separate incremental from non-incremental performance. hence, we give the non-incremental word error rate (and sentence error rate) and calculate incremental r-correctness relative to the asr’s final results. additionally, we observed that the asr often lagged behind at word boundaries: the previous word was usually extended over a few hundred milliseconds, before the next word started. to account for such errors, we also measure p-correctness. table 2 shows the results for our evaluation. the asr’s non-incremental wer is 18.8 % (which is a fair value for spontaneous speech (finke et al. 1997)), and of all incremental hypotheses, 30.9 % are correct w. r. t. the final hypothesis, and 53.1 % are at least a prefix of what would be correct. still, we see a substantial edit overhead of 90.5 %, meaning that there were ten times as many edits as there would be in a perfect system. timing metrics were calculated in absolute wall-clock time from the left (fo) respectively right (fd) word boundaries. their means, standard deviation and median are also given in table 2. on average, the correct hypothesis about a word becomes available 276 ms after the word has started (fo). with a mean word duration of 378 ms this means that information becomes available after roughly 3/4 of the word have been spoken. on average, a word becomes final (i. e. is not changed any more) when it has just about ended (mean(fd) = 0.004). we also analysed the correction time for r-correct words (fd−fo). of all words, 58.6 % were immediately correct. the distribution of correction times for correct words is given in figure 8, giving the percentage of words with correction times equal to or lower than the time on the x-axis. while this starts at the initial 58.6 % of words that were immediately correct, it rises above 90 % for a correction time of 320 ms and above 95 % for 550 ms. inversely this means that we can be certain to 90 % (or 95 %) that a current correct hypothesis about a word will not change anymore once it has not been revoked for 320 ms (or 550 ms respectively). (notice that we should calculate correction time for all and not just for r-correct word ius. this is, however, a computationally complex task, as in that case timings of all intermediate hypotheses have to be computed. our experience is that wrong hypotheses die off quickly, so that a curve of correction times for all hypothesized words – not just correct ones – should be even steeper.) 129 baumann, buss and schlangen 40 50 60 70 80 90 100 0 0.2 0.4 0.6 0.8 1 1.2 p e rc e n ta g e o f w o rd s t h a t a re f in a l correction time in s figure 8: distribution of correction times (fd−fo). we concluded from this experiment that our metrics give us useful information about the performance of the component. for example, we now know when (on average) we can expect incremental output in relation to when it has been spoken, and how certain we can be about the ‘finality’ of this output. 5.1.2 stability of the incremental measures to test the dependency of our measures on details of the specific setting, such as audio quality and language model reliability, we varied these factors systematically, by adding white noise to the audio and changing the language model weight relative to the acoustic model (lm weight was 8 in the experiment mentioned above). we varied the noise to produce signal to noise ratios ranging from hardly audible (−20 db), through annoying noise (−10 db) to barely understandable audio (0 db). 0 20 40 60 80 100 2 5 8 11 lm weight r-correctness p-correctness edit overhead wer figure 9: correctness, eo and non-incremental word error rate (wer) with varied language model weights. 0 20 40 60 80 100 orig -20 -15 -10 -5 0 signal to noise ratio in db r-correctness p-correctness edit overhead wer figure 10: correctness, eo and non-incremental word error rate (wer) with additive noise. 130 evaluation and optimisation of incremental processors figure 9 gives an overview of the performance with different lm weights and figure 10 with degraded audio signals. overall, we see that r-correctness and eo remain remarkably stable with different lm and am performance and correspondingly degraded non-incremental wer. a tendency can be seen that higher lm weights result in higher correctness and lower eo. a higher lm weight leads to less influence of acoustic events which dynamically change hypotheses, while the static knowledge from the lm becomes more important, thus reducing jitter. we conclude that non-incremental and incremental measures are highly decoupled and both give fairly independent views of the processor’s overall performance across the variations of the setting. 5.2 evaluation of incremental reference resolution in this section we evaluate a component for incremental reference resolution. we define incremental reference resolution (irr) as the task of inferring the referent referred to in an utterance, while this utterance is going on. the component we evaluate uses a data-driven methodology, and was trained (and cross-validated) on a corpus of 684 utterances, each referencing one of the puzzle pieces in descriptive language. the irr processor consumes word ius. for evaluation, these were derived from text transcriptions of the test corpora used. employing a referent-specific language model trained with the sri lm package (stolcke 2002), the processor arrives at a likelihood belief distribution over the twelve pieces in the domain after each word iu. this distribution was updated with each incoming iu and the top-most referent was evaluated against the gold standard. in terms of the metrics defined here, the study determined edit overhead, fo and fd. additionally, we looked at how correctness improved with utterance completion. notable experimental conditions included using n-best lists and various language models trained with and without pseudowords for hesitations and silences to test robustness against disfluencies. in addition, an adaptive threshold strategy was employed such that a new belief decision is only made if the maximal value after the current word iu is above a certain threshold, where this threshold is reset every time this condition is met. the complete set of experimental results is displayed in table 3. best results for fo and fd were found at 30.43 % (n-best with disfluencies) and 70.89 % (adaptive threshold, no disfluency) of utterance completion respectively. correctness was topped at 37.81 % using n-best processing with disfluencies. eo was low at 69.61 % using adaptive threshold strategy. these figures allow some initial performance analysis along the incremental dimensions. timing metrics fo and fd show promising results, as the processor was able to correctly hypothesise user table 3: first occurrence, first decision, edit overhead and correctness for our incremental reference resolution component with (w/ h) and without hesitations (w/o h) used by the processor. model: n-best rnd-nb adapt max random measures w/ h w/o h w/ h w/o h w/ h w/o h fo 30.4 % 33.7 % 29.6 % 53.9 % 55.3 % 46.6 % 49.3 % 42.6 % fd 87.7 % 85.0 % 97.1 % 71.2 % 70.9 % 96.1 % 94.3 % 98.4 % eo 93.5 % 90.7 % 96.7 % 69.6 % 67.7 % 92.6 % 89.4 % 93.2 % correctness 37.8 % 36.8 % 23.4 % 23.0 % 26.6 % 17.8 % 20.2% 7.8% 131 baumann, buss and schlangen intention on average as early as after a third of the input utterance and offer a stable hypothesis after roughly two-thirds (albeit under different experimental conditions). also, reference resolution edit overhead stayed far below that of the asr, discussed above. we conclude that the metrics were useful in describing the performance of this component, and allow us to perform meaningful comparisons between the variations of the model. 5.3 evaluation of incremental semantics construction in addition to the statistical irr component described above, we evaluated a rule-based semantic chunker (described in (atterer and schlangen 2009)) on this task. the chunker also consumes word ius and, based on chunks defined in a grammar, fills slots in a semantic frame. we present results from two evaluations of this component, one in which the input words came from corpus transcriptions and another in which they were taken from actual asr output. 5.3.1 evaluation with ideal input the initial text-only evaluation using data from task-oriented corpora in the puzzle domain described above provided timing and similarity measures for all slots in a semantic frame (piece involved, action, possibly parameters of that action). fo and fd were determined for each utterances’ frame interpretation. the frame fo was below 40 % of utterance completion for almost all utterances, with gold standard transcriptions filling at least one slot. fd was generally achieved between 85 % utterance completion and the utterance’s end.9 the results from the study are re-represented in figure 11 as recall, precision and f-score. the gold standard available for evaluation did not include linkage information between individual input and output ius, so precise reasoning about each input iu’s performance with respect to correctness or timing was impossible. for this reason, we reverted to using f-score which provides a natural workaround for this shortcoming. portability of experimental results is one of our aims in re-presenting this data here. as such, this representation provides a point of comparison to (sagae et al. 2009), who present a similar approach employing similar measures. while training of the incremental semantic models for the two approaches varies, the experimental findings are very similar. both models exhibit an initial f-score of approximately 20 % progressing to an asymptotical final of just below 80 %, exhibiting similar gains along the way. this similarity provides some preliminary conclusions about the nature of stability and accuracy of partial complex nlu hypotheses. building on the work by sagae et al. (2009), devault et al. (2009) train a classifier to decide on when the interpretation of an utterance is likely to not improve any further. this is similar in spirit to our output optimisation techniques discussed in section 6. 5.3.2 evaluation with asr output we also evaluated the chunker with actual asr output (and including asr timing information) using a sub-domain in which users were asked to describe one of the puzzle pieces using descriptive language.10 we collected an evaluation corpus in a wizard of oz study; the gold standard for the 9. note that these numbers are based on word counts rather than wall clock timing, which would be less meaningful for a word-based processor. 10. since we are only interested in one of the slots filled by the chunker, the one referring to the piece in question, the task may perhaps more aptly be described as chunk-based reference resolution. 132 evaluation and optimisation of incremental processors 0 0.2 0.4 0.6 0.8 1 0 20 40 60 80 100 utterance completion (relative position in %) recall precision f-score figure 11: recall, precision and f-score of the semantics component at different positions in the utterance (from atterer and schlangen (2009), metrics recalculated). asr was hand-transcribed, while gold for the chunker was derived from meta-information (each scene of the study was associated with a single puzzle piece.) the one-best case in that study is available as a baseline, resulting in fo at 51.2 %. this is in line with some of the findings from evaluating our irr component in section 5.2. this observation is relevant in several respects. the fo data is a measure of response time. by seeing similar results in both components for responsiveness, we validate our hypothesis that something can be known about a user’s intention early on in the utterances. in addition, as we noted when comparing our results with (sagae et al. 2009), we now have a rough indication what this responsiveness might be (or at least an operating point on which to improve, see section 6 below). lastly, since these two components can be charged with solving overlapping tasks (in our case resolving references quickly), and since our intention here is to enable informed comparisons of two or more components, we now have the tools to do so from a timing perspective. 5.4 evaluation with incremental asr n-best results the asr results of the study described in the previous subsection also included n-best lists, which we used to gauge the overall effect, beneficial or adverse, of considering recognition alternatives on the three dimensions of incremental evaluation: similarity, timing, and evolution. for each incremental asr result obtained we measured wer, as well as correctness as a concept error rate, cer, as per the reference resolution task.11 notice that we made cer time adjusted, i. e. an unfilled concept was counted as a small error early on in the utterance, but as a full error towards the end. we then calculated oracle and anti-oracle scores for wer and cer. an oracle error rate for an n-best list is the error rate of the best hypothesis contained in the list; correspondingly, an anti-oracle score is the worst score in the list. this gave us potential gains in wer, cer, and fo, associated with being able to identify the best hypothesis among the n , as well as the risk associated with picking the worst instead. in addition, by observing the output for all n over time, we were able to observe 11. for performance reasons we limited the asr’s beam width and set n to a maximum of 100,000. 133 baumann, buss and schlangen 0 0.05 0.1 0.15 0.2 0.25 0.3 0 20 40 60 80 100 o ra c le c e r utterance completion (% of total time) incremental oracle concept error rate for different n n=1 n=11 n=79 figure 12: incremental oracle concept error rate for selected n in n-best processing. the effect of n-best processing on edit overhead, as presumably the number of unnecessary changes would increase. results from ranking the n-best list for oracle wer were sobering. using the recognition results associated with oracle wer, provided little or no gain in cer. however, there were timing benefits: we observed a 20 % potential gain in fo from a baseline 51.2 % to an oracle 41.0 % (i. e. asr results in the n-best list existed that resolved references on average that much sooner than a one-best result would have). however, this gain came with a significant (150-fold) increase in edit overhead, which is clearly unacceptable and results from the sporadically huge n-best lists during recognition. we repeated the evaluation with capped n-best lists. already low values of maxn promised similar (but smaller) gains in correctness, the same reduction in fo, but a significant reduction in edit overhead (e. g. at a low maxn of 11, it halved to 60-fold over the one-best baseline). graphs for various settings of maxn and corresponding oracle cer are shown in figure 12. note that cer rises because we employed an adjusted error as mentioned above. the final drop can be attributed to a syntactic phenomenon in the data set. without a method for determining the best result in an n-best list the benefits just outlined remain hypothetical. we will now discuss methods that are available for improving incremental processors. 6. optimising the output of incremental processors in the previous section, we showed some example incremental processors and how they fared in evaluations of their respective tasks. in this section, we demonstrate simple (incremental) postprocessing techniques to improve the output of an incremental processor. these techniques are aimed mainly at the reduction of edit overhead, as this was one of the major problems that came up in our evaluations. so far, they are mostly only tested on asr output (where the problem of edit overhead is most severe), but apply equally to other incremental processors.12 the graphs presented in these subsections are similar in spirit to zilberstein’s (1996) 12. this is especially true for other kinds of ‘input’ processors, that receive input from the outside world (e. g. gesture and face recognition). 134 evaluation and optimisation of incremental processors performance profiles allowing to analyse where some delay can hopefully result in the largest improvements of an overall system. finally, we give a theoretical discussion of how post-processing is helpful even if edit overhead is not an issue, as is typically the case in what we call ‘beat-driven’ systems. 6.1 right context incremental processing is hard due to the fact that a processor is expected to output a result as soon as some input becomes available. obviously, some results will be wrong, as the correct final output can not yet be determined from the very beginning (cmp. the discussion in section 4.2.3 above). a simple strategy to improve results is to allow the processor some right context which it can use as input, but for which it does not yet have to provide any output. for example, typical asr systems use this strategy internally at word boundaries (with very short right contexts) in order to restrict the language model hypotheses to an acoustically plausible subset (ortmanns and ney 2000). in the experiment described here, we allow the asr a larger right context of size ∆ by taking into account at time t the output of the asr up to time t− ∆ only. that is, what the asr hypothesises about the right context is considered to be too immature and is discarded, while the hypotheses about the input up to t− ∆ have the benefit of a lookahead up to t. this reduces the jitter, which is found mostly to the very right of the incremental hypotheses. thus, we expect to reduce the edit overhead in proportion with ∆. on the other hand, allowing the use of a right context leads to the current hypothesis lagging behind the gold standard. (correspondingly, fo increases by ∆.) obviously, using only information up to t− ∆ has averse effects on correctness as well, as this measure evaluates the word sequences up to time t which may already contain more words (those in the right context). to account for this lag (which is known in advance), we also report a discounted r-correctness. figure 13 details the results for the data from section 5.1 with right context between 1.5 s and −0.2 s. (the x-axis plots ∆ as negative values, with 0 being “now”. results for a right context (∆) of 1.2 can thus be found 1.2 to the left of 0, at −1.2.) we see that at least in the discounted measure, fixed lag performs quite well at improving both the processor’s correctness and eo. this is due to the fact that asr hypotheses become more and more stable when given more right context. however even for fairly long lags, late edits cannot be ruled out entirely. to illustrate the effect of a system that does not support editing of hypotheses but immediately commits itself to an hypothesis we plot the fixed wer that would be reached by such a system after a right context of ∆. as can be seen in the figure, it is extremely high for low right contexts (exceeding 100 % for ∆ ≤ 400 ms) and remains substantially higher than the non-incremental wer even for fairly large right contexts. actually, the wer plot by wachsmuth et al. (1998) looks very similar. as expected, the analysis of timing measures shows an increase with larger right contexts with their mean values quickly approaching ∆ (or ∆−mean word duration for fd), which are the lower bounds when using right context. correspondingly, the percentage of immediately correct hypotheses increases with right context reaching 90 % for ∆ = 580 ms and 98 % for ∆ = 1060 ms. finally, we can extend the concept of right context into negative values, predicting the future, as it were. by choosing a negative right context, in which we extrapolate the last hypothesis state by ∆ into the future, we can measure the correctness of our hypotheses correctly predicting the near future. the graph shows that 15 % of our hypotheses will still be correct 100 ms in the future and 10 % will 135 baumann, buss and schlangen 0 20 40 60 80 100 -1.6 -1.4 -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 0.2 right context in s (scale shows larger right contexts towards the left) (strict) r-correctness discounted r-corr. p-correctness edit overhead wer figure 13: correctness, eo and fixed-wer for varying right contexts ∆. 0 20 40 60 80 100 -1 -0.8 -0.6 -0.4 -0.2 0 smoothing in s (scale shows larger smoothings towards the left) (strict) r-correctness discounted r-corr. p-correctness edit overhead figure 14: correctness and edit overhead for varying smoothing lengths. still be correct for 170 ms. unfortunately, there is no way to tell apart hypotheses that will survive and those which will soon be revised. 6.2 hypothesis smoothing in the previous method we suppressed wrong edits by avoiding the recognition jitter that the asr produces in the ‘youngest’ parts of its recognition results. in this section, we look at the diachronic evolution of the results and use the age of an hypothesized increment as a cue. (this is similar in spirit to using correction time to determine the certainty that an iu will last.) we simply only pass on an iu if it reaches a certain age n , that is, it is part of n consecutive hypotheses. to illustrate the process with n = 2 we return to figure 2. neither of the words “bis”, or “es” would ever be output, because they are only present for one time-interval each. edits would occur at the following times: ⊕(nimm) at t6, ⊕(bitte) at t10 (only then is “bitte” the result of two consecutive hypotheses) and ⊕(das) at t12 and so on. with a smoothing factor of n = 2, no words are revoked in the example, because all revocations last for only 1 time-step. yet, in general, changes can of course still occur, namely if it is called for by at least n consecutive time-steps. for this strategy, as can be seen in figure 14, edit overhead falls rapidly, reaching 50 % (for each edit necessary, there is one superfluous edit, eo parity) with only 110 ms and 10 % with 320 ms. (the same thresholds are reached through the use of right context at 530 ms and 1150 ms respectively as shown above.) likewise, the prefix correctness improvements are better than when using right context, indicating the effectiveness of the technique. at the same time, r-correctness is poor. this is due to correct hypotheses being held back too long due to the hypothesis sequence being interspersed with wrong hypotheses (which only last for few consecutive hypotheses), resetting the counter until the add message (for the prevalent and potentially correct word) is sent. still, the positive effects of hypothesis smoothing outweigh its disadvantages, as we reach eo parity with n at 110 ms instead of ∆ = 530 ms leading to an increase in fo of 140 ms and in fd of 67 ms. when comparing to increases of at least 530 ms using right context, results are rather impressive. 136 evaluation and optimisation of incremental processors 6.3 optimising for beat-driven systems when we described the advantages of the above methods, we put a special focus on edit overhead, as this is a very important metric in our event-based spoken dialogue system architecture (schlangen and skantze 2009, schlangen et al. 2010). some incremental spoken dialogue systems (devault et al. 2009, raux and eskenazi 2009) follow a different approach to incremental processing, which we here call the beat-driven approach. in such systems, processing is done repeatedly (at a fixed or flexible rate) on the partial input available so far. while the system and its components may or may not know about the fact that they are repeatedly processing partial (and unfolding) input, they do not usually take that into account while processing. instead, they completely reprocess all input at every beat.13 obviously, for such a system edit overhead is not an issue, because all intermediate results are recomputed at every beat regardless of previous input. in this section, we want to show how an incremental processor may optimise its output for a beat-driven consumer. again, we take as an example our incremental speech recognition component. our processor, as explained above, is suited to only output messages once new information is available. in the previous subsections, we have shown ways of increasing the reliability of such messages, and have mostly been concerned with reducing the edit overhead from messages that have to be revoked and changed. here we are concerned with the increased correctness (and prefix-correctness) of the output when the above improvements are carried out. consider a system with a beat β of 200 ms, which polls the asr at every 200 ms for its most recent result. in such a setting correctness remains the same (as—on average—the results at every 200 ms will be just as good or bad as all other incremental results). however, new output increments will be delayed on average by β/2 because they are only registered on the next beat (which is, on average, β/2 into the future). thus, as timing is already worse due to querying results only when the beat falls, adding an additional small delay through the use of smoothing or right context as described above may be well justified by the resulting increase in correctness. 7. conclusions and future work in this article we have presented a general approach to representing the output of what we call incremental processors, that is, components in incremental dialogue processing systems. we have discussed how gold standards can be created in the same format, to enable evaluation based on the comparison of actual and ideal output. sometimes, resources for producing a fully incremental gold standard may not be available; we have discussed ways to approximate one in these cases. we have then presented three types of metrics: (1) timing metrics, (2) similarity metrics and (3) diachronic metrics, measuring incrementality in terms of speed, accuracy and change over time, respectively. presenting our own evaluations of various incremental modules, we have demonstrated how these metrics can usefully characterise performance and how they can guide optimisation techniques that improve performance without changes to the internal workings of the modules. we hope that these metrics and methods will be of use in the wider community and improve comparability of results and the exchangeability of modules. 13. we do not claim that one or the other method is superior to the other, at least not from a system point of view. (we do think, however, that repeated re-processing is psycho-linguistically rather implausible.) while starting from previous results potentially saves processing cycles, it also incurs book-keeping overheads. also, the beat-driven approach allows to easily integrate originally non-incremental components into an incremental system. 137 baumann, buss and schlangen our discussion focused on incremental processors that are tasked with input recognition and interpretation. we did not discuss how these metrics can be applied to the evaluating of processors that are concerned with ‘later’ aspects of dialogue, such as producing system behaviour and generating output. there are some recent efforts in this direction, such as (skantze and hjalmarsson 2010, buß and schlangen 2010) and (buß et al. 2010). how exactly the metrics presented here relate to these efforts is an open question we currently explore. also missing from our discussions is the issue of how the characterisation of the performance of components relates to that of systems, or more concretely, how improvements of the former can lead to improvements of the latter. zilberstein (1996) analyses how performance profiles for anytime algorithms can be used to find an optimal resource allocation when combining several components to form a complex system. in incremental processing, a similar analysis of how incremental metrics (e. g. timing) aggregate in a complex system and how and where to best apply optimisation techniques in the system would make for very helpful future work. the profiles which we report for the optimization techniques in section 6 can be seen as a first step in this direction. finally, feed-back loops as in wirén’s (1992) full incrementality may incurr additional challenges. on a less technical account, the ‘success’ of spoken dialogue systems is usually measured in terms such as user satisfaction or task success. the prediction is that that there is at least an indirect connection between such subjective measures and the performance metrics discussed in this article, since user-facing behaviours that are enabled by incremental processing (e. g., fast turn-taking, backchannels; see discussion in (buß and schlangen 2010)) should benefit from improvements in incremental processing. the exact nature of this connection, however, and the place of the metrics in system-evaluation frameworks like (walker et al. 1997) or (möller and ward 2008) will have to be explored in future work. acknowledgements this work was funded by a dfg grant in the emmy noether programme. we wish to thank the anonymous reviewers for their very helpful comments. references gregory aist, james allen, ellen campana, carlos gallo, scott stoness, mary swift, and michael tanenhaus. incremental dialogue system faster than and preferred to its nonincremental counterpart. in proceedings of the 29th annual conference of the cognitive science society, pages 761–766, nashville, usa, august 2007. james allen, donna byron, myroslava dzikovska, george ferguson, lucian galescu, and amanda stent. an architecture for a generic dialogue shell. natural language engineering, 6(3):213–228, 2000. jan willers amtrup. incremental speech translation. springer verlag, berlin, germany, 1999. michaela atterer and david schlangen. rubisc – a robust unification-based incremental semantic chunker. in proceedings of srsl 2009, the 2nd workshop on semantic representation of spoken language, pages 66–73, athens, greece, march 2009. 138 evaluation and optimisation of incremental processors michaela atterer, timo baumann, and david schlangen. no sooner said than done? testing the incrementality of semantic interpretations of spontaneous speech. in proceedings of interspeech, pages 1855–1858, brighton, uk, september 2009. timo baumann, michaela atterer, and david schlangen. assessing and improving the performance of speech recognition for incremental systems. in proceedings of naacl-hlt, pages 380–388, boulder, usa, may 2009a. timo baumann, okko buß, michaela atterer, and david schlangen. evaluating the potential utility of asr n-best lists for incremental spoken dialogue systems. in proceedings of interspeech, pages 1031–1034, brighton, uk, september 2009b. manuela boros, wieland eckert, florian gallwitz, günther görz, gerhard hanrieder, and heinrich niemann. towards understanding spontaneous speech: word accuracy vs. concept accuracy. in proceedings of the 4th icslp, pages 1009–1012, philadelphia, usa, october 1996. okko buß and david schlangen. modelling sub-utterance phenomena in spoken dialogue systems. in proceeding of pozdial, the 14th international workshop on the semantics and pragmatics of dialogue, pages 33–41, poznań, poland, june 2010. okko buß, timo baumann, and david schlangen. collaborating on utterances with a spoken dialogue system using an isu-based approach to incremental dialogue management. in proceedings of the sigdial conference, pages 233–236, tokyo, japan, september 2010. john carroll, ted briscoe, and antonio sanfilippo. parser evaluation: a survey and a new proposal. in proceedings of lrec, pages 447–454, granada, spain, may 1998. thomas dean and mark boddy. an analysis of time-dependent planning. in proceedings of aaai-88, pages 49–54, cambridge, usa, august 1988. david devault, kenji sagae, and david traum. can i finish? learning when to respond to incremental interpretation results in interactive dialogue. in proceedings of the sigdial conference, pages 11–20, london, uk, september 2009. michael finke, petra geutner, hermann hild, thomas kemp, klaus ries, and martin westphal. the karlsruhe-verbmobil speech recognition engine. in proceedings of icassp, pages 83–86, munich, germany, april 1997. wolfgang finkler. automatische selbstkorrektur bei der inkrementellen generierung gesprochener sprache unter realzeitbedingungen. dissertationen zur künstlichen intelligenz. infix verlag, sankt augustin, germany, 1997. carlos gómez gallo, gregory aist, james allen, william de beaumont, sergio coria, whitney gegg-harrison, joana p. pardal, and mary swift. annotating continuous understanding in a multimodal dialogue corpus. in proceeding of decalog, the 11th international workshop on the semantics and pragmatics of dialogue, pages 75–82, trento, italy, june 2007. markus guhe. incremental conceptualization for language production. lawrence erlbaum associates, mahwah, usa, 2007. 139 baumann, buss and schlangen yulan he and steve young. semantic processing using the hidden vector state model. computer speech and language, 19(1):85–106, 2005. silvan heintze, timo baumann, and david schlangen. comparing local and sequential models for statistical incremental natural language understanding. in proceedings of the sigdial conference, pages 9–16, tokyo, japan, september 2010. bernd hildebrandt, hans-jürgen eikmeyer, gert rickheit, and petra weiß. inkrementelle sprachrezeption [incremental language understanding]. in kogwis: proceedings der 4. fachtagung der gesellschaft für kognitionswissenschaft, pages 19–24, bielefeld, germany, september 1999. melvin j. hunt. figures of merit for assessing connected-word recognisers. speech communication, 9(4):329–336, 1990. yoshihide kato, shigeki matsubara, and yasuyoshi inagaki. stochastically evaluating the validity of partial parse trees in incremental parsing. in proceedings of the acl workshop on incremental parsing: bringing engineering and cognition together, pages 9–15, barcelona, spain, july 2004. anne kilger and wolfgang finkler. incremental generation for real-time applications. technical report rr-95-11, dfki, saarbrücken, germany, 1995. staffan larsson and david r. traum. information state and dialogue management in the trindi dialogue move engine toolkit. natural language engineering, 6(3–4):323–340, 2000. sebastian möller and nigel g. ward. a framework for model-based evaluation of spoken dialog systems. in proceedings of the 9th sigdial workshop on discourse and dialogue, pages 182–189, columbus, usa, june 2008. nist website. rt-03 fall rich transcription, 2003. url http://www.itl.nist.gov/iad/894. 01/tests/rt/2003-fall/index.html. stefan ortmanns and hermann ney. look-ahead techniques for fast beam search. computer speech and language, 14(1):15–32, 2000. kishore papineni, salim roukos, todd ward, and wei-jing zhu. bleu: a method for automatic evaluation of machine translation. in proceedings of acl, pages 311–318, philadelphia, usa, july 2002. antoine raux and maxine eskenazi. a finite-state turn-taking model for spoken dialog systems. in proceedings of naacl-hlt, pages 629–637, boulder, usa, may 2009. d. raj reddy, lee d. erman, richard d. fennell, and richard b. neely. the hearsay-i speech understanding system: an example of the recognition process. ieee transactions on computers, c-25(4):422–431, 1976. ehud reiter and anja belz. an investigation into the validity of some metrics for automatically evaluating natural language generation systems. computational linguistics, 35(4):529–558, 2009. 140 http://www.itl.nist.gov/iad/894.01/tests/rt/2003-fall/index.html http://www.itl.nist.gov/iad/894.01/tests/rt/2003-fall/index.html evaluation and optimisation of incremental processors kenji sagae, gwen christian, david devault, and david traum. towards natural language understanding of partial speech recognition results in dialogue systems. in proceedings of naacl-hlt, pages 53–56, boulder, usa, may 2009. david schlangen and gabriel skantze. a general, abstract model of incremental dialogue processing. in proceedings of eacl, pages 710–718, athens, greece, march 2009. david schlangen and gabriel skantze. a general, abstract model of incremental processing. dialogue and discourse, 2(1):83–111, 2011. david schlangen, timo baumann, and michaela atterer. incremental reference resolution: the task, metrics for evaluation, and a bayesian filtering model that is sensitive to disfluencies. in proceedings of the sigdial conference, pages 30–37, london, uk, september 2009. david schlangen, timo baumann, hendrik buschmeier, okko buß, stefan kopp, gabriel skantze, and ramin yaghoubzadeh. middleware for incremental processing in conversational agents. in proceedings of the sigdial conference, pages 51–54, tokyo, japan, september 2010. gabriel skantze and anna hjalmarsson. towards incremental speech generation in dialogue systems. in proceedings of the sigdial conference, pages 1–8, tokyo, japan, september 2010. gabriel skantze and david schlangen. incremental dialogue processing in a micro-domain. in proceedings of eacl, pages 745–753, athens, greece, march 2009. andreas stolcke. srilm an extensible language modeling toolkit. in proceedings of the 7th icslp, pages 901–904, denver, usa, september 2002. sven wachsmuth, gernot a. fink, and gerhard sagerer. integration of parsing and incremental speech recognition. in proceedings european signal processing conference, pages 371–375, rhodes, greece, september 1998. marilyn a. walker, diane j. litman, candace a. kamm, and alicia abella. paradise: a framework for evaluating spoken dialogue agents. in proceedings of acl and eacl, pages 271–280, madrid, spain, july 1997. willie walker, paul lamere, philip kwok, bhiksha raj, rita singh, evandro gouvea, peter wolf, and joe woelfel. sphinx-4: a flexible open source framework for speech recognition. technical report smli tr2004-0811, sun microsystems inc., 2004. mats wirén. studies in incremental natural language analysis. unpublished doctoral dissertation, linkoping university, sweden, 1992. steve j. young, nh russell, and jhs thornton. token passing: a simple conceptual model for connected speech recognition systems. cambridge university engineering department technical report cued/f-infeng/tr, 38, 1989. shlomo zilberstein. using anytime algorithms in intelligent systems. ai magazine, 17(3):73–83, 1996. 141 introduction previous work our notion of incremental processing incremental processors representing incremental data evaluating incremental processors two types of gold standards for evaluation evaluation with incremental gold standards evaluation with non-incremental gold standards metrics for evaluation of incremental processors similarity metrics timing metrics diachronic metrics interrelations between metrics summary of metrics example evaluations evaluation of incremental asr capturing the performance of incremental speech recognisers stability of the incremental measures evaluation of incremental reference resolution evaluation of incremental semantics construction evaluation with ideal input evaluation with asr output evaluation with incremental asr n-best results optimising the output of incremental processors right context hypothesis smoothing optimising for beat-driven systems conclusions and future work c:/users/eleni/documents/talks/dialogue discourse/ddpaper020511.dvi dialogue and discourse 2(1) (2011) 199–233 doi: 10.5087/dad.2011.109 incrementality and intention-recognition in utterance processing eleni gregoromichelaki and ruth kempson {eleni.gregor, ruth.kempson}@kcl .ac.uk department of philosophy king’s college london, strand, london wc2r 2ls, uk matthew purver mpurver@dcs.qmul .ac.uk department of computer science queen mary university of london, mile end road, london e1 4ns, uk gregory j. mills gjmills@stanford.edu department of psychology stanford university, 450 serra mall, stanford, california 94305, usa ronnie cann r.cann@ed.ac.uk school of philosophy, psychology and language sciences university of edinburgh, dugald stewart building, 3 charles street edinburgh, eh8 9ad, uk wilfried meyer-viol wilfried.meyer viol @kcl .ac.uk department of philosophy king’s college london, strand, london wc2r 2ls, uk patrick g. t. healey ph@dcs.qmul .ac.uk department of computer science queen mary university of london, mile end road, london e1 4ns, uk editor: hannes rieser and david schlangen abstract ever since dialogue modelling first developed relative to broadly gricean assumptions about utterance interpretation (clark, 1996), it has remained an open question whether the full complexity of higher-order intention computation is made use of in everyday conversation. in this paper we examine the phenomenon ofsplit utterances, from the perspective ofdynamic syntax, to further probe the necessity of full intention recognition/formation in communication: we do so by exploring the extent to which the interactive coordination of dialogue exchange can be seen as emergent from low-level mechanisms of language processing, without needing representation by interlocutors of each other’s mental states, or fully developed intentions as regards messages to be conveyed. we thus illustrate how many dialogue phenomena can be seen as direct consequences of the grammar architecture, as long as this is presented within anincremental, goal-directed/predictivemodel. keywords: dialogue, gricean accounts of communication, higher-order intention recognition, incrementality, intentions, mind-reading, plans, predictivity, split-utterances c©2011 e. gregoromichelaki, r. kempson, m. purver, g. j. mills, r. cann, w. meyer-viol, p. g. t. healey submitted 1/2010; accepted 1/2011; published online 5/2011 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey 1. introduction: rethinking intentionalism in communication ever since the first attempts at modelling communication relative to broadly gricean assumptions, it has remained an open question what level of complexity of higher-orderintention computation is made use of in everyday conversation. in this paper, we argue that theinteractive coordination of dialogue exchange can be seen as emergent from the mechanisms of language processing, without either needing representation by interlocutors of each other’s mentalstates or fully developed intentions as regards messages to be conveyed. this conclusion is controversial in that it is not commensurate with a broad swathe of recent pragmatic theorising. higher-order intention recognition forms the underpinning to grice’s account of non-natural meaning (meaningnn ) and the subsequent communication models that have been based on it. central to all such accountsis the assumption that understanding by a hearer involves recognition of the particular proposition a speaker intended to express, via their recognition of that intention. though the conceptual and psychological problems higher-order intention recognition gives rise to are wellknown, responses to these criticisms have been muted. either the problems are ignored altogether; or they have resulted in unsubstantiated weakenings of the stringent requirements such recognition places onthe recovery of meaning in communication even though purportedly retaining the central tenets of the gricean paradigm. in this paper, having introduced general philosophical and psychological re-evaluations of the status of higher-order intention recognition, we turn to an additional consideration, the problems raised for gricean views by the phenomenon ofincrementalityin both comprehension and production as manifested in conversational dialogue. our particular focus is the phenomenon of so-calledsplit utterances, commonly seen in dialogue, in which speakers and hearers reverse roles mid-utterance. to deal effectively with the analysis of such shared productions, we turn to a model in which incrementality is a core property of the grammar formalism. under this assumption, we first show how, with “syntax” re-defined to be the incremental and monotonic growth ofsemantic representation, the split utterance phenomenon is straightforwardly both predictable and explainable. we then argue that, relative to this model, recognition of the content of speakerintentions is not a necessary condition for human interaction. hence, we will conclude, it is not an intrinsic property of communication. 1.1 intention recognition in communication and dialogue grice’s account of communication (published as grice, 1975), based onthe notion of “meaningnn ”, was the point of departure for many subsequent pragmatic models (see levinson, 1983; bach, 1997; bach and harnish, 1982; cohen et al., 1990, a.o.).1 it characterised communication as essentially involving rationality and cooperation, displayed by the requirement that cooperative interlocutors must be guided by reasoning about mental states: speaker’s meaning, whose recovery is elevated as the fundamental criterion for successful communication, involves the speaker at least (a) having the intention of producing a response (e.g. belief) in the addressee (i.e. having a thought about the addressee’s thoughts) and (b) also having a second order intention regarding the addressee’s belief about the speaker’s second order thought (in order to capture the presumed fulfillment of the communicative intention by means of its recognition). under this definition, speakers have (at least) fourth order thoughts and hearers must recover speaker’s meaning through reasoning about these 1. note that our arguments here do not necessarily concern grice’s philosophical account, in so far as it is seen by some as just normative, but its employment in subsequent (psychological/computational) models of communication/pragmatics. 200 incrementality and intention -recognition in utterance processing thoughts. early on, philosophers like strawson (1964) and schiffer (1972) severally presented scenarios where the criterion of higher-order intention recognition was satisfied even though this still was not sufficient for the cases to be characterised as instances of “communication” (as opposed to covert manipulation, sneaky intentions etc.). this led to the postulation of successively higher levels of intention recognition as a prerequisite for communication, and an attendant concept of “mutual knowledge” of speaker’s intentions, both of which were recognised as facing a charge of infinite regress (see e.g. sperber and wilson, 1995, 256-77). although in applications of this account in psychological implementations it is not necessary to assume that explicit reasoning takes place online, nevertheless, an inferentially-driven account of communication on this basis has to provide a model that explicates the concept of ‘understanding’ as effectively analysed through a logical system that implements these assumptions (see e.g. allott, 2005). so, even though such a system can be based on heuristics that short-circuit complex chains of inference (grice, 2001, 17), the logical structure of the derivation of an output has to be transparentif the implementation of that model is to be appropriately faithful (see e.g. grice, 1981, 187 for the calculability of implicatures). agents that are not capable of grasping this logical structure independently cannot be taken to be motivated by such computations, except as an idealisation pending a more explicit account. on the other hand, ignoring in principle the actual mechanisms that implement such a system as a competence/performance issue or an issue involving marr’s (marr, 1982) computational vs the algorithmic and implementational levels (see e.g. stone (2005); stone (2004); geurts (to appear):ch4) does not shield one from charges of psychological implausibility: if the same effects can be accounted for with standard psychological mechanisms without appeal to the complex model then, by occam’s razor, such an account would be preferable, especially if subtle divergent predictions can be uncovered (see e.g. horton and gerrig, 2005). the controversial notion of ‘intention’ as a psychological state has beenexplicated in terms of hierarchical planning structures (bratman, 1990), a view generally adopted in ai models of communication (cohen et al., 1990). as the gricean individualistic view of speaker’s intention being the sole determinant of meaning underestimates the role of the hearer, dialogue models have turned to bratman’s account ofjoint intentionsto model participant coordination. in this account, joint intentions arise through the composition of appropriately coordinated individual intentions and a network of mutual beliefs. in this respect, the notion of gricean conversational “cooperation” features prominently in h. clark’s account of communication: dialogue involves (intentional) joint actions built on the coordination of individual actions based on shared beliefs (common ground) (see e.g. clark, 1996). hence, a strong gricean element still underlies reasoningabout speakers’ intentions and meaning even though now supported by an account in terms of joint action and conversational structure. thus, even here, the move from individualistic accounts of action, planning and intention to joint action and coordination in dialogue sees the latter as derivative and raises important philosophical and psychological issues that challenge received viewsof meaning and language. 1.2 re-evaluating intentionalism the first type of challenge comes from views that have emerged under theinfluence of late wittgensteinian ideas on language or the refutation of any rigid distinction between natural and non-natural meaning. prominent amongst these are millikan’s teleosemantic approach to language content (millikan, 2005) and brandom’s social-inferential account of communication (brandom, 1994) which severally target aspects of gricean and neo-gricean conceptions of communication. 201 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey millikan argues against gricean approaches to communication from a naturalistic point of view. she argues that the standard gricean view, with its heavy emphasis on mind-reading which is demonstrably not achievable by small children (see below sections 1.3, 1.4), turns the process of language acquisition, a heavily context-dependent process, into a mystery. unlike the gricean conception of meaningnn which rules out causal effects on the audience, e.g. involuntary responses in the hearer, millikan’s account, to the contrary, examines language and communication on the basis of phenomena studied by evolutionary biology, with linguistic understanding seen as analogous to direct perception rather than reasoning:2 objects of direct ordinary perception, e.g. vision, are not less abstract than linguistic meanings. both require contextual filling in through processing of the incoming data in order to be comprehended; yet, in the case of ordinary perception, this processing obviously does not require considering someone’s intention. an analogous assumption can then be made as regards linguistic understanding, so that the resolution of underspecified input in context would not require considering interlocutors’ mental states as a necessary ingredient. millikan then provides an account of linguistic meaning in a continuum with natural meaning based on the function that linguistic devices have been selected to perform (their survival value). these functions are defined through what linguistic entities are supposed to do (not what they normally do or are disposed to do) so that “function”, in millikan’s sense, becomes a normative notion. norms of language, “conventions”, are uses that had survival value, andmeaning is thus equated with function. in contrast then to bratman’s account of intentional action whichsees the planning structures involved as distinctive of rational agents, distinguishing them from entities exhibiting merely purposive behaviour (see e.g. bratman, 1999, 5), in millikan’s naturalistic perspective, function, i.e. meaning, does not depend upon speaker intentions. nonetheless, speakers indeed can be conceived as behaving purposefully in producing tokens of linguistic devices (as hearts and kidneys behave purposefully) but without representing hearers’ mental statesor having intentions about hearers’ mental states (see also csibra and gergely, 1998; csibra, 2008). similarly, hearers understand speech through direct perception of what the speech is about without necessary reflection on speaker intentions. of course, adults can, and often do, use reflections about the interlocutor’s mental states; but this is not a necessary ingredient for meaningful interaction. gricean mechanisms, that is, can be invoked but only as derivative or in cases of failure of the normal functioning of the primary mechanisms involved in the recovery of meaning, such as deception etc. from thisperspective, what the schiffer and strawson scenarios show is that gricean assumptions are on the wrong footing as a foundation for accounts of communication: generalising from these elaborate cases to cases of ordinary interaction is like taking hallucinations as the basis of an account ofveridical perception (for a rejection of this view in the domain of perceptual experiences see e.g. mcdowell, 1982). it is then no wonder that similar paradoxes are generated, e.g. themutual knowledge paradox(clark and marshall, 2002) according to which interlocutors have to compute an infinite series of beliefs in finite time. the dilemma here is that there is plenty of evidence foraudience designin language production, a type of cooperative behaviour, posing the problem of how to model the interlocutors’ abilities allowing them to achieve this during online processing. but the solution tosuch problems ideally should not replicate that problematic structure (see e.g. clark andmarshall (2002), who assume that interlocutors carry around detailed models of the people they know which they consult when they come to interact with them). replacing such accounts with a psychological per2. the strict dichotomy between “meaningnn ” and “showing” has also been disputed within relevance theory (see e.g. wharton, 2003). 202 incrementality and intention -recognition in utterance processing spective that focuses on the mechanisms involved can undercut the intractability of such solutions by invoking independently established low-level memory mechanisms that provide explanation of how people appear to achieve “audience designed” productions withoutin fact constructing explicit models of the interlocutor (see e.g. horton and gerrig (2005) where retrieval of ordinary episodic memory traces predicts both conformity and deviation from the dictates of the common ground idealisation in experimental settings). moreover, by taking seriously the linguistic resources available to the interlocutors, research in conversational analysis has revealed that when these low-level mechanisms fail there are dedicated socially-controlled devices for repairing coordination (see also clark, 1996; ginzburg, forthcoming), devices which allow for a form ofexternalised inference as regards the interlocutors’ purposes. an alternative account of communication combining gricean and millikanesqueperspectives is that of recanati (2004), which makes gricean higher-order intention recognition a prerequisite only for implicature reconstruction. for what he terms “primary processes”, on the other hand, recanati adopts millikan’s account of understanding-as-direct-perception forthe pragmatic processes that are involved in the determination of the truth-conditional content of an underspecified linguistic signal. these processes are blind and mechanical relying on ‘accessibility’ so that no inference or reflection of speaker’s intentions and beliefs is required. it is only ata second stage, for the derivation of implicatures, that genuine reasoning about mental states comes into play. brandom (1994) also eschews the individualistic character of accountsof meaning espoused by the gricean perspective, as part of his rationalist programme for semantics/pragmatics, and, more generally, philosophy. but unlike millikan (and recanati), brandom analyses meaning/intentionality as arising out of linguistic social practices, with meaning, beliefs and intentionsall accounted for in terms of the linguistic game of giving and asking for reasons, a view adoptedin the domain of computational semantics by kibble (2006). the guiding principle behind such social, non-intentionalist explanations of communication and dialogue understanding is to replace mentalist notions such as ‘belief’ with public, observable practical and propositional ‘commitments’, in order to resolve the problems arising for dialogue models associated with the intersubjectivity ofbeliefs and intentions, i.e. the fact that such private mental states are not directly observable and available to the interlocutors. a further motivation arises from the fact that it has been shown that beliefs, goals and intentions underdetermine what “rational” agents will do in conversation: social obligations or conversational rules may in fact either displace beliefs or intentions as the motivation for agents’ behaviour or enter as an additional explanatory factor (traum & allen 1994). brandom’s account presents an inferentialist view of communication which seeks to replace mentalist notions such as belief with public, observable practical and propositional commitments. underthis view (as in asher & lascarides 2008), commitment does not imply ‘belief’ in the usual sense. a speaker may publicly commit to something which she does not believe. and ‘intention’ can be cashed out as the undertaking of a practical commitment or a reliable disposition to respond differentially to the acknowledging of certain commitments.3 from our point of view, the advantage of such non-individualistic, externalist accounts (see also burge, 1986) is that, in not giving supremacy to an exclusively individualist conception of psychological processes, they break the presumed exhaustive dichotomy between behaviourist and mentalist accounts of meaning and behaviour (see e.g. preston, 1994) orcode vs. inferential models 3. an intermediate position is presented by lascarides and asher (2009); asher and lascarides (2008) who also appeal to a notion ofpublic commitmentassociated with dialogue moves but which they link to a parallel cognitive modelling component based on inference about private mental states (see alsotraum and allen, 1994; poesio and traum, 1997). 203 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey of communication (see e.g. krauss and fussell, 1996). instead, ascribing contents to behaviours is achieved by supra-individual social or environmental structures, e.g. conventions, “functions”, practices, routinisations, that act as the context that guides agents’ behaviour. the mode of explanation for such behaviours then does not enforce a representational component, accessible to individual agents, that analyses such behaviours in folk-psychological mentalistic terms, to be invoked as an explanatory factor in the production and interpretation of social action. individual agents instead can be modelled as operating through low-level mechanistic processes without necessary rationalisation of their actions in terms of mental state ascriptions (see e.g. barr (2004) for the establishment of conventions and pickering and garrod (2004) for coordination). this view is consonant with recent results in neuroscience indicating that notions like intentions, agency, voluntary action etc. can be taken as post hoc confabulations rather than causally efficacious (work by benjamin libet, john bargh and read montague, for a survey see wegner, 2002): according to these results, when a thought which occurs to an individual just prior to an action, is seen as consistent with that action, and no salient alternative causes of the action are accessible, the individual will experience conscious will and ascribe agency to themselves. accordingly, when examining human interaction, and more specifically dialogue, notions like intentions and beliefs may enter into common sense psychological explanationsthat the participants themselves can invoke and manipulate, especially when the interaction does not run smoothly. as such, theydo operate as resources that interlocutors can utilise explicitly to account fortheir own and others’ behaviour. in this sense, such notions constitute part of themetalanguage participants employ to make sense of their actions in conscious, often externalised reflections (see e.g. heritage (1984); mills and gregoromichelaki (2010); healey (2008), section 2.2below). cognitive models that elevate such resources to causal factors in terms of plans, goals etc. either risk not doing justice to low-level mechanisms that implement the epiphenomenal effects they describe, or they frame their provided explanations as competence/computational level descriptions(see e.g. stone, 2005, 2004). the stance such models take may be seen as innocuous preliminary idealisation, but this is acceptable only in the absence of either emerging internal inconsistency oralternative explanations that subsume the phenomena under more general assumptions. for example, there are well-known empirical/conceptual problems with the reduction of agent coordination in termsof bratman’s joint intentions (searle, 1990; gold and sugden, 2007);4 and there are also psychological/practical puzzles in cognitive/computational implementations in that the plan recognition problemis known to be intractable in domain-independent planning (chapman, 1987). but, more pertinently for our concerns, cashing out communicative intentions in causal terms via the planning metaphor (bratman, 1990) ignores the fact that the kinds of representations interlocutors actually employ to perform and interpret action do not explicitly deal with intentions or plans (unless these are explicit, conscious deliberations). as argued by suchman (1987/2007); agre and chapman(1990), instead of taking plans and intentions as causal factors inside the agents’ head guiding their action, they should be seen as arising as explicit articulations of antecedent conditions and consequences of past or future action that account for it in a way that can be made sense of by the agents themselves or the interpreters of their behaviours. in that respect, plans and intention attribution have a genuine explanatory role to play in human cognition and interaction but we see no reason to assign similar status, in addition, to mechanisms that are formulated in non-mentalistic, mechanistic terms to which agents have no conscious access. from this perspective, thesemechanisms do not display 4. in addition such accounts of coordination are not general enough inthat they are discontinuous with explanations of collective actions, in e.g. crowd coordination, individuals walking past each other on the sidewalk, etc. 204 incrementality and intention -recognition in utterance processing identical functional roles with folk-psychological concepts (cf stone 2004) and such metaphorical appropriation of such notions in fact obscures the actual function that explicit use of plan and intention attributions play in the agents’ cognition.5 these conclusions can be further substantiated on the basis of empirical and psychological evidence to which we now turn. 1.3 re-evaluating intentionalism: empirical evidence buttressing these foundationalist arguments is a range of psycholinguistic research suggesting that recognition of intentions is an unduly strong psychological condition to imposeas a prerequisite to effective communication. first, there is the problem of autism and related disorders. autism, despite being reliably associated with inability (or at least markedly reduced capacity) to envisage other people’s mental states, is not a syndrome precluding first-language learning in high-functioning individuals (gl̈uer and pagin, 2003). secondly, language acquisition across childrenis established well before the onset of ability to recognise higher-order intentions (wellman et al., 2001), as evidenced by the so-called ‘false-belief task’ which necessitates the child distinguishing what they believe from what others believe (perner, 1991). given that language-learning takes place very largely through the medium of conversational dialogue, these results appear to show that at least communication with and by children cannot rely on higher-order intention recognition. there is also very considerable independent evidence that even though adults are able to think about other people’s perspectives, they are significantly influenced by their own point of view (egocentrism) (keysar, 2007). this suggests that the complex hypotheses requiredby gricean reasoning in communication may not reliably be constructed by adults either.6 this is corroborated by an increasingly large body of research demonstrating that gricean “common ground” is not a necessary building block in achieving coordinative communicative success: speakers regularly violate shared knowledge at first pass in the use of anaphoric and referential expressions which supposedly demonstrate the necessity of established common ground (keysar, 2007, a.o.).7 in accordance with these results, it is a regular occurrence in conversation that both speakers and hearers may elect not to make use of what is well established shared knowledge. on the one hand, in selecting an interpretation, a hearer may fail to check against consistency with what they believe the speaker could have intended (as in (1) where b construes the question in flagrant contradiction to what she knows a knows): (1) a: why don’t you have bean chili? b: beef? youknow i’m a vegetarian [natural data] furthermore, the speaker’s choice of anaphoric expression, supposedly restricted to well established shared knowledge, is regularly made in apparent neglect of what the hearer might take as salient: 5. in addition, it has been argued that use of such folk-psychologicalconstructs are culture/occasion-specific (du bois, 1987; duranti, 1988), hence should not be seen as underpinning general cognitive abilities. 6. indeed, it is useful to note that even adults fail the false belief task, if itis a bit more complex (birch and bloom, 2007). 7. though ‘audience design’ and coordination effects are regularly observed in experiments (see e.g. hanna et al., 2003), these can be shown to result from general memory-retrieval mechanisms rather than as based on some common ground calculation based on metarepresentation or reasoning (see horton and gerrig, 2005; pickering and garrod, 2004). 205 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey (2) a having read out newspaper headline about brown and obama, upon reading next headline provides as follow-on: a: they’ve received 10,000 emails. b: brown and obama? a: no, the camerons. [natural data] given this type of example, checking in parsing or producing utterances that information is jointly held by the dialogue participants the perceivedcommon groundcannot be a necessary condition on such activities.8 one might want to characterise (1)-(2) as dysfunctional uses of language, impaired performance etc. but, firstly, there is psycholinguistic evidence that such neglect of common ground does not significantly impede successful communication and is not even detected by participants (engelhardt et al., 2006, a.o.). secondly, if indeed such data are set aside as unsuccessful acts of communication, one is left without an account of how people manage tounderstand what each other has said in these cases. but it is now well-documented that “miscommunication” not only provides vital insights as to how language and communication operate (schegloff, 1979), but also facilitates dialogue coordination: as healey (2008) shows, the local processes involved in the detection and resolution of misalignments during interaction lead to significantly more positive effects on measures of successful interactional outcomes (see also brennan and schober, 2001). in addition, these localised procedures lead to more gradual, group-level modifications, which in turn account for language change. therefore, the gricean and neo-gricean focus on detecting speaker meaning as the sole criterion of communicative success misrepresents the goals of human interaction: miscommunication (which is an inevitable ingredient of interlocutors that do not share a priori common ground) and the specialised repair procedures made available by the structured linguistic and interactional resources available to interlocutors are the sole means that can guarantee intersubjectivity and coordination; and, as saxton (1997) shows, in addition, such mechanisms, in the form of negative evidence and embedded repairs (see also clark and lappin, 2011), crucially mediate language acquisition (see also goodwin, 1981, 170-171). 1.4 the weakening of gricean assumptions such evidence has led to a move within relevance theory (rt) (sperber and wilson, 1995) weakening further the gricean assumptions (breheny, 2006). the relevance-theoretic view of communication is that the content of an utterance is established by a hearer relative to what the speaker could have intended (relative also to a concept of ‘mutual manifestness’ of background assumptions). this explanation involves meta-representation of other people’s thoughts, but the process of understanding is effected by a mental module enabling hypothesis construction about speaker intentions. as noted by rt researchers, along with the communicated propositions, the context for interpretation falls under the speaker’s communicative intention and the hearer selects it (in the form of a set of conceptual representations) on this basis. so, even though, unlike common ground, mutual manifestness of assumptions are in principle computable by conversational participants, and the interpretation process is not a rational one in the sense of grice, it still remains the case that speaker meaning and intention are the guiding interpretive criteria which are implemented on mechanisms that have evolved to effect mind-reading. for this reason, breheny argues that children in the initial 8. we are not claiming here that explanations for such phenomena cannot be given within standard gricean models, especially since reasoning about speaker’s intention is a form of non-demonstrative inference which can be expected to go wrong. the issue is whether such cases should be seen as deviantor not. 206 incrementality and intention -recognition in utterance processing stages of language acquisition communicate relative to a weaker ‘naive-optimism’ strategy in which some context-established interpretation is simply presumed to match the speaker’s intention, only coming to communicate in the full sense substantially later (see also tomasello, 2008). in effect, this presents a non-unitary view of communication, which, based on the occasional sophistication that adult communicators exhibit, radically separates the abilities of adult communicators from those of children and high-functioning autistic adults. on the other hand, given the known intractability of notions like planning recognition and common ground/mutual knowledge computation, computational models of dialogue, even when based on generally clarkian theories of common ground, have also largely been developed without explicit high-order meta-representations of other parties’ beliefs or intentions except where dealing with complex dialogue domains (e.g. non-cooperative negotiation, traum etal., 2008). with algorithmically defined concepts such asdialogue gameboard, qud, (ginzburg, forthcoming; larsson, 2002) and default rules incorporating rhetorical relations (lascarides and asher, 2009; asher and lascarides, 2008), the necessity for rational reconstruction of inferential intention recognition is largely sidestepped (though see lascarides and asher (2009); asher and lascarides (2008) for discussion). even models that avow to implement gricean notions (see e.g. stone, 2005, 2004) have significantly weakened the gricean reconstruction of the notion of “communicative intention” and meaningnn positing instead representations whose content does not directly reflectthe logical structure (e.g. reflexive or iterative intentions) required by a genuine gricean account. but, in our view, this is not the notion of rationality that grice envisaged. and, as we saidearlier (section 1.2), we see no reason to confuse the postulates of such models with the psychological constructs of the gricean account. in fact, in many respects, these models are directly compatible with the view expressed here, namely, the need for low-level mechanistic explanations ofjoint action based on skills for collaboration and procedural knowledge. 2. incrementality in dialogue another set of major challenges to gricean models of communication arise fromthe radical incrementality of processing in dialogue, and the incremental emergence of ‘joint intentions’ at the level of ‘joint projects’ (bangerter and clark, 2003). 2.1 split utterances the incrementality of on-line processing is now uncontroversial. it has been established for some considerable time now that language comprehension operates incrementally;and, standardly, psycholinguistic models assume that partial interpretations are built more or less ona word-by-word basis (see e.g. sturt and crocker, 1996). more recently, language production has also been argued to be incremental (kempen and hoenkamp, 1987; levelt, 1989; ferreira,1996). guhe (2007) further argues for the incremental conceptualisation of observed events resulting in the generation of preverbal messages in an incremental manner guiding semantic and syntacticformulation. in all the interleaving of planning, conceptual structuring of the message, syntactic structure generation and articulation, incremental models assume that information is processed as itbecomes available, reflecting the introspective observation that the end of a sentence is not planned when one starts to utter its beginning (guhe et al., 2000). in accordance with this, in dialogue, evidence for radical incrementality is provided not merely by the fact that participants incrementally “ground” each other’s contribution (allen et al., 2001) throughback-channelcontributions likeyeah, mhm, etc. but 207 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey also by the observation that people clarify, repair and extend each other’s utterances, even in the middle of an emergent clause: (3) context: friends of the earth club meeting a: so what is that? is that er... booklet or something? b: it’s a book c: book b: just ... talking about al you know alternative d: on erm... renewable yeah b: energy really i think...... a: yeah [bnc:d97] in fact, such completions and continuations have been viewed by herb clark, among others, as some of the best evidence for cooperative behaviour in dialogue (clark, 1996, 238). but even though, indeed, such joint productions demonstrate the communicators’ skill to collaboratively participate in communicative exchanges, this ability to take on or hand over utterances raises the problem of the status of intention-recognition within human interaction. firstly, on the gricean assumption that pragmatic inference in dialogue operates on the basis of reasoning based on evidence of the interlocutor’s intention, delivered by fixing the semantic propositional structure licensed by the grammar, the data in (3) cannot be easily explained, except as causing serious disruptions in normal processing. secondly, on the assumption that communication necessarily involves recognising the propositional content intended by the speaker, there would be an expected cost for the original hearer in having to infer or guess this content before the original sentence is complete, and for the original speaker in having to modify their original intention, replacing it with that of another in order to understand what the new speaker is offering. but, wholly against this expectation, interlocutors very straightforwardly shift out of the parsing role and into the role of producer and vice versa as though they had been in that newly adopted role all along. indeed, it is the case that such interruptions do sometimes occur when the respondent appears to have guessed what they think was intended by the original speaker, what have been calledcollaborative completions: (4) conversation from a and b, to c: a: we’re going to ... b: bristol, where jo lives. (5) a: are you left or b: right-handed. but this is not the only possibility: as (6)-(7) show, such completions by no means need to be what the original speaker actually had in mind: (6) morse: in any case the question was suspect: a very good question inspector [morse, bbc radio 7] (7) daughter: oh here dad, a good way to get those corners out dad: is to stick yer finger inside. daughter: well, that’s one way. [from lerner (1991)] 208 incrementality and intention -recognition in utterance processing in fact, such continuations can be completely the opposite of what the original speaker might have intended as in what we will callhostile continuationsor devious suggestionswhich are nevertheless collaboratively constructed from a grammatical point of view: (8) (a and b arguing:) a: in fact what this shows is b: that you are an idiot (9) (a mother, b son) a: this afternoon first you’ll do your homework, then wash the dishes and then b: you’ll give me£10? furthermore, as all of (4)-(10) show, speaker changes may occur at any point in an exchange (purver et al., 2009), even very early, as illustrated by (10), with the clarification becoming absorbed into the final in-effect collaboratively derived content: (10) a: they x-rayed me, and took a urine sample, took a blood sample. er,the doctor b: chorlton? a: chorlton, mhmm, he examined me, erm, he, he said now they were on about a slide 〈unclear〉 on my heart. [bnc: kpy 1005-1008] this phenomenon has consequences for accounts of both utterance understanding and utterance production. on the one hand, incremental comprehension cannot be based primarily on guessing speaker intentions: for instance, it is not obvious why in (6)-(9), the addressee has to have guessed the original speaker’s (propositional) intention/plan before they offer their continuation.9 on the other hand, speaker intentions need not be fully-formed before production: the assumption of fullyformed propositional intentions guiding production will predict that all the cases above where the continuation is not as expected, as in (6)-(9), would have to involve some kind of revision or backtracking on the part of the original speaker. but this is not a necessaryassumption: as long as the speaker is licensed to operate with partial structures, they can start anutterance without a fully formed intention/plan as to how it will develop (as the psycholinguistic models in anycase suggest) relying on feedback from the hearer to shape their utterance (goodwin,1979). the importance of feedback in co-constructing meaning in communication has been already documented at the propositional level (the level of speech acts) within conversational analysis (ca) (see e.g. schegloff, 2007). however, it seems here that the same processes can operate sub-propositionally, but only relatively to grammar models that allow the incremental, sub-sentential integrationof cross-speaker productions. we turn to one such model next. 2.1.1 modelling the incrementality of split utterances the challenge of modelling the full word-by-word incrementality required in dialogue has recently been taken up, not merely within the dynamic syntax framework, a matter to whichwe return to 9. these are cases not addressed by devault et al. (2009), who otherwise offer a method for getting full interpretation as early as possible. lascarides and asher (2009); asher and lascarides (2008) also define a model of dialogue that partly sidesteps many of the issues raised in intention recognition. but, inadopting the essentially suprasentential remit of sdrt, their model does not address the step-by-step incrementality needed to model split-utterance phenomena. 209 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey in due course, but also by poesio and rieser (2010) (p&r henceforth).p&r set out a dialogue model for german, defining a thorough, fine-grained neo-gricean model of dialogue interactivity that builds on an ltag grammar base. their primary aim is to modelcollaborative completions, as in (4) and (5), in cooperative task-oriented dialogues where take-over by the hearer relies on the remainder of the utterance taken to be understood or inferrable from mutualknowledge/common ground.their account is an ambitious one in that it aims at modelling the generation and realisation of joint intentions which accounts for the production and comprehension ofco-operative completions. the p&r model hinges on two main areas: the assumption of recognition of interlocutors’ intentions according to shared joint plans (bratman, 1992), and the use ofincremental grammatical processing based on ltag. with respect to the latter, this account relies on the assumption of a string-based level of syntactic analysis, for it is this which provides the top-down, predictive element allowing the incremental integration of such continuations. this assumption, however, would seem to impede a more general analysis, since there are cases where split utterances cannot be seen as an extension by the second contributor to the proffered string of words/sentence: (11) eleni: is this yours or yo: yours. [natural data] (12) with smoke coming from the kitchen: a: i’m afraid i burnt the kitchen ceiling b: but have you a: burned myself? fortunately not. in (11), the string of words that the completion yields is not at all what eitherparticipant takes themselves to have constructed, collaboratively or otherwise. and in (12) also, even though the grammar is responsible for the dependency that licenses the reflexive anaphormyself, the explanation for b’s continuation in the fourth turn of (12) cannot be string-based as thenmyselfwould not be locally bound (its antecedent isyou). moreover, in ltag, p& r’s selected syntactic framework, words are defined in terms of syntactic/semantic pairings, relative to a given head, with adjuncts as a means of splitting these. yet, as (7)-(12) indicate, utterance take-over can take place without a head having occurred prior to the split (see also purver et al 2009, howes et al thisvolume), and even across split dependencies (in (13) between an npi and its triggering environment): (13) a: have you mended b: any of your chairs? not yet. given that such dependencies are defined grammar-internally, the grammar has to be able to license such split-participant realisations. but string-based grammars cannot account straightforwardly for many types of split utterances except by defining each part as sententialin its own right. furthermore, if the attempt is to reconstruct speaker’s intentions as part of the interpretation recovered, as p&r explicitly advocate, there is the additional problem that such fragments can play multiple roles at the same time (in (5), (7), (11): question/completion/acknowledgment/answer). not only that but the ca sequential structures (speech acts) normally taken to underpin coherence among propositional turn units, in fact, also operate within such collaborative constructions. for example, such completions might be explicitly invited by the speaker thus forming aquestionanswer pair: 210 incrementality and intention -recognition in utterance processing (14) a: and you’re leaving at ... b: 3.00 o’clock (15) a: and they ignored the conspirators who were b: geoff hoon and patricia hewitt [radio 4, today programme, 06/01/10] (16) jim: the holy spirit is one who〈pause〉 gives us? unknown: strength. jim: strength. yes, indeed.〈pause〉 the holy spirit is one who gives us?〈pause〉 unknown: comfort. [bnc hdd: 277-282] (17) george: cos they〈unclear〉 they used to come in here for water and bunkers you see. anon 1: water and? george: bunkers, coal, they all coal furnace you see, ... [bnc, h5h: 59-61] (18) anon 1: yeah, the all-weather joyce: gallops? anon: yeah, wh what we call the all-weather gallop. [bnc, hdh: 169-171] within the p&r model, such multifunctionality would not be capturable except as a case of ambiguity or by positing hidden constituent reconstruction that has to be subjectto some non-monotonic build-and-revise strategy that is able to apply even within the processing ofan individual utterance. in addition, in fact, in some contexts, invited completions have been argued to exploit the vagueness of the speech act involved to avoid overt elicitation of information (ferrara, 1992): (19) (lana = client; ralph = therapist) ralph: your sponsor before ... lana: was a woman it has to be said that the p&r account is not intended to cover such data, asthe setting for their analysis is one in which participants are assigned a collaborative task with a specific joint goal, so that joint intentionality is fixed in advance and hence anticipatory computationof interlocutor’s intentions can be fully determined; but such fixed joint intentionality is decidedlynon-normal in dialogue and leaves any uncertainty or nondeterminism in participants’ intentions an open challenge. nonetheless, by employing a dynamic view of the grammar, the p&r account marks a significant advance in the analysis of such phenomena. 2.1.2 fragments as incomplete sentences? relative to any other grammatical framework, dialogue exchanges involvingincremental split utterances of any type are even harder to model, given the near-universal commitment to a static performance-independent methodology. first of all, in almost all standard grammar frameworks, it is usually the sentence/proposition that is the unit of syntactic/semantic analysis. fragments then are assigned sentential analyses with semantics provided through ellipsis resolution involving abstraction operations as in dalrymple et al. (1991) (see e.g. purver, 2004; ginzburg and cooper, 2004; ferńandez, 2006). the abstraction is defined over a propositional contentprovided by the previouscontext (in ginzburg’s terms the previously establishedquestion under discussion) to yield appropriate functors to apply to the fragment. of course, multiple optionsof appropriate 211 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey “antecedents” for elliptical fragments are usually available (one for eachpossible abstract). in consequence, some parsing mechanism that is defined to make reference to such a grammar-provided account of ellipsis must appeal to general pragmatic models having to do with recognizing the speaker’s intention in order to select a single appropriate interpretation. but the intention recognition required for disambiguation is unavailable in sub-sentential split utterances in all but the most task-specific domains. this is because, in principle, attribution to any party of recognition of the speaker’s intention to convey some specific propositional content is unavailable until the appropriate propositional formula is established. this is particularly clear where abstraction is required too early in the emergent proposition for there to beany appropriate abstract definable from context as that for which clarification is sought: (10) a: they x-rayed me, and took a urine sample, took a blood sample. er,the doctor b: chorlton? a: chorlton, mhmm, he examined me, erm, he, he said now they were on about a slide 〈unclear〉 on my heart. [bnc: kpy1005-1008] here, the only abstracts that are provided by the context in which b’s clarificatory request of “chorlton?” occurs are (informally) ‘λ x. x took a blood sample’, ‘λy. y-x-rayed me’, ‘λ x. x took a urine sample’, but none of these is the intended basis on which the fragment is uttered: what a presumes is that b can recover the interpretation as a request for clarification as to the identity of the doctor, and this is not of propositional type.10 so such data constitute counter-examples for this style of account: at best, it remains incomplete, needing some other explanation for such early-placed clarifications. such collaboration without necessary recovery of gricean intentions applies not only at the level of the sentence/turn but also at higher levels of discourse organisation (‘joint projects’) as we will discuss immediately below: people begin to interact in order to jointly achieve the completion of a cooperative task without having figured out what they are expected todo as would be predicted by planning models. instead, they expect that, by engaging in the task, interactional routines will emerge that will guide their actions. 2.2 emergent intentions: experimental evidence while core pragmatic research has largely left on one side the phenomenonof collaborative construction of propositional information, the emergence of propositional contents in dialogue has been documented over many years in conversation analysis (ca) (see e.g. lerner 2004). but both ca empirical research (schegloff, 2007) and psycholinguistic experimentssuggest that the same phenomenon can also be observed at higher levels of discourse organisation, the level of ‘joint projects’ (bangerter and clark, 2003). by probing the process of coordinationin task-oriented dialogue experiments it can be demonstrated that notions of joint intentions and plans emerge gradually in a regular manner, rather than guiding utterance production and interpretation throughout. maze-game experiments (see garrod and anderson, 1987) provide a context in which, despite the high-level shared goal for the participants, the lower-level intentions/subplans over which they 10. it may be that ginzburg and cooper (2004)’sconstituentclarification abstraction approach could be successfully employed at such early stages, as it requires only the presence of a recognised syntactic constituent rather than a full sentential proposition. but as currently formulated, it cannot apply to examples like (10) where the clarificational fragmentchorlton is lexically distinct from its antecedentthe doctor. 212 incrementality and intention -recognition in utterance processing have to coordinate are not given from the start. in experimental studies using the chat-tool methodology of healey et al. (2003), this emergence of intentions during the course of the exchange was probed. this methodology allows controlled manipulations to be applied to the dialogue: in this case, artificial clarification requests were inserted by the server into the maze-game task which only one participant (a below) could see. a’s response and the subsequent acknowledgment by the server were also not visible to the other participant:11 (20) a: i’m at the top (target turn) server: top? (artificial turn by server) a: yes (response by a) server: thanks (artificial ack. by server) in such exchanges, highly formulaic conventions emerge, revealing the high levels of coordination among participants. at late stages of a series of games, participants may havedeveloped highly efficient elliptical exchanges like the following: (21) a: 4,5 2,6 1,4 b: 1,2 3,4 7,1 a: 1,2 b: 4,5 a: 1,2 from mills (2007) the interpretation of these fragments crucially relies on the rich sequential structure of the joint project as even homonymous fragments like1,2 above acquire distinct interpretations depending on their position in the sequence (e.g. the second1,2 is interpreted as “i can get to 1,2” whereas the third as “i am now on 1,2”, see mills and gregoromichelaki (2010)). clarification requests surreptitiously inserted into the dialogue can then be interpreted in various ways as revealed by the participants’ responses: sometimes receiving standard interpretations vis-a-vis content (ginzburg and cooper, 2004; purver, 2004; schlangen, 2004) as in (22); but sometimes as querying “what the turn is doing” in the sequence in which it occurs (drew, 1997), as in (23):12 (22) a: 5, 6 (target turn) server: what?/5?(artificial clarification) a: 5 along 6 across (23) a: 5, 6 (target turn) server: what?/5?(artificial clarification) a: go to 5 across, 6 is my switch 11. the chat tool is an experimental resource for carrying out investigation of dialogue, allowing fine-grained interventions over the communicative features of the interaction. participants communicate through a familiar (instantmessenger-like) text-based interface. however, instead of passing turns directly to the appropriate chat clients, each turn is routed via a server. this information can then be used to trigger specific experimental interventions. for example, an artificial clarification request might be issued that appearsto originate from another participant. the recipient responds to the clarification, and the server produces an acknowledgement, neither of which are seen by the other participant. subsequent turns are then transmitted as normal. it has been shown that this can be done without disruption to the dialogue or detection by the participants. 12. for reasons of space, the following are amalgamations of real examples in the data. 213 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey the range of responses to such artificial requests were examined. thisrevealed a differential pattern in cr responses in early vs. late games. during the first few mazes, when the participants are relatively inexperienced in the task, crs are interpreted as querying the referential import of the constituent concerned. at late stages of the interaction, however, both fragment and “what” crs are interpreted significantly more frequently as concerning the intention or plan behind the target utterance, i.e. as questioning what the target-turn as a whole “is doing” in the sequence. these results can be interpreted as follows: empirical ca analyses of the sequential coherence of conversation emphasise the importance of the turn-by-turn organisation of dialogue whichallows juxtaposition of displays of participant understandings and provides structures fororganised repair. rather than interlocutors having to figure out each other’s mental states and plans through metarepresentational means, conversational organisation provides the requisite structure forcoordination. similarly, as garrod and anderson (1987) observe, in maze-game experiments explicit negotiation is neither a preferential nor an effective means of coordination. if it occurs at all, it usually happens after participants have already developed some familiarity with the task. hence, the interactive alignment model developed by pickering and garrod (2004) emphasizes the importance of tacit co-ordination and implicit common ground as the primary means of coordination. the establishment of routines and the significance of repair as externalised inference are also noted by pickering and garrod. the hypothesis that these implicit means, rather than intention recognition, were theprimary method of coordination was further probed here by inserting artificial clarificationsregarding intentions (why?) and observing the responses they receive at initial and later stages of around of games (see mills and gregoromichelaki, 2008/in prep): (24) first few mazes: a: go to the top right of the maze server: why? (artificial clarification) a: dunno/no idea (25) a: can you get to the top of the maze? server: why? (artificial clarification) a: can you get to the top of the maze / try it here too we observe that, at early stages, individuals display little recognition of specific intentions/plans underpinning their own utterance and explicit negotiation is either ignored or more likely to impede (see also mills, 2007; healey, 1997). this is because participantshave not yet figured out the structure of the task, hence they do not have yet developed a metalanguage involving plan and intention attribution in order to explicitly negotiate their purposes. this implies that discursive constructs such as intentions need to emerge, even in such task-oriented joint projects. initially, participants seem to follow trial-and-error strategies to figure out what thetask involves. these strategies and the routines participants develop lead, at later stages of the maze-game, to highly coordinated, efficient interaction, as we saw earlier in (21), where, oncetask expertise is established, participants’ utterances become highly contracted fragments (telescoping). as familiarity with the task and expertise increases, participants seem to disambiguate artificial clarification requests more and more as concerning “intentions” and plans: 214 incrementality and intention -recognition in utterance processing (26) later mazes: a: 5, 6 server: what?/5?/why? (artificial clarification) a: because you’ve got to go there/you asked me to go there these results appear to undermine both accounts of co-ordination that rely on an a priori notion of (joint) intentions and plans (see also clark 1996) and also accounts which rely on some kind of strategic negotiation/agreement to mediate coordination. instead, we take this as evidence that only at the late stages of a round of maze games can the presence of intentions and plans reliably guide the participants’ interpretations and actions. however, even at these late stages, it is not necessary to assume that participants follow some explicit plan or have explicit intentions withrespect to the interpretation of their fragmentary utterances.13 the formulaic structure of their exchanges and the embodied responses that have developed, underpinned by the participants’ increasing expertise with the task, makes it possible for them to still avoid having to work out each other’s intentions to disambiguate the fragments, notwithstanding the potential to resolve any arising confusion by explicit appeal to intentions or plans. in any case, it would not be desirable to assume that the participants do not “communicate” at early stages of the interaction when they still have not figured out what the task involves and how it’s structured. hence, even in such task-specific situations, joint intentionality is not guaranteedab initio but rather has to evolve incrementally with the increasing expertise.14 these observations seem consonant with an alternative approach to planning and intention-recognition according to which forming and recognising suchconstructs is a subordinated activity to the more basic processes that underlie people’s performance (see e.g. suchman, 1987/2007; agre and chapman, 1990). in accordance with this, in ordinary conversation, there is no guaranteethat there is a genuinely shared plan, or that the way the shared utterance evolves is what either party had in mind to say at the outset, indeed obviously not, as otherwise exchanges like the ones in (4) and (7) etc would appear otiose. instead, utterances are shaped genuinely incrementally and “opportunistically” according to feedback by the interlocutor (as already pointed out by clark 1996). grammatical integration of such joint contributions must therefore be flexible enough to allow such switches, with fragment resolutions occurring incrementally before computation of intentions is even possible. 3. dynamic syntax in response to the challenge that such data provide, we turn to dynamic syntax (ds: kempson et al 2001, cann et al 2005) to consider whether forms of correlation between parsing and generation, as they take place in dialogue, can provide a basis for modelling recovery of interpretation in communicative exchanges without reliance on recognition of specific intentional contents. we set out a model of parsing and production mechanisms that makes it possibleto show how, with speaker and hearer using incrementally the same mechanisms for construal,issues about interpretation choice and production decisions may be resolvable solely on the basis of feedback, without reflections on the other party’s mental state. as we shall see, according tothis account (purver et al 2006), what underpins the smooth shift in all joint endeavours of conversation is the incremental, 13. such fragments are argued to be assigned interpretations, throughthe routinisation of sets of actions, formulated as ad hoc idioms (‘ad-hoc concepts’, carston (2002)) (see mills and gregoromichelaki, 2008/in prep). 14. notably, the p&r data involve data collected after task training. 215 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey context-dependent processing shared by parsing and generation, and the tight coordination thereby achievable (similar assumptions underpin the model presented in stone, 2004, 2005, even though distinct conclusions are drawn there as to its implications with respect to the issue of intention recognition in communication). instead of data such as (4)-(19) being problematic, extensive use of mechanisms across interlocutors illustrates the advantages of a ds-style incremental, dynamic account over static models. the incremental licensing of word processing modelled by ds, directly provides for the construction of restricted, contextually salient structural frames within which fragment construal/generation takes place. from a parsing perspective, this allows narrowing down of the threatening multiplicity of interpretations by incrementally weeding out possibilities en route to some commonly shared understanding. but the features of incrementality, predictivity/goal-directedness and context-dependent processing are built into the grammar architecture itself, rather than being external factors imposed by parsing/production mechanisms: each successive processing step relies on a grammatical apparatus which integrates lexical input with essential reference to the contextin order to proceed. under this low-level licensing of incrementally expanding strings and their interpretations, no mechanisms trigger high-level decisions about speaker/hearer intentions as part of the grammar itself. rather, participants are modelled as gradually shaping propositional contents, on aword-by-word basis, drawing on subpersonal, synchronised mechanisms, without having to start with a fully-formed truth-evaluable content in mind. such a view is buttressed by the fact that, as(7)-(12) show, neither party in such role-exchanges are able to know the eventual joint proposition in advance. our ds-based claim then is that communication involves taking risks without requiring mindreading as an essential attribute: success in communication thus characteristically involves cycles of clarification/correction/extension/reformulation etc (“repair strategies”) as essential subparts of the exchange. when modelled non-incrementally, such strategies might lead to theimpression of nonmonotonic repair and the need to revise some otherwise stable context. but pursued incrementally, within a goal-directed architecture, as we shall see, these do not constitutecommunication breakdown or disfluencies, but the normal mechanism of context construction,hypothesised update, and confirmation (see also schegloff, 1979). by building on the assumption thatsuccessful communication may crucially involve subtasks of repair (ginzburg, forthcoming),mechanisms for informational update that underpin interaction can be defined without reliance on(meta-)representing contents of the interlocutors’ mental states as a precondition for successful communication. this is, emphatically, not to deny the rich human capacity for mind-reading but simply to argue that it is not a pre-requisite for effective communicative exchanges to take place. 3.1 dynamic syntax: the formalism ds is a procedure-oriented framework modelling sequential processing.as is displayed in (28) by way of illustration, the build up of interpretation for (27) is monotonic and strictlyword-by-word incremental: 216 incrementality and intention -recognition in utterance processing (27) bob saw someone (28) 0 ?ty(t), ♦ goal−prediction 7−→ 1 ?ty(t), ?〈↓〉ty(e), ?〈↓〉ty(e→ t) ?ty(e),♦ ?ty(e→ t) bob 7−→ 2 ?ty(t), ?〈↓〉ty(e→ t) ty(e), ι, x,bob′(x)ι, x,bob′(x)ι, x,bob′(x) ?ty(e→ t), ♦ saw 7−→ 3 ?ty(t), ?〈↓〉ty(e→ t) ty(e), ι, y, bob′(y) ?ty(e→ t) ?ty(e), ♦ ty(e→ (e→ t)), see′see′see′ someone 7−→ ... 7−→ completion 7−→ 4/tg y < x, ty(t),♦ see′(ǫ, x, person′(x))(ι, y, bob′(y)) ty(e), ι, y, bob′(y) ty(e→ t), see′(ǫ, x, person′(x)) ty(e), ǫ, x, person′(x)ǫ, x, person′(x)ǫ, x, person′(x) ty(e→ (e→ t)), see′ as (28) illustrates, the ds system provides mechanisms that enable the hearer to anticipate and therefore allow incremental word-by-word build-up of representationsof content paired with some word string. amongst such predictive steps are the construction by anticipation of a subjectpredicate schema (stages 0-1 above, with requirements for a subject andpredicate (?〈↓〉ty(e), ?〈↓ 〉ty(e → t)) imposed as a very first step (not illustrated here), and their immediate construction at the second). such a frame then makes possible the identification of the subject as some individual named bob, via processing of the wordbob (stage 2), and then successive steps of identifying the predicate and its internal argument to be paired with verb and object noun-phrase respectively (stages 3-4). these updates then provide input to the compilation by labelledtype-deduction of a propositional representation of content (stage 4). this then as a final step is subject to an algorithm 217 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey of evaluation determining how some assigned scope dependency choices are reflected in the constructed names (here the formulas < x indicates that the existential term binding the variablex is taken as dependent on the event-terms (see gregoromichelaki, 2011; cann, forthcoming). the mechanisms for tree growth and evaluation are identically available to speakers, hence in generation. the only essential difference in production is that the modelling of a speaker’s actions for tree-growth update involve a so-called “goal tree” (tree 4 in (28)) relative to which all intermediate construction steps have to be checked for commensurability, a checking step for which there may be no analogue in parsing (see section 3.2). the notion of incrementality in ds is closely related to another of its features, thegoal-directedness/ predictivityof both parsing and generation (demberg-winterfors, 2010, see also). at each stage of processing,structural predictionsare triggered that could fulfill the goals compatible with the input, in an underspecified manner. representations of the conceptual structure of messages are given as binary trees, formally encoded with the tree logicloft blackburn and meyer-viol (1994). loft is a modal logic with operators〈↑〉, 〈↓〉 〈↑∗〉, 〈↓∗〉 to define the relations of immediate and iterative domination, and to indicate node locations. what is novel about such trees is, on the one hand, that though they constitute a form of syntax, they are not inhabited by words ofthe language – they constitute structures inhabited by (lambda binding) formulae in the epsilon calculus, the selected semantic representation language. furthermore, the mechanisms that definesuch progressive tree construction constitute the sole concept of natural-language syntax whichthe ds grammar provides. the system is goal-directed; and trees are constructed, by starting (in thecontext-independent case) from a radically underspecified goal, theaxiom (the leftmost minimal tree in the illustration provided by (28), and proceeding throughmonotonic updatesof partial orstructurally underspecified trees until some tree is constructed from an input string in which all imposed goals and subgoals are met. every node in a complete tree bears annotations that include the semantic formulae and their type information. crucial for expressing the goal-directedness arerequirements, i.e. unrealized but expected node/tree specifications, indicated by ‘?’ in front of annotations. as the axiom and its immediate subsequent update tree development in (28) indicate, requirements mayalso take a modal form, e.g. the constraint?〈↓〉ty(e→ t), which is a constraint that the daughter be a formula of predicate type. requirements are essential to the tree-growth dynamics. all requirements must be satisfied if the construction process is to lead to a successful outcome, and, as indicated by the requirement for the predicate imposed at stage 2 in (28), these may not be satisfied until substantially later than the point at which they are imposed.15 updates are carried out by means of applying bothcomputationaland lexical actions, which introduce and update nodes, and move the pointer.computational actionsgovern general treeconstructional processes in a broadly top-down manner.16 lexical specifications, equally, induce actions that effect tree-development, providing annotations for nodes,in many cases also inducing the construction of further structure.17 in the update from stage 2 to 3 in (28), for example, the set of lexical actions for the wordseeis applied, yielding the predicate subtree and its annotations. sub15. the pointer,♦, indicates the ‘current node’ in processing, namely the one to be processed next, a constraint which governs word order. 16. this is the characterisation of incrementality adopted by some psycholinguists under the appellation ofconnectedness (sturt and crocker (1996)): an encountered word always gets ‘connected’ to a larger, predicted, tree. 17. for cases ofdislocation, ds employsunfixednodes (not developed in this paper) which is indeed a core notion: such nodes are initially assigned structurally underspecified positions that are subsequently updated (at the point of the gap, in transformational parlance): see kempson et al 2001, cann et al 2005 among others. 218 incrementality and intention -recognition in utterance processing sequent computational actions involve progressive labelled type-deduction decorating non-terminal nodes in the tree strictly bottom-up until the goal defined in the axiom is reached. indeed all actions, computational and lexical, are defined in the same tree-growth vocabulary, so there is free intercalation of the various types of process. thuspartial treesgrow incrementally, driven by procedures associated with particular words as they are encountered while conforming to top-down modal requirements on later development. central to the framework is the modelling of quantification as a process of termconstruction, using theepsilon calculusas the basic formula language (the epsilon calculus is the formal language that employs arbitrary-name terms in predicate logic natural-deduction proofs). all terms are of type e: epsilon terms, as illustrated by (ǫ, x, consultant′(x)). this term constitutes an arbitrary witness of the existentially quantified formula∃x.consultant′(x), as defined by the following equivalence: ψ(ǫ, x, ψ(x)) ≡ ∃x.ψ(x) notice how this equivalence yields the effect that an epsilon term invariablyreflects its containing environment (the predicateψ in the term’s restrictor is a duplicate of the predicate applying on the term). the construction of such terms is induced by actions which incrementally, in part lexically, specify and collect up scope constraints of the formx < y (to be understood as the term with variabley is dependent on the term with variablex). for example, indefinites project epsilon terms subject to the constraint that they are invariably dependent on either another quantifying expression or a term within the temporal specification; names, as iota terms (e.g.ι, x,bob′(x)), are, in contrast, taken to be epsilon terms of widest scope. a final algorithmic step yields the complex structure of the resulting terms as required in the equivalence: thusa consultant arrivedis assigned a propositional formula arrived′(ǫ, x, consultant′(x)) is evaluated as consultant′(a) ∧arrived′(a) where a = (ǫ, x, consultant′(x) ∧arrived′(x)) the overall dynamics is thus one of growth in names as well as in structure. more radical underspecification of formulae at intermediate stages, equally associated with a process of growth, is lexically licensed, for example by pronouns, whichact as simple place-holders for some possibly subsequent identification. these are defined as projecting a metavariable (notated asu,v etc) as a place-holder for some value to be assigned, with an associated type specification, for pronounsty(e). these invariably occur with an associated requirement for a fixed formula value (of the form?∃xfo(x)), making such provision of a value essential to a successful outcome. metavariables are substituted by other terms available in the context as part of the construction process, subject to locality restrictions differentiating e.g. pronouns andreflexives (for details see cann et al., 2005; kempson et al., 2001). a distinctive ds flavour lies in the license for a parse to proceed on the basis of such partial information. indeed, given the type specification but lack of formula value in the processing of a pronoun, the value for such metavariables may be able to established somewhat later, as for example in expletive uses of pronouns, as init is likely that geoff is wrong. in addition to the construction of individual predicate-argument structures, complex trees are obtained through a general tree-adjunction operation that licenses the construction of so-calledlink ed trees. these are pairs of trees sharing information in the form of a shared term, each such tree a 219 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey subdomain in which labelled type-deduction takes place as in the simple structures. these provide a grammar-induced structural form of context. the construction processes determining and then updating such partial tree representations are used to model a range of phenomena.18 for example, in taking definite nps to be anaphoric, we define the definite article as introducing a metavariable as a partial term, inducing also construction of alink transition to allow the construction of a tree providing possibly complex information as a constraint (“presupposition”) on the value to be substituted for the constructed partial term: (29) the man smokes. (30) the man 7−→ ?ty(t) ty(e),u,♦ ?ty(e→ t) ty(t),man’(u) ty(e),u ty(e → t),man′ l the structure on the subject node above is abbreviated as:ty(e),uman′(u). appositional structures, as ina consultant, a friend of jo’s, left, can equally be established as inducing a pair oflink ed structures. alink transition is defined with the effect shown in (31) from a node of typee in which a preliminary epsilon term has been constructed onto alink ed tree introduced with a requirement to develop a term using that very same variable:19 (31) a friend of jo’s 7−→ ?ty(t) ty(e), (ǫ, x, consultant′(x) ∧ friend′(jo′)(x)) ty(cn), (x,consultant′(x)) ty(cn → e), λp.ǫ, p ?ty(e → t) ty(e), (ǫ, x, friend′(jo′)(x)) ty(cn), (x, friend′(jo′)(x)) x friend′(jo′) jo′ friend′ ty(cn → e), λp.ǫ, p l a twinned evaluation rule then combines the restrictors of two such paired termsto yield a composite term on the main tree (unlike the p&r account, this does not involve ambiguityof the head 18. the canonical case is relative clause construal (kempson et al.,2001; cann et al., 2005), where some typee term once processed becomes the context for the projection of one suchlink ed structure, which, when completed, allows the pointer to return to that initial typee term, now enriched by the incorporation of information constructed upon such alink ed (adjunct) tree. 19. in (31), we abbreviate the annotation (ι, y, jo′(y)) to jo′ for simplicity. 220 incrementality and intention -recognition in utterance processing np according to whether a second or subsequent np follows). the fact that the first term has not been completed is no more than the term-analogue of the delaying tactic made available by expletive pronouns and extraposition-from-np constructions, whereby a parse can proceed from some type specification of a node (with attendant metavariable as its formula value),but without completing (evaluating) that formula. just as with expletives, this strategy allows term modification when the pointer returns from its sister node to that only partially constructed term immediately prior to compiling the decorations of its mother: (32) a man has won, someone you know. suchlink ed trees and their development set the scene for a general characterisation of context, ranging over possibly partial trees and their updates.contextin ds is defined as the storage of parse states, i.e., the storing of partial tree, word sequence parsed to date, plus the actions used in building up the partial tree. formally, a parse statep is defined as a set of triples〈t,w,a〉, where:t is a (possibly partial) tree;w is the associated sequence of words;a is the associated sequence of lexical and computational actions (cann et al., 2007). at any point in the parsing process, the contextc for a particular partial treet in the setp can be taken to consist of: a set of triplesp ′ = {. . . , 〈ti,wi, ai〉, . . .} resulting from the previous sentence(s); and the triple 〈t,w,a〉 itself, the subtree currently being processed. anaphora and ellipsis construal generally involve re-use of formulae, structures, and actions from the setc. all fragments illustrated above in (3)-(10) are processed by means of either extending the current tree, or by constructinglink ed structures with transfer of information among them so that one tree providesthe context for another. such fragments are licensed as wellformed by the grammar only relative to such contexts (cann et al., 2007; gargett et al., 2008; kempson et al., 2009). 3.2 parsing/generation coordination this architecture allows a dialogue model in which generation and parsing function in parallel, following exactly the same procedure in the same order. returning to (28), we now pick out the generation steps involved in producingbob saw mary, notated as (compressed) stages 0 to 4. as indicated earlier, generation of this utterance follows precisely the same actions and trees from left to right as in parsing, with the one additional filter, that the complete tree is available as agoal tree from the start (hence the labelling of the complete tree astg). the intuition this reflects is that the eventual message, in this simple context-independent case at least, is known in advance by the speaker and determines the choices to be made. what generation involves,in addition to parse steps, is reference totg to check whether each attempted generation stage (1, 2, 3, 4) is consistentwith it. according to this algorithm, asubsumptioncheck is carried out as to whether the current parse tree is monotonically extendible totg.20 the trees 1-3 are licensed because, for each of these, the subsumption relation totg is maintained. each time then the generator applies a lexical action, it is licensed to produce the word that carries that action only under successful subsumption check: at stage 3, for example, the generator processes the lexical action which results in the annotationsee′, and upon success and subsumption oftg license to generate the wordseeensues. for processing split utterances, two more consequences are pertinent.first, there is nothing to prevent speakers initially having only a partial structure to convey, i.e.tg may be apartial tree: 20. in fact, the goal tree for the speaker need only be subsumed by atleast one parse step, and in all nonfinal steps in the generation process non-trivially. 221 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey this is unproblematic, as all that is required by the formalism is monotonicity of treegrowth, and the subsumption check is equally well defined over partial trees. second,the goal treetg may change during generation of an utterance, as long as this change involves monotonic extension; and continuations/reformulations/extensions across speakers are straightforwardly modelled in ds by appending alink ed structure annotated with added material to be conveyed (preserving monotonicity) as in single speaker utterances: (33) a friend is arriving, with my brother, maybe with a new partner. such a model under which the speaker and hearer essentially follow the same sets of actions, each incrementally updating their semantic representations, allows the hearerto mirror the same series of partial trees as the producer, albeit not knowing in advance the content of the unspecified nodes. furthermore, not only can the same sets of actions be used for both processes, but also a large part of the parsing and generation algorithms is shared. in particular, the processing actions of both parsing and production involve the same progressive growth of partial tree representations, this being the only concept of “syntax” in the ds model. even the concept ofgoal tree, tg, may be shared between speaker and hearer, in so far as the hearer may havericher expectations relative to which the speaker’s input is processed, as in the processing of a clarification question. conversely, the speaker may have only a partial tree astg, relative to which they are seeking clarification. in general, as no intervening level of syntactic structure over the string isever computed, the parsing/generation tasks are more parsimonious in terms of representationsthan in other frameworks. additionally, the top-down architecture in combination with partiality allowsthe framework to be (strategically) more radically incremental in terms of interleaving planning and production than is possible within other frameworks. on the one hand, there is one less level of representation to be computed, so no need for a complex step-by-step correlation of syntactic and semantic output, and no recourse either to some externally imposed parser to ensure such correlation. on the other hand, the licensing of partial structures allows articulation before a completepropositional goal has been determined and, therefore, interlocutor suggestions can be integrated without the need for revision. 4. split utterances in dynamic syntax split utterances follow as an immediate consequence of these assumptions. for dialogues (7)-(12), a reaches a partial tree of what she has uttered through successive updates, while b as the hearer, follows the same updates to reach the same representation of what he has heard: they both apply the same tree-construction mechanism which is none other than their effectively shared grammar.21 this provides b with the ability at any stage to become the speaker, interruptingto continue a’s utterance, repair, ask for clarification, reformulate, or provide a correction, as and when necessary. according to ds assumptions, repeating or extending a constituent of a’s utterance by b is licensed only if b, the hearer now turned speaker, entertains a message to be conveyed (a newtg) that matches or extends in a monotonic fashion the parse tree of what he has heard. this message (tree) may of course be partial, as in (10), where b is adding a clarificational link ed structure to a still-partially parsed antecedent, or it may complete the tree as in (12) and elsewhere. importantly, in ds, both a and b can now re-use the already constructed (partial) parse tree in their immediate context as a point from which to begin parsing and generation,rather than having 21. a completely identical grammar is, of course, an idealisation but one that is harmless for current purposes. 222 incrementality and intention -recognition in utterance processing to rebuild an entirely novel tree or subtree. by way of illustration, we take a simplified variant of (12): (34) ann: did you burn bob: myself? no. here, of course, the reconstruction of the string as*did you burn myself?is unacceptable (at least with a reflexive reading ofmyself), illustrating the problem of purely syntactic accounts of split utterances. but under ds assumptions, with representations only of informational content, not of putative structure over strings of words, the switch of person is entirely straightforward. consider the partial tree induced by parsing a’s utterancedid you burnwhich involves a substitution of the metavariable projected byyouwith the name of the interlocutor/parser:22 (35) did you burn 7−→ ?ty(t), q ?ty(e), ty(e), u, ?∃xfo(x),bob′ ?ty(e→ t) ?ty(e),♦ ty(e→ (e→ t)), burn′ at this point, bob can complete the utterance with the reflexive as what such an expression does, by definition, is copy a formula from a local co-argument node onto the current node, just in case that formula satisfies the conditions set by the person and number of the uttered reflexive, in this case, that it names the current speaker: (36) myself 7−→ ?ty(t), q ty(e), bob′ ?ty(e→ t) ty(e), bob′ ty(e→ (e→ t)), burn′ hence the absence of a “syntactic” level of representation distinct fromthat of such semantic representations allows the direct successful integration of such fragments through the grammatical mechanisms themselves, rather than necessitating their analysis as sentential ellipsis. further, to illustrate how ds can sidestep the problems posed by abstraction accounts of ellipsis, we take a simplified version of (10): (37) a: the doctor b: chorlton? after processingthe doctorboth a and b share a context comprising a partial tree as follows: 22. the featureq on a decorated node is not taken to have a fixed speech-act content: given the range of acts achievable by interrogative structures (as diverse asyes-noquestions,wh questions, tag questions, exclamatives, etc.) we take interrogative forms to encode a direction by the speaker to the hearer for a particular type of coordination, here notated simply asq. 223 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey (38) a’s/b’s context: the doctor 7−→ ?ty(t) udoctor′(u) ?∃xfo(x),♦ ?ty(e→ t) at the next stage of processing, let’s assume that b fails to find a secure substitution for the metavariableu on the subject, thus being forced to request clarification if the requirementis to be satisfied. notice that at this point a and b’s contexts will diverge since a presumably knows who he’s referring to, i.e. has a substituend for the metavariable introduced by the definite description. now b’s goal tree for his request for clarification is: (39) b’s goal-treetg ?ty(t) udoctor′(u) (ι, x, chorlton′(x))doctor′(ι,x,chorlton′(x)) ?ty(e → t) 〈 l−1 〉 tn(n) (ι, x, chorlton′(x)), q l the link transition, which accommodates an additional property (that the individual being talked about is namedchorlton), takes the partial tree in (38) as its context. in this context, with the pointer at the subject node, the building of alink relation is licensed and this is duly constructed. now by uttering the wordchorlton? a new tree can be constructed for b which indeed subsumes the goal tree of (39): (40) b’s construction-tree link−adjunction 7−→ chorlton 7−→ ?ty(t) udoctor′(u),♦ ?ty(e → t) 〈 l−1 〉 tn(n) (ι, x, chorlton′(x)), ql now regular anaphoric substitution allows the metavariableu to be instantiated by the term (ι, x, chorlton′(x)), indeed essentially, as otherwise the two nodes will not be developed as involving any shared term. the result of this process will be exactly the tree in (39) and speaker and hearer context trees will be identical at this point. as illustrated here, the most recent (partial) parse tree constitutes the most immediately available local “antecedent” for fragment resolution; hence no separate computation or definition ofsalienceor speaker intention by the hearer is necessary for such incremental fragment construal. as in p&r, the mechanism is exactly thatof apposition, the building of a link ed structure, in this case, the result of that transition in its turn being used to provide a value for the metavariable place-holder associated with the definite article. as we saw, the hearer, b, may respond to what he has constructed during interpretation, anticipating a’s verbal completion as in (4) and (5). this is facilitated by the general predictivity/goaldirectedness of the ds architecture since the parser is always predictingtop-down goals (requirements) to be achieved in the next steps (see stage 2 of (28) or e.g. (38)). suchgoals are indeed what 224 incrementality and intention -recognition in utterance processing drives the search of the lexicon (lexical access) in generation, so a hearer who shifts to successful lexicon search before processing the anticipated lexical input providedby the speaker can become the generator and take over. in all the cases of split utterances, the original hearer is, indeed, using such anticipation to take over and offer a completion that, even though grammatically licensed, i.e. fitting the predicted structure of the context tree, might not necessarily be identical to the one the original speaker would have accessed had they been allowed to continuetheir utterance as in (7)(9).23 from this point of view, since both speakers and hearers are licensed tooperate with partial structures, speakers can start an utterance without a fully-formed intention/plan as to how it will develop (as the psycholinguistic models in any case suggest) relying on feedback from the hearer to shape their utterance: (41) a: oh. they don’t mean us to be friends, you see. so if we want to be . . . b: which we do a: then we must keep it a secret. [natural data] hence the assumption of underspecified partial speaker contents before the beginning of articulation allows genuine collaboration in the construction of utterances (goodwin, 1979), without necessarily having to resort to revision and backtracking. 5. summary evaluation with grammar mechanisms defined as inducing growth of information that is used symmetrically and incrementally in both parsing and generation, the availability of derivations for genuine dialogue phenomena, like split utterances, from within the grammar, shows how core dialogue activities can take place without any other-party meta-representation at all (thoughuse of reasoning over mental states is not precluded either). on this view, as we emphasised earlier, communication is not definitionally the full-blooded intention-recognising activity presumed bygricean and postgricean accounts. rather, speakers can, on this view, air propositional and other structures with no more than the vaguest of planning and commitments as to what they are going to say, expecting feedback to fully ground the significance of their utterance, to fully specify their intentions (see e.g. wittgenstein, 1953, 337). hearers, similarly, may signally fail to reconstruct putative intentions of their interlocutor as a filter on how to interpret the provided signal; instead, they are expected to provide evidence of howthey perceive the utterance in order to arrive at a joint interpretation. this view of dialogue, though not uncontentious, is one that has been extensively argued for, under distinct assumptions, in the ca literature. according to the proposed ds model of this phenomenon, the core ingredient of dialogue is incremental, context-dependent processing, implemented by a grammar architecture that reconstructs “syntax” as a goal-directed activity, able to seamlessly integrate with the joint activities people engage in. incrementality is afacilitator both for allowing entirely individualistic decisions as to what to say and how, and also for, nevertheless, making possible a joint activity in which an emergent structure canunfold through the 23. it might be argued that back-channels (mhm, yesetc), are problematic for this account. however, arguably, such signals do not encode recognition of intentional content even though this isoften the interpretation assigned in context: even the canonical content-agreement device,yes, is systematically used to signal merely shared attention or license for the interlocutor to continue without an explicit propositional content necessarily being available: (i) sue (opening conversation): tom/ (or knocking on tom’s door) tom (in response): yes? sue: are you busy? 225 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey requesting and receiving of feedback. the two properties that determine how such a weak underpinning can nevertheless yield the coordinative effect achieved in dialogue exchanges are the intrinsic predictivity/goal-directedness in the formulation of ds, and the factthat both parsing and production can have arbitrary partial goals, so that, in effect, both interlocutors are able to be building structures in tandem. in particular, because of the assumed partiality of goal trees (messages), speakers do not have to be modelled as having fully-formed messages, reflecting intentions with propositional content, at the beginning of the generation task, but can instead be viewed as relying on feedback to shape their unfolding utterance. as goal trees are expanded incrementally, completions/repairs/feedback by the other party can be monotonically accommodated, even though they might not represent what the speaker would have uttered if not interrupted. as long as what emerges as the eventual joint content is some compatible extension of the original speaker’s goal tree, itmay be accepted as sufficient for the purposes to hand. thus, in such an incremental model,repair procedures do not deviate from the normal processing mechanisms. in fact, the very possibility of some types of repair, e.g. mid-utterance self-repairs, requires the licensing of partial strings by the grammar (see e.g. ginzburg et al., 2007). but further than that, an incremental syntacticmodel licenses strings and their interpretations on a word-by-word basis and thus can naturally integrate any “repairs” via the already assumed progressive accumulation of information. hence “repair” phenomena naturally emerge as “coordination devices” (clark, 1996), devices exploiting mutually salient contexts for achieving coordination enhancement. and jointly constructed content that isestablished through cycles of “miscommunication” and “repair” is more securely coordinated (see e.g. healey, 2008, and section 1.3 above) and thus can form the basis of what each party considers shared cognitive context. in addition, the ca notion of ‘sequentiality’ that, in our view, can operate in dialogue both sub-propositionally (see (14)-(18)) and across turns through the turn-taking system has also the grammar as its most significant determinant (de ruiter et al., 2006). the dsmodel captures this naturally as the notion of ‘projection’ (schegloff, 1987) that underlies the possibility of harmonious turn-taking is integrated in the goal-directed/predictive architectureof the grammar, which requires the parser to constantly make assumptions as to what is licensed to follow, given an already established goal. overall then, given that such core ca notions, generally assumed to reflect the efficiency, social aptness and highly organised nature of conversation, can be modelled as consequences of the operation of a low-level mechanism like the grammar, the view of communication that emerges here does not require essential grounding in having to recognize speaker’s intentions, hence can be taken to be displayed equally by both young children and adults. one might argue against this view of communication that the phenomenon of conversational implicature, in which speakers may direct hearers to the construction of additional hypotheses to yield indirect inferential effects, necessitates an essentially meta-representational view of communication and explicit representation of a speaker’s intentions with respect to their interlocutor. however, there are alternative accounts of implicatures where, even though situatedinference is involved, the explanations do not necessarily invoke an interlocutor metarepresentational component (see e.g. gauker, 2001) (see also arundale, 2008; haugh, 2008). such inferences might necessitate overt modelling of the interlocutor but not essentially. accordingly, there is no restriction in the view proposed here on the types of representation participants may construct,so nothing precludes the construction of richer contexts to yield such effects.24 24. such richer contexts and consequent derived implications could bemodelled via the construction of appropriately term-sharinglink ed trees, whose mechanisms for construction are independently available in ds. 226 incrementality and intention -recognition in utterance processing this then enables a new perspective on the relation between linguistic ability and the use of language, constituting a position intermediate between the philosophical stances of millikan and brandom, and one which is notably close to that of recanati (2004). linguistic ability is grounded in the control of (low-level) mechanisms (see e.g. böckler et al., 2010) which enable the progressive construction of structured representations to pair with the overt signals of the language, used in conjunction with some generally available cognitive filter for determining particular choices made. the content of these representations is ascribed, negotiated and accounted for in context, via the interaction among interlocutors. constructing representations of the other participants’ mental states, though a possible means of securing communication, is by no means necessary. with dynamic syntax being a grammar formalism, we have not here had anything formal to say about the choice mechanism that selects interpretations from those made available through linguistic processing, although, given the millikan view of communication and the psycholinguistic evidence favouring low-level mechanisms (e.g. pickering and garrod 2004, horton and gerrig 2005, keysar 2007), we believe, along with recanati, that such a mechanism does not operate through the implementation of gricean assumptions. but, on this view, whatever the underpinnings of such a mechanism (e.g. relevanceas defined in sperber and wilson (1995); or rhetorical relations or the cognitive logic of lascarides and asher (2009); asher and lascarides (2008)),it interacts stepwise with the implementation of the resources for interaction that are provided by the grammar (see also ginzburg, forthcoming; cooper and ranta, 2008). hence we suggest, contra tomasello (2008), that we need to be exploring accounts of human communication as an activity involving emergent agent coordination without high-level mind-reading as a prerequisite skill. acknowledgments we gratefully acknowledge helpful discussions with alex davies, arash eshghi, chris howes, graham white, hannes rieser and robin cooper. we would also like to thankhannes rieser for editorial assistance and the reviewers, especially one of them, for detailedand incisive comments that vastly improved our clarity as regards the issues involved. this work issupported by the dynamics of conversational dialogue (dyndial) esrc-res-062-23-0962; leverhulme trust major research fellowship f00158bf for r. cann; marie curie iof fellowship(2010-2012) grant no: piof-ga-2009-236632-eris for gregory mills. references p.e. agre and d. chapman. what are plans for?robotics and autonomous systems, 6(1-2):17–34, 1990. j. allen, g. ferguson, and a. stent. an architecture for more realistic conversational systems. in proceedings of the 2001 international conference on intelligent user interfaces (iui), january 2001. n. allott. paul grice, reasoning and pragmatics.ucl working papers in linguistics, pages 217– 243, 2005. r.b. arundale. against (gricean) intentions at the heart of human interaction. intercultural pragmatics, 5(2):229–258, 2008. 227 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey n. asher and a. lascarides. commitments, beliefs and intentions in dialogue.proceedings of londial, pages 35–42, 2008. k. bach. the semantics-pragmatics distinction: what it is and why it matters.linguistische berichte, 8(1997):33–50, 1997. k. bach and r.m. harnish.linguistic communication and speech acts. mit press, 1982. a. bangerter and h.h. clark. navigating joint projects with dialogue.cognitive science, 27(2): 195–225, 2003. d.j. barr. establishing conventional communication systems: is common knowledge necessary? cognitive science, 28(6):937–962, 2004. sa birch and p. bloom. the curse of knowledge in reasoning about falsebeliefs. psychological science, 18(5):382, 2007. patrick blackburn and wilfried meyer-viol. linguistics, logic and finite trees. bulletin of the igpl, 2:3–31, 1994. a. böckler, g. knoblich, and n. sebanz. socializing cognition.towards a theory of thinking, pages 233–250, 2010. r.b. brandom.making it explicit: reasoning, representing, and discursive commitment. harvard univ pr, 1994. m. bratman. faces of intention: selected essays on intention and agency. cambridge univ pr, 1999. michael e. bratman. what is intention? in p. cohen, j. morgan, and m. pollack, editors,intentions in communication. mit press, 1990. michael e. bratman. shared cooperative activity.philosophical review, 101:327–341, 1992. richard breheny. communication and folk psychology.mind & language, 21(1):74–107, 2006. s.e. brennan and m.f. schober. how listeners compensate for disfluencies in spontaneous speech.journal of memory and language, 44(2):274–296, 2001. t. burge. individualism and psychology.the philosophical review, 95(1):3–45, 1986. ronnie cann. towards an account of the english auxiliary system. in r. kempson, gregoromichelaki e., and c. howes, editors,the dynamics of lexical interfaces, chapter 9. csli, forthcoming. ronnie cann, ruth kempson, and lutz marten.the dynamics of language. elsevier, oxford, 2005. ronnie cann, ruth kempson, and matthew purver. context and well-formedness: the dynamics of ellipsis. research on language and computation, 5(3):333–358, 2007. 228 incrementality and intention -recognition in utterance processing robyn carston.thoughts and utterances: the pragmatics of explicit communication.blackwell, 2002. d. chapman. planning for conjunctive goals.artificial intelligence, 32(3):333–377, 1987. a. clark and s. lappin.linguistic nativism and the poverty of the stimulus.wiley-blackwell, 2011. herbert h. clark.using language. cambridge university press, 1996. h.h. clark and c.r. marshall. definite reference and mutual knowledge.psycholinguistics: critical concepts in psychology, page 414, 2002. p.r. cohen, j.l. morgan, and m.e. pollack.intentions in communication. the mit press, 1990. r. cooper and a. ranta.natural languages as collections of resources. college publications, 2008. g. csibra. goal attribution to inanimate agents by 6.5-month-old infants.cognition, 107(2):705– 717, 2008. g. csibra and g. gergely. the teleological origins of mentalistic action explanations: a developmental hypothesis.developmental science, 1(2):255–259, 1998. m. dalrymple, s. m. shieber, and f. c. n. pereira. ellipsis and higher-order unification.linguistics and philosophy, 14(4):399–452, 1991. j.p. de ruiter, h. mitterer, and n.j. enfield. projecting the end of a speakers turn: a cognitive cornerstone of conversation.language, 82(3):515–535, 2006. v. demberg-winterfors.a broad-coverage model of prediction in human sentence processing. phd thesis, 2010. d. devault, k. sagae, and d. traum. can i finish? inproceedings of the sigdial 2009 conference, pages 11–20, london, uk, 2009. p. drew. open’class repair initiators in response to sequential sourcesof troubles in conversation. journal of pragmatics, 28(1):69–101, 1997. j.w. du bois. meaning without intention: lessons from divination.papers in pragmatics, 1(2), 1987. a. duranti. intentions, language, and social action in a samoan context.journal of pragmatics, 12 (1):13–33, 1988. p.e. engelhardt, k.g.d. bailey, and f. ferreira. do speakers and listeners observe the gricean maxim of quantity?journal of memory and language, 54(4):554–573, 2006. r. ferńandez. non-sentential utterances in dialogue: classification, resolution and use. phd thesis, king’s college london, university of london, 2006. k. ferrara. the interactive achievement of a sentence: joint productions in therapeutic discourse. discourse processes, 15(2):207–228, 1992. 229 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey v. ferreira. is it better to give than to donate: syntactic flexibility in languageproduction.journal of memory and language, 35:724–755, 1996. a. gargett, e. gregoromichelaki, c. howes, and y. sato. dialogue-grammar correspondence in dynamic syntax. inproceedings of the 12th semdial (londial), 2008. s. garrod and a. anderson. saying what you mean in dialogue: a study inconceptual and semantic co-ordination.cognition, 27:181–218, 1987. c. gauker. situated inference versus conversational implicature.noûs, 35(2):163–189, 2001. b. geurts.quantity implicatures. cambridge university press, to appear. j. ginzburg.the interactive stance: meaning for conversation. forthcoming. j. ginzburg and r. cooper. clarification, ellipsis, and the nature of contextual updates in dialogue. linguistics and philosophy, 27(3):297–365, 2004. j. ginzburg, r. ferńandez, and d. schlangen. unifying self-and other-repair. inproceeding of decalog, the 11th international workshop on the semantics and pragmatics of dialogue (semdial07), 2007. k. glüer and p. pagin. meaning theory and autistic speakers.mind & language, 18(1):23–51, 2003. n. gold and r. sugden. collective intentions and team agency.journal of philosophy, 104(3): 109–37, 2007. c. goodwin. the interactive construction of a sentence in natural conversation.everyday language: studies in ethnomethodology, pages 97–121, 1979. c. goodwin. conversational organization: interaction between speakers and hearers. academic press new york, 1981. e. gregoromichelaki. conditionals in dynamic syntax. in r. kempson, gregoromichelaki e., and c. howes, editors,the dynamics of lexical interfaces, chapter 8. csli, 2011. p. grice. logic and conversation.1975, pages 41–58, 1975. p. grice. presupposition and implicature.radical pragmatics, pages 183–199, 1981. p. grice.aspects of reason. oxford: clarendon press, ed. by r. warner edition, 2001. m. guhe. incremental conceptualization for language production. nj: lawrence erlbaum associates, 2007. m. guhe, c. habel, and h. tappe. incremental event conceptualizationand natural language generation in monitoring environments. inproceedings of the first international conference on natural language generation-volume 14, pages 85–92. association for computational linguistics, 2000. j.e. hanna, m.k. tanenhaus, and j.c. trueswell. the effects of commonground and perspective on domains of referential interpretation.journal of memory and language, 49(1):43–61, 2003. 230 incrementality and intention -recognition in utterance processing m. haugh. intention in pragmatics.intercultural pragmatics, 5(2):99–110, 2008. patrick healey. expertise or expert-ese: the emergence of task-oriented sub-languages. in m.d. shafto and p. langley, editors,proceedings of the 19th annual conference of the cognitive science society, pages 301–306, stanford, california, august 1997. stanford university. patrick healey. interactive misalignment: the role of repair in the development of group sublanguages. in r. cooper and r. kempson, editors,language in flux. college publications, 2008. patrick healey, matthew purver, james king, jonathan ginzburg, and gregory j. mills. experimenting with clarification in dialogue. inproceedings of the 25th annual meeting of the cognitive science society, boston, massachusetts, august 2003. j. heritage.garfinkel and ethnomethodology. polity, 1984. w.s. horton and r.j. gerrig. conversational common ground and memory processes in language production.discourse processes, 40(1):1–35, 2005. g. kempen and e. hoenkamp. an incremental procedural grammar for sentence formulation.cognitive science, 11(2):201–258, 1987. ruth kempson, wilfried meyer-viol, and dov gabbay.dynamic syntax: the flow of language understanding. blackwell, 2001. ruth kempson, eleni gregoromichelaki, and yo sato. incrementality, speaker/hearer switching and the disambiguation challenge. inproceedings of european association of computational linguistics proceedings., 2009. b. keysar. communication and miscommunication: the role of egocentric processes.intercultural pragmatics, 4(1):71–84, 2007. r. kibble. reasoning about propositional commitments in dialogue.research on language & computation, 4(2):179–202, 2006. r. m. krauss and s. r. fussell. social and psychological models of interpersonal communication. in e.t. higgins and a.w. kruglanski, editors,social psychlogy: handbook of basic principles, pages 655–701. guildford, 1996. s. larsson. issue-based dialogue management. phd thesis, g̈oteborg university, 2002. also published as gothenburg monographs in linguistics 21. a. lascarides and n. asher. agreement, disputes and commitments in dialogue. journal of semantics, 26(2):109, 2009. g.h. lerner. on the syntax of sentences-in-progress.language in society, pages 441–458, 1991. g.h. lerner. collaborative turn sequences. inconversation analysis: studies from the first generation, pages 225–256. john benjamins, 2004. w.j.m. levelt.speaking: from intention to articulation. mit pr, 1989. 231 gregoromichelaki, kempson, purver, m ills , cann , meyer-v iol , healey s. c. levinson.pragmatics. cambridge textbooks in linguistics. cup, 1983. d. marr. vision: a computational investigation into the human representation and processing of visual information. henry holt and co., inc. new york, ny, usa, 1982. j. mcdowell. criteria, defeasibility, and knowledge. inproceedings of the british academy, volume 68, pages 455–79, 1982. r.g. millikan. language: a biological model. oxford university press, usa, 2005. g. j. mills. semantic co-ordination in dialogue: the role of direct interaction.phd thesis, queen mary university of london, 2007. g. j. mills and e. gregoromichelaki. coordinating on joint projects. based on talk given at the coordination of agents workshop, nov 2008, kcl, 2008/in prep. g. j. mills and e. gregoromichelaki. establishing coherence in dialogue: sequentiality, intentions and negotiation. inproceedings of the 14th semdial, pozdial, 2010. j. perner.understanding the representational mind. mit press cambridge, ma, 1991. m. pickering and s. garrod. toward a mechanistic psychology of dialogue. behavioral and brain sciences, 27:169–226, 2004. m. poesio and h. rieser. completions, coordination, and alignment in dialogue. dialogue and discourse, 1(1):1–89, 2010. m. poesio and d. traum. conversational actions and discourse situations. computational intelligence, 13(3), 1997. b. preston. behaviorism and mentalism: is there a third alternative?synthese, 100(2):167–196, 1994. m. purver. the theory and use of clarification requests in dialogue. phd thesis, university of london, 2004. m. purver, chris howes, eleni gregoromichelaki, and patrick healey. split utterances in dialogue: a corpus study. 2009. f. recanati.literal meaning. cambridge univ pr, 2004. m. saxton. the contrast theory of negative input.journal of child language, 24(01):139–161, 1997. e.a. schegloff. the relevance of repair to syntax-for-conversation. syntax and semantics, 12: 261–286, 1979. e.a. schegloff. recycled turn beginnings.talk and social organization, pages 70–85, 1987. e.a. schegloff.sequence organization in interaction: a primer in conversation analysis i. cambridge univ pr, 2007. 232 incrementality and intention -recognition in utterance processing s.r. schiffer.meaning. oxford university press, usa, 1972. d. schlangen. causes and strategies for requesting clarification in dialogue. inproceedings of the 5th sigdial workshop on discourse and dialogue, boston, massachusetts, april 2004. association for computational linguistics. j. r. searle. collective intentions and actions. in philip r. cohen, j. morgan, and m. e. pollack, editors,intentions in communication, pages 401–415. mit press, 1990. d. sperber and d. wilson.relevance: communication and cognition (2nd editn). blackwell, oxford, 1995. m. stone. intention, interpretation and the computational structure of language. cognitive science, 28(5):781–809, 2004. m. stone. communicative intentions and conversational processes in human-human and humancomputer dialogue. in j. trueswell and m. tanenhaus, editors,approaches to studying worldsituated language use, pages 39–70. mit press, 2005. pf strawson. intention and convention in speech acts.the philosophical review, 73(4):439–460, 1964. p. sturt and m. crocker. monotonic syntactic processing: a cross-linguistic study of attachment and reanalysis.language and cognitive processes, 11:448–494, 1996. l.a. suchman.plans and situated actions: the problem of human-machine communication. cambridge university press, 1987/2007. m. tomasello.origins of human communication. mit, 2008. d. traum and j. allen. towards a formal theory of repair in plan execution and plan recognition. procedings of uk planning and scheduling special interest group, 1994. d. traum, s. marsella, j. gratch, j. lee, and a. hartholt. multi-party, multi-issue, multi-strategy negotiation for multi-modal virtual agents. in8th international conference on intelligent virtual agents., d.m. wegner.the illusion of conscious will. mit press, 2002. h.m. wellman, d. cross, and j. watson. meta-analysis of theory-of-mind development: the truth about false belief.child development, 72(3):655–684, 2001. t. wharton.pragmatics and the showing–saying distinction. phd thesis, ucl, 2003. l. wittgenstein.philosophical investigations, trans. g.e.m. anscombe. oxford: blackwell, 1953. 233 dialogue and discourse 5(1) (2014) 1-36 doi: 10.5087/dad.2014.101 corpus-driven semantics of concession: where do expectations come from?∗ livio robaldo robaldo@di.unito.it department of computer science university of torino torino, italy eleni miltsakaki elenimi@seas.upenn.edu institute for research in cognitive science university of pennsylvania philadelphia, usa editor: raquel fernández abstract concession is one of the trickiest semantic discourse relations appearing in natural language. many have tried to sub-categorize concession and to define formal criteria to both distinguish its subtypes as well as for distinguishing concession from the (similar) semantic relation of contrast. but there is still a lack of consensus among the different proposals. in this paper, we focus on those approaches, e.g. lagerwerf (1998), winter and rimon (1994), and korbayova and webber (2007), assuming that concession features two primary interpretations, “direct” and “indirect”. we argue that this two way classification falls short of accounting for the full range of variants identified in naturally occurring data. our investigation of one thousand concession tokens in the penn discourse treebank (pdtb) reveals that the interpretation of concessive relations varies according to the source of expectation. four sources of expectation are identified. each is characterized by a different relation holding between the eventuality that raises the expectation and the eventuality describing the expectation. we report a) a reliable inter-annotator agreement on the four types of sources identified in the pdtb data, b) a significant improvement on the annotation of previous disagreements on concession-contrast in the pdtb and c) a novel logical account of concession using basic constructs from hobbs (1998)’s logic. our proposal offers a uniform framework for the interpretation of concession while accounting for the different sources of expectation by modifying a single predicate in the proposed formulae. keywords: concession, contrast, discourse relations, penn discourse treebank 1. introduction our previous work on the semantic annotation of discourse relations in the penn discourse treebank (miltsakaki et al. (2008)) confirmed the well-known problem of achieving satisfactory interannotator agreement at the discourse level. as we move deeper into making semantic distinctions between eventualities participating in discourse relations, it becomes increasingly even more difficult to achieve reliability, even for well-studied relations such as concession. while the less fine ∗. financial support to eleni miltsakaki by the nsf iis-0803538 grant is gratefully acknowledged. the work of livio robaldo has been funded by the ateneo-san paolo project number to call03 2012 0046: “the role of visual imagery in lexical processing (rvilp)”. c©2014 livio robaldo and eleni miltsakaki submitted 01/14; accepted 12/14; published online 12/14 robaldo and miltsakaki semantic level described as comparison in the pdtb enjoyed high inter-annotator agreement, the distinction between its subtypes, i.e. contrast and concession, proved more challenging. a non-negligible 30% of tokens had to be adjudicated because in these instances one annotator picked concession and the other contrast. this was surprising given the reasonably clear semantic definition of the pdtb labels (cf prasad et al. (2008) or section 2 below). a careful analysis of the instances of disagreement as well as the examples studied so far in the literature revealed that a lot of linguistic variants could convey concession and contrast. of course, such a high variability makes the distinction between the two classes, or between their subtypes, rather unclear, especially when it is carried out by non-expert annotators. the research that we report in this paper is motivated by our belief that data driven representations can help us develop more precise formalization of semantic distinctions known to present challenges for annotators. precise semantic distinctions guided by naturally occurring data will help us to build semantic representations, the reliability of which can be empirically tested. we summarize our research goals below. (1) a. how can the study of discourse relations as attested in naturally occurring data improve our understanding of the semantics of discourse relations? b. what kind of semantic representation will allow covering the rich range of variants conveying concession and contrast? to address these questions, our research methodology is guided by a hybrid theoretical and empirical approach. we develop formal semantic representations of discourse relations based on an analysis of large scale empirical data. specifically, we analyze the semantic tagging and interannotator agreement of the discourse relations marked in the penn discourse treebank (prasad et al. (2008)). the penn discourse treebank 2.0 is, to date, the largest annotation effort at the discourse level, including approximately 40,000 annotations of discourse connectives and their arguments, sense labels, and speaker attribution. in the pdtb, sense labels are grouped in four basic types of semantic relations: a) temporal, b) contingency, c) comparison, and d) expansion. each category has types and subtypes. the full hierarchy of senses used in the pdtb is illustrated in prasad et al. (2008) and miltsakaki et al. (2008). as mentioned above, our focus in this paper is on the distinction between concession and contrast, the two subtypes of the comparison relation. as is natural when the body of the literature is large and coming from different disciplines, the interpretation of concessive relations has been addressed from several viewpoints. mann and thompson (1988)’s influential rhetorical structure theory (rst) views relations from a functional perspective. the proposed interpretation includes the speaker’s intention and the effect that the relation is intended to achieve on the hearer. grote et al. (1995) implement the theoretical insights from rst and other similarly minded proposals in the same vein into a real natural language generation (nlg) system designed to generate concessive sentences from formal representations of the speaker’s beliefs and communicative intentions. following moore and pollack (1992), we recognize the distinction between the intentional and informational levels of interpretation and find it problematic that the rst presumes a single relation between two discourse segments, thus conflating this distinction. in this spirit, our work extends prior work in sub-classifying discourse relations and developing formal representations of the identified classes. 2 corpus-driven semantics of concession our analysis revealed that concessive relations differ according to the source of expectation. specifically, we identified four distinct sources of expectation: causality, implication, correlation, and implicature. the reliability of the proposed categories was evaluated with a study of inter-annotator agreement. in addition to confirming the reliability of the proposed distinctions, we evaluated the merits of this proposal over existing denial-based approaches which treat all eventualities that trigger expectations uniformly. specifically, we extracted 200 problematic pdtb tokens that had previously been marked as tokens of inter-annotator disagreement between concession and contrast. these tokens were re-annotated by two new annotators. the high inter-annotator agreement in this challenging task provided further evidence for the validity of our proposal. finally, we developed a formal account of concession, grounded on our sub-classification, using basic semantic constructs from davidson (1967) and hobbs (1998). the resulting formulae are able to uniformly take into account the semantics of all variants of concessive statements identified in the literature. the paper is organized as follows. section 2 gives a brief overview of prior work on concession and contrast, while comparing the two classes. section 3 reviews logical accounts that have been proposed to model concessive interpretations, while section 4 highlights some important questions that have remained open. section 5 illustrates examples taken from pdtb that convinced us to further classify concession into four subtypes, depending on how expectations are created. on the other hand, section 6 reports the results of two empirical investigations carried out on pdtb instances that seems to support our analysis. in section 7, we present, briefly, the basic semantic constructs that we use from hobbs’s logic and outline in detail our semantic account for all but one source of expectation. the source of expectation we do not encompass in our approach is ‘implicature’. that requires pragmatic reasoning and so it is left for future work. we conclude in section 8. 2. concession and contrast concession is a particular relation holding between the interpretation of one clausal argument that creates an expectation and another clausal argument which denies it. in english, typical discourse connectives conveying concession are ‘but’, ‘although’, ‘however’, ‘yet’, and ‘nevertheless’. concessive discourse connectives are, of course, available in other languages (könig (1983)), which also have specialized words or even inflections to mark concessive relations, c.f. dascal and katriel (1977), horn (1989), and lagerwerf (1998). according to könig (1983), this diversity in the linguistic devices used to express concession suggests that the term ‘concessive’ does not only express a two-term relation, but also other possible rhetorical uses of the involved clauses. in the same spirit, grote et al. (1995) identify three rhetorical strategies a concessive construction may be built for: convincing the hearer, preventing false implicatures, and emphasizing surprising events. we investigate the interpretation of discourse connectives only, leaving outside other linguistic or non-linguistic cues that might be used to express concession. discourse connectives, in english and other languages, may use the same connective to express more than one type of relation. for instance, as observed in the pdtb, ‘but’ is used to express both contrast and concession. in line 3 robaldo and miltsakaki with prior work (lakoff (1971), spooren (1989), grote et al. (1995), and kehler (2002), among others), the pdtb adopts the following two definitions of contrast and concession1: (2) a. contrast applies when the connective indicates that the two sentence-arguments share a predicate or property and a difference is highlighted with respect to the values assigned to the shared property. examples are: i. [john paid $5 ] but [mary paid $10 ]. ii. [john likes music ] but [mary likes dancing]. iii. [i read two books ] while [you read only one ]. b. concession applies when the connective indicates that one of the arguments describes a situation a which “creates” an expectation c, while the other asserts (or implies) ¬c, i.e. it denies the expectation. examples are: i. although [john studied hard ], [he did not pass the exam ]. ii. [john married mary ] but [he loves another girl ]. iii. [we are going for a walk ] even if [it is raining ]. for now, let us indicate the relation between a and the expectation c as a kind of ‘default implication’, following winter and rimon (1994), and we will characterize this defeasible relation later in section 7.2. as pointed out above, contrast and concession are the two pdtb types of the higher level semantic category comparison. in the literature, prior work has described semantic classes that would fit under ‘comparison’ according to the pdtb sense tagset. for the connective “but”, specifically, there has been work analyzing its “corrective” or “rectification” use (dascal and katriel (1977), lang (1984), foolen (1991), von klopp (1994), winter and rimon (1994), and others). ‘rectification’ arises when the argument of the connective rewrites a predicate, e.g., “john is not american, but british”. in pdtb, similar cases are collapsed into the type contrast and are not our main focus2. in this paper, we are interested primarily in concession. the present section addresses the problem of distinguishing between concession and contrast. in the next section, we will address the central research question of the paper: identifying different subtypes of concession depending on how the expectation is created. later in the paper, we will provide some evidence that focusing on the source of expectation could also help distinguishing between concession and contrast. although the definitions in 2, as well as those used in most existing schemes of discourse relations, appear to be rather intuitive, it is possible that the distinction is sometimes hard because it is sensitive to the context. this has been investigated and argued in the work of lakoff (1971), anscombre and ducrot (1979), lang (1984), blakemore (1989), winter and rimon (1994), and spenader and lobanova (2009). an example, taken from winter and rimon (1994), is: (3) [john is quick ], but [bill is slow]. 1. all discourse connectives annotated in the pdtb have two arguments. in the examples shown in (2), argc is shown in boldface, argd in italics, and the discourse connective is underlined. 2. for a more detailed discussion of contrastive interpretations see izutsu (2008). 4 corpus-driven semantics of concession according to the instructions given in (2), an annotator should tag (3) as ‘contrast’. in fact, we may identify a common property ‘y is-a-quality-of x’, shared by the two arguments’ meaning representations, where x is either john or bill and y either one of the contrastive values ‘quick’ and ‘slow’. however, as discussed by winter and rimon (1994), (3) may be interpreted as concession in a context where john and bill belong to a sport team that is known to have only quick players. in this case, the first sentence in (3) is better interpreted as an assertion a which creates the expectation c that all players, including bill, are quick, while the second sentence in (3) explicitly asserts ¬c. as can be seen, forming a neat distinction between contrast and concession is even harder in real data. spenader and lobanova (2009) point out that the following example, marked as ‘concession’ in the rst corpus (carlson et al. (2001)), could be tagged as ‘contrast’ by an annotator who takes the brokerage operation and kidder as parallel elements and the profits/losses as the (symmetric) values with respect to which the parallel elements contrast. (4) [its 1,400-member brokerage operation reported an estimated $5 million loss last year ], although [kidder expects it to turn a profit this year]. according to winter and rimon (1994), the context is also responsible for the interpretation of certain concessive utterances, which, in different contexts, would be odd. for instance, (5.a) sounds odd if it is not interpreted in an appropriate context as in (5.b) (note that we mark the argument that creates the expectation and the one that denies it as argc and argd respectively). (5) a. #[i love venice ]argc , but [i would like to be there again ]argd . b. i always hated to visit again cities that i love. not in the case of venice: [i love it ]argc , but [i would like to be there again]argd . context and other pragmatic considerations are not only involved in the identification of the “default implication” that creates the expectation from the meaning of argc. once such implication is identified, the concessive interpretation may be triggered in different ways. sweetser (1990) identified three possible ways: content (or semantic) usage, epistemic usage, and speech act usage. these three classes have been recognized by many authors, among which lagerwerf (1998) and lang (2000). three examples of the classes, from lagerwerf (1998), are shown in (6). (6) a. [connors did not use kevlar sails ]argd, although [he expected little wind.]argc (content usage) b. [theo was not exhausted ]argd ,although [he was gasping for air. ]argc (epistemic usage) c. [mary loves you very much ]argd , although [you already know that. ]argc (speech act usage) in (6.a), argc creates the expectation via the pragmatic “default implication” if connor expects little wind, then he uses kevlar sails. based on the observation that theo is gasping for air, the speaker of (6.b) expects him to be exhausted. however, in (6.b), it cannot be deduced, or assumed, that if someone gasps for air, then he is exhausted. rather, the “default implication” has the opposite direction: if someone is exhausted, then he gasps for air. it is the fact of being exhausted that causes 5 robaldo and miltsakaki gasping for air, not the opposite. lagerwerf (1998) argues that, while content usage involves a “default implication” which is derived deductively, epistemic usage involves one that is derived abductively: from the observations to the causes3. finally, in (6.c) the expectation is denied by the illocution of argd, i.e., its speech act, rather than by its locutionary meaning. it is the fact that i tell you argd, and not argd itself, that is inconsistent with the expectation, created by the “default implication” if (i know that) you already know something, i do not tell it to you. note that it is possible to distinguish three further sub-cases of speech act usage: the concessive relation may involve the illocution of argd only, the one of argc only, or it may involve both. as argued in winter and rimon (1994), an occurrence of concession holding between the speech acts of both arguments is easily obtained by using imperative mood in both: (7) [take a chair ]argc, but [do not sit.]argd lagerwerf (1998) was the first who tried to collect all data and previous analyses, e.g., lakoff (1971) and sweetser (1990), and propose a classification of contrastive relations in natural language. he distinguished three main classes: semantic opposition, denial of expectation, and concessive opposition4. (8) a. semantic opposition, e.g., [greta was single] but [prince was married.] b. denial of expectation, e.g., examples (6.a-c). c. concessive opposition, e.g., shall we go to king tsin? [king tsin has great mu shu pork ]argc but [china first has good dim sum.]argd according to lagerwerf, all classes in (8) may involve speech acts in the arguments, not only denial of expectation (which we have seen in (6)). semantic opposition corresponds to the pdtb definition of contrast, e.g., (2.a) whereas both lagerwerf’s denial of expectation and lagerwerf’s concessive opposition are included in ptdb’s definition of concession. treating these as sub-types of concession is an approach that has also been adopted in several other works, winter and rimon (1994) and korbayova and webber (2007). we focus on this treatment in the next section. for the present discussion, the critical distinction is between lagerwerf’s semantic opposition and concessive opposition that we saw in the examples (8.a) and (8.c), respectively. with respect to (8.c), note how the preceding question “shall we go to king shin?” sets up a context that favors the concessive interpretation. this observation has been empirically verified by spooren (1989) who conducted an experiment with english speakers who showed a clear tendency to consider china first as the preferred restaurant. however, as discussed in lagerwerf (1998), an interpretation of semantic opposition is, also, possible for (8.c) if a different question is used to set up the context: 3. abduction is a form of non-monotonic defeasible reasoning. however, note that in case of concession the deductive counterpart is also defeasible, due to “default implication”. see also section 7.3. 4. in lagerwerf (1998), (8.c) is indeed termed as ‘concession’, not as ‘concessive opposition’. in pdtb 2.0 and other related work korbayova and webber (2007), the definition of ‘concession’ encompasses both lagerwerf’s ‘denial of expectation’ (6.a-c) and lagerwerf’s ‘concession’ (8.c). korbayova and webber (2007) rename lagerwerf’s ‘concession’ as ‘concessive opposition’ to avoid misunderstandings. in this paper, we follow korbayova and webber (2007)’s terminology. 6 corpus-driven semantics of concession (9) which restaurant is better? [king tsin has great mu shu pork ]argc but [china first has good dim sum.]argd indeed, lagerwerf (1998) argues that with the wh-question in (9), semantic opposition is the only available interpretation; two restaurants are compared with respect to their properties. from the discussion above it becomes clear that, in some cases, the crucial distinction between a concessive versus contrastive interpretation is entirely context dependent. different contexts, e.g., an introductory whversus yes/no question, enables different interpretations of the same statements, focusing on a symmetric versus asymmetric perspective with respect to the compared items. furthermore, both concessive and contrastive connectives compare two clausal arguments and highlight some kind of disagreement between the facts they each assert. contrast is a symmetric relation between two opposite facts. concession, on the other hand, features some kind of “directionality”: one of the two arguments asserts a fact that triggers an expectation and the other argument overrides it later. since context strongly influences the identification of the asymmetric roles of the arguments, in several cases the distinction could be rather subtle. although such instances are not pervasive, in the inter-annotator study that we report in section 6, we observed that in some cases, both a contrastive and a concessive interpretation could be built for the same token. 3. concessive interpretations in the previous section, we discussed prior work on concession with a focus on its difference from contrast. in this section, we review the literature on concession focusing on the various interpretations that have been proposed. lagerwerf (1998) identified two types of concession that he termed as denial of expectation and concessive opposition. a similar sub-categorization has also been outlined in prior work on the logical formalization of concessive relations, given in abraham (1991), winter and rimon (1994), korbayova and webber (2007). the corresponding formalizations model a direct and a less direct relation between the triggered expectation and the content of the textual span that denies it. korbayova and webber (2007) explain the distinction giving the two examples in (10). for simplicity, we will refer to them as instances of ‘direct concession’ (10.a) and ‘indirect concession’ (10.b). (10) a. although [greta garbo was considered the yardstick of beauty ]argc , [she never married ]argd . (denial of expectation ≡ direct concession) b. although [he does not have a car ]argc , [he has a bike ]argd . (concessive opposition ≡ indirect concession) in (10.a), a general “default implication” is presupposed, paraphrasable as beautiful women usually get married. because of this rule, argc directly triggers the expectation that greta garbo got married. this expectation is explicitly denied by argd. example (10.b) is different. in this case, “not having a car” does not imply “not having a bike”, i.e., no defeasible rule holds between the two arguments. according to lagerwerf (1998)’s terminology, in this case we can identify a ‘tertium comparationis’, i.e., a proposition entailed by argc with its negation entailed by argd. this proposition is presumably “he is mobile”. thus, two general rules can be identified from the arguments to the tertium comparationis: not having a 7 robaldo and miltsakaki car implies being less mobile and having a bike implies being mobile. as discussed in the previous section, the identification of the tertium comparationis strongly depends on context, e.g., it could be induced by a suitable introductory question. in the same spirit, sanders et al. (1992), sanders et al. (1993) argue that denial of expectation is ‘causal’, while concessive opposition is ‘additive’. the same conceptual difference has been sharpened in other works on discourse coherence, e.g., asher (1993) (‘structural relation’ versus ‘non-structural relation’), kehler (1994) (‘common topic’ versus ‘coherent situation’), and pander maat (1998) (‘causal relation’ versus ‘comparative relation’). how has this basic intuition on the relationship between expectation and denied expectation been formalized? francez (1995) proposes bilogic which uses two semantic structures, the standard and the actual world. the contrast between the two worlds gives rise to what we characterize here as concession and models the difference between concession and contradiction in terms of whether the statements are evaluated in different worlds. winter and rimon (1994) agree with francez (1995) on the basic intuition but propose to analyze concession as presupposition failure. they combine presupposition failure with the possibility and necessity operators of modal logic to define the semantics of what they classify as direct (‘although’, ‘even though’, ‘yet’, ‘nevertheless’) and indirect concessive connectives (‘but’)5. winter and rimon (1994)’s formal account of the two cases is shown in (11). in (11), p and q are the propositions denoted by argc and argd respectively, and � is the standard possibility operator defined in modal logic. in case of direct contrast (our direct concession), the expectation is identified by ¬q, while in case of indirect contrast (our indirect concession), a third proposition r is assumed to exist, which clearly refers to the tertium comparationis. its existence is implied by q while its negation is implied by p. (11) direct contrast: (p ∧ q) ∧ �(p→¬ q) indirect contrast: (p ∧ q) ∧ ∃r[�(p→¬ r) ∧ (q → r)] note that the ‘→’ is not the standard predicate logic operator of implication. while the details of its characterization are too complex to review here, note that the ‘→’ denotes an underspecified “cognitive reasoning” relation and it must not be confused with the standard first order logic entailment. in winter and rimon (1994), the standard first order logic entailment is formalized via the symbol ‘⇒’. several properties are asserted on ‘→’: it is reflexive, it holds that either ‘→’ is not transitive or that ‘⇒’ is not a special case of ‘→’, and, most importantly, ‘→’ is not defeasible. winter and rimon (1994) model defeasibility in terms of possible worlds, so that the consequent of ‘→’ holds when its antecedent holds, but not all cognitive implications that may be extracted from the sentence’s meaning are asserted in the current world. in order to assert that ‘p→¬ q’ (and ‘p→ ¬ r ’) only weakly hold, the implication is asserted as possible, via the modal operator ‘�’, while q (and ‘q→ r ’) are asserted as true. however, as winter and rimon acknowledge, some cases are problematic for such an account. for instance: (12) [john walks slowly ]argc , but [he walks. ]argd 5. for winter and rimon (1994), although, even though, yet, nevertheless have ‘restrictive’ meaning and only but is ‘non-restrictive’. this one-to-one correspondence between semantic descriptions and connectives breaks quickly when we look at empirical data. in pdtb, several connectives, including ‘but’ and other concessive connectives have more than one interpretation. the connective but, for example, has been annotated with seven sense tags. 8 corpus-driven semantics of concession suppose (12) is uttered in a context where john had a surgical operation. p may be taken as the eventuality “john walks slowly”, q as “john walks”, and ¬r as the expectation “the operation was not a success”. however, “john walks slowly” is clearly a particular case of “john walks” (p⇒q) and from this we infer that p→r holds in the current state of information. in other words, “john walks slowly” cognitively implies “the operation was a success”, which is clearly not the case. in order to handle such inconsistencies, winter and rimon (1994) have to propose further restrictions on ‘→’, in terms of possible worlds. in our account, we propose a formal solution using non-monotonic (default) reasoning, instead of possible world semantics. such a solution, which is also advocated by winter and rimon (1994) as a plausible alternative to their account, does not suffer from the problem exemplified in (12). winter and rimon’s ‘direct contrast’ appears to be a viable formalization of lagerwerf’s ‘denial of expectation’. however, among the three subcases of denial of expectation demonstrated above in (6), winter and rimon (1994) consider examples of content usage and speech act usage only, while remaining silent with respect to occurrences of epistemic usage. let us now turn our attention to lagerwerf and how he formalizes his intuitions about noncontent ‘denial of expectation’. lagerwerf (1998) proposes associating the sentences (6.b-c), copied here as (13.a-b) for convenience, with the formulae (14.a-b), respectively. (13) a. [theo was not exhausted ]argd , although [he was gasping for air. ]argc (epistemic usage) b. [mary loves you very much ]argd , although [you already know that. ]argc (speech act usage) (14) a. ∀x[gfb(x) > b(i, exh(x))] b. ∀x[k(i, k(y, x)) > � ¬t(i, y, x)] let us, first, consider the formula in (14.b). k(a, x) means that the agent a knows x while t(a1, a2, x) asserts that the agent a1 tells x to the agent a2. i and y are two constants referring to the speaker and the hearer respectively. thus, t(i, y, x) means that the speaker tells x to the hearer. ‘>’ is the defeasible implication operator defined in asher and morreau (1991). it corresponds to winter and rimon’s operator ‘→’ when it is asserted within the scope of the possibility operator ‘�’. formula (14.b) states that if someone knows something, i need not say it. this formalization directly mirrors the intuition about speech act usage and we agree with that. obviously, the parallel with winter and rimon’s is obtained by assuming p→¬q = ∀x[k(i, k(y, x)) > �¬t(i, y, x)]. on the other hand, in (14.a), gfb and exh are predicates denoting, respectively, the set of individuals gasping for air and the exhausted ones, while b(i,exh(x)) is an epistemic operator asserting that the speaker believes x to be exhausted. in the next section, we discuss the main challenges that these frameworks still need to address and we clarify our data-driven approach to meeting these challenges. 4. challenges in the previous section, we looked at basic issues in building the semantics of concession and some of the basic logical accounts that were proposed in the literature, among which winter and rimon (1994), lagerwerf (1998), and korbayova and webber (2007). 9 robaldo and miltsakaki in this section, we will look a little closer at the challenges that the concession data present to these accounts. we start with a summary of the key points of these accounts: (15) a. these approaches focus on the distinction between ‘direct concession’ and ‘indirect concession’. in ‘direct concession’, the expectation raised by one argument is explicitly denied by the other. in ‘indirect concession’, a ‘tertium comparationis’ must be first identified, i.e., an intermediate proposition entailed by one argument, whose negation is entailed by the other. b. the expectation or the tertium comparationis is triggered by some kind of default implication. a proper formalization of such a default implication has been mostly neglected in literature. in some approaches, e.g., sanders et al. (1993), it has been argued that it is a causal relation in case of ‘direct concession’ and a comparative one in case of ‘indirect concession’. c. as argued in lagerwerf (1998), a logical account of concession needs to be general enough to include several variants, featuring expectations that correspond to the speech acts of the arguments and/or an abductive, rather than deductive, use of the default implication. drawing from our observations in the instances of concession in the pdtb, it becomes clear that any theory that recognizes all and only two concessive interpretations will fall short when accounting for real data. on the contrary, we argue that the effort to successfully characterize how expectations are created, rather than how they are denied, is critical. in other words, we recognize an underlying general principle similar to the abc-scheme proposed in grote et al. (1995), which has been the starting point of korbayova and webber (2007), where we simply assert that argd is inconsistent with the raised expectation. contrary to korbayova and webber (2007), who further develop that principle by specifying subtypes of concession according to how the expectation is denied, we develop, instead, a deeper analysis on how the expectation is created because, in so doing, we are able to characterize more accurately the “default implication” mentioned in (15.b). of course, we are still able to distinguish between direct or indirect concession, but our approach does not advocate any direct correspondence of logic formulae to direct/indirect concession. it cannot be maintained that in all cases of direct concession the expectation is created by a causal rule. we have seen this even in simple examples such as (16.a-b). unless we assume an ad-hoc context, it would be odd to assert that being a penguin “causes” not flying and that the fact that john will do his report “causes” the fact that he will do it at home. (16) a. [penguins are birds ]argc . nevertheless [they do not fly. ]argd b. [john will do his report ]argc but [he will finish it at home. ]argd it would, also, be hard to try to identify a tertium comparationis in all cases of indirect concession. indeed, there are cases in which a tertium comparationis is not there to be identified. lagerwerf (1998) suggests that an easy way to identify a tertium comparationis is by presenting the utterance as an answer to an appropriate question. but what could be the appropriate questions for utterances (17.a-b)? 10 corpus-driven semantics of concession (17) a. although [john ate a lot of pizza ]argc , [he did not eat it all. ]argd b. [open the computer case ]argc , but [do not touch the wires. ]argd it is rather hard to interpret (17.a) as an answer to a particular question. perhaps we may think about a context in which someone asks is there some pizza left?, and the speaker replies with (17.a). in such a case, a tertium comparationis analysis seems possible for (17.a) : argc is interpreted by the hearer as a negative comment to the prospect of eating, and argd as a (stronger) positive one. in the case of (17.b), it is even harder to find a context which involves a tertium comparationis because the example may be uttered in any context in which someone gives instructions on how to open the computer case. to account for all the observed instances of concession, including (17.a-b), we need a more general definition of indirect concession. in our account, all cases in which argc is insufficient or irrelevant with respect to the satisfaction of speaker’s intentions are classified under the term ‘concession-implicature’. in (17.a), argc is irrelevant with respect to the satisfaction of speaker’s intentions, i.e., communicating to the hearer that there is some pizza left, and could lead the latter to conclude that there is nothing for him to eat. analogously, in (17.b), the command in argc is insufficient, and could lead the hearer to take wrong actions. the speaker adds then further specifications by uttering argd. on the other hand, in (8.c), argc is interpreted by the hearer as a preference for king tsin over china first, which does not meet the speaker’s intentions. however, the fact that the latter has another preference, and so a tertium comparationis may be identified, in our view is simply a special instance. if the sentence was modified to “king tsin has great mu shu pork, but i do not want to talk about that”, the tertium comparationis would be less easy to identify. finally, we agree with lagerwerf (1998) that a proper logical account of concession must take into account speech acts and abductive use of the relation triggering the expectations, but we find his formalization of the epistemic usage problematic for two reasons. consider again the example in (13), repeated in (18) for convenience. (18) a. [theo was not exhausted ]argd ,although [he was gasping for air. ]argc b. ∀x[gfb(x) > b(i, exh(x))] in (18.b), gfb and exh are predicates denoting, respectively, the set of individuals gasping for air and the exhausted ones, while b(i,exh(x)) is an epistemic operator asserting that the speaker believes x to be exhausted. however, it is somehow odd to assert that the speaker believes someone to be exhausted given that he is gasping for air. the defeasible rule in (18.a) is general and therefore does not apply specifically to any particular speaker. therefore, in the formalization, i should be most properly substituted by universal quantification over all possible believers. this is in line with spooren (1989), sanders (1994) and pander maat (1998), who identify different subjectivities that may be ascribed to the statements in a discourse. in particular, pander maat (1998) conducts a corpus analysis showing that three perspectives of subjectivity must be distinguished: ‘objective perspective’ (the statements are objective facts that are taken to be acceptable by the speaker), ‘speaker perspective’ (the belief of the statement is ascribed to the speaker), and ‘other perspective’ (the belief of the statement is ascribed to persons other than the speaker). pander maat (1998) proposes a revision 11 robaldo and miltsakaki of the hierarchy of discourse relations provided by sanders et al. (1992) that includes a new feature specifying the perspective configuration. secondly, the formula in (18.b) does not adhere to lagerwerf’s intuition (lagerwerf (1998), pp.41-42) that epistemic usage of concession involves a defeasible rule which applies abductively, i.e., getting from observations to causes. in fact, the rule does not appear in his formulae, e.g., (18.b). furthermore, the “default implication” is also defeasible in case when it is used deductively, as in content usage. therefore, if it should be asserted that the speaker believes the pre-conditions when he observes the effects, it should also be asserted that he believes the effects when he observes the pre-conditions. in our view, lagerwerf (1998)’s valid intuition must be formalized exactly as it is stated: the formula must include an explicit defeasible rule corresponding to “being exhausted causes gasping for air”. separately, the formula asserts that the rule yields the expectation abductively. in other words, in the formulae the assertion of the defeasible rule must remain orthogonal to its usage. to sum up, in this and the previous sections we have attempted to give a comprehensive review of the literature on concession and the challenges that any logical account of concession will need to address. one could view these challenges as a purely theoretical exercise in semantic theory and continue to work on them on a theoretical basis. one might argue that these theoretical challenges should be mostly irrelevant to human annotators whose task is to identify and annotate concession in naturally occurring data. while, indeed, it may not be surprising that there are theoretical challenges to be addressed in the formal treatment of concession, we were intrigued by the fact that what, for a human, seemed to be a fairly straightforward definition of concession (denial of expectation) yielded surprisingly high disagreement among annotators. indeed, almost 30% of pdtb tokens annotated as concession by one annotator were annotated as contrast by the other. close investigation of the data helped us realize that viewing the relation with a focus on the denial of expectation made it hard for the annotators to discern the triggers of expectation. analyzing the sources of expectation, the different types of relation that trigger them and the inferences that they allow helped them identify the relations with much improved consistency. our work bridges the gap between corpus data and logic and our methodological approach is doing so by starting at the bottom. we started by looking at discourse connectives in the pdtb, and then built up more abstract models for deriving appropriate inferences. in the next section, we will look at concession in the pdtb and present a data-driven analysis of the different inferences that are triggered from the range of the sources of expectations attested in the annotations of concession. 5. concession in the pdtb: where do expectations come from? the pdtb corpus contains 1193 annotated instances of concession associated with an explicit discourse connective. in order to identify the possible defeasible relations involved in concessive relations in real data, 1000 of these 1193 instances have been analyzed. table 1 shows the distribution of all the tokens in the pdtb that were labelled6 as concession (or any of its subtypes7) and contrast (or any of its subtypes). there were a total of 1193 tokens 6. a full description of the sense tags used in the pdtb is given in prasad et al. (2008) and miltsakaki et al. (2008). 7. the pdtb distinguishes two subtypes of concession: “expectation” and “contra-expectation”. when the clausal argument that syntactically bounds the discourse connective creates the expectation, the pdtb instance is labelled as “expectation”. otherwise, it is labelled as “contra-expectation”. note that this sub-categorization is orthogonal to 12 corpus-driven semantics of concession labelled as concession. the most common concessive connective is ‘but’ with 508 tokens (42% of all concessive labels), followed by ‘although’ with 154 tokens (13% of all concessive labels). the connective ‘but’ is, also, very common in relations marked as ‘contrast’, which may have contributed to the confusion between ‘concession’ and ‘contrast’ that we noted earlier. connective concession although 154 (13%) but 508 (42.5%) even if 35 (3%) even though 72 (6%) however 77 (6.5%) nevertheless 19 (1.5%) nonetheless 17 (1.5%) still 82 (7%) though 84 (7%) while 83 (7%) yet 32 (2.5%) other 30 (2.5%) total 1193 connective contrast although 157 (4.06%) but 2422 (62.8%) by contrast 27 (0.7%) even though 21 (0.51%) however 355 (9.3%) meanwhile 37 (0.95%) on the other hand 35 (0.9%) still 96 (2.45%) though 131 (3.42%) while 427 (11.16%) yet 53 (1.32%) other 76 (2.43%) total 3856 table 1: concession and contrast labels in pdtb 2.0 it seems surprising that concession, despite having a straightforward definition, was so frequently confused with contrast. for all the tokens that were annotated as either concession or contrast by one of the pdtb annotators, there was almost 30% disagreement. in most cases of disagreement, at least one annotator would choose contrast over concession because they would prefer to construct a contrastive interpretation between the created expectation and its denial, failing to see that the involved predicates were not symmetric. this section reports several pdtb instances tagged as concession, out of 1000 selected ones, and shows that it is possible to identify four types of semantic relation that give rise to the asymmetry characterizing concession: causality, non-monotonic implication, correlation, and implicature. the next four subsections discuss each category with corpus examples and outline their meanings. support for the proposed classification of the identified sources of expectations is given by the results of an inter-annotator study that we conducted asking the annotators to label the data with the new categories (cf. section 6). 5.1 causality in sweetser (1990), sanders et al. (1992), lagerwerf (1998), and also in prasad et al. (2008), it is assumed that in all cases of concession the expectation comes from a (defeasible) causal rule. the next subsections argue that this is true in most, but not all, cases. an example, taken from the pdtb, in which the expectation is created by a defeasible causal rule is shown in (19). the one addressed in this paper, i.e. “direct” and “indirect” concession. the latter concerns the way the expectation is denied, regardless of the clausal argument that creates it. 13 robaldo and miltsakaki (19) although [they represent only 2% of the population]argc , [they control nearly one-third of discretionary income ]argd . in (19), argc asserts that “they” represent a very low percentage of population. that creates the expectation that they control a (proportionally) low percentage of income. where does this expectation come from? the obvious answer is that our world knowledge includes a general causal rule representing that a low percentage of population causes control of a small amount of income, which instantiates on argc and creates the expectation. this causal rule is, however, defeasible, i.e., its effect may be falsified or canceled, as it is done by argd in (19). more cases of concession triggered by a causal rule are shown in (20). (20) a. although [imports account for less than 1% of beer sales in japan]argc , [asahi breweries ltd., which has been gaining share with its popular dry beer, plans to fend off japanese competitors by pouring $1.06 billion into facilities to brew 50% more beer]argd . (causality) b. [this meeting “put in motion” procedural steps that would speed up both of these functions]argc . but [ no specific decisions were taken on either matter]argd . (causality) c. [...that hung over parts of the factory]argd even though [exhaust fans ventilated the area]argc . (causality) d. [a sanwa bank spokesman denied that the finance ministry played any part in the bank’s decision]argc . still [mr. utsumi may have a hard time convincing market analysts who have rightly or wrongly believed that the ministry played a role in orchestrating recent moves by japanese banks]argd . (causality) e. [an undistinguished college student, who dabbled in zoology until he concluded that he couldn’t stand cutting up frogs, mr. corry wanted to work for a big company “that could do big things”]argc but [after joining the tax department of a usx subsidiary 30 years ago, he set the modest goal of becoming tax manager by the age of 46.]argd (causality) in (20.a), the low percentage of sales in japan should cause asahi breweries ltd. to invest somewhere else. similarly, in (20.b), “the procedural steps triggered by the meeting” (defeasibly) causes “taking important decisions in both of these functions” and, in (20.c), the fans should blow away whatever it was that hung over it. in (20.d), the declarations of the sanwa bank spokesman create the expectation that mr. utsumi may be on the safe side. finally, (20.e) is particularly interesting in that the expectation is created by a conjunction of two different causes. the fact that mr. corry was an undistinguished college student and the fact that he had the intention to work for a big company defeasibly cause the fact that he got smart professional results. 5.2 implication the previous subsection presented some examples of concession where the expectations are created via abstract (defeasible) causal relations that instantiate on argc. and, it has been pointed out that many past proposals assume that the expectation is always triggered by a causal rule. the data in the pdtb reveal that not all occurrences of concession involve causality. an interesting instance is shown in (21): (21) [the prime minister,]argd [whose hair is thinning and gray and whose face has a perpetual pallor,]argc nonetheless [continues to display an energy, a precision of thought and a willingness to say publicly what most other asian leaders dare say only privately]argd . 14 corpus-driven semantics of concession argc describes two properties featured by the prime minister, which do not appear to cause the negation of argd, or a tertium comparationis related to it. rather, the description recalls in our minds some kind of prototypical old and tired man, of which the prime minister would be an instantiation. the expectation stems from the fact that the prime minister inherits all typical properties of such a prototype, among which the one of having a lazy and indolent attitude. default inheritance from a prototype is clearly a defeasible implication, i.e., its consequent may be overridden as it is done by argd in (21). these considerations are well-known by researchers working on default logics. consider the typical example shown in (22). (22) [penguins are birds.]argc . nevertheless [they do not fly]argd . in (22), argc suggests that penguins have the property of flying, which they inherit from the prototype of ‘bird’. this expectation is explicitly denied by argd. in many pdtb occurrences argc evokes a kind of prototype of which some properties are overridden by argd. some are reported in (23): (23) a. [so far, all the studies have concluded that ru-486 is safe.]argc . but [“safe” in the definition of marie bass of the reproductive health technologies project, means “there’s been no evidence so far of mortality”]argd . (implication) b. [david is a pragmatist.]argc . but [mr. dinkins’s sense of pragmatism often comes across more as an insider’s determination not to upset the political apple cart]argd . (implication) c. although [working for u.s. intelligence]argc , [mr. noriega was hardly helping the u.s. exclusively ]argd . (implication) d. although [insider trading has long been criminal]argc , [it has never been statutorily defined ]argd . (implication) e. [you can do all this ]argd even if [ you’re not a reporter or a researcher or a scholar or a member of congress]argc . (implication) in (23.a-b) it is easy to see the inheritance by default that creates the expectation. the concepts of “safe” and “pragmatic” respectively used in the sentences are not exactly the ones that are standarly assumed, i.e., the prototypes. argd specifies the prominent differences with respect to such a prototype, i.e., what properties are overridden. in many cases, the prototype from which the canceled expectations are inherited is not so easy to identify. in those cases rather than thinking in terms of “inherited properties”, it is more convenient to think in terms of “necessary conditions” to which the prototype must adhere. for instance, in (23.c), working for u.s. exclusively is perceived as a necessary condition for working for u.s. intelligence. in other words, by reading argd in (23.c) we perceive that mr. noriega is arguably breaking some kind of rule required by his role. similarly, in (23.d), it seems that, in order to claim that “something is criminal”, it is necessary that “it is defined as such by the law”. finally, in the context of (23.d) it defeasibly holds that whoever can do all this must be either a reporter, or a scholar, or a researcher, etc. 15 robaldo and miltsakaki 5.3 correlation the annotators of the empirical study presented below in section 6 chose ‘causality’ or ‘defeasible entailment’ for about 70% of the occurrences of concession taken from the pdtb. the remaining cases seem to involve different relations. consider for instance (24): (24) [the treasury will raise 10 billion in fresh cash by selling 30 billion of securities . . . ]argc . but [rather than sell new 30-year bonds, the treasury will issue 10 billion of 29 year, nine-month bonds]argd . in (24), it does not seem that there is a general causal rule at stake. the fact that the treasury will raise money cannot be the cause of the way it will actually do it. arguing for the existence of a prototype evoked by argc also seems hard, though more compatible than the causal interpretation (cf. next subsection). it seems that in examples such as (24) the expectation is created on the basis of the history of the previous similar situations. in the context, it is assumed that there are two events that usually correlate. argc describes one of the two, and we expect the other one to co-occur based on the fact that in several similar previous situations they did so. accordingly, the third source of expectation has been termed as ‘correlation’. archetypal cases of correlation are all examples where argc describes a kind of trend and argd an eventuality that diverges from that trend. (25) shows some examples taken from the pdtb. (25) a. although [the notes held at a price of 92 to 93 immediately after the reset,]argc , [they started falling soon afterward ]argd . (correlation) b. [sales of the heart drug tpa were $43.6 million, better than last year’s depressed third period when the company sold just $29.1 million of the drug ]argc . but [tpa sales fell below levels for this year’s first and second quarter sales of $48 million ]argd . (correlation) c. [the ldp won by a landslide in the last election, in july 1986 ]argc . but [less than two years later, the ldp started to crumble, and dissent rose to unprecedented heights ]argd . (correlation) a variant of this pattern encompasses occurrences where argd describes an eventuality that sounds “surprising” together with the one described by argc (cf. könig (1983)). (26) shows some of such instances. in (26.a), it is “surprising” that wedtech got rolling so late, given its start date. similarly, in (26.b) and (26.c), it is surprising that mr. collor remains ‘the favorite’ and that the journal did not mention the reserve fund and the creators of the money-fund concept. (26) a. although [started in 1965 ]argc , [wedtech didn’t really get rolling until 1975 ]argd . (correlation) b. [the favorite remains fernando collor de mello, a 40-year-old former governor of the state of alagoas ]argc . but [after building up a commanding lead, the moderate to conservative mr. collor has slipped to about 30% in the polls from a high of about 43% only a few weeks ago ]argd . (correlation) c. [actually, about two years ago, the journal listed the creation of the money fund as one of the 10 most significant events in the world of finance in the 20th century ]argc . but [the reserve fund, america’s first money fund, was not named, nor were the creators of the money-fund concept, harry brown and myself ]argd . (correlation) 16 corpus-driven semantics of concession is correlation a source of expectation that is really distinct from the others? as argued above, concession may stem from causality or implication if the context includes a general causal or entailment rule that creates the expectation. it may then be observed that, in those cases, the event that triggers the expectation and the event that describes the expectation co-occur. let us look at the following simpler example of correlation: (27) [john will finish his report ]argc , but [he’ll do it at home ]argd . from (27), we infer that john usually does not finish his reports at home, and the present occasion constitutes an exception to this general trend. but it may be argued that there is a particular (unknown) reason why john never does his reports at home. maybe his home is too noisy or the reports must be returned by the end of the work day. these reasons might cause the fact that john does not finish his reports at home. similarly, in (24) we may think of a “prototypical treasury” that always raises money in the same way, namely by selling new 30-year bonds. although such considerations might indicate that causality and non-monotonic implication often entail correlation, in our view they should be kept distinct for two reasons. first, precisely because we do not know if there is a particular hidden reason why john does his reports at the office, we should not assert its existence, unless we believe that this is the inference that the reader draws from the text, which is clearly not the case. secondly, it has been attested beyond doubt that there are instances of concession for which no causal rule or defeasible entailment can be construed. there are, also, examples involving a causal rule, for which it cannot be asserted that the cause co-occurs with the effect. (28.a-b) from winter and rimon (1994) and grice (1961) are cases in point: (28) a. [take a chair ]argc, but [do not sit.]argd (correlation) b. [she is poor ]argc but [she is honest. ]argd (causality) in (28.a), we cannot infer a causal rule or prototype stating that encouraging someone to take a chair “causes” or “entails” an invitation to sit on it. maybe correlation could best model instances of concession involved in speech act usage but we have not conducted a study for speech acts specifically to support any claims. conversely, in (28.b) the expectation is created via a causal rule: poverty may be the cause driving people to criminal activity such as stealing. but, it would be wrong to infer from that causal relation that poor people tend to be dishonest, i.e., a correlation relation. 5.4 implicature there are cases of concession in which the expectation is created by the pragmatics of the conversation. as mentioned earlier, while all occurrences belonging to this class express indirect contrast, not all of them involve a tertium comparationis. for this reason, we associate this class with a broader definition. concession is triggered via implicature whenever argc is insufficient or irrelevant to the speaker’s intention. it could lead the hearer to draw unintended inferences. it seems that in such cases argc violates a gricean maxim, grice (1975). argd adds to argc the relevant information that the speaker wants to convey. the examples discussed in the introduction are repeated in (29). 17 robaldo and miltsakaki (29) a. a: shall we take this room? b: [it has a beautiful view ]argc but [it is very expensive. ]argd b. although [john ate a lot of pizza ]argc , [he did not eat it all. ]argd c. [open the computer case ]argc , but [do not touch the wires. ]argd in (29.a), argc could be interpreted by the hearer as “ i (the speaker) want to rent this room”, which is not the speaker’s intention. in (29.b), argc is irrelevant with respect to the satisfaction of speaker’s intentions, i.e., communicating to the hearer that there is some pizza left, who in the context might be looking for something to eat. similarly, in (17.c), the command in argc could lead the hearer into thinking that the permission to which opening the computer case extends is unconstrained. below are some examples of concession via implicature taken from the pdtb: (30) a. although [it is not the first company to produce the thinner drives ]argc , [it is the first with an 80-megabyte drive ]argd . b. [also, exxon went down 3/8 to 45 3/4 and allied-signal lost 7/8 to 35 1/8 ]argc , even though [the companies’ results for the quarter were in line with forecasts ]argd . in (30.a), argc does not create any expectation that is inconsistent with argd. argd, simply, it conveys an achievement that is worth noticing in this context. similarly, in (30.c) argc reports some data about the stock value trend of exxon and alliedsignal. argd simply stops the potential inference that their results, which are indeed independent from the stock value, were not in line with the forecasts. the pdtb does not include enough implicature examples for analysis, so clearly more work is needed before a satisfactory treatment of this category can be offered. 6. studies of inter-annotator agreement in this section we report two inter-annotator agreement studies that we conducted to evaluate a) the reliability of distinguishing four sources of expectations in the semantic description of concession and b) the impact of the new analysis of concession on the, previously low, inter-annotator agreement between contrast and concession in the pdtb. it must be pointed out that these annotation experiments ought to be considered only as preliminary studies of our claims, i.e., concession is more characterized by how expectations are created rather than by how they are denied. on the other hand, in order to obtain reliable annotations we will need precise guidelines with linguistic examples and subsequent adjudication steps as suggested by versley and gastel (2013). since annotating discourse relations is a rather difficult task, versley and gastel (2013) propose a set of linguistic tests that annotators should use in order to tag difficult non-archetypal cases. for such cases, annotators are required to perform paraphrases of the utterance, insertion/substitution operations of either the connective or its argument, etc. and check which aspects of the overall meaning are changed and which are not. the check should make annotators able to select the proper sense label. with respect to the ambiguity between contrast/concession, versley and gastel (2013) propose linguistic tests aiming at testing the symmetric/asymmetric role of the arguments (cf. versley and gastel (2013), section 4.1). 18 corpus-driven semantics of concession furthermore, since the quality of the annotations obviously does not only depend on the clarity of the guidelines, but also on how the annotators are able to apply them, versley and gastel (2013) suggest using a set of quantitative tests to subsequently inter-adjudicate the annotations. this is particularly strategic for discourse relations, for which all annotation schemes proposed so far in the literature appear to be intuitive with respect to sample cases, but it is not so when applied to real data, due to the strong context-sensitivity of discourse connectives (cf. (2) above). in the same spirit, spenader and lobanova (2009) uses χ2 to check statistically significant correlations between lexical markers and their senses. this and similar methods could be used for “filtering” discourse markers that are intuitively associated with certain senses but that, empirically, are not. for instance, spenader and lobanova (2009) found out that “however”, standardly taken to be a marker of contrast, is indeed equally used in cause-effect relations. nevertheless, the creation of such a reliable corpus is beyond the goal of the present paper, and it will deserve a new separate paper. the key point of our paper, we stress again, is to provide a logical formalization of concessive relations alternative to the ones proposed by winter and rimon (1994), lagerwerf (1998), korbayova and webber (2007), and others. these proposals are essentially grounded on the analysis of sample sentences while our formalization is mainly guided by an empirical analysis of real data stored in the pdtb. 6.1 annotation of expectation sources we conducted an empirical analysis on 1000 pdtb tokens of explicit connectives annotated as ‘concession’. two trained annotators, one of the authors and a post-doctoral researcher in linguistics, tagged each token with one of the four sources of concession identified above. the postdoctoral researcher received a short tutorial about the different sources of expectation as explained and had the option to use ‘other’ if none of the suggested labels were appropriate. the option ‘other’ was not used by either annotator. the most common source of expectation comes from causal relations (41.6%), followed by implication (28.7%), correlation (19.4%) and implicature (10.3%). source although but total causality 65 248 416 (41,6%) implication 45 125 287 (28,7%) correlation 31 87 194 (19,4%) implicature 13 48 103 (10,3%) table 2: distribution of the four sources of concession. the kappa statistic for inter-annotator agreement yielded 0.8 agreement, indicating that the defined categories are reliable8. in the formula below, pr(a) is the percentage of agreement (85% of 8. since the mid-1990s, when we saw an increased interest in producing semantic and discourse level annotations to linguistic corpora, it has been widely recognized that the highly subjective nature of semantic and pragmatic interpretations could yield unreliable annotations. when two annotators disagree, either one of the two annotators is wrong or the annotation schema, often a set of tag categories, is not capturing a reliable characterization. semantic and discourse annotation efforts are renowned for the struggle to identify reliable categories that would minimize interannotator disagreement (among others, carletta (1996), di eugenio (2000), and poesio and artstein (2008)). for example, in the development of the rst corpus, carlson et al. (2001) used professional language analysts with prior 19 robaldo and miltsakaki the 1000 cases considered) while the percentage of each tag, i.e., pr(e), is equal to 25%, as there are four possible sources of concession. κ = pr(a)−pr(e) 1−pr(e) = 0.85−0.25 1−0.25 = 0.8 after the computation of inter-annotator agreement, there was a brief adjudication effort that resulted in resolving any disagreements so we could compute the distribution of labels. in the cases of disagreement, we did not observe any interesting pattern to report. table 2 shows the distribution of the four labels for the most common connectives conveying concession, i.e., ‘but’ and ‘although’. 6.2 annotation of concession vs contrast making a reliable distinction between contrast and concession was the most challenging annotation task in the pdtb, exhibiting relatively low inter-annotator agreement. in a total of 4319 instances of explicit connectives that were annotated as either concession or contrast, there was agreement in 3057 cases, i.e., 70.8%. for that reason, we conducted a second inter-annotator agreement study focusing only on the pdtb tokens that were annotated as either concession or contrast. specifically, we extracted 200 tokens of disagreement, i.e., tokens that one annotator had labelled as concession and the other as contrast. we trained two annotators, not the authors, to perform the task. both annotators were linguistics students who attended a two-hour seminar on the distinctions between the different sources of expectation. the definition of contrast remained the same as in the original pdtb annotation. for each token, they were instructed to choose one of four annotation labels: a) contrast, b) concession, c) comparative, and d) other. they were allowed to use the label comparative when they could not decide between contrast and concession and other when they thought that the example belonged to a different semantic class. the results are reported in table 3. the distributions of the two annotators are almost equal. but, of course, this does not mean that we obtained almost 100% agreement. indeed, there are only six instances that have been labelled as ‘contrast’ by both annotators. for all other instances, either both annotators chose ‘concession’ or they assigned different labels. each annotator used the label ‘other’ a single time, but not for the same instance. the label ‘comparative’ has never been selected. therefore, annotators agreed on 161 tokens (80.5%), most of which (155 tokens) have been labelled by both annotators as ‘concession’. the kappa score is: k = pr(a)−pr(e) 1−pr(e) = 0.805−0.25 1−0.25 = 0.74 however, this kappa cannot be compared with the statistics of the original pdtb annotators, as they had more labels to choose from when they performed the annotations. but this is not critical for our experience in data annotation and they only achieved kappa 0.60 for annotating rst-style discourse relations (including concession) reaching maximum kappa 0.75 after the annotators had worked together for a week. explaining to annotators what to do does not guarantee agreement even if they are trained. inter-annotator studies are, therefore, crucial for the evaluation of the reliability of the suggested semantic categories. 20 corpus-driven semantics of concession label student1 student2 contrast 25 24 concession 174 175 comparative 0 0 other 1 1 table 3: distribution of concession vs contrast. purposes because we are only interested in evaluating the possible gain of analyzing concession in terms of the four sources of expectation on the annotation of contrast versus concession. the strong preference towards concession was indeed expected. we recall that the 200 instances were selected among those that were ambiguous between concession and contrast in the original pdtb annotation. intuitively, it is somehow unlikely that such doubtful cases were conveying a symmetrical relation, which should be rather easy to identify. in other words, it is possible that the pdtb annotators could not reliably identify the underlying relations that gave rise to expectations. focusing on the types of relations that give rise to expectations made it clearer that unlike concession, an asymmetrical relation, contrast involves a symmetrical relation between a common shared predicate receiving different values. concession, on the other hand, is always an asymmetrical relation, which relies on understanding the underlying relation of two events, not mentioned explicitly. understanding the nature of the relation that gives rise to an expectation (causality, implication, correlation, implicature) highlights the asymmetry inherent in concession. therefore, in the same spirit as versley and gastel (2013), section 4.1, who propose linguistic tests aiming at testing the symmetric/asymmetric role of the arguments for disambiguating between contrast and concession, focusing on the sources of expectation could be perhaps taken as a semantic/pragmatic test for the very same task. looking at the instances of persisting disagreement, we observed that in most of these cases, it was possible to construct both a contrastive and a concessive interpretation. consider for example, token (31). in this case, both a concessive and a contrastive interpretation can be built. a contrastive interpretation can be built by juxtaposing the predicates aware but not responsible. a concessive interpretation can be built if the reader assumes that the president waldheim knew about the killings before they happened and did nothing to prevent them. in this context, asserting that he was not responsible for the killings creates the expectation that he was not aware of them. (31) [london has concluded that austrian president waldheim wasn’t responsible for the execution of six british commandos in world warii ]argc , although [he probably was aware of the slayings]argd . it is possible that in some cases, better understanding of the context might help in disambiguating the intention of the author. on the other hand, it is also possible that both interpretations are entertained by the reader. since we did not give the annotator the option to annotate with a double tag concession-contrast, we do not know if such a tag would be used. while further studies would be required to evaluate the impact of the proposed analysis on a bigger scale, these results offer strong support in favor of looking closer at sources of expectation when analyzing concession. with these encouraging results, we set out to develop a logical account that 21 robaldo and miltsakaki would most elegantly capture the semantics of concession while making appropriate distinctions for the identified (semantic) sources of expectation. 7. semantics of concession this section proposes a logical account for the occurrences of concession in which the source of the expectation is either causality, non-monotonic implication, or correlation. a proper formal treatment of concession via implicature is seen as the object of future work. rather than designing new ad-hoc logical constructs to handle the semantics of concession, we make the effort to formalize our insights using an existing logical framework, if possible. the framework that allowed us to give the most elegant account has been defined in hobbs (1998) and several other earlier publications by the same author9. hobbs defines a wide-coverage logic for natural language semantics based on the notion of reification davidson (1967), bach (1981). it implements a fairly large set of linguistic and semantic concepts including sets, composite entities, scales, change, causality, time, event structure, etc., into an integrated first order logical formalism. hobbs’ framework includes all ingredients needed to properly represent the concepts introduced in the previous sections, in particular the possibility of defining defeasible relations. in addition, hobbs’ modular logic can be used to study the semantics of the connectives independently of the semantics of the arguments argc and argd. finally, we also show that hobbs’ framework is a suitable choice for the easy integration of other insights offered in the literature, such as lagerwerf (1998)’s, and the extension to the semantics of other discourse connectives. interestingly, lagerwerf (1998)’s work as well as several other researchers’ work on discourse semantics, is based on the taxonomy of coherence relations proposed by sanders et al. (1992), which is in turn based on hobbs’ notion of “discourse coherence” hobbs (1991). the following subsection briefly describes hobbs’ logical framework, with a particular focus on the ingredients needed to handle concessive relations. our proposal for a logical account of concession will be illustrated in 7.2. 7.1 hobbs’ logical framework hobbs (1998) proposed a wide coverage logical framework for nl semantics centered on the notion of reification. reification allows a wide variety of complex natural language (nl) statements to be expressed in predicate logic. nl statements are formalized such that events, states, etc., correspond to constants or quantifiable variables of the logic. in other words, the states and events denoted by these constants as well as the variables are things in the world. hobbs uses the term ‘eventuality’ to denote the reification of both a state or an event. hobbs distinguishes two parallel sets of predicates: primed and unprimed. the unprimed predicates are standard first order predicates commonly used in logical representations. for example, (give a b c) asserts that a gives b to c in the real world. the primed predicate represents the reification of the corresponding unprimed relation. the expression (give′ e a b c) says that e is a giving event by a of b to c. eventualities may be possible or actual. in hobbs, this distinction is represented via a unary predicate rexist that holds for eventualities really existing in the world. to give 9. see http://www.isi.edu/∼hobbs/csknowledge-references/csknowledge-references.html and http://www.isi.edu/∼hobbs/csk.html. 22 corpus-driven semantics of concession an example cited in hobbs, if i want to fly, my wanting really exists, but my flying does not. this is represented as: (rexist e) ∧ (want′ e i e1) ∧ (fly′ e1 i) eventualities can be treated as the objects of human thoughts. reified eventualities are inserted as parameters of such predicates as believe, think, want, etc. reification can be applied recursively. the fact that john believes that jack wants to eat an ice cream is represented as an eventuality e such that it holds10: (rexist e) ∧ (believe′ e john e1) ∧ (want′ e1 jack e2) ∧ (eat′ e2 jack ic) ∧ (icecream′ e3 ic) hobbs’ logic distinguishes between specific eventualities, like “fido is barking”, and general or abstract types of eventualities, like “dogs bark”. they are not treated as radically different kinds of entities. at some level, they are both eventualities that can be the content of thoughts. to this end, the logical framework includes the notion of typical element (from hobbs (1995) and hobbs (1998)). the typical element of a set is the reification of the universally quantified variable ranging over the elements of the set (cf. mccarthy (2002)). typical elements are first-order individuals. their introduction is motivated by the need of moving from the standard set theoretic notation in predicate logic: (forall (x) (iff (member x s) (p x))) to a simple statement that p is true of a “typical element” of s. in hobbs’ notation, the typical element t of a set s satisfies the predicate (typelt t s) . the principal property of typical elements is that all properties asserted on them are inherited by the members of their corresponding sets. it is important not to confuse the concept of a typical element with the standard concept of “prototype”, which allows for defeasibility, i.e., properties that are not inherited by all of the real members of the set. asserting a predicate on a typical element of a set is logically equivalent to the multiple assertions of that predicate on all elements of the set. these considerations lead to the distinction between eventuality types and eventuality tokens. the logic defines the following concepts, for which we omit formal details: a. eventuality types (also known as abstract eventualities): eventualities that involve at least one typical element among their arguments or arguments of their arguments. b. partially instantiated eventuality types (aka partial instances): a particular kind of eventuality type resulting from substituting the typical elements of some of its (sub-)arguments with other typical elements corresponding to proper subsets. c. eventuality tokens (also known as instances): a particular kind of partially instantiated eventuality type with no typical elements in the arguments or sub-arguments11. 10. the formula expresses the de re reading of the sentence, where e1, e2, e3, john, jack, ic are first order constants respectively referring to the three eventualities, the two boys, and an ice cream. 11. actually, ‘instance’ is a term with a broader meaning. there are instances of typical elements that are not eventualities. for simplicity in this paper we assume ‘instances’ and ‘eventuality tokens’ to be synonymous. 23 robaldo and miltsakaki in order to assert that an eventuality e is a, possibly partial, instance of another abstract eventuality ea, hobbs introduces the predicate (partialinstance ea e). another predicate (instance ea e) specifies that e is a total instantiation of ea. it is a consequence of universal instantiation: any property that holds of an eventuality type is true of any (partial) instance of it. we omit here the axioms that formally assert is-a inheritance between eventuality types and their instances. every relation on eventualities, including logical operators, causal and temporal relations, and even tense and aspect, may be reified into another eventuality. for instance, by asserting (imply′ e e1 e2), we reify the implication from e1 to e2 into an eventuality e and e is, then, thought of as “the state holding between e1 and e2 such that whenever e1 really exists, e2 really exists too”. on the other hand, negation is represented as (not′ e1 e2): e1 is the eventuality of e2’s not existing. the predicates imply′ and not′ are defined to model the concept of ‘inconsistency’. in the next subsection, we show how this concept can be used to construct a uniform account of ‘direct’ and ‘indirect’ contrast. two eventualities e1 and e2 are said to be inconsistent if and only if they (respectively) imply two other eventualities e3 and e4 such that e3 is the negation of e4. the definition is as follows12: (32) (forall (e1 e2) (iff (inconsistent e1 e2) (and (eventuality e1) (eventuality e2) (exists (e3 e4) (and (imply e1 e3) (imply e2 e4)(not’ e3 e4)))))) the concept of reification used in hobbs’ logic is suitable for the study of the semantics of discourse connectives because it allows focusing on their meaning while leaving underspecified details about the eventualities involved. in the case of concession, this amounts to identifying the two eventualities that respectively create and deny the expectation in argc/argd , and define the semantics of concessive relations on them. (32) is an example of ‘axiom schema’. in this logic, an ‘axiom schema’ provides one or more different axioms for each predicate p. axioms determine the expressivity and the computational complexity of the logic. however, the axioms defined in the current version of the logic do not guarantee that the logic is recursively enumerable or computationally tractable. in a real system, we envision handling this problem by defining ontologies for specific domains and making queries to these domains. in what follows, we will briefly illustrate three basic concepts from hobbs’ logic that we utilize in our proposed semantics of concession, namely causality, defeasible implication and likelihood. 7.1.1 causality hobbs’ logic adopts a defeasible account of causality, originally proposed in hobbs (1993). this distinguishes between the monotonic notion of ‘causal complex’ and the non-monotonic, defeasible notion of ‘cause’. as hobbs (1993) explains, when we flip a switch to turn on a light, we say that flipping the switch “caused” the light to turn on. but for this to happen, many other factors need to be satisfied: the bulb is good, the switch is connected to the bulb, there is power in the city, etc. the set of all the states and events that are necessary for the event e to take place as a result, are 12. hobbs defined several axioms to determine the semantics of the predicates. those axioms make use of some metaoperators, e.g. if(f), exists, and forall. 24 corpus-driven semantics of concession called the ‘causal complex’ of e. in a causal complex, the majority of participating eventualities are normally true and therefore presumed to hold. in the light bulb case, it is normally true that the bulb is not burnt out, the wiring is in good condition and the power is on, so the conditions are presumed to hold. what cannot be presumed to hold is whether the switch is on or off. eventualities that cannot be assumed to be true under normal contexts are commonly identified as causes (cf. kayser and nouioua (2009)). based on these ontological grounds, hobbs represents causality in terms of two predicates: (cause′ c e1 e2) and (causalcomplex s e2). the predicate cause says that c is the state holding between e1 and e2 such that the former is a non-presumable cause of the latter. the predicate causalcomplex says that s is the set of all presumable or non-presumable eventualities that are involved in causing e2, including e113. in order to preserve defeasibility, the real existence of the effect e2 does not depend on the real existence of the cause e1 and the causal rule c. in other words, the truth of (rexist e1) and (rexist c) does not imply that (rexist e2) is also true. (rexist e2) is true just in case all the eventualities in the causal complex of c2 really exist, as asserted by the following axiom: (forall (s e) (if (and (causalcomplex s e) (forall (e1)(if (member e1 s) (rexist e1))) ) (rexist e) )) it must be pointed out that in practice we can never specify all the eventualities in a causal complex. for instance, consider the following toy example of concession: (33) although [john studied hard]argc , [he did not pass the exam ]argd . in (33) the expectation “john passed the exam” is created by a defeasible general causal rule “studying hard causes passing exams” that instantiates on the present context. nevertheless, john did not pass the exam. there was a particular unknown reason why he did not, despite his hard studying. determining all context-dependent co-causes that had to be in place in order to properly trigger the causal rule would be clearly impossible. of course, this amounts to saying that nl sentences may be properly interpreted even if causal complexes are unknown. therefore, to conclude, in most cases the causal complex exists, but it is not possible to infer it. 7.1.2 defeasible implication defeasibility does not hold only for causal rules. most of our everyday knowledge is non-monotonic, i.e., only approximately correct. for example, knowing that birds fly allows us to infer that if tweety is a bird, then tweety can fly. this conclusion will be defeated later when we learn that tweety is actually a penguin and therefore does not fly. the example illustrates that we need to be careful about how we model knowledge of the world. hobbs, following mccarthy (1980), models common sense implication via monotonic implication (meta-operator if), but allows for defeasibility via the introduction of the underspecified predicate etc in the antecedent of the implication. 13. this is asserted by the following axiom: (forall (e1 e2) (if (cause e1 e2) (exists (s) (and (causalcomplex s e2) (member e1 s))))) 25 robaldo and miltsakaki (forall (x) (if (and (bird x) (etc)) (fly x))) the formula says that if x is a bird and has other unspecified properties encoded as etc (i.e., x’s wings are robust enough), then x can fly. in other words, the formula describes the prototype of bird with respect to the property of flying. etc is a conjunction of eventualities that are true for the prototype and allow for the property of flying. for non-flying birds, at least one of those properties does not hold. although etc is left underspecified in the formulae, its precise definition depends, and so needs to be indexed, on the corresponding predicate, e.g., bird, in the example above. in order to set up a uniform formal account of concession, we need to introduce a new predicate that denotes non-monotonic implications. let us term this new predicate as ‘nonmonotonicif ’. the predicate nonmonotonicif must have the same syntactic structure as the predicate cause: it must relate two eventualities e1 and e2, and it may be reified into a new eventuality. as for e1 and e2 they can be abstract eventualities or instances. obviously, (nonmonotonicif e1 e2) is true iff e1 defeasibly implies e2. the definition of nonmonotonicif is reported in (34). an eventuality e1 defeasibly implies e2 if and only if for each partial instance of e1 there is a partial instance of e2 for which the meta-predicate if, augmented with an opportune etc predication, holds. of course, if and only if e1 and e2 are two instances, it is necessary to assume that the predicate partialinstance denotes a reflexive relation, i.e., that every eventuality is a partial instance of itself. p1 and p2 are the predicates indicating the types of the eventualities e1 and e2 respectively. hobbs defines a meta-predicate pred to relate an eventuality with the unique predication that describes it: (pred p e) states that p is the predicate whose reification is e. (34) (forall (e1 e2) (iff (nonmonotonicif e1 e2) (forall (e ′1) (if (partialinstance e ′1 e1) (exists (e ′2) (and (partialinstance e ′2 e2) (forall (x1 x2 . . . xn) (if (and (pred p1 e ′1)(p1 x1 x2 . . . xn)(etc)) (and (pred p2 e ′2)(p2 x1 x2 . . . xn))))))))) ) 7.1.3 likelihood eventualities exist in a platonic universe of possible individuals: entities, states and events. as said above, if they happen to actually occur in the real world, that is one of their properties, and we express it with the predicate rexist. real existence is one mode of existence but there are others, too. the eventuality could be part of someone’s beliefs but not occur in the real world. it could be merely possible or likely but not real. it could, also, be unlikely or impossible. an especially important modality is “happening at a particular time”. possibility is one common judgment we make about eventualities in situations of uncertainty. likelihood is another. likelihood is intended as the common sense notion of the mathematical version of probability. mathematically defined probability is a special case of common sense likelihood. likelihood is a qualitative notion intended to model the vague probability judgements we make in everyday life, as when we say that it’s likely to rain or that the train may be late. 26 corpus-driven semantics of concession likelihoods are members of a partially ordered scale of likelihoods. for hobbs, such a scale s satisfies the predicate (likelihoodscale s). the likelihood of an eventuality e is with respect to an implicit set of constraints c defining the sample space. an eventuality c may be defined as a single eventuality ec that reifies the conjunction of all the constraints such that (and ′ ec e1 . . . en) is true, where e1, . . . , en are the eventuality-constraints. the likelihood of e is given in the context where the predicate (rexist ec) is assumed to hold. with the formula (likelihood d e c), where d is a number, e an eventuality, and c a set of constraints, we assert that d is the likelihood of e’s really existing, if the set of eventualities in c really exist and d belongs to the contextually relevant likelihood scale s. we say that a certain eventuality e is ‘likely’ when a set of eventualities c holds, iff the likelihood of e given c is a qualitative value belonging to the highest part of the contextually relevant likelihood scale. (forall (e c) (iff (likely e c) (exists (s d s1) (likelihood d e c) (likelihoodscale s) (belong d s1) (high s1 s)) )) likelihood is connected to other modalities via additional axioms. if the likelihood of an eventuality e with constraints c is the top of the likelihood scale, then e is necessary given c, i.e., it is implied from the latter. if the likelihood of e is the bottom of the likelihood scale, then it is not possible given c. in the next subsection, we use the above definitions of causality, non-monotonic implication and likelihood to define the semantics of concession with respect to the different sources of expectation that we identified in our analysis in the pdtb corpus. 7.2 a refined logical account of concession in this section, we propose logical formulae in hobbs’ logic that represent the meaning of concessive relations using a uniform representation that is minimally adjusted to reflect the interpretation of the different sources of expectation. as discussed in section 1, previous approaches of concession, e.g., winter and rimon (1994) and lagerwerf (1998), mostly focus on how the expectation is denied. the distinction between ‘denial of expectation’ and ‘concessive opposition’ is in this spirit. in ‘denial of expectation’, argd directly denies the expectation. in ‘concessive opposition’, argd entails the negation of the expectation. in this line of work, not a lot of attention has been paid to how argc creates the expectation. in most cases, it is simply assumed that the expectation is created from argc via an underspecified entailment ‘→’. our approach is different in that it focuses on characterizing how the expectation is created rather than how it is denied. in formal terms, we propose to represent semantic concessive relations via the following general pattern in hobbs’ logic: (35) (exist (sac sc ec ee ed) ∧ (partialinstance sc sac) ∧ (φ sc ec ee) ∧ (rexist ec) ∧ (rexist ed) ∧ (inconsistent ee ed) ) 27 robaldo and miltsakaki where φ is a generic predicate referring to the underspecified entailment that creates the expectation; it corresponds to winter&rimon’s ‘→’. according to our analysis, φ can be any of the predicates cause’, nonmonotonicif’, or likely’ as defined by hobbs’ account and illustrated in the previous section. the eventuality ee corresponds to the created expectation. it is created from the eventuality ec, conveyed by argc, via the relation φ. the eventuality ed is conveyed by argd. the formula in (35) asserts that ee and ed are inconsistent, via the predicate (inconsistent ee ed). according to the definition in (32), whether the inconsistency is achieved directly or indirectly remains underspecified. (inconsistent ee ed) comes out true iff ee implies a third, possibly different, eventuality, and ed its negation. obviously, it may be either the case that ed is already the negation of ee (denial of expectation), or that it implies it (concessive opposition). the eventuality sac is a general abstract (defeasible) rule. the formula asserts that the reification of φ, i.e. the eventuality sc, is a more specific instantiation of sac . both ec and ed really exist in the context, as asserted in (35) via the predicate rexist . on the contrary, sac and sc do not necessarily exist in the real world; as exemplified below in (45), they could exist only in the speaker’s beliefs. the general pattern in (35) covers the cases of causality, non-monotonic implication, and correlation. in (36) we repeat the toy examples used earlier to demonstrate the three semantic classes. (36) a. although [john studied hard]argc , [he did not pass the exam ]argd . (causality) b. [penguins are birds ]argc . nevertheless [they do not fly]argd . (implication) c. [john will do his report ]argd , but [he will do it at home]argc . (correlation) the formulae in hobbs’ logic that represent examples (36.a-c) are shown in (37), (38), and (39) respectively. note that, with respect to the general pattern in (35), the three formulae below differ only in the predicate φ. in (37), it has been substituted by cause’, in (38) by nonmonotonicif’, and in (39) by likely’. in the next subsection, we will show that (35) is also able to account for the generalizations identified above in section 4. (37) (exist (sac sc ec ee ed) (partialinstance sc sac) ∧ (cause’ sc ec ee) ∧ (rexist ec) ∧ (rexist ed) ∧ (inconsistent ee ed) ) ec = “john studied hard” sac = “studying hard causes passing exams” ee = “john passed the exam” ed = “john did not pass the exam” 28 corpus-driven semantics of concession (38) (exist (sac sc ec ee ed) (partialinstance sc sac) ∧ (nonmonotonicif’ sc ec ee) ∧ (rexist ec) ∧ (rexist ed) ∧ (inconsistent ee ed) ) ec = “penguins are birds” sac = “birds fly” ee = “penguins fly” ed = “penguins do not fly” (39) (exist (sac sc ec ee ed) (partialinstance sc sac) ∧ (likely’ sc ec ee) ∧ (rexist ec) ∧ (rexist ed) ∧ (inconsistent ee ed) ) ec = “john will do his report” sac = “john does not usually do his reports at home” ee = “john will not do his report at home” ed = “john will do his report at home” 7.3 abductive usage of defeasible rules lagerwerf (1998) identified three different sub-cases of concession. the three examples given in (36) are cases of ‘denial of expectation content’, according to lagerwerf’s terminology. the other two subcases are ‘denial of expectation epistemic’ and ‘denial of expectation speech act’. for convenience, we give the examples again in (40): (40) a. [mary loves you very much ]argd , although [you already know that]argc . (denial of expectation speech act) b. [theo was not exhausted ]argd , although [he was gasping for air]argc . (denial of expectation epistemic) in (40.a) the expectation is denied by the illocution of argd, i.e., its speech act, rather than by its locutionary meaning. it is the fact that i tell you argd, and not argd, that is inconsistent with the expectation, created by the rule “if i know that you already know something, i won’t’ say it to you again”. the formalization of occurrences of concession involving speech acts are simple in winter and rimon (1994) and lagerwerf (1998). the default implication is asserted on the speech act associated with the proposition rather than the proposition itself. speech acts do not raise any problems in our approach, either, and so we will not go into the details. consistent with hobbs’s logic, the speech act of an eventuality is simply reified into a new eventuality, and the defeasible rules are asserted on the latter. the interesting case is ‘denial of expectation epistemic’, shown in (40.b). in this example, there is a causal rule “being exhausted causes gasping for air”, and the expectation is created abduc29 robaldo and miltsakaki tively, i.e., by observing that theo was gasping for air, it may be concluded that he was exhausted. lagerwerf (1998) formalizes this intuition in predicate logic as follows14: ∀x[gfb(x) > b(i, exh(x))] where gfb and exh are predicates denoting the set of individuals gasping for air and the set of exhausted individuals, respectively. b(y,φ) is an epistemic operator asserting that y believes φ, i refers to the speaker, and ‘>’ is the defeasible implication operator defined in asher and morreau (1991). as discussed in section 4, lagerwerf’s intuition is correct in our view, but the proposed formalization seems to deviate from the heart of the intuition. as discussed above in (18.a), the defeasible rule is general and therefore does not apply specifically to any particular speaker. therefore, in the formalization, i should be most properly substituted by a universal quantification over all possible believers. similar remarks may be found in pander maat (1998), who argue that three different perspectives of subjectivity may be ascribed to the belief of the statements in a discourse: ‘objective’, ‘speaker’, and ‘other’ perspective. the latter holds when the belief of a statement is ascribed to people other than the speaker. consider the following examples of ‘denial of expectation’ taken from the pdtb15: (41) a. although [it adopted a poison-pill defense]argc , [the board believed that mr. icahn is more interested in talking the stock price higher than acquiring usx]argd . (causality; denial of expectation epistemic) b. [sure, price action is volatile and that’s scary ]argc . but [all-in-all stocks are still a good place to be ]argd . (implication; denial of expectation epistemic) c. [program trading increases volatility ]argc , but [i don’t think it should be banned ]argd . (causality; denial of expectation content) in (41.a), by observing that it adopted a particular strategy, the board, and not the speaker, should believe that mr. icahn is interested in acquiring usx. on the other hand, in (41.b), just like in (40.b), the statements and the conclusions that may be drawn are presented as true objective facts: by observing that price action is volatile everyone should conclude that all-in-all stocks are no longer a good place to be. finally, (41.c) highlights that the belief of a statement may be ascribed to the speaker also when the defeasible rule is used deductively (lagerwerf’s ‘denial of expectation content’). in (41.c), by observing that program trading increases volatility, the speaker, but not necessarily everybody else, believes that it must be banned. it should be made clear, however, that the assertion of the proper perspectives is orthogonal to the semantics of concession. even the general defeasible rules may, in principle, be asserted in the speaker’s beliefs only. consider the following toy example: (42) although [i watered my feet every morning for one month]argc , [i did not get taller ]argd . (causality) 14. on the other hand, as shown above in (14.b), the defeasible rule used in (40.a) could be formalized as: ∀x[k(i, k(y, x)) >� ¬t(i, y, x)], where k(a, x) is an epistemic operator asserting that a knows x, t(a1, a2, x) is true if a1 tells x to a2, i and y are two constants respectively referring to the speaker and the hearer. 15. we did not correct typos that appeared in the corpus. (41.a) should be corrected to “. . . mr. icahn is more interested in taking the stock . . .” . 30 corpus-driven semantics of concession from the example in (42), we understand that the speaker believes that “watering the feet causes growing up”. in hobbs’, such an abstract causal rule is reified into an eventuality sac. then, in order to ascribe it to the speaker’s subjectivity only, a separate predicate like (b i sac) can be conjoined to the whole formula. in other words, in all other examples seen above, each defeasible rule has been always taken as a true fact, but obviously in the case that the hearer does not believe it, it may be asserted as a speaker’s belief only. in our view, lagerwerf (1998)’s intuition must be formalized exactly as it is stated. in (40.b), the defeasible causal rule is “being exhausted causes gasping for air”, and it yields the expectation abductively. thus, the formula is the one in (43); the only difference with respect to the formulae associated above with occurrences of concession where the expectation is created via causality is the assertion of (cause’ sc ee ec) in place of (cause’ sc ec ee). (43) (exist (sac sc ec ee ed) (partialinstance sc sac) ∧ (cause’ sc ee ec) ∧ (rexist ec) ∧ (rexist ed) ∧ (inconsistent ee ed) ) ec = “theo was gasping for air” sac = “being exhausted causes gasping for air” ee = “theo was exhausted” ed = “theo was not exhausted” if the causes/effects of the causal rule or the causal rule itself are believed by the speaker or by someone else, this is separately asserted via additional conjuncts. the formulae corresponding to (41.a) and (42) are shown in (44) and (45) respectively (in the formulae, b is a constant referring to the board and i is a constant referring to the speaker). (44) (exist (sac sc ec ee ed) (partialinstance sc sac) ∧ (cause’ sc ee ec) ∧ (rexist ec) ∧ (b b sac) ∧ (b b ed) ∧ (inconsistent ee ed) ) ec = “it adopted a poison-pill defence” sac = “mr. icahn’s intention of acquiring usx causes adopting a poison-pill defence” ee = “mr. icahn wants to acquire usx” ed = “mr. icahn does not want to acquire usx” (45) (exist (sac sc ec ee ed) (partialinstance sc sac) ∧ (cause’ sc ec ee) ∧ (rexist ec) ∧ (rexist ed) ∧ (b i sac) ∧ (inconsistent ee ed) ) ec = “i watered my feet” sac = “watering the feet causes getting taller” ee = “i got taller” ed = “i did not get taller” 31 robaldo and miltsakaki 8. conclusions 8.1 sources of expectation we presented an empirical analysis of concession based on the annotations of concessive connectives in the penn discourse treebank. in concessive relations, one argument gives rise to an expectation which is then denied in the second argument. in this paper, we have argued that a proper account of concession should be grounded on how the expectation is created. specifically, we identified four sources of expectation: causality, implication, correlation, and implicature. in causality, the created expectation is causally related to the eventuality that creates it. in implication, the expectation is created on the basis of a specific property associated with the eventuality that creates the expectation and in correlation the expectation is related to the eventuality that creates it via co-occurrence. we termed “implicature” the source of expectation that involves pragmatic inferencing. we leave a formal account of this semantically complex category for future work. to test the reliability of these categories, we conducted an inter-annotator agreement study on one thousand concessive tokens in the pdtb. the high kappa score confirms that the categories can be distinguished reliably. to evaluate the practical merits of our approach over previous accounts of concession, we conducted a second inter-annotator study. in this study, two linguistic students annotated 200 instances from pdtb that had been annotated as concession by one annotator and contrast by the other. we trained the new annotators with the new definition of concession, explaining the four sources of expectations. there was more than 80% agreement in this dataset which is a very significant improvement in making reliable distinctions between contrast and concession. 8.2 formal treatment of concession following earlier work by lagerwerf (1998), we refined the semantics of concession using basic constructs from the logic proposed in hobbs (1998). central to hobbs’ proposal is the notion of reification which allows complex natural language statements to be modelled in predicate logic. we propose that every type of concession presupposes a general rule that holds in the context. we propose a logical account that models the semantics of concession by defining the rule for each type of concession. in the case of causality, we infer a defeasible causal relationship between the eventuality expressed in one argument and the eventuality of the expectation. in implication, we infer that the eventuality expressed in one argument non-monotonically entails the eventuality of the expectation. in correlation we infer that the eventuality expressed in the argument that creates the expectation is likely to co-occur with the eventuality of the expectation. the proposed logic formulae differ by the kind of predicate describing the abstract rule (cause’, nonmonotonicif’, likely’). 8.3 impact and future work we have shown that by identifying the different sources of expectation in concession not only are we able to characterize more precisely the events that give rise to expectations but we obtain more reliable semantic distinctions between concession and contrast. we were able to obtain empirical evidence for this claim by a new inter-annotator study that showed significant inter-annotator agree32 corpus-driven semantics of concession ment for tokens that were previously confusing (tokens of annotator disagreement between contrast and concession). the proposed account of concession and its logical treatment is an improvement over some previous accounts which are insufficient in capturing the range of variants identified in naturally occurring data. we maintain that the identified sources of expectation in concession and their logical treatment adequately demonstrated how the study of discourse relations as attested in naturally occurring data can help us improve our understanding of the semantics of discourse relations. an important aspect of concession, the source of expectation, was overlooked in previous approaches but became apparent when studying the data. thus, addressing our second question “what kind of semantic representation will allow covering the rich range of variants conveying concession and contrast?”, we concluded that defining predicates which take as arguments reified eventualities (and even speech acts) is critical for handling discourse level semantic relations. hobbs’s proposed semantics for natural language has proven to be especially well suited for articulating a uniform model of concession while accounting for the range of variants in a simple and straightforward manner. for future work, we need to delve deeper in the cases in which the created expectation involves pragmatic reasoning. moreover, we need to further test the ambiguity between contrast and concession that, in the present paper, was studied only with respect to 200 tokens that were identified as problematic in the pdtb. finally, although the identification of the four sources of concession was empirically tested against a larger set of occurrences (1000 pdtb tokens), in order to see how our results are generalizable, we advocate further experiments on corpora pertaining to different genres and in other languages. while we believe that our approach is in the right direction, defining an important step towards processing discourse problems automatically, the proposed semantics cannot be readily implemented in current state-of-the-art inference systems. significantly more work would be required to integrate the proposed semantics in a real system, e.g., tacitus system hobbs (1986), montazeri and hobbs (2011), ovchinnikova et al. (2011), which implements hobbs’ logic. on the other hand, most current systems are based on shallow features. a recent proposal along this line is the one of meyer and popescu-belis (2012). they trained a statistical classifier on pdtb data for disambiguating discourse connectives, among which discourse connectives conveying concession and contrast. the classifier involves a large set of syntactic and semantic features and it is used for enhancing the performances of a separate statistical machine translation system. classifiers based on shallow features could also benefit from our work. for instance, the overall performances of meyer and popescu-belis (2012)’s classifier could perhaps be improved by including semantic features specifically aimed at identifying the sources of concessive relations. references w. abraham. discourse particles in german: how does their illocutive force come about? in w. abraham, editor, discourse particles in german, pages 203–252. john benjamins, amsterdam/philadelphia, 1991. j.c. anscombre and o. ducrot. deux mais en francais? lingua, 43:23–40, 1979. n. asher. reference to abstract objects in discourse. dordrecht, kluwer, 1993. 33 robaldo and miltsakaki n. asher and m. morreau. commonsense entailment: a modal theory of nonmonotonic reasoning. in proc. of the 12th international joint conference on artificial intelligence volume 1, pages 387–392, san francisco, ca, usa, 1991. morgan kaufmann publishers inc. e. bach. on time, tense, and aspect: an essay in english metaphysics. in p. cole, editor, radical pragmatics, pages 63–81. academic press, new york, 1981. diane blakemore. denial and contrast: a relevance theoretic analysis of ‘but’,. linguistics and philosophy, 12(1):15–37, 1989. j. carletta. assessing agreement on classification tasks: the kappa statistic. computational linguistics, 22:249–254, 1996. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in proceedings of the second sigdial workshop on discourse and dialogue volume 16, sigdial ’01, pages 1–10, stroudsburg, pa, usa, 2001. association for computational linguistics. m. dascal and t. katriel. between semantics and pragmatics: the two types of but hebrew ’aval’ and ’ela’. theoretical linguistics, 4:143–172, 1977. d. davidson. the logical form of action sentences. in nicholas rescher, editor, the logic of decision and action. university of pittsburgh press, 1967. b. di eugenio. on the usage of kappa to evaluate agreement on coding tasks. in in proceedings of the second international conference on language resources and evaluation, pages 441–444, 2000. a. foolen. polyfunctionality and the semantics of adversative conjunctions. multilingua, 10-12: 79–92, 1991. n. francez. contrastive logic. logic journal of the igpl, 3(5):725–744, 1995. p. grice. the causal theory of perception. in proc. of the aristotelean society, vol. 35, pages 121–152, 1961. p. grice. logic and conversation. in p. cole and j. l. morgan, editors, syntax and semantics: vol. 3: speech acts, pages 41–58. academic press, san diego, ca, 1975. b. grote, n. lenke, and m. stede. ma(r)king concessions in english and german. discourse processes, 24(1):87–117, 1995. j. r. hobbs. on the coherence and structure of discourse. technical report 85-37, stanford, ca: stanford university, center for the study of language and information, 1991. j. r. hobbs. toward a useful notion of causality for lexical semantics. journal of semantics, 22: 181–209, 1993. j.r. hobbs. overview of the tacitus project. computational linguistics, 12(3), 1986. 34 corpus-driven semantics of concession j.r. hobbs. monotone decreasing quantifiers in a scope-free logical form. in in semantic ambiguity and underspecification, pages 55–76. 1995. j.r. hobbs. the logical notation: ontological promiscuity. in chapter 2 of discourse and inference. 1998. available at http://www.isi.edu/∼hobbs/disinf-tc.html. l.r. horn. a natural history of negation. cambridge university press, cambridge, 1989. m. izutsu. contrast, concessive, and corrective: towards a comprehensive study of opposition relations. the journal of pragmatics, 40:646–675, 2008. d. kayser and f. nouioua. from the textual description of an accident to its causes. artificial intelligence, 173:1154–1193, 2009. a. kehler. interpreting cohesive forms in the context of discourse inference. in proc. of the 32th meeting of the association of computational linguistics, 1994. a. kehler. coherence, reference, and the theory of grammar. csli publications, 2002. e. könig. concessive, connectives, and concessive sentences: cross-linguistic regularities and pragmatic principles. in j.a. hawkins, editor, explaining language universal, pages 145–166. blackwell, london, 1983. i. korbayova and b. webber. interpreting concession statements in light of information structure. in bunt h. and muskens r., editors, interpreting concession statements in light of information structure, pages 145–172. kluwer academic publishers, 2007. l. lagerwerf. causal connectives have presuppositions: effects on coherence and discourse structure. phd thesis, the hague: holland academic graphics, the netherlands, 1998. r. lakoff. if’s, and’s, and but’s about conjunction. in in c.j. fillmore and d.t. langendoen, editors, studies in linguistic semantics. holt, reinhart and winston, new york, 1971. e. lang. the semantics of coordination. john benjamins b.v., amsterdam, 1984. e. lang. adversative connectors on distinct levels of discourse: a re-examination of eve sweetser’s three-level approach. in bernd / couper-kuhlen kortmann, editor, cause, condition, concession, contrast, pages 235–256. berlin: mouton de gruyter, 2000. w. c. mann and s. a. thompson. rhetorical structure theory: toward a functional theory of text organization. text, 8(3):243–281, 1988. j. mccarthy. circumscription: a form of nonmonotonic reasoning. artificial intelligence, 13: 27–39, 1980. j. mccarthy. epistemological problems of artificial intelligence. in proc. of the international joint conference on artificial intelligence, pages 1038–1044, cambridge, massachusetts, 2002. t. meyer and a. popescu-belis. using sense-labeled discourse connectives for statistical machine translation. in proc. of the joint workshop on exploiting synergies between information retrieval and machine translation (esirmt) and hybrid approaches to machine translation (hytra), pages 129–138, 2012. 35 robaldo and miltsakaki e. miltsakaki, l. robaldo, a. lee, and a. joshi. sense annotation in the penn discourse treebank. in proc. of computational linguistics and intelligent text processing, pages 275–286, cambridge, massachusetts, 2008. n. montazeri and j.r. hobbs. elaborating a knowledge base for deep lexical semantics. in j. bos and s. pulman, editors, in proc. of 9th international workshop on computational semantics, pages 195–204, 2011. j. moore and m. pollack. a problem for rst: the need for multi-level discourse analysis. computational linguistics, 18:537–544, 1992. e. ovchinnikova, n. montazeri, t. alexandrov, j.r. hobbs, m.c. mccord, and r. mulkar-mehta. abductive reasoning with a large knowledge base for discourse processing. in proc. of the ninth international conference on computational semantics (iwcs 2011), pages 225–234, 2011. h. pander maat. classifying negative coherence relations on the basis of linguistic evidence. journal of pragmatics, 30(2):177–204, 1998. m. poesio and r. artstein. anaphoric annotation in the arrau corpus. in proc. of language resources and evaluation (lrec08), 2008. r. prasad, e. miltsakaki, n. dinesh, a. lee, a. joshi, b. webber, and l. robaldo. the penn discourse treebank 2.0. annotation manual. technical report ircs-06-01, institute of research in cognitive science, university of pennsylvania, 2008. j. sanders. perspective in narrative discourse. phd thesis, tilburg university, the netherlands, 1994. t.j.m. sanders, w.p.m. spooren, and l.g.m. noordman. toward a taxonomy of coherence relations. discourse processes, 15(1), 1992. t.j.m. sanders, w.p.m. spooren, and l.g.m. noordman. coherence relations in a cognitive theory of discourse representation. cognitive linguistics, 4(2), 1993. j. spenader and a. lobanova. reliable discourse markers for contrast relations. in proc. of the eighth international conference on computational semantics (iwcs-8), pages 210–221, 2009. w. spooren. some aspects of the form and interpretation of global contrastive coherence. phd thesis, nijmegen university, the netherlands, 1989. e. sweetser. from etymology to pragmatics. metaphorical and cultural aspects of semantic structure. cambridge university press, cambridge, 1990. y. versley and a. gastel. linguistic tests for discourse relations in the tba-d/z treebank of german. dialogue & discourse, 4(2):142173, 2013. a. von klopp. but and negation. nordic journal of linguistics, 17(1), 1994. y. winter and m. rimon. contrast and implication in natural language. journal of semantics, 11: 365–406, 1994. 36 dialogue and discourse 7(1) (2016) 1-49 doi: 10.5087/dad.2016.101 evaluation in discourse: a corpus-based study farah benamara benamara@irit.fr irit, université de toulouse. 118 route de narbonne, 31062, toulouse, france. nicholas asher asher@irit.fr irit-cnrs. 118 route de narbonne, 31062, toulouse, france. yvette yannick mathieu ymathieu@linguist.univ-paris-diderot.fr bat olympes de gouges, 8 place paul ricoeur, case courrier 7031, 75205 paris cedex 13, france. vladimir popescu popescu@irit.fr irit, université de toulouse. 118 route de narbonne, 31062, toulouse, france. baptiste chardon baptiste.chardon@synapse-fr.com synapse développement. 33 rue maynard, 31100 toulouse. editor: massimo poesio abstract this paper describes the casoar corpus, the first manually annotated corpus exploring the impact of discourse structure on sentiment analysis with a study of movie reviews in french and in english as well as letters to the editor in french. while annotating opinions at the expression, sentence, or document level is a well-established task and relatively straightforward, discourse annotation remains difficult, especially for non experts. therefore, combining opinion and discourse annotations pose several methodological problems that we address here. we propose a multi-layered annotation scheme that includes: the complete discourse structure according to the segmented discourse representation theory, the opinion orientation of elementary discourse units and opinion expressions, and their associated features (including polarity, strength, etc.). we detail each layer, explore the interactions between them, and discuss our results. in particular, we examine the correlation between discourse and semantic category of opinion expressions, the impact of discourse relations on both subjectivity and polarity analysis, and the impact of discourse on the determination of the overall opinion of a document. our results demonstrate that discourse is an important cue for sentiment analysis, at least for the corpus genres we have studied. 1. introduction sentiment analysis has been one of the most popular applications of natural language processing for over a decade both in academic research institutions and in industry. in this domain, researchers analyze how people express their sentiments, opinions and points of view from natural language data such as customer reviews, blogs, fora and newspapers. opinions concern evaluations expressed by a holder (a speaker or a writer) towards a topic (an object or a person). an evaluation is characterized by a polarity (positive, negative or neutral) and a strength that indicates the opinion degree of positivity or negativity. example (1), extracted from our corpus of movie reviews, illustrates these phenomena1. in this review, the author expresses three opinions: the first two are explicitly lexical1. this example has been extracted from metacritic website as it is, including typos and english errors. c©2016 benamara, asher, mathieu, popescu, chardon submitted 04/14; accepted 11/15; published online 01/16 benamara, asher, mathieu, popescu, chardon ized opinion expressions (underlined in the example) whereas the last one (in italic) is an implicit positive opinion since it contains no subjective lexical cues. (1) what a great animated movie. i was so thrilled by seeing it that i didn’t movie a single second from my seat. from a computational perspective, most current research examine the expression and extraction of opinion at two main levels of granularity: the document and the sentence2. at the document level, the standard task is to categorize documents globally as being positive or negative towards a given topic (turney (2002); pang et al. (2002); mullen and nigel (2004); blitzer et al. (2007)). in this classification problem, all opinions in a document are supposed to be related to only one topic3. overall document opinion is generally computed on the basis of aggregation functions (such as the average or the majority) that take as input the set of explicit opinions scores of a document and output either a polarity rating or an overall multi-scale rating (pang and lee (2005); lizhen et al. (2010); leung et al. (2011)). at the sentence level, on the other hand, the task is to determine the subjective orientation and then opinion orientation of sequences of words in the sentence that are determined to be subjective or express an opinion (yu and vasileios (2003); riloff et al. (2003); wiebe and riloff (2005); taboada et al. (2011)). this second level also assumes that a sentence usually contains a single opinion. to better compute the contextual polarity of opinion expressions, some researchers have used subjectivity word sense disambiguation to identify whether a given word has a subjective or an objective sense (akkaya et al. (2009)). other approaches identify valence shifters (viz. negations, modalities and intensifiers) that strengthen, weaken or reverse the prior polarity of a word or an expression (polanyi and zaenen (2006); shaikh et al. (2007); moilanen and pulman (2007); choi and cardie (2008)). the contextual polarity of individual expressions is then used for sentence as well as document classification (kennedy and inkpen (2006); li et al. (2010)). we believe that viewing opinions in a text as a simple aggregation of opinion expressions identified locally is not appropriate. in this paper, we argue that discourse structure provides a crucial link between local and document levels and is needed for a better understanding of the opinions expressed in texts. to illustrate this assumption, let us take the example (2), extracted from our corpus of french movie reviews. (2) contains four opinions: the first three are strongly negative while the last one (introduced by the conjunction but in the last sentence) is positive. a bag of words approach would classify this review as negative, which is contrary to intuitions for this example. (2) les personnages sont antipathiques au possible. le scénario est complètement absurde. le décor est visiblement en carton-pâte. mais c’est tous ces éléments qui font le charme improbable de cette série. the characters are unpleasant. the scenario is totally absurd. the decoration seems to be made of cardboard. but, all these elements make the charm of this tv series. discourse structure can be a good indicator of the subjectivity and/or the polarity orientation of a sentence. in particular, general types of discourse relations that link clauses together like parallel, contrast, result and so on from theories like rhetorical structure theory (rst) (mann and 2. there is also a third level of granularity not detailed here which is the aspect or feature level where opinions are extracted according to the target domain features (liu (2012)). 3. of course, this assumption is debatable. for instance in forums, blogs and news, opinions are related to several topics. 2 evaluation in discourse thompson (1988)) or segmented discourse representation theory (sdrt) (asher and lascarides (2003)) furnish important clues for recognizing implicit opinions and assessing the overall stance of texts. for instance4, sentences related by the discourse relations parallel or continuation often share the same subjective orientation like in mary liked the movie. her husband too. here, parallel (triggered by the discourse marker too) holds between the two sentences and allows us to detect the implicit opinion conveyed by the second sentence. polarity is often reversed in case of contrast and usually preserved in case of parallel and continuation. result on the other hand does not have a strong effect on subjectivity and polarity is not always preserved. for instance, in your life is miserable. you don’t have a girlfriend. so, go see this movie, the positive polarity of the recommendation follows the negative opinions expressed in the first two sentences. in case of elaboration, subjectivity may not be preserved, in contrast to polarity (it would be difficult to say the movie was excellent. the actors were bad). finally, attribution plays a role only when its second argument is subjective, as in i suppose that the employment policy will be a disaster. in this case, depending on the reported speech act used to introduce the opinion, attribution affects the degree of commitment of the author and the holder (asher (1993); prasad et al. (2006)). discourse-based opinion analysis is an emerging research area (asher et al. (2008); taboada et al. (2008, 2009); somasundaran (2010); zhou et al. (2011); heerschop et al. (2011); zirn et al. (2011); polanyi and van den berg (2011); trnavac and taboada (2010); mukherjee and bhattacharyya (2012); lazaridou et al. (2013); trivedi and eisenstein (2013); wang and wu (2013); hogenboom et al. (2015); bhatia et al. (2015)). studying opinion within discourse gives rise to new challenges: what is the role of discourse relations in subjectivity analysis? what is the impact of the discourse structure in determining the overall opinion conveyed by a document? does a discourse based approach really bring additional value compared to a classical bag of words approach? does this additional value depend on corpus genre? the casoar project (a two year dga-rapid project (2010-2012) involving toulouse university and an nlp company synapse développement) aimed to address these questions by gathering and analyzing a corpus of movie reviews in french and in english as well as letters to the editor in french. it extended our earlier work where segmented discourse representation theory (sdrt) (asher and lascarides (2003)) was used to study opinion within discourse (asher et al. (2008, 2009)). before moving to real scenarios that rely on automatic discourse annotations, we first wanted to measure the impact of discourse structure on opinion analysis in manually annotated data. while annotating opinions at the expression, sentence or document level achieved a relatively good interannotator agreements, at least for explicit opinion recognition, and opinion polarity (wiebe et al. (2005); toprak et al. (2010)), annotation of complete discourse structure is a more difficult task, especially for non experts (carlson et al. (2003); afantenos et al. (2012)). combining opinion and discourse annotations poses several methodological problems: the choice of the corpus in terms of genre and document length, the definition of the annotation model, and the description of the annotation guide so as to minimize errors, etc. a second point was more challenging: what is the most appropriate level to annotate opinion in discourse? should we annotate opinion texts using a small set of discourse relations? or should we use a larger set? should discourse annotations annotators be simply asked to follow their intuitions after having been given a gloss of the discourse relations to be used, or should we provide them with a precise description of the structural constraints regarding the underlying discourse theory? 4. in this paragraph, assertions are based on our own observations of the data. they have however been empirically validated in this corpus study, as shown later in this paper. 3 benamara, asher, mathieu, popescu, chardon we developed a multi-layered annotation scheme that includes: the complete discourse structure according to sdrt, opinion orientation of elementary discourse units and opinion expressions, and their associated features. in this paper, we detail each layer, explore the interactions between them and discuss our results. in particular, we examine: the correlation between discourse and semantic category of opinion expressions focusing on the role of evaluation to identify discourse relations, the impact of discourse relations on both subjectivity and polarity analysis, and the impact of discourse on the determination of the overall opinion of a document. our results demonstrate that discourse is an important cue for sentiment analysis, at least for the corpus genres we have studied. the paper is organized as follows. section 2 gives some background on annotating sentiment and discourse, and provides a brief introduction to sdrt, our theoretical framework. section 3 presents our corpus. section 4 details the annotation scheme, annotation campaign, and reliability of the scheme. section 5 gives our results. we end the paper by a discussion where we highlight the main conclusions of our corpus-based study and discusses the portability and applicability of the annotation scheme. 2. background 2.1 existing corpora annotated with sentiment there are several existing annotated resources for sentiment analysis. each resource can be characterized in terms of the corpus used, the basic annotation unit and annotation levels. in this section, we overview main existing resources according to these three criteria. 2.1.1 data several authors have focused on annotating a single corpus genre like movie reviews (pang and lee (2004)), book reviews (read and carroll (2012)), news (wiebe and riloff (2005)), political debates (somasundaran et al. (2007); somasundaran and wiebe (2010)) and blogs (liu et al. (2009)). well known resources include mpqa (wiebe et al. (2005)), the jdpa-corpus (kessler et al. (2010)) and the darmstadt-corpus (toprak et al. (2010)). multi-domain sentiment analysis has been explored in blitzer et al. (2007) with a corpus of product reviews taken from amazon.com5. compared to english, few resources have been developed for other languages. in french, the blogoscopy corpus (daille et al. (2011)) is composed of 200 annotated posts and 612 associated comments. there is also bestgen et al. (2004)’s dataset composed of 702 sentences extracted from a newspaper6. in spanish, the tass corpus7 is composed of 70,000 tweets annotated with global polarity as well as an indication of the level of agreement or disagreement of the expressed sentiment within the content. in german, the mlsa8 (clematide et al. (2012)), is a publicly available corpus composed of 270 sentences manually annotated for objectivity and subjectivity. finally for italian, the senti-tut corpus9 includes sentiment annotations of irony in tweets (bosco et al. (2013)). multilingual sentiment annotation has also been explored: the emotiblog corpus consists of labeled blog posts in spanish, italian and english (boldrini et al. (2012)), mihalcea et al. (2007) manually 5. http://www.cs.jhu.edu/˜mdredze/datasets/sentiment/ 6. https://sites.google.com/site/byresearchoa/home/ 7. http://www.daedalus.es/tass2013/about.php 8. http://iggsa.sentimental.li/index.php/downloads/ 9. http://www.di.unito.it/tutreeb/sentitut.html 4 http://www.cs.jhu.edu/~mdredze/datasets/sentiment/ https://sites.google.com/site/byresearchoa/home/ http://www.daedalus.es/tass2013/about.php http://iggsa.sentimental.li/index.php/downloads/ http://www.di.unito.it/tutreeb/sentitut.html evaluation in discourse annotated 500 sentences in english, romanian, and spanish10. finally, banea et al. (2010) automatically annotated english, arabic, french, german, romanian, and spanish news documents. in this paper, we aim to annotate opinion in discourse in multi-genre documents (movie reviews and news reactions) in french and movie reviews in english. to our knowledge, no one has conducted a corpus-based study across genres and languages that analyzes how opinion and discourse interact at different levels of granularity (expression, discourse unit and the whole document). thus, there is almost no extent work for us to compare ourselves to other. even though several annotation schemes already exist for the expression/phrase level (mpqa, jdpa-corpus, darmstadt-corpus, mlsa), the descriptive analysis investigating the interaction between sentiment and discourse is novel. 2.1.2 basic annotation unit state-of-the art opinion annotation campaigns take the expression (a set of tokens), sentence or document as their basic annotation unit. however, annotating opinion in discourse required to move to start with elementary discourse unit (edu) which is the intermediate level between the sentence and the document. indeed, the sentence level is not appropriate for analyzing opinions in discourse, since, in addition to objective clauses, a single sentence may contain several opinion clauses that can be connected by rhetorical relations. moving to the clause level is also not appropriate, since several opinion expressions can be discursively related as in the movie is great but too long where we have a contrast relation introduced by the marker but. therefore, we need to move to a fine-grained and semantically motivated level, the edu. annotating edus not quite corresponding to either sentences or clauses has been standard in discourse annotation efforts for many years (see section 2.2 for an overview). however, annotating sentiment within edus is still marginal. among the few annotated sentiment corpora at the edu level, we cite asher et al. (2009), who analyzed explicit opinion expressions within edus. somasundaran et al. (2007) used a similar level in order to detect the presence of sentiment and arguing in dialogues. zirn et al. (2011) performed subjectivity analysis at the segment level. they used a corpus of product reviews segmented using the hilda tool11, an rst discourse parser. lazaridou et al. (2013) used the slseg software package12 to segment their corpus into edus following rst. the corpus was then used to train a joint model for unsupervised induction of sentiment, aspect and discourse information. in this paper, documents are segmented according to sdrt principles. 2.1.3 annotation levels our annotation scheme is multi-layered and includes: the complete discourse structure, segment opinion orientation, and opinion expressions. at the document level, we propose to annotate the document overall opinion as well as its full discourse structure following the sdrt framework. global opinion annotation resembles previous document level annotation (pang and lee (2004)). to the best of our knowledge, this is the first sentiment dataset that incorporates discourse structure annotation. 10. http://www.cse.unt.edu/˜rada/downloads.html#msa 11. http://nlp.prendingerlab.net/hilda 12. http://www.sfu.ca/˜mtaboada/research/slseg.html 5 http://www.cse.unt.edu/~rada/downloads.html#msa http://nlp.prendingerlab.net/hilda http://www.sfu.ca/~mtaboada/research/slseg.html benamara, asher, mathieu, popescu, chardon at the segment level, we propose to associate to each edu a subjectivity type (among four main types: explicit evaluative, subjective non evaluative, implicit, and objective) as well as polarity and strength. segment opinion type mainly follows wiebe et al. (2005); toprak et al. (2010) and liu (2012). wiebe et al. (2005) already proposed an expression-level annotation scheme that distinguishes between explicit mentions of private states, speech events expressing private states, and expressive subjective elements. toprak et al. (2010), following wiebe et al. (2005), distinguished in their annotation scheme (consumer reviews) between explicit opinions and facts that imply opinions. finally, liu (2012) has also observed that subjective sentences and opinionated sentences (which are objective or subjective sentences that express implicit positive or negative opinions) are not the same, even though opinionated sentences are often a subset of subjective sentences. in this work, we propose, in addition, to study what are the correlations between segment opinion types and the overall opinion on the one hand (cf. section 5.3.4), and between segment types and rhetorical relations (cf. sections 5.3.2 and 5.3.3). the opinion expression is the lowest level and focuses on annotating all the elements associated to an opinion within a segment: (1) the opinion span, excluding operators (negation, modality, intensifier, and restrictor), (2) opinion polarity and strength, (3) opinion semantic category, (4) topic span, (5) holder span, and (6) operator span. our annotation at this level is very similar to state of the art annotation schema at the expression level (e.g. mpqa, jdpacorpus, darmstadt-corpus, mlsa corpus). however, in addition, we explore the link between discourse and opinion semantic category of subjective segments (cf. section 5.3.1). 2.2 existing corpora annotated with discourse the annotation of discourse relations in language can be broadly characterized as falling under two main approaches: the lexically grounded approach and an approach that aims at complete discourse coverage. perhaps the best example of the first approach is the penn discourse treebank (prasad et al. (2008)). the annotation starts with specific lexical items, most of them conjunctions, and includes two arguments for each conjunction. this leads to partial discourse coverage, as there is no guarantee that the entire text is annotated, since parts of the text not related through a conjunction are excluded. on the positive side, such annotations tend to be reliable. pdtb-style annotations have been carried out in a variety of languages (arabic, chinese, czech, danish, dutch, french, hindi and turkish). complete discourse coverage requires annotation of the entire text, with most, if not all, of the propositions in the text integrated in a structure. it includes work from two main different theoretical perspectives, either intentionally or semantically driven. the first perspective has been investigated within rhetorical structure theory, rst (mann and thompson (1988)), whereas the second includes segmented discourse representation theory, sdrt (asher and lascarides (2003)), and the discourse graph bank model (wolf and gibson (2006)). rst annotated resources exist in basque, dutch, german, english, portuguese and spanish. corpora following sdrt exist in arabic, french and english. to get a complete structure for a text, three decisions need to be made: • what are the elementary discourse units (edu)? • how do elementary units combine to form larger units and attach to other units? • how are the links between discourse units labelled with discourse relations? 6 evaluation in discourse many theories such as rst take full sentences or at least tensed clauses as the mark of an edu. sdrt, as developed in (asher and lascarides (2003)), was largely mute on the subject of edu segmentation, but in general also followed this policy. concerning attachment, most discourse theories define hierarchical structures by constructing complex segments (cdus) from edus in recursive fashion. rst proposes a tree-based representation, with relations between adjacent segments, and emphasizes a differential status for discourse components (the nucleus vs. satellite distinction). captured in a graph-based representation, with long-distance attachments, sdrt proposes relations between abstract objects using a relatively small set of relations. identifying these relations is a crucial step in discourse analysis. given two discourse units that are deemed to be related, this step labels the attachment between the two discourse units with discourse relations such as elaboration, explanation, conditional, etc. for example in [this is the best book]1 [that i have read in along time.]2 we have elaboration(1, 2). their triggering conditions rely on the propositional contents of the clauses a proposition, a fact, an event, a situation –the so-called abstract objects (asher (1993)) or on the speech acts expressed in one unit and the semantic content of another unit that performs it. some instances of these relations are explicitly marked i.e., they have cues that help identifying them such as but, although, as a consequence. others are implicit i.e., they do not have clear indicators, as in i didn’t go to the beach. it was raining. in this last example to infer the intuitive explanation relation between the clauses, we need detailed lexical knowledge and probably domain knowledge as well. in this paper, we aim to annotate the full discourse structure of opinion documents following a semantically driven approach, as done in sdrt. 2.3 overview of the segmented discourse representation theory (sdrt) sdrt is a theory of discourse interpretation that extends kamp’s discourse representation theory (drt) (kamp and reyle (1993)) to represent the rhetorical relations holding between edus, which are mainly clauses, and also between larger units recursively built up from edus and the relations connecting them. sdrt aims at building a complete discourse structure for a text or a dialogue, in which every constituent is linked to some other constituent. we detail below the three steps needed to build this structure, namely: edu determination, attachment, and relation labelling. 2.3.1 edu determination we follow the principles defined in the annodis project13 (afantenos et al. (2012)). in annodis, an edu is mainly a sentence or a clause in a complex sentence that typically corresponds to verbal clauses, as in [i loved this movie]a [because the actors were great]b where the clause introduced by the marker because, indicates a cutting point. we have here the relation explanation(a, b). an edu can also correspond to other syntactic units describing eventualities, such as prepositional and noun phrases, as in [after several minutes,]a [we found the keys on the table]b where we have two edus related by frame(a, b). in addition, a detailed examination of the semantic behavior of appositives, non restrictive relative clauses and other parenthetical material in our corpora, revealed that such syntactic structures also contributed edus14. such constructions provide semantic contents that do 13. this project aimed at building a diversified corpus of written french texts enriched with a manual annotation of discourse structures. the resource can be downloaded here http://w3.erss.univ-tlse2.fr/annodis 14. in rst, embedding is handled by the “same unit” relation. to a much more limited extent, pdtb also allows for nominalizations to be arguments to relations. 7 http://w3.erss.univ-tlse2.fr/annodis benamara, asher, mathieu, popescu, chardon not fall within the scope of discourse relations or operators between the constituents in which they occur. in example (3), we see that the apposition in italic font does not or at least needs not fall within the scope of the conditional relation on a defensible interpretation of the text. such “nested” edus are a useful feature in sentiment analysis as edus conveying opinions may be isolated from surrounding ”objective” material, as in the movie review in (4). finally, concerning attributions, we segment cases like “i say that i am happy” into two edus: “i say” and “that i am happy”. (3) if the former president of the united states, who has been all but absent from political discussions since the 2008 election, were to weigh in on the costs of the economic shutdown, the radical republicans might be persuaded to vote to lift the debt ceiling. (4) [the film [that distressed me the most] is cry of freedom]. in addition to this definition, we observe in our corpora that several opinion expressions (often conjoined np or ap clauses) could be linked by discourse relations. we thus resegment such edus into separate units. annodis segmentation principles were then refined in order to take into account the particularities of opinion texts. for example, the following sentence: [the movie is long, boring but amazing] is segmented as follows: [the movie is long,]1 [boring]2 [but amazing]3 with continuation(1,2) and contrast([1,2],3), [1,2] being a complex discourse unit. even if segments 2 and 3 do not follow the edu standard definition (they are neither sentences nor clauses), we believe that such fine-grained segmentation will facilitate polarity analysis at the sentence level. during the annotation of edus, we consider that argument naming generally follows the linear order in the text. in case of embedding, the main clause is annotated first. for instance in (4), we have: [the film [that distressed me the most]2 is cry of freedom]1. 2.3.2 attachment decision in sdrt, a discourse representation for a text t is a structure in which every edu of t is linked to some (other) discourse unit, where discourse units include edus of t and complex discourse units (cdus) built up from edus of t connected by discourse relations in recursive fashion. proper sdrss form a rooted acyclic graph with two sorts of edges—edges labeled by discourse relations that serve to indicate rhetorical functions of discourse units, and unlabeled edges that show which constituents are elements of larger cdus. sdrt allows attachment between non adjacent discourse units and for multiple attachments to a given discourse unit15, which means that the structures created are not always trees but rather directed acyclic graphs. these graphs are constrained by the right frontier principle that postulates that each new edu should attach either to the last discourse unit or to one that is super-ordinate to it via a series of subordinate relations and complex segments. one of the most important feature that makes sdrt an attractive choice for studying the effects of discourse structure on opinion analysis is the scope of relations. for instance, if an opinion is within the scope of an attribution that spans several edus, then knowing the scope of the attribution will enable us to determine who is in fact expressing the opinion. similarly, if there is a contrast that has scope over several edus in its left argument, this can be important to determine the overall 15. in sdrt, several discourse relations can hold between two constituents if they are of the same type, i.e., either all coordinating or all subordinating and their semantics effect are compatible. so semantics puts important constraints on relations. consider the example: [john kissed mary.]1 [she then slapped him]2 [and his wife did too, at the same time.]3. in this example, segment 1 and 2 are related by both a narration and a result. we have thus the following annotation: narration(1, [2, 3]) ∧ result(1, [2, 3]) ∧ parallel(2, 3), [2,3] being a cdu. 8 evaluation in discourse contribution of the opinions expressed in the arguments of the contrast. to get this kind of information, we need to have discourse annotations in which the scopes of discourse relations are clear and determined for an entire discourse graph. example (5) taken from the annodis corpus (afantenos et al. (2012)) illustrates what are called long distance attachments16. (5) [suzanne sequin passed away saturday at the communal hospital of bar-le-duc,]3 [where she had been admitted a month ago.]4 [she would be 79 years old today.]5 [. . . ] [her funeral will be held today at 10h30 at the church of saint-etienne of bar-le-duc.]6. a causal relation like result, or at least a temporal narration holds between 3 and 6, but it should not scope over 4 and 5 if one does not wish to make sequin’s admission to the hospital a month ago and her turning 79 a consequence of her death last saturday. 2.3.3 relation labelling sdrt models the semantics/pragmatics interface using discourse relations that describe the rhetorical roles played by utterances in context, on the basis of their truth conditional effects on interpretation. relations are constrained by: semantic content, pragmatic heuristics, world knowledge and intentional knowledge. they are grouped into coordinating relations that link arguments of equal importance and subordinating relations linking an important argument to a less important one. this semantic characterization of discourse relations has two advantages for our study: first, the semantics of discourse relations makes it more straightforward to study their interactions with the semantics of subjective expressions, and secondly the semantic classification in sdrt leads to a smaller taxonomy of discourse relations than that given in rst, enabling an initial study of the interaction of discourse structure and opinion to find generalisations. additionally, the fact that in sdrt multiple relations may relate one discourse unit to other discourse units allows us to study more complex interactions than it would be possible in the other theories. figure 1 gives an example of the discourse structure of the example (6), familiar from asher and lascarides (2003). in this figure, circles are edus, rectangles are complex segments, horizontal links are coordinating relations while vertical links represent subordinating relations. (6) [john had a great evening last night.]1 [he had a great meal.]2 [he ate salmon.]3 [he devoured lots of cheese.]4 [he then won a dancing competition.]5 3. the casoar corpus we selected data according to four criteria: document genre, the number of documents per topic, document length and the type of opinion conveyed in the document. to better capture the dependencies between discourse structure and corpus genre, the annotation campaign should be conducted on different types of online corpora, each with a distinctive style and audience. for each corpus, topics (a movie, a product, an article, etc.) have to be selected according to their related number of documents or reviews. our hypothesis was that the more attractive a topic is (i.e., it aroused a great 16. for a discussion of long-distance discourse relations in rst, see (marcu (2000)). 9 benamara, asher, mathieu, popescu, chardon figure 1: an example of a discourse graph. number of reactions), the more opinionated the reviews are. in addition, the number of positive and negative documents has to be balanced. given that discourse annotation is time consuming and error prone, especially for long texts where long distance attachments are frequent, documents should not be too long. on the other hand, documents should have an informative discourse structure and hence should not be too short either. finally, the data should contain explicit opinion expressions as well as implicit opinions. one of our aims was to measure how these kinds of opinions are assessed in discourse. given these criteria, we chose to build our own corpus and not to rely on existing opinion datasets. indeed, in french, the only existing and freely available opinion dataset (the blogoscopy corpus daille et al. (2011)17) was not available when we began our annotation campaign. in english, there are several freely available corpora already annotated with opinion information. among them, we have studied four resources: the well known mpqa (wiebe et al. (2005)) corpus18, the sentiment polarity dataset and the subjectivity dataset19 (pang and lee (2004)), and the customer reviews dataset20(hu and liu (2004)). we chose not to build our discourse based opinion annotation on the top of mpqa for two reasons. first, text anchors which correspond to opinion in mpqa are not well defined since each annotator is free to identify expression boundaries. this is problematic if we want to integrate rhetorical structures into the opinion identification task. secondly, mpqa often groups discourse indicators (but, because, etc.) with opinion expressions not leading to any guarantee that text anchors will correspond to a well formed discourse unit. the sentiment polarity dataset consists of 1,000 positive and 1,000 negative processed reviews annotated at the document level. however, it was not appropriate for our purposes because the documents in this corpus are very long (more than 30 sentences per document) which would have made the annotation of the discourse structure too hard. on the other hand, the subjectivity dataset 17. http://www.lina.univ-nantes.fr/?blogoscopie,762.html 18. http://mpqa.cs.pitt.edu 19. www.cs.cornell.edu/people/pabo/movie-review-data 20. http://www.cs.uic.edu/˜liub/fbs/sentiment-analysis.html 10 http://www.lina.univ-nantes.fr/?blogoscopie,762.html http://mpqa.cs.pitt.edu www.cs.cornell.edu/people/pabo/movie-review-data http://www.cs.uic.edu/~liub/fbs/sentiment-analysis.html evaluation in discourse contains 5,000 subjective and 5,000 objective processed sentences. only sentences or snippets containing at least 10 tokens were included along with their automating labelling decision (objective vs. subjective), as shown in (7). since sentences are short (at most 3 discourse units), this corpus also did not meet our criteria. finally, the customer reviews dataset consists of annotated reviews of five products (digital camera, cellular phone, mp3 player and dvd player), extracted from amazon.com. this corpus provides only target and polarity annotations at the sentence or the snippet level focusing on explicit opinion sentences (cf. (8) and (9) where [u] indicates that the target is not lexicalized (implicit)). (7) nicely combines the enigmatic features of ’memento’ with the hallucinatory drug culture of ’requiem for a dream . ’ (8) camera[+2] ## this is my first digital camera and what a toy it is... (9) size[+2][u] ## it is small enough to fit easily in a coat pocket or purse. to conclude, none of the above pre-existing annotated corpora firs our objectives. we thus built the following corpora, summarized in table 1: • the french data are composed of two corpora: french movie/product reviews (fmr) and french news reactions (fnr). the movie reviews were taken from allocine.fr, book and video game reviews from amazon.fr, and restaurant reviews from qype.fr. the news reactions, extracted from lemonde.fr, are reactions to articles from the politics and economy sections of the “le monde” newspaper. we selected those topics (movies, products, articles) that are associated to more than 10 reviews/reactions. in order to guarantee that the discourse structure is informative enough, we also filtered out documents containing less than three sentences. in addition, for fmr, we balanced the number of positive and negative reviews according to their corresponding general evaluation (i.e., stars21). for fnr, reactions that are responses to other reactions were removed. • the english data are movie reviews (emr) from metacritic22. the choice of movie reviews is motivated firstly by the fact that this genre is widely used in the field and secondly, by our aim to compare how opinions are expressed in discourse in different languages (movie reviews were also selected for the french annotation campaign). the selection procedure (number of reviews per movie, number of sentences per review) was the same as for the one used in french data selection. number of documents selected topics fmr 180 films (6), books and video games (6), restaurants (13), tv series (20) fnr 131 politics (5), economy (6), international (2) emr 110 films (11) table 1: characteristics of our data. 21. the star scale was 1-5 and neutral reviews (3-star) were equally distributed in the positive/negative class. 22. http://www.metacritic.com 11 http://www.metacritic.com benamara, asher, mathieu, popescu, chardon 4. methods 4.1 annotation scheme the annotation scheme is multi-layered, and includes: (1) the complete discourse structure according to sdrt, (2) opinion orientation of edus, and (3) opinion expressions, and their associated features. each level has its own annotation manual and annotation guide, as described in the next sections. in the remainder of this paper, all the examples are extracted from our corpora. examples from emr are given in english while examples from fmr and fnr are given in french along with their direct english translation (when possible). note however that there are substantial semantic differences between the two languages. 4.1.1 the document level in this level, annotators were asked to give the document overall opinion towards the main topic using a five-level scale, where 0 indicates a very bad (negative) opinion and 4 a very good (positive) one. then, annotators have to build the discourse structure of the document following the sdrt principles. our discourse annotation scheme was inspired from an already existing manual elaborated during the annodis project, a french corpus where each document was annotated according to the principles of sdrt. this manual gives a complete description of the semantics of each discourse relation along with a listing of possible discourse markers that could trigger any particular relation. however, the manual did not provide any details concerning the structural postulates of the underlying theory. this was justified, since one of the objectives of the annodis project was to test the intuitions of the naive annotators relevant to these issues. in casoar however, we aimed at testing the intuitions of naive annotators on how discourse interacts with opinion. we therefore modified the annodis manual in order to make precise all the constraints annotators should respect while building the discourse graph. in particular, we made explicit the constraints concerning segment attachment and accessibility of complex segments. we stipulated in the manual that each segment in the graph should be connected and that the attachment should normally follow the reading order of the document and the right frontier principle (cf. section 2.3). cdu constraints detailed how edus can be grouped to form complex units. figure 2 shows an example of a complex discourse unit constraint. suppose [1,2] and [2,3] are cdus. figures on the right and in the middle are correct configurations whereas the one on the left is not allowed for two main reasons: an edu cannot belong to two distinct cdus (as the edu 2 in the cdus [1,2] and [2,3]) and the head of a cdu23 cannot appear as a second argument of a relation. during the writing of this manual, we faced another decision: (1) should we annotate opinion texts using a small set of discourse relations or (2) should we use a larger set (i.e., the 19 relations already used in the annodis project). the first solution is more convenient and has already been investigated in previous studies. for example, in asher et al. (2008), we experimented with an annotation scheme where lexically-marked opinion expressions and the clauses involving these expressions are related to each other using five sdrt-like rhetorical relations: contrast and correction (introduced by signals such as: although, but, contradict, protest, deny, etc.), support that 23. the head of a cdu is the first edu that composes it. for example, 1 is the head of the cdu [1,2]. 12 evaluation in discourse figure 2: a cdu constraint. groups together explanation and elaboration, result (usually marked by so, as a result) which indicates that the second argument is a consequence or the result of the first argument, and finally, continuation. somasundaran (2010) proposed the notion of opinion frames as a representation of documents at the discourse level in order to improve sentence-based polarity classification and to recognize the overall stance. two sets of homemade relations were used: relations between targets (same and alternative relations) and relations between opinion expressions (reinforcing and non-reinforcing relations). finally, trnavac and taboada (2010) examined how some nonveridical markers and two types of rhetorical relations (conditional and concessive) contribute to the expression of appraisal in movie and book reviews. in our case, we chose not to use a predefined small set of rhetorical relations selected according to our intuitions because we did not know in advance what were the most frequent relations occuring in opinion texts and how this frequency was correlated with corpus genre. of course, this choice made it harder to do the annotations. but we think that this was a necessary step to investigate the real effects of discourse relations on both polarity and subjectivity as well as to evaluate the impact of discourse structure when assessing document overall opinion. among the set of 19 relations used in the annodis project, we focused our study on 17 relations that involve entities from the propositional content of the clauses24. these relations are grouped into coordinating relations (contrast, continuation, conditional, narration, alternative, goal, result, parallel, flashback) and subordinating relations (elaboration, e-elab, correction, frame, explanation, background, commentary, attribution). table 2 provides a detailed list of these relations along with their definitions. in this table, α and β stand respectively for the first and the second argument of a relation. (c) and (s ) represent respectively coordinating and subordinating relations. annotators were asked to link constituents (edus or cdus) through whichever discourse relation they felt appropriate, from our list above. in addition to this set of 17 relations, we also added the relation unknown in case annotators were not able to decide which relation is more appropriate to link two constituents. 4.1.2 the segment level for each edu in a document, annotators were asked to annotate its subjectivity orientation as well as its polarity and strength. subjectivity orientation. it can belong to five categories: 24. meta-talk (or pragmatic) relations that link the speech acts expressed in one unit and the semantic content of another unit that performs it were discarded. 13 benamara, asher, mathieu, popescu, chardon discourse relations definitions causality explanation (s) the main eventuality of β is understood as the cause of the eventuality in α goal (s) β describes the aim or the goal of the event described in α result (c) the main eventuality of αis understood to cause the eventuality given by β structural parallel (c) α and β have similar semantic structures. the relation requires α and β to share a common theme continuation (c) α and β elaborate or provide background to the same segment contrast (c) α and β have similar semantic structures, but contrasting themes or when one constituent negates a default consequence of the other logic conditional (c) α is a hypothesis and β is the consequence. it can be interpreted as: if α then β alternation (c) α and β are related by a disjunction reported speech attribution (s) relates a communicative agent stated in α and the content of a communicative act introduced in β exposition/narration background (s) β provides information about the surrounding state of affairs in which the eventuality mentioned in α occurs narration (c) α and β introduce an event and the main eventualities of α and β occur in sequence and have a common topic flashback (c) is equivalent to narration(β,α). the story is told in the opposite temporal order frame (s) α is a frame and β is on the scope of that frame elaboration elaboration (s) β provides further information (a subtype or part of) about the eventuality introduced in α entity-elaboration (s) β gives more details about an entity introduced in α commentary commentary (s) β provides an evaluation of the content associated with α correction correction (s) α and β have a common topic. β corrects the information given in the segment α table 2: sdrt relations in the casoar corpus. • se – segments that contain explicitly lexicalized subjective and evaluative expressions, [one of the best films i’ve ever seen in my life.] • si – segments that do not contain any explicit subjective cues but where opinions are inferred from context, as in [this is a definite choice to be in my dvd collection,] [and should be shared by fathers to their sons for generations.] 14 evaluation in discourse • o – segments that contain neither a lexicalized subjective term nor an implied opinion. they are purely factual, as in [i went to the cinema yesterday.] • sn – subjective, but non-evaluative segments used to introduce opinions. in general, these segments contain verbs used to report the speech and opinions of the author or others, as in the first segment in [i have no doubt][that this movie is excellent]. the opinion polarity (positive, negative, or neutral) is given by the verb complements. it is important to note that the sn category does not cover the cases of neutral opinion. • sei – that contain both explicit and implicit evaluations on the same topic or on different topics. for instance, [fantastic pub !]a [the pretty waitresses will not hesitate to drink with you]b, segment b contains two opinions, one explicit, towards the waitress, and the other one implicit, towards the pub. polarity. it can have five different values: positive, negative, neutral, both, and no polarity. neutral indicates that the positivity/negativity of the segment depends on the context, as in [this movie is poignant]. both means that the segment has a mixed polarity as in [this stupid president made a wonderful talk]. finally, no polarity concerns segments do not convey any evaluation (i.e., o and sn segments). strength. several types of scales have been used in sentiment analysis research, going from continuous scales (benamara et al. (2007)) to discrete ones (taboada et al. (2011)). in our case, we think that the chosen scale has to ensure a trade off between a fine-grained categorisation of subjectivity and the reliability of this categorization with respect to human judgments. for our annotation campaign, we chose a discrete 3-point scale, [1, 3] where 1 indicates a weak strength. objective segments (o) are associated by default to the strength 0. 4.1.3 the opinion expression level after segment annotation, the next step is to identify within each edu at least one of these elements: the opinion expression span, opinion topic, opinion holder, and operators that interact locally with opinion expressions. once all these elements are identified, annotators have to link every operator, topic and holder to its corresponding opinion expression using the scope relation. this relation aims to link: an operator to an opinion expression under its scope, a holder to its associated opinion expression, and an opinion expression to its related topic. since most opinion expressions reflect the writer’s point of views (i.e., the main holder), we decided not to annotate the scope relation in this case so as not to make the annotation more laborious. operators as well as topics are linked to the opinion in their scope only if several opinion expressions are present in an edu. we detail below the annotation scheme. opinion expression span. within each edu, annotators can identify zero (in case of si and o segments), one or several non overlapping opinion spans. an opinion span is composed of subjective tokens (adjectives, verbs, nouns, or adverbs), excluding operators25. its annotation includes: a polarity (positive, negative, and neutral), a strength (on a discrete 3-point scale, cf. above), a semantic category and a subcategory. according to the opinion categorization described in asher et al. (2008), each opinion expression can belong to four main categories: reporting which provides, at 25. operators are annotated separately. the idea is to capture both the prior and contextual polarity of opinion expressions. contextual polarity is annotated at the segment level while prior polarity at the expression level. 15 benamara, asher, mathieu, popescu, chardon least indirectly, a judgment by the author on the opinion expressed, judgment which contains normative evaluations of objects and actions, advice which describes an opinion on a course of action for the reader, and sentiment-appreciation containing feelings and appreciations. subcategories include, for example, inform, assert, evaluation, recommend, fear, astonishment, blame, etc. topics and holders. they are textual spans within a segment that are associated with a type. the opinion topic can have three types: main indicating the main topic of the document, such as “the movie”, part of in case of features related to the main topic, such as “the actors”, “the music”, and finally other when the topic has no ontological relation with the main topic, for example “theater” in the movie was great. shame that the theater was dirty. also, we distinguish between two types of holders: main that stands for the author’s review and other (as in my mother loved the movie). operators. finally, we deal with four types of operators: (i) negations that may affect the polarity and the strength of an expression, (ii) modals used to express the degree of belief of the holder, (iii) intensifiers used to strengthen (we use the operator int+) or weaken (int-) the prior polarity of a word or an expression, and (iv) restrictors that narrow the scope of the opinion in the sense that the positivity and/or negativity of the expression can be evaluated only under certain conditions, as in the restaurant is very good for children. operators have to be annotated when opinion expressions are under their scope as well as in case of implicit segments when appropriate. 4.1.4 a complete example figure 3 gives the annotation at the opinion expression and the segment level of the review (10), taken from emr. in this figure, we provide for each opinion expression its polarity and strength. similarly, we associate for each segment a triple that indicates its type (among: se, si, 0, sn, and sei), polarity (among: +, –, neutral, both, and no polarity), and strength (in a three level scale). figure 4 provides the associated discourse graph. (10) [i saw this movie on opening day.]1 [went in with mixed feelings,]2 [hoping it would be good,]3 [expecting a big let down]4 [(such as clash of the titans (2011), watchmen etc.).]5 [this movie was shockingly unique however.]6 [visuals, and characters were excellent.]7 4.2 annotation procedure 4.2.1 data preparation in order to avoid errors in determining the basic units (which would thus make the inter-annotator agreement study problematic), we decided to discard the segmentation from the annotation campaign. instead, edus were automatically identified. to train our segmenter, two annotators manually annotated a subset of fmr (henceforth fmr′) by consensus. this yields a total of 130 documents and 1,420 edus, among which 1.33% were embedded. automatic segmentation was carried out by adapting an already existing sdrt-like segmenter (afantenos et al. (2010)), built on the top of the annodis corpus26. the features used in afantenos et al. (2010) include the distance from sentence boundaries, the dependency path, and the chunk start/end. since we used a different syntactic parser, we modified certain features accordingly, and 26. the corpus used for training the parser was composed of 47 documents extracted from l’est rébublicain newspaper. this corpus is mainly objective and contains 1,400 edus, among them 10% were nested. 16 evaluation in discourse figure 3: annotation of (10) at the segment and the opinion expression level. figure 4: discourse annotation of (10). discarded others. we performed a two-level segmentation. first, we constructed a feature vector for each word token, which is classified into: right for words starting an edu, left for tokens ending an edu, nothing for words completely inside an edu, and both for tokens which constitute the only word of an edu. once all edus were found, subjective edus that contain at least one token belonging to our subjective lexicon27 are filtered out because they are good candidates for a further segmentation. the proportion of such edus in fmr′ was relatively small (around 12%). this second step was performed using symbolic rules which are mainly based on discourse connectives and punctuation marks. 27. our lexicon is manually built and is composed of 270 verbs, 632 adjectives, 296 nouns, 594 adverbs, 51 interjections. 17 benamara, asher, mathieu, popescu, chardon right left nothing r p f r p f r l n (e1) 0.858 0.913 0.885 0.872 0.894 0.883 0.976 0.967 0.972 (e2) 0.791 0.927 0.853 0.752 0.917 0.827 0.978 0.926 0.952 (e3) 0.925 0.942 0.933 0.941 0.952 0.946 0.982 0.977 0.980 table 3: evaluation of the classifier in terms of precision (p), recall (r), and f-measure (f). fmr′ fmr′lex r p f r p f boundaries 0.976 0.968 0.972 0.961 0.977 0.969 edu recognition 0.821 0.732 0.774 0.751 0.772 0.762 table 4: evaluation of the symbolic rules in terms of precision (p), recall (r) and f-measure (f). our discourse segmentation followed a mixed approach using both machine learning and rulebased methods. we first evaluated the classifier and then the symbolic rules. we performed a supervised learning using maximum entropy model28 in order to classify each token into right, le f t, nothing or both classes as described above. we conducted three evaluations: (e1) a 10-fold cross validation on the annodis corpus in order to compare our results to the ones obtained by afantenos et al. (2010); (e2) training on annodis and testing on fmr′ to see to what extent our set of features was independent of the corpus genre; (e3) a 10-fold cross validation on fmr′. table 3 shows our results for the right, le f t, and nothing boundaries, in terms of precision (p), recall (r), and f-measure (f). our results for the configuration (e1) are similar to those obtained by afantenos et al. (2010) on annodis. the best performance was achieved when training on our data (i.e., the configuration (e3)). table 4 shows the results of the symbolic rules when applied on the outputs of the configuration (e3). results concern both segment boundaries (averaged over all the four classes) and the recognition of an edu as a whole with a begin boundary and its corresponding end. we evaluated both on fmr′ when subjective edus are given by manual annotation and on fmr′lex when they are automatically identified using our lexicon. again, our rules performed very well. this tool was used to automatically segment fmr and fnr documents. the resulting segmentation was manually corrected when necessary29. we did not design an automatic segmenter for english and segmentation in emr was performed manually by two annotators by consensus. 4.2.2 annotation campaign we managed two annotation campaigns. the french one was the first and took six months. the english campaign came second and lasted three months. fmr and fnr was doubly annotated by three french native speakers while emr was annotated by two english native speakers. french annotators were undergraduate linguistic students while english ones were teachers. annotators 28. http://www.cs.utah.edu/˜hal/megam/ 29. we mainly corrected unbalanced bracketing. to this end, we designed a script that recognizes if for each begin bracket, there is a corresponding end bracket. if not, we manually ensured correct bracketing. we also checked if the other segmentation cases that we defined were correctly handled. overall, manual correction was very fast. 18 http://www.cs.utah.edu/~hal/megam/ evaluation in discourse benefited from a complete and revised annotation manual as well as an annotation guide explaining the inner workings of the glozz platform30, our annotation tool. since documents are already segmented, annotators first had to click on each edu, specified its category, polarity, and strength (see section 4.1.2), and then could isolate, within each edu, spans of text corresponding to the annotation scheme described in section 4.1.3. discourse annotation was performed by inserting relations between selected constituents using the mouse. when appropriate, edus were grouped to form cdus using glozz schemata. glozz also provides a discourse graph as part of its graphical user interface which helps the annotator to better capture the discourse structure while linking constituents. figure 5 illustrates how a document, extracted from emr, is annotated under glozz. the first segment includes the spans this and movie annotated as main topics, definitely and all time annotated as intensifier operators and the best annotated as an opinion expression. the annotation associated to the first segment is shown in the features structure on the right. segment 2 and 3 are related with a continuation relation, and the structure continuation(2, 3) is grouped into a cdu (the blue circle in the figure). figure 5: the annotation of an english movie review under the glozz platform. the french annotation proceeded in two stages. first, the annotation of the movie reviews; then, the annotation of news reactions. for each stage, we performed a two-step annotation where an intermediate analysis of agreement and disagreement between the three annotators was carried out. annotators were first trained on 12 movie reviews and then they were asked to annotate separately 168 documents from fmr. then, they were trained on 10 news reactions. afterwards, they continued to annotate separately 121 documents from fnr. the training phase for fmr was longer than for fnr since annotators had to learn about the annotation guide and the annotation tool. similarly, the english annotation campaign was done in two steps. annotators were trained on 10 emr and then the rest of the corpus (100 documents) was annotated separately. the time needed to annotate entirely one text was about 1 hour. 30. www.glozz.org 19 www.glozz.org benamara, asher, mathieu, popescu, chardon during training, we noticed that annotators often made the same errors. at the segment and the opinion expression level, these errors included: segments labelled as opinionated (se and sei) with no opinion expression inside; o or si segments with an opinion expression inside; o and sn segments with a prior polarity; opinion expressions with no associated semantic category, etc. for example, if one annotator considered the following segment i am a huge fan of tintin to be subjective, he should annotate the span fan as being an opinion expression. some of the discourselevel errors include: violation of the right frontier constraint, cycles, overlapping cdus, segments not attached to the discourse graph, etc. to ensure that the annotations were consistent with the instructions given in the manual, we designed a tool to automatically detect these errors. among all the provided annotations, 15% of the french documents contained errors at the segment and the opinion level vs. 12% for the english documents. the annotators were asked to correct their errors before continuing to annotate new documents. with respect to discourse structure, just a few french documents were ill-formed. however, the english annotators felt uncomfortable with discourse annotation, and their annotations were full of errors. we retrained them but finally decided to annotate discourse in emr by consensus. 4.3 reliability of the annotation scheme in this section, we report on inter-annotator agreements at the document, segment, and opinion expression levels. all statistics have been computed using the irr library under r31. 4.3.1 at the document level recall that the document annotation level consists of two tasks: assigning to each document an overall opinion (on a discrete five-level scale) and then a discourse structure. agreements have been computed on 152 fmr documents, 100 emr, and 120 fnr. agreements on overall opinion. we used two different measures. first, cohen’s kappa which assesses the amount of agreement between annotators. second, pearson’s correlation that measures the linear correlation between two vectors variables: the annotators’ overall opinions (variable 1) and the original overall opinions as given by allociné or metacritic users (variable 2). the aim is identify whether the first variable tends to be higher (or lower) for higher values of the other variable. pearson’s correlation gives a value between [−1,+1] where +1 indicates a total positive correlation, 0 no correlation, and 1 total negative correlations. table 5 gives our results in terms of cohen’s kappa when overall opinion has to be stated on the five level scale 0 to 4 (kappa multi-scale), the weighted kappa (weighted kappa multi-scale), and the kappa after collapsing the ratings 0 to 2 and 3 to 4 into respectively positive and negative ratings (kappa polarity). compared to a non weighted version, weighted kappa allows to compute agreements on ordinal labels. hence, a disagreement of 0 vs. 4 is much more significant that a disagreement of 1 vs. 2. we also give the average pearson’s correlation between the overall opinion given by our annotators and the overall ratings already associated to each movie review documents32. our results are good in movie reviews in polarity rating and weighted kappa but moderate in multi-scale rating, with a lower value obtained for news reactions. this shows that news reactions 31. https://cran.r-project.org/web/packages/irr/irr.pdf 32. correlations are given only for fmr and emr documents since in news reactions (fnr), authors are not asked to give the overall opinion of their comments. 20 https://cran.r-project.org/web/packages/irr/irr.pdf evaluation in discourse kappa multi-scale weighted kappa multi-scale kappa polarity pearson correlation fmr 0.48 0.66 0.68 0.83 emr 0.53 0.72 0.70 0.79 fnr 0.40 0.51 0.55 – table 5: inter-annotator agreements on document overall opinion rating. are more difficult to annotate. finally, when evaluating the correlation between the annotators’ overall opinions and the authors overall scores, we observe that correlations are good. agreements on discourse structure. as described in section 4.1, discourse annotation depends on two decisions: a decision about where to attach a given edu, and a decision on how to label the attachment link via discourse relations. two inter-annotator agreements have thus to be computed and the second one depends on the first because agreements on relations can be performed only on common links. for attachment, we obtained an f-measure of 69% for fmr and 68% for fnr assuming attaching is a yes/no decision on every edus pair, and that all decisions are independent, which of course underestimates the results. when commonly attached pairs are considered, we get a cohen’s kappa of 0.57 for the full set of 17 relations for fmr and 0.56 for fnr, which is moderate. here again, this kappa is computed without an accurate analysis of the equivalence between rhetorical structures33. figure 6 shows two discourse annotations for the french movie review in example (11). we observe that the annotator (on the left) formed more cdus than the other annotator (on the right) which causes both attachment and relation labeling errors. our goal being to study the effects of discourse on opinion analysis, a detailed analysis of inter-annotator attachment agreements is out of the scope of this study and is left for future work. figure 6: two discourse annotations of example (11). overall, our results are higher than those obtained by annodis (66% f-measure for attachment and a cohen’s kappa of 0.4 for relation labeling) mainly for two reasons. first, our annotation manual was more constrained since we provided annotators a detailed description of how to build 33. see (afantenos et al. (2012)) for an interesting discussion on the difficulty on how to compare rhetorical structures, especially when cdu are have to be taken into account. 21 benamara, asher, mathieu, popescu, chardon the discourse structure. second, our documents are smaller (an average of 20 edus compared to 55 edus in annodis) which implies less long distance attachments. (11) [bonne série.]1 [petits épisodes plus ou moins bien ficelés]2 [(mais n’est-ce pas le cas dans les autres aussi ?).]3 [le tout tenant en une 20e de minutes...]4 [rapide]5 [et sans temps morts.]6 [good tv series.]1 [small serials more or less well done]2 [(but it isn’t the case in the others too ?)]3 [all within 20 minutes time...]4 [fast]5 [and without time out.]6 4.3.2 at the segment/opinion level table 6 shows the inter-annotator agreements on segment opinion type, segment polarity and segment strength averaged over all the annotators. agreements have been computed on 1706 fmr segments, 1260 emr, and 1060 fnr. when computing these statistics for segment polarity, we have discarded the neutral category since we do have few instances of it in our data. in addition, since the both category means that the segment conveys at the same time negative and positive opinions, we decided to count it only once by conflating it with the positive category. similarly, we have also counted the sei class (which indicates that segments contain both implicit and explicit opinions) with se. fmr emr fnr kappa on segment opinion type 0.66 0.60 0.50 kappa on segment polarity 0.76 0.71 0.48 kappa on segment strength 0.35 0.27 0.27 weighted kappa on segment strength 0.49 0.43 0.34 table 6: inter-annotator agreements on segment opinion type, polarity, and strength per corpus genre. we observe that the inter-annotators agreements are better for movie reviews than for news reactions and that fmr achieves the best scores. we get very good kappa measures for both explicit opinion segments se (0.74) and the polarity (positive and negative) of a segment in french movie reviews (respectively 0.78 and 0.77). we get similar results in english with as an example a kappa of 0.67 for the se class and a kappa of 0.75 and 0.74 for respectively positive and negative segment opinion type. these results are in agreement with state-of-the-art results obtained in contemporary annotation campaigns (see e.g. wiebe et al. (2005)). the kappa for the sn class is also very good: 0.74 in fmr and 0.64 in emr. finally, the agreements for the si and o classes were respectively 0.56 and 0.63 in fmr, and 0.52 and 0.58 in emr. they are moderate because annotators often fail to decide whether a segment is purely objective and thus if it conveys only facts or if a segment expresses an implicit opinion. here are two examples illustrating annotators disagreement on segment opinion type: (12) [as mentioned elsewhere,]1 [the romance in the movie was painful]2 [but helped tie things up at the end.]3 [good way to burn 2h of your life and 15$]4. 22 evaluation in discourse (13) [the production company had any idea how to market this film.]1 [the trailer looks like a non-stop action thriller set in a train station,]2 [when in fact it is far slower]3 [but wellpaced,]4 [and its best moments come away from the station.]5 in (12), one annotator (a) considered that segments 3 and 4 conveyed positive implicit opinions towards the movie while the second annotator (b) has labeled these segments as explicit by selecting the spans good way and tie things up as being positive opinion expressions. in (13), (a) and (b) agreed to put the segments 4 and 5 into the se category but disagreed on the category of the first three segments: for (a), segments 1 and 2 are implicit negative segments whereas for (b) they are purely objective. similarly, for (b) segment 3 is objective and for (a) it is an explicit opinion because it contains the word slower which has been annotated as a negative opinion expression. the difficulty to discriminate between explicit, implicit, and objective segments can also be explained by the lower kappa measure obtained for no polarity with 0.60 in emr and 0.68 in fmr compared to the kappa obtained on positive and negative segment polarity. this difficulty is, we believe, an artifact of the length of the texts. indeed, the longer a text is, the greater the difficulty for human subjects to detect discourse context. however, the study of this hypothesis falls out of the scope of this paper and is therefore left for future work. nonetheless, these results are good in the range of state-of-the-art research reports in distinguishing between explicit and implicit opinions. for instance, toprak et al. (2010) obtained a kappa of 0.56 for polar fact sentences which are close to our si category. in fnr, our results were moderate for the se and sn classes (respectively 0.56 and 0.58) and weak for the si and o classes (respectively 0.48 and 0.40). we have the same observations for the agreements on segment polarities where we obtain moderate kappas on all the three classes (positive, negative, and no polarity). this shows that the newspaper reactions were more difficult to annotate because the main topic is more difficult to determine (even by the annotators) – it can be one of the subjects of the article, the article itself, its author(s), a previous comment or even a different topic, related to various degrees to the subject of the article. implicit opinions, very frequent, can be of a different nature: ironic statements, jokes, anecdotes, cultural references, suggestions, hopes and personal stances, especially for political articles. here is an example of implicit segments extracted from fnr. annotators disagreed on how to annotate the first segment: for (a), 1 is negative implicit while for (b) it is explicit (with the spans vraiment/really and plaindre/pity annotated respectively as an operator and an opinion expression): (14) [les enseignants sont-ils vraiment à plaindre ?]1 [avec 6 mois de vacances par an]2 [et la possibilité de prendre une retraite à 45 ans dans certains cas...]3 [are teachers really to be pitied ?]1 [with 6 months vacation per year]2 [and the opportunity to retire at 45...]3 finally, the kappa for segment strength averaged over the scale [0, 3] is bad. however, the kappas are good on the extreme values of this scale, and moderate when using a weighted measure. for example, we get a kappa of 0.67 and 0.58 in respectively fmr and emr on the strength 0 vs. 0.4 in fnr. these results confirm that multi-scale polarity annotation is a difficult task, as already observed in similar annotation schema (cf. toprak et al. (2010)). we think that low agreements were mainly due to the annotation manual that failed to clearly explain strength annotation. indeed, for the same “basic” opinion expression, we got different annotations. for example, in similar contexts, the adjective good got different scores (+1 or +2). we think that the manual can be 23 benamara, asher, mathieu, popescu, chardon improved by explicitly stating the prior score of “basic” expressions (e.g., good (+1), brilliant (+2) and exceptional (+3)) and then asking annotators to score new expressions by comparing their strength to these expressions. 5. results we give now the results of the annotation campaign focusing on quantitative results on each annotation level, and more importantly on the impact of discourse on sentiment analysis. 5.1 quantitative analysis at the document level our discourse annotations contain a total of 3,453 discourse relations for fmr, 1,740 for fnr and 1,677 relations for emr. we analyzed our results according to two main axis: the distribution of relations per corpus genre and the importance of cdus for sentiment analysis. 5.1.1 distribution of discourse relations per corpus genre figure 7 shows these distributions, sorted according to their frequency in fmr, from the most frequent (on the left) to the less frequent one (on the right). the frequencies of each discourse relation across corpus genres are statistically different from what would be expected by chance using the χ2 test. note however that the difference between the observed and the expected frequencies of conditional were not statistically significant. in this figure, we discarded the frequencies of the relations flashback and unknown for two reasons. first, flashback was highly infrequent in all the corpora (0.12%, 0.06% and 0% for respectively fmr, fnr, and emr) and second, the relation unknown was not used in emr since the discourse annotation in this corpus has been performed by consensus. it is however interesting to note that this relation was more frequent in fmr (around 2.06%) than in fnr (0.69%) mainly because the annotators were more experienced with respect to the “reviews” corpus (annotated first). overall, the frequencies can be grouped into three classes: (1) continuation, elaboration and commentary (more than 10%), (2) contrast, entity-elaboration, result, explanation, attribution and frame (from 3% to 10%) and (3) correction, goal, narration, parallel, background, conditional and alternation (less than 3%). we noticed that some relations are more present in certain corpora. for instance, commentary, entity-elaboration, explanation, attribution, frame, goal, parallel and alternative are more frequent in news reactions than in reviews. the frequencies of parallel, alternative and frame are consistent with a logically more structured discourse for news reactions than for movie reviews. also goal and explanation are more frequent which confirms that fnr contains more argumentative structures than in reviews. the same goes for the attribution relation, which denotes that in fnr people tend to make reference to what other people said., e.g. the president thinks that..., or even that people tend to be more reserved when stating opinions, e.g. i guess that this is a good measure, unlike in the reviews, where people might tend to be more categorical, e.g. this movie is great, without modalizing the statement. also, entity-elaboration is more frequent in fnr (more than 10%), which confirms that news reactions are multi-topic opinion documents. another interesting comparison between corpus genres is the frequency of commentary, more frequent in news reactions where commentaries are often ironic. finally, the proportions of elaboration, contrast, background, narration and result in the en24 evaluation in discourse figure 7: the distribution of discourse relations per corpus genre. glish corpus were higher compared to the two other corpora, may be because english reviews tend to be more verbose. 5.1.2 importance of cdus we have also analyzed the ratio of complex segments to the total number of rhetorical relation arguments in our annotations. figures 8, 9, 10 show the proportions of relations between edus, between an edu and a cdu, and between cdus, sorted according to the increasing frequencies of relations between edus (all the relations are shown except unknown and flashback). first, we see that some relations are local and tend to appear more often between edus (more than 70%), as in example (15) taken from emr. in news reactions, these local relations have the same distributions except for attribution and conditional which link simple segments in 60% of cases. this is more salient for background with only 45% of instances. we will see in section 5.3 that some of these local relations are very important for sentiment analysis while others can simply be ignored. background and commentary have different behaviors in english reviews compared to french documents: background seems to be more local in french documents whereas commentary tends to be more local in english reviews. on the other hand, the following relations often have cdus in at least one of their arguments: elaboration, explanation, frame, result, contrast, correction, narration and commentary. for example, correction concerns cdus in most of 55% of cases. this relation links segments sharing a common topic and such that the second argument corrects 25 benamara, asher, mathieu, popescu, chardon the information given in the first argument (which is often at a long distance attachment) (see the correction in example (16)). another interesting behavior comes from the contrast relation. contrary to our expectations, only 40% of instances of this relation link edus in all the corpora. example (17) illustrates a contrast with scope over two cdus. (15) [one of the worst movies ever !]1 [it’s just terrible !]2 explanation(1,2) (16) [the day before,]1 [i went to see this movie,]2 [i thought]3 [i knew]4 [what awesome was,]5 [but i was so wrong.]6 frame(1,2) background([1,2],[3,4,5,6]) continuation(3,4) attribution([3,4],5) correction([3,4,5],6) (17) [the dialogue is stodgy]1 [and the drama slows the pace,]2 [but the violent action]3 [and the imaginative look make it fun to watch.]4 continuation(1,2) continuation(3,4) contrast([1,2],[3,4]) 5.2 quantitative analysis at the segment and opinion expression level the total number of annotated segments was 3,825 for fmr, 2,071 for fnr and 2,578 for emr. the histogram in figure 11 gives a comparative analysis of how segments are distributed over the five classes (i.e., se (explicit opinion), si (implicit opinion), o (objective), sn (subjective non evaluative) and sei (explicit and implicit segment)). a similar analysis is given in figure 12, this time for segment polarity (i.e., positive, negative, neutral, no polarity and both). the frequencies of each segment opinion type and each segment polarity type across corpus genres are statistically different from what is expected by chance using the χ2 test. we observed that the frequencies of the segments containing implicit opinions (si) depend on the corpus genre: for fmr and emr, frequencies are less important (respectively 26.5% and 24.5%) compared to fnr (47.1%). moreover, in the three corpora, the purely objective segments are not very widespread (less than 20% of all segments). the same goes for segments that contain at the same time an explicit and an implicit opinion (sei), with a yet lower frequency for enr. as for the subjective non-evaluative segments (sn), they are rather infrequent as well, especially in french and english movie reviews. however, they are slightly more numerous for fnr, which shows that the reported speech constructions are more frequent in reactions to newspaper articles than in movie reviews. another interesting genre bias concerns the polarity of the segments: whereas in french movie reviews positive segments are a majority in spite of balancing the corpus between overall positive and overall negative documents (in terms of their star counting), this is not the case for the reactions to newspaper articles, where negative segments are a majority. in emr however, segment polarity distribution is more balanced than for fmr. we also observe that non evaluative segments (mainly from the objective and the subjective non evaluative segment type) are more numerous in 26 evaluation in discourse figure 8: the distributions of discourse relations in fmr according to the type of their arguments. english reviews than in french reviews. finally, the proportion of both and neutral are a minority in all the corpora (respectively less than 3% and 2%). the last segments in examples (18) and (19) respectively illustrate segments from the both and neutral category. (18) [i am very torn about this film,] [as i think] [it contains some really bad directing by a great director.] (19) [as some one commented already] [it is a combo of ”black beauty” and ”all quiet on the western front.”] within evaluative segments (i.e., se, sei and sn), 2,329 opinion expressions were annotated for fmr, 743 for fnr and 1,610 for emr. among explicit segments (i.e., se and sei), 97% contain a single opinion expression for fmr and emr vs. 94% for fnr. this confirms the usefulness of the per-segment analysis since this simplifies opinion fusion with respect to a per-sentence analysis for instance. we further discuss this important result in section 6. the semantic categories of opinion expressions are similarly distributed for fmr and emr with around 3% for advice, and between 5 and 8% for reporting. however, we observe that in english movie reviews, most opinion expressions are from the sentiment-appreciation category (48.2% vs. 24.2% for french) while, in fmr, opinion expressions are mostly judgments and evaluations (66.4% vs. 36.4% for english). as expected, we get different distributions of semantic categories for fnr, with a greater number of reporting (27.5%) and advice expressions (6.9%) and no instances of the sentiment-appreciation category. 27 benamara, asher, mathieu, popescu, chardon figure 9: the distributions of discourse relations in fnr according to the type of their arguments. concerning the annotations of topics and holders, the total number was respectively: 2,939 and 754 for fmr, 1,915 and 262 for fnr, and 1,981 and 499 for emr. for movie reviews, topics are mainly from the part of category (around 60%) whereas few of them are out of topic (other) (around 10%). however in fnr, we observe a different distribution: the number of topics from the main category are lower (around 9%) whereas the number of other topic are greater (around 19.4%). for the holders, we get similar distributions over all the corpora: 2/3 of annotated holders are from the main category. lastly, we also noticed the importance of opinion operators: 1,371 for fmr, 924 for emr and 488 for fnr. at least one such operator is present in 32% of subjective segments in news reactions vs. 40% for movie reviews. these operators are also present in implicit segments (18% for the french corpus vs. 25% for the english documents and 17% for news reactions) which indicates that valence shifter terms are good cues for detecting implicit opinions. the distribution of operators per category is shown in figure 13. most of them are intensifiers. restrictors are from different types: they can be temporal (as some in [some scenes are beautifully shot] and at times in [it can be entertaining at times]) or topic restrictions as in [this movie is made for 10 year old kids.]. in our previous work on using discourse in sentiment analysis, we have annotated opinion semantic categories at the segment level in movie reviews and letters to the editor in english and french. our past results, reported in (asher et al. (2008)), showed that the distribution of semantic categories in these corpora are comparable to those observed in the corpora annotated in this current 28 evaluation in discourse figure 10: the distributions of discourse relations in emr according to the type of their arguments. study. as far as the semantic categories are concerned, we can conclude that our observations are valid to french and english movie reviews and news reactions in general. we believe that our results on segment polarity and segment type can also be generalized. more annotations are however needed to validate this assertion. 5.3 impact of discourse on sentiment analysis in this section, we attempt to answer the challenges mentioned in the introduction of this paper: what is the role of discourse relations in subjectivity analysis? what is the impact of the discourse structure in determining the overall opinion conveyed by a document? does a discourse based approach really bring additional value compared to a classical bag of words approach? does this additional value depend on corpus genre? to this end, we explored the interactions between the discourse, the segment, and the opinion expression annotation layer. in particular, • section 5.3.1 investigates the correlation between discourse and opinion semantic category of subjective segments (mainly from the se, sei and the sn category). recall that an opinion expression can belong to four semantic categories, namely: sentiment-appreciation, judgment, advice and reporting. our aim is to analyze to what extent semantic categories of opinion expressions can be an indicator for predicting discourse relations. 29 benamara, asher, mathieu, popescu, chardon figure 11: frequencies of segments per opinion type. • section 5.3.2 focuses on the impact of discourse on subjectivity analysis. can discourse relations be used to predict subjectivity orientation of elementary discourse units? • section 5.3.3 analyzes the impact of discourse on polarity analysis. can discourse relations be used to predict polarity of elementary discourse units? • section 5.3.4 studies the impact of segment opinion type and segment polarity on the determination of the document overall opinion. do segments with implicit opinions contribute to the author’s global opinion on the main topic of the document? this section details experiment aspect addressing each of these challenges while section 6 summarizes the conclusions answering these questions. 5.3.1 discourse and opinion semantic categories we tested two hypotheses: (h1) there is an association between the relative position of segments within the document and the semantic category of the opinion expressions they contain. if a correlation is found, then the position can be used for example to identify the semantic category of segments conveying implicit opinions. (h2) there is an association between discourse relations and the semantic categories of the opinion expressions that appear within the relation arguments. 30 evaluation in discourse figure 12: frequencies of segments per polarity type. position of segments vs. semantic categories. table 7 gives the proportions (in percent) of opinion semantic categories according to the relative position of the segment they belong to. we considered two positions: beginning and end of the document. to compute them, we simply divided a document into 3 parts (beginning, middle, end). the first two segments being the beginning while the last two the end. in the table 7, the configurations begin-x (resp. end-x) stand for segments containing an opinion expression from an x category. when using the χ2 test, the hypothesis (h1) is confirmed at p < 0.05. we see that the proportion of the advice category is higher when expressions of this type appear in segments at the end of the document. the proportion of the other categories is relatively stable. this increase is more impressive in reviews (more than 10%) than in news reaction (around 5%) which confirms that users in reviews tend to end their reviews by expressions of recommendations, hopes, or suggestions. discourse relations vs. semantic categories of their arguments. for each corpus, we constructed three contingency tables: • (t1) gives the number of discourse relations that have a right argument containing an opinion expression from a given semantic category. for each discourse relation r and for each semantic category c ∈ {s entiment − appreciation, judgment, advice,reporting}, we counted all the pairs r(se c, all) where se c is an se segment containing an opinion expression from a category c and all stands for an edu whatever its type (i.e., se, sei, o, sn or si). 31 benamara, asher, mathieu, popescu, chardon figure 13: the distribution of operator per category. emr fmr fnr begin-reporting 0.10 0.06 0.32 begin-judgment 0.46 0.64 0.60 begin-sentiment-appreciation 0.43 0.29 0.00 begin-advice 0.01 0.01 0.08 end-reporting 0.09 0.02 0.22 end-judgment 0.38 0.58 0.65 end-sentiment-appreciation 0.43 0.26 0.00 end-advice 0.10 0.14 0.13 table 7: proportions (in percent) of opinion semantic categories according to the relative position of the segment they belong to. • in table (t2), we do the same by counting all the pairs r(all, se c). • table (t3) provides the frequencies for each relation r and the frequencies of r(se c, se c). tables 8, 9, and 10 give respectively the results of (t1), (t2) and (t3) for the french movie reviews corpus. the tables associated to the other two corpora looked similar. 32 evaluation in discourse advice right sentiapp right reporting right judgment right elaboration 3 49 13 170 attribution 5 16 15 27 goal 0 3 0 7 continuation 10 164 16 511 frame 1 7 9 13 conditional 2 4 0 2 e-elab 2 35 9 89 parallel 1 5 0 13 explanation 2 44 14 134 result 16 82 18 102 background 1 4 1 7 narration 0 5 1 9 commentary 24 95 21 175 alternative 2 2 2 4 correction 1 4 3 17 contrast 5 47 14 136 table 8: frequency of discourse relations that have a right argument containing an opinion expression from a given semantic category. advice left sentiapp left reporting left judgment left elaboration 5 78 18 184 attribution 9 13 47 15 goal 0 5 0 9 continuation 12 155 26 509 frame 0 2 1 7 conditional 0 3 0 1 e-elab 4 34 8 99 parallel 0 6 2 15 explanation 0 46 26 110 result 20 64 8 138 background 0 6 2 0 narration 0 5 2 12 commentary 4 76 14 193 alternative 2 3 0 7 correction 2 7 1 23 contrast 6 41 15 153 table 9: frequency of discourse relations that have a left argument containing an opinion expression from a given semantic category. given the frequencies in these tables, the hypothesis (h2) was rejected using the χ2 test. for each corpus genre, there is no statistically significant relationship between discourse relations and the opinion category of their arguments. however, in the french corpora, after removing the relations goal, conditional, frame, background and attribution from the contingency table (t1), the 33 benamara, asher, mathieu, popescu, chardon advice same sentiapp same reporting same judgment same elaboration 0 3 0 10 attribution 1 0 0 0 goal 0 0 0 2 continuation 4 25 3 134 frame 0 0 0 1 conditional 0 1 0 0 e-elab 0 6 0 25 parallel 0 3 0 7 explanation 0 6 6 34 result 0 6 0 16 background 0 0 0 0 narration 0 1 0 4 commentary 0 7 0 17 alternative 0 1 0 3 correction 0 0 0 4 contrast 0 6 0 47 table 10: frequency of discourse relations that have arguments containing opinion expressions from the same semantic category. association between discourse relations and opinion category of right arguments was significant at p < 0.05 using the χ2 test34. for emr, the association is significant when removing the same set of relations as above and when discarding, in addition, the categories advice and reporting. in (t2), the association between discourse relations and left arguments was significant when removing the advice category and the same set of relations as above except attribution. finally, for (t3), we get a statistically significant association when removing both the same set of relations as above and the categories advice and reporting. overall, the absence of a strong correlation between discourse relations and opinion categories can be due to the categories themselves that were not adequate to capture that relations well. to confirm or reject hypothesis (h2), it would be interesting to conduct a similar study using different categories. concerning the distribution of relations with regard to the opinion semantic category, the proportion of attribution relations is relatively high when the first argument of this relation is from the reporting category. we also have instances from continuation and elaboration. similarly, the proportion of result is high when its second argument contains an advice expression. examples like (20) are very frequent in our reviews corpora (here segments 4 and 5 contain explicit recommendations to see the movie and they are related to the first part of the document by a result relation): 34. note that the χ2 test cannot be computed if some frequencies are less than 5. to overcome this problem, some relations that have similar semantic effects on opinion were grouped, like contrast with correction, continuation, parallel with alternative, etc. 34 evaluation in discourse (20) [it it is the best adventure movie of our time]1 [and whats in bonus its an awesome joy-full adventure for all ages.]2 [its a full family entertainer.]3 [so go]4 [and watch the movie]5 [and uncover the secret of hugo cabret.]6 on the other hand, several advice expressions in the emr corpus are related with conditional, like in (21) where the author recommends the movie under certain conditions: (21) [if you’re after a film]1 [that doesn’t employ too much thinking]2 [and is enjoyable to watch]3 [i would recommend going to see this.]4 finally, advice can also come under a commentary, as in the news reaction in (22): (22) [quand on en est à emprunter de l’argent frais]1 [pour payer les intérêts des emprunts précédents,]2 [c’est que le mur se rapproche.]3 [un conseil :]4 [achetez de l’or...]5 [when we are borrowing money]1 [to pay past loan interest,]2 [it means that the wall is approaching.]3 [an advice:]4 [buy gold...]5 5.3.2 discourse relations and subjectivity analysis we also assessed, for each relation instance linking edus only (except for unknown and flashback since they have the lowest frequencies), whether they preserve subjectivity or not. we computed statistics on the stability of the subjectivity class (for the (se, se), (si, si) and (si, se) pairs) or the variation of the stability class (for the (o, other) pairs, where “other” spans the set of subjectivity classes, other than o). figures 14, 15 and 16 summarize our results. the relations in these figures are sorted according to the decreasing frequencies of subjectivity preservation. the subjectivity preservation frequencies of each discourse relation across corpus genres are statistically different from what is expected by chance using the χ2 test. note however that the difference between the observed and the expected frequencies of attribution and conditional were not significant. we observe that our predictions (as stated in the introduction) are by and large confirmed. some relations preserve subjectivity in all corpora (with more than 70% of instances): continuation, parallel, alternative, contrast, elaboration, explanation, commentary, result, narration, and contrast. for some relations, the preservation is more salient for reviews than for news reaction. for instance, commentary preserves subjectivity in 80% of cases in fnr vs. between 60 and 72% for reviews where examples like (23) are less frequent. result however gets a different distribution, with more than 80% preservation in reviews vs. 70% in reactions. (23) [j’ai découvert la vie de piaf,]1 [on a l’impression d’être avec elle tout le long du film.]2 [i discovered piaf’s life,]1 [i felt that i was with her all along the movie]2 other relations do not preserve subjectivity across our corpora: background, attribution and frame. in news reaction, attribution preserves subjectivity in 50% of cases whereas in reviews the proportion is about 20%. this might be because examples like [the chairman thought] [that it rained in his town yesterday] are more frequent in the first corpus genre (movie reviews) than in the second (news reactions) where attributions are more often used to introduce opinions and point of views. subjectivity preservation in the case of frame is about 40% in french document vs. 87% in english reviews because in french corpora, this relation often relates non evaluative segments to evaluative ones. correction seems to preserve subjectivity in reviews (60% in english 35 benamara, asher, mathieu, popescu, chardon figure 14: discourse relations and subjectivity in emr. reviews and 83% in french reviews) but not in news reactions where the proportion is about 50%. we observe the contrary for conditional and entity-elaboration where subjectivity preservation is more frequent in news reactions. indeed, in fmr, consequences are often objective even when their corresponding conditions are evaluative as shown in (24). (24) [c’est long,]1 [froid,]2 [pas bon.]3 [si vous y allez une fois]4 [ce sera bien la seule]5. [it’s long,]1 [cold,]2 [not good.]3 [if you go once]4 [it will be the only time]5. continuation(1,2) continuation(2,3) result([1,2,3],[4,5]) conditional(4,5) 5.3.3 discourse relations and polarity analysis we finally computed similar statistics for the polarities, but between subjective (sn, se, sei, si) edus only: the (+, +) and (–, –) for stability and (+, –) for polarity change. we assess in figures 17, 18 and 19 the behavior of our relations with respect to polarity preservation and non-preservation. only relations preserving subjectivity are taken into account (background, attribution, frame and entity-elaboration have been discarded35). they are presented by decreasing order of polarity 35. relations that do not preserve subjectivity are necessarily relations that do not preserve polarity. 36 evaluation in discourse figure 15: discourse relations and subjectivity in fmr. preservation frequencies. the polarity preservation frequencies of each discourse relation across corpus genres are statistically different from what is expected by chance using the χ2 test. note however that the difference between the observed and the expected frequencies of conditional, correction and contrast were not significant. as far as polarity is concerned, our hypotheses seem by and large verified as well, for all corpora. however, contrary to expectations, contrast seems to change polarity in reviews but not in news reactions. in reactions, this can be explained by examples of the type: [the economical situation is grim,] [but the cultural life is grim as well] where there is the but connective linking the two segments, which makes the annotators place a contrast between the two segments. however, in this particular case it would be more appropriate to link the two segments by the parallel relation or with both parallel and contrast36, which is possible in sdrt and provides the right semantics for such relations (asher (1993)). note however that the frequencies of contrast and correction in all the corpora were not significant. we need more annotations to establish the relationships between these relations and polarity analysis. 5.3.4 segment type, segment polarity, and overall opinion we investigated whether implicit opinion segments contribute to the author’s global opinion on the main topic of the document. we have computed the pearson’s correlation between the global 36. when preparing the gold standard, we reconsidered the relation labels only in 5% of the cases. 37 benamara, asher, mathieu, popescu, chardon figure 16: discourse relations and subjectivity in fnr. opinion score (on a scale going from 0 for a strongly negative opinion, to 4 for a strongly positive opinion) and the subjectivity class and polarity of the segments. more specifically, for each of the three corpora, we have constructed a vector with the global opinion scores for all the annotated document instances37. then, another set of four vectors has been built for each corpus, with the counts of segments of a given subjectivity class and polarity: se pos for explicit positive opinion segments (se and sei) class with a positive polarity; se neg for explicit negative opinion segments; si pos for implicit positive opinion segments (si class with positive polarity); and si neg for implicit negative opinion segments. similarly, we have computed the correlation between the overall opinion and segment polarity regardless of their types: all pos for positive segments and all neg for negative segments. in addition, we have measured the correlation between the overall opinion vector and the average segments scores (given between −3 and +3) of each document (all avg). the results are shown in table 11, averaged over all the annotators. in movie reviews (fmr and emr) there is a better correlation between global opinion score and explicit subjective segment counts (of both positive and negative polarities – for negative polarities, a good correlation means a negative pearson’s correlation of high absolute value) than between global opinion score and implicit subjective segment counts. in fnr, a different behavior is observed: the correlation is better for segments which contain implicit opinions. this brings us to the conclusion that the importance of implicit opinions varies, depending on the corpus genre: in 37. if one input document has been doubly annotated, we thus obtained two document annotation instances. 38 evaluation in discourse figure 17: discourse relations and polarity in emr. fmr emr fnr se pos 0.54 0.72 0.3 se neg -0.64 -0.74 -0.19 si pos 0.42 0.59 0.42 si neg -0.45 -0.67 -0.40 all pos 0.64 0.61 0.52 all neg -0.6 -0.57 -0.38 avg all 0.19 0.17 0.10 table 11: correlations between overall opinion and segment opinion type/polarity. movie review, more direct and sometimes terse, explicit opinions are better correlated to the global opinion score, whereas in news reactions, implicit opinions are more important. this could indicate a tendency to “conceal” negative opinions as apparently objective statements, which can be related to social conventions (politeness, in particular) (pang and lee (2008)). now, when we have grouped segments by polarity (cf. all pos and all neg), we observe that the correlation with positive segments are better compared to those with negative polarity. the politeness bias is more salient in news reactions than in movie reviews where users tend to express their opinions in a more positive way. finally, we see that correlations in all avg are the lowest, which confirms that overall opinions is not only a simple aggregation of opinions taken in isolation. a more elaborated way of aggregation is needed. 39 benamara, asher, mathieu, popescu, chardon figure 18: discourse relations and polarity in fmr. 6. discussions 6.1 interim conclusions in this paper, we aimed at measuring the impact of discourse on sentiment analysis with a study of three corpora: french and english movie reviews as well as french news reactions. here are the main conclusions of our corpus-based study: (a) segment-based opinion analysis is more appropriate to study opinions in discourse. our results showed that more than 90% of segments contain only one opinion expression. this demonstrates that the segment level will make polarity analysis easier compared to the sentence or the clause level. in addition, our automatic discourse segmentation is feasible and yielded very good results. (b) complex discourse units (cdus) are an important part of the discourse structure of a document. in the whole corpora, our results showed that the proportion of relations involving cdus is higher compared to the proportion of relations linking edus. in particular, we observed that the arguments of the relations contrast, elaboration, and result are cdus in more than 55% of cases. cdus related with a frame are more frequent in movie reviews (more than 57%) whereas those related with a commentary are more frequent in the french corpora (more than 64%). these results demonstrate 40 evaluation in discourse figure 19: discourse relations and polarity in fnr. that cdus are important for assessing the overall opinion of a document. (c) implicit opinions are important. our results showed that the importance of implicit opinions varies, depending on the corpus genre: for movie reviews, explicit opinions are better correlated to the global opinion score, whereas for news reactions, implicit opinions are more important when negative opinions are concerned. (d) semantic categories of opinion expressions can be good indicators for identifying some discourse relations. indeed, we observed that the discourse relations contrast, continuation, narration, alternative, result, parallel, elaboration, entity-elab, correction, explanation, and commentary are correlated with the semantic categories (reporting, judgment, advice, and sentimentappreciation) of the opinion expression within their arguments. (e) discourse relations can be grouped according to their effects on the opinion orientation of elementary discourse units. we studied 17 discourse relations that involve entities from the propositional content of the clauses: 9 coordinating relations (contrast, continuation, conditional, narration, alternative, goal, result, parallel, flashback) and 8 subordinating relations (elaboration, e-elab, correction, frame, explanation, background, commentary, attribution. among these relations, some can be grouped according to their similar effects on both subjectivity and po41 benamara, asher, mathieu, popescu, chardon larity analysis: correction and contrast, elaboration and explanation, continuation, parallel, narration, and alternative. table 12 summarizes the effects of these relations. for a given relation (or group of relations), “ √ ” (resp. “x”) indicates that the relation preserves (resp. does not preserve) subjectivity (resp. polarity) in more than 75% of cases in at least two corpora. this table shows that some relations have no effect at all on sentiment analysis: frame, goal, background, conditional and flashback while others impact on subjectivity analysis, on polarity analysis or influence both these two tasks. these results confirm that discourse relations can help in identifying segments conveying implicit opinions or retrieving segment contextual polarity which, for instance can be very useful in identifying ironic statements. discourse relation frequency in at least two corpora impact on sentiment analysis subjectivity analysis polarity analysis parallel ≤ 5% and ≥ 1% √ √alternative ≤ 5% and ≥ 1% continuation ≥ 5% narration ≤ 5% and ≥ 1% explanation ≤ 5% and ≥ 1% √ √ elaboration ≥ 5% commentary ≥ 5% x √ contrast ≥ 5% √ √ correction ≤ 5% and ≥ 1% e-elab ≥ 5% x √ result ≥ 5% √ √ attribution ≤ 5% and ≥ 1% √ x frame ≤ 5% and ≥ 1% x x goal ≤ 5% and ≥ 1% x x background ≤ 5% and ≥ 1% x x conditional ≤ 5% and ≥ 1% x x flashback ≤ 1% x x table 12: discourse relations and sentiment analysis: interim conclusions. 6.2 portability of the annotation scheme the results reported in this study were obtained on manually annotated discourse structures when the annotation scheme was instantiated on two corpus genres: movie/product reviews and news reactions. these corpora have similar characteristics: they are texts and not discussions/dialogues (remember that letters to the editor that responded to other letters were removed from fnr), they are relatively small (less than 30 edus per document), opinions are about one main topic and its related subtopics and are the viewpoints of one holder (mainly the author of the review). more important, the overall opinion is the result of a bottom-up aggregation process, from local opinions at the segment level to the global opinion at the document level. however, several other corpus genres do not meet these characteristics. some are author-oriented like blogs where all the documents (posts and comments) are associated to the blogs’ owners, others are both multi-topic and multi-holder documents like news articles, while others are composed of follow-up opinions as in discussion forums. to what extent is the casoar annotation scheme portable to these other sources of opinion? 42 evaluation in discourse concerning blogs, we believe that our scheme can be easily applied. blog comments are generally short, they are the point of view of one author towards the main topic of the blog article which is quite similar to news reactions. for news documents, things are more complicated since several viewpoints by several opinion holders are mentioned. consider the following scenario. the author introduces and elaborates on a topic, ‘switches’ to other topics or reverts back to an older topic. this is known as discourse popping where a change of topic is signaled by the fact that the new information does not attach to the prior clause, but rather to an earlier one that dominates it (asher and lascarides (2003)). in this case, our three-level annotation scheme needs to be adapted. though the discourse annotation model incorporates discourse pops, their effects on topics for opinions is presently not taken into account. discourse pops often indicate shifts in topic, and so, instead of one topic, we will have to deal with many. at the expression level, we have to take this multi-topicality into account, by modifying the annotation of topic spans. at the segment level, we would have to link each opinion expression to its topic. at the document level, the notion of overall opinion has to evolve towards (topic, holder) overall opinion scores. each score can be computed using a bottomup aggregation procedure over a discourse sub-graph focusing only on those segments that convey the opinions on a specific holder. this procedure needs however to be tested on news documents to show its feasibility. finally, adapting our scheme to discussion forums will require to us adapt our scheme to handle dialogues. a thorough linguistic analysis of the link between opinion and discourse in dialogue will be very interesting. 6.3 towards discourse-based sentiment analysis the casoar corpus is a first step towards automated discourse-based opinion analysis. we have already used a subset of this corpus in order to investigate how discourse can help in different sentiment analysis stages. in benamara et al. (2011), we investigated how discursive features could improve subjectivity analysis. we automatically distinguished between subjective non-evaluative (sn) and objective segments (o) and between implicit (si) and explicit opinions (se), by using both local and global context features. chardon et al. (2013) exploited the french gold standard corpus to determine what are the best strategies that need to be implemented to automatically compute a document overall opinion. here we have made a complementary, in depth multi-lingual and multigenre analysis of a new corpus study for english and provided new results concerning the french corpus. a final issue is how to validate our results on automatically parsed data. since review style documents are relatively short, we believe that building such a discourse parser becomes easier. as far as we know, the only existing powerful discourse parser based on sdrt is the one that has been developed on the top of the annodis corpus (muller et al. (2012)). this parser achieves between 47 and 66% accuracy on the structure for the full set of 17 relations. we plan to adapt this parser to opinion texts. in particular, given our observations (cf. table 12), we propose to discard certain relations from the learning process and to group others according to their similar effect on both subjectivity and polarity analysis. this will reduce the number of relations to be predicted to 10 instead of 17 actually which, we believe, will make our discourse parser more reliable. 43 benamara, asher, mathieu, popescu, chardon 7. conclusion in this paper, we presented the casoar corpus, a multi-layered annotation scheme for analyzing opinion in discourse that includes: the complete discourse structure according to the segmented representation discourse theory, the opinion orientation of elementary discourse units and opinion expression annotation. for each layer, we presented the annotation model, annotation guide, and results of its annotation campaign. we explored the interactions between these different layers—in particular, the impact of discourse structure on the overall opinion of a document and implicit opinions, the link between discourse and opinion semantic category, and the role of discourse relations on both subjectivity and polarity analysis. our results demonstrate that opinion and discourse structure are strongly related and that discourse is an important cue for sentiment analysis, at least for the corpus genres we have studied. acknowledgments this work was supported by a dga-rapid project under grant number 0102906143. we also thank our annotators: simon leva, nicolas nouno, anny soubeille, lisa petersen, and julie hall for their efforts during the annotation campaign. the authors also thank erc grant 269427 for research support. we finally thank the editor and three anonymous reviewers for their constructive comments, which helped us to improve the manuscript. references stergos afantenos, pascal denis, philippe muller, and laurence danlos. learning recursive segments for discourse parsing. in proceedings of the seventh language resources and evaluation conference (lrec), pages 3578–3584, 2010. stergos afantenos, nicholas asher, farah benamara, myriam bras, cecile fabre, mai ho-dac, anne le draoulec, philippe muller, marie-paul pery-woodley, laurent prevot, josette rebeyrolles, ludovic tanguy, marianne vergez-couret, and laure vieu. an empirical resource for discovering cognitive principles of discourse organisation: the annodis corpus. in proceedings of the eighth international conference on language resources and evaluation (lrec), 2012. cem akkaya, janyce wiebe, and rada mihalcea. subjectivity word sense disambiguation. in proceedings of empirical methods in natural language processing (emnlp), pages 190–199, 2009. nicholas asher. reference to abstract objects in discourse. kluwer, dordrecht, 1993. nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. nicholas asher, farah benamara, and yvette yannick mathieu. distilling opinion in discourse: a preliminary study. in proceedings of international conference on computational linguistics (coling), pages 7–10, 2008. nicholas asher, farah benamara, and yannick mathieu. appraisal of opinion expressions in discourse. linguisticae investigationes 32:2, 32(2):279–292, 2009. 44 evaluation in discourse carmen banea, rada mihalcea, and janyce wiebe. multilingual subjectivity: are more languages better. in proceedings of the international conference on computational linguistics (coling), pages 28–36, 2010. farah benamara, carmine cesarano, antonio picariello, diego reforgiato, and v. s. subrahmanian. sentiment analysis: adjectives and adverbs are better than adjectives alone. in in proceedings of the international conference on weblogs and social media (icwsm), 2007. farah benamara, baptiste chardon, yannick mathieu, and vladimir popescu. towards contextbased subjectivity analysis. in proceedings of the international joint conference on natural language processing (ijcnlp), pages 1180–1188, 2011. yves. bestgen, cédrick. fairon, and laurent. kevers. un barométre affectif effectif. in g. purnelle, c. fairon, and a. dister (eds.), actes des septiéme journées internationales d’analyse statistique des données textuelles, pages 182–191, 2004. parminder bhatia, yangfeng ji, and jacob eisenstein. better document-level sentiment analysis from rst discourse parsing. in proceedings of the conference on empirical methods in natural language processing, (emnlp), 2015. john blitzer, dredze mark, and pereira fernando. biographies, bollywood, boom-boxes and blenders: domain adaptation for sentiment classification. in proceedings of the annual meeting of the association for computational linguistics (acl), pages 440–447, 2007. ester boldrini, alexandra balahur, patricio martnez-barco, and andrés montoyo. using emotiblog to annotate and analyse subjectivity in the new textual genres. in data mining and knowledge discovery, volume 25, issue 3, pages 603–634, 2012. cristina bosco, viviana patti, and andrea bolioli. developing corpora for sentiment analysis: the case of irony and senti-tut. in ieee intelligent systems, special issue on knowledge-based approaches to content-level sentiment analysis. 28:2, 2013. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in jan van kuppevelt and ronnie smith, editors, current directions in discourse and dialogue, pages 85–112. kluwer academic publishers, 2003. baptiste chardon, farah benamara, yvette yannick mathieu, vladimir popescu, and nicholas asher. measuring the effect of discourse structure on sentiment analysis. in proceedings of the computational linguistics and intelligent text processing (cicling), pages 25–37, 2013. yejin choi and claire cardie. learning with compositional semantics as structural inference for subsentential sentiment analysis. in proceedings of the conference on empirical methods in natural language processing (emnlp), pages 793–801, 2008. simon clematide, stefan gindl, manfred klenner, stefanos petrakis, robert remus, josef ruppenhofer, ulli waltinger, and michael wiegand. mlsa. a multi-layered reference corpus for german sentiment analysis. in proceedings of the eight international conference on language resources and evaluation (lrec), 2012. 45 benamara, asher, mathieu, popescu, chardon béatrice daille, estelle dubreil, laura monceaux, and mathieu vernier. annotating opinionevaluation of blogs : the blogoscopy corpus. in language resources and evaluation springer, 45(4), pages 409–437, 2011. bas heerschop, frank goossen, alexander hogenboom, flavius frasincar, uzay kaymak, and franciska de jong. polarity analysis of texts using discourse structure. in proceedings of the 20th acm international conference on information and knowledge management, pages 1061–1070, 2011. alexander hogenboom, flavius frasincar, franciska de jong, and uzay kaymak. using rhetorical structure in sentiment analysis. commun. acm, 58(7):69–77, 2015. minqing hu and bing liu. mining and summarizing customer reviews. in proceedings of the tenth acm sigkdd international conference on knowledge discovery and data mining, pages 168– 177, 2004. hans kamp and uwe reyle. from discourse to logic. dordrecht, 1993. alistair kennedy and diana inkpen. sentiment classification of movie and product reviews using contextual valence shifters. computational intelligence, 22(2):110–125, 2006. jason s. kessler, miriam eckert, lyndsay clark, and nicolas nicolov. the 2010 icwsm jdpa sentment corpus for the automotive domain. in 4th int’l aaai conference on weblogs and social media data workshop challenge (icwsm-dwc 2010), 2010. angeliki lazaridou, ivan titov, and caroline sporleder. a bayesian model for joint unsupervised induction of sentiment, aspect and discourse representations. in proceedings of the 51st annual meeting of the association for computational linguistics (acl), pages 1630–1639, 2013. cane wing-ki leung, chi-fai chan stephen, fu-lai chung, and grace ngai. a probabilistic rating inference framework for mining user preferences from reviews. world wide web, 14(2):187–215, 2011. shoushan li, sophia y. m. lee, ying chen, chu-ren huang, and guodong zhou. sentiment classification and polarity shifting. in proceedings of the 23rd international conference on computational linguistics (coling), pages 635–643, beijing, china, 2010. bing liu. sentiment analysis and opinion mining (introduction and survey). morgan claypool publishers, 2012. feifan liu, li bin, and liu yang. finding opinionated blogs using statistical classifiers and lexical features. in proceedings of the third international aaai conference on weblogs and social media (icwsm-2009), 2009. qu lizhen, georgiana ifrim, and gerhard weikum. the bag-of-opinions method for review rating prediction from sparse text patterns. in proceedings of the international conference on computational linguistics (coling), pages 913–921, 2010. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text, vol. 8, no. 3., 1988. 46 evaluation in discourse daniel marcu. the theory and practice of discourse parsing and summarization. mit press, cambridge, ma, usa, 2000. rada mihalcea, carmen banea, and jan wiebe. learning multilingual subjective language via cross-lingual projections. in association of computational linguistics (acl), 2007. karo moilanen and stephen pulman. sentiment composition. in proceedings of recent advances in natural language processing (ranlp), pages 378–382, 2007. subhabrata mukherjee and pushpak bhattacharyya. sentiment analysis in twitter with lightweight discourse analysis. in proceedings of the international conference on computational linguistics (coling), pages 1847–1864, 2012. tony mullen and collier nigel. sentiment analysis using support vector machines with diverse information sources. in proceedings of the conference on empirical methods in natural language processing (emnlp), pages 412–418, 2004. p. muller, afantenos, s., denis p., and n. asher. constrained decoding for text-level discourse parsing. in proceedings of the international conference on computational linguistics (coling), pages 1883–1900, 2012. bo pang and lillian lee. a sentimental education: sentiment analysis using subjectivity summarization based on minimum cuts. in proceedings of the 42nd annual meeting on association for computational linguistics (acl), pages 271–278, 2004. bo pang and lillian lee. seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales. in association of computational linguistics (acl), pages 115–124, 2005. bo pang and lillian lee. opinion mining and sentiment analysis. foundations and trends in information retrieval, 2(1-2):1–135, 2008. bo pang, lillian lee, and shivakumar vaithyanathan. thumbs up?: sentiment classification using machine learning techniques. in proceedings of the conference on empirical methods in natural language processing (emnlp), pages 79–86, 2002. livia polanyi and martin van den berg. discourse structure and sentiment. in data mining workshops (icdmw), pages 97–102, 2011. livia polanyi and annie zaenen. contextual valence shifters. in computing attitude and affect in text: theory and applications, volume 20 of the information retrieval series, pages 1–10, 2006. rashmi prasad, nikhil dinesh, alan lee, aravind joshi, and bonnie webber. annotating attribution in the penn discourse treebank. in proceedings of the workshop on sentiment and subjectivity in text, sst ’06, pages 31–38. association for computational linguistics, 2006. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of the sixth international conference on language resources and evaluation (lrec), 2008. 47 benamara, asher, mathieu, popescu, chardon jonathon read and john carroll. annotating expressions of appraisal in english. in language resources and evaluation., 46(3), pages 421–447, 2012. ellen riloff, janyce wiebe, and theresa wilson. learning subjective nouns using extraction pattern bootstrapping. in proceedings of the seventh conference on natural language learning, conll 2003, held in cooperation with hlt-naacl, pages 25–32, 2003. al mostafa shaikh, helmut prendinger, and ishizuka mitsuru. assessing sentiment of text by semantic dependency and contextual valence analysis. in proceedings of the 2nd international conference on affective computing and intelligent interaction, pages 191–202, 2007. swapna somasundaran. discourse-level relations for opinion analysis. phd thesis, university of pittsburgh, 2010. swapna somasundaran and janyce wiebe. recognizing stances in ideological on-line debates. in in proceedings of the workshop on computational approaches to analysis and generation of emotion in text. north american association for computational linguistics (naacl), pages 116–124, 2010. swapna somasundaran, josef ruppenhofer, and janyce wiebe. detecting arguing and sentiment in meetings. in proceedings of the sigdial workshop on discourse and dialogue, pages 26–34. association for computational linguistics, 2007. maite taboada, voll kimberly, and brooke julian. extracting sentiment as a function of discourse structure and topicality. in school of computing science technical report 2008-20, 2008. maite taboada, julian brooke, and manfred stede. genre-based paragraph classification for sentiment analysis. in proceedings of the sigdial 2009 conference: the 10th annual meeting of the special interest group on discourse and dialogue, sigdial ’09, pages 62–70, 2009. maite taboada, brooke julian, tofiloski milan, kimberly voll, and manfred stede. lexicon based methods for sentiment analysis. in computational linguistics, 2011. 37(2), pages 267–307, 2011. cigdem toprak, niklas jakob, and iryna gurevych. sentence and expression level annotation of opinions in user-generated discourse. in proceedings of the 48th annual meeting of the association for computational linguistics (acl), pages 575–584, morristown, nj, usa, 2010. rakshit s. trivedi and jacob eisenstein. discourse connectors for latent subjectivity in sentiment analysis. in human language technologies: conference of the north american chapter of the association of computational linguistics (hlt-naacl), pages 808–813, 2013. radoslava trnavac and maite taboada. the contribution of nonveridical rhetorical relations to evaluation in discourse. language sciences, 34 (3):301–318, 2010. peter d. turney. thumbs up or thumbs down?: semantic orientation applied to unsupervised classification of reviews. in proceedings of annual meeting of the association for computational linguistics (acl), pages 417–424, 2002. fei wang and yunfang wu. exploiting hierarchical discourse structure for review sentiment analysis. in proceedings of the international conference on asian language processing (ialp), pages 121–124, 2013. 48 evaluation in discourse janyce wiebe and ellen riloff. creating subjective and objective sentence classifiers from unannotated texts. in proceedings of the international conference on intelligent text processing and computational linguistics (cicling), lecture notes in computer science, pages 486–497, 2005. janyce wiebe, theresa wilson, and claire cardie. annotating expressions of opinions and emotions in language. language resources and evaluation, 39(2–3):165–210, 2005. florian wolf and edward a. gibson. coherence in natural language: data stuctures and applications. mit press, 2006. hong yu and hatzivassiloglou vasileios. towards answering opinion questions: separating facts from opinions and identifying the polarity of opinion sentences. in proceedings of the conference on empirical methods in natural language processing (emnlp), pages 129–136s, 2003. lanjun zhou, binyang li, wei gao, zhongyu wei, and kam-fai wong. unsupervised discovery of discourse relations for eliminating intra-sentence polarity ambiguities. in proceedings of the conference on empirical methods in natural language processing (emnlp), pages 162–171, 2011. cäcilia zirn, mathias niepert, heiner stuckenschmidt, and michael strube. fine-grained sentiment analysis with structural features. in proceedings of the 5th international joint conference on natural language processing (ijcnlp), pages 336–344, 2011. 49 introduction background existing corpora annotated with sentiment data basic annotation unit annotation levels existing corpora annotated with discourse overview of the segmented discourse representation theory (sdrt) edu determination attachment decision relation labelling the casoar corpus methods annotation scheme the document level the segment level the opinion expression level a complete example annotation procedure data preparation annotation campaign reliability of the annotation scheme at the document level at the segment/opinion level results quantitative analysis at the document level distribution of discourse relations per corpus genre importance of cdus quantitative analysis at the segment and opinion expression level impact of discourse on sentiment analysis discourse and opinion semantic categories discourse relations and subjectivity analysis discourse relations and polarity analysis segment type, segment polarity, and overall opinion discussions interim conclusions portability of the annotation scheme towards discourse-based sentiment analysis conclusion dialogue & discourse 10(1) (2019) 1–19 doi: 10.5087/dad.2019.101 reinforcement adaptation of an attention-based neural natural language generator for spoken dialogue systems matthieu riou matthieu.riou@alumni.univ-avignon.fr ceri-lia, avignon université avignon, france bassam jabaian bassam.jabaian@univ-avignon.fr ceri-lia, avignon université avignon, france stéphane huet stephane.huet@univ-avignon.fr ceri-lia, avignon université avignon, france fabrice lefèvre fabrice.lefevre@univ-avignon.fr ceri-lia, avignon université avignon, france editor: vera demberg submitted 02/2018; accepted 01/2019; published online 02/2019 abstract following some recent proposals to handle natural language generation in spoken dialogue systems with long short-term memory recurrent neural network models (wen et al., 2016a), this work first investigates a variant thereof with the objective of a better integration of the attention sub-network. our second objective is to propose and evaluate a framework to adapt the nlg module on-line through direct interactions with users. the basic way to do so is to lead the users to utter alternative sentences rephrasing the expression of a particular dialogue act. to add such a new sentence to its model, the system can rely on automatic transcription, which is costless but error-prone, or ask the user to transcribe it manually, which is almost flawless but costly. to optimise this choice, we investigate a reinforcement learning approach based on an adversarial bandit scheme. the bandit reward is defined as a linear combination of expected payoffs, on the one hand, and costs of acquiring the new data provided by the user, on the other hand. we show that this definition allows the system designer to find the right balance between improving the system performance, for a better match with the user’s preferences, and limiting the burden associated with it. finally, the actual benefits of the system are assessed through a human evaluation, showing that the progressive inclusion of more diverse utterances increases user satisfaction. keywords: natural language generation, recurrent neural network, adversarial bandit, on-line learning, user adaptation 1. introduction in a spoken dialogue system, the natural language generation (nlg) component aims to produce an utterance from a system dialogue act (da) decided by the dialogue manager. for instance, the c©2019 matthieu riou, bassam jabaian, stéphane huet, and fabrice lefèvre this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). riou, jabaian, huet and lefèvre system da: inform(name=bar_metropol, type=bar, area=north, food=french) may generate the utterance bar metropol is a bar in the northern part of town serving french food. traditional nlg systems use patterns and rules to generate system answers. recently, several proposals have emerged to address the data-driven language generation issue (see for instance rieser and lemon, 2011, chap. 9) . they can be roughly grouped into two main categories: neural translation of dialogue acts and utterance language models. in the latter group, generation is embedded into the whole process of interaction and each new system utterance is sampled from a neural network conditioned by the history of the dialogue (e.g., serban et al., 2016) . in the first group, a more classical compositional approach has been followed, consisting in translating a targeted da (or meaning representation) into a surface form (e.g., wen et al., 2015b) with a recurrent network model close to the seq2seq model (bahdanau et al., 2014). this work is in line with previous studies showing that the transfer between texts and das can be directly handled by a general language translation approach (jabaian et al., 2016) or inverted semantic parsers (konstas and lapata, 2013). in all these cases, a difficulty remains: a huge amount of data is required. we propose to address this difficulty hampering the practical development of such models by combining the current template-based approach with the on-line training of a neural nlg model. some corpus extension methods are also possible (e.g., manishina et al., 2016) but they do not allow a simultaneous adaptation to the user’s preferences. the overall scheme consists in bootstrapping a first version of the model based on a corpus built with some simple templates and a small information database (to help fill in the template placeholders with values). this model sets up a first version of the dialogue system; once operational, the initial system is used to collect new training data while interacting with users. it should be noted that at this critical step of development, users should still be under the control of the designers (they can be designers themselves or colleagues), as it can be hazardous to let the general public directly access such a functionality without any efficient means to counterbalance the effect of the on-line adaptation. this difficult and sensitive point will be addressed more thoroughly in future work. the objective is to maintain the additional workload of the user resulting from the system’s requests at an admissible level. indeed, to collect new data for its model, the system will have to decide at each turn whether: 1. it should ask the user for an alternative to its current answer, 2. it can use the automatic transcription of the user’s input directly or ask for additional processing. basically, such processing will consist in manual corrections of the transcription, but ideally this step could also be handled vocally, which could be rather tedious if done properly. this paper is organised as follows: after presenting related work in section 2, we define our novel nlg model in section 3. section 4 describes the framework we propose to adapt the model on-line through direct interactions. section 5 provides an experimental study with automatic and human evaluations of our approach. we conclude our discussion and propose further perspectives in section 6. 2. related work template-based models still constitute the mainstream method used in the nlg field for commercial purposes. they rely on hand-crafted rules and linguistic resources and turn out to produce good-quality utterances for repetitive and specific tasks (rambow et al., 2001). for this reason, 2 reinforcement adaptation of an attention-based neural natural language generator the nlg component has long received less attention in dialogue system research than spoken language understanding or dialogue management components for instance. however, recent studies have tried to alleviate two main drawbacks of the template-based models: the lack of scalability to large open domains and the frequent repetition of identical and mechanical utterances (see gatt and krahmer, 2018, for a recent survey of the current trends in the nlg field) . one example is to build upon stylistic generation with psychological underpinnings to adjust to the user’s personality dynamically (mairesse and walker, 2010). data-driven and stochastic approaches have been devised to increase maintainability and extensibility. oh and rudnicky proposed to use a set of word-based n-gram language models (lms) to over-generate a set of candidate utterances, from which the final form is selected (oh and rudnicky, 2002). mairesse and young extended this model by introducing factors built over a coarse-grained semantic representation to build phrase-based lms (mairesse and young, 2014). more recently, wen et al. have proposed several models based on recurrent neural networks (rnns) (wen et al., 2015a,c,b; mei et al., 2016). some recent extensions include the proposition of dušek and jurčíček of a seq2seq model with attention to produce both strings and deep syntax trees in a joint generation, replacing the classical pipeline (dušek and jurčíček, 2016). evaluations made by human judges show that these systems are able to generate high-quality utterances which are also more linguistically varied than template-based approaches. the use of recurrent encoder-decoder nns has also been investigated to build end-to-end dialogue systems in a non-goal-directed context, for which large corpora are available (serban et al., 2016), or selective generation from weather forecasting and sportscasting datasets (mei et al., 2016). our proposal is related to two of the generation models proposed by wen et al.: 1. the semantically conditioned lstm-based model (sclstm) introduces an additional control cell into the long short-term memory (lstm) to decide for each generated word what information to retain for the remaining part of the utterance (wen et al., 2015a); 2. the rnn encoder-decoder architecture with an attention mechanism encodes the dialogue act into a distributed vector representation with attention screening over slot-value pairs updated after each generated word. after that, a decoder eventually produces a word sequence with an lstm network (wen et al., 2015b). stochastic models still require extensive work to produce corpora for new domains. novikova et al. proposed a crowd-sourcing framework to collect data for nlg (novikova et al., 2016). wen et al. presented an incremental recipe to deal with the domain adaptation problem for rnn-based generation models (wen et al., 2016b). they used counterfeited data synthesised from an out-ofdomain dataset to fine-tune their model on a small set of in-domain utterances. we still aim at reducing the burden to produce new data, not to adapt to another domain like walker et al. (2007), but to generate more diverse utterances better adapted to the user’s preferences. to this end, a reinforcement learning approach based on an adversarial bandit scheme is applied (auer et al., 2002). this approach has been used previously in dialogue systems for language understanding (ferreira et al., 2015, 2016). here, we propose a protocol to adapt the rnn-based model on new utterances that vary from the training dataset, taking into account the cost for the user to provide these examples. other approaches based on active learning have been used in nlg. for instance, mairesse et al. (2010) included an active learning protocol in their language generator in order to optimise the data collection process, using a model which can determine the next semantic input to annotate based on its estimated certainty about the correctness of its output. likewise, fang et al. (2017) proposed a deep reinforcement learning algorithm capable of learning an active learning strategy from data in 3 riou, jabaian, huet and lefèvre order to decide whether or not to annotate each utterance. but our work is the first to propose the use of an adversarial bandit algorithm to support the decision-making process for the active learning in an nlg data collection. the use of bandit algorithms has already been investigated in various active learning protocols, for example in recommendation systems by li et al. (2010), or in dialogue systems where they have been applied to automatically update spoken language understanding models deployed in a spoken service that evolved with time (gotab et al., 2009). 3. a combined-context lstm for language generation this section presents the generation model proposed in this paper. it is based on two previous models: the semantically conditioned lstm (wen et al., 2015a) and the attention-based rnn encoder-decoder (wen et al., 2015b). after a detailed description, the proposed model is compared with the reference models to point out precisely where the expected benefits of the combined model lie. then the training and decoding processes are described. 3.1 model description our model is based on the same recurrent neural architecture as (wen et al., 2015a). the overall principle is to generate each new element of the word sequence conditioned on the previouslygenerated one, a hidden (recurrently updated) vector and a context-information vector. in practice, a 1-hot encoding wt−1 of a token1 wt−1 is input to the model at each time step t conditioned on a recurrent hidden layer ht−1, from which the probability distribution of the next token wt is defined. to ensure that the generated utterance represents the intended meaning, an additional context vector dt, encoding the dialogue act and its associated slot-value pairs, is also provided at each step t. as in the attention-based encoder-decoder, the decoding process is performed through a standard lstm (fig. 1b), which is fed by an additional vector at representing the information on which the model currently focuses (fig. 1a). at is called the local da embedding with attention. the set of relations between all the vectors involved in the lstm cell is: it = sigmoid(wwiwt−1 +whiht−1 +waiat) (1) ft = sigmoid(wwfwt−1 +whfht−1 +wafat) (2) ot = sigmoid(wwowt−1 +whoht−1 +waoat) (3) ĉt = tanh (wwcwt−1 +whcht−1 +wacat) (4) ct = ft � ct−1 + it � ĉt (5) ht = ot � tanh (ct) (6) where it, ft,ot ∈ [0, 1]n are input, forget and output gates respectively, ĉt and ct are proposed and true cell values at time t, and � denotes element-wise multiplication. subsequently, the next token wt is picked up, either through argmax or sampling, on the output distribution formed as: p (wt|wt−1, wt−2, ...w0,at) = softmax(whoht) (7) wt ∼ p (wt|wt−1, wt−2, ...w0,at) (8) 1. the same terminology as in wen et al. (2015a) is used since the input text is also delexicalised: the slot values (e.g.“chinese” for slot food) are replaced in the input by their corresponding slot tokens (e.g. slot_food). 4 reinforcement adaptation of an attention-based neural natural language generator lstm decoder ht wt−1 ht−1 r a dt−1 at dt wt−1 ht−1at−1 z0 ht−1 at−1 (a) global architecture of the combined-context lstm. ciĉ wt−1 ht−1 at wt−1 ht−1at f ct−1 wt−1 ht−1 at ct o wt−1 ht−1at ht (b) details of the lstm decoder. figure 1: combined-context lstm where who is the output weight matrix. as with the attention-based encoder-decoder model, the local da embedding at is computed from the global da embedding dt, which represents the information remaining to express in the rest of the generation. to define the initial global da d0, each slot-value pair is embedded as a vector z0,i: z0,i = si + vi (9) where si and vi are the i-th slot and value pair of the dialogue act, each represented by a 1-hot representation. then the complete dialogue act is represented by: d0 = act0 ⊕ ∑ i z0,i (10) where act0 is a 1-hot representation of the act type and ⊕ stands for vector concatenation. therefore, the global dialogue act dt, corresponding to the remaining information to deliver, is obtained at each time step by: dt = actt ⊕ ∑ i zt,i . (11) actt and zt,i are updated with the parts of si and vi that remain in the global dialogue act dt according to the reading gate r (fig. 1a): dt = rt � dt−1 (12) rt = sigmoid(wwrwt−1 +whrht−1 +warat−1) . (13) the local da embedding at, representing the information we want to focus on at step t, is formed as: at = actt ⊕ ∑ i ωt,izt,i (14) 5 riou, jabaian, huet and lefèvre where ωt,i is the weight of i-th slot-value pair computed by the attention mechanism a: ωt,i = softmax(βt,i) (15) βt,i = qt. tanh (whmht−1 +wmmz0,i +wamat−1) (16) q and ws being parameters to learn. 3.2 comparison with the reference models the generation model proposed here combines the semantically conditioned lstm (wen et al., 2015a) and the attention-based rnn encoder-decoder (wen et al., 2015b). each of these models proposes a way to process the semantic information represented as a da to produce an utterance. without delving into details (for which we strongly advise to refer to the original papers), we briefly recall the structure of the two models in figure 2 and try to summarise their main differences w.r.t. the processing of their input data, the dialogue acts. the sclstm reading-gate handles the da by choosing what information to retain or discard at each step as illustrated in figure 2 (a). for this purpose, at each step the reading-gate takes as input the last turn’s remaining unprocessed information dt−1 and outputs dt, conditioned on the previous word wt−1 and the lstm state ht−1. conversely, the attention mechanism takes as input the initial da d0 at each step, and outputs the information to process at, conditioned on the initial da d0 and the previous lstm state ht−1. however, it loses the progression of unprocessed information (figure 2 (b)). subsequent to this, for both models, an lstm decoder generates the next lstm state ht from which the next word wt is picked up (using equation 7), conditioned on the previous word wt−1 and the lstm state ht−1. each model offers some advantages and inconveniences. a plus is that the sclstm is less inclined to forget slots as the reading gate informs on the remaining unprocessed information. a disadvantage is that it tends to deliver incoherent and ungrammatical sentences in order to deliver all the slots at all costs. for example, for the following input da: inform(name=restaurant ducroix, kids_allowed=no, phone=4153917195, postcode=94111, address=690 sacramento street) an sclstm generates: the address of restaurant ducroix is 690 sacramento street child and allowed and is 4153917195 and the postcode is 94111. where concepts kids_allowed and phone are present but wrongly formulated. for its part, the encoder-decoder uses the attention mechanism to select the part of the da that should be considered by the lstm decoder at each step. thus the system can better process each slot locally. but it has no dedicated mechanism to ensure that all slots have been processed at the end of sentence. for example, for the following da: inform(name=thep phanom thai, address=400 waller street, postcode=94117, phone=415431256) an attention-based rnn encoder-decoder proposes: the address for thep phanom thai restaurant is 400 waller street and the postcode is 94117. 6 reinforcement adaptation of an attention-based neural natural language generator lstm decoder ht wt−1 ht−1 rdt−1 dt dt wt−1 ht−1 (a) lstm decoder ht wt−1 ht−1 a d0 at d0 ht−1 (b) figure 2: semantically conditioned lstm (a) and attention-based rnn encoder-decoder (b). which is grammatically correct but in which the phone number is missing. our objective is to combine the advantages of both models, using a reading gate and an attention mechanism to sequentially process the da. thus, the system is less inclined to forget or misprocess some slots, and as a consequence should improve its bleu score and slot error rate. therefore, in practice, the major novelty of our model lies in the computation of at in the attention mechanism (see equation 14), which takes as input the current da dt (decomposed in actt and zt,i) instead of the initial da d0. this current da is obtained from the output of a reading-gate rt (see equations 12 and 13). besides, the previous attention’s output at−1 is added as a parameter in both the reading-gate rt and the attention’s weights ωt,i (see equations 15 and 16). 3.3 training and decoding the objective function used to train the weights of the network computes the cross-entropy between the predicted token distribution pt and the actual token distribution yt: f (θ) = ∑ t (pt t log (yt)) + ‖dt ‖+ t−1∑ t=0 (η ∗ ξ‖dt+1−dt‖) . (17) following wen et al. (2015c), an l2 regularisation term is introduced as well as a second regularisation term2 required to control the reading gate dynamics. we optimise the parameters with stochastic gradient descent and back propagation through time. early-stopping on a validation set prevents over-fitting. the decoding is split into two steps: 1. during an over-generation phase, the system is used to generate several utterances for the given da, by randomly picking the next token on the output distribution, and 2. in a subsequent re-ranking phase, each utterance is ranked on the basis of a score r calculated as: r = −(f (θ) + λerr) (18) where λ is a trade-off constant, set to 10, and err is the slot error rate. err = (p + q)/n with n the total number of da slots, and p, q the number of missing and redundant slots in the proposed utterance, compared to the input da. 2. t is the total number of steps, η = 10−4, ξ = 100. 7 riou, jabaian, huet and lefèvre 4. on-line interactive problem neural nlg can give good results, but it requires a large amount of annotated data to be trained in order to have an efficient model with diversity in its outputs. several examples of utterances for each da are then required to train the model. in order to reduce the cost of collecting such a corpus, the following on-line learning protocol was set up. we propose to proceed in two main steps: 1. a bootstrap corpus, consisting of references generated from templates, is used to train a generation model; 2. this learned model generates new utterances and the users are required to propose better or varied alternatives. in order to reduce the effort on the user’s side and to avoid useless actions, we propose to rely on an adversarial bandit algorithm to decide whether the system should prod the user considering the expected gain and cost of its action or not. 4.1 static case once the system generates the utterance, the system can choose one action (from a probability distribution) among a set i of m actions. in this preliminary setup, we consider a case where m = 3 and i can be defined as: i := {skip, askdictation, asktranscription}. let i ∈ i be the action index. we assume that the user effort φ(i) ∈ n can be measured by the time needed to perform action i. the actions and associated user efforts are: • skip: skip the refinement process. the cost of this action is always set to 0 (φ(skip) = 0). • askdictation: refine the model, taking into account an alternative utterance proposed by the user and transcribed automatically with an asr system (φ(askdictation) = 1). • asktranscription: ask the user to transcribe the correction or the alternative utterance. two different costs are considered for this action: – un-normalised cost: φ(asktranscription) = 1 + l – normalised cost: φ(asktranscription) = 1 + l lmax with l the length of the proposed utterance, and lmax the maximum possible length (set to 40 words in our experiments). then the gain of the chosen action is estimated as follows: • skip: nothing is learned, gain is 0 (g(skip) = 0). • askdictation: we propose to compute the gain as the remaining margin of the bleu-4 score that would have been obtained by the utterance generated by the system, using the userproposed utterance as a reference, noted bleugen/prop. to take into account the potential errors added by the asr system, the gain is penalised by the estimations of wer and err: g(askdictation) = (1− bleugen/prop)× (1−wer)× (1− err) . 8 reinforcement adaptation of an attention-based neural natural language generator the global wer expresses the confidence we can have in the bleu-4 measure (as it is based on erroneous utterances), while the slot error rate err penalises utterances that do not contain the required semantic information, due to asr errors. • asktranscription: asking the user to manually transcribe the utterance prevents asr errors. therefore, the gain estimate only considers the bleu-4 score of the utterance generated by the system, using the user-proposed sentence as a reference (g(asktranscription) = 1 − bleugen/prop). finally, a loss function is defined l(i) ∈ [0, 1] such that the system, through an optimisation process, seeks to maximise the gain measure g(i) and to minimise the user effort φ(i) jointly: l(i) = α(1− g(i))︸ ︷︷ ︸ system improvement +(1− α) φ(i) φmax︸ ︷︷ ︸ user effort (19) very importantly α allows to weight the payoff w.r.t. the cost, allowing the designer to influence the system’s behavior depending on the targeted operational conditions (from fast improvement, no matter the cost, down to slow improvement to preserve users’ efforts). 4.2 adversarial bandit case the following scenario for the adversarial bandit problem is considered: the system produces a sentence then chooses an action it ∈ i. once the action it is performed, the system computes: (a) the gain estimate gt(it) with the user collaboration, (b) the user effort φt(it) and (c) the current loss. the goal of the bandit algorithm is then to find i1, i2, . . . , so that for each t , the system minimises the total loss as expressed in the previous section. every n iterations, the user-proposed utterances are added to the training corpus, and the model is updated on this extended corpus. at the same time, we compute the loss function for each bandit’s choice, and update its policy. 5. experimental study in this section, the improvement of the combined-context lstm over sclstm and encoderdecoder is measured (section 5.1). then the on-line learning protocol is evaluated on simulated data in section 5.2. in order to evaluate whether (or not) the on-line learning approach has an impact on real users’ subjective appreciation of systems, a human evaluation is made in section 5.3. finally, in section 5.4, the impact of the wer simulation is evaluated on a smaller dataset, collected with a real asr. 5.1 system comparison a first set of experiments was conducted on the sf restaurant corpus, described in wen et al. (2015c) and freely accessible.3 it contains 5 191 utterances, for 271 distinct das. with each da, the corpus associates a template-generated utterance and several utterances in natural english proposed by humans, each utterance being delexicalised. 3. https://www.repository.cam.ac.uk/handle/1810/251304 9 https://www.repository.cam.ac.uk/handle/1810/251304 riou, jabaian, huet and lefèvre system bleu-4 err (%) sclstm 0.722* 0.66 encoder-decoder 0.697 0.65 combined-context lstm 0.711 0.24** * significant w.r.t the encoder-decoder by the t-test (p-value< 0.01) ** significant by the t-test (p-values< 0.001) table 1: results on the top 5 hypotheses. candidate reference sclstm encoder-decoder combined-context lstm sclstm 0.857 0.883 encoder-decoder 0.847 0.824 combined-context lstm 0.849 0.840 table 2: bleu-4 cross-comparison of the three systems. our model and both the sclstm and the attention-based rnn encoder-decoder were implemented using the tensorflow library4 and were trained on a corpus split into 3 parts: training, validation and testing (3:1:1 ratio), using only the human-proposed utterance references. the three systems were compared using two metrics: the bleu-4 score (papineni et al., 2002) and the slot error rate (err). the bleu-4 value validates the utterance generation, especially grammaticality, but is often not seen as a useful measure of content quality for nlg (reiter and belz, 2009). to remedy that deficiency, err, which concentrates only on the semantic contents but with more accuracy, is also computed. for each example, we over-generated 20 utterances and kept the top 5 hypotheses for evaluation. each hypothesis has been processed as an independent output sentence to evaluate, and so averaged during the bleu-4 computation. multiple references for each da were obtained by grouping delexicalised utterances with the same da specification, and then “relexicalised” with the proper values. as can be seen in table 1, the bleu-4 score of the combined-context lstm falls between the two other systems (roughly 0.01 gap between each) but the measured differences were not statistically significant (p-value> 0.01 for each pair of systems,5 except between the sclstm and the encoder-decoder with a p-value= 0.002). however, the slot error rate is reduced by one third by our new model w.r.t. the two other systems, the improvement being statistically significant between sclstm and combined-context lstm (p-value< 0.001). this means that, while the new model does not really achieve to learn more diverse responses, it offers a better coverage of the expressed concepts resulting in fewer omitted concepts, which is the first purpose of an nlg system. to evaluate the differences in sentence generation, bleu-4 has been used to compare the outputs of each system with the two others. as can be seen from table 2, all cross-system bleu-4 scores are pretty high, around 0.80-0.90. this indicates that systems tend to produce comparable sentences, and as a consequence we did not investigate further the possibility of combining these different neural models into one. 4. https://www.tensorflow.org 5. statistical significance was computed using a two-tailed student’s t-test between each pair of systems. 10 https://www.tensorflow.org reinforcement adaptation of an attention-based neural natural language generator (a) un-normalised cost (b) normalised cost figure 3: evolution of bleu-4 score, as a function of the cumulated learning cost. 5.2 on-line adaptation evaluation for the evaluation of the on-line adaptation procedure, the same corpus is used again. but this time, training, validation and testing parts follow a 2:1:1 ratio (to maintain a test set of minimal size despite a smaller corpus). the combined-context system is used to train an initial bootstrap model on the training set, using the template-generated utterance references. the validation corpus was used for early stopping, again with the template-generated references. then, we simulated the on-line learning on the same training set, using this time the enclosed human-proposed references. the model and bandit updates were learned every 400 utterances. wer was simulated by randomly inserting errors (confusion, deletion, insertion) into the corpus examples until a pre-defined global wer was reached. this wer simulation put aside the idiosyncratic properties of the used language and asr system. but realistic models are very complex to develop and never really satisfactory, notably because different asrs make different errors. for this reason, a rather simple random model for simulation has been chosen, and a test was carried out with an actual asr system afterwards in the user trials. the models are trained on delexicalised utterances, which allows the computation of err. we note that the err score can be reduced when the value of the slot-value pair does not appear in the surface form of utterances (for example with the “dont_care” value). in the on-line learning setup, it can raise new issues if a user proposed an alternative surface form which does not contain the wanted value, resulting in a higher err score. the initial model, trained on the template-generated part of the training corpus, obtains a high bleu-4 score, 0.802, when tested on the template-generated part of the test, but this value is dramatically reduced to 0.397 on the human-proposed references. this tends to confirm that even a well-trained model does not compete with the diversity of possible responses occurring in a conversation in natural language. figure 3 plots the bleu-4 score as a function of the learning cost, the simulated wer being set to 5%. bleu-4 is obtained by testing the model against the human-proposed subset of the test. the learning cost is computed as the sum of the costs of all choices made by the bandit during the learning. different configurations are tested: the forced ‘askdictation’ choice (fcdic) and the forced ‘asktranscription’ choice (fctrans). besides, the bandit is tested with two α values: 0.5 11 riou, jabaian, huet and lefèvre figure 4: evolution of bleu-4 score, as a function of the training data size. (α5) and 0.7 (α7), each displayed with normalised and un-normalised costs. the second value reduces the influence of the cost, allowing the system to increase the effort asked to the user. each curve is composed of seven points; the first one corresponds to the score of the system before online learning, the six others are computed after each block of 400 utterances. the cost is cumulative over all previous blocks. with the un-normalised cost (figure 3a), we can observe that the bandit succeeds in reducing the cost of learning up to a certain amount of training data (after a cumulated cost of 7 500, asktranscription outperforms all other configurations). after using all the training data, α5 and α7 reach both 0.476 bleu, an intermediate value between 0.464 for fcdic and 0.503 for fctrans. asktranscription costs much more than askdictation, therefore, at first, the bandit learns better than fctrans by balancing between the two choices. but after the first two blocks, the increase reduces until both the α5 and α7 curves pass below fctrans. a higher α value tends to favour the ask actions over skip, and asktranscription over askdictation. the normalised cost has been tested with the same configuration. the results (figure 3b) are quite similar. asktranscription still outperforms all other configuration and each setup reaches the same bleu-4 score as with the un-normalised cost. nevertheless, the cost is much lower for α5 (2 085), α7 (2 348) and asktranscription (3 034). the normalised cost reduces the gap between the estimated costs of a dictation and a transcription; thus, it tends to favour more the asktranscription. figure 4 displays the bleu-4 scores as a direct function of the training data size. the results are consistent with previous conclusions, since both α5 and α7 achieve their learning with a reduced amount of data, and performances between the forced ask choices. the bandit was also tested with a higher wer (20%). at this rate, the system no longer learns from the choice askdictation, errors overwhelming improvement. the forced choice askdictation gives a bleu-4 score of 0.383, less than the initial system. an analysis of the learned policy shows that with a low wer (5%), the bandit globally explores both ask learning choices, and presents at the end a slight preference for askdictation. with a higher wer (20%), the bandit favours asktranscription (chosen almost half of the time at the last iteration), due to more utterances with a high slot error rate and therefore a lower gain. 12 reinforcement adaptation of an attention-based neural natural language generator 5.3 human evaluation objective evaluations using automatic metrics like bleu-4 do not necessarily reflect real users’ preferences (callison-burch et al., 2006). in particular, naturalness is really hard to formalise (as in general, more socially-oriented qualities of a system cannot be easily captured with current evaluation modes (e.g., curry et al. (2017) or perez-beltrachini and gardent (2017)) in order to better evaluate whether or not the results of the on-line learning approach induce a better appreciation of the system by the users, a human evaluation was conducted. any of the enhanced models could have been used as the objective here is to confirm that the observations in the simulated environment (bleu and eer scores) transferred well to subjective user’s appraisal, that is: does the adaptation procedure, whatever the exact model used, improve the quality of the generation process? however, due to the cost of the experiment, only the adapted model α5 was chosen to be compared with the initial model (only referred to as ‘the adapted model’ hereafter). five annotators were recruited to automatically evaluate generated utterances. for each example, a dialogue act and the three best sentences generated by each system have been presented to the annotators. the six utterances were randomly ordered, and there was no indication on the system that built them. the annotators have been asked to give three scores to each utterance, all rating from 1 to 3: • informativeness evaluates whether the information given by the dialogue act is well expressed in the generated utterance: – 3: all information given by the dialogue act is conveyed and no additional information is introduced; – 2: minor information is missing or there is extra information not present in the dialogue act; – 1: any other case. • syntax evaluates the level of syntactic correctness of an utterance: – 3: the utterance is correct; – 2: there are small, hardly audible imperfections; – 1: there are some clear mistakes in the utterance. • naturalness evaluates whether the utterance sounds like a potential human production: – 3: the utterance could have been pronounced by a human in this situation; – 2: the utterance is correct but unfit to the situation, or sounds synthetic; – 1: even by correcting the grammatical errors if needed, the utterance could never have been pronounced by a human. to reduce the burden of the annotators, duplicate sentences were merged. in addition to the evaluation of each sentence, they were asked to indicate their preferred one. to evaluate annotator agreement using the fleiss kappa metrics, the first 20 examples were shared over all annotators. a total of 471 evaluations were collected. the κ for the global annotator agreement is 0.550. the task on which the annotators less agreed is rating naturalness, with a κ of 0.468 to be compared with 0.594 and 0.576 respectively for informativeness and syntax correctness. 13 riou, jabaian, huet and lefèvre initial system adapted system global score 2.356 2.425∗ informativeness 2.528 2.509 syntax 2.272 2.383∗ naturalness 2.267 2.383∗ ∗ p < 0.001 table 3: average scores for each system. statistical significance was computed using a two-tailed student’s t-test between the two systems. act type system all informativeness naturalness syntax inform initial 2.50 2.39 2.39 2.42 adapted 2.42 2.52 2.37 2.39 inform_only_match initial 2.33 2.50 2.50 2.00 adapted 1.94 2.00 2.00 1.83 ?inform_no_match initial 2.48 2.78 2.33 2.34 adapted 2.19 2.23 2.07 2.05 ?select initial 2.16 2.52 2.01 1.89 adapted 2.05 2.52 2.37 2.33 ?request initial 2.65 2.82 2.65 2.47 adapted 2.63 2.77 2.57 2.55 ?reqmore initial 2.27 2.67 2.07 2.07 adapted 2.62 2.47 2.73 2.67 ?confirm initial 2.02 2.27 1.91 1.88 adapted 2.05 2.33 1.94 1.88 goodbye initial 1.71 1.70 1.70 1.72 adapted 2.59 2.60 2.59 2.57 table 4: average scores for each system w.r.t. the act type of the dialogue act. as shown in table 3, the adapted system obtains a significantly higher global average score than the initial system. more specifically, the adapted system obtains significantly higher scores for naturalness and syntax, and the slight decrease in informativeness is not significant. table 4 indicates the variation of scores w.r.t. act type. contrary to what one might think, it can be observed that the scores are quite regular over all act types, event though they each represent very variable levels of complexity. on the contrary, table 5 shows that both systems tend to have higher overall scores for dialogue acts of low or medium length (2 or 3 slots). with more slots, the scores tend to decrease as the utterances become more complicated and more conducive to errors with automatic generation. when the annotators were asked to vote for their favourite utterance, they mainly voted for the best utterance of each system, with a significant preference for the adapted system (table 6). but mostly the second and third best propositions of the adapted systems were also selected quite often, unlike the second and third best propositions of the initial system. this tends to confirm that the adapted system can generate more satisfying sentences with a variability greater than the initial system. 14 reinforcement adaptation of an attention-based neural natural language generator # slots system all informativeness naturalness syntax 0 slot initial 1.74 1.76 1.72 1.74 adapted 2.59 2.60 2.60 2.58 1 slot initial 2.38 2.59 2.34 2.21 adapted 2.38 2.55 2.31 2.27 2 slots initial 2.58 2.80 2.50 2.46 adapted 2.64 2.75 2.58 2.58 3 slots initial 2.47 2.72 2.30 2.39 adapted 2.32 2.48 2.26 2.29 4 slots initial 2.28 2.35 2.24 2.27 adapted 1.71 1.74 1.69 1.70 5 slots initial 1.75 2.00 1.67 1.58 adapted 1.66 1.50 1.75 1.58 table 5: average scores for each system w.r.t. the number of slots in the dialogue act. rank initial system adapted system 1 111 (22.0%) 143 (28.5%) 2 51 (10.2%) 103 (20.6%) 3 18 (3.6%) 75 (15.0%) total 180 (35.9%) 321 (64.1%)∗ ∗ p < 0.001 table 6: number of selected sentences by annotators w.r.t. their rank. statistical significance was computed by means of a two-tailed binomial test. 5.4 on-line adaptation using real asr data evaluation to evaluate in practice the on-line framework, and in particular the impact of the wer during the interactions, a data collection has been carried out with true users. to collect the data, a dialogue act with some possible references was presented to the user for each example. then, the user has been asked to dictate an alternative sentence corresponding to the dialogue act. to facilitate the experiment deployment, the capabilities of speech recognition available in the browser chrome, with the google asr api, were used. with the sentence transcribed, the user had the possibility to manually correct the automatic output if needed. both the transcribed and corrected utterances were collected, as well as the confidence score of the automatically transcribed sentences. 426 pairs of transcriptions (automatic, manual) have been collected this way.6 the transcribed utterances present a mean confidence score of 0.86 and a mean wer (between transcribed and corrected utterances) of only 2.42%. a new model was adapted, through the on-line adaptation experimentation described in section 5.2 using the new collected data instead of the former human-proposed references. the new corpus was divided in 300 examples for training and 126 for testing. the same initial bootstrap model was used, but this time it was updated, as well as the bandit, every 50 utterances due to the 6. all data used in this study are available upon request to the authors. 15 riou, jabaian, huet and lefèvre smaller size of the corpus. to enhance the gain estimation, the estimated wer was replaced by the confidence score of the transcription: g(askdictation) = (1− bleugen/prop)× confidence score× (1− err) . this new adapted model has been tested on the same corpus as for the first experiment, by comparing the generated utterances to the references generated by patterns, and to the references proposed by human annotators including the corrected references of the new collected oral data. results are similar to the experiments done with the simulated wer. when using patterns as references, bleu-4 is reduced by 0.10 (from 0.829 to 0.727), while it slightly decreases by 0.03 (from 0.482 to 0.451) when compared to human references. the lowest scores observed with this last setup with human references can be explained by the large number of human utterances from the initial corpus that have not been learned in this experiment. the bandit algorithm avoids having the system interfere too often with the user. during the entire learning, it asked for transcriptions 53% of the time, compared to only 23% for dictation, and it did not ask anything (skipped) 23% of the time. in this way, the cumulated cost of the learning has been divided by two w.r.t. a system that would always ask for transcription (from 4 243 to 2 430) without decreasing too much learning performance (bleu-4 score is 0.458 for the forced transcription system against 0.444 with the bandit). however, it has a lower performance than a system that would always ask for dictation, which obtains a bleu-4 score of 0.451 for a cumulated cost of 300. to allow the system to better estimate whether it has to ask or not for an oral alternative, and whether or not a transcription is needed, the context has to be taking into account by the bandit (the nature of the dialogue act, complexity...). this could be done with the same protocol by using a contextual bandit (auer et al., 2002) instead of the adversarial bandit. however, such a setup is very likely to converge faster than the non-contextual version and thus limit the exploration steps, which the adversarial bandit maintains steadily. in any case, a comparison of the variants is planned to obtain more insights on their true performance. 6. conclusion in this paper we have investigated an attention-based neural network for natural language generation, combining two systems proposed by wen et al.: the semantically conditioned lstm-based model (sclstm) and the rnn encoder-decoder architecture with an attention mechanism. while not improving the bleu score globally, this model outperforms them on the slot error rate, preventing the semantic repetitions or omissions in the generated utterances. we then proposed a protocol to adapt a bootstrapped model using on-line learning. results obtained on a simulated experiment have been confirmed and completed with real users, providing new proposals to the system and assessing the qualities of the adapted system’s hypotheses. the bandit algorithm has been shown to allow the system to balance between improving the system’s performance and the additional workload it implies for the user. it also leads to a system considered more varied by the users. in future work, we will study how to improve the system’s learning ability by taking the context into account before making a choice, with recourse to a contextual bandit. more importantly, the natural language generator has to be evaluated within an entire dialogue system to definitely confirm the practical interest of the overall approach. 16 reinforcement adaptation of an attention-based neural natural language generator 7. acknowledgements this work has been partially carried out within the labex blri (anr-11-labx-0036). references peter auer, nicolò cesa-bianchi, yoav freund, and robert e. schapire. the nonstochastic multiarmed bandit problem. siam journal on computing, 32(1):48–77, 2002. dzmitry bahdanau, kyunghyun cho, and yoshua bengio. neural machine translation by jointly learning to align and translate. sep 2014. url http://arxiv.org/abs/1409.0473. chris callison-burch, miles osborne, and philipp koehn. re-evaluating the role of bleu in machine translation research. in eacl, pages 249–256, 2006. amanda cercas curry, helen hastie, and verena rieser. a review of evaluation techniques for social dialogue systems. in proceedings of the 1st acm sigchi international workshop on investigating social interactions with artificial agents, isiaa 2017, pages 25–26, 2017. ondřej dušek and filip jurčíček. sequence-to-sequence generation for spoken dialogue via deep syntax trees and strings. in proceedings of the 54th annual meeting of the association for computational linguistics (volume 2: short papers), pages 45–51, berlin, germany, 2016. meng fang, yuan li, and trevor cohn. learning how to active learn: a deep reinforcement learning approach. in proceedings of the 2017 conference on empirical methods in natural language processing, pages 595–605, 2017. emmanuel ferreira, bassam jabaian, and fabrice lefèvre. zero-shot semantic parser for spoken language understanding. in proceedings of interspeech, 2015. emmanuel ferreira, alexandre reiffers-masson, bassam jabaian, and fabrice lefèvre. adversarial bandit for online interactive active learning of zero-shot spoken language understanding. in proceedings of icassp, 2016. albert gatt and emiel krahmer. survey of the state of the art in natural language generation: core tasks, applications and evaluation. journal of artificial intelligence research, 61:65–170, 2018. pierre gotab, frédéric béchet, and géraldine damnati. active learning for rule-based and corpusbased spoken language understanding models. in ieee workshop on automatic speech recognition & understanding, 2009, asru 2009, pages 444–449. ieee, 2009. bassam jabaian, fabrice lefèvre, and laurent besacier. a unified framework for translation and understanding allowing discriminative joint decoding for multilingual speech semantic interpretation. computer speech & language, 35:185–199, 2016. ioanis konstas and mirella lapata. a global model for concept-to-text generation. journal of artificial intelligence research, 48:305–346, 2013. lihong li, wei chu, john langford, and robert e. schapire. a contextual-bandit approach to personalized news article recommendation. in proceedings of the 19th international conference 17 http://arxiv.org/abs/1409.0473 riou, jabaian, huet and lefèvre on world wide web, www ’10, pages 661–670, new york, ny, usa, 2010. acm. isbn 9781-60558-799-8. doi: 10.1145/1772690.1772758. url http://doi.acm.org/10.1145/ 1772690.1772758. françois mairesse and marilyn a. walker. towards personality-based user adaptation: psychologically informed stylistic language generation. user modeling and user-adapted interaction, 20 (3):227–278, august 2010. françois mairesse, milica gašić, filip jurčíček, simon keizer, blaise thomson, kai yu, and steve young. phrase-based statistical language generation using graphical models and active learning. in proceedings of the 48th annual meeting of the association for computational linguistics, acl ’10, pages 1552–1561, stroudsburg, pa, usa, 2010. association for computational linguistics. url http://dl.acm.org/citation.cfm?id=1858681.1858838. françois mairesse and steve young. stochastic language generation in dialogue using factored language models. computational linguistics, 40(4):763–799, 2014. elena manishina, bassam jabaian, stéphane huet, and fabrice lefèvre. automatic corpus extension for data-driven natural language generation. in proceedings of lrec, may 2016. hongyuan mei, mohit bansal, and matthew r. walter. what to talk about and how? selective generation using lstms with coarse-to-fine alignment. in proceedings of naacl-hlt, 2016. jekaterina novikova, oliver lemon, and verena rieser. crowd-sourcing nlg data: pictures elicit better data. in proceedings of the 9th international natural language generation conference, pages 265–273. association for computational linguistics, 2016. alice h. oh and alexander i rudnicky. stochastic natural language generation for spoken dialog systems. computer speech & language, 16(3–4):387–407, 2002. kishore papineni, salim roukos, todd ward, and wei-jing zhu. bleu: a method for automatic evaluation of machine translation. in proceedings of the 40th annual meeting on association for computational linguistics, 2002. laura perez-beltrachini and claire gardent. analysing data-to-text generation benchmarks. in proceedings of the 10th international natural language generation conference, santiago de compostelle, spain, september 2017. owen rambow, srinivas bangalore, and marilyn walker. natural language generation in dialog systems. in proceedings of hlt, 2001. ehud reiter and anja belz. an investigation into the validity of some metrics for automatically evaluating natural language generation systems. computational linguistics, 35(4):529–558, 2009. verena rieser and oliver lemon. reinforcement learning for adaptive dialogue systems: a datadriven methodology for dialogue management and natural language generation. theory and applications of natural language processing. springer-verlag new york inc, 2011. iulian v. serban, alessandro sordoni, yoshua bengio, aaron courville, and joelle pineau. building end-to-end dialogue systems using generative hierarchical neural network models. in proceedings of aaai conference on artificial intelligence, 2016. 18 http://doi.acm.org/10.1145/1772690.1772758 http://doi.acm.org/10.1145/1772690.1772758 http://dl.acm.org/citation.cfm?id=1858681.1858838 reinforcement adaptation of an attention-based neural natural language generator marilyn walker, amanda stent, françois mairesse, and rashmi prasad. individual and domain adaptation in sentence planning for dialogue. journal of artificial intelligence research, 30(1): 413–456, november 2007. tsung-hsien wen, milica gašić, dongho kim, nikola mrkšić, pei-hao su, david vandyke, and steve young. stochastic language generation in dialogue using recurrent neural networks with convolutional sentence reranking. in proceedings of sigdial, 2015a. tsung-hsien wen, milica gašić, nikola mrkšić, lina m. rojas-barahona, pei-hao su, david vandyke, and steve young. toward multi-domain language generation using recurrent neural networks. in proceedings of nips workshop on machine learning for spoken language understanding and interaction, 2015b. tsung-hsien wen, milica gašić, nikola mrkšić, pei-hao su, david vandyke, and steve young. semantically conditioned lstm-based natural language generation for spoken dialogue systems. in proceedings of emnlp, 2015c. tsung-hsien wen, milica gašić, nikola mrkšić, lina m. rojas-barahona, pei-hao su, stefan ultes, david vandyke, and steve young. a network-based end-to-end trainable task-oriented dialogue system. technical report, university of cambridge, 2016a. url https://arxiv. org/abs/1604.04562. tsung-hsien wen, milica gašić, nikola mrkšić, lina m. rojas-barahona, pei-hao su, david vandyke, and steve young. multi-domain neural network language generation for spoken dialogue systems. in proceedings of naacl-hlt, 2016b. 19 https://arxiv.org/abs/1604.04562 https://arxiv.org/abs/1604.04562 introduction related work a combined-context lstm for language generation model description comparison with the reference models training and decoding on-line interactive problem static case adversarial bandit case experimental study system comparison on-line adaptation evaluation human evaluation on-line adaptation using real asr data evaluation conclusion acknowledgements dialogue & discourse 10(1) (2019) 87–135 doi: 10.5087/dad.2019.104 how compatible are our discourse annotation frameworks? insights from mapping rst-dt and pdtb annotations vera demberg vera@coli.uni-saarland.de department of computer science department of language science and technology saarland university merel c.j. scholman m.c.j.scholman@coli.uni-saarland.de department of language science and technology saarland university fatemeh torabi asr ftorabia@sfu.ca department of linguistics, simon fraser university editor: barbara di eugenio submitted 07/18; accepted 05/19; published online 06/19 abstract discourse-annotated corpora are an important resource for the community, but they are often annotated according to different frameworks. this makes joint usage of the annotations difficult, preventing researchers from searching the corpora in a unified way, or using all annotated data jointly to train computational systems. several theoretical proposals have recently been made for mapping the relational labels of different frameworks to each other, but these proposals have so far not been validated against existing annotations. the two largest discourse relation annotated resources, the penn discourse treebank and the rhetorical structure theory discourse treebank, have however been annotated on the same texts, allowing for a direct comparison of the annotation layers. we propose a method for automatically aligning the discourse segments, and then evaluate existing mapping proposals by comparing the empirically observed against the proposed mappings. our analysis highlights the influence of segmentation on subsequent discourse relation labelling, and shows that while agreement between frameworks is reasonable for explicit relations, agreement on implicit relations is low. we identify several sources of systematic discrepancies between the two annotation schemes and discuss consequences for future annotation and for usage of the existing resources. keywords: coherence relations, discourse annotation, mapping, 1. introduction several sizeable discourse annotated resources have been created – most notably the penn discourse treebank (pdtb; prasad et al., 2008) and the rst treebank (rst-dt; carlson et al., 2003) for english – and more discourse annotation projects are currently under way in various languages (see, for example, oza et al., 2009; stede & neumann, 2014a). however, there is as of yet no consensus on a common discourse relation labelling scheme. existing discourse frameworks share c©2019 vera demberg, merel c.j. scholman and fatemeh torabi asr this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). demberg, scholman and asr basic notions of what a coherence relation is, and many of them make relation sense distinctions that are based on similar underlying ideas, but frameworks differ in how they define discourse relational arguments, in terms of constraints on resulting discourse structure (e.g., whether it has to be a tree), and in whether they are “lexically grounded” (like pdtb-style annotations), semantically driven (like sdrt), or contain a combination of semantic and intentional relations (rst-dt). this makes it difficult to study discourse relations across resources annotated according to different schemes, or across languages. in automatic discourse relation classifiers, it also limits the extent to which all available resources can be used effectively for training classifiers. this situation has long been recognized as a problem: in the early nineties, hovy & maier (1995) taxonomized the more than 400 relations that have been proposed in different frameworks as a hierarchy of roughly 70 discourse relations. more recently, several large initiatives have addressed this issue as well: the cost initiative textlink1 was aimed at organizing the properties of discourse relations and encouraging the use of a single taxonomy for subsequent discourse annotation, as well as for searches in existing corpora. in this context, some concrete proposals have been made for how discourse relations may be mapped onto one another (bunt & prasad, 2016; benamara & taboada, 2015; chiarcos, 2014; sanders et al., 2018). bunt & prasad (2016) developed an iso standard for coherence relations, in which they propose a new set of coherence relations that are central in many frameworks, and relate each of these to existing labels in other frameworks. the proposed set of labels can also be used for mapping labels to each other. in a similar line of work, benamara & taboada (2015) proposed a unified set of 26 discourse relations based on distinctions made by several styles of the rst and sdrt frameworks. they compared discourse relation labels between the frameworks based on their definitions and their overall frequencies of occurrence, but they did not have any data available that was annotated according to both frameworks. they therefore did not evaluate whether the actual annotations of the two frameworks would correspond to one another. benamara & taboada (2015)’s work highlights differences in granularity between frameworks by identifying certain labels that exist in one framework but do not have a corresponding label in the other framework. such relations with no correspondence across taxonomies need more consideration when using an intermediate framework. a possible solution for this mapping issue is to create an intermediate representation for mapping between frameworks, rather than creating a new framework. chiarcos (2014) was the first to attempt this: he developed an ontology to integrate rst-dt, pdtb and ontonotes annotations within a higher-level, more general framework. in this framework, the rst-dt and pdtb labels are assigned new labels with respect to the more general relation senses in both schemata. as part of a deliverable for textlink, sanders et al. (2018) created a different version of an intermediate representation. they worked out a mapping for discourse relations from various frameworks to a set of properties, such that each coherence relation can be described in terms of its properties or “dimensions” (such as “causal” vs. “additive”, “positive” vs. “negative”, and “subjective” vs. “objective”). through this intermediary representation in terms of properties, coherence relations can be mapped onto one another. as a community, we now find ourselves in a situation where several alternative proposals have been made for mapping coherence relation labels onto one another, but they have not been evaluated. these mappings were mainly proposed based on relation definitions, and we do not know whether alternative proposals for mapping relations are equivalent. moreover, we do not know how 1. http://textlink.ii.metu.edu.tr 88 http://textlink.ii.metu.edu.tr how compatible are our discourse annotation frameworks? well any of these proposals live up to the annotation of actual data. it is possible, for instance, that coherence relation labels should correspond to one another according to the annotation guidelines, but that differences in the operationalizations of frameworks (i.e., how annotators are asked to proceed for deciding on a relation label) lead to slightly different usages of these labels in actual annotations. annotation manuals are also often (necessarily) incomplete – they list clear cases or prototypical examples of certain types of coherence relations, but annotators learn through training and discussion how to deal with various kinds of difficult cases. the effect of such implicit knowledge on annotation practice may also lead to discrepancies between actual annotations compared to theoretically posited correspondences. a difference between annotations in the pdtb vs. rst-dt frameworks that runs even deeper is that rst-dt annotation aims to reflect in its relational annotations the intention of the author, whereas pdtb annotations focus on the logical semantic relation between discourse segments. rst theory explicitly distinguishes between intentional, semantic and textual relations in their definition of coherence relations. intentional relations express the writer’s communicative intentions: annotations relate to the writer’s goal or intended effect of each segment of a text with respect to the neighbouring segments, whereas semantic relations express information such as causality or temporal sequentiality. textual relations are linear relations (such as list and disjunction). the distinction between semantic and intentional levels has a long history in the discourse coherence community, and are also referred to as informational vs. intentional (moore & pollack, 1992), subject matter vs. presentational (mann & thompson, 1988), propositional vs. illocutionary (sanders & spooren, 1999), and ideational vs. interpersonal relations (hovy & maier, 1995). in this article, we will use the terms ideational and intentional relations. when searching for a discourse phenomenon across several corpora, it may therefore not be sufficient to rely on the theoretically mapped labels. instead, additional insights for which labels to consider or exclude can be gained from learning which labels correspond to one another empirically based on a large set of annotated instances. this article therefore also aims to elucidate the extent to which aspects of operationalization or training may affect labelling decisions during annotation. the pdtb 2.0 (prasad et al., 2008, 2014a) and rst-dt (carlson et al., 2003) corpora represent an excellent opportunity for addressing these questions, as they have been partly annotated on the same texts (meaning that there is overlap between the annotations of the corpora). we will therefore focus on the pdtb to rst-dt label mappings that have been proposed by various researchers, and compare them against the correspondences between the annotations of actual instances found in the corpora. in the ideal case, we would expect to find (i) that the different mapping schemes are consistent with each other, i.e. they propose the same set of equivalences between relations and (ii) that the proposed theoretical equivalences also hold for actual annotated data, i.e. if a given text is annotated according to two different schemes, and there is a mapping between these schemes, the actual annotations should correspond to one another as specified by the theoretical mappings. the current study thus extends previous work by mapping existing pdtb and rst-dt annotations onto one another, and by comparing them to theoretically posited correspondences between relations. this allows us to identify systematic differences between the annotations. such discrepancies could be caused by differences in the respective operationalizations in discourse annotations, or differences in annotation goals (annotating semantic relationships between discourse segments vs. annotating the intended communicative function of a segment with respect to another one). the empirical mapping approach taken here can also provide valuable insight for training future automatic discourse relation classifiers. automatic discourse relation classification has seen an increase in attention in recent years, with discourse relation labels having been shown to help improve 89 demberg, scholman and asr on down-stream tasks such as machine translation (meyer & popescu-belis, 2012; popescu-belis, 2016), question answering (jansen et al., 2014; sharp et al., 2015) and sentiment analysis (somasundaran et al., 2009; zhou et al., 2011; zirn et al., 2011). progress on this topic has been made possible through the large-scale annotation of text corpora such as the pdtb. a better understanding of how labels from different annotation schemes relate to one another may also help researchers to develop methods that can better exploit simultaneously the annotations from different corpora. first attempts at doing this using a multi-task setup have been proposed (e.g., liu et al., 2016), but these approaches could potentially profit from insight in how the annotations relate to one another, and what types of discrepancies there are. this article first provides background on the rst-dt and pdtb frameworks (section 2), and then proceeds to laying out the proposed mappings between rst-dt and pdtb 2.0 relations according to the three recent approaches which specified such mappings (chiarcos, 2014; bunt & prasad, 2016; sanders et al., 2018) in section 3. section 4 discusses challenges due to differences in discourse segmentation between rst-dt and the pdtb, and describes the alignment algorithm for mapping rst-dt annotations to the penn discourse treebank. results of the discourse relation label mapping are discussed in section 5, and compared to the theoretically posited correspondences. section 3.5 discusses the results from our mapping to a previous approach which used a similar methodology, albeit in a simplified setting and much smaller scale (rehbein et al., 2016). finally, we discuss implications for annotation as well as automatic discourse processing in section 6. our article makes the following contributions: • we propose a method for aligning rst-dt and pdtb 2.0 annotations. • we evaluate how well existing proposals for mapping discourse relation labels correspond to the mapping between existing rst-dt and pdtb 2.0 annotations. • we analyse how compatible rst-dt and pdtb 2.0 annotations are. • we identify sources of systematic discrepancies between annotations according to the two annotation schemes, and discuss their consequences for future annotation, corpus search, and the training of automatic discourse relation labellers. • we identify coherence relations for which human annotation is informative and beneficial, as well as cases for which it is unclear whether manual annotation is sufficiently consistent to be useful. • we provide an aligned discourse corpus where both pdtb 2.0 and rst-dt annotations can be queried simultaneously. 2. background in this section, we describe the notions underlying the two discourse relation annotation frameworks and their corresponding corpora that are mapped in this article, namely the pdtb 2.0 and the rstdt. this background will provide the necessary information for understanding the reasons behind differences in segmentation and discourse relation sense labelling that we find in our study. 90 how compatible are our discourse annotation frameworks? 2.1 rhetorical structure theory discourse treebank (rst-dt) the framework that is used to annotate the rst-dt (carlson & marcu, 2001) is based on the rhetorical structure theory (rst) as proposed by mann & thompson (1988). there are different implementations of rst annotation, including for example the basque rst treebank (iruskieta et al., 2013), the potsdam commentary corpus (stede & neumann, 2014b), and the cstnews corpus (cardoso et al., 2011). these corpora all follow the overall style of rst annotation, but may differ in how exactly they define discourse segments, what exact set of relation labels is chosen, and how nuclearity is interpreted or operationalized (c.f. stede, 2008). for the current study, we focus on the rst-dt style of rst. basic premises rst-dt is a descriptive theory of discourse relations, originally developed to guide computational text generation (taboada & mann, 2006). relations in rst-dt can be of semantic, intentional, or textual nature (carlson et al., 2003), and are explicitly classified into these three categories. a fundamental constraint on rst annotation is that each part of a text has to be included into the overall discourse structure, and that the discourse structure has to be arranged into a tree structure. a second essential characteristic of rst-dt is the assignment of nuclearity: texts spans are characterized as nuclei or satellites (every relation has to consist of at least one nucleus). the nucleus is the more central part of a relation in the text (with respect to its intentional discourse structure), while the satellite is supportive of the nucleus (see example (1-a)). some relations have symmetrically important arguments by definition. these relations consist of two nuclei rather than a nucleus and a satellite (see example (1-b)), and are referred to as multinuclear relations. (1) a. [but even on the federal bench, specialization is creeping in,]nucleus [and it has become a subject of sharp controversy on the newest federal appeals court.]satellite — elaboration-additional, wsj 0601 b. [that isn’t much compared with what bill cosby makes, or even connie chung for that matter (...)]nucleus [but the money isn’t peanuts either, particularly for a news program.]nucleus — contrast, wsj 0633 segmentation and annotation process in order to annotate a text in rst-style, each document is first decomposed into non-overlapping sequential text spans, called elementary discourse units (edus). edus generally consist of clauses, but attributions, relative clauses, nominal postmodifiers, and phrases that begin with a strong discourse marker are also considered edus in rst-dt (see carlson & marcu, 2001, p.3). after determining the edus of a text, nuclearity is assigned to the nodes and adjacent spans are linked together via rhetorical relations. in order to assign nuclearity, annotators should consider the writer’s intentions (i.e., what does the writer want to achieve?). determining nuclearity can therefore rarely be done without taking the context of the relation into consideration. nuclearity assignment is determined simultaneously with the assignment of a discourse relation (carlson et al., 2003). the discourse relations are linked recursively, thereby creating a hierarchical tree structure. rst-dt relations can hold between two (or more) non-overlapping text spans. the tree structure in rst annotations does not allow crossing or embedded edus. in order to deal with these limitations and the restriction that edus cannot overlap, a same unit tag was introduced in rstdt, which allows annotators to express that an edu is discontinuous. to illustrate this, consider 91 demberg, scholman and asr example (2) below. the clause when implemented (unit 2) is embedded in another clause. as a result, unit 1a cannot be connected to its other half, unit 1b, because unit 2 cannot be skipped. the tag same unit can be applied in this situation to express that the units 1a and 1b in fact make up one discontinuous segment. (2) ... [that it will,]unit 1a [when implemented,]unit 2 [provide significant reduction in the level of debt and debt service owed by costa rica.]unit 1b — same-unit, wsj 0624 annotators are instructed to annotate the writer’s goal of each segment of a text with respect to the neighbouring segments and the resulting hierarchical structure of the entire document. discourse structure and relational inventory carlson et al. (2003) distinguish 78 relation labels, partitioned into 16 classes that share some type of rhetorical meaning (see appendix a for a list of rst-dt’s relational inventory, and see the manual (carlson & marcu, 2001) for the definitions). the inventory is data-driven, based on analysis of the rst-dt corpus. rst-dt’s relational inventory can be divided into ideational and intentional relation labels: relations such as cause or restatement belong to the ideational group of relations, evaluation or purpose belong to the intentional group (hovy & maier, 1995). some rst-dt classes contain relations that are not considered to be coherence relations in other approaches such as pdtb 2.0; examples include the attribution relations (which is also annotated in pdtb, but not considered a coherence relation) and cohesion relations (cases where coherence between discourse segments is not achieved through a specific coherence relation, but rather through cohesion). such cases are annotated as elaboration relations in rst-dt. pdtb treats them as a separate type of discourse relations (namely entrel). relational definitions of rst-dt’s classes are based on functional and semantic criteria, and not on signals, because the creators argued that no unambiguous signal for any relation was found (taboada & mann, 2006); e.g., a connective such as but can mark different types of relations. 2.2 penn discourse treebank (pdtb)-style annotation the pdtb 2.0 corpus (prasad et al., 2008) is the largest manually annotated discourse relation corpus available at the moment. the framework that was used to annotate the corpus is referred to as pdtb as well. the framework has also been used to create new corpora in other languages and genres, such as arabic (al-saif & markert, 2010), italian (tonelli et al., 2010), and chinese (zhou & xue, 2015). this has resulted in different styles of pdtb annotation, but they can be considered to be interoperable (cf. prasad et al. 2014b). here, we use the term pdtb to refer to the framework that corresponds to the original pdtb-style annotation. the pdtb research group is set to release an enriched, enlarged and modified version of the corpus (pdtb 3.0, prasad et al. 2018), for which they have also adapted the framework. compared to the pdtb 2.0, pdtb’s 3.0 relational hierarchy has been simplified and extended. for example, in pdtb 3.0, the types conjunction and list are merged, subtypes of condition and contrast are removed, and relation types such as manner, purpose and negative condition are added. the level-3 (subtype) senses are now restricted to differences in directionality. regarding the annotation procedure, the pdtb 3.0 takes a more systematic approach to annotating multiple labels for a single relation (webber et al., 2016). 92 how compatible are our discourse annotation frameworks? basic premises pdtb-style annotation (prasad et al., 2007, 2008, 2014a) is characterized by two basic premises. first, it has a theory-agnostic approach to annotation, thereby making no commitment to what kinds of high-level structures may be created from the low-level annotation of relations (prasad et al., 2008). second, pdtb follows a lexically-grounded approach to discourse relation representation, meaning that they focus on annotating lexical items that can signal discourse relations. pdtb distinguishes between explicit and implicit discourse relations. explicit relations are marked with a coordinating conjunction, subordinating conjunction or a discourse adverbial, which we will jointly refer to as discourse connectives in this article. implicit relations, on the other hand, are not marked with a discourse connective. instead, annotators are asked to insert a connective they think would best fit, and annotate the coherence relation with the inserted connective. three additional labels were employed for marking cases where an implicit connective could not be inserted. altlex (alternative lexicalization) applies to coherence relations for which insertion of a connective leads to a perception of relation reduncancy (e.g., because including the connective therefore would sound odd and be redundant in a sentence starting with for this important reason). entrel is not a coherence relation as such, but rather marks cases which are connected only through cohesion rather than a specifiable discourse relation. norel is used when no discourse or entity relation holds. this label is necessary because of the way that annotations for implicit relations in pdtb was performed: annotators were asked to label all adjacent sentences (see also next section); sometimes adjacent sentences may however not stand in a direct relation to one another, because they belong to two different larger discourse segments. for these cases the norel label could be assigned. segmentation and annotation process relations in the pdtb have two and only two arguments, referred to as arg1 and arg2. these arguments can be continuous or discontinuous. in the case of explicit relations, the argument that is syntactically bound to the connective is labeled as arg2; the other argument is arg1, and may be adjacent or non-adjacent with arg2. annotators were instructed to first identify explicit connectives based on a list of discourse cues; they then identified the discourse relational arguments. the selection of these arguments is restricted by the “minimality principle,” according to which only as much material should be included in the argument as is minimally required and sufficient for the interpretation of the relation. any material that is relevant but not “mininally necessary” for interpretating the relation is marked as supplementary information. supplementary material is annotated for approximately 4% of pdtb 2.0 relations. after identifying explicit connectives and their arguments, a relation label is assigned, as in example (3-a). in pdtb 2.0, implicit discourse relations have only been annotated between adjacent sentences within paragraphs, as well as between complete clauses delimited by a semi-colon (“;”) or colon (“:”) (see also prasad et al., 2017). in a first round of annotation, a connective was inserted, and the relation label was then assigned in a subsequent step, see example (3-b). because arguments of implicit relations have often been annotated as complete sentences or clauses of sentences with colons or semi-colons, the annotations have a slightly different pattern in segmentation for implicit relations compared to explicit relations (this observation will become important for the design of the mapping algorithm). (3) a. although [that may sound like an arcane maneuver of little interest outside washington],arg2 [it would set off a political earthquake].arg1 — comparison.concession.expectation, wsj 0609 93 demberg, scholman and asr b. [mr. carpenter denies the speculation]arg1 [to answer the brokerage question, kidder, in typical fashion, completed a task-force study].arg2 — temporal.synchronous, wsj 0604 discourse structure and relational inventory the framework distinguishes 43 relation labels (see appendix a for the labels, and prasad et al. 2007 for their definitions). these labels are organised in a hierarchy consisting of three levels: (i) class is the top level, which contains the four major semantic classes; (ii) type is the second level, which further refines the semantics of the class levels; and (iii) subtype is the most fine-grained level, which defines the semantic contribution of each argument. when an annotator was uncertain of the more fine-grained senses of subtype, s/he could choose the higher level type, which was also beneficial for inter-annotator agreement (prasad et al., 2008). the pdtb taxonomy contains a few pragmatic labels (e.g., contingency.pragmatic cause), but the focus of pdtb is on annotating ideational relations, rather than interpersonal relations. the frameworks differs in this respect from rst-dt. another important aspect distinguishing pdtb 2.0 annotations from rst-dt annotations is that pdtb annotators were allowed to assign several labels to the same relational arguments, when they found that multiple concurrent discourse relations held between the arguments. this is important for our evaluation of correspondences between assigned relation labels later on, as we decided to evaluate the correspondence in terms of the pdtb label that is most similar to the rst-dt label. in terms of discourse structure, an important difference between pdtb-style annotation and rst-style annotation is that in pdtb-style annotation, it is not necessary for all parts of a text to be connected in an overall discourse structure. in fact, the minimality principle leads to many partial sentences not being part of any discourse relational argument. furthermore, the restriction of only annotating implicit relations between adjacent sentences in pdtb 2.0 means that sentence-internal or cross-paragraph coherence relations may be missed (this was addressed in pdtb 3.0; also see a recent extension of pdtb annotation to vps, webber et al. 2016). the bottom-up annotation of pdtb based on explicit connectives and adjacent sentences also entails that there is no guarantee that pdtb annotations follow a tree structure. in fact, lee et al. (2006) report that discourse structure is a lot more variable than syntax, exhibiting nested, crossed and other “non-tree like” configurations. lee et al. (2008) extends this analysis by focussing on cases where two different coherence relations share relational arguments of the form { x conn1 ( y } conn2 z), where x, y and z are discourse relational arguments and conn1 and conn2 are connectives; they come to the conclusion that such non-tree structures are quite common in discourse. 3. theory-based proposals for mapping rst-dt and pdtb 2.0 relations a mapping of discourse relation labels between different frameworks can be achieved by determining the correspondence of relation labels from one framework to the other directly (e.g., pdtb 2.0 to rst-dt, rst-dt to sdrt, sdrt to pdtb 2.0). another option is to map all frameworks to an intermediary representation, such that the mapping between any two frameworks can be obtained via this intermediary representation. this approach has the advantage of being more general in case many frameworks should be mapped onto one another: rather than creating a new mapping between the new framework and all other frameworks, researchers only have to create a single mapping from the new framework to the intermediary framework. another advantage is that it produces a 94 how compatible are our discourse annotation frameworks? figure 1: simplified illustration of olia’s hierarchical mapping approach. candidate for a single set of relational labels potentially suitable for future annotation: the intermediary relation representation. all three of the mapping approaches discussed below propose such an intermediary representation. the mapping approach by benamara & taboada (2015) could not be included here, because it focused on comparing rst and sdrt and has not yet completed the mapping to pdtb labels.2 3.1 mapping according to the olia reference model chiarcos (2014) mapped the pdtb and rst-dt schemes onto each other as part of the ontologies of linguistic annotation (olia,). olia provides a terminology repository that can be used to facilitate the conceptual interoperability of annotations (see appendix b or the olia website3 for more details). this is done using an intermediate level of representation that mediates between several existing frameworks. the intermediate representation is formalised as subclassof descriptions. to illustrate this, consider figure 1, which illustrates a hierarchical mapping of pdtb’s condition. in olia, this relation type is characterised as a subclass of semantic condition relations, which is in turn a subclass of condition relations. this can then be mapped onto rst-dt’s class of condition, which has the same superclasses. chiarcos (2014) argues that ontologies are able to represent more fine-grained nuances of meaning, and to quantify the number of shared descriptions between annotations of different frameworks (chiarcos, 2014). the table in appendix f shows the proposed correspondences (indicated by an ‘o’) between pdtb 2.0 and the rst-dt according to the proposal in olia. 3.2 mapping via unifying dimensions the unifying dimensions mapping (unidim) was proposed by sanders et al. (2018) with the goal of mapping labels from different frameworks onto each other using an interlingua (see appendix c). in the unidim proposal, relation labels are not mapped to intermediate labels; rather, they are described in terms of their characteristics, or values on certain dimensions. the set of unifying dimensions is an extended version of the dimensions originally proposed as the cognitive approach to coherence relations (ccr; sanders et al., 1992). the original ccr distinguishes four cognitive dimensions that apply to every relation, namely polarity, basic operation, source of coherence, and order of the segments. for example, a reason relation would be represented as a relation with positive polarity, causal basic operation, objective source of coherence, and backward order of the segments. as pdtb and rst-dt make some distinctions which cannot be represented in terms of only these four dimensions, sanders et al. (2018) extended ccr to account for more fine-grained properties of relations. 2. personal communication november 2017 3. http://www.acoli.informatik.uni-frankfurt.de/resources/discourse/ 95 http://www.acoli.informatik.uni-frankfurt.de/resources/discourse/ demberg, scholman and asr the intermediate representation in terms of these dimensions allows for mapping relation labels from one framework into the representation as dimensions, and from this representation to the second framework. the method of describing relations in terms of their characteristics makes it easier to identify similarities and differences between relations. for example, similarities between relations can be described in terms of how many of their characteristics are identical: cause and concession differ in one dimension (polarity) but they are both types of causal, objective, nonconditional relations (in the case of concessions, the expected result has not occurred, or a result occurred even though the usual cause was not present). appendix f shows the proposed correspondences (indicated by a ‘u’) according to the proposal by sanders et al. (2018). 3.3 mapping according to the iso standard proposal bunt & prasad (2016) provide a mapping of frameworks that is based on a different system. they proposed an international standard (iso standard) for coherence relation annotation, which consists of a set of 20 core relations that are commonly found in some form in existing approaches (see appendix d). they did not aim to provide a fixed and exhaustive set of coherence relations; rather, they aimed at providing an open, extensible set of relations. bunt & prasad (2016) propose that the iso standard can be used for future annotation efforts, as well as for mapping between annotations using different frameworks. to this end, they provided mappings of these iso relations to other frameworks, including pdtb and rst-dt. from these proposed correspondences to iso standard candidate relations, we can infer how the pdtb 2.0 and rst-dt relations correspond to one another. the full table of hypothesized correspondences according to the iso proposal are marked by ‘i’ in appendix f. 3.4 discussion of agreement and discrepancies between proposed mappings appendix f shows the grid of proposed correspondences between the schemes. while we can see many cells that include ‘o’ , ‘u’ and ’i’ (olia, unidim, and iso, respectively), indicating that all approaches agree that these relation labels should correspond to one another, we can also see some relational labels that have only been linked by one of the proposals. we have identified three main reasons for these discrepancies, which we discuss below: differences in granularity of mapping schemes, differences in how concepts (in particular, order and subjectivity) are defined by the frameworks, and differences in the interpretation of definitions in the annotation guidelines. granularity of the proposals the first source of discrepancies is the granularity of the intermediary schemes. in principle, we would like to obtain a one-to-one mapping between the labels of one framework to another one. however, this is impossible if one framework makes more fine-grained relational distinctions than the other one, or if distinctions between relations don’t correspond to one another. in this case, a one-to-many or many-to-many mapping will be necessary. from the perspective of desigining and intermediary mapping scheme, it is methodologically most preferable to have a scheme that does not “conflate” several relational labels, i.e. two different labels from a source scheme should never be mapped onto the same intermediary label. when designing a mapping scheme that interfaces with several frameworks at the same time, we necessarily obtain a one-to-many mapping from source framework to intermediary representation, and many-to-one mapping from intermediary representation to target framework. mapping schemes like iso, that 3. mappings from the paper for certain labels were updated; personal communication january 2018. 96 how compatible are our discourse annotation frameworks? was not designed for mapping but as a new set of relation labels, may already employ a many-tomany mapping from the source framework to the intermediary iso representation, and hence only allow for a “coarser” mapping. for instance, the causal relation labels in the iso-based mapping are coarser than the distinctions in rst-dt and pdtb, because the proposed iso standard doesn’t differentiate between semantic and pragmatic causal relations (a distinction also known as objective vs. subjective or content vs. epistemic); i.e. causal relations where the link can be established based on the semantics of the two arguments and causal relations where the author posits a causal link. olia and unidim do distinguish between semantic and pragmatic relations, and therefore do not map certain rst-dt labels that the manual mentions are semantic (e.g., reason and explanation-argumentative) to pdtb’s pragmatic causal label justification, whereas iso does map semantic rst-dt labels to justification. an example of a relational category that is more fine-grained in the iso proposal than in the other frameworks is exception, which corresponds to pdtb’s exception, but has no equivalent mapping to an rst-dt relation. as a result, pdtb’s exception cannot be mapped to a corresponding rst-dt label based on the iso proposal. in unidim, however, exception is mapped to contrast, antithesis and preference, i.e. a set of more general labels with similar characteristics to pdtb’s exception. the same goes for rst’s means relations. certain labels can also not be mapped by olia. for example, rst-dt’s evaluation, comment and definition relations are part of the superclass assessment, which doesn’t occur in the pdtb inventory; furthermore, the relations background and circumstance are part of the superclass background, which also doesn’t occur in pdtb. olia also doesn’t map any comparison rst-dt relations. labels could be mapped using their supersuperclass, but this would not be very informative (e.g., it would result in mapping the background label to all pdtb expansion labels). in temporal relations marked with connectives such as before and after, pdtb assigns the label temp.async.succession to relations marked with after, and the label temp.async.precedence to relations marked with before as a subordinating conjunction. similarly, rst uses the labels temporal-after and temporal-before for these instances. unidim and olia therefore proposed to map temp.async.succession onto temporal-after and temp.async.precedence to temporal-before. iso on the other hand more generally maps both of pdtb’s temporal.async relations onto all asynchronous temporal relations in rst. discrepancies in granularity occur because of differences in the goals of the mapping frameworks, and different systems of mapping. unidim was designed to be able to map labels, and, as a result, the interlingua can be used to map relational labels even when there is no direct equivalent. iso, on the other hand, was proposed as a new set of relations; i.e. mapping was not its primary goal. the mapping that is provided in bunt & prasad (2016) mainly includes direct correspondences. olia also has the goal to identify all correspondences between frameworks, but because of the hierarchical structure of the interlingua, a relation label that is part of a unique superclass in one framework can often not be mapped properly to another framework. subjective vs. epistemic relations a somewhat difficult issue is the notion of subjectivity. a relation can be objective, subjective or epistemic. rst distinguishes between objective and subjective relations with labels such as cause and result on the one hand and reason, evidence and explanation-argumentative on the other hand, while pdtb distinguishes between non97 demberg, scholman and asr rst-dt label mapping to pdtb according to: olia unidim iso comparison not mapped conjunction contrast antithesis contrast concession concession & contrast elab.-objectexpansion specification entrel attribute & generalization background not mapped conjunction entrel & asynchronous circumstance not mapped conjunction, synchronous synchronous & asynchronous table 1: overview of the differences between the olia, unidim and iso-based mappings caused by different interpretations of definitions. epistemic relations contingency.cause.reason and contingency.cause.result and epistemic relations such as contengency.pragmatic cause.justification. olia proposed to map rst-dt’s subjective labels such as evidence label only to pdtb’s epistemic label pragmatic cause, which does however not do justice to pdtb’s more restrictive notion of subjectivity as epistemic relations only. a similar problem occurs for rst-dt’s manner, means, problemsolution and enablement. difference in interpretations of definitions the third source of discrepancies is rooted in different interpretations of relational definitions in the annotation manuals, and hence represents the theoretically most interesting case for comparison between proposals. in the remainder of this section, we will focus our discussion on these cases. table 1 provides an overview of the differences between the olia, unidim and iso-based mapping according to this type of discrepancies. rst-dt’s comparison first, the proposals differ in their mapping of rst-dt’s comparison relations. the manual states that the two segments of a comparison relation are not in contrast with each other (carlson & marcu, 2001, p. 50). based on this description, unidim mapped comparison to pdtb conjunction. iso, however, mapped comparison to pdtb’s contrast relational class. finally, olia mapped rst-dt’s comparison to the superclass non-contrastive comparison. however, none of the labels in the pdtb have a superclass non-contrastive comparison, and therefore rst-dt’s labels do not have a correspondence in pdtb. the mapped data will be able to indicate how the comparison label was used in practice rst-dt’s antithesis the frameworks also disagree on the mapping to contrastive and concessive relations. the former is a relation of semantic opposition; the latter contains a denial of expectation. to illustrate the difference between the two types of relations, consider the following examples: 3. differences that are not discussed in this section include for instance: unidim’s mapping of rst-dt hypothetical to pdtb pragmatic condition (in addition to condition); iso’s mapping of rst-dt evaluation to the general expansion class; olia’s mapping of rst-dt summary relations to equivalence (in addition to specification and generalization). 98 how compatible are our discourse annotation frameworks? (4) dylan used to live in washington dc, but now he lives in baltimore. (5) dylan lives in baltimore, but he works in washington dc. example (4) presents a simple contrast between where dylan used to live and where he lives now. the relation could also have been expressed by the connective whereas, which is a typical marker of contrastive relations. in example (5), the first segment informs you where dylan lives now. this segment presupposes an implicit expectation that dylan also works there, but the second segment denies this expectation: dylan works in a different city. this is typical of concession relations: one argument creates an expectation of a cause or consequence, which is denied by the other argument. the relation could also have been expressed with typical markers of concessive relations, such as even though or nevertheless. the distinction between contrast and concession is relatively difficult to make, even for trained annotators (see, for example, robaldo & miltsakaki, 2014; zufferey & degand, 2013). the frameworks all disagree on the mapping of the rst-dt label antithesis to pdtb contrast and concession. rst-dt’s annotation manual states that antithesis is a contrastive relation, but some of the examples that are provided are concessive relations. in unidim, antithesis is therefore mapped to both contrastive and concessive pdtb relational labels, whereas in iso, antithesis is mapped only to concessive labels. in olia, antithesis is mapped onto contrast only. rst-dt’s elaboration-object-attribute the proposals differ in their mapping of rstdt’s elaboration-object-attribute label: olia maps it to general expansion relations, unidim maps it to pdtb’s specification and generalization relations, while the iso-based proposal maps elaboration-object-attribute to entrel. rst-dt’s background and circumstance finally, unidim and iso differ in their treatment of rst-dt’s background and circumstance. iso maps background to pdtb entrel, whereas unidim maps background to conjunction and temporal.asynchronous based on the description of background in the manual: “the satellite is not the cause / reason / motivation of the situation presented in the nucleus ... the events represented in the nucleus and the satellite occur at distinctly different times” (carlson & marcu, 2001, p. 47). regarding circumstance relations, unidim and iso agree on mapping these to synchronous relations, but unidim also maps to asynchronous and conjunction labels. the disagreements that are attributable to different interpretations of the definitions in the manual will be systematically analysed with respect to actual annotations, and discussed in section 5, in order to determine which of the proposed correspondences are justified by the actual data. these results can then also be used to clarify annotation guidelines for future research, and highlight which definitions are particularly susceptible to inconsistent annotation. 3.5 previous work on empirical evaluation of mapping coherence relation annotations we are aware of three other efforts (rehbein et al., 2016; scheffler & stede, 2016a; polakova et al., 2017) to systematically evaluate the mapping of discourse annotations from different frameworks on the same text. polakova et al. analyse the same resource as us: the pdtb and rst-dt corpora. their study focuses on the question of how implicit relations are signalled. to this end, they identify a subset of 472 implicit relations in pdtb that have same argument spans and matching labels in 99 demberg, scholman and asr the rst annotation, and analyse what types of additional signals are present in these relations based on the annotations of such signals in the rst signalling corpus (das & taboada, 2018). polakova et al. (2017) find that a large proportion of the pdtb implicit relations are signalled by semantic signals expressed in specific lexical chains in the relational arguments. another frequent pattern observed among these implicit relations were parallel syntactic constructions between relational arguments, as well as the “unsure” label. they conclude that implicit relations cannot be easily annotated automatically based on signals such as the ones annotated in the rst signalling corpus, as they are hard to identify and the subset of semantic lexical chains often falls out of well-defined semantic relations such as synonymy, antonymy etc. scheffler & stede (2016b) compared pdtb 3.0-style annotations and rst-style annotations on the german potsdam commentary corpus (pcc; stede & neumann, 2014b). the pcc includes 1104 explicit connectives annotated in pdtb-style but lacks pdtb-style annotations for implicit relations. scheffler & stede (2016b) propose a simple method for mapping the pdtb 3.0-style and rst-style annotations for explicitly marked relations in the potsdam commentary corpus (pcc) onto one another, in order to empirically observe commonalities and differences in the annotations of discourse structure between the two approaches. they do not, however, compare relation labels for the mapped relations. the alignment algorithm used in scheffler & stede (2016b) compares the spans of the relation argument annotations, and distinguishes different segmentation constellations. they observe that the majority (84%) of instances in their corpus consists of cases that are easy to map, including exact match of discourse relational arguments (41%) and “boundary match” (39%), where a relation is annotated between two adjacent text spans, and the boundary between the two arguments is identical. they however also report difficult cases of non-local relations where the segment boundaries differ (13%) (as in the example of the contrast relation in figure 3 below), and cases in which no match is possible (3%). as we will discuss in more detail in section 4, we also find segmentation and alignment to be a challenging first step in aligning the annotations of the pdtb and rst-dt corpora. rehbein et al. created an english corpus of spoken discourse containing pdtb 3.0 and ccr annotations for every relation. they segmented the texts first, and then proceeded to assign sense labels according to both schemes for that given segmentation. this procedure thus avoided challenges related to differences in segmentation between frameworks. after annotation, the relation labels were mapped onto one another directly, and the correspondence between pdtb and ccr annotations was evaluated. rehbein et al. (2016) reported three systematic biases introduced in the operationalizations of pdtb 3.0 and ccr, which lead to differences in annotations in some areas: a first observation holds that pdtb’s additive relations expansion.instantiation, expansion.restatement.specification and expansion.restatement.equivalence were quite often (30% of relations) annotated as causals in ccr. this finding is consistent with an observation we made in the present analysis, where we found that pdtb expansion.instantiation and expansion.restatement are often annotated with a causal label in rst, see section 5.3. as noted by blakemore (1997) and carston (1993), and shown in scholman & demberg (2017), these types of discourse relations are often ambiguous, as examples can at the same time also serve as evidence for a claim. the second category of systematic disagreements concerns comparison.contrast and comparison.concession relations: among the negative relations, annotators often disagreed on the causal vs. additive basic operation. this was partly due to a slightly different definition of what con100 how compatible are our discourse annotation frameworks? stitutes a concession, but note that distinguishing between contrastive and concessive discourse relations is a well-attested difficulty (see, for example, robaldo & miltsakaki, 2014; zufferey & degand, 2013). again, the same difficulty is obvious in the mapping between rst-dt and pdtb, as discussed in section 5.1. as a third pattern of disagreements, rehbein et al. (2016) report effects of operationalization of annotation procedures: some instances marked by but were annotated as positive polarity relations in pdtb, but as negative in ccr (including instances marked with but also). these discrepancies were systematic and due to an annotation instruction – as a rule, all relations that can be marked with but are annotated as negative polarity relations in ccr. while this specific pattern is not relevant for the mapping between pdtb and rst-dt, we note that annotation operationalizations by different frameworks might have substantial effects on the annotations. 4. data, segmentation and automatic alignment pdtb 2.0 and rst-dt annotations overlap for 385 newspaper articles in sections 6, 11, 13, 19 and 23 of the wall street journal corpus. the annotation of the rst-dt involved more than a dozen of people and several phases of revision. the average inter-annotator agreement (final results for 6 taggers) on span detection, nuclearity assignment and relation sense annotation was 86.8%, 80.7%, and 72%, respectively (carlson et al., 2003).4 the pdtb 2.0 reports an inter-annotator agreement of 94%, 84%, and 80% for the class, type and subtype levels respectively, and pdtb’s discourse segments were identified with an agreement (exact string match) of 90.2% for explicit relations and 85.1% for implicit relations (prasad et al., 2008). our investigation will be based on the intersection of the pdtb and the rst-dt, with annotations from both frameworks included as different annotation layers. 4.1 segmentation comparing annotations of the two corpora is not a trivial task, because annotations not only differ in the label sets that were used, but also in terms of segmentation. firstly, there are discrepancies in what is considered an ”elementary discourse unit” in rst-dt vs. what is considered a discourse relational argument or an attribution in pdtb 2.0. note that in the remainder of this paper, we use the term “segment” to refer to the text elements that are part of a relation; that is, the arguments in pdtb and the edus, nuclei and satellites in rst-dt. there are also differences in the discourse structure: rst annotates discourse trees spanning the whole document, while pdtb 2.0 only annotates relations between adjacent sentences and relations marked by an explicit connective. this results in a considerably lower number of pdtb relations than rst-dt relations for the same text. we therefore use pdtb relations as a starting point in alignment, with the goal of identifying for each pdtb relation the corresponding relation label in the rst annotation. pdtb’s minimality principle (cf. section 2.2) and rst’s tree structure (cf. section 2.1) influence the result of the segmentation and annotation steps. an annotation alignment process must therefore take into account systematic differences arising from the respective segmentation procedures. in the automatic alignment step, our goal is to map as many discourse relation labels as possible in order to get a maximally complete picture regarding how well the annotations correspond 4. these numbers are the result of averaging over the inter-annotator agreement scores reported for every two annotators in table 2 of carlson et al. (2003). 101 demberg, scholman and asr figure 2: pdtb and rst-dt annotations for a paragraph of wsj 1172. 1 refers to arg1 in pdtb; 2 refers to arg2. n refers to nucleus in rst; s refers to satellite. (a-d) refer to rst-dt’s edus. to one another. at the same time, we must only map those labels where annotators inferred the same relation – if the rst-dt annotators annotated a relation holding between two text segments, and the pdtb annotators marked a relation between two different segments, these labels should not be recorded as valid alignments, and labels hence shouldn’t be compared. to illustrate this, consider example 2, which presents the pdtb (left) and rst-dt (right) annotations for a fragment of a wall street journal article. pdtb segmented arg1 differently than rst-dt, leading to a difference in interpretation. pdtb considers segments (c-d) as the result of the event in segment (b), whereas rst-dt focused on the different opinions expressed in segments (a-b) and (c-d). the disagreement between the two labels (result vs. list, respectively) does not stem from annotator disagreement regarding the label, but from a more fundamental difference in segmentation. such cases should therefore not be included in the evaluation of mapped labels. the alignment algorithm proposed in this article (see section 4.2 below) is more complex than the one proposed in scheffler & stede (2016b), in order to better address those cases for which there are differences in segmentation between annotation layers. the core idea of how valid alignments can be identified even in the face of mismatches between relational arguments builds on the strong nuclearity hypothesis (marcu, 2000), which was used for rst-dt annotation. the strong nuclearity hypothesis states that when a relation is postulated to hold between two spans of text, it should also hold between the nuclei of these two spans. note that the notion of “nuclearity” has had several slightly different interpretations throughout the conception and further development of rst, as laid out in stede (2008). independent of the theoretical discussion about the intentions behind nuclearity as such, the specific notion used in rst-dt annotation is helpful for determining relation alignment. rst-dt’s nuclearity was assigned in an instance-by-instance decision for identifying those parts of a discourse relation which are crucial for that relation to hold; it was not used as a general property of relation types, as in some other variants of rst annotation. in that sense, the strong nuclearity annotation guideline from marcu (2000) is related to the minimality principle used in pdtb 2.0 annotations: both help to indicate which segments of the text are central to establishing the discourse relation. to illustrate this, consider the contrast relation in figure 3; the strong nuclearity hypothesis means that if the relation holds between (6-7) and (8-9), it should also hold between (6) and (8), but not between (7) and (8) or (7) and (9). in the following, we will use the expression nucleus path to refer to the path between a complex argument of a high-level relation, and the single edu which one ends up with if always following the path down the segments annotated as nucleus. 102 how compatible are our discourse annotation frameworks? figure 3: rst discourse structure, figure 1.1 from marcu (2000). figure 4: pdtb and rst annotations for a section of wsj 1146. 4.1.1 segmentation constellations we will now go through the different segmentation and alignment constellations using examples, to explain where challenges in alignment lie, and how these are dealt with by our alignment procedure. for ease of reference, we will here adopt the pdtb distinction between explicitly marked and implicit relations, even when referring to rst-dt annotations. pdtb relations with adjacent arguments the simplest case is an exact match between the discourse relational arguments for the two annotation layers. additionally, there can be cases where the argument spans largely overlap but differ in their exact boundaries. consider figure 4: in pdtb 2.0, segments (a-b) and (c) are connected in a temporal.synchrony relation. in rst-dt, a temporal-same-time relation was annotated between segments (b) and (c). even though the spans differ in whether (a) is included, they clearly correspond to each other. as neither of the discourse relational arguments of the synchrony relation in pdtb nor the temporal-sametime relation in rst-dt is complex (i.e., no other relations are embedded under either of their arguments), it is straightforward to decide which relation labels should correspond to each other. pdtb relations with non-adjacent arguments we also frequently encounter more complex cases, where the pdtb arguments are not directly adjacent to one another. this can happen both for explicitly marked relations and for implicit relations (when the sentences are adjacent but the chosen spans do not cover the complete sentence). whenever the pdtb arguments are not adjacent, we will either have a mismatch between the size of the discourse segments in that the rst-dt edu 103 demberg, scholman and asr is larger than the pdtb argument, or in that the rst-dt argument is complex, i.e. it consists of other relations. in figure 4, this is the case for the restatement relation: in pdtb, this relation holds between segments (a) and (d), whereas in rst-dt, the relation holds between segments (a-c) and (d). for deciding whether the pdtb expansion.restatement and rst-dt restatement relations should be aligned, we rely on the strong nuclearity hypothesis. it says that the complex relation between (a-c) and (d) should also hold between the nucleus of (a-c), hence (a) and (d). we can then infer an exact match between discourse relational arguments (a) and (d) between the two annotation layers, and map the labels onto one another. such cases also occur among explicitly marked relations. since rst-dt relations are annotated in a hierarchical tree structure, relations connecting non-adjacent sentences will have large discourse relational arguments. consider example 3 again: rst-dt annotates a contrast relation with segments (6-7) as one nucleus and (8-9) as the other nucleus. in the pdtb annotation, because of the minimality principle of marking discourse relational arguments, the relation marked by but has segment (8) as its arg2 and segment (6) as its arg1. similarly, the evidence relation between segments (2-3) and (4-9) would differ in terms of its argument boundaries in pdtb annotation, as segments (6-9) would typically not be included in the arg2 of the implicit relation. nevertheless, an alignment of relations is possible given nuclearity annotation, and labels from such cases are included in the mapping. relations with inconsistent nuclearity there are however also cases for which labels should not be mapped due to discrepancies in what relation the annotators intended to label. these cases can typically be identified through inconsistencies between the discourse relational arguments annotated by pdtb and the nuclearity assignment annotated in rst-dt. to illustrate this, consider the passage in figure 5. in this example, one would have to map pdtb relation contrast to rst-dt’s consequence, if one were to only take into account maximal overlap of discourse segments, but ignore nuclearity: the difference in span size (pdtb’s annotation excludes segments (a) and (e)) affects the interpretation of the relation. in pdtb’s annotation, the relation holds between the state’s action and the farmers’ actions. in rst-dt’s annotation, on the other hand, the annotated relation holds between the state’s action and the consequences of that action. the relation labels therefore correspond to different interpretations. such cases are flagged automatically because pdtb’s arg2 cannot be safely traced to the nucleus of the satellite of rst-dt’s relation, due to an intervening multinuclear relation. in order to investigate how often relations with intervening multinuclear relations (i.e., relations in which one of the segments consists of a larger tree branch including a multinuclear relation) occur in the data and to what extent they pose a problem for automatic alignment, we extracted all instances that contain an intervening multinuclear relation. in total, 892 relations (13% of the data) have one or more intervening multinuclear relations. however, not all of these pose a potential risk to alignment – cases where the multi-nuclear relation is the relation to be compared, and multinuclear relations that are not on the nucleus path, are safe to map. we found that 295 multinuclear cases were flagged as potentially violating strong nuclearity. we manually checked 50 of these flagged instances and found that almost all of them indeed do not represent valid alignments, and should therefore be excluded from the mapping analysis. internal relations the segmentation granularity between the two frameworks can differ, which can lead to a specific instance from the finer-grained framework not being mapped to a label in the coarser framework. for example, two rst-dt edus could occur internally within a sentence, 104 how compatible are our discourse annotation frameworks? figure 5: pdtb and rst-dt annotations for a paragraph of wsj 1146. note: only the pdtb annotation that is relevant for this example is included. figure 6: pdtb and rst-dt annotations for a paragraph of wsj 1962. without this relation being annotated in the pdtb annotation layer. figure 3 shows an example where the background relation between segments (2) and (3) has no corresponding pdtb annotation. similarly, there are also cases where a connective is annotated as an explicit marker for a relation in the pdtb 2.0 annotation, but both arguments of this relation are part of the same rst edu, as in example (6): pdtb annotated the connective because whereas in rst-dt, the entire sentence was considered as one edu. internal relations can inherently not be aligned, and hence were excluded from our mapping analysis. (6) [that’s] because [municipal-bond interest is exempt from federal income tax – and from state and local taxes too, for in-state investors.] — contingency.cause.reason, wsj 0689 rst-dt’s same-unit centrally embedded relations (where the arg1 is discontinuous and arg2 is located inside the span of arg1) can be successfully mapped by our algorithm if the rst-dt annotation contains a same-unit annotation that links the discontinuous parts of the rst segments mapping to the pdtb arg1. this is illustrated in figure 6. if a same-unit annotation in rst-dt is identified, the label of the corresponding pdtb relation is mapped to the label of the relation below the same-unit relation (96 instances). we manually verified a subset of mapped relations to make sure that this heuristic is valid in practice (see section 4.3). for same-unit relations where both of the rst-dt segments contain multiple edus (44 instances), the corresponding relation could not be determined automatically in a reliable way. these cases are flagged and excluded from the mapping analysis. cohesion the pdtb relation label entrel is used to mark cohesion when no specific coherence relation can be identified. as cohesion is not defined as a type of coherence relation, the 105 demberg, scholman and asr entrel label is not included in the unidim mapping. however, rst-dt contains several labels (elaboration-additional, definition, background) that tend to be annotated when there is cohesion but no easily identifyable coherence relation. in our analysis, we therefore decided to include the entrel label, to check whether it actually coincides with rst-dt’s labels signalling cohesion. 4.2 alignment algorithm our procedure for mapping annotations provided by the two frameworks takes the pdtb relations as a starting point, because there are fewer relations annotated in the pdtb 2.0 than in the rst-dt. the alignment procedure is aimed at determining the optimal mapping of each pdtb relation to an rst relation. thereby, our goal is to identify as many valid correspondences between annotations as possible, but at the same time minimise mapping “noise” by also identifying those cases in which we cannot be sure that the annotators identified the same underlying relation.5 our mapping algorithm involves two major steps: 1. identifying for every pdtb discourse relation those rst-dt segments (edus or sub-trees containing more than one edu) that best correspond to the pdtb segments arg1 and arg2 separately. 2. identifying the rst-dt relation label that describes the relation between the arg1-equivalent and arg2-equivalent spans. in the first step, for each pdtb argument, we iterate over all rst edus in the source file and select the one with maximum overlap (common characters) and minimum margin (extra characters). we then determine whether a pdtb argument should be aligned to more than a single rst edu by iterating over all rst relation annotations (sub-trees spanning over several edus) using the same criteria. having identified the closest matching rst-annotated text spans for both arguments of the pdtb relation (arg1-equivalent and arg2-equivalent rst spans), we move to step 2 to find the lowest rst relation within the discourse tree that contains the two rst spans obtained in the previous step in different arguments. in this step, several error flags are set for manual investigation of possibly invalid alignments, which we briefly introduced in the previous section and will discuss in more detail below. to illustrate the alignment procedure, consider the example shown in figure 7. the first step of the alignment algorithm is to identify the corresponding text segment for pdtb’s arg2 (here, segment a) and arg1 (here, segments c-d). in order to do that, the algorithm would first compare all edus separately to the pdtb argument spans to identify the ones with most character overlap, and then move on to comparing larger spans. this would mean that it would first identify (c) as a matching edu for pdtb’s arg1, and then replace this by the better-matching combination of the span comprising both edus (c) and (d). we can see that in this example, segments (c-d) are conjoined in a condition relation in rst-dt, and are part of an attribution relation with segment (b), as well as a list relation with segments (e-g). finally, there is a relation connecting segment (a) to segments (b-g). pdtb’s 5. the alignments will be made available. 106 how compatible are our discourse annotation frameworks? figure 7: pdtb and rst-dt annotations for a paragraph of wsj 0619. note: only the pdtb annotation that is relevant for this example is included. segment (e-g) is simplified, it actually consists of multiple rst-dt relations. arg1 is hence embedded in several relations as part of a tree branch in rst-dt, including one multinuclear relation (the list relation). as part of its second step, the automatic algorithm proposes an alignment of the pdtb concession label to the rst concession label, because rst’s concession relation is the lowest relation which includes the pdtb arg2-equivalent span (a) and the arg1-equivalent span (c-d) in separate arguments. however, it would also automatically flag this instance due to the intervening multi-nuclear list relation: the multi-nuclear relation prevents it from unambiguously identifying the nucleus path from the nucleus of rst’s concession relation to the relevant textspan that corresponds to pdtb’s arg1. the algorithm flags all instances for which the mapping was potentially problematic, based on various criteria. first, relations for which pdtb arguments are discontinuous (e.g., containing text spans which are marked as not belonging to the relation, or centrally embedding the other pdtb argument, or overlapping with it) are flagged to allow for further manual checking. second, relation mappings that are inconsistent with the strong nuclearity hypothesis (marcu, 2000) were flagged automatically for exclusion from further analysis. for instance, imagine that pdtb had annotated a relation between segments (a) and (d) in the example in figure 5. such a constellation would be detected as violating strong nuclearity because (a) is the satellite of the relational argument that forms a relation with a span including segment (d). the labels for such a pair of relations should then not be compared. additionally, we added flags for relations that contain intervening multinuclear relations, for relations that were originally labelled as same-unit, and for relations that contain an intervening rst-dt attribution relation. 4.3 results of the alignment procedure in total, we were able to include 76% of pdtb relations from the joint corpus into our mapping analysis (a total of 5141 relations). 52% of these relations (a total of 2662) have directly corresponding argument spans, for which argument spans are exactly identical or differ only with respect to punctuation or inclusion/exclusion of connective in a segment. the remaining 48% (2489 instances) of the data included in the mapping analysis consists of relations for which the rst-dt tree is more complex than the pdtb relation. in other words, at least one of the pdtb arguments mapped onto an rst-dt relation that consists of multiple 107 demberg, scholman and asr figure 8: pdtb and rst-dt annotations for a paragraph of wsj 1176. note: only the pdtb annotation that is relevant for this example is included. rst-dt edus. in order to investigate whether these more complex relations are aligned correctly, we randomly selected 100 instances and evaluated whether the algorithm was justified in mapping pdtb and rst-dt labels for these instances. we found that 95 relations were mapped successfully, while 5 instances were unjustified. in these five cases, the nucleus of a larger rst-dt span matched pdtb’s argument, but the annotators did not evaluate the same type of relation. this was largely due to one of the segments (usually arg2) being part of a larger branch in rst-dt. even though this segment was the nucleus of that branch, it still blocked a stronger interpretation. to illustrate this, consider the example in figure 8. pdtb annotated a contrast relation between segments (b) and (d) (with segments (a) and (c) included as attribution), whereas rst-dt annotated an elaboration-additional relation between segments (a-b) and (c-d). the inclusion of the two attribution relations in rst-dt changes the interpretation of the two segments: the relation is focused on the speaker saying something and then adding to that (expressed by an elab.-additional relation), rather than on the content of what the speaker is saying (expressed by a contrast relation). although both the pdtb and the rst-dt annotators selected the same arguments for these relations, the labels might not always match because attribution can (in a subset of cases) “block” an interpretation. to quantify the risk related to intervening rst attributions, we counted the occurrence of attribution relations in the otherwise successfully mapped relations and found that 595 pdtb relations (12%) in total have at least one rst attribution relation in either of the matched segments; 49 (less than 1% of the data) have two or more intervening attributions (like the example given above). note that out of these 49 instances, only some exhibit the problem of attribution leading to different annotated labels between the two corpora; we therefore decided to include these instances in our analysis. 4.3.1 instances that could not be successfully aligned 24% (1621 instances) of pdtb relations were flagged by at least one of our flags indicating difficult cases. that is, these relations were automatically aligned, but these instances exhibit e.g. violation of the strong nuclearity principle, such that there is a higher chance that the annotated relation labels do not correspond to one another (i.e., annotators had different discourse structures or interpretations in mind). to provide a more quantitative idea of the cases that were excluded from the mapping, we again randomly selected and manually evaluated 100 instances. we found that 76 items were correctly flagged as unjustified mappings. out of these, 15 instances consisted of entrel labels. finding many entrel labels among the flagged instances is expected, since entrel is annotated between 108 how compatible are our discourse annotation frameworks? figure 9: pdtb and rst-dt annotations for a paragraph of wsj 0604. note: only the pdtb annotation that is relevant for this example is included. *the full label is elaboration-objectattribute. figure 10: pdtb and rst-dt annotations for a paragraph of wsj 1315. segment (c-e) is simplified. adjacent segments that do not have a stronger reading. often, these occur on the boundary of larger rst-dt spans (that is, discourse segments that function at higher rst-dt tree levels). looking at the total set of flagged mappings, we find that entrel makes up a large portion of this set: it occurs 333 times, which amounts to 19% of all flagged mappings. additionally, pdtb’s norel occurs 32 times in the flagged mappings (norel is used for adjacent arguments between which no discourse relation holds). the 76 correctly flagged relations also included four cases of same-unit relations which could not be mapped automatically because both segments of these same-unit relations consisted of multiple edus. an additional four same-unit relations were correctly mapped by choosing the relation label below the same-unit relation. after inspection of additional samples of same-unit relations for which the label could be unambiguously resolved automatically, we decided to include all of the unambiguously resolvable same-unit relations into our analyses, while same-unit relations where both segments have multiple edus are excluded. the remainder of 24 relations do in fact represent valid mappings. this illustrates that our algorithm prefers high reliability of the mapping over full coverage. in more than half of these instances (13 cases), the rst-dt annotation seems to be inconsistent with the strong nuclearity principle6. figure 9 illustrates this: pdtb annotated a synchrony relation between segment (b) and (c). rst-dt annotated a circumstance relation between segment (a-b) and (c), but segment (a) is in fact annotated as the nucleus. in other words, the nucleus path could not be traced back to pdtb’s arg1, and therefore the relation was flagged. however, the circumstance relation does not actually hold between the kidder name is one of only six or seven and when considering a merger deal; it holds between what pdtb marks as arg1 and arg2. 6. note that difficulties in consistently annotating nuclearity have been pointed out before (see e.g., stede, 2008) 109 demberg, scholman and asr ideational intentional explicit implicit explicit implicit 0 25 50 75 100 relation type p er ce nt ag e pdtb and rst−dt labels mismatch match figure 11: agreement between theoretically posited and practically found label mappings. the analysis is split out by explicitly marked vs. implicit relations and intentional vs. ideational relations. another typical case for automatic flagging by our algorithm occurred in cases where pdtb’s annotation constraint for annotating only adjacent implicit relations caused a mismapping: this can happen when two adjacent sentences convey a similar message and are followed by a third sentence which is arg2. in those cases, the second but not the first sentence has to be annotated as arg1 in pdtb. figure 10 presents an example of this: segments (a) and (b) present similar content. pdtb selects segment (b) as arg1 because of their segmentation principles for implicit relations. rstdt, however, selected segment (a) as the nucleus and (b) as the satellite. with this structure, the nuclearity path cannot be followed to (b) and the relation is flagged. we conclude that our manual inspection of a representative sample of instances confirms that our algorithm finds the right alignment between corpora in the majority of cases and thus provides us with reliable data for cross-corpora analysis of relation annotation. 5. correspondence between mapped relation labels our analysis of correspondences between mapped labels is based on a total of 5141 pdtb labels that could be mapped automatically with high confidence (see section 4.3). for any relation that carried two pdtb labels (in case the annotators thought that both relations held), we selected the label that is most similar to the corresponding rst-dt relation. a first overview of the results is given in figure 11, which summarizes the levels of agreement between actual annotations in the discourse corpora with the theoretically posited mapping (the present graph is based on the unidim mapping; specific differences between mapping proposals are however discussed in more detail in section 5.1). the figure shows that agreement between the empirically observed and theoretically predicted label correspondences is far from perfect. second, it shows that there is a big difference in agreement between explicit and implicit relations: for explicitly marked relations, empirical data is consistent with theoretical mappings in more than 70% of cases, while it is consistent with theoretical predictions for implicits in less than 50% of cases. we also show separately here the correspondence in mappings for ideational and intentional relations. following hovy & maier (1995), we classified each rst-dt label as ideational vs. intentional. this 110 how compatible are our discourse annotation frameworks? temp. cont. comp. expansion rst-dt label pd t b la be l sy nc h. a sy nc h. c au se c on di tio n c on tr as t pr ag m .c on tr. c on ce ss io n c on ju nc tio n c ho se n al t. e xc ep tio n in st an tia tio n sp ec ifi ca tio n l is t e nt r el total comparison 2 1 1 1 52 2 24 2 2 87 antithesis 1 3 186 2 37 15 5 2 1 2 254 elab.-object-attr. 6 7 1 1 15 background 10 13 13 2 19 2 4 21 84 circumstance 90 83 39 22 15 4 21 5 15 294 table 2: alignment of rst-dt relations for those labels that were identified as theoretically interesting cases of discrepancies between the three mapping proposals. split is theoretically motivated by the observation that pdtb focuses more on ideational relations; hence we might expect a higher agreement on ideational relations compared to intentional ones. we indeed find such a pattern, especially for explicit relations; however, a closer look at these cases shows that the effect is mostly driven by the fact that rst’s concession relations (which are classified as intentional relations) are often annotated as contrast relations in pdtb. we will get back to a more detailed analysis of the different types of discrepancies and the reasons for why they occur below. in the following sections, we will first discuss results concerning the discrepancies in theoretical mappings between frameworks as introduced in section 3, and then analyse in more detail the cases where theoretical predictions differ from empirical observations of labels assigned by the two frameworks in the corpus. as results differ strongly between explicit and implicit relations, we will discuss these relations separately: explicit relations are discussed in section 5.2 and implicit relations in section 5.3. 5.1 analysis of theoretically interesting discrepancies in expected mappings section 3 laid out theoretically interesting discrepancies between proposed mappings, relating to highlighting the rst-dt relations comparison, antithesis, elaboration-object-attribute, background and circumstance. table 2 shows the mapping of relations corresponding to these labels. we find that the empirical data confirms the usage of the label antithesis primarily for contrast relations rather than concessives; we would suggest that any new annotation efforts using the antithesis label should clarify its intended use in the annotation guidelines; the same holds for the label comparison, which would need to be more clearly defined for future use to avoid misinterpretation. the mapping of relations background and circumstance to pdtb is difficult, because these labels are more general, used when additional information needs to be provided to allow the reader/listener to understand. therefore, these segments have a function within the overall goal of the text, but do not seem to stand in a consistent semantic relation to the segment that they are supposed to provide additional context for. we next discuss each of the cases in relation to the proposed mappings in more detail. regarding rst-dt’s comparison, pdtb does not have a directly corresponding label. sanders et al. (2018) mapped this label to conjunction, because the rst-dt annotation manual ex111 demberg, scholman and asr plicitly defines comparisons as non-contrastive. however, bunt & prasad (2016) mapped it to contrast. looking at the empirical results (see table 2), we find that approximately one third of rst-dt comparison instances are annotated as conjunction, while two thirds (63%) are annotated as contrast in pdtb; we specifically note a high proportion of pdtb contrast.juxtaposition labels for these instances. this distribution is similar for explicit and implicit instances of comparison. typical markers for explicit instances are while, but and however. we therefore conclude that rst-dt’s comparison should be mapped to both pdtb conjunction and contrast. a second area of disagreement was the label antithesis, which unidim proposed to map to contrast or concession, but iso only to concession, and olia only to contrast. the empirical data clearly shows that the vast majority of antithesis relations (73%) map onto pdtb’s contrast relations, while 15% map to concession relations. these results are hence best in line with the olia and unidim proposals. next, we consider rst-dt’s elaboration-object-attribute. olia proposed a correspondence with general expansion labels, unidim proposed a mapping to specification and generalization, and iso proposed a mapping to pdtb’s entrel. the data only contained 15 instances of elaboration-object-attribute, and the majority of these (seven instances) were mapped to conjunction, confirming olia’s mapping. only one instance was annotated as specification. six cases were mapped to pdtb’s asynchronous label, which none of the frameworks predicted. closer inspection of these cases revealed that the modifying clause (which makes rst-dt annotate elaboration-object-attribute) often has a temporal aspect to it, as in example (7), thereby making pdtb more likely to annotate a temporal relation. the closest corresponding labels for elaboration-object-attribute therefore seem to be the more general conjunction label, as well as the temporal asynchronous type. (7) lawmakers would avoid [putting many spending projects into legislation in the first place for fear of the embarrassment] [of having them singled out for a line-item veto] later. — rst-dt elab.-object-attribute, pdtb asynchronous, wsj 0609 finally, we consider the labels background and circumstance. unidim mapped background to pdtb conjunction and asynchronous, whereas iso mapped it to entrel. as shown in table 2, background relations are annotated as a variety of relations: 23% are annotated as pdtb conjunction, 12% as asynchronous, and 25% as entrel, but cause and contrast (both 15%) also occur frequently. regarding circumstance, unidim mapped it to pdtb conjunction, synchronous and asynchronous, and iso only mapped it to synchronous. table 2 shows that circumstance is also annotated as a wide range of pdtb relations, indeed including conjunction (7%), synchronous (31%) and asynchronous (28%), but also cause (13%) and condition (7%). rst-dt’s background and circumstance therefore seem to be more general labels than both unidim and iso predicted. specifically, the rst-dt manual states that these labels are not causal, but in reality, they can be ambiguous. such cases are often marked with ‘as’ and ‘when’ (as in example (8)), which are known to be ambiguous (see e.g. asr & demberg, 2013). it is thus not surprising that rst-dt might annotate such relations as background or circumstance, whereas the pdtb would annotate these as cause. this could also in part be due to the lack of corresponding label in pdtb. 112 how compatible are our discourse annotation frameworks? (8) [but even that niche is under attack] as [several wall street firms pulled back from program trading last week under pressure from big investors.] — rst-dt circumstance, pdtb cause.reason, wsj 604 we conclude that the empirical mapping can provide insights for correcting or adjusting the theoretical proposals, for instance in the case of the elaboration-object-attribute relation. for other relations, we find that several proposals capture some of what actually happens in practical annotations, reinforcing the idea that oftentimes, the best mapping we can get between existing relation labels is not a one-to-one mapping, and that some labels, like background and circumstance capture a textual function which cannot be easily described as a specific semantic relation. we now turn to the findings for the mapping of other labels to see if the two frameworks’ annotations are compatible. 5.2 mapping for explicitly marked relations table 3 displays the mapping of pdtb annotations onto rst-dt annotations for explicitly marked discourse relations that occurred more than 30 times in total. in the table, relation mappings which were suggested by all three proposals are considered “expected” mappings and are marked in the table by underlined, bold numbers. mappings of labels for which at least two of three proposals agree are indicated by underlined numbers. note that for some labels, several entries in a row or column are marked as “expected”. this is because a one-to-one mapping is not always possible, as some labels have multiple matching candidate labels in another framework due to differences in the granularity of distinctions between frameworks (see section 3). generally, we find that deviations from expected mappings are often related to the relation begin marked with an ambiguous connective (e.g., as, when, but, while). these ambiguous connectives are in some instances interpreted differently in the two frameworks. furthermore, we find systematic differences in the operationalization between frameworks, which lead to mismatches in annotations for relations such as list. we will next take a closer look at the correspondence between theoretical mappings and empirical matches of labels for each of the major discourse relation classes (temporals, causals, contrastives, additives). temporals the results show that most (81%) of the explicitly marked relations that were classified as synchronous by pdtb were tagged as rst-dt temporal-same-time or circumstance. there are however also some cases where annotations deviated from expected mappings for temporal pdtb labels. these include cases where one of rst-dt’s causal labels (specifically, explanation-argumentative or consequence) was annotated. closer inspection revealed that frequent connectives in these relations, which did not receive a temporal sense label in rst-dt, were as and when. these connectives are known to be ambiguous markers (see e.g. asr & demberg, 2013), and are frequent among temporal relations such as circumstance and temporalsame-time in rst-dt. we hence find that there are some instances containing these ambiguous connectives which could not be consistently disambiguated between frameworks, with one framework labelling these instances as temporal and the other as causal. we will analyse the annotation of ambiguous connectives in more detail in section 5.2.1 pdtb’s temporal asynchronous relations generally also map well to their corresponding rst-dt classes (79%). the most notable unexpected pattern consists of pdtb temporal relations marked with until often being classified as condition in rst. this mismatch could be indicative of 113 demberg, scholman and asr temp. cont. comp. expansion rst-dt label pd t b la be l sy nc h. a sy nc h. c au se c on di tio n c on tr as t c on ce ss io n c on ju nc tio n in st an tia tio n l is t total temp.-same-time 63 1 1 1 4 5 75 temporal-after 48 1 2 54 sequence 2 29 1 1 19 52 circumstance 89 79 29 21 7 4 10 259 result 1 3 30 2 5 41 consequence 5 6 45 4 1 16 1 95 explanation-arg. 6 39 7 2 1 57 reason 2 67 72 condition 5 13 1 104 2 1 1 182 contrast 1 1 160 23 17 208 concession 5 4 3 101 53 4 182 antithesis 1 3 170 37 10 243 comparison 2 1 26 2 9 41 elaboration-add. 4 3 1 30 8 122 3 3 192 example 1 1 3 29 35 list 2 4 1 17 1 303 47 377 total 187 197 215 129 533 132 527 32 51 table 3: alignment of explicit discourse relation classes for which n >30. numbers indicate how many instances occurred in our high-confidence mapping; values are represented as colors. underlined, bold numbers indicate the mapping predicted to all three proposals; numbers that are only underlined (not bold) are predicted by two of three proposals. inconsistencies in disambiguation of this marker, or could be more systematically related to rst-dt annotating the intention in subjective relations, while pdtb annotations stay closer to the semantic relation (scholman & demberg, 2017). for rst-dt’s sequence class, we find that a substantial portion is annotated as pdtb’s conjunction (37%), i.e. annotations by the different frameworks disagree for these instances whether the relation is temporal or not. this was not predicted by any of the mapping proposals. a closer look at these instances reveals that they are mostly marked by the underspecified marker and, and can indeed have a somewhat temporal aspect to them, as example (9) illustrates. (9) [the farmer leaves] and [the naczelnik shuts his door.] — rst-dt sequence, pdtb conjunction, wsj 1146 causals explicit causal and conditional pdtb relation labels generally map well onto causal and conditional rst-dt labels; a finding that can be attributed to the relative definiteness of causal markers (we will see a different results for implicit relations in section 5.3). rst-dt distinguishes more types of causal relations, which results in causal pdtb relations being distributed among the various causal rst-dt classes. unexpected mappings – e.g., pdtb causals and conditionals annotated as rst-dt’s circumstance (12% and 16%, respectively) – were found to occur again for instances marked by the ambiguous connectives as and when, respectively. 114 how compatible are our discourse annotation frameworks? contrastives the majority of pdtb’s contrast relations was mapped to rst-dt’s contrast and antithesis relations (62%), as expected. however, we also found that a substantial portion (19%) of pdtb’s contrast relations are annotated as rst-dt’s concession; and some also as rst-dt’s elaboration-additional (6%). these cases were often marked by the connective but, which is an ambiguous connective. a closer look at the subtypes of pdtb concession reveals that relations annotated as pdtb’s concession.expectation map quite well (54%) onto rst-dt’s concession relations, while concession.contra-expectation relations are often annotated as contrast in rst, especially when marked with the connective but. we note that the distinction between concession and contrast relations is known to be difficult in discourse relation annotation (robaldo & miltsakaki, 2014). it is possible that the observed differences stem from differences in interpretation between annotators, and slight biases in the frameworks. overall, pdtb has a stronger bias than rst-dt towards assigning the contrast label: the majority of three rst-dt relational labels, namely contrast, antithesis and concession, are mapped to pdtb’s contrast. additives finally, looking at pdtb’s expansion relations, we find that a majority of relations annotated as pdtb’s conjunction is annotated as rst-dt’s list (57%), which was not expected based on the theoretical definitions of these relations. closer inspection shows that the high number of cases annotated as rst-dt list stems from the fact that pdtb annotation guidelines say that lists have to be “defined in the prior discourse” (prasad et al., 2008, p.37), as stocks are up does in example (10). unannounced lists cannot be annotated as list relations in pdtb. such a criterion is however not applied in rst-dt. we find that pdtb’s list and instantiation relations map well onto rst-dt’s list and example relations respectively. (10) those stocks are up because [their earnings are up] and [their dividends are up.] — rst-dt list, pdtb list, wsj 0681 we also observe a substantial amount of noise, i.e. 26% of conjunction relations have a temporal, causal or contrastive label in rst; some of these cases contain the connectives but or while, again indicating that connective disambiguation may not always be consistent between frameworks. a final interesting observation regards the annotation of the connective unless: these instances are annotated as as alternative.disjunctive relations in pdtb, but as condition in rst-dt (note that rst-dt does not have a label corresponding to pdtb’s disjunctive). although these relation types might seem to differ greatly, in reality they can be compatible. consider example (11), which was annotated as pdtb alternative.disjunctive and rst-dt condition. this is a negative conditional relation, but also can be considered disjunctive: two possibilities are evoked but only one of them can hold. (11) [essentially, he can’t make any hostile moves,] unless [he makes a tender offer at least $300 a share.] — rst-dt condition, pdtb alternative.disjunctive, wsj 1305 more generally, the finding that there are differences in the interpretation of ambiguous connectives may seem somewhat surprising, since disambiguation of these connectives is one of the prime motivation for annotating explicitly marked discourse relations in the first place. we will therefore next analyse the annotation of connectives in more detail, before turning to the mapping of implicit relations. 115 demberg, scholman and asr 5.2.1 analysis of annotation for ambiguous connectives given that the disagreements between the pdtb and rst-dt can be related to different interpretations of specific ambiguous connectives (as discussed in section 5.2), we studied the agreement on annotating ambiguous connectives in more detail. after all, it is not surprising that annotators can reach high agreement on relations with an unambiguous connective such as if for condition; given the empirical findings from previous annotation work in english, this could be done automatically in the future. rather, the value of additional manual annotation comes from disambiguating between relations when the connective can mark different relations (or when a relation is not explicitly marked, see section 5.3). methods to investigate the agreement on connectives, we tested the independence of pdtb and rst-dt annotations by calculating a separate χ2 test for each connective. for connectives where the distribution violated the assumptions of the χ2 test (less than 5 expected observations in a cell), we instead used the non-parametric fisher’s exact test. these analyses reveal whether the annotations agreed more with each other than can be expected based on the distribution of the connectives in the data. the results will provide insight into whether the content of the discourse relation arguments had been taken into account for the actual annotations of the relation instances, or whether the distinction is potentially too difficult or too subtle for human annotators to make reliably. note that this analysis is more strict than the usual kappa for inter-annotator agreement, because we use the distribution of relations per connective (i.e. which relations a connective can mark, and how often it does so for each relation type in the text at hand), which we are not normally known. we report some representative results for connectives where the null hypothesis of pdtb and rst-dt annotations for a connective being independent from one another. here, independence would mean that the content of the actual content of an instance is not predictive of which label was chosen by the one vs. the other framework; in the ideal case, we would expect strong nonindependence: given a pair of segments, the labels chosen by one and the other framework should correspond to one another and hence be deterministic not random. results the connective while is an example of a case where labels corresponded well to one another, i.e., where the null hypothesis of labels being independent could be rejected with high confidence: we find that annotators could reliably distinguish between the temporal.synchronous vs. contrast / comparison reading of while; the annotations from the two frameworks almost always agreed on the reading (p < 0.0001). there are however also connectives for which we find that the observed distributions are similar to random distributions. an example for this are connectives that are ambiguous between more similar discourse relations (contrast vs. concession), such as but, although and however. according to a χ2 test for but and fisher’s exact tests for although and however, the meaning distributions did not significantly differ from a random distribution of these sense labels (given the marginals). calculating κ values in this strict reading (i.e. taking for granted that but cannot mark causals or temporals and only testing agreement on different subtypes of negative relations) corresponds to κ < 0.1 for these relation label distinctions between frameworks. to summarize, we find that there are some ambiguous connectives for which manual annotation from the two frameworks reliably agreed on how the connective should be disambiguated and hence provide valuable additional information. this was mostly the case for when the alternative readings of the connective strongly differ from one another. however, we found no conclusive evidence 116 how compatible are our discourse annotation frameworks? that most subtle distinctions could be made reliably – on these cases, annotations from the two frameworks often don’t agree with one another more than would be expected by random assignment of labels that a connective can occur with (given the distribution of these connectives). this lack of agreement between humans may also provide a partial explanation for why automatic discourse relation disambiguation is difficult – it is unclear whether the training data that these distinctions are trained on is fully consistent internally, and/or the distinction may be so subtle in a substantial number of real data cases that even humans find it hard to agree. we will next move on to the analysis of agreement between frameworks on implicit discourse relation annotation. 5.3 mapping of implicit discourse relations the overall picture of annotation agreement between the two frameworks looks a lot more problematic for implicit than for explicit relations. to get an idea of what underlying causes these differences stem from, we decided to provide two perspectives on the results: once from the pdtb view, which provides an overview of how rst labels are distributed for a given pdtb relation (table 4a) and once from the rst view, showing how the pdtb labels are distributed for a given rst relation (table 4b). to keep tables readable, we only include those labels that occurred more than twenty times in the data. mapping of relation annotations for implicit relations, as seen from the pdtb perspective, is shown in table 4a. expected correspondences in the table are again indicated by underlined numbers. the colours in the table indicate the percentage of correspondence to rst-dt labels, with darker shades indicating a higher proportion of instances with a certain pdtb label falling into that rst-dt category. for example, 7 instances of pdtb temporal.asynchronous relations are annotated as rst-dt background. as these 7 instances represent more than 10% of the data on temporal asynchronous relations, the number is shaded in light green. on the other hand, the 7 counts of comparison.contrast relations labelled as rst’s consequence represent less than 10% of the 219 pdtb comparison.contrast relations; the entry is therefore not shaded. a first striking observation is that the agreement between frameworks is a lot worse than for explicit relations. while we saw a diagonal line of shaded cells that largely overlapped with theoretically expected correspondences, we can see that a substantial proportion of instances from almost all pdtb classes were annotated as rst-dt’s elaboration-additional. table 4b shows the alignment of implicit relations from the rst-dt perspective. many cells here are shaded green, which indicates that the annotations for many of the rst-dt relations have a wide variety of pdtb labels. generally, a stronger relation (biasing away from annotating simple additive labels) tended to be chosen by pdtb annotators; this bias can likely be linked to the pdtb annotation instructions of first inserting an explicit connective that fits the relation, and then assigning a label. wherever the rst-dt annotators chose a label other than elaboration-additional (the predominant label assigned to the implicit cases), the annotations matched with their pdtb equivalents relatively well. we will next analyse in more detail the mapping results for each of the four major relation classes. temporals temporal relations were not consistently identified between frameworks. only roughly one third of pdtb’s asynchronous.precedence relations had the expected rst label sequence, and most of pdtb’s asynchronous.succession relations were labelled as elaboration-additional in rst-dt. the rst-dt temporal-after label is very rarely annotated among implicit relations. 117 demberg, scholman and asr causals for pdtb’s cause.reason relations, less than 40% of instances were annotated as one of the expected causal classes by rst-dt annotators (expected classes were explanationargumentative, reason, evidence or interpretation), and only a very small percentage of pdtb’s cause.result were annotated as consequence, result or cause-result relations in rst. instead, most of these relations are annotated as elaboration.additional, showing that pdtb annotators tended to choose a “stronger” label than rst annotators for these instances. rst-dt’s causal relations consequence and reason map relatively well onto pdtb’s causal relations. however, other causal rst-dt labels (evidence and explanation-argumentative) are often mapped onto the additive pdtb labels instantiation and specification. this difference can be attributed to a fundamental difference between the approaches: pdtb annotates the lower-level ideational relations between arguments, while rst-dt focuses more on the intentional level. scholman & demberg (2017) show that often two functions can be identified in these specific relations, both illustrated in example (12): a segment can provide an example or a specification of something mentioned in the first argument (in this case, the track record), as well as providing evidence for a previously stated claim (e.g., that the track record stands out). these double functions are reflected in the mapping. (12) [in the cornucopia of go-go apples, the fuji’s track record stands out.] [during the past 15 years, it has gone from almost zilch to some 50% of japan’s market] — rst-dt evidence, pdtb specification, wsj 1128 contrastives only a minority of pdtb contrast relations was also labelled as a contrastive relation in rst-dt (17% for rst-dt contrast, 10% for comparison). instead, the majority maps to elaboration-additional and even list labels. from the perspective of rst-dt’s contrast relations, we observe that they were for the most part annotated as contrastive relations in pdtb, usually using the underspecified contrast label (rather than one of its subtypes). additives one pdtb implicit relation that matches well with rst-dt annotation is list, for which 84% of instances were annotated as rst-dt’s list relation. for rst-dt’s list relation, we see that the largest proportion of its instances (44%) are annotated as conjunction in pdtb; as mentioned earlier, this problem is partially due to the guideline in pdtb that lists have to be announced. taking a look at pdtb entrel labels, we can see that these are predominantly annotated as elaboration-additional and list in rst-dt. nevertheless, we also observe a number of different annotations on the rst side. we analysed a randomly selected subset of these instances and found that a common reason for the richer rst-dt labels is that these entrel relations are in fact high-level relations in rst-dt. the stronger interpretation is in those cases often due to the effect of additional content (outside the nucleus path), which facilitates the stronger interpretations. in our view, it is arguable whether entrel labels should really be considered to be mapped to rstdt labels, as the task given to the annotators differed markedly for these instances: in the pdtb case, the annotator task was to label the relation holding between two adjacent sentences, whereas the task of the rst-dt annotators was to join high-level discourse segments and describe their relation. 118 how compatible are our discourse annotation frameworks? rst-dt’s comment relation shows an almost uniform distribution across pdtb labels, consistent with other communicative functions that are not represented in the pdtb annotation scheme, such as background and circumstance, which have been discussed in section 5.1. finally, we observe that rst-dt’s elaboration-general-specific relations are often labelled as pdtb’s restatement (55%). this correspondence was predicted according to the expected mappings. 21% of instances were labelled as instantiation, which can be attributed to the subtlety of the distinction between these two labels. discussion on correspondence of annotations for implicit relations we find that the level of agreement in labels for implicit relations is a lot lower than for explicit ones. these results raise the question of how these very substantial differences in annotations of implicit relations can be explained. we think that the discrepancy can be attributed largely to the differences in annotation guidelines and operationalizations for implicit discourse relations. pdtb’s connective-driven approach biases against annotating simple additive relations when a connective can be inserted and hence an additional stronger interpretation of the discourse relation is available. rst-dt prescribes a different strategy: annotators are asked to annotate the writer’s intentions. the resulting low agreement between pdtb and rst-dt on implicit relations have implications for the reliability and validity of these annotations. we will expand on this point in the discussion. 6. discussion in the current paper, we evaluated how well theoretical proposals for mapping discourse relational labels correspond to the mapping between existing rst-dt and pdtb 2.0 annotations. some of the most important findings include: (i) rst-dt and pdtb agree more on explicit than implicit relations. the differences in annotation procedure and operationalization contribute greatly to this. (ii) among explicit relations, the ambiguity of connectives is a major source of disagreement. (iii) certain relations, such as concession and contrast, are inherently difficult to tease apart; the pdtb uses the contrast label more often while rst-dt more frequently uses the concession label. these differing label usages then lead to discrepancies in the mapping. the study consisted of three main steps: alignment of relations, evaluation of theoretically interesting discrepancies between the proposals, and evaluation of the mapped labels for explicit and implicit relations. we will discuss each one in turn. automatic alignment in order to evaluate existing annotations, we proposed an automatic mapping algorithm for pdtb and rst-dt discourse relation annotations and applied it to a segment of the wsj corpus that contains annotations from both frameworks. the algorithm proposed in this article allowed us to align 76% of pdtb discourse relation annotations to corresponding rst annotations. our manual error analysis shows that the alignment algorithm is highly accurate; it also correctly identifies instances where annotations cannot be aligned due to more fundamental differences in how the discourse is analysed. our study highlights the importance of discourse segmentation. segmentation has a strong effect on determining the scope and argument structure of a discourse relation (see also hoek et al., 2017). the differences in segmentation may hold interesting insights about effects of operationalization of discourse segmentation on discourse annotation, which could be explored in future work to refine annotation processes both for manual annotation and automatic processing. 119 demberg, scholman and asr temp. cont. comp. expansion rst-dt label pd t b la be l a sy nc h. c au se c on tr as t c on ju nc tio n a lte rn at iv e in st an tia tio n r es ta te m en t l is t e nt r el sequence 16 1 2 4 3 background 7 12 10 16 2 4 21 circumstance 1 10 8 11 6 15 consequence 4 16 7 11 2 2 1 4 evidence 12 2 12 2 29 28 6 explanation-arg. 2 114 16 16 2 35 51 1 22 reason 1 20 1 3 1 2 3 result 1 12 1 4 3 1 5 evaluation 7 9 11 2 1 12 interpretation 16 3 9 2 8 8 contrast 1 2 35 2 2 3 comparison 1 24 14 4 2 2 antithesis 15 5 2 1 2 elaboration-add. 29 168 73 221 10 36 151 7 266 example 9 1 6 64 21 2 elab.-gen.-spec. 8 11 17 44 18 list 6 24 29 120 2 6 13 74 30 restatement 5 2 1 2 12 1 comment 16 9 8 4 8 total 75 499 277 511 26 202 406 88 463 (a) colours encode percentage agreement from pdtb perspective, i.e. darker colours show that most instances of a pdtb relation type occurred in that specific rst-dt class. temp. cont. comp. expansion rst-dt label pd t b la be l a sy nc h. c au se c on tr as t c on j. a lte rn . in st an t. r es ta t. l is t e nt r el total sequence 16 1 2 4 3 27 background 7 12 10 16 2 4 21 72 circumstance 1 10 8 11 6 15 51 consequence 4 16 7 11 2 2 1 4 49 evidence 12 2 12 2 29 28 6 92 explanation-arg. 2 114 16 16 35 51 1 22 267 reason 1 20 1 3 1 2 3 31 result 1 12 1 4 3 1 5 28 evaluation 7 9 11 2 1 12 46 interpretation 16 3 9 2 8 8 49 contrast 1 2 35 2 2 3 56 comparison 1 24 14 4 2 2 44 antithesis 15 5 2 1 2 26 elaboration-add. 29 168 73 221 10 36 151 7 266 991 example 1 9 1 6 64 21 2 106 elab.-gen.-spec. 8 11 17 44 18 99 list 6 24 29 120 2 6 13 74 30 311 restatement 5 2 1 2 12 1 25 comment 16 9 8 4 8 49 (b) colours encode percentage agreement from rst-dt perspective, i.e. darker colours show that most instances of a rst-dt relation type occurred in that specific pdtb class. table 4: alignment of implicit discourse relation classes for which n >20. numbers indicate how many instances occurred in our high-confidence mapping. 120 how compatible are our discourse annotation frameworks? evaluation of mapping proposals as a result of the annotation alignment, we are able to offer a more complete picture of how annotations from the two frameworks relate to one another in practice. we compared actual annotations to expected correspondences that were determined based on three recent proposals for mapping discourse relations onto one another. our aim in evaluating three different mapping proposals was not to identify the “best” proposal; rather, the comparison between proposed correspondences and empirical coocurrences of annotated labels is helpful for achieving a deeper understanding of how certain definitions in the annotation guidelines were applied in practical annotation, and can help to decide between alternative proposals for mappings. the observed mismatches between the proposals can furthermore be used for clarifying annotation guidelines in future annotation projects; for example, the three proposals differed in their interpretation of rst-dt’s antithesis label, which indicates that the definition of this label could be expanded on. we found high numbers of disagreement between observed and expected annotations for those relations that did not have a direct correspondence in the other scheme, including labels that seem to be often used as blankets for ambiguous cases. examples of such relations include rst’s background and circumstance relations. the three proposals treated these relations differently from each other based on the definitions in the annotation manual, but generally, they were mapped to temporal or additive labels. however, the empirical mapping showed that many of these relations were also annotated as causals or even contrastives in pdtb. we see two possible explanations for what causes this: either the mapping scheme (and possibly the annotation manuals) would have to be revised to more clearly or exhaustively describe the relations, or there is a function to the relation that is not reflected in the pdtb scheme, and should be considered to be added to relation schemes, again possibly by annotating both the ideational and intentional functions separately. the current work addressed efforts in the community to create a standard typology of relations (consider frameworks such as the pdtb and rst-dt, but also international standards and intermediary approaches, such as those proposed by benamara & taboada, 2015; bunt & prasad, 2016; chiarcos, 2014; hovy & maier, 1995). this paper has highlighted that there are significant differences in the predicted correspondences of these proposals, and that we are still far from the goal of creating a standard relational inventory. the european cost initiative textlink, which ran from april 2014 until april 2018, was aimed at unifying the numerous, scattered linguistic resources on discourse structure. the deliverables of this cost action include a database of discourseannotated corpora, a database of lexicons, and a web portal that provides access to resources and a small suite of search, visualization and dissemination capabilities. one of textlink’s original goals was to create a single taxonomy that could be used in future annotation efforts. this proved to be inconceivable. instead, the unifying dimensions approach (sanders et al., 2018) was created as a final deliverable.the results from the current study show that the unidim approach was relatively successful in mapping between pdtb and rst-dt, but our empirical analysis also shows that framework-specific guidelines and operationalizations can cause mismappings that likely no intermediate approach can successfully deal with – the annotations done in the different frameworks simply do not always correspond to one another. the present article provides a detailed analysis that can help researchers to be aware of what types of relations they might find if searching the interoperable corpus originally annotated in one framework using a specific label from another framework. it also points out ways in which the research community can design the annotation guidelines in future annotation to reduce such discrepancies. 121 demberg, scholman and asr evaluation of mapped explicit and implicit relations after evaluating the theoretical discrepancies between the proposals, we looked at the mapped data for explicit and implicit relations separately. the most striking observations were a lower than expected level of agreement on annotations for implicit relations, and low agreement on more fine-grained distinctions for explicitly marked relations. in order to get more insight into the issue of difference in agreement between implicit and explicit relations, we recommend that future annotation efforts (corpus annotation as well as other tasks) report agreement on implicit and explicit relations separately. while the rate of noncorrespondences does seem problematic, we were able to identify several patterns that lead to these observed disagreements. specifically, many of the differences can be traced back to different operationalizations employed during annotation, as well as to the different goals of pdtb 2.0 vs. rst-dt annotation. first, the operationalization of discourse annotations may have a strong effect on the resulting annotations: in the pdtb annotation process, annotators were asked to annotate implicit relations by first identifying a discourse connective that would fit the relation, and then in a second step annotate the relation sense. it seems that this practice encourages annotators to assign more specific relation labels to implicit relations than rst-dt’s annotation procedure does. we therefore find that most implicit relations receive the rst-dt label elaboration-additional and a more specific pdtb label. for future annotation efforts, it is important that this consequence of annotation operationalization is taken into account. we cannot determine whether rst-dt annotators relied too heavily on the absence of a connective, which may have biased them to annotate the elaboration-additional relation more often, or whether the insertion of connectives as a task made some of the interpretations of relations stronger than they were without the connective, i.e. whether the operationalization to insert a connective may have changed the inferred relation in some cases. to answer this question, we suggest systematic annotation experiments for measuring effect of annotation instructions. second, the framework-specific goals have led to a focus on different levels of analysis of discourse relations, namely the ideational and the intentional level. ideational relations describe the semantic relation between the information conveyed in the consecutive elements of a coherent discourse (cf. moore & pollack, 1992). intentional relations on the other hand involve the writer’s attempts to affect the addressee’s beliefs, attitudes, desires etc. by means of language (cf. hovy & maier 1995; see also crible & degand 2017; redeker 1990). this distinction is relevant for the data used in the current study, because the goals of rst-dt annotations and pdtb 2.0 annotations differ with respect to these functions. while annotators in rst-dt were instructed to annotate the writer’s goal or intended effect of each segment of a text with respect to the neighbouring segments, pdtb annotators were asked to assess the relation between relational arguments, with a strong focus on the role of connectives (lexically-driven approach). the analysed empirical mappings provide support for the idea to annotate several levels of discourse relational arguments when appropriate. for instance, a segment can be an example for something that was said in the other segment, but it may at the same time serve as evidence for a claim (see also carston, 1993; blakemore, 1997). some of these patterns are systematic and go beyond the pdtb approach of allowing to annotate multiple labels for the same relation: pdtb annotators were not asked to systematically try to annotate all relations that hold. we conclude that future research should explore in more depth whether it is possible to devise an annotation procedure that allows researchers to identify both functions in a systematic way (see also crible & degand, 2017). 122 how compatible are our discourse annotation frameworks? finally, regarding the annotation of explicit relations, we found that disagreements are often related to ambiguous connectives such as as, but and while. we analysed whether the annotations of ambiguous connectives agreed more with each other than can be expected based on the distribution of connectives in the data. we found that the pdtb and rst corpus annotations agreed well for relations marked by connectives that can mark very different types of relations (e.g., while can mark a causal or a temporal relation), but they disagreed often on annotation of connectives that mark similar types of relations (e.g., but can mark a contrastive or a concessive relation). we believe that these cases warrant further study in order to better understand why the disagreements occur, and what the implications should be (e.g., a less fine-grained distinction among discourse relations if the present distinction cannot be made reliably, or an improved operationalization of the annotation process in order to achieve more agreement?). future directions the mapped annotations will be made available online so that other researchers can profit from the aligned corpus. we see several possible directions of research for which this mapped data can be useful. first, for theoretical studies, the data can serve to further investigate the frameworks and the effects of their operationalizations on the annotations. especially the cases that could not be aligned due to differences in structuring the discourse, and cases where different labels were chosen, are interesting from this viewpoint. some of the mismatches that occur between the pdtb and rst-dt annotations are systematic; for example, certain causal labels in rst-dt are often annotated as additive labels in pdtb, and rst’s contrast is often annotated in pdtb as concession. the mapping reveals these patterns and can therefore function as a starting point for other experiments that investigate these systematic mismatches. second, the mapping can prove to be useful for future annotations. the patterns of matches and mismatches that can be observed in the data can function as input for defining future annotation guidelines. the mapped data reveals which relation types may be particularly relevant for carrying several functions, and hence displayed less agreement between frameworks. the labels and definitions agreed upon across frameworks can be considered well-established, but for other types of relations, our mapping indicates that definitions may need to be refined in future efforts (for example, pdtb’s and rst’s contrast and concession, but also rst’s comparison deserves more consideration). our detailed results can also inform the ongoing discussion on identifying a set of labels for discourse relation annotation, which has been a long-lasting issue causing a lot of controversy in the literature. third, the mapped data can contribute towards automated discourse parsing efforts. discourse relation annotations have been used as training data in all recent efforts in automatic discourse relation classification. the classification of implicit discourse relations has received the bulk of the attention and work, given that classification of explicit relations was found to be relatively easy and accurate (pitler et al., 2008). implicit discourse relation classification has recently also been the subject of two conll shared tasks (xue et al., 2015, 2016), with accuracies just over 40% f-score on implicit relation sense labelling for an 11-way classification. important questions to be considered in the light of the mapping results in this article relate to how these classification results can be interpreted in the light of the difficulty of the implicit relation classification task. how can we make sure that consistency is improved for training automatic discourse relation classifiers? can and should we train classifiers separately for ideational vs. intentional discourse relation levels? should classifiers be evaluated by taking into account several possible labels for a relation, so that either the pdtb label or the corresponding rst label would be considered correct? or should 123 demberg, scholman and asr we weigh differently classification mismatch for categories that humans don’t commonly replace for one another versus those that are more interchangeable? the alignment data can also be used directly to select easy vs. difficult relation instances for training and evaluation of automatic relation identification systems. finally, we would like to emphasize that some of the methodological decisions for the present mapping are specific to the two exact frameworks we worked with, rst-dt and pdtb2.0. other instances of rst-style annotation may treat nuclearity differently, in which case some of the assumptions of our alignment algorithm would not necessarily generalize to those annotations. furthermore, pdtb-style annotation also differs between languages; different pdtb resources may use different relation inventories; nevertheless, we would expect that some of the fundamental observations we made (such as the effect of operationalization like the usage of implicit connectives during annotation) would transfer to those other pdtb-style resources. references al-saif, a., & markert, k. (2010). the leeds arabic discourse treebank: annotating discourse connectives for arabic. in proceedings of the 7th international conference on language resources and evaluation (lrec) (pp. 2046–2053). valletta, malta. asr, f. t., & demberg, v. (2013). on the information conveyed by discourse markers. in proceedings of the fourth annual workshop on cognitive modeling and computational linguistics (pp. 84–93). benamara, f., & taboada, m. (2015). mapping different rhetorical relation annotations: a proposal. in proceedings of the fourth joint conference on lexical and computational semantics,* sem (pp. 147–152). blakemore, d. (1997). restatement and exemplification: a relevance theoretic reassessment of elaboration. pragmatics & cognition, 5, 1–19. bunt, h., & prasad, r. (2016). iso-dr-core (iso 24617-8): core concepts for the annotation of discourse relations. in proceedings 12th joint acl-iso workshop on interoperable semantic annotation (isa-12) (pp. 45–54). cardoso, p. c., maziero, e. g., jorge, m. l., seno, e. m., di felippo, a., rino, l. h., nunes, m. g., & pardo, t. a. (2011). cstnews – a discourse-annotated corpus for single and multi-document summarization of news texts in brazilian portuguese. in proceedings of the 3rd rst brazilian meeting (pp. 88–105). carlson, l., & marcu, d. (2001). discourse tagging reference manual. carlson, l., marcu, d., & okurowski, m. e. (2003). building a discourse-tagged corpus in the framework of rhetorical structure theory. in current and new directions in discourse and dialogue (pp. 85–112). springer. carston, r. (1993). conjunction, explanation and relevance. lingua, 90, 27–48. 124 how compatible are our discourse annotation frameworks? chiarcos, c. (2014). towards interoperable discourse annotation. discourse features in the ontologies of linguistic annotation. in proceedings of the ninth international conference on language resources and evaluation (lrec 2014) (pp. 4569–4577). crible, l., & degand, l. (2017). reliability vs. granularity in discourse annotation: what is the trade-off? corpus linguistics and linguistic theory, . das, d., & taboada, m. (2018). rst signalling corpus: a corpus of signals of coherence relations. language resources and evaluation, 52, 149–184. hoek, j., evers-vermeul, j., & sanders, t. j. (2017). segmenting discourse: incorporating interpretation into segmentation? corpus linguistics and linguistic theory, advance online publication. hovy, e. h., & maier, e. (1995). parsimonious or profligate: how many and which discourse structure relations. unpublished manuscript, . iruskieta, m., aranzabe, m. j., de ilarraza, a. d., gonzalez, i., lersundi, m., & de lacalle, o. l. (2013). the rst basque treebank: an online search interface to check rhetorical relations. in 4th workshop rst and discourse studies (pp. 40–49). jansen, p., surdeanu, m., & clark, p. (2014). discourse complements lexical semantics for nonfactoid answer reranking. in acl (1) (pp. 977–986). lee, a., prasad, r., joshi, a., dinesh, n., & webber, b. (2006). complexity of dependencies in discourse: are dependencies in discourse more complex than in syntax? in proceedings of the 5th international workshop on treebanks and linguistic theories (tlt), prague, czech republic (pp. 79–90). lee, a., prasad, r., joshi, a., & webber, b. (2008). departures from tree structures in discourse: shared arguments in the penn discourse treebank. in proceedings of the constraints in discourse iii workshop (pp. 61–68). liu, y., li, s., zhang, x., & sui, z. (2016). implicit discourse relation classification via multitask neural networks. in proceedings of the thirtieth aaai conference on artificial intelligence (aaai-16) (pp. 2750–2756). mann, w. c., & thompson, s. a. (1988). rhetorical structure theory: toward a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8, 243–281. marcu, d. (2000). the theory and practice of discourse parsing and summarization. mit press. meyer, t., & popescu-belis, a. (2012). using sense-labeled discourse connectives for statistical machine translation. in proceedings of the joint workshop on exploiting synergies between information retrieval and machine translation (esirmt) and hybrid approaches to machine translation (hytra) (pp. 129–138). association for computational linguistics. moore, j. d., & pollack, m. e. (1992). a problem for rst: the need for multi-level discourse analysis. computational linguistics, 18, 537–544. 125 demberg, scholman and asr oza, u., prasad, r., kolachina, s., sharma, d. m., & joshi, a. (2009). the hindi discourse relation bank. in proceedings of the third linguistic annotation workshop (law) (pp. 158–161). association for computational linguistics. pitler, e., raghupathy, m., mehta, h., nenkova, a., lee, a., & joshi, a. k. (2008). easily identifiable discourse relations. technical report. polakova, l., mirovsky, j., & synkova, p. (2017). signalling implicit relations: a pdtb-rst comparison. dialogue & discourse, 8, 225–248. popescu-belis, a. (2016). manual and automatic labeling of discourse connectives for machine translation. in textlink–structuring discourse in multilingual europe second action conference károli gáspár university of the reformed church in hungary budapest, 11–14 april, 2016 (p. 16). prasad, r., dinesh, n., lee, a., miltsakaki, e., robaldo, l., joshi, a. k., & webber, b. (2008). the penn discourse treebank 2.0. in proceedings of the international conference on language resources and evaluation (lrec). citeseer. prasad, r., forbes-riley, k., & lee, a. (2017). towards full text shallow discourse relation annotation: experiments with cross-paragraph implicit relations in the pdtb. in proceedings of the 18th annual sigdial meeting on discourse and dialogue (pp. 7–16). prasad, r., miltsakaki, e., dinesh, n., lee, a., joshi, a. k., robaldo, l., & webber, b. (2007). the penn discourse treebank 2.0 annotation manual. prasad, r., webber, b., & joshi, a. (2014a). reflections on the penn discourse treebank, comparable corpora, and complementary annotation. computational linguistics, . prasad, r., webber, b., & joshi, a. (2014b). reflections on the penn discourse treebank, comparable corpora, and complementary annotation. computational linguistics, 40, 921–950. prasad, r., webber, b., & lee, a. (2018). discourse annotation in the pdtb: the next generation. in proceedings 14th joint acl-iso workshop on interoperable semantic annotation (pp. 87–97). redeker, g. (1990). ideational and pragmatic markers of discourse structure. journal of pragmatics, 14, 367–381. rehbein, i., scholman, m. c. j., & demberg, v. (2016). annotating discourse relations in spoken language: a comparison of the pdtb and ccr frameworks. in proceedings of the tenth international conference on language resources and evaluation (lrec). portoroz, slovenia: european language resources association (elra). robaldo, l., & miltsakaki, e. (2014). corpus-driven semantics of concession: where do expectations come from? dialogue & discourse, 5, 1–36. sanders, t. j. m., demberg, v., hoek, j., scholman, m. c. j., torabi asr, f., zufferey, s., & evers-vermeul, j. (2018). unifying dimensions in discourse relations: how various annotation frameworks are related. corpus linguistics and linguistic theory, ahead of print, 1–71. 126 how compatible are our discourse annotation frameworks? sanders, t. j. m., & spooren, w. (1999). communicative intentions and coherence relations. in w. bublits, u. lenk, & e. ventola (eds.), coherence in spoken and written discourse (pp. 235– 250). john benjamins. sanders, t. j. m., spooren, w. p. m. s., & noordman, l. g. m. (1992). toward a taxonomy of coherence relations. discourse processes, 15, 1–35. scheffler, t., & stede, m. (2016a). adding semantic relations to a large-coverage connective lexicon of german. in proceedings of the tenth international conference on language resources and evaluation (lrec 2016). scheffler, t., & stede, m. (2016b). mapping pdtb-style connective annotation to rst-style discourse annotation. in proceedings of the 13th conference on natural language processing (konvens 2016). scholman, m. c. j., & demberg, v. (2017). examples and specifications that prove a point: identifying elaborative and argumentative discourse relations. dialogue & discourse, 8, 56–83. sharp, r., jansen, p., surdeanu, m., & clark, p. (2015). spinning straw into gold: using free text to train monolingual alignment models for non-factoid question answering. in hlt-naacl (pp. 231–237). somasundaran, s., namata, g., wiebe, j., & getoor, l. (2009). supervised and unsupervised methods in employing discourse relations for improving opinion polarity classification. in proceedings of the 2009 conference on empirical methods in natural language processing: volume 1-volume 1 (pp. 170–179). association for computational linguistics. stede, m. (2008). rst revisited: disentangling nuclearity. subordinationversus coordinationin sentence and text, (pp. 33–59). stede, m., & neumann, a. (2014a). potsdam commentary corpus 2.0: annotation for discourse research. in proceedings of the ninth international conference on language resources and evaluation (lrec 2014) (pp. 925–929). stede, m., & neumann, a. (2014b). potsdam commentary corpus 2.0: annotation for discourse research. in lrec (pp. 925–929). taboada, m., & mann, w. c. (2006). rhetorical structure theory: looking back and moving ahead. discourse studies, 8, 423–459. tonelli, s., riccardi, g., prasad, r., & joshi, a. k. (2010). annotation of discourse relations for conversational spoken dialogs. in proceedings of the 7th international conference on language resources and evaluation (lrec) (pp. 2084–2090). marrakesh, marocco. webber, b., prasad, r., lee, a., & joshi, a. (2016). a discourse-annotated corpus of conjoined vps. in proceedings of the 10th linguistic annotation workshop (law-x) (pp. 22–31). xue, n., ng, h. t., pradhan, s., prasad, r., bryant, c., & rutherford, a. (2015). the conll-2015 shared task on shallow discourse parsing. in conll shared task (pp. 1–16). 127 demberg, scholman and asr xue, n., ng, h. t., rutherford, a., webber, b., wang, c., & wang, h. (2016). conll 2016 shared task on multilingual shallow discourse parsing. proceedings of the conll-16 shared task, (pp. 1–19). zhou, l., li, b., gao, w., wei, z., & wong, k.-f. (2011). unsupervised discovery of discourse relations for eliminating intra-sentence polarity ambiguities. in proceedings of the conference on empirical methods in natural language processing (pp. 162–171). association for computational linguistics. zhou, y., & xue, n. (2015). the chinese discourse treebank: a chinese corpus annotated with discourse relations. language resources and evaluation, 49, 397–431. zirn, c., niepert, m., stuckenschmidt, h., & strube, m. (2011). fine-grained sentiment analysis with structural features. in ijcnlp (pp. 336–344). zufferey, s., & degand, l. (2013). annotating the meaning of discourse connectives in multilingual corpora. corpus linguistics and linguistic theory, 1, 1–24. 128 how compatible are our discourse annotation frameworks? appendix a. pdtb 2.0 and rst-dt tagsets temporal contingency comparison expansion synchronous asynchronous precedence succession cause pragmatic cause condition pragmatic condition reason result justification hypothetical general unreal present unreal past factual present factual past relevance implicit assertion contrast pragmatic contrast concession pragmatic concession juxtapositon opposition expectation contra-expectation conjunction instantiation restatement alternative exception list specification equivalence generalization conjunction disjunction chosen alternative figure 12: hierarchy of relation senses in pdtb (prasad et al., 2008). figure 13: tagset of relation senses in rst-dt (carlson & marcu, 2001). 129 demberg, scholman and asr appendix b. information on the olia repository the olia format is a machine-readable format that uses several intermediary hierarchies of relations. as a result of its format, it cannot be easily represented compactly as a tagset or something similar. the olia website7 does, however, provide two figures to illustrate the approach. these figures are replicated here. figure 14 presents an example mapping of several instances annotated by both the pdtb and rst-dt. figure 15 presents a more detailed example mapping of the condition relation labels in the pdtb and rst-dt. figure 14: example mapping of the relations in a text fragment carrying rst-dt and pdtb 2.0 annotations. figure 15: example mapping of the condition relation labels in pdtb and rst-dt using olia. 7. http://www.acoli.informatik.uni-frankfurt.de/resources/discourse/ 130 http://www.acoli.informatik.uni-frankfurt.de/resources/discourse/ how compatible are our discourse annotation frameworks? appendix c. unifying dimension definitions figure 16: the unifying dimensions and features, and their values (sanders et al., 2018) 131 demberg, scholman and asr appendix d. iso tagset figure 17: tagset of relation senses in the iso proposal (bunt & prasad, 2016) 132 how compatible are our discourse annotation frameworks? appendix e. pseudocode of the pdtb-rst mapping algorithm procedure findrela/onmappings(file){ for pdtbrela-on in file do: // find equivalent rst spans to arg1 and arg2 arg1equivalent := findspanmatch(pdtbrela-on.arg1); arg2equivalent := findspanmatch(pdtbrela-on.arg2); bestrela-onmatch := null; // find the rst rela/on at the lowest (smallest) sub-tree spanning over the two rst spans for rstrela-on in file do: if rstrelaion.covers(arg1equivalent, arg2equivalent) and (bestrela-onmatch == null or rstrela-on.spansize < bestrela-onmatch.spansize) then: bestrela-onmatch := rstrela-on; // checks for replacement and/or flagging: //#1: check for same-unit relabon same-unitcheck(arg1equivalent, arg2equivalent, bestrela-onmatch); //#2: check for intervening alribu/on rela/ons aqribu-oncheck(arg1equivalent, arg2equivalent, bestrela-onmatch); //#3: check for the strong nuclearity principle nuclearitycheck(arg1equivalent, arg2equivalent, bestrela-onmatch); // ader all checks are done, the mapping will be added to the pool allmappings.add(pdtbrela-on, bestrela-onmatch) return allmappings; } procedure findspanmatch(pdtbarg, file){ // a span in rst can be a single edu or a mul/-edu sub-tree argequivalent := null; maxoverlap := 0; minmargin := 100000; for rstspan in file do: if (rstspan.overlap(arg1) > maxoverlap) or (rstspan.overlap(arg1) == maxoverlap and rstspan.margin(arg1) < minmargin) then: argequivalent := rstspan; } figure 18: pdtb-rst mapping algorithm pseudocode 133 demberg, scholman and asr appendix f. theoretical mapping between rst-dt and pdtb 2.0 labels according to iso, olia, and unifying dimensions pdtb→ temporal contingency comparison asynch. cause pragm. prag. contrast concess. cause cond. rst-dt ↓ sy nc hr on ou s pr ec ed en ce su cc es si on r ea so n r es ul t ju st ifi ca tio n c on di tio n r el ev an ce im pl .a ss er t. c on tr as t ju xt ap os . o pp os iti on pr ag m .c on tr. e xp ec ta tio n c on tr aex p. pr ag m .c on c. background u u circumstance u, i u u cause o, u, i o, u, i i cause-result o, u, i o, u, i i result o, u, i o, u, i i consequence o, u, i o, u, i i comparison i i i preference u u u u i i analogy proportion u u u condition o, u, i o, u o, u hypothetical o, u, i u u contingency o, u, i otherwise u, i contrast o, u, i o, u o, u o, u concession o, u, i o, u, i o, u antithesis o, u o, u o, u o, u u, i u, i u el-additional el-gen.-spec. el-part-whole el-proc.-step el-object-attr. el-set-mem. example definition purpose o o, u, i enablement u, i o o evaluation u u conclusion o, u o o, u o o o comment evidence u, i u, i o, u, i expl.-argum. o, u, i o, u, i i reason o, u, i o, u, i i list disjunction summary restatement temp.-before o, u, i i temp.-after i o, u, i t.-same-time o, u, i sequence o, u, i u, i inverted-seq. u, i o, u, i means u u o o problem-sol. u u o o not mapped i i i i table 5: proposed mappings between rst-dt and pdtb 2.0 labels (expansion labels on next page). ‘i’ indicates proposed correspondence according to the iso proposal, ‘o’ corresponds to olia, and ‘u’ corresponds to the unifying dimensions proposal. 134 how compatible are our discourse annotation frameworks? pdtb→ expansion restatement alternative rst-dt ↓ e xp an si on c on ju nc tio n in st an tia tio n sp ec ifi ca tio n e qu iv al en ce g en er al iz . c on ju nc tiv e d is ju nc tiv e c ho se n al t. e xc ep tio n l is t e nt r el n ot m ap pe d background u i o circumstance u o cause cause-result result consequence comparison u o preference u u o analogy u, i o proportion u, i o condition hypothetical contingency otherwise o o o contrast u concession antithesis i u el-additional o u, i i el-gen.-spec. o, u, i o, u, i el-part-whole o u, i u, i el-proc.-step o u, i u, i el-object-attr. o u u i el-set-mem. o u, i example o, u, i definition i u u o, i purpose enablement evaluation i u u o conclusion i i comment i u u u o evidence expl.-argum. reason list o, u, i disjunction o, u, i o, u, i o, u summary o, u, i o o, u, i restatement o, u, i temp.-before temp.-after t.-same-time sequence inverted-seq. means i problem-sol. i not mapped o, i u table 6: proposed mappings between rst-dt and pdtb 2.0 expansion labels, including a general expansion category. ‘i’ indicates proposed correspondence according to the iso proposal, ‘o’ corresponds to olia, and ‘u’ corresponds to the unifying dimensions proposal. note: labels that were excluded by at least two frameworks were not included in this table. these labels are rstdt’s question-answer, statement-response, topic-comment, comment-topic, rhetorical-question, topic-shift, topic-drift, attribution. general labels for the types condition, contrast, and expansion were included as subtypes because the three proposals mapped to these more general labels. 135 introduction background rhetorical structure theory discourse treebank (rst-dt) penn discourse treebank (pdtb)-style annotation theory-based proposals for mapping rst-dt and pdtb 2.0 relations mapping according to the olia reference model mapping via unifying dimensions mapping according to the iso standard proposal discussion of agreement and discrepancies between proposed mappings previous work on empirical evaluation of mapping coherence relation annotations data, segmentation and automatic alignment segmentation segmentation constellations alignment algorithm results of the alignment procedure instances that could not be successfully aligned correspondence between mapped relation labels analysis of theoretically interesting discrepancies in expected mappings mapping for explicitly marked relations analysis of annotation for ambiguous connectives mapping of implicit discourse relations discussion pdtb 2.0 and rst-dt tagsets information on the olia repository unifying dimension definitions iso tagset pseudocode of the pdtb-rst mapping algorithm theoretical mapping between rst-dt and pdtb 2.0 labels according to iso, olia, and unifying dimensions journal of machine learning research-microsoft word template dialogue and discourse 3(2) (2012) 1˗9 doi: 0.5087/dad.2012.201 ©2012 paul piwek and kristy boyer varieties of question generation: introduction to this special issue paul piwek p.piwek@open.ac.uk centre for research in computing the open university walton hall milton keynes, mk7 6aa, united kingdom kristy elizabeth boyer keboyer@ncsu.edu department of computer science north carolina state university 890 oval dr. campus box 8206, raleigh, nc 27695-8206, united states introduction questions are a topic that is ideally placed at the intersection of research on dialogue and discourse: in most dialogue, questions are the driving force, setting its direction, whereas answers to questions require discourse processing (e.g., contextual resolution in the case of short answers) or can be viewed as a piece of discourse themselves (e.g., an explanation in response to a whyquestion). current work on questions in both formal and computational semantics can be traced back to early explorations into the application of modern logic to questions. in particular, cohen (1929) proposed that the content of a question can be expressed as an open formula with one or more unbound variables. this idea reverberates in more recent work on the (computational) semantics of questions (e.g., ginzburg & sag, 2001; piwek, 1998; scha, 1983); at the same time, an alternative view of questions, originally proposed by c.l. hamblin, has taken hold among formal semanticists; according to this view, the content of a question corresponds to the set of its answers (see groenendijk & stokhof, 1997). the past two decades have also seen a recognition of the centrality of questions in theories of dialogue – e.g., ginzburg‟s questions under discussion (ginzburg, 2012). in the field of natural language processing, questions have been extensively studied as part of the task of question-answering. question-answering research initially focused on answering questions from databases and knowledge representations (e.g., green et al, 1961; bronnenberg et al., 1979), but in the past two decades has refocused on retrieving answers from text – e.g., in 1999 the evaluation of question-answering systems became part of the text retrieval conference (trec) series. simultaneously, there has been a strand of research on advisory dialogue systems – e.g., winograd‟s shrdlu (winograd, 1972) and more recently the denk cooperative assistant (ahn et al., 1995) – concentrating on theoretical issues by working with applications in restricted domains (cf. hirschman and gaizauskas, 2001). all the aforementioned systems were primarily aimed at responding to the user‟s questions; even most advisory dialogue systems would only ask a question if there was some issue, which the user introduced, that needed clarification. 1 1 for an exception, see the seminal work of power (1979) on machine-machine dialogue (rather than human-machine, dialogue), which involves dialogue agents asking questions as a part of joint plan construction. the questions are used by the agents to elicit answers that address gaps in their domain knowledge. piwek and boyer 2 the rich tradition of research into questions in both formal/computational semantics and natural language processing has focused on the interpretation of questions and how to answer or respond to them. until recently, there were few studies looking into the conditions under which a question gets asked, i.e., the principles and/or processes that underlie the generation of questions. in natural language processing and more specifically the subfield of natural language generation (mcdonald, 1993; reiter & dale, 2000; evans et al., 2002) research has been limited almost exclusively to generating text consisting of declarative sentences. an exception is work specifically on clarification questions in natural language processing (e.g., kievit et al., 2001; purver, 2004) and also stent‟s (2001) work on language generation for dialogue systems. in formal semantics, one of the few exceptions is wisniewski‟s monograph „the posing of questions‟ (wisniewski, 1995). according to wisniewski „arriving at a question‟ is analogous to coming to a conclusion, i.e., as „some premises [being] involved and some inferential thought processes [taking] place‟ (wisniewski, 1995: xi). in contrast with formal/computational semantics and natural language processing, question generation has a substantial history in education and psychology, see olney et al. (this volume) for an overview of work in this tradition which goes back at least as far as piaget (1952). particularly influential has been the work of art graesser and collaborators in which specific mechanisms for question generation are proposed. for example, graesser et al. (1992) present 22 different mechanisms that are grouped into four categories: correction of knowledge deficit, monitoring common ground, social coordination of action, and control of conversation and attention. recently, there has been a broader uptake of research into question generation. since 2008, researchers from different communities, ranging from, but not limited to, discourse analysis, dialogue modeling, formal semantics, intelligent tutoring systems, natural language generation, natural language understanding, and psycholinguistics, have met annually at the question generation workshop. 2 at the 3 rd workshop in 2010, the first question generation shared task and evaluation campaign (see rus et al., this volume) took place. with a significant body of work accumulating on question generation, it seems timely to collect a representative sample of high quality efforts in a special issue. the original call for the issue attracted 26 initial submissions. as a result of several selection stages, 7 papers were eventually accepted. before briefly summarizing the contributions of these papers, the next section sets out to define the concept of question generation and to identify different types of question generation. section 2 introduces the papers in this issue. we conclude with some remarks on the prospects and future direction of question generation as a topic of research. 1 question generation tasks as computational problems the emerging question generation (qg) community has adopted the following definition of question generation (cf. rus et al., 2008): “question generation is the task of automatically generating questions from various inputs such as raw text, database, or semantic representation. question generation is regarded as a discourse task involving the following four steps: (1) when to ask the question, (2) what the question is about, i.e. content selection, (3) question type identification, and (4) question construction.” (rus, n.d.) 2 the 1st workshop took place in arlington (hosted by the nsf), the 2nd in brighton (co-located with the 14th international conference on artificial intelligence in education), the 3rd in pittsburgh (co-located with the tenth international conference on intelligent tutoring systems) and the 4th event, in 2011, was organized as an aaai symposium in arlington. varieties of question generation 3 this definition views qg as the task of constructing algorithms that transform inputs to certain outputs. in computer science, an algorithm is “a tool for solving a well-specified computational problem.” (cormen et al., 2009:5) the definition of qg raises an entire family of computational problems, rather than a single one: it gives examples of different types of input, leaves the relation between the input and the output undetermined, and says of the output only that it should be a question. in other words, the definition of question generation gives rises to a large variety of question generation problems. in order to contextualise the papers in the current issue, we aim to chart some of this variety. following piwek et al. (2008), we characterize a specific question generation problem by the answers to the following three questions: 1. what is the input? 2. what is the output? 3. what is the relation between the input and the output? let us start with the output, since the constraints on this seem most stringent. according to the definition, the output should be a question. even here, there is, however, scope for several interpretations. one possible interpretation of this is that the output should syntactically qualify as a question, i.e., an interrogative sentence. this seems, however, to be overly restrictive, and a more plausible interpretation is that the output should be pragmatically a question, i.e., a request for information. requests for information can take many different forms, including the utterance of an interrogative sentence (“how late is it?”), the utterance of a declarative sentence with a question intonation (“mary is at home?”) and the utterance of an imperative (“tell me where i can find the off switch.”). of course, the imperative “tell me where i can find the off switch.” features an embedded question (“where i can find the off switch”), but not all embedded questions signal the presence of a pragmatic question. for example, the assertion “i know where the off switch is” is not a request for information, in contrast with “tell me where i can find the off switch.” additionally, not all pragmatic questions contain overt embedded questions; consider, e.g. “tell me the time” as one way for asking what time it is. even though we believe that the appropriate concept of a question for the purpose of the question generation community is that of a pragmatic question, i.e., a request for information, most work so far has focused on the more narrow problem of generating interrogative sentences. this holds true, for example, for the contributions to this issue. additionally, the work reported here concentrates on written questions in a natural language (specifically, english and french). many alternative question generation problems can, however, be envisaged if we change the output medium (auditory, visual, tactile, …) or modality (natural language, gesture, knowledge representation, …). for example, a spoken question (medium: auditory; modality: natural language) could consist of a declarative sentence that is uttered with question intonation or which is uttered in a context which leads the addressee to infer that it is a question rather than an assertion (see beun 1990 on „declarative questions‟). requests for information can also be multimodal (modality: natural language & gesture) as in „who is that?‟ (accompanied by a pointing act). it may even be possible to ask a question using facial expressions only (e.g., raising an eyebrow at the right point in a conversation). turning our attention to the input side, we have the same wealth of possibilities. again, the ground covered so far in qg research seems to be limited. the work in this issue focuses on input in the form of written text and knowledge representations. by varying medium and modality we arrive at many further possibilities: spoken language input, diagrams, presentations by virtual agents, etc. the input also need not consist of declarative, asserted, information; for example, the input could itself be a question. input and output do not need to share medium and modality: the input could be diagram and the output a written (or spoken) question about that diagram. additionally, even when the piwek and boyer 4 medium or modality is constant across input and output there may be significant differences. for example, input and output could be in natural language, but with the input language german and the output language french. this would give rise to a cross-language question generation problem. in third place, and perhaps most importantly, there is the relation between the input and the output. here again, research on qg has focused mainly on one particular relation: the situation where the output question is answered by the input (e.g., input = “john bought five cakes”, and output = “how many cakes did john buy?”) there has, however, also been some interest in question reformulation, where the output question is a clearer, better or different formulation of the input (which is a question or a search query), see marciniak (2008). wisniewski (1995) pioneered the notion that an output question can be raised by an input, as in the following example. the input „if john tried to check in after 12:00, he missed the flight. did john miss the flight?‟ raises the question „did john try to check in after 12:00?‟ finally, one further relation between input and output is that of clarification: the output question requests information about the form, content, or intended use of the input (e.g., input = “remove that block” and output = “the green one?”). there are many more relations, some of which are implicit in the qg mechanisms that are described in graesser et al. (1992). the papers in this issue deal with declarative (asserted) input, either in natural language (french or english) or a knowledge representation formalism (concept maps), and output consisting of interrogative sentences, which are answered by the input. this significantly narrows down the set of qg problems that are addressed. the algorithms for solving these problems vary, however, in a number of ways. we conclude this section, by focusing on one particular dimension along which they differ, which is loosely based on the vauquois triangle, or pyramid, from research in machine translation (vauquois, 1968). machine translation (mt) algorithms take a source language text and map it to a target language text. three approaches can be distinguished known as direct, transfer and interlingua mt. in direct mt, the transformation rules operate directly on the strings of the source text to obtain the target text. for transfer, an intermediate (syntactic or semantic) structure is constructed for the input text; this is mapped to the corresponding structure in the target language; the structure for the target text is then mapped to the target language text itself. in contrast with the transfer approach, the interlingua approach is based on a single (conceptual) structure to which the source text is mapped and from which, in turn, the target text is generated. in text-to-text qg, the input is a declarative sentence. usually, a pre-processing step is involved which divides this sentence up into sentences of a size that is appropriate for generating questions. similar to the hierarchy in mt (direct, transfer and conceptual), there is a hierarchy of qg approaches. in principle, it is possible to carry out qg strictly on the string level (figure 1.a). to our knowledge, there are, however, no systems that directly implement this approach. rather, most systems carry out syntactic processing and some semantic processing to arrive at an intermediate representation, which may or may not include the original source text. the amount of semantic processing varies from only partial analysis (e.g. recognition of named entities) to full processing into a semantic representation language. the approaches share the use of transformation rules. there are, however, two different kinds of transformation rule: 1) rules that map the representation of the input text, which includes the input text itself, directly to the target question without any further intermediate representation (figure 1.b), and 2) rules that map the representation of the input text, which may or may not include the input text itself, to a representation of the target question (figure 1.c and d). the transformation rules of the second kind require a further step that consists of generating the target question text from the target question representation. at the top of the pyramid in figure 1 we have transformations from a semantic representation of input text to a semantic representation of the target question. note the absence of an equivalent to interlingua mt: in mt, the conceptual representation for the content of the target and source language text is the same (it resides at the peak of the pyramid); in varieties of question generation 5 contrast, there is no shared conceptual representation (i.e., the pyramid has no peak) that captures both the content of the declarative source text and the interrogative target question. we should add that it is possible to use transformation rules that overgenerate and are followed by a post-processing step that weeds out ill-formed outputs using an independent measure, e.g. by ranking generated question with an appropriate language model (see heilman & smith, 2009; 2010). figure 1 the question generation pyramid 2 overview of papers in this issue rather than try to replicate the excellent abstracts that precede each of the papers, in this section we contrast and compare the papers using some of the distinctions that were introduced in the previous section. we group the seven papers into four broad categories, primarily based on domain of application: 1) application-neutral generation of questions from sentences, 2) question generation in educational settings, 3) question generation for virtual agents, and 4) question generation shared tasks and evaluation. 2.1 application-neutral generation of questions from sentences the paper by yao, bouma & zhang entitled “semantics-based question generation and implementation” falls firmly in the category question generation from declarative sentences. the main contribution of this paper is to demonstrate the feasibility of semantics-based qg. the open source mrsqg system is the first of its kind, implementing a semantic rewriting approach (see figure 1.d). it relies on parsing the english input sentence (after some preprocessing) to a structure in minimal recursion semantics (mrs; copestake et al., 2005). a rewriting step then maps the mrs for the declarative sentence into an mrs for a corresponding question. language generation is deployed to turn this mrs into an english output question. both parsing and language generation are based on existing open source tools for mrs. the rewriting step is in principle language independent: it maps the language neutral mrs for the input to an mrs for the output question. the system performed very well in the first question generation shared task and evaluation campaign (qgstec; rus et al., this volume). the qgstec was set up to be application neutral, using texts on a variety of topics and evaluation measures that ranked approaches based on generic criteria (such as syntactic and semantic correctness of the generated questions). piwek and boyer 6 the paper by bernhard, de viron, moriceau & tannier (“question generation for french: collating parsers and paraphrasing questions”) relies less on semantics and more on syntax: in particular, it draws on the outputs of two different syntactic parsers in combination with named entity recognition. in terms of figure 1, the approach falls somewhere between c. and d. in this system, generation concerns removing information from the output question‟s syntax tree and application of elisions (e.g., from que to qu′). additionally, the system can generate reformulations based on substitution of different question words. the paper is a first in that it introduces an application-neutral questions-from-sentences generator for french. the authors describe a range of ways in which their system was evaluated. these include human judgments on acceptability of generated question, use of post-edited questions as a target reference against which generated questions are compared using the human-targeted translation edit rate (hter; snover et al., 2006), recall in terms of the clef question answering campaign corpus, and extrinsic evaluation in terms of how well generated questions can be parsed correctly automatically. 2.2 question generation in educational settings olney, graesser & person (“question generation from concept maps”) present a method for generating questions for tutorial dialogue. this involves automatically extracting concept maps from textbooks. in contrast with the aforementioned qg from sentences approaches, this approach does not deal with the input text on a sentence-by-sentence basis only. rather, various global measures (based on frequency measures and comparison with external ontologies) are applied to extract an optimal concept map from the textbook. the template-based generation of questions from the concept maps allows for questions at different levels of specificity to enable various tutorial strategies, from asking more specific questions (which indicated precisely to the student which information they are expected to provide) to a student who is struggling, to the use of less specific questions to stimulate extended discussion. evaluation of the approach follows the qgstec methodology (rus et al., this volume), enriched with an evaluation of the pedagogical value of questions. liu, calvo & rus (“g-asks: an intelligent automatic question generation system for academic writing support”) introduce a system which helps students to develop their writing skills. it focuses on writing around citations. the approach is template-based, and takes as input individual sentences. in contrast with most other approaches in this issue, the relation between the input sentence and the question is not one of answerhood. rather, the questions generated by the system typically ask for evidence or support for the claim that is made by the input sentence (e.g., input = “cannon (1927) challenged this view mentioning that physiological changes were not sufficient to discriminate emotions.”; output = “why did cannon challenge this view mentioning that physiological changes were not sufficient to discriminate emotions?”). 2.3 question generation for virtual agents the papers in this section are concerned with the generation of questions for the benefit of dialogue agents, and more specifically virtual agents, i.e., dialogue agents that have a computeranimated embodiment. work on such agents faces the challenges of portability and scalability: how to create agents that answer and/or ask questions in new or very large domains. yao, tosch, chen, nouri, artstein, leuski, sagae &traum (“creating conversational characters using question generation tools”) use question generation technology to automatically construct a repository of question-answer pairs from text. a virtual agent can use such a repository to respond to user questions by retrieving an answer from the repository that answers a question that is similar to the user‟s question. yao et al. have built a tool that incorporates the question transducer of heilman and smith (2009) and perform three experiments to gauge the prospects varieties of question generation 7 of this approach for automatically extracting question-answer repositories. their experiments address the issue whether the approach is feasible at all for creating new dialogue characters, and also whether it can be used to enrich manually created repositories (without degrading their quality too much). in contrast with yao et al., mendes, curto & coheur (“question generation based on lexico-syntactic patterns learned from the web”) describe work on a virtual agent which primarily asks rather than answers questions. their work is about automatically harvesting suitable questions for such an agent. the the-mentor platform starts from a set of seed question/answer pairs. it learns patterns between questions and answers based on information obtained from querying the web with queries that have been automatically constructed from the seed questions. the patterns are used to generate fresh questions from new input sentences. evaluation of the qg patterns was performed on the qgstec 2010 development set and a wikipedia page, rating grammatical and semantic correctness. in terms of the qg pyramid (fig. 1), the approach is closest to b. 2.4 question generation shared tasks and evaluation the final paper by rus, wyse, piwek, lintean, stoyanchev & moldovan provides “a detailed account of the first question generation shared task evaluation challenge”. the qgstec was held in 2010 as part of the 3 rd workshop on question generation (boyer & piwek, 2010) and involved two application-neutral tasks: a. question generation from paragraphs and b. question generation from sentences. five teams participated in the qgstec, with approaches ranging from those based on shallow surface features and syntax to deep semantics. human judges evaluated generated questions on several criteria: relevance, question type, syntactic correctness and fluency, ambiguity and variety. additionally, for the qg from paragraphs task, the task was to generate questions with different scope (from broad, paragraph level, to phrase level or less). 3 concluding remarks this special issue presents a range of question generation problems and approaches that are reasonably representative of the current direction of qg research. this research is somewhat biased towards textual input, possibly partly as a result of the recent question generation shared task and evaluation campaign (qgstec; rus et al., this volume). for the qgstec, raw text was chosen as the input format to avoid ruling out participants due to specific theoretical commitments that other input formats (e.g., knowledge representation formalisms) might entail. with semantic web notations (such as owl) gaining widespread acceptance, there is, however, good reason to revisit this decision and investigate whether future challenges could also include tasks based on non-linguistic input. an attractive aspect of this is that genuine knowledge representation inputs allow deeper reasoning about the content of questions and, consequently, also about decisions regarding which question to ask. this, in turn, would make it possible to stimulate a strand of deep question generation in contrast with, for example, the current „generating questions from sentences‟ task which stimulates primarily research that focuses on direct shallow relations between the input text and the output question. though all the approaches in this volume deal with question generation as a computational problem, they draw on other disciplines, notably linguistics, computational semantics, psychology and education. there is, however, less evidence that current work builds on the extensive literature in the field of formal and computational semantics that deals with the semantics and pragmatics of questions (groenendijk & stokhof, 1997; ginzburg & sag, 2001). here, there seems to be a further opportunity, for example, to draw on notions such as partial and conditional piwek and boyer 8 answerhood, to gain a deeper and richer understanding of the relation between questions and their answers. acknowledgements the call for this special issue followed on from the 3 rd workshop on question generation. in response to the call for statements of intent, we received 26 abstracts. some of the submissions, but not all, were based on papers presented at the workshop. in a first round, the special issue editors provided authors with feedback on the suitability of their proposed topic for the special issue. in response, 14 full papers were submitted that then went through a two-stage reviewing process. eventually, 7 papers were accepted for the special issue. each paper was reviewed by three reviewers. paper authors were asked to review one other submission. additionally, we are grateful to the following reviewers: greg aist, lee becker, nadjet bouayad-agha, marc cavazza, mark core, barbara di eugenio, aldabe itziar, rodger kibble, james lester, françios mairesse, smaranda muresan, rodney nielsen, rashmi prasad, and mariët theune. also, we acknowledge the help of anton dil, chris fox, julius goth, david king, chris mitchell, and alistair willis with proofreading the contributions to this issue. finally, we would like to thank the managing editors of dialogue & discourse, in particular, jonathan ginzburg (editor-in-chief) and amanda stent for their support throughout the preparation of this issue. references ahn, r.m.c., beun, r.j., borghuis, t., bunt, h.c. & van overveld, c.w.a.m. (1995). the denk-architecture: a fundamental approach to user-interfaces. artificial intelligence review journal, 8(2): 431-445. beun, r.j. (1990). the recognition of dutch declarative questions. journal of pragmatics, 14:3956. boyer, k. and p. piwek (eds.) (2010). proceedings of the 3 rd workshop on question generation, part of the 10 th conference on intelligent tutoring systems (its 2010), carnegie mellon university, pittsburgh pa., june 18, 2010. bronnenberg, w.j., h.c. bunt, j. landsbergen, p. medema, r. scha, w.j schoenmakers and e. van utteren (1979). the question-answering system phliqa 1. in: l. bolc (ed.), natural communication with computers. macmillan. cohen, f.s. (1929). what is a question? the monist 39: 350-364. copestake, a., d. flickinger, c. pollard, and i. sag (2005). minimal recursion semantics: an introduction. research on language & computation, 3(4): 281-332. cormen, t., c. leiserson, r. rivest and c. stein (2009). introduction to algorithms. mit press. evans, r., p. piwek and l. cahill (2002). what is nlg? proceedings of international natural language generation conference inlg02, new york, usa, 1-3 july 2002. ginzburg, j. & i. sag (2001). interrogative investigations. csli publications/university of chicago press. ginzburg, j. (2012). the interactive stance: meaning for conversation. oxford university press. graesser, a. c., person, n. k., & huber, j. d. (1992). mechanisms that generate questions. in: t. e. lauer, e. peacock, & a. c. graesser (eds.), questions and information systems (pp. 67187). erlbaum. green, b. f., wolf, a. k., chomsky, c. and laughery, k. (1961). baseball: an automatic question answerer. in: proceedings western joint computer conference 19, pp. 219-224. varieties of question generation 9 groenendijk, j. and m. stokhof (1997). questions. in: van benthem, j. and a. ter meulen (eds.), handbook of logic & language. north-holland. heilman, m. and n. smith (2009). question generation via overgenerating transformations and ranking. technical report cmu-lti-09-013, carnegie mellon university language technologies institute. heilman, m. and n. smith (2010). good question! statistical rnaking for question generation. in: proc. naacl/hlt 2010, los angeles. hirschman, l. and r. gaizauskas (2001). natural language question answering: the view from here. natural language engineering 7(4): 275-300. kievit, l., p. piwek, r.j. beun and h. bunt (2001). multimodal cooperative resolution of referential expressions in the denk system. in: bunt, h. and r.j. beun, cooperative multimodal communication, lecture notes in artificial intelligence series 2155, springer, berlin/heidelberg, pp.197-214. marciniak, t. (2008). language generation in the context of yahoo! answers. in: v. rus and a. graesser (eds.), online proceedings of 1 st question generation workshop, september 25-26, 2008, nsf, arlington, va. mcdonald, d. (1993). issues in the choice of a source for natural language generation. computational linguistics, 19(1):191–197. piaget, j. (1952). the origins of intelligence. international university press. piwek, p. (1998). logic, information & conversation. phd thesis, eindhoven university of technology. url: http://alexandria.tue.nl/extra2/9803039.pdf piwek, p., h. prendinger, h. hernault, and m. ishizuka (2008). generating questions: an inclusive characterization and a dialogue-based application. in: rus, v. and a. graesser (eds.), online proceedings of 1 st question generation workshop, september 25-26, 2008, nsf, arlington, va. power, r. (1979). the organisation of purposeful dialogues. linguistics, 17, 107–152. purver, m. (2004). the theory and use of clarification requests in dialogue. phd thesis, university of london, 2004. reiter, e. and r. dale. (2000). building natural language generation systems. cambridge university press, cambridge, uk. rus, v. (n.d.) question generation main website. questiongeneration.org (accessed 19 january 2012). rus, v., a. graesser, a. stent, m. walker & m. white (2007). text-to-text generation. in: dale, r. & m. white (eds.), shared task and comparative evaluation in natural language generation, workshop report, arlington, va. rus, v., z. cai and a. graesser (2008). question generation: example of a multi-year evaluation campaign. in: rus, v. and a. graesser (eds.), online proceedings of 1 st question generation workshop, september 25-26, 2008, nsf, arlington, va. scha, r. (1983). logical foundations for question answering. ph.d. dissertation, university of groningen, m.s. 12.331, philips research laboratories, eindhoven, the netherlands. snover, m., b. dorr, r. schwartz, l. micciulla, and j. makhoul (2006). a study of translation edit rate with targeted human annotation. in: proceedings of association for machine translation in the americas, 2006. stent, a. (2001). dialogue systems as conversational partners: applying conversation acts theory to natural language generation for task-oriented mixed-initiative spoken dialogue. unpublished phd thesis. university of rochester. vauquois, b. (1968). a survey of formal grammars and algorithms for recognition and transformation in machine translation. in: ifip congress-68, pp. 254-260. winograd, t. (1972) understanding natural language. academic press, new york. wisniewski, a. (1995). the posing of questions: logical foundations of erotetic inferences. synthese library / volume 252. kluwer academic publishers. http://alexandria.tue.nl/extra2/9803039.pdf dialogue & discourse 10(1) (2019) 20–33 doi: 10.5087/dad.2019.102 using gestures to resolve lexical ambiguity in storytelling with humanoid robots catelyn scholl cjscholl@uwm.edu computer science university of wisconsin-milwaukee susan mcroy mcroy@uwm.edu computer science university of wisconsin-milwaukee editor: manfred stede submitted 02/2018; accepted 01/2019; published online 02/2019 abstract gestures that co-occur with speech are a fundamental component of communication. prior research with children suggests that gestures may help them to resolve certain forms of lexical ambiguity, and might also help them learn homophones. to test this idea in the context of human-robot interaction, the effects of iconic and deictic gestures on the understanding of homophones was assessed in an experiment where a humanoid robot told a short story containing pairs of homophones to small groups of young participants, accompanied by either expressive gestures or no gestures. both groups of subjects completed a pretest and post-test to measure their ability to discriminate between pairs of homophones and we calculated aggregated precision. the results show that the use of iconic and deictic gestures aids in general understanding of homophones, providing additional evidence for the importance of gesture to the development of children’s language and communication skills. keywords: lexical ambiguity, speech prosody, homophones, iconic gestures, deictic gestures 1. introduction we have been using humanoid robots to teach children about computers, showing them how programs can control how robots move, speak, or play music. we have noticed that children across many age groups were more attentive to the robot sessions than to other demonstrations of computers. we have also noted that these children were not especially interested in complex spoken interaction with the robot on its own and have wondered about the potential impact on attention or learning if we were to add coordinated gestures to the robot’s spoken behaviors. it is well known in human interaction that using gesture is not just a matter of naturalness; body movement or gestures are an important form of communication, especially for young children or new learners of a second language (kidd and holler, 2009; ray, 2015). at some point, teachers might want to incorporate robots into their classrooms as a way of enhancing student’s language skills. thus, it would be important to assess the potential benefit of adding gestures to robot-human communication with children. c©2019 catelyn scholl and susan mcroy this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). using gestures to resolve lexical ambiguity in storytelling with humanoid robots our study attempts to measure the potential benefit of adding gesture to robot speech to help discriminate among homophones, where the target population is young children who are not yet fully fluent in their own language. the assessment involves storytelling, which would be a normal activity for this age group and also something that one might want to incorporate in classroom activities. homophones occur when two words are pronounced alike but have different meanings, derivations or spellings. an example of a homophone would be blue, the color, and blew, the past form of the verb blow. given this task, we then consider what types of gestures might be helpful and how these gestures can be timed appropriately with speech. 1.1 background gestures come in many types. some gestures are used instead of speech and some are meant to enrich it. for example, ekman and friesen (1969) define five types of gestures for communication of content and emotion, as well as for non-communicative functions. these types include emblems (for directly replacing some words), illustrators (for shaping what is being said), affect displays (for showing emotion), regulators (for controlling the flow of conversation) and adapters (for helping to relieve the speaker’s own tension). if one focuses on co-speech gestures, the gestures that accompany and enrich speech rather than replace it, these are most often divided into four subtypes, following mcneill (1992). these four subtypes are: iconic (for representing object attributes), deictic (for relating one entity to another by pointing), metaphoric (for relating an abstract idea to a concrete one by creating a picture) and beat gestures (for keeping the rhythm of the speech, without expressing content.) furthermore, according to mcneill, these sorts of gestures and speech involve the same mental process and thus are naturally temporally and semantically related. indeed among children less than 18 months old, vocabulary acquisition has been found to be better when words were accompanied by appropriate iconic gestures than when used without gesture or with an uncoordinated gesture (see namy et al., 2000; zammit and schafer, 2011). gestures can help convey meaning by helping listeners to resolve lexical ambiguity or to better understand and acquire unfamiliar vocabulary. gestures may also enhance attention by making the communication livelier or more memorable. the use of gesture is one of the first forms of intentional co-ordination of behavior that people learn. gesture appears to precede early verbal language development and it has been surmised that the use of pointing gestures may act as a precursor to initial acquisition of single words, so as children point at things, people tell them what they are called. children also combine multiple gestures representing different concepts before producing two-word utterances, (kidd and holler, 2009). past work with young children and the use of gesture during storytelling has focused on the use of deictic and iconic gestures, as they convey meaning but do not require a complex mapping between abstract and concrete concepts. for example, kidd and holler (2009) considered children’s own use of gesture while retelling a story containing homonyms, which is where the same word can have two very different meanings, for example bat as an animal versus as an object. their study involved three to five -year-old children. for the study, the researchers created short story books containing age-appropriate homonyms such as bat, mouse, and glass. the stories were read to the children and then the children were asked to retell the story while the researchers considered how they resolved the homonym ambiguity. the results revealed that the youngest children, 3-year-olds, primarily used nouns and deictic gestures as their means of disambiguation. the 4-year-olds added more iconic gestures. the 5-year-olds mainly relied on speech alone. 21 scholl and mcroy another study, involving children who were telling stories, investigated the use of gesture in resolving lexical ambiguity among english-as-a-second-language (esl) learners (ray, 2015), the results suggest that these learners frequently used gestures when telling stories, and younger age groups used gesture more frequently than older groups. the main difference between children who are second language learners versus first language speakers seems to be that second language learners, especially those who appeared to have the most difficulty with speaking the second language, used iconic gestures more frequently than deictic gestures. no work has been done to measure the use or impact of gestures on lexical disambiguation for robot-human communication with children. one study, by kory westlund et al. (2017), assessed the importance of speech prosody on children’s acquisition of new vocabulary and their ability to recall and retell a story. the study compared the impact of “flat” (non-prosodic) versus “expressive” (with prosody and gestures) storytelling. their findings suggest that children were more effective in using newly acquired vocabulary after listening to the expressive robot. furthermore, during a delayed recall test, children who listened to the expressive story had less difficulty retelling the story than those who had heard the flat storytelling. there is another gesture-related study with robots (see van dijk et al., 2013) that considers the impact of robot-generated iconic gestures during interaction with older adults. that study found that the gestures helped in the retention of information. gestures may help all listeners (old and young) in retaining information about a story including the semantic context necessary to understand lexical ambiguities such as homophones. 1.2 overview of the present study the goal of our work was to assess the hypothesis that robot-to-human speech accompanied by gesture would improve children’s ability to disambiguate homophones more than without gestures. we conducted an experiment using humanoid robots to tell stories to preschool and early elementary age children and assessed the children’s ability to discriminate homophones both before and after experiencing the story. our experiments included only iconic and deictic gestures because they carry the most concrete semantic information and thus are likely to be the most understandable to young children. we selected an existing story from the children’s literature appropriate for the ages of the subjects recruited for the study. to make the storytelling more natural, in both conditions we implemented the tempo aspects from an existing rule-based model of speech prosody for storytelling when converting the text to speech (theune et al., 2006). this approach ensures that the storytelling is not too fast for the anticipated skill level of listeners and that the boundaries of phrases and sentences are clear. our implementation does not put any stress on any of the words, including the homophones, that might draw attention to them. 2. methods we used a humanoid robot, programmed to read the published children’s story dear deer: a book of homophones, (barretta, 2010). table 1 includes some sample sentences from this work. where possible, we added simultaneous movements by the humanoid robot corresponding to different related gestures. for assessment, we created a pretest and a post-test using a subset of the vocabulary from the story and wrote our own simple sentences. after approval by our institutional review board, we recruited subjects from a university-run daycare center and conducted our experiment on site. more detail is provided below. 22 using gestures to resolve lexical ambiguity in storytelling with humanoid robots the moose loves mousse. he ate eight bowls. have you seen the ewe? she’s been in a daze for days. the toad was towed to the top of the seesaw, so he could see the sea. the whale was allowed to wail aloud. the bear had to pause to bare his big paws. hey, the elephant threw a pail through the big bale of hay! have you read about the red fox who blew blue bubbles? table 1: sample sentences from dear deer 2.1 the robot framework two separate presentations were created to tell the story. the first presentation was for the control group, which contained only speech for the homophone pairs. the second presentation, for the experimental group, contained speech and gestures that corresponded with some of the homophones. we used a nao robot and its accompanying software, called choreograph, to implement the story teller. choreograph software allows one to create timed sequences of actions by the robot, including simultaneous physical motions and speech. some poses, such as sitting, are provided as part of the library. choreograph software also supports something called animation mode, which allows one to define new actions using the real robot like a puppet. in this mode, one can manually move any part or set of parts and save the joint settings as an object, called a “box”. then, one can create a sequence of boxes called a “timeline” that combines predefined behaviors or poses (such as “say” or “sit”) and ones created using animation mode. one can also control the duration of each behavior, so that a particular movement or spoken phrase can be synchronized (e.g. by having them start at the same time and have approximately the same duration). once set, the robot will perform the story the same way each time. each presentation (speech and gesture) was created using a combination of functions from the choreograph library and some additional python scripts. specifically we used the existing functions to control text-to-speech, but added some python scripts to calculate the correct timing for speech and pauses, based on theune’s model. a few physical motions (like sitting, and opening or closing a hand) were used as predefined, but most required the use of animation mode to approximate the gesture. then all the actions were manually placed in a sequence. adjustments to the timing of the gestures in a timeline were made by watching the robot tell the story and manually adjusting the timing to provide the most natural observer experience. the final parameterized timelines (with and without gestures) were saved as separate objects, that when initiated, cause the nao robot to present the complete story without additional interaction. in both conditions, the robot begins in its initial position (a crouch), and moves to a sitting position to tell the story. 2.2 speech prosody to implement more natural sounding phrasing when converting text from the short story to speech, we used part of a rule-based framework for storytelling prosody (theune et al., 2006). this work provides rules for a global storytelling speaking style, which specify a consistent and natural set of 23 scholl and mcroy speech variations, where the goal is to better emulate natural storytelling versus read speech. we used only the rules for controlling the timing of phrases and pauses, both to seem more natural and also to help clarify the syntax. we implemented these rules directly in the text-to-speech api for nao robots, adjusting the duration of syllables. all other aspects of speech, such as pitch or stress on vowels or syllables, were controlled by the standard api. we did not implement theune’s other rules for pitch or intensity as these aspects seemed adequate for our purposes. we did not make any specific adjustments for homophones, as in some sentences of the story nearly every word is a homophone (e.g. he ate eight bowls. and the whale was allowed to wail aloud.). the timing rules, as we implemented them, are summarized below: 1. a rate of about 3.6 syllables per second was used for the tempo. this was found by counting the syllables in a few trial sentences and averaging the time it took to read the sentences. 2. the duration of accented vowels was 1.5 times their average duration. this was done using nao’s text-to-speech api. 3. for pauses, after each phrase, we added a break of a length of 0.4s and between sentences we added a break of length of 1.3s. these pauses were also created using nao’s text-to-speech api. 2.3 gestures to accompany the homophones the children’s story dear deer: a book of homophones (barretta, 2010), is a story created with the intention of introducing homophones, by incorporating pairs of homophones with pictures and sentence content to help discriminate the respective meanings. telling the story typically takes an adult three to five minutes (see jean, 2015). the story is 241 words long and has a reading lexile level of 530l and an atos reading level of 2.1.i which means that it is considered appropriate for a seven year old to read independently (teachingbooks.net, 2018). the story contains the 32 homophone pairs listed in table 2, shown with their primary part of speech. each sentence of the story contains from one to three homophone pairs (most contain two). most pairs occur only once except for “dear-deer” and “aunt-ant”, which are each repeated once. manual review of the pairs suggested that for 11 of the pairs at least one homophone sense used in the story had a possible iconic or deictic gesture and one pair had two, one for each sense. four of the homophone senses involved actions for which there is an associated body part (hear-ear, seeeye, ate-mouth, wail-eyes). three were a body part (hair, feet, paws). three seemed naturally deictic (here, you, him). one other was an animal with a distinctive feature (moose-animal with big antlers) and the other an adjective associated with a body part (hoarse-throat). the specific gestures were chosen to be ones that the authors felt a human reader might use, based on the lexical meaning of the homophone and the typical way that such meanings are depicted. past research suggests that deictic and iconic gestures are used by even very young children and might include gestures like pointing towards a foot to indicate the body part, feet (see iverson et al., 1994). the gestures chosen were also ones that could be performed given the capabilities of the humanoid robot. the flexibility and motor control of the robot limit it to moving its arms or hands in different directions (left, right, up, down), at various angles, and can be positioned manually to reach towards one of its parts (head, feet, eyes, mouth, ears, neck, torso) and either holding the appendages still (hold, point), rotating them slightly to simulate rubbing (rotate) or letting them 24 using gestures to resolve lexical ambiguity in storytelling with humanoid robots allowed (v) aloud (adv) hey (adv) hay (n) ate (v) eight (n) him (n) hymn (n) aunt (n) ant (n) horse (n) hoarse (adj) bear (n) bare (adj) kneaded (v) needed (v) bee (n) be (v) know (v) no (adv) blew (n) blue (adj) mood (n) mooed (v) choose (v) chews (v) moose (n) mousse (n) daze (n) days (n) news (n) gnus (n) dear (adj) deer (n) paws (n) pause (n,v) doe (n) dough (n) read (v) red (adj) feat (n) feet (n) sea (n) see (v) flew (v) flu (n) tale (n) tail (n) flea (n) flee (v) threw (v) through (prep) hair (n) hare (n) toad (n) towed (v) hear (v) here (pron) whale (n) wail (n,v) herd (n) heard (v) you (pron) ewe (n) table 2: pairs of homophones, with parts of speech for each, where n is noun, v is verb, pron is pronoun, prep is preposition, adj is adjective, and adv is adverb. drop to end the gesture (move). if there was no plausible, achievable gesture then the arms of the robot would remain down and still during that part of the sentence. overall, 11 pairs were chosen and 12 corresponding iconic or deictic gestures were selected. table 3 provides a dictionary of gestures for the homophones that were identified as having a corresponding gesture and a brief description of the gesture that was used. the robot was programmed to implement the gesture associated with the word so that it would occur at approximately the same time as the corresponding word was spoken. the timing of the gestures was verified by videorecording the robot telling the story with the gestures and asking adult observers to view the video (on youtube) and judge the naturalness of the timing of the gestures. each judge was asked to rate the timing according to the following rubric: the timing of the gestures was poor they were distracting to watch. the story would be better with no gestures at all. the timing of the gestures was okay a few times they were noticeably too early or too late. the timing of the gestures was good they seemed natural, but at least once i felt that the timing could be improved. the timing of the gestures was great none of them was noticeably out of sync with what was being said. judges were also asked to comment on what issues with timing might be improved. no major issues were revealed. 25 scholl and mcroy homonym gesture type gesture hear-here iconic hold hand near ear hear-here deictic move hand toward torso moose-mousse iconic hold hands up like antlers ate-eight iconic move hand toward mouth you-ewe deictic point hand away him-hymn deictic point hand away hoarse-hoarse iconic/deictic hold hand near neck feat-feet iconic/deictic point hand toward feet sea-see iconic hold hands above eyes whale-wail iconic rotate hands near eyes paws-pause iconic/deictic hold hands near torso hair-hare iconic/deictic hold hands near head table 3: gesture dictionary 2.4 recruitment, experimental procedures, and data analysis we recruited participants from two different classrooms within the children’s learning center of uw-milwaukee, a large public university in the midwestern united states. parental consent forms were distributed to parents of all enrolled children between the ages of 4 and 8 before conducting the research study. for each child to participate, the parent of a child was required to sign and return the permission letter and the child had to give his or her own verbal consent. children in the study were randomly assigned to one of two groups: either the control (without gestures) or the experimental condition (with gestures). the robot performed the story to the two groups separately on two consecutive friday afternoons, during their regularly scheduled time at the children’s center. children in both groups were first given a pretest and afterward a post-test, while sitting at small tables of three to four people with a staff member at each table to help them. the pretests and posttests both comprised the same set of ten items each of which included two full-color cartoon-style images, selected to depict two distinct words that are homophones, similar to the illustrations used in published versions of the story. table 4 includes the sentences and a brief description of the images from which the children could choose. the children were read the sentences by an adult and asked to circle the image that they felt best corresponded to the meaning of the sentence. after the pretest, the children moved to a nearby area where they were were seated on the floor near the robot. there, they were introduced to the researcher and the robot. then they heard the robot tell the story which took about three minutes. after the story, the children returned to the tables with the caregivers for the post-tests. precision was measured for each condition by calculating the number of correct responses divided by the total number of responses for each condition, aggregated over the entire group and the complete set of items. to account for possible differences between the two groups besides the condition, we calculated two measures of significance: one between the post-test values of precision for the control and experimental groups and one between the pretest and post-test values within the experimental group alone. we also looked at the differences between the number of incorrect versus 26 using gestures to resolve lexical ambiguity in storytelling with humanoid robots sentence pair of images 1. the sea was as smooth as glass. ocean waves, eye 2. he ate twice the amount that you did. boy eating sandwich, number eight 3. she didn’t come here to talk to me. girl pointing to floor, hand cupping ear 4. she felt the hair rising on the back of her neck. hair without face, rabbit 5. hey, where are you going? boy waving, bale of yellow hay 6. a whale can be very loud underwater. boy with tears, whale spouting water 7. we had some mousse for dessert. walking moose, dish of swirled brown food 8. you again? generic sheep, hand pointing towards reader 9. the teacher’s criticisms were enough to make anyone red. girl reading book, splat of red color 10. i was driving along the road when a deer jumped out at me out of the blue. boy blowing candles, splat of blue color table 4: test sentences with descriptions of image choices correct responses for individual items, aggregated over the entire group. any response that had two circles was counted as incorrect based on the assumption that the child was unable to disambiguate the lexical ambiguity of the sentence. 3. results twenty-two adults from a university course on natural language processing viewed a video of the robot storytelling and assessed the timing of the gestures. most (18 of 22) rated the timing of the gestures as natural (“great”: 14; “good”: 4 ; “okay”: 2; “poor”: 2). among those who gave it a lower rating, the duration of two of the gestures was cited, including two who said pointing to the feet was too slow and three who said pointing to the eyes was too quick. however, no major problems with the gestures was revealed. the children in the homophone study ranged from ages three to seven. subjects in the control group had ages ranging from three to five, while subjects in the experimental group had ages ranging from four to seven. 13 students were tested in the control group pretest and 12 students were tested in the control group post-test. thus the aggregated number of responses for the control pretest was 130 and for the control post-test it was 120. 8 students were tested in the experimental group, i.e. the group with gesture, both pretest and post-test. thus for both the pretest and post-test there were 80 responses. table 5 shows the pretest and post-test precision for both the control (no gesture) and experimental (with gestures added) condition. the total number of correct over the total number of responses is shown in parentheses. pretest con post-test con pretest exp post-test exp 0.762 (99/130) 0.633 (76/120) 0.875 (70/80) 0.925 (74/80) table 5: pretest and post-test precision for control (con) and experimental (exp) conditions 27 scholl and mcroy in the control group, aggregated precision dropped by 0.109, which is 14.3 percent; in the experimental group, the aggregated precision increased by 0.050, which is 5.7 percent. the difference in aggregated precision between the control and experimental groups for the post-test is significant (p < .001). the increase in aggregated precision within the experimental group is promising, but insufficient to reject the null hypothesis (.1 < p < .2). looking at individual items, for the control group, the number of correct responses increased for two items (here-hear and red-read), stayed the same for two items (eight-ate and moose-mousse) and decreased for 6 items. by contrast, for the experimental group, the number of correct responses increased for 4 items, stayed the same for 6 items, and never decreased. (charts capturing the complete results of all the pretests and post-tests for both conditions are provided in appendix a.) in the control group (hearing the story without gestures), there were several homophones that a majority of the children found difficult. in the pretest, children struggled most with moose-mousse, which was the only pair that had more incorrect responses than correct responses (9 incorrect, 4 correct), followed by read-red (6 incorrect, 7 correct), and then here-hear (4 incorrect, 9 correct). (none of these pairs involves a syntactically ambiguous word; for the target audience both “moose” and “mousse” might be rare.) the children appeared to excel at understanding wail-whale (100 percent of the students got this correct), sea-see, and blue-blew. overall, the children in the pretest control group had 99 correct responses and 31 incorrect responses. the control group post-test had significantly different results. children still struggled with moose-mousse (9 incorrect, 3 correct). the results showed improvement among the homophone here-hear (22 percent increase), but a slight decrease in most other homophone pairs. in the post-test control group, there were 77 correct responses and 43 incorrect responses. in the pretest of the experimental group, children had some difficulty with here-hear, heyhay, and mouse-mousse. despite having some mistakes, however, none of the pairs were chosen incorrectly more than correctly and none involves a syntactic ambiguity. there were a total of 6 incorrect homophone pairs (sea-see, eight-ate, here-hear, hey-hay, moose-mousse, and readred), with 10 incorrect responses total. the children did have a few perfect scores as well for hair-hare, ewe-you, blue-blew, and whale-wail. in the post-test, the experimental group showed more uniform improvement. there was an overall decrease in the number of incorrect homophone pairs, which was 4 pairs and 6 incorrect responses total out of 80 responses. the children had 100 percent accuracy for six of the ten items: eight-ate, hair-hare, whale-wail, ewe-you, read-red, and blew-blue. two of the items answered correctly in the post-test had been among those that were incorrect in the pretest (eight-ate, read-red); neither involves a syntactic ambiguity. pairs that were still sometimes resolved incorrectly involved words that might be more unusual and also harder to gesture (hey-hay, here-hear, moose-mousse); only one was not also discriminated by the syntactic context (moose-mousse). 4. discussion the results of the study suggest that adding coordinated gestures can have a positive impact on children’s ability to distinguish between homophones in the context of robot storytelling, and potentially more broadly. our work also suggests that studying competency in understanding homophones is an interesting measure of human language development. as we noted earlier, much past work with children focuses on children’s own use of gestures to discriminate meaning (see kidd and holler, 2009; ray, 2015), which would suggest that children 28 using gestures to resolve lexical ambiguity in storytelling with humanoid robots can understand gestures at young ages. studies directly related to children’s perceptions of gesture have considered either vocabulary acquisition or object labeling. vocabulary acquisition among infants (ages 9 to 15 months), measured as the number of new words produced or recalled, has been found to be larger when mothers use iconic gestures versus uncoordinated gestures, e.g. (zammit and schafer, 2011). recall of object labels among infants is also improved by parents’ use of iconic gestures (namy et al., 2000). similarly, the work by kory westlund et al. (2017) found that among preschoolers (average age of five years) vocabulary acquisition was better when social robots used gestures. none of this prior work has directly addressed the challenge of understanding homophones. our results reveal a clear benefit to gestures in homophone discrimination. precision in the experimental group, which saw the gestures, increased from an already high value of 0.875 to a near-perfect precision of 0.925. for only two items (here-hear and hay-hey) did more than one child make a mistake on the post-test. this improvement is in sharp contrast to the results in the control group where the responses seemed somewhat random and actually declined from pretest to post-test. the only item where the performance of the control group improved was on arguably the hardest item (here-hear), which had the most errors in the pretest of the experimental group. we did not make any adjustments to either the pitch or intensity of the homophones, as it would not be possible to do so uniformly given the sometimes high density of homophones within a sentence. however, in the future we would like to look at increasing the pitch or intensity of homophones that are not surrounded by other homophones, especially those that occur at the beginning or end of syntactic phrases. this method of emphasis seems to be used by human storytellers reading dear deer. also the theune et al. (2006) model suggests rules for these, although the researchers themselves did not evaluate their impact at the time due to limitations in their text-to-speech software. we acknowledge several limitations of the study. the number of children in both groups was small. also, the number in the experimental group was smaller than in the control. we found predicting attendance ahead of time can be difficult, as attendance at a university-based daycare center can vary significantly, due to schoolwork or travel, which affects both student and faculty parents. we suspect that having a larger number of children in the control group, including more children who were at the low end of the age range and fewer at the high end, likely contributed to some of the variability we saw in the pretest to post-test results for the control versus the experimental group. the children in both groups were relatively young and this factor, along with possible excitement or fatigue, might have affected their ability to understand or follow instructions. to reduce this risk, the children were seated at different tables in groups of around four, with a staff person at each table to explain the procedure and remind the children not to discuss their answers. we did observe, however, that occasionally the children would not comply and had to be reminded. thus, while some sharing of answers might have occurred, it was likely limited to a few items, and only among children within a table group. the results for any particular pair of homophones may have been influenced by a number of potential linguistic effects that would occur when reading an existing story to children. the homophones might vary in their difficulty for children, even more than for adults, either because some words might be more or less familiar to an individual child (i.e. people’s lexicons grow as they interact with language) or because the words vary in their linguistic features (e.g. some are nouns, some are verbs, some can be both). the homophones chosen to be gestured were based on the feasibility of coming up with a coordinated gesture given the limitations of the robot and the phrasing of the 29 scholl and mcroy story. the test sentences were chosen such that only one of the answers would be semantically appropriate; in most cases (6 out of 10) only one answer would have correct syntax, based on judgment of an adult native speaker of english, with expertise in linguistics (the second author). thus, the linguistic context surrounding each homophone (either in the story or in the test sentences) could also potentially have had an impact. whether a given child is aware of any of these aspects is something that would be hard to predict, however, as children’s ability to recognize linguistic features has been found to improve as they mature (clahsen et al., 2007). measuring understanding individually and dynamically, such as through event-related brain potentials (erp), might be interesting, but it would be a very different study. such measurement would also not contribute to the motivating goal of our work, which is to see if adding gestures to a humanoid robot reading a story in a classroom setting would be educationally worthwhile. that goal was met, as the change in performance from pretest to post-test was much better in the experimental group than in the control. the results on the pretests were very different between the control and the experimental groups, with the experimental group having higher overall precision on the pretest. this difference might be explained by the natural variability in ability among new language learners or by differences in the distribution of ages within the two groups. we observed a slightly higher proportion of older children (more six or seven year olds) in the experimental group than in the control. this could be a factor, as prior research suggests that some children’s ability to detect lexical ambiguity does not emerge until around six years of age, as reported by (kidd and holler, 2009); however they did find some subjects as young as three could understand and perform gestures appropriately. some of these factors, such as the age of the children, are unavoidable if one wants to study children’s language development. however, in a future study, one could do a similar experiment using esl learners, as these subjects would be more mature and less prone to outside distraction. some of the limitations could be addressed by conducting a broader study with assignments to groups conditioned on the results of pretests conducted ahead of time. another aspect to explore would be the types of gestures used. using humanoid robots offers the opportunity to control exactly what gestures are produced when, and to assure they are repeated exactly the same way each time. currently, we do not know if the semantics of iconic and deictic gestures is more important than just the fact that using multiple modalities, in an apparently coordinated way, creates an emotional response that enhances learning and memory. it may also help to keep young children’s attention to other supporting details. in a future study, it would be interesting to compare the effect of semantic gestures to the use of coordinated, but simpler, gestures. simple beat gestures, such as moving arms up, down, or forward, in time with the phrasing, might be sufficient to enhance engagement. (one of the judges who viewed our video of the robot gestures reported that it seemed unnatural when the robot was not using any gestures.) if coordinated beat gestures are enough, then, in addition to assessing the impact on understanding homophones, one could explore how robot gesture affects the retention of narrative details, using a story that has distinct characters, setting, or events. 5. conclusion this paper describes an experiment conducted in the context of human-robot interaction, to assess the impact of iconic and deictic gestures on the understanding of homophones by young children. we used a humanoid robot and its associated software framework to present a short story containing pairs of homophones to small groups of children, accompanied by either these expressive gestures 30 using gestures to resolve lexical ambiguity in storytelling with humanoid robots or no gestures. what we found is that the use of iconic and deictic gestures appears to aid these language learners in their general understanding of homophones. this work thus provides valuable additional evidence for the importance of gesture to the development of children’s language and communication skills. appendix a: pretest and post-test results for reading study 0 2 4 6 8 10 12 14 pre-test without gesture correct incorrect 0 2 4 6 8 10 12 post-test without gesture correct incorrect 31 scholl and mcroy 0 1 2 3 4 5 6 7 8 9 pre-test with gesture correct incorrect 0 1 2 3 4 5 6 7 8 9 post-test with gesture correct incorrect acknowledgments the authors thank the staff and families of uw-milwaukee children’s learning center for the use of their facility and their assistance and enthusiastic support for the study. 32 using gestures to resolve lexical ambiguity in storytelling with humanoid robots references gene barretta. dear deer: a book of homophones. square fish, 2010. a previous hard copy edition was published by holt & company in 2007. harald clahsen, monika lück, and anja hahne. how children process over-regularizations: evidence from event-related brain potentials. journal of child language, 34(3):601–622, 2007. doi: 10.1017/s0305000907008082. paul ekman and wallace v. friesen. the repertoire of nonverbal behavior: categories, origins, usage, and coding. semiotica, 1(1), 1969. jana m. iverson, olga capirci, and m. cristina caselli. from communication to language in two modalities. cognitive development, 9(1):23 – 43, 1994. issn 0885-2014. doi: 10.1016/0885-2014(94)90018-3. angelina jean. dear deer, by: gene barretta, read by: angelina jean, 2015. url https: //www.youtube.com/watch?v=jyamp36d8s0. accessed august 2, 2018. evan kidd and judith holler. children’s use of gesture to resolve lexical ambiguity. developmental science, 12(6):903–913, 2009. jacqueline m. kory westlund, sooyeon jeong, hae w. park, samuel ronfard, aradhana adhikari, paul l. harris, david desteno, and cynthia l. breazeal. flat vs. expressive storytelling: young children’s learning and retention of a social robot’s narrative. frontiers in human neuroscience, 11:295, 2017. issn 1662-5161. doi: 10.3389/fnhum.2017.00295. david mcneill. hand and mind: what gestures reveal about thought. university of chicago press, chicago, il, 1992. laura namy, linda acredolo, and susan goodwyn. verbal labels and gestural routines in parental communication with young children. journal of nonverbal behavior, 24(2):63–79, 2000. elizabeth m. ray. gestures used by esl children to resolve lexical ambiguity. master’s thesis, ohio university, 2015. teachingbooks.net. website, 2018. url https://www.teachingbooks.net/tb.cgi? tid=16401. search key: dear deer. accessed august 2, 2018. mariët theune, koen meijs, and dirk heylen. generating expressive speech for storytelling applications. ieee transactions on audio, speech and language processing, 14(4):1137–1144, 2006. elisabeth t. van dijk, elena torta, and raymond h. cuijpers. effects of eye contact and iconic gestures on message retention in human-robot interaction. international journal of social robotics, 5(4):491–501, nov 2013. issn 1875-4805. doi: 10.1007/s12369-013-0214-y. maria zammit and graham schafer. maternal label and gesture use affects acquisition of specific object names. j child lang, 38(1):201–221, 2011. doi: 10.1017/s0305000909990328. 33 dialogue & discourse 9(2) (2018) 35–79 doi: 10.5087/dad.2018.202 towards integration of cognitive models in dialogue management: designing the virtual negotiation coach application andrei malchanau andrei.malchanau@lsv.uni-saarland.de spoken language systems group saarland university, germany volha petukhova v.petukhova@lsv.uni-saarland.de spoken language systems group saarland university, germany harry bunt harry.bunt@uvt.nl tilburg center for communication and cognition tilburg university, the netherlands editor: massimo poesio submitted 06/2017; accepted 11/2018; published online 01/2019 abstract this paper presents an approach to flexible and adaptive dialogue management driven by cognitive modelling of human dialogue behaviour. artificial intelligent agents, based on the act-r cognitive architecture, together with human actors are participating in a (meta)cognitive skills training within a negotiation scenario. the agent employs instance-based learning to decide about its own actions and to reflect on the behaviour of the opponent. we show that task-related actions can be handled by a cognitive agent who is a plausible dialogue partner. separating task-related and dialogue control actions enables the application of sophisticated models along with a flexible architecture in which various alternative modelling methods can be combined. we evaluated the proposed approach with users assessing the relative contribution of various factors to the overall usability of a dialogue system. subjective perception of effectiveness, efficiency and satisfaction were correlated with various objective performance metrics, e.g. number of (in)appropriate system responses, recovery strategies, and interaction pace. it was observed that the dialogue system usability is determined most by the quality of agreements reached in terms of estimated pareto optimality, by the user’s negotiation strategies selected, and by the quality of system recognition, interpretation and responses. we compared human-human and human-agent performance with respect to the number and quality of agreements reached, estimated cooperativeness level, and frequency of accepted negative outcomes. evaluation experiments showed promising, consistently positive results throughout the range of the relevant scales. keywords: dialogue management, cognitive agent technology, intelligent tutoring system, spoken/multimodal dialogue system 1. introduction the increasing complexity of human-computer systems and interfaces results in an increasing demand for intelligent interaction that is natural to users and that exploits the full potential of spoken and multimodal communication. much of the research in human-computer system design has been c©2018 andrei malchanau, volha petukhova, and harry bunt this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). malchanau, petukhova and bunt technique example task dialogue phenomena handled finite state script long-distance calling user answers questions frame based getting train arrival and departure information user asks questions, simple clarifications by the system information state update travel booking agent flexible shifts between pre-determined topics/tasks refined grounding mechanisms plan based kitchen design consultant dynamically generated topic structures, e.g. negotiation dialogues agent based disaster relief management different modalities, e.g. planned world and actual world collaborative planning and acting probabilistic approaches various information-seeking tasks, dialogue policies design, i.e. learning negotiation games combined with the most approaches mentioned above chat-oriented; retail ‘chat commerce’ question-answering skills interactive pattern matching psychotherapies, personal assistant social interactive aspects /template-based table 1: state-of-the-art techniques for task-oriented dialogue system. conducted in the area of task-oriented systems, especially for information-seeking dialogues concerning well-defined tasks in restricted domains – see table 1 for the main paradigms used for dialogue modelling in domains of varying complexity. many existing systems represent a set of possible dialogue state transitions for a given dialogue task. dialogue states are typically defined in terms of dialogue actions, e.g. question, reply, inform, and slot-filling goals. states in a finite state transition network are often used to represent the dialogue states (bilange, 1991; dahlbäck and jönsson, 1998). some flexibility has been achieved when applying statistical machine learning methods to dialogue state tracking (williams et al., 2013). statistical dialogue managers were initially based on markov decision processes (young, 2000) where given a number of observed dialogue events (often dialogue acts), the next event is predicted from the probability distribution of the events which have followed these observed events in the past. partially observable markov decision processes (williams and young, 2007) model unknown user goals by an unknown probabilistic distribution over the user states. the pomdp approach is considered as the state-of-the-art in task-oriented spoken dialogue systems, see young et al. (2013). however, when dealing with real users, the defined global optimisation function poses important computational difficulties. recently, deep neural networks have gained a lot of attention (henderson et al, 2013; 2014). hierarchical recurrent neural networks have also been proposed to generate open domain dialogues and build end-to-end dialogue systems trained on large amounts of data without any detailed specification of information states (serban et a., 2016). the real challenge for end-to-end frameworks is however the decision-taking problem related to the dialogue management for goal-oriented dialogues. statistical and end-to-end approaches require really large amounts of data, while offering a rather limited set of dialogue actions (kim et al., 2015). while such dialogue systems may perform well on simple information-transfer tasks, they are mostly unable to handle real-life communication in complex settings like, for example, multi-party conversations, tutoring sessions and debates. more conversationally plausible dialogue models are based on rich representations of dialogue context for flexible dialogue management, e.g. information-state updates (isu, traum et al., 1999; bunt, 1999; bos et al., 2003; keizer et al., 2011). other approaches to dialogue processing and management are built as full models of rational agency accounting for planning and plan recognition (cohen and perrault, 1979; carberry, 1990; sadek, 1991). plan construction and inference are activities that can however easily get very complex and become computationally intractable. alternatively, dialogue plans and strategies can be learned and adapted through reinforcement learning (sutton and barto, 1998). however, this seems to require even greater amounts of data, henderson et al. (2008). 36 towards integration of cognitive models in dialogue management the research community is currently targeting more flexible, adaptable, open-domain multimodal dialogue systems. advances are made in modelling and managing multi-party interactions, e.g. for meetings or multi-player games, where approaches developed for two-party dialogue have to be extended in order to model phenomena specific to multi-party interactions. nevertheless, simple command/control and query/reply systems prevail. some dialogue systems developed for research purposes allow for more natural conversations, but they are often limited to a narrow manually crafted domain and to rather restricted communication behaviour models, e.g. often modelled on information retrieval tasks. in some cases, these restrictions are imposed deliberately by the researchers to be able to investigate a limited set of dialogue phenomena without having to deal with unrelated details. however, this reduces the practical realism of the dialogue system. expectations of the users of today are rather high, requiring a real-time engagement with highly relevant personalized content that mimics human natural behaviour and is able to adapt to changing user needs and goals. nowadays, there is a growing interest in artificial intelligence (ai)-powered conversational systems that are able to learn and reason, to facilitate realistic interactive scenarios with realistic assets and lifelike, believable characters and interactions. ai models may represent rather complex research objects. despite their acknowledged potential, generating plausible ai models from scratch is challenging. for instance, cognitive models were successfully integrated into intelligent tutoring and intelligent narrative systems, see paiva et al. (2004); riedl and stern (2006); vanlehn (2006); ritter et al. (2007); lim et al. (2012). since such models produce detailed simulations of human performance encompassing many domains such as learning, multitasking, decision making, and problem solving, they are also perfectly capable to play the role of a believable human-like agent in various human-agent settings. although the abilities of cognitive agents continue to improve, human-agent interaction is often awkward and unnatural. the agents most of the time cannot deliver human-like interactive behaviour, but deal well with task actions thanks to the use of well-defined computational cognitive task models. this paper presents an approach to the incorporation of cognitive task models into information state update (isu) based dialogue management in multimodal dialogue systems. such integration has important advantages. the isu methodology has been applied successfully to a large variety of interactive tasks, e.g. information seeking (keizer et al., 2011), human-robot communication (peltason and wrede, 2011), instruction giving (lauria et al., 2001), and controlling smart home environments (bos et al, 2003). several isu development environments are available, such as trindikit (larsson and traum, 2000), dipper (bos et al., 2003) and flores (morbini et al., 2014). the isu approach provides a flexible computational model for understanding and generation of dialogue contributions in term of effects on the information states of the dialogue participants. isu models account for the creation of (shared) beliefs and mechanisms for their transfer, and have welldefined machinery for tracking, understanding and generation of natural human dialogue behaviour. cognitive modelling of human intelligent behaviour, on the other hand, enables deep understanding of complex mental task processes related to human comprehension, prediction, learning and decision making. threaded cognition (salvucci and taatgen, 2008) and instance-based learning (gonzalez and lebiere, 2005) models developed within the act-r cognitive architecture (anderson, 2007) are used to design a cognitive agent that can respond and adapt to new situations, in particular to a communicative partner changing task goals and strategies. the agent is equipped with theory of mind skills (premack and woodruff, 1978) and is able to use its task knowledge not only to determine its own actions, but also to interpret the human partner’s actions, and to adjust its behaviour to whom it interacts with. in this way, we expect to achieve flexible adaptive dialogue 37 malchanau, petukhova and bunt system behaviour in dynamic non-sequential interactions. the integrated cognitive agent does not only compute the most plausible task action(-s) given its understanding of the partner’s actions and strategies, provides alternatives and plans possible outcomes, but it also knows why it selects a certain action and can explain why its choices lead to the specific outcome. this enables the agent to act as a cognitive tutor, supporting the development of the (meta)cognitive skills of a human learner. finally, the agent can be built with rather limited real or simulated dialogue data: it is supplied with initial state-action templates encoding domain knowledge and the agent’s preferences, and the agent further learns from the collected interactive experiences. the present study investigates the core properties of cognitive models that underlie human task-related and interactive dialogue behaviour, shows how such models provide a basis for dialogue management and can be integrated into a dialogue system, and assesses the resulting system usability. as the use and evaluation case, our simulated agents and human actors participate in (meta)cognitive skills training within negotiation based scenarios. this paper is structured as follows. section 2 discusses cognitive modelling, with a focus on human interactive multitasking, learning and adaptive behaviour. we briefly discuss the act-r architecture and provide details on an instance-based cognitive model that we used as a basis for designing an agent’s decisions-making processes and generation of task-related actions. section 3 describes an interactive learning scenario for the development of metacognitive skills in a multiissue bargaining setting. we provide an overview of existing approaches and systems for cognitive tutoring tasks, as well as dialogue systems used in negotiation domains. we specify tasks and actions performed by negotiators, negotiation structures procedures and negotiation strategies. the data collection scenario is outlined and the semantic annotations of the data are discussed. section 4 specifies a multi-agent dialogue manager architecture that makes use of a dynamic multidimensional context model and incorporates a cognitive task agent plus various interaction control agents trained on the annotated data. section 5 presents the virtual negotiation coach, outlining the system architecture and providing important details for key modules.1 section 6 reports on the system evaluation, where users’ subjective perception of effectiveness, efficiency and satisfaction were correlated with various objective performance metrics. evaluation results are also provided with respect to the number and quality of agreements reached, estimated level of cooperativeness, and acceptance of negative outcomes, as well as the subjective assessment of the skill training effects. section 7 summarises our findings and outlines future research. 2. cognitive modelling cognitive models have been used for decades to explain and model human intelligent behaviour, and have been successful in capturing a wide variety of phenomena across multiple domains such as decision making (marewski and link, 2014), memory (nijboer et al., 2016), problem solving (lee et al., 2015), task switching (altmann and gray, 2008), user models in tutoring applications (ritter et al., 2007), and neuroimaging data interpretation (borst and anderson, 2015). one of the most widely researched cognitive architecture is act-r, see anderson (2007), a theory and platform for building models of human cognition, which accounts for hundreds of empirical 1note that details on the integrated system are provided to enable the evaluation of the designed dialogue manager. since it is difficult to evaluate the dialogue manager as a separate module, its evaluation is performed as part of the userbased evaluation of the integrated dialogue system with negotiation, tutoring and interactive capabilities. the detailed description of the full dialogue system functionality is thus out of scope of this study. 38 towards integration of cognitive models in dialogue management results obtained in the field of experimental psychology. act-r proposes a hybrid architecture that combines a production system to capture the sequential, symbolic structure of cognition, with a sub-symbolic, statistical layer to capture the adaptive nature of cognition. since available cognitive models produce detailed simulations of human (multi-)task performance, they are also of interest for playing a role in a multi-agent setting. this application is exploited in this study. it is of chief importance that our artificial agents exhibit plausible human behaviour, notably a human-like way of learning and interacting. this means that such an agent makes decisions and takes actions that humans might also make and take, but also that the agent is influenced by its experiences and builds representations of the people it interacts with. thus, the agent should be able to (1) learn by collecting a variety of experiences, through instruction and feedback, and through monitoring and reasoning about its own behaviour and that of others; (2) adapt its interactive behaviour to a human dialogue partner’s knowledge, intentions, preferences and competences; and (3) process and perform several actions related to the interactive tasks and the roles it should play, e.g. as a partner or as a tutor. 2.1 models of human learning: reinforcement and instance-based learning human learning involves acquiring and modifying knowledge, skills, strategies, beliefs, attitudes, and behaviors. learning may involve synthesizing different types of information (schunk, 2012). learning is a relatively permanent change in behavior as a result of experience (gross, 2016). people learn from their successes and failures, from observing situations around them, and from imitating the behaviour of others (bandura, 2012). two widely used learning models are reinforcement learning (rl) and instance-based learning (ibl). reinforcement learning is a formal model of action selection where the utility of different actions is learned by attending to the reward structure of the environment. generally speaking, rl works in a trial-and-error fashion attempting various actions and recording the reward gained by those actions, see sutton and barto (1998). one of the limitations of rl as a model of human decision making becomes apparent in environments where goals change. this may happen, for example, due to changes in the environment or to newly obtained knowledge of the environment, e.g. you need to mail a letter, you searched online for the closest post office, but on your way to it you see a street mailbox, so you drop the letter in there. initial goal changes may occur due to the understanding and evaluation of partner behaviour. this often happens in negotiations where a negotiator may revise his initial offers and make concessions dependent on the interpretation of partner behaviour concerning these goals. rl models make decisions based solely on the learned state-action utilities. rewards are set a priori, are fixed and never revisited. if the goal changes, the utilities representing the reward structure from the initial goal become irrelevant at best, and subversive at worst (veksler et al., 2012). recently, serious efforts have been undertaken to solve this issue combining concurrent learning (co-learning) of the system policy training and the policy trained against simulated users. for instance, georgila et al. (2014) and xiao and georgila (2018) showed that in negotiation setting agents using multi-agent rl techniques are able to adapt to the human users, also in situations which were not observed during training. humans, by contrast, employ their knowledge of the environment and their interactive partners to make decisions for achieving new goals, e.g. acting from experience or by association. our memories are retrieved based on their recency and frequency of use (anderson and schooler, 1991) and strategies are adapted with increasing task experience (siegler and stern, 1998). 39 malchanau, petukhova and bunt human learning often occurs as a result of experience. decisions are made by finding a prior experience (an instance) that is similar to the current situation and/or most recently, frequently used under comparable conditions, see logan (1988); gonzalez and lebiere (2005). an instance consists of a representation of the current state of the world (what do i know, what do i know about others, what am i asked, what can i do, what has happened before), and an action to be taken in that situation (give information, run tests, examine something, reason about others, change attitude, etc.). information is encoded in an instance as a state-actions template specifying decisions about which activity to engage in and how to move from one activity to the other. initial templates can be designed (pre-programmed) by experts and/or modelled as the result of dialogue corpus analysis. an ibl agent can start an interaction with an (almost) empty template, request information from the partner and add it to the memory as the interaction proceeds. newly created (partially) filled instances are stored in a human-like memory that models forgetting, similarity and blending of experiences. the most active instance is retrieved. activation is based on history (e.g. frequency and recency) and on similarity (e.g. how similar the instance is, given the context), see section 4.2.1 for the specification of instances and activation functions for our interactive settings. an agent can be also trained by giving it a set of instances (learning-by-instruction), which it can refine and/or augment in actual interaction (learning-by-doing and learning-by-feedback). rl is a useful paradigm where the possible strategies are relatively clear. if the underlying interaction structure is very flexible, unclear or absent (i.e. hard to derive on the basis of the system’s behaviour), ibl based models have advantages, see also arslan et al. (2017). for instance, whenever a new goal is given, the ibl model will employ its stored knowledge (instances) to make informed goal-directed decisions. it does not need to learn the reward structure through trial-anderror; rather, the decision what action will be performed is based on the computed activation level, e.g. similarity between a past experience and the given current goal. moreover, feedback can be used in ibl to create an instance that contains the correct solution, i.e. the model will add an instance of another strategy, whereas the rl model will punish the strategies that lead to a wrong solution. strategy selection, which is implicit in rl, is explicit in the ibl model which makes it particularly suitable for tutoring applications. ibl is moreover robust to missing information due to the partial matching component in the act-r activation function, e.g. when the agent does not have access to the same information as his partner. we applied the instance-based learning approach to create flexible cognitive agents, also because it requires far less experience than machine learning methods that learn bottom-up, and the agent’s decision-taking behaviour incrementally improves as its set of instances increases in size. instance-based learning takes the middle ground between expert systems, in which knowledge typically lacks flexibility, and bottom-up machine learning, which requires extensive training data, and in which decisions are reached in an opaque manner. 2.2 adaptive interactive behaviour interactive systems and interfaces tailored towards specific users have been demonstrated to outperform traditional systems in usability. nass et al. (2005) present an in-car user study with a “virtual passenger”. experimental results indicate that subjective and objective criteria, such as driving quality, improve when the system adapts its voice characteristics to the driver’s emotion. nass and li (2000) confirm in the study of spoken dialogues in a book shop that similarity attraction is important for personality expression: matching the users’ degree of extroversion strongly influenced trust and attributed intelligence. 40 towards integration of cognitive models in dialogue management these observations have triggered the development of interactive systems that model and react to the users’ traits and states, for example by adapting the interaction based on language generation techniques (mairesse and walker, 2005). in gnjatovic and rösner (2008) a gaming interface is based on emotional states computed from the interaction history and actual user command. nasoz and lisetti (2007) describe a user modelling approach for an intelligent driving assistant, which derives the best system action in terms of driving safety, given estimated driver states. the above approaches adapt locally, i.e. the adaptation decision is made at turn level with very limited context and thus with no or very limited foresight. reinforcement learning has emerged as a promising approach for long-term considerations. while early studies (walker et al., 1998; singh et al., 2002) used rl to build strategies for simple systems, more complex paradigms are represented by statistical models, see frampton and lemon (2009). however, when users with different personalities in different states are systematically confronted with a learning system, most studies resort to user simulation: janarthanam and lemon (2009) simulate users of different levels of expertise, lópez-cózar et al. (2009) simulate users with different levels of cooperativeness, and georgila et al. (2010) simulate interactions of old and young users. these studies demonstrate that the simulation of different user types is expected to lead to strategies which adapt to each user type. however, adaptivity has been not achieved at the level of dynamically changing goals within one dialogue. rewards that are used in dialogue policy learning and optimizations are fixed a priori. human learning however does not only involve strengthening of existing knowledge, compilation of new rules, collection of episodic experiences to improve future decisions, etc., but often requires more explicit reasoning, assessing why a particular solution worked or not, and manipulating the task representation accordingly this process is called ’metacognition’. in this study, metacognition plays two major roles: (1) it guides and regulates system task behaviour; and (2) it improves a participant’s learning by triggering reasoning about one’s own and partner behaviour. metacognitive processes concern reasoning about other people’s intentions and knowledge. mastering metacognitive skills is important in language use (van rij et al., 2010) and in playing knowledge games (meijering et al., 2012). a more elaborate form of these reasoning skills is important in collaboration, negotiation and other social and interpersonal skills. people with welldeveloped metacognitive skills are more concerned that their interactions will go well, and are able to flexibly modify their actions during interaction in order to better adapt to the dynamics of the situation, typically by using other people’s behaviour as a guide to their own (ickes et al., 2006). they are also better able to accomplish their goals, which appears to be the result of their superior planning skills (jordan and roloff, 1997). metacognitive skills can be trained by humans and learned by a system. when learning, humans also observe their partners’ behaviour. in addition to using experiences to determine its own decisions, an interactive agent can use them to interpret and reason about the behaviour of others (i.e. humans). the ability to understand that other people have mental states, with desires, beliefs and intentions, which can be different from one’s own, is called theory of mind (tom; premack and woodruff, 1978). in our application, the tom methodology has been used to design agents that can infer, explain, predict and correct a partners’ negotiation behaviour and negotiation strategies. 2.3 multitasking in human-computer interaction a dialogue system has at least three core tasks: (1) to monitor user dialogue behaviour; (2) to understand user dialogue contributions; and (3) to react adequately. participation in a dialogue is thus 41 malchanau, petukhova and bunt a complex activity. participants do not only need to exchange certain information, instruct another participant, negotiate an agreement, discuss results or plan future actions, etc., but among other things dialogue participants also share information about the processing of each others messages, elicit feedback, manage the use of time, take turns, and monitor contact (allwood, 2000). they often use linguistic and nonverbal elements to address several interactive and task-related aspects at the same time. during interaction, a dialogue system is usually in the role of “speaker” (or “sender’)’ or in the role of “addressee” (also called “hearer” or ”recipient”). the system may also play the role of a side-participant who witnesses a dialogue without participating in it, see clark (1996). a dialogue system’s tasks depend also on the application domain in relation to the role(-s) it plays, e.g. as a full-fledged interactive partner with equal responsibilities as a human one, as an assistant, adviser or mediator, as a passive observer, as a tutor or coach, and so on. for our virtual negotiation coach application we identified the following key roles: • observer: system observes dialogue sessions between two or more humans and keeps track of human-human dialogue without actively participating in it; • experiencer: system actively plays the role of one of the interaction participants, i.e. sender and addressee; • mirror: system re-plays the user’s performance in a human-system dialogue in real time. the user observes his own performance and has the opportunity to terminate, re-enter and re-play the dialogue session from any point; • tutor or coach: system provides feedback from ongoing formative or summative assessment of the user performance in one or more tutoring sessions (mory, 2004). the system may play multiple roles simultaneously and/or interchangeably. in most existing approaches to dialogue management the dialogue manager (dm) is able to handle one particular dialogue task at a time. most human activities however are essentially multitasking. for example, driving a car consists of two main processes: one that keeps the car in the middle of the driveway by looking at the road ahead of the car while operating the steering wheel and the gas and brake pedals, and a second process that monitors the traffic environment (e.g., is there a car behind you). thus, human cognition can be conceptualized as a set of parallel cognitive modules (e.g. vision, declarative memory, working memory, procedural memory, manual control, vocal control, etc.). as long as multiple tasks do not need the same resources at the same time, these tasks can be carried out in parallel without interference. in the case of the driving example, if the driver is given an additional task, for example to operate a cell phone, he will abandon the monitoring task due to lack of resources. threaded cognition, as the theory of parallel execution of tasks, has been proposed to explain human multitasking behaviour: why and when certain tasks may be performed together with ease, and which combinations pose a difficulty, what types of multitasking are disruptive, and when are they most disruptive. threaded cognition models have been used in a wide spectrum of multitasking experiments (salvucci and taatgen 2008; 2010). this theory has been built on top of the act-r cognitive architecture. we designed a multi-threaded dialogue manager with integrated multitasking cognitive agent which, along with being an active dialogue participant with monitoring, understanding and reacting tasks, is capable of providing feedback on partner performance and which can reason about its own and a partner’s behaviour, and suggest alternative actions. 42 towards integration of cognitive models in dialogue management 3. interactive training of metacognitive skills 3.1 interactive learning and tutoring cognitive tutoring systems aim to support the development of metacognitive skills. examples of such systems are described in bunt and conati (2003); azevedo et al. (2002); gama (2004); aleven et al. (2006) and baker et al. (2006). these systems rely on artificial intelligence and cognitive science as a theoretical basis for analysing how people learn (roll. et a., 2007). research by chi et al. (2001) revealed that the interactivity of human tutoring drives its effectiveness. interactive learning is a modern pedagogical approach that has devolved out of the hyper-growth in the use of digital technology and virtual communication. interactive learning is a promising and powerful way to develop metacognitive skills. in this study, the interactivity of a tutoring system is achieved through the use of multimodal dialogue. while many intelligent tutoring dialogue systems have been developed in the past (litman and silliman, 2004; riedl and stern, 2006; core et al., 2014; moore et al., 2005; paiva et al., 2004), to the best of our knowledge no existing cognitive tutoring system makes use of natural spoken and multimodal dialogue. metacognitive skills are domain-independent and should be applicable in any learning domain and in a variety of different learning environments, but despite their transversal nature, metacognitive skills training can only be practiced within certain domains and activity types. some systems have been developed successfully for the domains of mathematics, physics, geometry, biology and computer programming (metatutor, azevedo et al., 2009; rus et al., 2009: harley et al., 2013). for negotiation, metacognition has been empirically proven to be important since it significantly improves decision-making processes (aquilar and galluccio, 2007). for many existing human-computer negotiation systems, interactions are typically modelled as a sequence of competitive offers where partners claim a bigger share for themselves. valuable work has been done on well-structured negotiations where a few parties interact with fixed interests and alternatives, see e.g. traum et al. (2008), georgila and traum (2011), guhe and lascarides (2014), efstathiou and lemon (2015). in many real-life negotiations, parties negotiate not over one but over multiple issues, see e.g. cadilhac et al. (2013), where they have interests in reaching agreements about several issues, and their preferences concerning these issues are not completely identical (raiffa et al., 2002a). negotiators may have partially competitive and partially cooperative goals, and may make trade-offs across issues in order for both sides to be satisfied with the outcome. parties can delay making a complete agreement on the first discussed issue, e.g. they postpone making an agreement or make a partial agreement, until an agreement is reached on the second one. they can revise their past offers, accept or decline any standing offer, make counteroffers, etc. we consider such complex strategic negotiations as multi-issue integrative bargaining dialogues, see petukhova et al. (2016 and 2017). we aim at modelling these interactions with the main goal to train metacognitive skills. comparable work has been performed on modelling socalled semi-cooperative multi-issue bargaining dialogues, see (lewis et al., 2017), who proposed an approach to end-to-end training of negotiation agents using a dataset of human-human negotiation dialogues, and applying reinforcement learning. their study presents a new form of planning ahead where possible complete dialogue continuations are simulated dialogue rollout. our approach also allows to compute the best alternative move at each negotiation stage and plan ahead the complete negotiation. we compute about 420 outcomes per scenario, for 9 scenarios in total, each featuring different participant preference profiles. additionally, for tutoring purposes the model provides an explanation for all alternative choices and how they lead to what outcomes. the two approaches 43 malchanau, petukhova and bunt differ with respect to the amount of data/resources used (our 50 vs 5808 dialogues); scenario complexity (4 issues, 16 values and 9 different preference profiles in our scenario vs 3 types of items and 6 objects in lewis et al., 2017); and modalities modelled (multimodal vs typed conversations). in our study, we explicitly model various negotiation strategies, while in lewis et al. (2017), evidence of such strategies is observed, e.g. compromising or deceiving, and are implicitly learned but not considered by design. 3.2 models of multi-issue bargaining three main types of negotiations can be distinguished: distributive, joint problem-solving and integrative2. distributive negotiation means that any gain of one party is made at the expense of the other and vice versa; any agreement divides a fixed pie of value between the parties, see e.g. walton and mckersie (1965). the goal of joint problem-solving negotiations is, by contrast, to work together on an equitable and reasonable solution: negotiators will listen more and discuss the situation longer before exploring options and finally proposing solutions. the relationship is important for joint problem solving, mostly in that it helps trust and working together on a solution, see beach and connolly (2005). in integrative bargaining, parties bargain over several goods and attributes, search for an integrative potential (interest-based bargaining or win-win bargaining, see fisher an ury, 1981). this increases the opportunities for cooperative strategies that rely on maximizing the total value of the negotiated agreement (enlarging the pie) in addition to maximizing one’s own value at the expense of the partner (dividing the pie). the different types of negotiation are manifest mainly in how parties create and claim values. negotiation starts with the anchoring phase, in which participants introduce negotiation issues and options. they also obtain and provide information about preferences, establishing jointly possible values contributing to the zone of possible agreement (zopa, sebenius, 2007). participants may bring up early (tentative) offers, typically in the form of suggestions, and refer to the least desirable events ‘create value’. the actual bargaining occurs in the ‘claim value’ phase, potentially leading to adaptation, adjustment or cancelling the originally established zopa actions. patterns of concessions, threats, warnings, and early tentative commitments are observed here. distributive negotiations are more ‘claiming values’, while joint problem-solving negotiations are more ‘value creating’ interactions, and integrative negotiations are a mix of ‘creating and claiming values’ negotiations (watkins, 2003a). in distributive negotiations the size of the zopa is mostly determined by the ‘bottom lines’ of the opposite parties, which are formed by their respective best alternatives to a negotiated agreement (batna), see fisher and ury (1981). in integrative bargaining the zopa is mainly determined by the number of possible pareto optimal outcomes. pareto optimality reflects a state of affairs when there is no alternative state that would make any partner better off without making anyone worse off. after establishing the zopa, negotiators may still cancel previously made agreements and negotiations might be terminated. negotiation outcome is the phase associated with the “walk-away” positions for each partner. finally, negotiators can move to the secure phase summing up and restating negotiated agreements or termination outcomes. at this stage, strong commitments are expressed and weak beliefs concerning previously made commitments and agreements are strengthened. participants take decisions to move on with another issue or re-start the discussion. figure 1 2a fourth type of negotiations is bad faith, where parties only pretend to negotiate, but actually have no intention to compromise. such negotiations often take place in a political context, see cox (1958) 44 towards integration of cognitive models in dialogue management depicts the general negotiation structure as described in watkins (2003) and sebenius (2007), and observed in our data described in the next section. n ew issu es/ c h an ge g am e bargaining final arrangement deadlock constraints & opportunities feedback & adaptation anchoring zone of possible agreement ’create value’ ‘claim value’ outcome termination secure figure 1: negotiation phases associated with negotiation structure, based on watkins (2003); sebenius (2007). the negotiation outcome depends on the setting, but also on the agenda and the strategy used by each partner (tinsley et al., 2002). the most common strategy of novice negotiators observed is issue-by-issue bargaining (see data collection below). parties may start with what they think are the ‘toughest’ issues, where they expect the most sharply conflicting preferences and goals, or they may start to discuss the ‘easiest’, most compatible options. sometimes, however, negotiators bring all their preferences on the table from the very beginning. this increases the chance to reach a pareto efficient outcome, since a participant can explore the negotiation space more effectively, being able to reason about each others’ goals, see e.g. stevens et al. (2016b). defensive behaviour, i.e. not revealing preferences, but also being misleading or deceptive, i.e. not revealing true preferences, results in missed opportunities for value creation, see e.g. watkins (2003); lax and sebenius (1992). it has also been observed that as a rule it is easier for a negotiator to bargain down, i.e. to start with his highest preference and if this is not accepted by the partner, go down and discuss sub-optimal options, than it is to bargain in, i.e. to reveal his minimum goal and go up, offering preferences that are not necessarily shared by the partner. all the aspects mentioned above may influence negotiators’ strategies. traum et al. (2008), who also consider a multi-issue bargaining setting, but viewed as a multi-party problem-solving task, define strategies as objectives rather than the orientations that lead to them. they distinguish seven different strategies: find issue, avoid, attack, negotiate, advocate, success and failure. other researchers define negotiation strategies closely related to conflict management styles, i.e. the overall approach for conducting a negotiation. five main strategies are observed: competing (adversarial), collaborating, compromising, avoiding (passive aggressive), and accommodating (submissive), see raiffa et al. (2002a); tinsley et al. (2002). as in integrative negotiation, where the negotiators strive to achieve a delicate balance between cooperation and competition (lax and sebenius, 1992), we define two basic negotiation strategies: cooperative and non-cooperative. cooperative negotiators share information about their preferences with their opponents, are engaged in problem-solving behaviours and attempt to find mutually beneficial agreements (de dreu et al., 2000). a cooperative negotiator prefers the options that have the highest collective value. if not enough information is available to make this determination, a cooperative negotiator will elicit this information from his opponent concerning. a cooperative negotiator will not engage in positional bargaining3 tactics, instead, he will attempt to find issues where a trade-off is possible. 3positional bargaining involves holding on to a fixed set of preferences regardless of the interests of others. 45 malchanau, petukhova and bunt o all outdoor smoking allowed o no smoking in public transportation o no smoking in public transportation and parks o no smoking in public transportation, parks and open air events scope o flyer and billboard campaign in shopping district o anti-smoking posters at all tobacco sales points o anti-smoking television advertisements o anti-smoking advertisements across all traditional mass media campaign o no change in tobacco taxes o 5% increase in tobacco taxes o 10% increase in tobacco taxes o 15% increase in tobacco taxes o 25% increase in tobacco taxes taxation o police fines for minors in possession of tobacco products o ban on tobacco vending machines o police fines for selling tobacco products to minors o identification required for all tobacco purchases o government issued tobacco card for tobacco purchases enforcement figure 2: preference card: example of values in four negotiated issues presented in colours: brighter orange colours indicated increasingly negative options and brighter blue colours increasingly positive options. when incorporated into the graphical interface, partners’ offers were visualized with red arrow (system) and green one (user). non-cooperative negotiators prefer to withhold their preferences in fear of weakening their power by sharing too much, or they may not reveal true preferences deceiving and misleading the partner. these negotiators focus on asserting their own preferred positions rather than exploring the space of possible agreements (fisher and ury, 1981). a negotiator agent using this strategy will rarely ask an opponent for preferences, and will often ignore a partner’s interests and requests for information. instead, a non-cooperative negotiator will find his own ideal offer, state it, and insist upon it in the hope of making the opponent concede. he will threaten to end the negotiation or will make very small concessions. the non-cooperative negotiator will accept an offer only if he can gain a lot from it. we also model a neutral (or cautious) strategy. neutral actions describe behaviours that are not indicative of either strategy above. to sum up, our approach is based on the cognitive negotiation model of integrative multi-issue bargaining, which incorporates potentially different beliefs and preferences of negotiation partners, learns to reason about these beliefs and preferences, and accounts for changes in participants’ goals and strategies. 3.3 collection and annotation of negotiation data for adequate modelling of human-like multi-issue bargaining behaviour, a systematic analysis of collected and semantically annotated human-human dialogue data was performed. the collected and analysed data also served for the ibl instance template definition as well as for the training agent’s negotiation behaviour, e.g. various classifiers were built using this data, see section 5. the specific setting considered in this study involved a real-life scenario about anti-smoking legislation in the city of athens passed in 2015-2016. after a new law was enacted, many cases of civil disobedience were reported. different stakeholders came together to (re-)negotiate and improve the legislation. the main negotiation partner was the department of public affairs of the city council who negotiated with representatives of small businesses, police, insurances, and others. 46 towards integration of cognitive models in dialogue management dialogue act relative frequency (in %) dialogue act relative frequency (in %) communicative function modality/ qualifier communicative function modality/ qualifier propositionalquestion 2.0 suggest 10.0 checkquestion 2.2 addresssuggest 1.4 setquestion 10.3 acceptsuggest 2.0 choicequestion 0.6 declinesuggest 1.7 inform −> 30.3 offer −> 16.7 . . . non-modalized 41.3 . . . conditional 28.3 . . . prefer 30.4 . . . tentative 35.0 . . . disprefer 3.1 . . . final 36.7 . . . acquiesce 3.0 addressoffer 0.6 . . . need 2.0 acceptoffer −> 5.8 . . . able 19.0 . . . tentative 47.6 . . . unable 1.2 . . . final 52.4 agreement 10.3 declineoffer tentative 2.0 disagreement 4.1 table 2: distribution of task-related dialogue acts in the analysed multi-issue bargaining dialogues. the anti-smoking regulations were concerned with four main issues: (1) smoke-free public areas (scope); (2) tobacco tax increase (taxation); (3) anti-smoking program promotion (campaign); and (4) enforcement policy and police involvement (enforcement), see figure 2. each of these issues involves four to five most important negotiation values with preferences representing negotiation positions, i.e. preference profiles. nine cases with different preference profiles were designed. the strength of preferences was communicated to the negotiators through colours. brighter orange colours indicated increasingly negative options; brighter blue colours increasingly positive options. in the data collection experiments, each participant received the background story and a preference profile. their task was to negotiate an agreement which assigns exactly one value to each issue, exchanging and eliciting offers concerning 〈issue;value〉 options. participants were randomly assigned their roles. they were not allowed to show their preference cards to each other. no further rules on the negotiation process, order of discussion of issues, or time constraints were imposed. they were allowed to withdraw or re-negotiate previously made agreements within a session, or terminate a negotiation. 16 subjects (young professionals aged between 19 and 25 years) participated in the experiments. the resulting data collection consists of 50 dialogues of a total duration of about 8 hours, comprising approximately 4.000 speaking turns (about 22.000 tokens). the recorded speech was transcribed, segmented and annotated with iso 24617-2 dialogue act information. the iso 24617-2 taxonomy (iso, 2012; see also bunt et al., 2010) distinguishes 9 dimensions, addressing information about a certain task; the processing of utterances by the speaker (auto-feedback) or by the addressee (allo-feedback); the management of difficulties in the speaker’s contributions (own-communication management) or that of the addressee (partner communication management); the speaker’s need for time to continue the dialogue (time management); the allocation of the speaker role (turn management); the structuring of the dialogue (dialogue structuring); and the management of social obligations (social obligations management). additionally, to capture the negotiation task structure, task management acts are introduced. these dialogue acts explicitly address the negotiation process and procedure. this includes utterances for coordinating 47 malchanau, petukhova and bunt negotiation move relative frequency (in %) offer 75.0 counteroffer 12.4 exchange 6.6 concession 1.2 bargainin 0.4 bargaindown 1.2 deal 2.4 terminate 0.8 table 3: negotiation moves and their relative frequencies in the annotated multi-issue bargaining corpus. the negotiators’ activities (e.g., “let’s go issue by issue”) or asking about the status of the process (e.g., “are we done with the agenda?”). task management acts are specific for a particular task and are often similar in form but different in meaning from discourse structuring acts, which address the management of the interaction, e.g. “to sum up ...”, “let’s move to a next round”. at the negotiation task level, human-computer negotiation dialogue is often modelled as a sequence of offers. the offers represent participants’ commitments to a certain negotiation outcome. in human negotiation, however, offers as binding commitments are rare and a larger variety of negotiation actions is observed, see raiffa et al. (2002b). participant actions are focused mainly on obtaining and providing preference information. a negotiator often states his preferences without expressing (strong) commitments to accept an offer that includes a positively evaluated option, or to reject an offer that includes a negatively evaluated option. to capture these variations, we distinguished five levels of commitment using the iso 24617-2 dialogue act taxonomy4 and its superset dit++5: (1) zero commitment for offer elicitations and preference information requests, e.g. by questions; (2) the lowest non-zero level of commitment for informing about preferences, abilities and necessities, e.g. in the form of modalized answers and informs; (3) an interest and consideration to offer a certain value, i.e. suggestions; (4) weak (tentative) or conditional commitment to offer a certain value; and (5) strong (final) commitment to offer a certain value, see petukhova et al., 2017. to model negotiation behaviour with respect to preferences, abilities, necessity and acquiescence, and to compute negotiation strategies as accurately as possible, we define several modal relations between the modality ‘holder’ (typically the speaker of the utterance) and the target which consists of the negotiation move (and its arguments), see lapina and petukhova (2017). additionally, to facilitate structuring the interaction and enable participants to interpret partner intentions, dynamically changing goals and strategies efficiently, we defined a set of qualifiers attached to offer acceptances or rejections and agreements, tentative or final. semantically, dialogue acts correspond to update operations on the information states of the dialogue participants. they have two main components: (1) the communicative function, that specifies how to update an information state, e.g. inform, question, and request, and (2) the semantic content, i.e. the objects, events, situations, relations, properties, etc. involved in the update, see bunt (2000), bunt (2014a). negotiations are commonly analysed in terms of certain actions, such as offers, counter-offers, and concessions, see watkins (2003), hindriks et al. (2007). we consid4for more information see bunt (2009); visit also http://dit.uvt.nl/\#iso_24617-2 5http://dit.uvt.nl/ 48 towards integration of cognitive models in dialogue management iso 24617-2 dimension relative frequency (in %) task 47.6 task management 10.3 autofeedback 18.7 allofeedback 2.3 turn management 6.6 time management 6.6 discourse structuring 4.6 own communication management 2.1 partner communication management na social obligation management 1.2 table 4: distribution of dialogue acts per iso 24617-2 dimension in the multi-issue bargaining corpus. ered two possible ways of using such actions, also referred to as ‘negotiation moves’, to compute the update semantics in negotiation dialogues. one is to treat negotiation moves as task-specific dialogue acts. due to its domain-independent character, the iso 24617-2 standard does not define any communicative functions that are specific for a particular kind of task or domain, but the standard invites the addition of such functions, and includes guidelines for how to do so. for example, a negotiation-specific kind of offern function could be introduced for the expression of commitments concerning a negotiation value.6 another possibility is to use negotiation moves as the semantic content of general-purpose dialogue acts. for example, a negotiator’s statements concerning his preference for a certain option can be represented as in f orm(a,b,3o f f er(x ;y )). we chose the latter possibility and specified 8 basic negotiation moves, whose distribution in the analysed data is shown in table 3. to sum up, the designed negotiation dialogue model accounts for several types of action performed by negotiators: (1) task-related dialogue acts expressing negotiation preferences and commitments; (2) qualified (‘modalized’) actions expressing participants’ negotiation strategies, see table 2; (3) negotiation moves specifying events and their arguments, see table 3; and (4) communicative actions to control the interaction, see table 4. a detailed specification of negotiation update semantics can be found in petukhova et al. (2017). semantic annotations were performed by three trained annotators who reached a good interannotator agreement in terms of cohen’s kappa of 0.71 on average, when performing segmentation and annotation simultaneously. in total, the corpus data contains more than 18.000 annotated entities. annotations were delivered in iso diaml format (iso 24617-2, 2012),.diaml files consisting of primary data in tei-compliant representation, with 24617-2 dialogue act annotations. the collected data and annotations is part of the metalogue multi-issue bargaining (mib) corpus (petukhova et al., 2016) which is released through ldc.7. 6negotiation ‘offers’ may have a more domain-specific name, e.g. bid for selling-buying bargaining. 7please visit https://catalog.ldc.upenn.edu/ldc2017s11 49 malchanau, petukhova and bunt 4. multi-agent dialogue manager: functional design and technical integration as act-r based computational cognitive models of threaded cognition and ibl can be used to design cognitive agents that simulate task-related behaviour showing close to human decision-making performance. if such agents have theory of mind (tom) skills they can exhibit metacognitive capabilities that are beneficial for better understanding and adequate modelling of adaptive and proactive task behaviour. they cannot yet deliver natural human-like interactive performance, but combining them with interactive agents based on advanced computational dialogue models opens new possibilities. inspired by the distinction that can be made between task control actions and dialogue control actions (bunt, 1994), we explored these possibilities by integrating a cognitive task agent into the isu-based dialogue manager as part of a dialogue system. in the dialogue system design community, involving both theorists and practitioners, a clean separation into two layers is observed. one layer deals with the task at hand, and the other with the communicative performance itself, see e.g. lemon et al. (2003). to design task managers (agents), detailed task analysis, originally proposed by annett et al. (1971), is often performed. the method, in which a task is described in terms of a hierarchy of operations and plans, has been used successfully to simulate human decision-making processes. in dialogue management, it has also been deployed in the form of hierarchical task decomposition and expectation agenda generation within the ravenclaw framework (bohus and rudnicky, 2003) and tested successfully in several systems. examples include the use of a tree-of-handlers in the agenda communicator (xu and rudnicky, 2000), of activity trees in witas (lemon et al., 2001), and of recipes in collagen (rich et al., 1998). however, models based on task hierarchies, agendas, recipes and trees are rather static and are difficult to apply for non-linear (multi-branching) or non-sequential interactions, like multi-issue barganing dialogues. a more flexible approach is the plan-based approach. for instance, in the trips system (allen et al., 2001) a task manager is implemented that relies on planning and plan recognition, and coordinates actions with a conversational manager. plan construction and inference are activities that can easily get very complex, however, and become computationally intractable. multi-agent architectures have been proposed for adaptive and flexible human-computer interaction, e.g. in the jaspis speech application (turunen et al., 2005), in the open agent architecture (martin et al., 1999), and in galaxy-ii (seneff et al., 1998). an isu-based approach to dialogue management has been used to handle multiple aspects (‘dimensions’) simultaneously, see keizer et al. (2011); petukhova (2011); malchanau et al. (2015), separating task control acts and various classes of dialogue control acts. the dialogue manager tracks updates in multiple dimensions of the participants’ information states, as the effect of processing incoming dialogue acts, and generates multiple task control acts and dialogue control acts in response. in order to capture the dynamics related to frequently changing participants’ interactive and strategic goals, we propose a flexible adaptive form of multidimensional dialogue management inspired by cognitive models of multitasking, learning and cognitive skills transfer. to this end, we designed a cognitive task agent and integrated it as part of an isu-based multidimensional dialogue manager (dm). the dm receives data in the form of the recognized dialogue acts, updates the information state, and generates output. 50 towards integration of cognitive models in dialogue management 4.1 information state: the multidimensional context model according to the isu approach, dialogue behaviour, when understood by a dialogue participant, evokes certain changes in the participants’ information state or ‘context model’. since we deal with several different interactive, task-related and tutoring aspects, an articulate context model should contain all the information considered relevant for interpreting such rich dialogue behaviour in order to enable the system to generate an adequate reaction playing the role of a negotiator or that of a tutor. an articulate dialogue model and context model have been proposed by bunt (1999). complexities of natural human dialogue are handled by analysing dialogue behaviour as having communicative functions in several dimensions, as discussed above.  functionalsegment(fs) :  start : 〈tokenindex|time point〉 end : 〈tokenindex|time point〉 verbatim : 〈tokenindex|time points = ‘token1′, . . .〉 prosody : 〈duration, pitch,energy, . . .〉 nonverbal : 〈  head〈elementindex|time points = expression1, . . .〉 hands〈elementindex|time points = expression1, . . .〉 f ace〈elementindex|time points = expression1, . . .〉 posture〈elementindex|time points = expression1, . . .〉 〉 sender : 〈participant〉 dial acts(das) : {〈  dimension(d) : 〈dim〉 comm f unction(cf) : 〈c f 〉 sem content(sc) : 〈content〉 sender/speaker : 〈participant〉 addressee(−s) : {〈participant〉} f unc dependency : [ antecedent : {〈da〉} ] f b dependency : [ antecedent : {〈fs〉} ] rhetorical relation : [ antecedent : {〈da〉} type : 〈elaborate| . . .〉 ]  〉}   figure 3: feature structure representation of a functional segment. adopted from petukhova, 2011. the proposed context model has five components: (1) linguistic context (lc) with information about (a) ’dialogue history’; (b) ’latest segment’ in the form of functional segment to which one or multiple dialogue acts are assigned (see fig. 3), and (c) ’dialogue future’ (or ’planned state’); (2) semantic context (semc) containing information about the task/domain; (3) cognitive context (cc) representing information about the current and expected participants’ processing states; (4) perceptual/physical context (pc) having information about the perceptible aspects of the communication process and the task/domain; (5) social context (socc) containing information about current speaker’s beliefs about his own and his partner’s social obligations and rights. each of these five components contains the representation of three parts: (1) the speaker’s beliefs about the task, about the processing of previous utterances, or about certain aspects of the interactive situation; (2) the addressee’s beliefs of the same kind, according to the speaker; and (3) the beliefs of the same kind which the speaker assumes to be shared (or ’grounded’) with the addressee. a context model for multi-party dialogues is more complex, containing representations of the speaker’s beliefs about contexts of more than one addressee and possibly also of other participants (e.g. of the audience in a debate). figure 4 shows the context model with its component structure. each of the model parts can be updated independently while other parts remain unaffected. for instance, the linguistic context is updated when dealing with linguistic/multimodal behavioural aspects and some interactive aspects, such as turn management; in the cognitive context participants’ processing states are modelled, as well as aspects related to time and own communication management (e.g. speech production errors). the semantic context contains representations of task-related 51 malchanau, petukhova and bunt  lingcontext :  speaker :  dialogue history : { 〈previous segments fig. 3 〉 } latest segment : [ fs : fig. 3 state = opening|body|closing ] dialogue f uture : plan : [ candidates : 〈list das〉 order : 〈ordered list das〉 ]  partner : 〈partner linguistic context〉 (according to speaker) shared : 〈shared linguistic context〉  semcontext :  speaker task model : 〈belie f s〉 partner task model : 〈belie f s〉 (according to speaker) shared task model : 〈mutual belie f s〉  cogcontext :  speaker own proc state :  proc problem : yes|no problem input : fs time need : negligible|small|substantial  partner proc state : 〈partner cognitive context〉(according to speaker) shared : 〈shared cognitive context〉  perccontext :  speaker : [ own presence : positive|negative own readiness : positive|negative ] partner : 〈partner perceptual context〉(according to speaker) shared : 〈shared perceptual context〉  soccontext :  speaker : [ interactive pressure : none|greet|apology|thanking| . . . reactive pressure : {〈dialogue acts〉} ] partner : 〈partner social context〉(according to speaker) shared : 〈shared social context〉   figure 4: feature structure representation of the context model. adopted from malchanau et al., 2015. actions, in our scenario a participant’s negotiation moves and their arguments, partners’ negotiation strategies, and the system’s tutoring goals and expectations on a trainee’s learning progress. 4.2 cognitive task agent the cognitive task agent (cta) operates on a structured dynamic semantic context as described above, identifies the partner’s task-related goals, and uses a strategy to compute its next negotiation move. it interprets and produces negotiation actions based on the estimation of partner’s preferences and goals. the agent adjusts its strategy according to the perceived level of the opponent’s cooperativeness. currently, the agent distinguishes three strategies: cooperative, non-cooperative and neutral. the agent starts neutrally, requesting the partner’s preferences. if the agent believes the opponent is behaving cooperatively, it will react with a cooperative negotiation move. for instance, it will reveal its preferences when asked for, it will accept the opponent’s offers, and propose concessions or cross-issues trade-offs. it will use modality triggers of liking and ability. if the agent experiences the opponent as non-cooperative, it will switch to non-cooperative mode. it will stick to its preferences and insist on acceptance by the opponent. it will repeatedly reject the opponent’s offers using modal expressions of inability, dislike and necessity. it will rarely make concessions. it will threaten to withdraw reached agreements and/or terminate negotiation. such meta-strategies for strategy adjustment are observed in human negotiation and coordination games, see kelley and stahelski (1970), smith et al. (1982). we explain in some detail how this is implemented. 4.2.1 instance design: creation, activation and retrieval the agent’s negotiation moves and their arguments are encoded as ‘instances’, represented as a set of slot-value pairs corresponding to the agent’s preference profile. information encoded in an instance concerns beliefs about agent’s and partner’s preferences (state of the negotiation and conditions), and agent’s and estimated partner’s goals (actions), see table 5. the agent assumes that the partner’s preferences are comparable to his, but values may differ. at the beginning of the 52 towards integration of cognitive models in dialogue management information type explanation source strategy the strategy associated with the instance negotiationmove, modality my-bid-value-me the number of points the agent’s bid is worth to the agent preference profile my-bid-value-opp the number of points that the agent believes its bid is worth to the user opp-bid-value-me the number of points the user’s bid is worth to the agent opp-bid-greater true if the user’s bid is at least as much as the agent’s current bid, false otherwise next-bid-value-me the number of points that the next best option is worth the next best option is defined as the option closest in value to the current one (not including those that are worth more than the current option.) overall-value the total value of all options that have been agreed upon so far. historythis is a measure of how the negotiation is going. if it is negative, negotiation is likely to result in an unacceptable outcome. my-move the move that the agent should take in this context. planned future table 5: structure of an instance in the cognitive task agent, adopted from stevens et al. (2016a). interaction, the agent may have no or weak assumptions about the partner’s preferences. as the interaction proceeds the agent builds up more knowledge about the partner’s negotiation options. the agent achieves this by taking the perspective of its partner and using its own knowledge to evaluate the partner’s strategy, i.e. apply tom skills. the agent’s memory holds three sets of preference values: the agent’s own preferences (zero tom), the agent’s beliefs about the user’s preferences (first-order tom), and the agent’s beliefs about the user’s beliefs about the agent’s preference values (second-order tom). when a negotiation move and its arguments are recognized, the information is passed to the cta. the agent constructs a retrieval instance and fills in as many slots as it can with the received details and the current context. subsequently, the cta updates its own representation of the negotiation state by retrieving the most active instance from its declarative memory. an instance i that is used most recently and most frequently gets the highest activation value, which is derived from the following equation, see bothell (2004): ai = ln( n ∑ j=1 t−d j )+logistic(0,s) where n is the number of times an instance i has been retrieved in the past; t represents the amount of time that has passed since the jth presentation or creation of the instance, and d is the rate of activation decay.8 the rightmost term of the equation represents noise added to the activation level, where s controls the noise in the activation levels and is typically set at about 0.25, consistent with the value used in lebiere et al. (2000). thus, the equation effectively describes both the effects of recency more recent memory traces are more likely to be retrieved, and frequency if a memory trace has been created or retrieved more often in the past it has a higher likelihood of being retrieved. an instance does not have to be a perfect match to a retrieval request to be activated. act-r can reduce its activation according to the following formula used to compute partial matching pi, see bothell (2004): pi = ∑ l pmli where mli indicates the similarity value between the relevant slot value in the retrieval request (l) and the corresponding slot instance i summed over all slot values in the retrieval request. p denotes the mismatch penalty and reflects the amount of weighting given to the matching, i.e. when p is 8in the act-r community, 0.5 has emerged as the default value for the parameter d over a large range of applications, anderson et al. (2004). 53 malchanau, petukhova and bunt higher, activation is more strongly affected by similarity. we set the constant p high at 5, consistent with the value used in lebiere et al. (2000).9 the agent will thus be able to retrieve past instances for reasoning even when a particular situation has not been encountered before. partial matching, combined with activation noise, allows for flexibility in the agent’s behaviour. the agent will not rigidly make the exact same moves every time. for example, suppose the cta retrieves the following instance: instance-a strategy cooperative the opponent’s strategy is cooperative my-bid-value-me 4 the agent’s current offer is worth 4 points to him opp-bid-value-me 1 the opponent’s offer is worth 1 point to the agent opp-bid-greater true the opponent’s offer is equal or greater than agent’s current bid next-bid-value-me 2 the next best option for the agent is worth 2 points opp-move concede opponent changed its offer to one that was less valuable to him my-move concede the agent repays the opponent by also selecting a less valuable option two pieces of information will be extracted from these instances: the strategy of the user (cooperative) and an estimate of the user’s preference for the options mentioned in the move (1,true). if there are other good options available, a cooperative negotiator will explore those options first before insisting on his current position, so from this behaviour the agent infers that it is dealing with a cooperative negotiator with positive preferences on at least two issues. now the agent uses its own context to choose an appropriate response to the user. depending on how the user has acted, and what the agent knows (guesses) about the user’s preferences, the agent chooses to respond cooperatively, i.e. to concede. 4.2.2 multitasking behaviour the cta can reason about the overall state of the negotiation task, and attempts to identify the best negotiation move for the next action. it computes: (1) the agent’s counter-move, and (2) feedback sharing the agent’s beliefs about the user’s preferences and the user’s negotiation strategy. the agent may propose a strategically better alternative move that the user could have taken and explain ‘why’. as the result, the system is able to play simultaneously or interchangeably the four roles specified in section 2.1: observer, negotiator, mirror and tutor. in the observer mode, the agent monitors and keeps track of all performed own and partner’s actions and logs them. the created log files are used to evaluate the participants’ performance and for system improvement (see section 6). as a mirror, the agent’s monitoring and interpretation results are immediately displayed to the user. these displays include a transcript of the agent’s and user’s utterances (as recognized by the system), the agent’s perceived cooperativeness level and the recognized partner’s preferences. the agent’s and partner’s most recent offers and estimated partner’s preferences are also flagged in the dynamically updated preference card (fig. 2). the latter can have certain tutoring effects as well, since it may activate a user’s monitoring, reflection and regulating strategies, but also trigger a user’s corrective actions in case of agent processing failures. as a negotiator, the agent takes into account the recognized partner negotiation strategy, the agent’s preferences, and the estimation of those of the partner, and computes the most appropriate next negotiation move. this leads to relevant updates in the semantic context that give rise to goals 9to disable partial matching p can be set at 0. 54 towards integration of cognitive models in dialogue management id utterance dm input dm information state dm output (wording) da d;cf sc=nm(i;v )/ cta state dm update11 da for id [dependence] modality /decision generation a1 what do you want for scope da1 p1=offer(1;?v) neutral/elicit wants(a,know(a, p1)) task;setquestion c1 i think it would be fine bel(c,wants(a,know(a, p1))) if we stop smoking p2=offer(1;b)/ bel(c,2p2) in public transportation da2 task;answer[da1] prefer cooperative wants(c,know(a,2p2)) a2.1 okay bel(a,wants(c,know(a,2p2)) a2.2 i would go for that point cooperative/ bel(a,¬2p2)) da3 p2=offer(1;b) agree(a,1b) wants(a,know(c,3p2)) task;agreement a2.3 i prefer that we do not cooperative/ increase taxes da4 p3=offer(2;a) offer(a,2a) wants(a,know(c,2p3)) task;inform c2.1 this is not possible da5 task; p3=offer(2;a) non-cooperative/ bel(c,wants(a,know(c,2p3))) for us disagreement[da4] unable reject(c,2a) wants(c,know(a,¬3p3)) c2.2 i would like 5% tax increase at least da6 task;inform a3 sorry i was distracted apology & could you repeat da7 bel(a,¬interpreted(a,c2.2)) autonegative c3 sure da8 autopositive[da7] bel(c, interpreted(c,a3)) i prefer 5% tax increase da9 task;inform p4=offer(2;b) non-cooperative/ bel(c,wants(a,know(c,¬3p3)) prefer offer(c,2b) wants(c,know(a,2p4)) a4.1 okay da10 autopositive[da9] p4=offer(2;b) bel(a, interpreted(a,c3)) autopositive a4.3 i will give you 5% if you p4=offer(2;b) bel(a,wants(c,know(a,2p4)) bel(a,3p4) agree to ban all tobacco p5=offer(4;b) neutral/ bel(a,2p5) vending machines da11 p6=offer( exchange(a, bel(a,3p6) task;offer (2;b)→ (4;b)) (2b ∧ 4b)) wants(a,know(c,3p6)) c4 i think i can live with that da12 task; p6=offer( cooperative/agree(c, bel(c,wants(a,know(c,3p6)) agreement[da11] (2;b)→ (4;b)) (2b→ 4b)) bel(c,3p6) table 6: example of a negotiation dialogue with processing and generation by the dialogue manager. (a = agent (business representative); c = human negotiator (city councilor); da = dialogue act; d = dimension; cf = communicative function; sc= semantic content; nm = negotiation move; i = issue; v=value; bel = believes; 3 = possible; 2 = preferable) to perform a certain dialogue act, e.g. tentative agreement. other contexts may be also updated in parallel and goals are created to perform, for example, turn-taking (linguistic context) and feedback (cognitive context) actions, see next section. the dialogue manager passes dialogue act list for generation, < da1 = turntake,da2 = positiveautofeedback,da3 = task;agreement >, where da1 is decided to be generated implicitly, da2 non-verbally by a smiling and nodding avatar and verbally by ‘okay’, and da3 is generated by the utterance ‘i can live with it’. as a tutor, the agent shares its beliefs about the current negotiation state and its planned continuation, e.g. may offer strategically better user negotiation moves leading to higher quality negotiation outcomes in terms of pareto efficiency. after each action, the agent is also able to provide an explanation why decisions are made to perform certain actions. at the end of each negotiation session summative feedback is generated in terms of estimated pareto optimality, degree of cooperativeness, and acceptance of negative outcomes. this type of feedback accumulates across multiple consecutive negotiation rounds. the execution of these shared and varied tasks is expected to have positive effects both on user and system performance, enabling activation and improvement of metacognitive processes. moreover, since these processes do not require additional resources (memory, processing and control), but are model-inherent belief creation and transfer processes and characteristics (instance slots), multiple tasks related to various roles can be executed by the dm in parallel without interference. 55 malchanau, petukhova and bunt 4.2.3 dialogue manager state update: example table 6 provides an example of a dialogue between an agent a playing the role of the business representative and a human negotiator c in the role of the city councilor. the cta starts neutrally. a elicits an offer from c on the first issue and does this in the form of a set question. the understanding that a certain dialogue act is performed leads to corresponding context model updates.10 if the partner reacts to the agent’s elicitation by sharing his preferences in c1, he is evaluated by the agent as being cooperative. the agent’s preferences are not identical but not fully conflicting either: it is possible for the agent to agree with the opponent’s preferences accepting his offer in a2.2, where a believes that the offer made in c1 is not the most preferred one but still acceptable/possible for a.11 the cta stays in the cooperative mode. if the negotiator’s preferences differ from the options proposed by the partner, he may refuse to accept the partner offer as in c2.1 and may offer another value which is more preferable for him, i.e. perform a counter-offer move ( c2.2 repeated in c3 after the agent signaled that his processing was unsuccessful. the cta interprets the partner’s strategy as being non-cooperative and switches his strategy to neutral, proposing to exchange offers (in a4.2) that still aim at the better deal for himself. if this is again rejected, the agent will apply the noncooperative strategy and insist on his previous proposal expressed in a2.2, otherwise he will either elicit an offer for the next issue or propose an offer himself. the agent computes the partner’s negotiation strategy using the linguistic modality expressed in the partner’s utterance and the type of the dialogue act performed. the collected data was used to train classifiers in the supervised setting to make such predictions, see section 5. to assess the minimal amount of data required to detect the partner’s negotiation strategy reliably, a series of learnability experiments was performed. to achieve an accuracy higher than 75%, about 1300 training instances are used. it was noticed the classifier performance further benefits from adding more training data. an accuracy of 83% was achieved on a training set comprising 3800 instances, so twice as many as in the first iteration and consuming almost the entire human-human mib corpus. the system showing this performance was evaluated, see section 6. follow-up experiments indicated that adding more data (e.g. evaluation and simulated data) further improves the classification performance, although not significantly, gaining 1% in accuracy when adding additional 1000 instances. 4.3 dialogue control task actions account for less than half of all actions in our negotiation data, see table 4. other frequently occurring acts are concerned with task management, discourse structuring, feedback and social obligations. along with moving towards a final set of agreements, negotiators need to take care how to optimally structure and manage the negotiation and the interaction. in multi-issue bargaining, negotiators have a variety of task management strategies. they may discuss issues sequentially or bargain simultaneously about multiple options, making trade-offs across issues. they may withdraw and re-negotiate previously reached agreements. all these decisions require explicit communicative actions. the task management acts are recognized and generated by the system, and are modelled as part of the system’s semantic context containing, along with the information about the speaker’s beliefs about the negotiation domain, information concerning task progress 10a detailed specification of dialogue act update semantics is provided in bunt (2014b) and petukhova (2011). 11 we provide here a simplified representation of the participants’ information states as tracked and updated by the dm. the full specification of participants’ information states and their updates can be found in petukhova et al. (2017). 56 towards integration of cognitive models in dialogue management processing level latest dialogue act previous dialogue act planned dialogue act communicative function negotiation move perception unknown unknown any request repeat and/or autonegative interpretation unknown offer(x) any accept(offer(x)) or reject(offer(x)) unknown offer(unknown) any question(offer(?)) and/or autonegative any unknown any request repeat or rephrase and/or autonegative unknown unknown accept(offer(x)) or reject(offer(x)) question(offer(?)) or inform(offer(y)) unknown unknown offer(x) question(offer(y)) question unknown accept(offer(x)) or reject(offer(x)) inform(offer(y)) table 7: decision-making support for the system’s feedback strategies concerning perception and interpretation of task-related actions, and expected dialogue continuation. note: x 6=y. and success. a task planner as part of the task manager (see fig. 3) takes care of updates and generation processes of this type. acts related to negotiators’ perception of the partner’s physical presence and readiness to start, continue or terminate the interaction as well as participants’ beliefs concerning the availability and properties of communicative and perceptual channels are modelled as part of the perceptual context. dialogue behaviour addressing these aspects is important, in particular, these actions are considered for generation, since the system’s multimodal behaviour related to contact management is embodied by a virtual character (full body avatar). the contact manager takes care of updates and the generation of these acts. a participant’s beliefs concerning the interaction structure (i.e. history, present and future states) and beliefs concerning topic shifts are modeled as a part of the linguistic context; the discourse structuring module takes care of the updates and generation specific for the interaction management and monitoring. 4.4 validity checking, repair and clarification strategies for an interactive system it is important to know that its contributions are understood and accepted by the user, as well as to signal the system’s processing of the same kind. conversation is a bilateral process that is, a joint activity, and speaking and listening are not autonomous processes conversational partners monitor their own processing of the exchanged utterances as well as the processing done by the others, see clark and krych (2004) for discussion. given the bilateral nature of conversation, interlocutors can construct and provide feedback on both their own processing (auto-feedback) as and on that by the other (allo-feedback). feedback is crucial for successful communication. feedback can be provided at different levels of processing the communicative behaviour of interlocutors. allwood et al. (1993) and clark (1996) notice that interlocutors need to establish contact and gain or pay attention to each others behaviour, in order be involved in conversation. a speaker’s behaviour needs to be perceived (i.e. heard, seen) or identified (clark, 1996). perceived behaviour should be interpreted, i.e. interlocutors should be able to extract the meaning of each other’s behaviour. the constructed interpretation needs to be evaluated against one’s information state: if it is consistent with the current information state it 57 malchanau, petukhova and bunt processing level latest dialogue act previous dialogue act validity planned dialogue act communicative function negotiation move evaluation inform terminate any valid stop negotiation accept offer(x) final offer(x) valid inform(deal(x)) any other than accept offer(x) final offer(x) invalid auto/allofeedback: question(?offer(x)) accept deal(x)) inform(deal(x)) valid discoursestructuring: topicshift taskmanagement:suggest(next issue) discoursestructuring: closing reject or accept offer(x) suggest(offer(x)) valid inform(offer(y)); question(offer(y)) reject or accept offer(x) inform(offer(x)) valid inform(offer(y)); question(offer(y)) question offer(x) inform(deal(x)) valid inform(deal(x)) reject offer(x) reject(offer(x)) valid reject(offer(x)) final offer or accept offer(x) accept(offer(x)) valid accept(offer(x)) inform suggest offer offer(x) accept(offer(y)) valid if x= ¬y accept(offer(?x)) or reject(offer(x)) inform offer(y) inform(deal(x)) invalid interpret as accept(deal(x)) if x=y otherwise reject(offer(x)) accept offer(x) any other than suggest or inform(offer(x)) invalid interpret as inform(offer(x)) or suggest(offer(x)) if x=y otherwise generate autoor allonegative reject offer(x) any other than suggest or inform(offer(x)) invalid interpret as inform(offer(y)) or suggest(offer(y)) if x¬y otherwise generate autoor allonegative accept offer(x) inform(terminate) invalid interpret as accept(terminate) and generate autoor allonegative reject offer(x) inform(terminate) invalid interpret as accept(terminate) and/or generate autoor allonegative inform deal(x) inform(terminate) invalid question(offer(?)) and/or autoor allonegative inform suggest offer offer(x) accept(offer(y)) invalid if x=y question(offer(?x)); accept(offer(y)) and/or autoor allonegative inform suggest offer offer(x) reject(offer(y)) invalid if x=y question(offer(?x)); reject(offer(y))and/or autoor allonegative inform deal(x) reject(offer(x)) invalid reject(offer(x)); question(offer(?x)) and/or autoor allonegative table 8: decision-making support for the system’s recovery and clarification strategies concerning evaluation of task-related actions, and expected dialogue continuation. in this table, valid stands for the state that can be recovered from the available information, otherwise invalid state that cannot be automatically recovered and requires activation of the clarification strategy. note: x 6=y. 58 towards integration of cognitive models in dialogue management processing level latest dialogue act previous dialogue act preferences validity planned dialogue act communicative function negotiation move execution inform suggest offer offer(x) any negative valid reject(offer(x)) and/or inform(offer(y)) and/or autonegative inform suggest offer offer(x) any positive valid accept(offer(x)) and/or inform(offer(y)) and/or autonegative inform suggest offer offer(x) any neutral valid accept(offer(x)) and/or inform(offer(y)) inform deal(x) no accept(offer(x)) any invalid reject(deal(x)); question(offer(?x)) and/or autonegative inform deal(x) reject(offer(x)) any invalid reject(deal(x)); question(offer(?x)) and/or autonegative inform terminate no final offer(x) any invalid question(offer(?x)) and/or autonegative table 9: decision-making support for the system’s feedback strategies concerning execution of task-related actions. in this table, valid stands for the state that can be recovered from the available information, otherwise invalid state that cannot be automatically recovered and requires activation of the clarification strategy. note: x 6=y can be incorporated into that state; if it is inconsistent, this can be reported as negative feedback. the incorporation of new information, and the performance of other mental and physical actions in response to communicative behaviour is called the execution or application (bunt, 2000). a speaker may provide feedback (feedback giving) or elicit feedback (feedback eliciting). as for positive feedback acts, explicitly signalled acceptances are generated, either verbally or non-verbally. we also consider generation of multimodal expressions of implied and entailed positive feedback (see bunt 2007; 2012) for strategic reasons, e.g. to provide more certainty due to potentially erroneous automatic speech recognition output. detected difficulties and inconsistencies in recognition, interpretation, evaluation and execution need to be resolved immediately if these problems are serious enough to impede further task performance; such problems are reported accordingly. problems due to deficient recognition and interpretation are frequent in spoken human-computer dialogue systems, but rarely observed in the collected human-human dialogue data. good news however is that humans generally exhibit certain re-occurring behavioural patterns when their processing fails. for our scenario and dialogue setting we incorporated observations and analyses of other available dialogue resources such as the human-human ami and hcrc maptask corpora (carletta, 2006; anderson et al., 1991), and human-human and human-computer dbox quiz game data (petukhova et al., 2014; 2015). observations from human-human and human-computer dialogues resulted in the definition of feedback strategies at the level of perception (recognition) and interpretation mostly comprising corrections and requests to repeat or rephrase (table 7), at the level of evaluation reporting inconsistencies/(in)validity due to certain logical constraints, given the grounded negotiation history (table 8), and at the level of execution reporting inability to accept an offer or to reach an agreement due to the negotiator’s preference profile (table 9). certain system processing flaws can be recovered from the information available to the system, some problems are too severe to continue the dialogue successfully and trigger feedback acts (clarification requests). in total, about 30 clarification and recovery strategies have been defined and evaluated (see also section 6). 59 malchanau, petukhova and bunt candidate dialogue acts dialogue acts for presentation update & generation processes/threads interpretation manager/ fusion fission/generation manager consistency evaluation/conflict resolution process manager 'internal clock'/timer 'information state'/contexts linguistic semantic cognitive perceptual social task manager auto-/allo feedback social obligations management turn management discourse structuring cognitive task agent (cta) task planner contact management figure 5: cognitive task agent (grey box) incorporated into the dialogue manager architecture: fused dialogue act information is passed to the dialogue manager from the interpretation manager for context model update and next action(-s) generation which are ‘fissed’ in different output modalities; both processes are regulated by the process manager. information concerning successes and failures in the processing of a partners’ dialogue contributions are modelled as part of the cognitive context (see fig. 3). thus, dialogue control acts present an important part for any interaction. in a shared cultural and linguistic context, choices concerning the frequency of such actions and the variety of expressions are rather limited. conventional forms are mostly used to greet each other, to apologize, to manage the turns and the use of time, to deal with speaking errors, and to provide or elicit feedback. models of dialogue control behaviour once designed can therefore be applied in a wide range of communicative situations. the use of task-related dialogue acts, by contrast, is more applicationspecific. the separation between task-related and dialogue control actions is therefore not only a cost-effective solution, but also allows designing flexible architectures and combinations of different modelling approaches and techniques, resulting in more robust and rich system behaviour. 4.5 dialogue manager architecture the above considerations have resulted in a dialogue manager consisting of multiple agents corresponding currently to six iso 24617-2 or dit++ dimensions12: the task manager with the integrated cta and task planner for task control, the auto/allo feedback agent, the turn manager, the discourse structuring manager, the contact manager, and the social obligations manager. the dialogue manager (dm) is designed as a set of processes (‘threads’) that receive data, update the information state, and generate output. additionally, consistency checking and conflict resolution is performed to avoid that the context model would be updated with inconsistent or 12the set of agents may in future be extended to include all nine iso 24617-2 dimensions and possibly other additional dimensions. 60 towards integration of cognitive models in dialogue management conflicting information and incompatible dialogue acts are generated, see also petukhova (2011). figure 5 presents the overall dm architecture. first, data are received from the fusion/interpretation module. next, the information state (‘context model’) is updated based on the received input. the process manager decides what parts of the context model to update. following receiving and updating, the output based on the analysis of the information state is generated. the output presents the ordered list of dialogue acts which is sent to the fission module, see next section for complete dialogue system architecture. 5. the virtual negotiation coach: design and evaluation as a proof of concept, and for assessing the potential value of the integration of a cognitive agent into a dialogue manager, we designed the virtual negotiation coach (vnc), an interactive system with the functionality described in the scenario for data collection (section 3.2). the vnc gets a speech signal, recognizes and interprets it, identifies relevant actions and generates multimodal actions, i.e. speech and gestures of a virtual negotiator and positive and negative visual feedback for tutoring. figure 6 shows the vnc architecture and processing workflows. speech signals are recorded from multiple sources, such as wearable microphones, headsets for each dialogue participant, and an all-around microphone placed between participants. the speech signals serve as input for two types of further processing: automatic speech recognition (asr), leading to lexical, syntactic, and semantic analysis, and prosodic analysis concerned with voice quality, fluency, stress and intonation of speech. the kaldi-based asr component incorporates acoustic and language models developed using various available data sources: the wall street journal wsj0 corpus13, hub4 news broadcast data14, the voxforge corpus15, the librispeech corpus16 and ami project data17. in total, about 759 hours of data has been used to train an acoustic model. the collected in-domain negotiation data is used as language model adaptation. the background language model is based on a combination of different corpora, like the approach taken to train the acoustic model. the asr performance is measured at 34.4% word error rate (wer), see singh et al. (2017)18. the asr outputs a single best word sequence without any scores. prosodic properties were computed automatically using praat (boersma and weenink, 2009) such as minimum, maximum, mean, and standard deviation of pitch, energy, voicing and speaking rate.19 the asr output is used by the negotiation moves and dialogue act classifiers. negotiation moves specify events and their arguments represented as negotiationmove(issue;value). conditional random field models for sequence learning (crf, lafferty et al. (2001)) are trained to predict three types of classes (move, issue and value) and their boundaries in asr n-best strings: 13https://catalog.ldc.upenn.edu/ldc93s6a 14https://catalog.ldc.upenn.edu/ldc98s71 15http://www.voxforge.org/ 16http://www.openslr.org/12/ 17http://groups.inf.ed.ac.uk/ami/corpus/ 18it should be noticed that the asr performance has been measured when interacting with non-native english speakers, who significantly varied in language skills level and speech fluency, some having a rather strong greek accent. 19we computed both raw and normalized versions of these features. speaker-normalized features were obtained by computing z-scores (z = (x-mean)/standard deviation) for the feature, where mean and standard deviation were calculated from all functional segments produced by the same speaker in the debate session. we also used normalizations by the first speaker turn and by prior speaker turn. 61 malchanau, petukhova and bunt speech avatar (tts) mics (tascam, headsets) wav voice quality & prosody analysis recognition layer input devices automatic speech recognition (asr) interpretation layer negotiation moves & modality classification dialogue acts & relations classification output processing layer fission multi-agent dialogue manager, fig 5 (context model, fig 4 ) core layer gesture avatar output layer visual feedback fusion figure 6: architecture of the virtual negotiation coach system. from bottom to top, signals are received through input devices, further recognized by tailored processing modules. after interpretation concerned with negotiation moves, modality and dialogue act classification, semantic representations from different modalities and modules are fused as dialogue acts. fused dialogue act information is passed to the dialogue manager for context model update and next action generation. the generated system response is rendered or ‘fissed’ in different output modalities. adopted with extensions and adjustments from van helvert et al., 2016 . negotiation move, issue, preference value. a ten-fold cross-validation using 5000 words of transcribed speech from the negotiation domain yielded an f-score of 0.7 on average. for the recognition of the intentions encoded in participants’ utterances various machine learning techniques have been applied, such as support vector machine (svm, boser et al., 1992), logistic regression (yu et al., 2011), adaboost (zhu et al., 2009), and the linear support vector classifier (vapnik, 2013). f-scores ranging between 0.83 and 0.86 were obtained, which corresponds to state-of-the-art performance, see amanova et al. (2016). the incremental tokenand chunk-based dialogue act crf-classifiers showed a performance of .80 f-scores on average, see ebhotemhen et al. (2017). after extensive testing, a non-incremental svm-based classifier has been integrated into the vnc system. the svm-based modality classifiers show accuracies in the range between 73.3 and 82.6% lapina and petukhova (2017). finally, information from the linguistic context related to the dialogue history has been used to ensure context-dependent interpretation of dialogue acts. additionally, the trainee has a choice to select options using a graphical interface as depicted in figure 2. as task progress support, partner offers and possible agreements are visualized with red (system) and green arrows (user). 62 towards integration of cognitive models in dialogue management the system’s fusion module currently fuses interpretations from two modules obtaining full semantic representations of user speech contributions. in the future, we will extend the system to other non-verbal modalities by integrating modern sensing technology at the input level. given the dialogue acts provided by the dialogue manager, the fission module generates responses splitting their parts into different modalities, such as avatar20 and voice (tts21) for negotiation actions, and visual feedback for tutoring actions. the latter includes a representation of the negotiators’ current cooperativeness, visualized by happy and sad face emoticons. at the end of each negotiation session, summative feedback is generated specifying the number of points gained or lost for each partner, the number of negative agreements, and the pareto optimality of the reached agreements. all messages exchanged between modules are in the standard tei and iso diaml formats. 6. evaluation it is generally not a trivial task to evaluate the performance of a dialogue manager as a single module due to its dependency on the quality of its potentially erroneous inputs. the performance of a dm is often evaluated as a part of the integrated dialogue system in a user-based fashion, by letting end users assess their interaction with the system. such assessment is typically based on the satisfaction of the users with the completion of the task. for example, paradise, one of the most widely-used evaluation models (walker et al., 1997), predicts user global satisfaction given a set of parameters related to task success and dialogue costs. satisfaction is calculated as the arithmetic mean of nine judgements on different quality aspects rated on 5-point likert scales. subsequently, the relation between task success and dialogue costs parameters and the mean human judgement is estimated carrying out a multivariate linear regression analysis. another way to evaluate a dialogue system is on the basis of interaction with computer agents that substitute human users and emulate user behaviour, see e.g. lópez-cózar et al. (2006). the various types of users and system factors can be systematically manipulated, e.g. interactive, dialogue task and error recovery strategies. several sets of parameters have been recommended for spoken dialogue system evaluation, ranging from a single bleu score metric for end-to-end system evaluation (wen et al., 2017), to seven parameters related to the entire dialogue (duration, response delay, number of turns) defined in fraser (1998) and 52 parameters in möller (2004) to meta-communication strategies (number of help requests, correction turns), to the system’s cooperativity (contextual appropriateness of system utterances), to the task which can be carried out with the help of the system (task success, solution quality), as well as to the speech input performance of the system (word error rate, understanding error rate). as for measuring satisfaction, various questionnaires have been proposed: nine satisfaction questions defined within paradise (walker et al., 2000); 44 evaluative statements of the subjective assessment of speech system interfaces (sassi) questionnaire (hone and graham, 2001); 53 evaluative statements in revu (report on the enjoyment, value, and usability, dzikovska et al., 2011); 24 bipolar adjective pairs defined in the godspeed questionnaire (bartneck et al., 2009); 122 evaluative statements in the questionnaire for user interface satisfaction (quis version 7.0, 20commercial software of charamel gmbh has been used, see reinecke (2003) 21vocalizer of nuance, http://www.nuance.com/for-business/text-to-speech/vocalizer/ index.htm, was integrated. 63 malchanau, petukhova and bunt evaluation criteria humanhumanhuman computer number of dialogues 25 (5808) 185 (na) mean dialogue duration (in turns per dialogue) 23 (6.6) 40 (na) agreements (%) 78 (80.1) 66 (57.2) pareto optimal (%) 61 (76.9) 60 (82.4) negative deal (%) 21 (na) 16 (na) cooperativeness rate (%) 39 (na) 51 (na) table 10: comparison of human-human and human-agent negotiation behaviour. adopted from petukhova et al. (2017). in brackets the best results reported by lewis et al. (2017) for comparison. na stands for not applicable, i.e. not measured. chin et al., 1988). the absence of standard performance metric sets and questionnaires for dialogue system evaluation makes it difficult to compare the results from different studies, and the various existing dialogue system evaluation results exhibit great differences. one of the common practices is to evaluate an interactive system or user interface by measuring usability, using well-defined observable and quantifiable metrics (see iso 9241-11 and iso/iec 9126-4 standards for usability metrics for effectiveness, efficiency and satisfaction). for this purpose, the usability perception questionnaire was constructed assessing eight main factors: task completion and quality, robustness, learnability, flexibility, likeability, ease of use and usefulness of the application,. the questionnaire has sufficient internal consistency reliability (cronbach’s alpha of 0.87) and comprises 32 evaluative statements22. using this questionnaire, we collected human judgements concerning the system performance in 28 evaluation sessions, with 28 participants aged 25-45, all professional politicians or governmental workers. nine negotiation scenarios were used, based on different negotiator preference profiles, see petukhova et al. (2016). participants were assigned a councilor role and a random scenario. the questionnaire allows human judgements to be linked to the performance of certain modules (or module combinations), see table 11. user judgements were presented in 5-point likert scales. the usability of the vnc system was measured in terms of effectiveness, efficiency and satisfaction. previous research suggests that there are differences in perceived and actual performance (nielsen, 2012): performance and perception scores are correlated, but they are different usability metrics and both need to be considered when conducting quantitative usability studies. in our design, subjective perception of effectiveness, efficiency and satisfaction were correlated with various performance metrics and interaction parameters to assess their impact on the qualitative usability properties. we computed bi-variate correlations to determine possible factors impacting user perception of the system usability and the performance metrics and interaction parameters derived from logged and annotated evaluation sessions. as performance metrics, system and user performance related to task completion rate23 and its quality24 were computed. we also compared system negotiation performance with human per22the usability questionnaire is available at https://docs.google.com/forms/d/e/ 1faipqlsf1h110uoflmgaqtt0hacbd7t0nihqlokoi1qs7o28wlzizsw/viewform 23 we consider the overall negotiation task as completed if parties agreed on all four issues or parties came to the conclusion that it is impossible to reach any agreement. 24 overall task quality was computed in terms of number of reward points the trainee gets at the end of each negotiation round and summing up over multiple repeated rounds; and pareto optimality (coefficient from 0 to 1) which reflects a state of affairs when there is no alternative state that would make any partner better off without making anyone worse off. 64 towards integration of cognitive models in dialogue management formance on the number of agreements reached, the ability to find pareto optimal outcomes, the degree of cooperativeness, and the number of negative outcomes25. it was found that participants reached a lower number of agreements when negotiating with the system than when negotiating with each other, 66% vs 78%. participants made a similar number of pareto optimal agreements (about 60%). human participants show a higher level of cooperativity when interacting with the system, i.e. 51% of the actions are perceived as cooperative. this may mean that humans were more competitive when interacting with each other. a lower number of negative deals was observed for human-agent pairs, 21% vs 16%. users perceived their interaction with the system as effective when they managed to complete their tasks successfully reaching pareto optimal agreements by performing cooperative actions but avoiding excessive concessions. our results differ from those reported in lewis et al. (2017) for both the human-human and the human-agent setting, see table 10. however, as noticed above, due to differences in tasks, scenario and interactive setting it is hard to draw clear comparative conclusions. nevertheless, we can conclude that the implemented cta is capable of making decisions and performing actions similar to those of humans. no significant differences in this respect were observed between human-human and human-system interactions. as for efficiency, we assessed temporal and duration dialogue parameters, e.g. time elapsed and number of system and/or user turns to complete the task (or a sub-task) and the interaction as a whole. we also measured the system response time, the silence duration after the user completed his utterance and before the system responded. weak negative correlation effects have been found between user perceived efficiency and system response delay, meaning users generally found the system reaction and the interaction pace too slow. dialogue quality is often assessed measuring word and sentence error rates (walker et al., 1997; lópez-cozár et al., 2006) and turn correction ratio (danielli and gerbino, 1995). many designers have noticed, however, that it is not so much how many errors the system makes that contributes to its quality, but rather the system’s ability to recognize errors and recover from them. this contributes to the perceived system robustness and is appreciated by users. users also value if they can easily identify and recover from their own mistakes. all system’s processing results were visualized to the user in a separate window, which contributes to the system observability. the repair and recovery strategies used by the system and the user were evaluated by two expert annotators, whose agreement was measured in terms of kappa. repairs were estimated as the number of corrected segments, recoveries as the number of regained utterances which were partially failed at recognition and understanding, see also danieli and gerbino (1995). while annotators agreed that repair strategies were applied adequately, longer dialogue sessions due to frequent clarifications are undesirable. the vnc is evaluated to be relatively easy to interact with (4.2 likert points). however, users found an instruction round with a human tutor prior to the interaction useful. most users were confident enough to interact with the system on their own, some of them however found the system too complex and experienced difficulties in understanding certain concepts/actions. a performance metric which was found to negatively correlate with system learnability is user response delay, the silence duration after the system completed its utterance and the user proposed a relevant dialogue continuation. nevertheless, the vast majority of users learned how to interact with the system and 25 we considered negative deals as flawed negotiation action, i.e. the sum of all reached agreements resulted in an overall negative value meaning that the trainee made too many concessions and selected mostly dispreferred bright ‘orange’ options (see figure 2). 65 malchanau, petukhova and bunt usability metric perception performance rassessment metric/parameter value effectiveness mean rating score effectiveness 4.08 task completion rate23; in % 66.0 .86* (task completeness) effectiveness (task quality) reward points24; mean, max.10 5.2 .19 user’s action error rate (uaer, in %)25 16.0 .27* pareto optimality24; mean, between 0 and 1 0.86 .28* cooperativeness rate; mean, in % 51.0 .39* efficiency (overall) mean rating score efficiency 4.28 system response delay (srd); mean, in ms 243 -.16 interaction pace; utterance/min 9.98 -.18 dialogue duration; in min 9:37 -.21 dialogue duration; average, in number of turns 56.2 -.35* efficiency (learnability) 3.3 (mean) user response delay (urd); mean, in ms 267 -.34* efficiency (robustness) 3.2 (mean) system recovery strategies (srs); correctly activated (cohen’s κ) 0.89 .48* user recovery strategies (urs); correctly recognized (cohen’s κ) 0.87 .45* efficiency (flexibility) 3.8 (mean) proportion spoken/on-screen actions; mean, in % per dialogue 4.3 .67* satisfaction (overall) aggregated per user ranging between 40 and 78 asr word error rate; wer, in % 22.5 -.29* negotiation moves recognition; accuracy, in % 65.3 .39* dialogue act recognition; accuracy, in % 87.8 .44* correct responses (cr)29; relative frequency, in % 57.6 .43* appropriate responses (ar)28; relative frequency, in % 42.4 .29* table 11: summary of evaluation metrics and obtained results in terms of correlations between subjective perceived system properties and actions, and objective performance metrics (r stands for pearson coefficient; * = statistically significant (p < .05) complete their tasks successfully in consecutive rounds. we observed a steady decline in user response delays from round to round.26 users appreciated the system’s flexibility. the system offered the option to select continuation task actions using a graphical interface on a tablet in case the system processing failed entirely. the use of concurrent multiple modalities was positively evaluated by the users. it was always possible for users to take initiative in starting, continuing and wrapping up the interaction, or leave these decisions to the system. at each point of interaction, both the user and the system were able to re-negotiate any previously made agreement.27 as overall satisfaction, the interaction was judged to be satisfying, rather reliable and useful, however, less natural (2.76 likert points). the latter is largely attributed to rather tedious multimodal generation and avatar performance. system actions were judged by expert annotators as appropriate28, correct29 and easy to interpret. other module-specific performance parameters reflect commonly used metrics derived using reference annotations such as various types of error rates, accuracy, and κ scores measuring agreement between the system performance and human annotations of the evaluation sessions. recognition and interpretation mistakes turned out to have moderate negative effects on user satisfaction. table 11 summarizes the results. 26for now, this is just a general observation; this metric will be taken into consideration in future test-retest experiments. 27performance metrics related to initiative and task substituitivity aspects and their impact on the perceived usability will be an issue for the future research. 28 a system action is appropriate given the context if it introduces or continues a repair strategy. 29 a system action is considered as correct if it addresses the user’s actions as intended and expected. these actions exclude recovery actions and error handling. 66 towards integration of cognitive models in dialogue management session recordings, system recognition and processing results, as well as the generated feedback were logged and converted to .anvil format in order to be able to use the anvil video analysis tool30 to view, browse, search, replay and edit negotiation sessions. anvil allows for automatic generation of some summative feedback about one or multiple sessions. moreover, applied prediction models can be evaluated by the negotiators and tutors on the fly, and edited and corrected annotated data can be used to retrain the system. with the satisfaction questionnaire we were also able to evaluate the system’s tutoring performance. participants indicated that system feedback was valuable and supportive. however, they expected more visual real-time feedback and more explicit summative feedback on their learning progress. most respondents think that the system presents an interesting form of skills training, and would use it as part of their training routine. 7. limitations and future work we have presented an approach to dialogue management that integrates a cognitive task agent able to reason about the goals and strategies of human partners, and to successfully engage in a negotiation task. this agent leverages established cognitive theories, namely act-r and instance-based learning, to generate plausible, flexible behaviour in this complex setting. we also argued that separate modelling of task related and dialogue control actions is beneficial for current and future dialogue system designs. the implementation introduced a theoretical novelty in instance-based learning for theory of mind skills and integrating this in the dialogue management of a tutoring system. the cognitive task agent used instance knowledge not only to determine its own actions, but also to interpret the human user’s actions, allowing it to adjust its behaviour to its mental image of the user. this work was successful: human participants who took part in evaluation experiments were not able to discern human users from simulated task agents (see also stevens et al. (2016b)), and an agent using theory of mind prompted users to use that themselves. our evaluation results suggest that the dialogue system with the integrated cognitive agent technology delivers plausible negotiation behaviour leading to reasonable user acceptance and satisfaction. the work presented here has certain limitations. instance templates in the instance-based learning model, slots, values and preferences for both partners were largely pre-programmed, which limits their general applicability. in the future, the agent will learn from real human-human dialogues, e.g. extract negotiation issues and values, and assess their importance. we will also enable the collaborative creation and real-time interactive correction, (re-)training and generation of agents by domain experts and target users. we aim to design authoring tools supporting agent learning and re-training across different situations. furthermore, we successfully integrated cognitive, interaction and learning models into a baseline proof-of-concept system. more research is needed on the connections between the cognitive models and the interaction and learning models, and overall mechanisms need to be further specified that underlie communication strategies depending on information about the current state of the task, participant (learning) goals, a participant’s affected state, and the interactive situation/environment. negotiation is more than the exchange of offers, decision making or problem solving; it involves a wide range of aspects related to feelings, emotions, social status, power, and interpersonal relations, context and situation awareness. for instance, tentative cooperative actions can engender a positive reaction and build trust over time, while social barriers can trigger interactive processes that often 30ww.anvil-software.de 67 malchanau, petukhova and bunt lead to bad communication, polarization and conflict escalation (sebenius, 2007). such dynamics may be observed in negotiations involving participants of different genders, races, or cultures (nouri et al., 2017). aspects related to social and interpersonal relations like dominance, power, politeness, emotions and attitudes deserve substantially more attention. finally, recent advances in digital technologies open new possibilities for us to interact with our environment, as well as for our environment to interact with us. everyday artefacts which previously were not aware of the environment at all are turning into smart devices and smart toys with sensing, tracking or alerting capabilities. this offers many new ways for real-time interaction with highly relevant, social and context-aware agents in multimodal multisensory environments which, in turn, enables designing rich immersive interactive experiences. an immersive and highly personalised coaching experience can be achieved by elaborate analysis and effective use of interaction data, applying advanced affective signal processing techniques and rich domain knowledge. a dialogue model that includes a comprehensive account of the user’s feelings, motivations, and engagement will form a foundation for a new generation of interactive tutoring systems. a direction that is not yet fully explored is to optimise in a system for the user’s feelings, motivation and engagement, as opposed to optimise for pure functional efficiency. 8. acknowledgments the research reported here was partly funded by the eu fp7 metalogue project, under grant agreement number 611073. references vincent aleven, bruce mclaren, ido roll, and kenneth koedinger. toward meta-cognitive tutoring: a model of help-seeking with a cognitive tutor. international journal of artificial intelligence in education, 16:101–128, 2006. james allen, george ferguson, and amanda stent. an architecture for more realistic conversational systems. in proceedings of the 6th international conference on intelligent user interfaces, pages 1–8. acm, 2001. jens allwood. an activity-based approach to pragmatics. abduction, belief and context in dialogue, pages 47–81, 2000. jens allwood, joakim nivre, and elisabeth ahlsén. on the semantics and pragmatics of linguistic feedback. journal of semantics, 9(1):1–26, 1992. erik altmann and wayne gray. an integrated model of cognitive control in task switching. psychological review, 115(3):602, 2008. dilafruz amanova, volha petukhova, and dietrich klakow. creating annotated dialogue resources: cross-domain dialogue act classification. in proceedings of the 9th international conference on language resources and evaluation (lrec 2016). elra, paris, 2016. anne anderson, miles bader, ellen bard, elizabeth boyle, gwyneth doherty, simon garrod, stephen isard, jacqueline kowtko, jan mcallister, jim miller, et al. the hcrc map task corpus. language and speech, 34(4):351–366, 1991. 68 towards integration of cognitive models in dialogue management john anderson. how can the human mind occur in the physical universe? new york, ny: oxford university press, 2007. john anderson and lael schooler. reflections of the environment in memory. psychological science, 2(6):396–408, 1991. john anderson, daniel bothell, michael byrne, scott douglass, christian lebiere, and yulin qin. an integrated theory of the mind. psychological review, 111(4):1036, 2004. john annett and neville stanton. task analysis. crc press, 2000. francesco aquilar and mauro galluccio. psychological processes in international negotiations: theoretical and practical perspectives. springer science & business media, 2007. burcu arslan, niels taatgen, and rineke verbrugge. five-year-olds systematic errors in secondorder false belief tasks are due to first-order theory of mind strategy selection: a computational modeling study. frontiers in psychology, 8, 2017. roger azevedo, amy witherspoon, amber chauncey, candice burkett, and ashley fike. metatutor: a metacognitive tool for enhancing self-regulated learning. in roberto pirrone, roger azevedo, and gautam biswas, editors, cognitive and metacognitive educational systems: papers from the aaai fall symposium (fs-09-02), 2002. roger azevedo, amy witherspoon, arthur graesser, danielle mcnamara, amber chauncey, emily siler, zhiqiang cai, vasile rus, and mihai lintean. metatutor: analyzing self-regulated learning in a tutoring system for biology. in aied, pages 635–637, 2009. ryan baker, albert corbett, kenneth koedinger, and ido roll. generalizing detection of gaming the system across a tutoring curriculum. in intelligent tutoring systems: 8th international conference, its 2006, jhongli, taiwan, june 26-30, 2006, volume 4053 of lecture notes in computer science, pages 402–411. springer, 2006. albert bandura. social cognitive theory. handbook of social psychological theories, 2012:349–373, 2011. christoph bartneck, dana kulić, elizabeth croft, and susana zoghbi. measurement instruments for the anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety of robots. international journal of social robotics, 1(1):71–81, 2009. lee beach and terry connolly. the psychology of decision making: people in organizations. sage, 2005. eric bilange. a task independent oral dialogue model. in proceedings of the fifth conference of the european chapter of the association for computational linguistics, pages 83–88, berlin, germany, 1991. association for computational linguistics. paul boersma and david weenink. praat: doing phonetics by computer. computer program. available at http://www.praat.org/, 2009. 69 malchanau, petukhova and bunt dan bohus and alexander rudnicky. ravenclaw: dialog management using hierarchical task decomposition and an expectation agenda. in proceedings of the 8th european conference on speech communication and technology (eurospeech 2003), geneva, switzerland, 2003. jelmer borst and john anderson. using the act-r cognitive architecture in combination with fmri data. in an introduction to model-based cognitive neuroscience, pages 339–352. springer, 2015. johan bos and tetsushi oka. an inference-based approach to dialogue system design. in proceedings of the 19th international conference on computational linguistics-volume 1, pages 1–7. association for computational linguistics, 2002. johan bos, ewan klein, oliver lemon, and tetsushi oka. dipper: description and formalisation of an information-state update dialogue system architecture. in proceedings of the 4th sigdial workshop on discourse and dialogue, pages 115–124, 2003. bernhard boser, isabelle guyon, and vladimir vapnik. a training algorithm for optimal margin classifiers. in proceedings of the fifth annual workshop on computational learning theory, pages 144–152. acm, 1992. dan bothell. act-r 6.0 reference manual, working draft, 2004. andrea bunt and cristina conati. probabilistic student modelling to improve exploratory behaviour. user modeling and user-adapted interaction, 13(3):269–309, 2003. harry bunt. context and dialogue control. think quarterly 3(1), pages 19–31, 1994. harry bunt. dynamic interpretation and dialogue theory. in m. taylor, f. neel, and bouwhuis d., editors, the structure of multimodal dialogue ii, pages 139–166. john benjamins, amsterdam, 1999. harry bunt. dialogue pragmatics and context specification. in h. bunt and w. black, editors, abduction, belief and context in dialogue: studies in computational pragmatics, pages 81–105. john benjamins, amsterdam, 2000. harry bunt. multifunctionality and multidimensional dialogue act annotation. in communication action meaning, a festschrift to jens allwood. e. ahlsèn et al. (ed.), pages 237–259. göteburg university press, 2007. harry bunt. the dit++ taxonomy for functional dialogue markup. in h. heylen, c. pelachaud, r. catizone, and d. traum, editors, proceedings of the aamas 2009 workshop ‘towards a standard markup language for embodied dialogue acts’ (edaml 2009), pages 13–25, budapest, 2009. harry bunt. the semantics of feedback. in proceedings of the 16th workshop on the semantics and pragmatics of dialogue (semdial 2012), pages 118–127, 2012. harry bunt. annotations that effectively contribute to semantic interpretation. in computing meaning, volume 4. springer, dordrecht, 2014a. 70 towards integration of cognitive models in dialogue management harry bunt. a context-change semantics for dialogue acts. in johan bos harry bunt and stephen pulman, editors, computing meaning, volume 4. springer, dordrecht, 2014b. anais cadilhac, nicholas asher, farah benamara, and alex lascarides. grounding strategic conversation: using negotiation dialogues to predict trades in a win-lose game. in proceedings of the conference on empirical methods in natural language processing (emnlp), pages 357–368, 2013. sandra carberry. plan recognition in natural language dialogue. acl-mit press series in natural language processing. bradford books, mit press, cambridge, massachusetts, 1990. jean carletta. announcing the ami meeting corpus. the elra newsletter, 11(1):3–5, 2006. michelene chi, stephanie siler, heisawn jeong, takashi yamauchi, and robert hausmann. learning from human tutoring. cognitive science, 25(4):471–533, 2001. john chin, virginia diehl, and kent norman. development of an instrument measuring user satisfaction of the human-computer interface. in proceedings of the sigchi conference on human factors in computing systems, pages 213–218. acm, 1988. herbert h clark. using language. cambridge university press, 1996. herbert h clark and meredyth a krych. speaking while monitoring addressees for understanding. journal of memory and language, 50(1):62–81, 2004. philip cohen and raymond perrault. elements of a plan-based theory of speech acts. cognitive science, 3(3):177–212, 1979. mark core, chad lane, and david traum. intelligent tutoring support for learners interacting with virtual humans. in r. sottilare, a. graesser, x. hu, and b. goldberg, editors, design recommendations for intelligent tutoring systems, volume 2, pages 249–257. u.s. army research laboratory, orlando, fl, usa, 2014. archibald cox. the duty to bargain in good faith. harvard law review, pages 1401–1442, 1958. niels dahlbaeck and arne jonsson. a coding manual for the linköping dialogue model. unpublished manuscript, 1998. morena danieli and elisabetta gerbino. metrics for evaluating dialogue strategies in a spoken language system. in proceedings of the 1995 aaai spring symposium on empirical methods in discourse interpretation and generation, volume 16, pages 34–39, 1995. carsten de dreu, laurie weingart, and seungwoo kwon. influence of social motives on integrative negotiation: a meta-analytic review and test of two theories., 2000. myroslava dzikovska, johanna moore, natalie steinhauser, and gwendolyn campbell. exploring user satisfaction in a tutorial dialogue system. in proceedings of the 12th annual meeting of the special interest group on discourse and dialogue (sigdial 2011), pages 162–172. association for computational linguistics, 2011. 71 malchanau, petukhova and bunt eustace ebhotemhen, volha petukhova, and dietrich klakow. incremental dialogue act recognition: tokenvs chunk-based classification. in proceedings of the 18th annual conference of the international speech communication association (interspeech), stockholm, sweden, 2017. ioannis efstathiou and oliver lemon. learning non-cooperative dialogue policies to beat opponent models:the good, the bad and the ugly. proceedings of the 19th workshop on the semantics and pragmatics of dialogue (semdial 2015 godial), page 33, 2015. roger fisher and william ury. getting to yes: negotiating agreement without giving in. harmondsworth, middlesex: penguin, 1981. matthew frampton and oliver lemon. recent research advances in reinforcement learning in spoken dialogue systems. the knowledge engineering review, 24(4):375–408, 2009. norman fraser. assessment of interactive systems. in handbook of standards and resources for spoken language systems, pages 564–615. mouton de gruyter, 1998. claudia gama. metacognition in interactive learning environments: the reflection assistant model. in j. c. lester, r. m. vicario, and f. paraguacu, editors, intelligent tutoring systems: 7th international conference, its 2004, maceió, alagoas, brazil, august 30 september 3, 2004, volume 3220 of lecture notes in computer science, pages 668–677. springer, 2004. kallirroi georgila and david traum. reinforcement learning of argumentation dialogue policies in negotiation. in twelfth annual conference of the international speech communication association, 2011. kallirroi georgila, maria wolters, and johanna moore. learning dialogue strategies from older and younger simulated users. in proceedings of the 11th annual meeting of the special interest group on discourse and dialogue, pages 103–106. association for computational linguistics, 2010. kallirroi georgila, claire nelson, and david traum. single-agent vs. multi-agent techniques for concurrent reinforcement learning of negotiation dialogue policies. in proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: long papers), volume 1, pages 500–510, 2014. milan gnjatovic and dietmar rösner. emotion adaptive dialogue management in human-machine interaction. citeseer, 2008. cleotilde gonzalez and christian lebiere. instance-based cognitive models of decision-making. in d. zizzo and a. courakis, editors, transfer of knowledge in economic decision making. macmillan, 2005. richard gross. psychology: the science of mind and behaviour. hodder education, 2016. markus guhe and alex lascarides. persuasion in complex games. proceedings of the 18th workshop on the semantics and pragmatics of dialogue (semdial 2014 dialwatt), page 62, 2014. jason harley, françois bouchet, and roger azevedo. aligning and comparing data on emotions experienced during learning with metatutor. in international conference on artificial intelligence in education, pages 61–70. springer, 2013. 72 towards integration of cognitive models in dialogue management james henderson, oliver lemon, and kallirroi georgila. hybrid reinforcement/supervised learning of dialogue policies from fixed data sets. computational linguistics, 34(4):487–511, 2008. matthew henderson, blaise thomson, and steve young. deep neural network approach for the dialog state tracking challenge. in proceedings of the 14th annual meeting of the special interest group on discourse and dialogue (sigdial 2013), pages 467–471, 2013. matthew henderson, blaise thomson, and steve young. robust dialog state tracking using delexicalised recurrent neural networks and unsupervised adaptation. in spoken language technology workshop (slt), 2014 ieee, pages 360–365. ieee, 2014. koen hindriks, catholijn jonker, and dmytro tykhonov. analysis of negotiation dynamics. in international workshop on cooperative information agents, pages 27–35. springer, 2007. kate s hone and robert graham. subjective assessment of speech-system interface usability. in seventh european conference on speech communication and technology, 2001. william ickes, renee holloway, linda stinson, and tiffany hoodenpyle. self-monitoring in social interaction: the centrality of self-affect. journal of personality, 74(3):659–684, 2006. srinivasan janarthanam and oliver lemon. a two-tier user simulation model for reinforcement learning of adaptive referring expression generation policies. in proceedings of the 10th annual meeting of the special interest group on discourse and dialogue (sigdial 2009), pages 120– 123. association for computational linguistics, 2009. jerry jordan and michael roloff. planning skills and negotiator goal accomplishment the relationship between self-monitoring and plan generation, plan enactment, and plan consequences. communication research, 24(1):31–63, 1997. simon keizer, harry bunt, and volha petukhova. multidimensional dialogue management. in a. van den bosch and g. bouma, editors, imix book. springer, 2011. harold kelley and anthony stahelski. social interaction basis of cooperators’ and competitors’ beliefs about others. journal of personality and social psychology, 16(1):66, 1970. seokhwan kim, luis fernando d’haro, rafael banchs, jason williams, and matthew henderson. dialog state tracking challenge 4, 2015. john lafferty, andrew mccallum, fernando pereira, et al. conditional random fields: probabilistic models for segmenting and labeling sequence data. in proceedings of the 18th international conference on machine learning (icml 2001), pages 282–289, san francisco, ca, usa, 2001. morgan kaufmann publishers inc. valeria lapina and volha petukhova. classification of modal meaning in negotiation dialogues. in proceedings of the 13th joint acl-iso workshop on interoperable semantic annotation (isa13), pages 59–70, montpellier, france, 2017. stafan larsson and david traum. information state and dialogue management in the trindi dialogue move engine toolkit. natural language engineering, 6(3-4):323–340, 2000. 73 malchanau, petukhova and bunt stanislao lauria, guido bugmann, theocharis kyriacou, johan bos, and a klein. training personal robots using natural language instruction. ieee intelligent systems, 16(5):38–45, 2001. david lax and james sebenius. the manager as negotiator: the negotiators dilemma: creating and claiming value. dispute resolution, 2:49–62, 1992. christian lebiere, dieter wallach, and rl west. a memory-based account of the prisoners dilemma and other 2x2 games. in proceedings of international conference on cognitive modeling, pages 185–193, 2000. hee seung lee, shawn betts, and john anderson. learning problem-solving rules as search through a hypothesis space. cognitive science, 2015. oliver lemon, anne bracy, alexander gruenstein, and stanley peters. the witas multi-modal dialogue system i. in proceedings of the 7th european conference on speech communication and technology (eurospeech 2001), 2001. oliver lemon, lawrence cavedon, and barbara kelly. managing dialogue interaction: a multilayered approach. in proceedings of the 4th sigdial workshop on discourse and dialogue, 2003. mike lewis, denis yarats, yann n dauphin, devi parikh, and dhruv batra. deal or no deal? end-to-end learning for negotiation dialogues. arxiv preprint arxiv:1706.05125, 2017. mei yii lim, joão dias, ruth aylett, and ana paiva. creating adaptive affective autonomous npcs. autonomous agents and multi-agent systems, 24(2):287–311, 2012. diane litman and scott silliman. itspoke: an intelligent tutoring spoken dialogue system. in demonstration papers at hlt-naacl 2004, pages 5–8. association for computational linguistics, 2004. gordon logan. toward an instance theory of automatization. psychological review, 95(4):492, 1988. ramón lópez-cózar, zoraida callejas, and michael mctear. testing the performance of spoken dialogue systems by means of an artificially simulated user. artificial intelligence review, 26(4): 291–323, 2006. ramon lópez-cózar, gonzalo espejo, zoraida callejas, ana gutiérrez, and david griol. assessment of spoken dialogue systems by simulating different levels of user cooperativeness. methods, 1(8):9, 2009. françois mairesse and marilyn walker. learning to personalize spoken generation for dialogue systems. in interspeech, pages 1881–1884, 2005. andrei malchanau, petukhova volha, bunt harry, and klakow dietrich. multidimensional dialogue management for tutoring systems. in proceedings of the 7th language and technology conference (ltc 2015), poznan, poland, 2015. julian marewski and daniela link. strategy selection: an introduction to the modeling challenge. wiley interdisciplinary reviews: cognitive science, 5(1):39–59, 2014. 74 towards integration of cognitive models in dialogue management david martin, adam cheyer, and douglas moran. the open agent architecture: a framework for building distributed software systems. applied artificial intelligence, 13(1-2):91–128, 1999. ben meijering, hedderik van rijn, niels taatgen, and rineke verbrugge. what eye movements can tell about theory of mind in a strategic game. plos one, 7(9):e45961, 2012. sebastian möller. quality of telephone-based spoken dialogue systems. springer science & business media, 2004. david moore, yufang cheng, paul mcgrath, and norman powell. collaborative virtual environment technology for people with autism. journal of the hammill institute on disabilities, 20(4): 231243, 2005. fabrizio morbini, david devault, kenji sagae, jillian gerten, angela nazarian, and david traum. flores: a forward looking, reward seeking, dialogue manager. in natural interaction with robots, knowbots and smartphones, pages 313–325. springer, 2014. edna mory. feedback research revisited. handbook of research on educational communications and technology, 2:745–783, 2004. fatma nasoz and christine lisetti. affective user modeling for adaptive intelligent user interfaces. in human-computer interaction. hci intelligent multimodal interaction environments, pages 421–430. springer, 2007. clifford nass and kwan min lee. does computer-generated speech manifest personality? an experimental test of similarity-attraction. in proceedings of the sigchi conference on human factors in computing systems, pages 329–336. acm, 2000. clifford nass, ing-marie jonsson, helen harris, ben reaves, jack endo, scott brave, and leila takayama. improving automotive safety by pairing driver emotion and car voice emotion. in chi’05 extended abstracts on human factors in computing systems, pages 1973–1976. acm, 2005. jakob nielsen. user satisfaction vs. performance metrics. nielsen norman group, 2012. menno nijboer, jelmer borst, hedderik van rijn, and niels taatgen. contrasting single and multicomponent working-memory systems in dual tasking. cognitive psychology, 86:1–26, 2016. elnaz nouri, kallirroi georgila, and david traum. culture-specific models of negotiation for virtual characters: multi-attribute decision-making based on culture-specific values. ai & society, 32(1): 51–63, 2017. ana paiva, joao dias, daniel sobral, ruth aylett, polly sobreperez, sarah woods, carsten zoll, and lynne hall. caring for agents and agents that care: building empathic relations with synthetic agents. in proceedings of the third international joint conference on autonomous agents and multiagent systems-volume 1, pages 194–201. ieee computer society, 2004. julia peltason and britta wrede. the curious robot as a case-study for comparing dialog systems. ai magazine, 32(4):85–99, 2011. 75 malchanau, petukhova and bunt volha petukhova. multidimensional dialogue modelling. phd dissertation. tilburg university, the netherlands, 2011. volha petukhova, martin gropp, dietrich klakow, anna schmidt, gregor eigner, mario topf, stefan srb, petr motlicek, blaise potard, john dines, et al. the dbox corpus collection of spoken human-human and human-machine dialogues. in proceedings of the 9th international conference on language resources and evaluation (lrec 2014). european language resources association (elra), 2014. volha petukhova, harry bunt, andrei malchanau, and ramkumar aruchamy. experimenting with grounding strategies in dialogue. in proceedings of the godial 2015 workshop on the semantics and pragmatics of dialogue, goteborg, sweden, 2015. volha petukhova, christopher stevens, harmen de weerd, niels taatgen, fokie cnossen, and andrei malchanau. modelling multi-issue bargaining dialogues: data collection, annotation design and corpus. in proceedings 9th international conference on language resources and evaluation (lrec 2016). elra, paris, 2016. volha petukhova, harry bunt, and andrei malchanau. computing negotiation update semantics in multi-issue bargaining dialogues. in proceedings of the semdial 2017 (saardial) workshop on the semantics and pragmatics of dialogue, saarbrücken, germany, 2017. david premack and guy woodruff. does the chimpanzee have a theory of mind? behavioral and brain sciences, 1(04):515–526, 1978. howard raiffa, john richardson, and david metcalfe. negotiation analysis: the science and art of collaborative decision making. harvard university press, 2002a. howard raiffa, john richardson, and david metcalfe. negotiation analysis: the science and art of collaborative decision making. harvard university press, 2002b. alexander reinecke. designing commercial applications with life-like characters. lecture notes in computer science, pages 181–181, 2003. charles rich and candace sidner. collagen: a collaboration manager for software interface agents. user modeling and user-adapted interaction, 8:3:149–184, 1998. mark riedl and andrew stern. believable agents and intelligent story adaptation for interactive storytelling. in international conference on technologies for interactive digital storytelling and entertainment, pages 1–12. springer, 2006. steven ritter, john anderson, kenneth koedinger, and albert corbett. cognitive tutor: applied research in mathematics education. psychonomic bulletin & review, 14(2):249–255, 2007. ido roll, vincent aleven, bruce mclaren, and kenneth koedinger. can help seeking be tutored? searching for the secret sauce of metacognitive tutoring. in r. luckin, kr koedinger, and j. greer, editors, artificial intelligence in education: building technology rich learning contexts that work, volume 158 of frontiers in artificial intelligence and applications, pages 203–210. ios press, 2007. 76 towards integration of cognitive models in dialogue management vasile rus, mihai lintean, and roger azevedo. automatic detection of student mental models during prior knowledge activation in metatutor. international working group on educational data mining, 2009. david sadek. dialogue acts are rational plans. in proceedings of the esca/etrw workshop on the structure of multimodal dialogue, pages 19–48, maratea, italy, 1991. dario salvucci and niels taatgen. threaded cognition: an integrated theory of concurrent multitasking. psychological review, 115(1):101, 2008. dario salvucci and niels taatgen. the multitasking mind. oxford university press, 2010. dale h schunk. learning theories an educational perspective sixth edition. pearson, 2012. james sebenius. negotiation analysis: between decisions and games. advances in decision analysis: from foundations to applications, page 469, 2007. stephanie seneff, ed hurley, raymond lau, christine pao, philipp schmid, and victor zue. galaxy-ii: a reference architecture for conversational system development. in proceedings of the 5th international conference on spoken language processing, 1998. iulian serban, alessandro sordoni, yoshua bengio, aaron courville, and joelle pineau. building end-to-end dialogue systems using generative hierarchical neural network models. in aaai, volume 16, pages 3776–3784, 2016. robert siegler and elsbeth stern. conscious and unconscious strategy discoveries: a microgenetic analysis. journal of experimental psychology: general, 127(4):377, 1998. mittul singh, youssef oualil, and dietrich klakow. approximated and domain-adapted lstm language models for first-pass decoding in speech recognition. in proceedings of the 18th annual conference of the international speech communication association (interspeech), stockholm, sweden, 2017. satinder singh, diane litman, michael kearns, and marilyn walker. optimizing dialogue management with reinforcement learning: experiments with the njfun system. journal of artificial intelligence research, 16:105–133, 2002. leasel smith, dean pruitt, and peter carnevale. matching and mismatching: the effect of own limit, other’s toughness, and time pressure on concession rate in negotiation. journal of personality and social psychology, 42(5):876, 1982. christopher stevens, harmen de weerd, fokie cnossen, and niels taatgen. a metacognitive agent for training negotiation skills. in proceedings of the 14th international conference on cognitive modeling (iccm 2016), 2016a. christopher stevens, niels taatgen, and fokie cnossen. instance-based models of metacognition in the prisoner’s dilemma. topics in cognitive science, 8(1):322–334, 2016b. richard sutton and andrew barto. reinforcement learning: an introduction, volume 1(1). mit press cambridge, 1998. 77 malchanau, petukhova and bunt catherine tinsley, kathleen o’connor, and brandon sullivan. tough guys finish last: the perils of a distributive reputation. organizational behavior and human decision processes, 88(2): 621–642, 2002. david traum, johan bos, robin cooper, staffan larsson, ian lewin, colin matheson, and massimo poesio. a model of dialogue moves and information state revision. trindi project deliverable d2.1, 1999. david traum, stacy marsella, jonathan gratch, jina lee, and arno hartholt. multi-party, multiissue, multi-strategy negotiation for multi-modal virtual agents. in international workshop on intelligent virtual agents, pages 117–130. springer, 2008. markku turunen, jaakko hakulinen, k-j raiha, e-p salonen, anssi kainulainen, and perttu prusi. an architecture and applications for speech-based accessibility systems. ibm systems journal, 44(3):485–504, 2005. joy van helvert, volha petukhova, christopher stevens, harmen de weerd, dirk börner, peter van rosmalen, jan alexandersson, and niels taatgen. observing, coaching and reflecting: metalogue a multi-modal tutoring system with metacognitive abilities. eai endorsed transactions on future intelligent educational environments, 16(6), 2016. jacolien van rij, hedderik van rijn, and petra hendriks. cognitive architectures and language acquisition: a case study in pronoun comprehension. journal of child language, 37(3):731– 766, 2010. kurt vanlehn. the behavior of tutoring systems. international journal of artificial intelligence in education, 16(3):227–265, 2006. vladimir vapnik. the nature of statistical learning theory. springer science & business media, 2013. vladislav veksler, christopher myers, and kevin gluck. an integrated model of associative and reinforcement learning. technical report, air force research lab wright-patterson afb oh, 2012. marilyn walker, diane litman, candace kamm, and alicia abella. paradise: a framework for evaluating spoken dialogue agents. in proceedings of the 8th conference on european chapter of the association for computational linguistics, pages 271–280. association for computational linguistics, 1997. marilyn walker, jeanne fromer, and shrikanth narayanan. learning optimal dialogue strategies: a case study of a spoken dialogue agent for email. in proceedings of the 36th annual meeting of the association for computational linguistics and 17th international conference on computational linguistics-volume 2, pages 1345–1351. association for computational linguistics, 1998. marilyn walker, candace kamm, and diane litman. towards developing general models of usability with paradise. natural language engineering, 6(3-4):363–377, 2000. richard walton and robert mckersie. a behavioral theory of labor negotiations: an analysis of a social interaction system. cornell university press, 1965. 78 towards integration of cognitive models in dialogue management michael watkins. analysing complex negotiations. harvard business review, december, 2003. tsung-hsien wen, david vandyke, nikola mrksic, milica gasic, lina rojas-barahona, pei-hao su, stefan ultes, and steve young. a network-based end-to-end trainable task-oriented dialogue system. in proceedings of the 15th european chapter of the association for computational linguistics (eacl 2017), valencia, spain, 2017. jason williams and steve young. partially observable markov decision processes for spoken dialog systems. computer speech & language, 21(2):393–422, 2007. jason williams, antoine raux, deepak ramachandran, and alan black. the dialog state tracking challenge. in proceedings of the 10th annual meeting of the special interest group on discourse and dialogue (sigdial 2013), pages 404–413, 2013. gang xiao and kallirroi georgila. a comparison of reinforcement learning methodologies in twoparty and three-party negotiation dialogue. in proceedings of the 31st international florida articial intelligence research society conference, pages 217–220, 2018. wei xu and alexander rudnicky. task-based dialog management using an agenda. in proceedings of the anlp-naacl 2000 workshop on conversational systems, pages 42–47, 2000. steve young. probabilistic methods in spoken–dialogue systems. philosophical transactions of the royal society of london a: mathematical, physical and engineering sciences, 358(1769): 1389–1402, 2000. steve young, milica gašić, blaise thomson, and jason d williams. pomdp-based statistical spoken dialog systems: a review. proceedings of the ieee, 101(5):1160–1179, 2013. hsiang-fu yu, fang-lan huang, and chih-jen lin. dual coordinate descent methods for logistic regression and maximum entropy models. machine learning, 85(1-2):41–75, 2011. ji zhu, hui zou, saharon rosset, trevor hastie, et al. multi-class adaboost. statistics and its interface, 2(3):349–360, 2009. 79 dialogue and discourse 3(2) (2012) 147-175 doi: 0.5087/dad.2012.207 question generation based on lexico-syntactic patterns learned from the web sérgio curto sergio.curto@l2f.inesc-id.pt spoken language systems laboratory l2f/inesc-id r. alves redol, 9 2o– 1000-029 lisboa, portugal ana cristina mendes ana.mendes@l2f.inesc-id.pt spoken language systems laboratory l2f/inesc-id instituto superior técnico, technical university of lisbon r. alves redol, 9 2o– 1000-029 lisboa, portugal luı́sa coheur luisa.coheur@l2f.inesc-id.pt spoken language systems laboratory l2f/inesc-id instituto superior técnico, technical university of lisbon r. alves redol, 9 2o– 1000-029 lisboa, portugal editor: paul piwek and kristy elizabeth boyer abstract the-mentor automatically generates multiple-choice tests from a given text. this tool aims at supporting the dialogue system of the falacomigo project, as one of falacomigo’s goals is the interaction with tourists through questions/answers and quizzes about their visit. in a minimally supervised learning process and by leveraging the redundancy and linguistic variability of the web, the-mentor learns lexico-syntactic patterns using a set of question/answer seeds. afterwards, these patterns are used to match the sentences from which new questions (and answers) can be generated. finally, several filters are applied in order to discard low quality items. in this paper we detail the question generation task as performed by the-mentor and evaluate its performance. keywords: question generation, pattern learning, pattern matching 1. introduction nowadays, interactive virtual agents are a reality in several museums worldwide (leuski et al. 2006, kopp et al. 2005, bernsen and dybkjær 2004). these agents have educational and entertainment goals – “edutainment” goals, as defined by adams et al. (1996) – and communicate through natural language. they are capable of establishing small social dialogues, teaching some topics for which they were trained, and asking/answering questions about these topics. following this idea, the falacomigo project invites tourists to interact with virtual agents through questions/answers and multiple-choice tests.1 agent edgar (figure 1) is the face of the project, currently operating in the palace of monserrate (moreira et al. 2011). 1. a multiple-choice test is composed of several multiple-choice test items, each item consisting of a question, its correct answer and a group of incorrect answers (also called distractors). c©2012 sérgio curto, ana cristina mendes, luı́sa coheur submitted 2/11; accepted 1/12; published online 3/12 curto, mendes and coheur figure 1: agent edgar the efforts of enhancing the portuguese cultural tourism started with duarte digital (mendes et al. 2009), a virtual agent based on the diga dialog framework (martins et al. 2008), which engages in inquiry-oriented conversations with users, answering questions about a famous piece of portuguese jewelry, custódia de belém. the approach based on multiple-choice tests has also been recently explored in the lisbon pantheon, where a virtual agent assesses the visitors’ knowledge about the monument. however, in all these cases, the involved questions/answers and multiplechoice tests are hand crafted, which poses a problem every time a different monument or piece of art becomes the focus of the project or a different language is to be used. therefore, one of falacomigo’s main objectives is to find a solution to automatically generate questions, answers and distractors from texts describing the monument or piece of art in study, without needing to have experts hand coding specific rules for this purpose. the-mentor (mendes et al. 2011) is our response to this challenge, as it receives as input a set of seeds – that is, natural language question/answer pairs – and uses them to automatically find, in the web, lexico-syntactic patterns capable of generating new question/answer pairs, as well as distractors from a given a text (a sentence, a paragraph or an entire document). contrary to other systems, the process of building these patterns is automatic, based on the previous mentioned set of seeds. thus, as done in many information extraction frameworks – such as dare (domain adaptive relation extraction)2(xu et al. 2010) – our approach is based on minimally supervised learning, where a set of seeds is used to bootstrap patterns that will be used to find the target type of information. figure 2 illustrates some of the information used by the-mentor. the first line shows the syntactic structure of a seed question; the second line depicts one learned lexico-syntactic pattern. different colors are used to identify the matching between the generated pattern and the text tokens. the last line shows an example of a multiple-choice test item created by the-mentor.3 in this paper, we present our approach to the text-to-question task (as identified by rus et al. (2010a)) and we thoroughly describe and evaluate the factoid question generation as performed by the-mentor. the remainder of this paper is organized as follows: in section 2 we overview related work. in sections 3 and 4 we detail the pattern learning and question generation subtasks, respectively. in section 5 we evaluate the-mentor. finally, in section 6 we present the main conclusions and point out future work directions. 2. http://dare.dfki.de/. 3. a first version of the-mentor is online at http://services.l2f.inesc-id.pt/the-mentor. in http://qa.l2f.inesc-id.pt/ one can find the corpora and scripts used/developed for the evaluation of the-mentor. the code of the-mentor will be made available in the near future. 148 question generation with the-mentor whnp vbd np np vbd by answer one of my favorite pieces is moonlight sonata composed by beethoven. another one is mozart’s requiem. qwho composed moonlight sonata ? beethoven mozart question analysis generated pattern sentences in the target document. extracted q/a pair and distractor. generated test item figure 2: information manipulated by the-mentor. 2. related work question generation (qg) is, nowadays, an appealing field of research. three recent workshops exclusively dedicated to qg and a first shared evaluation challenge, in 2010, with the goal of generating questions from paragraphs and sentences (rus et al. 2010a) have, definitively, contributed to the increase of interest in this topic. areas such as question-answering (qa), natural language generation or dialogue systems contribute to qg with techniques and resources, such as corpora with examples of questions and grammars designed to parse/generate questions. nevertheless, the specificities of qg raise many issues, namely the taxonomy that should be considered, the process of generation or the systems’ evaluation. these and many others questions led to the actual research lines. in the following we present some recent discussions and achievements related to the different research targets in qg. one of the decisions that needs to be made has to do with choosing the question taxonomy that will be used during the qg. as said by forăscu and drăghici (2009), in the context of a qg campaign, it is important to have a clear inventory of the most used question types. forăscu and drăghici (2009) survey several taxonomies proposed in the literature and point to some available resources. boyer et al. (2009) propose a hierarchical classification scheme for tutorial questions, in which the top level identifies the tutorial goal and the second level the question type, sharing many categories with previous work on classification schemes. despite the differences between this and other taxonomies existing in the literature, it should be clear that the decision of choosing a taxonomy is directly related to the goal of the task in hands. another aspect that needs to be carefully addressed in qg is related to the information sources in use. in fact, if it is obvious that the decision of using a certain resource should be influenced by the specific goal of the qg system (for instance, if the target are questions about politics, the information source should be of that domain), it is also true that different texts in the same domain can involve different difficulty levels of processing and, thus, lead to better or worse questions. for instance, a text about art in wikipedia is probably easier to process than a text about the same subject written by experts for experts. since wikipedia articles are to be read by everybody, they usually have an accessible vocabulary and simple syntax; a text written by experts for experts tends to use specific vocabulary and they are not written with the goal of being understood by everyone. therefore, a rule-based qg system will probably have more difficulty when generating plausible questions from 149 curto, mendes and coheur the latter. therefore, in some qg systems (and as done in many other natural language processing tasks), texts are simplified in a pre-processing step, before being used for qg. for instance, the approach of heilman and smith (2010) – where an algorithm for extracting simplified sentences from a set of syntactic constructions is presented – texts are simplified before being submitted to the qg module. authors show that this process results in a more suitable generation of factoid questions. kalady et al. (2010) also report a text pre-processing step before the question generation stage, where anaphora resolution and pronoun replacement is performed. still related to the type of text to be used as input, a change in the nature of the information source can lead to extensive changes in the qg system, as reported by chen et al. (2009), the authors of project listen’s reading tutor, where an intelligent tutor tries to scaffold the self-questioning strategy that readers usually use when reading a text. chen et al. (2009) explain the changes they had to do in their approach when they moved from narrative fiction to information texts. concerning resources, many systems perform qg over wikipedia (heilman and smith 2009). however, becker et al. (2009) execute their qg system over the full option science system (foss) and wyse and piwek (2009) draw the attention of the community to the openlearn repository, which covers a wide range of materials (in different formats) and is only authored by experts. in the latter, a qg system is implemented, ceist, running on the openlearn repository. corpora from this repository were also used in the evaluation of the first question generation shared task and evaluation challenge (qgstec 2010). in fact, another important resource for the qg community are the corpora provided by the coda project (piwek and stoyanchen 2010) where, in addition to a development/test set with questions and answers, a dialogue corpus is provided, where segments are aligned with monologue snippets (with the same content). the coda project team targets a system for automatically converting monologue into dialogue. in their approach, a set of rules maps discourse relations marked in the (monologue) corpus to a sequence of dialogue acts. also, as suggested by ignatova et al. (2008) sites such as yahoo! answers4 or wikianswers5 can have an important role in qg, namely in tasks such as generating high quality questions from low quality ones. regarding the qg process, different approaches have been followed. the usage of patterns to bridge the gap between the question and the sentence in which the answer can be found is a technique that has been extensively used in qa and in qg. the main idea behind this process is that the answer to a given question will probably occur in sentences that contain a rewrite of the original question. in the qa track of the trec-10, the winning system – described in (soubbotin 2001) – presents an extensive list of surface patterns. the majority of systems that target qg also follow this line of work and are based on handmade rules that rely on pattern matching to generate questions. for instance, in (chen et al. 2009), after the identification of key points, a situation model is built and question templates are used to generate questions from those situation models. the ceist system described in (wyse and piwek 2009) uses syntactic patterns and a tool called tregex (levy and andrew 2006) that receives as input a set of hand-crafted rules and matches these rules against the parsed text, generating, in this way, questions (and answers). kalady et al. (2010) bases their qg process in up-keys – that are significant phrases in documents – in parse tree manipulation and named entity recognition (ner). up-keys are used to generate definitional questions; tregex is, again, used for parse tree manipulation (as well as tsurgeon (levy and andrew 2006)); named entities are used to generate factoid questions, according to their semantics. for instance, if a named 4. http://answers.yahoo.com/ 5. http://wiki.answers.com/ 150 question generation with the-mentor entity of type people is detected, a who-question is triggered (if some conditions are verified). heilman and smith (2009) also consider a transformation process in order to generate questions. considering the evaluation of qg systems, several methods have been proposed. chen et al. (2009), for instance, classify the generated questions as plausible or implausible, according to their grammatical correctness and if they make sense in the context of the text. becker et al. (2009) give a detailed evaluation of their work, where automatically generated questions are compared against questions generated by humans (experts and not experts). they concluded that questions from human tutors outscore their system. the computer-aided environment for generating multiple-choice test items, described in mitkov et al. (2006) also proposes a similar evaluation, where multiplechoice tests generated by humans are compared with the ones generated by their system. the authors start by identifying and extracting key-terms from the source corpora, using regular expressions that match nouns and noun-phrases; afterwards, question generation rules are applied to sentences with specific structures; finally, a filter assures the grammatical correctness of the questions, although in a post-editing phase, results are revised by human assessors. this system was later adapted to the medical domain (karamanis et al. 2006). the work described by heilman and smith (2009) should also be mentioned, as they propose a method to rank generated questions, instead of classifying them. 3. learning patterns in the-mentor, a set of question/answer pairs is used to create the seeds that will be the input of a pattern learning process. its pattern learning algorithm is similar to the one described by ravichandran and hovy (2002). for a given question category, their algorithm learns lexical patterns (instead, our patterns also contain syntactic information) between two entities: a question term and the answer term. these entities are submitted to altavista and, from the 1000 top retrieved documents, the ones containing both terms are kept. the patterns are the longest matching substrings extracted from these documents, given that the question term and the answer term are replaced by the tokens and , respectively. for instance, for the category birthdate, a learned pattern is was born in . some examples of this line of work are the qa systems described by brill et al. (2002) and figueroa and neumann (2008). the former bases the performance of the qa system askmsr on manually created rewrite rules which are likely substrings of declarative answers to questions; the latter searches for possible answers by analysing substrings that have similar contexts of already known answers and uses genetic algorithms in the process. brill et al. (2002) also explain the need to produce less precise rewrites, since the correct ones did not match any document. we also feel this need and the-mentor also accepts patterns that do not match all the components of the question. in the following section we describe the component responsible for learning lexico-syntactic patterns in the-mentor. we start by explaining how the-mentor builds seed/validation pairs that originate the patterns. afterwards, we present the actual process of pattern learning and describe the three types of learned patterns: strong, inflected and weak. finally, we end with a description of how the learned patterns are validated. 3.1 building seed/validation pairs the-mentor starts by building a group of seed/validation pairs based on a set of questions and their respective correct answers (henceforward, a pair composed of a question and its answer will 151 curto, mendes and coheur be denoted as question/answer (q/a) pair). figure 3 shows examples of q/a pairs given as input to the-mentor. “in which city is the wailing wall?; jerusalem” “in which city is wembley stadium?; london” “in which sport is the cy young trophy awarded?; baseball” “in which sport is the davis cup awarded?; tennis” “who painted the birth of venus?; botticelli” “who sculpted the statue of david?; michelangelo” “who invented penicillin?; alexander fleming” “where was the first zoo?; china” “where was the liberty bell made?; england” “how many years did rip van winkle sleep?; twenty” “how many years did sleeping beauty sleep?; 100” figure 3: questions and their correct answers used as input to the-mentor. a seed/validation pair is composed by two q/a pairs: the first pair is used to query the selected search engine and extract patterns from the retrieved results – the seed q/a pair; the second pair validates the learned patterns by instantiating them with the content of its question – the validation q/a pair. for the pattern learning subtask, the questions in a seed/validation pair are required to have the same syntactic structure and prefix, and the answers are supposed to be of the same category. for instance, [“how many years did rip van winkle sleep?; twenty”/“how many years did sleeping beauty sleep?; 100”] is a seed, while [“who pulled the thorn from the lion’s paw?; androcles”/“what was the color of christ’s hair in st john’s vision?; white”] is not. in order to create seed/validation pairs from the given set of q/a pairs, the-mentor performs three steps: 1. all questions are syntactically analyzed. this step allows to group q/a pairs that are syntactically similar and avoiding the creation of seed composed by pairs like “how many years did rip van winkle sleep?; twenty” and “who sculpted the statue of david?; michelangelo” (see table 1). in this work, we used the berkeley parser (petrov and klein 2007) trained on the questionbank (judge et al. 2006), a treebank of 4,000 parse-annotated questions; index 0 1 2 3 tag whnp vbd np vb content how many years did rip van winkle sleep index 0 1 2 3 tag whnp vbd np content who sculpted the statue of david table 1: syntactic constituents identified by the parser for two input questions. 2. the content of the first syntactic constituent of the question (the question’s prefix) is collected. this is a first step towards a semantically-driven grouping of q/a pairs, allowing the creation of pairs like [“what color was the maltese falcon?; black”/“what color was moby dick?; 152 question generation with the-mentor white”]. however, it also leads to wrongly grouped pairs like [“what was the language of nineteen eighty four?; newspeak”/“what was the color of christ’s hair in st john’s vision?; white”] (the question’s prefixes are what color and what, respectively); 3. all questions are classified according to the category of the expected answer. this step allows us to group the q/a pairs when the answers share the same semantic category, such as “who painted the birth of venus?; botticelli” and “who sculpted the statue of david?; michelangelo”, both expecting a person’s name as answer. in this work, a machine learning-based classifier was used, and fed with features derived from a rule-based classifier (silva et al. 2011). the used taxonomy is li and roth’s two-layer taxonomy (li and roth 2002), consisting of a set of six coarse-grained categories (abbreviation, entity, description, human, location and numeric) and fifty fine-grained ones. this taxonomy is widely used by the machine learning community (silva et al. 2011, li and roth 2002, blunsom et al. 2006, huang et al. 2008, zhang and lee 2003), because the authors have published a set of nearly 6,000 labeled questions (the university of illinois at urbana-champaign dataset), freely available on the web, making it a very valuable resource for training and testing machine learning models. after these steps, the q/a pairs are divided into different groups according to their syntactic structure, prefix and category, and the-mentor builds seed/validation pairs from all combinations of two q/a pairs in every group. thus, if a group has n q/a pairs, the-mentor builds ( n 2 ) = n! 2(n− 2)! seed/validation pairs. the following are examples of built seed/validation pairs, grouped according to the syntactic structure and category of their questions: location:city –whpp vbz np [“in which city is the wailing wall?; jerusalem”/“in which city is wembley stadium?; london”] human:individual – whnp vbd np [“who painted the birth of venus?; botticelli”/“who sculpted the statue of david?; michelangelo”] [“who sculpted the statue of david?; michelangelo”/“who invented penicillin?; alexander fleming”] [“who invented penicillin?; alexander fleming”/“who painted the birth of venus?; botticelli”] entity:word – whnp vbz np [“what is the last word of the bible?; amen”/“what is the hebrew word for peace used as both a greeting and a farewell?; shalom”] human:individual – whnp vbd np [“who was don quixote’s imaginary love?; dulcinea”/“who was shakespeare’s fairy king?; oberon”]. 153 curto, mendes and coheur 3.2 query formulation and passage retrieval the-mentor captures the frequent patterns that relate a question to its correct answer from the web. the extraordinary dimensions of this particular document collection allow the existence of certain properties that are less expressive in other built document collections: the same information is likely to be replicated multiple times (redundancy) and presented in various ways (linguistic variability). the-mentor explores these two properties – redundancy and linguistic variability – to learn patterns from a seed/validation pair. the process of accessing the web relies on submissions to a web search engine of several queries. for a given seed q/a pair, different queries are built from the permutations of the set composed by: (a) the content of the phrase nodes (except the wh-phrase) of the question, (b) the answer and (c) a wildcard (*). the-mentor uses phrase nodes instead of words (as done by ravichandran and hovy (2002)) since they allow us to reduce the number of permutations and, thus, of query submissions to the chosen search engine; also, they usually represent a single unit of meaning, and therefore should not be broken down into parts (except for verb phrases). for instance, considering the question who sculpted the statue of david?, the corresponding parse tree and phrasal nodes are depicted in figure 4 (we added a numeric identifier to the tree nodes to help their identification.). it does not make sense to divide the noun phrase the statue of david, as it would generate several meaningless permutations, such as statue the sculpted michelangelo of * david. the wildcard is used as a placeholder for one or more words and allows diversity in the learned patterns.6 for instance, the wildcard in the query "michelangelo * sculpted the statue of david" allows the sentence michelangelo has sculpted the statue of david to be present in the search results, matching the verb has. one query is built from one permutation enclosed in double quotes, which will constrain the search results to contain the words in the query in that exact same order, without any other change. for example, the query "sculpted the statue of david * michelangelo" requires the presence in the search results of the words exactly as they are stated in the query. here the wildcard stands for one or more extra words allowed between the tokens david and michelangelo. the total number of permutations (queries submitted) is n! − 2(n − 1)!, in which n is the number of elements to be permuted. finally, the query is sent to a search engine.7 both "michelangelo * sculpted the statue of david" and "the statue of david sculpted * michelangelo" are examples of two queries submitted to the search engine for the seed q/a pair “who sculpted the statue of david?; michelangelo”. the brief summaries, or passages, retrieved by the search engine to the submitted query are then used to learn patterns. being so, the passages8 are broken down into sentences, parsed (we have used, once again, the berkeley parser, this time trained with the wall street journal) and if there is a text fragment that matches the respective permutation, we rewrite it as a pattern. the pattern is composed of the lexico-syntactic information present in the matched text fragment, thus, expressing the relation between the question and its answer present in the sentence. 6. the wildcard is not allowed as the first or the last element of the permutation and, thus, of the query. 7. in this work, we use google. however, there is no technical reason that hinders the-mentor from using other search engine. 8. for now, we ignore the fact that passages are often composed of incomplete sentences. 154 question generation with the-mentor (a) 1 root 2 sbarq 16 . ? 5 sq 6 vp 8 np 12 pp 14 np 15 nnp david 13 in of 9 np 11 nn statue 10 dt the 7 vbd sculpted 3 whnp 4 wp who (b) 3 whnp 4 wp who 7 vbd sculpted 8 np the statue of david figure 4: (a) parse tree of the question who sculpted the statue of david?. (b) phrase nodes of the parse tree in (a). algorithm 1 describes the pattern learning algorithm, which includes the procedures to build the query, the call to the search engine and the reformulation of text fragments into patterns. consider again the question who sculpted the statue of david?, its flat syntactic structure “[whnp who] [vbd sculpted] [n the statue of david]”, its correct answer michelangelo, the sentence michelangelo has sculpted the statue of david (that matches the permutation "michelangelo * sculpted the statue of david") and its parse tree depicted in figure 5. a resulting pattern would be np{answer} [has] vbd np, where: the tag {answer} indicates the position of the answer; the syntactic information is encoded by the tags of the parsed text fragment; and, the lexical information is represented by the tokens between square brackets. note that the lexical information that composes the patterns refers to the tokens learned from the sentence, but that are not present in the seed question, nor in the answer. 155 curto, mendes and coheur algorithm 1 pattern learning algorithm procedure pattern-learning(seed-pair : question-answer pair) patterns← [] phrase-nodes← get-phrase-nodes(seed-pair.question.parse-tree) for each permutation in permute({phrase-nodes, ∗, seed-pair.answer}) do query ← enclose-double-quotes(permutation) results← search(query) for each sentence in results.sentences do if matches(sentence, permutation) then pattern← rewrite-as-pattern(sentence, phrase-nodes) patterns← add(patterns, pattern) end if end for end for return patterns end procedure 1 root 2 s 17 . . 5 vp 7 vp 9 np 13 pp 15 np 16 nnp david 14 in of 10 np 12 nn statue 11 dt the 8 vbn sculpted 6 vbz has 3 np 4 nnp michelangelo figure 5: parse tree of the sentence michelangelo has sculpted the statue of david. 3.3 strong, inflected and weak patterns patterns that withhold the content of all phrases of the seed question are named strong patterns. these are built by forcing every phrase (except the wh-phrase) to be present in the patterns. this is the expected behaviour for nounand prepositional-phrases that should be stated ipsis verbis in the sentences from where the patterns are learned. however, the same does not apply for verb156 question generation with the-mentor phrases. for instance, the system should be flexible enough to learn the pattern for the question who sculpted the statue of david? from the sentence michelangelo finished sculpting the statue of david in 1504, even though the verb to sculpt is inflected differently. therefore, another type of patterns is allowed – the inflected patterns – where the main verb of the question is replaced by each of its inflections and the auxiliary verb (if it exists) is removed from the permutations that build the queries sent to the search engine. being so, the next queries are also created and submitted to the search engine: "michelangelo * sculpting the statue of david", "michelangelo * sculpt the statue of david" and "michelangelo * sculpts the statue of david". this allows the system to find patterns that are not in the same verbal tense of the question. a final type of patterns – the weak patterns – is also possible. in order to learn these patterns, only the nounand prepositional-phrases of the question, the answer and the wildcard are submitted to the search engine, that is, verb-phrases are simply discarded. following the previous example, both "michelangelo * the statue of david" and "the statue of david * michelangelo" are the two queries that might result in weak patterns. although they do not completely rephrase the question, these patterns are particularly interesting because they can capture the relation between the question and the answer. for instance, the pattern np [,] [by] np{answer} should be learned from the sentence the statue of david, by michelangelo, even if it does not include the verb (in this case, sculpted). these patterns are different from the strong and inflected patterns, not only because of the way they were created, but also because they trigger distinct strategies to question generation, as we will see. finally, we should mention that the-mentor patterns are more complex than the ones presented, as they are explicitly linked to their seed question by indexes, mapping the position of each one of their components into the seed question components. for the sake of simplicity, we decided to omit these indices.9 the syntax of the patterns, however, is not as sophisticated as the ones allowed by tregex (levy and andrew 2006), since it does not allow so many node-node relations. 3.4 patterns’ validation the first q/a pair of each seed/validation pair – the seed q/a pair – is necessary for learning strong, inflected and weak patterns, as explained before. however, these patterns have to be validated by the second q/a pair. this validation is required because, although many of the learned patterns are generic enough to be applied to other questions, others are too closely related to the seed pair and, therefore, too specific. for instance, the pattern np [:] nnp{answer} [’s] [renaissance] [masterpiece] [was] vbn (learned from the sentence statue of david: michelangelo’s renaissance masterpiece was sculpted from 1501 to 1504.) is specific to a certain cultural movement (the renaissance) and will only be applied in sentences that refer to it. therefore, the validation q/a pair of the seed/validation pair is parsed, matched against the previously learned patterns and sent to the search engine. for a candidate pattern to be validated, its components are instantiated with the content of the question’s respective syntactic constituents. the result of this instantiation is enclosed in double quotes and sent to the search engine. a score that measures the generality of the pattern is given by the ratio between the number of retrieved text snippets and the maximum number of text snippets retrieved by the search engine. if this score is above a certain threshold, the pattern is validated. 9. we will refer to these indices in section 4, which is dedicated to the question generation subtask. 157 curto, mendes and coheur for example, consider the seed/validation pair [“who sculpted the statue of david?; michelangelo”/“who invented penicillin?; alexander fleming”], that whnp vbd np is the parsing result of the question in the second q/a pair and that np vbd by {answer} is a candidate pattern. the result of the instantiation of the pattern, later submitted to the search engine, is: "penicillin invented by alexander fleming". this strategy is possible because the-mentor forces both questions of a seed/validation pair to be syntactically similar (whnp vbd np is also the parsing result of the question in the first q/a pair). therefore, the components of the candidate pattern can be associated with the syntactic constituents of the second q/a pair. algorithm 2 summarizes this step. in this work, the maximum number of text snippets was 16 and the threshold was set to 0.25. algorithm 2 pattern validation algorithm procedure pattern-validation(patterns : patterns to validate, validation-pair : questionanswer validation pair) threshold← x validated-patterns← [] for each pattern in patterns do query ← rewrite-as-query(validation-pair, pattern) query ← enclose-double-quotes(query) results← search(query) score← results.size/max-size if score ≥ threshold then validated-patterns← add(validated-patterns, pattern) end if end for return validated-patterns end procedure table 2 shows a simplified example of a set of patterns (with different types) learned for questions with syntactic structure whnp vbd np and category human:individual. question human:individual-whnp vbd np score pattern type 0.625 {answer} [’s] np weak 0.9375 {answer} [painted] np weak 0.25 {answer} [began] vbg np inflected 1.0 {answer} [to] vb np inflected 0.625 np vbd [by] {answer} strong 1.0 {answer} vbd np strong table 2: examples of learned patterns with respective score and type. 158 question generation with the-mentor 4. question generation in this section we describe the component responsible for generating questions in the-mentor. we start by presenting our algorithm for matching the lexico-syntactic patterns. afterwards, we show how question generation is performed depending on the type of learned patterns and we finish with a description of the filters used to discard generated questions of low quality. 4.1 pattern matching the subtask dedicated to the question (and answer) generation takes as input the previously learned set of lexico-syntactic patterns and a parsed target text (a sentence, paragraph or a full document) from which questions should be generated. afterwards, the patterns are matched against the syntactic structure of the target text. for a match between a fragment of the target text and a pattern to occur, there must be a lexicosyntactic overlap between the pattern and the fragment. thus, two conditions must be met: • the fragment of the target text has to share the syntactic structure of the pattern. that is, in a depth-first search of the parse tree of the text, the sequence of syntactic tags in the pattern has to occur; and • the fragment of the target text has to contain the same words of the pattern (if they exist) in that exact same order. this means that each match is done both at the lexical level – since most of the patterns include surface words – and at the syntactic level. for that purpose, we have implemented a (recursive) algorithm that explores the parsed tree of a text in a top-down, left-to-right, depth-first search. this algorithm tests if the components of the lexico-syntactic pattern are present in any fragment of the target text. for example, consider the inflected pattern np{answer} [began] vbg np, a sentence in november 1912, kafka began writing the metamorphosis in the target text and its parse tree (figure 6). in this case, there are two text fragments in the sentence that syntactically match the learned pattern, namely those conveyed by the following phrase nodes: • 5 np 12 vbd 15 vbg 16 np – november 1912 began writing the metamorphosis • 9 np 12 vbd 15 vbg 16 np – kafka began writing the metamorphosis given that the pattern contains lexical information (the token began), we test if this information is also in the text fragment. being so, there is a lexico-syntactic overlap between the pattern and each of the text segments. in subsection 4.3 we will explain how we discard the first matched text fragment. 4.2 running strong, inflected and weak patterns after the matching, the generation of new questions (and extraction of their respective answers) is straightforward, given that we keep track of the q/a pairs that originated each pattern and the links between them. 159 curto, mendes and coheur 1 root 2 s 11 vp 13 s 14 vp 16 np 18 nn metamorphosis 17 dt the 15 vbg writing 12 vbd began 9 np 10 nnp kafka 8 , , 3 pp 5 np 7 cd 1912 6 nnp november 4 in in figure 6: parse tree of the sentence in november 1912, kafka began writing the metamorphosis. the strategies for generating new q/a pairs differ according to the type of pattern and go as follows: strong patterns – there is a direct unification of the text fragment that matched a pattern with the syntactic constituents of the question that generated the pattern. the wh-phrase from the seed question is used. figure 7 shows graphically the generation of a question for a strong pattern and all the information involved in this subtask. the colours identify components (either text or syntactic constituents) that are related. in the-mentor, these associations are done by means of indices that relate the constituents of the syntactic structure of a question to a learned pattern. that is, a pattern np2 vbd1 by np{answer} is augmented with indices that refer to the constituents of the seed question: whnp0 vbd1 np2. when the match between a pattern and a fragment in the target text occurs, we use those indices to build a new question, with the respective content in the correct position. inflected patterns – the same strategy used for the strong patterns is applied. however, the verb is inflected with the tense and person existing in the seed question and the auxiliary in the question is also used. weak patterns – there is a direct unification of the text fragment that matched a pattern with the syntactic constituents of the question that generated it. for all the components that do not appear in the fragment, the components in the question are used. the verbal information of strong and inflected patterns is then used to fill the missing elements of the weak pattern. that is, since weak patterns are learned by discarding verbal 160 question generation with the-mentor phrases in the query submitted to the search engine, sometimes questions cannot be generated because a verbal phrase is needed. therefore, strong and inflected patterns, which have the same seed pair in their origin as the weak pattern, will contribute with their verbal phrases. whnp vbd np np vbd by answer who composed moonlight sonata ? ; beethoven question analysis generated pattern sentence in the target document. matched text fragment. new q/a pair source q/a pair one of my favorite pieces is moonlight sonata composed by beethoven. who sculpted the statue of david ? ; michelangelo figure 7: question generation by the-mentor. 4.3 filtering to discard low quality q/a pairs, three different filters are applied: semantic filter – a first filter forces the semantic match between the question and the answer, discarding the q/a pairs where the answer does not comply with the question category. for that, at least one of words in the answer has to be associated with the category attributed to the question. by doing so, the-mentor is able to rule out the incorrectly generated q/a pair “who wrote the metamorphosis?; november 1912”. this filter is based on wordnet’s (fellbaum 1998) lexical hierarchy and on a set of fifty groups of wordnet synsets, each representing a question category, which we have manually created in a previous work (silva et al. 2011). to find if a synset belongs to any of the predefined groups, we use a breadth-first search on the synset’s hypernym tree. thus, a word can be directly associated with a higher-level semantic concept, which represents a question category. for example, the category human:individual is related with the synsets person, individual, someone, somebody and mortal. given that actor, leader and writer are hyponyms of (at least) one of these synsets, all of these words are also associated with the category human:individual. this result is important to validate the following q/a pair “who was françois rabelais?; an important 16th century writer”. here, the answer agrees with the semantic category expected by the question (human:individual) since the word writer is associated with that category. therefore, this q/a pair is not ruled out. 161 curto, mendes and coheur search filter – a different filter instantiates the permutation that generated the pattern using the elements of the extracted fragment. it is again sent to the search engine enclosed in double quotes. if a minimum number of results is returned, the question (and answer) is considered to be valid; if not, the fragment is filtered out. anaphora filter – a final filter discards questions with anaphoric references, by using a set of regular expressions. those which we have empirically verified that will not result in quality questions are also filtered out, for example what is it?, where is there? or what is one?. 5. evaluation in this section we present a detailed evaluation of the-mentor’s main steps. we start by evaluating the pattern learning subtask and then we test the-mentor in the qgstec 2010 development and test sets10 and also in the leonardo da vinci wikipedia page11. the former allows us to evaluate the-mentor in a general domain; the latter allows the evaluation of the-mentor in a specific domain – a text about art, as this domain is the ultimate target of the-mentor. we conclude the section with a discussion about the attained results. 5.1 pattern learning regarding the pattern learning subtask, 139 factoid questions and their respective answers were used in our experiments (this same set was used in (mendes et al. 2011)). some of these were taken from an on-line trivia game, while others were hand crafted. examples of such questions (with their respective answers) are “in what city is the louvre museum?; paris”, “what is the capital of russia?; moscow” or “what year was beethoven born?; 1770”. the set of q/a pairs was automatically grouped, according to the process described in section 3.1, resulting in 668 seed/validation pairs. afterwards, the first q/a pair of each seed/validation pair was submitted to google (see algorithm 1), and the 16 top ranked snippets were retrieved and used to learn the lexico-syntactic patterns’ candidates. a total of 1348 patterns (399 strong, 729 inflected and 220 weak) were learned for 118 questions (from the original set of 139 questions). the relation between the number of learned patterns (y axis) and the number of existing questions for each category (x axis) is shown in figure 8. as it can be seen, the quantity of the learned patterns is usually directly proportional to the number of seeds; however, there are some exceptions. table 3 makes the correspondence between the questions and their respective categories. it shows that the highest number of patterns were found for category human:individual, which was the category with more seed/validation pairs. nevertheless, we can also notice that the category entity:language (e.g., which language is spoken in france?) had a small number of pairs (4) and led to a large number of patterns (195), all belonging to the type strong and inflected. also, categories location:city and location:state, despite having a similar number of seed/validation pairs, gave rise to a very disparate number of learned patterns: 115 and 69 learned patterns for 11 and 12 questions, respectively. the achieved results suggest that the pattern learning subtask depends not only on the quantity of questions given as seed, but also on their content (or quality). however, after analysing the 10. http://questiongeneration.org/qg2010. 11. http://en.wikipedia.org/wiki/leonardo_da_vinci. 162 question generation with the-mentor figure 8: number of patterns per number of questions (for each category). learned patterns # questions category strong inflected weak total 2 location:mountain 10 8 0 18 3 entity:sport 0 0 1 1 4 entity:language 53 142 0 195 5 entity:currency 15 10 0 25 11 location:city 44 67 4 115 12 location:state 23 30 16 69 23 numeric:date 37 94 30 161 24 location:country 77 112 20 209 27 location:other 63 120 39 222 28 human:individual 77 146 110 333 total 399 729 220 1348 table 3: distribution of the number of questions per category by types of learned patterns. original set of 139 questions, there is nothing in their syntax or semantics that can give us a clue about the reasons why some questions resulted in patterns and others did not. for instance, the q/a pair “in what country were the 1948 summer olympics held?; england” resulted in a pattern; however, no pattern was learned for the q/a pair “in what country were the 1964 summer olympics held?; japan” and “in what country were the 1992 summer olympics held?; spain”. note that all these three q/a pairs were used as seed and validation (recall that the-mentor creates 6 seed/validation pairs from 3 q/a pairs). 163 curto, mendes and coheur the quality of a question is difficult to estimate, since very similar questions can trigger a very disparate number of patterns. the learning of patterns relies on the large number of occurrences of certain information in the information sources, which depends on several factors, like the characteristics of the used information sources and the actuality/notoriety of the entities and/or events in the question. for instance, our feeling is that it is probable that more patterns are learned if the question asks for the birthplace of leonardo da vinci, rather than the birthplace of josé malhoa (a portuguese painter), since the former is more well known than the latter. however, the achieved results are also influenced by the type of corpora used and this situation would probably not occur if the patterns were to be learned from corpora composed of documents about portuguese art (instead of the web). finally, the algorithm we employ to learn patterns also depends on the documents the search engine considers relevant to the posed query (and the ranking algorithms used at the moment). currently, this is a variable that we cannot control. despite our intuition about the desirable properties of a question and the used information sources to our pattern learning approach, this is certainly a topic that deserves further investigation and should be a direction of future work. 5.2 question generation 5.2.1 evaluation criteria in order to evaluate the questions12 generated by the-mentor, we follow the approach of chen et al. (2009). these authors consider the questions to be plausible if they are grammatically correct and if they make sense regarding the text from which they were extracted; they are considered to be implausible, otherwise. however, and due to the fact that the system does not handle anaphora and can generate anaphoric questions (e.g., when was he sent to the front?), we decided to add another criteria: if the question is plausible, but contains, for instance, a pronoun, and context is needed in order to understand it. in conclusion, in our evaluation, each question is marked with one of the following tags: • pl: for plausible, non-anaphoric questions. that is, if the question is well formulated at the lexical, syntactical and semantical levels and makes sense in the context of the sentence that originate it. for instance who is the president of the queen international fan club? is a question marked as pl; • impl: for implausible questions. that is, if the question is not well formulated in lexical, syntactical or semantical terms, or if it could not be inferred from the sentence that originate it. as examples, questions such as where was france invented? and who is linux? are marked as impl. • ctxpl: for plausible questions, being given a certain context. that is, if the sentence is well formulated in lexical, syntactical and semantical terms, but contains pronouns or other references that can only be understood if the user is aware of the context of the question. an example of a question marked as ctxpl is who lost his only son?. 12. although we have described the-mentor’s strategy to create new q/a pairs, here we only evaluate the generated questions and ignore the appropriateness of the extracted answer. 164 question generation with the-mentor 5.2.2 open domain question generation in a first evaluation we used the development corpus from qgstec 2010 task b (sentences), which contains 81 sentences extracted from wikipedia, openlearn and yahooanswers. three experiments were carried out to evaluate the filtering process described in section 4.3; experiments were also conducted taking into account the involved type of patterns. in a first experiment all the filters were applied – all filters (semantic filter + search filter + anaphora filter) – which resulted in 29 generated questions. in the second experiment only the filter related with the question category were applied – semantic filter. this filtering added an extra set of 103 questions to the previous set of 29 questions. finally, a third experiment ran with no filters – no filters. in addition to the 29+103 questions, 1820 questions were generated. nevertheless, for evaluation purposes, only the first 500 were evaluated. results are shown in table 4.13 all filters (total: 29) strong patterns inflected patterns weak patterns pl impl ctxpl total pl impl ctxpl total pl impl ctxpl total 8 10 3 21 1 1 0 2 2 4 0 6 semantic filter (new: 103) strong patterns inflected patterns weak patterns pl impl ctxpl total pl impl ctxpl total pl impl ctxpl total 4 6 1 11 3 3 1 7 1 81 3 85 no filters (new: 1820, evaluated: 500) strong patterns inflected patterns weak patterns pl impl ctxpl total pl impl ctxpl total pl impl ctxpl total 4 41 1 46 2 42 4 48 3 402 1 406 total pl total impl total ctxpl 28 590 14 table 4: results from the qgstec 2010 corpus the first conclusion we can take from these results is that the three approaches – all filters, semantic filters and no filters – have almost equally contributed to the final set of (28) plausible questions (all filters (11), semantic filter (8) and no filters (9)). as expected, strong and inflected patterns are more precise than weak patterns, and their use results in 36-50% of plausible questions, considering the all filters and semantic filter approach (8/21 and 4/11 for strong patterns, 1/2 and 3/7 for inflected patterns). moreover, and also as expected, when no filter is applied, weak patterns over-generate implausible questions (402 impl questions in 406). regarding the number of questions generated by each sentence, in 42 questions that were labelled as pl or ctxpl, 20 had the same source sentence. thus, only 22 from the 81 sentences of qgstec 2010 resulted in pl or ctxpl sentences. considering these 22 questions, 12 were generated from sentences from openlearn, 10 from the wikipedia and 2 from yahooanswers. as there were around 24 sentences from yahooanswers and wikipedia (and the rest was from open13. to profit from the tregex engine, we repeated this same experiment by mapping our patterns into tregex format and the same set of questions were obtained. 165 curto, mendes and coheur learn), yahooanswers seems to be a less useful source for qg with the-mentor. however, more experiments need to be carried out to corroborate this observation. a final analysis was made on the distribution of the 29 questions generated with all filters regarding their categories. results can be seen in figure 9, where all the remaining categories had 0 questions associated and are not in the chart. figure 9: evaluation of generated questions according to their category (all filters applied). category location:state was omitted since only one impl question was generated from one inflected pattern. apparently, a category such as location:other is too general, thus resulting in many implausible questions. more specific categories such as human:individual, location:city or location:country attain more balanced results between the number of plausible and implausible questions generated. also, looking again at table 3, all the categories that resulted in a low number of patterns did not produce any plausible question. 5.2.3 comparing the-mentor with other systems the previous experiments using the development corpus of qgstec 2010 permitted us to evaluate the system and to chose the setting where the-mentor achieved the best results, that is, the setting that we would use if we were participating in the qgstec 2010 task b (sentences). however, it is not possible to fully simulate our participation in this challenge, since the-mentor is not prepared to receive as input a sentence and the question types that should be generated (the-mentor only works with the sentence from which to generate the questions). therefore, our results are judged with only three parameters used in the challenge: (a) relevance, (b) syntactic correctness and fluency, and (c) ambiguity (more details about the evaluation guidelines can be found in rus et al. (2010b)). our human evaluator started by studying the guidelines and the last twenty questions evaluated by qgstec 2010 evaluators. then, our annotator evaluated the first fifty questions generated by the 166 question generation with the-mentor participating systems, and the cohen’s kappa coefficient (cohen 1960) was computed to calculate the agreement between our evaluator and the evaluators of qgstec 2010 task b (as their scores are available). tables 5, 6 and 7 summarize these results for each of the evaluation parameters under consideration. it should be said that, as the qgstec scores are presented as the average between each of the evaluator’s scores, a normalization of values was made. the first column expresses our evaluation and the first row the qgstec 2010 scores (for instance, considering the first cell, it means that our evaluator gave score 1 and their evaluators gave score 1 or 1.5 to 39 questions). relevance 1 or 1.5 2 or 2.5 3 or 3.5 4 total 1 39 0 0 0 39 2 3 2 3 0 5 3 3 2 1 0 5 4 0 0 1 0 1 total 45 4 1 0 50 table 5: inter-annotator agreement for the relevance parameter. syntax 1 or 1.5 2 or 2.5 3 or 3.5 4 total 1 15 0 0 0 15 2 9 10 0 0 19 3 0 4 10 0 14 4 0 0 1 2 2 total 24 14 10 2 50 table 6: inter-annotator agreement for the syntactic correctness and fluency parameter. ambiguity 1 or 1.5 2 or 2.5 4 total 1 41 0 0 41 2 3 2 0 5 3 0 2 2 4 total 16 21 4 50 table 7: inter-annotator agreement for the ambiguity parameter. regarding relevance, the inter-annotator agreement is considered fair (0.38); in what concerns the other two variables, they are considered substantial (respectively, 0.62 and 0.65). afterwards, our annotator evaluated the questions generated by the-mentor in the sentences set. it generated 25 questions using our best predicted setting, that is, it ran with strong and inflected patterns and with the semantic filter. table 8 shows the scores attributed by our annotator to the questions generated by the-mentor according to each one of the parameters. recall that lower scores are better. 167 curto, mendes and coheur relevance syntax ambiguity #1 21 12 15 #2 1 3 4 #3 3 6 6 #4 0 4 n.a. total 25 25 25 table 8: results achieved by the-mentor. table 9 shows the comparison between the results achieved by the-mentor and the results achieved by the other systems participating at qgstec 2010.14 system relevance syntax ambiguity good questions a 1.61 ± 0.74 2.06 ± 1.01 1.52 ± 0.63 181/354 (51%) b 1.17 ± 0.48 1.75 ± 0.74 1.30 ± 0.41 122/165 (74%) c 1.68 ± 0.87 2.44 ± 1.06 1.76 ± 0.69 83/209 (40%) d 1.74 ± 0.99 2.64 ± 0.96 1.96 ± 0.71 44/168 (26%) the-mentor 1.28 ± 0.68 2.08 ± 1.88 1.64 ± 0.86 14/25 (56%) table 9: comparison between the results achieved by the-mentor and the other participating systems of qgstec 2010. results are similar if we compare relevance, syntactic correctness and fluency and ambiguity. nevertheless, our results are clearly worse than other systems if we compare the number of questions generated. however, as said before, the-mentor is not prepared to generate questions of different types for each sentence, thus the number of questions generated is low. 5.2.4 in-domain question generation we also tested the-mentor in a text about art, since this is the application domain of falacomigo. we have chosen the leonardo da vinci page from wikipedia, with 365 sentences.15 this time, motivated by the previously achieved results, we decided not to use weak patterns. once again, plausible questions such as where is leonardo’s earliest known dated work? were marked as impl, if the sentences from where they were generated do not contain reference to its answer (leonardo’s earliest known dated work is a drawing in pen and ink of the arno valley, drawn on august 5, 1473). however, we marked as pl questions that did not have an explicit answer in the sentence responsible for their generation, but it could be inferred from the sentence. an example is the question in which country was leonardo educated?, as italy can be inferred from born the illegitimate son of a notary, piero da vinci, and a peasant woman, caterina, at vinci in the region of florence, leonardo was educated in the studio of the renowned florentine painter, verrocchio. the question who created 14. for all systems, the number of good questions is calculated by taking into account only the three mentioned parameters. 15. this includes section titles and image captions. 168 question generation with the-mentor the cartoon of the virgin? was also marked as impl, because it lacks content (the name of the cartoon is the virgin and child). finally, we marked as impl questions with grammatical errors, like what sustain the canal during all seasons?. since there was almost no difference between the questions generated when all the filters were applied and questions generated with the semantic filter (the latter contributed with one additional question, that was marked as impl), this filtering process was removed from the results presented in table 10. all filters (total: 41) strong patterns inflected patterns pl impl ctxpl total pl impl ctxpl total 14 13 4 31 2 8 0 10 no filters (total: 219) strong patterns inflected patterns pl impl ctxpl total pl impl ctxpl total 25 40 4 69 21 121 7 149 total pl total impl total ctxpl total 62 182 15 259 table 10: results from leonardo da vinci’s wikipedia page. the number of generated questions per sentence in this experiment is slightly smaller than the number of generated questions per sentence in the previous experiment. this time 259 questions were generated from 365 sentences and previously there were 135 questions generated (ignoring those originated in weak patterns) from 81 sentences: that is, a ratio of 1.4 and 1.6 generated questions per sentence, respectively. in contrast, the percentage of questions marked as pl or ctxpl was slightly higher with leonardo da vinci’s page: 77 questions in 259, instead of 32 in 135. thus, the percentage of successfully generated questions is between 24% and 30%, if weak patterns are ignored. 5.2.5 inter-annotator agreement two human annotators evaluated the questions generated from the leonardo da vinci page, having the-mentor ran with all the filters enabled. the annotators were not experts and had no previous training. the guidelines were defined and explained before the annotation and the annotators did not interact during this process. again, the cohen’s kappa coefficient (cohen 1960) was used to calculate the inter-annotator agreement. table 11 summarizes the inter-annotator agreement results. annotators agreed in the evaluation of 37 of the 41 questions: 15 pl, 21 impl and 1 ctxpl. there was no agreement in 4 questions. the question who was “lionardo di ser piero da vinci”?, generated from his full birth name was ”lionardo di ser piero da vinci”, meaning ”leonardo, (son) of (mes)ser piero from vinci” was classified as pl by one annotator and as impl by the other. this is an example of a question well formulated at the lexical, syntactic and semantic level and, according to an annotator, it could be posed to a user of the-mentor; on the contrary, the other annotator considered the nonexistence of support from the sentence where it was generated as 169 curto, mendes and coheur pl impl ctxpl total pl 15 0 0 15 impl 1 21 3 25 ctxpl 0 0 1 1 total 16 21 4 41 table 11: inter-annotator agreement. sufficient reason to classify it as implausible (following the guidelines). the remaining 3 questions were classified as impl by one annotator and as ctxpl by the other. one example of such questions is who wrote an often-quoted letter? generated from at this time leonardo wrote an often-quoted letter to ludovico, describing the many marvellous and diverse things that he could achieve in the field of engineering and informing the lord that he could also paint. therefore, the relative observed agreement among raters, pr(a), is 0.90, the probability of chance agreement, pr(e), is 0.46 and, finally, the kappa coefficient (k)16 is 0.82, which is, as considered by some authors an almost perfect agreement. the guidelines for this evaluation were set ahead. our main goals during the definition of the guidelines were to make them simple and intuitive (to minimize uncertainty during the annotation process), general enough to be used in other evaluations and to allow comparisons between qg systems, but also adapted to the expected output of the-mentor (for example, given that the system does not solve anaphora, it generates anaphoric questions and in those situations the tag ctxpl is available). in our opinion, these characteristics of the guidelines led to the high agreement between annotators. also, the evaluation of certain questions was a straightforward process that simply did not raise any doubts: for instance, it was obvious to mark as impl the question in which continent is a drawing, taken from leonardo’s earliest known dated work is a drawing in pen and ink of the arno valley, drawn on august 5, 1473. however, it is also visible the different perceptions about the quality of the generated questions by the different annotators. whereas one annotator classified the question by itself and its usefulness when presented alone to a human user, the other classified it taking in consideration the entire task of qg: for instance, despite the lexical, syntactic and semantic correction of the generated question, if the answer can not be found in the supporting sentence, the annotator marked it as implausible. 6. conclusions and future work we described the qg task as performed by the-mentor, a platform that automatically generates multiple-choice tests. the-mentor takes as input a set of q/a pairs, constructs seeds and learns lexico-syntactic patterns from these seeds. three types of patterns are possible: strong, inflected and weak patterns. the first type imposes more constraints, since the seed elements need to be present in the learned patterns; the second adds some flexibility, allowing seeds and patterns not to agree in the verb forms; the third only imposes noun-phrases from the seeds to be present in the patterns. these patterns are used to generate questions, being given a text as input. the filters applied both to the learned patterns and to the generated questions were also described. moreover, 16. k = (pr(a)-pr(e))/(1-pr(e)). 170 question generation with the-mentor we have evaluated the-mentor on the qgstec 2010 corpus, as well as on a wikipedia page. results show that: • there is a linear relation between the number of seeds with a certain category and the number of resulting patterns, although there are exceptions; • only questions with a large amount of patterns contribute to a successful qg; • questions with coarse-grained categories mostly generate implausible questions; • weak patterns over-generate implausible questions; • the percentage of successfully generated questions is between 24% and 30%, if weak patterns are ignored; • simulating the evaluation of the-mentor with the qgstec 2010 test set, we managed to obtain average results in what concerns relevance, syntactic correctness and fluency, and ambiguity, but the-mentor generates significantly fewer questions than the other systems. there is still plenty of room for improvement in the-mentor. however, there are two points worth mentioning regarding the achieved results. on one hand, some questions were marked as impl but could be easily converted to pl questions: for instance, the question who is linux? could be transformed into what is linux?. on the other hand, some questions were marked as impl uniquely because the answer did not appear in the sentence that originated them (for instance who invented the hubble space telescope?). however, it could be worth presenting such questions to the user of the-mentor, depending on the application goal (for instance, if the-mentor is used to help a tutor on the creation of tests). in this case, the judgement of the correctness of the user’s answer has to be done by other means, either by a human evaluator (e.g., the tutor) or by using other strategies. regarding future work directions, we have identified a set of research points that will be the target of our next efforts, namely: • given the co-references that usually appear in a text which, in addition to complex syntactic constructions, make it harder to process, we intend to follow the approaches of heilman and smith (2010) and kalady et al. (2010) and pre-process the texts used to learn patterns and generate the question/answer pairs. with this, we hope to obtain more plausible questions. • one of the main causes for obtaining such an amount of implausible questions is the use of a coarse-grained question classification. we have seen that questions classified as location:other originate a large number of implausible questions, which is not the case of questions classified as location:country or location:city. thus, this taxonomy needs to be altered and enriched with more fine-grained categories. • we will also add more questions to certain categories and see if this extra set can lead to plausible questions of that category. • we will move from factoid to definition questions, although this will probably led us to different strategies. 171 curto, mendes and coheur • the learning process can become more flexible if we allow the use of synonyms in the queries submitted to the search engine. in fact, by allowing different verb forms we have created the inflected patterns that have significantly contributed to the qg of plausible questions. thus, a next step will be to expand queries during the learning process. • still considering the learning process, the validation step needs to be revised, as some of the patterns that are considered to be too specific, could be a good source of questions. for instance, a sentence such as the statue of david was sculpted around 1504 by michelangelo will never originate a pattern to a question such as who sculpted the statue of david? because of the specificity of the appearing date. however, if we are able to tag 1504 as a date, that sentence could originate a pattern, that would be able to trigger questions from sentences like mount rushmore national memorial was sculpted around 1935 by gutzon borglum. that is, not only syntactic categories and tokens will be taken into account in the pattern learning process, but named entities should also be considered. • we will also follow heilman and smith (2009) ideas and re-rank our questions. in fact, even if we consider the hypothesis of presenting a user with 100 questions from which 80 are implausible, it is better to have the most probable plausible questions in the first positions. • finally, we are currently porting the-mentor to portuguese and analysing the possibility of using dependency grammars, as they will allow the relation of long distance constituents. acknowledgments this work was supported by fct (inesc-id multiannual funding) through the piddac program funds, and also through the project falacomigo (projectovii em co-promoção, qren n 13449) that supports sérgio curto’s fellowship. ana cristina mendes is supported by a phd fellowship from fundação para a ciência e a tecnologia (sfrh/bd/43487/2008). references elizabeth s. adams, linda carswell, amruth kumar, jeanine meyer, ainslie ellis, patrick hall, and john motil. interactive multimedia pedagogies: report of the working group on interactive multimedia pedagogy. sigcse bull., 28(si):182–191, 1996. lee becker, rodney d. nielsen, and wayne h. ward. what a pilot study says about running a question generation challenge. in the 2nd workshop on question generation, 2009. niels ole bernsen and laila dybkjær. domain-oriented conversation with h.c. andersen. in proc. workshop on affective dialogue systems, pages 142–153. springer, 2004. phil blunsom, krystle kocik, and james r. curran. question classification with log-linear models. in proc. 29th annual international acm sigir conference on research and development in information retrieval (sigir ’06), pages 615–616. acm, 2006. kristy elizabeth boyer, william lahti, robert phillips, michael wallis, mladen vouk, and james lester. an empirically derived question taxonomy for task oriented tutorial dialogue. in the 2nd workshop on question generation, 2009. 172 question generation with the-mentor eric brill, susan dumais, and michele banko. an analysis of the askmsr question-answering system. in proc. acl-02 conference on empirical methods in natural language processing, emnlp ’02, pages 257–264. association for computational linguistics, 2002. wei chen, gregory aist, , and jack mostow. generating questions automatically from informational text. in the 2nd workshop on question generation, 2009. jacob cohen. a coefficient of agreement for nominal scales. educational and psychological measurement, 20(1):37, 1960. christiane fellbaum, editor. wordnet: an electronic lexical database. mit press, 1998. alejandro g. figueroa and günter neumann. genetic algorithms for data-driven web question answering. evol. comput., 16(1):89–125, 2008. corina forăscu and iuliana drăghici. question generation: taxonomies and data. in the 2nd workshop on question generation, 2009. michael heilman and noah smith. ranking automatically generated questions as a shared task. in the 2nd workshop on question generation, 2009. michael heilman and noah smith. extracting simplified statements for factual question generation. in the 3rd workshop on question generation, 2010. zhiheng huang, marcus thint, and zengchang qin. question classification using head words and their hypernyms. in proc. conference on empirical methods in natural language processing, pages 927–936, 2008. kateryna ignatova, delphine bernhard, and iryna gurevych. generating high quality questions from low quality questions. in proc. workshop on the question generation shared task and evaluation challenge, 2008. john judge, aoife cahill, and josef van genabith. questionbank: creating a corpus of parseannotated questions. in proc. 21st international conference on computational linguistics and the 44th annual meeting of the association for computational linguistics (acl-44), pages 497– 504. association for computational linguistics, 2006. saidalavi kalady, ajeesh elikkottil, and rajarshi das. natural language question generation using syntax and keywords. in the 3rd workshop on question generation, 2010. nikiforos karamanis, le an ha, and ruslan mitkov. generating multiple-choice test items from medical text: a pilot study. in proc. fourth international natural language generation conference (inlg ’06), pages 111–113. association for computational linguistics, 2006. stefan kopp, lars gesellensetter, nicole krämer, and ipke wachsmuth. a conversational agent as museum guide: design and evaluation of a real-world application. in proc. 5th international working conference on intelligent virtual agents, pages 329–343, 2005. anton leuski, ronakkumar patel, and david traum. building effective question answering characters. in proc. 7th sigdial workshop on discourse and dialogue, pages 18–27, 2006. 173 curto, mendes and coheur roger levy and galen andrew. tregex and tsurgeon: tools for querying and manipulating tree data structures. in proc. 5th international conference on language resources and evaluation, 2006. xin li and dan roth. learning question classifiers. in proc. 19th international conference on computational linguistics, pages 1–7. association for computational linguistics, 2002. filipe martins, ana mendes, márcio viveiros, joana paulo pardal, pedro arez, nuno j. mamede, and joão paulo neto. reengineering a domain-independent framework for spoken dialogue systems. in software engineering, testing, and quality assurance for natural language processing, an acl 2008 workshop, pages 68–76. workshop of acl, 2008. ana cristina mendes, rui prada, and luı́sa coheur. adapting a virtual agent to users’ vocabulary and needs. in proc. 9th international conference on intelligent virtual agents, lecture notes in artificial intelligence, pages 529–530. springer-verlag, 2009. ana cristina mendes, sérgio curto, and luı́sa coheur. bootstrapping multiple-choice tests with the-mentor. in proc. 12th international conference on intelligent text processing and computational linguistics (cicling), pages 451–462, 2011. ruslan mitkov, le an ha, and nikiforos karamanis. a computer-aided environment for generating multiple-choice test items. nat. lang. eng., 12(2):177–194, 2006. catarina moreira, ana cristina mendes, luı́sa coheur, and bruno martins. towards the rapid development of a natural language understanding module. in proc. 10th international conference on intelligent virtual agents, pages 309–315. springer-verlag, 2011. slav petrov and dan klein. improved inference for unlexicalized parsing. in human language technologies 2007: the conference of the north american chapter of the association for computational linguistics; proc. main conference, pages 404–411. association for computational linguistics, 2007. paul piwek and svetlana stoyanchen. question generation in the coda project. in the 3rd workshop on question generation, 2010. deepak ravichandran and eduard hovy. learning surface text patterns for a question answering system. in proc. 40th annual meeting on association for computational linguistics (acl ’02), pages 41–47. association for computational linguistics, 2002. vasile rus, brendan wyse, paul piwek, mihai lintean, svetlana soyanchev, and cristian moldovan. the first question generation shared task evaluation challenge. in proc. 6th international natural language generation conference, pages 251–257, 2010a. vasile rus, brendan wyse, paul piwek, mihai lintean, svetlana soyanchev, and cristian moldovan. overview of the first question generation shared task evaluation challenge. in proc. 3rd workshop on question generation, 2010b. joão silva, luı́sa coheur, ana mendes, and andreas wichert. from symbolic to sub-symbolic information in question classification. artificial intelligence review, 35:137–154, 2011. 174 question generation with the-mentor martin m. soubbotin. patterns of potential answer expressions as clues to the right answers. in proc. 10th text retrieval conference, pages 293–302, 2001. brendan wyse and paul piwek. generating questions from openlearn study units. in the 2nd workshop on question generation, 2009. feiyu xu, hans uszkoreit, sebastian krause, and hong li. boosting relation extraction with limited closed-world knowledge. in proc. 23rd international conference on computational linguistics: posters, coling ’10, pages 1354–1362. association for computational linguistics, 2010. dell zhang and wee sun lee. question classification using support vector machines. in proc. 26th annual international acm sigir conference on research and development in information retrieval, pages 26–32. acm, 2003. 175 dialogue and discourse 3(1) (2012) doi: 0.5087/dad.2012.101 concept type prediction and responsive adaptation in a dialogue system svetlana stoyanchev sstoyanchev@cs.columbia.edu columbia university, computer science department 450 computer science building, 1214 amsterdam avenue, new york, ny 10027-7003, usa amanda j. stent stent@research.att.com at&t labs – research 180 park avenue, florham park, nj 07932, usa editor: gregory aist abstract responsive adaptation in spoken dialogue systems involves a change in dialogue system behavior in response to a user or a dialogue situation. in this paper we address responsive adaptation in the automatic speech recognition module of a spoken dialogue system. we hypothesize that information about the content of a user utterance may help improve speech recognition. we use a two-step process to test this hypothesis: first, we automatically predict the task-relevant concept types likely to be present in a user utterance using features from the dialogue context and from the output of first-pass recognition of the utterance; and then, we adapt the speech recognizer’s language model to the predicted content of the user’s utterance and run a second pass of speech recognition. we show that: (1) it is possible to achieve high accuracy in determining presence or absence of particular concept types in a post-confirmation utterance; and (2) 2-pass speech recognition with concept type classification and language model adaptation can lead to improved speech recognition performance for post-confirmation utterances. keywords: speech recognition, dialog structure, error handling 1. introduction there are many possible sources of error in dialogue system processing, but the source that is most immediately obvious to the user is speech recognition errors. furthermore, despite years of research on robust handling of errorful input, most dialogue systems still cannot adapt their behavior when the user responds in an unexpected way to a speech recognition error. einstein said that the definition of insanity is doing the same thing and expecting a different result. in the presence of a speech recognition error, users frequently vary their input (changing the volume, pitch, speaking rate, words, syntax, or concepts) (litman et al. 2006, oviatt et al. 1998, shin et al. 2002). however, the dialogue system typically does not change its expectations about the form of a response, leading to cascading errors, low task completion rates and poor levels of user satisfaction. in this paper we propose a method that permits a dialogue system to responsively adapt its speech recognition behavior in the face of unexpected user input by using the dialogue context and information from the current user utterance. the proposed method uses two-pass speech recognition, with a concept type predictor that predicts the presence of task-relevant concepts in the user’s utterance. the prediction happens between the two speech recognition passes. in the first pass the c©2012 svetlana stoyanchev and amanda stent submitted 11/10; accepted 2/12; published online 2/12 speech recognizer uses a generic language model; then, the first-pass speech recognition results and other features are fed into the concept type predictor; and finally, the predicted concept types are used to select an adapted language model to be used in the second-pass speech recognition. we evaluate this method at one crucial point in a dialogue: post-confirmation utterances, or user utterances made in response to system confirmation prompts (section 4.2). confirmation prompts are yes/no questions a system produces to solicit user confirmation that it has correctly recognized an earlier input. an expected behaviour in response to a confirmation prompt is a yes or no answer. however, our observations show that users often exhibit unexpected behaviour in post-confirmation utterances, providing extra information and specifying new task-relevant concepts, causing frequent failures in speech recognition. it is particularly important for systems to recognize post-confirmation prompts correctly because recognition errors at these dialogue locations lead to cascading error subdialogues, frustrating the user and negatively impacting task success. we hypothesize that information about the content of the current user utterance as well as the dialogue history may lead to improved speech recognition. we use a two-step process to test this hypothesis. first, we test concept type prediction (section 6), and second, we test speech recognition with an adapted language model (section 7). in the concept type prediction experiment, we automatically predict the expected content of post-confirmation user utterances using features from the dialogue context and from the output of first-pass speech recognition. we show that it is possible to achieve high accuracy in determining the presence or absence of particular concept types in post-confirmation utterances. in the speech recognition experiment we adapt the speech recognizer’s language model to the predicted content of the user’s utterance and evaluate the accuracy of second-pass speech recognition. we show that two-pass speech recognition with concept type classification and language model adaptation can lead to improved speech recognition performance. the rest of this paper is structured as follows: in section 2, we define the notion of responsive adaptation and motivate this research. in section 3, we outline related work. in section 4, we describe the system and data that we used and in section 5, our experimental method. in sections 6 and 7 we present our experimental results for the concept type prediction and speech recognition experiments. finally, in sections 8 and 9 we conclude and present ideas for future work. 2. motivation adaptation is a natural and effective behavior, widely used by humans in spoken dialogue with each other and with computers (brennan and clark 1996, kraljic et al. 2008, garrod and anderson 1987, branigan et al. 2004, dubey et al. 2006b). however, most dialogue systems do not adapt their behavior. furthermore, when a dialogue system does adapt the adaptation is typically egocentric or directive. egocentric adaptation takes place when a dialogue system changes its behavior in response to a change in its own internal state (e.g. when a dialogue system changes its responses because of its internal state of misunderstanding (hockey et al. 2003)). directive adaptation takes place when a dialogue system attempts to cause a change in user behavior (e.g. by requesting input in a different modality or different manner (filisko and seneff 2005), or by modifying output to cause changes in the user’s input style (kruijff-korbayova and kukina 2008)). by contrast, responsive adaptation in spoken dialogue systems involves a change in dialogue system behavior in direct response to a user input or a dialogue situation. the change can be manifested in any of the components of a dialogue system: natural language understanding (nlu), natural language generation (nlg), dialogue management (dm), or speech recognition (asr). for example, a system 2 figure 1: dialogue systems recognition and interaction may modify its parsing lexicon and grammar based on the words and constructs used by the user (dubey et al. 2006a), or may modify its generation strategy based on a model of the user’s expertise or certainty (fukubayashi et al. 2006, forbes-riley and litman 2009). because speech recognition is the component closest to the user, the impact of responsive adaptation is most easily explored in the context of adaptation in the speech recognizer. the motivation for our two-stage adaptive recognition approach is drawn from human language processing, in which the context of an utterance helps conversation partners disambiguate speech. for example, “take this train”, “take the strain” and even “take this drain” can in normal conversational speech only be distinguished from each other by context1. a dialogue system faces a similar challenge; for example, the place name “admore” may be misrecognized as the more frequently occurring common noun “morning” in the let’s go! dialogue system. speech recognizers use two models: an acoustic model and a language model. the acoustic model, which maps acoustic frequency features to lexical units, is generated from speech data with aligned transcriptions. the language model scores sequences of lexical units using a model of the likelihood of their component n-grams (word sequences of length n). language models can be statistical (generated from a text corpus) or grammar-based (generated from a manually constructed context free grammar). this means that the goodness of fit between user utterances and the dataset or grammar used to generate the language model affects the performance of the speech recognizer. word error rates for commercial state-of-the-art open-domain speaker-independent speech recognition technology are around 25%-30% (riccardi and hakkani-tür 2003). noisy conditions, speaker accent, and out-of-vocabulary speech are among many factors that may increase the frequency of recognition errors. the performance of speech recognition is also partly dependent on the type of input the system is designed to recognize (see figure 1). limited input dialogue systems require the user to respond to each system prompt using only the concepts and keywords currently requested by the system. by contrast, flexible input dialogue systems allow the user to respond to system prompts with longer phrases and sentences and specify information other than that currently re1. we thank alistair conkie for this example. 3 quested. speech recognition (asr) accuracy in limited input systems is better than in flexible input systems (danieli and gerbino 1995, smith and gordon 1997). however, task completion rates and times can be better in flexible input systems (chu-carroll and nickerson 2000, smith and gordon 1997). researchers have shown that user training improves performance of limited input systems, while prompt design improves performance of flexible input systems. for example, tomko and rosenfeld showed that trained users communicating with a limited input dialogue system achieve better speech recognition than users communicating with a corresponding flexible input dialogue system (tomko and rosenfeld 2006). sheeder and balogh showed that in flexible input dialogue systems prompts can be formulated to maximize speech recognition accuracy and reduce the number of speech recognition timeouts (sheeder and balogh 2003). it is now common practice to adapt speech recognizers to the type, context or style of input speech (bellegarda 2004). language model adaptation has been used to improve automatic speech recognition performance in automated meeting transcription (tur and stolcke 2007), speech-driven question answering (stoyanchev et al. 2008), broadcast news recognition (gildea and hofmann 1999), and spoken dialogue systems (tur et al. 2005). language models in dialogue systems can be adapted to the dialogue state (riccardi and gorin 2000, esteve et al. 2001), the topic (iyer and ostendorf 1999, gildea and hofmann 1999), or the speaker (tur 2007). however, typically the language model is adapted based on dialogue system behavior (e.g. the topic of the system’s prompt) rather than on user behavior. language model adaptation in our work is based on user behavior; in particular, the content of the user’s current utterance and the dialogue context. 3. related work on adaptation in speech recognition language model adaptation is a technique for improving speech recognition performance. it involves adjusting probabilities in the language model or selecting the data for building the language model. the goal of this adaptation is to make the model better fit the language in input utterances, leading to improved speech recognition. riccardi and gorin (2000) describe an approach to language model adaptation in which the language model is conditioned on the current state of the dialogue system, leading to reductions in word error rate. it has now become standard practice to use dialogue state-specific language models (bechet et al. 2004), and the system we used for our experiments follows this approach. iyer and ostendorf (1999) describe an approach to language model adaptation based on topic rather than on dialogue state. by using a weighted combination of topic-specific language models, they obtained a 4.5% reduction in word error rate on the wall street journal text corpus, but only a 1.2% relative reduction in word error rate on the switchboard spoken dialogue corpus. in other work on topic-based language model adaptation, martins et al. (2010) dynamically adapt a language model for broadcast news recognition over time by using documents retrieved from the web. co-constraining speech recognition and natural language understanding (nlu) has been shown to benefit both processes. young (1994) uses output from the nlu along with acoustic model probabilities to detect misrecognized words on a second pass through the recognizer. bigi et al. (2004) describe another two-pass approach to speech recognition. the authors use terms from firstpass speech recognition to retrieve matching documents, and then interpolate a generic language model with a model built on the retrieved documents. language model adaptation can be most easily done for statistical language models. however, many dialogue systems use grammar-based language models, which can be created in the absence of 4 training data. grammar-based and statistical language modeling can be combined to improve speech recognition performance. in some research, a probabilistic grammar is used directly (e.g. jurafsky et al. (1995), knight et al. (2001)). by contrast, gorrell et al. (2002) and hockey et al. (2003) use a combination of grammar-based and statistical speech recognition in a two-pass approach. first, the user’s utterance is passed through a grammar-based language model (lm). using a threshold on confidence level, the system either accepts the utterance or passes it to a statistical lm. in our experiments we also perform language model adaptation in a two-stage speech recognition approach. instead of performing language model interpolation based on document retrieval, we perform language model selection based on concept type prediction. our concept type predictor uses features similar to those used by gabsdil and lemon (2004) and litman et al. (2006). we specifically address speech recognition of user utterances following system confirmation prompts. in the cases when a user attempts to correct the system in a post-confirmation utterance, we observe a significant increase in word error rate for the user’s utterance, which may lead to a cascading error sequence of misunderstandings. by adapting the system’s language model to the predicted task-related concepts in the user’s utterance, we achieve improved speech recognition, limiting the likelihood of cascading errors. this method operates independently of the type of language model used or the amount of training data in the language model. 4. system in our experiments we used the let’s go! dialogue system (raux et al. 2005). in this section we briefly describe let’s go!, and then discuss the task-related concepts and confirmation types present in the system. 4.1 let’s go system description let’s go! is a telephone-based dialogue system maintained and deployed at carnegie mellon university. it provides information about bus routes in pittsburgh. the system is reachable through the local port authority number outside of business hours (human operators answer the phone lines during business hours). it consequently receives calls from a diverse population of real users. let’s go! was developed using the olympus distributed dialogue framework and has the architecture shown in figure 2. system components run as separate applications communicating through a central hub. speech recognition is done by the pocket sphinx speech recognizer (huggins-daines et al. 2006). speech recognition output is parsed by phoenix, a robust parser which allows the system to skip unknown words and perform partial parsing (ward and issar 1994a). the dialogue manager was developed using ravenclaw (bohus and rudnicky 2003); in ravenclaw, dialogue structure is defined as a graph whose nodes are minimal dialogue components (typically individual exchanges), and whose edges indicate dialogue flow. the rosetta template-based generator is used for response generation. the speech synthesis component is the open source free tts system. other research olympus-based dialogue systems developed at cmu include the room line, larri (bohus and rudnicky 2002), and sbu calendar (stenchikova et al. 2007) systems2. the let’s go! domain and data were used in the first and second spoken dialogue system challenges (black et al. 2010). 2. the sbu calendar system was the first dialogue system based on the ravenclaw/olympus framework to be developed outside of cmu. 5 figure 2: let’s go! system architecture system: welcome to the cmu lets go! bus information system. how may i help you? user: i’d like to go to forbes and murray. system: to forbes and murray. where are you leaving from? user: carnegie mellon. system: from carnegie mellon. there are five buses running between carnegie mellon and forbes and murray: the 501, the 59u, the 61a, the 61b, and the 61c. what else can i do for you? user: when is the next one? system: the next 61c going to forbes and murray will leave carnegie mellon at 5:13 pm. table 1: sample dialogue with let’s go! in the two datasets we analyzed from 2005 and 2006, let’s go! received on average 40 calls per day. average call length was 12.9 turns, but there was a large standard deviation in call length. a 2005 call analysis showed a raw speech recognition word error rate of 68% (raux et al. 2005). the task success rate was estimated at 43%. table 1 shows a sample dialogue with the system. to accommodate the diverse user population and noisy speaking conditions, let’s go! is designed as a flexible-input, linear system-initiative dialogue following an initial open prompt (how may i help you?). in order to provide the user with route information, let’s go! elicits values for four task-related concepts: a departure location, a destination, a departure time, and optionally a bus route number. each concept value provided by the user is explicitly confirmed by the system. 6 concept type example user utterance place i need to go from oakland:p time leaving at four p. m.:t bus i need 28x:b table 2: examples of concept-containing user utterances to the let’s go! system. concept annotations: :p indicates place, :t indicates time, and :b indicates bus. figure 3: dialogue states and language models used in let’s go! let’s go! adapts its language model turn-by-turn according to the system’s dialogue state. it has four dialogue states corresponding to the information it elicits: first-query, place, time, and confirm (see figure 3)3. the state-specific language models are trained on user utterances from the corresponding dialogue states from previous dialogues with the system. for example, the place lm is built from responses to the where are you leaving from? prompt. the place lm is more likely to correctly recognize typical user responses with location-relevant vocabulary such as leaving, from, going, to and place names4. it is less likely to correctly recognize responses containing other task-relevant concepts (e.g. at four). 4.2 concepts and confirmations let’s go! is a flexible input system. it allows users to specify any combination of concepts in each state. for example, in response to a first query prompt the user can specify all of the information about the desired route (e.g. going from downtown to oakland at four p.m.), or only part of the information (e.g. leaving from downtown). in response to the place prompt, where are you leaving from?, users are likely to specify a place concept value; however, they can also take task initiative and specify values for other concepts. users can even specify a concept for which there is no state. although there is no state corresponding to a bus route request (as the system does not ask the user 3. there is also next-query which is similar to first-query and is omitted from the diagram. 4. language models used by let’s go! are hierarchical. concept values such as place names are stored in a dictionary. if the training data contains an utterance with a place concept, the relevant concept values are listed in a subsidiary language model inserted at the location of the place concept. 7 system’s confirmation question user response response type going to wood street. did i get that right? yes positive confirmation leaving from downtown. did i get that right? no, oakland rejection & correction leaving from waterfront, is this correct? yes and go to oakland topic change leaving from robinson. is this correct? from polish hill correction going to regent square. is this correct? no, braddock avenue rejection & correction the 61a. did i get that right? wondering when the next bus is topic change table 3: example answers to system confirmation prompts figure 4: word error rate on post-confirmation user utterances for a bus route explicitly), a bus route can be specified in responses to other prompts. for example, after the how may i help you? prompt the user may respond i want to take a 28x.5 users are particularly likely to specify non-requested concept values in responses to system confirmation prompts. let’s go!, like most dialogue systems, explicity confirms user-provided taskrelated concepts. the user’s response to a confirmation prompt such as leaving from waterfront? may consist of a simple confirmation (e.g. yes), a simple rejection (e.g. no), a correction (e.g. no, braddock avenue) or a topic change (e.g. no, leave at 7 or yes, and go to oakland). table 3 contains more examples of post-confirmation user utterances to the let’s go! system. the user’s response type has implications for further system processing. in particular, because let’s go! uses state-specific language models, corrections and topic changes are more likely to be misrecognized, leading to cascading errors and negatively affecting task completion rates and user satisfaction. in our analysis of let’s go! data from 2005, users specify a concept in 18% of post-confirmation utterances. 15.6% of post-confirmation utterances in the 2005 dataset contain a place concept, 3.2% 5. a bus route can also be automatically inferred by the system from the user’s start and destination locations. 8 figure 5: two-pass automatic speech recognition contain a time concept, and 6.4% contain a bus concept (see concept type features in table 4). because utterances with a concept are not well represented in the confirm language model, recognition is likely to fail on utterances containing a concept. even though such utterances are relatively infrequent, they are disproportionately important. as figure 4 shows, in let’s go! the word error rate on post-confirmation let’s go! utterances containing a concept is 10% higher than on utterances without a concept. our previous analysis of the communicator corpus (walker et al. 2002) shows that the probability of a consecutive error (when a sequence of utterances is misrecognized) is significantly higher than the probability of an initial error (stoyanchev 2009). correct prediction of the content of post-confirmation user utterances can lead to improved speech recognition, fewer and shorter sequences of speech recognition errors, and improved dialogue system performance. in short, post-confirmation utterances present an opportunity for responsive adaptation that is likely to have a positive impact on dialogue success rates. 5. experimental approach 5.1 two-pass speech recognition we adopt the two-pass recognition architecture previously introduced by young (1994) and illustrated in figure 5. in the first pass, the input utterance is processed using the confirm language model. recognition may fail on concept words such as oakland or 61c, but is likely to succeed on closed-class words (e.g. yes, no). then, the concept type predictor uses acoustic, lexical and dialogue history features to determine the task-related concept type(s) likely to be present in the utterance. in the second recognition pass, any utterance containing a concept type is re-processed using a concept-specific language model. 9 event 2005 2006 num % num % total dialogues 2411 1430 total confirm utts 9098 100 9028 100 confirms utts with a concept 2194 24 1635 18.1 dialogue state total confirm place system utts 5548 61 5347 59.2 total confirm bus system utts 1763 19.4 1589 17.6 total confirm time system utts 1787 19.6 2011 22.3 concept type features user’s post-confirm utts with place 1416 15.6 1007 11.2 user’s post-confirm utts with time 296 3.2 305 3.4 user’s post-confirm utts with bus 584 6.4 323 3.6 lexical features user’s post-confirm utts with ‘yes’ 4395 48.3 3693 40.9 user’s post-confirm utts with ‘no’ 2076 22.8 1564 17.3 user’s post-confirm utts with ‘i’ 203 2.2 129 1.4 user’s post-confirm utts with ‘from’ 114 1.3 185 2.1 user’s post-confirm utts with ‘to’ 204 2.2 237 2.6 acoustic/prosodic features feature mean stdev mean stdev duration (seconds) 1.341 1.097 1.365 1.242 energy (rms mean) 0.037 0.033 0.055 0.049 f0 mean 183.0 60.86 185.7 58.63 f0 max 289.8 148.5 296.9 146.5 table 4: statistics on post-confirmation utterances 5.2 experimental data the data we used for concept type prediction and language model adaptation comes from the first two months of let’s go! system operation in 2005 (2411 dialogues), and one month in 2006 (1430 dialogues). researchers at carnegie mellon transcribed this data and labeled the presence of task-relevant concepts by hand. in the annotated transcripts, the following concepts are labeled: neighborhood, place, time, hour, minute, time-of-day, and bus. we collapsed these concepts into three concept types: time, (including time, hour, minute, part of the day, and day of the week), place (including place and neighborhood), and bus (see table 2). table 4 shows statistics on post-confirmation user’s utterances in let’s go! for the 2005 and 2006 datasets. perhaps because of system improvements and user experience, the two data sets have some interesting differences. most confirmation prompts in both data sets are for a place (61% and 59.2% respectively). however, in the 2005 dataset the system prompted for bus concept confirmation more often than in the 2006 dataset (19.4% vs. 17.6%). perhaps some users figured out that bus is not a required piece of information; start and end locations are sufficient for the system to figure out the bus route. 10 figure 6: a confirm-type baseline approach to language modeling there are also differences in user responses to confirmation prompts. the proportion of responses containing yes, no, and/or a concept all dropped between 2005 and 2006. this may be caused by users in the 2006 dataset using more variation when responding to confirmation prompts. we also observe some differences in duration of user utterances between the two datasets. this may be due to improvement in automatic detection of the end of user speech. finally, the 2006 dataset shows higher energy rms mean measure; this may be due to change in the hardware settings. because of these differences between the two datasets, we used only the 2006 dataset for the concept type classification experiments. we used the 2005 dataset to build language models, as was done in the live version of the let’s go! system that was used to collect the 2006 dataset. in our speech recognition experiment, only the language models used for recognizing post-confirmation user utterances are different from the language models used in the 2006 system. 6. predicting concept type experiment in this section we describe our experiments on concept type prediction. first, we describe three methods for concept type prediction, two baseline methods and a machine learning method. second, we present experimental results comparing our machine learning method to the two baseline methods, and comparing the performance of different feature sets for concept type prediction. we present experimental results for both transcribed speech (best possible performance) and automatically recognized speech (real-world performance). in section 7.2 we present speech recognition results for two-pass speech recognition using the three concept type prediction methods. 6.1 baseline methods 6.1.1 no-concept baseline prediction our first baseline method predicts no concept for all utterances. since the majority of post-confirmation utterances in our data do not contain a concept, the overall accuracy of this prediction method is 82%. however, it is not useful for improving speech recognition on utterances containing a concept as its prediction for these utterances is always incorrect. 11 system’s confirm state place bus time 2005 dataset user confirm place 0.86 0.13 0.01 user confirm bus 0.18 0.81 0.01 user confirm time 0.07 0.01 0.92 2006 dataset user confirm place 0.87 0.10 0.03 user confirm bus 0.34 0.64 0.02 user confirm time 0.15 0.13 0.71 table 5: system’s confirmation state vs. user concept type 6.1.2 confirm-type baseline prediction our second baseline method uses the concept type being confirmed by the system (figure 6) as the concept type for the input user utterance. for example, if the system requests confirmation of a place, this method predicts that the user’s post-confirmation utterance will contain a place concept. there are two problems with the second baseline approach. first, the majority of utterances (82% in the 2006 dataset) do not contain any concept, so the overall accuracy of this prediction method is less than 18%. second, users may attempt topic changes in post-confirmation utterances using a different concept than the one confirmed. table 5 shows a confusion matrix for confirmation prompt concept type and post-confirmation utterance concept type. for example, in the 2006 dataset, after a system confirmation prompt for a bus, a bus concept is used in only 64% of concept-containing user utterances. 6.1.3 machine learning method we use decision trees to classify each post-confirmation user utterance according to the concept type(s) it contains (place, time, bus or none). we experimented with using the system confirmation concept type feature (the same one used in the second baseline method described above), lexical features, prosodic features, and dialogue history features. all of these features (except for those derived from the transcript) are available at run-time and can be used in a live system. the concept prediction performance on lexical features from transcripts is used to estimate the best possible performance of the lexical feature set. the concept prediction performance on lexical features from the output of the first-pass recognition is the realistic performance given noisy input. the features are outlined in table 6 and described below. system confirm-type feature (dia) the system confirm-type feature is the same one used in the confirm-type baseline method. it indicates the concept type requested in the confirmation prompt and takes the values place, bus, or time. acoustic features (raw) acoustic features are extracted from the raw audio of the user utterances. we use utterance maximum pitch (f0 max), energy (rms), duration, and the difference between f0 max in the first and second halves of the utterance. we selected these features based on the work of litman et al. (2006) on detecting speech recognition errors; we anticipated that these features would help distinguish corrections and rejections from confirmations. we used pratt (boersma and weenink) scripts to automatically extract these features from the audio. 12 feature type feature source feature description system confirm-type (dia) system log system’s confirmation prompt concept type (confirm time, confirm place, or confirm bus) acoustic (raw) raw speech f0 max; rms max; rms mean; duration; difference between f0 max in first half and in second half dialogue history (dh1, dh3) 1-3 previous utterances system’s dialogue states of previous utterances (first query, place, time, confirm place, confirm time, or confirm bus); [transcribed speech only] concept(s) that occurred in user’s utterances (yes/no for each of the concepts place, bus, time) lexical (lex) transcript/first-pass recognition output presence of specific lexical items; number of tokens in utterance; [transcribed speech only] string edit distance between current and previous user utterances asr confidence score (asr) first-pass recognition output speech recognizer confidence score concept type match (ctm) first-pass recognition output presence of concept-specific lexical items table 6: features for concept type prediction dialogue history features (dh) we use either one or three utterances of dialogue history (dh1, dh3). these features capture information about the dialogue state history (sh) and concept history (ch). in dh1, the system dialogue state for, and the concept type(s) present in, the previous user utterance are recorded. in dh3, the system dialogue states for, and the concept types presented in, the three previous user utterances are recorded. the system dialogue state values can be first query, place, time, confirm place, confirm time, or confirm bus. this feature is extracted from the system log. the concept history features are extracted from the annotated, transcribed user speech and are represented as triples of binary values indicating the presence of place, time, or bus concepts in the user’s utterance. table 7 shows the values of dialogue state (dia), state history, and concept history features on an extract from a let’s go! dialogue. user utterance #4 is a response to a confirmation prompt about a bus route number, so its dia feature value is confirm bus. the value for sh1 is first query, the value of the preceding system state corresponding to system utterance #1. the value for ch1 contains bus because the previous user utterance (#2) mentioned a bus. for utterance #4 there are no user utterances more than one back, so sh2, sh3, ch2, and ch3 are undefined. lexical features (lex) lex features include non-concept-value words and bigrams from the user’s current utterance, such as go, leave, to, or from. these features can be highly indicative of the presence or absence of any concept type, as well as of the presence of a particular concept type. for example, going to may be highly correlated with a place concept and leaving at may be correlated with a time concept. we explored two methods for identifying the most salient lexical features: manual identification and mutual information-based identification. both of these methods select a set of salient lexical features that are then used for concept type prediction. • manual approach: we manually selected five lexical features: yes (indicates a confirmation), no (indicates a rejection), to and from (indicate presence of concept types place and time), and i (indicates a complete sentence). these features were selected based on a heuristic estimate of their importance and their high relative frequency in the corpus. 13 # speaker utterance dia state history concept history 1 s (first query) what can i do for you? 2 u i want to catch the 28x 3 s (conf) the 28x. did i get that right? 4 u (post conf) yes. from the airport to downtown conf bus sh1=first query sh2=∅ sh3=∅ ch1=bus ch2=∅ ch3=∅ 5 s (conf) leaving from the airport. is this correct? 6 u (post conf) yes. conf place sh1=confirm bus sh2=first query sh3=∅ ch1=place ch2=bus ch3=∅ 7 s (conf) okay. going to downtown. is this correct? 8 u (post conf) yes. conf place sh1=confirm place sh2=confirm bus sh3=first query ch1=∅ ch2=place ch3=bus table 7: dialogue state and history features example • mutual information approach: this method was successfully used by gorin et al. (1997) in a call routing system for detecting salient phrases. we used it to select lexical features according to the mutual information between potential feature and concept types (manning et al. 2008). we extracted lexical features (unigrams and bigrams) from the transcribed user utterances. we removed all words that realize concept values (e.g. 61c, squirrel hill), as these are likely to be misrecognized in the first pass recognition of a post-confirmation utterance. we then computed the mutual information between each remaining potential lexical feature and each concept type, and selected the features with the highest mutual information scores. we computed the mutual information score i for each lexical feature t and each concept type class c ∈ { place +, place -, time +, time -, bus +, bus -} as follows: i = ntc n ∗ log2 n ∗ntc nt. ∗n.c + n0c n ∗ log2 n ∗n0c n0. ∗n.c + nt0 n ∗ log2 n ∗nt0 nt. ∗n.0 + n00 n ∗ log2 n ∗n00 n0. ∗n.0 where ntc= number of utterances where t co-occurs with c, n0c= number of utterances with c but without t, nt0= number of utterances where t occurs without c, n00= number of utterances with neither t nor c, nt.= total number of utterances containing t, n.c= total number of utterances containing c, and n = total number of utterances. table 8 shows several lexical features with high mutual information for each concept type. for example, the feature to co-occurs with the concept place in 217 utterances (ntc), and occurs without the concept place in only 39 utterances (nt0), so presence of this feature in an utterance is indicative of presence of a place. the feature yes, on the other hand, occurs without the concept place in 3652 utterances and with the concept place in only 41 utterances, so it is indicative of absence of place. 14 features n0c nt0 n00 ntc info. measure place yes 964 3652 2501 41 0.127 to 788 39 6114 217 0.069 from 828 25 6128 177 0.058 going 891 14 6139 114 0.038 bus yes 307 3678 3158 15 0.036 the 232 80 6756 90 0.036 the next 297 26 6810 25 0.0089 time yes 167 3690 3298 3 0.022 at 151 26 6962 19 0.0085 on 166 23 6965 4 0.0008 table 8: mutual information for selected features. we tried two methods for selecting features with the highest mutual information. in the first method, we selected for each concept type the 50 features with the highest mutual information. in the second method, we selected for each concept type the 30 features with the highest mutual information that occurred at least 20 times in the training data6. asr confidence score feature (asr) we used the speech recognizer’s confidence score for the first-pass recognition for each utterance (using the generic-confirm language model). concept type match features (ctm) the ctm features indicate whether a user’s utterance matches a concept value for a particular concept type. we tokenized all concept values (names of bus stops, places, buses, and times). each automatically recognized user utterance was matched to the bag of tokens for each of the concepts, and the result assigned to one of the three binary features ctm place, ctm bus, and ctm time. for example, the ctm place feature is set to true when a recognized utterance matches a part of one of the place concept values, such as street or avenue. for transcribed speech there is a one-to-one correspondence between presence of the concept and the ctm feature, so this feature alone gives 100% concept prediction accuracy. consequently, we only evaluate this feature for recognized speech. we hypothesize that the ctm feature will improve cases where part of (but not the whole) concept value is recognized in first-pass recognition. so, if in the utterance madison avenue, avenue (but not madison), is recognized in first-pass recognition, the ctm feature can flag the utterance for place, helping the classifier to correctly assign the place type to the utterance. then, in second-pass recognition the utterance will be decoded with a place concept-specific language model, potentially improving speech recognition performance. 6.2 experimental results in this section we present experimental results for concept type prediction for both baseline methods and for the machine learning method. we performed a series of 10-fold cross-validation experiments 6. we aimed to select an equal number of features for each concept type, while ensuring that each feature had mutual information in the top 25%. 30 was an empirically derived threshold for the number of lexical features to satisfy these conditions. 15 measure description formula pre+ precision of predicting presence of a concept tp/(tp+fp) rec+ recall of predicting presence of a concept tp/(tp+fn) f+ f-measure for predicting presence of a concept 2*[rec+]*[pre+] / ([pre+] + [rec+]) acc overall accuracy (tp+tn)/(tp+tn+fp+fn) switch+ error due to misclassification of utts with concept with an incorrect concept 1-(tp/all utts with concept) switch error due to misclassification of any utt with an incorrect concept 1-((tp+fp)/all utts) table 9: measures of concept prediction. tp=true positives, tn=true negatives, fp=false positives, fn=false negatives to examine the impact on concept type prediction of different methods and of different feature combinations. we trained three binary classifiers for each experiment, one for each concept type, i.e. we separately classified each post-confirmation utterance as place + or place -, time + or time -, and bus + or bus -. we used weka’s implementation of the j48 decision tree classifier (witten and eibe 2005).7 we report overall classification performance separately for feature combinations using lexical features from transcribed speech (table 11) and from automatically recognized speech (table 13). performance for each concept type is reported in table 12 for transcribed speech and table 14 for recognized speech. the results on transcribed speech give us an idea of the best possible performance on concept type classification. the results on recognized speech provide a realistic estimate of the performance in a live dialogue system. table 9 outlines each of our performance measures and describes how they are computed. for each experiment, we report precision (pre+) and recall (rec+) for determining presence of each concept type, and overall classification accuracy for each concept type (place, bus and time). we do not report precision or recall for determining absence of each concept type. because in the data 82.2% of utterances do not contain any concepts (see table 4), precision and recall for determining absence of each concept type are above .9 in each of the experiments. we also report overall pre+, rec+, f-measure (f+), and classification accuracy across the three concept types. to get an estimate of the potential impact on speech recognition performance, we also report the percentage of switch+ errors and switch errors. switch+ errors are the proportion of utterances containing a concept ca that are classified as containing a different concept cb . utterances containing bus classified incorrectly as time/place, time as bus/place, and place as bus/time are counted as switch+ errors. in the second pass of speech recognition these utterances will be decoded with a language model built for a concept different from the concept in the utterance and will be likely to have a higher word error rate. the switch error rate is the proportion of all utterances misclassified 7. we used decision trees because they gave good performance on our data set compared with other classification methods, and because they permit examination of the features in the learned models. 16 features classification accuracy rec+ acc lexmanual5 0.55 0.89 lextopmi50 0.52 0.88 lexfreq30 0.56 0.89 raw+dh+lexmanual5 0.57 0.89 raw+dh+lex50 0.56 0.89 raw+dh+lexfreq30 0.62 0.90 table 10: comparing approaches to selection of lexical features. concept type classification accuracy is reported on lexical features from recognized speech. best overall values in each group are highlighted in bold. as containing one of the concepts. switch errors include all of the switch+ errors and also errors on utterances with no concept classified as place, bus or time. utterances classified as containing one of the three concept types are subject to second-pass recognition using a concept-specific language model. utterances that are classified correctly as containing a particular concept type (rec+ represents proportion of correctly classified utterances with a concept) will be subject to second-pass recognition using a more appropriate language model. speech recognition performance on these utterances may improve in the second pass of speech recognition. on the other hand, utterances that are incorrectly classified as containing a particular concept type (switch+) will be subject to second-pass recognition using a poorly-chosen language model. this is a severe error that is likely to cause speech recognition performance to suffer. this means that we want to maximize rec+ and minimize switch+ errors. 6.2.1 performance of baseline methods the no-concept baseline achieves overall classification accuracy of 82% but rec+ of 0 (see table 11). switch+ on the no-concept baseline is 0 because all utterances are classified as ‘no concept’. misclassifications of utterances with a concept as ‘no concept’ are not counted as switch+ errors. (utterances with a concept misclassified as none will be decoded with the same generic confirm language model in the second pass of recognition. the word error rate from second-pass recognition will be the same as from first-pass recognition.) the confirm-type baseline achieves rec+ of .79, but overall classification accuracy of only 14%. switch+ is .17. 6.2.2 performance of machine learning method in this section, we explore the impact of different feature sets on performance of the machine learning method. all results are averaged 10-fold cross-validation results, and all models use the dia feature. for simplicity, we will call the model trained on lex features the lex model, the model trained on raw features, the raw model, and so on. we determine significance of the difference between conditions using the inference on proportions test with bonferroni correction for rec+ (which is the proportion of utterances with concepts that were correctly classified). 17 features overall pre+ rec+ f+ acc switch+ switch error error no concept baseline 0 0 0 0.82 0 0 confirm-type baseline 0.14 0.79 0.24 0.14 0.170 0.723 features from current utterance raw 0.67 0.34 0.45 0.85 0.064 0.040 lex 0.87 0.72 0.79 0.93 0.073 0.032 lex raw 0.88 0.70 0.78 0.93 0.074 0.030 +features from dialogue history dh1 lex 0.88 0.81 0.84 0.95 0.055 0.029 dh3 lex 0.89 0.78 0.83 0.94 0.052 0.026 table 11: overall concept type classification results: transcribed speech (all models include feature dia). best overall values in each group are highlighted in bold. features place time bus pre+ rec+ acc pre+ rec+ acc pre+ rec+ acc no concept baseline 0 0 .86 0 0 0.81 0 0 .92 confirm-type baseline 0.87 0.85 0.86 0.64 0.54 0.58 0.71 0.87 0.78 features from current utterance raw 0.65 0.53 0.92 0.25 0.01 0.96 0.38 0.07 0.96 lex 0.81 0.88 0.96 0.77 0.48 0.98 0.83 0.59 0.98 lex raw 0.83 0.84 0.96 0.75 0.54 0.98 0.76 0.59 0.98 +features from dialogue history dh1 lex 0.85 0.91 0.97 0.72 0.63 0.98 0.89 0.83 0.99 dh3 lex 0.85 0.87 0.97 0.72 0.59 0.98 0.92 0.82 0.99 table 12: concept type classification results for each concept: transcribed speech (all models include feature dia). lexical feature selection approaches we compare the performance of the machine learning method using three approaches to selecting lexical features: (a) manual selection, (b) lex50, automatically selecting the 50 features with the highest mutual information; and (c) lexfreq30, automatically selecting the 30 features with the highest mutual information that occur at least 20 times in the training data. as table 10 shows, the lexfreq30 feature set achieves the highest classification accuracy and rec+, both on its own and combination with other feature sets. the prosodic (raw) and dialogue history (dh) feature sets lead to additional improvements in performance. therefore, in the experiments described later in this section, all lex features are selected using the lexfreq30 approach.8 features from the current utterance (raw, lex, lex raw) we look at the performance of simple models using only the dialogue state (dia) feature, with lexical (lex) and acoustic/prosodic (raw) features from the current utterance. a model trained on raw features alone achieves rec+ 8. switch+ errors are not reported here as they did not differ across the lexical feature selection approaches. 18 features overall pre+ rec+ f+ acc switch+ switch error error no concept baseline 0 0 0 0.82 0 0 confirm-type baseline 0.14 0.79 0.24 0.14 0.170 0.723 features from current utterance raw 0.67 0.34 0.45 0.85 0.064 0.040 lex 0.75 0.56 0.64 0.89 0.099 0.049 lex raw 0.76 0.60 0.67 0.90 0.103 0.051 +features from dialogue history dh1 lex raw 0.77 0.60 0.67 0.90 0.082 0.046 dh3 lex raw 0.77 0.62 0.68 0.90 0.072 0.046 +features specific to recognized speech asr dh3 lex raw 0.77 0.62 0.68 0.90 0.072 0.045 ctm dh3 lex raw 0.85 0.74 0.79 0.93 0.039 0.029 ctm asr dh3 lex raw 0.85 0.74 0.79 0.93 0.042 0.030 table 13: overall concept type classification results: recognized speech (all models include feature dia). best overall values in each group are highlighted in bold. features place time bus pre+ rec+ acc pre+ rec+ acc pre+ rec+ acc no concept baseline 0 0 .86 0 0 0.81 0 0 .92 confirm-type baseline 0.87 0.85 0.86 0.64 0.54 0.58 0.71 0.87 0.78 features from current utterance raw 0.65 0.53 0.92 0.25 0.01 0.96 0.38 0.07 0.96 lex 0.70 0.70 0.93 0.67 0.15 0.97 0.65 0.62 0.98 lex raw 0.70 0.72 0.93 0.66 0.38 0.97 0.68 0.57 0.98 +features from dialogue history dh1 lex raw 0.71 0.68 0.93 0.68 0.38 0.97 0.78 0.63 0.98 dh3 lex raw 0.71 0.70 0.93 0.67 0.42 0.97 0.79 0.63 0.98 +features specific to recognized speech asr dh3 lex raw 0.71 0.70 0.93 0.69 0.42 0.97 0.79 0.63 0.98 ctm dh3 lex raw 0.82 0.82 0.96 0.86 0.71 0.99 0.76 0.68 0.98 ctm asr dh3 lex raw 0.82 0.81 0.96 0.86 0.69 0.99 0.76 0.68 0.98 table 14: concept type classification results for each concept type: recognized speech (all models include feature dia). 19 concept average # non-concept average # average # type words in utt words in concept chars in concept place 1.29 2.2 12.8 bus 1.63 2.9 10 time 1.73 1.7 6.6 table 15: length of user utterances containing concept of 0.34 and overall accuracy of 0.85 (see table 11). this model performs surprisingly well, beating both baselines in overall accuracy (0.85 vs. 0.82 & 0.14 for the no-concept & confirm-type baselines, both differences significant at p < .001). however, this model only works well for place concepts. as shown in table 14, the rec+ for the raw model is 0.53 for the place concept, but only 0.01 and 0.07 for the time and bus concepts. this result indicates that utterances containing values for the place concept, but not the time or bus concepts, contain prosodic information that can be used for determining presence of a concept. one possible reason for this difference in performance may be the lack of training data for the time and bus concepts (see table 4). another reason may be the difference in duration of the concept values. table 15 shows the average number of non-concept words in an utterance, the average number of words in a concept value in an utterance, and the average number of characters in a concept value in an utterance9. realizations of values for the time concept are much shorter than realizations of values for the place and bus concepts 10. the lex model for both transcribed (table 11) and recognized (table 13) speech achieves significantly higher rec+ than the raw model (0.72 & 0.56 vs. 0.34) and overall accuracy (0.93 & 0.89 vs. 0.85, all differences significant at p < .001). despite higher rec+, for recognized speech, the lex model has significantly more switch+ errors than the raw model (0.099 vs 0.064, p < .001). this means that lex and raw models differ in the type of errors that they make. the lex model makes more errors that involve mislabeling utterances with a concept by a different concept while the raw model makes more errors by mislabeling an utterance with a concept as no concept or an utterance without a concept as containing a concept. this suggests that combining the two models may improve the performance of concept prediction. for transcribed speech, the lex raw model does not perform significantly differently from the lex model in terms of overall accuracy, rec+, or switch+ errors. however, for recognized speech, lex raw achieves significantly higher rec+ (0.60) and overall accuracy (0.90) than lex (rec+ 0.56 and acc 0.89, p < .001). lexical features from transcribed speech are very good indicators of concept type. however, lexical features from recognized speech are noisy, so concept type classification for recognition output can be improved by using acoustic/prosodic (raw) features. prediction accuracy varies widely across concept types. figure 7 depicts rec+ for place, time, and bus concepts using lex from transcribed speech, lex from recognized speech, and lex raw from recognized speech. we achieve highest rec+ for the place concept for each of the feature combinations. this may be partially due to the fact that we have more training data for the place concept than for the other concepts, and partially due to more informative lexical features (e.g. to, from) in utterances containing values for the place concept. the time concept has the lowest rec+, and the biggest drop in performance due to recognition errors (difference between lex on 9. we use the number of characters to approximate the number of syllables. 10. the most common value for the time concept, now, is 3 characters long. 20 figure 7: value of rec+ for each concept (graphical illustration of the results in tables 12 and 14) transcribed and lex on recognized speech). however, prosodic features have the biggest impact on rec+ for the time concept, improving rec+ from a low 0.15 to 0.38. overall, models containing only features from the current utterance perform significantly worse than the confirmation state baseline in terms of rec+ (p < .001). however, they have significantly higher accuracy and fewer switch+ errors (p < .001) . features from the dialogue history (dh1, dh3) next, we add features from the dialogue history to our best-performing models so far. for transcribed speech (table 11), a model with one utterance of history (dh1 lex) performs significantly better than lex in terms of rec+ (0.81 vs 0.72), overall accuracy (0.95 vs. 0.93), and switch+ errors (0.055 vs. 0.073, p < .001). a model with three utterances of history (dh3 lex) performs significantly worse than dh1 lex in terms of rec+ (0.78 vs. 0.81 p < 0.05). for recognized speech (table 13), neither dh1 lex raw nor dh3 lex raw is significantly different from lex raw in terms of rec+ or overall accuracy. however, both dh1 lex raw and dh3 lex raw do perform significantly better than lex raw in terms of switch+ errors (.082 and .072 vs. .103, p < .05). there are no significant performance differences between dh1 lex raw and dh3 lex raw. features specific to recognized speech (asr, ctm) finally, we add the speech recognition confidence score (asr) and concept type match (ctm) features to models trained on recognized speech. we hypothesized that the classifier can use the recognizer’s confidence score to decide whether an utterance is likely to have been misrecognized. however, asr dh3 lex raw is not significantly different from dh3 lex raw in terms of rec+, overall accuracy or switch+ errors. this agrees with the findings of lemon and konstas (2009) that asr confidence scores show lower information gain than other features when classifying recognition hypothesis quality. by contrast, adding the concept type match (ctm) feature to dh3 lex raw leads to a large improvement on all measures: a 12% absolute increase in rec+ (from .62 to .74), a 3% absolute increase in overall accuracy (from .90 to .93), and decreases in switch+ errors (from .072 to .042), all statistically significant at p < .001. there are no statistically significant differences between ctm dh3 lex raw and ctm asr dh3 lex raw. 21 6.2.3 summary and discussion in this section we evaluated different models for concept type prediction. the best performing transcribed speech model, dh1 lex, significantly outperforms the confirm-type baseline on overall accuracy and on switch+ and switch errors (p < .001), and is not significantly different on rec+. the best performing recognized speech model, ctm dh3 lex raw, significantly outperforms the confirm-type baseline on overall accuracy and on switch+ and switch errors, but is significantly worse on rec+ (p < .001). the best transcribed speech model achieves significantly higher rec+ and overall accuracy than the best recognized speech model (p < .01), but is not significantly different in terms of switch+ errors. although the best-performing concept type prediction models achieve very high accuracy, high rec+ rates, and low rates of switch+ errors, speech recognition is a noisy process. therefore, in order to confirm that concept type prediction can lead to improved speech recognition performance, we ran a speech recognition experiment. 7. speech recognition experiment in this section we look at the impact of concept type prediction on speech recognition performance in let’s go! data. we hypothesized that speech recognition performance for post-confirmation utterances containing a concept can be improved with the use of concept-specific language models. in this set of experments, we (1) compare the existing generic confirm language model used in let’s go! with the proposed concept-specific language model adaptation strategy; (2) compare two methods for selecting user utterances for building language models; and (3) evaluate the impact of different methods of concept type prediction on concept-specific language model adaptation. 7.1 method we used the pocketsphinx speech recognition engine (huggins-daines et al. 2006) with genderspecific telephone-quality acoustic models built for communicator (rudnicky et al. 2000). we trained trigram language models using 0.5 ratio discounting with the cmu language modeling toolkit (xu and rudnicky 2000)11. we built stateand concept-specific language models from the let’s go! 2005 data. the language models are hierarchical and encode semantic information (ward and issar 1994b), smoothing probabilities for the concepts not used in the data. we evaluate speech recognition performance on post-confirmation user utterances from the 2006 let’s go! dataset. each experiment varies in 1) the language model used for the final recognition pass and 2) the method of selecting a language model for use in second-pass recognition. 7.1.1 language models we used the language model types outlined in table 16. the generic-confirm model is trained on all utterances in the 2005 dataset that were produced in the confirm dialogue state. this corresponds to the approach used in the let’s go! 2006 system. the confirm-type models are trained using all utterances from the 2005 dataset that were produced in the confirm dialogue state following confirm place, confirm bus and confirm time system confirmation prompts respectively. the 11. we used the same speech recognizer, acoustic models, language modeling toolkit, and language model building parameters that were used in the live let’s go! system (see raux et al. (2005)). 22 concept type prediction method language models data used for building language models no-concept generic-confirm all post-confirmation utterances confirm-type confirm-place post-confirmation utts after confirm place confirm-time post-confirmation utts after confirm time confirm-bus post-confirmation utts after confirm bus concept-based concept-place post-confirmation utts with place concept concept-confirm post-confirmation utts with time concept concept-confirm post-confirmation utts after bus generic-confirm all post-confirmation utterances table 16: speech recognition experiment summary: language models concept type prediction method prediction decision based on no-concept no prediction confirm-type confirm-place post-confirmation utts after confirm place confirm-time post-confirmation utts after confirm time confirm-bus post-confirmation utts after confirm bus concept-based concept-place classifier predicts place concept concept-confirm classifier predicts time concept concept-confirm classifier predicts bus none classifier predicts none or multiple concepts table 17: speech recognition experiment summary: choosing language models concept-based models are trained on all utterances from the 2005 dataset that were produced in the confirm dialogue state and contain a mention of a place, bus or time respectively. we used the three methods for choosing language models presented in section 6 and outlined in table 17. the first, no-concept method simply uses one model for recognizing all utterances. for the second method (confirm-type), we use the confirm-type baseline concept type prediction method to choose one of the three confirm-type models. for the third method (concept-based), we use one of the classifiers described in the section 6. the classifier outputs place, time, bus, or no concept, and we use the corresponding concept-based model. 7.1.2 recognizers we report results for seven experimental conditions (see table 18). the experimental conditions vary in method of building and choosing language models. in experimental conditions 1 3, recognition was done in a single pass. in the baseline experimental condition (1), we used the generic-confirm language model to process all post-confirmation utterances. in the 1-pass confirm experimental condition (2) we used the confirm-type method for building and choosing language models. we built confirm-place, confirm-bus and confirm-time language models to recognize post-confirmation utterances produced following a confirm place, confirm bus and confirm time prompt respectively12. in the 1-pass concept experimental condition (3) we used the concept-place, 12. as shown in tables 4 and 5, some, but not all, utterances in a confirmation state contain the corresponding concept. 23 exp num predict lm build lm overall concept utterances # pass method method wer wer concept recall 1 1-pass baseline baseline 38.49% 49.12% 50.75% 2 1-pass confirm-type confirm-type 38.83% 48.96% 51.36% 3 1-pass confirm-type concept-type 46.47% ** 50.73% * 52.9% † 4 2-pass dh3 lex raw concept-type 38.48% 47.56% ** 53.2% † 5 2-pass asr dh3 lex raw concept-type 38.51% 47.99% * 52.7% 6 2-pass ctm asr dh3 lex raw concept-type 38.42% 47.86% * 52.6% 7 2-pass oracle concept-type 37.85% ** 45.94% ** 54.91% ** table 18: speech recognition results. ** indicates a statistically significant difference (p<.01). * indicates a statistically significant difference (p<.05). † indicates a near-significant trend in difference (p<.1). significance for wer is computed using paired t-tests. significance for concept recall is computed as an inference on proportions. concept-bus and concept-time language models to recognize post-confirmation utterances produced following a confirm place, confirm bus and confirm time prompt respectively. in experimental conditions 4 7 we used the 2-pass recognition method outlined in figure 5. we performed first-pass recognition of post-confirmation utterances using the generic confirm language model. then, we ran the output of the first pass through a concept type classifier. finally, we performed second-pass recognition using the concept-place, concept-bus or concept-time language models if the utterance was classified as place, bus or time respectively13. we experimented with the three classification models with highest overall rec+ when trained on recognized speech: dh3 lex raw (4), asr dh3 lex raw (5), and ctm asr dh3 lex raw (6). to get an idea of “best possible” performance, we also report 2-pass oracle (7) recognition results, assuming an oracle classifier that always outputs the correct concept type for an utterance. 7.2 experimental results 7.2.1 comparison of models in table 18 we report average per-utterance word error rate (wer) on post-confirmation utterances, average per-utterance wer on post-confirmation utterances containing a concept, and average concept recall rate (percentage of correctly recognized concepts) on post-confirmation utterances containing a concept. in slot-filling dialogue systems like let’s go!, the concept recall rate largely determines the potential of the system to understand user-provided information and continue the dialogue successfully. therefore, our goal is to maximize concept recall and minimize wer on concept-containing utterances, without causing overall wer to decline. 13. we treated utterances classified as containing more than concept type as none. in the 2006 data, only 5.6% of utterances with a concept contain more than one concept type. 24 transcript generic model concept-specific model hypothesis hypothesis 1 no ardmore no five morning no ardmore 2 no leaving from hobart and murray no leaving from four a m no leaving from for murray 3 eleven o’clock d braddock o’clock eleven o’clock 4 fifth and dinwiddie fifty the 1a fifth and dinwiddie table 19: examples utterances with improved speech recognition performance in the second pass of the speech recognizer for the dh3 lex raw prediction model. as table 18 shows, the 1-pass confirm-type (2) and 1-pass concept-type (3) experimental recognizers perform better than the baseline recognizer (1) in terms of concept recall, but worse in terms of overall wer. most of these differences are not statistically significant. however, the 1-pass concept-type recognizer (3) has significantly worse overall and concept utterance wer than the baseline recognizer (p < .01). the 1-pass concept-type recognizer would in practice follow the performance of the confirm-type prediction method, which has the highest rate of switch+ (17%) and switch (72%) errors (see table 11). this is because with the confirm-type prediction method all utterances without a concept (82%) are decoded with a language model built on utterances with a concept. the switch+ error rate for this method indicates that 17% of utterances containing a concept type would be classified as containing a different concept type and decoded with a language model built for that different concept type, leading to poorer speech recognition performance. all of the 2-pass recognizers (4-7) use automatic concept prediction and achieve significantly lower concept utterance wer than the baseline recognizer (p < .05). differences between these recognizers in overall wer and concept recall are not significant. the 2-pass oracle recognizer (7) shows the best possible improvement from using concept-type language models. it achieves significantly higher concept recall and significantly lower overall and concept utterance wer than the baseline recognizer (p < .01). it also achieves significantly lower concept utterance wer than any of the 2-pass recognizers that use automatic concept prediction (p < .01). 7.2.2 analysis of improvements and errors we further analysed the effect of concept type prediction of the 2-pass dh3 lex raw model. table 19 shows examples of utterances where a correctly predicted concept-specific model improved speech recognition performance. in examples #1 and #2 the user corrects the system by making a negation and specifying a place concept. a generic model incorrectly predicts a time concept in both cases while the concept-specific model correctly recognizes a place concept in #1 and partially recognizes a place concept in #2. examples #3 and #4 show improvement of speech recognition leading to correct time and place concept detection. while the overall speech recognition and concept detection rate improves, it is possible for the recognition to degrade when a concept prediction model predicts an incorrect concept. these errors correspond to switch+ errors in table 13. the dh3 lex raw prediction model has 7.2% switch+ errors. table 20 shows examples of incorrectly recognized utterances following two-stage recognition with the dh3 lex raw prediction model and concept-specific language models. notice that example #5 shows concept type prediction leading to recognition of a different concept 25 transcript generic model concept-specific model hypothesis hypothesis 5 1a i need the 1a from 1a any one leave from we’re me what leave from mall 6 no conway no morning no port authority 7 early morning trolley morning from leave morning table 20: examples utterances with degraded speech recognition performance in the second pass of the speech recognizer for the dh3 lex raw prediction model. type (bus number 1a vs. a place), while in examples #6 and #7 neither speech recognition process leads to correct recognition, which is likely to happen if an utterance is not spoken clearly or contains background noise. in future work, we may attempt to further reduce the switch+ error rate by considering multiple speech recognition hypotheses from first-pass and second-pass recognition in parallel. 7.2.3 summary our results with two-pass recognition show that it is possible to use knowledge of (or predictions about) the concepts in a user’s utterance to improve speech recognition. our results with the onepass concept-type recognizer condition show that this cannot be effectively done by assuming that the user will always address the system’s question; instead, one must consider the user’s actual utterance and the discourse history (as in the dh3 lex raw model). 8. discussion and future work in this set of experiments, we looked at user responses to system confirmation prompts in taskoriented spoken dialogue. we explained how these post-confirmation utterances are of outsize importance in spoken dialogue, because they may contain unrequested task-relevant concepts that are likely to be misrecognized, leading to cascading errors and reduced user satisfaction. we then examined one type of responsive adaptation: a task-oriented dialogue system adapting its speech recognizer to handle user responses to system confirmation prompts. we showed that by using acoustic, lexical, dialogue state and dialogue history features, we are able to predict the presence of task-relevant concept types in first-pass speech recognition output for post-confirmation utterances with 93% accuracy. we also showed that use of a concept type predictor can lead to improvements in two-pass speech recognition performance in terms of wer and concept recall. of course, any possible improvements in speech recognition performance are dependent on (1) the performance of concept type classification; (2) the accuracy of first-pass speech recognition; and (3) the accuracy of second-pass speech recognition. for example, with the general language model, we get a fairly high overall wer of 38.49%. in future work, we will systematically vary the wer of both the firstand second-pass speech recognizers to further explore the interaction between speech recognition performance and concept type classification. the improvements the two-pass recognizers achieve have quite small local effects (up to 3.18% absolute improvement in wer on utterances containing a concept, and less than 1% on postconfirmation utterances overall) but may have larger impacts on dialogue completion times and task 26 completion rates, as they reduce the number of cascading recognition errors in the dialogue (shin et al. 2002). furthermore, we could use knowledge of the concept type(s) contained in user utterances to improve dialogue management and response planning (bohus 2007). in future work, we will look at (1) extending the use of concept-type classifiers to utterances following any system prompt; and (2) the impact of these interventions on overall metrics of dialogue success. although this work was carried out with a flexible input dialogue system where a user can speak longer phrases and sentences, the proposed approach may also be applied to fixed input systems where the vocabulary accepted by a system is limited. even in fixed input systems, a user may attempt corrections, clarifications and topic shifts in post-confirmation utterances. to avoid situations of repetitive “sorry i did not understand you” prompts, a system may attempt to predict the concept type(s) present in post-confirmation utterances and adapt its grammar or language model to decrease the chance of errors due to out-of-vocabulary speech. our results also have implications for unconstrained open-domain dialogue systems, such as turing test candidate systems. although in an unconstrained dialogue, topics and vocabulary may shift dramatically over time, pairs of consecutive utterances are related to one another and a flow can be traced throughout most coherent dialogues (schegloff and sacks 1973). this property of communication allows open-domain systems to use dialogue history to adapt their models for the upcoming utterances making recognition and understanding more tractable. an alternative method for maximizing recognition performance in dialogue systems is to guide the user to use only desired vocabulary and syntax. in separate research, we are looking at directive adaptation in spoken dialogue systems (stent et al. 2006, stoyanchev and stent 2009), exploring the potential impact on the user of micro-level design decisions in system prompt construction. 9. conclusions adaptation is an important feature of successful spoken dialogue. most dialogue systems use either egocentric adaptation (adaptation of system behavior in response to changes in system internal state) or directive adaptation (adaptation of system behavior intended to cause changes in user behavior). in this paper, we looked at responsive adaptation, adaptation in direct response to user behavior. we examined adaptation in a dialogue situation that is particularly obvious and frustrating for users: speech recognition errors in user utterances in response to system confirmation prompts. we showed that responsive adaptation has the potential to reduce the frequency and severity of these types of error, leading to improved dialogue outcomes. of course, there are many other ways to apply adaptation in dialogue systems: for example, a system may modify the type of help it provides in response to different types of processing error (hockey et al. 2003), or may modify the type of feedback it provides in response to user indications of uncertainty (forbes-riley and litman 2009), or may adjust its words and syntactic choices to match those of the user to seem more “natural” (dubey et al. 2006a). each type of adaptation on its own may have only a small impact, but together they have the effect of creating dialogue systems that are easier to use. 10. acknowledgments this research was done while both authors were at stony brook university in stony brook, ny, usa. this material is based upon work supported by the national science foundation under grant 27 no. 0325188. we thank our collaborators on this project, particularly drs. susan brennan and marie huffman. we also thank the developers of the let’s go! system, particularly drs. maxine eskenazi and antoine raux, for providing access to their system and data. references f. bechet, g. riccardi, and d. hakkani-tür. mining spoken dialog corpora for system evaluation and modeling. in proceedings of the conference on empirical methods in natural language processing (emnlp), 2004. j. r. bellegarda. statistical language model adaptation: review and perspectives. speech communication special issue on adaptation methods for speech recognition, 42:93–108, 2004. b. bigi, y. huang, and r. de mori. vocabulary and language model adaptation using information retrieval. in proceedings of interspeech, 2004. a. w. black, s. burger, b. langner, g. parent, and m. eskenazi. spoken dialog challenge 2010. in proceedings of the spoken language technology conference, 2010. p. boersma and d. weenink. praat. http://www.fon.hum.uva.nl/praat/. d. bohus. error awareness and recovery in task-oriented spoken dialog systems. phd thesis, carnegie mellon university, 2007. d. bohus and a. rudnicky. larri: a language-based maintenance and repair assistant. in proceedings of multi-modal dialogue in mobile environments, 2002. d. bohus and a. rudnicky. ravenclaw: dialog management using hierarchical task decomposition and an expectation agenda. in proceedings of eurospeech, 2003. h. branigan, m. pickering, and a. cleland. syntactic coordination in dialogue. cognition, 75: b13–b25, 2004. s. brennan and h. clark. conceptual pacts and lexical choice in conversation. journal of experimental psychology, 22(6):1482–1493, 1996. j. chu-carroll and j. nickerson. evaluating automatic dialogue strategy adaptation for a spoken dialogue system. in proceedings of the meeting of the north american chapter of the association for computational linguistics (naacl), 2000. m. danieli and e. gerbino. metrics for evaluating dialogue strategies in a spoken language system. in proceedings of the aaai spring symposium on empirical methods in discourse interpretation and generation, 1995. a. dubey, f. keller, and p. sturt. integrating syntactic priming into an incremental probabilistic parser, with an application to psycholinguistic modeling. in proceedings of the meeting of the association for computational linguistics (acl), 2006a. a. dubey, p. sturt, and f. keller. parallelism in coordination as an instance of syntactic priming: evidence from corpus-based modeling. in proceedings of the conference on empirical methods in natural language processing (emnlp), 2006b. 28 y. esteve, f. bechet, a. nasr, and r. de mori. stochastic finite state automata language model triggered by dialogue states. in proceedings of eurospeech, 2001. e. filisko and s. seneff. developing city name acquisition strategies in spoken dialogue systems via user simulation. in proceedings of the sigdial workshop on discourse and dialogue (sigdial), 2005. k. forbes-riley and d. litman. adapting to student uncertainty improves tutoring dialogues. in proceedings of the 14th international conference on artificial intelligence in education (aied), 2009. y. fukubayashi, k. komatani, t. ogata, and h. okuno. dynamic help generation by estimating user’s mental model in spoken dialogue systems. in proceedings of interspeech, 2006. m. gabsdil and o. lemon. combining acoustic and pragmatic features to predict recognition performance in spoken dialogue systems. in proceedings of the meeting of the association for computational linguistics (acl), 2004. s. garrod and a. anderson. saying what you mean in dialogue: a study of conceptual and semantic coordination. cognition, 27(2):181–218, 1987. d. gildea and t. hofmann. topic-based language models using em. in proceedings of eurospeech, 1999. a. l. gorin, g. riccardi, and j. h. wright. how may i help you? speech communication, 23: 113–127, 1997. g. gorrell, i. lewin, and m. rayner. adding intelligent help to mixed initiative spoken dialogue systems. in proceedings of the international conference on spoken language processing (icslp), 2002. b. hockey, o. lemon, e. campana, l. hiatt, g. aist, j. hieronymus, a. gruenstein, and j. dowding. targeted help for spoken dialogue systems: intelligent feedback improves naive users’ performance. in proceedings of the meeting of the european chapter of the association for computational linguistics (eacl), 2003. d. huggins-daines, m. kumar, a. chan, a. black, m. ravishankar, and a. rudnicky. pocketsphinx: a free, real-time continuous speech recognition system for hand-held devices. in proceedings of the international conference on acoustic, speech and signal processing (icassp), 2006. r. iyer and m. ostendorf. modeling long distance dependencies in language: topic mixtures versus dynamic cache model. ieee transactions on speech and audio processing, 7(1):30–39, 1999. d. jurafsky, c. wooters, j. segal, a. stolcke, e. fosler, g. tajchman, and n. morgan. using a stochastic context-free grammar as a language model for speech recognition. in proceedings of the international conference on acoustic, speech and signal processing (icassp), 1995. s. knight, g. gorrell, m. rayner, d. milward, r. koeling, and i. lewin. comparing grammar-based and robust approaches to speech understanding: a case study. in proceedings of eurospeech, 2001. 29 t. kraljic, a.g. samuel, and s.e. brennan. first impressions and last resorts: how listeners adjust to speaker variability. psychological science, 19:332–338, 2008. i. kruijff-korbayova and o. kukina. the effect of dialogue system output style variation on users’ evaluation judgments and input style. in proceedings of the sigdial workshop on discourse and dialogue, 2008. o. lemon and i. konstas. user simulations for context-sensitive speech recognition in spoken dialogue systems. in proceedings of the meeting of the european chapter of the association for computational linguistics (eacl), 2009. d. litman, j. hirschberg, and m. swerts. characterizing and predicting corrections in spoken dialogue systems. computational linguistics, 32:417–438, 2006. c. d. manning, p. raghavan, and h. schütze. introduction to information retrieval. cambridge university press, 2008. c. martins, a. teixeira, and j. neto. dynamic language modeling for european portuguese. computer speech and language, 24:750–773, 2010. s. oviatt, j. bernard, and g. a. levow. linguistic adaptations during spoken and multimodal error resolution. language and speech, 41:419–442, 1998. a. raux, b. langner, a. black, and m eskenazi. let’s go public! taking a spoken dialog system to the real world. in proceedings of eurospeech, 2005. g. riccardi and a.l. gorin. stochastic language adaptation over time and state in natural spoken dialog systems. ieee transactions on speech and audio processing, 8(1):3–10, 2000. g. riccardi and d. hakkani-tür. active and unsupervised learning for automatic speech recognition. in proceedings of eurospeech, geneva, switzerland, september 2003. a. rudnicky, c. bennett, a. black, a. chotomongcol, k. lenzo, a. oh, and r. singh. task and domain specific modelling in the carnegie mellon communicator system. in proceedings of the international conference on spoken language processing (icslp), 2000. e. schegloff and h. sacks. opening up closings. semiotica, 8:289–327, 1973. t. sheeder and j. balogh. say it like you mean it: priming for structure in caller responses to a spoken dialog system. international journal of speech technology, 6(2):103–111, 2003. j. shin, s. narayanan, l. gerber, a. kzemzadeh, and d. byrd. analysis of user behavior under error conditions in spoken dialogs. in proceedings of the international conference on spoken language processing (icslp), 2002. r. smith and s. gordon. effects of variable initiative on linguistic behavior in human-computer spoken natural language dialogue. computational linguistics, 23(1):141–168, 1997. s. stenchikova, b. mucha, s. hoffman, and a. stent. ravencalendar: a multimodal dialog system for managing a personal calendar. in proceedings of the human language technology conference (hlt), 2007. 30 a. stent, s. stenchikova, and m. marge. dialog systems for surveys: the rate-a-course system. in proceedings of the ieee/acl spoken language technology workshop (slt), 2006. s. stoyanchev. impact of responsive and directive adaptation on local dialog processing. phd thesis, stony brook university, 2009. s. stoyanchev and a. stent. lexical and syntactic priming and their impact in deployed spoken dialog systems. in proceedings of the meeting of the north american chapter of the association for computational linguistics (naacl), 2009. s. stoyanchev, d. hakkani-tür, and g. tur. name-aware speech recognition for interactive question answering. in proceedings of the international conference on acoustic, speech and signal processing (icassp), 2008. s. tomko and r. rosenfeld. shaping user input in speech graffiti: a first pass. in proceedings of the sigchi conference on human factors in computing systems, 2006. g. tur. extending boosting for large scale spoken language understanding. machine learning, 69 (1):55–74, 2007. g. tur and a. stolcke. unsupervised language model adaptation for meeting recognition. in proceedings of the international conference on acoustic, speech and signal processing (icassp), 2007. g. tur, d. hakkani-tür, and r. e. schapire. combining active and semi-supervised learning for spoken language understanding. speech communication, 45(2):171–186, 2005. m. walker, a. rudnicky, r. prasad, j. aberdeen, e. bratt, j. garofolo, h. hastie, a. le, b. pellom, a. potamianos, r. passonneau, s. roukos, g. sanders, s. seneff, and d. stallard. darpa communicator: cross-system results for the 2001 evaluation. in proceedings of the international conference on spoken language processing (icslp), 2002. w. ward and s. issar. recent improvements in the cmu spoken language understanding system. in proceedings of the human language technology conference (hlt), 1994a. w. ward and s. issar. integrating semantic constraints into the sphinx-ii recognition search. in proceedings of the international conference on acoustic, speech and signal processing (icassp), 1994b. i. witten and f. eibe. data mining: practical machine learning tools and techniques. morgan kaufmann, san francisco, 2nd edition, 2005. w. xu and a. rudnicky. language modeling for dialog system. in proceedings of the international conference on spoken language processing (icslp), 2000. s. young. detecting misrecognitions and out-of-vocabulary words. in proceedings of the international conference on acoustic, speech and signal processing (icassp), 1994. 31 dialogue & discourse 11(1) (2020) 89–121 doi: 10.5087/dad.2020.104 modelling structures for situated discourse nicholas asher nicholas.asher@irit.fr cnrs, aniti, irit toulouse julie hunter jhunter@linagora.com linagora, toulouse kate thompson cthompson@linagora.com linagora, toulouse editor: manfred stede submitted 01/2019; accepted 03/2020; published online 03/2020 abstract in this paper, we argue that modelling situated discourse requires not only allowing nonlinguistic events to enter into discourse relations with speech act contents, but also modelling semantic interactions between nonlinguistic events themselves. in an evolving nonlinguistic context, these interactions can give rise to a rich semantic, nonlinguistic structure that is relevant for the interpretation of conversational moves. examining how these nonlinguistic structures interact with structures determined by dialogue moves reveals new types of discourse structure and a novel perspective on discourse threads and goals. we motivate our arguments with a study of a corpus of situated multiparty chats developed for the stac project1 and annotated for discourse structure in the style of segmented discourse representation theory (sdrt; asher and lascarides, 2003). the stac corpus is not only a rich source of data on strategic conversation, but also the first corpus that we are aware of that provides discourse structures for multiparty dialogues situated within a virtual environment. the corpus was annotated in two stages: we initially annotated the chat moves only, but later decided to annotate interactions between the chat moves and non-linguistic events from the virtual environment. this two-step procedure allows us quantify various ways in which adding information from the nonlinguistic context affects dialogue structure.2 keywords: discourse structure, situated discourse, nonlinguistic context, nonlinguistic events 1. introduction the study of discourse structure, in particular rhetorical structure, on texts is now a well entrenched cottage industry in computational linguistics. several discourse annotated corpora exist including the penn discourse treebank (pdtb; prasad et al., 2008) and the rst discourse treebank (rstdt; carlson et al., 2002)—the only large corpus with full discourse structure for texts—as well as smaller annotated corpora such as discor (baldridge et al., 2007) or annodis (afantenos et al., 2012). the stac project (hunter et al., 2015b; asher et al., 2016) extends this work to dialogue 1. strategic conversation, erc grant n. 269427. 2. we gratefully acknowledge support from erc grant 269427 (strategic conversation), bpi france grant p169201 (linto-assistant vocal open-source respectueux des données personnelles pour l’entreprise), and 3ia grant aniti anr-19-pi3a-0004 (aniti – artificial and natural intelligence toulouse institute). we would also like to thank four anonymous reviewers and the editor for their extensive and very helpful comments. c©2020 nicholas asher, julie hunter, and kate thompson this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). asher, hunter and thompson annotation on a corpus of chats from an online version of the game settlers of catan. the stac corpus is, as far as we know, the only corpus with full relational, discourse structure for dialogues, and the only corpus with a complete set of full discourse annotations for situated dialogues. in this paper we characterize in a largely informal but precise way how the game events in our corpus interact with chat moves. to do so, we exploit two sets of annotations for our corpus, which we call the chat-only annotations and the situated annotations. the former contain discourse structures determined by only chat moves from the corpus, while the latter, which were begun only once the chat-only annotations were complete, contain full situated discourse structures that account for both game events and chat moves. in the situated annotations, both chat moves and game events contribute arguments to rhetorical relations, which allows us to account for the flexible dynamics of natural discourse situations. the two sets of annotations enable us to clarify and quantify the influence that a situated environment can have on a linguistic message and to make a detailed, empirical comparison of the structures in the chat-only and the situated annotations. as noted in tenbrink et al. (2013), little information is available concerning the systematic annotation of situated dialogues, and there are few annotated corpora. this paper goes towards filling these gaps. the analysis developed in this paper goes considerably beyond models of the nonlinguistic context in terms of deixis and reference, in which nonlinguistic entities are understood as crucial for the interpretation of discourse structures, but not for their construction (kaplan, 1989; rickheit and wachsmuth, 2006; kranstedt et al., 2004; kruijff et al., 2010). it also extends tenbrink et al. (2013)’s work on the navspace corpus, in which nonlinguistic actions are treated as contributing dialogue acts but the topic of their contributions to overall rhetorical structures is not broached. our aim in this paper is to show that in modelling situated discourse, we need to do more than allow nonlinguistic events to contribute arguments to discourse relations. we need to model structural relations between nonlinguistic events as they dynamically develop, because these relations are often relevant for interpreting dialogue moves, and we need to model the higher level structures that these relations give rise to. doing so, as we explain, uncovers new types of discourse structures and novel ways of thinking about discourse threads and discourse goals. in positing that nonlinguistic events contribute to larger discourse structures, this work echoes claims in lascarides and stone (2009), in which coverbal gestures can have discourse functions that are parasitic on verbal dialogue acts. we consider a much wider range of discourse interactions, however, and consider the nature of situated discourse in more general terms. we also examine a very different kind of interaction in which nonlinguistic events drive discourse development. we provide here strong empirical evidence for theoretical claims from hunter et al. (2018), but we go further by examining in detail the influence of the surrounding environment on discourse structure, describing in particular its effects on global discourse structures rather than local discourse relations alone. section 2 describes our corpus, the annotation process, and three points about multiparty discourse. section 3 argues that the game events in our corpus determine a rich structure of their own and that bringing this structure together with discourse structures determined by chat moves gives rise to new types of discourse structures. section 4 then quantifies the differences between the chat-only and situated annotations. this comparison gives a fuller picture of the role of nonlinguistic structures in the overall interpretation of our situated interactions. 90 comparing discourse structures between linguistic and situated messages 2. the settlers corpus in this section, we give an overview of the settlers corpus and certain features of the annotations. more quantitative details can be found in the appendix. section 2.1 explains the difference between the two sets of annotations, while section 2.2 lays out the general approach adopted to annotate the corpus. section 2.3 discusses three complications for building discourse structures in the chat-only annotations and our approach to them. these problems result from the multiparty setup and are not normally accounted for by theories of rhetorical structure. 2.1 building the corpus the settlers corpus consists of a series of chats taken from an online version of the game the settlers of catan that have been annotated for discourse structure in the style of segmented discourse representation theory or sdrt (asher, 1993; asher and lascarides, 2003). the settlers of catan is a multiparty, win-lose game in which players use resources such as wood and sheep to build roads, settlements, and cities on a game board. players acquire resources in various ways, including trading with other players and rolling the dice. as shown in figure 1, the game board is divided into hexes, each associated with a certain type of resource and a number between 2-6 or 8-12. a dice roll of, say, a 4 and a 2 gives any player with a building on a hex marked “6” one or more resources associated with that hex. rolling a 7 triggers a series of moves: the current player must move a game piece known as “the robber” to a hex of her choice and then steal a resource from a player with a building on that hex. the robber will stay on that hex until moved in another turn, and its presence will continue to impact the game by blocking resource distributions for the occupied hex. figure 1: a snapshot of the game board for the settlers of catan 91 asher, hunter and thompson to construct the settlers corpus, we modified an online, open-source version of catan to include a chat window. figure 1 illustrates the game interface and provides a snapshot of the way the game board looks to a particular player, simon. simon’s resources are shown in the upper lefthand rectangle, but as in the physical version of the game, simon cannot see the resources of his opponents. figure 1 shows that simon is preparing to make a trade via a trade panel: he has prepared an offer to the (red) player din, but has not yet clicked “register trade”. once simon registers his trade, din’s response, whether he accepts or rejects the trade, will be described in the game window, which records many game events that are public to all players. finally, the chat window allows players to chat, and prior chat is recorded in the history window. to encourage discussion, players were instructed to negotiate trades in the chat interface before executing an agreed trade through the trade panel. chat-only annotations. the settlers corpus was developed as a part of a project whose original aim was to study the discourse structure of strategic dialogue, in which interlocutors can have divergent discourse goals and therefore fail to be entirely cooperative. as such, the annotations were originally limited to the chat moves of the corpus, as it is during trade negotiations that strategic reasoning was assumed to affect discourse structure. in total we annotated the complete chat transcripts of 46 games using the glozz annotation tool (widlöcher and mathet, 2009) to build what we refer to as the chat-only annotations. to make the annotation process more manageable, we originally divided the chat history for each game into a number of (what we refer to as) dialogues, each of which involved one or more bargaining sessions. the criteria for dividing the dialogues were not precise, though there was the obvious goal of keeping each trading session intact. typically dialogues contained just one negotiation session with one player leading the bargaining. however, occasionally annotators linked elements in one negotiation session with elements in another, and in that case, we considered the linked sessions as contributing to a single dialogue. players used the chat interface to discuss numerous aspects of the game state, and it ultimately became clear during the first annotation campaign that much of the chat conversation was related to the game board and game events in intricate and semantically significant ways that merited further study. this observation triggered a second round of annotations in which the chats were re-annotated in light of publicly observable game events, descriptions of which were extracted from the game logs. we refer to the outcome of this annotation campaign as the situated annotations. situated annotations. figure 2 is an example from the situated annotations that illustrates how chat moves interact with game events (click for graph).3 every turn in our corpus, whether it is a chat move or a game event, is assigned a turn number, and all turns are automatically recorded and aligned in a game log for each game. turn numbers are indicated in the left column of figure 2. game messages that were added in a later stage for the situated annotations were assigned sequence identifiers in order to preserve the numbering of the chat and game events that were present in the first stage. (note that some game events were assigned numbers in the initial stage, although they were not shown to annotators.) each turn is also identified with an agent, as shown in the second column of figure 2. for chat moves, the agent is the player who typed the chat message (e.g., gwfs for turn 434);4 visually observable game events and states are either described in server messages, 3. in addition to the graphs provided in this paper, we add a link for each example that allows the viewer to see the example in a larger context and to compare the situated and chat-only graphs. go to https://www.irit.fr/stac/stac game graphs/index.html for instructions on how to read the graphs (see the “read me” document) and to view the graphs for every game in the annotated corpus. 4. small capitals indicate user names that have been abbreviated to save space or preserve anonymity. 92 https://www.irit.fr/stac/stac_game_graphs/s1_league1_game3/superdoc_17/dialogues.html https://www.irit.fr/stac/stac_game_graphs/index.html comparing discourse structures between linguistic and situated messages figure 2: the edges in our graphs are color coded for relation labels; e.g., all of the green arrows above represent result relations. many of which were visible to all players via the game window, or reconstructed (by our team) using information from the user interface (ui). in figure 2, william plays a monopoly card, which allows him to steal all instances of a resource of his choice that are possessed by the other players. in turn 433.0.4, he steals all of the wheat, and gwfs comments on this move in turn 434. ljay adds a smiley face in 440. there is some ambiguity as to whether the smile is a comment on the theft itself or on gwfs’ comment in 439; we opted for the latter interpretation, as the smile is added just after 439. in 441, ljay comments on the result of the theft: that william has 13 resources.5 2.2 annotating the corpus the chats in our settlers corpus were annotated for discourse structure in the style of sdrt, but there are numerous other theories of discourse structure for texts which might have been employed: rhetorical structure theory (rst; mann and thompson, 1987), the linguistic discourse model (ldm; polanyi et al., 2004), the discourse graphbank model (wolf and gibson, 2005), discourse lexicalized tree adjoining grammar (dltag; forbes et al., 2003), and the penn discourse treebank (pdtb; prasad et al., 2008). each one of these has, or at least can give rise to, an annotation model for discourse structure, and all of the theories agree on how the process of annotation for discourse structure should be approached. 5. for a more in-depth description of the corpus, see asher et al. (2016) and hunter et al. (2018), from which some of the foregoing description borrows. to see the webpage that was created to inform players about playing in our league, go to: http://homepages.inf.ed.ac.uk/mguhe/socl/. to view the annotations for the corpus, visit: https://www.irit.fr/stac/corpus.html. the annotation manual for the chat-only annotations can be found here: https://www.irit.fr/stac/stac-annotation-manual.pdf. 93 http://homepages.inf.ed.ac.uk/mguhe/socl/ https://www.irit.fr/stac/corpus.html https://www.irit.fr/stac/stac-annotation-manual.pdf asher, hunter and thompson the annotation process begins by segmenting each text or dialogue to be annotated into a set of what we will call elementary discourse units or edus, which serve as the basic building blocks of discourse structures. edus are typically clauses but may also include material that is in the periphery of the main predication in a clause (cf. afantenos et al., 2012). the next step in the process is to figure out how each edu should be related to other discourse units in the discourse representation. this requires solving two interrelated problems: the attachment problem and the labelling problem. the attachment problem concerns where each edu is attached in an incoming discourse structure. a typical case is that an edu attaches to another discourse unit as the argument of a rhetorical relation, but it can also, at least in sdrt, first be added to a group of discourse units that collectively provide the argument to a discourse relation.6 in this latter scenario, the discourse units that work together to provide a discourse argument form what is called a complex discourse unit or cdu in sdrt. the labelling problem then involves associating each discourse attachment with a label that denotes a discourse relation such as elaboration, explanation, narration, question answer pair, and so on. in a coherent conversation, each edu should serve a rhetorical function, which means that each edu, apart from the first one, should be attached to the incoming discourse via a rhetorical relation; in other words, the discourse should have a weakly connected structure. sdrt thus represents the structure of a discourse as a weakly connected graph with directed edges—for a relation instance explanation(edun,edum), for example, it is the content of edum that explains that of edun, not the other way around. sdrt further posits that graphs are acyclic, as every edu is expected to advance the conversation, not make a loop in which it is ultimately related to itself. in contrast to many other theories, including dltag, ldm, and rst, however, sdrt does not require that its graphs have a tree structure. to create our corpus, the chats were segmented automatically and manually corrected. because of the nature of chat, this was a relatively simple task. the annotation of discourse structure, on the other hand, was a large effort. the first stage was carried out by four “naive” annotators who received training over 22 negotiation dialogues, which included 560 turns in all. annotators were first asked to classify the segments as various kinds of dialogue acts. this involved labelling them according to surface form—question, command, or assertion—and to categories more specific to the game—-offer, counteroffer, accept (offer), refusal (offer), and other. these features proved to be important predictors for automatically learning discourse structure (afantenos et al., 2015; perret et al., 2016). annotators were then asked to choose an attachment point for each chat move except the very first one—so long as they were able to find an intuitive attachment point—and to label the attachments with at least one out of 16 possible labels for rhetorical relation types. further details on the annotation process can be found in the annotation manual (see footnote 5). after training, we evaluated our annotators on a small set of new dialogues in our pilot corpus. using an exact match criterion of success, the inter annotator agreement score was a kappa of 0.72 on attachment only on structures and 0.58 on attachment and labelling for doubly annotated dialogues—a respectable score given the complexity of handling both attachment and labelling together (asher et al., 2016). after these pilot annotations, four expert annotators continued the annotation process, and annotations of attachments and relations in the chat-only annotations went through a multi-step review process. the experts iterated passes over the annotations from the naive annotators as well as new 6. not all theories of discourse structure explicitly countenance the construction of larger units; in rst (mann and thompson, 1987), the step of assigning nuclearity values to different nodes could probably be so extended, but the possibility is not explicitly discussed. 94 comparing discourse structures between linguistic and situated messages figure 3: a truly non-treelike structure annotations, improving the data and debugging it by checking for violations of constraints, such as acyclicity, in the annotated structures. we had at least five stages of revision. we used the trello system (trello.com) to maintain our annotations, to keep track of the revisions, and to establish an agreement by consensus among four experts as to what should constitute our gold annotations. 2.3 three points about multiparty dialogue the multiparty nature of our corpora gave rise to various phenomena, some already noted in the literature on multiparty dialogue, that one finds rarely if at all in single authored text. we comment on three phenomena here: the presence of non-treelike discourse structures (cf. afantenos et al., 2015; perret et al., 2016; asher et al., 2016), overlapping conversational threads (cf. crystal, 2011; bartlett, 2014; afantenos et al., 2015; asher et al., 2016), and subjective interpretations (cf. lascarides and asher, 2009; ginzburg, 2012). the first two will be relevant to the characterization of our situated discourse structures in the sections that follow; the third we decided to ignore in our study for reasons that we explain below. a non-treelike structure is one containing a discourse unit or du with more than one incoming arrow, but within this category, we distinguish between quasi-treelike dus and truly non-treelike dus. a quasi-treelike du has two incoming arrows (with two different labels) from the same source du. for example, in uttering a sentence of the form “p but then q”, the speaker indicates that the content of q is linked to that of p by both a relation of contrast and a relation of sequence (contrast and narration in sdrt); the edu determined by q is thus quasi-treelike. a truly non-treelike du, on the other hand, has two or more incoming arrows from different source dus. turn 239 in figure 3 is an example of a truly non-treelike du (click for graph). in turn 234, gwfs makes an offer, to which he receives three negative replies (235, 236, 238). he responds in 239 with the acknowledgement “kk” (=“okay cool”), which is intuitively aimed at all three negative responses. the label qap is short for “question-answer pair” and types the connections between 234 and each of the three replies. truly non-treelike structures are fairly frequent in our corpora; out of a total of 12,588 dus in the chat-only annotations, 928 or about 7% were truly non-treelike. our annotation framework countenances both quasi and truly non-treelike structures, so this aspect of multiparty dialogue did not pose a problem for us, but we note that frameworks like that of rst (mann and thompson, 1987) and ldm (polanyi et al., 2004) do not countenance such structures. dialogue threads are created when groups of two or more dialogue participants engage in two or more connected exchanges (connected via instances of discourse relations) that are developed 95 https://www.irit.fr/stac/stac_game_graphs/s2_leaguem_game2/superdoc_10/dialogues.html asher, hunter and thompson simultaneously but independently in the sense that they are not discourse connected. figure 4 contains at least 3 threads, which we have represented with solid, dashed and dotted lines. note that the lines in this graph do not reflect rhetorical relations but only thread membership. to see the full graph with relations, click here. as figure 4 shows, the threads are developed simultanefigure 4: three discourse threads ously: for example, gwfs starts a new discussion in 167 that develops over seven segments, ending with 178b, and throughout this discussion, a sequence of trade negotiations initiated in 165—before gwfs started a new thread—develops and continues until well after 178b. the threads are independent in the sense that while two threads might share a common root (not pictured here), once they are created, there are no rhetorical relations that connect edus across threads. threads call for a revision of assumptions about the projectivity of discourse structure (like those assumed in rst; carlson and marcu, 2001) and salience-driven constraints on discourse attachment as formulated for monologue (polanyi, 1985; asher, 1993), as developing a new thread does not block the accessibility of edus from a previously initiated thread.7 such constraints do not even hold for individual speakers: gwfs, for example, is free to engage in all three threads simultaneously. in section 3 we will compare “standard” multi-party threads to similar structures generated in the situated annotations. the third phenomenon that we address in this section, which we chose to ignore in our annotations, is the possibility that interlocutors might interpret a discourse differently, leading to differing, and even contradictory, representations of a conversation. amplifying on lascarides and asher (2009); venant et al. (2014); venant and asher (2015) or on the more fundamental theoretical analyses of asher and paul (2018), we should create an individual discourse graph for each participant 7. for more on how such salience constraints are affected in our multiparty chats, see hunter et al. (2015a). see hunter et al. (2018) for a discussion of how to extend such constraints to situated discourse. 96 https://www.irit.fr/stac/stac_game_graphs/s1_league1_game3/superdoc_7/dialogues.html comparing discourse structures between linguistic and situated messages (cf. ginzburg, 2012). doing this for our corpora, however, would have required building three or four discourse graphs for each game, which would have precluded a more extensive, global view of discourse structures across a significant corpus. we therefore decided to annotate the multiparty dialogue from the perspective of a third party observer who infers a structure based on players’ public commitments to content derived from their contributions. this choice had various consequences for our annotation decisions. it meant, for instance, that we ignored personal server messages, including those displayed during trades and steals. when a player moves the robber and steals a resource from another player, for example, the history window will display a general message of the sort “william stole a resource from gwfs”. at the same time, gwfs and william will each receive a personal and more detailed message, such as “william stole a wheat resource from you” and “you stole a wheat resource from gwfs,” respectively. we disregarded all such personal messages. another effect of our decision to ignore subjectivity was that it left subjective contributions such as comments (linked by the comment relation) incomplete and possibly ambiguous. imagine that a player p utters a comment, say “thanks” or “sorry” or “oucho”, with content c. were we to have individual discourse graphs for each player, p’s graph would reflect commitment to c, while the other players’ graphs would be updated to reflect (at most) their commitment to p’s commitment to c. as we opted for a single graph, we had to choose between having all players commit to c, or having all players—including p—commit to the content that p commits to c.8 we chose the latter, weaker option, because it does not force us to ascribe attitudes to players without evidence that the players hold those attitudes. the downside, however, is that it yields an incomplete picture of players’ commitments, as it prevents us from capturing how other players respond to p’s utterance of c and whether they accept c. this in turn yields a less satisfying analysis of acknowledgments of comments, in which the content of a comment is established as common ground. using only one graph had an even more significant impact on the treatment of correction, and it led us to confine the use of correction almost entirely to self-corrections or corrections by the server of a player’s attempted but forbidden action (as a player in this case is forced to correct her action if game play is to continue). corrections between different players signal disagreements and thus leave room for differing points of view: a speaker who utters a correction with content q commits to the falsity of some content p, but the speaker of p might not be willing to accept the correction and hence might not commit to p’s being false. the closest that we could come to accurately representing such a disagreement between two different players without using individual graphs and commitment slates (portner, 2004) was to use a contrast relation that captures the fact that one player commits to p but another player commits to q. while this might be accurate, it, too, leads to incomplete representations: it misses the truth conditional effects of a correction on which the contents of the first term of the relation are taken, by at least one player, to be false. despite the incompleteness entailed by our choice to use a single discourse graph for each game in our corpus, it was the preferable option. in most cases, we simply lacked the information necessary to decide on each player’s commitment to each chat move. in other words, not only would building individual graphs have greatly complicated the annotation process, but in many cases, it would have failed to solve the incompleteness problem. the result would not have justified the cost. 8. the semantics of most discourse relations, including comment, is based on dynamic conjunction; this allows us to treat a comment as merely entailing common commitments to the content that the agent of the comment expresses a certain attitude. 97 asher, hunter and thompson 3. moving to situated dialogue to make the situated annotations, we posited that game events can contribute what we call elementary event units or eeus, whose contents serve as independent arguments to rhetorical relations just as the contents of edus or linguistic speech acts do (cf. also tenbrink et al., 2013). allowing eeus to contribute arguments to discourse relations, however, opens up the possibility that two eeus could be rhetorically related to each other. in fact, what we found in our corpus is that inferring such relations is often necessary for interpreting chat moves and the content of the overall game-chat exchanges. these relations between eeus in turn determine higher-level structures of their own. our central aim in this paper is to show the necessity of modelling these higher-level structures for (often nonlinguistic) events occurring in the larger situation and to explore how these structures interact with structures determined by conversational exchanges. this is what makes our task challenging: we are not simply trying to model the fact that a single nonlinguistic event can have an impact on discourse interpretation; we are trying to model interactions between two somewhat independent structures. accordingly, this section addresses two questions: (i) what types of relations and structures do we see holding of eeus in our corpus? (ii) what types of higher-level structures result from interactions between eeus and edus in our corpus? section 3.1 focuses on (i); section 3.2, on (ii). the setup of our corpus greatly facilitates study of these questions because a large portion of the game events are described in server messages that are temporally aligned with all of the chat moves. a major hurdle to studying the discursive contribution of nonlinguistic events is that the nonlinguistic context does not contain an analogue to a linguistic clause; nonlinguistic information is often presented in a steady stream, which means it is up to interpreters to determine how to individuate events. moreover, given that the same event can generally be described in different ways depending on one’s purposes, an interpreter must also determine what propositional content to associate with a nonlinguistic event. in other words, there is not only a more challenging segmentation problem for nonlinguistic information, but also a classification or conceptualization problem. because the game logs for our corpus contain descriptions of many game events, however, our corpus allows us to largely bypass both the segmentation and conceptualization problems for the game events. this frees us up to focus on how such events influence the structure and interpretation of the larger chat exchanges. the fact that the server provided descriptions of game events might suggest that eeus are not so nonlinguistic after all. hunter et al. (2018) address this concern in detail but we note here that some events were represented only visually to players. information about these events—which included ending of turns and selection of an addressee for a trade among others—had to be extracted from the user interface (ui) and assigned a content that could be subsequently annotated. a more general point, though, is that game events coming packaged with descriptions would be problematic if we were studying the conceptualization problem for nonlinguistic content. but in this empirical study we are not; we want to determine how the game events, once they are segmented and conceptualized, influence discourse construction and interpretation. for this it suffices that the interactions between discourse moves and game moves in our corpus mirror those that we could expect from people playing a nonvirtual version of settlers of catan. 98 comparing discourse structures between linguistic and situated messages 3.1 the conceptualization and structure of game events nonlinguistic events, discounting gestural events produced as a part of a conversation, differ from speech acts in the sense that they must be associated with contents that are true in the external world; eeus describe events that actually happened. as a result, the external world imposes a natural order over eeus: if an eeu εn appears before an eeu εm in a linear ordering of edus and eeus to be annotated for a situated discourse, then in general, we can infer that the event described in εn happened before that described by εm; at least, we know that it didn’t happen after it. in this, eeus differ considerably from edus, whose linear presentation can depart from the order inferred over the events described in chat moves: in bill cried because jane insulted him, for example, the event that is described first is understood to have happened second. the fact that the external world imposes its own temporal ordering over such nonlinguistic events might suggest that the structures that these events give rise to are more constrained than those for speech act events. this is to some extent true. relations like contrast, conditional, parallel, clarification-question, and question-elaboration, do not appear in our corpus between game events (though as indicated in table 9 in the appendix, some of them can relate eeus to edus). this is to be expected: it is not clear how an event that is not a speech act or some other communicative act (such as a gestural event) could convey a conditional dependency, a contrast or parallel, or follow-up or clarification question. in a multimodal corpus containing communicative nonlinguistic events, the possibilities would be different: a puzzled facial expression can convey a request for clarification, and with conventional gestures, a wide range of relations may be available. nevertheless, the structures formed by game events in our corpus are significantly richer than a mere sequential ordering. moreover, these structures can be influenced by chat moves—the chat moves providing “inverse information” in the sense of katagiri et al. (2006) and perry (1986) about the nonlinguistic context. sorting out the structural possibilities was thus an imposing task, and all the more so given the sheer size of the situated annotations: our choices about which nonlinguistic event types to consider yielded 31,811 eeus in the situated annotations, in contrast to 12,588 edus in the chat-only annotations. in what follows, we detail what we found, starting at the level of individual relations before moving on to larger structures. 3.1.1 relations relevant for the interpretation of the situated annotations many pairs of eeus were in fact linked by the sequence relation, where sequence(εn, εm) holds just in case the event described by εn took place before that described by εm. but eeus were also often linked by either result, which indicates that the first eeu describes the cause of the second, or continuation, which has the semantics of dynamic ‘&’ and simply requires that its arguments be true without commitment to temporal order. in figure 5 (click for graph), j’s comment in 207 is clearly a comment on 204, as it is about her dice roll, but it is also about 205 and, importantly, about the causal relationship between 204 and 205. her dice roll sucked precisely because it resulted in a resource distribution for her opponents while yielding nothing for her, so a result relation needs to be represented between 204, on the one hand, and the complex unit composed of the two segments for the two distribution events in 205, on the other. in addition, the two eeus representing the distribution events in 205 need to be related via continuation, as they occur simultaneously. eeus are also frequently arguments to question-answer pair (qap) and elaboration relations. figure 6 illustrates both cases (click for graph). because the addressee of a trade offer (e.g., gwfs in 51.1 of figure 6) was not identified in the server messages (e.g., 51) for our game, we 99 https://www.irit.fr/stac/stac_game_graphs/pilot14/superdoc_13/dialogues.html https://www.irit.fr/stac/stac_game_graphs/s1_league1_game5/superdoc_2/dialogues.html asher, hunter and thompson figure 5: result and continuation relations between eeus figure 6: elaboration and qap relations between eeus had to extract this information separately, yielding a separate turn that more fully specifies an offer. we decided to link these pairs of eeus via elaboration, whose semantics in sdrt require that the second argument specify properties of the first. thus in figure 6, we get elaboration(51, 51.1), and the fact that 51 and 51.1 work together to fully specify the offer is reflected by grouping them in a cdu. as hunter et al. (2015a) argue, a natural way of understanding the relation between the cdu [51, 51.1] and the subsequent trade in 52 is as a qap. the majority of linguistic offers in our corpus are in fact expressed as questions (see, for example, 165 in figure 4), but even nonlinguistic offers in effect present a pair of alternatives: to trade or not to trade. a trade like that described in 52 then functions as a “yes”, and a refusal or rejection via the trade panel functions as a “no”. table 1 shows the distribution of relation labels for links between eeus and/or cdus containing only eeus (i.e., with no edus at any level of constituency). 100 comparing discourse structures between linguistic and situated messages type situated eeus only question answer pair 929 continuation 8546 elaboration 547 result 12273 correction 118 background 45 sequence 5695 total 28153 table 1: relations between eeus in the situated annotations the few instances of background in the corpus are very systematic: each time a player wins the game, the server emits a message that reports the number of rounds that were played in the game and how long the game took. these messages were consistently attached with background to the message announcing that a player had won the game (for games that were played through the end). instances of correction involve cases in which a player tries to make an illicit move and the system blocks the move, such as when a player tries to make a trade but lacks the necessary resources. 3.1.2 complex structures relevant for the situated annotations a second way in which the structure over eeus in our situated annotations departs from a mere sequential linear ordering, aside from involving relations other than sequence, is that it was often natural or even necessary to group eeus together as cdus to capture the full content of a game. the exchanges in figures 5 and 6, discussed above, provide two examples. another case is when a player uses the trade interface to make successive offers to different players and then ends her turn after each of the offers results in a refusal. in such cases, annotators judged that the accumulation of refusals brought about the player’s decision to abandon her strategy and end her turn, and thus we grouped the series of failed offers in a cdu and related this cdu to the end turn move via result. we also decided to use cdus to group robber-related events, as players often commented on the complex events as a whole. as explained in section 2.1, a roll of a 7 in settlers of catan brings out a game piece known as the “robber” and triggers a complex series of events that includes at least moving the robber and choosing a player to steal from, and depending on the configuration of the game board at that time, possibly other events as well. figure 7 provides an example (click for graph). the decision to systematically represent robber events with cdus, in which each sub-event causes the next sub-event, yielded the following relation instances for the robber events in figure 7: result(278,[278.1,278.2]) and result(278.1,278.2). the situated annotations have a large number of cdus, many of them quite large: 5777 eeu-only cdus compared to 1450 edu-only cdus in the chat-only annotations. table 8 in the appendix provides more detail on the cdus in both the chat-only and situated annotations. while the larger discourse context can lead annotators to group eeus into complex units, influence goes in the other direction as well: giving annotators access to the game events revealed that certain chat moves were working together in semantically significant ways that were not obvious in the absence of the nonlinguistic context. turns 71-74 in figure 8, for example, were not grouped 101 https://www.irit.fr/stac/stac_game_graphs/s1_league1_game5/superdoc_18/dialogues.html asher, hunter and thompson figure 7: a typical robber sequence in a cdu in the chat-only annotations, but had to be grouped together in the situated annotations (click for graph). in turn 71, t.k. expresses interest in trading with other players, but in turns figure 8: an edu-only cdu introduced after consideration of game events 72-74, all of the other players reject his offer. in turn 75, t.k. goes to a port to get the wood he needs to build the road that he builds in 76. because a trade from a port is far more “expensive” than a trade with another player—trades with other players are usually 1:1 or 2:1 at most, while a trade with a port is 3:1—annotators judged that t.k. only pursued this trade because his prior trading attempt was unsuccessful. that is, they judged that the trade was a result of the entire failed negotiation in 71-74. 3.2 interactions between chat moves and game events in the situated annotations, links between chat moves and game events run in both directions: the comment “oucho” in figure 7 is a reaction to william’s stealing a resource in 280, but in figure 8, t.k.’s trade with the port is a reaction to a failed negotiation exchange. in this subsection, we take a look at what kinds of structures arise from interactions between chat moves and game events and how they relate to discourse threads and discourse goals. 102 https://www.irit.fr/stac/stac_game_graphs/s1_league1_game3/superdoc_4/dialogues.html comparing discourse structures between linguistic and situated messages 3.2.1 asymmetric and interleaved structures once we started to annotate the server and ui messages, the conception of how dialogues should be individuated changed: the idea of using bargaining sessions as a rough guideline gave way to a turn-based criterion so that the situated exchanges in our corpus are broken down by turns that begin when a player gets the dice and end when she ends her turn and passes the dice to the next player. there are cases in which we find links across turn boundaries, as when a player comments on a move from the immediately preceding turn, but in general, breaking down the interactions in the games by turns worked out well. figure 9, which represents the structure of an eeu only dialogue, gives an idea of what kind of structure a typical turn has (click for graph). jon’s getting the dice in 238.0.2 results in his rolling a 6 and a 2 in 239, which in turn results in the set of resource distributions detailed in turn 240 (grouped into a cdu). the ui then updates information about the resources of all four players, and jon ends his turn. figure 9: a standard eeu-only structure the central role of game development often leads to an asymmetric semantic dependence of chat moves on the game moves in the situated annotations. this observation leads to the characterization of new types of discourse structure. consider the example in figure 10. the example begins with dave getting the dice and ends with dave finishing his turn. from the first to the last move, there is a continuous succession of game events: dave rolls the dice, moves through a sequence of robber events, builds a road, tries to trade with the other players, buys a development card, and finally ends his turn. during all of this, he has a short exchange with tomm about the resource that he stole (in 96, 98-100, and 102). this exchange with tomm would not be interpretable without considering how the game is developing: we would be left to wonder what was unkind and what led to dave’s getting one of tomm’s resources. by contrast, were we to ignore this interaction, the development and interpretation of the game would remain completely intact and unchanged. in figure 10, we use a blue line to connect the moves central to game development and a magenta line to connect chat moves figuring in exchanges that asymmetrically depend on this development. the resulting subgraph indicated in blue is the backbone or what we will call the core of the structure in figure 10. the set of outlying nodes connected in magenta determine two peripheral structures, the second of which is the exchange between tomm and dave described above. clearly, ignoring the peripheral structures has no affect on the connectedness or the interpretation of the core, as the outlying nodes are all connected via outgoing edges, but ignoring the core leads to unconnectedness and uninterpretability for the peripheral structures. 103 https://www.irit.fr/stac/stac_game_graphs/pilot03/superdoc_9/dialogues.html asher, hunter and thompson figure 10: a dialogue and its asymmetric structure. the bold blue line traces the core, which runs through the eeus contributed by server and ui messages, including a cdu for the robber sequence in 94.0.1 and 95.0.1–95.0.3, and a cdu containing the chat moves 101 and 103-105 involved in a trade negotiation. the peripheral structures, indicated by magenta edges, consist of nodes which do not belong to the core and their incoming edges (click for graph). note that chat events that figure in trade negotiations are considered core moves in figure 10. the core of this structure therefore forms what we call an interleaved structure. an interleaved structure is set off not by its structural properties but by the types of its nodes: it is simply a multimodal graph that contains nodes for both edus and eeus. in this case, eeus feed into edus, but unlike in the asymmetric structures in our corpus, edus also feed into eeus. in interleaved structures, chat moves break into the intuitive structure of game events, so the latter cannot be seen as forming an autonomous linear structure of their own. this exposes a further type of complication for the idea that game events can be represented with a temporal linear ordering to which the linguistic context can simply be appended. moreover, interleaved structures can also figure in novel, non-treelike structures, as illustrated by figure 11 (click for graph). in turn 256, nelsen offers to trade wood for sheep or ore and both kersti and tyrant lord respond with interest (turns 257/261 and 260, respectively). their verbal acceptances result in two trade offers by nelsen, represented with two cdus, [262,262.0.1] and [264,264.0.1], whose members are related via elaboration. the cdu [264,264.0.1] is related to tyrant lord’s acceptance in 263 via sequence, 104 https://www.irit.fr/stac/stac_game_graphs/pilot01/superdoc_5/dialogues.html https://www.irit.fr/stac/stac_game_graphs/s2_league3_game1/superdoc_15/dialogues.html comparing discourse structures between linguistic and situated messages figure 11: an example of a truly non-treelike cdu [264,264.0.1], with edges from 261 and 263 but it also has an incoming arrow from 261 labelled as a result. the situated annotations add about 300 more non-treelike structures to the total for the chat-only annotations discussed in section 2.3. interactions between game events and chat moves also give rise to mixed cdus, containing both edus and eeus, that are interleaved into the larger game structure. one example is when a player reacts to a nonlinguistic trade offer with a chat move such as, “sorry, i don’t have wood” before rejecting the offer through the trade interface. formal definitions for cores, peripheries, interleaved and asymmetric structures are given in the appendix. in our two corpora, we have considered only maximal cores, also defined in the appendix, that start with the initial du of a dialogue and end with the last du with respect to the textual, or in our case time stamp, ordering. table 10 in the appendix gives statistical details concerning asymmetric and interleaved structures. 3.2.2 discourse threads and goals asymmetric structures arise anytime a discourse splits into two threads that branch off of the same node.9 this is a commonplace occurrence at dinner parties, for instance, when a group is chatting and then two people in the group decide to continue the conversation in slightly different directions, leading to a (perhaps temporary) split of the group into subgroups that each follow a different branch of the preceding conversation. in such a case, we could in principle choose either continuation of 9. for the purposes of this paper, we define asymmetric structures as involving two threads that never rejoin (which means that truly non-treelike structures do not count as asymmetric structures). this is a delicate point: in hunter et al. (2018), we discuss an example that violates this assumption, and it might be that in face-to-face conversations speakers can use expressions such as “we were just discussing the same thing” to bring together two groups of conversationalists and two conversation threads. we suspect, however, that these violations are very limited and generally need to be marked explicitly, so we do not think that our simplified definition is problematic for our present purposes. 105 asher, hunter and thompson the conversation as defining the core of the discourse structure, yielding two possible asymmetric structures; the notion of a core is a thematic and functional one. the graphs in our situated annotations reveal threads of a more constrained form. first, if we look at the graph for any one of our games, we find one thread that runs through the entire game—the thread that contains the game events and perhaps some interleaved chat moves. second, the threads that make up the periphery contain typically 2.28 nodes on average, with a range of 50 nodes (see table 10 for more details), whereas a split in a regular conversation can lead to an extended discussion.10 third, the structures in our peripheries contain no outgoing links into the core. this is a part of what it is to be a peripheral structure, but what is interesting about our corpus is that there are so many “bushes” of structure that make up the periphery for each game. finally, while the interactions that make up the peripheries in our situated annotations are less extended than full conversations, they are longer than the peripheral structures that we find in single-authored text. appositive relative clauses, for example, generally contribute peripheral structures of one to two discourse units that attach to the main discourse via a relation such as background, but do not play a central role in the progression of a discourse (venant et al., 2013). these structural characteristics correlate with interesting features of the interactions in our corpus. the continuity of the series of game events as well as the brevity of the interactions in the periphery and the tendency of the players to return to focus on core events reflects the fact that playing and winning the game is clearly the leading goal and one that drives the interaction in the corpus; linguistic exchanges are secondary to this goal. furthermore, if we look at the content of the chat exchanges and how they relate to the core, we find that they are largely reactive, being frequently attached to eeus via relations such as comment (see table 2 for details). these features correlate with the limited size of the peripheral structures and the fact that the structures do not feed back into the core—it is unsurprising that a commentary would fail to divert attention entirely away from the game and that it would be inert with regard to game development. other types of threads might have a different discourse function that would be reflected through different structural features. a clarification request, for example, might temporarily take conversational participants away from a question that has been asked, but then feed back into it by having a direct impact on the nature of the answer that will ultimately be given to the main question (cf. ginzburg, 2012). while the goal of this paper is not to develop an account of exactly how the shape of a discourse structure or the distribution of various discourse relations relates to the nature of different discourse goals, we emphasize that there is an important connection between these topics that would be worthy of exploring in future work. to give another brief example, gwfs was the player who ultimately won the competition that we set up to build our corpus, and we suspect that part of his strategy for winning was to introduce conversational threads that were less closely related to the main events of the game than most of the other peripheral exchanges, especially those led by other players. it was almost as though he was trying to distract his competition. table 3 shows that gwfs initiated more peripheral structures than other players. figure 4, provides an illustration of gwfs’s behavior: ljay 10. conversational threads in face-to-face conversation often subdivide interlocutors into mutually exclusive groups. in our games, by contrast, two different threads can involve the same set of speakers. this is in part due to the task-based setup—even in face-to-face interactions, people can chat while also coordinating on a task that might require occasional discussion. it is also in part due to the chat environment—two people can easily carry on two conversational threads even in the absence of a task that they are jointly performing. it would be interesting in future work to see how different kinds of threads associate with different constraints on group membership, but we cannot go into that topic here. 106 comparing discourse structures between linguistic and situated messages relation type # relation type # comment 1262 continuation 40 question answer pair 514 result 38 acknowledgement 355 parallel 29 elaboration 134 background 25 clarification question 130 sequence 15 explanation 108 correction 5 contrast 58 narration 3 q elab 57 alternation 1 conditional 1 table 2: breakdown of relations that connect asymmetric structures to a core in the situated annotations initiates a trade negotiation, and gwfs replies to ljay but then asks, “so how do people know about the league?,” setting off a discussion that is independent of the game at hand. in the chat-only and situated annotations, we related gwfs’s question to his reply to ljay via background, but this is arguably unsatisfying. his question is not directly related to ljay’s attempt to trade or even to the particular game that they are playing. intuitively, he is “popping” up to a much larger, implicit topic that involves the information that they are playing this game and the other games as a part of a league (a pop that is signalled by his use of “so”). this is not an issue that we attempted to tackle when building our corpus, nor did it seem to us to be frequent enough to pose a significant problem for our annotation approach, but the hypothesis that gwfs’s strategy would be reflected by the way that his contributions influence the overall shape of the discourse would be an interesting topic for future investigation, and a potential point of contact with work on implicit topics and questions under discussion (ginzburg, 2012; roberts, 2012). emitter # emitter # gwfs 351 gramos 71 inca 164 raefbrisbin 67 ljay 116 dmm 62 ztime 102 t.k. 61 skinnylinny 98 zorburt 57 sabercat 98 somdechn 51 william 81 shawnus 51 cardlinger 80 nelsen 51 tomm 51 table 3: number of peripheral structures started per speaker (cutoff at 50) 107 asher, hunter and thompson 4. preservation of linguistic structure we have emphasized that modelling the contribution of nonlinguistic events to a larger discursive interaction involves modelling the interaction between different structures. not only does the nonlinguistic context do more than add information via grounding or reference, it does more than merely contribute contents that function like dialogue acts. we must consider what kinds of structures nonlinguistic events give rise to in their own right and how these structures interact with linguistic structures. section 3.2.1 took on the first task and showed how the game events in our corpus form a rich structure of their own, while section 3.2.2 tackled the second task and revealed new kinds of situated discourse structures that reveal information about how the different edus and eeus involved contribute to discourse goals. section 4 aims to flesh out the claim about the importance of nonlinguistic structure by comparing the effects of adding game events in the situated annotations at the level not only of individual events and relation instances, but also at the structural level. while preservation of relation instances shows that our decisions about these instances were relatively robust even when limited to information present only in the linguistic turns, the figures for structural preservation show conclusively that we cannot take the situated annotations to be a conservative extension of the chat-only annotations. this is a familiar point when performing an error analysis on the output of a learning program for discourse structure: a respectable f1 score at the level of instances might not mean much in terms of structure preservation; the predicted structures might still be useless as discourse structures, as argued in ferracane et al. (2019). this comparison is made possible by the fact that the chat-only annotations in our corpus were completed in their entirety before we began the situated annotations and by the fact that the data set used for the situated annotations is a minimal extension of that used for the chat-only ones. that is, the two data sets differ only in that one includes game events and the other does not, enabling us to identify and circumscribe the semantic contribution of nonlinguistic events at the discourse level. the data sets thereby form something like a minimal pair, as illustrated by (1) and (2): (1) louise read any book. (2) louise read every book. minimal pairs are a useful tool in semantics for characterizing the semantic contribution of a certain type of expression. (1) and (2), for example, are useful for studying the behavior of polarity items, and their semantic value is only helped, not hindered, by the ungrammaticality of (1). while our pair of data sets differs importantly from minimal pairs used in formal semantics, as pointed out by one reviewer, a similar point can be made: rather than making it semantically irrelevant, the lack of game events in our chat-only data set renders it invaluable for our purposes. all discourses take place in a larger context involving shared knowledge and presuppositions, if not a shared visual scene, and no corpus can account for all contextual information. even our final corpus is impoverished in the sense that it does not include information on most game states because we could not find a way to extract information about them from the ui in a concise way that also would reflect their persistence rather than make them look like punctual events. what is important is that a corpus be adequate for the task for which it is employed. in our case, we do not need the most completely annotated corpus possible; we need one that characterizes methodically the influence of information about game events on the construction and interpretation of discourse structures for the chat moves. our two data sets and their associated annotations are perfectly tailored for this task. 108 comparing discourse structures between linguistic and situated messages we begin in section 4.1 with a simple comparison of the chat-only annotations and the situated annotations to show how much information was added in the extension and how many relations were preserved. section 4.2 then looks at how structures were preserved—or not—as we moved from the chat-only annotations to the situated annotations. 4.1 preservation of relations both naive and expert annotators for the chat-only annotations attempted to relate each noninitial chat move to another part of the chat in order to capture the discursive function of the former. when they were unable to find a reasonable attachment point, the result was an “orphan” linguistic move; that is, a move with no incoming link. for example, gwfs’s comment “oucho” in figure 7, repeated below as figure 12, was left as an orphan in the chat-only annotations, but once the game events in figure 12 were included it became clear that turn 279 was a reaction to and a commentary on william’s stealing a resource in 278.2 (click for graph). figure 12: turn 279 was an orphan in the chat-only annotations in all, the chat-only annotations contained 1501 orphans, all of which were assigned incoming links after we introduced eeus into our annotations.11 this led to a significant difference between the number of semantic links between the two sets of annotations: while the chat-only annotations contain 12,271 links, the situated annotations contain 3,591 links that relate an edu (or a cdu containing at least one edu) with an eeu (or a cdu containing at least one eeu). in many cases, adding information about game events not only led to new relations, but it also led annotators to revise judgments about relations between chat moves. table 4 shows how many relations from the chat-only annotations persisted when we moved to the situated annotations. as indicated in the table, around 20% of the relations in the chat-only annotations changed or disappeared in the situated annotations. in addition, relations were added between chat moves in the situated annotations, so that the persisting chat-only annotations made up 72% of the relations between chat moves in the situated annotations. the latter thus constitute a non-negligeable revision of the chat-only annotations. still, one might argue that linguistic structure was largely preserved; almost three quarters of the chat-only annotations were appropriate for the final corpus. had we been 11. as explained in section 2.1, the chats for the corpus were subdivided into dialogues (individuated roughly by player turns) in order to make annotation more tractable. while annotators were instructed to relate each non-initial dialogue chat move to another (dialogue internal or external) move where possible, they were also allowed to relate a dialogue initial chat move to another move if they felt it was the right thing to do. this instruction probably led to a bias towards finding antecedents for non-initial dialogue moves that was absent for initial moves. the number of orphans given above includes every dialogue-initial chat move of a non-chat-initial dialogue. if we consider only non-initial dialogue orphans, the total is 400. 109 https://www.irit.fr/stac/stac_game_graphs/s1_league1_game5/superdoc_18/dialogues.html asher, hunter and thompson relations between chat moves (i.e., edus and cdus composed only of edus) # linguistic games 12271 situated games 13605 persisting 10206 strictly persisting (same relation type) 9829 table 4: persisting relations gauging the success of a machine learning algorithm for learning discourse structure, this precision score would have counted as a success. figure 13 illustrates the difference between measuring preservation at the level of relation instances and the level of structures. the graph on the left is from the chat-only annotations and the graph on the right, from the situated annotations (click for full graphs).12 four out of five links from the left graph are preserved in the right graph; the only link that was removed in the shift to the situated annotations is a continuation link between edus 425, the top edu in the linguistic graph and edu 427, represented in dark green. however, this link does not assign a real function to edu 425, as it is naturally understood as a commentary on something, but not edu 427. in the situated annotation, 425 comments on a cdu of game events or states (represented by the second red node) in the incoming context, and not 427. and this comment serves to explain why the author of 425 makes the offer he does in 427. this change breaks up a connected structure in the linguistic graph into two pieces in the situated annotation, and the two sub-structures figure in significantly different parts of the overall game structure. we also get information about the temporal relation between the top two nodes in the left graph and the bottom four: the blue arrow along the right side of the right graph represents sequence, which imposes a temporal order on its arguments. this one small change in fact changes the structure and content of the conversation substantially. 4.2 preservation of substructures to measure structure preservation across annotations, we check for elementary embeddings of chatonly dialogue structures into the situated structures of the corresponding dialogues, where an elementary embedding preserves all relations, functions, and designated objects in a (chang and keisler, 1973). more precisely: definition 1 an elementary embedding is a one-to-one function f from the domain a of a structure a to the domain b of a structure b such that for any relation r, function g and designated object a in the signature of a, a |= r(b, c) iff b |= r(f(b), f(c)); a |= g(b) = c iff b |= g(f(b)) = f(c), and f(a) = b, where b is the designated object of b. when such an embedding f exists, f(a) is an elementary substructure of b. in fact, in our situation, there is a canonical embedding f that is the identity function on discourse units in the linguistic structure; so to say that f(a) is an elementary substructure of b is just to say that a is an elementary substructure of b. 12. in figure 13, we adopt the more minimalist format used for our clickable graphs to facilitate a visual comparison of the two structures under discussion. the more detailed graphs used throughout the rest of the paper would have been too large to place side by side. 110 https://www.irit.fr/stac/stac_game_graphs/s2_league4_game2/superdoc_26/dialogues.html comparing discourse structures between linguistic and situated messages figure 13: all but one of the relations in the chat-only structure (left) are preserved in the situated structure (right), but the former is broken into two pieces when we move to the situated annotations our discourse graphs are first-order structures with a designated object—in other words they are pointed models, as can be seen from definition 2 for sdrt graphs. definition 2 a discourse graph g is a tuple (v,e1, e2, `,last), where v is a set of edus and cdus; e1, a set of edges in v 2 representing discourse attachments; e2, a set of edges that relate each cdu to its members; ` : e1 → p(rl), a function that labels the discourse attachments from e1 with sets of discourse relation labels taken from rl; and last , a label for the last unit in v relative to textual order. to check for the relevant elementary embeddings from our chat-only structures into situated structures, we ignore the element last from definition 2 so that we consider embeddings only on quadruples (v,e1, e2, `). last is used to define constraints on how a discourse can dynamically evolve, but in checking the preservation on a finished structure, we can (and must) ignore it. let l be the weakly connected linguistic substructure for a dialogue d (henceforward our structures will pertain to whole dialogues), and let s be the situated structure for d that includes all the nodes of l. we say that l is elementarily preserved in s just in case l is an elementary substructure of s. where major changes to the chat-only annotations occur in the situated annotation, l will not survive as an elementary substructure of s . our results show that substructure preservation is a much stronger condition on persistence of information than the preservation of relation instances; we get 72% preservation of relation instances, but only 496 out of the 1137 dialogue structures in the chat-only annotations, or a little over 43%, are preserved as elementary substructures of the corresponding s structures. 4.3 preservation of subtypes building on discussions of discourse-central or at-issue content (potts, 2005; roberts, 2012; simons et al., 2010), much current work in formal semantics and pragmatics asks how the use of a certain linguistic form or expression indicates the role of a particular discourse move in achieving 111 asher, hunter and thompson discourse goals. the results from our corpus study add a new dimension to this discussion because the information that is driving discourse development and the achievement of discourse goals in our corpus can often not be predicted at the level of individual linguistic moves. it is only by looking at the game-chat structures as a whole and the way that substructures like cores and peripheries function within these larger structures that we can see how individual moves inside of those structures address discourse goals. with this in mind, it is important to consider how subtypes of discourse structures—e.g., core and periphery types—are preserved between our two sets of annotations. when a core structure is preserved by an s embedding, this indicates that while the linguistic features of individual chat moves might not have been sufficient to reflect the role of these moves in achieving discourse goals, they were at least helpful in determining the role of these moves in linguistic substructures whose relevance to discourse development and goals was clear. linguistic core structures that remain core structures in the situated annotations form an integral part of the game and address the main point of the overall discourse, which is to play and win the game. on the other hand, preservation of a peripheral structure indicates that linguistic information at the level of substructures was sufficient for determining the discursive functions of the individual moves in these substructures and that these moves were not integral to the main point. more formally, we say that an l substructure a of type τ is τ preserved under the canonical s embedding f into an s structure b just in case a is a substructure of a structure of type τ in b. for example, if we consider the structure a for a core of a dialogue d in the chat-only annotations, then we say that a is core-preserved under an s embedding f just in case a is a substructure of the core of b. when there is an s embedding of an l structure a, which may contain both core and peripherytype structures, such that all substructures of a maintain their type under the embedding, we call this perfect type preservation. perfect type preservation strictly entails elementary preservation, core-type and periphery-type preservation, while core-type and periphery-type preservation together entail pefect type preservation. in general, perfect type preservation is rare in our corpus. tables 11 and 12 in the appendix provide the details, but we summarize the main results here. out of 296 core-only l structures, which were in some sense the simplest case, only 23% had an s embedding that was core-type preserving (and hence perfect-type preserving). type preservation results were much worse for asymmetric structures; only 4% exhibited perfect type preservation; 7% of all l structures exhibited periphery-type preservation and 11,5% exhibited core-type preservation embedding. the general moral we draw from this study is that at least in our situated annotations, information relevant to the main goals of the conversational participants is not reliably signalled by linguistic means, either at the level of individual speech acts or at the level of relations. in fact our discussion in section 3.2.2 suggests that linguistic chat in our corpus may reflect a secondary goal to distract other players from their main goals, something that the subtype preservation results support at least indirectly. in any case, elementary and subtype preservation will prove, we think, useful tools for analyzing the interplay between player goals, strategies and linguistic and nonlinguistic actions. 5. conclusions in this paper, we have surveyed and compared two sets of annotations of a corpus culled from an online version of the game the settlers of catan: the chat-only annotations, which take only chat moves into account, and the situated annotations, which also include descriptions of visually presented events from the games during which the players were chatting. considering the nature 112 comparing discourse structures between linguistic and situated messages of the situated annotations and how they compare with the chat-only annotations allowed us to illuminate and measure a variety of ways in which information from the nonlinguistic context can influence the content and structure of a discourse. our study provides new data and statistics to substantiate claims from hunter et al. (2018) that modelling discourse situated in a dynamically evolving nonlinguistic context requires attributing a rich structure to that context, and that this structure can interact with purely linguistic structures in new and interesting ways. the nature and extent of nonlinguistic influence evidenced in the comparison of our chat-only and situated annotations supports the claim that nonlinguistic information is relevant for far more than reference or domain restriction; nonlinguistic events can contribute entire propositions to discourse content without being picked out by a linguistic expression or deictic act. but we have also showed that we must go further than allowing nonlinguistic events to contribute arguments that link them rhetorically to the contents of speech acts: the structural and semantic relations that link nonlinguistic events to each other can also influence a linguistic message. what’s more, as we showed in section 3, looking at how the nonlinguistic and linguistic structures interact can give us a more complete understanding of what is driving discourse development and how speech acts and nonlinguistic events contribute to achieving overall discourse goals. finally, section 4 reinforced the importance of considering nonlinguistic structures by showing that when we consider structure, rather than merely relation instances, we get a different and more complete picture of how the game events in our corpus influenced the overall content of our annotations. future work will involve generalizing the results described in this paper to other types of corpora. the chats in our corpora are mostly directed towards the competitive task at hand (winning the game), the nonlinguistic events that show up in our discourse structures are highly standardized due to game rules, and the virtual environment leads to a relatively controlled and simplified nonlinguistic context. these features simplified the annotation of relations between the chats and the game state, which allowed us to carry out the comparisons described in this paper. other situated conversations might take place in a much more complex and varied environment. still, we believe that the basic points that we have used our corpora to make about situated discourse are largely general. in any case, at this point in empirical research on situated discourse, we suspect that a certain amount of standardization is a necessary feature of events that convey discourse functions beyond causal or sequential discourse relations, functions like answering or posing a question, for example. we would like to pursue this question in future research. references stergos afantenos, nicholas asher, farah benamara, myriam bras, cécile fabre, lydia-mai hodac, anne le draoulec, philippe muller, marie-paule péry-woodley, laurent prévot, josette rebeyrolle, ludovic tanguy, marianne vergez-couret, and laure vieu. an empirical resource for discovering cognitive principles of discourse organisation: the annodis corpus. in the eight international conference on language resources and evaluation, pages 2727– 2734, istanbul, turkey, may 2012. european language resources association (elra). url http://www.lrec-conf.org/proceedings/lrec2012/index.html. stergos afantenos, eric kow, nicholas asher, and jérémy perret. discourse parsing for multi-party chat dialogues. in the 2015 conference on empirical methods in natural language processing, pages 928–937, lisbon, portugal, september 2015. association for computational linguistics. url http://aclweb.org/anthology/d15-1109. 113 http://www.lrec-conf.org/proceedings/lrec2012/index.html http://aclweb.org/anthology/d15-1109 asher, hunter and thompson nicholas asher. reference to abstract objects in discourse. kluwer academic publishers, 1993. nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. nicholas asher and soumya paul. strategic conversation under imperfect information: epistemic message exchange games. logic, language and information, 27(4):343–385, 2018. nicholas asher, julie hunter, mathieu morey, farah benamara, and stergos afantenos. discourse structure and dialogue acts in multiparty dialogue: the stac corpus. in the tenth international conference on language resources and evaluation, pages 2721–2727, portorož, slovenia, may 2016. european language resources association (elra). url http://www.juliejhunter.com/ uploads/3/9/6/1/39617901/lrec2016 final.pdf. jason baldridge, nicholas asher, and julie hunter. annotation for and robust parsing of discourse structure on unrestricted texts. zeitschrift fur sprachwissenschaft, 26:213–239, 2007. tom bartlett. analyzing power in language: a practical guide. routledge, london, 2014. lynn carlson and daniel marcu. discourse tagging reference manual. isi technical report isi-tr545, 54:56, 2001. lynn carlson, mary ellen okurowski, and daniel marcu. rst discourse treebank ldc2002t07. linguistic data consortium, university of pennsylvania, 2002. c.c. chang and h. jerome keisler. model theory. north holland publishing, 1973. david crystal. internet linguistics. routledge, london, 2011. elisa ferracane, greg durrett, junyi jessy li, and katrin erk. evaluating discourse in structured text representations. in the 57th annual meeting of the association for computational linguistics, pages 646–653, florence, italy, july 2019. association for computational linguistics. url https://www.aclweb.org/anthology/p19-1062. katherine forbes, eleni miltsakaki, rashmi prasad, anoop sarkar, aravind k. joshi, and bonnie l. webber. d-ltag system: discourse parsing with a lexicalized tree-adjoining grammar. journal of logic, language and information, 12(3):261–279, 2003. jonathan ginzburg. the interactive stance: meaning for conversation. oxford university press, oxford, 2012. julie hunter, nicholas asher, eric kow, jérémy perret, and stergos afantenos. defining the right frontier in multi-party dialogue. in the semantics and pragmatics of dialogue (godial), pages 95–103, gothenburg, sweden, august 2015a. url https://flov.gu.se/digitalassets/1537/ 1537599 semdial2015 godial proceedings.pdf. julie hunter, nicholas asher, and alex lascarides. integrating non-linguistic events into discourse structure. in the 11th international conference on computational semantics (iwcs), pages 184–194, london, england, april 2015b. url http://www.aclweb.org/anthology/w15-0123. julie hunter, nicholas asher, and alex lascarides. a formal semantics for situated conversation. semantics & pragmatics, 11(10), 2018. url https://doi.org/10.3765/sp.11.10. early access. 114 http://www.juliejhunter.com/uploads/3/9/6/1/39617901/lrec2016_final.pdf http://www.juliejhunter.com/uploads/3/9/6/1/39617901/lrec2016_final.pdf https://www.aclweb.org/anthology/p19-1062 https://flov.gu.se/digitalassets/1537/1537599_semdial2015_godial_proceedings.pdf https://flov.gu.se/digitalassets/1537/1537599_semdial2015_godial_proceedings.pdf http://www.aclweb.org/anthology/w15-0123 https://doi.org/10.3765/sp.11.10 comparing discourse structures between linguistic and situated messages david kaplan. demonstratives. in j. almog, j. perry, and h. wettstein, editors, themes from kaplan. oxford, 1989. yasuhiro katagiri, mayumi bono, and noriko suzuki. conversational inverse information for context-based retrieval of personal experiences. in jsai’05: proceedings of the 2005 international conference on new frontiers in artificial intelligence, pages 365–376. jsai, march 2006. alfred kranstedt, peter kühnlein, and ipke wachsmuth. deixis in multimodal human computer interaction: an interdisciplinary approach. in a. camurri and g. volpe, editors, gesture-based communication in human-computer interaction, pages 112–123. springer, 2004. geert-jan m. kruijff, pierre lison, trevor benjamin, henrik jacobsson, hendrik zender, ivana kruijff-korbayová, and nick hawes. situated dialogue processing for human-robot interaction. in h. i. christensen, g. m. kruijff, and j. l. wyatt, editors, cognitive systems, pages 311–364. springer berlin heidelberg, berlin, heidelberg, 2010. url http://dx.doi.org/10.1007/ 978-3-642-11694-0 8. alex lascarides and nicholas asher. agreement, disputes and commitment in dialogue. journal of semantics, 26(2):109–158, 2009. alex lascarides and matthew stone. a formal semantic analysis of gesture. journal of semantics, 26(4):393–449, 2009. william c. mann and sandra a. thompson. rhetorical structure theory: a framework for the analysis of texts. international pragmatics association papers in pragmatics, 1:79–105, 1987. jérémy perret, stergos afantenos, nicholas asher, and mathieu morey. integer linear programming for discourse parsing. in the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 99– 109, san diego, california, june 2016. association for computational linguistics. url http://www.aclweb.org/anthology/n16-1013. john perry. perception, action, and the structure of believing. in r. grandy and r. warner, editors, philosophical grounds of rationality, pages 333–361. oxford, 1986. livia polanyi. a theory of discourse structure and discourse coherence. in w. h. eilfort, p. d. kroeber, and k. l. peterson, editors, papers from the general session at the 21st regional meeting of the chicago linguistics society. chicago linguistics society, 1985. livia polanyi, chris culy, martin van den berg, gian lorenzo thione, and david ahn. a rule based approach to discourse parsing. in the 5th sigdial workshop on discourse and dialogue, pages 108–117, cambridge, massachusetts, usa, april 30 may 1 2004. association for computational linguistics. url http://www.aclweb.org/anthology/w04-2322. paul portner. the semantics of imperatives within a theory of clause types. in semantics and linguistic theory 14, pages 235–252, evanston, illinois, u.s.a., may 2004. url https: //journals.linguisticsociety.org/proceedings/index.php/salt/article/view/2907. christopher potts. the logic of conventional implicatures. number 7. oxford university press on demand, 2005. 115 http://dx.doi.org/10.1007/978-3-642-11694-0_8 http://dx.doi.org/10.1007/978-3-642-11694-0_8 http://www.aclweb.org/anthology/n16-1013 http://www.aclweb.org/anthology/w04-2322 https://journals.linguisticsociety.org/proceedings/index.php/salt/article/view/2907 https://journals.linguisticsociety.org/proceedings/index.php/salt/article/view/2907 asher, hunter and thompson rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in the sixth international conference on language resources and evaluation, pages 2961 – 2968, marrakech, morocco, may 2008. elra. url http://www.lrec-conf.org/proceedings/lrec2008/pdf/754 paper.pdf. gert rickheit and ipke wachsmuth. situated communication. walter de gruyter, 2006. craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. semantics and pragmatics, 5(6):1–69, 2012. mandy simons, judith tonhauser, david beaver, and craige roberts. what projects and why. in semantics and linguistic theory 20, pages 309–327, vancouver, british columbia, canada, april 29 may 1 2010. url https://journals.linguisticsociety.org/proceedings/index.php/salt/ article/view/2584. thora tenbrink, kathleen eberhard, hui shi, sandra kuebler, and matthias scheutz. annotation of negotiation processes in joint-action dialogues. dialogue and discourse, 4(2):185–214, 2013. antoine venant and nicholas asher. dynamics of public commitments in dialogue. in the 11th international conference on computational semantics, pages 272–282, london, england, 2015. association for computational linguistics. url http://www.aclweb.org/anthology/w/ w15/w15-0131. antoine venant, nicholas asher, philippe muller, pascal denis, and stergos d. afantenos. expressivity and comparison of models of discourse structure. in proceedings of sigdial 2013, pages 2–11, metz, france, august 2013. url https://www.aclweb.org/anthology/w13-40.pdf. antoine venant, nicholas asher, and cedric degremont. credibility and its attacks. in the 18th workshop on the semantics and pragmatics of dialogue, pages 154–162, edinburgh, scotland, september 2014. url http://www.macs.hw.ac.uk/interactionlab/semdial/semdial14.pdf. antoine widlöcher and yann mathet. la plate–forme glozz : environnement d’annotation et d’exploration de corpus. in la 16e conférence traitement automatique des langues naturelles, senlis, france, june 2009. atala, lipn. url http://talnarchives.atala.org/taln/taln-2009/ taln-2009-court-023.pdf. florian wolf and edward gibson. representing discourse coherence: a corpus based study. computational linguistics, 31(2):249–287, 2005. 116 http://www.lrec-conf.org/proceedings/lrec2008/pdf/754_paper.pdf https://journals.linguisticsociety.org/proceedings/index.php/salt/article/view/2584 https://journals.linguisticsociety.org/proceedings/index.php/salt/article/view/2584 http://www.aclweb.org/anthology/w/w15/w15-0131 http://www.aclweb.org/anthology/w/w15/w15-0131 https://www.aclweb.org/anthology/w13-40.pdf http://www.macs.hw.ac.uk/interactionlab/semdial/semdial14.pdf http://talnarchives.atala.org/taln/taln-2009/taln-2009-court-023.pdf http://talnarchives.atala.org/taln/taln-2009/taln-2009-court-023.pdf comparing discourse structures between linguistic and situated messages 6. appendix 6.1 formal definitions for asymmetric structures (hunter et al., 2018) let e(x, y) mean that the edge e connects its initial point x to its end point y. definition 3 let g = (v,e1, e2, `,last) be a discourse graph and let i be the initial du in v with respect to the textual order. a subgraph g′ = (v ′, e′1, e ′ 2, ` g � e′1,last) of g forms a core c just in case: (i) {i,last} ⊆ v ′; (ii) the transitive closure of e′1 induces a transitive, asymmetric ordering r over v ′ in which for every element a, other than last and i, r(i, a) and r(a,last). note that any maximal chain over v , defined in the standard way, is a core, and any set of maximal chains over v forms a core as well. we call a core c of a graph g maximal just in case there is no substructure a of g such that p(a,c) 6= ∅ and a is also a core of g. definition 4 let g be a discourse graph, and g′ a subgraph of g. let end(e) be the endpoints of an edge e and end(e) = {x : ∃e ∈ e.x ∈ end(e)}. let \ stand for set-theoretic difference, and let end(e1 \ e′1) = v p (for periphery). then g − g′ =defn (v p , e1 \ e′1, e2 � v p , ` � (e1 \ e′1), x), with x the last element in v \ v ′ ordered by a linear ordering ≺ over v . note that g−g′ may not be a discourse graph in our sense in that it is no longer weakly connected; nevertheless, it corresponds to a set of discourse graphs. note also that for a given discourse graph g and substructure g′, g′ and g − g′ may share nodes but form a partition over the set of relation instances or arcs in e1. definition 5 p(g,c), the periphery of a structure g with respect to a core c, is such that p(s,c) = g− c. definition 6 an asymmetric structure g is a graph with a core c such that c 6= g. 6.2 additional tables table 5: table of truly non-treelike structures linguistic situated truly non-treelike structures 928 1207 # truly non-treelike structures per dialogue mean 0.82 0.47 max 13 14 min 0 .tex 0 117 asher, hunter and thompson table 6: table of basic comparisons linguistic situated overview total games 45 45 total dialogues 1137 2593 total dus (including cdus) 14041 52050 # dialogues per game max dialogues 45 120 min dialogues 2 13 mean dialogues 25.27 57.62 total turns 10596 33940 # turns per dialogue max turns 120 165 min turns 1 6 mean turns 9.88 14.08 # edus (or eeus) per turn max per turn 4 4 min per turn 1 1 mean per turn 1.19 1.31 # edus (or eeus) per dialogue max per dialogue 158 204 min per dialogue 1 5 mean per dialogue 11.07 17.12 table 7: table of argument types of relations linguistic situated edu edu 9521 8741 edu cdu 1454 2025 cdu edu 995 1316 cdu cdu 301 2776 edu eeu 987 eeu edu 462 eeu eeu 19808 eeu cdu 3324 cdu eeu 3619 118 comparing discourse structures between linguistic and situated messages table 8: table of cdus linguistic situated total # cdus 1450 7651 # depth-1 1394 7368 # depth-2 54 270 # depth-3 2 13 cdu composition edu only 1450 1825 eeu only 0 5777 both 0 49 # dus contained in cdus mean 2.14 2.77 max 6 39 min 2 2 table 9: the distribution of relation labels for links in the chat-only annotations and the situated annotations, with a breakdown for the situated annotations as follows: (i) relations between eeus (and eeu-only cdus), (ii) relations between edus (and edu-only cdus), and (iii) relations connecting chat and game moves. type linguistic situated (i) (ii) (iii) question answer pair 2914 3865 929 2904 32 comment 2037 2529 0 1644 881 acknowledgement 1543 1771 0 1576 195 continuation 1194 9391 8546 841 5 elaboration 1044 1631 .tex547 1054 30 q elab 645 661 0 639 21 contrast 537 539 0 532 7 explanation 527 567 0 519 48 clarification question 456 529 0 404 125 result 445 13741 12273 395 1074 correction 232 293 118 169 7 parallel 212 193 0 192 1 conditional 155 154 0 153 0 alternation 141 128 0 128 0 narration 103 69 0 66 3 background 86 134 45 86 3 sequence 0 6863 5695 12 1159 total 12271 43058 28153 11314 3591 119 asher, hunter and thompson as background to table 10, around 38% of our 2595 situated structures are asymmetric; these structures average two peripheral structures attached to the core, with the maximum being 22 peripheral structures for one dialogue. the table makes clear the plethora of periphery type structures in the annotations. the remaining 1617 or 62% of our situated dialogues have what we call a core-only structure. around 35% of the s structures for dialogues in the situated annotations (902 out of 2595) are interleaved. table 10: table of asymmetric and interleaved structures linguistic situated total periphery-type structures 2129 2545 # periphery-type structures per dialogue mean 1.87 0.98 max 18 22 min 0 0 size of periphery-type structures max # nodes 39 51 min # nodes 1 1 mean # nodes 1.86 2.28 interleaved structures total # cores 1137 2593 total # cores containing edus 1137 900 % of core dus that are edus mean 1 0.29 max 1 0.92 min 1 0.02 percentage dialogue edus contained in core mean 0.68 0.21 max 1 1 min 0.02 0 table 11: table for persisting core-only l structures total persisting core-only l structures 161 perfect type preservation 70 elementary preservation core in s periphery 77 core in both s core and s periphery 14 120 comparing discourse structures between linguistic and situated messages table 12: table for persisting asymmetric l structures total persisting asymmetric l structures 335 perfect type preservation 95 core type preservation l core in s core // l periph in s core & periph 33 or 28% l core in s core // l periph in s-core 62 periphery type preservation l core in s core & periph // l periph in s periph 30 l core in s periph and l periph in s periph 35 elementary preservation l core in s core & periph // l periph in s core 28 l core in s periph // l periph in s core & periph 7 l core and l periph in s core and s periph 40 elementary preservation with inverted types l core in s periph // l periph in s core 5 121 introduction the settlers corpus building the corpus annotating the corpus three points about multiparty dialogue moving to situated dialogue the conceptualization and structure of game events relations relevant for the interpretation of the situated annotations complex structures relevant for the situated annotations interactions between chat moves and game events asymmetric and interleaved structures discourse threads and goals preservation of linguistic structure preservation of relations preservation of substructures preservation of subtypes conclusions appendix formal definitions for asymmetric structures hunter:etal:2018 additional tables journal of machine learning research-microsoft word template dialogue and discourse 6(1) (2015) 1-25 doi: 10.5087/dad.2015.101 ©2015 yasuko obana and michael haugh submitted 08/14; accepted 02/15; published online 02/15 co-authorship of joint utterances in japanese yasuko obana yobana@kwansei.ac.jp school of science and technology, kwansei gakuin university 2-1, gakuen, sanda city, hyogo prefecture japan, 669-1337 michael haugh m.haugh@griffith.edu.au school of languages and linguistics, nathan campus, griffith university, nathan, qld 4111, australia editor: raquel fernández abstract the paper introduces a type of joint utterance construction in japanese, in which two independent sentential-level units are amalgamated, which has hitherto received little attention in the literature. unlike traditional joint utterance construction where one speaker maintains authority over the syntactic structure of the forthcoming continuation and the other accedes to this, thereby constituting a single tcu (turn constructional unit), our examples demonstrate that both speakers can have authority over the syntactic design of joint utterances. we call such collaborative utterances „co-authored joint utterances‟ in this paper. the uniqueness of co-authored joint utterances lies in their syntactic architecture. while syntactic and semantic continuity are successfully achieved in constructing co-authored joint utterances, they represent a co-joined structure in which two sentential-level units are involved with their shared part constituting a point of amalgamation. in analysing co-authored joint utterances, we examine how they can be treated in relation to the distinction between tcu (turn constructional unit) continuation and new tcus. due to the particularities of the syntactic architecture of co-authored joint utterances, their existence raises questions about the way in which this distinction is currently operationalised, because despite being syntactically an incremental continuation, and so seemingly a tcu continuation, the co-authored joint utterance implements an action beyond what was initially instantiated by the antecedent of that joint utterance, and so arguably constitutes a new tcu. key words: joint utterance, japanese, co-authorship, tcu 1 introduction the present paper introduces a type of joint utterances in japanese, in which two separate „sentential-level units‟ 1 are amalgamated by sharing a part of each component. joint utterances traditionally refer to “a domain of practices by which a speaker produces an utterance that is designed to grammatically continue … an ongoing utterance initiated by another speaker” 1 although the term „sentence‟ is not normally associated with spoken discourse, we use the term, „sentential-level unit‟, in this paper when focusing on the syntactic analysis of a component that consists of a 'subject and predicate'. this is because many examples of co-authored joint utterances contain two „sentential-level‟ units, which is quite distinct from the structure of traditional joint utterances. obana and haugh 2 (hayashi, 2003: 1). according to such accounts, the first speaker controls the syntactic construction and the second speaker completes it by adding modifiers or subordinate clauses, or by taking over and expanding upon the first speaker‟s (unfinished) utterance. however, there is another type of joint utterance that has previously received little attention, in which both speakers equally control the syntactic structure of their own utterance. we call this phenomenon a „coauthored joint utterance‟ in this paper. this paper also examines the applicability of the distinction between tcu (turn construction unit) continuation and new tcus in relation to co-authored joint utterances. while the current distinction between tcu continuation and new tcus grew out of schegloff‟s (1996) initial claims about “increments” in english talk-in-interaction (e.g., couper-kuhlen & ono, 2007; ford, et al., 2002; ono & couper-kuhlen, 2007), its validity has since come into question (couperkuhlen, 2012; krekoski, 2012; cf. sidnell, 2012). drawing from our analysis of co-authored joint utterances, we build on these claims in suggesting that while the distinction between tcu continuation and new tcus may indeed serve as a resource for participants in many instances, it is not a distinction that is necessarily defensible across all cases of incremental expansion of prior units by other speakers. our main argument, in brief, is as follows. in the structure of co-authored joint utterances, since the second speaker‟s contribution is latched onto or “parasitic” on the antecedent, it appears to constitute an “other continuation” of the prior speaker‟s turn construction unit (tcu) (sidnell, 2012: 316). however, the resultant joint utterance implements an action that is distinct from that initially accomplished through the antecedent, and so, in this respect, it also constitutes a new tcu (ford & thompson, 1996; schegloff, 1996). due to these characteristics, the existence of these co-authored joint utterances raise questions about the way in which the distinction between tcu continuation and new tcus is currently operationalised in conversational analysis (ca). as couper-kuhlen (2012: 276) points out, this distinction is primarily a grammatical one given it “hinges on whether the material added after a point of possible completion is syntactically dependent on the prior unit or syntactically independent from it”, despite tcus themselves now generally being defined in terms of “pragmatic or action projection” (fox, et al., 2013: 732). however, on that account, a co-authored joint utterance seems to involve an instance of new tcu (given it implements an action distinct from that accomplished by the antecedent), on the one hand, yet, on the other hand, it is accomplished through what appears to be an instance of a tcu continuation (given that extension is recognisably syntactically dependent on the antecedent). the distinction between tcu continuation and new tcus, therefore, appears difficult to operationalise in the case of co-authored joint utterances due to the particularities of the way in which they are constructed. 2 co-construction and joint utterances in this section, we briefly discuss previous studies of joint utterances, and also examine the semantic contributions to the construction of these utterances. 2.1 previous studies of joint utterances joint utterances can be treated as a particular type of „co-construction‟ following initial work by jacoby and ochs (1995). co-construction involves all sorts of conversational phenomena in which participants jointly create a continuous flow of talk-in-interaction. however, in this paper, we focus on “the co-construction of syntactic units, namely, practices by which participants … co-authorship of joint utterances in japanese 3 complete a sentence-in-progress started by another participant” (hayashi & mori, 1998: 77). in other words, we focus on what is often loosely called a „sentence‟, given it is basically comprised of subject and predicate (see footnote 1), in cases where two or more participants incrementally take turns in its construction. a joint utterance is thus a single syntactic unit of talk that is collaboratively produced. joint utterances were first noted by sacks (1992a, 1992b), and since then the topic has attracted increasing attention (e.g., bolden, 2003; haugh, 2010; hayashi, 1999, 2003; hayashi & mori, 1998; kim, 1999; lerner, 1991, 1996; lerner & takagi, 1999; liddicoat, 2004; local, 2005; ono & thompson, 1996; rühlemann, 2007; sidnell, 2012; szczepek, 2006; tanaka, 1999, 2000). they are generally divided into two broad types: completions and expansions. completion types encompass instances where “the next speaker completes a syntactic unit that the first speaker has left incomplete” (rühlemann, 2007: 100), or more specifically, “a practice whereby a participant produces an utterance that is grammatically fitted to the ongoing trajectory of another participant‟s utterance-in-progress and which brings that other participant‟s utterance to completion” (hayashi, 2003: 25). expansion types of joint utterance, on the other hand, are defined as instances where “the first speaker articulates an utterance that is syntactically complete and the next speaker expands the first speaker‟s utterance into a longer syntactic unit” (rühlemann, 2007: 100), by adding a subordinate clause, a prepositional phrase or an adverbial (phrase). therefore, whether joint utterances are completions or expansions, the first speaker controls the syntactic architecture of the incipient joint utterance and the second speaker maintains it. in the following excerpt, rühlemann (2007: 100) provides examples of joint utterances that arise through both completion and expansion (the arrows here point to the continuations by the second speaker of the antecedent [howes, et al., 2011: 287] produced by the prior speaker, either through completion or expansion). (1) british national corpus: kbp2506 1. nina: no, they die down ( ) 2. clarence: mm. mm. 3. nina: most of the ones that we brought seem to have erm 4. clarence: survived. 5. nina: survived. which i‟m glad. 6. clarence: mm. mm. in (1), in response to nina indicating through the hesitation token erm that she is struggling to find the right word (line 3), clarence completes the utterance begun by nina in the subsequent turn (line 4). this constitutes an example of a completion joint utterance. nina then further expands upon the previously co-constructed utterance with the addition of a relative clause which i’m glad (line 5), which constitutes an instance of an expansion joint utterance. in other words, through both completion and expansion the two speakers can be seen to be co-constructing a single, complex syntactic unit (rühlemann, 2007: 101). 2.2 semantic contributions to constructing joint utterances it should be noted that even if a joint utterance is successfully created, the second speaker‟s prediction of the projected trajectory that can be inferred from the antecedent may or may not be obana and haugh 4 treated as intended by the first speaker (or what is retrospectively claimed to have been intended). in extreme cases, joint utterances do not have to involve such a prediction, and yet can be successfully constituted by two (or more) participants. in this sub-section, we examine the semantic contributions made by participants in constructing a joint utterance. in the case of example (1) above, clarence‟s use of „survived‟ appears to match the word nina was searching for because nina repeated the same word just after clarence uttered it. however, in example (2), izumi‟s completion of tomoko‟s utterance in line 2 is apparently not what tomoko was originally going to express (or was in the process of thinking of what to say) given the way she responds to it; nonetheless, izumi‟s utterance turns out to be more appropriate as tomoko accepts it (a, soo ka = „oh, that‟s it.‟). 2 (2) 1 tomoko: nanka higashino keigo no nandemo= somehow of everything “somehow, higashino keigo‟s (books) all…” 2 izumi: =hitto shi te masu mon ne. success do te polite md md “(are a) great success, aren‟t they?” 3 tomoko: a soo ka. oh so md “oh, that‟s it (i see).” the two interactants, tomoko and izumi, have been talking about recent popular culture, ranging from tv programmes to novels. higashino keigo is a mystery writer, who is one of the most popular novelists in recent years. in this excerpt, we find a continuation of talk about his novels, in which tomoko seemingly intended to comment on his novels. we do not know what tomoko originally intended to say. however, what can be observed here is that izumi takes over tomoko‟s utterance, and syntactically completes it. izumi‟s utterance was apparently not what tomoko originally intended to say due to tomoko‟s subsequent utterance, a, sooka (oh, i see), which shows her surprise. nonetheless, tomoko appears to accept izumi‟s continuation. hayashi (2003: 173-204) discusses this kind of semantic redirection steered by the second speaker in completing a joint utterance. for example, 2 examples (2) and (8) ~ (12) discussed in this paper are from our own data, which consists of 18 interviews (each lasting 20 minutes) and two business meetings (each lasting 30 minutes), transcribed using a basic set of conventions that are listed at the end of this paper. these examples are transcribed to a level sufficient for showing how joint utterances are created from a syntactic perspective, given our focus is primarily on the syntactic architecture of joint utterances, but not, we concede, to a level sufficient for a full multimodal or ca investigation of how participants interactionally achieve co-constructions. the latter is the focus of a number of previous studies, some of which are cited near the beginning of section 2.1. our contention is that those kinds of interactional analyses of co-constructions can be usefully informed by closer consideration of their syntactic architecture, including cases of co-authored joint utterances, which we discuss in this paper. co-authorship of joint utterances in japanese 5 (3) 1 asami: uun (ikkai) unto ne: ima uh-uh once uhm fp now “uh-uh, (once), let‟s see, now…” 2 ichiman gosen gurai dashiteru. 10,000 5,000 about pay “…((we) pay about 15,000.” 3 (0.2) 4 chika: dake de sumu n [yaro:?] only cp get away with n tag “((you)) get away with only about 15,000, right?” (example (4) in hayashi, 2003: 177) in example (3), dake (only) in line 4 of chika‟s utterance, which is termed an “utterance-initial postposition” by hayashi, is grammatically latched onto the phrase, ichiman gosen gurai (around 15,000 yen), in line 2 of asami‟s utterance. by providing further examples of utterance-initial postpositions, hayashi asserts that “the availability of particular grammatical resources … allows participants to accomplish a particular type of action, i.e., redirection of the trajectory of an utterance, in a particular, perhaps language-specific way” (hayashi, 2003: 173). although there may be questions on whether this kind of grammatical latching can actually be strictly speaking categorised as a joint utterance 3 since the utterances in lines 2 and 4 are not syntactically connected, but rather only a part inside the utterance in line 2 is taken as zero in line 4, it does indeed constitute a type of co-construction in a broader sense (but not a joint utterance in a strict sense). however, putting aside such definitional intricacies, our primary claim here is that joint utterances do not have to be anticipatorily constructed, but the second speaker has free choice as to where his/her semantic contribution should steer the continuation of the prior utterance. more extreme examples of joint utterances are cited by gregoromichelaki, et al. (2011), which they call “hostile continuations”, on the basis of which they assert that “incremental comprehension cannot be based primarily on guessing speaker intention” (gregoromichelaki, et al., 2011: 209). for example: (4) [morse, bbc radio 7] morse: in any case the question was suspect: a very good question inspector (gregoromichelaki, et al., 2011: 208) (5) (a and b arguing:) a: in fact what this shows is b: that you are an idiot. (gregoromichelaki, et al., 2011: 209) 3 hayashi (2003: 173) calls this type of construction “a particular utterance design in joint utterance construction”. obana and haugh 6 obviously, the completions in both examples (4) and (5) diverge from the first speaker‟s probable intention, thereby showing the second speaker‟s freedom to create his/her own continuation. in fact, the second speaker semantically controls the joint utterance by bringing a hostile effect into the interaction. therefore, the second speaker does not have to predict the first speaker‟s intention, and yet through syntactic continuation can co-construct a joint utterance. such hostile continuations demonstrate that “it is not obvious why … the addressee has to have guessed the original speaker‟s (propositional) intention/plan before they offer their continuations” (gregoromichelaki, et al., 2011: 209). at the same time, “speaker intentions need not be fullyformed before production”, thus, “as long as the speaker is licensed to operate with partial structures, they can start an utterance without a fully formed intention/plan as to how it will develop” (ibid, 209). we agree with gregoromichelaki, et al.‟s (2011: 226) view that “speakers do not have to be modelled as having fully-formed messages”. instead, utterances can be expanded incrementally as the conversation proceeds, and “as long as what emerges as the eventual joint content is some compatible extension of the original speaker‟s goal tree (message), it may be accepted as sufficient for the purposes to hand” (ibid, 226). likewise, the cases of co-authored joint utterances we examine here are not always the result of prediction by the second speaker of the first speaker‟s putative intention, but can be the result of the second speaker‟s own agentive continuation of the antecedent. in other words, the second speaker is not bound down by inferences about the first speaker‟s original putative intentions or the action trajectory implemented by the first speaker‟s utterance, but can go beyond it to achieve a distinct action trajectory or even a new direction of activity in the given discourse. what this sub-section has illustrated is that joint utterances do not necessarily require the initial semantic direction or action trajectory implemented by the first speaker to be maintained by the second speaker. thus, what makes a particular utterance a joint utterance is syntactic continuity passed on from one speaker to another. in most studies to date, joint utterances are typically taken to involve the completion of an ostensibly unfinished unit of talk with further talk that is syntactically continuous, but because the subsequent completion is made by more than one participant, the term „joint utterance‟ has been given, in order to distinguish it from other types of co-constructed utterances. 4 likewise, the examples of co-authored joint utterances introduced in this paper apparently preserve syntactic continuity. however, unlike examples of completion or expansion joint utterances, which constitute a single sentential-level unit, co-authored joint utterances are derived from two independent sentential-level units, a part of which is amalgamated, thereby creating a unique syntactic design. 3 co-authored joint utterance constructions in this section, we analyse examples of co-authored joint utterances, and by examining how they are constructed, we highlight significant features of such utterances, as well as discuss how they differ from other types of co-constructions. 3.1 analysis of co-authored joint utterances 4 another type of joint utterance that has received attention is what are termed expansion joint utterances. however, both completion and expansion types of joint utterance constitute a single „sentential-level unit‟. co-authorship of joint utterances in japanese 7 examples of joint utterances in previous studies generally refer to either completions or expansions, in which it is the first speaker who initiates and controls the syntactic structure of the joint utterance while the second speaker maintains it by adding further units through either completion or expansion of the antecedent. those examples present a (completed) simple sentential-level unit, which has been jointly constructed by participants. in this section, however, we discuss co-authored joint utterances in japanese, in which two syntactically independent, sentential-level units are partially amalgamated, and tied in such a way that the resultant structure presents unique structural intricacy that goes beyond the scope of most extant accounts of grammar. 5 in the case of co-authored joint utterances, therefore, both speakers have authority over the syntactic design (as well as the semantic direction) of the resultant utterance. in example (6), for instance, we can observe how the two participants take turns to construct a continuous joint utterance, in which two independent sentential-level units are integrated into one through sharing the same object. (6) (bts 6 ) 1 f02: minna kai-tai-hoodai everybody buy-want-freely “everybody bought ( ) freely” 2 f01: fudan zettai kawa-nai mono toka usually never buy-not things such as “things like what (they) usually never buy” 3 f02: gamanshi-te-iru mono o resist-te-prog thing acc “things (they) had resisted (to buy)” 4 f01: zettai ire-te-ta yo, are. certainly put-te-past md indeed “(they) surely put (such things into the basket), indeed.” two female friends refer to their sports training camp and start talking about how generous their coach was because he gave 10,000 yen for the students to buy whatever they liked in a shop. this excerpt describes how they went shopping. the four utterances in example (6) are all connected syntactically to create a co-authored joint utterance. the syntactic architecture of these interconnected four utterances is shown in two alternative ways in figure 1 below. 5 however, dynamic syntax (cann, et al., 2005; kempson, et al., 2001) and poesio and riser‟s (2010) dialogue model represent important exceptions to this trend, as does recent work in interactional linguistics (e.g., auer, 2009; laury & ono, 2014). 6 bts refers to data taken from the corpus of spoken japanese compiled by mayumi usami and her team at tokyo university of foreign studies, which is transcribed using the „basic transcription system‟ (bts). this excerpt and the other examples indicated as taken from that corpus are presented as they were transcribed using the bts, although we use the romanised form of japanese in this paper for the sake of reader accessibility. how dynamic syntax (cann, et al., 2005; kempson, et al., 2001), which gregoromichelaki, et al. (2011) argue is well equipped to handle joint utterances more generally, would treat co-authored joint utterances is an intriguing question, but lies outside the scope of this current paper. obana and haugh 8 1 f02: subject predicate 1 2 f01: object 1 joint utterance co-authored joint utterance 3 f02: object 2 4 f01: predicate 2 subject object predicate 1 (object1, object 2) predicate 2 (nb: a b = „a‟ modifies „b‟) figure 1. the syntactic architecture of the utterances in example (6). figure 1 illustrates that f02‟s first utterance (line 1) consists of a subject and predicate (predicate 1), and that f01‟s first utterance (line 2) consists of an object (object 1), while f02‟s second utterance (line 3) constitutes an additional object (object 2), which apparently functions as a postpositioned insertion into the syntactic structure in line 1, namely as the argument of predicate 1. this means that the utterances in lines 1 to 3 constitute a traditional joint utterance with the objects as increments (schegloff, 1996: 73f). however, the utterance in line 4 brings up another predicate (predicate 2), which equally takes the same object (objects 1 + 2) as its argument (but realised as zero). this means that the object here is shared between predicates 1 and 2. the subject (minna, „everybody‟), which appears in line 1, persists throughout excerpt (6) and is thus omitted in line 4. the two participants, by adding utterances incrementally, successfully construct a co-authored joint utterance in which the two sentential-level units (one unit comprising the utterances in lines 1 to 3 and the other in lines 2 to 4) are merged into one component with the argument (i.e., the object in lines 2 and 3) being shared between those two sentential-level units. this is not a mere completion, where the second speaker takes over finishing off the first speaker‟s utterance. nor is it an expansion, 7 whereby the second speaker adds peripheral units to complement the first speaker‟s utterance. while both speakers control their own syntactic design, two distinct units are amalgamated into one continuous component. these continuous turns also achieve wellcoordinated semantic continuity. co-authored joint utterances such as example (6), which show the participants‟ relayed participation, 8 are most likely to occur when the participants are in close accord, appearing to share the same (or similar) feelings about their common experiences. in a similar way, example (7) illustrates an instance where the two female participants agree that they as seniors (senpai) in their club prefer not to be involved in preparing for the forthcoming christmas party, and wish 7 in example (6), objects 1 and 2 are in coordination, and both fill in the object position. 8 we note in passing that this kind of relayed participation may be parasitic in part on the phenomenon of “clause chaining” that single speakers have been observed to accomplish in japanese conversations (see laury and ono, 2014: 571, and references therein). co-authorship of joint utterances in japanese 9 that they could just turn up at the party and pay the fee, although in reality senpai are obliged to work hard to prepare for the party. their shared feelings create a kind of relayed conversation, in which a co-authored joint utterance is incrementally accomplished. (7) (bts) 1 f14: senpai toshite motehayasa-re-tai senior as hail-pass-want “(i) want to be hailed as a senpai” 2 aa, senpaai toka it-te, mujookenni. oh senpai etc. say-te unconditionally “(they) say unconditionally such as „wow, senpai‟.” 3 f13: ki-te-kure-ta-n-desu kaa toka it-te come-te-receive-past-nomi-polite md like say-te “like (they) say, „(you) have kindly come (to the party).” 4 f14: ne md “indeed” 5 f13: irasshai toka it-te welcome like say-te “(they) say such as „welcome‟.” 6 f14: soo soo soo soo yep yep yep yep “yep, yep, yep, yep.” 7 f13: okane atsume-rare-te-mitai money collect-pass-te-try(want) “(i) once want to have (my) money collected (that way).” the units in lines 1 and 2 constitute an inverted sentence (with the unit in line 1 as a main clause), and the utterances in lines 3 and 5 are an expansion added to the main clause in line 1, which creates a traditional, expansion-type joint utterance. however, the subordinate clauses in lines 2, 3 and 5 at the same time modify the main clause in line 7. this means that the two sentential-level units (one unit comprising the utterances in lines 1~3 and 5 and the other in lines 2, 3, 5 and 7) are amalgamated with the subordinate clauses in lines 2, 3 and 5 being shared, and so form a coauthored joint utterance. this is illustrated in figure 2 below. 1 f14: main clause ① 2 subordinate clause (1) 3 f13: subordinate clause (2) 5 subordinate clause (3) 7 main clause ② figure 2. the syntactic relations between the utterances to construct joint utterances joint utterance co-authored joint utterance obana and haugh 10 figure 2 indicates that the two main clauses, ① and ②, are linked with the three subordinate clauses, (1) ~ (3); the subordinate clauses thus belong to both of the main clauses. semantically, the two participants keep the same line of thought by adding information incrementally, creating a kind of relayed conversation. syntactically, the second speaker (f13) initially contributed to an expansion of f14‟s utterance (a typical example of an “increment”: schegloff, 1996: 73f) in lines 3 and 5, thereby creating a traditional joint utterance. however, the main clause in line 7 reframes the whole discourse into a co-authored joint utterance where the two main clauses, ① and ②,are involved and fused with the antecedent subordinate clauses, (1) ~ (3). one might argue that co-authored joint utterances can be constructed only when the participants know exactly what they are talking about due to their shared experience as indicated in examples (6) and (7). however, example (8) shows that shared knowledge (or experience) is not sine qua non to the achievement of co-authored joint utterances. this is because they invariably arise through the grammatical conversion of a prior syntactic unit, and so involve more than simply the sharing of common experiences or stances. example (8) is taken from an interview which tomoko (who is a student) conducted with another student, rie. they knew each other by name, and so tomoko first asked where rie is living and then, rie asked where tomoko is living. tomoko said the name of a place, but apparently rie did not hear it clearly. the excerpt below follows on from this. (8) 1 tomoko: wakaru? [hooryuuji] hooryuuji= understand horyuji “(did you) get it? horyuji, horyuji) 2 rie: [hai hai] “yes, yes” 3 =no chikaku desu ka? of nearby polite q “(are you living) near (horyuji temple)?” 4 tomoko: eki. station “(near) the (horyuji) station” 5 rie: metcha kakkoii desu ne. very cool polite md “very cool, indeed.” 6 tomoko: (.) tte iwa reru yooninat ta otona ninat te. quote say pass become past adult become te “(so i am) told since (i) became grown-up.” the utterances in lines 1 and 3 are syntactically merged into a continuous unit as shown in figure 3. co-authorship of joint utterances in japanese 11 1 tomoko: hooryuuji. (one-word completed utterance) 3 rie: ø no chikaku desu ka. (noun) figure 3. the conversion of a one-word complete utterance into a noun in another utterance tomoko answered that she is living in horyuji. horyuji can refer to the name of a place (or that of a railway station) in nara, but at the same time it is also the name of a historically famous temple. tomoko meant the former, as it turns out, but rie initially interpreted the reference as the latter. this is why rie asked whether tomoko is living near the temple (line 3). the two utterances constitute a co-authored joint utterance, in which tomoko‟s predicate noun (hooryuuji) is shared between the two independent utterances. tomoko completes her utterance by producing one word with a descending tone, but this utterance is taken advantage of and utilized as a noun (but empty in the syntactic slot) in rie‟s subsequent utterance (line 3), which combines with a possessive particle, no, leading to the creation of a completely new sententiallevel structure (i.e., an interrogative). unlike examples (6) and (7), which are semantically driven in the same direction by adding further information step-by-step in a kind of „relay‟, example (8) involves a different orientation in the case of the continuation of the antecedent of the joint utterance. although rie‟s utterance does not deviate from trajectory at the macro level of discourse (i.e., talking about where they are currently living), it redirects the action trajectory of the prior utterance by creating a question to tomoko through a continuation of tomoko‟s prior utterance. tomoko‟s utterance in line 1 is framed as an answer to a prior question raised by rie, but is reframed as a part of a question by rie through a syntactic continuation of tomoko‟s utterance. the utterances in lines 5 and 6 constitute another co-authored joint utterance, whose syntactic relations are shown in figure 4. 5 rie: metcha kakkoii desu ne (completed sentential-level unit) 6 tomoko: ø tte iwareruyooninatta,… (quotation clause) figure 4. quotation conversion rie commented that tomoko‟s place is really cool because horyuji is historically famous. rie‟s utterance is completed with the mood marker, ne (indeed), which indicates a possible end of rie‟s turn. however, tomoko incorporates this utterance into her new utterance, by converting the former into a quotation clause in tomoko‟s subsequent utterance. tomoko replied that since she entered adulthood, she has often received such comments. syntactically, the quotation clause is omitted (i.e. a zero anaphor indicated as „ø‟ in figure 4), and in order to achieve syntactic continuation, tomoko‟s utterance employs anastrophe by inverting the two clauses, otonaninatte („since i was grown up‟) and ø tte iwareruyooninatta („so [i] have been told‟). semantically, tomoko shifts rie‟s frame into a different one, that is, from an assessment into a quotation about obana and haugh 12 which tomoko takes a particular stance. yet due to the syntactic continuity, tomoko claims equal authorship with rie, which gives rise to a co-authored joint utterance. this kind of „quotation conversion‟ often occurs in our data, especially in business meetings. for example: (9) fukuda: kyoo juuni kuru? today within come “(does it) come sometime today?” (.) doi: toyuu yotei desu. quote schedule polite “(that it will come today) has been scheduled.” (10)fukuda: sore moo tyuumon ni dashi ta? that already order to dispatch past “(did they) order that already?” (.) endo: toyuu hookoku o uke te i masu quote report acc receive te prog polite “(i) received a report (that it has already been ordered).” both (9) and (10) are excerpts from a business meeting in which fukuda as a senior person checks tasks each junior member was supposed to complete. in both examples, the first speaker‟s completed utterance with a raising tone, here framing it as a question, is converted into a quotation clause through the continuation by the second speaker. this means that while two sentential-level units are merged into one component, both speakers not only control their own syntactic design but also proffer two different action trajectories, that is, a question and an answer. it should be noted that this conversion is possible only when the first speaker‟s sentence structure is affirmative with a raising tone. if the first speaker‟s structure is formulated in an interrogative form with a question marker, ka or no, the second speaker cannot create a joint utterance. this means that in spite of the raising tone in the first speaker‟s utterance, which forms a question, the second speaker focuses only on the syntactic trajectory and adds further units, which serve as an answer to the first speaker‟s enquiry. in sum, joint utterance construction has traditionally been regarded as allowing “the original speaker to maintain his/her authority over a turn even when completed by another” (lerner, 2004: 225). however, the examples of co-authored joint utterances that we have discussed indicate that both speakers contribute talk with different syntactic designs and yet this constitutes a syntactically continuous construction by sharing a part across the two independent sententiallevel units. a syntactically shared part between the two sentential-level units modifies the head of each utterance as shown in example (6), modifies the main clause as shown in (7), or is grammatically converted into a different unit in the second speaker‟s continuation as shown in examples (8) to (10). it may seem in our analysis thus far that we have emphasised syntactic aspects of co-authored joint utterances at the expense of a consideration of the role of semantic links in the construction of co-authored joint utterances. indeed, co-authored joint utterances cannot be constructed without semantic considerations, and the semantic contribution of utterances is important to co-authorship of joint utterances in japanese 13 consider in analysing how the units in the co-authored joint utterance are related to one another. for example, the co-authored joint utterances in examples (6) and (7) constitute a continuous relay as if one person were uttering a single utterance, even though the utterance arises over multiple turns by the two participants. in examples (8) to (10), both speakers exhibit their own authorship in interaction, and semantic divergence occurs because each participant projects their own stance with respect to the action trajectory in question. however, it is rather the unique syntactic architecture (i.e., syntactic amalgamation) of co-authored joint utterances that distinguishes them from other types of co-construction. co-constructions involve all sorts of cases in which the participants jointly create a continuous flow of talk-in-action. they may reveal only fragments of utterances, offer only semantic connections between utterances, and/or consist of independent utterances which nonetheless achieve a semantically relevant flow. on the other hand, co-authored joint utterances, which we argue constitute yet another different type of coconstruction, exhibit the syntactic amalgamation of two independent sentential-units, and the resultant construction is so tightly-entwined that it is arguably structurally inseparable. 3.2 anaphoric relations in co-authored joint utterances the main reason why co-authored joint utterances are attainable in japanese is that the japanese language allows ellipsis to occur in all sorts of syntactic positions. this is particularly conspicuous in conversation. for instance, the utterance in line 1 of example (6) can stand on its own with its object omitted although the verb (kau = „to buy‟) is potentially transitive. in examples (8) to (10), the second speaker takes over the first speaker‟s utterance and utilizes it syntactically as zero in the former‟s continuation of the first speaker‟s utterance, and yet it is naturally accepted in conversation. as a matter of fact, the participants take advantage of the omitted units, which come into play in building up a co-authored joint utterance. because of this, one may wonder if co-authored joint utterances can be analysed with reference to anaphoric relations alone. if not, as we are arguing, then what distinguishes co-authored joint utterances from other elliptic co-constructions? to further consider this question, let us look at example (11), which appears at first glance to be a syntactic continuation that is collaboratively constituted. (11) (bts) f02: sugokat-ta ne are wa. terrific-past md that top “terrific, that was.” f01: omoshirokat-ta, un. fun-past hmm “(that was) interesting, indeed.” this excerpt follows on from example (6), and involves the participants commenting on the experiences they shared in the camp. f02‟s utterance is syntactically scrambled between subject and predicate, and completed with a subject. f01‟s utterance starts without a subject, but it is evident that the subject is the same as that of f02‟s utterance. the subject, are wa („that‟), appears to be the merging point between the two utterances. from this perspective, example (11) can be treated as an instance of a co-authored joint utterance. however, since are wa does not receive any grammatical conversion but remains as the same subject in f01‟s utterance, it is also obana and haugh 14 possible to assume that the subject in f01‟s utterance is not a merging point, but merely a case of ellipsis. that is, two separate utterances are displayed, but because the subject has already occurred in f02‟s utterance, it is omitted for the sake of establishing topic continuity 9 (givón 1983; hinds, 1978, 1982, 1983, 1984). therefore, inverted structures like f02‟s utterance in example (11) are somewhat more controversial as candidate instances of co-authored joint utterances. it is difficult to conclusively determine whether it should be treated as a co-authored joint utterance or as two separate, distinct units with an anaphoric relation. to avoid such controversies, it is possible to limit the scope of co-authored joint utterances to those which exhibit grammatical conversion at the merging point (except relay-like structures as shown in examples (6) and (7); this will be discussed below). for instance, in example (8), hooryuuji (horyuji temple) as a one-word completed utterance in line 1 is converted into a simple noun in a noun phrase in line 3. in example (9), fukuda‟s utterance, which is a completed utterance as a question, is converted into a subordinate clause in doi‟s utterance. on the other hand, example (11) does not display such a grammatical conversion since the same unit is merely shared between the two utterances. we thus regard this phenomenon as a simple case of ellipsis, as shown in figure 5. f02: predicate 1 + subject f01: ø + predicate 2 figure 5. the subject realised as a zero anaphor in the subsequent utterance grammatical conversion may be a solution to distinguish co-authored joint utterances from other elliptic cases like example (11). however, grammatical conversion does not occur in examples (6) and (7) although we claim these examples should be considered co-authored joint utterances. it might therefore be argued that examples (6) and (7) can be explained as forms of ellipsis. indeed, we admit that examples (6) and (7) involve anaphoric relations with zero units shared between the utterances, and co-authored joint utterances are obtained because they take advantage of the phenomena of ellipsis in japanese. however, unlike other utterances bearing ellipsis, coauthored joint utterances bring in syntactic continuity which builds up a unique structure: the merging of two independent sentential-units. this unorthodox syntactic architecture distinguishes co-authored joint utterances from other zero-anaphoric utterances. in other words, while elliptic utterances in general only indicate how zeros refer back to relevant antecedents, co-authored joint utterances go further and exploit zeros to create an amalgamated syntactic continuation. 4 joint utterances and the role of syntax in turn continuation turn constructional units (tcus) are one of major features often highlighted to explain turntaking systems in ca. sacks, et al. (1974) originally defined a tcu as a recognisably complete 9 zero anaphora is considered to be on top of the coding topic accessibility. in example (11), for instance, are wa (with wa as a topic marker) is assigned as a topic in the discourse, and in subsequent structures, the topic is omitted as evidence of “the unmarked form of topic continuity” (hinds, 1983:49). co-authorship of joint utterances in japanese 15 unit in conversation. tcus can be words, phrases, clauses and sentences, and participants predict through their talk where they expect a given tcu is coming to its point of possible completion. while it is not always acknowledged in ca, sacks, et al. (1974) in fact relied heavily on syntactic constructions in conceptualising tcus, because a new tcu (by another speaker, for example) does not start arbitrarily at any point in the previous utterance, but in the vicinity of the projectable completion point of a certain syntactically self-contained unit (ford, et al., 1996, 2013). in this respect, traditional joint utterance construction is an interesting case because a single unit is collaboratively produced. lerner (1991) considers joint utterance constructions to nevertheless constitute single tcus, terming such cases a type of “compound tcu”, because it “projects in its course, and prior to the onset of a final component, that a two-part unit is underway” (lerner & takagi, 1999: 53). this assumption seems to be plausible because a traditional joint utterance constitutes a single syntactic unit, as well as constituting a single unit of action 10 , except that it is contributed to by more than one participant. however, the instances of co-authored joint utterances we have examined suggest that the proposed distinction between tcu continuation and new tcus (couper-kuhlen & ono, 2007; ford, et al., 2002; schegloff, 1996) may be more difficult to maintain than has been assumed to date. the first issue is that co-authored joint utterances consist of two independent sentential-level units; the first component is already completed, and yet the second component continues on to attach another syntactically autonomous sentential-level unit (although taking advantage of a part of the first component). this means that the second component is not simply an “increment” (schegloff, 1996), but constitutes another form of syntactic continuation. the second issue is that co-authored joint utterances often involve cases where the second speaker‟s continuation deviates from the incipient trajectory of the first speaker‟s utterance, and creates a new semantically or pragmatically independent unit, and yet this is not subsequently contested or resisted by the first speaker. these observations thus cast doubt on the applicability of the distinction between tcu continuation and new tcus in the case of co-authored joint utterances, given there is clearly syntactic continuation at play here (and so is recognisable as an instance of tcu continuation to participants), and yet the action instantiated through the joint utterance can differ from the incipient action trajectory of the antecedent (and so a co-authored joint utterance is recognisable as an instance of a new tcu to participants). in this section, we first review previous studies on tcu continuation in further detail, and then examine whether or not the distinction between tcu continuation and new tcus can apply to our examples of co-authored joint utterances. 4.1 previous studies of increments and tcu continuation the distinction between tcu continuations and new tcus is closely related to the finding that in natural conversation, not all transition-relevance places (trps) necessarily lead to speaker 10 however, if action is taken into consideration, gregoromichelaki, et al.‟s (2011) “hostile continuations” and hayashi‟s (2003) “utterance-initial postpositions” discussed in section 2.2 would be problematic as they redirect the action initiated by the prior speaker, thereby, giving rise to two units of action in the single (compound) tcu. equally problematic are purver et al.‟s (2010) “split utterances” where speaker/addressee pronouns change as the speaker transition occurs (e.g., a: “did you give me back” b: “your penknife?...”: purver, et al., 2010: 43), and gregoromichelaki, et al.‟s (2011) „question-answer‟ type joint utterance (e.g., a: “are you left or” b: “right handed” : gregoromichelaki, et al., 2011: 208). obana and haugh 16 change 11 , a claim that is now well established in ca (clayman, 2013; sacks, et al., 1974; schegloff, 1996). this is partly because utterances in conversation may contain repetitions, repairs or retrospective add-ons by the current speaker, through which a potential trp is circumvented in spite of the projectable completion of a tcu. for example, the same speaker keeps his/her turn and continues talking after apparently having completed his/her utterance due to a particular array of syntactic and prosodic features that include compressing the trp or bridging it through pivots (clayman, 2013). the term “increment” was initially used by schegloff (1996) to refer to a syntactically extended component added to the host tcu. an increment is grammatically continuous to the prior-host tcu, and so is syntactically dependent on the host which itself stands as a complete tcu. therefore, increments are a form of tcu continuation (schegloff, 2001). on the other hand, a new tcu is pragmatically independent, and is claimed to not be related to the prior unit (couper-kuhlen, 2012: 276), although schegloff (1996: 76) also discusses cases where recognisably new tcus are “grammatically continuous with what preceded”. ford, et al. (2002: 16) follow suit, referring to an increment as a “non-main clause after a possible point of turn completion”. in the following example taken from two students‟ conversation, we can find evidence of the existence of these kinds of incremental continuations in japanese. (12) 1 tomoko: watashi kekkoo yotchaun desu yo. i quite get sick polite md “i often feel sick, actually.” 2 hiroko: densha de? train on “on the train?” 3 tomoko: furansu no tgv desae yotta koto aru-n desu yo. france of tgv even get sick-past thing exist-nomi polite md “(i) have been sick even on tgv in france.” 4 hiroko: e:::? really “really?” 5 tomoko: ha ha ha sonnani yurenai noni laugh so much shake though “(laugh) though (that train) does not rattle around so much” in example (12), the utterance in line 5 is an incremental expansion of the utterance in line 3 which is the host component. the two participants were talking about travelling, and tomoko told hiroko that she does not particularly like travelling by train because she often feels sick (line 1), and she felt sick even on the tgv (line 3). she then added „though [that train] does not rattle around so much‟ (line 5). although the utterance in line 3 is a complete tcu and hiroko‟s interjection enters straight after line 3, the utterance in line 5 evidently constitutes a continuation of the tcu in line 3 due to its semantic and syntactic dependence on the tcu in line 3. therefore, 11 for example, selting (2000) points out that in story-telling contexts, speakers often present multiple tcus within a single turn, (temporarily) suspending the sequential relevance of trps. co-authorship of joint utterances in japanese 17 the component in line 5 is analysable as an increment, and thus the utterances in lines 3 and 5 can be considered to be an example of tcu continuation. in this way, then, tcu continuations have been determined mainly based on syntactic continuity, where add-ons are incremental to and syntactically dependent on the host component (e.g. couper-khulen & ono, 2007; ford, et al., 2002; schegloff, 1996, 2001), although couperkhulen & ono (2007) and ford, et al. (2002) also admit the importance of prosody and other pragmatic features in analysing increments. however, the central role placed on syntax in dealing with tcu continuation has been challenged by other researchers, who take a more holistic view of tcus and include semantic, prosodic and other pragmatic features (e.g. auer, 2007; ford, 2004; ford, et al., 1996; krekoski, 2012; luke & zhang, 2007; sidnell, 2012). streeck and hartge (1992) even include extralinguistic features such as gesture, gaze, facial expressions in locating trps. let us look at a few examples. luke and zhang (2007) go into detail in describing different types of increments, and argue that certain “insertables” in chinese, although they appear to be syntactically incremental and thus dependent, should be considered free constituents. for example, when cai („just‟, an adverb) is retrospectively added to an apparently completed component, it can syntactically be placed back into this host component. however, luke and zhang (2007: 622f) consider this adverb to nevertheless constitute a new tcu because the host component ends with the descending tone and a slight pause. thus, they argue that syntactic continuation alone cannot determine whether the unit concerned is continuous to the host component (thus, constituting one continuous tcu) or distinctive from the host (thus, accounting for two tcus). they conclude that prosody should be taken into consideration. if prosody should be considered in distinguishing tcu continuation from new tcus, as argued by luke and zhang (2007), the utterance in line 5 of example (12) above should be regarded as a new tcu because the utterance in line 3 ends with a descending tone and stands as a complete tcu in its own right, which thus allows the other speaker to enter the interaction straight after the utterance in line 3. in a similar way, luke and zhang‟s (2007) claim would undermine krekoski‟s (2012) argument about increments as they would presumably consider the utterance in line 3 of example (13) to constitute a new tcu. (13) 1 m: [suki da kara nani mo] iemasen kedo ne. like therefore nothing say:can:not but fp “he likes it so nothing can be said” 2 k: ..n=. “mhm” 3 m: . . . kotchi wa. this side wa “(by) me (lit. „(as for) this side‟)” (adapted from krekoski, 2012: 301) however, krekoski‟s (2012) claim is that m‟s talk in line 3 is an increment relative to prior talk in line 1. he argues that although the utterance in line 1 ends with a descending tone and “marks clearly syntactic closure”, the subsequent talk in line 3 shows that “the speaker opts to continue the previous possible completed turn with the increment” and so concludes that “line 3 represents obana and haugh 18 an extension of the prior action, and is retrospectively oriented” (krekoski, 2012: 301). however, according to luke and zhang (2007), this insertable utterance should be considered a new tcu. one may argue that luke and zhang‟s (2007) insertable is a mere add-on (or an adjunct syntactically), whereas the add-on in line 3 of example (13) is a filler which is syntactically missing in the host (line 1). this may be a crucial point that distinguishes between increments and new tcus, although neither luke and zhang (2007) nor krekoski (2012) refer to this criterion. if this criterion is employed, what could be treated as a mere increment in line 5 in example (12) should be considered a new tcu. however, this creates further complications. according to schegloff (2001) and ford, et al. (2002), increments are grammatically continuous to the main (host) clause, which can be integrated into the host component; that is, increments are non-main clauses as extensions to the host clause. however, if prosody and pause are to be incorporated, as argued by luke and zhang (2007), the current definition of increments should be re-examined; otherwise, different scholars would present different analyses of the same phenomena when distinguishing between tcu continuation and new tcus. krekoski (2012) further argues that even syntactically independent utterances can be considered continuations as a result of prosodic effects. for example, (14) 1 w: ..datte tsukuru dake ja sumanai mon. but prepare only dewa end:neg fp “but (it) doesn‟t end only with preparing (meals),” 2 .. katazukeru mon tidy fp “(they/we/i) tidy up (the house).” 3 … kaimono mo suru mon. shopping also do fp “(they/we/i) also do the shopping.” (adapted from krekoski, 2012: 308) although example (14) consists of three completely independent clauses, prosodic patterns are repetitive with mon (a mood marker indicating the end of a clause) carrying the same declining tone in all utterances in lines, 1 to 3, with a short pause in each clause. therefore, krekoski (2012: 309) concludes that the utterances in lines 2 and 3 are “strongly and retrospectively oriented and pragmatically bound together, and seem to be serving as a continuation of action initiated with the clause ending” in line 1. this means that krekoski treats “prosodic coherence linking the material together” (krekoski, 2012: 311) as an indicator of tcu continuation even if the utterances are all syntactically independent. krekoski‟s (2012) idea directly challenges the definition of increments by schegloff (1996) because increments are supposed to be syntactically bound to the host component. according to schegloff‟s definition, continuations like example (14) should not be considered increments but other types of continuation. krekoski‟s assumption also appears to undermine lerner‟s (2004) treatment of tcu continuation, because lerner excludes coordinated sentences linked with and from the category of tcu continuations since each clause in the coordination is complete and could potentially achieve a distinct unit. in this respect, lerner (2004) is apparently in line with schegloff (1996) who places value on syntactic dependence in determining whether something constitutes a tcu continuation. however, if krekoski‟s (2012) approach is employed, co-authorship of joint utterances in japanese 19 coordination should be regarded as a single tcu because it binds two clauses prosodically as well as pragmatically, and the speaker‟s „intention‟ to continue talking after the first clause can be clearly observed. 4.2 co-authored joint utterances and tcu continuations previous studies of tcu continuation seem to show that in spite of general acceptance of the term, its definition is not clear enough to explain many phenomena, and therefore, researchers resort to or add on different features such as prosody, epistemic values or putative speaker intentions to distinguish between tcu continuation and new tcus. however, this has led different scholars to arriving at different conclusions, and controversy over the distinction between tcu continuation and new tcus is still ongoing. the examples of co-authored joint utterances examined in this paper serve to further complicate this issue because they intersect with all the potential controversies discussed in section 4.1. first, because two sentential-level units are amalgamated in the case of co-authored joint utterances, the concept of „increment‟ does not apply because increments are supposed to be syntactically dependent on the host component. second, while both speakers implement an independent utterance with its own action trajectory and claim equal authorship, the construction as a whole is syntactically continuous. this is a new phenomenon that has not been considered to date, because studies of tcu continuation have focussed on two (or more) components in which one is dominant and the other(s) syntactically dependent. third, in building up co-authored joint utterances, the second speaker does not necessarily maintain the same trajectory as that of the first speaker; the former has freedom to exercise control over the directionality as long as the overall flow of conversation remains intact. finally, prosody and inferences about the first speaker‟s putative intentions with respect to the antecedent are often irrelevant to the second speaker‟s continuation, and yet the whole construction is semantically as well as syntactically continuous. the question thus arises: is it really necessary to determine whether co-authored joint utterances involve tcu continuation or consist of two independent tcus? we believe it is not because phenomena of co-authored joint utterances go beyond the syntactic mechanisms that are generally invoked when making this distinction. for example, let us return to consider how example (6) can be parsed. 1 f02: subject predicate 1 2 f01: object 1 3 f02: object 2 4 f01: predicate 2 figure 1. example (6) figure 1 shows that by taking advantage of a zero argument in the object position in line 1, two speakers add on objects 1 and 2 to predicate 1 (of the first speaker, f02), and then the speaker, f01, in line 4, further takes advantage of these add-ons to link to a new predicate (predicate 2). although two sentential-level units (the utterances in lines 1 and 4 with incremental units in lines 2 and 3) are joined, there is no cut-off point syntactically; it is a symmetrically conjoined unit that obana and haugh 20 is tightly knitted into one continuous unit. semantically the two participants cooperate in building up continuous discourse by adding information each in turn. originally line 1 is treated as a complete, independent clause, and therefore, the utterances in lines 2 and 3 can be considered increments. however, the utterance in line 4 is not a simple add-on but a predicate which takes advantage of objects 1 and 2, forming another distinct sentential-level unit. therefore, figure 1 presents a tightly-knitted construction, which is created by amalgamating two independent sentential-level units, and yet gives rise to a semantically coherent unit. this kind of construction is not a matter of „conversational turns‟ but a jointly-interlaced or fused component that cannot be parsed any longer once it is formulated in that way. the grammatical conversions shown in examples (8) to (10) present more intricate issues. let us look at figure 4 again, which we have altered slightly for the subsequent discussion below. rie: metcha kakkoii desu ne (completed sentence) tomoko: ø tte iwareruyooninatta, otona ni nat-te. (quotation clause) (rie: [it‟s] really cool. tomoko: [so i am] told, since [i] became grown-up.) figure 4. quotation conversion rie comments on the place where tomoko is living, which is uttered as a complete utterance with ne (a mood marker which projects possible completion of the turn, i.e., a trp). however, this independent sentential-level unit is used as a quotation (clause) in tomoko‟s next utterance. in that way, rie‟s complete sentential-level unit is converted into a quotation clause in tomoko‟s construction, although it is realised as zero in reality. grammatical conversions are possible because japanese is a head-final language and also allows ellipsis to occur in all sorts of syntactic slots. the first speaker‟s component is used as an embedded clause in the second speaker‟s component, and the second speaker‟s main clause occurs on the right-hand node in syntax, as is shown in figure 7. figure 7. the interlocked structure derived from two independent sentential-level units the second speaker, therefore, not only takes advantage of the first speaker‟s component ① and grammatically converts it into a different role, but also declares syntactic authorship by placing ① into a subordinate clause in the component ②. on the other hand, the occurrence of ② the first speaker‟s complete sentence ① φ(① realized as zero and embedded) + main clause the second speaker‟s complete sentence ② co-authorship of joint utterances in japanese 21 depends on that of ①. this means that while the speaker of ② declares equal authorship with the speaker of ①, the occurrence of ② is not totally independent from the context initiated by ①. in a similar way, from the perspective of action trajectory, ② offers a new direction by shifting the original frame maintained by ① to a different one (i.e., shifting from a comment to a quotation), and yet both ① and ② are closely bound together in the whole discourse because ② is an extension of the action instantiated by ①. therefore, we argue that co-authored joint utterances present features that go beyond those encapsulated by the current distinction between tcu continuation and new tcus. they are not like traditional tcu continuations in which increments are added to the host component. they are co-created by two speakers through two independent sentential-level units, and yet their syntactic construction is so tightly-knitted that they are structurally inseparable. co-authored joint utterances are something speakers produce jointly as if weaving (independent) strands into a single tapestry. in conversational interaction, two (or more) participants perform dynamic acts through the toand-fro of fragments, retrospective add-ons, interruptions and synchronisms, sudden stops and unexpected silences. turns are thus constantly extemporaneous. ford, et al. (1996) and ford (2004) assert that tcus are emergent and cannot be pre-defined because conversation involves “the management of simultaneously unfolding facets of action, sound production, gesture, and grammar produced by multiple participants” (ford, 2004: 27). this view of tcus undermines, however, the emphasis placed on syntactic structure in making the distinction between tcu continuation and new tcus, and thus the implicit reliance on syntax in defining tcus themselves, a point of critique which has been noted by others working in ca (ford, et al., 1996, 2013). what is more, there are a lot more significant features at possible points of tcu completion, or trps, where the transition from one speaker to another can take place; for example, rules and restrictions when amalgamating two syntactic units into one component, semantic diversities and changes in speaker stance between the participants, and characteristics of the japanese language which make it possible to create co-authored joint utterances. therefore, analysing an utterance as either a tcu continuation or two distinct tcus would “miss building an account of what people are doing in interaction” because various practices, “syntactic, pragmatic prosodic, gestural, can be drawn upon in a wide variety of ways to frame conversational actions as nearing, or not nearing, completion, and thus displaying participants‟ understanding of whether or not it is someone else‟s talk” (ford, et al., 1996: 450). there is another point to be noted. we assume that in analysing interaction in languages such as japanese which allow ellipsis in all sorts of places in utterances, we would face more difficulties in distinguishing new tcus from tcu continuations. couper-kuhlen (2012: 298) argues that “in languages where so-called „zero-arguments‟ abound, … tcu continuations and new tcus would be all the more difficult to distinguish”. the examples of co-authored joint utterances we have presented in this paper are indeed just such an example that takes advantage of zero arguments, thereby enabling two sentential-level units to be conjoined. these two independent units are intertwined as if original strands are no longer autonomous strands but woven into a tapestry. it is our contention that in the case of such tightly-knitted phenomena, the distinction between tcu continuation and new tcus is no longer significant for either participants or analysts. obana and haugh 22 5 conclusion in this paper, we have introduced a type of joint utterance, in which two independent sententiallevel units are conjoined by each sharing a part of the other unit. we call these „co-authored joint utterances‟ because both speakers equally control their own utterance; they are not a mere expansion of the host component or completion of the first speaker‟s incomplete utterance. syntactically, two independent sentential-level units are amalgamated, and the omitted units (in the second speaker‟s component) function as the merging point for the whole construction. the co-authored joint utterance is unique due to two sentential-level units being involved and partially merged on the same syntactic plane. we have also examined the applicability of the distinction between tcu continuation and new tcus in relation to examples of co-authored joint utterances. we have argued that the construction of co-authored joint utterances goes beyond a matter of „conversational turns‟ as examined through the lens of tcus, because although two participants are involved, the construction is fused and inseparable. it is a jointly interlaced unit for which the distinction between tcu continuation and new tcu no longer appears significant for those participants. terms used in morphological gloss acc – accusative case marker conj – conjunction cop – copula, da and its conjugated forms md – mood marker nomi – nominalizer pass – passive forms, -reru/-rareru past – past tense, -ta polite – polite forms, masu and desu prog – progressive auxiliary q – question marker, ka quote – quotation from, to and its variations such as -tte, -toiu te – the form which bridges between a verb and an auxiliary top – topic marker, wa transcription conventions = talk „latched‟ onto previous speaker‟s talk : stretching of sound . falling intonation [ ] overlapping talk capitals markedly louder talk (.) micro-pause funding acknowledgement: this work was supported by grant-in-aid for scientific research (kakenhi), the ministry of education, culture, sports, science and technology, japan (grant no. kiban (c) – 23520537) co-authorship of joint utterances in japanese 23 references peter auer (2007). why are increments such elusive objects? – an afterthought. pragmatics, 17: 647-658. peter auer (2009). projection and minimalistic syntax in interaction. discourse processes, 46: 180-205. galina b. bolden (2003). multiple modalities in collaborative turn sequences. gesture, 3: 187212. ronny cann, ruth kempson and lutz marten (2005). the dynamics of language. oxford: elsevier. steven e. clayman (2013). turn-constructional units and the transition-relevance place. in sidnell, j and stivers, t. (eds.) handbook of conversation analysis. malden, ma: wileyblackwell, pp.150-166. elizabeth couper-kuhlen (2012). turn continuation and clause combinations. discourse processes, 49: 273-299. elizabeth couper-kuhlen and ono tsuyoshi (2007). incrementing in conversation: a comparison of methods in english, german and japanese. pragmatics, 17(4): 513-552. cecilia e. ford (2004). contingency and units in interaction. discourse studies, 6(1): 27-52. cecilia e. ford, barbara a. fox, and sandra a. thompson (1996). practices in the construction of turns: the “tcu” revisited. pragmatics, 6(3): 427-454. cecilia e. ford, barbara a. fox, and sandra a. thompson (2002). constituency and the grammar of turn increments. in: ford, c.e., fox, b. aa, and thompson, s.a. (eds.) the language of turn and sequence. oxford: oxford university press, pp.14-38. cecilia e. ford, barbara a. fox, and sandra a. thompson (2013). units and/or action trajectories? the language of grammatical categories and the language of social action. in: szczepek, r.b. and raymond, g. (eds.) units of talk – units of action. amsterdam: john benjamins, pp.13-55. barbara a. fox, makoto hayashi, and robert jasperson (1996). resources and repair: a crosslinguistic study of syntax and repair. in: ochs, e., schegloff, e.a., and thompson, s.a. (eds.) interaction and grammar. cambridge: cambridge university press, pp.185-237. barbara a. fox, sandra a. thompson, cecilia e. ford, and elizabeth couper-kuhlen (2013). conversation analysis and linguistics. in: sidnell, j. and stivers, t. (eds.) handbook of conversation analysis. malden, ma: wiley-blackwell, pp.726-740. eleni gregoromichelaki, ruth kempson, matthew purver, gregory j. mills, ronnie cann, wilfried meyer-viol, and patrick g. healey (2011). incermentality and intention-recognition in utterance processing. dialogue and discourse, 2(1): 199-233. michael haugh (2010). co-constructing what is said in interaction. in: eniko, n.t. and bibok, k. (eds.) the role of data at the semantics-pragmatics interface. berlin: degruyter mouton, pp.349 – 380. makoto hayashi (1999). where grammar and interaction meet: a study of co-participant completion in japanese conversation. human studies 22: 475-499. makoto hayashi (2003). joint utterance construction in japanese conversation. amsterdam/philadelphia: john benjamins publishing company. makoto hayashi and junko mori (1998). co-construction in japanese revisited: we do“finish each other‟s sentences”. in: akatsuka, n., hoji, h., iwasaki, s., sohn, s.o., and strauss, s. (eds.) japanese/korean linguistics, volume 7. stanford/california: csli publications, pp.77-93. obana and haugh 24 christine howes, matthew purver, patrick g. healey, gregory j. mills, and eleni gregoromichelaki (2011). on incrementality in dialogue: evidence from compound contributions. dialogue and discourse, 2(1): 279-311. sally, jacoby and elinor oshs (1995). co-construction: an introduction. research on language and social interaction, 28: 171-183. talmy givón (1983). topic continuity in discourse: an introduction. in: givón t. (ed.) topic continuity in discourse (typological studies in language. vol. 13). london: john benjamins publishing company, pp.1-14. john hinds (1978). anaphora in japanese conversation‟. in: hinds j. (ed.) anaphora in discourse. edmonton: linguistic research inc., pp.136-179. john hinds (1982). ellipsis in japanese. carondale and edmonton: linguistic research inc. john hinds (1983). topic continuity in japanese. in: givón t. (ed.) topic continuity in discourse. amsterdam: john benjamins publishing company, pp.43-93. john hinds (1984). topic maintenance in japanese narratives and japanese conversational interaction. discourse processes 7: 465-482. ruth kempson, wilfried meyer-viol and dov gabbay (2001). dynamic syntax: the flow of language understanding. oxford: blackwell. jun-ju kim (1999). the co-construction of utterances in korean in face-to-face conversation between friends. crossroads of language, interaction, and culture, 1: 61-75. ross krekoski (2012). causal continuations in japanese. discourse processes, 49: 300-313. ritva laury and tsuyoshi ono (2014). the limits of grammar: clause combining in finnish and japanese conversation. pragmatics, 24: 561-592. gene lerner (1991). on the syntax of sentences-in-progress. language in society, 20: 441-458. gene lerner (1996). on the „semi-permeable‟ character of grammatical units in conversation: conditional entry into the turn space of another speaker. in: ochs, e., schegloff, e.a., and thompson, s.a. (eds.) interaction and grammar. cambridge: cambridge university press, pp.238-276 gene lerner (2004). on the place of linguistic resources in the organization of talk-in-interaction: grammar as action in prompting a speaker to elaborate. research on language and social interaction, 37: 154-184. gene lerner and tomoyo takagi (1999). on the place of linguistic resources in the organization of talk-in-interaction: a co-investigation of english and japanese grammatical practices. journal of pragmatics, 31: 49-75. anthony j. liddicoat (2004) the projectability of turn constructional units and the role of prediction in listening. discourse studies, 6(4): 449-469. john local (2005). on the interactional and phonetic design of collaborative completions. in: hardcastle, w. and beck, j. (eds.) figure of speech: a festschrift for john laver. mahwah, nj: lawrence erlbaum, pp.263 – 282. kwang-kwong luke and wei zhang (2007). retrospective turn continuations in mandarin chinese conversation. pragmatics 17: 605-635. tsuyoshi ono and elizabeth couper-kuhlen (2007). increments in cross-linguistic perspective: introductory remarks. pragmatics, 17(4): 505-512. tsuyoshi ono and sandra a. thompson (1996). interaction and syntax in the structure of conversational discourse: collaboration, overlap, and syntactic dissociation. in: hovy, e.h. and scott, d.r. (eds.) computational and conversational discourse: burning issues – a interdisciplinary account. berlin: springer-verlag, pp.67-96. co-authorship of joint utterances in japanese 25 massimo poesio and hannes rieser (2010). completions, coordinations, and alignment in dialogue. dialogue and discourse, 1: 1-89. matthew purver, eleni gregoromichelaki, wilfried meyer-viol and ronnie cann (2010). splitting the “i”s and crossing the “you”s: context, speech acts and grammar. proceedings of semdial (pozdial) 2010: 43-50. christoph rühlemann (2007) conversation in context – a corpus-driven approach. london: continuum. harvey sacks (1992a). lectures on conversation (vol.1). oxford: blackwell. harvey sacks (1992b). lectures on conversation (vol.2). oxford: blackwell. harvey sacks, emanuel a. schegloff, and gail jefferson (1974). a simplest systematic for the organization of turn-taking for conversation. language, 50: 696-735. emanuel a. schegloff (1996). turn organization: one intersection of grammar and interaction. in: ochs, e., schegloff, e.a., and thompson, s.a. (eds.) interaction and grammar. cambridge: cambridge university press, pp.52 – 133. emanuel a. schegloff (2001). conversation analysis: a project in process – “increments”. forum lecture, lsa linguistic institute, uc santa barbara. margaret selting (2000). the construction of units in conversational talk. language in society, 29: 477-517. jack sidnell (2012). turn-continuation by self and other. discourse processes, 49: 314-337. jürgen streeck and ulrike hartge (1992). gestures at the transition place. in: auer, p. and di luzio, a. (eds.) the contextualization of language. amsterdam/philadelphia: john benjamins publishing company, pp.135-157. beatrice szczepek (2006). prosodic orientation in english conversation. basingstoke: palgrave macmillan. tanaka hiroko (1999). turn-taking in japanese conversation: a study in grammar and interaction. amsterdam/philadelphia: john benjamins publishing company. tanaka hiroko (2000). turn-projection in japanese talk in interaction. research on language and social interaction 33: 1-38. dialogue and discourse 4(1) (2013) doi: 10.5087/dad.2013.101 clarification and generalized quantifiers∗ robin cooper cooper@ling.gu.se department of philosophy, linguistics and theory of science university of gothenburg box 200 405 30 göteborg, sweden editor: raquel fernández abstract we attempt to show that a classical concern of formal semantics, quantification, interacts in an interesting way with dialogue strategies when viewed from the perspective of our approach to semantics using type theory with records (ttr). ttr analyzes semantic content in terms of structured semantic objects containing subcomponents and we argue that these components influence what a dialogue participant can take up in response to a reprise of part of a previous utterance of a quantified sentence. the discussion builds on previous work by purver and ginzburg who introduce the reprise content hypothesis and use it to argue for a particular analysis of quantification. in previous work we contrasted their approach with a more classical generalized quantifier analysis. in the present paper we synthesize the two approaches and suggest that this gives us the best account of the reprise phenomena associated with quantification. keywords: clarification, quantification, reprise content hypothesis, type theory 1. introduction an important concern of classical formal semantics (montague, 1974; dowty et al., 1981) is the analysis of quantified sentences such as (1). (1) a. every child loves a toy b. no child likes brussel sprouts c. the father gave the child a drink the italicized determiners in (1) can be analyzed in terms of the existential and universal quantifiers of classical logic. barwise and cooper (1981); keenan and stavi (1986) introduced into linguistics the semantic treatment of generalized quantifiers, that is, a broader set of quantifiers which include those which cannot be treated in terms of the existential and universal quantifiers. examples of such quantifiers are italicized in (2). ∗. this research was supported in part by vr project 2009-1569, semantic analysis of interaction and coordination in dialogue (saicd). i would like to thank jonathan ginzburg and staffan larsson for useful discussion. this material is a revised and extended version of cooper (2010) and an earlier version was presented at the logic and language technology seminar in gothenburg. i would like to thank the audience for lively discussion, particularly dag westerståhl for pointing out an embarassing error. i would also like to thank three anonymous referees for detailed comments which have led to significant improvements in the paper. c©2013 robin cooper submitted 04/12; accepted 02/13; published online 03/13 cooper (2) a. most children eat baked beans b. few children are enthusiastic about vegetables c. many children have lots of friends peters and westerståhl (2006) provide an in depth study of generalized quantifiers, covering a large part of the literature on them. the leading idea of the semantic analysis of generalized quantifiers is that they express relations between sets. for example, (2a) expresses that the intersection of the set of children and the set of those who eat baked beans is a set which contains “most” children. what constitutes most can vary from context to context. the minimum requirement seems to be that the intersection contain more than half of the children. (2b) says that the intersection of the set of children and the set of those who are enthusiastic about vegetables is a set containing “few” members where again what counts as few is determined by context (including the nature of the predicates compared). similar remarks can be made about many. an advantage of this proposal for the non-classical quantifiers which are not definable in terms of the existential and universal quantifiers is that the same view in terms of comparison of sets can be taken of the classical quantifiers as well. thus the indefinite article corresponding to an existential quantifier can be analyzed as requiring that the intersection between the two sets is non-empty. its negation no can be regarded as requiring that the intersection is empty. the universal quantifier (represented by every) requires that all the members of the first set are in the intersection (that is, the first set is a subset of the second set). the generalized quantifier view of the comparison of sets seems intuitively psychologically appealing when we are considering complete sentences. however, a question arises as to what the utterance of a quantified noun-phrase on its own should represent. this becomes relevant when we consider clarification requests in dialogue which consist of a single noun-phrase. the classical proposal for this in the semantics literature is that noun-phrases represent families of sets, that is the set of sets which enter into the appropriate relation with the set represented by the common noun-phrase. a variant of this proposal (originating with montague) is that noun-phrases represent functions from properties (that is, in montague’s terms functions from possible worlds to sets) to truth-values: the characteristic function of a set of properties. such proposals make the compositional semantics of sentences containing quantified noun-phrases work out, but seem less attractive as candidates for the content of lone noun-phrases. a dialogue participant who utters a reprise clarification request many children? seems to be raising issues about children, or the number of children in a certain context, and not about a set of sets or a function from properties to truth-values. a standard reply that classical semanticists might give to this kind of objection is that the meanings we assign to constituents are not meant to represent what people are talking about directly but are rather a mathematical model which enables us to obtain compositionally meanings which model the kind of truth conditions and entailments which speakers associate with utterances. while this may be a reasonable answer in the context of a semantics oriented towards the analysis of complete sentences, it still leaves open the question of what a speaker who utters a sentence fragment during the course of a dialogue (e.g. a lone noun-phrase as a clarification request) is actually talking about. purver and ginzburg have addressed this problem in a series of works (purver and ginzburg, 2004; ginzburg and purver, 2008; ginzburg, 2012). they propose an alternative semantics for noun-phrases which exploits the notion of witness set from barwise and cooper (1981). intuitively a witness set is a set which will make a quantifier into something true. for example, a witness set for many children is a set which contains many children: the set of people who smoke may not 2 clarification and generalized quantifiers be such a set whereas the set of people who sleep at least eight hours a night may be such a set. this would mean that the sentence many children smoke would be false whereas the sentence many children sleep at least eight hours a night would be true. on this kind of analysis the content of a lone noun-phrase many children would be based on the notion of witness set. this seems more intuitive when we ask what the utterer of a lone noun-phrase is actually talking about. the answer is “a set containing many children”. in the clarification request perhaps the speaker is asking for a characterization of some set or sets which contain many children. we will explore the details of purver and ginzburg’s proposal below. in connection with this analysis purver and ginzburg introduce a principle of interpretation for reprise utterances, that is utterances which repeat part of the preceding utterance. they propose the reprise content hypothesis (rch) and use it to support their proposal for the analysis of quantified noun-phrases. rch comes in two versions and is stated in ginzburg (2012) as rch (weak) a fragment reprise question queries a part of the standard semantic content of the fragment being reprised. rch (strong) a fragment reprise question queries exactly the standard semantic content of the fragment being reprised. they argue for the strong variant and then use this to draw consequences for the semantic content of quantified noun phrases in general, claiming that this provides a strengthening of the constraints placed on semantic interpretation by compositionality. this represents an important step in connecting the traditional concerns of formal semantics with the more recent concerns concerning clarification which have arisen in dialogue semantics. this paper is an attempt to show that the connection might be even tighter that purver and ginzburg originally suggested. in cooper (2010) we argued that a more classical generalized quantifier analysis, recast in terms of type theory with records, accounts in an explanatory way for certain aspects of the reprise clarification data that purver and ginzburg cite. in this paper we will repeat those arguments and in addition argue for an account that unifies the analysis of purver and ginzburg and cooper (2010). we will first consider (in section 2) the anatomy of generalized quantifiers, presenting first a ttr version of a classical generalized quantifier analysis and the latest version of the purver and ginzburg proposal presented by ginzburg (2012) which we will revise slightly in order to account for monotone decreasing and non-monotone quantifiers. we will show how the two approaches can be unified into a single analysis. we will then look at some theoretical possibilities for how generalized quantifiers might be clarified (section 3). we will review the data concerning the clarification of quantifiers that purver and ginzburg presented (section 4) concentrating mainly on the kinds of clarifications they involve. (purver and ginzburg concentrated on the types of clarification requests.) we will conclude that the data are consistent with the hypotheses introduced in the previous section. finally (in section 5), we will draw some conclusions about the nature of rch. 2. the anatomy of generalized quantifiers the anatomy of quantified propositions can be characterized using ttr (cooper, 2005, 2012) and the analysis of non-dynamic generalized quantifiers presented in cooper (2004) as the type in (3). 3 cooper (3)  restr : ppty scope : ppty cq : q(restr,scope)  (3) is a record type. it has three fields: • the first field corresponds to the restriction or first argument of the quantifier, represented by the common noun phrase following the determiner in a noun phrase. in a record which is of the type (3) there must be a field with the same label, ‘restr’ containing an object of the type ppty, that is a property. we will spell out what we mean by property below. • the second field represents the scope or second argument of the quantifier, corresponding to the verb phrase in a sentence when the quantifier is in subject position. this is also required to be a property. • the third field represents a constraint requiring that a certain quantifier relation q hold between the two properties. for example, if q is the existential quantifier relation (corresponding to the english determiner a or singular count some) then the relation will hold just in case there is an object which has both the restriction and the scope property. q(restr,scope) also represents a type. we can think of it as the type of witnesses for the quantifier relation holding between the two properties. so in the case of the existential a witness would be something which has both properties ‘restr’ and ‘scope’. if there is no such object then this type will be empty.1 we use the idea from intuitionistic type theory (martin-löf, 1984) that the intuitive notion of proposition is represented by a type (known under the slogan “propositions as types”).2 the idea is that the proposition is “true” if there is something of the type and “false” if the type is empty. a particular example of a quantified proposition will be a refinement of the type in (3). if we represent the property of being a thief informally as ‘thief ’, the property corresponding to broke in here last night as ‘bihln’ and the existential quantifier relation as ∃, then the type representing the “proposition” corresponding to a thief broke in here last night would be (4). (4)  restr=‘thief ’ : ppty scope=‘bihln’ : ppty c∃ : ∃(restr,scope)  where the values in the ‘restr’ and ‘scope’ fields are restricted to be ‘thief’ and ‘bihln’ respectively.3 1. thinking of q(restr,scope) as a type of witness objects goes against claims that i have made orally and in work in progress that types constructed with predicates (so-called “ptypes”) should be thought of as types of situations. i believe that the discussion in this paper could be recast in terms where quantificational types like q(restr,scope) are construed as situation types, but will leave this for future research. 2. this idea is also known as the curry-howard isomorphism. 3. technically, the type ppty has been restricted to be the singleton type which contains exactly the property ‘thief ’ in the restriction field and the property ‘bihln’ in the scope field. the notation used in these fields, [ `=a:t ] , is used as a convenient way of representing [ `:ta ] , where if t is a type and a is an object of any type, then ta is a type such that b : ta iff b : t and b = a. (note that this is a revision of the notion of singleton type that we have used previously in, for example, cooper, 2012.) 4 clarification and generalized quantifiers we make this precise by using the relations between sets from classical generalized quantifier theory as presented, for example, in barwise and cooper (1981). we take ppty to be an abbreviation for a function type, as given in (5). (5) ppty abbreviates the type ( [ x:ind ] →rectype) that is, the type of functions from records with a field labelled ‘x’ for an individual to record types (corresponding to the intuitive notion of proposition). in order to relate properties to sets we first introduce a notation for the set of objects of a type. (6) the extension of type t , [̌t ], is the set {a | a : t}. if p is a property, that is p :ppty, we will talk of the set of objects which have p . we will call this the property extension, or p-extension, of p . the definition of this uses the notion of the extension of a type. (7) the p-extension of property p , [↓p ], is the set {a | ∃r[r : [ x:ind ] ∧ r.x = a ∧ [̌p (r)] 6= ∅]}. that is, the p-extension of p is the set of objects a which occur in the x-field of some record r4 such that the extension of the type p (r) is non-empty. suppose that q is a generalized quantifier relation between properties. we will use q∗ to represent the corresponding relation between sets from classical generalized quantifier theory. we require (8). (8) the type q(p1, p2) is non-empty iff the relation q∗ holds between [↓ p1] and [↓p2]. for example, if q is the existential quantifier relation, this will require that q(p1, p2) is a non-empty type just in case the p-extensions of the two properties have a non-empty overlap, [↓p1]∩[↓p2] 6= ∅. we need to go further and say what the objects of type q(p1, p2) are. (9) is an option that will work for all conservative quantifier relations. to say that a quantifier relation q is conservative means that q∗ will hold between two p-extensions [↓ p1] and [↓ p2] just in case it holds between [↓p1] and the intersection of [↓p1] and [↓p2]. we will assume with peters and westerståhl (2006) that all natural language quantifier relations are conservative.5 now we can make a precise proposal for what the objects of the type with the quantifier relation will be: (9) a : q(p1, p2) iff q∗ holds between [↓p1] and [↓p2] and a = [↓p1] ∩ [↓p2] (9) says that an object a will be of the type q(p1, p2) just in case the classical quantifier relation q∗ holds between the p-extensions of the two properties and in addition a is the intersection of those p-extensions. this definition builds on the notion of witness set from barwise and cooper (1981). it claims that for any p1 and p2 the objects of the type q(p1, p2) are witness sets for the quantifier. now let us return to the record type (4). the type theory will require that an object will be of this type if it is a record containing at least three fields with the labels in the type (labels may only 4. the notation r.x refers to the object in the x-field of r. 5. this means that only is not considered to represent a quantifier relation. 5 cooper occur once in a record or record type) and values in those fields of the types required by the record type. thus if there is no witness for the quantifier constraint type ‘c∃’ then there will not be anything of the record type (4) either. what makes this crucially a generalized quantifier approach to quantified propositions is the use of the quantifier relation which holds between two properties. there is another aspect to montague’s treatment of noun-phrases which is sometimes referred to as a generalized quantifier approach (by purver and ginzburg, among others). this is the psychologically troubling aspect of the treatment of generalized quantifiers discussed on p. 2f. this involves λ-abstraction over properties in the compositional treatment of noun phrase interpretations according to montague’s style (montague, 1974, chapter 8: ‘the proper treatment of quantification in ordinary english’). so, for example, the content of the noun phrase a thief will be a function from properties to record types where the scope field has been abstracted over: (10) λp :ppty (  restr=‘thief ’ : ppty scope=p : ppty c∃ : ∃(restr,scope) ) the use of the λ-calculus in (10) can be regarded as a kind of glue to get the compositional semantics to work out. (this is the kind of view of the λ-calculus as a glue language which is presented by blackburn and bos, 2005.) if you have another way to engineer the compositional semantics then you could abandon the λ-abstraction used in (10) but still use the generalized quantifier notion of relations between sets. now let us consider (11). (11) q-params: [ x:{ind} r:most(x,student) ] cont:left(q-params.x)  this a representation for most students left proposed by ginzburg (2012) in his ttr recasting of the purver and ginzburg approach to quantification. it requires that there be a set ‘x’ which is what barwise and cooper (1981) would call a witness set for the quantifier ‘most(student)’, that is, some set containing most students. in addition it requires that the predicate ‘left’ holds (collectively) for that set (which we may interpret as the predicate ‘left’ holding individually of each member of the witness set).6 this analysis is, then, also a generalized quantifier analysis. it differs from the previous one in that it emphasizes the witness set and uses a different relation between sets for the quantifier relation, namely a relation between a witness set and the set corresponding to what we called the restriction previously. the witness quantifier relation is ‘most’ in (11). this analysis works well for monotone increasing quantifiers.7 however, as purver and ginzburg (2004) point out, it is more problematic with monotone decreasing quantifiers8 since there you have to check the 6. note that the notion of witness set for a quantifier introduced by barwise and cooper is related though slightly different from the notion of witness set for a quantified sentence which we discussed above. the witness sets for a quantifier are potentially witnesses for the whole quantified sentence. a witness set for the sentence will be a witness set for the quantifier, but a witness set for the quantifier will not necessarily be a witness set for the sentence. 7. a monotone increasing quantifier relation q is one such that q(p1, p2) implies q(p1, p3) if [↓p2] is a subset of [↓p3]. 8. a monotone decreasing quantifier relation q is one such that q(p1, p2) implies q(p1, p3) if [↓ p3] is a subset of [↓p2]. 6 clarification and generalized quantifiers witness set against the restriction and the scope in a different way. in that paper they go through a number of different options for solving the problem, finally coming to a preference for treating monotone decreasing quantifiers as the negation of monotone increasing quantifiers, partly because it facilitates a treatment of complement anaphora. that suggests to me that the representation for few students left corresponding to the analysis in (11) should be something like (12). (12) c:¬( q-params: [ x:{ind} r:many(x,student) ] cont:left(q-params.x) )  that is, something that requires that there is no set x containing many students such that x (collectively) left. now the only way i can think of to engineer the compositional semantics to achieve (12) is to have something along the lines of (13) corresponding to the noun phrase. (13) λp :ppty ( c:¬( q-params: [ x:{ind} r:many(x,student) ] cont:p (q-params.x) ) ) but this involves exactly the montagueesque λ-paraphernalia that purver and ginzburg wish to avoid. however, as before, if you have an alternative way of engineering the compositional glue then you can apply it here as well while still maintaining the anatomy of quantification based on the witness quantifier relation. a curiosity here, though, is that the negation and the λ-abstraction would have to include the q-params field in their scope rather than being within the content field. this would be the only way in which the negation could get wider scope than many. but q-params according to ginzburg (2012) is at the top level of the sign, on the same level as phonology. it seems like it would be preferable to have an analysis where the negation is contained within the content-field. perhaps more difficult is the fact that the purver-ginzburg analysis also has difficulties with nonmonotone quantifiers such as an even number of students where it is not so clear that the negation strategy is available. this separation of the glue function of the λ-calculus and the analysis of quantified utterances in terms of generalized quantifier relations between sets leads me to suppose that purver and ginzburg’s objection is not so much to generalized quantifiers as such as to the use of montague’s λ-calculus based approach to compositional semantics. for these reasons i think the following variant of purver and ginzburg’s analysis, which fixes the monotonicity problem, might be considered as a friendly amendment which is consonant with their aims of emphasizing the role of witness sets. for quantifier predicates q we introduce a new predicate q† which takes a single property argument. q†(p ) is the type of witness sets for the quantifier in the sense of barwise and cooper (1981). the objects belonging to q†(p ) are characterized in (14). (14) a : q†(p ) iff a ⊆ [↓p ] and q∗ holds between [↓p ] and a. our proposed revision to the purver and ginzburg analysis is given in (15). (15) [ q-params : [ w : q†(p1) ] cont : [ cq=q-params.w : q(p1, p2) ] ] 7 cooper according to this most students left will correspond to (16) (using ‘student’ and ‘left’ as informal representations of the appropriate properties). (16) [ q-params : [ w:most†(student) ] cont : [ cq=q-params.w:most(student,left) ] ] it is easy to see that the quantifier anatomies presented in (3) and (15) are not really alternatives that conflict with each other. a first attempt to combine them is the more detailed anatomy given in (17). (17)  q-params: [ w:q†(cont.restr) ] cont : restr :ppty scope :ppty cq=q-params.w:q(cont.restr,cont.scope)   the advantage of (17) over (3) is that it makes explicit a role for a witness set which is determined only by the quantifier relation and a single property. this becomes important for the interpretation of noun-phrases where the second property is not determined. our preference is to represent the unsaturated nature of an np using montague’s λ-technology. if there is another way to treat compositionality it could presumably be used with the analysis in (17). however, (17) is not quite correct in that it claims that the content of the utterance is a record providing two properties and a witness set for the quantifier relation holding between them. this would not sufficiently distinguish intuitively different contents. consider (18). (18) a. two students left b. few students left both (18a) and (18b) could be true in a domain where there are fifty students and exactly two of them left, say margaret and billy. the set {margaret, billy} is an appropriate witness set for both (18a) and (18b) being the intersection of the set of students and the set of leavers and being a set which contains two students as well as being a set which contains few students. thus the records corresponding to cont in (17) would be the same (even though the types we are assigning to them are distinct), although intuitively the two sentences have different content, though they can be used to describe the same situation. the content should be the record type rather than the record. in general if we have been tempted to have a field as in (19) (19) [ ` : t ] and discover that the object in the `-field of a record of this type should be t itself rather than an object of type t , we should change (19) to (20). (20) [ `=t : type ] following this strategy on (17) would yield (21), which, however, introduces another problem. (21)  q-params: [ w:q†(cont.restr) ] cont= restr :ppty scope :ppty cq=q-params.w:q(cont.restr,cont.scope) :rectype  8 clarification and generalized quantifiers (21) is not a well-formed type since the path ‘cont.restr’ which is used in the ‘q-params’-field and the embedded type is no longer defined. our move to a manifest field requiring the type itself as value makes the labels unavailable within the larger type. however, it is possible to use paths from the larger type to define the type in the manifest field. a solution to this is to place the restriction field in the ‘q-params’-field and adjust path-names accordingly, as in (22). (22) q-params: [ restr:ppty w :q†(q-params.restr) ] cont= [ scope :ppty cq=q-params.w:q(q-params.restr,scope) ] :rectype  this example shows the disadvantage of the inexact abbreviatory notation for dependent fields that we are using. within the ‘cq-field’ in the embedded type the lables ‘q-params.w’ and ‘qparams.restr’ refer to the larger embedding type whereas ‘scope’ refers to the embedded type but there is nothing to signal this in the notation. in the more explicit (but less readable) representation explained in cooper (2012) (22) would be (23). (23)  q-params: [ restr:ppty w :q†(q-params.restr) ] cont:〈λv1:{ind} (λv2:ppty (rectype scope:ppty cq :〈λv3:ppty (q(v2,v3)v1), 〈scope〉〉 )), 〈q-params.w,q-params.restr〉〉  while (23) is the correct full representation and should remain as the official notation it is not exactly perspicuous. a compromise perhaps is to use ‘⇑’ before a path-name to indicate that it “takes its scope” in the next higher embedding record type, as in (24). (24) q-params: [ restr:ppty w :q†(q-params.restr) ] cont= [ scope :ppty cq=⇑q-params.w:q(⇑q-params.restr,scope) ] :rectype  our proposal for the analysis of the noun-phrase most students is thus (25) (using ‘student’ as an abbreviation for the property of being a student and quant as the type (ppty→rectype), the type of quantifiers, corresponding to nps). (25)  q-params: [ restr=student:ppty w :most†(q-params.restr) ] cont= λp :ppty ( scope=p :ppty cmost=⇑q-params.w:most(⇑q-params.restr, scope) ):quant  9 cooper our intention in having ‘q-params’ external to the content is that it should represent a kind of referential reading (donnellan, 1966), perhaps related to what barwise and perry (1983) would call a value-loaded reading, where the witness set and the restriction property is fixed by context.9 the idea would be then that in compositional interpretation the information in the ‘q-params’-field should percolate up to the top. in order to achieve this the ‘q-params’-field associated with a constituent such as a verb-phrase or a sentence should be of the type which is the result of merging the types required for ‘q-params’-fields of its daughters. this means that the labels ‘restr’ and ‘w’ need to be changed to provide unique identifiers in order not to clash with ‘q-params’ information coming from other noun-phrases. for now we will assume that each noun-phrase comes along with a unique identifier i which can be subscripted to the labels. thus (25) will become (26). (26)  q-params: [ restri=student:ppty wi :most†(q-params.restri) ] cont= λp :ppty ( scope=p :ppty cmost=⇑q-params.wi:most(⇑q-params.restri, scope) ):quant  this view of q-params will require that (26) is not the only reading associated with most students. consider examples where it is very hard or impossible to obtain referential (or even de re) readings for embedded noun-phrases. (27) a. sam doesn’t claim that most students are enthusiastic b. everybody claimed that most students are enthusiastic just as it seems very hard to give most students wide scope in (27) in order to obtain a de re reading (“most students are such that sam doesn’t claim that they are enthusiastic”, “most students are such that everybody claimed that they are enthusiastic”), it seems hard to have a wide scope existential quantification over a witness set for the quantifier (“there is a set containing most students such that sam doesn’t claim that they constitute a proof that most students are enthusiastic”, “there is a set containing most students such that everybody claimed that they are enthusiastic”). compare this with the corresponding affirmative sentence in (28). (28) sam claims that most students are enthusiastic here there does seem to be an intuitive de re reading: “most students are such that sam claims that they are enthusiastic”. this is a reading where sam need not be willing to affirm that most students are enthusiastic. she can make individual claims about a number of students, that they are enthusiastic, without having made a claim that most students are enthusiastic. similarly, though 9. this is different to what is proposed in ginzburg (2012) where the ‘q-params’-field is external to the content at the noun-phrase level but becomes part of the content by compositional processes at the sentence level, thus obtaining a non-referential reading. techniques for wide-scope interpretation (such as storage) would be needed in order to get the referential reading of the noun-phrase in the sentence. note also that we are generalizing the notion of referential reading to apply to more than just definite descriptions. a thorough discussion of referential readings involving witness sets in this way is beyond the scope of this paper. 10 clarification and generalized quantifiers we are less used to talking about it, there seems to be a reading: “there is a set containing most students such that sam claims that they constitute a proof that most students are enthusiastic. this is a reading where there is some particular set containing most students about which sam makes a particular claim, namely that it is a witness set for most students are enthusiastic. on such a reading it is possible to continue the sentence with (29) (29) . . . namely all the first years registered for the course this is not a possible continuation for (27a). a proposal for the non-referential (attributive or “value-free”) reading is (30). (30)  q-params:rec cont= λp :ppty (  restri=student:ppty wi:most†(restri) scope=p :ppty cmost=wi:most(restri,scope) ):quant  where rec is the type of records with no constraints on their fields.10 both (26) and (30) are meant to be part of larger sign types of the kind proposed by ginzburg (2012), for example, as sketched in (31). 10. this type can also be represented as ‘[ ]’. 11 cooper (31) a. referential reading phon:“most students” cat=np:cat . . . q-params: [ restri=student:ppty wi:most†(q-params.restri) ] cont= λp :ppty ( scope=p :ppty cmost=⇑q-params.wi:most(⇑q-params.restri, scope) ):quant  b. non-referential reading phon:“most students” cat=np:cat . . . q-params:rec cont= λp :ppty (  restri=student:ppty wi:most†(restri) scope=p :ppty cmost=wi:most(restri,scope) ):quant  the details of the rest of the sign type (indicated by ‘. . . ’ in (31)) are not important for the current discussion, except that we assume that there will either be a constituents-field as discussed in ginzburg (2012) or a more traditional daughters-field as in classical hpsg as discussed in a ttr setting in cooper (2008). 3. potential clarifications we introduce a hypothesis concerning potential responses to clarification requests: (32) clarification request response hypothesis a. a response to a clarification request must address a path in the type corresponding to content of the clarification request b. there is a strong tendency for a response to a clarification request to be a major constituent such as a noun-phrase or sentence. our main hypothesis is that what can be addressed by a clarification in response to a clarification request are paths within the type corresponding to the content of the clarification request. from a theoretical point of view this is natural since paths in a sign can be regarded as parameters of an utterance whose values can be questioned. (33) shows the paths represented in (31). 12 clarification and generalized quantifiers (33) a. referential reading phon cat . . . q-params q-params.restri q-params.wi cont b. non-referential reading phon cat . . . q-params cont the ‘. . . ’ correspond to what we have not made precise in (31). in this will be included paths to the constituents of the noun-phrase (for example, if we adopt the daughters proposal of cooper 2008 corresponding to the daughters feature in standard hpsg, see, for example, ginzburg and sag 2000). according to our hypothesis these paths represent what can be addressed by an answer to a clarification question.11 note that we make different predictions for the referential and nonreferential readings. namely, that on the referential reading the restriction and witness are available whereas they are not on the non-referential reading. it is a significant advantage of this analysis that we can distinguish referential and non-referential readings of a lone noun-phrase and that it is not, for example, dependent on the scope that it takes within a complete sentence. the scope of the quantifier (corresponding to the verb-phrase if the noun-phrase were to be placed in the subject position of a sentence) cannot be addressed as it does not lie on a path within the type. the witness role can be addressed on a referential reading but not a non-referential one. the restriction, if it corresponds to a constituent, can be addressed on either reading, whereas the restriction in a noun-phrase without a constituent corresponding to the restriction (e.g. everybody) can only be addressed on the referential reading (via the q-params:wi-path). this makes some precise and intuitive theoretical predictions, although it can be difficult to tease apart the readings on the basis of actual data which greatly underdetermine the interpretation. one problem, of course, is that the notion of referentiality of quantified noun-phrases is notoriously slippery. added to this is the fact that in a dialogue game-board analysis there is no requirement that both dialogue participants have the same interpretation or, if we take underspecification into account, even have decided individually on a particular interpretation. viewing interpretation from the perspective of uncertainty for dialogue participants can actually help us understand why referentiality can seem so problematic when viewed in terms of classical non-dialogical semantics. consider (34). 11. it seems to us unintuitive that the q-params path would be addressed, that is the collection of all the q-params as opposed to the individual q-params represented by the paths q-params.restri and q-params.wi which are provided in the referential reading but not in the non-referential reading. this, together with the unique subscripting needed on the labels within q-params, suggests to us that the q-params field is ultimately not quite what is needed for this analysis. but as of yet we do not have an alternative analysis to offer. 13 cooper (34) a: a thief broke in here last night b: a thief? a: a. my ex-husband, actually (witness) b. burglar wearing a mask (restriction) c. got in through the bedroom window (scope) d. two thieves, actually (content) there is nothing in a’s first utterance which indicates to b whether a thief is to be referentially interpreted or not. there may of course be other circumstances which indicate this which are not recorded here. a may be pointing at a particular person or a photograph. there may have been previous dialogue which made it perfectly clear that a knows who it was who broke in. but if this is the beginning of the dialogue and there is no pointing or a sufficiently rich context, then b cannot know whether a’s use of a thief is referential or not, that is, b cannot know whether a has an indepedent way of identifying the thief, as given by the instantiated q-params. b’s clarification request can be seen as a step towards obtaining information that will indicate whether there is support for a referential reading or not. certain answers to the clarification request will require a referential reading. at the point at whichb utters the clarification requestb does not know whether a’s utterance of a thief was intended referentially or not. similarly, a cannot know whether b’s clarification request was intended referentially or not, although given that b is asking the question it seems reasonable to assume that b does not have an independent way of identifying the thief.12 however, what a is clarifying is her original use of the noun-phrase and she is therefore at liberty to take this on either a referential or a non-referential reading. (34a) addresses the issue of an appropriate witness for the noun-phrase that b presents in the clarification request. as such it requires a referential reading of the noun-phrase. at this point in the dialogue b knows that a has an independent way of identifying the thief. this does not mean that the original intended interpretations of the clarification request or the initial utterance by a were referential – only that it is possible to construe them as referential, that is, that there is support for the claim that a has an independent way of identifying the individual in question. compare (34a) with (35) where a referential reading is unavailable for the initial utterance. (35) a: sam doesn’t think that a thief broke in here last night b: a thief? a: my ex-husband, actually (35) is hard to interpret as a coherent dialogue. a’s second utterance can at best be interpreted as ignoring b’s clarification request and suggesting her own theory of who she thinks or knows broke in. the negation makes all the difference here. compare it with (36). (36) a: sam thinks that a thief broke in here last night b: a thief? a: my ex-husband, actually here a’s response to b’s clarification request can be interpreted as a clarification of the belief which a is attributing to sam. it is at this point in the dialogue that we realize that a thief in a’s 12. except for the perhaps unusual context where b knows perfectly well who it is but is questioning the use of the description thief for this person. 14 clarification and generalized quantifiers first utterance must be interpreted referentially (even though it may be de dicto, apparently giving us the conclusion that in a’s view sam believes that a’s ex-husband is a thief). (34b) involves a clarification of a content corresponding to a syntactic constituent (thief ) and therefore, according to our hypothesis, does not affect referentiality. however, the second part of our hypothesis (32b) comes into play. the common noun phrase presented is not a major constituent and this seems to make it difficult to interpret as a response.13 the difficulty becomes even clearer when we embed it beneath a negated attitude verb as in (37). (37) a: sam doesn’t think that a thief broke in here last night b: a thief? a: burglar wearing a mask it is hard, though perhaps not impossible, to interpret a’s second turn as denying that sam has a de dicto belief “a burglar wearing a mask broke in here last night”. the reading we are after seems much facilitated by (38) where the negation and the whole noun-phrase is repeated. (38) a: sam doesn’t think that a thief broke in here last night b: a thief? a: not a burglar wearing a mask, at any rate such readings seem easier to obtain in the non-negated case in (39). (39) a: sam thinks that a thief broke in here last night b: a thief? a: burglar wearing a mask (34c) seems quite difficult to interpret as a response to the clarification question.14 while the scope of the quantifier in a’s first utterance corresponds to a constituent verb-phrase broke in here last night it does not correspond to a constituent of the clarification request and neither is there a path in the sign-type of the clarification request which corresponds to the scope. this is not to say that dialogues of this form may not occur, but if they do, the turn following the clarification would not be interpreted as a response to the clarification request, rather as a continuation which ignores the clarification request. (34d) addresses the entire content of the noun-phrase, available on a path in the sign-type and also a constituent in the sense that it corresponds to the complete clarification request. however, it updates only the quantifier relation and determiner constituent in that it just repeats the restriction. two without thieves is also possible but this should probably be interpreted as providing a complete 13. as one of the reviewers points out, this condition does not play a role if the clarification question itself is not a major constituent, for example just the common noun thief. this might suggest a tendency for the category of the clarification to match the category of the clarification question. 14. one of the reviewers pointed out that the clarification question might be interpreted as a request for a repeat or clarification of what followed the words a thief. in this case (34c) is an appropriate response. such a clarification request is addressing the phonology (indicating that the speaker didn’t hear the rest of the utterance). it’s not clear to me whether there is a special intonation associated with such questions or whether dialogue participants rely on gestural cues to indicate this kind of question. languages seem to differ in the degree to which such questions are available. my intuition is that such questions are more frequent in swedish than in english, for example. this could be checked in dialogue corpora by investigating the answers given to noun-phrase reprise questions. 15 cooper noun-phrase (quantifier) rather than a determiner (quantifier relation) in line with (32b). if the determiner is not one that is capable of being a complete noun-phrase the response becomes much less acceptable. thus the examples in (40) sound downright ungrammatical as responses to the clarification request. (40) a. the, actually b. every, actually these would have been fine if the clarification request had been did you say “a thief”? with emphasis on a, where it is explicit that the phonology is being addressed. we can compare the examples of responses to reprise clarification requests to examples of nonreprise clarification requests relating to a quantifier as given in (41). (41) a: somebody broke in here last night b: a. (not) your ex-husband? (witness) b. burglar wearing a mask? (restriction) c. got in through the bedroom window? (scope) d. just one person? (content) much the same remarks hold for these clarification requests as for the responses to the reprise clarification requests. the witness clarification request seems to force a referential reading, not on a’s original utterance but at the point at which the clarification request is uttered. a can go on to deny that a referential reading was intended by saying something like (42). (42) all i know is that there was a noise in the kitchen the restriction clarification request sounds kind of elliptical and a repetition of the determiner a would probably be preferable in accordance with (32b) adjusted to apply to such clarification requests. it is also the case that there is no common noun constituent which this clarification request can address. it again seems to force a referential reading (since that is the only way that the restriction can be made available) and a is at liberty to deny that such a reading is intended as in the witness case. but if a answers yes to this question then there is a commitment to a referential reading on the part of both participants at this point even if the intended original utterance by a was a non-referential reading. our hypothesis would predict that the scope clarification request is more available since, in contrast to a response to the noun-phrase clarification request of the previous examples, the scope is available as a constituent. my intuition is that it is more available although the vp-constituent does not appear to count as a major constituent and there appears to be a preference for a complete sentence as in (43) to express this (and a corresponding question as to whether it should count as a clarification request at all). (43) did they get in through the bedroom window? finally the clarification request addressing the content focusses on the quantifier relation and can be expressed as just one?. as before it seems that this is to be seen as a complete noun-phrase rather than just a determiner. 16 clarification and generalized quantifiers this suggests that the hypothesis (32) could be generalized to include this kind of clarification. in addition there seem to be other data that behave in a similar way. (44) presents examples of apposition, parentheticals or repairs.15 how we classify them seems to depend on the semantic relation between the original noun-phrase and its appositive. (44) a. a thief, my ex-husband, actually, broke in here last night (witness, appositive) b. a thief, ??(a) burglar wearing a mask, broke in here last night (restriction, appositive/parenthetical) c. ?∗a thief, got in through the bedroom window, broke in here last night (scope) d. ?a thief, he got in through the bedroom window, broke in here last night (scope, parenthetical) e. a thief, two thieves, actually, broke in here last night (content, repair) the judgements discussed above are rather subtle and there are a number of predictions that are difficult to test against real data. however, if we compare these examples with (45) and (46), where the potential clarifications and clarification questions are not directly addressing an available path associated with the previous turn, there is a robust intuition that you have to work a lot harder or be embedded in a rich context to interpret them as coherent. (45) a: a thief broke in here last night b: a thief? a: a. maroon b. maroon sweater c. police d. scar over the left eye (46) a: somebody broke in here last night b: a. maroon? b. maroon sweater? c. police? d. scar over the left eye? it is hard to give examples of impossible dialogues since there is no notion of grammaticality as there is with single sentences. what we can examine is the most likely interpretation given what we gather about the context from what we know about the dialogue. for example, let us consider how we might interpret (46). (46a) seems hard to interpret at all unless, for example maroon is being used (innovatively) as a way of characterizing skin-colour, in which case it would be a clarification relating to the restriction, though, of course, non-preferred since it does not represent a complete noun-phrase. a natural way of interpreting (46b) would be as elliptical for wearing a maroon sweater which would in effect coerce it to be a clarification of the restriction, again non-preferred. depending on the political situation in the country the dialogue is about, (46c) might be interpreted 15. apposition was suggested by one of the reviewers of this paper. 17 cooper as a restriction clarification, i.e. was it the police who broke in?, or as a very elliptical way of asking whether a called the police. this latter interpretation could be facilitated, for example, if a and b routinely talked about break-ins and had a checklist of questions which they normally asked, among them whether the police was called. in this case, of course, (46c) would not be a clarification of the quantifier. finally, (46d) is most naturally interpreted as elliptical for with a scar over the left eye, making it as a clarification of the restriction. similar remarks can be made about the appositive/parenthetical/repair cases in (47). (47) a. a thief, maroon, broke in here last night b. a thief, maroon sweater, broke in here last night c. a thief, police, broke in here last night d. a thief, scar over the left eye, broke in here last night except for (47d) all these examples sound pretty incoherent unless embedded in a rich context of the kind we discussed for the clarification examples. this seems consistent with a view of parentheticals as a discourse level phenomenon rather than a syntactic phenomenon. a central question is to what extent similar facts can be observed about generalized quantifiers in general and whether different classes of quantifiers behave differently with respect to the availability of clarification interpretations. consider (48) a: most thieves are opportunists http://www.accessmylibrary.com/coms2/ summary_0286-33299010_itm, accessed 18th january, 2010 b: most thieves? a: a. successful ones (witness/restriction) b. bide their time (scope) c. 80%, actually (content, with focus on the quantifier relation) here the witness and restriction clarifications appear to collapse since a witness set for the quantifier has to be a subset of thieves (i.e. the restriction) which contains most thieves. however, (48a) does appear to be ambiguous between an interpretation corresponding to “successful thieves are opportunists” (a witness reading) and “most successful thieves are opportunists” (a restriction reading). thus while the form of the clarification is the same its interpretation is ambiguous between a witness clarification and a restriction clarification. however, there would seem to be a preference for the restriction clarification interpretation to be represented by the complete noun-phrase most successful ones since for the restriction clarification reading successful ones would have to be parsed as a common noun phrase. for the witness reading successful ones would be parsed as a complete noun-phrase. as with the previous examples we have discussed the scope response (48b) appears to be inappropriate as a response to the clarification request whereas addressing the whole content with focus on the quantifier relation is fine. 18 clarification and generalized quantifiers 4. towards an empirical investigation we have not attempted to make an empirical investigation of the phenomena we have discussed in this paper. rather we have tried to clarify some of the formal semantic issues and sharpen intuitions relating to the examples. the judgements are subtle and part of the point is that fine-grained semantic distinctions are left underdetermined in dialogue data. because of this, one may wonder whether there is any hope of doing empirical work in this area at all. it is not straightforward to find relevant examples in corpora and one has to rely on making subtle distinctions in interpretation which are not really possible to annotate. experimental methods, where you have more control over the data produced, may prove more tractable but again there is the problem that even the interpretation of data obtained in an experimental situation may be underdetermined. this notwithstanding, i will try to argue in this section that the prospects are not entirely bleak and will support this view by looking at the data that purver and ginzburg have already collected. let us try to summarize what we have from the previous discussion that could be investigated in a dialogue corpus. for a noun-phrase clarification request there are three kinds of clarification available: • witness • restriction, i.e. a common noun-phrase (dispreferred) • content (a complete noun-phrase possibly with restriction or quantifier relation focus) a restriction clarification in the form of a common noun phrase is dispreferred apparently because of a preference for clarifications to be “major” constituents. instead there is a tendency for the clarification to be a noun-phrase which addresses the complete content but may focus on either the restriction (common noun phrase, if there is one) or the quantifier relation (determiner, if there is one). the purver-ginzburg data divides into witness clarifications, one dubious case of a possible restriction clarification and content clarifications possibly focussed either on the restriction or the quantifier relation. the majority of the cases they cite are content clarifications focussed on the restriction. there are no examples of clarifications other than of these three types in their data. we give details of the relevant examples below. 4.1 witness clarifications (49) unknown: and er they x-rayed me, and took a urine sample, took a blood sample. er, the doctor unknown: chorlton? unknown: chorlton, mhm, he examined me, erm, he, he said now they were on about a slide 〈unclear〉 on my heart. mhm, he couldn’t find it. bnc file kpy, sentences 1005–1008 purver and ginzburg (2004) 19 cooper (50) terry: richard hit the ball on the car. . . . nick: what ball? terry: james [last name]’s football. bnc file kr2, sentences 862, 865–866 purver and ginzburg (2004) intuitively both of these examples appear to be witness clarifications, although one might argue that this status is unclear. (50) might be arguably a content clarification focussed both on the restriction (ball→football) and the quantifier relation if we analyze james [last name]’s as a determiner representing a quantifier relation. 4.2 restriction clarifications there is one example in their data which could be an instance of restriction clarification. (51) george: you want to tell them, bring the tourist around show them the spot sam: the spot? george: where you spilled your blood bnc file kdu, sentences 728–730 purver and ginzburg (2004) here we have interpreted the clarification as additional material which is to be added as a modifier to the restriction, that is the spot where you spilled your blood. however, another interpretation is available where where you spilled your blood is a noun-phrase in its own right or perhaps an embedded question as in show them where you spilled your blood. this would then make the example a case of content clarification (or perhaps even witness clarification). as this is the only potential example of a restriction clarification and our hypothesis is that restriction clarifications are dispreferred perhaps these interpretations are more likely. 4.3 content clarifications with restriction focus the most common kind of clarification in the data are ones where the entire noun phrase is repeated with the additional modifier inserted, as in the following examples: (52) terry: richard hit the ball on the car. nick: what car? terry: the car that was going past. bnc file kr2, sentences 862–864 purver and ginzburg (2004) 20 clarification and generalized quantifiers (53) anon 1: in those days how many people were actually involved on the estate? tommy: well there was a lot of people involved on the estate because they had to repair paths. they had to keep the river streams all flowing and if there was any deluge of rain and stones they would have to keep all the pools in good order and they would anon 1: the pools? tommy: yes the pools. that’s the salmon pools anon 1: mm. bnc file k7d, sentences 307–313 purver and ginzburg (2004) (54) eddie: i’m used to sa-, i’m used to being told that at school. i want you 〈pause〉 to write the names of these notes up here. anon 1: the names? eddie: the names of them. anon 1: right. bnc file kpb, sentences 417–421 purver and ginzburg (2004) (55) nicola: we’re just going to beckenham because we have to go to a shop there. oliver: what shop? nicola: a clothes shop. 〈pause〉 and we need to go to the bank too. bnc file kde, sentences 2214–2217 purver and ginzburg (2004) (56) is different in that it is the dialogue participant who contributes the original clarification request who provides alternative restrictions. note that in this case the restrictions do not correspond to a syntactic constituent in the original utterance (nothing). (56) anon 1: er are you on any sort of medication at all suzanne? nothing? suzanne: no. nothing at all. anon 1: nothing? no er things from the chemists and cough mixtures or anything 〈unclear〉? bnc file h4t, sentences 43–48 purver and ginzburg (2004) in (57) we have a case where a modifier in the original utterance is replaced by a new modifier in the clarification, thus changing what was said non-monotonically, not merely further specifying what was said. 21 cooper (57) elaine: what frightened you? unknown: the bird in my bed. elaine: the what? audrey: the birdie? unknown: the bird in the window. bnc file kbc, sentences 1193–1197 purver and ginzburg (2004) the whole of the restriction can be replaced in this way. (58) mum: what it ever since last august. i’ve been treating it as a wart. vicky: a wart? mum: a corn and i’ve been putting corn plasters on it bnc file ke3, sentences 4678–4681 purver and ginzburg (2004) even though a different noun is chosen to express the restriction it can nevertheless be a refinement of the original utterance. in (59) the natural interpretation is the director is a woman. (59) stefan: everything work which is contemporary it is decided katherine: is one man? stefan: no it is a woman katherine: a woman? stefan: a director who’ll decide. bnc file kcv, sentences 3012–3016 purver and ginzburg (2004) (60) seems to be a case where the speaker is searching for the right noun to express the restriction. (60) unknown: what are you making? anon 1: erm, it’s a doit’s a log. unknown: a log? anon 1: yeah a book, log book. bnc file knv, sentences 188–191 purver and ginzburg (2004) 4.4 content clarifications with quantifier relation focus in the data that purver and ginzburg present there appear to be two clear examples of content clarifications with quantifier relation focus. 22 clarification and generalized quantifiers (61) anon 2: was it nice there? anon 1: oh yes, lovely. anon 2: mm. anon 1: it had twenty rooms in it. anon 2: twenty rooms? anon 1: yes. anon 2: how many people worked there? bnc file k6u, sentences 1493–1499 (purver and ginzburg (2004) cite it without the last turn) we included the final turn to strengthen the interpretation that it is the quantifier relation which is being clarified. it seems hardly likely that the restriction rooms is in need of clarification. (62) marsha: yeah that’s it, this, she’s got three rottweilers now and sarah: three? marsha: yeah, one died so only got three now 〈laugh〉 bnc file kp2, sentences 295–297 purver and ginzburg (2004) 4.5 other content clarifications (63) is a little difficult to classify. (63) richard: no i’ll commute every day anon 6: every day? richard: as if, er saturday and sunday anon 6: and all holidays? richard: yeah 〈pause〉 bnc file ksv, sentences 257–261 purver and ginzburg (2004) it might be interpreted as if it involves a discussion of whether the restriction day is to mean weekdays or all days of the week and whether it is to include holidays. an alternative analysis might classify it as a clarification with quantifier relation focus, that is, a discussion as to whether it really is every day that is meant. 5. conclusion we have examined the nature of generalized quantifiers in the light of clarifications given as answers to clarifications which consists of a single noun-phrase. whereas purver and ginzburg focus on the nature of the clarification request, we focus on the nature of the clarification and thereby indirectly illuminate the content that should be associated with the clarification request. we agree with purver and ginzburg that the notion of witness should play a central role in explaining these interactions but have argued that their witness-based analysis of generalized quantifiers should be combined with a more classical approach and that this makes certain predictions about what can be addressed by the clarification. a review of the data that purver and ginzburg 23 cooper present seems to be consistent with our combined analysis concerning the availability of witness, restriction and whole content readings associated with the clarification. issues of referentiality are less clear, partly because the phenomenon is inherently underdetermined from the dialogue and the dialogue participants may have different views on the referentiality of a particular noun-phrase utterance – and indeed a dialogue participant may leave the referentiality of a noun-phrase underspecified. a clear prediction is that a witness clarification offers the possibility of interpreting a noun-phrase referentially. this appears to follow trivially from the nature of witnesses. perhaps there is not much more that can be said in terms of specific examples. what does this say about purver and ginzburg’s reprise content hypothesis? the analysis we have given is consistent with rch in that it seems to allow a fragment reprise question to query exactly the standard semantic content of the fragment being reprised. for the most part it appears to be this content which is being addressed in the clarification which responds to the fragment reprise question although there may be focus on either the restriction or the quantifier relation. the clarification may also provide a witness for this content. the fact that the clarification addresses parameters identified by paths in the clarification request does not of itself affect the rch (which is concerned only with reprises). thus the rch remains intact but what we consider to be the content of a noun-phrase reprise is richer than the content proposed by purver and ginzburg in that it combines both the witness-based analysis and the more classical analysis based on the quantifier relation. one may consider the clarifications (that is, the responses to the clarification requests) as responding to certain aspects of the noun-phrase content (witness, restriction, quantifier relation). thus while the strong version of the rch holds intact for the reprise clarification request, the response to this request may address part of the meaning of the reprise. one of the advantages of using ttr is that you get structured semantic contents. instead of the unstructured sets and functions of classical model theoretic semantics, you get articulated record types with labels pointing to various components. what the clarifications discussed in this paper seem to show is that speakers pick up on these meaning components even when they are not represented by separate syntactic constituents. it seems to me that this is an important part of a semantic theory of dialogue which should allow us to make detailed predictions about the nature of cohesion in dialogue. references jon barwise and robin cooper. generalized quantifiers and natural language. linguistics and philosophy, 4(2):159–219, 1981. jon barwise and john perry. situations and attitudes. bradford books. mit press, cambridge, mass., 1983. patrick blackburn and johan bos. representation and inference for natural language: a first course in computational semantics. csli studies in computational linguistics. csli publications, stanford, 2005. robin cooper. dynamic generalised quantifiers and hypothetical contexts. in ursus philosophicus, a festschrift for björn haglund. department of philosophy, university of gothenburg, 2004. url http://www.phil.gu.se/posters/festskrift/. 24 clarification and generalized quantifiers robin cooper. austinian truth, attitudes and type theory. research on language and computation, 3:333–362, 2005. robin cooper. type theory with records and unification-based grammar. in fritz hamm and stephan kepser, editors, logics for linguistic structures, pages 9–34. mouton de gruyter, 2008. robin cooper. generalized quantifiers and clarification content. in paweł łupkowski and matthew purver, editors, aspects of semantics and pragmatics of dialogue. semdial 2010, 14th workshop on the semantics and pragmatics of dialogue. polish society for cognitive science, poznań, 2010. robin cooper. type theory and semantics in flux. in ruth kempson, nicholas asher, and tim fernando, editors, handbook of the philosophy of science, volume 14: philosophy of linguistics, pages 271–323. elsevier bv, 2012. general editors: dov m. gabbay, paul thagard and john woods. keith s. donnellan. reference and definite descriptions. the philosophical review, 75(3):281–304, 1966. david dowty, robert wall, and stanley peters. introduction to montague semantics. reidel (springer), 1981. jonathan ginzburg. the interactive stance: meaning for conversation. oxford university press, oxford, 2012. jonathan ginzburg and matt purver. quantfication, the reprise content hypothesis, and type theory. in lars borin and staffan larsson, editors, from quantification to conversation. university of gothenburg, gothenburg, 2008. jonathan ginzburg and ivan a. sag. interrogative investigations: the form, meaning, and use of english interrogatives. number 123 in csli lecture notes. csli publications, stanford, california, 2000. e. l. keenan and j. stavi. natural language determiners. linguistics and philosophy, 9:253–326, 1986. per martin-löf. intuitionistic type theory. bibliopolis, naples, 1984. richard montague. formal philosophy: selected papers of richard montague. yale university press, new haven, 1974. ed. and with an introduction by richmond h. thomason. stanley peters and dag westerståhl. quantifiers in language and logics. oxford university press, 2006. matt purver and jonathan ginzburg. clarifying noun phrase semantics. journal of semantics, 21 (3):283–339, 2004. 25 brown-schmidt&hanna_final dialogue and discourse 2 (2011) 11-33 received 1/2010; first letter 6/2010; accepted 12/2010; doi: 10.5087/dad.2011.102 final revision 1/2011; published online 5/2011 ©2010 sarah brown-schmidt and joy e. hanna talking in another person's shoes: incremental perspective-taking in language processing sarah brown-schmidt brownsch@illinois.edu department of psychology and beckman institute university of illinois at urbana-champaign champaign, il, 61820, usa joy e. hanna joy.hanna@oberlin.edu department of psychology oberlin college oberlin, oh, 44074, usa editors: david schlangen, hannes rieser abstract language use in conversation is fundamentally incremental, and is guided by the representations that interlocutors maintain of each other’s knowledge and beliefs. while there is a consensus that interlocutors represent the perspective of others, three candidate models, a perspective-adjustment model, an anticipation-integration model, and a constraint-based model, make conflicting predictions about the role of perspective information during on-line language processing. here we review psycholinguistic evidence for incrementality in language processing, and the recent methodological advance that has fostered its investigation—the use of eye-tracking in the visual world paradigm. we present visual world studies of perspective-taking, and evaluate each model's account of the data. we argue for a constraint-based view in which perspective is one of multiple probabilistic constraints that guide language processing decisions. addressees combine knowledge of a speaker’s perspective with rich information from the discourse context to arrive at an interpretation of what was said. understanding how these sources of information combine to influence interpretation requires careful consideration of how perspective representations were established, and how they are relevant to the communicative context. keywords: conversation, eye-tracking, perspective, processing 1 introduction recent years have seen a surge of interest in how language is processed in real time and in real space. this interest has, in part, been stimulated by the advent of modern eye-tracking devices that can monitor moment-by-moment language processing in settings that mimic many aspects of real world language use. these experimental techniques have enabled researchers to address central questions about the incrementality of the language processing system. words occur over time, at an average of 150-200 words per minute (tauroza & allison, 1990; also see levelt, 1989, p. 22). as a result, many words and phrases are ambiguous, at least temporarily, causing a proliferation of possible syntactic and semantic interpretations as they accrue. thus a critical problem for theories of language processing is to explain how listeners build a representation of an on-going sentence when the meaning of many words is unclear until later in the sentence. tanenhaus (2004, p. 376) describes (and ultimately criticizes) an early view of language understanding as a catch-up game in which the listener held each word in a memory buffer, waiting for subsequent information that could disambiguate the words. instead, evidence from over a decade of studies using on-line methods have revealed that: (a) words are integrated brown-schmidt and hanna 12 immediately into the on-going sentence, with listeners making provisional commitments to interpretations of lexical (allopenna, magnuson, & tanenhaus, 1998; dahan, magnuson, tanenhaus, & hogan, 2001), referential (hanna, tanenhaus & trueswell, 2003; chambers et al., 2002), syntactic (tanenhaus, spivey-knowlton, eberhard, & sedivy, 1995), and other ambiguities; and (b) listeners make sophisticated predictions about upcoming material (altmann & kamide, 1999, 2007, 2009; de goede et al., 2009). this evidence has led to further work addressing questions about the timing with which various sources of linguistic and non-linguistic information, from lexical frequency and neighborhood effects (magnuson, dixon, tanenhaus, & aslin, 2007) to speaker identity (brown-schmidt, 2009a; metzing & brennan, 2003), are used during the incremental process of language use. in this chapter, we focus on the use of one type of non-linguistic information in incremental language comprehension—information about others’ knowledge and beliefs. as we shall see, studying perspective-taking in language processing requires methodologies that afford the use of naturally produced speech in rich, interactive contexts, while at the same time providing detailed information about the time-course of production and comprehension. in the last 10-15 years, the application of eye-tracking methodology to dialog tasks has stimulated significant debate and theorizing about when considerations of the speaker’s perspective play a role during processing. three primary candidate models of the role of perspective during on-line comprehension have emerged: a perspective-adjustment model (keysar, barr, balin, & paek, 1998), an anticipationintegration model (barr, 2008), and a constraint-based model (hanna, et al., 2003). as we shall see, the key differences between these models are in their proposals for when perspective guides interpretation decisions. in order to set the stage for our discussion of this issue, and our perspective on it, the remainder of this section will provide an overview of how words are interpreted incrementally in rich contexts. we then turn to the question of how knowledge of speaker perspective modulates this process. much of the support for incremental interpretation of language comes from studies that use the visual world paradigm (vwp) developed by tanenhaus and his colleagues (tanenhaus, et al., 1995; also see cooper, 1974; pechmann, 1989). in the typical visual-world experiment, a participant addressee is presented with a display (either in the real world or on a computer screen) and is asked to execute a command that involves manipulating objects in the display. meanwhile, her eye movements are monitored, either with a head-mounted or a remote eye-tracking device, at a resolution between 30 and 2000 samples per second, depending on the type of equipment. people typically freely attend to and look at objects as they are mentioned (tanenhaus, magnuson, dahan, & chambers, 2000), particularly when they have to reach for or click on the objects. further, listeners tend to gaze at objects that share semantic (yee & sedivy, 2006) or visual features (dahan & tanenhaus, 2005; huettig & altmann, 2005) with a mentioned object (altmann & kamide, 2007). thus, eye movements provide a relatively unobtrusive measure of an addressee’s linguistic interpretation. researchers can calculate the timing and duration of looks to the various objects in a display, and measures include the amount of time it takes before a participant fixates on the target object for the first time, the amount of time it takes before a participant fixates on the target object for the last time (usually just before they select the object), and the changes over time in the likelihood of looking at the target and other potential target objects (known as competitors), as the critical linguistic expression unfolds. the vwp initially focused on the participant addressee’s comprehension processes while following directions from a speaker, who very often was the experimenter or a confederate participant who was using a script or trained in what to say. other work extended this paradigm to language production (griffin & bock, 2000; meyer, sleiderink, & levelt, 1998), and more recently to the comprehension of spontaneous, interactive speech (hanna & brennan, 2007; brown-schmidt & tanenhaus, 2008; brown-schmidt, campana, & tanenhaus, 2005). incremental perspective taking 13 the typical design in the vwp makes use of definite referring expressions (e.g., the starred yellow square), that refer to a single unique object in the display, but that are temporarily ambiguous among several objects when produced over time by the speaker (e.g., when there are several starred objects). these studies often manipulate and examine the point of disambiguation (pod) of the referring expression; that is, the point at which the interpretation of the expression is disambiguated given linguistic or other information available to the addressee in the visual world of the experiment. this methodology has not only allowed researchers to examine the basic nature of the incremental interpretation of noun and other linguistic phrases, but has also proven to be a powerful tool for examining the influence of higher-level sources of information on reference resolution, such as the pragmatics of how the objects in the display can be manipulated (chambers et al., 2002; hanna & tanenhaus, 2004) as well as information from the context of being engaged in a conversation, such as the knowledge or perspective of the speaker and whether it matches that of the addressee (keysar, barr, balin, & brauner, 2000; hanna, et al., 2003). we will further explore the latter work in this paper, but will first describe an example vwp experiment that illustrates the fundamentally incremental nature of language processing. in one of the first studies to use the vwp, eberhard, spivey-knowlton, sedivy, and tanenhaus (1995) showed participants displays with four objects, including colored squares and rectangles, some of which were starred. on critical trials, participants heard instructions to touch the starred yellow square. the critical manipulation was when, during the noun phrase (underlined), the display afforded unique identification of the target. when the display contained only a single starred object, the target could be identified at the word starred. when all four objects were starred, but only one was yellow, the target could be identified at yellow. in a final condition, the display contained a starred yellow square (the target) and a starred yellow rectangle, along with two different color shapes, meaning that the target could only be identified at the final word square. eberhard, et al.’s experiment allowed for a critical test of the incrementality hypothesis: if listeners interpret the test sentence incrementally, narrowing in on an identification of the target referent with each new word, then there should be significant differences in how quickly listeners look at the target referent in the three conditions. the results were consistent with this hypothesis; listeners were significantly faster to identify the target when the scene provided an earlier point of disambiguation. reference resolution also happened incredibly quickly, in a fashion that was incrementally time-locked to the words in the noun phrase. taking into account the fact that it takes about 200 ms to program and launch a saccadic eye movement (hallet, 1986), looks to the target object were programmed immediately after or even before the end of the disambiguating word. the eberhard, et al. findings demonstrate that during language processing, interpretation takes place incrementally, taking into account both the unfolding linguistic input as well as the domain of objects against which that input is being interpreted. it also means that information that has the potential to change the domain of interpretation (or referential domain, cf. salmon-alt & romary, 2000 or chambers, et al., 2002) can be manipulated and examined within the vwp as well. this is important, because the process of understanding a sentence is more complex than simply matching words onto entities in the world. instead, addressees use a variety of sources of information to resolve ambiguity and form meaningful representations of sentences as they unfold in time. employing these sources of information to constrain language processing may require accessing information from long-term memory, holding multiple pieces of information in working memory, or complex inferencing procedures. thus, a critical question is whether the language processing system is able to access and use various sources of information at a speed quick enough to affect interpretation processes as they occur in real time. the diagnostic of these effects is the same as it was for the first demonstrations of incrementality: determining whether a given source of information can affect the domain of interpretation for a referring expression by including or excluding candidate referents, and therefore changing the point of disambiguation brown-schmidt and hanna 14 for that expression. this area of research has generated many noteworthy findings, among them the fact that all of the following factors constrain the domain of interpretation incrementally with each successive word (or partial word, see dahan, et al., 2001): action affordances (chambers, et al., 2002); verb semantic restrictions (dahan & tanenhaus, 2004; altmann & kamide, 1999); fluency of referential expressions (arnold, hudson kam, & tanenhaus, 2007); scalar implicatures (sedivy, tanenhaus, chambers, & carlson, 1999); speaker reliability (grodner & sedivy, in press); speaker identity (metzing & brennan, 2003); speaker eye gaze (hanna & brennan, 2007); and event knowledge (nieuwland, otten, & van berkum, 2007). below we consider in detail one factor that has been proposed to fundamentally define the domain of interpretation for interlocutors in conversational settings – knowledge about what information is and is not in common ground. 2 common ground and perspective-taking common ground is the set of mutual knowledge and beliefs shared among conversational participants, and is thought to be formed on the bases of community membership, linguistic interactions (known as linguistic co-presence), and physical environments (known as physical copresence) (clark, 1992; 1996; clark & marshall, 1978, 1981). a large body of work by herbert clark and his colleagues has focused on the nature of common ground, which they view as a core component of language use, affecting many basic aspects of the form and process of reference (e.g., clark & wilkes-gibbs, 1986; wilkes-gibbs & clark, 1992). successful communication often requires keeping track of what is and is not in common ground with your conversational partner; for example, the experimenter in eberhard et al. (1995) could not successfully refer to the starred yellow square if the experimental participant could not see that shape in the display. indeed, clark has argued that common ground defines the relevant domain of interpretation (clark, 1992; 1996), and as such plays a central role in language processing. however, the question of whether common ground can restrict the domain of interpretation incrementally and rapidly, and the degree to which it restricts possible interpretations, was relatively unexplored until recently, and remains controversial. this is partly due to the methods available for studying language use in conversation. the typical common ground experiment employs some type of referential communication task (based on the original by krauss & weinheimer, 1966), where two people must work together to construct or arrange objects in a display without being able to see each other or each other’s display, although they are able to communicate freely in ways that maintain many aspects of naturally occurring conversation. usually, the speaker or director knows the goal arrangement of objects, and gives directions to the addressee or matcher to move and manipulate them in order to get their displays to match. before the advent of head-mounted eye tracking devices, researchers using referential communication tasks were limited to measures of language processing on a relatively coarse time-scale, such as the number of speaking turns taken by the director and matcher, the number of confusions or errors indicated by an incorrect object selection, and the total time it took to complete the task (e.g., see clark & wilkes-gibbs, 1986). while these measures provided a wealth of data about the nature of common ground, they were not able to provide information about an addressee’s interpretation on the moment-by-moment basis needed to assess the effects of common ground on reference resolution and domain restriction. in addition, prior to the development of the vwp, the techniques that were available for studying on-line processing, such as reading-time measures, or event-related potentials, were not easily adaptable for use with real-world contexts and in conversational settings. the role of common ground in language processing was also neglected due to the preponderance of processing models that limited the hypothesized role of contextually based sources of information, in order to explain how a complex process such as language incremental perspective taking 15 comprehension could occur so rapidly and seemingly effortlessly. early models of syntactic parsing proposed that this efficiency is due to the encapsulation of syntactic processes from other sources of information, such as discourse context, which were thought to require resourceintensive processing (ferreira & clifton, 1986; see discussion in trueswell & tanenhaus, 1994). sentences were thought to be processed first by a fast, syntactic parser, with more complex sources of information, such as the number of referents in the discourse context, playing a role only during a later revision stage. this early view was in direct conflict with the view of clark and colleagues that the context in which language occurred was central to language itself (clark, 1992, 1996). the crucial turning point in our understanding of the role of context in real-time language comprehension was the finding by tanenhaus, et al., (1995), that rich information from a scene can eliminate or dramatically reduce (see novick, thompson-schill, & trueswell, 2008) syntactic ambiguity. this, and subsequent experiments using the vwp (chambers, tanenhaus, & magnuson, 2004; chambers, et al., 2002; spivey, tanenhaus, eberhard, & sedivy, 2002; trueswell, sekerina, hill, & logrip, 1999; novick, et al., 2008; eberhard, et al., 1995) have demonstrated that as listeners interpret an utterance, information from the discourse context, such as the number and features of the potential referents, immediately constrains the interpretation of the sentence, and that there is clearly no early, context-free processing stage. subsequent research extended these findings, demonstrating that contextual information is combined with a variety of sources of information including verb bias, thematic fit, prosody, etc. to guide language processing decisions over time (e.g., beun & cremers, 1998; brown-schmidt & tanenhaus, 2008; niewland, et al., 2007; hanna & tanenhaus, 2004; arnold & griffin, 2007; metzing & brennan, 2003; wilson & garnsey, 2009, watson, tanenhaus, & gunlogson, 2008, arnold, et al., 2007; dahan & tanenhaus, 2004; sedivy, et al., 1999; altmann & kamide, 1999). if common ground is central to language (clark, 1992, 1996; clark & wilkes-gibbs, 1986), and serves as the primary context for language understanding, then the clear prediction would be that common ground guides the incremental processing of language. as we shall see, soon after the development of the vwp, researchers began to extend the paradigm to examine whether common ground constrained referential processing (e.g., keysar, et al., 1998). before we review this literature, it is important to first consider what the role of common ground would be. for every pair of individuals, some of their knowledge is likely to be shared, and some of their knowledge will be private. one role might be using information about what is shared and what is private to circumscribe the domain of interpretation of a referring expression. for example, if a speaker were to say that’s a lovely ring!, the addressee should constrain the domain of interpretation to a ring that is visible to both speaker and addressee, and not, say, a toe-ring concealed by the addressee’s shoe. likewise, if the speaker were to ask what just happened?, a cooperative addressee (grice, 1975) would be expected to provide information not already known to the speaker. as we shall see, studies of the role of common ground in incremental language understanding typically employ the vwp to examine situations in which a speaker and addressee interact with a set of entities in a context in which only some of those entities are in common ground. common ground is typically manipulated through either visual co-presence (whether an object can be seen by both partners), or linguistic co-presence (whether an object or piece of information has been mentioned in the context of both partners). importantly, by creating situations in which some task-relevant objects are not in common ground—that is, they are in one partner’s privileged ground—researchers can examine whether speakers and addressees take this perspective difference into account when producing or understanding the language of their partner (see schober & brennan, 2003). this ability to take a perspective difference into account is often described as ‘perspective-taking’ (cf. baron-cohen, leslie, & frith, 1985; schober, 1993; epley, keysar, van boven, & gilovich, 2004), however, it is important to note that in many cases, for brown-schmidt and hanna 16 example the interpretation of an informational question, appropriate use of common ground information requires attention to the privileged ground. thus, we define perspective-taking in language understanding not as the ability to adopt the perspective of one’s conversational partner, but rather, the ability to appropriately attend to information that is either shared, or not shared with one’s partner, depending on the context. in what follows, we review the key empirical findings regarding the role of common ground in incremental understanding, as well as the theories that have been developed to account for these findings. 3 studies of on-line perspective-taking empirical investigations of whether common ground guides on-line language processing have yielded results that appear to be contradictory. some results show dramatic failures to use common ground when comprehending (keysar, lin, & barr, 2003; keysar, et al., 1998; keysar, et al., 2000) or producing (horton & keysar, 1996) language, while others show that common ground is used only prior to receiving a critical linguistic stimulus, but not during its interpretation (barr, 2008). these findings are contradicted by yet another group of results showing very early use of common ground (hanna, et al., 2003; heller, grodner, & tanenhaus, 2008; brown-schmidt, gunlogson, & tanenhaus, 2008; brown-schmidt, 2009a,b; nadig & sedivy, 2002; metzing & brennan, 2003). the emerging consensus is that common ground only sometimes influences language processing decisions, which leaves two key open questions: 1) when does the language processing system have access to information about perspective?; and 2) why do people appear to only sometimes show sensitivity to this information? with respect to the first question, the divisions among the primary candidate theories reflect earlier debates regarding the role of semantic and contextual information in syntactic parsing (e.g. clifton & ferreira, 1989; tanenhaus, et al., 1995), with questions such as the timing of effects and resourcelimitation issues taking center stage. with respect to the second question, these inconsistencies have been attributed to differences in experimental design characteristics, such as the type of ambiguity examined (barr, 2008; keysar, et al., 2003), or the way common ground was established (hanna, et al., 2003), but other factors, such as culture (wu & keysar, 2007), inhibition control (brown-schmidt, 2009b; nilsen & graham, 2009), and mood (converse, lin, keysar, & epley, 2008; van berkum, 2009), are also thought to play a role. there are three prominent candidate models of on-line perspective-taking, a perspectiveadjustment model, an anticipation-integration model, and a constraint-based model. the critical difference between the competing theories lies in the proposed timing of perspective effects. given recent improvements in methods for investigating the time-course of processing in contexts that afford perspective differences and the manipulation of perspective-taking, it would appear that an empirical resolution would be straightforward. in contrast, results from the most recent on-line research in this area are in conflict, and a comprehensive analysis of work in this area is lacking. as will become clear, these models are motivated by fundamentally distinct views of the nature of common ground, a by-product of which is critical differences in how common ground is experimentally manipulated. consideration of such differences may allow a reconciliation of the contradictory findings. in what follows, we outline the basic claims of the primary candidate theories, and present the evidence that is used to support them. ultimately, we will argue that the bulk of this evidence is well accounted for within in a constraint-based processing framework in which perspective is one of multiple, partial constraints on language processing. we will show that findings that apparently show failures or delays in the use of perspective actually highlight the way in which the language processing system is continuously sensitive to multiple, partial constraints, including perspective. in the process, we argue that careful consideration of how common ground is incremental perspective taking 17 established in experimental frameworks, and whether it is in conflict with other constraints, is critical to understanding how perspective information is used during language processing. 4 perspective adjustment perhaps the earliest proposal for how perspective information is incorporated into the incremental interpretation of utterances was the perspective-adjustment model (keysar, et al., 1998; keysar, et al., 2000; keysar, et al., 2003; keysar, 2007). this model was motivated by the assumptions that taking another person’s perspective is resource intensive, and that most of the time perspective-taking is unnecessary because the perspectives of the speaker and addressee will overlap. the model proposed that the initial interpretation of a referring expression is unrestricted by common ground, only considering information available from the egocentric perspective of the addressee, and that a second monitoring process checks for violations of common ground, adjusting the initial interpretation as necessary (keysar, et al., 1998). early support for this model came from studies showing that, when addressees interpreted their partners’ referring expressions, private information (the ‘privileged ground’) competed for interpretation (keysar, et al., 1998; keysar, et al., 2000; keysar, et al., 2003). the most prominent demonstration of what was termed an egocentric-first processing strategy comes from keysar, et al. (2003). participants in their first experiment followed instructions to manipulate objects, such as ‘pick up the tape’, in contexts that included a cassette tape in common ground and a roll of scotch tape in privileged ground. during interpretation of tape, participants were five times more likely to fixate the privileged ground scotch tape compared to a control condition where the privileged ground item was a battery. more strikingly, they reported that participants frequently reached for the privileged ground tape: 71% of participants attempted to reach for the privileged ground item on at least 1 out of 4 trials; 46% reached on at least 2 of 4 trials. this type of result is surprising in the context of the view that emphasizes common ground as a core component of meaning (e.g. clark, 1992, 1996), and suggests that even in the ultimate interpretation of an expression, common ground does not always play a central role. what these results do not demonstrate, however, is what role common ground does have during interpretation. recent arguments in favor of the perspective-adjustment model have focused on the claim that taking another person’s perspective is resource intensive. if so, this would suggest that it should be too difficult to routinely integrate perspective into initial language processing decisions. evidence that perspective-taking is resource intensive comes primarily from measures and manipulations of working memory. for example, lin, keysar, and epley (2010; also see kronmüller & barr, 2007, horton & keysar, 1996) used a design similar to keysar, et al. (2003) and found that individuals with fewer working memory resources (due to lower working memory as measured in a span task, or due to an external load manipulation) looked at the competitor (e.g., tape) more than a control object (e.g., battery). similarly, epley, morewedge, and keysar (2004) found that children, who presumably have limited executive function, considered competitors more than adults. this evidence suggests that the resolution of competition in language processing is not always a resource-free process. what this evidence does not reveal, however, is whether perspective-taking is always a resource-intensive process, or whether perspective information is integrated into early on-line processing decisions, in spite of possible resource requirements. these issues are closely related to the question of what role, if any, common ground has in these experiments. we return to these issues in some detail in section 7. brown-schmidt and hanna 18 5 anticipation integration more recently, barr (2008) proposed a two-stage anticipation-integration account in which the language processing system has access to common ground in anticipation of (i.e., prior to) a referring expression, but when interpreting that expression, common ground does not play a role. according to this proposal, two distinct linguistic processes – anticipation and integration – show differential sensitivities to common ground. on this view, prior to hearing a referring expression (e.g., the buckle), an addressee might anticipate that her partner would refer to an object in common ground. thus, unlike the perspective-adjustment model, common ground does routinely play a role in language processing, however this role is limited to the expectations that listeners form prior to the speaker’s production of that expression. much like the perspective-adjustment model, the anticipation-integration model proposes that during the processing of a referring expression, common ground is not integrated into the processing of that expression. support for this view comes from the results of three experiments in which addressees listened to instructions to manipulate various objects, some of which were in common ground with the speaker, and some of which were in the addressee’s privileged ground. critically, addressees tended to prefer to fixate common ground objects prior to hearing a referring expression, but during interpretation of that expression, looks to common and privileged objects increased at the same rate. for example, in barr (2008) experiment 1, participants followed prerecorded instructions that they were led to believe were produced by a speaker in another room. participants saw a screen with four pictures. a target picture (buckle) and two unrelated pictures were ostensibly shared with the speaker; a competitor picture was either shared or in the participant’s privileged ground, and either had the same initial phonemes as the target (bucket), or did not (ladder). prior to the onset of buckle, participants were significantly more likely to fixate the competitor when it was in common vs. privileged ground. this baseline difference between the two conditions was argued to be evidence for the use of common ground in the anticipation of an expression. then, during interpretation of buckle, the increase in the number of fixations to the competitor bucket (i.e., the slope function relating fixation likelihood to time) was not affected by whether the competitor was in common or privileged ground. this result is consistent with the hypothesis that perspective does not guide interpretation of the noun. a second experiment which was explicitly non-interactive (subjects knew they were listing to recordings) showed weaker anticipatory effects. again, during buckle, fixations to the competitor increased at a similar rate, regardless of ground. in a third experiment, participants were led to believe they were interacting with two other participants, a male and a female, from whom they heard pre-recorded instructions. participants saw screens with three pictures, one of which was ostensibly in common ground with both speakers, one in common ground with the male speaker, and one in common ground with the female. prior to the critical instruction, participants were more likely to look at a competitor if it was in common ground with the current speaker, however common ground played no role in the overall increase in fixations to the competitor during interpretation of the noun. the results from these three experiments present a significant challenge to the view supported by clark and his colleagues that common ground is a central component to language, defining the context within which comprehension takes place. instead, they appear to suggest that language processing consists of at least two components, only one of which, anticipation, is sensitive to common ground. a key strength of the empirical findings is that the experiments do show effects of common ground during the anticipation period; this shows that participants were sensitive to the manipulation. however, a concern (that barr raises) is whether these results are representative of processing in natural situations. after all, participants did not truly form common ground with other individuals. instead, they were told what was shared with the two speakers, and then listened to pre-recorded instructions. a key question, then, is whether addressees would show incremental perspective taking 19 effects of common ground during incremental interpretation of language in more natural, interactive settings in which common ground is collaboratively established. 6 constraint-based models according to proponents of constraint-based models of perspective-taking in language processing (nadig & sedivy, 2002; hanna, et al., 2003; hanna & tanenhaus, 2004; heller, et al., 2008; brown-schmidt, et al., 2008; brown-schmidt, 2009b), perspective is only one of many partial constraints on language processing. thus, on this view, evaluating whether an individual was sensitive to perspective when interpreting a referential ambiguity requires considering the other sources of information that may have influenced this process. how strongly a given constraint is predicted to bias interpretation in favor of one outcome over another depends critically on the relative strength of the evidence provided for the possible outcomes, and how strongly that constraint is weighted relative to other constraints. as an example, imagine a situation in which you and a friend go to the store, buy some roquefort cheese, and jointly store it in the refrigerator. unbeknownst to your friend, some fresh mozzarella is also in the fridge. if your friend were to later ask you can you grab some of that cheese for me?, it would be clear that she meant the roquefort, because only the roquefort was established in common ground. in contrast, imagine your friend were simply reading from a recipe book, the next ingredient is cheese. in this case, which cheese you select from the fridge would be less influenced by the fact that the roquefort is in common ground, than which type of cheese would go best with the dish. thus, depending on characteristics of the discourse context and the linguistic expression, common ground may be more or less relevant to the understanding of language and associated behavioral goals. support for constraint-based accounts of perspective-taking comes from a growing number of findings that perspective information is used to reduce ambiguity. for example, participants in hanna, et al. (2003)’s first experiment interpreted live instructions, produced by a confederate speaker, such as put the blue triangle on the red one, in contexts that contained one blue triangle and two red triangles. the critical manipulation was whether both red triangles were in common ground, or if one was in common ground and the other was in the addressee’s privileged ground. in this situation, a variety of constraints are relevant to interpretation of the critical noun phrase, the red one. the most relevant constraints are common ground and lexical information: the use of the definite determiner the indicates that the referent should be in common ground (or uniquely identifiable in the discourse context, see roberts, 1993), and the lexical information in red indicates that the referent should be red. the constraint-based view predicts that these constraints should combine to influence identification of the referent, thus there should be competition between the red triangles that is diminished somewhat when only one is in common ground. the results were consistent with these predictions: when both red triangles were in common ground, during the interpretation of red, participants showed a significant competition effect, with equivalent fixations to the two red triangles. in contrast, when one of the red triangles was in privileged ground, participants primarily looked at the common ground triangle within the first few hundred milliseconds of processing. critically, however, as predicted by constraint-based processing, there was still a lexical competition effect; participants were significantly more likely to fixate the privileged ground competitor triangle when it was red compared to a condition in which it was yellow. hanna et al. concluded that common ground acts immediately, but probabilistically, to restrict the domain of interpretation; this means that objects that are visually available and salient, and that match the referring expression, can interfere with reference resolution even if they are in privileged ground, but that common ground acts immediately to probabilistically restrict the addressee’s domain of interpretation to those objects that can be referred to by the speaker. this observation of the simultaneous effects of many constraints is consistent with a large body of work showing that multiple, sometimes competing sources of brown-schmidt and hanna 20 information combine to influence on-line processing (novick, et al., 2008; britt, 1994; snedeker & yuan, 2008; hanna & tanenhaus, 2004; garnsey, pearlmutter, myers, & lotocky, 1997; wilson & garnsey, 2009). one objection to these results was that participants might have been cued to use perspective information due to the fact that the instructions would have been globally ambiguous without taking perspective into account (barr, 2008; also see keysar, et al., 2003). the concern, then, was that the results do not reflect typical processing, based on an assumption that global ambiguities are infrequent in natural language. however, the results have since been replicated with adults in constructions with temporary ambiguities, i.e. ambiguities that are resolved linguistically before the end of the phrase (heller, et al., 2008; also see brown-schmidt, et al., 2008; brown-schmidt, 2009b), thus the conclusion that participants used perspective information to constrain interpretation of these expressions seem warranted. further, results similar to hanna, et al.’s have also been observed in young children (nadig & sedivy, 2002; nilsen & graham, 2009; scott, brown-schmidt, fisher, & baillargeon, 2009), and in individuals with amnesia (rubin, brownschmidt, duff, tranel, & cohen et al., 2009), suggesting that use of common ground may not be exceptionally resource-intensive, contrary to some claims (converse, et al., 2008; keysar, 2007; lin, et al., 2010). sometimes perspective-taking actually requires attending to privileged information rather than ignoring it. consider the case of an informational question. a felicitous informational question typically asks about information in the addressee’s privileged ground. that is, speakers tend to ask questions when they don’t know the answer but they believe that the addressee might. thus from the addressee’s perspective, a perspective-appropriate interpretation of what an informational question is asking about would be information in the addressee’s privileged ground. in previous work, perspective-taking has typically been defined as interpreting imperatives or indirect requests as referring to common ground referents. in contrast, brown-schmidt and colleagues (brown-schmidt, et al., 2008; brown-schmidt 2009b) have demonstrated use of perspective information during the on-line interpretation of informational wh-questions. for example, brown-schmidt (2009b) asked participants what’s above the cow that’s wearing shoes?, in contexts that contained two cows (one with shoes, one with glasses). the critical manipulation was whether the animal above the competitor cow (with glasses) was in common ground. following cow, participants’ preference to fixate the privileged ground target was significantly higher when the animal above the competitor was in common ground. thus, addressees showed sensitivity to perspective by directing attention towards privileged ground entities when interpreting temporarily ambiguous questions. this finding places doubt on claims that perspective-taking involves serial adjustment away from the egocentric perspective (i.e., epley, keysar, et al., 2004), as in some cases, perspective-taking actually requires attending to privileged information. 7 reconciling previous findings in this section, we will show that each of the major findings on the time-course of perspectivetaking during language processing can be accounted for under the constraint-based processing framework. in contrast, the major alternative theories, perspective-adjustment and anticipationintegration, fail to account for key findings that support the constraint-based view. let us first consider findings by keysar, et al.’s (2003) experiment, that participants were more likely to fixate a privileged-ground competitor when it matched the referential description tape than when it did not (i.e., scotch tape vs. battery). while this result was interpreted as evidence against the use of perspective information, according to the constraint-based view, this result shows nothing more than evidence of lexical competition effects in interpretation. participants should always fixate a referent that matches the critical referring expression more incremental perspective taking 21 than one that does not, since lexical information is a very good cue to speaker meaning. the extent of interference from privileged ground competitors in this experiment was further amplified by a meaning dominance effect: each of the competitor objects (e.g., scotch tape) were explicitly designed to be a better match to the critical referring expression, tape, than the common ground target (e.g., cassette tape). similar problems are present in several other studies that purport to show evidence for difficulties in perspective-taking (e.g., converse, et al., 2008; epley, et al., 2004; keysar, et al., 1998; keysar, et al., 2000; lin, et al., 2010; wu & keysar, 2007). for example, in keysar, et al. (2000), the director referred to the bottom block in contexts with three blocks, two in common ground and one in privileged ground, each on a different row of the display. however, in the critical conditions, the privileged object was always the best perceptual match for the referring expression. in this example, the block in privileged ground was always the bottom-most block, and therefore matched the referring expression, the bottom block, the best. similar to the keysar, et al. (2003) results, participants were not only just as likely to look at the blocks in privileged and common ground, but they actually initially preferred to look at the block in privileged ground, a result that keysar, et al. (2003) used to support their view that initial interpretation is egocentric. under these circumstances, it is difficult to know whether common ground also played a role, since this result only shows that common ground did not completely rule out consideration of the privileged ground object when there was lexical competition between two interpretations, and a lexical bias in favor of a privileged competitor. what role might perspective information have in these situations? according to constraintbased models, perspective information would have offered some support to the common-ground interpretation of the critical expression. critically, however, the experiments by keysar, et al. (2003) and keysar, et al. (2000) lacked the crucial comparison condition that would have revealed the common ground effect—one in which the critical competitor is in common ground (e.g., the scotch tape or the bottom-most block). this key condition would likely have shown that participants are less likely to fixate a competitor when it is privileged versus common (i.e., the common ground effect observed in hanna, et al., 2003). thus, experimental designs such as the one used by keysar, et al. (2003; also converse, et al., 2008; epley, morewedge, et al., 2004; keysar, et al., 1998; keysar, et al., 2000; lin, et al., 2010; wu & keysar, 2007) are not designed in such a way that perspective effects could be revealed, even if they were to occur. similar problems confound the interpretation of findings that individuals with fewer resources—individuals with low, or taxed working memory (lin, et al., 2010), or children (epley, morewedge, et al., 2004) were more likely to gaze at a privileged ground object when it was a strong lexical competitor, than when it was not. while these results were taken as evidence that perspective-taking requires mental effort, in our view, what they actually show is that resolving lexical competition requires resources. the crucial comparison that would test the use of perspective in the face of resource limitations—a condition in which both the target and lexical competitor were common ground—was missing. several other concerns limit the impact of these findings. one is the assumption that if perspective-taking is resource-intensive, then it cannot guide on-line processing. this has not been demonstrated experimentally, and the fact that many studies have now shown successful use of common ground on-line with adults (hanna, et al., 2003; heller, et al., 2008; brown-schmidt, et al., 2008; brown-schmidt, 2009b), and even children (nadig & sedivy, 2002; nilsen & graham, 2009; matthews, lieven, & tomasello, 2010), suggests that even if the premise is true, the conclusion is false. a further consideration is that measures and manipulations of working memory (also kronmüller & barr, 2007; horton & keysar, 1996) may be more informative about language experience than capacity (see macdonald & christiansen, 2002; wells, et al., 2009). finally, in some cases computing and accessing representations of common ground may be highly automatized, due to associations in memory between partners and shared information (horton & gerrig, 2005a,b; horton, 2007), or low-level cues to joint knowledge, such as shared gaze (hanna & brennan, 2007). brown-schmidt and hanna 22 let us now consider the studies reported in barr (2008) which showed evidence of perspective-taking in anticipation of hearing a critical expression, but not during its interpretation. unlike the studies by keysar and colleagues, these experiments did include the critical comparison condition in which the competitor was in common ground. the conclusion drawn from the results of these experiments, that perspective does not affect on-line interpretation of the critical expressions, is incompatible with the constraint-based prediction. further, the fact that there was a significant perspective effect prior to the critical expressions suggests that this was not simply a null effect due to a weak manipulation of perspective. however, there are several reasons to treat these results with caution. our first objection concerns the claim that there are two separate processing stages— anticipation and integration—and the finding that only the former is sensitive to perspective. if the baseline preference to fixate common ground entities were truly an anticipatory effect, it would suggest that the processing system always anticipates common ground referents prior to a critical expression. however, this can’t possibly hold for normal conversation, because, as brown-schmidt and colleagues (2008) pointed out, whether shared or private information is relevant depends on utterance form (e.g. whether it is a statement or a question). a processing mechanism that focused attention on common ground information prior to each linguistic act would frequently make the wrong prediction. indeed, in the only study to examine perspectivetaking in completely unscripted conversation between naïve partners (brown-schmidt, et al., 2008, experiment 2), addressees were equally likely to gaze at common and privileged entities prior to the onset of critical expressions. given this, it would seem that the ‘anticipatory’ effects may simply reflect processing of the initial words in the speaker’s request; in barr’s first experiment, they were click on the. as participants heard these words, they were interpreted as a request to perform an action on some object, with the constraint that it should be a shared object. the second objection has to do with the finding that perspective played no role in interpretation of the critical word bucket, which was temporarily ambiguous between the shared target and the privileged competitor, ‘buckle’. according to constraint-based theories, at least two constraints should play a role in the interpretation: perspective, which should have supported the ‘bucket’ interpretation; and lexical information, which should have temporarily supported both ‘buckle’ and ‘bucket’ interpretations. why, then, did fixations to buckle and bucket increase at the same rate? we suspect the answer may lie in the fact that participants did not actually form common ground in this experiment. recall that according to classic accounts, common ground forms as individuals collaboratively establish what information is jointly known through an interactive grounding process (brennan & clark, 1996). in each of the studies that have shown significant effects of common ground in on-line interpretation, participants interacted with live partners with whom they were able to collaboratively form common ground (e.g., hanna, et al., 2003; nadig & sedivy, 2002; heller, et al., 2008; brown-schmidt, et al., 2008; brown-schmidt, 2009a,b; metzing & brennan, 2003). in contrast, in barr’s (2008) experiments, participants never interacted with live partners, and never engaged in grounding procedures. instead, in the first experiment, a marking on a card indicated which object was shared with the speaker. in the second experiment, participants knew they were listening to pre-recorded instructions and again saw markings on cards to cue which object was shared. finally, in the third experiment, participants listened to recordings from what they were led to believe were two different speakers. this experiment contained a pseudo-grounding procedure in which the participant overheard the experimenter and one of the pre-recorded voices ‘discuss’ what was in common ground. this design would be expected to produce weak common ground effects as it is well known that overhearers do not form common ground in the same way as full participants (schober & clark, 1989; also see wilkes-gibbs & clark, 1992). consistent with this assertion is recent evidence from brown-schmidt (2009a) that addressees show immediate sensitivity to joint knowledge in interactive settings only. incremental perspective taking 23 re-framing these findings in such a way, it becomes clear that the constraint-based framework would predict the observed findings given the weak manipulation of common ground in conjunction with a powerful lexical competition effect: prior to the critical noun, participants incrementally interpreted the utterance click on the... as referring to shared objects, due to the absence of competing constraints and a weak perspective effect. then, during interpretation of the critical noun buckle, lexical constraints were in conflict with a relatively weak perspective effect. these lexical constraints overwhelmed perspective information, resulting in an equivalent increase in fixations to the target and competitor1. because of the weak perspective manipulation, then, the experiment was not designed in such a way to distinguish between anticipationintegration and constraint-based accounts. a final consideration is that individuals may vary in their sensitivity to perspective information, or the degree to which they are able to successfully combine competing sources of information to arrive at an appropriate interpretation of the speaker’s meaning. taking into account such individual differences may provide a partial explanation for the large amount of variability in perspective-taking findings, and for why some participants fail to use perspective information in the face of competing constraints. for example, inhibition control, which is known to predict children’s performance in theory-of-mind tasks (carlson & moses, 2001; chasiotis, kiessling, hofer, & campos, et al., 2006; hughes & ensor, 2005) also predicts the use of perspective information during on-line language processing in both children (nilsen & graham, 2009), and adults (brown-schmidt, 2009b), with individuals who score higher on conflict inhibition tasks showing better ability to rule out perspective-inappropriate interpretations of their partner’s speech. the locus of the effect is unknown, but may be related to improved inhibition of the perspective-inappropriate interpretation. alternatively, individuals with higher inhibition control may simply be more skilled at resolving conflicts between multiple constraints (also see novick, et al., 2008). mood and culture may also play a role. for example, converse, et al., (2008) found that participants in a happy mood were more likely to adopt perspectiveinappropriate interpretations, and wu & keysar (2007) found that chinese participants were less likely to adopt perspective-inappropriate interpretations compared to their american counterparts. on the constraint-based view, the fact that factors such as these modulate whether an addressee will entertain a perspective-inappropriate response suggests that it is unlikely that there exists a perspective-free processing stage, but instead that the likelihood that perspective will guide interpretation is dependent on a variety of factors. 8 evaluating constraint-based models in this section, we will consider what sorts of experimental manipulations would provide critical tests of the constraint-based model. the constraint-based approach proposes that perspective 1 it is worth noting that there are additional concerns regarding the statistical support for these anticipation and integration effects. barr (2008) used a regression approach, and equated anticipatory effects with condition differences at the intercept, and integration effects with slope differences, that is, differences between the conditions in the change of the likelihood of a target fixation over time. this statistical approach of modeling intercept effects separately from slope effects has been questioned (tanenhaus, et al., 2008), as slopes and intercepts are likely to be non-independent (even in log-odds space). the independence assumption requires the population of trials on which there was not a baseline target fixation to be representative and comparable across the conditions being compared, as the rise in slope following the critical noun is largely driven by shifts in fixations from non-targets to the target. however, if only a subset of the trials (or participants) are likely to be affected by the perspective manipulation (e.g., due to individual differences in mood [converse, et al., 2008], inhibition control [brown-schmidt, 2009b], etc.), those datapoints would be systematically excluded from the slope measure, as they would have already shown sensitivity to the manipulation at baseline. an alternative is to eliminate baseline (intercept) effects by analyzing the sub-set of trials without a baseline target fixation, however this approach can result in significant data loss (e.g., brown-schmidt, 2009b). alternative designs, such as those in which attention is drawn to an unrelated object at baseline (see hanna, et al., 2003) may be required to resolve these issues. brown-schmidt and hanna 24 information, along with multiple other partial constraints, combine to influence interpretation processes. a growing number of results are consistent with the constraint-based view, including findings that addressees use information about the perspective of their partner to guide on-line understanding (nadig & sedivy, 2002; hanna, et al., 2003; hanna & tanenhaus, 2004; heller, et al., 2008; brown-schmidt, et al., 2008; brown-schmidt, 2009b; or character in a narrative, ferguson, scheepers, & sanford, 2010), as well as findings that addressees are sensitive to partner-specific referring history when interpreting referring expressions (brown-schmidt, 2009a, metzing & brennan, 2003; brennan & hanna, 2009; matthews, et al., 2010). in the previous section, we have shown that the constraint-based model of perspective-taking can account for the wide range of findings that were once thought to be inconsistent with the constraint-based approach. as we demonstrated, considerations of how perspective was established, and what other constraints were relevant, resolve such inconsistencies. in the past, implemented constraint based models have been used to generate testable predictions about the strength of verb bias, thematic fit, and discourse context effects, and the subsequent predictions for interpretation of attachment ambiguities (spivey & tanenhaus, 1998; mcrae, spivey-knowlton, & tanenhaus, 1998; tanenhaus, spivey-knowlton, & hanna, 2000). a similar extension to common ground, however, proves more difficult, because quantifying the degree to which common ground picks out a particular referent is not straightforward. a more tangible goal may be to manipulate the relative strength of common ground, an approach that was successfully applied by brown-schmidt (2009a) to the question of whether collaborativelydefined expressions are part of partner-specific common ground. a similar approach could be taken to examine the role of perspective information in on-line processing. generally speaking, the constraint-based view predicts that the strength of perspective representations should directly modulate their effectiveness, with the strongest effects of common ground predicted when it is established interactively. if a weak common ground manipulation were pitted against a strong competing constraint (e.g., see barr, 2008), the perspective effect is expected to be small or possibly eliminated. for example, if participants were instructed to put the star below the bucket in a context that included a common-ground ‘bucket’, and a ‘buckle’ that was either common or privileged, the constraint-based account would predict that during interpretation of bucket, participants should experience less competition from the buckle when it was privileged, compared to common. however, the magnitude of this competition reduction should be directly modulated by the strength of the common ground manipulation, with larger reductions in interactive conversation compared to non-interactive settings. a perhaps unsatisfying implication of such predictions is that in cases where the constraintbased view predicts significant effects of ground and interactivity, the competing models predict null effects during on-line interpretation. another way to tease apart these models would be to examine situations in which the theories make opposite predictions. a candidate test could involve examination of spatial perspective-taking: consider the fact that in a face-to-face conversation, one’s egocentric ‘left’ is one’s partner’s ‘right’. in face-to-face task-based conversation, dialog partners often negotiate such terms by agreeing to re-define these terms from a particular perspective (although sometimes neutral perspectives are adopted, or a more skilled partner accommodates to a less skilled partner, c.f., schober, 1993; 1995). for example, consider the following exchange, from naïve participants on either side of a 3-d display (brown-schmidt, et al., 2008, experiment 2): 1. yeah what's to the left of that? 2. ahhh 1. or to the right? 2. pig with sneakers… it's always gonna be from your perspective pig with sneakers incremental perspective taking 25 now consider a similar situation in which partners play a game in which they sit face-to-face on either side of a display with animals, and direct each other to re-arrange shapes in the display (see figure 1). imagine that the partners jointly define “right” from partner 1’s perspective, thus making this new definition common ground. if partner 1 were to subsequently say take the star and put it below the dog on the right, in a context containing two dogs, one on the left side of the display, and one on the right, it would be possible to examine partner 2’s incremental interpretation of the expression. according to constraint-based models, at right, partner 2 should immediately access the collaboratively-defined meaning of right as defined from partner 1’s perspective, and interpret the expression as referring to the dog wearing the purse. lexical competition from the previous word dog, and the overall frequency in the language of using right to refer to the egocentric right would be competing constraints, partially supporting interpretation of the expression as the dog wearing the flower. according to the constraint-based view, these competing constraints would manifest as increased fixations to the competitor dog (wearing the flower) compared to the cows in the scene. figure 1. example display from partner 2’s perspective. partner 1 would see the mirrorimage reverse. in contrast, on the perspective-adjustment model, when partners have opposite perspectives, the initial interpretation of right should be egocentric, as the dog wearing the flower. similarly, on the anticipation-integration account, during the on-line interpretation (integration) of right, perspective should not be relevant, thus the expression should be interpreted as referring to the dog wearing the flower. a crucial test of the constraint-based account would come from a case in which partner 1 violated common ground, and used the spatial term from partner 2’s perspective: take the star and put it below the dog on the right wearing the flower. on perspective-adjustment and anticipation-integration accounts, this utterance is consistent with the egocentric perspective, and should be easy to interpret. in contrast, on the constraint-based account, this sentence should create a garden-path effect, and should be confusing. a comparison condition would come from a case in which the partners jointly agreed to define right as from partner 2’s perspective. according to the constraint-based account, the critical sentence should be significantly easier to understand when right was defined from 2’s perspective. there should be no difference according to the competing theories, since both would predict that interpretation of right from partner 2’s perspective should be relatively easy. brown-schmidt and hanna 26 finally, while empirical investigations such as the one we suggested will be important as we evaluate and improve models of on-line processing, it is important to point out that much of the research and theorizing about the on-line use of perspective has focused on issues of timing. one consequence of this near-singular focus on the timing of perspective effects is that little is known about the nature of perspective representations themselves. attention to such matters, however, is necessary in order to build explicit models of perspective-taking, and may be important for fleshing out predictions regarding the difficulty of rapidly engaging perspective representations during on-line interpretation. there are at least two views of how perspective representations might be encoded. early psychological views of common ground postulated that interlocutors form rich representations of each other’s knowledge that go well beyond whether a given object is visually co-present or not. for example, clark and marshall (1978, 1981) proposed that interlocutors store rich, diary-like information about events and the people they experienced those events with. they argued that information enters the common ground through a variety of different routes including not only what is jointly visible, but also what has been said, what can be jointly inferred, and what is general knowledge in the community. speakers and addressees alike access these representations of jointly experienced events, dialogs, and community knowledge and use this information to support the use of language in conversational settings. whether representations as rich as the ones that clark and marshall postulated are actually used on-line in conversation is a topic of current debate. while there has been some study of the use of these different types of common ground in language production and in the ultimate understanding of an utterance (fussell & krauss, 1991, 1992; issacs & clark, 1987; also see schober & brennan, 2003), most if not all of the existing research on the on-line use of common ground has focused on simple visual or linguistic co-presence (e.g. keysar, et al., 1998; keysar, et al., 2003; hanna, et al., 2003; brown-schmidt, et al., 2008; nadir & sedivy, 2002; brown-schmidt, 2009b). another possibility, suggested by hanna and colleagues (hanna, et al., 2003; also see horton, 2007), is that sensitivity to common ground may be largely supported by low-level information sources such as eye-gaze (see hanna & brennan, 2007; richardson, dale, & kirkham, 2007), gesture and body position, and possibly even very basic co-occurrence information (horton & gerri, 2005a, b; horton, 2007). understanding whether the use of common ground requires accessing rich, event-based representations (clark & marshall, 1978, 1981), or sensitivity to low-level cues (hanna & brennan, 2007; horton & gerri, 2005a, b), or both, will likely provide crucial insights into the mechanisms by which this information is integrated into the on-line processing of utterances. 9 conclusions in this article, we focused on one type of contextual information, specifically the common-ground status of the entities in the discourse context. like earlier debates about the role of context in syntactic processing decisions, the crucial theoretical question was not whether common ground played a role in processing, but when, during the time-course of understanding an utterance, common ground was relevant. a second important question that has emerged is why there is so much variability between individuals and across contexts in the use of common ground. as we have discussed, the candidate theories—the perspective-adjustment, anticipation-integration, and constraint-based models—primarily differ in terms of when, in time, common ground is thought to guide interpretation of referring expressions in dialog. thus, like earlier debates about the role of context in syntactic ambiguity resolution, the precise time-scale of eye-tracking data in combination with the ability to monitor eye movements in rich discourse contexts has made the vwp the dominant methodology for examining the role of common ground in on-line language understanding. the now large number of studies using this paradigm have generated many insights, and puzzles, about how common ground information guides on-line language understanding. incremental perspective taking 27 here we have argued for a constraint-based model of the role of common ground in on-line language understanding. we proposed that common ground, or more generally, perspective information, is always potentially available to language processing decisions, but that the likelihood an addressee will adopt a perspective-appropriate interpretation of her partner’s utterance will depend on the strength and relevance of the perspective representation, and whether perspective is in conflict with other constraints. this view can account for a large number of results in the literature, including findings that as an addressee follows instructions from a dialog partner, she is likely to consider all entities that match her partner’s referring expression (keysar, et al., 1998; keysar, et al., 2000; keysar, et al., 2003), with a preference for those in common ground (hanna, et al., 2003, nadir & sedivy, 2002). this view is also consistent with findings that as an addressee interprets an informational wh-question, she is likely to interpret that question as asking about information only the addressee knows about (brown-schmidt, et al., 2008; brown-schmidt, 2009b). while perspective is always available, in some circumstances it may be less relevant to language processing decisions. for example, if the addressee is listening to pre-recorded speech (e.g. barr, 2008), as in the pre-recorded messages played in airports, or if the speaker is not herself the author of the ideas, as when reading a recipe aloud or giving a survey (see schober & conrad, 2008), the addressee may rightly infer that the speaker’s knowledge state is of little relevance to understanding what is being said. in these circumstances, according to the constraint-based model, the addressee would be expected to weight other sources of information, such as lexical meaning or thematic fit, more heavily (see brownschmidt, 2009a). in conclusion, using language in conversation requires speakers and listeners to quickly incorporate many different sources of information into language production and comprehension processes. here, we focused primarily on the task of the addressee. as she listens to her partner’s speech, she must rapidly and incrementally interpret each word and utterance, making many provisional commitments about the intended meaning of each word or structure. as has been clearly articulated before, we arrive at meaning in language not through the linguistic stimulus alone, but by combining it with information gleaned from our background knowledge and a rich discourse context (clark, 1992, 1996). for scholars of language processing, the domain of inquiry, then, is to examine what sources of information are relevant, and the mechanisms by which they are integrated into our understanding of language. acknowledgements this material is based partially upon work supported by the national science foundation under grant no. bcs-10-19161 to the first author, and grant no. iis-07-13287 to the second author. the authors would like to thank three anonymous reviewers for their insightful comments and suggestions. j. e. hanna would like to thank timothy dustin, sarah mundell, barbara percival, and anne wildman for their contributions. some of the material in this paper was presented by j.e. hanna at the workshop on incrementality in verbal interaction (june, 2009) at zif (centre for interdisciplinary research), university of bielefeld, germany. references gerry t. m. altmann and yuki kamide (1999). incremental interpretation at verbs: restricting the domain of subsequent reference. cognition, 73:247-264. gerry t. m. altmann and yuki kamide (2007). the real-time mediation of visual attention by language and world knowledge: linking anticipatory (and other) eye movements to linguistic processing. journal of memory and language, 57:502-518. brown-schmidt and hanna 28 gerry t. m. altmann and yuki kamide (2009). discourse-mediation of the mapping between language and the visual world: eye movements and mental representation. cognition, 111:5571. paul d. allopenna, james s. magnuson, and michael k. tanenhaus (1998). tracking the time course of spoken word recognition: evidence for continuous mapping models. journal of memory and language, 38:419-439. jennifer e. arnold and zenzi m. griffin (2007). the effect of additional characters on choice of referring expression: everyone counts. journal of memory and language, 56:521-536. jennifer e. arnold, carla l. hudson kam, and michael k. tanenhaus (2007). if you say thee uh you are describing something hard: the on-line attribution of disfluency during reference comprehension. journal of experimental psychology, learning, memory and cognition, 33:914-930. simon baron-cohen, alan m. leslie and uta frith (1985). does the autistic child have a “theory of mind”? cognition, 21:37-46. dale j. barr (2008). pragmatic expectations at linguistic evidence: listeners anticipate but do not integrate common ground. cognition, 109:18-40. robbert-jan beun and anita h. m. cremers (1998). object reference in a shared domain of conversation. pragmatics & cognition, 6:121-151. susan e. brennan and herbert h. clark (1996). conceptual pacts and lexical choice in conversation. journal of experimental psychology: learning, memory & cognition, 22:482493. m. anne britt (1994). the interaction of referential ambiguity and argument structure in the parsing of prepositional phrases. journal of memory and language, 33:251-283. sarah brown-schmidt (2009a). partner-specific interpretation of maintained referential precedents during interactive dialog. journal of memory and language, 61:171-190. sarah brown-schmidt (2009b). the role of executive function in perspective-taking during online language comprehension. psychonomic bulletin and review, 16:893-900. sarah brown-schmidt and michael k. tanenhaus (2008). real-time investigation of referential domains in unscripted conversation: a targeted language game approach. cognitive science, 32:643-684. sarah brown-schmidt, christine gunlogson, and michael k. tanenhaus (2008). addressees distinguish shared from private information when interpreting questions during interactive conversation. cognition, 107:1122-1134. sarah brown-schmidt, ellen campana, and michael k. tanenhaus (2005). real-time reference resolution by naïve participants during a task-based unscripted conversation. approaches to studying world-situated language use: bridging the language as product and language as action traditions, eds. john c. trueswell and michael k. tanenhaus, pages 153-171. mit press, cambridge, massachusetts. stephanie m. carlson and louis j. moses (2001). individual differences in inhibitory control and children’s theory of mind. child development, 72:1032–1053. athanasios chasiotis, florian kiessling, jan hofer, and domingo campos (2006). theory of mind and inhibitory control in three cultures: conflict inhibition predicts false belief understanding in germany, costa rica and cameroon. international journal of behavioral development, 30:249-260. craig g. chambers, michael k. tanenhaus, kathleen m. eberhard, hana filip, and greg n. carlson (2002). circumscribing referential domains during real-time sentence comprehension. journal of memory and language, 47:30-49. craig g. chambers, michael k. tanenhaus, and james s. magnuson (2004). actions and affordances in syntactic ambiguity resolution. journal of experimental psychology: learning, memory and cognition, 30:687-696. incremental perspective taking 29 herbert h. clark (1992). arenas of language use. university of chicago press, chicago, illinois. herbert h. clark (1996). using language. cambridge university press, cambridge, united kingdom. herbert h. clark and catherine marshall (1978). reference diaries. theoretical issues in natural language processing (vol. 2), ed. david l. waltz, pages 57-63. association for computing machinery, new york, new york. herbert h. clark and catherine marshall (1981). definite reference and mutual knowledge. elements of discourse understanding, eds. aravind k. joshi, bonnie l. webber, ivan a. sag, pages 10-63. cambridge university press, cambridge, united kingdom. herbert h. clark and deanna wilkes-gibbs (1986). referring as a collaborative process. cognition, 22:1-39. charles clifton, jr. and fernanda ferreira (1989). ambiguity in context. language and cognitive processes, 4:77-103. benjamin a. converse, shuhong lin, boaz keysar, and nicholas epley (2008). in the mood to get over yourself: mood affects theory-of-mind use. emotion, 8(5):725-730. roger m. cooper (1974). the control of eye fixation by the meaning of spoken language: a new methodology for the real-time investigation of speech perception, memory, and language processing. cognitive psychology, 6:84-107. delphine dahan, james s. magnuson, michael k. tanenhaus, and ellen m. hogan (2001). subcategorical mismatches and the time course of lexical access: evidence for lexical competition. language and cognitive processes, 16(5/6):507-534. delphine dahan and michael k. tanenhaus (2004). continuous mapping from sound to meaning in spoken-language comprehension: immediate effects of verb-based thematic constraints. journal of experimental psychology, learning, memory and cognition, 30:498513. delphine dahan and michael k. tanenhaus (2005). looking at the rope when looking for the snake: conceptually mediated eye movements during spoken-word recognition. psychonomic bulletin & review, 12:453-459. dieuwke de goede, petra van alphen, emma mulder, yvonne blokland, josé kerstholt, and jos j. a. van berkum (2009). the effect of mood on anticipation in language comprehension: an erp study. poster presented at the 16th annual meeting of the cognitive neuroscience society (cns-2009), san francisco. kathleen m. eberhard, michael j. spivey-knowlton, julie c. sedivy, and michael k. tanenhaus (1995). eye-movements as a window into spoken language comprehension in natural contexts. journal of psycholinguistic research, 24:409-436. nicholas epley, boaz keysar, leaf van boven, and thomas gilovich (2004). perspective taking as egocentric anchoring and adjustment. journal of personality and social psychology, 87(3):327-339. nicholas epley, carey k. morewedge, and boaz keysar (2004). perspective taking in children and adults: equivalent egocentrism but differential correction. journal of experimental social psychology, 40:760-768. heather j. ferguson, christoph scheepers, and anthony j. sanford (2010). expectations in counterfactual and theory of mind reasoning. language and cognitive processes, 25:297346. fernanda ferreira and charles c. clifton, jr. (1986). the independence of syntactic processing. journal of memory and language, 25:348-368. susan r. fussell and robert m. krauss (1991). accuracy and bias in estimates of others’ knowledge. european journal of social psychology, 21:445-454. susan r. fussell and robert m. krauss (1992). coordination of knowledge in communication: effects of speakers’ assumptions about what others know. journal of personality and brown-schmidt and hanna 30 social psychology, 62:378-391. susan m. garnsey, neal j. pearlmutter, elizabeth meyers, and melanie a. lotocky (1997). the contributions of verb bias and plausibility to the comprehension of temporarily ambiguous sentences. journal of memory and language, 37:58-93. zenzi m. griffin and kathryn bock (2000). what the eyes say about speaking. psychological science, 11(4):274-279. herbert p. grice (1975). logic and conversation. syntax and semantics 3: speech arts, eds. peter cole and jerry l. morgan, pages 41-58. academic press, new york, new york. daniel grodner and julie sedivy (in press). the effects of speaker-specific information on pragmatic inferences. the processing and acquisition of reference, eds. n. pearlmutter & e. gibson. mit press: cambridge, massachusetts. peter e. hallett (1986). eye movements. handbook of perception and human performance (vol. 1), eds kenneth r. boff, lloyd kaufman, and james p. thomas, pages 10.1-10.112. wiley, new york, new york. joy e. hanna and susan e. brennan (2007). speakers' eye gaze disambiguates referring expressions early during face-to-face conversation. journal of memory and language, 57:596-615. joy e. hanna, michael k. tanenhaus, and john c. trueswell (2003). the effects of common ground and perspective on domains of referential interpretation. journal of memory and language, 49:43-61. joy e. hanna and michael k. tanenhaus (2004). pragmatic effects on reference resolution in a collaborative task: evidence from eye movements. cognitive science, 28:105-115. daphna heller, daniel grodner, and michael k. tanenhaus (2008). the role of perspective in identifying domains of reference. cognition, 108:831-836. william s. horton (2007). the influence of partner-specific memory associations on language production: evidence from picture naming. language and cognitive processes, 22:11141139. william s. horton and richard j. gerrig (2005a). conversational common ground and memory processes in language production. discourse processes, 40(1):1-35. william s. horton and richard j. gerrig (2005b). the impact of memory demands on audience design during language production. cognition, 96:127-142. william s. horton and boaz keysar (1996). when do speakers take into account common ground? cognition, 59:91-117. falk huettig and gerry t. m. altmann (2005). word meaning and the control of eye fixation: semantic competitor effects and the visual world paradigm. cognition, 96:b23-b32. claire hughes and rosie esnor (2005). executive function and theory of mind in 2 year olds: a family affair? developmental neuropsychology, 28:645-668. ellen a. issacs and herbert h. clark (1987). references in conversation between experts and novices. journal of experimental psychology: general, 116:26-37. boaz keysar, dale j. barr, jennifer a. balin, and timothy s. paek (1998). definite reference and mutual knowledge: process models of common ground in comprehension. journal of memory and language, 39:1-20. boaz keysar, dale j. barr, jennifer a. balin, and jason s. brauner (2000). taking perspective in conversation: the role of mutual knowledge in comprehension. psychological science, 11(1):32-38. boaz keysar, shuhong lin, and dale j. barr (2003). limits on theory of mind use in adults. cognition, 89:25-41. boaz keysar (2007). communication and miscommunication: the role of egocentric processes. intercultural pragmatics, 4:71-84. incremental perspective taking 31 robert m. krauss and sidney weinheimer (1966). concurrent feedback, confirmation, and the encoding of referents in verbal communication. journal of personality and social psychology, 4(3):343–346. edmundo kronmüller and dale j. barr (2007). perspective-free pragmatics: broken precedents and the recovery-from-preemption hypothesis. journal of memory and language, 56:436455. williem j. m. levelt (1989). speaking. mit press, cambridge, massachusetts. shuhong lin, boaz keysar, and nicholas epley (2010). reflexively mindblind: using theory of mind to interpret behavior requires effortful attention. journal of experimental social psychology, 46:551-556. maryellen c. macdonald and morten h. christiansen (2002). reassessing working memory: comment on just and carpenter (1992) and waters and caplan (1996). psychological review, 109(1):35-54. james s. magnuson, james a. dixon, michael k. tanenhaus, and richard n. aslin (2007). the dynamics of lexical competition during spoken word recognition. cognitive science, 31:133156. danielle matthews, elena lieven, and michael tomasello (2010). what’s in a manner of speaking? children’s sensitivity to partner-specific referential precedents. developmental psychology, 46:749-760. ken mcrae, michael j. spivey-knowlton, and michael k. tanenhaus (1998). journal of memory and language, 38:283-312. charles metzing and susan e. brennan (2003). when conceptual pacts are broken: partner specific effects on the comprehension of referring expressions. journal of memory and language, 49:201-213. antje s. meyer, astrid m. sleiderink, and willem j. m. levelt (1998). viewing and naming objects: eye movements during noun phrase production. cognition, 66:b25-b33. aparna s. nadig and julie c. sedivy (2002). evidence of perspective-taking constraints in children's on-line reference resolution. psychological science, 13(4):329-336. mante s. nieuwland, marte otten, and jos j. a. van berkum (2007). who are you talking about? tracking discourse-level referential processing with event-related potentials. journal of cognitive neuroscience, 19(2):228-236. elizabeth s. nilsen and susan a. graham (2009). the relations between children’s communicative perspective-taking and executive functioning. cognitive psychology, 58:220249. jared m. novick, sharon l. thompson-schill, and john c. trueswell (2008). putting lexical constraints in context into the visual-world paradigm. cognition, 107:850-903. thomas pechmann, t. (1989). incremental speech production and referential overspecification. linguistics, 27:89–110. daniel c. richardson, rick dale, and natasha z. kirkham (2007). the art of conversation is coordination. psychological science, 18(5):407-413. craige roberts (2003). uniqueness in definite noun phrases. linguistics and philosophy, 26:287350. rachael d. rubin, sarah brown-schmidt, melissa c. duff, daniel tranel, and neal j. cohen (2009). common ground representations in hippocampal amnesia. poster presented at the society for neuroscience conference, chicago, il. susanne salmon-alt and laurent romary (2000). generating referring expressions in multimodal contexts. paper presented at the workshop on coherence in generated multimedia inlg 2000, mitzpe ramon, israel. michael f. schober (1993). spatial perspective-taking in conversation. cognition, 47:1-24. michael f. schober (1995). speakers, addressees, and frames of reference: whose effort is brown-schmidt and hanna 32 minimized in conversations about locations? discourse processes, 20:219-247. michael f. schober and susan e. brennan (2003). processes of interactive spoken discourse: the role of the partner. handbook of discourse processes, eds. a. c. graesser, m. a. gernsbacher, & s. r. goldman, pages 123-164. lawrence erlbaum, hillsdale, new jersey. michael f. schober and herbert h. clark (1989). understanding by addressees and overhearers. cognitive psychology, 21:211-232. michael f. schober and frederick g. conrad (2008). survey interviews and new communication technologies. in frederick g. conrad and michael f. schober (eds.), envisioning the survey interview of the future, pages 1-30. wiley, new york, new york. rose scott, sarah brown-schmidt, cynthia fisher, and renee baillargeon (2009). 2.5-year-olds consider others’ false beliefs when interpreting their speech. paper presented at the society for research on child development annual conference, denver, co. julie c. sedivy, michael k. tanenhaus, craig g. chambers, and greg n. carlson (1999). achieving incremental semantic interpretation through contextual representation. cognition, 71:109-147. jesse snedeker and sylvia yuan (2008). effects of prosodic and lexical constraints on parsing in young children (and adults). journal of memory and language, 58:574-608. michael j. spivey and michael k. tanenhaus (1998). syntactic ambiguity resolution in discourse: modeling the effects of referential context and lexical frequency. journal of experimental psychology: learning, memory, and cognition, 24:1521-1543. michael j. spivey, michael k. tanenhaus, kathleen m. eberhard, and julie c. sedivy (2002). eye movements and spoken language comprehension: effects of visual context on syntactic ambiguity resolution. cognitive psychology, 45:447-481. michael k. tanenhaus, michael j. spivey-knowlton, kathleen m. eberhard, and julie c. sedivy (1995). integration of visual and linguistic information in spoken language comprehension. science, 268:1632-1634. michael k. tanenhaus, michael j. spivey-knowlton, and joy e. hanna (2000). modeling thematic and discourse context effects on syntactic ambiguity resolution within a multiple constraints framework: implications for the architecture of the language processing system. in matthew w. crocker, martin pickering, and charles clifton, jr. (eds.). architecture and mechanisms of the language processing system. cambridge university press, cambridge, united kingdom. michael k. tanenhaus, james s. magnuson, delphine dahan, craig chambers (2000). eye movements and lexical access in spoken-language comprehension: evaluating a linking hypothesis between fixations and linguistic processing. journal of psycholinguistic research, 29:557-580. michael k. tanenhaus (2004). on-line sentence processing: past, present and, future. the on-line study of sentence comprehension: erps, eye movements and beyond, eds. manuel carreiras and charles clifton, jr., pages 371-392. psychology press, new york, new york. michael k. tanenhaus, austin f. frank, t. florian jaeger, anne pier salverda, mikhail masharov (2008, march). the art of the state: mixed effects regression modeling in the visual world. paper presented at the 21st annual cuny conference on human sentence processing, chapel hill, north carolina. steve tauroza and desmond allison (1990). speech rates in british english. applied linguistics, 11:90-105. john c. trueswell and michael k. tanenhaus (1994). toward a lexicalist framework for constraint-based syntactic ambiguity resolution. perspectives on sentence processing, eds. charles clifton, jr., lyn frazier, and keith rayner, pages 155-179. lawrence erlbaum associates, hillsdale, new jersey. incremental perspective taking 33 john c. trueswell, irena sekerina, nicole m. hill, and marian l. logrip (1999). the kindergarten-path effect: studying on-line sentence processing in young children. cognition, 73:89-134. jos j. a. van berkum (2009). language and affect: how comprehension depends on how we feel and what we care about. 15th annual conference on architectures and mechanisms for language processing (amlap 2009). barcelona, spain. duane g. watson, michael k. tanenhaus, christine gunlogson (2008). interpreting pitch accents in on-line comprehension: h* vs l h*. cognitive science, 32:1232-1244. justine b. wells, morten h. christiansen, david s. race, daniel j. acheson, maryellen c. macdonald (2009). experience and sentence processing: statistical learning and relative clause comprehension. cognitive psychology, 58:250-271. deanna wilkes-gibbs and herbert h. clark (1992). coordinating beliefs in conversation. journal of memory and language, 31:183-194. michael p. wilson and susan m. garnsey (2009). making simple sentences hard: verb bias effects in simple direct object sentences. journal of memory and language, 60:368-392. shali wu and boaz keysar (2007). the effect of culture on perspective taking. psychological science, 18(7):600-606. eiling yee and julie c. sedivy (2006). eye movements reveal transient semantic activation during spoken word recognition. journal of experimental psychology: learning, memory and cognition, 32:1-14. dialogue & discourse 7(3) (2016) 47–64 doi: 10.5087/dad.2016.302 dialog history construction with long-short term memory for robust generative dialog state tracking byung-jun lee bjlee@ai.kaist.ac.kr school of computing, college of engineering korea advanced institute of science and technology 291 daehak-ro, yuseong-gu, daejeon, republic of korea kee-eung kim kekim@cs.kaist.ac.kr school of computing, college of engineering korea advanced institute of science and technology 291 daehak-ro, yuseong-gu, daejeon, republic of korea editor: jason d. williams, antoine raux, and matthew henderson submitted 04/15; accepted 02/16; published online 04/16 abstract one of the crucial components of dialog system is the dialog state tracker, which infers user’s intention from preliminary speech processing. since the overall performance of the dialog system is heavily affected by that of the dialog tracker, it has been one of the core areas of research on dialog systems. in this paper, we present a dialog state tracker that combines a generative probabilistic model of dialog state tracking with the recurrent neural network for encoding important aspects of the dialog history. we describe a two-step gradient descent algorithm that optimizes the tracker with a complex loss function. we demonstrate that this approach yields a dialog state tracker that performs competitively with top-performing trackers participated in the first and second dialog state tracking challenges. keywords: dialog state tracking, dialog management, spoken dialog system (sds), hidden information state (his), long short term memory (lstm). 1. introduction spoken dialog systems (sds) are a rapidly growing subject of research with the grand challenge of developing intelligent systems capable of understanding conversational commands by human users. the generic architecture of a spoken dialog system is shown in figure 1: the system decodes the user voice into the list of hypotheses in a form of action type and slot-value pairs (see table 1 for details), which is defined as the spoken language understanding (slu). the system then analyzes the hypotheses to decide how to respond to the user, which is referred to as dialog management (dm). the dm can be further decomposed into two sub-tasks: the dialog state tracking that identifies the user goal and the dialog policy optimization that finds the appropriate response. the actual vocal response is produced by the speech generation. due to the inevitable misinterpretation of the user utterance (i.e. ambiguity) in slu, robust dm is essential in any dialog system. since the task of determining an appropriate response based on the history of slu data aligns well with reinforcement learning framework, a large number of previous studies were based on the markov decision process (mdp, levin et al. 1998; daubigney c©2016 byung-jun lee and kee-eung kim this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). byung-jun lee and kee-eung kim figure 1: a diagram of general spoken dialog systems et al. 2012) and the partially observable markov decision process (pomdp, williams and young 2007; young et al. 2010; gasic and young 2014). one of the most representative work is the sdspomdp (young, 2006). it treats slu data as noisy and partial observations of the underlying state of dialog that encapsulates sufficient information for dm: user’s intention, utterance and the dialog history. actions of the sds-pomdp are defined as the abstract system responses, one of which will be selected to maximize the expected accumulation of rewards. in this fashion, the sds-pomdp provides a unified decision theoretic model for dm. given a pre-determined dialog strategy, sds-pomdp tracks the dialog state using a bayesian filter. this is often referred to as a generative dialog state tracker, which facilitates identifying core components that can be engineered individually through standard statistical machine learning techniques. however, the performance of the optimized tracker still critically depends on what information is captured from dialog for bookkeeping. if an important aspect of the dialog is not captured and kept appropriately, it is evident that the tracker would not be so accurate in inferring the user’s intention. in most of the previous work, this bookkeeping was done through heuristically selecting the information and embedding it into the dialog history (young et al., 2010). on the other hand, we note that there are a number of recent approaches using discriminative models, which do not involve explicit bookkeeping mechanism (metallinou et al., 2013; lee, 2013; williams, 2014). they use features that are defined over the finite window of previous dialog turns. a more recent work using recurrent neural network (henderson et al., 2014b) enables learning with features that are defined over the window of arbitrary length. in these approaches, the bookkeeping mechanism can be viewed as implicit and parameterized, which is optimized through training. this in general made the tracker perform much better than generative ones. in this paper, we present a generative dialog state tracker that uses long-short term memory (lstm), which is one of the deep-learning architectures for recurrent neural network that overcomes the major drawback of original recurrent neural network called vanishing gradients (bengio et al., 1994; hochreiter and schmidhuber, 1997). with this approach, we can preserve the intuitiveness and compositionality of generative dialog state trackers while autonomously bookkeeping appropriate information over time, which results a boost in performance. this is especially useful in practical applications, where we typically have some prior knowledge on the characteristics of the dialog domain and speech and language processing unit. 48 dialog history construction for robust dialog state tracking our method extends the bayesian filter in sds-pomdp so that the dialog history is replaced with the embedding vector computed by lstm. we then propose a two-step optimization algorithm to deal with the local convergence when training the model. our tracker performed better than any tracker from the dialog state tracking challenge 1 (dstc1, williams et al. 2013)1 and on par with top-performing trackers from the dialog state tracking challenge 2 (dstc2, henderson et al. 2014a)2. this paper is an extension of kim et al. (2013); lee et al. (2014) using lstm for the dialog history. the rest of the paper is organized as follows. section 2 describes the bayesian filtering model of the sds-pomdp, with the additional techniques we adopted in this work. section 3 describes the component probability models outlined in section 2. the optimization algorithm designed for our tracker is explained in section 4. in section 5, the experiment setup, the result and analyses are described. we conclude the paper with discussions in section 6. 2. bayesian filtering for the dialog state tracking we start from the basic setting of the sds-pomdp framework: in each turn of the dialog, the system executes system action a, and the user with goal g responds to the system with user action u. the slu module processes the user action and generates the result as an n -best list (observation, in other words) o = [〈ũ1, f1〉, . . . , 〈ũn , fn 〉] of the hypothesized user action ũi and its associated confidence score fi. table 1 shows a dialog where the slu generates n -best lists from noisy user actions and the system responds to the user after tracking true user goals from the slu output. here, the dialog state in each turn is defined as s = (u, g, h), where h is the dialog history encapsulating additional information needed for tracking the dialog state (williams, 2008). because the system cannot exactly identify the user goal, it maintains a posterior probability distribution over user goals, called a belief. the belief over the dialog states is updated by bayesian filtering: b′(u′, g′, h′) = ηpr(o|u′, g′, h′, a′, a) ∑ u,g,h pr(u′, g′, h′|u, g, h, a, a′)b(u, g, h) (1) where η is the normalization constant. with reasonable independence assumptions, the joint probabilities in the equation can be factorized and approximated into simpler forms. further assuming that the user goal does not change, we obtain pr(o|u′, g′, h′, a′, a) ≈ pr(o|u′) (2) pr(u′, g′, h′|u, g, h, a′, a) ≈ pr(u′|g′, h′, a′) pr(g′|u′, g, h, a) pr(h′|u, g, h, a) (3) ≈ pr(u′|g′, h′, a′) pr(h′|u, g, h, a)i(g = g′) b′(u′, g′, h′) = ηpr(o|u′) pr(u′|g′, h′, a′) ∑ u,h pr(h′|u, g, h, a)b(u, g, h) (4) where pr(o|u′) is the observation model, pr(u′|g′, h′, a) is the user action model, pr(h′|u, g, h, a) is the history model. according to the formulation, the posterior computation has to be carried out for all possible user goals in order to obtain the normalizing constant η. this is not feasible for 1. http://research.microsoft.com/en-us/events/dstc 2. http://camdial.org/ mh521/dstc/ 49 byung-jun lee and kee-eung kim table 1: an example of a dialog depicted in dstc1 system east pittsburg bus schedules. say a bus route, like 28x, or say i’m not sure. hello(), request(route), example(route=28x), example(route=dont know) user 61a slu 0.77 inform(route=61a) 0.12 inform(route=61) 0.01 inform(route=61d) system okay, 61a. to change, say go back. where are you leaving from? impl-conf(route=61a), example(act=goback), request(from) user downtown slu 0.59 inform(from.desc=downtown) 0.10 inform(from.desc=from downtown) system okay, downtown. you can always say go back. and where are you going to? impl-conf(from.desc=downtown), example(act=goback), request(to) user east pittsburgh east pittsburgh slu 0.25 inform(to.desc=pittsburgh) 0.20 inform(to.desc=east pittsburgh) 0.01 inform(to.desc=to pittsburgh) system pittsburgh. is that right? expl-conf(to.desc=pittsburgh) user no slu 0.71 negate() 0.07 affirm() system oh, sorry, where are you going to? sorry(), request(to) user east pittsburgh slu 0.97 inform(to.desc=east pittsburgh) system okay, east pittsburgh. do you want times for the next few buses? say yes or no. impl-confirm(to.desc=east pittsburgh) , expl-confirm(act=want-next-bus) example(act=affirm), example(act=negate) user yes slu 0.97 affirm() system ok, at 7:45 pm today, there is a 61 a, from fifth avenue and market street, downtown. schedule(result.from.time.hour=7, result.from.time.minute=45, result.from.time.ampm=pm, result.from.date.day=today, result.route=61a, result.from.desc=fifth avenue and market street, downtown real dialog domains that have a large number of user goals as in dstc1 and dstc2 domains. as such, we adopt the goal partitioning approach in his (young et al., 2010) that groups user goals into equivalence classes and assigns a single probability for each class. the equivalence classes are called partitions and denoted by ψ. beginning with one root partition with the probability 1, partitions are split whenever the distinction is required by observations, i.e. a user action hypothesis from the slu output. this confines 50 dialog history construction for robust dialog state tracking the possible goal states to the values that have been appeared at least once as an slu hypothesis, and provides scalability without a significant loss in accuracy when the coverage of n-best list is extensive enough to include the true user action. by defining belief refinement model pr(ψ′|ψ) to be a probability mass ratio that the partition splits, the bayesian filtering equation (4) becomes: b′(u′, ψ′, h′) = ηpr(o|u′) pr(u′|ψ′, h′, a′) ∑ u,h pr(h′|u, ψ, h, a) pr(ψ′|ψ)b(u, ψ ⊃ ψ′, h) (5) where ψ ⊃ ψ′ denotes that ψ is a parent partition (superset) of ψ′. even with the partitioned approach, the total number of partitions grows very fast as dialog turn progresses (exponential in number of observed goal types, polynomial in number of observed goals), and it is required to give certain limit to the number of partitions. we used the incremental partition recombination algorithm (williams, 2010), which recombines the less important partitions in each turn if the number of partitions exceeds the threshold. 3. designing component probability models the equation (5) needs a model for each component probability term, i.e. observation model, user action model, history model, and belief update model. although assigning simple statistics to each component probability model (young et al., 2010; kim et al., 2013) can be a valid choice, lee et al. (2014) has shown that designing each probabilities appropriately and optimizing it results in much better performance. in following subsections, each component probability model is described. 3.1 observation model the observation model pr(o|u) represents the probability of structured observation o = [〈ũ1, f1〉, . . . , 〈ũn , fn 〉] given user action u. here, we assume that the observation model only depends on the type of user action u and its associated confidence score fi to obtain a simple model. then, we obtain pr(o|u = ũi) = k · (exp(wtype(ũi))fi + exp(btype(ũi))) (6) where wtype(u) is the weight associated with specific user action type of u and btype(u) is bias term for the same user action type, which are exponentiated for preserving non-negativity. k is the normalizing constant, which can be ignored since it is subsumed by the constant η in the belief update equation (5). the user action type here is domain-specific attribute of slu hypothesis, which is given in prior (e.g. inform, affirm, negate in table 1). refer to dstc1 or dstc2 handbook for full description of user and system action types. 3.2 belief refinement model the belief refinement model pr(ψ′|ψ) defines the split ratio of partitions when the partition is required to split due to the observations. according to the definition, the full specification of the belief refinement model corresponds to the prior distribution of partitions, which is summation of prior probability of each individual user goals included in partitions. we define the prior on goals pr(g) as a smoothed empirical distribution pr(g) = ε #turns having goal g #turns + (1− ε) 1 #individual goals (7) 51 byung-jun lee and kee-eung kim with the prior probability of each individual user goals, we can compute belief refinement model by summing the prior probabilities of user goals to get the ratio pr(ψ′|ψ) = ∑ g′∈ψ′ pr(g′)∑ g∈ψ pr(g) (8) if the ψ is divided to ψ′ and other partitions this turn by observation, pr(ψ′|ψ) = 0 otherwise. 3.3 user action model and history model in previous approaches (young et al., 2010; lee et al., 2014), the user action model pr(u′|ψ′, h′, a′) only utilized the features of current turn and relied on manually defined bookkeeping methods for h′. this is potentially disadvantageous compared to trackers that are able to learn which information to maintain over multiple dialog turns (metallinou et al., 2013; lee, 2013; williams, 2014; henderson et al., 2014b). this disadvantage becomes apparent in complex domains like dstc2, where the conversation topics change over time. in dstc2, in which the domain was a restaurant information system, the user usually starts with providing search constraints and then proceeds to the stage of choosing the restaurant to be informed. learning such phase transition in dialogs where the user acts differently is therefore crucial for designing robust user action model. recurrent neural networks provide a natural model for handling such problems, as they are able to model any dynamic sequences with complex behaviors by learning to choose what information to keep. adopting recurrent neural network in dialog state tracking has proven its success in henderson et al. (2014b), while our approach is more focused on learning the dependencies over dialog turns in a generative state tracker, encoding them as the dialog history. instead of the conventional recurrent neural network, we adopt one of the variants, namely long-short term memory (lstm, bengio et al. 1994; hochreiter and schmidhuber 1997), which effectively mitigates the vanishing gradient problem in training recurrent neural networks. 3.3.1 lstm architecture used for user action model the input to lstm is a feature constructed from u, ψ and a. we used every combinations of [user action type, system action type, partition-action consistency] that appears in the training set as a feature set, which results in maximum 1 of |u| · |a| · 4 coding3. in the following equations of lstm, • φ(ut, ψt, at) is 1 of |u| · |a| · 4 coding feature vector, which is the input to lstm. • ct−1, ct are the cell values stored in lstm. • yt−1, yt are the outputs of lstm, and also stored. • wp,wi,wf ,wc,wo, ui, uf , uc, uo, vi, vf , vo, wd, wl are weight matrices (or vectors) to be learned. • bi, bf , bc, bo are bias vectors. 3. since we only considered the feature combinations that appear in the training set, the vector length is much less than |u| · |a| · 4 because of the nonexistent system action-user action combination. also, due to the dual or more action types, there are the cases where it is not exactly 1-of-k coding. 52 dialog history construction for robust dialog state tracking figure 2: the long-short term memory architecture used to model user action model. from the usual lstm architecture, a linearly compressing input layer and direct edge from input to output is added. the gray colored cell vector is maintained through the time. the input layer is passed through the input gate to compute a new cell value. the cell value is computed by adding input and the previous cell if not forgotten, as follows. it = σ(wiφ(u, ψt, at) + uiyt−1 + vict−1 + bi) (9) c̃t = tanh(wcφ(u, ψt, at) + ucyt−1 + bc) (10) ft = σ(wfφ(u, ψt, at) + ufyt−1 + vfct−1 + bf ) (11) ct = it ∗ c̃t + ft ∗ ct−1 (12) where ∗ denotes the elementary multiplication between vectors. output is then computed by the output gate: ot = σ(woφ(u, ψt, at) + uoyt−1 + voct + bo) (13) yt = ot ∗ tanh(ct) (14) we finally get the user action model value pr(ut|ψt, ht, at) by adding the direct edge from feature function and using a softmax function, pr(ut|ψt, ht, at) = exp(wtl yt(φ(ut, ψt, at)) + wtd φ(ut, ψt, at))∑ u exp(wtl yt(φ(u, ψt, at)) + wtd φ(u, ψt, at)) . (15) 53 byung-jun lee and kee-eung kim where yt on the right hand side of the equation is implicitly dependent on ht via the information stored in lstm. in practice, however, we cannot sum up the lstm outputs for every u. instead, we only evaluate the lstm for ũis from n -best list, and evaluate ũofflist for the remaining actions. 3.3.2 history model and m -best lstms according to the model above, what we are maintaining as a history state ht is the cell layer and the output layer of the last turn’s lstm, in which the sequence [u1, a1, ψ1, . . . , ut−1, at−1, ψt−1] is implicitly embedded. the history model pr(ht|ut−1, ψt−1, ht−1, at−1) is therefore deterministic, and the bayesian filtering equation is then b′(ψ′, h′) = η ∑ u′ pr(o|u′) pr(u′|ψ′, h′, a′) ∑ h,u pr(h′|u, ψ, h, a) pr(ψ′|ψ)b(u, ψ ⊃ ψ′, h) = η ∑ u′ pr(o|u′) pr(u′|ψ′, h′, a′) ∑ h,u i(h′ = [u, ψ, a, h]) pr(ψ′|ψ)b(u, ψ ⊃ ψ′, h) = η ∑ u′ pr(o|u′) pr(u′|ψ′, h′, a′) pr(ψ′|ψ)b(ψ ⊃ ψ′, h ⊃ h′) (16) where h ⊃ h′ denotes that h is a parent history of h′ (i.e. h′ = {h, a, u, ψ}). this implies that we should maintain |h||ψ| partitions and |h| lstms each turn to exactly compute posterior probability over user goals. however, it is impractical due to the exponentially increasing size of history. to gain a better insight, the following equation is what to be evaluated for the dialog state tracking problem: b′(ψ′) = η ∑ u′ pr(o|u′) pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ ⊃ ψ′, h ⊃ h′) (17) it can be seen that we are taking the weighted average of lstms over the joint belief b(u ∈ h′, ψ ⊃ ψ′, h ⊃ h′) for every h′. similar to the partition recombination technique (williams, 2010) that limits the number of partition to certain constant, we can approximate it by limiting the number of histories by m (equivalently, m lstms). since the aggregation of histories are not possible, other than m histories with the best belief are ignored. a belief update with m -lstms is performed, as specified in the following algorithm 1. 54 dialog history construction for robust dialog state tracking data: system action a, observation o = {(ũ1, f1), ..., (ũn , fn )}, partitions ψ = {ψ1, ...ψp }, lstms l = {l1, ..., lm}, histories h = {h1, ..., hm} and the belief b(ψ, h) ∀ψ, h result: partitions ψ′, lstms l′, histories h ′ and the belief b′ ψ′ = {}; h ′ = {} for ũi in o do for ψj in ψ do if ũi need to split ψj then ψ′ = ψ′ ∪ {ψj ∧ ũi, ψj ∧ ¬ũi}; calculate pr(ψj ∧ ũi|ψj) by equation (8); for hk in h do b(ψj ∧ ũi, hk) = pr(ψj ∧ ũi|ψj)b(ψj , hk); b(ψj ∧ ¬ũi, hk) = b(ψj , hk)− b(ψj ∧ ũi, hk) b(ψj , hk) = 0; end end end end b′(ψ, h) = 0 ∀ψ, h; for ψi in ψ′ do for ũj in o do calculate pr(o|u = ũi) by equation (6); for hk in h do h′k = [ũj , a, ψi, hk]; h ′ = h ′ ∪ {h′k}; calculate pr(ũj |ψi, hk, a) by equation (15) and lstm lk; b′(ψi, h ′ k) = b′(ψi, h ′ k) + pr(o|u = ũi) pr(ũj |ψi, hk, a)b(ψi, hk) end end end if |h ′| > m then order h ′ by decreasing ∑ ψ b ′(ψ, h); for hk in h ′, k > m do h ′ = h ′ − {hk}; b′(ψ, hk) = 0 ∀ψ end end l′ = updated l to contain embedded histories of h ′; if |ψ′| > p then run partition recombination according to williams (2010) end b′(ψ, h) = b′(ψ,h)∑ ψ,h b ′(ψ,h) ∀ψ, h algorithm 1: updating the belief over partitions and histories 55 byung-jun lee and kee-eung kim 4. optimization in this paper, we use l2 metric as the loss function since it is found to be most influential to dialog system performance (lee, 2014). the model is hence optimized to minimize the l2 loss function l(w) = 1 2 n∑ i=1 t∑ t=1 ∑ ψ∈ψi,t (r(ψ)− b(ψ))2 (18) where i sums overn training instances, t sums over t turns of each training instance andψi,t is the group of partitions in the instance i, turn t. r(ψ) is the binary label with value 1 if and only if the partition ψ contains the true user goal. in the rest of the section, gradient methods used to optimize with respect to the loss function are described. 4.1 cascading gradient taking the insight from the bptt algorithm (mozer, 1989), we can unfold the update formula through time and calculate the gradient with respect to weight vectors. for instance, considering a parameter w from the observation model, ∂l/∂w can be obtained by: ∂l ∂w = n∑ i=1 t∑ t=1 ∑ ψ∈ψi,t (b(ψ)− r(ψ)) ∂b(ψ) ∂w (19) ∂b′(ψ′) ∂w = ∂η ∂w ∑ u′∈o pr(o|u′) pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ ⊃ ψ′, h ⊃ h′) + η ∑ u′∈o ∂ pr(o|u′) ∂w pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ ⊃ ψ′, h ⊃ h′) + η ∑ u′∈o pr(o|u′) pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)∂b(ψ ⊃ ψ ′, h ⊃ h′) ∂w = η ∑ u′∈o ∂ pr(o|u′) ∂w pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ ⊃ ψ′, h ⊃ h′) − b(ψ′) ∑ ψ̄∈ψi,t η ∑ u′∈o ∂ pr(o|u′) ∂w pr(ψ′|ψ̄) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ̄ ⊃ ψ′, h ⊃ h′) + η ∑ u′∈o pr(o|u′) pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)∂b(ψ ⊃ ψ ′, h ⊃ h′) ∂w − b(ψ′) ∑ ψ̄∈ψi,t η ∑ u′∈o pr(o|u′) pr(ψ′|ψ̄) ∑ h′ pr(u′|ψ′, h′, a′)∂b(ψ̄ ⊃ ψ ′, h ⊃ h′) ∂w (20) we call the above cascading gradient since ∂b′(ψ)/∂w requires computation of the gradient in the previous dialog turn ∂b(ψ)/∂w, and hence reflects the temporal impact of the parameter change throughout the dialog turns. once we obtain the gradients, we can simultaneously update all the parameters with any algorithm. in this paper, l-bfgs was used. 56 dialog history construction for robust dialog state tracking 4.2 initialization using simple gradient the l-bfgs using cascading gradient, however, seriously suffers from the local convergence since the cost function is very complex in parameter space, mostly due to the repeated appearance of features throughout dialog turns. the randomized initialization had a limited effect because of the large dimensionality of parameter space. if we ignore the gradient from previous turn and treat each turn as individual training instance, the gradient becomes, for example of a parameter w from the observation model, ∂b′(ψ′) ∂w = ∂η ∂w ∑ u′∈o pr(o|u′) pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ ⊃ ψ′, h ⊃ h′) + η ∑ u′∈o ∂ pr(o|u′) ∂w pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ ⊃ ψ′, h ⊃ h′) = η ∑ u′∈o ∂ pr(o|u′) ∂w pr(ψ′|ψ) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ ⊃ ψ′, h ⊃ h′) − b(ψ′) ∑ ψ̄∈ψi,t η ∑ u′∈o ∂ pr(o|u′) ∂w pr(ψ′|ψ̄) ∑ h′ pr(u′|ψ′, h′, a′)b(ψ̄ ⊃ ψ′, h ⊃ h′) (21) since it is an approximation that only catches the direct impact of the features and makes loss function much simpler, it takes the parameters close to the global optimum, increasing the chance of l-bfgs converging to the good local optimum. our optimization algorithm is hence composed of two steps, first initializing with the simple gradient and then training with the cascading gradient. 5. experiment and analysis table 2: description of the dstc1 datasets used for the tracker (bus information system) dataset calls similarity slot train1a 1013 similar to train2 9 train2 678 similar to train1a 9 train3 779 distinct from train1a and 2 5 test1 765 very similar to train1a and 2 9 test2 983 similar to train1a and 2 5 test3 1037 very similar to train3 5 in the experiments, we used datasets from both dstc1 and dstc2. three labeled training datasets (train1a, train2, train3) and four test datasets (test1, test2, test3, test4) were included in dstc1 whereas two labeled training datasets (dstc2 train, dstc2 dev) and one test dataset (dstc2 test) were included in dstc2. description of the datasets is provided in table 2 and table 3. test4 is omitted in this paper since the significant number of missing or incorrect labels were found. we can group the datasets based on their characteristics. datasets train1a, train2, test1, and test2 have many turns in each call. however, they include only one hypothesis in each dialog turn 57 byung-jun lee and kee-eung kim table 3: description of the dstc2 datasets used for the tracker (restaurant information system) dataset calls similarity slot dstc2 train 1612 very similar to dev 4 dstc2 dev 506 very similar to train 4 dstc2 test 1117 similar to train and test 4 (1-best slu output) and the user goal rarely changes. these datasets can be said to be relatively easy since there is not much information to examine in inferring the user goal. on the other hand, datasets train3, test3, dstc2 test, dstc2 dev and dstc2 test are considered harder. these datasets contain dialogs that are closer to real world conversation with latent relations among system actions and user actions and dialog flows. dependencies that span over multiple dialog turns are more frequent and they even change over turns in dstc2 datasets. for each target test dataset, we chose the training datasets that are known to be similar. specifically, train1a and train2 are used to train the tracker that tests test1 and test2, whereas train3 is used for test3 and dstc2 train, dstc2 dev are used for dstc2 test. we only used the slu data for observations (i.e. ignored asr information), and our result is therefore compared with teams that only used slu data. note that the evaluations of our algorithm are taken after the release of the test set while the other teams’ results are evaluated before the release of the test set. we measured the tracker performance according to the following evaluation metrics used in dstc14: • accuracy (acc) measures the rate of the most likely hypothesis h1 being correct. • average score (avgp) measures the average of scores assigned to the correct hypotheses. • l2 follows the definition in the equation (18). • mean reciprocal rank (mrr) measures the average of 1/r, where r is the minimum rank of the correct hypothesis. • roc equal error rate (eer) is the sum of false accept (fa) and false reject (fr) rates when fa rate=fr rate. • roc.v1.p measures correct accept (ca) rate when there are at most p% false accept (fa) rate. n(fa), n(cr), n(ca) and n(fr) are the number of false accepts (fa), correct rejects (cr), correct accepts (ca), and false rejects (fr) respectively. additionally, nd is the total number of data instances. the evaluation takes the most likely hypothesis h1 and its score s1 and compares them with the threshold θ and the ground truth user goal h∗. each evaluation increments appropriate counters by  n(ca)++ if s1 ≥ θ and h1 = h∗ n(cr)++ if s1 < θ and h1 6= h∗ n(fa)++ if s1 ≥ θ and h1 6= h∗ n(fr)++ if s1 < θ and h1 = h∗ 4. http://research.microsoft.com/apps/pubs/?id=169024 58 dialog history construction for robust dialog state tracking the ca rate is defined as n(ca) nd and the fa rate is defined as n(fa) nd . roc.v2.p from the dstc1 metrics is not included in the experiment. with datasets, the baseline trackers that work in a simple deterministic manner was included. for example, the most basic one picks the slu hypothesis with the highest probability so far for each goal slot. more detailed description is given in dstc2 handbook. 5.1 dstc1 result table 4 compares the performance of our tracker against that of the most competitive entries from dstc1 participants. base stands for the baseline tracker and x-y stands for team x entry y . the following is the summary description provided by the teams: • team 1 : deep neural network (henderson et al., 2013) • team 6 : feature-rich discriminative model (lee and eskenazi, 2013) • team 9 : generative dialog state tracker with basic statistics (kim et al., 2013) the following abbreviations are used for our tracker with different history models and optimization algorithms: • c: direct edge only user action model optimized with cascading gradient. • sc: direct edge only user action model initialized with simple gradient and optimized with cascading gradient. • lc: lstm history based user action model optimized with cascading gradient. • lsc: lstm history based user action model initialized with simple gradient and optimized with cascading gradient. we choose the trackers with the minimum training objective function among 100 random seeds. the number of partitions is limited to 10, and the number of histories (lstm, equivalently) is limited to 3. 50 units are used for all (input, cell, output) layers of lstms. this setting of parameters enable the real-time tracking of dialog state, while preserving most of the information required. overall, the trackers initialized with simple gradients outperform the trackers without the initialization process. it can be seen that the local convergence problem of cascading gradient is most severe in test3, as the problem gets more complex even though the other domains also show the clear difference in scores. the trackers optimized with two different gradients outperform the scores of other teams participated in dstc1. in case of history construction using lstm, dstc1 test1 and test2, which are categorized as easy dataset, are not showing any performance gain from the history construction. this is expected due to the fact that the dialogs of test1 and test2 are so simple that keeping extra information from dialog history would not help. the results of test1 by the trackers with the lstms are even worse than the trackers without the lstms since the increased number of parameters raises the problem of over-fitting. on the other hand, distinct difference between trackers with and without lstms is observed in test3. although the accuracies are similar (0.891↔0.898), l2 scores show the clear decrease of 0.189→ 0.172 with lstms when optimized with two different gradients. 59 byung-jun lee and kee-eung kim table 4: results of the trackers of dstc1 and our tracker using various models and optimization algorithms, evaluated by the average of all slots. the bold face denotes top score in each evaluation metric. base 1-1 2-2 3-1 4-1 5-1 6-3 7-1 9-4 c sc lc lsc dstc1 test 1 accuracy 0.712 0.832 0.807 0.808 0.737 0.795 0.867 0.783 0.822 0.819 0.870 0.817 0.868 avgp 0.733 0.774 0.771 0.807 0.737 0.787 0.823 0.762 0.794 0.796 0.841 0.798 0.837 l2 0.377 0.319 0.322 0.273 0.372 0.300 0.246 0.335 0.290 0.287 0.223 0.285 0.229 mrr 0.797 0.875 0.858 0.846 0.813 0.852 0.900 0.843 0.878 0.873 0.910 0.870 0.907 roc.v1.eer0.244 0.126 0.246 0.243 0.737 0.122 0.118 0.147 0.143 0.183 0.146 0.173 0.145 roc.v1.05 0.622 0.723 0.672 0.601 0.196 0.710 0.763 0.650 0.720 0.677 0.748 0.681 0.744 dstc1 test 2 accuracy 0.546 0.646 0.707 0.683 0.635 0.622 0.790 0.652 0.705 0.834 0.856 0.834 0.865 avgp 0.573 0.550 0.629 0.684 0.634 0.615 0.714 0.649 0.651 0.824 0.846 0.823 0.842 l2 0.603 0.633 0.503 0.446 0.517 0.535 0.386 0.492 0.476 0.245 0.213 0.247 0.220 mrr 0.650 0.717 0.792 0.756 0.713 0.722 0.843 0.744 0.797 0.882 0.900 0.879 0.904 roc.v1.eer0.192 0.197 0.394 0.144 0.635 0.212 0.159 0.189 0.219 0.129 0.115 0.124 0.128 roc.v1.05 0.431 0.487 0.516 0.452 0.164 0.480 0.660 0.479 0.490 0.761 0.793 0.764 0.786 dstc1 test 3 accuracy 0.789 0.793 0.843 0.819 0.819 0.779 0.835 0.790 0.847 0.819 0.891 0.823 0.898 avgp 0.751 0.725 0.757 0.787 0.784 0.701 0.752 0.755 0.740 0.805 0.862 0.809 0.875 l2 0.352 0.369 0.323 0.291 0.295 0.395 0.334 0.337 0.343 0.271 0.189 0.266 0.172 mrr 0.835 0.851 0.883 0.853 0.853 0.828 0.890 0.841 0.887 0.854 0.919 0.859 0.924 roc.v1.eer0.189 0.164 0.154 0.273 0.124 0.171 0.148 0.119 0.129 0.123 0.101 0.122 0.095 roc.v1.05 0.565 0.647 0.681 0.724 0.702 0.623 0.687 0.701 0.738 0.749 0.840 0.752 0.855 5.2 dstc2 result similar to the previous section, we compare our tracker to other trackers submitted to dstc2. we do not include trackers that additionally use asr (automatic speech recognition) information in the comparison; we restricted comparison to trackers that only used nlu data as ours for fair comparison. the following is the summary description provided by the teams: • team 1 : linear-chain conditional random field (kim and banchs, 2014) • team 4 : recurrent neural network (henderson et al., 2014b) • team 6 : maximum-entropy markov model (ren et al., 2014) • team 7 : combined model of rule-based, maximum-entropy and deep neural network (sun et al., 2014) 60 dialog history construction for robust dialog state tracking table 5: results of the trackers of dstc2 and our tracker using various models and optimization algorithms, evaluated on the joint slot. the bold face denotes top score in each evaluation metric. base 1-0 3-0 4-3 6-2 7-4 9-0 4 rep c sc lc lsc dstc2 test accuracy 0.719 0.601 0.729 0.737 0.718 0.735 0.499 0.735 0.703 0.728 0.710 0.741 avgp 0.678 0.503 0.659 0.636 0.638 0.673 0.522 0.641 0.587 0.662 0.605 0.677 l2 0.464 0.649 0.452 0.406 0.437 0.433 0.760 0.402 0.465 0.408 0.441 0.394 mrr 0.779 0.661 0.763 0.804 0.772 0.787 0.608 0.801 0.756 0.783 0.780 0.809 roc.v1.eer0.332 0.096 0.320 0.461 0.432 0.349 0.313 0.323 0.327 0.352 0.346 0.331 roc.v1.05 0.256 0.382 0.249 0.208 0.226 0.243 0.000 0.223 0.245 0.251 0.236 0.285 evaluation is taken over the joint slot5, which was the featured metric in dstc2, that checks the joint correctness of every goal slot. lstm parameter configuration identical to dstc1 experiment was also used for this experiment. since this is a complex dialog similar to test3 dataset in dstc1, the initialization with simple gradient showed a significant performance improvement in both models, with and without lstms. on the other hand, lstm yielded additional significant performance improvement than in dstc1 datasets, successfully capturing more complex dependencies over dialog turns. while c, sc, lc yielded competitive scores with other teams, lsc scored the highest among the trackers using slu. note that team 4 by henderson et al. (2014b) adopted a recurrent neural network and results a very similar scores (0.737↔0.741). for a detailed comparison, we replicated the work of team 4 referring to the description in the henderson’s thesis (henderson, 2015), and presented above as 4 rep. it can be seen that 4 rep is also achieving a very similar score, and the differences in three systems(4-3, 4 rep, and lsc) do not seem significant if the randomness of the learning algorithm is considered. however, we found out that our system requires much less computational cost when compared to the replicated system. the model itself used ten times smaller number of parameters (3 × 105 parameters used for 4 rep while 3 × 104 parameters used for lsc) to achieve similar or better performance due to the careful choice of component probability models, encoding prior knowledge. 6. conclusion and discussion in this paper, we proposed a robust generative dialog state tracker by using long-short term memory (lstm) for the dialog history. based on the bayesian filtering in the sds-pomdp framework, we designed each component probability model to capture important dependencies in dialog state tracking. lstm is used for user action model, aimed at learning complex system-user action dependencies over time, and has exhibited its performance improvment in complex domains where bookkeeping of important dialog aspects is important for accurate tracking. we also proposed the two-step optimization algorithm that optimizes the tracker. the performance of a local optimization algorithm used to train the tracker was highly dependent on the initial 5. the performance gain between trackers hence cannot be directly compared with the gains from dstc1. it is harder to get better score in joint slot than the averaged score on individual slots. 61 byung-jun lee and kee-eung kim solution, as the objective function is highly nonlinear in the parameter space. we therefore introduced a preliminary optimization stage where we use the gradient descent with approximated gradients calculated by ignoring the dependencies over time steps. the solution from the preliminary optimization was used as the initial solution for the second stage, where we calculated the exact gradients by dynamic programming. this two-step algorithm consistently improved the performance over single-step optimization approach with exact gradient and random initial solution. we have demonstrated the performance of the tracker and the effectiveness of the optimization algorithm by comparing to the state-of-the-art trackers submitted to dialog state tracking challenges 1 and 2. acknowledgments this work was partly supported by the ict r&d program of msip/iitp [14-824-09-014, basic software research in human-level lifelong machine learning (machine learning center)] and national research foundation of korea (grant# 2012-007881). references yoshua bengio, patrice simard, and paolo frasconi. learning long-term dependencies with gradient descent is difficult. neural networks, ieee transactions on, 5(2):157–166, 1994. lucie daubigney, matthieu geist, senthilkumar chandramohan, and olivier pietquin. a comprehensive reinforcement learning framework for dialogue management optimization. selected topics in signal processing, ieee journal of, 6(8):891–902, 2012. milica gasic and steve young. gaussian processes for pomdp-based dialogue manager optimization. ieee/acm transactions on audio, speech and language processing (taslp), 22(1): 28–40, 2014. henderson, blaise thomson, and jason williams. the second dialog state tracking challenge. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), 2014a. matthew henderson, blaise thomson, and steve young. deep neural network approach for the dialog state tracking challenge. in proceedings of the 14th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 467–471, 2013. matthew henderson, blaise thomson, and steve young. word-based dialog state tracking with recurrent neural networks. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), page 292, 2014b. matthew s henderson. discriminative methods for statistical spoken dialogue systems. phd thesis, university of cambridge, 2015. sepp hochreiter and jürgen schmidhuber. long short-term memory. neural computation, 9(8): 1735–1780, 1997. 62 dialog history construction for robust dialog state tracking daejoong kim, jaedeug choi, kee-eung kim, jungsu lee, and jinho sohn. engineering statistical dialog state trackers: a case study on dstc. in proceedings of the 14th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 462–466, 2013. seokhwan kim and rafael e banchs. sequential labeling for tracking dynamic dialog states. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), page 332, 2014. byung-jun lee, woosang lim, daejoong kim, and kee-eung kim. optimizing generative dialog state tracker via cascading gradient descent. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), page 273, 2014. sungjin lee. structured discriminative model for dialog state tracking. in proceedings of the 14th annual meeting of the special interest group on discourse and dialogue (sigdial), 2013. sungjin lee. extrinsic evaluation of dialog state tracking and predictive metrics for dialog policy optimization. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), page 310, 2014. sungjin lee and maxine eskenazi. recipe for building robust spoken dialog state trackers: dialog state tracking challenge system description. in proceedings of the 14th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 414–422, 2013. esther levin, roberto pieraccini, and wieland eckert. using markov decision process for learning dialogue strategies. in proceedings of the 1998 ieee international conference on acoustics, speech and signal processing (icassp), volume 1, pages 201–204, 1998. angeliki metallinou, dan bohus, and jason williams. discriminative state tracking for spoken dialog systems. in proceedings of the 51st annual meeting of the association for computational linguistics (acl), pages 466–475, 2013. michael c mozer. a focused back-propagation algorithm for temporal pattern recognition. complex systems, 3(4):349–381, 1989. hang ren, weiqun xu, and yonghong yan. markovian discriminative modeling for dialog state tracking. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), page 327, 2014. kai sun, lu chen, su zhu, and kai yu. the sjtu system for dialog state tracking challenge 2. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), page 318, 2014. jason williams, antoine raux, deepak ramachandran, and alan black. the dialog state tracking challenge. in proceedings of the 14th annual meeting of the special interest group on discourse and dialogue (sigdial), 2013. jason d williams. exploiting the asr n-best by tracking multiple dialog state hypotheses. in proceedings of the 2008 interspeech, pages 191–194, 2008. 63 byung-jun lee and kee-eung kim jason d williams. incremental partition recombination for efficient tracking of multiple dialog states. in proceedings of the 2010 ieee international conference on acoustics, speech and signal processing (icassp), pages 5382–5385, 2010. jason d williams. web-style ranking and slu combination for dialog state tracking. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), page 282, 2014. jason d williams and steve young. partially observable markov decision processes for spoken dialog systems. computer speech & language, 21(2):393–422, 2007. steve young. using pomdps for dialog management. in proceeding of 2006 ieee workshop on spoken language technology (slt), pages 8–13, 2006. steve young, milica gašić, simon keizer, françois mairesse, jost schatzmann, blaise thomson, and kai yu. the hidden information state model: a practical framework for pomdp-based spoken dialogue management. computer speech & language, 24(2):150–174, 2010. 64 dialogue & discourse 9(2) (2018) 1-34 doi: 10.5087/dad.2018.201 asymmetries between interpretation and production in catalan pronouns laia mayol laia.mayol@upf.edu universitat pompeu fabra editor: vera demberg submitted 03/18; accepted 11/18; published online 12/18 abstract the literature on romance null-subject languages has often postulated a division of labor between null and overt pronouns: nulls prefer to retrieve an antecedent in subject position, whereas overts prefer an antecedent in a lower syntactic position (carminati, 2002). however, recent research on english pronouns (rohde and kehler, 2014) has shown grammatical function alone cannot explain pronoun interpretation. according to this model, pronoun interpretation and production are sensitive to different sets of factors and, instead of being mirror images of each other, are related probabilistically in a bayesian fashion. this paper tests this model with catalan data from two discourse-completion experiments to study the structural, grammatical factors and the semanticopragmatic factors that affect the interpretation and production of null and overt pronouns. our main result is that both null and overt pronouns present asymmetries regarding their interpretation and production: (1) the production of null pronouns is affected mainly by grammatical factors (they are more likely to be used to refer to subjects), but their interpretation is also influenced by semantico-pragmatic factors (rhetorical relations and the verb’s lexical semantics), and (2) while overt pronouns have a strong interpretation bias towards the object, they are not the preferred form to refer to the object. keywords: pronouns, anaphora, null pronouns, overt pronouns, reference, rhetorical relations, catalan 1. introduction the production and interpretation of referring expressions is essential for successful communication: speakers need to choose one particular way to refer to the entities they want to talk about; hearers need to assign a discourse referent to the referring expressions in a discourse. a longstanding idea is that more reduced anaphoric expressions (such as pronouns) tend to be used for more accessible referents, while more complex expressions will be used to refer to less accessible referents (ariel, 1990; givón, 1983; gundel et al., 1993). this raises the question of which specific factors contribute to accessibility, and whether the same factors drive both production and interpretation. a set of factors that are well-known to influence pronoun interpretations are structural, grammatical factors. in particular, subjects make particularly good antecedents (crawley et al., 1990; arnold, 1999), as shown by (1), from kehler et al. (2008). although the semantic content of both sentences is identical, the preferred antecedent of the pronoun changes: the preferred interpretation of the pronoun is that it refers to the antecedent in subject position of the previous clause. c©2018 laia mayol this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). mayol (1) a. bush narrowly defeated kerry, and special interests promptly began lobbying him. [him=bush] b. kerry was narrowly defeated by bush, and special interests promptly began lobbying him. [him=kerry] however, in other occasions, structural factors are overridden by semantico-pragmatic factors and the subject preference is not observed, as illustrated by the minimal pairs in (2). while we still observe a subject preference for (2-a), it disappears in (2-b). this difference is partly due to the verb type: both apologize and scold are so-called implicit causality verbs (caramazza et al., 1977; brown and fish, 1983; stevenson et al., 1994) which attribute the cause of the event to one of the two referents (the subject in the case of apologize, the object in the case of scold). this is precisely the referent to which the pronoun is more likely to refer. these semantic biases are in force in restricted discourse conditions, in particular when the second sentence gives a cause for the event in the first sentence (i.e. when there is a rhetorical relation of explanation between the two sentences). a different rhetorical relation may create a different bias, as shown in (2-c), in which the event of the second sentence follows the event of the first sentence and a narration is established. (2) a. mike apologized to joe because he was late.[he=mike] b. mike scolded joe because he was late. [he=joe] c. mike scolded joe and then he left. [he=mike] rohde and kehler (2014) present a model which aims to capture both structural and (non-structural) semantico-pragmatic factors. in their model, semantico-pragmatic factors (including verb type or rhetorical relations) affect only pronoun interpretation, while structural factors affect both pronoun interpretation and production. thus interpretation and production are fundamentally asymmetrical and this asymmetry can be modeled using bayes’ rules (see section 2.3 for a more full-fledged explanation). while some research has been devoted to study how structural and semantico-pragmatic factors interact in languages such as english, romance languages such as catalan, the language under study in this paper, have an additional aspect to consider given their dual system of pronominal forms: these languages have both overt and null pronouns (rigau, 1986). moreover, the interaction of structural and semantico-pragmatic factors has not been thoroughly studied for either of the two types of pronouns. this paper aims to fill this gap and answer the following questions: 1) how do the structural and semantico-pragmatic factors affect both null and overt pronouns (nulls and overts, henceforth) in catalan? 2) can the behavior of nulls and overts be modeled by a probabilistic model such as the one proposed by rohde and kehler (2014)? this paper is structured as follows. section 2 presents the relevant background on pronoun interpretation and production. sections 3 and 4 present two discourse-completion studies whose aim is to study the production and interpretation of pronouns in catalan. the quantitative data that emerges from these experiments are useful to describe how the two pronouns are used and understood and, moreover, gives us the necessary information to put the probabilistic model to a test. finally, section 5 ends the article with some conclusions. 2 asymmetries in catalan pronouns 2. background 2.1 effect of grammatical factors on pronoun interpretation as mentioned in the introduction, it has recurrently been proposed in the literature that grammatical factors affect the interpretation of pronouns: that is, the grammatical function of a referent can make it a more or less likely antecedent for a pronoun. particularly well-studied is the subjectassignment strategy (crawley et al., 1990), which postulates that a referent mentioned in subject position becomes a more likely antecedent for a pronoun than referents in other syntactic positions. the literature on romance language has used this idea to explain the interpretation of nulls and overts, and many studies have postulated a division of labor between both types of forms. for instance, in several corpus studies, it has been found that nulls and overts carry different biases: a null pronoun prefers a subject antecedent, and an overt pronoun a non-subject antecedent (cameron, 1992; silva-corvalán, 1977). carminati (2002) called this asymmetry the position of antecedent hypothesis (pah), as defined in (3). (3) position of antecedent hypothesis: nulls prefer to retrieve an antecedent in the highest subject position, whereas overts prefer an antecedent in a lower syntactic position. carminati found support for the pah in questionnaire and reading-time experiments for italian. in her experiment 1, participants were presented with two-sentence discourses, in which the second sentence contained a potentially ambiguous null (see (4-a)) or overt (see (4-b)). then they had to answer a question about the referent of such pronoun (4-c). she found clear opposed biases between null and overts: nulls pronouns preferred to retrieve the antecedent in subject position (‘marta’ in the examples), and overts the antecedents in object position (‘piera’ in the examples). (4) a. marta scriveva frequentemente a piera quando ∅ era negli stati uniti. “marta wrote frequently to piera when ∅ was in the united states.” b. marta scriveva frequentemente a piera quando lei era negli stati uniti. “marta wrote frequently to piera when she was in the united states.” c. who was in the states? support for the pah has also been found for other romance languages such as spanish (alonsoovalle et al., 2002; keating et al., 2016) and catalan (mayol and clark, 2010).1 other researchers have proposed that the asymmetry should not be cast in syntactic terms, but in information-structural terms, such that nulls tend to indicate topic continuity, and overts indicate topic change (vallduvı́ (1992) for catalan, samek-lodovici (1996) and dieugenio (1998) for italian). both theories make the same predictions in most cases, given that the referent in the subject position is usually also the topic, but would differ in the cases in which the topic is realized by another grammatical function. in some previous work (mayol, 2010b), i showed that the syntactic preferences of nulls are not a byproduct of their informational structure preferences, but that the two levels interact and the pah needs to be refined in order to capture the pronouns’ preferences. nulls have a simple preference 1. as an anonymous reviewer points out, the support for the pah in alonso-ovalle et al. (2002) is only partial: while nulls displayed a strong subject preference, overts did not have a strong bias to either antecedent. 3 mayol for subject antecedents, regardless of whether they are topics or not. overts have a more complex preference for low-salience (non-subject, non-topics) antecedents.2 another important difference that has been pointed out in the literature between the two types of pronouns in romance languages is that overts often carry a contrastive flavor, absent in nulls. this has lead some authors to treat overts as a counterpart to stressed pronouns in languages like english, and to account for their use as ways to encode focus (luján, 1985, 1999) or contrastive topics (mayol, 2010a). having briefly examined the role of grammatical factors in pronominal references, let us now turn to the less-studied effect of semantico-pragmatic factors. 2.2 effect of semantico-pragmatic on the interpretation of pronouns in the last decade it has become clear that grammatical factors alone cannot explain pronoun interpretation. take for instance, example (5) (which extends example (1) discussed above). (5-a) and (5-b), as mentioned, have been used to argue for the subject-assignment strategy, since the use of the active or passive voice alters the referent assignment. however, why is this strategy not active in (5-c)? in order to account for cases like this, it has been proposed that pronouns refer to an antecedent with the same grammatical function (smyth, 1994): a subject pronoun would be biased to a subject antecedent, and an object pronoun to an object antecedent. unfortunately, this hypothesis cannot account for (5-d), in which a subject pronoun is interpreted as referring to the previous object. (5) a. bush narrowly defeated kerry, and special interests promptly began lobbying him. [him=bush] b. kerry was narrowly defeated by bush, and special interests promptly began lobbying him. [him=kerry] c. bush narrowly defeated kerry, and romney absolutely trounced him.[him=kerry] d. bush narrowly defeated kerry and he quickly demanded a recount. [him=kerry] kehler (2002) notices that the rhetorical relations involved in (5) are not the same and argues that rhetorical relations affect pronoun interpretation. examples (5-a) and (5-b) are examples of occasions. the rhetorical relation of occasion allows a speaker to signal a narration event, in which a set of events are temporally ordered: the initial state of the second sentence is equated with the final state of the first one. thus, in an occasion relation, the most salient referents at the end state of the first utterance become likely antecedents for upcoming pronouns. the most salient referent will be the topic of the sentence (the referent about which the sentence is about), which typically corresponds to the subject. hence, many pronouns appearing in occasion relations are subjectbiased. that is, the discourse in (5-a) is about what bush did and so the pronoun refers to him. similarly (5-b) is about what kerry did, so the pronoun refers to him. in contrast, the discourses in (5-c) and (5-d) are not occasions. (5-c) is a case of parallel, in which commonalities and differences between entities are highlighted: both sentences express a similar eventuality in which either bush or romney defeat kerry and, thus, the pronoun is naturally assigned to the previous object. finally, (5-d) is a result and, therefore, a causal link is established between the two sentences, so that the second sentence expresses a consequence of the first. our world-knowledge pushes the antecedent 2. since the two types of categories (syntactic and informational) largely overlap, this work will focus on the study of syntactic functions. 4 asymmetries in catalan pronouns assignment to the object, given that it is the one who loses an election the one who typically demands a recount.3 the conclusion from the previous discussion is that rhetorical relations greatly affect pronoun interpretation and come with their own biases. moreover, rhetorical relations often interact with other semantic factors, such as verb type. this was partially illustrated in example (2) with implicit causality verbs: some verbs carry their own biases in cases of explanation (when the event in the second utterance explains why the event of the first utterance occurred). another well-studied semantic class of verbs is the transfer of possession verbs (tpv, henceforth), such as ‘give’, which express the event of an object being transferred from a source argument to a goal argument (stevenson et al., 1994; arnold, 2001; kehler et al., 2008; rosa and arnold, 2017). both the source and the goal can function as subjects or objects depending on the verb, as illustrated in (6). previous research (stevenson et al., 1994; arnold, 2001) used the discourse-completion paradigm and showed that goal referents are more accessible, particularly when they are in subject position (6-a), while the bias is milder when the goal is not the subject (6-b). (6) a. bob received a book from peter. he ... b. peter handed a book to bob. he ... kehler et al. (2008) replicated this finding: in items in which the goal was not the subject (as in (7-b)), the pronoun did not show a clear interpretation bias (half of the time participants used the pronoun to refer to peter; the other half to refer to bob). however, once the data was grouped by rhetorical relations very clear patterns emerged: occasions and results were clearly biased to the non-subject, while elaborations (a relation in which both sentences describe the same event) and explanations were clearly biased towards the subject. why should this be so? as mentioned, pronouns occurring in an occasion relation are biased towards the most salient referent at the end state of the event in the first utterance, because in an occasion the events are temporally ordered: the end state of the first event is followed by another event, so whatever referent is salient at the end of the first event is likely to be picked up again in subsequent events. while this referent is usually the subject, kehler et. al. (2008), propose that in a context with a tpv, the goal argument is the most salient one, since it is the one in focus at the end of the event. that is, if the speaker is narrating a series of events and she mentions that peter handed a book to bob, a very likely way to continue the discourse would be to explain what bob did with the book: by the end of the first sentence, the most salient referent is bob. the strong non-subject bias of the result relation can be explained in similar terms: if the speakers wants to talk about the consequences of peter giving a book to bob, she will probably continue talking about bob. in contrast, other relations do not focus on the end state of the event in the first utterance: for instance, elaboration and explanations. these relations will be biased towards the subject. for instance, if the speakers want to elaborate on how or why the handing took place, she will probably continue talking about peter and not about bob. thus, we can conclude that rhetorical relations are biased towards a particular discourse referent: in the case of a tpv context, elaborations and explanations are biased towards the subject, and occasions and results are biased towards the non-subject. another semantico-pragmatic factor that has been shown to influence pronoun interpretation is the aforementioned verbs’ implicit causality. implicit causality verbs (icvs) have been used 3. an anonymous reviewer points out that, although the pronoun can be coerced to refer to the object, s/he finds (5-d) fairly unnatural. see section 2.3 for more discussion on the interaction between structural and pragmatic factors. 5 mayol recurrently to study the role of semantic factors in the interpretation of pronouns (caramazza et al., 1977; brown and fish, 1983; fukumura and van gompel, 2010; rohde and kehler, 2014). icvs attribute the cause of the event they denote either to the subject (icv1s, henceforth) or to the object (icv2s, henceforth)4, and, therefore, impose different biases on the pronouns that follow them. for instance, the verb ‘surprise’ is an icv1 and, therefore, (7-a) will more likely be interpreted as conveying that it was anne, the subject, who aced the exam and that this this surprised mary. in contrast, ‘congratulate’ is an icv2, and (7-b) will more likely be interpreted as conveying that it was mary, the object, who aced the exam and that is why anne congratulated her. (7) a. anne surprised mary because she aced the math exam. b. anne congratulated mary because she aced the math exam. 2.3 relationship between interpretation and production there is a current debate in the literature about the relationship between interpretation and production, and particularly about whether interpretation and production are affected by the same set of factors. in particular, the discussion revolves around whether the semantico-pragmatic factors that affect pronoun interpretation (see previous section) also shape production. on the one hand, arnold (1998, 2001) has argued that both interpretation and production are affected by semantico-pragmatic factors. this idea is captured by her expectancy hypothesis, according to which more accessible entities are more likely to be mentioned again in discourse, which in turn increases their probability of being pronominalized. this hypothesis has been tested with transfer of possession verbs (arnold, 2001; rosa and arnold, 2017): in these studies, participants used pronouns more often to refer to the goal than to refer to the source. the effect was particularly strong for non-subjects; in contrast, it was either weak or non-existent for subjects (in arnold (2001) and in experiments 2 and 3 in rosa and arnold (2017)). on the other hand, other studies have found that semantic-pragmatic factors did not affect the choice of anaphoric form. fukumura and van gompel (2010) used implicit causality verbs and found that the rate of pronominalization towards either of the arguments was constant regardless of whether the verb was biased to the subject or the object. a similar finding is reported in rohde and kehler (2014). rohde and kehler (rohde and kehler, 2014; kehler and rohde, 2013) have argued for a model in which pronoun production and interpretation are not mirror images of each other, but are sensitive to different factors. this could explain an interesting asymmetry found in the discourse-completion study by stevenson et al. (1994). as mentioned, when participants were forced to continue with a pronoun, as in (8-a), it was not clearly biased towards either of the antecedents. in contrast, when they could choose which referring expression to use to continue the discourse, as in (8-b), participants mostly used a pronoun to refer to the subject and a name to refer to the object. (8) a. peter handed a book to bob. he ... b. peter handed a book to bob. if speakers mostly use a pronoun to refer to the subject, why doesn’t the pronoun display a stronger subject preference? the answer in rohde and kehler (2014) is that pronoun production and inter4. fukumura and van gompel (2010) use the terminology ‘stimulus-experiencer (se)’ and ‘experiencer-stimulus (es)’. i will use the more neutral icv1 and icv2, given that some of the arguments of the verbs i will use are not either stimuli or experiencers, but rather agents and themes, like for instance the arguments of the verb congratulate in (7b). 6 asymmetries in catalan pronouns pretation are not mirror images of each other, but are sensitive to different factors. in particular, while pronoun interpretation is affected by both structural and semantico-pragmatic factors, pronoun production is not affected by the latter. this explains the asymmetry in the two conditions of stevenson et al. (1994): although speakers mostly use pronouns to refer to the previous subject, hearers can easily assign a pronoun to the non-subject if the semantico-pragmatic factors (i.e. the context, rhetorical relation or the verb’s implicit causality, for instance) push in that direction. rohde and kehler (2014) show that, although this asymmetry between production and interpretation may not be intuitive, it is, in fact, expected if production and interpretation are related probabilistically, through bayes’ rule, illustrated in (9). (9) p(referent | pronoun) = p(pronoun | referent)p(referent) p(pronoun) the term in the left-hand side, p(referent | pronoun), is the probability to refer to a particular referent given that a pronoun has been used. it, thus, represents the point of view of the interpreter: he has heard a pronoun and needs to assign it a referent. the term p(pronoun | referent), in contrast, represents the point of view of the speaker: given that she wants to refer to a particular referent, should she use a pronoun? the crucial point is that these two probabilities are not mirror images of each other, given that a third probability plays a role: p(referent). this is the probability that a particular referent will be mentioned regardless of the form.5 the production bias, p(pronoun | referent), is basically affected only by grammatical factors: i.e. there is a production bias towards pronominalizing the previous subject referent. in contrast, the interpretation bias is affected both by the grammatical factors that affect the production bias and the factors that affect whether a particular referent will be mentioned next (i.e. p(referent)). the latter factors are predicted not to affect production. rohde and kehler (2014) showed how this model can explain the behavior of pronouns in english in a variety of contexts and that it is superior to other models: (i) the one they call expectancy model (after the expectancy hypothesis in arnold (1999)), which assumes that the interpretation bias of a pronoun towards a referent (that is, p(referent | pronoun)) is the probability that the referent is mentioned again (p(referent)) and (ii) the one they call mirror model, which assumes that the bias of a pronoun towards a referent is the probability that such referent will be pronominalized (p(pronoun | referent)). one of the goals of this paper is to examine whether this model can accurately account for the behavior of nulls and overts in catalan. to this end, two discourse-completion studies were carried out. in both experiments, the context sentence contains a verb which manipulates the context in a particular way: experiment 1 uses transfer of possession verbs and experiment 2 implicit causality verbs. the reasons for this choice were two-fold: first, each verb type is expected to interact with rhetorical relations in a particular way and, second, since they have been studied extensively, we will be able to compare our results with previous research. 5. the denominator, p (pronoun), is the probability that a pronoun is used and can be computed by summing the terms in the numerator by all possible referents (subject and object). this term contributes a constant factor which normalizes the values so that they are probabilities and, therefore, sum to 1. it will not be discussed further in the rest of the paper. 7 mayol 3. experiment 1: transfer of possession verbs experiment 1 is a discourse-completion study in which the context sentence used tpvs, which, as mentioned, usually trigger many continuations about the goal argument. in particular, only verbs that locate the goal in indirect object6 position are used so that the typical biases are reversed and we obtain a context biased towards the object. this atypical context is useful to study the interaction between grammatical and semantico-pragmatic factors. 3.1 methods materials this experiment replicates in catalan the aforementioned experiments in stevenson et al. (1994) and kehler et al. (2008), including a condition for the overt pronoun. thus, the experiment has three conditions, as can be seen in (10): (10) a. condition 1: null prompt. el pere li va passar un llibre al robert. ∅ ... peter passed a book to robert. ∅ ... b. condition 2: overt prompt. el pere li va passar un llibre al robert. ell ... peter passed a book to robert. he ... c. condition 3: free prompt. el pere li va passar un llibre al robert. ... peter passed a book to robert. ... the first two conditions will provide interpretation data for nulls and overts respectively: participants will need to interpret the pronouns and provide a completion coherent with their interpretation. the last condition (‘free condition’, henceforth) will provide production data, since participants are free to choose whatever form they deem appropriate for the subject position of the completion. the experiment included 18 critical items, all containing a tpv. in each sentence, two referents of the same gender were mentioned; half of the items contained two male proper names and the other half two female proper names. the source of the transfer event always appeared in subject position, the goal appeared as the indirect object, and the transferred object as the direct object. three lists were built, so that each participant only saw each item in one of the conditions. each list also contained 18 fillers, with non-tpvs and with prompts containing a connector or a temporal expression. the full list of critical items with their english translations can be seen in appendix 1. based on the previous literature, we formulate the following hypotheses: 1. nulls are expected to receive more subject interpretations than overts. 2. type of prompt is expected to affect rhetorical relations. if indeed nulls trigger more subject interpretations, we expect more subject-biased rhetorical relations (explanation and elaboration) in the null condition than in the overt condition. 3. pronoun interpretation is expected to be affected by rhetorical relations. we expect occasions and results to be highly object-biased and elaboration and explanation to be subject-biased. 6. in this section, whenever i use the label ‘object’ i am referring to the indirect object. 8 asymmetries in catalan pronouns 4. pronoun production is not expected to be affected by rhetorical relations. in the free condition, the rate of nulls for a particular referent is expected to be consistent across different rhetorical relations. procedure the data was collected in an online survey site (https:app.surveygizmo.eu). first, the participants read the instructions in which the procedure was explained. then, each item was presented once at a time and the participants were asked to write a single, complete sentence as a completion. they were told to write the first completion that came to mind, avoiding humor. one of the difficulties of working with a null prompt is that participants needed to understand what a null is. in order to achieve this, the instructions contained a brief informal explanation. the instructions explained the difference between a null subject and an overt subject (expressed by means of a pronoun or another phrase) and contained examples of each type of subject. there were several examples in which the null referred both to the previous subject or to the previous object, in order not to bias the participants. participants ninety participants took part in the experiment. they were all native speakers of catalan and students at the universitat pompeu fabra. they were entered in a raffle to win a gift certificate. 3.2 results a total of 1620 completions (18 items * 90 participants) were collected. two judges, the author of the paper and a linguistics graduate student at upf, coded the antecedent of the subject of the completion into one of the following categories: subject (if it referred to the source), object (if it referred to the goal), joint (plural reference to both the subject and the object), other (if it referred to the transferred object or to something else) or unsure (if the pronoun could be understood as referring to more than one referent). the results reported in this paper concern a subset of the collected 1620 completions, specifically those in which both judges agreed that the subject of the completion unambiguously referred to either the subject or the object. we exclude from the analysis those completions in which the judges disagreed about the coding, or the completions coded as join, other and unsure. also discarded are the completions in which a null was not used in the null condition (1.3% (n=23) of the data). i take this low percentage as evidence that participants understood what a null is. in total, 1098 completions were analyzed. these completions were further coded according to (i) the rhetorical relation between the two sentences and (ii) what type of referring expression was used in subject position in the free condition: null, overt, proper name, etc. in this case, any disagreement was individually discussed and the judges agreed on a decision. (11) summarizes the typology of rhetorical relations assumed in the paper (adapted from kehler and rohde (2013)): (11) a. explanation: infer that the second sentence describes a cause for the eventuality described in the first sentence. “john handed a book to bob. he no longer had a use for it.” b. elaboration: infer that both sentences provide descriptions of the same eventuality. “john handed a book to bob. he did so slowly and carefully.” 9 mayol c. occasion: infer a change of state from the second sentence, taking its initial state to be the final state of the eventuality described in the first sentence. “john handed a book to bob. he began reading it.” d. result: infer that the first sentence describes a cause or reason for the eventuality described in the second sentence. “john handed a book to bob. he thanked him for the gift.” e. violated expectation: infer that the second sentence describes an unexpected result of the eventuality described in the first sentence. “john handed a book to bob. he showed no interest in reading it.” f. parallel: infer that the first and second sentences express similar eventualities, as if each provides a partial answer to a common question. “john handed a book to bob. he gave a magazine to him as well.” to test for the statistical significance of the results, mixed-effect logistic regressions were performed, using r (r core team, 2013) and lme4 (bates et al., 2015). all models contained items and participants as random effects. unless otherwise noted, the models also contained random slopes for all predictors and interactions (barr et al., 2013). in case of non-converging models, the random effects structure was simplified, as specified for each individual model. likelihood ratio tests are used to compare mixed-effects models different only in the presence or absence of the fixed effect in question. in the models with prompt, which is a 3-level predictor, as a fixed effect, free is treated as the baseline. in addition, we use pairwise comparisons and report the coefficient estimates (β estimates) and p-value for each combination. figure 1: subject and object continuations by prompt figure 1 shows the percentage of subject and object reference in the three prompt conditions. in the three conditions we find an object bias, which is milder in the null condition (62%), and greater in the free (76%) and, particularly, overt conditions (90%). to test for a main effect of prompt, a likelihood-ratio test was conducted between mixed-effects models different only in the presence or absence of a fixed main effect of prompt (see more details about the non-reduced model in appendix 2, table 7). in both models, pronoun reference (subjects vs. object) was the dependent variable. the likelihood-ratio test showed a main effect of prompt (χ2 = 105.99, p< .001). pairwise comparisons shows that the differences are significant in all three combinations: free-null (β = -1.04, p< .001), 10 asymmetries in catalan pronouns free-overt (β = 1.20, p= .006) and null-overt (β = 2.25, p< .001). thus, hypothesis 1 is borne out: nulls are more subject-biased than overts, although overall both are object-biased in this context. figure 2 breaks down the data of the free condition, showing the subject and object bias of those completions in which participants chose to use either a null or an overt. we can observe that the results in figures 1 and 2 are comparable: while overts display a strong bias towards the object (87%), nulls show a much milder bias (57%). we compared two models in which pronoun reference (subject vs. object) was the dependent variable and which differed only in the presence or absence of a fixed main effect of pronoun type. a likelihood-ratio test showed the effect of pronoun type is significant (χ2 = 23.90, p< .001). figure 2: subject and object continuations by pronoun type in the free condition figure 3: distribution of referring expressions in the free condition let us now examine what kind of referring expressions participants chose in the free condition to refer to the previous subject and object. the data is summarized in figure 3.7 for the subject, the preferred form is the null (72%) followed by a proper name (22%). for the object, the reverse order is found (30% of nulls and 56% of proper names). the overt pronoun is only the third most used form for both cases, accounting for only between 7% and 14% of the data. to test for 7. we exclude from the analysis the cases in which other forms, such as noun phrases or demonstratives were used, which accounted for less than 5% of the data. 11 mayol the statistical significance of the data, we grouped the data in two binary categories: (i) choice of null vs. not-null as a referring expression and (ii) choice of name vs. not-name as a referring expression. we compared two models in which referring expression (null vs. not-null) was the dependent variable and which differed only in the presence or absence of a fixed main effect of pronoun reference (subject vs. object) (see more details about the non-reduced model in appendix 2, table 8). a likelihood-ratio test showed the effect of pronoun reference is significant (χ2 = 63.44, p< .001). we further compared two models in which referring expression (name vs. not-name) was the dependent variable and which differed only in the presence or absence of a fixed main effect of pronoun reference (subject vs. object) (see more details about the non-reduced model in appendix 2, table 9). a likelihood-ratio test showed that the effect of pronoun reference is significant (χ2 = 63.63, p< .001). the asymmetry between interpretation and production observed by english has been replicated in the catalan data for nulls. although nulls displayed a mild interpretation bias towards the object, they display a very strong production bias towards the subject. when participants could choose a form to refer to the subject, they overwhelmingly chose a null pronoun. however, when they had to interpret a pronoun, the pragmatic biases came into play and overwhelmed the production bias. given that it was more likely that the object (and not the subject) was mentioned next, the pronoun was interpreted with an object bias. in the next section, i discuss how this data fits in the bayesian model proposed by rohde and kehler (2014). the data has also uncovered another asymmetry concerning overts. although interpreters have a very strong bias to interpret the overt pronoun as referring to the object, it is clearly not the preferred form to refer to the object. more discussion of this asymmetry is postponed to section 3.4. let us now examine the effect of rhetorical relations. first, table 1 shows the distribution of rhetorical relations in each condition. we can take the results of the free condition as the neutral results for tpv contexts: that is, as an estimate of which kind of rhetorical relations we are mostly likely to encounter after a tpv context. the data shows that, in a tpv context, we can expect a fair share of occasions and explanations, followed by results, elaborations and violated expectations.8 now, observe how forcing participants to use either a null or an overt causes a switch in the percentage of observed rhetorical relations. null prompts cause a rise of the rhetorical relations with subject bias, elaboration and explanation, and a decrease of the rhetorical relations with object bias, occasion and result. the opposite result is found for overts. free null overt explanation 25 33 17 elaboration 18 32 9 occasion 27 15 46 result 15 10 20 violated expectation 12 9 6 parallelism 3 1 2 total 100 100 100 table 1: % rhetorical relations by prompt. 8. since parallelism occurred very rarely, it will not be discussed further. 12 asymmetries in catalan pronouns figure 4: distribution of rhetorical relations by prompt in order to see the pattern more clearly, rhetorical relations can be grouped in two categories: subject-biased relations, which include elaboration and explanation, and object-biased relations, which include result and occasion. figure 4 shows the distribution of these two categories by prompt. while in the free condition, subject and object-biased relations are evenly split, nulls clearly favor subject-biased relations and overts clearly favor object-biased relations. in order to test for statistical significance, two models were compared in which rhetorical relation bias (subject-biased vs. object-biased, as described above) was the dependent variable and which differed only in the presence or absence of a fixed main effect of prompt type (null, overt, free) (see more details about the non-reduced model in appendix 2, table 10). a likelihood-ratio test showed the effect of prompt type is significant (χ2 = 142.92, p< .001). pairwise comparisons shows that the differences are significant in all three combinations: free-null (β = 0.95, p< .001), free-overt (β = -1.30, p< .001) and null-overt (β = -2.25, p< .001). thus, hypothesis 2 is also borne out: type of prompt affects the distribution of rhetorical relations. figure 5: subject continuations by rhetorical relation bias and pronoun type let us now turn to the question of whether pronoun interpretation is affected by rhetorical relations. figure 5 shows the proportion of continuation about the subject by rhetorical relation bias in conditions 1 and 2 (with nulls and overts). it can again be seen how rhetorical relations greatly affect the percentage of subject bias and the split between subject and object-biased relationships. 13 mayol object-biased relations show almost no reference to the subject, regardless of whether we find a null or an overt. in contrast, subject-biased relations show a greater proportion of continuations about the subject. two models were compared in which pronoun reference was the dependent variable and which differed only in the presence or absence of a fixed main effect of rhetorical relation bias (subjectbiased vs. object-biased) (see more details about the non-reduced model in appendix 2, table 11). a likelihood-ratio test showed the effect of rhetorical relation is significant (χ2 = 350.12, p< .001). hypothesis 3 is borne out: pronoun interpretation is affected by rhetorical relations. while the results of object-biased relations is very homogenous, it is worth taking a closer look at the data for subject-biased relations. table 2 summarizes the proportion of continuation about the subject in explanation and elaboration by pronoun type. while elaborations shows a high percentage of subject references, the behavior of explanation is somewhat unexpected. although the proportion of subject continuations is higher than for object-biased relations, it is not clearly subject-biased either. this unexpected result will be discussed in section 3.4. null overt explanation 29 32 elaboration 72 50 table 2: proportion of continuations about the subject in explanations and elaborations by pronoun type the results so far have confirmed that pragmatic factors affect pronoun interpretation. that is, pronouns biases are different across different rhetorical relations. for instance, nulls are mostly interpreted as referring to the subject in cases of elaborations, while they are interpreted as referring to the object in cases of occasions. now, we should take a look at the production data. remember that the hypothesis is that the pronouns biases for a particular referent should be similar across different rhetorical relations. figure 6 shows the percentage of nulls referring to the subject in three rhetorical relations.9 we compared two models in which referring expression (null pronoun vs. no pronoun) was the dependent variable and which differed only in the presence or absence of a fixed main effect of rhetorical relation (elaboration being the baseline; see more details about the non-reduced model in appendix 2, table 12). a likelihood-ratio test showed the effect of rhetorical relation is not significant (χ2 = 0.25, p= .88). thus, hypothesis 4 is borne out: the production of nulls remains constant and it is not affected by whether the rhetorical relation is subject biased or not. 3.3 applying the bayesian model let us now examine the predictions of the bayesian model, presented in section 2.3. in particular, we will examine the production and interpretation of the null pronoun to refer to a subject antecedent.10 we can, thus, rewrite the equation in (9) as (12). 9. we only present the cases in which there were at least 10 data points per rhetorical relation, which was not the case for result, violated expectation and parallelism. 10. an analysis for the overt pronoun will not be presented. there are not enough data points for a reliable estimation, given that it was scarcely produced in the free conditions 14 asymmetries in catalan pronouns figure 6: proportion of nulls referring to the subject by rhetorical relation (12) p (subject | null) = p (null | subject)p (subject) p (null) all these parameters can be estimated using data from the discourse-completion study. first, p (subject | null) is the probability that, given that a null pronoun has been used, it will be interpreted as referring to the subject. this probability, which can be estimated using the data of the null condition, is 0.38 (see figure 1). second, the reverse probability, p (null | subject), is the probability that a null pronoun will be used given that the referent is the subject. this probability, which can be estimated using data from the free condition, is 0.72 (see figure 3). third, the overall probability that the subject is the referent (regardless of the form used) is 0.24 (estimated using data from the free condition, see figure 1). finally, the denominator is the probability that a null pronoun is used, which can be calculated by summing the terms in the numerator by all possible referents (in this case subject and object). all the parameters are summarized in (13). the predicted probability is 0.44, not far from the observed probability (0.38). in order to test whether the correlation between observed and predicted probabilities is statistically significant we carried out a linear regression test over item means (that is, we computed the mean observed and predicted probabilities for each item).11 the correlation is significant (adjusted r2= 0.40, p= .007). (13) a. observed p (subject | null) = 0.38 b. p (null | subject) = 0.72 c. p (subject)= 0.24 d. p (null) = p (null | subject) ∗ p (subject) + p (null | object) ∗ p (object) = 0.40 e. predicted p (subject | null)= 0.43 the bayesian account outperforms other models, such as the mirror model, which equatesp (subject | null) with the probability that a null is used given that the referent is the subject (p (subject | null), or the expectancy model, which equates p (subject | null) with the probability that the subject is named again (p (subject)). in the case of the mirror model the correlation is not significant (ad11. we excluded items for which in the free condition participants did not refer to one of the two referents. that was the case for 1 of the 18 items. 15 mayol justedr2= -0.06, p= .71); in the case of the expectancy model the correlation is significant (adjusted r2= 0.37, p= .009), but slightly lower than the one obtained with the bayesian model. 3.4 discussion experiment 1 mostly confirms the model by kehler and rohde (2014) about the role of rhetorical relations. different pronoun prompts raise or lower the probability of particular rhetorical relations, which in turn carry their own interpretation biases. for instance, occasion and result are strongly object-biased and elaboration is subject-biased. in contrast, pronoun production is not affected by such pragmatic factors: the rate of pronominalization is similar in occasion and elaboration, although they display opposite interpretation biases. one of the results that does not fit precisely with their account is the low percentage of subject references with a null pronoun in the case of explanation, which was only 29%. if explanation is subject-biased and nulls are subject-biased, why do 71% of nulls in explanation refer to objects? i believe the answer is that two different types of explanations are in play for subject and object references. (14) shows some examples of subject-referring explanations, while (15) shows some examples of object-referring explanations. (14) a. l’elena the elena li dat va regalar gave un a llibre book a to la the mercè. mercè. ∅ ∅ sap knows que that li dat encanten please aquesta this mena kind de of regals. gifts. ‘elena gave a book to mercè. (she) knows she adores this kind of gift.’ b. el the joan joan va donar gave una a joguina toy a to l’enric. the enric. ∅ ∅ va preferir preferred cedir to give in abans before que that barallar-se fight un one altre other cop. time ‘joan gave a toy to enric. (he) preferred to give in rather than getting into a fight again. (15) a. el the pere pere li dat va passar passed la the clau key anglesa english al to the marc. marc. ∅ ∅ la it necessitava needed per to obrir open un a calaix drawer enorme. huge. ‘pere passed the wrench to marc. (he) needed it to open a huge drawer.’ b. l’alba the alba li dat va deixar lent el the cotxe car a to la the maria. maria. ∅ ∅ havia had d’anar to go als to the pirineus pyrenees a to veure see la the seva her famı́lia. family. ‘alba lent her car to maria. (she) had to go to the pyrenees to see her family.’ the object-referring completions answer a question like ‘why did the object referent need the transferred possession?’ and typically include transferred possessions which are used to achieve something: a wrench is used to do something; a car is used to go somewhere. in a more fine16 asymmetries in catalan pronouns grained typology of rhetorical relations, such examples could be coded as conveying purpose, rather than explanation.12 in contrast, the subject-referring completions usually involved sentences with transferred possessions such as ‘book’ and ‘toy’. in this case, since it is not very informative to write about why the object referent needs a book or a toy, the completions rather explained why the subject transferred the possession to the object. in fact, according to kehler et al. (2008), the biases in the relations that express cause or effect (explanation and result, respectively) “will depend on the semantics incorporated in the passage and the referent to which causality or consequentiality is most likely to be attributed in a particular context” (page 26). thus, explanations are not intrinsically subject-biased: in some cases it is more likely to form a causal link between the two sentences with the subject antecedent, and in other cases with the object antecedent (see bott and solstad (2014) for a typology of explanations which attempts to make this idea more precise). the data has also uncovered a strong asymmetry between the production and interpretation of the overt pronoun. upon hearing an overt pronoun, a hearer will readily interpret it as referring to the object, even though the probability of using an overt pronoun to refer to an object is very low. this data shows that that the division of labor between nulls and overts is only partial: present in interpretation, but not in production. we, thus, see again that production and interpretation are not mirror images of each other. the fact that overts come only in third place in terms of the referring expressions chosen, after proper names and nulls, suggests that, although they do display a clear object bias, their role in discourse is not to indicate such a bias, but rather to signal other relevant properties. contrast seems the obvious candidate. as mentioned in section 2, overts often convey a contrastive flavor and are compulsory when they encode focus or contrastive topics. in order to examine the function of overts more closely, we can examine their occurrence in the free conditions; that is, those completions in which participants freely choose to use a pronoun. there were 38 such occurrences, of which 87% had object reference and 13% subject reference. in those cases, we do indeed find some examples in which overts are conveying a contrastive topic, see (16), or a focus, see (17), in which the overt is in the postverbal contrastive focus position. (16) l’elena the elena li dat va regalar gave un a llibre book a to la the mercè. mercè. ella she en in canvi change li dat va regalar gave una a rosa. rose. “elena gave a book to mercè. she gave him a rose instead.” (17) el the sergi sergi li dat va facilitar supplied totes all les the eines tools a to l’àlex. the àlex. aixı́ thus no not havia had de to buscar-les search them ell. he. “sergi supplied all the tools to àlex. thus, he was not the one that had to look for them”. these cases, albeit possible, are not by any means representative of most of the data: that is, it is not the case that pronouns were mostly used to convey contrastive topic or focus. there is, however, a wider notion of contrast that may be useful to understand the data. table 3 shows the distribution of rhetorical conditions among the occurrences of overts in the free condition. if we compare this distribution to the one of the whole free condition (see table 1), we can see that the most striking difference is the higher percentage of violated expectations. in a violated expectation there is some kind of contrast between what is expected to happen and what really 12. i thank an anonymous reviewer for this observation. 17 mayol explanation 29 occasion 29 violated expectation 21 result 13 parallelism 8 table 3: distribution of rhetorical conditions for overts in the free condition happened, as illustrated in (18) with some of the completions of the experiment. so, although these cases are not cases of foci or contrastive topics, they do convey contrast at a discourse level. the use of overts is, thus, favored by discourse contrastivity, at least in this context, although this is not a necessary condition for their appearance. the production of overts will be further discussed in connection to experiment 2, which is presented in the next section. (18) a. la the gemma gemma li dat va subministrar supplied tot all el the material material necessary necessary a to l’elisena. the elisenda. tot all i and això, this, ella she no not li dat va thanked. agrair. “gemma supplied all the necessary material to elisenda. however, she did not thank her. b. la the sra. ms. molins molins li dat va cedir bestowed la the seva her col·lecció collection de of segells stamps a to la the lluı̈sa. lluı̈sa. però but ella she no not sabia know què what fer-ne. do part. “ms. molins bestowed her stamp collection on lluı̈sa. but she did not know what to do with it. 4. experiment 2: implicit causality verbs experiment 2 is also a discourse-completion study, which uses implicit causality verbs (icvs). as mentioned, icvs are useful to study the role of semantico-pragmatic factors, since they attribute the cause of the event they denote either to the subject (icv1s) or to the object (icv2s), which has been shown to affect how the pronouns that follow these verbs are interpreted. 4.1 methods materials experiment 2 is similar to experiment 1, but it contains two factors: (i) verb type (icv1 and icv2) and (ii) prompt type (null, overt, free prompt (free condition)). there were, thus, 6 conditions: (19) illustrates the three conditions with an icv1 and (20) the three conditions with an icv2. (19) a. condition 1: icv1 + null la núria va sorprendre la maria. ∅ ... ‘núria surprised maria. ∅ ...’ 18 asymmetries in catalan pronouns b. condition 2: icv1 + overt la núria va sorprendre la maria. ella ... ‘núria surprised maria. she ...’ c. condition 3: icv1 + free la núria va sorprendre la maria. ... ‘núria surprised maria. ...’ (20) a. condition 4: icv2 + null la núria va felicitar la maria. ∅ ... ‘núria congratulated maria. ∅ ...’ b. condition 5: icv2 + overt la núria va felicitar la maria. ella ... ‘núria congratulated maria. she ...’ c. condition 6: icv2 + free la núria va felicitar la maria. ... ‘núria congratulated maria. ...’ the experiment contained 30 items: 15 with an icv1 and 15 with an icv2. in each of the items, two referents of the same gender were mentioned, half containing masculine names and the other half feminine names. the first referent always appeared as subject and the second one as either direct object, indirect object or predicative complement. from now on, i will be referring the non-subject argument as ‘object’. three lists were constructed, so that each participant only saw each item with one of the prompts. the lists also contained 20 fillers with non-icv verbs and with a connector or a temporal expression as a prompt. the full list of critical items with their english translations can be seen in appendix 1. based on the previous literature, we formulate the following hypotheses: 1. nulls are expected to receive more subject interpretations than overts. 2. given their lexical semantics, icvs are expect to trigger a high number of explanation completions. 3. pronoun interpretation is expected to be affected by pragmatic factors. for explanations, we expect a contrast between icv1s and icv2s: icv1s are predicted to be subject-biased and icv2s to be object-biased. we do not expect other rhetorical relations to be sensitive to the contrast between icv1 and icv2. as we found in experiment 1, elaboration is expected to be subject-biased, and occasion and result are expected to be object-biased, regardless of verb type. 4. pronoun production is not expected to be affected by rhetorical relations. in the free condition, the rate of pronominalization for a particular referent is expected to be similar in icv1s and icv2s. procedure the same procedure described in the previous experiment was followed. participants seventy-eight participants took part in the experiment. none of the participants had participated 19 mayol in experiment 1. they were all native speakers of catalan and students at the universitat pompeu fabra and were entered in a raffle to win a gift certificate. 4.2 results a total of 2340 completions (30 items * 78 participants) were collected. the data was coded following the same procedure explained for experiment 1. the antecedent of the referent first mentioned in the completion was coded into one of the following categories: subject, object, joint (plural reference to both the subject and the object), other or unsure (if the pronoun could be understood as referring to more than one referent). the results are based on a subset of the completions, in which both judges agreed that the subject unambiguously refereed to the previous subject or the previous object. the same exclusions reported for the previous experiment were carried out: all those completions in which the judges did not agree or the subject was coded as unsure, joint or other. also discarded were the cases in which a null pronoun was not used in the conditions icv1+null and icv2+null, which amounted to 0.68% (n=16) of the data. in total, 1934 completions were analyzed. these completions were coded according to the rhetorical relation between the two sentences and type of referring expression (null, overt, proper name, etc.) used in the two free conditions. as in experiment 1, any disagreement between the judges was individually discussed and the judges agreed on a decision. the statistical analysis followed the same methods previously described for experiment 1. let us start by discussing hypothesis 1, which predicts that nulls should be more subject-biased than overts. figure 7 shows that this is indeed the case both for icv1s and icv2s. we compared two models in which pronoun reference (subject vs. object) was the dependent variable and which differed only in the presence or absence of a fixed main effect of pronoun type (null vs. overt). a likelihood-ratio test showed that the effect of pronoun type is significant (χ2 = 175.84, p< .001). if we take the data in the free conditions, a similar pattern emerges, as can be seen in figure 8. again, in a likelihood-radio test comparing two models which differ only in the presence or absence of a fixed main effect of pronoun type, the effect of pronoun type is significant (χ2 = 21.78, p< .001). figure 7: proportion of continuations about the subject by verb type (icv1 vs. icv2) and prompt type (null vs. overt) the two figures above have also uncovered another interesting pattern: there are more subject references with icv1s than with icv2s, both with overt and nulls. as a result, nulls show a clear bias towards the subject with icv1s, and overts a strong bias towards the object with icv2s. in the other two combinations (null + vc2, and overt + vc1), the two biases conflict with each 20 asymmetries in catalan pronouns figure 8: proportion of continuations about the subject in the free condition by verb type and form type other and, as a result, there is no clear tendency. a model with verb type and pronoun type as fixed effects was compared to a model with an interaction between pronoun type and verb type (to achieve convergence, both models only contained items and participants as random effects; see more details about the non-reduced model in appendix 2, table 13). a likelihood-ratio test showed that the interaction between pronoun type and verb type is significant (χ2 = 7.98, p< .005). we can understand these results as a by-product of both hypotheses 2 and 3: we are expecting icv contexts to trigger many explanations, and explanation is the rhetorical relation which imposes different biases (icv1s towards the subject and icv2s towards the object). thus the prevalence of explanations is what is responsible for the pattern observed in the tables above. recall that the prediction is that the difference in biases in icv1 and icv2 contexts should only be present in explanation and, therefore, should disappear with other relations. let us start by seeing whether icvs really triggered a high number of explanation continuations (hypothesis 2). table 4 shows the distribution of rhetorical relations: explanation does indeed dominate, followed by elaboration and result. hypothesis 2 is indeed borne out. explanation 61 elaboration 19 result 14 violated expectation 3 occasion 2 parallelism 1 table 4: distribution of rhetorical relations we can now examine whether pronoun interpretation is affected by pragmatic factors. figures 9 and 10 show the subject bias of nulls and overts respectively by verb type and rhetorical relation.13 two models were compared in which pronoun reference was the dependent variable. one of the models had rhetorical relation, verb type and pronoun type as fixed effects, while the other had an interaction between the three fixed effects (elaboration being the baseline; to achieve convergence, both models only contained items, participants and pronoun type as random effects; see more details about the model with the interaction in appendix 2, table 14). a likelihood-ratio test showed that the 13. we eliminate from the analysis the rhetorical relations which account for less than 5% of the data. 21 mayol figure 9: proportion of continuations about the subject with nulls (by rhetorical relation and verb type) figure 10: proportion of continuations about the subject with overts (by rhetorical relation and verb type) interaction is significant (χ2 = 175.54, p< .001). having seen that there is a significant interaction, let us examine the data in specific rhetorical relations. for each rhetorical relation, two models were compared in which pronoun reference was the dependent variable and which differed only in the presence or absence of a fixed main effect of verb type. this comparison was done both for the data in the null and in the overt condition. as expected, in both cases, likelihood-ratio test showed the effect of verb type was significant for explanations: icv1s are subject biased and icv2s are not (in the null condition, χ2 = 30.80, p< .001; in the overt condition, χ2 = 42.21, p< .001). elaboration is subject-biased in both types of contexts, both with nulls and overts: the effect of verb type is not significant for overts (χ2 = 2.24, p = .52), but it is significant for nulls (χ2 = 25.63, p< .001). result yields an unexpected significant contrast between icv1 and icv2 for the null pronoun condition, (χ2 = 14.46, p< .001) while no contrast arises with the overt pronoun (χ2 = 1.88, p= 0.59). i will discuss this unexpected contrast in section 4.4. with the exception of the behavior of result with nulls, hypothesis 3 is also borne out: pronoun interpretation is affected by pragmatic factors. finally, let us turn to production data by examining what referring expression participants used to refer to subject and object depending on whether the context sentence contained an icv1 or an 22 asymmetries in catalan pronouns icv2.14 figures 11 and 12 summarize the data. the main observation is that, in both graphs, the distribution of referring expressions is fairly similar. in figure 11 we can observe, that to refer to the subject, participants overwhelmingly used a null pronoun regardless of whether the context was icv1 or icv2. in figure 12 we can see that, to refer to the object, the preferred form was also the null pronoun, but the percentage of proper names increased as well. whether the verb was icv1 or icv2 does not seem to make a difference. again we find that in both cases overts were scarcely produced. figure 11: distribution of referring expressions to refer to the subject figure 12: distribution of referring expressions to refer to the object we compared two models in which referring expression (null vs. not-null) was the dependent variable. one of the models had verb type and reference as fixed effects, while the other had an interaction between the two fixed effects (to achieve convergence, both models only contained items, participants and reference as random effects; see more details about the non-reduced model in appendix 2, table 15). a likelihood-ratio test showed the effect of the interaction is not significant (χ2 = 0.25, p = .61). in fact, as it can be seen in 15, only reference is significant, while neither verb type nor the interaction between verb type and reference are significant. thus, hypothesis 4 is borne out: production is not affected by rhetorical relation, which greatly affects interpretation. 14. as in experiment 1, cases in which a noun phrase or a demonstrative was used are excluded from the analysis. 23 mayol 4.3 applying the bayesian model let us examine again how well the bayesian model can predict the observed interpretation biases of the null pronoun to refer to a subject antecedent. the interpretation biases observed in the data and the ones predicted by the bayesian models (computed as explained in 3.3 for experiment 1) are summarized in (21). it can be seen how the bayesian model adequately captures the tendencies in the data. a linear model analysis was performed over item means, and the correlation between the expected and the observed probabilities is significant (adjustedr2= 0.57, p< .001). the bayesian account outperforms the the mirror model, which equates p (subject | null) with the probability that a null is used given that the referent is the subject (p (subject | null), or the expectancy model, which equates p (subject | null) with the probability that the subject is named again (p (subject)). in the case of the mirror model the correlation is not significant (adjusted r2= -0.02, p= .65); in the case of the expectancy model the correlation is significant (adjusted r2= 0.44, p < .001), but lower than the one obtained with the bayesian model. (21) a. observed p (subject | null) = 0.63 b. p (null | subject) = 0.89 c. p (subject)= 0.50 d. p (null) = p (null | subject) ∗ p (subject) + p (null | object) ∗ p (object) = 0.70 e. predicted p (subject | null)= 0.64 4.4 discussion experiment 2 clearly showed how the interpretation of nulls and overts is influenced both by grammatical and pragmatic factors. null subject bias increases in explanations with an icv1 verb, while it decreases with an icv2 verb. the opposite it true for overts: they have a strong object bias in explanations with an icv2 verb, which decreases in an icv1 context. in contrast, pronoun production is not affected by pragmatic factors. the rate of use of nulls in explanations remains constant regardless of whether the verb was icv1 or icv2. a surprising result of experiment 2 was the bias shown by nulls in result relations. we expected results to not be sensitive to verb type and, considering what was found with transfer of possession verbs, to display an object bias. instead, we found an object bias with icv1s, and a subject bias with icv2s. this behavior is actually not so surprising if we take into account that many verbs also display biases attributing the consequences of the event to one of their arguments. this bias is usually called ‘implicit consequentiality’15 (stewart et al., 1998; crinean and garnham, 2006; pickering and majid, 2007). many of the icvs used the experiments are actually psychological verbs (47%). in those cases, the subject position of an icv1-sentence is occupied by the stimulus and the object by the experiencer, as in ‘intimidate’ or ’terrify’ (see the first sentences in (22)). this pattern is reversed in icv2s, such as ‘hate’ or ‘fear’ (see the first sentences in (23)). when a result relation was expressed, it most often conveyed the consequences for the experiencer: this amounts to object references for icv1s and subject references for icv2s. the second sentences in (22) and (23) show typical completions of both cases. thus, our hypothesis that, in results, we would find an object bias (like we did for tpvs) was not borne out: instead, what we find is that results display an experiencer bias (see crinean and garnham (2006) for more discussion on the relationship between 15. i thank an anonymous reviewer for pointing out this concept to me. 24 asymmetries in catalan pronouns implicit causality, implicit consequentiality and thematic roles). table 5 shows the subject bias in results by the thematic role of the subject: if the subject was an experiencer, the pronoun had a categorical subject preference, while if it was a stimulus, it had a strong object preference. in the cases of non-psychological verbs, we find the expected object preference. (22) icv1 a. el the pol pol intimida intimidates l’àlex. the àlex. ∅ ∅ sempre always que that el him veu sees ∅ ∅ marxa leaves ràpid. quickly. “pol intimidates àlex. whenever (he) sees him, he leaves quickly.” b. la the pilar pilar té has la the blanca blanca atemorida. terrified. ∅ ∅ sempre always arriba arrives plorant crying a at casa. home. “pilar terrifies blanca. (she) always comes home crying.” (23) icv2 a. en the guillem guillem té has por fear del o the nicolau. nicolau. per for això, this, ∅ ∅ no not va goes a to l’escola. the school. “guillem is afraid of nicolau. this is why, (he) does not go to school.” b. la the candela candela té has enveja envy de of la the júlia. júlia. per for això, this. ∅ ∅ vol wants deixar-la leave her en in ridı́cul. embarrassment. “candela is envious of júlia. this is why, (she) wants to ridicule her. c. el the julià julià odia hates el the nil. nil. ∅ ∅ m’ha dat has conessat confessed que that un one dia day ∅ ∅ el him matarà. kill. “julià hates nil. (he) has confessed to me that one day (he) will kill him.” agent 35 experiencer 100 stimulus 6 table 5: % of subject references in results by thematic role of the subject experiment 2 also confirmed the low probability of using an overt to refer to either antecedent. again, although presumably speakers could use an overt to signal they do not want to refer to the most prominent referent, they usually do not do that and use a proper name instead. in order to understand why an overt is used, we can examine those completions in which an overt was chosen in the free conditions. there were 49 such cases: they mostly refer to the object (80%), consistent with what occurred in the overt conditions, and mostly occur with icv2s (63%), which is what we would expect considering their object bias. the rhetorical relations found in those cases are shown in table 6. the main finding is that there is an increase in the percentage of results and a decrease in the percentage of elaborations. although we do find a few examples of violated expectation (see the examples in (24)), they do not account for a significant amount of the data, unlike what we found in experiment 1. the reason is probably that the items in experiment 1 favored completions in which participants narrated the expected (occasions) or unexpected (violated expectations) events that followed the transfer of possession. in contrast, the items in experiment 2 favor mainly explanations about the event of the first sentence, or results for the experiencer/theme. in all the examples in which the overt was 25 mayol used to express a result, it referred to the object, which was either the experiencer or the theme (see examples in (25)). explanation 58 elaboration 10 result 24 violated expectation 6 parallelism 2 table 6: distribution of rhetorical relations free + overt (24) a. l’andreu the andreu va demanar gave disculpes apologies al to the joan. joan. tot everything i and això, this, ell he no not va acceptar-les. accepted them “andreu apologized to joan. however, he did not accept them.” b. l’iris the iris confia trusts en in la the sı́lvia. sı́lvia. i and ella she la her va trair. betrayed. “iris trusts in sı́lvia. and she betrayed her.” (25) a. l’esteve the esteve va espantar scared el the roger. roger. ell he el him va empenyer pushed com as a a venjança. revenge. “esteve scared roger. he pushed him in revenge.” b. l’adam the adam va elogiar praised en the mateu. mateu. ell he es refl va posar turned vermell. red. “adam praised mateu. he blushed.” 5. conclusion this paper has uncovered two main asymmetries concerning pronouns in catalan: (i) the interpretation of nulls is affected by several pragmatic factors, which do not influence production, and (ii) while overt pronouns show a strong interpretation bias towards the object, they are not frequently produced with this goal. thus, the data supports only partially the idea that there is a division of labor between null and overt pronouns: while nulls are clearly the default pronouns in catalan, the role of the overts is severely restricted. overall the data supports the model put forward by rohde and kehler (2014), according to which interpretation and production are not affected by the same set of factors. in their words, although it might be natural to expect “that speakers will employ pronouns in just those contextual circumstances in which the intended referent will be favored by the comprehender’s own biases” (rohde and kehler, 2014, p. 924), this is not what our data shows. the results are compatible with a model in which production is affected by structural factors (i.e. nulls are subject biased and overts are non-subject biased) and insensitive to semantico-pragmatic factors. according to fukumura and van gompel (2010), this difference between structural and semanticopragmatic factors arises because only the former contributes to an entity’s accessibility. this makes 26 asymmetries in catalan pronouns sense if we consider that while speakers can manipulate the structure of a sentence (for instance, choosing the passive form instead of the active form) depending on the relative accessibility of the entities in their discourse model, no such choice occurs with semantico-pragmatic factors. the speaker will use an icv1 or an icv2 depending on the meaning she wants to communicate; she will not choose one verb or another depending on the accessibility of the entities. the same reasoning can be applied to rhetorical relations: a speaker will link two utterances in her discourse by an occasion or by an elaboration depending on what is relevant for the discourse. it is unlikely that she will choose to elaborate on a previous utterance just to create a bias towards a certain entity. now, by uttering an elaboration a bias is certainly created and the interpreter can use this cue, but the reason to utter an elaboration (as opposed to, say, an occasion) is not to create the bias. more generally, the picture that seems to emerge is one in which the interpreter combines multiple cues to assign reference (in particular, the grammatical bias associated with the choice of referential form with the prior probability of who will be mentioned), while the speaker ignores some of the cues that could potentially shape production. the result that production ignores factors which do affect interpretation challenges the audience design hypothesis (clark, 1996), the idea that speakers always plan their utterance with the hearer in mind. it is instead compatible with a view that speakers for the most part use their own model to plan their utterances, while ignoring the hearer’s (see fukumura and van gompel (2012) for experimental evidence showing that speakers use their own discourse model, and not the hearer’s, when producing referential expressions). a possible, although at this stage speculative, explanation of why speakers should ignore some cues which are useful for hearers would be that production is a more costly process than interpretation. a speaker has the burden to plan and produce the utterance, which includes selecting the appropriate lexical items within a huge lexicon, giving them a grammatical structure and articulating the relevant sounds. thus, language production is undoubtedly hard (see, for instance, macdonald (2013) for a model in which linguistic form is derived from the attempts by the speakers to mitigate utterance planning difficulties). in contrast, the task of the hearer is relatively simpler and that could be why the burden of integrating pragmatic information is passed exclusively to him. this would also be compatible with the finding that interpreters accommodate the needs of the speaker (from perspective-taking (duran et al., 2011), to phonetic adaptations (kraljic et al., 2008) or lexical and syntactic ambiguity resolution (macdonald et al., 1994)). apart from providing support to the bayesian approach with data from a language other than english16, this study also contributes to the characterization of overts in romance null-subject languages. although we replicated the non-subject interpretation bias often discussed in the literature, a very striking result of our experiments is that overts are rarely produced to fulfill this goal. we explored the possibility that overts, like stressed pronouns in english, are only used in strictly contrastive uses, but this idea was not supported by the data. in fact, although, as mentioned, overts were not used often, in some of the cases in which they were chosen, their function seemed to be precisely to refer to low-salience referents (which is fully compatible with their interpretation bias). but this raises the question: if overts can be used to refer to low-salience referents, why don’t speakers do it more often? i do not have a full answer to this question, but the difference in distribution of rhetorical relations seems to suggest they play an important role in licensing overts. for future work, corpus-based research is planned so that the behavior of overt pronouns can be further studied and the predictions of the bayesian hypothesis can be tested with naturally-occurring 16. similar studies were conducted in japanese (ueno and kehler, 2016), but the predictions of the bayesian approach were not tested. 27 mayol data. a second venue for future work includes the use of richer contexts in the experiments. while our results support those who argue that predictability does not affect form production (fukumura and van gompel, 2010; rohde and kehler, 2014), recent research has made the opposite point. in particular, in rosa and arnold (2017), the effects of predictability on form production were seen more clearly in tasks with very rich contexts, in which participants were asked to describe a series of pictures which told a coherent story. thus, further work is needed to clarify the relationship between predictability and form production. finally, we are also planning to further examine tpvs, comparing those in which the source argument is in subject position with those in which it is not, to corroborate that semantico-pragmatic factors do not play a role in determining referential form. we expect that we should not find a difference in pronominalization rate depending on whether the subject is the source or the goal, in the same way the we found similar pronominalization rates for rhetorical relations with opposed biases. acknowledgements i am grateful to the three anonymous reviewers and to the associate editor for their constructive and detailed comments, suggestions and criticisms, which greatly contributed to improve this paper, as well as to the audiences of the venues where previous versions of this research was presented: glif, amore, universitat de les illes balears, university of vienna and xprag 2017. i would also like to thank aina obis for her work in annotating the data, andrew kehler for many suggestions during the early stages of this project and hannah rohde for help in generating the graphs. the research underlying this article has been partially supported by projects ffi2015-66732-p and ffi2015-67991-p, funded by the ministry of economy, industry and competitiveness and the european regional development fund (feder, ue). appendix 1 experiment 1 el joan li va portar un got d’aigua al robert. (‘joan brought a glass of water to robert.’) el pere li va passar la clau anglesa al marc. (‘pere passed a wrench to marc.’) en roger li va lliurar el treball al professor ramos. (‘roger submitted his work to professor ramos.’) en pep li va enviar el seu cv al toni. (‘pep sent his cv to toni.’) la marta li va donar una samarreta a la ruth. (‘marta gave a t-shirt to ruth.’) l’elena li va regalar un llibre a la mercè. (‘elena gifted a book to mercè.’) la jèssica li va servir l’arròs a la carme. (‘jèssica served the rice to carme.’) l’eva li va tornar la grapadora a l’esther. (‘eva returned the stapler to esther.’) el martı́ li va proporcionar medicines al miquel. (‘martı́ provided medicines to miquel.’) la rosa li va vendre una postal a la dolors. (‘rosa sold a postcard to dolors.’) el jordi li va dur el sopar a l’ernest. (‘jordi brought dinner to ernest.’) la sra. molins li ha cedit la seva col·lecció de segells a la lluı̈sa. (‘mrs. molins brought her stamp collection to lluı̈sa.’) l’adrià li va entregar un sobre a l’albert. (‘adrià delivered an envelope to albert.’) l’alba li ha deixat el cotxe a la marina. (‘alba lent her car to marina.’) la teresa li va acostar la sal a la núria. (‘teresa moved the salt closer to núria.’) 28 asymmetries in catalan pronouns el sergi li va facilitar totes les eines a l’àlex. (‘sergi provided all the tools to àlex.’) la gemma martı́nez li va subministrar el material necessari a l’elisenda. (‘gemma martı́nez supplied the necessary material to elisenda.’) el joan li va donar una joguina a l’enric. (‘joan gave a toy to enric.’) experiment 2 icv1: l’andreu va demanar disculpes al joan. (‘andreu apologized to joan.’) la mar va ofendre l’irene. (‘mar offended irene.’) la raquel va enganyar l’aurora. (‘raquel deceived aurora.’) l’àngels va humiliar la marga. (‘àngels humiliated marga.’) l’abel va fer enfadar el cesc. (‘abel annoyed cesc.’) la sara fa riure molt la laia. (‘sara amuses laia.’) el miquel treu el gerard de polleguera. (‘miquel bothers gerard.’) l’amanda va deixar la montse bocabadada. (‘amanda amazed montse.’) l’esteve fa posar nerviós el david. (‘esteve makes david nervous.’) el quim va decebre el pep. (‘quim disappointed pep.’) l’esteve va espantar el roger. (‘esteve scared roger.’) en lluı́s va sorprendre el vı́ctor. (‘lluı́s surprised vı́ctor.’) el jordi té el pau absolutament captivat. (‘jordi captivated pau.’) en pol intimida l’àlex. (‘pol intimidates àlex.’) la pilar té la blanca atemorida. (’pilar frightens blanca.’) icv2: la marina va consolar la susanna. (‘marina comforted susanna.’) la isabel va felicitar la júlia. (‘isabel congratulated júlia.’) la noemı́ va ajudar la clara. (‘noemı́ helped clara.’) el dı́dac es va burlar del xavier. (‘dı́dac mocked xavier.’) l’alı́cia va calmar la xènia. (‘alı́cia calmed xènia.’) l’àdam va elogiar en mateu. (‘àdam praised mateu.’) la cati va esbroncar la conxita. (‘cati told conxita off.’) en tomàs va renyar el josep antoni. (‘tomàs scolded josep antoni.’) l’aleix va donar les gràcies al felip. (‘aleix thanked felip.’) l’isaac va corregir el joel. (‘issac corrected joel.’) en guillem té por del nicolau. (‘guillem fears nicolau.’) la candela té enveja de la joana. (‘candela is jealous of joana.’) el julià odia el nil. (‘julià hates nil.’) l’iris confia en la sı́lvia. (‘iris trusts sı́lvia.’) la laura valora la xènia. (‘laura values xènia.’) 29 mayol appendix 2 variable estimate error z-value p-value null 1.04 0.27 3.85 0.0001 overt -1.20 0.39 -3.05 0.002 table 7: experiment 1 results. pronoun reference (subjects, object) by prompt (null, overt, free) model: reference ∼ prompt + (prompt|subject)+(prompt|item) variable estimate error z-value p-value subject -22.85 9.79 -2.33 < .01 table 8: experiment 1 results. form of referring expression (null, not-null) by reference (subject, object) model: nullbinary ∼ reference + (reference|subject)+(reference|item) variable estimate error z-value p-value subject 30.07 6.87 4.37 < .001 table 9: experiment 1 results. form of referring expression (name, not-name) by reference (subject, object) model: namebinary ∼ reference + (reference|subject)+(reference|item) variable estimate error z-value p-value null -0.95 0.24 -3.92 < .0001 overt 1.30 0.25 5.18 < .0001 table 10: experiment 1 results. rhetorical relations (subject-biased, not subject-biased) by prompt (free, null, overt) model: rhetoricalbinary ∼ prompt + (prompt|subject) + (prompt|item) variable estimate error z-value p-value object-biased -4.54 0.38 -11.886 < .0001 table 11: experiment 1 results. reference (subject, object) by rhetorical relation (subject-biased, not subject-biased) model: reference ∼ rhetoricalbinary + (rhetoricalbinary|subject)+(rhetoricalbinary|item) 30 asymmetries in catalan pronouns variable estimate error z-value p-value explanation 0.97 2.82 0.34 0.73 occasion 0.31 4.64 0.06 0.94 table 12: experiment 1 results. form of referring expression (null, non-null) by rhetorical relation (elaboration, explanation, occasion) model: nullbinary ∼ rhetorical + (rhetorical|subject)+(rhetorical|item) variable estimate error z-value p-value overt -1.20 0.18 -6.6 < .0001 icv2 -1.20 0.26 -4.46 < .0001 overt*icv2 -0.75 0.26 -2.82 0.004 table 13: experiment 2 results. reference (subject, object) by pronoun (null, overt) * verb type (icv1, icv2) model: reference ∼ pronoun*type + (1|subject)+(1|item) variable estimate error z-value p-value icv2 1.49 0.63 2.33 0.01 explanation 1.19 0.42 2.81 0.004 result -3.02 0.73 -4.08 < .0001 overt -1.10 0.45 -2.44 0.01 icv2*explanation -4.55 0.66 -6.83 < .0001 icv2*result 1.12 0.99 1.13 0.25 icv2*overt -0.84 0.75 -1.12 0.26 explanation*overt 0.25 0.55 0.45 0.64 result*overt 0.03 0.93 0.03 0.97 icv2*explanation*overt 0.11 0.92 0.12 0.90 icv2*result*overt -2.30 1.36 -1.68 0.09 table 14: experiment 2 results. reference (subject, object) by pronoun (null, overt) * verb type (icv1, icv2) * rhetorical relation (elaboration, explanation, result) model: reference ∼ pronoun*type*rhetorical + (pronoun|subject)+(pronoun|item) variable estimate error z-value p-value icv2 -0.66 0.65 -1.02 0.30 subject -3.09 0.78 -3.92 < .0001 icv2*subject 0.52 1.04 0.50 0.61 table 15: experiment 2 results. form of referring expression (null, non-null) by reference (subject, object) * verb type (icv1, icv2) model: formbinary ∼ reference*type + (reference| subject) + (reference| item) 31 mayol references alonso-ovalle, l., solera, s. f., frazier, l. and clifton, c. (2002), ‘null vs. overt pronouns and the topic-focus articulation in spanish’, journal of italian linguistics 14:2, 151–169. ariel, m. (1990), accessing noun-phrase antecedents, london: routledge. arnold, j. e. (1999), reference form and discourse patterns., phd thesis, stanford university. arnold, j. e. (2001), ‘the effect of thematic roles on pronoun use and frequency of reference continuation’, discourse processes 31(2), 137–162. barr, d. j., levy, r., scheepers, c. and tily, h. j. (2013), ‘random effects structure for confirmatory hypothesis testing: keep it maximal’, journal of memory and language 68(3), 255–278. bates, d., mächler, m., bolker, b. and walker, s. (2015), ‘fitting linear mixed-effects models using lme4’, journal of statistical software 67(1), 1–48. bott, o. and solstad, t. (2014), from verbs to discourse: a novel account of implicit causality, in c. f.-h. b. hemforth, b. mertins, ed., ‘psycholinguistic approaches to meaning and understanding across languages’, springer, pp. 213–251. brown, r. and fish, d. (1983), ‘the psychological causality implicit in language’, cognition 14(3), 237–273. cameron, r. (1992), pronominal and null subject variation in spanish: constraints, dialects, and functional compensation, phd thesis, university of pennsylvania. caramazza, a., grober, e., garvey, c. and yates, j. (1977), ‘comprehension of anaphoric pronouns’, journal of verbal learning and verbal behavior 16, 601–609. carminati, m. n. (2002), the processing of italian subject pronouns, phd thesis, university of massachusetts. clark, h. h. (1996), using language, cambridge university press: cambridge. crawley, r. a., stevenson, r. j. and kleinman, d. (1990), ‘the use of heuristic strategies in the interpretation of pronouns’, journal of psycholinguistic research 19(4), 245–264. crinean, m. and garnham, a. (2006), ‘implicit causality, implicit consequentiality and semantic roles’, language and cognitive processes 21(5), 636–648. dieugenio, b. (1998), centering in italian, in a. k. j. m. walker and e. prince, eds, ‘centering theory in discourse’, oxford university press, pp. 114–137. duran, n. d., dale, r. and kreuz, r. j. (2011), ‘listeners invest in an assumed other’s perspective despite cognitive cost’, cognition 121(1), 22–40. fukumura, k. and van gompel, r. p. (2010), ‘choosing anaphoric expressions: do people take into account likelihood of reference?’, journal of memory and language 62(1), 52–66. 32 asymmetries in catalan pronouns fukumura, k. and van gompel, r. p. (2012), ‘producing pronouns and definite noun phrases: do speakers use the addressee’s discourse model?’, cognitive science 36(7), 1289–1311. givón, t. (1983), topic continuity in discourse, john benjamins publishing company. gundel, j. k., hedberg, n. and zacharski, r. (1993), ‘cognitive status and the form of referring expressions in discourse’, language pp. 274–307. keating, g. d., vanpatten, b. and jegerski, j. (2016), ‘online processing of subject pronouns in monolingual and heritage bilingual speakers of mexican spanish’, bilingualism: language and cognition 19(01), 36–49. kehler, a. (2002), coherence, reference and the theory of grammar, stanford, ca, csli publiacions. kehler, a., kertz, l., rohde, h. and elman, j. l. (2008), ‘coherence and coreference revisited’, journal of semantics 25(1), 1–44. kehler, a. and rohde, h. (2013), ‘a probabilistic reconciliation of coherence-driven and centeringdriven theories of pronoun interpretation’, theoretical linguistics 39(1-2), 1–37. kraljic, t., samuel, a. g. and brennan, s. e. (2008), ‘first impressions and last resorts: how listeners adjust to speaker variability’, psychological science 19(4), 332–338. luján, m. (1985), binding properties of overt pronouns in null pronominal languages, in p. k. w. h. eilforth and k. peterson, eds, ‘proceedings of the chicago linguistics society’, vol. 21, pp. 424–438. luján, m. (1999), expresión y omisión del pronombre personal, in v. demonte and i. bosque, eds, ‘gramática descriptiva de la lengua española’, espala-calpe, madrid, pp. 1275–1316. macdonald, m. c. (2013), ‘how language production shapes language form and comprehension’, frontiers in psychology 4, 226. macdonald, m. c., pearlmutter, n. j. and seidenberg, m. s. (1994), ‘the lexical nature of syntactic ambiguity resolution.’, psychological review 101(4), 676. mayol, l. (2010a), ‘contrastive pronouns in null-subject romance languages’, lingua 120(10), 2497–2514. mayol, l. (2010b), ‘refining salience and the position of antecedent hypothesis: a study of catalan pronouns’, university of pennsylvania working papers in linguistics 16(1), 15. mayol, l. and clark, r. (2010), ‘pronouns in catalan: games of partial information and the use of linguistic resources’, journal of pragmatics 42(3), 781–799. pickering, m. j. and majid, a. (2007), ‘what are implicit causality and consequentiality?’, language and cognitive processes 22(5), 780–788. r core team (2013), r: a language and environment for statistical computing, r foundation for statistical computing, vienna, austria. url: http://www.r-project.org/ 33 mayol rigau, g. (1986), some remarks on the nature of strong pronouns in null-subject languages, in i. bordelois, h. contreras and k. zagona, eds, ‘generative studies in spanish syntax’, foris, dordrecht. rohde, h. and kehler, a. (2014), ‘grammatical and information-structural influences on pronoun production’, language, cognition and neuroscience 29(8), 912–927. rosa, e. c. and arnold, j. e. (2017), ‘predictability affects production: thematic roles can affect reference form selection’, journal of memory and language 94, 43–60. samek-lodovici, v. (1996), constraints on subjects: an optimality theoretic analysis, ph.d. thesis, rutgers university. silva-corvalán, c. (1977), a discourse study of word order in the spanish spoken by mexicanamericans in west los angeles, master’s thesis, university of california, los angeles. smyth, r. (1994), ‘grammatical determinants of ambiguous pronoun resolution’, journal of psycholinguistic research 23(3), 197–229. stevenson, r. j., crawley, r. a. and kleinman, d. (1994), ‘thematic roles, focus and the representation of events’, language and cognitive processes 9(4), 519–548. stewart, a. j., pickering, m. j. and sanford, a. j. (1998), implicit consequentiality, in ‘proceedings of the twentieth annual conference of the cognitive science society’, lawrence erlbaum associates, nj, pp. 1031–1036. ueno, m. and kehler, a. (2016), ‘grammatical and pragmatic factors in the interpretation of japanese null and overt pronouns’, linguistics 54(6), 1165–1221. vallduvı́, e. (1992), the informational component, garland, new york. 34 dialogue and discourse 3(2) (2012) 43–74 doi: 0.5087/dad.2012.203 question generation for french: collating parsers and paraphrasing questions delphine bernhard dbernhard@unistra.fr lilpa, université de strasbourg, france louis de viron louis@earlytracks.com earlytracks sa, belgium véronique moriceau veronique.moriceau@limsi.fr limsi-cnrs, univ. paris-sud, orsay, france xavier tannier xavier.tannier@limsi.fr limsi-cnrs, univ. paris-sud, orsay, france editors: paul piwek and kristy elizabeth boyer abstract this article describes a question generation system for french. the transformation of declarative sentences into questions relies on two different syntactic parsers and named entity recognition tools. this makes it possible to further diversify the questions generated and to possibly alleviate the problems inherent to the analysis tools. the system also generates reformulations for the questions based on variations in the question words, inducing answers with different granularities, and nominalisations of action verbs. we evaluate the questions generated for sentences extracted from two different corpora: a corpus of newspaper articles used for the clef question answering evaluation campaign and a corpus of simplified online encyclopedia articles. the evaluation shows that the system is able to generate a majority of good and medium quality questions. we also present an original evaluation of the question generation system using the question analysis module of a question answering system. keywords: question generation, syntactic analysis, syntactic transformation, paraphrasing, question answering 1. introduction question generation (qg) has been addressed recently from different perspectives and for different application domains: dialogue systems, intelligent tutoring, or automatic assessment. most recent methods perform text to text generation, i.e. they transform declarative sentences into their interrogative counterpart. it thus constitutes the inverse operation to question answering (qa) which aims at retrieving answers for a given question based on a collection of text documents. qg is actually a complex task, which relies on a large variety of resources and natural language processing tools: named entity recognition, syntactic analysis, anaphora resolution and text simplification. however, while several methods have been proposed for the english language, there is no equivalent system for the french language. the system presented in this article aims at closing this gap. the main contributions of the article are as follows: c©2012 delphine bernhard, louis de viron, véronique moriceau and xavier tannier submitted 3/11; accepted 1/12; published online 3/12 bernhard, de viron, moriceau and tannier • we present a question generation system for french which adapts methods proposed in the context of question generation for the english language • we evaluate how the use of different syntactic parsers and named entity recognition tools for analysing the source sentence affects the quantity and quality of the generated questions • we propose two mechanisms for generating different surface forms of the same question: reformulation of the interrogative part and question nominalisation. • we present a novel evaluation of the question generation system using the question analysis module of a question answering system. in the next section, we detail related work on the topic of question generation from text. in section 3, we detail the typology of the questions generated by our system. we then describe the system in section 4 and explain our method for question paraphrasing in section 5. we evaluate our question generation system and discuss the results in section 6. finally, we conclude and give the perspectives of this work in section 7. 2. state of the art 2.1 question generation from expository text question generation from text consists in automatically transforming a declarative sentence into an interrogative sentence. the question thus generated targets one part of the input sentence, e.g. the subject or an adverbial adjunct. this research domain has been the focus of increasing interest, owing to the recent organisation of an international challenge aimed at comparing question generation systems for english (rus et al., 2010; rus et al., this volume). automatic question generation from text has two main application domains: (i) dialogue and interactive question answering systems and (ii) educational assessment. in the first application context, question generation has been used for automatically producing dialogues from expository texts (prendinger et al., 2007; piwek and stoyanchev, 2010). the dialogues thus generated may be presented in the form of written text or thanks to virtual agents with speech synthesis. concrete uses of such dialogues include presenting medical information, such as patient information leaflets and pharmaceutical notices, or educational applications. in a related domain, question answering, the quality of the interactions between the system and a user can be improved if the qa system is able to predict some of the questions that the user may wish to ask. harabagiu et al. (2005) describe a question generation method for interactive qa which first identifies entities relevant for a specific topic and then applies patterns to obtain questions. automatically generated questions are then presented to the user by selecting those which are most similar to the question asked initially. the second application context of qg systems, question generation for educational assessment, has been investigated for many years (wolfe, 1976). indeed, test writing is a very time-consuming task and the availability of a question generation system thus reduces the workload for instructors. educational assessment applications rely on question generation methods for producing open or multiple-choice questions for text comprehension (wolfe, 1976; gates, 2008) or summative assessment (mitkov et al., 2006). depending on the system, the generated questions may correspond to manually-defined templates (brown et al., 2005; wang et al., 2008), or be less constrained (mitkov et al., 2006; gates, 2008; heilman and smith, 2009). 44 question generation for french automatic question generation systems for the english language usually proceed according to the following steps: (i) perform a morphosyntactic, syntactic, semantic and/or discourse analysis of the source sentence, (ii) identify the target phrase for the question in the source sentence, (iii) replace the target phrase with an adequate question word, (iv) make the subject agree with the verb and invert their positions, (v) post-process the question to generate a grammatical and well-formed question (wolfe, 1976; gates, 2008; heilman and smith, 2009; kalady et al., 2010). given this procedure, question generation requires that the input sentence be at least morpho-syntactically analysed. it is often also useful to have additional information about semantics, such as named entities (person, organisation, country, town) or the distinction between animates and inanimates. this approach has been shown to yield good results for the english language. we therefore apply the same method to the french language. prior to question generation, we use two different syntactic analysers and two different named entity recognition methods to analyse the input text. our goal in doing so is to evaluate the impact of the prior analysis tools on the quality and quantity of generated questions. 2.2 question paraphrasing as a related task to question generation from text we consider question paraphrasing. indeed, the same question may be asked with different surface forms. for instance, the following questions can be considered as paraphrases, since they have the same meaning and expect the same answer, while presenting alternate wordings: how many ounces are there in a pound?, what’s the number of ounces per pound?, how many oz. in a lb.? the ability to identify question paraphrases is useful for question answering on databases of question-answer pairs, e.g. faqs or social q&a sites (tomuro and lytinen, 2004; zhao et al., 2007; bernhard and gurevych, 2008). in this case, the task of question answering is boiled down to the problem of finding question paraphrases in a database of answered questions. in the context of question generation, the potential uses of automatic question paraphrasing are manifold. we will detail three of these applications: (i) vary the form of the question for educational assessment, (ii) induce different phrasings of answers and (iii) improve recall in automatic qa. first, one drawback of automatic question generation from text is that generated questions may be too close to the original sentence, as generation usually relies on transforming the syntactic structure from declarative to interrogative, without changing the words used (apart maybe from modifications in the main verb’s inflection). however, for the automatic generation of educational assessment questions from study material, it is necessary that the generated questions be different from the surface form of the input clause, so as not to give extraneous cues to students and make questions too easy to answer (karamanis et al., 2006). second, different formulations of the same question induce different phrasings of answers, varying in their level of precision (granularity) or elaborateness. this is due to the phenomenon of lexical entrainment which has been observed in human-human dialogues (garrod and anderson, 1987) and human-machine dialogues (gustafson et al., 1997; stoyanchev and stent, 2009; parent and eskenazi, 2010): people tend to adapt to the vocabulary used by their interlocutor, be it a human or a machine. what is more, interrogative words determine the semantic type of the expected answer. they may be rather vague (e.g., “where”) and thus determine a large spectrum of potential answers subsumed by the broad semantic type related to the question word (e.g., “location”), or more precise, in which case they determine a restricted type of answer (e.g., “town”, “road”, “region”, etc.) 45 bernhard, de viron, moriceau and tannier (garcia-fernandez, 2010; garcia-fernandez et al., 2010). as an example, the two questions below have different interrogative words and request answers with differing degrees of precision: adverbial question: où est-ce que se trouve la joconde ? (where is the mona lisa?) • vague answer: en france • mid-range answer: à paris • precise answer: au musée du louvre determinative question: dans quelle ville se trouve la joconde ? (in what city is the mona lisa located?) • answer: à paris, en france a basic question can thus be reformulated to induce different answers. we make use of this observation in our question generation system by varying the interrogative part of the question (see section 5.1). finally, automatic question paraphrasing can be used to reformulate questions given as an input to automatic question answering systems, in order to facilitate answer retrieval and improve recall. tomuro (2003) present paraphrasing patterns for questions, aimed at identifying question reformulations for answer retrieval from faqs. rinaldi et al. (2003) maps user questions to a logical form, which makes it possible to resolve syntactic variations between question and answer. term variations involving synonyms and morpho-syntactic variants are associated with a unique concept identifier. duboue and chu-carroll (2006) describe a more shallow technique based on machine translation to automatically paraphrase questions, which were then fed to their qa system. in our qg system, we perform automatic question paraphrasing by transforming verbal constructions involving action verbs into the corresponding nominal construction (see section 5.2). 3. typology of generated questions numerous studies have been devoted to interrogative sentences, especially in linguistics, cognitive psychology and information science (pomerantz, 2005). as a consequence, several taxonomies of question types have been proposed, depending on the target application or the type of discourse under study. in the field of nlp, research on this topic has been fostered by the development of question answering systems. question types defined within large-scale evaluation campaigns for question answering, namely trec1 or clef,2 focus on factoid questions (person, time, location, etc.), definition questions and list questions (dang et al., 2006; giampiccolo et al., 2007). more complex questions types such as why or how have been more seldomly dealt with, even though they have been studied in recent research (diekema et al., 2003; moriceau et al., 2010). in our system, we focus on factoid questions (what, when, where, who) targeted at concept completion and knowledge elicitation, closed questions, which require a yes/no answer, and some definition questions (lehnert, 1978; graesser and person, 1994). the expected response for our questions is either a boolean (yes/no) or a short phrase from the input sentence. for the time being, 1. http://trec.nist.gov/ 2. http://www.clef-campaign.org/ 46 question generation for french we do not deal with quantity (how much? how many?) or measurement (what size? what length?) questions. questions are generated based on syntactic transformations applied to a source sentence. we have therefore identified which syntactic constituents are best suited as targets for automatically generated questions, based both on our own intuition and linguistic studies (langacker, 1965; grevisse, 1975). these syntactic constituents are the expected answers for the generated questions. next, we have classified these constituents based on their grammatical function. table 1 details the results of this preliminary study. in addition to the question categories described in this typology, the qg system also generates boolean questions which do not target a specific constituent in the source sentence. table 1: typology of question types and related grammatical functions. function example constituent subject s: terminer ses études est l’objectif de john. to complete his studies is john’s goal. noun phrase pronoun q: quel est l’objectif de john ? what is john’s goal? infinitive clause that/what clause direct object s: j’ai vu le père de marie. i saw mary’s father. noun phrase q: qui as-tu vu ? whom did you see? pronoun subordinate clause indirect object s: j’ai offert ce cadeau à mon frère. i offered this present to my brother. noun phrase pronoun q: à qui as-tu offert ce cadeau ? to whom did you offer this present? subject complement s: la peinture est très belle. the painting is beautiful. noun phrase adjective q: comment est la peinture ? what is the painting? locative adverbial adjunct q: je viens du cinéma. i come from the movie theater. noun phrase pronoun q: d’où viens-tu ? where you do come from? temporal adverbial adjunct s: mon avion décolle à 13h. my plane takes off at 1 pm. noun phrase pronoun q: à quelle heure décolle ton avion ? when does your plane take off? appositive s: le président français, nicolas sarkozy, ... the french president, nicolas sarkozy, ... noun phrase q: qui est nicolas sarkozy ? who is nicolas sarkozy? parenthesised acronym s: l’organisation des nations unies (onu) ... the united nations (un) ... noun phrase q: que signifie onu ? what does un mean? 47 bernhard, de viron, moriceau and tannier 4. description of the question generation system in this section, we detail the question generation system. the input is a french sentence. the output consists of all the questions which were generated by the system for this input sentence. 4.1 syntactic analysis and named entity recognition the question generation system proceeds by transforming declarative sentences into questions. the transformations rely on a syntactic analysis of the input sentences. in order to alleviate possible errors stemming from erroneous parses, we use the analyses of two different syntactic parsers: xip (aït-mokhtar et al., 2002) and bonsai (candito et al., 2010). xip (xerox incremental parsing) is a robust parser which also performs named entity recognition and dependency parsing. bonsai relies on the berkeley parser adapted for french and trained on the french treebank. xip only produces a shallow constituent tree, as the other type of information is available in the form of dependency relationships.3 we therefore pre-process the tree in order to incrementally group phrases which belong together and thus obtain complete syntactic groups for question generation. in practice, we grouped elements which, together, realize the same grammatical function in the sentence. figure 1 illustrates the process of grouping constituents in the parse tree produced by xip. figure 1: example rule for grouping constituents. goal : group a noun phrase and noun complement under an np node (ex: le chat de jean dort.) • tregex : np=nom $+ pp=cn [< −1 nom (the pp is moved under the np node) top sent . sc fv dort pp jeande np chatle top sent . sc fv dort np pp jeande chatle head of the np while xip combines syntactic analysis with named entity tagging, bonsai does not perform any named entity recognition. we therefore combined bonsai’s output with the following information: 3. the current version of the system does not make use of information about dependency relations. as suggested by one reviewer, this type of information would however be helpful for qg. 48 question generation for french • identification of numerical expressions: day, month, year, date, period, duration, percentage, length, temperature, financial amount, weight, volume, physical measure. numerical expressions are identified with a tool has been originally developed for a question answering system (ferret et al., 2002) and which relies on regular expressions. • identification of named entities: organisations (association, firm, institution), locations (astronym, building, city, country, geonym, hydronym, region, supranational location, way), people (celebrity, dynasty, ensemble, ethnonym, firstname, pseudoanthroponym, job, family), events (feast, history, manifestation). these named entities are identified using a simple list-based approach. we used two different named entity databases: prolexbase (tran and maurel, 2006; bouchou and maurel, 2008) which is a lexical resource for proper names, and some of the lists constituted for the system described in elkateb-gara (2005). we also used the french wiktionary4 to retrieve lists of terms related to family (mother, father, etc.) as well as lists of jobs. a further list of jobs was extracted from the onisep website.5 the size of each list is detailed in appendix c. we also used the melt pos tagger combined with the lefff lexicon to obtain morpho-syntactic tags (denis and sagot, 2009). the syntactic trees produced by xip and bonsai are processed with tregex and tsurgeon (levy and andrew, 2006).6 tregex makes it possible to explore syntactic trees with a specific regular expression language, relying on node relations in the tree. tsurgeon performs transformations on the syntactic tree, based on specific nodes identified with tregex patterns. these tools have been used previously for qg, in other systems (gates, 2008; heilman and smith, 2009; kalady et al., 2010). 4.2 sentence simplification prior to question generation, the input sentences are simplified. sentence simplification has been shown to be a key asset for several natural language processing applications, including question generation. sentence simplification can occur before question generation, in order to ease the generation process. for instance, heilman and smith (2010b) describe a factual statement extractor system which identifies simplified statements within complex sentences. the system relies on the sentence’s structure in order to identify splitting points and remove the least significant elements.7 simplification can also occur at the end of the generation process, to improve the quality of automatically generated questions (gates, 2008). our simplification process relies on tregex patterns associated with tsurgeon operations. the sentences are first parsed using the bonsai analyser. then, several simplification rules are applied. table 2 details the rules with some simplification examples in french. in order to prevent information loss due to the simplification process, non-simplified sentences are also provided as input to the question generation system. 4. http://fr.wiktionary.org 5. http://www.onisep.fr 6. http://nlp.stanford.edu/software/tregex.shtml. 7. the system is downloadable online: http://www.ark.cs.cmu.edu/mheilman/qg-2010-workshop/ [visited august 23, 2011]. 49 bernhard, de viron, moriceau and tannier table 2: sentence simplification rules. simplification before after break up coordinated propositions le trafic des trains à grande vitesse ice a été provisoirement suspendu entre francfort et paris, et les trains ont été redirigés via strasbourg. le trafic des trains à grande vitesse ice a été provisoirement suspendu entre francfort et paris. les trains ont été redirigés via strasbourg. break up colonseparated propositions je reste chez moi : il pleut. je reste chez moi. il pleut. remove commaseparated sentenceinitial np texte de "compromis", cette nouvelle loi européenne a l’ambition de mettre l’ue "à la pointe dans la protection animale", selon les mots du commissaire européen john dalli. cette nouvelle loi européenne a l’ambition de mettre l’ue "à la pointe dans la protection animale", selon les mots du commissaire européen john dalli. remove commaseparated sentencefinal pp texte de "compromis", cette nouvelle loi européenne a l’ambition de mettre l’ue "à la pointe dans la protection animale", selon les mots du commissaire européen john dalli. texte de" compromis", cette nouvelle loi européenne a l’ambition de mettre l’ue "à la pointe dans la protection animale". remove commaseparated sentenceinitial adverb or pp à ce stade, l’enclos du piton de la fournaise demeure accessible au public. l’enclos du piton de la fournaise demeure accessible au public. remove pp surrounded by commas en france, pour permettre une meilleure fréquentation de ses trains, la sncf a mis en service le tgv en 1981. en france, la sncf a mis en service le tgv en 1981. remove a pp in a sequence of pps la première locomotive à vapeur a été inventée par richard trevithick en angleterre en 1804. la première locomotive à vapeur a été inventée en angleterre en 1804. remove subordinate clauses cette pratique fut énormément utilisée sur la loire, pour la remontée de nantes à orléans, voire plus en amont si les conditions le permettaient. cette pratique fut énormément utilisée sur la loire, pour la remontée de nantes à orléans, voire plus en amont. remove parentheticals en europe, il est cultivé dans la plaine du pô (italie). en europe, il est cultivé dans la plaine du pô. 50 question generation for french 4.3 question generation the generation of questions is decomposed in two steps: 1. identification of the target of the question, based on the question typology. 2. transformation of the source sentence into a question. the first step relies on manually defined tregex expressions, which are detailed in appendix b. it consists in verifying that the potential target constituents are in the source sentence. for instance, in order to generate a question about the subject, the system locates the subject in the source sentence and labels it with a special tag in the syntactic tree. once the target constituents have been identified, tsurgeon transformation rules are applied to the sentence. if no target constituent is found in the source sentence, then no question is generated. figure 2 shows how questions concerning locations are generated. the generation of other question types follows the same kind of procedure. 4.4 challenges for french question generation generating questions in french implies two challenges that are not present in english: • subject-verb inversion • modifications in the verbal form, due to the syntactic transformations. 4.4.1 subject-verb inversion langacker (1965) describes two types of subject-verb inversions: simple and complex inversions. simple inversion the subject and the verb are simply inverted. this inversion is applied if the subject is a pronoun (1) and if the question is about the direct object (2) or the subject complement (3). (1) il mange une pomme (eng: he eats an apple)→ que mange-t-il? (eng: what does he eat?) (2) jean mange une pomme (eng: john eats an apple)→ que mange jean? (eng: what does john eat?) (3) jean est électricien (john is an electrician)→ qu’est jean? (what is john?) for the compound tenses, the subject is placed between the auxiliary and the participle if it is a pronoun (4) and after the entire verbal form if it is not a pronoun (5). (4) nous avons travaillé (eng: we have worked)→avons-nous travaillé (eng: have we worked?) (5) jean a mangé une pomme (eng: john has eaten an apple)→ qu’a mangé jean? (eng: what did john eat?) in order to perform simple inversions, we use regular expressions which shift the verbal form before the subject. if the verb is a compound, the subject is inserted between the auxiliary and the rest of the form. then, hyphens are added (see example 4) and elisions are performed if necessary (see example 5). 51 bernhard, de viron, moriceau and tannier figure 2: generation of questions about locations. top sent . pp angleterre +country +locationen sc fv inventéeétéa np locomotivela top sent . sc fv inventéeétéa np locomotivela question angleterre +country +locationen (a) (b) où (c) a-t-elle été inventée 1. identification of the locative adverbial adjunct: pp=loc[< abandonner vmn---- abandon ncms the nominalisation process is based on tregex patterns which identify specific question structures including action verbs. overall, we defined 13 patterns. once such a structure has been identified, the verb’s lemma is looked up in a nominalisation resource and the question is subsequently transformed into its nominalised version. the main difficulty lies in selecting the correct nominalisation. indeed, for some verbs, there are several nominalisations (bauer, 2003). for instance, given the question “en quelle année a été tourné le film "war games" ?” (eng: in which year was the movie "war games" shot?), there are several alternative deverbal nouns derived from the verb tourner: tournage (eng: filming, shooting), tour (eng: turn) and tournée (eng: tour). only the first of these alternatives is correct. we therefore have to perform ambiguity resolution to select the best noun corresponding to the verb. for this, we use the web, such as proposed by sumita et al. (2005) for validating distractors in multiplechoice cloze items. for each candidate deverbal noun, we form several web queries including words situated in the immediate left and right context of the deverbal noun. we issue these queries to the yahoo search engine and sum the number of hits. the deverbal noun which gets the largest amount of hits is eventually selected. 6. evaluation the evaluation of our question generation system aims at answering the following questions: 1. what is the quality of the automatically generated questions? what amount of post-editing is necessary for producing perfect questions out of them? 2. what is the recall of the system? 3. how well does question nominalisation perform? 55 bernhard, de viron, moriceau and tannier 4. how well does automatic question analysis for question answering perform for automatically generated questions? we first present the evaluation datasets and then seek answers for all of the above questions. 6.1 evaluation datasets for the evaluation of the qg system, we used two different datasets for the source sentences: • 40 sentences from the clef french corpus, containing answers to questions for the 2005 and 2006 challenges. these sentences have been randomly selected among sentences containing answers to definition questions and factoid questions expecting an answer of type person, organisation, date and place. the clef french corpus is composed of about 177,000 wellformed news articles from le monde and ats 1994 and 1995 (about 2 gb). the sentences of these documents are supposed to be well-written and syntactically correct. clef 2005 questions are factoid and definition questions only; clef 2006 introduced a few list questions that we did not use in our evaluation. • 40 sentences from wikimini8 and vikidia,9 which are simplified versions of the french wikipedia targeted at children. we randomly selected the articles and we took the first sentence of each article, which we assume would contain the most important information. sentence simplification prior to question generation, the input sentences were automatically simplified using the simplification method described in section 4.2. the simplification process was iteratively applied to simplified sentences in order to further simplify sentences. this iterative process is stopped when none of the sentences can be simplified anymore. for our corpora, the simplification process was terminated after two iterations (simpl. 1 and simpl. 2). table 3 displays the statistics for the evaluation datasets before and after simplification. the figures indicate the number of sentences produced by each simplification step. for instance, for the clef dataset, 31 simplified sentences are produced out of the 40 source sentences. as expected, the simplification leads to shorter sentences. moreover, sentences in the clef corpus are, on average, longer than those in the vikidia-wikimini dataset. table 3: statistics for the evaluation datasets before and after simplification clef vikidia-wikimini source simpl. 1 simpl. 2 source simpl. 1 simpl. 2 # sentences 40 31 8 40 19 3 # tokens 978 665 167 567 222 20 avg. sentence length 24.45 21.45 20.88 14.18 11.68 6.67 generated questions overall, 869 questions were generated for the clef dataset out of which 646 were unique. for the vikidia-wikimini dataset, we obtained 431 questions, out of which 275 8. http://fr.wikimini.org 9. http://fr.vikidia.org 56 question generation for french were unique.10 figure 5 displays the number of questions generated (a) depending on the syntactic analyser used and (b) depending on the source sentence used (before and after simplification). the amount of questions generated is larger for sentences analysed with bonsai. however, this discrepancy can be partly explained by the fact that the rules for generating acronym questions have been only implemented for bonsai analyses. acronym questions account for 45 questions in the clef dataset. strikingly, the overlap between questions generated based on both syntactic analysers is rather low: the identical questions account for 130 questions in clef and 115 questions in vikidiawikimini. this justifies the use of different syntactic parsers for question generation, as this makes it possible to obtain a larger variety of question types. figure 5: number of questions generated. bonsai only xip only both 0 50 100 150 200 250 300 350 400 450 309 207 130 84 76 115 393 283 245 clef vikidiawikimini sum (a) depending on the syntactic analyser (unique questions) source simpl. 1 simpl. 2 0 100 200 300 400 500 600 700 800 414 356 99 278 131 22 692 487 121 clef vikidiawikimini sum (b) depending on the simplification step figure 6 displays the repartition of question types for both corpora. interestingly, no acronym or appositive questions were generated for the vikidia-wikimini dataset. this is due to the greater simplicity of the vikidia-wikimini sentences which rarely include appositive constructions or parenthesised acronyms. the proportion of locational and temporal questions is also larger for the clef dataset, here again due to the bigger complexity of the source sentences which more often contain prepositional phrases corresponding to adverbial adjuncts. 6.2 intrinsic evaluation of question quality the acceptability of the generated questions was manually evaluated by native speakers. for each corpus, we randomly selected 100 automatically generated questions, including questions generated from simplified sentences. the questions were annotated by two independent annotators each. a 3-point scale was used for the annotation: 1. unworthy questions, which were directly rejected because they were too inconsistent or ungrammatical to be redeemed. 10. duplicates are due to identical questions generated at different simplification steps or based on the analysis of one of our two syntactic analysis tools. 57 bernhard, de viron, moriceau and tannier figure 6: repartition of question types for the evaluation datasets. 45 16 70 40 59 158 3 174 81 clef 59 60 20 28 2 85 21 vikidia-wikimini acronym appositive closed direct object indirect object locational subject complement subject temporal 2. questions which were not perfect but could be improved, by performing some manual changes. 3. perfect questions, which could be accepted without any change. table 4 summarises the evaluation results for both corpora while table 5 lists examples of generated questions with their corresponding score. the annotators agreed in 70% of the cases for the clef corpus and in 67% of the cases for the vikidia-wikimini corpus. the corresponding kappa values (0.537 for clef and 0.484 for vikidia-wikimini) indicate a moderate level of agreement. however, disagreements mostly concern adjacent categories in our annotation and cases where one annotator considers a question as perfect while the other rejects it are rare (9 cases for clef and 4 cases for vikidia-wikmini). the performance of the qg system is the lowest for direct object and indirect object questions. closed questions and acronym questions perform best. this result is unsurprising as these questions are produced by rather simple generation rules. question quality is also globally better for the simpler vikidia-wikimini corpus. if we compare both syntactic parsers, xip obtains better scores for the clef corpus, but also leads to the generation of less questions than bonsai. for the vikidiawikimini corpus, on the contrary, bonsai performs slightly better on average. in order to further evaluate the quality of the generated questions, we asked the human annotators to post-edit the automatically generated questions that were neither perfect, nor rejected, which corresponds to a score of 2. the goal of the manual post-edition was to create targeted references for the questions, based on the questions generated by the system. an example of an automatically generated question and the result after post-editing by both annotators is shown below: question qu’est-ce qui est une ville de suisse et un port fluvial situé sur le rhin, qui coupe la ville en deux ? post-edit 1 quelle ville de suisse et port fluvial est coupé en deux par le rhin ? post-edit 2 qu’est ce qui est une ville de suisse et un port fluvial situé sur le rhin ? post-edition has been shown to be a valid evaluation method for natural language generation systems (sripada et al., 2005) and the generation of multiple-choice test items (mitkov et al., 2006). 58 question generation for french table 4: question quality clef vikidia-wikimini # quest. annot. 1 annot. 2 # quest. annot. 3 annot. 4 all 100 2.20 2.05 100 2.27 2.38 xip 43 2.37 2.19 64 2.28 2.41 bonsai 73 2.15 2.05 79 2.41 2.61 acronym 6 3.00 2.83 0 – – appositive 3 2.33 2.33 0 – – closed 6 2.33 2.33 22 2.73 2.77 direct object 7 1.57 1.43 21 1.86 1.76 indirect object 5 1.80 1.60 5 2.40 2.60 location 34 2.23 2.09 10 2.20 2.60 subject 26 2.11 1.88 31 2.35 2.71 subject complement 0 – – 1 1.00 1.00 temporal 13 2.31 2.23 10 2.00 1.60 table 5: example questions with their scores score source sentence question 1 une fable est un court récit qui contient une morale, généralement à la fin. à quelle année une fable est-elle un court récit qui contient une morale, généralement ? eng: a fable is a short story that contains a moral, usually at the end. eng: at which year a fable is a short story that contains a moral, usually? 2 la circonférence d’un cercle est la longueur de sa ligne de contour. que la circonférence d’un cercle est-elle ? eng: the circumference of a circle is the length of its contour line. eng: what the circumference of a circle is it? 3 le tabac a été découvert par christophe colomb en amérique. qu’est-ce qui a été découvert par christophe colomb en amérique ? eng: tobacco was discovered by christopher columbus in america. eng: what was discovered by christopher columbus in america? 59 bernhard, de viron, moriceau and tannier the manually post-edited questions were used as targeted reference to compute human-targeted translation edit rate (hter). the translation edit rate measure has been developed to evaluate automatic machine translation (snover et al., 2006). it corresponds to the number of edits needed to transform the output of a machine translation system into a reference translation, normalised by the number of words in the reference translation. the possible edit operations are the insertion, deletion and substitution of single words as well as changes in the position (shifts) of word sequences. we applied this measure to pairs made up of an automatically generated question (hypothesis) and the post-edited question (reference), using the ter-plus evaluation tool.11 table 6 displays the average hter for both corpora, as well as the average number of inserts, deletes, substitutions and shifts. average (avg.) hter corresponds to the average hter values obtained while global hter corresponds to the total number of edits divided by the total numbers of words for all evaluated question pairs. note that both values are multiplied by 100 here. table 6: hter values. clef vikidia-wikimini annot.1 annot. 2 annot. 3 annot.4 avg. sum avg. sum avg. sum avg. sum ins 3.83 153 3.70 174 1.08 56 1.04 25 del 0.05 2 0.13 6 0.69 36 0.88 21 sub 0.48 19 0.98 46 1.19 62 0.83 20 shift 0.00 0 0.00 0 0.29 15 0.04 1 avg. global avg. global avg. global avg. global hter 24.64 22.74 30.32 25.11 30.81 26.67 29.64 23.51 the obtained average hter values are roughly equivalent across both corpora, which indicates that the amount of post-edition work is about the same. in the clef corpus, the majority of the edits were due to insertions in the automatically generated sentences, which shows that the automatically generated questions tend to include extraneous material and are longer than what would be desirable. this is often due to indirect speech, when the reporting clause is kept in the question, e.g. “où l’écrivain américain charles bukowski a-t-il succombé mercredi à une pneumonie, à l’âge de 73 ans, a annoncé jeudi un de ses proches ?” (eng: where did the american writer charles bukowski die on wednesday from pneumonia, at the age of 73, did one his relatives announce on thursday?) in the vikidia-wikimini corpus, the amount of insertions and substitutions tend to be roughly equivalent. moreover, the average number of insertions is lower than in the clef corpus, which shows that the generated questions do not include as much extraneous material, most certainly due to the shorter length of the source sentences. 6.3 comparison with original clef questions we also present a task-oriented experiment aimed at assessing the recall of the system. for this purpose, we automatically compare the questions generated by our system for the clef corpus 11. http://www.umiacs.umd.edu/~snover/terp/ 60 question generation for french with the corresponding questions of the clef campaign12. the goal of this study is to verify to what extent the system is able to produce questions identical or close to the meaning of the questions used in the clef evaluation campaign. for the clef 2005 subset, the system was able to generate 3 identical questions and 7 close questions out of 28. for the clef 2006 subset, the system was able to generate 2 identical questions and 1 close questions out of 12. two examples of automatically generated questions and the expected clef question are shown below: source sentence l’organisation mondiale de la santé (oms) a contredit mardi ses précédentes déclarations rassurantes sur l’épidémie d’ebola au zaïre. clef question qu’est-ce que l’oms ? generated question qu’est-ce que le oms ? source sentence francesco de lorenzo avait été ministre de la santé de 1989 à 1993. clef question quel politicien libéral était ministre de la santé en italie de 1989 à 1993 ? generated question qui avait été ministre de la santé de 1989 à 1993 ? despite the large number of generated questions, the system still does not cover all the questions which could be generated from a sentence. this is due to the following problems: • inferences and external knowledge. in this case, the generation of the expected questions would require inferencing capabilities. for instance, given a sentence such as le japon (...) estime que la création d’un sanctuaire pour les baleines est sans fondement scientifique. (eng: japan believes that the creation of a sanctuary for whales is not scientifically based), the generation of the question quel pays est contre la création d’un sanctuaire pour les baleines en antarctique ? (eng: which country is against the creation of a sanctuary for whales in antarctica?) requires some external knowledge about japan’s position on whaling. this is beyond the capabilities of our qg system in its current state. • missing question generation rules. valuable information for generating questions is often contained in grammatical constituents which are not taken into account by our question typology, in particular noun postmodifiers: – prepositional phrases: passage à la monnaie unique en 1999 → quelle est la date du passage à la monnaie unique ?, (eng: changeover to the single currency→ what is the date of the changeover to the single currency?) – relative clauses: le dalaï lama, qui a obtenu le prix nobel de la paix en 1989→quand le dalaï lama a-t-il été nommé prix nobel de la paix ?, (eng: the dalai lama, who obtained the nobel peace prize in 1989→when was the dalai lama nominated for the nobel peace prize?) this analysis provides important insights for the future developments of the question generation system in particular concerning the question types which should be tackled with the highest priority. 12. recall that the source sentences used for our evaluation correspond to sentences containing the answer to one question from either the clef 2005 or clef 2006 evaluation campaign. 61 bernhard, de viron, moriceau and tannier 6.4 evaluation of question nominalisation we applied the question nominalisation rules presented in section 5.2 to the list of french questions for the clef 2005 and clef 2006 evaluation campaigns, totalling 200 questions each. we did not use the questions generated by our qg system here since errors in the generated questions could have a detrimental effect on the nominalisation module. overall, the system was able to generate 35 nominalisations for the clef 2005 dataset and 30 nominalisations for the clef 2006 dataset. about 1/3 of the nominalised questions were judged to be well formed. for the remainder of the questions, the nominalisation errors can be categorised in 6 different classes, exemplified in table 7. note that the percentages do not sum to 100 as several problems may be present in one and the same question. table 7: question nominalisations and problem categories. category original question nominalised question percent ok quand le roi talal de jordanie abdiqua-t-il ? quelle est la date d’abdication du roi talal de jordanie ? 32 nominalisation is impossible quelle multinationale française a changé son nom pour celui de groupe danone ? quelle est la multinationale de changement de son nom pour celui de groupe danone ? 40 incomplete question quand l’australie, la suède et la finlande entreront-elles dans l’union européenne ? quelle est la date d’entrée de l’australie, la suède et la finlande ? 28 bad nominalisation quelle église a ordonné des femmes prêtres en mars 1994 ? quelle est l’église d’ordre de des femmes prêtres en mars 1994 ? 22 bad question word à qui appartient le milan ac ? quelle est l’appartenance du milan ac ? 20 no noun needed de combien d’étoiles se compose notre galaxie ? quelles sont les étoiles de composition de notre galaxie ? 20 pronominal verb dans quel pays se trouve euskirchen ? quel est le pays de trouvaille d’euskirchen ? 18 these results might seem disappointing at first sight. however, some of the problems should be easy to solve by extending the nominalisation rules, in particular for solving the issue of incomplete questions, or by using additional constraints, for a better selection of the adequate question word or deverbal noun. some of the problems would however require a more complete resource than verbaction, which only relates deverbal action nouns to their verbs and does not provide additional information, e.g. the nominalisation’s argument structure. such kind of information is included in the english nomlex lexicon (macleod et al., 1998), which has been used to produce nominalisation patterns for information extraction (meyers et al., 1998). it would also be helpful to possess resources about other kinds of deverbal nouns, such as deverbal agent nouns, e.g. invent, inventor. this would make it possible to further diversify the nominalisation transformation rules, in order to encompass questions such as the following: “who invented the television?” → “who is the inventor of the television?”. 62 question generation for french 6.5 extrinsic evaluation with a qa system in the last part of the evaluation, we used the automatically generated questions as input for fidji, an open-domain qa system for french and english (moriceau and tannier, 2010). this system combines syntactic information with traditional qa techniques such as named entity recognition and term weighting in order to validate answers through different documents. this particular evaluation consisted in estimating whether fidji was able to analyse the generated questions correctly (question type and expected answer type). we could have evaluated the full answer extraction process, but fidji uses xip for answer extraction in a similar way that we used it for question generation. there would then be a bias since fidji would have found the sentences which have been used for question generation. this experiment aims at investigating two possibilities: 1. an automatic evaluation of the quality of questions, if we can identify that the qa system fails at analysing bad questions but succeeds for good questions. 2. an evaluation of the robustness of the qa system, if it still manages to analyse correctly imperfect questions. question analysis in fidji turns the question into a declarative sentence where the answer is represented by the ‘answer’ lemma. it also aims to identify: • the syntactic dependencies, • the expected type(s) of the answer (named entity type), • the question type: – factoid questions are introduced by a specific wh word or by a specific trigger: ∗ questions in ‘qui’ (‘who’) expect a person or organization named entity (ne). ∗ questions in ‘quand’ (‘when’) expect a date, as well as ‘à quelle date’ (‘at which date’), etc. temporal nes can be made more specific (a year, a day, etc.). ∗ questions in ‘combien’ (‘how much/many’) expect a number, as well as ‘quelle vitesse/température’ (‘what speed/temperature’), etc., with more specific number nes. ∗ questions in ‘où’ (‘where’) expect a location, as well as ‘dans quelle ville’ (‘in which city’), etc. ∗ questions in ‘quel’ (‘what/which’) expect a specific answer type which may not have a corresponding ne type, as ‘quelle déclaration’ (‘which declaration’). – definition questions: ‘qu’est-ce que...’ (‘what is’), ‘que signifie’ (‘what is the meaning of ’), ‘qui est’ (‘who is’), . . . – boolean questions: ‘est-ce que...’, ‘... est-il’ (yes/no question triggers), . . . – complex questions are called ‘comment’ (‘how’) and ‘pourquoi’ (‘why’), but can be introduced by other means, such as ‘pour quelle raison’ (‘for what reason’), ‘dans quel but’ (‘for what purpose’), ‘de quelle manière’ (‘in which manner’). . . 63 bernhard, de viron, moriceau and tannier – list questions are those containing explicitly a plural answer type (e.g. which planets...?, who are the...?, etc.). all the items above are determined by using the syntactic structure of the question. we added to xip’s encrypted grammars some semantic lexical resources: about 200 nouns representing persons (such as teacher, minister or astronaut), organizations, locations, nationalities and numerical value units (currencies, physic units...). we also added to xip’s grammars about 250 grammar rules to deal with the analysis of question syntactic constructions. only those added rules are used for question analysis in fidji. for example: question: quel premier ministre s’est suicidé en 1993 ? (which prime minister committed suicide in 1993?) dependencies: date(1993) person(answer) subj(se suicider, answer) attribut(answer, ministre) attribut(ministre, premier) question type: factoid expected answer type: person these last two features (question type and expected answer type) are then compared with manually assessed types. we performed this evaluation on the same 100 generated questions coming from the clef corpus and evaluated in section 6.2. when both annotators did not agree on the quality of the question (see table 4), we chose the “worst case” for this evaluation (i.e. the lowest value). the first information is that fidji still finds question and answer types, even if the question is really bad (73% of type-1 questions, against 87% of types 2 and 3). contrary to our initial hypothesis, an automatic evaluation of question quality based on fidji’s question analysis is thus impracticable. we then compared the system’s ability to analyse the question correctly, when the quality score was either 2 or 3.13 the results are presented in table 8. 61% of “medium” questions are still correctly analysed, against 73% for perfect questions. this means that a qa system like fidji is robust enough to deal with imperfect questions generated by an automatic system. as a next step, it would be interesting to check whether this finding holds for questions by humans which have been shown to be often ill-formed (ignatova et al., 2008; sitbon et al., 2008). table 8: analysis of generated questions by fidji. perfect medium unworthy questions (“3”) questions (“2”) questions (“1”) total total 26 41 33 100 correct analysis 19 (73%) 25 (61%) / 48 incorrect analysis 7 (27%) 16 (39%) / 19 13. value 1 represents unworthy questions, for which finding the proper types makes no sense at all, since the questions are usually incomprehensible. 64 question generation for french the reasons for fidji’s incorrect analyses are quite different between types 2 and 3 questions. failures on perfect questions are solely due to correct constructions that were unknown to fidji but frequently used as question word reformulations by the generation system, mainly ‘à quel endroit. . . ’ (‘in which place. . . ’ instead of ‘where’) and ‘quelle personne. . . ’ (‘which person’ instead of ‘who’). the automatic generation of variations in question words thus led to identifying gaps in fidji’s question analysis module. for example: question: à quoi correspond le sigle oms ? (what do initials oms stand for?) is correctly analysed as a definition question with the relation acronym(oms, answer), while the construction: question: que représente le sigle oms ? (what do initials oms represent?) is not recognized and leads to an unknown type with dependency obj(represent, answer). on the other hand, failures on type-2 questions come, in addition, from syntactic errors at the beginning of the generated question (as the use of ‘qu’est-ce qui’ (‘what’) when seeking for a person name or ‘pendant quand’ (‘during when’) for an event). this is the case of the two following questions, where the type is not found: question: *qui est 46 ans ? (∼ *who has 46 year-old?) question: *pendant quand le président de l’anc, a-t-il symbolisé la lutte pour le pouvoir de la majorité noire d’afrique du sud ? (∼ *during when the president of anc has been the symbol of...?) 7. conclusion and perspectives we have presented a question generation system for the french language. the system proceeds by transforming declarative sentences into their interrogative counterparts, focusing on one grammatical constituent of the source sentence. this is, to our knowledge, the first available system of this type for french. the evaluation shows that using two different syntactic analysis and named entity recognition tools leads to an increased variety in the questions generated, while preserving the average quality of the questions for different analyzers. we also proposed a question nominalisation procedure which generates nominal structures out of questions containing action verbs. finally, we evaluated the quality of the questions generated by applying the question analysis module of a question answering system, fidji. this analysis showed that fidji is rather robust with respect to ill-formed questions. while the system presented in this article is in principle not targeted at a single application, its expected uses are interactive question answering and educational assessment. we plan to integrate the question paraphrasing system into a qa system in order to improve the answer retrieval process. we are also interested in the automatic generation of educational assessment exercises, e.g. from wikipedia. we envision several directions for future work. first, a better automatic simplification of the source sentences would make it possible to improve the quality of the generated questions. second, we also plan de develop a question classification module, similar to the one proposed by heilman and smith (2010a), in order to distinguish between well-formed and ill-formed questions. 65 bernhard, de viron, moriceau and tannier finally, we would like to stress one important limitation of question generation as implemented in our system, as well as many other similar tools. the input considered for generation is a single sentence, which may or may not have been simplified. we believe that an interesting direction for future work would be the ability to generate transversal questions, relying on pieces of information extracted from several sentences within the same paragraph or document. consider for instance the following sentences: barack obama, born august 4, 1961 honolulu, hawaii, is the 44th and current president of the united states. (...) he was named the 2009 nobel peace prize laureate. the following question contains information extracted from both sentences: which president of the united states was named the nobel peace prize laureate in 2009? this would require moving beyond sentence-level analysis towards discourse-level analysis and in particular the identification of co-reference chains. acknowledgments at the time of this work louis de viron was a student at the université catholique de louvain, belgium and delphine bernhard was at limsi-cnrs, orsay, france. this research was partly supported by the quaero program, funded by oseo. we would like to thank anne garcia-fernandez, anne-laure ligozat and anne-lyse minard for their valuable help in evaluating the system. appendix a. list of tools and resources this appendix lists the tools and resources used in our question generation system. for the xip tool, we did not use the web service described here but our own copy of the tool which was kindly provided by the xerox corporation. tool description availability melt pos tagging free, downloadable from https://gforge. inria.fr/projects/lingwb/ bonsai syntactic analysis free, downloadable from http://alpage. inria.fr/statgram/frdep/fr_stat_ dep_bky.html xip syntactic analysis and named entity recognition web service, see http://open.xerox. com/services/xipparser tregex and tsurgeon pattern matching in trees free, downloadable from http://nlp. stanford.edu/software/tregex. shtml table 9: list of tools. 66 question generation for french resource description availability lefff morposyntactic lexicon downloadable from http://atoll. inria.fr/~sagot/lefff.html lists of proper names ne lexicon built from prolexbase (http://www.cnrtl. fr/lexiques/prolex/) and in-house resources constituted for the system described in elkateb-gara (2005) list of terms related to family lexicon used for person identification retrieved from http://fr.wiktionary. org (category: “lexique en français de la famille”) lists of jobs lexicon used for person identification retrieved from http://fr.wiktionary. org (categories: “métiers en français”, “métiers du secteur primaire en français”, “métiers du secteur secondaire en français”, “métiers du secteur tertiaire en français”) and http://www.onisep.fr/ table 10: list of resources. appendix b. list of tregex rules b.1 questions about the subject that/what clause: /s(sub|c)/=subj [<>,/sent|top/] infinitive clause: /vpinf|iv/=subj [!<<, /prep|pp/ ] & [$++ /vn|fv/=verb] & [>>, /sent|top/=root ] noun phrase: np=subj $++ (/vn|fv/=verb[< /v|verb/=verbalform & ?<-1 vpp=vpp1 & ?< (verb=vpp2 < /partpas/ < /vmod/)]) & [>>, /sent|top/ | $(ponct $pp)] clitic pronoun: cls=subj $++ (v=verb) & [!$++ clo] & [>>, /vn/] & [< /cln|clit/] b.2 questions about the direct object subordinate clause: /sc|ssub/=obj[< top=sent]) | $(vn=verb[< sent=sent])] pronoun, rule 1: /np|cls/=subj [$++ (fv=verb < (pron=cod $+ (pron=coi[<> top=sent) | > (vn=verb < (clo=cod $+ (clo=coi[< sent=sent)] pronoun, rule 2: 67 bernhard, de viron, moriceau and tannier /np|cls/=subj [$++ (fv=verb < (pron=coi $+ (pron=cod[<> top=sent) | > (vn=verb < (clo=coi $+ (clo=cod[< sent=sent)] pronoun, rule 3: /np|cls/=subj [$++ (fv=verb < (pron=cod[<> top=sent) | > (vn=verb < (clo=cod[< sent=sent)] pronoun, rule 4: sc[<< (np=subj $++ (fv=verb < (pron=cod[< top=sent & ! << /coorditems/ | $-(vn=verb[< sent=sent] b.3 questions about the indirect object pp=coi [<< /prep|p/=prep & < np=obj] & [>+(sent) (sent=sent < (vn=verb << v=verbalform $-np=subj $++ pp)) | $-(sc=sent < (fv=verb << verb=verbalform $-np=subj))] b.4 questions about the subject complement ap=obj [$(sc[<< (fv=verb[<< verb=verbalform] & $-np=subj) & > top=sent]) | $(vn=verb[<< v=verbalform & > sent=sent] $-np=subj)] b.5 questions about locative adverbial adjuncts pp=loc [<< /prep|p/=prep & << /lieu|location/] & [>+(sent) (sent < (vn=verb $-np=subj )) | $-(sc < (fv=verb $-np=subj))] b.6 questions about temporal adverbial adjuncts pp=time [<< /prep|p/=prep & << /time|date/] & [>+(sent) (sent < (vn=verb $-np=subj )) | $-(sc < (fv=verb $-np=subj))] b.7 questions about appositives people, rule 1: punct < /,/ . (np=person << /person/ . (punct < /,/)) people, rule 2: np . (ponct < /,/ . (/np|pp/=person << /person/ . (ponct < /,/))) people, rule 3: np=person << /person/ . (punct < /,/ . (np=full . /punct|sent/ ) ) 68 question generation for french people, rule 4: /np|pp/=person << /person/ << /firstname/ . (ponct < /,/ . (/np|pp/=full !<< /person/ . /ponct|sent/ ) ) dates: np=event . (/ponct|punct/ < /,/ . (np=date << /date/ . (/ponct|punct/ < /,/))) b.8 questions about acronyms np=full < (ponct << /lrb/ << np=abbrev << /rrb/) b.9 closed questions np=subj $+ /fv|vn/=verb > /sent|sc/ appendix c. details for resources type description size association association name 81 astronym name of astronomical object 27 building famous building names 88 celebrity famous people names 4,142 city city names 41,992 country country names 450 demonym-association names of association members 8 demonym-astronym names of astronomical objects inhabitants 8 demonym-building names of building dwellers 8 demonym-city names of city residents 63,602 demonym-country names of country inhabitants 1,755 demonym-ethnonym related to an ethnic group 5 demonym-geonym names of geographical locations inhabitants 52 demonym-hydronym names of body of water inhabitants 24 demonym-region names of region inhabitants 2,985 demonym-supranational names of supranational entity inhabitants 180 dynasty dynasty names 57 ensemble group names 14 ethnonym ethnic group members 199 family family member names 137 feast celebration names 11 firm firm names 23,290 firstname given names 23,864 table 11: size of the resources used for list-based named entity tagging and person names tagging (part 1). 69 bernhard, de viron, moriceau and tannier type description size geonym geographical location names 217 history famous historical events 210 hydronym body of water proper names 4,383 institution institution names 68 manifestation famous events 5 meteorology meteorological phenomena 1 object famous object names 1 organization organization names 1,008 prof job names 2,854 pseudoanthroponym e.g. c-3po 3 region region names 2,666 supranational names of supranational entities 91 vessel names of famous vessels 4 way names of famous streets, roads and squares 16 work names of famous books and texts 84 table 12: size of the resources used for list-based named entity tagging and person names tagging (part 2). references salah aït-mokhtar, jean-pierre chanod, and claude roux. robustness beyond shallowness: incremental deep parsing. natural language engineering, 8(3):121–144, 2002. laurie bauer. introducing linguistic morphology. georgetown university press, 2003. 2nd edition. laurence benetti and gilles corminboeuf. les nominalisations des prédicats d’action. cahiers de linguistique française, 26:413–435, 2004. delphine bernhard and iryna gurevych. answering learners’ questions by retrieving question paraphrases from social q&a sites. in proceedings of the 3rd workshop on innovative use of nlp for building educational applications held in conjunction with acl 2008, pages 44–52, columbus, ohio, usa, june 19 2008. delphine bernhard, bruno cartoni, and delphine tribout. a task-based evaluation of french morphological resources and tools, a case study for question-answer pairs. lilt (linguistic issues in language technology), 5(2), 2011. béatrice bouchou and denis maurel. prolexbase et lmf: vers un standard pour les ressources lexicales sur les noms propres. traitement automatique des langues, 49(1):61–88, 2008. jonathan c. brown, gwen a. frishkoff, and maxine eskenazi. automatic question generation for vocabulary assessment. in proceedings of the conference on human language technology and empirical methods in natural language processing, pages 819–826, 2005. 70 question generation for french marie candito, joakim nivre, pascal denis, and enrique henestroza anguiano. benchmarking of statistical dependency parsers for french. in coling 2010, pages 108–116, beijing, china, august 2010. hoa trang dang, jimmy lin, and diane kelly. overview of the trec 2006 question answering track. in proceedings of the fifteenth text retrieval conference (trec 2006), 2006. pascal denis and benoît sagot. coupling an annotated corpus and a morphosyntactic lexicon for state-of-the-art pos tagging with less human effort. in proceedings of the pacific asia conference on language, information and computation, 2009. anne r. diekema, ozgur yilmazel, jiangping chen, sarah harwell, elizabeth d. liddy, and lan he. what do you mean? finding answers to complex questions. in proceedings of the aaai spring symposium: new directions in question answering, 2003. pablo ariel duboue and jennifer chu-carroll. answering the question you wish they had asked: the impact of paraphrasing for question answering. in proceedings of the human language technology conference of the naacl, companion volume: short papers, pages 33–36, 2006. faiza elkateb-gara. l’organisation des connaissances, approches conceptuelles, chapter extraction d’entités nommées pour la recherche d’informations précises, pages 73–82. l’harmattan, 2005. olivier ferret, brigitte grau, martine hurault-plantet, gabriel illouz, christian jacquemin, laura monceaux, isabelle robba, and anne vilnat. how nlp can improve question answering. knowledge organization, 29:135–155, 2002. anne garcia-fernandez. génération de réponses en langue naturelle orales et écrites pour les systèmes de question-réponse en domaine ouvert. phd thesis, université paris sud 11 orsay, 2010. anne garcia-fernandez, sophie rosset, and anne vilnat. macaq : a multi annotated corpus to study how we adapt answers to various questions. in proceedings of the seventh conference on international language resources and evaluation (lrec’10), 2010. simon garrod and anthony anderson. saying what you mean in dialogue: a study in conceptual and semantic co-ordination. cognition, 27(2):181 – 218, 1987. donna m. gates. automatically generating reading comprehension look-back strategy questions from expository texts. master’s thesis, carnegie mellon university, may 2008. danilo giampiccolo, anselmo peñas, christelle ayache, dan cristea, pamela forner, valentin jijkoun, petya osenova, paulo rocha, bogdan sacaleanu, and richard sutcliffe. overview of the clef 2007 multilingual question answering track. in alessandro nardi and carol peters, editors, working notes for the clef 2007 workshop, budapest, hungary, 19-21 september 2007. arthur c. graesser and natalie k. person. question asking during tutoring. american educational research journal, 31(1):104–137, 1994. maurice grevisse. le bon usage. grammaire francaise avec des remarques sur la langue francaise d’aujourd’hui. duculot, gembloux, 10e edition, 1975. 71 bernhard, de viron, moriceau and tannier j. gustafson, a. larsson, r. carlson, and k. hellman. how do system questions influence lexical choices in user answers. in proceedings of eurospeech ’97, pages 2275–2278, 1997. sanda harabagiu, andrew hickl, john lehmann, and dan moldovan. experiments with interactive question-answering. in proceedings of the 43rd annual meeting of the association for computational linguistics, pages 205–214, 2005. nabil hathout and ludovic tanguy. webaffix: discovering morphological links on the www. in proceedings of the third international conference on language resources and evaluation, pages 1799–1804, las palmas de gran canaria, espagne, 2002. elra. nabil hathout, fiammetta namer, and georgette dal. many morphologies, chapter an experimental constructional database : the mortal project, pages 178–209. cascadilla press, 2002. michael heilman and noah a. smith. question generation via overgenerating transformations and ranking. http://www.cs.cmu.edu/ mheilman/papers/heilman-smith-qg-tech-report.pdf, 2009. michael heilman and noah a. smith. good question! statistical ranking for question generation. in proceedings of naacl/hlt, 2010a. michael heilman and noah a. smith. extracting simplified statements for factual question generation. in proceedings of qg2010: the third workshop of question generation, pittsburgh, 2010b. kateryna ignatova, delphine bernhard, and iryna gurevych. generating high quality questions from low quality questions. in proceedings of the workshop on the question generation shared task and evaluation challenge, arlington, va, september 2008. saidalavi kalady, ajeesh elikkottil, and rajarshi das. natural language question generation using syntax and keywords. in proceedings of qg2010: the third workshop of question generation, pittsburgh, 2010. nikiforos karamanis, le an ha, and ruslan mitkov. generating multiple-choice test items from medical text: a pilot study. in proceedings of the fourth international natural language generation conference, pages 111–113, 2006. ronald w. langacker. french interrogatives: a transformational description. language, 41(4): 587–600, october-december 1965. wendy g. lehnert. the process of question answering: a computer simulation of cognition. lawrence erlbaum associates, hillsdale, n.j., 1978. roger levy and galen andrew. tregex and tsurgeon: tools for querying and manipulating tree data structures. in 5th international conference on language resources and evaluation, 2006. xin li and dan roth. learning question classifiers. in proceedings of the 19th international conference on computational linguistics, pages 1–7, 2002. catherine macleod, ralph grishman, adam meyers, leslie barrett, and ruth reeves. nomlex: a lexicon of nominalizations. in proceedings of euralex’98, pages 187–193, 1998. 72 question generation for french adam meyers, catherine macleod, roman yangarber, ralph grishman, leslie barrett, and ruth reeves. using nomlex to produce nominalization patterns for information extraction. in proceedings of the workshop on the computational treatment of nominals, pages 25–32, 1998. ruslan mitkov, le an ha, and nikiforos karamanis. a computer-aided environment for generating multiple-choice test items. natural language engineering, 12(2):177–194, 2006. véronique moriceau and xavier tannier. fidji: using syntax for validating answers in multiple documents. information retrieval, special issue on focused information retrieval, 13(5):507– 533, 2010. véronique moriceau, xavier tannier, and mathieu falco. une étude des questions "complexes" en question-réponse. in actes de la conférence traitement automatique des langues naturelles (taln 2010, article court), montréal, canada, july 2010. gabriel parent and maxine eskenazi. lexical entrainment of real users in the let’s go spoken dialog system. in interspeech-2010, pages 3018–3021, 2010. paul piwek and svetlana stoyanchev. generating expository dialogue from monologue: motivation, corpus and preliminary rules. in proceedings of the 2010 annual conference of the north american chapter of the association for computational linguistics, pages 333–336, 2010. jeffrey pomerantz. a linguistic analysis of question taxonomies. journal of the american society for information science and technology, 56(7):715–728, 2005. helmut prendinger, paul piwek, and mitsuru ishizuka. automatic generation of multi-modal dialogue from text based on discourse structure analysis. in proceedings 1st ieee international conference on semantic computing (icsc-07), pages 27–36, irvine, ca, usa, september 2007. fabio rinaldi, james dowdall, kaarel kaljurand, michael hess, and diego mollá. exploiting paraphrases in a question answering system. in proceedings of the second international workshop on paraphrasing, pages 25–32, 2003. v. rus, b. wyse, p. piwek, m. lintean, stoyanchev s., and c. moldovan. the first question generation shared task evaluation challenge. in proceedings of the 6th international natural language generation conference (inlg 2010), dublin, ireland, 2010. laurianne sitbon, patrice bellot, and philippe blache. evaluating robustness of a qa system through a corpus of real-life questions. in proceedings of the sixth international language resources and evaluation (lrec’08), marrakech, morocco, 2008. matthew snover, bonnie j. dorr, richard schwartz, linnea micciulla, and john makhoul. a study of translation edit rate with targeted human annotation. in proceedings of amta, 2006. somayajulu g. sripada, ehud reiter, and lezan hawizy. evaluating an nlg system using postediting. in proceedings of the international joint conference on artificial intelligence (ijcai 2005), pages 1700–1, 2005. 73 bernhard, de viron, moriceau and tannier svetlana stoyanchev and amanda stent. lexical and syntactic priming and their impact in deployed spoken dialog systems. in proceedings of the 2009 annual conference of the north american chapter of the association for computational linguistics, companion volume: short papers, pages 189–192, 2009. eiichiro sumita, fumiaki sugaya, and seiichi yamamoto. measuring non-native speakers’ proficiency of english by using a test with automatically-generated fill-in-the-blank questions. in proceedings of the second workshop on building educational applications using nlp, pages 61–68, june 2005. noriko tomuro. interrogative reformulation patterns and acquisition of question paraphrases. in proceedings of the second international workshop on paraphrasing, pages 33–40, 2003. noriko tomuro and steven lytinen. retrieval models and q&a learning with faq files. in mark t. maybury, editor, new directions in question answering, chapter retrieval models and q&a learning with faq files, pages 183–194. aaai press, 2004. mickaël tran and denis maurel. prolexbase : un dictionnaire relationnel multilingue de noms propres. traitement automatique des langues, 47(3):115–139, 2006. weiming wang, tianyong hao, and wenyin liu. automatic question generation for learning evaluation in medicine. in advances in web based learning – icwl 2007, volume 4823 of lecture notes in computer science, pages 242–251. springer berlin / heidelberg, 2008. john h. wolfe. automatic question generation from text an aid to independent study. sigcue outlook, proceedings of the sigcse-sigcue joint symposium on computer science education, 10(si):104–112, 1976. shiqi zhao, ming zhou, and ting liu. learning question paraphrases for qa from encarta logs. in proceedings of the 20th international joint conference on artificial intelligence, pages 1795– 1800, hyderabad, india, january 6-12 2007. 74 journal of machine learning research-microsoft word template dialogue & discourse 8(1) 132–150 doi: 10.5087/dad.2017.105 ©2017 natalia levshina and liesbeth degand this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). just because: in search of objective criteria of subjectivity expressed by causal connectives natalia levshina natalia.levshina@uni-leipzig.de leipzig university institute of british studies liesbeth degand liesbeth.degand@uclouvain.be université catholique de louvain institute for language and communication editor: raquel fernández submitted 04/2016; accepted 01/2017; published online 02/2017 abstract the connective because can express both highly objective and highly subjective causal relations. in this, it differs from its counterparts in other languages, e.g. dutch, where two conjunctions omdat and want express more objective and more subjective causal relations, respectively. the present study investigates whether it is possible to anchor the different uses of because in context, examining a large number of syntactic, morphological and semantic cues with a minimal cost of manual annotation. we propose an innovative method of distinguishing between subjective and objective uses of because with the help of information available from an english/dutch segment of a parallel corpus, which is accompanied by a distributional analysis of contextual features. on the basis of automatic syntactic and morphological annotation of approximately 1500 examples of because, every english sentence is coded semi-automatically for more than twenty contextual variables, such as the part of speech, number, person, semantic class of the subject, modality, etc. we employ logistic regression to determine whether these contextual variables help predict which of the two causal connectives is used in the corresponding dutch sentences. our results indicate that a set of semantic and syntactic features that include modality, semantics of referents (subjects), semantic class of the verbal predicate, tense (past vs. non-past) and the presence of evaluative adjectives, are reliable predictors of the more subjective and objective uses of because, demonstrating that this distinction can indeed be anchored in the immediate linguistic context. the proposed method and relevant contextual cues can be used for identification of objective and subjective relationships in discourse. keywords: because, causal connectives, objective and subjective causality, parallel corpus 1 theoretical background and aims of this study this paper deals with the distinction between subjective and objective uses of the english causal connective because, and proposes a method to distinguish between these two uses in discourse. the issue of subjective versus objective (causal) connectives has received a lot of attention in linguistics, going back to seminal studies like rutherford (1970), van dijk (1979) and sweetser just because 133 (1990). these studies demonstrate that connectives like because seem to have systematically different patterns of meaning and use. more recently, the distinction between objective and subjective causal relations has been put forward in studies of coherence relations and their linguistic markers in different languages (canestrelli, 2013; degand & fagard, 2012; pander maat & sanders, 2001; sanders & sweetser, 2009; stukker & sanders, 2012; zufferey & cartoni, 2012). objective causal relations express causality between events in the real world, as in (1), whereas subjective causal relations express the speaker’s motivation for mental conclusions or speech acts, as in (2)1. (1) the tennis match was cancelled because it had been raining too much. (2) apparently it has been freezing, because the geraniums are dead. this distinction is grounded in sweetser’s (1990) seminal trichotomy establishing (causal) relations in the content domain, epistemic domain and speech-act domain, illustrated with her examples (1990: 77–78) in (3-5), respectively. thus, content use is based on the cause-andeffect relationships in the real world; epistemic use introduces the speaker’s reason for making a conclusion, and speech act use expresses the motivation for the speaker’s performing a particular speech act, e.g. asking a question in (5). (3) john came back because he loved her. (4) john loved her, because he came back. (5) since you are so smart, when was george washington born? an alternative classification of causal relations and connectives is in terms of a speaker involvement scale, where speaker involvement “refers to the degree to which the present speaker is implicitly involved in the construal of the causal relation … speaker involvement increases with the degree to which both the causal relation and the related units are constituted by the assumptions and actions of the present speaker.” (pander maat & degand, 2001: 214). the proposed scale for causal relations is given in (6) (based on pander maat & degand 2001: table 1): (6) non-volitional < volitional < causal epistemic < noncausal epistemic < speech act broadly speaking, non-volitional and volitional causal relations correspond to the content domain2; causal and noncausal epistemic to the epistemic domain3, and speech act to the speech act domain, obviously. for more details on the nature of these finer-grained distinctions, we refer the reader to the original publications (pander maat & degand, 2001; degand & pander maat, 2003). more important for our present purposes is the distinct operationalization of subjectivity in the approaches presented, namely a categorical (based on sweetser’s trichotomy) vs. a scalar one 1 degand and fagard (2012) distinguish between objective, subjective and intersubjective causal relations, thus broadly mapping sweetser’s (1990) original domain distinctions. here we propose to follow the bipartite classification in objective content-based causal relations, on the one hand, and subjective relations, on the other hand, the latter including both reasoning-based and speech-act-based causal relations (see also, pander maat & sanders, 2001; sanders & spooren, 2015). 2 the distinction between volitional and non-volitional relations was introduced in rhetorical structure theory (mann & thompson, 1988), and further developed in later work on causal connectives (degand, 2001; pit, 2006). a non-volitional (causal) relation comes about without the intervention of a volitional actor, while a volitional relation results from the will of a conscious, active, participant (sanders & pander maat, 2001). 3 in pander maat and degand’s (2001) proposal, causal epistemic relations are grounded in deductive reasoning, while noncausal epistemic relations are grounded in abductive reasoning. this distinction is not followed here. levshina and degand 134 (pander maat & degand, 2001). it is the categorical approach that we will follow here based on the typical mapping between connective use and (causal) domain. in pander maat and degand’s words, a (causal) “connective encodes a certain speakerinvolvement level, which it contributes to the interpretation of its discourse environment. when this level is too low or too high to be combined with the level allowed for by the discourse environment, the use of the connective is inappropriate.” (p. 230). in other words, there are (semantic) constraints on the relational context in which a given connective can appear. these constraints are of course language-specific (cf. sanders & sweetser, 2009). while the english connective because can be used to express both subjective and objective causal (backward) relations,4 as illustrated in examples (1-4) (see also couper-kuhlen, 1996; ford, 1994; kac, 1972; knott & dale, 1994; knott & sanders, 1998; schleppegrell, 1991), other languages have specialized connectives to express different types of causal relations (pit, 2007; stukker & sanders, 2012; zufferey & cartoni, 2012), and dutch is a case in point. several corpus-based studies have established a systematic relationship between the different types of causal relations and the connectives used to express these relations. more specifically, adopting a categorical point of view, the connective want is a typical marker of subjective relations, and the connective omdat typically expresses objective relations (e.g. degand & pander maat, 2003; pit, 2006; sanders & spooren, 2009).5 these are strong tendencies which have been confirmed both for spoken, spoken-like and written data (spooren, sanders, huiskes & degand, 2010; sanders & spooren, 2015), even if there is no one-to-one relation between the connective and the type of causal relation (see sanders & spooren, 2013). while the distinction between subjective and objective causal relations is conceptually fairly straightforward to grasp, categorizing authentic data in terms of one or the other appears to be a lot less straightforward, not the least because the criteria used to classify less prototypical examples vary (cf. sanders, vis & broeder 2012 on the reusability of causal connective corpus studies). spooren and degand (2010) report kappa agreement values of .60 or lower for variables that are central to determining the degree of subjectivity of a causal relation marked by want. in their discussion of the sources of disagreements, they distinguish two distinct cases: (i) disagreements resulting from “real ambiguities”, i.e. segments which can receive several interpretations; such disagreements are probably inevitable, and (ii) disagreements that are in fact coding errors resulting from a misinterpretation of the categorization variables (pp. 250-251). it is the latter type of disagreements, which are avoidable, which we would like to tackle in this study. the present paper thus pursues the following goals. from a theoretical perspective, we want to find contextual cues that are strongly associated with objective and subjective uses of causal connectives, focusing on the connective because, which can express all of these meanings. this also has a practical use: such cues will help annotators in the future to code the subtle semantic distinctions and make the annotation more reliable. to reach these goals, we propose a semi-automatic approach based on an innovative combination of the data from a parallel corpus and the methodology of distributional semantics. the cross-linguistic data enables one to use a more specific language to make predictions about another language that does not make certain distinctions. in this study, we disambiguate the functions of the english because with the help of dutch, a language which makes the distinction 4 in a backward causal relation, the causal segment is the host of the causal connective and is in general preceded by the consequence segment, as in (1): segment1 = the match was cancelled (effect or consequence) because segment2 = it has been raining (cause). 5 dutch also has a highly specific connective doordat (and its forward counterpart daardoor) specialized to express nonvolitional (highly objective) causal relations. this connective is however a lot less frequent (degand, 2001; pit, 2003). with only 62 occurrences, it was also relatively rare in our data (see section 2.1). just because 135 between more objective and more subjective causal relations by means of different connectives, most importantly, omdat (more objective) and want (more subjective). for this purpose, we use aligned english and dutch sentences from the opus version of the europarl corpus (koehn, 2005; tiedemann, 2012), which contains proceedings of the european parliament. examples (7-9) illustrate the principle. in (7), the english because segment is aligned with dutch omdat, expressing objective causal relations (non-volitional in 7a and volitional in 7b). (7) a. opus europarl 33356150 listen to and see the agony of a doctor who has to tell a non-smoker that she has cancer because she breathed the smoke of someone who thought that smoking looked cool on a screen or on a page. stelt u zich de pijn voor die een arts voelt wanneer hij een nietrookster moet meedelen dat zij kanker heeft omdat zij de rook heeft ingeademd van iemand die vond dat roken “cool” stond op televisie of in de pers. b. opus europarl 2458129 i voted for the atkins report because it is extremely important for pensioners… ik heb voor het verslag-atkins gestemd omdat het voor de gepensioneerden (...) zeer belangrijk is... in contrast, want expresses epistemic subjective causality, as in (8a), where the subordinate clause provides an explanation for the speaker’s inference presented in the main clause, and subjective speech-act relations, as in (8b), where the subordinate clause contains the reason for asking the question in the main clause. again, both segments are aligned with english because: (8) a. opus europarl 33885879 mr henderson knows that because he is on the council. dat weet de heer henderson want die zit in de raad... b. opus europarl 2439113 is it available, because i have noticed that many meps have not seen this text? is hij ter beschikking, want ik merk dat heel veel collega’s deze tekst niet gezien hebben. the second component of our method involves the use of distributional semantics (wittgenstein, 1953; harris, 1954; firth, 1957, etc.). according to this approach, semantic properties of words and constructions are closely linked with contextual environments where these words or constructions occur. moreover, different senses of a word or construction will be observed in different types of contexts. in other words, these senses will have different distributional properties. this idea goes back to structuralist semantics (e.g. apresjan, 1966) and has been more recently implemented in automatic algorithms of word sense disambiguation (e.g. pedersen, 2006). we will use a bottom-up approach and employ contextual variables, which can be coded semilevshina and degand 136 automatically with the help of syntactic and morphological information about the english sentences with because. these variables will be investigated with the help of logistic regression analysis in order to select those contextual features that can be used for distinguishing between objective and subjective uses of because, which correspond to omdat and want, respectively. the idea of using distributional clues for subjectivity analysis is not new. in fact, there has been some work on subjectivity word sense disambiguation (swsd) in computational linguistics (e.g. akkaya et al., 2009). consider (9) as an illustration. the noun alarm has an objective meaning in (9a) and a subjective one in (9b): (9) a. the alarm went off. b. his alarm grew. (akkaya et al., 2009: 191) the task for a swsd algorithm is to determine whether the word is used subjectively or objectively on the basis of contextual clues. although the type of subjectivity and objectivity in the lexicon is different from the one established at the clausal level, this task is essentially rather similar to ours. however, there are important differences. first, the purpose of swsd is to classify words, sentences or texts as correctly as possible. the process of deciding for subjective or objective meaning is ultimately a black box. in our study, we develop a method of automatic disambiguation, too. our primary goal, however, is to learn which contextual features can help us discriminate between the subjective and objective uses of because. second, the contextual cues in our study are at a higher level of abstraction than those in swsd. we employ diverse syntactic and morphological information about the clauses that are connected by the conjunction, as well as the semantic classes of the subjects and predicates, since these features have been demonstrated to work well for disambiguation of discourse relationships expressed by connectives (e.g. pitler & nenkova, 2009). in contrast, the swsd approach is usually based on more lexically specific clues (surrounding words). the rest of the paper is organized as follows. section 2 presents the data (parallel corpus) and contextual variables. section 3 reports the results of the statistical analysis. finally, section 4 summarizes the findings and suggests some directions for further research. 2 data and contextual variables this section describes the data and variables that were used in this study. 2.1 data source and extraction procedure the data come from europarl, a collection of the european parliament proceedings in 21 languages of the european union (koehn, 2005), which constitutes a part of the opus corpus (tiedemann, 2012). this corpus was chosen because the proceedings contain many uses of causal clauses in various functions, such as defending one’s political position, justification of requests, explaining why one chooses a particular wording, or introducing the causes and consequences of some socially relevant events. the opus query engine was used to extract 5,000 examples of contexts that contained because in the english version.6 these contexts were manually checked. we kept only those sentences where because was used as a subordinate conjunction (thus excluding the preposition because of) and corresponded to omdat or want in the dutch version. the direction of translation was not taken into account. the initial large sample contained a large number of repetitions. this is why we had to remove all repeated sentences. we also discarded the contexts in the following cases: 6 see http://opus.lingfil.uu.se (last access 01.02.2017). http://opus.lingfil.uu.se/ just because 137  causal connectives with adverbial modifiers in dutch, e.g. p, juist omdat “just because” q; p, niet in de laatste plaats omdat “not least because” q;  paired conjunctions in dutch, e.g. p, niet (alleen) omdat “not (only) because” q, maar (ook) omdat “but (also) because” r;  the subordinate clause with the causal connective in the preposed position in dutch (omdat q, p). in all these cases, only omdat can be used, want being excluded for syntactic reasons. this is why it would not make sense to include these contexts in the sample. as a result of this cleaning procedure, the final sample contained 1521 examples in total, 798 with omdat and 723 with want. these sentences (in the english version) were first analysed manually, so that the clauses that represented the cause and effect were extracted. after that, the clauses were parsed automatically with the help of the stanford parser (klein & manning, 2003). as a result, we obtained the syntactic dependencies and morphological information (part of speech) about every word in a clause. with the help of a python script written specifically for this purpose, this morphological and syntactic information was used for data annotation. the variables are presented in section 2.2. the annotation was manually checked. 2.2 contextual variables this section describes the contextual variables that were used for distinguishing between the uses of because that correspond to omdat and want in the dutch segment of europarl. all of these variables, except for the one describing the dutch connectives, represent the structural and semantic properties of the english sentences. first, we describe the variables related to the entire sentence. many of these variables were inspired by corpus analytic work on the expression of subjectivity and objectivity in language, mainly pit (2003, 2006) and torres cacoullos and schwenter (2005). the motivation for most of the variables is specified below. note however, that a decisive criterion for including a variable was whether it was possible to code it automatically on the basis of the available syntactic and morphological information from the parser. an exception is the semantic coding of the subjects, but in the future, we hope to be able to perform semantic classification automatically, too, with the help of the state-of-the-art word sense disambiguation methods and semantic resources, such as wordnet.  dutch connective: whether the dutch equivalent of because was omdat or want.  the presence or absence of direct address in the sentence with because. for example, (10) contains direct addresses mr president and ladies and gentlemen: (10) opus europarl 37571406 mr president, ladies and gentlemen, i voted against the thyssen report, because it is detrimental in many respects to certain business sectors and of course to people starting up in business. if the speaker uses a direct address, he or she highlights the interactive character of discourse, which makes it more subjective.  the presence or absence of evaluative adjectives (positive or negative). this variable was added because manual analysis has shown that subjective causal relations tend to cooccur with evaluative adjectives and/or adverbs, while objective causal relations do not (pander maat & degand, 2001; pit, 2006; sanders & spooren, 2015). the variable was levshina and degand 138 coded with the help of sentiwordnet (baccianella, esuli & sebastiani, 2010). in this data base, which has the structure similar to one of wordnet (fellbaum, 1998), words have scores along the positive, negative and objective dimensions. we considered the word emotionally charged if its scores either on the negative or positive scales were greater than zero (the scores were averaged across the meanings). consider (11): (11) opus europarl 94416 colleagues, i am in a very difficult position because i cannot change the agenda. this sentence contains the adjective difficult, which has negative scores 0.75 (in the sense ‘hard’) and 0.625 (in the sense ‘unmanageable’, as in difficult child). we expect the more emotionally charged contexts to be more subjective. the coding was done automatically; no word sense disambiguation was performed. the aim was to see how far we can go with a fully automatic approach.  the presence or absence of evaluative adverbs. the logic and the procedure were the same as above.  explicit references to conceptualizer and his or her conceptualization of causal relationships. in other words, the mental process of establishing a causal connection becomes explicit. this information could be present in structures like i/we/x think/believe/consider/feel/trust… that p because q, where q gives a reason for my/our/x’s thinking, believing, etc. that p. according to langacker (1985), an explicit reference to the conceptualizer objectifies him or her. usually, omdat is used in these cases, as in (11), whereas want is used with implicit conceptualizers (cf. pit, 2003; sanders & spooren, 2009). compare the examples in (12) and (13): (12) opus europarl 2347145 (…) i think that it is imperative to discuss plans for extending it and to look at pricing again, because we simply have to get a hold on the situation. (…) ik denk dat het dringend noodzakelijk is opnieuw over uitbreidingsplannen en prijzen te praten omdat we de situatie gewoon onder controle moeten krijgen. (13) opus europarl 5709553 this is necessary because it is the only way to achieve an increasingly balanced labour market. dat evenwicht is ook nodig, want alleen op die manier kunnen we zorgen voor een evenwichtige groei van de werkgelegenheid. the variables that are listed below were coded for the main and subordinate clauses separately: just because 139  part of speech of the subject of the clause. the values are ‘noun’ (including nominalizations, e.g. the rich), ‘pronoun’ (all possible pronouns) or ‘no subject’ (in case there are no subjects). the subject of a clause very often corresponds to the causally primary participant (or cp, pit, 2006) which plays an important role in the causal conceptualization of the state or event. according to pit (2006: 162) “cps referred to by a nominal are more objective (less deeply perspectivized) than cps referred to by a pronominal” (cf. langacker, 1985: 126–127). furthermore, subject coreferentiality (morphologically coded by pronouns) has been shown to reflect subjective use (torres cacoullos & schwenter, 2005);  grammatical person of the subject: ‘1st’, ‘2nd’, ‘3rd’ or ‘no subject’. this variable is motivated by the assumption that first and second person participants are more subjective than third person participants;  grammatical number of the subject: ‘singular’, ‘plural’ or ‘no subject’. one can expect subjective contexts to be more associated with singular cognizers;  semantic class of the subject: ‘animate’ (including people, animals and organizations), ‘inanimate’ (all the rest) or ‘no subject’. the assumption is that subjective contexts are associated with animate subjects more than with inanimate ones;  tense of the finite predicate: ‘present’, ‘past’, ‘future’, ‘other’ (modals and imperatives), ‘no predicate’ (in case there is no finite predicate), where present tense is assumed to mark more subjective relations (in the “here and now”) (pit, 2003; pander maat & degand, 2001);  voice of the finite predicate: ‘active’, ‘passive’ or ‘no predicate’. one would expect passive forms to be more associated with objective contexts (biber 1988);  the presence of a modal verb in the predicate: ‘yes’, ‘no’ or ‘no predicate’ (see above), where modal verbs would signal greater subjectivity;  polarity: ‘positive’ or ‘negative’ (when the clause contains a negation). we expect negative polarity contexts to be more subjective because negation is central in political debates, where speakers correct or reject different proposals;  semantic class of the verbal predicate: ‘mental’ (verbs of perception, desire, thinking, resolution, etc.), ‘social’ (verbs of communication), ‘other’ (all other verbs) or ‘no predicate’. mental and social verbs are expected to be more frequently present in subjective contexts. the semantic class annotation of the subjects was tested first on a small sample of 200 observations by the co-authors. for the main clauses, cohen’s kappa was 0.89, and for the subordinate clauses, it was 0.79. from that we concluded that the annotation schema was reliable enough and coded the entire sample. due to much lower kappa scores for the verb classes in the pilot study, it was decided to code the verbs following a closed list approach. the list was based on levin’s (1993) classes, where the mental verbs included the classes “declare”, “conjecture”, “see”, “sight”, “peer”, “feel”, “admire”, “marvel”, “want”, “long”, positive and negative judgement verbs, “assess”, “investigate”, whereas the communication verbs included the classes “tell”, “snap/cackle”, “cable”, “talk”, “chitchat”, “say”, “complain” and “advise”. for an illustration of the coding schema, consider the example in (14): (14) opus europarl 18558329 we want a new framework agreement because without the ability to scrutinise your commission effectively, we cannot do our job properly. levshina and degand 140 we willen een nieuwe kaderovereenkomst, want zonder het vermogen uw commissie doeltreffend te controleren, kunnen wij ons werk niet goed doen. the sentence has the following values: 1. dutch conjunction: want 2. variables coded for the entire english sentence:  direct address: no  evaluative adjective: yes (new)  evaluative adverbs: no  explicit reference to conceptualization of causation: no 3. variables coded for the main and subordinate clauses separately 3.1. the properties of the main clause  part of speech of the subject of the main clause: pronoun  person of the subject of the main clause: 1st  number of the subject of the main clause: plural  semantic class of the subject of the main clause: animate  tense of the finite predicate of the main clause: present  voice of the finite predicate of the main clause: active  modality of the main clause: no  polarity of the main clause: positive  semantic class of the verbal predicate of the main clause: mental 3.2. the properties of the subordinate clause  part of speech of the subject of the subordinate clause: pronoun  person of the subject of the subordinate clause: 1st  number of the subject of the subordinate clause: plural  semantic class of the subject of the subordinate clause: animate  tense of the finite predicate of the subordinate clause: present  voice of the finite predicate of the subordinate clause: active  modality of the subordinate clause: yes  polarity of the subordinate clause: negative  semantic class of the verbal predicate of the subordinate clause: other the relevance of these variables for the disambiguation of subjective and objective causal relations was investigated in the statistical analyses that are presented in section 3. 3 quantitative analyses: data transformation and logistic regression this section first describes the data transformation procedures and next reports the results of our logistic regression modelling. 3.1 data transformation regression analysis is sensitive to low-frequency values and strong associations between predictors. these factors can seriously undermine the quality of a model. this is why some of the just because 141 initial variables were conflated, as well as some values of the variables, on the basis of standard regression model diagnostics (see levshina, 2015: ch. 12). the transformed variables were the following:  semantic class, part of speech and person of the subject, both in the main and subordinate clause. the resulting variables were called simply ‘subject’ and contained the values ‘speech act participants (sap)’ (i.e. 1st and 2nd person pronouns), ‘animate’ (all other animate subjects) and ‘inanimate’.  tense of the verbal predicate, both in the main and subordinate clauses. to optimize the analyses, we recoded the variable as ‘past’ and ‘non-past’ (present, future, other, no predicate) on the basis of the model diagnostics. all cases with incomplete sentences without verbal predicates or subjects were removed, because they produced data sparseness and made the model suboptimal. in total, we had 1512 observations left. 3.2 logistic regression model we performed a binary logistic regression analysis. this is a method for modelling the binary outcome (here, omdat or want) that can be predicted from other variables, which are usually called predictors. here, the predictors were the variables that describe the english contexts. we fitted a multiple regression model, where the effect of each predictor of interest was measured while controlling for the other predictors. we used the free statistical software r (r core team, 2015) with an add-on package rms (harrell, 2015). the model structure was determined by using the following procedure. first, a full model with all pairwise interactions was defined on the basis of bidirectional (backward and forward) stepwise selection based on all predictors and all possible pairwise interactions between them. interactions are observed when the effect of two or more variables on the outcome is non-additive. stepwise selection means that the algorithm tries to add (forward selection) or remove (backward selection) a variable or interaction term one by one, until no further improvement can be made. the criterion for improvement was aic (akaike information criterion), which shows how well a model fits the data, while at the same time giving advantage to more parsimonious models with a smaller number of predictors. after that, usual diagnostic tests were performed. unfortunately, a bootstrap validation revealed that the model suffered from severe overfitting. in particular, the optimism, which is commonly used as a diagnostic statistic, was about 0.38 in the slope and 0.11 in r2 (cf. harrell, 2001). these levels were too high to be tolerated. all that means that the model would be useless when applied to new data and has thus little scientific value. for a detailed explanation of the statistical procedures, see levshina (2015: ch. 12). to solve that problem, we tried to minimize the number of interactions by excluding all nonsignificant ones and those that do not change the direction of the predictors’ effects. this did not help to solve the problem of overfitting, even when a penalty was applied for shrinkage (ibid.). the only model with acceptable optimism was the one with main effects only, selected on the basis of backward elimination with aic as the criterion (see table 1). still, we had to apply a penalty to shrink the estimates (with the penalty factor of 6.8) in order to correct for the undue optimism. we also tested possible interactions between the predictors manually and inspected them graphically, but they did not reveal significant cross-over effects and therefore were not included in the final model. although some variables are clearly related (e.g. only animate subjects can take mental verbs as predicates), logistic regression is known to be robust with regard to some correlations between predictors. moreover, there were no symptoms of strong multicollinearity, since all vif (variance levshina and degand 142 inflation factor) scores, which are used traditionally for diagnostics, were below 5. all this means that we do not have reasons for concern regarding the quality of the model. the predictive power of the final model (i.e. how well it could discriminate between the contexts that corresponded to want and omdat in dutch), was modest, with the concordance index c = 0.67 and nagelkerke’s r2 = 0.13. as a rule of thumb, the c value should be at least 0.7 for the model to be considered good. the accuracy, defined as the proportion of correct predictions made by the model, was 0.62. this is greater than the baseline level of 0.52, which would be the probability of making a correct prediction if one always selected the more frequent response, i.e. omdat. in our view, this result can still be regarded as satisfactory because we try to predict the use of a dutch connective from its equivalent contexts in english, where numerous other factors play a role (see discussion in section 4). parameter coefficient se p-value intercept -1.56 0.23 < 0.001 sap subject in main clause (in contrast with animate) 0.81 0.18 < 0.001 inanimate subject in main clause (in contrast with animate) 0.85 0.18 < 0.001 singular subject in main clause 0.23 0.13 0.089 mental verb in main clause 0.35 0.14 0.012 verb of communication in main clause 0.05 0.21 0.826 non-past tense in main clause 1.03 0.19 < 0.001 passive predicate in main clause -0.30 0.22 0.17 modal verbs in main clause 0.55 0.14 < 0.001 sap subject in subordinate clause (in contrast with animate) 0.43 0.16 0.009 inanimate subject in subordinate clause (in contrast with animate) 0.48 0.15 < 0.001 singular subject in subordinate clause 0.28 0.13 0.029 passive predicate in subordinate clause -0.27 0.17 0.112 modal verb in subordinate clause 0.48 0.16 0.003 evaluative adjective(s) 0.19 0.11 0.076 table 1. logistic regression coefficients and their standard errors and p-values. the final regression estimates are presented in table 1. the coefficients in the second column of the table are log-odds ratios. an exception is the intercept, which represents logarithmically transformed odds of want against omdat in the reference level context, i.e. when all variables have the default values, or the values opposite to those displayed in the table. positive coefficients (logodds ratios) show that this value of a variable increases the odds of want in the dutch sentence in comparison with omdat. negative coefficients, in contrast, indicate that the value increases the likelihood of omdat in comparison with want. the other columns in the table contain the standard errors, which give an idea of how variable the estimates may be, as well as the p-values, which are used in frequentist statistics to determine if the effect is statistically significant. a p-value below the conventional level of 0.05 serves as an indication that the effect is not due to chance alone. the values between 0.05 and 0.1 are often considered as marginally significant. in a pilot study, it makes sense to try to interpret the marginally significant variables (here, between 0.05 and 0.2), as well, so that they can be more carefully investigated in larger-scale follow-up studies. the estimates suggest the following. both in the main and subordinate clauses, the preferences are rather similar. first, the presence of modal predicates, singular subjects (marginally significant just because 143 in the main clause), sap subjects (or first and second person subjects) and inanimate subjects in comparison with animate 3rd person subjects increase the odds of want in the dutch version. moreover, there are marginally significant effects of passive predicates, both in the main and subordinate clauses. the passive forms tend to increase the chances of omdat to be found in the dutch sentence. non-past tense forms and mental verbs (as opposed to verbs of communication and all other verbs) in the main clause increase the odds of want. there is also a marginally significant effect of the presence of evaluative adjectives, which boost the chances of want, too. the effects of saps as subjects, evaluative adjectives, mental verbs, non-past and nonpassive verbs are theoretically interpretable. most of these features are typical of involved nonabstract, non-technical communication (biber, 1988), which is characterized by a high degree of subjectivity. this kind of communication is contrasted with informative and abstract, technical types of discourse. the distinction between involved and informational and abstract communication closely corresponds, in our opinion, to the distinction between subjectivity and objectivity of discourse. in involved communication, the speaker and the hearer, their beliefs, attitudes and intentions are in the centre of attention. in contrast, abstract technical communication distances from the speaker and hearer’s personal experiences, focusing instead on the properties and events of the external world. in addition, the use of modals is a typical feature of explicit marking of the speaker’s own point of view or, alternatively, of argumentative discourse designed to persuade the addressee (idem.: 111). therefore, the use of modals is indicative of subjective communication. it is more difficult to explain, however, why the plural and 3rd person animate subjects disfavour want. a close inspection of the individual contexts suggests that these are often the names of social and political groups and entities, including members of a party (15a) and representatives of countries (15b). (15)a. opus europarl 912705 mr president, let me repeat what i said yesterday, namely that the french socialists will not take part in the vote because they believe that this is not the correct procedure, since the only environmental directives referred to are those on wild birds and on natura 2000 and a more balanced approach should have been taken. b. opus europarl 689846 that really is a brave decision, because some member states will clearly find themselves towards the bottom of the league table, which is something nobody likes. in these contexts, the people are conceptualized as political groups, rather than as individual subjects, which explains why the more objective causal connective is used. the regression model data also enable us to compute the so-called fitted values of every example in the data set, i.e. the predicted probabilities of want and omdat based on the values of the predictors and their coefficients in the regression model. the two sentences with top highest predicted values of want are given in (16a), where the predicted probability of want was 80%, and (16b), where the predicted probability of the connective was 78%. this means that the sentences contain many features that boost the probability of want. not surprisingly, the corresponding dutch sentences contained the predicted connective. (16) a. opus europarl 25188486 levshina and degand 144 well, mr. belder and colleagues from the ppe-de and alde groups, either your homework has not been done properly, or we must congratulate the magical powers of the commission, because two years ago it must have eaten some chinese fortune cookie which said that in september 2006 parliament would make such a call to initiate a structured dialogue. welnu, geachte heer belder en collega’s van de ppe-de-fractie en de aldefractie, of u heeft uw huiswerk niet goed gedaan, of we moeten dankbaar zijn voor de magische krachten van de commissie, want twee jaar geleden moet de commissie een of ander chinees gelukskoekje hebben gegeten waarop stond dat het parlement haar in september 2006 zou vragen het initiatief te nemen voor een gestructureerde dialoog. b. opus europarl 23706492 there must be a typing error in the commission’s speech, because i would have thought you would very happily have looked forward to countries introducing more stringent legislation in order to achieve the kyoto objectives. verder vraag ik mij af of er geen tikfout in de getallen van de commissie zit, want u zou het toch met vreugde hebben begroet als landen striktere wetgeving invoeren om de doelstellingen van kyoto te bereiken? the examples in (16a) and particularly in (16b) are cases of an epistemic use in english. interestingly, the dutch version of (15b) contains a reported question in the main clause ik vraag mij af of… “i ask myself whether…” and a question in the subordinate clause with the modal particle toch, which reflects the speaker’s desire to be reassured or confirmed. for omdat, the sentences with the highest predicted scores (96% and 95%, respectively), are shown in (17a) and (17b). again, the dutch versions contain the predicted conjunction omdat. (17) a. opus europarl 20477436 these countries were deprived of their sovereignty for many decades because they had no partner prepared to perform the duties of an ally without hesitation. ...die gedurende decennia verschrikkelijk hebben geleden en hun soevereiniteit kwijt waren, omdat zij niet over een partner beschikten die er niet voor terugschrok zijn verplichtingen als bondgenoot na te komen. b. opus europarl 16019572 a constituent of mine bought a flight online with the notorious ryanair but when they went to collect the ticket they were denied access because they had an international student id card and were refused boarding on the grounds that it was out of date. … toen deze persoon het ticket ophaalde mocht hij niet aan boord omdat hij een internationale studentenkaart had die verlopen was. just because 145 in both cases, the subjects of the main clauses are non-agentive, since they are affected by someone else’s actions. the causes are either historical consequences, as in (16a), or impersonal rules (16b). thus, the causal relationships are construed as objective. 3.3. cases of mismatches it is also instructive to study the cases of mismatches, where the dutch connective used in a particular context has a low predicted probability. this could help us identify the reasons why the prediction is far from being perfect. e.g. whether there are additional contextual variables that need to be taken into account, or whether there are some structural differences between english and dutch that distort the picture. in order to identify the cases where the predictions made by the model differed the most from the observed dutch connective, we performed the following. first, we computed the predicted scores, as was shown in section 3.2. the greater the score, the higher the probability of want. next, we binarized the observed categories, with omdat having the value 0 and want corresponding to 1. after that, we computed the differences between the binarized outcome and the predicted scores. the observations with the greatest absolute differences between the observed and predicted scores were examined. the overwhelming majority of the examples with the greatest mismatch scores contain want, but the model predicts omdat with a very high probability. example (18) shows the observation with the greatest mismatch, where omdat was predicted with the probability of almost 90%. (18) opus europarl 24326485 i would like to stress, however, that the decisions leading to such a situation were probably not taken by women, because there are practically no women in the places where decisions are taken on security policy or at negotiation tables. naar alle waarschijnlijkheid waren het echter geen vrouwen die de beslissingen namen die tot die situatie hebben geleid, want op plekken waar wordt besloten over veiligheidsbeleid en aan de onderhandelingstafels is vrijwel geen vrouw te vinden. in spite of the fact that the english sentence contains such omdat-favouring features, as the passive predicate, inanimate subject in the main clause and past tense, it expresses, however, the speaker’s conjecture. the sentence contains the adverb probably, which expresses epistemic modality. the subordinate clause provides the speaker’s grounds for making this conjecture. this marker corresponds to the phrase naar alle waarschijnlijkheid “by all odds” in dutch. thus, epistemic modality markers may be a new variable that should be added to the list of markers. another example of a mismatch is provided in (19). again, the sentence has a high probability of omdat (almost 87%), but the dutch translation contains want. this sentence expresses the speaker’s evaluation of someone else’s actions (‘x was right to do y because…’), and provides the reason for this evaluation. in dutch, the evaluation is present, as well, although the structure is somewhat different (‘x rightly did y because…’), with the adverb terecht “rightly”. (19) opus europarl 18486909 mr schulz was right to draw attention to commissioner vitorino s important role, because the excellent result has been achieved partly thanks to his input and influence. levshina and degand 146 de heer schulz heeft terecht gewezen op de belangrijke rol van commissaris vitorino, want mede door zijn inbreng en invloed is een heel goed resultaat geboekt. the other examples of mismatch where the subjectivity of a sentence is not captured by the model are similar to the ones provided above. they show that subjectivity can be expressed by a rich variety of linguistic strategies. a full account of such strategies and their automatic extraction remains a task for future research, however. although some of these subjectivity indicators are lexical and easy to list, e.g. english probably and dutch terecht “rightly”, many other expressions are periphrastic and more difficult to capture automatically. for example, the phrase ‘x was right to do y because…’, allows for a large number of possible modifications, such as x was completely right not to openly criticize this idea or nobody can accuse me of having been wrong to do so. another possible reason of mismatches is a non-perfect equivalence between the english and dutch versions. consider (20), where the speech act (directive) in english becomes a modal clause with moeten “must” mitigated by ik vrees dat… “i’m afraid that…” in dutch. (20) opus europarl 38368559 let us try to redirect the situation because your reply has also been very general. ik vrees dat we iets dieper op de zaak zullen moeten ingaan, mijnheer de commissaris, want uw antwoord blijft nogal algemeen. “i’m afraid we must go somewhat deeper into the matter, mr. commissioner, because your reply is still very broad.” however, in the overwhelming majority of our examples, the features of the dutch sentences are faithfully reflected in the english contexts, and vice versa. the lack of full structural and semantic correspondence is thus not the main factor that can explain the mismatches. 4. conclusions and outlook using the data from a parallel corpus, we have managed to predict the choices between the more objective dutch causal connective omdat and the more subjective want on the basis of the contextual properties of the english sentences with because. we found that such semantic and syntactic features as modality, semantics of referents (subjects), semantic class of the verbal predicate, tense (past vs. non-past) and the presence of evaluative adjectives are significantly associated with the use of omdat or want in the dutch sentence. this means that our pilot study has shown promising results in disambiguation between objective and subjective uses of causal connectives. in view of the high cost of manual annotation of discourse connectives, we would like to suggest that the semantic and syntactic features we identified as predictors of subjective and objective contexts may be used to automatically annotate (or rather pre-annotate) the data. in a second step, this automatic annotation would then be verified or corrected by manual analysis thus gaining time and analysis effort. we hope that these features will be applied in new research on other discourse markers. if we aim at facilitating the annotation process, its reliability should also be taken into account. in addition to being work-intensive, manual annotation of discourse connectives is known to give rise to fairly low interrater agreement (spooren & degand, 2010). an just because 147 interesting question is therefore whether the automatic annotation compares to the manual one, or would even outrank it. a way to find out is to manually code part of the data in order to calculate agreement between the automatic and the manual annotation. however, it is worthwhile to question whether the manual annotation is really the more reliable than the automatic one. this is an issue which will have to be left for future research. another area in which our approach could also be useful is objective measuring of the degree of subjectification and intersubjectification in grammaticalization, for instance by uncovering the contextual linguistic features that progressively anchor emerging subjective meanings (see torres cacoullos & schwenter, 2005). yet, judging from the modest discriminating power of the model, the model is far from being perfect. the most important reason, as suggested by our analysis of mismatches between the observed connectives and the ones predicted by the model, is our operationalization of the subtle pragmatic and semantic functions with the help of very coarse-grained and discrete contextual features. the strategies of expressing subjectivity in discourse are very diverse and present a challenge for automatic identification. a comprehensive list of such markers and constructions is not available, to the best of our knowledge, and their grammatical structures vary greatly crosslinguistically. moreover, the mismatches can be explained by the usage itself. the semantics of causal connectives, similar to that of many other linguistic categories, has a prototypical structure, with the core functions and periphery (stukker, 2005). when a causal connective is used nonprototypically, this may be due to the speaker’s intention to change the construal of a causal relationships for rhetorical purposes (stukker & sanders, 2012). such modulations can be detected only on the basis of a careful contextual analysis, as it is done, for example, in sanders and spooren (2013). yet, we are convinced that the gains of our approach outweigh its limitations. namely, our approach provides objective criteria for subjectivity and thus helps the linguist to avoid circularity in his or her semantic and pragmatic analyses. notably, the best predictors in our analysis are the usual suspects in analysis of register variation when it comes to the dimensions of involved and non-abstract non-technical communication as opposed to informational and abstract, technical discourse (biber, 1988). one might wonder if the other features associated with involved communication in general, such as the use of general emphatics and discourse particles, hedges and amplifiers, or the features associated with abstract, technical discourse (e.g. conjuncts, past participial clauses and adverbial subordinators) might be relevant predictors in distinguishing between more subjective and objective uses of discourse markers. these interesting questions are left for future research. acknowledgements the present study was a part of the project “mapping the causative continuum: a multivariate typological investigation of causative constructions based on a multilingual parallel corpus”, which was carried out by the first author at the university of louvain in louvain-la-neuve (2013–2015) and which was funded by the belgian research foundation f.r.s – fnrs. the first author is grateful to these two institutions for the generous financial, administrative and scientific support. all usual disclaimers apply. references cem akkaya, janyce wiebe and rada michalcea (2009). subjectivity word sense disambiguation. proceedings of the 2009 conference on empirical methods in natural levshina and degand 148 language processing, pages 190–199, singapore, 6-7 august 2009. available online at http://www.aclweb.org/anthology/d09-1000 (last access 01.02.2017). juri apresjan (1966) analyse distributionelle des significations et champs sémantiques structures. langage, 1(1):44–74. stefano baccianella, andrea esuli and fabrizio sebastiani (2010). sentiwordnet 3.0: an enhanced lexical resource for sentiment analysis and opinion mining. in: nicoletta calzolari, khalid choukri, bente maegaard, joseph mariani, jan odijk, stelios piperidis, mike rosner & daniel tapias (eds.), proceedings of the 7th international conference on language resources and evaluation (lrec'2010), 17–23 may, malta. european language resources association (elra). douglas biber (1988). variation across speech and writing. cambridge university press, cambridge. doi: 10.1017/cbo9780511621024. anneloes canestrelli (2013). small words, big effects? subjective versus objective causal connectives in discourse processing. lot, utrecht. elisabeth couper-kuhlen (1996). intonation and clause-combining in discourse: the case of because. pragmatics, 6(3):389-426. liesbeth degand (2001). form and function of causation. a theoretical and empirical investigation of causal constructions in dutch. peeters, leuven. liesbeth degand and benjamin fagard (2012). competing connectives in the causal domain: french car and parce que. journal of pragmatics, causal connectives in discourse: a crosslinguistic perspective, 44(2):154-68. liesbeth degand and henk pander maat (2003). a contrastive study of dutch and french causal connectives on the speaker involvement scale. in arie verhagen & j. m. van de weijer (ed.), usage-based approaches to dutch, pp. 175-199. lot, utrecht. christiane fellbaum (ed.) (1998). wordnet: an electronic lexical database. mit press, cambridge, massachusetts. john r. firth (1957). a synopsis of linguistic theory 1930–1955. in: studies in linguistic analysis, pp. 1–32. blackwell, oxford. cecilia e. ford (1994). dialogic aspects of talk and writing: because on the interactive-edited continuum. text interdisciplinary journal for the study of discourse, 14:531-554. frank e. harrell, jr. (2001). regression modeling strategies. with applications to linear models, logistic regression, and survival analysis. springer, new york. frank e. harrell, jr. (2015). rms: regression modeling strategies. r package version 4.4-0. available online at http://cran.r-project.org/package=rms. (last access 01.02.2017). zelig harris (1954). distributional structure. word, 10(23):146–162. michael b. kac (1972). clauses of saying and the interpretation of because. language, 48(3): 626-632. dan klein and christopher d. manning (2003). accurate unlexicalized parsing. in proceedings of the 41th annual meeting of the association for computational linguistics, pages 423-430. available online at http://nlp.stanford.edu/manning/papers/unlexicalized-parsing.pdf (last access 01.02.2017). alistair knott and robert dale (1994). using linguistic phenomena to motivate a set of rhetorical relations. discourse processes, 18(1):35-62. alistair knott and ted j.m. sanders (1998). the classification of coherence relations and their linguistic markers: an exploration of two languages. journal of pragmatics, 30(2):135-175. phillip koehn (2005) europarl: a parallel corpus for statistical machine translation. in proceedings of the 10th machine translation summit, september 12–16, 2005, phuket, thailand. available online at http://www.statmt.org/europarl/ (last access 01.02.2017) ronald w. langacker (1985). observations and speculations on subjectivity. in john haiman (ed.) iconicity in syntax, pages 109-150. john benjamins, amsterdam. http://www.aclweb.org/anthology/d09-1000 http://cran.r-project.org/package=rms http://nlp.stanford.edu/manning/papers/unlexicalized-parsing.pdf http://www.statmt.org/europarl/ just because 149 beth levin (1993). english verb classes and alternations: a preliminary investigation. university of chicago press, chicago. natalia levshina (2015). how to do linguistics with r: data exploration and statistical analysis. john benjamins, amsterdam. william mann and sandra thompson (1988). rhetorical structure theory: toward a functional theory of text organization. text interdisciplinary journal for the study of discourse 8(3): 243–281. henk pander maat and ted j.m. sanders (2001). subjectivity in causal connectives: an empirical study of language in use. cognitive linguistics, 12(3):247-273. ted pedersen (2006). unsupervised corpus-based methods for wsd. in eneko agirre & philip edmonds (eds.), word sense disambiguation: algorithms and applications, pp. 133–166. springer, new york. mirna pit (2003). how to express yourself with a causal connective. subjectivity and causal connectives in dutch, german and french. amsterdam, john benjamins. mirna pit (2006). determining subjectivity in text: the case of backward causal connectives in dutch. discourse processes, 41(2):151-74. mirna pit (2007). cross-linguistic analyses of backward causal connectives in dutch, german and french. languages in contrast, 7:53-82. emily pitler and ani nenkova (2009). using syntax to disambiguate explicit discourse connectives in text. in proceedings of the acl-ijcnlp 2009 conference short papers (aclshort '09), pages 13-16. association for computational linguistics, stroudsburg, pennsylvania. r core team (2015). r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. available online at https://www.r-project.org/. (last access 01.02.2017) william e. rutherford (1970). some observations concerning subordinate clauses in english. language, 46(1):97-115. ted j.m. sanders and wilbert p.m. spooren (2009). causal categories in discourse – converging evidence from language use. in ted j.m. sanders & eve sweetser (eds.) causal categories in discourse and cognition, pages 205–246. mouton de gruyter, berlin. ted j.m. sanders and wilbert p.m. spooren (2013). exceptions to rules: a qualitative analysis of backward causal connectives in dutch naturalistic discourse. text and talk, 33(3):377–398. ted j.m. sanders & wilbert p.m.s. spooren (2015). causality and subjectivity in discourse: the meaning and use of causal connectives in spontaneous conversation, chat interactions and written text. linguistics, 53(1):53-92. ted j.m. sanders and eve sweetser (2009). introduction: causality in language and cognition – what causal connectives and causal verbs reveal about the way we think. in ted sanders & eve sweetser (eds.) causal categories in discourse and cognition, pages 1–18. mouton de gruyter, berlin. ted j.m. sanders, kirsten vis and daan broeder (2012). project notes on the dutch project discan. eighth joint acl iso workshop on interoperable semantic annotation, pisa. mary j. schleppegrell (1991). paratactic because. journal of pragmatics, 16:323-367. wilbert p.m. spooren and liesbeth degand (2010). coding coherence relations: reliability and validity. corpus linguistics and linguistic theory, 6(2):241-66. wilbert p.m. spooren, ted j.m. sanders, mike huiskes and liesbeth degand (2010). subjectivity and causality: a corpus study of spoken language. in john newman & sally rice (eds.) empirical and experimental methods in cognitive/functional research, pages 256-270. csli publications, stanford, california. ninke stukker (2005). causality marking across levels of language structure. a cognitive semantic analysis of causal verbs and causal connectives in dutch. phd diss. universiteit utrecht. https://www.r-project.org/ levshina and degand 150 ninke stukker and ted j.m. sanders (2012). subjectivity and prototype structure in causal connectives: a cross-linguistic perspective. journal of pragmatics, 44(2):169-190. eve sweetser (1990). from etymology to pragmatics. metaphorical and cultural aspects of semantic structure. cambridge university press, cambridge. jörg tiedemann (2012). parallel data, tools and interfaces in opus. in nicoletta calzolari, khalid choukri, thierry declerck, mehmet uğur doğan, bente maegaard, joseph mariani, asuncion moreno, jan odijk & stelios piperidis (eds.), proceedings of the 8th international conference on language resources and evaluation (lrec'2012), istanbul, turkey, pages 2214 – 2218. european language resources association (elra). available online at http://www.lrecconf.org/proceedings/lrec2012/pdf/463_paper.pdf (last access 01.02.17). rena torres cacoullos and scott a. schwenter (2005). towards an operational notion of subjectification. in rebecca t. cover et yuni kim (eds.). proceedings of the 31st annual meeting of the berkeley linguistics society: general session and parasession on prosodic variation and change, pages 347-358. berkeley linguistics society, berkeley, california. teun a. van dijk (1979). pragmatic connectives. journal of pragmatics, 3:447-456. ludwig wittgenstein (1953). philosophical investigations. blackwell, oxford. sandrine zufferey and bruno cartoni (2012). english and french causal connectives in contrast. languages in contrast, 12(2):232–250. http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_paper.pdf http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_paper.pdf dnd_rohlfing_final2011_ms dialogue and discourse 2 (2011) received 9/10; accepted 6/11; doi: 0.5087/dad.2011.201 published online 7/11 ©2011 katharina j. rohlfing exploring "associative talk": when german mothers instruct their two year olds about spatial tasks katharina j. rohlfing kjr@uni-bielefeld.de emergentist semantics bielefeld university, citec universitätsstr. 21–23 bielefeld, germany editor: matthew w. crocker abstract in this study, maternal semantic input was analyzed during a task, in which german mothers instructed their two-year-old children to put two objects together in a particular way. in the setting, the spatial relation (on and under) and the canonicality of these relations (canonical such as ‘an iron on an ironing board’ and noncanonical like ‘a cup under a table’) were varied and related to children’s spatial cognitive skills. two kinds of discourse strategies are proposed that characterize mothers’ semantic input in this task: bring-in and follow-in. for the analysis, an automatic procedure was developed, in which the amount of words spent on a strategy was related to the overall word amount. the data suggest that the canonicality of the task can change the discourse: bring-in strategies dominated the discourse in the under tasks with canonical spatial relations while in the more difficult non-canonical tasks, mothers used follow-ins significantly more often than in the canonical tasks. together, the results of this study shed light on the process of an on-line adaptation of the mother to her child and give us insight into how a situated understanding in a task-oriented discourse emerges. keywords: mother-child discourse, semantic input, understanding of prepositions 1 introduction mothers talk about events to their children. this has been studied extensively in two different areas of research: language acquisition and memory development. one intriguing topic in language acquisition is how mothers structure their discourse, and how children's emerging communication skills are associated with maternal conversational style (for example, hart & risley, 1982; tomasello & todd, 1983). in this area, the research goal is to find out how mothers scaffold the language-learning process. studies in memory development, in turn, examine how parental conversations about events influence the ways in which children store these events. here, the goal is to identify conversational styles that foster developmental changes in memory skills (for example, haden, ornstein, rudek, & cameron, 2009). the aim of the present article is to combine both strings of research when analyzing a task-oriented dialogue. in a task, a mother uses instructions because she wants her child to achieve a goal. however, the way she shapes her instructions takes two aspects into account: the child's attention and the child's knowledge about the objects, their spatial relations and events involved. therefore, the following sections review the conversational strategies from the area of language acquisition for organizing a child's attention and the strategies from the area of memory development on children's acquired knowledge about events. katharina j. rohlfing 2 children's lexicon acquisition and mothers' linguistic and communicative style several studies have proposed that mothers' speech and their communicative styles influence children's word learning (for example, hart & risley, 1982; tomasello & todd, 1983), and it has been shown convincingly that the amount of talking and the richness of lexical variation in caregiver input influences children's lexical and grammar development (for example, hart & risley, 1982; hoff & naigles, 2002; huttenlocher, vasilyeva, cymerman, & levine, 2002; masur, flynn, & eichrost, 2005; tamis-lemonda, bronstein, & baumwell, 2001). some crosslinguistic variations in maternal speech and its relation to children's lexical acquisition have also been reported. for example, choi (2000) has shown that the distribution of nouns and verbs in caregivers' talk differs between englishand korean-speaking mothers, and that these distributions relate to children's vocabulary composition in early years. not only the amount and quality of the caregiver's input but also how the caregiver responds to and incorporates the child's attention is a key element in talking to children (estigarribia & clark, 2007; tomasello & todd, 1983). della corte, benedict, and klein (1983) as well as tomasello and todd (1983) have investigated maternal conversation style in terms of socialinteraction variables. tomasello and todd (1983) measured mother's style of attention regulation when addressing their 12to 13-month-old children, and analyzed how this style correlated with referential-expressive differences in the children's early lexical acquisition. they found that the maternal style—directing attention toward an event the child is not focusing on—was associated with more personal-social words and fewer nominals in the child's lexical development. this style was judged not to be optimal for establishing joint attention (see also tomasello & farrar, 1986), because "now it is the child who must discern where the adult's attention is focused" (tomasello & todd, 1983: 200). in addition, an experimental study by tomasello and farrar (1986) has shown that 17-month-old children learned words better when they were presented within a joint episode; that is, when the child's attention was already focused on the labeled object (see also mundy & gomes, 1998). a further study by akhtar, dunham, and dunham (1991) examined mothers' utterances in terms of prescriptive commands directed toward the child or descriptive statements describing events that either followed or directed the child's focus of attention. for example, a "followprescriptive" occurred when a mother and her child were engaged in building a tower and the mother said "give me the block!" referring to the one the child was holding and looking at (akhtar et al., 1991). whereas maternal speech was examined when the children were 13 months old, the children's lexical development was measured at the age of 22 months. interestingly, only the follow-prescriptives correlated significantly with the child's productive vocabulary at the age of 22 months. hence, the study suggests that giving commands rather than descriptions to 13month-old children in the context of joint focus may be beneficial in the early stages of a child's vocabulary development (akhtar et al., 1991). in sum, the studies presented above have shown that establishing joint attention is a key element in communicative exchanges, and that the way in which mothers align their speech with their child's focus of attention accounts for individual differences in learning. however, lieven (1994) observed that in some cultures, mothers tend to organize their children's attention rather than merely responding to it. thus, joint attention should not be limited to a behavioral strategy. instead, following estigarribia and clark (2007), who studied 40 dyads of adults talking to children at the mean age of 18 to 36 months, establishing joint attention is an interactive process in which it seems to be important for children to understand "why the adult is trying to get their attention" (estigarribia & clark, 2007: 811). as reported here, research in language acquisition has concentrated on joint attention in the form of conversational strategies for organizing a child's attention. however, as i shall show below, research on memory development points to other means of achieving a rather top-down organization of attention. exploring "associative talk" 3 children's memory development and mothers' conversational style the options available to a mother when talking about events have been examined in studies investigating the effects of maternal input on children's memory for events. memory is a crucial cognitive component for early development: it allows children to recall and talk about past experiences (ornstein & haden, 2001). several studies have proposed that language plays an essential role in the development of memory (bauer & wewerka, 1995; boland, haden, & ornstein, 2003; haden et al., 2009; reese, haden, & fivush, 1993; see, for nonverbal recall, mcguigan & salmon, 2006). hayne and herbert (2004) modeled some new actions with objects for 18-month-old children who were not yet fluent speakers. during the demonstration, one group of children received a narration about the event goals and individual target actions. this verbal description gave the full names of the objects and full verbs, for example: "we can use these things to make a rattle. push the ball into the cup . . ." (hayne & herbert, 2004: 131). another group received an "empty narration" containing deictic terms like "let's have a look at this. then we have this bit . . . ." this style of verbal behavior maintained the infant's attention, but contained no additional information about either the target actions or the event goals. after a 4week delay, the infants' memory of the previously demonstrated actions with the objects was tested. results showed that infants who had been given full narrations exhibited superior retention of the events compared to infants given empty narrations (hayne & herbert, 2004). in a study with older children (30–46 months), reese et al. (1993) analyzed how mothers construct events so they can be better memorized. they identified two conversational styles: lowelaborative and high-elaborative. a high-elaborative style (also called "high-eliciting" style in haden et al., 2009: 121) is characterized in terms of (a) eliciting discussions of past events; (b) frequently asking wh-questions; (c) elaborating between what is happening in the here and now and what a child might already know about the event (for example, adding information or associating it with previously experienced events to guide children's memory); (d) encouraging children to talk about aspects of the events that seem to interest them (see also akhtar et al., 1991; tomasello & farrar, 1986; tomasello & todd, 1983); (e) repetitions; and (f) providing positive evaluations of children's responses (boland et al., 2003; haden, ornstein, eckerman, & didow, 2001). a low-elaborative style in contrast, is characterized by mothers talking to their children about practical matters and focusing on the who and what (bauer & wewerka, 1995). boland and colleagues (2003) tested the positive effects of the high-elaborative style on toddlers' ability to remember events. they first trained some mothers to use elaborative conversational styles; other untrained mothers served as controls. trained mothers learned some conversational techniques in order to provide input that focuses on children's attention and, thus, presumably increases their understanding of events. the mothers were then asked to apply these techniques in joint play events with their children, such as fishing or camping. the children were interviewed about the events both one day after the session and after a 3-week interval. they were asked to tell, for example, what they had experienced on the camping trip with their mother and to answer other questions concerning the details of the event. recall of instances in which a component of the event was named were coded, and the percentage of recall was compared across children with trained versus untrained mothers. results suggested that conversational interaction focusing children's attention on salient aspects of an event enhances their understanding of the event. such interaction helps to establish a richly detailed and organized representation of the experience (bauer & wewerka, 1995). hence, mother-child interaction is linked in important ways to what is encoded and subsequently remembered (haden et al., 2009). associative talk: background knowledge about events a closer look at the above shown elaborated style reveals that a bundle of potential factors is involved in the positive effect on memory. the next step is to tease the factors apart and katharina j. rohlfing 4 investigate them separately. the effect of some of the factors is reported in the literature: for example, the positive effects of asking wh-questions (for example, bauer & wewerka, 1995; walsh & blewitt, 2006), of follow-prescriptives of a child's actual behavior (akhtar et al., 1991; dunham & dunham, 1992; rosenthal rollins, 2003; tomasello & farrar, 1986; tomasello & todd, 1983), and of positive evaluations (della corte et al., 1983). however, we know little about the content of these types of input, in particular, about the range of options mothers have for elaborating or associating their input with the knowledge the child might already possess about the event. for example, it has been shown that toddlers know early about putting things into a container and on a surface, but they get to know later in the development how to put things behind or under (casasola, 2008). mothers have to consider the cognitive development of their children when talking as an attempt to solve a task. therefore, the approach presented here aims at exploring what ornstein, haden, and hedrick (2004: 382) referred to in their review as "associative talk." how does association work? the temporal concurrence or consecutiveness of units seems to be crucial for the development of relationships between them. according to strube (1984), two processes can induce a search of memory: one based on common features or feature patterns; the other, on a specific context. features or feature patterns are linked when they occur in close temporal proximity. they can be seen as paradigmatic relationships. the strategy based on a specific context requires a joint perception and action context and retrieves syntagmatic relationships (de saussure, 1931). research on children's event memory has not yet specified these association processes. instead, it is proposed that an event is jointly constructed and structured in a conversation between mother and child (haden et al., 2001). in order to explore the "associative talk", this article focuses on the discursive options that mothers have when trying to scaffold their child's understanding of a given task and to introduce the relevant features of a jointly structured event into the situation. unlike the studies in memory development, however, the aim of the study is not to test what the participant remembers about an event. instead, the goal is to explore the associative talk occurring within a task. this task is a mother requesting her child to produce a particular spatial configuration between two objects. for example, she asks her child to put a cup on a table. such a task might be linked to an event (for example, having a tea party) as long as the child has such an event representation in her or his memory and can apply this knowledge to perform the task. the analysis of associative talk applies to the semantic content of the discourse – that according to masur, flynn, and eichorst (2005: 88) is influential in lexical development because it maintains or motivates their children's attention to events. however, up to now, it has not been a focus of research. some research on semantic content has been carried out in cognitive linguistics. for example, sinha and jensen de lópez (2000) have suggested that the linguistic way of structuring events is based on cultural knowledge about objects (artifacts). sinha (1983: 269) refers to the child's knowledge that can be useful for a task as "background knowledge." its involvement in the construction of meaningful events has been stressed from a developmental perspective by sinha (1983) and from a general perspective on concept formation by barsalou (2002). it can be defined as knowledge that enables a child to make inferences that facilitate the comprehension and establishment of a coherent representation in memory. freeman, lloyd, and sinha (1980) have shown that the background knowledge about the canonical orientation of an object and the canonical relationships between objects impacts on children's performance in a task. for example, children understand an instruction such as "put a block on a cup" poorly when an inverted (that is, noncanonical) orientation of a cup is being requested. canonicality (sinha, 1983), that is, the background knowledge about the functional relationship, influences not only the relationship between two objects but also the way one object is handled. a cup, for example, must be held in an appropriate way (that is, the opening facing upward) to fulfill its function. canonicality is established according to "canonical rules" (sinha, exploring "associative talk" 5 1983: 276) that reflect the social interests and values of a particular society and mediate its cultural practices. in the following study, canonicality was used as an independent variable having an influence on the task difficulty. from previous research, it is known that children at the age of 20 to 26 months have to put a greater effort into a noncanonical goal (clark, 1973; rohlfing, 2005). even though it has been shown that children are sensitive to the semantic principles of spatial categories of their target language from at least 18 months of age (choi et al., 1999), other findings indicate that two-year-olds' proficiency in understanding spatial terms is highly situation and task-dependent (for example, clark, 1973; freeman, lloyd & sinha, 1980; sinha & jensen de lópez, 2000; rohlfing, 2005), and therefore still under strong influence of nonlinguistic factors such as object's function and geometry. these findings on the sensitivity to language-specific spatial semantics on the one hand and the sensitivity to nonlinguistic factors concerning the materiality of objects on the other hand put forward the idea of a continuous and complex interaction among cognition, perception and language in spatial tasks (bowerman & choi, 2001; casasola, 2008). the canonicality in the tasks below will therefore likely affect this interaction. the present study hypothesizes that maternal input will differ as a function of the canonicality of the requested relation. mothers are expected to make more use of children's background knowledge and, thus, relate more often to shared past events when requesting a canonical relationship as opposed to requesting a noncanonical relationship. this is examined by extending analyses of the conversational style (for example, haden et al., 2001) involved in "associative talk" and proposing a more detailed characterization. 2 method participants nineteen german-speaking mother-child pairs participated in this study. the 10 girls and 9 boys were aged 22–24 months (m = 22.6 months, sd = 21 days). at this particular age, children are able to engage in dyadic task-oriented interactions on the one hand, on the other hand, previous research have shown that 22 to 24 month-olds are biased towards some canonical spatial configurations (or some geometric features) and have difficulties in understanding verbal requests for noncanonical goals (for example, clark, 1973; rohlfing, 2005), which gives an opportunity to investigate the impact of task difficulty on the dialogue. participants were selected from a subject pool of interested parents answering an advertisement in the local newspaper of a north german city. all children were being raised in a monolingual german-speaking environment. mothers were not paid for their participation, but their children were given a book as a present. stimuli the stimuli were objects whose relations varied in two ways: in terms of canonicality (canonical vs. noncanonical relation) and the kind of relation (on vs. under). for each task, two relations were prepared with and without an animated trajector. all stimuli are presented in figure 1. a canonical relation refers to the most common function between two particular objects. for an iron and the ironing board, the canonical relation is the iron going on the ironing board. the objects used for a canonical relation in this study are depicted in the top row of figure 1. in contrast to the canonical relation, a noncanonical relation was defined as a relation that is possible and plausible with the objects involved but does not relate to their customary function (for example, a cup under a table). when choosing objects for a canonical or noncanonical relation, it is important to be aware of the fact that canonicality also involves appropriate katharina j. rohlfing 6 orientations of the objects for a relation. the orientation of an object, in turn, can be culturespecific (sinha & jensen de lópez, 2000) and depends on the child's personal experience (rohlfing, rehm, & goeke, 2003). this study considers four noncanonical relations: a spoon on a cup, a rabbit on a hutch, as well as a cup under a table and a horse under a bridge (previous studies had shown that children at the age studied are very likely to put a horse on or across the bridge particularly when a staircase leads to the top of the bridge). the noncanonical relations are depicted in the bottom row of figure 1. iron on the ironing board chair under the awning canonical relations girl on the chair girl under the umbrella spoon on the cup cup under the table noncanonical relations rabbit on the hutch horse under the bridge figure 1: photos shown to the mothers and labels for the different tasks. the relations on and under were chosen because of the different level of knowledge children demonstrate in understanding them: whereas the preposition on is reported to be understood very early, the understanding of under is relatively poor at the age of 20 to 24 months (for example, clark, 1973; rohlfing, 2005). exploring "associative talk" 7 procedure and language survey sessions lasted about half an hour. at the beginning, the experimenter engaged the child in free play at a small table while the mother filled out a language survey containing items on the child's understanding and production of 49 spatial terms. more specifically, in this survey, the mother was asked whether the child understands spatial terms for actions (for example, open, put, hide), relations between objects (for example, in, on, under, to), nouns (for example, front, inside, top), and other deictic terms (for example, here, where). these words were chosen because of their semantic relevance for this experiment. kickert (2008) has shown that the german version of this language survey correlates strongly with the scores on elfra-2 (grimm & doil, 2000), the german adaptation of the macarthur-bates communicative development inventories, cdi. for example, the productive vocabulary in the language survey presented here correlated very significantly with the productive vocabulary in elfra-2 (r = .88, p < .001). after a few minutes, all toys and books were removed from the table. the mother sat at the table next to the child, but at a 90º angle. the experimenter presented pairs of objects one by one to the child. the experimenter told the child that she was going to show some new toys and that they would all play a game together. next, the mother was shown a photo depicting a relation (see figure 1). she then proceeded to instruct her child. the data was transcribed from this time on. the mother was told that the relation on the photo was the target relation, and she was instructed to feel free to use verbal and nonverbal means to get her child to understand and perform the task. however, we asked mothers not to actually perform the relationship on the objects directly, for example, not to put a spoon on a cup. if a child still did not understand the task after several instructions, the experimenter moved on to the next pair of objects. data coding data were transcribed using an xml-format program called martha. this program was specially developed for this study, and its customized structure made the transcription process simple but appropriate for this analysis. mothers' verbal behavior was transcribed and coded in an xml format on four levels: lexicon, sentence reference, discourse, and nonverbal behavior. the present paper reports on the discourse coding. the discourse was quantified by counting all words uttered by the mother in every transcript. this word total then represented 100 % of the discourse. next, the number of words used for a particular discourse strategy (see the category system below) was related to the overall word count. this revealed the percentage involvement of a particular strategy within the overall discourse. this new procedure is more objective, especially in comparison to previous practices such as (a) taking the raw frequencies into account, because with this new method, it is possible to take the interpersonal variability into account, that is, the fact that the dialogues between mothers and their children were of different length (some mothers talked more than others) as no time constraint was given in the task; thus, the proportions seem to give a better picture of the involvement of the different strategies in the discourse than the raw frequencies; (b) first counting sentences—and having to make difficult decisions on whether one-word utterances like "good!" or syntactically incomplete utterances like "this way!" are sentences in child-directed speech— and then calculating the number of discourse strategies per sentence (rohlfing & choi, 2004) or (c) counting discourse strategies as a specific sentence type alongside other types such as explicit instructions (choi & rohlfing, 2010). child's behavior was coded in terms of the task performance: a child was successful if she or he put the objects together in a requested manner; a child was not successful if answered with another relation than requested. for example, in the case that a mother requested a noncanonical relationship between a horse and the bridge (put it under!), and the child performed a canonical relationship, this task performance was coded as not successful. katharina j. rohlfing 8 category system for the discourse the category system was based on a schema developed in rohlfing and choi (2004). it was assumed that a task-directed instruction has the format put x on/under y! and explicitly directs the child to put two things together (choi & rohlfing, 2010). when the instruction diverged from this format, the mothers' verbal behavior was analyzed for its discourse strategy. as the left-hand column in table 1 shows, utterances were assigned either to bring-in or to follow-in strategies. each type of discourse then has subtypes. their descriptions are given in table 1. the right-hand column gives examples of these strategies with the crucial parts for the category assignment highlighted in bold. strategy description prominent example bring-in story about the trajector object a further description of an event related to the trajector object containing the object's function a tea-cup! story about the landmark object a further description of a landmark object is provided containing the object's function this is a beach umbrella! story about the situation the whole event is named, which suggests the involvement of both objects in a target relation let’s have a tea party! comparison the task is compared to a similar situation that the child has already experienced or knows like at home, you put the cup with water always on the table paraphrase the objects' names, the relation or the action are paraphrased to evoke associations with the target relation put the horse under the bridge, put him in the hole! follow-in describing comments on what the child is actually doing yeah, you are going over the bridge! inhibiting negative comments on what the child is actually doing; sometimes precedes contrast now, don't stir! there is nothing inside. indirection the child's actual behavior is lead to a point, from which a task-directed instruction is given; it may concern the orientation or the role of an involved object or the child’s attention turn the bridge over! show mummy the spoon! contrast the target relation is contrasted with what the child is currently doing [in this example, the first sentence is coded labeling while the second is coded as contrast to the first one that’s on the bridge. put the horse under the bridge! noun uptake mother expands on child's suggestion; if the child labels an object, the mother picks up this noun and uses it for the task instruction child: steps! mother: can you put the horsey under the steps? table 1: coding system for the discourse. five types of the bring-in strategy were identified. they comprise what ornstein et al. (2004) term "associative talk" or what haden et al. (2009: 121) refer to as "statement elaborations" and explain as "declarative comments that provided new information about the event, but did not call for the child to respond (for example, 'we saw lots of dinosaurs at the museum.')." in all of the bring-in strategies, the mother introduced a particular event frame to the dialog (for example, "tea party") that was familiar to the child. for example, when a mother asked "will you have some exploring "associative talk" 9 tea?" for the task of putting a cup on the table, the utterance was coded as a bring-in discourse strategy, because the mother was evoking a particular event. evoking the event may help the child to understand the requested spatial relation. a bring-in can be achieved by talking about past events directly in terms of telling a story about the trajector ("storytr" strategy), the landmark ("storylm"), or the situation ("storysit"). these stories provide an elaboration of the function of the objects, thus allocating an object in context. a fourth type of bring-in strategy is comparison in which the mother compares the current situation with a specific, personally experienced event. finally, the fifth type of bring-in is paraphrase in which the mother paraphrases a preceding utterance that may be difficult for the child to understand by introducing another notion that is more familiar. for example, in the task in which the horse had to be put under the bridge, the preposition under was replaced by the preposition in ("in the middle" or "in the hole"). in the example given in table 1, a mother replaced the preposition under with the preposition in and said "in the hole," evoking the hollow space under the bridge as a containment, so that the child could perceive the target location and follow the task better. the follow-in strategies were defined on the sole basis of maternal discourse (and not through nonlinguistic aspects such as the child's eye gaze as in tomasello & farrar, 1986) as utterances that refer to children's situated attention on the physical and spatial properties of the objects themselves. strategies identified as describing, indirection, contrast, and noun uptake use language as a tool to directly instruct the desired event on the basis of the action or the object the child is attending to. for example, when a mother said "you are going over the bridge," she was describing what the child was doing and therefore this utterance was assigned as a follow-in (describing). the strategy inhibiting characterized cases in which a mother gave negative comments on what the child was actually doing such as "now, don't move!" transcript analyzed as 01 c: [manipulates] 02 m: guck mal, das soll ein sonnendach sein. look, this should be a sunroof 03 m: da damit die sonne da nicht hinkommt. so the sun does not come here bring-in: story about the situation 04 m: und der sonnenstuhl kann jetzt u:nter dem dach stehen. and the sun chair can now stand under the roof bring-in: story about the trajector 05 c: [manipulates] 06 m: den stuhl the chair… 07 m: ja:? so liegt der stuhl ja. yes? now lays the chair. follow-in: describing 08 c: [manipulates] 09 m: kannst du den auch hinstellen? wie bei uns auf der terasse. can you also put it upright? like at our home at the awning? follow-in: indirection bring-in: comparison 10 m: ja genau. yes, exactly. table 2: example of transcript analysis: words spent on strategies are marked in bold katharina j. rohlfing 10 however, some overlaps of strategy categories were possible within this system. in one case, a mother said with reference to the horse: "und jetzt läuft's ja oben drüber. kann das auch unten durchkrabbeln? [now, it goes on top. can it crawl under as well?]" in the latter utterance, one word was coded as a bring-in strategy, because the mother paraphrased the relation "under the bridge" by the action of crawling. however, this utterance also implied a follow-in strategy, because the mother contrasted the requested relation with the child's actual doing. the xmlbased analysis tool took account of the involvement of both strategies in the overall discourse. because of the overlap, the two strategies are considered as two dependent variables and are subjected to two separate statistical tests. as stated in the data coding section, the total number of words the mother used in a task was calculated first. then the number of words used for a particular strategy was calculated as percentage of a particular strategy involved in the whole discourse. the transcript in table 2 provides an example. 3 results children's task performance children performed an average of six tasks correctly (from a minimum of four and a maximum of eight, sd = 1.6). an inspection of the percentage of children performing the requested relation (see table 3) revealed that the tasks varied in level of difficulty. table 3 reports both the percentage of successes and the median number of instructions needed for the child to succeed. it indicates that canonical relationships were easier to perform (with canonical on easier than under) than noncanonical ones. in addition, a higher percentage of both bring-in and follow-in strategies combined was found in noncanonical settings: on average, 18.7 % of the mothers' discourse took the form of strategies in noncanonical tasks compared to 14.9 % in canonical tasks. relation percentage of successful performance median of instructions needed canonical iron / ironing board girl / chair 89 100 7 3 on noncanonical cup / spoon rabbit / hutch 42 79 17 13 canonical chair / awning girl / umbrella 74 89 13 8 under noncanonical pot / table horse / bridge 68 68 12 15 table 3: children's task performance and the number of instructions. exploring "associative talk" 11 task performance, children's age, and children’s spatial lexicon since the children in the sample had an age span of 3 months (from 22 to 24 months), it was possible that their performance was dependent on their age. however, no evidence was found for this assumption (children's age did not correlate significantly with children's performance: r = .00, df = 19, ns). alpha was set at .05 for all statistical tests. the next step was to examine possible correlations with the children's abilities in the spatial lexicon (as reported by mothers in the language survey). children's task performance correlated significantly with reported productive spatial lexicon (r = .52, df = 19, p < .03) but not with the reported receptive spatial lexicon (r = .28), suggesting that a more advanced reported productive spatial lexicon helped children to solve the task successfully. the reported production of spatial terms related particularly to performance on noncanonical tasks (r = .48, df = 19, p < .04). thus, children's productive language capabilities seemed to be involved in their task performance, which, in turn, influenced the length and quality of the discourse that the mother provided to the child. this indirect influence was examined with an analysis of covariance (s. below) taking the reported productive lexicon as covariate. discourse strategies and canonicality overall, the total number of words provided by the mothers to their children across all eight tasks was 10039. the average number of words from a mother to her child per task was 528 (sd = 199), with 19.2 % of this discourse being provided in the form of bring-in strategies and 14.4 % in the form of follow-in strategies. the greater presence of bring-ins was statistically significant, t(18) = 2.37, p < .03. this indicated the relevance of bring-in strategies as an important component of the discourse. since bring-in and follow-in strategies can overlap (s. section on data coding and category system), two separate 2 x 2 ancovas were performed on the percentage of words constituting bring-in and follow-in discourse strategies with spatial relation (on and under) and canonicality of the relation (canonical and noncanonical) as within subject factors and the reported productive spatial lexicon of the children as covariate. as shown in figure 2, the frequency of bring-in strategies was higher in the canonical settings than in the noncanonical settings, but only for the under tasks. the statistical analysis failed to attain any significance. one explanation seems evident when looking at the data in table 3: in the canonical on tasks, children displayed ceiling effects in their performance and mothers did not need to instruct a lot. this is in line with previous research (clark, 1973; rohlfing, 2001). since in this task, little discourse was needed anyway, the comparison with the noncanonical task seems to be futile. a paired t test performed only on the under tasks, revealed a statistical difference, t(18) = 1.82, p = .03 (one-tailed), according to which more bring-ins were produced in canonical tasks than in noncanonical tasks verifying the one-directional hypothesis raised for the analysis: mothers were expected to make more use of children's background knowledge and, thus, bring-in more often shared past events when requesting a canonical relationship as opposed to requesting a noncanonical relationship. as shown in figure 2, the frequency of follow-in strategies was significantly higher in the noncanonical settings than in the canonical settings for both relations, which was confirmed by the statistical analysis revealing a main effect for canonicality f(1,17) = 9.18, p < 0.01, eta2 = 0.35. this main effect was further investigated by means of post hoc pairwise bonferronicorrected (.05/2) t tests indicating a difference in canonicality between canonical vs. noncanonical on tasks t(18) = -3.47, df = 18, p = 0.003 and canonical vs. noncanonical under tasks t(18) = 3.11, df = 18, p = 0.006. in addition, the ancova analysis revealed also a main effect for relation f(1,17) = 3.17, p < 0.01, eta2 = 0.16. accordingly, mothers made use of more follow-ins when instructing for under relations rather than for on. together, these results match the katharina j. rohlfing 12 analysis of the task difficulty (s. results on children's task performance) suggesting that the more difficult the task was, the more use of follow-ins the mothers made. figure 2: percentage of (above:) bring-ins and (below:) follow-ins involved in the discourse with regard to canonicality of the on and under spatial relation. while follow-ins are used more often in noncanonical tasks of both spatial relations, the dominance of the bring-ins in canonical relations can be seen only for the more difficult under spatial relation. in sum, the data provides support for the hypothesis that the type of maternal discourse could change as a function of the canonicality of a spatial relationship: generally, in canonical and noncanonial conditions, mothers drew on their children's background knowledge such as familiar events related to the spatial task. when more discourse is needed (which is the case in the more difficult under tasks), mothers more often brought in shared past events when requesting a canonical relationship as opposed to requesting a noncanonical relationship. when instructing for a noncanonical relation, in turn, mothers followed-in more and thus focused more directly on the spatial task itself. recall that the ancovas were conducted to examine whether the strategies mothers used in their discourse varied as a function of the children's language production. however, the outcome did not support the hypothesis that children with a less advanced spatial lexicon received a different type of discourse input than those with a more advanced one. that is, the use of bring-in or follow-in strategies did not relate to the children's level of spatial lexicon. this suggests that exploring "associative talk" 13 strategies were task-dependent and selected on-line as a function of the canonicality (familiarity) of spatial relationships. the question remains whether the choice of a particular strategy led to successful performance by the child. however, no statistically relevant relation between mothers' use of particular discourse strategies and children's performance was found. there was only a marginally negative correlation between children's performance and follow-in discourse strategies in noncanonical settings, indicating that children who were not successful in the noncanonical tasks were followed-in more by their mothers. this lack of a significant relation suggests that even though the organization of discourse seems task-dependent, particular strategies do not necessarily help children to solve the task. subcategories of discourse strategies the subcategories of each type of discourse strategy (see table 1) were also analyzed, starting with the frequency of each subtype in canonical versus noncanonical settings (see table 4). the huge individual differences become apparent when looking at the standard deviations. discourse strategy canonical noncanonical comparison m sd m sd t(18) bring-in story tr story lm story sit comparison paraphrase 7.9 1.3 12.8 12.3 3.3 10.7 3.2 14.4 10.4 15.6 15.2 1.5 6.7 6.5 18.2 10.3 2.4 8.1 6.8 13.1 -2.0* -0.2 1.3 2.1* -6.0*** follow-in describing inhibiting indirection contrast noun uptake 3.9 2.1 16.1 2.6 0.4 5.6 4.5 16.4 4.9 0.9 29.2 4.0 18.6 22.1 2.6 21.3 4.7 19.8 15.9 4.86 -5.2*** -1.4 -0.4 -5.1*** -1.9 table 4: the percentage involvement of each of the particular strategies type in the overall discourse; the right column reports the results of the comparison between canonical and noncanonical tasks: the t-values and their statistical significance (*p < 0.5, **p < 0.1, ***p < .001). table 4 shows the distribution of the subtypes of the bring-in strategy in canonical and noncanonical settings. altogether, significantly more comparisons were used in canonical settings, t(18) = 2.05, p = .05, with mothers relating these settings to known past events or known situations. fewer such comparisons were possible in noncanonical settings in which the task seemed to deviate from known situations. according to further paired t tests, mothers invented more stories about the trajector object in noncanonical settings, t(18) = 2.05, p = .05, by saying, for example, that the rabbit wanted to see the sun or to escape. in addition, significantly more katharina j. rohlfing 14 paraphrasing was used in the noncanonical settings, t(18) = -6.02, p < .001. for example, in the task of putting the rabbit on the hutch, the hutch was often paraphrased as dach [roof] to evoke the function of an on. the subtype paraphrase included attempts to use another noun not only for the reference objects such as the hutch or the bridge, but also for the spatial relations such as under. this relation was paraphrased by verbs such as "to hide" as in und jetzt möcht sich das pferdchen unter der leiter verstecken [and now the horsey would like to hide under the ladder] or "to crawl" as in kann das auch unten durchkrabbeln? [can the horse crawl under as well]. as these examples show, it was especially the animate trajector objects that were often personalized by providing modal verb forms like "want" or "can" with the effect of strengthening the story character. further findings on the subtype paraphrase revealed that the preposition under was also paraphrased by using other prepositions such as through (durch die treppe durch [through the stairs]), or the preposition on was paraphrased by crossways (quer). the frequent use of paraphrase for noncanonical relationships suggests that mothers were aware that their children had difficulties in understanding the under relation (see fernyhough, 1996) and tried to use other similar terms with which the children might be more familiar. here, the association seemed to be achieved by the feature or feature pattern resemblance. hence, more paradigmatic relationships were required. table 4 shows also the subtypes of the follow-in strategy in canonical versus noncanonical settings. two differences between the two types of conditions emerged and attained statistical significance (paired t tests): these were describing, t(18) = -5.2, p < .001, and contrast, t(18) = 5.14, p < .001. other differences did not attain a statistical significance. since in noncanonical settings, children spent a lot of time performing the canonical relation (for example, putting the rabbit in the hutch or putting the horse over the bridge), mothers seemed to guide their children to the pursued task by first describing what they were doing and then contrasting that with what they should have been focusing on. this provided linguistic contrasts, as can be seen in the following examples: a mother to her 22 months old son in the bridge task: so geht das pferd die treppe runter und jetzt möchte das pferd unter der treppe hergehen. [the horse walks down the stairs and now it wants to go under the stairs] a mother to her 23 months old daughter in the hutch task: jetzt haste den in den stall getan, ne? stell’ den hasen doch mal aufs dach! [now you put it in the hutch, right? put the rabbit on the roof!] a mother to her 23 months old daughter in the table task: eine tasse. die kommt auf den tisch eigentlich, ne? kannst du die denn auch mal u:nter den tisch stellen? [a cup. usually, it goes on a table, right? can you put it also under the table?] a mother to her 22 months old daughter in the cup task: nein, nicht hinein! auf die tasse drauf legen. [no, not inside! put it on top of the cup!] a mother to her 22 months old son in the hutch task: nicht in den stall hinein, sondern o:ben drauf! [not inside the hutch, but on top of it!] exploring "associative talk" 15 4 discussion this study has been guided by the question of how mothers—being sensitive to their 2-year-old children's cognitive biases towards some spatial configurations between objects discernable in poor understanding of some requests for spatial relations—provide their young interlocutors with alternative perspectives on the shared situation. i propose that in their discursive behavior, mothers have at least two options at their disposal: on the one hand, they can bring-in past events; on the other hand, they can follow-in from what the child is actually doing. up to now, much research has been devoted to investigating follow-in strategies and in this study, their semantics has been analyzed by focusing what is being said to the child within this joint focus. regarding the fact that more follow-ins were used in the noncanonical tasks, an important question is whether in this case, follow-ins reflect a strategy at all. an alternative explanation is that mothers just spend a lot of time telling their children not to put the objects in a more obvious or preferred relation. however, mothers do a lot more than just preventing their children from the canonical relation. to specify the different forms of follow-ins in terms of their semantics is an extension to previous research (for example, carpenter, nagell & tomasello, 1998) that considers this behavior on an attention level only. more importantly however, this exploration study shows that bring-in strategies constitute an even greater part of the overall discourse with children and need to be studied further. which strategy will be used depends on the given situation: in canonical settings, children's background knowledge about events is brought-in and predominantly restricts the discourse. in contrast, children tend to be followed-in in the noncanonical tasks. in these noncanonical tasks, mothers exert lot of effort to redirect their children's attention away from the spatial relation (or the object's feature) to which the children seem biased. why might different types of discourse strategies be more beneficial for one task or another? according to the present data, past experiences can be brought-in by evoking a familiar "story" about a particular object or situation. a given situation like a tea party or breakfast can also be associated with past experiences in which the same type of object has been handled in a particular way, and by analogy, children can perform the task. thus, in a canonical situation, it seems sufficient for the mothers to introduce stories/situations, and the children do not need explicit spatial expressions to perform the task. the specification of the bring-in strategies is a clear development of what ornstein et al. (2004) have called "associative talk," because the identified subcategories point to different association methods, that is, ways of specifying what is being associated with what and by which stylistic options in a specific task. using these strategies, mothers guide their children to relate the present event to past experiences and event knowledge; in this way, bring-in strategies are wellsituated in canonical tasks. they enhance the child's comprehension of the task and align the mother and the child's memory. follow-in strategies work differently. they primarily address children's present attention and manipulative behavior rather than their background knowledge. the finding that mothers make an effort to redirect their children's attention is in line with research on coordinated joint attention (carpenter, nagell, & tomasello, 1998; dunham & dunham, 1992; hollich, hirsh-pasek, & golinkoff, 2000), in which the positive effect of follow-in linguistic input during a child's action is already well-documented. vocabulary acquisition is facilitated when, during interactions, caregivers describe aspects of the infant's current focus of attention (dunham & dunham, 1992: 414; rosenthal rollins, 2003). the research presented here adds to these results: even though there is a stronger presence of this type of discursive behavior in situations in which children develop comprehension problems, follow-in strategies pursue the same goal as the bring-in strategies; both are used to establish a shared view of the situation by providing verbal katharina j. rohlfing 16 information (see fernyhough, 1996). however, the follow-in strategies focus linguistically on the spatial actions themselves. the potential smoothness of the transition from bring-in to follow-in strategies in the discourse can be seen in the strategy named contrast (see examples 1–4 above). when following this strategy, mothers first follow-in and label an actual event, for example, "that's over the bridge!" and then immediately suggested an alternative event, for example, "can you put the horsey under?" the alternative event for the latter example can then be paraphrased in better known words (for example, "hide the horsey!," that can evoke the appropriate knowledge (about what one does when hiding) and lead to the correct response. this interaction between the two types of discourse strategy is characteristic for establishing a shared view on the situation (known also as grounding process) in this particular task: the competent speaker interweaves background event knowledge while describing the child's actual action. concerning the question which strategies lead to a successful task solution, the data presented here provide little evidence that the children's performance relates to their mothers' conversational style. there are no direct correlations between discourse strategies and children's success in solving the tasks. there is also no relation between children's reported lexical competence and mothers' use of strategies. together, these results suggest that strategies are chosen on-line as problem-solving alternatives and are not necessarily related to the level of language acquisition skills. however, the next step in pursuing these correlations more directly will be to conduct a study in which mothers are trained to use specific strategies. with regard to different discoursive styles, choi and rohlfing (2010) have recently looked at cross-cultural differences and compared the discourse style of north-american mothers to that of korean mothers. they have reported that korean mothers make more effort in general to evoke background knowledge in conversations with their children; korean mothers refer more to their children's knowledge about objects and past events than north american mothers do. interestingly, this difference does not relate to children's vocabulary. one aspect that might contribute to more background knowledge being provided by korean mothers is the fact that spatial relations are expressed by means of verbs. verbs, more than nouns, promote expressions of events. these crosslinguistic differences may reflect cultural differences in mother-child interaction in general and the use of associative talk in particular. this is also a topic for future research. acknowledgments the research reported in this paper was made possible by a dilthey fellowship (research initiative focus on the humanities) from the volkswagen foundation. i would like to thank lars schillingmann for creating a new analysis tool for discourse. i am grateful to soonja choi and karla mcgregor for support and to three anonymous reviewers for their comment on an earlier version of this paper. many thanks also to members of the dialoglab for their help with the study and to all the mothers and their children who participated in it. references akhtar, n. dunham, f., & dunham, p. (1991). directive interactions and early vocabulary development: the role of joint attentional focus. journal of child language 18: 41–49. barsalou, l. w. (2002). being there conceptually: simulating categories in preparation for situated action. representation, memory and development. essays in honor of jean mandler, eds. n. l. stein, p. j. bauer, m. rabinowitz, 1–15. mahwah, new jersey/london: lawrence erlbaum associates, publishers. bauer, p. j., & wewerka, s. s. (1995). oneto two-year-olds' recall of events: the more expressed, the more impressed. journal of experimental child psychology 59: 475–496. exploring "associative talk" 17 boland, a. m., haden, c. a., & ornstein, p. a. (2003). boosting children's memory by training mothers in the use of an elaborative conversational style as an event unfoldes. journal of cognition and development 4: 39–65. bowerman, m., & choi, s. (2001). shaping meanings for language: universal and languagespecific in the acquisition of spatial semantic categories. language acquisition and conceptual development, eds. m. bowerman & s. c. levinson, 475–511. cambridge: cambridge university press. casasola, m. (2008). the development of infants' spatial categories. current directions in psychological science 17: 21–25. carpenter, m., nagell, k., & tomasello, m. (1998). social cognition, joint attention, and communicative competence from 9 to 15 months of age. monographs of the society for research in child development, 63 (4, serial no. 255). choi, s. (2000). caregiver input in english and korean: use of nouns and verbs in book-reading and toy-play contexts. journal of child language 27: 69–96. choi, s., mcdonough, l., bowerman, m. & mandler, j. m. (1999). early sensitivity to languagespecific spatial categories in english and korean. cognitive development 14: 242–268. choi, s., & rohlfing, k. j. (2010). discourse and lexical patterns in mothers' speech during spatial tasks: what role do spatial words play? japanese / korean linguistics, volume 17, eds. s. iwasaki, h. hoji, p. m. clancy & s.-o. sohn, 117–133. standford: csli publications. clark, e. v. (1973). non-linguistic strategies and the acquisition of word meanings. cognition 3: 161–82. della corte, m., benedict, h., & klein, d. (1983). the relationship of pragmatic dimensions of mothers' speech to the referential-expressive distinction. journal of child language 10: 35– 43. de saussure, f. (1931). grundfragen der allgemeinen sprachwissenschaft. berlin: walter de gruyter. dunham, p., & dunham, f. (1992). lexical development during middle infancy: a mutually driven infant-caregiver process. developmental psychology 18: 414–420. estigarribia, b., & clark, e. v. (2007). getting and maintaining attention in talk to young children. journal of child language 34: 799–814. fernyhough, c. (1996). the dialogic mind: a dialogic approach to the higher mental functions. new ideas in psychology 14: 47–62. freeman, n. h., lloyd, s., & sinha, c. g. (1980). infant search tasks reveal early concepts of containment and canonical usage of objects. cognition 8: 243–263. grimm, h., & doil, h. (2000). elternfragebogen für die früherkennung von risikokindern (elfra-2). göttingen: hogrefe. haden, c. a., ornstein, p. a., eckerman, c. o., & didow, s. m. (2001). mother-child conversational interactions as events unfold: linkages to subsequent remembering. child development 72: 1016–1031. haden, c. a., ornstein, p. a., rudek, d. j., & cameron, d. (2009). reminiscing in the early years: patterns of maternal elaborativeness and children's remembering. international journal of behavioral development 33: 118–130. hart, b. m., & risley, t. r. (1982). how to use incidental teaching for elaborating language. lawrence, ks: h & h enterprises. hayne, h., & herbert, j. (2004). verbal cues facilitate memory retrieval during infancy. journal of experimental child psychology 89: 127–139. hollich, g., hirsh-pasek, k., & golinkoff, r. (2000). breaking the language barrier: an emergentist coalition model of word learning. monographs of the society for research in child development, 65 (3, serial no. 262). hoff, e., & naigles, l. (2002). how children use input to acquire a lexicon. child development 73: 418–433. huttenlocher, j. vasilyeva, m., cymerman, e., & levine, s. (2002). language input and child syntax. cognitive psychology 45: 337–374. katharina j. rohlfing 18 kickert, k. (2008). unterstützen des präpositionserwerbs bei 2-jährigen kindern — eine interventionsstudie [supporting the acquisition of prepositions in 2-year-old children: an intervention study]. unpublished diploma thesis, department of psychology, bielefeld university, germany. lieven, e. v. (1994). crosslinguistic and crosscultural aspects of language addressed to children. input and interaction, eds. c. gallaway, & b. j. brichards, 56–73. cambridge: cup. masur, e. f., flynn, v., & eichorst, d. l. (2005). maternal responsive and directive behaviours and utterances as predictors of children's lexical development. journal of child language 32: 63–91. mcguigan, f., & salmon, k. (2006). the influence of talking on showing and telling: adult-child talk and children’s verbal and nonverbal event recall. applied cognitive psychology 20: 365– 381. mundy, p., & gomes, a. (1998). individual differences in joint attention skill development in the second year. infant behavior and development 21: 469–482. ornstein, p. a., & haden, c. a. (2001). memory development or the development of memory. current directions in psychological science 10: 202–205. ornstein, p. a., haden , c. a., & hedrick, a. m. (2004). learning to remember: socialcommunicative 3 exchanges and the development 4 of children’s memory skills. developmental review 24: 374–395. reese, e., haden, c. a., & fivush, r. (1993). mother-child conversations about the past: relationships of style and memory over time. cognitive development 8: 403–430. rohlfing, k. j. (2001). no preposition required. the role of prepositions for the understanding of spatial relations in language acquisition, applied cognitive linguistics i: theory and language acquisition, eds. m. pütz, s. niemeier & r. dirven, 230–247. berlin: mouton de gruyter. rohlfing, k. j. (2005). learning prepositions. perspectives on language learning and education 12: 13–17. rohlfing, k. j., rehm, m., & goecke, k.u. (2003). situatedness: the interplay between context(s) and situation. journal of cognition and culture 3: 132–157. rohlfing, k. j., & choi, s. (2004). getting you to understand. how mothers instruct their children to put two things together. paper presented at the iada (international association for dialogue analysis) conference, chicago, 30. march – 3. april. rosenthal rollins, p. (2003). caregivers’ contingent comment to 9-month-old infants: relationships with later language. applied psycholinguistics 24: 221–234. saussure, f. de (1931). grundfragen der allgemeinen sprachwissenschaft [course in general linguistics] (h. lommel, trans.). berlin: walter de gruyter. sinha, c. (1983). background knowledge, presupposition and canonicality. concept development and the development of word meaning, eds. t. seiler & w. wannenmacher, 269–296. berlin: springer. sinha, c., & jensen de lópez, k. (2000). language, culture and the embodiment of spatial cognition. cognitive linguistics 11: 17–41. strube, g. (1984). assoziation. der prozeß des erinnerns und die struktur des gedächtnisses. berlin, heidelberg, new york, tokyo: springer-verlag. tamis-lemonda, c. s., bornstein, m. h., & baumwell, l. (2001). maternal responsiveness and children’s achievement of language milestones. child development 72: 748–767. tomasello, m., & farrar, m. j. (1986). joint attention and early language. child development 57: 1454–1463. tomasello, m., & todd, j. (1983). joint attention and lexical acquisition style. first language 4: 197–212. walsh, b. a., & blewitt, p. (2006). the effect of questioning style during storybook reading on novel vocabulary acquisition of preschoolers. early childhood education journal 33: 273–278. dialogue & discourse 8(1) (2017) 66–105 doi: 10.5087/dad.2017.103 a corpus-driven approach to discourse organisation: from cues to complex markers marie-paule péry-woodley pery@univ-tlse2.fr clle, université de toulouse cnrs, ut2j, france lydia-mai ho-dac hodac@univ-tlse2.fr clle, université de toulouse cnrs, ut2j, france josette rebeyrolle rebeyrol@univ-tlse2.fr clle, université de toulouse cnrs, ut2j, france ludovic tanguy tanguy@univ-tlse2.fr clle, université de toulouse cnrs, ut2j, france cécile fabre cecile.fabre@univ-tlse2.fr clle, université de toulouse cnrs, ut2j, france editor: maite taboada submitted 03/2016; accepted 12/2016; published online 01/2017 abstract this paper reports on an experiment implementing a data-intensive approach to discourse organisation. its focus is on enumerative structures envisaged as a type of textual pattern in a sequentiality-oriented approach to discourse. on the basis of a large-scale annotation exercise calling upon automatic feature mark-up alongside manual annotation, we explore a method to identify complex discourse markers seen as configurations of cues. the presentation of the background to what is termed “multi-level annotation” is organised around four issues: linearity, complexity of discourse markers, top-down processing, granularity and the multi-level nature of discourse structures. in this context, enumerative structures seem to deserve scrutiny for a number of reasons: they are frequent structures appearing at different granularity levels, they are signalled by a variety of devices appearing to work together in complex ways, and they combine a textual role (discourse organisation) with an ideational role (categorisation). we describe the annotation procedure and experimental framework which resulted in nearly 1,000 enumerative structures being annotated in a diversified corpus of over 600,000 words. the results of two approaches to the rich data produced are then presented: firstly, a descriptive survey highlights considerable variation in length and composition, while showing enumerative structure to be a basic strategy resorted to in all three sub-corpora, and leads to a granularity-based typology of the annotated structures; secondly, recurrent cue configurations—-our “complex markers”—-are identified by the application of data mining methods. the paper ends with perspectives for further exploitation of the data, in particular with respect to the semantic characterisation of enumerative structures. keywords: discourse structures, discourse markers, corpus linguistics, corpus annotation, data mining c©2017 marie-paule péry-woodley, lydia-mai ho-dac, josette rebeyrolle, ludovic tanguy, cécile fabre this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). a corpus-driven approach to discourse organisation: from cues to complex markers 1. introduction texts can be seen as the result of squeezing complex hierarchical structures into a largely linear format. understanding a text entails constructing a representation of the underlying structures. a major challenge in the study of written discourse is to identify the signals which guide readers in the process of constructing this representation. depending on one’s theoretical underpinning and focus, signals may be seen as discourse construction devices, as metadiscourse, as reading or processing instructions, as traces of the writers cognitive processes, or as cues revealing the authors intentions. the study presented in this paper sets up a data-intensive methodology whereby signals “emerge” from the systematic analysis of a large set of annotated structures. its aim is the empirical characterisation of configurations of cues signalling a particular discourse pattern: enumerative structures. as these structures can concern text spans of any size, the perspective is described as multi-level. the study relies on the systematic annotation of structures in a corpus of french language texts, and on the application of data mining methods to detect emergent complex discourse markers. while in terms of methodology it belongs in corpus linguistics and natural language processing, its theoretical foundations are to be found in functional linguistics, in psycholinguistics and in research on the visual dimension of texts. we chose to start from what may be seen as the most basic among the notions called upon to account for text/discourse organisation: linearisation, continuity vs. discontinuity (the fundamental question behind discourse segmentation), and discourse patterns. the arguments for this “back to basics” approach are given in the next section, organised around four issues: the linearity constraint, the non-discrete nature of discourse markers, the importance of top-down processing, granularity and the multi-level nature of discourse structures. these constitute the foundation for the choice of enumerative structures for annotation, the rationale for which is given in section 3, followed in section 4 by the annotation model and method, from corpus preparation procedures to the manual annotation of structures and cues. in section 5, a descriptive survey of nearly 1,000 annotated structures leads to a proposal for a granularity-based typology, and to an analysis of genre-related variations. finally, section 6 presents the recurrent cue configurations made apparent by the application of data mining techniques to the rich annotated data. 2. multi-level annotation in the annodis project: preliminaries the research presented here started with the annodis annotation project, which can be described as a large-scale discourse-level annotation experiment calling upon different discourse models and different genres of written french language texts1. the project comprised two distinct approaches, respectively labelled bottom-up and multi-level. bottom-up and multi-level annotations were applied to different corpora, for reasons which are explained in 4.2.2, but a set of texts was annotated in both frameworks to allow a direct comparison of the approaches (annodis duo, see 5.3). the bottom-up annotation, conducted according to segmented discourse representation theory (asher, 1993), focused on the identification of rhetorical relations. the multi-level annotation—-the main focus of this paper—-took a less well-chartered path, which the present section aims to describe and justify in the form of basic propositions underlying the choice of objects to annotate, the annotation method, as well as the questions asked of the annotated corpus. 1. project funded by the humanities and social sciences programme of the french national research agency (anr appel corpus 2007) (see péry-woodley et al., 2011; afantenos et al., 2012, for complete descriptions). the annodis resource is available under a creative commons licence http://redac.univ-tlse2.fr/corpus/ annodis. 67 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre 2.1 linearisation is a problem language is linear, while mental representations are not (or not necessarily). this, as many authors have pointed out (levelt, 1981; gernsbacher, 1995, 1997; heurley, 1997, inter alia), can be seen as problematic insofar as “in text, a multidimensional discourse model is squeezed into a linear form. linearity requires the writer to produce each textual unit in turn, and processing constraints demand short units of meaning. yet the mental representation on which the discourse is based is not a succession of facts or ideas which can each be expressed in one sentence. this is where discourse organisation comes in...” (ho-dac and péry-woodley (2009), echoing levelt (1981)). few approaches to expository discourse, however, focus on linearisation and its inescapable consequence—-sequentiality, i.e. the segmentation of discourse into subsequent text spans. goutsos (1996, 1997) is one author who argues for a theory of sequential relations, in which he sees “an autonomous source of text connectivity” (ibid. 502). he describes most approaches as favouring a what-perspective—-“what is taking place in discourse” (“propositional or semantic content—-over a how-perspective—-“which would focus on the structuring rather than the individual units of text” (goutsos, 1996, p. 503). the macrostructure or story grammar approach to discourse coherence (van dijk, 1980) is an example of the what-perspective, as are, largely, models relying on the notion of rhetorical relations (rhetorical structure theory (mann and thompson, 1988), segmented discourse representation theory (asher, 1993), inter alia). goutsos’ howperspective has its roots in functional linguistics, in the notion of “information packaging” (chafe, 1976, 1994), in halliday’s textual metafunction (halliday, 1977/1983), in the notion of textual strategy (enkvist, 1985; virtanen, 1992). a number of researchers in the field of automatic text generation, especially in the rst “sphere”, have developed models based on a distinction related to the one goutsos proposes: in particular virbel, via his text architecture model based on the notion of “textual object” (virbel, 1989, 2015; lemarié et al., 2008), and power et al., who argue that what they call “abstract document structure” is a separate descriptive level in the analysis and generation of written texts (power et al., 2003). a distinction does exist within rst between subject-matter and presentational relations, a distinction which taboada and mann (2006) associate with goutsos’ proposals [p. 443]. but power et al. go further by questioning the ambiguity in rst between text spans and the meanings of these spans in the attribution of relations, and call for a clear distinction between document structure and rhetorical structure, which they claim are as distinct as syntax from semantics (power et al., 2003, p. 245). these authors have in common a central concern with text segmentation, and how it is signalled (i.e. with what signalling devices). they do not primarily focus on the nature of relations between text segments—-the central concern for models of discourse organisation based on discourse relations. in goutsos’ model, the two fundamental relations between text spans are simply continuity and discontinuity (or shift), text being seen as a “periodic alternation of transition and continuation spans” (goutsos, 1996, p. 501). continuity applies by default, and therefore can be implicit, whilst discontinuity requires some form of signalling, i.e. the presence of linguistic devices that “function as cues to the reader”, “help[ing] the reader assign the utterance in which they occur to a continuation or a transition space” (goutsos, 1996, p. 517). the next section sketches out the conception of discourse organisation signals which is embodied in the present study. 68 a corpus-driven approach to discourse organisation: from cues to complex markers 2.2 signalling discourse organisation is “a struggle between different forces” at any given time, linguistic choices in text production are influenced by several principles concurrently at play. these choices are, in enkvist’s words, “the outcome of a conspiracy or a struggle between the different forces that affect the linearization of discourse” (enkvist, 1985, p. 321). we shall use an example to explain why enkvist’s image seems relevant: example 1 from a policy oriented text published by ifri (french institute for international relations)2 new budgetary cuts [section heading] let us now look at the effect of the crisis as things stand at present. [...] in the united kingdom , the defence budget, which amounted to 44.5 billion in addition to spending relating to external operations in afghanistan and iraq, is to be cut by [...] in germany , debate is raging over whether or not to abolish national service, which would reduce troop numbers from 250,000 to [...] in austria , a question mark is hanging over the military service and most of the country’s tanks have been withdrawn from service. [...] in greece, the defence budget will be amputated by [...] on reading example 1, one is immediately aware of the presence of a number of paragraphinitial adverbials (in the united kingdom, in germany, etc.). time and space adverbials are amongst potential sequentiality cues which have also been considered good segmentation markers by researchers working in a what-perspective: they can be associated with topic shifts (piérard and bestgen, 2006), and they introduce a new interpretation criterion projecting forward (charolles et al., 2005). adopting a how-perspective, we would argue that what is significant about these adverbials is that there are four of them in relatively short succession, exhibiting strong parallelism: paragraphinitial prepositional phrases (in + name of country) followed by a comma. together, they form an identifiable pattern, and recognising this pattern is in our view very much part of understanding what the text is about. now, a series of paragraph-initial sequencers (firstly, secondly,...) instead of adverbials, or adverbials of time instead of space, would create a functionally similar pattern. a sequence of four non-initial, non-detached adverbials, on the other hand, would definitely not realise the same text strategy (virtanen, 2004), and would not be perceived in the same way. in this perspective, features that are often overlooked in discourse organisation research must be fully taken into account, in particular layout and punctuation: the paragraph breaks, as well as the commas following each prepositional phrase, are clearly determining features. the pattern at once delineates the items as discontinuous and brings them together—-because of their parallelism—-into a higher level span (see figure 2 below). viewed from a what-perspective, the four adverbials in example 1 introduce spatial criteria which are essential for the interpretation of subsequent text: everything said after adverbial n and before adverbial n+ 1 only applies in (is only true for) the spatial area designated by the adverbial. we find unconvincing the dichotomy sometimes found in the literature on textual metadiscourse between propositional and non-propositional textual material (or primary and secondary discourse) (see ho-dac et al., 2012). the adverbials in example 1 are both at once: they have both an ideational and a textual role, and they are noteworthy for both the whatand the howperspectives. this duality takes us back to enkvist’s remark about “a struggle between different forces”: given that at any one time several processes are concurrently going on in text (producing content, organising text), signalling devices must be expected to be largely multifunctional, since they may be shared by several processes. 2. http://www.ifri.org/downloads/europevisions7ojehin.pdf 69 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre the perspective sketched out here challenges the view of discourse organisation as signalled in text primarily via specialised (lexical) discourse markers. along the lines of authors such as marcu (marcu, 2000, 2006), we would argue that discourse markers are likely to be more eclectic and less discrete, i.e. to come in the form of bundles of cues in which most of the time no single element is either necessary or sufficient. if, as suggested in the analysis of example 1, and in line with a number of authors (virtanen, 1992, 2004; hasselgård, 2010), adverbials only function as sequentiality markers when associated with other features (e.g. positional or punctuational features) and/or occurring in a series (ho-dac and péry-woodley, 2009), the search for discourse markers is redefined as a search for recurrent cue configurations. a similar approach is adopted in recent research on the signalling of implicit discourse relations (taboada and das, 2013). in their study, taboada and das stress that “we need to move beyond the signalling by discourse markers (...) in order to understand how relations are processed, and in order to extract them automatically” (idem, p.250). the authors propose annotating a wide range of cues in order to discover combined and multiple signals of discourse relations (idem, p.250), an objective close to our own search for complex discourse markers. 2.3 top-down processing influences discourse interpretation discourse markers are usually seen as bearers of instructions. for instance, connectives carry the instruction to link two text segments (or the propositional meanings they convey) via a particular relation. this conception is rooted in a bottom-up view of discourse interpretation as a unit by unit construction. we focus here on movement in the opposite direction, proposing that in text processing—-particularly in the case of expository text—-there is also an immediate perception of high-level signals, a gestalt-like grasp of large-scale textual patterns, which in turn influences the step by step interpretation process (asher et al., to appear). this interest in top-down processing relates to research in neighbouring fields: cognitive psychologists and psycholinguists have studied the effect of headings and sub-headings, and of layout features such as paragraph breaks, on reading comprehension and recall (lorch and lorch, 1996; lemarié et al., 2012, 2008; heurley, 1997); computational approaches to document generation and understanding, in their concern with linearisation (cf. section 2.1), have sought to define principles governing the physical presentation of text (power et al., 2003; bateman et al., 2001; virbel, 1989, 2015). we use the term “textual pattern” to suggest that this top-down processing may have more to do with pattern recognition than with a compositional meaning-construction process. signalling is part and parcel of the definition of textual patterns, since they are characterised by their ability to be readily perceived by readers. there can be no such thing as an implicit textual pattern. 2.4 discourse structures are multi-level in our attempt to draw attention to top-down processing, we referred to high-level signals and largescale textual patterns. more precisely, the signals and textual patterns in question are typically multi-level and apply recursively, and these are properties which are of great interest to us. the textual pattern which is the focus of this paper, enumerative structures, will be seen to range from a few lines to whole sections of text, and to allow several levels of embedding. this multi-level property is clearly related to how the text spans delimited by discourse structures interact with document structure segmentation (sections and sub-sections, paragraphs). example 1 illustrates such an interaction between two modes of organisation, where place adverbials and paragraph segmentation 70 a corpus-driven approach to discourse organisation: from cues to complex markers may be seen as signalling the items of an enumeration. the approach therefore implies attention to document structure (power et al., 2003), and its signals. visual signals of document structure are considered as fully-fledged discourse features with the potential to combine with other cues to form what we call complex discourse markers. the previous section has allowed us to clarify our general objective in the light of the basic tenets of our approach. we can now turn to the method specifically set up to identify these complex discourse markers, which involves automatic feature-tagging (section 4.1.1), manual annotation of textual patterns (section 4.1.2) and data mining techniques to identify correlations (section 6). the method is designed for long expository texts which differ noticeably from newspaper articles in terms of discourse organisation (cf. section 4.2.2). as described in section 6.3, this method has made it possible to identify complex discourse markers made up of cues appearing in series or in specific patterns. we also insist on the role of genre in determining what cues are used, or in shifting the balance of interpretation of particular sets of cues. the first stage in the description of the method is to present the context—-the annodis multi-level annotation experiment. 3. an annotation experiment to implement a multi-level approach to discourse in order to observe diverse structuring modes, including at high levels of organisation, we devised an annotation experiment to be carried out on lengthy non-narrative texts, organised into three distinct sub-corpora so as to allow potential genre-related variation to emerge. in accordance with the approach presented above, the annotation is not based on predefined markers: the identification of cue configurations functioning as discourse markers is an expected outcome of the analysis of the annotated data. however, our manual annotation relies on extensive pre-processing, in particular the systematic pre-marking of selected features, in an approach inspired by biber (biber, 1988; biber et al., 2007). figure 1 gives an overview of the methodology developed for this experiment. the association of exhaustive nlp-generated linguistic information (pre-marked features) with human intuitions (manual annotation of structures and cues) produces rich data, opening the way for new investigations using corpus-linguistics or data driven methods. two multi-level structures have been annotated according to this methodology within the annodis project—-topical chains and enumerative structures—-, but the present paper deals solely with the latter. in line with the approach outlined in section 2, the annotation project described here differs in several major ways from previous discourse annotation initiatives, such as the penn discourse treebank (pdtb prasad et al., 2008), the rst (rhetorical structure theory) treebank (carlson et al., 2003), or the discourse graphbank (wolf et al., 2004). the pdtb’s focus is low-level discourse structure (elementary predicate-argument relations) and it is grounded in a lexicalised approach to discourse (role of discourse connectives as predicates). though based on different models of discourse relations, with varying views on the role of lexical connectives, the rst discourse treebank and the discourse graphbank share similar objectives, in line with a what-perspective more than a how-perspective, to return to the distinction introduced in section 2.1. the fact that all these annotation projects use only news material, mostly from the wall street journal, is also revealing of how they differ from the experiment presented here, as will be made clear in the description of our experimental framework in section 4.2. 71 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre figure 1: a data-intensive method for the study of discourse organisation and discourse signalling 3.1 why enumerative structures? having set out the context in which the annotation experiment was designed, we will now focus on one of the two multi-level structures selected for annotation: enumerative structures beyond the sentence level, as illustrated by example 2: example 2 from wikipedia (english): “global warming” (retrieved 2014-10-09) examples of impacts include: • food: crop production will probably be negatively affected in low latitude countries, while effects at northern latitudes may be positive or negative. global warming of around 4.6 ◦c relative to pre-industrial levels could pose a large risk to global and regional food security. • health: generally impacts will be more negative than positive. impacts include: the effects of extreme weather, leading to injury and loss of life; and indirect effects, such as undernutrition brought on by crop failures. our interest in enumerative structures as textual patterns deployed at different discourse organisation levels is initially rooted in systemic functional linguistics: halliday’s description of text as “the unit of the semantic process” (halliday, 1977/1983, p. 63) encourages the formulation of hypotheses on how perception of high-level structures may influence text interpretation at a more local level. in this context, we are developing an approach to “texture” which takes into account visual aspects of text construction, aspects considered by power et al. (2003) as part of “document structure”. along the lines defined by these authors—-pursuing nunberg’s reflection on “text grammar” (nunberg, 1990)—-and by researchers inspired by virbel’s model for “text architecture” (luc 72 a corpus-driven approach to discourse organisation: from cues to complex markers et al., 2000; luc and virbel, 2001), we argue for a linguistic status for what power et al. (2003) call “the graphical component”, on a par with lexico-syntactic cues. after luc and virbel (2001), we describe enumerative structures as textual objects resulting from a textual act whereby text is arranged (visually or through other devices) so that the reader becomes aware of this textual arrangement. the associated semantics is that the reader is led to interpret the enumerated elements (i.e. the items) as similar in some respect, and therefore as constituting a segment homogeneous in terms of a “co-enumerability criterion”. the co-enumerability criterion may be lexically expressed, as in example 2 (examples of impacts), or realised more indirectly. two peripheral elements may contribute to this textual arrangement: a trigger which announces the enumeration and/or a closure. enumerating appears thus as a very basic way of organising text, and a generic one in the sense that it can be resorted to for a wide range of semantic or rhetorical functions. despite this basic character, enumerating as a text construction strategy has not elicited much interest among discourse linguists: the “relative neglect” noted by schiffrin in 1994, and described by her as “a surprising oversight” (schiffrin, 1994, p. 378) still seems to apply. there are, on the other hand, quite a few studies focusing on specific linguistic elements playing a role in enumerating, in particular lexical item introducers, which have been variously named “linear integration markers” (turco and coltier, 1988; jackiewicz, 2005), “sequencers” (hempel and degand, 2008) and “serial markers” (bras and schnedecker, 2013). mostly concerned with the semantic description and classification of such markers (numerical: firstly, etc.; temporal: subsequently, finally, etc.; spatial: in the first place, etc.), these studies tend to leave out non-lexical cues such as visual devices and document structure. research on “enumerable” nouns (tadros, 1985, p. 6) or “shell nouns” (francis, 1994; schmid, 2000), so-called because of their underspecified meaning, also constitutes a relevant related field of investigation. such nouns are seen as announcing (“predicting” to use tadros’ term) subsequent specification in the following text, and, in the case of enumerative structures, naming the co-enumerability criterion which provides the rationale for enumerating. in contrast with these two groups of studies which mostly take specific markers as their starting point, our interest is in the text-structuring role of enumerating, and in the diverse ways in which these structures are signalled (ho-dac et al., 2010). seen from this angle, the “markers” selected in the studies mentioned above are cues amongst others, playing a role in multiple-cue signalling devices—-complex discourse markers. the existing research, however, as well as providing numerous insights, raises a number of issues of interest to our project, issues which underlie the questions we are going to ask of our annotated data. from our text organisation perspective, as distinct from the discourse markers perspective in which studies of item introducers were carried out, there is no reason to give special treatment to lexical markers or to distinguish them from the various other ways in which enumerating may be signalled. as mentioned earlier, we take full account of document structure and include among potentially relevant text features the visual devices which organise the text on the page, delimiting different spans of text: typographical variations, layout (indentation, line spacing, paragraph breaks, bullets, headings). as visual devices can be seen as pulling enumerative structures towards the textual component while lexico-syntactic cues seem better able to contribute to the ideational component, a central objective of this study is to examine how these various cues interact, and what variables may have an impact on these interactions. another underlying question concerns the nature of the relation between the enumeration proper, i.e. the items, and the peripheral elements: trigger and closure. viewing the relationship between the co-enumerability criterion expressed in the trigger (or closure) and each of the items in the enumer73 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre ation in terms of taxonomy-based hypernym-hyponyms relations is clearly too narrow: expressions of co-enumerability may be used to gather together world entities, but also textual objects (section, chapter), rhetorical functions (examples, as in example 2), or other forms of textual organisation such as steps in a chronology, stops in an itinerary, etc. within a larger goal of exploring how the textual and the ideational components are enmeshed, enumerative structures deserve scrutiny for their ability to organise text and categorise content at one and the same time. 3.2 an annotation model for enumerative structures our annotation model was designed to allow a moderately open-ended annotation task, aiming to leave some leeway for possible off-model annotators’ intuitions. according to this model, an enumerative structure (es) extends beyond a single sentence3 and is made up of three segments: (1) the trigger: an optional introductory segment, (2) the enumeration, defined as a list of at least two co-items, (3) the closure: an optional closing segment. figure 2 gives a schematic representation of this definition of ess. it shows how the elements entering into this linear arrangement are in fact different kinds of nested text spans linked via a number of possible short and long distance discourse relations4. figure 2: enumerative structure representation following luc et al. (2000), enumerating is described as a textual act which asserts the coenumerability of the listed “entities” by transposing it textually. this “textual transposition” can take many forms, the most obvious being when items are separate paragraphs with bullet points. the nature of the co-enumerability may be made explicit, as in example 3 below, where the two text segments are presented as similar in that they are both “criticisms” (a moralistic criticism [une critique moraliste] and a deterministic criticism [une critique déterministe])5. the co-enumerability 3. in accordance with this definition, the sentence-level es which ends the second item of example 2 was not included in the annotation. in some cases, however, annotators decided to include single sentence ess. 4. the identification and annotation of these relations were outside the scope of this annotation exercise. 5. all examples from this point on are extracts from the annotated french-language corpus. for each example, the expressions which are central to the commentary are given in the text in an english translation followed by the original french (in square brackets). for complete english translations of the french-language examples, see appendix 1. 74 a corpus-driven approach to discourse organisation: from cues to complex markers criterion may be expressed in the trigger, in a prospective element (two types of criticisms [deux types de critiques]), and/or in the closure, in an encapsulation (these two criticisms [ces deux critiques]) (cf. conte, 1996; sinclair, 1983). the enumeration is the only necessary element in this structure. example 3 from wiki sub corpus6 (wik2 libertese coder3 1254325598390) en effet, contre la liberté indépendance, il existe au moins deux types de critiques : es trigger une critique moraliste : cette liberté relève de la licence, i.e. de l’abandon au désir. or, il n’y a pas de liberté sans loi (rousseau, emmanuel kant), car la liberté de tous serait en ce sens contradictoire : [...] item 1 on remarque que dans cette conception philosophique de la liberté, les limites ne sont pas des limites contraignant la liberté de la volonté humaine ; ces limites définissent en réalité un domaine d’action où la liberté peut exister, ce qui est tout autre chose. une critique déterministe : s’abandonner à ses désirs, n’est-ce pas leur obéir, et dès lors un tel abandon ne relève-t-il pas d’une forme déguisée de déterminisme ? nous serions alors victimes d’une illusion de libre arbitre : [...] item 2 nietzsche reprendra cette critique : aussi longtemps que nous ne nous sentons pas dépendre de quoi que ce soit, [...] ces deux critiques mettent en lumière plusieurs points importants. closure [...] from example 3 onward, the formatting of examples obeys the following conventions: the two right-hand columns delimit each annotated es and its components; horizontal lines in the lefthand column indicate paragraph breaks in the original, i.e. each boxed segment corresponds to a complete paragraph. in example 3 for instance, the trigger is in a paragraph, the items consist of two paragraphs each, and the closure starts a sixth paragraph. where excessively long paragraphs were cut, this is signalled by [...] (items 1 and 2). when a component covers only part of a paragraph, this is indicated as in example 3 for the closure. the reference associated with each example (e.g. wik2 libertese coder3 1254325598390) is its identifier in the annodis resource. it can be searched for in the resource using the annodis browser7. 4. annotating enumerative structures: annotation procedure and experimental framework biber et al. (2007) propose a step-by-step method for corpus-based studies of discourse, providing a detailed account of seven steps seen as necessary in order to arrive at generalisable descriptions of discourse structure in corpora. these steps may be carried out in two possible orders: either topdown (a priori communicative/functional categories provide the basis for manual text segmentation) or bottom-up (starting with automatic segmentation based on lexical cohesion). in both cases, the segmentation stage leads to a linguistic characterisation based on the analysis of the distribution of textual features, according to the methodology initially set up by biber to produce an emergent text typology (biber, 1988). in the bottom-up approach, the communicative/functional categories are derived from the linguistic characterisation (identification of clusters, which are then given a functional interpretation). our approach is fundamentally grounded in biber’s use of systematic feature-marking and analysis, but can be seen as proposing a third way with respect to biber et al. 6. the composition of our corpus is described in section 4.2.2 below. 7. http://redac.univ-tlse2.fr/corpus/annodis/me_download/annodis_se.xml 75 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre (2007). in accordance with the top-down approach, the functional units under study are determined and defined a priori. this is the model which is embodied in the annotation manual, and sketched in section 3.2 above. but in a perspective akin to biber et al’s bottom-up approach, the human annotation process is guided by the pre-marking of features which emerge from previous studies as potentially relevant for the identification and description of enumerative structures. the next section outlines these two major steps in the annotation procedure. section 4.2 then fills in the detail of how they were carried out in practice: it describes the annotation interface, the corpus and the annotation model. 4.1 a two-step annotation procedure 4.1.1 automatic pre-marking of features prior to manual annotation, a systematic pre-marking of potentially relevant features was automatically carried out on the pos-tagged and syntactically parsed texts8, relying on local grammars and making use of specifically designed lexicons. the selection of features calls upon previous research (see section 3.1) and covers a wide variety of linguistic phenomena, both visual (punctuation, layout) and lexico-syntactic. the set of pre-marked features is organised in table 1 below into seven types, which constitute the basis for the analyses which will be presented in section 6. feature type feature description typography and layout indentation, line space, paragraph breaks, bullets and numbering, headings punctuation patterns [:] preceding [;] or [,] [:] in final paragraph position sequencers (sentence-initial—s.i.) first, secondly, on the other hand, the third x circumstance adverbials (s.i.) including lexemes expressing time, place, categories: since 1956, in austria, in linguistics, prospective elements specific np patterns including specific lexemes: the following elements, two types of criticisms encapsulations pre-verbal demonstrative nps with a numeral as determiner: these two criticisms, connectives (s.i.) moreover, to sum up, in contrast table 1: the set of pre-marked features the inclusion of sentence-initial circumstance adverbials in this set of features is based on charolles’ framing hypothesis (charolles et al., 2005): such adverbials have the potential to project forward an interpretation criterion, and thus define the initial boundary of a discourse frame, i.e. a text segment clustering around a specific interpretation criterion. they were pre-marked as potential item introducers, as were sentence-initial connectives, which are potential sequencers. the automatic detection of such sentence-initial cues proceeded in two steps: the first was to detect all syntactically detached elements occurring before the grammatical subject; the second to attribute 8. the pos-tagging was performed by treetagger (schmid, 1995), and the parsing by syntex (bourigault et al., 2005; bourigault, 2007). 76 a corpus-driven approach to discourse organisation: from cues to complex markers to each detached element a syntactico-semantic function (circumstantial adverbial, sequencer, other connective). prospective elements consist in simple and fairly unambiguous cataphoric patterns: xxx as follows (.:) prep the following number xxxs. in addition to these cataphoric patterns, prospective elements also include plural noun phrases where a classifier, or “shell-noun” (cf. section 3.1), occurs9. as for encapsulations, two patterns associated with shell-nouns were used (conte, 1996; schmid, 2000): plural demonstrative noun phrases and noun phrases introduced by the semi-determiner such (tel(le)s). example 4 reproduces (3) with all pre-marked features shown in bold (prospective elements, encapsulations and punctuation). bullets and other layout features are not highlighted, as they are, by definition, visible. example 4 from wiki sub corpus (wik2 libertese coder3 1254325598390) en effet, contre la liberté indépendance, il existe au moins deux types de critiques : es trigger une critique moraliste : cette liberté relève de la licence, i.e. de l’abandon au désir. or, il n’y a pas de liberté sans loi (rousseau, emmanuel kant), car la liberté de tous serait en ce sens contradictoire : [...] item 1 on remarque que dans cette conception philosophique de la liberté, les limites ne sont pas des limites contraignant la liberté de la volonté humaine ; ces limites définissent en réalité un domaine d’action où la liberté peut exister, ce qui est tout autre chose. une critique déterministe : s’abandonner à ses désirs, n’est-ce pas leur obéir, et dès lors un tel abandon ne relève-t-il pas d’une forme déguisée de déterminisme ? nous serions alors victimes d’une illusion de libre arbitre : [...] item 2 nietzsche reprendra cette critique : aussi longtemps que nous ne nous sentons pas dépendre de quoi que ce soit, [...] ces deux critiques mettent en lumière plusieurs points importants. closure [...] 4.1.2 manual annotation of structures and cues pre-marked features were meant to act as flags to guide annotators in the identification of sporadic discourse units, leading them away from linear reading towards a more global view of text. the manual annotation consisted of two main tasks: first, delimiting and labelling the components of ess which were detected (trigger, co-items, closure and the co-enumerability criterion if explicitly stated); second, marking-up features considered as relevant cues, by either validating pre-marked features or annotating and labelling complementary cues. in (3) for example, after delimiting the components and identifying the co-enumerability criterion (criticisms [critiques] expressed in trigger and closure), the annotators validated all the pre-marked features and went on to annotate as an extra cue the parallelism between the two nps (a moralistic criticism / a deterministic criticism) introducing the items. once identified, additional cues were labelled according to the categories defined in the annotation guidelines, i.e. the categories used for premarking with the addition of syntactic parallelism10. 9. a list of 64 classifiers was manually adapted from schmid (2000) to french. 10. this lexico-syntactic feature was not pre-marked as its automatic identification is not yet operational due to the high computational complexity of the task. 77 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre if no predefined label fitted, the annotators were invited to create descriptive labels. as a consequence, a proportion of annotator-added cues form a heterogeneous set of non-categorised features (e.g. coreferential expression, trigger repetition, apposition, named entity...). each feature marked-up as relevant (whether or not it was pre-marked) becomes what we call an “es-cue” , i.e. a linguistic feature which, in combination with others, participates in the signalling of enumerative structures, and hence of discourse organisation. it is through the identification of recurring configurations of es-cues that we propose to define complex markers (cf. section 6). 4.2 experimental framework 4.2.1 the annotation interface with annotators having to annotate textual zones of varying sizes, as well as deal with discontinuity and possible overlaps with previously delimited zones, the annotation task was highly complex and required an efficient purpose-built annotation interface. the design of the glozz interface (widlöcher and mathet, 2012) reflects two major requirements concerning text visualisation and the annotation procedure itself. the text visualisation interface has to take into account the output of the pre-marking procedure, including xml encoding of the text layout and formatting. the annotation interface must offer a panel of user-friendly editing tools for delimiting and characterising units; it must also facilitate navigation in the text being annotated. the solution adopted is to offer two views of the text: a main view for annotation and a global view to get a large-scale vision of the text (cf. glozz premarked documents in figure 3). through these two modes of access to text, the interface encourages the annotator to combine a top-down and a bottom-up approach during reading, so as to be able to see local cues as well as global structures. in order to ensure that the grasp of texts by annotators is as ecological as possible, and to reduce the inevitable processing difference between annotating and reading, the main view must present texts as real documents with major aspects of layout preserved. all these requirements were taken on board in the design of the glozz annotation platform11. 4.2.2 the annodis corpus our approach to discourse imposes constraints on the selection of texts for the corpus. contrary to previous discourse annotation programmes (cf. section 2), we opted for lengthy expository texts, first because they tend not to be structured around a major referent—-as is often the case in narratives—-, secondly because they favour complex organisation and are therefore more likely to contain different structuring modes (including complex document structure). another consideration was that corpus linguistics methods require fairly large volumes of texts and sufficient numbers of annotated structures if some generalisation of observations is to be possible (cf. piérard and bestgen, 2006). finally, because we consider genre as a feature to be taken into account in the definition of complex discourse markers (ho-dac and péry-woodley, 2009; taboada and das, 2013), we compiled a diversified corpus enabling contrastive analyses. considering all these criteria, the texts selected for inclusion in the annodis corpus combine three genres of lengthy expository texts: web-encyclopaedia articles (nearly 200,000 words from the french wikipedia in its version of june 18, 2008), scientific papers (proceedings of congrès mondial de linguistique française 2008, about 135,000 words) and reports in the field of interna11. http://glozz.free.fr/ 78 a corpus-driven approach to discourse organisation: from cues to complex markers figure 3: annodis corpus preparation: from original text to pre-marked document ready for annotation tional relations (from the french institute for international relations, over 180,000 words). these three sub-corpora are respectively named wiki, ling and geop. the total number of texts (83) was set in accordance with the time constraints on the annotation programme. given our objectives, special care was taken in the preparation of the corpus: not only were all the texts xml encoded in conformity with the tei-p5 norm, but it was imperative that the visual appearance of the texts be preserved, which is also a departure from previous experiments. semiautomatic procedures were set up to annotate and encode the specific layouts signalling textual objects: title, headings with their level, paragraphs, lists and citations. figure 3 gives a schematic view of the corpus preparation process. 4.2.3 the annotation exercise the manual annotation was programmed in two stages, beginning with an exploratory phase dedicated to the evaluation of the task’s feasibility, which led to a series of improvements in the procedure: clarification of the protocol, simplification of the annotation model, changes in the visualisation parameters, and correction of the annotation guide. then came the annotation task itself. three undergraduate linguistics students were selected as neutral (non-specialist) annotators. the 83 texts of our corpus were split into 3 sets. there was a training phase during which four texts were jointly annotated and the three annotators were encouraged to compare and share their annotations. this training phase led to an improved, stable version of the annotation guide (colléter et al., 2012). after training, six new texts were annotated by the three coders, and these annotations were used for measuring inter-annotator agreement. measuring inter-annotator agreement in this case means checking whether two annotators have identified the same ess in the target text. we considered that there was agreement on a given es when annotators a and b had selected the same text span, and identified exactly the same items 79 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre within this span. if two annotated ess differed only in terms of trigger and/or closure (these units being optional) while respecting the previous conditions, they were considered identical. overall agreement between two coders for each text was measured via the f-score. the nature of the task ruled out traditional agreement measures (such as cohen’s kappa) because es marking is not a categorisation task. in a task such as ours, as hripcsak and rothschild (2005) explain, there is no access to a negative count, i.e. we cannot take into account the fact that both annotators agreed that there are no ess in a particular span of text. for the evaluation of a marking task, the f-score is the measure which is most commonly used (see e.g. brants (2000) for syntactic annotation). in our case the f-score is based on the number of ess identified by both annotators and the overall number of ess identified by each, as formulated below: f = (2×# of conjointess) #ofcodera ess+#ofcoderb ess this score was measured for every pair of annotators over the 6 texts (2 from each subcorpus), each having been annotated by three different coders. the overall average f-score is 0.67 (sd 0.21), meaning that over two ess out of three marked up by one coder were also marked up by the other coder. this value was considered sufficient for the final annotation phase to be launched, whereby the remaining texts were distributed among the three annotators, each text being dealt with only once. in a final stage, disagreements were post-annotated, and adjudicated versions of the ten multiannotated texts from the training and evaluation phases were produced12. as observed in colléter et al. (2012), disagreements mostly concern small and/or isolated ess, as well as structures which may be considered as contrasts or chronologies rather than ess. the data collected at each stage is available on-line (original documents, texts prepared for annotation, pre-adjudicated versions, etc.)13, together with a technical report which includes the annotation manual together with coders’ testimonies and adjudication details (colléter et al., 2012). the exploitation of the annotations has so far been carried out in two ways: manually, by means of an exploration interface14, and automatically, via data mining techniques. 5. analysing the annotated corpus: enumerative structures (ess) as a basic strategy the rich annotated data resulting from the annotation exercise just described can now be examined for answers to the issues and questions raised in sections 1 and 2. we start with a descriptive survey of the frequency, length and distribution of ess in the corpus, which provides the basis for a quantitative assessment of their importance as a text construction strategy (section 5.1), and a structural characterisation in terms of cardinality (number of items) and composition (presence/absence of a trigger and a closure) (section 5.2). in section 5.3 we compare bottom-up and multi-level approaches, taking advantage of the annotation of ess in terms of discourse relations in a sub-corpus. we finally delve deeper into characteristics which are directly relevant to two major discourse organisation issues: enumerative structures are multi-level structures, capable of organising textual material at any level of granularity from entire sections to the sub-sentential level (see the typology in section 5.4); enumerative structures are text-segmenting patterns as well as content-structuring 12. a reflection on the annotation exercise is presented in ho-dac and péry-woodley (2014). 13. http://redac.univ-tlse2.fr/corpus/annodis/me_download/index_en.html 14. http://redac.univ-tlse2.fr/corpus/annodis/me_download/annodis_me_browser.html 80 a corpus-driven approach to discourse organisation: from cues to complex markers categorisation devices, highlighting the interweaving of the textual and ideational components (section 6). 5.1 frequency, length and distribution of annotated ess table 2 summarises the results of a first survey of the annotations, showing ess to be a basic strategy frequently resorted to by writers in different genres of expository texts. there is an average of 12 ess per text in our corpus (range: 2 to 34), and ess have an average length of 429 words, with considerable variation (from 8 to 8,666 words). the “text coverage” value is the proportion of a given text appearing in at least one es15: on average, 44.6% of a text’s words are contained in ess; in some cases, text coverage is over 90%. ess are present in all three sub-corpora with significant variations which will be presented in the next section together with variations regarding composition. sub-corpus texts, n ess, n mean n of ess per text mean length (words / es) text coverage (%) wiki 28 401 14 455 55.5 ling 25 297 12 452 46.8 geop 30 293 10 369 32.8 total 83 991 12 429 44.6 table 2: frequency and coverage of annotated ess in annodis and all three sub-corpora the next sub-sections aim to flesh out this initial picture via analyses of the composition ess and of their interaction with discourse relations and document structure. 5.2 composition of annotated ess table 3, an overall view of the composition of ess in the corpus, shows that only a small proportion is complete with respect to the canonical three-part model—-trigger, items, closure16. sub-corpus trigger items closure completeness n % cardinality (mean n of items) n % complete ess (%) minimalist ess (%) wiki 300 74.8 4.1 36 9.0 7.5 23.7 ling 230 77.4 2.9 46 15.5 12.1 19.2 geop 209 71.3 2.9 49 16.7 12.3 24.2 total 739 74.6 3.4 131 13.2 10.3 22.5 table 3: es composition in annodis and all three sub-corpora in example 5, the completeness of the structure combines with a profusion of signalling devices (in bold). 15. as ess can be nested, a portion of text can be contained in several ess. this phenomenon, though fairly frequent, was not taken it into account at this stage of the analysis. 16. complete ess have both a trigger and a closure. minimalist ess have neither trigger nor closure. 81 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre example 5 from ling sub-corpus (ling kleiberse coder3 1254143156093) 2.2 deux manières de nier la polysémie es trigger une réponse possible est [...]la polysémie en tant qu’association de plusieurs sens à une même forme lexicale se trouve niée de deux manières apparemment paradoxales : -ad’une part, les vocables donnés comme polysémiques se voient en quelque sorte “monosémisés” par [...] item 1 -bd’autre part, de façon tout à fait inverse aux tentatives de monosémisation, on fait proliférer les sens [...] item 2 les positions -aet -bne sont qu’apparemment paradoxales : il n’y a aucune contradiction, d’un côté, à [...] closure the trigger is clearly marked: it consists of the heading (2.2. two ways of denying polysemy [deux manières de nier la polysémie]) and the sentence following it. the prospective element in the heading, made obvious by a numeral determiner, two ways [deux manières], is reiterated in the first sentence. the items are then signalled in four complementary ways: each one makes up a paragraph, they are introduced by a dash, sequentially labelled with letters, and start with a correlative adverbial which stresses the parallelism between the two assertions (on the one hand – on the other hand [d’une part – d’autre part]). the closure ends the enumeration with an encapsulating noun phrase (positions -aand -b[les positions -aet -b-]). this example shows how different es-cues can reinforce one another, giving the es high visibility. although table 3 gives the average number of items per es as 3.4, it is worth noting that 42% of ess contain only two items, whilst rare extreme cases may comprise up to 48 items17. cardinality (number of items) and length (number of words) are positively correlated, though at a marginal level (r=0.14). closures are rare (13.2%), whereas most ess start with a trigger (74.6%). given that either trigger or closure can express the co-enumerability criterion, trigger-less ess could have been thought more likely to have a closure, but cross-tabulation of the data does not confirm this hypothesis: only 3% of trigger-less ess have a closure. to end this general picture, complete ess are fairly rare (around 10%) while over 22% of ess are minimalist, i.e. composed only of items. no significant correlation has been established between completeness and length or cardinality. looking at tables 2 and 3, interesting variations across our sub-corpora begin to emerge. the largest ess, both in length and cardinality, are found in encyclopaedia articles (wiki), where they also cover a larger part of the text than in other sub-corpora (455 words/es; 55.5% of total text). at the other end of the spectrum, international relation reports (geop) contain fewer and shorter ess (369 words/es) which cover a much smaller proportion of the text’s surface (32.8% of total text). these variations across sub-corpora are all statistically significant (p < 0.001, kruskal-wallis test), but a clearer understanding of the text structuring role of ess is needed to assess their linguistic significance. this is what the next two sections work towards by bringing into the picture two distinct sets of annotation, discourse relations and document structure, in order to arrive at a better characterisation of ess’ discourse function. 5.3 interaction between discourse relations and ess as mentioned in section 2, the annodis project also involved a bottom-up annotation of discourse relations, and a part of the annodis resource, labelled “annodis duo”, was annotated with both ess and discourse relations. the model and method for the annotation of discourse relations originate in segmented discourse representation theory (sdrt): coders started by segmenting 17. only 4 ess are made up of more than 15 items. 82 a corpus-driven approach to discourse organisation: from cues to complex markers texts into elementary discourse units (edus) and, after reaching mutual agreement, associated them with discourse relations, building up complex discourse units (cdus) until they arrived at a complete hierarchical representation of the text. sixteen discourse relations were annotated18, a selection which represents a compromise between informativeness and reliability of the annotation process. the selection constitutes a consensual set of relations which are shared by most discourse models, or correspond to well-defined subgroups in fine-grained theories (hovy, 1990), as well as to the level of grain adopted for the penn discourse tree bank (prasad et al., 2008)19. ess annotated in the annodis duo resource20 contain an average of 24.5 edus per es (between 3 and 65). because cdus are recursively nested, the raw number of cdus per es is not relevant without a more qualitative analysis. looking at the discourse relations annotated on the borders of triggers and items, i.e. relations associated to edus starting or ending triggers and items, the following associations may be observed: • 75% of triggers end with an edu linked forward to another segment via elaboration*21 and/or frame relations, • 94% of items start with an edu attached to another segment via an elaboration* relation associated in 35% of cases with a simultaneous continuity relation, • when considering only initial items, 92% of ess have an initial item where the starting edu is associated to an elaboration* relation. the fact that most ess can be described in terms of just two discourse relations, i.e. elaboration* and continuity, confirms that the structure can legitimately be regarded as a functional unit, regardless of the diverse forms in which it occurs. moreover, each of these relations seems to have a specific role in the structure: elaboration* between the trigger and items, and continuity between items, as example 5 illustrates. these observations also support the sdrt model developed in bras et al. (2008) and vergez-couret et al. (2008) for explaining long distance attachments between trigger segments and cdus introduced by pairs of discourse markers such as d’abord/ensuite (first/then). as a consequence, a new relation called enumeration was introduced by vergez-couret et al. (2012) as defined in figure 4. according to its authors, the enumeration relation was introduced so that analysts would be able “to juggle between constituents describing semantic content and constituents describing discourse packaging while entity-elaboration would not have allowed so” (vergez-couret et al., 2011). the authors clearly sense that two different types of text-building process are at play, which they want to account for while keeping them apart, hence the “juggling”. this term calls to mind the doubts expressed in goutsos (1996, p. 257) as to the possibility of dealing with essentially textual relations in ideational terms: 18. explanation, goal, result, parallel, contrast, continuation, alternation, attribution, background, flashback, frame, temporal-location, elaboration, entity-elaboration, comment 19. like the annotation guide for multi-level structures, the guide produced for the annotation of discourse relations is freely available (muller et al., 2012). it provides an intuitive introduction to discourse segments, including the question of embedding of discourse segments to form cdus; a list of detailed instructions describing how to handle segmentation; and a semantic definition of each discourse relation with examples and potential markers. 20. 26 ess: 15 in wiki, 5 ling and 6 in geop. 21. entity-elaboration and elaboration relations are merged for this analysis under the label elaboration*. 83 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre co-items (πb, πc) together introduce a complex constituent (π) which is attached to the trigger (πa) by the enumeration relation. a coordinating relation (by default continuation) is inferred between the co-items. figure 4: discourse structure for a classical enumerative structure (vergez-couret et al., 2012) “ideational analyses of texts have identified relations of joint, list or sequence (hoey, 1979; mann and thompson, 1988), whose status is clearly not so prominently ideational as textual. more generally, it is doubtful whether essentially presentational relationships like enumeration or listing can be couched in ideational terms at all. [...] the insistence on recognising a semantic relation between every single text segment comes into conflict with the occurrence of purely descriptive, propositionally loosely or arbitrarily related chunks of text.” a similar question was raised in an earlier study on enumerating by luc et al. (1999), who argued that their initial representation within virbel’s text architecture model (cf. section 2.1) needed complementing by a rhetorical structure theory representation, while pointing out the inadequacy of tree-like structures to represent ess and stressing the importance of visual clues, largely overlooked by researchers working within rst. regarding the latter argument, virbel et al. (2005, p.234) denounce linguistics’ blindness to visual cues : “linguistics, just as—-to a lesser extent—-information science, has long been ‘blind’ to the role of visual properties of written language, while other research fields (anthropology, history of texts, cognitive and experimental psychology) did point to the fundamental importance of these properties from the viewpoint of cognition.”22. among the visual properties overlooked by linguists, including discourse linguists, are titles and headings, whose role in text processing has been the object of much study in cognitive psychology (lorch and lorch, 1996; lemarié et al., 2008, 2012), but which are difficult to integrate within a discourse relations approach. the present study was designed from the outset to deal with the corpus not just as texts but as documents, whose layout structure is meaningful. the next section 22. “les sciences du langage comme, dans une moindre mesure, celles de l’information sont restées longtemps “aveugles” au rôle des propriétés visuelles du langage inscrit, alors que d’autres recherches (anthropologie, histoire des textes, psychologie cognitive et expérimentale) avaient signalé l’importance fondamentale de ces aspects du point de vue cognitif.” 84 a corpus-driven approach to discourse organisation: from cues to complex markers focuses on an annotation layer dedicated to the documents’ layout structure, which appears to be particularly well-suited to characterising the annotated ess in their diversity. 5.4 enumerative structures are multi-level: a granularity-based typology as mentioned in section 4.2.2, the layout structure of documents was annotated according to teip5 encoding. the textual objects considered here are sections, headings, lists and paragraphs. they are used as features in order to account for the variety of annotated ess in terms of length and composition. because they enter into hierarchical relationships, these layout units also provide a scale for describing ess’ granularity level (cf. section 2.4). we observed earlier that in terms of completeness, ess show variations that are not explained either statistically by length or cardinality (cf. sections 5.1 and 5.2) or by distinct discourse relations (cf. section 5.3). in contrast, granularity level appeared as the most informative variable for the classification of these structures (see ho-dac et al., 2010). the interaction between ess and the document’s layout structure, presented in table 4, provides the basis for a granularity-based typology of ess. this granularity-based typology emerged as the optimal way of clustering the annotated ess according to quantifiable variations in their form and composition. moreover, it gives us a way of organising the data by distinguishing classes of objects likely to make use of different signalling modes, as described in section 6. each type is now defined more precisely in terms of its typographical and layout features. • type 1 corresponds to multi-section ess, in which items are sections with a visible heading, as in example 6 below. • type 2 ess are prototypical formatted lists where each item is signalled by a bullet or number, as in examples 2 and 3 above. • ess which extend over more than one paragraph but do not belong to either of the previous types are type 3. these type 3 ess contain at least one paragraph break which may occur between two components (e.g. between trigger and first item or, as in example 8 between final item and closure as well as between items), with no specific constraints on the position and/or number of paragraph breaks. • finally, type 4 stands for ess contained within a single paragraph, see example 10 and 11. it is the most frequent across the whole corpus. es type description nb of ess mean length nb % (nb of words per es) type 1 multisection 126 12.7 1,858 type 2 bulleted list 244 24.6 184 type 3 multiparagraph 216 21.8 449 type 4 (intra)paragraph 405 40.9 120 all types 991 100 429 table 4: granularity-based typology of ess there is considerable variation in the distribution of types across sub-corpora, as shown in figure 5. wiki ess are the most strongly associated with visual layout: in 19% of cases, the items 85 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre are headed sections (type 1) and in 36.4% they are formatted lists (type 2). such emphasis on visual properties is to be expected in texts designed to be read on screen. it is in marked contrast with linguistics papers and international relations reports, where ess aligned on visual layout are fairly rare (fewer than 10% of type 1 ess) and type 4 ess i.e. low-level structures without visual properties, are the most frequent (45.5% in ling and 61.4% in geop against 22.4% in wiki). type 3 ess, multi-level structures without headings or bullets as item introducers, are the most stable across corpora (between 20% and 23%). these variations support our assumption that genre must be considered a relevant feature for the characterisation es and also for the definition of complex discourse markers. for example, (sub)headings may be considered as es-cues only in specific genres or text types. figure 5: granularity-based es types across sub-corpora as mentioned above, the proposed typology emerged as the best way to explain the variation in the form and composition of ess. an overview of how es composition varies across types is given in table 5, followed by descriptions of the four es types. es type trigger items closure completeness sub-corpus % cardinality (mean nb of items) % complete ess (%) minimalist ess (%) type 1 84.1 3.4 4.0 4.0 15.9 type 2 96.3 4.1 13.1 13.2 3.7 type 3 60.6 3.7 20.8 16.2 34.7 type 4 65.9 2.8 12.1 7.4 29.3 all types 74.6 3.4 13.2 10.3 22.5 table 5: typology of ess composition of annotated ess 5.4.1 type 1: multisection ess type 1 ess cover several sections of a document, with section headings signalling co-items. the least frequent in all three sub-corpora, they are, as can be expected, the longest ess in our corpus, averaging 1,858 words (with an enormous range: 252 to 8,666 words), yet their cardinality is close to the average (3.4 items per es). the few type 1 ess which have a closure (4%) are all complete 86 a corpus-driven approach to discourse organisation: from cues to complex markers ess i.e. they also have a trigger. indeed, most type 1 ess have a trigger (84%) which is generally a heading of the next level up and announces the enumeration both visually (via document structure) and semantically. example 6 shows a type 1 es with a trigger and 2 items: example 6 from wiki sub-corpus (wik2 julescesarse coder2 1254907327695) 6. les conquêtes amoureuses de césar es trigger 6.1. les femmes de la haute société romaine item 1 d’après l’historien latin suétone, césar séduit de nombreuses femmes tout au long de sa vie et plus particulièrement celles issues de la haute société romaine. il aurait ainsi séduit postumia, la femme de servius sulpicius, lollia, [...] césar entretient des relations particulières avec servilia caepionis, [...] le penchant de césar pour les plaisirs de l’amour semble également attesté par [...] 6.2. les reines item 2 césar a des relations amoureuses avec eunoé, femme de bogud, roi de mauritanie. cependant, sa relation avec cléopâtre vii est restée plus célèbre. [...] the relation between a section heading and its sub-headings could arguably be seen as an inclusion relation similar to that of co-items in an enumerative structure. yet it is not the case that all headed sections including headed sub-sections can be classed as ess, and indeed most were not identified as ess by the annotators. what marks out the annotated type 1 ess is the presence of a semantic criterion linking the items, in other words the fact that they function on both the ideational level and the textual level: in example 6, the first level heading, cæsar’s amorous conquests, provides this semantic criterion which unites under the category cæsar’s amorous conquests upper-class roman women (les femmes de la haute société romaine) and queens (les reines). 5.4.2 type 2: bulleted lists type 2 ess are characterised by the presence of bullets or numbers signalling each item. they have the highest cardinality (4.1 items/es), but are significantly shorter (184 words/es, p < 0.001), their constituent items being generally restricted to short phrases, as in example 7 below. there are exceptions, however, such as example 3 above, where some items cover several paragraphs. triggers are almost systematically present: 95% in wiki, 97% in geop, and 100% in ling. the corollary is a tiny percentage of minimalist ess. example 7 from wiki sub-corpus (wik2 telecommunicationsse coder2 1255513359128) parmi les principaux organismes de normalisation-standardisation mondiaux, citons : es trigger l’etsi : european telecommunication standards institute ou institut européen des normes de télécommunication ; item 1 l’itu : international telecommunication union ou union internationale des télécommunications ; item 2 l’ietf : internet engineering task force ; item 3 l’atm forum ; item 4 l’ansi : american national standard institute ; item 5 l’ieee : institute of electrical and electronics engineers. item 6 5.4.3 type 3: multiparagraph ess type 3 ess stretch over at least two paragraphs, with no headings or bullets, as illustrated in example 8. this example is a case of nesting: es2 is embedded in es1. the larger es (es1) is type 3, with two paragraph breaks: one between the two items, the other before the closure. this example illustrates the role of paragraph-initial position in the signalling of ess: each item-paragraph 87 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre starts with a sequencer (a first observation [une première observation]; a second observation [une deuxième observation]). these sequencers are echoed by the two item-introducing expressions in the embedded type 4 es (es2) (in the first case [dans le premier cas]; the second position [la seconde position]). example 8 from ling sub-corpus (ling kleiberse coder3 1254142826046) une première observation est à faire à ce niveau. on constate dans l’abondante littérature sur la multiplicité des [...] es1 item 1 une deuxième observation concerne le niveau où s’exerce la critique du fait polysémique. item 2 si on part de la conjonction définitionnelle provisoire i et ii -, la polysémie peut être remise en cause, soit en critiquant i -, soit en critiquant ii -. es2 trigger dans le premier cas, celui où i est faux, mais où ii subsiste, les relations de ii sont à porter au crédit de la construction [...] item 1 la seconde position, celle où l’on conserve i -, mais où l’on refuse ii -, revient à transformer un cas de polysémie en un cas d’homonymie. elle est, c’est significatif, beaucoup moins [...] item 2 nos deux observations tirent dans la même direction : elles montrent que c’est avant tout le point i -, celui de [...] closure [...] precise alignment of elements (trigger, items, closure) with paragraphs is not mandatory for this type. example 9 shows a complete enumerative structure where the trigger and the first two items share one paragraph, while the last item and the closure appear in a separate one. example 9 from geop sub-corpus (geop 19se coder1 1253605625609) par ailleurs, une guerre contre l’irak pouvait se faire selon trois scénarios. es trigger le premier consistait à renouveler l’expérience de 1991 (sans doute avec une coalition amoindrie); il nécessitait des mois de préparation et posait de réels problèmes de politique intérieure aux etats-unis. item 1 le deuxième scénario consistait à répéter l’expérience de décembre 1998, à savoir [...] item 2 le troisième scénario consistait à envoyer sur place plusieurs commandos de services spéciaux chargés de liquider le dictateur [...] item 3 de fait, la première option semblait être la seule permettant de poursuivre l’effort de 1991, en poussant [...] closure type 3 ess are average in length, with slightly above average cardinality. whilst the frequency of triggers is markedly low (61%), closures are much more frequent than elsewhere, particularly in ling (27%) and geop (32%). despite this comparatively high frequency of closures, type 3 includes the highest proportion of minimalist ess: over a third have neither trigger nor closure. these minimalist type 3 ess are characterised by the presence of series of es-cues in paragraphinitial position (see section 5.2.2 above). it may also be noted that only in types 3 and 4 do we find ess which have a closure and no trigger, as es1 in example 8. 5.4.4 type 4: intraparagraph ess type 4 ess, which are contained within a paragraph, are the most frequent. the es in (10) below and the nested structure (es2) in (8) are examples of this type. unsurprisingly, type 4 ess have the smallest mean length (120 words/es); they also have significantly fewer items than other types: over half are 2-item ess (against 29% for type 2, 34% for type 1 and 41% for type 3). the presence of triggers and closures is slightly below average. as a consequence, complete ess are fairly rare (7%), in contrast with minimalist ess (29%), illustrated in example 10. 88 a corpus-driven approach to discourse organisation: from cues to complex markers example 10 from geop sub-corpus (geop 11se coder1 1254301361468) [...] entre 1949 et 1970, la part de la demande couverte par le pétrole importé est passée de 10 % à 23 %. es item 1 entre 1978 et 1985, les importations ont fortement baissé, tant en valeur absolue (3,8 mb / j) que relative (16 points de part de marché). deux facteurs expliquent ce phénomène : le développement du champ géant de prudhoe bay en alaska, et la chute de la demande pétrolière liée au second “choc pétrolier” de 1979 et à la récession économique. item 2 à partir de 1985, la part du pétrole importé dans la couverture de la demande n’a cessé d’augmenter, jusqu’à aujourd’hui. item 3 to summarise, the major variations accounted for by the granularity-based typology are as follows: • type 1 ess are significantly longer; • type 2 ess, with higher cardinality and shorter length, have a trigger most of the time and as a consequence are rarely minimalist ess; • type 3 ess have significantly more often a closure, but are also more minimalist than the others; • type 4 ess are the shortest in length and cardinality with a high proportion of minimalist ess. this typology will be used in the next section to organise the data by distinguishing classes of objects likely to make use of different signalling modes. 6. mining the annotated corpus for configurations of es-cues in this final section, we move on to the search for recurring configurations of es-cues as a way of identifying the complex markers signalling ess. prior to this phase of the analysis, es-cues (i.e. validated features) had to be organised into relevant categories. we added syntactic parallelism, encountered in various forms in examples 1, 4 and 5, as a frequent annotator-added cue which had been identified from the outset as a potential es-cue but could not be pre-marked for technical reasons. we now describe this re-classification, which accounts for the differences between table 1 (section 3) and table 6 below. in addition to the main annotation task, identifying ess and their constitutive elements, the annotators were asked to mark up the cues which they identified as signalling these structures (cf. 3.1.2). the resulting corpus contains 4,052 individual annotated cues which were explicitly identified as es-cues either through pre-marked feature validation or through manual addition; to these must be added 500 headings, systematically counted as es-cues when occurring in a trigger or when item-initial. it should be stressed that identifying cues is considerably more difficult than identifying ess, and at this stage we have no inter-annotator agreement measure on this task. a number of problems were encountered, some of which originate in the pre-marking procedure—-any text processing program inevitably generates both noise and silence—-, others in the level of linguistic competence required of the annotators, or in semantic difficulties inherent in some of the cues. due to these limitations, our analysis will be restricted to the identification of the global behaviour of es-cues. the goal of this section is twofold: 89 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre 1. to examine frequencies and distributions for the different kinds of es-cues; 2. to identify recurrent cue configurations as a first step towards the definition of es markers (cf. section 3.1). table 6 below lists the categories of es-cues taken into account. the abbreviations in bold are used throughout the remainder of this section. nb of cues description example in trigger 443 triggerlex.: prospective elements and other lexical features ex. 3 (deux types de critique), 5 and 6 302 triggerpunct.: punctuation patterns ex. 3, 5 and 7 in item 595 paral.: syntactic parallelisms ex. 3 (item-initial nps) 628 seq.: sequencers and connectives ex. 5 and 8 649 adv.: circumstance adverbials ex. 10 433 itemhead.: headings at the beginning of items ex. 6 1065 bullets: bullets and numbering ex. 3 246 itempunct.: punctuation patterns 88 itemothers: other lexical features in closure 103 closurelex.: encapsulations, connectives and other lexical features ex. 3: ces deux critiques ex. 9: to sum up / en somme table 6: categories of es-cues for analysis (after re-classification) annotator-added cues, except syntactic parallelisms, were counted as “triggerlex.”, “closurelex.” or “itemothers” according to their host component. we are aware that these categories are excessively broad. we will in particular need to isolate prospective elements and encapsulations in order to investigate the expression of the co-enumerability criterion. a semantic characterisation of the expression of the co-enumerability criterion is required for a finer functional classification of ess. 6.1 description of cues in es components tables 7 and 8 provide the detail of the distribution of annotated cues for each component. distributions are given both globally and according to es types. all values are percentages, and are relative to the frequency of the corresponding element: out of the 131 ess with a closure (i.e. out of 13.2% of ess, cf. table 5) 78.6% have a lexical cue. percentages do not add up to 100 for triggers and items, as they each can have between zero and several cues (of different kinds). 6.1.1 trigger and closure cues trigger and closure are almost systematically signalled by a cue: over 75% of these components were associated by the annotators with at least one es-cue. two categories vary considerably in frequency across types: explicit lexical elements, which potentially announce the co-enumerability criterion (triggerlex.), and punctuation marks (a final 90 a corpus-driven approach to discourse organisation: from cues to complex markers trigger closure % with cue es type (nb.) nb. triggerlex. triggerpunct. nb. % with closurelex. type 1 (126) 106 41.5 3.8 5 80.0 type 2 (244) 235 55.3 76.6 32 68.8 type 3 (216) 131 74.8 19.8 45 75.6 type 4 (405) 267 64.0 34.5 49 87.8 all types (991) 739 59.9 40.9 131 78.6 table 7: distribution of trigger and closure cues colon, triggerpunct.), which have a purely textual role. characteristic trigger punctuation is most frequent in type 2 ess, part of a well-established pattern for introducing lists, seen in examples 3, 5 and 7. punctuation cues are also fairly frequent in type 4 ess: example 11 illustrates how punctuation is instrumental in signalling such intraparagraph ess, with a colon as a trigger cue, and final commas reinforcing the parallelism between items. example 11 from geop sub-corpus (geop 27se coder2 1282829750411) [...] les phénomènes terroristes prolifèrent au croisement de quatre grandes circulations : es trigger celle des mots et des images (qui permet de bricoler des solidarités entre des sociétés très différentes), item 1 celle des capitaux (qui autorise la mise sur pied de logistiques performantes), item 2 celle des armes (qui ouvre sans cesse le champ des dangers futurs), item 3 et celle des hommes. item 4 [...] closures are characterised by the strong presence of lexical cues (closurelex.). it must however be kept in mind that this component is rare in our annotated data (cf. table 3), with the consequence that these percentages correspond to very few cases and the results cannot be extrapolated. 6.1.2 item cues table 8 shows that all predefined categories of item cues were indeed found in the annodis resource, and that conversely very few item cues are found in the “itemothers” miscellaneous category. over 15% of ess contain at least one of the most common lexical cues i.e. a sequencer, an adverbial or a parallelism. among these lexical cues, only parallelisms are distributed more or less equally in all es types. as a consequence, parallelism is the cue which combines most frequently with the visual cues inherent to type 1 and 2 ess (itemhead. or bullets). the other lexical cues are on the contrary extremely rare in type 1 and type 2. this lack of variety in the signalling of items creates a stark contrast between types 1 and 2 on the one hand, and types 3 and 4 on the other, the latter displaying a greater complexity of organisation associated with a wide range of cues. the distribution of item cues in types 3 and 4 presents an interesting contrast: type 3 ess favour circumstance adverbials over sequencers (47.8% and 26.3% respectively), whereas in type 4 sequencers (34.4%) prevail over adverbials (19.3%). an explanation for this difference may be found in the organising role of circumstance adverbials (cf. 1.2): in order to function as discourse segmentation markers, these must be paragraph-initial, as in example 12 below where 3 place adverbials occupy paragraph-initial position (in the united states [aux états-unis], in germany [en allemagne], in spain [en espagne]. 91 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre es type nb of % with items itemhead. bullets itempunct. seq. adv. paral. itemothers type 1 434 100.0 0.5 0.0 1.6 6.0 17.1 4.1 type 2 995 0 100.0 na 2.0 2.1 16.8 2.4 type 3 802 0 7.2 2.6 26.3 47.8 15.6 4.2 type 4 1134 0 5.3 16.8 34.4 19.3 20.2 1.1 all types 3365 13.1 31.1 7.3 18.7 19.3 17.7 2.6 table 8: distribution of item cues (cue category per es type) example 12 from wiki sub-corpus (wik2 attentats11septse coder1 1254125810843) aux états-unis, la seule personne à avoir été jugée jusqu’à présent pour son implication directe avec les attentats du 11 septembre est le français zacarias moussaoui. arrłté moins d’un mois avant les attaques, il a été accusé par les autorités fédérales américaines d’avoir eu connaissance des attentats à venir mais de n’avoir pas communiqué ses informations. le 3 mai 2006, au terme de deux mois de procès, il a été reconnu coupable par le jury du tribunal fédéral d’alexandria en virginie de six chefs d’accusation de complot en liaison avec les attentats terroristes du 11 septembre et condamné à la prison à perpétuité, sans possibilité de remise de peine. es item 1 en allemagne, le marocain mounir al-motassadeq arrłté le 28 novembre 2001, est condamné une première fois à quinze ans de prison en 2003 pour complicité dans ces attaques. remis en liberté en février 2006 après que sa condamnation a été cassée, il voit sa première peine confirmée par le tribunal de hambourg le 8 janvier 2007. item 2 en espagne, le syrien imad eddin barakat yarkas, chef de la cellule locale d’al-qaida est arrłté le 13 novembre 2001, inculpé de conspiration en vue des attentats de septembre 2001. il est condamné le 26 septembre 2005 à vingt-sept ans de prison. item 3 this positional constraint is incompatible, by definition, with intraparagraph type 4 ess. sequencers on the other hand are specialised in the signalling of ess, and seem therefore to be more independent from positional constraints, which could explain why they are particularly suited to these low-level ess, as illustrated by example 8. in addition, punctuation item cues are fairly frequent in type 4, as are punctuation trigger cues. example 13 illustrates such a combination of punctuation and lexical cues in a type 4 es where items are separated by a semicolon and the last one introduced by the connective enfin / finally. example 13 from geop sub-corpus (geop 16se coder1 1255425907703) [...] mais les discussions sont occultées par des positions idéologiques : ‘la subvention est intrinsèquement néfaste’, ‘la pac est intouchable’, ‘les ped sont quoi qu’il arrive victimes d’un système injuste’ ... positions contredites par les pratiques. es trigger tout le monde subventionne, même les pays les plus vertueux, d’une façon qui peut fausser les échanges ; item 1 la pac est en constante révision, et son coût n’est pas élevé (0,5 % du pib européen) ; item 2 enfin, il est faux que les ped aient tout à gagner d’une disparition totale des subventions, tant est grand l’avantage comparatif des plus gros producteurs agricoles, qui ne sont pas des ped. item 3 6.2 cue associations the examples make it clear that most ess are signalled concurrently by several kinds of es-cues. the previous section gave an insight into the frequencies of individual es-cues without taking into 92 a corpus-driven approach to discourse organisation: from cues to complex markers account their co-occurrence. yet our hunch, as stressed in section 2, is that textual patterns are not just signalled by discrete clearly identifiable dedicated markers, but by configurations of es-cues functioning as complex discourse markers. in such a perspective, a discourse function should not be attributed to a particular lexical expression—-on the basis of a specific semantic or pragmatic value—-but rather to this expression when it occurs in a particular context or configuration (cf. the pattern formed by the series of paragraph-initial adverbials in example 1). the annodis resource now provides us with data to investigate this hypothesis, and this is what we attempt below using the notion of cueset. this is not an easy task, however, as we want to allow for flexibility while hoping to catch recurring patterns. the identification of textual patterns when reading can be conceptualised in terms of pattern recognition: a threshold is reached when there are enough converging cues to push interpretation towards the identification of the pattern in question. a more satisfactory approach to cuesets would involve attributing weights to individual es-cues in order to account for the fact that several weak cues may do the same work as one strong cue. this is our horizon, with the work presented here as a first exploration of the data in this direction. a cueset is the set of cue categories occurring in an es. as the purpose of these sets is to help identify frequent cue associations, we apply the following simplifications: 1. for item cues, a single occurrence suffices for the cue to be included in the set, there is no need for the cue to appear in every item; 2. cue frequency within an es is not taken into account, and is reduced to a simple binary value of presence/absence. the main reasons for these simplifications are the potential incompleteness of item marking (e.g. firstly in the first item not followed by other sequencers), and the inherent difficulty of the cueannotation task. the number of different associations was calculated for the whole collection of ess, and for each es type studied independently. of all theoretically possible configurations23, over half were actually observed: 113 distinct cuesets were identified for all 991 ess. this result is interpreted as meaning, on the one hand, that ess are signalled by a variety of cue configurations—-for example a lexical cue in the trigger followed by a series of sequencers, or a combination of sequencers and adverbials; and on the other hand, that certain specific cue associations recur, while others are not found. in order to identify the most frequent cuesets, we focus here on the 14 cuesets occurring at least 20 times across types and corpora, which represent 63% of all ess. among the most frequent cuesets, we find clusters made of cues which have not been the focus of much attention in studies of enumerating i.e bullets, punctuational patterns and more interestingly lexical cues in the trigger (cf. example 7). in almost all frequent cuesets there is at least one trigger cue (a punctuational or a lexical one) associated with all possible item cues (i.e. headings, bullets, sequencers, adverbials, parallelisms). the most frequent cueset is the combination triggerpunct.+triggerlex.+bullets which occur 83 times i.e. in 8.4% of ess. the fairly similar cueset triggerpunct. + bullets recur only 40 times, which means that visual cues are usually combined with lexical ones. the same kind of combination is found with the cueset triggerpunct.+triggerlex.+ itempunct. this finding supports our view of signalling as a struggle between different forces (cf. section 2.2): 23. there are 210 theoretically possible configurations since all 10 es-cues listed in table 6 may signal ess. 93 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre visual devices can be seen as pulling enumerative structures towards the textual component while lexico-syntactic cues seem better able to contribute to the ideational component. cuesets including the much-studied sequencers are also very frequent. two types were observed with approximately equal frequency: cuesets composed of sequencers only as in example 8 (73 cuesets, 7.4% of ess); and cuesets which combine sequencers with other cues such as a lexical cue in the trigger (as in example 9), parallelisms or adverbials (87 cuesets, 8.8% of ess). cuesets made up purely of adverbials (examples 9 and 11) are as frequent as those made up purely of sequencers (74 cuesets, 7.5% of ess). but cuesets mixing adverbials with other kind of cues are fairly rare (only 23 with lexical cue in the trigger and 20 with sequencers). all the cuesets described are fairly stable across sub-corpora except for two: those made up of itemhead. and paral. which only recur in scientific papers and those made up of adverbials which occur primarily in encyclopaedia articles and never in scientific papers. a number of specific configurations have been identified, which, when correlated with es types, can be summarised as follows: • type 1 ess are typically signalled by a sequence of same level headings, with an upper level heading acting as a trigger, and occurs in documents where layout and visual formatting play a prominent role. • type 2 ess are typically signalled by a sequence of bulleted items, almost systematically introduced by a punctuational cue (final colon in the preceding paragraph), and/or a lexical cue in the trigger which may carry semantic information on the co-enumerability criterion. • type 3 ess are typically signalled by a contiguous series of paragraphs with circumstance adverbials (or, less likely, sequencers) in initial position, which have both a textual and an ideational role. such structures seem to be highly genre-sensitive. • type 4 ess can be described as a single paragraph containing a series of sequencers, with a high probability of a colon marking the end of the trigger, or a prospective element indicating the co-enumerability criterion. in contrast with type 2, type 4 ess reflect the ideational dimension of discourse organisation more than the textual dimension. 6.3 towards complex discourse organisation markers identification in order to validate and formalise the cuesets observed above, we used a common data mining technique for identifying recurrent associations between pairs of cues by extracting the association rules i.e. the logical implication rules between cues (agrawal et al., 1993). this method ensures that all possible cases are systematically examined. the association rules are of the form: x1 & x2 ... xn→ y (where xi and y are cue types), meaning that most (at least 75%) of ess having x1, x2 and xn as cues also have y. this technique identified the following rules (in order of decreasing systematicity): 1. bullets→ triggerpunct. 2. itempunct.→ triggerpunct. 3. triggerpunct. & bullets→ triggerlex. 94 a corpus-driven approach to discourse organisation: from cues to complex markers 4. itemhead.→ triggerlex. 5. closurelex. → triggerlex. rules 1 and 2 indicate that whenever the items are bulleted (type 2 es) or have a punctuation mark (essentially type 4 ess), they have a trigger with a punctuation cue (colon), and vice-versa. this reflects the coherence of patterns of punctuation marks in ess. rules 3 and 4 merely confirm that ess of types 1 and 2 have a high proportion of triggers (table 5), and that most of these triggers contain a lexical cue (table 7). such a finding can be interpreted as showing that even in apparently purely visual i.e. textual ess, a lexical cue somewhere will ensure the presence of the ideational dimension. rule 5 links the existence of an encapsulation to that of a prospective element. again, this must be interpreted with the knowledge that both these cues are quite systematic in triggers and closures. rule 5 says that ess with a closure generally have a trigger. in other words, closures are not used to compensate for the absence of a trigger. by lowering the tolerance of the association rules system, more rules can be made to emerge, although they are known to be much less reliable and systematic. one interesting point is that, even with a low threshold, no rule involving circumstance adverbials emerges. this negative result confirms that this cue category is much less likely to work in association with others. most of these results were predicted and explained in previous sections, which suggests that no other obvious specific cue associations can be identified as a result of our annotation exercise. 7. conclusion enumerative structures were selected as the focus of this study as a way of throwing new light on linearisation and segmentation, discourse phenomena which are particularly difficult to analyse empirically. we described enumerative structures in sections 2 and 3 as a generic multi-level device for organising text. according to our broad functional definition, they are textual patterns assembling text spans which are made to appear as similar in a given respect, thereby forming a higher-level segment homogeneous in this particular respect. they arrange into linear format text segments which are ideationally discontinuous but functionally equivalent and interchangeable. their signalling calls upon a great diversity of cues working together. as such, enumerative structures constitute good handles for analysing how writers cope with “the unbearable linearity of texts”. our objective of proposing a data-intensive methodology for the study of linearisation and segmentation imposed certain requirements in terms of ease of detection and annotation. as their function depends on their being readily detectable, they constitute a good object for annotation. the sizeable annotated resource described here, which has been made available to the research community, is characterised by a number of original features: • it is composed of highly-structured long expository texts in three different genres (as opposed to short news material); • its mark-up combines nlp-based exhaustive techniques and human intuitions; • the visual characteristics of the texts have been encoded so as to provide a presentation respectful of the original layout; 95 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre • the annotation guidelines were designed so as to bring together under a functional umbrella objects that are linguistically more diverse than in previous studies: type 4 (intraparagraph ess) is shown to be just one realisation, accounting for no more than 40% of annotated ess. the paper summarises the results of the first analyses of the annotated data. the evidence-based and quantified typology we propose encapsulates the major results of our analyses so far: it provides a broad picture of a device realising a basic textual strategy. the preliminary analyses presented here only scratch the surface of what the annotated corpus allows. where cues have been counted together, e.g. in the analysis of triggers and closures, finer analyses are needed to take into account the specific contribution of each type of cue, in particular in the case of expressions of the co-enumerability criterion. qualitative studies are necessary for the analysis of the rhetorical and semantic functions of enumerative structures in text, opening the way for the study of correlations between such functions and cue configurations, and for the exploration of the differences between sub-corpora. one important issue concerns the nature of the relation between the items and the “classifier” which introduces and links them. does an enumerative structure reveal a pre-existing categorisation or can it “discursively create” such knowledge, as suggested by schiffrin (1994, p. 396) or luc et al. (2000)? the latter hypothesis, a constructivist one more in keeping with our textual approach, makes the expression of the co-enumerability criterion worth studying as potentially revealing not just of pre-existing knowledge structures, but of a writer’s discourse strategy. a systematic study of expressions of the co-enumerability criterion is under way (rebeyrolle and péry-woodley, 2014). it suggests that only in very few cases are these expressions linked to the enumerated items by a hypernym-hyponym relation, i.e. a taxonomic relation which is discourse-independent (< 10%). first results show that the expressions of the co-enumerability criterion would generally be better described as “text-bound” labels (francis, 1994), in terms of shell nouns or signalling nouns (flowerdew, 2003; flowerdew and forest, 2015). the study develops a model associating semantic and textual properties of linguistic expressions of co-enumerability, proposing that semantic characteristics situate enumerative structures containing these expressions on a cline from mainly textual (metadiscursive) to mainly ideational (stable, discourse-independent categorisation). on this basis, a classification of ess in terms of their discourse function will be put forward, to be compared with existing taxonomies of relevant discourse relations and analyses based on them (e.g. joint, list, sequence in rst). cue configurations should also be examined further in relation to layout (es type) and composition: for example minimalist ess (neither trigger nor closure) are markedly more numerous in types 3 and 4, which is also where adverbials are most frequent as item markers, and may compensate for the absence of expression of the co-enumerability criterion in a prospective or encapsulating element (rebeyrolle and péry-woodley, 2014). we wish to look further into these trade-offs as examples of how ideational and textual metafunctions are interwoven in these structures. ess should also be examined in context, within the linearity of text: interactions between ess (nested and in sequence), interactions between ess and other textual structures (including annotated topical chains). finally, the markers identified should be tested and refined for the automatic detection of ess, with potential applications in automatic text synthesis and document navigation. 96 a corpus-driven approach to discourse organisation: from cues to complex markers references stergos afantenos, nicholas asher, farah benamara, myriam bras, cécile fabre, lydia-mai hodac, anne le draoulec, philippe muller, marie-paule péry-woodley, laurent prévot, josette rebeyrolle, ludovic tanguy, marianne vergez-couret, and laure vieu. an empirical resource for discovering cognitive principles of discourse organisation: the annodis corpus. in nicoletta calzolari, khalid choukri, thierry declerck, mehmet uğur doğan, bente maegaard, joseph mariani, jan odijk, and stelios piperidis, editors, proceedings of the eight international conference on language resources and evaluation (lrec’12), pages –, istanbul, turkey, may 2012. european language resources association (elra). url https: //hal.archives-ouvertes.fr/hal-00976087. rakesh agrawal, imielinski tomasz, and arun swami. mining association rules between sets of items in large databases. acm sigmond record, 22(2):207216, 1993. nicholas asher. reference to abstract objects in discourse. kluwer, 1993. nicholas asher, farah benamara, myriam bras, lydia-mai ho-dac, and philippe muller. annodis and related projects: case studies on the annotation of discourse structure. in nancy ide and james pustejovsky, editors, the handbook of linguistic annotation. springer, berlin, to appear. john bateman, thomas kamps, jörg kleinz, and klaus reichenberger. towards constructive text, diagram, and layout generation for information presentation. computational linguistics, 27(3): 409–449, 2001. douglas biber. variation across speech and writing. cambridge university press: cambridge, massachusset, 1988. douglas biber, ulla connor, and thomas a. upton. discourse on the move, using corpus analysis to describe discourse structure, volume 28 of studies in corpus linguistics. john benjamins publishing company: amsterdam/philadelphia, 2007. didier bourigault. un analyseur syntaxique opérationnel : syntex. mémoire d’hdr, université de toulouse, 2007. didier bourigault, cécile fabre, cécile frérot, marie-paule jacques, and silvia ozdowska. syntex, analyseur syntaxique de corpus. in actes des 12èmes journées sur le traitement automatique des langues naturelles, 2005. thorsten brants. inter-annotator agreement for a german newspaper corpus. in proceedings of the second international conference on language resources and evaluation (lrec), athens, greece, 2000. myriam bras and catherine schnedecker. dans un (premier+second+nième) temps vs en (premier+second+nième) lieu : qu’est ce qui fait la différence? langue française, 179:89–108, 2013. myriam bras, laurent prévot, and marianne vergez-couret. quelles relations de discours pour les structures énumératives? in bernard laks jacques durand, benot habert, editor, congrès 97 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre mondial de linguistique française cmlf’08, pages 1945–1964. edp sciences, institut de linguistique française, paris, 2008. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in jan van kuppevelt and ronnie smith, editors, current directions in discourse and dialogue, pages 85–112. kluwer academic publishers, 2003. wallace l. chafe. subject and topic, chapter givenness, contrastiveness, definiteness, subjects, topics, and point of view, pages 25–55. new york/san francisco/london: academic press, 1976. wallace l. chafe. discourse consciousness and time: the flow and displacement of conscious experience in speaking and writing. university of chicago press: chicago, 1994. michel charolles, anne le draoulec, marie-paule péry-woodley, and laure sarda. temporal and spatial dimensions of discourse organisation. journal of french language studies, 15(2):203– 218, 2005. maud colléter, cécile fabre, ho-dac lydia-mai, marie-paule péry-woodley, josette rebeyrolle, and ludovic tanguy. la ressource annodis multi-échelle : guide d’annotation et bonus. carnets de grammaires, 20:1–63, 2012. url https://hal.archives-ouvertes.fr/ hal-00983076. marie-elisabeth conte. anaphoric encapsulation. belgian journal of linguistics: coherence & anaphora, 10:1–9, 1996. nils erik enkvist. a parametic view of word order. in e.szer, editor, text connexity text coherence: aspects methods results, pages 320–336. helmut buske: hamburg, 1985. john flowerdew. signalling nouns in discourse. english for specific purposes, 22(4):329–346, 2003. john flowerdew and richard w forest. signalling nouns in academic english. cambridge university press, 2015. gill francis. labelling discourse: an aspect of nominal-group lexical cohesion. in m. coulthard, editor, advances in written text analysis, pages 83–101. london & new york: routledge, 1994. morton ann gernsbacher. the structure building framework: what it is, what it might also be, and why. in b. k. britton and a. c. graesser, editors, models of text understanding, pages 289–311. lawrence erlbaum associates: mahwah, new jersey, 1995. morton ann gernsbacher. two decades of structure building. discourse processes, 23(3):265–304, 1997. dionysis goutsos. modeling discourse topic: sequential relations and strategies in expository text, volume 59. greenwood publishing group, 1997. dyonisos goutsos. a model of sequential relations in expository test. text, 16(4):501–533, 1996. 98 a corpus-driven approach to discourse organisation: from cues to complex markers michael a.k. halliday. text as semantic choice in social contexts. in j. webster, editor, the collected works of m.a.k. halliday (volume 2): linguistic studies of text and discourse, page 2381. london: continuum, 1977/1983. hilde hasselgård. adjunct adverbials in english. cambridge: cambridge university press, 2010. susanne hempel and liesbeth degand. sequencers in different text genres: academic writing, journalese and fiction. journal of pragmatics, 40:676–693, 2008. laurent heurley. processing units in written texts: paragraphs or information blocks? in j. costermans and m. fayol, editors, processing interclausal relationships: studies in the production and comprehension of text, pages 179–200. lawrence erlbaum associates: mahwah, new jersey, 1997. lydia-mai ho-dac and marie-paule péry-woodley. a data-driven study of temporal adverbials as discourse segmentation markers. discours, 4:http://discours.revues.org/5952, june 2009. doi: 10. 4000/discours.5952. url https://hal.archives-ouvertes.fr/hal-00979739. lydia-mai ho-dac and marie-paule péry-woodley. annotation des structures discursives : l’expérience annodis. in franck neveu, peter blumenthal, linda hriba, annette gerstenberg, judith meinschaefer, and sophie prévost, editors, 4e congrès mondial de linguistique française (cmlf 2014), pages 2647 – 2661, berlin, germany, july 2014. doi: 10.1051/shsconf/ 20140801286. url https://hal.archives-ouvertes.fr/hal-01068119. lydia-mai ho-dac, marie-paule péry-woodley, and ludovic tanguy. anatomie des structures énumératives. in traitement automatique des langues naturelles, page (publication numérique), montréal, canada, 2010. url https://halshs.archives-ouvertes. fr/halshs-00509189. lydia-mai ho-dac, cécile fabre, marie-paule péry-woodley, josette rebeyrolle, and ludovic tanguy. an empirical approach to the signalling of enumerative structures. discours, 10:(publication en ligne), 2012. url https://halshs.archives-ouvertes.fr/ halshs-00954182. michael hoey. signalling in discourse. discourse analysis monographs 6. english language research, birmingham university, birmingham, 1979. eduard h. hovy. parsimonious and profligate approaches to the question of discourse structure relations. in proceedings of the fifth international workshop on natural language generation, pages 128–136, dawson, pa, june 1990. george hripcsak and adam s. rothschild. agreement, the f-measure, and reliability in information retrieval. journal of the american medical informatics association: jamia, 12(3):296–298, 2005. agatha jackiewicz. les séries linéaires dans le discours. langue française, 148:95–110, 2005. julie lemarié, robert frederick lorch, hélène eyrolle, and jacques virbel. sara: a text-based and reader-based theory of signaling. educational psychologist, 43(1):27–48, 2008. 99 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre julie lemarié, robert frederick lorch, and m.-p. péry-woodley. understanding how headings influence text processing. discours, special issue on signalling discourse organisation – multidisciplinary approaches to discourse 2010 (mad 10), 10, 2012. url http://discours. revues.org. willem j.m. levelt. the speaker’s linearization problem. philosophical transactions royal society london, b295:305–315, 1981. robert frederick lorch and elizabeth pugzles lorch. effects of organizational signals on free recall of expository text. journal of educational psychology, 88(1):38, 1996. christophe luc and jacques virbel. le modèle d’architecture textuelle : fondements et expérimentation. verbum, 23(1):103–123, 2001. christophe luc, mustapha mojahid, jacques virbel, claudine garcia-debanc, and marie-paule péry-woodley. a linguistic approach to some parameters of layout: a study of enumerations. in aaai 1999 fall symposia ”using layout for the generation, understanding or retrieval of documents, pages 20–29, north falmouth, massachussetts, 1999. christophe luc, mustapha mojahid, marie-paule péry-woodley, and jacques virbel. les énumérations : structures visuelles, syntaxiques et rhétoriques. in actes de cide 2000 (colloque international sur le document électronique), pages 21–40, 2000. william c mann and sandra a thompson. rhetorical structure theory: toward a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8(3):243–281, 1988. daniel marcu. the rhetorical parsing of unrestricted texts: a surface-based approach. computational linguistics, 26(3):395–448, 2000. daniel marcu. automatic discourse parsing. in k.brown, editor, encyclopedia of language and linguistics, pages 649–654. elsevier, oxford, 2nd edition, 2006. philippe muller, marianne vergez-couret, laurent prévot, nicholas asher, farah benamara, myriam bras, anne le draoulec, and laure vieu. manuel d’annotation en relations de discours du projet annodis. technical report 21, carnets de grammaires, clle-erss, 2012. geoff nunberg. the linguistics of punctuation. csli, lecture notes, university of chicago press, 18, 1990. marie-paule péry-woodley, stergos afantenos, lydia-mai ho-dac, and nicholas asher. le corpus annodis, un corpus enrichi d’annotations discursives. tal, 52(3):71–101, 2011. sophie piérard and yves bestgen. validation d’une méthodologie pour l’étude des marqueurs de la segmentation dans un grand corpus de textes. tal, 47(2):89–110, 2006. richard power, donia scott, and nadjet bouayad-agha. document structure. computational linguistics, 2(29):211–260, 2003. 100 a corpus-driven approach to discourse organisation: from cues to complex markers rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in nicoletta calzolari, khalid choukri, bente maegaard, joseph mariani, jan odjik, stelios piperidis, and daniel tapias, editors, proceedings of the sixth international language resources and evaluation (lrec’08), marrakech, morocco, may 2008. european language resources association (elra). url http: //www.lrec-conf.org/proceedings/lrec2008/. josette rebeyrolle and marie-paule péry-woodley. énumration et structuration discursive. in franck neveu, peter blumenthal, linda hriba, annette gerstenberg, judith meinschaefer, and sophie prévost, editors, 4e congrès mondial de linguistique française (cmlf 2014), pages 3183–3196, berlin, germany, july 2014. deborah schiffrin. making a list. discourse processes, 17(3):377405, 1994. hans-jörg schmid. english abstract nouns as conceptual shells: from corpus to cognition. berlin, new york: mouton de gruyter, 2000. helmut schmid. treetagger— a language independent part-of-speech tagger. institut für maschinelle sprachverarbeitung, universität stuttgart, 43:28, 1995. john sinclair. planes of discourse. the twofold voice: essays in honour of ramesh mohan. salzburg: universitt salzburg., 1983. maite taboada and debopam das. annotation upon annotation: adding signalling information to a corpus of discourse relations. dialogue and discourse, 4(2):249–281, 2013. maite taboada and william c mann. rhetorical structure theory: looking back and moving ahead. discourse studies, 8(3):423–459, 2006. angela tadros. prediction in text. number 10. english language research, 1985. gilbert turco and danielle coltier. des agents doubles de l’organisation textuelle, les marqueurs d’intégration linéaire. pratiques, 57:57–79, 1988. teun a van dijk. story comprehension: an introduction. poetics, 9(1):1–21, 1980. marianne vergez-couret, laurent prévot, and myriam bras. interleaved discourse, the case of two-step enumerative structures. in constraints in discourse iii, 2008. marianne vergez-couret, myriam bras, laurent prévot, laure vieu, and caroline atallah. discourse contribution of enumerative structures involving” pour deux raisons”. in constraints in discourse 2011, page online, 2011. marianne vergez-couret, laurent prévot, and myriam bras. how different information sources interact in the interpretation of interleaved discourse: the case of two-step enumerative structures. discours, 11:(publication en ligne), 2012. doi: 10.4000/discours.8743. url http: //discours.revues.org/8743. jacques virbel. the contribution of linguistic knowledge to the interpretation of text structures. in structured documents, pages 161–180. cambridge university press, 1989. 101 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre jacques virbel. textual enumeration. in texts, textual acts and the history of science, pages 221–266. springer, 2015. jacques virbel, claudine garcia-debanc, thierry baccino, lartitia carrio, corinne dominguez, christian jacquemin, christophe luc, mustapha mojahid, marie-paule péry-woodley, and sabine schmid. approches cognitives de la spatialisation du langage. de la modélisation de structures spatiolinguistiques des textes l’expérimentation psycholinguistique : le cas d’un objet textuel, l’énumération. cognitique, agir dans l’espace, éditions de la maison des sciences de l’homme, 2005. tuija virtanen. discourse functions of adverbial placement in english: clause-initial adverbials of time and place in narratives and procedural place descriptions. abo akademi university press: abo, 1992. tuija virtanen. point of departure: cognitive aspects of sentence-initiale adverbials. in t. virtanen, editor, approaches to cognition through text and discour, pages 79–97. berlin/new york: mouton de gruyter, 2004. antoine widlöcher and yann mathet. the glozz platform: a corpus annotation and mining tool. in proceedings of the 2012 acm symposium on document engineering, pages 171–180. acm, 2012. florian wolf, edward gibson, amy fisher, and meredith knight. discourse graphbank. linguistic data consortium, philadelphia, 2004. 102 a corpus-driven approach to discourse organisation: from cues to complex markers appendix nb1. the aim of this appendix is to give readers access to the french-language examples so they can follow the arguments developed in the text. we have therefore opted for translations which remain close to the original french, sometimes at the expense of the quality of the resulting english. nb2. many of the examples given cover in reality large stretches of text and have been cut, sometimes extensively, for the sake of clarity and brevity (cuts are indicated by [...] in the text). the reference associated with each example is its identifier in the annodis resource. it can be searched for in the resource using the annodis browser (http://redac.univ-tlse2.fr/corpus/ annodis/me_download/annodis_se.xml). example 3 from wiki sub corpus (wik2 libertese coder3 1254325598390) against (the idea of) freedom as independence, there are at least two types of criticisms: a moralistic criticism: this freedom would be a form of licentiousness, i.e. surrender to ones desires. now, there is no freedom without law (rousseau, emmanuel kant), as freedom for everyone would be a contradiction in terms: [...] one notices that in this philosophical conception of freedom, the limits are not limits that constrain the freedom of human will; a deterministic criticism: is surrendering to ones desires not a way to obey them? and does such a surrender not amount in the end to a hidden form of determinism? we would in this case be under an illusion of free choice: [...] nietzsche picks up this criticism: as long as we do not feel dependent on something, [...] these two criticisms throw light on several important points. [...] example 5 from ling sub-corpus (ling kleiberse coder3 1254143156093) 2.2 two ways to negate polysemy a possible reply is [...]. [...] polysemy as the association of several meanings to one lexical form is negated in two apparently paradoxical ways: -aon the one hand, word forms given as polysemous become in a way “monosemised” by [...] -bon the other hand, in a manner which is quite the opposite of the monosemisation attempts, meanings are made to proliferate [...] positions -aand -bare only apparently paradoxical: there is no contradiction between, on the one hand [...] example 6 from wiki sub-corpus (wik2 julescesarse coder2 1254907327695) vi. cæsar’s amorous conquests vi.1. women from roman high society according to the roman historian suetonius, cæsar won over many women in the course of his life, in particular women belonging to roman high society. it is said that he won the love of postumia, servius sulpicius’ wife, of lollia, [...] 103 péry-woodley, ho-dac, rebeyrolle, tanguy, fabre cæsar had a special relationship with serviliacaepionis, [...] evidence of cæsar’s taste for the pleasure of love also comes from [...] vi.2. queens cæsar had an affair with euno, wife of bogud, the king of mauritania. however, his relationship with cleopatra has remained most famous. example 7 from wiki sub-corpus (wik2 telecommunicationsse coder2 1255513359128) among the major international normalisation-standardisation bodies, let us mention: • l’etsi: european telecommunication standards institute ou institut europen des normes de tlcommunication; • l’itu: international telecommunication union ou union internationale des tlcommunications; • l’ietf: internet engineering task force; • l’atm forum; • l’ansi: american national standard institute; • l’ieee: institute of electrical and electronics engineers. example 8 from ling sub-corpus (ling kleiberse coder3 1254142826046) a first observation must be made at this stage. one notices in the abundant literature on the multiplicity of [...] a second observation concerns the level at which the criticism of the polysemous fact is conducted. if one starts from the temporary definitional conjunction -iand -ii-, polysemy can be questioned, either by a criticising -i, or by criticising -ii-. in the first case, where -iis false but -iisubsists, the relations in -iican be credited to the construction [...] the second position, conserving -ibut refusing -ii-, amounts to transforming a case of polysemy into a case of homonymy. this position is, and this is significant, much less [...] our two observations go in the same direction: they show that it is first and foremost point -i, the one [...] example 9 from geop sub-corpus (geop 19se coder1 1253605625609) besides, a war against iraq could take place according to three scenarios. the first consisted in reproducing the 1991 campaign (probably with a reduced coalition); it required months of preparation and confronted the us with real domestic policy problems. the second scenario consisted in repeating the 1998 campaign, i.e. [...] the third scenario consisted in sending several special services commandos there, with the order to remove the dictator [...] indeed, the first option seemed to be the only one that made it possible to pursue the 1991 effort, while pushing [...] 104 a corpus-driven approach to discourse organisation: from cues to complex markers example 10 from geop sub-corpus (geop 11se coder1 1254301361468) [...] between 1949 and 1970, [...] ; [...] the share of demand covered by imported oil rose from 10% to 23%. between 1978 and 1985, imports went down sharply, in absolute terms (-3.8mb/d) as well as in relative terms (-16 points of market share). two factors explain this phenomenon: the development of the giant oil field at prudhoe bay in alaska, and the fall in oil demand linked to the second “oil crisis” in 1979 and to the economic recession. since 1985, the share of imported oil in covering demand has continuously increased up to today. example 11 from geop sub-corpus (geop 27se coder2 1282829750411) [...] terrorist events proliferate at the crossroad between four great circulations: the circulation of words and images (which makes it possible to cobble together solidarities between very different social groups), the circulation of capital (which allows the setting up of efficient logistics), the circulation of weapons (which keeps opening up the prospects for future dangers), and the circulation of men. [...] example 12 from wiki sub-corpus (wik2 attentats11septse coder1 1254125810843) in the us, the only person until now to have been judged for direct implication in the 9/11 attacks is the frenchman zacarias moussaoui. arrested less than a month before the attacks, he has been accused by the american federal authorities of having had knowledge of the forthcoming attacks and of not having communicated his information. on may 3rd 2006, after a two-month trial, he was found guilty by the jury of the federal tribunal of alexandria in virginia on six charges of conspiracy linked to the terrorist attacks of september 11th and sentenced to life imprisonment without the possibility of parole. in germany, the morrocan mounir al-motassadeq, arrested on november 28th 2001, is sentenced first to 15 years in prison in 2003 for complicity in these attacks. freed in february 2006 after his conviction was quashed, he saw his initial sentence confirmed by the tribunal of hamburg on january 8th 2007. in spain, the syrian imad eddin barakat yarkas, chief of the local al-qaida cell, is arrested on november 13th 2001, charged with conspiring towards the september 2001 attacks. on september 26th 2005 he receives a twenty-seven year prison sentence. example 13 from geop sub-corpus (geop 16se coder1 1255425907703) [...] but the discussions are obscured by ideological positions: “subsidies are intrinsically bad”, “you can’t touch the cap”, “developing countries will always be the victims of an unfair system” ... positions which are not borne out by practices. all countries, even the most virtuous, use subsidies in a manner which may distort trade; the cap is constantly under revision, and its cost is not high (0.5% of european gnp); finally, it is not true that developing countries stand to gain from a total end to subsidies, given the massive comparative advantage of the biggest agricultural producers, which are not developing countries. 105 microsoft word d&d revision4_final.docx dialogue & discourse 11(1) 1-39 ©2020 clare patterson and petra b. schumacher this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). the timing of prominence information during the resolution of german personal and demonstrative pronouns clare patterson clare.patterson@uni-koeln.de university of cologne department of german language and literature i, linguistics petra b. schumacher petra.schumacher@uni-koeln.de university of cologne department of german language and literature i, linguistics editor: massimo poesio submitted 12/2018; accepted 02/2020; published online 03/2020 abstract german personal and demonstrative pronouns have distinct preferences in their interpretation; personal pronouns are more flexible in their interpretation but tend to resolve to a prominent antecedent, while demonstratives have a strong preference for a non-prominent antecedent. however, less is known about how prominence information is used during the process of resolution, particularly in the light of twostage processing models which assume that reference will normally be to the most accessible candidate. we conducted three experiments investigating how prominence information is used during the resolution of gender-disambiguated personal and demonstrative pronouns in german. while the demonstrative pronoun required additional processing compared to the personal pronoun, prominence information did not affect resolution in shallow conditions. it did, however, affect resolution under deep processing conditions. we conclude that prominence information is not ruled out by the presence of stronger resolution cues such as gender. however, the deployment of prominence information in the evaluation of candidate antecedents is under strategic control. keywords: prominence, anaphora, pronoun resolution, demonstrative, shallow processing, eye-tracking 1 introduction personal pronouns are referentially ambiguous, but they tend to refer to antecedents that are prominent in the prior discourse. this generalisation has led to a great deal of interest in how an entity in discourse becomes more or less prominent. many different linguistic and non-linguistic factors have been found to contribute to prominence, including (but not limited to) grammatical role (crawley & stevenson, 1990; crawley et al., 1990; gordon et al., 1993; järvikivi et al., 2005; kaiser & trueswell, 2008), thematic role (schumacher et al., 2016; stevenson et al., 1994) and its relation to implicit causality (garvey, caramazza & yates, 1974), word order (clark & sengul, 1979; gernsbacher & hargreaves, 1988; järvikivi et al., 2005), and information structure (almor, 1999; colonna et al., 2012; kaiser & trueswell, 2008). many languages have a rich set of pronominal devices which are sensitive to different factors. for example, spanish, italian, turkish, japanese and greek have null pronouns which tend to refer to prominent entities, but also have one or more overt pronouns that refer to less prominent antecedents.1 german, like dutch and norwegian, does not have null pronouns but does have various demonstratives that can function like personal pronouns by referring to animate entities. while personal pronouns in german tend to refer to prominent antecedents and 1 although the division of labour between null and overt pronouns appears to differ between languages (filiaci et al., 2014). doi: 10.5087/dad.2020.101 patterson and schumacher 2 demonstratives to less prominent ones, the precise division of labour between german personal and demonstrative pronouns is under-researched and still under debate. in this paper we focus on the german personal pronoun “er” and one demonstrative pronoun, “der”. it has long been acknowledged that different anaphoric lexical forms signal different antecedent preferences. for example, gundel et al. (1993) developed the givenness hierarchy, which relates the use of anaphoric forms to a particular function based on the cognitive status of the antecedent. in this hierarchy, the use of demonstratives signals antecedents that are activated but not in focus, whereas unstressed personal pronouns or zero pronouns refer to something that is in focus (related to the notion of discourse topic). furthermore, work on english demonstratives has shown that demonstrative pronouns such as this differ from it not only in the cognitive status or salience of the referent but also the type (and also complexity) of the referent (brown-schmidt et al., 2005; çokal et al., 2018), and links to the idea of a form-specific account (kaiser & trueswell 2008). this approach specifies that different referential forms not only seek different types of antecedent but are also differently sensitive to factors affecting the salience or cognitive status of an antecedent, and that accounting for processing of different forms is not possible with a single scale. however, in the current study, we restrict ourselves to two pronominal forms that are both used as independent pronouns referring to animate np entities. the challenge lies accounting for the resolution preferences and division of labour of these particular forms, which is assumed to be related to prominence. 1.1 personal and demonstrative pronouns in german except for certain instances in colloquial speech, german requires overt pronouns. the least marked form is the unstressed personal pronoun, which is inflected for gender and number as well as case (e.g., nominative singular: er/sie/es; nominative plural: sie). more marked variants include the stressed personal pronoun and three different types of demonstrative pronouns (der/die/das, dieser/diese/dieses, jener/jene/jenes). the latter convey greater distance between the speaker and the referent (however this contrast is diminishing; himmelmann, 1997). the availability of the two proximal forms has been associated with differences in pragmatic function (such as contrast and topic marking) among others (ahrenholz, 2007). crucially, demonstrative pronouns in german – unlike english this – productively refer to animate entities (see example [1] below).2 furthermore, the demonstrative pronoun under investigation here refers to entity-denoting referents and not to propositional content, contrary to this and that in english (see çokal et al., 2018; see also brownschmidt et al., 2005 for reference to composite entities). the german demonstrative pronoun for propositional content has a different morphology (dies, das). in the following, we concentrate on the unstressed personal pronoun er and the demonstrative pronoun der in german.3 to illustrate the different resolution preferences of the two pronoun types, consider the minidiscourse shown in (1) (adapted from bosch & umbach, 2007). the second sentence could contain a personal pronoun, shown in bold in (1a), or a demonstrative pronoun, shown in bold in (1b): (1) serena wollte mit venus tennis spielen. serena wanted to play tennis with venus. (1a) doch sie war krank. but she was sick. (1b) doch die war krank. but she was sick. 2 in english it is possible to use a demonstrative such as “this” to refer to human or animate entities in very limited contexts, when they serve as subjects of the verb “to be” in specification contexts such as “this is the woman who saved my life”, but in general they can only refer to inanimates when used independently (stirling & huddleston, 2002). using a demonstrative this/that instead of she in example 1, for instance, would not be possible in english. 3 note that the use of der as a demonstrative pronoun is fully grammatical in german; it is very frequent in spoken german (with some regional variation) and less frequent, though still attested, in written german (for corpus results see portele & bader, 2016). prominence during german pronoun resolution 3 while the pronouns in (1a) and (1b) are both technically ambiguous, there is a tendency for the personal pronoun in (1a) to refer to the more prominent referent serena and for the demonstrative pronoun in (1b) to refer to the less prominent referent venus. furthermore, the demonstrative pronoun yields more robust preferences, being more strongly associated with a non-prominent entity, while the personal pronoun tends to be more flexible in its interpretation (bosch, katz, & umbach, 2007; schumacher et al., 2016). in particular, it has been a challenge to identify which aspects of prominence are relevant when it comes to the resolution of personal and demonstrative pronouns in german. one proposal is that grammatical role is important: bouma and hopp (2006, 2007) claim that subjecthood is an important factor in determining prominence for the personal pronoun er, and further that subjecthood is a more important factor than order of mention. when it comes to the comparison between personal and demonstrative pronouns, bosch et al. (2003) and bosch et al. (2007) claim that personal pronouns refer to subjects, while the demonstrative der refers to non-subjects. however, this position was revised (bosch & umbach, 2007) to say that der avoids reference to antecedents that are topics. kaiser (2011a) found that personal pronouns are more flexible than demonstratives, which had a strong object bias in her sentence completion experiment. the experiment also showed that coherence relations did not modulate the interpretation of demonstratives to the same extent as personal pronouns. in kaiser’s study, the contribution of grammatical role, topichood and thematic role was not explored. but the notion that topichood is crucial to understanding the difference between personal and demonstrative pronouns has been taken up and refined by hinterwimmer (2015), and is confirmed by experimental data from wilson (2009), who claims that demonstratives are sensitive to topichood (while personal pronouns are sensitive to both topichood and grammatical role). furthermore, bosch and hinterwimmer (2016) incorporate the notion of topichood into the semantic representation of the two pronoun types, with er and der sharing the same representation except for an “avoid topic” feature for der. effectively, this analysis makes a requirement on the demonstrative to seek out a non-prominent antecedent. this leaves the personal pronoun underspecified and accounts for its greater flexibility in referential preferences while accounting for the difference between the two pronouns. evidence from schumacher and colleagues (schumacher et al., 2015; 2016; 2017) shows that the findings from garvey, caramazza and yates (1974) and stevenson et al. (1994) regarding the importance of thematic role for pronoun resolution also applies to german pronouns. further, they were able to isolate the contribution of thematic role, grammatical role, and word order, and claim that thematic role in some circumstances takes priority over grammatical role in the resolution of german personal and demonstrative pronouns, in making antecedents with the (proto-) agent role more prominent than antecedents with the (proto-) patient role.4 while an initial study in german by wilson (2009) using actives and passives to manipulate thematic role showed a null result, as wilson (2009) acknowledges, this may have been due to the grammatical role hierarchy and thematic role hierarchy pointing in opposite directions. schumacher and colleagues were able to isolate the contribution of thematic role in german by manipulating verb type, contrasting active accusative and dative experiencer verbs. with active accusative verbs, the nominative-marked argument has the thematic role of (proto-) agent, and the accusative-marked argument has the role of (proto-) patient. with dative experiencer verbs, the dative-marked argument has the thematic role of (proto-) agent and the nominative-marked argument has the role of (proto-) patient. this verb type contrast enabled the hierarchy of grammatical role (nominative subject = most prominent) to be separated from the hierarchy of thematic role (agent = most prominent), because in the dative experiencer verbs, the two hierarchies are not aligned. in schumacher et al. (2016), an antecedent-selection experiment and two sentence completion experiments, the personal pronoun tended to be resolved to arguments with the thematic role of agent and the demonstrative to arguments with the patient role, irrespective of the grammatical role hierarchy. (note that manipulations of word order showed that grammatical role was also a relevant parameter, but it appears to be less important than thematic role.) importantly, in schumacher’s approach, the relevance of thematic role for pronoun resolution does not completely rule out the contribution of other factors such as grammatical role. indeed, grammatical role may be an important factor when the grammatical role of the pronoun is also considered, where grammatical role parallelism may exert an influence (sauermann & gagarina, 2017). rather, it is proposed that these factors form prominence hierarchies which interact to determine the appropriate referent. personal and 4 see dowty (1991) for a discussion of proto-agent and proto-patient roles. patterson and schumacher 4 demonstrative pronouns appear to respond differently to the prominence hierarchies, with personal pronouns referring to the most prominent antecedent and demonstrative pronouns avoiding reference to the most prominent antecedent, referring instead to an antecedent with a lower prominence ranking. the differing preferences between personal and demonstrative pronouns are also evident during online processing (see schumacher et al., 2017 for visual world eye tracking data). furthermore, an event-related brain potential study (schumacher et al., 2015) revealed two different neurocognitive profiles for personal and demonstrative pronouns in german. the latter evoked a more pronounced negative deflection followed by an enhanced positivity. these two effects have been associated with differing discourse functions. the negative deflection for the demonstrative (compared to the personal pronoun) was associated with the cost of linking of the pronoun with an antecedent. in the case of the demonstrative der, this follows the requirement for a less prominent antecedent, with the cost arising either from the process of excluding a prominent antecedent, or from the retrieval of a less prominent entity (see also burkhardt, 2006; nieuwland & van berkum, 2006; streb, hennighausen, & rösler, 2004 for enhanced negativity effects when anaphor–antecedent linking is more demanding; see çokal et al., 2018 for increased processing demands associated with propositional antecedents compared to np antecedents). the enhanced positivity effect, on the other hand, was associated with forward-looking discourse processes. the demonstrative pronoun signals the potential for a thematic shift, unlike the personal pronoun which serves to maintain the existing topic. the shift in attention for the demonstrative elicits an enhanced positivity, as previously observed for topic shift and contrastive focus (hirotani & schumacher, 2011; hung & schumacher, 2014; wang & schumacher, 2013). until now, studies contrasting the online processing of personal and demonstrative pronouns in german have used paradigms where the pronouns are ambiguous. when no other cues are available, the calculation of prominence hierarchies may be particularly important in the pronoun resolution process, because they may be the only factor guiding antecedent choice (particularly in the case of the demonstrative, with the requirement to seek out a non-prominent antecedent). however, the role of prominence information is less clear when pronouns are unambiguous, that is, when only a single entity from the prior discourse is a suitable candidate antecedent. in these situations, prominence information, rather than guiding the choice of antecedent, may be used instead to evaluate the felicity of the antecedent that has already been selected. does prominence information still play an important role in these situations, and is the information used immediately? the current paper examines how prominence hierarchies are used when german personal and demonstrative pronouns are disambiguated by gender cues. 1.2 gender, processing depth and two-stage pronoun resolution many pronoun resolution studies have made use of gender as a disambiguating cue (badecker & straub, 2002; felser & cunnings, 2012; sturt, 2003, among many others). however, the precise way in which gender cues interact with accessibility cues has been debated. (in the current paper we equate the notion of prominence with accessibility: we assume that the most prominent discourse entity is cognitively the most accessible (for discussions about accessibility see ariel, 1990; arnold, 1998)). there have been claims that gender is used at an early stage of pronoun resolution, followed by accessibility at a later stage (crawley et al., 1990; ehrlich, 1980). it has also been claimed, however, that gender information is not always used immediately (mcdonald & macwhinney, 1995) or is used strategically (garnham, oakhill, & cruttenden, 1992). arnold et al. (2000) investigated gender and accessibility cues to pronoun resolution in two visual world eye tracking experiments, manipulating accessibility via order of mention and pronominalisation. they found that both gender and accessibility contribute immediately to the resolution of the pronoun. these conflicting claims have been somewhat clarified by more recent proposals for two-stage pronominal resolution (rigalleau, caplan, & baudiffier, 2004; stewart, holler, & kidd, 2007), as outlined below. away from the debate about prominence/accessibility and pronoun resolution, a different line of research has been concerned with whether this process is always completed. for example, greene, mckoon and ratcliffe (1992), in a series of probe recognition tasks, demonstrated that a unique referent for a pronoun is not always identified during processing. this finding accords with the shallow processing hypotheses according to which linguistic input is not always fully processed leading to underspecified and sometimes even incorrect interpretations (ferreira, bailey, & ferraro, 2002; ferreira & patson, 2007; sanford & sturt, 2002). with respect to pronouns, several studies have prominence during german pronoun resolution 5 shown that whether full interpretation of a pronoun takes place depends on the level of engagement that participants have in the task; for example, love and mckoon (2011) demonstrated that simply increasing the length of a story text from four to eleven lines increased participants’ engagement to the extent that pronoun resolution was successfully completed. a shallow level of processing may be induced in an experimental setting where participants are required to read a number of short texts which are unrelated to each other. incremental models of pronoun resolution which assume two processing stages can incorporate the possibility that pronouns may be left underspecified during processing. this has been explicitly incorporated in rigalleau et al.’s (2004) two-stage pronoun resolution model. this model also accounts for the conflicting findings discussed above for the timing of gender and accessibility information. in the model, the first stage of pronoun resolution consists of co-indexation. gender cues from the pronoun are automatically checked against the most accessible potential antecedent(s) in the prior discourse. the indexing at this stage does not represent a full commitment to a particular antecedent, rather, it is an automatic process that checks whether resolution is possible by identifying potential antecedents. the second stage involves disengagement: the activation of competing antecedents is reduced so that a single referent for the pronoun is identified. importantly, the disengagement stage is under strategic control. strategic control was described in garnham et al. (1992) as a process that participants could “turn on and off, depending on how they perceive their task”, and by rigalleau et al. (2004) as “requiring time and attention”. this second disengagement stage, then, is unlikely to take place under shallow processing. the model therefore assumes that under deep processing conditions a unique antecedent can be identified, whereas under shallow conditions the resolution is left underspecified, with only potential matching antecedents identified. this two-stage model is similar to the bonding and resolution model (garrod & sanford, 1990; garrod & terras, 2000; sanford, garrod, lucas, & henderson, 1983), in which candidate antecedents are identified in the bonding stage, and then semantic fit, plausibility and contextual appropriateness are checked in a second stage. both these models have in common that the first stage involves coindexation or bonding to the most accessible candidate(s)5,6. this implicitly assumes that the relative accessibility of potential antecedents has already been computed. for instance, if grammatical role contributes to antecedent accessibility, then the subject of a sentence would be more accessible than an object, leading to higher activation values for the subject (irrespective of whether it is later pronominalized or not). when a pronoun is subsequently encountered, the subject, being highly accessible, would be checked for gender match with the pronoun in the first stage of resolution. while this assumption fits with pronouns that are usually resolved to the most prominent antecedent, such as personal pronouns in english or german, it is less clear how this works when the pronoun is usually resolved to a less prominent antecedent, as in german demonstrative pronouns. stewart et al. (2007) extend rigalleau et al.’s (2004) two-stage model to ambiguous pronouns. in their study, depth of processing was manipulated by means of comprehension questions. they found that under deep processing (where participants are engaged with the task), ambiguous pronouns take longer to process than unambiguous ones. they attribute this ambiguity cost to the longer disengagement process, where participants have to strategically engage in order to deselect one of the potential pronouns. under shallow conditions, there is no cost of ambiguity. evidence from their experiment 2 points to ambiguous pronoun resolution in shallow conditions being delayed until disambiguating information becomes available, and no initial co-indexation stage. however, even under shallow conditions, the pronouns were nonetheless eventually resolved. the role of prominence hierarchies, as they are used in the resolution of german personal and demonstrative pronouns, has so far not been addressed in relation to two-stage processing models. if we equate prominence information to the notion of accessibility, then prominence hierarchies should contribute to stage 1, where co-indexation with the most accessible antecedent(s) takes place. but since we know that demonstrative pronouns have a requirement to avoid the most accessible candidate (while personal pronouns are flexible, or tend to be resolved to the most accessible candidate), the felicity of linking german pronouns to more or less accessible (prominent) candidates must be 5 garrod and sanford (1990) call this the antecedent that is in the reader’s attentional focus. 6 in this paper we concentrate on rigalleau et al.’s (2004) model and its extension by stewart et al. (2007) since they explicitly incorporate the notion of deep and shallow processing. patterson and schumacher 6 evaluated at some point during processing. this assessment process seems to fit better into the description of stage 2. 1.3 current study the current study investigates the processing of the german personal and demonstrative pronouns with respect to the relative prominence of potential antecedents and looks at how this process unfolds when the pronouns are disambiguated based on gender cues. the purpose of experiment 1 was to test the processing of the two pronoun types in german in environments where the referent was unambiguous, with the goal of gaining insight into the timecourse over which the two pronouns are processed and the timecourse of implementing the prominence hierarchies of grammatical role and thematic role. the underlying assumption was that the pronouns would be fully interpreted. foreshadowing the findings of experiment 1, we found that the pronouns did not appear to be fully interpreted, and we therefore carried out experiments 2 and 3 to further investigate the interaction of gender cues and the prominence hierarchies under deep processing conditions. 2 experiment 1 experiment 1 made use of the verb type distinction between active accusatives and dative experiencer verbs, as per previous studies by schumacher and colleagues (schumacher et al., 2015, 2016, 2017) (see introduction). using these two verb types allows the prominence hierarchy for grammatical role to be distinguished from that of thematic role. the purpose of experiment 1 was, firstly, to observe the timecourse over which the two prominence hierarchies interact, and secondly, to compare the processing of personal and demonstrative pronouns. in active accusative verbs the grammatical role and thematic role hierarchies are aligned, with both hierarchies pointing to the nominative marked agent argument as the most prominent antecedent. for example, in (2), the first np (henceforth np1) der reporter/die reporterin (“the reporter”) is nominative marked (highest grammatical role) and is the proto-agent (highest thematic role), while the second np (henceforth np2) die sprecherin/den sprecher “the spokesperson” is accusative marked (lower-ranked grammatical role) and is the protopatient (lower thematic role). in the dative experiencer verbs, conversely, the hierarchies are not aligned. as can be seen in (3), the np1 dem trainer/der trainerin “the trainer” is marked for dative case (lower grammatical role) and is the proto-agent (highest thematic role), while the np2 die lehrerin/der lehrer “the teacher” is nominative marked (highest grammatical role) and is the protopatient (lower thematic role). if the thematic role hierarchy were the only relevant parameter in determining prominence of antecedents, there should not be any difference between the two verb types when the pronouns are processed. if, however, both prominence hierarchies influence pronoun resolution (as suggested in schumacher et al., 2015, 2016, 2017) we expect to see a processing cost when the two prominence hierarchies are not aligned, as in the dative experiencer verbs. this may be demonstrated in later resolution of pronouns for dative experiencer verbs compared to active accusative verbs. furthermore, in experiment 1 we wished to test whether resolution of the demonstrative would be prolonged compared to the personal pronoun. this is based on findings that the neurocognitive profiles of the two pronouns differ (schumacher et al., 2015). it is assumed that the demonstrative, because of the requirement to pick out a less prominent referent and because of the change in expectations about the upcoming discourse, makes higher processing demands than the personal pronoun. if so, the timecourse of resolution for the demonstrative pronoun may be prolonged. in order to test these assumptions, we used a gender-match paradigm, which allowed us to detect the resolution of the pronouns. the two arguments of the verb (np1 and np2), which served as potential referents for the pronoun, were of different grammatical gender. the pronoun matched in gender with only one of the two potential antecedents, allowing us to manipulate resolution to either antecedent. here, we predict that the preferences for the personal and the demonstrative pronoun should differ. following previous findings, personal pronouns should be preferentially resolved to antecedents with a proto-agent role (always np1), or should not have a strong preference, given that some studies have shown that the personal pronoun is more flexible in its interpretation. the demonstratives, conversely, should prefer reference to the antecedent with a proto-patient role (always np2), given the lexical-semantic requirement to resolve the demonstrative to a non-prominent antecedent. this means that processing should be easier when the personal pronoun refers to np1 compared to np2, and for the demonstrative processing should be easier when referring to the np2 prominence during german pronoun resolution 7 compared to np1. this should give rise to an interaction between pronoun (personal/demonstrative) and antecedent (np1/np2). this interaction should be delayed for the dative experiencer verbs compared to the active accusative verbs if the non-aligned prominence hierarchies lead to additional processing. these predictions are based on the assumption that both gender information (information about gender match between the pronoun and the referent) and prominence information (resolution preferences based on prominence hierarchies) would contribute to the processing of the pronouns. 2.1 materials experimental items were short german texts comprising two sentences each. the first sentence contained two animate referents (np1 and np2) in the main clause. the second sentence contained a masculine pronoun (er or der, “he”).7 32 items contained an active accusative verb in the main clause of sentence 1, and 32 items contained a dative experiencer verb in the main clause of sentence 1, giving a total of 64 experimental items across the two verb types.8 within each verb type two factors were systematically manipulated: pronoun (er versus der); and antecedent (np1 versus np2). the factor antecedent was manipulated via gender match: the gender of the pronoun matched either the gender of the np1 or the np2 (the subordinate clause did not contain any masculine nouns). in the active accusative sentences np1 was always in nominative case, and in the dative experiencer sentences the np1 was always in dative case (i.e. the base argument order). an example of the conditions is given in (2) and (3) below (note that the idiomatic translation given at the end of each item set covers all four conditions). in addition, the following aspects of the materials were controlled in order to minimise the variation in eye-movements between and across conditions: there were always four words following the pronoun; the word preceding the pronoun was always aber or doch (“but”, “yet/still”). nouns in np1 position and np2 position did not differ (overall) in frequency, and were matched per item in letter length and syllable length. (2) active-accusative verbs 2a. er, np1 der reporter wollte die sprecherin befragen, weil die konferenz ausfiel. aber er hatte dann keine zeit. der reporter woll-te die sprecher-in befragen, the.m.sg.nom reporter.m want.pst-3sg the.f.sg.acc spokesperson-f interview.inf weil die konferenz aus-fiel. because the.f.sg.nom conference.f out-fall.pst aber er hat-te dann kein-e zeit. but he have.pst-3sg then no-f time.f 2b. er, np2 die reporterin wollte den sprecher befragen, weil die konferenz ausfiel. aber er hatte dann keine zeit. die reporter-in woll-te den sprecher befragen, the.f.sg.nom reporter.f want.pst-3sg the.m.sg.acc spokesperson.m interview.inf weil die konferenz aus-fiel. because the.f.sg.nom conference.f out-fall.pst aber er hat-te dann kein-e zeit. but he have.pst-3sg then no-f time.f 7 only masculine singular forms were used; the feminine pronouns are ambiguous for case and number. 8 note that there are only a limited number of dative-experiencer verbs in german. in the current experiment six verbs were used and repeated with different contexts to obtain 32 items: auffallen (“to catch so. eye”), entgehen (“to escape/evade so.”), missfallen (“to displease so.”), imponieren (“to impress so.”), behagen (“to please so.”), gefallen (“to please so.”). patterson and schumacher 8 2c. der, np2 die reporterin wollte den sprecher befragen, weil die konferenz ausfiel. aber der hatte dann keine zeit. die reporter-in woll-te den sprecher befragen, the.f.sg.nom reporter.f want.pst-3sg the.m.sg.acc spokesperson.m interview.inf weil die konferenz aus-fiel. because the.f.sg.nom conference.f out-fall.pst aber der hat-te dann kein-e zeit. but he.dem have.pst-3sg then no-f time.f 2d. der, np1 der reporter wollte die sprecherin befragen, weil die konferenz ausfiel. aber der hatte dann keine zeit. der reporter woll-te die sprecher-in befragen, the.m.sg.nom reporter.m want.pst-3sg the.f.sg.acc spokesperson-f interview.inf weil die konferenz aus-fiel. because the.f.sg.nom conference.f out-fall.pst aber der hat-te dann kein-e zeit. but he.dem have.pst-3sg then no-f time.f “the reporter wanted to interview the spokesperson, because the conference was cancelled. but he didn’t have time then.” (3) dative-experiencer verbs 3a. er, np1 dem trainer hatte die lehrerin imponiert, und zwar seit dem letzten sportfest. doch er wollte das nicht zugeben. dem trainer hat-te die lehrer-in imponiert, the.m.sg.dat trainer.m have.pst-3sg the.f.sg.nom teacher-f impress.ptcp und zwar seit dem letzt-en sportfest. and indeed since the.n.sg.dat last-n.sg.dat sports.day.n doch er woll-te das nicht zugeben. but he want.pst-3sg that not admit.inf 3b. er, np2 der trainerin hatte der lehrer imponiert, und zwar seit dem letzten sportfest. doch er wollte das nicht zugeben. der trainer-in hat-te der lehrer imponiert, the.f.sg.dat trainer-f have.pst-3sg the.m.sg.nom teacher.m impress.ptcp und zwar seit dem letzt-en sportfest. and indeed since the.n.sg.dat last-n.sg.dat sports.day.n doch er woll-te das nicht zugeben. but he want.pst-3sg that not admit.inf prominence during german pronoun resolution 9 3c. der, np2 der trainerin hatte der lehrer imponiert, und zwar seit dem letzten sportfest. doch der wollte das nicht zugeben. der trainer-in hat-te der lehrer imponiert, the.f.sg.dat trainer-f have.pst-3sg the.m.sg.nom teacher.m impress.ptcp und zwar seit dem letzt-en sportfest. and indeed since the.n.sg.dat last-n.sg.dat sports.day.n doch der woll-te das nicht zugeben. but he.dem want.pst-3sg that not admit.inf 3d. der, np1 dem trainer hatte die lehrerin imponiert, und zwar seit dem letzten sportfest. doch der wollte das nicht zugeben. dem trainer hat-te die lehrer-in imponiert, the.m.sg.dat trainer.m have.pst-3sg the.f.sg.nom teacher-f impress.ptcp und zwar seit dem letzt-en sportfest. and indeed since the.n.sg.dat last-n.sg.dat sports.day.n doch der woll-te das nicht zugeben. but he.dem want.pst-3sg that not admit.inf “the teacher had impressed the trainer, in particular since the last sports day. but he didn’t want to admit it.” the 64 experimental items were interspersed with 96 fillers (64 containing feminine pronouns) and divided over 8 lists in a latin-square design. 2.2 participants data was collected from 35 native german speakers (10 male), age range 19-40 years, of whom 32 were included in the analysis. (two participants were excluded for excessive track loss, and one participant was excluded due to low accuracy (<70%) on the comprehension questions.) no participant reported any language disorders. all participants gave their consent and received a small fee or course credit for participation. 2.3 procedure participants were seated with their eyes 70cm from the computer screen displaying the text, with their head supported by a chin-rest and forehead-rest. texts were displayed on the screen in a black font (20pt) on a white background using the courier new font. texts were displayed over two lines. the line break was in the subordinate clause of the first sentence so that the second sentence always started in the middle of the second line, away from the line break. participants were asked to read sentences silently from computer screen at their normal reading rate and to answer comprehension questions on gamepad, while their eye movements were recorded using the eyelink 1000 (sr research) desktop mount. comprehension questions (y/n) followed half of experimental items and approx. quarter of the fillers (58 questions in total). six questions directly probed the referent of the pronoun, and five no-questions probed aspects of the pronoun sentence. the experiment took around 45 minutes, and participants were paid a small amount or given course credit for their participation. patterson and schumacher 10 2.4 analysis 2.4.1 regions of interest for the analysis, the sentences were divided into regions. the pronoun region consisted of the first two words of the second sentence: aber er or aber der (“but he”). the spillover region consisted of the two words following the pronoun. 2.4.2 data cleaning procedure if an individual fixation was shorter than 80ms, it was merged with a nearby fixation (within 1 degree visual angle). if there was no nearby fixation to merge it with, the fixation was deleted. individual fixations longer than 1200ms were deleted. in a given trial, any regions that were skipped during firstpass reading were removed from the analysis and counted as missing data. skipping rates were 6.15% in the pronoun region and 2.83% in the spillover region. 2.4.3 data analysis procedure in the two regions the following reading measures were calculated: first-pass times (summed duration of fixations in a region before exiting it for the first time); right-bound reading times (summed duration of fixations in a region before exit to a later region); rereading times (summed duration of fixations minus the first-pass times; when no rereading took place, this was counted as missing data, making the rereading times a contingent measure); and total viewing times (summed duration of all fixations in a region). linear mixed-effects models to analyse the data per verb type/region/measure using the package lmertest version 2.0-33 (kuznetsova, brockhoff, & christensen, 2016) in the r statistical program (r core team, 2017). the data was transformed before submitting it to the model. the transformation was determined with the box-cox procedure (box & cox, 1964) using recommendations from osborne (2010). this resulted in a reciprocal square-root transformation for the pronoun region and a log transformation in the spillover region.9 the fixed part of each model contained the sum-coded factors pronoun (der; er) and antecedent (np1; np2) and the (centred) trial number, as well as the interactions between pronoun and antecedent. (the trial number was included to control for effects over the course of the experiment; the output was interpreted in an exploratory analysis, see section 2.7.) the random part of the model contained random intercepts for participant and item. the inclusion of by-item and by-participant random slopes for pronoun and antecedent were determined using the function repca from the package repsychling (baayen, bates, kliegl, & vasishth, 2015). 2.5 predictions assuming that (proto-) agenthood plays an important role in the selection of an antecedent during pronoun resolution in german (schumacher et al., 2015; 2016; 2017), in both the active accusative and the dative experiencer items, it should be easier to process er when it refers to np1 (proto-agent), compared to np2 (proto-patient). conversely, it should be easier to process der when it refers to np2 compared to np1. this should lead to an interaction of pronoun and antecedent: er-np1 will have shorter reading times than er-np2 and der-np2 will have shorter reading times than der-np1. if the misalignment of prominence hierarchies in the dative experiencer verbs leads to additional processing demands, the interaction of pronoun and antecedent in these items should appear in a later measure (rereading times) or a later region (spillover region) than the interaction for the active accusative items. building on the claim that the demonstrative pronoun is more rigid in its interpretive preferences and rejects the most prominent entity, enhanced processing costs are predicted. if the demonstrative pronoun der exerts more processing demands than the personal pronoun er, we will see overall higher reading times for der than er: this should be detected in the cumulative measure (total viewing times) and may be visible in either earlier measures (first-pass times; right-bound reading times) or later ones (rereading times). here, it is also possible that the processing for der is prolonged such that the expected np2-np1 difference for der appears in later measures or a later region than the np2-np1 difference for er. 9 note that this results in the effect directions being reversed for the pronoun region. prominence during german pronoun resolution 11 2.6 results 2.6.1 comprehension questions one participant scored below 70% in the comprehension questions and was removed from the analysis. for the remaining participants (n=32), overall accuracy was 83% (range 72-90%). 2.6.2 eyetracking results the means for each region and measure in the active-accusative items are shown in table 1. table 2 shows the outcome of the statistical models (effects of trial were included only as a control measure and are therefore not shown in this table – but see the exploratory analysis below). first-pass times right-bound reading times rereading times total-viewing times pronoun region mean mean mean mean der, np1 339 (196) 355 (223) 357 (192) 458 (289) der, np2 332 (186) 342 (203) 299 (156) 406 (251) er, np1 308 (169) 316 (182) 308 (244) 376 (250) er, np2 293 (154) 305 (167) 269 (145) 358 (194) spillover region mean mean mean mean der, np1 360 (181) 401 (198) 379 (328) 491 (276) der, np2 371 (181) 405 (207) 360 (211) 485 (271) er, np1 359 (189) 385 (201) 379 (259) 477 (278) er, np2 355 (195) 375 (203) 329 (213) 446 (246) table 1. means (in ms) for the pronoun and spillover regions in the active accusative items in experiment 1. standard deviations shown in parentheses. patterson and schumacher 12 first-pass times right-bound reading times rereading times total viewing times effect estimate (se) t -value estimate (se) t -value estimate (se) t -value estimate (se) t -value pronoun region pronoun -1.239e-03 (4.124e-04) -3.005** -1.344e-03 (3.794e-04) -3.542** -2.072e-03 (1.401e-03) -1.480 -1.710e-03 (5.080e-04) -3.366** antecedent -3.901e-04 (3.180e-04) -1.227 -3.309e-04 (3.232e-04) -1.024 -1.555e-03 (8.974e-04) -1.733(*) -7.343e-04 (4.254e-04) -1.726(*) pronoun x antecedent 1.841e-04 (3.179e-04) 0.579 2.380e-05 (3.231e-04) 0.074 -8.863e-04 (8.936e-04) -0.992 -5.329e-04 (3.525e-04) -1.512 spillover region pronoun 1.634e-02 (2.080e-02) 0.786 3.137e-02 (1.661e-02) 1.888(*) 1.585e-02 (3.190e-02) 0.497 3.187e-02 (1.497e-02) 2.128* antecedent -6.690e-03 (1.517e-02) -0.441 2.092e-03 (1.558e-02) 0.134 1.481e-02 (3.186e-02) 0.465 1.676e-02 (1.496e-02) 1.120 pronoun x antecedent -1.390e-02 (1.386e-02) -1.003 -8.734e-03 (1.323e-02) -0.660 -1.652e-02 (3.149e-02) -0.525 -2.910e-03 (1.496e-02) -0.195 table 2. model outputs for first-pass times, right-bound reading times, rereading times and total viewing times in the pronoun and spillover regions, activeaccusative items, experiment 1. *** p < 0.001; ** p < 0.01; * p < 0.05; (*) p < 0.1. prominence during german pronoun resolution 13 for the active accusative items, there is a main effect of pronoun in the pronoun region in first-pass and right-bound reading times, with der taking longer than er. the same effect is also reflected in the total viewing times for both the pronoun and the spillover regions. additionally, there are marginal effects of antecedent in rereading and total-viewing times in the spillover region, with longer reading times when the pronouns refer to np1. the means for each region and measure in the dative-experiencer items are shown in table 3. table 4 shows the outcome of the statistical models. first-pass time right-bound reading time rereading time total-viewing time pronoun region mean mean mean mean der, np1 361 (228) 383 (255) 370 (263) 488 (319) der, np2 364 (225) 371 (232) 391 (261) 495 (311) er, np1 305 (177) 319 (195) 349 (278) 395 (296) er, np2 311 (172) 321 (195) 281 (168) 372 (226) spillover region mean mean mean mean der, np1 387 (216) 419 (232) 373 (259) 512 (294) der, np2 353 (174) 405 (199) 354 (201) 493 (246) er, np1 365 (190) 386 (192) 350 (211) 474 (251) er, np2 353 (195) 384 (224) 376 (239) 469 (276) table 3. means (in ms) for the pronoun and spillover regions in the dative experiencer items in experiment 1. standard deviations shown in parentheses. patterson and schumacher 14 first-pass times right-bound reading times rereading times total viewing times effect estimate (se) t -value estimate (se) t -value estimate (se) t -value estimate (se) t -value pronoun region pronoun -1.857e-03 (3.913e-04) -4.745*** -1.887e-03 (4.279e-04) -4.410*** -2.675e-03 (9.579e-04) -2.793** -3.026e-03 (3.894e-04) -7.772*** antecedent 2.985e-04 (3.398e-04) 0.879 -2.698e-05 (3.332e-04) -0.081 -9.335e-04 (8.681e-04) -1.075 5.799e-05 (4.144e-04) 0.140 pronoun x antecedent -2.141e-04 (3.401e-04) -0.630 -3.014e-04 (3.334e-04) -0.904 1.586e-03 (8.696e-04) 1.824(*) -5.772e-06 (3.712e-04) -0.016 spillover region pronoun 1.391e-02 (1.400e-02) 0.994 3.350e-02 (1.639e-02) 2.044* -5.552e-03 (3.052e-02) -0.182 3.277e-02 (1.607e-02) 2.039(*) antecedent 2.700e-02 (1.399e-02) 1.930(*) 1.062e-02 (1.333e-02) 0.797 -3.780e-03 (3.059e-02) -0.124 1.155e-02 (1.458e-02) 0.792 pronoun x antecedent 7.419e-03 (1.398e-02) 0.531 -4.347e-03 (1.332e-02) -0.326 1.381e-02 (3.053e-02) 0.452 -3.578e-03 (1.457e-02) -0.245 table 4. model outputs for first-pass times, right-bound reading times, rereading times and total viewing times in the pronoun and spillover regions, dative experiencer items, experiment 1. *** p < 0.001; ** p < 0.01; * p < 0.05; (*) p < 0.1. prominence during german pronoun resolution 15 for the dative experiencer items, there is a main effect of pronoun in all measures of the pronoun region, with der taking longer than er. the same effect is also reflected in the right-bound reading times in the spillover region. additionally, there is a marginal main effect of antecedent in the firstpass times of the spillover region with longer reading times for np1, and in the pronoun region rereading times there is a marginal pronoun by antecedent interaction. 2.7 exploratory analysis there is a persistent main effect of pronoun throughout the experiment, which could be associated with differences in resolution processes but also more low-level differences between er and der. there are several word-level factors which may make the demonstrative harder to process than the personal pronoun. on its own, the word der is ambiguous between a definite article and a demonstrative, with the demonstrative pronoun being the less frequent of these uses; it is less frequent (in its pronominal usage) than the personal pronoun, and it tends to appear more in the spoken than in the written modality. we would expect frequency and ambiguity to affect reading times, but if they are the sole basis for the pronoun effect we would expect them to diminish over the course of the experiment. this is because participants tend to adapt very quickly to the local conditions of an experimental setting; participants should get used to der appearing as a demonstrative pronoun leading to a lowering of the associated processing cost. this was directly tested in this exploratory analysis (not pre-planned) by checking whether trial number interacted with the effect of pronoun. the same statistical models as in the main analysis were rerun, this time adding the interaction with trial. outputs for the pronoun by trial interaction are shown in appendix a. there was only one pronoun by trial interaction (rereading times, spillover region for the active accusatives). otherwise there were no interactions of the pronoun effect with trial. 2.8 discussion main effects of pronoun, with longer reading times for der, are seen throughout the experiment across both active accusative and dative experiencer verbs, most strongly visible in the pronoun region but also reflected in the spillover region. this effect is seen in both early and later eyetracking measures. the expected pronoun by antecedent interaction was not detected, except marginally in the rereading times for the dative experiencer verbs. based on findings from schumacher et al. (2015), we expected to find that the demonstrative pronoun would have a processing cost relative to the personal pronoun. this was seen in the main effects of pronoun throughout the experiment. in order to rule out that these effects are due to lowlevel differences in frequency or ambiguity of form, we checked whether the pronoun effects diminish over the course of the experiment. given that the pronoun effect largely did not interact with trial (as shown in the exploratory analysis), we find it unlikely that the pronoun effect comes from low-level differences between the two pronouns. our prediction of a processing cost for der was based on the distinct discourse functions of the two pronouns. one function is the requirement for the demonstrative pronoun to pick out a non-prominent referent as its antecedent; this is associated with higher processing demands, either because a prominent antecedent has to be excluded (in contrast to the unspecific resolution instruction associated with the personal pronoun), or because there is a cost for retrieving a less accessible antecedent. the other function is the change in expectations about the upcoming discourse, because demonstratives are associated with a topic shift. however, both of these functions pertain to the demonstrative referring to the non-prominent antecedent (np2 in experiment 1). but in our experiment, reference to np1 and np2 was systematically manipulated. therefore the discourse functions would only explain the processing cost for the demonstrative if there was also an interaction between pronoun and antecedent. an alternative explanation for the cost for the demonstrative which is not associated with its eventual referent is the unexpectedness of a form which signals a referential shift instead of thematic maintenance, regardless of the eventual referent (see portele & bader, 2016 for production data in german that show a low likelihood of demonstrative pronoun continuations). an alternative explanation for the extra cost associated with the der is the one letter length difference between the der region (7 characters) versus the er region (6 characters). we think that this is unlikely to be the sole reason for the cost associated with der, given the persistence and strength of the effect. length differences are likely to give rise to differences in landing position and gaze patterson and schumacher 16 duration (rayner, 1998) but are not normally also associated with increased processing downstream, especially when the difference is only one letter. as discussed in the introduction, previous studies investigating the interpretation of personal and demonstrative pronouns have shown that personal pronouns tend to refer to entities with a protoagent role (or to be flexible in their interpretation) while demonstratives tend to refer to entities with a proto-patient role. we expected that these interpretive preferences would be reflected in the reading record as an interaction between pronoun and antecedent. however, this pattern was not detected. we are therefore unable to assess whether the alignment of grammatical role and thematic role prominence hierarchies, which are aligned in the active accusative verbs but not aligned in the dative experiencer verbs, had an impact on processing times for pronoun resolution. the lack of a pronoun by antecedent interaction is surprising, given how robust the interpretation preferences have been in previous offline studies and in visual-world eye-tracking (bosch et al., 2007; schumacher et al., 2016, 2017). by using gender cues to resolve the pronoun to one antecedent or another, participants in the current experiment were presented with sentences that should have been infelicitous (in particular the der-np1 condition, but possibly also er-np2 requiring the less preferred resolution site), yet this was not evident in the reading times. even if the personal pronoun is much more flexible than the demonstrative in interpretive preferences (bosch et al., 2007; schumacher et al., 2016), which may minimise the difference between the personal pronoun conditions, there should still have been an effect for the demonstrative. we consider two (related) possible reasons why the interaction was not detected. one possibility is the way in which gender cues interact with cues that arise from the discourse structure (the prominence hierarchies). when presented with an ambiguous pronoun (no gender cues), as in previous studies, the act of resolving the pronoun may make use of any and all available cues to guide resolution. this means that information from the relative prominence of antecedents becomes important because it is able to guide resolution. conversely, when gender cues are sufficient to identify a single antecedent, as in the current experiment, prominence cues may be less important because they are not determining reference in any way (cf. crawley et al., 1990; ehrlich, 1980; garnham & oakhill, 1985; arnold et al., 2000 for personal pronouns). it is possible that gender cues completely overrule the prominence hierarchies because they are sufficient to identify a single referent, therefore no further processing is needed. in this way, even if the personal pronoun refers to the least prominent antecedent or the demonstrative to the most prominent, it will not affect the processing so long as a unique antecedent has been identified through gender cues. note that this possibility, if correct, would have implications for a two-stage model of pronoun resolution. it implies the following: if, at the indexing stage (stage 1), a single antecedent had been identified through gender cues, then evaluation of the link at stage 2, at least with respect to prominence cues, is effectively cancelled. the detection of a pronoun by gender interaction rests on the assumption that the pronouns are fully processed during the experiment. however, it is possible that participants were engaged in shallow processing. in the current experiment, the participants had to frequently answer comprehension questions, but most of these questions did not probe the referent of the pronoun. previous research has shown that, under similar conditions, participants did not always fully resolve pronouns (garnham et al., 1992; greene et al., 1992; rigalleau et al., 2004). an initial indexing stage may take place (and make use of gender cues, as discussed in the introduction), but a more strategic evaluation of candidate antecedents may not take place. under both scenarios, the use of information from prominence hierarchies is restricted or blocked. but from the current results it is unclear whether this is due to the task demands, or whether strong cues such as gender more generally limit the use of other, more subtle cues to pronoun interpretation. in order to decide between these two possibilities, experiment 2 was carried out. 3 experiment 2 experiment 2 was an untimed rating task, using a subset of materials from experiment 1. by asking participants to rate each item, we expected them to engage in deeper, more strategic processing than in experiment 1. this enabled us to test whether the presence of gender cues effectively blocks the use of prominence information, or whether prominence information can affect pronoun resolution if the task demands encourage deeper processing. prominence during german pronoun resolution 17 3.1 materials the materials consisted of 32 experimental items (16 active accusative, 16 dative experiencer) and 28 fillers which were a subset of the experimental and filler items from experiment 1. each experimental item appeared in 4 conditions, as per experiment 1. to ensure sufficient variation in the verbs for the dative experiencer conditions, each verb used in experiment 1 appeared at least twice, and four verbs (auffallen (“to catch so. eye”), entgehen (“to evade/escape so.”), missfallen (“to displease so.”), gefallen (“to please so.”)) appeared three times. the 28 fillers were a subset of the fillers from experiment 1, 16 of which contained a feminine pronoun (sie or die) to balance out the masculine pronouns appearing in the experimental items; the remaining 12 fillers did not contain any pronoun. half of the fillers were adjusted to make them sound unnatural or implausible. this was to encourage participants to use a wide range in the rating scale. 3.2 participants data was collected from 42 native speakers of german (8 male), age range 18-38 years; one participant was bilingual (german/english). no participant reported any language disorders. all participants gave their consent and received course credit for participation. 3.3 procedure the experimental items were mixed with the fillers and distributed over four lists in a latin-square design. each participant saw only one list. access to the experiment was via a link, and participants completed it remotely via the qualtrics survey platform (qualtrics, provo, ut). participants were instructed to read each text carefully and then to rate how good each text sounds from 1 sehr schlecht (“very bad”) to 7 sehr gut (“very good”). evaluations were given by clicking on stars under each text. 3.4 predictions if the presence of gender cues completely blocks the use of information from the prominence hierarchies, there should be no difference in ratings whether the pronouns refer to np1 or to np2. if, however, prominence information is used despite the presence of gender cues, then we expect an interaction of pronoun and antecedent. specifically, er-np1 should be rated significantly higher than er-np2; and der-np2 should be rated significantly higher than der-np1. 3.5 data analysis scores from 1-7 for each participant and item combination (including fillers) were converted to zscores, as is generally recommended for judgment data (schütze & sprouse, 2014) to account for participants’ variation in the use of the scale. the filler items were analysed to ensure that the task had been carried out as expected. mean z-scores for the adjusted (implausible) filler items were lower than the scores for the unadjusted items (-0.40 versus 0.83), and this difference was significant when tested in a linear mixed-effects model (t=7.391, p<.001). in addition, each participant’s mean z-score for the adjusted fillers was lower than their mean z-score for the unadjusted items, confirming that they all performed the task as expected. z-scores for the experimental items were then analysed using a linear mixed-effects model using the package lmertest version 2.0.33 (kuznetsova et al., 2016) in the r statistical program (r core team, 2017). the fixed part of the model contained the sum-coded factors pronoun (er; der), antecedent (np1; np2), and verb type (active accusative; dative experiencer), and the (centred) trial number, as well as the interactions between them. the random part of the model contained random intercepts for participant and item. the inclusion of random slopes for pronoun, antecedent and verb type were determined using the function repca from the package repsychling (baayen et al., 2015). 3.6 results table 5 shows the mean raw and z-scores per condition and verb type. patterson and schumacher 18 active accusative verbs dative experiencer verbs raw scores z-scores raw scores z-scores er, np1 4.40 (2.03) 0.19 (0.88) 3.30 (1.88) -0.35 (0.84) er, np2 4.08 (2.00) 0.05 (0.80) 3.43 (1.80) -0.33 (0.84) der, np1 3.53 (2.00) -0.24 (0.88) 2.68 (1.77) -0.68 (0.72) der, np2 4.58 (1.90) 0.27 (0.88) 3.18 (1.83) -0.44 (0.76) table 5. mean raw scores and z-scores per condition and verb type for experiment 2. standard deviations shown in parentheses. the model showed main effects of pronoun (t=-3.340, p=.001), antecedent (t=-2.966, p=.005) and verb type (t=5.517, p<.001). there was a significant interaction of pronoun and antecedent (t=-5.427, p<.001) and a significant interaction of pronoun, antecedent and verb type (t=-2.583, p=.010). data was split by verb type and pronoun to follow up the interactions (p-values were bonferroni-corrected to adjust for the four tests). for the active accusative verbs, the personal pronoun er showed no effect of antecedent (t=1.633, p=.107), whereas for the demonstrative pronoun der there was a significant main effect of antecedent (t=-4.660, p<.001), with significantly lower scores for the np1 condition versus np2. for the dative experiencer verbs, the personal pronoun er showed no effect of antecedent (t=-0.108, p=.915), whereas for the demonstrative pronoun der there was a main effect of antecedent (t=-2.339, p=.025), with lower scores for the np1 condition versus np2. 3.7 discussion based on the results of experiment 1, it was possible that the presence of gender cues that identified a single antecedent blocked the use of information from prominence hierarchies. the purpose of experiment 2 was to test whether or not this is the case in an untimed experiment in which participants were asked to evaluate the items from experiment 1. experiment 2 showed that participants were sensitive to prominence hierarchies despite the presence of gender cues; while ratings did not differ significantly for the personal pronoun er, they did for the demonstrative der, such that reference to the more prominent antecedent (np1) was rated lower than reference to the less prominent antecedent (np2). this finding is in line with previous studies which have shown stronger preferences for der compared to er, and that der is preferentially resolved to an antecedent that is not the most prominent one. scores for sentences containing the demonstrative pronoun were not in general lower than those containing the personal pronoun. this suggests that in general, the demonstrative pronoun does not sound unnatural in these contexts compared to the personal pronoun. we interpret this as further evidence that the longer reading times for the demonstrative in experiment 1 are due to a specific processing burden, not captured by an offline rating task, rather than a general dispreference for the demonstrative. furthermore, the scores for the active accusative sentences were higher overall than those for the dative experiencer sentences. this could indicate that the misalignment of thematic role and grammatical role prominence hierarchies in dative experiencer verbs, as predicted for experiment 1 but not confirmed, does indeed have an impact on resolution. however, it is not possible in this experiment to rule out alternative explanations such as a dispreference for having a dative argument in first position or the lower overall frequency of these verbs compared to the active accusatives. importantly, the findings of experiment 2 rule out the possibility that gender cues completely block the use of prominence information. we attribute the difference between the findings for experiments 1 and 2 to the difference in task: in experiment 1, participants were not required to resolve the pronoun to answer (most of) the comprehension questions, nor were they asked to evaluate the sentence in any way. these are the conditions under which shallow processing may take place. in experiment 2, the rating task meant that participants had to engage in deeper processing in order to make a meaningful judgment. it therefore seems likely that the lack of prominence effects in experiment 1 was due to shallow processing, and that prominence hierarchies do indeed play a role when participants are forced to engage more deeply with pronoun resolution. this result further suggests that the use of prominence information for evaluation of reference is under strategic control. prominence during german pronoun resolution 19 while prominence information may be used during an offline evaluation, the timecourse over which the interaction between gender cues and prominence information plays out remains unclear. note that the two-stage processing models discussed in the introduction do not specify whether prominence information is only part of indexing, in that it contributes to the accessibility of certain candidates, or whether the prominence hierarchies also contribute more strategically to the evaluation of candidate antecedents, based on preferences for candidates ranked higher (in the case of personal pronouns) or lower (in the case of demonstrative pronouns) in the prominence hierarchies. we carried out experiment 3 to investigate the timing of prominence information during pronoun resolution. 4 experiment 3 the findings of experiments 1 and 2 seem to indicate that the use of prominence information in german pronoun resolution is at least partly under strategic control of the participants, with its effects showing up during deeper processing but not under shallow processing conditions. the purpose of experiment 3 was to investigate the timing of prominence information during the pronoun resolution process, by comparing the processing of pronouns that are (temporarily) ambiguous versus unambiguous with respect to gender cues under deep processing conditions. in this experiment, comprehension questions directly probed the referent of the pronoun in half the experimental items, such that deep processing should take place throughout the experiment. we expect that, during deep processing, prominence information will be used to evaluate candidates during stage 2 of pronoun resolution, resulting in the pronoun by antecedent interactions that we expected, but did not find, in experiment 1. furthermore, we expect that the timing of the use of prominence information will differ between (temporarily) ambiguous conditions and unambiguous conditions. in the unambiguous conditions, gender cues point to a single potential antecedent during the indexing stage, so evaluation of the link with respect to prominence (stage two) can begin straight away. in the ambiguous conditions, gender cues do not narrow down the set of potential antecedents for the pronoun. disambiguating information becomes available further on in the sentence, so we expect that at this later point, a full stage two evaluation takes place in the ambiguous conditions.10 this means that pronoun by antecedent interactions should be seen earlier in unambiguous conditions compared to (temporarily) ambiguous conditions where the gender cues do not identify a single candidate antecedent. 4.1 materials experimental items were 64 short german texts comprising two sentences each. the first sentence comprised a main and subordinate clause, with two animate referents (np1 and np2) in the main clause and an active-accusative main verb. the second sentence started with a time adverbial (e.g. später “later”, kurz darauf “shortly afterwards”, danach “afterwards”) followed by an auxiliary verb and a masculine pronoun (er or der “he”) as the subject of the main verb, followed by a buffer region containing an adverbial phrase. following the buffer, either the np1 or the np2 from the first sentence was repeated as the object; the repeated np could not therefore be the referent of the pronoun. the buffer containing the adverbial phrase before the repeated np was always at least 10 characters long, to ensure that the disambiguating information from the repeated np was not visible when the pronoun was initially encountered. three factors were manipulated: pronoun (er versus der); antecedent (np1 versus np2) and disambiguation (early versus late), giving rise to eight conditions. in the early disambiguation conditions, the referent of the pronoun was disambiguated through gender match, as per experiments 1 and 2; the gender of the two nps always differed (one masculine, one feminine). the position of the gender matching (masculine) antecedent was manipulated (np1 or np2) to create the two antecedent conditions. in the late disambiguation conditions, the gender of np1 and np2 was always the same 10 we remain ambivalent on precisely what happens at the point of the pronoun in the (temporarily) ambiguous conditions. following stewart et al. (2007), it is possible that participants decide on a likely referent for the pronoun even without disambiguating cues under deep processing. if so, this would have to be re-evaluated at the disambiguation point. if, on the other hand, participants refrain from singling out a possible referent early on, and wait until disambiguating information becomes available (behaviour which may be reinforced throughout the experiment as they discover multiple items in which the pronoun is disambiguated downstream), then stage two evaluation would take place at the disambiguation point. in either scenario, then, some kind of stage two evaluation takes place at the disambiguation point; this is the effect that we are interested in picking up. patterson and schumacher 20 (masculine) so that both matched the pronoun, creating a temporary ambiguity about the referent of the pronoun. the correct referent of the pronoun only became clear at the repeated np. the position of the correct referent (np1 or np2) was manipulated to create the two antecedent conditions. an example of the conditions is given in (4) to (5) below. (note that the idiomatic translation given at the end of each item set covers all four conditions.) (4) early disambiguation 4a. er-np1 der trainer hat die spielerin getroffen, um die turnschuhe abzugeben. danach hat er überraschenderweise die spielerin zur abschiedsfeier eingeladen. der trainer hat die spieler-in getroffen, the.m.sg.nom trainer.m have.3sg the.f.sg.acc player-f meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat er überraschenderweise die spieler-in have.3sg he surprisingly the.f.sg.acc player-f zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp 4b. der-np1 der trainer hat die spielerin getroffen, um die turnschuhe abzugeben. danach hat der überraschenderweise die spielerin zur abschiedsfeier eingeladen. der trainer hat die spieler-in getroffen, the.m.sg.nom trainer.m have.3sg the.f.sg.acc player-f meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat der überraschenderweise die spieler-in have.3sg he.dem surprisingly the.f.sg.acc player-f zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp 4c. er-np2 die spielerin hat den trainer getroffen, um die turnschuhe abzugeben. danach hat er überraschenderweise die spielerin zur abschiedsfeier eingeladen. die spieler-in hat den trainer getroffen, the.f.sg.nom player-f have.3sg the.m.sg.acc trainer.m meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat er überraschenderweise die spieler-in have.3sg he surprisingly the.f.sg.acc player-f zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp prominence during german pronoun resolution 21 4d. der-np2 die spielerin hat den trainer getroffen, um die turnschuhe abzugeben. danach hat der überraschenderweise die spielerin zur abschiedsfeier eingeladen. die spieler-in hat den trainer getroffen, the.f.sg.nom player-f have.3sg the.m.sg.acc trainer.m meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat der überraschenderweise die spieler-in have.3sg he.dem surprisingly the.f.sg.acc player-f zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp “the {trainer/player} met the {player/trainer} in order to hand in the plimsolls. afterwards he surprisingly invited the player to a farewell party.” (5) late disambiguation 5a. er-np1 der trainer hat den spieler getroffen, um die turnschuhe abzugeben. danach hat er überraschenderweise den spieler zur abschiedsfeier eingeladen. der trainer hat den spieler getroffen, the.m.sg.nom trainer.m have.3sg the.m.sg.acc player.m meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat er überraschenderweise den spieler have.3sg he surprisingly the.m.sg.acc player.m zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp 5b. der-np1 der trainer hat den spieler getroffen, um die turnschuhe abzugeben. danach hat der überraschenderweise den spieler zur abschiedsfeier eingeladen. der trainer hat den spieler getroffen, the.m.sg.nom trainer.m have.3sg the.m.sg.acc player.m meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat der überraschenderweise den spieler have.3sg he.dem surprisingly the.m.sg.acc player.m zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp patterson and schumacher 22 5c. er-np2 der spieler hat den trainer getroffen, um die turnschuhe abzugeben. danach hat er überraschenderweise den spieler zur abschiedsfeier eingeladen. der spieler hat den trainer getroffen, the.m.sg.nom player.m have.3sg the.m.sg.acc trainer.m meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat er überraschenderweise den spieler have.3sg he surprisingly the.m.sg.acc player.m zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp 5d. der-np2 der spieler hat den trainer getroffen, um die turnschuhe abzugeben. danach hat der überraschenderweise den spieler zur abschiedsfeier eingeladen. der spieler hat den trainer getroffen, the.m.sg.nom player.m have.3sg the.m.sg.acc trainer.m meet.ptcp um die turnschuh-e abgeben. danach in.order the.f.pl.acc plimsoll.m-pl hand.in.inf afterwards hat der überraschenderweise den spieler have.3sg he.dem surprisingly the.m.sg.acc player.m zu-r abschiedsfeier eingeladen. to-f.dat farewell.party.f invite.ptcp “the {trainer/player} met the {player/trainer} in order to hand in the plimsolls. afterwards he surprisingly invited the player to a farewell party.” the 64 experimental items were interspersed with 96 fillers and divided over 8 lists in a latin-square design. the number and type of referents in both sentences of the fillers was varied to prevent participants building expectations about the position and type of referents in the experimental items. 72 of the fillers contained feminine pronouns sie or die (“her”). 4.2 participants data was collected from 65 native german speakers (10 male), age range 18-35 years, of whom 54 were included in the analysis based on their accuracy (>75%) in the comprehension questions. no participant reported any language disorders. all participants gave their consent and received a small fee or course credit for participation. 4.3 procedure participants were seated with their eyes 60cm from the computer screen displaying the text, with their head supported by a chin-rest and forehead-rest. texts were displayed on the screen in a black font on a grey background using the courier new font. participants were asked to read sentences silently from computer screen at their normal reading rate and to answer comprehension questions by pressing a button on a keyboard, while their eye movements were recorded using the eyelink 1000 (sr research) desktop mount. comprehension questions (y/n) followed 42 of the experimental items; 32 questions directly probed the referent of the pronoun. the experiment took around 50 minutes, and participants were paid a small amount or given course credit for their participation. prominence during german pronoun resolution 23 4.4 analysis 4.4.1 regions of interest for the analysis, the sentences were divided into regions. the pronoun region always appeared on the second line of text. the analysis was carried out on the pronoun region, which always contained the auxiliary verb and the pronoun, the buffer region which contained the adverbial phrase, and the repeated np region. 4.4.2 data cleaning procedure the data cleaning procedure was the same as in experiment 1, except that the maximum length for an individual fixation was set to 1000ms. skipping rates were 12.3% in the pronoun region, 2.9% in the buffer region and 0.6% in the repeated np region. 4.4.3 data analysis procedure the following reading measures were calculated: first-pass times; right-bound reading times; rereading times; and total viewing times (see experiment 1 for definitions). these reading measures were analysed in linear mixed-effects models per region/measure using the package lmertest version 2.0.33 (kuznetsova et al., 2016) in the r statistical program (r core team, 2017). before submitting it to the model the data was transformed. the transformation was determined with the box-cox procedure (box & cox, 1964) using recommendations from osborne (2010). the fixed part of the model contained the sum-coded factors pronoun (der, er); antecedent (np1, np2); ambiguity (early disambiguation, late disambiguation) and the trial number (centred), as well as the interactions between pronoun, antecedent and ambiguity. the random part of the model contained random intercepts for participant and item. the inclusion of by-item and by-participant random slopes for pronoun and antecedent were determined using the function repca from the package repsychling (baayen et al., 2015). where interactions were followed with pairwise comparisons, p-values were bonferroni-corrected. 4.5 predictions under deep processing conditions (i.e. all conditions in this experiment), it should be easier to process er when it refers to the more prominent np1 (proto-agent), compared to the less prominent np2 (proto-patient). conversely, it should be easier to process der when it refers to np2 compared to np1. this should lead to an interaction of pronoun and antecedent: er-np1 will have shorter reading times than er-np2 and der-np2 will have shorter reading times than der-np1. our interest lies in identifying when such interactions take place. if prominence information is used evaluatively as soon as a unique antecedent is identified (i.e. at stage two in two-stage processing models), then the timing of the above prominence effects will differ between the early and late disambiguation conditions: in the early disambiguation conditions, where gender cues identify a single antecedent, prominence effects should be visible in earlier measures (first-pass times, right-bound reading times) in the pronoun and buffer regions. in contrast, in the late disambiguation conditions, prominence effects should be delayed until the disambiguating information is read, i.e. in the repeated np region (early or late measures) and possibly also in later measures in the pronoun and buffer regions. additionally, if the demonstrative pronoun der exerts more processing demands than the personal pronoun er, as seen in experiment 1, we will see overall higher reading times for der compared to er: this should be detected in the cumulative measure (total viewing times) and may be visible in either earlier or later measures. furthermore, there may be an overall processing cost for late disambiguation conditions compared to early disambiguation conditions, because of the additional strategic engagement needed to deselect one of the potential antecedents (stewart et al., 2007). 4.6 results 4.6.1 comprehension questions overall accuracy on the comprehension questions was 89%, sd 6, range 76-100. accuracy on experimental items was also 89%, sd 6, range 76-100. patterson and schumacher 24 4.6.2 skipping rates skipping rates for the critical region (auxiliary + pronoun) were 12.3%; for the buffer region 2.9%; and for the repeated np region 0.6%. 4.6.3 eyetracking results the means for each region and measure are shown in table 6. table 7 shows the outcome of the statistical models. results are described in more detail in the sections below; total viewing times are not discussed, since they simply reflect the effects that are found in the early or the late measures. prominence during german pronoun resolution 25 first-pass times right-bound reading times rereading times total viewing times pronoun region mean mean mean mean early disambiguation der-np1 262 (164) 334 (189) 426 (296) 488 (305) der-np2 265 (167) 330 (181) 360 (213) 471 (254) er-np1 243 (133) 287 (140) 312 (195) 364 (219) er-np2 244 (128) 299 (150) 353 (263) 394 (247) late disambiguation der-np1 289 (182) 347 (178) 408 (266) 514 (301) der-np2 283 (177) 330 (192) 413 (413) 505 (326) er-np1 243 (121) 288 (146) 315 (203) 380 (218) er-np2 249 (129) 300 (148) 352 (232) 402 (243) buffer region early disambiguation der-np1 342 (225) 413 (259) 468 (389) 562 (393) der-np2 338 (205) 418 (241) 406 (293) 536 (334) er-np1 337 (196) 383 (223) 401 (309) 477 (328) er-np2 323 (190) 369 (225) 482 (447) 509 (391) late disambiguation der-np1 337 (195) 408 (237) 493 (349) 592 (378) der-np2 329 (185) 391 (240) 490 (481) 582 (437) er-np1 328 (210) 358 (219) 420 (291) 511 (335) er-np2 341 (203) 374 (231) 476 (359) 549 (386) repeated np region early disambiguation der-np1 359 (172) 406 (192) 445 (361) 524 (320) der-np2 381 (209) 419 (226) 431 (325) 535 (326) er-np1 360 (171) 381 (191) 378 (341) 479 (303) er-np2 369 (178) 399 (193) 450 (478) 534 (397) late disambiguation der-np1 343 (155) 382 (185) 478 (350) 552 (552) der-np2 351 (185) 399 (212) 509 (446) 599 (430) er-np1 338 (160) 372 (189) 426 (308) 514 (318) er-np2 368 (368) 394 (185) 569 (482) 600 (464) table 6. means (in ms) of the first-pass times, right-bound reading times, rereading times and total viewing times for each region (pronoun, buffer, repeated np) per condition in experiment 3. standard deviations shown in parentheses. patterson and schumacher 26 first-pass times right-bound reading times rereading times total viewing times effect estimate (se) t -value estimate (se) t -value estimate (se) t -value estimate (se) t -value pronoun region pronoun 3.770e-02 (8.826e-03) 4.271*** 5.418e-02 (9.430e-03) 5.745*** 9.093e-02 (1.621e-02) 5.611*** 1.188e-01 (1.079e-02) 11.013*** antecedent -1.927e-03 (8.826e-03) -0.218 -8.169e-04 (7.557e-03) -0.108 -5.330e-03 (1.504e-02) -0.354 -9.930e-03 (9.069e-03) -1.095 ambiguity -2.166e-02 (8.830e-03) -2.453* -7.241e-03 (7.560e-03) -0.958 -1.110e-02 (1.501e-02) -0.739 -2.323e-02 (9.072e-03) -2.560* pronoun x antecedent 3.586e-03 (8.824e-03) 0.406 1.977e-02 (7.556e-03) 2.617** 3.135e-02 (1.504e-02) 2.084* 1.722e-02 (9.068e-03) 1.899(*) pronoun x ambiguity -1.365e-02 (8.829e-03) -1.546 -4.315e-03 (7.560e-03) -0.571 -4.501e-03 (1.501e-02) -0.300 -2.851e-03 (9.072e-03) -0.314 antecedent x ambiguity -1.607e-03 (8.831e-03) -0.182 -7.636e-03 (7.561e-03) -1.010 1.437e-02 (1.503e-02) 0.956 -9.435e-03 (9.073e-03) -1.040 pronoun x antecedent x ambiguity -2.884e-03 (8.832e-03) -0.327 -9.151e-03 (7.562e-03) -1.210 6.598e-03 (1.505e-02) 0.438 3.454e-04 (9.074e-03) 0.038 buffer region pronoun 1.057e-02 (9.010e-03) 1.173 4.866e-02 (8.140e-03) 5.978*** 3.246e-02 (2.348e-02) 1.382 5.866e-02 (1.071e-02) 5.475*** antecedent 3.037e-03 (7.836e-03) 0.388 3.662e-03 (7.302e-03) 0.501 -2.306e-03 (1.738e-02) -0.133 -5.138e-03 (8.921e-03) -0.576 ambiguity -3.900e-03 (7.843e-03) -0.497 1.816e-02 (7.308e-03) 2.485* -3.487e-02 (1.736e-02) -2.008* -3.026e-02 (8.929e-03) -3.389*** pronoun x antecedent 4.203e-03 (7.836e-03) 0.536 -2.314e-04 (7.302e-03) -0.032 5.360e-02 (1.737e-02) 3.086** 1.871e-02 (8.921e-03) 2.097* pronoun x ambiguity 3.265e-04 (7.852e-03) 0.042 -2.554e-03 (7.316e-03) -0.349 -1.495e-02 (1.734e-02) -0.862 -4.873e-04 (8.939e-03) -0.055 antecedent x ambiguity 9.518e-03 (7.845e-03) 1.213 5.272e-03 (7.309e-03) 0.721 -2.966e-03 (1.745e-02) -0.170 1.818e-03 (8.931e-03) 0.204 pronoun x antecedent x ambiguity -1.350e-02 (7.845e-03) -1.721(*) -1.885e-02 (7.311e-03) -2.579** 9.377e-03 (1.742e-02) 0.538 -4.709e-03 (8.932e-03) -0.527 repeated np region pronoun -4.701e-03 (6.821e-03) -0.689 1.797e-02 (6.304e-03) 2.851** 2.836e-02 (1.788e-02) 1.586 2.314e-02 (9.212e-03) 2.511* prominence during german pronoun resolution 27 antecedent -1.666e-02 (6.817e-03) -2.444* -1.904e-02 (6.300e-03) -3.022** -3.720e-02 (1.792e-02) -2.075* -3.173e-02 (8.127e-03) -3.904*** ambiguity 2.318e-02 (6.823e-03) 3.397*** 1.862e-02 (6.307e-03) 2.953** -7.743e-02 (1.793e-02) -4.318*** -2.759e-02 (8.135e-03) -3.391*** pronoun x antecedent 6.493e-03 (6.817e-03) 0.953 8.691e-03 (6.301e-03) 1.379 4.277e-02 (1.795e-02) 2.383* 1.694e-02 (8.128e-03) 2.084* pronoun x ambiguity 7.937e-03 (6.827e-03) 1.163 9.674e-03 (6.310e-03) 1.533 2.521e-02 (1.788e-02) 1.410 5.829e-03 (8.140e-03) 0.716 antecedent x ambiguity 4.094e-04 (6.824e-03) 0.060 4.071e-03 (6.307e-03) 0.645 1.313e-02 (1.790e-02) 0.734 7.581e-03 (8.135e-03) 0.932 pronoun x antecedent x ambiguity -1.226e-02 (6.822e-03) -1.797(*) 1.895e-03 (6.305e-03) 0.301 -9.399e-03 (1.794e-02) -0.524 5.094e-03 (8.133e-03) 0.626 table 7. model outputs for all measures (first-pass times, right-bound reading times, rereading times and total viewing times) in the pronoun, buffer and repeated np regions. *** p < 0.001; ** p < 0.01; * p < 0.05; (*) p < 0.1. patterson and schumacher 28 4.6.3.1 early measures, pronoun region in first-pass times and right-bound reading times in the pronoun region there is a main effect of pronoun, with der conditions taking longer than er conditions. in first-pass times there is a main effect of ambiguity, with late disambiguation conditions taking longer than early disambiguation conditions. in the right-bound reading times there is an interaction of pronoun and antecedent: follow-up pairwise comparisons did not reach significance. 4.6.3.2 early measures, buffer region in the buffer region the right-bound reading times show a main effect of pronoun, with der conditions taking longer than er conditions, and a main effect of ambiguity, with early disambiguation conditions taking longer than late disambiguation conditions. these effects are qualified by a significant interaction of pronoun, antecedent and ambiguity. follow-up pairwise comparisons did not reach significance. 4.6.3.3 early measures, repeated np region in the repeated np region the right-bound reading times show a main effect of pronoun, with der conditions taking longer than er conditions. there is a main effect of antecedent in both first-pass times and right-bound reading times, with np2 conditions taking longer than np1 conditions. there is a main effect of ambiguity in first-pass times and right-bound reading times, with early disambiguation conditions taking longer than late disambiguation conditions. in the first-pass times there is a marginal interaction between pronoun, antecedent and ambiguity. follow-up pairwise comparisons reveal that in the late disambiguation conditions, first-pass times are significantly shorter when er refers to np1 compared to np2 (t = -2.912, p = 0.01). 4.6.3.4 later measures, pronoun region in the pronoun region there is a main effect of pronoun in rereading times, with der taking longer than er. in rereading times there is a significant interaction of pronoun and antecedent. follow-up pairwise comparisons do not reach significance. 4.6.3.5 later measures, buffer region there is a main effect of ambiguity in rereading times, with late disambiguation conditions taking longer than early disambiguation conditions. and there is a significant interaction of pronoun and antecedent in rereading times. follow-up pairwise comparisons reveal that rereading times are significantly longer when der refers to np1 compared to np2 (t = 2.276, p = 0.05), and that rereading times are marginally shorter when er refers to np1 compared to np2 (t = -2.123, p = 0.07). 4.6.3.6 later measures, repeated np region there is a main effect of antecedent in the rereading times, with np2 conditions taking longer than np1 conditions. there is also a main effect of ambiguity in rereading times, with late disambiguation conditions taking longer than early disambiguation conditions. and finally there is a significant interaction of pronoun and antecedent in the rereading times. pairwise comparisons reveal that er has significantly longer rereading times when referring to np2 compared to np1 (rereading t = -3.052, p < 0.01). 4.7 discussion in this experiment, in which deep processing was encouraged by frequently probing the referent of the pronoun, there are three main findings: increased processing times for the demonstrative prominence during german pronoun resolution 29 pronoun der compared to the personal pronoun er; increased processing times when the pronoun was ambiguous; and pronoun by antecedent interactions. the main effect of pronoun, with increased reading times for the demonstrative, was evident in right-bound reading times in all three regions, and first-pass and rereading times in the pronoun region. the effect here appears not only on the pronoun itself but two regions downstream, in both earlier and later measures. this finding looks very similar to the main effect of pronoun in experiment 1. there, the effect was attributed to the unexpectedness of a form that signals a referential shift, as opposed to the personal pronoun which signals referential maintenance (regardless of the actual referent). but there are two important differences. firstly, some of the main effects of pronoun are qualified by an interaction with antecedent. secondly, the conditions of experiment 1 gave rise to shallow processing, while in the current experiment, deeper processing was encouraged. so it is possible that the pronoun effect here arises for different reasons than in experiment 1. considering again the discourse functions of the demonstrative, previous studies have attributed the increased processing for the demonstrative to the fact that it must retrieve a less prominent, and therefore less accessible, antecedent from the previous discourse, and also the forward function of topic shift for the upcoming discourse (schumacher et al., 2015, 2017). in the present experiment this would only apply to the der-np2 conditions where the antecedent is less accessible, which could account for cases where there is also an interaction of pronoun and antecedent. this still leaves some main effects of pronoun (without interactions) unexplained. we propose the following: while the der-np1 condition has the advantage that a more prominent antecedent is being retrieved, it has the disadvantage that resolving to np1 clashes with the requirement that the demonstrative refers to a non-prominent antecedent. in effect, in the der-np1 condition participants are forced to resolve the pronoun to a more prominent, but less felicitous, antecedent. this clash may result induce a processing cost, reflected in longer reading times. this would mean that there is a processing cost associated with both the der-np2 conditions (accessibility) and the der-np1 conditions (infelicity qua the antiagent preference). another finding from experiment 3 were main effects of ambiguity. it was predicted, based on stewart et al.’s (2007) study, that ambiguous pronouns (late disambiguation conditions) would take longer than unambiguous pronouns (early disambiguation conditions), because of the longer time required for disengagement during the second processing stage. but the picture here looks to be a little more complex. the predicted effect was found in the first-pass times in the pronoun region, but the pattern is then reversed in early measures at the buffer and repeated np regions, with early disambiguation conditions taking longer; it reverses back again in the later measure in these two regions, where the late disambiguation conditions take longer. stewart et al. (2007) attribute the effect to the increased strategic processing required to reduce activation of one of the potential antecedents; in other words, participants have to actively decide between two (or more) potential referents and this requires additional attention. we suggest that this applies when the pronoun is first encountered (reflected in early measures at the pronoun region). but, precisely because a single candidate antecedent has been identified in the early disambiguation conditions, the integration of that antecedent can begin earlier than in late disambiguation conditions, and this integration is reflected in longer processing times in early measures in the buffer region. this may be due to participants reading on more quickly in the late disambiguation conditions as they try to obtain more information, or it may be due to a slowdown in reading the early disambiguation conditions while integration starts. tentative evidence that integration has already started for early disambiguation conditions at that point is that the pronoun by antecedent interaction differs over ambiguity (resulting in a pronoun by antecedent by ambiguity interaction at that point). however, it is difficult to be sure about what is happening here because pairwise differences are not significant. at this stage, there would be no need to do additional processing for the ambiguous pronoun because the relevant disambiguating information is not yet available. when it does become available at the repeated np region, there is a (later) slowdown for the late patterson and schumacher 30 disambiguation conditions while the antecedent is integrated, resulting in the main effect of ambiguity in that region. there is strong evidence that prominence information affects pronoun resolution during deep processing. there are a number of pronoun by antecedent interactions, and in particular the reading times for the personal pronoun are slower when it refers to a less prominent antecedent. the trend is for reading times of the demonstrative to be slower when it refers to a more prominent antecedent, although the pairwise differences are significant only in one measure. these effects confirm findings from earlier studies, that the prominence information affects the personal and demonstrative pronouns differently (bosch et al., 2007; schumacher et al., 2015, 2016, 2017). with regard to the timing of the prominence information, the picture is more mixed. we expected that, under deep processing, the use of the prominence information in evaluating pronoun-antecedent links (i.e. at stage 2 of the two-stage model) would lead to antecedent by pronoun interactions. furthermore, we expected that the interactions would occur earlier in the early disambiguation conditions than in the late disambiguation conditions, because we assumed that stage 2 would begin when disambiguating information became available (for the early disambiguation conditions, this is at the pronoun region and for the late disambiguation conditions this is at the repeated np region). there is limited evidence for this timing distinction in right-bound reading times in buffer region, where there is a three-way interaction of pronoun, antecedent and ambiguity, indicating that the pronoun by antecedent interaction differs between early and late disambiguation conditions; however, the pairwise comparisons were not significant. at the repeated np region, first-pass times reveal a prominence effect in the expected direction for the personal pronoun in the late disambiguation conditions. this could indicate that the evaluative use of prominence information may be delayed until disambiguating information becomes available, but once it does become available, prominence information is deployed rapidly. on the other hand, there is also evidence for prominence information affecting resolution at some distance from the point at which a single candidate antecedent could be identified: later measures in the buffer and repeated np region show pronoun by antecedent interactions that do not differ between ambiguous and unambiguous conditions, indicating that prominence information is affecting both early and late disambiguating conditions similarly. for the early disambiguation conditions, this is at some distance from the point at which a single candidate could be identified. it is possible that the availability of the repeated np in both early and late disambiguation conditions (effectively creating a second unambiguous cue to pronoun resolution in the early disambiguation conditions) encouraged participants to delay resolving the pronoun until the repeated np had been encountered in some cases. this would add further weight to the argument that prominence information is deployed strategically during pronoun resolution. a final observation from this experiment is that when prominence information was used, more differences were found for the personal than the demonstrative pronoun. differences for the demonstrative pronoun may have been masked by overall increased reading times in those conditions. it was suggested above that, for the processing of the demonstrative pronoun, both the position on the prominence hierarchy (i.e. accessibility) and the felicity of referring to a particular position on the hierarchy may have had an effect on processing times. this fits with the finding of more differences for the personal pronoun compared to the demonstrative. for the personal pronoun, the er-np1 condition is advantageous for both felicity and accessibility, and the ernp2 condition is disadvantageous for both. if felicity and accessibility are combined, the difference between er-np1 and er-np2 is larger than between der-np1 and der-np2. 5 general discussion we conducted three experiments investigating how prominence information is used during the resolution of gender-disambiguating german personal and demonstrative pronouns. prominence during german pronoun resolution 31 demonstrative pronouns are particularly interesting because, unlike personal pronouns, they tend to be resolved to antecedents that are not the most prominent entities in the prior discourse. this is a challenge for processing models which are built on the assumption that pronouns always seek prominent antecedents. in comparison to many previous studies, we presented unambiguous pronouns while still manipulating prominence. experiment 1 was an eye-tracking during reading experiment in which gender features of the pronoun and antecedent made the pronouns unambiguous. the original assumption was that the pronouns would be fully resolved, enabling us to explore the prominence hierarchies of thematic and grammatical roles during the resolution of personal and demonstrative pronouns. but, while demonstrative pronouns elicited higher reading times than personal pronouns throughout the experiment, we found that the prominence of the antecedents did not make any difference to participants’ reading behaviour. in experiment 1 the comprehension questions rarely probe the referent of the pronoun, creating conditions under which participants may have engaged in shallow processing. during shallow processing, pronouns may not be fully resolved. if this was indeed the case in experiment 1, as we argue below, it would suggest that prominence information is not part of the automatic co-indexation process that has been claimed to form the first part of pronoun resolution (rigalleau et al., 2004; stewart et al., 2007). however, this seeming discrepancy may arise from differing concepts of prominence and how prominence information is used. we address this point below. given that gender cues in experiment 1 made available only one potential referent from the context sentence, the increased reading times for the demonstrative pronoun under shallow conditions remain surprising. we find an explanation of low-level factors such as frequency and ambiguity unconvincing given the stability of the effect over the course of the experiment, and we suggest instead that there is a cost in encountering a form which signals referential shift rather than topic maintenance, regardless of whether a single referent is eventually identified. experiment 2 was a rating task using a subset of materials from experiment 1. the task involved participants having to reflect on each sentence in turn and make a judgment about it, and as such encouraged a deeper level of processing. even though participants were not explicitly asked to make judgments about the referents of the pronouns, the results showed that the prominence information did affect participants’ rating of the items; specifically, the items containing demonstrative pronouns received lower ratings when the pronoun referred to a prominent antecedent. this is in line with previous research showing that demonstrative pronouns are preferentially resolved to non-prominent antecedents, and that the personal pronoun is somewhat more flexible in its interpretation preferences (bosch & umbach, 2007; schumacher et al., 2016; for similar findings with dutch demonstratives, see kaiser, 2011b). the finding that prominence information affected ratings despite the presence of gender cues rules out the possibility that gender cues block the use of weaker resolution cues such as prominence, as long as the task encourages deep processing. furthermore, the differing role of prominence information in experiments 1 and 2, which differ in task and depth of processing, suggests that the evaluative use of prominence information is under strategic control during pronoun resolution. in experiment 3, an eye-tracking during reading experiment, we created conditions for deep processing by increasing the number of comprehension questions that directly probed the referent of the pronoun. under these conditions we manipulated pronoun type, reference to a more or less prominent antecedent, and the ambiguity of the pronoun in order to assess the timecourse of processing associated with prominence information. as in experiment 1, demonstrative pronouns elicited longer reading times than personal pronouns. in addition, there were interactions of pronoun and antecedent which suggested that prominence information was being used during resolution. this demonstrates that under deeper processing conditions prominence information has an effect that is not captured under shallow processing, and is therefore likely to be under strategic control. the precise timing of the prominence information patterson and schumacher 32 was less clear, however. interactions with ambiguity suggested that prominence information was (sometimes) deployed rapidly as soon as a single candidate antecedent could be identified. but there was also evidence that the use of the prominence information was somewhat delayed. it could be the case that participants differed in how rapidly they deployed the prominence information, or that individual participants changed their strategy over the course of the experiment. in sum, we would like to argue that shallow processing was involved in experiment 1, but not experiments 2 and 3, and on this premise we claim that the combined results of the three experiments together demonstrate that the evaluative use of prominence information during pronoun resolution is under strategic control. in order to substantiate this claim, it is therefore important to first establish that in experiment 1, unlike experiments 2 and 3, shallow processing was involved. after conducting experiment 1 we considered two alternative explanations for not finding the expected interaction of pronoun and antecedent. firstly, it could be the case that the particular materials we used did not give rise to the expected prominence effect. however, given that experiment 2 used a subset of the same materials and did show the effect, this possibility can be dismissed. a second possibility we considered was that only ambiguous pronouns show prominence effects, and that the presence of disambiguating gender features overrules the need for considering prominence. this possibility was again ruled out by the results of experiment 2, which showed prominence effects despite presenting gender disambiguated pronouns. finally, we consider task effects. experiment 2, unlike experiment 1, required participants to make a metalinguistic judgment about each item. such an evaluative task is very likely to enhance participant engagement with the presented material. we would further argue that our experimental set up for experiment 1, where comprehension questions rarely targeted the reference of the pronoun, were precisely the conditions under which shallow processing was shown to take place in stewart et al. (2007), and we followed their set-up using more frequent, targeted comprehension questions to enhance the depth of processing for experiment 3. within the framework of a two-stage model of pronoun resolution such as put forward by rigalleau et al. (2004) and elaborated by stewart et al. (2007), the first, automatic stage of pronoun resolution proceeds whether or not processing is shallow or deep. this stage involves indexing potential antecedents using gender features. under shallow processing conditions, the process stops at this point. under deeper processing conditions, resolution proceeds to stage 2, which involves strategic processing. here, links between the pronoun and the candidate antecedents are evaluated and the activation of the non-antecedent is reduced. we propose that prominence information is used at stage 2 to evaluate candidate antecedents. as such, prominence information may not be deployed in shallow processing, where pronoun resolution is stopped after the first, automatic stage. the question of whether prominence information is used at stage 1, stage 2 or both may well depend on the precise notion of prominence information and how that information is used. rigalleau et al. (2004) invoke the notion of accessibility in order to determine which antecedents get checked (for gender cues) at stage 1 and which do not. they follow greene et al. (1992) in assuming that the available antecedents are in the “focus of attention”. in their own experiments, rigalleau and colleagues create scenarios where two potential referents are available in the preceding text, with one at quite some distance from the pronoun (approx. 35 words) and one much closer (6 words); they assume that the close but not the distant referent is accessible. they also imply (p. 919) that appearing in the previous sentence results in enough activation to warrant being the focus of attention. as such, some kind of prominence information must feed into stage 1, such that certain antecedents are checked and others are not, but this could equate to a simple distance metric or labelling all antecedents accessible if they appear in the preceding sentence. stewart et al. (2007) similarly assume that both referents in a typical agent–patient or subject– object configuration are potentially indexable (both having appeared in the previous sentence). in a scenario such as ours in experiment 1, then, where both potential antecedents appear in the prominence during german pronoun resolution 33 preceding sentence and where the less prominent antecedent is linearly closer to the pronoun, it is hard to argue that only the most prominent/accessible referent (np1) is indexed; it seems more likely that both are indexed. if both are indeed indexed at stage 1, this means that the relative prominence between an agent and a patient in the preceding sentence, which has been shown to have robust influence on the eventual pronoun resolution preferences of personal and demonstrative pronouns, must be irrelevant at stage 1, under the model that rigalleau, stewart and colleagues propose. we therefore suggest that the type of prominence information that we are interested in, i.e. the relative prominence of an agent compared to a patient, is deployed at stage 2, under deep processing conditions. we should also emphasise that we are talking about the evaluative use of this prominence information in unambiguous resolution, i.e. when gender features only match one antecedent. for example, a demonstrative pronoun refers to the more prominent antecedent; while the reference is clear, referring to a less prominent antecedent may not be considered felicitous and it is this evaluative use of prominence information we are interested in. as pointed out by a reviewer, placing prominence at stage 2 would appear to contradict previous literature which claims that prominence or accessibility information is considered early, notably arnold et al. (2000). (earlier studies claiming an important role for prominence (e.g. hudson d’zmura & tanenhaus, 1998; gordon et al., 1993) do not deploy methodologies with a timecourse fine-grained enough to detect the differences that we are interested in). arnold and colleagues’ results, however, do show a difference between ambiguous and unambiguous cases. in the unambiguous conditions, there are no differences between referring to a more or less prominent antecedent in the time-window they examined, which would be expected if prominence were being used to evaluate the candidate antecedent. so while their results certainly reflect the authors’ main claim that gender information is used early, it is not clear that their results wholly contradict our claim that evaluative use of prominence takes place at stage 2. furthermore, most of the studies showing effects of prominence/accessibility on pronoun resolution have also involved some kind of judgment or evaluation task, which can trigger deep processing and therefore involve stage 2. many studies contained offline experiments involving referent selection or sentence continuation tasks, or a judgment task (bosch et al., 2007; bouma & hopp, 2006; 2007; colonna et al., 2012; crawley & stevenson, 1990; kaiser, 2011a; kaiser & trueswell, 2008; schumacher et al., 2015; 2016; stevenson et al., 1994). self-paced reading experiments were combined with referent selection tasks (bosch et al., 2007; crawley et al., 1990; hudson d’zmura & tanenhaus, 1998; in gordon et al., 1993 spr is combined with a true/false task that is not described in detail). visual world experiments normally involved a picture-matching or picture judgment task (arnold et al., 2000; kaiser & trueswell, 2008; kaiser, 2011b; exceptions are järvikivi et al., 2005 and schumacher et al., 2017 in which the accompanying continuation task appeared infrequently, rather than after every trial). in the eeg experiment reported in schumacher et al., 2015, a comprehension question was presented after every trial, and some questions required pronoun resolution. one further, persistent finding was that the demonstrative pronoun was more costly to process than the personal pronoun. as discussed in experiment 1, we originally expected to find that the demonstrative pronoun would have a processing cost relative to the personal pronoun because of its distinct discourse function, following schumacher et al. (2015). the requirement for the demonstrative pronoun to pick out a non-prominent referent as its antecedent is associated with higher processing demands; the demonstrative also has a forward-looking topic shift function which is also associated with higher costs. but both of these discourse functions pertain to the demonstrative referring to the non-prominent antecedent. in experiment 1, reference to np1 and np2 was systematically manipulated. the discourse functions would only explain the processing cost for the demonstrative if there was also an interaction between pronoun and antecedent, which we did not find. our alternative explanation, then, is that the form itself, which signals a referential shift instead of thematic maintenance, was unexpected and therefore increased processing times, regardless of the eventual referent. patterson and schumacher 34 as also discussed in experiment 1, while there are several word-level factors such as frequency and ambiguity that make the demonstrative pronoun more effortful to process than the personal pronoun, these factors alone seem insufficient to explain why the effect for the demonstrative pronoun does not diminish during the experiments as participants become familiar with the local experimental context. we suggested in experiment 3, when there was deeper processing, that the cost for the demonstrative may have had a different source, arising from the combination of two discourse factors, accessibility and felicity. when the demonstrative refers to a less prominent antecedent, there is a cost of retrieving a less accessible entity. when it refers to the more accessible antecedent, there is a cost associated with retrieving a less preferred antecedent for the demonstrative. thus both conditions for the demonstrative pronoun are associated with a processing penalty. the combination of accessibility and felicity can explain why prominence effects for the personal pronoun were easier to detect in experiment 3 than prominence effects for the demonstrative. note that this fits with the finding from experiment 2 that ratings for the demonstrative were affected by prominence information. in the rating task, the cost of retrieving a more or less accessible antecedent would not be visible, since it is not a task that can measure processing. but it does measure felicity: there is a cost when the demonstrative refers to an infelicitous antecedent. the personal pronoun, being more flexible in its referent, was not penalised to the same extent for referring to a dispreferred antecedent. the relevance of prominence information to pronoun resolution processes is particularly important for languages such as german where the distinction between different pronoun types may be largely dependent on their resolution preferences. our results align with previous findings that demonstratives require additional processing resources compared to personal pronouns, and that personal pronouns preferentially refer to more prominent antecedents while demonstratives avoid prominent antecedents. further, we have shown that prominence information is not ruled out by the presence of stronger resolution cues such as gender. however, the deployment of prominence information in the evaluation of candidate antecedents is under strategic control and therefore does not take place under shallow processing conditions. acknowledgements we are grateful to simon napierala, daniela mertzen, hilde penner and julia plechatsch for assistance with materials and experimental set-up. to janna drummer, claudia kilter, daniela mertzen, jana mewe and brita rietdorf for data collection. thanks to franziska kretzschmar for analysis advice. claudia felser not only generously provided laboratory space but also offered advice and assistance in the planning stages in adapting the experimental design to an eyetracking during reading paradigm, and advice on conceptual issues. finally, we gratefully acknowledge that this research was funded by the german research foundation (dfg) as part of the collaborative research center 1252 “prominence in language” – project-id 281511265 – in the project c07 “forward and backward functions of discourse anaphora” at the university of cologne, department of german language and literature i, linguistics. references ahrenholz, b. (2007). verweise mit demonstrativa im gesprochenen deutsch: grammatik, zweitspracherwerb und deutsch als fremdsprache. berlin: de gruyter. almor, a. (1999). noun-phrase anaphora and focus: the informational load hypothesis. psychological review, 106(4), 748–765. ariel, m. (1990). accessing noun-phrase antecedents. routledge. arnold, j. (1998). reference form and discourse patterns (phd thesis). stanford university. prominence during german pronoun resolution 35 arnold, j. e., eisenband, j. g., brown-schmidt, s., & trueswell, j. c. (2000). the rapid use of gender information: evidence of the time course of pronoun resolution from eyetracking. cognition, 76(1), b13–b26. http://dx.doi.org/10.1016/s0010-0277(00)00073-1 badecker, w., & straub, k. (2002). the processing role of structural constraints on the interpretation of pronouns and anaphors. journal of experimental psychology: learning, memory and cognition, 28, 748–769. bosch, p., & hinterwimmer, s. (2016). anaphoric reference by demonstrative pronouns in german. in empirical perspectives on anaphora resolution (pp. 193–212). de gruyter. bosch, p., katz, g., & umbach, c. (2007). the non-subject bias of german demonstrative pronouns. in anaphors in text: cognitive, formal and applied approaches to anaphoric reference (pp. 145–164). bosch, p., rozario, t., & zhao, y. (2003). demonstrative pronouns and personal pronouns: german der versus er. in proceedings of the eacl 2003. budapest. bosch, p., & umbach, c. (2007). reference determination for demonstrative pronouns. in d. bittner (ed.), proceedings of conference on intersentential pronominal reference in child and adult language. berlin. bouma, g., & hopp, h. (2006). effects of word order and grammatical function on pronoun resolution in german. in ambiguity in anaphora (pp. 5–13). bouma, g., & hopp, h. (2007). coreference preferences for personal pronouns in german. in d. bittner & n. gagarina (eds.), intersentential pronominal reference in child and adult language (pp. 53–74). berlin. box, g. e., & cox, d. r. (1964). an analysis of transformations. journal of the royal statistical society. series b (methodological), 211–252. brown-schmidt, s., byron, d. k., & tanenhaus, m. k. (2005). beyond salience: interpretation of personal and demonstrative pronouns. journal of memory and language, 53, 292–313. https://doi.org/10.1016/j.jml.2005.03.003 burkhardt, p. (2006). inferential bridging relations reveal distinct neural mechanisms: evidence from event-related brain potentials. brain and language, 98(2), 159–168. https://doi.org/10.1016/j.bandl.2006.04.005 clark, h., & sengul, c. j. (1979). in search of referents for nouns and pronouns. memory & cognition, 7(1), 35–41. https://doi.org/10.3758/bf03196932 çokal, d., sturt, p., & ferreira, f. (2018). processing of it and this in written narrative discourse. discourse processes, 55(3), 272–289. https://doi.org/10.1080/0163853x.2016.1236231 colonna, s., schimke, s., & hemforth, b. (2012). information structure effects on anaphora resolution in german and french: a crosslinguistic study of pronoun resolution. linguistics, 50(5), 991–1013. crawley, r. a., & stevenson, r. j. (1990). reference in single sentences and in texts. journal of psycholinguistic research, 19(3), 191–210. https://doi.org/10.1007/bf01077416 crawley, r. a., stevenson, r. j., & kleinman, d. (1990). the use of heuristic strategies in the interpretation of pronouns. journal of psycholinguistic research, 19(4), 245–264. https://doi.org/10.1007/bf01077259 dowty, d. (1991). thematic proto-roles and argument selection. language, 67(3), 547–619. ehrlich, k. (1980). comprehension of pronouns. quarterly journal of experimental psychology, 32(2), 247–255. https://doi.org/10.1080/14640748008401161 felser, c., & cunnings, i. (2012). processing reflexives in a second language: the timing of structural and discourse-level constraints. applied psycholinguistics, 33(3), 571–603. https://doi.org/10.1017/s0142716411000488 ferreira, f., bailey, k. g. d., & ferraro, v. (2002). good-enough representations in language comprehension. current directions in psychological science, 11(1), 11–15. https://doi.org/10.1111/1467-8721.00158 patterson and schumacher 36 ferreira, f., & patson, n. d. (n.d.). the ‘good enough’ approach to language comprehension. language and linguistics compass, 1(1-2), 71–83. https://doi.org/10.1111/j.1749818x.2007.00007.x filiaci, f., sorace, a., & carreiras, m. (2014). anaphoric biases of null and overt subjects in italian and spanish: a cross-linguistic comparison. language, cognition and neuroscience, 29(7), 825–843. https://doi.org/10.1080/01690965.2013.801502 garnham, a., & oakhill, j. (1985). on-line resolution of anaphoric pronouns: effects of inference making and verb semantics. british journal of psychology, 76(3), 385–393. https://doi.org/10.1111/j.2044-8295.1985.tb01961.x garnham, a., oakhill, j., & cruttenden, h. (1992). the role of implicit causality and gender cue in the interpretation of pronouns. language and cognitive processes, 7(3–4), 231–255. https://doi.org/10.1080/01690969208409386 garrod, s., & sanford, a. j. (1990). referential processes in reading: focusing on roles and individuals. in comprehension processes in reading (pp. 465–485). hillsdale, nj: lawrence erlbaum associates, inc. garrod, simon, & terras, m. (2000). the contribution of lexical and situational knowledge to resolving discourse roles: bonding and resolution. journal of memory and language, 42(4), 526–544. https://doi.org/10.1006/jmla.1999.2694 garvey, c., caramazza, a., & yates, j. (1974). factors influencing assignment of pronoun antecedents. cognition, 3(3), 227–243. https://doi.org/10.1016/0010-0277(74)90010-9 gernsbacher, m. a., & hargreaves, d. j. (1988). accessing sentence participants: the advantage of first mention. journal of memory and language, 27(6), 699–717. gordon, p. c., grosz, b. j., & gilliom, l. a. (1993). pronouns, names, and the centering of attention in discourse. cognitive science, 17(3), 311–347. https://doi.org/10.1207/s15516709cog1703_1 greene, s. b., mckoon, g., & ratcliff, r. (1992). pronoun resolution and discourse models. journal of experimental psychology: learning, memory, & cognition, 18, 266–283. gundel, j. k., hedberg, n., & zacharski, r. (1993). cognitive status and the form of referring expressions in discourse. language, 69(2), 274–307. baayen, h., bates, d., kliegl, r., & vasishth, s. (2015). repsychling: data sets from psychology and linguistics experiments. himmelmann, n. p. (1997). deiktikon, artikel, nominalphrase: zur emergenz syntaktischer struktur. tübingen: niemeyer. hinterwimmer, s. (2015). a unified account of the properties of german demonstrative pronouns. in p. grosz, p. patel-grosz, & i. yanovich (eds.), the proceedings of the workshop on pronominal semantics at nels 40 (pp. 61–107). glsa publications. hirotani, m., & schumacher, p. b. (2011). context and topic marking affect distinct processes during discourse comprehension in japanese. journal of neurolinguistics, 24(3), 276– 292. http://dx.doi.org/10.1016/j.jneuroling.2010.09.007 hudson d’zmura, s., & tanenhaus, m. k. (1998). assigning antecedents to ambiguous pronouns: the role of the center of attention as the default assignment. in m. a. walker, a. k. joshi, & e. f. prince (eds). centering theory in discourse, 199–226. oxford: oxford university press. hung, y.-c., & schumacher, p. b. (2014). animacy matters: erp evidence for the multidimensionality of topic-worthiness in chinese. brain research, 1555, 36–47. https://doi.org/10.1016/j.brainres.2014.01.046 järvikivi, j., van gompel, r. p. g., hyönä, j., & bertram, r. (2005). ambiguous pronoun resolution: contrasting the first-mention and subject-preference accounts. psychological science, 16(4), 260–264. https://doi.org/10.1111/j.09567976.2005.01525.x prominence during german pronoun resolution 37 kaiser, e. (2011a). on the relation between coherence relations and anaphoric demonstratives in german. in i. reich, e. horch & d. pauly (eds), proceedings of sinn & bedeutung 15, 337–351. saarbrücken, germany: saarland university press. kaiser, e. (2011b). salience and contrast effects in reference resolution: the interpretation of dutch pronouns and demonstratives. language and cognitive processes 26(10), 1587–1624. http://dx.doi.org/10.1080/01690965.2010.522915 kaiser, e., & trueswell, j. c. (2008). interpreting pronouns and demonstratives in finnish: evidence for a form-specific approach to reference resolution. language and cognitive processes, 23(5), 709–748. https://dx.doi.org/10.1080/01690960701771220 kuznetsova, a., brockhoff, p. b., & christensen, r. h. b. (2016). lmertest: tests in linear mixed effects models. retrieved from https://cran.r-project.org/package=lmertest love, j., & mckoon, g. (2011). rules of engagement: incomplete and complete pronoun resolution. journal of experimental psychology: learning, memory, and cognition, 37(4), 874–887. mcdonald, j. l., & macwhinney, b. (1995). the time course of anaphor resolution: effects of implicit verb causality and gender. journal of memory and language, 34(4), 543–566. https://doi.org/10.1006/jmla.1995.1025 nieuwland, m. s., & van berkum, j. j. a. (2006). when peanuts fall in love: n400 evidence for the power of discourse. journal of cognitive neuroscience, 18(7), 1098–1111. https://doi.org/10.1162/jocn.2006.18.7.1098 osborne, j. w. (2010). improving your data transformations: applying the box-cox transformation. practical assessment, research and evaluation, 15(12). portele, y., & bader, m. (2016). accessibility and referential choice: personal pronouns and dpronouns in written german. discours. revue de linguistique, psycholinguistique et informatique. a journal of linguistics, psycholinguistics and computational linguistics, 18, 1–39. r core team. (2017). r: a language and environment for statistical computing. vienna, austria: r foundation for statistical computing. retrieved from https://www.rproject.org/ rayner, k. (1998). eye movements in reading and information processing: 20 years of research. psychological bulletin, 124(3), 372–422. rigalleau, f., caplan, d., & baudiffier, v. (2004). new arguments in favour of an automatic gender pronominal process. the quarterly journal of experimental psychology section a, 57(5), 893–933. https://doi.org/10.1080/02724980343000549 sanford, a.j., garrod, s., lucas, a., & henderson, r. (1983). pronouns without explicit antecedents? journal of semantics, 2(3–4), 303–318. https://doi.org/10.1093/semant/2.34.303 sanford, anthony j., & sturt, p. (2002). depth of processing in language comprehension: not noticing the evidence. trends in cognitive sciences, 6(9), 382–386. https://doi.org/10.1016/s1364-6613(02)01958-7 sauermann, a., & gagarina, n. (2017). grammatical role parallelism influences ambiguous pronoun resolution in german. frontiers in psychology, 8, 1205. https://doi.org/10.3389/fpsyg.2017.01205 schumacher, p. b., backhaus, j., & dangl, m. (2015). backwardand forward-looking potential of anaphors. frontiers in psychology, 6, 1746. https://doi.org/10.3389/fpsyg.2015.01746 schumacher, p. b., dangl, m., & uzun, e. (2016). thematic role as prominence cue during pronoun resolution in german. in empirical perspectives on anaphora resolution (pp. 213–240). berlin, boston: de gruyter. schumacher, p. b., roberts, l., & järvikivi, j. (2017). agentivity drives real-time pronoun resolution: evidence from german er and der. lingua, 185, 25–41. http://dx.doi.org/10.1016/j.lingua.2016.07.004 patterson and schumacher 38 schütze, c. t., & sprouse, j. (2014). judgment data. research methods in linguistics, 27–50. stevenson, r. j., crawley, r. a., & kleinman, d. (1994). thematic roles, focus and the representation of events. language and cognitive processes, 9(4), 519–548. https://doi.org/10.1080/01690969408402130 stewart, a. j., holler, j., & kidd, e. (2007). shallow processing of ambiguous pronouns: evidence for delay. the quarterly journal of experimental psychology, 60(12), 1680– 1696. https://doi.org/10.1080/17470210601160807 stirling, l., & huddleston, r. (2002). deixis and anaphora. in r. huddleston & g. pullum (authors), the cambridge grammar of the english language (pp. 1449-1564). cambridge: cambridge university press. https://doi.org/10.1017/9781316423530.018 streb, j., hennighausen, e., & rösler, f. (2004). different anaphoric expressions are investigated by event-related potentials. journal of psycholinguistic research, 33, 175–201. sturt, p. (2003). the time-course of the application of binding constraints in reference resolution. journal of memory and language, 48(3), 542–562. wang, l., & schumacher, p. (2013). new is not always costly: evidence from online processing of topic and contrast in japanese. frontiers in psychology, 4, 363. https://doi.org/10.3389/fpsyg.2013.00363 wilson, f. (2009). processing at the syntax–discourse interface in second language acquisition. doctoral thesis, university of edinburgh. prominence during german pronoun resolution 39 appendix a table a1. model outputs for trial by pronoun interaction in all measures of the pronoun and spillover region for active accusative and dative experiencer verbs, experiment 1. first-pass times right-bound reading times rereading times total viewing times effect estimate (se) t-value estimate (se) t-value estimate (se) t-value estimate (se) t-value pronoun region, active accusatives -1.468e-06 (5.414e-06) -0.271 -2.548e-06 (5.499e-06) -0.463 4.808e-06 (1.493e-05) 0.322 -2.577e-06 (6.002e-06) -0.429 spillover region, active accusatives 8.257e-05 (2.354e-04) 0.351 -1.186e-05 (2.250e-04) -0.053 -1.333e-03 (5.082e-04) -2.622** -3.502e-04 (2.544e-04) -1.376 pronoun region, dative experiencers -9.095e-07 (5.009e-06) -0.182 -2.295e-06 (4.913e-06) -0.467 1.965e-05 (1.252e-05) 1.569 6.140e-06 (5.541e-06) 1.108 spillover region, dative experiencers 7.414e-05 (2.046e-04) 0.362 -2.228e-04 (1.947e-04) -1.144 -8.129e-05 (4.431e-04) -0.183 -1.174e-04 (2.131e-04) -0.551 microsoft word g-asksforqgspecialissue0210.docx dialogue and discourse 3(2) (2012) 101–124 doi: 0.5087/dad.2012.205 ©2012 ming liu, rafael a. calvo and vasile rus. submitted 2/11; accepted 1/12; published online 3/12 g-asks: an intelligent automatic question generation system for academic writing support ming liu ming.liu@sydney.edu.au rafael a. calvo rafael.calvo@sydney.edu.au school of electrical and information engineering university of sydney sydney nsw 2006 australia vasile rus vrus@memphis.edu department of computer science university of memphis, memphis tn38152 usa editors: paul piwek and kristy elizabeth boyer abstract many electronic feedback systems have been proposed for writing support. however, most of these systems only aim at supporting writing to communicate instead of writing to learn, as in the case of literature review writing. trigger questions are potentially forms of support for writing to learn, but current automatic question generation approaches focus on factual question generation for reading comprehension or vocabulary assessment. this article presents a novel automatic question generation (aqg) system, called g-asks, which generates specific trigger questions as a form of support for students' learning through writing. we conducted a large-scale case study, including 24 human supervisors and 33 research students, in an engineering research method course and compared questions generated by g-asks with human generated questions. the results indicate that g-asks can generate questions as useful as human supervisors (‘useful’ is one of five question quality measures) while significantly outperforming human peer and generic questions in most quality measures after filtering out questions with grammatical and semantic errors. furthermore, we identified the most frequent question types, derived from the human supervisors’ questions and discussed how the human supervisors generate such questions from the source text. general terms: automatic question generation, natural language processing, academic writing support 1 introduction when students are asked to write a literature review or an essay, the purpose is often not only to develop disciplinary communication skills, but also to learn and reason from multiple documents. this involves skills such as sourcing (i.e., citing sources as evidence to support arguments) and information integration (i.e., presenting the evidence in a cohesive and persuasive way). in learning through writing, students need to consider trigger questions and monitor their understanding. most students, however, fall short in these metacognitive skills (graesser & person 1994). afolabi (1992) identified some of the most common problems that students have when writing a literature review, including not being sufficiently critical, lacking synthesis, and liu,calvo and rus 102 not discriminating between relevant and irrelevant materials. simple generic questions have been used to address these problems. these are questions such as “have you clearly identified the contributions of the literature reviewed?” and “did you connect the literature with the research topic by identifying its relevance?” reynolds and bonk (reynolds & bonk 1996) showed that students given generic trigger questions in a writing activity perform better than those who received no trigger questions. however, generic questions may not always support the process of writing on specific topics. more content-related questions need to be asked, and most academics would ask such questions in the process of providing feedback to students. in the field of automatic question generation (aqg), most systems (heilman & smith 2009; rus, cai, & graesser, 2007; wolfe, 1976) focus on the text-to-question task where a set of content-related questions are generated based on a given text. usually, the answers to the generated questions are contained in the text. for example, heilman and smith presented an aqg system to generate factual questions with an ‘overgenerating and ranking’ strategy based on natural language processing (nlp) techniques, such as name entity recognizer and whmovement rules, and a statistical ranking component for scoring questions based on features. the target applications of such systems are reading comprehension and vocabulary assessment. these are significantly different from academic writing, which is our target application. this is the first project, to our knowledge, that contributes a system for generating content specific questions that support writing. targeting this type of student-generated content has significant challenges distinct from those found in the literature for generating questions to support reading skills, where the content is made of textbooks and expertly written material. we contribute a description of how questions can be automatically generated from students’ texts that is grounded on taxonomies developed in the writing research literature. we use the following citation categories based on work by lehnert et al. (1990) relevant to literature review papers: opinion, result, aim of study, system, method, and application. table 1 shows examples of automatically generated questions based on the citation category. for example, if a student (citer) cites an opinion in his academic writing thus: “cannon (1927) challenged this view mentioning that physiological changes were not sufficient to discriminate emotions”, g-asks will generate trigger questions asking the student for evidence regarding the opinion. examples of these are shown in the row corresponding to the opinion category in table 1. a goal of this study is to investigate how human experts generate their trigger questions from the source text, as a form of feedback, to support academic writing, what types of trigger questions are commonly used by human experts in writing, and how useful these questions are. in addition, this project contributes a novel question generation technique based on sentence classification, particularly citation sentences which are common and informative elements in academic writing. this technique is used to develop and evaluate a tutoring system that scaffolds students' reflections on their academic writing with content-related trigger questions automatically generated from citations using nlp techniques. our basic approach to generate trigger questions is to first automatically extract citations from students' compositions together with key content elements. then, the citations are classified using a machine learning approach, and questions are generated based on a set of templates and content elements. experiments based on a bystander turing test, a version of the original turing test (person & graesser 2002), revealed that human evaluators have moderate difficulties distinguishing questions generated by the system from questions produced by humans (liu et al. 2010). the bystander turing test (person & graesser 2002) asks participants to rate if particular dialog moves, such as hint and prompt, in tutoring transcripts are generated by the computer system or human tutors. it also measures the quality of generated dialog moves by asking the participants to give a score. g-asks 103 category question source sentence opinion why did cannon challenge this view mentioning that physiological changes were not sufficient to discriminate emotions? (what evidence is provided by cannon to prove the opinion?) does any other scholar agree or disagree with cannon? cannon (1927) challenged this view mentioning that physiological changes were not sufficient to discriminate emotions. result does davis objectively show that this classification accuracy gets higher from about 70 % up to 98 % while actors express emotions and computers perform the...? (how accurate and valid are the measurements?) how does it relate to your research question? this classification accuracy gets higher from about 70% up to 98% while actors express emotions and computers almost perform the same on classifying 5-7 emotions (davis, 2001). system in the study of macdonald, why does workbench tool provide feedback on spelling, style and diction by analyzing english prose and suggesting possible improvements? what are the strength and limitations of the system? does it relate to your research question? the writer’s workbench tool provides feedback on spelling, style and diction by analysing english prose and suggesting possible improvements (macdonal et al, 1982). application why did hunter use fbg arrays as tunable elements for high-speed signal correlating of grating-based processors? could the problem have been approached more effectively from another perspective? does it relate to your research question? hunter (2003) used fbg arrays as tunable elements for high-speed signal correlating of grating-based processors. method why did ghosh develop an on-line algorithm capability of covering a complete range of faults from benign fault to faults of...? what are the strengths and limitations of this approach? to achieve this, ghosh (2004) developed an on-line algorithm capability of covering a complete range of faults from benign fault to faults of an unrestricted nature. aim why does gawlik conduct this study to investigate biomass conversion in water at pressure and temperature ranges of 30-50 mpa and 330-410 c? (what is the research question formulated by gawlik? what is gawlik's contribution to our understanding of the problem under study? ) gawlik (2003) investigated biomass conversion in water at pressure and temperature ranges of 30-50 mpa and 330-410°c. table 1: an example of content-related trigger questions produced by g-asks liu,calvo and rus 104 the approached presented in this article improves on a former approach we developed in a previous study (liu et al. 2010). in the previous approach, we proposed a combined tregex expression rule-based approach with sentiment analysis to classify the citation sentences, and evaluated it with a pilot study on a small group of subjects. tregex (levy & andrew 2006) is a powerful pattern matching technique which can denote the relations between syntactical tree nodes, such as noun phrase (np) or verb phrase (vp), from the syntactic tree derived from a sentence by using a sentence parser, such as the stanford parser (klein & manning 2003). for example, tregex pattern np < nn $ vp is denoted as an np immediately dominate an nn and sister vp. six conceptual citation categories were used in that study: opinion, aim, result, method, system and other. the study’s result shows that the accuracy on citation extraction reaches 60% in 145 citation sentences. one of the biggest advantages of the tregex expression pattern-matching rule is that it can match deep syntactic features of a sentence, such as predicate verb, subject and object. for example, if the predicate verb matches ‘argue,’ ‘challenge,’ or ‘claim,’ then the conceptual category of this citation sentence is ‘opinion’. we also used sentiwordnet (esuli & sebastiani 2006) to check if a sentence contains sentiment words. if it contains sentiment words, then it also is considered as ‘opinion’. however, creating rules is labor intensive, and this approach is not scalable. here, we describe machine learning techniques that have significantly improved the citation classifier’s performance. furthermore, we have conducted a full-scale study, in a real course, with writing activities that involved 57 subjects, including 24 supervisors and 33 postgraduate students. the remainder of the paper is organized as follows. section 2 provides a review of the literature with a focus on writing support and aqg systems. it also describes several question classification schemas relevant to our work. section 3 presents the major steps of our approach while section 4 details a case study we conducted to assess the quality of the generated questions using the bystander turing test. section 5 discusses the obtained results and suggests lines of future work. 2 related work nlp techniques have been used to develop a number of tutoring and feedback systems for academic writing support. section 2.1 reviews some of the writing support systems. section 2.2 focuses on systems that generate questions automatically. section 2.3 summarizes question classification schemas while section 2.4 presents automatic citation classification work. 2.1 automated feedback systems for writing support computational approaches to writing support have focused primarily on assessment and less on providing automatic feedback on writing (shermis & burstein 2002; williams & dreher 2004). despite a variety of initiatives to improve the quality of automatic feedback, the effectiveness of proposed systems remains to be proven and further research is needed. meanwhile, providing timely and appropriate feedback at key stages of the writing process remains a manual task and therefore a serious challenge for university lecturers. some of the early systems include writers workshop (anderson 2005), developed at bell laboratories, and editor (thiesmeyer & theismeyer 1990), developed at rochester institute of technology. both systems focus on grammar and style. studies of the impact of editor (beals 1998) concluded that the pedagogical benefits of grammar and style checking are limited. it could also be argued that these systems only aim at supporting writing to communicate as opposed to writing to learn. sak, a writing tutoring system (wiemer-hastings & graesser 2000) developed at the university of memphis, is based on the notion of voices that speak to the writer during the process of composition. sak uses avatars to associate each voice with a face and personality. each avatar provides feedback on a different aspect of the composition, pointing out the strong g-asks 105 and weak parts of the text but without correcting it. sak uses latent semantic analysis (lsa) to calculate the average distance between consecutive sentences and provide feedback on the overall coherence of the text. lsa is a technique used to measure the semantic similarity of texts (landauer et al. 2007). sak can also analyze the purpose of a sentence, identifying clusters of topics amongst student’s writings so that when the topic of a new composition is not identified students can be asked for an explanation or reformulation. sourcer’s apprentice intelligent feedback system (saif) (britt et al. 2004) is an automated feedback tool for writing essays which can be used to detect plagiarism, uncited quotations, lack of citations, and limited content integration problems. once a problem is detected, saif can give helpful feedback to the student as shown in table 2. saif also uses lsa techniques for plagiarism detection, computing the similarity between each essay sentence and the source sentences in the lsa space. for finding citations, saif uses a regular expression pattern matching technique to detect the explicit citations by recognizing phrases containing author name (e.g. according to, as stated in, state). evaluations showed that saif provides feedback that encourages more explicit citations in students’ essays. however, saif only addresses some basic problems for sourcing and integration. moreover, it requires a large number of source documents to build the lsa semantic space and a large number of predefined pattern matching rules. problem feedback prompts student to: 1a. unsourced copied material (plagiarism) reword plagiarism and model proper format. 1b. unsourced copied material (quotation) explicitly credit source and model proper format. 2. explicit citations explicitly make a minimum of 3 citations. 3. distinct sources mentioned cite at least 2 different sources. 4. excessive quoting paraphrase more instead of relying on quotations too heavily. 5.integration from multiple sources include a more complete coverage of the documents in set. table 2: types of problems saif addresses and the intended goal of feedback glosser is an automated feedback system that provides academic writing support for college students (villalon et al. 2008; calvo & ellis 2010). it uses textual data mining and computational linguistics algorithms to analyse various features of texts based on which feedback is provided to student writers. the feedback is in the form of generic trigger questions (adapted to each course) and document features that relate to each set of questions. for example, by analysing the words in each paragraph, glosser can measure how related two adjoining paragraphs are. if the paragraphs are too unrelated, this can indicate lack of lexical cohesiveness which glosser will flag. glosser (1.0) provides feedback on four aspects of the writing: structure, coherence, topics, and concept visualisation. glosser does not address sourcing directly, but four trigger questions are provided: (1) are the ideas used in the essay relevant to the question? (2) are the ideas developed correctly? (3) does this essay simply present the academic references as facts, or does it analyze their importance and critically discuss their usefulness? (4) does this essay simply present ideas or facts, or does it analyze their importance? liu,calvo and rus 106 2.2 computational approaches to natural question generation one of the first automatic question generation systems proposed for supporting learning activities was autoquest (wolfe 1976). in this case, as in most of the current research, questions are generated from external sources that students read (as opposed to write). the purpose of these questions is to help novices to learn english. the approach used in autoquest is similar to that of kunichika et. al. (2001) who proposed an aqg approach based on both the syntactic and semantic information extracted from the original text. their educational context was the assessment of grammar and reading comprehension around a story. the extracted syntactic features include subject, predicate verb, object, voice, tense, and sub-clause. the semantic information contains three semantic categories, noun, verb and preposition, which are used to determine the interrogative pronoun for the generated question. for example, in the noun category, several noun entities can be recognized including person, time, location, organization, country, city, and furniture. in the verb category, bodily actions, emotional verbs, thought verbs, and transfer verbs can be identified. it also extracts semantic relations related to the time, location, and other semantic categories, when an event occurs. because this technique extracts substantial syntactic and time/space semantic information from sentences, the generated questions can be quite sophisticated, leading to very good writing support. evaluations showed that 80% of the questions were considered by experts as appropriate for novices learning english and 93% of the questions were semantically correct (kunichika et al. 2001). for vocabulary assessment, there are recent attempts to automatically generate multiplechoice closed questions. in theory, they take reading materials and generate questions by removing some words from a source sentence. the two major issues in automatic multiple-choice question generation are: 1) to determine which words to remove from the source sentence and 2) to choose the wrong alternatives or distracters. coniam (1997) determined the words to remove by selecting every nth-word in the text to be a test item and distracters are produced by choosing the same part of speech (e.g. noun, verb or adjective) and similar word frequency in a tagged corpus. mitkov and ha (2003) determined the words to remove by choosing the key terms, which are noun phrases with a frequency over a certain threshold. the distracters (e.g. hypernyms and hyponyms of the term) were selected by consulting wordnet (fellbaum 1998), which is a lexical database that groups english nouns, verbs, adjectives and adverbs into synonym sets or synsets. a synset is linked to other synsets with various relations, including synonym, antonym, hypernym, hyponym and other semantic relations. they demonstrated that automatic generation and manual correction of questions can be more time-efficient than manual question creation alone. autotutor, developed by graesser et al. (person & graesser 2002) at the university of memphis, is an intelligent tutoring system (its) that improves students’ knowledge in computer literacy and newtonian physics through an animated agent asking a series of deep reasoning questions that follow the graesser-person taxonomy (graesser & person 1994). in each of these subjects a set of topics have been identified. each topic contains a focal question, a set of good answers, and a set of anticipated bad answers (misconceptions). the system initiates a session by asking a focal question about a topic and the student is expected to write an answer containing 510 sentences. initially, the system used a set of predefined hints or prompts to elicit the correct and complete answer. more recently, the hints and prompts are automatically generated (rus, cai, & graesser, 2007) and then human experts validate them instead of students evaluating the quality of hints and prompts as in person and graesser’s study. the authors showed that autotutor’s questioning approach had a positive impact on learning with an effect size on a pretest post-test study of approximately 0.8 standard deviation units in the areas of computer literacy and newtonian physics. however, the tutor system is domain dependent and requires a large number of human resources to predefine the content of each topic. g-asks 107 2.3 question taxonomy question taxonomies are developed based on analysis of human questions in tutoring situations, classroom teaching, or even technical manuals. the types of question are often related to bloom’s taxonomy (bloom 1984), which is a framework to assess cognitive processes. deeper questions are highly correlated to higher-level processes in bloom’s taxonomy. question taxonomies have been proposed according to different application domains, including computational modeling of question answering as a cognitive process (lehnert 1978), analyzing students’ questions in a dialog between a human and intelligent tutoring agent (acker et al. 1991), analyzing tutors’ questions in human tutorial dialog (graesser & person 1994; nielsen et al. 2008; boyer et al. 2009). the most well known question taxonomy was one proposed by graesser and person (1994) based on their two studies about human tutors and students’ questions during tutoring sessions in a college research method course and middle school algebra course. six trained human judges coded the questions in the transcipts, obtained from the tutoring sessions, on four dimensions: question identification, degree specification (e.g high degree means questions contain more words that refer to the elements of desired information), question-content category, and question generation mechanism (the reasons for generating questions include knowledge deficit in the learner own knowledge base, common ground between dialogue participants, social actions among dialogue participants, and conversation control ). they defined following 18 question categories according to the content of information sought rather than on the interrogative words (i.e. why, how, where, etc). 1. verification: invites a yes or no answer. 2. disjunctive: is x, y, or z the case? 3. concept completion: who? what? when? where? 4. example: what is an example of x? 5. feature specification: what are the properties of x? 6. quantification: how much? how many? 7. definition: what does x mean? 8. comparison: how is x similar to y? 9. interpretation: what does x mean? 10. causal antecedent: why/how did x occur? 11. causal consequence: what next? what if? 12. goal orientation: why did an agent do x? 13. instrumental/procedural: how did an agent do x? 14. enablement: what enabled x to occur? 15. expectation: why didn’t x occur? 16. judgmental: what do you think of x 17. assertion: 18. request/directive after analyzing 5,117 questions in the research methods and 3,174 questions in the algebra sample, they found four frequent question categories: verification, instrumental-procedural, concept completion, and quantification questions. they stated that some questions could belong to more question categories. for example, the question “did the drug dosage decrease the anxiety?” is ambiguous since it belongs to verification and antecedent questions. but, this type of question should not be construed as a weakness in the classification scheme because polythetic classification schemes can be used. in our study, we adapted the question taxonomy proposed by graesser and person to analyze questions generated by human experts. section 4.6 describes it in more detail. liu,calvo and rus 108 2.4 citation classification and extraction citations are commonly used in research documents. similar to question taxonomies, different citation classification schemas and methods have been proposed motivated by different purposes. lehnert et al. (1990) presented a citation taxonomy based on a corpus of machine learning research papers, defining 18 conceptual reference categories including system, method, concept, result, fact, criticism, example, application, proposal, problem and argument, and 3 relationships between these categories including, similarity, difference and flagship. the purpose of this research project was to summarize a scientific research paper in terms of underlying research trend. they used a rule-based sentence parser, called circus, to automatically fill out predefined frames containing reference categories and relationships slots. it was reported that the system correctly classified 75% of citations sentences. however, their evaluation is not quite convincing since they only used 28 citations sentences from 2 research papers. this rule-based approach does not generalize well to unseen data, and designing the parsing rules is very time-consuming. because our work is similar to this study, we adapted some reference categories from lehnert et al., but we used a an approach with greater generalization power based on a machine learning method proposed by tefuel (2006). comparing to a rule-based approach proposed by lehnert et al., tefuel proposed a supervised machine learning approach to automatically classify the citation types in order to improve the impact factor calculations and citation indexer. they defined 12 mutually exclusive categories including weak (weakness of previous researches), contrast, base (current work is based on other research), use (current work uses other method), similarity (current work is similar to other work) and neutral (neutral description of cited work). the feature set contains 12 features that record the presence of 892 cue phrases identified by annotators, which include agent type (the authors of the paper or everybody else), action type (e.g. aim of study), location, verb tense and voice. they tested the approach using 2,829 citations from 116 articles, randomly selected from acl (association for computational linguistics) conferences. to report the classification results, teufel and colleagues (2006) used a macro-averaging f1-score (0.57) and the ibk (k=3) algorithm, which is an alternative version of the k-nearest neighbour. there are several drawbacks of this approach and methodology: the distribution of citation categories is skewed and the proposed cue-phrases feature is too coarse-grained. since different automated citation classification methods were proposed, citation extraction and analysis have become a great interest to researchers. powley and dale (2007) introduced terminologies to describe the variety of citations styles, such as textual citation (it uses authoryear pair to refer to an entry in the reference list) and index citation (it uses numbers). they further divided the textual citation styles into 4 categories. examples of textual citations are provided below in bold face: 1. textual syntactic: citations form a syntactic part of the sentence. e.g. levin (1993) provides a classification of over 3000 verbs. 2. textual parenthetical: citations are enclosed in parentheses. e.g. two current approaches to english verb classifications are wordnet (miller et al., 1990) … 3. prosaic: people name is used to refer to an earlier citation. e.g. levin groups verbs based on an analysis of their syntactic properties . . . 4. pronominal: pronoun is used to refer to an earlier citation. e.g. her approach reflects the assumption that the syntactic behavior of a verb is determined. finally they used regular-expression-based heuristics to automatically extract citations from a research paper. the citation extraction recall was reported as 0.99 based on a collection of papers from 2000-2005 containing 294 citations. in our study, we will focus on textual citation extraction by using regular-expression-based heuristics. g-asks 109 3 system design and architecture g-asks has been integrated into our iwrite web application (calvo et al. 2010), which allows students to write and submit their assignments, and provides them with a complete solution for supporting the write-review-feedback cycle of a writing activity. iwrite used automatic feedback tools for students to do revision, such as glosser described in the related works section. these tools are based on tml (http://sourceforge.net/projects/tml-java/), a multipurpose text mining library that provides functionalities for the pre-processing of documents, such as tokenization, stemming, stop-word removal, sentence segmentation and building latent semantic space. it maintains three corpora, adding each new document, at the sentence, paragraph, and document level to an apache lucene database, which is commonly used in information retrieval and text mining tasks. compared to a traditional database, such as mysql, the speed of lucene fulltext search and indexing is much faster. figure 1 shows the g-asks system architecture and its integration to iwrite. the input to gasks is a literature review paper stored at the sentence level after the preprocessing in iwrite and the output is a set of generated questions used by iwrite, which delivers the questions to the student author. iw rite w eb a pplication human teachers give trigger questions as feedbacks students submit literature review papers system deliveries human and automated feedbacks g-asks: automated question generation system stage 1:citation extraction, pronoun resolution, parsing, sentence simplification. stage 2: citation classification with naïve bayes classifier stage 3: qg with rule-based approach article stored as sentence level in lucene db question generated by system question template repository based on graesser and person’s question taxonomy figure 1: the g-asks system architecture integrated in iwrite web application the question generation process follows 3 major stages. stage 1: citation extraction. in this stage, all the sentences are retrieved from the lucene database and citation sentences are extracted, parsed and then simplified. in our current approach, we are only interested in textual citations, defined by powley and dale, containing names as this fits best with our goal of generating trigger questions for writing support. thus, the index citation style is ignored. a pattern matching technique is used to extract the textual syntactic and textual parenthetical citation styles. the regular expression code is shown below. \([a-za-z]*\s*\d{4}\)|\([p.]+\s*\d{1,4}\)|\([a-za-z]+\s*[a-za-z]* \s*[a-za-z]*\w*\d{4}|\([^)]*\d{4}\s?\) as you can see, this regular expression code would match the textual syntactic style by \([a-zaz]*\s*\d{4}\)|\([p.]+\s*\d{1,4}\) or match the textual parenthetical \([a-za-z]+\s*[a-za-z]* \s*[a-za-z]*\w*\d{4} . a state of the art named entity tagger (ner), lbj (ratinov & roth 2009), is used to identify citations with prosaic style, and a simple pronoun resolver, finding the nearest name entity appearing before the pronoun, was used to identify citations with pronominal style. liu,calvo and rus 110 once citations are extracted using the previous approach, sentence simplification is performed which involves splitting compound and complex sentences and also removing phrase types such as appositives, non-restrictive relative clauses, and participial modifiers. after the sentence simplification is performed, we parse the simple citation sentences to get syntactic features, including subject, main verb, and auxiliary verb (e.g. be, am, will, have and can) predicate, voice and tense, which are essential to perform subject-auxiliary inversion. we used the stanford parser (klein & manning 2003) and tregex (levy & andrew 2006) for parsing sentence and sentence simplification separately. the tregex expressions in table 3 and 4 are used to simplify a sentence and extract subject, predicate verb and predicate from that sentence. description of transformation tregex expression a simple sentence, containing one subject and a compound verb, is split into two sentences: 1 subject + predicate1; 2 subject+ predicate2. cc=conjuct $+ vp=predicate1 $ vp=predicate2 > (vp=predicateparent > (s > root) $ np=subject) a compound sentence, containing two or three independent clauses joined by a coordinator, is split into two or three sentences: s1, s2, s3. cc=conj $+ (s=s1 > (s=smain > root)) $-(s=s2 > (s > root)) | $ (s=s3 > (s > root)) a complex sentence, containing an independent clause joined by one dependent clauses, is split into two sentences: s1, np+vp sbar < in < s =s1 $ (np=np > (s > root) $ vp=vp) remove the apposition by delete app, lead and trail. rrc|pp|sbar|vp|np=app $/,/=lead $+ /,/=trail !$ cc !$ conjp remove non-restrictive clauses by delete comma and modifier. root=root << (vp !< vp < (/,/=comma $+ /[^`].*/=modifier)) table 3: examples of tregex expression rules used in stage 1 to split and compress complex sentences syntactic feature tregex expression subject np > (s > root) predicate verb /^vb/ > ( vp > ( s >root)) predicate vp > (s > root) table 4: examples of tregex expression rules used to extract syntactic features. stage 2: citation classification. the goal of this stage is to identify the citation category (described next) for each citation candidate retrieved based on the citation style detection rules described above. the description of each citation category is shown below: 1. aim: to present the aim of an author’s study, e.g. bunescu et al. focused on extracting named entities from natural language documents. 2. opinion: to express the opinion of an author, e.g. reiter and dale (1997) state that template-based systems are more difficult to maintain and update. g-asks 111 3. result: to report the result of an author’s study, e.g. mccallum and nigam (1998) show that the multivariate bernoulli model performs well with small vocabularies. 4. method: to describe a method, algorithm, technique, model, or framework proposed by an author, e.g. bi-normal separation is a relatively new feature selection method introduced by forman (2003). 5. system: to describe a system, e.g. autoslog (riloff, 1996) is a dictionary construction system that creates extraction patterns automatically using heuristic rules. 6. application: to apply a method/system to a field, e.g. kappa statistics k will be used to measure the reliability (siegel and castellan, 1988). we implemented a statistical citation classifier using a machine learning approach. we represent each citation as a vector of 17 generic features. as a training set we used 504 citations from 45 academic papers. the features are described in the following: cue phrases. this is similar to teufel’s approach (2006) that defines a verb cluster as a feature for a category. however, our approach provides more information to identify a feature. we call it a cue phrase group that includes a noun cluster, an adjective words cluster and an adverb cluster. according to hyland’s study (1994), reporting verbs are widely used in academic writing and each reporting verb can be used for different reasons (aim of study, opinion and result) in a citation. for example, a reporting verb list in the opinion verb cluster contains argue, claim, view, reason, explain, emphasize and etc. the opinion noun verb cluster includes opinion, view, claim, limitation and etc. some verbs or nouns can be used by more than one category and we define these verb or noun clusters in a shared cue phrase group. for instance, both the opinion and result category share the cue phrase group containing suggest, note, point to, observe and etc. in other words, the shared cue phrase group feature gives weight to both opinion and result category. we define 12 cue phrase group features: 1) aim cue phrase group (verb, noun and adjective clusters), 2) opinion cue phrase group (verb and noun clusters), 3) shared opinion, 4) result cue phrase group (verb cluster), 5) result cue phrase group (result verb and none cluster), 6) shared aim & result cue phrase group (verb cluster), 7) application cue phrase group (verb and noun cluster), 8) method cue phrase group (noun cluster), 9) system cue phrase group (noun cluster), 10) shared system & method cue phrase group (verb and noun cluster), 11) own (no cluster) and 12) other cue phrase group (verb, noun cluster and adjective cluster). sentiment feature. this is a binary feature to check if the citation sentence contains sentiment words with polarity that is either positive or negative. the sentiwordnet (esuli & sebastiani 2006), a publicly available lexical resource for opinion mining, is an extension of wordnet and has three categories for a word sentiment with some magnitude: positive, negative and neutral. we utilize this resource to identify sentiment words. negation feature. we define the following four cue phrase groups (70 words in total) to detect negation in a citation sentence: 1. traditional negation words, such as not, no, never, neither, nor, none and not only. 2. restrictive adverbs, such as few, little, rarely, seldom, hardly, scarcely, barely. 3. verbs with negative meaning, such as fail, deny, avoid. 4. adjectives with prefix in-, dis-, unand non, such as insufficient, imbalance, uncommon, nonassessable and insignificant. these words are obtained from the frequent academic word lists (coxhead 2000). syntactic features. we use the voice and tense features from teufel’s study (2006). other features. we use the length of a citation sentence and the numeric features which indicate if the citation sentence contains numeric characters. stage 3: generation. the final stage of our approach is the actual trigger question generation module. it uses a template-based approach. once the semantic and syntactic features extracted liu,calvo and rus 112 from a citation match the predefined patterns in our repository of templates, the corresponding questions are generated. table 5 shows the 6 rules defined in our rule repository. for example, a citation is extracted in step 1: cannon (1927) challenged this view mentioning that, physiological changes were not sufficient to discriminate emotions. in step 2, the citation classifier categorizes it as opinion. step 3 applies rule 1 to generate the following question to trigger the student’s reflection by asking for the evidence for other person’s opinion: why did cannon challenge this view mentioning that physiological changes were not sufficient to discriminate emotions? (what evidence is provided by cannon to prove the opinion?) does any other scholar agree or disagree with cannon? in order to fill in this question template shown in the rule 1 of table 5, the subject_auxiliary_inversion operation occurs where the auxiliary precedes a subject. the key to this implementation is to find the auxiliary verb. we used predefined tregex patterns to implement this. for example, this tregex pattern " md > ( vp > ( s > root )) " is used to find the modal auxiliary, such as can, could, may, might, need and ought while the pattern "/^vb/ > ( vp > ( s > root)) < (are|is|am|was|were|has|have|had|do|will|would|should|) is used to find the regular auxiliary, such as be, have/has, do, will, would, shall, should and had. if we couldn’t find both auxiliary in the sentence, we will set the auxiliary verb as do, does or did depending on the tense of main verb and singular or plural of the main verb. rule category question template 1 opinion why +subject_auxiliary_inversion()? what evidence is provided by +subject+ to prove the opinion? do any other scholars agree or disagree with +subject+? 2 result subject_auxiliary_inversion()? is the analysis of the data accurate and relevant to the research question? how does it relate to your research question? 3 system in the study of +subject+, why +subject_auxiliary inversion()? what are the strength and limitations of the system? does it relate to your research question? 4 application why+subject_verb_inversion()? could the problem have been approached more effectively from another perspective? does it relate to your research question? 5 method in the study of +subject+, why +subject_auxiliary _inversion()? which dataset does +subject+ use for this experiment? what are the strengths and limitations of this approach? 6 aim why does +subject+ conduct this study to +predicate+? what is the research question formulated by +subject+? what is +subject+s contribution to our understanding of the problem? table 5 six rules and template-based questions. 4 evaluation of the automatic question generation system to evaluate the ability of g-asks to generate high quality and effective questions for supporting academic writing, we compared questions generated by the system to those produced by humans. like the bystander turing test conducted by person and graesser(2002), in this study judges (student writers) rated the quality of each question according to different measures and were also asked to ascertain whether the question was generated by a human or a system. however, there are some differences between the tests carried out by person and graesser and our two evaluations. we used the academic writing task to generate questions while they used a snippet of g-asks 113 a tutorial dialog. furthermore, our judges were the writers of the content while in person and graesser’s study judges did not know the content before the experiment. 4.1 participants and procedure we conducted a study with 57 participants (33 phd students-writers and 24 supervisors). the students were enrolled in a research methods course from the faculty of engineering at the university of sydney. each student submitted a research proposal as part of this subject (and their phd requirements) to the iwrite web site. each proposal was read by a peer, who is another phd student from this course, and the supervisor, who is supervising the student-writer of this proposal, both providing feedback in the form of questions. having been informed about this experiment, each student was asked to rate the quality of questions generated from his/her literature review paper. these questions were produced by the supervisor, a peer, and by g-asks, combined with a sample of generic questions shown in table 6. generic questions 1 did your literature review cover the most important relevant works in your research field? 2 did you clearly identify the contributions of the literature reviewed? 3 did you identify the research methods used in the literature reviewed? 4 did you connect the literature with the research topic by identifying its relevance? 5 what were the author's credentials? were the author's arguments supported by evidence? table 6: examples of generic question type. each question producer (supervisor, peer, g-asks or generic question) generated a maximum of 5 questions. therefore, each student evaluated 20 questions at most. students were then given the following five quality measures, some of which were used in similar studies by heilman and smith (heilman & smith 2009), to evaluate the questions using a likert scale where 1 was ‘strongly disagree’ and 5 ‘strongly agree’: 1. this question is correctly written (qm1). 2. this question is clear (qm2). 3. this question is appropriate to the context (qm3). 4. this question makes me reflect about what i have written (qm4). 5. this is a useful question (qm5). we received 5 ratings each for 615 questions under each quality measure. quality measures 1, 2 and 3 focus on ‘acceptability’ of the generated question (whether it is grammatically correct, not vague, and makes sense according to the context) while the quality measures 4 and 5 focus on the ‘usefulness’ of the generated question, whether it is helpful to trigger reflection. 4.2 citation extraction evaluation and result first, we evaluated g-asks’s performance with respect to citation extraction ability and semantic correctness of generated questions. the training set consisted of 45 papers as described above. the 33 literature review papers used for testing contained 534 citations, out of which 469 were extracted (see table 7). for example, this citation “heesang et al. (2007), based on ant colony behavior, proposed a new path-flow routing algorithm for the backbone network of the next generation networks.” is extracted by using our predefined regular expression for textual syntactic style. however, the citation is not valid in our case because the name entity recognizer couldn’t identify the heesang as a person name, and the author name is required in our question liu,calvo and rus 114 generation process. in addition to these 469 extracted citations, 20 non-citations were wrongly extracted as citations because of lbj ner tagger errors. for example, this citation “one of the examples is the unbalanced mach zehnder interferometer filter which based on the cascade of two couplers.” was wrongly identified as a citation since the name entity recognizer wrongly identified the mach zehnder as a person name. number of citations 534 extracted citations 469 citation extraction rate 88% table 7: citation extraction rate. within these 469 citations, 21 generated questions that had serious grammatical errors because of the semantic parser and sentence splitter’s performance. therefore, we only evaluated the statistical citation classifier using 448 citations. 4.3 citation classification performance this experiment examined whether our current method (using machine learning) achieves higher accuracy compared to liu et al.’s (2010) rule-based approach when running the same conditions. we also evaluated the performance on different learning algorithms. because in the previous study we found that the application category was frequent, we added this new citation category in our study. the testing dataset contained 448 citations which belong to one of the seven categories. we used balanced f1-score, precision and recall to measure the classifier’s performance. the f1 of each class is computed by the following formula: (1) the precision for a class is the number of true positives (i.e. the number of citation sentences correctly labeled as belonging to the positive class) divided by the total number of elements labeled as belonging to the positive class, while the recall for a class is the number of true positives divided by the total number of elements that actually belong to the positive class. table 8 shows the accuracy of the rule-base classifier compared to our method in terms of six common categories. model category the rule-based classifier current model (naïve bayes) p r f p r f opinion 0.49 0.44 0.47 0.76 0.75 0.75 result 0.90 0.52 0.66 0.76 0.93 0.84 system 0.80 0.32 0.46 0.85 0.41 0.55 method 0.73 0.28 0.4 0.78 0.75 0.76 aim 0.85 0.46 0.6 0.71 0.74 0.72 other 0.68 0.19 0.3 0.49 0.66 0.56 average 0.74 0.37 0.48 0.73 0.71 0.7 table 8: the accuracies of liu et al’s rule-based classifier and our method. p stands for precision, r recall and f f-score. 1 2 precision recallf precision recall ∗ ∗ = + g-asks 115 running the same conditions, our current machine learning approach achieved higher accuracy than the rule-based classifier based on the tregex expression described in the related work section. although the tregex expression rules are good at extracting the semantic features of a sentence, such as predicate verb, subject and object and matching the well-defined cue phrases, this method still has a problem when dealing with sentences with more complex syntax. we propose that a machine learning approach can handle this very well because only the individual np or vp are considered. we also evaluated the citation classifier’s performance with different classification algorithms. table 9 shows the citation classifier’s performance and how well each class was predicated by each classifier. we can see that aim, result, opinion, and application can be identified quite well using defined features, except in some cases where the function of a citation is implicit. for example, this is an implicit citation from the aim category: there has been much interest recently in devising strong game-theoretic strategies. we also observe that the system category is difficult to identify, because it mainly depends on the noun cluster feature containing words such as, system, tool, and agent, and a verb cluster feature including words such as devise, present, develop, and propose. when the citation contains a technical term and doesn’t contain any verb in the shared verb cluster, it would be difficult to identity. for example, the circsim-tutor [12] teaches cardiovascular physiology by describing… in this case, it is hard to identify circsim-tutor as a system name. classifier category naïve bayes support vector machine j48 decision tree p r f p r f p r f opinion 0.76 0.75 0.75 0.41 0.52 0.46 0.42 0.68 0.52 result 0.76 0.93 0.84 0.79 0.84 0.81 0.76 0.77 0.77 system 0.85 0.41 0.55 0.91 0.1 0.18 0.35 0.8 0.48 application 0.69 0.89 0.78 0.56 0.88 0.68 0.73 0.67 0.7 method 0.78 0.75 0.76 0.39 0.85 0.53 0.57 0.8 0.67 aim 0.71 0.74 0.72 0.56 0.78 0.65 0.65 0.55 0.6 other 0.49 0.66 0.56 0.85 0.01 0.02 0.73 0.19 0.3 average 0.72 0.73 0.71 0.60 0.66 0.48 0.6 0.64 0.58 table 9: the citation classifier’s performance with different learning algorithms. p stands for precision, r recall and f f-score. 4.4 question quality evaluation and result a total of 615 questions were generated based on the 33 literature review papers. table 10 shows that the supervisors generated 142 questions (107 citation related questions and 35 non-citation related), while the peers generated 151 questions (133 citation related questions and 18 noncitation related). the non-citation related questions addressed other types of writing feedback, such as clarity, organization, and referencing. liu,calvo and rus 116 question producer number of questions question types supervisor 142 citation related (107) non-citation related (35) peer 151 citation related (133) non-citation related (18) g-asks 161 (randomly sampled from 469 questions) 161 citation related generic system 161 161 generic total 615 562 questions were evaluated table 10. number of questions produced from 33 literature review papers in order to make it comparable, we only used the 562 citation-related questions by supervisors and peers and the 161 questions by the system as well as 161 generic questions. table 11 shows average scores of the questions according to different quality measures. supervisors’ questions (average score: 4.4) outscored g-asks questions (average score: 3.8), peer generated questions (average score: 3.77), and generic questions (average score: 3.73). quality measures producer qm1: correctness qm2: clarity qm3: context appropriate qm4: trigger reflection qm5: usefulness average g-asks 3.94 3.91 3.76 3.69 3.69 3.80 supervisor 4.57 4.53 4.45 4.26 4.20 4.40 peer 3.93 3.96 3.77 3.62 3.56 3.77 generic 4.04 3.92 3.65 3.49 3.57 3.73 table 11: comparisons of normalized mean scores a one-way anova was conducted to examine whether differences in the average score, as well as in each quality measure, were statistically significant. the anova yielded a significant difference in average (f(3,558)=12.15, p<0.05), qm1 (f(3,558)=9.254, p<0.05), qm2 (f(3, 558)=8.863, p<0.05), qm3 (f(3, 558)=12.552, p<0.05), qm4 (f(3, 558)=11.103, p<0.05), and qm5 (f(3, 558)=7.913, p<0.05). follow-up fishers’ least significant difference (lsd) tests with 95% confidence interval were performed to determine which pairs of treatments differed from one another. table 12 shows the mean differences (md) and lsd results from which we conclude that the questions from supervisors significantly outscored g-asks in all the measures. there were no statistically significant differences between questions generated by the peer and gasks, or between generic question and the g-asks. mean difference criteria g-asks vs supervisor g-asks vs peer g-asks vs gq qm1 md =0.632 md =0.006 md =0.099 lsd=0.259 lsd=0.246 lsd=0.236 qm2 md =0.626 md =0.056 md =0.012 lsd=0.264 lsd=0.251 lsd=0.239 qm3 md =0.691 md =0.009 md =0.112 lsd=0.270 lsd=0.256 lsd=0.245 qm4 md =0.572 md =0.065 md =0.199 lsd=0.267 lsd=0.254 lsd=0.243 qm5 md =0.516 md =0.126 md =0.118 lsd=0.280 lsd=0.266 lsd=0.255 average md =0.607 md =0.026 md =0.063 lsd=0.237 lsd=0.225 lsd=0.215 table 12: fisher's least significant difference (lsd) tests with 95% confidence interval. g-asks 117 because the automatically generated questions were randomly sampled, some had grammatical and semantic errors due to the performance of the semantic parser, the citation classifier and the lbj name entity tagger. this may have degraded the scores of the g-asks questions. we performed further analysis on the quality of the aqg system by excluding 35 questions with grammatical and semantic errors. as shown in table 13, supervisors’ questions still got the highest score in each quality measure, but the g-asks’s questions now takes the second place with an average score of 4.10. we performed post-hoc analysis using lsd tests as we did before. table 14 shows that the questions from the g-asks significantly outscored generic questions in each quality measure while outperforming peers’ questions in quality measures 1, 3, 4, 5 and average. as expected, the g-asks outscored generic questions because the contentrelated questions were more helpful than the generic questions. moreover, the system questions significantly outperformed human peer questions, which might be explained by the following factors. first, peers may not be familiar with the topic of the literature review paper which he or she reviewed so the peer cannot generate very useful questions. secondly, peers tend to generate conceptual questions which are not deep enough to trigger reflection. for example, this peer’s question what is the beginning process of mic that shown by little and lee? is only concerned with asking the student-author to identify the concept of the beginning process of mic. the difference between supervisor generated questions and g-asks questions for qm5 was not significant. this result indicates that questions produced by the g-asks system were perceived to be as useful as questions from human supervisors. this positive result might be explained by two factors. first, the system’s questions are useful because they are content-related and have semantic meaning. second, students may have intended to give high scores to questions which they thought were from supervisors, where actually they were from the system. in table 16, we can see that 21% (34 out of 161) of the g-asks questions have been wrongly identified as supervisor questions. criteria question producer qm1 qm2 qm3 qm4 qm5 average g-asks 4.26 4.18 4.06 3.99 4.02 4.10 supervisor 4.57 4.53 4.45 4.26 4.21 4.40 peer 3.93 3.96 3.77 3.62 3.56 3.77 generic 4.04 3.92 3.65 3.49 3.57 3.73 table 13: comparisons of normalized mean scores after removing questions with grammatical and semantic errors mean difference(md) criteria g-asks vs supervisor g-asks vs peer g-asks vs gq qm1 md =0.308 md =0.330 md =0.225 lsd=0.242 lsd=0.230 lsd=0.219 qm2 md =0.350 md =0.220 md =0.263 lsd=0.248 lsd=0.236 lsd=0.225 qm3 md =0.385 md =0.297 md =0.418 lsd=0.258 lsd=0.246 lsd=0.235 qm4 md =0.269 md =0.368 md =0.501 lsd=0.257 lsd=0.245 lsd=0.234 qm5 md =0.182 md =0.460 md =0.452 lsd=0.269 lsd=0.256 lsd=0.245 average md =0.299 md =0.335 md =0.372 lsd=0.226 lsd=0.215 lsd=0.206 table 14: fisher's least significant difference (lsd) tests with 95% confidence interval after removing questions with grammatical and semantic errors. liu,calvo and rus 118 rule criteria aim application method opinion result system qm1 4.08 4.35 4.43 4.10 4.42 4.10 qm2 4.00 4.30 4.36 3.90 4.27 4.40 qm3 3.88 4.13 4.14 3.81 4.24 4.20 qm4 3.88 4.04 4.07 3.90 4.03 4.10 qm5 3.92 4.13 4.00 3.81 4.09 4.30 average 3.95 4.19 4.20 3.90 4.21 4.22 table 15: comparisons of scores for each rule. the quality of each generation rule was also evaluated. table 15 shows the average scores. the rule for system got the highest average score (4.22). rules for result, application, and method obtained similar scores (above 4.19), while rules for aim and opinion obtained lower scores (below 4.0). because we used a five-point likert scale (3 means ‘neither agree nor disagree’ while 4 means ‘agree’), these scores for rules (system, result, application and method) indicate that evaluators agreed that those questions are correctly written, clear, appropriate, helpful to reflect and useful. the students’ evaluation results about the quality of the aim and opinion’s question template type were not as good as other types, but it was almost close to ‘agree with the good quality of questions’. 4.5 human perception evaluation and result for each of the questions, participants were asked to guess whether it was written by the supervisor, a peer, the g-asks (an intelligent computer system), or whether it was a generic question. we use the balanced f-score described in formula 1 to evaluate the classification. table 16 shows the participants' average performance on the classification, which found that they achieved an f-score of 0.49 on the supervisor category, 0.46 on the peer category, 0.34 on the gasks category, and 0.60 on the generic question category. generic questions were the easiest to identify, as we expected, while questions produced by g-asks were the most difficult. interestingly, 21% (34 out of 161) of the g-asks questions were wrongly identified as being questions from supervisors, and 40% (64 out of 161) were wrongly identified as questions from peers. there are two major reasons for this result: 1. the questions generated by the system are specific, especially related to the citations. 2. similar to peers and supervisors, the g-ask questions used abstract concepts (especially for application, result, system and method citations). such questions are also quite similar to our questions templates. for example, people often would like to ask studentwriters to critically evaluate the method/theory/system and explain why this method is useful to solve a particular problem. however, the human questions are more concise and correctly written than the system. some system questions with long length and grammatical errors could be easily identified. real prediction supervisor peer g-asks generic supervisor 74(52%) 40(27%) 34(21%) 11(7%) peer 41(29%) 82(54%) 64(40%) 16(10%) g-asks 14(10%) 23(15%) 51(32%) 51(32%) generic 13(9%) 6(4%) 12(7%) 83(51%) total 142 151 161 161 table 16: human classification result on authorship of questions. the accuracy is shown in percentage in brackets. g-asks 119 4.6 question types evaluation and result question types are important for the application of automatic question generation, for example it can help us to define the question types which the systems should generate from the source text. there were 142 questions generated by human supervisors from 33 literature review papers written by the students. these questions included 35 surface questions, which concern presentation issues such as formating, spelling, grammatical errors, and some generic questions. because we are only concerned about specific questions, the remaining 107 questions are analyzed as follows. two human annotators were asked to independently annotate these 107 questions generated by human supervisors. the annotation was based on 18 frequent question types in graesser and pearson’s taxonomy (graesser & person 1994). cohen’s kappa measures the agreement between two annotators: pr( ) pr( ) 1 pr( ) a e e κ − = − (1) where pr (a) is the relative observed agreement among annotators and pr (e) is the overall probability of random agreement. as a result, the cohen’s kappa coefficient is 0.57 (n=13; n=107; k=2), indicating moderate reliability considering the relatively large number of categories. annotation of verification, causal and procedural questions were considered to be relatively reliable with more than 74% agreements between annotators, versus only 23% for judgmental questions, 69% for concept questions and 50% for comparison questions. table 17 shows seven frequent question types in our dataset. the left column gives the definition of each question type while the right column shows an example question generated by human supervisors. type 1 and 2 questions were classified as simple/shallow, 3 as intermediate, and 4-7 as deep questions in relation to bloom’s taxonomy of cognitive domain (bloom 1984). question type examples 1 verification: implied yes/no/ answers (shallow) is it possible to reuse some of previous routing techniques, for example those used in cellular networks, in the ngmn? 2 concept: who, when ,what, where? (shallow) can you give more details about the generalized beam theory? 3 comparison: how is x similar to y? (intermediate) in lim and nethercot, how well did the numerical results compare with the experimental results? 4 causal antecedent: what event causally led to an event? why network coding in [13] can increase the system throughput? 5 causal consequence: what is the consequence of an event? (deep) what is the likely consequence of the nonlinear stressstrain curve on the local-overall interaction buckling behavior of stainless steel structural members? 6 procedural: what instrument or plan allows an agent to accomplish a goal? (deep) how does the formation of mechanical twins provide corrosion resistance? 7 judgmental: what do you think of x?(deep) how do you see the generalized beam theory being applied in your project? table 17: graesser and pearson’s question taxonomy with examples of questions from academic supervisors table 18 shows the frequency of each of question type. as we can see, concept, causal, and procedural questions were more frequent than judgmental and verification questions. liu,calvo and rus 120 question category frequency concept 28 causal antecedent/consequence 28 procedural 23 comparison 2 judgmental 13 verification 11 other types: feature specification 1, goal orientation 1 table 18: frequency of the question type table 19 shows the average scores from students’ evaluations of the questions. verification, concept and procedural questions obtained slightly higher scores than causal and judgmental questions. however, there were no statistically significant differences between scores from these questions types (f (4,100) =2.162, p>0.05). question type average score standard deviation concept 4.12 0.993 causal 3.88 1.021 procedural 4.34 0.742 judgmental 3.92 1.145 verification 4.75 0.430 table 19: the average score of the question type from table 18, we see that human supervisors like to generate simple questions in addition to deep questions. it indicates that conceptual questions are as important as procedural or causal questions, which should be considered when designing the question templates. in order to investigate how the question was generated from the source text, we classified the source of questions into the following four abstract levels: lexicon level: the question is generated from a key term or concept. e.g., can you give more details about the generalized beam theory, explain their advantages and disadvantages? the generalized beam theory is the key concept, which is asked to be analyzed critically. sentence level: the question is generated from a single sentence without domain knowledge. e.g., in lim and nethercot, how well did the numerical results compare with the experimental results? the source sentence was: lim and nethercot compared the numerical tests results with finite element simulation results. in this case, the source sentence reports comparative tests and the trigger question is about how good the results are. discourse level: the question is generated from more than one sentence with inference process and domain knowledge. e.g., on what basis did besson et al 1997 compare the solid state and liquid state production of pyrazine? the source sentences were besson et al. (1997) demonstrated that pyrazine can be produced by solid state fermentation. they stipulate the concentrations which are much higher than that of liquid state fermentation. in this case, we should combine two sentences together with some domain knowledge and infer that pyrazine can be produced from both solid state and liquid state, and the solid state is better. g-asks 121 background knowledge: the question is generated based on the domain knowledge which is not expressed in the writing. e.g., what is the power range for each type of the wind turbine? in this case, we should know that the power range is one property of wind turbine. source level num. of questions and question type distribution(only show question type more than 2 times ) lexicon level 12 questions include 7 concept, 5 judgmental sentence level 49 questions include 20 causal, 10 procedural, 10 concept, 5 judgmental, 4 verification. discourse level 14 questions include 4 procedural, 3 verification background knowledge 32 questions include 9 concept, 7 procedural, 4 causal, 3 verification. table 20 frequency of source level for question generation table 20 shows that the number of questions generated at the lexicon and sentence levels take 57% (61 out of 107) while the discourse level takes 13.1% and background knowledge level takes 29.9%. the dominant question types were concept and judgmental in the lexicon level; causal, procedure, and concept in the sentence level; procedural and verification in the discourse level; and concept and procedure in the background level. this indicates promising opportunities for generating questions from lexicon and sentence level by using current nlp technologies. 5 conclusion and future work this article presented an intelligent automatic question generation system, as a feedback tool used in the iwrite web application, which generates contextualized trigger questions from citations to support literature review writing. the trigger questions are aimed at improving learning during academic writing as opposed to only focusing on writing as a communication skill. in order to evaluate the system’s performance and to analyze human expert generated questions, we compared automatically generated questions with human-generated and generic questions using a bystander turing test. in contrast to the citation classification with tregex rule-based approach used in the previous study, for the present study we used a machine learning approach based on some useful features, such as cue phrases. the result shows the new approach outperformed the rule-based approach across 5 citation categories. however, from the result of the basic system performance (see section citation extraction rate and citation classification performance), we can see the bottlenecks of the pipeline system were in the ner (name entity recognizer) tagger, the sentence parser and the statistical citation classifier. the lbj ner tagger was primarily trained on news text corpora, and this might have affected its performance on academic articles. as for the performance of the sentence parser and citation classifier, there is room for improvement. these issues caused the system to generate semantically or syntactically erroneous questions and hence decreased the overall quality of the system. however, these issues can be overcome by a ranking function which would reduce the probability of questions with semantic or syntactic errors being selected. in order to get more insights on how human experts generate specific trigger question, we analyzed the human supervisors’ generated questions for literature review writing support. six frequent question types based on graesser and person’s question taxonomy were identified. this can help us design question templates. the results show that 57% of the questions generated at liu,calvo and rus 122 the lexicon and sentence levels were without complex inference processing, which indicates that many potential questions can be exploited by using current nlp techniques. future work will focus on using ‘the overgenerate-and-rank’ approach, which has been applied previously in the natural language generation community (langkilde & knight 1998; walker et al. 2001). as shown in table 20, conceptual questions are also important; these questions were rated well by students (achieving an average score of 4.12) and placed second among all question types. we are working to generate conceptual question types based on key concepts in a single document. acknowledgment the authors would like to thank our colleagues jorge villalon and stephen o'rourke for the development of the tml java library. this project was partially supported by a university of sydney ties grant, an australian research council discovery project grant (dp0986873) and google research award for measuring the impact of feedback on the writing process. this research was supported in part by the institute for education sciences (r305a100875) and national science foundation (0938239) through grants awarded to dr. vasile rus. references liane acker, james lester, art souther and bruce porter, eds. (1991). generating coherent explanations to answer students' questions. intelligent tutoring systems: evolutions in design, psychology press, new york. jeff anderson (2005). mechanically inclined:building grammar, usage, and style into writer's workshop. stenhouse publishers portland, me. timothy j. beals (1998). between teachers and computers: does text-checking software really improve student writing? the english journal, 87(1): 67-72. benjamin bloom (1984). taxonomy of educational objectives book i: cognitive domain. addison wesley, boston. kristy e. boyer, william lahti, robert phillips, michael wallis, mladen vouk and james lester (2009). an empirically-derived question taxonomy for task-oriented tutorial dialogue. in proceedings of the 2nd workshop on question generation, pages: 9-16, brighton. m. anne britt, peter wiemer-hastings, aaron a. larson and charles a. perfetti (2004). using intelligent feedback to improve sourcing and integration in students' essays. international journal of artificial intelligence in education, 14(3): 359-374. rafael a. calvo and robert a. ellis (2010). students' conceptions of tutor and automated feedback in professional writing. journal of engineering education, 99(4): 427-438. rafael a. calvo, stephen t. o'rourke, janet jones, kalina yacef and peter reimann (2010). collaborative writing support tools on the cloud. ieee transactions on learning technology, 4(1): 88-97. david coniam (1997). a preliminary inquiry into using corpus word frequency data in the automatic generation of english language cloze tests. the computer assisted language instruction consortium journal, 14(2): 15-33. averil coxhead (2000). a new academic word list. teachers of english to speakers of other languages journal quarterly, 34(2): 213-238. andrea esuli and fabrizio sebastiani (2006). sentiwordnet: a publicly available lexical resource for opinion mining. in proceedings of the 5th conference on language resources and evaluation, pages: 417-422, genoa, italy. g-asks 123 christiane fellbaum (1998). wordnet: an electronic lexical database. the mit press, cambridge, ma. arthur c. graesser and natalie k. person (1994). question asking during tutoring. american educational research journal, 31(1): 104-137. michael heilman and noah a. smith (2009). good question! statistical ranking for question generation. in proceedings of the 2010 annual conference of the north american chapter of the association for computational linguistics, pages: 609-617, stroudsburg, pa. k. hyland (1994). academic attribution: citation and the construction of disciplinary knowledge. applied linguistics, 20(3): 341-367. dan klein and christopher d. manning (2003). fast exact inference with a factored model for natural language parsing. in proceedings of the international conference in advances in neural information processing systems, pages: 3-10, cambridge, ma. hidenobu kunichika, tomoki katayama, tsukasa hirashima and akira takeuchi: (2001). automated question generation methods for intelligent english learning systems and its evaluation. in proceedings of the international conference on computers in education, pages: 1117-1124, seoul, korea. thomas k. landauer, danielle s. mcnamara, simon dennis and walter kintsch (2007). handbook of latent semantic analysis. lawrence erlbaum. irene langkilde and kevin knight (1998). generation that exploits corpus-based statistical knowledge. in proceedings of the 36th annual meeting of the association for computational linguistics and 17th international conference on computational linguistics, pages: 704-710, montreal, quebec. wendy lehnert, claire cardie and ellen riloff (1990). analyzing research papers using citation sentences. in proceedings of the twelfth annual conference of the cognitive science society, pages: 511-518, cambridge, ma. wendy g. lehnert (1978). the process of question answering a computer simulation of cognition. hillsdale. l. erlbaum associates, hillsdale, nj. roger levy and galen andrew (2006). tregex and tsurgeon: tools for querying and manipulating tree data structures. in proceedings of the fifth international conference on language resources and evaluation, pages: 2231-2234, genoa, italy. ming liu, rafael a. calvo and vasile rus (2010). automatic question generation for literature review writing support. in proceedings of the tenth interational conference on intelligent tutoring systems, pages: 45-54, pittsburgh, usa. ruslan mitkov and le an ha (2003). computer-aided generation of multiple-choice tests. in proceedings of the hlt-naacl 03 workshop on building educational applications using natural language processing, pages: 17-22, morristown, nj. rodney d. nielsen, jason buckingham, gary knoll, ben marsh and leysia palen (2008). a taxonomy of questions for question generation. in proceedings of the 1st workshop on question generation., pages: 15-22, arlington, virginia. natalie k. person and arthur c. graesser (2002). human or computer? autotutor in a bystander turing test. in proceedings of the 6th international conference on intelligent tutoring systems, pages: 821--830, london, uk. brett powley and robert dale (2007). evidence-based information extraction for high-accuracy citation extraction and author name recognition. in proceedings of the 8th riao international conference on large-scale semantic access to content, pages: 15-22, pittsburgh, pa. lev ratinov and dan roth (2009). design challenges and misconceptions in named entity recognition. in proceedings of the thirteenth conference on computational natural language learning, stroudsburg, pa. thomas h. reynolds and curtis jay bonk (1996). computerized prompting partners and keystroke recording devices: two macro driven writing tools. educational technology research and development, 44(3): 83-97. liu,calvo and rus 124 mark d. shermis and jill c. burstein (2002). automated essay scoring: a cross-disciplinary perspective. routledge, mahwah, nj. simone teufel, advaith siddharthan and dan tidhar (2006). automatic classification of citation function. in proceedings of the international conference on empirical methods in natural language processing, pages: 103-110, sydney, australia. elaine c. thiesmeyer and john e. theismeyer (1990). editor:a system for checking usage, mechanics, vocabulary, and structure. new york: modern language association, raleigh, nc. jorge villalon, paul kearney, rafael a. calvo and peter reimann (2008). glosser: enhanced feedback for student writing tasks. in proceeding of eighth ieee international conference on advanced learning technologies, pages: 454-458, santander,spain. marilyn a. walker, owen rambow and monica rogati (2001). spot: a trainable sentence planner. in proceeding of the north american chapter of the association for computational linguistics on language technologies, pages: 1-8, pittsburgh, pennsylvania. peter wiemer-hastings and arthur c. graesser (2000). select-a-kibitzer: a computer tool that gives meaningful feedback on student compositions. interactive learning environments, 8(2): 149--169. robert williams and heinz dreher (2004). automatically grading essays with markit©. in proceedings of informing science conference, pages: 0693-0700, rockhampton, queensland. john h. wolfe (1976). automatic question generation from text an aid to independent study. in proceedings of the acm sigcse-sigcue technical symposium on computer science and education, pages: 104-112, new york. ilkin-sturt.dvi dialogue and discourse 2(1) (2011) 35–58 doi: 10.5087/dad.2011.103 active prediction of syntactic information during sentence processing∗ zeynep ilkin zeynepilk@gmail .com psychology university of edinburgh 7 george square edinburgh eh8 9jz patrick sturt patrick.sturt@ed.ac.uk psychology university of edinburgh 7 george square edinburgh eh8 9jz editor: abstract we describe an eye-tracking experiment that tested the effect of syntactic predictability on skipping rates during reading. we found that plural noun phrases wereskipped more often than singular noun phrases, in syntactic contexts which induced a high expectation for a plural. we interpret this effect as evidence that the plural noun phrase has been predicted ahead of time. the results indicate that the examination of skipping rates might be a useful tool for the investigation of syntactic prediction effects. keywords: incrementality, parsing, eye-tracking, reading, prediction 1. introduction successful language comprehension requires the incremental integration of partial linguistic information. for example, in dialogue, part of an utterance may be inaudible due to noise in the environment; a speaker’s utterance may stop (or be continued by an interlocutor) before it is formally complete; or a speaker may re-start an utterance mid-stream. in order for communication to be successful in these situations, the processing mechanism must be extremely flexible: incomplete pieces of structure may have to be integrated in order to create partial interpretations, or missing input items may have to be inferred from the information available. in contrast, previous theories of human parsing have often lacked this type of flexibility, due to the adoption of bottom-up processing stategies. such theories assume that the basic building blocks of interpretation are complete syntactic constituents, orthat the building of structure requires the presence of a licensing phrasal head in the input (mulders, 2002; pritchett, 1991). the prediction is that there is a limit to the degree to which partial linguistic information can be used, because some crucial part of a constituent, for example, its lexicalhead, needs to be recognised before that ∗. this research was presented at the cuny sentence processing conference, in davis, ca, 2009; at the zif workshop on incrementality and verbal interaction, in bielefeld, germany, 2009; and at the european conference on eye movements, in southampton, uk, 2009. we thank the audiencesof these conferences, as well as three anonymous reviewers, for their comments. c©2011 zeynep ilkin and patrick sturt submitted 1/2010; accepted 12/2010; published online 5/2011 ilkin and sturt constituent can be integrated into the previous syntactic context. such theories would require extra mechanisms to explain the ease with which partial utterances are integrated in dialogue, and they are also incompatible with empirical evidence in reading research. for example, evidence from the processing of japanese, a verb final language, shows thatlong-distance dependencies can be formed into a subordinate clause before the verbal head of that clause is recognised in the input (aoshima et al., 2004), and analogous effects have been found in english (lee, 2004). these findings would not be expected if the bottom-up input of the verb were necessary for the formation of the dependency. in contrast to the data-driven character of processing in bottom-up models, recent work has argued for a greater role for top-down prediction in human parsing. the idea is that comprehenders activate representations of possible continuations of thesentence, and these are constantly updated as each new word comes in (see, for example, levy (2008) for a probabilistic model incorporating this idea). predictive processing has an obvious relevancefor explaining the flexibility of dialogue processing (pickering and garrod, 2007): some such processmust underlie the ability of interlocutors to complete each other’s utterances. moreover, incases where bottom-up information is unavailable due to noise in the environment, comprehendersmay be forced to rely on top-down information. despite the intuitive appeal of predictive processing for modelling human linguistic performance, unambiguous experimental evidence for prediction has been surprisingly elusive. as theoretical processing models become more sophisticated, theybegin to make fine-grained predictions about the time-course of prediction, and about the level of representation that is predicted. it then becomes increasingly important to develop experimental techniques that test prediction at such finegrained levels. the problem is that, although a large body ofwork has demonstrated processing facilitation for words that are expected given the context,it is not always possible to interpret such results as evidence for the prediction of the word ahead of time. for example, studies have shown that lexical decision times are faster when the target word is predictable given the context than when it is not (relative to controls). schwanenflugel and shoben (1985) recorded lexical decision times for the final (underlined) word in visually presented stimuli such as the following, where the word was preceded by a separately presented context: (1) a. high constraint: expected word the worker was criticised by his boss b. neutral baseline xxx xxxxxx xxx xxxxxxxxxx xx xxx boss c. high constraint: unexpected word the worker was criticised by his manager d. neutral baseline xxx xxxxxx xxx xxxxxxxxxx xx xxx manager lexical decision responses to the expected word (bossin this case), were faster in the high constraint context relative to a baseline context consisting of a row of x’s, while responses to the unexpected word (manager) did not differ from the baseline context. it is possible to interpret such findings in terms of predictive mechanisms. according to such an interpretation, the highly constraining context leads to the activation of the expected word’s lexical entry ahead of time, thus allowing the recognition of the word to occur faster when it eventually appears in the input. 36 syntactic prediction however, another possible interpretation of results such as these is that people build up a situation model based on the context sentence, and the expected word issimply easier to integrate into this situation model than the unexpected word (traxler and foss,2000). crucially, this explanation does not rely on the pre-activation of the test word, and thus the experiment does not afford unambiguous evidence of prediction. similar alternative interpretations can be made for demonstrations of predictive processing in other domains, such as syntax and reference. van gompel and liversedge (2003) recorded participants’ eye-movements while they read sentences containing cataphoric pronouns, such as the following: (2) a. congruent: masculine when he was fed up, the boy visited the girl very often. b. congruent feminine when she was fed up, the girl visited the boy very often. c. incongruent: masculine when she was fed up, the boy visited the girl very often. d. incongruent feminine when he was fed up, the girl visited the boy very often. the stimuli always included a gender-marked pronoun in a preposed adverbial clause (he or she), which was either congruent or incongruent with the genderof the main clause subject (the boyor the girl in this case). eye-movement measures showed that participants slowed down at or immediately after the main clause subject in the incongruent conditions relative to the congruent conditions. van gompel and liversedge (2003) argued that the processor forms the referential dependency before the gender information of the main clausesubject has been computed, resulting in a slow-down in the incongruent conditions, where the dependency has to be abandoned. however, such findings are also compatible with a highly predictive processing strategy, according to which the main clause subject position is predicted ahead of time,while the initial subordinate clause is being processed. the referential dependency is then formedbetween the cataphoric pronoun and this predicted subject position. this then causes processing difficulty in the incongruent conditions, due to the gender mismatch between the features of the predicted subject position and those of the head nounboy/girl . this type of predictive strategy has been suggested by kazanina et al. (2007), and also by kreiner et al. (2008), both of whom report similareffects. kazanina et al. (2007) additionally showed that the difference between congruentand incongruent conditions disappears when the subject is ruled out as an antecedent position by binding principles, as in (3). (3) a. principle c it seemed worrisome to him that john/ruth was gaining so muchweight, but matt didn’t have the nerve to comment on it. b. no constraint it seemed worrisome to his family that john/ruth was gainingso much weight, but matt didn’t have the nerve to comment on it. in (3a), binding principle c (chomsky, 1981) rules out the subordinate clause subject (john/ruth) as an antecedent ofhim, while this is not the case in (3b). kazanina et al. (2007) demonstrated a 37 ilkin and sturt slow-down in self-paced reading times for the incongruentruthrelative to the congruentjohn in (3b), but no such congruency effect was found in (3a), where the dependency was not licit in terms of binding theory. kazanina et al. (2007) interpreted theseresults in terms of a predictive processing strategy which they termactive search, in which the relevant position for the antecedent phrase is predicted ahead of time, but where the referential link is only made where this predicted phrase is a licit antecedent in terms of binding theory. kreiner et al. (2008) also argued for a predictive processing strategy on the basis of experimental data on cataphoric reference. in their experiment 2, kreiner et al. (2008) recorded eye-movements while participants read sentences like those in (4): (4) a. stereotypical gender: congruent after reminding himself about the letter, the minister immediately went to the meeting at the office. b. stereotypical gender: incongruent after reminding herself about the letter, the minister immediately went to the meeting at the office. c. definitional gender: congruent after reminding himself about the letter, the king immediately went to the meeting at the office. d. definitional gender: incongruent after reminding herself about the letter, the king immediately went to the meeting at the office. the design of kreiner et al. (2008) incorporated both stereotypical gender-biased nouns (e.g.minister in (4a,b) is likely to be a man, but does not have to be), and also definitional gender nouns (e.g. king in (4c,d) has to denote a male by definition). the stereotypical or definitional gender could either be congruent or incongruent with a preceding reflexive, which had to be co-referential with the main clause subject. kreiner et al. (2008) found a congruency effect for the definitional gender conditions, with longer reading times immediately following king in (4d) than in (4c). however, no such difference was found between the two stereotype gender conditions (4a) and (4b). kreiner et al. (2008) interpreted this result in terms of a predictive processing strategy in which the dependency betweenhimself/herselfand the main clause subject was formed in advance, constraining the gender of the subject before its appearance in the input. thecongruency effect for the definitional nouns can be explained on the assumption that the gender information for these nouns is retrieved at lexical access, and this causes a slow-down due to gender mismatch in the incongruent condition. the lack of such a congruency effect for the stereotypical nouns can be explained on the assumption that the gender information for these nouns is determined through fit with the context, or via stereotype inference, rather than through the retrieval of a gender feature at lexical access. the predictive dependency formation strategy results in the gender of the main clause subject being determined before the stereotype role name is processed in the input. since the gender is already available, there is no need for a stereotype inference, and the lack of mismatch effect for (4a) and (4b) can be explained. however, although the notion of predictive dependency formation is compatible with the results of both kazanina et al. (2007) and kreiner et al. (2008), thisis not the only possible explanation. in both cases, it is possible that the referential dependency was formed only after the head noun 38 syntactic prediction of the antecedent noun phrase had been processed in the input. the congruency effect reported by kazanina et al. (2007) (see (3) above) could have arisen if the parser did not predict the position of the antecedent of the cataphoric pronoun in advance. according to this non-predictive account, the processor would form the referential dependency when the critical word of the antecedent phrase (john/ruth) has been processed. for example, the processor might attempt to update co-reference relations whenever any potentially referential noun phrase is read. it might be at this point that the structural details relevant to binding constraints are taken into account, and referential dependencies are considered only where allowed by binding theory. onthis account, the gender congruency effect that kazanina et al. (2007) obtained could be explained in a number of ways; for example, as suggested by van gompel and liversedge (2003), it might bethe case that the referential dependency is formed after the part-of-speech information ofthe antecedent is retrieved, but before the gender information has been retrieved, due to architectural constraints on the timecourse of the availability of linguistic information. this would lead tothe processing difficulty in the incongruent condition, since the dependency is formed before the genderis checked. an alternative possibility is that the processing difficulty arises from competition among simultaneously active constraints, as in constraint satisfaction models of parsing (mcrae et al., 1998); in the case of the incongruent condition, the gender constraint would lead to a bias against the referential dependency, while other constraints, such as, for example, the subject preference for pronoun antecedents would lead to a bias in favour of the referential dependency. the competition between these constraints would explain the processing difficulty. the results of kreiner et al. (2008) (see example (4) above) could also be explained in terms of a non-predictive processing strategy. again, accordingto this account, it is assumed that the processing of the potential referential antecedent (in this case, the main clause subject) triggers an update of co-reference information. it is at this point thatthe processor registers the possibility of a co-reference relation between the cataphoric reflexiveand the main clause subject (in fact, in this case, the co-reference is grammatically obligatory).the congruency effect for the definitional gender nouns can be explained as before, if we assume either an architectural constraint on the temporal order of linguistic information, or alternatively, competition among simultaneously active constraints. the lack of a congruency effect for the stereotype nouns could be explained if the referential dependency is formed, and the gender of the mainclause subject determined, before any attempt to execute the stereotype gender inference. the above literature review has examined three examples of effects that have been interpreted in terms of predictive processing, where alternative non-predictive accounts are possible. the common factor in all of these studies, as well as many others in the literature, is that the experimental manipulation allows processing differences to be observed only at, or following, the predicted element. for example schwanenflugel and shoben (1985) tested predictionby measuring lexical decision times to the predicted word. similarly, kazanina et al. (2007) andkreiner et al. (2008) tested prediction by measuring reading times on a phrase whose gender feature had been putatively predicted. in all such cases, the evidence for prediction consists of facilitation of the processing of the predicted element and/or processing difficulty for the unpredicted element. there are at least two different ways in which one can make a more watertight case for prediction. the first is an argument based on the speed with which therelevant effects are observed. for example, lau et al. (2006) conducted an event related potential (erp) study in which an elan effect (early left anterior negativity) was elicited for a (normally) ungrammatical sequence consisting of a possessor followed by the wordof (e.g. . . . max’s of . . .). lau et al. (2006) manipulated 39 ilkin and sturt whether or not the sequence was licensed by the possibility of ellipsis (e.g.although erica kissed mary’s mother, she did not kiss dana’s . . .). the results showed that the licit and illicit conditions began to diverge at around 200 msec after the onset of the critical stimulus. taking into account previous estimates of the timecourse of lexical access at around 100-200 msec (sereno and rayner, 2003), the authors argued that the early occurrence of the elan effect was unlikely to have been observed unless the ellipsis had been predicted ahead of time. a second type of evidence that has been used to argue for predictive processing has involved demonstrations that the predicted element affects processing before it has been processed in the input. in the following paragraphs, we will describe two recent lines of research which take this approach. in an event related potential (erp) study, delong et al. (2005) examined the processing of sentences such as (5), with a word-by-word visual presentation: (5) a. the day was breezy so the boy went outside to fly a kite. b. the day was breezy so the boy went outside to fly an airplane. in (5), the context induces a strong expectation for the wordkite, as verified by a norming study reported by delong et al. (2005). this expectation is fulfilled in (5a), but not in (5b), whereairplane is not the most strongly expected continuation. in the study, delong et al. (2005) sought to demonstrate that the specific word formkite had been predicted before its appearance in the input. the study exploited the fact that the form of the english indefinite differs according to whether the following word begins with a vowel (an) or a consonant (a). in (5), the expected wordkite begins with a consonant, which is consistent with the definite article a, but inconsistent with the forman. the erp analysis focused on responses to the indefinite article. at the indefinite article, a greater negative deflection was found foran in (5b) than fora in (5a), and this effect was consistent with the n400 erp component, which is known to be sensitive to the degree of contextual expectation of a word. moreover, the authors demonstrated that the size of this n400 effect correlated reliably with the degree of expectation as measured by the prior norming results. thus, the authors found clear evidence of anticipation of a specific wordbeforethat word had been encountered in the input. using a similar experimental logic, van berkum et al. (2005)examined the erp responses to dutch inflected pre-nominal adjectives as a function of the degree of expectation of an up-coming noun. they found a positive deflection in conditions where the pre-nominal adjective did not agree in gender with the predicted noun, relative to conditions where the adjective agreed with the predicted noun. this is therefore also evidence that the prediction is active before the predicted element is encountered in the input. a second line of research showing anticipation before the predicted element comes from the visual world paradigm. in this technique, participants look at a depicted scene while they listen to a spoken sentence. the relative proportions of looks to various target and distractor objects in the depicted scene are analysed over time, as a function of manipulations in the spoken sentence. altmann and kamide (1999) showed that people began to look atobjects that would be expected given the sentential context, even before the relevant object was mentioned in the spoken sentence. for example, in a sentence likethe boy will eat the cake, participants’ looks to a depicted cake began to increase soon around the offset of the wordeat, relative to a condition where the spoken sentence did not lead to the expecation of the wordcake(i.e. the boy will move the cake). this experiment, as well as others using the visual world paradigm (kamide et al., 2003) have shown that people can use linguistic knowledge to anticipate the mention of visually presented objects ahead of 40 syntactic prediction time. however, it is currently not known whether these typesof effects rely on the presence of the visually presented objects in the scene. the presentation of a small number of potential referents at the start of each trial might effectively narrow down the possibilities for anticipation to a degree that allows the anticipatory effects to be observed, but such effects might not occur when the visual objects are not present. to summarise, the above literature review has highlighted some of the difficulties involved in interpreting experimental evidence for prediction in sentence processing. the most convincing evidence comes from studies in which effects of prediction are demonstrated before the predicted element is encountered in the string. it is also possible, given assumptions about the time-course of processes involved, to argue for predictive mechanisms based on the early appearance of congruency effects, as argued by lau et al. (2006). in the experiment reported below, we examine the effect of cataphoric pronouns on the prediction of features of their antecedents, as did kreiner et al. (2008) and kazanina et al. (2007). however, we use a new source of evidence, namely word-skipping in reading, to examine the prediction of syntactic information in sentence processing. we argue that the use of this measure makes a more convincing case for predictive effects in parsing. 2. experiment the experiment used eye-tracking in reading to test the extent to which people maintained expectations of plural morphological information on a noun phrase before the phrase had been read. as well as the usual eye-movement measures, which involve measurement of fixation time on particular words or phrases of interest, here we make crucial use of the measurement ofskippingrates as a measure of predictability. in eye-movement research, the term “skipping” refers to the phenomenon whereby a word, or a sequence of words is not fixated directly during initial reading, but instead, the reader makes an eye-movement that “skips” over the word.the “skipping” eye-movement is launched from material preceding the critical word, and it lands on material that follows the critical word. around one in three words are skipped in normal readingin english (brysbaert et al., 2005). previous studies have already established a link between lexical predictability and word skipping (balota et al., 1985; rayner et al., 2004). for example, rayner et al. (2004) recorded eye-movements while participants read sentences like (6): (6) a. predictable: before warming the milk, the babysitter took the infant’sbottleout of the travel bag. b. unpredictable: to prevent a mess, the caregiver checked the infant’sbottlebefore leaving. off-line tests established that the wordbottle was a very frequent continuation of the sentence prefix in (6a), and that it was a very rare continuation in (6b). in the eye-tracking experiment, readers skipped over the critical wordbottle more often in the predictable condition than in the unpredictable condition. in other words, the proportion oftrials in whichbottle received a fixation during initial reading was greater in the unpredictable condition than the predictable condition. in the e-z reader model of eye-movement control (reichle et al., 2003) skipping occurs in cases where, during a fixation on word n, attention shifts to word n+11, and lexical processing for n+1 1. word n+1 refers to the word immediately to the right of word n. 41 ilkin and sturt is initiated. if the first stage of lexical processing of n+1 is completed quickly enough, an eyemovement is programmed to word n+2, and the previously programmed eye-movement to word n+1 is cancelled. the eye-movement launched from word n to word n+2 results in the skipping of word n+1. in the study of rayner et al. (2004), the higher skipping rates for the predictable condition relative to the unpredictable condition can be explained if we assume that the critical word is partially activated ahead of time in the predictable condition, due to the contextual constraint. when the reader is fixating the word immediately preceding the critical wordbottle, the increased activation ofbottleallows the first stage of lexical processing of this word to becompleted relatively quickly, increasing the probability of skipping the word relative to the unpredicted condition. in the current experiment, we used a similar way to measure predictability with eye-movements, except that in this case, we measure the prediction of a syntactic feature, namely a plural number feature on a phrase, rather than the prediction of a specific word. the stimuli were designed to give a very high expectation fora plural noun phrase. this was done using cleft sentences with reciprocal anaphors, as in (7): (7) a. it was to each other that the girls from the school said that the children from next door wanted us to give advice. b. it was to each other that the girl from the school said that the children from next door wanted us to give advice. due to binding constraints, the reciprocaleach otherrequires a locally c-commanding plural noun phrase as its antecedent. in the cleft construction in (7), this requirement is satisfied via an unbounded dependency, witheach otheras the dependent element. we assumed that this would lead to a strong expectation for a plural subject of thethat-clause, as inthe girls in (7a). however, due to the fact that this requirement is licensed via an unbounded dependency, the plural antecedent may be more deeply embedded in the structure, and the subject of the that-clause is not grammatically constrained to be the antecedent. thus, it is grammaticallypossible for the subject of thethat-clause to be plural, as in (7a), or singular, as in (7b). in both cases, the globally correct antecedent is found more deeply embedded in the sentence (us). although both (7a) and (7b) are both grammatical, only (7a) satisfies the strong expectation for the subject ofthethat-clause to be plural. thus, we expected (7b) to elicit processing difficulty relative to (7a). however, as we have seen in the review above, such evidence of processing difficulty at or following the predicted element would not necessarily imply that the plural feature of the subject is predicted ahead of time. in this experiment, we therefore measured eye-movement behaviour that was initiatedbefore the critical noun phrase is fixated in the string, in additionto the standard eye-movement measures based on reading time. specifically, we measured the proportion of trials in which readers skipped the phrasethe girls in initial reading, without fixating it. if the skipping rates forthe girl(s) vary by condition, this can be interpreted as evidence that some relevant information from this phrase has been processed by the cognitive system parafoveally, without the reader fixating it. given that any skipping of thephrasethe girls must have been launched from a fixation position earlier than the determiner the, and also given that the two conditions differ only in the morphology of the second word of the phrase (in this case, the presence vs. absence of the plural marker “s”), any such result must mean that this morphological information has been processed while the reader is fixating a position that is at least two words to the left of the predicted head noun. in cases where the number feature is highly predictable, we assume that the plurality of the subject noun phrase is computed ahead of time, facilitating the parafoveal processing 42 syntactic prediction of the phrase, and increasing the probability that it is skipped. we will return to the mechanisms by which skipping might be affected by syntactic parsing mechanisms in the general discussion. a second aim of the experiment was to assess the role of pre-verbal dependency formation in english. recall that aoshima et al. (2004) showed evidence for long distance dependency formation in advance of the verb in japanese sentence processing, which they interpreted as evidence against the head-driven strategy for japanese. the current experiment allows us to conduct a similar test for english, which, unlike japanese, has a predominantly head-initial constituent order. specifically, the experiment allows us to test whether the dependency betweeneach otherandthe girls in (7) is made in advance of the verbsaid. this result would imply considerable postulation of structure in advance of bottom-up evidence, given that the dependency betweeneach otherandthe girlsrequires (a) an appropriate binding configuration, and (b) an unbounded dependency, in whichto each other is associated with an underlying indirect object position. 2.1 method 2.1.1 participants thirty two native speakers of english who were members of theuniversity of edinburgh community participated to the experiment. they were each paid£4 to participate in the experiment. all had normal or corrected to normal vision, and all were naı̈ve to the purpose of the experiment. 2.1.2 materials experimental materials consist of 28 sets of sentences like(8) in a 2×2 factorial design, orthogonally manipulating the form of the initial pp (conjunctionvs. reciprocal) and the number of the subject of thethat-clause. for simplicity of exposition, we will refer to the reciprocal plural and singular conditions asmatchandmismatchrespectively. we will refer to the two experimental factors aspp-typeandnumberrespectively. (8) a. reciprocal: plural (match) it was to each other that the girls from the school said that the children from next door wanted us to give advice. the children liked us. b. reciprocal: singular (mismatch) it was to each other that the girl from the school said that thechildren from next door wanted us to give advice. the children liked us. c. conjunct: plural it was to john and mary that the girls from the school said thatthe children from next door wanted us to give advice. the children liked us. d. conjunct: singular it was to john and mary that the girl from the school said that the children from next door wanted us to give advice. the children liked us. the two conjunct conditions were included in the design for experimental control. they incorporated a conjoined noun phrase instead of a reciprocal. the reason for this is that the skipping rates for a plural noun likegirls might differ from those for a singular noun likegirl , irrespective of whether or not the noun has been predicted, for example, due to length differences. the conjoined conditions allow us to measure the skipping rates for exactly the same critical nouns (girl /girls), but 43 ilkin and sturt without the highly predictive context. moreover, it is alsopossible that the fact that the conjoined phrase is semantically plural could control for the possibility that skipping rates might be affected by low-level priming of number information; for example, itmight be the case thateach other primes a plural noun simply because both elements bear a plural feature, and not because the plural noun is predicted.2 therefore, the conjoined conditions give the required degree of experimental control in order to observe a prediction-related effect on skipping rates. this effect should show up as an interaction, such that plural noun phrases are skippedmore often than singular noun phrases when preceded byeach otherthan when preceded by the conjoined phrase, but this effect should be absent, or reversed for the singular nouns. all experimental items were followed by a second short sentence to prevent wrap up effects at the end of the first sentence. the 28 sets of items were distributed among 4 lists in a latin square design. each participant was randomly assigned to a list. 28 experimental sentences were intermixed with 84 filler items. sixteen of the experimentalitems and 44 of the filler items were followed by a comprehension question. all questions were answerable as ‘yes’ or ‘no’. the correct answers were counterbalanced, with half of the correct answers ‘yes’ and half ‘no’. 2.2 results and discussion prior to analysis, trials with track loss were eliminated. fixations less than 80 msec in duration and within one character of the previous region were incorporated into the neighbouring fixation. remaining fixations that were less than 80 msec were deleted. in the following presentation of the results, we will give the results relating to word skipping first, and this will be followed by the results for standard reading-time based measures. for the purposes of the data analysis, each experimental sentence was initially divided into the following regions: start of sentence (it was to) anaphor (each other) complementizer (that) first subject (the girls) first subject modifier (from the school) main verb (said that) second subject (the boys) second subject modifier (from york) embedded vp (requested us) end of sentence (to play piano) for the data analyses reported below,probability of first-pass fixationis defined as the probability that the region received a fixation before subsequent regions were fixated. note that this measure is directly related to skipping rates, since if a region is skipped, by definition, it cannot have received a first-pass fixation; and, conversely, if a region receives one or more first-pass fixation, it cannot have been skipped. therefore, prob(fixation)=1-prob(skipping). table 1 gives the means and standard errors for the probability of initial fixation, for the first subject region. 2. however, this assumes that such priming is based onsemanticnumber, as the nouns in the conjoined condition are each morphologically singular. 44 syntactic prediction first subject (the girl/girls) probability of fixation (%) match 87 (3.00) mismatch 96 (1.21) conj-plur 94 (1.85) conj-sing 89 (1.95) table 1: probability of first-pass fixation for the first subject region (the girl(s)) estimate std. error z value p value (intercept) 3.01095 0.26115 11.529< 2e-16 *** number 0.31602 0.28721 1.100 0.271198 pp-type 0.02925 0.28317 0.103 0.917743 number×pp-type 1.99467 0.56662 3.520 0.000431 *** table 2: results of statistical analysis for probability offixation in the first subject region. key: “***”: p < .001. model estimates are given in terms of the log-odds of the probability: log odds(p)= ln( p 1−p ) the measure of probabilty of fixation involves binary response data (either a trial received a first-pass fixation or it did not). therefore, we use a logistic mixed effect regression for this measure (jaeger 2008; see also baayen et al. in press and bates and sarkar, 2007). the regression model included random intercepts for both subjects and items. thepattern of significance was not affected by the inclusion of extra parameters for random slopes, nor was model fit statistically improved by the inclusion of these parameters. we therefore report analyses for the simpler models, which include random intercepts only. prior to analysis, the predictor variables were centered, such that categorical factors were transformed into numerical codeswith a mean of zero and a range of 1. this procedure reduces collinearity between variables, and, in combination with sum coding of contrasts, also allows a straightforward interpretation of main effects and interactions in a way that is analogous to that of analysis of variance. the calculation of the significance of each effect is based on the wald z-test. the results for the probability of one or more fixations in thefirst subject region are given in table 2: the main result was a reliable interaction between pp-type and number, such that there were fewer fixations (and therefore more skipping) on the plural when it was preceded byeach otherthan when it was preceded by a conjoined phrase (87% vs. 94%),while the reverse effect was found for the singular phrase (96% vs. 89%). pairwise comparisons revealed that both of these two contrasts were reliable (p’s < .05). moreover, there were reliably fewer fixations on the plural np than the singular np when it was preceded byeach other(p’s < .05), while the fixation rates for singular and plural nps did not differ significantly when preceded byand, though the numerical pattern was in the opposite direction, with more fixations for the plural than the singular. this pattern of data is exactly as expected if skipping ratesare increased by the predictability of a plural np. in the conjunct conditions, where the is no particular reason to predict a plural np, fixation rates are statistically similar for singular and plural nps, although the tendency is for more 45 ilkin and sturt fixations on the plural, which is expected due the extra length of the plural. the fact that there were differences from baseline for both the match and mismatch reciprocal conditions suggests that there may have been both facilitatory and inhibitory processes atwork in this experiment. if this is the case, it would mean that the probability of skipping is not only increased (relative to baseline) when the number of the head noun matches the prediction, but also that it isdecreased, again relative to baseline, when it mismatches the prediction; in other words, the mismatch leads to extra fixations on the critical noun phrase. one potential objection to our interpretation of the skipping results is that the critical effect was based on only a very small number of skips over the course of the experiment. clearly, as the very high rates of fixation show, a skip of a two word region is avery rare event, and most of our participants gave only one or two skips (or none at all) out ofa maximum of seven in any given condition. thus, the effect reported above is based on differences among a relatively small number of trials. it will therefore be important to replicate the present findings with the same design, but a much larger number of items. however, it is also true that, even though very few skips were being made in our experiment, the critical interactive pattern was quite robust over our participant sample. this is shown, firstly, by the fact that the fit of the logistic regression model was not improved by adding random slope parameters to allow the size ofthe critical interaction to vary by participant (p > .9). moreover, we can gain an idea about the generality of the pattern by looking at the interaction for each individual participant. we defined an interaction score based on the proportions of fixations in each condition, which is then calculated individually for each participant: [prop(match)−prop(mismatch)]−[prop(conjplur)−prop(conjsing)]. this score should have a negative value for participants who show the predicted interactive pattern, a positive value for participants who show an interactive pattern opposite to that predicted, and be zero for those who show no interactive pattern. it turns out that 20 of our 32 participants showed the predicted pattern, with only four showing the opposite pattern (binomial sign test:p = .0008), and eight showing no interactive pattern (of whom six had ceiling levels of fixation, at 100% for all conditions). thus, although the statistical pattern of the interaction is based on only a relatively small number of trials, the pattern was very general across participants, for thosewho did make two-word skips. however, the fact that six of our participants were at ceiling levels of fixation in all four conditions suggests that there may be individual differences in the propensity to skip two words at a time. we now turn to the reading time measures. as demonstrated above, initial data analysis revealed that the first subject region (e.g.the girls) differed reliably by condition in the probability of firstpass fixation, making the interpretation of reading time measures hard to interpret. for both reading time measures (first pass and regression path), we thereforepooled the first and second subject regions with their corresponding modifiers (e.g.the girls from the school). these larger regions received equal rates of first pass fixation by condition, withall conditions receiving a first pass fixation 100% of the time in both regions. the reading time measures are defined as follows. firstpass reading time is the sum of all fixation durations from thefirst entry into the region from the left, until the first exit of the region,either to the left or to the right. regression-path time is the sum of all fixation durations from the first entry into the region from the left, until the first exit of the regionto the right. note that regression path times may contain fixations outside the region following regressions out of the region. for both reading time measures, cases where the region is skipped in first-pass reading are treated as missing data, rather than contributing a zero value to the mean. as mentioned above, the fact that probabilities of initial fixations differed among conditions motivated the use of larger analysis regions for the readingtime measures. 46 syntactic prediction we analyzed the reading time measures starting from the anaphor region until the end of the initial sentence: anaphor (each other) complementizer (that) first subject and the modifier (the girls from the school) main verb (said that) second subject and the modifier (the boys from york) embedded vp (requested us) end of sentence (to play piano) means and standard errors for both first-pass reading timesand regression path times for each region are given in table 3. linear mixed effects regression models were computed for each region and each measure. thet-statistic from these models is reported in table 2.2. the mixed effects regressions included both experimental factors and their interaction as predictors, and also included random intercepts for subjects and items. random slope parameters corresponding to main effects or the interaction were included only when justifiedby a significant improvement in model fit, based on a log-likelihoodχ2 test. experimental variables were centered before analysis. the main region we were interested in is the first subject plusmodifier region. as mentioned above, reading time data for the first subject region alone (determiner plus noun) are hard to interpret due to skipping differences3. here, if the reader is expecting to form the dependency as soon as possible, then in the reciprocal singular conditionthere should be longer fixation durations. the reader should be surprised to encounter a singular noun where a plural is expected. also again in reciprocal singular condition after a failure to form the dependency at the first possible subject position, if the parser continues to actively search for a possible subject np to complete the dependency this should be visible in the following regions.the parser should initiate the dependency formation again at other possible subject positions,which could be interpreted as potential antecedents. the results from the first 2 regions, the anaphor region (each other) and the complementizer region (each that) are not of theoretical interest in this experiment. the only significant difference between conditions was in anaphor region, where the fixation durations was significantly longer when the sentences started with two conjoined propernames, compared with when it started with ’each other’, both in first-pass reading times (means,367 vs. 312 ms.; respectively) and in regression-path times (means, 493 vs. 352 ms.; respectively). this effect is probably due to differences in length and/or semantic complexity between the two conditions, and will not be discussed any further. there were no other significant differences in either measure for either the anaphor or complementizer regions. we will now discuss each followingregion separately for each measure. the first subject and the modifier (the girls from the school): for the first-pass reading times there was no significant main effect of sentence initial pp orthe number of the first subject np. however the interaction between the pp and the number of the np was significant. there were longer fixations when the sentence included a reciprocal andthe np was singular, compared to 3. for completeness, means are reported here for the first subject np region (i.e. excluding modifier). given the skipping differences, we give means excluding skipped trials followed by means inlcuding skipped trials as zero values (in parentheses) first pass: match: 269ms (239ms); mismatch: 263ms (252ms); conj-plural: 299ms (283ms); conjsingular: 277ms (252ms). regression path: match 343ms (305ms); mismatch: 321ms (306ms); and-plural: 375ms (356ms); and-singular: 421ms (380ms). 47 ilkin and sturt measure anaphor comp. np1+mod. main verb np2+mod. emb. vp end of sent. fpt match 312 (15) 225 (12) 599 (34) 361 (21) 697 (37) 338 (24) 738 (36) mismatch 311 (17) 242 (17) 632 (30) 349 (16) 695 (40) 366 (23) 717 (39) conj-plur 374 (22) 261 (16) 608 (28) 339 (17) 668 (36) 347 (21)724 (37) conj-sing 360 (17) 240 (15) 573 (37) 356 (19) 689(42) 334 (23)716 (35) rpt match 385 (24) 261 (17) 741 (41) 424 (25) 1059 (62) 417 (44) 1054 (101) mismatch 402 (25) 303 (30) 775 (47) 451 (30) 1199 (80) 489 (42)883 (71) conj-plur 615 (37) 287 (17) 772 (41) 443 (45) 1074 (62) 438 (36) 875 (69) conj-sing 601 (40) 293 (31) 805 (50) 339 (33) 1031 (54) 454 (44) 849 (60) table 3: means and standard errors for the first-pass time and regression-path time measures for each region anaphor comp. np1 main np2 emb. end measure +mod. verb +mod. vp of sent. fpt pp type t = −3.07∗ < 1 < 1 < 1 < 1 < 1 < 1 number t = < 1 < 1 < 1 < 1 < 1 < 1 < 1 pp× number t = < 1 < 1 2.01∗ < 1 < 1 1.36ns < 1 rpt pp type t = −5.88∗ < 1 −1.24ns < 1 1.58ns < 1 1.54ns number t = < 1 < 1 −1.32ns < 1 1.07ns 1.27ns −1.29ns pp× number t = < 1 < 1 < 1 < 1 2.05∗ < 1 −1.42ns table 4: t-values based on linear mixed effect model of reading time data. key: “ns”: nonsignificant;“∗”: p < .05. the expression “< 1” is a shorthand for|t| < 1. significance at α = .05 is based on the criterion of|t| > 2. 48 syntactic prediction when it included conjunction and the first np was singular (means, 632 vs, 573 ms.; respectively: p < .05). however, there were no reliable differences for the plural np’s (p > .1). in regression path times, there was no reliable effect of sentence initial pp type, number of the first subject np or the interaction between them. the results for first pass time are consistent with the skipping rates for the first subject, and with the interpretation that the readers were expecting to have aplural noun at the first subject region only when sentences started with a reciprocal. in such cases, the parser may have been actively constructing the dependency as soon as the reciprocal was encountered, with the expectation to complete the dependency at the first subject position. for the main verb region (said that) there were no reliable effects of either of the two experimental factors, or of the interaction between them, eitherfor the first pass time or for regression path time. it seems that the processing difficulty for the reciprocal singular condition (as observed in the first pass times) does not spill over to this region. for the second subject np and the modifier (the boys from york), in first-pass time, there were no reliable effects. however in regression-path times, theinteraction between the two factors was significant. the pairwise comparisons revealed that thereciprocal-singular condition was read more slowly than the conjunct singular (mismatch) condition (reciprocal-singular, 1199 vs. conjunct-singular, 1031),p < .05), but there was no such corresponding effect for the plural conditions (p > .1). for the reciprocal singular condition, this region was the first region that would allow the parser to complete the unbounded dependency afteran unsuccessful attempt at the first subject np region. the longer fixation durations for the regression path times could indicate that the parser is reattempting to form the dependency at this region. considering that the regression path time measure includes fixations that are made to the leftof the region, the longer fixation durations could indicate the reactivation of the sentence initial pp and it’s attempted integration with the second subject np. there were no remaining significant effects in the reading time measures. 3. general discussion the experiment reported above showed that the parser actively projects the syntactic position and grammatical features associated with a potential antecedent of the reciprocal before the antecedent has been reached in the input. evidence for this comes not only from fixation times, but also from skipping rates. the experiment showed that readers had a tendency to skip over a two word noun phrase when the morphology on the head noun matched the predicted number of the phrase. these results have implications for theories of human parsing as well as theories of eye-movement control. we will consider each of these sets of implications in turn below. 3.1 implications for theories of human parsing the results reported here are consistent with a model in which the antecedent of a cataphoric dependency can be predicted, and assigned syntactic features, ahead of time, as proposed by kazanina et al. (2007) and kreiner et al. (2008). however, unlike those earlier studies, the present experiment established the relevant congruency effect in skipping rates, at a positionbeforethe critical word had been fixated in the input, thus lending further evidence for prediction. in the following paragraphs, we consider the parsing processes that might underlie the observed effects on skipping rates and 49 ilkin and sturt fixation durations that we observed. these results have implications for parsing theories because they clearly show that early sentence initial information is used for projecting upcoming structure. in the introduction to the experimental section above, we assumed that the prediction of the plural np leads to the computation of plural features in advance. this then leads to a facilitation of the parafoveal processing of the plural np, resulting in theobserved skipping rate differences. if this is so, one important question is what type of processing actually takes place in the parafovea. one possibility is that lexical access occurs for both the determiner theand the head noungirl(s) while the reader is fixating to the left of this region. both words ofthis two-word phrase are accessed, and integrated with the predicted syntactic information. a second possibility appeals to the notion of shallow processing. according to this account, the phrase is only partially processed parafoveally. we assume that the determinerthe is processed fully, as it is a short, high frequency functionword. however, one possibility is that the head noungirl(s) is processed via a heuristic which involves checking the “s” plural marker, without necessarily going through a full lexical access process. according to this account, the parser predicts the plurality of the subject noun phrase, and also guesses the internal structure, assuming a highly probabledeterminer-noun combination. the parser predicts that the head noun should be marked for plurality, which in english usually means that an “s” morpheme is expected at the end of the word.4 according to this account, the skip-trials include cases where the plural prediction is satisfied through the parafoveal processing of the “s” morpheme at the end of the head noun, without the parser accessing the content of the noun. thus, the reader will end up with only a partial interpretation of the phrase,because the content of the head noun will not have been retrieved, but will perceive the sentenceas grammatical, because the morphology matches the requirements for the number feature. there are previous experiments that show that lexical decisions are facilitated if the syntactic category of the word is predictable from previous sentence context (e.g. wright and garrett (1984)). this result is consistent with a syntactic early predictive mechanism that is operative before full lexicalaccess. based on the current data, it is not straightforward to distinguish between these two possibilities, but both are consistent with a predictive processing mechanism. in the second account, if lexical access has not yet occurred when the skip takes place, then this is by definition a case where behavior associated with prior prediction occurs before full bottom-up input has taken place, similarly to the case of delong et al. (2005) and van berkum et al. (2005). in the first account, where lexical access is assumed to take place even in the presence of a skip, one canargue for prior prediction on the basis of the extreme earliness of the effect, following a similar argument to that of lau et al. (2006). after all, this is a case where information on the critical word is being processed from a point that is at least two words to the right of a fixation. it is hard to imagine a congruence effect occuring so early without prior prediction. of course the results we obtained also could be explained by non-predictive mechanisms, like the integration of the first subject np with the sentence initial unbounded dependency. however the fact that we observed very early facilitation of processing of plural nouns in a predictive context makes this possibility unlikely. the results of the experiment imply that the parser processes at least the morphology of the parafoveal word if it is indicated by previous context that there should be a plural noun. 4. of course, although this is a highly likely possibility, it is not the only one. for example, the head noun might not be morphologically marked for plurality (e.g.the couple, the boy and the girl, etc, or the head noun might not occur adjacent to the determiner (e.g.the shop assistants). 50 syntactic prediction the results suggest that the parser not only expects a pluralsubject np, but also that it projects the place where this antecedent ofeach otheris anticipated (at the first possible subject np position). in this respect these results are consistent with other experiments (staub and clifton, 2006), which suggest that the parser predicts the arrival of a particularstructure if it is indicated by previous sentence context. for example staub and clifton (2006) showed that presence ofeither facilitated the fixation durations on phrases that start withor. similarly in our experiment the parser using information from the dative reciprocal, actively predicted the site and the morphology of the upcoming subject np. in staub and clifton’s studyeither induces a very high phrasal expectancy for the wordor. in our case the reciprocal restricts the possibilities but does not direct the reader to a specific word. however along with the possible site of the antecedent noun phrase readers also predicted specific morphological information. as we mentioned in the introduction, the sentence initial reciprocal is unbounded and it does not have to have its subject at the first position. however without waiting for bottom-up information confirming the place of the dependency site, the parser triesto construct the dependency at the first possible site, thus indicating an active, or eager search strategy (kazanina et al., 2007). moreover, results from skipping proportions and fixation times on firstsubject position indicate that the reader forms the dependency between the unbounded np and its subject before encountering any verb information. preverbal processing of dependencies has been shown in head final languages (aoshima et al., 2004; kamide et al., 2003). however the current experiment shows early processing of dependencies without verb information in english where generally the verb comes early in the sentence. this generalizes the previous evidence against head-driven parsing. the early prediction of morphological information is consistent with recent results from meg studies reported by dikker et al. (2009) and dikker et al. (2010). in their meg study, dikker et al. (2009) showed a very early response for the nouns (around 130 ms) if there is a mismatch between the expected syntacticcategory and the category marking closed class morphemes. dikker et al. (2010) showed that theparser not just reacts to closed class morphemes, but also to the form typicality of the words. in this study they compared nouns, with typical closed class category marking morphemes (like farmer) with nouns with typical properties of nouns, (like movie) and with neutral nouns, whose form does not have a clear bias towards either being a noun or a verb. they found a very early sensitivity to typical nouns and nouns with closed class morphemes, around 100 ms. these very early effects are hard to explain by full lexical processing of the words. the authors claim that their results not only indicate that the system predicts words categories, but that it also predicts the form features that are associated with that category. the skipping patterns obtained in our experimentare also in line with the results obtained by dikker et al. as we mentioned earlier, it might be easier for the system to pick up the cues for plurality because sentence initial reciprocal highly constrains the features of the subject noun phrase. the first pass times on the first subject region, showed that readers slowed down in the mismatch condition, indicating that the parser attempts to form the dependency between the sentence initial reciprocal and its possible antecedent at the first possibleposition, and this effect was replicated in the skipping probabilities. at the second subject position, there was a similar interaction effect in regression path time. this could indicate that readers attempt once more to form the dependency, this time with the second subject np, in conditions where theinitial attempt with the first subject has failed, due to a mismatch of features. if so, the increased reading times in the mismatch condition 51 ilkin and sturt could indicate the cost of retrieving sentence initial reciprocal and and integrate it with possible antecedent. 3.2 implications for models of eye-movement control there are three aspects of our results that have consequences for models of eye-movement control in reading. first, the results show that morphological information from the head noun must have been processed parafoveally, and second, this parafoveal information was accessed from a point at least two words to the right of the current fixation; in other words from word n+2. finally, the results imply that syntactic prediction can influence saccade target location. all of these points have implications for models of eye-movement control in reading, and we will discuss them separately. in discussing models of eye-movement control, we will concentrate on two models, one of which, swift, allows parallel processing of multiple words aroundthe fixation (kliegl et al., 2006), and the other, e-z reader allows only serial processing of words(reichle et al., 2003). of course, these are not the only two models of eye-movement control in the literature. however, these models are useful because they make clear and testable predictions, and because they represent very different approaches to the phenomena in question. current models of eye-movement control do not have specific mechanisms that allow for morphological pre-processing in the parafovea, and indeed, experimental work in eye-movement control has in general failed to show evidence that morphological information is accessed from a parafoveal word (inhoff, 1989; kambe, 2004; lima, 1987). these earlierstudies used display change techniques in which a word changes dynamically, as the reader’s eye crosses an imaginary boundary in the text. the studies manipulated whether a morpheme from word n+1 was available for preview in parafoveal vision while the reader was fixating word n. thestudies failed to find a facilitatory effect of morphological preview for fixations on word n+1. in other words, word n+1 was not read any faster when one of its morphemes had been previouslyavailable for preview, relative to when a non-morpheme string had been available for preview (see also hyönä et al. (2004), bertram and hyönä (2007) and hyönä and pollatsek (1998) for studies failing to find parafoveal effects of morphology in reading finnish compound nouns). however,none of these studies considered cases where this morphological information was highly predictable from the preceding syntactic context. in contrast, the experiment reported in the present paper not only manipulated the predictability of the morpheme, but also included conditions where the morphology both matched and mismatched the prediction. it might therefore be the case that morphological information is processed parafoveally only when this information is highly predictable from the context, as it was in our experiment reported here, and both match and mismatch conditions might be required in order to observe the effect. evidence consistent with our findings comes from a study in hebrew reported by deutsch et al. (2005), who showed that morphological information can extracted parafovealy in hebrew in cases where the parafoveal word has a common verbal pattern with the target word. deutsch et al. (2005) also showed that if the parafoveal preview word was syntactically incongruent with the target word, then an inhibition effect in processing was observed. interestingly, semantically biasing the context did not lead to parafoveal processing of the root morphemes.these results are important because they show processing of morphological information in the parafovea and these effects arise from syntactic context, but not when semantic information was manipulated. in a similar way in our experiment the sentence initial reciprocal constrains theupcoming information in several respects. 52 syntactic prediction after reading the reciprocal, the parser expects to find the antecedent of the initial reciprocal phrase at the first possible position, which is immediately after the complimentizerthat. in addition, the parser expects to have a plural and an animate (probably a np related to a person) phrase. all of these constraints are possibly projected after reading thereciprocal, leading to facilitation of early processing of the projected noun, before fixating it. an eye-tracking study by underwood et al. (1990) showed thatreaders tend to look at the informative part of the word. if high information content isat the beginning of the word, as in the case ofengagement, readers tend to look at the beginning of the word, compared to words like underneath. they claim that this effect could only be explained if the morphological information is processed parafoveally. there is some evidence for processing of semantic (yan et al., 2009) and to some extent morphological (yen et al., 2008) information parafoveally in chinese. for example yen et al. (2008) showed either real words or pseudo words in the parafovea, and all words had the same initial morpheme with the target word. they showed that there were shorter fixations for the real words that have the same initial morpheme with the target word comparedto pseudo words. the second aspect of our results that has implications for models of eye-movement control is the fact that the relevant parafoveal information was extracted from word n+2. at first sight, this appears to be more compatible with models in which multiple words around the fixation can be processed in parallel, such as the swift model (kliegl et al., 2006), rather than models where word identification occurs serially, as in the e-z reader model (reichle et al., 2003). however, in the e-z reader model, although words are processed serially, attention can move to word n+1 while word n is being fixated. if word n+1 is identified early enough, then the planned saccade to n+1 can be cancelled, and attention re-allocated to n+2. thus, it is still technically possible in the e-z reader model for processing of word n+2 to begin while word n is being fixated. in practice, within the e-z reader model, this happens only under very specific circumstances, for example, when word n+1 is a very short or high frequency word. in our case, n+1 is the very high frequency wordthe, so it is possible that some aspects of the head noun could havebeen processed during the fixations before the article, even in the e-z reader model. in fact, however, evidence for parafoveal processing of word n+2 is very limited in the eyemovement literature. using a display change technique, angele et al. (2008) and rayner et al. (2007) failed to find any evidence for preview benefit for n+2 (i.e. parafoveal viewing of n+2 did not lead to faster fixations when n+2 was subsequently fixated). however as rayner et al. (2007) mention, in both of these experiments both n+1 and n+2 were relatively long words, apart from the second experiment of rayner et al. (2007), where n+1 and n+2 were 3 or 4 letters long. in contrast, kliegl et al. (2007) again in a display change experiment, found evidence for parafoveal processing of n+2, when the intervening word was 3 letters long. however the effect was observed in fixation durations on words n+1 and n. also risse et al. (2008) found a frequency effect on single fixation durations only when n+1 was 3 or 2 letters longand n+2 was not longer than 4 letters. in our experiment, where the nouns were preceded bythree letter article,the, there may well have been some parafoveal processing of the noun following the article. such processing may be facilitated if the article and noun are treated as a “word group” (radach, 1996) for the purpose of eye-movement control. radach (1996) showed that initial fixations on words that are preceded by a three letter word showed a single normal distribution overthe whole combined two word string, consistent with the grouping of the two words together into asingle visual object. drieghe et al. (2008) found that these combined normal distributions are limited to situations where the three-letter 53 ilkin and sturt word preceding the noun is an article, although they argue that the apparent normal distribution may be better described as bimodal. however, in our stimuli, thecritical noun was always preceded by the determinerthe, so, if theword groupinghypothesis is correct, the two words were treated as a single unit, increasing the effectiveness of parafoveal preview of n+2. indeed, even if readers were not grouping the article and the noun together, it is still possible that having a highly constraning syntactic context in addition to obvious parafoveal cues might have facilitated processing of the noun. the final aspect of our results that has consequences for models of eye-movement control is the fact that skipping rates were affected by syntactic predictability. there is evidence in the literature that words that are predictable from context are skipped more often than those that are not (ehrlich and rayner, 1981; rayner and well, 1996; white et al., 2005).this fact can be accounted for in swift (kliegl et al., 2006), but also in e-z reader (reichle et al., 2003), as mentioned in the introduction to the present paper. in swift lexical activation for words builds up gradually, and starts to decline after it reaches the maximum. saccades aredirected towards the word word with highest activation. if lexical processing is facilitated on word n+1 then the exitibility starts to decline before the saccade program can be finalized. if the activation level for n+2 exceeds n+1, then the saccade is directed to n+2, resulting in a skip. in swift the word can be skipped even if it is not fully lexically processed. however this explanation still does not account for the skipping rates in the present experiment. in our study, we manipulated not the predictability of a specific word, but the predictability of an abstract feature of the word, namely its plurality. neither e-z reader nor swift include mechanisms that allow for this moreabstract type of predictability to influence skipping rates. 4. conclusion overall, the results of the experiment showed a strong case for predictive mechanisms in the formation of long distance dependencies, using eye movement measures. also the results indicate a strong facilitation of parafoveal processing of words in a syntactically biased context. clearly, further experiments are need to follow up these findings, and to examine in more detail what types of information are extracted from the parafovea and under what types of conditions. however the experiment presented here addresses important and novel issues about the nature of the mechanisms and the time course of the effects of syntactic prediction. references g. altmann and y. kamide. incremental interpretation at verbs: restricting the domain of subsequent reference.cognition, 73:247–264, 1999. b. angele, t.j. slattery, j. yang, r. kliegl, and k. rayner. parafoveal processing in reading: manipulating n+ 1 and n+ 2 previews simultaneously.visual cognition, 16(6):697–707, 2008. s. aoshima, c. phillips, and a. weinberg. processing filler-gap dependencies in a head-final language.journal of memory and language, 51:23–54, 2004. r. h. baayen, d. j. davidson, and d. m. bates. mixed effects modeling with crossed random effects for subjects and items.journal of memory and language, in press. 54 syntactic prediction d. a. balota, a. pollatsek, and k. rayner. the interaction ofcontextual constraints and parafoveal visual information in reading.cognitive psychology, 17:364–390, 1985. d. bates and d. sarkar.lme4: linear mixed-effects models using s4 classes, 2007. r package version 0.9975-11. r. bertram and j. hyönä. the interplay between parafovealpreview and morphological processing in reading.eye movements: a window on mind and brain, pages 391–407, 2007. m. brysbaert, d. drieghe, and f. vitu. word skipping: implications for theories of eye movement control in reading. in g. underwood, editor,cognitive processes in eye guidance, pages 53–77. oxford university press, oxford, 2005. n. chomsky.lectures on government and binding. foris, 1981. k. delong, t. urbach, and m. kutas. probabilistic word pre-activation during language comprehension inferred from electrical brain activity.nature neuroscience, 8(8):1117–1121, 2005. a. deutsch, r. frost, a. pollatsek, and k. rayner. morphological parafoveal preview benefit effects in reading: evidence from hebrew.current issues in morphological processing, page 341, 2005. s. dikker, h. rabagliati, and l. pylkkänen. sensitivity tosyntax in visual cortex.cognition, 110 (3):293–321, 2009. s. dikker, h. rabagliati, t.a. farmer, and l. pylkkänen. early occipital sensitivity to syntactic category is based on form typicality.psychological science, 21(5):629, 2010. d. drieghe, a. pollatsek, a. staub, and k. rayner. the word grouping hypothesis and eye movements during reading.journal of experimental psychology. learning, memory, andcognition, 34 (6):1552, 2008. s.f. ehrlich and k. rayner. contextual effects on word perception and eye movements during reading.journal of verbal learning and verbal behavior, 20(6):641–655, 1981. j. hyönä and a. pollatsek. reading finnish compound words:eye fixations are affected by component m orphemes.journal of experimental psychology: human perception and performance, 24(6):1612–1627, 1998. j. hyönä, r. bertram, and a. pollatsek. are long compound words identified serially via their constituents? evidence from an eye-movement-contingent display change study.memory and cognition, 32(4):523, 2004. a.w. inhoff. parafoveal processing of words and saccade computation during eye fixations in reading. journal of experimental psychology: human perception and performance, 15(3):544–555, 1989. t. f. jaeger. categorical data analysis: away from anovas (transformation or not) and towards logit mixed models.journal of memory and language, 59:434–446, 2008. 55 ilkin and sturt g. kambe. parafoveal processing of prefixed words during eyefixations in reading: evidence against morphological influences on parafoveal preprocessing. perception and psychophysics, 66(2):279–292, 2004. y. kamide, g. t. m. altmann, and s. l. haywood. the time-course of prediction in incremental sentence processing: evidence from anticipatory eye movements. journal of memory and language, 49:133–156, 2003. n. kazanina, e. lau, m. lieberman, m. yoshida, and c. phillips. the effect of syntactic constraints on the processing of backwards anaphora.journal of memory and language, 56:384–409, 2007. r. kliegl, a. nuthmann, and r. engbert. tracking the mind during reading: the influence of past, present, and future words on fixation durations.journal of experimental psychology: general, 135(1):13–35, 2006. r. kliegl, s. risse, and j. laubrock. preview benefit and parafoveal-on-foveal effects from word n 2. journal of experimental psychology, 33(5):1250–1255, 2007. h. kreiner, p. sturt, and s. garrod. processing definitionaland stereotypical gender in reference resolution: evidence from eye-movements.journal of memory and language, 58:239–261, 2008. e. lau, c. stroud, s. plesch, and c. phillips. the role of prediction in rapid syntactic analysis. brain and language, 98:74–88, 2006. m-w lee. another look at the role of empty categories in sentence processing (and grammar). journal of psycholinguistic research, 33:51–73, 2004. r. levy. expectation-based syntactic comprehension.cognition, 106(3):1126–1177, 2008. s. d. lima. morphological analysis in sentence reading.journal of memory and language, 26(1): 84–99, 1987. k. mcrae, m. j. spivey-knowlton, and m. k. tanenhaus. modeling the influence of thematic fit (and other constraints) in on-line sentence comprehension. journal of memory and language, 38:283–312, 1998. i. mulders. transparent parsing: head-driven processing of verb-finalstructures. lot, utrecht, 2002. m. pickering and s. garrod. do people use language production to make predictions during comprehension?trends in cognitive sciences, 11:105–110, 2007. b. l. pritchett. head position and parsing ambiguity.journal of psycholinguistic research, 20(3): 251–270, 1991. r. radach.blickbewegungen beim lesen: psychologische aspekte der determination von fixationspositionen. waxmann verlag, 1996. k. rayner and a.d. well. effects of contextual constraint oneye movements in reading: a further examination.psychonomic bulletin & review, 1996. 56 syntactic prediction k. rayner, j. ashby, a. pollatsek, and e. d. reichle. the effects of frequency and predictability on eye fixations in reading: implications for the e-z reader model. journal of experimental psychology: human perception and performance, 30(4):720–732, 2004. k. rayner, b.j. juhasz, and s.j. brown. do readers obtain preview benefit from word n+ 2? a test of serial attention shift versus distributed lexical processing models of eye movement control in reading. journal of experimental psychology: human perception and performance, 33(1): 230–245, 2007. e. reichle, k. rayner, and a. pollatsek. the e-z reader modelof eye-movement control in reading: comparisons to other models.behavioral and brain sciences, 26:446–526, 2003. s. risse, r. engbert, and r. kliegl. eye-movement control inreading: experimental and corpusanalysis challenges for a computational model.cognitive and cultural influences on eye movements, pages 65–91, 2008. p. j. schwanenflugel and e. j. shoben. the influence of sentence constraint on the scope of facilitation for upcoming words.journal of memory and language, 24:232–252, 1985. s. sereno and k. rayner. measuring word recognition in reading: eye-movements and event-related potentials.trends in cognitive science, 7:489–493, 2003. a. staub and c. clifton. syntactic prediction in language comprehension: evidence from either ... or. journal of experimental psychology: learning, memory and cognition, 32:425–436, 2006. m. j. traxler and d. j. foss. effects of sentence constraint on priming in natural language comprehension.journal of experimental psychology: learning, memory, andcognition, 26:1266– 1282, 2000. g. underwood, s. clews, and j. everatt. how do readers know where to look next local information distributions influence eye fixations.quarterly journal of experimental psychology, 42(1): 39–65, 1990. j. j. a. van berkum, c. m. brown, p. zwisterlood, v. koojman, and p. hagoort. anticipating upcoming words in discource: evidence from erps and readingtimes. journal of experimental psychology-learning memory and cognition, 31(3):443–467, 2005. r. p. g. van gompel and s. p. liversedge. the influence of morphological information on cataphoric pronoun assignment.journal of experimental psychology-learning memory and cognition, 29:128–139, 2003. s.j. white, k. rayner, and s.p. liversedge. the influence of parafoveal word length and contextual constraint on fixation durations and word skipping in reading. psychonomic bulletin & review, 12(3):466, 2005. b. wright and m. garrett. lexical decision in sentences: effects of syntactic structure.memory & cognition, 12(1):31–45, 1984. m. yan, e. m. richter, h. shu, and r. kliegl. readers of chinese extract semantic information from parafoveal words.psychonomic bulletin & review, 16(3):561–566, 2009. 57 ilkin and sturt m. yen, j. tsai, o.j.l. tzeng, and d. l. hung. eye movements and parafoveal word processing in reading chinese.memory & cognition, 36(5):1033–1045, 2008. 58 dialogue & discourse 8(1) (2017) 31–65 doi: 10.5087/dad.2017.102 training end-to-end dialogue systems with the ubuntu dialogue corpus ryan lowe ryan.lowe@cs.mcgill.ca school of computer science, mcgill university nissan pow nissan.pow@mail.mcgill.com school of computer science, mcgill university iulian vlad serban julianserban@gmail.com diro, université de montréal laurent charlin lcharlin@gmail.com school of computer science, mcgill university chia-wei liu chia-wei.liu@mail.mcgill.ca school of computer science, mcgill university joelle pineau jpineau@cs.mcgill.ca school of computer science, mcgill university editor: amanda stent submitted 04/2016; accepted 12/2016; published online 01/2017 abstract in this paper, we analyze neural network-based dialogue systems trained in an end-to-end manner using an updated version of the recent ubuntu dialogue corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words1. this dataset is interesting because of its size, long context lengths, and technical nature; thus, it can be used to train large models directly from data with minimal feature engineering. we provide baselines in two different environments: one where models are trained to select the correct next response from a list of candidate responses, and one where models are trained to maximize the loglikelihood of a generated utterance conditioned on the context of the conversation. these are both evaluated on a recall task that we call next utterance classification (nuc), and using vector-based metrics that capture the topicality of the responses. we observe that current end-to-end models are 1. this work is an extension of a paper appearing in sigdial (lowe et al., 2015). this paper further includes results on generative dialogue models, more extensive evaluation of the retrieval models using vector-based generative metrics, and a qualitative examination of responses from the generative models and classification errors made by the dual encoder model. experiments are performed on a new version of the corpus, the ubuntu dialogue corpus v2, which is publicly available: https://github.com/rkadlec/ubuntu-ranking-dataset-creator. the early dataset has been updated to add features and fix bugs, which are detailed in section 3. c©2017 ryan lowe, nissan pow, iulian vlad serban, laurent charlinn, chia-wei liu and joelle pineau this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). lowe, pow, serban, charlinn, liu and pineau unable to completely solve these tasks; thus, we provide a qualitative error analysis to determine the primary causes of error for end-to-end models evaluated on nuc, and examine sample utterances from the generative models. as a result of this analysis, we suggest some promising directions for future research on the ubuntu dialogue corpus, which can also be applied to end-to-end dialogue systems in general. 1. introduction deriving statistical models that can naturally and coherently converse with humans is one of the cornerstone problems of artificial intelligence. until recently, such models required significant handengineering of features, and thus could only generate a limited number of responses and be deployed in constrained situations. recent advances in neural network-based language models have begun to make feasible the idea of learning an entire dialogue model directly from conversational data, with humans only specifying the model hyper-parameters. however, significant work needs to be done before these models can be implemented in practice with high confidence. in this paper we consider the problem of building dialogue agents in an end-to-end manner. we define end-to-end systems, contrary to modular systems, as those that are trained directly from conversational data to optimize a single objective function (see section 1.1). we use the recently released ubuntu dialogue corpus, which consists of almost one million two-person (dyadic) conversations extracted from the ubuntu chat logs, which provide technical support for various ubunturelated problems. dialogues in the corpus are multi-turn and unstructured, as there is no a priori logical representation for the information exchanged during the conversation. this is in contrast to recent systems which focus on structured dialogue tasks, using slot-filling representations (williams et al., 2013; henderson et al., 2014a; singh et al., 2002). the creation of such a large, unstructured dialogue dataset was motivated by observations of progress in various sub-fields of ai. in particular, it has been argued that this progress can be attributed to three major factors: 1) the public distribution of very large rich datasets (deng et al., 2009), 2) the availability of substantial computing power, and 3) the development of new training methods for neural architectures, in particular leveraging unlabeled data. we conduct an analysis of several dialogue models that can be used in conjunction with the ubuntu dialogue corpus. we first consider classification models, which are trained to select the correct next response of a conversation from a list of candidate responses. we use a baseline model that calculates term frequency-inverse document frequency (tf-idf) between the context and each response, and compare it to a dual encoder (de) model using both recurrent neural networks (rnns) and long short-term memory (lstms). next, we present encoder-decoder models that are trained to generate an utterance given the context. we consider both the traditional lstm language model, which corresponds to the encoder and decoder having tied weights, and the recently proposed hierarchical recurrent encoder-decoder (hred) (serban et al., 2016), which has a second recurrent network that encodes utterance-specific information, and is thus able to model longer-term dependencies in the context. we evaluate these models on the task of next utterance classification (nuc), where the model ranks a list of candidate responses by how likely they are to have followed the context. we also evaluate using vector-based metrics to determine the quality of generated responses, in terms of semantic similarity to the ground-truth next utterance. we observe that the state-of-the-art models, the lstm dual encoder and hred, outperform the baselines on all metrics. 32 training end-to-end dialogue systems finally, we conduct a qualitative analysis to determine the main sources of error for the de model. we find that the most common errors are a lack of understanding of the semantics of the responses, which includes missing key words that are copied between the context and target response, and a lack of higher-level inference. there are also a number of cases where the model would benefit from explicitly incorporating the turn-taking structure of dialogue, and using some source of external knowledge for ubuntu terminology. an examination of the responses produced by the generative models reveals similar shortcomings; while the models are able to generate reasonable responses, they are often generic or lack a semantic understanding of the context. it is clear that end-to-end systems are not close to solving a domain as complex as ubuntu. we hope that this analysis can help guide future research on the ubuntu dialogue corpus, and the development of end-to-end dialogue systems. 1.1 motivation for end-to-end dialogue systems it is important to define specifically what is meant by an ‘end-to-end’ dialogue system. we begin with the standard architecture for a dialogue system, which incorporates a speech recognizer, language interpreter, state tracker, response generator, natural language generator, and speech synthesizer. in the case of text-based (written) dialogues, the speech recognizer and speech synthesizer can be left out. although some previous literature on dialogue systems identifies only the state tracker and response selection components as belonging inside the dialogue manager (young, 2000), we adopt a broader view where the language interpreter and natural language generator are also part of the dialogue manager. figure 1: an end-to-end dialogue system replaces the traditional components of a dialogue system with a single statistical model. when we speak of an ‘end-to-end’ dialogue system, we mean a single system that can be used to solve each of these four aspects simultaneously (see figure 1). typically this is a system that takes as input the history of the conversation and is trained to optimize a single objective, which is a function of the textual output produced by the system and the correct (ground truth) response. this is in contrast to the ‘modular’ system approach to dialogue systems, where each component of figure 1 is trained separately, and either takes a more structured input, such as a set of dialogue acts, or is trained to maximize an intermediary objective, such as slot-filling. more formally, we define 33 lowe, pow, serban, charlinn, liu and pineau a modular dialogue system as a system where two or more elements (sub-components or system parameters) are optimized with respect to two or more different objective functions (e.g. where the state tracker is trained to minimize the cross-entropy error of predicting the slot-value pairs, and where the response generator is trained to maximize the conditional log-likelihood of the correct response given the slot-value pairs). thus, any machine learning-based dialogue system which is not a modular dialogue system is an end-to-end dialogue system. note that, according to this definition, whether a system is end-to-end is independent of how it is evaluated. both retrieval and generative models can be end-to-end so long as they are trained using a single objective function. similarly, end-to-end models can be evaluated using intermediary tasks such as nuc, which do not evaluate the ability of the models to generate new utterances unseen in the training set. of course, in order to evaluate the full capability of the models it is best to evaluate their outputs in a setting as realistic as possible; however, this this difficult to do automatically when there is no notion of task completion (liu et al., 2016). examples of end-to-end dialogue systems in the recent literature involve neural network-based approaches that are fully differentiable, and are usually trained to maximize the log-likelihood of the generated utterance conditioned on some conversational context (serban et al., 2016; vinyals and le, 2015). these systems learn off-line through examples of human-human dialogues, and thus learn to emulate the behaviour of agents in the training corpus. however, differentiability and offline learning are not strict prerequisites for end-to-end dialogue systems, and other methods could be devised. modular dialogue systems have been historically preferred over end-to-end systems. this is because such modular systems are easier to train, require less data, and so far have been shown to achieve better results in practice, albeit typically for highly structured tasks. it is also easier to manually program each component specifically to obey certain task constraints or to solve one or more isolated tasks. however, there are significant advantages to end-to-end dialogue systems that make investigating them worthwhile. in particular, modular dialogue systems are restricted to taskspecific domains, and often require significant human feature engineering, including pre-defining the state and action spaces of the model. although this can work well for narrow domains, it does not necessarily generalize to general-purpose dialogue. on the other hand, end-to-end models do not require a pre-defined state or action space representation; instead, these representations are learned directly from conversational data. once an end-to-end model architecture is specified, all that is needed to have the system learn to converse about another domain is to provide new training data for that domain. as the amount of available dialogue data grows and more general-purpose conversational systems are desired, we believe that training end-to-end models without hand-crafted features will yield better performance. however, despite these advantages it is not clear whether end-to-end approaches will work for any dialogue domain. for example, in complex negotiation domains where thorough analysis is required and little data is available, end-to-end systems may not be able to learn successful strategies for every combination of preferences and goals. thus, it is crucial to conduct further research on end-to-end dialogue systems in order to determine the domains in which they are most effective. 1.2 paper outline the paper is structured as follows. in section 2, we detail some relevant dialogue datasets that are available, and give an overview of existing end-to-end dialogue systems. in section 3, we describe 34 training end-to-end dialogue systems the ubuntu dialogue corpus, including corpus statistics, how it was collected, and any changes made in the updated version. we then go on to describe the response ranking models on the ubuntu dialogue corpus in section 4, and the response generation models in section 5. we give our results in each of these sections, including our models’ performance on next utterance classification and embedding similarity. we also provide a qualitative error analysis of the mistakes made by the classification model in section 4.6. finally, we conclude in section 6, and discuss potential extensions and limitations of the ubuntu dialogue corpus, evaluation metrics, and future directions for end-to-end dialogue systems. 2. related work we briefly review existing dialogue datasets, and some of the more recent learning architectures used for both structured and unstructured dialogues. this is by no means an exhaustive list, but surveys resources most related to our contributions. 2.1 dialogue datasets the switchboard dataset (godfrey et al., 1992), and the dialogue state tracking challenge (dstc) datasets (williams et al., 2013, 2016) have been used to train and validate dialogue management systems for interactive information retrieval. in the case of the dstc, the problem is typically formalized as a slot filling task, where agents attempt to predict the goal of a user during the conversation. these datasets have been significant resources for structured dialogues, and have allowed major progress in this field, though they are relatively small compared to datasets currently used for training neural architectures for language-related tasks. recently, a few datasets have been used containing unstructured dialogues extracted from twitter2. ritter et al. (2010) collected 1.3 million conversations; this was extended in sordoni et al. (2015b) to take advantage of longer contexts by using a-b-a triples. shang et al. (2015) used data from a similar chinese website called weibo3. however to our knowledge, these datasets have not been made public, and furthermore, the post-reply format of such microblogging services is perhaps not as representative of natural dialogue between humans as the continuous stream of messages in a chat room. in fact, ritter et al. (2010) estimate that only 37% of posts on twitter are ‘conversational in nature’, and 69% of their collected data contains exchanges of only length 2 (ritter et al., 2010). we hypothesize that the interaction patterns of chat-room style messaging are more closely correlated to human-human dialogue than micro-blogging websites, or forum-based sites such as reddit. part of the ubuntu chat logs have previously been aggregated into a dataset, called the ubuntu chat corpus (uthus and aha, 2013b). however that resource preserves the multi-participant structure and thus is less amenable to the investigation of more traditional two-party conversations. also weakly related to our contribution is the problem of question-answer systems. several datasets of question-answer pairs are available (boyd-graber et al., 2012), however these interactions are much shorter than what we seek to study. for a comprehensive survey of both available dialogue datasets and prevalent models, see serban et al. (2015). 2. https://twitter.com/ 3. http://www.weibo.com/ 35 lowe, pow, serban, charlinn, liu and pineau 2.2 learning architectures for end-to-end dialogue systems most dialogue research has historically focused on structured slot-filling tasks (schatzmann et al., 2005). various approaches were proposed, yet few attempts leverage more recent developments in neural learning architectures. a notable exception is the work of henderson et al. (2014b), which proposes an rnn structure, initialized with a denoising autoencoder, to tackle the dstc 3 domain. work on end-to-end dialogue systems was recently pioneered by ritter et al. (2011), who proposed a response generation model for twitter data based on ideas from statistical machine translation. in particular, they consider a model that ‘translates’ from the context of a conversation to the associated response. this is shown to give superior performance to previous information retrieval (e.g. nearest neighbour) approaches (jafarpour et al., 2010). this idea was further developed by sordoni et al. (2015b) to exploit information from a longer context, using a structure similar to the recurrent neural network encoder-decoder model (cho et al., 2014). this achieves rather poor performance on a-b-a twitter triples when measured by the bleu score (a standard for machine translation), yet performs comparatively better than the model of ritter et al. (2011). their results were also verified with a human-subject study. a similar encoder-decoder framework for dialogue is presented by shang et al. (2015) and vinyals and le (2015). this model also uses one rnn to transform the input to some vector representation, and another rnn to ‘decode’ this representation to a response by generating one word at a time. the model from shang et al. (2015) was also evaluated in a human-subject study, although on a smaller scale compared to sordoni et al. (2015b). a hierarchical version of the encoder-decoder framework has also recently been proposed (serban et al., 2016). this model consists of two rnns stacked on top of each other: one ‘sentencelevel’ rnn encodes each utterance into a fixed length vector, while a ‘conversation-level’ rnn takes as input each utterance vector and outputs a vector that summarizes the conversation so far. this is mapped back to text using a recurrent decoder. this improves over the traditional encoderdecoder frameworks in both word perplexity and word error rate, particularly when bootstrapped with word embeddings derived from distributional semantics. however, the model has not been evaluated in any human-subject studies. another approach, taken in traum et al. (2015), uses information retrieval techniques to map user questions to systems responses in the domain of time-offset interaction. since the natural language interpreter, dialogue response selection, and natural language generator model are all combined, this can also be seen as a form of end-to-end dialogue system. inaba and takahashi (2016) also propose an end-to-end retrieval model, however they use neural networks to select a response from a fixed dataset. this is similar to the model used by lowe et al. (2015). our work is also inspired by nio et al. (2014) whose model, although rule-based, is not composed of modules, as it retrieves a response to the context based on cosine similarity. this is in turn related to the work on example-based dialogue modeling (lee et al., 2009). there has been some work on combining end-to-end dialogue models with auxiliary information regarding the persona or participant role of each person in the dialogue. luan et al. (2016) investigate several models that incorporate participant roles, using topic-modelling based approaches with lda. li et al. (2016a) use an embedding for each separate speaker in the conversation, which is used to condition the decoder in an lstm model. they achieve improvements in both perplexity and bleu on a twitter dataset. 36 training end-to-end dialogue systems there has also been interesting work using deep reinforcement learning for end-to-end dialogue generation. li et al. (2016b) propose using a deep q-network (dqn) for dialogue generation, using a set of reward functions designed to increase the diversity of generated responses. zhao and eskenazi (2016) similarly use a deep recurrent q-network (drqn) to replace the conventional nlu, state tracking, and dialogue policy modules for task-oriented dialogue. one of the most effective task-oriented end-to-end systems is presented by wen et al. (2016), who train an end-to-end system on a small dataset of restaurant recommendations. they show that they are able to achieve a higher task completion rate than a modular baseline, and have significantly higher scores in naturalness, comprehension, preference, and performance. overall, these models highlight the potential of end-to-end learning architectures for interactive systems. however, much work remains before these can be implemented with confidence in a variety of settings. 3. the ubuntu dialogue corpus there are several factors that motivated the creation of the ubuntu dialogue corpus. in particular, there was a lack of large, multi-turn, publicly available dialogue datasets. in addition to providing a dataset that satisfied these constraints, we wanted the dataset to be two-way (dyadic), as opposed to multi-participant chat, and we desired a task-specific domain. all of these characteristics are satisfied by the ubuntu dialogue corpus. 3.1 ubuntu chat logs the ubuntu chat logs refer to a collection of logs from ubuntu-related chat rooms on the freenode internet relay chat (irc) network. this protocol allows for real-time chat between a large number of participants. each chat room, or channel, has a particular topic, and every channel participant can see all the messages posted in a given channel. many of these channels are used for obtaining technical support with various ubuntu issues. as the contents of each channel are moderated, most interactions follow a similar pattern. a new user joins the channel, and asks a general question about a problem they are having with ubuntu. then, another more experienced user replies with a potential solution, after first addressing the ‘username’ of the first user. this is called a name mention (uthus and aha, 2013a), and is done to avoid confusion in the channel — at any given time during the day, there can be between 1 and 20 simultaneous conversations happening in some channels. in the most popular channels, there is almost never a time when only one conversation is occurring; this renders it particularly problematic to extract dyadic dialogues. a conversation between a pair of users generally stops when the problem has been solved, though some users occasionally continue to discuss a topic not related to ubuntu. despite the nature of the chat room being a constant stream of messages from multiple users, it is through the fairly rigid structure in the messages that we can extract the dialogues between users. figures 2 and 3 show an example chat room conversation from the #ubuntu channel as well as the extracted dialogues, which illustrates how users usually state the username of the intended message recipient before writing their reply (we refer to all initial questions and replies as ‘utterances’). for example, it is clear that users ‘taru’ and ‘kuja’ are engaged in a dialogue, as are users ‘old’ and ‘bur[n]er’, while user ‘ pm’ is asking an initial question, and ‘livecd’ is perhaps elaborating on a previous comment. 37 lowe, pow, serban, charlinn, liu and pineau time user utterance 03:44 old i dont run graphical ubuntu, i run ubuntu server. 03:45 kuja taru: haha sucker. 03:45 taru kuja: ? 03:45 bur[n]er old: you can use “ps ax” and “kill (pid#)” 03:45 kuja taru: anyways, you made the changes right? 03:45 taru kuja: yes. 03:45 livecd or killall speedlink 03:45 kuja taru: then from the terminal type: sudo apt-get update 03:46 pm if i install the beta version, how can i update it when the final version comes out? 03:46 taru kuja: i did. sender recipient utterance old i dont run graphical ubuntu, i run ubuntu server. bur[n]er old you can use “ps ax” and “kill (pid#)” kuja taru haha sucker. taru kuja ? kuja taru anyways, you made the changes right? taru kuja yes. kuja taru then from the terminal type: sudo apt-get update taru kuja i did. figure 2: example chat room conversation from the #ubuntu channel of the ubuntu chat logs (left), with the disentangled conversations for the ubuntu dialogue corpus (right). time user utterance [12:21] dell well, can i move the drives? [12:21] cucho dell: ah not like that [12:21] rc dell: you can’t move the drives [12:21] rc dell: definitely not [12:21] dell ok [12:21] dell lol [12:21] rc this is the problem with raid:) [12:21] dell rc haha yeah [12:22] dell cucho, i guess i could just get an enclosure and copy via usb... [12:22] cucho dell: i would advise you to get the disk sender recipient utterance dell well, can i move the drives? cucho dell ah not like that dell cucho i guess i could just get an enclosure and copy via usb cucho dell i would advise you to get the disk dell well, can i move the drives? rc dell you can’t move the drives. definitely not. this is the problem with raid :) dell rc haha yeah figure 3: example of before (left) and after (right) the algorithm adds and concatenates utterances in dialogue extraction. since rc only addresses dell, all of his utterances are added, however this is not done for dell as he addresses both rc and cucho. 3.2 dataset creation in order to create the ubuntu dialogue corpus, first a method had to be devised to extract dyadic dialogues from the chat room multi-party conversations. the first step was to separate every message into 4-tuples of (time, sender, recipient, utterance). given these 4-tuples, it is straightforward 38 training end-to-end dialogue systems figure 4: plot of number of conversations with a given number of turns. both axes use a log scale. to group all tuples where there is a matching sender and recipient. although it is easy to separate the time and the sender from the rest, finding the intended recipient of the message is not always trivial. 3.2.1 recipient identification while in most cases the recipient is the first word of the utterance, it is sometimes located at the end, or not at all in the case of initial questions. furthermore, some users choose names corresponding to common english words, such as ‘the’ or ‘stop’, which could lead to many false positives. in order to solve this issue, we create a dictionary of usernames from the current and previous days, and compare the first word of each utterance to its entries. if a match is found, and the word does not correspond to a very common english word4, it is assumed that this user was the intended recipient of the message. if no matches are found, it is assumed that the message was an initial question, and the recipient value is left empty. 3.2.2 utterance creation the dialogue extraction algorithm works backwards from the first response to find the initial question that was replied to, within a time frame of 3 minutes. a first response is identified by the presence of a recipient name (someone from the recent conversation history). the initial question is identified to be the most recent utterance by the recipient identified in the first response. all utterances that do not qualify as a first response or an initial question are discarded; initial questions that do not generate any response are also discarded. we additionally discard conversations longer than five utterances where one user says more than 80% of the utterances, as these are 4. we use the gnu aspell spell checking dictionary. 39 lowe, pow, serban, charlinn, liu and pineau # dialogues (human-human) 936,000 # utterances (in total) 7,100,000 # words (in total) 100,000,000 min. # turns per dialogue 3 avg. # turns per dialogue 7.71 avg. # words per utterance 10.34 median conversation length (min) 6 training set dialogues 898,000 validation/test set dialogues 19,000 training set examples unspecified table 1: properties of ubuntu dialogue corpus. note that any number of training examples can be specified during creation of the training set. depending on the desired number of examples, multiple passes are made through the dataset, where each pass samples a new context stochastically from each dialogue. very large training sets are possible, yet they will have overlapping examples. typically not representative of real chat dialogues. finally, we consider only extracted dialogues that consist of 3 turns or more to encourage the modeling of longer-term dependencies. to alleviate the problem of ‘holes’ in the dialogue, where one user does not address the other explicitly, as in figure 3, we check whether each user talks to someone else for the duration of their conversation. if not, all non-addressed utterances are added to the dialogue. an example conversation along with the extracted dialogues is shown in figure 3. note that we also concatenate all consecutive utterances from a given user. we do not apply any further pre-processing (e.g. tokenization, stemming) to the data as released in the ubuntu dialogue corpus. however the use of pre-processing is standard for most nlp systems, and was also used in our analysis (see section 4). 3.2.3 special cases and limitations it is often the case that a user will post an initial question, and multiple people will respond to it with different answers. in this instance, each conversation between the first user and the user who replied is treated as a separate dialogue. this has the unfortunate side-effect of having the initial question appear multiple times in several dialogues. however the number of such cases is sufficiently small compared to the size of the dataset. another issue to note is that the utterance posting time is not considered for segmenting conversations between two users. even if two users have a conversation that spans multiple hours, or even days, this is treated as a single dialogue. however, such dialogues are rare. we include the posting time in the corpus so that other researchers may filter as desired. 3.3 dataset statistics table 1 summarizes properties of the ubuntu dialogue corpus. one of the most important features of the ubuntu chat logs is its size. this is crucial for research into building dialogue managers based 40 training end-to-end dialogue systems on neural architectures. another important characteristic is the number of turns in these dialogues. the distribution of the number of turns is shown in figure 4. it can be seen that the number of dialogues and turns per dialogue follow an approximate power law relationship. 3.4 test set generation we set aside 2% of the ubuntu dialogue corpus conversations to form a test set that can be used for evaluation of response selection algorithms5. compared to the rest of the corpus, this test set has been further processed to extract a pair of (context, response, flag) triples from each dialogue. the flag is a boolean variable indicating whether or not the response was the actual next utterance after the given context. the response is a target (output) utterance which we aim to correctly identify. the context consists of the sequence of utterances appearing in the conversation prior to the response. we create a pair of triples, where one triple contains the correct response (i.e. the actual next utterance in the dialogue), and the other triple contains a false response, sampled randomly from elsewhere within the test set. the flag is set to 1 in the first case and to 0 in the second case. an example pair is shown in table 2. to make the task harder, we can move from pairs of responses (one correct, one incorrect) to a larger set of wrong responses (all with flag=0). in our experiments below, we consider both the case of 1 wrong response and 10 wrong responses. context response flag well, can i move the drives? i guess i could just 1 eot ah not like that get an enclosure and copy via usb well, can i move the drives? you can use “ps ax” 0 eot ah not like that and “kill (pid #)” table 2: test set example with (context, reply, flag) format. the ‘ eot ’ tag is used to denote the end of a user’s turn within the context, and the ‘ eou ’ tag is used to denote the end of a user utterance without a change of turn. since we want to learn to predict all parts of a conversation, as opposed to only the closing statement, we consider various portions of context for the conversations in the test set. the context size is determined stochastically by uniform sampling6: c =∼ unif(2, t− 1). here, parameter t is the actual length of that dialogue (thus the constraint that c ≤ t − 1). in practice, this leads to short test dialogues having short contexts, while longer dialogues are often broken into a combination of short, medium, and long contexts. 5. note that, contrary to the original ubuntu dialogue corpus, the updated version separates the training, validation, and test sets by time. that is, the training set consists of conversations that started from 2004 to approximately april 27, 2012; the validation set consists of dialogues starting from april 27 to august 7, 2012; and the test set has dialogues from august 7 to december 1, 2012. this mimics the training of dialogue systems in practice, where we only have access to data in the past, and want to answer user queries in the future. 6. note that this is a different formula than the original ubuntu dialogue corpus, which sampled from a decreasing distribution. the new formula is simpler and leads to longer sampled contexts, which we consider desirable. 41 lowe, pow, serban, charlinn, liu and pineau note that, except for the plot in figure 7, all experiments, results, and analysis in this paper will refer to the updated ubuntu dialogue corpus v2. 4. response classification architectures to provide further evidence of the value of our dataset for research into neural architectures for dialogue managers, we provide performance benchmarks using two different training and evaluation criteria: response classification, and response generation. we first consider response classification architectures, which attempt to distinguish between valid and invalid next responses to the context of a conversation. these are trained on the task of best response selection, which we call next utterance classification (nuc). this can be achieved by processing the data as described in section 3.4, without requiring any human labels. this classification task is an adaptation of the recall and precision metrics previously applied to dialogue datasets (schatzmann et al., 2005). note that retrieval models trained on the task of nuc are still end-to-end, as the natural language understanding, dialogue planning, and generation modules are combined, and the system is learned with a single supervision signal. these models can be used to ‘generate’ the next utterance in a conversation by retrieving the most probable next utterance from the entire training set, given the context. thus, we can also evaluate these models using several generative metrics, that compare the selected response to the ground-truth response. we carry this out in section 4.5. we consider one naive model and two neural network-based retrieval models. the approaches considered are: tf-idf, and models using recurrent neural networks (rnn) and long short-term memory (lstm). prior to applying each method, we perform standard pre-processing of the data using the nltk7 library and twitter tokenizer8 to parse each utterance. we use generic tags for various word categories, such as names, locations, organizations, urls, and system paths. to train the rnn and lstm architectures, we process the full training ubuntu dialogue corpus into the same format as the test set described in section 3.4, extracting (context, response, flag) triples from dialogues. for the training set, we sample the responses in the same way described in section 3.4. one can generate any number of training examples by iterating several times through the training data. negative responses are selected at random from the rest of the training data. we note that for all models presented in this paper, the entire context of the dialogue that is available (i.e. after the context length sampling procedure in section 3.4 used to create the dataset) is taken into account, and not just the most recent utterance. for the response classification architectures, this is done by concatenating all context utterances together. we note that the models proposed below do not explicitly take into account ordinal information. the reasons for doing this are two-fold. first, training neural networks using classification for ranking tasks is well-established in the literature (bordes et al., 2014), and is both simple to implement and effective in practice. second, in the ubuntu dialogue corpus we do not have supervised ordinal data for the relative quality of next responses given a context. more advanced methods could consider some way to approximate this ordinal information, such that a neural network model could be explicitly trained as a ranking system; however, this is beyond the scope of this paper. 7. www.nltk.org/ 8. http://www.ark.cs.cmu.edu/tweetnlp/ 42 training end-to-end dialogue systems 4.1 tf-idf term frequency-inverse document frequency is a statistic that intends to capture how important a given word is to some document, which in our case is the context (ramos, 2003). it is a technique often used in document classification and information retrieval. the ‘term-frequency’ term is simply a count of the number of times a word appears in a given context, while the ‘inverse document frequency’ term puts a penalty on how often the word appears elsewhere in the corpus. the final score is calculated as the product of these two terms, and has the form: tfidf(w, d,d) = f(w, d)× log |d| |{d ∈ d : w ∈ d}| , (1) where f(w, d) indicates the number of times word w appeared in context d and the denominator represents the number of dialogues in which the word w appears. for classification, the tf-idf vectors are first calculated for the context and each of the candidate responses. given a set of candidate response vectors, the one with the highest cosine similarity to the context vector is selected as the output. for recall@k, the top k responses are returned. 4.2 rnn dual encoder recurrent neural networks are a variant of neural networks that allows for time-delayed directed cycles between units (elman, 1990). this leads to the formation of an internal state of the network, ht at time step (word index) t, which allows it to model time-dependent data. the internal state is updated at each time step as some function of the observed variables xt, and the hidden state at the previous time step ht−1, with w x and w h matrices associated with the input and hidden state: ht = f(w hht−1 +w xxt). (2) rnns have been the primary building block of many current neural models for language-related tasks (sutskever et al., 2014; sordoni et al., 2015b), which use rnns as encoders and decoders; in this case, the first rnn is used to encode the given context, and the second rnn generates a response by using beam-search, where its initial hidden state is biased using the final hidden state from the first rnn. we detail such models in section 5. however, in this section, we are concerned with classification of responses, and thus using a decoder rnn for generation is not strictly necessary (and thus is not used in the model shown in figure 5). in this section we build upon the approach in (bordes et al., 2014), which has also been recently applied to the problem of question answering (yu et al., 2014), and use rnns for classification rather then generation. we use a siamese network consisting of two rnns with tied weights to produce the embeddings for the context and response, that we call the dual-encoder (de) model. given some input context and response, we compute their embeddings — c, r ∈ rd, respectively — by feeding the word embeddings one at a time into its respective rnn. word embeddings are initialized using the pretrained vectors (common crawl, 840b tokens from (pennington et al., 2014)), and fine-tuned during training. the hidden state of the rnn is updated at each step, and the final hidden state represents a summary of the input utterance. using the final hidden states from both rnns, we then calculate the probability that this is a valid pair: p(flag = 1|c, r,m) = σ(ctmr + b), (3) 43 lowe, pow, serban, charlinn, liu and pineau figure 5: diagram of the dual encoder (de) model. the rnns have tied weights. c, r are the last hidden states from the rnns. ci, ri are word vectors for the context and response, i < t. we consider contexts up to a maximum of t = 160. where the bias b and the matrix m ∈ rd×d are learned model parameters. this can be thought of as a generative approach; given some input response, we generate a context with the product c′ = mr, and measure the similarity to the actual context using the dot product. this is converted to a probability with the sigmoid function. the model is trained by minimizing the cross entropy of all labeled (context, response) pairs (yu et al., 2014): l = − ∑ n log p(flagn|cn, rn,m) (4) where ||θ||2f is the frobenius norm of θ = {m, b}. a diagram of the de model can be seen in figure 5. for training, we used a 1:1 ratio between true responses (flag = 1), and negative responses (flag=0) drawn randomly from elsewhere in the training set. the rnn architecture is set to 1 hidden layer with 100 neurons (optimized over {10, 50, 100, 200, 300}), and a learning rate of 0.0001 (optimized over {0.1, 0.01, 0.001, 0.0001}). the w h matrix is initialized using orthogonal weights (saxe et al., 2013), while w x is initialized using a uniform distribution with values between -0.01 and 0.01. we use the first-order stochastic gradient optimization procedure adam (kingma and ba, 2014) with the default parameters, using gradients clipped to 10 and a batch size of 512 (optimized over {128, 256, 512}). we found that weight initialization as well as the choice of optimizer were critical for training the rnns. 4.3 lstm dual encoder in addition to the rnn model, we consider the same architecture but change the hidden units to long-short term memory (lstm) units (hochreiter and schmidhuber, 1997). lstms were introduced in order to model longer-term dependencies. this is accomplished using a series of gates that determine whether a new input should be remembered, forgotten (and the old value retained), or used as output. the error signal can now be propagated back much further using the gates of the lstm unit. this helps overcome the vanishing gradient and exploding gradient problems in standard rnns, where the error gradients would otherwise decrease or increase at an exponential rate. for this model, we used 1 hidden layer with 200 neurons, a learning rate of 0.001, and a batch size of 256 (optimized over the same values as the rnn). we again use the default adam settings, 44 training end-to-end dialogue systems and initialize the forget gate bias of the lstm to 2.0. the hyper-parameter configuration (including number of neurons) was optimized independently for rnns and lstms using a separate validation set extracted from the training data. 4.4 evaluation metrics we consider two types of evaluation metrics: retrieval metrics, and generative metrics. these metrics are applicable to both models trained on the task of nuc, detailed in this section, and the generative models introduced in section 5. in particular, they offer two ways of automatically evaluating dialogue systems trained in an end-to-end manner. for retrieval, we evaluate using recall@k (denoted r@1 r@2, r@5 below), which has often been used in language tasks. here the agent is asked to select the k most likely responses, and it is correct if the true response is among these k candidates. only the r@1 metric is relevant in the case of binary classification (as in the table 2 example). although a language model that performs well on these retrieval metrics is not guaranteed to achieve good performance on utterance generation, we hypothesize that improvements on a model with regards to the classification task will eventually lead to improvements for the generation task. see section 6 for further discussion of this point. we also consider generative metrics that compare the generated or retrieved utterance to the ground-truth next utterance. in general, this is a hard open problem (liu et al., 2016). we use methods based on word embeddings that have recently been proposed for use in evaluating nontask oriented dialogue systems, when no task completion signal is available. these metrics use external word embeddings trained via distributional semantics, such as glove, to determine how close the generated utterance is to the ground truth next utterance. we note that these metrics do not necessarily correlate strongly with human judgement (liu et al., 2016); here, we consider them to be measures of the topicality of the retrieved responses. if the generated response and groundtruth response are semantically similar, then the vector-based metrics should be higher, as word embeddings themselves contain semantic information (mikolov et al., 2013). it is because of this interpretation that we prefer the vector-based metrics over word-overlap metrics such as bleu. the embedding average score approximates the compositional embedding of each utterance by taking an average of the word vectors that compose the utterance. the utterance embedding similarities are then compared using cosine similarity. the greedy matching score, originally used to analyze semantic similarity between sentences in intelligent tutoring systems (rus and lintean, 2012), matches the most similar word in the generated utterance to the actual utterance using cosine similarity of the word embeddings. the vector extrema score was proposed by (forgues et al., 2014) for dialogue systems. instead of averaging each word embedding, this approach takes the elementwise maximum (or minimum) of each component in the word vectors composing an utterance. this results in utterance embeddings of the same size of each word vector, which can again be compared using cosine similarity. these metrics are able to capture some aspects of dialogue that are not present in bleu score. we note that, to keep the assumptions of independent and identically distributed (i.i.d.) training and test data examples valid, it is important to have the word embeddings used for the metrics trained on a different corpus than the task corpus, as argued by liu et al. (2016). this preserves the statistical independence between the task and each performance metric, and alleviates the possibility of spurious and potentially misleading correlations between data examples. 45 lowe, pow, serban, charlinn, liu and pineau 4.5 experimental results we examine the performance of the models using both retrieval and vector based metrics, as shown in tables 3 and 4. for nuc, the models were evaluated using both 1 (1 in 2) and 9 (1 in 10) false examples.9 retrieval metrics method 1 in 2 r@1 1 in 10 r@1 1 in 10 r@2 1 in 10 r@5 tf-idf 74.9% 48.8% 58.7% 76.3% dual encoder w/rnn units 77.7% 37.9% 56.1% 83.6% dual encoder w/lstm units 86.9% 55.2% 72.1% 92.4% table 3: results for the three algorithms using various recall measures for binary (1 in 2) and 1 in 10 (1 in 10) next utterance classification %. generative metrics method embedding average greedy matching vector extrema tf-idf 0.536 0.370 0.342 dual encoder w/ lstm units 0.650 0.413 0.376 table 4: results for tf-idf and the de model with lstm units on the embedding average, greedy matching, and vector extrema scores. these scores provide an estimate of the topic consistency of the generated responses. context “any apache hax around ? i just deleted all of path which package provides it ?”, “reconfiguring apache do n’t solve it ?” ranked responses flag 1. “does n’t seem to, no” 1 2. “you can log in but not transfer files ?” 0 figure 6: example showing the ranked responses from the lstm. each utterance is shown after pre-processing steps. we observe that the dual encoder with lstm units outperforms both the dual encoder with rnn units and tf-idf on all evaluation metrics. it is interesting to note that tf-idf actually outperforms the rnn on the recall@1 case for the 1 in 10 classification. this is most likely due to the limited ability of the rnn to take into account long contexts, which can be overcome by using the lstm. an example output of the lstm where the response is correctly classified is shown in figure 6. we also show, in figure 7, the increase in performance of the lstm as the amount of data used for training increases. this confirms the importance of having a large training set. 9. the performance metrics recall@2 and recall@5 are not relevant in the binary classification case. 46 training end-to-end dialogue systems figure 7: the lstm (with 200 hidden units), showing recall@1 for the 1 in 10 classification, with increasing dataset sizes up to 120k dialogues. note that this was calculated using the old version of the ubuntu dialogue corpus, and thus the recall@1 values are higher than those in table 3. 4.6 qualitative error analysis there are a large number of technical challenges that must be solved in order to construct a system that can provide adequate responses in a dialogue. in fact, almost all common challenges in natural language processing are present in some form or another in the dialogue problem. these include, but are not limited to: coreference resolution, lexical semantics, discourse coherence and cohesion, natural language understanding, natural language generation, compositional semantics, the turn taking structure of dialogue, and more. further, it is often necessary to have some technical knowledge about the subject matter being discussed. it is clear that current end-to-end dialogue systems are not able to adequately address all these problems, yet precisely which aspects of conversation are the most prevalent sources of errors remains relatively unknown. this is particularly true for neural network models for dialogue, which have only recently come into prominence. we undertake the task of evaluating an end-to-end dialogue system, the de model with lstm units, on the ubuntu dialogue corpus for the nuc task. we hope that an understanding of the most common errors made by this model can help inform future work on neural dialogue systems, particularly on the ubuntu dialogue corpus. we conduct an error analysis with three participants10 evaluating a total of 100 randomly chosen errors made by the dual encoder. for each error made by the model11, we consider what abilities the model would need to have in order to answer the question correctly. we classify these into several 10. participants were graduate students in computer science, who had familiarity with both dialogue systems and the ubuntu domain. 11. we consider an error to be any example where the correct response is not the top 1 response ranked by the model. 47 lowe, pow, serban, charlinn, liu and pineau categories: using knowledge, understanding tone and style of the responses, a better understanding of the semantic similarity of phrases, and explicitly considering the turn-taking nature of dialogue. since we are evaluating a classification model, we do not consider problems associated with natural language generation. note that an error can be classified into multiple categories, if they are each necessary to answer the question correctly. in addition to classifying the errors made by the model, we qualitatively evaluate the difficulty of the questions on a scale from 1-5. a rating of 1 on the difficulty scale means that the question is easily answerable by all humans. a 2 indicates moderate difficulty, which should still be answerable by all humans but only if they are paying attention. a 3 means that the question is fairly challenging, and may either require some familiarity with ubuntu or the human respondent paying very close attention to answer correctly. a 4 is very hard, usually meaning that there are other responses that are nearly as good as the true response; many humans would be unable to answer questions of difficulty 4 correctly. a 5 means that the question is effectively impossible: either the true response is completely unrelated to the context, or it is very short and generic. finally, we evaluate the appropriateness of the response chosen by the model for each question on a scale from 1-3. a score of 1 indicates that the chosen response is completely unreasonable given the context. a 2 means that the response chosen was somewhat reasonable, and that it’s possible for a human to make a similar mistake. a 3 means that the model’s response was more suited to the context than the actual response. we note that such an analysis is partially dependent on the model and the domain. the endto-end system from wen et al. (2016) achieves very strong performance, and thus would not face exactly the same problems as the models we present here. however, this is because the model was trained on a very narrow dataset of restaurant recommendations, and thus the space of generated responses is comparatively small. we believe that other conversational models trained on large, complex datasets are likely to encounter the same problems that we present here. in the ubuntu domain, questions where using external knowledge would be helpful for the model involve technical terminology. in most cases, the correct response contains the name of a command or process that is related to one stated in the context; however, the model is unable to link the two together. an example of such a question is shown in figure 8. in this case, the context of the conversation is about file searching in ubuntu, and the correct response (in italics) mentions the locate command. this response would have been assigned a higher probability if it was able to determine the meaning of the locate command. there are other examples where the model may be incapable of taking into account the specific tone or style of the users in the conversation. for instance, a speaker may use many emoticons, have poor english grammar skills and be prone to misspelling words, use frequent abbreviations, or use a particularly formal tone. being able to spot these distinct language features could lead the model to improve its performance in terms of selecting the actual next response. an example of this is shown in figure 9. in the context, speaker a appends his question with an (unnecessary) smiley face. thus, it is more likely that the candidate response with multiple smiley faces is the correct response. one of the most important challenges in natural language is understanding the semantics of phrases. classification dialogue models can make errors due to an inability to detect semantic similarities between sentences, or due to the detection of spurious similarities. this category covers the general case where the topic of the model’s response is clearly different than the topic of the context and true response. in this category, we also define two special cases: one where there is a 48 training end-to-end dialogue systems context: speaker a: is there anything i can do to make ubuntus filesearching faster ? i am running from an ssd and it ’s still painfully slow eou speaker b: searching how ? eou speaker a: hitting the search button from nautilus eou searching systemwide eou binary probability candidate responses 0.10 i tend to use the locate command . eou 0.38 i‘m not that into it , but it has to be session in one , or track in one or something to have the rw funtion eou 0.06 np eou 0.56 probably just junky firefox eou i bet you have a tonne of addons eou that all takes resources eou what other apps are you running ? eou how much ram frees after you close skype ( if its convenient ) 0.48 except i get an invite eou 0.44 installing from source on ubuntu isn’t a great idea imo . but look for a make uninstall option eou 0.16 oh i get it , thanks a lot eou 0.21 the python one eou i think cron may be able to do that .. to restart a task if it dies out prematurely eou well you can show your #python script and people may suggest the best way to *overcome and premature *unknown** .. eou if it ’s a buggy script then you’d expect it to be very problematic with anything starting it eou 0.22 or killall ftl* eou 0.54 any help with custom msg eou figure 8: example where the model would benefit from using an external knowledge source. correct answer is in italics, and the model’s selected answer is in bold. note that the probabilities do not sum up to 1, as they are binary probabilities – the model considers each candidate response independently. context: speaker a: whats the best rdp software for ubuntu ? i want to be able to rdp into my ubuntu desktop from my ubuntu laptop :) eou speaker b: then just use vnc . eou binary probability candidate responses 0.15 what software , do you have any links to show how to to it ? : ) eou ur a beast ! ty again :) eou 0.04 wrong place stop it eou 0.33 could use puppet :) eou 0.29 lol eou 0.17 i’ve installed mine from “ additional drivers ” eou 0.36 xp yeah that its been a long time since i last used my vpn server eou 0.00 yes , and where does it say it ’s released , or you can buy it , in actually anything about it eou **unknown** ’s not released eou it ’s a concept canonical are working on / trying to create eou 0.75 it ’s that “ persistence ” stuff ? what do you mean ? eou 0.00 sudo chown root : root /tmp && sudo chmod 1777 /tmp eou 0.00 python eou python eou figure 9: example where the model would benefit from understanding the tone and style of the speakers. correct answer is in italics, and the model’s selected answer is in bold. 49 lowe, pow, serban, charlinn, liu and pineau context: speaker a: how do i move programs and all their dependencies automatically ? eou speaker b: you mean between two ubuntu installations ? eou binary probability candidate responses 0.04 i have a chroot partition i’ve been using ldd and doing it manually but it is quite slow eou 0.00 if your not a geek then you won’t understand eou would you advise your grandmother to try and install linux ? eou 0.02 that ’s where im stumped ... an older kernel made no difference , whilst an older release of ubuntu did eou 0.71 well , everything with indicators is basically dbus eou 0.01 it ’s cool that you help people who run free software ;) bye eou 0.00 not by default , but it can . /var holds a lot of temp stuff like logs and debs , you don’t need those cluttering your ssd and using write cycles eou also , move your web cache to ramdisk to make it fast as well as not use your hdd at all :) eou its a disk space ... in ram eou 0.59 yes , but they speak http so i could use the browser as a low-level access tool for browing repos and i would like to do that , but that doesn’t seem to work . eou 0.61 sure eou 0.97 i believe it uses gdm but i’m not sure . the login manager thing looks the same as the stock ubuntu 12.10 one eou 0.14 ok , ty eou figure 10: example where the model would benefit from the ability to conduct high-level inference to better understand the semantic similarity between context and correct response. correct answer is in italics, and the model’s selected answer is in bold. direct word copying between the context and true response that the model failed to detect, and one where some high-level inference is required to answer correctly. an example of the latter case is shown in figure 10; speaker a asks how to install some programs automatically, and the correct response states that they had previously been ‘doing it manually’. thus, if the model was able to infer that a person who has asked to perform an operation automatically could previously have been doing it manually, it would have assigned a higher probability to the correct response. finally, we consider errors where the model is unable to account for the turn-taking structure of dialogue. for the ubuntu dialogue corpus, interactions between users usually take a certain form, where one user is asking for help and the other user is providing answers. thus, it is important to consider the role of the current user when selecting the correct response; indeed, there has been preliminary work in this direction (luan et al., 2016; li et al., 2016a). we also consider a special case of this error, when the last utterance in the context asks a question and the response chosen by the model is not answering any question at all. for example, in figure 11, the final utterance asks the question ‘no mm’s?’. the response selected by the model begins with ‘thanks’, which is clearly not a reasonable response to a question. an example of the general turn-taking error is shown in figure 12. this depicts a typical dialogue between two users in the ubuntu dialogue corpus: speaker a is having trouble with their brightness keys, and speaker b is trying to help them. the model must predict the next response of speaker a. in the first response, the user states that they are appreciative of the help being given, which fits with speaker a’s behaviour in the context; thus, it is more likely to be the correct response. we examined 100 randomly selected errors of the de model on the ubuntu corpus to compute the number of errors in each category; the results are shown in table 13. we can first note that there is significant progress to be made for classification models on the ubuntu dialogue corpus; over half (60%) of the errors made by the model can be considered feasible for the majority of humans (1-3 on the difficulty rating). however, the number of questions that every human could 50 training end-to-end dialogue systems context: speaker a: i can’t seem to get audio working as a non-root user . has anyone ever had this problem ? eou speaker b: alsamixer to the rescue eou speaker a: alsamixer shows everything turned on , and looks exactly the same for my normal user as it does for root eou speaker b: no mm’s? eou binary probability candidate responses 0.62 correct eou 0.92 true but how will he find my new ip so easily if i get it changed ? all i do is programming c and check my mail usually eou 0.68 yes strange . then omit the dash altogether , try giving set default sink :/ eou 0.93 thanks eou where is the db app i cannot locate it ( sorry to be such a noob ! ) eou 0.03 i’m switching the location to my on board ssd drive that ’s embedded to the laptops board . i just haven’t been using the storage so i figure i could try and utilize the space while the ram being 8 gig ’s itself i see no problem with the switch . do you understand what i’m doing . i’m only asking here so i don’t go screwing up and save myself hours of headaches eou 0.03 you’ll love it eou i am joking , but you will probably enjoy learning about it eou well it ’s a step up from opening your hard drive up and using a magnet eou 0.44 yeah , 512 is plenty eou 0.20 i use it on a number of machines with no problems . just this one . eou modprobe pulls up a variety of mouse drivers eou 0.62 so the issue is **unknown** **unknown** . gz ’ is different from the same file on the system ” but i don’t have any idea why/what that means , sorry . best of luck . eou 0.89 http://www.geforce.com/hardware/desktop-gpus/geforce-gtx-680m comes with optimus technology . so i think it has an onboard intel card eou figure 11: example where the model selects an inappropriate response to a question. correct answer is in italics, and the model’s selected answer is in bold. context: speaker a: hi eou i have a problem with fn keys for brightness with my laptop and nvidia propertiary driver eou speaker b: what make and model laptop ? ¿ eou speaker a: sony vaio vgn fz31z eou and im using nvidia propertiary driver version current ( recommended one ) eou speaker b: try the boot option : acpi backlight=vendor eou speaker a: i have added acpi backlight for vendor i have updated grub but the keys are not working eou my grub cmd line linxu default : quiet splash acpi backlight=vendor eou speaker b: try the boot option : acpi osi=linux eou speaker a: ok i must remove the acpi backlihgt/ eou speaker b: i’d also report a bug eou could try quantal livecd to see if the newer kernel plays nicer eou speaker a: i think that is a nvidia problem with the propertiary eou speaker b: possibly , or it could be acpi based eou speaker a: ok thank you i must remove the previous about the vendor ok ? eou speaker b: could try both and then just one eou binary probability candidate responses 0.38 thanks for the help . trying now . is there any other same bug report for vaio/ eou 0.20 it ’s actually ubuntu support , since i’m using ubuntu , isn’t it ? eou 0.49 yes eou the usb disk will just be seen as a hard disk , install to it eou 0.50 if you do unattended-upgrades -d , that might tell you a few things ? eou 0.58 does this have ’ open terminal here ’ and ’ 2pane mode ’ options ? eou found terminal option , just looking for 2pane eou 0.71 it ’s cool eou 0.02 it ’s like hotel internet eou http://www.fdlinux.com/networksetuphowto.html eou 0.01 i’ll check what it means in google . thank you . eou 0.36 i never liked it ... for thin versions , i use fluxbox or some other window manager eou 0.49 so do i just paste that code in to the beginning of the script ... ? eou sorry experienced linux user , very very novice coder ;-p eou figure 12: example where the model does not take into account the roles of the participants in the dialogue. correct answer is in italics, and the model’s selected answer is in bold. 51 lowe, pow, serban, charlinn, liu and pineau difficulty rating (1-5) number of errors % of errors impossible (5) 19 19% very difficult (4) 21 21% difficult (3) 22 22% moderate (2) 25 25% easy (1) 13 13% model response rating (1-3) very reasonable (3) 14 14% somewhat reasonable (2) 37 37% unreasonable (1) 49 49% error category tone and style 8 9% knowledge 18 20% semantic similarity 45 49% word copying 11 12% high-level inference 16 18% turn-taking structure 20 22% answering questions 6 7% figure 13: qualitative evaluation of the errors from the de model. note that counts for parent categories (semantic similarity and turn-taking structure) include the counts for the child categories. error categories are not classified for impossible questions and are not mutually exclusive, thus totals may not add up to 100. answer unconditionally is small, as technical language can often be confusing for people who are unaccustomed to it. the other questions are roughly uniformly distributed over the remaining levels of difficulty, from moderate to impossible. we also note that there are a large number of cases (49%) where the response retrieved by the model was completely unreasonable given the context, which further indicates that there is room for improvement in these models. it is also interesting to examine the distribution of errors across the examples. as can perhaps be expected, the most common form of error was a lack of understanding of the semantics of the responses. what is more surprising is that there is a significant number of examples where the model failed to observe that there was a key word shared between the context and the correct response; this could be because there are often common words between the context and false responses in the training set, and the model is unable to distinguish between words that are relatively unimportant and those that carry significant semantic meaning. thus, there is much progress to be made in dialogue systems by working on the general problem of natural language understanding. there are many examples where the model could be improved by explicitly accounting for the turn-taking structure of dialogue, as there were often instances where the model selected a response that was not suited to the current speaker. in several cases, the model also needed some form of external knowledge base in order to answer the question correctly. note that the number of such examples in table 13 refers to instances where the correct response mentions a ubuntu term that is related but not identical to the terminology in the context; if this were to be extended to all questions where technical vocabulary is mentioned, the number would be significantly higher. finally, there is a small number of cases where a better understanding of the tone of the dialogue would help the model, however this does not seem to be the best direction for future research. 52 training end-to-end dialogue systems figure 14: diagram of the rnn architecture for dialogue modeling. note that utterances in the context are concatenated together before being fed into the rnn. 5. generative response architectures in order to aid progress towards the goal of building fully generative conversational models, we present baseline models for generating responses conditioned on the context of the conversation for the ubuntu dialogue corpus. it should be noted that the format of the dataset can easily be altered to support training in this manner: one can simply remove all (context, response, flag) triples with flag = 0, and be left with only the valid (context,response) pairs. 5.1 generative recurrent neural language model we first describe the standard recurrent neural network language model (rnn-lm), as shown in figure 4.6, a neural network model which is used to predict the next word in a sequence of words (mikolov et al., 2010). the model observes the dialogue word-by-word and updates its hidden state ht at time step (word index) t. given a hidden state ht the model outputs a probability distribution over all words in the vocabulary. formally, the model computes the hidden state as described in equation (2): ht = f(w hht−1 + w xxt), where, ht−1 is the previous hidden state, xt is the current input word, f(·) is a non-linear activation function such as tanh, and w x and w h are model parameters. we will take ŷn = {w1, ..., wt } as the target output sequence for the nth training example, given the input sequence x̂n = {x1, ..., xt }. the conditional distribution over each output symbol is computed in a similar manner, and depends on the current hidden state ht: p(wt|wt−1, ..., w1) = exp(w o wt · ht)∑ w exp([w o]w · ht) , (5) where w o is the output matrix, and w o w indicates the row of the w o matrix corresponding to the output index for word w. the model is trained with teacher forcing, meaning the input xt to the network is the previous ground-truth output word wt−1. the model is trained in an end-to-end fashion by gradient descent to maximize the conditional log-likelihood of input-output pairs from the training set, {x̂n, ŷn}: max θ 1 n n∑ n=1 log pθ(ŷn|x̂n), (6) 53 lowe, pow, serban, charlinn, liu and pineau figure 15: diagram of the hred model. note that each utterance in the context is encoded with a separate ‘utterance-level’ encoder, which is then fed into a ‘context-level’ encoder. where θ are the parameters of the model, including w x,w h,w o, and the corresponding biases. thus, the model learns a probability distribution over all output sequences, p(ŷ1, ..., ŷt ). for dialogue response generation, the model is conditioned on the previous dialogue context and used to generate a response, i.e. the next utterance in the dialogue. such a model could be used as a full dialogue system, as defined in section 1.1, to carry out a conversation with a user. 5.2 hierarchical recurrent encoder-decoder one problem with directly applying a standard rnn language model to modeling dialogues is that it does not take into account the turn-taking nature of conversations. it is well known that recurrent neural networks have trouble learning long-term dependencies (bengio et al., 1994), a problem only partially alleviated with lstms. thus, if a long context is fed into the encoder, it is possible for the model to put a large weight on only the most recent utterance. in order to investigate models that are able to retain state over long conversations, we implement the recently proposed hierarchical recurrent encoder-decoder (hred) sordoni et al. (2015a). while this model was initially proposed for context-sensitive query suggestion, it has been adapted for dialogue response generation on a dataset of movie subtitles (serban et al., 2016). the hred model builds on the traditional encoder-decoder model (cho et al., 2014). the main addition is a second encoder, the utterance-level encoder, that takes as input the fixed-length vectors produced by the lower-level encoder, which we refer to as the word-level encoder. instead of letting the decoder take as input the fixed-length vectors from the word-level encoder, the decoder takes as input the output of the utterance-level encoder. intuitively, the utterance-level encoder summarizes the history of the conversations into a single vector, which is more sensitive to previous utterances in the conversation. this provides a more powerful architecture as it is now possible for the model to encode order-dependent patterns inherent in the turn-taking nature of dialogue. as before, the 54 training end-to-end dialogue systems model is trained end-to-end to maximize the log-likelihood of the generated utterance. a diagram of the hred model can be seen in figure 5.1. to summarize: the rnn-lm (figure 4.6) uses a single rnn (or lstm) to encode the entire history of the dialogue, which consists of all utterances in the context concatenated together. it then uses the same rnn (i.e. an rnn with the same parameters) to decode the prediction utterance. the encoder-decoder model (not shown here) augments this with a second rnn with different parameters for the decoder. all context utterances are still concatenated in the encoder, and thus it is difficult to model long-term dependencies for utterances that occur earlier in the dialogue. this problem is alleviated with the hred model (figure 5.1), which does not concatenate the context utterances: each is encoded with a separate utterance-level encoder, whose output is fed into an additional context-level encoder. the output of the context-level encoder depends on all the utterances in the context, and is fed into the decoder. each rnn has separate parameters. generative metrics embedding average greedy matching vector extrema lstm-lm 0.561 0.425 0.380 hred 0.617 0.452 0.408 tf-idf 0.536 0.370 0.342 dual encoder w/ lstm units 0.650 0.413 0.376 table 5: results for both the generative and retrieval models on the embedding average, greedy matching, and vector extrema scores. these scores provide an estimate of the topic consistency of the generated responses. retrieval metrics 1 in 2 r@1 1 in 10 r@1 1 in 10 r@2 1 in 10 r@5 lstm-lm 58.9% 19.6% 33.1% 61.4% hred 61.8% 21.5% 35.8% 64.5% tf-idf 74.9% 48.8% 58.7% 76.3% dual encoder w/rnn units 77.7% 37.9% 56.1% 83.6% dual encoder w/lstm units 86.9% 55.2% 72.1% 92.4% memn2n (dodge et al., 2015) — 63.72% — — rnn-cnn (baudiš and šedivỳ, 2016) 91.1% 67.2% 80.9% 95.6% ensemble (kadlec et al., 2015) 91.5% 68.3% 81.8% 95.7% r-lstm (xu et al., 2016) 88.9% 64.9% 78.5% 93.2% table 6: results for both the generative and retrieval models using various recall measures for binary (1 in 2) and 1 in 10 (1 in 10) next utterance classification %. we include state-of-theart results from more recent papers. 55 lowe, pow, serban, charlinn, liu and pineau 5.3 experimental results we now examine the performance of the generative models using the vector-based metrics defined in section 4.4. the results for the rnn language model with lstm units (lstm-lm) and the hred model can be seen in figure 5. as expected, the hred model outperforms the baseline lstm model across all of the metrics. however, it is interesting to note that in a direct comparison with the dual encoder model, the hred model has a higher score in 2 out of the 3 metrics considered. thus, it is likely that the hred model is generating responses that are more semantically similar to the ground-truth response, and is better at staying on topic. we also present the results for these models on the nuc task. it is possible to apply the generative models to the nuc task, as all that is required is the ability to assign probabilities to sequences of utterances. again, the hred model predictably outperforms the lstm-lm model on all metrics. this coincides with the results of (serban et al., 2016) on a different dataset, using different metrics. we can also see that the generative models perform much worse than the models explicitly trained to retrieve utterances from a list. this is to be expected, as the retrieval models were trained explicitly on the nuc task, while the generative models were not. because of the discrepancy in training objectives, we do not recommend the use of nuc for comparing generative models with retrieval models. however, we believe that nuc is very useful for comparing generative models with other generative models, and retrieval models with retrieval models; we justify this further in section 6.3. 5.4 examples of generated responses in order to obtain a better understanding of the quality of responses from the generative model, we provide a table of examples from the lstm-lm and hred models in table 7. we chose several representative samples that provide insight into the behaviour, and the limitations, of the generative models.12 first, we note that there are several situations where the generative models, particularly the hred model, produce reasonable responses. for instance, in example 8), the models are able to generate a coherent and useful response concerning the installation of grub. the models achieve this despite having no knowledge of ubuntu other than what it has observed in the training corpus. the hred model is also able to generate reasonable responses when it is playing the role of the non-expert who is seeking help. this is the case in examples 1), 3), and 6). however, the responses from the model are fairly homogeneous; it is simply stating that it does not know how to solve the problem. this is a commonly observed behaviour in the generated samples of the model. given a context, it is often reasonable to simply state ‘i don’t know’, or ‘i’m not sure’, and indeed these are likely responses in the training set given many different contexts. this is the problem of generic responses, and has previously been observed in end-to-end generative dialogue models (vinyals and le, 2015). there are other instances where the lack of semantic understanding is clear. in example 2), the hred makes a relevant statement about the command ‘sudo apt-get’, stating that the user should try to run this command. however, it fails to understand that the last user utterance (‘i tried that, it comes the same error’), refers to the user having already tried this exact command. it is difficult for these models to integrate information from multiple utterances (li et al., 2016b), and to understand the concept of coreference. 12. the full set of model samples are available at: www.iulianserban.com/files/ubuntudialoguecorpus.zip 56 training end-to-end dialogue systems context response 1) i set up my hd such that i have to type a passphrase to access it at boot. how can i remove that passwrd, and just boot up normal. i did this at install, it works fine, just tired of having reboots where i need to be at terminal to type passwd in. help? → backup your data, and re-install without encryption “ might ” be the easiest method hred: how would i do that? lstm: how do i do that? 2) hi, everyone. i download the googleearth. deb, and want to install it with : sudo dpkg -i something. deb, it tells me : google-earth-stable depends on lsb-core (>= 3.2 ) ; however : package lsb-core is not installed. so i just install the lsb-core, but it comes error : what should i do? → the error suggests running “ sudo apt-get -f install ” → i tried that, it comes the same error hred: try “ sudo aptget -f install ” lstm: 3) i’ve a 2.0 tb hard drive. it no longer boots from the hard drive, but i can see it from gpart and it has a “ boot ” flag. how do i recover it? → the boot flag doesn’t mean anything in linux why do you not have a backup? you can mount the partition in livecd and access the data easily ( assuming the partition is healthy ) → i have one hard drive for the moment ( well, i don’t count the 250gb drive ) besides, i don’t have two sata cables i will make a backup as soon as i get another nice sized hard drive and a sata cable and how do i find out if the partition is healthy? → boot to ubuntu livecd and try and mount the partition hred: how do i do that? lstm: on the live cd? 4) bottom line is usually if you got 64bit why not use it? ( unless there is something that doesn’t work on 64bit, which is amitedly not much ) → well, now clip2net is not working... just seems ubuntu runs more seamlessly on 32-bit and after reading this article, im wondering if 64 is a bit hyped? hred: i have no experience with 64bit, so i don’t know what you mean. lstm: , , , , 5) how would i enable vim syntax highlighting on a default install of ubuntu 12.04? → have you insalled vim instead of the default vim-tiny? → i installed vim-runtime but still no highlighting hred: yes lstm: yes 6) its comming from my computer 15 people using same router and they have no issues, also it happened recently, before it was ok → yes but your router may have optional settings for dns which if not set to the same dns will override your network interface... it ’s an option to consider hred: how do i do that? lstm: that is what i was looking for, thanks 7) it looks as if the network ’s dhcp server is dead : “ < warn > ( eth0) : dhcpv4 request timed out ” → but it isn’t. windows works hred: that ’s what i thought, but i don’t know how to fix it lstm:? 8) i’ve been having trouble with a new 3tb harddrive and a fresh ubuntu install. i install ( as gpt ) and the install looks ok. but then get grub “ error : out of disk ” after reboot. i have an october 2011 pc with updated bios and i do see the 3tb. google searches haven’t been to helpful so far. has anyone seen grub trouble when installing on 3tb harddrives? → have you tried installing grub again and updating it? sudo grub-install sudo update-grub? → i guess i could try that. so i would have to install grub from an older ubuntu install? hred: yes, you can install grub to the mbr of your hard drive lstm: grub is installed on the mbr of the hard drive table 7: ubuntu model examples. the → token indicates a change of turn, and is the unknown token. finally, we also observed that the lstm language model in general produced many poor responses, consisting either of a single punctuation mark or a series of unknown tokens. this is seen in examples 2), 4), and 7). on the contrary, the hred model rarely produced such responses; thus, there is some indication that by improving the model architecture, we will produce models that generate longer and more interesting responses. 57 lowe, pow, serban, charlinn, liu and pineau 6. discussion this paper presents the ubuntu dialogue corpus v2, a large dataset for research in unstructured multi-turn dialogue systems. we describe the construction of the dataset and its properties. the availability of a dataset of this size opens up several interesting possibilities for research into dialogue systems based on rich neural-network architectures. we present results demonstrating use of this dataset to train end-to-end rnn-based models, and critically evaluate the errors they make. we find that, while these models hold promise for building non-task oriented dialogue systems, they still make many obvious errors, and there is significant room for improvement. next, we outline several interesting directions for future work. 6.1 conversation disentanglement our approach to conversation disentanglement consists of a small set of rules. more sophisticated techniques have been proposed, such as training a maximum-entropy classifier to cluster utterances into separate dialogues (elsner and charniak, 2008). however, since we are not trying to replicate the exact conversation between two users, but only to retrieve plausible natural dialogues, the heuristic method presented in this paper may be sufficient. this seems supported through qualitative examination of the data, but could be investigated with a more formal evaluation. 6.2 non-task oriented model evaluation it may seem unconventional that, given the technical nature of the ubuntu dialogue corpus and the fact that it involves interactions where the end goal is solving a user’s problem, we are treating our models as non-task oriented, meaning that we do not incorporate a supervised task completion or user satisfaction signal during training or evaluation. the reasons for this are purely practical; in general, training large, end-to-end goal-driven models is very difficult as it requires the collection of a large amount of task completion data. annotating data in this way on a large scale is extremely expensive, and is usually only feasible for technical support channels at large corporations, which are rarely released publicly. indeed, the ubuntu dialogue corpus has no such labelled task completion data, and thus cannot be analyzed in the task-oriented setting for the time being. obtaining such signals automatically remains an open problem. on the other hand, training non-task oriented dialogue systems such as chatbots only requires conversational data, which can be obtained and shared publicly on a large scale. we believe that significant progress in dialogue systems can be made in this manner, as there remains many unsolved problems as illustrated in section 4.6. 6.3 automatic evaluation of dialogue systems a crucial part of research in building dialogue systems concerns the problem of evaluation. in the goal-oriented setting, when there is a supervised task completion signal available with the data, methods for automatic evaluation are well-established, such as paradise (walker et al., 1997) and memo (möller et al., 2006). an overview of such methods can be found in jokinen and mctear (2009) and hastie (2012). however, in the non-goal oriented setting we consider here, evaluation is more difficult. this is particularly true for end-to-end systems, as there is no way to measure the accuracy of the state tracking module using tasks such as slot filling, since they are not modular systems. indeed, there 58 training end-to-end dialogue systems are several reasons for wanting to move away from the slot filling metrics that have become common for modular systems. in slot filling, the set of candidate outputs (states) is identified a priori through knowledge engineering, and is typically rather small in comparison to the set of responses considered by nuc. further, it has been speculated that state-of-the-art state-tracking models (henderson et al., 2014b; williams, 2014) are achieving close to human-level performance. thus it is desirable to move beyond this domain into more difficult problems. to do this, it is crucial to investigate ways to evaluate models in the non-goal oriented setting that do not require supervised test data for the internal modules of a system. researchers have previously proposed measuring word perplexity and word classification error rate, as these are widely applied in the language modeling and automatic speech recognition community (serban et al., 2016; vinyals and le, 2015; pietquin and hastie, 2013). however, these metrics cannot be computed for retrieval models. researchers have also proposed to use word overlap metrics from machine translation (galley et al., 2015; sordoni et al., 2015b). however, such metrics based on word overlaps suffer from severe sparsity issues, since it is unlikely that any sequence of words will be identical in both the generated and reference responses. these have shown to correlate poorly with human judgements when only a single ground-truth response is available (liu et al., 2016), and they have at best a mediocre correlation when multiple ground-truth responses are available (galley et al., 2015). furthermore, it has been argued that such metrics mainly focus on pronouns and punctuation marks when applied to non-task oriented dialogue datasets (serban et al., 2016). while the word embedding metrics used here do not have a strong correlation with human judgement, they have an additional interpretation of measuring the semantic similarity between the generated and reference responses (as argued in section 4.4), which is why we favour them over word-overlap scores such as bleu. however, we reiterate that none of these metrics measures the coherence of the generated responses, and this remains an important direction for future work. another option for evaluating dialogue systems trained in an end-to-end manner is using an alternative task such as next utterance classification. while this does not directly compare the generated response of the system to the ground-truth response, there are several reasons for preferring the recall metric: 1. it is a more difficult task than slot filling, and thus will require further development of more sophisticated dialogue systems in order to solve the task. 2. it does not suffer from the same problems as the word overlap metrics, as it does not have to directly compare the quality of a generated response to the ground-truth response, an inherently noisy process. instead, it measures the model’s capacity to pick out the correct response from a list of responses. 3. performance using the recall metrics is easily interpretable, and can easily be compared to human performance. indeed, this has been recently done by lowe et al. (2016), who show that human performance on this task is above the performance for the dual encoder model on the ubuntu dialogue corpus, as well as on movie and twitter corpora. thus, there is room for improvement for models on this task. 4. the task is consistent with the end goal of building dialogue systems that can converse naturally with humans. more precisely, models that are able to generate good responses should also be able to pick good responses from a list of candidates, as in nuc. 59 lowe, pow, serban, charlinn, liu and pineau 5. it is easy to alter the task difficulty in a controlled manner. we demonstrated this by moving from 1 to 9 false responses, and by varying the recall@k parameter. in the future, instead of choosing false responses randomly, one could consider selecting false responses that are similar to the actual response (e.g. as measured by tf-idf cosine similarity). a dialogue model that performs well on this more difficult task should also manage to capture a more fine-grained semantic meaning of sentences, as compared to a model that naively picks replies with the most words in common with the context such as tf-idf. in fact, when the set of candidate responses for the model to choose from is close to the size of the dataset (e.g. all utterances ever recorded), then nuc becomes close to the response generation case. given the above points, we believe that evaluating models with the nuc task is very useful for the time being. however, we believe that caution should be used when comparing retrieval models to generative models using nuc, as the retrieval models are directly trained on the task of nuc, rendering it an unfair comparison. 6.4 future research directions for end-to-end systems given the analysis performed in section 4.6, we postulate several interesting directions for future research on end-to-end dialogue systems, particularly on the ubuntu dialogue corpus. an important challenge in dialogue systems is the ability to understand the turn-taking structure of dialogue. this is a significant source of errors for the dual encoder model. some progress in this direction has been made for end-to-end dialogue systems (luan et al., 2016; li et al., 2016a), using approaches derived from topic modelling or by explicitly modelling each user with a continuousvalued vector. however, this is still an open problem. this is related to the issue of end-to-end dialogue personalization, which involves building end-to-end dialogue systems that are tailored to a particular user and that evolve over time as the user’s preferences change. the largest source of errors from the analysis in section 4.6 was in the failure to understand the semantic similarity between the context and response. this falls under the more general problem of natural language understanding, which arises in many nlp tasks. this will require adjustments in the architecture of end-to-end models to render them more suited to processing language. it is possible that insights can be derived from architectures developed on more targeted language understanding tasks, such as the cnn/ daily mail reading comprehension dataset (hermann et al., 2015), where attention-based models have achieved strong performance. in order to be able to correctly answer questions regarding ubuntu and solve the user’s problem, dialogue models will inevitably require some knowledge of the ubuntu domain. this will most likely be achieved by using some source of external knowledge, in addition to the knowledge that is present in the dialogue of the ubuntu dialogue corpus. thus, an important direction for research is the investigation of methods that incorporate external knowledge sources with end-to-end dialogue systems. this applies more generally to any end-to-end system that is developed for the goaloriented setting, and may require imposing additional structure on the output space of the model. there is promising work in this direction from wen et al. (2016), however methods must be derived that are effective in a larger and more general setting than restaurant recommendation. a common problem that has been observed when training generative end-to-end models that maximize the log-likelihood of the conversational response is that these models tend to produce generic responses at test time. this has been observed empirically (vinyals and le, 2015; serban et al., 2016), and was also seen in some of the lstm and hred examples presented in section 60 training end-to-end dialogue systems 5.4. this has been investigated in (li et al., 2015), where the authors construct an objective function based on mutual information that promotes diversity, however they achieve only modest improvements. this is a large impediment for building end-to-end systems that can have interesting and engaging interactions with users. finally, an important direction for future research is building large-scale datasets that allow the training of goal-oriented systems. the ubuntu domain is particularly suited for training goaloriented systems, however this is not yet possible on the ubuntu dialogue corpus as there are no supervised task completion signals, as mentioned in section 6.2. building models that can approximate such signals is challenging, yet it may be necessary in order to develop systems that can solve users’ problems in a meaningful way in a domain as complex as ubuntu. acknowledgments the authors would like to profusely thank rudolf kadlec and martin schmid, who produced the script for generating the updated version of the ubuntu dialogue corpus. we gratefully acknowledge financial support for this work by the samsung advanced institute of technology (sait) and the natural sciences and engineering research council of canada (nserc). we would like to thank michael noseworthy and nicolas angelard-gontier for their input into this paper, and the anonymous reviewers for their helpful feedback. references p. baudiš and j. šedivỳ. sentence pair scoring: towards unified framework for text comprehension. arxiv preprint arxiv:1603.06127, 2016. y. bengio, p. simard, and p. frasconi. learning long-term dependencies with gradient descent is difficult. neural networks, ieee transactions on, 5(2):157–166, 1994. a. bordes, j. weston, and n. usunier. open question answering with weakly supervised embedding models. in proceedings of the meeting on machine learning and knowledge discovery in databases, 2014. j. boyd-graber, b. satinoff, h. he, and h. daume. besting the quiz master: crowdsourcing incremental classification games. in proceedings of the conference on empirical methods in natural language processing, 2012. k. cho, b. van merrienboer, c. gulcehre, f. bougares, h. schwenk, and y. bengio. learning phrase representations using rnn encoder-decoder for statistical machine translation. in proceedings of the conference on empirical methods in natural language processing, 2014. j. deng, w. dong, r. socher, l.j. li, k. li, and l. fei-fei. imagenet: a large-scale hierarchical image database. in proceedings of the conference on computer vision and pattern recognition, 2009. j. dodge, a. gane, x. zhang, a. bordes, s. chopra, a. miller, a. szlam, and j. weston. evaluating prerequisite qualities for learning end-to-end dialog systems. in proceedings of the international conference on learning representations, 2015. 61 lowe, pow, serban, charlinn, liu and pineau j. l. elman. finding structure in time. cognitive science, 14(2):179–211, 1990. m. elsner and e. charniak. you talking to me? a corpus and algorithm for conversation disentanglement. in proceedings of the annual meeting of the association for computational linguistics, 2008. g. forgues, j. pineau, j.m. larchevêque, and r. tremblay. bootstrapping dialog systems with word embeddings. in proceedings of the workshop on modern machine learning and nlp, nips, 2014. m. galley, c. brockett, a. sordoni, y. ji, m. auli, c. quirk, m. mitchell, j. gao, and b. dolan. deltableu: a discriminative metric for generation tasks with intrinsically diverse targets. arxiv preprint arxiv:1506.06863, 2015. j.j. godfrey, e.c. holliman, and j. mcdaniel. switchboard: telephone speech corpus for research and development. in proceedings of the international conference on acoustics, speech and signal processing, 1992. helen hastie. metrics and evaluation of spoken dialogue systems. in data-driven methods for adaptive spoken dialogue systems, pages 131–150. springer, 2012. m. henderson, b. thomson, and j. williams. the second dialog state tracking challenge. in proceedings of the meeting of the special interest group on dialogue and discourse (sigdial), page 263, 2014a. m. henderson, b. thomson, and s. young. word-based dialog state tracking with recurrent neural networks. in proceedings of the meeting of the special interest group on dialogue and discourse (sigdial), 2014b. k. hermann, t. kocisky, e. grefenstette, l. espeholt, w. kay, m. suleyman, and p. blunsom. teaching machines to read and comprehend. in proceedings of the annual conference on neural information processing systems, 2015. s. hochreiter and j. schmidhuber. long short-term memory. neural computation, 9(8):1735–1780, 1997. michimasa inaba and kenichi takahashi. neural utterance ranking model for conversational dialogue systems. in proceedings of the meeting of the special interest group on discourse and dialogue (sigdial), 2016. s. jafarpour, c. burges, and a. ritter. filter, rank, and transfer the knowledge: learning to chat. advances in ranking, 10, 2010. k. jokinen and m. mctear. spoken dialogue systems. morgan claypool, 2009. r. kadlec, m. schmid, and j. kleindienst. improved deep learning baselines for ubuntu corpus dialogs. in proceedings of the workshop on spoken language understanding, nips, 2015. d.p. kingma and j. ba. adam: a method for stochastic optimization. proceedings of the international conference on learning representations, 2014. 62 training end-to-end dialogue systems cheongjae lee, sangkeun jung, seokhwan kim, and gary geunbae lee. example-based dialog modeling for practical multi-domain dialog system. speech communication, 51(5):466–484, 2009. j. li, m. galley, c. brockett, j. gao, and b. dolan. a diversity-promoting objective function for neural conversation models. in proceedings of the meeting of the north american chapter of the association for computational linguistics, 2015. j. li, m. galley, c. brockett, j. gao, and b. dolan. a persona-based neural conversation model. in proceedings of the association for computational linguistics, 2016a. j. li, w. monroe, a. ritter, and d. jurafsky. deep reinforcement learning for dialogue generation. in proceedings of the conference on empirical methods in natural language processing, 2016b. c.-w. liu, r. lowe, i.v. serban, m. noseworthy, l. charlin, and j. pineau. how not to evaluate your dialogue system: an empirical study of unsupervised evaluation metrics for dialogue response generation. in proceedings of the conference on empirical methods in natural language processing, 2016. r. lowe, n. pow, i. serban, and j. pineau. the ubuntu dialogue corpus: a large dataset for research in unstructured multi-turn dialogue systems. in proceedings of the meeting of the special interest group on dialogue and discourse (sigdial), 2015. r. lowe, i. serban, m. noseworthy, l. charlin, and j. pineau. on the evaluation of dialogue systems with next utterance classification. in proceedings of the meeting of the special interest group on dialogue and discourse (sigdial), 2016. y. luan, y. ji, and m. ostendorf. lstm based conversation models. arxiv preprint arxiv:1603.09457, 2016. t. mikolov, m. karafiát, l. burget, j. cernockỳ, and s. khudanpur. recurrent neural network based language model. in proceedings of interspeech, 2010. t. mikolov, i. sutskever, k. chen, g. s. corrado, and j. dean. distributed representations of words and phrases and their compositionality. in proceedings of the annual conference on neural information processing systems, 2013. s. möller, r. englert, k. engelbrecht, v. hafner, a. jameson, a. oulasvirta, a. raake, and n. reithinger. memo: towards automatic usability evaluation of spoken dialogue services by user error simulations. in proceedings of interspeech, 2006. lasguido nio, sakriani sakti, graham neubig, tomoki toda, mirna adriani, and satoshi nakamura. developing non-goal dialog system based on examples of drama television. in natural interaction with robots, knowbots and smartphones, pages 355–361. springer, 2014. j. pennington, r. socher, and c.d. manning. glove: global vectors for word representation. in proceedings of the conference on empirical methods in natural language processing, 2014. o. pietquin and h. hastie. a survey on metrics for the evaluation of user simulations. the knowledge engineering review, 28(01):59–73, 2013. 63 lowe, pow, serban, charlinn, liu and pineau j. ramos. using tf-idf to determine word relevance in document queries. in proceedings of the international conference on machine learning, 2003. a. ritter, c. cherry, and w. dolan. unsupervised modeling of twitter conversations. in proceedings of the meeting of the north american chapter of the association for computational linguistics, 2010. a. ritter, c. cherry, and w. dolan. data-driven response generation in social media. in proceedings of the conference on empirical methods in natural language processing, 2011. v. rus and m. lintean. a comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics. in proceedings of the seventh workshop on building educational applications using nlp, 2012. a.m. saxe, j.l. mcclelland, and s. ganguli. exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arxiv preprint arxiv:1312.6120, 2013. j. schatzmann, k. georgila, and s. young. quantitative evaluation of user simulation techniques for spoken dialogue systems. in proceedings of the meeting of the special interest group on dialogue and discourse (sigdial), 2005. i. v. serban, r. lowe, l. charlin, and j. pineau. a survey of available corpora for building datadriven dialogue systems. arxiv preprint arxiv:1512.05742, 2015. i. v. serban, a. sordoni, y. bengio, a. c. courville, and j. pineau. building end-to-end dialogue systems using generative hierarchical neural network models. in proceedings of the thirtieth aaai conference on artificial intelligence, 2016. l. shang, z. lu, and h. li. neural responding machine for short-text conversation. arxiv preprint arxiv:1503.02364, 2015. s. singh, d. litman, m. kearns, and m. walker. optimizing dialogue management with reinforcement learning: experiments with the njfun system. journal of artificial intelligence research, 16:105–133, 2002. a. sordoni, y. bengio, h. vahabi, c. lioma, jakob grue s., and j. y. nie. a hierarchical recurrent encoder-decoder for generative context-aware query suggestion. in proceedings of the acm international on conference on information and knowledge management, pages 553–562, 2015a. a. sordoni, m. galley, m. auli, c. brockett, y. ji, m. mitchell, j.y. nie, j. gao, and w. dolan. a neural network approach to context-sensitive generation of conversational responses. in proceedings of the meeting of the north american chapter of the association for computational linguistics, 2015b. i. sutskever, o. vinyals, and q. le. sequence to sequence learning with neural networks. in proceedings of the annual conference on neural information processing systems, 2014. d. traum, k. georgila, r. artstein, and a. leuski. evaluating spoken dialogue processing for timeoffset interaction. in proceedings of the meeting of the special interest group on dialogue and discourse (sigdial), 2015. 64 training end-to-end dialogue systems d.c. uthus and d.w. aha. extending word highlighting in multiparticipant chat. in proceedings of the florida artificial intelligence research society conference, 2013a. d.c. uthus and d.w. aha. the ubuntu chat corpus for multiparticipant chat analysis. in proceedings of the aaai spring symposium on analyzing microtext, 2013b. o. vinyals and q. le. a neural conversational model. arxiv preprint arxiv:1506.05869, 2015. m. walker, d. litman, c. kamm, and a. abella. paradise: a framework for evaluating spoken dialogue agents. in proceedings of the meeting of the european chapter of the association for computational linguistics, 1997. h. wang, z. lu, h. li, and e. chen. a dataset for research on short-text conversations. in proceedings of the conference on empirical methods in natural language processing, 2013. m. wen, t.and gasic, n. mrksic, l. rojas-barahona, p. su, s. ultes, d. vandyke, and s. young. a network-based end-to-end trainable task-oriented dialogue system. arxiv preprint arxiv:1604.04562, 2016. j. williams, a. raux, d. ramachandran, and a. black. the dialog state tracking challenge. in proceedings of the meeting of the special interest group on dialogue and discourse (sigdial), 2013. j. williams, a. raux, and m. henderson. the dialog state tracking challenge series: a review. dialogue & discourse, 7(3):4–33, 2016. j. d. williams. web-style ranking and slu combination for dialog state tracking. in proceedings of the meeting of the special interest group on discourse and dialogue (sigdial), 2014. z. xu, b. liu, b. wang, c. sun, and x. wang. incorporating loose-structured knowledge into lstm with recall gate for conversation modeling. arxiv preprint arxiv:1605.05110, 2016. s. j. young. probabilistic methods in spoken–dialogue systems. philosophical transactions of the royal society of london a: mathematical, physical and engineering sciences, 358(1769): 1389–1402, 2000. l. yu, k. m. hermann, p. blunsom, and s. pulman. deep learning for answer sentence selection. arxiv preprint arxiv:1412.1632, 2014. m.d. zeiler. adadelta: an adaptive learning rate method. arxiv preprint arxiv:1212.5701, 2012. tiancheng zhao and maxine eskenazi. towards end-to-end learning for dialog state tracking and management using deep reinforcement learning. in proceedings of the meeting of the special interest group on discourse and dialogue (sigdial), 2016. 65 dialogue and discourse 2(1) (2011) 235–277 doi: 10.5087/dad.2011.110 an incremental model of anaphora and reference resolution based on resource situations massimo poesio university of essex and università di trento hannes rieser universität bielefeld editor: david schlangen abstract notwithstanding conclusive psychological and corpus evidence that at least some aspects of anaphoric and referential interpretation take place incrementally, and the existence of some computational models of incremental reference resolution, many aspects of the linguistics of incremental reference interpretation still have to be better understood. we propose a model of incremental reference interpretation based on loebner’s theory of definiteness and on the theory of anaphoric accessibility via resource situations developed in situation semantics, and show how this model can account for a variety of psychological results about incremental reference interpretation. 1. introduction evidence from both corpora and behavioral experiments suggests that at least some aspects of the interpretation of referring expressions are incremental. for instance, in the following fragment from the trains corpus of dialogues collected at the university of rochester by the trains project (allen et al. 1995),1 the repair in utterance 10.1 is clearly initiated because participant s has started processing the definite description the engine at avon before m’s utterance is complete, and has identified the actual referent of the definite description (engine e1). (1) 9.1 m: so we should 9.2 : move the engine 9.3 : at avon 9.4 : engine e 9.5 : to 10.1 s: engine e1 11.1 m: e1 12.1 s: okay 13.1 m: engine e1 13.2 : to bath 13.3 : to / 13.4 : or 13.5 : we could actually move it 1. the methodology employed in compiling the transcripts, including the methodology used for segmenting utterances into utterance units approximately expressing prosodic boundaries, is described in (gross et al. 1993). c©2011 massimo poesio and hannes rieser submitted 1/2010; accepted 3/2011; published online 5/2011 poesio and rieser to dansville to pick up the boxcar there 14.1 s: okay substantial behavioral evidence from the last fifteen years conclusively supports the intuitions gained by studying such corpus data. in particular, studies using the visual world paradigm have shown that subjects following spoken instructions to manipulate objects in a “visual world” fixate to the relevant objects as soon as the phonetic prefix is completely unambiguous (tanenhaus et al. 1995; eberhard et al. 1995; tanenhaus and trueswell 2005). for example, in a visual world containing an apple and a towel, subjects will fixate on the towel as soon as they hear the first syllable of the word towel. although accounts of incremental interpretation of referring expressions have been developed— e.g., (stoness et al. 2004; schlangen et al. 2009; dubey 2010)—the relation between psychological evidence and existing linguistic theories of anaphora and reference is still poorly understood. in this paper, we propose a theory of the incremental interpretation of referring expressions in terms of a theory of the linguistics of such expressions based, on the one hand, on the theory of definiteness developed by loebner (1987); on the other, on the theory of anaphoric accessibility via resource situations developed in situation semantics (barwise and perry 1983; gawron and peters 1990; poesio 1993; cooper 1996). the paper builds on previous work in the ptt framework (poesio and traum 1997; poesio and rieser 2010), but makes two novel contributions. first, it consolidates into a single proposal a number of ideas about definites, resource situations, and anaphora developed over the years within ptt but never integrated into a coherent whole. secondly, it provides an explicit account of the main psychological evidence about reference interpretation gathered using the visual world paradigm. the structure of the paper is as follows. in section 2 we summarize the main psychological evidence about incrementality and reference. in section 3 we provide a quick summary of the aspects of ptt from (poesio and rieser 2010; poesio to appear) essential for the present proposal. in section 4 we provide a novel and unified account of the semantics and pragmatics of referring expressions. finally, in section 5 we show how the proposal can explain the evidence discussed in section 2. a survey of related literature and a discussion follow. 236 incremental anaphora and reference through resource situations 2. psychological evidence on incrementality and reference in this section we discuss the key evidence about incremental interpretation in general, and in particular the evidence about incremental reference resolution that our theoretical proposals were designed to explain. 2.1 combinatorial explosion and incrementality given the number of phonetically, lexically, syntactically, semantically and pragmatically distinct readings identified by theoretical linguists for most natural language expressions, it is a wonder that such expressions can be understood at all. yet, people appear able to process them rapidly and without apparent effort. the explanation for this puzzle involves a variety of factors, but clearly, one of the key ingredients of the solution is the fact that people appear to process natural language expressions incrementally, immediately making choices about their interpretation before proceeding to the next input segment. the initial evidence about the incremental nature of human language processing came from research on human parsing, in the form of the phenomenon of garden paths observed by bever (1970). garden paths are sentences such as those in (2), which are perfectly grammatical, but which subjects nevertheless find odd because the ambiguity between a reduced relative reading and a matrix verb reading of the verbs raced, floated etc. is immediately resolved in favor of the matrix verb interpretation, thus forcing the reader to a reanalysis step later on. (2) a. the horse raced past the barn fell. b. the boat floated down the river sank. shortly after the initial findings by bever, psychological evidence was found suggesting that semantic interpretation processes are incremental, as well. well-known cross-modal priming experiments suggested that lexical access proceeds by first immediately activating all senses of an ambiguous word-form, and then immediately discarding all but the chosen one (swinney 1979; seidenberg et al. 1982). in these experiments, the subjects were presented with texts such as the one in (3); half of the time a disambiguating context was provided (the string spiders, roaches and other). swinney found priming effects for both ant and spy at [1], even with a strongly disambiguated context; but only for ant at [2]. (3) rumour had it that for years the government building had been plagued with problems. the man was not surprised when he found several (spiders, roaches, and other) bugs [1] in the corner [2] of his room. using similar methods, corbett and chang (1983) found that anaphora resolution, as well, involved the rapid activation of all (genderand number-) matching antecedents of anaphors like she in (4a), as could be verified by cross-modal testing of the activation of the antecedents at [1]. all but one of these however were dropped by point [2] in different-gender conditions, but not in same-gender conditions as in (4b). (4) a. karen poured a drink for bob and then karen, she [1] put [2] the bottle down. b. karen poured a drink for emily and then karen, she [1] put [2] the bottle down. 237 poesio and rieser 2.2 parallelism the original data from bever led to the development of so-called garden-path theory (frazier 1979, 1987) and numerous other incremental models of parsing based on the assumption that interpretations were generated in a serial fashion, one at a time (abney 1991; shieber and johnson 1993; milward 1994). however, the cited results about lexical access could only be explained in terms of parallel processing (marslen-wilson 1975), and indeed they led to the development of the so-called cohort model of lexical access (marslen-wilson 1987). the results by corbett and chang about pronominal interpretation, as well, suggest a parallel model. in recent years, the predominant view has been that parallel processing is the rule in the case of syntax as well (gibson 1991; jurafsky 1996; pearlmutter and mendelsohn 1999); recent evidence on the relation between parsing and lexical disambiguation also suggests a parallel model (macdonald et al. 1994). 2.3 incrementality in reference: the visual world paradigm as discussed above, early evidence that anaphora resolution is incremental was provided by the cross-modal priming experiments by corbett and chang (1983). these results were confirmed and much strengthened by work using the so-called visual world paradigm (tanenhaus et al. 1995; eberhard et al. 1995; arnold et al. 2000; tanenhaus and trueswell 2005). figure 1: the visual world paradigm: ambiguous and unambiguous visual world situations in the visual world paradigm, subjects looking at scenes such as those in figure 1 (this is figure 4 from (tanenhaus et al. 2004)) and wearing head-mounted eye trackers hear instructions such as (5). through the eye trackers it is possible to measure the percentage of fixations to each object in the visual scene millisecond by millisecond; a concentration of fixations on a given object from a certain point on (typically around 400ms after the onset of the target referring expressions) provides very good evidence that that object has been identified as the referent of the expression. (5) put the apple on the towel in the box. eberhard et al. (1995) and allopenna et al. (1998) showed that this concentration of fixations on the referent occurs as soon as an unambiguous prefix has been processed—i.e., in a visual scene in which there is only one object whose name begins with ap-, fixations begin to concentrate on that object as soon as that prefix has been processed, without waiting for the rest of the head noun. in fact, eberhard et al. (1995) showed that in cases in which the referring expression contains unambiguous modifiers, subjects do not even wait until the head noun before beginning to concentrate their attention—i.e., in a situation in which there is only one red object, fixations concentrate on that object 400ms after the onset of red. 238 incremental anaphora and reference through resource situations figure 2: the materials used by arnold et al. to study pronouns with the visual world paradigm the first studies of reference using the visual world paradigm focused on the interpretation of nominals. arnold et al. (2000) extended the use of the paradigm to the study of the interpretation of pronouns. the subjects of arnold et al. listened to two-sentence texts while viewing one of the four pictures in figure 2. the first sentence of the text contained either two same-gender referents (donald / mickey) or two different-gender ones (donald / minnie), where the second sentence contained either a masculine or a feminine pronoun, as in (29). (6) a. donald is bringing some mail to [mickey / minnie] while a violent storm is beginning. b. he’s / she’s carrying an umbrella, and it looks like they’re both going to need it. arnold et al. found both a gender and a first-mention effect. in the different gender contexts, fixations would concentrate on the unambiguous referent of the pronoun already after 400ms, and so they would in the same gender context when reference was to the first mention entity. in the same gender, second mention reference context, however, the percentage of fixations on the first and second mentioned entity was the same. finally, a series of visual world paradigm experiments including, among others, (altmann and kamide 1999; chambers et al. 2002; brown-schmidt et al. 2005) found substantial empirical evidence for the focus shift principles proposed in (grosz 1977; poesio 1993) on the basis of the analysis of data from task-oriented dialogues, and later studied by beun and cremers (1998). these 239 poesio and rieser focus shifting effects have now become known as effects of referential domain restriction (brownschmidt et al. 2005). chambers et al. (2002) found that after hearing instruction (7) in a visual scene containing a number of containers some of which are big enough to fit the cube whereas others aren’t, the subjects’ attention quickly concentrates on the containers into which the cube can fit. this type of referential domain restriction / focus shifting effect is now known as task compatibility. (7) pick up the cube. put it in . . . the experiments by chambers et al. took place in controlled experimental situations. the experiments discussed in (brown-schmidt et al. 2005; brown-schmidt and tanenhaus 2008), by contrast, involved subjects performing tasks in fairly ecologically valid situations; but these studies, as well, found focusing effects—in particular, effects both of task compatibility in the sense of chambers et al. and proximity (greater salience of closest objects). 2.4 interaction between reference and parsing crain and steedman (1985) and altmann and steedman (1988) observed that many classical garden path sentences such as (2) or (5) involve an ambiguity between a reading in which a constituent (raced past the barn, on the towel) is interpreted as a modifier of a definite np and a second reading in which it is interpreted as part of the main clause. they also observed that the fact that this second reading is initially preferred—thus originating the garden path—might be due to the lack of a second object in the context (a second horse, or a second apple) that would justify the use of the modification; and hypothesized that the garden path effect might be reduced, or eliminated, in contexts in which this object is present. the visual world paradigm offered the opportunity for a very convincing verification of this hypothesis, reported in tanenhaus et al. (1995) and spivey et al. (2002). the subjects in these experiments were presented either with the visual context on the left in figure 1, in which only one apple is present, or with the context on the right, in which there are two. they then heard the instruction in (5) on page 230. a much greater proportion of fixations on the incorrect destination (the towel) was observed in the situation in which only one apple was present. 240 incremental anaphora and reference through resource situations 3. a short introduction to ptt ptt (poesio and traum 1997; poesio and muskens 1997; matheson et al. 2000; poesio and rieser 2010) is a theory of dialogue semantics and dialogue interpretation developed to explain how utterances are incrementally interpreted in dialogue, considering both their semantic impact and their impact on aspects of dialogue interaction traditionally considered as outside the scope of semantic theory. much like sdrt (asher and lascarides 2003), ptt is a dynamic theory of language interpretation based on drt (kamp and reyle 1993), hence designed to formalize the linguistics of anaphora and reference; but it incorporates ideas about conversation and the construction of the common ground from the work of clark (1996) and from situation semantics (barwise and perry 1983; cooper 1996; ginzburg to appear). in this section we briefly introduce the aspects of the theory that are relevant for our discussion of incremental interpretation in dialogue; in the next section we will discuss specifically reference and anaphoric interpretation. for more details on ptt, including a complete fragment for german, see (poesio and rieser 2010). 3.1 compositional drt ptt is implemented in compositional drt, a reconstruction due to muskens (1996) of drt in terms of a standard type logic to which two new types have been added: the type of discourse referents π and the type of states s. discourse referents are used to model the dynamics of context in the same way as they are used in drt, i.e., in the sense that each noun phrase introduces a new discourse referent. states are used to model contexts themselves, and the way they are modified by natural language sentences; they are the object-language equivalent of the assignments used to formalize the semantics of drss in (kamp and reyle 1993). this dynamics is mediated by drss, which in muskens’ type logic are relations between states. a function v : π → (s → e) provides the mapping from discourse referents and states to entities, in the sense that v(x)(i) specifies the ‘value’ of discourse referent x at state i. muskens specifies a translation for all drt constructs in terms of this type logic. the most important translations are for drs conditions—general relation and equality predicate is —drss, and drs composition, as follows: (8) a. r{x1 . . . xn} is short for λi.r(v(x1)(i), . . . v(xn)(i)) b. x1 is x2 is short for λi.v(x1)(i) = v(x2)(i) c. [x1 . . . xn|φ1 . . . φm] is defined as λiλj(i[x1...xn]j ∧ φ1(j) . . . ∧ φm(j)) where i[x1...xn]j is short for i and j differ at most over [x1 . . . xn]. d. k;k’ is defined as λiλj(∃kk(i)(k) ∧k ′(k)(j)) for example, the type-logic translation of the drs in (9a) is shown in (9b). (9) a. [x,w, y, z, s, s′|engine(x),avon(w), s : at(x,w),boxcar(y), s′ : hooked-to(z, y), z is x] b. λiλji[x,w, y, z, s, s′]j∧[engine(x)](j)∧[avon(w)](j)∧[s : at(x,w)](j)∧[boxcar(y)](j)∧ [s′ : hooked-to(z, y)](j) ∧ v(z)(j) = v(x)(j) 3.2 the discourse situation ptt is an information state theory of dialogue (larsson and traum 2000; stone 2004; ginzburg 2011) in which the participants in a conversation maintain an information state about the conver241 poesio and rieser sation consisting of private information together with a conversational score including ‘grounded’ (clark 1996) and semi-public information. in ptt, as in situation semantics, the conversational score consists of a record of all actions performed during the conversation, i.e., what in situation semantics is called the discourse situation (barwise and perry 1983; ginzburg to appear). according to this view, the common ground in an ordinary conversation does not consist only of the content of assertions, but it is a general record of actions the actions that were performed, including actions whose function is to acquire, keep, or release a turn, to signal how the current utterance relates to what has been said before, or to acknowledge what has just been uttered. (bunt (1995) called these actions dialogue control acts.) the discourse situation also contains information about non-verbal actions such as pointing. poesio and traum (1997) argued that the discourse situation-oriented view of the conversational score from situation semantics could be formalized using the tools already introduced in drt (kamp and reyle 1993)—specifically, in muskens’s compositional drt 1996. speech acts— conversational events, in ptt terms—and non verbal actions are treated just like any other event; conversational events and their propositional contents can serve as the antecedents of anaphoric expressions. for instance, poesio and rieser (2010) hypothesize that the two directives in (10) (an edited version of two turns from the bielefeld toyplane corpus) result in the update to the common ground in (11).2 (10) inst: so jetzt nimmst du eine orangene schraube mit einem schlitz so now you take a orange screw with a slit cnst: ja ok inst: und steckst sie dadurch, von oben, daß also die drei festgeschraubt werden dann and you put it through from above so that the three get fixed (11) [ k1.1, up1.1, ce1.1,k2.1, up2.1, ce2.1| k1.1 is [e, x, x′|screw(x), orange(x), slit(x′),has(x, x′), e : grasp(cnst, x)], up1.1 : utter(inst,”so jetzt nimmst du ... ”), sem(up1.1) is k1.1, ce1.1 : directive(inst&cnst,cnst,k1.1) generate(up1.1, ce1.1) k2.1 is [z, e′, s, w, y|z is x, e′ : put-through(cnst, z, hole1), w is wing1, y is fuselage1, s : fastened(w, y)], up2.1 : utter(inst,”und steckst sie ... ”), sem(up2.1) is k2.1, ce2.1 : directive(inst,cnst,k2.1) generate(up2.1, ce2.1)] 2. we name discourse referents as follows: names with a k prefix for drss, with a ce prefix for conversational events, with a u prefix for utterances, with a up prefix for phrasal utterances. as far as content is concerned, we use an e prefix for discourse referents denoting events and the last letters of the alphabet x, x′, y, y′, z, z′, w, w′ for other types of discourse referents. 242 incremental anaphora and reference through resource situations (11) records the occurrence of two conversational events, ce1.1 and ce1.2, both of type directive (matheson et al. 2000) whose propositional contents are separate drss specifying the interpretation of the two utterances in (10). the contents of conversational events are associated with propositional discourse referents (discourse referents whose values are drss) k1.1 and k2.1, as in (poesio and muskens 1997) and in a number of other theories of the common ground, most notably sdrt (asher and lascarides 2003). it is further assumed in ptt that dialogue acts are generated (pollack 1986) by locutionary acts (austin 1962) which we represent here as events of type utter. non-verbal actions are also viewed in ptt as conversational events, albeit of a different type. so for instance an act of pointing by agent dg would lead to the following update of both agents’ information state: (12) [pe1.1|pe1.1 : point(dg,α)] where α is what dg is pointing at. (determining experimentally what is α is the main question addressed by (lücking et al. to appear), as discussed in section 4.4.) the extent to which speakers take the common ground into account while referring has been challenged by studies such as (horton and keysar 1996). such studies suggest the need to develop a more nuanced theory of the information state maintained by speakers and how it affects conversational behavior than those developed in response to the original work by clark and colleagues (clark and marshall 1981). and indeed, one of the key differences between ptt and other theories of the common ground that build on drt is that ptt includes an explicit account of the process by which information becomes part of the common ground, based on the theory of grounding proposed by traum (1994) and on a theory of the information state in conversations developed in (poesio and traum 1997; matheson et al. 2000; poesio and rieser 2010) according to which the information state involves a combination of public, semi-public and private information. in this paper however, for reasons of simplicity, we will omit any discussion of the information state and grounding, and ignore the interaction between incremental interpretation and grounding; see (poesio and rieser 2010) for a discussion. 3.3 the ingredients of incrementality, i: micro conversational events it is assumed in ptt that the conversational score is incrementally updated whenever a verbal or non-verbal event is perceived (poesio 1995a). in particular, each word incrementally updates the discourse situation with a locutionary act of type utter and with syntactic expectations about the occurrence of more complex utterances as hypothesized in lexicalized tree adjoining grammar (ltag) (schabes 1988; abeille and rambow 2000), that lends itself to a very natural account of the process by which syntactic interpretations are constructed incrementally (sturt and crocker 1996). for instance, an utterance of definite article the results in the conversational score being updated with the occurrence of an utterance udet of syntactic category det (a micro conversational event (mce) (poesio 1995a)) and with the expectation that this utterance will be part of an utterance of an np which will also include an utterance un ′ of syntactic category n′. mces are characterized by lexical, syntactic, semantic and discourse information in the form of features.3 one type of syntactic information about mces is syntactic constituency; every mce 3. there is a clear relation between mces in ptt and signs in theories such as hpsg, see (poesio to appear) for discussion. 243 poesio and rieser u that is not the root of a tree has a mother node u′. we indicate this with the notation used by muskens (2001) to indicate direct subconstituency in his logic of trees: u ↑ u′ we will assume here the additional features of mces in table 1. cat specifying the syntactic category of a mce gen specifying the gender of mces num specifying the number sem specifying its (conventional) semantics do specifying the discourse referent introduced by the np table 1: features of mces. the lexical semantics of words that update the discourse model and of anaphoric expressions is as proposed in compositional drt (muskens 1996), according to the grammar fragment discussed in (poesio and rieser 2010). the sem value of phrasal utterances is obtained compositionally via defeasible inference rules that by default assign, for instance, to an utterance of an np like unp above the conventional semantics sem(unp ) resulting from the application of sem(udet) to sem(un ′), but that can be overridden e.g., in the case of metonymy or as in the case of anaphoric expressions, as discussed below (poesio and traum 1997; poesio to appear; poesio and rieser 2010). we will mostly encode the information associated with mces in the compact format illustrated by the following example, representing the update resulting from observing an utterance of determiner the and by the following lexical access. (13) [ udet, unp , un ′ | unp :np xxxxxx ������ udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un ′ :n’ ] 3.4 the ingredients of incrementality, ii: defeasible reasoning the evidence on incremental interpretation and parallel hypothesis generation discussed in section 2 suggests that utterance interpretation is a form of defeasible inference in which competing hypotheses are activated, one of which is rapidly selected, whereas the other ones are discarded (poesio 1994, 1995a,b, to appear). this view is also taken in sdrt for aspects of interpretation such as intention recognition or anaphora resolution; in ptt it is assumed that all aspects of utterance interpretation are defeasible, from lexical access to parsing and semantic composition, as already assumed in hobbs’ interpretation as abduction framework (hobbs et al. 1993) and in the great majority of recent computational linguistics work on disambiguation (hwang and schubert 1993; jurafsky 1996; asher and lascarides 2003; bod et al. 2003; jurafsky and martin 2009). a good case can be (and has been) made that the defeasible inferences involved in language interpretation are a form of statistical inference (macdonald et al. 1994; jurafsky 1996), and most recent theories of interpretation in computational linguistics are of this type. however, it is still an open problem 244 incremental anaphora and reference through resource situations how to combine the logics used in formal semantics with statistical inference,4 so ptt follows the more traditional approach adopted by virtually all theoretical approaches attempting to combine a theory of performance with a theory of semantic competence based on formal semantics in using a form of logic to model defeasible inference (perrault 1990; hobbs et al. 1993; hwang and schubert 1993; poesio 1994, 1995b; lascarides and asher 1991; asher and lascarides 2003). specifically, in ptt interpretation is modeled in terms of prioritized default logic (pdl) (brewka and eiter 2000; horty 2007). for instance, lexical access is modelled in ptt as a default theory—a set of defeasible inference rules (specifically, prioritized default rules) specifying the alternative lexical interpretations accessed by encountering an utterance of a given phonetic form. these alternative lexical interpretations are alternative hypotheses about how to update the discourse situation upon hearing that utterance, where each update adds to the discourse situation the information exemplified by (13)— that is, the lexical and ltag information about that use of the word form. such hypotheses about updates are produced by (normal) lexical default rule like the rule lex-the below specifying one of the lexical interpretations of english definite article the. lex-the states that if an utterance of the was observed, and it is consistent to hypothesize that this utterance is to be interpreted as the utterance of the determiner of an np (we will get to the semantics in a moment), then do so.5 u : utter(a, “the”) : [ unp , un′ | unp : np hhh ��� u : det the un ′ : n ′ ] lex-the [ unp , un′ | unp : np hhh ��� u : det the un ′ : n ′ ] homonyms like stock or bank are associated with multiple such defaults; when these words are encountered, all the extensions of the default theory specifying the current information state (i.e., all the inferential closures of the theory obtained using defaults which are consistent) are immediately computed, and if one has a higher priority than the others, that interpretation is chosen and the others remove; else an ambiguity is detected (poesio 1996, to appear). we will at times use the standard simplified notation for normal defaults: lex-the u : utter(a, “the”) ⇒ [ unp , un′ | unp : np hhh ��� u : det the un ′ : n ′ ] 4. for preliminary work on the matter, see, e.g., (hwang and schubert 1993). more recently, markov logic networks have been proposed for this purpose (richardson and domingos 2006). 5. most of the defaults discussed in this paper are open defaults—i.e., default schemata. we use capital letters to indicate the variables in the open default—in this example, u , unp , and un′ are all variables. 245 poesio and rieser the ptt view of the interpretive processes that follow lexical access such as syntactic interpretation (parsing) is very much inspired by work in grammatical frameworks like categorial grammar (pereira 1990; carpenter 1998) in that syntactic interpretation is also viewed as an inferential process. parsing in ptt is a process during which hypotheses about the results of lexical access combine together in phrasal hypotheses through default inference. such phrasal hypotheses are viewed as hypotheses about utterances of phrases: e.g., the occurrence of contiguous utterances of syntactic category det and n results in a phrasal hypothesis about the occurrence of an utterance of category np. these hypotheses are the result of a second set of defeasible inference rules that encode syntactic competence. ptt is more unusual in proposing that semantic composition, as well, is a form of defeasible reasoning: i.e., that the semantic value sem of utterances corresponding to non-terminal nodes like unp in the example of the is specified by default rules which may compete with other defaults (poesio to appear). the original motivation for this hypothesis are data about metonymy, and in particular the theory proposed by nunberg (1995). nunberg identifies two types of metonymy: deferred indexical reference and predicate transfer. we are particularly interested in the second of these, illustrated by the utterance in (14), to be imagined uttered by a customer handing his key to an attendant at a parking lot: (14) i am parked out back. nunberg argues that in this example, am parked out back is not interpreted as denoting the predicate that holds of objects that are parked out back, but a predicate that applies to human beings whose car is parked out back. i.e., that two different hypotheses about the interpretation of am parked out back compete. (15) λy.(∀x[car-of(y) = x]→parked-out-back(x)). nunberg argues that predicate transfer is ‘. . . a phrasal phenomenon that works in concert with the process of semantic composition’ and is subject to the same constraints; e.g., composition has to apply in a certain order. his description of semantic composition makes it sound very much like a defeasible reasoning process: “. . . one way of dealing with [these cases] would be to permit transfer to take place independently on any simple or complex predicate or term, and then filter the output via constraints charged with maintaining consistency . . . ” (p. 121) the conclusion drawn in (poesio to appear) is that semantic composition rules, as well, are prioritized defaults. nunberg’s observations can be explained by hypothesizing that at least two defaults apply in this case to derive the sem value of the vp node from the meanings of its constituents: a low-priority one, binary semantic composition (bsc), assigning to a constituent a meaning on the basis of the meaning of its constituents, and a higher-priority one, pt-bin-sem-comp, that applies whenever there is a predicate transfer function g mapping φ into a predicate υ (e.g., g could be the transfer function mapping predicates like parked-out-back into predicates like (15)). we won’t discuss here pt-bin-sem-comp (see (poesio to appear) for details), only the latest version of bsc, proposed in (poesio and rieser 2010), which is a direct implementation of type-driven semantic composition. the default specifies that if u1, u2 and u3 are utterances, u1 and 246 incremental anaphora and reference through resource situations u2 are direct constituents of u3, and the semantic value of u1 is a function taking as values objects of the type of the semantic value of u2, then one hypothesis about the semantic value sem(u3) of u3 is that it results from the application of the semantic value of u1 to the semantic value of u2. u1 ↑ u3, u2 ↑ u3, sem(u1) is φ〈α,β〉, sem(u2) is ψα : sem(u3) is φ(ψ) bsc sem(u3) is φ(ψ) (poesio to appear) also postulates an additional default for percolating the meaning up in the case of nodes with a single constituent, unary semantic composition (usc).6 u1 ↑ u2, sem(u1) is φα : sem(u2) is φ usc sem(u2) is φ we will show in the rest of the paper that incremental interpretation provides further evidence for the hypothesis that semantic composition is defeasible: specifically, we will see that the defaults that produce hypotheses about the interpretation of referring expressions, called principles for anchoring resource situations, are in fact semantic composition defaults. 3.5 the ingredients of incrementality, iii: parallelism and pruning as discussed in section 2, the view is taking hold that incremental processing should be explained not in terms of serial interpretation as in frazier’s garden path model, but in terms of parallel models in which alternative hypotheses are generated in parallel. these hypotheses are sometimes only entertained very briefly before pruning (as in the cases of lexical access first studied by (swinney 1979)); in others, these hypotheses survive until the end of the sentence (as in the cases of pronoun interpretation studied by (corbett and chang 1983)).7 this process of generating multiple hypotheses in parallel is naturally modelled in terms of extension generation over the discourse situation. language interpretation is initiated when the occurrence of a new utterance u is recorded in the information state. at this initial stage, the interpretation of u is h-underspecified in the sense of (poesio to appear)—i.e., the discourse situation does not specify the value of sem(u), or its syntactic properties. we can formally characterize the state of the language processor after observing u in terms of the extensions of a default theory generated by prioritized rules like the ones we have discussed. what is still missing to have a complete account of results such as swinney’s is an explanation of the second crucial ingredient of parallel search theories, pruning: i.e., how the language processor decides which extensions to keep and which ones to throw away. the ptt account of this is quite simple: at the end of each process of extension generation according to the currently active set of pdl rules, only the extensions with highest priority survive; the other ones get pruned. if this hypothesis is correct, at the end of each round of hypothesis generation the processor may find itself in one of two situations. if there is only one remaining extension, the 6. the actual formulation of the defaults is slightly more complex due to the need to ensure that u1 is the only child of u2. we assume binary trees only (poesio 1994) 7. in other cases yet, multiple hypotheses about the interpretation survive even after end-of-sentence processing: this is what happens in cases of deliberate ambiguity, which is fairly common both in political language and in poetry. we called this situation perceived ambiguity in (poesio 1996). 247 poesio and rieser processor commits itself to that hypothesis,8 as in the simplest cases of lexical access in which all hypotheses but the one with highest priority get pruned. this single extension may then represent either a fully specified interpretation, an h-underspecified interpretation, or a p-underspecified interpretation (poesio to appear; poesio et al. 2006)—an interpretation which is underspecified in the sense that a more general sense for the word is chosen, as in the cases of lexical polysemy discussed by (frazier and rayner 1990). however, it’s also possible that more than one extension remains, because more than one conflicting default inference rule with the same priority was activated. at this point different things may happen. the results by (corbett and chang 1983) indicate that in some cases of pronoun resolution the conflicting extensions are kept around until the end of the sentence, but then all but one are pruned at that point. 9 4. resource situations, anaphora, and reference in this section we develop the new treatment of definites and anaphoric expressions in ptt proposed in (poesio and rieser 2009), to account for the data about incremental reference interpretation. there are two distinctive aspects in this proposal with respect to the standard treatments of anaphora and reference proposed in drt and sdrt. the first novel aspect is the adoption of the ‘functional’ interpretation of definite nps due to loebner (1987), obtaining a picture of the interpretation of anaphoric expressions with many points in common with the treatment proposed e.g., in (chierchia 1995). the second is the idea of accessibility via resource situations from situation semantics, leading to a unified treatment of anaphoric and deictic reference. 4.1 a reconstruction of loebner’s theory of definiteness according to loebner, what all definites have in common is that they are terms, i.e., functions (in the sense that a skolem function is a function) that may take a different number of arguments, but all have a value of type e. thus, a definite the p is licensed either because predicate p is semantically functional, as in classical examples like the king of france, or because p is turned into a function by a modifier, as in the first point to make is that..., or because p is pragmatically coerced into a function by resolving it. translated into standard logics,10 the idea is that proper noun jack denotes the (0-argument) function ιx.(x = j), whereas the definite description the pope would denote the 1-argument function λsιx.(x = pope(s)(x)), 8. this view that committing to an hypothesis is not simply a matter of computing the extensions of a defeasible theory, loosely inspired by kyburg’s work on acceptance (e.g., (kyburg 1974)) and pollock’s work on a cognitive architecture based on defeasible reasoning (pollock 1990) clearly requires development, but embedding defeasible jumping to conclusions in a more explicit account of belief maintenance and revision will be required anyway to account for the process by which interpretations are reanalyzed and repaired, a major issue for theories of incremental interpretation. for a probabilistic account of such processes, see (jurafsky 1996; schlangen et al. 2009). 9. finally, in the cases of perceived ambiguity, the conflicting extensions stay around even after the end of the sentence. in other words, a different sort of pruning seems to take place after the first round of hypothesis generation; this second phase of pruning eliminates some interpretations in the corbett and chang cases, but not in the case of perceived ambiguity. 10. loebner only provided an informal discussion of his theory. 248 incremental anaphora and reference through resource situations taking a situational or temporal argument s. as just sketched, loebner’s theory would not account for the dynamic properties of definites. the first aspect of our proposal is to combine loebner’s proposals with the treatment of definites in drt, allowing definites such as jack or the chair to update context. this is done by assigning to jack the cdrt semantics jack⇒ λp.([y|y is ιx.[|x is j]];p (y)) i.e., the set of properties of discourse referent y, where y is the unique object that is equal to constant term j denoting jack ( is is the equality condition), whereas definite descriptions like the chair translate as follows: (16) the chair⇒ λp.([y|y is ιx.[|chair(x)]];p (y)) this treatment of definites is implemented in ptt by hypothesizing that a definite article (e.g., english the) results in the update to the information state in (17), which combines an ltag-style prediction of an elementary tree with the cdrt semantics just discussed. the update to the discourse situation in (17) specifies that an utterance udet of the word the has been observed, and that as a result of lexical access, this utterance has been hypothesized to be part of the realization of utterance unp of an np part of which has not yet been observed, and has been assigned a sem value of λpλp ′([y|y is ιxp (x)];p ′(y)). (17) [ udet, unp , un ′ | unp :np xxxxxx ������ udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un ′ :n’ ] proper names and pronouns update the discourse situation by adding to it a record of the utterance of a complete np. the update resulting from proper name jack is as in (18). we’ll discuss pronouns in section 4.5. (18) [ upn , unp | unp :np upn :pn jack: λp [y|y is ιx[|x is j]];p (y) ] 4.2 resource situations in traditional formal semantics, a sharp distinction is made between anaphoric and referential interpretations of expressions such as demonstratives. in an utterance of the sentence this chair was hand-made by an artisan accompanied by a pointing gesture to a chair (the demonstration), this chair is interpreted as direct reference to the chair. by contrast, in the sentence hannes bought a chair in the centre of rovereto. this chair was hand-made by an artisan, this chair is anaphoric. according to kaplan (1978), this contrast indicates that demonstrative this chair is semantically ambiguous. the claim that demonstratives are ambiguous has been challenged by barwise and perry (1983) and, more recently, by gundel et al. (1993) in corpus linguistics and by semanticists such as roberts 249 poesio and rieser (2002). barwise and perry proposed that referring expressions like this chair in the example above are not ambigous, but depend for their intepretation on a resource situation: a situation (in the sense of (barwise 1989)) containing the object in question that may or may not coincide with the described situation (see also (ginzburg to appear)). in the case of the chair being used deictically, the resource situation is the visible situation; when it is used anaphorically, it is the described situation. the demonstration is a cue to which resource situation should be used. this proposal was developed in (gawron and peters 1990; poesio 1993; cooper 1996; poesio and muskens 1997). poesio (1993) proposed a theory of resource situation identification based on prioritized default rules called principles for anchoring resource situations, subsequently revised in (poesio 1994). one of the proposed principles, pars1, produces an hypothesis that (parts of) the visual scene s are a possible choice of resource situation when they have been made salient (e.g., as the result of instructions that direct the attention to those parts of the scene). pars1 if a speaker uses a referring expression the p, the speaker intends the mutual attention of the conversational participants to be focused on the situation s, and the visible situation contains an object of type p, then the listener may hypothesize that s is the resource situation for the p. a second principle, pars2, makes anaphoric reference possible, licensing the choice of the current discourse situation s described by core speech act (csa, (traum and hinkelman 1992)) c as a possible resource situation for a definite reference the p whenever an object of type p has been mentioned.11 pars2 if the current discourse topic is the situation or situation kind s that includes a discourse marker z of type p, a definite np of the form the p may be taken to refer to z. the proposal in (poesio 1993) was formulated in terms of episodic logic, a logic with situations (hwang and schubert 1993). poesio and muskens (1997) recast the resource situation proposal in terms of (compositional) drt. they proposed that resource situations are contexts—drss— and that all anaphoric expressions contain an implicit variable over contexts; it is this variable that supplies the value for the discourse referent. if we combine these ideas about resource situations with the loebnerian semantic for definites introduced previously we obtain the single semantic interpretation for definite the chair in (19). (19) the chair⇒ λp ′.([y|y is ιx.k; [|chair(x)]];p ′(y)) according to the semantics in (19), definite the chair gets a presuppositional interpretation requiring the identification of a resource situation k in the context in which an object of type chair is particularly salient. note that k is used presuppositionally—i.e., it is a free variable, just like context variables in rooth’s analysis of focus-sensitive particles (rooth 1992). the definite gets a deictic interpretation when k is identified with a context specifying (parts of) the visual scene; an anaphoric one when k gets identified with the content of a previous speech act. within this new framework, the principles for anchoring resource situations proposed in earlier ptt work can be reformulated as coercion rules: semantic composition rules alternative to the 11. the original formulation of the principle was more general allowing the formulation of such hypotheses also when the previously mentioned object was of type p′ with p′ a lexical prime of p (either a synonym or a hyponym) but we will simplify matters here. 250 incremental anaphora and reference through resource situations default ones discussed earlier, bsc and usc, that take a non-functional nominal predicate p (e.g., chair) and turn it into a presuppositional predicate λx.k; [|p(x)] that is pragmatically functional in the sense of loebner wrt a resource situation k. for instance, such a coercion rule turns the np interpretation in (17) into the following one: (20) unp :nphhhhhhh ((((((( udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un ′ : n ′ : λxk; [|p(x)]: un : n : p we will consider pars1 first, and discuss pars2 in the next section. pars1 states that the presence of an object z of type p in a situation in mutual visual attention kmva is grounds to hypothesize that kmva is the resource situation of a definite description the p and z is the referent of the definite description, i.e., to coerce the interpretation of nominal predicate p to the following predicate which clearly is functional in that it only is true of z: λxkmva; [|p(x)]; [|x is z] this is implemented by formulating pars1 as a default rule (of higher priority than usc seen before) proposing an alternative specification of the semantic value of the un ′ utterance—one in which kmva occurs as resource situation. in the formulation of pars1 below we use a simpler linear notation for representing syntax trees, omitting utterance names where there is no risk of confusion. we also use the notationk |= φ, for k a drs and φ a condition, to indicate that condition k entails φ in the sense that all pairs of assignments i, j that verify k must also verify φ. ∀i, jk(i)(j)→ φ(i)(j) finally, we adopt a very simple formalization of the notion of mutual visual attention— hypothesizing a distinguished variable msoa specifying the current mutual situation of attention (see (grosz 1977); see also (poesio 1993) for discussion), whose value is constantly updated as mutual attention shifts as the effect of focus shift principles (see section 5). k being the value of msoa implies mutual belief that k is mutually seen:12 msoa is kmva → bela,b(a,b, seea,b(a,b,k)) with this notation pars1 is as follows. an utterance of definite the p in a discourse situation in which msoa is kmva and kmva contains an object z of type p leads to hypothesize that z may be the intended referent of the p if it is consistent to assume so. pars1 [unp : np [det the: λqλq′([y|y is ιxq(x)];q′(y))] [un′ : n ′ [n : λx[|p(x)]]]] msoa is kmva kmva |= p(z) ⇒ [unp : np [det the: λqλq′([y|y is ιxq(x)];q′(y))] [un′ : n ′ : λxkmva; [|p(x)]; [|x is z] [n : λx[|p(x)]]]] 12. as mentioned earlier, please refer to (poesio and rieser 2010) for a discussion of the information state in ptt and the modalities that qualify its different aspects. 251 poesio and rieser notice that the uniqueness requirement on z, proposed in (poesio 1994), has been dropped. this is because in case kmva contains more than one object of type p multiple competing hypotheses— i.e., multiple extensions of a default theory with the same priority—would be generated, and as a result, the processor cannot commit to any of them in the sense discussed in section 3.5. this prediction is confirmed by the results of eye-tracking experiments in which multiple objects in the visual situation are briefly considered and maintained until a single interpretation can be obtained (tanenhaus and trueswell 2005). notice also that the formulation of pars1 above only works, stricly speaking, for definite descriptions. this is in keeping with the assumption that distinct interpretation processes apply to each type of referring expression, widely shared among linguists (gundel et al. 1993), psycholinguists (garrod 1994) and computational linguists (sidner 1979; passonneau 1993; hoste 2005; poesio and kabadjov 2004). we’ll make the simplifying assumption in this paper that demonstratives and definite descriptions have the same semantics, and they only differ in that the pars3 principle governing interpretation via pointing proposed by poesio and rieser (2009) and discussed later in this section only applies to demonstratives, whereas versions of both pars1 and pars2 apply to both (modulo the triggering condition). the semantics we propose for pronouns however is different, as are the interpretation principles; we’ll get back to pronouns after discussing anaphoric accessibility. 4.3 anaphoric accessibility via resource situations before discussing pars2 we need to address two apparent problems with the account of incremental reference in discourse situations introduced so far. the first of these problems is an issue for all theories of anaphoric interpretation that do not make the simplifying assumption that discourse structure is completely flat. under the anaphoric accessibility rules of drt, one would conclude that viewing the common ground as a discourse situation leads to the prediction that anaphoric antecedents introduced in previous core speech acts are not accessible, because they are in the scope of the (intentional) operators. thus for example the antecedent for the pronoun sie in (10) (repeated below for convenience), the orange screw with a slit introduced in the first utterance, would be expected to be inaccessible under the view of the common ground in (11), as the screw would be included in proposition k1.1 (the content of the first speech act ce1.1) whereas the pronoun would be part of drs k2.1 (the content of the second speech act, ce2.1). (10) inst: so jetzt nimmst du eine orangene schraube mit einem schlitz so now you take a orange screw with a slit cnst: ja ok inst: und steckst sie dadurch, von oben, daß also die drei festgeschraubt werden dann and you put it through from above so that the three get fixed (11) [ k1.1, up1.1, ce1.1,k2.1, up2.1, ce2.1| k1.1 is [e, x, x′|screw(x), orange(x), slit(x′),has(x, x′), e : grasp(cnst, x)], up1.1 : utter(inst,”so jetzt nimmst du ... ”), sem(up1.1) is k1.1, 252 incremental anaphora and reference through resource situations ce1.1 : directive(inst&cnst,cnst,k1.1) generate(up1.1, ce1.1) k2.1 is [z, e′, s, w, y|z is x, e′ : put-through(cnst, z, hole1), w is wing1, y is fuselage1, s : fastened(w, y)], up2.1 : utter(inst,”und steckst sie ... ”), sem(up2.1) is k2.1, ce2.1 : directive(inst,cnst,k2.1) generate(up2.1, ce2.1)] but as we said, an explanation for this apparent problem has been available for many years. as argued by reichman (1985); grosz and sidner (1986); webber (1991); asher and lascarides (2003), and others, accessibility in dialogue depends on discourse structure: the antecedents which are accessible are those introduced by utterances belonging to the same discourse segment. in the formulation of grosz and sidner, discourse structure depends on intentional structure: utterance u1 belongs to the same discourse segment as utterance u2 if the discourse intention of u2 satisfactionprecedes the discourse intention of u1, whereas it belongs to a subordinate segment if its discourse intention is dominated by the discourse intention of u2. this account was developed most extensively in sdrt (asher and lascarides 2003), in which the veridicality axiom ensures that proposition k1 provides the context for the interpretation of proposition k2 whenever the speech act with content k1 is related by one of a small number of discourse relations to the speech act with content k2. in (poesio and traum 1997), the effect of intentional structure on accessibility was also explained in terms of axioms similar to veridicality, but the formulation of the semantics of definites proposed in this paper, which ‘brings the context in’ in the form of the resource situation, suggests a solution that does not involve such axioms. we propose instead an explanation for anaphoric accessibility based on a new formulation of the principle governing anaphoric resolution of resource situations, pars2. this new version of pars2 is as follows. let the referring expression that is to be interpreted, unp , be a constituent of utterance u , and let u generate a core speech act c ′, jointly performed by conversational participants a and b. let c be a core speech act also jointly performed by a and b with content kdt —we indicate this using the notation c : csa(a,b,kdt) —and let c dominate or satisfaction-precede core speech act c ′. we use the notation accessible(c,c ′) to indicate that c either dominates or satisfaction-precedes c ′ in the sense of (grosz and sidner 1986) (see (poesio 1993, 1994; poesio and traum 1997) for details). then pars2 hypothesizes that content kdt is the resource situation for definite unp . pars2 [unp : np [det the: λpλp ′([y|y is ιxp (x)];p ′(y))] [un′ : n ′ [n: λx[|p(x)]]]] unp ↑ u , generates(u,c′), c : csa(a,b,kdt), accessible(c,c′), kdt |= p(z) ⇒ [unp : np [det the: λpλp ′([y|y is ιxp (x)];p ′(y))] [un′ : n ′ : λxkdt; [|p(x)]; [|x is z] [n : λx[|p(x)]]]] 253 poesio and rieser notice that this proposal amounts to the claim that there are two separate attentional structures: one depending on visual attention (implemented here in terms of the msoa situation) and one depending on accessibility.13 there is still an open issue with the current proposal: how can antecedents become accessible in the sense just discussed during incremental interpretation of anaphoric expressions, when the illocutionary force of the utterance to which the anaphoric expression belongs may not yet have been detected? we postpone discussing this issue to the next section. in the rest of this section we complete the discussion of demonstratives and introduce our treatment of pronouns. 4.4 demonstratives and pointing kaplan (1978) did not actually propose that all referring expressions are ambiguous, only demonstratives; and he did not view all references to objects in the visual situation as directly referring, only those cases which expressed a demonstration, usually by pointing. in this section we will see that even the data about demonstratives accompanied by pointing do not require stipulating an ambiguity, reprising the arguments from (poesio and rieser 2009). the main claim of poesio and rieser is that the findings from lücking et al. (to appear) suggest that pointing is just another way for anchoring resource situations. using a marker-based optical tracking system, lücking et al. (to appear) measured in detail the precision with which the pointing cone projected by an index finger or gaze (lücking et al. 2010; pfeiffer 2010) uniquely identifies an object in a visual scene. they concluded that pointing is fuzzy: in most demonstrations the projected ray fails the target. this led them to suggest what they called the inf heuristic: inf (inf) an object is referred to by pointing only if 1. the object is intersected by the pointing cone and 2. the distance of this object from the central axis of the cone is less than any other object’s distance within this cone. inf succeeds in 96% of the cases, which led poesio and rieser to formulate what they called strong prag hypothesis: strong prag hypothesis a pointing gesture refers to the one object selected by an appropriate inference from the set of objects covered by the pointing cone extending from the index finger. this led poesio and rieser to introduce a third principle for anchoring resource situations, pars3— an implementation of the inf heuristic. the formulation of the principle proposed here (a slight variant of that proposed in our earlier paper), in addition to coercing the resource situation for the 13. as one of the anonymous reviewers pointed out, this version of anaphoric accessibility may seem overly restrictive in that it doesn’t account for cases in which the antecedent of an anaphoric expression is cumulatively constructed out of referents introduced in separate utterances. e.g., in k1: mary (x) entered the store. k2: soon after john (y) reached her. they (x+y) were looking for a present for susan, the antecedent for they is the sum of john and mary, introduced in separate drss k1 and k2. we argue that such cases are not particularly problematic but do require to impose constraints like asher and lascarides’ veridicality, imposing that interpretation of x in k1 and k2 is consistent. an alternative explanation proposed in (poesio 1994) was to stipulate the construction of a propositional structure out of the content of the single utterances as proposed by webber (1991). 254 incremental anaphora and reference through resource situations demonstrative unp to the area of the visual situation kpointing that is covered by the pointing cone generated by pointing act p that temporally overlaps with unp , also proposes as interpretation for the demonstrative np the object z which is closest to the pointing axis of the cone. (we won’t get here into the best formulation of the notion of ‘nearest to’ and simply hide the details in a function nearest-to; we also assume that every pointing action p has a ‘pointing axis’ again without entering into any details.) pars3 [unp : np [det this: λpλp ′([y|y is ιxp (x)];p ′(y))] [un′ : n ′ [n: λx[|p(x)]]]], p : point(a,kpointing), overlaps(unp , p ), kpointing |= p(z), z is nearest-to(pointing-axis(p )), ⇒ [unp : np [det this: λpλp ′([y|y is ιxp (x)];p ′(y))] [un′ : n ′ : λxkpointing; [|p(x)]; [|x is z] [n : λx[|p(x)]]]] notice that as formulated pars3 only applies to demonstratives with this. we believe a version of the default may exist for demonstratives with that, but probably not for definite descriptions. 4.5 resource situations and pronouns up until now we have only been concerned with full nominals. we conclude this section by discussing pronouns, beginning with their semantics. the resource situation idea suggests the following about pronouns. it has often been argued that, syntactically, pronouns in english are like determiners. the translation proposed for pronouns such as german sie in (21) makes pronouns behave semantically like determiners, as well. (21) unp :np upro:pro sie:λpλp ′([y|y is ιxk;p (x)];p ′(y)) this translation is based on the idea that whereas the definite article may be licensed by a semantically functional, but non anaphoric, predicate, pronouns must always be pragmatically licensed— i.e., there must be some highly salient resource situation k containing a highly salient object. furthermore, pronouns require a contextual property restricting the interpretation of the referent y: resolving a pronoun amounts to identifying such restriction. one obvious candidate is an identity property—i.e. a property of the form λw([|w is z]) where z is a discourse entity. according to the treatment just sketched, resolving sie in (10) involves identifying the content of the first directive in (11), k1.1, as the resource situation for the pronoun, and discourse entity z as the antecedent (i.e., applying the result to the identity property λw([|w is z]). as said above, our theory of anaphoric interpretation is based on the assumption that the interpretation rules—the principles for anchoring resource situations—depend very much on the form of the referring expression. we assume therefore that the interpretive steps just discussed are the result of a principle very much like pars2, but which applies specifically to pronouns. we call 255 poesio and rieser the principle pars2pro. in first approximation, here is a version of the principle that is exactly as pars2 except that is triggered by the occurrence of a pronoun instead of by the occurrence of a definite description. this version of the principle generates a semantic interpretation for any pronominal form pro which is part of the performance of an utterance u generating core speech act c ′ such that an antecedent for pro is available as part of the content of speech act c accessible from c ′. pars2pro [unp : np [pro pro: λpλp ′([y|y is ιxk;p (x)];p ′(y))]] unp ↑ u , generates(u,c′), c : csa(a,b,kdt), accessible(c,c′), kdt |= p(z) ⇒ [unp : np : λp ′([y|y is ιxkdt;p (x)]; [y is z];p ′(y)) [pro pro: λpλp ′([y|y is ιxk;p (x)];p ′(y))]] this principle produces the interpretation in (21′). (21’) unp :np:λp ′[y|y is ιxk1.1; [|x is z]];p ′(y) udet:det sie:pro:λpλp ′([y|y is ιxk;p (x)];p ′(y)) we will propose a revised formulation of this principle, including also an agreement check, in section 5.4, after presenting our view of how surface anaphoric expressions like pronouns are interpreted. 5. accounting for the evidence about incremental reference resolution in this section we discuss how the proposals about modelling incremental language processing as a process of defeasible inference and about the semantics of anaphoric expressions discussed in the previous section can account for the evidence about the incremental interpretation of anaphoric expressions coming from the psycholinguistics literature. 5.1 basics: incremental resolution of references to the visual scene we begin by showing how the proposed theory accounts for the fundamental results concerning incrementality in reference to objects in the visual world from tanenhaus et al. (1995) and eberhard et al. (1995). let us consider the two types of visual world situation studied by tanenhaus and colleagues and shown in figure 1. in the situation on the left there is only one apple; in the situation on the right there are two. subjects looking at either visual situation hear the instruction “put the apple on the towel in the box.” the first word utterance in the instruction, of verb put, leads to updating the discourse situation first by recording the occurrence of an utterance of word “put”, and then by the results of lexical access—which, in the ltag framework adopted here, means predicting the utterance not only of a vp, but of a whole sentence, as discussed in detail in (poesio and rieser 2010). to keep the syntactic trees in this section manageable we will therefore omit representing this part of the phrasal structure. let us instead consider in more detail what happens according to ptt when processing the next words in the instruction, forming the referring expression the apple. 256 incremental anaphora and reference through resource situations perceiving an utterance of determiner the results in the discourse situation being updated with the observation of an utterance of determiner the, as in (22). (22) [ uthe : utter(a, “the′′)] this update leads in turn to the parallel activation of all lexical access defaults associated with word form the, and in particular lexical default lex-the discussed in section 3.4. this in turn leads to the update in (17) here repeated for convenience. (we are not concerned here with wordsense disambiguation, but anyway we can assume it’s not a major issue in the case of this utterance.) (17) [ udet, unp , un ′ | unp :np xxxxxx ������ udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un ′ :n’ ] next (or while the interpretation inferences activated by the are taking place), perceiving the utterance of noun apple leads to the update in (23), which in turn leads again to lexical access, i.e, to the parallel activation of all lexical defaults, and to the selection of the interpretation in (24) through wordsense disambiguation processes which are not our concern here.14 (23) [ uapple : utter(a, “apple′′)] (24) [ un | un :n apple:λx[|apple(x)] ] parsing then results in the interpretation in (25), through one of the basic ltag operations, substitution of un into unp . (25) [ udet, unp , un ′ | unp :np``````̀ udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un ′ :n′ un :n apple:λx[|apple(x)] ] in both visual world scenarios in figure 1 the entire visual world—called here kvisual—is in the mutual focus of attention msoa = kvisual so that the subject can use pars1 to assign kvisual as the resource situation for the definite. in the case of a single apple—let us call that a1 —pars1 can only be applied once, producing the single hypothesis in (26). this interpretation is fully specified: the discourse situation provides syntactic 14. two senses are listed for apple in wordnet—the fruit sense and the tree sense—but this is presumably a case of polysemy which would be handled in ptt by assuming a p-underspecified lexical interpretation, see (poesio to appear). 257 poesio and rieser and semantic interpretations for all phrasal and lexical mces. the processor can therefore commit to the interpretation in (26), which results in the concentration of fixations on the target word observed in such situations (see e.g., the introduction to such results at pages 13–14 of (tanenhaus and trueswell 2005)). (committing to this hypothesis also prevents further attachments, as discussed below.) (26) unp :np:λp ′[y|]; [|y is ιxkvisual; [|apple(x)]; [|x is a1]];p ′(y)hhhhhhhhh ((((((((( udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un ′ :n′ : λxkvisual; [|apple(x)]; [|x is a1] un :n apple:λx[|apple(x)] in (26), the interpretation for unp is the set of properties that hold of discourse referent y such that y is the only object inkvisual that is an apple and is equal to a1. by contrast, in the situation illustrated on the right of figure 1, pars1 can be applied twice to produce two hypotheses concerning the sem value of utterance un ′ —as in (26), and the interpretation which is identical to the one (26) in all respects except that the cohort apple is chosen (let us call this second apple a2). this leads to the fixations being divided between the two apples.15 however both of these interpretations satisfy the uniqueness requirement on the interpretation of the definite. as a result, subjects cannot commit to a single extension, and therefore cannot assign an interpretation to the np utterance unp , as discussed in section 3.5, and therefore have to backtrack, so that when the rest of the instruction, . . . on the towel in the box, comes in, hypothesis (25) is still open to modification, which results in the lack of garden path effect in this case, as discussed in section 5.3. 5.2 incremental establishment of referential domains the key difference between the resource situation view of domain restriction and standard theories of quantifier domain restrictions such as those proposed, e.g., in (partee 1991; rooth 1992), in which any contextually salient property p can serve to restrict the domain of a quantifier, is the idea that the domain is restricted to the objects of a situation—a spatially and possibly temporally limited set of objects. in our original work (poesio 1993, 1994) this stronger formulation of domain restriction (at least for definite descriptions), inspired by grosz (1977), was motivated by the fact that restriction domain shifts in the trains dialogues appeared to be tied in with locations on the map that participants in the experiments are looking at, which represents a highly simplified ’world’ with a few towns represented as circles and connected by lines representing railways—i.e., moving a train to a town seemed to restrict the domain of interpretation to the area of the map around that town. this led to the formulation of the hypothesis that the trains world had a structure, in the sense that at the very least each landmark in the world identified a sub-situation that could serve as msoa; it was also possible that larger sub-situations could be identified. and whereas grosz 15. numbers of fixations in visual world studies are computed across a number of subjects, which, as pointed out by one of the reviewers of this papers, leads to the question whether each subject divides her/his attention between the extensions, or whether half of the subjects fixate on one interpretation whereas the other half fixate on the second. the matter clearly requires more empirical evidence, but judging from our experience with ambiguity perception (e.g., (poesio and artstein 2005)), we feel that the second explanation is more likely, and that individual subjects choose to fixate between the two equally likely candidates randomly as discussed in (poesio 1994). 258 incremental anaphora and reference through resource situations (1977) had identified focus shifting principles tied to the structure of the task, we identified a new, spatially related (visual) focus shifting principle that we called follow-the-movement: follow-the-movement part of the intended effect of an utterance instructing an agent to move an object from one location to another is to make the terminal location of the movement the new mutual situation of attention. one of the great opportunities offered by the development of the visual-world methodology was the possibility to investigate in a proper empirical fashion the relation between shifts in the visual focus and reference interpretation, and indeed a key line of research in this type of work has been the study of incremental focus shifting, under the name rapid restriction of referential domains (chambers et al. 2002; brown-schmidt et al. 2005). the experiments by brown-schmidt et al. discussed in section 2 are the experimental setting closest to that of the trains dialogues. although the experimenters did not directly test specific focus-shift principles, the results clearly confirm the hypothesis that spatial landmarks identify sub-situations from the attentional point of view. the preliminary results of a more direct test of follow-the-movement just undertaken in our lab (in preparation), and (more indirectly) of the generation experiments in (zender 2010), also appear to confirm the existence of that focus shift principle. we hypothesize therefore that at least in the simplified type of visual scenes used in visual world experiments or in the trains dialogues, each landmark l identifies a visual subsituation kl. having made this assumption, follow-the-movement translates into the following default: if a intends a,b to move object c to landmark l, a also intends sub-situation kl to be the new msoa. follow-the-movement inta(move({a,b},c,l)),⇒ inta(msoa is kl) the data on referential domain restriction by chambers et al., however, seem to indicate that the interpretation domain can also be restricted not according to spatial location, but according to what brown-schmidt et al. call task compatibility: after hearing pick up the cube. put it in . . . , attention is focused on the set of containers into which the cube can fit. this suggests that a more general formulation of domain restriction principles is needed than we proposed in our previous work, one in which the resource situation need not be spatially defined. we claim that the new formulation of resource situations in terms of drss proposed in this paper is exactly what is needed and covers these types of referential domain restriction as well. specifically, we propose that upon hearing the utterance pick up the cube. put it in . . . , the discourse situation is updated by introducing a new resource situation kfit−cube thus defined: (27) [ kfit−cube|kfit−cube is [x, y, s|x is ιz.pick-up(subj, z), container(y), s : fit-in(x, y)], ] and this new resource situation then becomes the msoa, ensuring that pars2pro would only suggest those objects as antecedents of following references to containers. 5.3 the interaction between reference and parsing next, let us discuss how the proposal presented in this paper can account for the interaction between reference and parsing studied in (crain and steedman 1985; altmann and steedman 1988; 259 poesio and rieser tanenhaus et al. 1995; spivey et al. 2002). let us consider again the two types of experimental settings constrasted in the study by tanenhaus and colleagues (figure 1, left and right) which we just discussed focusing exclusively on reference interpretation. let us begin again with the visual world context on the left, in which there is only one apple. as just discussed, in this situation the processor can use pars1 to choose the currentmsoa,kvisual, as the resource situation, and choose the only apple in kvisual, that we called a1, to produce the single hypothesis in (26), repeated below for convenience. (26) unp :np:λp ′[y|]; [|y is ιxkvisual; [|apple(x)]; [|x is a1]];p ′(y)hhhhhhhhh ((((((((( udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un ′ :n’:λxkvisual; [|apple(x)]; [|x is a1] un : n apple:λx[|apple(x)] as discussed above, this interpretation is fully specified: the syntactic and semantic interpretations of each phrasal and lexical mce are fixed by the discourse situation. the subject can therefore commit to the interpretation in (26), preventing further attachments. so when the subject hears next the utterance of a pp, on the towel, the only available interpretation is as an argument of put, leading to a garden path. by contrast, in the case of the visual situation on the right of figure 1, the hypotheses produced using pars1 ((26) and the interpretation which is identical to (26) in all respects except that apple a2 is chosen) could not be committed to as they did not satisfy the uniqueness restriction imposed by the determiner, and therefore subjects have to backtrack to hypothesis (25). this means that when the next part of the instruction, . . . on the towel, comes in, this hypothesis is still open to modification—in fact, it requires the meaning of un ′ to be restricted in order to find a discourse referent satisfying the uniqueness condition. we argue that this requirement is what makes adjunction of . . . on ... to un ′ preferred over substitution as second argument of put. as a result of this adjunction we obtain: unp :nphhhhhhhh (((((((( udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un′ :n’ pppp ���� u2 n′ :n’ un : n apple:λx[|apple(x)] upp :pp hhh ��� up : p on u2 np : np at this point a second definite np is uttered, the towel. a crucial point in need for an explanation about this example is the fact that this definite np is felicitous in a context in which there are two towels. we claim that this is another case of task compatibility leading to rapid referential domain adaptation, just as in the cases studied by chambers et al. and whose analysis we presented in the previous section. i.e., we claim that upon hearing on, the discourse situation is updated by changing the msoa to a new resource situation containing the objects on which an apple is placed, as follows. 260 incremental anaphora and reference through resource situations (28) [kapple−on,msoa| kapple−on is [x, y, s|x is ιz.[|apple(z),put(subj, z, y)], object(y), s : on(x, y)], msoa is kapple−on ] notice that the new resource situation only contains one towel, the towel in the top left quadrant, that we will call t1. kapple−on can now be chosen as resource situation for the towel via pars1; this makes the definite felicitous.16 the interpretation resulting from this first application of pars1 is shown below. according to this interpretation, the apple on the towel gets interpreted as the apple on the unique towel in the visual scene that has an apple on it. unp :nphhhhhhhh (((((((( udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un′ :n’ xxxxxx ������ u2 n′ :n’ un : n apple:λx[|apple(x)] upp :pp:λx[z|]; [|z is ιwkapple−on; [|towel(w)]; [|w is t1]]; [|on(x, z)] as this update makes the meaning of un ′ functional, pars1 can now apply to choose the visual situation kvisual as resource situation for the first definite, identifying a1 on towel t1, as shown below. unp :nphhhhhhhh (((((((( udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) un′ :n’: λxkvisual; [|apple(x)]; [|x is a1]; [z|]; [|z is ιwkapple−on; [|towel(w)]; [|w is t1]]; [|on(x, z)] xxxxxx ������ un′ :n’ un : n apple:λx[|apple(x)] upp :pp:λx [z|]; [|z is ιwkapple−on; [|towel(w)]; [|w is t1]]; [|on(x, z)] as this interpretation is fully specified, the subject can commit. 5.4 incremental interpretation of anaphoric reference via pronouns whereas the experiments by tanenhaus et al. and by eberhard et al. were only concerned with references via full nominals to entities in the visual situation, the visual world methodology has also been shown in experiments such as those by arnold et al. (2000) to confirm earlier evidence (by, e.g., (corbett and chang 1983)) that pronouns are interpreted incrementally, as well. 16. the idea that the towel in this case is felicitous in virtue of being interpreted as, essentially, ’the towel that an apple is on’ is reminiscent of the interpretation for definites proposed by webber in her thesis (webber 1979). 261 poesio and rieser we now discuss how the new version of pars2 for pronouns proposed in section 4.5, pars2pro, can explain how an interpretation is assigned to the pronouns in the experiments discussed by arnold et al., repeated here. (29) a. donald is bringing some mail to [mickey / minnie] while a violent storm is beginning. b. he’s / she’s carrying an umbrella, and it looks like they’re both going to need it. as we are not concerned with speech act interpretation and discourse structure recognition in this paper, we will just make some assumptions here about the results of these interpretation processes. we believe that a fuller account could be developed building on the detailed analysis of these interpretation processes proposed by asher and lascarides (2003) in the sdrt framework, which shares many assumptions with ptt. the first utterance, (29a), generates a core speech act of type assert, that we will call ce1. overall, the update resulting from the first utterance is then as in (30a), where a is the experimenter and b the experimental subject. in this example we have included in the description of the discourse situation, in addition to full utterance up1, two of its subconstituents: the micro conversational events up1.1 of uttering name “donald” and up1.2 of uttering name “mickey,” in both cases omitting syntactic and semantic information for these mces except for their gender. the update resulting from the variant with minnie instead of mickey, in (29b), is the same as (29a) except that in this case x3 refers to minnie and the gender of up1.2 is feminine. (30) a. [ k1, up1, ce1, up1.1, up1.2| k1 is [e1, x1, x2, x3|x1 is donald,mail(x2), x3 is mickey, e1 : bring(x1, x2, x3)], up1.1 : utter(a,”donald”), gen(up1.1) is masc, up1.1 ↑ up1, up1.2 : utter(a,”mickey”), gen(up1.2) is masc, up1.2 ↑ up1, up1 : utter(a,”donald is bringing some mail to mickey”), sem(up1) is k1, ce1 : assert(a,b,k1) generate(up1, ce1) ] b. [ k1, up1, ce1, up1.1, up1.2| k1 is [e1, x1, x2, x3|x1 is donald,mail(x2), x3 is minnie e1 : bring(x1, x2, x3)], up1.1 : utter(a,”donald”), gen(up1.1) is masc, up1.1 ↑ up1, up1.2 : utter(a,”minnie”), gen(up1.2) is fem, up1.2 ↑ up1, up1 : utter(a,”donald is bringing some mail to minnie”), sem(up1) is k1, 262 incremental anaphora and reference through resource situations ce1 : assert(a,b,k1) generate(up1, ce1) ] the pronoun (he or she) uttered at the beginning of the second utterance (utterance up2.1) is interpreted as the beginning of an utterance up2 generating a second core speech act ce2 whose precise type we do not know yet.17 we show the update resulting from he in (31a), that resulting from she in (31b). (31) a. [ k2, up2, ce2| utterance(up2),ce2 : csa(a,b,k2), generate(up2, ce2),up2.1 : utter(a,”he”), up2.1 ↑ up2, gen(up2.1) is masc ] b. [ k2, up2, ce2| utterance(up2),ce2 : csa(a,b,k2), generate(up2, ce2)up2.1 : utter(a,”she”), up2.1 ↑ up2, gen(up2.1) is fem ] according to the theory of resource interpretation anchoring developed in section 4, in the case of references to the visual situation it doesn’t matter that the core speech act of whose realization the referring expression is part, or its connection with the rest of the discourse structure, hasn’t yet been identifed at the time the referring expression is uttered, because the interpretation of the referring expression only depends on the visual attentional state, as opposed to the linguistic attentional state. on the other hand this does matter in the case of anaphoric references, like pronoun sie in (10) or the pronouns in the example under discussion. this is because principles pars2 and pars2pro, in order to choose the content of a previous core speech act c as resource situation, require that core speech act to dominate or satisfaction-precede the core speech act containing the anaphor. the question is, how can anaphora resolution proceed prior to recognizing the intention behind the utterance being produced? as in sdrt, and consistent with the general view of interpretation discussed in section 3, discourse structure recognition is viewed in ptt as a defeasible inference process. this means that hypotheses about discourse structure and accessibility are produced before complete information is available. in fact, in ptt it is assumed that such hypotheses tend to be produced very early on the basis of fairly superficial information, and possibly revised later. in cases like the example under discussion, we assume that the accessibility of (the content of) ce1 is hypothesized before knowing the content of ce2—possibly even before knowing its illocutionary force. there are two possible time points at which this accessibility hypothesis may be produced. it could be produced immediately, on coherence grounds: i.e., in circumstances in which it would seem that a story is being told, simply assume by default that the next utterance is going to tell the next episode in the story, by way of a default that would be like a highly underspecified version of asher and lascarides’s narration (asher and lascarides 2003). alternatively, the accessibility hypothesis might be produced as a byproduct of establishing a preliminary link between the pronoun and one of the antecedents. we will only pursue here this second possibility, as in this way we can also spell out more fully the view of anaphoric processing adopted in ptt. (anyway, more empirical evidence about the precise time point of discourse structure identification is needed before being able to decide which of the possibilities is more likely, or whether the interpretation results from a combination of the two factors.) 17. as pointed out by one of the reviewers, a full account of incremental processing would also require an account of the process by which the beginning of an utterance is recognized. 263 poesio and rieser the treatment of ‘surface’ anaphora resolution (hankamer and sag 1976) we propose here is based, on the one hand, on the proposal by garrod (1994) and garrod and sanford (1994) that the resolution of these types of anaphoric reference consists of distinct bonding and resolution stages; on the other, on centering theory (grosz et al. 1995; poesio et al. 2004). according to garrod and sanford, in the initial bonding stage a link is made between the anaphoric expression and one or more candidate antecedents in the discourse context, on the basis of superficial information. in the subsequent resolution phase, the link made in the bonding stage is evaluated, recomputed if necessary, and integrated into the semantic interpretation. what is meant by ‘superficial information’ has never been spelled out in detail by garrod and sanford, but we propose here that the ‘superficial level’ is the micro conversational events level hypothesized in ptt: i.e., that at least some of the defeasible rules for anaphora resolution establish bonding links between the micro conversational events introducing discourse antecedents. we further propose that these mces are the forward looking centers (cfs) of centering, another notion whose linguistic characterization has never been spelled out fully (see also (poesio 1994, to appear)). these bonding links between cfs, in turn, lead to hypothesizing dominance or satisfactionprecedes links between the core speech acts generated by the utterance of which the anaphoric cf is a constituent and the utterance which includes the antecedent cf. this link, finally, enables pars2pro—our proposal concerning the ‘resolution’ stage. this theory is implemented by assuming, first of all, that the local focus is recomputed after what we will call here c-utterance, for ‘centering utterance’, as proposed in centering theory.18 this translates into hypothesizing that the discourse situation has a distinguished discourse marker (in the sense of cdrt) called ccu (for ’current c-utterance’), whose value changes after every sentence; we also hypothesize that end-of-sentence processes include updating the discourse situation with several statements of the form cf-utt(u, u’) indicating that u′ is a cf in c-utterance u. second, we stipulate that at least some of the mechanisms for interpreting pronouns operate at the surface level: specifically, that (one of) the default rules for pronoun resolution, that we call here pro-match, creates a bonding link between the utterance of a pronoun and one of the cfs of the current c-utterance, provided that their agreement features match. we write bond(u,u′) to indicate that u′ is bonded to u: in the case of this default rule, that the utterance upro of a pronoun realized in ccu un+1 is bonded to the utterance unp of a cf realized in un and which agrees with the pronoun in gender person and number.19 cat(upro) is pro, cf-utt(un+1,upro), cf-utt(un,unp), ccu is un+1, agr-match(unp,upro) : bond(unp,upro) pro-match bond(unp,upro) 18. the notion of ’utterance’ is used in centering to indicate the amount of language after which the local focus gets updated. we identify here ’utterances’ in the centering sense with sentences, on the basis of the results in (poesio et al. 2004). 19. although we only propose here this treatment for surface anaphors, we suspect a similar dd-match rule may generate anaphoric interpretation hypotheses for definite descriptions on the basis of head similarity. 264 incremental anaphora and reference through resource situations the establishment of bonding links is one trigger for further inference processes that hypothesize dominance / satisfaction precedes relations between the core speech acts generated by the two utterances, if they haven’t already been established by coherence assumptions or by previous intention recognition processes. the following default, acc-from-bond, hypothesizes that conversational actionce1 is accessible in the sense discussed in section 4.3 from conversational actionce2 if a pronoun that is part of the realization of ce1 is bonded to a cf that is part of the realization of ce2. cat(upro) is pro, cf-utt(un+1,upro), cf-utt(un,unp), bond(unp,upro), generates(un+1,ce2), generates(un,ce1) : accessible(ce1,ce2) acc-from-bond accessible(ce1,ce2) after establishing accessibility through acc-from-bond, or possibly through other shallow coherence-inference methods, the resolution stage can begin and pars2pro can be applied to identify the resource situation. we can finally produce the revised version of pars2pro promised earlier in the paper. whereas pars2 for definites only depends on the existence of an object of the appropriate type in the resource situation, according to pars2pro the identification of an antecedent for a pronoun utterance upro depends on having established (via agreement matching) a bonding link with a forward-looking centeru1 np in the context. the pronoun is not just resolved to any referent in the content of a context accessible from the core speech act being produced, but to the described object do of u1 np . pars2pro [upro : np [pro pro: λpλp ′([y|y is ιxk;p (x)];p ′(y))]] upro ↑ u , generates(u,c), bond(unp , upro), unp ↑ u ′, generates(u ′, c′), c : csa(a,b,kdt), accessible(c′, c), do(unp ) is z, kdt |= p(z) ⇒ [upro : np λp ′([y|y is ιxkdt; [|y is z]];p ′(y)) [pro pro: λpλp ′([y|y is ιxk;p (x)];p ′(y))]] let us now return to the data from arnold and consider first (31a) uttered in the discourse situation following (30b), in which the two cfs are of different gender. as only mce up1.1 (the utterance of “donald”) matches up2.1 in gender, pro-match can only produce one hypothesis about the interpretation of up2.1: that it bonds to up1.1—i.e., to the update in (32). (the same happens when “she” is uttered, with both contexts in (30).) (32) [ |bond(up1.1, up2.1)] this hypothesis is immediately committed to, resulting in acc-from-bond being triggered. this in turn leads to the following update: (33) [ |accessible(ce1, ce2)] which in turn triggers pars2pro, resulting in the interpretation in (34) 265 poesio and rieser (34) unp :npλp ′[y|y is ιxk1; [|x is z]];p ′(y) udet:det sie:λpλp ′([y|y is ιxk;p (x)];p ′(y)) it is not clear from the results of arnold et al. whether the concentration of the fixations on the target starting from around 400msec after the onset of the pronoun is the result of bonding or of resolution; more experimental evidence is needed to resolve the issue. let us now consider the case of (30a), in which both up1.1 and up1.2 match up2.1 in gender. as a result, pro-match can be activated in two different ways, producing the two distinct hypotheses in (35) (35) a. [ |bond(up1.1, up2.1)] b. [ |bond(up1.2, up2.1)] each of these hypotheses in turn activates acc-from-bond—the updates resulting from this default are however identical (and identical with (33)). arnold et al.’s results are that in case the same-gender target is the first mentioned entity, fixations quickly concentrate on the target, whereas if the target is the second-mentioned entity, the subjects look at both the target and the competitor in the same amount. this situation is reminiscent of the situation with lexical interpretation and scope access discussed in (poesio 1994, 1996), and suggests that a stronger default than pro-match is at play in the case of first-mention entities. when this default, shown below and that we call pro-match-fm, is triggered, it overrides the weaker pro-match; otherwise a conflict between weaker defaults is obtained, which typically results in a toss-up between the alternatives. cat(upro) is pro, cf-utt(un+1,upro), cf-utt(un,unp), first-mention(un,unp), cu is un+1, agr-match(unp,upro) : bond(unp,upro) pro-match-fm bond(unp,upro) (a more elegant formalization would of course be available in a framework with evidence accumulation.) 5.5 reference interpretation prior to hearing a complete head noun eberhard et al. (1995) and allopenna et al. (1998) showed that interpretation processes begin much earlier than discussed until now: they begin as soon as an unambiguous phonetic prefix has been uttered. allopenna et al. (1998) also showed that in the case of the interpretation of referring expressions, this unambiguous phonetic prefix need not be part of the head noun—in a situation in which there is a single red object, and click on the red triangle is uttered, fixations concentrate on that object as soon as adjective red has been perceived, without waiting to hear triangle. a proper account of the incremental effect of sub-word prefixes would require an implementation in ptt of a theory of sub-word based lexical access such as the cohort model (marslen-wilson 1987) or the trace model (mcclelland and elman 1986), so we will not attempt that here. we will however discuss how the present proposal accounts for how hearing an adjective affects reference resolution. 266 incremental anaphora and reference through resource situations the interpretation process resulting from an utterance of the is as discussed above, and results again in update (17). upon hearing red, the discourse situation is updated with the expectation of encountering an n’, as in (36). after parsing, this interpretation is adjoined to the syntactic interpretation of the, resulting in the updated interpretation in (38). (36) [ ured : utter(a, ”red”)] (37) [ u1 n ′ , u 2 n ′ | u1 n ′ : n ′ xxxxx ����� ured:adj red: λpλx[|red(x)];p (x) u2 n ′ : n ′ ] (38) [ | unp :nphhhhhhhhhh (((((((((( udet:det the:λpλp ′[y|y is ιxp (x)];p ′(y) u1 n ′ : n ′ xxxxx ����� ured:adj red: λpλx[|red(x)];p (x) u2 n ′ : n ′ ] the evidence from eberhard et al. suggests that, at least in interpretive contexts like the visual world scenarios, updates like (38) are sufficient to trigger reference resolution. this translates into an hypothesis that a version of pars1 exists activated by the observation of an utterance of an adjective with semantic interpretation λpλx([|p(x)];p (x)). this version of pars1, that we call pars1adj , is shown below. this version again requires a visual situationkmva to be in the mutual focus of attention, but unlike the versions of the default proposed earlier, it is sufficient for predicate p to hold of object z, even if the predicate is not the head of an np. the default updates the discourse situation by restricting the interpretation of the np through adjoining. pars1adj [unp : np [det the: λpλp ′([y|y is ιxp (x)];p ′(y))] [u1 n′ : n ′ [u1 adj : adj:λpλx[|p(x)];p (x)] u2 n′ : n ′ ]] msoa is kmva kmva |= p(z) ⇒ [unp : np [det the: λpλp ′([y|y is ιxp (x)];p ′(y))] [u1 n′ : n ′ [u2 adj :adj:λpλxkmva; [|p(x)]; [|x is z];p (x) [uadj :adj: λpλx[|p(x)];p (x)]] u2 n′ : n ′]] this hypothesis raises several issues which we believe could be addressed by further experimental work. first of all, there is the question of the extent to which interpretation in a visual world context is the same as interpretation in other contexts, already raised by britt (1994) (see (tanenhaus and trueswell 2005)). i.e., is a rule like pars1adj only available in contexts in which a visual situation is available? or perhaps only when the subject is required to do something with the objects in the situation? we also hope that this example may explain why we believe that formulating the interpretation processes in more detail is going to show that current theories about incremental reference interpretation are still open. for instance, our formulation raises the question of whether there are in fact 267 poesio and rieser separate versions of pars1—i.e., different ways of using the visual situation for different types of expressions—or a single one. one may also wonder whether the principle proposed is only valid for intersective adjectives, or for all types, or even for all types of modifiers including for instance nominal premodifiers.20 6. related literature we are not aware of any other account of the psychological results about incremental reference using the visual world paradigm in terms of a dynamic semantics, but there has been a lot of research relevant to providing such an account. in this section we will first of all discuss other work on incremental interpretation and formal grammar; then alternative theories of the dynamics of dialogues; finally, some recent computational models of the incremental interpretation of reference. 6.1 related linguistic formalisms modulo the re-interpretation of trees in terms of mces, the grammar formalism used in ptt is (deliberately) very standard both in terms of ltag analysis and in terms of cdrt semantics; in particular, it is closely related to muskens’ logical description grammar (ldg) (muskens 2001), an earlier proposal to combine ltag with cdrt. the main differences concern the semantics of definites (muskens’ analysis is not based on loebner’s account or on resource situations). also muskens is not particularly concerned with incrementality; if at all, his formalism is more motivated by ideas about the role of underspecification in formal grammar. the opposite is true of dynamic syntax (cann et al. 2005), one of the few formal grammatical formalisms taking the incrementality of interpretation as the central fact that a theory of grammar has to explain.21 the main difference between ptt and dynamic syntax lies in the treatment of anaphora. contrary to what one could expect from the name, dynamic syntax is not based on a dynamic approach to the common ground in the sense of drt or dynamic logic. its concerns are mainly at the sentence level, and therefore, although it includes a proposal concerning the semantics of pronouns, it does not provide an account of which antecedents are available for them.22 6.2 other theories of the common ground in dialogue the two best-known theories of semantics in dialogue are sdrt (asher and lascarides 2003) and ginzburg’s kos (ginzburg 2011). like ptt, sdrt is an extension of drt developed to account for the pragmatics of the common ground—in particular, for the effect of discourse structure on language interpretation. it thus provides a highly developed account of rhetorical relations and the process by which they get established, also based on the assumption that this is a process of defeasible inference. it does not however provide a theory of how interpretation proceeds incrementally, and it would not be easy to incorporate a treatment like ptts, for although it would be quite simple to include the equivalent of micro conversational events in sdrt’s picture of the common ground, one of the fundamental 20. one limitation of pars1adj as formulated here is that it only applies to definite descriptions with a single adjective. this is not however a real limitation as it can easy be remedied by generalizing the rule to make it sensitive to the occurrence in the discourse situation of any np containing an adjectival phrase (adjp) rather than a single adjective. 21. another being combinatorial categorial grammar (steedman 2001). 22. a more detailed comparison between ptt and dynamic syntax can be found in (poesio and rieser 2010). 268 incremental anaphora and reference through resource situations assumptions of the theory is that the processes of grammatical interpretation and discourse interpretation are completely distinct—in fact, they are ruled by distinct logics. kos is built like ptt on the view of the common ground developed in situation semantics, and as such it incorporates very similar views about the presence of non-sentential utterances in the common ground (e.g., (fernandez 2006)) but it is not built on a logical formalism designed to account for the anaphoric properties of utterances and until recently it did not incorporate an extensive treatment of anaphora. a fairly detailed comparison between ptt and both these semantical formalisms can be found in (poesio and rieser 2010). 6.3 computational accounts of incremental reference a computational implementation could eventually provide a large-scale test of the predictions of a model such as the one proposed in this paper. in recent years, the first computational models of incremental reference resolution have appeared. although still relatively simple from a linguistic perspective, these models give us hope that computational modelling could soon become a tool in the study of incremental reference resolution. stoness et al. (2004) propose an account of the incremental interaction of reference resolution with parsing implemented in an actual spoken dialogue system. the account is based on the hypothesis that the reference resolution module is called upon every time that the parser identifies an np, and attempts to find a referent for it in the knowledge base. the ability to resolve it adds to the score of that particular parsing interpretation, which may lead to it being chosen over the alternatives. stoness et al. showed that this may result in improvements in parsing performance. schlangen et al. (2009) propose a model of incremental reference resolution based on a bayesian filtering model which captures quite directly the visual world scenario. each object r in the visual scene is associated with a probability p (r|w1:n) that words w1 . . . wn are referring to that object. this probability is incrementally updated after every word. schlangen et al. proposed an evaluation metric for this task and methods for learning these probabilities. finally, dubey (2010) implemented a computational model of reference interpretation consisting of a probabilistic parser, a probabilistic coreference resolver, and a pragmatics processor modelling coherence constraints. (the first two models are trained on actual data, the latter hand-coded.) he tested the model by simulating the garden path data finding a good match between predictions of the model and the experimental results. 6.4 models of visual attention a proper account of the effect of visual salience on the interpretation of referring expressions will require a more detailed theory of visual attention than the one assumed here. a proposal in such direction was made by kelleher (kelleher et al. 2005; kelleher 2007). 269 poesio and rieser 7. discussion the proposal presented in this paper is, as far as we know, the only full account of the incremental interpretation of anaphoric and referential expressions taking into account the findings of both psycholinguistics and formal linguistics. the proposal makes a few clear predictions which should be possible to verify experimentally, including: • that definite descriptions are not ambiguous between an anaphoric and referential interpretation; • that it should be possible to categorize anaphors according to whether their resolution takes place at the surface level or at the deeper level. the main limitation of the present proposal is that it does not provide a full list of the principles governing focus shift and resource situation anchoring, and that the model of defeasible reasoning adopted here is very simple. a natural development of the theory would be to provide an account couched in terms of a probabilistic model that could be learned from data (jurafsky 1996; bod et al. 2003). a second limitation of the present work is that it is not integrated with an account of the grounding process such as that developed by traum (1994), whose interaction with the present model of incremental interpretation was discussed in (poesio and rieser 2010). we think this development would be especially interesting at the light of the evidence from, e.g., keysar et al. (2000) suggesting that reference does not only involve information in the common ground. finally, the current version of the theory also doesn’t take full advantage of the formalization in terms of a defeasible logic to provide an account of reanalysis (fodor and ferreira 1998) and repairs (ferreira et al. 2004). to do so however would require embedding the theory into a fully formalized account of belief revision. acknowledgments we gratefully acknowledge the support of the crc 673 alignment in communication at universität bielefeld, germany, the university of essex’s school of computer science and electronic engineering, and the center for mind and brain sciences, università di trento, to the research leading to this paper. we also wish to thank david schlangen, the editor of this paper, and the anonymous reviewers for their detailed comments and many insightful suggestions. references a. abeille and o. rambow, editors. tree adjoining grammars. csli, 2000. s. abney. parsing by chunks. in r. berwick, s. abney, and c.tenny, editors, principle-based parsing, pages 257–278. kluwer, dordrecht, 1991. j. f. allen, l. k. schubert, g. ferguson, p. heeman, c. h. hwang, t. kato, m. light, n. martin, b. miller, m. poesio, and d. r. traum. the trains project: a case study in building a conversational planning agent. journal of experimental and theoretical ai, 7:7–48, 1995. 270 incremental anaphora and reference through resource situations p. d. allopenna, j. s. magnuson, and m. k. tanenhaus. tracking the time course of spoken word recognition: evidence for continuous mapping models. journal of memory and language, 38: 419–439, 1998. g. t. m. altmann and y. kamide. incremental interpretation of verbs: restricting the domain of subsequent reference. cognition, 73:247–264, 1999. g. t. m. altmann and m. steedman. interaction with context during human sentence processing. cognition, 30:191–238, 1988. j. e. arnold, j. g. eisenband, s. brown-schmidt, and j. c. trueswell. the immediate use of gender information: eyetracking evidence of the time-course of pronoun resolution. cognition, 76:b13– b26, 2000. n. asher and a. lascarides. the logic of conversation. cambridge university press, 2003. j. l. austin. how to do things with words. harvard university press, cambridge, ma, 1962. j. barwise. the situation in logic. csli lecture notes. university of chicago press, 1989. j. barwise and j. perry. situations and attitudes. the mit press, 1983. r. beun and a. cremers. object reference in a shared domain of conversation. pragmatics and cognition, 6(1/2):121–152, 1998. t. bever. the cognitive basis for linguistic structure. in cognition and the development of language. wiley, new york, 1970. r. bod, j. hay, and d. jannedy, editors. probabilistic linguistics. mit press, 2003. g. brewka and t. eiter. prioritizing default logic. in s. hölldobler, editor, intellectics and computational logic: papers in honor of wolfgang bibel. kluwer, 2000. m. a. britt. the interaction of referential ambiguity and argument structure in the parsing of prepositional phrases. journal of memory and language, 33:251–283, 1994. s. brown-schmidt and m. tanenhaus. real-time investigation of referential domains in unscripted conversation: a targeted language game approach. cognitive science, 32:643–684, 2008. s. brown-schmidt, e. campana, and m. k. tanenhaus. real-time reference resolution by naive participants. in j. c. trueswell and m. k. tanenhaus, editors, approaches to studying worldsituated language use, pages 153–171. mit press, 2005. h. c. bunt. dialogue control functions and interaction design. in r.j. beun, m. baker, and m. reiner, editors, dialogue in instruction, pages 197–214. springer verlag, 1995. r. cann, r. kempson, and l. marten. the dynamics of language: an introduction. elsevier, 2005. b. carpenter. type-logical semantics. mit press, 1998. url http://mitpress.mit.edu/ book-home.tcl?isbn=0262531496. 271 poesio and rieser c. g. chambers, m. k. tanenhaus, k. m. eberhard, h. filip, and g. n. carlson. circumscribing referential domains during real-time language comprehension. journal of memory and language, 47:30–49, 2002. g. chierchia. dynamics of meaning. anaphora, presupposition and the theory of grammar. university of chicago press, 1995. h. h. clark. using language. cambridge university press, cambridge, 1996. h. h. clark and c. r. marshall. definite reference and mutual knowledge. in a. joshi, b. webber, and i. sag, editors, elements of discourse understanding. cambridge university press, new york, 1981. r. cooper. the role of situations in generalized quantifiers. in s. lappin, editor, handbook of contemporary semantic theory, chapter 3, pages 65–86. blackwell, 1996. a. corbett and f. chang. pronoun disambiguating: accessing potential antecedents. memory and cognition, 11:283–294, 1983. s. crain and m. steedman. on not being led up the garden path: the use of context by the psychological syntax processor. in d. r. dowty, l. karttunen, and a. m. zwicky, editors, natural language parsing: psychological, computational and theoretical perspectives, pages 320–358. cambridge university press, new york, 1985. a. dubey. the influence of discourse on syntax: a psycholinguistic model of sentence processing. in proc. of the acl, uppsala, sweden, 2010. k. eberhard, s. spivey-knowlton, j. sedivy, and m. tanenhaus. eye movements as a window into real-time spoken language processing in natural contexts. journal of psycholinguistic research, 24:409–436, 1995. r. fernandez. non-sentential utterances in dialogue: classification, resolution and use. phd thesis, department of computer science, king’s college, london, 2006. f. ferreira, e. f. lau, and k. g. d. bailey. disfluencies, language comprehension, and tree adjoining grammars. cognitive science, 28:721–749, 2004. j. d. fodor and f. ferreira, editors. reanalysis in sentence processing. kluwer, 1998. l. frazier. on comprehending sentences: syntactic parsing strategies. phd thesis, university of connecticut, 1979. available via indiana university linguistic club. l. frazier. sentence processing: a tutorial review. in m. coltheart, editor, attention and performance xii: the psychology of reading, pages 559–586. erlbaum, hove, 1987. l. frazier and k. rayner. taking on semantic commitments: processing multiple meanings vs. multiple senses. journal of memory and language, 29:181–200, 1990. s. c. garrod. resolving pronouns and other anaphoric devices: the case for diversity in discourse processing. in c. clifton, l. frazier, and k. rayner, editors, perspectives in sentence processing. lawrence erlbaum, 1994. 272 incremental anaphora and reference through resource situations s.c. garrod and a. j. sanford. resolving sentences in a discourse context. in m. a. gernsbacher, editor, handbook of psycholinguistics, chapter 20, pages 675–698. academic press, 1994. j. m. gawron and s. peters. anaphora and quantification in situation semantics, volume 19 of lecture notes. csli, 1990. e. gibson. a computational theory of human linguistic processing: memory limitations and processing breakdown. phd thesis, carnegie mellon university, pittsburgh, 1991. j. ginzburg. the interactive stance: meaning for conversation. oxford, 2011. j. ginzburg. situation semantics: from indexicality to metacommunicative interaction. in p. portner, c. maierborn, and k. von heusinger, editors, the handbook of semantics. de gruyter, to appear. to appear. d. gross, j. allen, and d. traum. the trains 91 dialogues. trains technical note 92-1, computer science dept. university of rochester, june 1993. b. j. grosz. the representation and use of focus in dialogue understanding. phd thesis, stanford university, 1977. b. j. grosz and c. l. sidner. attention, intention, and the structure of discourse. computational linguistics, 12(3):175–204, 1986. b. j. grosz, a. k. joshi, and s. weinstein. centering: a framework for modeling the local coherence of discourse. computational linguistics, 21(2):202–225, 1995. (the paper originally appeared as an unpublished manuscript in 1986.). j. k. gundel, n. hedberg, and r. zacharski. cognitive status and the form of referring expressions in discourse. language, 69(2):274–307, 1993. jorge hankamer and ivan sag. deep and surface anaphora. linguistic inquiry, 7(3):391–426, 1976. j. r. hobbs, m. stickel, p. martin, and d. edwards. interpretation as abduction. artificial intelligence journal, 63:69–142, 1993. w. s. horton and b. keysar. when do speakers take into account common ground? cognition, 59: 91–117, 1996. j. f. horty. defaults with priorities. journal of philosophical logic, 36:367–413, 2007. v. hoste. optimization issues in machine learning of coreference. phd thesis, university of antwerp, 2005. c. h. hwang and l. k. schubert. episodic logic: a comprehensive, natural representation for language understanding. minds and machines, 3:381–419, 1993. d. jurafsky. a probabilistic model of lexical and syntactic access and disambiguation. cognitive science, 20:137–194, 1996. d. jurafsky and j. h. martin. speech and language processing. prentice hall, 2nd edition, 2009. 273 poesio and rieser h. kamp and u. reyle. from discourse to logic. d. reidel, dordrecht, 1993. d. kaplan. on the logic of demonstratives. journal of philosophical logic, 8:81–98, 1978. j. d. kelleher. attention driven reference resolution in multimodal contexts. artificial intelligence review, 25:21–35, 2007. j.d. kelleher, f. costello, and j. van genabith. dynamically updating and interrelating representations of visual and linguistic discourse. artificial intelligence, 167:62–102, 2005. b. keysar, d. j. barr, j. a. balin, and j. s. brauner. taking perspective in conversation: the role of mutual knowledge in comprehension. psychological science, 11:32–37, 2000. h. e. kyburg. logical foundations of statistical inference. d. reidel, 1974. s. larsson and d. r. traum. information state and dialogue management in the trindi dialogue move engine toolkit. natural language engineering, 6:323–340, 2000. a. lascarides and n. asher. discourse relations and defeasible knowledge. in proc. acl-91, pages 55–63, university of california at berkeley, 1991. s. loebner. definites. journal of semantics, 4:279–326, 1987. a. lücking, k. bergmann, f. hahn, s. kopp, and h. rieser. the bielefeld speech and gesture alignment corpus (saga). in proc. of the lrec workshop on multimodal corpora: advances in capturing, coding and analyzing multimodality, 2010. a. lücking, t. pfeiffer, and h. rieser. pointing and reference reconsidered. submitted, to appear. m. c. macdonald, n. j. pearlmutter, and m. s. seidenberg. lexical nature of syntactic ambiguity resolution. psychological review, 101(4):676–703, 1994. w.d. marslen-wilson. sentence perception as an interactive parallel process. science, 189:226– 228, 1975. w.d. marslen-wilson. functional parallelism in spoken word recognition. cognition, 25:71–102, 1987. c. matheson, m. poesio, and d. traum. modeling grounding and discourse obligations using update rules. in proc. of the first annual meeting of the north american chapter of the acl, seattle, april 2000. j. l. mcclelland and j. l. elman. interactive processes in speech perception: the trace model. in d. e. rumelhart and j. l. mcclelland, editors, parallel distributed processing, volume 2. mit press, 1986. d. milward. dynamic dependency grammar. linguistics and philosophy, 17:561–605, 1994. r. a. muskens. combining montague semantics and discourse representation. linguistics and philosophy, 19:143–186, 1996. 274 incremental anaphora and reference through resource situations r. a. muskens. talking about trees and truth conditions. journal of logic, language and information, 10(4):417–455, 2001. g. d. nunberg. transfers of meaning. journal of semantics, 12(2):109–132, 1995. b. h. partee. topic, focus and quantification. in proc. salt-91, 1991. r. j. passonneau. getting and keeping the center of attention. in m. bates and r. m. weischedel, editors, challenges in natural language processing, chapter 7, pages 179–227. cambridge university press, 1993. n. j. pearlmutter and a. a. mendelsohn. serial versus parallel sentence comprehension. manuscript in revision, available at http://www.psych.neu.edu/people/ njp/papers, 1999. f. c. n. pereira. categorial semantics and scoping. computational linguistics, 16(1):1–10, march 1990. c. r. perrault. an application of default logic to speech act theory. in p. r. cohen, j. morgan, and m. e. pollack, editors, intentions in communication, chapter 9, pages 161–185. the mit press, cambridge, ma, 1990. t. pfeiffer. understanding multimodal deixis with gaze and gesture in conversational interfaces. phd thesis, university of bielefeld, 2010. m. poesio. a situation-theoretic formalization of definite description interpretation in plan elaboration dialogues. in p. aczel, d. israel, y. katagiri, and s. peters, editors, situation theory and its applications, vol.3, chapter 12, pages 339–374. csli, stanford, 1993. m. poesio. discourse interpretation and the scope of operators. phd thesis, university of rochester, department of computer science, rochester, ny, 1994. m. poesio. a model of conversation processing based on micro conversational events. in proceedings of the 17th annual conference of the cognitive science society, pages 698–703, pittsburgh, july 1995a. m. poesio. disambiguation as (defeasible) reasoning about underspecified representations. in p. dekker and m.stokhof, editors, proc. of the tenth amsterdam colloquium, pages 607–625. illc, december 1995b. m. poesio. semantic ambiguity and perceived ambiguity. in k. van deemter and s. peters, editors, semantic ambiguity and underspecification, chapter 8, pages 159–201. csli, stanford, ca, 1996. m. poesio. incrementality and underspecification in semantic interpretation. lecture notes. csli, stanford, ca, to appear. to appear. m. poesio and r. artstein. the reliability of anaphoric annotation, reconsidered: taking ambiguity into account. in a. meyers, editor, proc. of acl workshop on frontiers in corpus annotation, pages 76–83, june 2005. 275 poesio and rieser m. poesio and m. a. kabadjov. a general-purpose, off the shelf anaphoric resolver. in proc. of lrec, pages 653–656, lisbon, may 2004. m. poesio and r. muskens. the dynamics of discourse situations. in p. dekker and m. stokhof, editors, proceedings of the 11th amsterdam colloquium, pages 247–252. university of amsterdam, illc, december 1997. m. poesio and h. rieser. anaphora and direct reference: empirical evidence from pointing. in proc. of diaholmia, the 13th workshop on the semantics and pragmatics of dialogue, pages 35–43, stockholm, june 2009. m. poesio and h. rieser. completions, coordination, and alignment in dialogue. dialogue and discourse, 1(1):1–89, 2010. doi: 10.5087/dad.2010.001. m. poesio and d. traum. conversational actions and discourse situations. computational intelligence, 13(3):309–347, 1997. m. poesio, r. stevenson, b. di eugenio, and j. m. hitzeman. centering: a parametric theory and its instantiations. computational linguistics, 30(3):309–363, 2004. m. poesio, p. sturt, r. arstein, and r. filik. underspecification and anaphora: theoretical issues and preliminary evidence. discourse processes, 42(2):157–175, 2006. m. e. pollack. inferring domain plans in question-answering. phd thesis, department of computer and information science, university of pennsylvania, 1986. j. l. pollock. how to build a person. mit press, 1990. r. reichman. getting computers to talk like you and me. the mit press, cambridge, ma, 1985. m. richardson and p. domingos. markov logic networks. machine learning, 62:107–136, 2006. c. roberts. demonstratives as definites. in k. van deemter and r. kibble, editors, information sharing, pages 89–196. csli, 2002. m. rooth. a theory of focus interpretation. natural language semantics, 1:75–116, 1992. y. schabes. new parsing strategies for tree adjoining grammars. in proc. of the 12th international conference on computational linguistics (coling), 1988. d. schlangen, t. baumann, and m. atterer. incremental reference resolution: the task, metrics for evaluation, and a bayesian filtering model that is sensitive to disfluencies. in proc. of sigdial, london, uk, 2009. m. s. seidenberg, m. k. tanenhaus, j. leiman, and m. bienkowski. automatic access of the meanings of ambiguous words in context: some limitations of knowledge-based processing. cognitive psychology, 14:489–537, 1982. s. shieber and m. johnson. variations on incremental interpretation. journal of psycholinguistic research, 22(2):287–318, 1993. 276 incremental anaphora and reference through resource situations c. l. sidner. towards a computational theory of definite anaphora comprehension in english discourse. phd thesis, mit, 1979. m. j. spivey, m. k. tanenhaus, k. m. eberhard, and j. c. sedivy. eye movements and spoken language comprehension: effect of visual context on syntactic ambiguity resolution. cognitive psychology, 45:447–481, 2002. m. steedman. the syntactic process. mit press, 2001. m. stone. intention, interpretation and the computational structure of language. cognitive science, 28(5):781–809, 2004. s. c. stoness, j. tetreault, and j. allen. incremental parsing with reference interaction. in proc. acl workshop on incremental parsing, pages 18–25, barcelona, 2004. p. sturt and m. crocker. monotonic syntactic processing: a cross-linguistic study of attachment and reanalysis. language and cognitive processes, 11(5):449–494, 1996. d. a. swinney. lexical access during sentence comprehension: (re)consideration of context effects. journal of verbal learning and verbal behavior, 18:545–567, 1979. m. k. tanenhaus and j. c. trueswell. eye movements as a tool for bridging the language-as-product and language-as-action traditions. in j. c. trueswell and m. k. tanenhaus, editors, approaches to studying wordl-situated language use, pages 3–37. mit press, 2005. m. k. tanenhaus, m. spivey-knowlton, k. m. eberhard, and j. c. sedivy. integration of visual and linguistic information in spoken language comprehension. science, 268:1632–1634, june 1995. m. k. tanenhaus, c. g. chambers, and j. e. hanna. referential domains in spoken language comprehension: using eye movements to bridge the product and action traditions. in j. m. henderson and f. ferreira, editors, the interface of language, vision, and action: eye movements and the visual world, pages 279–317. psychology press, 2004. d. r. traum. a computational theory of grounding in natural language conversation. phd thesis, university of rochester, department of computer science, rochester, ny, july 1994. d. r. traum and e. a. hinkelman. conversation acts in task-oriented spoken dialogue. computational intelligence, 8(3), 1992. special issue on non-literal language. b. l. webber. a formal approach to discourse anaphora. garland, new york, 1979. b. l. webber. structure and ostension in the interpretation of discourse deixis. language and cognitive processes, 6(2):107–135, 1991. h. zender. situated production and understanding of verbal references to entities in large-scale space. phd thesis, university of saarbruecken, 2010. 277 d&d: dialogue act classification, insta...re | di eugenio | dialogue & discourse vasishthdrenhausdd.dvi dialogue and discourse 1(2) (2011) 59-82 doi: 10.5087/dad.2011.104 locality in german shravan vasishth vasishth@uni-potsdam.de department of linguistics university of potsdam potsdam, 14476, germany heiner drenhaus drenhaus@coli.uni-saarland.de computational linguistics and phonetics university of saarland saarbruecken, 66123, germany editor: david schlangen and hannes rieser abstract three experiments (self-paced reading, eyetracking and an erp study) show that in relative clauses, increasing the distance between the relativized noun and the relative-clause verb makes it more difficult to process the relative-clause verb (the so-called locality effect). this result is consistent with the predictions of several theories (gibson, 2000; lewis and vasishth, 2005), and contradicts the recent claim (levy, 2008) that in relative-clause structures increasing argument-verb distance makes processing easier at the verb. levy’s expectation-based account predicts that the expectation for a verb becomes sharper as distance is increased and therefore processing becomes easier at the verb. we argue that, in addition to expectation effects (which are seen in the eyetracking study in first-pass regression probability), processing load also increases with increasing distance. this contradicts levy’s claim that heightened expectation leads to lower processing cost. dependencyresolution cost and expectation-based facilitation are jointly responsible for determining processing cost. keywords: sentence processing; locality; surprisal; expectation-based sentence comprehension; eye-tracking; eeg; self-paced reading. 1. introduction a well-known claim in the psycholinguistic literature (chomsky, 1965; just and carpenter, 1980, 1992; gibson, 2000) states that completing a dependency between two linguistic units (such as a verb and an argument) is partly a function of the distance between them. an example is the selfpaced reading study by grodner and gibson (2005), which showed increasing reading time at the verb supervised as a function of the distance between the subject nurse and the verb: (1) a. the nurse supervised the administrator while. . . b. the nurse from the clinic supervised the administrator while. . . c. the nurse who was from the clinic supervised the administrator while. . . chomsky (1965, 13-14) was perhaps the first to propose that the reduced acceptability of sentences containing a “nesting of a long and complex element” arises from “decay of memory.” in related work, just and carpenter (1980, 1992) directly address dependency resolution in sentence c!2011 shravan vasishth and heiner drenhaus submitted 1/2010; accepted 12/2010; published online 5/2011 vasishth, and drenhaus comprehension in terms of memory retrieval (similar early approaches are the production-system based models of anderson et al. 1977). just and carpenter developed a model of integration that involved activation decay (as a side-effect of capacity limitations) as a key determinant of processing difficulty. for example, under the rubric of distance effects, they describe the constraints on dependency resolution as follows (just and carpenter, 1992, 133): the greater the distance between the two constituents to be related, the larger the probability of error and the longer the duration of the integration process. the explanation for the distance effect in terms of activation decay was taken a great deal further in the syntactic prediction locality theory or splt (see gibson (1998, 9) for a historical overview of the connection between decay and distance) and, more recently, the dependency locality theory or dlt gibson (2000). the dlt proposes (among other things) that the cognitive cost of assembling a dependent with a head is partly a function of the number of new intervening discourse referents that are introduced between the dependent and the head; see figure 1 for an example. in effect, the dlt discretizes the concept of activation decay in the dlt complexity metric (gibson, 2000, 103). the predictions of splt and dlt find quite good empirical support from online experiments involving english (e.g., gibson and thomas 1999, grodner and gibson 2005, warren and gibson 2005) and also chinese (hsiao and gibson, 2003). at least one offline study involving japanese is also consistent with the splt’s (the precursor of dlt) predictions (babyonyshev and gibson, 1999).             figure 1: a schematic illustration of dlt’s predictions for multiply embedded structures. integration costs are labeled along the arcs that define the argument-head dependencies, computed by counting the number of intervening discourse referents. another component of the theory is storage cost; these costs are presented under each verb for illustration. the storage costs is computed by counting the number of heads predicted at each point. as mentioned above, locality cost is characterized by the dlt in terms of the number of discourse referents intervening between the dependent and the head. one may ask: what is so special about the number of new discourse referents? why not count the number of intervening syntactic nodes, words, letters, syllables, etc.? the rationale within the dlt is that building discourse referents is computationally costly; independent evidence for this idea comes from studies showing that the accessibility of the intervening discourse referent (as defined by the accessibility hierarchy) can modulate retrieval difficulty (warren, 2001; warren and gibson, 2005). one interesting empirical problem is that, apart from the grodner and gibson (2005) results, locality does not seem to have much empirical support (indeed, jaeger et al. (2008) have recently 60 locality in german presented evidence that the locality constraint may not apply even in english, the language that has the most-attested instances of locality effects). konieczny (2000) presented an important counterexample from german to the locality hypothesis. in a self-paced reading study involving centerembedded relative clauses, he showed that increasing argument-head distance, analogous to example 1 above, resulted in faster reading time at the verb, not slower (as predicted by locality-based accounts). konieczny’s explanation for the result was that the strength of prediction for the upcoming verb increases if more intervening material is present between the dependent and the head (he calls this the anticipation hypothesis). recently, levy (2008) has proposed a related explanation for such antilocality effects. under this view, antilocality effects could be explained by assuming that the material intervening between the dependent and head could serve to sharpen the expectation for the upcoming verb. as levy (2008, 1144) puts it: “more preverbal dependents gives [sic] the comprehender more information with which to predict the final verbs identity and location, and comprehension should therefore be easier.” this expectation hypothesis is the suggested explanation for antilocality effects seen in german (konieczny, 2000) and hindi (vasishth and lewis, 2006). an interesting prediction of levy’s expectation-based account is that english should also show antilocality effects. indeed, levy quotes a self-paced reading study by jaeger and colleagues (also see the further studies in jaeger et al. 2008) which confirmed this prediction. jaeger et al. (2008) presented participants with sentences like 2. they found that reading time at the verb bought is faster as distance between the dependent player and the verb is increased. jaeger and colleagues argue that this speedup occurs because ”. . . end of [the relative clause (rc)] and hence presence of matrix verb becomes more probable after each additional pp in rc.” this is another example of how sharpened expectation of an upcoming verb results in faster reading time at the verb. (2) a. the player [rc that the coach met at 8 o’clock ] bought the house . . . b. the player [rc that the coach met by the river at 8 o’clock ] bought the house . . . c. the player [rc that the coach met near the gym by the river at 8 o’clock ] bought the house . . . given the recent findings from english of jaeger and colleagues, one might wonder whether locality can be dismissed altogether as a possible contraint on processing difficulty. however, there are at least two problems with dismissing the locality effect. first, van dyke and lewis (2003) present indirect evidence for locality. they conducted a self-paced study involving sentences such as 3. one factor was ambiguity (presence/absence of the sentential complement that), and another was distance between an argument (here, the noun student) and verb (was standing). (3) a. the assistant forgot that the student was standing in the hallway. b. the assistant forgot the student was standing in the hallway. c. the assistant forgot that the student who knew that the exam was important was standing in the hallway. d. the assistant forgot the student who knew that the exam was important was standing in the hallway. 61 vasishth, and drenhaus the ambiguity manipulation ensures that reanalysis takes place at was standing – the np student must be reanalyzed as the subject of a sentential complement rather than the object of forgot. the distance manipulation ensures that the reattachment of the np as subject of was standing is affected by locality. the reanalysis requires an integration between the verb and the argument, which is either near or distant from the verb. consequently, if a significantly greater reanalysis cost is observed in the intervening-items conditions 3c,d than in the non-intervening-items conditions 3a,b, this would be a locality effect, and it would be independent of spillover confounds because the comparison is no longer a direct one between conditions with differing regions preceding the critical verb. the interaction was in fact observed in the van dyke and lewis study (i.e., the difference between 3c,d and 3a,b was significant). this result suggests that distance (or decay) can adversely affect processing. the second problem with doing away with locality explanations is the fact that the experiment by grodner and gibson (2005) discussed above clearly shows locality effects; levy’s expectationbased account cannot explain this effect. here, it is legitimate to question the replicability of the grodner and gibson result; after all, if their finding cannot be reproduced, perhaps the data should be doubted, not the expectation-based account. however, bartek et al. (2010) present four replications of the grodner and gibson result using self-paced reading and eyetracking; this leads us to question the expectation-based account. to summarize the discussion so far, both locality and antilocality effects have been attested in english, but to our knowledge only antilocality effects have been seen in german and hindi. the goal of the present study was to determine whether locality effects can be observed in german in a signature syntactic configuration where levy’s expectation-based processing account predicts antilocality effects. specifically, we investigated structures like (4), where the expectation for a verb should necessarily get sharper as dependent-head distance increases. (4) a. die the mutter mother von of paula paula und and die the schwester sister von of sophie sophie gruessten greeted den the direktor, director den whom maria maria und and franziska franziska ignoriert ignored hatten. had ‘the mother of paula and the sister of sophie greeted the director whom maria and franziska had ignored.’ b. paula paula und and die the schwester sister von of sophie sophie gruessten greeted den the direktor, director den whom maria maria und and die the mutter mother von of franziska franziska ignoriert ignored hatten. had ‘paula and the sister of sophie greeted the director whom maria and the mother of franziska had ignored.’ c. paula paula und and sophie sophie gruessten greeted den the direktor, director den whom die the schwester sister von of maria maria und and die the mutter mother von of franziska franziska ignoriert ignored hatten. had ‘paula and the sister of sophie greeted the director whom maria and the mother of franziska had ignored.’ 62 locality in german here, once a relative clause begins after the argument direktor, the appearance of the verb in the relative-clause verb is guaranteed (as in the stimuli in the experiments by konieczny (2000). the prediction of the expectation-based account for these sentences is exactly the same as for the studies by konieczny’s (2000) study and jaeger et al’s (2008) studies, where increasing the length of the relative clause renders processing easier at the verb: delaying the appearance of such a clause-final verb should sharpen the expectation for a verb. by contrast, the locality-based explanations predict a slower reading time and greater processing difficulty at the verb as distance is increased. we present three experiments that evaluated these predictions. 2. experiment 1: self-paced reading 2.1 participants fifty-one students at the university of potsdam, all of them native speakers of german, took part in the study for course credit or payment. 2.2 method a self-paced reading comprehension experiment (just et al., 1982) was carried out in german at the university of potsdam, germany. thirty target sentences with three conditions each (see examples 4 above) were presented in a counterbalanced manner, with 104 filler sentences pseudo-randomly interspersed between the target sentences. the counterbalancing meant that each participant saw each sentence only once. the experiment was run on a macintosh computer using the linger software developed by doug rohde (http://tedlab.mit.edu/!dr/linger/). all items for this and other experiments reported in the paper are available from the first author. participants read the introduction to the experiment on the computer screen. in order to read each word of a sentence successively in a moving window display, they had to press the space bar. the word seen previously was masked. each word was shown separately. at the end of each sentence subjects had to answer yes-no comprehension questions in order to ensure that they would try to comprehend the sentences; no feedback was given as to whether the response was correct or not. one concern was that participants might develop a question-answering strategy without paying attention to the entire sentence; therefore, questions were designed to probe different argument-verb relations in the sentences. all stimulus items and accompanying questions, along with the expected correct answers, are available from the authors. 2.3 results figure 2 shows all the reading times at all positions; note that differences in reading time at regions preceding position 16 are not comparable because the words differ in each condition. we compared reading time at three regions of interest: position 16, the word preceding the critical word ignoriert, ‘ignored’; position 17, the critical word; and position 18, the word following the critical word (see 4). we defined no specific prediction for the pre-critical region, but we included it in the analysis because this was the first region that was comparable across the sentences. for the other two regions we predicted increased processing difficulty with increasing distance (position 18 was expected to show spillover effects possibly originating at the critical region, see mitchell (1984)). the mean rts and standard errors for the three regions analyzed are shown in table 1. 63 vasishth,anddrenhaus 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 400500600700800 w or d po sit io n reading time (ms) m ea n rt s b y w or d po sit io n ba se lin e n on −l oc al 1 n on −l oc al 2 fi gu re 2: m ea n re ad in g tim e (m s) by w or d po sit io n in ex pe rim en t1 (s el f-p ac ed re ad in g) ,w ith sta nd ar d er ro rb ar s. th is fig ur e sh ow sm ea ns an d sta nd ar d er ro rs fo rt he un re du ce d da ta se t; th e da ta an al ys is w as ca rri ed ou ta fte rr em ov in g ex tre m e va lu es ;s ee m ai n te xt fo r de ta ils .t he m ea ns fo rt he po sit io ns 16 –1 8 ar e sh ow n in ta bl e 1. 64 locality in german a linear mixed model was fit with items and participants as random intercepts, and helmert contrasts (venables and ripley, 1999) were defined to investigate the effect of the locality manipulation. helmert contrast coding was defined so that the first contrast involved comparing condition 4a with the average of 4b and 4c, and the second contrast compared 4b and 4c. this contrast coding has the advantage that a single statistical model gives us complete information about the relationship between the three conditions’ means (the alternative would have been to carry out three pairwise t-tests, which would reduce the chances of detecting a true effect; standard statistics textbooks discuss this point, but also see vasishth and broe (2011) for a detailed discussion). in the model fits shown below, the coefficients are the differences between the means compared. in order to reduce non-normality in the residuals, reading times less than 80 ms and greater than 2000 ms were removed from the data (this resulted in the removal of 0.05% of the data).1 the data analysis was carried out on log-transformed reading times for reasons discussed in baayen and milin (2010). as summarized in table 2, helmert contrasts showed that the average of conditions (b) and (c) was slower than (a) in all three regions; and that condition (c) was slower than (b) in the pre-critical and post-critical regions. table 1: means (ms) and standard errors for the three regions analyzed. these are means with extreme values removed; see text for details. condition pre-critical critical post-critical a 444 (10) 514 (14) 526 (13) b 462 (11) 538 (14) 539 (15) c 497 (13) 542 (14) 597 (19) table 2: results of linear mixed model fit for experiment 1. items and participants were crossed random factors. the asterisk represents statistical significance at ! = 0.05. region comparison coef. se t-value pre-critical a vs b,c 0.0702 0.0204 3.4 * b vs c 0.0628 0.0235 2.7 * critical a vc b,c 0.0526 0.0219 2.4 * b vs c 0.0110 0.0252 0.4 post-critical a vs b,c 0.05162 0.0262 2.0 * b vs c 0.0816 0.0302 2.7 * 1. retaining all data points in the data analysis resulted in no effect at the critical region (the verb) for the comparison a vs b,c (t=1.8), and no effect in the post-critical region for the comparison a vs b, c (t=1.5). 65 vasishth, and drenhaus 2.4 discussion we found increased reading times in the critical and post-critical regions as distance between the verb (ignoriert) and the argument (direktor) is increased. interestingly, we also see increased reading time as a function of distance in the region preceding the verb; we discuss one possible implication of this result below. the data clearly find support for the idea that increasing head-dependent distance increases processing load at the head. thus, the result is consistent with locality predictions of models like gibson’s dlt and lewis and vasishth’s act-r cue-based retrieval account. it is problematic for the expectation-based processing proposal by levy, that increasing head-dependent distance in relative clauses renders the head more predictable and therefore easier to process. regarding the locality pattern seen in the pre-critical region, one possibility (a speculation at this point) is that the verb phrase (which is predicted from the moment that the relative clause begins) is already retrieved when the noun preceding the verb is processed; this assumption is consistent with levy’s expectation-based account. once retrieved, the integration phase (retrieving the subject direktor) could already begin while the noun is being processed. assuming, as the dlt and the cuebased retrieval theory do, that this retrieval is costlier in the non-local conditions, the locality effect could appear even before the verb is processed. clearly, this proposal needs further investigation. focusing now on only the regions for which we had well-defined a priori predictions, although we found clear evidence for locality accounts, one question arises: can both locality accounts and the expectation-based account co-exist? after all, as discussed above, evidence exists for both claims, and boston et al. (2010) have shown that both the cue-based retrieval account of lewis and vasishth (2005) and the expectation-based account can independently explain reading data. we return to this question later in the paper. our next step is to attempt to replicate the above result using a different method. replication is vital in order to verify the robustness of the finding, and a different method is necessary at the very least because it is important to determine whether the result in experiment is method-dependent or not. in this context, eyetracking is a potentially interesting method because, unlike self-paced reading, both first-pass and non-first pass measures can be investigated, and the reading task is more natural than self-paced reading. 3. experiment 2: eyetracking 3.1 participants thirty-five native-speaker students at the university of potsdam took part in the experiment, either for course credit or for payment. 3.2 method participants were seated approximately 50 cm from a 19-inch color lcd monitor with 1024 " 768 pixel resolution; twenty-eight pixels equaled about one degree of visual angle. the eyetracker used was an sr research eyelink 1000 eyetracker running at 500 hz sampling rate with 0.01 degree tracking resolution and a gaze position accuracy of < 0.5 degree. although viewing was binocular, only data from the right eye was used in analyses. for stability, participants were asked to place their head on a chin-rest and a forehead-rest. participants were instructed to avoid large shifts in position throughout the experiment. responses were recorded by a 7-button 66 locality in german microsoft sidewinder game pad. the presentation of the materials and the recording of the responses was controlled by a notebook windows pc running eyetrack, which was interfaced with the eyetracker. each stimulus sentence was presented on two lines with the line break placed consistently after the comma (direktor,). the space between the two lines was 221 pixels. this space was introduced in order to prevent participants from previewing the second line of the target sentence parafoveally while reading the first line. each participant was randomly assigned one of three stimuli lists which comprised different item-condition combinations according to a latin square. the trials per session were randomized individually per participant. at the start of the experiment, the experimenter performed the standard eyelink calibration procedure, which involves participants looking at a grid of thirteen fixation targets in random succession. calibration was repeated during the session if the experimenter noticed that measurement accuracy was poor (e.g., after strong head movements or a change in the participant’s posture). each trial consisted of the following steps. first, a fixation target appeared 30 pixels distant from the left screen border, the position where the text would be aligned. the stimulus was presented only after the subject fixated on this target. after the participant had finished reading the sentence, they fixated a dot in the lower right corner of the screen and simultaneously pressed a specific button on the game pad. this triggered the presentation of a simple comprehension question which the participant had to answer either with the left (‘no’) or the right trigger button (‘yes’). the text stimuli were presented using a courier new font, printed in black on a white background. the characters (including spaces) were all the same width, approximately 9 pixels or 0.32 degrees of visual angle. the presentation software automatically recorded the coordinates of rectangular interest areas around each word; fixations were then associated with words according to whether their coordinates fell within a word’s interest area. the upper and lower boundaries were 15 pixels above and below the top and bottom of the line of text. the line of text was 17 pixels high. 3.3 materials we used sixty items, each with the three conditions shown in (4). in addition, forty filler sentences were pseudo-randomly interspersed with the target items. we did not use a larger number of filler items because in our experience more than 100 items results in significant participant fatigue. 3.4 dependent measures in eyetracking as mentioned earlier, eyetracking data has the great advantage that it provides a record of both first-pass and second pass reading; however, the richness of data is not necessarily an improvement over simpler methods such as self-paced reading; a central problem with calculating statistics on all available eyetracking dependent measures (as is often done in psycholinguistics) is that they tend to be highly correlated, making much of the statistical calculation uninformative. for completeness, in appendix a we provide summary statistics of the major dependent measures, where we also include the definitions of the various dependent measures used. since we were asked by a reviewer to compute statistics for all dependent measures, we did so, but we found no effects except in dependent measures related to re-reading. in this section we provide only two sets of critical analyses, which relate to first-pass processing (first-pass regression probability) and second-pass processing (re-reading probability and re-reading 67 vasishth, and drenhaus time). first-pass regression probability at any word n is the probability of the eye moving leftward to a preceding word n # k (where k $ 1), after the word n has been fixated at least once. rereading probability for a word n is the probability of revisiting that word after having having made a first-pass through that word (in other words, it is total reading time minus first-pass reading time; see appendix a for definitions). re-reading time is the amount of time spent revisiting a word after having having made a first-pass through that word. as also discussed in appendix a, there are two ways to define re-reading time; one (the standard approach) includes zeroes (sturt, 2003) and the other only looks at pure second-pass time, excluding zero re-reading time (vasishth et al., 2010). we will refer to the former as re-reading time, and the latter as second-pass time. re-reading time, as defined above, is considered the standard definition presumably because it also includes non-re-reading, i.e., the full dataset. it appears that second-pass time (as defined above) is never presented in papers; presumably, researchers believe that this measure yields no information about processing difficulty. this assumption is probably well-founded when the proportion cases where re-reading occurred is small (if nothing else, the low power that would result would make null results meaningless), but it makes little sense when a large proportion of re-reading is present in the data. two published papers where re-reading proportion was high are vasishth et al. (2008) and vasishth et al. (2010). in the latter work, the authors systematically compared secondpass time as defined above with self-paced reading data from identical materials across german and english. these studies showed that spr reading time and second-pass time had comparable results; in other words, it may be a mistake to ignore second-pass time as an informative dependent measure in eyetracking. for this reason, we include second-pass time in the analyses below. in the present dataset, re-reading proportion was low, and therefore power is low. nevertheless, as shown below, the results are informative. re-reading probability is to our knowledge not used in the literature, but it is highly correlated to re-reading time (trivially so) and is useful in the present paper because it allows us to compare probabilities in both first-pass (regression probability) and second-pass, rather than probabilities in first-pass and reading times in second pass. the broad conclusions remain unchanged when we use re-reading time instead of re-reading probability, as shown in the next section. 3.5 results table 3 shows first-pass regression probability, table 4 shows re-reading probability, and table 5 shows re-reading time and second-pass time. interestingly, first-pass regression probability shows lower regression proportions in conditions (c) versus (b) at the pre-critical region as well as the critical region; this is consistent with the expectation-based account, but not with locality based accounts. however, as shown in table 4, second-pass or re-reading probability shows evidence consistent with locality-based predictions: re-reading probabilities are significantly higher in the post-critical region in the comparison (b,c) versus (a) and the comparison (b) versus (c). in addition, re-reading probability is marginally higher in the (b,c) versus (a) comparison in the pre-critical and critical regions, with p-values somewhat greater than 0.05. finally, as shown in table 5, log re-reading time showed locality effects in the pre-critical region in the comparison involving (b,c) versus (a), and in the post-critical region for the comparison involving (c) versus (b); in the post-critical region the (b,c) versus (a) comparison was essentially significant, with t=1.98 (the distinction between a t-value of 1.98 and the critical t-value of 2 is 68 locality in german table 3: analyses for first-pass regression probability (experiment 2). region contrast coef. se z-score p-value pre-critical a vs b,c -0.0624 0.2750 -0.23 0.82 b vs c -1.0574 0.3251 -3.25 <0.01 * critical a vs b,c -0.0813 0.2392 -0.34 0.734 b vs c -0.5074 0.2827 -1.79 0.073 post-critical a vs b,c -0.0326 0.2481 -0.13 0.90 b vs c 0.0345 0.2881 0.12 0.90 table 4: analyses for re-reading probability (experiment 2). region contrast coef. se z-score p-value pre-critical a vs b,c 0.331 0.190 1.75 0.08 b vs c -0.321 0.206 -1.56 0.12 critical a vs b,c 0.2914 0.1632 1.78 0.074 b vs c 0.2841 0.1828 1.55 0.120 post-critical a vs b,c 0.3509 0.1702 2.06 0.040 * b vs c 0.4627 0.1875 2.47 0.014 * small enough to consider this effect significant). second-pass reading time showed a locality effect in the pre-critical region in the comparison (c) versus (b); this was in spite of the low power that resulted from removing 82.7% of the cases where no re-reading occurred (in the critical region 69.3% involved no re-reading and in the post-critical region 76.3%). 3.6 discussion since re-reading probability and re-reading time deliver essentially the same results, we focus our discussion on the former. the central findings in the eyetracking study are that (i) first-pass regression probability shows a facilitation in processing with increasing distance, but only in the pre-critical and critical regions (in the latter, the effect is marginal, p=0.07); and re-reading probability shows locality cost at all three regions, with marginal effects in the pre-critical (p=0.08) and critical regions (p=0.07), and a statistically significant effect in the post-critical region. thus, in experiment 2, we find evidence favoring the expectation-based account in a first-pass measure, and evidence favoring locality accounts in a measure that includes second pass. this result is in harmony with the idea, proposed in the discussion section of experiment 1, that both expectation-based facilitation and locality cost play a role in determining processing cost but that these two factors operate at different stages of processing. specifically, at the region (the noun) preceding the verb, it is plausible that the parser already retrieves the verb-phrase as soon as the noun is processed, and that this process completes faster in the long-distance conditions due to the higher expectation for the upcoming verb phrase—this is completely in line with levy’s (2008) proposal. once the verb phrase is retrieved, a dependency must be established between the subject and the verb phrase; here, the dependency completion cost is affected by decay. since this process occurs after the verb phrase is retrieved, the locality cost expresses itself in second-pass reading 69 vasishth,anddrenhaus ta bl e 5: a na ly se sf or lo g re -re ad in g tim e an d lo g se co nd -p as s tim e (e xp er im en t2 ). lo g rr ts lo g se co nd -p as st im e re gi on co nt ra st co ef . se t-v al ue co ef . se t-v al ue pr ecr iti ca l a vs b, c 0. 26 43 0. 13 17 2. 01 * -0 .0 11 55 0. 08 71 -0 .1 b vs c -0 .2 31 4 0. 15 21 -1 .5 2 0. 16 06 0. 07 98 2. 0 * cr iti ca l a vs b, c 0. 27 50 5 0. 15 19 1. 81 0. 01 82 0. 06 98 0. 3 b vs c 0. 27 54 3 0. 17 53 1. 57 -0 .0 67 0 0. 07 67 -0 .9 po stcr iti ca l a vs b, c 0. 28 56 0. 14 41 1. 98 * -0 .0 74 9 0. 06 93 -1 .1 b vs c 0. 42 27 0. 16 64 2. 54 * 0. 05 19 0. 07 48 0. 7 70 locality in german measures. a similar proposal was made by sommerfeld et al. (2007), based on eyetracking and erp data from experiments involving locality manipulations (the design in those studies was different from the present paper’s). we turn now to a third study, where we use the eeg methodology to replicate these results in a different setting; as discussed above, our motivation for doing the eeg study was to determine whether the results are robust and whether they can be replicated in methods other than self-paced reading and eyetracking. 4. experiment 3: event-related potentials 4.1 participants twenty-two undergraduate students from the university of potsdam participated in the erp study. 4.2 materials there were a total of 360 sentences (180 critical items and 180 of unrelated filler sentences). the sentences were split in three versions to avoid repetition of lexical material. each subjects saw 60 critical sentences (see example 4) intermixed with 60 unrelated sentences which makes a total of 120 sentences. the sentences were presented in a pseudo-randomized order. 4.3 procedure twelve training sentences (4 in each of the critical conditions, see above) were presented to the participants. after this training set, the 60 critical sentences and the 60 unrelated filler sentences were randomly presented in the center of a 17” computer screen, with 400 ms (plus 100 ms interstimulus interval) for each word. 500 ms after the last word of each sentence a comprehension question was presented on the screen for 2500 ms. the questions were designed to have 50% yes answers and 50% no answers, respectively. the task for the subjects was to answer the questions within a maximal interval of 3000 ms by pressing one of two buttons. 1000 ms after their response, the next trial began. the eeg was recorded by means of 25 ag/agcl electrodes with a sampling rate of 250hz (impedances < 5k!) and were referenced to the left mastoid (re-referenced to linked mastoids offline). the horizontal electro-oculogram (eog) was monitored with two electrodes placed at the outer canthus of each eye and the vertical eog with two electrodes above and below the right eye. 4.4 data preprocessing only the trials that did not have artifacts were selected for the erp analysis. the data were filtered with 0.2 hz (high pass) to compensate for drifts. single subject averages were computed in a 1000 ms window relative to the onset of the critical item (ignoriert ‘ignored’) and aligned to a 200 ms pre-stimulus baseline. one time window was analyzed: 300-500 ms. 4.5 results the erp patterns from the onset of the critical item (the verb ignoriert ’ignored’, onset at 0 ms) up to 1000 ms thereafter are displayed for the electrodes fc5, cp5 and c3, in figure 3, which shows the grand average erps for the control condition compared to the two more complex conditions at one 71 vasishth, and drenhaus electrode. visual inspection shows that the more complex conditions (4b,c) are more negative-going compared to the control condition (4a). table 6: results of linear mixed model fit for experiment 3. participants were random factors; the time window is 300-500 ms, and the electrodes are fc5, c3, cp5. random slopes by subject were fit for the two helmert contrasts described in the text. contrast coefficient se t-value a bs b,c -1.184 0.373 -3.17 * b vs c 0.596 0.431 1.38 as summarized in table 6, for three electrodes (fc5, c3, and cp5), we see a significant negativity for the contrast (b,c) vs (a). the coefficient in table 6 is negative because of that fact that the erp response is negative-going; by contrast, in the spr and eyetracking studies, the coefficients had a positive sign because more complex conditions tended to show longer reading times and higher re-reading probability. the erp pattern and topology is reminiscent of the left anterior negativity, which has been argued to be triggered in long-distance dependency resolution (kluender and kutas, 1993a). if only the electrode fc5 is analyzed, we see only a marginal negativity (t=-1.9); in addition, no effects were seen in other analyses involving the three frontal (f3, fz, f4), central (c3, cz, c4), and posterior (p3, pz, p4) electrodes. 4.6 discussion in the erp study, we find a negative-going potential in the non-local versus local conditions; this pattern is seen in the 300–500 ms window from the onset of the verb ignoriert, ignored. the negativity has a fronto-central distribution; the topology is not dissimilar from the left anterior negativity, which has been seen in long-distance wh-dependency resolution (kluender and kutas, 1993b,a). in these studies it was interpreted as indexing a “looking back” process which was triggered by the attempt to integrate syntactic information with material occurring earlier in the structure. additionally, other researchers have found a negativity on the verb when the parser was checking for an appropriate subject and was trying to integrate this information (osterhout and holcomb, 1992; king and kutas, 1995; vos et al., 2001). thus, consistent with the self-paced reading study and the eyetracking experiment, the erp study also provides evidence for greater integration cost as argument-head distance increases. 5. general discussion we begin by discussing the evidence from the three experiments for the locality hypothesis, and then we turn to the evidence in the eyetracking data for expectation-based processing. the evidence for locality is summarized across the three methods in table 7. it is clear from table 7 that we see greater processing difficulty in the non-local conditions mostly at the critical and post-critical regions; in the spr study we see a slowdown in the non-local conditions in the pre-critical region as well. the latter effect could of course be a type i error; 72 locality in german   figure 3: erp voltage averages for the three conditions: control condition (solid), the more complex condition (dotted), and the most complex condition (dashed) at the left fronto-central electrode fc5. time onset of the critical stimulus (the verb) at 0 s. negativity is plotted upwards. for presentation purposes only, erps were filtered off-line with 8 hz low pass. 73 vasishth, and drenhaus table 7: a summary of the locality effects found across the three experiments. only statistically significant effects are shown. experiment 1 (spr) region comparison coef. se t-value pre-critical a vs b,c 0.0702 0.0204 3.4 * b vs c 0.0628 0.0235 2.7 * critical a vc b,c 0.0526 0.0219 2.4 * post-critical a vs b,c 0.0516 0.0262 2.0 * b vs c 0.0816 0.0302 2.7 * experiment 2 (eyetracking) contrast coefficient se z-value post-critical a vs b,c 0.3509 0.1702 2.06 * b vs c 0.4627 0.1875 2.47 * experiment 3 (erp) contrast coefficient se t-value critical a bs b,c -1.184 0.373 -3.17 * indeed, we did not expect an effect before the critical region in spr. however, as mentioned earlier, if we assume that the effect is actually present, a plausible explanation would be that the parser may have retrieved the predicted verb phrase as soon as the pre-critical region is processed, and a retrieval of the grammatical subject may have been started as soon as the parser retrieves the verb phrase (or even while it is retrieving the verb phrase—nothing speaks against parallel execution of the prediction-integration steps). such a mechanism could account for the slowdown seen in the pre-critical region. taken together, these findings deliver considerable evidence consistent with the locality hypothesis (gibson, 2000; lewis and vasishth, 2005) and go directly against the expectation-based account of levy (2008). specifically, the statement by levy (2008, 1144) that “more preverbal dependents gives [sic] the comprehender more information with which to predict the final verbs identity and location, and comprehension should therefore be easier” is incorrect. it may well be true that the interveners give more information about the final verb’s identity and location, but it is not true that comprehension is necessarily easier: the costs associated with integration processes cannot be avoided.2 why did our materials show a locality effect when previous work (e.g., konieczny (2000)) has consistently shown antilocality effects? we believe that this comes from a particular property of our materials. all the items contain multiple instances of proper names in all three conditions; this makes the retrieval of direktor (the argument of the verb) more difficult across all conditions due to generally high interference lewis and vasishth (2005); van dyke (2007). the presence of multiple candidate noun phrases that could be a subject of the verb ignoriert could result in this generally increased difficulty in retrieving the correct noun at the verb. thus, in these stimuli, integration cost dominates over expectation-based facilitation. 2. for an alternative explanation for locality effects in terms of good enough processing, see christianson and luke (2010). our data are consistent with this view as well; the central idea is that in long-distance dependencies the parser builds only partial trees, not complete structures as gibson (2000) and lewis and vasishth (2005) assume. this may well be the correct way to characterize locality effects. 74 locality in german note that we do not deny a role for expectation-based facilitation of the type that levy advocates. in fact, the eyetracking record provides direct evidence for the expectation-based view. specifically, in first-pass regression probability we see, in the pre-critical region a difference between condition (c) and (b), such that regression probability is lower in the non-local condition. we have independent evidence that regression probability has widely been acknowledged to index increased processing load. for example, boston et al. (2008) and boston et al. (2010) show, using the potsdam sentence eyetracking corpus (kliegl et al., 2006) and two distinct types of grammars (probabilistic context free grammars and dependency grammars) derived from a treebank corpus, that both expectation cost (quantified as surprisal) and retrieval cost independently explain reading times. if the above explanation is correct, this suggests that expectation-based facilitation and locality interact in an interesting manner: expectation plays a dominant role only when working memory load is relatively low. a related idea has been suggested independently by gibson (2007); he suggested that expectation-based costs may dominate in relatively simple structures and/or more frequently occurring constructions, and distance-based costs may dominate in more complex and/or less frequently occurring constructions. in sum, the present work limits the extent to which the expectation-based account can explain processing difficulty to very specific situations. this is consistent with the remark by levy (2008) that expectation-based accounts are not an alternative account of dependency-resolution phenomena, but rather an additional constraint. one consequence of treating integration cost and expectation-based facilitation as two (largely) orthogonal factors is that one can dominate over the other depending on factors such as working memory load. further support for this view comes from recent work by boston et al. (2010), as discussed above. the contribution of this paper is to provide novel empirical evidence from three methods demonstrating how integration cost can be detected independent of any advantage due to sharpened expectations for upcoming word-types. in this appendix, we provide detailed summary statistics for experiment 2. we begin by providing definitions of the most common dependent measures used in reading research. first fixation duration (ffd) is the first fixation during the first pass, and has been argued to reflect lexical access costs inhoff (1984). gaze duration or first pass reading time (fprt) is the summed duration of all the contiguous fixations in a region before it is exited to a preceding or subsequent word; inhoff (1984) has suggested that fprt reflects text integration processes, although rayner and pollatsek (1987) argue that ffd and fprt may reflect similar processes and could depend on the speed of the cognitive process. first-pass regression probability is the probability of the eye moving leftward from a currently fixated word to a preceding word; it has been argued to index increased processing load. right-bounded reading time (rbrt) is the summed duration of all the fixations that fall within a region of interest before it is exited to a word downstream; it includes fixations occurring after regressive eye movements from the region, but does not include any regressive fixations on regions outside the region of interest. rbrt may reflect a mix of late and early processes, since subsumes first-fixation durations. re-reading time (rrt) is the sum of all fixations at a word that occurred after first pass; rrt has been assumed to reflect the costs of late (integration) processes (gordon et al., 2006, 1308). most researchers (e.g., birch and rayner (1997), sturt (2003)) include zero re-reading times in their calculation of re-reading time; this appears to be the standard method in eyetracking research. 75 vasishth, and drenhaus a minority (vasishth et al. (2010, 2008), and possibly also gordon et al. (2006)) exclude zero reading times in re-reading time. the approach, although not standard, has been shown to yield results comparable to self-paced reading times in english and german (vasishth et al., 2010) (to our knowledge there has been no similar cross-method comparison of the ‘standard’ re-reading time, i.e., re-reading time including zeroes). in the present paper, we use a binary version of re-reading time, re-reading probability. this is a binary response variable, 1 for re-reading, 0 for no re-reading. this measure is essentially identical to using the ‘standard’ version of re-reading time except that the response is a binomial variable. using a binomial version of re-reading has the advantage in the present study that we can compare two probabilities rather than probabilities and reading time: regression probability (a first-pass measure) and re-reading probability (a second-pass measure). another measure that may be related to late processing is regression path duration, which is the sum of all fixations from the first fixation on the region of interest up to, but excluding, the first fixation downstream from the region of interest. finally, total reading time (trt) is the sum of all fixations on a word. table 8: re-reading probability. word pos (a) (b) (c) 1 0.06 (0.01) 0.26 (0.02) 0.28 (0.02) 2 0.38 (0.03) 0.11 (0.02) 0.09 (0.02) 3 0.13 (0.02) 0.21 (0.02) 0.44 (0.03) 4 0.32 (0.02) 0.44 (0.03) 0.64 (0.03) 5 0.05 (0.01) 0.15 (0.02) 0.25 (0.02) 6 0.12 (0.02) 0.42 (0.03) 0.55 (0.03) 7 0.36 (0.03) 0.56 (0.03) 0.11 (0.02) 8 0.07 (0.01) 0.24 (0.02) 0.25 (0.02) 9 0.35 (0.03) 0.49 (0.03) 0.37 (0.03) 10 0.52 (0.03) 0.07 (0.01) 0.11 (0.02) 11 0.2 (0.02) 0.29 (0.02) 0.26 (0.02) 12 0.42 (0.03) 0.07 (0.01) 0.02 (0.01) 13 0.09 (0.01) 0.13 (0.02) 0.1 (0.02) 14 0.26 (0.02) 0.33 (0.03) 0.31 (0.02) 15 0.06 (0.01) 0.1 (0.02) 0.04 (0.01) 16 0.22 (0.02) 0.28 (0.02) 0.24 (0.02) 17 0.29 (0.02) 0.31 (0.02) 0.36 (0.03) 18 0.22 (0.02) 0.24 (0.02) 0.31 (0.02) 76 locality in german table 9: first-fixation duration. word pos (a) (b) (c) 1 44 (5) 116 (7) 110 (6) 2 164 (5) 48 (5) 49 (5) 3 78 (6) 97 (6) 203 (7) 4 195 (7) 194 (6) 243 (6) 5 30 (4) 71 (6) 99 (7) 6 109 (7) 193 (7) 211 (6) 7 198 (6) 249 (6) 64 (6) 8 63 (6) 106 (7) 104 (7) 9 192 (7) 230 (6) 167 (6) 10 241 (6) 55 (6) 81 (6) 11 108 (8) 140 (6) 185 (7) 12 227 (7) 44 (5) 29 (4) 13 56 (6) 101 (7) 102 (7) 14 128 (6) 181 (6) 169 (6) 15 39 (4) 79 (6) 53 (5) 16 170 (7) 170 (6) 165 (6) 17 222 (5) 238 (6) 232 (6) 18 188 (6) 195 (6) 187 (6) table 10: first-pass reading time. word pos (a) (b) (c) 1 86 (7) 209 (14) 200 (10) 2 234 (11) 74 (5) 76 (5) 3 109 (7) 144 (7) 254 (9) 4 244 (9) 292 (12) 361 (12) 5 48 (5) 103 (7) 154 (7) 6 141 (8) 255 (10) 409 (19) 7 307 (13) 351 (11) 136 (9) 8 87 (6) 168 (8) 180 (9) 9 246 (9) 391 (16) 261 (12) 10 353 (11) 154 (10) 121 (7) 11 152 (8) 232 (10) 229 (9) 12 394 (16) 63 (5) 46 (5) 13 147 (9) 133 (8) 135 (8) 14 199 (8) 235 (10) 208 (8) 15 56 (5) 109 (7) 76 (6) 16 207 (8) 203 (8) 204 (8) 17 288 (9) 297 (8) 291 (8) 18 211 (6) 220 (7) 216 (7) 77 vasishth, and drenhaus table 11: regression path duration. word pos (a) (b) (c) 1 86 (7) 209 (14) 200 (10) 2 291 (14) 105 (9) 113 (9) 3 136 (10) 171 (11) 337 (18) 4 325 (19) 389 (21) 458 (22) 5 78 (10) 148 (15) 239 (15) 6 161 (10) 369 (24) 871 (46) 7 367 (17) 413 (17) 138 (10) 8 115 (11) 274 (27) 248 (13) 9 311 (16) 931 (51) 361 (16) 10 423 (19) 170 (18) 157 (13) 11 236 (17) 364 (18) 280 (13) 12 815 (44) 93 (12) 60 (7) 13 152 (11) 175 (13) 171 (15) 14 304 (14) 323 (24) 277 (20) 15 92 (10) 142 (12) 100 (10) 16 270 (22) 265 (15) 242 (14) 17 335 (17) 363 (18) 331 (15) 18 269 (17) 290 (20) 280 (19) table 12: re-reading time. word pos (a) (b) (c) 1 19 (5) 132 (16) 119 (15) 2 167 (16) 32 (6) 23 (5) 3 36 (7) 89 (14) 217 (20) 4 166 (18) 267 (25) 362 (22) 5 13 (4) 43 (6) 82 (10) 6 40 (8) 209 (20) 292 (22) 7 196 (21) 307 (22) 37 (7) 8 18 (4) 73 (9) 75 (10) 9 148 (15) 260 (20) 162 (18) 10 268 (21) 25 (5) 26 (5) 11 62 (8) 119 (15) 104 (12) 12 222 (21) 15 (3) 4 (2) 13 25 (5) 33 (5) 32 (7) 14 100 (14) 132 (14) 129 (14) 15 14 (3) 24 (4) 10 (3) 16 75 (10) 113 (13) 100 (12) 17 114 (15) 123 (14) 131 (12) 18 73 (9) 74 (9) 99 (11) 78 locality in german table 13: total reading time. word pos (a) (b) (c) 1 105 (10) 341 (24) 318 (20) 2 401 (20) 106 (9) 99 (8) 3 145 (11) 233 (17) 471 (24) 4 410 (20) 559 (29) 723 (26) 5 60 (7) 146 (10) 237 (14) 6 182 (12) 464 (23) 700 (29) 7 503 (25) 658 (25) 172 (14) 8 105 (8) 241 (13) 255 (14) 9 394 (18) 652 (25) 423 (22) 10 621 (24) 179 (12) 147 (10) 11 214 (13) 351 (18) 333 (16) 12 616 (26) 78 (7) 50 (5) 13 171 (12) 166 (10) 166 (11) 14 298 (16) 367 (18) 336 (17) 15 69 (7) 133 (9) 85 (7) 16 283 (14) 316 (17) 303 (16) 17 403 (18) 420 (16) 422 (15) 18 284 (12) 294 (12) 315 (14) 79 vasishth, and drenhaus references j. r. anderson, p. kline, and c. lewis. a production system model for language processing. in m. just and p. carpenter, editors, cognitive processes in comprehension. lawrence erlbaum associates, hillsdale, nj, 1977. r.h. baayen and p. milin. analyzing reaction times. to appear in international journal of psychological research, 2010. m. babyonyshev and e. gibson. the complexity of nested structures in japanese. language, 75(3): 423–450, 1999. b. bartek, r. l. lewis, s. vasishth, and m. smith. in search of on-line locality effects in sentence comprehension. submitted, 2010. s. l. birch and k. rayner. linguistic focus affects eye movements during reading. memory and cognition, 25(5):653–60, 1997. m. f. boston, j. t. hale, u. patil, r. kliegl, and s. vasishth. parsing costs as predictors of reading difficulty: an evaluation using the potsdam sentence corpus. journal of eye movement research, 2(1):1–12, 2008. marisa f. boston, john t. hale, shravan vasishth, and reinhold kliegl. parallelism and syntactic processes in reading difficulty. language and cognitive processes, 2010. n. chomsky. aspects of the theory of syntax. mit press, cambridge, ma, 1965. k. christianson and s. g. luke. the ham stands alone: a good enough processing account of locality and local coherence. in revision, 2010. e. gibson. locality and anti-locality effects in sentence comprehension. workshop on processing head-final languages, 2007. e. gibson. dependency locality theory: a distance-based theory of linguistic complexity. in alec marantz, yasushi miyashita, and wayne o’neil, editors, image, language, brain: papers from the first mind articulation project symposium. mit press, cambridge, ma, 2000. e. gibson. linguistic complexity: locality of syntactic dependencies. cognition, 68:1–76, 1998. e. gibson and j. thomas. memory limitations and structural forgetting: the perception of complex ungrammatical sentences as grammatical. language and cognitive processes, 14(3):225–248, 1999. p. c. gordon, r. hendrick, m. johnson, and y. lee. similarity-based interference during language comprehension: evidence from eye tracking during reading. journal of experimental psychology: learning memory and cognition, 32(6):1304–1321, 2006. d. grodner and e. gibson. consequences of the serial nature of linguistic input. cognitive science, 29:261–290, 2005. f. hsiao and e. gibson. processing relative clauses in chinese. cognition, 90:3–27, 2003. 80 locality in german a. w. inhoff. two stages of word processing during eye fixations in the reading of prose. journal of verbal learning and verbal behavior, 23(5):612–624, 1984. f. t. jaeger, e. fedorenko, p. hofmeister, and e. gibson. expectation-based syntactic processing: antilocality outside of head-final languages. in cuny sentence processing conference, north carolina, 2008. m.a. just and p. a. carpenter. a capacity theory of comprehension: individual differences in working memory. psychological review, 99(1):122–149, 1992. m.a. just and p.a. carpenter. a theory of reading: from eye fixations to comprehension. psychological review, 87(4):329–354, 1980. m.a. just, p. a. carpenter, and j. d. woolley. paradigms and processes in reading comprehension. journal of experimental psychology: general, 111(2):228–238, 1982. j.w. king and m. kutas. who did what and when? using word-and clause-level erps to monitor working memory usage in reading. journal of cognitive neuroscience, 7(3):376–395, 1995. r. kliegl, a. nuthmann, and r. engbert. tracking the mind during reading: the influence of past, present, and future words on fixation durations. journal of experimental psychology: general, 135:12–35, 2006. r. kluender and m. kutas. bridging the gap: evidence from erps on the processing of unbounded dependencies. journal of cognitive neuroscience, 5(2):196–214, 1993a. r. kluender and m. kutas. subjacency as a processing phenomenon. language and cognitive processes, 8(4):573–633, 1993b. l. konieczny. locality and parsing complexity. journal of psycholinguistic research, 29(6):627– 645, 2000. r. levy. expectation-based syntactic comprehension. cognition, 106:1126–1177, 2008. r. l. lewis and s. vasishth. an activation-based model of sentence processing as skilled memory retrieval. cognitive science, 29:1–45, may 2005. d. c. mitchell. an evaluation of subject-paced reading tasks and other methods of investigating immediate processes in reading. in d. e. kieras and m.a. just, editors, new methods in reading comprehension research. erlbaum, hillsdale, n.j., 1984. l. osterhout and p.j. holcomb. event-related brain potentials elicited by syntactic anomaly. journal of memory and language, 31(6):785–806, 1992. k. rayner and a. pollatsek. eye movements in reading: a tutorial review. attention and performance xii: the psychology of reading, pages 327–362, 1987. e. sommerfeld, s. vasishth, p. logačev, m. baumann, and h. drenhaus. a two-phase model of integration processes in sentence parsing: locality and antilocality effects in german. in proceedings of the cuny sentence processing conference, la jolla, ca, 2007. 81 vasishth, and drenhaus p. sturt. the time-course of the application of binding constraints in reference resolution. journal of memory and language, 48:542–562, 2003. j. van dyke. interference effects from grammatically unavailable constituents during sentence processing. journal of experimental psychology: learning memory and cognition, 33(2):407– 30, 2007. j. van dyke and r. l. lewis. distinguishing effects of structure and decay on attachment and repair: a cue-based parsing account of recovery from misanalyzed ambiguities. journal of memory and language, 49:285–316, 2003. s. vasishth and m. broe. the foundations of statistics: a simulation-based approach. springer, berlin, 2011. s. vasishth and r. l. lewis. argument-head distance and processing complexity: explaining both locality and antilocality effects. language, 82(4):767–794, 2006. s. vasishth, s. bruessow, r. l. lewis, and h. drenhaus. processing polarity: how the ungrammatical intrudes on the grammatical. cognitive science, 32(4), 2008. s. vasishth, k. suckow, r. l. lewis, and s. kern. short-term forgetting in sentence comprehension: crosslinguistic evidence from head-final structures. language and cognitive processes, 25(4): 533–567, 2010. w. n. venables and b. d. ripley. modern applied statistics with s-plus. springer, new york, 1999. s.h. vos, t.c. gunter, h.h.j. kolk, and g. mulder. working memory constraints on syntactic processing: an electrophysiological investigation. psychophysiology, 38(01):41–63, 2001. t. c. warren. understanding the role of referential processing in sentence complexity. phd thesis, massachusetts institute of technology, cambridge, ma, 2001. t. c. warren and e. gibson. effects of np-type on reading english clefts. language and cognitive processes, pages 89–104, 2005. 82 dialogue & discourse 7(4) (2016) 1-35 doi: 10.5087/dad.2016.401 reports in discourse julie hunter juliehunter@gmail.com irit, université paul sabatier toulouse, france & glif, universitat pompeu fabra barcelona, spain editor: jonathan ginzburg submitted 08/2014; accepted 06/2016; published online 06/2016 abstract attitude or speech reports in english with a non-parenthetical syntax sometimes give rise to interpretations in which the embedded clause, e.g., john was out of town in the report jill said that john was out of town, seems to convey the main point of the utterance while the attribution predicate, e.g., jill said that, merely plays an evidential or source-providing role (urmson, 1952). simons (2007) posits that parenthetical readings arise from the interaction between the report and the preceding discourse context, rather than from the syntax or semantics of the reports involved. to my knowledge, however, no account of these discourse interactions has been developed in formal semantics. research on parenthetical reports within frameworks of rhetorical structure has yielded hypotheses about the discourse interactions of parenthetical reports, but these hypotheses are not semantically sound. the goal of this paper is to unify and extend work in semantics and discourse structure to develop a formal, discourse-based account of parenthetical reports that does not suffer the pitfalls faced by current proposals in rhetorical frameworks. 1 keywords: speech reports, parenthetical reports, evidential reports, discourse structure, discourse connectives 1. introduction in the absence of further context, a speaker2 who utters (1) would naturally be understood as using (1b) to propose an explanation for why john did not come to her party. (1) a. john didn’t come to my party. b. jill said he was out of town. the content of jill said in (1b) does not participate, at least not directly, in this explanation—while it is possible to imagine a scenario in which john didn’t come to the party because jill said he was out of town, this reading of (1b) would require a more elaborate discourse than (1). it is rather the 1. i would like to thank márta abrusán, pascal amsili, nicholas asher, laurence danlos, eric kow, mandy simons, participants of the crest international workshop on formal and computational semantics at kyoto, and participants at the nijmegen workshop backgrounded reports, including organizers corien bary and emar maier, for helpful discussions on the issues discussed in this paper. i also thank my three anonymous dialogue & discourse reviewers for their thorough and constructive comments. this work has been supported by the french agency agence nationale de la recherche (anr-12-cord-0004) and the european research council (grant 269427). 2. throughout the paper, i will use speaker to refer to the agent of either a spoken or written utterance/sentence token. ©2016 julie hunter this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). hunter embedded clause, he was out of town that is understood as a (possible) explanation of john’s absence. informally, we can say that the embedded clause seems to contribute the main point of (1b) while the attribution predicate, jill said, merely serves to provide a source or evidence for the embedded content. following hooper (1975), simons (2007) and urmson (1952), i will call the use of say in (1b) a parenthetical use and i will call a report in which the embedding verb is used parenthetically a parenthetical report. now compare (1) with (2), which involves a standard, non-parenthetical use of the same report. (2) a. john is mad at jill. b. jill said he was out of town, c. so i didn’t invite him to my party. the report in (2b) is identical to that in (1b), yet the report does not make the same contribution to (2) as it does to (1). in (2b), the attribution predicate is integral to the explanation of (2a): the speaker is not suggesting that john is mad at jill because he was out of town, but that he is mad at jill because she said something (which led to his being excluded from the speaker’s party). thus, while the embedded clause alone seems to make the main discourse contribution of (1b), the attribution predicate of (2b), or perhaps the report as a whole, crucially participates in the report’s main contribution to (2). the discursive difference between (1b) and (2b) brings with it a difference in semantic entailments. to the extent that a speaker of (1) is committed to the possibility that john didn’t come to the party because he was out of town, she must also be committed to the possibility that john was out of town (at the relevant time). that is, she must be committed to at least the possibility that the embedded content of (1b) is true. there is no such requirement of commitment to the embedded content of (2b): (2b) can be used to explain (2a) even in a context in which it is common knowledge that john was not out of town. if we assume that parenthetical reports like (1b) have the same syntactic structure as their nonparenthetical counterparts—which i, following simons (2007), will—then we must conclude that the discursive and semantic differences between (1b) and (2b) do not arise from the syntactic and semantic features of the report alone, but rather from the interaction between these features and other discourse moves. accordingly, i will introduce the more specific term discourse parenthetical reports to talk about parenthetical reports akin to (1b). this will help to distinguish these reports from those whose parenthetical readings are marked syntactically (see §5 for further discussion of syntactic parentheticals).3 the goal of this paper is to develop a formal, discourse-based model of discourse parenthetical reports. while parenthetical reports have been discussed at length in hooper (1975), rooryck (2001), simons (2007), and urmson (1952) and it has been noted that parenthetical readings arise from the discourse function of parenthetical reports, there is as of yet no formal model of their discourse function or of how this function affects the semantic entailments of reports in different discourse contexts. work in formal semantics has tended to bring observations about parenthetical readings back to bear on outstanding problems in formal semantics. urmson, for example, was concerned with the implications of parenthetical readings for the semantics of attitude reports. simons uses parenthetical reports to argue that the presuppositional behavior of factive verbs is not determined by their lexical semantics: given the right discourse context, even the content in the scope of a factive verb can carry the main point of the report—i.e., the report can have a parenthetical reading—and so the 3. i would like to thank an anonymous reviewer for suggesting the term discourse parenthetical. 2 reports in discourse embedded content can fail to be presupposed. this paper takes for granted the syntax and semantics of reports, and in particular, the assumption that reports with the surface structure of (1b)/(2b) have the same semantics and syntax regardless of whether they are interpreted parenthetically or not. the aim is rather to formalize the notion of discourse function at work in parenthetical readings and to explain how it accounts for the different entailments that arise from parenthetical and nonparenthetical readings. simons provides the most developed discussion of the discourse function of parenthetical reports that i am aware of in the formal semantics and pragmatics literature, but as her focus is on the behavior of different embedding verbs, she limits her discussion to how reports behave in question/answer sequences: (3) a: why didn’t john come to my party? b: jill said (thinks, suspects, imagines, supposes, heard...) that he’s out of town. in these sequences, simons proposes that the discourse function of the embedded clause of a discourse parenthetical report, e.g. (3b), is to answer the question posed by the preceding move, e.g. (3a). she says: “whatever proposition communicated by the response constitutes an answer (complete or partial) to the question is the main point of the response” (p. 1036). as simons explicitly acknowledges (p. 1035), however, this criterion provides only a limited picture of the discourse function of parenthetical reports. theories of rhetorical structure such as rhetorical structure theory (rst; mann and thompson (1988)) and segmented discourse representation theory (sdrt; asher and lascarides (2003)) have more developed notions of discourse structure and discourse function. the discourse function of an utterance u is given by the discourse relation that connects the content c of u to the surrounding discourse: if c is related to some other content c′ via an explanation relation, then c’s discourse function is to explain the eventuality described by c′; if c is related to c′ via a narration relation, then its discourse function is to push the narrative forward by describing the next (discourse relevant) eventuality that occurs after that described by c′, and so on. the structure of the discourse is then determined by the collection of relation instances in the discourse. this notion of discourse function has been applied to discourse parenthetical reports in multiple efforts to annotate newspaper texts for rhetorical structure (dinesh et al., 2005; hardt, 2013; hunter et al., 2006). the parenthetical/non-parenthetical distinction generally becomes relevant when annotators, who are in many cases untrained in linguistics, decide that only the embedded content of a report is relevant to the discourse relation that connects the report to the preceding discourse. (4), an attested example with the form of (1), provides an illustration. (4) london serves increasingly as a conduit for program trading of u.s. stocks. market professionals said london has several attractions. first, the trading is done over the counter ... second, it can be used to unwind positions before u.s. trading begins... (pdtb, file 0097). (4) is originally from the wall street journal, but figures in the corpus for the penn discourse tree bank (pdtb; prasad et al. (2007b)). in this example, pdtb annotators chose the content of the boldface text as the proposed explanation for the claim expressed by the sentence in italics; the attribution predicate of the report, market professionals said, was excluded from the explanans. in other words, annotators took the report in (4) to be discourse parenthetical. 3 hunter while discourse parenthetical reports might not be discussed as such in these annotation efforts, the annotation schemes that they have developed in response to discourse parenthetical readings in effect yield the following definition: a report r is discourse parenthetical just in case it is the embedded clause of r, not the attribution predicate, that enters into a discourse relation with some element from the discourse preceding r.4 simons’ proposed constraint for question/answer pairs can be seen as a particular case of this more general notion: a report r is discourse parenthetical just in case the embedded clause of r alone provides the second argument for an instance of the relation question/answer pair, where q provides the first argument, for some question q in the incoming discourse.5 of course, simons’ constraint could be generalized in different ways. an alternative approach might involve defining discourse function and structure within a question under discussion account (ginzburg, 2012; roberts, 2012; simons et al., 2010), though i will not develop such an approach here. the model that i develop builds on the more general notion of discourse function offered by rhetorical theories, but rectifies certain semantic problems inherent in the annotation-driven solutions that have been offered. in particular, as i explain in §3, extant proposals do not account for the fact, well-known from semantic discussions of parenthetical reports, that a speaker who utters a discourse parenthetical report need not in general be fully committed to the embedded content of that report. nor do they account for the fact that a speaker must nevertheless be at least tentatively committed to the embedded content. finally, these accounts prevent the attribution predicate and the embedded clause of a report from being simultaneously relevant to the discourse outside of the report, at least when they play distinct rhetorical functions. these accounts in effect require annotators to treat either one or the other clause as the “main point” of the report, but that requirement is too strong given the data on reports in discourse. in a nutshell, my proposal is that a discourse parenthetical report is one in which the embedded clause is rhetorically connected to the discourse preceding the report, as in extant accounts, but unlike extant accounts, the embedded clause is discursively and semantically subordinate to the attribution predicate. as a result of these rhetorical connections, discourse parenthetical reports introduce modal discourse relations between the embedded clause and the preceding discourse, and the attribution predicate can be rhetorically related to the preceding discourse, although it must be related by a different discourse relation. it should be noted that the account that i offer is designed to model the contribution of discourse parenthetical reports to discourse structure, but is not intended as a complete semantic account of parenthetical reports. it is meant to complement, rather than replace, studies on other factors that affect the interpretation of speech or attitude reports such as the lexical semantics of embedding verbs, e.g. simons (2007), or the influence of certain kinds of world knowledge on our judgements about the reliability of a discourse parenthetical report, e.g. de marneffe et al. (2012). it should also be a useful supplement to annotation-based work on rhetorical structure. determining the impact of speech and attitude reports on discourse structure and interpretation is of utmost importance. for tasks such as automated discourse parsing, text summarization, and the automatic recognition of textual entailment, for example, we need to be able to draw reliable inferences from a discourse as 4. constraints on how a bit of content can attach to the incoming discourse context are given by independent principles that vary from theory to theory. all that is important here is that the embedded clause is being attached to a representation of the discourse prior to the report, rather than to its own attribution predicate. 5. different frameworks might have different names for this relation; i’ve opted here for the terminology of sdrt, but that is not important. however, given the semantics of question/answer pair (qap) in sdrt (asher and lascarides, 2003), it is important to note that this criterion will be equivalent to simons’ because for a discourse unit to serve as the second argument to an instance of qap, it must provide a complete or partial answer to the first argument. 4 reports in discourse a whole. an important piece in this puzzle is figuring out how speakers make use of other peoples’ speech acts and attitudes to perform their own speech acts and convey their own attitudes. i begin in §2 by providing more detail on rhetorical theories and reviewing three proposals for the annotation of discourse parenthetical reports within different rhetorical frameworks. while these accounts differ in various ways, they all share the core idea that discourse parenthetical reports should be modelled by attaching the embedded clause to the preceding discourse. in §3, i show that these accounts as they stand come into conflict with semantic facts about discourse parenthetical reports and with independent principles of rhetorical structure. in §4 i show how the notion of discourse function adopted in rhetorical theories can be developed into a more more general, consistent account of discourse parenthetical reports. §5 takes a look at how syntactic parentheticals fit into the picture developed in §4. §6 concludes the discussion. 2. rhetorical theories extant rhetorical theories, including rst, sdrt, d-ltag (webber, 2004), and other annotation methods, including those for the pdtb and the copenhagen discourse tree bank (buch-kromann and korzen, 2010), differ in critical ways, and in the more theoretical discussion below, we will need to focus on a single theory. nevertheless, the import of the current study should be of interest for research on rhetorical structure more generally. while the three accounts that i describe below are worked out using very different annotation methodologies, the heart of the proposals is the same, and as such, parts of the discussion to follow will be relevant to all of them. where we will need a particular theory is in working out the details of my concerns and developing a positive proposal. building a representation of the rhetorical structure of a given discourse requires performing three tasks: segmenting the discourse into minimal units, attaching each discourse unit to some other unit in the structure, and labelling each discourse attachment with a rhetorical relation. when it comes to the annotation of discourse parenthetical reports, the approaches outlined in the pdtb (dinesh et al., 2005), the copenhagen dependency treebank (cdt) (buch-kromann et al., 2011), and sdrt (hunter et al., 2006; reese et al., 2007) all follow a similar recipe for accomplishing these tasks. first, a report r that is intuitively discourse parenthetical is segmented such that the embedded clause contributes its own discourse unit u. second, u is attached to another unit u′ in the discourse structure s built from the set of utterances prior to that of r. third, the relation between u′ and u is labelled with the discourse relation that intuitively would have related u′ and u had u not been the argument of an attribution predicate. finally, if desired, the discourse function of the attribution predicate can be modelled by attaching it to u with a special discourse relation that indicates its evidential (or emotive, etc.) discourse function. the pdtb, for example, would indicate the discourse contribution of (1) as follows: (1) a. john didn’t come to my party. b. implicit = because jill said he was out of town. following the conventions of the pdtb manual (prasad et al., 2007b), content that contributes to the first argument of a discourse connective is placed in italics, while content that contributes to the second argument is placed in boldface. the connective, when explicit, is underlined; implicit connectives are marked as in (1b). we can see from the representation of (1) (which echoes the actual pdtb annotation of the found example (4)) that the embedded clause forms its own segment and it alone serves as the argument that ties the report to the discourse context prior to (1b) (prasad et al., 5 hunter 2007a).6 this segmentation choice contrasts with the segmentation for non-parenthetical reports, in which the entire report, attribution predicate plus embedded clause, is treated as a single discourse segment. the pdtb does not take a stand on the discourse contribution of the attribution predicate in discourse parenthetical reports; information about the sources and spans of attributions is stored alongside the annotation of a discourse in the pdtb but there is no connective posited to capture the discourse function of this information. in the cdt, the annotation of syntactic structure and discourse structure is done using a single dependency graph so that there is a high level of uniformity between discourse and syntactic structure (buch-kromann and korzen, 2010). one of the rare violations of this uniformity is allowed for discourse parenthetical reports, in which only the embedded clause is treated as a discourse argument. instances of relations involving discourse parenthetical reports (and other syntax/discourse mismatches) are marked with a ‘*’: an asterisk to the left of a discourse connective signals that the discourse/syntax mismatch can be found in the first argument of the relation; an asterisk to the right signals that the mismatch lies in the second argument (hardt, 2013). (1), for example, would be annotated along the following lines: explanation∗(1a, 1b). like the pdtb, the cdt does not take a stand on the discourse contribution of the attribution predicate in a discourse parenthetical report. a third proposal was offered by hunter et al. (2006) for sdrt as part of the project discor7. hunter et al., in contrast to the pdtb group, always treat the attribution predicate and embedded clause of a report as separate segments, regardless of whether the report has a discourse parenthetical reading or a non-parenthetical one. the different readings of reports are distinguished by the fact that the embedded clause of a discourse parenthetical report is the clause that is attached to the incoming discourse structure, as for the pdtb and cdt. in addition, discourse parenthetical and non-parenthetical reports are distinguished by different discourse relations that connect the attribution predicate of the report to the embedded clause: attribution is used for non-parenthetical reports and source, for parenthetical reports.8 in attribution, the embedded clause is discourse subordinate (see asher and lascarides (2003)) to the attribution predicate, echoing the syntactic structure of the report. source, however, reverses the arguments of attribution so that the attribution predicate is discourse subordinate to the syntactically embedded clause. furthermore, source and attribution differ in their entailments: when two arguments are related by source, the text entails the content of both arguments, whereas when two contents are related by attribution, the text only entails that the agent of the attribution stand in the relation indicated by the embedding verb to the content of 6. the pdtb group argue that the relation of attribution is one that holds between an individual and an abstract object. as discourse relations are taken to relate abstract objects only, attribution is not treated as a discourse relation. so the difference between discourse parenthetical reports and non-parenthetical reports is that in the latter, the whole report is treated as a single unit that can figure in an argument for a discourse relation whereas in the former, only the embedded content contributes to an argument. in their words: “a discourse relation may hold either between the attributions (and the agents of attributions) themselves or only between the abstract object arguments of the attribution...” (p. 20) 7. discor, discourse structure and co-reference resolution, was an nsf projet designed to study the relation between co-reference resolution and discourse structure. most of the texts annotated for discourse structure were texts from the message understanding conference (muc) 6, that were already annotated for co-reference. 8. the relation source was originally introduced by hunter et al. (2006) under the name evidence. the name was changed so that evidence could be used for a different evidential relation. i have chosen to use the name source here because it is consistent with later sdrt annotations, e.g. reese et al. (2007), that incorporated hunter et al.’s approach. 6 reports in discourse the embedded clause.9 an example of source is provided in (1h), following the segmentation of (1) below: (1) a. [john didn’t come to my party.]α b. [jill said]β [he was out of town.]γ (1h) explanation(α, γ), source(γ, β) two further accounts of speech reports that deserve mention here, though i will not discuss them in the rest of the paper, are those of carlson and marcu (2001) and redeker and egg (2006), both developed in rhetorical structure theory (mann and thompson, 1988). carlson and marcu suggest that all reports be annotated with a relation that they call attribution, which is structurally similar to hunter et al.’s source in that the attribution predicate in a report is treated as the satellite and the reported content serves as the nucleus. redeker and egg (2006) criticizes carlson and marcu’s approach and offers an account that reverses the arguments of attribution so that the reported content is the satellite and the attribution predicate, the nucleus. the reason why i will not pursue either of these accounts in this paper is that neither makes a distinction between discourse parenthetical reports and non-parenthetical reports.10 the attribution predicate is simply treated as the satellite of attribution in carlson and marcu (2001) and as the nucleus in redeker and egg (2006). what we’re interested in for the purposes of this paper is accounts that advocate a solution specifically for discourse parenthetical readings of reports. the treatment of discourse parenthetical reports in the pdtb, the cdt, and hunter et al. (2006) all have in common the idea that discourse parenthetical reports are best modelled by attaching the embedded clause directly to the incoming discourse and that this attachment pattern distinguishes them from non-parenthetical reports, in which it is the attribution predicate that is attached to the incoming discourse. accordingly, i classify these accounts as attachment solutions. however, the claim that discourse parenthetical reports show different attachment patterns in discourse does not entail that discourse parenthetical reports and non-parenthetical reports should be distinguished syntactically, and i will take it for granted that the difference is not a syntactic one—except, of course, when the parenthetical verb appears in a syntactic parenthetical as in mary will be late, john said (see §5 for a discussion of syntactic parentheticals). i will not defend this position here, because my focus is on the discourse contribution of parenthetical reports, regardless of their syntactic structure.11 the point is that claims about the discourse structure of a chunk of discourse do not automatically entail claims about the syntactic structure of the constituents in that chunk. in fact, of the frameworks introduced here, the one that assumes the closest tie between syntactic structure and discourse structure is that for the cdt, and even the cdt treats discourse parenthetical reports as involving a syntax/discourse mismatch. in other words, they assume that the discourse contribution of a parenthetical report does not mirror its syntactic form. 9. for attribution(α, β) or source(β, α), the content of α will entail: x said (thought,...) p for some agent x, and β will specify the content of p in the sense that p will denote a subset of the worlds in the proposition denoted by kβ , i.e. the drs that results from the processing of β. 10. however, see matthiessen and thompson (1988) for an interesting discussion about the fact that whether a given clause is ultimately considered to be a main clause or a subordinate clause depends on features of the discourse in which that clause is used. 11. see simons (2007) for arguments that reports in which an embedding verb appears in the main clause, as in john said (that) mary will be late, have the same syntactic structure regardless of whether the embedding verb is used parenthetically or not. 7 hunter 3. conflicting criteria in the ensuing discussion, i will adopt sdrt as my theoretical framework. to fully model the rhetorical contribution of reports, we need to take a stand on the discourse function of both the embedded clause and the attribution predicate, and we need a theory that offers independent principles to guide our theoretical choices on this matter. sdrt satisfies this criterion, but not all rhetorical frameworks do. the pdtb annotation method, for instance, is designed to be theory-neutral, and so cannot provide a theoretical framework by design. moreover, part of its theory-neutral approach is to annotate isolated pairs of discourse arguments and the connectives that relate them; there is no goal to describe the discourse contribution of every discourse unit. as a result, the pdtb can remain agnostic about the role of the attribution predicate in discourse parenthetical reports. sdrt is also largely motivated by semantic and pragmatic concerns and is the only rhetorical theory to provide a semantics for its discourse relations and to deliver fully interpretable logical forms for discourse. many of the issues that we confront with discourse parenthetical reports are semantic and pragmatic, having to do with tracking the entailments of reports in a discourse or exploring aspects of their behavior that cannot be traced back to their syntactic structure. thus a theory like sdrt is preferable to the framework of the cdt, which is wedded to a strong correspondence between syntactic structure and discourse structure, and even to rst, which does not emphasize the semantic interpretation of its discourse structures. i will also limit the range of report verbs in my study. there are many factors that influence the interpretation of reports aside from rhetorical structure: lexical semantics, world knowledge, perhaps focus, and so on. to study the interaction of rhetorical structure and reports, we need to minimize the influence of these other factors as much as possible. most of my discussion will therefore be centered around embedding verbs like say and other speech report verbs that are non-factive and so do not indicate a particular level of author commitment to the content in their syntactic scope. such speech report verbs are common in the corpora that i am pulling from12 and easily give rise to both discourse parenthetical and non-parenthetical readings. also common are third person reports, and my discussion will focus largely on these as well, as it is with third person reports that questions about author commitment and entailments of reported content really become tricky. with these caveats in place, i turn now to an argument that the attachment-based solution outlined in section 2 leads to conflicts with semantic facts about discourse parenthetical reports and with independent principles of rhetorical theories. 3.1 hedged commitments a speaker will often use a discourse parenthetical report to weaken or “hedge” her commitment to the embedded content of the report (simons, 2007). in (1), for example, if the speaker were sure that john had been out of town, then the simplest solution would be to say so directly; the fact that she does not suggests that she is not entirely certain that he was out of town. of course, we can imagine scenarios in which the use of a discourse parenthetical report is compatible with full speaker commitment to the embedded content, but what’s important is that a discourse parenthetical report does not itself require such commitment. this fact about discourse parenthetical reports comes into conflict with the constraint of veridicality imposed by many discourse relations. a relation r is veridical just in case the truth of an 12. the majority of the verbs in the corpus annotated for discor were evidential verbs, e.g. say, as opposed to, for example, verbs indicating emotions, e.g. regret. 8 reports in discourse instance of r entails the truth of that instance’s arguments. boolean conjunction, for example, is veridical: a formula of the form p ∧ q can only be true in a model m if both p and q are true in m . conditional relations, by contrast, are not veridical: a formula of the form p→ q can be true even if both of its arguments are false. sdrt incorporates these basic logical connections into the semantics of its discourse relations—an arguably reasonable move for any semantic theory of rhetorical structure. the relation continuation, for instance, conjoins two discourse units, and so each instance of continuation will entail both of its arguments. contrast also has conjunction as the foundation of its semantics, as does narration; these relations are therefore veridical as well. explanation is yet another veridical relation: a discourse unit u cannot truly explain another discourse unit u′ unless u′ and u both describe eventualities that held in the world of evaluation. to capture this dependence, a discourse formula of the form explanation(u′, u) can only be true in a modelm in sdrt if u and u′ are true in m . when a speaker performs a speech act that presents a discourse unit u as standing in a relation r to another unit u′, the content of this speech act, r(u′, u), is added to the logical form for the discourse (for anyr, u, u′). ifr is veridical, then the speaker takes on a commitment to the content of bothu andu′, i.e. the content of her discourse entails bothu andu′.13 note that this does not ensure that either u or u′ will be true in the relevant model—speakers are not infallible—it only ensures the speaker’s commitment to their truth. here is where the conflict between veridicality and hedged commitments arises: attachment solutions attempt to model the discourse function of discourse parenthetical reports by attaching the embedded content of a report to the incoming discourse with the same relation that would have been used had the embedded content not been embedded. where the relation is veridical, this entails author commitment to the embedded content—commitment that the speaker may not be ready to take on. what’s more, the relation at issue always will be veridical. if a discourse parenthetical report seems to provide an argument to a non-veridical relation, such as alternation or conditional, the result is that the attribution predicate will be understood as scoping over the discourse relation, as in (5): (5) if [john finishes his housework]α, then [linda said]β [he’ll come to the party.]γ if linda has reasons for thinking that john will come to the party that have nothing to do with him finishing his housework, then the choice of antecedent in (5) is unmotivated. the natural interpretation is therefore one in which linda is committed to the conditional as a whole. hunter et al. (2006) avoid the conflict between veridicality and hedged commitments by requiring that source be veridical. this means that source can only be used in cases in which annotators (interpreters) judge that the speaker is committed to the embedded content. recall hunter et al.’s analysis of (1) repeated here: (1) a. [john didn’t come to my party.]α b. [jill said]β [he was out of town.]γ (1h) explanation(α, γ), source(γ, β) in fact, while i provided this example to illustrate the structural features of source, this annotation would have only been allowed by hunter et al. if the context provided reason to accept jill’s report as totally reliable. 13. this picture of commitment follows original sdrt, but is complicated by issues such as embedded commitments and disagreements about commitments, topics which are handled in more recent work on sdrt. see venant et al. (2014) and venant and asher (2016). 9 hunter in the corpora used by the pdtb, the cdt, and discor, which consist mainly of newspaper articles, treating source as a veridical relation often yields the right results. this is because in many of the articles, the content is so uncontentious and the sources so reliable that there is little reason to question the embedded content of the reports or the author’s commitment to it. (6b) is discussed by dinesh et al. (2005) as an example of contrast with a discourse parenthetical report (their example (12)): (6) a. at the same time, the new brunswick, n.j., company said negotiations about pricing and volumes of product had collapsed between it and its exclusive distributor in the u.s., national medical care inc. ... yesterday, the [delmed] spokeswoman said sales of delmed products through the exclusive arrangement with national medical accounted for 87% of delmed’s 1988 sales of $21.1 million. b. the current distribution arrangement ends in march 1990, although delmed said it will continue to provide some supplies of the peritoneal dialysis products to national medical, the spokeswoman said. (6a) provides two of the three sentences that precede (6b) in the pdtb (file 0970). (6b) is presented as being factual: delmed and the spokeswoman should be reliable sources given their relation to delmed, there is nothing contentious about the content that they have reported (the content that dinesh et al. mark as contributing to argument 2 of although), and the author gives no signal in the rest of the article that s/he is not fully committed to this content. in examples like (6b), the take-home message seems unchanged if we remove information about the writer’s sources. (7) illustrates the same point, but with an elaboration or instance relation. while this example is not from the pdtb, i will use their conventions for marking the arguments for ease of exposition. (7) a. another firm, ureveal, thinks that it’s cracked the code. charles “bucky” clarkson, ureveal’s chairman and ceo, said that software such as his makes it easier to to parse all those government reports and organize the data so that analysts can get more out of it, and more quickly. he also claims that the software is so simple to use that (gasp!) even liberal arts majors can use it. jokes about “soft majors” aside, the idea of data analysis tools easy for anyone to use is compelling because it frees up data scientists to do more specialized work. b. it can also bring in specialist knowledge from people who aren’t data scientists. clarkson said, for example, that deploying these kinds of data analysis tools in hospitals have allowed doctors to spot trends that they would have otherwise missed, by analyzing their observation notes in conjunction with other electronic medical records. one doesn’t have to look far, however, even in the realm of newspaper articles, to see that discourse parenthetical reports are used widely in contexts in which speaker commitment to the embedded content is not ensured. a paradigm example is when a journalist uses discourse parenthetical reports to report the opinions of two parties who disagree with each other. for instance, in example (8), a toy variant of (13), which is discussed below, the speaker uses one report to express the point of view of the employees and another report to express the point of view of the boss, who directly disagrees with the employees. 10 reports in discourse (8) a. there was an explosion at the factory. b. the employees said that it was the fault of the boss c. but the boss said that it was the fault of the employees. we cannot infer speaker commitment to the embedded clause of either report by looking at (8) alone, but both reports are nevertheless used discourse parenthetically to offer possible explanations of the explosion. this is a problem for hunter et al.’s account because the discourse parenthetical use of the reports calls for annotating the reports with source, but the fact that the embedded content of the reports is not entailed precludes an annotation with source. a further problem with hunter et al.’s account is that even in cases such as (6b) and (7b), in which author commitment to the embedded content of a discourse parenthetical report can be inferred, commitment is almost always inferred by using world knowledge to reason about the reliability of the source(s) cited in the attribution predicate of the report. this kind of world-knowledge based reasoning, central to de marneffe et al. (2012)’s study of veridicality, is generally independent of the reasoning used to determine rhetorical structure. in other words, hunter et al.’s source/attribution distinction does not reflect a rhetorical distinction, and thus it fails to use the notion of discourse function from sdrt (or any rhetorical theory) to model the intuitive discourse function of discourse parenthetical reports. 3.2 no commitment veridicality entails that the embedded clause of (1b) cannot be related to (1a) via explanation, at least if there is any doubt about the speaker’s commitment to the truth of this clause’s content. this is intuitively correct: if the speaker is not entirely confident about the claim that john was out of town, she cannot be confident about the implicature14 that john’s being out of town explains why he wasn’t at the party. still, in uttering (1b) she makes salient the possibility that john didn’t come to the party because he was out of town and performs something like an explanation—a hedged explanation, if you will. this intuition plays an important role in motivating attachment-based treatments of discourse parenthetical reports. however, some reports that seem in other ways to be discourse parenthetical carry no requirement of speaker commitment to the embedded clause and are not intuitively used to offer the content of the embedded clause as even a potential explanation or answer, etc. this happens when a speaker explicitly denies the content of the embedded clause or otherwise reveals that she thinks it is false. simons (2007) discusses some such examples using question/answer pairs. (9) a. which course did louise fail? b. henry, falsely, thinks that she failed calculus. (simons, example (18)) these examples are complicated. at first glance, the report (9b) seems discourse parenthetical: the content of the embedded clause appears to be what makes the report relevant to the incoming discourse because it is this content that potentially provides an answer to (9a). on the other hand, in (9b), the speaker cannot be taken as actually offering the proposition louise failed calculus as even a potential answer to the question posed in (9a); in fact, she makes it clear that she is committed to 14. a discourse relation that is not explicitly marked but inferred on the contents of its arguments is cancellable, although many instances will be very difficult to cancel. 11 hunter that proposition’s not being an answer. as a result, relating the clause embedded under thinks to (9a) via the relation question-answer pair would seem inappropriate.15 one could also use a negated speech or attitude verb (didn’t say, doesn’t think) or a negative verb (doubts, is skeptical) to block the associated speech act. even more interestingly, from the perspective of rhetorical structure, is that we can undermine an intuitively discourse parenthetical reading by stringing multiple discourse units together to form multi-part responses. for instance, a speaker can disengage herself from the embedded content of a report with a comment, e.g. (10), or a contrast, e.g. (11) and (12): (10) henry said she failed calculus. he always gets things wrong! (11) henry said she failed calculus, but he’s wrong. (12) henry said she failed calculus, but anna disagrees. in fact, it can take many discourse turns to determine a speaker’s commitment to the embedded content of a report. the excerpt below is from an article in the new york times.16the first two paragraphs of the article discuss an earthquake that hit prague, oklahoma in 2011 and the destruction that it caused. the excerpt provides the third and fourth paragraphs: at a packed town hall meeting days later, ms. cooper said, state officials called the shocks, including a 5.7 tremor that was oklahoma’s largest ever, “an act of nature, and it was nobody’s fault.” many scientists disagree. they say those quakes, and thousands of others before and since, are mainly the work of humans, caused by wells used to bury vast amounts of wastewater from oil and gas exploration deep in the earth near fault zones. and they warn that continuing to entomb such huge quantities risks more dangerous tremors — if not here, then elsewhere in the state’s sprawling well fields. these paragraphs work together to present possible explanations for the earthquake introduced in paragraphs 1 and 2. let’s simplify the example for the sake of our discussion. (13) a. in november 2011, a 5.0 magnitude earthquake shook prague, oklahoma. b. state officials said that the earthquake was an act of nature that was nobody’s fault. c. however, many scientists argued that the quake was caused by wells used to bury vast amounts of wastewater from oil and gas exploration deep in the earth near fault zones. 15. in (9), we can infer a negative answer to (9a), namely that louise did not fail calculus (groenendijk and stokhof, 1984). however, i would like to distinguish (9b) from a report such as (9b’): “henry thinks/said that she didn’t fail calculus". in (9b’), the speaker is really offering the embedded content as a negative answer to (9a), albeit a tentative negative answer. in (9b), the speaker has an answer that is completely independent of what henry thinks or says—that’s why she can judge that henry is wrong. the point of bringing henry into it, then, is not so much to use him as a source or as evidence for a potential answer, but to indicate that the speaker is aware that henry might answer the question differently and thinks that henry is confused. thus, henry’s commitments are rhetorically central in (9) in a way that they are not in (9b’). 16. ‘as quakes rattle oklahoma, fingers point to oil and gas industry’, by r. a. oppel jr. and m. wines. the new york times, april 3, 2015. 12 reports in discourse (13b) and (13c) oppose two viewpoints, much like (12). at this point in the article, we can only take the author to be presenting two possible explanations.17 the rest of the (lengthy) article makes it clear that the authors side with the scientists, however. the following paragraph, taken from the same article, reveals this endorsement: but in a state where oil and gas are economic pillars, elected leaders have been slow to address the problem. and while regulators have taken some protective measures, they lack the money, work force and legal authority to fully address the threats. at this point, the discussion shifts from trying to find an explanation for the earthquake to a discussion of why the local government has been so slow to address the problem brought up by the scientists. the assertions are no longer in the scope of reports, so we are dealing here with the authors’ commitments. the use of definites such as the problem and the threats presuppose the existence of the problem/threat that that the scientists introduce in (13c). the larger discussion of why the local government has been so slow in addressing the problem also presupposes that the question of what the problem is has been answered. as this paragraph reflects the authors’ commitments, we can infer that they have sided with the scientists, and thus have no commitment to the embedded content of (13b). we can make a similar point with (1). imagine an utterance of (1) followed by either (14) or (15) (but not both): (14) she must be covering for him, though, because i spotted them together at the market this morning. (15) but he was only an hour away. he could have come if he’d wanted to. there must be some other reason. in judging (1) alone, an interpreter would be entitled to infer that the speaker is committed to the possibility of john’s having been out of town and is using this possibility to provide a possible explanation of why john didn’t come to her party. however, once the speaker continues with either (14), which denies that the potential explanandum holds, or (15), which denies that the potential explanandum actually explains john’s absence, an interpreter is no longer justified in inferring even a hedged explanation relation between (1a) and (1b). the attitudes that the speaker takes towards the truth of the embedded content of a report and its discourse function are not revealed by considering only the discourse move immediately preceding the report. starting definitions of discourse parenthetical reports, like that offered by simons (2007) or the attachment-based proposals offered by various rhetorical frameworks, are driven by the intuition that the embedded contents of these reports play a certain discourse function. intuitions are primed by considering pairs consisting of a report and a single, preceding discourse unit. the problem brought out by examples like (10)-(13) is that if we place these pairs in a larger discourse context, our intuitions about the discourse function of the report, and what inferences are licensed by it, 17. the fact that scientists are often taken to be more reliable sources than state officials with regard to causes of natural disasters might lead a reader of (13) to suspect that the writers side with the scientists and wish to adopt their explanation. however, we do not get that from looking at the discourse structure of (13) alone. this kind of world knowledge, discussed by de marneffe et al. (2012), also plays an important role in the interpretation of reports in discourse, but the focus of this article is on how the rhetorical structure of discourse, and the anaphoric connections between discourse units, affects interpretation. 13 hunter can change significantly. but then what should we say about the reports in such cases? are they discourse parenthetical or not? what features shall we use to decide? attachment-based accounts provide no clear answer. simons’ discussion waffles on these questions as well. of examples like (9) (her (18)), she says that: “main point content cannot be identified with the content of either the subordinate clause alone or the main clause. rather, main point content emerges from the interaction between the subordinate clause content and the attitudes to that content expressed by the other predicates used.” this undermines her definition of discourse parenthetical reports. in the beginning, she defines a discourse parenthetical report as one whose embedded clause conveys the main point of the utterance, but later, she seems to suggest that reports like (9b) should count as discourse parenthetical. although she notes this tension, she does not go on to work it out, as her focus is on other issues, so the ensuing discussion does not point to a path down which rhetorical accounts of discourse parentheticals can go. 3.3 distinct functions in the previous sub-section, we discussed examples in which a report is offered as a response to a single discourse unit (e.g. which course did louise fail?), but in which the full discourse contribution of the report—the information that determines what conclusions can be drawn from the report—can only be understood by looking at chunks of discourse involving multiple discourse units. in this section, we note that the other direction is possible as well: sometimes a single report is a coherent response to more than one discourse unit. in other words, the embedded clause and the attribution predicate might each have a discourse function relative to the incoming discourse, but these functions are distinct and we might not catch both of them if we only consider a single incoming discourse unit (such as a question). let’s return to (13), but alter it slightly to bring it more in line with the original excerpt. (16) a. in november 2011, a 5.0 magnitude earthquake shook prague, oklahoma. b. state officials said that the earthquake was an act of nature that was nobody’s fault. c. many scientists disagree. d. they argue that the quake was caused by wells used to bury vast amounts of wastewater from oil and gas exploration deep in the earth near fault zones. (16c) describes the scientists’ attitudes, and (16d) intuitively elaborates on this description. yet (16d) can only be an elaboration on (16c) if we treat the attribution predicate as rhetorically relevant; the embedded clause alone does not talk about the scientists’ attitudes, but only about the object of their attitudes. thus locally—that is, looking at the report in (16d) as a response or follow up to (16c) alone—the main point of the report seems to be to provide more information about what the scientists hold and the report is not obviously construed as discourse parenthetical. however, as we saw in the last subsection, if we look at the way (16b-d) (or the associated part of the original excerpt) is functioning in the larger discourse, we see that the embedded clause of (16d) serves a higher discourse function, namely that of conveying a possible explanation of why the earthquake occurred. on this level, it is the embedded content that seems more rhetorically important or that carries the main point. 14 reports in discourse once again, we see that the effect of reports on discourse structure can be subtle and diffuse and the intuitions about main point and discourse relevance triggered by looking at pairs of utterances, e.g. (1), do not carry us through to a full account of reports in discourse. 4. modeling reports in discourse the discussion from §3 highlights three issues that a theory of discourse parenthetical reports needs to address. first, because of hedged commitments, the use of a discourse parenthetical report can be taken to be at most a speech act that conveys the possibility of a rhetorical connection between two eventualities. second, because of examples of no commitment, even an act of hedging seems too strong in many cases—a speaker might use a report that at first seems discourse parenthetical only to later reveal a lack of commitment to the embedded content. third, because of distinct functions, both parts of a discourse parenthetical report can enter into discourse relations with the larger context; in other words, both can serve a discourse function in the sense defined by a theory of rhetorical structure. as these issues remain unresolved, our working definition of a discourse parenthetical report remains vague: a discourse parenthetical report is one in which the embedded clause plays a major role in making that report relevant to the preceding discourse. the aim of this section is to clarify what it is for a report to be discourse parenthetical and to develop a model of the discourse contributions of such reports. the model that i favor is presented in §4.2. but first, §4.1 lays out my reasons for rejecting a very different, but seemingly attractive, alternative model. 4.1 separate layers of information? the problems of hedged commitments and no commitment stem largely from a conflict between the intuitive discourse function of a discourse parenthetical report and the kind of speaker commitments that such a function normally entails. an appealing hypothesis, then, is that discourse structures should track information about discourse relations and information about a speaker’s attitude to those relations in two separate dimensions (cf. potts (2005)’s multi-dimensional account of (not-) at-issue content). the pdtb annotation approach suggests, although it does not entail, one such model. recall that the content of the attribution predicate of a discourse parenthetical report never contributes to the argument of a discourse connective in the pdtb. nevertheless, this information is stored alongside the annotations and could in theory be exploited during the interpretation of the annotations. such a two-dimensional approach might also work to model simon’s claim that the two clauses of a discourse parenthetical report “convey two different types of content” (simons 2007, p. 1053), although because simons does not provide a model or a formal notion of discourse function, she is not committed to such an approach. the problem with a two-dimensional model that treats the attribution predicate and the embedded clause as conveying two different types of content is that it does not square with the observation of distinct functions. if both parts of a discourse parenthetical report can enter into discourse relations, i.e., both parts make the same type of contribution to discourse structure, then whatever differences may exist between the types of content conveyed by the two clauses, it is not a difference that motivates a model in which the attribution predicate contributes the whole of its content to one 15 hunter dimension or layer of discourse structure while the embedded clause contributes the whole of its content to another.18 an alternative approach would be to posit two layers of information, one that tracks rhetorical relations and one that tracks commitment to those relations, but allow the attribution predicate to contribute to both layers. danlos and rambow (2011) proposes an account along these lines. the general idea is to annotate the relations that intuitively hold between discourse units without regard for the speaker’s commitment to the content of those units. then, as each labelled attachment is constructed, formulas are added to the annotation that contain information about the speaker’s attitude towards each argument of that relation. these formulas take the form: f(e, s) =mod, pol, where s is an agent of a propositional attitude, e is the eventuality that serves as the object of the attitude, and f is a propositional attitude function that maps a pair (e, s) to a judgment about e. following saurí and pustejovsky (2009), these judgments have a modal component (mod)—-which can take the value certain (ct), probable (pr), possible (ps), or unknown (u)—and a polarity component (pol)—positive (+), negative (-), or unknown (u). finally, each discourse relation in the annotation structure is subscripted with a source. in general, the source will be the writer or speaker, but in the case of reports, the source can be the source of the reported content. to illustrate, in (17), the narration relation that intuitively relates bill and jane’s attending dinner and their going dancing would be attributed to bill. (17) [jane had dinner at my place on thursday,]α [and bill said that]βatt [she went dancing afterwards.]β this would (skipping a few details that are irrelevant here) yield the following, two-part annotation: (i) attributionwr(βatt, β), narrationbill(α, β) (ii) f(eα,wr) = ct + ∧ f(eβ,wr) = uu ∧ f(eβ, bill) = ct+ ‘attributionwr(βatt, β)’ means that the writer is committed to an attribution relation between βatt and β and ‘narrationbill(α, β)’ means that bill is committed to a narration relation between α and β. ‘f(eα,wr) = ct+’ means that the writer is certain that eα holds (similarly for ‘f(eβ, bill) = ct+’), and ‘f(eβ,wr) = uu’ means that both the truth of eβ and the author’s attitude towards its truth are unknown.19 while this proposed annotation captures many intuitive features of (17), the proposed account fails to make the right predictions. first, (17) supports neither the inference that bill is committed to narration(α, β) nor the inference that he is committed to eβ having occurred; it justifies at most the conclusion that the speaker is committed to bill’s being committed to these things—the speaker could be making all of this up just to get jane in trouble. second, there is a question of why we 18. maier and bary (2015), following work by potts (2005), develops a different kind of two-dimensional model in which the embedded clause of a discourse parenthetical report contributes to one dimension and the entire content of the report (including the embedded clause again) contributes to another. maier and bary do not explore a discourse-based account of parenthetical readings, opting rather for an ambiguous semantics for reports that distinguishes between parenthetical (two-dimensional) and non-parenthetical (one-dimensional) reports. for this reason, i will not explore their account in the current paper; nevertheless, i suspect that the troubles brought out by distinct functions will apply to their account as well, as such examples count against a simple binary distinction in which only the matrix clause or only the embedded clause is active in a given interpretation. 19. it is unclear whether the speaker’s attitude towards the attribution predicate is recorded. it does not appear to be, though it is not obvious in danlos and rambow (2011) why it would be excluded, since it figures in discourse relations. 16 reports in discourse would want to track bill’s commitments in the first place. the units βatt and β are only relevant to α insofar as the speaker is committed to there being a relation between α and some part of the report. discourse structures track speech acts of the speaker because it is the speaker who chooses the discourse. it is also unclear why annotations of discourse structure must track the distinction between a speaker’s thinking an eventuality is possible and her thinking that it is probable. distinctions like this will in general be entailed by the semantics of the embedding verb chosen (know is factive, etc.) or by the way that the report is used in the larger discourse context. this information doesn’t require an extra layer of annotation on discourse structure. plus, even if we wanted to make such modal distinctions, they could be captured by appending modal operators to the relations (e.g., ♦explanation(α, β)). we need modal relations in any case to handle responses like: well, he was out of town, so maybe that’s why he didn’t come, in which the speaker expresses her commitment to the potential explanans but hedges her commitment to the relation itself. tracking commitments with modalized relations also entails the requisite commitments to the arguments themselves: if a speaker is committed to ♦r(α, β) for a veridical relation r, then veridicality entails that her level of commitment to β be at least ♦β, which is consistent with either full commitment to β or commitment only to ♦β. thus, by putting the modal on the relation, we get the right results for both the relation and the arguments. by contrast, merely putting the modal on the argument itself (which is equivalent to what danlos and rambow propose) does not achieve what is intuitively required—a modal relation. the import of the report in (1) is not correctly captured as explanation(1a, ♦he was out of town), but as ♦explanation(1a, he was out of town). the more general point is that if a speaker uses a discourse parenthetical report to hedge her commitment to a discourse relationr(α, β), then, as explained in §3.1 and §3.2, she is not performing a speech act whose content is r(α, β). annotating (1) with the formula explanation(1a, he was out of town) and then separately adding the information that the speaker is not fully committed to this content simply isn’t the right thing to do in a rhetorical theory, because there is no speech act of explanation that has taken place; the speech act just is a proposal of a possible explanation. one has to take the semantics of speech acts into account in calculating what speech act took place; consequently, the veridical nature of veridical relations and their discourse function cannot actually be separated as a two-dimensional or two-layer theory would propose. the natural thing to do within a rhetorical theory is to see the way that speakers use discourse parenthetical reports as speech acts in their own right by assigning them a different kind of discourse relation and then associating these relations with their own semantics. i develop such an account in the next subsection. 4.2 an integrated model i develop my account of reports in discourse by extending sdrt and modifying attachment-based analyses of discourse parenthetical reports to handle the three issues outlined in §3. let’s begin with the problem of distinct functions, that is, the fact that both parts of a discourse parenthetical report r can be rhetorically relevant by simultaneously being linked to distinct discourse units in the discourse graph for the discourse preceding r. to capture these data, we must allow both parts of a report to contribute to arguments of discourse relations, which in turn requires that we distinguish the attribution predicates of all reports from their embedded clauses by segmenting the two parts as separate discourse units. thus a given report r will contribute two segments to discourse: the 17 hunter attribution predicate, whose content entails ∃p.x said (thinks...) that p for some agent x supplied by the report, and the embedded clause, which specifies the content of p. segmenting reports this way is somewhat controversial, because the attribution predicate does not denote a precise eventuality on its own; it merely tells us that something was said (or thought, etc.). only when combined with the embedded clause does it adequately specify the requisite speech event (or attitude). this gives some merit to the pdtb and cdt choice to not treat reports as contributing two distinct segments despite the fact that reports do convey two eventualities: that of the saying or thinking, etc. and that described by what was said or thought, etc. the benefits of segmentation outweigh the complications, however. we have too much evidence that both clauses of a discourse parenthetical report can be discursively relevant to ignore one or the other. plus, the idea of creating separate discourse units for intimately intertwined contents is already required for treating other phenomena such as presupposition in sdrt. in the case of presupposition, there is no syntactic unit that circumscribes a discourse unit for the presupposition; nevertheless, a separate unit for the presupposition is semantically motivated and required in the discourse structure.20 the next step in modelling the discourse function of discourse parenthetical reports is to determine how the attribution predicate and embedded clause are related to one another. recall that hunter et al. propose a special relation, source, to handle discourse parenthetical readings. in contrast, i propose that all reports be annotated with attribution, regardless of whether they receive a parenthetical reading or not. that is, given two segments, α and β, where α labels the attribution predicate of a report r and β labels the embedded clause, i propose that r contributes the following structure to a discourse graph: α β attribution structurally, attribution is a subordinating relation, which reflects the syntactic structure of discourse parenthetical reports. for its semantics, suppose we have an instance of attribution(α, β) for some discourse units α and β, and let kα be the drs that represents the content of α, and kβ , the drs that represents the content of β (see kamp and reyle (1993)). kα should entail that some agent x stands in some attitude a to a proposition p. for example, the attribution predicate of jill said john was out of town entails the content said(jill,p) for some proposition p (i.e. jill said something). the formula attribution(α, β) is then true just in case the content of kα is true and the content of kβ specifies the content of the proposition p. or in dynamic semantic terms, update with attribution(α, β) at a world w and assignment f extends f to an assignment g that classically satisfies attribution(α, β) at w just in case w and f likewise support update with kα and the resulting assignment g assigns p a subset of the w-accessible worlds w′ that support update with kβ given the assignment g. more formally: • (w, f)jattribution(α, β)k(w, g) iff g ⊃ f and ∃x∃p.kα |= akα(x, p) and (w, f)jkαk(w, g), andg(p) ⊆ {(w′, k) : (w′, g)jkβk(w′, k)} 20. thanks to laure vieu for discussions on this point. 18 reports in discourse where akα is an attitude predicate determined by the embedding verb, e.g. said, thought, etc., contributed by α. to capture the discourse function of discourse parenthetical reports, we need to model the relation between the report and the incoming discourse context. i propose that this relation has three features, which come together to make a report discourse parenthetical. first, the embedded clause must be related directly to a discourse unit introduced in the discourse preceding the report. this feature is common to all attachment-based solutions. second, the relation must be distinct from the relation via which the attribution predicate is related to the incoming context (if the attribution predicate is so related). third, the relation connecting the embedded clause to the preceding context will be a modal relation indicating a hedged commitment from the speaker.21 the function of (1b), for example, is to provide a possible explanation of john’s absence from the party. modal relations are triggered by reports as follows. suppose we have an instance of attribution with the form attribution(α, β) and suppose that there is a discourse link whose second argument is β and whose first argument, γ, is found in the discourse preceding α. the link between γ and β will be labelled ♦r for some discourse relation r. more precisely, we can add the rule pa, for parenthetical attributions, to the logic of sdrt: pa: (∃e.e(α, β) ∧ l(e) = attribution)→ (∃γ∃e′(γ (‘ok. ’) from the dialogue corpus 5 classification of discourse markers this section describes a set of machine learning experiments aiming at automatically distinguishing between discourse markers, disfluencies, and segments that are neither discourse markers nor disfluencies (sus), using acoustic-prosodic features. in order to do so, we have performed intra-domain and cross-domain multiclass classification experiments using three sets of acoustic-prosodic features. we have performed an analysis of the most relevant prosodic features that can be used to characterize the discourse markers, and ranked the most prominent acoustic-prosodic features in the three classes. 5.1 prosodic parameters and machine learning methods our experiments are based on three sets of features, automatically extracted using the opensmile toolkit (eyben et al., 2010), a toolkit capable of extracting a very wide range of acoustic-prosodic features that has been successfully applied to a number of paralinguistic classification tasks, including disfluency prediction (schuller et al., 2013). the first set of features, henceforth referred as is13, is derived from the interspeech 2013 paralinguistic challenge, and corresponds to 6125 speech features calculated by applying segment-level statistics (means, moments, distances) over a set of energy, spectral and voicing related frame-level features. the other two sets of features were recently introduced by eyben et al. (2016), and correspond to gemaps – geneva minimalistic acoustic parameter set for voice research and affective computing, and egemaps – an extended version of gemaps. gemaps comprises a total of 62 acoustic-prosodic features, whereas its extended version, egemaps, comprises a set of 88 features. these small sets of features are well-known for their usefulness in a wide range of paralinguistic tasks, and are derived from: frequency related parameters (for example, pitch, jitter, and formant 1, 2, and 3 frequency); energy related parameters (such as shimmer, loudness, and harmonics-to-noise ratio – hnr); spectral parameters (like alpha cabarrão, moniz, batista, ferreira, trancoso and mata 94 ratio, spectral slopes, and harmonic differences); loudness, pitch (for voiced and unvoiced regions), and temporal features. in the scope of this work, we have applied a wide range of machine learning methods by means of the open source toolkit weka (hall et al., 2009) and the best results were consistently achieved by logistic regression (lr) and support vector machines (svm), the latter using the sequential minimal optimization (smo) algorithm. for that reason, results reported here are for these two methods. in most cases, our baseline consists of selecting the most frequent class, which was also achieved by applying the zeror classification method from weka. experiments reported on the paper are based on 5-fold cross-validation, thus covering all the data both for training and for evaluation. we have also performed 10-fold cross-validation experiments, but the corresponding experiments take about twice the time and results turned out to be very similar. the parameter c (complexity) used in svms defines the complexity of the model. the default value (c=1.0) is commonly used in similar experiments, and proved to be a good choice for the two smaller feature sets. however, for the large feature set (is13), the default value usually leads to poor performance, while taking several days to run, sometimes more than a week. for that reason, based in a previous work with disfluencies, we have set c to 0.01 for this feature set, which also proved to be suitable for the features in use while achieving a good speed performance. we did not tune this parameter in our cross-validation experiments, but we have tested other values of c for all feature sets in order to make sure that the parameter was correctly picked. results here presented were evaluated using the following standard performance metrics: precision, recall, f-measure, and accuracy (makhoul et al., 1999): precision = 𝑇𝑃 𝑇𝑃 + 𝐹𝑃 , recall = 𝑇𝑃 𝑇𝑃 + 𝐹𝑁 f-measure = 2 × 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 ∗ 𝑅𝑒𝑐𝑎𝑙𝑙 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 + 𝑅𝑒𝑐𝑎𝑙𝑙 , accuracy = 𝑇𝑃 + 𝑇𝑁 𝑇𝑃 + 𝐹𝑃 + 𝐹𝑁 + 𝑇𝑁 , where tp is the number of hits (true positives), tn is the number of correct rejections (true negatives), fp is the number of false alarms (false positives), and fn is the number of misses (false negatives). our experiments are also measured in terms of the kappa statistic, a chancecorrected measure of agreement between the classifications and the true classes. a value close to zero indicates that results could be achieved almost by chance; whereas a value close to 1.0 means that the model is adequate to the problem. 5.2 intra-domain classification the number of instances and distribution of discourse markers, disfluencies, and sus is quite different for both corpora: the dialogue corpus contains 6381 sus, 1834 disfluencies, and 723 discourse markers, and the lectures corpus contains 16273 sus, 6618 disfluencies, and 1359 discourse markers. this unbalanced data presents an additional challenge for the classification approaches that may lead to biased results towards the most common class, namely sus. table 3 presents the automatic classification results for each corpus, together with the baselines of 71.4% and 67.1%, corresponding to the most frequent class for each corpus. the table shows an overall better performance for the university lectures, despite the baseline being lower, which can be partially explained by the higher number of speakers in the dialogues (7 vs. 20 speakers), corresponding to a higher speaker variation. that is also reflected in the kappa values, all of them above 0.60 for the lectures. the acoustic-prosodic features in use allow for very significant improvements relative to the baseline: 20% for the university lectures and 13% for the dialogues. the best results were achieved using the svms and the is13 feature set, achieving accuracies of about 84% and 87%. however, the two other smaller feature sets also proved to be a good choice for discriminating discourse markers, especially when used in combination with lr. on the other discourse markers in european portugese 95 table 3 classification results for unbalanced data hand, lr consistently achieved quite low performance when applied to the is13 feature set while also taking several days to run. the gemaps and their extended version achieved acceptable performance while taking minutes to run, instead of hours and even days. the last 6 columns show individual performance results for discourse markers and disfluencies. with respect to te dialogues, we achieved 85.9% precision and 94.7% recall for sus, 77.1% precision and 55.7% recall for disfluences, and 77.8% precision and 62.5% recall for discourse markers, revealing that discourse markers are easier to classify than disfluencies in general. that is not the case for the university lectures, where we achieved 17% precision and 62.3% recall for discourse markers and 83.1% precision and 73.4% recall for disfluencies. it is interesting to notice that the model created using gemaps and svms for dialogues achieved 100% precision with a very low recall. that is because the model classified only one event as a discourse marker and it turned out to be correct. table 4 confusion matrix for unbalanced data, achieved using the is13 feature set and svms as previously mentioned, the data is considerably unbalanced, thus making it more difficult for the machine learning method to classify discourse markers. table 4 shows the confusion matrix for the two corpora, revealing that su is, in fact, the class selected more often, followed by disfluencies and, finally, discourse markers. in order to prevent biasing the models towards the most frequent class, we conducted the remaining experiments over a balanced version of the data. we considered the frequency of the discourse markers, the least frequent class in each corpus, as the number of samples to select for each class. so, for dialogues we selected 723 samples of each class and for university lectures, this number was extended to 1359. the balanced version of the data was achieved using the filter spreadsubsample available in weka. the corresponding results are presented in table 5. dialogues lectures classified as => su disf dm su disf dm su 6041 264 76 su 15349 773 151 disfluency (disf) 759 1022 53 disf 1642 4859 117 discourse marker (dm) 231 40 452 dm 297 215 847 cabarrão, moniz, batista, ferreira, trancoso and mata 96 table 5. classification results for balanced corpora once again, svms with the is13 feature set achieve the best performance. notice that the accuracy baseline is now 33.3% and that these results should not be directly compared with the results from table 3. kappa values are high for both corpora, revealing that we achieved good models for the three structures. the last six columns show an impressive f-measure of 82.9%87.5% for discourse markers. notice however that discourse markers here being considered are exclusively turn-initial while disfluencies include all possible types of disfluency in distinct positions of the corpus. the discourse marker classification performance is now better than for disfluencies, even on the university lectures. 5.3 cross-domain classification in order to verify how robust our classification is across domains, we conducted additional experiments in a cross-domain evaluation scenario, using the training data from one corpus and using the other corpus for testing. thus, to test if the model trained with the university lectures would generalize for the dialogues, we trained the models with the balanced data from lectra and tested on coral, and vice-versa. table 6 shows the corresponding results. it is known that by using out-of-domain data the performance decreases so, as expected, results are worse than the ones achieved using data from the same domain. in fact, the accuracy performance decreases about 11%-12% absolute. however, the performance achieved suggests that data from university lectures can still be used to classify the events on the dialogues and vice-versa. the egemaps feature set achieved about 71% and 66% accuracy in university lectures and dialogues, respectively, an impressive result when considering the small size of the feature set. the kappa statistic is between 0.666 and 0.717, also suggesting strong and suitable cross-domain models. an interesting additional result is that, when using balanced data, our results suggest that discourse markers are easier to identify than disfluencies, in both corpora, either using intra-domain or cross-domain models. disfluencies are recognized as being one of the hardest structures to detect amongst the structural metadata events (moniz et al., 2016). it is also recognized that the intonation of discourse markers are marker specific and they may share the melodic contours of a neutral declarative in ep, h+l* l%, but they are uttered with a wide pitch range, very distinctive from other units, and do tend to have three main prosodic patterns as previously stated. discourse markers in european portugese 97 table 6. cross-domain classification results for balanced corpora table 7 presents four confusion matrices, achieved with the balanced data, showing the most common correct and incorrect automatic classifications. all the matrices show a strong diagonal, indicating that our models, both inter-corpora and cross-corpora, perform correct classification most of the time. there is also clear evidence that discourse markers are better identified with indomain data, in line with what was previously said concerning the distributional patterns of such events, i.e., they may occur exclusively in one corpus or be more productive in one corpus than in the other. when using cross-domain models, the tendency is to classify discourse markers as disfluencies and, to a smaller degree, to classify them as sus. this trend may also be explained by the shared properties between discourse markers and certain types of disfluencies, mostly filled pauses, behaving as vocalic supports with plateau contours (h* h-/%). table 7. confusion matrices for intra-domain and cross-domain results achieved with egemaps, balanced data, and smo with c=1 intra-domain dialogues lectures classified as => su disf dm su disf dm su 504 105 114 su 1128 116 115 disfluency (disf) 141 448 134 disf 252 899 208 discourse marker (dm) 75 44 604 dm 51 121 1187 cross-domain lectures => dialogues dialogues => lectures classified as => su disf dm su disf dm su 553 109 61 su 985 220 154 disfluency (disf) 187 440 96 disf 229 944 186 discourse marker (dm) 171 113 439 dm 94 316 949 cabarrão, moniz, batista, ferreira, trancoso and mata 98 5.4 discussion of relevant features and acoustic-prosodic patterns of structural metadata events the automatic discrimination of structural metadata events is a complex task, due to the diversity of acoustic-prosodic and distributional patterns of such events and also to the prosodic information they may share. however, despite the complexity of such structures, they have discriminative prosodic behaviors captured by acoustic correlates. the literature for portuguese points out to an array of features relevant for the description of metadata events. the data-driven approaches followed in this work allow us to reach a structured set of basic features towards the disambiguation of such events beyond the established evidences for portuguese. this study is a contribution to the analysis and discriminative behavior of discourse markers in ep (to the best of our knowledge, completely absent in our literature), and adds to the previous studies on structural metadata events a layer more to the understanding of the acoustic-prosodic behavior of such structures. the automatic discrimination between classes of structural metadata events is feasible because discourse markers have class properties, as will be detailed bellow. sentence-like units and disfluencies were previously discriminated (moniz et al., 2016). disfluencies are mostly characterized by: two identical contiguous words; both energy and pitch increases in the following word and (mostly) a plateau contour on the preceding word; and a higher confidence level for the following word than for the previous word. this set of features reveals that repetitions are being identified, that repair regions are characterized by prosodic contrast marking (increases in pitch and energy) between disfluency-fluency repair (as in moniz et al., 2012 and moniz, 2013), and also that the first word of the repair has a higher confidence score. since repetitions are more frequent in dialogues (22% of all disfluencies vs. 16% in lectures), the feature identical contiguous words has a significantly higher impact in dialogues. full stops are described by a falling contour in the previous word; a plateau energy slope in the previous word; the duration ratio between the previous and the following words; and previous word higher confidence score. this characterization is the one that most resembles neutral statements in portuguese, with the canonical contour h+l* l% (frota, 2001), associated with terminus value. question marks in lectures are characterized by two main patterns: a rising contour in the current word and a rising/rising energy slope between previous and following words; and a plateau pitch contour in the previous word and a falling energy slope in the previous word. the rising patterns are not surprising, since interrogatives are cross-language perceived as having a rising contour (hirst & di cristo, 1998). the falling pitch contours have also been ascribed for different types of interrogatives, especially whquestions in portuguese. commas are the event characterized by the fewest prosodic features, being mostly identified by morpho-syntactic features. however, in dialogues they are better classified. the two most relevant features are: identical contiguous words and mostly plateau energy and pitch shapes between words. the first feature is associated with emphatic repetitions, comprising several structures, namely: (i) affirmative or negative backchannels (sim, sim, sim ‘yes, yes, yes’) and (ii) repetition of a syntactic phrase, such as a locative prepositional phrase (para cima, para cima ‘above, above’). they are used for precise tuning with the follower and for stressing the most important part of the instruction. emphatic repetitions are annotated with commas separating the repeated item(s) and account for 1% of the total number of words in dialogues. although not a disfluent structure, if they were accounted as disfluent words, they would represent 16.7% of all disfluent items. as for the features energy and pitch plateau shapes between words, they are linked to lists of enumerated names of the landmarks in a given map, at the end of a dialogue. regarding regular words, the most salient features are related to the absence of silent pauses, explained by the fact that, contrary to the other events, regular words within phrases are connected. the presence of a silent pause is a strong cue to the assignment of a structural metadata event. discourse markers in european portugese 99 adding discourse markers to this characterization, they are described as: having different prosodic patterns: (i) as a major intonational phrase (ip) (see figure 3); (ii) as an intermediate intonational phrase (ip); (iii) deaccented and functioning as a clitic or an initial vocalic support (see figure 4), with plateau contours and f0 values lower than the following prosodic unit, being, therefore, uttered in an intermediate tonal space; (iv) with reduced f0 slopes relatively to the following prosodic unit (e.g., então, ‘so’ and pronto, ‘ok’). moreover, the pitch range of discourse markers behaves as a continuum from a very compressed to a very wide range. table 8. top 25 most influent features table 8 presents the most relevant features extracted from the set of egemaps features with in-domain data. the choice of this particular set is motivated by: (i) the substantial dimensionality reduction regarding the opensmile features; and (ii) the faster classification of structural metadata events, with comparable results. the “*” indicates the most informative features, considering their weights extracted with logistic regression models – the higher number of “*” the higher the relevance. the ranking is distributed per corpus, encompassing all the structural metadata events, and also per the most striking differences between discourse markers and disfluencies. we note that the most informative features are the pitch and energy related ones. considering the prosodic characterization described above, this result may be interpreted as pointing out to the different f0 slopes in the production of discourse markers, an evidence more of the continuum in the pitch range of such structures. cabarrão, moniz, batista, ferreira, trancoso and mata 100 figure 3: example of the discourse marker agora (‘now’) in the excerpt: agora, isto que aqui está é apenas um conjunto de classes (‘now, what we have here is just a set of classes’), from the university lectures corpus the discourse marker agora (‘now’ – figure 3) tends to be accented, with high f0 range within the accented syllable, and with similar or higher f0 values than the adjacent prosodic constituents. this shows that there is an effort to mark this word and to distinguish it from the adjacent contexts. on the other hand, the markers portanto (‘ok’ – figure 4) and ok are mainly unaccented, present plateau contours, and have f0 values lower than the following prosodic unit, being, therefore, uttered in an intermediate tonal space, almost as a vocalic support for uttering the following prosodic units. both então (‘so’), pronto (‘ok’) and portanto (ok) present reduced f0 slopes, even though the latter tends to be unaccented and the first one is mainly accented. figure 4: example of the discourse marker portanto (ok) in the excerpt: portanto, contornaste o solar dos mil amores por fora (‘ok, you’ve passed the manor of a thousand loves from the outside’), from the dialogues corpus discourse markers in european portugese 101 the prosodic contours can be distinct accordingly to their distribution, i.e., initial, medial or final positions. since we’re targeting the turn-initial ones, the prosodic patterns of such markers highly depend on the marker selected, meaning if it is mostly a deaccented portanto (‘ok’) vs. a prominent agora (‘now’), starting a new topic with a wide range of pitch and energy. the fact that discourse markers are associated with different pragmatic functions may also be an explanation for the variation found. we can hypothesize that the markers that have a function similar to disfluencies, like stalling, may share with them some prosodic properties, meaning the plateau contours contrasting with the rises in the following prosodic constituents. other discourse markers, such as agora (‘now’), introduce a new topic and are prosodically prominent. to confirm this preliminary analysis, we still need to perform a more detailed tonal and acoustic study of the discourse markers and adjacent prosodic contexts. the acoustic-prosodic characterization of discourse markers, in particular, and structural metadata events, in general, is still a matter of debate. structural metadata events can be equated to discourse markers as a broad linguistic class, which encompasses both disfluencies and discourse markers5. the linguistic literature on discourse markers points out to several strategies when accounting for disfluencies in such a broader class: (i) either consider fillers and reformulation markers as a sub-type of discourse markers, but always stating that they have a distinct nature, although not clarifying truly this nature with empirical evidence; or (ii) consider them as completely different and not part of discourse markers. we should add a third strategy, a flawed or limbo classification, a fuzzy one with no clear criteria for either the inclusion or exclusion of such events into a broader class of discourse markers as a whole. it is our belief that data-driven studies based exclusively on acoustic-prosodic features shed light on the classification of discourse markers, since they are generally defined as syntactically detached structures with no propositional content, and bring to light the discriminative prosodic behavior of disfluencies as a legitimate subtype of discourse markers. in future work, we aim at tackling prosodic parameters per discourse markers subtypes, bridging the acoustic-prosodic properties to the discourse derived sub-categorization of this broad class. 6 conclusions and future work this work presented our first attempt to describe discourse markers in three different corpora in ep, namely university lectures, map-task dialogues, and a collection of tweets. our main goal was to analyze the type of discourse markers used in the various domains. our results showed that the selection of discourse markers is domain and speaker dependent. even in the same corpus, there are speakers that tend to use the same discourse marker and there are those who vary amongst several structures. we also found that the most frequent discourse markers are similar in all three corpora. however, in the collection of tweets there are discourse markers that do not occur in the other two corpora. these are discourse markers typically used in more informal contexts, and are clearly produced by teenagers and young adults. in this multidisciplinary study, comprising both a linguistic perspective and a computational approach, discourse markers are also automatically discriminated from other structural metadata events, namely sentence-like units and disfluencies. as for their distributional patterns, our results showed that markers and disfluencies tend to co-occur in the dialogue corpus, but have a complementary distribution in the university lectures. 5 in this perspective, punctuation marks are not accounted for, since they are part of sentence-type structures of different illocution values. cabarrão, moniz, batista, ferreira, trancoso and mata 102 we used three acoustic-prosodic feature sets and machine learning to automatically distinguish between discourse markers, disfluencies and sus. our in-domain experiments achieved an accuracy of about 87% in university lectures and 84% in dialogues, in line with our previous results. the gemaps, and especially the egemaps features, recently introduced for voice research and affective computing and commonly used for other paralinguistic tasks, achieved good performance on our data, while taking minutes to run instead of hours and even days. our results suggest that turn-initial discourse markers are usually easier to classify than disfluencies, a result also previously reported in the literature. we replicated the multiclass experiments in a cross-domain evaluation in order to evaluate how robust are the models across domains. the results achieved are about 11%-12% lower than when in-domain data is used, but we have concluded that data from one domain can still be used to classify the same events in the other. these results allow us to hypothesize that it will be possible to classify discourse markers in out-of-domain data. overall, despite the complexity of this task, these are very encouraging state-of-the-art results for multiclass classification. both discourse markers and disfluencies are described in the literature as sharing some properties (goldwater et al., 2010; liu et al., 2006). ultimately, using exclusively acoustic--prosodic cues, discourse markers can be fairly discriminated from disfluencies and sus. in order to better understand the contribution of each feature, we have also analyzed the impact of the features both in dialogues and university lectures. pitch features, namely pitch slopes, are the most relevant ones for the distinction between discourse markers and disfluencies. these features are in line with the wide pitch range of discourse markers, in a continuum from a very compressed pitch range to a very wide one, expressed by total deaccented material or h+l* l* contours, with upstep h tones. in future work, we intend to do a more exhaustive analysis of the most prominent acoustic-prosodic features and also to do a tonal analysis of the discourse markers and adjacent prosodic contexts, to understand the relevance of the discourse marker in the classification process. we also intend to include discourse markers in the language models for ep already trained with other structural metadata events, which will result in enriched automatic transcriptions, and to integrate the classifiers in spoken dialogue systems. furthermore, we also intend to study morpho-syntactic features of the discourse markers, as well as paralinguistic events, such as laughs, nodding, and hand gestures. we would also tackle the cross-language analysis of discourse markers, in order to verify how specific or universal are their behavior. acknowledgements this work was supported by national funds through fundação para a ciência e a tecnologia (fct) with reference uid/cec/50021/2013, and under phd grant sfrh/bd/96492/2013, and post-doc grant sfrh/pbd/95849/2013. references jean-michel adam (2008). a lingüística textual. introdução à análise textual dos discursos. são paulo: cortez editora. karin aijmer (2013). understanding pragmatic markers: a variational pragmatic approach. edinburgh: edinburgh university press. karin aijmer, ad foolen and anne-marie simon-vandenbergen (2006). pragmatic markers in translation: a methodological proposal. approaches to discourse particles, 1:101-114. fernando batista (2011). recovering capitalization and punctuation marks on speech transcriptions. phd thesis, instituto superior técnico, lisbon. discourse markers in european portugese 103 fernando batista, helena moniz, isabel trancoso, nuno mamede and ana isabel mata (2012). extending automatic transcripts in a unified data representation towards a prosodic-based metadata annotation and evaluation. journal of speech sciences, 2:115-138. fernando batista, pedro dos santos lopes curto, isabel trancoso, alberto abad, jaime rodrigues ferreira, eugénio alves ribeiro, helena moniz, david martins de matos, ricardo ribeiro (2016). spa: web-based platform for easy access to speech processing modules. in lrec, european language resources association (elra), pages 3886-3892, doi: isbn: 978-2-9517408-9-1, portoroz, slovenia. kate beeching (2002). gender, politeness and pragmatic particles in french. amsterdam: john benjamins. kate beeching (2014). just a suggestion pragmatic marker just/e in french and english. oral communication in the international workshop pragmatic markers, discourse markers and modal particles: what do we know and where do we go from here?, università dell'insubria, como (italy). margarita borreguero and araceli lópez (2010). los marcadores del discurso y la variación lengua hablada vs. lengua escrita. los estudios sobre marcadores del discurso en español, hoy. madrid: arco/libros: 415-496. gaspar brogueira, fernando batista, joão paulo carvalho, helena moniz (2014). expanding a database of portuguese tweets. in slate'14 3rd symposium on languages, applications and technologies, schloss dagstuhl, vol. 4569, series openaccess series in informatics (oasics), pages 275-282, bragança, portugal. vera cabarrão, helena moniz, jaime ferreira, fernando batista, isabel trancoso, ana isabel mata, sérgio curto (2015). prosodic classification of discourse markers. in the scottish consortium for icphs 2015 (ed.), proceedings of the 18th international congress of phonetic sciences, pages 14-17, glasgow, uk: the university of glasgow. maria antónia coutinho (2009). marcadores discursivos e tipos de discurso. linguistic studies 2:193-210. annika denke (2009). nativelike performance. pragmatic markers, repair and repetition in native and non-native english speech. vdm publishing. inês duarte (2000). língua portuguesa. instrumentos de análise. lisbon: universidade aberta. florian eyben, martin wöllmer, bjorn schuller (2010). opensmile the munich versatile and fast open-source audio feature extractor. acm multimedia (mm), acm, firenze, italy. florian eyben, klaus r. scherer, björn w. schuller, johan sundberg, elisabeth andré, carlos busso, laurence y. devillers, julian epps, petri laukka, shrikanth narayanan, and khiet truong (2016). the geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing. ieee transactions on affective computing 7, no. 2:190202. kerstin fischer (1998). discourse particles, turn-taking, and the semantics-pragmatics interface. revue de sémantique et pragmatique, 8:111–137. kerstin fischer (2000). from cognitive semantics to lexical pragmatics: the functional polysemy of discourse particles. mouton de gruyter, berlin/new york. bruce fraser (1988). types of english discourse markers. acta linguistica hungarica, 38:19-33. bruce fraser (1990). an approach to discourse markers. journal of pragmatics, 14:383–395. bruce fraser (1999). what are discourse markers? journal of pragmatics, 31:931–952. bruce fraser (2009). topic orientation markers. journal of pragmatics, 41:892-898 tiago freitas, maria celeste ramilo (2003). o actual estatuto da palavra portanto. in actas do xviii encontro da associação portuguesa de linguística, pages 357-369, lisbon. cabarrão, moniz, batista, ferreira, trancoso and mata 104 sónia frota (2000). prosody and focus in european portuguese: phonological phrasing and intonation. new york, garland publishing. sónia frota (2014). the intonational phonology of european portuguese. in sun-ah jun (ed.) prosodic typology ii: 6-42. oxford: oxford university press. sharon goldwater, dan jurafsky, christopher d. manning (2010). which words are hard to recognize? prosodic, lexical, and disfluency factors that increase speech recognition error rates. speech communication, 52:181–200. agustin gravano, julia hirschberg, and štefan beňuš. (2012). affirmative cue words in taskoriented dialogue. computational linguistics, 38(1):1-39. mark hall, eibe frank, geoffrey holmes, bernhard pfahringer, peter reutemann and ian h. witten (2009). the weka data mining software: an update. sigkdd explorations, 11(1):10-18. peter heeman and james allen (1999). speech repairs, intonational phrases and discourse markers: modeling speakers’ utterances in spoken dialogue. computational linguistics, 25:527–571. julia hirschberg and diane j. litman (1993). empirical studies on the disambiguation of cue phrases. computational linguistics, 19 (3):501–530. daniel hirst and albert di cristo (1998). intonation systems: a survey of twenty languages. cambridge university press. andreas h. jucker and sara w. smith (1998). and people just you know like “wow‟: discourse markers as negotiating strategies. in jucker, andreas h. and yael ziv (eds). discourse markers: description and theory, 57:171. amsterdam: john benjamins. philipp koehn (2005). europarl: a parallel corpus for statistical machine translation. in mt summit, vol. 5, pages 79-86. matthew lease and mark johnson (2006). early deletion of fillers in processing conversational speech. in proceedings of hlt-naacl 2006 (human language technology conference of the north american chapter of the association of computational linguistics), companion volume: short papers, pages 73–76. new york, ny. rivka levitan, stefan benus, agustin gravano, and julia hirschberg (2015). entrainment and turn-taking in human-human dialogue. in aaai spring symposium on turn-taking and coordination in human-machine interaction, stanford, ca. antónio calado lopes, david martins de matos, vera cabarrão, ricardo ribeiro, helena moniz, isabel trancoso, ana isabel mata (2015). towards using machine translation techniques to induce multilingual lexica of discourse markers. http://arxiv.org/abs/1503.0914. ana cristina macário lopes (1997). então: elementos para uma análise semântica e pragmática. actas do xii encontro nacional da apl, vol. 1:177-189. lisbon: colibri. ana cristina macário lopes (2009). justification: a coherence relation. pragmatics, 19(2): 223239. ana cristina macário lopes (2014a). aliás: a contribution to the study of a portuguese discourse marker. in piera molinelli & chiara ghezzi, (eds.), discourse and pragmatic markers from latin to the romance languages, (9). oxford university press. ana cristina macário lopes and patrícia amaral (2006). from time to discourse monitoring: agora e então in european portuguese. in corneille, b. and delbecque, n. (eds.), topics in subjectification and moralisation. belgian journal of linguistics, (20):3-18. ana cristina macário lopes and sara sousa (2014b). the discourse connectives ao invés and pelo contrário in european contemporary portuguese. journal of portuguese linguistics, 13(1):3-27. ana cristina macário lopes (2016). discourse markers 24. in wetzel, leo, menuzzi, sergio & costa, joão (eds.). the handbook of portuguese linguistics: 441-456. new york: wileyblackwell. discourse markers in european portugese 105 yang liu, elizabeth shriberg, andreas stolcke, dustin hillard, mari ostendorf and mary harper (2006). enriching speech recognition with automatic detection of sentence boundaries and disfluencies. in ieee transactions on audio, speech, and language processing, vol. 14, n. 5, pages 1526-1540. john makhoul, francis kubala, richard schwartz and ralph weischedel (1999). performance measures for information extraction. in proceedings of the darpa broadcast news workshop, pages 249-252. herndon, va. ana isabel mata and helena moniz (2016). prosódia, variação e processamento automático. in martins, ana maria & ernestina carrilho (eds). manual de linguística portuguesa (vol. 16). mrl series. de gruyter. amália mendes (2013). organização textual e articulação de orações. in raposo, eduardo b. p., m. fernanda bacelar do nascimento, m. antónia coelho da mota, luísa segura, amália mendes (orgs) gramática do português, vol. ii: 1691-1755. lisboa: fundação calouste gulbenkian. helena moniz (2006). contributo para a caracterização dos mecanismos de (dis)fluência no português europeu. ma thesis, university of lisbon. helena moniz (2013). processing disfluencies in european portuguese. phd thesis, university of lisbon. helena moniz, fernando batista, isabel trancoso, ana isabel mata da silva (2012). prosodic context-based analysis of disfluencies. in interspeech 2012, isca, pages 1961-1964. portland, oregon, u.s.a. helena moniz, fernando batista, ana isabel mata, isabel trancoso (2014). speaking style effects in the production of disfluencies. speech communication, vol. 65:20-35. helena moniz, jaime rodrigues ferreira, fernando batista, isabel trancoso (2015). disfluency detection across domains. in diss 2015, edinburgh, scotland, u. k. helena moniz, fernando batista, ana isabel mata, isabel trancoso (2016). towards automatic language processing and intonational labeling in european portuguese. in m. armstrong, n. c. henriksen, & m. m. vanrell (eds.), intonational grammar in ibero-romance: approaches across linguistic subfields, pages 295-324. philadelphia, usa: john benjamins. joão neto, hugo meinedo, márcio viveiros, renato cassaca, ciro martins and diamantino caseiro (2008) broadcast news subtitling system in portuguese. in proceedings of icassp’08, pages1561–1564. las vegas, usa. mari ostendorf, benoît favre, ralph grishman, dilek hakkani-tür, mary harper, dustin hillard, julia hirschberg, heng ji, jeremy g. kahn, yang liu, sameer maskey, evgeny matusov, hermann ney, andrew rosenberg, elizabeth shriberg, wen wang, and chuck wooters (2008). speech segmentation and spoken document processing. in ieee signal processing magazine, pages 59-69. janet pierrehumbert and julia hirschberg (1990). the meaning of intonational contours in the interpretation of discourse. intentions in communication:271-311. ana pimentel (2012). os marcadores conversacionais no ensino de português língua estrangeira: um estudo de caso. ma thesis, faculty of letters, university of oporto. andrei popescu-bellis, sandrine zufferey (2011). automatic identification of discourse markers in dialogues: an in-depth study of like and well. computer speech & language 25 (3):499-518. gisela redeker (1990). ideational and pragmatic markers of discourse structure. journal of pragmatics, 14:367-381. cabarrão, moniz, batista, ferreira, trancoso and mata 106 kenneth brian samuel (1999). discourse learning: an investigation of dialogue act tagging using transformation-based learning. phd thesis, university of delaware. deborah schiffrin (1987). discourse markers. cambridge university press, cambridge. deborah schiffrin (2001). discourse markers: language meaning and context. in deborah schiffrin, deborah tannen and heidi hamilton (eds.) handbook of discourse analysis. oxford: basil blackwell. lawrence schourup (1999). discourse markers. lingua, 107: 227-65. björn schuller, stefan steidl, anton batliner, felix burkhardt, laurence devillers, christian müller, and shrikanth narayanan (2013). paralinguistics in speech and language state-ofthe-art and the challenge. computer speech & language, 27, no. 1: 4-39. elizabeth shriberg (1994). preliminaries to a theory of speech disfluencies. phd thesis, university of califórnia. elizabeth shriberg (2001). to "errrr" is human: ecology and acoustics of speech disfluencies. journal of the international phonetic association, 31:153–169. augusto soares da silva (2006). the polysemy of discourse markers: the case of pronto in portuguese. journal of pragmatics, 38:2188-2205. isabel trancoso, maria do céu viana, inês duarte, and gabriela matos (1998). corpus de diálogo coral. in proceedings of propor'98, porto alegre, brasil. isabel trancoso , rui martins, helena moniz, ana isabel mata, and maria do céu viana (2008). the lectra corpus classroom lecture transcriptions in european portuguese. in proceedings lrec'08, marrakech, morocco. hudinilson urbano (2003). marcadores conversacionais. in preti, d. (org.). análise de textos orais. 6a edição:93-116. são paulo: humanitas. maria do céu viana, sónia frota, isabel falé, flaviane fernandes, isabel mascarenhas, ana isabel mata, helena moniz & marina vigário (2007). towards a p_tobi. in http://www.ling.ohio-state.edu/~tobi/. dialogue & discourse 7(3) (2016) 65–88 doi: 10.5087/dad.2016.303 recurrent polynomial network for dialogue state tracking∗ kai sun ks985@cornell.edu department of computer science cornell university, ithaca, ny 14853, usa qizhe xie cheezer@sjtu.edu.cn key laboratory of shanghai education commission for intelligent interaction and cognitive engineering speechlab, department of computer science and engineering shanghai jiao tong university, shanghai, 200240, china kai yu† kai.yu@sjtu.edu.cn key laboratory of shanghai education commission for intelligent interaction and cognitive engineering speechlab, department of computer science and engineering shanghai jiao tong university, shanghai, 200240, china editor: jason d. williams, antoine raux, and matthew henderson submitted 04/15; accepted 02/16; published online 04/16 abstract dialogue state tracking (dst) is a process to estimate the distribution of the dialogue states as a dialogue progresses. recent studies on constrained markov bayesian polynomial (cmbp) framework take the first step towards bridging the gap between rule-based and data-driven approaches for dst. in this paper, a novel hybrid framework – recurrent polynomial network (rpn) is proposed to further improve the combination of rule-based and data-driven approaches. rpn’s unique structure enables the framework to have all the advantages of cmbp including efficiency, portability and interpretability. additionally, rpn achieves more properties of data-driven approaches than cmbp. rpn was evaluated on the data corpora of the second and the third dialog state tracking challenge (dstc-2/3). experiments showed that rpn can significantly outperform both traditional rule-based approaches and data-driven statistical approaches with similar feature set. compared with the state-of-the-art data-driven statistical dst approaches with a lot richer features, rpn is also competitive. keywords: statistical dialogue management, dialogue state tracking, recurrent polynomial network 1. introduction a task-oriented spoken dialogue system (sds) is a system that can interact with a user to accomplish a predefined task through speech. it usually has three modules: input, output and control, shown in figure 1. the input module consists of automatic speech recognition (asr) and spoken language ∗. this work was supported by the program for professor of special appointment (eastern scholar) at shanghai institutions of higher learning and the china nsfc project no. 61222208. †. corresponding author c©2016 kai sun, qizhe xie, and kai yu this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). sun, xie and yu understanding (slu), with which the user speech is converted into text and semantics-level user dialogue acts are extracted. once the user dialogue acts are received, the control module, also called dialogue management accomplishes two missions. one mission is called dialogue state tracking (dst), which is a process to estimate the distribution of the dialogue states, an encoding of the machine’s understanding about the conversion as a dialogue progresses. another mission is to choose semantics-level machine dialogue acts to direct the dialogue given the information of the dialogue state, referred to as dialogue decision making. the output module converts the machine acts into text via natural language generation and generates speech according to the text via text-tospeech synthesis. dialogue management is the core of a sds. traditionally, dialogue states are assumed to be observable and hand-crafted rules are employed for dialogue management in most commercial sdss. however, because of unpredictable user behaviour, inevitable asr and slu errors, dialogue state tracking and decision making are difficult (williams and young, 2007). consequently, in recent years, there is a research trend from rule-based dialogue management towards statistical dialogue management. partially observable markov decision process (pomdp) framework offers a wellfounded theory to both dialogue state tracking and decision making in statistical dialogue management (roy et al., 2000; zhang et al., 2001; williams and young, 2005, 2007; thomson and young, 2010; gašić and young, 2011; young et al., 2010). in previous studies of pomdp, dialogue state tracking and decision making are usually investigated together. in recent years, to advance the research of statistical dialogue management, the dst problem is raised out of the statistical dialogue management framework so that a bunch of models can be investigated for dst. user text-to-speech synthesis natural language generation automatic speech recognition spoken language understanding dialogue manager dialogue state tracking input module control module output module waveforms waveforms words words dialogue acts dialogue acts figure 1: diagram of a spoken dialogue system (sds) most early studies of pomdp on dst were devoted to generative models (young et al., 2010), which learn joint probability distribution over observation and labels. fundamental weaknesses of generative model was revealed by the result of (williams, 2012). in contrast, discriminative state tracking models have been successfully used for sdss (deng et al., 2013). compared to generative models where assumptions about probabilistic dependencies of features are usually needed, discriminative models directly model probability distribution of labels given observation, enabling rich features to be incorporated. the results of the dialog state tracking challenge (dstc) (williams et al., 2013; henderson et al., 2014b,a) further demonstrated the power of discriminative statisti66 recurrent polynomial network for dialogue state tracking cal models, such as maximum entropy (maxent) (lee and eskenazi, 2013), conditional random field (lee, 2013), deep neural network (dnn) (sun et al., 2014a), and recurrent neural network (rnn) (henderson et al., 2014d). in addition to discriminative statistical models, discriminative rule-based models have also been investigated for dst due to their efficiency, portability and interpretability and some of them showed good performance and generalisation ability in dstc (zilka et al., 2013; wang and lemon, 2013). however, both rule-based and statistical approaches have some disadvantages. statistical approaches have shown large variation in performance and poor generalisation ability due to the lack of data (williams, 2012). moreover, statistical models usually have more complex model structure and features than rule-based models, and thus can hardly achieve efficiency, portability and interpretability as rule-based models. as for rule-based models, their performance is usually not competitive to the best data-driven statistical approaches. additionally, since they require lots of expert knowledge and there is no general way to design rule-based models with prior knowledge, they are typically difficult to design and maintain. furthermore, there lacks a way to improve their performance when training data are available. recent studies on constrained markov bayesian polynomial (cmbp) framework take the first step towards bridging the gap between rule-based and data-driven approaches for dst (sun et al., 2014b; yu et al., 2015). cmbp formulate rule-based dst in a general way and allow data-driven rules to be generated. concretely, in the cmbp framework, dst models are defined as polynomial functions of a set of features whose coefficients are integer and satisfy a set of constraints where prior knowledge is encoded. the optimal dst model is selected by evaluating each model on training data. yu et al. (2015) further extended cmbp to real-coefficient polynomial where the real coefficients can be estimated by optimizing the dst performance on training data using grid search. cmbp offers a way to improve the performance when training data are available and achieves competitive performance to the state-of-the-art data-driven statistical approaches, while at the same time keeping most of the advantages of rule-based models. nevertheless, adding features to cmbp is not as easy as to most data-driven statistical approaches because on the one hand, the features usually need to be probability related features, on the other hand, additional prior knowledge is needed to constrain the search space. for the same reason, increasing the model complexity, such as by using higher-order polynomial, by introducing hidden variables, etc. also requires additional suitable prior knowledge to be introduced to limit the search space. moreover, cmbp can hardly fully utilize the labelled data because in practice its polynomial coefficients are set by grid search. in this paper, a novel hybrid framework, referred to as recurrent polynomial network (rpn), is proposed to further improve the combination of rule-based and data-driven approaches for dst. although the basic idea for transforming rules to neural networks has been there for many years (cloete and zurada, 2000), few work has been done for dialogue state tracking. rpn can be regarded as a kind of human interpretable computation network, and its unique structure enables the framework to have all the advantages of cmbp including efficiency, portability and interpretability. additionally, rpn achieves more properties of data-driven approaches than cmbp. in general, rpn has neither restriction to feature type, nor search space issue to be concerned about, so adding features and increasing the model complexity are much easier in rpn. furthermore, rpn can better explore the parameter space than cmbp with labelled data. the dstcs have provided the first common testbed in a standard format, along with a suite of evaluation metrics to facilitate direct comparisons among dst models (williams et al., 2013). to evaluate the effectiveness of rpn for dst, both the dataset from the second dialog state tracking challenge (dstc-2) which is in restaurants domain (henderson et al., 2014b) and the dataset from 67 sun, xie and yu the third dialog state tracking challenge (dstc-3) which is in tourists domain (henderson et al., 2014a) are used. for both of the datasets, the dialogue state tracker receives slun -best hypotheses for each user turn, each hypothesis having a set of act-slot-value tuples with a confidence score. the dialogue state tracker is supposed to output a set of distributions of the dialogue state. in this paper, only joint goal tracking, which is the most difficult and general task of dstc-2/3, is of interest. the rest of the paper is organized as follows. section 2 discusses ways of combining rule-based and data-driven approaches. section 3 formulates rpn. the rpn framework for dst is described in section 4, followed by experiments in section 5. finally, section 6 concludes the paper. 2. combining rule-based and data-driven approaches broadly, it is straightforward to come up with two possible ways to combine rule-based and datadriven approaches: one starts from rule-based models, while the other starts from data-driven models. cmbp takes the first way, which is derived as an extension of a rule-based model (sun et al., 2014b; yu et al., 2015). inspired by the observation that many rule-based models such as models proposed by wang and lemon (2013) and zilka et al. (2013) are based on bayes’ theorem, in the cmbp framework, a dst rule is defined as a polynomial function of a set of probabilities since bayes’ theorem is essentially summation and multiplication of probabilities. here, the polynomial coefficients can be seen as parameters. to make the model have good dst performance, prior knowledge or intuition is encoded to the polynomial functions by setting certain constraints to the polynomial coefficients, and the coefficients can further be optimized by data-driven optimization. therefore, starting from rule-based models, cmbp can directly incorporate prior knowledge or intuition into dst, while at the same time, the model is allowed to be data-driven. more concretely, assuming that both slot and value are independent, a cmbp model can be defined as bt+1(v) = p ( bt(v), p+ t+1(v), p−t+1(v), p̃+ t+1(v), p̃−t+1(v), 1, brt ) s.t. constraints (1) where bt(v), p+ t (v), p−t (v), p̃+ t (v), p̃−t (v), brt are all probabilistic features which are defined as below: • bt(v): belief of “the value being v at turn t” • p+ t (v): sum of scores of slu hypotheses informing or affirming value v at turn t • p−t (v): sum of scores of slu hypotheses denying or negating value v at turn t • p̃+ t (v) = ∑ v′ /∈{v,none} p + t (v′) • p̃−t (v) = ∑ v′ /∈{v,none} p − t (v′) • brt : belief of the value being ‘none’ (the value not mentioned) at turn t, i.e. brt = 1 −∑ v′ 6=none bt(v ′) 68 recurrent polynomial network for dialogue state tracking and p(·) is a multivariate polynomial function1 p(ι0, · · · , ιd) = ∑ 0≤k1≤···≤kn≤d gk1,··· ,kn ∏ 1≤i≤n ιki (2) where d + 1 is the number of input variables, n is the order of the polynomial, gk1,··· ,kn is the parameter of cmbp. order 3 gives good trade-off between complexity and performance, hence order 3 is used in our previous work (sun et al., 2014b; yu et al., 2015) and this paper. since both slot independence and value independence are assumed in this paper, the belief of the joint goal (slot1 = v1, slot2 = v2, . . . , slotn = vn ) is n∏ i=1 b(vi). the constraints in equation (1) encode all necessary probabilistic conditions (yu et al., 2015). for instance, 0 ≤ p+ t (v) + p̃+ t (v) ≤ 1 (3) 0 ≤ bt(v) ≤ 1 (4) the constraints in equation (1) also encode prior knowledge or intuition (yu et al., 2015). for example, the rule “goal belief should be unchanged or positively correlated with the positive scores from slu” can be represented by ∂p ( bt(v), p+ t+1(v), p−t+1(v), p̃+ t+1(v), p̃−t+1(v), 1, brt ) ∂p+ t+1(v) ≥ 0 (5) the definition of cmbp formulates a search space of rule-based models, where it is easy to employ data-driven criterion to find a rule-based model with good performance. considering cmbp is originally motivated from bayesian probability operation which leads to the natural use of integer polynomial coefficients (g ∈ z), the data-driven optimization can be formulated by an integer programming program (sun et al., 2014b; yu et al., 2015). additionally, cmbp can also be viewed as a data-driven approach. hence, the polynomial coefficients can be extended to real numbers. the optimization of real-coefficient can be done by first getting an integer-coefficient cmbp and then performing hill climbing search (yu et al., 2015). 3. recurrent polynomial network recurrent polynomial network, which is proposed in this paper, takes the other way to combine rulebased and data-driven approaches. the basic idea of rpn is to enable a kind of data-driven model to take advantage of prior knowledge or intuition by using the parameters of rule-based models to initialize the parameters of data-driven models. computational networks have been researched for decades from very basic architectures such as perceptron (rosenblatt, 1958) to today’s various kinds of deep neural networks. recurrent computational networks, a class of computational networks which have recurrent connections, have also been researched for a long time, from fully recurrent networks to networks with relatively complex structures such as long short-term memory (lstm) (hochreiter and schmidhuber, 1997). like common neural networks, rpn is a data-driven approach so it is as easy to add features and try 1. the notation of ∑ 0≤k1≤···≤kn≤d is shorthand for series of n nested sums over ranges bounded, i.e.∑ 0≤k1≤d ∑ k1≤k2≤d . . . ∑ kn−1≤kn≤d . 69 sun, xie and yu complex structures in rpn as in neural networks. however, compared with common neural networks which are “black boxes”, an rpn can essentially be seen as a polynomial function. hence, considering that a cmbp is also a polynomial function, the encoded prior knowledge and intuition in cmbp can be transferred to rpn by using the parameters of cmbp to initialize rpn. in this way, it combines rule-based models and data-driven models. a recurrent polynomial network is a computational network. the network contains multiple edges and loops. each node is either an input node, which is used to represent an input value, or a computation node. each node x is set an initial value u(0)x at time 0, and its value is updated at time 1, 2, · · · . both the type of edges and the type of nodes decide how the nodes’ values are updated. there are two types of edges. one type, referred to as type-1, indicates the value updating at time t takes the value of a node at time t − 1, i.e. type-1 edges are recurrent edges, while the other type, referred to as type-2, indicates the value updating at time t takes another node’s value at time t. except for loops made of “type-1” edges, the network should not have loops. for simplicity, let ix be the set of nodes y which are linked to node x by a type-1 edge, îx be the set of nodes y which are linked to node x by a type-2 edge. based on these definitions, two types of computation nodes, sum and product, are introduced. specifically, at time t > 0, if node x is a sum node, its value u(t)x is updated by u(t)x = ∑ y∈ix wx,yu (t−1) y + ∑ y∈îx ŵx,yu (t) y (6) where w, ŵ ∈ r are the weights of edges. similarly, if node x is a product node, its value is updated by u(t)x = ∏ y∈ix u(t−1)y mx,y ∏ y∈îx u(t)y m̂x,y (7) where mx,y and m̂x,y are integers, denoting the multiplicity of the type-1 edge −→yx, and the multiplicity of the type-2 edge −→yx respectively. it is noted that only w, ŵ are parameters of rpn while mx,y, m̂x,y are constant given the structure of an rpn. 𝑎 𝑏 𝑐 𝑑 1.0 0.5 figure 2: a simple example of rpn. the type of nodes a, b, c, d are input, input, product, and sum respectively. edge −→ dd is of type-1, while the other edges are of type-2. m̂a,c = 2, m̂b,c = m̂c,d = md,d = 1. let u(t), û(t) denote the vector of computation nodes’ values and the vector of input nodes’ values at time t respectively, then a well-defined rpn can be seen as a polynomial function as 70 recurrent polynomial network for dialogue state tracking below. u(t) = p ( û(t),u(t−1), 1 ) (8) where p is defined by equation (2). for example, for the rpn in figure 2, its corresponding polynomial function is (u(t)c , u (t) d ) =p ( u(t)a , u (t) b , u (t−1) c , u (t−1) d , 1 ) = ( (u(t)a ) 2 u (t) b , 0.5u (t−1) d + (u(t)a ) 2 u (t) b ) (9) each computation node can be regarded as an output node. for example, for the rpn in figure 2, node c and node d can be set as output nodes. 4. rpn for dialogue state tracking as introduced in section 1, in this paper, the dialogue state tracker receives slu n -best hypotheses for each user turn, each hypothesis having a set of act-slot-value tuples with a confidence score. the dialogue state tracker is supposed to output a set of distributions over the joint user goal, i.e., the value for each slot. for simplicity and consistency with the work of sun et al. (2014b) and yu et al. (2015), slot and value independence are assumed in the rpn model for dialogue state tracking2, though neither cmbp nor rpn is limited to the assumptions. in the rest of the paper, bt(v), p+ t (v), p−t (v), p̃+ t (v), p̃−t (v) are abbreviated by bt, p + t , p − t , p̃ + t , p̃ − t respectively in circumstances where there is no ambiguity. 4.1 structure before describing details of the structure used in the real situations, to help understand the corresponding relationship between rpn and cmbp, let’s first look at a simplified case with a smaller feature set and a smaller order, which is a corresponding relationship between the rpn shown in figure 3 and 2-order polynomial (10) with features bt−1, p+ t , 1: bt = 1− (1− bt−1)(1− p+ t ) = bt−1 + p+ t − p + t bt−1 (10) recall that a cmbp of polynomial order 2 with 3 features is the following equation (refer to equation (2)): p(ι0, ι1, ι2) = ∑ 0≤k1≤k2≤2 gk1,k2 ∏ 1≤i≤2 ιki (11) 2. for dstc-2/3 tasks, one slot can have at most one value, i.e. 0 ≤ ∑ v 6=none bt(v) ≤ 1. since value independence is assumed, to strictly maintain that relation, the belief is rescaled to ensure the sum of the belief of each value plus the belief of ‘none’ to be 1 when the belief is being output. actually, to enable rpn strictly maintain that relation, our original design of rpn had a “normalization” step when passing the belief from turn t to turn t+1. the “normalization” step will rescale the belief to make the sum of the belief of each value plus the belief of ‘none’ to be 1. our later experiment, however, demonstrated that there was no significant performance difference between the rpn with and without the “normalization” step. therefore, in practice, for simplicity, the “normalization” step can be omitted, value independence can be assumed, and the only thing needed is to rescale the belief when it is being output. 71 sun, xie and yu 𝑏𝑡−2 𝑃𝑡−1 + 1 𝑏𝑡−1 −10 1 0 1 0 𝑏𝑡−1 𝑃𝑡 + 1 𝑏𝑡 −10 1 0 1 0 𝑏𝑡 𝑃𝑡+1 + 1 𝑏𝑡+1 −10 1 0 1 0 figure 3: a simple example of rpn for dst. the rpn in figure 3 has three layers. the first layer only contains input nodes. the second layer only contains product nodes. the third layer only contains sum nodes. every product node in the second layer denotes a monomial of order 2 such as (bt−1) 2, bt−1, p + t and so on. every product node in the second layer is linked to the sum node in the third layer whose value is a weighted sum of value of product nodes. with weight set according to coefficients in equation (10), the value of sum node in the third layer is essentially the bt in equation (10). like the above simplified case, a layered rpn structure shown in figure 4 is used for dialogue state tracking in our first trial which essentially corresponds to 3-order cmbp, though the rpn framework is not limited to the layered topology. recall that a cmbp of polynomial order 3 is used as shown in the following equation (refer to equation (2)): p(ι0, · · · , ιd) = ∑ 0≤k1≤k2≤k3≤d gk1,k2,k3 ∏ 1≤i≤3 ιki (12) 𝑏𝑡−2 𝑃𝑡−1 + 𝑃𝑡−1 − 𝑃𝑡−1 + 𝑃𝑡−1 − 1 𝑏𝑡−1 𝑃𝑡 + 𝑃𝑡 − 𝑃𝑡 + 𝑃𝑡 − 1 𝑏𝑡 𝑃𝑡+1 + 𝑃𝑡+1 − 𝑃𝑡+1 + 𝑃𝑡+1 − 1 𝑏𝑡−1 𝑏𝑡 𝑏𝑡+1 𝑤, 𝑤 𝑤, 𝑤 𝑤, 𝑤 figure 4: rpn for dst. let (l, i) denote the index of i-th node in the l-th layer. the detailed definitions of each layer are as follows: • first layer / input layer: input nodes are features at turn t, which corresponds to variables in cmbp in section 2. i.e. 72 recurrent polynomial network for dialogue state tracking – u (t) (1,0) = bt−1 – u (t) (1,1) = p+ t – u (t) (1,2) = p−t – u (t) (1,3) = p̃+ t – u (t) (1,4) = p̃−t – u (t) (1,5) = 1 while 7 features are used in previous work of cmbp (sun et al., 2014b; yu et al., 2015), only 6 of them are used in rpn with feature brt−1 removed3. since our experiments showed the performance of cmbp would not become worse without feature brt−1, to make the structure more compact, brt−1 is not used in this paper for rpn. in accordance to this, cmbp mentioned in the rest of paper does not use this feature either. • second layer: the value of every product node in the second layer is a monomial like the simplified case. and every product node has indegree 3 which is corresponding to the order of cmbp. every monomial in cmbp is the product of three repeatable features. correspondingly, the value of every product node in second layer is the product of values of three repeatable nodes in the first layer. every triple (k1, k2, k3)(0 ≤ k1 ≤ k2 ≤ k3 ≤ 5) is enumerated to create a product node x = (2, i) in second layer that nodes (1, k1), (1, k2), (1, k3) are linked to. i.e. îx = {(1, k1), (1, k2), (1, k3)}. and thus u(t)x = u (t) 1,k1 u (t) 1,k2 u (t) 1,k3 . and different node in the second layer is created by a distinct triple. so given the 6 input features, there are 5∑ k1=0 5∑ k2=k1 5∑ k3=k2 1 = ( 6+3−1 3 ) = 56 nodes in the second layer. to simplify the notation, a bijection from nodes to monomials is defined as: f : {x|x is the index of a node in the 2nd layer} → {(k1, k2, k3)|0 ≤ k1 ≤ k2 ≤ k3 ≤ d} (13) f(x) = (k1, k2, k3)⇐⇒ u (t) 2,i = u (t) 1,k1 u (t) 1,k2 u (t) 1,k3 (14) where d + 1 = 6 is the number of nodes in the first layer, i.e. input feature dimension. • third layer: the value of sum node x = (3, 0) in the third layer is corresponding to the output value of cmbp. every product nodes in the second layer are linked to it. node x’s value u(t)3,0 is a weighted sum of values of product node u(t)2,i where the weights correspond to gk1,k2,k3 in equation (12). 3. brt is the belief of value being ‘none’, whose precise definition is given in section 2. 73 sun, xie and yu with only sum and product operation involved, every node’s value is essentially a polynomial of input features. and just like recurrent neural network, node at time t can be linked to node at time t+ 1. that is why this model is called recurrent polynomial network. the parameters of the rpn can be set according to cmbp coefficients gk1,k2,k3 in equation (12) so that the output value is the same as the value of cmbp, which is a direct way of applying prior knowledge and intuition to data-driven models. it is explained in detail in section 4.4. 4.2 activation function in dst, the output value is a belief which should lie in [0, 1], while values of computational nodes are not bound by certain interval in rpn. experiments showed that if weights are not properly set in rpn and a belief bt−1 output by rpn is larger than 1, then bt may grow much larger because bt is the weighted sum of monomials such as (bt−1) 3. belief of later turns such as bt+10 will tend to infinity. therefore, an activation function is needed to map bt to a legal belief value (referred to as b′t) in (0, 1). 3 kinds of functions, the logistic function, the clip function, and the softclip function have been considered. a logistic function is defined as logistic(x) = l 1 + e−η(x−x0) (15) it can map r to (0, 1) by setting l = 1. however, since basically the rpn designed for dialogue state tracking does similar operation as cmbp which is motivated from bayesian probability operation (yu et al., 2015), intuitively we expect the activation function to be linear on [0, 1] so that little distortion is added to the belief. as an alternation, a clip function is defined as clip(x) =  0 if x < 0 x if 0 ≤ x ≤ 1 1 if x > 1 (16) it is linear on [0, 1]. however, if b′t = clip(bt), bt 6∈ [0, 1] and l is the loss function, ∂l ∂bt = ∂l ∂b′t ∂b′t ∂bt = ∂l ∂b′t × 0 = 0 (17) thus, ∂l ∂bt would be 0 whatever ∂l ∂b′t is. this gradient vanishing phenomenon may affect the effectiveness of backpropagation training in section 4.5. so an activation function softclip(·) is introduced, which is a combination of logistic function and clip function. let ε denote a small value such as 0.01, δ denote the offset of sigmoid function such that sigmoid (ε− 0.5 + δ) = ε. here the sigmoid function refers to the special case of the logistic function defined by the formula sigmoid(x) = 1 1 + e−x (18) the softclip function is defined as softclip(x) ,  sigmoid (x− 0.5 + δ) if x ≤ ε x if ε < x < 1− ε sigmoid (x− 0.5− δ) if x ≥ 1− ε (19) 74 recurrent polynomial network for dialogue state tracking -1 -0.5 0 0.5 1 1.5 2 -1 0 1 2 -1 -0.5 0 0.5 1 1.5 2 -1 0 1 2 -1 -0.5 0 0.5 1 1.5 2 -1 0 1 2 (a) 𝑓(𝑥) = 𝑐𝑙𝑖𝑝(𝑥) (b) 𝑓(𝑥) = 1 1+𝑒−5(𝑥−0.5) (c) 𝑓 𝑥 = 𝑠𝑜𝑓𝑡𝑐𝑙𝑖𝑝(𝑥) 0 0.01 0.02 0.03 -0.15 0.01 0.03 (d) partial enlargement of 𝑠𝑜𝑓𝑡𝑐𝑙𝑖𝑝(⋅) figure 5: comparison among clip function, logistic function, and softclip function softclip : r→ (0, 1) is a non-decreasing, continuous function. however, it is not differentiable when x = ε or x = 1− ε. so we defined its derivative as follows: ∂softclip(x) ∂x ,  ∂sigmoid(x−0.5+δ) ∂x if x ≤ ε 1 if ε < x < 1− ε ∂sigmoid(x−0.5−δ) ∂x if x ≥ 1− ε (20) it is like a clip function. however, its derivative may be small on some inputs but is not zero. figure 5 shows the comparison among clip function, logistic function, and softclip function. according to the result of an experiment done on the dstc-2 dataset where rpns using different activation functions are trained dstc2trn and tested on dstc2dev, softclip function has demonstrated better performance than both clip and logistic function4, and is thus used in the rest of the paper. with the activation function, a new type of computation node, referred to as activation node, is introduced. activation node only takes one input and only has one input edge of type-2, i.e |îx| = 1 and ix = ∅. the value of an activation node x is calculated as u(t)x = softclip ( u(t)x ) (21) where x denotes the input node of node x. i.e. îx = {x}. the activation function is used in the rest of the paper. figure 6 gives an example of rpn with activation function, whose structure is constructed by adding an activation function to the rpn in figure 4. 4. the accuracy and l2 of rpns with clip, logistic, and softclip function are (0.779, 0.329), (0.789, 0.352), (0.790, 0.317) respectively. in particular, logistic functions with several different η are evaluated and the result reported here is the best one. 75 sun, xie and yu 𝑏𝑡−2 𝑃𝑡−1 + 𝑃𝑡−1 − 𝑃𝑡−1 + 𝑃𝑡−1 − 1 𝑏𝑡−1 𝑃𝑡 + 𝑃𝑡 − 𝑃𝑡 + 𝑃𝑡 − 1 𝑏𝑡 𝑃𝑡+1 + 𝑃𝑡+1 − 𝑃𝑡+1 + 𝑃𝑡+1 − 1 𝑏𝑡−1 𝑏𝑡 𝑏𝑡+1 figure 6: rpn for dst with activation functions 4.3 further exploration on structure adding features to cmbp is not easy because additional prior knowledge is needed to add to keep the search space not too large. concretely, adding features can introduce new monomials. since the trivial search space is exponentially increasing as the number of monomials, the search space tends to be too large to explore when new features are added. hence, to reduce the search space, additional prior knowledge is needed, which can introduce new constraints to the polynomial coefficients. for the same reason, increasing the model complexity also requires additional suitable prior knowledge to be added to limit the search space not too large in cmbp. in contrast to that, since rpn can be seen as a data-driven model, it is as easy as most data-driven statistical approaches such as rnn to add new features to rpn and use more complex structures. at the same time, no matter what new features are used and how complex the structure is, rpn can always take advantage prior knowledge and intuition which is discussed in section 4.4. in this paper, both new features and complex structure are explored. adding new features can be done by just adding input nodes which correspond to the new features, and then adding product nodes corresponding to the new possible monomials introduced by the new features. in this paper, for slot s, value v at turn t, in addition to f0 ∼ f5 which are defined as bt−1(v), p+ t (v), p−t (v), p̃+ t (v), p̃−t (v), and 1 respectively, 4 new features are investigated. f6 and f7 are features of system acts at the last turn: for slot s, value v at turn t, • f6 , canthelp(s, t, v) ∪ canthelp.missing slot value(s, t) =1 if the system cannot offer a venue with the constraint s = v or the value of slot s is not known for the selected venue, otherwise 0. • f7 , select(s, t, v) =1 if the system asks the user to pick a suggested value for slot s, otherwise 0. f6 and f7 are introduced because user is likely to change their goal if given machine acts canthelp(s, t, v), canthelp.missing slot value(s, t) and select(s, t, v). f8 and f9 are features of user acts at the current turn: for slot s, value v at turn t, • f8 , inform(s, t, v) = 1 if one of slu hypotheses from the user is informing slot s is v, otherwise 0. 76 recurrent polynomial network for dialogue state tracking • f9 , deny(s, t, v) =1 if one of slu hypotheses from the user is denying slot s is v, otherwise 0. f8 and f9 are features about slu act type, introduced to make system robust when the confidence scores of slu hypothesis are not reliable. in this paper, the complexity of evaluating and training rpn for dst would not increase sharply because a constant order 3 is used and number of product nodes in the second layer grows from 56 to 220 when number of features grows from 6 to 10. in addition to new features, rpn of more complex structure is also investigated in this paper. to capture some property just like belief bt of dialogue process, a new sum node x = (3, 1) in the third layer is introduced. the connection of (3, 1) is the same as (3, 0), so it introduces a new recurrent connection. the exact meaning of its value is unknown. however, it is the only value used to record information other than bt of previous turns. every other input features except bt are features of current turn t. compared with bt, there are fewer restrictions on the value of (3, 1) since its value is not directly supervised by the label. hence, introducing (3, 1) may help to reduce the effect of inaccurate labels. the structure of the rpn with 4 new features and 1 new sum node, together with new activation nodes introduced in section 4.2 is shown in figure 7. 𝑓1 (𝑡−1) 𝑏𝑡 𝑏𝑡+1 𝑓2 (𝑡−1) 𝑓8 (𝑡−1) 𝑓9 (𝑡−1) 𝑏𝑡−1 𝑓1 (𝑡) 𝑓2 (𝑡) 𝑓8 (𝑡) 𝑓9 (𝑡) 𝑓1 (𝑡+1) 𝑓2 (𝑡+1) 𝑓8 (𝑡+1) 𝑓9 (𝑡+1) figure 7: rpn with new features and more complex structure for dst 4.4 rpn initialization like most neural network models such as rnn, the initialization of rpn can be done by setting each weight, i.e. w and ŵ, to be a small random value. however, with its unique structure, the initialization can be much better by taking advantage of the relationship between cmbp and rpn which is introduced in section 4.1. when rpn is initialized according to a cmbp, prior knowledge and constraints are used to set rpn’s initial parameters as a suboptimum point in the whole parameter space. rpn as a datadriven model can fully utilize the advantages of data-driven approaches. rpn is better than real cmbp while they both use data samples to train parameters. in the work of yu et al. (2015), realcoefficient cmbp uses hill climbing to adjust parameters that are initially not zero and the change of parameters are always a multiple of 0.1. rpn can adjust all parameters including parameters initialized as 0 concurrently, while the complexity of adjusting all parameters concurrently is nearly 77 sun, xie and yu the same as adjusting one parameter in cmbp. besides, the change of parameters can be large or small, depending on learning rate. thus, rpn and cmbp both are combining rule-based models and data-driven ones, while rpn is a data-driven model utilizing rule advantages and cmbp is a rule model utilizing data-driven advantages. in fact, given a cmbp, an rpn can achieve the same performance as the cmbp just by setting its weights according to the coefficients of the cmbp. to illustrate that, the steps of initializing the rpn in figure 7 with a cmbp of features f0 ∼ f9 is described below. first, to ensure that the new added sum node x = (3, 1) will not influence the output bt in rpn with initial parameters, ŵx,y is set to 0 for all y. so node x’s value u(t)x is always 0. next, considering the rpn in figure 7 has more features than cmbp does, the weights related the new features should be set to 0. specifically, suppose node x is the sum node in the third layer in rpn denoting bt before activation and node y is one of the product nodes in the second layer denoting a monomial, if product node y is products of features f6, f7, f8, f9 or the added sum node, then node y’s value is not a monomial in cmbp, then weights ŵx,y should be set to 0. finally, if product node y is the product of features f0 ∼ f5, suppose the order of cmbp is 3, then f(y) = (k1, k2, k3) defined in equation (13) should satisfy 0 ≤ k1 ≤ k2 ≤ k3 ≤ 5. weights ŵxy should be initialized as gk1,k2,k3 which is the coefficient of fk1fk2fk3 in cmbp. thus, wx,y = { gk1,k2,k3 if x = (2, 0) and f(x) = (k1, k2, k3) 0 otherwise (22) for rpn of other structures, the initialization can be done by following similar steps. experiments show that after training, there are only a few weights larger than 0.1, no matter using cmbp or random initialization. 4.5 training rpn suppose t (d) is the number of turns in dialogue d, h(d, s, t) is the set of values corresponding to slot s appearing in slu hypothesis in turn t in dialogue d, bd,s,t(v) is the output belief of value v in dialogue d and ld,s,t(v) is the indicator of goal s = v being part of joint goal at turn t in the label of dialogue d. the cost function is defined as l = 1∑ d t (d) ∑ d t (d)−1∑ t=0 ∑ s ∑ v∈∪ti=0h(d,s,i) (bd,s,t(v)− ld,s,t(v))2 (23) training process of a mini-batch can be divided into two parts: forward pass and backward pass. forward pass for each training sample, every node’s value at every time is evaluated first. when evaluating u(t)x , values of nodes in ix and îx should be evaluated before. the computation formula should be based on the type of node x. in particular, for a layered rpn structure, we can simply evaluate ut1x1 earlier than ut2x2 if t1 < t2 or t1 = t2 and x1’s layer number is smaller than x2’s. backward pass backpropagation through time (bptt) is used in training rpn. let error of node x at time t be δ(t)x = ∂l ∂u (t) x . if a node x is an output node, then δ(t)x should be initialized according to its label lt and output value u(t)x , otherwise δ(t)x should be initialized to 0. after a node’s error 78 recurrent polynomial network for dialogue state tracking foreach mini batch m do initialize ∆wxy = 0,∆ŵx,y = 0 for every x, y initialize the value of recurrent node at turn 0 as 0 foreach training dialogue d, slot s, value v in mini batch m do t ← the number of turns of current training sample /* forward pass */ for t← 1 to t do for d← 1 to 4 do foreach node x in time t, layer d do evaluate u(t)x if x is output node then δ (t) x ← 2(u (t) x − lt) else δ (t) x ← 0 /* backward pass */ for t← t to 1 do for d← 4 to 1 do foreach node x in time t, layer d do foreach node y ∈ îx do /* node y is linked to node x by a "type-2" edge */ δ (t) y ← δ (t) y + δ (t) x ∂u (t) x ∂u (t) y foreach node y ∈ ix do /* node y is linked to node x by a "type-1" edge */ δ (t−1) y ← δ (t−1) y + δ (t) x ∂u (t) x ∂u (t−1) y for t← 1 to t do for d← 1 to 4 do foreach sum node x in time t, layer d do foreach node y ∈ îx do ∆ŵxy ← ∆ŵxy + αδ (t) x u (t) y foreach node y ∈ ix do ∆wxy ← ∆wxy + αδ (t) x u (t−1) y foreach edge(x, y) do wxy ← wxy −∆wxy ŵxy ← ŵxy −∆ŵxy algorithm 1: training algorithm of rpn for dst 79 sun, xie and yu δ (t) x is determined, it can be passed to δ(t−1)y (y ∈ ix) and δ(t)y (y ∈ îx). error passing should follow the reversed edge’s direction. so the order of nodes passing error can follow the reverse order of evaluating nodes’ values. when every δ(t)x has been evaluated, the increment on weight ŵxy can be calculated by ∆ŵxy = α ∂l ∂ŵxy = α t∑ i=1 ∂l ∂u (t) x ∂u (t) x ∂ŵxy = α t∑ i=1 δ(t)x u(t)y (24) where α is the learning rate. ∆wxy can be evaluated similarly. note that only wxy and ŵxy are parameters of rpn. the complete formula of evaluating node value u(t)x and passing error δ(t)x can be found in appendix. in this paper, mini-batch is used in training rpn for dst with batch size 8. in each training epoch, ∆wxy and ∆ŵxy are calculated for every training sample and added together. the weight wxy and ŵxy is updated by wxy = wxy −∆wxy (25) ŵxy = ŵxy −∆ŵxy (26) the pseudocode of training is shown in algorithm 1. 5. experiment as introduced in section 1, in this paper, dstc-2 and dstc-3 tasks are used to evaluate the proposed approach. both tasks provide training dialogues with turn-level asr hypotheses, slu hypotheses and user goal labels. the dstc-2 task provides 2118 training dialogues in restaurants domain (henderson et al., 2014b), while in dstc-3, only 10 in-domain training dialogues in tourists domain are provided, because the dstc-3 task is to adapt the tracker trained on dstc-2 data to the new domain with very few dialogues (henderson et al., 2014a). table 1 summarizes the size of datasets of dstc-2 and dstc-3. task dataset #dialogues usage dstc2trn 1612 training dstc-2 dstc2dev 506 training dstc2eval 1117 test dstc-3 dstc3seed 10 not used dstc3eval 2265 test table 1: summary of data corpora of dstc-2/3 the dst evaluation criteria are the joint goal accuracy and the l2 (henderson et al., 2014b,a). accuracy is defined as the fraction of turns in which the tracker’s 1-best joint goal hypothesis is 80 recurrent polynomial network for dialogue state tracking correct, the larger the better. l2 is the l2 norm between the distribution of all hypotheses output by the tracker and the correct goal distribution (a delta function), the smaller the better. besides, schedule 2 and labelling scheme a defined in (henderson et al., 2013) are used in both tasks. specifically, schedule 2 only counts the turns where new information about some slots either in a system confirmation action or in the slu list is observed. labelling scheme a is that the labelled state is accumulated forwards through the whole dialogue. for example, the goal for slot s is “none” until it is informed as s = v by the user, from then on, it is labelled as v until it is again informed otherwise. it has been shown that the organiser-provided live slu confidence was, at best, a noisy signal (zhu et al., 2014; sun et al., 2014a). hence, most of the state-of-the-art results from dstc-2 and dstc-3 used refined slu (either explicitly rebuild a slu component or take the asr hypotheses into the trackers (williams, 2014; sun et al., 2014a; henderson et al., 2014d,c; kadlec et al., 2014; sun et al., 2014b)). in accordance to this, except for the results directly taken from other papers (shown in table 5 and 6), all experiments in this paper used the output from a refined semantic parser (zhu et al., 2014; sun et al., 2014a) instead of the live slu provided by the organizer. for all experiments, mse is used as the training criterion. mse is chosen because of two reasons: (i) mse can directly reflect l2 performance which is one of the main metrics in dstcs. (ii) experiment has shown that some other criterion such as cross-entropy loss cannot lead to better performance. for both dstc-2 and dstc-3 tasks, dstc2trn and dstc2dev are used, 60% of the data is used for training and 40% for validation, unless otherwise stated. validation is performed every 5 epochs. learning rate is set to 0.6 initially. during the training, learning rate is halved each time the performance does not increases. training is stopped when the learning rate is sufficiently small, or the maximum number of training epochs is reached. here, the maximum number of training epochs is set to 40. l2 regularization is used for all the experiments. the parameter of l2 regularization is set to be the one leading to the best performance on dstc2dev when trained on dstc2trn. l1 regularization is not used since it cannot yield better performance. 5.1 investigation on rpn configurations this section describes the experiments comparing different configurations of rpn. all experiments were performed on both the dstc-2 and dstc-3 tasks. as indicated in section 4.4, an rpn can be initialized by a cmbp. table 2 shows the performance comparison between initialization with a cmbp and with random values. in this experiment, the structure shown in figure 6 is used. the performance of initialization with random values reported here is the average performance of 10 different random seeds, and their standard deviations are given in parentheses. the random scheme used is the one with the best performance on dstc2dev when trained on dstc2trn among various kinds of random initialization schemes. the performance of the rpn initialized by random values is compared with the performance of the rpn initialized by the integer-coefficient cmbp. here, the cmbp has 12 non-zero coefficients and has the best performance on dstc2dev when trained on dstc2trn. it can be seen from table 2 that the rpn initialized by the cmbp coefficients outperforms the rpn initialized by random values moderately on dstc2eval and significantly on dstc3eval (p-value< 0.05). this demonstrates the encoded prior knowledge and intuition in cmbp can be transferred to rpn to improve rpn’s performance, which is one of rpn’s advantage, combining rule-based models and 81 sun, xie and yu initialization dstc2eval dstc3eval acc l2 acc l2 random 0.753 (0.008) 0.468 (0.020) 0.633 (0.005) 0.667 (0.026) cmbp 0.756 0.373 0.648 0.553 table 2: performance comparison between the rpn initialized by random values, and the rpn initialized by the cmbp coefficients on dstc2eval and dstc3eval. data-driven models. in the rest of the experiments, all rpns use cmbp coefficients for initialization. since section 4.3 shows that it is convenient to add features and try more complex structures, it is interesting to investigate rpns with different feature sets and structures, as shown in table 3. it can be seen that while no obvious correlation between the performance and different configurations of feature sets and structures can be observed on dstc2eval and dstc3eval, rpns with new features and new recurrent connections have achieved slightly better performance on accuracy, though the performance difference is not significant. in the rest of the paper, both new features and new recurrent connections are used in rpn, unless otherwise stated. feature set new recurrent dstc2eval dstc3eval connections acc l2 acc l2 f0 ∼ f5 no 0.756 0.373 0.648 0.553 f0 ∼ f9 0.757 0.374 0.650 0.557 f0 ∼ f5 yes 0.756 0.373 0.648 0.553 f0 ∼ f9 0.757 0.374 0.650 0.549 table 3: performance comparison among rpns with different configurations on dstc2eval and dstc3eval. 5.2 comparison with other dst approaches the previous subsection investigates how to get the rpn with the best configuration. in this subsection, the performance of rpn is compared to both rule-based and data-driven statistical approaches. to make fair comparison, all data-driven models together with rpn in this subsection use similar feature set. altogether, 2 rule-based trackers and 3 data-driven trackers were built for performance comparison. • maxconf is a rule-based model commonly used in spoken dialogue systems which always selects the value with the highest confidence score from the 1st turn to the current turn. it was used as one of the primary baselines in dstc-2 and dstc-3. • hwu is a rule-based model proposed by wang and lemon (2013). it is regarded as a simple, yet competitive baseline of dstc-2 and dstc-3. 82 recurrent polynomial network for dialogue state tracking type system dstc2eval dstc3eval acc l2 acc l2 rule maxconf 0.668 0.647 0.548 0.861 hwu 0.720 0.445 0.594 0.570 data-driven dnn 0.719 0.469 0.628 0.556 maxent 0.710 0.431 0.607 0.563 lstm 0.736 0.418 0.632 0.549 hybrid cmbp 0.755 0.372 0.627 0.546 rpn 0.756 0.373 0.648 0.553 table 4: performance comparison among rpn, rule-based and data-driven approaches with similar feature set on dstc2eval and dstc3eval. the performance of cmbp in the table is the performance of the rpn which has been initialized but not been trained. • dnn is a data-driven statistical model using deep neural network model (sun et al., 2014a) with probability feature as rpn. since dnn does not have recurrent structures while rpn does, to fairly take into account this, the dnn feature set at the tth turn is defined as⋃ i∈{t−9,··· ,t} { p+ i (v), p−i (v), p̃+ i (v), p̃−i (v) } ∪ { p̂ (t) } where p̂ (t) is the highest confidence score from the 1st turn to the tth turn. the dnn has 3 hidden layers with 64 nodes per layer. • maxent is also a data-driven statistical model using maximum entropy model (sun et al., 2014a) with the same input features as dnn. • lstm is another data-driven statistical model using long short-term memory model (hochreiter and schmidhuber, 1997) with the same input features as rpn. it has similar structure as the dnn model (sun et al., 2014a) except for its hidden layers using lstm blocks. the lstm used here has 3 hidden layers with 100 lstm blocks per layer. it can be observed that, with similar feature set, rpn can outperform both rule-based and datadriven approaches in terms of joint goal accuracy. statistical significance tests were also performed assuming a binomial distribution for each turn. rpn was shown to significantly outperform both rule-based and data-driven statistical approaches at 95% confidence level. for l2, rpn is competitive to both rule-based and the data-driven statistical approaches. 5.3 comparison with state-of-the-art dstc trackers in the dstcs, the state-of-the-art trackers mostly employed data-driven statistical approaches. usually, richer feature set and more complicated model structures than the data-driven statistical models in section 5.2 are used. in this section, the proposed rpn approach is compared to the best submitted trackers in dstc-2/3 and the best cmbp trackers, regardless of fairness of feature selection and the slu refinement approach. rpn is compared and the results are shown in table 5 and table 83 sun, xie and yu system approach rank acc l2 baseline* rule 5 0.719 0.464 williams (2014) lambdamart 1 0.784 0.735 henderson et al. (2014d) rnn 2 0.768 0.346 sun et al. (2014a) dnn 3 0.750 0.416 yu et al. (2015) real cmbp 2.5 0.762 0.436 rpn rpn 2.5 0.757 0.374 table 5: performance comparison among rpn, real-coefficient cmbp and best trackers of dstc2 on dstc2eval. baseline* is the best results from the 4 baselines in dstc2. 6. note that structure shown in figure 7 with richer feature set and a new recurrent connection is used here. note that, in dstc-2, the williams (2014)’s system employed batch asr hypothesis information (i.e. off-line asr re-decoded results) and cannot be used in the normal on-line model in practice. hence, the practically best tracker is henderson et al. (2014d). it can be observed from table 5, rpn ranks only second to the best practical tracker among the submitted trackers in dstc-2 in accuracy and l2. considering that rpn only used probabilistic features and very limited added features and can operate very efficiently, it is quite competitive. system approach rank acc l2 baseline* rule 6 0.575 0.691 henderson et al. (2014c) rnn 1 0.646 0.538 kadlec et al. (2014) rule 2 0.630 0.627 sun et al. (2014b) int cmbp 3 0.610 0.556 yu et al. (2015) real cmbp 1.5 0.634 0.579 rpn rpn 0.5 0.650 0.549 table 6: performance comparison among rpn, real-coefficient cmbp and best trackers of dstc3 on dstc3eval. baseline* is the best results from the 4 baselines in dstc3. it can be seen from table 6, rpn trained on dstc-2 can achieve state-of-the-art performance on dstc-3 without modifying tracking method5, outperforming all the submitted trackers in dstc3 including the rnn system. this demonstrates that rpn successfully inherits the advantage of good generalization ability of rule-based model. considering the feature set and structure of rpn are relatively simple in this paper, future work will investigate richer features and more complex structures. 6. conclusion this paper proposes a novel hybrid framework, referred to as recurrent polynomial network, to improve the combination of the rule-based approaches and data-driven approaches. with the ability 5. the parser is refined for dstc-3 (zhu et al., 2014). 84 recurrent polynomial network for dialogue state tracking of incorporating prior knowledge into a data-driven framework, rpn has the advantages of both rule-based and data-driven approaches. experiments on two dstc tasks showed that the proposed approach not only is more stable than many major data-driven statistical approaches, but also has competitive performance, outperforming many state-of-the-art trackers. since the rpn in this paper only used probabilistic features and very limited added features, the performance of rpn can be influenced by how reliable the slu’s confidence scores are. therefore, future work will investigate the influence of slus on the performance of rpn, and rich features for rpn. moreover, future work will also address applying rpn to other domains, such as the bus timetables domain in dstc-1, and theoretic analysis of rpn. references ian cloete and jacek m zurada. knowledge-based neurocomputing. mit press, 2000. li deng, jinyu li, jui-ting huang, kaisheng yao, dong yu, frank seide, michael seltzer, geoff zweig, xiaodong he, jason williams, et al. recent advances in deep learning for speech research at microsoft. in acoustics, speech and signal processing (icassp), 2013 ieee international conference on, pages 8604–8608. ieee, 2013. milica gašić and steve young. effective handling of dialogue state in the hidden information state pomdp-based dialogue manager. acm transactions on speech and language processing (tslp), 7(3):4, 2011. matthew henderson, blaise thomson, and jason williams. dialog state tracking challenge 2 & 3. 2013. matthew henderson, blaise thomson, and jason d. williams. the third dialog state tracking challenge. in proceedings of ieee spoken language technology workshop (slt), december 2014a. matthew henderson, blaise thomson, and jason d williams. the second dialog state tracking challenge. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 263–272, philadelphia, pa, u.s.a., june 2014b. association for computational linguistics. url http://www.aclweb.org/anthology/w14-4337. matthew henderson, blaise thomson, and steve young. robust dialog state tracking using delexicalised recurrent neural networks and unsupervised adaptation. in proceedings of ieee spoken language technology workshop (slt), december 2014c. matthew henderson, blaise thomson, and steve young. word-based dialog state tracking with recurrent neural networks. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 292–299, philadelphia, pa, u.s.a., june 2014d. association for computational linguistics. url http://www.aclweb.org/ anthology/w14-4340. sepp hochreiter and jürgen schmidhuber. long short-term memory. neural computation, 9(8): 1735–1780, 1997. 85 sun, xie and yu rudolf kadlec, miroslav vodoln, jindrich libovick, jan macek, and jan kleindienst. knowledgebased dialog state tracking. in proceedings 2014 ieee spoken language technology workshop, south lake tahoe, usa, december 2014. sungjin lee. structured discriminative model for dialog state tracking. in proceedings of the sigdial 2013 conference, pages 442–451, metz, france, august 2013. association for computational linguistics. url http://www.aclweb.org/anthology/w/w13/w13-4069. sungjin lee and maxine eskenazi. recipe for building robust spoken dialog state trackers: dialog state tracking challenge system description. in proceedings of the sigdial 2013 conference, pages 414–422, metz, france, august 2013. association for computational linguistics. url http://www.aclweb.org/anthology/w/w13/w13-4066. frank rosenblatt. the perceptron: a probabilistic model for information storage and organization in the brain. psychological review, 65(6):386, 1958. nicholas roy, joelle pineau, and sebastian thrun. spoken dialogue management using probabilistic reasoning. in proceedings of the 38th annual meeting on association for computational linguistics, pages 93–100. association for computational linguistics, 2000. kai sun, lu chen, su zhu, and kai yu. the sjtu system for dialog state tracking challenge 2. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 318–326, philadelphia, pa, u.s.a., june 2014a. association for computational linguistics. url http://www.aclweb.org/anthology/w14-4343. kai sun, lu chen, su zhu, and kai yu. a generalized rule based tracker for dialogue state tracking. in proceedings of ieee spoken language technology workshop (slt), december 2014b. blaise thomson and steve young. bayesian update of dialogue state: a pomdp framework for spoken dialogue systems. computer speech & language, 24(4):562–588, 2010. zhuoran wang and oliver lemon. a simple and generic belief tracking mechanism for the dialog state tracking challenge: on the believability of observed information. in proceedings of the sigdial 2013 conference, pages 423–432, metz, france, august 2013. association for computational linguistics. url http://www.aclweb.org/anthology/w/w13/w13-4067. jason williams, antoine raux, deepak ramachandran, and alan black. the dialog state tracking challenge. in proceedings of the sigdial 2013 conference, pages 404–413, metz, france, august 2013. association for computational linguistics. url http://www.aclweb.org/ anthology/w/w13/w13-4065. jason d williams. challenges and opportunities for state tracking in statistical spoken dialog systems: results from two public deployments. selected topics in signal processing, ieee journal of, 6(8):959–970, 2012. jason d williams. web-style ranking and slu combination for dialog state tracking. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 282–291, philadelphia, pa, u.s.a., june 2014. association for computational linguistics. url http://www.aclweb.org/anthology/w14-4339. 86 recurrent polynomial network for dialogue state tracking jason d williams and steve young. scaling up pomdps for dialog management: the“summary pomdp”method. in automatic speech recognition and understanding, 2005 ieee workshop on, pages 177–182. ieee, 2005. jason d williams and steve young. partially observable markov decision processes for spoken dialog systems. computer speech & language, 21(2):393–422, 2007. steve young, milica gašić, simon keizer, françois mairesse, jost schatzmann, blaise thomson, and kai yu. the hidden information state model: a practical framework for pomdp-based spoken dialogue management. computer speech & language, 24(2):150–174, 2010. kai yu, kai sun, lu chen, and su zhu. constrained markov bayesian polynomial for efficient dialogue state tracking. ieee/acm transactions on audio, speech, and language processing, 23(12):2177–2188, december 2015. bo zhang, qingsheng cai, jianfeng mao, eric chang, and baining guo. spoken dialogue management as planning and acting under uncertainty. in interspeech, pages 2169–2172, 2001. su zhu, lu chen, kai sun, da zheng, and kai yu. semantic parser enhancement for dialogue domain extension with little data. in proceedings of ieee spoken language technology workshop (slt), december 2014. lukas zilka, david marek, matej korvas, and filip jurcicek. comparison of bayesian discriminative and generative models for dialogue state tracking. in proceedings of the sigdial 2013 conference, pages 452–456, metz, france, august 2013. association for computational linguistics. url http://www.aclweb.org/anthology/w/w13/w13-4070. appendix derivative calculation using mse as the criterion, δ(t)x = ∂l ∂u (t) x is initialized as following6: δ(t)x = { 2(u (t) x − ld,s,t(v)) if x is a output node calculating bd,s,t(v) 0 otherwise (27) suppose node x is an activation node and f(·) = softclip(·), let y = x, δ(t)y = ∂l ∂u (t) y = ∂l ∂u (t) x ∂u (t) x ∂u (t) y = δ(t)x ∂f(u (t) y ) ∂u (t) y (28) 6. the symbols used in this section such as î , i , m̂ , m follow the definitions in section 3. 87 sun, xie and yu suppose node x = (d, i) is a sum node, then when node x passes its error, the error of node y ∈ îx is updated as δ(t)y = δ(t)y + ∂l ∂u (t) x ∂u (t) x ∂u (t) y = δ(t)y + δ(t)x ŵx,y (29) similarly, error of node y ∈ ix is updated as δ(t)y = δ(t)y + ∂l ∂u (t) x ∂u (t) x ∂u (t−1) y = δ(t)y + δ(t)x wx,y (30) suppose node x = (d, i) is a product node, then when node x passes its error, error of node y ∈ îx is updated as δ(t)y = δ(t)y + ∂l ∂u (t) x ∂u (t) x ∂u (t) y = δ(t)y + δ(t)x m̂x,yu (t) y m̂x,y−1 ∏ z∈îx−{y} u(t)z m̂x,z ∏ z∈ix u(t−1)z mx,z (31) similarly, error of node y ∈ ix is updated as δ(t)y = δ(t)y + ∂l ∂u (t) x ∂u (t) x ∂u (t−1) y = δ(t)y + δ(t)x mx,yu (t−1) y mx,y−1 ∏ z∈îx u(t)z m̂x,z ∏ z∈ix−{y} u(t−1)z mx,z (32) 88 3727-7564-3-ed dialogue & discourse 8(2) 84-104 doi: 10.5087/dad.2017.204 ©2017 mindaugas mozuraitis, daphna heller this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). discourse coherence and the interpretation of accented pronouns mindaugas mozuraitis mozuraitismindaugas@gmail.com cancer care ontario canada daphna heller daphna.heller@utoronto.ca department of linguistics university of toronto canada editor: massimo poesio submitted 11/2016; accepted 08/2017; published online 10/2017 abstract it has been assumed since akmajian and jackendoff (1970) that accenting (or stressing) a pronoun, namely making it prosodically prominent, changes its interpretation. however, more recent work suggests that not all accented pronoun receive an interpretation that is different from that of their unaccented counterpart. one proposal is that the alternative is blocked if it does not result in a plausible sentence (taylor et al., 2013). in a series of three experiments that use an offline comprehension task, we show, first, that plausibility cannot account for the lack of reversal in all cases. we also rule out two other hypothesis about why the alternative interpretation is blocked: when it competes with a strong default, and when it is lower in salience than the default. instead, we conclude that coherence relations indirectly constrain the interpretation of accented pronouns. specifically, the difference between parallel and result coherence is that the former, but not the latter, explicitly provides the alternative to the accented event which drives reversal. this means that reversal is only achieved under special circumstances. like other accented constituents, accented pronouns simply involve alternatives for their interpretation (krifka, 2008). our findings indicate that the most readily available alternative is the negation of the accented event; this indirectly gives rise to the effect of ‘surprisal’ previously associated with accented pronouns. keywords: pronouns, discourse coherence, anaphora, accenting, prosody 1 introduction pronouns (e.g., she, he, it) have received much attention in the (psycho)linguistics literature. this is because the meaning they encode is minimal (e.g., gender, number, animacy), and their interpretation depends in large part on the integration of information from the discourse context (i.e., the preceding linguistic information), and potentially other sources of information. for example, the pronoun he in (1b) requires a referent that is male, singular, and a person. (1) a. william called oliver in the morning. b. then, he went to school. linguistically, the pronoun he in (1b) is ambiguous: it can be interpreted as either of the entities mentioned in prior discourse (i.e., william and oliver in 1a). however, there is a strong intuitive preference to interpret the pronoun as referring to william; we will call this the default interpretation throughout the paper. pronoun interpretation has been accounted for using the discourse coherence and accented pronouns 85 concept of accessibility (or salience). specifically, it has been widely assumed that, during the processing of discourse, listeners create a hierarchy of entities (ariel, 1990; gundel, hedberg, & zacharski, 1993; grosz, joshi, & weinstein, 1995; arnold, 2010, and many others). this hierarchy has been argued to depend on how entities were previously mentioned in the discourse, including their syntactic position, their thematic role, and also how recently they have been mentioned (arnold, 2010). in this approach, the default interpretation of a pronoun would be the most accessible (or salient) entity in the discourse context. for example, processing (1a) would result in a hierarchy of two entities, where william is ranked higher than oliver; this is because william was the subject of the first sentence. next, when processing (1b), he will be interpreted as the higher-ranking entity in this hierarchy, namely william. it is well known, however, that the notion of accessibility alone cannot account for the default interpretation of pronouns in all cases. for example, after processing (2a), william will rank higher than oliver on an accessibility scale, just like after (1a), and yet the default interpretation of him in (2b) is oliver. similarly, (3a) would render bill more accessible than john, but the default interpretation of he in (3b) is actually john. (2) a. william hit oliver. b. then, rob slapped him. (adapted from smyth, 1994) (3) a. bill passed the comic to john. b. then, he read it. (adapted from stevenson, knott, oberlander, & mcdonald, 2000) these kinds of examples have led to the suggestion that pronoun interpretation should not be viewed as a process that operates on referents and their accessibility status alone; instead it should be seen as a by-product of a general mechanism of establishing coherence in discourse (hobbs, 1978, 1979; hobbs, stickel, appelt, & martin, 1993; stevenson, crawley, & kleinman, 1994; stevenson et al., 2000; kehler, 2002; de hoop, 2004; wolf, gibson, & desmet, 2004; jasinskaja, kölsch, & mayer, 2007; kehler, kertz, rohde, & elman, 2008, among others). the idea is that the referent assigned to a pronoun is a result of a more general process of reasoning about how the meaning of the current sentence can be integrated with prior discourse. in (2), if him is interpreted as referring back to the previous object, this will enhance the similarity of the event in the second sentence to the event in the first sentence (cf. hobbs, 1979; smyth, 1994; chambers, & smyth, 1998; kehler, 2002); this similarity relation between sentences has been described as a discourse relation of “parallel” coherence (e.g., kehler, 2002). for (3), what connects the second sentence (i.e., reading a comic book) with the first sentence (i.e., giving a comic book) is world knowledge that people normally read a comic book after being given one rather than after giving one (cf. hobbs, 1979; stevenson et al., 2000; kehler, 2002). this relation between sentences has been described as the discourse coherence of “result”, because of the causal relationship between the two events (e.g., stevenson et al., 2000; kehler, 2002). the focus of the current paper is the interpretation of accented pronouns (also called stressed pronouns), namely pronouns that are acoustically more prominent. it has been widely accepted since akmajian and jackendoff (1970) that an accented pronoun receives a different interpretation than its unaccented counterpart (i.e., an interpretation different from the default interpretation). for example, in (2) above, the unaccented pronoun him is intuitively interpreted as oliver (the default interpretation), while the accented counterpart him is intuitively interpreted as william (a non-default interpretation, see also ariel, 1990; hirschberg & ward, 1991; kameyama 1994, 1999; cahn, 1995; nakatani, 1997). this pattern of “reversal” has been accounted for with different mechanisms in the literature. for example, kameyama’s (1999) complementary preference hypothesis operates on a set of ‘currently salient’ entities. in this system, the interpretation of an accented pronoun is achieved by first computing the referent for the unaccented pronoun, and then substituting it with a different entity from the set of ‘currently salient’ entities. but reversal has also been derived within coherence-based accounts, such as mozuraitis and heller 86 kehler (2005): here accenting does not target entities but rather events, indicating that events unfold contrary to expectations, which is achieved by assigning the accented pronoun a referent that is different from the default interpretation. thus, despite relying on different mechanisms, the two accounts will both predict reversal in many cases that accented pronouns will receive a different interpretation than their unaccented counterparts, namely reversal. however, not all accented pronouns actually exhibit reversal. for example, de hoop (2004) presents a set of naturally-occurring examples from written texts, where the accented pronouns receive the same interpretation as their unaccented counterparts (see also german, 2009 for similar results in controlled experiments). two of her examples are given in (4): (4) a. jack and mary are good friends. he is from louisiana. b. ‘i left oxhead in the road outside. will you see him safe for me? take a guardor four.’ ‘yes, alexander.’ he went off in a blaze of gratitude. there was a felt silence; antipatros was looking oddly under his brows. ‘alexander. the queen your mother is in the theatre. had she not better have a guard?’ in (4a), he is intuitively interpreted as jack and in (4b) she intuitively interpreted as the queen; both are the same interpretation that would be assigned to an unaccented pronoun in the same position, or the default interpretation. note that in both contexts there is only one referent that fits the gender features of the pronoun; that is, an alternative to the default interpretation is simply not available. de hoop suggests that accenting is licensed if it can signal contrast, but this contrast need not come from a different interpretation for the pronoun itself: in (4a) jack’s being from louisiana is contrasted with mary not being from louisiana, and in (4b) the queen having a guard is contrasted with her not having one1. we come back to this idea in the general discussion, but in the meantime, we may wish to draw a new generalization: the default interpretation is assigned to an accented pronoun if no alternative is available. this idea is pursued by taylor, stowe, redeker, and hoeks (2013). in an offline judgement experiment, they examine discourses like (5a) where the coherence relation is one of parallelism, and discourses like (5b) where the coherence relation is causal, that is result. (5) a. sandra called monica, and roger emailed her/her. b. michelle trained beth, and anne paid her/her. taylor et al., find that in parallel discourses the accented pronoun received the non-default interpretation (i.e., a different interpretation from the unaccented pronoun), but in the result discourses there was no such reversal (cf. tavano & kaiser, 2008). in a separate experiment, they find that the sentence with the non-default referent was significantly more plausible in the parallel discourses (i.e., where the two sentences describe similar events) than in the result discourses (i.e., where the second event can be taken to be caused by the first). linking the two findings, they claim that accenting a pronoun leads to a change in its interpretation only if the alternative interpretation is plausible. this proposal fits well with frameworks that subsume pronoun interpretation under the drive to make the discourse coherent (e.g., hobbs, 1979; kehler, 2005). however, because in taylor et al.’s study coherence and plausibility were confounded, we cannot rule out the possibility that it is coherence and not plausibility that is responsible for the presence or absence of reversal. the goal of the current study is to identify the reason for why some accented pronouns receive an alternative interpretation (i.e., showing reversal), while others do not. we start where 1 we do not go into de hoop’s analysis in more detail because it cannot account for the default interpretation of unaccented object pronouns in parallel discourses (such as [2] above) which are at the heart of our paper. this is because her analysis follows centering theory (grozs et al., 1995) in relying on the notion of “continued topic” for pronoun interpretation, but this notion makes the wrong prediction for the parallel discourses; see chambers and smyth (1998) for evidence and discussion. discourse coherence and accented pronouns 87 taylor et al. (2013) left off, trying to dissociate discourse coherence and plausibility in experiment 1; to this end, we use discourses that are matched on the plausibility of the two possible interpretation, but differ in coherence (parallel vs. result). since we find reversal in parallel discourses but not in result discourses, we conduct two additional experiments that aim to explore other generalizations for the discourse situations under which reversal takes place. experiment 2 considers the strength of the preference for the default interpretation, and experiment 3 considers the syntactic position of antecedents. like de hoop (2004) and taylor et al. (2013), we assume that an accented pronoun will receive a different interpretation if one is available, and explore the circumstances under which the alternative interpretation may not be available. 2 experiment 1 the goal of experiment 1 was to dissociate effects of discourse coherence from effects of plausibility. we created discourses where the final sentence contained an ambiguous (object) pronoun: consider the target sentence in table 1. to manipulate discourse coherence, we changed the preceding sentence: see table 1 (we will continue to call it the preceding sentence for short). to create parallel coherence, we used a preceding sentence that described an event similar to the event in the target sentence (e.g., mailed a souvenir and sent a postcard). as mentioned previously, a number of studies have shown that in parallel discourses the default interpretation of an object pronoun is the previous object (e.g., smyth, 1994; chambers, & smyth, 1998). to create result coherence, we used a preceding sentence that described an event that could be taken as the cause of the event in the target sentence (e.g., hobbs, 1979; kehler, 2002). in these cases, the default interpretation of an unaccented object pronoun would be the previous subject (e.g., wolf, gibson, & desmet, 2004; kehler et al., 2008). introduction sentence the animals went on a school exchange across the globe. preceding sentence (parallel) pig mailed elephant a souvenir. preceding sentence (result) pig gave elephant his address. target sentence then, bear sent him/him a postcard. question who did bear send a postcard? table 1. sample discourses in experiment 1, with two possible coherence relations. two aspects of the materials are worth pointing out. first, in contrast to taylor et al. (2013), the target sentence that contains the critical pronoun is the same for both parallel and result coherence relations (cf. 5 above). this allowed us to keep the event described in the target sentence constant across coherence relations, which is important because some verbs could bias more towards parallel or result relations independent of the preceding linguistic context. this also allowed us to use the exact same token, and hence the same prosody, across the manipulation; this is important because some prosodic cues preceding or following the pronoun could bias towards one interpretation or another. note that in order to keep the target sentence identical across the manipulation we used a connective that does not explicitly specify the relationship between the sentences (i.e., then); this meant that listeners had to infer the coherence relations from the events described (cf. tavano & kaiser, 2008). second, because our goal was to dissociate coherence and plausibility, we aimed to create materials where, for both parallel coherence and result coherence, the two possible interpretations of the pronoun would be equally plausible – we tested the materials to verify this. mozuraitis and heller 88 2.1 method 2.1.1 participants fifty-two undergraduate students at the university of toronto, all native speakers of english, participated in exchange for $5. participants were tested individually in a session that lasted approximately half an hour. nine additional participants were tested but were excluded from analysis because their accuracy on unambiguous comprehension questions in filler trials was below 85% (see materials). 2.1.2 materials (design and norming) sixteen three-sentence discourses were created, with some adapted from chambers and smyth (1998): see table 1 for an example (the full set of items is provided in the supplementary material folder). the first sentence (introduction sentence) introduced the scenario, always referring to the animals as a group. the second sentence (preceding sentence) mentioned two of the three animals using their name (e.g., bear, cat, elephant, etc.), one in subject position (the preceding subject) and the other in direct object position (the preceding object). the third and final sentence in the discourse (target sentence) mentioned a previously-unmentioned animal by name in subject position, and contained a linguistically-ambiguous pronoun in the object position. there were two experimental manipulations. coherence manipulated the discourse relation between the preceding sentence and the target sentence by changing the preceding sentence while keeping the target sentence identical across the manipulation. in parallel conditions, the preceding sentence depicted a similar event to the one depicted in the target sentence by using similar verbs (e.g., mailed a souvenir and sent a postcard). in this case, the default interpretation of an unaccented object pronoun was expected to be biased towards the preceding object (see e.g. chambers, & smyth, 1998). in result conditions, the preceding sentence described an event that could be taken as the cause of the event in the target sentence. in this case, the default interpretation of the unaccented object pronoun was expected to be biased towards the preceding subject (e.g., kehler et al., 2008). we verified that the parallel and result discourses did not differ with respect to the plausibility of the potential antecedents for the pronoun. in other words, our goal was to equate materials for plausibility, such that bear sent pig a postcard and bear sent elephant a postcard were both plausible continuations, for both coherence relations. to this end, we had participants recruited on amazon’s mechanical turk rank the discourses with a full name on a 1-7 scale, with each participant judging only one discourse – the complete details are given in appendix a. we analyzed the data using a mixed-effects linear regression model with plausibility rating (1-7) as the dependent variable, and coherence (parallel vs. result) and antecedent (preceding subject vs. preceding object) as fixed factors (as none of the slopes significantly improved the model, so the final model only included a random intercept for items2: see more under statistical modelling). the model revealed that a main effect of coherence (parallel: m = 5.88, sd = 1.49 vs. result: m = 4.87, sd = 1.98, β = -0.99, se = 0.20, χ2 = 24.45, p < 0.001), indicating that, overall, participants perceived the parallel discourses as more plausible than the result discourses. there was no main effect of antecedent (ps > 0.5), indicating that subject and object antecedents were equal in plausibility. most important, the coherence x antecedent interaction was not significant (parallel: 5.93 vs. 5.82; result 4.95 vs. 4.80; ps > 0.9), indicating that the preceding subject and the preceding object were equally plausible in parallel and result discourses. in other words, this finding confirms that our parallel and result discourses did not differ in terms of the plausibility of the two potential antecedents (cf. taylor et al., 2013). 2 because each participant judged only one story, there were no participant dependencies and hence no random effects for participants. discourse coherence and accented pronouns 89 the coherence manipulation described above was crossed with the accent manipulation in a 2x2 within-participants design. the accent manipulation determined whether the pronoun in the target sentence was unaccented (or unstressed) or accented (with a narrow focus accent). sentences were recorded in their entirety by a (female) native speaker of canadian english who was instructed to produce a phonological distinction of accenting the pronouns (no acoustic manipulation was performed). figure 1 plots two example sentences in the two conditions. as verification that accented pronouns indeed differed from their unaccented counterparts, we measured the duration of the pronouns and their mean pitch (f0). indeed, the accented pronouns were longer (483 ms, sd = 49 vs. 287 ms, sd = 37) and their average pitch was higher (211 hz, sd = 9.9 vs. 170 hz, sd = 8.8). figure 1. sample f0 tracks for the target sentences with an unaccented pronoun (top panel) and an accented pronoun (bottom panel). the four versions of each discourse (parallel-unaccented, parallel-accented, resultunaccented, and result-accented) were assigned to one of four presentation lists, such that each list contained an equal number of discourses in each experimental condition, and each participant encountered a specific discourse only once. each list also included thirty-four filler discourses. the fillers were 2 to 4 sentence long. furthermore, the filler discourses differed in structure from the experimental discourses: some contained ambiguous pronouns in subject position, some contained no pronouns at all, and some contained accented names (this was done in order to draw mozuraitis and heller 90 attention away from accented object pronouns). the filler discourses were interspersed within the experimental discourses, such that no more than two experimental discourses occur in sequence (adjacent experimental discourses were never in the same condition). each of the critical and filler discourses was followed by a comprehension question. in the experimental discourses, the question targeted the interpretation of the ambiguous pronoun, and the answer was used as the dependent variable (see again table 1). in the fillers, some questions were ambiguous whereas others were not. we used the twenty-one unambiguous questions in the filler trials as a measure that participants were paying attention; we only included in analysis those participants who answered at least eighteen questions (or more than 85%) correctly. 2.1.3 procedure each participant was randomly assigned to one of the four presentation lists, with an equal number of participants in each list. participants were told that the experiment investigated how people understood stories, and they should listen carefully so they could answer the question that would follow each story. an introductory screen displayed all six animals, e.g., “on the upper left is elephant, to his left is bear … “; the pronoun his was used in order to ensure that all animals will be taken to be masculine. two practice trials were used to familiarize participants with the procedure. on each trial, pictures of the three animals appeared on the screen. after a delay of 100 ms, the discourse played over speakers. when the discourse ended, the question was displayed visually on the screen, and participants were instructed to answer it by clicking on one of the three depicted animals. 2.1.4 statistical modelling given the categorical nature of the dependent variable, the data were analyzed using mixedeffects logistic regression models with participants and items as crossed, independent, random effects, implemented in package lme4 of the statistical software r 3.2.2 (bates, maechler, bolker, & walker, 2015; r core team, 2015). the independent variables were coded using deviation coding: parallel and unaccented were coded as -.5, whereas result and accented were coded as .5. pair-wise comparisons were conducted by recoding the different levels of the independent variables following west, aiken, and krull (1996). we used models with the structure of random effects that was supported by the data. to select the model with the appropriate structure of random effects, we used a backwards-selection method (cf. andrews, & lo, 2013; fine, & jaeger, 2013; taft, & krebs-lazendic, 2013; von bastian, & oberauer, 2013). specifically, we started from a model that included the full structure of random effects supported by the design, namely random slopes for the two fixed effects and their interaction for both participants and items, and random intercepts for participants and items. we then eliminated those random effects (specifically, random slopes) that did not improve the performance of the model, starting with the interaction and following a backwards-selection procedure. at minimum, all models included random intercepts for both participants and items. note that when conducting pairwise comparisons, the same structure of random effects was used. in the results section, we report the final structure of random effects used. 2.2 results the mean performance on comprehension questions in the unambiguous filler trials was 95%. the main dependent variable was how participants interpreted the ambiguous pronoun, as measured by their response to the comprehension question. we coded whether participants chose the character corresponding to the subjects of the preceding sentence (e.g., pig in the example in table 1) or the object of the preceding sentence (e.g., elephant in the example in table 1): table 2 provides the results in terms of the mean likelihood to choose the preceding object. let us first discourse coherence and accented pronouns 91 consider the choice of antecedents when the pronoun was unaccented. in parallel discourses, the previous object was chosen 65% of the time, confirming our prediction, in accordance with previous findings in the literature, that the preferred interpretation of an object pronoun is the previous object. in result discourses, the previous object was chosen only 18% of the time; that is, the subject was the preferred antecedent, again in accordance with previous findings in the literature. our main question, however, is what happens when the pronoun is accented. in the parallel discourses, the likelihood of choosing the preceding object was now lower (41% compared with 65%): the preferred antecedent is now the subject. in contrast, in result discourses, accenting did not seem to affect the preferred interpretation of the pronoun (20%, compared with 18%). discourse coherence previous object choices parallel unaccented 65% accented 41% result unaccented 18% accented 20% table 2. the mean likelihood of choosing the preceding object in experiment 1. our main question is whether accenting has a different effect depending on discourse coherence, which should be reflected in an interaction between discourse relation and accenting. but because the default interpretation of the pronoun in the unaccented cases is different, we cannot address this question by looking at object choices directly (this could result in an interaction that arises because of the default interpretation, and not change due to accenting). thus, for statistical analysis, we use a different dependent variable, noting whether participants chose the default antecedent for that context (i.e., the preferred antecedent for an unaccented pronoun). in parallel discourses, where the preceding object is the default antecedent, choosing the object is coded as 1 and choosing the subject is coded as 0 (yielding the same coding as in table 2). in result discourses, in contrast, where it is the preceding subjects that is the default antecedent, choosing the previous object will now be coded as 0 and choosing the subject is coded as 1 (this reverses the coding from table 2). figure 2 plots this new dependent variable: because the analysis used is logistic regression, the data is plotted in logit space. we fitted a mixed-effects logistic regression model (dependent variable: default interpretation = 1; non-default interpretation = 0) with coherence (parallel vs. results) and accent (unaccented vs. accented) as fixed factors. in the final model selected (see under statistical modelling), the random effect structure included a random slope for coherence for both participants and items, as well as random intercepts for both participants and items. the model revealed a main effect of accenting (β = 0.72, se = 0.18, z = 4.04, p < 0.001), indicating that participants were overall more likely to choose the default antecedent when the pronoun was not accented; this is expected if accenting changes the interpretation of pronouns. there was also a main effect of coherence (β = 1.66, se = 0.42, z = 3.96, p < 0.001), indicating that participants were overall more likely to choose the default referent for result discourses. most important, these main effects were qualified by a significant coherence x accent interaction (β = 1.07, se = 0.35, z = 3.02, p = 0.003), indicating that accenting had a different effect on pronoun interpretation depending on discourse coherence. specifically, accenting significantly changes the choice of antecedent in the parallel conditions (65% vs. 41%; β = 1.26, se = 0.23, z = 5.42, p < 0.001), but not in the result conditions (18% vs. 20%, p = 0.485). mozuraitis and heller 92 figure 2. the mean likelihood of choosing the default entity as a function of coherence and accent, experiment 1. this is presented in the logit space, in accordance with our logistic regression analysis (zero corresponds to 50% where there is no preference for either antecedent). error bars represent ±1 se estimated for pair-wise comparisons. we further asked whether the significant effect of accenting in the parallel condition can be characterized as a reversal of the preferred interpretation, namely whether the preferred antecedent changed. to address this question, we examined whether the likelihood of choosing the default antecedent in each condition differed significantly from chance; this was done using the error terms derived for the critical pair-wise comparisons above. indeed, in the parallelunaccented condition, participants chose the default referent significantly more than chance (65%, z = 3.44, p < 0.001), and in parallel-unaccented significantly less than chance (41%, z = 2.37, p = 0.018). these results indicate that accenting did not just have a significant effect on pronoun interpretation in parallel discourses, but this effect is one where the preferred antecedent is reversed3. 2.3 discussion first, these results replicate previous findings that accenting a pronoun has a different effect, depending on the coherence relations in the discourse. like tavano and kaiser (2008) and taylor et al. (2013), we find that accenting a pronoun reverses its interpretation in parallel discourses (i.e., changes which antecedent is preferred), but not in result discourses (i.e., the preferred antecedent stays the same). the current results extend previous findings by showing that (i) this effect is found even when the second event is consistent across coherence relations (cf. taylor et al., 2013), and (ii) even when the coherence relations are not explicitly given and have to be inferred by listeners (cf. tavano, & kaiser, 2008). this serves as further evidence against the generalization, originally due to akmajian and jackendoff (1970), that what accenting does to pronouns is reverse their interpretation. as such, these results also provide evidence against theories aimed to account for a pattern of reversal, such as kameyama’s (1999) complementary preference hypothesis. more important, these results extend previous findings in that they allow us to assess taylor et al.’s (2013) proposal that reversal will depend on the plausibility of the alternative interpretation. recall that in our materials the plausibility of the alternative interpretation did not differ from the plausibility of the default interpretation, for both parallel and result discourses 3 for completeness, we note that in both result discourses, participants chose the previous subject significantly more than chance (unaccented: z = 6.76, p < 0.001; accented: z = 6.14, p < 0.001). discourse coherence and accented pronouns 93 (even though result continuations were overall less plausible) – see again under materials. thus, while our findings in the main task are similar to taylor et al.’s, they cannot be interpreted the same way: taylor et al.’s proposal would wrongly predict reversal in our result discourses. having shown that plausibility is not responsible for the lack of reversal, the remainder of the paper examines other reasons for why parallel and result discourses respond differently to the accenting of pronouns. we continue to work with taylor et al.’s logic that (i) accenting normally reverses the interpretation of a pronoun, and (i) the reason reversal is not obtained in result discourses is that the alternative interpretation is not available. what could make the alternative not available in result discourses? one possibility we consider is that the alternative interpretation is not available if there is strong preference for the default interpretation. recall that in the unaccented condition, the preference for the default interpretation was significantly stronger in the result discourses than in the parallel discourses (82% vs. 65%; β = 1.23, se = 0.45, z = 2.48, p = 0.013). it is therefore possible that the lack of reversal was because the alternative interpretation was blocked due to the default interpretation being strongly preferred. we test this hypothesis in experiment 2. 3 experiment 2 the goal in experiment 2 is to examine whether the strength of preference for the default interpretation affects whether accenting creates a pattern of reversal. we test the hypothesis that a strong preference for the default interpretation renders the alternative interpretation unavailable. we focus on discourses with parallel relations, manipulating how much the default interpretation will be preferred – see table 3. this is achieved by changing the introduction sentence, which either creates an expectation that the one animal has a special role (strong preference), or leaves it open as to which animal will participate in which role (weak preference). for example, in table 3 the parallel-strong introduction sentence implies that there is one animal that went abroad, which should lead listeners to expect that actions such as mailing a souvenir and sending a postcard would be directed towards this particular character. this contrasts with the parallel-weak introduction sentence which is neutral with respect to which character may be the recipient of mail (this case is identical to the parallel discourse from experiment 1; see again table 1). introduction (parallel-weak) the animals went on a school exchange across the globe. introduction (parallel-strong) the animals were missing their friend who went abroad on a school exchange. preceding sentence pig mailed elephant a souvenir. target sentence then, bear sent him/him a postcard. question who did bear send a postcard? table 3. sample discourses in experiment 2 with two possible introduction sentences. if the alternative becomes unavailable when the default is strongly preferred, we would expect weak and strong parallel discourses to respond differently to accenting. in the strong parallel discourses, the alternative interpretation will not be available and thus an accented pronoun will receive the same interpretation as its unaccented counterpart. in the weak parallel discourses, the alternative interpretation will be available, just like in experiment 1. if instead we find that both parallel discourses respond similarly to accenting, this will provide evidence against the preference hypothesis. mozuraitis and heller 94 3.1 method 3.1.1 participants fifty-two undergraduate students at the university of toronto, all native speakers of english, participated in exchange for $5. none of these participants had participated in experiment 1. six additional participants were tested, but excluded from the final analyses because they did not meet criterion (85%) in answering the unambiguous comprehension questions. 3.1.2 materials the discourses were adapted from the parallel discourses in experiment 1. there were two experimental manipulations. strength of preference (weak vs. strong) manipulated the extent to which the default antecedent was preferred for the unaccented pronoun. this was achieved by changing the introduction sentence. to create a strong preference, the introductory sentence was changed to imply that one character is more likely to be the object of the events in the preceding and the target sentences. in the example in table 3, it is implied that there is one animal that is away and thus should be will be receiving mail, which should lead listeners to expect that it will be the object of mailed a souvenir and sent a postcard4. the weak preference versions of the discourses were neutral in that respect (these are identical to the parallel discourses used in experiment 1). to ensure that our manipulation did not also alter the plausibility of the alternative interpretation, we tested the plausibility of the two continuations as in experiment 1 by having participants from amazon’s mechanical turk rate the target sentence for plausibility on a scale of 1-7– see again appendix a. indeed, the mixed-effects regression model with plausibility rating as dependent variable and with strength of preference and antecedent as independent variables as well as random intercept for items revealed no significant main effects or interactions (ps > 0.07), indicating that the two parallel discourses, as well as both antecedents, were perceived to have the same level of plausibility. the rest of the design of the materials for the main task was identical to experiment 1. 3.1.3 procedure the procedure was identical to experiment 1. 3.1.4 statistical modelling the statistical modeling was as in experiment 1. 3.2 results the average performance on the comprehension questions of the unambiguous filler trials was 95%. table 4 provides the mean likelihood of choosing the previous object. the pattern of results for unaccented pronouns suggests that our manipulation was successful: the strength of preference manipulation caused a stronger preference (weak: 65% vs. strong: 81%). this puts us in a good position to assess whether accenting is affected by the strength of the preference. we observe that when the pronoun was accented, both parallel conditions responded similarly, 4 for some speakers, this introduction sentence is ambiguous between the desired reading where the animals were missing one friend (i.e., the wide scope reading), and a second reading where each animal was missing a different friend (i.e., the narrow scope reading); consider the full list of introduction sentences in the supplementary material folder. note, first, that if the narrow scope reading is available, it should create the opposite bias than the one intended: this is because it will encourage interpreting the target sentence as an event that involves a different animal (e.g., bear sending a postcard to someone other than elephant). the fact that our manipulation worked and the introduction sentence increased the bias as intended (81% vs. 65%) indicates that narrow scope reading is, at the very least, much less salient for most speakers. discourse coherence and accented pronouns 95 namely with a reversal pattern (weak: 45% vs. strong: 44%). that is, the strength of the bias toward the default interpretation does not seem to affect the availability of the alternative interpretation. discourse version previous object choices weak preference unaccented 65% accented 45% strong preference unaccented 81% accented 44% table 4. the mean likelihood of choosing the previous object in experiment 2. for purposes of statistical analysis, we again used the dependent variable of whether participants chose the default antecedent (i.e., the preferred antecedent for an unaccented pronoun), in order to have a comparable analysis to experiment 1 (unlike in experiment 1, here the default antecedent was the same for both conditions). figure 3 plots – in logit space – the mean proportions of choosing the default antecedent, or the previous object. figure 3.the mean likelihood (in logit space) of choosing the default antecedent as a function of strength of bias and accenting, experiment 2 (chance is 0). error bars represent ±1 se estimated for pair-wise comparisons. a mixed-effects logistic regression model was fitted to data, with strength of preference (weak vs. strong) and accenting (unaccented vs. accented) as fixed factors. the random effect structure in the final model included an uncorrelated random intercept and a random accent slope for participants, and a random intercept for items (see again the model selection procedure in 2.1.4). the model revealed a main effect of strength of preference (β = 0.55, se = 0.19, z = 2.99, p = 0.003), indicating that participants were overall more likely to choose the default antecedent (i.e., the previous object) in the strong preference condition than in the weak bias condition (63% vs. 55%). there was also a significant main effect of accent (β = 1.84, se = 0.34, z = 5.48, p < 0.001), indicating that participants were overall more likely to choose the default antecedent mozuraitis and heller 96 when the pronoun was unaccented (73% vs. 45%), as would be expected if accenting reversed the interpretation of pronouns. finally, the strength of preference x accent interaction was also significant (β = 1.25, se = 0.37, z = 3.34, p < 0.001), indicating that the effect of accenting depended on the strength of the preference for the default interpretation. pairwise comparisons showed that participants were significantly more likely to choose the default antecedent for the unaccented pronoun when the preference was strong, indicating that our manipulation was successful (81% vs. 65%: β = 1.18, se = 0.28, z = 4.23, p < 0.001). next, when examining the effect of accenting, we find that in both conditions an accented pronoun led participants to choose the previous object significantly less (strong: 81% vs. 44%, β = 2.46, se = 0.40, z = 6.17, p < 0.001; weak: 65% vs. 45% β = 1.21, se = 0.37, z = 3.30, p < 0.001), suggesting a pattern of reversal. a final question is whether the effect of accenting can be characterized as reversal; to address this question, we ask whether the likelihood of choosing an antecedent is different from chance. when the pronoun was unaccented, the likelihood of choosing the previous object was significantly above chance in both the weak preference conditions (65%, z = 2.44, p = 0.015) and in the strong preference condition (81%, z = 5.30, p < 0.001). when the pronoun was accented, the likelihood of choosing the previous object was numerically below chance in both conditions (45% vs. 44%), but neither reached significance (p=.395 and p=.285, respectively). this means that accenting changed the preferred antecedent, but it leads to a situation of ambiguity and not strictly reversal. 3.3 discussion the current experiment was designed to test the hypothesis that accenting a pronoun will not change its interpretation if the alternative interpretation becomes unavailable when it is strongly dispreferred in the context. we tested this hypothesis by focusing on discourses with parallel coherence, comparing the effect of accenting across discourses that were designed to create a weak or a strong bias towards the default reading. since the results confirm that our bias manipulation was successful (the unaccented pronoun was more likely to be interpreted as the previous object in the strong bias condition), we were able to evaluate our main hypothesis: in the accented condition, we found that accenting (numerically) reversed the preferred interpretation, independent of the strength of the bias. this pattern provides evidence against our hypothesis that a strong bias will render the alternative interpretation unavailable to be the antecedent of an accented pronoun. instead, we observed that both discourses with parallel coherence show a (numerical) reversal pattern in response to accenting. having ruled out a second idea for what makes the alternative interpretation unavailable, our final step is to consider a third possible reason. we observe that what distinguishes the result discourses (exp. 1) from the two types of parallel discourses (exp. 1 and 2) is the syntactic position of the default antecedent: in the discourses that showed reversal, the default antecedent was the previous object, whereas in the discourses that did not show reversal, the default antecedent was the previous subject. since it is well known that subject antecedents have a privileged status (e.g., they serve as topics and also more salient), it is possible that result discourses resist reversal because that would require switching to a less-salient object antecedent, whereas parallel discourses allow reversal, because here the switching is to a more salient antecedent. we test this possibility in experiment 3. 4 experiment 3 the goal of experiment 3 is to examine the hypothesis that accenting would reverse the interpretation of a pronoun only if it changes from an object antecedent, which is less salient, to a subject antecedent, which is more salient, but not vice versa. to this end, we adapted the result discourses from experiment 1 minimally by changing the preceding sentence, such that result discourse coherence and accented pronouns 97 coherence would be best implied with the preceding object, not the preceding subject. if reversal was not observes for result discourses in experiment 1 because it is not possible to switch to a less-salient alternative (i.e., an object antecedent instead of a subject antecedent), then the new result-object discourses should show reversal. this is because, like the parallel discourses, those would require switching to a more salient alternative, namely a subject antecedent. 4.1 method 4.1.1 participants fifty-two undergraduate students at the university of toronto, all native speakers of english, participated in exchange for $5. none of these participants had participated in experiments 1 or 2. nine additional participants were tested, but excluded from analysis, because they did not meet criterion (85%) in answering the unambiguous comprehension questions. 4.1.2 materials the discourses were adapted from experiment 1: the preceding sentence from the result discourses was changed, such that the result coherence will be best implied when the antecedent of the pronoun is the previous object – see table 5. as in experiment 1, we tested the materials in order to confirm that all versions of discourses used in the current experiment allowed plausible readings with both antecedents (e.g., “then, bear sent pig/elephant a postcard”) – see the supplementary material folder for a list of all items. the mixed-effects regression model with plausibility rating as dependent variable and coherence and antecedent as fixed factors, as well as random intercept for items, revealed that participants perceived the result discourses as overall less plausible (parallel: m = 5.32, sd = 1.82 vs. result-object: m = 5.88, sd = 1.493; β = -0.55, se = 0.19, χ2 = 8.25, p = 0.004). critically, however, there was no the main effect of antecedent, and the coherence x antecedent interaction was also not significant (p s > 0.2). this indicates that both the parallel and the result version of the discourses were equally plausible with both possible antecedent (5.82 vs. 5.93 and 5.47 vs. 5.17, respectively). that is, the change from result-subject to result-object did not affect the plausibility of the two continuations. the rest of the design was as in experiment 1. introduction sentence the animals went on a school exchange across the globe. preceding sentence (parallel) pig mailed elephant a souvenir. preceding sentence (result-object) pig missed elephant who left for india. target sentence then, bear sent him/him a postcard. question who did bear send a postcard? table 5. sample discourses in experiment 3 with two possible coherence relations. 4.1.3 procedure the procedure was identical to experiments 1 and 2. 4.1.4 statistical modelling the statistical modeling was as in experiments 1 and 2. 4.2 results the average performance on the comprehension questions of the unambiguous filler trials was 95%. table 6 provides the mean likelihood of choosing the previous object as the referent for the linguistically-ambiguous pronoun. when the pronoun was unaccented, it was interpreted as the mozuraitis and heller 98 preceding object 66% of the time in parallel discourses (cf. 65% for the exact same materials in both experiment 1 and experiment 2), and 71% of the time in the result-object discourses. this confirms that our manipulation was effective: in both the parallel discourses and the result-object discourses the default antecedent was now the previous object. as expected from previous experiments, in the parallel discourses, accenting led to reversal: the previous object was now dispreffered at 37% (cf. 41% in experiment 1 and 45% in experiment 2, for the exact same materials). interestingly, the result discourses here did not exhibit the reversal effect: the previous object was still preferred at 57%. discourse version previous object choices parallel unaccented 66% accented 37% result-object unaccented 71% accented 57% table 6. the mean likelihood of choosing the preceding object in experiment 3. here again our inferential analysis uses the dependent variable of the default antecedent, which is the previous object in both parallel and result-object discourses. figure 4 plots the mean proportions of choosing the default antecedent for the critical pronoun, in logit space. figure 4. the mean likelihood (in logit space) of choosing the default entity as a function of accent and coherence, experiment 3 (0 is chance). error bars represent ±1 se estimated for pair-wise comparisons. we fitted a mixed-effects logistic regression model (dependent variable: default interpretation = 1; non-default interpretation = 0) with coherence (parallel vs. result-object) and accent (unaccented vs. accented) as fixed factors. the random effect structure supported by the data was a random intercept for participants and a random intercept for items (see again the model selection process under 2.1.4). the model revealed a main effect of coherence (β = 0.58, se = 0.15, z = 3.74, p < 0.001), reflecting that, overall, participants were more likely to choose the discourse coherence and accented pronouns 99 default antecedent (i.e., previous object) in the result condition than in the parallel condition (64% vs. 51%). there was also a main effect of accent (β = 0.99, se = 0.16, z = 6.34, p < 0.001), indicating that participants were overall more likely to choose the default antecedent when the pronoun was unaccented (68% vs. 47%). importantly, these main effects were qualified by a significant coherence x accent interaction (β = 0.66, se = 0.31, z = 2.13, p = 0.033), indicating that accenting affected pronoun interpretation differently depending on coherence relations. specifically, although accenting had a significant effect on the choice of antecedent in both the parallel and the result-object conditions (parallel: 66% vs. 37%; β = 1.32, se = 0.22, z = 6.00, p < 0.001; result-object: 71% vs. 57%; β = 0.66, se = 0.22, z = 3.02, p = 0.002), the difference in parallel discourses was more pronounced than in the result discourses (β = 1.38 vs. β = 0.69). we also tested the question of reversal, namely whether the preference is different from chance. in parallel discourses, an unaccented pronoun was interpreted as the previous object significantly more than chance (66%; z = 2.97, p = .003), and an accented pronoun significantly less than chance (37%; z = 2.78, p = .005), replicating the reversal pattern we observed in experiment 1 (and numerically in both parallel discourses in experiment 2). in result-object discourses, an unaccented pronoun was interpreted as the preceding object significantly more than chance (71%; z = 4.02, p < .001); the accented counterpart was also numerically above chance (57%), although this difference did not reach significance (z = 1.43, p = 0.153). 4.3 discussion we find that the parallel discourses exhibited a reversal pattern, but the result-object discourses do not. the fact that these discourses respond differently to accenting despite both having the preceding object as the default antecedent suggests that accenting is not simply sensitive to the syntactic position of antecedents (although it is interesting to note the contrast between a significant effect of accenting in result-object and the lack of one in result-subject). 5 general discussion in three experiments, we investigated why in some discourses accenting a pronoun leads to a reversal of its preferred interpretation, while in others it does not. experiment 1 examined taylor et al.’s (2013) proposal that reversal is blocked when the alternative interpretation is not plausible. our results provide evidence against this proposal: when equated on the plausibility of the alternative, parallel discourses still showed reversal whereas result discourses did not. experiments 2 and 3 investigated two other hypotheses that share the same logic, namely that reversal is not observed if the alternative interpretation is not available for some reason. experiment 2 examined the hypothesis that the alternative is not available if the default interpretation is strongly preferred. counter this idea, we found that a parallel discourse with a strong preference for the default interpretation nonetheless exhibited reversal. experiment 3 examined the hypothesis that the alternative is not available if it requires switching to a less salient entity, namely from a subject antecedent to an object antecedent. counter this idea, we found that a result discourse where the default interpretation is the preceding object resists reversal, just like result discourses where the default interpretation is the preceding subject (it is worth noting, though, that accenting had a significant effect only in the object). note that the results of experiment 3 provide further evidence again the hypothesis considered in experiment 2: the absence of reversal in result-object discourses is another case where the bias is weak, but it nevertheless does not show reversal. these results add to previous findings showing that not all cases of accented pronouns are interpreted as an alternative referent (e.g., de hoop, 2004; tavano, & kaiser, 2008; german, 2009; taylor et al., 2013). as such, they constitute further evidence that the original observation by akmajian and jackendoff (1970) that accented pronouns receive the reverse interpretation of their unaccented counterparts, which has been widely adopted in the literature, is an overmozuraitis and heller 100 generalization. we therefore argue against mechanisms that try to account for the interpretation of accented pronouns by substituting the default antecedent with a different antecedent from the list of ‘currently salient’ entities (e.g., kameyama, 1999). although such proposals could explain the patterns we observed for parallel discourses, they make the wrong prediction for result discourses, where accenting the pronoun did not reverse its interpretation (as well as for other cases presented in de hoop, 2004 and german, 2009). furthermore, the current study also provides evidence against (three different instantiations of) a possible modification to this classical generalization: an accented pronoun will receive an alternative interpretation, unless this interpretation is blocked. the first reason for why the alternative interpretation would not be available was proposed by taylor et al. (2013): if this reading does not lead to a plausible continuation. but the result discourses in experiments 1 and 3 we show that a lack of reversal can be observed even when the alterative interpretation is plausible. we also proposed – and rejected – two other possible reasons. first, we considered the possibility that the alternative interpretation will not be available if it has to replace a default interpretation which is strongly preferred: the strong-bias parallel discourses in experiment 2 exhibited (numerical) reversal despite showing a strong bias for the default, and the result-object discourses in experiment 3 did not exhibit reversal despite only a weak bias for the default. the second reason we considered was that the alternative interpretation will not be available if it requires switching to less salient antecedent (e.g., switch from a subject to an object). the evidence against this possibility comes from the result-object discourses in experiment 3, which did not show reversal despite the default antecedent being the object. let us take stock. we found reversal in two kinds of parallel discourses, repeated in (6), and no reversal in two different types of result discourse, repeated in (7). (6) a. parallel-weak the animals went on a school exchange across the globe. pig mailed elephant a souvenir. then, bear sent him/him a postcard. b. parallel-strong the animals were missing their friend who went abroad on a school exchange. pig mailed elephant a souvenir. then, bear sent him/him a postcard. (7) a. result-subject the animals went on a school exchange across the globe. pig gave elephant his address. then, bear sent him/him a postcard. b. resultobject the animals went on a school exchange across the globe. pig missed elephant who left for india. then, bear sent him/him a postcard. we thus propose that the different behavior of accented pronouns in these cases is linked to discourse coherence. but how? to answer this question, we need to consider the role of accenting beyond pronouns. it has been accepted since rooth (1985) that the interpretation of accenting involves alternatives (for a review, see krifka, 2008). while the alternative can be narrowly the antecedent alone, an accented pronoun can mark accenting on whole verb phrase (or the predicate of the sentence), and so the interpretation would involve alternatives to sent x a postcard5. we propose that parallel and result discourses differ in the alternatives that are explicitly available in the discourse context. specifically, when alternative are called upon in order to interpret sent him 5 this point has been made by kehler (2005), but he assumes that the alternative computed would involve a different antecedent for the pronoun, and thus this is essentially equivalent to computing alternatives for the antecedent alone. discourse coherence and accented pronouns 101 a postcard, the parallel discourse contains an explicit alternative, namely the verb phrase of the preceding sentence (e.g., mailed elephant a souvenir). in order for the new verb phrase to contrast with this alternative (cf. de hoop, 2004), the pronoun has to receive a different interpretation (otherwise, the two verb phrase will be similar). in the result discourses, however, no explicit alternative is available against which sent him a postcard can be interpreted, and so listeners have to compute an appropriate alternative on the fly. our findings suggest that the most immediate way to compute (or accommodate) an alternative is using negation: the alternative to sent x a postcard is not(sent x a postcard). here the contrast is not between events with different participants, but rather between the event happening or not happening. we propose that this contrast gives rise to an inference that the event with the accenting is surprising or unexpected, a property that was proposed as part of meaning of accenting by both kehler (2005) and german (2009). what does this account of accented pronouns in parallel and results discourses mean for the interpretation of accented pronouns more generally? the interpretation of an accented pronoun requires contrasting the interpretation of the constituent in which the pronoun is embedded with an alternative. if an appropriate alternative is already explicitly available in the discourse context, this alternative will be used. if not, an alternative will be constructed (or accommodated), using minimal change to the current constituent, which is negation. reversal will be observed only under special circumstances, when information in the discourse context can serve as the (contrasting) alternative with a non-default interpretation of the pronoun. it is worth noting in this context that in all experimental work on the interpretation of accented pronoun, reversal was only observed for parallel discourses (avrutin, lubarsky, & greene, 1999; venditti, stone, nanda, & tepper, 2002). the other extreme is when the discourse context contains no information that can serve as the alternative: in this case, the alternative will be constructed (and accommodated) by applying negation to the asserted constituent. of course, many cases are in-between: the discourse context contains information that can be readily used to construct an alternative (e.g., in de hoop’s (4a) is it that mary is not from louisiana). our proposal is thus a departure from the standard view, originally due to akmajian and jackendoff (1970), that reversal is the unmarked interpretation of accented pronouns. acknowledgements this work was supported by grants from the social sciences and humanities research council to d. heller. we are grateful to natalie muradian for help with stimuli preparation and testing. mozuraitis and heller 102 appendix a: materials test the goal of materials test was to confirm that the versions of discourses used in experiments 1, 2, and 3 allowed equally plausible readings with both antecedents (e.g., elephant and bear) in the object position of the last sentence (e.g., “then, bear sent pig/elephant a postcard”). a.1. method participants. a total of 595 participants (330 male, 265 female) were recruited via amazon’s crowdsourcing platform mechanical turk over the course of 30 days. participants were paid $0.25 each. participants’ age was restricted to 18-50 years (mean = 31, sd = 7.45). all reported finishing high school, with 416 having attended or completed college and 79 having attended graduate school. additional 48 participants were excluded because (i) they were slow to complete the task (3.5 sds above the mean or 54 seconds, n = 8), (ii) they were extremely fast to complete the task (less than 5 seconds, n = 35), or because they indicated that at the age of 5 they spoke more than 2 languages (n = 5). materials and design. the discourses used were the same as the ones used in experiments 1, 2, and 3. thus, there were four versions of each discourse: parallel (as in experiments 1, 2, and 3), parallel–strong (as in experiment 2), result-subject (as in experiment 1), result–object (as in experiment 3). in addition, we manipulated whether the object of the target sentence included the animal that was the preceding subject or the preceding object. in contrast to experiments 1, 2, and 3, we used proper names and not pronouns (e.g., “then, bear sent pig/elephant a postcard” see table 7 for an example). this resulted in 8 versions for each of the 16 experimental discourses. instructions: in this simple task, you will read three sentences, and be asked to judge whether it makes sense for the last event (green sentence) to follow the earlier ones (blue sentences): see an example below. read the sentences once, and choose a number on the scale that reflects your gut feeling about how the events fit together. plausibility judgment task: does it make sense for the last event (green sentence) to follow the earlier ones (blue sentences)? ---------------------------------------- the animals went on a school exchange across the globe. pig mailed elephant a souvenir. then, bear sent elephant a postcard. ---------------------------------------- makes no sense at all 1 2 3 4 5 6 7 makes perfect sense table 7. example of the instructions and the plausibility judgment task. procedure. the entire experiment was conducted online using amazon’s mechanical turk. the whole experiment (the instructions, the sensibility judgment task, and the demographic questionnaire) was completed as a single assignment (“hit”) that turkers complete for payment. before beginning the assignment, interested turkers were presented with a description of the sensibility judgment task, an example of the task, and an electronic consent form. table 7 provides an example of the description of the sensibility judgment task, and an example of the task. turkers who agreed to participate were randomly assigned to 1 of the 128 discourses (4-5 discourse coherence and accented pronouns 103 participants per discourse). each participant rated the sensibility of a single discourse to prevent against any learning and/or fatigue effects. however, across all participants, each of the 16 experimental discourses appeared in all 8 conditions approximately equal number of times (4-5 times). after completing the plausibility judgment task, participants completed the demographic questionnaire asking about age, gender, education, and language background. the entire assignment lasted approximately 3 minutes. each participant was exposed only to 1 passage (between-participant manipulation). across the norming experiment, each passage used in the main experiments appeared in each of the 8 conditions (within-item manipulation). references akmajian, a., & jackendoff, r. (1970). coreferentiality and stress. linguistic inquiry, 1, 124– 126. andrews, s, & lo, s. (2013). is morphological priming stronger for transparent than opaque words? it depends on individual differences in spelling and vocabulary. journal of memory and language, 68, 279–296. ariel, m. (1990). accessing noun-phrase antecedents. london: routledge. arnold, j. e. (2010). how speakers refer: the role of accessibility. language and linguistic compass, 4, 187–203. avrutin, s., lubarsky, s., & greene, j. (1999). comprehension of contrastive stress by broca’s aphasics. brain and language, 70, 163–186. von bastian, c. c., & oberauer, k. (2013). distinct transfer effects of training different facets of working memory capacity. journal of memory and language, 69, 36–58. bates, d., maechler, m., bolker, b., & walker, s. (2015). fitting linear mixed-effects models using lme4. journal of statistical software, 67, 1–48. cahn, j. (1995). the effect of pitch accenting on pronoun referent resolution. acl '95 proceedings of the 33rd annual meeting on association for computational linguistics. stroudsburg, pa, usa. chambers, g. c., & smyth, r. (1998). structural parallelism and discourse coherence: a test of centering theory. journal of memory and language, 39, 593–608. fine, a. b. and jaeger, t. f. 2013. evidence for implicit learning in syntactic comprehension. cognitive science, 37, 578–591. german, j. s. (2009). prosodic strategies for negotiating reference in discourse. (unpublished doctoral dissertation). northwestern university, evanston, illinois. grosz, b., j., joshi, a. k., & weinstein, s. (1995). centering: a framework for modeling the local coherence of discourse. computational linguistics, 21, 203-225. gundel, j., hedberg, n., & zacharski, r. (1993). cognitive status and the form of referring expressions in discourse. language, 69, 274-307. hirschberg, j., & ward, g. (1991). accent and bound anaphora. cognitive linguistics, 2, 101121. hobbs, j. r. (1978). resolving pronoun references. lingua, 44, 311-338. hobbs, j. r. (1979). coherence and coreference. cognitive science, 3, 67–90. hobbs, j. r., stickel, m. e., appelt, d. e., & martin, p. (1993). interpretation as abduction. artificial intelligence, 63, 69–142. de hoop, h. (2004). on the interpretation of stressed pronouns. in r. blutner and h. zeevat (eds.), optimality theory and pragmatics (pp. 25–41). palgrave macmillan. new york. jasinskaja, e., kölsch, u., & mayer, j. (2007). nuclear accent placement and other prosodic parameters as cues to pronoun resolution. in branco, a. (ed.) anaphora: analysis, algorithms and applications (pp. 1-14). berlin/heidelberg: springer. kameyama, m. (1994) stressed and unstressed pronouns: complementary preferences. in bosch, mozuraitis and heller 104 p. and van der sandt, r. (eds.), focus and natural language processing (pp. 475–484). institute for logic and linguistics, ibm, heidelberg. kameyama, m. (1999), ‘stressed and unstressed pronouns: complementary preferences’. in p. bosch and r. van der sandt (eds.), focus: linguistic, cognitive, and computational perspectives (pp. 306–321). cambridge: cambridge university press. kehler, a. (2002). coherence, reference, and the theory of grammar. stanford, ca: csli publications. kehler, a. (2005). coherence-driven constraints on the placement of accent. in e. georgala and j. howell (eds.), proceedings of the 15th conference on semantics and linguistic theory (salt-15). clc publications. cornell university. 98–115. kehler, a., kertz, l., rohde, h., & elman, j. l. (2008). coherence and coreference revisited. journal of semantics, 25, 1–44. krifka, m. (2008). basic notions of information structure. acta linguistica hungarica, 55(3), 243–276. nakatani, c. h. (1997). the computational processing of intonational prominence: a functional prosody perspective, phd thesis, harvard university, cambridge, massachusetts. r core team (2015). r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. rooth, m. (1985) association with focus, phd thesis, glsa, university of massachusetts, amherst. smyth, r. (1994). grammatical determinants of ambiguous pronoun resolution. journal of psycholinguistic research, 23, 197–229. stevenson, r. j., crawley, r. a., & kleinman, d. (1994). thematic roles, focus, and the representation of events. language and cognitive processes, 9, 519–48. stevenson, r. j., knott, a., oberlander, j., & mcdonald, s. (2000). interpreting pronouns and connectives: interactions among focusing, thematic roles, and coherence relations. language and cognitive processes, 15, 225–262. taft, m., & krebs-lazendic, l. (2013). the role of orthographic syllable structure in assigning letters to their position in visual word recognition. journal of memory and language, 68, 85– 97. tavano, e., & kaiser, e. (2008). effects of stress and coherence on pronoun interpretation. poster presented at the 21st annual cuny conference on human sentence processing, chapel hill, nc. taylor, r. c., stowe, l. a., redeker, g., & hoeks, j. c. j. (2013). comprehension of marked pronouns in spanish and english: object anaphors cross-linguistically. the quarterly journal of experimental psychology, doi:10.1080/17470218.2013.773356 venditti, j. j., stone, m., nanda, p., & tepper, p. (2002). discourse constraints on the interpretation of nuclear-accented pronouns. in proceedings of the 2002 international conference on speech prosody. aix-en-provence, france. west, s. g., aiken, l. s., & krull, j. l. (1996). experimental personality designs: analyzing categorical by continuous variable interactions. journal of personality, 64, 1–48. wolf, f., gibson, e., & desmet, t. (2004). discourse coherence and pronoun resolution. language and cognitive processes, 19, 665–675. dialogue and discourse 2(1) (2011) 143-170 doi: 10.5087/dad.2011.107 incremental interpretation and prediction of utterance meaning for interactive dialogue david devault devault@ict.usc.edu usc institute for creative technologies 12015 waterfront drive playa vista, ca 90094 usa kenji sagae sagae@ict.usc.edu usc institute for creative technologies 12015 waterfront drive playa vista, ca 90094 usa david traum traum@ict.usc.edu usc institute for creative technologies 12015 waterfront drive playa vista, ca 90094 usa editor: david schlangen, hannes rieser abstract we present techniques for the incremental interpretation and prediction of utterance meaning in dialogue systems. these techniques open possibilities for systems to initiate responsive overlap behaviors during user speech, such as interrupting, acknowledging, or completing a user’s utterance while it is still in progress. in an implemented system, we show that relatively high accuracy can be achieved in understanding of spontaneous utterances before utterances are completed. further, we present a method for determining when a system has reached a point of maximal understanding of an ongoing user utterance, and show that this determination can be made with high precision. finally, we discuss a prototype implementation that shows how systems can use these abilities to strategically initiate system completions of user utterances. more broadly, this framework facilitates the implementation of a range of overlap behaviors that are common in human dialogue, but have been largely absent in dialogue systems. 1. introduction human spoken dialogue is highly interactive, including feedback on the speech of others while the speech is progressing (so-called “backchannels” (yngve 1970)), monitoring of addressees and other listener feedback (nakano et al. 2003), fluent turn-taking with little or no delays (sacks et al. 1974), and overlaps of various sorts, including collaborative completions (goodwin 1979), repetitions and other grounding moves (clark and schaefer 1987), and interruptions. interruptions can be either to advance the new speaker’s goals (which may not be related to interpreting the other’s speech) or in order to prevent the speaker from finishing. few of these behaviors can be replicated by current spoken dialogue systems. most of these behaviors require first an ability to perform incremental interpretation, and second, an ability to predict the final meaning and timing of the utterance. c©2011 david devault, kenji sagae, and david traum submitted 1/2010; accepted 12/2010; published online 5/2011 devault, sagae and traum most spoken dialogue systems wait until the user stops speaking before trying to understand and react to what the user is saying. in particular, in a typical dialogue system pipeline, it is only once the user’s spoken utterance is complete that the results of automatic speech recognition (asr) are sent on to natural language understanding (nlu) and dialogue management, which then triggers generation and synthesis of the next system utterance. while this style of interaction is adequate and perhaps even preferred for some applications (funakoshi et al. 2010), it enforces a rigid pacing that can be unnatural and inefficient for mixed-initiative dialogue. to achieve more flexible turntaking with human users, for whom turn-taking and feedback at the sub-utterance level is natural and common, the system needs to engage in incremental processing, in which interpretation components are activated, and in some cases decisions are made, before the user utterance is complete. incremental interpretation enables more rapid response, since most of the utterance can be interpreted before utterance completion (skantze and schlangen 2009). it also enables giving early feedback (e.g., head nods and shakes, facial expressions, gaze shifts, and verbal backchannels) to signal how well things are being perceived, understood, and evaluated (allwood et al. 1992). for some responsive behaviors, one must go beyond incremental interpretation and predict some aspects of the full utterance before it has been completed. for behaviors such as complying with the evocative function (allwood 1995) or intended perlocutionary effect (sadek 1991), grounding by demonstrating (clark and schaefer 1987), or interrupting to avoid having the utterance be completed, one must predict the semantic content of the full utterance from a partial prefix fragment. for other behaviors, such as timing a reply to have little or no gap, grounding by saying the same thing at the same time (called “chanting” by hansen et al. (1996)), performing collaborative completions (clark and wilkes-gibbs 1986), or some corrections, it is important not only to predict the meaning, but also the form and timing of the remaining part of the utterance. we have begun to explore these issues in the context of the dialogue behavior of virtual humans (rickel and johnson 1999) (also called embodied conversational agents (cassell et al. 2000)) for multiparty negotiation role-playing (traum et al. 2008). in these kinds of systems, human-like behavior is a goal, since the purpose is to allow a user to practice this kind of dialogue with the virtual humans in training for real negotiation dialogues. the more realistic the characters’ dialogue behavior is, the more kinds of negotiation situations can be adequately trained for. we discuss these systems further in section 2. in the remainder of the paper, we present the progress we have made to date in implementing responsive behaviors in our systems. this paper summarizes and expands on several previously reported results. in sagae et al. (2009), we presented our first results at prediction of semantic content from partial speech recognition hypotheses, looking at length of the speech hypothesis as a general indicator of semantic accuracy in understanding. sections 3 and 4 summarize this previous work, and also provide additional details and examples for our predictive model. section 3 presents our basic approach to understanding complete user utterances, and section 4 shows how we have extended this approach to enable predictive, incremental understanding of partial utterances. in devault et al. (2009), we extended our previous work by incorporating additional features of real-time incremental interpretation to develop a more nuanced prediction model that can accurately identify moments of maximal understanding within individual spoken utterances. we summarize this work in section 5. in section 6, we explore the potential value of this new ability using a prototype implementation that collaboratively completes user utterances when the system becomes confident about how the utterance will end. this section provides a more detailed evaluation and discussion of the characteristics of this prototype utterance completion capability than was present in 144 incremental interpretation figure 1: saso-en negotiation in the cafe: dr. perez (left) looking at elder al-hassan. devault et al. (2009). finally, this paper adds a new, expanded discussion of all of these techniques in section 7, including discussion of their limitations, potential role in system building, and future research directions. we conclude in section 8. we believe the predictive models presented in this paper will be more broadly useful in implementing responsive overlap behaviors such as rapid grounding using completions, confirmation requests, or paraphrasing, as well as other kinds of interruptions and multi-modal displays. 2. domain setting the case study we present in this paper is taken from the saso-en virtual human system (hartholt et al. 2008, traum et al. 2008). this system is designed to allow a trainee to practice multi-party negotiation skills by engaging in face to face negotiation with virtual humans. the scenario involves a negotiation about the possible re-location of a medical clinic in an iraqi village. a human trainee plays the role of a us army captain, and there are two virtual humans that he negotiates with: doctor perez, the head of the ngo clinic, and a local village elder, al-hassan. the doctor’s main objective is to treat patients. the elder’s main objective is to support his village. the captain’s main objective is to move the clinic out of the marketplace, ideally to the us army base. figure 1 shows 145 devault, sagae and traum the doctor and elder in the midst of a negotiation, from the perspective of the trainee. figure 2 presents a sample dialogue from this domain. 1 c hello doctor perez. 2 d hello captain. 3 e hello captain. 4 c thank you for meeting me. 5 e how may i help you? 6 c i have orders to move this clinic to a camp near the us base. 7 e we have many matters to attend to. 8 c i understand, but it is imperative that we move the clinic out of this area. 9 e this town needs a clinic. 10 d we can’t take sides. 11 c would you be willing to move downtown? 12 e we would need to improve water access in the downtown area, captain. 13 c we can dig a well for you. 14 d captain, we need medical supplies in order to run the clinic downtown. 15 c we can deliver medical supplies downtown, doctor. 16 e we need to address the lack of power downtown. 17 c we can provide you with power generators. 18 e very well captain, i agree to have the clinic downtown. 19 e doctor, i think you should run the clinic downtown. 20 d elder, the clinic downtown should be in an acceptable condition before we move. 21 e i can renovate the downtown clinic, doctor. 22 d ok, i agree to run the clinic downtown, captain. 23 c excellent. 24 d i must go now. 25 e i must attend to other matters. 26 c goodbye. 26 d goodbye. 26 e farewell, sir. figure 2: successful negotiation dialogue between c, a captain (human trainee), d, a doctor (virtual human), and e, a village elder (virtual human). the system has a fairly typical set of processing components for virtual humans or dialogue systems, including asr (mapping speech to words), nlu (mapping from words to semantic frames), dialogue interpretation and management (handling context, dialogue acts, reference and deciding what content to express), nlg (mapping frames to words), non-verbal generation, and synthesis and realization. the doctor and elder use the same asr and nlu components, but have different 146 incremental interpretation modules for the other processing, including different models of context and goals, and different output generators. 3. understanding complete user utterances in the saso-en system we begin by reviewing the technical approach we have used to understand complete user utterances in saso-en. we first define the nlu task in section 3.1, and then describe the data used in our experiments in section 3.2. we present our nlu model in section 3.3, and summarize our results in section 3.4. we will turn to the understanding of partial user utterances in section 4. 3.1 the natural language understanding task the nlu module used in the saso-en virtual human system takes the output of asr as input, and produces domain-specific semantic frames as output. for all the results presented in this paper, we have used sonic (pellom 2001) as the asr component. alternative asr components are compatible with our techniques, however. our techniques require only for the asr to provide the top-ranking text hypothesis for the utterance, and for the asr to be able to provide its top hypothesis incrementally, as user speech progresses and additional audio is captured. indeed, in more recent work, we have begun to use the pocketsphinx asr (huggins-daines et al. 2006) as a substitute for sonic. given the top-ranking text hypothesis from the asr component, the nlu module analyzes the utterance into a domain-specific semantic frame representation. these frames are intended to capture much of the meaning of the utterance, although a dialogue manager further enriches the frame representations with pragmatic information (traum 2003). the nlu output frame representation is an attribute-value matrix (avm), where the attributes and values represent semantic information that is linked to a domain-specific ontology and task model (hartholt et al. 2008). complicating the nlu task is the relatively high word error rate (0.54) in asr of user speech input, given conversational speech in a complex domain and an untrained broad user population. figure 3 shows an example of nlu input and output for an utterance where the user attempts to address complaints about lack of power in the proposed location for the clinic. in the figure, we show both the ideal, hypothetical condition in which the asr output perfectly matches the user utterance (we are prepared to give you guys generators for electricity downtown) as well as the actual condition in which the asr output includes several errors. given either of these asr results as input, the nlu output is the same in our system, which illustrates the desired robustness to asr errors. the figure shows the nlu frame, which is an avm, linearized using a path-value notation. the linearized semantic frame in this example corresponds to the avm shown in figure 4. 3.2 data the work described in the following sections is based on a corpus of user data consisting of 4,500 spoken user utterances, which were spread across a number of different dialogue sessions in the saso-en system. each user utterance was transcribed, and each of the complete utterance transcripts was manually annotated with a “gold standard” semantic frame. this annotated frame is viewed as the desired output from the nlu module for that complete utterance. utterances that were judged to be out-of-domain (13.7% of the corpus) were assigned to a “garbage” frame, with no semantic content. 147 devault, sagae and traum .mood declarative .sem.agent captain−kirk .sem.event deliver .sem.modal.possibility can .sem.speechact.type offer .sem.theme power−generator .sem.source us−army .sem.type event .mood declarative .sem.agent captain−kirk .sem.event deliver .sem.modal.possibility can .sem.speechact.type offer .sem.theme power−generator .sem.source us−army .sem.type event we are prepared to give you guys generators for electricity downtown captured user speech we up apparently give you guys generators for a letter city don town ideal processing actual processing nlu output: (semantic frame) asr output: (text string, input to nlu) figure 3: example of nlu input and output for a specific user utterance.  mood : declarative sem :  type : event agent : captain− kirk event : deliver theme : power − generator source : us− army modal : [ possibility : can ] speech− act : [ type : offer ]   figure 4: avm utterance representation. 148 incremental interpretation for training, development and evaluation of data-driven nlu models, the corpus was divided as follows. approximately 10% of the utterances were set aside for evaluation, and another 10% were designated as the development test corpus for the nlu module. the remaining 80% was used for training the nlu module. the training set included 136 distinct frames. note that the development and test sets were chosen so that all the utterances from the same dialogue session were kept in the same set, but sessions were chosen at random for inclusion in the development and test sets. 3.3 nlu as multiclass classification with mxnlu the saso-en nlu module, mxnlu (sagae et al. 2009), is based on a data-driven approach to language understanding that treats the task as multiclass classification. we use a supervised machine learning technique to learn a mapping from user utterances to semantic representations. more specifically, mxnlu uses maximum entropy (me) models (berger et al. 1996), where entire semantic frames are treated as classes, and features used for utterance classification are derived from the text string produced by asr. our task is then to estimate a model that determines the conditional probability that the user has expressed the meaning represented by a specific frame y, given that asr output x was observed, which we denote by p(y|x). the model is estimated using the maximum entropy framework (see berger et al. (1996) for details) according to a set of training examples {(x1, y1), (x2, y2), ..., (xn, yn)}, and has the following parametric form: p(y|x) = 1 z(x) exp (∑ i λifi(x, y) ) (1) where z(x) = ∑ y exp( ∑ i λifi(x, y)) is a normalizing constant determined by the requirement that ∑ y p(y|x) = 1 for all x, each fi is a feature function, and the parameters λi can be thought of as “weights” that reflect the importance of the corresponding feature fi. the features used by mxnlu come from a fixed set of templates used to generate feature functions similar to the one below: fj(x, y) = { 1 if y = y′ and x includes the word “generators” 0 otherwise although it is possible to define different types of features that are specific to each individual classes (frames) y′, in practice mxnlu considers the same types of features for every frame. this means that the me model includes features such as the one shown above for every y′ corresponding to each specific frame observed in the training examples. each feature function fi is then intended to capture a specific event observed in the output of asr (such as the occurrence of the word “generators”, or the occurrence of the word sequence “we can”), and the corresponding parameter λi is intended to capture the relationship between that event and a specific output frame. the events captured by the templates used to generate features for mxnlu are: each word in the input string (bag-of-words representation of the input), each bigram (consecutive words), each pair of any two words in the input, and the number of words in the input string. table 1 illustrates what information about a specific text string is captured by the feature functions used in mxnlu. note that the features generated for the bigram “we can” and the pair of words “we can” are distinct, and each is treated separately by the classifier, according to their respective λi parameters. the training examples used to create the me model used by mxnlu come from the corpus of user data described in section 3.2, which includes asr output from audio files containing recorded 149 devault, sagae and traum asr output (nlu input) we can help you words we, can, help, you bigrams ( we), (we can), (can help), (help you), (you ) pairs of words (we can), (we help), (we you), (can help), (can you), (help you) length 4 table 1: information about a specific utterance, represented as a text string produced by asr, captured by the feature functions used in mxnlu. the special tokens and mark the beginning and end of the utterance, respectively. utterances directed at the system and corresponding semantic representations. each pair of user utterance (represented as the output of asr) and semantic frame (represented as a linearized avm) is used to create an individual training example. each training example is composed of a feature vector (generated using the feature templates listed above) and a corresponding class (i.e. an entire semantic frame, which is treated as an atomic entity, rather than compositionally). given the set of training examples, the maximum entropy classifier should learn, for example, that when the word “generators” appears in the output of asr, the correct output frame is likely to be one that includes the value power-generator. although such a mapping between a specific feature of the nlu input and individual keys or values in nlu output frames cannot be made explicitly in our nlu model, since frames are never decomposed into individual keys and values, it can still hold implicitly, under the assumption that input instances that include the word “generators” are associated with frames that include the value power-generator in the training data. another consequence of our choice not to decompose frames is that the process of mapping utterances to semantic representations is divorced from the complexity and other details of the frame representation. for example, the frame language could allow for co-indexation of values and reentrant structures (although these are not used in saso-en frames), which mxnlu would have no problems dealing with. however, mxnlu would not treat these phenomena productively, but rather would simply memorize each instance as a part of a specific frame, which it could reproduce at run time. in other words, mxnlu can only produce a finite number of output frames, which correspond exactly to the set of frames present in the training data. although this is a clear theoretical limitation of the approach of classifying utterances into entire frames at once, the precise impact of this limitation on the efficacy of this nlu approach depends on specific characteristics of the dialogue system and the semantic representation it uses. in our corpus of user data, only 136 distinct frames were observed for the 3,500 user utterances in the training set, and the development set contained 150 incremental interpretation no frames that had not been observed in the training set. this indicates that the lack of productive ability in the nlu approach should have little negative impact, if any, in saso-en, which is a system that supports relatively rich natural language interaction in a limited domain. 3.4 nlu performance on complete asr output although mxnlu produces entire frames as output, we evaluate nlu performance by looking at precision and recall of the attribute-value pairs (or frame elements) that compose frames. precision represents the portion of frame elements produced by mxnlu that were correct, and recall represents the portion of frame elements in the gold-standard annotations that were proposed by mxnlu. by using precision and recall of frame elements, we take into account that certain pairs of frames are more similar than others and also allow more meaningful comparative evaluation with nlu modules that construct a frame from sub-elements or for cases when the actual frame is not in the training set. when mxnlu is trained on complete asr utterances in the training set, and tested on complete asr utterances in the development test set, the f-score of frame elements is 0.76, with precision at 0.78 and recall at 0.74. to gain insight on what the upperbound on the accuracy of the nlu module might be, we also trained the classifier using features extracted from gold-standard manual transcription (instead of asr output), and tested the accuracy of analyses of gold-standard transcriptions (which would not be available at run-time in the dialogue system). under these ideal conditions, the nlu f-score is 0.87. training on gold-standard transcriptions and testing on asr output produces results with a lower f-score, 0.74. 4. understanding partial user utterances in saso-en in this section, we refine the non-incremental nlu module described in section 3 in a way that enables the incremental understanding and prediction of utterance meaning during user speech. there is a growing body of work on incremental processing in dialogue systems. some of this work has demonstrated overall improvements in system responsiveness and user satisfaction; e.g. aist et al. (2007), skantze and schlangen (2009). several research groups, inspired by psycholinguistic models of human processing, have also been exploring technical frameworks that allow diverse contextual information to be brought to bear during incremental processing; e.g. kruijff et al. (2007), aist et al. (2007). while this work often assumes or suggests it is possible for systems to understand partial user utterances even before those utterances are complete, this premise has generally not been given detailed quantitative study (though see schlangen et al. (2009), heintze et al. (2010)). in this section, we demonstrate and explore quantitatively the extent to which our saso-en dialogue system can anticipate what an utterance means, on the basis of partial asr results, before the utterance is complete. a useful starting point is to observe the length, measured in words, of complete spoken user utterances in our corpus. figure 5 shows the utterance length distribution in the development set. roughly half of the utterances in our data contain six words or more, and the average utterance length is 5.9 words. since the asr module is capable of sending partial results to the nlu module even before the user has finished an utterance, in principle the dialogue system can start understanding and even responding to user input as soon as enough words have been uttered to give the system some indication of what the user means, or even what the user will have said once the utterance 151 devault, sagae and traum 0 10 20 30 40 50 60 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 n um be r of u tte ra nc es ( ba rs ) 0 50 100 150 200 250 300 350 400 450 c um ul at iv e nu m be r of ut te ra nc es ( lin e) length n (words) at most n words exactly n words figure 5: length of utterances in the development set. is completed. to measure the extent to which our nlu module can predict the frame for an input utterance when it sees only a partial asr result with the first n words, we examine two aspects of nlu with partial asr results. the first is correctness of the nlu output with partial asr results of varying lengths, if we take the gold-standard manual annotation for the entire utterance as the correct frame for any of the partial asr results for that utterance. the second is stability: how similar the nlu output with partial asr results of varying lengths is to what the nlu result would have been for the entire utterance. the most straightforward way to perform nlu of partial asr results is simply to process the partial utterances using the nlu module trained on complete asr output. however, as we will show, better results may be obtained by training separate nlu models for analysis of partial utterances of different lengths. to train these separate nlu models, we first ran the audio of the utterances in the training data through our asr module, in 200 millisecond increments, and recorded all the partial asr results for each utterance. then, to train a model to analyze partial asr results containing n words, we used only those partial asr results in our training set containing n words (unless the complete asr result contained less than n words, in which case we simply used the complete asr result). in some cases, multiple partial asr results for a single utterance contained the same number of words, and we used the last partial result with the appropriate number of words.1 we trained separate nlu models for n varying from one to ten. 1. at run-time, this can be closely approximated by taking the partial utterance immediately preceding the first partial utterance of length n + 1. 152 incremental interpretation 0 10 20 30 40 50 60 70 80 1 2 3 4 5 6 7 8 9 10 all fsc or e trained on partials up to length n (words) length n + context trained on complete asr results trained on partials up to length n figure 6: correctness for three nlu models on partial asr results up to n words. figure 6 shows the f-score for frames obtained by processing partial asr results up to length n using three variants of mxnlu. the dashed line is our baseline nlu model, trained on complete utterances only (model 1). the solid line shows the results obtained with length-specific nlu models (model 2), and the dotted line shows results for length-specific models that also use features that capture dialogue context (model 3). in these experiments, we used unigram and bigram word features extracted from the most recent system utterance to represent context, but found that these context features did not improve nlu performance. our final nlu approach for partial asr hypotheses is then to train separate models for specific lengths, using hypotheses of that length during training (solid line in figure 6). as seen in figure 6, there is a clear benefit to training nlu models specifically tailored for partial asr results. training a model on partial utterances with four or five words allows for relatively high f-score of frame elements (0.67 and 0.71, respectively, compared to 0.58 and 0.66 when the same partial asr results are analyzed using model 1). considering that half of the utterances are expected to have more than five words (based on the length of the utterances in the training set), allowing the system to start processing user input when four or five-word partial asr results are available provides interesting opportunities. targeting partial results with seven words or more is less productive, since the time savings are reduced, and the gain in accuracy is modest. 153 devault, sagae and traum 0 10 20 30 40 50 60 70 80 90 100 1 2 3 4 5 6 7 8 9 10 s ta bi lit y f -s co re length n of partial asr output used in model 2 figure 7: stability of nlu results for partial asr results up to length n . the context features used in model 3 did not provide substantial benefits in nlu accuracy. it is possible that other ways of representing context or dialogue state may be more effective. this is an area we are currently investigating. finally, figure 7 shows the stability of nlu results produced by model 2 for partial asr utterances of varying lengths. this is intended to be an indication of how much the frame assigned to a partial utterance differs from the ultimate nlu output for the entire utterance. this ultimate nlu output is the frame assigned by model 1 for the complete utterance. stability is then measured as the f-score between the output of model 2 for a particular partial utterance, and the output of model 1 for the corresponding complete utterance. a stability f-score of 1.0 would mean that the frame produced for the partial utterance is identical to the frame produced for the entire utterance. lower values indicate that the frame assigned to a partial utterance is revised significantly when the entire input is available. as expected, the frames produced by model 2 for partial utterances with at least eight words match closely the frames produced by model 1 for the complete utterances. although the frames for partial utterances of length six are almost as accurate as the frames for the complete utterances (figure 6), figure 7 indicates that these frames are still often revised once the entire input utterance is available. 5. detecting points of maximal understanding in this section, we present a strategy that uses machine learning to more closely characterize the performance of a maximum entropy based incremental nlu module, such as the mxnlu module described in section 4. our aim is to identify strategic points in time, as a specific utterance is occurring, when the system might react with confidence that the interpretation will not significantly 154 incremental interpretation improve during the rest of the utterance (devault et al. 2009). this reaction could take several forms, including providing feedback, or, as described in section 6 an agent might use this information to opportunistically choose to initiate a completion of the user’s utterance. 5.1 motivating example figure 8 illustrates the incremental output of mxnlu as a user asks, elder do you agree to move the clinic downtown? our asr processes captured audio in 200ms chunks. the figure shows the partial asr result after the asr has processed each 200ms of audio, along with the f-score achieved by mxnlu on each of these partials. note that the nlu f-score fluctuates somewhat as the asr revises its incremental hypotheses about the user utterance, but generally increases over time. for the purpose of initiating an overlapping response to a user utterance such as this one, the agent needs to be able (in the right circumstances) to make an assessment that it has already understood the utterance “well enough”, based on the partial asr results that are currently available. we have implemented a specific approach to this assessment which views an utterance as understood “well enough” if the agent would not understand the utterance any better than it currently does even if it were to wait for the user to finish their utterance (and for the asr to finish interpreting the complete utterance). concretely, figure 8 shows that after the entire 2800ms utterance has been processed by the asr, mxnlu achieves an f-score of 0.91. however, in fact, mxnlu already achieves this maximal f-score at the moment it interprets the partial asr result elder do you agree to move the at 1800ms. the agent therefore could, in principle, initiate an overlapping response at 1800ms without sacrificing any accuracy in its understanding of the user’s utterance. of course the agent does not automatically realize that it has achieved a maximal f-score at 1800ms. to enable the agent to make this assessment, we have trained a classifier, which we call maxf, that can be invoked for any specific partial asr result, and which uses various features of the asr result and the current mxnlu output to estimate whether the nlu f-score for the current partial asr result is at least as high as the mxnlu f-score would be if the agent were to wait for the entire utterance. 5.2 machine learning setup to facilitate the construction of our maxf classifier, we identified a range of potentially useful features that the agent could use at run-time to assess its confidence in mxnlu’s output for a given partial asr result. these features are exemplified in figure 9, and include: k, the number of partial results that have been received from the asr; n , the length (in words) of the current partial asr result; entropy, the entropy in the probability distribution mxnlu assigns to alternative output frames; pmax, the probability mxnlu assigns to the most probable output frame; nlu, the most probable output frame (represented for convenience as fi , where i is an integer index corresponding to a specific complete frame). we also define maxf (gold), a boolean value giving the ground truth about whether mxnlu’s f-score for this partial is at least as high as mxnlu’s f-score for the final partial for the same utterance. in the example, note that maxf (gold) is true for each partial where mxnlu’s f-score (f (k)) is ≥ 0.91, the value achieved for the final partial (elder do you agree to move the clinic downtown). of course, the actual f-score f (k) is not available at run-time, and so cannot serve as an input feature for the classifier. 155 devault, sagae and traum utterance time (ms) n l u f ! sc o re (e m p ty ) (e m p ty ) a ll e ld er e ld er d o y o u e ld er t o y o u d e ld er d o y o u a g re e e ld er d o y o u a g re e to e ld er d o y o u a g re e to m o v e th e e ld er d o y o u a g re e to m o v e th e e ld er d o y o u a g re e to m o v e th e cl in ic t o e ld er d o y o u a g re e to m o v e th e cl in ic d o w n e ld er d o y o u a g re e to m o v e th e cl in ic d o w n to w n e ld er d o y o u a g re e to m o v e th e cl in ic d o w n to w n 2 0 0 4 0 0 6 0 0 8 0 0 1 0 0 0 1 2 0 0 1 4 0 0 1 6 0 0 1 8 0 0 2 0 0 0 2 2 0 0 2 4 0 0 2 6 0 0 2 8 0 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 partial asr result figure 8: incremental interpretation of a user utterance. 156 incremental interpretation maxf model training features partial asr result f (k) k n entropy pmax nlu maxf (gold) (empty) 0.00 1 0 2.96 0.48 f82 false (empty) 0.00 2 0 2.96 0.48 f82 false all 0.00 3 1 0.82 0.76 f72 false elder 0.00 4 1 0.08 0.98 f39 false elder do you 0.83 5 3 1.50 0.40 f68 false elder to you d 0.50 6 3 1.31 0.75 f69 false elder do you agree 0.83 7 4 1.84 0.35 f68 false elder do you agree to 0.83 8 5 1.40 0.61 f68 false elder do you agree to move the 0.91 9 7 0.94 0.49 f10 true elder do you agree to move the 0.91 10 7 0.94 0.49 f10 true elder do you agree to move the clinic to 0.83 11 9 1.10 0.58 f68 false elder do you agree to move the clinic down 0.83 12 9 1.14 0.66 f68 false elder do you agree to move the clinic downtown 0.91 13 9 0.50 0.89 f10 true elder do you agree to move the clinic downtown 0.91 14 9 0.50 0.89 f10 true figure 9: features used to train the maxf model. our general aim, then, is to train a classifier, maxf, whose output predicts the value of maxf (gold) as a function of the input features. to create a data set for training and evaluating this classifier, we observed and recorded the values of these features for the 6068 partial asr results in a corpus of asr output for 449 actual user utterances.2 we chose to train a decision tree using weka’s j48 training algorithm (witten and frank 2005).3 to assess the trained model’s performance, we carried out a 10-fold cross-validation on our data set.4 we present our results in the next section. 5.3 results we will present results for a trained decision tree model that reflects a specific precision/recall tradeoff. in particular, given our aim to enable an agent to sometimes initiate overlapping speech, while minimizing the chance of making a wrong assumption about the user’s meaning, we selected a model with high precision at the expense of lower recall. various precision/recall tradeoffs are possible in this framework; the choice of a specific tradeoff is likely to be system and domaindependent and motivated by specific design goals. we evaluate our model using several features which are exemplified in figure 10. these include maxf (predicted), the trained maxf classifier’s output (true or false) for each partial; 2. this corpus was not part of the training data for mxnlu. 3. of course, other classification models could be used. 4. all the partial asr results for a given utterance were constrained to lie within the same fold, to avoid training and testing on the same utterance. 157 devault, sagae and traum maxf model evaluation features k f (k) ∆f (k) t (k) maxf (predicted) 1 0.00 -0.91 2.6 false 2 0.00 -0.91 2.4 false 3 0.00 -0.91 2.2 false 4 0.00 -0.91 2.0 false 5 0.83 -0.08 1.8 false 6 0.50 -0.41 1.6 false 7 0.83 -0.08 1.4 false 8 0.83 -0.08 1.2 false 9 (= kmaxf) 0.91 0.00 (=∆f (kmaxf)) 1.0 true 10 0.91 0.00 0.8 true 11 0.83 -0.08 0.6 false 12 0.83 -0.08 0.4 false 13 0.91 0.00 0.2 true 14 0.91 0.00 0.0 true figure 10: features used to evaluate the maxf model. kmaxf, the first partial number for which maxf (predicted) is true; ∆f (k) = f (k) − f (kfinal), the “loss” in f-score associated with interpreting partial k rather than the final partial kfinal for the utterance; t (k), the remaining length (in seconds) in the user utterance at each partial. we begin with a high level summary of the trained maxf model’s performance, before discussing more specific impacts of interest in the dialogue system. we found that our trained model predicts that maxf = true for at least one partial in 79.2% of the utterances in our corpus. for the remaining utterances, the trained model predicts maxf = false for all partials. the precision/recall/f-score of the trained maxf model are 0.88/0.52/0.65 respectively. the high precision means that 88% of the time that the model predicts that f-score is maximized at a specific partial, it really is. on the other hand, the lower recall means that only 52% of the time that f-score is in fact maximized at a given partial does the model predict that it is. for the 79.2% of utterances for which the trained model predicts maxf = true at some point, figure 11 shows the amount of time in seconds, t (kmaxf), that remains in the user utterance at the time partial kmaxf becomes available from the asr. the mean value is 1.6 seconds; as the figure shows, the time remaining varies from 0 to nearly 8 seconds per utterance. this represents a substantial amount of time that an agent could use strategically, for example by immediately initiating overlapping speech (perhaps in an attempt to improve communication efficiency), or by exploiting this time to plan an optimal response to the user’s utterance. however, it is also important to understand the cost associated with interpreting partial kmaxf rather than waiting to interpret the final asr result kfinal for the utterance. we therefore analyzed 158 incremental interpretation utterance time remaining (seconds) f re q u en cy 0 2 4 6 8 0 10 20 30 figure 11: distribution of t (kmaxf). the distribution in ∆f (kmaxf) = f (kmaxf)− f (kfinal). this value is at least 0.0 if mxnlu’s output for partial kmaxf is no worse than its output for kfinal (as intended). the distribution is given in figure 12. as the figure shows, 62.35% of the time (the median case), there is no difference in f-score associated with interpretingkmaxf rather thankfinal. 10.67% of the time, there is a loss of -1, which corresponds to a completely incorrect frame at kmaxf but a completely correct frame at kfinal. the converse also happens 2.52% of the time: mxnlu’s output frame is completely correct at the early partial but completely incorrect at the final partial. the remaining cases are mixed. while the median is no change in f-score, the mean case is a loss in f-score of -0.1484. this is the mean penalty in nlu performance that could be paid in exchange for the potential gain in communication efficiency suggested by figure 11. 6. completing user utterances to illustrate one use of the techniques described in the previous sections, we have implemented a prototype module that performs user utterance completion. this allows an agent to jump in during a user’s utterance, and say a completion of the utterance before it is finished. this type of completion is often encountered in human-human dialogue, e.g., skuplik (1999) as reported by poesio and rieser (2010) found 126 completions in a corpus of 3675 utterances. completions may be used, for example, for grounding or for bringing the other party’s turn to a conclusion. in this section, we present and discuss a prototype implementation that builds on the incremental understanding and 159 devault, sagae and traum ∆f (kmaxf) range percent of utterances -1 10.67% (−1, 0) 17.13% 0 62.35% (0, 1) 7.30% 1 2.52% mean(∆f (kmaxf)) -0.1484 median(∆f (kmaxf)) 0.0000 figure 12: the distribution in ∆f (kmaxf), the “loss” associated with interpreting partialkmaxf rather than kfinal. prediction models we have presented above, and which has allowed us to equip one of our virtual humans, doctor perez, with an ability to perform utterance completions. the decision to complete another speaker’s utterance is potentially a complex one. in our prototype implementation, we have incorporated the mxnlu and maxf classifiers into a relatively simple model of this decision-making, as a way to begin to explore the value and limitations of these classifiers in the decision to complete a user’s utterance. the process proceeds as follows. the first step is for the agent to recognize whether it has reached a moment of maximal understanding for the utterance. in our prototype, doctor perez never attempts to complete an utterance until maxf becomes true. this is a heuristic that means doctor perez will never try to complete an utterance if it thinks its understanding could be improved by waiting. (note that even if the classifier judges maxf to be true, it is not guaranteed that the utterance has been understood correctly. the maxf classifier may have made a mistake, and the utterance would in fact be understood better with additional user speech. or, it may be the case that the utterance will never be understood well, and the maxf classifier has detected this unfortunate state of affairs before the utterance concludes.) as discussed in section 5, maxf often (but not always) becomes true before the user has completed the utterance. nlu is performed on partial asr hypotheses as they become available, and maxf decides whether the agent’s understanding of the current partial hypothesis is likely to improve given more time. once maxf indicates that the agent’s understanding is already maximized, our prototype takes the current partial asr hypothesis, and attempts to generate text to complete it in a way that is fluent and agrees with the predicted meaning of the utterance the user has in mind. the generation of the surface text for completions takes advantage of the manual transcriptions in the corpus of utterances used to train the nlu module. for each frame that the agent understands, our training set contains a set of user utterances that correspond to the meaning in that frame. at the point where the agent is ready to formulate a completion, mxnlu has already predicted a frame for the user’s utterance (even though it is still incomplete). we then consider only the set of known utterances that correspond to that frame as possible sources of completions. as a simple distance 160 incremental interpretation metric, we compute the word error rate (wer) between the current partial hypothesis for the user’s utterance and a prefix of each of these known utterances. in our prototype, these prefixes have the same length as the current partial asr hypothesis. we then select the utterance whose prefix has the lowest wer against the current partial asr hypothesis. as a final step, we look in the prefix of our selected utterance for the last occurrence of the last word in the partial asr, and if such a word is found, we take the remainder of the utterance as the agent’s completion. considering only the set of utterances that correspond to the frame predicted by mxnlu makes it likely that the completion will have the appropriate meaning. since the completion is a suffix of a transcript of a previous user utterance, and this suffix follows the last word uttered by the user, it is likely to form a fluent completion of the user’s partial utterance. we applied this process to 449 user utterances.5 for 356 of these utterances (79.2%), maxf becomes true at some point during the utterance. of these 356, our prototype is able to generate a (non-empty) completion for 190 utterances (53.3% of the maxf utterances), and fails to generate a completion for 166 utterances (46.6% of the maxf utterances). thus, overall, our prototype is able both to identify a moment of maximal understanding and also to generate an utterance completion for 42.3% (190 of 449) of user utterances. we now discuss the completions that our prototype generates. we provide some representative examples of its completions in tables 2 and 3. these tables divide the generated completions into cases when the nlu output frame at kmaxf is perfectly correct vs. those in which it is either partially or completely incorrect. the case of perfectly correct nlu output at kmaxf, which is exemplified in table 2, occurs in 103 (54.2%) of the 190 cases for which our completion process is able to produce a completion of the user’s utterance. these cases reflect the ideal situation in which the maxf classifier allows the agent to detect that it has reached a moment of maximal understanding, and further, the nlu has in fact correctly understood the utterance. the alternative case, in which the nlu output at kmaxf is either partially or completely incorrect, is exemplified in table 3. this case includes 87 (45.8%) of the 190 utterances in which our implementation is able to produce a completion of the user’s utterance. in these cases, the generated completions exhibit a number of weaknesses, including incorrect predictions about the user’s meaning (for example, predicting a vehicle when the user was going to say a well for the village) as well as disfluent completions caused by mistakes in the asr prefix; see the examples marked full transcript, in which an incorrect asr prefix can affect the system’s completion. together, these examples highlight both strengths as well as limitations in our prototype implementation of utterance completions. table 2 includes a number of cases in which this system is able to complete the user’s utterance in a manner that seems fluent and consistent with the user’s communicative intention. on the other hand, together with the examples of table 3, in which maxf is true but the nlu output is not completely accurate, these results suggest that a more sophisticated decision-making process would be beneficial in deciding to complete a user’s utterance. in particular, maxf detects that the system’s understanding is unlikely to improve by waiting, but it does not guarantee that the system’s understanding is actually correct. (the nlu output might be incorrect, and likely to remain incorrect.) thus, it could be useful to train the system to identify situations in which it is sure that its understanding is in fact highly accurate, before risking attempting a com5. this was performed in such a way that the to-be-completed utterances were never included in the training process for either mxnlu or maxf. 161 devault, sagae and traum partial asr result at kmaxf predicted completion actual user completion f-score at kmaxf yes we’d like to move you out of here move your location 1 we need to move the clinic soon to move your clinic 1 we can provide you with supplies you need with supplies 1 we can have locals to move you locals help with the move 1 we can provide transportation to move the patient there (empty) 1 we should move this facility move the clinic right now 1 we will bring trucks then move it to the trucks and move them to the new safe clinic full transcript: we will bring trucks in move your patients and your supplies 1 it is not safe here we can’t protect you safe here 1 there are supplies where we are going there 1 table 2: some examples of generated completions when the predicted nlu frame is correct. partial asr result at kmaxf predicted completion actual user completion f-score at kmaxf we have to move to a safer place move the clinic because it is not safe here 0.33 we can provide a vehicle a well for the village 0.71 we have supplies available to move this clinic 0.38 so we are trying to help these people also will move the clinic downtown away from the insurgent activity 0.1 if we can give you transportation full transcript: we can not protect you here 0.4 yes i see agree 0 doctor would the market isn’t a safe place for a clinic v full transcript: doctor what do you think 0 we can i can help you with supplies full transcript: we cannot protect you 0.3 would like that you move your clinic full transcript: we’d like to talk to you both about the clinic 0 table 3: some examples of generated completions when the predicted nlu frame is incorrect, or partially incorrect. 162 incremental interpretation figure 13: a graphical user interface for interactive utterance completion. pletion of the user’s utterance. we are pursuing this extension, and possibilities for combining the maxf judgment with a decision about nlu correctness, in ongoing work. finally, it is worth emphasizing that even when an agent has the ability to generate a completion, clearly a number of broader strategic considerations could be relevant in deciding whether to do so. determining a broader dialogue policy that results in natural behavior with respect to the frequency of completions for different types of agents is a topic under current investigation. 6.1 toward completion of live user utterances the above analysis of our prototype utterance completion capability is based on off-line, batch mode completion of pre-captured user utterances. in more recent work, we have developed a run-time demonstration that allows the utterance completion capability to be explored interactively (sagae et al. 2010). to help visualize the agent’s decision to complete an utterance, we have developed a simple gui, depicted in figure 13, that highlights important state changes within our mxnlu and maxf incremental nlu models, and also provides some control over the timing of utterance completions. 163 devault, sagae and traum the figure shows an example in which the character judges that the user’s partial utterance we really have to move could be completed by your clinic.6 the agent’s incremental reasoning each 200ms is depicted in the gui, in chronological order, from the top of the gui to the bottom. first, the gui shows the evolution of incremental asr results during the user’s speech; see the lines marked ‘asr’, with a white background. the gui also includes an entry each time there is a change in the semantic frame predicted by mxnlu; see the lines marked ‘nlu’, highlighted in blue. for compactness, the gui presents an english language gloss of the predicted semantic frame, such as the captain wants to move the clinic, rather than the full avm structure. the gloss utterance not understood is used to represent the “garbage” frame. finally, the gui also indicates those moments when the maxf classifier judges that maxf is true. in this example, maxf becomes true when the asr output is we really have to. at this point, an attempt could be made to generate an utterance completion. however, it is very likely that this would result in the character barging in and interrupting the user’s speech, or talking simultaneously with the user. while this aggressive behavior is interesting, we have found it undesirable in most cases. as we have not yet developed a refined model of the decision to interrupt ongoing user speech with utterance completions, we have instead provided an interactive control which requires a threshold amount of silence, after maxf has become true, before an utterance completion will be initiated by the character. this control, seen at the top of the figure and marked “completion delay (milliseconds)”, is set here to 600 milliseconds. in this example, the user does indeed pause for 600 milliseconds after saying we really have to move. (this can be seen by the repeated partial asr results in the figure.) after this threshold period has elapsed, the character initiates its completion and utters your clinic. 7. discussion and future work in this section, we present some discussion and context for the techniques we have presented above. we address issues of generality, alternative approaches, integration into dialogue system design, limitations, and planned future work. 7.1 generality and training requirements because we use data-driven techniques, which draw on domain-specific utterance data to train a domain-specific nlu capability, we can expect that our techniques will transfer to some extent into new dialogue domains where similar utterance data is available or can be acquired. unfortunately, due to the numerous and varied empirical details that may affect nlu performance, it is difficult to specify in advance, with precision, the amount of data that is required in order for acceptable performance to be achieved in a new domain with our techniques. in our saso-en domain, we achieve a relatively high f-score (0.76) for nlu, given the complexity of the task and the high word error rate in asr (0.54). this is accomplished using a training set of about 3,500 annotated user utterances. however, even in situations where collecting and annotating a training corpus of this size may be impractical, it is still possible for the multiclass classification approach to produce high quality nlu results. when mxnlu is trained using a re6. note that the text your clinic is drawn from previous user utterances. for our prototype completion capability, we have not taken the step of replacing user pronouns (your clinic) with corresponding character pronouns (my clinic). 164 incremental interpretation duced training set of 1,000 utterances, its f-score is about 0.7, with small gains resulting from the addition of more training material. of course, the performance of mxnlu or a similar module based on multiclass classification depends on several factors that are specific to each dialogue system. these include asr performance, the size of the domain-specific vocabulary that users will employ, the number of different concepts and semantic frames the system is expected to understand, and the target user population. in general, it is important that there be reliable features in the asr output, similar to the features displayed in table 1, that the nlu can (learn to) use to identify the correct interpretation. the extent to which such features exist for a domain will depend on many empirical details. see also heintze et al. (2010) for a useful comparative analysis of incremental nlu performance in different domains. 7.2 alternative approaches to incremental understanding the approach to incremental understanding we have presented in this paper is fundamentally predictive. as the user is speaking, the nlu module attempts at each moment to predict the complete semantic frame that will be the correct analysis of their entire utterance. alternative approaches to incremental nlu are possible. for example, the incremental nlu might not be predictive, and instead try to identify only those frame elements that have been directly expressed by the user in their speech so far. or, as a variant on this approach, the incremental nlu could also try to identify which specific words or expressions in the user’s partial utterance have expressed which frame elements (heintze et al. 2010). these alternative approaches solve different problems than our predictive model solves, and we view them as providing complementary information to incremental dialogue systems. as one example of how this information might be combined, in a situation where a confident prediction could not be made of the user’s complete meaning, a system could nevertheless begin to reason about and respond to those aspects of the user’s meaning about which it is confident. 7.3 system architecture to date, implemented dialogue systems have generally attempted to understand user utterances only after the conclusion of the user’s speech. the kinds of modular architectures that have been used to implement these systems have often assumed that a sequential pipeline of processing steps follows the conclusion of the user’s utterance.7 it may not be straightforward to substitute incremental modules, such as our mxnlu and maxf classifiers, into these pipelines, unless the other modules are also modified to respond appropriately to the incremental processing results; see schlangen and skantze (2009) for an approach to this general problem. for example, a dialogue manager might be modified to delay the initiation of any response to a user utterance until maxf is true. such a policy could prove useful in some systems, but it is undoubtedly a simplistic approach to a complex problem. our prototype implementation of utterance completion, which initiates completion only once maxf is true, was designed to help explore the potential value of these incremental processing models. however, there remain a number of questions about how system architectures should be adjusted to take maximum advantage of them. 7. for example: user speech→ asr→ nlu→ dm→ nlg→ tts 165 devault, sagae and traum 7.4 limitations of the utterance completion method the utterance completion capability presented in section 6 has some limitations. because each completion text used by the character is extracted verbatim from the set of existing utterance transcripts in our nlu training corpus, the set of possible completion texts is finite. although the set of possible completion texts is not so small (as there are thousands of utterance transcripts to draw from), it could nevertheless be argued that our prototype supports little or no linguistic creativity by the character as it performs the completions. a more general capability would productively generate the completion text, perhaps using a semantic analysis of the user’s partial utterance (identifying the subset of frame elements the user has directly expressed in their speech) and a productive grammar (to allow the character to express linguistically the predicted frame elements which the user has not yet expressed). it would be interesting to explore the development of such a refined completion capability in future work. we should note however that our current prototype does support linguistic creativity on the user’s side, as our mxnlu and maxf models, which underlie the completion capability, do accept arbitrary user speech. indeed, the analysis and results for our completion capability, presented in section 6, were confined to completions of new user utterances which were not drawn from the training set of either the mxnlu or maxf models. finally, as mentioned in section 6, the decision to complete a user’s utterance is in reality a complex turn-taking decision, which should be based on a range of additional factors not taken into account in our prototype. we are just beginning to explore these issues in our ongoing work. 7.5 backchannels we are currently exploring the use of incremental processing for providing backchannels, using both verbal and non-verbal channels. the incremental nlu and speech results are used by the dialogue manager, which attempts to perform reference resolution against the internal domain model, and extracts information about whether or not the agent understands the utterance, agrees with it, and the agent’s emotions (gratch and marsella (2004)) toward the described state or action, or the participants. this information, along with the words, semantic frame, and maxf information are provided to the verbal and non-verbal generators (devault et al. (2008), lee and marsella (2006)), which decide whether to produce a backchannel and the form (e.g. repeating words, saying “uh huh”, “yes” or “no”, head nods or shakes, and facial expressions). 7.6 evaluation in this paper we have presented several kinds of evaluation for our techniques. in section 3.4, we analyzed the performance of mxnlu on complete asr output, in terms of its recall and precision of semantic frame elements. in section 4, we presented a refined performance analysis for several alternative approaches to training mxnlu to predict complete semantic frames from only partial asr output. in section 5.3, we presented an analysis of our maxf classifier’s performance, in terms of precision and recall of points of maximal understanding, as well as the potential time saved and penalty in f-score. finally, in section 6, we presented statistics on the frequency with which our completion prototype is able to generate completions for new user utterances, as well as the frequency with which mxnlu’s prediction at the moment of completion was correct vs. incorrect. 166 incremental interpretation what we have not evaluated to date, however, is what progress we have made toward the higherlevel goal of improving the responsiveness and naturalness of interacting with virtual human characters. while this remains our long-term motivation, there remains more work to be done before it will be possible to demonstrate the usefulness in live dialogue sessions of an utterance completion capability, or other overlap behaviors that draw on our techniques. in particular, we expect that this will require that virtual humans engage in more detailed decision-making about when and how to initiate an overlapping response based on incremental processing results. beyond utterance completions and back-channels, our virtual humans have a number of additional options for overlapping responses, including eye gaze, facial displays, head nods or head shakes, and other non-verbal behavior such as gesture. as we explore these options, we also expect that more detailed consideration of the fine-grained timing of overlapping responses than we have presented here will likely prove essential to achieving natural behavior. here we have presented several fundamental techniques for the incremental interpretation and prediction of utterance meaning, which we believe are likely to prove valuable in implementing this more complex decision-making. we are eagerly exploring the issues involved in translating these techniques into live capabilities in our ongoing work. 8. conclusion we have presented a framework for interpretation of partial asr hypotheses of user utterances, and high-precision identification of points within user utterances where the system has reached maximal understanding of the intended meaning. our initial implementation of an utterance completion ability for a virtual human serves to illustrate the capabilities of this framework, but only scratches the surface of the new range of dialogue behaviors and strategies it allows. acknowledgments the work described here has been sponsored by the u.s. army research, development, and engineering command (rdecom). statements and opinions expressed do not necessarily reflect the position or the policy of the united states government, and no official endorsement should be inferred. we would also like to thank anton leuski for facilitating the use of incremental speech results, and the ict dialogue group and david schlangen for helpful discussions, and our anonymous reviewers for helpful suggestions on the presentation of the paper. references gregory aist, james allen, ellen campana, carlos gomez gallo, scott stoness, mary swift, and michael k. tanenhaus. incremental dialogue system faster than and preferred to its nonincremental counterpart. in proc. of the 29th annual conference of the cognitive science society, 2007. jens allwood. an activity based approach to pragmatics. technical report (gptl) 75, gothenburg papers in theoretical linguistics, university of göteborg, 1995. jens allwood, joakim nivre, and elisabeth ahlsen. on the semantics and pragmatics of linguistic feedback. journal of semantics, 9, 1992. 167 devault, sagae and traum adam l. berger, stephen d. della pietra, and vincent j. d. della pietra. a maximum entropy approach to natural language processing. computational linguistics, 22(1):39–71, 1996. justine cassell, joseph sullivan, scott prevost, and elizabeth churchill, editors. embodied conversational agents. mit press, cambridge, ma, 2000. herbert h. clark. arenas of language use. university of chicago press, 1992. herbert h. clark and edward f. schaefer. collaborating on contributions to conversation. language and cognitive processes, 2:1–23, 1987. herbert h. clark and deanna wilkes-gibbs. referring as a collaborative process. cognition, 22: 1–39, 1986. also appears as chapter 4 in clark (1992). david devault, david traum, and ron artstein. making grammar-based generation easier to deploy in dialogue systems. in fifth inlg conference, 2008. david devault, kenji sagae, and david traum. can i finish? learning when to respond to incremental interpretation results in interactive dialogue. in the 10th annual sigdial meeting on discourse and dialogue (sigdial 2009), 2009. kotaro funakoshi, mikio nakano, kazuki kobayashi, takanori komatsu, and seiji yamada. nonhumanlike spoken dialogue: a design perspective. in proceedings of the sigdial 2010 conference, pages 176–184, tokyo, japan, september 2010. charles goodwin. the interactive construction of a sentence in natural conversation. in g. psathas, editor, everyday language: studies in ethnomethodology, pages 97–121. ervington press, new york, 1979. jonathan gratch and stacy marsella. a domain-independent framework for modeling emotion. journal of cognitive systems research, 2004. brian hansen, david novick, and stephen sutton. prevention and repair of breakdowns in a simple task domain. in proceedings of the aaai-96 workshop on detecting, repairing, and preventing human-machine miscommunication, pages 5–12, 1996. arno hartholt, thomas russ, david traum, eduard hovy, and susan robinson. a common ground for virtual humans: using an ontology in a natural language oriented virtual human architecture. in european language resources association (elra), editor, proc. of the sixth international language resources and evaluation (lrec’08), marrakech, morocco, may 2008. silvan heintze, timo baumann, and david schlangen. comparing local and sequential models for statistical incremental natural language understanding. in the 11th annual meeting of the special interest group in discourse and dialogue (sigdial 2010), 2010. david huggins-daines, mohit kumar, arthur chan, alan w. black, mosur ravishankar, and alex i. rudnicky. pocketsphinx: a free, real-time continuous speech recognition system for hand-held devices. in proceedings of icassp, 2006. 168 incremental interpretation geert-jan m. kruijff, pierre lison, trevor benjamin, henrik jacobsson, and nick hawes. incremental, multi-level processing for comprehending situated dialogue in human-robot interaction. in language and robots: proc. from the symposium (langro’2007). university of aveiro, 12 2007. jina lee and stacy marsella. nonverbal behavior generator for embodied conversational agents. in jonathan gratch, michael young, ruth aylett, daniel ballin, and patrick olivier, editors, iva, pages 243–255. springer, 2006. isbn 3-540-37593-7. yukiko i. nakano, gabe reinstein, tom stocky, and justine cassell. towards a model of face-toface grounding. in acl, pages 553–561, 2003. bryan pellom. sonic: the university of colorado continuous speech recognizer. in university of colorado, tech report #tr-cslr-2001-01, boulder, colorado, 2001. massimo poesio and hannes rieser. completions, coordination, and alignment in dialogue. dialogue and discourse, 1:1–89, 2010. url http://elanguage.net/journals/ index.php/dad/article/view/91. jeff rickel and w. lewis johnson. virtual humans for team training in virtual reality. in proceedings of the ninth international conference on artificial intelligence in education, pages 578–585. ios press, 1999. harvey sacks, emanuel a. schegloff, and gail jefferson. a simplest systematics for the organization of turn-taking for conversation. language, 50:696–735, 1974. m. david sadek. dialogue acts are rational plans. in proceedings of the esca/etr workshop on multi-modal dialogue, 1991. kenji sagae, gwen christian, david devault, and david r. traum. towards natural language understanding of partial speech recognition results in dialogue systems. in short paper proceedings of naacl hlt, 2009. kenji sagae, david devault, and david r. traum. interpretation of partial utterances in virtual human dialogue systems. in the 11th annual conference of the north american chapter of the association for computational linguistics (naacl-hlt 2010 demonstration), 2010. david schlangen and gabriel skantze. a general, abstract model of incremental dialogue processing. in proc. of the 12th conference of the european chapter of the acl, 2009. david schlangen, timo baumann, and michaela atterer. incremental reference resolution: the task, metrics for evaluation, and a bayesian filtering model that is sensitive to disfluencies. in the 10th annual sigdial meeting on discourse and dialogue (sigdial 2009), 2009. gabriel skantze and david schlangen. incremental dialogue processing in a micro-domain. in proceedings of eacl 2009, pages 745–753, 2009. kristina skuplik. satzkooperationen. definition und empirische untersuchung. technical report 1999/03 of sfb 360, bielefeld, 1999. 169 devault, sagae and traum david traum. semantics and pragmatics of questions and answers for dialogue agents. in proc. of the international workshop on computational semantics, pages 380–394, january 2003. david traum, stacy marsella, jonathan gratch, jina lee, and arno hartholt. multi-party, multiissue, multi-strategy negotiation for multi-modal virtual agents. in proc. of intelligent virtual agents conference iva-2008, 2008. ian h. witten and eibe frank. data mining: practical machine learning tools and techniques. morgan kaufmann, 2005. victor h. yngve. on getting a word in edgewise. in papers from the sixth regional meeting, pages 567–78. chicago linguistic society, 1970. 170 journal of machine learning research-microsoft word template dialogue & discourse 11(1) 62-88 doi: 10.587/dad.2020.103 ©2020 yipu wei, jacqueline evers-vermeul and ted j.m. sanders this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). the use of perspective markers and connectives in expressing subjectivity: evidence from collocational analyses yipu wei weiyipu@pku.edu.cn school of chinese as a second language, peking university yiheyuan road 5, 100871, beijing, china jacqueline evers-vermeul j.evers@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands ted j.m. sanders t.j.m.sanders@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands editor: manfred stede submitted 05/2019; accepted 03/2020; published online 03/2020 abstract this study explores how subjectivity is expressed in coherence relations, by means of a distinctive collocational analysis on two chinese causal connectives: the specific subjective kejian ‘so’, used in subjective argument-claim relations, and the underspecified suoyi ‘so’, which can be used in both subjective argument-claim and objective cause-consequence relations. on the basis of both horn’s pragmatic relation and quality principles and the uniform information density theory, we hypothesized that the presence of other linguistic elements expressing subjectivity in a discourse segment should be related to the degree of subjectivity encoded by the connective. in line with this hypothesis, the association scores showed that suoyi is more frequently combined with perspective markers expressing epistemic stance: cognition verbs and modal verbs. kejian, which already expresses epistemic stance, co-occurred more often with perspective markers related to attitudinal stance, such as markers of expectedness and importance. the paper also pays attention to similarities and differences in collocation patterns across contexts and genres. keywords: collocational analysis, connectives, perspective marking, stance, subjectivity 1 introduction in everyday communication, speakers and writers often express their conclusions and feelings. for instance, instead of merely reporting objective causal relations between events in the real world, as in (1a), they frequently utter subjective relations, which involve someone’s reasoning (langacker 1990, pander maat & sanders 2000, verhagen 2005), as illustrated in (1b). subjective relations are not observable in the real world; one needs to take into account another person’s (e.g., the speaker’s or another agent’s) perspective (sanders et al. 2009, 2012) to process the reasoning, and thus one needs to track the source of information. in other words, subjective relations concern the degree of the use of perspective markers and connectives in expressing subjectivity 63 involvement of a locutionary agent or a subject of consciousness (finegan 1995, lyons 1977, sanders et al. 2009). (1) a. this restaurant is decorated with several art works of mondriaan, so it attracts lots of fans of modern art. b. this restaurant is decorated with several art works of mondriaan, so its owner must be a fan of modern art. in order to communicate in a coherent way, speakers choose words to express the relations between consecutive discourse segments (sanders et al. 1993: 94, cf. also sanders & spooren 2007, schilperoord & verhagen 1998). for instance, they can use connectives such as so and therefore to provide the reader with information on the type of coherence relation to be established, in this case a causal one (britton 1994, graesser & mcnamara 2011, mak & sanders 2010, van silfhout et al. 2014, 2015). such information facilitates the reading process. it triggers faster processing of information immediately following the connective (cain & nash 2011, cozijn et al. 2011, sanders & noordman 2000, van silfhout et al. 2014, 2015) compared to the processing of that same information in unmarked relations. as examples (1a) and (1b) illustrate, english so can be used in objective and subjective causal relations. it only marks the causal nature of the relation, and does not indicate the degree of subjectivity of the relation. however, certain connectives in other languages do code information about subjectivity. for example, some connectives are only used for objective relations, such as dutch daardoor ‘as a result’ and chinese yin’er ‘as a result’, as is illustrated in the dutch (2a) respectively chinese (3a) translations of (1a). by contrast, the dutch connectives want ‘because’ and dus ‘so’ (degand & pander maat 2003, sanders & spooren 2015, spooren et al. 2010, stukker & sanders 2008, verhagen 2005), and mandarin chinese kejian ‘so’ prototypically express subjective coherence relations (li et al. 2013). this is illustrated by the dutch (2b) respectively chinese (3b) counterparts of the subjective relation in (1b). just like english so in example (1a) and (1b), some connectives in other languages leave the subjectivity information underspecified, i.e. they can be used for both subjective and objective relations (e.g. chinese suoyi ‘so’ in example (3a) and (3b)). (2) dutch a. dit restaurant is versierd met diverse kunstwerken van mondriaan, daardoor trekt het veel fans van moderne kunst. b. dit restaurant is versierd met diverse kunstwerken van mondriaan, dus de eigenaar moet wel een fan zijn van moderne kunst. (3) chinese a. zhe jia canguan zhuangshi zhe hao ji fu mengteli’an de huazuo, yin’er/ suoyi ta xiyin le henduo xiandai yishu mi. this cl restaurant decorate asp(ipfv) cl mondrian mod painting, as a result/ so 3sg attract asp(pfv) many modern art fan. b. zhe jia canguan zhuangshi zhe hao ji fu mengteli’an de huazuo, kejian/ suoyi ta de zhuren keneng shi yi ge xiandai yishu mi. this cl restaurant decorate asp(ipfv) cl mondrian mod painting, in conclusion/ so 3sg mod owner probably be a cl modern art fan. the degree of subjectivity expressed by connectives is found to affect the processing of coherence relations. for instance, the dutch subjective connective want ‘because’ leads to longer processing times directly after the connective compared to the dutch objective connective omdat ‘because’ (canestrelli et al. 2013). such processing effects can be attributed to the difficulty of interpreting wei, evers-vermeul and sanders 64 subjectivity: the reader needs to track the source of information to interpret subjectivity. specific subjective connectives such as want ‘because’ instruct the reader at an early stage that there is a coherence relation, and that the relation is subjective, before the entire sentence is processed. in terms of the information density, subjective connectives encode more information compared to underspecified connectives. as the choice of connectives in examples (1) to (3) and the accompanying processing results illustrate, speakers and writers continuously have to decide how informative they should be in order to provide sufficient cues for others to comprehend them. at the same time, they should also avoid being too wordy. this tension has been systematically described by horn’s framework for pragmatic inference: his q (quality) principle describes the need to ‘make your contribution sufficient’, the r (relation) principle describes the need to ‘make your contribution necessary’ (horn 1984: 13). according to horn (1984), speakers should find a balance between the speakerbased economy (saving the speaker’s production efforts) and the hearer-based economy (saving the hearer’s processing efforts). a highly similar point has been made by the uniform information density theory (uid), which is about the speakers’ strategy of choosing between alternative linguistic forms at several levels of linguistic representations: phonetic, syntactic, pragmatic, etc. (frank & jaeger 2008, jaeger 2010, levy & jaeger 2007). the uid suggests that speakers modulate their word choice according to the amount of information in the utterance: full linguistic forms are more often used at the point where the content conveyed by the form is unexpected in its context, i.e. the point with a low probability and a high information density (for details, see frank & jaeger 2008). for instance, connectives can be omitted if the information they convey is highly predictable given other linguistic cues in the context (asr & demberg 2015). through such modulation of word choices, the density of information of the utterance is kept at a uniform level – a roughly equal amount of information at each unit of the sentence (levy & jaeger 2007). the uid theory echoes horn’s pragmatic theory in the sense that both theories predict a modulated process of word selection to optimize communication. in terms of discourse relations and connectives, these theoretical discussions raise the question as to which information is exactly conveyed by connectives, and how that information may become predictable given other cues in the context. hence, it is worthwhile to explore, as is done in the current paper, which linguistic markers also provide information on the degree of subjectivity of a relation, and would thereby allow for a division of labor between connectives and segment-internal elements (see hoek 2018; hoek et al., 2018). if other markers already indicate the degree of subjectivity, this will reduce the need of information on subjectivity to be expressed at the connective. this seems to be the case for expressions such as probably, surprisingly and according to peter, which are addressed as markers of stance (biber et al. 1999, conrad & biber 2000), evaluation markers (bednarek 2006, 2009, thompson & hunston 2000), or appraisals (eggins & slade 1997, martin 2000). conrad and biber (2000) suggest three sub-types of stance markers (see bednarek 2006, bednarek 2009 and thompson & hunston 2000 for similar classifications): i. epistemic stance, which indicates how certain the speaker or writer is, or where the information comes from (e.g. probably, according to the president). ii. attitudinal stance, which indicates feelings or judgements about what is said or written (e.g. surprisingly, unfortunately). iii. style stance, which indicates how something is said or written (e.g. honestly, briefly.) (conrad & biber 2000: 57) stance markers introduce the viewpoint of the speaker or other agents, and hence can be termed as perspective markers (sanders & redeker 1996). perspective markers expressing epistemic stance show overlap with specific subjective connectives. both indicate subjective reasoning, either from the speaker or from a character. canestrelli et al. (2013) and traxler et al. (1997) found that the the use of perspective markers and connectives in expressing subjectivity 65 processing effects of connectives are influenced by epistemic stance markers: by adding volgens peter ‘according to peter’ to the first clause connected in a subjective relation, as in example (2c), the extra processing time associated with the subjective connective want ‘because’ disappears. (2) c. volgens peter is de eigenaar van dit restaurant een fan van moderne kunst, want het restaurant is versierd met diverse kunstwerken van mondriaan. according to peter the owner of this restaurant is a fan of modern art, because the restaurant is decorated with several art works of mondrian. in terms of horn’s pragmatic theory, the reader/hearer has obtained sufficient information about the degree of subjectivity by the introduction of epistemic perspective markers. upon encountering the subjective connective the reader/hearer does not have to establish an entirely new subjective mental representation, but rather only has to make a link to an already established mental representation introduced by the perspective marker in the first clause. in other words, epistemic stance markers in the first clause make it clear that the first clause is a claim and thereby create the expectation that the next clause will be an argument for this claim. the empirical findings of canestrelli et al. (2013) and traxler et al. (1997) suggest an overlap between specific subjective connectives and perspective markers in their function of instructing readers on the degree of subjectivity of the relation. the question is whether this holds true for perspective markers in general, including all types of stance markers, or only pertains to markers of epistemic stance. epistemic stance markers explicate the dimension of reliability/certainty and evidentiality, which directly introduces a source of information. however, attitudinal stance markers and style stance markers introduce a source in an indirect way: by indicating attitudes, feelings and styles of writing/speaking that can be attributed to a source. although all three types of stance markers presuppose a source of information, they differ in the way in which this source of information is involved. how these perspective markers overlap with connectives marking different degrees of subjectivity may shed light on the relation between subjectivity and perspective marking. in this paper, we investigate this issue in natural language data. starting from the assumption that language users will tend to avoid a doubling of information in terms of marking subjectivity in discourse relations, we may expect authors/speakers to observe some pragmatic strategies (e.g. apply horn’s r principle or try to produce an information flow with a uniform information density) to achieve a successful communication (both sufficient and necessary). avoiding repetition of information in the same dimension fits the r principle as well as the uid. therefore, in natural language data we may expect connectives marking different degrees of subjectivity to vary in their co-occurrence patterns with perspective markers. in corpus linguistics, the method of collocational analysis (evert 2008; gries & stefanowitsch 2004) provides insightful information on the context of given linguistic elements. it measures the association strengths between words or expressions, and produces a list of important collocates in attraction or repulsion with a target word. collocational analysis can advance our knowledge about the properties of a connective on the basis of its contextual features. we therefore conducted a corpus-based study using collocational analyses to examine the use of connectives and perspective markers in discourse, aiming to answer the following research questions: 1) do connectives of different subjectivity degrees differ in their types of collocates? 2) more specifically, do connectives differ in the types of perspective markers they co-occur with? we focused on two chinese causal connectives for which we could derive hypotheses from the literature. kejian ‘so’ is mostly used in the epistemic domain (li et al., 2013), indicating that the causal reasoning arises from someone’s mind; it encodes the epistemic stance apart from its wei, evers-vermeul and sanders 66 discourse function of causally connecting two segments. such subjectivity information is underspecified with the generic connective suoyi ‘so’, which can be used in both objective and subjective relations (li et al. 2013). on the basis of horn’s theory of speaker economy, kejian can be expected to co-occur less with perspective markers of the epistemic stance than suoyi. since neither the specific subjective kejian nor the generic connective suoyi encode attitudinal or style stance, no differences in collocation tendencies are expected between the connectives for the other two types of perspective markers. 2 method we conducted a series of distinctive collocates analyses on the two chinese causal connectives suoyi ‘so’ and kejian ‘so’, with the aim to investigate the contextual features of the two connectives. regular collocational analyses allow researchers to calculate association strengths between target words and their collocates. distinctive collocates analyses (church et al., 1991) are a specific type of collocational analysis: they allow for a direct comparison of the contexts of two semantically similar words (a word pair), identifying collocates that prefer to appear in the context of one word over the other word from the pair. with this type of analysis, words with high association scores are not associated with the target word in a general sense, but only if they are attracted more to this target word than to a reference context (i.e., in this study the alternative connective). this type of analyses has become especially popular for lexical alternatives in specific constructions (i.e., distinctive collexeme analysis or distinctive collostructional analysis, see gries & stefanowitsch 2004; stefanowitsch & gries 2003). in the current study, we use this method to identify words that tend to ‘sit’ in the context of suoyi more often than in the context of kejian and vice versa, paying special attention to linguistic elements expressing subjectivity. 2.1 sample of texts we used a balanced modern chinese corpus: the ccl corpus (zhan et al. 2003), which covers a variety of written texts: fiction, newspapers, conferences, translated literature, blogs, etc. the total size of the ccl corpus is 581,794,456 characters. we only investigated actual texts, which lead to the exclusion of dictionaries, and to make sure all the texts were homogeneous in terms of mode (written), we excluded the sources of oral texts (spoken), and tv (written to be spoken), etc. from the remainder of the corpus, we selected texts from three types of genres: narrative genres on the one hand, and informative and argumentative genres on the other. narrative genres included literature, drama, biographies and fiction magazines; informative and argumentative genres included newspapers, legal documents, academic works of natural science and social sciences, governmental reports and other texts labeled as practical writing. the argumentative and informative texts were collapsed as the ‘non-narrative genre’, because of the low number of argumentative texts available in the ccl corpus. from the afore-mentioned parts of ccl, we then generated two raw datasets: text files containing all the sentences with the words suoyi or kejian, with a search scope of 200 characters to the left and 200 to the right. this scope was much wider than the length of a sentence so that we would have enough contexts for the analysis on the intended discourse unit. in line with the parameters of collocation (gries 2013), we first decided to investigate words as the linguistic units of collocates. because natural chinese texts do not have spaces between words, we used the chinese word segmentation tool nlpir-ictclas (zhang et al., 2003; tag: ict_pos_map_second) to separate the word boundaries of characters in the text. in this segmentation system, white spaces were added between words, and words were tagged based on their semantic types. meanwhile, punctuations such as commas, full stops, parentheses, colons were also marked with tags. the word segmentation tool thereby generated segmented and annotated texts for later analysis. the use of perspective markers and connectives in expressing subjectivity 67 in terms of the distance between collocates, a collocate did not need to be directly adjacent to the connective. any words appearing within one clause before or one clause after the connectives were considered collocates. instead of adapting an arbitrary number of words as the context, we set the context of the target word in such a way that it was meaningful at the discourse level: discourse clauses were taken as the units for analysis. 2.2 sample of connective fragments from the two segmented datasets of all sentences containing suoyi or kejian, we compiled a sample of connective fragments. this step was necessary, because suoyi does not only occur as a connective, but can also be used in an inversion construction zhisuoyi ‘why there is a consequence of’. for the word kejian, we can observe a clear grammaticalization process in progress (liu & yao 2011; q. zhang 2012). there are cases in which kejian is used as a verb, sometimes resulting in modified constructions such as qingxi kejian ‘clearly can see’, and there are cases where kejian is clearly a connective, or where the use of kejian is ambiguous. in order to exclude the clear verbal cases of kejian and all the inversion constructions zhisuoyi, we conducted a two-step screening process in our sampling. first, we restricted the sample of target items to cases preceded by a punctuation marker (namely comma, full stop, question mark, exclamation mark, semicolon, or ellipsis) in the software antconc_3.4.4.0 (anthony 2016). this screening process filtered out verbal uses of kejian such as qingxi kejian ‘clearly can see’, as well as cases of kejian which are preceded by prepositional phrases such as youci kejian ‘from this can see’. after the rough automatic screening process, 67,147 sentences with suoyi and 3,902 sentences with kejian were included for further analyses. we then manually checked the remaining sentences marked by kejian, in order to exclude all other verbal instances of kejian. the verbal status of kejian could easily be derived from the absence of the main verb in the clause headed by kejian. for example, in (4), interpreting kejian as a connective with the meaning ‘so’ would only leave a noun phrase as the remainder of the second clause: the status of german cars in the minds of chinese. by contrast, interpreting kejian as a verb ‘can see’, results in a grammatical clause, because in chinese, the subject can be dropped. hence, only full sentences such as (5) were included in the analyses of the connective use of kejian. (4) deguo chan de dazhong, aodi and benchi zhanyou hen da de bili, kejian deguo chan de qiche zai zhongguoren xinmuzhong de diwei. germany produce mod volkswagen, audi and benz occupy very big mod proportion, kejian ‘from this can see’/*kejian ‘so’ germany produce mod car in chinese mind mod status. the german products volkswagen, audi and benz take a big proportion (of chinese market), from this we can see/*so the status of german cars in the mind of chinese people. (5) yi ge neng zhide yi tou niu de jiaqian, kejian nashihou shiliu zai woguo haishi xihan wu. one cl can worth one cl cow mod price, kejian ‘so’ that-time pomegranate in ourcountry still-is rare thing. one (pomegranate) was worth the price of a cow, so pomegranate was still very rare in our country at that time (in ancient china). all in all, the automatic and manual screening process excluded 20,096 cases of suoyi and 10,900 cases of kejian. table 1 shows the resulting distribution of suoyi and kejian in the narrative and non-narrative texts in the sample. wei, evers-vermeul and sanders 68 narrative texts non-narrative texts connectives retrieved from ccl used for analysis (%) retrieved from ccl used for analysis (%) suoyi 34,641 29,077 (83.94%) 52,445 37,913 (72.29%) kejian 2,494 752 (30.15%) 11,688 2,530 (21.65%) total 37,135 29,829 (80.33%) 64,133 40,443 (63.06%) table 1. distributions (and percentage of actually used cases) of suoyi and kejian in two genres 2.3 three sets of distinctive collocates analyses the actual collocate analyses were conducted using the software r (r core team 2015) with the r package mclm_0.1 (speelman 2018). the method of distinctive collocates analysis was applied three times. we first applied it to a context of one clause before and one clause after the connective, irrespective of genre. to obtain a proper context containing exactly one clause before and one clause after the connective, we automatically searched the closest punctuation markers (including comma, full stop, question mark, exclamation mark, semicolon, and ellipsis) around the target (a connective preceded by a punctuation marker). with this first analysis, we obtained a general picture of the words in collocation with one connective compared to the other. second, we explored the collocates of the two connectives in their preceding context and following context separately, so that contextual features could be located more precisely. however, the distinctive collocates of suoyi versus kejian may be different depending on the genre they appear in, because the narrative genre is supposed to be more descriptive (e.g., describing events and actions), while the non-narrative genre is expected to be more argumentative. therefore, in the third analysis, we took genre into account, distinguishing the collocational patterns in the narrative genre on the one hand, and in the informative and argumentative genres on the other. the attraction and repulsion strength between a given word and the target connectives are measured by association scores (evert 2008, gries 2013), which are calculated on the basis of observed frequencies (o11, o12, o21, o22) and expected frequencies (e11, e12, e21, e22) in a contingency table (table 2). for each word (target word) in the corpus that appeared at least once in the context of kejian or suoyi, the mclm r package computes o11 (i.e. target word instances in the target context), o12 (non-target words in the target context), o21 (target word instances in nontarget contexts), and o22 values (occurrence of non-target words in non-target contexts), as well as the corresponding expected frequencies. observed frequencies expected frequencies target-word non-target word totals target-word non-target word suoyi-context o11 o12 r1 e11=r1c1/n e12=r1c2/n kejian-context o21 o22 r2 e21=r2c1/n e22=r2c2/n totals c1 c2 n table 2. contingency table (adapted from evert 2008: 1231) in the current study, we selected g2 (the log-likelihood measure, 2∑ij oij log(oij/eij), evert 2008), a statistical measure that is one of the most frequently used measures in collocational analyses. it is robust for differences in sample size, and compares observed frequencies and expected frequencies for each of the words taking into account the amount of evidence. we selected the top 100 items ranked according to g2 values (see the appendix). since g2 reports association strengths without indication of their direction, the top 100 collocates contains both words in strong attraction with the target word suoyi (i.e. in repulsion to kejian) and words in strong repulsion to suoyi (i.e. in attraction with the reference word kejian). the dir (direction) values provided by the r package the use of perspective markers and connectives in expressing subjectivity 69 were used to judge whether a word was attracted to suoyi (positive) or repelled by suoyi and hence attracted by kejian (negative). the delta-p value (o11/(o11+o12) o21/(o21+o22), gries 2013) is an effect size measure and measures the difference between the observed frequency of the target word in one context and that in the other context. in this study, delta-p value was used as a secondary criterion for the collocates: the words in attraction to suoyi all needed to be above the threshold of 0, and the words in repulsion to suoyi (the collocates of kejian) needed to be below this threshold (<0). the delta-p measure was applied because it is considered more psycholinguistically realistic, as it takes into account the directionality of the collocation: ‘whether the word1 is more predictive of word2 or the other way round’ (gries 2013: 141). an efficient and common way to interpret the outcomes of a collocational study is to cluster the collocates manually and draw meaningful interpretations based on these clusters (see gries & stefanowitsch 2010). in the current study, we were able to identify seven clusters within the top 100 collocates: pronouns, communication verbs, cognition verbs, modal verbs, and three types of perspective markers, namely exclamatory adverbials, expressions of expectation and expressions of importance. the results section will focus on these items; a full list of collocates for each distinctive collocates analysis can be found in the appendix. 3 results in this section, we discuss the results of the three distinctive collocates analyses. section 3.1 illustrates the general collocation patterns of the two connectives. section 3.2 compares the collocates of the two connectives in the clause preceding the connective and the clause following the connective. a genre-specific analysis shown in section 3.3 reveals the collocations in different genres. 3.1 general analysis the top 100 collocates (either attracted by suoyi or attracted by kejian) were categorized according to their semantic types. since our goal was to find out whether language users avoid overlap in the expression of subjectivity in their utterances, we checked the top 100 for linguistic elements that can be related to subjectivity and perspective marking. table 3 shows the collocates of suoyi that are relevant to our discussion, with their observed and expected frequencies and the g2 scores indicating the distinctiveness of particular collocates in the context of suoyi compared to the context of kejian. from the top 100, certain types of words stood out as significant collocates of suoyi, the connective that is underspecified in terms of subjectivity. an important cluster is formed by pronouns of all types (singular and plural, 1st, 2nd and 3rd person). on the one hand, pronouns can be linked to objective relations in which actors carry out certain actions for certain reasons. on the other hand, they can be used in subjective relations in which the pronouns refer to the individuals whose perspective is presented. therefore, we are not sure whether the higher number of occurrences of pronouns in the context of suoyi compared to the context of kejian should be attributed to a contextual feature of the objective relations that suoyi can express, or to the tendency to avoid doubling of subjectivity information in the context of kejian. this is much clearer for the other clusters that are attracted by suoyi, but repulsed by kejian: communication verbs, cognition verbs and modal verbs. both communication verbs and cognition verbs can express the epistemic stance of the speaker, to be specific, the evidentiality of the information. modal verbs indicate the author’s/character’s degree of certainty towards the proposition, which is also one of the dimensions of epistemic stance. this observation can be accounted for in terms of subjectivity. with suoyi in the sentence, the subjectivity information is underspecified. if subjectivity needs to be expressed, cognition verbs (marking evidentiality) and modal verbs (marking certainty) are used to help readers/hearers track the source of information. wei, evers-vermeul and sanders 70 collocates frequency (obs. vs exp.) g2 pronouns wo ‘i/me’ 22802: 21785 1189.91 ta ‘she/her’ 8972: 8602 380.85 ni ‘you’(singular) 8177: 7840 345.64 ta ‘he/him’ 24340: 23734 318.83 women ‘we/us’ 9212: 8864 314.23 tamen ‘they’ 7183: 6908 253.13 ziji ‘self’ 6104: 5905 146.47 nimen ‘you’(plural) 1129: 1080 53.62 communication verbs shuo ‘say’ 12301: 12050 103.46 gaosu ‘tell’ 856: 818 43.12 cognition verbs xiang ‘think’ 3990: 3829 159.58 zhidao ‘know’ 3333: 3199 131.28 renwei ‘believe’ 2794: 2699 74.07 xiwang ‘hope’ 1213: 1161 55.65 juede ‘feel’ 1576: 1516 55.04 pa ‘be afraid of’ 901: 863 39.06 modal verbs hui ‘would’ 8270: 8030 152.17 neng ‘can’ 8883: 8714 64.20 keneng ‘may’ 2711: 2628 55.84 yinggai ‘should’ 1526: 1469 51.18 keyi ‘can’ 4025: 3933 43.46 bixu ‘have to’ 2123: 2059 42.70 table 3. important collocates of suoyi from top 100 cognition and modal verbs would be repetitive for readers/hearers, however, in kejian contexts. kejian already implies someone is making the inference (normally, the speaker), so the use of cognition verbs and modal verbs would be a repetition of information on subjectivity. the strong association between suoyi and the communication verb shuo ‘say’ is indicative of the pattern we try to establish. an alternative explanation would be that this collocation is due to the high frequency of the expression suoyi shuo ‘so (i) say’. this expression has been segmented as two separate words by nlpir-ictclas, but in combination, it functions as a discourse marker that expresses the epistemic stance of the speaker. however, the cases in which suoyi and shuo are not intervened by any other linguistic elements, only account for 6.36% of the data (782 out of 12301 instances). this leaves many instances in which the communication verbs contributed to the expression of the epistemic stance of the speaker, as in example (6). still, our data show that communication verbs were not exclusively used in epistemic contexts; they could also be used for reporting an objective description of real-world events, as in example (7). therefore, we cannot be sure of the reason for the collocation of communication verbs and suoyi. this collocation pattern could be due to the speaker/author’s strategy to avoid repetition of subjectivity information in subjective relations, just as for the cases with cognition verbs. alternatively, communication verbs could be a feature of the context typical of the objective relations expressed by suoyi. the use of perspective markers and connectives in expressing subjectivity 71 (6) ta meiyou wudao yiyang meili de dongzuo, wangwang dongzuo de kaishi jius shi dadou de jieshu, suoyi li xiaolong shuo, jiequandao juedui bu shiyi biaoyan. it (jiequandao ‘jeet kune do’, a type of chinese kong fu) neg:have dance alike mod motion, often motion mod start just cop fight mod end, conj name said, jiequandao absolutely neg suit performance. it (jeet kune do) does not have beautiful motions like a dance; the start of a motion is often the end of a fight, so li xiaolong said jeet kune do is absolutely not suitable for performances. (7) ta shuo ta jiu shi changqi jianchi xialai, suoyi bai toufa zhijin dou bi ta de jiemeimen shao de duo. 3sgf said 3sgf just cop long:time insist down, conj till:now all compare 3sgf mod sisters less mod much. she said that she just kept (the good habit) for a long time, so up till now she has much less white hair compared to her sisters. some of the words in the top 100 list were repelled by suoyi and should therefore be seen as distinctive for kejian instead of suoyi, as illustrated in table 4. as mentioned in section 2, we included all kinds of indications of subjectivity, irrespective of their grammatical categories. the noun jiazhi ‘value’ was clustered with the adjective zhongyao ‘important’, because jiazhi ‘value’ is often associated with evaluations that are made from a person’s perspective. collocates frequency (obs. vs exp.) g2 exclamatory adverbials duome ‘how much’ 147: 235 263.63 hedeng ‘how much’ (literary) 33: 75 164.21 expressions of importance jiazhi ‘value’ 768: 843 84.06 zhongyao ‘important’ 1423: 1489 42.90 expressions of expectation jing ‘surprisingly’ 245: 272 33.95 table 4. important collocates of kejian from top 100 the exclamatory adverbials, expressions of expectation and expressions of importance can be related to subjectivity: they indicate that someone’s feeling or evaluation is involved, and that the hearer/reader is not merely dealing with a description of real-world facts. these collocational patterns indicate that language users do not necessarily avoid a doubling of information, as both kejian and these collocates express that subjectivity is involved. however, from this list of collocates of kejian, it can also be derived that language users do pay attention to the type of subjectivity information, in other words how the perspective of a speaker/character is involved. while the important collocates of suoyi (cognition verbs, communication verbs and modal verbs) could be related to epistemic stance marking, the important collocates of kejian – expectation markers and importance markers – can be related to attitudinal stance marking. hence, there is no doubling of epistemic stance marking information, the crucial type of subjectivity expressed by the connective kejian. 3.2 collocational analysis on different clauses given the general information on the contextual features in the analysis across clauses and genres, we obtained a basic understanding of the types of collocates that appear in the context of kejian and suoyi. however, we do not know from the overall analysis where these collocates appeared exactly wei, evers-vermeul and sanders 72 – do they appear in the clause preceding the connective, or do they appear in the clause following the connective? by precisely identifying the locations of different types of collocates, we can be more informed on how language users combine different linguistic cues to express subjectivity in discourse. moreover, for further psycholinguistic experiments, collocation distributions by clause provide insights into how linguistic stimuli should be designed to closely reflect authentic linguistic data. the current section therefore elaborates on the distribution of collocates in different clauses. table 5 and table 6 show the collocates of suoyi and kejian we derived from the top 100 in preceding clauses and in following clauses. preceding clause following clauses collocates frequency (obs. vs exp.) g2 frequency (obs. vs exp.) g2 pronouns wo ‘i/me’ 12056: 11547 497.78 10777: 10277 740.49 ta ‘she/her’ 5023: 4816 193.22 3959: 3796 192.48 ni ‘you’(singular) 4056: 3878 182.67 4128: 3974 155.00 ta ‘he/him’ 13497: 13154 171.33 10870: 10609 145.81 women ‘we/us’ 4721: 4534 166.11 4505: 4349 142.19 tamen ‘they’ 3826: 3685 113.05 3372: 3240 144.12 ziji ‘self’ 3501: 3382 86.09 2617: 2537 61.23 nimen ‘you’(plural) --601: 575 32.74 communication verbs shuo ‘say’ --4720: 4549 166.63 jiao ‘call’ --866: 824 65.78 ting ‘listen’ --542: 516 38.84 chengwei ‘be stated as’ --442: 421 33.77 wen ‘ask’ --413: 393 30.92 cognition verbs mingbai ‘understand’ 372: 353 24.87 -- zhidao ‘know’ 2452: 2342 117.21 -- xiang ‘think’ 2168: 2072 98.91 1825: 1761 59.68 renwei ‘believe’ 1898: 1833 46.96 903: 870 32.49 pa ‘be afraid of’ 666: 636 32.87 -- liaojie ‘understand’ 560: 533 31.11 -- juede ‘feel’ 903: 870 26.87 675: 648 29.90 xiwang ‘hope’ --713: 680 46.61 gan ‘dare’ --521: 497 32.69 modal verbs hui ‘would’ 4535: 4404 76.03 3747: 3639 76.24 keneng ‘may’ 1562: 1507 40.97 -- yinggai ‘should’ 501: 478 25.46 -- yiding ‘must’ 898: 858 41.32 -- xiande ‘seem’ --273: 259 29.38 neng ‘can’ --4621: 4513 57.71 bixu ‘have to’ --1452: 1410 29.55 table 5. important collocates of suoyi in preceding and following clauses most of the general collocation patterns also held in the analysis per clause except for the communication verbs. in both preceding and following clauses, pronouns, cognition verbs, modal verbs co-occurred with suoyi. most of these perspective markers may serve as the supplement of the use of perspective markers and connectives in expressing subjectivity 73 subjectivity information supplied by suoyi, regardless of whether they appear before or after the connective. examples (8) and (9) illustrate the combined use of suoyi ‘so’ and the perspective marker renwei ‘believe’, which can appear in both the clause before and the clause after the connective. (8) rensheng nande you tongtongkuaikuai xiangshou de rizi, ni cuoguo le, jianglai lao le shi, xiang xiangshou dou meiyou nengli, yanli bugou, meiyou yachi, tingjue you buhao, ni xiang qu xiangshou yixia, ye libucongxin, suoyi wo renwei yinggai chen nianqing de shihou, jishi xingle. life difficult have joyful enjoy mod day, you miss asp(pfv), future old asp(pfv) time, want enjoy even neg:have ability, eye neg:enough, neg:have tooth, hearing also neg:good, you want go enjoy a:bit, also incapable, conj i believe should when young mod time, in:time enjoy:life. it’s difficult to have joyful days to enjoy. if you miss them, you cannot enjoy them anymore when you get old: with failing eyesight, few teeth, and defective hearing, it is impossible to enjoy even for a little bit, so i believe (we) should enjoy life at youth. (9) ta zi renwei wancheng qingzang gaoyuan de senlin hangkong kance renwu shi ta yiburongci de zeren, suoyi ta bu gu ziji de shenti. 3sgm self believe complete qinghai-tibet plateau aviation survey task cop 3sgm unshirkable duty, conj 3sgm neg care self mod health. he believes that completing the aviation survey task in qinghai-tibet plateau is his unshirkable duty, so he does not care about his own health. an important difference with the general collocation pattern is that the communication verbs appeared as important collocates of suoyi only in the clauses following this connective. this means that for the co-occurrence with such reportative verbs, no significant difference between suoyi and kejian can be found in the first clause. a clear-cut difference between the collocates of kejian in preceding and following clauses is suggested in table 6. exclamatory adverbials and expressions of importance were only distinctive for kejian in the clauses following this connective. this finding may be due to the tendency to express an evaluation in the second clause in a forward causal relation: the evaluation of importance is expressed in the second clause based on the events/phenomena described in the first clause, as is illustrated in (10). (10) huaiyun hou jiaolü bu’an de muqin geng rongyi nanchan he shengchu yichang de haizi, kejian yunqi zhong zhuyi xinli weisheng shi duome zhongyao. pregnant after anxious disturbed mod mother more easy dystocia and deliver abnormal mod infant, conj pregnancy middle pay:attention:to mental health is so:much important. mothers who are anxious and disturbed after pregnancy are more likely to suffer dystocia and deliver abnormal infants, so paying attention to mental health is very important during pregnancy. expressions of expectation only appeared as important collocates of kejian in the preceding clause. these linguistic elements express an attitude of the speaker towards the situation described in the first clause, such as in example (11): the author is surprised by the fact that wang jian, a general, won the battles both in the south and in the north. wei, evers-vermeul and sanders 74 preceding clause following clauses collocates frequency (obs. vs exp.) g2 frequency (obs. vs exp.) g2 exclamatory adverbials duome ‘how much’ --67: 153 343.15 hedeng ‘how much’ (literary) --16: 60 207.58 xiangdang ‘considerably’ --269: 303 49.06 expressions of expectation jing ‘surprisingly’ 110: 143 69.67 -- juran ‘unexpectedly’ 44: 59 34.71 -- jingran ‘surprisingly’ 30: 41 27.07 -- expressions of importance zhongyao ‘important’ --757: 841 109.15 juzuqingzhong ‘crucial’ --3: 9 29.39 jiazhi ‘value’ 432:463 26,61 336:380 64.20 communication verbs cheng ‘state’ 231: 265 47.84 -- yue ‘say’ (formal) 42: 57 35.93 -- yan ‘speak’(formal) 165: 185 25.00 -- table 6. important collocates of kejian in preceding and following clauses (11) wang jian jingran neng zai nanbei liang fang de zuozhan zhong dou qusheng, kejian qi zai yongbing fangmian yingdang shi shuyu quanfangwei de wujiang. wang jian (a general in chinese history) surprisingly can at south:north two side mod battle in all win, kejian 3sgm at military aspect should cop belong:to extensive mod general. surprisingly wang jian won the battles both in the south and in the north, so he should be a general with extensive military capabilities. in contrast to the general collocation pattern, some instances of communication verbs were found as important collocates of kejian instead of suoyi in the preceding clause. most of them are formal expressions, which are more characteristic of formal contexts such as informative and argumentative texts. compared to the findings in table 5 and 6, communication verbs can be collocates of either suoyi or kejian, depending on the formalities encoded in different specific communication verbs. formal communication verbs such as cheng ‘state’ and yue ‘say’ patterned with kejian, while informal communication verbs such as shuo ‘say’ patterned with suoyi. therefore, it is not possible to identify a uniform pattern in the co-occurrence of communication verbs in relation to the degree of subjectivity expressed by the connective. 3.3 collocational analysis on different genres the results discussed so far may be the result of a confound with the genre preference of the connectives under investigation. suoyi is a generic connective that can be used for all types of genres, while kejian is not frequent in narrative texts (cf. table 1). moreover, several of the collocate clusters found in section 3.1 and 3.2 may be a side-effect of genre preferences as well. for example, communication verbs can be expected to appear more in the narrative genre, just like pronouns. therefore, communication verbs and pronouns may pattern with suoyi simply because they all share the preference for the narrative genre. to neutralize the influence of genre as a confounding factor, we further examined the collocation of the two connectives in different genres, the use of perspective markers and connectives in expressing subjectivity 75 namely narratives and non-narratives. the collocation distributions of these two connectives with other linguistic elements in different genres are summarized in table 7 and table 8. narratives non-narratives collocates frequency (obs. vs exp.) g2 frequency (obs. vs exp.) g2 pronouns wo ‘i/me’ 16329: 16043 241.62 6473: 6104 407.51 women ‘we/us’ 3503: 3434 69.49 5709: 5424 257.24 tamen ‘they’ 3744: 3673 67.21 3439: 3279 129.18 ta ‘she/her’ 7614: 7514 58.08 1358: 1290 62.14 ta ‘he/him’ 16328: 16196 43.53 8012: 7868 37.47 ni ‘you’(singular) 6439: 6368 33.57 1738: 1632 129.07 ziji ‘self’ 3202: 3155 31.68 2902: 2789 72.55 ta ‘it’ --4099: 3979 54.42 nimen ‘you’(plural) --339: 316 34.80 communication verbs shuo ‘say’ --6623: 6433 83.64 chengwei ‘be stated as’ --444: 418 28.38 gaosu ‘tell’ --313: 293 27.29 cognition verbs xiang ‘think’ 2776: 2733 30.99 1214: 1156 49.56 zhidao ‘know’ 2504: 2463 30.56 829: 794 24.15 renwei ‘believe’ --1928: 1837 76.42 xiwang ‘hope’ --713: 673 44.12 juede ‘feel’ --497: 470 26.93 modal verbs hui ‘would’ 4043: 3990 30.82 4227: 4078 84.80 neng ‘can’ 3258: 3223 16.11 5625: 5477 58.89 bixu ‘have to’ 545: 533 15.05 1578: 1511 47.52 zhineng ‘can only’ 240: 233 13.56 -- keneng ‘may’ 876: 860 13.46 1835: 1758 54.43 yinggai ‘should’ --898: 853 41.34 keyi ‘can’ --2492: 2415 37.27 table 7. important collocates of suoyi in different genres pronouns were observed as important collocates of suoyi in both types of genres, which indicated that this collocation pattern is not a side-effect of the genre preference of suoyi. cognition verbs also appeared as important collocates of suoyi in both types of genres. although the exact collocates differ per genre, they all expressed the same cognitive state of knowing and thinking. in addition, modal verbs were still distinctive collocates for suoyi in both narratives and nonnarratives. therefore, we may infer that the collocation of the generic connective suoyi with cognition verbs and modal verbs is not due to genre differences. we did find a difference with the general collocation pattern, however. contrary to our hypothesis communication verbs were found not to be significant collocates of suoyi in narratives, although they were still important collocates in non-narratives (559 cases (8.44%) of which were instances of suoyi immediately followed by shuo). apparently, suoyi and kejian do not differ in their preference for co-occurring with communication verbs in the narrative genre, but only in the non-narrative genre. wei, evers-vermeul and sanders 76 narratives non-narratives collocates frequency (obs. vs exp.) g2 frequency (obs. vs exp.) g2 exclamatory adverbials duome ‘how much’ 102: 124 62.78 45: 112 238.76 hedeng ‘how much’ (literary) --15: 53 152.93 expressions of expectation juran ‘unexpectedly’ 68: 79 25.59 -- jing ‘surprisingly’ 140: 151 16.78 105: 124 25.97 guoran ‘as expected’ 28: 34 16.69 -- expressions of importance jiazhi ‘value’ 660: 723 57.04 table 8. important collocates of kejian in different genres even though the ratios of observed versus expected frequencies differ from the ones in the general analysis, the top 100 items still display similar collocation patterns for kejian. as table 8 indicates, exclamatory adverbials and expressions of expectations stayed distinctive for kejian in both narratives and non-narratives, although there were some differences per item. these perspective markers on the attitudinal stance dimension of expectedness are more associated with kejian rather than suoyi across genres. expressions of importance only appeared as important collocates of kejian in non-narrative genres. 4 general conclusion and discussion the current study explored whether language users try to avoid doubling of subjectivity information in discourse, specifically in coherence relations. on the basis of distinctive collocates analyses, we examined whether the chinese connectives kejian and suoyi, which differ in the degree of subjectivity they express, differed in their types of collocates, and especially if they differed in the types of perspective markers they co-occurred with. in line with our predictions, the degrees of subjectivity encoded in the two connectives was related to the type of linguistic cues in their contexts. in section 4.1, we will summarize and discuss the general patterns we found; in section 4.2, we will discuss our main findings per clause (preceding or following the connective) and genre, and in section 4.3, we discuss the limitations of our study and put forward some suggestions for future research. 4.1 general collocation patterns in line with pragmatic principles and uid in general, the underspecified connective suoyi ‘so’, which can express both subjective and objective relations, patterned with more occurrences of cognition verbs and modal verbs in comparison to the specific subjective connective kejian ‘so’. in the context of kejian, we found more exclamatory adverbials, expressions of importance and expressions of expectation compared to the context of suoyi as a reference level. the collocation results showed that perspective markers as a general type of linguistic cues marking subjectivity can be used in combination with either of the two causal connectives. however, if perspective markers are specifically categorized into sub-types with regards to various dimensions of subjectivity, different collocation patterns surfaced. suoyi turned out to collocate with epistemic stance markers more often, while kejian co-occurred with attitudinal stance markers. the collocation pattern of epistemic stance markers and suoyi is consistent with horn’s pragmatic theory of relation principle (reducing the speaker’s production effort) and quality principle (reducing the hearer’s comprehension effort). from the perspective of the r principle, if subjectivity information on the epistemic stance (including (un)certainly and evidentiality) is already specified in the connective kejian, epistemic stance markers in the context of the connective the use of perspective markers and connectives in expressing subjectivity 77 are redundant, i.e. not efficient from the speaker economy account. suoyi, by contrast, does not provide sufficient information on the epistemic stance, and the use of epistemic stance markers therefore provides valuable information that compensates the lack of subjectivity information in suoyi. the q principle is observed and hearers/readers’ comprehension process should be facilitated. the collocation results can also be well explained by the uniform information density theory account. with the two alternative connectives expressing discourse coherence, the presence of epistemic stance markers (e.g. cognition verbs, modal verbs) makes the content of the context highly expectable (high probability and low information), which is why it is more likely to have an underspecified connective, suoyi in this case. on the other hand, utterances with fewer occurrences of epistemic stance markers make the content conveyed by the context unexpected (low probability and high information). in this sense, the use of a specific connective is preferred. the prevalence of epistemic stance markers in the context of suoyi and their lower co-occurrence with kejian fit the need for a uniform information density throughout the sentence in terms of subjectivity. optimal information density is realized in this way. however, speakers/authors did not avoid overlap in the expression of subjectivity at all costs. some attitudinal stance markers such as jingran ‘surprisingly’ and zhongyao ‘important’, which also indicate the involvement of a speaker responsible for an evaluation, occurred as important collocates in the context of kejian. both epistemic stance markers and attitudinal stance markers express that a source of information is involved. apparently, in their use of kejian, which also indicates that a source of information is involved, speakers and writers do not avoid overlap with that same information provided by attitudinal stance markers. however, kejian does not overlap with attitudinal stance markers in the subjectivity dimension it expresses, i.e. in expressing how the source of information is involved. the fact that kejian patterns with attitudinal but not with epistemic stance markers indicates that language users try to avoid overlap in the expression of these dimensions. the connective kejian and epistemic stance markers both indicate how certain the speaker/writer is about the information, while attitudinal stance markers express the attitude or feelings of a person towards the information. in terms of uid, the combination of attitudinal stance markers and kejian does not create high information density in the utterance. taken together, the two observations above help to explain why attitudinal stance markers and kejian were found in collocation; they show a kind of agreement of subjectivity at the discourse level, jointly contributing to a subjective context. 4.2 collocation patterns in different genres and clauses to test for potential genre influences on the results, we performed collocational analyses on narratives and non-narratives separately. these analyses showed that communication verbs patterned with suoyi in the non-narrative genre, but not in narratives. this asymmetry might be due to different usage patterns of communication verbs in these genres. as illustrated in section 3.1, communication verbs can be used to express an epistemic stance, as in example (6), or as a reportative verb to introduce a description of real-world events, as in example (7). in the nonnarrative genre, we would expect a higher frequency of epistemic communication verbs. the fact that communication verbs stood out as important collocates of suoyi in this genre is in line with the avoidance of doubling of information (as illustrated in table 7): suoyi has more needs of epistemic markers to strengthen the epistemic nature of the utterance than kejian, which encodes such information by itself. in narratives, with their abundance of descriptions of real-world events, however, we would expect a higher number of reportative communication verbs, which do not create a doubling of information with the information provided by kejian when they are used to report objective events in one of the clauses connected by kejian. this might explain why communication verbs do not stand out as collocates of suoyi in narratives, although this explanation needs to be corroborated in future corpus research on the actual usage of different types of communication verbs in narrative and non-narrative genres. wei, evers-vermeul and sanders 78 apart from communication verbs, all other types of collocates of suoyi in the general analysis still surfaced as important in both narrative and non-narrative genres. although the individual collocates of each cluster slightly vary per genre, the collocation between cognition verbs, modal verbs and pronouns with suoyi was robust across genres. as for kejian, exclamatory adverbials and expressions of expectation also appeared in both narratives and non-narratives as important collocates, which suggests that the perspective markers related to expectations were indeed an important contextual feature of kejian. in order to locate the positions of each type of collocates in causal relations, we analyzed the preceding clauses and the following clauses of connectives separately. most of the perspective markers as collocates of suoyi appeared in both the clauses preceding the connective and the clauses following it, except for communication verbs. the collocates of kejian, however, differed in the position they appeared in contexts. expressions of expectation appeared as important collocates of kejian in the preceding clause, which makes sense because in a subjective relation with an argument-claim structure, expressions of expectation such as jingran ‘surprisingly’ in example (11) mark the speaker’s surprisal – either about the propositional content of the clause preceding the connective, or about the fact as an argument-claim relation as a whole. both exclamatory adverbials and expressions of importance tended to appear with kejian in the clause following the connective. these perspective markers served as expressions of the speaker’s attitude towards the claim presented in the second segment of the relation. communication verbs exhibited very different collocation patterns depending on the clause they occurred in. in the preceding clause, formal communication verbs did not surface as important collocates of suoyi, but rather patterned with kejian more often (example (11)). in the following clause, the tendency was reversed – communication verbs only surfaced as collocates of suoyi such as in example (6). as example (12) shows, the communication verb yue ‘say’ co-occurred with kejian mainly in very formal texts, in which kejian was found more often. therefore, such collocation pattern could be attributed to an effect of formality. on the basis of the current explorative study, we cannot draw a decisive conclusion on this issue. further studies on the use of communication verbs in different contexts are needed, especially to find out whether our ideas about the formality and about the objective versus the subjective use of communication verbs can be corroborated. (12) (master linji, a buddhism master) da yue: ruguo yi kou qi bu lai, zhe routi haiyou ganqing ma? kejian qinggan buzai routi shang, er zai lingxing shang. (master linji) answer say: if one cl breath neg come, this body have emotion? conj emotion neg at body, but at spirituality on. (master linji) said in response: if one doesn’t breathe anymore, does the body still have emotions? so emotion is not in the body, but rather in the spirit. 4.3 future studies and conclusion in closing, we would like to discuss some limitations of this study. first, the current study only comprises a small set of connectives, which enabled us to provide an in-depth analysis. future studies could extend this set to include other causal connectives. interesting candidates for further distinctive collocates analyses seem to be yushi and yin’er (both meaning ‘so/therefore’); li et al. (2013) have shown that these causal connectives differ in the area of volitionality (i.e., is the causal relation an intentional one or not?). it would be interesting to see whether collocation patterns vary with this feature as well. second, the sentences in causal relations were retrieved directly from the corpus without any manual annotation of the relation type, which would have taken even more time and effort. kejian is mainly used for subjective relations, while suoyi is generic (li et al. 2013). this means that the sample of suoyi contexts contained both subjective and objective relations, while the contexts of kejian mainly consisted of subjective relations. the unbalanced distribution of relations in the the use of perspective markers and connectives in expressing subjectivity 79 contexts of the two connectives may be a confounding factor. for instance, the fact that pronouns were distinctive for suoyi may be a feature of objective relations, because the descriptions of events and acts in objective relations may involve the use of pronouns. nonetheless, the major findings such as the fact that modal verbs and cognition verbs are important collocates of suoyi are not characteristic of objective relations at all. we would expect stronger distinctive collocation patterns of these expressions with suoyi if we had limited the scope of the investigation to subjective relations only. more fine-grained analyses are expected to shed a clearer light on this issue. third, collocational analyses only provide rough tendencies in the word use in the context of a target word. it cannot support any decisive inferences, such as the predictability of one word given the other word. to further investigate the relation between connectives and their collocates, one could refer to regression analyses to investigate whether the presence of certain words in the context correlates with the presence of specific connectives, or one could opt for experimental research to investigate the effects of perspective markers on the processing of connectives. a fourth limitation is that collocational analyses cannot distinguish word forms with the same syntactic tag but with multiple meanings. a relevant case in our corpus concerns modal verbs, some of which can be used in either a deontic or an epistemic way. although both types of modals can have a subjective/interpersonal function (lyons, 1977), they differ in the meaning they encode deontic modals express obligations and permissions, while epistemic modals concerns believes (foley & van valin, 1984; johnson-laird & ragni, 2019; verstraete, 2001). our claims about modal verbs as linguistic means to express perspective would pertain to epistemic modals in particular. examining whether the modal verbs in our corpus are used in the deontic or the epistemic way, however, would require a manual screening process. this seems to be an interesting avenue for further research. lastly, for practical reasons we made no distinction between argumentative and informative genres. these two genres have certain features in common in which they differ from narratives. for instance, both argumentative and informative genres have the author as the illocutionary force in most of the cases, while in narrative texts other characters are also frequently involved as the illocutionary force. however, argumentative genres also differ from informative genres in several respects – the use of communication verbs, for instance, may be different between argumentative texts and informative texts. separate analyses of the argumentative genres and the informative genres can provide a more refined picture. despite these limitations, we can conclude that the explorative approach in the current study has produced a number of interesting insights into the way in which subjectivity is expressed in discourse. what is more, this study has illustrated the relation between connectives and perspective markers, demonstrating that this will be a productive area for future research to pursue. acknowledgements the first author’s work was funded by the chinese scholarship council (grant number 2013077220042, 2013) and the starting grant of peking university (no. 7101502221). we thank dr. dirk speelman for his advice on the methodology and research design. wei, evers-vermeul and sanders 80 appendix. top 100 collocates per analysis 1-25 26-50 51-75 76-100 words dir words dir words dir words dir wo ‘i/me’ 1 zhan ‘occupy’ -1 juede ‘feel’ 1 hao ‘good’ 1 yinwei ‘because’ 1 laodong ‘work’ -1 bingfei ‘not’ -1 zai ‘again’ 1 youyu ‘since’ 1 cai ‘only’ 1 nimen ‘you’(plural) 1 yizhi ‘always’ 1 ta ‘she/her’ 1 yi ‘already’ -1 gei ‘give’ 1 pa ‘be afraid of’ 1 ni ‘you’(singular) 1 jiazhi ‘value’ -1 yinggai ‘should’ 1 zong’e ‘total amount’ -1 ta ‘he/him’ 1 woguo ‘our country’ -1 dan ‘but’ 1 di ‘emperor’ -1 women ‘we/us’ 1 gen ‘and’ 1 jizai ‘record literally’ -1 zengzhang ‘increase’ -1 duome ‘how much’ (excl. mood) -1 lai ‘come’ 1 ziben ‘capital’ -1 yin ‘because’ 1 tamen ‘they’ 1 rang ‘let’ 1 you ‘further’ 1 juan ‘roll’ -1 jiu ‘just’ 1 renwei ‘believe’ 1 danshi ‘but’ 1 men (plural affix) 1 yao ‘want’ 1 shihou ‘time’ 1 zai ‘at’ 1 zhege ‘this one’ 1 bu ‘no’ 1 na ‘that’ 1 tebie ‘especially’ 1 wei ‘for’ -1 dou ‘all’ 1 bisai ‘race’ 1 de (particle) 1 shangnian ‘last year’ -1 hedeng ‘how much’ (literary)(excl. mood) -1 shengyu jiazhi ‘surplus value’ -1 xianzai ‘now’ 1 yige ‘one’ 1 xiang ‘think’ 1 tai ‘too’ 1 gongren ‘worker’ -1 bujin ‘not only’ -1 hui ‘would’ 1 qi ‘that’ -1 bili ‘proportion’ -1 ju ‘according to’ -1 zhi ‘of’ -1 de ‘of’ 1 keyi ‘can’ 1 mei (negation) 1 ziji ‘self’ 1 di (adverbial particle) 1 gaosu ‘tell’ 1 cidian ‘dictionary’ -1 meiyou ‘have not’ 1 neng ‘can’ 1 zhongyao ‘important’ -1 zhunbei ‘prepare’ 1 zhidao ‘know’ 1 zhe ‘imperfective aspect marker’ 1 bixu ‘must’ 1 kaishi ‘begin’ 1 qu ‘go’ 1 zuo ‘make’ 1 shangpin ‘goods’ -1 shengchan ‘produce’ -1 hen ‘very’ 1 yu ‘left’ -1 zhi ‘reach’ -1 fazhan ‘develop’ -1 yuan ‘chinese dollar’ -1 daojiao ‘taoism’ -1 ba (disposal constr.) 1 ren ‘people’ 1 shuo ‘say’ 1 keneng ‘may’ 1 shu ‘book’ -1 jing ‘surprisingly’(short) -1 le (perf. aspect marker) 1 xiwang ‘hope’ 1 jingji ‘economy’ -1 zhao ‘find’ 1 table 1. top 100 collocates ranked by g2 (general context) the use of perspective markers and connectives in expressing subjectivity 81 1-25 26-50 51-75 76-100 words dir words dir words dir words dir yinwei ‘because’ 1 yu ‘left’ -1 dui ‘correct’ 1 yinyong ‘quote’ -1 youyu ‘since’ 1 juan ‘roll’ -1 zhi ‘reach’ -1 deng ‘wait’ -1 wo ‘i/me’ 1 jizai ‘record literally’ -1 shou ‘first’ -1 bijiao ‘compare’ 1 shi ‘is’ 1 dan ‘but’ 1 jin ‘now’ -1 chuan ‘pass’ -1 de ‘of’ 1 yijing ‘already’ 1 juran ‘unexpectedly’ -1 gongzuo ‘work’ 1 ta ‘she/her’ 1 yin ‘because’ 1 ju ‘sentence’ -1 yan ‘how’(literary) -1 ni ‘you’ (singular) 1 cheng ‘state’ -1 zengzhang ‘increase’ -1 jiu ‘just’ 1 bu ‘no’ 1 danshi ‘but’ 1 yishang ‘above’ -1 da ‘reach’ -1 ta ‘he/him’ 1 renwei ‘believe’ 1 na ‘that’ 1 jingran ‘surprisingly’ -1 women ‘we/us’ 1 zong’e ‘total amount’ -1 tongji ‘statistics’ -1 juede ‘feel’ 1 hen ‘very’ 1 zheng ‘straight’ 1 san ‘three’ -1 shihou ‘time’ 1 yuan ‘chinese dollar’ -1 shu ‘book’ -1 pa ‘be afraid of’ 1 jiazhi ‘value’ -1 meiyou ‘have not’ 1 yiding ‘for sure’ 1 zhe ‘imp.asp.marker’ 1 shi ‘family name’ -1 zhan ‘occupy’ -1 keneng ‘may’ 1 chanzhi ‘output value’ -1 dongxi ‘things’ 1 zhidao ‘know’ 1 buguo ‘but’ 1 ju ‘according to’ -1 wen ‘literal’ -1 tamen ‘they’ 1 gai ‘change’ -1 chuban ‘first edition’ -1 you ‘further’ 1 xiang ‘think’ 1 gen ‘and’ 1 yige ‘one’ 1 zai ‘at’ 1 tai ‘too’ 1 zhege ‘this one’ 1 nian ‘year’ -1 yinggai ‘should’ 1 zhi ‘of’ -1 ren ‘people’ 1 shi ‘poem’ -1 xing ‘gender’ 1 dou ‘all’ 1 xihuan ‘like’ 1 liaojie ‘understand’ 1 chutu ‘unearthed’ -1 ziji ‘self’ 1 que ‘but’ -1 ji ‘collection’ -1 ce ‘volume’ -1 hui ‘would’ 1 bili ‘proportion’ -1 jian ‘see’ -1 ru ‘if’ -1 yao ‘want’ 1 yue ‘say’ (literary) -1 zhe ‘this’ 1 yin ‘print’ -1 jing ‘surprisingly’ -1 woguo ‘our country’ -1 bisai ‘race’ 1 yan ‘speak’(literary) -1 wei ‘for’ -1 shengyu jiazhi ‘surplus value’ -1 yi ‘upon’ -1 mingbai ‘understand’ 1 table 2. top 100 collocates in the preceding clause, ranked by g2 wei, evers-vermeul and sanders 82 1-25 26-50 51-75 76-100 words dir words dir words dir words dir wo ‘i/me’ 1 qi ‘that’ -1 jingji ‘economy’ -1 guanxi ‘relation’ -1 duome ‘how much’ (excl. mood) -1 gei ‘give’ 1 de (particle) 1 renwei ‘believe’ 1 shi ‘is’ -1 rang ‘let’ 1 dao ‘reach’ 1 kaishi ‘begin’ 1 jiu ‘just’ 1 jiao ‘call’ 1 gen ‘and’ 1 you ‘further’ 1 hedeng ‘how much’ -1 jiazhi ‘value’ -1 shen ‘very’ -1 juda ‘huge’ -1 (literary, excl. mood) ta ‘she/her’ 1 zhi ‘of’ -1 shehui ‘society’ -1 wen ‘ask’ 1 shuo ‘say’ 1 ziji ‘self’ 1 ting ‘listen’ 1 yizhi ‘always’ 1 yi ‘already’ -1 fazhan ‘develop’ -1 shen ‘deep’ -1 tebie ‘especially’ 1 ni ‘you’(singular) 1 xiang ‘think’ 1 zuo ‘make’ 1 juede ‘feel’ 1 qu ‘go’ 1 neng ‘can’ 1 qiye ‘company’ -1 zaiyu ‘lie in’ -1 ta ‘he/him’ 1 daojiao ‘taoism’ -1 bing ‘and’ -1 bixu ‘must’ 1 yao ‘want’ 1 yi ‘one’ 1 shangpin ‘goods’ -1 zai ‘again’ 1 tamen ‘they’ 1 shihou ‘time’ 1 qianli ‘potential’ -1 juzuqingzhong ‘crucial’ -1 women ‘we/us’ 1 bian ‘convenient’ 1 zhao ‘find’ 1 miansha ‘cotton’ -1 bingfei ‘not’ -1 ci ‘times’ 1 zuoyong ‘function’ -1 xiande ‘seem’ 1 zhongyao ‘important’ -1 qing ‘please’ 1 shengyu jiazhi ‘surplus value’ -1 zhe ‘imp. aspect marker’ 1 cai ‘only’ 1 shi ‘time’ 1 dang ‘at’ 1 liang ‘good’ -1 lai ‘come’ 1 bisai ‘race’ 1 bujin ‘not only’ -1 ziben ‘capital’ -1 dangshi ‘that time’ -1 woguo ‘our country’ -1 de ‘of’ -1 li ‘inside’ 1 di (adverbial particle) 1 xiangdang ‘considerably’ -1 jiang ‘be about to’ 1 yongxin ‘attentively’ -1 le (perf. asp. marker) 1 bei (passive) 1 yinsu ‘reason’ -1 zhidu ‘system’ -1 laodong ‘work’ -1 xiwang ‘hope’ 1 diwei ‘status’ -1 meiyou ‘have not’ 1 dou ‘all’ 1 na ‘that’ 1 chengwei ‘be stated as’ 1 dai ‘take’ 1 ba (disposal constr.) 1 xianzai ‘now’ 1 nimen ‘you’(plural) 1 dui ‘correct’ -1 hui ‘would’ 1 yijing ‘already’ -1 gan ‘dare’ 1 zaoyi ‘very early already’ -1 table 3. top 100 collocates in the following clause, ranked by g2 the use of perspective markers and connectives in expressing subjectivity 83 1-25 26-50 51-75 76-100 words dir words dir words dir words dir wo ‘i/me’ 1 meiyou ‘have not’ 1 zai ‘again’ 1 ren ‘mercy’ -1 yinwei ‘because’ 1 jie ‘session’ -1 gen ‘and’ 1 men (plural affix) 1 jiu ‘just’ 1 juran ‘unexpectedly’ -1 qin ‘diligent’ -1 jie ‘all’ -1 women ‘we/us’ 1 wu ‘no’ -1 xi ‘to be informed of’ -1 an ‘we’(oral) -1 tamen ‘they’ 1 cai ‘only’ 1 liang ‘good’ -1 fuhe ‘conform to’ -1 zhi ‘of’ -1 jian ‘see’ -1 mingming ‘apparently’ -1 lin yuxiang (name) -1 duome ‘how much’ (excl. mood) -1 xiang ‘mutural’ -1 bingfei ‘not’ -1 zhi ‘only’ 1 ta ‘she/her’ 1 de ‘of’ 1 haocheng ‘be known as’ -1 diren ‘enermy’ 1 yao ‘want’ 1 ba (disposal construction) 1 shang ‘even’ -1 dou ‘all’ 1 qi ‘that’ -1 lajiao ‘peper’ -1 stalin (name) -1 he ‘and’ 1 ta ‘he/him’ 1 que ‘but’ -1 shouji ‘collect’ -1 ding ‘vessel’ -1 youyu ‘since’ 1 qu ‘go’ 1 bujin ‘not only’ -1 shenme ‘what’ (literary) -1 rang ‘let’ 1 zou ‘walk’ 1 jing ‘surprisingly’ -1 beibi ‘mean’ -1 kejian ‘so’ -1 zhe ‘imperfective aspect marker’ 1 guoran ‘as expected’ -1 shengguo ‘outrace’ -1 ni ‘you’(singular) 1 yuanyi ‘would like’ 1 ji ‘avoid’ -1 yongxin ‘attentively’ -1 ziji ‘self’ 1 gei ‘give’ 1 neng ‘can’ 1 xinli ‘psychology’ -1 zu ‘block’ -1 hourou ‘next generation’ -1 shen ‘deep’ -1 zhineng ‘can only’ 1 shi ‘family name’ -1 danshi ‘but’ 1 kandao ‘see’ 1 keneng ‘may’ 1 zai ‘at’ 1 biaodian ‘punctuation’ -1 xie ‘some’ 1 zhou enlai (name) -1 xiang ‘think’ 1 shen ‘very’ -1 zhong ‘heavy’ -1 anquanju ‘security office’ -1 hui ‘would’ 1 du ‘poison’ -1 fangyang ‘dialect’ -1 juejin ‘tunnelling’ -1 zhidao ‘know’ 1 di (adverbial particle) 1 la ‘spicy’ -1 shen congwen (name) -1 renyi ‘righteousness’ -1 chunqiu ‘spring and autumn’ -1 bixu ‘must’ 1 yongtu ‘purpose’ -1 shihou ‘time’ 1 bu ‘no’ 1 dao ‘reach’ 1 wufa ‘have no mean’ 1 yuan ‘source’ -1 lai ‘come’ 1 zi ‘character’ -1 zhi ‘reach’ -1 table 4. top 100 collocates in narrative genres, ranked by g2 wei, evers-vermeul and sanders 84 1-25 26-50 51-75 76-100 words dir words dir words dir words dir youyu ‘since’ 1 laodong ‘work’ -1 di ‘emperor’ -1 yi ‘also’ -1 yinwei ‘because’ 1 ta ‘she/her’ 1 qu ‘go’ 1 xunlian ‘train’ 1 wo ‘i/me’ 1 neng ‘can’ 1 yin ‘because’ 1 chuban ‘first edition’ -1 women ‘we/us’ 1 shengyu jiazhi ‘surplus value’ -1 xing ‘gender’ 1 juede ‘feel’ 1 duome ‘how much’ (exclamatory mood) -1 jiazhi ‘value’ -1 jizai ‘record literally’ -1 ziben ‘capital’ -1 hedeng ‘how much’ (literary, excl. mood) -1 dan ‘but’ 1 chuan ‘pass’ -1 chanzhi ‘output value’ -1 yao ‘want’ 1 keneng ‘may’ 1 nimen ‘you’(plural) 1 chang ‘space’ 1 dou ‘all’ 1 ta ‘it’ 1 bingfei ‘not’ -1 miansha ‘cotton’ -1 tamen ‘they’ 1 xiang ‘think’ 1 canjia ‘participate’ 1 chongdie ‘overlap’ -1 ni ‘you’(singular) 1 bixu ‘must’ 1 gongren ‘worker’ -1 xianzai ‘now’ 1 de ‘of’ 1 dui ‘group’ 1 shangnian ‘last year’ -1 jing ‘surprisingly’ -1 bisai ‘race’ 1 aoyunhui ‘olympic game’ 1 zai ‘at’ 1 he ‘and’ 1 bu ‘no’ 1 daojiao ‘taoism’ -1 kaishi ‘begin’ 1 yizhi ‘always’ 1 hui ‘would’ 1 xiwang ‘hope’ 1 di (adverbial particle) 1 shu ‘book’ -1 yuan ‘chinese dollar’ -1 zuo ‘make’ 1 zhunbei ‘prepare’ 1 renhe ‘any’ 1 shuo ‘say’ 1 woguo ‘our country’ -1 zong’e ‘total amount’ -1 ce ‘volume’ -1 jiu ‘just’ 1 bijiao ‘compare’ 1 ji ‘collection’ -1 xianling ‘shilling’ -1 zhan ‘occupy’ -1 yinggai ‘should’ 1 juan ‘roll’ -1 shengchanziliao ‘production means’ -1 hen ‘very’ 1 feichang ‘very much’ 1 you ‘further’ 1 danshi ‘but’ 1 renwei ‘believe’ 1 cai ‘only’ 1 shijian ‘time’ 1 jinxing ‘going on’ 1 meiyou ‘have not’ 1 tebie ‘especially’ 1 zui ‘most’ 1 zhidao ‘know’ 1 ziji ‘self’ 1 shi ‘is’ 1 chengwei ‘be stated as’ 1 fubai ‘corruption’ -1 zhi ‘of’ -1 tai ‘too’ 1 zhexie ‘these’ 1 yiban ‘ordinary’ 1 yi ‘already’ -1 ta ‘he/him’ 1 bili ‘proportion’ -1 bei (passive) 1 yu ‘left’ -1 keyi ‘can’ 1 gaosu ‘tell’ 1 lai ‘come’ 1 table 5. top 100 collocates in non-narrative genres, ranked by g2 the use of perspective markers and connectives in expressing subjectivity 85 references laurence anthony (2016). antconc (version 3.4.4.0) [computer software]. tokyo, japan, waseda university. retrieved from http://www.laurenceanthony.net/. fatemeh torabi asr and vera demberg (2015). uniform information density at the level of discourse relations: negation markers and discourse connective omission. in proceedings of the 11th international conference on computational semantics, pages 118-128, association for computational linguistics, london. monika bednarek (2006). evaluation in media discourse: analysis of a newspaper corpus. journal of quantitative linguistics, 17(3), 253-256. monika bednarek (2009). dimensions of evaluation: cognitive and linguistic perspectives. pragmatics & cognition, 17(1), 146-175. douglas biber, stig johansson, geoffrey leech, susan conrad, and edward finegan (1999). longman grammar of spoken and written english. london, longman. bruce k. britton (1994). understanding expository text: building mental structures to induce insights. in morton a. gernsbacher (ed.), handbook of psycholinguistics (1st ed.): 641-674. san diego, ca, academic press. kate cain and hannah m. nash (2011). the influence of connectives on young readers’ processing and comprehension of text. journal of educational psychology, 103(2), 429-441. anneloes r. canestrelli, willem m. mak, and ted j.m. sanders (2013). causal connectives in discourse processing: how differences in subjectivity are reflected in eye-movements. language and cognitive processes, 28(9), 1394-1413. reinier cozijn, leo g.m. noordman, and wietske vonk (2011). propositional integration and world-knowledge inference: processes in understanding because sentences. discourse processes, 48(7), 475-500. kenneth w. church, william gale, patrick hanks, and donald hindle (1991). using statistics in lexical analysis. in uri zernik (ed.) lexical acquisition: exploiting on-line resources to build up a lexicon: 115-164. hillsdale, nj, lawrence erlbaum associates. susan conrad and douglas biber (2000). adverbial marking of stance in speech and writing. in susan hunston and geoff thompson (eds.), evaluation in text: authorial stance and the construction of discourse: 56-73. oxford, oxford university press. liesbeth degand and henk l.w. pander maat (2003). a contrastive study of dutch and french causal connectives on the speaker involvement scale. in arie verhagen and jereon van de weijer (eds.), usage based approaches to dutch: 175-199. utrecht, lot. suzanne eggins and diana slade (1997). analysing casual conversation. london/ washington, cassell. stefan evert (2008). corpora and collocations. in anke lüdeling and merja kytö (eds.), corpus linguistics. an international handbook: 1212-1248. berlin, mouton de gruyter. edward finegan (1995). subjectivity and subjectivisation: an introduction. in dieter stein and susan wright (eds.), subjectivity and subjectivisation: linguistic perspectives: 1-15. cambridge, cambridge university press. william foley and robert van valin (1984). functional syntax and universal grammar. cambridge: cambridge university press. austin f. frank and t. florian jaeger (2008). speaking rationally: uniform information density as an optimal strategy for language production. in bradley c. love, kateri mcrae, and vladimir m. sloutsky (eds.), proceedings of the 30th annual conference of the cognitive science society, pages 933-938, cognitive science society, london. arthur c. graesser and danielle s. mcnamara (2011). computational analyses of multilevel discourse comprehension. topics in cognitive science, 3(2), 371-398. stefan th. gries (2013). 50-something years of work on collocations. international journal of corpus linguistics, 18(1), 137-166. wei, evers-vermeul and sanders 86 stefan th. gries and anatol stefanowitsch (2004). extending collostructional analysis: a corpusbased perspective on ‘alternations’. international journal of corpus linguistics, 9(1), 97129. stefan th. gries and anatol stefanowitsch (2010). cluster analysis and the identification of collexeme classes. in sally rice and john newman (eds.), empirical and experimental methods in cognitive/functional research: 73-90. stanford, csli publications. laurence r. horn (1984). toward a new taxonomy for pragmatic inference. in d. schiffrin (ed.), meaning, form, and use in context: linguistic applications: 11-42. washington, d.c., georgetown university press. t. florian jaeger (2010). redundancy and reduction: speakers manage syntactic information density. cognitive psychology, 61(1), 23-62. jet hoek (2018). making sense of discourse: on discourse segmentation and the linguistic marking of coherence relations. phd thesis, utrecht university. utrecht, lot. available online at https://www.lotpublications.nl/documents/509_fulltext.pdf. jet hoek, sandrine zufferey, jacqueline evers-vermeul, and ted j.m. sanders (2018). the linguistic marking of coherence relations: interactions between connectives and segmentinternal elements. pragmatics and cognition, 25(2), 275-309. p.n. johnson-laird and marco ragni (2019). possibilities as the foundation of reasoning. cognition, 193, 103950. ronald w. langacker (1990). subjectification. cognitive linguistics, 1(1), 5-38. roger levy and t. florian jaeger (2007). speakers optimize information density through syntactic reduction. in bernhard schölkopf, john platt, and thomas hofmann (eds.), advances in neural information processing systems (1st ed.): 849-856. cambridge, mass, mit press. fang li, jacqueline evers-vermeul, and ted j.m. sanders (2013). subjectivity and result marking in mandarin: a corpus-based investigation. chinese language and discourse, 4(1), 74-119. yahui liu and xiaopeng yao (2011). kejian de qingtaihua yu guanlianhua – jian lun hanyu liang lei shijueci de yanhua chayi [on the modalitization and relevantization of ‘kejian’]. hanyu xuebao [chinese linguistics], 4, 27-33. john lyons (1977). semantics. cambridge, cambridge university press. willem m. mak and ted j.m. sanders (2010). incremental discourse processing: how coherence relations influence the resolution of pronouns. in martin everaert, tom lentz, hannah de mulder, øystein nilsen, and arjen zondervan (eds.), the linguistics enterprise: from knowledge of language to knowledge in linguistics: 167-182. amsterdam, john benjamins. james r. martin (2000). beyond exchange: appraisal systems in english. in susan hunston and geoff thompson (eds.), evaluation in text: authorial stance and the construction of discourse: 56-73. oxford, oxford university press. henk l.w. pander maat and ted j.m. sanders (2000). domains of use or subjectivity? the distribution of three dutch causal connectives explained. in elizabeth couper-kuhlen and bernd kortmann (eds.), cause, condition, concession and contrast: cognitive and discourse perspectives: 57-82. amsterdam, john benjamins. r core team. (2015). r: a language and environment for statistical computing (version 3.1.3) [computer software]. vienna, austria, r foundation for statistical computing. retrieved from: http://www.r-project.org/ ted j.m. sanders and leo g.m. noordman (2000). the role of coherence relations and their linguistic markers in text processing. discourse processes, 29(1), 37-60. josé sanders and gisela redeker (1996). perspective and the representation of speech and thought in narrative discourse. in gilles fauconnier and eve e. sweetser (eds.) spaces, worlds, and grammar: 290–317. chicago, university of chicago press. https://www.lotpublications.nl/documents/509_fulltext.pdf the use of perspective markers and connectives in expressing subjectivity 87 ted j.m. sanders, josé sanders, and eve e. sweetser (2009). causality, cognition and communication: a mental space analysis of subjectivity in causal connectives. in ted j.m. sanders and eve e. sweetser (eds.), causal categories in discourse and cognition: 19-60. berlin, mouton de gruyter. josé sanders, ted j.m. sanders, and eve e. sweetser (2012). responsible subjects and discourse causality. how mental spaces and perspective help identifying subjectivity in dutch backward causal connectives. journal of pragmatics, 44(2), 191-213. ted j.m. sanders and wilbert p.m.s. spooren (2007). discourse and text structure. in dirk geeraerts and hubert cuyckens (eds.), handbook of cognitive linguistics: 916-941. oxford, oxford university press. ted j.m. sanders and wilbert p.m.s. spooren (2015). causality and subjectivity in discourse: the meaning and use of causal connectives in spontaneous conversation, chat interactions and written text. linguistics, 53(1), 53-92. ted j.m. sanders, wilbert p.m.s. spooren, and leo g.m. noordman (1993). coherence relations in a cognitive theory of discourse representation. cognitive linguistics, 4(2), 93-133. joost schilperoord and arie verhagen (1998). conceptual dependency and the clausal structure of discourse. in j.-p. koenig (ed.), discourse and cognition: bridging the gap: 141-163. chicago, university of chicago press. wilbert p.m.s. spooren, ted j.m. sanders, mike huiskes, and liesbeth degand (2010). subjectivity and causality: a corpus study of spoken language. in sally rice and john newman (eds.), empirical and experimental methods in cognitive/functional research: 241255. stanford, csli publications. dirk speelman (2018). mastering corpus linguistics methods: a practical introduction with antconc and r. hoboken, nj: wiley. anatol stefanowitsch, and stefan th. gries (2003). collostructions: investigating the interaction of words and constructions. international journal of corpus linguistics, 8(2), 209-243. ninke stukker and ted j.m. sanders (2008). another(‘s) perspective on subjectivity in causal connectives: a usage-based analysis of volitional causal relations. in liesbeth degand, cathrine fabricius-hansen, and wiebke ramm (eds.), linearisation and segmentation in discourse: 113-122. oslo, oslo university. geoff thompson and susan hunston (2000). evaluation: an introduction. in susan hunston and geoff thompson (eds.), evaluation in text: authorial stance and the construction of discourse: 1-27. oxford, oxford university press. matthew j. traxler, anthony j. sanford, joy p. aked, and linda m. moxey (1997). processing causals and diagnostics in discourse. journal of experimental psychology: learning, memory and cognition, 23(1), 88-101. gerdineke van silfhout, jacqueline evers-vermeul, willem m. mak, and ted j.m. sanders (2014). connectives and layout as processing signals: how textual features affect students’ processing and text representation. journal of educational psychology, 106(4), 1036-1048. gerdineke van silfhout, jacqueline evers-vermeul, and ted j.m. sanders (2015). connectives as processing signals: how students benefit in processing narrative and expository texts. discourse processes, 52(1), 47-76. jean-christophe verstraete (2001). subjective and objective modality: interpersonal and ideational functions in the english modal auxiliary system. journal of pragmatics, 33(10), 1505-1528. arie verhagen (2005). constructions of intersubjectivity: discourse, syntax, and cognition. oxford, oxford university press. weidong zhan, rui guo, and yirong chen (2003), the ccl corpus of chinese texts: 700 million chinese characters, the 11th century b.c. present, available online at the website of center for chinese linguistics (abbreviated as ccl) of peking university, http://ccl.pku.edu.cn:8080/ccl_corpus. wei, evers-vermeul and sanders 88 hua-ping zhang, hong-kui yu, de-yi xiong, and qun liu. (2003). hhmm-based chinese lexical analyzer nlpir-ictclas. in proceeding of the second sighan workshop on chinese language processing, pages 184-187, sapporo, japan. qi zhang (2012). ‘kejian’ de yufahua [the grammaticalization of ‘kejian’]. journal of hubei university of education, 29(5), 28-32. the use of perspective markers and connectives in expressing subjectivity: evidence from collocational analyses 1 introduction 2 method 2.1 sample of texts 2.2 sample of connective fragments 2.3 three sets of distinctive collocates analyses 3 results 3.1 general analysis 3.2 collocational analysis on different clauses 3.3 collocational analysis on different genres 4 general conclusion and discussion 4.1 general collocation patterns in line with pragmatic principles and uid 4.2 collocation patterns in different genres and clauses 4.3 future studies and conclusion acknowledgements appendix. top 100 collocates per analysis references dialogue & discourse 7(3) (2016) 1–3 doi: 10.5087/dad.2016.305 introduction to the special issue on dialog state tracking jason d. williams jason.williams@microsoft.com microsoft research one microsoft way redmond, wa, 98052, usa antoine raux araux@fb.com facebook 1 facebook way menlo park, ca 94025, usa matthew henderson matthen@google.com google 1600 amphitheatre parkway mountain view, ca 94043, usa editor: david schlangen in the core of most task-oriented conversation systems is a component called a dialog state tracker, which estimates the user’s goal given all of the dialog history so far. for example, in a weather information system, the dialog state might indicate the location the user is interested in (seattle, london, beijing), and for which date (today, tomorrow, this saturday). dialog state tracking is difficult because automatic speech recognition (asr) and spoken language understanding (slu) errors are common, and can cause the system to misunderstand the user. even so, state tracking is crucial because the system relies on the estimated dialog state to choose actions – for example, which weather forecast to provide. as conversational systems become a part of daily life – examples include apple’s siri, google now, and cortana from microsoft – interest from the research community in dialog state tracking has grown considerably. yet to date, work in dialog state tracking has been spread among a number of workshops and conferences. with this backdrop, this special issue has sought to provide thorough coverage of the state-of-the-art in this specialized field. of 7 papers submitted, 4 were ultimately accepted for publication. one of the main catalysts for increased in dialog state tracking recently has been a series of shared tasks called the dialog state tracking challenge, and all papers in this special issue draw on these tasks. in these challenge tasks, detailed logs of human-computer dialogs are released, and teams build trackers that attempt to infer the correct dialog state at each turn. results from all the systems that participate in each challenge have been published in special sessions at conferences and workshops (williams et al., 2013; henderson et al., 2014b,a). the first paper reviews work in dialog state tracking generally, then provides a cohesive review of the first three instances of the dialog state tracking challenge, including a detailed descrition of the data, a compilation of results, and an analysis of the technical advances resulting from the first three challenges. the second, third, and fourth papers explore emerging methods for dialog state tracking in some depth. c©2016 jason d. williams, antoine raux, and matthew henderson this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). williams, raux, and henderson the second paper, spectral decomposition method of dialog state tracking via collective matrix factorization by julien perez, observes that the task of dialog state tracking can be cast as a matrix factorization problem. this novel approach yields state-of-the-art performance on dialog data from the second dialog state tracking challenge. the third paper, dialog history construction with long-short term memory for robust generative dialog state tracking by byung-jun lee and kee-eung kim, adopts an established generativestyle state update, but estimates two of the core components using a “long short-term memory” (lstm) neural network. the authors observe state-of-the-art performance on dialog data from the first and second dialog state tracking challenge. the fourth paper, recurrent polynomial network for dialogue state tracking by kai sun, qizhe xie, and kai yu, introduces a computational recurrent network archicture based on polynomials. the advantage of this approach is that it allows expert knowledge to be encoded in the network, which substantially reduces the search space and thus the number of dialogs required to train the model. the authors observe state-of-the-art performance on dialog data from the second and third dialog state tracking challenges. together these three papers provide a concrete illustration of the roles of machine learning – including both generative and discriminative methods – and the role of expert knowledge, for example hand-crafted rules. we thank prof david schlangen, editor-in-chief of this journal, for his constant support and attentiveness, as well as the editorial board of the dialogue and discourse journal for adopting this special issue. we also thank the reviewers very much for their careful work. as this special issue goes to press, we note that the dialog state tracking challenge series has continued, with the fourth dialog state tracking challenge having recently concluded kim et al. (2016), and planning for the fifth and sixth underway now. this is an exciting period for spoken dialogue systems, both in terms of theoretical advances and commercial deployments. we hope this special issue provides a useful platform for its papers’ authors, and helps support the continuing advance of the state-of-the-art in this domain. jason d. williams, microsoft research antoine raux, facebook mathew henderson, google references matthew henderson, blaise thomson, and jason d. williams. the third dialog state tracking challenge. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014a. matthew henderson, blaise thomson, and jason d williams. the second dialog state tracking challenge. in proc sigdial conf on discourse and dialogue, philadelphia, usa, 2014b. seokhwan kim, luis fernando dharo, rafael e. banchs, jason d. williams, and matthew henderson. the fourth dialog state tracking challenge. in proc intl workshop on spoken dialog systems (iwsds), saariselka, finland, 2016. 2 introduction to the special issue on dialog state tracking jason d williams, antoine raux, deepak ramachadran, and alan black. the dialog state tracking challenge. in proc sigdial conf on discourse and dialogue, metz, france, august 2013. 3 dialogue & discourse 7(3) (2016) 34 –46 doi: 10.5087/dad.2016.304 spectral decomposition method of dialog state tracking via collective matrix factorization julien perez julien.perez@xrce.xerox.com xerox research center europe grenoble, 38000, france editor: jason d. williams, antoine raux, and matthew henderson submitted 04/15; accepted 02/16; published online 04/16 abstract the task of dialog management is commonly decomposed into two sequential subtasks: dialog state tracking and dialog policy learning. in an end-to-end dialog system, the aim of dialog state tracking is to accurately estimate the true dialog state from noisy observations produced by the speech recognition and the natural language understanding modules. the state tracking task is primarily meant to support a dialog policy. from a probabilistic perspective, this is achieved by maintaining a posterior distribution over hidden dialog states composed of a set of context dependent variables. once a dialog policy is learned, it strives to select an optimal dialog act given the estimated dialog state and a defined reward function. this paper introduces a novel method of dialog state tracking based on a bilinear algebric decomposition model that provides an efficient inference schema through collective matrix factorization. we evaluate the proposed approach on the second dialog state tracking challenge (dstc-2) dataset and we show that the proposed tracker gives encouraging results compared to the state-of-the-art trackers that participated in this standard benchmark. finally, we show that the prediction schema is computationally efficient in comparison to the previous approaches. 1. introduction the field of autonomous dialog systems is rapidly growing with the spread of smart mobile devices but it still faces many challenges to become the primary user interface for natural interaction through conversations. indeed, when dialogs are conducted in noisy environments or when utterances themselves are noisy, correctly recognizing and understanding user utterances presents a real challenge. in the context of call-centers, efficient automation has the potential to boost productivity through increasing the probability of a call’s success while reducing the overall cost of handling the call. one of the core components of a state-of-the-art dialog system is a dialog state tracker. its purpose is to monitor the progress of a dialog and provide a compact representation of past user inputs and system outputs represented as a dialog state. the dialog state encapsulates the information needed to successfully finish the dialog, such as users’ goals or requests. indeed, the term “dialog state” loosely denotes an encapsulation of user needs at any point in a dialog. obviously, the precise definition of the state depends on the associated dialog task. an effective dialog system must include a tracking mechanism which is able to accurately accumulate evidence over the sequence of turns of a dialog, and it must adjust the dialog state according to its observations. in that sense, it is an essenc©2016 julien perez this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). tial componant of a dialog systems. however, actual user utterances and corresponding intentions are not directly observable due to errors from automatic speech recognition (asr) and natural language understanding (nlu), making it difficult to infer the true dialog state at any time of a dialog. a common method of modeling a dialog state is through the use of a slot-filling schema, as reviewed in williams and young (2007). in slot-filling, the state is composed of a predefined set of variables with a predefined domain of expression for each of them. the goal of the dialog system is to efficiently instantiate each of these variables thereby performing an associated task and satisfying the corresponding intent of the user. various approaches have been proposed to define dialog state trackers. the traditional methods used in most commercial implementations use hand-crafted rules that typically rely on the most likely result from an nlu module as described in yeh et al. (2014). however, these rule-based systems are prone to frequent errors as the most likely result is not always the correct one. moreover, these systems often force the human customer to respond using simple keywords and to explicitly confirm everything they say, creating an experience that diverges considerably from the natural conversational interaction one might hope to achieve as recalled in williams (2014). more recent methods employ statistical approaches to estimate the posterior distribution over the dialog states allowing them to represent the uncertainty of the results of the nlu module. statistical dialog state trackers are commonly categorized into one of two approaches according to how the posterior probability distribution over the state calculation is defined. in the first type, the generative approach uses a generative model of the dialog dynamic that describes how the sequence of utterances are generated by using the hidden dialog state and using bayes’ rule to calculate the posterior distribution of the state. it has been a popular approach for statistical dialog state tracking, since it naturally fits into the partially observable markov decision process (pomdp) models as described in young et al. (2013), which is an integrated model for dialog state tracking and dialog strategy optimization. using this generic formalism of sequential decision processes, the task of dialog state tracking is to calculate the posterior distribution over an hidden state given an history of observations. in the second type, the discriminative approach models the posterior distribution directly through a closed algebraic formulation as a loss minimization problem. statistical dialog systems, in maintaining a distribution over multiple hypotheses of the true dialog state, are able to behave robustly even in the face of noisy conditions and ambiguity. in this paper, a statistical type of approach of state tracking is proposed by leveraging the recent progress of spectral decomposition methods formalized as bilinear algebraic decomposition and associated inference procedures. the proposed model estimates each state transition with respect to a set of observations and is able to compute the state transition through an inference procedure with a linear complexity with respect to the number of variables and observations. roadmap: this paper is structured as follows, section 2 formally defines transactional dialogs and describes the associated problem of statistical dialog state tracking with both the generative and discriminative approaches. section 3 depicts the proposed decompositional model for coupled and temporal hidden variable models and the associated inference procedure based on collective matrix factorization (cmf). finally, section 4 illustrates the approach with experimental results obtained using a state of the art benchmark for dialog state tracking. 2. transactional dialog state tracking the dialog state tracking task we consider in this paper is formalized as follows: at each turn of a task-oriented dialog between a dialog system and a user, the dialog system chooses a dialog act d to express and the user answers with an utterance u. the dialog state at each turn of a given dialog is defined as a distribution over a set of predefined variables, which define the structure of the state as mentioned in williams et al. (2005). this classic state structure is commonly called slot filling and the associated dialogs are commonly referred to as transactional. indeed, in this context, the state tracking task consists of estimating the value of a set of predefined variables in order to perform a procedure or transaction which is, in fact, the purpose of the dialog. typically, the nlu module processes the user utterance and generates an n-best list o = {< d1, f1 >, . . . , < dn, fn >}, where di is the hypothesized user dialog act and fi is its confidence score. in the simplest case where no asr and nlu modules are employed, as in a text based dialog system as proposed in henderson et al. (2013) the utterance is taken as the observation using a so-called bag of words representation. if an nlu module is available, standardized dialog act schemas can be considered as observations as in bunt et al. (2010). furthermore, if prosodic information is available by the asr component of the dialog system as in milone and rubio (2003), it can also be considered as part of the observation definition. a statistical dialog state tracker maintains, at each discrete time step t, the probability distribution over states, b(st), which is the system’s belief over the state. the general process of slot-filling, transactional dialog management is summarized in figure 1. first, intent detection is typically an nlu problem consisting of identifying the task the user wants the system to accomplish. this first step determines the set of variables to instantiate during the second step, which is the slotfilling process. this type of dialog management assumes that a set of variables are required for each predefined intention. the slot filling process is a classic task of dialog management and is composed of the cyclic tasks of information gathering and integration, in other words – dialog state tracking. finally, once all the variables have been correctly instantiated, a common practice in dialog systems is to perform a last general confirmation of the task desired by the user before finally executing the requested task. as an example used as illutration of the proposed method in this paper, in the case of the dstc-2 challenge, presented in henderson et al. (2014b), the context was taken from the restaurant information domain and the considered variables to instanciate as part of the state are {area (5 possible values) ; food (91 possible values) ; name (113 possible values) ; pricerange (3 possible values)}. in such framework, the purpose is to estimate as early as possible in the course of a given dialog the correct instantiation of each variable. in the following, we will assume the state is represented as a concatenation of zero-one encoding of the values for each variable defining the state. furthermore, in the context of this paper, only the bag of words has been considered as an observation at a given turn but dialog acts or detected named entity provided by an slu module could have also been incorporated as evidence. two statistical approaches have been considered for maintaining the distribution over a state given sequential nlu output. first, the discriminative approach aims to model the posterior probability distribution of the state at time t + 1 with regard to state at time t and observations z1:t. second, the generative approach attempts to model the transition probability and the observation probability in order to exploit possible interdependencies between hidden variables that comprise the dialog state. figure 1: prototypical transactional dialog management process, also called slot-filling dialog management figure 2: generative dialog state tracking using a factorial hmm 2.1 generative dialog state tracking a generative approach to dialog state tracking computes the belief over the state using bayes’ rule, using the belief from the last turn b(st−1) as a prior and the likelihood given the user utterance hypotheses p(zt|st), with zt the observation gathered at time t. in the prior work williams et al. (2005), the likelihood is factored and some independence assumptions are made: bt ∝ ∑ st−1,zt p(st|zt, dt−1, st−1)p(zt|st)b(st−1) (1) figure 2 depicts a typical generative model of a dialog state tracking process using a factorial hidden markov model proposed by ghahramani and jordan (1997). the shaded variables are the observed dialog turns and each unshaded variable represents a single variable describing the task dependent variables. in this family of approaches, scalability is considered as one of the main issues. one way to reduce the amount of computation is to group the states into partitions, as proposed in the hidden information state (his) model of gasic and young (2011). other approaches to cope with the scalability problem in dialog state tracking is to adopt a factored dynamic bayesian network by making conditional independence assumptions among dialog state components, and then using approximate inference algorithms such as loopy belief propagation as proposed in thomson and young (2010) or a blocked gibbs sampling as in raux and ma (2011). to cope with such limitations, discriminative methods of state tracking presented in the next part of this section aim at directly model the posterior distribution of the tracked state using a choosen parametric form. 2.2 discriminative dialog state tracking the discriminative approach of dialog state tracking computes the belief over a state via a trained parametric model that directly represents the belief b(st+1) = p(ss+1|st, zt). maximum entropy has been widely used in the discriminative approach as described in metallinou et al. (2013). it formulates the belief as follows: b(s) = p (s|x) = η.ew tφ(x,s) (2) where η is the normalizing constant, x = (du1 , d m 1 , s1, . . . , d u t , d m t , st) is the history of user dialog acts, dui , i ∈ {1, . . . , t}, the system dialog acts, dmi , i ∈ {1, . . . , t}, and the sequence of states leading to the current dialog turn at time t. then, φ(.) is a vector of feature functions on x and s, and finally, w is the set of model parameters to be learned from annotated dialog data. according to the formulation, the posterior computation has to be carried out for all possible state realizations in order to obtain the normalizing constant η. this is not feasible for real dialog domains, which can have a large number of variables and possible variable instantiations. so, it is vital to the discriminative approach to reduce the size of the state space. for example, metallinou et al. (2013) proposes to restrict the set of possible state variables to those that appeared in nlu results. more recently, lee et al. (2013) assumes conditional independence between dialog state variables to address scalability issues and uses a conditional random field to track each variable separately. finally, deep neural models, performing on a sliding window of features extracted from previous user turns, have also been proposed in henderson et al. (2014c). of the current literature, this family of approaches have proven to be the most efficient for publicly available state tracking datasets. in the next section, we present a decompositional approach of dialog state tracking that aims at reconciling the two main approaches of the state of the art while leveraging on the current advances of low-rank bilinear decomposition models, as recalled in ma et al. (2014), that seems particularly adapted to the sparse nature of dialog state tracking tasks. 3. spectral decomposition model for state tracking in slot-filling dialogs in this section, the proposed model is presented and the learning and prediction procedures are detailed. the general idea consists in the decomposition of a matrix m , composed of a set of turn’s transition as rows and sparse encoding of the corresponding feature variables as columns. more precisely, a row of m is composed with the concatenation of the sparse representation of (1) st, a state at time t (2) st+1, a state at time t + 1 (3) zt, a set of feature representating the observation. in the considered context, the bag of words composing the current turn is chosen as the observation. the parameter learning procedure is formalized as a matrix decomposition task solved through alternating least square ridge regression. the ridge regression task allows for an asymmetric penalization of the targeted variables of the state tracking task to perform. figure 3 illustrates the collective matrix factorization task that constitutes the learning procedure of the state tracking model. the model introduces the component of the decomposed matrix to the form of latent variables {a,b,c}, also called embeddings. in the next section, the learning procedure from dialog state transition data and the proper tracking algorithm are described. in other terms, each row of the matrix corresponds to the concatenation of a ”one-hot” representation of a state description at time t and a dialog turn at time t and each column of the overall matrix m corresponds to a consider feature respectively of the state and dialog turn. such type of modelization of the state tracking problem presents several advantages. first, the model is particularly flexible, the definition of the state and observation spaces are independent of the learning and prediction models and can be adapted to the context of tracking. second, a bias by data can be applied in order to condition the transition model w.r.t separated matrices to decompose jointly as often proposed in multi-task learning as described in caruana (1996) and collective matrix factorization as detailed in kumar bokde et al. (2015). finally, the decomposition method is fast and parallelizable because it mainly leverages on core methods of linear algebra. from our knowledge, this proposition is the first attend to formalize and solve the state tracking task using a matrix decomposition approach. figure 3: spectral state tracking, collective matrix factorization model as inference procedure 3.1 learning method for the sake of simplicity, the {b,c} matrices are concatenated to e, and m is the concatenation of the matrices {st, st+1, zt} depicted in figure 3. equation 3 defines the optimization task, i.e. the loss function, associated with the learning problem of latent variable search {a,e}. min a,e ||(m −ae)w ||22 + λa||a||22 + λb||e||22 , (3) where {λa, λb} ∈ r2 are regularization hyper-parameters and w is a diagonal matrix that increases the weight of the state variables, st+1 in order bias the resulting parameters {a,e} toward better predictive accuracy on these specific variables. this type of weighting approach has been shown to be as efficient in comparable generative discriminative trade-off tasks as mentioned in ulusoy and bishop (2006) and lasserre and bishop (2007). an alternating least squares method that is a sequence of two convex optimization problems is used in order to perform the minimization task. first, for known e, compute: a∗ = argmin a ||(m −ae)w ||22 + λa||a||22 , (4) then for a given a, e∗ = argmin e ||(m −ae)w ||22 + λb||e||22 (5) by iteratively solving these two optimization problems, we obtain the following fixed-point regularized and weighted alternating least square algorithms where t correspond to the current step of the overall iterative process: at+1 ← (et t wet + λai)−1et t wm (6) et+1 ← (at t at + λbi)−1at t m (7) as presented in equation 6, thew matrix is only involved for the updating ofa because only the subset of the columns of e, representing the features of the state to predict, are weighted differently in order to increase the importancd of the corresponding columns in the loss function. for the optimization of the latent representation composing e, presented in equation 7, each call session’s embeddings stored in a hold the same weight, so in this second step of the algorithm, w is actually an identity matrix and so does not appear. 3.2 prediction method the prediction process consists of (1) computing the embedding of a current transition by solving the corresponding least square problem based on the two variables {st, zt} that correspond to our current knowledge of the state at time t and the set of observations extracted from the last turn that is composed with the system and user utterances, (2) estimating the missing values of interest, i.e. the likelihood of each value of each variable that constitutes the state at time (t+ 1), st+1, by computing the cross-product between the transition embedding calculated in (1) and the corresponding column embeddings of e, and of the value of each variable of st+1. more precisely, we write this decomposition as m = a.et (8) where m is the matrix of data to decompose and . the matrix-matrix product operator. as in the previous section, a has a row for each transition embedding, and e has a column for each variablevalue embedding in the form of a zero-one encoding. when a new row of observationsmi for a new set of variables state si and observations zi and e is fixed, the purpose of the prediction task is to find the row ai of a such that: ai.e t ≈ mt i (9) even if it is generally difficult to require these to be equal, we can require that these last elements have the same projection into the latent space: ati .e t .e = mt i .e (10) then, the classic closed form solution of a linear regression task can be derived: ati = mt i .e.(e t .e)−1 (11) ai = (et .e)−1.et .mi (12) in fact, equation 11 is the optimal value of the embedding of the transition mi, assuming a quadratic loss is used. otherwise it is an approximation, in the case of a matrix decomposition of m using a logistic loss for example. note that, in equation 11, (et .e)−1 requires a matrix inversion, but for a low dimensional matrix (the size of the latent space). several advantages can be identified in this approach. first, at learning time, alternative ridge regression is computationally efficient because a closed form solution exists at each step of the optimization process employed to infer the parameters, i.e the low rank matrices, of the model. second, at decision time, the state tracking procedure consists of (1) computing the embedding a of the current transition using the current state estimation st and the current observation set zt and (2) computing the distribution over the state defined as a vector-matrix product between a and the latent matrix e. finally, this inference method can be partially associated to the general technique of matrix completion. but, a proper matrix completion task would have required a matrix m with missing value corresponding to the exhausive list of the possible triples st, st+1, zt, which is obviously intractable to represent and decompose. 4. experimental settings and evaluation in a first section, the dialog domain used for the evaluation of our dialog tracker is described and the different probability models used for the domain. in a second section, we present a first set of experimental results obtained through the proposed approach and its comparison to several reported results of approaches of the state of the art. 4.1 restaurant information domain we used the dstc-2 dialog domain as described in williams et al. (2013) in which the user queries a database of local restaurants by interacting with a dialog system. the dataset for the restaurant information domain were originally collected using amazon mechanical turk. a usual dialog proceeds as follows: first, the user specifies his personal set of constraints concerning the restaurant he looks for. then, the system offers the name of a restaurant that satisfies the constraints. user then accepts the offer, and requests for additional information about accepted restaurant. the dialog ends when all the information requested by the user are provided. in this context, the dialog state tracker should be able to track several types of information that composes the state like the geographic area, tracker joint goal baseline 0.69 focus 0.74 hwu 0.75 hwu+ 0.72 rule-based 0.73 maxent 0.72 rnn 0.75 cmf-350 0.79± 0.03 table 1: accuracy of the proposed model on the dstc-2 test-set the food type, the name and the price range slots. in this paper, we restrict ourselves to tracking these variables, but our tracker can be easily setup to track others as well if they are properly specified. the dialog state tracker updates its belief turn by turn, receiving evidence from the nlu module with the actual utterance produced by the user. in this experiment, it has been chosen to restrict the output of the nlu module to the bag of word of the user utterances in order to be comparable the most recent approaches of state tracking like proposed in henderson et al. (2013) that only use such information as evidence. one important interest in such approach is to dramatically simplify the process of state tracking by suppressing the nlu task. in fact, nlu is mainly formalized in current approaches as a supervised learning approach. the task of the dialog state tracker is to generate a set of possible states and their confidence scores for each slot, with the confidence score corresponding to the posterior probability of each variable state w.r.t the current estimation of the state and the current evidence. finally, the dialog state tracker also maintains a special variable state, called none, which represents that a given variable composing the state has not been observed yet. for the rest of this section, we present experimental results of state tracking obtained in this dataset and we compare with state of the art generative and discriminative approaches. 4.2 experimental results as a comparison to the state of the art methods, table 1 presents accuracy results of the best collective matrix factorization model, with a latent space dimension of 350, which has been determined by cross-validation on a development set, where the value of each slot is instantiated as the most probable w.r.t the inference procedure presented in section 3. in our experiments, the variance is estimated using standard dataset reshuffling. the same results are obtained for several state of the art methods of generative and discriminative state tracking on this dataset using the publicly available results as reported in sun et al. (2014). more precisely, as provided by the state-of-the-art approaches, the accuracy scores computes p(s∗t+1|st, zt) commonly name the joint goal. our proposition is compared to the 4 baseline trackers provided by the dstc organisers. they are the baseline tracker (baseline), the focus tracker (focus), the hwu tracker (hwu) and the hwu tracker with original flag set to (hwu+) respectively. then a comparison to a maximum entropy (maxent) proposed in lee and eskenazi (2013) type of discriminative model and finally a deep neural network (dnn) architecture proposed in sun (2014) as reported also in sun et al. (2014) is presented. 5. related work as depicted in section 2, the litterature of the domain can mainly decomposed into three family of approaches, rule-based, generative and discriminative. in previous works on this topics, williams (2007) formally used particle filters to perform inference in a bayesian network modeling of the dialog state, williams (2008) presented a generative tracker and showed how to train an observation model from transcribed data, williams (2010) grouped indistinguishable dialog states into partitions and consequently performed dialog state tracking on these partitions instead of the individual states, thomson and young (2010) used a dynamic bayesian network to represent the dialog model in an approximate form. so, most attention in the dialog state belief tracking literature has been given to generative bayesian network models until recently as proposed in paek and horvitz (2000) and thomson and young (2010). on the other hand, the successful use of discriminative models for belief tracking has recently been reported by williams (2012) and henderson et al. (2013) and was a major theme in the results of the recent edition of the dialog state tracking challenge. in this paper, a latent decomposition type of approach is proposed in order to address this general problem of dialog system. our method gives encouraging results in comparison to the state of the art dataset and also does not required complex inference at test time because, as detailed in section 3, the tracking algorithm hold a linear complexity w.r.t the sum of realization of each considered variables defining the state to track which is what we believe is one of the main advantage of this method. secondly collective matrix factorization paradigm also for data fusion and bias by data type of modeling as successfully performed in matrix factorization based recommender systems koren et al. (2009). 6. conclusion in this paper, a methodology and algorithm for efficient state tracking in the context of slot-filling dialogs has been presented. the proposed probabilistic model and inference algorithm allows efficient handling of dialog management in the context of classic dialog schemes that constitute a large part of task-oriented dialog tasks. more precisely, such a system allows efficient tracking of hidden variables defining the user goal using any kind of available evidence, from utterance bagof-words to the output of a natural language understanding module. our current investigation on this subject are the beneficiary of distributional word representation as proposed in mikolov et al. (2013) to cope with the question of unknown words and unknown slots as suggested in henderson et al. (2014a). in summary, the proposed approach differentiates itself by the following points from the prior art: (1) by producing a joint probability model of the hidden variable transition in a given dialog state and the observations that allow tracking the current beliefs about the user goals while explicitly considering potential interdependencies between state variables (2) by proposing the necessary computational framework, based on collective matrix factorization, to efficiently infer the distribution over the state variables in order to derive an adequate dialog policy of information seeking in this context. finally, while transactional dialog tracking is mainly useful in the context of autonomous dialog management, the technology can also be used in dialog machine reading and knowledge extraction from human-to-human dialog corpora as proposed in the fourth edition of the dialog state tracking challenge. references the sjtu system for dialog state tracking challenge 2. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial). association for computational linguistics, 2014. harry bunt, jan alexandersson, jean carletta, jae-woong choe, alex chengyu fang, koiti hasida, kiyong lee, volha petukhova, andrei popescu-belis, laurent romary, claudia soria, and david traum. towards an iso standard for dialogue act annotation. in proceedings of the seventh international conference on language resources and evaluation (lrec’10). european language resources association (elra), may 2010. rich caruana. algorithms and applications for multitask learning. in proc. 13th international conference on machine learning, pages 87–95. morgan kaufmann, 1996. milica gasic and steve young. effective handling of dialogue state in the hidden information state pomdp-based dialogue manager. tslp, 7(3):4, 2011. zoubin ghahramani and michael i. jordan. factorial hidden markov models. machine learning, 29(2-3):245–273, 1997. m. henderson, b. thomson, and s. j. young. robust dialog state tracking using delexicalised recurrent neural networks and unsupervised adaptation. in proceedings of ieee spoken language technology, 2014a. matthew henderson, blaise thomson, and steve young. proceedings of the sigdial 2013 conference, chapter deep neural network approach for the dialog state tracking challenge, pages 467–471. association for computational linguistics, 2013. matthew henderson, blaise thomson, and jason d. williams. the second dialog state tracking challenge. in proceedings of sigdial. acl association for computational linguistics, june 2014b. matthew henderson, blaise thomson, and steve young. word-based dialog state tracking with recurrent neural networks. in in proceedings of sigdial, 2014c. yehuda koren, robert m. bell, and chris volinsky. matrix factorization techniques for recommender systems. ieee computer, 42(8):30–37, 2009. dheeraj kumar bokde, sheetal girase, and debajyoti mukhopadhyay. role of matrix factorization model in collaborative filtering algorithm: a survey. corr, abs/1503.07475, 2015. julia lasserre and christopher m. bishop. generative or discriminative? getting the best of both worlds. bayesian statistics, 8:3–24, 2007. donghyeon lee, minwoo jeong, kyungduk kim, seonghan ryu, and gary geunbae lee. unsupervised spoken language understanding for a multi-domain dialog system. ieee transactions on audio, speech & language processing, 21(11):2451–2464, 2013. sungjin lee and maxine eskenazi. recipe for building robust spoken dialog state trackers: dialog state tracking challenge system description. in proceedings of the sigdial 2013 conference, pages 414–422, metz, france, august 2013. association for computational linguistics. rick ma, nafise barzigar, aminmohammad roozgard, and samuel cheng. decomposition approach for low-rank matrix completion and its applications. ieee transactions on signal processing, 62(7):1671–1683, 2014. angeliki metallinou, dan bohus, and jason williams. discriminative state tracking for spoken dialog systems. in association for computer linguistics, pages 466–475. the association for computer linguistics, 2013. tomas mikolov, wen tau yih, and geoffrey zweig. linguistic regularities in continuous space word representations. in lucy vanderwende, hal daumé iii, and katrin kirchhoff, editors, hltnaacl, pages 746–751. the association for computational linguistics, 2013. diego h. milone and antonio j. rubio. prosodic and accentual information for automatic speech recognition. ieee transactions on speech and audio processing, 11(4):321–333, 2003. tim paek and eric horvitz. conversation as action under uncertainty. in uai ’00: proceedings of the 16th conference in uncertainty in artificial intelligence, stanford university, stanford, california, usa, pages 455–464. morgan kaufmann, 2000. antoine raux and yi ma. efficient probabilistic tracking of user goal and dialog history for spoken dialog systems. in interspeech, pages 801–804. isca, 2011. kai sun, lu chen, su zhu, and kai yu. a generalized rule based tracker for dialogue state tracking. in slt, pages 330–335. ieee, 2014. blaise thomson and steve young. bayesian update of dialogue state: a pomdp framework for spoken dialogue systems. computer speech & language, 24(4):562–588, 2010. i. ulusoy and c. m. bishop. comparison of generative and discriminative techniques for object detection and classification. in toward category-level object recognition, pages 173–195, 2006. jason williams, antoine raux, deepak ramachandran, and alan black. the dialog state tracking challenge. in proceedings of the sigdial 2013 conference, pages 404–413, metz, france, august 2013. association for computational linguistics. jason d. williams. using particle filters to track dialogue state. in ieee workshop on automatic speech recognition & understanding, asru 2007, kyoto, japan, december 9-13, 2007, pages 502–507, 2007. jason d. williams. exploiting the asr n-best by tracking multiple dialog state hypotheses. in interspeech, pages 191–194. isca, 2008. jason d. williams. incremental partition recombination for efficient tracking of multiple dialog states. in icassp, pages 5382–5385, 2010. jason d. williams. challenges and opportunities for state tracking in statistical spoken dialog systems: results from two public deployments. j. sel. topics signal processing, 6(8):959–970, 2012. jason d. williams. web-style ranking and slu combination for dialog state tracking. in proceedings of sigdial. acl association for computational linguistics, june 2014. jason d. williams and steve young. partially observable markov decision processes for spoken dialog systems. computer speech & language, 21(2):393–422, 2007. jason d. williams, pascal poupart, and steve young. factored partially observable markov decision processes for dialogue management. in in 4th workshop on knowledge and reasoning in practical dialog systems, pages 76–82, 2005. peter z. yeh, benjamin douglas, william jarrold, adwait ratnaparkhi, deepak ramachandran, peter f. patel-schneider, stephen laverty, nirvana tikku, sean brown, and jeremy mendel. a speech-driven second screen application for tv program discovery. in carla e. brodley and peter stone, editors, proceedings of the twenty-eighth aaai conference on artificial intelligence, july 27 -31, 2014, québec city, québec, canada, pages 3010–3016. aaai press, 2014. isbn 9781-57735-661-5. steve young, milica gasic, blaise thomson, and jason d. williams. pomdp-based statistical spoken dialog systems: a review. proceedings of the ieee, 101(5):1160–1179, 2013. dialogue and discourse 4(2) (2013) 185-214 doi: 10.5087/dad.2013.209 annotation of negotiation processes in joint action dialogues thora tenbrink t.tenbrink@bangor.ac.uk school of linguistics and english language bangor university bangor, gwynedd ll57 2dg, uk kathleen eberhard keberhar@nd.edu department of psychology university of notre dame notre dame, in 46556, usa hui shi shi@informatik.uni-bremen.de fb3 faculty of mathematics and computer science university of bremen 28334 bremen, germany sandra kübler skuebler@indiana.edu department of linguistics indiana university bloomington, in 47405, usa matthias scheutz matthias.scheutz@tufts.edu department of computer science tufts university medford, ma 02155, usa editors: stefanie dipper, heike zinsmeister, bonnie webber abstract situated dialogue corpora are invaluable resources for understanding the complex relationships among language, perception, and action. accomplishing shared goals in the real world can often only be achieved via dynamic negotiation processes based on the interactants’ common ground. in this paper, we investigate ways of systematically capturing structural dialogue phenomena in situated goal-directed tasks through the use of annotation schemes. specifically, we examine how dialogue structure is affected by participants updating their joint knowledge of task states by relying on various non-linguistic factors such as perceptions and actions. in contrast to entirely languagebased factors, these are typically not captured by conventional annotation schemes, although isolated relevant efforts exist. following an overview of previous empirical evidence highlighting effects of multi-modal dialogue updates, we discuss a range of relevant dialogue corpora along with the annotation schemes that have been used to analyze them. our brief review shows how gestures, action, intonation, perception and the like, to the extent that they are available to participants, can affect dialogue structure and directly contribute to the communicative success in situated tasks. accordingly, current annotation schemes need to be extended to fully capture these critical additional aspects of non-linguistic dialogue phenomena, building on existing efforts as discussed in this paper. keywords: annotation schemes; dialogue corpora; situatedness; joint action; dialogue modelling c⃝2013 thora tenbrink et al. submitted 03/12; accepted 12/12; published online 07/13 tenbrink, eberhard, shi, kübler, and scheutz 1. introduction dialogue corpora are of great utility to a diverse set of researchers, ranging from those who intend to test theories of language use in human social situations (e.g., conversation analysis, discourse analysis) to researchers developing artificial agents that can effectively interact with humans via natural language (human-computer and human-robot interaction). while the identification of factors that affect the efficiency of language use are common to both research agendas, the latter community is particularly interested in corpora of task-oriented dialogues given that artificial agents are typically designed to be directed by humans for performing various tasks. over the years, an increasing number of task-oriented dialogue corpora has been developed, with varying task and interaction complexities. among the most long-standing and well-known corpora are the map task corpus (anderson et al., 1991) and the trains corpus (heeman and allen, 1995) which are both based on relatively simple tasks and domains (e.g., drawing or planning routes on a two-dimensional map) with no nonverbal response types. more recent corpora (discussed in detail below) are based on more complex situated tasks and domains where interactants have to manipulate real-world objects, move through environments, and overall perform more naturalistic perceptions and actions. our focus in this paper is on annotation schemes that capture structures of negotiation processes in joint action dialogues, building on speech-act related theories (e.g., allen and perrault, 1980; cohen and levesque, 1980) to describe dialogic actions in discourse (i.e., dialogue acts as specialized speech acts referring to dialogue actions). this structure-oriented approach contrasts with another line of research on automatic dialogue control, which has produced important results for negotiation processes using information state systems. these focus primarily on belief states on the part of each interactant rather than structural patterns, following grosz and sidner (1986). a typical target for research in this area concerns information seeking dialogues (e.g., bohlin et al., 1999; larsson and traum, 2000), addressing the negotiation of alternative solutions (although some applications in multimodal scenarios have been proposed as well, e.g., lemon et al., 2002). in contrast, joint action scenarios do not primarily focus on the exchange of information but rather on achieving shared goals that require physical actions. in such settings, the negotiation of alternative solutions and the conveyance of facts are often less important compared to updating common ground with respect to perceptions, actions, and action outcomes. we will argue that this intricate interplay between perceptions, actions, and linguistic exchanges requires fine-grained annotation schemes in order to model the organization of the dialogue adequately. in section 2, we thus start with a discussion of dialogue aspects that are relevant for annotations of task-oriented dialogue, starting from clark’s influential view (clark, 1996; clark and wilkesgibbs, 1986) of dialogue as a joint activity. in the case of task-oriented settings, the dialogue itself is embedded within a larger joint activity corresponding to the task’s prescribed goals, and the verbal dialogue is a means by which the participants coordinate their task-relevant (non-communicative) actions. however, as our discussion highlights, the relative role of language is affected by the extent to which the interactants share their perceptions of the task domain, of the actions performed within it, and of each other. that is, shared perceptions provide a reliable basis for the interactants to update their common ground, which includes shared knowledge about the task state and the remaining goals. this reduces the need to rely on dialogue for the updating process; hence annotation of nonverbal communicative actions (e.g., gaze, gestures, head nods, etc.) is important for capturing the participants’ coordination processes. conversely, the less shared perception there is, the greater the 186 annotation of joint action dialogues participants’ reliance on dialogue for updating their common ground; this enhances the need for annotating further layers of the dialogue that convey crucial information (e.g., intonational patterns, types of referential expressions, discourse markers, syntactic complexity, types of dialogue acts, etc.). with these considerations in mind, section 3 reviews different annotation schemes proposed for corpora that differ with respect to the level of shared perception as well as the situatedness of the task and the complexity of the task goals. we specifically consider the usability of the annotation schemes for investigating the coordination and negotiation processes, i.e., the ease of updating common ground, in task-oriented dialogue. section 4 then provides a brief discussion of the layers of annotation that are required for fully capturing activity coordination and goal negotiation processes as they occur in joint action scenarios. we conclude by a brief discussion of the challenges ahead and consequences for dialogue annotation and corpus data collection. 2. common ground in task-oriented dialogue we adopt clark’s (1996) view of dialogue as a joint activity, which he defines as involving two or more participants coordinating their actions in order to achieve a mutual dominant goal. in the case of task-oriented dialogue, the dialogue is a joint activity embedded within a larger joint activity corresponding to the prescribed goal of the task. the task’s goal is specified with respect to a circumscribed domain of objects, the actions that are to be performed on the objects, the constraints on the actions, and the participants’ responsibilities in performing the actions. thus, the dialogue is important for coordinating the participants’ performance of non-communicative actions involving task-relevant objects. achieving the task’s dominant goal requires a hierarchy of nested subgoals, and corresponding joint activities. this hierarchy is central to the participants’ common ground, which is the knowledge that they believe they share about their joint activity. at any given point, the common ground includes the joint activity’s initial state, its current state, and the mutually known events and perceptions that have led to the current state (clark, 1996). the updating of common ground is fundamental to the participants’ coordinating both the sequencing and timing of their actions and goals. the ease or accuracy of the updating determines the efficiency with which they perform the overall joint activity, for instance, in terms of accuracy, time needed, number of subgoals that were achieved, etc. an overall aim of analyzing corpora of task-oriented dialogue then is to identify factors that affect the participants’ coordination of their joint activity, and hence the efficiency of their performing the task. because coordination is a dynamic process that involves updating information in common ground, important factors will include those that affect the ease and/or accuracy of the updating. the greatest efficiency occurs with face-to-face interaction in which the participants share perception of each other’s performance of task-relevant actions. this shared perception provides reliable and immediate updating of the task’s status in common ground, with less reliance on dialogue for the updating. moreover, face-to-face interactions enable the use of deictic gestures and other visual cues that can focus the participants’ attention and facilitate coordinated behavior. for example, the ability to monitor each others’ gaze allows for joint attention processes to aid the coordination of the taskrelevant actions (e.g., bangerter, 2004; bard et al., 1996; brennan et al., 2008; neider et al., 2010). in particular, a speaker’s gaze at an object when referring to it can facilitate an addressee’s estab187 tenbrink, eberhard, shi, kübler, and scheutz lishment of reference (e.g., hanna and brennan, 2007), and the addressee’s gaze at an object can provide evidence of his or her understanding to the speaker. the facilitation is reflected in less interruption in speech and fewer turns (boyle et al., 1994; o’malley et al., 1996), due to the availability of non-verbal signals of understanding (e.g., head nods) and the use of mutual gaze for regulating turn-taking (e.g., kendon, 1967; allwood et al., 2007). substantial facilitation can be achieved even when only the interlocutors’ faces are perceivable, e.g., when interactants are co-located but blocked from seeing task-relevant objects and actions (e.g., brennan et al., 2012). however, the beneficial effects of shared face perception do not appear to extend to remote interactants who share face perception via video conferencing. in fact, video conferencing increases the number of turns (o’malley et al., 1996). if a delay of responses is involved, this will affect the dialogue even if it is only slight (sellen, 1995; fischer and tenbrink, 2003). since multimodal layers of interaction thus evidently affect how people perceive a situation and act within it, the structure of interaction should be fundamentally affected by the availability of multimodal channels. however, although dialogic negotiation processes have been approached from various perspectives, to our knowledge no annotation scheme has been proposed that systematically captures negotiation processes and keeps track of common ground via a comprehensive structural integration of non-verbal dialogue contributions. in the following, we examine promising efforts in relevant directions. 3. annotation schemes for various layers of task-oriented corpora task-based dialogues are comprised of many layers of task-relevant information exchanges, including the critical layer of dialogue acts. existing annotation schemes focusing on this aspect include dit++ (bunt, 2009), verbmobil (alexandersson et al., 1998), as well as damsl (allen and core, 1997; jurafsky et al., 1997) and hcrc (carletta et al., 1997). other schemes highlight further layers of dialogues. the well-developed annotation scheme tobi (beckman and hirschberg, 1994; beckman et al., 2005) captures features of intonation (e.g., in the columbia games corpus, gravano, 2009), and nonverbal forms of communication such as gestures are captured, for instance, by form (martell, 2002; martell and kroll, 2007) and the mpi gesture annotation (kipp et al., 2007). the discussion in this section will start with a brief description of the human communication research centre (hcrc) annotation scheme and the dialogue act markup in several layers (damsl) scheme. they were initially developed for the map task corpus (anderson et al., 1991) and the trains corpus (heeman and allen, 1995), respectively, but have also been applied to other corpora, e.g., the maze task (kowtko et al., 1993) and switchboard (jurafsky et al., 1997). while both schemes illustrate the kinds of dialogue act classifications that reflect coordination of communicative actions, they differ with respect to the communication levels they recognize: the hcrc scheme distinguishes transaction, game, and move levels; damsl uses communication management, task management, and task levels. based on an extensive study of coding outcomes using these two annotation schemes, stirling et al. (2001) provide a detailed analysis of the extent to which the schemes can capture relations between dialogue structure and prosody patterns. since we are concerned with joint action tasks in the current paper, our focus lies on individual dialogue moves and task-relevant dialogue acts as part of the negotiation process. also, we focus on the types of phenomena that can potentially be highlighted by annotation schemes, rather than attempting a direct comparison and evaluation based on extensive recording procedures. 188 annotation of joint action dialogues the following subsections provide descriptions of a non-exhaustive list of corpora and annotation schemes. we present examples of how the speakers update common ground, and discuss these in light of annotations of dialogue acts, intonation, and gestures. the corpora were selected in part because they are well-known (in the case of the map task and trains corpus), and in part because they contribute new types of annotation, or new insights concerning the specific phenomena relevant for joint action tasks. we conclude the description of each corpus and annotation example with a brief evaluation of its usability for describing updating mechanisms for common ground in task-oriented scenarios. 3.1 map task corpus and the hcrc coding scheme the map task (anderson et al., 1991) represents a widely used paradigm that has been employed in various versions. typically it involves two participants who both have a map, which the other person cannot see. the maps contain line drawings of labeled objects (e.g., rocks, bridge, mountain, etc.) which can serve as landmarks. the maps are not identical and the participants are informed of this. both maps are marked with a start point and an end point. one person’s map has a line indicating a route from the start point to the end point. the person with this map is the instruction giver, and the other person is the follower. the task is for the follower to draw the same route on his or her map as the one indicated on the giver’s map by following the giver’s instructions. the hcrc annotation scheme1 (carletta et al., 1997) consists of three interdependent levels of coding. the basic dialogue acts are coded as conversational moves, which indicate the communicative purpose of an utterance. each utterance is coded with a single move, which is classified as either initiation or response. initiation moves include instructions (instruct), explanations (explain), and questions (query-yn, query-w) and two types of requests: confirmation of accurate understanding (check) and confirmation of a readiness to continue (align). the initiation moves elicit response moves, which include acknowledgments (acknow), replies to questions (reply-y, reply-n, reply-w), and clarifications (clarify). the game level identifies the set of moves that fulfill the purpose of an overall initiation move. for example, an instruction may initiate a game that is ended by an acknowledgment of the completed action. games can be embedded within games, as in the case when an instruction move is followed by a question-reply sequence (game) reflecting the need for clarifying information before the instruction can be completed. the transaction level distinguishes sets of games that concern performing a task-relevant action from those that concern managing the task (e.g., reviewing transactions, which refer to previously completed actions, and overview transactions, which refer to the overall task). the normal transactions correspond to subtasks (segments) of the overall task. the identification of the subtasks requires taking into account the task-relevant actions performed by the participants. in the case of the map task, the transactions are labeled with the beginning and ending points on the map that correspond to the segments of the route that are incrementally drawn on the map by the follower. table 1 gives an example of a normal transaction consisting of a dialogue sequence that corresponds to the follower marking the first segment of the route on his or her map. the transaction involves two instruction games and an embedded check game in which the follower checks their understanding of the giver’s preceding instruction. instruct, acknow, check, reply-y, clarify are moves. 1. more information about the corpus and the annotation can be found at http://groups.inf.ed.ac.uk/ maptask/. 189 tenbrink, eberhard, shi, kübler, and scheutz id speaker utterance game move utt2 giver starting off we are above a caravan part game 1 (instruct) instruct utt3 follower mmhmm acknow utt4 giver we are going to go due south straight south and then we’re going to g– turn straight back round and head north past an old mill on the right hand side game 2 (instruct) instruct game 3 (check), embedded 1 utt5 follower due south and then back up again check utt6 giver yeah reply-y utt7 giver south and then straight back up again with an old mill on the right and you’re going to pass on the left-hand side of the mill clarify utt8 follower right okay acknow table 1: an example dialogue from the map task corpus. usability for task-oriented scenarios. three hierarchical levels of dialogue structure can be distinguished in this annotation scheme: conversational moves, games, and transactions. the dialogue’s structure closely reflects the task’s goal structure because the dominant goal involves minimal action, which is performed by the follower (i.e., the follower draws a line indicating a route on the map). the follower’s incremental drawing of the route is included as a layer of annotation, which is used to segment the dialogue at the transaction level of coding. this annotation also includes instances when the follower crossed out a portion of the route and drew a new one, reflecting the correction of inaccurate information in common ground. the online corpus also has other layers of linguistic annotation (e.g., disfluencies, part-of-speech tagging) and a layer of gaze annotation for a subset of participants. the latter indicates whether the participant was looking at the map or at the other participant during the task. however, the gaze annotation has not been used to investigate the dynamic process of coordinating the joint actions. rather, the effects of gaze have been investigated by comparing overall differences in various aspects of the dialogue annotations (e.g., number of turns, number of words, frequency of types of disfluencies) in eye-contact vs. no eye-contact conditions (branigan et al., 1999; boyle et al., 1994; bard et al., 2007). 3.2 the trains corpus and damsl annotation scheme the trains corpus (heeman and allen, 1995, 1994) consists of dialogues between two participants in a problem-solving task involving shipping goods by train to various cities. the participants were informed that the goal was to build a computer system that can aid people in problem-solving tasks (see allen et al., 1995; sikorski and allen, 1997; traum, 1996). one participant had the role of system user and the other had the system role. the participants were asked to act normally rather than mimic an automatic system. they were given a map showing engines, box cars, and factories, and their locations along a single rail track. the user was given various tasks (e.g., construct a plan for shipping a box car of oranges to bath by 8 am). the system had further information that was important for the user’s task such as timing and engine capacity. no actions were directly executed in the task. since the knowledge about the task was partially shared and information was distributed between participants, the trains corpus contains rich negotiation dialogues and frequent initiative shifts. 190 annotation of joint action dialogues id speaker utterance dialogue act utt1 user we could use the train at dansville. open-option utt2 system could we use one at avon instead? open-option, hold(utt1) utt3 user no, i want it for something else. assert, reject(utt2) utt4 system how about the one at corning then? open-option, hold(utt1-utt3) utt5 user okay. assert, accept(utt4) utt6 system okay. accept(utt1-utt5) table 2: an example dialogue from the trains corpus. the damsl scheme (core and allen, 1997; allen and core, 1997) involves three orthogonal layers of coding. the layer that codes the communicative purpose or functions of utterances (dialogue acts) has two categories: forward-looking functions and backward-looking functions (similar to the initiation-response dichotomy in the hcrc annotation scheme). forward-looking functions affect the subsequent interaction as in the case of questions (info-request), instructions (action-directive), suggestions (open-option), statements (assert), etc. the backwardlooking functions concern the utterance’s relation to the previous discourse, e.g., answers to questions (answer), signals of agreement (accept) or understanding (acknowledge), etc. an important feature of the coding at this level (contrasting with the hcrc scheme) is that utterances may simultaneously perform more than one function. the second layer is the information level, which distinguishes utterances that are about the performance of task-relevant actions (task), from those that concern planning, coordinating, or clarifying the task-relevant actions (task-management), and those that concern the management of the dialogue itself (e.g., acknowledgments, opening, closings, repairs, signals of delays, etc.). the third layer (communicative status) distinguishes complete utterances from incomplete, abandoned, or unintelligible utterances. table 2 shows an example from the trains corpus (adapted from allen and core, 1997)2 that illustrates the dialogue acts at the communicative purpose layer, which is the most relevant layer for the discussion of functional dialogue structures. the trains scenario differs from the map task in that responsibilities and knowledge are divided between the speakers. although the user is responsible for deciding the planned actions involving which trains pick up, move, and deliver cargo, the system can suggest plans as well as reject plans proposed by the user due to constraints which are known to the system, but not to the user. we provide two examples to illustrate how this particular aspect affects dialogue structure. table 2 shows a portion of a dialogue from the trains corpus that has two embedded subdialogues. the dialogue begins with an open-option tag, which is a forward-looking function corresponding to a suggested course of action. the first embedded subdialogue is tagged by hold(utt1), which is a backward-looking function that indicates a suspension of a response to the suggestion in the first utterance. the suspension is due to the system suggesting an alternative course of action (openoption), which is rejected by the user in utt3. the second embedded subdialogue results from the system suggesting a second alternative course of action in utt4. this suggestion is accepted by the user in utt5, with the system’s acceptance in utt6 closing the task that began with utt1. the more complex example in table 3 further illustrates the negotiation of shared information in this joint task. we take a closer look at the underlying intentions behind dialogue acts as far as they 2. see http://www.cs.rochester.edu/research/cisd/resources/damsl/ for the annotation details. 191 tenbrink, eberhard, shi, kübler, and scheutz id speaker utterance dialogue act utt7 user the orange warehouse where i need the oranges from is in corning assert utt8 system right accept(utt7) utt9 user so i need is it possible for one of the engines would it be faster for an engine to come from elmira or avon info-request utt10 system uh elmira is a lot closer answer(utt9), assert utt11 user what time would engine two and three leave elmira info-request utt12 system um well they’re not scheduled yet answer(utt11), reject(utt11) utt13 system but we can send them at any time we want answer(utt11), open-option utt14 user okay accept(utt13) utt15 system uh so + if we sent them right away it’d get there at at um at two a.m. open-option, offer, assert utt16 user at corning hold(utt15), info-request utt17 system yeah answer(utt16), assert utt18 user and how long would it take to get from corning to bath info-request, commit utt19 system uh two hours answer(utt18), assert utt20 user how long would it take to load the oranges from the warehouse into the engine info-request, commit utt21 system uh well we can’t load oranges into an engine we need a boxcar to load them into answer(utt20), reject(utt20), assert utt22 user mm-hm accept(utt21) utt23 user so can i dispatch an engine and a boxcar from elmira simultaneously to corning info-request, offer utt24 system uh yeah yeah answer(utt23), assert utt25 system we can uh connect an engine to the boxcar and then take have the engine take the boxcar to corning accept, commit, assert utt26 user so it’ll be two hours to corning assert utt27 system right accept(utt26) table 3: an example dialogue from the trains corpus, showing the negotiation of shared information. 192 annotation of joint action dialogues can be inferred from the speakers’ information status. the purpose of the assert dialogue act in utt7 is to convey the information that oranges are needed (which is unknown to the system), while at the same time establishing the orange warehouse in corning as common ground. both aspects of this dual-purpose assertion are accepted by the system in utt8. utt9 contains a request for timing information, which is simply responded to by the system in utt10. the following request in utt11 by the user signals a misconception of the kinds of information available to the system, leading to a rejection in utt12. in utt20, the user poses an information request containing a commitment to the future action “to load the oranges from the warehouse into the engine” in a way that signals (via presupposition) a misconception of the task. this aspect, which is not captured directly by the annotation, is clarified by the system in utt21, leading to a dialogue utterance that is simultaneously a rejection of the request and an assertion of facts, which in the following leads to new considerations and requests for information on the part of the user. altogether, the dialogue is dominated by the need to confirm and exchange information based on what is mutually known or needs to be conveyed to the dialogue partner. the damsl annotation reflects the negotiation identified here to some extent via the following emerging structures: assert and accept (e.g., utt7 and utt8, utt26 and utt27), info-request and answer (e.g., utt9 and utt10, utt18 and utt19), and info-request and answer + reject as a less well established type of adjacency pair (utt11 and utt12, utt20 and utt21) that can be traced back to the fact that information is distributed across participants. usability for task-oriented scenarios. the damsl coding scheme is clearly devised for precisely the kind of scenario found in the trains corpus, which involves no shared action or perception, distributed resources of information on both sides, and a joint task. in contrast to the hcrc scheme, the three layers used in damsl are not hierarchical, but orthogonal. rather than providing further detail about a type of verbal interaction, they highlight three different aspects of the same verbal interaction (forward or backward looking functions, information level, and communicative status). like hcrc, all of the annotation pertains to the verbal interaction, rather than taking further aspects into account such as the checking of information from the database. however, since most of the system’s utterances in table 3 rely on the information solely available in the system’s database, such a database check must have been a frequent action on the part of the interactant playing the part of the system. this affects the development of the dialogue since a substantial part of the dialogue consists of the updating of common ground based on information from external resources available only to one of the interactants, accessed in reaction to the other interactant’s contributions. in other scenarios, existing beliefs may be negotiated in a balanced way, or information updates may be based on changes in the real world. 3.3 schober: accounting for the addressee’s abilities in various papers, schober (1993, 1995, 1998, 2009) addresses how speakers modify their spatial perspectives for the benefit of their addressee with respect to visually shared scenes (seen from different perspectives). schober’s approach is to provide a systematic analysis along with illustrative examples, where annotation is focused on the conceptual aspect of perspective taking (which is of less concern for our current purposes). the dialogue in table 4 is an example from schober (2009, p. 31) and involves speakers with mismatched abilities. we have added a damsl-type annotation to this example in order to highlight the communicative acts. the task (a referential communication task) was to identify one out of four dots arranged around a sketched airplane shown on a screen. 193 tenbrink, eberhard, shi, kübler, and scheutz id speaker utterance dialogue act utt1 director okay my plane is pointing towards the left assert utt2 matcher okay accept(1) utt3 director and the dot is directly at the tail of it assert utt4 director like right at the back of it assert utt5 matcher okay mine is pointing to the right acknowledge(3,4), assert utt6 director oh yours is pointing to the right acknowledge repeat-rephrase(5) utt7 matcher yeah accept(6) utt8 director so your dot should be on the left assert utt9 director because my dot is on the right completion(8) utt10 director in back of it assert utt11 director so your dot should be at the left reassert(8) utt12 director at the back of it right reassert(10) utt13 matcher yeah accept(8,10,11,12) utt14 director yeah accept(13) utt15 matcher but if it is the same but if it the same dot-right? wait a minute, assert, correction, if my your plane is pointing to the left *[something] * comm. manag., repeat-rephrase(1) utt16 director *my* plane is pointing to the left reassert(1) utt17 matcher mm-hm comm. manag. utt18 matcher and that dot and the dot that’s highlighted is the one assert all the way in the back of it utt19 matcher like behind the tail assert utt20 matcher yes, so so my dot is gonna be accept(18,19), assert utt21 director so my dot is on the right reassert utt22 director and yours should be on the left right reassert, completion(20) info-request utt23 matcher yeah accept, answer(22) utt24 director okay *so your * acknowledge(23), assert utt25 matcher *right behind the tail* okay completion, accept(24) utt26 matcher okay comm. manag. table 4: an example dialogue from schober (2009). director and matcher had different views on this scene, but the other person’s view was indicated on the screen. however, this fact did not keep participants from negotiating each other’s view in order to reach their discourse goal. in this example, director and matcher jointly try to agree on which dot is the correct one. this requires some negotiation, to which both interactants contribute although only the director has access to the information about the goal dot. the matcher’s contribution is to provide information about the state of understanding, as if thinking aloud for the benefit of the director, who can use this information to support the thought process. in particular, the director states the direction of the airplane (from his or her point of view) in utt1 in order to describe where the dot is located (utt3). instead of mentally rotating the airplane based on this description and marking the dot on the display, the matcher merely states his or her own view on the scene in utt5. subsequently, the director describes the location of the dot from the matcher’s perspective (utt8 to utt12). the matcher exhibits confusion in utt15 and tries to use the object-centered view to understand the location of the dot. then the director again uses different perspectives (utt21 and utt22) to describe the dot, which is accepted by the matcher. this iterative exchange, required by this particular pair of 194 annotation of joint action dialogues speakers to establish common ground related to their different perception of the scene, is in contrast with the following very simple exchange, which achieves the same subgoal: director: it’s behind the plane. matcher: okay. the negotiation effects are reflected in our damsl-type annotation by frequent cases of reassert, repeat-rephrase, and completion, along with many cases of assert contributed by both speakers. thus, a simple instruction is not deemed sufficient by the speakers in this case; the main point of negotiation is to find out how the two views can be matched so as to interpret the instruction correctly. this (conceptual) task is jointly achieved by both interactants, so that the matcher can follow the instruction of marking the goal dot (which is a simple action once the conceptual matching process has been completed). accordingly, neither instruction nor action are ever explicitly mentioned throughout this dialogue. usability for task-oriented scenarios. this dialogue corpus is particularly well-suited for highlighting how interactants achieve common ground by taking each other’s view on the scene into account. the negotiation of spatial perspectives is tightly integrated in the overall task. by highlighting dialogue acts, the annotation in table 4 reflects the structure of the perspective clarifications needed to achieve common ground about spatial locations. however, a thorough understanding of the interactants’ intentions and conceptual perspectives is only possible by accounting for the associated pictures as perceived by the participants. 3.4 airplane: continuing each other’s thoughts the collaborative research center sfb 360 situated artificial communicators in bielefeld (germany) (rickheit and wachsmuth, 2006) explored a scenario in which participants had to build a toy airplane from its parts3. while this scenario was used to address diverse challenges in the artificial intelligence area including human-robot interaction, we focus here on a human-human dialogue setting involving language-based instruction. participants were separated by a screen and could not see each other. one of them, the instructor, had a completed toy airplane, while the other (the constructor) had a set of separate parts. thus, the setting involved sufficient complexity of actions to involve a high amount of negotiation. poncin and rieser (2006) discuss a brief dialogue example in much detail in order to establish how speakers manage to negotiate actions, and in particular, how the constructor completes some of the instructor’s utterances, shown in table 5 (annotations by original authors, augmented by the dialogue acts in damsl for purposes of comparison). from the example in table 5, it is obvious that the negotiation of actions can lead to a high involvement by the participants with less knowledge, who make informed guesses about the possible steps of action – even to the extent that they complete the instructor’s sentences. according to poncin and rieser (2006) this is only possible because of the information embedded in the directives along with a high amount of shared common ground based on the previous actions and background knowledge. moreover, they point to the important contribution of prosody in the interpretation of the speakers’ joint achievement in this exchange. in particular, “cnst’s completion ‘copies’ the global, rising-falling pitch movement of inst’s preceding utterance”, and “inst’s repair of the proposed 3. the corpora without annotation are available at http://www.sfb360.uni-bielefeld.de/ transkript/. 195 tenbrink, eberhard, shi, kübler, and scheutz id speaker utterance annotation dialogue act utt1 inst so, jetzt nimmst du inst’s proposal action-directive well, now you take utt2 cnst eine schraube, cnst’s acceptance of inst’s proposal and accept(1) a screw, her proposal for a continuation offer utt3 inst eine orangene mit einem schlitz inst’s non-acceptance of cnst’s proposal reject-part(2) an orange flat-head and his repair by extension action-directive utt4 cnst ja. cnst’s acceptance of inst’s repair accept(3) yes. table 5: example dialogue from the airplane scenario. completion is realized using some kind of contrasting prosody” (poncin and rieser, 2006, p. 730). this analysis is one of few that aim to capture prosodic phenomena thoroughly (but see also purver and kempson, 2004). although the authors only discuss two specific examples, they claim (based on an investigation of the whole corpus) that the phenomenon is widespread and can be generalized. to our knowledge, no systematic analysis or annotation scheme capturing these effects is available. the damsl dialogue acts that we added in the right column of table 5 capture the outcome, not the process of updating common ground itself. usability for task-oriented scenarios. this interaction scenario raises an important issue that future annotation schemes will need to address, namely, which aspects in the joint task will lead to shared common ground that is sufficiently established to enable the less informed person to complete the more informed person’s sentences. as part of this process, the contribution of prosodic features is crucial for the update of common ground based on the immediate recognition and integration of subtle prosodic cues. 3.5 dollhouse: pro-active cooperation by an instructee an extensive corpus collected by coventry, andonova, and tenbrink (first cited in tenbrink et al., 2008) involved pairs of participants who were asked to furnish a dollhouse. only directors had full visual information about positions of objects; matchers were given the task of placing the objects into an empty dollhouse based on the directors’ instructions. since directors could not see the matchers’ dollhouses and actions, information needed to be communicated verbally. since the task did not involve shared action, it could be assumed that matchers listened to and followed the director’s commands without much negotiation. however, as shown by tenbrink et al. (2008), this was not the case; matchers were, in fact, rather active in discussing spatial locations despite their role as recipient of information. in particular, the matchers’ contributions of new spatial content could serve to clarify a global aspect of the current situation, to disambiguate an ambiguous description by explicitly mentioning options, or to specify an object’s position. the latter could be achieved by relating it to an(other) object already placed, or by suggesting another spatial term to describe the spatial relationship. thus, matchers actively participated in the joint effort of furnishing the dollhouse. table 6 shows an example from the dollhouse corpus; we provide annotations according to the damsl scheme. the example illustrates the engagement of the matcher in the identification and placement of an object (here: a washbasin) in the dollhouse. at the beginning, in d311-37 there is a clarification question about the director’s current spatial focus, since the director started a new dis196 annotation of joint action dialogues id speaker utterance dialogue act d311-36 director dann is’ das nächste ding du hast ähm assert then the next thing is you have um d311-37 matcher noch immer im selben raum? info-request still in the same room? d311-38 director genau. answer(d311-37) correct. d311-39 director links neben der toilette hast du diese blaue stellwand, diese zwischenwand. assert to the left next to the toilet you have this blue movable wall, this partition. d311-40 matcher ja. accept(d311-39) yes. d311-41 director und dahin, praktisch da links daneben davon is’ das waschbecken angesiedelt. action-directive and behind it, basically to the left beside it, the washbasin is located. d311-42 matcher das waschbecken mit dem spiegel? info-request the washbasin with the mirror? d311-43 director exakt. answer(d311-42) exactly. d311-44 director und mit diesem handtuchhalter daneben. action-directive and with this towel rail beside it. assert d311-45 matcher ja genau. accept(d311-44) yes, exactly. d311-46 matcher und das kommt ähm wohin? info-request and ah where to put it? d311-47 director in die lücke, in die lücke zwischen den beiden wänden an der linken hauswand, unter das dach praktisch. answer(d311-46), in the gap, in the gap between the two walls on the left outside wall, under the roof basically. action-directive d311-48 director okay? info-request okay? d311-49 matcher ach ach so in die lücke da. assert oh oh so into the gap there. offer d311-50 director genau, in die lücke. accept(d311-49) right, into the gap. action-directive d311-51 matcher wo die wand is’? info-request where the wall is? d311-52 director exakt, zwischen die beiden wände. answer(d311-51) exactly, between the two walls. action-directive table 6: example dialogue from the dollhouse corpus. 197 tenbrink, eberhard, shi, kübler, and scheutz course subgoal. rather than asking an open question, the matcher makes an informed guess, which is confirmed in d311-38. then the director establishes common ground via an assertion about the current status (d311-39). this serves as the basis for the actual object placement description in d311-41. d311-42 reveals uncertainty on the part of the matcher concerning the identity of the object to be placed; again the matcher makes an informed guess, which is confirmed and further enhanced by the director in d311-43 and d311-44. this subdialogue about the object’s identity is closed by the matcher’s acceptance in d311-45, who then returns in d311-46 to the previously started negotiation of the object’s placement. again, the matcher is actively involved by partially repeating (thereby confirming) instructions, or by making informed guesses as in d311-51. the director, in turn, elicits confirmation that the matcher is still on track, as in d311-48. this high level of engagement of the matcher may be traced back to two specific features of the scenario: that (1) the matcher actively needs to manipulate real world objects when furnishing the dollhouse, which requires sufficient understanding of the instruction to perform the associated action, rather than just identifying one of various possibilities without further consequences in the given task setting; and that (2) each single object placement represents a complex subgoal within the overall task of furnishing the dollhouse (objects need to be identified, placed, and oriented correctly in relation to the other objects within the same house). this overall complexity provides ample opportunity for the matcher to integrate his or her knowledge about previously reached subgoals with the actual goal of placing a particular object. usability for task-oriented scenarios. similar to the airplane scenario discussed in the previous subsection, the dollhouse scenario allows for the less informed person (the matcher) to contribute to the continuous update of common ground by making relevant suggestions. in particular, their perception of the scene allows for a number of spatial inferences to be integrated towards a reasonable interpretation of the director’s intentions. paralleling our observations for schober’s scenario in subsection 3.3, the dialogue-act-based annotation captures the underlying structural elements on a superficial level, but does not actually account for the perception-based negotiation efforts contributed by the interaction partners. to capture this latter aspect, their different perceptions on the spatial scene need to be considered. 3.6 navspace: actions as dialogic contributions the navspace corpus (tenbrink et al., 2010) serves as an example of the effects of the instructor’s ability to perceive task-relevant actions performed by the instructee. the interlocutors were guiding a wheelchair avatar through a schematized spatial environment shown on a computer screen. only the instructor knew about the goal position of the avatar, and only the instructee could move the avatar, using a joystick. the instructee was either a human or a dialogue system; instructors were always humans. communication was achieved via a text-based chat interface. participants in the human-human condition shared the same view on the scene including the avatar’s position, although they were not co-located and could, therefore, not use gestures or facial expressions to communicate intent. while the dialogue system was only equipped to respond (affirmatively) by saying “ok” and moving the avatar, or (in the case of problems) by asking a clarification question or rejecting the instruction, the human instructee’s responses were more complex and varied. in contrast to all other scenarios described so far, this setting allows for actions as responses to instructions, which are frequently found in the navspace corpus, particularly in the human-human interaction situation. unlike the dialogue system, humans did not regularly say “ok” before moving 198 annotation of joint action dialogues id speaker action utterance dialogue act 10-52-16 instructor weiter gerade aus action-directive continue straight ahead 10-52-17 instructee action accept(10-52-16) 10-52-26 instructor stopp action-directive stop 10-52-35 instructee action offer 10-52-58 instructor rückwärts zum vorigen raum links reject(10-52-35) backwards to the previous room on the left action-directive 10-53-2 instructee action accept(10-52-58) 12-48-26 instructor eine weiter action-directive one further 12-48-30 instructee bis zur ersten oder zweiten biegung? info-request until the first or the second turn? hold(12-48-26) 12-48-32 instructee action accept(12-48-26) 12-48-35 instructee ok accept(12-48-26) 12-48-35 instructee action accept(12-48-26) 12-48-36 instructor zweite action-directive second answer(12-48-30) 12-48-56 instructee action accept(12-48-36) table 7: example dialogues from the navspace corpus. the avatar, and human instructors interrupted actions if the move direction was incorrect. neither damsl nor the hcrc annotation scheme cover action in lieu of (verbal) dialogic contributions since they were created for different kinds of corpora and research questions. therefore, in order to annotate dialogue corpora with actions, some extensions are required. here we propose to introduce an action layer parallel to the task layer in damsl, which (like language-based dialogue acts) may serve forward-looking and backward-looking functions. the dialogue acts conveyed by actions are less rich than in language-based utterances. the most frequent action acts with forward-looking functions are offer and request, and with backward-looking functions accept and hold. table 7 gives examples. in the first example in table 7, instructions given in language (10-52-16 and 10-52-58) are accepted via actions (10-52-17 and 10-53-2, respectively). furthermore, the instructee uses an action to provide an offer in 10-52-35, which is rejected by the instructor in 10-52-58. in the second example in table 7, a request for additional information (12-48-30) is posed in parallel to a movement action that accepts the previous instruction (12-48-32). in 12-48-35, acceptance is signaled in parallel via language and action. usability for task-oriented scenarios. this scenario illustrates the importance of actions as proper contributions to the dialogue. as shown in the example dialogues in table 7, action and language are frequently interleaved temporally, and they provide meaningful information to advance the dialogue individually or jointly. generally, if actions are perceptually accessible to both dialogue partners (independent of whether they are physically located in the same room), actions contribute directly to the dialogue processes and structures, just as utterances do. actions can be used directly to update common ground, via reactions that are appropriate at a specific state in the dialogue. furthermore, both speech and actions constitute common ground and can be referred to in later utterances by both 199 tenbrink, eberhard, shi, kübler, and scheutz dialogue participants. annotation schemes clearly need to account for these effects, and we have provided here a first suggestion for how this can be accomplished in the damsl scheme. 3.7 crest: lexical and intonational expressions of uncertainty the indiana cooperative remote search task (crest) corpus (eberhard et al., 2010) is a corpus of approximately 8 minute dialogues recorded from 16 dyads performing a cooperative search task in which one person (director), who was located in a room away from the search environment, remotely directed the other person (searcher) through the environment by a hands-free telephone connection. the environment consisted of six connected rooms leading to a long winding hallway. neither the director nor the searcher was familiar with the environment prior to the task. the director guided the searcher through the environment with a map. the environment contained a cardboard box, 8 blue boxes, each containing three colored blocks, 8 pink boxes, and 8 green boxes. the locations of the cardboard box, blue boxes, and pink boxes were shown on the map; however, 3 of the 8 blue boxes’ locations were inaccurate, and the participants were informed of this. at the beginning of the experimental session, the director and searcher were told that the searcher was to retrieve the cardboard box and empty the blocks from the blue boxes into it. the searcher also was to report the locations of the green boxes to the director, who was to mark them on the map. they were told that instructions for the pink boxes would be given sometime during the task. five minutes into the task, the director and searcher were interrupted and the director was told that each of the 8 blue boxes contained a yellow block, and the searcher was to place a yellow block into each of the eight pink boxes. in addition, they had three minutes in which to complete all of the tasks. a timer that counted down the remaining time was placed in front of the director. the dyads’ performance was scored with respect to the number of boxes out of 24 for which the designated task was completed. the average score was 12 with a range of 1-21. the dyads’ verbal interactions were recorded along with video recordings of the searcher’s movement through the environment and the director’s marking the map with the green boxes. the verbal interactions were orthographically transcribed and annotated for conversational moves using the hcrc scheme. in the crest corpus, the director’s map, which was a two-dimensional floor plan, was a less reliable source of knowledge about the task domain (search environment) than the searcher’s direct perceptual experience of that domain. the disparity in the reliability of the knowledge was reflected in the directors producing twice as many requests for information (i.e., query-yn, query-w, check, align) than the searchers. in addition, about a third of the directors’ unsolicited descriptions of new elements in the environment, coded as explain moves, conveyed uncertainty via hedging expressions (e.g., “i think”, “it looks like”, “there should be”, etc.). these utterances were identified by coding them as explain/hedged. examples are given in the top half of table 8. notice that the searcher’s acknowledgments to the explain/hedged moves were affirmatives (i.e., “yes”, “right”), which confirm the accuracy of the common ground. there also were instances in which the directors’ explain moves appeared to convey uncertainty by ending with rising intonation, similar to questioning intonation. the location of the rising intonation was indicated with a question mark in the transcriptions. an example is shown in the bottom half of table 8. specifically, utt7 was coded as explain/query-yn because its declarative form and its unsolicited description of new elements in the environment are consistent with the explain code. however, the final rise in intonation was consistent with a request for confirmation of the accuracy of the description, or the query-yn code. ordinarily, the latter code would be assigned, but a consideration of the larger 200 annotation of joint action dialogues id speaker utterance dialogue act utt43 d3 okay ready utt44 d3 and straight in front of you should be: filing cabinets explain/hedged utt45 s3 yes acknow utt44 d3 okay ready utt46 d3 so: between the second cubicle on the right and the filing cabinets there should be: kind of like a space to walk through explain/hedged utt47 s3 right acknow utt48 d3 so go though there instruct utt49 s3 kay acknow utt5 d4 and through the first door instruct utt6 s4 okay acknow utt7 d4 and you’ll come to like a platform with some steps? explain/query-yn utt8 s4 yes acknow utt9 d4 and you’re gonna wanna turn to the right? instruct utt10 s4 yes acknow utt11 d4 and go straight ahead through that door instruct utt12 s4 yes acknow table 8: example dialogues from the crest corpus for explain/hedged and explain/queryyn moves. context made the interpretation of the final rise in intonation ambiguous. specifically, like utt7, utt9, which is an instruct move, ended with the same rising intonation, whereas utt11, which is also an instruct move, ended with falling intonation. this pattern is consistent with “list intonation”, which occurred when directors or searchers gave a complex instruction or description in installments. in the example in the table, utt7 utt12 constitute a segment of a larger dialogue in which the director directs the searcher through three connected rooms and down a hallway to retrieve a cardboard box. the first two utterances, utt5 and utt6, end the first segment in which the searcher was directed from the first room to the second room. thus, utt7 utt12 involved directing the searcher to the third room. like a yes-no question, the rising intonation at the end of each installment is followed by a pause for an acknowledgment from the addressee. like backchannels, the acknowledgments produced in this context perform a continuer function, i.e., i hear you, please continue (e.g., gardner, 2002). furthermore, the typical forms for this function are “mhm”, “uh huh”, and “yeah”, which also are associated with a confirmatory function. like the examples of explain/hedged moves in the top half of the table, the explain/query-yn example is followed by the affirmative acknowledgment “yes”, which may indicate the searcher’s sensitivity to the possible request. usability for task-oriented scenarios. the explain/hedged and explain/query-yn moves in the crest corpus illustrate the reliance on dialogue for updating common ground and coordinating joint actions in a scenario where the interactants communicated remotely. in the case of the explain/hedged move, the director’s hedged description of an aspect of the environment reflected his or her reliance on a map, which was a less reliable source of information compared to the searcher’s direct perceptual experience of the environment. the explain/query-yn move demonstrates how intonation can be ambiguous with respect to whether it reflects the communica201 tenbrink, eberhard, shi, kübler, and scheutz tive action being performed (i.e., a request for confirmation) or the dialogue structure (i.e., an installment in a sequence of moves for completing a segment of the task, c.f. with “uptalk”). this distinction might be captured in a layer of finer-grained annotation of the intonation, such as the tobi labeling scheme (beckman and hirschberg, 1994) (see below). however, regardless of whether the director’s rising intonation was intended to be a request for confirmation, the searcher’s affirmative acknowledgment provided this confirmation, allowing the common ground to be updated accordingly (e.g. safarova, 2006). 3.8 intonation and turn-taking in the columbia games corpus the columbia games corpus (gravano, 2009) is a corpus of 12 dyadic conversations recorded during two collaborative tasks, namely games played on separate computer screens. the games involved locating cards based on spoken communication, and lasted on average about 45 minutes. the participants were seated in a room, facing each other so that their own computer screen could not be seen by the other participant. additionally, there was an opaque curtain separating them so that they could not see each other. in such a scenario, speakers have to rely on spoken language and intonational cues to keep track of the conversation and manage common ground while constantly recurring to the perceptual information shown on their separate screens. ford and thompson (1996) showed that, while most intonationally complete utterances are also syntactically complete, about one half of syntactically complete utterances are intonationally incomplete, signalling the continuation of an ongoing speaker turn. this demonstrates the prominent role of intonation for dialogue structure. gravano (2009) aimed at identifying the precise ways in which these turn taking processes operate on the basis of subtle intonational cues conveyed during speech. the annotation in the columbia games corpus includes self-repairs, non-word vocalization (laughs, coughs, breaths, and the like), and a detailed intonation analysis using the tobi labeling scheme (beckman and hirschberg, 1994). tobi consists of four tiers: an orthographic tier, which includes the orthographic transcription of the recording, a break index (bi) tier, a tonal tier, and a miscellaneous tier, which can be used to mark disfluencies, etc. the break index tier annotates the end of each word for the strength of its association with the next word. the tonal tier describes a phonological analysis of the utterance’s intonation pattern. this annotation use two distinct tones, h(igh) and l(ow), with additional diacritics to mark pitch accents or downstep. the corpus was also annotated for turn-taking based on categories suggested by beattie (1982). the annotation scheme distinguishes between overlap, interruption, butting-in, smooth switch, pause interruption, backchannel with overlap, and backchannel. according to this analysis, speakers were more likely to switch turns following subtle speaker cues such as a lower intensity level, a lower pitch level, and a point of textual completion. table 9 shows an example4 with tobi annotation, where speaker b uses contrastive stress to draw attention to two distances between symbols, the distance between the ruler and the blue crescent, and between the blue and the yellow crescent. usability for joint action scenarios: as illustrated in previous subsections, intonation patterns can prove crucial for the updating of common ground. the tobi annotation scheme provides a systematic solution for capturing intonation and turn-taking aspects. although (to our knowledge) it has not been used to identify common ground updating processes, it stands to reason that speakers 4. we thank a. gravano for providing the example. 202 annotation of joint action dialogues speaker b: bi: tone: but there’s 1 h* speaker a: bi: # 3p speaker b: bi: tone: more 1 h* space 1 !hbetween 3 the 1 ruler 1 l+h* hand 3 the 1 blue 1 crescent 1 h* h* # 4 than h* there 1 is 1 hthe 3 blue 1 crescent 1 h* hand 3 the 1 yellow 1 l* crescent 1 l-l% # 4 speaker a: bi: tone: huh h* l-l% # 4 table 9: example from the columbia games corpus. use intonational clues to conclude when common ground has been reached sufficiently for current purposes, or when they need to step in so as to ask a clarification question, leading to the turn-taking patterns identified by gravano (2009). 3.9 basic gesture annotation mcneill (2000) developed an annotation scheme for gestures which was used for the rapport corpus5. this corpus contains face-to-face dialogues with dyads of people who know each other well. in an example given in mcneill (2002), one person watched a tweety and sylvester cartoon and had to tell the story to another person. both participants had been told that the listener would have to tell the story to yet another person. gestures were not mentioned in the instructions. table 10 shows the dialogue and gesture annotation. square brackets indicate the beginning and end of the gesture, and boldface marks the gesture stroke – “the phase with semantic content and the quality of effort” (mcneill, 2002, footnote 11). the annotations specify whether one hand or both are used, and whether the gesture was symmetrical or asymmetrical in the latter case. it also describes the movement (e.g., “move down”) and the number of repetitions (e.g., 2x). in the example, the speaker mostly uses gestures to illustrate motion events. in utterance (1), for example, “going up the inside” is accompanied by an upward gesture of the right hand. another approach to gesture annotation called form was used for video recordings of monologues by martell (2002). the annotation, in the form of annotation graphs, captures different body parts in a complex procedural model (see figure 1). form uses fine-grained labels for gestures such as “upperarm.location” and “handandwrist.movement”. thus, it captures a wider range of movements in a more standardized way than the scheme proposed by mcneill (2000). usability for task-oriented scenarios. gestures often accompany utterances to provide supporting or additional information, which is used to establish common ground. the annotation schemes for gestures exemplified here capture the nature of gestures, which is an essential first step. however, no information can be derived about the updating function in the dialogue. establishing the role of gestures in the context of dialogue is central for situated interaction scenarios that incorporate relevant gestures. 5. http://mcneilllab.uchicago.edu/corpora/rapport_corpus.html 203 tenbrink, eberhard, shi, kübler, and scheutz id utterance gesture (1) he tries going [up the inside of the drainpipe and] 1hand: rh rises up 3x (2) tweety bird runs and gets a bowling ba[ll and drops it down the drainpipe #] symmetrical: 2 similar hands move down (3) [and / as he s coming up] asymmetrical: 2 different hands, lh holds, rh up 2x (4) [and the bowling ball s coming d]] asymmetrical: 2 different hands, rh holds, lh down (5) [own he ssswallows it] asymmetrical: 2 different hands, rh up, lh down (6) [ # and he comes out the bottom of the drai] 1hand: lh comes down (7) [npipe and he s got this big bowling ball inside h]im symmetrical: 2 similar hands move down (8) [and he rolls on down] [into a bowling all] symmetrical: 2 similar hands move forward 2x (9) [ey and then you hear a sstri]ke # symmetrical: 2 similar hands move apart table 10: example dialogue from mcneill (2002). figure 1: an example for gesture annotation in form (martell, 2002). 204 annotation of joint action dialogues 3.10 route gestures striegnitz et al. (2009) examined the use of gestures to support route directions, focusing on gestures associated with landmarks. in a close examination of five route dialogues, they identified gestures that indicated route perspective (i.e., the perspective of the person following the route), gestures consistent with a map-based survey-perspective, gestures which locate the object with respect to the speaker’s actual position and orientation, and gestures that indicate the landmark’s shape. all of these clearly contribute valuable spatial information, enriching the spoken language. for instance, the utterance “and it’s really big” is accompanied by a gesture that indicates the landmark’s horizontal extent. this information is not included in the dimension-neutral size term “big”. striegnitz et al. (2009) coded the route dialogues using a damsl type annotation scheme, identified the gesture information separately, and related gesture types to utterance categories. most of the gestures accompanied statements: typically those that mentioned a landmark, but also some that did not. these were further specified with respect to their role in the dialogue, such as plain statements or those that serve the function of a response, a query, an elaboration, or a redescription. usability for task-oriented scenarios. although no further specification of the annotation scheme was proposed by striegnitz et al. (2009) (nor were any dialogue examples given), the proposed scheme still demonstrates the necessity to systematically account for speech-accompanying gestures, as they play an important role in the updating of common ground. 3.11 rolland: multimodal interaction with a mobile robot the rolland multimodal interaction study (first cited in anastasiou, 2012)6 addressed the interaction between a powered wheelchair called rolland and a user who was asked to carry out a set of simple tasks with rolland. participants (i.e., users) were told that rolland could understand spoken language as well as gestures by hands and arms, but would only react through driving actions. the study was a “wizard-of-oz” setup, i.e., the wheelchair was remote-controlled by a human operator, unbeknownst to the participants. table 11 includes an example dialogue in which the user asks rolland to come to the bedside, so that he could get into the wheelchair without much effort. we provide a damsl-type annotation here. the user starts by using spoken language only (until utt140), but remains dissatisfied with rolland’s reactions. consequently, he adds gestures (utt141, utt143 and utt145) that enhance the spoken utterances, so as to instruct rolland more pointedly. although still unhappy with the result, he finally accepts it in utt147. interestingly, at one point in the interaction, the gesture is inconsistent with the spoken utterance. in utt143, the user asks rolland to drive a little bit backward. along with this, however, the hand points to the goal location beside the bed – seemingly contradicting the spoken command. arguably, the utterance refers to a more specific level of granularity (the immediate action to be performed) than the gesture, which points to the higher-level goal location. as a result, rolland’s action in utt144 is ambiguous with respect to the performed dialogue act. on the one hand, the pointing gesture in utt143 is accepted by the action of driving towards the bed. the spoken utterance in utt143 (driving backwards), on the other hand, is rejected by the same action. 6. this study was carried out by dimitra anastasiou and daniel c. vale in the collaborative research center sfb/tr 8 spatial cognition in bremen, funded by the dfg. thanks to the authors for allowing us to gain insight into their work, and cite aspects central to our current focus. 205 tenbrink, eberhard, shi, kübler, and scheutz id part. utterance action gesture dialogue act utt132 user rolland komme bitte zum bett hierhin action-directive rolland please come here to the bed utt133 user hier wo ich sitze action-directive here where i’m sitting utt134 rolland driving to bed accept(132,133) utt135 user er ist mir auf den fuß gefahren, okay assert accept(134) he drove over my foot, okay utt136 user komme etwas näher zum bett action-directive come a bit closer to the bed utt137 rolland driving to bed accept(136) utt138 user noch etwas näher zum bett action-directive once more closer to the bed utt139 rolland driving to bed accept(138) utt140 user weiter zu mir action-directive further to me utt141 user etwas näher zu mir two-handed configuration action-directive a bit closer to me open hand shapes utt142 rolland driving to bed accept(140,141) utt143 user ein bisschen zurück fahren pointing with reject-part(142) drive a little bit backward index finger action-directive utt144 rolland driving to bed ambiguous utt145 user rolland ein bisschen zurück fahren pointing with reject(144) rolland drive a little bit backward index finger action-directive utt146 rolland driving backward accept(145) utt147 user okay ich versuche mal so accept(146) okay i try like this table 11: example dialogue from the rolland multimodal interaction corpus. 206 annotation of joint action dialogues usability for joint action scenarios: although this scenario is clearly restricted by the limited interaction capabilities of the wheelchair, the example illustrates the tightly integrated yet independent role of gesture, just as in human face-to-face interaction. gestures can elaborate and expand language to establish common ground, address different aspects, or appear as incongruent with the verbally conveyed content, leading to further dialogue structure complexities. interlocutors may ignore gesturally conveyed content, or react verbally or non-verbally. annotation schemes need to account for these procedures. our damsl-based annotation example provides a first suggestion of how this might be accomplished. 4. layers of annotation the above review of insights gained in joint task settings highlighted a range of factors that are crucial for the negotiation of shared goals in situated dialogue, showing how speakers manage to update their knowledge of the current state of the task in their common ground. in the following, we propose four potential additional layers for annotation schemes (depending on the research question at hand). generally, for each layer, an aspect can be a direct contribution to dialogue if it is accessible to both interactants to the same degree, while it affects the dialogue more indirectly if such access is not shared. for example, an action can only serve as a direct response to a directive if this is also perceived by the director; if the action is not perceived as such, the actor will typically acknowledge the directive verbally in addition to acting upon it, so as to update the director’s state of knowledge. along these lines, sharedness of these layers turns out to be a major factor in any analysis of dialogue. although we were able to use existing schemes like damsl to express some of these effects, there is a clear need for further extensions of annotation formalisms to be able to represent the intricate interplay between linguistic and non-linguistic interactions in joint action scenarios. 4.1 intonation as demonstrated by several examples in the previous subsections, intonation plays an important role in common ground updating processes. speakers use intonational cues to determine when meaningful fragments in an utterance have been completed or when clarification questions may be asked, thus aiding turn-taking and contributing to dialogue structure. intonation can also highlight meanings, convey the significance of an utterance, and provide feedback about the acknowledgment or rejection of a previous utterance. critically, intonational cues are used and picked up automatically by interactants and directly affect the pragmatic implicatures and the dialogue flow. intonational cues thus play a prominent role in spoken dialogues. when speakers cannot see each other they can compensate for missing cues conveyed by gestures or facial expressions. we are not aware of any current dialogue scheme that combines dialogue structure annotation with intonation patterns such as those identified by the tobi annotation scheme. conceivably, the annotation of intonation can be pragmatically reduced to the most relevant aspects that directly affect meaning interpretations and dialogue structure. this would involve adding a further layer to an existing dialogue structure annotation scheme, with the possibility to override meaning interpretations and dialogue moves in other layers (e.g., what might otherwise be coded as “acknowledgement” could end up as a “yn-question” based on intonational information). 207 tenbrink, eberhard, shi, kübler, and scheutz 4.2 gestures gestures can substantially supplement verbal information. this is most clearly demonstrated when they contribute spatial information, for instance in referential phrases, thereby substantially enhancing the common ground updating process by contributing aspects that may not be verbalised at all. similar processes are active whenever speakers have visual access to their interaction partner such as in face-to-face communication. dialogue structure annotation schemes can be straightforwardly enriched by an additional layer capturing gestures, as exemplified in table 11 above. the level of granularity of gesture annotations will depend on the research issue at hand. 4.3 perception of the task domain situated tasks involve perceptual access to the task domain. even if perception is not shared, the task domain information that each of the speakers has access to is central to the coordination of actions and accumulation of common ground. in the dollhouse scenario in subsection 3.5, for instance, the matcher is able to make informed guesses about positions of objects because verbal instructions can be compared to the arrangement of objects in the perceivable scene. similarly, many verbal contributions in the crest scenario (subsection 3.7) directly build on the non-joint perceptions of both interlocutors and thereby affect dialogue structure. while the results of these effects are captured by dialogue structure annotations, the procedures as such can only be fully understood if the relevant perceptions are also taken into account. we suggest adding a layer of scene perception to the annotation that can be used to capture relevant perceptual aspects that speakers draw from. in this way, the functions of particular dialogue acts can be interpreted more reliably. moreover, in the case of non-shared scene perception, cases of miscommunication can be better identified and accounted for based on the discrepancy between the interactants’ access to the task domain. 4.4 actions the ability to perceive task-relevant actions provides reliable and timely information for updating common ground about the status of the task. the navspace example in subsection 3.6 illustrates how using the same categories for coding task-relevant actions and dialogue structure captures the complementary role of actions and dialogue moves in the process of updating common ground. the interaction between actions and dialogue becomes more complicated when speakers are engaged in a joint task without directly sharing perceptual access to action outcomes, as exemplified by the crest example (subsection 3.7). in such scenarios, speakers will communicate action outcomes in some cases, but assume that they can be inferred in other cases. as a result, common ground representations among interactants may start to diverge and become inconsistent, which will eventually result in dialogue interactions solely dedicated to resolving the inconsistencies and re-establishing common ground. to account for these effects and reliably interpret the function of dialogue acts (e.g., to reestablish coordination and common ground), we recommend keeping track of the speakers’ actions by adding a corresponding layer to the dialogue annotation. the specification of this layer (e.g., richness of action descriptions, temporal extension, etc.) will depend on the purpose of the task and dialogue analysis. 208 annotation of joint action dialogues 5. conclusion in this paper, we reviewed existing dialogue corpora across various interaction settings to investigate the different linguistic and non-linguistic aspects that affect how interactants negotiate joint actions and update their common ground in task-based dialogues. we specifically examined existing analyses and established annotation schemes in the literature to determine the extent to which they are able to identify and capture relevant aspects of action negotiation and updating of common ground. this is particularly important as interactants in joint activities will make use of any information, including perceptions and action available to them (e.g., perceptions about the task domain, gestures by other interactants, or actions on task-relevant objects). current annotation schemes typically fail to account for these features of situated task scenarios. in other words, while clark’s influential work (clark, 1996; clark and wilkes-gibbs, 1986) has led to a widely acknowledged view of dialogue as joint activity that involves more than just the dialogue interactions, this recognition is not yet systematically or coherently reflected in dialogue annotation schemes. for information-seeking dialogues, relevant insights have been gained about clarification phenomena (e.g., purver et al., 2003; rieser and moore, 2005). however, dialogue structure analysis for situated task scenarios is more complex, as dialogue patterns differ substantially from those identified in purely language-based interaction settings. a better understanding of the intricate processes involved in human-human dialogues as part of situated joint activities is not only central to a better understanding of human natural language interactions, but also critical for research in human-computer or human-robot interaction (e.g., alexandersson et al., 1998; green et al., 2006; allen et al., 2001; shi et al., 2010). to pursue this line of research, rich annotations of dialogue corpora are required that, in addition to linguistic annotations, include interlocutors’ perceptions, intonation, gestures (where appropriate), actions, and any other relevant factors that contribute to building up and negotiating common ground. only with these additional annotations will it be possible to determine and build computational models of the intricate interplay between linguistic and non-linguistic aspects in task-based dialogues. 6. acknowledgments this work was in part funded by the dfg, sfb/tr 8 spatial cognition, project i5-[diaspace], to the first and third author, and by onr muri grant #n00014-07-1-1049 to the second and last author. we also wish to thank elena andonova, john a. bateman, kenny r. coventry, nina dethlefs, juliana goschler, cui jian, robert j. ross, and kavita e. thomas for collaboration on dialogue corpora and related issues, and christoph broschinski for relevant inspiration. we are also grateful to the anonymous reviewers, who made constructive suggestions that helped us to improve and focus this paper. references jan alexandersson, bianka buschbeck-wolf, tsutomu fujinami, michael kipp, stephan koch, elisabeth maier, norbert reithinger, birte schmitz, and melanie siegel. dialogue acts in verbmobil-2 (second edition). verbmobil report 226, university of the saarland, saarbrücken, germany, 1998. 209 tenbrink, eberhard, shi, kübler, and scheutz james allen and mark core. draft of damsl: dialogue markup in several layers. technical report, university of rochester, 1997. url http://www.cs.rochester.edu/research/ speech/damsl/revisedmanual/. james allen and c. raymond perrault. analyzing intention in utterances. artificial intelligence, 15 (3):143–178, 1980. james allen, donna byron, myroslava dzikovska, george ferguson, lucian galescu, and amanda stent. toward conversational human-computer interaction. ai magazine, 22(4):27–37, 2001. james f. allen, george ferguson, brad miller, and eric ringger. trains as an embodied natural language dialog system. in embodied language and action: papers from the 1995 fall symposium. aaai technical report fs-95-05, 1995. jens allwood, stefan kopp, karl grammer, elisabeth ahlsén und elisabeth oberzaucher, and markus koppensteiner. the analysis of embodied communicative feedback in multimodal corpora: a prerequisite for behavior simulation. journal on language resources and evaluation, 41(2-3):325–339, 2007. special issue on multimodal corpora for modeling human multimodal behaviour. dimitra anastasiou. a speech and gesture spatial corpus in assisted living. in proceedings of the 8th international conference on language resources and evaluation (lrec), istanbul, turkey, 2012. anne anderson, miles bader, ellen gurman bard, elizabeth boyle, gwyneth doherty, simon garrod, stephen isard, jaqueline kowtko, jan mcallister, jim miller, cathy sotillo, henry thompson, and regina weinert. the hcrc map task corpus. language and speech, 34(4):351–366, 1991. adrian bangerter. using pointing and describing to achieve joint focus of attention in dialogue. psychological science, 15(6):415–419, 2004. ellen gurman bard, catherine sotillo, anne h. anderson, henry s. thompson, and martin m. taylor. the dciem map task corpus: spontaneous dialogue under sleep deprivation and drug treatment. speech communication, 20(1-2):71–84, 1996. ellen gurman bard, anne h. anderson, yiya chen, hannele b.m. nicholson, catriona havard, and sara dalzel-job. let’s you do that: sharing the cognitive burdens of dialogue. journal of memory and language, 57(4):616–641, 2007. geoffrey beattie. turn-taking and interruption in political interviews: margaret thatcher and jim callaghan compared and contrasted. semiotica, 39(1/2):93–114, 1982. mary beckman and julia hirschberg. the tobi annotation conventions. technical report, the ohio state university, 1994. mary beckman, julia hirschberg, and stefanie shattuck-hufnagel. the original tobi system and the evolution of the tobi framework. in sun-ah jun, editor, prosodic models and transcription: towards prosodic typology. oxford university press, 2005. 210 annotation of joint action dialogues peter bohlin, robin cooper, elisabet engdahl, and staffan larsson. information states and dialog move engines. electronic transactions in ai, 3(9), 1999. elizabeth a. boyle, anne h. anderson, and alison newlands. the effects of visibility on dialogue and performance in a cooperative problem solving task. language and speech, 37(1):1–20, 1994. holly p. branigan, robin j. lickley, and david mckelvie. non-linguistic influences on rates of disfluency in spontaneous speech. in proceedings of the 14th international congress of phonetic sciences, san francisco, ca, 1999. susan e. brennan, xin chen, christopher a. dickinson, mark b. neider, and gregory j. zelinsky. coordinating cognition: the costs and benefits of shared gaze during collaborative search. cognition, 106(3):1465–1477, 2008. susan e. brennan, gregory j. zelinsky, joy e. hanna, and kelly j. savietta. eye gaze cues for coordination in collaborative tasks. in duet 2012: dual eye tracking in cscw, seattle, wa, 2012. harry bunt. the dit++ taxonomy for functional dialogue markup. in proceedings of the amaas 2009 workshop towards a standard markup language for embodied dialogue acts, budapest, hungary, 2009. jean carletta, stephen isard, amy isard, gwyneth doherty-sneddon, jacqueline kowtko, and anne anderson. the reliability of a dialogue structure coding scheme. computational linguistics, 23 (1):13–31, 1997. herbert h. clark. using language. cambridge university press, 1996. herbert h. clark and deanna wilkes-gibbs. referring as a collaborative process. cognition, 22: 1–39, 1986. philip r. cohen and hector j. levesque. speech acts and the recognition of shared plans. in proceedings of the third biennial conference, canadian society for computational studies of intelligence, victoria, bc, canada, 1980. mark g. core and james f. allen. coding dialogs with the damsl annotation scheme. in aaai fall symposium on communicative action in humans and machines. aaai press, cambridge, ma, 1997. kathleen eberhard, hannele nicholson, sandra kübler, susan gunderson, and matthias scheutz. the indiana “cooperative remote search task” (crest) corpus. in proceedings of the seventh international conference on language resources and evaluation (lrec), valetta, malta, 2010. kerstin fischer and thora tenbrink. video conferencing in a transregional research cooperation: turn-taking in a new medium. in jana döring, h. walter schmitz, and olaf schulte, editors, connecting perspectives. videokonferenz: beiträge zu ihrer erforschung und anwendung, aachen, 2003. shaker. 211 tenbrink, eberhard, shi, kübler, and scheutz cecilia ford and sandra a. thompson. interactional units in conversation: syntactic, intonational, and pragmatic resources for the management of turns. in elinor ochs, emanuel schegloff, and sandra a. thompson, editors, interaction and grammar, pages 134–184. cambridge university press, 1996. rod gardner. when listeners talk: response tokens and listener stance. john benjamins publishing co., philadelphia, 2002. agustı́n gravano. turn-taking and affirmative cue words in task-oriented dialogue. phd thesis, columbia university, 2009. anders green, helge hüttenrauch, elin anna topp, and kerstin severinson eklundh. developing a contextualized multimodal corpus for human-robot interaction. in proceedings of the international conference on language resources and evaluation (lrec), genoa, italy, 2006. barbara j. grosz and candace l. sidner. attention, intentions and the structure of discourse. computational linguistics, 12(3):175–204, 1986. joy e. hanna and susan e. brennan. speakers’ eye gaze disambiguates referring expressions early during face-to-face conversation. journal of memory and language, 57(4):596–615, 2007. peter a. heeman and james allen. the trains 93 dialogues. technical report, computer science department, the university of rochester, 1995. url http://www.cs.rochester.edu/ research/speech/trains.html. peter a. heeman and james allen. tagging speech repairs. in arpa workshop on human language technology, pages 187–192, plainsboro, nj, 1994. dan jurafsky, liz shriberg, and debra biasca. switchboard swbd-damsl shallow-discoursefunction annotation coders manual, draft 13. technical report rt 97-02, institute for cognitive science, university of colorado at boulder, 1997. adam kendon. some functions of gaze-direction in social interaction. acta psychologica, 26: 22–63, 1967. michael kipp, michael neff, and irene albrecht. an annotation scheme for conversational gestures: how to economically capture timing and form. journal on language resources and evaluation, 41(3-4):325–339, 2007. jacqueline kowtko, amy isard, and gwyneth doherty. conversational games within dialogue. technical report hcrc/rp-31, human communication research centre, university of edinburgh, 1993. staffan larsson and david traum. information state and dialogue management in the trindi dialogue move engine toolkit. natural language engineering, pages 323–340, 2000. special issue on best practice in spoken language dialogue systems engineering. oliver lemon, alexander gruenstein, and stanley peters. collaborative activities and multitasking in dialogue systems. traitement automatique des langues (tal), 43(2):131–154, 2002. special issue on dialogue. 212 annotation of joint action dialogues craig martell. form: an extensible, kinematically-based gesture annotation scheme. in proceedings of the international conference on spoken language processing, denver, co, 2002. craig martell and joshua kroll. corpus-based gesture analysis: an extension of the form dataset for the automatic detection of phases in gesture. international journal of semantic computing, 1 (4):521–536, 2007. david mcneill. language and gesture. cambridge university press, cambridge, 2000. david mcneill. gesture and language dialectic. acta linguistica hafniensia, 34(1):7–37, 2002. mark b. neider, xin chen, christopher a. dickinson, susan e. brennan, and gregory j. zelinsky. coordinating spatial referencing using shared gaze. psychonomic bulletin and review, 17(5): 718–724, 2010. claire o’malley, steve langton, anne anderson, gwyneth doherty-sneddon, and vicki bruce. comparison of face-to-face and video-mediated interaction. interacting with computers, 8(2): 177–192, 1996. kristina poncin and hannes rieser. multi-speaker utterances and co-ordination in task-oriented dialogue. journal of pragmatics, 38:718–744, 2006. matthew purver and ruth kempson. incrementality, alignment and shared utterances. in proceedings of the 8th workshop on the semantics and pragmatics of dialogue (catalog), pages 85–92, barcelona, spain, 2004. matthew purver, jonathan ginzburg, and patrick healey. on the means for clarification in dialogue. in ronnie smith and jan van kuppevelt, editors, current and new directions in discourse and dialogue, pages 235–255. kluwer academic publishers, dordrecht, 2003. gert rickheit and ipke wachsmuth. situated communication. mouton de gruyter, berlin, 2006. verena rieser and johanna d. moore. implications for generating clarification requests in taskoriented dialogues. in proceedings of the 43rd annual meeting of the acl, ann arbor, mi, 2005. marie safarova. rises and falls: studies in the semantics and pragmatics of intonation. phd thesis, university of amsterdam, 2006. michael f. schober. spatial perspective taking in conversation. cognition, 47(1):1–24, 1993. michael f. schober. how addressees affect spatial perspective choice in dialogue. in patrick l. olivier and klaus-peter gapp, editors, representation and processing of spatial expressions, pages 231–245. lawrence erlbaum associates, mahwah, new jersey, 1998. michael f. schober. spatial dialogue between partners with mismatched abilities. in kenny coventry, thora tenbrink, and john bateman, editors, spatial language and dialogue, pages 23–39. oxford university press, 2009. michael f. schober. speakers, addressees, and frames of reference: whose effort is minimized in conversations about location? discourse processes, 20(2):219–247, 1995. 213 tenbrink, eberhard, shi, kübler, and scheutz abigail j. sellen. remote conversations: the effects of mediating talk with technology. humancomputer interaction, 10:401–444, 1995. hui shi, robert j. ross, thora tenbrink, and john bateman. modelling illocutionary structure: combining empirical studies with formal model analysis. in a. gelbukh, editor, proceedings of the 11th international conference on intelligent text processing and computational linguistics (cicling 2010), lecture notes in computer science, berlin, 2010. springer. march 21-27. iasi, romania. teresa sikorski and james f. allen. a task-based evaluation of the trains-95 dialogue system. in ecai workshop on dialogue processing in spoken language systems, pages 207–219. springer, budapest, hungary, 1997. lesley stirling, janet fletcher, ilana mushin, and roger wales. representational issues in annotation: using the australian map task corpus to relate prosody and discourse structure. speech communication, 33:113 – 134, 2001. kristina striegnitz, paul tepper, andrew lovett, and justine cassell. knowledge representation for generating locating gestures in route directions. in kenny coventry, thora tenbrink, and john bateman, editors, spatial language and dialogue, pages 147–165. oxford university press, 2009. thora tenbrink, elena andonova, and kenny coventry. negotiating spatial relationships in dialogue: the role of the addressee. in proceedings of londial – the 12th semdial workshop, london, uk, 2008. thora tenbrink, robert j. ross, kavita e. thomas, nina dethlefs, and elena andonova. route instructions in map-based human-human and human-computer dialogue: a comparative analysis. journal of visual languages and computing, 21(5):292–309, 2010. david r. traum. conversational agency: the trains-93 dialogue manager. in susann luperfoy, anton nijholt, and gert veldhuijzen van zanten, editors, dialogue management in natural language systems, pages 1–11. universiteit twente, enschede, 1996. 214 completions_camera_ready_repaired dialogue and discourse 1 (2010) 1-89 received 2/09; first letter 7/09; accepted 7/09; doi: 10.5087/dad.2010.001 final revision submitted 12/09; published online 2/10 ©2010 massimopoesio and hannes rieser completions, coordination, and alignment in dialogue massimo poesio massimo.poesio@unitn.it centre for mind/brain sciences and disi università di trento c. so bettini 31 rovereto (tn), italy hannes rieser hannes.rieser@uni-bielefeld.de fakultät für linguistik und literaturwissenschaft universität bielefeld universitätsstraße 25 33615 bielefeld, germany editor: jonathan ginzburg abstract collaborative completions are among the strongest evidence that dialogue requires coordination even at the sub-sentential level; the study of sentence completions may thus shed light on a number of central issues both at the `macro’ level of dialogue management and at the `micro’ level of the semantic interpretation of utterances. we propose a treatment of collaborative completions in ptt, a theory of interpretation in dialogue that provides some of the necessary ingredients for a formal account of completions at the ‘micro’ level, such a theory of incremental utterance interpretation and an account of grounding. we argue that an account of semantic interpretation in completions can be provided through relatively straightforward generalizations of existing theories of syntax such as lexical tree adjoining grammar (ltag) and of semantics such as (compositional) drt and situation semantics. at the macro level, we provide an intentional account of completions, as well as a preliminary account within pickering and garrod’s alignment theory. keywords: completions, coordination, incremental interpretation, ptt, grounding 1 introduction utterances such as 1.2, 1.3, and 2.2 in the following fragment of a transcript from the bielefeld toy plan corpus of task-oriented dialogues (skuplik, 1999) are examples of the constructions that clark (1996) called collaborative completions. (1.1) 1.1 inst so, jetzt nimmst du [pause] well, now you take 1.2 cnst eine schraube a screw 1.3 inst eine <-> orangene mit einem schlitz. an <-> orange one with a slit 1.4 cnst ja yes poesio and rieser 2 2.1 inst und steckst sie dadurch, also and you put it through there, let’s see 2.2 cnst von oben from the top 2.3 inst von oben, daß also die drei festgeschraubt werden dann from the top, so that the three bars get fixed then 2.4 cnst ja yes that agents cooperate is a fundamental assumption in many theories of dialogue (see, e.g., the papers in cohen et al, 1990, or (clark, 1996)). that cooperation requires coordination is a central theme particularly in the work of clark (e.g., 1996) and garrod (e.g., (garrod & anderson, 1987)). collaborative completions such as those in the example dialogue are among the strongest evidence for the argument that dialogue requires coordination even at the subsentential level (clark, 1996; garrod and anderson, 1987; pickering and garrod, 2004). more generally, studying collaborative completions may shed light on a number of central issues for models of the semantics and pragmatics of dialogues, both at the `macro’ level of dialogue management and at the `micro’ level of the semantic interpretation of utterances. at the macro level, this type of data may be used to compare competing claims about coordination—i.e., whether it is best explained with an intentional model like clark’s and the models discussed in cohen et al’s book, or with a model based on simpler alignment mechanisms like pickering and garrod’s. at the micro level, completions are clear evidence that intention recognition in dialogue proceeds incrementally, and may provide insights about semantic composition. in this paper we propose a treatment of collaborative completions in ptt (poesio and traum, 1997, 1998; poesio and muskens, 1997; matheson, poesio and traum, 2000) a theory of interpretation in dialogue incorporating ideas from (compositional) drt (kamp and reyle, 1993; muskens, 1996), situation semantics (barwise and perry, 1983), and lexical tree adjoining grammar (ltag) (abeillé and rambow, 2000; joshi 2004). our main ambition is to demonstrate that it is possible to provide an analysis of completions covering a great many of the syntactic, semantic, and pragmatic properties of these constructions, but relying for the most part on already established and independently motivated formal devices incorporated in ptt, supplemented by either a theory of intentions such as tuomela’s (2000) or an alignment-based account such as pickering and garrod’s (2004). we also argue that ptt provides certain ingredients that are essential for a formal account of completions yet are missing from existing accounts, such as (purver et al, 2006) –e.g, an explicit account of grounding (clark and schaefer, 1989; traum, 1994), required to understand the interaction of completions with grounding. finally, we explore in some detail the debate between ‘intention-based’ and `alignment-based’ models of dialogue, providing both a more traditional ‘intentional’ account and an alignment-based treatment of the phenomenon. the structure of the paper is as follows. in section 2 we briefly summarize clark’s theory of coordination in dialogue and our main empirical evidence on completions, data from the bielefeld toy plane corpus. in section 3 we introduce ptt; readers already familiar with the theory may skip this section, or perhaps consult only those parts they need to understand section 5. (additional formal details are provided in the appendix.) . in section 4 we present an intentional account of completions. the first part of this section is background, introducing recent ideas about intentions and shared plans that are required to explain completions; whereas the second part provides an explanation of the example based on these ideas. section 5 is the main section of the paper; in it, we use the notions introduced in the previous two sections to provide an completions, coordination and alignment in dialogue 3 intentional account of completions in ptt taking into account the incremental nature of the phenomenon and its interaction with grounding. . in section 6, we provide a preliminary treatment of continuations in terms of the pickering and garrod alignment model. finally, in section 7 we discuss alternative theories of dialogue including kos (ginzburg, 2009), sdrt (asher and lascarides, 2003) and dynamic syntax (kempson et al, 2001; cann et al, 2005) and compare our analysis with that proposed by purver et al. 2 coordination, grounding, and completions 2.1 coordination and grounding in dialogue the theory of interpretation in dialogue developed in this paper relies heavily on the views on dialogue developed by herbert clark and colleagues, and summarized, e.g., in (clark, 1992; 1996). clark points out that the traditional view of speech acts as developed primarily by searle (e.g., 1969) and exposed in (levinson 1983) ignores the role of listeners. he views conversation as a form of joint activity like playing football or playing in a string quartet (clark 1996, ch. 3), in which participants engage in joint projects at all communicative levels, from uttering sounds to performing illocutionary acts. he argues that speech acts are really joint projects between the speaker and the listener (clark 1996, ch. 5). in order for these joint communicative activities to be successful, agents need to coordinate both on matters of timing (e.g., to avoid speech overlap) and on matters of content, developing together what clark calls joint construal (joint interpretation) of these actions, that need not be the interpretation the speaker originally intended. this coordination is only made possible by the common ground (stalnaker, 1978) between the participants; yet clark points out that the establishment of a common ground is not automatic (“the common ground isn’t just there, ready to be exploited” – clark 1996, p. 116). common ground has to be established with each agent with whom we interact; this is particularly the case for that part of the common ground that has to do with the present conversation. the participants to a joint action need to establish the mutual belief that they have succeeded well enough for current purposes (principle of joint closure, clark 1996, p. 226). this establishment does not occur by default, but through a process called grounding in which positive evidence of understanding at different levels is required (clark and wilkes-gibbs, 1986; clark and schaefer, 1989; clark and brennan, 1991). such evidence may consist of signals of understanding like nods or “uh huh”, of presuppositions of understanding like taking up the relevant part of the joint project—e.g, answering the question just asked—etc. this process is structured around contributions to discourse—signals successfully understood. contributions are organized into a presentation phase, in which conversant a presents the signal to b, and an acceptance phase, in which b gives evidence that she believes she understands what a meant (clark & schaefer, 1987, 1989). presentations and acceptances organize dialogues according to a `collateral’ structure which is separate and independent from the types of structure traditionally studied in computational linguistics (grosz and sidner, 1986; asher and lascarides, 2003), which are focused on what clark calls the ‘official business’ of the conversation; yet they are still (metacommunicative) acts (allwood et al, 1993; bunt, 1995;clark 1996). traum (1994) developed this view by providing a systematic theory of the grounding process articulated around grounding acts, which is the basis of the account of grounding adopted in this paper. traum’s account also provides a solution to the problem of infinite regress incurred by the clark and schaefer formulation of grounding: if every presentation needs an acceptance, how can dialogues ever end? we will discuss traum’s solution in section 3. 2.2 coordination on utterances and completions the typical conversational utterance is not a flawless delivery by the speaker of a complete sentence. speakers generally do not plan utterances in their entirety, and as a result often realize poesio and rieser 4 mid-way that their delivery is being less than ideal (schegloff et al, 1977; clark, 1992, 1996; alwood et al, 1993; traum, 1994; ginzburg, 2009, inter alia). in addition, speakers need to coordinate with listeners and ensure that the signal is understood. as a result, the typical utterance has a non-linear structure with interruptions, and restarts, and with collateral utterances whose objective is to ensure grounding. when a speaker cannot produce a faultless delivery, a disruption may ensue in the form of a repair . such disruptions have a fixed structure that includes a suspension point, often indicated by a pause, a word cut-off, or a lengthened syllable; a hiatus during the period in which the speaker plans how to resume, often occupied with fillers or editing sequences; and a resumption performing an operation that clark calls replacement , which may involve simply continuing what came before the suspension point, repeating it, substituting it, deleting it, or adding new material to it (clark 1996, p. 264). the listener, as well, may interrupt the performance of an utterance, either by initiating a repair when a problem is perceived (see clark’s ‘principle of repair’, clark 1996, p. 284) or by initiating a collaborative completion (lerner, 1987; wilkes-gibbs, 1986) providing their own completion of the utterance, as in (2.1) (from (lerner, 1987), reported in (clark 1996, p. 238)) or a truncation –interrupting the speakers because they think they already understand the rest of the utterance (conversely, a speaker may decide to fade out an utterance when she perceives that the rest isn’t needed.) (2.1) marty: now most machines don’t record that slow. so i’d wanna-when i make a tape, josh: be able to speed it up marty: yeah completions are the focus of this paper. as we will see in section 4, completions may be performed for a variety of reasons, including signaling understanding, or the desire to be cooperative. 2.3 completions in the bielefeld toy plane corpus 2.3.1 the bielefeld toy plane corpus the bielefeld toy plane corpus (btpc) is a collection of 22 filmed, speech recorded and transcribed construction dialogues between two agents, the instructor and the constructor, whose task is to interactively construct a “baufix” toy airplane (see figure 2.1). the instructor (inst) explains to the constructor (cnst) how to assemble the airplane. the airplane that cnst is meant to produce in the conversation in question is shown in figure 2.1 (a), whereas the state of the assembly at the beginning of (1.1) is shown in figure 2.1 (b). (a) the baufix model constructor has to assemble in the dialogue under examination. (b) constructor's state of assembly at the beginning of (1.1) figure 2.1. the baufix plane to be produced by constructor and its state at the beginning of (1.1). completions, coordination and alignment in dialogue 5 instructor and constructor sit at separate tables, have the same collection of “baufix”components (and know this), and can communicate freely. the dialogues were recorded in different conditions so as to set up different contexts for the production and understanding of referring expressions. one dimension of variation concerned visibility: total screen, face to face, half-screen allowing eye contact. different conditions concerning the instructions to be followed by the instructor were also tested. (in some cases the instructor had to direct the constructor on the basis of a building plan, in others using an already completed model.) the corpus consists of 3675 contributions, counting everything except non-verbal events like groans or laughter. 2.3.2 collaborative completions in the btpc: some statistics skuplik (1999) carried out a corpus study of collaborative completions in the btpc. she classified a contribution as as a collaborative completion (or, to use her term, sentence cooperation) if at least two dialogue participants participate in its production. following wilkes-gibbs’s dissertation (1986), skuplik distinguished between two types of sentence cooperations: completions proper, when a sub-sentential structure is filled up by obligatory constituents; and continuations, when material gets added to an already existing sentence. (we will keep using the better-known term completion to cover both completions and continuations, except in this section.) skuplik identified 126 sentence cooperations (54 completions (43%) and 72 continuations (57%)) and classified them along the dimensions (1) producer, (2) type of phrase, (3) grammatical function, (4) syntactic category of completing utterance, (5) resulting syntactic construction, (6) gapping construction, (7) acceptance or denial with respect to part added by other agent, (8) wording of acceptance or denial, (9) indications for change of speaker. the statistics skuplik obtained for the dimensions we are primarily interested in (1, 2, 4, 7, and 9) were as follows. (the figures below are for completions and continuations together, except where noted.) (1) producer: in 79% of the cases it was cnst who produced the completing or the continuing part; inst provided the expansions in only 21% of the cases. (2) type of phrase: 61% of sentence cooperations are complete phrases (german “satzglieder”); (4) syntactic category of completing utterance: prepositional phrases (37%) are the most common form of completing utterance; they are followed by noun phrases (24%), adverbial phrases (7%), nouns (6%) and fragments with finite verbs (5%). (7) acceptance or denial of added part by other agent: 84% of the completing or continuing utterances were accepted by the other participant. 41% of the completing or continuing parts were not explicitly accepted or denied. explicit acceptance is indicated by e.g. ja/yes (28%) and other affirmative particles. in addition, acceptance can be indicated by various forms of resumption or by paraphrase. (9) structural clues for change of speaker: only in 31% of the sentence cooperations the change of speaker is indicated by prosodic or other means such as various forms of hesitation; in the remaing case no clue is apparent. 2.3.3 collaborative completions in the btpc: some statistics the statistical evidence collected by skuplik suggests that (1.1) is a typical illustration of collaborative completions in the bielefeld toy plane corpus. 79% of completions are performed by cnst. cnst’s completion in this example, eine schraube, is an obligatory np (as in 30% of the cases), yielding a single complete construction (i.e.,. german “satzglied”, as in 46% of the cases), making up a sentence if merged with inst’s production (as happens in 50% of the cases). in 31% of sentence cooperations a request for a turn release is signalled by prosodic means such poesio and rieser 6 as lengthening of german du and level tone (37% for completions in the strict sense). here the evidence is not conclusive that the completing part is not accepted by inst (9%), who extends it with an <-> orange one with a slit (9%). inst’s contribution is in turn accepted by cnst. in the second joint construction, 2.1-2.4, we find a use of (german) also in 2.1, which might indicate that a change of the speaker role will be accepted by current speaker. german also frequently indicates a planning pause. however, this time we have a continuation by cnst, forming a single complete sentence unit of category advp, adding finally up to a sentence without extraposition. inst acknowledges by resuming the phrase and extending it with a subordinate clause. in 1.2 inst’s extension acts as a repair, in 2.2 it provides the description of a causal consequence in the domain. 3 incremental meaning composition in dialogue: the ptt approach an essential prerequisite of any account of completions is a theory of semantic interpretation in dialogue explaining how the meaning of fragmentary utterances performed by different speakers is incrementally combined to derive more complex interpretations. our own account is based on such a theory, ptt (poesio, 1994; poesio and traum, 1997; poesio and muskens, 1997; poesio and traum, 1998; matheson, poesio, and traum, 2000) , a theory of dialogue semantics and dialogue interpretation that originated from work on the trains-93 system (allen et al, 1995). ptt was developed to explain how utterances are incrementally interpreted in dialogue, crucially considering both their semantic impact (e.g., how the occurrence of the pronoun “sie” in utterance 2.1 of our example dialogue is interpreted) and their impact on other aspects of dialogue interaction traditionally considered as outside the scope of semantic theory (e.g., the role of the two “ja”s in 1.4 and 2.4), building on the work of clark (1992, 1996) and on ideas from situation semantics (barwise and perry, 1983; cooper and poesio, 1994; cooper, 1996, ginzburg, 2009) . the first distinctive feature of ptt is the assumption—derived from ideas developed in situation semantics but also central to clark’s work, as seen in section 2—that the common ground doesn’t simply record the propositions asserted or the questions raised, but the whole variety of facts about the discourse situation shared between conversational participants (barwise and perry, 1983; see also ginzburg, 2009; ginzburg, to appear). among these facts is the occurrence of certain utterances of sub-sentential constituents in a certain order. furthermore, the theory assumes that the occurrence of these so-called micro-conversational events also leads to immediate updates of the participants’ information states (larsson and traum, 2000; stone, 2004) which in turn leads to the initiation of semantic and pragmatic interpretation processes. other facts recorded in the common ground include what clark called ‘meta-communicative’ acts: dialogue acts which have to do with coordination issues such as turn-taking and the process by which the common ground is established, or grounding (clark and schaefer, 1989; brennan, 2005; traum, 1994). a second feature of ptt that is key for the purposes of this paper is that it provides an explicit account of such meta-communicative acts, and in particular of grounding, building on (traum, 1994). it is assumed in ptt that new utterances result in the introduction of new discourse units—the formal correspondent of clark and schaefer’s (1989) contributions-which only become part of the common ground as a result of explicit or implicit acknowledgments, and may be cooperatively repaired or revised (as in utterances 1.3 and 2.3 of the example dialogue). these ideas are formalized using tools derived from discourse representation theory (drt) (kamp and reyle, 1993) – specifically, from muskens’ compositional drt (muskens, 1996), with additional axioms for specifying the anaphoric behavior of dialogue acts and a simple formalization of events based on (muskens, 1995). (a brief discussion of compositional drt can be found in appendix b.1.) this means that ptt shares many features with sdrt (asher, 1993; asher and lascarides, 2003); but as we will see, there are crucial differences between the two theories with respect to issues central to this paper such as incrementality and grounding. completions, coordination and alignment in dialogue 7 3.1 conversational events and discourse situations the idea that the shared `conversational score' in a conversation consists only of information about the propositional content of assertions, on which much modern work on the semantics of discourse rests, is clearly an idealization (stalnaker, 1978; barwise and perry, 1983; clark, 1992, 1996; allwood et al, 1993; traum, 1994; ginzburg, 2009, to appear). the participants in a conversation also share a great deal of information that they need to coordinate: e.g., whose turn it is to speak, how what is being said fits in within the structure of the rest of the conversation, and whether what has been said has been properly understood (clark, 1996). as a result, an ordinary conversation does not consist only of actions performed to assert or query a proposition, but also of actions whose function is to acquire, keep, or release a turn, to signal how the current utterance relates to what has been said before, or to acknowledge what has just been uttered. bunt (1995) proposed for these utterances the term dialogue control acts. (see also clark and schaefer, 1989; poesio and traum, 1997; ginzburg, 1997, 2009). the execution of these actions may involve both linguistic and non-linguistic tools. the linguistic tools include cue phrases such as so or (one sense of) okay; keep-turn signals such as filled pauses ( umm) or in a minute in the following fragment reported by coulthard (1977): a they have at their disposal enormous assets // and their policy b //look can i just come in on that// last year a //yes in a minute if you may and when i’m finished // then you’ll know b // yes i’m so sorry and grounding signals such as okay again, right or uhuh. non-verbal means include gaze, gestures such as nods or other head movements, and pointing. the context update potential of non-assertoric speech acts and of dialogue control acts is easy to formalize in terms of a speech act-based theory (bunt, 1995; traum and hinkelman, 1992; traum, 1994), particularly one in which speech acts are viewed as components of a joint project (clark, 1996). poesio and traum (1997) proposed that the conversational score consists of a record of all actions performed during the conversation, i.e., what in situation semantics is called the discourse situation (barwise and perry, 1983; cooper, 1992; ginzburg and sag, 2000; ginzburg, 2009, to appear). furthermore, poesio and traum argued that this view of the conversational score could be formalized using the tools already introduced in drt (kamp and reyle, 1993; muskens, 1996), because speech actsconversational events, in ptt termsare in many respects just like any other events, and because conversational events and their propositional contents can serve as the antecedents of anaphoric expressions. so, whereas the ordinary drt construction algorithm would assign to the text in (3.1.1) an interpretation along the lines of (3.1.2) – i.e., a single drs containing the merged propositional content of both assertions (using the syntax from muskens (1996) and his equality operator is for equality in drss) poesio and traum hypothesized that the common ground resulting from (3.1.1) would be as in (3.1.3). 1 (3.1.1) a. a: there is an engine at avon. b. b: it is hooked to a boxcar. (3.1.2) [x,w,y,z,s,s’| engine(x), avon(w), s: at(x,w), boxcar(y), s’:hooked-to(z,y), z is x] (3.1.3) [ce1,ce2| 1 in this paper, discourse referents will be named according to the following conventions. we will use terms with the prefix ce (ce1, ce2, etc) for conversational events; terms with the prefix u for utterances; terms denoting other events will be indicated by the prefix e. we will indicate terms denoting states by the prefix s; all other terms will have prefixes x, w, y, and z. poesio and rieser 8 ce1: assert(a,b,[x,w,s| engine(x), avon(w), s: at(x,w)]) ce2: assert(b,a,[y,z,s’| boxcar(y), s’:hooked-to(z,y), z is x])] (3.1.3) records the occurrence of two conversational events, ce1 and ce2, both of type assert (poesio and traum, 1998; matheson, poesio, and traum, 2000) whose propositional contents are separate drss specifying the interpretation of the two utterances in (3.1.1). the discourse entities ce1 and ce2 can serve as antecedents both of implicit and explicit anaphoric references. implicit anaphoric references include `backward’ acts like accepts, as we will see shortly; certain types of clarification requests (ginzburg, 2009, to appear), and grounding acts, as we will see later in the section. one example of explicit anaphoric reference is the following example, where the that uttered by b in 2 appears to be referring to the action of insulting, as opposed to the propositional content of a’s utterance (as in 2’) or to the locutionary act (as in 2’’). 1. a: you’re an idiot. 2. b: that was uncalled for. 2’ b: that’s not true. 2’’ b: i didn’t hear that. as in (kamp and reyle, 1993; muskens, 1995), a davidsonian treatment of events is assumed, in which each eventor state-describing predicate p such as hooked-to or assert has an additional argument for the event (or state). (we follow kamp and reyle’s (1993) notation and write e:p(x,y) rather than p(x,y,e) for these predicates.) one immediate advantage of this view is that it can be used to explain the updates to the common ground resulting from conversational events other than assertions, as well as to the all too common situation in which an utterance performs more than one type of conversational event (traum and hinkelman, 1992). even if we ignore the fact that interrogatives and imperatives have non-propositional contents (poesio and traum, 1997; ginzburg and sag, 2000, portner 2007)—in this paper, we will for simplicity assume that all contents of conversational events are propositional, as done, e.g., in sdrt (asher and lascarides, 2003)—clearly such contents cannot be viewed as providing restrictions on the same set of assignments / worlds as assertions. in (3.1.3’), for example, neither the content of the open-option conversational event generated by the first utterance in (3.1.1’) nor the content of the info-request resulting from (3.1.1’b) express statements about the current state of the world, and should not therefore be evaluated at that index, as they would be using standard drt semantics if the dialogue were to be assigned the interpretation in (3.1.2’). notice that (3.1.1’b) may be viewed as achieving at least two purposes: accepting the option proposed with ce1, and performing an info-request. notice also that accept is implicitly anaphoric to a previous conversational event (ce1), as is generally the case with backward-looking acts --one argument for assuming that conversational events introduce discourse markers just like normal events do. (3.1.1’) a. a: we should send an engine to avon. b. b: shall we use engine e3? (3.1.2’) [x,w,e, y,e’| engine(x), avon(w), e: send({a,b},x,w) engine(y), e3(y), e’:use({a,b},y)] (3.1.3’) [ce1,ce2,ce3| ce1: open-option(a,b,[x,w,e| engine(x), avon(w), e: send({a,b},x,w)]) ce2: accept(b,ce1) ce3: info-request(b,a,[y,e’| engine(y), e3(y), e’:use({a,b},y)])] completions, coordination and alignment in dialogue 9 assert, open-option and info-request in the example above are all examples of core speech acts (traum and hinkelman, 1992; poesio and traum, 1997, 1998) – conversational events that express the primary, domain-oriented intention the participant intends to convey. the repertoire of core speech acts assumed in ptt has been changing over the years; the later versions of the theory have been based on the damsl repertoire of dialogue acts (allen and core, 1997), and core speech acts for which formalizations are provided in ptt include, in addition to assert and open-option, the forward-looking acts statement, influencing-addressee-future-act, directive, committing-speaker-future-action, commit, and offer , and the backward-looking acts agreement, accept, answer, and reject (matheson, poesio and traum, 2000). in addition to core speech acts, utterances may also be used to perform dialogue control acts and grounding acts; we will see examples of these below. 2 in subsequent work (e.g., (poesio and muskens, 1997)), a revised view of the interpretation of speech acts was introduced in ptt, in which the contents of conversational events are associated with propositional discourse referents (discourse referents whose values are drss) as proposed, e.g., in sdrt (asher, 1993; asher and lascarides, 2003) and by geurts (1995). according to this view, the discourse situation resulting from the two utterances in (3.1.1) being interpreted as performing assertions (and nothing else) would be as in (3.1.4), where the propositional contents of ce1 and ce2 also become available for subsequent anaphoric reference. (3.1.4) [ce1,ce2,k1,k2| k1 is [x,w,s| engine(x), avon(w), s: at(x,w)], k2 is [y,u,s’| boxcar(y), s’: hooked-to(u,y), u is x], ce1: assert(a,b,k1), ce2: assert(b,a,k2)] this type of representation requires complicating the semantics in order to ensure wellfoundedness, but at least in principle, there are a number of ways of doing this (asher, 1993; geurts, 1995; poesio and muskens, 1997); and this representation has the advantage of providing antecedents for references to contents of conversational events—whether explicit as in, e.g., the case in which b follows ce1 with a denial like that’s not true, which in ptt would be interpreted as a reference to proposition k1; or implicit, as in the case of grounding acts. we will assume this type of interpretation here and exploit it in our formalization of grounding acts. finally, it is assumed in ptt that dialogue acts are generated (goldman, 1970; pollack, 1986) by locutionary acts (austin, 1962), which we represent here as events of type utter . these events are assumed to become part of the discourse situation as well, at least for a time;3 indeed, they play an important role in our account of incrementality, as discussed below. we further assume (a) that locutionary acts have a (conventional) semantics associated to them according to standard compositional semantics rules, as we will see in detail in the next subsection; (b) that this semantics is the value of a sem function (in fact, a family of functions sem[], sem[π], etc.); and (c) that the content of the core speech act generated by a locutionary act with semantics k is also k. after taking these additional assumptions into account, we obtain the picture of the information in the discourse situation after the second assertion in (3.1.1) shown in (3.1.5). 2 the version of ptt discussed in (poesio and traum, 1997) also assumes a class of dialogue acts called argumentation acts (traum, 1994) that capture some of the information expressed in sdrt by rhetorical relations. we will not have the space to discuss intentional structure and rhetorical structure in this paper, but these issues are discussed in the follow-up paper (poesio and rieser, in preparation). 3 an important simplification made in ptt is to ignore the issue of forgetting—i.e., the fact that information, particularly linguistic information, only remains `activated’ for a relatively short period. there is quite a lot of evidence that at least certain types of information disappear after a period (sacks, 1967) although the speed at which this happens is unclear (fletcher, 1994). a number of utterances in dialogue – so called informationally redundant utterances (walker, 1993) are also planned with the goal of preventing important information from being forgotten. poesio and rieser 10 (3.1.5) [u1,u2,ce1,ce2,k1,k2| u1: utter (a,”there is an engine at avon”), k1 is [x,w,s| engine(x), avon(w), s: at(x,w)], sem(u1) is k1, ce1: assert(a,b,k1), generate(u1,ce1), u2: utter (b,”it is hooked to a boxcar”), k2 is [y,u,s’| boxcar(y), s’: hooked-to(u,y), u is x], sem(u2) is k2, ce2: assert(b,a,k2), generate(u2,ce2)]] therefore, to a first approximation—i.e., simplifying the representation of contributions and leaving aside all sub-sentential locutionary acts (see next)—the interpretation of the example dialogue according to the view just presented is as follows. (3.1.6) [k1.1, up1.1, ce1.1, k2.1, up2.1, ce2.1 | k1.1 is [e,x,x3| screw(x), orange(x), slit(x3), has(x,x3), e:grasp(cnst, x)], utterance(up1.1), sem(up1.1) is k1.1, ce1.1:directive(inst&cnst,cnst,k1.1), generate(up1.1, ce1.1), k2.1 is[x6,e’,s,w,y| x6 is x, e’:put-through (cnst,x6,hole1), w is wing1, y is fuselage1, s: fastened(w,y), purpose(e’,s)], utterance(up2.1), sem(up2.1) is k2.1, ce2.1:directive(inst,cnst, k2.1) , generate(up2.1, ce2.1)] 3.2 micro conversational events a second distinctive feature of ptt, in particular with respect to sdrt (asher and lascarides, 2003), is that it takes as a central fact about dialogue that utterances are interpreted incrementally, as suggested by most psychological work on sentence processing (frazier, 1987; swinney, 1979; tanenhaus et al, 1995), and that many contributions to dialogue are fragmentary and nonsentential, as shown by corpus evidence (poesio, 1995; fernandez, 2006). as a result, one of the fundamental hypotheses underlying ptt is that the information state of a conversational participant is updated whenever any new event is perceived, including events such as sub-sentential or even sub-word utterances, and irrespective of whether this event is verbal or non-verbal: non-verbal events such as gestures or nods (mcneill, 1992; allwood et al, 1992; clark, 1996; rieser, 2004) update the information state, as well. psychological research suggests that such updates can take place every few milliseconds, and that perceiving a phoneme is sufficient to cause an update—in fact, for interpretive processes to begin (tanenhaus et al, 1995). here we will simply assume that the view of the discourse situation contained in the information state is updated at least after every word. we use the term micro conversational events (mces) to refer to events of uttering sub-sentential constituents (poesio, 1995). also, we will only discuss here updates caused by linguistic events, although other types of updates have been examined in past ptt work – e.g., to the focus of visual attention (poesio, 1993) and pointing gestures (rieser, 2004). this incremental update hypothesis is motivated not only by psychological findings about incremental interpretation in sentential utterances, but also by the fact that in dialogue many types of conversational acts hardly if ever require full sentences (clark, 1996; fernandez, 2006; ginzburg, 2009). among the core speech acts, answers to questions and clarification questions completions, coordination and alignment in dialogue 11 are often non-sentential. the utterances used to perform dialogue control acts such as taketurn , keep-turn and release-turn –actions whose function is to synchronize the two participants in the conversation as to who is holding the floor (sacks et al 1974; traum and hinkelman, 1992; traum, 1994; bunt 1995; clark, 1996)—and grounding acts, i.e., acts whose function is to keep the common ground synchronized between the two participants (grounding acts are discussed in greater detail below) are also typically non-sentential. these conversational actions are occasionally generated by sentential utterances that also generate a core speech act (e.g., the second utterance in (3.1.1)), but more commonly they are generated by single-word discourse markers—like okay, well, now, all of which may be used to perform a keep-turn dialogue control act, or indeed jetzt in 1.1 in the example dialogue (1.1), which arguably is used to perform a keep-turn function as well as (possibly) a temporal sequencing one. non-words such as filled pauses (umms and the like) may also be used (sacks et al, 1974; clark 1996). to account for the fact that non-sentential utterances may result in updates of the discourse situation it is necessary to have a theory of update in which the effects of these utterances can be modelled. in ptt, it is hypothesized that observing such utterances may result in adding to the discourse situation keep-turn events (as well as other possible events), via updates like those in (3.2.1): (3.2.1) well→ [u,ce|u: utter (a,”well”), ce: keep-turn(a), generate(u,ce)] umm→ [u,ce| u: utter (a,”umm”), ce: keep-turn(a), generate(u,ce)] in this paper we assume that the utterances of so at the beginning of 1.1 in the example dialogue (1.1) has primarily a dialog control function, so that observing its occurrence results in the following update to the discourse situation. (see section 5 and appendix b.2.3.) (3.2.2) so→ [u,ce|u: utter (a,”so”), ce: take-turn (a), generate(u,ce)] we instead assume that jetzt in the same utterance is interpreted solely as contributing to the specification of the propositional content of ce1.1, the core speech act generated in that contribution (see (3.1.6)). of course one could also hypothesize that jetzt has a keep-turn function instead, as in (3.2.2’), or perhaps both functions, but this is not essential for our purposes. (3.2.2’) jetzt→ [u,ce|u: utter (cnst,”jetzt”), ce: keep-turn(a), generate(u,ce)] 3.3 incremental interpretation with micro conversationa l events by combining the hypothesis that the information state is updated after every microconversational event with the view of semantic composition found in compositional drt (as opposed to that of standard drt) we can explain how utterances whose main function is to express part of the content of a core speech act, like the rest of the utterances in 1.1 nimmst, du, etc. do so incrementally, as well, instead of assuming that (3.1.5) or (3.1.6) are derived all at once after the entire mini-dialog in (3.1.1) or the example dialogue have been syntactically analyzed (as it would be necessary when assuming the construction algorithm from (kamp and reyle, 1993)). as shown by muskens (1996), using only tools already present in compositional drt one could already show how (3.1.5) could be derived compositionally and incrementally by concatenating separately produced interpretations for the two utterances in (3.1.1), as in (3.3.1): (3.3.1) [ce1, k1| k1 is [x,w,s| engine(x), avon(w), s: at(x,w)], ce1: assert(a,b,k1)] ; [ce2, k2 | k2 is [y,u,s’| boxcar(y), s’: hooked-to(u,y), u is x], ce2: assert(b,a,k2)] but if we accept the two assumptions that locutionary acts are recorded in the common ground, and that single word utterances can lead to updates as well, we can take advantage of compositional drt to explain how the utterances of nimmst and du can initiate syntactic and semantic interpretation before an entire sentence has been perceived. psychological research on priming at different levels (as summarized, e.g., in pickering and garrod, 2004) and work on poesio and rieser 12 clarification questions such as (ginzburg and cooper, 2004; purver and ginzburg, 2004) provide evidence concerning these updates. this evidence suggests that perceiving an utterance results in the record of the discourse situation being updated with the fact that this utterance just occurred, as well as the results of lexical access – i.e., that utterance’s syntactic classification, and its conventional meaning, which in ptt is identified with the compositional meaning as specified in compositional drt (poesio, 1995; poesio and traum, 1997; poesio and muskens, 1997; poesio, to appear). observing an utterance of the noun boxcar, for example, results in an update of the discourse situation which, in first approximation, we will represent as in (3.3.2). this update records the utterance of a new locutionary act u (as said above, we use discourse referents with prefix u to indicate utterances), of type utter (the type of locutionary acts), syntactically classifed as a noun, and with semantic content λx [|boxcar(x)]. (3.3.2) [u | utter (a,”boxcar”), noun(u), sem(u) is λx [|boxcar(x)]] (we will often use the abbreviated notations u:”boxcar”:noun to indicate the information added by the utterance of a word to the discourse situation, omitting its lexical semantics, and u:”boxcar”: λx [|boxcar(x)] that specifies its lexical semantics but omits its syntactic interpretation). a key assumption that ptt derives from psycholinguistic results on lexical access (swinney, 1979; tanenhaus et al, 1979) and shares with modern computational linguistics is that lexical access—in fact, all of utterance interpretation—is a process of defeasible inference in which competing hypotheses are activated, one of which is rapidly selected, whereas the other ones are discarded (poesio, 1994, 1995, to appear). a good case can be (and has been) made that the defeasible inferences that constitute language interpretation are a form of statistical inference (jurafsky, 1996), and all recent work in computational linguistics is based on this assumption. however, it is still an open problem how to combine the logics used in formal semantics with statistical inference (for preliminary work on the matter, see, e.g., (hwang and schubert, 1993)), so ptt follows the more traditional approach adopted by virtually all theoretical approaches attempting to combine a theory of performance with a theory of semantic competence based on formal semantics in using a form of logic to model defeasible inference (perrault, 1990; hobbs et al, 1992; hwang and schubert, 1993; poesio, 1994, 1995; asher and lascarides, 2003). specifically, in ptt interpretation is modeled in terms of prioritized default logic (pdl) (brewka 1989). we would like to be very clear however that we are not claiming here that natural language interpretation is a process of inference in pdl: only that pdl is the simplest and more standard theory of nonmonotonic inference that makes it possible to us to provide an explicit account of the properties of natural language interpretation in which we are interested here – namely, of the conclusions that we expect to be derived by the ‘real’ processes behind language interpretation (which are almost certainly specialized in carrying out certain types of inference and therefore, for instance, not likely to suffer from the efficiency problems of pdl). the lexicon in ptt is thus modeled as a default theory consisting of prioritized default rules specifying lexical updates activated by each utterance, such as (3.3.2).4 the ptt view of the interpretive processes that follow lexical access – syntactic interpretation (parsing) and semantic composition—is very much inspired by current work on grammar in frameworks like tree adjoining grammars (abeillé and rambow, 2000; joshi, 2004) for what concerns incremental interpretation, and categorial grammar (pereira, 1990; carpenter, 1994) in that syntactic interpretation is also viewed as an inferential process, the only difference being that in ptt interpretation is viewed as an inferential process resulting in updates of the discourse situation. 4 a more complex view of defeasible inference is adopted in sdrt, with multiple (default) logics for different types of interpretation. a much more radical reconsideration is however needed to incorporate insights from recent work in psycholinguistics and computational linguistics. for discussion see (poesio, to appear). completions, coordination and alignment in dialogue 13 specifically, parsing in ptt is a process during which hypotheses about the results of lexical access combine together in phrasal hypotheses through default inference. such phrasal hypotheses are viewed as hypotheses about utterances of phrases: e.g., the occurrence of contiguous utterances of type det and n results in a phrasal hypothesis about the occurrence of an utterance of type np. these hypotheses are the result of a second set of defeasible inference rules that encode syntactic competence.5 as in (poesio, 1995; poesio, 2001; poesio, to appear), we assume here the syntactic framework of lexicalized tree adjoining grammar (ltag) (schabes et al, 1988; schabes, 1990; sturt and crocker, 1996; abeillé and rambow 2000), since it lends itself to a very natural account of the process by which syntactic interpretations are constructed incrementally (sturt and crocker, 1996).6 in ltag, the lexical interpretations of words are elementary trees. in the case of sortal nouns like boxcar or schraube these trees are atomic and take the form in (3.3.3): (3.3.3) n | schraube in the case of words whose semantic interpretation takes arguments, such as verbs and determiners, the elementary trees are more complex, and already contain ‘attachment points’ for such arguments. the example in figure 3.3.1 illustrates the lexical interpretation of the determiner eine according to ltag. in ptt, the utterance of this determiner results in the discourse situation being updated not just with the observation that an utterance uspec occurred, but also with the expectation that uspec is going to be part of the performance of the utterance u of an np, of which uspec will occupy the specifier position, as well as the performance of an utterance u of type n and (possibly) of a complement ucompl. (this node is the attachment point for nominal arguments of relational nouns.) figure 3.3.1. the update to the discourse situation resulting from an observation of article eine the elementary trees introduced into the discourse situation by the performance of utterances of single words are combined by means of the two basic tag operations: substitution 5 this view of parsing as defeasible inference in prioritized default logic is not of course intended as an alternative to the modern view of parsing as statistical inference, but only as the simplest possible account of how the interpretations we are proposing could be obtained via defeasible inference. 6 the ltag framework has also been adopted in other modern frameworks concerned with semantic interpretation, such as muskens’ logical description grammar (muskens, 2001), although it is not used there to provide an account of incremental processing. uspec u u u ucompl u :np uspec :det u: n “ eine” sem(uspec)= λp’λp([y| ]; p’(y); p(y)) u:n ucompl poesio and rieser 14 and adjunction. substitution is the operation by which arguments—the `expected’ components of the syntactic structure of an utterance, such as the nominal head of the np expected as a result of the observation of eine—are `slotted into’ non-atomic elementary trees. for example, assuming an ltag translation for schraube analogous to that for boxcar, observing an utterance of schraube, and subsequently substituting the associated elementary tree into the interpretation in figure 3.3.1, results in the interpretation in figure 3.3.2. adjunction is the operation by which adjuncts and non-arguments are incorporated into the syntactic interpretation being inferred. in our fragment for the example dialogue, we propose that adjunction is used to incorporate appositions, i.e., to attach eine orangene, mit einem schlitz to eine schraube. we hypothesize for this apposition the syntactic structure in (3.3.4)(b); this structure is attached to the interpretation of np eine schraube in (3.3.4)(a) by ‘splitting’ the n’ node into two nodes, resulting in the structure in (3.3.4)(c). (3.3.4) (a) np (b) n’ (c) np eine n’ n’ npapp eine n’ schraube eineapp n’ n’ npapp nadj schraube eineapp nadj tag in general, and ltag in particular, have become widely used by psychologists as a syntactic framework in that adjunction provides an explanation for how modifiers can be incrementally attached to a single syntactic interpretation (sturt and crocker, 1996). ferreira and colleagues have argued that ltag also provides an ideal framework to account for disfluencies (ferreira et al, 2004). a complete ltag fragment for the example dialogue is provided in appendix b; the treatment of appositions assumed in this paper is discussed in b.4.2 uspec u u u ucompl u :np uspec :det u: n “ eine” sem(uspec) is λp’λp([y| ]; p’(y); p(y))] u:n ucompl “schraube” sem(u)= λv([ |screw(v)]] figure 3.3.2. the updates resulting from the observation of eine schraube, after substitution of the elementary tree for schraube into 3.3.1 completions, coordination and alignment in dialogue 15 the remaining aspect of the interpretation process, semantic composition—the process by which phrasal utterances receive an interpretation—is also viewed in ptt as an inference process (poesio, 1995; poesio, to appear), in accordance with the view adopted in categorial grammar and related work on `parsing as inference’ (pereira, 1990; carpenter, 1994), where the combination of utterances in larger utterances and the specification of the meaning of these larger utterances are provided by inference rules. ptt hypothesizes that semantic composition, as well, is the result of defeasible inferences over the drs obtained by concatenating the updates resulting from the utterances of single words (poesio, to appear). these default inference rules have the effect of the semantic composition rules introduced by muskens (1996) for compositional drt. for example, the rule binary semantic composition below specifies that if u1 and u2 are the (only) two constituents of u3 (we used ↑ to indicate dominance), one of them (say, u1) has a semantic interpretation of type 〈α,β〉, and the other has a semantic interpretation of type α, the semantic interpretation of u3 is derived by applying the semantic interpretation of u1 to that of u2. (cfr. muskens’ application rule (muskens, 1996, p. 166).) we also assume a unary semantic composition inference rule achieving the effect of muskens’ copying rule. binary semantic composition (bsc) : β α βα ψϕ ψ ϕ )()3( )2( ,)1( ,32,31 , = = = ↑↑ usem usem usem uuuu we can now return to our example dialogue (1.1) and see how it is interpreted according to ptt, using the ltag + compositional drt grammar from appendix b together with the microconversational events hypothesis. we will deal here with a slightly simplified version of the series of utterances that result in the first directive: jetzt nimmst du eine orangene schraube mit einem schlitz, ignoring so, merging cnst’s completion (eine schraube) with inst’s refashioning of it via an apposition (eine orangene … mit einem schlitz) (completions and refashioning are discussed in section 5; appositions in appendix b.2). each of these micro conversational events causes an incremental update of the discourse situation; the concatenated sequence of updates is shown in diagrammatic format in (3.3.5) and using linear notation in (3.3.5’), where we have used ↑ to indicate dominance. we used the prefix up to name phrasal utterances and the prefix ub to name xbar projections.7 (3.3.5) [u1.2, up1.1, up1.2 | up1.1:s u1.2:“jetzt“:advp up1.2:s sem(u1.2) is λp. now(p)]; [u1.3, up1.4 , up1.2’, up1.3, up1.5, up1.6| up1.2’:s up1.3:np up1.4:vp u1.3:”nimmst”:v up1.5:np up1.6:np sem(u1.3)= λqλx(q(λx’[e| e: grasp(x, x’)]))]; 7 the complete lexical and grammatical rules for the fragment of german we are considering, broadly based on muskens (1996), are given in appendix b. poesio and rieser 16 [u1.4, up1.5’ | up1.5’:np u1.4:”du”:pro sem(u1.4)= λp.p (you)]; [u1.5, up1.6’, ub1.6, u1.6’| up1.6’:np u1.5:”eine”:det ub1.6:nbar sem(u1.5)= λp’λp([y1|];p’(y1);p(y1)) u1.6’:noun]; [u1.8, ub1.6’, ub1.6’’ | ub1.6’:nbar u1.8:”orangene”:adj ub1.6’’:nbar sem(u1.8) = λpλz([ |orange(z)]; p(z))]; [u1.6| u1.6:utter (cnst,"schraube"), noun(u1.6), sem(u1.6)= λv([ |screw(v)]]; [u1.9, ub1.6’’’, ub1.6’’’’, up1.9, up1.10| ub1.6’’’’:nbar ub1.6’’’:nbar up1.9:pp u1.9:”mit”:prep up1.10:np sem(u1.9)= λp λy(p (λx[ |with(x,y)]))]; [u1.10, up1.10’, ub1.10, u1.11’| up1.10’:np u1.10:”einem”:det ub1.10:nbar sem(u1.10)= λp’λp([y2| ]; p’(y2); p(y2)) u1.11’:noun]; [u1.11| u1.11:utter (inst,"schlitz"), noun(u1.11), sem(u1.11)= λv([ |slit(v)]] completions, coordination and alignment in dialogue 17 (3.3.5’) [u1.2, up1.1, up1.2 | u1.2: utter (inst, “jetzt“), advp(u1.2), sem(u1.2) is λp. now(p), s(up1.1) , s(up1.2), u1.2 ↑ up1.1, up1.2 ↑ up1.1]; [u1.3, up1.4 , up1.2’, up1.3, up1.5, up1.6| u1.3: utter (inst,"nimmst"), verb(u1.3), sem(u1.3)= λqλx(q(λx’[e| e: grasp(x, x’)])), vp(up1.4), u1.3 ↑ up1.4, s(up1.2’), up1.4 ↑ up1.2’, np(up1.3), up1.3 ↑ up1.2’, np(up1.5), up1.5 ↑ up1.4, np(up1.6), up1.6 ↑ up1.4]; [u1.4, up1.5’ | u1.4: utter(inst,"du"), pro(u1.4), u1.4 ↑ up1.5’, np(up1.5’), sem(u1.4)= λp.p (you)]; [u1.5, up1.6’, ub1.6, u1.6’| u1.5: utter(inst,"eine"), det(u1.5), sem(u1.5)= λp’λp([y1| ]; p’(y1); p(y1)), np(up1.6’), nbar(ub1.6), noun(u1.6’), u1.5 ↑ up1.6’, u1.6’ ↑ ub1.6, ub1.6 ↑ up1.6’]; [u1.8, ub1.6’, ub1.6’’ | u1.8: utter (inst, “orangene”), adjp (u1.8), sem(u1.8) = λpλz([ |orange(z)]; p(z)), u1.8 ↑ ub1.6’, ub1.6’’ ↑ ub1.6’, nbar(ub1.6’), nbar(ub1.6’’)]; [u1.6| u1.6:utter (cnst,"schraube"), noun(u1.6), sem(u1.6)= λv([ |screw(v)]]; [u1.9, ub1.6’’’, ub1.6’’’’, up1.9, up1.10| u1.9:utter (inst,”mit”), prep(u1.9), sem(u1.9)= λp λy(p (λx[ |with (x,y)])), pp(up1.9), nbar(ub1.6’’’), nbar(ub1.6’’’’), np(up1.10), u1.9 ↑ up1.9, ub1.6’’’ ↑ ub1.6’’’’, up1.9 ↑ ub1.6’’’’, up1.10 ↑ up1.9]; [u1.10, up1.10’, ub1.10, u1.11’| u1.10:utter (inst,"einem"), det(u1.10), sem(u1.10)= λp’λp([y2| ]; p’(y2); p(y2)), np(up1.10’), nbar(ub1.10), noun(u1.11’), u1.10 ↑ up1.10’, ub1.10 ↑ up1.10’, u1.11’ ↑ ub1.10]; [u1.11| u1.11:utter (inst,"schlitz"), noun(u1.11), sem(u1.11)= λv([ |slit(v)]] after each of these updates, parsing and semantic composition inference rules apply to combine phrasal utterances together through substitution and adjunction and to assign an interpretation to such phrasal utterances. let us consider for instance how the updates due to the observation of “einem” (micro conversational event u1.10 in 3.3.5) and “schlitz” (mce u1.11) result in the hypothesis that an utterance of the np “einem schlitz” was observed, and in the assignment of an interpretation to that phrasal utterance. the hypothesis that the utterance of u1.10 and u1.11 in (3.3.5) is part of the utterance of an np leads to a substitution inference: the hypothesis that u1.11’ (expected after observing an utterance of einem) is the same as u1.11. the formulation of this hypothesis results in the following update of the discourse situation: [ | u1.11’ is u1.11] poesio and rieser 18 applying bsc to the discourse situation thus updated results in the assignment of the conventional meaning λ p.[y2| ];[| slit(y2)]; p(y2) to up1.10’, which results in the following update: [ | sem(up1.10’) is λ p.[y2| ];[| slit(y2)]; p(y2)] a similar process leads to hypothesizing the rest of the structure; the only difference is that attaching orangene and mit einem schlitz requires adjoining the pp resulting from processing mit einem schlitz into the nbar resulting from eine orangene: the result can be seen in figure 3.3.3. at the end of this process of unification of the phrasal utterances, we obtain the structure shown in figure 3.3.4, from which meanings have been omitted (but see appendix b and section 5). u u uspec u ucompl u :np uspec :np u:nbar u:nbar ucompl: pp “ orangene” “mit einem schlitz” sem(u)= λ p λz ([|orange(z)];p(z) λx([y|];[|slit(y)];[|with (x,y)] figure 3.3.3. the update resulting from the adjunction of ”mit einem schlitz” to “orangene” completions, coordination and alignment in dialogue 19 we will reiterate at this point one of the central claims of this paper, which is that many of the syntactic, semantic and pragmatic properties of completions can be accounted for relying on independently motivated formal devices incorporated in ptt. in this subsection we have shown, specifically, that explaining semantic composition in these constructions does not require developing novel syntactic or semantic formalisms, but only establishing a link between utterances in the discourse situation and syntactic phrases—a link already explicitly proposed in situation semantics, hpsg, and in ginzburg’s kos, for independent reasons—and translating lexical composition and parsing in terms of inferences on the discourse situation. we will next see that we do not need a novel treatment of grounding, before discussing two ways of explaining the reasons for making a completion in sections 4 and 5 (intentional analysis) and 6 (alignment analyses). to conclude, another bit of notation. in ptt, like in other dynamic theories of meaning, semantic interpretation and other inference processes generally result in adding to a drs new material, which in general can also include new discourse referents. for instance, the ptt equivalent of existential instantiation of predicate p in the context of drs k results in k being augmented with the drs [x | p(x)]. we will use the notation k+= k’ for indicating this operation of updating drs k with k’ when k is a propositional discourse referent and k’ is either a proposition or a proposition valued discourse referent. the effect of this operation is to change the value of k by adding k’. for instance, suppose the value of k is [y | q(y)]. then existential instantiation of p within k as above results in the update: k += [x | p(x)] figure 3.3.4. simplified representation of the complete syntactic structure assumed for the first contribution up1.1:s u1.2:”jetzt”:advp up1.4:vp up1.3:np u1.3:”nimmst”:v up1.5:npi up1.6:np u1.4:”du”:pro εi up1.2:s u1.5:”eine”:det ub1.6:n’ ub1.6’:n’ u1.8:”orangene”:adjp up1.10:np ub1.6’’:n’ up1.9:pp u1.6:”schraube”:n u1.9:”mit”:p u1.10:”einem”:det ub1.10: n’ u1.11:”schlitz”:n poesio and rieser 20 after which the value of k becomes [y,x | p(x), q(y)]. drs update is defined as follows.8 k += k’ let k be a proposition-valued discourse referent, and k’ be a proposition (drs) or a proposition-valued discourse referent. then k += k’ =def [ tmpk| tmpk is k];[ k| k is tmpk;k’] (where tmpk is an unused proposition-valued discourse referent). by the definition of drs and ; in cdrt (see appendix b.1) this is equivalent to: k += k’ =def λiλj ∃l i[tmpk] l ∧ l[k]j ∧ v(tmpk)(l) = v(k)(l) ∧ v(k)(j) = v([λi’ λj’ ∃l’ tmpk(i')(l’) ∧ k’(l’)(j’)])(j) 3.4 grounding and discourse units unlike other dynamic theories of interpretation, ptt does not rely on the assumption that every utterance automatically becomes part of the common ground; instead, it includes an explicit formalization of the grounding process discussed in section 2 (clark and schaefer, 1989; brennan, 1991; traum and hinkelman, 1992; traum, 1994; clark, 1996). many of the utterances in the example dialogue were most likely intended to play a role in this process; and completions themselves can be viewed as a particularly explicit form of acknowledgment, as discussed in section 2. the formalization of grounding developed in ptt is based on clark and schaefer’s (1989) proposal discussed in section 2, as modified by traum (1994). according to clark and schaefer, a conversation consists of a series of contributions which have to be acknowledged, possibly implicitly (thus becoming part of the common ground) or may need further clarifications and repairs. traum (1994) and matheson, poesio and traum (2000) developed this proposal by providing an account based on the assumption that grounding is achieved through a particular type of dialogue control acts called grounding acts. following (traum, 1994), we use the term discourse unit (du) to refer to a contribution. and again as in (traum, 1994), it is assumed in ptt that every utterance in a conversation either initiates a new du, continues a du, acknowledges a du, performs a repair , or requests the other participant to perform one of the grounding acts above. at any point in a conversation a conversational participant may begin a new contribution, i.e., initiate a new du; this new du gets added to the semi-public part of the information state. this is the case for example with 1.1 or 2.1 in the example dialogue. as we will discuss in detail in section 5, in 1.2 cnst simultaneously acknowledges the part of the new contribution which has already been introduced, grounding it—possibly in response to a perceived request from inst—and adds new material to the contribution. in 1.3, inst performs what clark and wilkes-gibbs called a refashioning of the contribution—adding new material which may also lead to a revision (e.g., to the choice of a new screw). (in ptt, this type of operation is viewed as a type of repair , as proposed by levelt (1989) and clark (1996). in fact, the introduction of new material in 1.2 is seen as a repair as well, as discussed in section 5.) finally, cnst acknowledges the remaining part of the contribution, which then is fully grounded, and accepts the directive specified by the full contribution.9 the accept itself is viewed in ptt as a second contribution. in 2.1, inst (implicitly) grounds the accept by initiating a new contribution (a new directive). as in (poesio and traum, 1997; matheson, poesio, and traum, 2000), we assume that discourse units are dynamic propositions about the discourse situation, i.e., drss containing the 8 a different definition of += was given in (poesio & traum, 1998). 9 notice that in ptt acknowledgments—a type of grounding act—and accepts—a backward act resulting in the speaker’s assuming the obligation to perform a certain action—are distinct. completions, coordination and alignment in dialogue 21 type of information about conversational events (whether micro conversational events, core speech acts, and other types of dialogue acts) that we have discussed in the rest of this section. in fact, in recent work we adopted the position that each of the micro-conversational update to a discourse situation we discussed earlier in this section constitutes a discourse unit, but here we will stick with the position adopted in (poesio and traum, 1997; poesio and traum, 1998), in which all the updates of the discourse situation related to the process of grounding a particular contribution are part of the same discourse unit. in (poesio and traum, 1997, 1998) the agent’s information state itself was formalized as a drs. this drs includes, in addition to information about the private attitudes of the agent, a record of the contributions (dus) so far, as well as a special drs g containing the material which has already been grounded. for example, the information state after interpreting and grounding the first contribution in (3.1.1), and after interpreting the second sentence (i.e., creating a du du2 for it) but before grounding it, would be as follows (ignoring micro-conversational events): (3.4.1) [du1,du2| du1 is [ce1, k1| u1: utter (a,”there is an engine at avon”), k1 is [x,w,s| engine(x), avon(w), s: at(x,w)], sem(u1) is k1, ce1: assert(a,b,k1), generate(u1,ce1)], g is [ce1, k1| u1: utter (a,”there is an engine at avon”), k1 is [x,w,s| engine(x), avon(w), s: at(x,w)], sem(u1) is k1, ce1: assert(a,b,k1), generate(u1,ce1)], du2 is [ce2,k2| u2: utter (b,”it is hooked to a boxcar”), k2 is [y,u,s’| boxcar(y), s’:hooked-to(u,y), u is x], sem(u2) is k2, ce2: assert(b,a,k2), generate(u2,ce2)]] there are two main reasons for viewing the information state as a drs. first of all, all grounding acts are implicitly anaphoric, in that they refer to particular dus, as we will see below. secondly, in compositional drt the modifications to g and the dus resulting from grounding acts can be modeled quite simply as updates to the values of discourse markers. for example, the effect of an acknowledgment on the information state was formalized by poesio and traum (1998) as replacing the previous value of g with a new drs which is the merge of g and du1 using the += operator just introduced: (3.4.2) g+= du1 in this paper, however, we will use a modal operator g to assert that a particular du is grounded. the main reason for the change is that this formulation is closer to that adopted in more recent work by traum (1999) in which grounding is not viewed as an all-or-nothing affair, but as a matter of degree. this type of theory is more easily modeled using one or more modal operators to specify grounding. the modal operator g expresses a stronger form of mutual knowledge than standard mk, in that everything that is grounded (acknowledged) is mutually known, but not vice versa. we will not provide a full axiomatization of g here, but we will require that everything that is grounded is mutually known: [ax-g-1] ∀ du g(du) → mk(du) poesio and rieser 22 according to this new view, the information state resulting from the first contribution in (3.1.1) being grounded, while the second one still isn’t, is as in (3.4.3), instead of as in (3.4.1): (3.4.3)[du1,du2| du1 is [ce1, k1| u1: utter (a,”there is an engine at avon”), k1 is [x,w,s| engine(x), avon(w), s: at(x,w)], sem(u1) is k1, ce1: assert(a,b,k1), generate(u1,ce1)], g(du1), du2 is [ce2,k2| u2: utter (b,”it is hooked to a boxcar”), k2 is [y,u,s’| boxcar(y), s’:hooked-to(u,y), u is x], sem(u2) is k2, ce2: assert(b,a,k2), generate(u2,ce2)]] according to this version of the theory, in the case of the example dialogue we obtain the following information state after the first contribution (the directive jointly produced by inst and cnst in 1.1 – 1.3) and the second (the acceptance produced by cnst in 1.4) are grounded, but before the second directive is grounded, again ignoring micro-conversational events and phrasal utterances (compare with (3.1.6)): (3.4.4) [du1.1, ce1.6, du1.4, ce2.2, du2.1 | du1 is [k1.1, up1.1, ce1.1, | k1.1 is [e,x,x3|screw(x), orange(x), slit(x3), has(x,x3), e:grasp(cnst, x)], utterance(up1.1), sem(up1.1) is k1.1, ce1.1:directive(inst&cnst,cnst,k1.1), generate(up1.1, ce1.1)] ce1.6 :ack(cnst,du1.1), g(du1.1), du1.4 is [ce1.7, s1.1 | ce1.7: accept(cnst,ce1.1), s1.1 : obl(cnst,k1.1) ], ce2.2 :ack(inst,du1.4), g(du1.4), du2.1 is [k2.1, up2.1, ce2.1 | k2.1 is [x6,e’,s,w,y| x6 is x, e’:put-through (cnst,x6,hole1), w is wing1, y is fuselage1, s: fastened(w,y), purpose(e’,s)], utterance(up2.1), sem(up2.1) is k2.1, ce2.1:directive(inst,cnst, k2.1 ) , generate(up2.1, ce2.1)]] notice that grounding acts ce1.6 and ce2.2 are not included in the discourse units. this is the solution proposed by traum (1994) to the ‘bottoming out’ problem present in clark and schaeffer’s work—explaining how information about the occurrence of grounding acts is unlike other information in the discourse situation in that it doesn’t seem to require grounding (else we would have an infinite regress). notice also that in all accounts of grounding derived from clark and schaeffer’s work, planning a core speech act really amounts to planning a contribution, i.e., opening a new du. hence such theories need to stipulate an additional step of intentional reasoning: that to intend to perform a core dialogue act ce it is to intend to make a contribution du with that act as content. this is expressed by the following schema. completions, coordination and alignment in dialogue 23 [ce-to-du-schema] : intcp1([ce | ce:core-act(cp1,cp2,φ)]) � intcp1([du | du is [ce | ce:core-act(cp1,cp2,φ)]]) we will see examples of application of the ce-to-du schema in section 5. grounding acts can be formalized in terms of operations on the information state. in the version of ptt adopted here, i. performing an init (du) simply means introducing a new du in the information state. we assume an init (du) every time a new contribution is initiated; such grounding acts are not explicitly recorded in the information state. ii. performing a cont(du) means adding to an existing du; again, we do not explicitly represent this grounding act in the information state. iii. acknowledging a du has the effect of grounding the du, i.e., adding g(du) to the information state; iv. repairing du1 with du2 means updating the discourse referent du1 by assigning to it du2 as value. v. requesting a grounding act means adding to the information state an obligation to address that request. this formulation does not specify preconditions for such acts—e.g., what it means for a contribution to be understood, and therefore when are acknowledgments warranted (clark’s grounding criterion ). clark points out that a contribution may fail to be understood at several levels: e.g., the addressee may have not heard what the speaker said, or may have heard it but not know its meaning. in this paper we assume that an utterance has been understood when values for all the functions that specify its linguistic classifications –phonetic, syntactic, and semantic (sem)—can be recovered either from the lexicon or from the context; but we will not provide update rules for grounding acts including such preconditions. (ginzburg’s (2009) proposals for information state update rules for grounding acts specifying such information and covering clarification requests as well are discussed in section 7.) matheson, poesio and traum (2000) discuss accepting in ptt for the case of core speech acts. update rules for core speech acts are conditional upon acceptance—meaning that for each core speech act there is a conditional update rule specifying the part of the update caused by that speech act that depends upon acceptance. for instance, for the case of directives, the conditional update rule specifies that in case ce is a directive by a to b with content k, then if b accepts ce, b assumes the obligation to bring it about that k: schematically, [up.directive ] ce:directive(a,b,k) � (accept(b,ce) � [o | o:obl(b,k)]) (notice that an update rule is not a material implication, but a rule to update the information state by concatenating new information.) we assume here that update of the information state as the result of grounding acts works in the same way, except that in the case of grounding acts, accepting leads to updates of the information state that affect dus. we use two such rules in this paper, up.repair for repairs, and up.ack for acknowledgments. to specify the update resulting from up.repair we use the following abbreviation: du1 ⇐ du2 let du1 be a propositional variable specifying a discourse unit, and du2 be a drs. then du1 ⇐ du2 =def λiλj i[du1]j ⇐ v(du1)(j) = v(du2)(i) the two update rules specifying the behavior of grounding acts, then, are as follows: [up.repair] ce:repair (a,du1,du2) � (accept(b,ce) � du1 ⇐du2) (with du1 a propositional discourse referent) [up.ack ] ce:ack(a,du) � (accept(b,ce) � [ | g(du)]) (with du a propositional discourse referent) poesio and rieser 24 so, for instance, the result of the ja in 1.4 is to add to the information state g(du1.1) when the acknowledgment is accepted (see (3.3.4)), whereas the result of accepting the repairs in 1.2 and 1.3 is to add to du1.1 the new material, as discussed in section 5. 3.5 the information state as already mentioned, ptt is an information-state based theory in the sense of (cooper et al 1999, larsson & traum 2000, stone 2004, ginzburg 2009). ginzburg (2009, chapter 4) discusses different views concerning what information states such a theory may model; ptt aims at modeling the information state of a single agent, assumed to consist of three main parts: a private part, with information available to the participant, but not introduced in the dialogue. this part includes private beliefs and intentions of that participant, as well as hypotheses about beliefs and intentions of other agents. a public part consisting of the dus that are assumed by that agent to have become part of the common ground. a semi-public part, consisting of the information introduced with contributions that haven’t yet been acknowledged. this information is not yet grounded, but it is accessible. so far, we have discussed the information about actions in the discourse situation. to conclude we briefly discuss information about private attitudes and social attitudes (obligations). in ptt, dialogue acts are performed to achieve intentions or to satisfy certain obligations. both the fact that one or more agents have a certain (possibly collective) intention, and that they are under certain obligations, may become part of the private, semi-private and public parts of the discourse situation, e.g., as a result of planning or of intention recognition (matheson, poesio, and traum, 2000). in previous work, only a partial formalization of obligations and intentions was given. in this earlier work, both obligations and intentions were viewed as relations between agents and action types: for example, the fact that agent a has the intention to perform a particular core speech act is captured by the presence in the discourse situation of intention (3.5.1). (3.5.1) i: intend(a,λce.ce: assert(a,b,k1)) we adopt here a more standard view of intentions and obligations as predicates indexed by the agent(s) holding their intention. we will also simplify matters concerning the dynamics of conversational events by assuming that intentions and obligations have as their contents propositions describing the state of affairs to be achieved –i.e., drss—rather than action types, as in (3.5.2), showing the intention by agent a to perform an assertion with content k1. (3.5.2) i: int a([ce|ce: assert(a,b,k1)]) a partial formalization of obligations was provided in (matheson, poesio and traum, 2000). as far as intentions are concerned, previous work on ptt only discussed the assumption, inherited from grosz and sidner (1986), that intentions may be related to each other in two ways: by a relation of dominance when satisfying a certain intention is part of the satisfaction of another, more complex, intention (more formally: intention i dominates intention i' if achieving i' is part of achieving i); and by a relation called satisfaction-precedes when satisfying an intention is a prerequisite for satisfying the second (intention i satisfaction-precedes intention i' if achieving i is a necessary prerequisite of achieving i’). as each intention is (directly) dominated by only one other intention, poesio and traum (1997) formalized dominance with a (partial) function dom mapping intention i2 to the intention to which it is subordinated, if any. more controversially, satisfaction-precedence was also formalized as a partial function sp(i2) = i1 mapping intention i2 to the intention i1 that must be achieved for i2 to be achievable. completions, coordination and alignment in dialogue 25 poesio and traum also used dom and sp to provide an account of discourse entity accessibility in ptt, but otherwise proposed no axioms for intentions. one of the goals of this paper is to use completions and continuations as a source of additional evidence on the properties of intentions, and how they affect interpretation; we discuss a few possibilities and our assumptions in sections 4 and 5. our views on the relation between intentionality and accessibility have changed in the meantime, but we will assume the definitions from poesio and traum (1997) here; a paper discussing these new ideas is in preparation (poesio and rieser, in preparation). 4 completions: an account based on shared plans in this paper we will propose two analyses of the example dialogue. we begin by presenting in this section and the next a mainstream, ‘intentional’ analysis based on hypotheses about the role of intentions and cooperation in communication developed in artificial intelligence, linguistics, philosophy, and psychology over the past thirty years. in section 6 we will then discuss a second analysis, based on the recent proposals by garrod and pickering on the basis of recent psychological results about interpretation, production and alignment in dialogue (pickering and garrod, 2004). 4.1 coordination, shared intentions, partial shared plans: a look at existing paradigms 4.1.1 assumptions: intention, cooperation, coordination, discourse plans most work on dialogue in artificial intelligence, philosophy and psychology in the last thirty or so years has been based on three main assumptions. the first assumption is that communication involves a great deal of intention recognition: the production of utterances is motivated by (implicit and explicit) intentions, and in order to communicate felicitously agents must recognize other agents’ intentions even when they do not intend to help these agents to achieve them. this hypothesis, generally associated with grice (grice 1969, grice 1975, later revised in grice 1991), has been the foundation of most theories of dialogue in artificial intelligence and computational linguistics, such as the work of allen and perrault (1980), cohen and levesque (1990a, 1990b), grosz and colleagues (grosz and sidner, 1986; grosz and kraus, 1996) sadek (sadek et al, 1994) asher and lascarides (2003) and stone (2004) among others; in philosophy—e.g., in work by bratman (1992) or tuomela (2000); and in psychology (clark, 1992, 1996; bara & tirassa, 2000) –and was thoroughly examined in the seminal book intentions in communication (cohen, morgan and pollack, 1990). 10 the theories of intentions developed in ai are typically also concerned with how such intentions can be achieved: hence many such theories, and particularly those of intention recognition, are formulated in terms of plans to achieve a particular goal. the intentional account we propose will be formulated in this way, as well. this hypothesis that intention recognition is central to communication is usually supplemented by two further hypotheses. much work on dialogue in artificial intelligence relies on the further assumption that at least in some contexts, communicating agents do not simply recognize other agents’ intentions; they are also cooperative, in the sense that they attempt to help other agents’ achieve their intentions even when they are not explicitly expressed. e.g., a genuinely helpful clerk at a ticket counter in a train station will not simply answer the question “from which platform does the train to montreal leave” by indicating the platform, if he / she knows that the train has been cancelled. this view is central both to clark’s theory of dialogue discussed in section 2 and to, e.g., allen’s and sadek’s theories of intention recognition in 10 in recent years, the role of intentions in communication has become of interest to developmental psychologists (homer & tamis-le monda, 2005) and neural scientists (bara & tirassa, 2000), as well. poesio and rieser 26 dialogue systems (allen and perrault, 1980; sadek, 1992). completions and continuations are viewed by, e.g., clark as some of the best evidence for cooperative behavior in dialogue (clark, 1996, p. 238). as discussed in section 2, the theory of dialogue developed by clark and summarized in (clark, 1992, 1996) is based in addition on a third hypothesis: that conversation is a form of joint activity just like playing football or playing in an orchestra, i.e., driven by joint intentions and in which it is necessary for agents to coordinate. clark uses “joint activities” as a foundational stratum on which coordination on content or process, on setting up common ground, on signalling, establishing (complex) joint projects, on communication using parallel tracks etc. is then based. clark defines joint project quite simply as “… a joint action, projected by one of its participants and taken up by the others” (clark (1996), p. 191). the notion of joint project or shared plan has been further developed in the artificial intelligence, especially by grosz and sidner (1990) and by grosz and kraus (1996). taken together, these three assumptions lead to the view on which our first analysis is based: that what happens in dialogue, particularly task-oriented dialogue, can be explained in terms of the joint intentions of dialogue participants and the shared plans developed to achieve them. in the case of the construction dialogues in the btpc, this view can be summarized as follows. inst and cnst have both shared and private domain plans. inst and cnst’s private domain plans overlap to a degree with the shared plan; the difference between the shared plan and the private plans may lead to discrepancies and negotiations (e.g., as when inst adds more details to cnst’s proposal in 1.3). simplifying things drastically, we will assume here that inst’s private domain plan is a fully specified plan for building the toy airplane (either instructions or a model). the shared domain plan is a partial plan to build a toy airplane, which at the beginning of a dialogue is virtually empty, but gets progressively refined through the construction dialogue. (more on this below.) cnst’s private plan is a refinement of the shared domain plan, likely to include at least local further specifications based on expectations. crucially, dialogue involves shared plans both at the domain level (how to build a toy airplane) and at the discourse level (how to convey a particular intention or plan) (litman and allen, 1990). for example, inst and cnst also share a discourse plan: that the conversation will consist of a series of instructions by inst to cnst aiming at building a toy airplane according to inst’s model. this view is clearly formulated, e.g., in the following quote from (grosz and sidner, 1990, p. 418): discourses may exhibit two types of collaborative behaviour: collaboration in the domain of discourse [...] and collaboration with respect to the discourse itself. although we cannot yet define [...] “collaboration with respect to a discourse”, it includes not only surface collaboration (such as coordinating turns in a dialogue) or use of appropriate referring expressions ([...]) but also collaborations related to the discourse purpose. for example, the participants collaborate to ensure that the utterances of the discourse itself provide sufficient information to make possible the satisfaction of the discourse purpose. in terms of shared plans, cnst’s completion in 1.2 can be interpreted as indicating that cnst has recognized inst’s intentions both at the domain level and at the discourse level, and one possible explanation (although not the only one, as we will see below) for her decision to complete is that she is being cooperative. our first, `intentional’ account of completions is based on this shared plans hypothesis. we follow here what is perhaps the best-known development of this idea, due to bratman in his work on shared cooperative activity (sca) (bratman, 1992) and further refined by tuomela (2000). these two theories are briefly reviewed next. completions, coordination and alignment in dialogue 27 4.1.2 bratman arguably the most influential modern account of intentions (private and shared) is that of bratman (1992). bratman’s theory aims at providing a formalization of the notion of “shared cooperative activity” (sca, p. 327) that underlies actions characterised as “being done together” like singing a duet, painting a house or starting an attack in a basketball game—i.e., joint actions in clark’s sense, i.e., most actions performed in dialogue, including completions. bratman identifies three main characteristics of scas: mutual responsiveness, commitment to the joint activity and commitment to mutual support. the conditions that according to bratman must hold in order for us to be said to perform a shared cooperative activity (sca) j are listed in (4.1.1). first of all, we need individual intentions to j together: (1)(a)(i) and (1)(b)(i). secondly, it must be the case that you and i develop these intentions because of the other’s intentions ((1)(a)(ii), (1)(b)(ii)). (1)(c) is intended to exclude forced cooperation, e.g. the “mafia sense” (1992, p. 333) of doing something together. fourth, the intentions must hold for the matching subplans involved and they must be stable. (an intention is minimally cooperatively stable if “there are cooperatively relevant circumstances in which the agent would retain that intention” (p. 338). this notion is similar to cohen and levesque’s (1990) notion of ‘commitment’.) finally, we must mutually know (1). mutual knowledge is used in the fixed point sense (fagin et al., 1995, p. 402; tuomela 2000, p. 78). (4.1.1) our j-ing is a sca only if (1) (a) (i) i intend that we j. (ii) i intend that we j in accordance with and because of meshing subplans of (1)(a)(i) and (1)(b)(i). (b) (i) you intend that we j. (ii) you intend that we j in accordance with and because of meshing subplans of (1)(a)(i) and (1)(b)(i). (c) the intentions in (1)(a) and in (1)(b) are not coerced by the other participant. (d) the intentions in (1)(a) and (1)(b) are minimally cooperatively stable. (2) it is common knowledge between us that (1). this first version of bratman’s characterization of an sca already captures some key elements of the `dialogue as a joint action’ view discussed in sections 2 and 4.1.1: commitment to joint activity, meshing subplans, interdependent intentions, mutual support, and connecting attitudes. parts (ii) of (4.1.1)(1)(a) and (4.1.1)(1)(b) mean that it is not sufficient for an action to be independently intended to be a joint action by me and you for it to be a sca; these intentions have to be causally connected. parts (ii) also make explicit reference to the shared plans required to achieve a joint intention, via the notion of meshing subplans, described by bratman as follows: “[...] our individual sub-plans concerning j-ing mesh just in case there is some way we could j that would not violate either of our subplans” (p. 332). (as we will see, inst and cnst’s private plans are meshing in this sense.) according to bratman, however, one ingredient is still missing from (4.1.1): mutual responsiveness. mutual responsiveness amounts to the following: “in an sca each participating agent attempts to be responsive to the intentions and actions of the other, knowing that the other is attempting to be similarly responsive” (p. 328). (as we will see below, mutual responsiveness, or more precisely, its development by tuomela into the hypothesis that participants in a sca may perform ‘unrequited contributory actions,’ provides us with one explanation for the performance of completions.) the full definition of sca is therefore as in (4.1.2). poesio and rieser 28 (4.1.2) for cooperatively neutral j, our j-ing is a sca if and only if (a) we j (b) we have the attitudes specified in (4.1.1) (1) and (2), and (c) (b) leads to (a) by way of mutual responsiveness (in the pursuit of our j-ing) of intention and in action. given bratman’s notion of sca, an intentional account of completions in the example dialogue would go as follows. the relevant j-ing at the stage of example dialogue (1.1) just preceding the completion in 1.2 is that inst and cnst we-intend that cnst fix the ‘wing’ to the ‘fuselage’ according to inst’s directives. the intended and mutually known meshing subplans are the domain plan for building the toy plane in front of inst, and the discourse plan that inst give the directives and cnst carry out the actions. (both plans are discussed in more detail in section 4.2.) in order then for the joining of wing and fuselage to count as a sca, the following must hold: (4.1.3) inst and cnst joining wing and fuselage is a sca only if: (1)(a) (i) inst intend that inst and cnst join wing and fuselage. (ii) inst intend that inst and cnst join wing and fuselage in accordance with and because of meshing subplans of (1)(a)(i) and (1)(b)(i). (b) (i) cnst intend that inst and cnst join wing and fuselage . (ii) cnst intend that inst and cnst join wing and fuselage in accordance with and because of meshing subplans of (1)(a)(i) and (1)(b)(i). (c) the intentions in (1)(a) and in (1)(b) are not coerced by the other participant. (d) the intentions in (1)(a) and (1)(b) are minimally cooperatively stable. (2) it is common knowledge between inst and cnst that (1). the necessary and sufficient conditions for the action to be an sca are specified by adding the mutual responsiveness requirement from (4.1.3), obtaining (4.1.4): (4.1.4) for cooperatively neutral joining of wing and fuselage, inst and cnst joining wing and fuselage is a sca if and only if: (a) inst and cnst join wing and fuselage (b) inst and cnst have the attitudes specified in (4.1.3) (1) and (2), and (c) (b) leads to (a) by way of mutual responsiveness (in the pursuit of inst and cnst’s joining of wing and fuselage) of intention and in action. we will discuss the meshing subplans at stake in some detail in section 4.2. 4.1.3 tuomela bratman’s (1992) concept of shared cooperative activity is not formalized in terms of a logic, but several such formalizations have appeared, most famously by cohen and levesque (1990a, 1990b). our own use of we-intention (intinst&cnst) in the rest of the paper will be based on the logical reconstruction of bratman’s theory by tuomela (2000). tuomela also proposes several revisions of bratman’s theory, that we briefly review here. one problem with bratman’s account of plan-based joint action identified by tuomela is that the meshing requirement should not be part of the definition of we-intention. rather, we should take it as an entailed conceptual presupposition along the lines of ‘if we have a joint intention to see to it that p, then as a rule we also have meshing subplans concerning p’. in addition, according to tuomela, condition (4.1.1)(1)(c) is too strong: coercion should be admitted as long as a subject’s intentional agency is not completely by-passed. tuomela proposes, therefore, a stripped-down characterization of the we-intention going into the construction of the joint of wing and fuselage at the point in which the first completion takes place that can be characterized in first approximation as follows. completions, coordination and alignment in dialogue 29 (4.1.5) inst and cnst we-intend that cnst join wing and fuselage is equivalent to: it is inst’s and cnst’s mutual knowledge11 that inst intends that cnst join wing and fuselage because cnst intends that cnst join wing and fuselage; and cnst intends that cnst join wing and fuselage because inst intends that cnst join wing and fuselage the two clauses in the scope of the mutual knowledge operator in (4.1.5) cover much of clauses (1a) and (1b) in (4.1.3), the ‘because’ connective providing the ‘meshing’ required by bratman, but without explicit mention of plans and of coercion. in tuomela’s notation (4.1.5) is expressed as follows, where ‘because’ gets translated by the reason relation /r, which is factual in the sense that x/ry → x ∧ y: (4.1.6) int inst&cnst (join (cnst, w, f)) ↔ mk ((int inst (join(cnst, w, f)) /r int cnst(join(cnst, w, f))) ∧ (int cnst(join (cnst, w, f)) /r int inst(join(cnst, w, f)))) according to (4.1.6) the we-intention intinst&cnst that cnst join w and f amounts to mutual knowledge that each agent’s intention that cnst join w and f is caused by the other agent’s intention that this be done. we said above that agents’ disposition to helping, of which at least some forms of cooperative completions are an illustration, is meant to be covered in bratman’s proposal by the notion of mutual responsiveness which, however, is not fully developed, as pointed out by bratman himself. tuomela (2000, p. 107) provides a more explicit explanation of the fact that inst gets extra help by cnst in formulating a directive for which there can’t be a sca yet. to do this, tuomela developed an alternative to bratman’s definition of intinst&cnst. this new definition cannot be fully discussed here. we will simply highlight the following condition from the definition (2000, p. 95): “[...] the participants are also assumed to be disposed willingly to perform unrequired contributory actions [our italics], thus being disposed to incur extra costs (this being rational as long as the costs of performing them are less than the gross gains accruing from their performance)”. tuomela’s notion of unrequired contributory action underlies the ‘intentional’ account of completions given here. briefly, the idea is that if cnst and inst have we-intentions and the associated shared plans, cnst can infer that a screw is needed at the point of the dialogue we are analyzing. she is therefore able to produce an unrequired contributory action by producing a completion, if circumstances demand it. (we will return on this shortly.) tuomela’s formalization of helping is of course related to the traditional notion of cooperativeness, as formalized, e.g., in cohen and levesque’s ‘cooperative’ axiom (cohen and levesque, 1990b) or in sadek’s theory (sadek, 1992). the distinctive aspect of tuomela’s notion of help is that it is embedded into joint intentions and actions. tuomela argues that participants acting on the basis of a cooperative attitude must be disposed to “strong” helping (p. 249).12 helping in the “full” sense means 11 in many contexts mutual knowledge will be too strong. we will however stick to mutual knowledge for the purposes of this paper. 12 tuomela distinguishes between `helping in the weak sense’ and `helping in the strong sense’. help in the weak sense means that agent a is responsive to agent b’s performance of her part of the sca, in order to successfully complete her own part; it is therefore a form of alignment. helping in the strong sense is supporting the other agent by performing part of their performance of an sca in an essential way. 1.2 and 2.2 in our example are examples of strong helping, whereas cnst’s acceptances are examples of helping in the weak sense. poesio and rieser 30 ‘helping in all circumstances in which help is contributive to the others’ part-performances’. it is exactly reliance on part-performances (making up a sca eventually) which distinguishes tuomela’s notion of help from cohen and levesque’s (1990b) and sadek’s (1992), which capture taking over of some agent’s intention. in the rest of the paper, we will use tuomela’s intinst&cnst predicate to express we-intentions, adopting the formalization in (4.1.6). 4.2 an intentional analysis of the example dialogue: first pass in order to maintain the complexity of the presentation manageable, we divided the presentation of our intentional analysis of the example dialogue between this section and the next. in this section we concentrate on we-intentions and shared plans, showing how the bratman / tuomela framework just introduced can explain how completions are produced, without however discussing how completions can be incrementally produced and interpreted, and how the common ground gets updated as a result. in fact, we abstract away from the details of the drt notation and of ptt and adopt a vanilla logical form. in the next section we present a more complete account in which these issues are considered as well, using the tools from ptt. as said above, our intentional analysis is based on the assumption that at the beginning of the conversation instructor and constructor have a we-intention –i.e., an sca in the sense of bratman and tuomela—that constructor assemble a toy airplane identical to the one that the instructor has. (as discussed in section 3, the argument of the we-intention is a drs, i.e., a proposition, but we adopt a simplified representation in this subsection.) (4.2.1) int inst&cnst ( toy-airplane(x) ∧ assemble(cnst, x)) the instructor (henceforth, inst) has a complete13 plan for assembling the toy airplane in figure 2.1 (a) shown in figure 4.2.1. the dialogue is driven by the goal of making this a shared plan by discussing it with the constructor (henceforth: cnst), so that cnst can then execute the relevant actions. (for the purposes of this paper, we’ll assume that the case in which inst receives no instructions, just a completed model of the airplane, can be subsumed under this case as well.) 13 arguably, a more plausible view is that inst doesn’t start with a complete plan for the construction of the model— instead, he/she develops her private plan incrementally as well. we do not think however that this more plausible, but also greatly more complex, view would have any implications for our account of completions. assemble toy airplane assemble fuselage assemble wing join wing and fuselage figure 4.2.1. inst’s private plan before the conversation. grasp-5h-bar grasp 3h bar grasp(screw1) join(5h-bar,3h-bar, screw1) ………………… …. grasp(orange-slit-screw align wing and fuselage insert-screwthrough-hole completions, coordination and alignment in dialogue 31 according to the intentional view, inst and cnst develop a shared plan to achieve this weintention of assembling the toy airplane. this shared plan is initially highly underspecified and simply specifies that inst and cnst are going to assemble a toy airplane; but as the conversation progresses, the shared plan gets progressively more refined as each subplan becomes an sca as well. (note that the agreed upon actions in the shared plan are immediately executed, mixing planning and execution, unlike in the trains conversations (gross et al, 1993; heeman and allen, 1995), for example.) at the point at which inst begins utterance 1.1 in (1.1), “so jetzt nimmst du eine …,” the goal of inst is to get cnst to produce the baufix plane in figure 2.1 (a), and the state of cnst’s assembly is as shown in figure 2.1 (b); we repeat figure 2.1 here for convenience as figure 4.2.2. cnst has built the tail and the rear part of the fuselage. she has already picked up a 7-hole bar which is to become the `wing’ (henceforth, w), has laid it across the `fuselage’ (henceforth, f), and has aligned the wing’s and the fuselage’s holes, which will make it possible for a fixing mechanism such as a screw to join wing and fuselage. each of these steps has been acknowledged by inst and cnst. (a full listing of the dialogue prior to the fragment being analyzed is in appendix a.) (a) the baufix model inst intends cnst to assemble. (b) constructor's state of assembly at the beginning of (1.1) figure 4.2.2. instructor's and constructor's respective situations at the beginning of (1.1). given this state of assembly, the partial shared plan at this point is presumably similar to the one shown in sketchy form in figure 4.2.3. the parts of inst’s private plan devoted to the assembly of fuselage and wing are now shared (in fact, they have already been executed). now inst and cnst have a we-intention that cnst join wing and fuselage; this we-intention is indicated in italics as it is currently at the top of the agenda. (more details on the shared plan in a moment.) poesio and rieser 32 at this point in the conversation cnst has a private plan as well, a refinement of the partial shared plan in figure 4.2.3. this private plan probably already contains the information that she will need to get a screw in order to join wing and fuselage, as sketchily shown in figure 4.2.4. however, cnst has lots of spare parts left, above all nine screws which she can in principle use for the join as shown in figure 4.2.5; hence, she cannot refine her own private plan any further. figure 4.2.5. screws available to cnst at the point when 1.1 is uttered. the intended screw is the orange screw with a slit. assemble toy airplane assemble fuselage assemble wing join wing and fuselage figure 4.2.3. the shared plan at the point the completion takes place. grasp(5h-bar) grasp( 3h bar) grasp(screw1) join(5h-bar,3h-bar, screw1) ………………… …. align wing and fuselage … assemble toy airplane assemble fuselage assemble wing join wing and fuselage figure 4.2.4. cnst’s private plan at the point the completion takes place. grasp(5h-bar) grasp( 3h bar) grasp(screw1) join(5h-bar,3h-bar,screw1) ………………… …. grasp(screw) align wing and fuselage ….. completions, coordination and alignment in dialogue 33 as said above, in the tuomela-derived logical notation just introduced, the currently active we-intention to join w and f can be expressed as follows: (4.2.2 ) int inst&cnst (join (cnst,w,f)) as shown in (4.1.6), repeated below for convenience, inst and cnst having a we-intention to join wing and fuselage, in tuomela’s framework, is equivalent to them mutually knowing that each of them has an intention to do this action because of the other agent’s intention to do it: (4.1.6) int inst&cnst (join (cnst, w,f)) ↔ mk ((int inst (join(cnst, w,f)) /r int cnst(join (cnst, w,f))) ∧ (int cnst(join (cnst, w,f)) /r int inst(join(cnst, w,f)))) in order to join two objects cnst needs, first of all, to find a screw and a nut. she must then stick the screw through the holes and screw it into the nut in order to produce a stable join of three bars (that we will call here the wing&fuselage-join). this partial domain plan, that we assume to be shared, can be summarized as follows:14 (4.2.3) join(agent, bar1,bar2) step 1: align(agent, bar1, bar2, hole1,hole2), step 2: grasp(agent, screw), step 3: grasp(agent, nut), step 4: put-through(agent, screw, hole1,hole2), step 5: screw-into(agent, screw, nut), as in (asher and lascarides, 2003), we assume that the relation between intentions and actions (here formalized in terms of plans) is non-monotonic: i.e., that a plan such as (4.2.3) is a default way of executing a particular action. (asher and lascarides formalization of intentions and actions does not involve the notion of plans.) this can be formalized as follows, where we use a generic normally involves operator to state that the conclusions are non-monotonic:15 (4.2.4) join(agent, bar1, bar2) normally involves b & c & d & e & f where (b) align(agent, bar1,bar2, hole1,hole2) (c) grasp(agent,screw), (d) grasp(agent, nut) (e) put-through(agent, screw, hole1,hole2), (f) screw-into(agent, screw, nut) in the conversations of the bielefeld corpus, things are a bit more complex, in two respects: (i) actions are performed by cnst under inst’s instruction; and (ii) actions are immediately executed. as it is not our goal here to provide an account either of the interleaving of planning 14 we are using a very simplified view of plans here, ignoring the distinction between preconditions –conditions that have to hold in order for the plan to be executable—and steps of the plan proper. normally, having the nut and the screw would be considered to be a precondition. (this difference is not unlike that between presupposition and assertion.) 15 (4.2.4) is meant to be a neutral notation which could be reconstructed in different nonmonotonic formalisms. in sdrt, the relation between intentions and plans would be expressed by a conditional, and normally involves would be asher and morreau’s >asher&morreau operator. in ptt defaults are inference rules, formulated using brewka’s (1990) prioritized version of reiter’s (1980) default logic). poesio and rieser 34 and execution, or of the interleaving of discourse plans and domain plans (see e.g., (litman and allen, 1990)), we simplify matters by assuming that in this domain, shared plans / meshing subplans combine discourse actions, domain actions, and execution. for the particular example under discussion, the resulting plan is assumed to be as shown in (4.2.5). according to this plan for cnst to join wing and fuselage, in order to jointly perform the action currently we-intended, inst has to produce the directives triggering these actions; each of these results in actions by cnst; and at the end, cnst must communicate that she completed the actions demanded. (4.2.5) join(cnst, w, f) normally involves b & c & d & e & f & g, where (b) 1. directive(inst, cnst, align(cnst, w, f, hole1, hole2)), 2. align(cnst, wing, fuselage, hole1, hole2), 3. assert(cnst, inst, aligned(cnst, w, f, hole1, hole2)) (c) 1. directive(inst, cnst, grasp(cnst, screw)), 2. grasp(cnst, screw), 3. assert(cnst, inst, grasped(cnst, screw )) (d) 1. directive(inst, cnst, grasp(cnst, nut)), 2. grasp(cnst, nut), 3. assert(cnst, inst, grasped(cnst, nut)) (e) 1. directive(inst, cnst, put-through (cnst, screw, hole1,hole2)), 2. put-through(cnst, screw, hole1,hole2), 3. tell(cnst, inst, put-through (cnst, screw, hole1,hole2)) (f) 1. directive(inst, cnst, screw-into(cnst, screw, nut)), 2. screw-into(cnst, screw, nut), 3. assert(cnst, inst, screwed-into(cnst, screw, nut)) (g) assert(cnst, inst, joined(cnst, w,f). we will assume throughout this section and the next that intention recognition for cnst amounts to recognizing the directives and performing the required actions.16 bratman’s and tuomela’s theories of intention make the stronger assumption that a weintention to bring-about a w&f join distributes over the conjuncts: i.e., having a we-intention to bring about an action entails we-intentions for all parts of the plan. with this assumption, we would then also derive from (4.2.1) and (4.2.5) (c) the following: (h) int inst&cnst (directive(inst, cnst, grasp(cnst, screw))) (i) int inst&cnst (grasp(cnst, screw)) (j) int inst&cnst ( assert (cnst, inst, grasped(cnst, screw))). notice that even with this assumption, inst and cnst would still have different private plans and different private intentions: inst, seeing the fully built up airplane model on his side, subscribes to (4.2.6) int inst(directive(inst, cnst, grasp(cnst, orange-slit-screw)))17 whereas for cnst we have (4.2.7) int cnst(grasp(cnst, screw)). (the difference between (4.2.6) and (4.2.7) is due to the fact that grasp(cnst, screw) can be satisfied by cnst in various ways due to the nine screws she has.) the point is that if we make the further assumption that int inst&cnst is distributive, we get (4.2.7) irrespective of any inferences 16 intention recognition could also be formalized in such a way as to bypass the directive recognition stage, and assuming simply that cnst recognizes inst’s intention of cnst performing a particular action. an account along these lines would not, however, differ from the one we are adopting for our purposes. 17 we are not being very precise concerning notation here. quantificational formulae are sometimes represented as constants for convenience’s sake. completions, coordination and alignment in dialogue 35 due to the verbal exchange: in other words, with this assumption, cnst already wants to grasp a screw before inst starts his directive. (cnst’s intention can be satisfied because there are screws on his side of the screen.) the only difference between the two plans would then be that cnst doesn’t know which particular elements have to be used. however, we will remain non-committal on whether cnst derives (4.2.7) from distributivity of we-intentions or from the directive and cooperativity. we now consider how an intentional account explains what may have caused cnst to produce a completion in 1.2. given that int inst(directive(inst, cnst, grasp(cnst, screw))) is the next step in the interleaved discourse / domain plan in (4.2.5) to achieve sca (4.2.2), we may ceteris paribus assume that by uttering well, now you grasp inst has started the production of the next directive in this shared plan. what about cnst’s contribution? according to the mutually intended step (d) of (4.2.5), cnst should by default wait until inst has fully produced his directive, and only then she should do the grasping. but she intervenes. so we must explain: a. the cause for the completion, including the information used for the intervention; b. how the old shared plan changes, and what the new shared plan looks like; c. the relation between old and new shared plan. let us start with the cause for the completion and the information used for the intervention. the step in the plan which is we-intended at this point in the conversation is: join(inst and cnst,w, f). inst begins to perform the next step, (c): directive(inst, cnst, grasp(cnst, orange-slit-screw)). however, cnst ‘jumps in’ before the directive is completed, and possibly without being prompted by some sort of request for help (as we saw in section 2, there is some evidence that in this particular example a request for completion may have been made using prosody, but the matter is not entirely clear, and these signals are not encountered with all completions). clearly, cnst has been trying to compare her expectations (based on the plan she assumes to be shared) with the actions she observes (as proposed, e.g., in the mosaic model (wolpert et al, 2003)) and she has been doing this incrementally. as a result, even before inst’s utterance of the directive is completed, cnst has already related the incomplete utterance she is observing to the next step in the shared plan, directive(inst, cnst, grasp(cnst, screw)). we will discuss in more detail how this can happen in the next section, in which ptt is used to provide a detailed account of incremental semantic interpretation in this example. as we said above, in a theory in which we-intentions are assumed to distribute over the subactions of plans, recognizing this directive is predicted to be fairly straightforward, as intention (4.2.6) would already be near the top of cnst’s agenda. otherwise, some sort of search through the space of possible plans to achieve (4.2.2) must be assumed. by comparing the part of the directive that has already been produced by inst with the directive she expects, cnst can hypothesize that what’s missing so far is an utterance of the fact that the object to be grasped is a screw. (there is no more information about this in the shared plan, and as there are nine screws, there are nine ways to make the plan more specific, so there is no way for cnst to make the recognition more precise, except by guessing. ) we see at least four ways of explaining why cnst follows up this recognition with a decision to utter “a screw”: a. responding to a request: cnst may have interpreted 1.1 (more precisely, its prosodic lengthening) either as a request to perform a continue (du), or as a request to acknowledge the du. in either case, in ptt it is assumed that the result is an obligation: obl(cnst,cont(du)) or obl(cnst, ack(du)). poesio and rieser 36 notice how an account of this type presupposes a theory of obligations and their discharge such as the theory in (matheson, poesio and traum, 2000). b. voluntary coordination-level control: even without having interpreted inst’s utterance as a request, cnst may nevertheless intend to signal her understanding of the directive: i.e., to acknowledge inst’s directive. acknowledging a directive cannot be done simply by repeating the part already uttered: this might be interpreted simply as a partial acknowledgment of only the part of the directive already performed. hence, cnst acquires the (private) intention to perform the missing part of the directive (as above). however, the directive is still considered by cnst as performed by inst only. such an account presupposes a theory of grounding acts like the one discussed in section 3. c. cooperativeness: cnst intends to help inst to perform the directive: directive(inst,cnst,grasp(cnst,screw)). we formalize this as cnst making performing the directive into a shared plan:18 int cnst&inst (directive(inst&cnst,cnst,grasp(cnst,screw))). this leads to cnst performing an “unrequired contributory action”: performing the part of the directive that is missing. as a result, cnst assumes the private intention: int cnst (utter (“a screw”)) d. blurting out: cnst feels that inst is taking too long to perform a simple directive. as soon as cnst recognizes the directive that inst is trying to perform, the plan (in terms of utterance actions) inst is using to do so, and what is still missing for the completion of this plan, cnst acquires the intention to utter the missing part, without necessarily making the directive into a joint action, and without necessarily intending to acknowledge inst’s contribution. (although both intentions could be there.) all these are valid explanations of what may have happened, but we won’t have enough space to pursue all of them in detail. in a way, explanation d. is the most interesting, but also the most difficult to formalize. explanations a.-c. all involve reasoning about information states—i.e., the type of explanation that is most in contrast with the alignment model proposed by pickering and garrod—so we will concentrate on them. further, explanations a. and b. are examples of the type of interaction most closely studied in previous work on ptt, particularly in (matheson, poesio and traum, 2000), so we will primarily concentrate on providing a detailed account of the cooperativeness explanation, c., providing however the bare bones of an account according to a. and b.—and in particular, of the interaction between such account and the grounding process. after the completion, as a result of inst and cnst’s coordinated action we get a complete directive: well, now you take a screw. however, there is a difference between this directive and the version in inst’s private plan: int inst(directive(inst, cnst, grasp(cnst, orange-slit-screw))) inst may react to cnst’s contribution by either a. accepting the completion (either because the complete speech act reflects inst’s initial intention, or because inst considers the new directive a valid alternative); b. rejecting it, c. performing a literal resumption, d. paraphrasing it, or 18 this notion of cooperativeness is different from cohen and levesque’s in two respects. first of all, cnst is deriving a joint intention from a shared plan, instead of deriving a private intention from inst’s plan. secondly, as we will see in section 5, cnst is adopting an intention to perform part of a plan whose execution would satisfy an obligation of inst, instead of performing an action from scratch. completions, coordination and alignment in dialogue 37 e. refashioning it, by modifying some aspects of it.. in this particular case, inst appears to be performing a refashioning, as discussed in sections 2 and 3 (clark and wilkes-gibbs 1986). inst knows that orange-slit-screw ⊂ screw; this can be captured by the following meaning postulate: orange-slit-screw(x) := screw(x) ∧ orange(x) ∧ ∃y(slit(y) ∧ with(y,x)). so he accepts some aspects of the contribution, but adds more material to ensure the right screw is identified. inst’s intention is of refashioning the directive currently under elaboration into the more specific directive: int inst(directive(inst, cnst, grasp(cnst, screw(x) ∧ orange(x) ∧ ∃y(slit(y) ∧ with(y,x))))). again, inst compares the directive already produced with his intended directive. a screw has already been produced; the difference between cnst’s ∃x(screw(x)) and inst’s ∃x(screw(x) ∧ orange(x) ∧ ∃y(slit(y) ∧ with (y,x)), is ∃x(orange(x) ∧ ∃y(slit(y) ∧ with (y,x)). we get the more economical int inst(directive(inst, cnst, grasp(cnst, orange(x) ∧ ∃y(slit(y) ∧ with(y,x))))) yielding an orange one with a slit. notice how performing the type of reasoning discussed here requires the ability to talk about incomplete speech acts, i.e., a theory such as ptt. we next show how ptt allows us to make the account just presented more precise. 5 an intentional analysis of the example dialogue: second pass in this section we analyze the example dialogue, utterance unit by utterance unit,19 supplementing the intentional analysis just provided with a discussion of how the utterance units are incrementally interpreted and produced according to the ptt view of interpretation, generation and the grounding process, on the basis of the grammar given in the cdrt fragment in appendix b. the analysis of these aspects we will provide is very detailed. on the other hand, ptt, unlike sdrt, does not provide (yet) a complete set of default rules specifying the inferential connection between utterances and their interpretation, so we will only discuss these inferences in an informal way. (see (poesio, 1994; poesio, 1996; poesio, to appear) for a discussion of the use of prioritized default logic in ptt and a formalization of other types of inferences within it.) 5.1 discourse situation updates up to the point just before 1.1 we discussed in section 4 how according to the intentional analysis, a conversation like the one from which fragment (1.1) is extracted is driven by inst and cnst’s we-intention to build a toy plane according to the model given to inst, which in the simplified view adopted here becomes the private plan in figure 4.2.1. by way of his instructions, inst turns the private plan into a progressively more specified shared plan interleaving discourse and domain actions as in the example in (4.2.5). as a result of cnst’s performing the domain actions in this shared plan, inst and cnst reach the situation whose relevant aspects are depicted in figure 4.2.1 (state of assembly) and figure 4.2.4 (leftover screws). in ptt, performing a directive amounts to performing a contribution which, ultimately, amounts to a series of micro-updates as in (3.3.5), each of which has to be properly grounded. the discourse situation just before (1.1) is a record of both core speech acts and micro conversational events performed up to that point. 19 following (traum, 1994, traum & nakatani, 1999, poesio and traum, 1997, poesio and traum, 1998) we use the term utterance unit to refer to basic units of dialogue processing, corresponding roughly to prosodic units. poesio and rieser 38 5.2 production of 1.1: so, jetzt nimmst du … we saw in section 4.2 that at this point in the dialogue, inst and const we-intend to join two specific objects, that we will call here wing1 and fuselage1.20 in ptt notation, and using drss as intentional contents, the fact that we-intention (4.2.2) is part of the common ground is captured by including (5.2.1) among the conditions characterizing the discourse situation.21 (5.2.1) i1.1a:int inst&cnst ([e1| e1:join (cnst,wing1,fuselage1)]) we also saw that in the btpc conversations such goals are achieved by following plans along the lines of (4.2.5), whose next step (step (c)) is a directive by inst to cnst to grasp a screw. in ptt terms, performing such a directive amounts to planning (i.e. intending) a conversational event ce1.1 of type directive and with content proposition k1.1: the occurrence of an event e of cnst grasping orange-slit-screw. (the event should take place in the future, but we will omit tense here for brevity). (5.2.2) i1.1b:int inst([ce1.1,k1.1| k1.1 is [x,e| x is orange-slit-screw, e:grasp(cnst, x)], ce1.1:directive(inst,cnst,k1.1)]) a couple of remarks about (5.2.2). notice, first, that embedded drss are essential to formulate this intention. secondly, inst’s intention is about a specific screw (represented here as the constant orange-slit-screw), instead of the more general intention to grasp an x which is a screw (in (5.2.2’)) which can be expected to be in the shared plan. (that inst’s intention is about a specific screw is suggested, e.g., by the subsequent repair.) (5.2.2’) i1.1b’:int inst([ce1.1,k1.1 |k1.1 is [x,e|screw(x),e:grasp(cnst,x)], ce1.1:directive(inst, cnst, k1.1)]) (in passing, note also that if we assume that we-intentions distribute over the plan in (4.2.5), cnst could derive (5.2.2’) from the non-specific joint intention in the shared plan shown in (5.2.2’’).) (5.2.2’’) i1.1b'’:int inst&cnst ([ce1.1,k1.1| k1.1=[x,e|screw(x), e:grasp(cnst, x)] , ce1.1:directive(inst, cnst, k1.1)]) assuming that intentions to perform domain actions lead to the adoption of domain plans to achieve such intentions via von wright’s (1963) practical syllogism,22 (5.2.2) leads to inst’s plan of performing an utterance generating the directive in the sense discussed in section 3.1, example (3.1.5). to simplify matters, we assume here, as in (poesio and traum, 1997), that this plan amounts to intending to perform an utterance up1.123 whose conventional meaning (the value of sem(up1.1)) is the same as the content of the directive. this intention is dominated by the intention i1.1b of performing a directive, ensuring that this intention is about the same ce1.1 and k1.1 as intention i1.1b (poesio and traum, 1997). (5.2.3) i1.1c:int inst([up1.1 | utterance(up1.1), sem(up1.1) = k1.1, generate(up1.1,ce1.1)]) (5.2.3) is the starting point for the generation process. and at this point, the hypothesis that the participants’ record of the discourse situation includes information about the occurrence of micro conversational events – events of uttering sub-sentential constituents—does begin to do 20 we remind the reader that in cdrt, unlike standard drt, there are regular constants; wing1 and fuselage1 are such constants, and so is orange-slit-screw. 21 we will generally assume de dicto (wide scope) intentions. 22 intending ϕ and believing that performing action χ has as a consequence ϕ leads to performing χ: intend (ϕ) ∧ bel(χ→ϕ) → do χ (from wright, georg henryk von. 1963. "practical inference." philosophical review, 72:159-179. _______. 1972. "on the so-called practical inference." acta sociologica, 15:39-53.) 23 terminological note: we use discourse referents named upi.j to indicate phrasal utterances, where i is the number of the turn. we use discourse referents named ui.j to indicate lexical utterances. completions, coordination and alignment in dialogue 39 some work. specifically, the mce hypothesis (i) connects intentions such as (5.2.3) with their syntactic realization, and (ii) provides us with an account of the information that is available to cnst when she starts planning her completion. in ptt, performing an utterance with the conventional meaning represented by k1.1 is viewed as an action (of type utter ). we can then view the results of surface realization, the process of determining the surface structure to be used to realize an utterance up1.1 with content k1.1 (reiter and dale, 1995), as a plan just like those produced in other types of planning. (the assumption being of course that specialized planners, called surface realizers in natural language generation, are responsible for this task.) as shown in (3.3.5), the plan chosen by inst for producing conventional meaning k1.1 in this particular occasion involves choosing up1.1 to be an utterance of type s. this action is decomposed in the performance of an utterance u1.2 of the temporal adverbial “jetzt” and an utterance of a second phrasal utterance of type s, here called up1.2. performing up1.2 involves performing a (phonetically silent) utterance up1.3 of type np24 and an utterance up1.4 of type vp, which in turn involves three subactions: an utterance u1.3 of type v (“nimmst”), a subject utterance up1.5 of type np (“du”), and a complement utterance up1.6 of type np (“eine orangene schraube mit einem schlitz”). another advantage of the micro-conversational events hypothesis is that we do not need to assume that every utterance that gets produced is a constituent of a single syntactic tree—a big `sign’ in the hpsg sense. in the case of 1.1, for example, it is plausible to view at least the first “so” (indicated here as u1.0), and possibly even “jetzt”, as separately planned utterances intended to achieve quite distinct goals from the other utterances which generate the directive. in this case, “so” would seem to realize a turn-taking action, whose goal is to take the turn and keep it (traum and hinkelman, 1992; traum, 1994; poesio and traum, 1997).25 the hesitation that may or may not have occurred at the end of 1.1 would play a turn-taking oriented goal, as well, and possibly a role in the grounding process. to show what such an analysis would be, we assume here this interpretation for “so,” whereas we interpret the utterance of “jetzt” as an utterance of type advp with a sentence-adjunct interpretation. omitting many of the details already shown in (3.3.5), including lexical interpretations and many sub-utterances, and using the notation u:“word“:cat to stand for: u:utter(a,“word“), cat(u) inst’s plan to satisfy i1.1c in (5.2.3) becomes a plan of performing the actions shown in tree format in figure 5.2.1 and in linear format in (5.2.4).26 intention i1.1d is dominated by i1.1c. (5.2.4) i1.1d:int inst([u1.0,ce1.0,up1.1,u1.2,u1.3,up1.2,up1.3,up1.4,u1.4,up1.5,up1.6| u1.0: utter (inst, “so”), ce1.0: take-turn (inst), generate(u1.0,ce1.0), s(up1.1) , sem(up1.1)= k1.1, u1.2:“jetzt“:advp, s(up1.2), u1.2 ↑ up1.1, up1.2 ↑ up1.1, np(up1.3), vp(up1.4), up1.3 ↑ up1.2, up1.4 ↑ up1.2, u1.3:“nimmst“:v, u1.3 ↑ up1.4, np(up1.5), up1.5 ↑ up1.4, u1.4:“du“:pro, u1.4 ↑ up1.5, up1.6:“eine orangene schraube mit einem schlitz“:np, up1.6 ↑ up1.4]) 24 according to ltag syntax, directives have a phonetically empty np in subject position. 25 “so” could also be interpreted as a ready signal as in the maptask scheme (carletta et al, 1997). 26 we have assumed for simplicity that the entire utterance is planned at once. however, in a more plausible theory, the generation process would be incremental, as well. poesio and rieser 40 the overall mental state of inst after acquiring intentions i1.1b, i1.1c, and i1.1d is portrayed in ‘box’ format in figure 5.2.1, where constituency relations are also displayed in a graphical format. as said in section 3, in the ptt model of contributions—derived from clark and schaefer’s proposals as modified by traum—the intention to perform a core dialogue act leads to an intention to make a new contribution, i.e., to open a new du, according to the ce-to-duschema here repeated: [ce-to-du-schema ] int cp1([ce | ce:core-act(cp1,cp2,φ)])� int cp1([du | du is [ce | ce:core-act(cp1,cp2,φ)]]) because of this schema, (5.2.2), (5.2.3) and (5.2.4) lead to (5.2.5): (5.2.5) i1.1e:int inst([du1.1inst | du1.1inst is [ce1.1,k1.1,up1.1,u1.0,ce1.0,u1.2,u1.3,up1.2,up1.3,up1.4, u1.4, up1.5,up1.6| k1.1 is [x,e| x is orange-slit-screw, e:grasp(cnst, x)], ce1.1:directive(inst, cnst,k1.1), utterance(up1.1), sem(up1.1) is k1.1, generate(up1.1,ce1.1), ..... remaining conditions from (5.2.4) ... ]]) figure 5.2.1. inst’s mental state after sentence planning. …. i1.1b:int inst ([k1.1,ce1.1| k1.1 is [e, x| x is orange-slit-screw,e:grasp(cnst,x)] ce1.1:directive(inst, cnst, k1.1)]) i1.1c:int inst ([up1.1| utterance(up1.1), sem(up1.1) = k1.1, generate(up1.1,ce1.1)]) i1.1d:int inst u1.0 ce1.0 u1.1 u1.2 u1.3 up1.2 up1.3 up1.4 u1.4 up1.5 up1.6 u1.0:utter (inst,“so"), ce1.0:take-turn (inst), generate(u1.0,ce1.0), ”eine orangene schraube mit einem schlitz” up1.1:s u1.2:”jetzt”:advp up1.4:vp up1.3:np u1.3:”nimmst”:v up1.5:npi up1.6:np u1.4:”du”:pro εi up1.2:s … i1.1b i1.1c i1.1d completions, coordination and alignment in dialogue 41 intentions such as i1.1e in (5.2.5) lead to inst performing the planned action, i.e., beginning a new contribution. inst starts executing this intention, succeeding to perform u1.2, u1.3, and u1.4. it’s not clear to us whether the lengthening observed at this point is meant to indicate a problem – e.g., that inst has forgotten which screws are unused, and therefore does not know what description would be adequate (dale and reiter, 1995; stone, 2004)— a request for acknowledgment, or whether it is just accidental. 5.3 the completion: 1.2, “eine schraube” the process that leads cnst to produce the completion begins when she observes inst performing the first four micro-conversational actions of his planned contribution. as discussed in section 3, according to ptt observing these utterances leads cnst to create a number of new discourse units, each of which represents “material to be grounded”; one ‘micro-du’ per utterance event is assumed. for simplicity, however, we represent here cnst’s interpretation of the utterance events in 1.1 as a single new discourse unit, that we will call du1.1cnst (to emphasize that it need not be the same as the du from inst’s perspective). the update to the discourse situation resulting from the observation of du1.1cnst is shown in linear format in (5.3.1) and in graphical format in figure 5.3.1. this new contribution minimally contains a record of the occurrence of each utterance, together with the results of lexical access and (incremental) parsing (see (poesio, 1995; poesio and muskens, 1997; poesio, to appear) for details). assuming that there are no interpretation problems, du1.1cnst would be a partial representation of the du planned by inst, as shown in (5.2.5). the difference between the two is that in the contribution observed by cnst, np up1.6 hasn’t yet been observed, but it is already expected.27 (in addition, a new dialogue act ce1.1 is presumably hypothesized, but at this point there is no evidence that it has already been classified as a directive, although that is likely as well given the information from word order and the pending we-intention.) (5.3.1) [ du1.1cnst | du1.1 cnst is [u1.0,ce1.1,u1.2,up1.1,u1.3,up1.2,up1.3,up1.4,up1.5,up1.6,u1.4 | u1.0: “so“:take-turn , s(up1.1), u1.2:“jetzt“:advp, s(up1.2), u1.2 ↑ up1.1, up1.2 ↑ up1.1, u1.3:“nimmst“:v, np(up1.3), vp(up1.4), up1.3 ↑ up1.2, up1.4 ↑ up1.2, u1.3 ↑ up1.4, np(up1.5), up1.5 ↑ up1.4, np(up1.6), up1.6 ↑up1.4, u1.4:“du“:np, u1.4 ↑ up1.5 ]] (5.3.1) is the starting point for cnst’s semantic interpretation process. the extent to which semantic construction and speech act interpretation take place immediately is still an open question (our views are discussed in more detail in (poesio, to appear)). the interpretation in (5.3.1) simply encodes the results of lexical access and preliminary syntactic interpretation, both of which are known to at least begin very early (swinney 1979, simpson 1994, frazier 1987, tanenhaus et al 1995). there is also evidence that the observation of the occurrence of an event of uttering an anaphoric np such as a pronoun is sufficient to start the processes by which such expressions are interpreted, as discussed in (poesio, 2001; poesio, to appear). how much more interpretation takes place? as discussed in section 4, under the intentional account the fact that a completion takes place is an indication that cnst somehow manages to recognize inst’s intention to perform a directive. (we represented this intention there using the simplified form int inst(directive(inst, cnst, grasp(cnst, screw))). ) this recognition is what prompts cnst to utter “eine schraube.” a more explicit ptt representation of the result of the first of these 27 we assume here that inst and cnst have the same mental grammar. poesio and rieser 42 inferential processes is shown in (a), whereas the intention to utter a screw would be represented as in (b). (a) int inst([ce1.1, k1.1 | k1.1=[x,e|screw(x), e:grasp(cnst, x)] ce1.1:directive(inst,cnst,k1.1)]) (b) int cnst([u|u:utter (cnst,”eine schraube”)]) (note that it’s hard to tell from the dialogue whether cnst has already identified a particular screw, so we will not assume anything in this respect here.) we consider two hypotheses concerning the way cnst reaches conclusion (a). as for the way recognition of (a) leads to the adoption of intention (b), we already saw in section 4 that this can be explained in a number of ways, and a number of these alternative hypotheses seem equally plausible, so that choosing among them would amount to mind reading. we will nevertheless attempt to show that several such hypotheses could be expressed in terms of ptt. 5.3.1 identifying intention (a) existential closure one hypothesis concerning the process through which cnst reaches conclusion (a) is based on the operation of existential closure proposed by chater, pickering and milward (1995). according to chater et al, existential closure takes place every time a new input is perceived; its purpose is to produce propositions out of partial syntactic interpretations, up1.1:s u1.2:”jetzt”:advp up1.4:vp figure 5.3.1. cnst’s interpretation after perceiving 1.1. du1.1 is up1.2:s up1.3:np εi u1.3:”nimmst”:v up1.5:npi up1.6:np u1.4:”du”:pro u1.0 ce1.0 ce1.1 up1.1 u1.2 u1.3 up1.2 up1.3 up1.4 u1.4 up1.5 up1.6 … du1.1 u1.0:utter (inst,“so"), ce1.0:take-turn (inst), generate(u1.0,ce1.0), completions, coordination and alignment in dialogue 43 so that the new input can be immediately evaluated against the current situation. in ptt terms, the existential closure hypothesis amounts to hypothesizing that cnst attempts to derive a proposition as the conventional meaning of up1.1 by existentially closing the missing argument, even when the only information at her disposal is what is presented in (5.3.1). more explicitly, the hypothesis is that as a result of observing the contribution represented in (5.3.1) as du1.1 cnst, cnst hypothesizes a new drs k1.1 as the argument of an intention attributed to inst: k1.1 is the proposition that an event of the constructor taking an object x yet to be determined will have to take place.28 (in ptt, the result of all inferences performed by cnst on the basis of (5.3.1), such as existential closure, is viewed as material to be grounded separately, i.e., as a separate du; for simplicity here we will however assume that all such inferences are added to the same du, du1.1 cnst.) we remind the reader that we specify drs updates using the notation k+= k’, defined as follows: (see section 3.2): k += k’: let k be a proposition-valued discourse referent, and k’ be a proposition (drs) or a proposition-valued discourse referent. then k += k’ =def [ tmpk| tmpk is k];[ k| k is tmpk;k’] (where tmpk is a new proposition-valued discourse referent) using this notation, the result of existential closure is the update of cnst’s view of the discourse situation in (5.3.2). (5.3.2) du1.1 cnst += [k1.1| k1.1 is [x,e|e:grasp(cnst,x)]]. this update of the information state changes the value of du1.1cnst by adding to it the existence of a new proposition k1.1 whose content is the existence of an event of grasping an unspecified object x. the recognition by cnst that an action of grasping something is being discussed is the basis for her inferring that the grasping action in du1.1cnst is the first step in performing the action currently we-intended. we recall that under the intentional account, inst and cnst are operating under the shared plan for bringing about the required join of the three pieces in (4.2.5), involving both domain and discourse actions. the next step in the ‘existential closure hypothesis’ is that inferring (5.3.2) leads cnst to infer that inst is performing the second directive (c.1) in the shared plan in (4.2.5). (the acceptance by cnst of that directive will result in cnst’s adopting the intention of performing the action indicated by the directive, (c.2) in (4.2.5).) 29 this inference leads to cnst updating her view of the content k1.1 of the directive with the information [|screw(x)]: (5.3.2’) du1.1 cnst += (k1.1 += [|screw(x)]) (which, by virtue of the definition of +=, is equivalent to updating the information state twice as follows: [tmpk| tmpk is du1.1 cnst]; [du1.1 cnst| du1.1 cnst is tmpk;[tmpk’| tmpk’ is k1.1][k1.1| k1.1 is tmpk’; [|screw(x)]]], i.e., update the value of du1.1 cnst within the information state by adding the condition that x is a screw to the value of k1.1 within du1.1 cnst). then cnst updates du1.1 cnst with (a), i.e., she attributes to inst a directive whose content is the updated k1.1 . this is shown in (5.3.3). notice that this intention is the intention discussed above in (5.2.2’). (as above, we assume cnst ascribes 28 as pointed out to us by jonathan ginzburg (p.c.), existential closure as specified by chater et al, and as implemented here, is a significant simplification in that in a dialogue not all core dialogue acts have a proposition as a content: the content of interrogatives are questions, and the content of directives presumably would be some form of action type. (for an ontology of the type of objects required in a theory of dialogue, see (ginzburg, 2009).) 29 in the ‘standard’ version of ptt, as presented, e.g., in (matheson, poesio, and traum, 2000), adopting this intention would be a way of addressing the obligation raised by the directive. as already said above, we are not concerned with obligations here. poesio and rieser 44 to inst a de dicto intention. also, as discussed in section 4, we leave it undetermined here whether this private intention attributed to inst is directly derived from we-intention (h) derived from (4.2.5) via distributivity of we-intention.) (5.3.3) du1.1 cnst += [i1.1b’|i1.1b’:int inst([ce1.1, k1.1| sem(up1.1) is k1.1, k1.1 is [x,e|screw(x),e:grasp(cnst,x)], generate(up1.1,ce1.1), ce1.1:directive(inst,cnst,k1.1)])] direct inference from the lexicon to the plan an alternative analysis of the inference process leading cnst to conclude (5.3.3) is that cnst recognized that inst is performing directive (c.1) in (4.2.5) directly from the lexical meaning of u1.3 in (5.3.1), which is a mention of a grasping action. in other words, mentioning a grasping action is sufficient by itself to activate step (c.1) of the shared plan in (4.2.5), without the step of existential closure resulting in (5.3.2). although the end result would again be (5.3.3), this second hypothesis has the advantage of avoiding the need to stipulate that existential closure takes place after each and every utterance; and there is increasing evidence for this type of ‘surface’ semantic reasoning (see, e.g., the work by ferreira et al on ‘good enough’ representations, (ferreira, ferraro, and bailey 2002)). 5.3.2 acquiring intention (b) to perform a completion as discussed in section 4, we can think of at least four possible explanations for the fact that cnst acquires intention (b) to perform a completion: responding to a request, voluntary coordination control (clark’s ‘collaborative completions’), cooperativeness, and ‘blurting out’. all these four explanations assume that cnst recognizes (5.3.1) as a partial plan for performing the contribution in (5.3.3). we discuss each in turn. cooperativity what we called the ‘cooperativity explanation’ in section 4 is that cnst decides to change the current contribution by modifying the directive into a joint action (and adding some details). revising a contribution in this way is a sort of refashioning (clark and wilkesgibbs, 1990, p. 481-486) of the du that represented the contribution. as discussed in section 3, refashioning is formalized in terms of the system of grounding actions proposed by traum (1994) and adopted in ptt as a repair grounding act: cnst proposes to replace du1.1cnst, according to which it is inst alone who performs the directive, with the concatenation of du1.1cnst and du1.2, du1.1cnst;du1.2 (shown in (5.3.4)), in which the directive is a joint action of inst and cnst. du1.1cnst;du1.2 still contains all the information in figure 5.3.1, including the information about utterances and the fact that sem(up1.1) = k1.1. (5.3.4) [i1.2a| i1.2a:int cnst([du1.2,ce1.2| du1.2 is [ce1.1, k1.1| k1.1 is[x,e|screw(x),e:grasp(cnst,x)], ce1.1:directive(inst&cnst,cnst,k1.1)] ce1.2:repair (cnst,du1.1 cnst,du1.1 cnst;du1.2)])] if the repair is accepted, (each participant’s view of contribution) du1.1 gets replaced by du1.2, as seen in section 3 and discussed below. (arguably, cnst is also acknowledging the partial contribution represented by du1.2; we will discuss this additional function of the completion shortly.) after deciding to turn the directive into a joint action, cnst has to decide how best to complete it—i.e., which additional utterance actions to perform in order to generate ce1.1 as in (5.3.4). again, cnst may have determined which utterances are missing from (5.2.4) in a number of ways. the simplest explanation is that a sort of structural alignment took place: cnst develops her own plan for performing the directive in (5.3.4), and then compares this plan with the actions she observed. cnst’s plan to perform ce1.1 is a series of utterances very much like those in (5.2.4), except that it includes the generic “eine schraube” instead of the more specific np “eine completions, coordination and alignment in dialogue 45 orangene schraube mit einem schlitz”, so the conventional meaning she associates to up1.1 is the generic proposition to grasp a screw. cnst’s utterance plan is shown in (5.3.5).30 (5.3.5) [ce1.1, u1.2, up1.1, u1.3, up1.2, up1.3, up1.4, up1.5, up1.6, u1.4 | s(up1.1) , sem(up1.1) is k1.1,31 u1.2:“jetzt“:advp, s(up1.2), u1.2 ↑ up1.1, up1.2 ↑ up1.1, u1.3:“nimmst“:v, np(up1.3), vp(up1.4), up1.3 ↑ up1.2, up1.4 ↑ up1.2, u1.3 ↑ up1.4, np(up1.5), up1.5 ↑ up1.4, np(up1.6), up1.6 ↑up1.4, u1.4:“du“:np, u1.4 ↑ up1.5, u1.5:“eine“:det, u1.6:“schraube“:n, u1.5↑ up1.6, u1.6 ↑ up1.6] having produced this utterance plan / syntactic structure, cnst compares it with the utterances she has observed inst performing.32 cnst then decides to include in the new contribution she is planning, du1.2 (shown in (5.3.4)), the missing information about np up1.6. the update to the information state with this intention, which is satisfaction-preceded by i1.2a, is shown in reduced form in (5.3.6). (5.3.6) [i1.2b| i1.2b:int cnst(du1.2 += [ | up1.6:“eine schraube“:np]), sp(i1.2b) = i1.2a] (by the definition of +=, (5.3.6) is equivalent to: [i1.2b| i1.2b:int cnst([tmpk| tmpk is du1.2]; [du1.2| du1.2 is tmpk; [ | up1.6:“eine schraube“:np]]), sp(i1.2b) = i1.2a]. (notice that the fact that i1.2a satisfaction-precedes intention i1.2b ensures that i1.2b refers to the same du1.2 as i1.2a, according to the grosz and sidner view of discourse accessibility as formalized in (poesio and traum, 1997). see section 3.5. accessibility in ptt is discussed more extensively in (poesio and rieser, in preparation).)) the fact that cnst decides to utter an np instead of producing a complete utterance such as, say, “ich nehme eine schraube” might also be explained in terms of cooperativity: reusing as much as possible of inst’s output would make ce1.1 a genuinely joint action. in case existential closure were assumed, alignment at the domain level (rather than at the syntactic one) could also be used to explain how cnst plans her contribution to the joint directive (see (4.2.5)(c)). that is, cnst might have decided to utter an np by comparing the content of the jointly intended directive in (5.3.4) –the proposition to grasp a generic screw—with the proposition of grasping some object in (5.3.2), derived from the partial directive that inst has 30 as correctly pointed out to us by a reviewer, assuming that cnst’s expectations about the utterance plan for the directive are so detailed as to include the intention to generate every single word is a big assumption, and could reveal problematic in the case of indexicals, for although in this case it seems correct to assume that the utterance plan contains 2nd person du and nimmst as they were uttered by inst, in case inst hesitated earlier, only producing so jetzt, it would seem more plausible to expect that cnst would produce 1st person nehme ich eine schraube rather than 2nd person nimmst du .... with respect to the first issue we will simply reiterate the point already made in sections 4 and 5.1—that the view of utterance planning adopted here is an extreme simplification—adding that in addition a much more complex theory of matching one’s expectations with the other participant’s production will also be required. such a theory will have to allow matching one participant’s first person verb with the other participant’s second person. with such a theory in place, the second issue will become unproblematic as well. 31 we are forcing the notation a bit here and use the is notation to indicate the condition λ i. sem(v(up1.1)(i)) = v(k1.1)(i). 32 this assumption that participants in a conversation always compare their expectations with what they observe is well accepted in cognitive science: for instance, it is at the basis of wolpert, doya and kazato’s (2003) ‘mosaic’ model of language comprehension, and also appears to be broadly consistent with pickering and garrod’s (2004) `alignment’ model (which will be discussed at greater length in a following section). poesio and rieser 46 performed by means of existential closure. cnst would then conclude that this latter proposition should be augmented with k1.1d: (5.3.7) k1.1d=[x|screw(x)] cnst could then decide to produce an utterance with k1.1d as content – i.e., an indefinite np, very much as in (5.3.6). (cfr. pickering and garrod’s “i/o constraint”.) acknowledgment completions are often produced in order to acknowledge an (as yet incomplete) du (see clark’s ‘collaborative completions’, clark 1996, p. 238). in other words, cnst’s intention in uttering “eine schraube” may have been to acknowledge (the complete version of) inst’s directive, instead of, or in addition to, repairing it, as proposed in (5.3.4). this possible interpretation of cnst’s intention is shown in (5.3.8). the content of intention i1.2a in (5.3.8) is an acknowledgment of the completed du1.1 cnst in (5.3.3) –i.e., not simply the utterance, but the utterance together with the results of inferences: (5.3.8) [i1.2a| i1.2a:int cnst([ce1.3| ce1.3:ack(cnst,du1.1 cnst)])] the effect of this acknowledgment is to make du1.1cnst grounded: i.e., to update the discourse situation with g(du1.1cnst) (see below). a few aspects of this interpretation of what happened in the sample dialogue are worth discussing. first of all, one might argue that what gets acknowledged by cnst is not just du1.1cnst, but the proposition resulting from the concatenation of du1.1cnst with the record of the completion taking place: du1.1 cnst += [ | up1.6:“eine schraube“:np] i.e., that what cnst actually intends is the following: (5.3.8’) [i1.2a| i1.2a:int cnst( (du1.1 cnst += [ | up1.6:“eine schraube“:np]); [ce1.3| ce1.3:ack(cnst,du1.1 cnst)])] which is equivalent to: [i1.2a| i1.2a:int cnst( [tmpk | tmpk is du1.1 cnst]; [du1.1 cnst| du1.1 cnst is tmpk; [ | up1.6:“eine schraube“:np]]; [ce1.3| ce1.3:ack(cnst,du1.1 cnst)]) (note that acknowledgments are not parts of discourse units (traum, 1994; matheson poesio and traum, 2000), so (5.3.8’) is a well-formed proposition as ce1.3 is not part of du1.1 cnst. ) secondly, it might be argued that an intention to acknowledge is part of cnst’s intentions even under the ‘cooperativity’ interpretation discussed above; i.e., that intention i1.2a in (5.3.4) should actually incorporate parts of (5.3.8), as follows: (5.3.4’) [i1.2a| i1.2a:int cnst([du1.2,ce1.2, ce1.3| du1.2 is [ce1.1, k1.1,up1.6| k1.1 is [x,e|screw(x),e:grasp(cnst,x)], ce1.1:directive(inst&cnst,cnst,k1.1), up1.6:“eine schraube“:np ], ce1.2:repair (cnst,du1.1 cnst, du1.1 cnst;du1.2), ce1.3:ack(cnst,du1.2)])] the distinction is subtle, which is as it should be as interpretation seems to proceed in the same way irrespective of the ‘true’ intention. the main remaining difference between the acknowledgment interpretation in (5.3.8’) and this new repair + acknowledgment interpretation in (5.3.4’) is that in the case of (5.3.8’), the directive would be still considered inst’s action. just as in the case of the cooperativity interpretation, in the case of acknowledgment interpretation in (5.3.8) cnst would have a number of ways to realize her intention. notice that a repetition of the partial directive would be appropriate as an acknowledgment of that part only (i.e., without making a completion). as in the case of the cooperativity explanation, to decide how to execute the acknowledgment cnst would have to determine what’s missing from inst’s completions, coordination and alignment in dialogue 47 directive either by a structural or by a domain-level match. either way, the result of this decision process is that cnst acquires an intention along the lines of (5.3.6) or (5.3.8’). responding to a request as pointed out before, another possible reason for the completion is that cnst interpreted the pause in 1.1 as a request. the system of grounding acts developed by traum (1994) and adopted in ptt (see section 3.6) includes three types of requests that inst may have intended to perform: request for acknowledgment, request for continuation, and request for repair. in all three cases, if cnst adopts this interpretation she will update her information state by adding an obligation to address the request (the mechanics of obligations in ptt is discussed in some detail in (matheson poesio and traum, 2000). after that, it will be this obligation (rather than cooperativity, or a voluntary desire to acknowledge) that will lead cnst to acquire intentions such as (5.3.4) or (5.3.8). once these intentions are acquired, however, the result will not be any different from that discussed above. blurting out finally, it is possible that cnst’s action is not motivated either by a decision to be cooperative at the domain level (i.e., by acquiring intention (5.3.4)) or by an intention to acknowledge (as in (5.3.8)), but by impatience – i.e., by a view that inst is taking too long to perform the action. according to this explanation, cnst quickly recognizes that inst has intention i1.1b’ in (5.3.3), and realizes what is missing by a structural or domain level match, just as explained above. however, cnst is also moved by other goals, such as completing the task quickly, and it is as a result of these goals that she acquires intentions (5.3.4) or (5.3.8). we do not have a full account of what these other goals might be and how they lead to the acquisition of one of these intentions, but we do expect these intentions to be expressible in terms of the formalism assumed here. 5.4 the repair: 1.3, “eine orangene mit einem schlitz” upon hearing the completion, inst’s view of the discourse situation is predicted by the analyses above to be as depicted in figure 5.4.1 (next page). inst’s behavior upon hearing the completion is consistent with him, first of all, recognizing that the np uttered by cnst is intended as a completion of the utterance inst himself previously performed: i.e., that he is meant to update du1.2 with the information in (5.4.1)—that the np uttered by cnst is meant to fill the position of the missing np in the tree in (5.2.4): (5.4.1) du1.2 += [ | up1.7 is up1.6] he then appears to ascribe to cnst, one the basis of the information in figure 5.4.1, one of the intentions in (5.3.4), (5.3.8), or (5.3.4’). in all of these cases cnst intends to perform a grounding act having as one of its arguments discourse unit du1.1 from (5.3.1) augmented with the information that a screw is needed (as well as, possibly, with changes to the directive, and/or information about the utterance). in the case of (5.3.4), the intention is of repairing du1.1; in the case of (5.3.8), the intention is to acknowledge it; in the case of (5.3.4’), to do both. (we might call this du the du ‘under discussion’, to borrow a term from ginzburg.) we will proceed here under the assumption that inst attributes to cnst the most complex intention, that in (5.3.4’)—i.e., the intention both to be cooperative and to acknowledge the contribution in du1.1 augmented with the information that the object that inst wants cnst to pick up is a screw. (in this way, we will cover most of the ideas that would be needed to analyze what happens in the dialogue under the other interpretations.) as a result, inst’s view of the discourse situation is updated as intended by cnst, as shown in (5.4.2). (5.4.2) du1.2 += [ce1.1, k1.1| k1.1 is [e,x|screw(x),e:grasp(cnst,x)], ce1.1:directive(inst&cnst,cnst,k1.1)]; [ce1.2, ce1.3| ce1.2:repair (cnst,du1.1 inst, du1.1 inst;du1.2), ce1.3:ack(cnst,du1.1 inst)] poesio and rieser 48 at this point, inst appears to ‘accept’ the repair proposed by cnst, but also to have decided that cnst’s description of the screw may not ensure that cnst chooses the correct screw (the one which we have called orange-slit-screw) and so decides to perform a continuation: an expansion of the description of the object to be grasped. also, inst may or may not ground du1.1, depending on whether he recognizes an acknowledgment. as said in section 3, we need two conditional update rules for grounding acts here: up.repair ce:repair (a,du1,du2) � (accept(b,ce) � du1 ⇐⇐⇐⇐ du2) up.ack ce:ack(a,du) � (accept(b,ce) � [ | g(du)]) (we remind readers that an update rule is not material implication, but a rule specifying how an information state is updated by concatenating new information (larsson and traum, 2000).) the result of inferring the intentions in (5.4.2), then, is that in case inst accepts the repair he will proceed to update his information state by replacing du1.1 with the du consisting of the concatenated contents of du1.1 and du1.2, as follows: (5.4.3) du1.1 ⇐⇐⇐⇐ du1.1;du1.2 which, as seen in section 3.5, is equivalent to the update: λiλj i[du1.1]j ∧ v(du1.1)(j) = v(du1.1;du1.2)(i). if inst also recognizes and accepts an acknowledgment, he will perform the further update: (5.4.4) [ | g(du1.1) ] inst then proceeds to plan his continuation. such expansions of an object’s description are also viewed in ptt as a type of refashioning; this time, of the du which is the result of the repair figure 5.4.1. inst’s view of the discourse situation after perceiving cnst’s completion in 1.2. du1.2 is … i1.1b i1.1c i1.1d du1.1 du1.2 up1.1:s u1.2:”jetzt”:advp up1.4:vp up1.3:np u1.3:”nimmst”:v up1.5:npi up1.6:np u1.4:”du”:pro up1.7 u1.5 u1.6 ub1.1 up1.7:np u1.5:”eine”:det u1.6:”schraube”:n ub1.1:n’ … i1.1d … … du1.1 is completions, coordination and alignment in dialogue 49 performed by cnst in 1.2, the just updated du1.1 inst. with his continuation, inst intends to repair du1.1 inst with a new du (called here du1.3) in which the argument of the directive, k1.1, is updated by adding to its previous content (in 5.4.2) the additional information about the screw x which inst believes it will be necessary in order for cnst to identify the appropriate screw. this intention is represented in (5.4.5) as intention i1.3a of performing a repair ce1.5, replacing du1.1 inst with the new du1.3, in which k1.1 is updated with the information in k1.3. (5.4.5) [i1.3a| i1.3a:int inst([du1.3, ce1.4, ce1.5 | du1.3 is du1.1 inst ; [k1.3 | k1.3 is [ | orange(x)];[ x3|slit(x3), has(x,x3)]]; (k1.1 += k1.3), ce1.4:ack(inst,du1.1 inst), ce1.5:repair (inst,du1.1 inst, du1.3)] (notice, incidentally, that at this point the content of directive ce1.1 (proposition k1.1) is no longer the content that inst had originally planned, shown in (5.2.5). i.e., the ptt account of this episode of collaboration in dialogue makes this core speech act a genuine case of ‘joint production’ in the sense of clark (1996).) furthermore, inst reaches the decision to realize the planned refashioning in (5.4.5) as an apposition, as it is often the case in dialogue. this makes a treatment of appositions essential for a theory of incremental dialogue processing, which has to be one based on a syntactic theory like ltag in which such incremental construction can be modeled. we return to appositions shortly. in the ptt framework there are at least two ways to explain how inst reaches intention i1.3a in (5.4.5). one explanation is that this intention is the result of plan matching: matching cnst’s version of the directive in (5.4.2) with inst’s originally planned directive in (5.2.5). according to this view, inst compared the content k1.1 of directive ce1.1 that would follow from cnst’s repair, with the content of ce1.1 according to his own original intention, shown in (5.2.5). as a result, inst identifies the need to produce the additional information in k1.3. a second explanation is that inst acquires the intention as the result of situation matching: inst considers the situations that may result from the execution of the content k1.1 of the directive as proposed by cnst, and realizes that there are 9 such situations, one for each of the screws in figure 4.2.4. this appears to be what pickering and garrod (2004) have in mind. (such simple planning can be done quite efficiently, as shown, e.g., by the work on planning in the trains project (allen and ferguson, 1994, ferguson et al, 1996).) as an aside, the fact that utterance unit 1.3 was actually produced in two prosodic units suggests that inst identifies the properties that have to be added to the description of the screw produced by cnst in order to uniquely identify the required screw incrementally, in two steps (levelt, 1989). first of all, inst realizes that the property [ | orange(x)] is required; then, that the property of having a slit, [ x3 | slit(x3), has(x,x3)] is also needed. each of these two properties is going to be independently realized using an apposition, further argument for the need for a theory of appositions. in ptt, this two-stage process would be modeled by attributing to inst two distinct intentions: an intention to enrich k1.1 with the information that screw x is orange (shown in (5.4.6)) and an intention to enrich this new description with the additional information that x has a slit, shown in (5.4.7). (5.4.6) [i1.3a’| i1.3a’:int inst([du1.3,ce1.4| du1.3 is du1.1inst ; [k1.3| k1.3 is [ | orange(x)]] ; (k1.1 += k1.3) , ce1.4:repair (inst, du1.1 inst, du1.3)])] poesio and rieser 50 (5.4.7) [i1.3a’’| i1.3a’’:int inst([du1.3’, ce1.4’| du1.3’ is du1.3 ; [ k1.3’| k1.3’ is [ x3 | slit(x3), has(x,x3)]; [ | orange(x)]] ; (k1.1 += k1.3’) , ce1.4’:repair (inst, du1.1 inst, du1.3’)]) we will however simplify matters in what follows and treat the whole of 1.3 as a single apposition. our treatment of appositions, and the derivation of the meaning of utterance 1.3, are discussed in appendix b.4; we just give a brief summary in this section. one assumption one is forced to make, given the properties of ltag, is that eine is ambiguous between the ‘normal’ interpretation and the form used in appositions, henceforth eineapp, to which is associated the auxiliary tree and lexical meaning in (5.4.8), which yields a type [[π],[π]] (<, > ) for the appositional np: (5.4.8) n+ n npapp detapp nadj + eineapp: λp’λpλx([v| ]; p(x); p’(x); [ | v is x] ) we further assume that the appositional np is syntactically adjoined to the previous utterance, resulting in the ltag analysis for eine schraube, eine orangene mit einem schlitz shown in (5.4.9). (5.4.9) np n’ npapp nadj n’ ppdat npdat det n detapp nadj pdat det n eine schraube eineapp orangene mit einem schlitz whose interpretation is shown in (5.4.10) (for the full derivation, see appendix b.3). (5.4.10) λp([x| ]; [v| ]; [ |screw(x)]; [ |orange(v)]; [ | v is x]; [y| ]; [ | slit(y)]; [ | has(x,y)]; p(x)) apart from the syntactic and semantic issues, the planning of (5.4.9) –or, more precisely, of its appositional part)—proceeds just as the planning of the completion discussed in section 5.2. completions, coordination and alignment in dialogue 51 5.5 acknowledging and accepting: 1.4, “ja” the next utterance in the example dialogue, 1.4 ,”ja”, illustrates both the fact that utterances can be used to achieve multiple communicative intentions at different levels—in this case, to perform both an acknowledgment act at the grounding level and a backward-looking core accept act – and the differences between these two types of backward-looking acts. at the grounding level, utterance 1.4 signals an acknowledgment of the refashioned discourse unit resulting from the repair that inst performed with utterance 1.3. i.e., utterance 1.4 appears to perform at the grounding level the same function as utterances 1.2 (in interpretations (5.3.4’) and (5.3.8)) and 1.3 (in interpretation (5.4.5)). as a result, both agents’ information state will be updated with the equivalent of (5.4.4), i.e., g(du1.1). at the task level, the function of 1.4 is to accept the (joint) directive resulting from the previous three utterances in this first contribution after the initial production by inst was refashioned first by cnst and then by inst, i.e., with the version of k1.1 in (5.4.5). the process leading to the planning of these two actions in response to the directive elaborated in the previous turns, and their effect on the discourse situation, have been the focus of several previous ptt papers (e.g., (matheson et al, 2000)) so we only discuss it briefly here. as discussed in section 3, in ptt every conversational action induces an obligation on the other participant to address that action. at this point in the conversation cnst has two obligations: (i) to address the directive: obl(cnst,address(cnst,ce1.1)) (ii) to perform a grounding action with respect to its content—informally, obl(cnst,grounding-act(ce1.1)) there are several ways to address a directive, two of which are accepting it and rejecting it. this is modeled in ptt by assuming that (part of) the formalization of actions such as directives consists of update rules leading to the conclusion that a participant has addressed a conversation action if certain types of responses have been produced. this in turn leads to the removal of the obligation. these conditional update rules take the following form: ce:directive(a,b,k) � (accept(b,ce) � address(b,ce)) ce:directive(a,b,k) � (reject(b,ce) � address(b,ce)). i.e., if a performs a directive to b, and b accepts that directive, then that counts as having addressed the obligation of addressing the directive. this removes the obligation to address the conversational action, but of course, accepting a directive then leads to a new obligation, to perform the action directed. this is formalized using another update rule which, in a schematic fashion, can be formalized as follows: ce:directive(a,b,k) � (accept(b,ce) � obl(b,k)) the same happens with obligation (ii), the obligation to perform a grounding act. the performance of any of the grounding acts in (traum, 1994) will discharge the obligation; the performance of any of the grounding acts will also lead to further updates – e.g., acknowledging a du results in that contribution being grounded, as discussed in section 3: ack(a,du) � g(du) in summary, as a result of the performance of 1.4, the discourse situation gets updated as in (5.5.1): (5.5.1) [ce1.6, du1.4| ce1.6:ack(cnst, du1.1 cnst), g(du1.1 cnst), du1.4 is [ce1.7, s1.1| ce1.7 : accept(cnst,ce1.1), s1.1 : obl(cnst,k1.1)]] where du1.1 cnst is the du resulting from the two repairs in 1.2 and 1.3: poesio and rieser 52 [ce1.1, k1.1 | k1.1 is [e,x,x3| screw(x), orange(x), slit(x3), has(x,x3), e:grasp(cnst, x)] ce1.1:directive(inst&cnst,cnst,k1.1)] this completes the first contribution in (1.1). 5.6 the rest of the conversation an analysis of the second contribution in the sample dialogue –utterances 2.1-2.4—at the level of detail just given for the first one would require a discussion of numerous additional aspects of ptt, including the treatment of anaphoric accessibility, and of a number of additional grammatical issues, such as our treatment of connectives and the german `satzklammer’. as it is just not possible to present all this material in a single journal paper, we will only provide a sketchy discussion of the main issues from the point of view of coordination, and defer the additional discussion to a second paper in preparation (poesio and rieser, in preparation). 2.1: “und steckst sie dadurch, also” with utterance 2.1 inst moves on to the next step in the shared plan in figure 4.2.3 and in (4.2.4), which turns into step (e) in (4.2.5). inst wants cnst to put the screw through the aligned holes of the longer bar serving as ‘fuselage’ and the shorter bar serving as ‘wings’ of the toy airplane. from a production / sentence planning perspective (levelt, 1989; dale and reiter, 2000), there are very few differences between the steps that inst goes through in planning utterance 2.1 and those that were involved in planning utterance 1.1. again, inst has the intention to perform a directive (let’s call it ce2.1); and again, inst decides to plan an utterance up2.1 generating directive ce2.1. (as said above, a full analysis of the particular realization of ce2.1 chosen by inst for the example dialogue would require a discussion of three linguistic phenomena: rhetorical relations (und, also), the so-called `satzklammer’ (steckst … dadurch ) and the pronoun sie, that raises the issue of anaphoric accessibility within micro conversational events. we defer this discussion to a second paper.) and again, cnst interrupts inst before the plan for up2.1 is completely executed. 2.2: “von oben” what happens at this point is very similar to what happened during the first contribution when cnst performed the completion in 1.2, and we believe our analysis of that completion applies with relatively few modifications to this case as well, except that there is very little evidence in this case that anything in inst’s performance could be interpreted by cnst as a request, so the other explanations (cooperativity, voluntary coordination control, and blurting out) look more plausible in this case than the request for acknowledgment. in fact, in this case our intuition tells us that there is a stronger sense that cnst is simultaneously acknowledging inst’s contribution and refashioning it. the main issue raised by this utterance is one of compositional semantics: how the meaning of the modifier von oben is integrated with the meaning of inst’s contribution. in the later paper we discuss two options: maintaining the simple treatment of vps as having type used so far, and going to the higher > type common in work on tense based on a parsonian semantics (parsons, 1990). 2.3: “von oben, daß also die drei festgeschraubt werden” with this utterance, inst achieves two objectives. first of all, he acknowledges cnst’s participation to the current du—this time by explicitly repeating cnst’s utterance. he then returns to the utterance plan which had been interrupted by 2.2, and continues the production of the purpose clause. the purpose clause also contains a second example of (in this case, deictic) reference: the reference to the three constituents of the airplane previously aligned. completions, coordination and alignment in dialogue 53 2.4: “ja” finally, the ja in 2.4 is produced (and interpreted) more or less like 1.4. we will just briefly remark that cnst seems not to have had any problems interpreting the definite description die drei. 5.7 recap this concludes our intentional analysis of the completion under discussion. in summary, we argued that it is possible to provide a fairly explicit analysis of four possible reasons for the completion by making the following assumptions: (i) that utterances are actions, and that planning an utterance is akin to planning any other action (which is not to say that the same planner is used!!) (ii) that inst and cnst have shared plans interleaving domain and discourse actions, as in (4.2.5); (iii) that the participants to a conversation—or at least, to a fairly structured task-oriented conversation like those in the bielefeld toy plane corpus—continuously generate expectations about what’s to come next in the dialogue on the basis of these interleaved and shared plans, and that they monitor each other’s utterances comparing them with these expectations so as to be able to recognize where they stand in the conversation; (iv) that they are able to do this comparison before an utterance is completed, perhaps via some form of existential closure; and finally (v) that as a result of that they can plan appropriate grounding acts; in particular, we have proposed that completions can be viewed as a form of repair. we would argue that none of these assumptions is ad-hoc, and that at least (i), (iv) and (v) are necessary to any account of the phenomenon at hand. in the following section we discuss an alternative account which does not make assumption (ii) and that adopts a different view of the monitoring process. 6 a non-intentional analysis pickering and garrod (2004) while encouraging psycholinguists to take up dialogue as a valid area of research, argued that many aspects of dialogue interpretation—particularly priming at the lexical, syntactic, and referential level (morton, 1969; branigan et al, 2000)—can be explained without resorting to an intentional analysis, and sketched instead a model of interaction centered around the notion of alignment between the dialogue participants’ models. although pickering and garrod’s alignment model is more of a programmatic proposal than a fully worked out theory, it has proven very influential (purver et al, 2006; rickheit and wachsmuth, 2008). as a contribution to the ongoing discussion concerning this model, in this section we briefly discuss which aspects of cooperations could be accounted by an alignment model, and how. 6.1 the pickering and garrod alignment model the starting point of pickering and garrod’s analysis is the consideration that successful communication between dialogue participants a and b requires them to go through the same stages of interpretation and production, illustrated in figure 6.1 (next page). pickering and garrod (henceforth: p&g) also note that priming effects can be observed at each of these levels of interpretation, from the phonetic (bard et al, 2000) to the lexical (garrod and anderson, 1987) to the syntactic (bock, 1986; branigan et al, 2000). finally, p&g discuss experiments suggesting that in order for dialogue to be felicitous, the participants need to produce aligned ‘situation models’. the term ‘situation model’ is widely used in the psycholinguistic literature on reading and text processing to indicate a variety of representations expressing information about the events and entities described in a narrative (theories of situation models have been proposed in, e.g., (sanford and garrod, 1981; johnson-laird, 1983; garnham, 2000). but what p&g have in mind is a much more general notion of situation model—and one which may change from task to task, often being developed on the spot by the participants to the task. for example, in one of the studies they discuss, by garrod and anderson (1987), the poesio and rieser 54 participants in a `cooperative maze game’ have to develop ways of describing each other’s position in a maze in order to follow a route. the subjects were found to develop `aligned’ views of the maze in the sense that they would jointly develop a protocol for referring to maze positions, and particular positions, confirming the results of work of clark and colleagues (clark and wilkes-gibbs, 1986; clark, 1996). figure 6.1. the levels of representation involved in (spoken) language comprehension. (from (pickering and garrod, 2004); reprinted with permission.) the conclusion that pickering and garrod draw from these observations is, first, that `aligning representations’ at all levels is a central aspect of dialogue processing; and that this alignment is the result of fairly simple comprehension mechanisms, rather than of intention-based mental state reasoning. (such `more complex’ types of processing are considered as an option, but not the basic form of processing available to dialogue participants.) a phenomenon that, according to p&g, can be predicted from the interactive alignment models is the creation of routines. p&g point out that dialogue is highly repetitious: e.g., aijmer (1996) suggests that up to 70% of words in the london-lund corpus occur as part of recurrent sequences of words. according to p&g, this follows from the alignment assumption: during production, participants in the dialogue will tend to choose expressions or syntactic structures which have already been produced as part of the present (or, perhaps, earlier) dialogues. 6.2 an alignment-based analysis of the example dialogue what would an account of sentence cooperations according to the interactive alignment model look like? we find it useful to divide this question in two parts: explaining how cnst reaches her conclusion about what inst is trying to achieve; and explaining how she decides to produce a completion. (we concentrate on the completion.) 6.2.1 interpreting inst’s actions the first question that has to be addressed is: what would the situation model be in a task in which the participants are not simply describing reality, but modifying it via actions? we think at least two answers are possible: completions, coordination and alignment in dialogue 55 (i) the situation model is a representation of the state that cnst has reached in assembling the toy plane: i.e., what has already been assembled, and which parts are still available. (ii) the situation model is the shared plan itself: i.e., which of the required actions have already been performed, and which still have to be. most, if not all, theories of planning view plans such as those discussed in the previous sections as representations –typically, tree structures—which can be manipulated by, e.g., decomposing an action in its subactions. (in fact, the idea that plans can be viewed as mental constructs is fairly recent, dating back to pollack’s dissertation (1986).) if we assume for this task a situation model of the second type, the process by which cnst would recognize inst’s actions under the iam would not be so very different from the process discussed above for the intentional model—it would simply require cnst to find a matching action in the plan, as done, e.g., by kautz (1987). we concentrate therefore on explaining how cnst could reach such a conclusion if the situation model was of type (i). one obvious difficulty is that inst and cnst’s situation models in this sense are not really `globally aligned’ in the sense of being identical until the very end of the conversation, as both inst and cnst are perfectly aware. it is however possible to argue that achieving the desired alignment is the entire point of instruction-giving conversations, which are not therefore an exception to the ia model. coming to what could be a `priming’ alternative to recognizing inst’s intentions by looking at a shared plan, we believe that a candidate can be found by generalizing the notion of `routine’ to non-linguistic actions. namely, it could be argued that during a typical conversation of the btpc, an aggregate formation routine is established: a repeated pattern of actions which together lead to the same result (namely, two parts of the toy airplane get joined). in the dialogue from which (1.1) is excerpted, inst first instructs cnst to join the `side rudder’ to the `fuselage’; then to join the tail to the fuselage; etc. after a while, cnst may be able to abstract over these repeated actions and produce the routine in figure 6.2. figure 6.2. routine created during a toy plane assemble dialogue once this routine becomes part of the common ground, an instance can be recognized through simple structural match, without need for explicitly reasoning with intentions. in the case under consideration, for example, cnst may be able to recognize that nimmst is an instance of the second step of the routine, and hence, that the missing argument must be a screw. she may also be able to identify the syntactic structure to use for the completion by looking at the realization of previous mentions of this type of object. aggregate formation parameters: c1, c2: constituents to be joined p1, p2: ports s: screw, with bolt b subactions: 1. align p1 of c1, with p2 of c2 2. get s 3. pass s through p1 and p2 4. fasten b poesio and rieser 56 6.2.2 deciding to produce a completion as the sketchy discussion above makes clear, in our opinion it would be possible to develop an alignment account of interpretation. what is not at all clear to us is how such a model could explain what happens next: i.e., why it is that cnst decides to produce a completion, instead of waiting for inst to finish his contribution. in other words, the alignment model as currently stated does not provide an account of dialogue management, an important aspect of what dialogue participants have to do (and one which had a significant influence in the design of ptt as currently conceived). 6.3 summary `traditional’—i.e., grice-inspired—models of dialogue processing tend to formalize all aspects of an agent’s behavior when involved in a dialogue in terms of intentions (see, e.g., cohen’s account of reference (cohen, 1979) or grosz and sidner’s theory of discourse structure (grosz and sidner, 1986)). we feel that the alignment model offers a very reasonable alternative to these gricean models as far as interpretation and production are concerned. however, intentions also play an important role in formalizations of dialogue management—how an agent decides what to do next—and we can’t find anything in the alignment model as currently stated which would explain how cnst reaches her decision to produce a completion, nor to explain what is it that inst is trying to do (a non-intentional alternative to speech acts, that is). 7 related work to our knowledge the only other proposal in the literature on dialogue explicitly concerned with completions is the paper by purver et al (2006), developed in the dynamic syntax framework (kempson et al. 2001, cann et al. 2005). we discuss dynamic syntax and purver et al’s proposal in this section, after however discussing the two other main theoretical frameworks for the semantics of dialogue, ginzburg’s kos (ginzburg and sag, 2000; ginzburg, 2009) and asher and lascarides’ sdrt (asher and lascarides, 2003), both of which provide theoretical insights that have been incorporated in our framework and that could perhaps be extended to provide a treatment of the phenomenon under consideration. 7.1 ginzburg’s theory of the dialogue game board arguably, the most developed modern theory of semantic interpretation in dialogue is ginzburg’s kos framework (ginzburg and sag, 2000; ginzburg, 2009). the theory was conceived from the start to deal with dialogue, and provides the widest coverage of dialogue-specific phenomena, including in particular grounding, clarification requests and non-sentential utterances; therefore it could probably be extended to provide an account of completions, at least for what concerns their interpretation. 7.1.1 general features of the model like ptt, kos is inspired by situation semantics (barwise and perry, 1983) and is founded on the assumption that many aspects of interpretation in dialogue can only be properly explained by providing an account of the discourse situation in its entirety. another similarity with ptt is that kos is based on an information state view of dialogue interpretation—and as in ptt, it is assumed that each participant’s information state – their dialogue gameboard (dgb)—contains both public and private information. in its basic form, a participant’s information state according to kos contains a record of speaker, addressee, the moves so far –whose first element is called the latestmove-the (shared) facts about the discourse situation, and a record of the questions under discussion, or qud. the information state is a record in type theory with records (cooper 2005). completions, coordination and alignment in dialogue 57 the qud field is one of the most distinctive features of this framework. its value is a stack (and in the general case a partial order) of the issues raised in a conversation. qud was hypothesized to account for the pragmatics of queries – and especially to explain how the propositional interpretation of short non-sentential answers is recovered—but , as we will see, it also plays a central role in kos’s treatment of grounding. simplifying matters considerably, when a asks b a query such as “who came?”, the dgb of both participants gets updated with the information that the latest move was a query: ask(a,b, λ x. come(x)). as a result, the issue raised by the query—in the case of a wh-query, an abstract: λ x. come(x)—gets added to the qud, as shown in figure 7.1.1. according to kos, the presence of this issue in qud is what enables b to provide a short answer to the question: say, “john”. the content j of this short answer can be combined with the content of the question that is maximal in qud to get a proposition, “john left”. when grounded, this proposition is added to facts. the effect of dialogue acts on the dgb—i.e., how the information state is updated as a result of, e.g., questions and answers-is modeled by conversational rules. a conversational rule specifies how information states that met certain preconditions , such as the presence in latestmove of an ask dialogue act, are updated as a result. for instance, the incrementation of qud with the issue resulting from a question is specified by a conversational rule called ask qud-incrementation (ginzburg 2009, p. 117).33 the overall series of updates to the information states of all conversational participants resulting from a change in the discourse situation is called a protocol; an example of protocol is the protocol for cooperative query exchange specified below (ginzburg 2009, p. 116). cooperative query exchange 1. latestmove.cont = ask(a,q) 2. a: push q onto qud; release turn 3. b: push q onto qud; take turn; make q-specific utterance. as this example shows, the updates specified by a protocol generally differ depending on the role of the conversational participant. after asking the question a will (normally) release the turn, whereas b will take it and perform a relevant utterance. as well as the most extensive treatment of dialogue phenomena, kos also comes with the most detailed fragment of any theory of dialogue, the hpsg grammar from ginzburg and sag (2000), subsequently modified in substantial ways in (purver, 2004; fernandez, 2006; ginzburg, 2009), inter alia. 7.1.2 grounding and clarifications (ginzburg, 2009) also includes a full treatment of the grounding process. the crucial advance of this proposal is that it provides a detailed account not just of successful communication, but also of failure in communication and how it is handled via clarification requests. ginzburg’s book 33 here and elsewhere page numbers refer to the last draft of the book available online. spkr: a addr: b moves: 〈….〉 latestmove: ask(a,b, λ x. come(x)) facts: 〈….〉 qud: 〈λ x. come(x)〉 figure 7.1.1. update of the dgb after asking a question. poesio and rieser 58 also contains an interesting theory regarding what is means to understand an utterance, and provides a formalization of the grounding criterion based on this theory. in its simplest form, the kos account of the process by which the dgb is updated as a result of utterance u is as follows: 1. if the utterance is understood, latestmove is updated with u (and the process of update continues as discussed above); 2. if the utterance is not completely understood, so-called crification ensues. note how utterances are not immediately added to latestmove upon being perceived. instead, they go to an information state field called pending, that contains ungrounded utterances (pending is very similar to the field called udus by matheson, poesio and traum (2000).) from there, utterances go to latestmove if grounded—but may also be revised as a result of the crification process. the content of pending are what ginzburg calls locutionary propositions (p. 148--149). a proposition in ttr is a type of record with the two fields [sit = s, sit-type = st] for which truth conditions are defined: the proposition above is true iff s is of type st, written s : st (ginzburg 2009, section 2.5.4). in locutionary propositions, sit-type is the utterance type resulting from parsing, whereas sit is the utterance. the situation type, in turn, is a record of type sign that encodes the same information encoded by signs in the earlier (ginzburg and sag, 2000): phonetic information in the phon field, syntactic category in the field cat, constituency information in the field constits, and semantic information. as in hpsg, the meaning of an utterance in kos is viewed as a function from contexts to contents, where contexts are a generalization of the notion of context due to kaplan (1989): whereas kaplan’s contexts were limited to a few values for indexicals like i, you, here, now , in the theory of contexts developed in situation semantics, particularly in the work of robin cooper (gawron and peters, 1990; cooper, 1992; cooper and poesio, 1994; cooper, 1996) a great deal of aspects of the interpretation are sought in contexts, including referents for proper names. hence, semantic information is encoded in two fields: the field c-params specifies the contextual parameters of a semantic object that have to be supplied by the context for the object to be interpretable, whereas the field cont specifies content proper—how these values and those supplied by lexical semantics are combined. for instance, the locutionary proposition corresponding to discourse unit du1 in (3.4.3) would be as follows: (7.1.1) [sit-type = [ phon = ”there is an engine at avon”, cat = s : syncat, constits = { there, is, an, engine, at, avon, an engine, at avon } : set (sign), c-params = [ spkr: ind, addr: ind, s0: sit, l: loc, a: ind, c1: named(a, “avon”)], cont = assert(spkr, addr, [sit = s0, sit-type = [x : ind, c1 = engine(x), c2 = at(x,a)]]) sit = [phon = “derisanenginatevn”, cat = s: syncat, constits = { u1(der), u2(is), u3(an), u4(engin),u5(at), u6(evn), u7(anengin), u8(atevn) }, c-params = [ spkr = a, addr = b, s0 = sit1, l = l1, a = avon, c1 = prop1] cont = ] comparing these propositions with the ptt encoding of the same contribution, it is clear that leaving aside the semantic differences between records and drss –potentially significant, but not entirely clear to us- the first main difference is that whereas in ptt the representation of the completions, coordination and alignment in dialogue 59 information state is more akin to what in situation semantics would be called a situation type, just as in drt and sdrt, the information state in kos consists of propositions specifying that indexical situational objects are of a certain situation type. it might be argued that the kos view is more natural, and a similar view of ‘anchored drt’ has also been advocated for drt by kamp e.g., in (kamp, 1990)—but the semantic differences become less clear with compositional drt, in which, unlike in vanilla drt, proper names and indexicals are treated as directly referential (i.e., are encoded as constants), not as existentially quantified. the one distinction which clearly seems to make a difference is the explicit encoding of situational parameters in kos. this explicitness is the basis for the formulation of the grounding criterion proposed by ginzburg and colleagues: that understanding an utterance means finding values for the contextually dependent parameters in c-params within the dgb (ginzburg, 2009, p. 153). in case of failure to ground—for instance, in case the referent of avon is unknown, or in case the addressee does not understand what is the relevant meaning of engine for this context—crification ensures. by contrast, in ptt we have failure of understanding in case an utterance can’t be assigned a meaning. the kos treatment is more general, offering for instance the potential to identify failure of understanding also in the case of failure of intention recognition, i.e., when no dialogue act can be assigned to an utterance. encoding contextual parameters in cdrt would not be particularly problematic, as they could be viewed as distinguished discourse referents (as indeed they were in (cooper and poesio, 1994)) but no attempt to compare the two types of grounding criterion and / or develop a version of ptt with contextual parameters has been made so far. the treatment of clarification requests in kos is based on the extensive empirical analysis of clarifications in the british national corpus carried out by purver (2004). purver classified clarification requests with respect to their form (identifying eight distinct types: wot, explicit, literal reprise, wh-substituted reprise, reprise sluice, reprise fragment, gap, and filler fillers corresponding to what we are calling here completions) and their content (identifying four distinct categories: repetition, clausal confirmation, intended content, and correction); many of these are treated in (ginzburg and cooper, 2004; ginzburg, 2009). according to kos, the process by which contextual parameters are instantiated begins by applying a contextual instantiation update rule (p. 153, 154) which updates the information state by ‘filling in’ parameters as soon as suitable instantiations are found. in case the instantiation of the parameters is only partial, clarification context update rules apply, of which two are discussed: parameter identification and parameter focussing. parameter identification applies whenever the value of one of the contextual parameters is missing and results in questions that can be paraphrased as “ what is the intended content of ui?” (e.g., “what do you mean by avon” or “who is bill?”. in kos terms, asking such a a question amounts to putting as maximal in qud a question of the form λ x mean(sprk, ui, x). parameter focusing produces confirmation requests on the basis of the value of the cont field. in both cases, the process simply involves putting an appropriate question on qud—in other words, clarification requests are supposed to work just like any other request. with these refinements, the process by which the information state is updated can be specified by the more complex protocol that follows (ginzburg, 2009, p. 180 – (96)): utterance processing protocol for agent a with is i: if a locutionary proposition lp = [ sit = u, sit-type = t] is maximal in pending: (a) if p is true, try to integrate p in a.dgb using a moves update rule (b) otherwise: try to accommodate p as a cr to latestmove (c) else: seek a witness for t by asking a clarification request. poesio and rieser 60 7.1.3 non-sentential utterances, self-initiated repair and incremental pending in chapters 6 and 7 of (ginzburg, 2009), two ingredients are added that make the kos framework appropriate for providing a treatment of completions. in chapter 6 an extensive treatment of non-sentential utterances is provided, based on the work of fernandez (2006), covering in particular short answers ( yes, no, john) and sluices of various types (as in, e.g., a: john left b: who?). that chapter discusses in detail how utterances providing a complete dialogue act can update and be interpreted with respect to the dgb even when syntactically they are not sentential. in chapter 7, and in particular in section 7.3, the possibility that pending can be incrementally updated (after every micro conversational event) is considered, which would provide a unified treatment of self-repair. such moves would make kos and ptt even more similar, making it virtually certain that solutions developed within one framework could be readily imported in the other.34 7.1.4 summary of the discussion to our knowledge no explicit treatment of the phenomenon of completions in its full generality has been provided in kos so far, but we don’t see any reason why the analysis provided in this paper could not be recast in the kos framework, once the differences between a treatment based on ttr and one based on drt are better understood. the main difference between the two frameworks is that ptt attempts to stay closer to mainstream formal semantics in terms of syntactic and logical underpinnings, whereas kos is closer to the unification-based world using hpsg as syntactic framework and ttr as semantic representation (hence, arguably, closer to treatments of dialogue in computational linguistics). it’s not clear to us however that such differences matter much in terms of the proposed treatment of completions; on the other end, the two frameworks share the basic assumptions about information state, common ground, and their update that we argued are crucial to the treatment provided here. arguably, the most interesting difference between the frameworks is in the grounding criterion, formulated in terms of lack of meaning for utterances in ptt, and of contextual parameters in kos. this is one area where the treatment proposed in kos could be adopted in ptt. 7.2 asher and lascarides’ logic of conversation 7.2.1 a general comparison between the design philosophies of loc and ptt asher and lascarides’ (2003) monograph contains a wealth of insights concerning the structure of dialogue and its systematic description. it is based on interesting ideas concerning the division of labour of, inter alia, information content, lexicon, cognitive states, and world knowledge. the formal apparatus is developed in great detail, integrating much current research in logics, linguistics, philosophy and ai, in a way echoing the “zeitgeist” in an attractive manner. an obvious similarity between loc and ptt concerns the underlying logic: both frameworks are derived from drt, of which they share the general logic assumptions about the dynamic nature of the common ground. ptt is more conservative, being based on compositional drt and hence, ultimately, on type theory in the sense of montague. loc is less traditional, both because meaning is characterized in terms of a logic more closely related to drt as formulated, say, in (kamp and reyle, 1993), and, above all, because a substantial effort has been made in loc to ‘modularize’ the process of interpretation and to identify distinct logical frameworks to characterize, for instance, semantic interpretation as opposed to meaning proper. no such effort has been made in ptt. 34 and indeed, purver (2004) provides an analysis of certain types of filler crs that might be considered cases of completions, as in a: is bill ... b: coming? completions, coordination and alignment in dialogue 61 another big difference, already pointed out in (asher and lascarides 2003, p. 304), is that ptt is speech-act and intention based, whereas loc’s theory of coherence rests on the concept of discourse relations such as elaboration, narration and the like and on various axioms like contrast guaranteeing/explaining local coherence. in ptt discourse relations and the like have played a very minor role so far, although argumentation acts are included among the dialogue acts covered in (poesio and traum, 1997). this is partly due to different preferences concerning the data used for setting up dialogue theory: loc sticks to the philosophy of language principle of using prototypical examples and pays less attention to the mapping from nl to logical forms, i.e. sdrss, which is more or less taken for granted, whereas ptt proceeds from natural data, especially from task-oriented dialogue, and tries to preserve nl surface using cdrt-style translations. so there are different guiding methodological principles here. discourse relations are a heavy machinery resting on rich intuitions concerning discourse structure and ptt is minimalistic in this respect. in the long run, ptt needs of course a theory of rhetorical structure-indeed, this is already required for the example transcript; this will be discussed in the follow-up paper. by contrast, in loc the idea of plan-basedness is boiled down to speech-act related goals (sargs), there are no large shared partial plan structures used; large structures emerge solely via discourse relations encapsulating recursively built up material. 7.2.2 completions in loc asher and lascarides state quite clearly that coordination of the sort discussed in this paper is beyond the scope of their book—e.g., when discussing their example (4), attributed to sacks, on p. 297, asher and lascarides say (footnote 3): ‘note that henry’s and mels’s utterances contribute a single proposition. we gloss over this here, however’. (4) a. joe: we were having an automobile discussion .... b. henry: discussing the psychological motives for c. mel: drag racing in the streets. nevertheless, loc has a lot to offer concerning the description of our example, repeated below, and appropriate extensions are not difficult to envisage. we will investigate this in some detail in this section. coordination in the production of single speech acts. two ingredients seem missing from loc in order to provide an account of such examples. in example (1.1), repeated here as (7.2.1), we have two cooperative dialogue moves, cnst’s production of a screw and inst’s other repair an <-> orange one with a slit, which is acknowledged by cnst. (7.2.1) inst: so, jetzt nimmst du well, now you grasp cnst: eine schraube a screw, inst: eine <-> orangene mit einem schlitz. an <-> orange one with a slit cnst: ja. yes. to handle cooperative productions one needs some sort of representational underpinning, such as joint intentions, or at least joint plan structures. in the analyses proposed in this paper shared plans, common intentions and common situations help to explain why alignment wrt completions occurs. misalignment arises simply as a consequence of the difference between shared plans and individual plans. in loc there is no real basis for the explanation of alignment, due to the lack of a principled treatment of fine grained coordination, despite the use of a cooperativity principle and claims to the contrary (p. 417). if one does not handle aligned structures, one misses a central feature of nl dialogue and neglects established insights, e.g. poesio and rieser 62 clark’s track-theory or his presentation-acceptance cycle for common productions. asher and lascarides’ notion of correction, which might be considered of relevance here at first sight (loc, pp. 345-354), solely rests on intuitions concerning denial and consistency, which is clearly not what one needs for the repair in (7.2.1), because there is no inconsistency there. furthermore, the decision to abstract away from explicit representation of utterances and speech acts means that some other mechanism will have to be introduced in order to account for non-sentential information state updates. indirect speech acts and acknowledgement in loc nevertheless, if we abstract from completion and coordination and take an idealised version of ex. 1, 1’, we can profitably ask what loc has to offer for that. as seen in sections 3 and 5, ptt has incorporated several general insights from sdrt concerning the representation of the content of dialogue acts, and more have been included in ptt’s treatment of anaphoric accessibility, discussed in the follow-up paper. in this section we discuss whether specific proposals about indirect speech acts and acknowledgments in sdrt apply to such examples. (1’) inst and cnst: well, now you grasp a screw, an <-> orange one with a slit. cnst: yes. the intuitions concerning an idealized (1’) which does not observe completion and coordination as phenomena sui generis are that inst and const’s turn is a request and const’s yes functions either as an acknowledgement, an acceptance act or an indication of the action carried out. we concentrate on the issues of request and acknowledgement here. before we can start that, a rough picture about the global theory levels going into loc might be of help (see figure 7.2.1): figure 7.2.1. interaction of logical modules in loc (from asher and lascarides (2003), p. 431, simplified) assuming that the first turn represents an unconventionalised indirect speech act, isa (loc, pp. 307-311), gricean style reasoning leads to the inference of an implicit speech act. according to loc we then get a complex type assertion• request for the first turn. complex types can be expanded either way, here we assume an expansion of request. where are we in the complex modularised structure in figure 7.2.1 after having applied gricean style resoning? presumably information lexicon (partial) description of content glue logic cognitive modelling cognitive states world knowledge completions, coordination and alignment in dialogue 63 within cognitive modelling. grammar alone will give us of course only an sdrs for the declarative sentence, which in linear notation would be as follows: [x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)]. what we want to have is, however, a request, i.e. it is the grasping that should be done. in loc-style notation this is (δ for do and the drs for the proposition which is requested): δ([x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)]). how one can proceed from grammatical information to cognitive modelling with respect to non-conventionalised indirect speech acts is not explained in detail in loc. we shall now see what we can do with δ([x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)]), i.e. which axioms of loc apply. the default request related goals (rrg) (loc, p. 394) states that if a request α is made, then the speaker’s goal is typically that the action aα it denotes be performed. this is indicated by the intention operator is(α). we obtain therefore: [x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)] ! > is ([x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)] ) (a[x, y| grasp (you, x), orange (x), screw(x), slit(y), with(y,x)]). we then need cooperativity, condition (a) (loc, p. 391), in order to model agents’ takeover of intentions: (a) ia(δ([x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)]) > ib(δ([x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)]) (where a = inst and b = cnst). following loc, p. 394, we can then replace is([x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)]) with sarginst([x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)], done(a[x, y| grasp (you, x), orange (x), screw(x), slit(y), with (y,x)]))). in addition, we can assume that cnst’s yes entails that he has accepted or achieved inst([x, y| grasp (you, x), orange (x), screw(x), slit(y), with(y,x)])’s sarg of ([x, y| grasp (you, x), orange (x), screw(x), slit(y), with(y,x)]). this being the case, the two turns, inst and cnst’s well, now you grasp a screw, an orange one with a slit and cnst’s yes, satisfy the semantic conditions of acknowledgement(αααα,ββββ) (loc, p. 466). 7.2.3 summary of the argument loc does not treat agents’ coordination in the production of single speech acts and is therefore not directly applicable. however, if one fuses inst’s and cnst’s contributions as in (1’), one can apply loc’s indirect speech act analysis and derive the speech-act related goal of inst’s, namely that he wants cnst to grasp a yellow slit bolt. in addition, one gets cnst’s take over of inst’s intention due to cooperativity as well as his acknowledgement. 7.3 dynamic syntax 7.3.1 the dynamic syntax framework the paradigms from psycholinguistics, linguistics and philosophy we have discussed so far all started from the assumption that agents’ joint verbal contributions have to be treated within a theory of dialogue. the ds framework, first summarized in kempson et al. (2001) started with the aim to ‘characterize structural properties of language in terms of the incremental process of building up interpretation from the sequence of words’ (kempson et al. (2001), p. ix). still, the poesio and rieser 64 empirical boundary of this endeavour was the sentence and, as far as we can see from the published work, at least in the beginning no extension toward dialogues such as e.g. the modelling of turn-constructional units or sequences of turns had been considered. however, in the course of the development of the paradigm since 2001 it turned out that ds modelling techniques for intra-sentential regularities such as anaphora, ellipses, relative clause construction or appositions generalize naturally to similar phenomena in dialogue. incrementality is, of course, the main methodological assumption ds shares with ptt (see section 3 and appendix b.2 in this paper). meanwhile (see gargett et al. (2008)) the state in ds theorizing appears to be as follows. there is a deep foundational debate among ds scholars whether some phenomena like, say, corrections or completions presuppose for their treatment a special paradigm, i.e. a dialogical one or whether these can be treated in a theory which is neutral vis à vis the sentence, monologue or dialogue distinction. the data of interest in this context are non-repetitive fragment forms of acknowledgements, clarifications and corrections. their example (7.3.1) below is very near the corpus data concerning completions which we discuss in this paper: (7.3.1) a: are you left or b: right-handed. there is a deep methodological issue here, namely which phenomena should be treated in which theoretical layer and how far sentence bound theories should be stretched. ds theoreticians maintain a twofold perspective: on the one hand they talk about a ds model of dialogue, and on the other hand they treat dialogue features within the frame of a proposition. we briefly look at these two aspects starting off with some of the basic assumptions and tools of ds, leaving, however, the wealth of technical details aside (concerning these, see cann, kempson, and marten (2005), purver, cann, and kempson, (2006) or gargett et al. (2008) for a short synopsis). basic assumptions and tools of ds (in the version of gargett et al. (2008)) ds is a parsing based approach using a strictly left-to-right interpretation of linguistic input. it works with growing tree structures and their decorations which are encoded using a special modal logic called loft—logic of finite trees (blackburn and meyer-viol, 1994). loft handles the usual relations on finite trees. underspecification and update are essential ingredients of ds. growth of information for tree structures and their decorations is conceived of as a monotonic process. the building up of tree structures is steered by the requirements of nodes: for example, an underspecified subject node in a tree would have to be equipped with the information that an expression of the type of individual is needed. requirements say which information must be got in order to yield a (more) complete structure. structures are built up by lexical or computational actions. computational actions include the introduction and the updating of structure. lexical actions introduce individual lexical items inducing nodes and decorations. partial trees grow incrementally, driven by the requirements of the words encountered. a pointer ◊ marks the state of the parsing process. complete individual trees correspond to predicate-argument structures. a tree adjunction operation is used for generating more complex structures via the fitting together of trees sharing one term. ‘importantly, adjunction as other forms of construction and update, can be employed to model how subsequent speakers may dynamically provide fragmentary extensions in response to the previous utterance’. (gargett et al. (2008), p. 45). ds uses structural and content under-specification, e.g. meta-variables for pronouns and subscripted placeholders for names both of which inform tree construction. meta-variables can be updated if the context yields an appropriate term for substitution. essential for the matters of this paper is the ds notion of context. a context in ds involves the storage of parse states, i.e. of the partial tree arrived at, the word sequence identified and the actions used. as a consequence, the context provides a tool for switching from parsing to generation. the switch from parsing to generation is needed in turn to model coordination of speakers in dialogue. completions, coordination and alignment in dialogue 65 central assumptions of ds concerning dialogue (i.e., what is called ‘the ds dialogue model’ in the some of the papers). parsing is the prime mechanism of ds, generation being defined upon it. generation and parsing are both seen as incremental processes. how is the parsing-generation relation modelled? whereas a hearer by needs builds up a succession of partial parse trees, a speaker is also equipped with a goal tree containing the information he wants to produce. the generation process is determined by the parsing process on the goal tree and a subsumption relation existing between the goal tree and the partial parse tree guaranteeing felicity of production, so to speak. the ds dialogue model takes account of the hearer’s parse tree as well as of the speaker’s goal and parse trees. as a consequence, hearer’s activities such as clarification requests, completions and acknowledgements can be handled in an on-line manner. above all, a hearer can start his generation process right from the parsed string resulting from the on-going production of the speaker. this technique is essential for completions and repairs as well as for the production of feed back indicators. ‘in particular, for split/joint utterances, this enables switch from hearer to speaker at any arbitrary point in the dialogue […]’ (gargett et al. (2008), p. 46). in sum, we see that we get the incremental tools to handle completions. in the next chapter we show how completions can be modelled in ds, starting with some remarks tying the ds methodology to general dialogue research. 7.3.2 completions in dynamic syntax: purver et al (2006) clark and wilkes-gibbs (1990) started from the hypothesis that in dialogue participants produce utterances cooperatively. this led to a sort of schema capturing presentation and mutual acceptance as a recursive process, described here informally and in an abridged version. clark and wilkes-gibbs’ idea was that e.g. in order to initiate a reference to, say, some tangram figure, an agent either presents a referring term x or invites one from the other agent. if a term is invited by the first agent, the other agent, being cooperative, presents one. if the presented term x is considered inadequate, then a revision x’ of x, an expansion y of x or a replacement z of x is provided. finally, a term x held adequate is accepted by both parties and mutual acceptance is inferred. observe that by the last step x enters the common ground. applying this device to example 1.2, we have a production so, jetzt nimmst du by inst, a presentation eine schraube by the other agent, cnst, most probably invited by inst’s lengthening. inst reacts with an expansion 1.3 eine orangene mit einem schlitz which gets mutually accepted, see cnst’s ja and inst’s continuation. in short, seen from the clark and wilkes-gibbs perspective, syntax production goes through three stages: presentation, repair by expansion and acceptance, as we saw in previous sections. as is clear from the clark and wilkes-gibbs approach, cooperatively produced utterances present a problem since they integrate production and parsing processes, normally treated as separate paradigms. recent extensions of dynamic syntax (ds, purver and kempson (2004), purver, cann, and kempson (2006)) have paved the way for modelling cooperative production of utterances by several agents in a theory of grammar framework achieving effects of the original clark and wilkes-gibbs proposal. purver et al. developed a tool which can toggle from parsing to generation states and vice versa and can this way also serve as a basis for modelling changes of speaker’s activity. at the heart of this account is the notion of a parser state p. p is a set of triples , t being a semantic tree under construction, w a sequence of words, and a a set of lexical or computational actions used for building up trees or singling out words. the notion of a parser state is then used for the definition of a context c of a particular tree t in the set p. c consists of: (a) a set of triples p’ = {…, , …} resulting from the previous sentence(s), and (b) the current triple . poesio and rieser 66 at the start of the parsing process p’ is empty and the context is identical to the current parser state p0 = . here t0 is the basic axiom with ?ty(t) a designated goal to be proved and ◊ serving as a pointer to the current stage of the parsing procedure. in p0 we have empty sequences for words and actions. defined in an analogous fashion, a generator state is a pair (tg, x). tg is a goal tree and x a set of pairs (s, p) with s standing for a candidate string and p for the associated parser state, a set of triples. at the beginning of the discourse x will consist of an empty candidate string and the standard initial parser state (0, p0). the context c for generation is defined as for parsing, as the set of triples p’ = {…, , …} and the current triple . purver et al. (2006) provide a model of shared utterances capturing tightly fitting completions and the change of roles among speaker and hearer35. the model is given along the following lines. the structural description of the first part so36, jetzt nimmst du which is incomplete by grammaticality standards can be used as an input for the parsing or the generation of the completion eine schraube. both productions together add up to a well-formed structure as in our example (1.1). since one can switch from parsing to generation and vice versa, at least a necessary condition for role switches is met by means of the respective contributions. the normal generation process starts with a generator state (tg, {0, p0}), where tg is the goal tree, 0 the place holder for the candidate string and p0 the initial parser state . as mentioned above, t0 stands for the basic axiom . on condition that a suitable goal tree, basically representing the meaning structure of the candidate string, is available for generation, a continuation can be produced using the structure arrived at. this is achieved replacing p0 with the parser state (a set of triples ), pt got from the structure produced, called transition state. now, assume speaker role a and the corresponding hearer role b. hearer b parses up to pt. if he has a suitable goal tree tg, he can set up a transition generator state gt = (tg,{(0, pt)}) and continue. thus gt can be directly used for generation. in order that this may happen, a monotonicity constraint has to be satisfied: the goal tree tg must be subsumed by one of the partial trees in pt. considering both sides, a and b, we have the following: at the point of transition a’s generator state gt’ contains the pair (st, pt’). st is the string produced so far and pt’ the corresponding parser state, the transition state for b. a is assumed to interpret b’s continuation in the context of his parser state pt’, into which the structures extracted from b’s productions will be integrated. by way of illustration cnst’s (= b’s) transition from hearer to speaker role is represented in figure 7.3.1. pt depicts cnst’s parsing procedure, the ? indicating that an object ty(e) is expected due to the type ty(e, (e, t)) of nehmen. jetzt is treated as an unfixed sentence adverb. cnst recognized the words [so,] jetzt, nimmst, du using actions a1, a2, and a3. he has a suitable goal tree to set up a transition generator state gt. the goal tree shows that the indexical must be changed. hence we have the logical form fo(jetzt’((nehmen’(ε, x, schraube’(x)))(fo(cnst)))). the slot for the string is empty, and this is what one expects, given the example: the string produced is eine schraube, it is fitted in as an object-np, two new actions a4, and a5 are needed to achieve that. 35 in discussing the completion in the ds framework based on (purver and kempson ( 2004)) and (purver et al. (2005)) we got help from matthew purver and ruth kempson, which we gratefully acknowledge. any remaining mistakes are ours, of course. 36 for so we assume a functional interpretation along the lines of the discussion in 5.2; so will hence not turn up in the ds trees any more. completions, coordination and alignment in dialogue 67 figure 7.3.1. modelling of completion jetzt nimmst du eine schraube involving role shift from hearer to speaker (simplified). poesio and rieser 68 figure 7.3.2. modelling of completion jetzt nimmst du eine schraube involving role shift from hearer to speaker (simplified). completions, coordination and alignment in dialogue 69 figures 7.3.1 and 7.3.2 only illustrate the completion part. there is also a suggestion concerning the description of appositions in (cann, kempson, marten (2005), ch. 7, p. 328, ch. 8, 8.4.1 and 8.67, p. 365), which could be used in order to graft the instructor’s eine orangene mit einem schlitz onto schraube. the apposition functions as an other-repair to cnst’s contribution. so far, repairs have not been treated in ds in a systematic way. the most explicit reconstruction of corrections known to us is in sdrt (asher and lascarides (2003), 345-373), which presupposes discourse relations but is not tied to incremental surface syntax. what is the difference between the ds model of completions and our account? the difference can essentially be traced back to the philosophy of dynamic syntax. by joining parsing and production ds provides a necessary condition for cnst’s production of some object np but this will not necessarily be eine schraube, it could be anything else, say a car, a black hole etc. the targeted production of language material is clearly plan-based. so, in order to make ds do that it would have to be tied to some theory of dialogue via a suitable interface. 8 conclusions theorists from clark to ginzburg to pickering and garrod all agree that perhaps the key test for a theory of dialogue is the extent to which it accounts for what is arguably the key difference between dialogue and monologue—the process by which the conversational participants achieve their communication goals by coordinating with each other--without assuming that this communication is always successful. this requirement on information-state theories has perhaps been best captured by ginzburg in the following passage (ginzburg 2009, p. 121): … one could argue that the basic criterion of adequacy of a semantic theory of dialogue is the ability to characterize for any utterance type the update that emerges in the aftermath of successful grounding … this is the early 21st century analogue of truth conditions to our knowledge, what was presented in this paper is the first formal account of a crucial aspect of this process of coordination in dialogue: the fact that coordination can be achieved at the level of the single utterance. most ideas have been developed elsewhere: our account builds, on the one hand, on clark’s proposals about coordination, as formalized by bratman and tuomela for shared plans, and by traum for grounding; on the other, on ideas about the dynamics of the common ground developed by stalnaker, lewis, and kamp; last but not least, on the view of the common ground developed in situation semantics by barwise, perry, cooper, and colleagues. as far as we know, however, this is the first attempt to put all these ideas together in a single account. this being the case, it is almost inevitable that our account is not going to be the ultimate word on the argument. we mentioned throughout the text a number of limitations of the present proposal (e.g., the need to integrate these ideas with the kind of statistical models of interpretation that are now becoming prevalent both in computational linguistics and in psycholinguistics) and many issues to be addressed in further research, ranging from empirical questions such as the proper formulation of the grounding criterion to formal questions concerning the consistency of the underlying logic. it’s virtually certain that many aspects of the theory will need to be revisited once these issues are studied in detail at the light of more empirical data. however, we believe that our approach of relying whenever possible on independently motivated tools and hypotheses such as ltag, compositional drt, prioritized default logic, the clark / traum view of grounding, the bratman / tuomela view of intentions, etc—are the best we could do to ensure that our proposals will scale up. some issues left open in this paper will be addressed in a followup paper (poesio and rieser, in preparation) that will extend the analysis to the phenomena poesio and rieser 70 observed in the second part of the example dialogue, in particular, providing an incremental account of anaphora. we remain noncommittal concerning whether an approach based on alignment and in which all types of intentional reasoning are avoided could provide an explanation of completions: a more detailed formulation of the alignment account will be required before that question can be properly addressed. we believe that some types of completions could be explained in this way: for instance, as one of our reviewers pointed out, the production of the second completion in the example dialogue, 2.2 (von oben) could be explained without reference to shared plans: cnst could simply have examined the state of her own knowledge, decided that she needed more information in order to execute the action, and performed a request. (there is ample evidence that conversational participants often decide how to act next on the basis of their own knowledge only–see, e.g., (bard et al, 2000).) we indicated some ways in which the alignment approach would need to be further developed in order to understand whether such explanations could work. conversely, however, it could be argued that in task-oriented dialogues the only ‘situation models’ that participants could attempt to align are their plans, and thus that alignment of representations would not eliminate the need for some sort of intentional representation and intention recognition. acknowledgements among the many people who helped us with comments and suggestions in the five years it took us to complete this paper we particularly wish to thank jonathan ginzburg and enric vallduvi (who, besides giving us lots of constructive suggestions through the years, invited us to present these ideas at an invited talk at catalog in barcelona in 2004); ruth kempson, with whom we had many useful discussions particularly about appositions and about the relationship between dynamic syntax and ptt, and who as coordinator of the leverhulme dialogue matters network gave us many opportunities to exchange our views with other colleagues working on similar ideas; david traum, who was instrumental in the development of ptt and gave us many suggestions about grounding; and jens allwood, herbert clark, raquel fernandez, pat healey, matt purver, david schlangen, ipke wachsmuth, for many many comments and discussions over the years. funding for massimo’s visits to bielefeld and hannes’ visits to essex and then trento was provided by sfb 360, “situated artificial communicators”, sfb 673, “alignment in communication”, the university of essex’s computer science department (now school of computer science and electronic engineering) and the university of trento ‘s center for mind / brain sciences. completions, coordination and alignment in dialogue 71 appendix a: full transcript of the dialogue up to the excerpt versuchspersonen-paar 21 instrukteur (i) m; konstrukteur (k) f bedingung: sicht blockiert, vorlage modell, beschre ibung/konstruktion simultan gesamtdauer 12: 45 anmerkung: i erwähnt 'flugzeug' zunächst nicht; k k ommt von selbst drauf. konstruktionsresultat i: aa=10.3ab=08.1ba=10.1bb=08.2c=02.1da=12.2db=01.1dc= 08.3ea=12.1 eb=01.2ec=07 ed=05 ee=13 ha=11.3hb=02.2hc=04 hd=03. 1he=11.1 la=14.1lb=15.1lc=09.1ld=16.1le=17.1ra=14.2rb=15.2rc =09.2rd=16.2 re=17.2va=11.2vb=03.2vc=06vd=03.3ve=03.4vf=10.2mutt ern=vl konstruktionsresultat k: aa=10.3ab=08.1ba=10.1bb=08.2c=02.1da=12.2db=01.1dc= 08.3ea=12.1 eb=01.2ec=07 ed=05 ee=13 ha=11.3hb=02.2hc=04 hd=03. 1he=11.1 la=14.1lb=15.1lc=09.2ld=16.1le=17.1ra=14.2rb=15.2rc =09.1rd=16.2 re=17.2va=11.2vb=03.2vc=06vd=03.3ve=03.4vf=10.2mutt ern=vr montiert hb parallel zu c bugwärts zeigend. 21i001 mit den sieben löchern. 21k001 ja. 21i002 da hast du zwei stück von 21k002 ja. 21i003 die kannst du erstmal zur seite lege n. dann ist eins mit eins, zwei, drei, vier, fünf löchern 21k003 {ja} 21i004 und eins mit d rei löchern 21k004 {ja} 21i005 so, <--> jetzt nimm das äh mit den f ünf löchern mal in die linke hand 21k005 mhm. 21i006 das andere in die rechte < -> und leg(e) sie so übereinander, daß sich zwei löcher überschneiden. 21k006 also die zwei mit den fünf löchern? 21i007 ja. 21k007 das sie sich wo 21i008 das kleine ist dann untendrunter so 21k008 hä? also das mit den drei löchern? 21i009 ja. 21k009 untendrunter. 21i010 ja. oh, das ist aber schwierig <-> jetzt überschneiden sich ja zwei löcher, ne? 21k010 zwei? 21i011 bei dir auch? 21k011 wie wann? <-> ich soll das in der mitte zusammenlegen überkreuzen 21i012 nein, nicht in der mitte so, daß es also wieder quasi eine <--> eine (ei)n teil so 21k012 ah ja, verlängerung mit zwei löcher. 21i013 mhm. 21k013 überschneiden. 21i014 also wenn du jetzt von oben draufguckst, has t du sechs löcher. 21k014 ja. 21i015 ja? 21k015 mhm. poesio and rieser 72 21i016 zwei zwei überschneiden sich, ist das richti g? 21k016 ja. 21i017 so und jetzt nimmst du die beiden <--> m die mit den äh <-> sieben {löchern} ja? 21k017 ja. 21i018 die die legst du so drauf <-> auf diese beid en, die sich überschneiden, daß sich das (ei)n kreu/, daß es sich ein kreuz erg ibt so. <-> weißt du? <-> legst du auf den tisch am besten. 21k018 das mit den drei löchern immer noch unten und dann leg(e) ich was? 21i019 ja. 21k019 (ei)n kreuz. 21i020 ja, daß es so jetzt überschneiden s ich quasi so die ähm eben haben sich ja nur die <--> beiden überschni tten, ja? 21k020 ja. 21i021 und jetzt überschneiden die sich dreimal. <> verstehst du? 21k021 nee. 21i022 m, dann machen wir das mal anders, warte mal <--> m, leg(e) das mal so auf den tisch mit den, daß die sich die bei, wie wir es eben hatten, daß wir sechs löch er so haben. 21k022 ja. 21i023 legst du so auf den tisch. 21k023 hab(e) ich schon. 21i024 und da ist noch ein anderes mit dr/, nimmst du noch ein anderes mit drei löchern, 21k024 {mhm} 21i025 und das schiebst du vor <-> von link s <--> drunter, daß sich auf der anderen seite 21k025 auch nur das zwei überschn/ 21i026 zwei überschneiden. 21k026 ah ja. 21i027 ja? 21k027 dann habe ich sieben löcher. 21i028 ja, ja, so können wir es liegen lassen. 21k028 mhm. 21i029 jetzt nimmst du so einen roten würfel. 21k029 was war mit dem {sieben} löchern, den wir ja gut. 21i030 ja? 21k030 ja. 21i031 und den legst du ganz links <--> drauf. 21k031 ganz links? 21i032 ja, auf das äußerste loch. 21k032 mhm. 21i033 so <-> jetzt nimmst du <-> eine <-> rote sch raube 21k033 mhm. 21i034 und schraubst es von unten fest. kannst du s o, mußt du so hochnehmen dafür. 21k034 (ei)ne kleine rote? 21i035 ja. 21k035 eckige? 21i036 ja, ja. 21k036 gut. <-> dann ist jetzt nur der drei er fest <--> erstmal 21i037 ja, genau, mhm. <-> kannst es aber wieder so hinlegen. ja, paß auf und dann nimmst du jetzt erstmal noch so(_ei)n teil mit fünf löchern. 21k037 ja. 21i038 und <-> schraubst es oben auf den würfel dra uf mit einer runden schraube mit (ei)nem schlitz drin. 21k038 mit der roten runden? 21i039 ja. hast du? 21k039 ja. completions, coordination and alignment in dialogue 73 21i040 so, jetzt nimmst du <-> (ei)ne gelbe runde 21k040 mhm. 21i041 und so(_ei)n ähm orangenes teil, so (_ei)n gewinde, mit gewinde. 21k041 ja. 21i042 ja? und schraubst das äh <-> dreier, die dre ierstange an die fünfer. 21k042 also so wie gehabt ne, diese untere? oder was? 21i043 ja, ja, also die dreier wird von unten an di e fünfer geschraubt. 21k043 mhm. 21i044 verstehst du? 21k044 ja. und wie rum ich das mache ist e gal, von unten die gelbe rein? 21i045 nee, von oben die gelbe. 21k045 und wo <-> wo von den beiden? ich h abe ja zwei löcher, wo ich es reinstecken kann. 21i046 also die gelbe schraube ist jetzt <--> rechts von dem roten würfel. 21k046 direkt dran. 21i047 oben. ja, ja. 21k047 ja. 21i048 und das ist quasi in dem ersten loch von dem fünfer, ne 21k048 mhm. 21i049 in dem ganz links äußersten und mittig 21k049 und dem mittleren vom dreier 21i050 mittig von den dreien. genau. und jetzt nimm st du die <-> eine eckige gelbe schraube und machst sie in das rechte <-> loc h daneben. 21k050 ah ja. 21i051 und unten wieder so(_ei)ne mutter vor. 21k051 mhm. 21i052 na. jetzt wird es schwieriger. <-> paß auf, jetzt nimmst du <--> jetzt legst du den anderen, dies dreierteil, ne dies ande re 21k052 ja. 21i053 von rechts, also das ist genau symmetrisch u nd mit dem <-> mit dem linken dreier, was du schon festgeschraubt hast. 21k053 ja. 21i054 so <--> jetzt nimmst du eine lange stange mi t sieben löchern drin, 21k054 ja. 21i055 die legst du jetzt auf <--> die fünfer. 21k055 auf die {obere} auf die untere? auf die obere? 21i056 quasi {ja} < sil: 4> wart(e) ähm m 21k056 was war denn jetzt mit der unteren dreier, d as habe ich noch nicht. 21i057 die kommt dadrunter, also legst du einfach, hältst du mal fest so 21k057 ja. 21i058 wenn die jetzt festgeschraubt wär(e), wär(e) die genauso wie die andere. 21k058 ja. 21i059 ja? 21k059 das ist klar. 21i060 jetzt nimmst du in die linke ha/, also das h ast du jetzt in der rechten hand, halt das mal so fest, oder legst es auf den tisch oder irgendwie so. 21k060 ja 21i061 ist jetzt ähm, legst du <--> das teil mit den <--> sieben löchern 21k061 ja. 21i062 so mhm so dadrauf, warte m al, wie beschreib(e) ich dir das jetzt, mhm 21k062 der obere oder der untere fünfer, wo soll ic h es hintun jetzt? 21i063 oben auf den <-> auf den fünfer. poesio and rieser 74 21k063 ja. 21i064 also nicht der auf den roten würfel ist. da ist ja auch einer mit fünf löchern auf dem roten würfel. 21k064 ja, also auf dem {unteren} fünf er. 21i065 wo links die beiden gelben schrauben drin si nd. 21k065 ja. 21i066 dann kommt eine, erst ein loch, das läßt du frei 21k066 mhm. 21i067 und dann kommt (ei)n loch, <-> wo du es draufpackst 21k067 ach so, also zwei 21i068 da legst du es drauf 21k068 überschneiden sich wieder. 21i069 genau, die überschneiden sich, genau wie (ei )n kreuz legst du es drauf und zwar mittig, ne 21k069 wie (ei)n kreuz nicht. 21i070 doch <-> ja so, was heißt (ei)n kreuz, eine kreuzung, wie eine kreuzung. 21k070 hä, welches überschneidet sich denn dann? 21i071 unten dieses ähm mit den fünf löchern, wo di e zwei gelben schrauben drin sind. 21k071 das zweite von rechts bei dem? 21i072 ja, genau, zweite von rechts. 21k072 und dann leg(e) ich das mittlerste von dem s iebener drüber. 21i073 genau 21k073 ah ja. 21i074 so machst du es und 21k074 ah ja gut 21i075 dadrunter liegt wiederum dies mit dem <-> dr ei löchern. 21k075 genau, das äußerste linke loch vom dreier. 21i076 {mhm} also die beiden von dem d reier, der da noch drunterliegt äh, da schneiden jetzt zwei <-> ja? 21k076 ja, das ist klar, das haben wir. completions, coordination and alignment in dialogue 75 appendix b: a grammar fragment for the example dialogue in this appendix we introduce the elementary trees associated with the words used in the fragment and discuss some of the more complex linguistic phenomena that occur in it, abstracting away from micro-conversational events. our ltag treatment is based on (abeillé and rambow, 2000; joshi 2004) and on frank’s view that elementary trees represent extended projections (frank, 2002). our semantics is based on compositional drt as presented in (muskens, 1996). some ideas from muskens’ (2001) logical description grammar, an integration of ltag with cdrt, were also taken into account. b.1. a brief overview of compositional drt the aim of muskens in developing compositional drt (muskens, 1996) was to develop a framework in which the dynamics of drt could be combined with compositional methods of meaning specification without developing an entirely new logic as done, e.g., in heim’s file change semantics (heim 1983) or groenendijk and stokhof’s dynamic predicate logic (groenendijk and stokhof, 1989). muskens uses the term grafting to describe his approach: extend a type logic like those standardly used in montagovian approaches with the types necessary to provide a reconstruction of the drt constructs and an (axiomatic) formalization of their properties. specifically, the logic proposed by muskens includes two new types: the type of discourse referents π and the type of states s. discourse referents are used to model the dynamics of context in the same way as they are used in drt, i.e., in the sense that each noun phrase introduces a new discourse referent. states are used to model contexts themselves, and the way they are modified by natural language sentences; they are the object-language equivalent of the assignments used to formalize the semantics of drss in (kamp and reyle, 1993). this dynamics is mediated by drss, viewed as relations between states. a function v: π � (s � e) provides the mapping from discourse referents and states to entities, in the sense that v(x)(i) specifies the `value’ of discourse referent x at state i. muskens specifies a translation for all drs constructs in terms of this type logic. the most important translations are shown in (b.1.1). (b.1.1) a. r{x1 … xn} is short for λi. r(v(x1)(i), … v(xn)(i)) b. x1 is x2 is short for λi. v(x1)(i) = v(x2)(i) c. [x1…xn| ϕ1 … ϕm] is defined as λiλj i [x 1 ... xn] j ∧ ϕ1(j) … ∧ ϕm(j) where i [x1...xn] j is short for i and j differ at most over [x1 … xn]. d. k;k’ is defined as λiλj ∃ k k(i)(k) ∧ k’(k)(j) for example, the type logic translation of the drs in (3.1.2) (repeated below as (b.1.2)(a)) is shown in (b.1.2)(b). (b.1.2) a. [x,w,y,z,s,s’| engine(x), avon(w), s: at(x,w), boxcar(y), s’:hooked-to(z,y), z is x] b. λiλj i [x,w,y,z,s,s’] j ∧ [engine(x)](j) ∧ [avon(w)](j) ∧ [s: at(x,w)](j) ∧ [boxcar(y)](j) ∧ [s’:hooked-to(z,y)](j) ∧ v(z)(j) = v(x)(j) in compositional drt, drss are the translations of sentences—the propositions. hence, type > is the type of propositions, and most translations of nl expressions are chains of functions whose ultimate value is an object of this type: e.g., common nouns have translations of type <π, >>. muskens introduces an abbreviated notation in which [t1 … tn] stands for >>>>: in this notation, [] stands for the type > of drss, [π] is the type <π, >> of common nouns, [[π][π]] is the type <<π, >>, <π, >>> of determiners, [[π]] is the type <<π, >>, >> of quantifiers, etc. we will use this abbreviated notation in what follows. poesio and rieser 76 the most important change to muskens’ semantics we make here is having two types of discourse referents and two v functions. although in the main body of the paper we only referred to a single v, in fact in addition to muskens’ referents of type π, to which a function v assigns values of type e, we also use referents of type πk, and a second function vk assigning to these referents values of type []. this of course requires a rather different semantics in which assignments have a higher type; see (poesio and muskens, 1997) for some discussion. b.2. ltag: basic trees b.2.1 nominal phrases as standard in ltag, we assume that nominal phrases (nps) are the projection of determiners rather than of nouns. the elementary trees for nouns do not therefore include a projection. we also assume the standard semantic treatment of nouns in cdrt, where nouns are translated as functions of type [π] (i.e., <π, >> see b.1). schraube, schlitz: n+ schraube: λx([ |screw(x)]) our treatment of determiners is also derived without modification from the analysis of determiners in ltag and compositional drt. although determiners are not viewed in ltag as heads of separate determiner phrases, but as heads of nps, the elementary trees associated with determiners do introduce an expectation of an n’, to which the elementary trees associated with nouns can be attached by substitution. semantically, determiners are given the standard cdrt type [[π],[π]]. both indefinite and definite determiners introduce new discourse referents, but the latter have a presuppositional component. eine: np+ det n’ eine: λp’λp([x| ]; p’(x); p(x)) einem: npdat + detdat n’dat einem: λp’λp([x| ]; p’(x); p(x)) we assume a `loebnerian’ treatment of definites, according to which definites have a uniqueness presupposition (see (poesio and rieser, in preparation)): die: npdat + detdat n’dat die: λp. λp’. ([y|y = ι x. p(x) ]; p ‘ (y); indexical and anaphoric pronouns are also analyzed as heads of nps. completions, coordination and alignment in dialogue 77 du: np2nd + du: λp.p (you) syntactically, adjectives in ltag introduce auxiliary (adjunct) trees, incorporated into nps by adjunction. semantically, they are treated as functions from predicates to predicates, of type [[π], π] (<<π, >>, <π, >>>). orange n+ adjp n orange: λp[π] λv([ |orange (v)]; p(v)) in german, adjectives can be nominalized, as in the case of orangene in 1.3. we treat these nominalized adjectives syntactically and semantically as nouns. b.2.2 prepositional phrases as discussed in section 3.4, prepositional phrases are treated in ltag as adjunct trees, incorporated into nps and vps by adjunction. semantically, prepositions are assumed to be functions from quantifiers to predicate modifiers, i.e., to have type [[[π]π]π]. mit n+ n ppdat pdat npdat mit: λp [[π],[π]] λx(p(λy[ | has (x,y)])) b.2.3 verbal phrases and sentences in ltag, sentences are viewed as projections of verbs, just as in hpsg. we’ll discuss the lexical interpretation of verbs in the example dialogue later, as they both require some discussion of aspects of german grammar. adverbs are interpreted as vp and s modifiers; we show here the elementary (auxiliary) tree involved in the interpretation of jetzt discussed in section 5.2. jetzt s+ advpinv sinv jetzt: λp. now(p) poesio and rieser 78 b.3. constructions occurring in the first turn: 1.1-1.3 the first turn of the example dialogue contains two linguistic phenomena whose treatment needs some discussion: inversion (the fact that the subject of nimmst occurs after the verb) and apposition. we discuss each in turn. b.3.1 inversion our syntactic treatment of subject-verb inversion is based on the standard treatment of inversions in german adopted, e.g., by most papers in (freidin, 1992), and is specifically modelled after the proposal by demena travis (1992). the elementary tree for the verb nimmst shown below includes an empty element in subject position, coindexed with the postverbal np. nimmst sinv vpinv + np v’ vfin2nd np2ndi + np+ εi nimmst: λxλq(q(λx’[e| e: grasp(x,x’)])) b.3.2 apposition the utterance of a noun phrase in 1.3 would generally be classified as an apposition, a form of parenthetical expressed with a noun phrase. appositives are extremely common in dialogue, reflecting the fact that utterance generation is incremental, as well; hence an account of this construction is essential for a theory like ptt. providing a unified account is however quite difficult, as nominal parentheticals have a wide range of forms and uses in addition to those illustrated by the example dialogue, as shown by the examples in (b.3.1). (see meyer (1992) for an extensive, if purely descriptive, investigation of appositions.) (b.3.1) westinghouse electric corp. is wooing michael h. jordan, a. former head of pepsico's international operations / 44 / from lawrenceville, nj b. the former head of pepsico’s international operations / a 36-year company veteran to be its new chief executive officer. as accounting for all these types of nominal parentheticals would be quite a challenge in itself, we only concentrate here on the class of appositions illustrated in the example dialogue, in which the apposition plays a restrictive role. our treatment of such constructions is consistent with the central claim of meyer (1992) that syntactically and semantically, such appositions are a type of nominal modification; we formalize this claim by proposing that semantically appositions are of type [[π] [π]], like pps and other nominal modifiers. (a reminder that nominal predicates, the type of montague grammar, in cdrt are functions from discourse referents to drss, <π, > >, abbreviated [π].) our main objective here is to explain how the apposition is incrementally integrated with the interpretation of the previous utterances, both from a syntactic perspective and from a compositional semantic point of view. this need to explain how appositions are incorporated from a syntactic perspective is one of our main reasons for adopting a syntactic formalism like ltag which allows for adjunction operations. the basic idea of our proposed ltag derivation for eine schraube … eine orangene is that the final interpretation is obtained by adjoining tree (b) for the apposition to tree (a) for the completions, coordination and alignment in dialogue 79 np, as shown schematically in (b.3.2): we assume that the syntactic interpretation for 1.3 (shown in (b)—we assume orangene is an adjective nominalization, as discussed below) is adjoined to the syntactic interpretation for 1.2 (in (a)), obtaining (c). (the pp mit einem schlitz is omitted for brevity.) (b.3.2) (a) np (b) n’ (c) np eine n’ n’ npapp eine n’ schraube eineapp n’ n’ npapp nadj schraube eineapp nadj as said above, we follow meyer in hypothesizing that the class of appositions we are concerned with semantically behave as nominal modifiers, of type [[π], [π]]. this still leaves open the question of how certain types of nps –particularly indefinites and definite nps—can serve as predicate modifiers. we take these to be cases of nps whose type has been lowered to that of predicates (farkas & de swart, 2003), and exploit the properties of muskens’ compositional drt to explain how the restriction expressed by the apposition is grafted onto the appropriate discourse entity. however, the implementation of this intuition is subject to a few constraints determined by ltag. ltag restricts the possible semantic analyses of nps used predicatively, because only lexical categories are allowed, and the only semantic operation which is allowed is application. such restrictions prevent treatments in which the predicative meaning of nps in appositions can be derived syncategorematically –say, by stipulating a structure as in figure b.4.1, with a semantic operation associated with npapp that `lowers’ the np-type interpretation for eine orangene, mit einem schlitz to one of type [[π], [π]]. in ltag, one has to stipulate instead that german eine is ambiguous between a normal determiner interpretation, shown in (b.3.3a), and the interpretation used in appositions, of type [[π], [[π], [π]]] (the cdrt translation of <, <, >>) shown in (b.3.3b). (b.3.3) a. eine: λp’λp([x| ]; p’(x); p(x)) b. eineapp: λp’λpλx([v| ]; p(x); p’(v); [ | v is x]) np ([[π]]) nbar ([π]) nbar ([π]) npapp ([[π] π]) λp λ u p (λ x [| u is x]) np ([[π]]) figure b.4.1. a syncategorematic treatment of appositions. poesio and rieser 80 the two interpretations are also associated with the distinct elementary trees in (b.3.4a) and (b.3.4b): (b.3.4) a. np+ det n eine: λp’λp([x| ]; p’(x); p(x)) b. n+ n npapp detapp nadj + eineapp: λp’λpλx([v| ]; p(x); p’(x); [ | v is x]) with these assumptions, and assuming standard elementary trees and semantic interpretations for the other words in utterance 1.3, we get the ltag analysis for eine schraube, eine orangene mit einem schlitz in (b.3.5). (b.3.5) np n’ n’ npapp n’ adj ppdat npdat det n detapp nadj pdat det n eine schraube eineapp orangene mit einem schlitz the interpretation of the np in (b.3.5) is arrived at through the derivation in (b.3.6). (b.3.6) eineapp: λp’λpλx([y| ]; p(x); p’(y); [ | y is x]) orangene: λp λy([ |orange(y)]; p(y)) einem: λp’λp([x| ]; p’(x); p(x)) schlitz: λy([ | slit(y)]) einem(schlitz): λp’λp([x| ]; p’(x); p(x))( λy([ | slit (y)] )) = completions, coordination and alignment in dialogue 81 λp([x| ]; [ | slit(x)] ; p(x)) mit: λpλx(p(λy[ | has(x,y)])) mit’(einem’(schlitz’)) : λpλx(p(λy[ | has(x,y)]))(λp([z| ]; [ | slit(z)] ; p(z))) = λx(λp([z| ]; [ |slit(z)]; p(z))(λy[ |has(x,y)])) = λx([z| ]; [ |slit(z)]; [ |has(x,z)]) /*renaming of variables = λx([y| ]; [ |slit(y)]; [ |has(x,y)]) orangene’(mit’(einem’(schlitz’))) : λp λz([ |orange(z)]; p(z))( λx([y| ]; [ |slit(y)]; [ |has(x,y)])) λz([ |orange(z)]; ([y| ]; [ |slit(y)]; [ |has(z,y)])) eineapp (orangene’(mit’(einem’(schlitz’)))): λpλx([z| ]; p(x); [ |orange(z)]; ([y| ]; [ |slit(y)]; [ |has(z,y)]); [ | x is z]) schraube’(eineapp((orangene’)(mit’(einem’(schlitz’))))): λx([z| ]; [ |screw(x)]; [ |orange(z)]; x = z; [y| ]; [ |slit(y)]; [ |has(x,y)]) eine’: λp’λp([x| ]; p’(x); p(x)) eine’(schraube’(eineapp((orangene’)(mit’(einem’(schlitz’)))))): λp’λp([x| ]; p’(x); p(x))( λz ([w| ]; [ |screw(z)]; [ |orange(w)]; [ | w is z]; [y| ]; [ |slit(y)] ; [ |has(z,y)] )) = λp([x| ]; [v| ]; [ |screw(x)]; [ |orange(v)]; [ | v is x]; [y| ]; [ |slit(y)]; [ |has(x,y)]; p(x)) we assume this treatment here. we also note, however, that a second solution is also possible within ptt by relaxing the strict restrictions on semantic operations that operate in ltag, and allow for type shifting –specifically, if we allowed the interpretation of nps to be shifted by applying to it the following operator: ntnm (np-to-noun-modifier) : λp λp λx(p(λy[ | x is y]) ; p(x)) which turns an np (type <, t>) into a noun modifier (type <, >). then it would be possible to derive the meaning of eine schraube, eine orangene as follows: (b.3.6) np λp [z| ]; [x| ]; [ |orange(x)]; [ |screw(x)]; [ | x is z]; p(z) eine n’ ( ntnm (λp([x| ]; [ |orange(x)]; p(x))) ( λy( [ |screw(y)]) = λ y ([x| ]; [ |orange(x)]; [ |screw(x)]; [ | x is y]) n’ np: λp([x| ]; [ |orange(x)]; p(x)) schraube eine nadj λp’λp([x| ]; p’(x); p(x)) λv[ |orange(v)] b.3.3 utterances with a dialogue control function many utterances, particularly non-sentential ones, do not contribute to the specification of the content of a contribution (core speech act), or make contributions both the ‘official business’ of the contribution and to its ‘collateral structure’ (clark, 1996) (see section 2). the first turn of the example dialogue contains two such utterances: so in 1.1 and ja in 1.4. the assumption in ptt is that at least some of these interpretations are contained in the lexicon, as well (which, we recall, is a default theory (poesio, 1995; poesio, to appear)) so that they are immediately accessed as the utterance is perceived and may be overridden by other interpretations. turn-taking dialogue acts as already discussed in section 3, the primary function of so in 1.1 appears to be of a keep-turn turn-taking act (traum & hinkelman, 1992; traum, 1994). this interpretation is encoded in the lexical update in (3.2.2), repeated here as (b.3.7). poesio and rieser 82 (b.3.7) so→ [u,ce|u: utter (a,”so”), ce: take-turn (a), generate(u,ce)] the repertoire of turn-taking acts in ptt includes the four acts take-turn , keep-turn, release-turn and assign-turn. as said in section 3.4, many other non-sentential utterances also are assumed to be associated with lexical updates involving turn-taking acts: among them, well, now, okay, and also non-words such as filled pauses (umm and the like), all of which appear to admit of an interpretation as keep-turn. of course, it is important to keep in mind that these discourse markers are (a) extremely ambiguous, particularly so if prosodic information is ignored (in which case competing updates would be possible) and (b) often play more than one function. these two issues are particularly evident in the case of okay, an utterance of which, depending on the context, may result in one or all of the updates below: okay→ [ce|ce: keep-turn(a)] okay → [ce|ce: accept(a,ce’)] okay → [ce|ce: acknowledge(a,ce’)] (see below for a discussion of acknowledge; note also that `backward-looking’ interpretations of utterances like okay include implicit references to a previous speech act ce’, one of the arguments supporting our hypothesis that conversational events introduce discourse markers in the common ground). and of course it could also be argued that assuming the same turn-taking translation for well, now, and umm is another big simplification. grounding acts ja in 1.4 is an explicit acknowledgment. again, it is assumed in ptt that at least some such interpretations are lexically encoded. the interpretation used in 1.4 is shown in (b.3.8). (b.3.8) ja → [u, ce, du | u: utter (a,”ja”), ce: ack(a,du), generate(u,ce)] completions, coordination and alignment in dialogue 83 references abeillé, a. and o. rambow, (eds) (2000). tree adjoining grammars. csli, stanford, ca. aijmer, k. (1996). conversational routines in english. longman. allen, j. f., schubert, l. k. , heeman, p. , hwang, c.-h., light, m., miller, b. , poesio, m. and traum, d. r. (1995). the trains project: a case study in defining a conversational planning agent, journal of experimental and theoretical ai. 7:7-48. allen, j. and core, m. (1997). draft of damsl: dialog act markup in several layers. october 18. allwood, j., nivre, j. and ahlsén, e. (1993). on the semantics and pragmatics of linguistic feedback. journal of semantics, 9(1):1-26 asher, n. (1993). reference to abstract objects in discourse. kluwer. asher, n. and lascarides, a. (2003). the logic of conversation. cambridge university press. austin, j. l. (1962). how to do things with words. oxford university press. bara, b. g. and tirassa, m. (2000). neuropragmatics: brain and communication. brain and language, 71(1):10-14. bard, e. g., anderson, a. h., sotillo, c., aylett, m. doherty-sneddon, g., & newlands, a. (2000). controlling the intelligibility of referring expressions in dialogue. journal of memory and language, 42 (1):1-22. barwise, j. & perry, j. (1983). situations and attitudes. mit press. blackburn, s. and meyer-viol, w. (1994). linguistic, logic and finite trees. bulletin of interest group in pure and applied logic, 2(1): 3-29. branigan, h.p., pickering, m.j. & cleland, a.a. (2000). syntactic coordination in dialogue. cognition, 75(2):13-25. branigan, h.p., pickering, m.j., stewart, a.j. & mclean, j.f. (2000). syntactic priming in spoken production:linguistic and temporal interference. memory & cognition, 28(8):12971302. bratman, m. (1992). shared cooperative activity, the philosophical review 101(2):327-341. bratman, m. (1993). shared intention. ethics 104(1):97-113. brennan, s. e. (2005). how conversation is shaped by visual and spoken evidence. in j. trueswell & m. tanenhaus (eds.), approaches to studying world-situated language use: bridging the language-as-product and language-action traditions, 95-129. mit press. brewka, g. (1989). nonmonotonic reasoning: from theoretical foundation towards efficient computation. phd dissertation, university of hamburg. brewka, g. (1989). preferred subtheories: an extended logical framework for default reasoning. proc. of ijcai, pages 1043-1048. brewka, g. (1990). handling incomplete knowledge in artificial intelligence. in is/ki, 11-29. bunt, h. (1995). dialogue control functions and interaction design. in: r.j. beun, m. baker & m. reiner (eds.) dialogue in instruction. springer verlag, 1995, 197 – 214. cann, r., kempson, r., marten, l. (2005), the dynamics of language. elsevier. carletta, j., isard, a., isard, s., kowtko, j. c., doherty-sneddon, g. and anderson, a. h. (1997). the reliability of a dialogue structure coding scheme. computational linguistics, 23(1):13-32. carpenter, b. (1994). quantification and scoping: a deductive account. in proceedings of the 13th west coast conference on formal linguistics. csli. chater, n.j., pickering, m.j., & milward, d. (1995). what is incremental interpretation? in d. milward & p. sturt (eds.), edinburgh working papers in cognitive science 11: incremental interpretation (pp. 1-22). centre for cognitive science, university of edinburgh clark, h. h. (1992). arenas of language use. the univ. of chicago press and csli . clark, h. h. (1996). using language. cambridge university press. poesio and rieser 84 clark, h. h. and wilkes-gibbs, d. (1990). referring as a collaborative process. in cohen, ph. r., j. morgan, and m.e. pollack, eds., intentions in communication, mit press, 1990: 461-493. clark, h. h., and brennan, s. a. (1991). grounding in communication. in l. b. resnik, j. m. levine, and s. d. teasley (eds.), perspectives on socially shared cognition, apa books, 127149. clark, h. h., and schaefer, e. f. (1987). collaborating on contributions to conversations. language and cognitive processes, 2 (1):19-41. clark, h. h., and schaefer, e. f. (1989). contributing to discourse. cognitive science, 13:259294. clark, h. h. and wilkes-gibbs, d. (1986). referring as a collaborative process. cognition, 22:139 cohen, p. r. and levesque, h. j. (1990a). persistence, intention, and commitment. in cohen, p. et al. (eds.), intentions in communication, pages 34-69. cohen, p. r. and levesque, h. j. (1990b). rational interaction as a basis for communication. in cohen, p. et al. (eds.), intentions in communication, pages 221-225. cohen, p.r., morgan, j., and pollack, m.e. (eds.) (1990). intentions in communication, mit press. cole, p. (ed.) (1978). syntax and semantics, 9: pragmatics. academic press. cooper, r. (1992). a working person's guide to situation theory 1992. in topics in semantic interpretation, ed. by s. l. hansen and f. soerensen, samfundslitteratur, frederiksberg, denmark cooper, r. and poesio, m. (1994). situation theory. in the fracas consortium (eds.), fracas deliverable d8, centre for cognitive science, edinburgh. cooper, r. (1996). the role of situations in generalized quantifiers. in lappin, s. (ed), handbook of contemporary semantic theory, blackwell. cooper, r., ericsson, s., larsson, s. and lewin, i. (2003). an information state update approach to collaborative negotiation. in kühnlein, p., rieser, h. and zeevat, h. (eds.): perspectives on dialogue in the new milennium, john benjamins. pages 271-287. cooper, r. (2005). type theory with records. university of gothenburg, ms. coulthard, m. (1977). an introduction to discourse analysis. longman. dale, r. and reiter, e. (1995) computational interpretations of the gricean maxims in the generation of referring expressions. cognitive science, 19(2):233-263. demena travis, l. (1992) parameters of phrase structure and verb-second phenomena. in freidin, r. (ed.) principles and parameters in comparative grammar. mit press. 339-365. fagin, r., halpern j.y., moses, y., and vardi, m.y. (1995). reasoning about knowledge. mit press. farkas, d. and de swart, h. (2003). the semantics of incorporation. csli. ferguson, g., allen, j. f., and miller, b. (1996). "trains-95: towards a mixed-initiative planning assistant," in proceedings of the third conference on artificial intelligence planning systems (aips-96), edinburgh, scotland, pages 70-77. fernandez, r. (2006). non sentential utterances in dialogue: classification, resolution and use. phd dissertation, king’s college, london. ferreira, f., ferraro, v., and bailey, k. g. d. (2002). good-enough representations in language comprehension. current directions in psychological science, 11:11-15. ferreira, f., lau, e. f., and bailey, k. g. d. (2004). disfluencies, parsing, and tree-adjoining grammar. cognitive science, 28:721-749. fletcher, c. (1994). levels of representation in memory for discourse. in gernsbacher, m. a. (eds), handbook of psycholinguistics, 1st edition, academic press. pages 589-607. frank, r. (2002). phrase structure composition and syntactic dependencies. the mit press. frazier, l. (1987). sentence processing: a tutorial review. in: coltheart, m. (ed.), the psychology of reading. erlbaum. pages 559-586. completions, coordination and alignment in dialogue 85 gargett, a., gregoromichelaki, e., howes, ch., sato, y. (2008). dialogue-grammar correspondence in dynamic syntax. in ginzburg, j., healey, p., sato, y. (eds.), proceedings of the 12th workshop on the semantics and pragmatics of dialogue (londial ’08). pages 43-51. garnham, a. (2001). mental models and the interpretation of anaphora, psychology press. garrod, s. and anderson, a. (1987). saying what you mean in dialogue: a study in conceptual and semantic coordination. cognition 27: 181-218 geurts, b. (1995). presupposing. phd dissertation, university of stuttgart. ginzburg, j. (1997). on some semantic consequences of turn taking. in: p. dekker et al (eds.) proceedings of the 11th amsterdam colloquium. ginzburg, j. (2009). the interactive stance: meaning for conversation. king's college, london ms. (currently under review). ginzburg, j. (to appear). situation semantics: from indexicality to metacommunicative interaction. to appear in p. portner, c. maierborn, von heusinger (eds), the handbook of semantics, de gruyter. ginzburg, j. and cooper, r. (2004). clarification, ellipsis, and the nature of contextual updates. linguistics and philosophy, 27(3):297-366. ginzburg, j. and fernandez, r. (2005). scaling up to multilogue: some benchmarks and principles. proc. of the 43rd meeting of the acl, ann harbor, michigan. ginzburg, j. and, sag, i. (2000). interrogative investigations. csli. goldman, a. j. (1970). a theory of human action. prentice-hall. grice, h. p. (1969). utterer's meaning and intention. philosophical review, 68:147-177. grice, h. p. (1975). logic and conversation. in cole, p. and morgan, j. (eds.) syntax and semantics, pages 1-58. grice, h. p. (1978). further notes on logic and conversation. in cole, p. (ed.), syntax and semantics, pages 113-127. groenendijk, j. and m. stokhof (1991). dynamic predicate logic. linguistics and philosophy 14: 39-100. gross, d., allen, j. f. and traum, d. r. (1993). the trains 91 dialogues. university of rochester, department of computer science, trains technical note 92-1. grosz, b. j. and sidner, c. l. (1986). attention, intention, and the structure of discourse. computational linguistics 12:175-204 grosz, b. j. and sidner, c. l. (1990). plans for discourse. in cohen, p. et al. eds., 417-445. grosz, b. j. and kraus, s. (1996). collaborative plans for complex group action. artificial intelligence. 86:269-357. heeman, p. and allen, j. f. (1995). the trains 93 dialogues. university of rochester, department of computer science, trains technical note 94-2. hobbs, j. r., stickel, m., martin, p., and edwards, d. (1992). interpretation as abduction. artificial intelligence journal, 63 :69–142. homer, b. d. and tamis le monda, c.s. (eds.) (2005). the development of social cognition and communication. psychology press. hwang, c.h. and schubert, l. k. (1993) meeting the interlocking needs of lf-computation, deindexing, and inference: an organic approach to general nlu. in proc. 13th int. joint conf. on artificial intelligence. johnson-laird, p. n. (1983). mental models: towards a cognitive science of language, inference, and consciousness. harvard university press. joshi, a. k. (2004). 2003 rumelhart prize special cognitive science issue honoring aravind k. joshi, cognitive science, 28(5). jurafsky, d. (1996). a probabilistic model of lexical and syntactic access and disambiguation. cognitive science 20: 137–194. kamp, h. and reyle, u. (1993). from discourse to logic. kluwer. poesio and rieser 86 kautz, h. a. (1987). a formal theory of plan recognition. phd dissertation, university of rochester. kempson, r., meyer-viol, w. and gabbay, d. (2001) dynamic syntax. blackwell. kranstedt, a., lücking, a., pfeiffer, th., rieser, h., wachsmuth, i. (2006). deictic object reference in task-oriented dialogue. in: rickheit, g. and wachsmuth, i. (eds.), situated communication. mouton de gruyter. pages 156-209. kroch, a. s. and santorini, b. (1992). the derived constituent structure of the west germanic verb-raising construction. in freidin, r. (ed.), pages 269-339. larsson, s. and traum, d. r. (2000). information state and dialogue management in the trindi dialogue move engine toolkit. natural language engineering, special issue on best practice in spoken language dialogue systems engineering. 6:323-340. lerner, g. h. (1987). collaborative turn sequences: sentence construction and social interaction. ph. d. dissertation, university of california, irvine levelt, j.m. (1989) speaking: from intention to articulation. mit press. levin, j. a. and moore, j. a. (1977). dialogue games: meta-communication structures for natural language interaction. isi/rr-77-53, information sciences institute, usc. levinson, s., c. (1983). pragmatics. cambridge university press. litman, d., j. and allen, j. f., (1987). a plan recognition model for subdialogues in conversation. cognitive science 11: 163-200. litman, d. j. and allen, j. f., (1990). discourse processing and commonsense plans. in cohen, p. et al. (eds.), pages 365-389. mann, w. c. (1988). dialogue games: conventions of human interaction. argumentation 2 : 511532 matheson, c., poesio, m. and traum, d. r. (2000). modeling grounding and discourse obligations using update rules. in proc. of the naacl. mcneill, d. (1992). hand and mind: what gestures reveal about thought. university of chicago press. morton j. (1969). interaction of information in word recognition. psychological review. 76: 165–178 muskens, r. (1995). tense and the logic of change. in u. egli, e.p. pause, c. schwarze, a. von stechow, and g. wienold, editors, lexical knowledge in the organization of language. pages 147-183. john benjamins. muskens, r. (1996). combining montague semantics and discourse representation. linguistics and philosophy 19:143-186 muskens, r. (2001). talking about trees and truth-conditions. journal of logic, language and information, 10(4):417-455. neumann, g. (1998). interleaving natural language parsing and generation through uniform processing. artificial intelligence. 99:121-163. parsons, t. (1990). events in the semantics of english: a study in subatomic semantics. current studies in linguistics series, vol. 19. mit press. pereira, f. (1990). categorial semantics and scoping. computational linguistics 16(1):1-10. perrault, c. r. (1990). an application of default logic to speech-act theory. in cohen, p. et al (eds.), intentions in communication. pages 161–187. perrault, c. r. and allen, j. f. (1980). a plan-based analysis of indirect speech acts. american journal of computational linguistics 6 (3-4): 167-182. pickering, m. j. and garrod, s. (2004). toward a mechanistic psychology of dialogue. behavioral and brain sciences 27:169-226. platzack, c. (1986). comp, infl, and germanic word order. in hellan, l. and koch christensen (eds.) topics in scandinavian syntax. d. reidel. pages 185-235. poesio, m. (1993). definite descriptions and the dynamics of mental states. in proc. aaai spring symposium on reasoning about mental states. completions, coordination and alignment in dialogue 87 poesio, m. (1994). discourse interpretation and the scope of operators. phd dissertation, university of rochester. poesio, m. (1995). a model of conversation processing based on micro conversational events. in proceedings of the 17th annual meeting of the cognitive science society. poesio, m. (1996). underspecification and the semantics / pragmatics interface, in proceedings of the workshop on the cognitive science of natural language processing. poesio, m. (2001). what psycholinguistics tells us about the semantics / pragmatics interface: the case of pronouns, in proc. amsterdam colloquium. poesio, m. (to appear.) incrementality and underspecification in utterance interpretation. csli. poesio, m. and muskens, r. (1997). the dynamics of discourse situations. proc. of the 11th amsterdam colloquium, pages 247-252. poesio, m. and rieser, h. (in preparation). an incremental model of anaphora and reference resolution based on resource situations. ms. poesio, m. and traum, d. r. (1997). conversational actions and discourse situations. computational intelligence, 13(3) :1-45. poesio, m. and traum, d. r. (1998). towards an axiomatization of dialogue acts. in proc. of twendial, pages 1-7. poesio, m. and traum, d. r., eds., (2000). proceedings of götalog 2000. fourth workshop on the semantics and pragmatics of dialogue, göteburg univ. pollack, m. (1986), inferring domain plans in question-answering. phd dissertation, university of pennsylvania. poncin, k. and rieser, h. (2006). multi-speaker utterances and co-ordination in task-oriented dialogue. journal of pragmatics 38:718-744. portner, paul (2007). imperatives and modals. natural language semantics, 15(4):351-383 purver, m., cann, r. and kempson, r. (2006), grammars as parsers: meeting the dialogue challenge. research on language and computation, 4(2-3):289-326. purver, m. and kempson, r. (2004). incremental parsing, or incremental grammar? in proceedings of the acl workshop on incremental parsing, pages 74-81, barcelona, july 2004. purver, m. and ginzburg, j. (2004) clarifying noun phrase semantics. journal of semantics 21(3):283-339. purver, m. (2004). the theory and use of clarification in dialogue. phd dissertation, king’s college, london. reiter, r. (1980). a logic for default reasoning. artificial intelligence, 13:81-132. rickheit, g. & wachsmuth, i. (2008). alignment in communication collaborative research center 673 at bielefeld university. künstliche intelligenz, heft 2-2008, pages 62-65 rieser, h. (2004). pointing in dialogue. in ginzburg, j. and vallduvi, e. (eds.) catalog ’04 . proceedings of the eigth workshop on the semantics and pragmatics of dialogue. department of translation and philology, universitat pompeu fabra, barcelona. rieser, h. (2008). aligned iconic gesture in different strata of mm-route-description dialogue. in: proceedings of londial, pages 167-174 rieser, h., kopp, s. and wachsmuth, i. (2007). speech-gesture alignment. in: isgs abstracts, integrating gestures. northwestern university, evanston,chicago, pages 25-27 rieser, h. and skuplik, k. (2000). multi-speaker utterances and coordination in task-oriented dialogue. in poesio, m. and traum, d. r. (eds.), proceedings of gotalog, pages 143-151. sacks, harvey (1967). the search for help: no one to turn to, in e. s. shneidman (ed.), essays in self-destruction, science house. pages 21-27. sacks, h., schegloff, e. and jefferson, g. (1974). a simplest systematics for the organization of turn-taking for conversation. language, 50(4):696-735. sadek, m.d. (1992). a study in the logic of intention. in: b. nebel, ch. rich, w. swartout (eds.): principles of knowledge representation and reasoning. proceedings of the third international conference. morgan & kaufmann publishers: san mateo, ca, pages 462-475 poesio and rieser 88 sadek, d. (1994). communication theory = rationality principles + communicative act models. in proc. of the aaai 94 workshop on planning for interagent communication. sanford, a.j. and garrod, s. (1981). understanding written language. john wiley. schabes, y., abeillé, a., & joshi, a.k. (1988). new parsing strategies for tree adjoining grammars. in proceedings of the 12th international conference in computational linguistics, pages 578– 583. schegloff, e.a., jefferson, g., and sacks, h. (1977). the preference for self-correction in the organization of repair in conversation. language, 53(2):361-382. schlangen, d. and a. lascarides (2003). the interpretation of non-sentential utterances in dialogue, in proceedings of the 4th sigdial workshop on discourse and dialogue, sapporo, japan. searle, j. r. (1969). speech acts. cambridge university press. searle, j. r. and d. vanderveken (1985). foundations of illocutionary logic. cambridge university press. schabes y. (1990). computational and mathematical properties of lexicalized grammars, phd dissertation, univ. of pennsylvania, philadelphia. skuplik, k. (1999). satzkooperationen. definition und empirische untersuchung. bielefeld univ., report 1999/03 of sfb 360. stalnaker, r. (1978). assertion. in syntax and semantics 9 (1978): 315-332. (reprinted in steven davis, pragmatics: a reader, new york and oxford: oxford university press, 1991.) stone, m. (2004) intention, interpretation and the computational structure of language. cognitive science 28(5):781-809. sturt, p. and crocker, m. (1996). monotonic syntactic processing: a cross-linguistic study of attachment and reanalysis. language and cognitive processes 11(5): 449–494. swinney, d. (1979). lexical access during sentence comprehension: (re) consideration of context effects. journal of verbal learning and verbal behavior, 18:645-659. tanenhaus m.k., leiman, j.l., and seidenberg, m.s. (1979). evidence for multiple stages in the processing of ambiguous words in syntactic contexts. journal of verbal learning and verbal behavior, 18:427-440. tanenhaus, m.k., spivey-knowlton, m.j., eberhard, k.m. & sedivy, j.e. (l995). integration of visual and linguistic information in spoken language comprehension. science, 268:1632-1634. traum, d. r. (1994). a computational theory of grounding in natural language conversation, ph.d. dissertation, computer science dept., u. rochester. traum, d. r. (1999). 20 questions on dialogue act taxonomies. in proceedings of amstelogue.. traum, d. r. , and hinkelman, e. a. (1992): conversation acts in task-oriented spoken dialogue. computational intelligence 8: 575-599 traum, d. r. (1999). computational models of grounding in collaborative systems. in working notes of aaai fall symposium on psychological models of communication, pages 124-131. traum, d. r. and nakatani, c. h. (1999). a two-level approach to coding dialogue for discourse structure: activities of the 1998 working group on higher-level structures, in proceedings of the acl'99 workshop towards standards and tools for discourse tagging, pages 101-108. traum, d. r. and larsson, d. (2003): the information state approach to dialogue management. in smith and kuppevelt (eds.): current and new directions in discourse & dialogue, kluwer academic publishers. pages 325-353 tuomela, r. (2000). cooperation. kluwer. walker, m. (1993). informational redundancy and resource bounds in dialogue. phd dissertation, university of pennsylvania. wilkes-gibbs, d. (1986). collaborative processes of language use in conversations. unpublished doctoral dissertation, stanford university, ca. completions, coordination and alignment in dialogue 89 wilkes-gibbs, d. (1995). coherence in collaboration: some examples from conversation. in: gernsbacher, m. a. and anderson, a. (eds.): coherence in spontaneous text. amsterdam: benjamins, pages 239-267 wolpert, d. m., doya, k. and kawato, m. (2003). a unifying computational framework for motor control and social interaction. phil. trans. r. soc. lond. b biol sci 358:593–602. master.dvi dialogue and discourse 4(2) (2013) 282-324 doi: 10.5087/dad.2013.212 a discriminative analysis of fine-grained semantic relations including presupposition: annotation and classification galina tremper tremper@cl.uni-heidelberg.de department of computational linguistics heidelberg university, germany anette frank frank@cl.uni-heidelberg.de department of computational linguistics heidelberg university, germany editors: stefanie dipper, heike zinsmeister, bonnie webber abstract in contrast to classical lexical semantic relations between verbs, such as antonymy, synonymy or hypernymy, presupposition is a lexically triggered semantic relation that is not well covered in existing lexical resources. it is also understudied in the field of corpus-based methods of learning semantic relations. yet, presupposition is very importantfor semantic and discourse analysis tasks, given the implicit information that it conveys. in this paper we present a corpus-based method for acquiring presupposition-triggering verbs along with verbal relata that express their presupposed meaning. we approach this difficult task using a discriminative classification method that jointly determines and distinguishes a broader set of inferential semantic relations between verbs. the present paper focuses on important methodological aspects of our work: (i) a discriminative analysis of the semantic properties of the chosen set ofrelations, (ii) the selection of features for corpus-based classification and (iii) design decisionsfor the manual annotation of fine-grained semantic relations between verbs. (iv) we present the results of a practical annotation effort leading to a gold standard resource for our relation inventory, and (v) we report results for automatic classification of our target set of fine-grained semantic relations, including presupposition. we achieve a classification performance of 55% f1-score, a 100% improvement over a best-feature baseline. keywords: presupposition, entailment, question-based annotation,automatic classification. 1. introduction computing lexical-semantic and discourse-level information is crucial in event-based semantic processing tasks. this is not trivial, because significant portions of content conveyed in a discourse may not be overtly realized. consider the examples (1.a) and (1.b), where (1.a) bears a presupposition that is overtly expressed in (1.b): (1) a. spain won the finals of the 2010 world cup. b. spain played the finals of the 2010 world cup. the presupposition expressed in (1.b) is implicitly encoded in (1.a), through lexical knowledge about the verbwin, and is thus automatically understood by humans who interpret (1.a), given their linguistic knowledge about the verbswin andplay. automatically acquiring this kind of lexical semantic information is one of the objectives of the presentwork. c©2013 galina tremper and anette frank submitted 04/12; accepted 01/13; published online 07/13 discriminative analysis of fine-grained semantic relations one reason for embedding the acquisition of presupposition-triggering verbs in a discriminative classification task is thatpresuppositionneeds to be carefully distinguished from other lexical relations, in particularentailment. the two relations are closely related, but crucially differ in specific aspects. consider the sentence pair in (2). (2) a. president john f. kennedy was assassinated. b. president john f. kennedy died. sentence (2.a) logically entails (2.b). but how does this differ from the presuppositional relation between (1.a) and (1.b)? generally speaking,entailmentis a strictly logical implication relation holding between propositionsp andq in such a way that wheneverp holds true,q also holds true. presupposition, by contrast, is a relation that may be perceived as holding between propositions, but is often viewed as a pragmatic relation holding between aspeaker and a proposition. crucially, presuppositionsare what a speaker assumes to hold true as a precondition for asentence to be true. our focus is on presuppositions as conventional implicatures, as opposed to conversational implicatures (levinson, 1983). there are a variety of linguistic sources for presuppositions, including possessive pronouns, definite reference, or cleft-/wh-constructions that triggerspecific presuppositions, such as possessive relations or existence. while these constitute a closed list, we are interested in lexically triggered presuppositions, mainly by verbs, that are grounded in the lexical meaning of the triggering predicates. examples are widespread, including aspectual verbssuch asbegin/start doing x– not having done x beforebut most importantly general action verbs such aswin – play, know – learn, find – lose, etc. thus, in this work, we concentrate on a notion of presupposition that is restricted to the lexical meaning relation holding between the presupposition-triggering verb and the verbal predicate of the evoked presupposition. but then again, how to distinguish between verb pairs that characterize lexically triggered presuppositions as in (1) from pairs of verbs that license a classical entailment relation as in (2)? the differences betweenpresuppositionandentailmentcan be studied using special presupposition tests (levinson, 1983). the most compelling one, whichwe will use throughout, is the negation test. it shows that presupposition is preserved under negation, while entailment is not. applied to (1) and (2), we note that (3.a), the negation of (1.a), still implies (1.b), while (3.b), the negation of (2.a), does not imply (2.b). this can be taken as evidence that win lexically presupposesplay, while assassinateanddie are lexical licensors for logical entailment. (3) a. spain didn’t win the finals of the 2010 world cup. b. president john f. kennedy wasn’t assassinated. the negation test not only helps us to distinguish these closely related verb relations. it also points to the distinct behavior of these relations in deriving implicit meaning from discourse, which is the main motivation underlying our work. if we encounter the verbwin in the intended meaning x wins the gamein some piece of discourse, we may inferx played the game– whether the phrase is negated or not. for a verb that stands in an entailment relation, by contrast, we need to make sure that the triggering verb is not in the scope of negation. so,x was killedimplies thatx died, but x wasn’t killeddoes not license this inference. 283 tremper andfrank similar to entailment, presuppositions are essentially grounded in world knowledge. at the same time, they are crucial for the computation of discoursemeaning and inference. this is exemplified in (4), a typical case of presupposition that introduces additional, implicit knowledge, by so-calledaccommodationbehavior (van der sandt, 1992; geurts and beaver, 2012). thepredicate lift licenses the presupposition that the ban on deep sea drilling that has been lifted had previously been imposed. because this presupposition is lexically triggered, it causes anyone unaware of this piece of world knowledge to infer that a moratorium on deep-water drilling had beenimposedfor the gulf of mexico some time before october 12, 2010, the publication date of the article. (4) the obama administration lifted its moratorium on deep-water drilling in the gulf of mexico tuesday, replacing it with what interior secretary ken salazar is calling a gold standard of safety standards for operators looking to drill in water depths greater than 500 feet.1 it is their relevance for discourse understanding and inference that motivates capturing lexical semantic relations in computational lexicons, to make themavailable for lexically driven inferences in nlp applications (frank and pádo, 2012). among these arethe major taxonomic lexical semantic relations, such asantonymy, synonymyor hypernymythat are grounded in linguistic tradition (lyons, 1977) and that form the core of lexical semantic resources such as wordnet (fellbaum, 1998). recent efforts in computational linguistics further aim toautomatically acquire lexical relations that determine linguistically licensed inferences, such as entailmentand other more fine-grained relations, which are not yet covered in sufficient detail andcoverage in the wordnet data base. chklovski and pantel (2004) were first to attempt the automatic classification of fine-grained verb semantic relations, such assimilarity, strength, antonymy, enablementand happens-before in verbocean. in the present paper we aim to extend the classification of semantic relations between verbs to lexical inferences licensed bypresupposition. to our knowledge, this has not been attempted before. we will address this task in a corpus-based discriminative classification task – by distinguishing presupposition from other semantic relations, in particularentailment, temporal inclusionandantonymy. our overall aim is to capture implicit lexical meanings conveyed by verbs, and to make this knowledge explicit for improved discourse interpretationby lexically induced inferences. this overall aim can be divided into two tasks: detecting and discriminating fine-grained semantic relations: we first detect and distinguish fine-grained semantic relations holding between verbs at the type level, to encode this lexical knowledge in lexical semantic resources. deriving implicit meaning from text: in a second step, we will apply this knowledge for the interpretation of discourse, at the context level, in order toenrich the overtly expressed content with implicit knowledge conveyed by presupposition, entailment, or other lexically supported semantic inferences. that is, when detecting a verb in a given piece of discourse that stands in a particular meaning relation with some other verb, we apply the learned lexical knowledge to enrich the discourse representation with this hidden meaning relation, by lexically driven inferences. through the inferred semantic knowledge we obtain densely structured semantic representations of discourse that can improve the quality of automatic semantic and discourse processing tasks, such as information extraction, text summarization, question-answering and full-fledged textual inferencing or natural language understanding tasks. 1. source: the christian science monitor, oct. 12, 2010. 284 discriminative analysis of fine-grained semantic relations the present paper concentrates on the first task. we present acorpus-based method for learning semantic relations between verbs, with a special interest in detecting verbs related by or triggering presuppositions. learning focused lexical semantic relations from corpora is a hard task. our main strategy for approaching this task is to design features forclassification that are able to discriminate presupposition from other lexical relations. a novel aspect of our work is that we employtype-based features that are derived fromlogical-semantic propertiesof the targeted lexical relations. as it turns out, the classification we aim to perform is even difficult for humans: the complex inference patterns that characterize the differences between the semantic relations we consider are difficult to discern using classical annotation schemes. wedevise a question-based annotation design that yields reliable annotation results. on the basis of the resulting annotated data set we will present first results for automatic discriminative classification of fine-grained semantic relations between verbs using alternative classification architectures. the structure of the paper is as follows: section 2 reviews related work. section 3 motivates the choice of our target set of semantic relations and studies their discriminative properties. section 4 discusses different annotation strategies and their difficulties and develops a question-based annotation scenario that yields improved annotation quality. in section 5 we present two classification experiments and the results we obtain. we present an error analysis and compare our results to related work. finally we summarize and present conclusionsin section 6. 2. related work semantic relation acquisition. significant progress has been made during the last decade in automatic detection of semantic relations between pairs of words, using corpus-based methods. the majority of approaches follow thedistributional hypothesis: semantically related words tend to occur in similar contexts (firth, 1957). two types of methods can be distinguished in this field.2 pattern-based methodsmake use of specific lexico-syntactic patterns that identify individual relations, e.g., thesuch aspatterns used by hearst (1992) to detect hyponymy (is-a) relations between nouns. similar techniques have been applied to detectmeronymyrelations (girju et al., 2006). in contrast,distributional methodsrecord co-occurring words in the surrounding context of a target word, and compute semantic relatedness between two target words using measures of distributional similarity such ascosineor jaccard(mohammad and hirst, 2012). the strength of pattern-based approaches is that particular relations can be identified with high precision, if effective relation-identifying patterns can be determined. often, however, pattern-based approaches are critically lacking recall. distributionalapproaches do not suffer from such coverage problems. but distributional measures of ‘similarity’ and‘relatedness’ are in general not specific enough to permit a clear-cut distinction of individual meaning relations (baroni and lenci, 2011). pantel and pennacchiotti (2006) propose a weakly supervised pattern-based bootstrapping algorithm, espresso, that addresses the recall problem. it admits generic patterns – high-recall, yet low-precision patterns – which may refer to more than one semantic class. in conjunction with espresso’s refined filtering methods, generic patterns yield high recall without loss of precision. in our approach, we will perform semantic relation classification in a different way, using features for classification that encode more abstractlinguistic propertiesof individual relation types. this way we avoid the fuzziness of distributional measures and, at the same time, compensate for the lack of discriminative surface patterns for the inferential relations we need to distinguish. 2. see frank and pádo (2012) for an overview. 285 tremper andfrank acquisition of (verb) inference rules. a related strand of work aims at the automatic acquisition of inference rules. engendered by the recognizing textual entailment (rte) challenges, the main goal is to identify inference relations holding between twopieces of text, such that one of them can be inferred from the other (dagan et al., 2009). the notion of inference that underlies the rte challenges is informally defined as themost probableinference that can be drawn from some text, relying on common human understanding of language and background knowledge. pekar (2008), aharon et al. (2010), berant et al. (2012) and weisman et al. (2012) extract broad inferential relations between verbs, without sub-classifying them into more fine-grained relation types, such aspresupposition, entailmentor cause. however, knowledge about the specific inferential properties of these relations is crucial for drawing correct inferences in a given context. distinguishing fine-grained semantic relations between verbs. only few attempts tried to further distinguish inferential relations between verbs. chklovski and pantel (2004) performed fine-grained semantic relation classification with verbocean. they built on work by lin and pantel (2001), who proposed a distributional measure that extracts highly associated verbs. chklovski and pantel (2004) took lin’s semantically associated verb pairs as a starting point and applied a semi-automatic pattern-based approach for determining fine-grained semantic relation types, includingsimilarity (synonyms or siblings),strength(synonyms or siblings, where one of the verbs expresses a more intense action),antonymy, enablement (a type of causal relation) andhappens-before. this inventory of semantic relations is different from ours. in contrast to verbocean, we do not considersynonymyandstrength. also, there is no direct mapping from their entailment relationsenablementandhappens-beforeto our target relations. inui et al. (2005) concentrate on the acquisition of causal knowledge. they sub-classify causal relations into the four types:cause, effect, preconditionandmeans, using the japanese connective marker tameas a contextual indicator. they distinguish two types of events: actions (act) and states of affairs (soa). for cause(soa1, soa2) and effect(act1 , soa2), soa2 happens as a result ofsoa1 or act1, respectively. withprecond(soa1 , act2), act2 cannot happen until soa1 has taken place. finally,means(act1, act2) involves two actions sharing agents and that can be paraphrased asact1 in order toact2. unlike inui et al. (2005) we do not distinguish subclasses of causal relations, but consider them as special cases ofentailment. important work on clarifying the implicative properties ofverbs has been presented by karttunen (2012). similar to our work, he tries to divide implicative constructions into different types, but in contrast to our work, he studies the relation between the implicative verb (phrase) and its complement clause. karttunen (2012) identifies different types of implicative signatures and classifies the verbs accordingly. for example,refuse tois a one-way implicative verb with the implicative signature+−: the entailment applies in affirmative contexts only, and consists in negating the complement clause. at present, this classification has not beenautomated. computing presuppositions. only little work is devoted to the computational treatment of presupposition. bos (2003) adopted the algorithm of van der sandt (1992) for presupposition resolution. his approach is embedded in the framework of drt (kamp and reyle, 1993). it requires heavy preprocessing and a lexical repository of presuppositional relations. clausen and manning (2009) compute presuppositions in a shallow inference framework called ‘natural logic’. their account is restricted to computing factivity presuppositions of sentence embedding verbs. in the field of corpus-based learning of semantic relations, the automatic acquisition of presupposition relations remains understudied. 286 discriminative analysis of fine-grained semantic relations 3. a corpus-based method for learning semantic relations we present a corpus-based method for learning semantic relations between verbs with a focus on verbs involved in lexically triggered presupposition relations. in order to better capture the specific properties of presuppositional relations, we embed this task in a discriminative classification setup. as target classes we initially consider five relation types:presupposition, entailment, temporal inclusion (which coverstroponymyandproper temporal inclusion), antonymyandsynonymythat we aim to differentiate, as well as a negative class of verb pairs related by some other relation, or that do not stand in any relation at all (other/unrelated).3 3.1 selection of target semantic relations this target set of relations is motivated by three criteria.first of all, we aim at a broad space of relation types, in order to acquire a wide spectrum of relations that bear inferential characteristics. for this reason, our selection encompasses the taxonomic relationshypernymy/troponymy, synonymyandantonymy, which have proven efficient in computational textual entailment and questionanswering tasks, as well as classical non-taxonomic inference relations. second, as our focus is on relations between verbs, the relations should be characteristic for verbs. finally, we need to choose relation types that are sufficiently discriminative to permit automatic subclassification using corpusbased methods. inferential relations (between verbs). lexical resources such as wordnet (fellbaum, 1998) or germanet (kunze and lemnitzer, 2002) cover the core taxonomic relationssynonymy(through the notion of synsets),antonymyandhypernymy/hyponymy. in the verbal domain,hypernymycorresponds to the special relationtroponymy(for instance,march – move, mutter – talk).4 these relations are clearly inferential: for synonymous verbsv1 andv2 and a propositionpv1 based onv1, we can inferpv2/v1 , i.e., the propositionpv2 that results from substitutingv1 with v2 and vice versa. antonymy allows us to infer¬pv2 from pv1 . for hypernymy or troponymy, we can inferpv2 from pv1 , but we cannot inferpv1 from pv2. cutting across these taxonomic relations, which apply to all major open class categories, we find inferential relations that are specific to verbs. these are based on temporal, causal, or inferential relations that are grounded in world knowledge about events: temporal inclusion, causation, entailmentor presupposition. temporal inclusion (sleep – snore) differs from troponymy in thatsnoring is not a special way ofsleepingbut merely an action that may occurwhile sleeping.causationcan be considered a special form of entailment that involves a physical or other external force that brings about a state of affairs:feed – eat, kill – die(carter, 1976). finally, we find a broad class of verbs that lexicallyentailor presupposeone another, such asbreathe – liveor win – play.5 they typically do not instantiate hierarchically related concepts as in troponymy, but can be characterized as ‘log3. in fact, we will excludesynonymylater on, for reasons relating to the specific corpus-based methods we apply. nevertheless we include it here, for the general discussionof the inferential properties of lexical-semantic relations. 4. while wordnet (fellbaum, 1998) makes use of thetroponymyrelation for verbs, germanet uses thehypernymy relation across the different word categories (henrich andhinrichs, 2010). 5. in what follows we adopt the commonly used convention, as e.g. in fellbaum (1998), that grounds inferential relations holding between propositions to their licensing verbs. i.e., we refer to pairs of verbsv1 andv2 that are able to license entailmentor presuppositionrelations between propositionspv1 and pv2 as standing in alexical entailmentand presuppositionrelation, respectively. 287 tremper andfrank ical consequences’ or ‘preconditions’ of each other and aregrounded in real-world knowledge. all of the latter relations are unidirectional, except for entailment, for which modus tollens holds.6 selecting target relations for classification. fellbaum (1998) establishes a hierarchy of inferential relations between verbs that distinguishes four types of lexical entailment:troponymyand proper temporal inclusion(which both involve a temporal inclusion relation between verbs) are distinguished frombackward presuppositionandcause(which do not involve temporal inclusion).7 this relation inventory is very fine-grained. in practice itis difficult to discriminate relation instances along the relevant criteria, such as ‘external force’ for causation, or proper temporal inclusion vs. coextensiveness, in order to discriminateproper temporal inclusionfrom troponymy. in fact, although fellbaum’s hierarchy distinguishes four relation types,backward presuppositionand proper temporal inclusionhave been grouped together (richens, 2008).8 in our approach we adopt a different relation hierarchy (seefigure 1). we adopt wordnet’s basic taxonomic relationssynonymy, antonymyand troponymy(as a special class ofhypernymy in the verbal domain). unlike wordnet, we rangecausationwith the more generalentailment relation. similar to wordnet we groupproper temporal inclusionwith troponymyas they share inferential properties, but distinguishentailment(inclusive ofcausation) from presuppositionsince these relations show distinct inferential behavior. the latter two classes differ from the former, as the verbs are involved in temporal sequence (precedence, overlap or succession).9 this leaves us with five relations that span a large range of inferential relations: taxonomic and non-taxonomic, symmetric and asymmetric, that we set out to distinguish using corpus-based classification. 6. we follow the classical definitions forpresuppositionandentailment, as given below: presuppositionis defined by strawson (1950) as follows: a statementa presupposesanother statementb iff: (a) if a is true, thenb is true; (b) ifa is false, thenb is true. condition (b) is known as the property ofpersistence under negationthat is characteristic for presupposition. the backward presuppositionrelation in wordnet is based on this definition, and like fellbaum (1998) we ground the presupposition relation holding between propositions to a lexical relation holding between the presupposition-triggering verb and the verbal predicate of the triggered presupposition. entailment, also referred to aslogical consequence, can be defined as follows: a semantically entailsb iff every situation that makesa true, makesb true. (levinson, 1983) similar to presupposition we consider only lexical entailment relations holding between verbs that determine entailment between propositionsa andb. 7. her terminology differs from the one adopted here, with ‘entailment’ being largely equivalent to our use of ‘inferential’. the structure of the wordnet entailment hierarchy isreproduced below. entailment + temporal inclusion − temporal inclusion + troponymy (coextensiveness) − troponymy (proper inclusion) backward presupposition cause (limp, walk) (snore, sleep) (succeed, try) (raise, rise) 8. by grouping(backward) presuppositionandcausetogether as special forms ofentailment, wordnet collapses two relation types with clearly distinct inferential properties, especially with regard to negation and cancellation (cf. section 1 and below discussion of (5)–(7), p. 289 and table 2,p. 290). moreover,causationcan be considered a special form ofentailment, while in this taxonomyentailmentis not represented as an individuated semantic relation type. 9. note that the distinction between proper temporal inclusion and cases of overlap in temporal sequence is difficult. however, we adhere to this distinction, as introduced by fellbaum (1998), as indeed we find clear differences in the inferential properties of these two types of verb relations. 288 discriminative analysis of fine-grained semantic relations verb semantic relations symmetric asymmetric synonymy antonymy “temporal inclusion” temporal sequence (<, o,>) (fix, repair) (go, stay) troponymy (is-a) proper(⊂) entailment presupposition (mutter, talk) temporal inclusion (buy, own),(arrive, depart) (win, play) (snore, sleep) (breathe, live) taxonomic non-taxonomic figure 1: hierarchy of inferential semantic relations, with selected classes printed in bold. ± temporal semantic example behavior under negation sequence relation (v1, v2) (v1, v2): ix : p±v1 cond p±v2 (v2, v1): ix : p±v2 cond p±v1 temporal i1: +� + i1: +� + precedence entailment (buy, own) i2: −� +e i2: +� −e (v1 precv2) i3: −� − i3: −� − v1 < v2 i4: ¬(+� −) i4: ¬(−� +) i1: +� + i1: +� + entailment (arrive, depart) i2: −� +e i2: +� −e temporal i3: −� − i3: −� − succession i4: ¬(+� −) i4: ¬(−� +) (v1 succv2) i1: +� + i1: +� + v2 < v1 presupposition (win, play) i2: −� +p i2: +� − i3: −� −c i3: −� − i4: ¬(+� −) i4: ¬(−� +) temporal i1: +� + i1: +� + overlap entailment (breathe, live) i2: −� +e i2: +� −e (v1 o v2) i3: −� − i3: −� − i4: ¬(+� −) i4: ¬(−� +) temporal i1: +� + i1: +� + inclusion (snore, sleep) i2: −� +p i2: +� − (proper t.i. i3: −� −c i3: −� − & troponymy) (mutter, talk) i4: ¬(+� −) i4: ¬(−� +) i1:¬(+� +) i1: ¬(+� +) antonymy (love, hate) i2: −� +t.n.d. i2: +� − − temporal i3:¬(−� −)t.n.d. i3: ¬(−� −)t.n.d. sequence i4: +� − i4: −� +t.n.d. i1: +� + i1: +� + synonymy (fix, repair) i2: ¬(−� +) i2: ¬(+� −) i3: −� − i3: −� − i4: ¬(+� −) i4: ¬(−� +) table 1: inferential properties of verb relation types.+/−: positive/negative polarity ofv1/v2. p indicatespersistence under negation; c: cancellation; e: exception; t.n.d: tertium non datur. 289 tremper andfrank inferential properties. table 1 details the inferential properties we find with instances of verb pairs instantiating the chosen relation types. these properties will establish important criteria for the automatic classification of verb relations into the target classes. we discriminate verb pairs (v1,v2) along two dimensions: theirtemporal sequence properties, in terms of the typical temporal relation holding between corresponding events (or no such relation), and theirinferential behavior, especially with regard to theirbehavior under negation. inferences that are found valid for the different subclasses are evaluated for both directions (i.e., withv1 or v2 as trigger verb) and are specified using modal conditional statements relating propositions involving the related verbs. we make use of epistemic conditionals forcharacterizing the inferential properties for different combinations of verb polarities, as the decisions for classification made by human annotators are best guided in terms of epistemic modal reasoning. in judging inferential patterns for related verb pairs, subjects consider whether possible situations that support the truth of an event referred to byv1 will also support the truth of an event referred to byv2. for each relation type we consider four inferential patterns (i1 to i4) using positive (+) and negative (−) polarity of the related verbs.10 an (epistemic) conditional thatnecessarily holds true (pv1 � pv2) corresponds tothe valid inferencethat wheneverpv1 is true in an (epistemically) accessible worldw, pv2 holds true inw. the weakerexistential reading(pv1 � pv2) is true if there is at least one (epistemically) accessible worldw wherepv1 is true that also supports the truth of pv2 . that is, we can conclude frompv1 thatpv2 may hold true or not:pv2 ∨ ¬pv2. ¬(pv1 � pv2) represents anegative inference, i.e., we cannot concludepv2 from pv1 . table 1 shows a clear contrast between symmetric and asymmetric relations. thesymmetric relations synonymyandantonymyshow symmetric inference patterns when applying forwards and backwards inferences (ix : pv1cond pv2 vs. ix : pv2cond pv1). for both relation types, the inferences reflect the core logical properties of the respective relations, allowing us to inferpv2 from pv1 for synonymy and¬pv2 from pv1 for antonymy (with obvious variations for different polarities).11 the asymmetric relations(presupposition, entailment, temporal inclusion) all pattern alike in terms of the forwards and backwards inferencesi1 andi4, which allow us to inferpv2 from pv1 in forward direction (withi4 the corollary ofi1 in the same direction) and¬pv1 from¬pv2 in backward direction. in forward direction, all asymmetric relation types permit us to concludepv2 ∨¬pv2 from ¬pv1, yet it is the inference typesi2 andi3 that mark the core of their differences. the inferential patternsi2 andi3, while superficially similar in forward direction, strictly divide entailment(e) (in all possible ways of temporal sequencing) frompresupposition(p) andtemporal inclusion (t), in that for entailment, applying common sense reasoning, we can infer¬pv2 from ¬pv1 as the ‘normal course of things’, while forpresuppositionand temporal inclusionwe can in general concludepv2 from¬pv1 , in line with the well-known inferential property of presuppositions that ‘survive under negation’ (levinson, 1983). that is, the corresponding inferencesi2 for entailmentandi3 for presuppositionandtemporal inclusionrepresent exceptional cases forentailment, and cancellation of presuppositions in the case ofpresuppositionandtemporal inclusion.12 10. the conditional statements used in table 1 to characterize valid inferences serve expository purposes only. we follow the definition of conditionals using a standard definition ofepistemic accessibility (see e.g. gamut (1991)). 11. note that for antonymy we adopt an idealized situation of‘tertium non datur’, that is, we only consider antonyms that realize the extreme ends of a scale, and ignore any intermediate values, such asbeing indifferent, for loveand hate. this assumption affects the inference patterns with negative antecedents for antonymy. 12. i3, in forward direction, with¬v1 as a trigger verb for entailment, represents a typical form of abductive inference that is subject to cancellation (similar toi3 for presupposition). karttunen (2012), following geis andzwicky (1971), calls such non-monotonic inferences ‘invited inferences’. 290 discriminative analysis of fine-grained semantic relations this can be shown by applying a number of paraphrase tests to verb pairs for the various relations, as illustrated in (5) to (7). the paraphrase pattern in (5) shows thatpv2 can be consistent with pv1 , but it does not discriminate the underlying differences between the relation types, nor does (6), which is designed to test for ‘persistence under negation’ as is typical for presuppositions. (5) you don’t/didn’tv1 but you (have)v2.13 (6) you don’t/didn’tv1, and this is because you didn’tv2 in the first place.14 however, (7), which explicitly refers to exceptional situations that do not correspond to the ‘normal course of events’, clearly establishes thatentailmentrelations are subject to exceptional conditions that can make the universal conditional fail (7.d–f), while for (7.a–c) the oddity of ‘exception catching paraphrases’ corroborates the behavior of presuppositionand temporal inclusion as being persistent under negation in their default interpretation. it is only by explicit cancellation, as in (6), that¬pv2 can be inferred from¬pv1. (7)a.–c.# you didn’twin/snore/mutter, so you didn’tplay/sleep/talkor you might haveplayed/ slept/talkedbut something exceptional happened so that you didn’twin/snore/mutter. (p,v2 < v1; t, v1 ⊂ v2; t, v1 ⊂ v2) d. you didn’t arrive, so you didn’tdepartor you might havedepartedbut something exceptional happened so that you didn’tarrive. (e,v2 < v1) e. you didn’t buy it, so you don’town it or you mightown it but something exceptional is the case so that you didn’tbuy it. (e, v1 < v2) f. he doesn’tbreathe, so he doesn’tlive / isn’t aliveor he mightlive / bealiveand something exceptional is the case so that he doesn’tbreathe. (e,v1 o v2) these differences are recorded in table 1 by marking forwards inferences under negation (i2) as subject to ‘exceptions’ (e) for all entailmentrelation types (withi2, in backward direction, as its inverse). in contrast,i2 is marked as the default inference (p: ‘persistence under negation’) 13. paraphrase instances forpresupposition(p),entailment(e) andtemporal inclusion(t): (i) you didn’t win, but you haveplayed. (p) (ii) you didn’t snore, but you haveslept. (t) (iii) you didn’t mutter, but you havetalked. (t) (iv) you didn’t arrive, but you havedeparted. (e) (v) you didn’t buy it, but youown it. (e) (vi) he doesn’tbreathe, but he (still)lives / is alive. (e) 14. example (v) is slightly anomalous, but this is not specific to theentailmentrelation, but rather due to temporal sequence properties, withv2 following v1, which does not conform to this specific pattern. (i) you didn’t win, and this is because you didn’tplay in the first place. (p) (ii) you didn’t snore, and this is because you didn’tsleepin the first place. (t) (iii) you didn’t mutter, and this is because you didn’ttalk in the first place. (t) (iv) you didn’t arrive, and this is because you didn’tdepartin the first place. (e) (v) # you didn’t buy it, and this is because you didn’town it in the first place. (e) (vi) he doesn’tbreathe, and this is because he doesn’tlive / isn’t alive in the first place. (e) 291 tremper andfrank inference patterns (v1,v2) relation temp.rel (v1,v2) ix : p±v1 opp±v2 example i1: +� + i buy – i own entailment v1 (<,o,>) v2 i2: −� +exception i don’t buy, but i (still) own (buy, own) i3: −� − i don’t buy, so i (normally) don’t own i4: ¬(+� −) presupposition v2 < v1 i1: +� + i win – i played (win, play) i2: −� +persistence i didn’t win but/when i played temp. inclusion i3: −� − cancellation i didn’t win – because i didn’t play (snore, sleep) v1 ⊂ / is-av2 i4: ¬(+� −) i1:¬(+� +) antonymy no temp. seq. i2: −� +tertium n.d. you don’t love – you hate (love, hate) i3:¬(−� −)tertium n.d. i4: +� you love – you don’t hate i1: +� + i fix – i repair synonymy no temp. seq. i2: ¬(−� +) (fix, repair) i3: −� − i don’t fix – i don’t repair i4: ¬(+� −) table 2: inference patterns and paraphrases for the different relation types. for presupposition, and similarly for both relation subtypes oftemporal inclusion: proper temporal inclusion (snore, sleep) andtroponymy(mutter, talk). conversely,i3 represents the case of ‘cancellation’ (c) for presuppositionand temporal inclusion, whereas it represents the ‘normal course of events’ forentailment. table 2 summarizes these outcomes, by aligning the inference patterns for the main relation types with the inference paraphrases they support as ‘normal’ or ‘invited’ inferences, or as inferences that must be marked as exceptions. 3.2 discriminating properties of semantic relations between verbs as can be seen from this analysis, the inferential properties of the chosen set of relations are complex and difficult to distinguish. however, their inferential properties go along with two dimensions: temporal sequence properties on the one hand and behavior with regard to negation on the other. temporal sequence. we observe that the taxonomic lexical semantic relationsantonymy, synonymyandtemporal inclusiontypically do not involve a temporal order. in contrast,presupposition relations between verbs do involve a temporal sequence. theevent that is presupposed, being considered as a precondition, typically precedes the event that triggers the presupposition. the verbs which stand in anentailmentrelation may or may not involve a temporal succession: the overtly realized verb can precede or succeed the entailed verb, but we also find events that are temporally overlapping, such aslive / be aliveandbreath. negation. another important aspect is the behavior of the different semantic relations under negation. presuppositionandtemporal inclusionare preserved under negation. this distinguishes them from entailmentandsynonymywhich do not persist under negation. 292 discriminative analysis of fine-grained semantic relations behavior under negation (v1, v2) (¬v1, v2) (¬v1,¬v2) (v1,¬v2) v1 precedesv2 e (e)e e temporal v1 succeedsv2 e (e)e e sequence p p (p)c v1 overlapsv2 e (e)e e t t (t)c no temporal a a sequence s s table 3: properties of the semantic relations: p(resupposition), e(ntailment), t(emporal inclusion), a(ntonymy), s(ynonymy);e: exceptions;c: cancellation. in fact, these temporal sequence and negation properties cross-classify and fully distinguish the selected semantic relation classes. this is schematicallyrepresented in table 3. the table reads as follows. we continue to usev1 as a placeholder for the trigger verb and v2 for the related verb.15 for the two dimensionsbehavior under negationandtemporal sequence we list the possible instantiations of these relation properties in terms of different combinations of negated and non-negated verb predicates and the different sequencing possibilities:v1 (typically) temporally precedes/succeeds/overlaps withv2, or no temporal sequence can be determined. within the table fields we record the relation types that support thecorresponding inference patterns. for thepresuppositionverb pair(win, play), for instance, the event of winning (v1) typically temporally succeeds the event of playing (v2). p(resupposition) therefore fills the second row. the presuppositional relation holds in case both events are asserted to hold true. p(resupposition) therefore fills the first column, marked(v1, v2). the event of not winning could be interpreted in two ways: its default interpretation: persistence under negation – you do not win although you’ve been playing(¬v1, v2), or else cancellation – you did not win because you did not play at all(¬v1,¬v2). but crucially, winning without playing(v1,¬v2) does not conform with the presuppositional relation between these verbs, so the respective field remains empty. for entailmentpairs (e) such as(kill, die) or (buy, own), we note thaty being killed entailsy being dead(v1, v2), but if y is not killed we do in general not conclude thaty is dead(¬v1, v2) – unless by considering other possible causes that may not be considered relevant in the situation at hand. thus, ify is not killed, we assume as default interpretation that (under normal circumstances) y is not dead (again – unless from some other cause)(¬v1,¬v2).16 both cancellation forpresupposition(c) and exceptional cases for inference under negated antecedents forentailment(e) are thus marked as exceptional inference patterns (indicated by brackets) that we do not assume to find frequently realized in corpus instances. 15. for the symmetric relationsantonymyandsynonymythere is no distinguished trigger verb. 16. this assumption is debatable, as only the inverse relation (¬v2,¬v1) is strictly entailed: ify is not dead,y has not been killed. however, as discussed above, we include this case as a typical form of abductive inference that is subject to cancellation as is presupposition whenever we encounter¬v1 as a trigger verb for entailment. note that nothing hinges on this assumption regarding the discriminative power of negation properties, as entailment differs from presupposition regarding persistence under negation. 293 tremper andfrank by examining these temporal and negation properties encoded in table 3, we find that they can be used to discriminate the considered semantic relation types: (i) presuppositionandentailment(whether or not temporally related) are distinguished on the basis of persistence under negation, which holds forpresuppositiononly. the same holds for temporal inclusionvs.entailment. (ii) temporal inclusionandpresuppositionbehave alike regarding negation properties, but can be distinguished in terms of temporal sequencing properties. (iii) entailmentbetween overlapping events is difficult to distinguish from(proper) temporal inclusionsolely on the basis of temporal properties. but due to their inferential behavior under negation, they can be clearly distinguished. (iv) antonymyclearly differs fromentailmentandpresuppositionwith respect to both properties, and fromtemporal inclusion, regarding negation properties. (v) finally, antonymyandsynonymyare opposites to each other regarding negation properties. according to this analysis, the observed temporal and negation properties could be used to discriminate four of the five semantic relation types.synonymyandentailmentare difficult to distinguish in cases whereentailmentdoes not involve a temporal sequence. however, as will become clear below, in our corpus-based classification approach, we will not be able to detect verb pair candidates for thesynonymyrelation. hence, we exclude this relation type for independent reasons and range it under the classunrelated. the remaining four relation types that will be subject to classification: presupposition, entailment, temporal inclusionandantonymywill be distinguished from a fifth class of unrelated verb pairs – which will include synonymous verbs, in case they (accidentally) are found to co-occur in corpus instances. 3.3 automatic classification of fine-grained semantic relations we pursue acorpus-basedsupervised classification approach to automatically detect and distinguish candidate verb pairs, given as types, as pertaining to one ofour target semantic relation types. to this end, we exploit the insights gained from the above analysis that yielded discriminating properties of these semantic relation types on the basis oftemporal sequenceand negation properties. in addition, we will employ a third dimension of contextualrelatedness, which records surface-level contextual relatedness properties of these semantic relations, using indicators such as embedding or coordinating conjunctions. these relatedness features will be utilized to distinguish semantically related fromunrelatedverb pairs, as we expect their contextual relatedness properties to be more diverse compared to semantically related verb pairs. moreover, contextual relatedness properties can be useful in cases where temporal or negation propertiesare difficult. selecting informative ‘contiguous’ corpus samples. for this approach we collect corpus samples of verb pairs co-occurring insinglesentences. even though co-occurrence in a single sentence bears high potential for the verbs being realized in a close syntagmatic relationship, this is not necessarily so. we therefore design a set of features that can be indicative of a close syntagmatic relationship between co-occurring verbs. we will refer to these features ascontiguity features.17 17. typical configurations of ‘contiguously related’ verbsare illustrated in (i). (i.a) replyingto the toast [..], dr julia kingsaidhow privileged the faculty was to have two active alumni associations. (i.b) you cansendus your comments by simplyclickingon this email. (i.c) this allows you toconnectanddisconnecteasily. 294 discriminative analysis of fine-grained semantic relations on the basis of a corpus study, we identified properties that can be indicative for contiguously related verbs in context: the distance between verbs, theiroccurrence in specific grammatical configurations as indicated by dependency relations or conjunctions, and co-referential binding of the arguments of both verbs. these features will be employed fordetecting contextual contiguity of verb pairs in specific contexts, and used to select context samples for classification that are informative for sub-classifying the semantic relations – including theunrelatedclass (see section 5.3.2). detecting type-based features for classification. our classification aims at assigning relation classes toverb pair types, and thus the feature vectors employed for classification must be defined accordingly at the type level. the temporal and negation properties we established as being discriminative for the chosen set of relations are equallytype-based. that is, they express properties we can identify in individual context samples, but not necessarily in all of them. in a corpus-based approach, we need to capture suchtype-basedproperties on the basis of individual classifications at the level of corpus samples, by observing and generalizing the information found with individual corpus samples. for our main classification features, this will be obtained in the following ways.18 in order to predicttemporal sequenceproperties as a type-level feature, we detect the temporal relation holding between individual verb pair occurrencesand compute the most prevalent temporal relation type for a given verb pair on the basis of these classifications, by applying an association measure such as point-wise mutual information (pmi). for determining thebehavior of inference under negationwe need to detect instances of all possible verb polarity combinations〈±v1,±v2〉 for different verb pairs in context. that is, we extract the information whether both verbs have positive/negative polarity, or whether the first verb has positive/negative polarity and the second verb has negative/positive polarity. from this token-level information we compute the probability for eachpolarity combination for any given verb pair. the obtained probabilities can be mapped to the negation properties of relations as displayed in table 3, where low probability of apolarity combination corresponds to unavailable or exceptional cases, and high probability manifests attested inference possibilities, under the respective relation. in order to obtain type-basedrelatednessfeatures, we raise twocontiguity features to the type level: verb distance and relating conjunctions. information about the average distance between verbs is crucial for distinguishing related and unrelated verb pair types. the distribution of conjunctions relating certain verb pairs can contribute indicative information for distinguishing specific semantic relations (e.g.,antonymyor temporal inclusion), or may indicate that the verbs are (probably) unrelated. finally, we measure the association between specific verb pairs on the basis of co-occurrence information manifested in a corpus, using pmi as association measure and use its strength as a type-based relatedness feature. supervised classification using manually labeled verb pairs (at the type level). we are going to perform supervised type-based classification using type-based feature vectors. that is, we need a training set of verb pairs annotated with the appropriate semantic relation (or the classunrelated) 18. detailed description of the features employed for classification is given in section 5.2. 295 tremper andfrank on the type level, i.e., for verb pairs out of context, and accordingly, we need a gold standard data set of unseen annotated verb pairs19 that can be used for testing. features for the type-based classification will be acquiredfor each verb pair in the training set, and similarly for the test set, using evidence gained from corpus sentences involving verb pairs that have been determined as beingcontiguouslyrelated. the features indicating the respective relation properties are acquired from the corpus samples and raised to the type level, as described above. in our experiments, the corpus samples will be drawn from a large web-based corpus, the ukwac corpus (baroni et al., 2009). at this step we excluded the synonymy relation, as even in such a large corpus, synonymous verbs usually do not occur contiguouslyin a single sentence. establishing annotated training and testing data sets. in order to build appropriate training and testing data sets, we cannot make use of existing resources such as wordnet or verbocean, as they assume different inventories of semantic relations (see section 3.1). we thus designed an annotation task for our target relation set, to construct training and testing data for the classification. 4. challenges of annotation annotating semantic relations, especially the relationspresuppositionandentailment, is a difficult task because of the subtlety of the tests and the involved decisions. in order to obtain reliable annotations it is important to define the task in an easy and accessible way and to give clear instructions to the annotators. for an initial annotation study we randomly selected a smallsample of 100 verb pairs for annotation. a further set of 250 verb pairs were annotated in a revised, question-based annotation setup. the resulting annotated data sets were used as development and gold standard test sets, respectively, for evaluating automatic semantic relation classificationin section 5. the verb pair candidates for annotation were chosen from the dirt collection (lin and pantel, 2001), a collection of automatically acquired semantically related verbs (see section 2, p. 284). 4.1 initial annotation strategies as a first take, we formulated two complementary annotation tasks: one was applied to verb pairs given as types out of context (type-based annotation) and another was applied to verb pairs presented in context (token-based annotation). we analyzed the difficulty of annotation in the respective annotation setups and examined to what degree these results correlate. in order to analyze the difficulty of annotation we gave each task to two annotators and computed the inter-annotator agreement between them.20 4.1.1 type-based annotation in this setup the verb pairs were presented to the annotatorswithout context. since some verbs can have more than one meaning and consequently verbs in a given verb pair can stand in more than one semantic relation, the annotators were allowed to assign more than one relation to each verb pair. 19. we restrict the notion of ‘gold standard’ data set to the subset of manually annotated verb pairs that we use for testing. 20. the annotators are trained computational linguistics students. they are native speakers of german with a high level of proficiency in english. the pairs of annotators which tookpart in the different annotation tasks are not always the same. only one student has taken part in both tasks and herannotations were taken to analyze the correlation between the different annotations. 296 discriminative analysis of fine-grained semantic relations semantic relation pattern example substitution in pattern presupposition v1 presupposesv2, win – play winningpresupposesplaying notv1 presupposesv2 not winningpresupposesplaying entailment v1 impliesv2, kill – die killing impliesdying notv1 doesn’t implyv2 not killing doesn’t implydying temporal v1 happens duringv2 or snore– sleep snoringhappens duringsleeping inclusion v1 is a special form ofv2 mutter– talk mutteringis a special form oftalking antonymy eitherv1 or v2, go– stay eithergoingor staying v1 is the opposite ofv2 goingis the opposite ofstaying other/unrelated none of the above jump– sing table 4: semantic relations and inference patterns for annotation. to support the annotators in their decisions, we provided them with a couple of inference patterns and examples for each semantic relation. this is shownin table 4. the inter-annotator agreement (iaa) for this task was 63% corresponding to a kappa21 value of k = 0.47. this can be taken as an indication of high difficulty when annotation of these semantic relations is performed out of context. 4.1.2 token-based annotation in a complementary setup, we tried to simplify the task by providing the annotators with verb pairs in their original contexts, consisting of single sentences. for this token-based annotation we chose the same 100 verb pairs and randomly selected 5 to 10 contextsfor each of them (there were 877 contexts overall). in contrast to type-based annotation, we only accepted a single relation label for a given verb pair. the inter-annotator agreement for this task was iaa = 77.4%,corresponding to a kappa value of k = 0.44. error analysis showed that the most important problems are not due to semantic relations which are difficult to distinguish (e.g.,presuppositionandentailment), but rather in determining whether or not there is a specific semantic relation between two verbs in a given context, i.e., the distinction between the ‘unrelated/other’ in contrast to the remaining semantic relation classes. 4.1.3 type-based vs. token-based annotation we examined the correlation between typeand token-based annotations by comparing the annotations of a single annotator for both annotation tasks.22 we chose only one annotator for this comparison, because we wanted to analyze how the decisions of one and the same annotator were affected by the different annotation setups.23 for 62% of the verb pair types we observe an overlap of labels, 28% of the verb pair types were assigned labelson the basis of the annotations in context which were not present on the type level, or else the type level label was not assigned in context, because of the small amount of contexts for a verb pair. for 10% of verb pair types we 21. cohen’s kappa; see cohen (1960). 22. only one annotator has taken part in both annotation tasks. 23. since in the initial task settings no translation of verbpairs was involved (cf. section 4.3), it was not possible to trace such differences across annotators. 297 tremper andfrank found conflicting annotations (e.g.,presuppositionandentailment). thus, for the most part (62%) the type-based annotation conforms with the ground truth obtained from token-based annotation. an additional 28% of verb pairs can be considered to be potentially correct. the divergences for these verb pairs could be explained by the random procedure of context extraction which does not always return appropriate contexts. they can also be explained by the difficulty for the annotator to consider all possible verb meanings for highly ambiguous verbs in type-based annotation. 4.2 a question-based annotation strategy using prototypical arguments our analysis of the two annotation setups clearly shows thatboth are difficult, yet in different ways. annotation on the type level is difficult because no indication is given about the intended meaning of the verbs. hence the annotators need to consider all possible combinations of meanings for any pairing of verbs. on the other hand, presenting the pairs in their original context does not make the decision much easier. this is because some sentences involve complex structure and interpretation difficulties, which require a lot of attention and time to annotate the individual examples. in general, the inference patterns offered to the annotators as decision criteria are rather involved, so they are sometimes difficult to check – with or without context. a general drawback of token-based annotation is that it is difficult to sample appropriate contexts for a balanced annotation set across the different relation types, and that annotation is necessarily time-consuming and expensive. in order to render the annotation task more reliable and lesstime-consuming, we need an annotation strategy that includes the positive elements of bothannotation strategies described above and that better supports the annotators in deciding on the applicability of the inference patterns. prototypical arguments in type-based annotation. one solution that captures positive aspects of typeand token-based annotation could be to have annotators considerverb pairs with prototypical argumentsinstead of offering them concrete sentences as disambiguating contexts. the argument abstractions could be represented by selectional preference classes. this offers the annotators hints on relevant readings to consider without them having to readand understand involved discourse snippets. at the same time, with a single reading of the verb in focus, the annotators do not need to consider and check pairs of verbs with multiple readings.evidently, annotation will proceed much quicker if it can be performed at the type level, even if different interpretation variants must be considered, based on selectional preference classes. question-based annotation. in order to support annotators in the verification of complexinference patterns, we develop aquestion scenarioto collect annotations. the idea is to guide the annotator step by step through the discriminative categorizingproperties, in particular temporal sequence and behavior under negation, using a cascade of case-adapted questions tailored to the verb pairs under investigation. the questions elicit the critical pieces of information needed to sub-classify the verb pair in question, according to the properties of relations displayed in table 3. a set of cascaded questions guide the annotator through all relevant decision criteria, where each question elicits only three possible answers:yes / no / maybe. in general, each annotation instance will be decided by three such consecutive questions. the collected answers can be used to distinguish between the target semantic relations and thusto annotate the data. we pursued both strategies: the use of prototypical arguments and question-based annotation, and applied them jointly in a third annotation task. examining the annotation quality obtained, we achieve considerable improvements, with an acceptable degree of inter-annotator agreement. 298 discriminative analysis of fine-grained semantic relations verb pairs with prototypical arguments semantic relation miss(person, person) – catch(person, person) unrelated(miss, catch) miss(person, train) – catch(person, train) antonymy (miss, catch) table 5: enriching verb pairs with prototypical arguments. expert vs. non-expert annotation. our annotators are trained computational linguistics students. since annotation is time-consuming and expensive, an obvious question is whether this simplified annotation setup – with annotation decisions broken down into more basic units – can make this difficult annotation task accessible for non-expert annotation. if so, we could collect larger sets of annotations using crowd-sourcing (munro et al., 2010). we will therefore compare the annotation quality obtained from linguistic experts to non-expert annotations. 4.2.1 integrating prototypical arguments in type-based annotation our analysis of problems in type-based and token-based annotation clearly showed that a general problem is the difficulty to capture verb interpretation dueto the ambiguity of verbs. the classification decisions crucially depend on verb interpretation and thus need to be controlled in the annotation task. further, we need to make sure annotators consider all relevant readings. both aspects are difficult to control in type-based annotation. in token-based annotation, annotators are often confronted with shades of meaning influenced by the specific context, which make decisions too case-specific and erroneous. we thus opt for a type-based annotation scheme that allows usto abstract away from concrete contexts and that at the same time allows us to control for verb ambiguity. this is achieved by offering prototypical arguments of the verbs, in terms of selectional preferences computed from corpora. the presentation of the verb pairs along with prototypical arguments helps the annotators focus on specific readings of the verbs, and thus avoid inconsistent annotations. an example is given in table 5 for the verb pairmissandcatch. when annotating this verb pair without context, two readings ofmissmay be considered:miss (1): feel or suffer from the lack of andmiss (2): fail to reach or get to. for the first reading, the annotator should determine the label unrelated, while for the second reading,antonymywould be the appropriate label. without control of context, the annotators could miss one orthe other reading, and we cannot trace which reading motivated the provided labels. presenting the verbs with prototypical arguments as generalizations directs the annotators to the appropriate interpretation and they can determine the corresponding label. since we record the arguments provided with the verbs, this kind of sense discrimination is available for both the learning and the classification process. it will also be crucial for inference in context, as it allows us to restrict inference of implied verb meanings to the appropriate interpretation of the trigger verb in a given context. for the computation of prototypical arguments of verb pairs, we apply resnik (1996)’s approach for computing selectional preference scores for verb arguments. with this we determine preference semantic classes as prototypical arguments insubject, objectandprepositional objectfunction. computing selectional association scores for verb pairs.resnik (1996) proposes an informationtheoretic measure to compute aselectional association scorebetween a predicatepi and a semantic 299 tremper andfrank classc that fills an argument ofpi as given in (8).24 he definesselectional preference strength s(pi) as the amount of information provided by the predicatepi for the posterior probability of co-occurring with some argument classc, compared to its prior probability. given this measure, he computes theselectional association scorebetween a predicate and a given particular classc by its relative contribution to the predicate’s overall selectional preference strength. (8) a(pi, c) = p (c|pi) log p (c|pi) p (c) s(pi) with s(pi) = ∑ c p (c|pi) log p (c|pi) p (c) since we are dealing with pairs of verbs, we slightly modify this measure to reflect the association of a classc with both verb predicatespi andpj, as stated in (9). (9) a(pi, pj, c) = p (c|pi,pj) log p (c|pi,pj ) p (c) s(pi,pj) with s(pi, pj) = ∑ c p (c|pi, pj) log p (c|pi,pj) p (c) we computed selectional preference scores for all verb paircandidates offered to the annotators, using the adapted measure in (9).25 prototypical arguments were selected manually from the arguments with the highest scores.26 controlling interpretation choices in the annotation task. having computed prototypical (preferential) argument classes for given verb pair candidates,these can be presented to the annotators as illustrated in table 5. however, in a number of cases prototypical arguments are notsufficient to clearly discriminate predicate interpretations. in order to detect such cases, we asked the annotators to translate the predicates into their mother language (if possible).27 examples of diverging interpretations are given in table 6, together with the labels the annotators assignedfor the interpretations they perceived. differences in translations were inspected manually. in case of divergences of interpretation, we not only record the actual interpretations chosen by the annotators, but also let the annotators re-annotate such verb pairs using the interpretation of their companion annotator as a constraint. this way we collect annotations for a maximum number of readings. 4.2.2 question-based annotation for classifying semantic relations the complex inference patterns that need to be considered inorder to distinguishentailment, presupposition, temporal inclusionandantonymymake the annotation difficult and error-prone. we therefore devise a question-based annotation setup that breaks down these complex annotation decisions into more basic units that are easier to decide. in a step-wise manner we elicit answers that 24. semantic classc is taken from a conceptional taxonomy. in our work we chose wordnet (version 3.0) as used in the nltk implementationhttp://nltk.org. 25. probabilities were estimated from sections 1 to 3 of the parsed ukwac corpus (baroni et al., 2009). parsing was performed using the stanford parser v1.6.4,http://nlp.stanford.edu/software/lex-parser.shtml. 26. we opted for manual selection for the time being, in ordernot to introduce noise into the annotation process. 27. in our experiment the annotation was done for english by native speakers of german, hence translation was to german. translation could also be into some other language (distinct from the language of the annotation task) as long as it is the same for both annotators. 300 discriminative analysis of fine-grained semantic relations verb pair review(person, material ) – teach(person, person) annotator translationv1 translationv2 relation assigned a1 bewerten (critique) unterrichten (teach)unrelated(review, teach) a2 wiederholen (reexamine) unterrichten (teach)temp. inclusion(review, teach) verb pair cry(person) – be scared(person) annotator translationv1 translationv2 relation assigned a1 schreien (yell) erschrecken (be scared)temp. inclusion(cry, be scared) a2 weinen (weep) sich fürchten (be afraid) unrelated(cry, be scared) table 6: capturing sense distinctions through translationto german. guide the annotators towards a classification using the discriminative properties we established in section 3: properties oftemporal sequenceandbehavior under negation. this question-based annotation scheme naturally extends the enhanced representation of verb pairs using prototypical arguments. in fact, it is dependent on this novel representation. using appropriate placeholders, we generate skeleton sentencesfor the target predicates and their prototypical arguments. these are presented to the annotators, and help them check and decide on the different relation properties that hold for the generated phrases. this novel presentation scheme can thus be considered a compromise between the context-less type-based annotation and the contextrich token-based annotation setups examined in section 4.1. our method is best illustrated using an example. figure 2 displays questions and answer possibilities for annotating the verb pairlearn– speak. using resnik’s selectional association scores, we determinepersonandlanguage as prototypical argument classes for this verb pair. from these abstract representations including predicate, prototypical arguments and prepositions, we generate sample phrases, as seen in questionq0.28 here we elicit translations to german for the given verbs in their typical argument context, to record the interpretations perceived by the annotators. questionq1 is designed to determine the temporal order in which the events typically occur. this question is offered in two ways: by generating the two verb phrases in the respective orders with appropriate temporal conjunctions (and then; at the same time). these options are supplemented with the corresponding fine-grained temporal relation types of allen (1983)’s classification.29 we target a coarse three-way distinctionbefore, afterand during that each encompasses several of allen’s relations. this was determined sufficient for classification and necessary for annotation, given that the annotators also consider borderline cases. using the graphical representations of these relations, we defined a mapping from allen’s relations to three coarse temporal relation classes that we offered to the annotators (cf. appendix ii).30 28. in order to generate natural phrases, we substitute someabstract classes likepersonwith proper names such as john, or language with spanish. 29. the annotation interface allows easy access to an overview of the relation inventory (cf. appendix i). we employed in particular the graphical representation of allen’s relations, which proved to be very helpful for the annotators in order to decide on the appropriate relation. 30. note that the coarse temporal relations ‘before(x,y)’ and ‘after(x,y)’ include the respective overlap conditions where y overlaps with the preceding/following x, whereas weassign ‘during(x,y)’ for all cases where x is fully 301 tremper andfrank q0: // characterizing the interpretation of the events: // please give a translation for the verbslearnandspeakin these readings: x: john learns spanish. translation: y: john speaks spanish. translation: q1: // determining the temporal order of events: // what is the typical order of the following events? a) john learns spanish and then he speaks spanish. x before y: {m, o,<} b) john speaks spanish and then he learns spanish. x after y: {mi, oi,>} c) john learns spanish and he speaks spanish at the same time.x during y: {s, si, f, fi, d, di, =} d) more than one order of events is possible. e) not sure (difficult to define) q2: // determining negation properties: x and y? // john learns spanish. will he speak spanish? a) yes (x and y) b) no (x and¬y) c) maybe (x and y or¬y) – persistence under negation→ presupposition q6: // determining negation properties:¬x and y? // john does not learn spanish. will he speak spanish? a) yes (¬x and y)→ none b) no (¬x and¬y) – cancellation→ presupposition c) maybe (¬x and¬y or y) → none result:pre(speak,learn) figure 2: annotation questions for the verb pairlearn – speak. the next set of questions is designed to elicit inference properties with respect to negation. the verb pairs are presented in sentence pairs consisting of a declarative statement involving the first verb and a subsequent question involving the second verb. this pair inquires whether the second sentence can be assumed to hold true given the first one is considered true.31 in case the annotator has selecteda) x before y, q2 will be chosen as a follow-up question, querying the dependence of y (= speak)’s truth on x (= learn) holding true:x and y?. here the annotator may chosea) yes: x and y if you learn a language, you will (be able to) speak it. but more realistically, he or she should choosec) maybe: x and y/¬y you may or may not be able to speak the language after having studied it.if the latter option is taken, the relation will be a candidate for presupposition (pre(y=speak, x=learn)) as answer c) establishes persistence under negation. at the same time, answer c) excludesentailment(ent(x=learn, y=speak)).32 given answer c) is selected forq2, we further check inference regarding the negation of x. thisis done in questionq6: ¬x and y?. included in y’s interval. these three coarse temporal relations are intended to correspond to the relations ‘precedes’, ‘succeeds’ and ‘overlap’ for temporally related events as used in table 1, p. 287. 31. the order in which x and y are presented as well as their temporal inflection is dependent on the answer to question q1. note further that depending on the relation being considered, x and y may change roles in being considered as trigger verbs, which fill the first argument of the relation. 32. this judgement is dependent on an interpretation oflearn as a non-accomplished process, in the meaning ofstudy. 302 discriminative analysis of fine-grained semantic relations figure 3: decision tree for question-based annotation pre(supposition), ent(ailment), t(e)mp(oral inclusion), ant(onymy), unr(elated). here, answerb) no: ¬x and ¬y (i.e., if you don’t learn a language, you will not speak it) indicates that cancellation of the presuppositionx=learn is valid if y=speakis false. overall, the three consecutive questions displayed in figure 2 establish the pairspeak – learn as an instance of presupposition, under an interpretation of learning as a process. an annotation decision tree. by extending this method to the full inventory of the targeted relation types, we establish a question-based annotation scenario that takes the form of a decision tree, as displayed in figure 3. we are able to differentiate the fiverelations using – in the default case – three questions per verb pair, by exploring their semantic properties, as summarized in table 3. the first questionq1 clarifies the temporal sequence properties of the examined verb pair. the answer to questionq1 also determines the order in which the consecutive tests forinference under negation are presented, e.g.,q3 presents x and y in a different order. this way we capture all relevant orders of verb pairs for the temporally sensitive relation typespresuppositionandentailment, in response to the temporal sequence properties detected in q1.33 questionsq2 to q5 (all at the same level of depth) follow the very same pattern. yet,they are dependent on the temporal properties established by the answer to questionq1, so the answers to these questions differ in view of the relation types they may indicate. similarly, questionsq6 to q8 are structurally equivalent, but given their dependence on the previous questions and answers they will trigger case-specific conclusions as to the predicted relation type. it should now be clear from the structure of the tree that for averb pair such asbuy – ownwe will obtain the classificationent(buy,own) by the following chain of questions and answers: (10) q1: which order? a)x before y → q2: x and y? a) yes:x and y → q6: ¬x and y? b) no: ¬x and ¬y34 33. for instance, the verb pairwin andplay cannot be classified aspresuppositionwith the verbs presented asx=win, y=play. this case is captured by response b) to questionq1, so that the inverted verb pair relation can be tested by q3 (the mirror ofq2), using inverted roles of x and y. 303 tremper andfrank questionsq2 andq3 and their follow-ups are triggered by verb pairs that involve a temporal sequence. they must be checked in both order variants to determineentailmentandpresupposition relations irrespective from the order in which the verb pairs are presented (see footnote 33). questionq4 discriminatesentailmentandtemporal inclusionby testing persistence under negation, similar to what is done forpresupposition. thus, we can establishsleep – snoreastmp(snore,sleep) vs. live – breathasent(live ,breath). antonymyis established for verbs that are not assumed to occur in sequence or concurrently, through answer d) toq1, which yields the value ‘undefined’ for temporal sequence. here it seems sufficient to test for complementarity, brought out by answer b) no toq5: x and y? for verb pairs such aslove – hate. additional questions for antonymy. for some verb pairs questionq1 yielded annotation differences depending on whether the annotators considered a syntagmatic or paradigmatic relation between the verbs. this was encountered in particular for verbpairs that qualify for bothantonymyand presuppositionrelations, such asopen – close, connect – disconnector accelerate – slow (down). therefore, we designed an additional question for the annotators, in case we encountered that one of them had annotated a pair withantonymy, while the other did not. the additional questions presented to the annotator that did not annotate antonymy in thefirst place (here, annotator 1) now focus explicitly on the antonymy relation. in case annotator 1 answers both questions withno, the verb pair will be annotated asantonymy. (11) additional questions targetingantonymy: annotator 1 pre(slow,accelerate) annotator 2 ant(slow,accelerate) → qant1: the car slows down. does this car accelerate? → qant2: the car accelerates. does this car slow down? additional questions for backward entailment. in some cases the entailment relation between verbs can be symmetric, as for the pairdepart – arrive. such pairs should be annotated as entailments in both directions. given the way we set up our hierarchical annotation scheme, each verb pair will only be assigned a single label. therefore, wedesigned additional questions for the annotators, to identify cases of symmetric entailment. these questions take the same form as the original questions (12), but the temporal order is reversed. the answersyes to the first question and no to the second question in (13) assign the backward entailment relation to the verb pair ent(depart,arrive). (12) standard questions targetingentailmentgenerated by the annotation system: ent(arrive,depart) q1: which order? john departs and then john arrives (x after y ) → q3: john departs. will he arrive? a)yes → q7: john doesn’t depart. will he arrive? b) no (13) additional questions targetingbackward entailment: → qent1 (= q2): john arrives. did he depart? yes → qent2 (= q6): john doesn’t arrive. did he depart?no 34. following our argumentation in section 3, we ask the annotators to consider the case of ‘what normally holds’ in a situation if¬x holds true and to disregard exceptional cases that are not relevant for the situation considered. 304 discriminative analysis of fine-grained semantic relations annotation interface. in order to hide the complexity of the decision process from the annotators, this decision tree was implemented in a web-based annotation interface that presents the annotator with novel questions depending on the answers given to the previous question. the annotators were given the possibility to go back and inspect or revise the answers given to previous questions. displays of the annotators’ views for the basic question types are given in the appendix. annotation quality. we evaluated the quality of annotation using this question-based annotation scheme, using 250 verb pairs selected from the dirt collection.35 as the novel annotation scheme is considerably simplified, we also tested it with non-expert annotators. for the two expert annotators we obtained an inter-annotator agreement (iaa) of 72% with a kappa value ofk = 0.64. this is considerably higher compared to the annotation quality we obtained using standard typeor token-based annotation.36 this result clearly indicates that the annotation task could be dramatically simplified, with a large improvement of inter-annotator agreement. however,the decisions to be made still seem too complex for non-expert annotators: we observe poor agreement between the non-expert annotator and either of the expert annotators: iaa = 60%,k = 0.46 and iaa = 64%,k = 0.49. thus, addressing this annotation task by crowd sourcing to non-experts does not seem to be an option in its current design. the distribution of the semantic relations in the final annotated data set is more or less equal. temporal inclusionis slightly under-represented (15%);entailmentandother/unrelatedare slightly over-represented (23% and 25%).37 5. classification of fine-grained semantic relations between verbs this section describes the classification architecture, employed feature sets and classification experiments for sub-classifying fine-grained semantic relations including presupposition. the performance of the classifiers is evaluated against the gold standard annotation set obtained using questionbased annotation, as described in section 4. as a reference for the subsequent description, figure 4 summarizes the classification architectures and feature sets for the experiments described below. 5.1 classification method our aim is to acquire verb pairtypesthat stand in a particular semantic relation from our selected relation inventory:presupposition, entailment, temporal inclusionandantonymy(section 3). the lexical knowledge acquired in this way will be used to enrichtextual occurrences of individually occurring trigger verbs with inferences on the basis of the learned verb relations. for this purpose we build a classifiercdiscr that automatically sub-classifies the relations holding between verb pair candidates into five classes: the four selected semantic relation types and a fifth class that captures verb pairs that stand in no or some other semantic relation not considered here. to classify the verb pairs according to our relation inventory we calculate type-based distributional features and use a supervised classification algorithm to build the model. the type-level 35. this set is distinct from the annotation set used in section 4.1. the annotation set produced in these initial experiments was used as development set in the classification experiments reported in section 5. 36. we did not perform separate evaluations of the impact of prototypical verb arguments and the break-down of annotation decisions in the question-based setting, due to the considerable annotation overhead this would have caused. 37. this does not reflect the natural distribution of these relations, due to some amount of pre-selection for the underrepresented classes. 305 tremper andfrank sample selection:ccntg: labels contiguous ([+contiguous]) corpus samples for feature extraction. fpath−len, fpath: length and form of path of grammatical functions betweenv1 andv2 fcoref : coreference relation holding between subjects/objects of v1 andv2 fdist−tok, fdist−verb: distance betweenv1 andv2 (in tokens and verbs) fconnectives: conjunction or direct grammatical function connectingv1 andv2 type-based classification:cdiscr : x → y assigns classification instancesx consisting of pairs of verb types (v1,v2) one labelr ∈ y. flat: classify instancesx ∈ x into 4 core relation types plus ‘u(nrelated)’:y = { e, p, t, a, u}. instance setx : verb pair typesx ∈ x (selected from dirt (lin and pantel, 2001)). hierarchical: 1st-stagecrel: crel classifies all input verb pairsx ∈ x as [± related]: [−related] if cnt([+contiguous])< cnt([−contiguous]) & temprel =undefined [+related] otherwise. 2nd-stagecdiscr: cdiscr classifies verb pairsx ∈ x classified as [+ related] bycrel target classesy ∈ { e, p, t, a}. feature vectors for classifiercdiscr in flat (5-way) and hierarchical (4-way) classification: compute feature vectors~fx = 〈f0, f1, . . . , fn〉 for all verb pair typesx ∈ x : feature type feature flat hier. typical temp. rel. f0: v ∈ { before, during, after, undef} x x polarity pairs f1 – f4: p (〈±v1,±v2〉 | v1, v2) x x f5: average distance betweenv1 andv2 in tokens x – relatedness f6: pmi for v1 andv2 in verb pairs (v1, v2): pmi(v1, v2) x – f7 – fn: cond. prob. for conjunctionsci: p (ci | v1, v2) x x baselines: cdiscr classifier:fconj: f7 – fn: conditional probability of conjunctionsc given (v1, v2) crel classifier:fconnectives: conjunction or grammatical function relatingv1 andv2. figure 4: summary of classification architectures and feature sets. feature vectors are calculated on the basis of a training setof corpus instances, i.e. sets of sentences involving pairs of verbs that are annotated on the type levelfor the relation that constitutes the classification target. the classifier learns weights for the features on the basis of the annotated training data and makes predictions for unseen verb pairs using the learned model. the performance of the classifier is tested against the set of verb relation labels defined in the gold standard data set. 306 discriminative analysis of fine-grained semantic relations classifier definition. we definea type-based classifiercdiscr: x → y that receives as input a set of instancesx ∈ x of verb pair types(v1, v2) and a set of the feature vectors~fx = 〈f1, f2, . . . , fn〉 calculated for any verb pairx under consideration.c returns one of the target class labelsr ∈ y. we experiment with two classification architectures:flat andhierarchical classification. in the flat classificationarchitecture, the classifiercdiscr distinguishes all five relation types including theunrelatedclass. inhierarchical classificationwe first partition the instance set of candidate verb pairs into two classes:relatedvs.unrelated, with the first class covering the four selected semantic relationsp(resupposition), e(ntailment), t(emporal inclusion)anda(ntonymy). in a second classification step,cdiscr performs 4-way flat classification for these four relation classes, taking as input the verb pair candidates that were classifiedas [+ related] by the first stage classifier. detailed information on the setup of these architectures isgiven in sections 5.3.4 and 5.3.5. 5.2 features for classification the discriminative semantic relation classifiercdiscr relies on the three groups of features motivated in section 3.2:temporal sequence, behavior under negationandcontextual relatedness. 5.2.1 temporal sequence our analysis of relation properties (cf. table 3) reveals that some of our target semantic relations involve a typical temporal order, while others do not. we make this property available for discriminative classification by defining a type-based featuretypical temporal orderwhich records the temporal relation that can be considered typical for a givenverb pair. we distinguish three basic temporal relationsbefore, after andduring, plus undefined, in case no typical temporal sequence can be determined. we obtain this information from a (token-based) temporal relation classifier. detecting and classifying temporal relations holding between verbs in context is a difficult task.38 in contrast to thetempevalchallenges (verhagen et al., 2010), we use a coarse relation inventory that is sufficient for our purposes. moreover, as our aim is to predict type-level temporal relation properties, we will rely on a subset of confident, i.e., reliable, token-level classifications. a token-based temporal relation classifier. for token-based temporal relation classification we define a variety of morpho-syntactic and semantic features,including tense, aspect, modality, auxiliaries, conjunctions, grammatical function paths, adverbial adjuncts, order of appearanceand verbnet classes (same/subsumed or different).39 40 this extends the feature set used by chambers et al. (2007) for temporal relation classification in context. we built a token-level temporal relation classifier that we trained on a set of manually annotated contexts, 200 contexts for each relation, using the three target temporal relations. using the above feature set we trained a 3-way bayesnet classifier41 for classification on the token level, with the target classesbefore, afterandduring. we evaluated this classifier using a set of manually annotated contexts, 20 contexts for each relation and achieved an f1-score of 84.3% on this set.42 38. see e.g. chambers et al. (2007), bethard and martin (2007), lee (2010). 39. the verbnet class feature is used as an indicator of temporal inclusion, in particular for the troponymy relation. 40. we use all verbnet classes except forothercos-45.4, which includes many opposite verbs, e.g.accelerate, slow. 41. we use the bayesnet algorithm implemented in weka (witten and frank, 2005). 42. it is difficult to compare the performance of this specially designed temporal relation classifier to results reported on the timebank corpus, because of the different temporal relation inventories used: while we are using a coarse set of relations, the relation set used in the tempeval challenges is more fine-grained (it distinguishes 6 relations). 307 tremper andfrank predicting a type-based ‘typical’ temporal relation. we predict a type-based ‘typical’ temporal relation for any pair of verbs, relying only on confident token-level classifications (threshold 0.75). the score for each relation is computed as the association between a verb pair (v1, v2) and the assigned temporal relation instances in context, by applying pmi (point-wise mutual information): pmi((v1, v2), t emp rel) = log p ((v1,v2),t emp rel) p ((v1,v2))p (temp rel) for any given verb pair we choose the relation that obtains the highest pmi score. if pmi does not indicate a typical temporal relation (we set a thresholdof 0.4, optimized on a held-out data set43), we assign the labelundefined. the quality of this type-level temporal relation classifierwas evaluated using the answers to the first question (q0) of our question-based annotation scenario as a gold standard.44 on this set it achieves an f1-score of 73%, with balanced precision and recall at 71% and 74%, respectively. 5.2.2 negation a token-based polarity labeler. to determine the behavior under negation for given verb pairs, we first need to correctly recognize the polarity of verbs in agiven context. we use a number of triggers to detect negative polarity contexts: negative particles (e.g.not/n’t); negative adverbs (e.g. never); negative adjectives (e.g.impossible) and negative verbs (e.g.refuse).45 in case we detect a single negation trigger for a verb in a sentence, the verb polarity isnegative. if we find a combination of triggers (e.g.never refuse) and the number of triggers is even, we assign the valuepositive, if the number is odd, the verb polarity isnegative. we also use a small set of adverbs that are able to switch a verb’s polarity in case it isnegative (e.g.badly, etc.). an example is given in (14). here, the negation trigger refers to the verbplay, but due to the combination with two negative triggers, we assign the polarity positive. (14) we wanted to win the third test as a matter of pride anddidn’t play badly but every time new zealand came into our 22 they scored. to evaluate the quality of the polarity labeler, we manuallyannotated the polarity of 200 verbs in context.46 on this set we achieve an f1-score of 85%, with 84% precision and 86% recall. computing type-based polarity co-occurrence features. for type-based classification of verb pair polarity co-occurrences, we compute a negation vector~fneg with four polarity co-occurrence features for the different combinations:〈±v1,±v2〉. we compute the values of these features using the conditional probability of a given polarity co-occurrence combination for any verb pair (v1, v2).47 another factor which influences our results positively is that we apply the classifier on contexts labeledcontiguous in our corpus preprocessing phase. that is, we compute temporal relations only for closely co-occurring verb pairs in contiguous syntactic contexts. 43. the held-out data consists of 100 manually annotated verb pairs. 44.q0: which is the typical order of the following events? 45. we employ a manually compiled list of trigger predicatescollected from various lexical resources. 46. a subset of 100 contexts of verb pairs that were previously annotated with a semantic relation on the context level. 47. the probability is computed using the set of verb pairs incontext that are labeled [+contiguous], see section 5.3.2. 308 discriminative analysis of fine-grained semantic relations ~fneg = 〈f0, f1, f2, f3〉, with: f0 = p (〈+v1,+v2〉|v1, v2) f1 = p (〈−v1,+v2〉|v1, v2) f2 = p (〈+v1,−v2〉|v1, v2) f3 = p (〈−v1,−v2〉|v1, v2) 5.2.3 relatedness although temporal relation and negation properties can be considered discriminative for identifying our core semantic relation types, they are not sufficient forcorpus-basedclassification. for example, in (15) win and losestand in an antonymy relation, but both verbs have positive polarity. so, the evidence found in the corpus does not always correspond to the properties captured in table 3. (15) winor lose, you pay nothing. detecting relatedness features for corpus-based classification. thus, we include a third dimension of features that record surface-level properties of underlying linguistic properties, as in this case, where semantic opposition is not expressed by opposite polarity, but via the conjunctionor. contextual relatedness features will prove particularly useful for distinguishingantonymyfrom other relation types, especially theunrelatedclass. recall also that the discriminative relation properties that we established do not include the necessary distinction betweensemantically relatedvs. unrelatedverb pairs. for theunrelatedclass, we find a broad variety of syntagmatic properties, while for the core semantic relations we find more characteristic contextual relatedness features. type-based relatedness features.as type-based syntagmaticrelatednessfeatures we employ surface-level information aboutdistanceand connectingconjunctionsbetween verbs,48 as well as distributional association measures, such as point-wise mutual information (pmi). these are raised to the type level in the following way (see also figure 4):49 fdist: average distance between two verbs in tokens within a sentence fpmi : pmi calculated for the two verbs in a given verb pair:pmi(v1, v2) fconj: conditional probabilities for conjunctionsci given specific verb pairs (p (ci|v1, v2)) 5.3 experiments and results 5.3.1 data sets all candidate verb pairs that are presented to the classifierare selected from a set ofsemantically related verbsfrom the dirt collection (lin and pantel, 2001). training set. as training set we employ a small number of seed verb pairs (3 to 6 for each semantic relation) that was used in previous experiments (tremper, 2010). we extended this data set with 30 additional verb pairs which were manually annotatedby two annotators using our novel question-based annotation method (see section 4). the overall set of 48 verb pairs yields a nearly uniform distribution of classes. 48. we manually grouped the most informative conjunctions to a set of 21 conjunction variants, collapsing, e.g. while/whilst, cause/because, to/in order to. strongest conjunctions areor, when, if, but, by. 49. all values are calculated on the set of verb pairs in context that are labeled [+contiguous] (see section 5.3.2), except for fpmi , which was calculated on the basis of the full ukwac corpus. 309 tremper andfrank test set. as our gold standard test set we employ the annotation set consisting of 250 verb pairs that was produced using the question-based annotation setup. the distribution of relations over the 250 verb pairs is as follows:presupposition: 18%,entailment: 23%, temporal inclusion: 15%,antonymy: 19%,other/unrelated: 25%. corpus instances. for the computation of type-based relation features we obtained corpus samples from the ukwac corpus (baroni et al., 2009). we extracted allsentences in which both verbs of a verb pair co-occur, considering sentences of up to 60 tokens in length. the number of contexts available for each verb pair ranges from 30 to about500 instances. 5.3.2 preprocessing: selecting informative samples forfeature extraction to avoid noise in the computation of type-based feature vectors we need to select informative corpus instances of co-occurring verbs that stand in a close syntagmatic relation. to this end, we perform a preprocessing step that selects context samples of co-occurring verbs that are contiguously related. we designed the following set ofcontiguity features that record different types of indicators for syntagmatic relatedness of co-occurring verbs. fpath−len, fpath: length and form of the path of grammatical functions relatingv1 andv2 fcoref : coreference relation holding between subjects and objects of v1 and v2 (coreferent subjects or objects; subj coreferent w/ object; no coreference) fdist−tok, fdist−verb: distance betweenv1 andv2 (in tokens and verbs) fconnectives: subordinating or coordinating conjunction, or else direct grammatical function connectingv1 andv2 using this feature set, we constructed a classifierccntg that labels verb pairs appearing in corpus sentences as [± contiguous]. the classifier was trained and tested on a manually annotated set of contexts involving our seed verb pairs (2343 contexts from which 90% were used for training and 10% for testing).50 best results were achieved using the j48 decision tree algorithm51 (f1-score: 79.3%). we applyccntg on the set of unlabeled verb pairs in context and select all contexts that were confidently labeled as [+contiguous] (above threshold 0.75) as corpus samples for computing the feature vectors for the relation classifiercdiscr. classifications obtained from the contiguity classifier are further used as a feature for the relatedness/non-relatedness classification in the hierarchical classification scenario (see section 5.3.5 for more detail). 5.3.3 learning algorithms for our main classification task we experimented with different classification algorithms and achieved best results using bayesnet. thus, unless noted otherwise,all results reported below were obtained using the bayesnet classifier implementation of weka (witten and frank, 2005). 5.3.4 experiment i: flat classification setup. experiment i performs classification using theflat classification architecture, which assigns class labels for all five relation classes including the unrelated class:y ∈ { p(resupposition), 50. inter-annotator agreement was 81%, with a kappa value of0.72. 51. weka implementation of the c4.5 decision tree algorithm(witten and frank, 2005) 310 discriminative analysis of fine-grained semantic relations semantic relation precision recall f1-score baseline f1-score presupposition 41% 45% 43% 25% entailment 47% 43% 44% 25% temporal inclusion 38% 47% 42% 26% antonymy 68% 71% 70% 47% other/unrelated 54% 53% 54% 12% all 50% 51% 51% 27% table 7: results for experiment i: flat classification (baseline: best feature:fconj: conjunctions). e(ntailment), t(emporal inclusion), a(ntonymy), u(nrelated)} (cf. figure 4). for each verb pair in our training and test sets we compute feature vectors as described in section 5.2. evaluation results. table 7 displays the results, evaluated against the test data set. the classifier performance is compared against a baseline that uses the best single featurefconj: conjunctions. the classifier outperforms the baseline for all relation types, with balanced precision and recall. precision is higher than recall forentailment. for presupposition, entailmentandantonymyrecall exceeds precision. with an overall f1-score of 51% the classification performance is still modest, however the difference between the chosen baseline and our model is significant (ρ < 0.05). note further that the average f1-score for the more complex inferential relations (p, e, t) is lower at around 43%, while forantonymyit is at 70%. 5.3.5 experiment ii: h ierarchical classification setup. as an alternative to flat classification, we investigate a hierarchical architecture that first separates related from non-related verb pairs, and subsequently sub-classifies related verb pairs into the four relation classes:p(resupposition), e(ntailment), t(emporal inclusion)anda(ntonymy). a binary classifier crel separatesrelated from non-related verb pairsusing as criterion (i) the ratio of contexts for a given verb pair classified as [±contiguous] by the contiguity classifier in sample selection (see section 5.3.2) and (ii) the typicaltemporal relation calculated by the type-based temporal relation classifier. we assign the label [−related] to a verb pair if the majority of contexts are annotated as [−contiguous] and there is no typical temporal relation for this verb pair (temprel =undefined). crel: classify all input verb pairsx ∈ x as [± related]: [−related] if count([+contiguous])< count([−contiguous]) & temprel =undefined [+related] otherwise. thesecond-stage discriminative relation classifiercdiscr takes as input all verb pairs classified as [+related] by the first-stage classifier and performs 4-way classification into the set of relation classesy = { p(resupposition), e(ntailment), t(emporal inclusion), a(ntonymy)}. since the unrelated class has already been separated in the first classification step, the classifier does not make use of the relatedness featuresf5: average distance between two verbs 311 tremper andfrank 1st-stage classifiercrel: related vs. unrelated classification precision recall f1-score baseline f1-score unrelated 82% 67% 74% 57% related 72% 84% 77% 54% 2nd-stage classifiercdiscr: 4-way semantic relation classification (oracle input) precision recall f1-score baseline f1-score presupposition 62% 50% 56% 30% entailment 53% 49% 51% 33% temp. inclusion 44% 62% 52% 25% antonymy 76% 80% 78% 63% all 59% 60% 59% 38% table 8: exp. iia: individual classifier performance for hierarchical classification (with oracle). baselines: best features: step 1:fconnectives: connectives; step 2:fconj: conjunctions. andf6: pmi(v1, v2), as these are designed to distinguish unrelated from related verb pairs. however, the conjunction features are considered useful for discriminating the core semantic relations, and are thus included as a feature in this classification step (cf. figure 4). evaluation setup. for experiment ii we perform evaluations for both classification steps, using adapted gold standard data sets: (i) classifications for thefirst-stage binary classifierare evaluated against a test data set compiled from the gold standard test set used in experiment i. itconsists of the set of all unrelated verb pairs (58) and the same amount of (randomly selected) related verb pairs. (ii) for the classifications for thesecond-stage classifiercovering 4 relation classes, the test set forms the subset of the standard test data set that consists of the related verb pairs only.52 we report two evaluations for hierarchical classification.for both, we use best-feature baselines for the individual classifiers: the best featurefconj for the discriminative relation classifiercdiscr as in experiment i, and the best featurefconnectives for the relatedness classifiercrel. experiment iia: individual classifier performance. table 8 analyzes the performance of the individual classifiers, where the second-stage classifier is based on perfect input, i.e. oracle classifications from the first-stage classifier. the binary classifier crel obtains an f1-score of 74% for the unrelated class, which clearly outperforms the best feature baseline by a margin of 14 points f1-score. while the related class is recognized with higher f1-score of 77%, we favour the results for the unrelated class,which is higher in precision (82% vs. 72%). generally, misclassifications of the first-stage classifier impede the overall performance of the cascaded classificationarchitecture, so while high precision is beneficial, the weaker recall (67%) could still impact the overall results. 52. the distribution in this reduced data set is:presupposition: 25%, entailment: 31%, temporal inclusion: 20%, antonymy: 24%. 312 discriminative analysis of fine-grained semantic relations semantic baseline flat classification hierarchical classification relation f1-score precision recall f1-score precision recall f1-score presupposition 25% 41% 45% 43% 50% 46% 48% entailment 25% 47% 43% 44% 44% 46% 45% temp. incl. 26% 38% 47% 42% 41% 47% 44% antonymy 47% 68% 71% 70% 72% 74% 73% unrelated 12% 54% 53% 54% 68% 63% 66% all 27% 50% 51% 51% 55% 55% 55% table 9: exp. iib: hierarchical classification (pipeline) –results contrasted with flat classification (baseline: best feature:fconj: conjunctions). evaluating theflat 4-way relation classifiercdiscr on oracle classificationswe obtain an overall performance of 59% f1-score.53 experiment iib: full hierarchical classification. table 9 presents the results for full hierarchical classification, with system input for the second-stage classifier.54 for convenience, the results are aligned with the results obtained for flat classificationin experiment i. with an overall f1-score of 55%, hierarchical classification significantly outperforms the baseline (ρ < 0.05). it also outperforms flat classification, but not significantly at a significance level of 5%. we observe performance gains for all relations, which are highest forpresupposition(+5 points f1-score) andunrelated/other (+12 points f1-score). again,antonymyperforms best. among the inferential relations,presupposition scores highest with 48% f1-score and the highest precision at 50%. 5.4 analysis of results 5.4.1 impact of individual features we measured the impact of individual feature classes on the results, using ablation testing for different feature groups (cf. figure 4):55 negationfeatures,temporal sequencefeatures andrelatedness features. as only theconjunctionsfeature was used in both settings, this was the only relatedness feature we omitted. the outcome, displayed in table 10, nicely underlines the observations made in our analysis of the relation properties. the results56 show that temporal sequence properties are the most important feature forentailment, presuppositionand temporal inclusion, whereas forantonymyand theunrelated/otherclass theconjunctionsfeature has the strongest effect. eliminating conjunctions causes an overall drop to 30% (35%) f1-score, forantonymyeven to 15% (14%). eliminating the temporal relation features incurs a drop to 32% (34%) with the biggest loss fortemporal inclusion: −12 (−11) points f1-score. eliminating the negation features shows only a small impactof about 3–5 points in f1-score. 53. these figures cannot be directly compared to the flat classification results of experiment i, which were computed over 5 classes (table 7). however, the overall tendencies are similar. 54. to enhance precision, we relied on the classifications for the unrelated class as input for the second-stage classifier cdiscr. 55. for hierarchical classification we performed the ablation testing only for the second-stage classifier. 56. in ablation testing, lower results indicate higher importance of the feature (group) that has been omitted. 313 tremper andfrank semantic exp i: flat classification exp iib: hierarchical classification relation all w/o neg w/o tmp w/o conj all w/o neg w/o tmp w/o conj presupposition 43% 37% 24% 35% 48% 41% 22% 34% entailment 44% 41% 14% 28% 45% 43% 14% 25% temp. incl. 42% 42% 12% 38% 44% 43% 11% 36% antonymy 70% 64% 64% 15% 73% 68% 59% 14% other/unrelated 54% 47% 45% 35% all relations 51% 46% 32% 30% 55% 52% 34% 35% table 10: ablation testing: f1-score results using different feature sets (exp i & iib). verb pair gold flat hierarchical (w/ prototypical arguments) standard classification classification abandon(person, do sth) – try(person, to do sth) pre pre pre fly(plane) – land(plane) ent ent ent multiply(person, numbers) – calculate(person, solution)tmp ent ent cry(person) – laugh(person) ant ant ant enter(person, house) – open(person, door) pre ent unr boil(water) – evaporate(water) ent unr unr steal(product) – take(product) tmp ant ant table 11: examples of correct and wrong classifications. although the weakest feature type, with overall 5 points loss in f1-score, the negation features clearly contribute to overall performance. interestingly, they have the strongest effect forpresupposition, with a drop of 6–7 points in f1-score. this clearly reflects the specific negation properties found with presupposition. this analysis corroborates that while the negation properties are very important for language understanding and logical inference, and proved effective as a guide for human annotation, a corpus-based classification approach needs to complement its effects, as human language often resorts to other means for expressing negative polarity, such as the use of conjunctions (or, whereas, etc.), or does not make it explicit at all. 5.4.2 classification examples and divergences table 11 displays examples of correct and wrong classifications for both architectures. the verb pairs are given with the prototypical arguments that were used for the gold standard annotation. 5.4.3 error analysis the most frequent errors we observe (especially for the flat architecture) are misclassifications between related and unrelated verb pairs and betweenpresuppositionand entailment. this points to weaknesses of contiguity features used in the contiguoussample selection step and of negation features used for the main classification. we also notice thatentailmentis often misclassified 314 discriminative analysis of fine-grained semantic relations as temporal inclusion. misclassifications betweenpresuppositionand temporal inclusionare rare compared to other relations. this indicates that the temporal sequence features are effective. as further major error sources we identified problems with verb ambiguity and coreference resolution. both of them affect the detection of semantic relations as being related vs. unrelated.57 finally, we identified errors in selecting contexts from theukwac corpus, which are used for computing the distributional features for the main classification. inspection of a small section of corpus samples shows that erroneous annotations of nouns oradjectives as verbs cause errors in the computation of the type-based feature vectors. we have solved the problem of erroneous annotations of adjectives as verbs by double checking the dependencies between verbs and nouns,58 but we still need to address the problem of erroneous annotations of nouns as verbs. regarding classification architectures, hierarchical classification outperforms flat classification for all relation types, and especially for theunrelatedclass. thus, the first-stage classifier that separates related from unrelated verbs implements a strongfilter. the pipeline architecture still suffers from a performance loss due to misclassifications ofthe first-stage classifier. this problem can be addressed in future work by using a joint classification approach. 5.5 comparison to related work related work on semantic relation classification differs from our approach in a variety of respects (see section 2). nevertheless we compare our results to whatcould be achieved there, to give an idea about the state of the art on comparable and related tasks. closest to our work is verbocean. chklovski and pantel (2004) apply a semi-automatic patternbased approach for extracting fine-grained semantic relations between verbs (similarity, strength, antonymy, enablementandhappens-before). this inventory is different from ours, especially it does not include relations such as entailment and presupposition with complex inferential behavior. for a sample of 100 automatically labeled verb pairs they determined a precision of 65.5%.59 results for recall and f1-score were not reported. the only common class of semantic relations used by both approaches isantonymyor opposition. we investigated the verb pairs which are labeled with this class for both systems, comparing to our gold standard test set. most antonyms are annotated byboth systems correctly. evaluating both systems against our test set yields 71% precision and 35% recall for verbocean. with 72% precision and 74% recall our system achieves better recall and overall more balanced results. examples of verb pairs which could not be found in verbocean are(hide, show)or (multiply, divide). some of the antonyms were annotated in verbocean with the classsimilar, e.g.(marry, divorce)or (play, work). we also find some verb pairs for which verbocean performs better than our system, e.g. (catch, miss). with an overall f1-score of 55% with balanced precision and recall obtained on a more difficult and more balanced data set, our results can beconsidered competitive. inui et al. (2005) perform classification of causal relations for japanese. they report high precision and recall results (95% precision forcause, precondandmeansrelations with 80% recall and 90% precision foreffectwith 30% recall). they emphasize that the framework can be applied to other languages, such as english, but no experiments are presented in the paper. 57. for coreference resolution we employed the stanford corenlp resolver (lee et al., 2011) – the system that performed best in the 2011 conll shared task on coreference resolution. 58. we check for the presence of the stanford parser dependency amod (adjectival modifier) between a verb and a noun (de marneffe et al., 2006) as an indicator of erroneous annotation. 59. only 2 and 8 pairs were evaluated forenablementandantonymy, respectively. 315 tremper andfrank pekar (2008) performs acquisition of verb entailment rules. his main focus is on detecting asymmetric relations between verbs and relating their argument positions, without trying to subclassify the obtained verb pairs into different relation types. his method is based on co-occurring verbs within locally coherent text and measures their asymmetric dependence using an information theoretic approach. a precision of 71% is reported, but no recall and f1-score. acquisition of verb entailment rules is also the aim of aharon et al. (2010). they acquire inference rules from the framenet resource using frame-to-frame relations and induce argument mappings for the related predicates. the obtained rules aretested against ace events. performance results are mixed, with modest precision and very low recall: precision: 55.1%, recall: 17.6%, f1-score: 24.6%. berant et al. (2010) explore graph optimization using integer linear programming (ilp) in order to find the best set of entailment rules under a transitivity constraint. the approach is restricted to entailment relations. they obtain balanced precision and recall at 69.6% and 67.3%, respectively. their work establishes that global methods outperform local methods for learning entailment relations. berant et al. (2012) offer extensive evaluationand further refinements of this method. weisman et al. (2012) use a large set of linguistically motivated features to acquire verb entailment rules. this feature set is designed to extract a wide spectrum of rules, therefore the system achieves a good recall of 71% with a moderate precision of 40%for the recognition of entailment rules. no attempt is made to distinguish the different relation types acquired by the system. 6. summary and conclusions in this contribution we presented a corpus-based approach for discriminative analysis and classification of fine-grained semantic relations between verbs. the set of relations we consider comprise the non-taxonomic inferential relationsentailment, presuppositionandtemporal inclusion, and the taxonomic relationsantonymyandtroponymy. we grouptemporal inclusionandtroponymygiven they have similar inferential properties, and excludedsynonymyas a result of the nature and technicalities of our corpus-based approach. to the best of our knowledge, we are the first to investigate presuppositionin a corpus-based lexical semantic relation acquisition task. the focus of this paper was to analyze the underlying properties of the selected relations, the design of features for a corpus-based learning approach, and to discuss possibilities for the annotation of such fine-grained semantic relations. we present experiments for automatic classification of the target relations with evaluation against the gold standard data set we constructed. in contrast to prior work, we present an in-depth analysis ofthe relations we aim to sub-classify, including a characterization of their inferential behavior. we determine a small set of differentiating properties relating to negation and temporal sequence properties. these do not only provide differentiating features for classification. they are also essential for appropriate inference in context, which is the ultimate goal of our work. inclusion of the presupposition relation is what clearly distinguishes our work from the state of the art in this area, which primarily focuses on the discovery of entailment relations proper. the acquired pairs of presupposition-triggering verbs and their presuppositional relata encode valuable commonsense knowledge about typical verb sequences and preconditions holding between events, such asplay – win, read – cite, learn – masteror hire – fire. these are not broadly covered in verb lexicons such as wordnet and only found with selected scenario frames in framenet (fillmore et al., 2003). related work by chambers and jurafsky (2008, 2009), which aims at acquiring 316 discriminative analysis of fine-grained semantic relations typical event sequences from large corpora, detects frequently occurring verb pairs, yet does not differentiate between fine-grained relation types. regneri et al. (2010) learn script-like knowledge using sequences of events they gathered from crowd-sourcing. however, script-like knowledge is only applicable to a small set of typical event chains. theirwork relies on gathering event sequences for pre-specified situation types. our approach is more general, as it is able to learn presupposition and other clearly distinguished inferential relations holding between any verb pairs, using a small set of manual annotations. finally, our work targets the temporal and inferential differences between the various relation types that are crucial for applying the learned relations in context and for drawing valid inferences. our analysis shows that the selected relations can be fully discriminated by their inferential and temporal properties. however, this does not mean that automatic or manual labeling of such verb relations is a trivial task. the classification of fine-grainedsemantic relations between verbs presents a major challenge, due to complicating factors such as verb ambiguity, coreference of arguments and the complexity and subtlety of the inference propertiesassociated with such relations. this was clearly brought out by our initial annotation experiments that followed traditional annotation strategies: type-based annotation forces annotators to considercomplex inferential patterns for (pairings of) different verb meanings out of context; token-based annotation is difficult because the contexts are often involved, with shades of meaning that make decisions difficult. moreover, acquiring sufficient numbers of context-based annotations is expensive, and it is difficult to ensure that all relevant readings are appropriately represented. we therefore designed a novel annotation setup that addresses the specific problems we identified: (i) controlling for verb readings and ambiguity, (ii) the need for abstraction from specific contexts and (iii) the need to reduce the complexity of the inferential patterns that need to be checked. the first two problems are addressed by providing verb pairs with prototypical arguments derived from selectional preference classes. from these representations we automatically generate skeleton sentences offered to the annotators. this restricts the interpretation of the verbs and at the same time provides sufficient generalization from particular contexts. the third problem is addressed by designing a question-based annotation scheme. the complex annotation decisions are broken down to basic decision units and are presented in the form of automatically generated skeleton phrases, with placeholders for prototypical arguments. with this novel setup, we obtain reliable inter-annotator agreement and are able to create a gold standard for evaluating fine-grained semantic relation classification. our novel question-based annotation scheme relieves the annotator from considering several non-trivial decisions in a single annotation step, and thus holds potential for crowd-sourcing the annotation task to non-experts, in order to acquire larger annotated data sets. however, presenting our task to a non-expert annotator did not confirm these expectations. more adaptations are needed to open up this task for crowd-sourcing. having successfully addressed the difficulties of manual annotation, we presented a method for corpus-based acquisition of fine-grained semantic relations between verbs, embedded in a discriminative classification task. the classification model is inspired by the temporal and inferential properties we established for the targeted relations, and are enhanced with corpus-based features designed to detect surface contiguity and semantic relatedness of verb co-occurrences. the classification makes use of type-based distributional features that are generalized from corpus samples. for this reason, the annotation of training andtest data sets can rely on type-based annotations that can be quickly acquired – now that the annotation process has been clarified. our 317 tremper andfrank classification model achieves good results with a small training set comprising about 10 verb pairs per relation. we proposed two classification architectures: flat and hierarchical classification. hierarchical classification outperforms flat classification by a margin of4 points in f1-score, though not significantly. both classification architectures achieve significant performance (up to 100% improvement) over a best-feature baseline. these results are still open for improvement, but with an overall performance of 55% f1-score we are able to show that – despite the considerable complexity of the task – both manual and automatic classification are feasible. the individual results indicate thatpresupposition, entailmentandtemporal inclusionare more difficult to classify thanantonymy; we also foundpresuppositionto outperformentailment, yielding higher precision. this effect might be due to the more prominent specific negation properties associated withpresupposition. closer investigation of the feature impact shows that temporal properties are most effective for the recognition of the inferential relationspresupposition, entailmentandtemporal inclusion, while relatedness features are strongest forantonymyand theunrelated class. the negation features are most effective for identifyingpresupposition. the analysis of the experiment results offers avenues for further enhancements. coming up with better solutions for sense disambiguation and coreferenceresolution could help to eliminate major sources of observed errors. elimination of noise in preprocessing could further improve the results. the hierarchical classification architecture still suffers from error propagation effects that could be reduced through a collective classification approach. finally, with only 10 seed verb pairs per relation our current model is weakly supervised. given thatwe do not require extensive annotations on the token level, adding more verb pairs for training couldfurther improve the results. in future work we will apply the learned relations to triggerverbs appearing in context to infer implicit information. for the proper usage of the acquired inference rules we need to disambiguate the candidates for trigger verbs. while prior and current work on textual inference focusses on entailment, we consider in particular the presupposition relation, which is ubiquitous in texts and subject to special conditions regarding temporal sequenceand negation properties. acknowledgements this work represents a completely revised and extended version of a paper presented at the 2011 dgfs-workshop“beyond semantics: corpus-based investigations of pragmatic and discourse phenomena”. we are grateful for comments and suggestions by the workshop participants, as well as valuable feedback by the anonymousreviewers and the editors. particular thanks go to our annotators: katarina boland, lukas funk, jan pawellek and carina silberer. references ben roni aharon, idan szpektor, and ido dagan. generating entailment rules from framenet. in proceedings of the acl 2010 conference short papers, pages 241–246, uppsala, sweden, 2010. james f. allen. maintaining knowledge about temporal intervals. incommunications of the acm, pages 832–843. acm press, november 1983. marco baroni and alessandro lenci. how we blessed distributional semantic evaluation. inproceedings of the gems 2011 workshop on geometrical models of natural language semantics, pages 1–10, edinburgh, uk, 2011. 318 discriminative analysis of fine-grained semantic relations marco baroni, silvia bernardini, adriano ferraresi, and eros zanchetta. the wacky wide web: a collection of very large linguistically processed web-crawled corpora. injournal of language resources and evaluation, volume 43(3), pages 209–226, 2009. jonathan berant, ido dagan, and jacob goldberger. global learning of focused entailment graphs. in proceedings of the 48th annual meeting of the association for computational linguistics (acl 2010), pages 1220–1229, uppsala, sweden, 2010. jonathan berant, ido dagan, and jacob goldberger. learningentailment relations by global graph structure optimization.computational linguistics, 38(1):73–111, 2012. steven bethard and james h. martin. cu-tmp: temporal relation classification using syntactic and semantic features. inproceedings of the 4th international workshop on semantic evaluations, semeval ’07, pages 129–132, prague, czech republic, 2007. johan bos. implementing the binding and accommodation theory for anaphora resolution and presupposition projection.computational linguistics, 29(2):179–210, 2003. nathanael chambers and dan jurafsky. unsupervised learning of narrative event chains. inproceedings of the acl/hlt 2008 conference, pages 789–797, columbus, ohio, 2008. nathanael chambers and dan jurafsky. unsupervised learning of narrative schemas and their participants. inproceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural languageprocessing of the afnlp, pages 602–610, suntec, singapore, 2009. nathanael chambers, shan wang, and dan jurafsky. classifying temporal relations between events. in proceedings of the 45th annual meeting of the association ofcomputational linguistics (acl 2007), prague, czech republic, 2007. timothy chklovski and patrick pantel. verbocean: mining the web for fine-grained semantic verb relations. inproceedings of the 2004 conference on empirical methods in natural language processing, pages 33–40, barcelona, spain, 2004. david r. clausen and christopher d. manning. presupposed content and entailments in natural language inference. inproceedings of the 2009 workshop on applied textual inference, aclijcnlp 2009, pages 70–73, suntec, singapore, 2009. jacob a. cohen. a coefficient of agreement for nominal scales. educational and psychological measurement, pages 37–46, 1960. ido dagan, bill dolan, bernardo magnini, and dan roth. recognizing textual entailment: rational, evaluation and approaches.natural language engineering, 15(special issue 04):i–xvii, 2009. christian fellbaum, editor.wordnet: an electronic lexical database. mit press, 1998. charles j. fillmore, christopher r. johnson, and miriam r. l. petruck. background to framenet. international journal of lexicography, 16(3):235–250, 2003. john r. firth. a synopsis of linguistic theory 1930–1955.studies in linguistic analysis, pages 1–32, 1957. 319 tremper andfrank anette frank and sebastian pádo. semantics in computational lexicons. in claudia maienborn, klaus heusinger, and paul portner, editors,semantics: an international handbook of natural language meaning, volume 3 ofhsk handbooks of linguistics and communication science series, pages 2887–2917. mouton de gruyter, 2012. l. t. f. gamut.logic, language, and meaning: introduction to logic. university of chicago press, chicago, 1991. michael l. geis and arnold m. zwicky. on invited inferences.linguistic inquiry, 2(4):561–566, 1971. bart geurts and david beaver. presuppostion. in claudia maienborn, klaus heusinger, and paul portner, editors,semantics: an international handbook of natural language meaning, hsk handbooks of linguistics and communication science series. mouton de gruyter, 2012. roxana girju, adriana badulescu, and dan moldovan. automatic discovery of part-whole relations. computational linguistics, 32:83–135, 2006. marti hearst. automatic acquisition of hyponyms from largetext corpora. inproceedings of the fourteenth international conference on computational linguistics (coling), pages 539–545, nantes, france, 1992. verena henrich and erhard hinrichs. standardizing wordnets in the iso standard lmf: wordnetlmf for germanet. inproceedings of the 23rd international conference on computational linguistics (coling), pages 456–464, beijing, china, 2010. takashi inui, kentaro inui, and yuji matsumoto. acquiring causal knowledge from text using the connective marker tame.acm transactions on asian language information processing(talip), 4(4):435–474, 2005. hans kamp and uwe reyle.from discourse to logic. introduction to modeltheoretic semantics of natural language, formal logic and discourse representation theory. kluwer, dordrecht, 1993. lauri karttunen. simple and phrasal implicatives. inproceedings of *sem: the first joint conference on lexical and computational semantics, pages 124–131, montréal, canada, 2012. claudia kunze and lothar lemnitzer. germanet representation, visualization, application. inproceedings of the 3rd international language resources and evaluation conference (lrec’02), pages 1485–1491, las palmas, canary islands, 2002. chong min lee. temporal relation identification with endpoints. inproceedings of the naacl/hlt 2010 student research workshop, pages 40–45, los angeles, california, 2010. heeyoung lee, yves peirsman, angel chang, nathanael chambers, mihai surdeanu, and dan jurafsky. stanford’s multi-pass sieve coreference resolution system at the conll-2011 shared task. inproceedings of the fifteenth conference on computational natural language learning: shared task, pages 28–34, portland, oregon, usa, 2011. stephen c. levinson.pragmatics. cambridge: cambridge university press, 1983. 320 discriminative analysis of fine-grained semantic relations dekang lin and patrick pantel. discovery of inference rulesfor question answering.natural language engineering, 7:343–360, 2001. john lyons.semantics. cambridge university press, 1977. marie-catherine de marneffe, bill maccartney, and christopher d. manning. generating typed dependency parses from phrase structure parses. inin proc. intl conf. on language resources and evaluation (lrec, pages 449–454, 2006. saif mohammad and graeme hirst. distributional measures ofsemantic distance: a survey. arxiv:1203.1858v1, 2012. first published in 2006. robert munro, steven bethard, victor kuperman, vicky tzuyin lai, robin melnick, christopher potts, tyler schnoebelen, and harry tily. crowdsourcing and language studies: the new generation of linguistic data. inproceedings of the naacl/hlt 2010 workshop on creating speech and language data with amazon’s mechanical turk, pages 122–130, los angeles, california, 2010. patrick pantel and marco pennacchiotti. espresso: leveraging generic patterns for automatically harvesting semantic relations. inproceedings of the 21st international conference on computational linguistics and 44th annual meeting of the association for computational linguistics, pages 113–120, sydney, australia, 2006. viktor pekar. discovery of event entailment knowledge fromtext corpora. computer speech & language, 22(1):1–16, 2008. michaela regneri, alexander koller, and manfred pinkal. learning script knowledge with web experiments. inproceedings of the 48th annual meeting of the association for computational linguistics (acl 2010), pages 979–988, uppsala, sweden, 2010. philip resnik. selectional constraints: an information-theoretic model and its computational realization. cognition, 61:127–159, 1996. tom richens. anomalies in the wordnet verb hierarchy. inproceedings of the 22nd international conference on computational linguistics (coling), pages 729–736, manchester, uk, 2008. rob van der sandt. presupposition projection as anaphora resolution. journal of semantics, 9: 333–377, 1992. peter f. strawson. on referring.mind, 59(235):320–344, 1950. galina tremper. weakly supervised learning of presupposition relations between verbs. inproceedings of the acl 2010 student research workshop, pages 97–102, uppsala, sweden, 2010. galina tremper and anette frank. extending semantic relation classification to presupposition relations between verbs. inproceedings of the dgfs workshop: “beyond semantics: corpusbased investigations of pragmatic and discourse phenomena” , göttingen, germany, 2011. marc verhagen, roser sauri, tommaso caselli, and james pustejovsky. semeval-2010 task 13: tempeval-2. inproceedings of the 5th international workshop on semantic evaluation, pages 57–62, uppsala, sweden, 2010. 321 tremper andfrank hila weisman, jonathan berant, idan szpektor, and ido dagan. learning verb inference rules from linguistically-motivated evidence. inproceedings of the 2012 joint conference on empirical methods in natural language processing and computational natural language learning, pages 194–204, jeju island, korea, 2012. ian h. witten and eibe frank.data mining: practical machine learning tools and techniques. morgan kaufmann publishers, 2nd edition, 2005. 322 discriminative analysis of fine-grained semantic relations appendix 1. web-based annotator interface figure 5: question-based annotation for verb pairlose – find. 323 tremper andfrank appendix 2. mapping of allen’s relations to coarse temporalrelation classes temporal relation class allen’s relation graphical representation before(x,y) x< y (strict precedence) x m y (x meetsy) x o y (x overlapsy) after(x,y) x > y (strict succession) x mi y (inverse of meets) x oi y (inverse of overlaps) during(x,y) x s y (x startsy) x f y (x finishesy) x d y (x during y) x = y (x equalsy) table 12: mapping of allen’s relations 324 journal of machine learning research-microsoft word template dialogue & discourse 10(2) 1-33 doi: 10.5087/dad.2019.201 ©2019 jet hoek, jacqueline evers-vermeul and ted j.m. sanders this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). using the cognitive approach to coherence relations for discourse annotation jet hoek jet.hoek@ed.ac.uk the university of edinburgh 3 charles street, edinburgh, eh8 9ad, united kingdom jacqueline evers-vermeul j.evers@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands ted j.m. sanders t.j.m.sanders@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands editor: maite taboada submitted 05/2018; accepted 08/2019; published online 10/2019 abstract the cognitive approach to coherence relations (sanders, spooren, & noordman, 1992) was originally proposed as a set of cognitively plausible primitives to order coherence relations, but is also increasingly used as a discourse annotation scheme. this paper provides an overview of new ccr distinctions that have been proposed over the years, summarizes the most important discussions about the operationalization of the primitives, and introduces a new distinction (disjunction) to the taxonomy to improve the descriptive adequacy of ccr. in addition, it reflects on the use of the ccr as an annotation scheme in practice. the overall aim of the paper is to provide an overview of state-of-the-art ccr for discourse annotation that can form, together with the original 1992 proposal, a comprehensive starting point for anyone interested in annotating discourse using ccr. keywords: discourse annotation, coherence relations, corpus annotation, cognitive approach to coherence relations 1 introduction annotating coherence relations refers to the process of attributing labels that best capture the relation inferred between two segments in a text to that relation. to annotate coherence relations, researchers make use of discourse annotation schemes. discourse annotation schemes differ greatly in the number of relations they distinguish, ranging from two (grosz & sidner, 1986) to 81 relations (carlson & marcu, 2001). this is in part due to the fact that there is disagreement about how many distinct coherence relations language users actually infer and how specific these relations are. on the other hand, these differences seem to be caused by the varying purposes of the annotation schemes and the research traditions they originate from. hoek, evers-vermeul and sanders 2 one approach to describing coherence relations that has been around for a while is the cognitive approach to coherence relations (ccr; sanders, spooren, & noordman, 1992, 1993). not originally designed as a discourse annotation approach, ccr defines four basic cognitive primitives that can be used to order the set of coherence relations language users infer between segments in a text. since its introduction, ccr has primarily been used as a basis for experimental and acquisition research on discourse coherence; this research includes both studies aimed to verify the cognitive relevance of ccr’s primitives and studies in which ccr’s primitives are used as a point of departure for researching discourse coherence (see sanders & evers-vermeul, 2019 for an overview). ccr is also increasingly used as a basis for discourse annotation, as is evidenced by the list of projects that have used ccr to annotate coherence relations included as appendix a. using ccr as a discourse annotation can be appealing for several reasons. since it consists of cognitively relevant primitives, ccr is applicable cross-linguistically.1 indeed, it has successfully been used in discourse annotation projects covering several different languages: dutch (e.g., evers-vermeul, 2005; spooren & sanders, 2008; stukker, 2005; vis, 2011), english (hoek, zufferey, eversvermeul, & sanders, 2017; rehbein, scholman, & demberg, 2016), german (pit, 2003), french (degand & pander maat 2001; pander maat & degand 2001; pit, 2003), spanish (santana, spooren, nieuwenhuijsen, & sanders, 2018), and mandarin chinese (li, evers-vermeul, & sanders, 2013; li, sanders, & evers-vermeul, 2016; xiao, li, sanders, & spooren, to appear). in addition, ccr’s primitives present a systematic approach to the categorization of coherence relations and have been shown to correspond to the distribution of connectives in various languages (e.g., knott & sanders, 1998; li, 2014; pit, 2003; sanders & spooren, 2015; wei, 2018). ccr’s individual primitives also make it attainable to employ naive annotators in annotation projects; scholman, evers-vermeul, and sanders (2016) show that undergraduate students can use a stepwise version of ccr to produce decent quality annotations without extensive training. not being entirely dependent on expert annotators helps cut down on time and expenses of traditional annotation projects and opens up the possibility of crowd-sourcing annotations. furthermore, ccr’s value combinations are often much more informative than end labels and can provide a better insight into annotator disagreements (demberg, scholman, & asr, 2019; see also section 4.1). finally, the ccr taxonomy is easily applied to only a subset of relations. when, for instance, only considering coherence relations involving some form of contrast, or relations signaled by because, it is clear which primitives and distinctions should be included in the annotation ‘tag set;’ for approaches that use end labels this is not necessarily as obvious. there also appear to be some downsides to using ccr as a discourse annotation approach. since ccr was designed to “identify the primitives in terms of which the set of coherence relations can be ordered,” it does not constitute a “complete descriptively adequate taxonomy of coherence relations” (sanders et al., 1992:4).2 since the original 1992 proposal, several additional distinctions have been proposed that aim to improve the descriptive adequacy of the taxonomy. however, these proposals are distributed over several individual papers. in addition, there appears to be some skewedness in how well the approach is developed for different types of relations; there has been a lot of debate on how to operationalize ccr’s primitives in the causal domain, but less so for other types of relations. in addition, several new distinctions have been proposed and frequently used within the domain of causal relations (e.g., volitionality, purpose), while fewer additional distinctions have been suggested for other types of relations. 1 the current ccr inventory is language-independent, unlike for instance the pdtb relation inventory, which was adapted for chinese (zhou & xue, 2015) and turkish (zeyrek & kurfali 2017). this does not rule out, however, the possibility to extend the inventory with language-specific or research-specific features. 2 see scholman (2019), chapter 5, for a discussion about cognitive plausibility and descriptive adequacy in discourse annotation. using ccr for discourse annotation 3 after outlining the basic considerations of ccr and the original ccr taxonomy in section 2, this paper provides an overview of new ccr distinctions that have been proposed over the years, summarizes the most important discussions about the operationalization of the primitives, and introduces a new distinction (disjunction) to the taxonomy to further improve the descriptive adequacy of ccr (section 3). finally, section 4 discusses several important issues that the use of the ccr as an annotation scheme in practice presents; taking note of these points can help in successfully using ccr in corpus annotation. overall, this paper thus provides an overview of state-of-the-art ccr for discourse annotation, and forms, together with the original 1992 proposal, a comprehensive starting point for anyone interested in annotating discourse using ccr. 2 the cognitive approach to coherence relations the original ccr taxonomy was proposed in sanders, spooren, and noordman (1992, 1993), and is very much in line with work by hobbs (1978, 1979, 1990) and kehler (1995, 2002), who also consider coherence relations to be cognitive entities, and approach coherence relations by formulating a limited set of organizing principles. sanders et al. (1992:2) define the concept of coherence relation as “an aspect of meaning of two or more discourse segments that cannot be described in terms of the meaning of the segments in isolation.” coherence relations are the reason that “the meaning of two discourse segments is more than the sum of the parts” (sanders et al., 1992:2). this basic property of coherence relations is referred to as the relational surplus; the criterion that ccr’s primitives have to be features of the relational surplus is the relational criterion. in ccr, discourse relations are considered to hold between segments that are minimally clauses (e.g., evers-vermeul, 2005; sanders & van wijk, 1996); we will refer to this as the clausal criterion. the clausal criterion is closely related to the basic definition of coherence relations in ccr, since clauses are the smallest grammatical units that can function meaningfully in isolation (see hoek, evers-vermeul, & sanders [2017] for a more elaborate discussion of the clausal criterion). ccr considers coherence relations to be cognitive constructs. its taxonomy is therefore intended to be cognitively plausible. for a distinction to meet the cognitive plausibility criterion, it should be observable in or make relevant predictions about language acquisition and language processing (sanders et al., 1992). in addition, evidence for cognitive plausibility can be drawn from the system of linguistic markers. knott and dale (1994) argue that the distinctions made by connectives and cue phrases are indicative of the distinctions made in the minds of language users (see also knott & sanders, 1998). because ccr defines coherence relations as cognitive constructs, the labels attributed to coherence relations in annotation should correspond to the relation that holds in the mental representation of the discourse, i.e., the inferred relation. if a relation is marked by a connective or cue phrase in the text, it may well be the case that the annotated relation does not correspond to what is explicitly signaled by the linguistic marker. (1), for example, is marked by the connective and, but the relation that is inferred is a causal relation: the not marrying is interpreted as a consequence, albeit jokingly, of the chips-eating. (1) should thus be annotated as a causal relation, not as an additive relation as the connective might suggest.3 3 with the exception of a few simple relations we constructed ourselves for the sake of clarity, the vast majority of examples in this paper were extracted from actual utterances, from either fictional or non-fictional sources. we opted to use real examples to give a more realistic illustration of the type of coherence relations you would encounter in annotation tasks than simplified, prototypical examples would give. the examples were not collected systematically but rather selected because of their suitability to illustrate specific properties of coherence relations. the source for each example is provided in appendix b. hoek, evers-vermeul and sanders 4 (1) i would ask a man to open the bag for me — men open most containers for me — but then [he would know i eat chips,]s1 and [he would never marry me.]s2 in focusing primarily on the relations that hold in the mental representation of a discourse, ccr’s approach to the depiction of coherence relations is distinctly different from ‘bottom-up’ annotation approaches that seem to place more focus on the linguistic markers of coherence relations, such as for example the penn discourse treebank (pdtb; prasad et al., 2008). the original ccr primitives meet all three of ccr’s criteria. they are properties of the relational surplus, thereby satisfying the relational criterion. they can also be used to describe relations that hold between clauses or larger discourse segments, thereby satisfying the clausal criterion. finally, all primitives are cognitively plausible. the difference between positive and negative relations, which are distinguished from each other by the polarity primitive (see section 2.1), can for instance be observed in processing (positive relations are processed faster than negative relations: clark, 1974; murray, 1997; wason & johnson-laird, 1971), language acquisition (positive relations are acquired earlier than negative relations: bates, 1976; bloom, et al., 1980; eisenberg, 1980; evers-vermeul & sanders, 2009), and the linguistic system (positive and negative relations are prototypically signaled by different connectives). the remainder of this section will give an overview of the four original ccr primitives: polarity, basic operation, source of coherence, and order of the segments. 2.1 polarity discourse relations hold between two propositions, expressed by s1, which refers to the first segment in the linear order of segments, and s2, which refers to the second segment. a relation with a positive value for polarity features p (antecedent) and q (consequent), as in (2). a relation has a negative value for polarity if it features a negative counterpart of p, not-p, or q, not-q, as in (3). (2) [we liked bob]s1 because [he was both different and apologetic.]s2 (3) [they … never failed to invite us to their houses]s1 although [they knew we would never come.]s2 in (2), s1 presents a consequence (q) of the cause (p) in s2. in (3), however, s1 is a contrastive consequence (not-q) of the cause (p) in s2; a logical consequence of knowing someone never takes your offer could be to stop inviting them. positive relations are often expressed with connectives such as and or because. negative relations are often signaled by connectives such as but or although. although positive relations can often be turned into negative relations by negating one of the arguments, it should be noted that relations with a negative value for polarity do not necessarily contain lexical negation, as is illustrated by (4). similarly, relations containing lexical negation can have a positive value for polarity, as can be seen in (5). (4) although [it’s inspired by the vinyl bars of japan,]s1 [this spot chooses accessibility over authenticity.]s2 (5) [i don’t make them a lot]s1 because [i don’t think it's fair to the other cookies.]s2 2.2 basic operation the category of basic operation takes two values: causal and additive. a relation is causal if there is an implication relation between the two arguments (p → q), as in (6). conditional relations, as in (7), also involve an implication relation and are categorized as having a causal basic operation under the original ccr proposal. using ccr for discourse annotation 5 (6) [phone service in the greater chicago area was tied up for two hours christmas eve]s1 because [some kid called a phone-in show to get a wife for his father.]s2 (7) if [there was a fan club]s1 [i’d be the president.]s2 a relation is additive if there is no causal relation between the segments and the only relation that can be inferred between the segments is p & q, as in (8). (8) [i’m worried]s1 and [i’m confused.]s2 2.3 source of coherence the main distinction made in the source of coherence of a discourse relation is between objective and subjective.4 a discourse relation is objective when its two segments are related by their locutionary meaning; the relation is observable in the real world, as in (9).5 subjective relations are related because of the illocutionary meaning of one or both of its segments; they involve the speaker’s reasoning, as in (10); subjective discourse relations are often a reason or motivation for a claim or conclusion. (9) [a harry potter festival that was supposed to take place near glasgow this summer has been cancelled,]s1 because [too many people wanted to go.]s2 (10) [knitted gifts are great]s1 because [they are timeless and will last forever if taken proper care of.]s2 a specific type of subjective relations are speech act relations. in a speech act relation one of the segments relates to the performance of the speech act in the other segment, for instance by offering a motivation or justification, as in (11), or by indicating the relevance of an utterance, as in (12). speech act relations can also hold between two speech acts, as in (13). (11) [how long are they going to take to cook?]s1 because [you’ve got twelve minutes to go.]s2 (12) [there is a wonderful theatre program,]s1 if [she’s interested in that.]s2 (13) [why would it take an unusual woman to keep him company?]s1 and [why was he wearing a russian astronaut on his lapel?]s2 the source of coherence values are highly comparable to sweetser’s (1990) domains of use, with objective relations corresponding to sweetser’s content relations, and subjective relations including both epistemic and speech act relations. other distinctions similar to ccr’s source of coherence can be found in halliday and hasan (1976) and martin (1992: internal vs. external), redeker (1990: ideational vs. rhetorical), mann and thompson (1988: subject matter vs. presentational matter), hovy and maier (1995: ideational vs. interpersonal and textual), pander maat (2002: content vs. epistemic and interactional), and van dijk (1977: semantic vs. pragmatic). 2.4 order of the segments discourse relations generally consist of two segments. the linearly first segment is always referred to as s1; the linearly second segment is always s2. the order of the segments feature refers to 4 in the original 1992 proposal, the values of source of coherence were called semantic and pragmatic. these were later renamed as, respectively, objective and subjective in pander maat and sanders (2000). 5 in using the term ‘real world,’ we do not only refer to the actual earth, but also to the ‘real world’ in for instance fictional settings. hoek, evers-vermeul and sanders 6 how p and q of the basic operation map onto s1 and s2. it takes two values: basic if s1 expresses p and s2 expresses q, as in (14), and non-basic if s1 expresses q and s2 expresses p, as in (15). (14) because [they live in sub-tropical climates,]s1 [african penguins have to cope with both cooling down on land and keeping warm in the water.]s2 (15) [i had to talk loud]s1 because [the movie was loud!]s2 the order of the segments is only relevant to causal relations, since additive relations are symmetrical in this respect. 3 extensions of the original ccr taxonomy polarity, basic operation, source of coherence, and order of the segments are the original four primitives of the ccr taxonomy. since the 1992 proposal, there has been a lot of discussion on how to operationalize the primitives, as well as proposals for new primitives or additional distinctions to ccr for discourse annotation. in this section, we provide an overview of the most important developments since the original ccr proposal. 3.1 additional distinctions within original primitives there have been proposals for additional distinctions within certain parts of the original ccr taxonomy. unlike the original primitives, these additional distinctions apply only to a (small) subset of coherence relations. the proposed distinctions allow annotators to make more fine-grained contrasts, thus improving the descriptive adequacy of the ccr taxonomy. additional distinctions have been proposed within the class of positive relations (temporality), the class of positive objective causal relations (volitionality and purpose) and within the class of negative relations (directness). it should be noted that many of the new distinctions have been proposed for only a subset of relations. this does not necessarily mean that the same distinction could not also be annotated for other types of relations. volitionality (section 3.1.2), for example, divides the subset of positive causal relations into volitional and non-volitional causal relations, a distinction that has been argued to be cognitive relevant on the basis of evidence from language processing, language acquisition, and linguistic systems. while volitionality could also be annotated for, for instance, conditional relations, there is no clear evidence that the distinction between volitional and non-volitional is as cognitively relevant within the class of conditional relations as it is within the class of causal relations. if such evidence were to be found, the volitionality distinction could easily be extended to apply to all implication relations with a positive value for polarity; the same holds for other distinctions as well. limiting additional distinctions to apply to only those subsets for which there is empirical evidence for the distinction’s cognitive relevance, helps to create a balance between descriptive adequacy on the one hand, and cognitive plausibility on the other. 3.1.1 temporality the ccr taxonomy has recently been proposed to be extended with a new distinction: temporality. the original ccr proposal considers temporal relations to be a subtype of positive additive relations; sanders, spooren, and noordman (1992:28) state that “the properties distinguishing temporal relations from other additive relations concern the referential meaning of the individual segments.” temporality is thus taken to be a propositional, rather than a relational feature of coherence relations, and, as such, does not meet all the criteria necessary to be adopted into the ccr taxonomy. evers-vermeul, hoek, and scholman (2017), however, argue that temporality does meet the relational criterion. they show that the temporal information in the propositional content of the segments is not always sufficient to establish a temporal coherence relation. in addition, they argue that the ordering of discourse segments in time can only be using ccr for discourse annotation 7 determined for a combination of the discourse segments; not for segments in isolation. as such, temporality is a feature of the relational surplus and meets the relational criterion. temporality is then argued to also meet all other ccr criteria. temporal relations can hold between clauses, and the relevance of temporality is observable in language processing, language acquisition, and in the connective inventory of several different languages. after discussing other options, evers-vermeul, hoek, and scholman (2017) argue that the best way of incorporating temporal relations in ccr is adding another primitive to the taxonomy. the proposed primitive distinguishes between relations that are ordered in time and relations that are not ordered in time. positive additive relations that are ordered in time are relations that are most prototypically referred to as ‘temporal relations.’ as is shown in figure 1, two additional steps make more fine-grained distinctions within the set of relations that are ordered in time: between sequential and synchronous relations and between sequential relations that are chronologically ordered and sequential relations that have an anti-chronological order. while not explicitly included in the original ccr taxonomy, the use of a primitive that includes additional distinctions relevant to only a subset of relations is in line with later proposals for additional distinctions, such as volitionality (see section 3.1.2). 1 temporal non-temporal 2 sequential synchronous 3 chronological anti-chronological figure 1. the three-step temporality primitive although it was not explicitly addressed in evers-vermeul, hoek, and scholman (2017) whether the temporality distinction is applicable to all coherence relations, we consider temporal order to be especially productive to relations with a positive value for polarity. with temporality as a separate primitive, two different types of order can be distinguished for causal and conditional relations: implication order, as depicted by the original order of the segments primitive, i.e., basic versus non-basic order, and temporal order, i.e. chronological versus anti-chronological order. these two orders will coincide for many relations, as in the conditional relation in (16). the relation has a basic order of the segments, i.e., s1 expresses p and s2 expresses q. it also has a chronological temporal order, i.e., the event expressed in s1, buying a railcard online, occurs before the event expressed in s2, replacing that railcard. sometimes, however, the two orders diverge, as in the conditional relation in (17), which has basic order, but anti-chronological order, since the event expressed by s1, gaining muscle, occurs after the event expressed by s2, lifting as heavy as possible. the idea that there is an underlying temporal order that is opposite to the order of the segments is underlined by the fact that the relation in (17) can be paraphrased as you should always lift as heavy as possible, because then you will gain muscle, while a similar construction cannot be used to paraphrase (16). (16) if [you bought your railcard online,]s1 [you will need to get a replacement online.]s2 (17) if [you want to gain muscle,]s1 [you should always lift as heavy as possible.]s2 the addition of the temporality feature has two important advantages. first, it opens up the possibility to investigate differences and similarities in the use of causal and conditional relations with a temporal order on the one hand and purely temporal relations on the other. second, it helps make finer-grained distinctions within the class of implication relations, with relations in which implication order and temporal order do not coincide corresponding to for instance relations annotated as enablement or problem-solution in the rst-dt (carlson & marcu, 2001) or as hoek, evers-vermeul and sanders 8 implicit assertion in the pdtb 2.0 (pdtb research group 2008). it should be noted that while the temporality feature (specifically the third temporal ordering step) can be annotated for causal and conditional relations, this does not imply that these relations in ccr would belong to the class of temporals distinguished in rst-dt or pdtb, since the implication relation is considered to be a more salient (stronger) feature of these relations. 3.1.2 volitionality it has been proposed that within the class of positive objective causal relations, a distinction can be made between volitional and non-volitional relations (e.g., pander maat & sanders, 2000; sanders et al., 1992; stukker, sanders, & verhagen, 2008; see also mann & thompson, 1988). volitional causal relations involve a thinking actor who is responsible for an event in the consequent of the relation, as in (18), where the making event in s1 is a volitional action. non-volitional causal relations do not involve a volitional action. in the relation in (19), for example, the consequent does not involve an agent; one fact leads to the other. it should be noted that some languages have dedicated connectives for non-volitional causal relations, such as daardoor ‘that is why’ and doordat ‘because of the fact that’ in dutch (e.g., stukker et al., 2008). (18) [i make them a lot]s1 because [i have this indescribable need to constantly have new pillows.]s2 (19) [the game has changed]s1 because [the way we communicate has changed.]s2 pander maat and sanders (2000) propose that volitional causal relations have something in common with subjective causal relations (see section 3.3). both types of relations involve a subject of consciousness (soc); a thinking entity involved in the relation. the main difference between volitional causal relations and subjective causal relations is that in subjective relations the soc is involved in the construal of the relation (see section 3.3.1), whereas in volitional causal relations, the soc is not. instead, the soc in a volitional causal relation is usually an agent. in addition, the soc in volitional relations is typically explicitly mentioned (onstage: see section 3.3.2), though it may also be inferable in for instance a passive construction, see also section 3.1.3. while the speaker is responsible for the action in s1 and the fact in s2, the causal relation does not stem from the speaker’s mind and is observable in the real world. non-volitional causal relations do not involve an soc at all. 3.1.3 purpose another distinction within the class of positive objective causal relations is purpose (sanders et al., 2018). purpose relations feature a volitional action for which the motivation is an intended result. in (20), for instance, the adding of the smell is done to achieve the intended result of people knowing when there is a gas leak. unlike the relation in (20), the relation in (21) does not feature an explicitly mentioned agent and instead uses a passive construction in s1. while the relation in (20) has an explicit agent (they), the agent in (21) is implicit in the passive construction in s1. since the agent in (21) is not absent but merely unmentioned, it can still be classified as a positive, objective causal relation specified for purpose. several different languages have connectives that typically express purpose relations, such as so that or in order to in english, or zodat ‘so that’ in dutch. (20) the gas is odorless, but [they add the smell]s1 so [you know when there’s a leak.]s2 (21) [services are being enhanced to remain open 24 hours]s1 so that [no one will have to stay on the streets during the cold snap.]s2 for causal relations specified for purpose, determining the order of the segments is not entirely straightforward (e.g., sanders et al., 2018; sanders et al., 1992). on the one hand, the relations in using ccr for discourse annotation 9 (20) and (21) are very similar to result relations (i.e., positive causal relations with basic order). on the other hand, they also bear similarities to volitional causal relations with a non-basic order like the one in (18), because the intended result is the motivation for executing the volitional action in the first place (see also reese et al., 2007:12-13). in ccr, the intended result in positive causal relations specified for purpose should be considered the consequent, q, while the intentional action should be considered the antecedent, p. the order of the segments in (20) and (21) is therefore basic. 3.1.4 directness pander maat (1998) evaluates the original ccr taxonomy with respect to negative relations. he argues that the original primitive inventory is insufficient to capture all major distinctions between relations with a negative value for polarity. on the basis of a corpus annotation study and using linguistic evidence, primarily from the dutch connective inventory, he proposes a new distinction to be applied to negative additive coherence relations: directness.6 pander maat (1998) poses that in negative additive relations, the two segments are compared to each other. this comparison is direct if “the propositions are themselves incompatible” (pander maat 1998:192); the propositional content of s1 is in direct contrast to the propositional content of s2. the comparison can also be indirect, in which case the results or conclusions on the basis of propositions are incompatible. direct, negative, objective, additive relations contain, for instance, a semantic contrast. in (22) the statements about neilia and jill are directly compared. in (23), on the other hand, an indirect, negative, objective, additive relation, it is not the segments themselves that are in contrast to each other, but rather the results of both segments (‘conflicting causal forces’); daily gains imply an improvement, but the second segment indicates a trend in the opposite direction. (22) [neilia would always be mommy,]s1 but [jill was mom.]s2 (23) [stock market notches daily gain,]s1 but [posts largest weekly drop since early 2016]s2 within negative, subjective, additive relations, directness mainly distinguishes between qualifications and concessions. sanders, spooren, and noordman (1992) categorize concessions as negative, subjective, additive relations. in their view (see also spooren, 1989), concessions are relations that feature two arguments in favor of opposing views, see figure 2.7 concessions are similar to relations with conflicting causal forces, as in (23), except for their source of coherence. i won’t eat the dish ≠ i will eat the dish [i don’t like vegetables,] but [i do love chicken.] figure 2. concession 6 pander maat (1998) also discusses perspective (same perspective versus perspective change) as a potential new distinction for negative relations. however, this distinction cannot be applied as systematically as all other ccr primitives and additional distinctions (see pander maat, 1998:194, figure 1), since it appears to be irrelevant to some types of negative relations, fixed for others, and only a true distinction for a few combinations of primitives/distinctions. in addition, the perspective distinction mainly appears to be a property of the propositional content of the segments, rather than a relational feature. we will therefore not discuss the perspective distinction at length here. 7 outside of ccr, concession is also often used to refer to negative causal relations, e.g., “although she studied hard, she failed the exam.” hoek, evers-vermeul and sanders 10 in concessions, the conclusion that can be drawn on the basis of the first segment is incompatible with the conclusion that can be drawn on the basis of the second segment. since the causality is not found between the segments, but rather between the segments and their associated inferences, the relation between s1 and s2 is not an implication relation and, as such, the relation in figure 2 is considered an additive relation. (24) and (25) are actual examples of concessions. in (24), the inference made on the basis of the first segment, “i won’t agree with you,” is in contrast with the inference made on the basis of the second segment “i will agree with you.” in (25), the contrast holds between “you can write it yourself” and “we will have someone else write it.” (24) “this is a beautiful house.” “thank you. i never know what to say when somebody says that. [you don’t want to agree]s1 but on the other hand, [it feels weird to disagree and say ‘no it’s a dump’.”]s2 (25) [i’m sure you would like to write the book yourself,]s1 but [your record is not what i might call promising, book-finishing-wise.]s2 in qualifications, the second segment “cancels the strongest interpretation of the first statement” (pander maat 1998:186). as is illustrated in figure 3, qualifications are similar to concessions, but the conclusion made on the basis of s2 directly contrasts with the propositional content of s1. while concessions are indirect, qualifications are thus direct. ≠ there are some vegetables i do like [i hate vegetables,] but [snap peas are okay.] figure 3. qualification (26) contains an actual example of a qualification relation. the proposition expressed in s1, “i don’t know any blind people,” is qualified by the statement that the speaker does know someone with a pretty severe eye condition, which implies that he does know someone who is practically blind. (26) [i, personally, don’t know any blind people,]s1 though [the guy i used to buy my newspaper from had pretty bad cataracts.]s2 pander maat (1998) further distinguishes four specific types of qualifications: simple qualification, exceptions, qualified denial, and denied intensification. the differences between the four types mainly refer to the direction of the qualification (weakening or intensifying) and to whether a stronger or weaker interpretation of the first segment should only be made to a certain extent or not at all. each subtype of qualification, however, follows the general relation configuration illustrated in figure 3. as such, the different qualification subtypes seem too fine-grained and segmentspecific to be incorporated into the ccr taxonomy by means of additional distinctions (this is also not something pander maat [1998] proposes). however, taking note of these specific variations may be helpful in recognizing qualifications during annotation. including the directness distinction within negative additive relations helps make the ccr taxonomy more descriptively accurate. in addition, pander maat (1998:199, table 1) demonstrates that the differences between direct and indirect negative relations can be observed in the dutch connective system, which suggests that the distinction is also cognitively plausible. finally, as pander maat (1998) points out, directness makes the ccr taxonomy more consistent, since the original 1992 proposal conflated source of coherence and directness for negative additive using ccr for discourse annotation 11 relations; the class of negative objective additive relations only included direct comparisons, while the class of negative subjective additive relations only included indirect comparisons. 3.2 proposing a new distinction: disjunction ccr was recently used as a tool to map other discourse annotation schemes onto each other (see sanders et al., 2018). the relation labels from rst, pdtb, and sdrt were ‘translated’ into ccr’s primitives, enabling a more accurate and straightforward comparison between the different frameworks than just comparing the end labels would have allowed. while ccr was able to capture the majority of distinctions, several extra features had to be formulated to ascribe a unique set of primitives and features to each relation label from a framework.8 most extra features were similar to the distinctions discussed in section 3.1 in that they were relevant to only a small subset of relations, and defined more specific instances of a certain relation type (e.g., list relations as a specific instance of positive additive relations). a notable exception was disjunction, a feature that distinguishes disjunction relations, in which the two segments are presented as alternatives, from other additive relations. whereas rst, pdtb, and sdrt all include disjunctions as a specific relation type, the original ccr taxonomy is unable to adequately capture the distinction between disjunctions and other types of additive relations. 3.2.1 disjunctions in ccr as the main reason for not including an “alternation relation” in their taxonomy, sanders, spooren, and noordman (1992:29) refer to the “unclear status of alternation.” while some of the existing approaches to discourse coherence treated disjunction relations as a distinct class of relations, for instance “on a par with conjoining, temporal, and implication,” as longacre (1983), others considered them a subcategory of additive relations (e.g., halliday & hasan, 1976). in addition, as sanders et al. (1992:29) point out, there was “also confusion about the nature of the alternation relation;” while some considered disjunctions to be primarily inclusive (e.g., longacre, 1983), others considered disjunctions to be primarily exclusive (e.g., gamut, 1982; levinson, 1983). here, we would like to argue in favor of including an additional distinction in the ccr taxonomy that can account for disjunction relations. not only would such a distinction improve the descriptive adequacy of the taxonomy, it also seems to meet all criteria set by the ccr approach. first of all, disjunction relations hold between clauses, see (27), thereby satisfying the basic clausal criterion. disjunction is also a feature of the relational surplus, since the meaning of the relation as a whole is more specific than just the segments in isolation; without disjunction, as in (27’), the two segments would not be considered alternatives and both segments would be considered to be true. (27) [you either know it]s1 or [you don’t.]s2 (27’) you know it // you don’t know it the final criterion that relational features have to meet before they can be included into the ccr taxonomy is cognitive plausibility. there seems to be ample linguistic evidence from connective inventories to suggest that disjunction is a cognitively plausible distinction, since many languages have connectives that prototypically mark disjunctions, for instance or or either or in english, of in dutch, oder in german, ou in french, and o in spanish. as discussed in section 2, other evidence related to the cognitive plausibility of features of coherence relations can be derived from language acquisition and language processing. disjunction at the discourse level, however, does not seem to have received a lot of attention in these fields. a notable exception is a self-paced reading study by 8 additional features were formulated if a distinction was made in at least two out of three frameworks. note that these additional features were not proposed as new distinctions within ccr, but as necessary tools for the purposes of the sanders et al. (2018) paper. hoek, evers-vermeul and sanders 12 staub and clifton (2006). this experiment compares reading times of disjunctions in past tense and future tense signaled by or or either or. staub and clifton (2006) find that readers benefit more from the presence of either in the past tense condition than in the future tense condition. this suggests that when encountering a connective indicating disjunction after the first segment, readers have to update the truth-conditional status of s1. this effect is much smaller, or even absent, in the future tense condition because the truth-conditional status of those segments is already uncertain. staub and clifton’s (2006) experiment thus shows that disjunction can affect language processing and, as such, provides additional evidence in favor of the cognitive plausibility of disjunction. while evidence pertaining to the cognitive plausibility of disjunction is fairly limited, there currently does not seem to be any evidence against disjunction being a cognitively plausible distinction at the discourse level. since disjunction currently appears to meet ccr criteria and would improve ccr’s descriptive adequacy, we suggest adopting disjunction as an additional primitive in ccr. further investigating the cognitive plausibility of disjunction at the discourse level seems a fruitful endeavor for future research; any counter-evidence this research may uncover should be taken into account in future evaluations of the disjunction primitive. 3.2.2 disjunction as a new distinction in ccr in line with the original sanders, spooren, and noordman (1992) paper, we consider disjunctions to be a specific type of additive relations. here, however, we propose to include disjunction as an additional distinction to the ccr taxonomy, applicable only to the class of additive relations. similar to the additional distinctions discussed in section 3.2, disjunction will carry the values alternative, in which case the segments are presented as alternatives, and not alternative, in which case the segments are not presented as alternatives. additive relations that are alternative are the relations prototypically referred to as the class of disjunctions; additive relations that are not alternative are all other types of additive relations. as mentioned in section 3.2.1, disjunctions can be exclusive, in which case the alternatives cannot hold at the same time, as in (27), or inclusive, in which case they can, as in (28). (28) [a little sweetener can take them [= waffles] from supper table to breakfast table]s1 or [even turn them into dessert.]s2 it is possible to distinguish between the inclusive and exclusive disjunctions using the polarity primitive (see also sanders et al., 2018). since the two segments can hold at the same time, inclusive disjunctions have a positive value for polarity: p & q. exclusive disjunctions, on the other hand, always involve the negative counterpart of either p or q: p & not-q or not-p & q. in (27), for instance, you know it, in which case you do not not know it, you do not know it, in which case you do not know it. the source of coherence primitive applies to disjunctions as it does to other types of coherence relations; disjunctions can be either objective or subjective. in (27), the segments present alternatives that hold in the real world. as such, (27) has an objective value for source of coherence. (29) presents two alternative opinions or claims, making the relation subjective; note that the two segments do not necessarily present real-world alternatives, since it could both be true that ‘this person’ is just stupid and has lost her mind, or that neither proposition holds. (30) is also subjective, since the disjunction holds between two speech acts, specifically between two questions. (29) either [this person has lost her presence of mind]s1 or [she is just stupid.]s2 (30) [are you just feeling lazy]s1 or [do you need a break?]s2 since disjunctions are considered to be a subtype of additive relations, the order of the segments primitive does not apply. it should be noted that ‘disjunctions’ are sometimes using ccr for discourse annotation 13 considered to include unless-relations (e.g., pdtb research group, 2007; reese et al., 2007: unless you know it, you don’t know it has a meaning highly similar to the relation in [27]). in ccr, relations marked by unless are categorized as negative conditional relations; this also holds for relations not specifically marked by unless but with a similar interpretation. 3.3 operationalizing source of coherence: segment-internal distinctions the distinction between objective and subjective relations (or a similar distinction) is, as mentioned in section 2.3, very common in theories about discourse and discourse annotation approaches. although researchers seem to agree on prototypical examples, the source of coherence of a relation can be difficult to determine in the practice of actual corpus annotation (e.g., sanders, 1997). a proposal to improve the application of this primitive in the annotation of real-world examples is to make use of paraphrase tests, in which the segments of the relation are inserted in a paraphrase that makes explicit either a subjective or objective reading, for instance the fact that p causes s’s claim/advice/conclusion that q can be used to test whether positive causal relations with a basic order are subjective (sanders, 1997). another practice that seems to facilitate determining the source of coherence of a relation is to consider the relation in its larger context, for example the whole text (sanders, 1997; sanders & spooren, 2013). it has been proposed that determining the source of coherence of a relation is difficult because while there are highly prototypical instances of objective and subjective relations, there are also many less prototypical examples (e.g., degand & sanders, 1999; sanders, 1997; stukker & sanders, 2012);9 non-prototypical examples are harder to classify than more prototypical examples. several papers explore what makes a relation prototypically subjective or objective. relevant features include the identity of the subject of consciousness, the explicit presence of the subject of consciousness, and the propositional attitude of the segments, each of which will be elaborated on in the rest of the section. using these individual features can facilitate the process of determining a relation’s source of coherence, as will be explained in section 4.3. at the same time, the individual features are also used as additional distinctions within the source of coherence primitive to examine connective profiles in a more fine-grained way (e.g., li, 2014; santana et al., 2018; xiao et al., to appear). while this section focuses on the additional distinctions pertaining to the source of coherence primitive that have been proposed in previous literature (which may all help in operationalizing this primitive), section 4.3 discusses an additional issue that appears to complicate the annotation of source of coherence: the distinction between source of coherence and truth value; since this issue does not come with additional distinctions that can be annotated, we discuss it in section 4, along with other issues that researchers may encounter when using ccr in corpus annotation. 3.3.1 identity of the subject of consciousness pander maat and sanders (2000) propose that subjective relations involve a subject of consciousness (soc) that is responsible for the construal of the relation; the relation stems from the soc’s mind (see also pander maat & degand, 2001; pit, 2003; j. sanders, sanders, & sweetser, 2012; sanders, j. sanders, & sweetser, 2009, among others). subjective causal relations, for instance, involve the soc’s reasoning, as in (31). as was mentioned in section 3.1.2, objective relations have either no soc (non-volitional relations) or an soc that is not responsible for the construal of the relation but is present as the agent of a volitional action (volitional relations). 9 some have even claimed that it involves fitting a scalar phenomenon into distinct categories (e.g., degand & pander maat, 2003; pander maat & degand, 2001). see stukker and sanders (2012) for an overview of this argument, as well as stukker and sanders’ argument in favor of a prototypicality account. hoek, evers-vermeul and sanders 14 (31) [it must have been turkey mating season in northern california]s1 because [we’ve never seen so many turkeys strutting around.]s2 in subjective coherence relations, the soc is usually the speaker (pander maat & sanders, 2000): either the speaker or author of the discourse, as in (31), or the speaker responsible for the content of a direct quote, as in (32). alternatively, the soc can be another actor in the discourse whose perspective is taken, as in (33). in (33), s2 is a conclusion made on the basis of information in s1. it does not explicitly say ‘so tarzan concludes that the natives must be very near,’ but it is clear that the conclusion is drawn by tarzan. tarzan is the thinking entity responsible for the construal of the relation and therefore the soc. in examples like (33), a third person actor temporarily becomes the speaker, although it would be even more accurate to say that in examples like these there is a ‘blend’ between the perspectives of the author or speaker and the discourse participant. (e.g., j. sanders et al., 2009, j. sanders & spooren, 1997). (32) “my intelligence can be very intimidating,” devos said. “and if [donald trump was a moron,]s1 [he would not want to be around people who are intelligenter than him.”]s2 (33) [tarzan] was startled. had he remained too long? quickly he reached the doorway and peered down the village street toward the village gate. the natives were not yet in sight, though [he could plainly hear them approaching across the plantation.]s1 [they must be very near.]s2 3.3.2 explicit presence of the subject of consciousness not only the identity of the soc, but also the extent to which the soc is explicitly present in the relation has been argued to bear on the subjectivity of a relation. langacker (1990, 1991, 2006) proposes that utterances with an explicitly mentioned, ‘onstage,’ speaker are more objective than utterances where the speaker is left implicit, or ‘offstage,’ since an explicitly mentioned speaker becomes itself the focus of attention. this view is applied to coherence relations by, for instance, pit (2003), sanders and spooren (2013, 2015), and stukker and sanders (2012), who show that relations with onstage socs, as in (34), are less prototypically subjective than relations in which the soc remains offstage, as in (34’). however, relations with an onstage speaker soc do tend to be considered to be subjective relations if the relation is centered around a subjective judgment, opinion, or conclusion (e.g., pander maat & degand, 2001; pander maat & sanders, 2000; pit, 2003; sanders & evers-vermeul, 2019; wei, 2018).10 (34) [i think all glitter should be banned,]s1 because [it’s microplastic.]s2 (34’) [all glitter should be banned,]s1 because [it’s microplastic.]s2 it should be noted that a subjective relation can explicitly mention someone whose identity corresponds to the identity of the soc, and still have an implicit soc. in (35), for instance, the soc is the speaker, but he is not explicitly mentioned in his role as soc (as would be the case in which i think was a bummer). instead, he is merely explicitly mentioned as an actor in the event in s2 that is used to motivate the judgment in s1. 10 as was mentioned in footnote 8, some researchers have proposed that subjectivity is a scalar, rather than a categorial notion (pander maat & degand, 2001; degand & pander maat, 2003). under this view, (34’) would be considered to be more subjective than (34). in this paper, we adopt a categorial view of the source of coherence primitive (while recognizing that relations can differ in how prototypically subjective they are), since this approach is most in line with the common annotation practice of assigning labels to relations. using ccr for discourse annotation 15 (35) i made it through the night without getting fired. [which was a bummer]s1 because [i had spent the days previous applying for new serving jobs through craigslist, just in case.]s2 3.3.3 propositional attitude of the segments a final feature of coherence relations that is relevant to its source of coherence is the propositional attitude of the segments (e.g., li, 2014; li et al., 2016; sanders & spooren, 2009, 2015; spooren & degand, 2010); are they, for instance, judgments, speech acts, or facts? subjective relations prototypically involve judgments or speech acts, while objective relations prototypically feature facts. for implication relations, the propositional attitude of the consequent, q, is most crucial (li, 2014). 3.4 state-of-the-art ccr for discourse annotation this section gave an overview of the most important developments in ccr since the original 1992 proposal when it comes to discourse annotation. figure 4 provides a schematic overview of stateof-the-art ccr for discourse annotation. the overview is a flowchart resulting in unique value combinations at the bottom of the scheme. as is indicated by the grey shading and the prominence of polarity, basic operation, and source of coherence, these are the only primitives relevant to all coherence relations. the distinctions in red squares are only relevant to the subset of relations below the primitive value to which they are attached. as such, they duplicate the set of relations below that primitive value. the numbers at the bottom of the scheme refer to the numbers in table 1, where a simple, prototypical example is provided for each value combination in ccr. the segment-internal distinctions for source of coherence discussed in section 3.3 are not explicitly incorporated in the scheme, but are considered to be part of the objective-subjective distinction within source of coherence. for temporality, we only included the temporal order step for positive causal relations; by definition, these relations contain an underlying sequential temporal order. hoek, evers-vermeul and sanders 16 f ig u re 4 . s ta te -o fth ear t c c r . p o s= p o si ti v e, n eg = n eg at iv e, c au s= ca u sa l, a d d = ad d it iv e, c o n d = co n d it io n al , o b j= o b je ct iv e, s u b j= su b je ct iv e, b = b as ic ,n b = n o n -b as ic , te m p = te m p o ra l, se q = se q u en ti al , sy n = sy n ch ro n o u s, c = ch ro n o lo g ic al , ac = an ti -c h ro n o lo g ic al , v o l= v o li ti o n al ,n o n -v o l= n o n -v o li ti o n al , p u rp = p u rp o se , al t= al te rn at iv e, d ir = d ir ec t, i n d ir = in d ir ec t using ccr for discourse annotation 17 1 positive causal objective basic chronological -volitional -purpose because it was raining, the streets were getting wet. 2 positive causal objective basic chronological +volitional -purpose because it was raining, jill brought her umbrella. 3 positive causal objective basic chronological +volitional +purpose joe put up a tarp over the party area to prevent everyone from getting wet. 4 positive causal objective non-basic anti-chronological -volitional -purpose the streets were getting wet because it was raining. 5 positive causal objective non-basic anti-chronological +volitional -purpose jill brought her umbrella because it was raining. 6 positive causal objective non-basic anti-chronological +volitional +purpose to prevent everyone from getting wet, joe put up a tarp over the party area. 7 positive causal subjective basic chronological the streets are wet, so it must be raining. 8 positive causal subjective basic anti-chronological to prevent everyone from getting wet, you should cover the party area with a tarp. 9 positive causal subjective non-basic chronological you should cover the party area with a tarp to prevent everyone from getting wet. 10 positive causal subjective non-basic anti-chronological it must be raining, since the streets are wet. 11 positive conditional objective basic chronological if it rains, jill will bring an umbrella. 12 positive conditional objective non-basic anti-chronological jill will bring an umbrella if it rains. 13 positive conditional subjective basic chronological if jill brought an umbrella, it must be raining. 14 positive conditional subjective basic anti-chronological if you want to prevent everyone from getting wet, you should cover the party area with a tarp. 15 positive conditional subjective non-basic chronological you should cover the party area with a tarp, if you want to prevent everyone from getting wet. 16 positive conditional subjective non-basic anti-chronological it must be raining, if jill brought an umbrella. 17 positive additive objective +temporal +sequence chronological -disjunction joe put up a tarp before it started to rain. 18 positive additive objective +temporal +sequence anti-chronological -disjunction before it started to rain, joe put up a tarp. 19 positive additive objective +temporal +synchronous -disjunction while it was raining, mona sat inside reading a book. 20 positive additive objective -temporal -alternative mona read a book. she also built legos with her son mike. 21 positive additive objective -temporal +alternative mike loves building with his legos or even just taking his lego creations apart. 22 positive additive subjective -temporal -alternative legos are great. reading is also wonderful. 23 positive additive subjective -temporal +alternative jill is usually described as being great company or even as being someone who makes any party a success. hoek, evers-vermeul and sanders 18 24 negative causal objective basic even though it was raining, the streets stayed dry. 25 negative causal objective non-basic the streets stayed dry, even though it was raining. 26 negative causal subjective basic even though it is raining, you should not bring an umbrella. 27 negative causal subjective non-basic you should not bring an umbrella, even though it is raining. 28 negative conditional objective basic unless the skies have cleared, we are bringing an umbrella. 29 negative conditional objective non-basic we are bringing an umbrella, unless the skies have cleared. 30 negative conditional subjective basic unless it is absolutely pouring down, you should not bring an umbrella. 31 negative conditional subjective non-basic you should not bring an umbrella, unless it is absolutely pouring down. 32 negative additive objective direct -alternative jill brought an umbrella, but her friend did not. 33 negative additive objective indirect -alternative the rain is making the streets wet, but the sun is drying them really quickly. 34 negative additive objective direct +alternative the whole party it was either drizzling or pouring down. 35 negative additive subjective direct +alternative every party last year was either really great or it was a total disaster. 36 negative additive subjective direct -alternative rain is the absolute worst, though the smell of a light drizzle after a sunny day is pretty wonderful. 37 negative additive subjective indirect -alternative going to that party sounds like fun, but it is pouring down outside. table 1. prototypical examples for each value combination in state-of-the-art ccr for discourse annotation. 4 ccr as an annotation scheme in practice in the introduction, we mentioned several advantages of using ccr for discourse annotation; it consists of cognitively plausible distinctions, is applicable cross-linguistically, and can be used by non-expert annotators. we also mentioned some potential problems researchers could run into when starting to use ccr, most of which we aim to solve in the current paper. we provided an overview of all proposed additional primitives and distinctions and gave a summary of several discussions that have been carried out over separate research papers. this eliminates the need to sift through many different research papers to create an overview of state-of-the-art ccr. in addition, we took inventory of the full ccr taxonomy to see if there were any potential extra distinctions that would be eligible to be adopted into ccr and would increase the approach’s descriptive adequacy. this led to our proposal for disjunction as a new distinction in ccr. in this section, we reflect on the use of ccr as an annotation scheme in practice, because the use of primitives presents certain challenges that an end-label approach does not. we discuss several issues that are relevant to take into account when implementing the ccr taxonomy in a discourse annotation project:11 different options for calculating inter-annotator agreement (section 11 insights are mainly based on a recent annotation project on english coherence relation in the europarl corpus (koehn, 2005; see hoek, zufferey, evers-vermeul, & sanders, 2017 for a more extensive overview of the annotation project). using ccr for discourse annotation 19 4.1), assumptions about the independence of primitives and distinctions, and the possibility of also using end labels when using ccr (section 4.2). in addition, we discuss an additional point concerning the operationalization of source of coherence. while the distinction between subjective and objective relations also exists in other discourse annotation frameworks, no framework applies this distinction as systematically to all types of relations as ccr (see sanders et al., 2018). when a decision about whether a relation is objective or subjective has to be made in another framework, this will be done by comparing the relation that has to be annotated to the objective and subjective counterpart of the candidate relation (e.g., considering the relation definitions of rst-dt’s cause versus pragmatic cause). in ccr, a relation’s source of coherence always has to be determined based on a general definition of the source of coherence primitive, which may be part of the reason why annotating this distinction is so difficult. in practice, a relation’s source of coherence appears to be commonly confused with its truth value. we clarify this distinction in section 4.3. taking note of the issues discussed in sections 4.1, 4.2, and 4.3 can help to successfully apply ccr –both the original primitives discussed in section 2 and the later proposed primitives and distinctions discussed in section 3– in corpus annotation. 4.1 calculating inter-annotator agreement when annotating coherence relations, researchers have to rely heavily on their own interpretation of the discourse, which is why discourse annotation is, at least to some extent, a subjective endeavor (e.g., spooren & degand, 2010). to demonstrate that annotation has been done reliably and reproducibly, researchers can report an inter-annotator agreement measure: either simple percentage agreement or a chance-corrected numerical index that indicates the amount of agreement between two independent coders (e.g., cohen’s kappa, krippendorf’s alpha, gwet’s ac1). for annotation efforts that make use of end labels to categorize coherence relations, the basis for calculating inter-annotator agreement is a confusion matrix like the one in table 2. coder 2 coder 1 end label 1 end label 2 end label 3 total end label 1 agree x x n end label 2 x agree x n end label 3 x x agree n total n n n n table 2. confusion matrix for annotation project using end labels since ccr uses primitives rather than end labels, calculating inter-annotator agreement requires additional consideration: while one option is to mimic the approach in table 2 and treat all primitive and distinction value combinations as end labels (e.g., ‘positive causal subjective non-basic,’ ‘negative additive subjective indirect’), it is also possible to calculate agreement for each primitive or distinction individually (e.g., basic operation: causal vs. additive). treating all primitive and distinction value combinations as end labels has the main advantage that it allows for a comparison of the inter-annotator agreement score to those from annotation efforts using another framework (this approach was for instance taken in rehbein et al., 2016).12 however, the ccr ‘end labels’ 12 note, however, that any comparison of inter-annotator agreement scores between annotation efforts using different frameworks should take into account the frameworks’ theoretical and methodological considerations, as well as the specifics of the annotation projects, all of which may lead to higher or lower inter-annotator agreement scores. for instance, annotation of explicit relations tends to lead to more agreement between coders than annotation of implicit relations (miltsakaki, prasad, joshi, & webber, 2004, prasad et al., 2008). in addition, the inter-annotator agreement could be influenced by whether or not a project involved describing the macro-structure of a text, as is common in for hoek, evers-vermeul and sanders 20 that are being compared are not entirely equivalent; relations have minimally three values (e.g., ‘positive additive objective’), but can have up to six values (e.g., ‘positive causal objective basic volitional purpose’). calculating agreement for each primitive or distinction, on the other hand, is much easier than taking an ‘end label’ approach. in addition, it generates a clear overview of where exactly confusions or disagreements arise, which can be extremely valuable for further annotator training. in both rehbein et al. (2016) and hoek et al. (2017), for instance, lowest agreement scores were reported for the source of coherence primitive (81.3%/κ=.63 and 75%-81%/κ=.52-.62, respectively; see also spooren & degand, 2010 for a discussion on reaching sufficient agreement on source of coherence); highest agreement in both studies was reached on order (86.2%/κ=.87 and 94%-100%/κ=.88-1.00, respectively).13 calculating agreement separately for each primitive or distinction makes it impossible, however, to check whether there is a systematic confusion between specific value combinations (e.g., ‘negative objective causal non-basic’ and ‘negative additive subjective indirect’), either because of annotator bias or because of a closer resemblance between two types of relations than the value combinations may suggest (see also section 4.2). in addition, annotations can be dependent on the annotation of the other primitives or distinctions, especially when it comes to distinctions relevant to only a subset of relations. if one coder categorizes a relation as causal, while the other one marks it as additive, the two coders do not have the same number or type of other primitives and distinctions to annotate; coder 1 will for instance have to determine whether the relation is conditional, while coder 2 has to make a decision on the disjunction distinction (see also scholman et al., 2016 on the interdependence of annotations in ccr). when using the full ccr taxonomy in an annotation project, it thus seems worth exploring the inter-annotator agreement both from the perspective of value combinations and for each individual primitive and distinction separately. the combination of both approaches will provide the most informative overview of annotations, as is also illustrated in section 4.2. when calculating interannotator agreement scores, it should be considered whether the annotation process, as well as the configuration in which they are being analyzed, match the assumptions of the inter-annotator agreement statistic used; an inter-annotator agreement statistic may be unequipped to be used for annotations that are not independent or annotations that involve an uneven number of steps.14 4.2 independence of primitives and the use of end labels in addition to primitives while ccr’s primitives are formulated as separate features, in practice the primitives seem to be slightly less independent than they may seem on the basis of the original taxonomy. first of all, the exact operationalization of a specific primitive or distinction can vary depending on other primitive values. as will be elaborated on in section 4.3, determining the source of coherence for conditional relations involves a frequent problem that is much less often encountered in other types of relations: distinguishing between subjectivity and truth value. in addition, agreeing on the basic operation of relations with a positive value for polarity tends to be much easier than instance rst and sdrt. for example, depicting the overall discourse structure of a text involves annotating many higher-order relations (which often appear to be implicit: e.g., hoek et al., 2017, patterson & kehler, 2013) and when annotating the same text, frameworks that connect all segments and chunks of segments into one top-level relation will annotate more relations than frameworks that do not (e.g., demberg et al., 2019). since differences between annotation frameworks or projects may thus influence inter-annotator agreement scores, differences in inter-annotator agreement scores should be interpreted with care; a lower score does not necessarily mean a worse performance by the coders. 13 it should be noted that hoek et al. (2017) did not annotate polarity, since relations were selected on the basis of their connectives, which were all judged to be unambiguous in their polarity. 14 discussing the basic assumptions of commonly used inter-annotator agreement statistics and relating them to the possible ways in which ccr annotations could be analyzed is beyond the scope of this paper, but see for instance zhao, liu, and deng (2013) for a comprehensive overview of the basic assumptions of many inter-annotator agreement statistics. for a more general discussion on inter-annotator agreement in discourse annotation, see for instance spooren and degand (2010) or hoek and scholman (2017). using ccr for discourse annotation 21 determining the basic operation of negative relations; distinguishing between positive additive and positive causal relations is simple compared to distinguishing between negative additive and negative causal relations. another indication that primitives may not always be entirely independent from each other is that annotations may reveal a relatively frequent confusion between two types of relations that differ in multiple values. based on the taxonomy, disagreement between relations that differ in only one value seems much more likely, and this type of confusion was indeed the most frequent type of disagreement in the annotation of the english relations in the parallel corpus (88% of all disagreements). the most common exception was a disagreement between annotators in which one annotator coded the relation as negative causal objective, (non-)basic, while the other coded it as negative additive subjective indirect, or vice versa.15 an example of such a relation can be found in (36). on the one hand, this relation could be analyzed as a negative objective causal relation, since setting targets and deadlines could plausibly lead to those targets and deadlines being met; the relation in (36) could then be analyzed as p leading to not-q. on the other hand, the relation could also be analyzed as a negative subjective indirect additive relation (concession); the conclusion that can be drawn on the basis of s1 is “we are doing great,” while the conclusion that can be drawn on the basis of s2 is “we are not doing so great.” (36) we learned from that programme that implementation was not good enough. we have a solid base of more than 200 legal acts in the environment. [we already have ambitious targets and deadlines in programmes,]s1 but [they have not all been met.]s2 in practice, it can sometimes be harder to distinguish between two types of coherence relations than would be expected on the basis of the primitive and distinction value combinations in the ccr taxonomy. the observation that in practice, primitives are slightly less independent than they may seem to be in the taxonomy makes comparing annotations between coders using value combinations worthwhile. in addition, it makes it useful to explore the operationalization of a specific primitive or distinction within a specific subset of relations, e.g., temporality within causal relations versus additive relations, or source of coherence within conditional versus causal versus additive relations. another possible solution is to use end labels in addition to the primitive value combinations. some types of relations, especially highly specific types of relations, seem to become easier to recognize after the researcher has become more familiar with relations that carry that specific combination of primitive and distinction values. it is for example very likely that inter-annotator agreement on relations with a negative value for polarity can be improved more by focusing on the exact difference between qualifications (negative subjective additive direct; see section 3.1.4) and concessions (negative additive subjective indirect; see section 3.1.4) than by further discussing the individual primitives. occasionally, it may thus seem easier to use end labels during annotation than individual primitives and distinctions. when encountering the relation in (37) in an annotation project using pdtb 2.0 (pdtb research group, 2007), the relation label that should be chosen is fairly straightforward: exception. using ccr, however, determining that (37) is a negative objective additive relation is, by comparison, much less obvious. similarly, attributing a label to a relation like the one in (38) when using carlson and marcu’s (2001) version of rst is simple: otherwise. arriving at an annotation in ccr is much more involved: a negative objective conditional relation with basic order. 15 distinguishing between negative additive and negative causal relations has also been reported as difficult or problematic on the basis of other annotation projects (e.g., robaldo & miltsakaki, 2014; zufferey & degand, 2017). hoek, evers-vermeul and sanders 22 (37) don’t let the internet fool you — making hard boiled eggs in the microwave oven is trouble. if you try to hard boil eggs in your microwave you’re likely to end up with a big mess to clean up. the rapid heat from the microwaves creates a lot of steam in the egg. [the steam has nowhere to go]s1 except [to explode out.]s2 (38) [when adding wine to a sauce, make sure you allow most of the alcohol to cook off;]s1 otherwise, [the sauce may have a harsh, slightly boozy taste.]s2 however, differences in how easy it is to annotate certain types of relations exist not just between ccr and annotation approaches with end labels, but between annotation approaches in general. (37) is simple to categorize using pdtb 2.0, but is much harder to label using carlson and marcu’s (2001) version of rst; (38) is straightforwardly labeled using carlson and marcu’s (2001) annotation scheme, but much more difficult to categorize using pdtb 2.0. in sum, it can be worthwhile exploring which distinctions can be more reliably made when using end labels in addition to the individual primitives when using ccr for discourse annotation. another benefit of using end labels to refer to specific combinations of primitive values is that end labels can make talking about specific relation types much more convenient. it is for instance much easier to talk about result relations than to repeatedly mention ‘positive objective basic order causal relations.’16 in such situations, the most obvious solution would be to define a relation type in terms of ccr primitives and distinctions and give it a single name to refer to the specific relation type. we took this approach ourselves in section 3.1.4 of this paper, where we used qualification to refer to negative additive subjective direct relations and concession to refer to negative additive subjective indirect relations. ccr’s primitive approach is thus not incompatible with the use of end labels. the original ccr proposal by sanders, spooren, and noordman (1992) already gives an overview of possible end labels that can be used to refer to specific combinations of primitive values. being aware of which specific value combinations correspond to which type of end labels also makes it easier to compare ccr to other discourse annotation approaches and existing literature on coherence relations. 4.3 source of coherence versus truth value a common source of confusion in annotating coherence relations in a corpus pertains to the relationship between source of coherence and truth value. objective relations are defined to hold between two events in the real world, but this does not mean that the relation that is established between the two segments is necessarily true. in (39), for instance, the relation signaled by because is a positive volitional objective causal relation in which an soc performs a volitional action for a specific reason. the relation as a whole, however, is a conclusion by the speaker, as is also indicated by so; the speaker makes a conclusion or claim about the unfolding of events in the real world. while the relation between s1 and s2 in (39) is an objective causal relation, the relation between the two combined segments and the rest of the discourse is subjective. we have found that in practice it can sometimes be difficult to distinguish between the source of coherence of the relation that is being annotated and the source of coherence at a higher discourse level. (39) so, [you’re really just apologizing]s1 because [you need my advice.]s2 (40) is a fragment extracted from the europarl corpus. the relation at the end of the fragment is highly similar to the one in (39), although it is slightly more complicated. 16 this example is not meant to imply that result relations in all discourse annotation frameworks correspond to positive objective basic order causal relations (they do not), nor to suggest that all existing end labels in other frameworks can be broken down neatly into ccr primitives (they cannot; see sanders et al. (2018) for a discussion of both issues); its only purpose is to point out that assigning an end label to a set of primitive values to refer to as a ‘short hand’ is perfectly compatible with ccr (see for instance sanders et al., 1992:12-16). using ccr for discourse annotation 23 (40) my group will also support the amendment, which other colleagues and i have signed in the name of the socialist group, for the deletion of paragraph 4. why do we do that? not because we necessarily disagree with the scientific committee on the issue of whether radiation can be safe for foodstuffs, but because it is not the whole story. the question is, why is this being done? is it really for the benefit of the consumer? is it something for which there is a consumer demand? if that was the case, we would not have as many cases as there are, certainly in my country, of illegal and covert irradiation. [this has been carried out]s1 because [they do not want consumers to know about it.]s2 the final sentence of (40) is a positive volitional objective relation, since it holds between an intentional act and a reason for that act. however, from the fragment it is clear that the discourse relation is the speaker’s answer to the question why is this being done? the relation is claimed to be true: the reason for the intentional act is invented or hypothesized by the speaker. this does not, however, mean that the relation itself becomes subjective. internally, the way in which s1 relates to s2 is objective, and without context, there would probably be no confusion. the subjective nature of the final sentence in (40) arrives from, and can be captured by, it as a whole being a claim and part of a subjective relation; the speaker claims that it is being done not for the benefit for the consumer, not because consumers demand it, but rather because consumers are preferred to not know about it. the distinction between source of coherence and truth value seems especially relevant to conditional relations, since they often seem to entail speaker involvement. conditionals, of which the content usually has not been realized, are often predictions. in the relation in (41), for example, the speaker announces what his party will do in a certain scenario. similar to the relations in (39) and (40), the relation between the two segments in (41) is objective, while the relation as a whole is a prediction. here too, the source of coherence within the relation is not the same as the source of coherence of the relations that holds at a higher discourse level, i.e., between the combined segments as a whole and the preceding discourse. (41) if [we find that any member of this house or their employees collaborated with the bbc in this farrago]s1 [we will expose them to the opprobrium of this house.]s2 removing the conditionality from the discourse relation helps when annotating the source of coherence of a conditional relation; if the resulting causal relation is objective, the conditional relation is also objective. without the conditionality, (41) would become ‘there has been a collaboration with the bbc, which is why we will expose them;’ a positive volitional objective causal relation. it is also not uncommon for conditional relations to express the speaker’s negative stance toward the antecedent actually taking place, and therefore toward the entire prediction (sometimes also called counterfactual or irrealis). in english, indicating that something is unlikely to come true can for instance be done by means of a distanced verb form, e.g., if we found that. speakers can also encode that the event did definitely not take place, e.g., if we had found that. even though negative stance seems to emphasize the presence of a speaker, it does not usually influence the source of coherence between the segments of the conditional relation. with negative stance added, the relation in (41) for example still expresses that if one real world event occurs, it leads to another real-world event. in general, an increased awareness of the difference between source of coherence and truth value and about the way in which the source of coherence at a higher discourse level can influence the way in which the source of coherence between two segments is perceived can help improve the quality and reliability of discourse annotation. hoek, evers-vermeul and sanders 24 5 conclusion the cognitive approach to coherence relations was originally proposed as a set of cognitively plausible primitives to order coherence relations, but is also increasingly used as an annotation scheme for classifying coherence relations, see appendix a. in this paper, we gave an overview of the most important developments within ccr from the point of view of discourse annotation. we discussed proposals for new primitives and additional distinctions, and summarized the discussion on how to operationalize an original primitive, source of coherence. in addition, we argued in favor of adding a new distinction to ccr: disjunction. finally, we discussed some practical issues we encountered during a recent annotation project using ccr. as a whole, this paper gives an overview of state-of-the-art ccr for discourse annotation. as such, it can be used, together with the original 1992 proposal, as a point of departure for anyone interested in annotating coherence relations using the cognitive approach to coherence relations. acknowledgements the first and second author’s work was funded by the snsf sinergia project modern (crsii2_147653). the second author’s work was also enabled by a grant awarded by the executive board of utrecht university to the anncor project, work package discourse annotation. using ccr for discourse annotation 25 appendix a: overview of corpora annotated with ccr hoek, evers-vermeul and sanders 26 appendix b: sources of examples (1) judkis, m. (2018, february 5). doritos is developing lady-friendly chips because you should never hear a woman crunch. the washington post. retrieved from https://www. washingtonpost.com/news/food/wp/2018/02/05/doritos-is-developing-lady-friendly-chipsbecause-apparently-you-should-never-hear-a-woman-crunch/?utm_term=.f0f41dcb8328 (2) sedaris, d. (1997). naked. new york, ny: little, brown and company. p.276 (3) jackson, s. (1962). we have always lived in the castle. new york, ny: viking press. p.21 (4) tchou, w. (2017, december 11). tokyo record bar’s riff on the speakeasy. the new yorker. retrieved from https://www.newyorker.com/magazine/2017/12/11/tokyo-record-bars-riff-onthe-speakeasy (5) crane, d., kauffman, m., astrof, j., sikowitz, m., chase, a., & ungerleider, i. (writers) & lazarus, p. (director). the one with the dozen lasagnas, s01e12 [television series episode]. in crane, d. & kauffman, m. (creators), friends. usa: warner bros. television & bright/kauffman/crane productions. (6) foster, g. (producer) & ephron, n. (director). sleepless in seattle [motion picture]. (7) krauss, n. (2005). the history of love. new york, ny: w.w. norton & company. p.77 (8) anderson, w. (producer & director). (2012). moonrise kingdom [motion picture]. (9) ‘wee’ harry potter fest cancelled because of popularity (2017, march 29). retrieved from http://www.bbc.co.uk/newsbeat/article/39431618/wee-harry-potter-fest-cancelled-becauseof-popularity (10) 8 things you could knit/crochet for your wedding (2014, may 7). retrieved from http://ihatecleaning.com.au/8-things-you-could-knitcrochet-for-your-wedding/ (11) s9e57 [television series episode]. in roddam, f. (creator), masterchef australia. australia: fremantlemedia australia, fremantlemedia, & shine australia. (12) ng, c. (2017). little fires everywhere. london: penguin press. p.14 (13) krauss, n. (2005). the history of love. new york, ny: w.w. norton & company. p.104 (14) penguin point (2018). retrieved from http://www.vanaqua.org/experience/exhibits-andgalleries/penguin-point (15) crane, d., kauffman, m., & curtis, m. (writers) & jensen, s. (director). the one where phoebe hates pbs, s05e04 [television series episode]. in crane, d. & kauffman, m. (creators), friends. usa: warner bros. television & bright/kauffman/crane productions. (16) frequently asked questions? (2019). retrieved from https://www.16-25railcard.co.uk/help/ faqs/53/ (17) perry, h. (2017, march). 7 secrets bodybuilders don’t want you to know. retrieved from https://bodyspartan.com/7-secrets-bodybuilders/ (18) luker g. (2014, august). how to use silhouette software to cut your own graphics. retrieved from https://www.theshabbycreekcottage.com/use-silhouette-software-cut-graphics.html (19) last, t.s. (2018, february 15). mayoral election: ‘the game has changed.’ retrieved from https://www.abqjournal.com/1134349/mayoral-election-the-game-has-changed-ex-whileknocking-on-doors-is-still-part-of-a-campaign-social-media-is-vital-today.html (20) crane, d., kauffman, m., & abrams, d. (writers) & mancuso, g. (director). the one where ross can’t flirt, s05e19 [television series episode]. in crane, d. & kauffman, m. (creators), friends. usa: warner bros. television & bright/kauffman/crane productions. using ccr for discourse annotation 27 (21) 104 extra beds for rough sleepers during cold snap (2018, february 26). retrieved from https://www.rte.ie/news/weather/2018/0226/943592-weather-snow/ (22) lozada, c. (2015, june 4). from ‘jill’ to ‘mom’ – inside jill biden’s relationship with beau and hunter. the washington post. retrieved from https://www.washingtonpost.com/ news/book-party/wp/2015/06/04/from-jill-to-mom-inside-jill-bidens-relationship-with-beauand-hunter/?utm_term=.da2586370201 (23) guadiano, a.m. & vlastelica, r. (2018, february 12). stock market notches daily gain, but posts largest weekly drop since early 2016. retrieved from https://www.marketwatch.com/ story/us-stock-futures-rise-as-dow-faces-worst-week-since-the-global-financial-crisis-201802-09 (24) kirshner, r. (writer) & clancy, s. (director), hay bale maze, s7e18 [television series episode]. in sherman-palladino, a. (creator), gilmore girls. usa: dorothy parker drank here productions, hofflund/polone, warner bros. television. (25) hill, n. (2016). the nix. london: picador. p.337 (26) sedaris, d. (2013). let’s explore diabetes with owls. new york, ny: little, brown and company. p. 21 (27) grinsteinner, k. (2018, april 13). ‘you either know it, or you don’t.’ retrieved from http://www.hibbingmn.com/news/local/you-either-know-it-or-you-don-t/article_171785423ebe-11e8-a10e-5ba84094a972.html (28) brant, k. (2018, february 28). wayward waffles: cornmeal and nontraditional additives set these bad boys apart from their all-flour relatives. arkansas online. retrieved from http://www.arkansasonline.com/news/2018/feb/28/wayward-waffles-20180228/ (29) philippines. (2016). retrieved from https://www.reddit.com/r/philippines/comments/3k4wr7/ alleged_fake_uber_car_roaming_around_the_metro/cuv20ju/ (30) are you just feeling lazy or do you need a break? (2008, december 15). retrieved from https://www.dumblittleman.com/are-you-just-feeling-lazy-or-do-you/ (31) petaluma, california (2017). retrieved from https://www.trover.com/d/1ewtk-petalumacalifornia (32) borowitz, a. (2017, october 6). devos defends trump: “would a moron hire me?” the new yorker. retrieved from https://www.newyorker.com/humor/borowitz-report/devos-defendstrump-would-a-moron-hire-me (33) burroughs, e.r. tarzan of the apes. retrieved from https://www.cs.cmu.edu/~rgs/tarz10.html. chapter 10. (34) andrews, r. (2017, november 29). scientist calls for glitter to be banned because it’s awful for the environment. retrieved from http://www.iflscience.com/environment/scientist-callsglitter-banned-awful-environment/ (35) casey, b. (2012, august 30). i had a face tattoo for a week. retrieved from https://www.vice. com/en_us/article/ppqm7v/i-had-a-face-tattoo-for-a-week (36) europarl (koehn, 2005): ep-01-05-30 (ep-year-month-day) (37) thomson, j.r. (2014, june 13). 13 things you should never put in the microwave. retrieved from https://www.huffingtonpost.com/2014/06/13/microwave-cooking-tips_n_5488231.html (38) hanson, c. (2017). why cooking with wine makes food taste better. retrieved from http://dish.allrecipes.com/cooking-wine-makes-food-taste-better/ hoek, evers-vermeul and sanders 28 (39) korsh, a. & cowan, j. (writers) & kumble, r (director). i want you to want me, s03e02 [television series episode]. in aaron korsh (creator), suits. usa: hypnotic, universal cable productions, & dutch oven. (40) europarl (koehn, 2005): ep 02-12-16 (41) europarl (koehn, 2005): ep-00-02-14 using ccr for discourse annotation 29 references elisabeth a. bates (1976). language and context: the acquisition of pragmatics. new york, academic press. birgit bekker (2006). de feiten verdraaid: over tekstvolgorde, talige markering en sprekerbetrokkenheid [twisted facts: about text order, linguistic marking and speaker involvement]. phd thesis, university of tilburg. lois bloom, margaret lahey, lois hood, karin lifter, and kathleen fiess (1980). complex sentences: acquisition of syntactic connectives and the semantic relations they encode. journal of child language 7:235-261. lynn carlson and daniel marcu (2001). discourse tagging reference manual. isi technical report isi-tr-545. herb h. clark (1974). semantics and comprehension. in t.a. sebeok (ed.), current trends in linguistics, vol. 12: linguistics and adjacent arts and sciences: 1291-1498. the hague, mouton.
 liesbeth degand (2001). form and function of causation: a theoretical and empirical investigation of causal constructions in dutch. leuven, peeters. liesbeth degand and henk l.w. pander maat (2003). a contrastive study of dutch and french causal connectives on the speaker involvement scale. in a. verhagen and j. van de weijer (eds.), usage based approaches to dutch: 175-199. utrecht, lot. liesbeth degand and ted j.m. sanders (1999). causal connectives in language use: theoretical and methodological aspects of the classification of coherence relations and connectives. levels of representation in discourse, working notes of the international workshop on text representation, pages 3-12. edinburgh, united kingdom. vera demberg, merel c.j. scholman, and fatimeh t. asr (2019). how compatible are our discourse annotations? insights from mapping rst-dt and pdtb annotations. dialogue & discourse 10(1):87-135. hanny n. den ouden (2004). prosodic realizations of text structure. phd thesis, university of tilburg. ann r. eisenberg (1980). a syntactic, semantic and pragmatic analysis of conjunction. stanford papers and reports on child language development 19:70-78. jacqueline evers-vermeul (2005). the development of dutch connectives: change and acquisition as windows on form-function relations. phd thesis, utrecht university. utrecht, lot. available online at https://www.lotpublications.nl/documents/110_fulltext.pdf. jacqueline evers-vermeul, jet hoek, and merel c.j. scholman (2017). on temporality in discourse annotation: theoretical and practical considerations. dialogue & discourse 8(2):1-20. jacqueline evers-vermeul and ted j.m. sanders (2009). the emergence of dutch connectives: how cumulative cognitive complexity explains the order of acquisition. journal of child language 36(4):829-854. jacqueline evers-vermeul and ninke m. stukker (2003). subjectificatie in de ontwikkeling van causale connectieven? de diachronie van daarom, dus, want en omdat. [subjectification in the development of causal connectives? the diachrony of daarom, dus, want and omdat.] gramma/ttt, tijdschrift voor taalwetenschap 9(2/3):111-139. l.t.f. gamut (1982). logica, taal en betekenis [logic, language and meaning] (vol.2). utrecht, spectrum. barbara j. grosz and candace l. sidner (1986). attention, intentions and the structure of discourse. computational linguistics 12:175-204. michael a.k. halliday and ruqaiya hasan. (1976). cohesion in english. london & new york, routledge. jerry r. hobbs (1978). why is discourse coherent? technical report 176, sri international. jerry r. hobbs (1979). coherence and coreference. cognitive science 3:67-90. hoek, evers-vermeul and sanders 30 jerry r. hobbs (1990). literature and cognition. stanford, ca, csli. jet hoek (2018). making sense of discourse: on discourse segmentation and the linguistic marking of coherence relations. phd thesis, utrecht university. utrecht: lot. available online at https://www.lotpublications.nl/documents/509_fulltext.pdf. jet hoek, jacqueline evers-vermeul, and ted j.m. sanders (2018). segmenting discourse: incorporating interpretation into segmentation? corpus linguistics and linguistic theory 14(2):357-386. jet hoek and merel c.j. scholman (2017). evaluating discourse annotation: some recent insights and new approaches. proceedings of the 13th joint iso-acl workshop on interoperable semantic annotation (isa-13), pages 1-13, toulouse, france. jet hoek, sandrine zufferey, jacqueline evers-vermeul, and ted j.m. sanders (2017). cognitive complexity and the linguistic marking of coherence relations: a parallel corpus study. journal of pragmatics 121:113-131. eduard h. hovy and elisabeth maier (1995). parsimonious or profligate: how many and which discourse structure relations? unpublished manuscript. andrew kehler (1995). interpreting cohesive forms in the context of discourse inference. phd thesis, harvard university. andrew kehler (2002). coherence, reference, and the theory of grammar. stanford: csli. alistair knott and robert dale (1994). using linguistic phenomena to motivate a set of coherence relations. discourse processes 18(1):35-62. alistair knott and ted j.m. sanders (1998). the classification of coherence relations and their linguistic markers: an exploration of two languages. journal of pragmatics 30(2):135-175. phillip koehn (2005). europarl: a parallel corpus for statistical machine translation. proceedings of the tenth machine translation summit (mt summit x), phuket, thailand. ronald w. langacker (1990). subjectification. cognitive linguistics 1(1):5-38. ronald w. langacker (1991). cognitive grammar. in f.g. droste and j.e. joseph (eds.), linguistic theory and grammatical description: nine current approaches: 275-306. amsterdam, john benjamins. ronald w. langacker (2006). subjectification, grammaticization, and conceptual archetypes. in a. athanasiadou, c. canakis, and b. cornillie (eds.), subjectification: 17-40. berlin, mouton de gruyter. stephen c. levinson (1983). pragmatics. cambridge, cambridge university press. fang li (2014). subjectivity in mandarin chinese. phd thesis, utrecht university. utrecht, lot. available online at https://www.lotpublications.nl/documents/365_fulltext.pdf. fang li, jacqueline evers-vermeul, and ted j.m. sanders (2013). subjectivity and result marking in mandarin: a corpus-based investigation. chinese language and discourse 4(1):74-119. fang li, ted j.m. sanders, and jacqueline evers-vermeul (2016). on the subjectivity of mandarin reason connectives: robust profiles or genre-sensitivity? in n.m. stukker, w.p.m.s. spooren, and g.j. steen (eds.), genre in language, discourse and cognition: 15-49. berlin, mouton de gruyter. robert e. longacre (1983). the grammar of discourse. new york, plenum press. william c. mann and sandra a. thompson (1988). rhetorical structure theory: toward a functional theory of text organization. text 8(3):243-281. james martin (1992). english text: systems and structure. amsterdam, john benjamins. eleni miltsakaki, rashmi prasad, aravind joshi, and bonnie webber (2004). the penn discourse treebank. in proceedings of the 4th international conference on language resources and evaluation (lrec 2004), pages 2237-2240, lisbon, portugal. john d. murray (1997). connectives and narrative text: the role of continuity. memory and cognition 25(2):227-236. henk l.w. pander maat (1998). classifying negative coherence relations on the basis of linguistic evidence. journal of pragmatics 30(2):177-204. using ccr for discourse annotation 31 henk l.w. pander maat (2002). tekstanalyse [text analysis]. bussum, coutinho. henk l.w. pander maat and liesbeth degand (2001). scaling causal relations and connectives in terms of speaker involvement. cognitive linguistics 12(3):211-246. henk l.w. pander maat and ted j.m. sanders (2000). domains of use or subjectivity? the distribution of three dutch causal connectives explained. topics in english linguistics 33:5782. gary patterson and andrew kehler (2013). predicting the presence of discourse connectives. proceedings of the 2013 conference on empirical methods in natural language processing, pages 914-923, seattle, wa, usa. ingrid persoon, ted sanders, hugo quené, and arie verhagen (2010). een coördinerende omdatconstructie in gesproken nederlands? tekstlinguïstische en prosodische aspecten [a coordinating omdat-construction in spoken dutch? textlinguistic and prosodic aspects]. nederlandse taalkunde 15:259-282. pdtb research group (2007). the penn discourse treebank 2.0 annotation manual. ircs technical report. mirna pit (2003). how to express yourself with a causal connective: subjectivity and causal connectives in dutch, german and french. amsterdam/new york, rodopi. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, avarind joshi, and bonnie webber (2008). the penn discourse treebank 2.0. proceedings of the 6th international conference on language resources and evaluation (lrec’08), pages 2961-2968, marrakech, morocco. gisela redeker (1990). ideational and pragmatic markers of discourse structure. journal of pragmatics 14(3):367-381. brian reese, julie hunter, nicholas asher, pascal denis, and jason baldridge. (2007). reference manual for the analysis and annotation of rhetorical structure. unpublished manuscript. university of texas at austin, austin, tx. ines rehbein, merel c.j. scholman, and vera demberg (2016). annotating discourse relations in spoken language: a comparison of the pdtb and ccr frameworks. proceedings of the 10th international conference on language resources and evaluation (lrec 2016), pages 10391046, portoroz, slovenia. livio robaldo and eleni miltsakaki (2014). corpus-driven semantics of concession: where do expectations come from? dialogue & discourse 5(1):1-36. josé sanders, ted j.m. sanders, and eve e. sweetser (2012). responsible subjects and discourse causality: how mental spaces and perspective help identifying subjectivity in dutch backward causal connectives. journal of pragmatics 44(2):191-213. josé sanders and wilbert p.m.s. spooren (1997). perspective, subjectivity, and modality from a cognitive linguistic point of view. in w. liebert, g. redeker, and l.r. waugh (eds.), discourse and perspective in cognitive linguistics: 85-112. amsterdam, john benjamins. ted j.m. sanders (1997). semantic and pragmatic sources of coherence: on the categorization of coherence relations in context. discourse processes 24(1):119-147. ted j.m. sanders, vera demberg, jet hoek, merel c.j. scholman, fatimeh t. asr, sandrine zufferey, and jacqueline evers-vermeul (2018). unifying dimensions in discourse relations: how various annotation frameworks are related. corpus linguistics and linguistic theory. first view online, doi: 10.1515/cllt-2016-0078 ted j.m. sanders and jacqueline evers-vermeul (2019). subjectivity and causality in discourse and cognition; evidence from corpus analyses, acquisition and processing. in ó. loureda, i. recio fernández, l. nadal, and a. cruz (eds.), empirical studies of the construction of discourse: 273-298. amsterdam/philadelphia, john benjamins. ted j.m. sanders, josé sanders, and eve e. sweetser (2009). causality, cognition and communication: a mental space analysis of subjectivity in causal connectives. in t.j.m. hoek, evers-vermeul and sanders 32 sanders and e.e. sweetser (eds.), causal categories in discourse and cognition: 19-60. berlin, mouton de gruyter. ted j.m. sanders and wilbert p.m.s. spooren (2009). causal categories in discourse: converging evidence from language use. in t.j.m. sanders and e.e. sweetser (eds.), causal categories in discourse and cognition: 19-60. berlin, mouton de gruyter. ted j.m. sanders and wilbert p.m.s. spooren (2013). exceptions to rules: a qualitative analysis of backward causal connectives in dutch naturalistic discourse. text & talk 33(3):377-398. ted j.m. sanders and wilbert p.m.s. spooren (2015). causality and subjectivity in discourse: the meaning and use of causal connectives in spontaneous conversation, chat interactions and written text. linguistics 53(1):53-92. ted j.m. sanders, wilbert p.m.s. spooren, and leo g.m. noordman (1992). toward a taxonomy of coherence relations. discourse processes 15(1):1-35. ted j.m. sanders, wilbert p.m.s. spooren, and leo g.m. noordman (1993). coherence relations in a cognitive theory of discourse representation. cognitive linguistics 4(2):93-133. ted j.m. sanders and carel h. van wijk (1996). pisa—a procedure for analyzing the structure of explanatory texts. text 16(1):91-132. andrea santana, wilbert p.m.s. spooren, dorien nieuwenhuijsen, and ted j.m. sanders (2018). subjectivity in spanish discourse: explicit and implicit causal relations in different contexts. dialogue & discourse 9(1):163-191. merel c.j. scholman (2019). coherence relations in discourse and cognition: comparing approaches, annotations, and interpretations. phd thesis, saarland university. available online at https://publikationen.sulb.uni-saarland.de/bitstream/20.500.11880/27370/1/disserta tionscholman.pdf merel c.j. scholman, jacqueline evers-vermeul, and ted j.m. sanders (2016). a step-wise approach to discourse annotation: towards a reliable categorization of coherence relations. dialogue & discourse 7(2):1-28. wilbert p.m.s. spooren (1989). some aspects of the form and interpretation of global contrastive coherence relations. phd thesis, k.u. nijmegen. wilbert p.m.s. spooren and liesbeth degand (2010). coding coherence relations: reliability and validity. corpus linguistics and linguistic theory 6(2):241-266. wilbert p.m.s. spooren and ted j.m. sanders (2008). the acquisition order of coherence relations: on cognitive complexity in discourse. journal of pragmatics 40(12):3-26. wilbert spooren, ted sanders, mike huiskes, and liesbeth degand (2010). subjectivity and causality: a corpus study of spoken language. in s. rice and j. newman (eds.), empirical and experimental methods in cognitive/functional research: 241-255. chicago, csli publications. adrian staub and charles clifton (2006). syntactic prediction in language comprehension: evidence from either...or. journal of experimental psychology: learning, memory, and cognition 32(2):425-436. ninke m. stukker (2005). causality marking across levels of language structure. phd thesis, utrecht university. utrecht, lot. available online at https://www.lotpublications.nl/documents/118_fulltext.pdf. ninke m. stukker and ted j.m. sanders (2012). subjectivity and prototype structure in causal connectives: a cross-linguistic perspective. journal of pragmatics 44(2):169-190. ninke m. stukker, ted j.m. sanders, and arie verhagen (2008). causality in verbs and in discourse connectives: converging evidence of cross-level parallels in dutch linguistic categorization. journal of pragmatics 40(7):1296-1322. eve e. sweetser (1990). from etymology to pragmatics: the mind-body metaphor in semantic structure and semantic change. cambridge, cambridge university press. teun a. van dijk (1977). text and context. new york, longman. using ccr for discourse annotation 33 rosie van veen (2011). the acquisition of causal connectives: the role of parental input and cognitive complexity. phd thesis, utrecht university. utrecht: lot. available online at https://www.lotpublications.nl/documents/286_fulltext.pdf. rosie van veen, jacqueline evers-vermeul, ted j.m. sanders, and huub van den bergh (2014). “why? because i’m talking to you!” parental input and cognitive complexity as determinants of children’s connective acquisition. in h. gruber and g. redeker (eds.), the pragmatics of discourse coherence: theories and applications: 209-242. amsterdam/philadelphia, john benjamins. kirsten vis (2011). subjectivity in news discourse: a corpus linguistic analysis of informalization. phd thesis, vu amsterdam. peter c. wason and philip johnson-laird (1971). psychology of reasoning: structure and content. cambridge, ma, harvard university press. yipu wei (2018). causal connectives and perspective markers in chinese: the encoding and processing of subjectivity in discourse. phd thesis, utrecht university. utrecht, lot. available online at https://www.lotpublications.nl/documents/482_fulltext.pdf. hongling xiao, fang li, ted j.m. sanders, and wilbert p.m.s. spooren (to appear). how subjective are mandarin reason connectives? a corpus study of spontaneous conversation, microblog and newspaper discourse. to appear in language and linguistics 22(1). deniz zeyrek and murathan kurfali (2017). tdb 1.1: extensions on turkish discourse bank. proceedings of the 11th linguistic annotation workshop (law xi), pages 76-81, valencia, spain. xinshu zhao, jun s. liu, and ke deng (2013). assumptions behind intercoder reliability indices. in charles t. salmon (ed.), communication yearbook 36: 419-480. new york, routledge. yuping zhou and nianwen xue (2015). the chinese discourse treebank: a chinese corpus annotated with discourse relations. language resources and evaluation 49(2):397-431. sandrine zufferey and liesbeth degand (2017). annotating the meaning of discourse connectives in multilingual corpora. corpus linguistics and linguistic theory 13(2):399-422. using the cognitive approach to coherence relations for discourse annotation 1 introduction 2 the cognitive approach to coherence relations 2.1 polarity 2.2 basic operation 2.3 source of coherence 2.4 order of the segments 3 extensions of the original ccr taxonomy 3.1 additional distinctions within original primitives 3.1.1 temporality 3.1.2 volitionality 3.1.3 purpose 3.1.4 directness [i hate vegetables,] but [snap peas are okay.] 3.2 proposing a new distinction: disjunction 3.2.1 disjunctions in ccr 3.2.2 disjunction as a new distinction in ccr it is possible to distinguish between the inclusive and exclusive disjunctions using the polarity primitive (see also sanders et al., 2018). since the two segments can hold at the same time, inclusive disjunctions have a positive value for polarity: p... the source of coherence primitive applies to disjunctions as it does to other types of coherence relations; disjunctions can be either objective or subjective. in ‎(27), the segments present alternatives that hold in the real world. as such, ‎(27) has... 3.3 operationalizing source of coherence: segment-internal distinctions 3.3.1 identity of the subject of consciousness 3.3.2 explicit presence of the subject of consciousness not only the identity of the soc, but also the extent to which the soc is explicitly present in the relation has been argued to bear on the subjectivity of a relation. langacker (1990, 1991, 2006) proposes that utterances with an explicitly mentioned,... (34) [i think all glitter should be banned,]s1 because [it’s microplastic.]s2 (34’) [all glitter should be banned,]s1 because [it’s microplastic.]s2 3.3.3 propositional attitude of the segments a final feature of coherence relations that is relevant to its source of coherence is the propositional attitude of the segments (e.g., li, 2014; li et al., 2016; sanders & spooren, 2009, 2015; spooren & degand, 2010); are they, for instance, judgment... 3.4 state-of-the-art ccr for discourse annotation this section gave an overview of the most important developments in ccr since the original 1992 proposal when it comes to discourse annotation. figure 4 provides a schematic overview of state-of-the-art ccr for discourse annotation. the overview is a ... 4 ccr as an annotation scheme in practice 4.1 calculating inter-annotator agreement 4.2 independence of primitives and the use of end labels in addition to primitives 4.3 source of coherence versus truth value 5 conclusion acknowledgements appendix a: overview of corpora annotated with ccr appendix b: sources of examples references dialogue & discourse issue 8(2) (2017) 105–128 doi: 10.5087/dad.2017.205 an empirical analysis of subjectivity and narrative levels in personal weblog storytelling across cultures reid swanson rswanson@ict.usc.edu usc institute for creative technologies playa vista, ca, 90094 andrew s. gordon gordon@ict.usc.edu usc institute for creative technologies playa vista, ca, 90094 peter khooshabeh khooshabeh@ict.usc.edu the army research lab west playa vista, ca, 90094 kenji sagae sagae@ucdavis.edu uc davis department of linguistics davis, ca 95616 richard huskey huskey.29@osu.edu the ohio state university school of communication columbus oh, 43210 michael mangus mangus@icb.ucsb.edu uc santa barbara institute for creative biotechnologies santa barbara, ca, 93106 ori amir oriacadem@gmail.com uc santa barbara institute for creative biotechnologies santa barbara, ca, 93106 rene weber renew@comm.ucsb.edu uc santa barbara department of communication santa barbara, ca, 93106 editor: barbara di eugenio submitted 12/2016; accepted 11/2017; published online 11/2017 abstract storytelling is a universal activity, but the way in which discourse structure is used to persuasively convey ideas and emotions may depend on cultural factors. because first-person accounts of life experiences can have a powerful impact in how a person is perceived, the storyteller may instinctively employ specific strategies to shape the audience’s perception. hypothesizing that some of the differences in storytelling can be captured by the use of narrative levels and subjectivity, we analyzed over one thousand narratives taken from personal weblogs. first, we compared stories from three different cultures written in their native languages: english, chinese and farsi. second, we examined the impact of these two discourse properties on a reader’s attitude and behavior toward the narrator. we found surprising similarities and differences in how stories are structured along these two c©2017 reid swanson, andrew s. gordon, peter khooshabeh, kenji sagae, richard huskey, michael mangus, oria amir, rene weber this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). analysis of subjectivity and narrative levels in storytelling across cultures dimensions across cultures. these discourse properties have a small but significant impact on a reader’s behavioral response toward the narrator. keywords: discourse analysis, narrative, storytelling, culture, persuasion 1. introduction personal stories, the kind we tell about our daily lives and experiences, are deeply connected to the way we perceive and understand the world (bruner, 1991; gerrig, 1993) and are ubiquitous across cultures. they are a rich and complex form of discourse that not only allow us to communicate our experiences, but are also used to evoke emotional responses and change people’s beliefs in targeted ways. one of the reasons they are so powerful is in their flexibility to adapt to the communicative and persuasive goals of the narrator. the same underlying set of events can be told as a story many different ways that preserves the basic semantic meaning, but could have substantially different content and make entirely different points. this is often referred to as the rashomon effect, in reference to the 1950 akira kurosawa film where a sexual encounter and death are witnessed by four characters each with a unique, and dramatically different, view of the same fundamental events. stories are not merely a sequential retelling of events, but critically provide the subjective evaluations why we should care, especially with regard to the plans, goals and motivations of the characters (labov and waletzky, 1967). therefore subjective language, which expresses opinions, emotions, thoughts, preferences, and other mental states of the narrator, is crucial for delivery of the intended interpretation of a personal story. similar to how a soundtrack can set a specific mood in film to heighten the emotional impact of the sights and sounds of a story, skilled rhetoric can also serve to enhance the impact of events depicted in writing. the discourse of a story is also characterized by narrative levels. a narrative is not only composed of discourse that takes place during the time of events, it also intertwines utterances that speak directly to the audience to provide necessary context as well as the narrated events taking place at the time of the experience. for example, a narrator might step outside the story for moment to provide additional information about the people in them to provide context for understanding their actions, especially if the narrator is aware that the audience has never met them. the study of narratology has identified many devices that can be used in this way to emphasize aspects of the story the narrator finds important and to highlight a particular point. our main goal for this research is to examine how different cultures construct personal stories using these devices and, if there are differences across cultures, test if stories that violate cultural norms impact the attitudes of the readers. there is some evidence that people embody at least some of their cultural values, which leads to different envisionments while reading and tell stories. for example, leung and cohen (2007) showed that american’s typically take the perspective of the main character or 1st person point of view corresponding to a high value of the self in social situations, while asian cultures more often take 3rd person point of view because of their high value in being seen by others as doing the right thing. these fundamental differences in how people across cultures envision and interpret a story may play a role in the use of narrative devices to construct stories and how a particular narrative is interpreted when reading or hearing one. 105 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber in this paper we focus on two of these types of narrative properties: subjectivity, which we define as portions of a story based on whether it expresses an unverifiable private mental state following the work of wiebe et al. (2004), and narrative levels, which categorizes content based on whether a clause describes events at the time of the story or something external to the story world (genette, 1980). we have chosen these properties because they are fundamental concepts central to many theories of narrative and they are simple enough that people can identify them in text. because the use of narrative levels and subjectivity is tied closely to how a storyteller shapes an audience’s perception of a set of events (genette, 1980), analysis of narratives along these two dimensions may provide insight on the differences in narrative structure across cultures and whether these structural properties matter to the reader’s receptiveness to the narrator and her message. narrative messages are often constructed with a particular cultural group in mind, either one’s own or another. when the message is targeted at another cultural group, for example public service announcements, translated literature and advertising, the subtle differences in the way subjective and narrative levels are used across these cultures could have unintended consequences in the way the target culture interprets the message. understanding the role of these types of narrative discourse properties could help improve the receptiveness of targeted messages across cultures and guide future automated narrative generation systems in discourse planning for targeted audiences. to examine these issues, we performed two exploratory experiments to address these questions based on the analysis of over one thousand narratives we collected for this study from three different cultures (chinese, american, and iranian). in the first, we found that there are both similarities in the way stories are structured according to these properties as well as significant differences. the most substantial differences occurred in the interactions between clause types rather than considering the distribution of each clause type independently. in the second, we examined the effects of these discourse properties on the attitudes and behavior of the participant toward the narrator for a subset of our cultures. in both cases we found that the subjective distribution of clauses has an impact in the attitudes of the reader toward the narrator. however, we were surprised that the attitudes of the chinese participants decreased monotonically as the distribution of subjective clauses increased despite our expectation that their attitudes should peak at the cultural norm and decrease as the distribution diverged from this value. 2. background 2.1 personal narratives in this research our focus is on written forms of personal narrative (text documents) extracted from internet weblogs. a personal narrative is broadly defined as the non-fiction stories that people share with each other about their own life experiences. they are primarily about a specific event in the life of the narrator that lasts for minutes or hours (not generalizations over years) and more than a chronological sequence of events. the narrator must also provide some evaluative context about why the events were important enough to share and why the audience should care. the phenomenal rise of personal weblogs has afforded new opportunities to collect and study electronic texts of personal narratives on a large scale. while blogging is popularly 106 analysis of subjectivity and narrative levels in storytelling across cultures associated with high-profile celebrities and political commentators, the typical weblog takes the form of a personal journal, read by a small number of friends and family (munson and resnick, 2011). as with the adoption of other forms of electronic communication, personal narratives in weblogs take on several new characteristics in adapting to a social media environment that is increasingly public and interconnected. eisenlauer and hoffman (eisenlauer, 2010) argue that the on-going technological development of weblog software has led to an increase of collaborative narration, moving the form further toward ochs and capps (ochs and capps, 2001) conception of the hypernarrative, where discourse is best understood as a conversation among multiple participants. langellier and peterson (langellier and peterson, 2004) characterize this collaborative narration as a form of public performance, creating a productive paradox between the insincerity needed to craft a good story and the sincerity of the blogger as a character in the narrated events. this productive paradox seen in weblog storytelling helps distinguish personal narrative from other narrative forms. similar to most narratives personal narratives consist in descriptions of events that are causally related, but have the additional requirement that the narrator was a participant in events. the expectation is that the narrator describes events that actually took place, namely it is a non-fiction story, but that some license to stretch the truth may be taken for dramatic or persuasive purposes. the genre of personal narrative also includes the stories told among family members while reviewing old photographs (chalfen, 1987), the accounts shared among coworkers in office environments (coopman and meidlinger, 1988), the testimonials of people in interviews (gordon, 2005), and the reflections of daily experiences of people written to private diaries (steinitz, 1997). personal stories found on weblogs share many of the fundamental properties of the other types of discourse in this genre, but are widely available on the web across many different cultures. 2.2 analysis of subjective language a personal story, almost by definition, is a mixture of describing what actually happened, i.e. the plot, and the narrator’s evaluations providing the reason why they are important. a large portion of these evaluations are subjective in nature that help to frame the story in terms of the narrator’s opinions, emotions and preferences that can shape the reader’s interpretation of events. exactly defining what it means for an utterance to be subjective is a challenging task in its own right, which is one reason this area has received little attention in the computational linguistics community until fairly recently. however, over the last 10 years, primarily due to the rise of online shopping and review sites, automatic sentiment analysis and opinion mining has become a rapidly growing area of interest. because of the challenging nature of this task, the problem is often greatly simplified and only a small subset of this overall space has been investigated. in this section we will describe several prominent approaches in the field and the definition we adopt for our work. to help reduce the complexity, the types of subjective content under consideration are often highly constrained and restricted to very narrow definitions of what subjectivity means for the task. for example, one of the most common formulations of the problem is to frame it as binary classification task to determine if a text is expressing a positive or negative attitude toward a particular topic. more complex formulations that move beyond simple 107 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber polarity, for example categorizing texts based on a range emotional states, is still much less common. so far, the bulk of this kind of research has been focused on the automatic analysis of only a few genres of discourse, with the most prominent example being reviews of movies, products, restaurants, etc. (e.g. (pang and lee, 2005; blitzer et al., 2007; snyder and barzilay, 2007)). reviews are an attractive target of sentiment analysis, as they are abundantly available online, they restrict language processing tasks to well-defined domains, and they necessarily express opinions that can often be binned into negative or positive categories relatively easily. within the context of automatically analyzing reviews, it is common to frame the task as sentiment polarity classification (“thumbs up” vs “thumbs down”), often aided by a preprocessing step that identifies subjective language, which pang and lee (2008) define simply as opinion-oriented language. another language genre where subjectivity and sentiment analysis has been studied extensively is news, where the identification of subjectivity is itself the target of analysis, rather than binary classification of sentiment polarity. in their work on subjectivity analysis, wiebe et al. (wiebe et al., 2004) take a broader view of subjective language, which they define as the expression of private states (quirk et al., 1985), which includes emotions, opinions, evaluations and speculations. a third major area of application of sentiment and subjectivity analysis, which has been growing rapidly, is user-generated content, including twitter, discussion boards, political weblogs, and youtube video reviews (e.g. (agarwal et al., 2011; melville et al., 2009; morency et al., 2011)). although far from exhaustive, the list of language genres mentioned above serves to illustrate how the goals of subjectivity analysis can vary widely when different types of content are considered. for example, in reviews it is more important to determine whether statements are positive or negative, while in news there is a greater focus on separating opinion from fact. even though goals and even definitions may vary, the most common types of application are related to fulfilling information needs or estimating public interest and opinion regarding specific issues, products, etc. in the case of narrative, however, analysis of subjectivity and sentiment can play a different type of role. characterizing the mental and emotional state of the narrator is crucial for analyzing the intent of a narrative beyond the facts and events of the story; narratives are often crafted with the explicit goal to have an emotional impact on the reader, sometimes more so than they are to convey a specific sequence of events. in contrast to the main role of subjectivity in reviews or editorial pieces, subjective language in narrative goes far beyond opinions. the expression of emotions, thoughts, preferences and other mental and emotional states is of primary importance. in our work, we adopt wiebe et al.’s notion of subjective language as the linguistic expression of private states (including opinions, evaluations, emotions, and speculations), which are experienced but are not open to external observation or verification by others. our main focus is on private states of the narrator, since we are dealing with personal narratives, which express the narrator’s point of view. while it can be tempting to define subjective language as the statement of opinions, in contrast to objective statement of facts, this would be an imprecise definition. for example, while the text segment i know her name may be considered a statement of fact by the narrator, it is a case of subjective language. the key issue here is not whether a statement is true or made with certainty or privileged knowledge, or even whether it can be considered a fact, and rather whether it 108 analysis of subjectivity and narrative levels in storytelling across cultures expresses a private state and not something that can be observed or measured objectively and externally. for example, while i felt sick is a subjective statement, since it cannot be observed externally, the statement i had a 102-degree fever is objective. similarly, it was hot yesterday is subjective (the narrator’s opinion), while it was 95 degrees yesterday is objective. a statement such as he was sad could be subjective because it expresses a private state of a third person, or because it expresses the narrator’s opinion or evaluation of a third person. this distinction is not necessary for our current work and introduces several complexities, such as decreasing inter-rater agreement. in either case the statement is considered subjective, which is sufficient for our work. on the other hand, he said he was sad is an objective statement, since it describes an event that can be observed externally (namely, the act of saying). 2.3 narrative levels when telling a story, the utterances of the narrator are situated in a particular time and location. while the primary purpose of a personal narrative is to describe events that occurred in the past, every utterance is not necessarily referencing this time frame. for example, in some instances the narrator may step outside the story to speak directly to the audience in the present time and location. consider, for example, a narrative that recounts events that include the narrator being afraid of a puppy and disliking dogs as a child, but also expresses the now adult narrator’s current embarrassment of this long abandoned fear and current fondness for dogs. the universe of the story, where the narrator is a child, is sometimes referred to as the diegetic (or intradiegetic) level, and the act of narration is performed at the extradiegetic level, where in this case the narrator is an adult addressing the reader. this example includes expression of several private states experienced by the narrator: as a character in the story (i.e. at the diegetic level), the narrator experiences fear at a specific moment, and holds a negative preference for dogs; in contrast, the narrator expresses the private states of embarrassment and positive preference for dogs at the time of storytelling (i.e. at the extradiegetic level). in our discussion we adopt the terms proposed by genette 1980 to speak of narrative levels in narrative. although genette proposes an additional narrative level, meta-diegetic, which concerns embedded narratives, this category is rarely used in personal narratives and was collapsed within the diegetic level. in other words, instead of performing a complete analysis of diegetic levels, we make only a binary distinction between the extradiegetic level and all other (intradiegetic) levels, with no distinction made in the levels of embedded narratives. an alternative way to characterize what we refer to as private states at the diegetic and extradiegetic levels is to use the notion of time points due to reichenbach (1947): the narrator might refer to private states at speech time (at the time of narration), or at the event or reference time. this distinction reflects the narrator’s exclusive advantage in framing the story to influence the audience’s interpretation and reaction. the impact of diegetic and extradiegetic material can be understood intuitively by considering the soundtrack in a movie. when watching a movie, we observe events taking place and a story unfolding, which may evoke emotion. external to the universe where the story takes place, however, we may also hear music (e.g. romantic music for a romantic 109 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber scene, or fast-paced music for a car chase), which sets a specific mood and serves to evoke or amplify emotional reactions. this music is at the extradiegetic level: it is audible to the audience only, and does not exist for the characters in the story. 3. experiment 1: narrative discourse structure across cultures it is clear that the content of stories will vary across cultures, reflecting the differences in daily activities and values each culture emphasizes. it is not as obvious if, or in what ways, the narrative levels and subjectivity vary across cultures. for our experiments we examined the narrative levels and subjectivity as they are expressed at the granularity of a clause and will subsequently refer to these two types of discourse features as narrative clauses. each of these clause types can completely partition the discourse and together are fundamental to the characterization of narrative as a genre. for brevity, we subsequently refer to these clause types as narrative clauses. in our first experiment we investigate this question through an empirical study that compares the distribution of narrative clauses across three different cultures over a large number of personal stories found on the web. in our first experiment we investigated the usage of narrative clauses across three different cultures and how the distributions of clauses differed among them. the basic approach proceeded in four stages. first, we trained an automated classifier to identify personal stories from internet weblogs in three different languages. second, we applied our classifier to automatically collect a large sample of personal stories across our three target cultures. third, we randomly sampled a subset of these documents and manually verified the documents were actually personal stories, which could be used for further annotation and analysis. fourth, we manually annotated the stories along the diegetic and subjective dimensions described above. fifth, we analyzed the distribution of clauses across the cultures to identify similarities and differences between them. 3.1 large scale personal narrative data collection the first step of our process required a large set of personal stories from the web across our three target cultures: american, chinese and iranian. there are several criteria that such a collection should posses for our experiments. first, the collection should be large enough to contain a representative sample of the types of activities people in that culture are likely to engage in. second, the process from which these stories are drawn should be similar across cultures, so that we do not introduce any unnecessary biases. since there is no existing repository that exclusively consists of personal narratives meeting these criteria, we started by collecting candidate weblog posts using an automated system to identify personal narratives from arbitrary weblog posts following the approach of gordon and swanson (2009). in this section we describe the automated approach for identifying our large, but noisy, collection of personal narratives. in section 4 we will describe our process for sampling a subset of this corpus for narrative clause annotation and, because the automated set is noisy, manually curating the documents to ensure only personal narratives were included. in gordon and swanson’s approach the problem of personal narrative identification is formulated as a standard document classification task where a weblog post is either labeled story if the majority of discourse on the page is a personal story adhering to our 110 analysis of subjectivity and narrative levels in storytelling across cultures definition in section 2.1 or not story if it is not. after developing a set of annotation guidelines based on the definition of a personal story, the agreement between annotators on this task for english narratives was 0.68 using cohen’s κ to adjust for chance agreement. 4,985 unique posts were labeled by 2 annotators with 5.4% of the posts being identified as personal stories. the labeled data was then used to train a statistical classifier. we used the standard classification measures precision, recall and f1-score to evaluate the performance of the system resulting in a precision of 0.66, recall of 0.48, and f1-score of 0.55. these measures attempt to provide more information than the accuracy of the system, which can be highly misleading when the class labels are unevenly distributed. precision measures how often the system gets the correct answer based on how many guesses it made and is defined as true positives true positives + false positives . recall measures how many positive examples the system identified out of all the possible positive examples and is defined as true positives true positives + false negatives . the f1-score is the harmonic mean between the two, which has a high value when both precision and recall are high and penalizes them when they are different from each other. despite the somewhat low performance, our results are near state-of-the art for this type of documentation classification problem and suitable for tasks tolerant to some noise or can incorporate manual filtering like this one. for the two experiments described in this work we used a new corpus of english stories following the same approach as gordon and swanson. this corpus was created by applying the classifier to every english weblog post identified by spinn3r.com from january 2010 to september 2014, resulting in 25 million unique stories. chinese stories were identified by following a similar procedure previously validated by gordon et al. 2013. instead of using spinn3r, we extracted posts from the weblog hosting provider sina (blog.sina.cn) identified by member lists and a search over the commoncrawl.org index. this identified 57,496 unique blogs and 995,661 total posts. topics discussed by users on this hosting service span the full breadth of chinese society, with increased attention on chinese popular culture and the entertainment industry. most posts appearing on sina’s weblog hosting service are written in standard chinese (mandarin) using the simplified chinese character set. we chose sina blogs primarily because they provide a public directory listing to the blogs of over ten thousand of their users. by crawling this directory, we gathered the urls for each user’s weblog. unfortunately, the xml rss feeds provided by sina for each of these weblogs contained only a subset (the latest five) of users’ posts. accordingly, we downloaded the textual content of all publicly accessible posts of each user as html files, and extracted the user-authored content of each page automatically. we applied a similar annotation process as described above on 6,999 unique posts with 11.5% being identified as stories. the resulting inter-rater agreement was 0.46 adjusting for chance using cohen’s κ. we believe the κ is low for several reasons. below are the three reasons we believe are causing most of the difficulty. first, posts often contain a mix of content, which can include at least one personal story. deciding exactly how much narrative content is sufficient to warrant labeling the post a story can be difficult. second, some posts have narrative like qualities, i.e., sequences of events, but have little or no causal or evaluative structure. these can be difficult for annotators especially for short posts that are under 5 or 6 sentences. third, our definition requires a story to be about a particular event spanning a well specified time frame and is exclusionary of generalizations 111 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber culture # labeled + story κ precision recall f1 corpus american 4,985 270 0.68 0.66 0.48 0.55 25m chinese 6,999 805 0.46 0.60 0.46 0.52 64k iranian 18,490 462 0.49 0.15 0.23 19k table 1: summary of the automated classifier results for each culture. # labeled indicates the total number of posts that were manually annotated. + story indicates the number of posts that were positively labeled as a personal story. corpus is the total number of stories identified by our automated classifier and used as the corpus to draw from in this experiment. over many episodes. however, these cases are not always easy to distinguish nor are they always mutually exclusive. an automated classifier was trained on the labeled data (precision=0.60, recall=0.46, and f1-score=0.52) and applied to the entire sina dataset. where there were discrepancies between automated labels and manually provided labels the manually provided labels were used. a total of 64,231 stories were identified. we also collected iranian stories in a similar fashion to the chinese using 12 popular iranian weblog hosting services. we applied a similar annotation process on 18,490 unique posts, which indicated 2.5% as personal stories. inter-rater agreement was assumed to be similar for the iranian population. these annotations were also used to create an automated story classifier for stories written in farsi (precision=0.49, recall=0.151, and f1score=0.23)1. using this classifier, we identified 132,641 unique blogs and extracted 778,745 posts. the automated classifiers used to create our base corpus of english, mandarin and farsi stories are available online2. a summary of the annotation and performance statistics for each culture is presented in table 1. the differences in narrative content and performance of the classifiers between the three cultures is substantially different and it would be interesting to explore the source of these differences in future work. we suspect that some of the differences in the distribution of stories across languages may not be cultural, but could simply be due to the different types of users each weblog hosting platform attracts. our classifiers rely heavily on simple lexical features, which were shown to work well for english where features such as past tense verbs and personal pronouns were highly indicative of narrative content. however, for example, chinese lacks verb tense, which could make recognizing past tense actions more difficult because multi-word or more complex constructions might be needed to detect them. this might also require more training data to achieve a similar level of performance. in all three languages the precision is higher than the recall indicating that our classifiers are often correct when it classifies a post as a story, however it fails to identify a large number 1. we are aware of the low recall of the farsi classifier, however, we have been unable to perform a thorough error analysis of the classifier at this time and we do not currently wish to speculate why the performance is lower than the others. because of the design of the experiments in part 2 the low performance will not be a concern for this work. 2. https://github.com/asgordon/storynonstory 112 analysis of subjectivity and narrative levels in storytelling across cultures of posts that do actually contain a personal story. while the performance of the classifiers, especially for chinese and farsi, is less than ideal, it is important to note that this is not a problem for this research because all the stories used in both our experiments go through an additional screening by our human annotators. 4. text annotation after collecting a large sample of candidate personal narratives using our classifiers we randomly selected documents from this corpus to create a smaller curated dataset of personal narratives which we annotated with the narrative level and subjective clause labels. we randomly selected documents from each of these corpora and followed the procedure described later in this section to produce segmented narratives annotated along these two dimensions. during the course of our annotation process we defined a detailed annotation scheme. the annotation scheme described here is the product of an iterative refinement involving a computational linguist and two annotators. the annotators, whose backgrounds are in linguistics and psychology, first acquired familiarity with basic concepts in narratology and computational analysis of narrative by reading the background chapter of computational modeling of narrative (mani, 2012). they then annotated a practice set of about 30 narratives, individually, but in frequent consultation. this process resulted in refinements to the annotation scheme and guidelines for dealing with borderline cases, which were recorded in an annotation manual. the final set of annotations were produced by applying our process to personal stories drawn from our corpus following a three-step process: filtering, segmentation, and narrative level and subjectivity annotation. our goal was to annotate 600 english, 300 mandarin, and 300 farsi stories for a total of 1,200. we collected fewer stories for mandarin and farsi because of the difficulty in finding suitable annotators in those native languages. after the collection effort a number of annotations were filtered for missing data and quality reasons, which left us with an odd number of stories and in the case of chinese, fewer than 300. 4.1 filtering because the precision of automatic story identification is not perfect, some posts selected automatically are not in fact stories. the first step in our annotation pipeline was to manually discard posts that were not personal stories. the same definition and guidelines for annotating personal stories from our prior work (gordon and swanson, 2009) was used to filter out non-stories. posts that contained multiple personal stories were also discarded by our annotators. 4.2 segmentation the next step was to break the text into the segments that will be labeled according to narrative level and subjectivity. text segments were determined automatically using a set of heuristics described below and consist of one or more clauses. once the documents were segmented, the annotators labeled each segment for subjectivity and narrative level. the use of clauses as the granularity for annotation was motivated by concerns both princi113 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber pled and practical in nature. perhaps the easiest segmentation strategy for identification of subjective passages in narrative is to consider each sentence as a target for labeling. sentences, however, are clearly too coarse grained, since a single sentence may express an unbounded number of objective and subjective statements through coordination. a more suitable strategy is to define segments in the spirit of the elementary discourse units (edus) in rhetorical structure theory (mann and thompson, 1988), as applied to entire texts by marcu (1999). instead of addressing the challenges of adapting full edu segmentation to the needs of our task, we opted to use a simplified segmentation scheme inspired by edus. we used clauses as the target of annotation with the application of rules and simple heuristics to prevent segmentation of certain types of subordinate clauses that tend not to be relevant to our annotation. practically, this means that discourse units must be continuous and non overlapping spans of text, which is not a requirement for the annotation guidelines used for rhetorical structure theory segmentation and parsing by marcu. for example, the non-finite subordinate clauses in he told her not to go and i like going to the movies are not split into segments separate from their matrix clauses. other examples of subordinated language that results in multi-clause segments include going to the movies is what i like to do on weekends (one segment with four clauses) and he said he would return (one segment with two clauses). our segmentation approach is based on identification of syntactic patterns in parse trees produced automatically by the stanford parser (klein and manning, 2003), and largely follows the edu segmentation approach described by tofiloski et al. (tofiloski et al., 2009), but without the full set of rules and lexical patterns necessary for complete edu segmentation according to the rst guidelines. the source code implementing our rules for segmentation and automatic classification is available online3. 4.3 narrative level and subjectivity annotation the last step was to assign two labels to each segment, one indicating narrative level (extradiegetic vs. diegetic), and the other a binary indicator of subjectivity. although grammatical tense is often indicative of narrative level, since events in the story are usually expressed in past tense, this is certainly not always the case (e.g. in a man walks into a bar, the event described is in the diegetic level, but described using the present tense). the notions of emotion and sentiment are certainly important aspects of narratives that are relevant to our overall goals, but we focused our efforts on the related notion of subjectivity as the expression of private states. this simplifies labeling of cases, such as reported speech and reported emotions. for example, in he said he was sad, we do not treat he was sad as an independent segment, since it is subordinated language. the single segment is labeled as objective, reflecting the saying event, even though it involves reporting of a private state. however, in i knew he was sad, there is a single segment and it is labeled as subjective, not because of the emotion reported, but because knowing is a private state. 3. the automated segmenter based on these heuristics is available at https://github.com/asgordon/ narrativeanalysis 114 analysis of subjectivity and narrative levels in storytelling across cultures 4.4 annotated narrative dataset we annotated these stories along our two dimensions following a similar set of guidelines developed by sagae et al. 2013 and rahimtoroghi et al. 2014, with the definitions of the clause types given below: 1. subjectivity (a) subjective clauses express private states, which include emotions, opinions, evaluations and speculations that are not open to external observation or verification by others. (b) objective clauses express states that can be externally observed and verified by others. 2. diegesis (a) diegetic clauses give information about events as they occurred in the world of the story. (b) extradiegetic clauses give information about the world in which the narrator is addressing the reader. upon completion of the process we had 617 english, 261 mandarin, and 323 farsi stories annotated along these 2 dimensions. most of the stories were labeled by a single annotator who matched the language and culture of the target narrative. a subset of the english language stories were annotated by raters from all three cultures to assess the interrater reliability treating it as a 4 category labeling task. raw pairwise agreement over 571 segments (40 stories, with 12 to 18 segments per story) on this four-way labeling task was 84%. we measured chance-corrected agreement over all annotations using krippendorf’s α and obtained a value of 0.73, which is generally considered acceptable agreement for these types of annotation tasks. because the agreement between annotators across all three cultures was high, our expectation was that agreement on labels on stories in their native languages would also be acceptable. to produce the final annotations, cases where the annotators disagreed were discussed between themselves and a final label was chosen through their discussion. table 2 presents a personal narrative annotated following the guidelines presented above4. all annotations discussed in this work were performed manually by our raters, however a tool was developed to segment and automatically label clauses along these two dimensions for these three languages. this tool is publicly available on github5. overall, these stories are more diegetic than extradiegetic, but the distribution is more equitable than one might expect (58% vs. 42%). they are also highly subjective expressing a private state 68% of the time. additionally, each clause type has an impact on the distribution of the other. for example, a clause is much more likely to be diegetic given the clause is not subjective p(diegetic = yes|subjective = no) = 0.77, whereas the likelihood of a clause to be diegetic given it is subjective p(diegetic = yes|subjective = yes) = 0.49. a summary of the clause distributions across all stories is provided in table 3. 4. because of issues regarding the expectation of privacy from bloggers (hayden, 2013) and the nature of the material in our narratives, we do not use examples taken from our corpus. 5. available at https://github.com/asgordon/narrativeanalysis 115 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber lvl sbj narrative clause e s i consider myself a very honest person, e s and i’ve always thought that truth is the best policy. e o i am a 42-year old mother of two, e s and my kids are the most important thing in my life. d o a few years ago my son tyler asked me if santa claus really existed. d o he was four at the time. e s oh boy! d s i just wasn’t ready for that question. d s it was a nice day outside, d o and i told him to go out and play. d o he came back only 15 minutes later d o and asked me again about santa. d o i pointed to his bike, d o and i asked him who gave it to him. d o he said santa did. d o i nodded d o and said well then, d s and he gave me a huge smile. d s i felt a little guilty at the time about lying to my child, e s but now i know that parenting is a balancing act. e s of course the truth is important, e s but nothing trumps a mother’s instinct. table 2: an example story segmented and annotated using the narrative clause labels. narrative levels (lvl): diegetic, extradiegetic and subjectivity (sbj): subjective, objective. subjective diegetic yes no sum p(d) p(s = y|d) yes 18700 13932 32632 0.58 0.83 no 19250 4074 23324 0.42 0.74 sum 37950 18006 55956 p(s) 0.68 0.32 p(d = y|s) 0.49 0.77 stories = 1,201 table 3: the frequencies of clauses for each of the possible combinations of types and their overall distributions. the column for p(d) indicates the marginal probability of the diegesis variable. p(s) is similarly displayed for subjectivity. the column p(d = y|s) contains the conditional probability that a clause is diegetic (yes) given the value of subjectivity either yes or no corresponding to the value in the appropriate row. the conditional probability p(s = y|d) is also provided. the total number of stories used to create this table is shown in the bottom right. 116 analysis of subjectivity and narrative levels in storytelling across cultures subjective language diegetic yes no sum p(d) p(s = y|d) mandarin yes 2414 2173 4587 0.50†‡ 0.53†‡ no 3400 1204 4604 0.50†‡ 0.74†‡ sum 5814 3377 9191 p(s) 0.63†‡ 0.37†‡ p(d = y|s) 0.42†‡ 0.64†‡ stories = 261 english yes 12144 8764 20908 0.58 0.58 no 12414 2553 14967 0.42 0.83 sum 24558 11317 35875 p(s) 0.68 0.32 p(d = y|s) 0.49 0.77 stories = 617 farsi yes 4142 2995 7137 0.66†‡ 0.58‡ no 3436 317 3753 0.34†‡ 0.92†‡ sum 7578 3312 10890 p(s) 0.70†‡ 0.30†‡ p(d = y|s) 0.55†‡ 0.90†‡ stories = 323 table 4: a comparison of the narrative clause frequencies and distributions across cultures using a format similar to table 3. a † indicates a significant difference (p� 0.01) from english and a ‡ indicates a significant difference (p� 0.01) from farsi. 4.5 cross-cultural comparison results we also performed an analysis of the clause distributions between cultures. we used a binomial test between each pair of cultures to assess significance between the distributions. the distributions were all significantly different (p� 0.01) for all distributions except p(s = y|d) between english and farsi. mandarin stories had the least amount of diegetic material and subjective material, whereas farsi stories were the most diegetic and subjective. there were also large differences in the conditional distributions. the likelihood that a clause is subjective given that it is extradiegetic was greater than 50% for all languages, however it was near certainty (92%) for farsi stories and only 74% for mandarin stories with english stories in between. in other words extra-diegetic material is less likely in farsi stories, but when there is an extra-diegetic clause it is almost certainly expressing an opinion or private mental state. in contrast, compared to the other languages, mandarin stories are more likely to contain extra-diegetic material that is objective and english stories split the difference between the two. a summary of the frequencies and distributions over all stories in each corpus is provided in table 4. 117 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber 5. experiment 2: impact on attitudes and behavior experiment 1 established that each culture structures its stories differently with regard to the distribution of narrative clause types. in this section we investigate whether these structural differences are embedded strongly enough in a cultural group that they could play a role in influencing the attitudes and behaviors of the reader. a narrator can craft a story for many communicative and persuasive purposes (schank, 1991) that might shape the reader’s attitude toward the narrator. one of the persuasive goals often found in personal narratives is to reaffirm the narrator’s identity and to project a desired self image to the audience. when a personal story is about a morally charged event, e.g., having an abortion, cheating on a test, or feeling guilty about stealing, the narrator will also often frame the story in a way that tries to convince the audience they acted appropriately. based on similarity-liking theory (byrne, 1997), which posits that people gravitate toward others they perceive as similar, our hypothesis is that stories whose narrative clause distribution matches the reader’s cultural norm will receive the most positive attitudes and decrease as the distribution diverges. our hypothesis maintains that the structural distribution of a story is not inherent to the genre of narrative, but that members of a cultural group adopt these types of features from the stories of others in their group. this adoption will also lead to a preference for stories that share structural properties that are similar to those of their group. this is similar to the work by (niederhoffer and pennebaker, 2002) who first demonstrated that readers can subconsciously pick up on certain stylistic properties of a discourse, such as word choice, and will alter their behavior by entraining their writing style to match the other party. here, we investigate whether the narrative discourse properties described above can also impact the attitudes and behaviors of the reader of a story. for example a story written in english, but structured similarly to a chinese story may not be as well received because it does not “feel” right to an american audience, i.e., is intuitively less similar. specifically, when narrative properties match cultural expectations we believe a reader will be more likely to: • perceive that the narrator belongs to the same culture. • have a positive attitude toward the narrator. • have a positive attitude toward the story. • find the narrator more trustworthy. • agree with the decisions or actions of the narrator. 5.1 design to test our hypotheses we designed an experiment to assess how the distribution of narrative clauses impacts the attitudes and behavior of readers as those distributions diverge from the cultural norm of the reader. the basic design of the experiment was to have a participant read a story from our annotated corpus in their native language, answer several follow up questions and play a 1-shot version of the trust game (berg et al., 1995). this was setup as a web-based survey across four pages: prequestionnaire, story, survey questions and trust game. although participants read stories drawn from their own native language and culture, the distribution of clauses varies widely within our dataset and is unknown to the reader. this allows us to test if the aggregate responses of the readers in a culture 118 analysis of subjectivity and narrative levels in storytelling across cultures differ based on changes in the underlying narrative clause distribution. the trust game, described below, is designed to more directly assess a person’s attitude toward another than self reported measures because it induces a behavioral response that requires a real valued concession on the part of the participant. there are many confounding factors that could also influence the responses of the readers, for example, the writing complexity, writing quality, word choice, length of the story and content. we deal with these issues by making certain assumptions about the data. primarily, we assume that these factors are independent from the narrative discourse style and therefore over a large sample of stories these factors will not substantially affect the aggregate results. while we recognize this assumption may not completely hold for our sample, we believe alternative experimental designs are at least equally problematic and much more difficult to conduct. for example, in a preliminary study prior this design we also considered an alternative approach based on a controlled experiment. in this preliminary design, a small number of stories were selected and then each was modified by an author to skew the distribution of clauses toward other cultural norms. however, we were immediately faced with several challenges that could not be reconciled. using a small sample of stories posed the first challenge with this approach. if the stories were not at least minimally engaging to the target audience, the responses to any follow on surveys were likely to be extremely noisy and uninformative. an extensive amount of pretesting would have been required to make this work and it was unclear how consensus would be reached to move beyond this phase. it was also challenging to adjust the distribution of an existing story without significantly altering the confounding factors, for example the content, document length and word choice. simply adding clauses to increase the relative frequency of one type resulted in stories of different length. it also often adds additional information not contained in the other versions, which would confound the results. by replacing one clause type with another, e.g., extradiegetic to diegetic, it was also possible to significantly alter the content of the story through omission. there were also difficulties in maintaining the same lexical difficulty, reading level and naturalness of the story. for these reasons we found it impractical to perform a small controlled study that we felt overcame these issues satisfactorily. modeling subjectivity and narrative levels: our primary hypothesis in this study is that as the distribution of narrative clauses diverges from a culture’s norm, the attitudes and behaviors of readers will be less positive toward the narrator. to model this situation we represent our independent variable (diegesis) as the percentage of clauses that are diegetic and the independent variable (subjectivity) as the percentage of clauses that are subjective. since we predict that the attitudes of the reader will become less positive as the diegesis and subjectivity diverge from the cultural norm, we model our independent variables d and s as the distance from the cultural average. for example, we see from table 4 the cultural norm for diegesis is 58%. so if a particular story in english is 30% diegetic then its distance from the cultural average would be -28%. if our hypothesis is correct we would expect to see a non-linear parabolic relationship where our dependent variables reach a maximum at the cultural norm and lower on either side. we will describe how we handle this expected non-linear relationship in the results section. prequestionnaire: the participant was first asked to complete a prequestionnaire that gathered several pieces of information about them. first, we gathered some basic 119 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber demographic information, such as their age, gender, home country, and level of education. to ensure our chinese participants were not overly westernized, we also included a subset of the stephenson acculturation inventory (stephenson, 2000) that was slightly modified for our american participants to keep the semantics consistent across both cultures6. story: after completing the prequestionnaire a participant was shown a story drawn from the collection of their native language. they were given as much time as they wanted, but once they progressed they were unable to revisit the story again. survey questions: on the next page we asked seven survey questions using a 7-point bipolar scale. q1. i enjoyed reading this story. q2. the narrator did the right thing in this situation. q3. given the situation described i would have done the same (or similar) thing as the author. q4. this story was written by someone like me or from my background. q5. given the opportunity i would be friends with the author. q6. i found this story interesting. q7. i have been in a situation exactly like this before. q1 and q6 were asked to gauge the reader’s general interest in the story. we used simple intuitive questions for this assessment because we believed that these direct questions would be nearly as reliable as other measures of interest or engagement used for narrative while keeping the survey short. in at least some of our stories, the reader will perceive that the narrator was faced with a difficult decision. q2 and q3 were intended to assess whether the narrator has convinced the reader that she chose reasonable actions under the circumstances. in instances in which readers provide strong opinions either for or against the narrator, these questions could enable future work based on morally charged narratives. one of our primary research questions was to determine if the discourse structure that varied substantially from the cultural norm of the reader would have a significant impact on the reader’s cultural affinity toward the narrator. we used q5 and especially q4 to measure this response. we also believed a reader’s response to the trust game or one of our other response variables might be moderated by how familiar the reader was with a given situation. for example, a reader intimately familiar with a difficult situation might empathize with the problem the narrator was describing, or alternatively a reader familiar with a boring situation might not pay attention to the details of the story making the judgments less reliable. we asked q7 for this purpose, although it was not used in the experiments described here. trust game: after answering the survey questions we had the participants play a 1-shot version of the trust game (berg et al., 1995), which is designed to measure a participant’s trust of an anonymous (or partially anonymous) partner. the game was designed to address a problem with self-reported measures of attitudes related to trust where participants will often report values that do not correspond to their actual behavior, because there are no actual consequences to their self reported decisions. in the game, attitudes are measured 6. we also asked the participants to take a standard need for cognition scale (cacioppo and petty, 1984), but did not use the results for this work. 120 analysis of subjectivity and narrative levels in storytelling across cultures not by what participants say they would do but how they actually behave by requiring them to give up (or believe they are giving up) part of their monetary compensation. for our version of the game game, we told the participants that we were in contact with the author and they could offer any portion of their participation compensation, from $0.00 to $1.50, to them. we would triple the amount specified and offer this new augmented value to the author. the author would then have a chance to reciprocate and give back any portion of this total value to the participant. the participant (in theory) has an opportunity to make more money than originally promised, however, standard economic models assuming rational behavior, based on the nash equilibrium, predict a participant should give $0.00. the amount given above $0.00 is stipulated to correlate with how much the participant trusts the other player. in our experiment we were not actually in contact with the author and no monetary transfer took place, although the participant was unaware of this until after the entire survey was complete. experiment: questions 1-5 of our survey and the monetary value specified in the trust game are our dependent variables that address our research questions and hypotheses. question 1-3 are intended to provide assessments of the reader’s attitude toward the message of the story, whereas question 4-5 address the reader’s attitude toward the narrator. we also believe that a reader’s level of engagement could moderate effects. we see this happening in two possible ways. first, if a reader is highly interested with the story then a discerning reader may be more attuned to the discourse properties of a text. alternatively, we could imagine a disengaged reader may be less influenced by the content of the story, which they are not engaged with, but discourse properties may still be subconsciously impacting their attitudes. we address this concern with a simple self reported assessment of engagement in question 6. the monetary value specified in the trust game is our final dependent variable that not only assess the reader’s attitude toward the narrator, but how willing they are to alter their behavior based on that attitude. 5.2 participant recruitment we recruited participants from two of our cultures (american’s and chinese) and hosted the survey using the crowdsourcing service crowdflower7. due to the difficulty of recruiting iranian nationals living in the united states online they were excluded from this portion of the study. participants were paid $1.50 for participating in the experiment, which took about 10-15 minutes to complete. for the chinese participants, we targeted native chinese nationals who are currently in the united states, but who are just visiting or have lived here less than 5 years. we randomly sampled 500 stories from our english corpus and 250 from our mandarin corpus. our target was to have each story read by 2 participants with each participant reading a single story. because of the difficulty in recruiting participants, especially chinese, our final design only included one rating per story as described in the next section. we were able to recruit a total of 742 american participants and 374 mandarin speaking participants. to help remove unreliable participants, we filtered participants using a number of criteria. 7. www.crowdflower.com 121 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber we first removed anyone who did not answer all of the required questions in the main part of our survey. this removed 76 english and 13 mandarin participants. crowdflower is typically used as a crowdsourcing platform for performing microtasks with objective verifiable outcomes. in a prototypical crowdflower task (unlike ours) a number of hidden “test” questions are included on the page, which the answer is known to the author but not the worker. these test questions are then used to create a quality profile of workers over all the tasks they perform. this score ranges from 0 (has answered no test questions correctly) to 1 (has answered all test questions correctly). although we did not have any of these test questions on our survey, we could still rely on the quality score of the participant based on the other tasks they have performed on crowdflower. to try to balance our concerns about obtaining highly reliable participants, while not filtering too many qualified subjects, we selected a threshold quality score of 0.8. applying this filter removed 322 english and 178 mandarin participants. the final criteria we used was the length of time the participant had been in the united states as determined by our prequestionnaire. the american participants were required to have been in the us for longer than 5 years, while the mandarin participants were required to have been in the us less than 5 years. this removed 22 american participants and 96 mandarin participants. after filtering we had 311 american stories rated by 322 participants (11 stories annotated by 2 raters) and 85 chinese stories rated by 89 participants (4 stories annotated by 2 raters). for the stories annotated by more than one participant we randomly selected an annotation from one of the raters, so that our final dataset consisted of 311 american stories and 85 chinese stories. 5.3 results we used the results of our questionnaire to build several linear models to test our hypotheses that stories whose clause distributions more closely matches the cultural norm of the participant will engender more trust toward the narrator. for these models we treated q6 and the narrative clauses distance from the norm as our independent variables. our dependent variables, each modeled with a separate linear model, were the results of the trust game and q3-q5. the results of other questions were not used in these experiments. our primary model described here is based on using the results of the trust game as our dependent variable. our model consists of 4 main effects: culture, diegetic distance from norm, subjective distance from norm, and q6. culture is a categorial variable with two values. d and s are continuous prediction variables representing the distance from the cultural norm as described above. q6 is an ordinal variable with 7 values ranging from strongly disagree (-3) to strongly agree (+3). our response variable is the continuous amount of money offered in the trust game. to account for the non-linear relationship expected between the narrative clause distances, we apply a second order polynomial transformation on the subjective and diegetic distance variables, which could both approximate a linear relationship or the type of nonlinear relationship we expect (i.e., a peak value of trust at the cultural norm and lower values for both high and lower distributions). we also include all two way interactions between the variables and the three-way interaction between culture×d×s. our model 122 analysis of subjectivity and narrative levels in storytelling across cultures variable df ss f p-value (intercept) 1 17.811 67.524 4.35e-15 culture 2 0.035 0.132 0.7169 d2 2 0.301 0.570 0.5660 s2 2 1.121 2.125 0.1210 q6 6 1.748 1.105 0.3592 culture:d2 2 0.639 1.211 0.2993 culture:s2 2 3.263 6.185 0.0023 d2 :s2 4 2.202 2.087 0.0821 d2 :q6 12 3.897 1.231 0.2598 s2 :q6 12 6.535 2.065 0.0187 culture:q6 6 2.290 1.447 0.1958 culture:d2 :s2 4 1.455 1.379 0.2407 r2 = 0.204 r2adj = 0.081 p = 0.0041 table 5: the anova results of our model with the degrees of freedom (df), sum of squares (ss), f-statistic (f) and p-value for each factor in the model. significant effects are highlighted in bold. is expressed in a slightly modified version of wilkinson notation (wilkinson and rogers, 1973) in equation (1) below. trust ∼ culture ∗d2 ∗ s2 + q6 + culture:q6 + d2:q6 + s2:q6 (1) a:b indicates the interaction between all factors/levels/categories of the variables a and b. a ∗b is a shorthand for the linear combination of a and b along with the interactions between them (i.e., a+b + a:b). in equation (1) we use the notation d2 and s2 to mean that a second order polynomial transformation has been applied to these terms, which is not standard, but helps the readability of the model and the terms in table 5. the polynomial is added in order to have sufficient power to model the expected non-linear parabolic relationship between our independent and response variables. we used a type iii sum of squares analysis of variance with divergence contrasts for the categorical variables and polynomial contrasts for the ordinal variable. table 5 summarizes the results of this analysis with the degrees of freedom, sum of squared errors and f statistic used to calculate the p-value for each term in the model. the table shows that none of the main effects of this model are significant (p-value < 0.05), but the interaction between culture and subjective distance (culture:s2) and the subjective distance and the reader’s interest (s2:q6) in the story were. this indicates that the distribution of subjective material has an effect on the level of trust in the narrator based on the culture and that effect of subjective material is moderated by the interest level of the reader. the model as a whole explains a relatively small overall amount of the variance (adjusted r2 = 0.081), but it is significant and non-trivial considering the model does not rely on the content of the stories at all. 123 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber we also applied models with the same independent variables to our other dependent variables q3-q5. in all cases we did not find any significant main effects or interactions that involved the narrative discourse properties, although all of the models were able to explain more of the variance than for trust (q3: adjusted r2 > 0.13, p < 0.0001; q4: adjusted r2 > 0.14, p << 0.0001; q5: adjusted r2 > 0.23, p << 0.0001). 6. discussion and conclusion stories are a rich method of discourse that both communicate informational content, but also subjective evaluative remarks that provide additional meaning to the audience. narrators implicitly and explicitly intertwine these types of discourse to compose stories that resonate with their intended audience. although these clause types could be combined arbitrarily, we see clear preferences for particular distributions and that these distributions vary across the three cultures we investigated. not only are the distributions different across cultures, but we also found that the clause distribution has some impact on a reader’s attitude and behavior toward the narrator. our results suggest that it is important to consider the discourse structure when authoring or translating a story for an audience from a specific culture/language. the attitudes expressed towards the narrator by american’s primarily fit with what we would expect from similarityliking theory and previous types of linguistic style matching. for this audience, a reader’s attitudes are most positive when the stories narrative clause distribution is near the cultural norm. however, the chinese results were not what we would expect. although the results were not significant, the trend for diegetic clauses is the opposite of what we would expect and the relationship for subjective clauses is linear, which was not predicted. we currently do not have a good explanation or hypothesized mechanism that could account for these distributions and a better understanding will require deep knowledge of chinese cultural factors related to expressions of opinion and trust. we believe there are a number of additional experiments that could provide additional insight. first, we would be interested in looking at additional culture’s, such as iranian, which were excluded due to difficulties in recruiting. in addition to providing more data, which would help verify the impact of narrative structure between cultures, but also help us determine if chinese is an exception to our similarity-liking hypothesis or another mechanism is more likely. while our web survey allowed us to increase the number of participants, it had two significant drawbacks. one, it introduced noise in into the responses from unreliable participants, and two it limited the types of behavioral interactions that we could perform (i.e., restricted to a 1-shot version of the trust game). moving to a more controlled environment would limit our ability to recruit participants, but would reduce the noise and allow us to study more complex behavioral responses. while, we believe it is infeasible to control for all the possible textual variables (e.g., length, content, difficulty, etc.), there are some steps that could be taken to control for some of them. for example, we could control for the content by selecting stories that are all based on a similar event, such as attending a wedding or going to a funeral. by controlling for the content we could potentially be more discriminative in isolating the narrative discourse structure. comparing stories about different topics might also shed light on what kinds of stories are most likely to be affected by discourse structure, for example stories about traumatic events versus mundane activities. 124 analysis of subjectivity and narrative levels in storytelling across cultures 7. acknowledgments the projects or efforts depicted were or are sponsored by the u. s. army. the content or information presented does not necessarily reflect the position or the policy of the government, and no official endorsement should be inferred. we would also like to thank heejung kim for her comments and suggestions. references apoorv agarwal, boyi xie, ilia vovsha, owen rambow, and rebecca passonneau. sentiment analysis of twitter data. in proceedings of the workshop on languages in social media, pages 30–38, stroudsburg, pa, 2011. association for computational linguistics. joyce berg, john dickhaut, and kevin mccabe. trust, reciprocity, and social history. games and economic behavior, 10(1):122–142, july 1995. issn 0899-8256. doi: 10. 1006/game.1995.1027. url http://www.sciencedirect.com/science/article/pii/ s0899825685710275. john blitzer, mark dredze, and fernando pereira. biographies, bollywood, boom boxes and blenders: domain adaptation for sentiment classification. in proceedings of the 45th annual meeting of the association of computational linguistics, 2007. jerome bruner. the narrative construction of reality. critical inquiry, 18(1):1–21, 1991. the university of chicago press. donn byrne. an overview (and underview) of research and theory within the attraction paradigm. journal of social and personal relationships, 14(3):417–431, june 1997. issn 0265-4075, 1460-3608. doi: 10.1177/0265407597143008. url http://spr.sagepub.com/ content/14/3/417. john t. cacioppo and richard e. petty. the need for cognition: relationships to attitudinal processess. in j. h. harvey, r. p. mcglynn, and j. e. maddux, editors, social perception in clinical and counseling psychology. texas tech university press, lubbock, tex., january 1984. isbn 978-0-89672-127-2. richard chalfen. snapshot version of life. bowling green state university popular press, bowling green, oh, 1987. stephanie j. coopman and katherine b. meidlinger. interpersonal stories told by a catholic parish staff. american communication journal, 1(3), 1988. volker eisenlauer. narrative revisited: telling a story in the new age of media, chapter once upon a blog. . . storytelling in weblogs, pages 79–108. john benjamins publishing company, 2010. gerard genette. narrative discourse: an essay in method. cornell university press, 1980. richard j. gerrig. experiencing narrative worlds: on the psychological activities of reading. yale university press, new haven, june 1993. isbn 978-0-300-05434-7. 125 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber andrew s. gordon. the fictionalization of lessons learned. ieee multimedia, 12(4):12–14, 2005. andrew s. gordon and reid swanson. identifying personal stories in millions of weblog entries. in third international conference on weblogs and social media, data challenge workshop, san jose, ca, may 2009. andrew s. gordon, luwen huangfu, kenji sagae, wenji mao, and wen chen. identifying personal narratives in chinese weblog posts. in intelligent narrative technologies (int6). papers from the 2013 aiide workshop, page 23, 2013. url http://www.aaai.org/ ocs/index.php/aiide/aiide13/paper/view/7419/7647. erika check hayden. guidance issued for us internet research. nature news, 496:411, 2013. dan klein and christopher d. manning. fast exact inference with a factored model for natural language parsing. in in advances in neural information processing systems 15 (nips), pages 3–10, cambridge, ma, 2003. mit press. william labov and joshua waletzky. narrative analysis: oral versions of personal experience. journal of narrative and life history, 7(1-4):3–38, 1967. kristin langellier and eric e. peterson. storytelling in daily life. temple university press, philadelphia, pa, 2004. angela k.-y leung and dov cohen. the soft embodiment of culture: camera angles and motion through time and space. psychological science, 18(9):824–830, september 2007. issn 0956-7976. doi: 10.1111/j.1467-9280.2007.01986.x. inderjeet mani. computational modeling of narrative, volume 5 of synthesis lectures on human language technologies. morgan & claypool publishers, 2012. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text, 8:243–281, 1988. daniel marcu. a decision-based approach to rhetorical parsing. in proceedings of the 37th annual meeting of the association for computational linguistics on computational linguistics, acl ’99, pages 365–372, stroudsburg, pa, usa, 1999. association for computational linguistics. isbn 1-55860-609-3. doi: 10.3115/1034678.1034736. prem melville, wojciech gryc, and richard d lawrence. sentiment analysis of blogs by combining lexical knowledge with text classification. in proceedings of the 15th acm sigkdd international conference on knowledge discovery and data mining, new york, ny, 2009. association for computing machinery. louis-philippe morency, rada mihalcea, and payal doshi. towards multimodal sentiment analysis: harvesting opinions from the web. in proceedings of the 13th international conference on multimodal interfaces, pages 169–176, new york, ny, 2011. association for computing machinery. 126 analysis of subjectivity and narrative levels in storytelling across cultures sean a. munson and paul resnick. the prevalence of political discourse in non-political blogs. in proceedings of the fifth international aaai conference on weblogs and social media, barcelona, spain, 2011. kate g. niederhoffer and james w. pennebaker. linguistic style matching in social interaction. journal of language and social psychology, 21(4):337–360, december 2002. issn 0261-927x, 1552-6526. doi: 10.1177/026192702237953. url http://jls.sagepub.com/ content/21/4/337. elinor ochs and lisa capps. living narrative: creating lives in everyday storytelling. harvard university press, cambridge, ma, 2001. bo pang and lillian lee. a sentimental education: sentiment analysis using subjectivity summarization based on minimum cuts. in proceedings of the 42nd annual meeting on association for computational linguistics, stroudsburg, pa, 2005. association for computational linguistics. bo pang and lillian lee. opinion mining and sentiment analysis. foundations and trends in information retrieval, 2(1-2):1–135, 2008. randolph quirk, sidney greenbaum, geoffrey leech, and jan svartvik. a comprehensive grammar of the english language. longman, new york, 1985. elahe rahimtoroghi, thomas corcoran, reid swanson, marilyn a. walker, kenji sagae, and andrew s. gordon. minimal narrative annotation schemes and their applications. in 7th workshop on intelligent narrative technologies, milwaukee, wi, june 2014. hans reichenbach. elements of symbolic logic. macmillan & co., new york, 1947. kenji sagae, andrew s. gordon, morteza dehghani, mike metke, jackie s. kim, sarah i. gimbel, christine tipper, jonas kaplan, and mary helen immordino-yang. a data-driven approach for classification of subjectivity in personal narratives. in mark a. finlayson, bernhard fisseni, benedikt löwe, and jan christoph meister, editors, 2013 workshop on computational models of narrative, volume 32 of openaccess series in informatics (oasics), pages 198–213, dagstuhl, germany, 2013. schloss dagstuhl–leibniz-zentrum fuer informatik. isbn 978-3-939897-57-6. doi: http: //dx.doi.org/10.4230/oasics.cmn.2013.198. url http://drops.dagstuhl.de/opus/ volltexte/2013/4145. roger c. schank. tell me a story: a new look at real and artificial memory. atheneum, january 1991. isbn 0-684-19049-4. benjamin snyder and regina barzilay. multiple aspect ranking using the good grief algorithm in human language technologies. in proceedings of the conference of the north american chapter of the association for computational linguistics, pages 300–307, 2007. rebecca steinitz. writing diaries, reading diaries: the mechanics of memory. communication review, 2:43–58, 1997. 127 swanson, gordon, khooshabeh, sagae, huskey, mangus, amir and weber m. stephenson. development and validation of the stephenson multigroup acculturation scale (smas). psychological assessment, 12(1):77–88, march 2000. issn 1040-3590. milan tofiloski, julian brooke, and maite taboada. a syntactic and lexical-based discourse segmenter. in proceedings of the acl-ijcnlp 2009 conference short papers, aclshort ’09, pages 77–80, stroudsburg, pa, usa, 2009. association for computational linguistics. janyce m. wiebe, theresa wilson, rebecca bruce, matthew bell, and melanie martin. learning subjective language. computational linguistics, 30(3):277–308, 2004. g. n. wilkinson and c. e. rogers. symbolic description of factorial models for analysis of variance. journal of the royal statistical society. series c (applied statistics), 22(3): 392–399, 1973. issn 0035-9254. doi: 10.2307/2346786. url http://www.jstor.org/ stable/2346786. 128 11384-article text-73357-1-6-20210118 dialogue & discourse 12(1) (2021) 1-20 doi: 10.5210/dad.2021.101 ©2021 alan garnham this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). opinion piece: how people structure representations of discourse alan garnham a.garnham@sussex.ac.uk school of psychology university of sussex brighton bn1 9qh, uk editor: patrick healey submitted 10/2018; accepted 01/2021; published online 03/2021 abstract mental models or situation models include representations of people, but much of the literature about such models focuses on the representation of eventualities (events, states, and processes) or (small-scale) situations. in the well-known event-indexing model of zwaan, langston, and graesser (1995), for example, protagonists are just one of five dimensions on which situation models are indexed. they are not given any additional special status. consideration of longer narratives, and the ways in which readers or listeners relate to them, suggest that people have a more central status in the way we think about texts, and hence in discourse representations, indeed, such considerations suggest that discourse representations are organised around (the representations of) central characters. this paper develops the idea of the centrality of main characters in representations of longer texts, by considering the way in which information is presented in novels, with l’éducation sentimentale by gustav flaubert as a case study. conclusions are also drawn about the role of representations of people in the representation of other types of text. other approaches to discourse and dialogue, and behaviour more generally, are considered in relation to the question of whether they account adequately for the central role of protagonists. keywords: mental model, situation model, discourse/text representation, protagonist 1. introduction the notion of a mental model of discourse (johnson-laird & garnham, 1980), which developed from earlier ideas outlined by johnsonlaird (1970), has proved a crucial one in the psychology of language. johnson-laird and garnham stressed the importance of a representation of not only part of a real or imaginary world being talked about, but also of the knowledge of the other participants in the discourse. this latter aspect of discourse models has received relatively little attention, at least in the psycholinguistic literature. the former aspect – the representation of situations in the real world and in imaginary words – has become a staple of psycholinguistic research, under the heads of mental models (e.g., johnson-laird, 1983; garnham, 1987, 2001) or situation models (e.g. van dijk & kintsch, 1983; zwaan & radvansky, 1998). indeed, it has been argued (garnham, 1996) that, in a general sense, the notion of a mental (or situation) model as the representation of the content of discourse follows from a marrian task analysis of discourse comprehension, which shows that its function is, at least in part, to convey or receive information about parts of the real world, imaginary worlds, or abstract domains. garnham (1996) further garnham 2 contrasts the domain of language with that of reasoning, where, in considering the task to be performed, it would appear that either model-based or rule-based processes could underlie reasoning from what is known, or assumed, to what in some sense (depending on the type of reasoning under consideration) follows from it. it is sometimes suggested, though more often informally than formally, that the notion of a mental model or situation model is an unclear one. i do not believe that this claim can still be justified. formal semanticists are broadly agreed on the semantic types needed to analyse the meanings of natural language utterances: entities, truth-values, eventualities, relations, and properties of various kinds (properties of objects, properties of eventualities, etc.). there have been disputes about which of these are basic, and there are different notions of what is meant by basic. in the parsimonious semantic landscape of richard montague (thomason, 1974) only entities and truth values are basic semantic types, whereas there are arguments for admitting other types, such as events, as basic (e.g. davidson, 1967; parsons, 1990). furthermore, psychological approaches to the question of what is basic may not produce the same answers as philosophical analyses. nevertheless, the broad principles of how to determine semantic types, and of how representations of situations (eventualities) might be constructed from the core semantics (including semantic types) of words, and the syntax of phrases and clauses is reasonably clear. mental models are more than just representations of the content of individual clauses. a key issue in mental models theory, traceable to its precursor in the work of bransford (e.g., bransford & franks, 1971), is the integration of information from different parts of a text (see, garnham, 2020 for a recent discussion). within both psycholinguistics and formal semantics there has been considerable progress in understanding how local relations between clauses or utterances are computed. these relations fall into two broad categories. first, anaphoric relations determine identities of either sense or reference between expressions of various kinds, primarily nps and vps. second, coherence relations, which may be signalled by clausal connectives (such as, “because”, “but”, “although”), or sometimes suggested by world knowledge, determine relations among the eventualities denoted by clauses. whether there is a specific set of coherence relations, or whether such relations are examples of more general types of relation (spatial, temporal, logical, causal, intentional, and moral, miller & johnson-laird, 1976) is a matter of debate (garnham, 1991). temporal relations, which are sometimes determined by connectives and sometimes by sequence of tenses (including aspectual information), share properties of both anaphoric and coherence relations. thus it has become clear since the 1960s that the primary purpose of text and discourse comprehension is not to produce a representation of the semantic structure of sentences in a text or utterances in a dialogue, but to extract information about situations in a world that the semantic structure encodes. so, in interpreting a text such as (1) two eventualities are recognized. (1) max confessed to bill because he wanted a reduced sentence. the first is an act of confession, with a male person max as the confessor and another male person bill as the person confessed to. the second is a state of desire (wanting). the person wanting is a male and the object of his desire is a reduced sentence. the state of wanting is presented as the cause of the confession. knowledge of the world suggests that sentence (1) is about someone charged with a crime and either about to plead guilty or who thinks he will be convicted for some other reason. this knowledge suggests that the male person wanting the reduced sentence is max. it also helps to disambiguate the word “sentence”. in this example the implicit causal bias (garvey & caramazza, 1974) of the verb “confess to” also supports the idea that it is max and not bill who wants the reduced sentence. “confess to” is a so-called np1 verb, which suggests that the subject of a simple active sentence containing the verb will be the primary cause of the eventuality it describes. how people structure representations of discourse 3 the above account appears to be basically correct, as far as it goes, though many details are lacking. the exact mechanism by which implicit causality has its effect and its relation to world knowledge remains to be clarified, for example (see, e.g., crinean & garnham, 2006; garnham, child, & hutton, 2020). another major issue, which i will return to indirectly later, is the extent to which representations are complete (oakhill, garnham, & vonk, 1989; ferreira, ferraro, & bailey, 2002; sanford & sturt, 2002). however, the basic idea that readers or listeners determine, from processing a text or discourse, the eventualities being presented and the relations between them, seems a good starting point for a theory of discourse comprehension. however, even if it is a good starting point, it does not seem to be a good finishing point, or even a satisfactory stopping off place. the reason for making this claim comes from an extension of the kind of marrian task analysis (marr, 1982) proposed by garnham (1996). at the local level, where information in single clauses is processed, and the relations among pieces of information in nearby clauses are computed, the account appears sound. however, if we ask what people are doing, or trying to do, when they process the kinds of text that they encounter in everyday life or enter into everyday conversations, it is immediately clear that they have a wide variety of purposes, which can be difficult to specify, but which clearly go beyond the who-did-what-to whom analysis of individual clauses and local relations between them. just as marr found in the case of vision, it is easier to specify the functions of lower levels of analysis than those of higher levels. for example, the function of word recognition, a lower level process, is to map a visual or auditory pattern, separated from the rest of the current visual or auditory input perhaps in part by the process of word recognition itself onto one of the words that one knows in the language currently in use. however, by beginning to think about some of the higher-level processes in language comprehension, it is possible to draw some conclusions about lower or intermediate levels. in particular, where readers or listeners may have strategic control over intermediate mechanisms, consideration of higher levels may shed light on how that strategic control is exercised. to be more specific, it is difficult not to recognize a clearly spoken or clearly printed word in one’s native language, unless one is deaf or blind, or one blocks one’s ears or closes one’s eyes or averts one’s gaze. in some sense, which we probably do not have an adequate handle on, the processes of word identification are automatic. processes of model building are subject to more subtle strategic control. for example, one can skim a text for an answer to a specific question and largely ignore parts of the text that are deemed irrelevant on the basis of rather superficial processing. furthermore, textual cues, such as (psychological) focusing, can be used by speakers and writers to guide listeners’ and readers’ attention to the most important parts of a text. of course, these general points may simply suggest that some parts of a text are processed more thoroughly than others, with the result that some parts of an overall situation model are better encoded and hence probably better remembered. other considerations, however, suggest that something more systematic is happening. 2. the central role of protagonists1 in this section i will present an impressionistic account of the central role of protagonists, particularly in novels, and i will then consider whether accounts of dialogue and discourse processing, particularly where they consider structure above that of clauses and local relations 1 the term “protagonist” is used in a number of related ways in this paper. an individual protagonist is clearly an entity. to say that protagonist is a dimension of situations, as is done in situation theory, is not to talk of individual protagonists, but to say that relations of some kind between individual protagonists in different eventualities can be used to connect those eventualities. situation model theory does not specify what relations allow eventualities to be linked via indices. however, for the protagonist index the most common relation is clearly that the same person in different eventualities can be used to link those eventualities. garnham 4 between them, properly explain this idea. i will then point out that other disciplines at least acknowledge this central role, even if they do not explain it. the crucial issue is not to establish that people are important in how we think about the world – of course they are, and it is not surprising that this idea is implicit in, for example, theories in social psychology. the question is why theories of dialogue and discourse do not properly incorporate this fact. there are potential dangers in focusing on particular types of text when developing a general theory of comprehension, as experience with story grammars showed (garnham, 1983). nevertheless, general considerations, from both common sense and from parts of psychology outside of psycholinguistics, suggest that what is important to most people most of the time is other people and not situations, or at least not situations for their own sake, but only because of people’s involvement in them. this general point suggests that representations of many types of texts, but perhaps most importantly texts such as narratives and newscasts, will be centred on representations of people. indeed, many such texts are naturally regarded as being “about” certain people. of course, in the case of a long and complex text, such as a novel, one cannot talk of the representation of the text, for a variety of reasons, most of which are obvious. however, in thinking about the representation of the content of an extended narrative that a reader maintains from one reading session to the next, and may use in thinking, in the meanwhile, about the narrative, it is natural to consider person-centred representations. by this claim, i do not simply mean that characters from texts are represented in situation models. all situation model theorists recognize this fact. for example, zwaan, langston, and graesser (1995) explicitly list protagonist as one of the five dimensions on which situation models are indexed – along with time, space, causality, and intentionality, and indices allow the creation of links within models. space has been particularly important in the history of situation model theory, largely because it was in spatial domains that it was easiest to show that situation models had different structures from the sentences describing them (potts, 1974, and many subsequent papers). however, purely spatial relations appear to be difficult to encode and may sometimes not be encoded at all (see below). time has fewer dimensions, and is linguistically more complex, but again there are profound differences between the linguistic and model-based representations of time and temporal relations (e.g., kamp, 1979; kamp & rohrer, 1983). causality, again, has been widely researched, though it is not always clearly differentiated from intentionality. intentionality is closely tied to people and their reasons for action or inaction. my general claim is that much of our thinking about the world is people-centred rather than situation-centred, and that this aspect of our thinking naturally carries over to the descriptions of the real or imaginary worlds we produce or describe in fiction, newscasts, everyday conversations and a variety of other forms of text and discourse. so, protagonists in narratives do not just provide one of five indices of situations. they are a central organizing feature of (persisting) mental representations of the content of texts. furthermore, existing theories of discourse representation do not do justice to this fact. my initial evidence for the centrality of protagonists will not be primarily experimental, though i will return later to other types of evidence that lead to, or at least are consistent with, the same conclusion. my initial evidence derives from a task analysis of narrative comprehension and from a consideration of the nature of texts, how people approach them, and what they hope or intend to get out of them (to use a deliberately neutral phrase – “learn”, for example, is too specific). 2.1 novels it is no coincidence that many novels are named after their principal protagonists, though there are numerous other ways of naming novels, for example via allusions to other literary works (wikipedia, 2020). indeed, there are precedents for naming after protagonists in earlier literature. homer’s odyssey almost falls into this category, though interestingly the iliad’s title derives, in a similar way, from a place name (ilium or troy). virgil’s aeneid has a title, which like that of the how people structure representations of discourse 5 odyssey, derives directly from that of the principal protagonist, aeneas. other early works with names derived from their principal characters include beowulf and the epic of gilgamesh, and the catalan works, blanquerna (llull, 1283) and tirant lo blanc (martorell, 1490). in more modern times, where the term “novel” is more clearly appropriate we have, among others, cervantes’ don quixote (1605), madame de la fayette’s la princesse de clèves (1678), defoe’s robinson crusoe (1719), prévost’s manon lescaut (1732), richardson’s pamela (1740), cleland’s fanny hill (1749), fielding’s tom jones (1749), voltaire’s candide (1759), goethe’s die leiden des jungen werthers (1774), and, somewhat tongue in cheek, given the late appearance of the eponymous “hero”, sterne’s tristram shandy (1759-1767). in the 19th century, many dickens novels have eponymous heroes, as do works by jane austen (emma), george eliot (adam bede, silas marner, felix holt, daniel deronda) and trollope (the warden, dr thorne, phineas finn), along with french works by balzac, flaubert, zola, and others, russian works, including pushkin’s verse novel eugene onegin (1833), oblomov (1858), anna karenina (1877). occasionally two protagonists give their name to a work, as in sir gawain and the green knight, rabelais’ gargantua and pantagruel (1532), flaubert’s bouvard et pécuchet (1881) or, more recently julian barnes’ arthur and george (2005). many of the names above are shortened versions of the full titles of the books, but they are the names by which those books are commonly known. for example, the full title of robinson crusoe is the life and strange surprising adventures of robinson crusoe of york, mariner: who lived eight and twenty years, all alone in an uninhabited island on the coast of america, near the mouth of the great river of oroonoque; having been cast on shore by shipwreck, wherein all the men perished but himself. with an account how he was at last as strangely deliver'd by pirates. written by himself. most of the books mentioned above, and many others, some named after principal characters, some not, are explicitly structured around the exploits of one individual. the rest of this section examines the implications of this idea, together with notions of what readers expect to take away (again, “learn” is both too strong and too weak a term) from their engagement with narrative texts, focusing on one novel in particular. it concludes by considering the implications of these ideas for the comprehension of texts of other kinds. 2.2 flaubert’s l’éducation sentimentale i will focus on one text, gustav flaubert’s l’éducation sentimentale (1869) best known in english through its translation, sentimental education, by robert baldick (1964), for penguin classics (reedited by geoffrey wall, 2004). though not named after its main character – and i will return to the title later – the action of the novel clearly centres on the exploits of frédéric moreau, if “action” and “exploits” are the correct terms for such an akratic character as frédéric. frédéric is returning from paris to his home in nogent-sur-seine when he encounters, and falls in love with, an older married woman, marie arnoux. this love, or perhaps infatuation is a better word, remains with him through the rest of the book, though its intensity fluctuates. frédéric returns to paris and cements his relations with mme arnoux’s husband, a man with a rather dubious business sense, in order to maintain contact with her. his various schemes for occupying himself mainly come to nothing, particularly after his receives an inheritance on the death of his uncle. the other women in his life are the courtesan rosanette bron (“the marshal”), louise roque, the daughter of a landowner in nogent, and madame dambreuse, the wife of a parisian banker. the novel plays out mainly in paris before and after the 1848 uprising. frédéric moves among a group of less clearly defined characters, friends and acquaintances with a variety of professions and political views. these characters are not entirely at ease with, and do not fully understand, the social and political changes they are living through. flaubert was deeply committed to making the historical, social and physical background to the novel as accurate as possible and at times the detail is almost overwhelming – a point i will return to later in relation to the model building of a reader garnham 6 of the novel. in a much-repeated quote, flaubert wrote in an 1864 letter to mademoiselle leroyer de chantepie: "i want to write the moral history of the men of my generation-or, more accurately, the history of their feelings. it's a book about love, about passion; but passion such as can exist nowadays--that is to say, inactive." flaubert clearly disapproves of much of what he sees in the society around him, but he eschews didacticism in favour of presenting the detailed and realistic account referred to above. at the end of the novel, in an ending that has generated much controversy, frédéric and his friend deslauriers, while reminiscing, decide that the best time of their lives was when they were first visiting a brothel in nogent. this ending raises, but does not answer, the question of whether frédéric has gained any insight from the education his sentiments have received since that time, or whether it is the reader who is intended to benefit from a sentimental education. in the course of reading sentimental education and in thinking about the book both during and after reading it, it is clear that a major strand of the representation centres on the fact that the book is, in some sense, about frédéric and presents his story. he is the prime example, in the book, of a man of flaubert’s generation, and although, as i have already said, there is no trace of didacticism in flaubert’s approach, the reader is clearly led to consider frédéric as representing a certain class of men in mid-19th century paris, and is presented with a very detailed view of one such man. in considering sentimental education, one does not think primarily of situations, and of frédéric being in them, but rather of frédéric and the situations he is in. or, rather, of frédéric and his relations to the other characters and to the events happening around him. none of the other people in the novel is characterized at the same level of detail as frédéric, and these other characters are mainly defined, in the reader’s mind, by their relationships to him. one does not think of them primarily according to the situations in which they find themselves, but in terms of their relations to frédéric. another important theme is how frédéric and (some of) the other characters change or do not change in response to the social and political upheaval around them. but again, it is natural to think of these changes in a character-focused way, and not by taking situations as the primary focus of change, with their effects on the characters as secondary. nevertheless, it was undoubtedly flaubert’s intention to present a detailed picture of paris in the mid nineteenth century. it is interesting to consider the comprehension of a text such as sentimental education from the perspective of zwaan et al.’s (1995) event-indexing model. as zwaan himself has shown (zwaan & van oostendorp, 1993), readers do not form strong representations of spatial information when reading naturalistic texts unless they are given special instructions. this finding fits with a least one reader’s (the present author’s) intuitions about what information is extracted from sentimental education and mentally represented. it also fits with the (incidental) finding of morrow’s research on spatial mental models (e.g., morrow, greenspan, & bower, 1987; morrow, bower, & greenspan, 1989) that learning spatial layouts is a difficult, time consuming process. there is a great deal of spatial detail in sentimental education, details of the layout of apartments and houses occupied by rosanette bron, the dambreuses, and the arnouxes, among others. and there are details of where in paris certain events happen. however, the overwhelming impression is that the details add to the texture of the narrative, rather than being encoded as spatial relations. it is almost never crucial to the action of the novel that exact spatial relationships are computed. in the case of locations within the city, if one knows these locations, one can map the events onto one’s mental map of paris, but otherwise it is not at all clear that they are encoded. as far as time is concerned, the narrative of sentimental education for the most part moves forward, unlike in some other novels, where the author is what zwaan (1996) calls deliberately inconsiderate and uses techniques such as flashback. there are gaps in the narrative, particularly the famous one at the beginning of part iii chapter v1, which begins “he travelled”. after frédéric quarrels with rosanette, loses contact with madame arnoux, rejects madame dambreuse, finds how people structure representations of discourse 7 that sophie has married deslauriers, and sees one of his old set of friends and acquaintances kill another, he leaves paris in peripatetic mode. in another place, flaubert moves time forwards and, surprisingly given his meticulous attention to detail, extends rosanette’s pregnancy to 25 months. the fact that most readers fail to notice this error, suggests that temporal relations are not always encoded in detail, just as spatial ones are not. however, zwaan, radvansky, hilliard, and curiel (1998) found that temporal relations are better encoded than spatial ones, in that temporal discontinuities are detected more easily than spatial ones. the causal and intentional indices of the event-indexing model are closely related to issues about the representation of protagonists. as mentioned earlier, in the psychological literature the distinction between causes and intentions is not always clearly made. in any case, in narratives, psychological causation (frédéric’s encounter with madame arnoux on the steamer ville-demontereau was the cause, or part of the cause, of his falling in love with her) is usually more important than physical causation. causes and intentions are almost certainly represented in relation to the characters that have them (intentions) or are affected by them (causes). they help us, as readers, to understand how those characters behave. furthermore, causes and intentions are likely to be represented more strongly in relation to main characters than in relation to subsidiary characters. another interesting aspect of main characters in narratives is how they define perspectives on the events of the narrative. the fact that they do so, again suggests that the representation is primarily character-based. narrative theory distinguishes between two aspects of perspective: narration and focalization (genette, 1980). narration is concerned with who is telling the story. sentimental education has an external narrator, who might be dubbed, albeit controversially, the author. most of its events, however, are seen, to a greater or lesser extent, from the point of view of frédéric, the main focalizer. there are also issues about the extent to which readers might or might not agree with frédéric’s apprehension of (or psychological point of view on) events. readers do not necessarily find him a sympathetic character, and do not always empathize with him. to some extent his actions are conditioned by his social context, but this context is not always seen as excusing his akrasia (though see the quote, above, from flaubert). 2.3 entertainment and edification so far, i have presented mainly impressionistic evidence for the idea that large-scale, longer-term representations of narrative text are character centred. i stated earlier that this conclusion also follows from a marrian task analysis of narrative comprehension, though it is the content of the task analysis that is crucial, not the fact that a task analysis has been performed. marr found that task analysis was easier at lower levels of vision, since the construction of high-level representations of the world around one can have many purposes, not all of which impact directly on visual processes. similarly, it is easier to give a task analysis of word recognition than of highlevel narrative comprehension. nevertheless, the primary reason why people read narratives (rather than secondary reasons such as the academic study of such texts) can be loosely glossed under two general heads: entertainment and edification, though these categories are neither clearly distinct nor mutually exclusive. entertainment, and here i am thinking not just of written text, but of film, television, radio and other media, typically involves portraying people in what are often called “situations”, though these situations are not the situations of situation model theory. they are part of the context against which the text is understood, like mid nineteenth century paris, in sentimental education. situation comedies, for example, often present a set of characters who live together, as a family or otherwise (the general situation in which the characters find themselves), and in each episode a more specific situation is explored. often the characters are unsubtly drawn, and they may be based on archetypal characters, and the individual episodes may be based on stock plots. archetypal characters and stock plots also appear in entertainment films, such as star wars. how do people represent the content, in the broadest sense, of texts, discourses, and items in other media garnham 8 that they have read, listened to or watched, partially or primarily for entertainment? on what do they base their conversations with family, friends, and colleagues about an episode of a situation comedy they have just seen, or that they saw last night? the characters are clearly important, as is the situation (in the broad sense). such situations can be glossed very broadly. for example, older versions of the wikipedia page on situation comedies (wikipedia, 2007) gives a list of plot formulas), which includes such examples as: • attempts to hide egregious mistakes or acts of weakness. • attempts to protect friends and family members from bad news. smaller incidents, which may be closer to the situations of situation model theory, can make an impact on a person’s representation of an episode of a situation comedy, but these incidents are not the primary organizing factors in a person’s memory representation. “texts” produced or consumed for edification are likely to be less crudely drawn. so, even if frédéric moreau is, in some sense, a representative figure, he is not an archetype or a stereotype. nevertheless, the primary memory representation for such text is likely to be similarly based on characters and context. it may be that the main difference between entertainment and edification is how deeply and subtly one (author, reader, or both) tries to relate the content of the text to other matters. and here we re-encounter one of the primary tenets of mental models theory: that the representations derived from texts and those derived from more direct observations, of various kinds, of the real world, are similar in form, and can easily be related to one another. 2.4 other types of text so far, i have talked about narrative text, primarily novels and, indeed, primarily a single novel, gustav flaubert’s sentimental education. but what of other types of text? as mentioned earlier, part, but only part, of the problem with the story grammar approach to the analysis and processing of texts was that it focused specifically on one type of text and did not readily generalize to other types. for example, some texts are not about people at all, so it is hardly likely that their long-term mental representations are person-centred. newscasts are often about situations, in the broad sense described above, rather than individuals, though these situations may be more or less specific (“the situation in the middle east”, “the situation in the palestinian territories”, “the situation in hebron”). nevertheless, some news stories are about individuals, and others are transformed, at least in some people’s minds, into stories about individuals (the oft-heard complaint that modern politics is about personalities rather than policies). note that the reverse rarely happens, which is further evidence for the primacy of people in representations of the content of text and of what happens in the world. furthermore, in much of what used to be the news media – newspapers that can no longer compete for speed of delivery of information with broadcast media and the internet – stories about individuals abound. these stories often appear to be about people whose primary role is to be a person about whom stories are written (a certain type of “celebrity”). what other types of text do people encounter in their day-to-day lives? biographies and autobiographies are, by definition, about individuals. and many other popular, or semi-popular, non-fiction titles focus on individuals. books on art may be about a particular artist, for example, and popular science books often devote a substantial number of pages to anecdotes about individual scientists, many of whom are interesting characters. history books may take their titles from specific historical figures. and so on. indeed, often books of these types either are narratives, or contain substantial narrative elements. there are, however, some types of text that are not narratives, equipment manuals, for example, which are notoriously difficult to write and interpret, though good examples can be found. such texts frequently need to supplement written material how people structure representations of discourse 9 with pictures and diagrams. and although they are not about people, they should take account, and have to take account if they are to be good, of how people interact with pieces of equipment whose functions they are reasonably familiar with, but whose operational details are not known to them. so, manuals for initial installation should, in conjunction with ways of packing and disabling equipment, try to prevent new owners from making their own, usually incorrect, assumptions about how a piece of equipment is going to work. and manuals for operation should be clearly organized around the functions that a piece of equipment has, and the internal structure of that set of functions. academics read abstract texts in the course of their work, as do many other professionals. some texts, such as reports of experiments, have highly and perhaps over-prescribed structures. on one level, these texts are not about people at all, and they only mention people in so far as they are associated with theories or studies relevant to those in the paper in which they are referenced. on another level, it is commonly said that the point of the methods section of an experimental report is to ensure that the study can be replicated by an independent person. to this end it should be written so it can be used as a series of (implicit) instructions on how to re-perform the experiment. it, therefore, needs to take account of what its intended audience will need to know and what can be assumed about what that audience already knows. and the general form of the methods section should have been designed (but perhaps was not) to take account of what we know about how people follow instructions. similarly, a recipe is a set of instructions for preparing a dish, and the structure of good recipes should reflect human thought processes and sensible working practices. for example, all the ingredients should be listed at the beginning, because this format makes it easy for the cook to make sure they have assembled, or know the location of, all the ingredients before they start cooking. presentations of theories are more removed from the kinds of text we have been talking about. they are about abstract entities – the theories and the theoretical constructs from which they are put together. comparatively little has been written about abstract texts in the situation models literature. however, abstract objects are not so different from concrete objects, except that they are abstract, and it is a reasonable assumption that human thought is primarily geared to concrete situations. thinking abstractly is generally difficult – abstract objects do not impinge on us in the same direct, immediate way that, in particular, other people (who are concrete objects in this sense) do. linguistically, abstract objects behave similarly to concrete ones. if i use the term “the law” in a particular context, i may be referring to a specific law, probably in my case one that is in force in england and wales. in different contexts, i will be referring to different law. or i could be using the term “the law” in a related sense to mean the body of law. i can follow “the law” with an identity of reference anaphoric pronoun “it” to refer to the same law again (in the same context), just as i can follow “the table” with “it”. some abstract objects are individual abstract objects (baddeley & hitch’s, 1974, theory of working memory, for example), but so are many concrete objects (nelson’s column in trafalgar square, london, for example). 3. protagonists and models of dialogue and discourse from what has been said so far, the failure to accommodate properly the role of main protagonists appears to be a phenomenon specific to the study of dialogue and discourse, and indeed one specific corner of this discipline – the theory of mental models or situation models (see section 4 for peoplecentred representations in other disciplines). in my own discipline of psycholinguistics there is a focus on comprehension rather than production. so, in the case of discourse models, the primary accounts are of how such models are built up during reading or listening. as we have seen, information in a clause leads to the representation of an eventuality (event, state, or process), a process that can be conceived of in a standard compositional framework, such as that of discourse representation theory (drt, kamp & reyle, 1993) or file change semantics (heim, 1983). alternatively, and perhaps related to the use of the term “situation model”, it can be thought of in the more loosely defined framework of barwise and perry’s (1983) situation semantics. garnham 10 it is, of course, widely recognized that dialogue and discourse has structure above the level of eventualities, and it would be natural to look for an account of the importance of main characters at this higher level. indeed, drt and file change semantics specifically address some issues about (anaphorically) relating references to entities in different eventualities. furthermore, and again as mentioned above, the different eventualities in a discourse must be related to one another as eventualities, at what can be regarded as a higher level of structure. the most common approach is to identify a set of relations that can hold between them. one view, consistent with the mental models doctrine that discourse models have the same general structure as models of the world constructed in other ways (e.g. from personal experience), is that there is a set of types of relation among eventualities, (spatial, temporal, logical, causal, intentional, and moral, miller & johnsonlaird, 1976). another view, exemplified for example in mann and thompson’s (1986) rhetorical structure theory (rst), is that there is a fixed set of specific relations that connect usually contiguous eventualities. in this context the relations are often referred to a holding between textual elements, called text spans or discourse segments. attempts to integrate these two levels of structure (eventualities and the relations among them) in a single framework have been attempted, most notable in asher and lascarides’s (2003) segmented discourse representation theory (sdrt). however, whether integrated with clauselevel structure or not, approaches based on discourse relations do not directly address the idea that protagonists are central in representations of (some) discourses, largely because they focus only on local relations between eventualities. of course, drt and file change semantics specifically address questions of co-referential anaphora. and some approaches to discourse coherence, particularly that of hobbs (1983) and those that derive from it (e.g. kehler et al., 2008), claim that co-reference is often established as a “side effect” of establishing coherence. nevertheless, there is no notion, in either of these approaches, that reference to the same person in multiple adjacent clauses has any special status. and clearly, a sequence of sentences (or utterances) and, hence, eventualities that have one or more entities in common between adjacent utterances does not necessarily form an interesting, or even a coherent discourse. it should also be noted that the general marrian approach, mentioned above, of carrying out a task analysis of discourse comprehension, does not of itself lead to the conclusion that protagonists have a special importance. it would be the result of carrying out the task analysis, not the fact that it is carried out, that would be the source of the solution. in other words, it must be independently assumed (or argued) that, at least in many cases, what is meant by conveying or receiving information about some world is writing or reading (or speaking or hearing) a story centred upon a central character. part, but not all, of what it means to be a central character is that there will be repeated reference to that person. related to this fact, pronouns, and np-anaphora more generally, are crucial in connecting discourse. a first step in going beyond individual anaphoric links is found in centering theory (grosz, joshi, & weinstein, 1995). centering theory’s rule 2, “sequences of continuation are preferred over sequences of retaining; and sequences of retaining are to be preferred over sequences of shifting” (grosz et al, 1995: 17) reflects the notion that repeated reference to the same person is expected, with that person remaining the focus of attention (continuation vs. retention). however, this notion far from fully captures the centrality of main characters over longer stretches of text. grosz (see, grosz & sidner, 1986) incorporated ideas that had already been developed in unpublished work on centering theory into an account of discourse structure that made use of the notions of attention and intention. centering theory provided mechanisms for local attention and particularly the search for referents for pronouns. the importance of intentions is that participants in a discourse are trying to achieve certain effects in each segment of the discourse. this idea is loosely related to the notion of performative utterances developed by austin (1962), which led to searle’s (1969) notion of speech acts. however, while grosz’s work, and subsequent work on dialogue inspired by it (e.g. rich & sidner, 1998; lochbaum, 1998; kraus, 2001; hirschberg & nakatani,1996) gives a key role to participants, and the intentions that they have in their roles as how people structure representations of discourse 11 agents in dialogue, in explaining its structure, and indeed in the understanding of the purpose of dialogues, the questions addressed in these theories are rather different from the ones about central characters in extended narratives addressed here. indeed, they are more akin to issues about the relation between writers and readers of novels. interestingly, and from a quite different perspective, that of conversation analysis, sacks (1995, vol. 2, part vii, lectures 9 -12, pp.470-494) makes a series of observations about how jokes (or one particular joke) is structured to take account of its recipient (the partner in the dialogue in which the joke is told, and, in some sense, regarded as an agent) but also about other aspects of how the joke is structured around protagonists in the story that constitutes the joke. however, insightful though these observations are, they do not resolve the issue of the central role of protagonists in stories. a different set of approaches to higher level structure in “stories” has its origins in structuralism. as the term “structuralist” suggests, these approaches, which vary considerably in tone and orientation, are only indirectly concerned with central characters. one such approach, which inspired psycholinguistic work on story grammars, is propp’s (1928/1968) morphology of the folk tale. in his analysis, propp identifies the roles of hero and villain in folk tales, though these are only two out of seven abstract character functions (villain, dispatcher, helper, prize/princess (or her father), donor, hero, false hero) that propp identifies as occurring in folktales. propp’s work provides a direct analysis of a narrow class of texts, though the structures he identifies can clearly have reflexes in more complex texts. indeed, propp focuses on russian folktales – perhaps folktales from different cultures have different characteristics. related to the analysis of folktales, but again with limited application to only some types of narrative text, is the idea of the monomyth or hero’s journey, discussed by campbell (1949), which clearly does focus on an individual character – the hero. the hero’s journey, with its (up to) 17 stages, is now generally thought of as an analysis, albeit a somewhat controversial and outdated one, within the broader discipline of narratology, though that term itself was not coined until much later (todorov, 1969). narratology draws on ideas from semiotics, and there has been a debate about whether a syntagmatic or a paradigmatic approach to the structure of narratives is the most fruitful one, a debate that has carried over into the psychological literature on story structure and story grammars. todorov’s approach was syntagmatic, and hence related to syntactic approaches to story grammar, whereas others (e.g. levi-strauss, 1958/1963) have suggested that a paradigmatic approach is more appropriate for understanding narratives. narratology has also spawned a computational branch (cavazza & pizzi, 2006), influenced by ai work on story generation and comprehension, for example the work on plot units by lehnert (1981). there are many other analyses within narratology and literary theory that identify high-level structure in discourses of certain kinds, for example the three-act analysis (setup, confrontation, resolution, sometimes satirized as beginning, middle and end) of plays and films, which can be traced back aristotle’s poetics. however, none of these approaches are specifically concerned with explaining the role of central characters, even though most, probably all, of them accommodate the idea that there are such characters. 4. other lines of evidence so far, i have done two things. on the one hand, i have presented impressionistic evidence that characters or protagonists have a central role in the mental representation of discourse and text. on the other hand, i have argued that attempts to identify structure in discourse above the level of the information conveyed by single clauses have not concerned themselves with explaining this idea, though it is unlikely that students of lengthy narratives would deny its truth. i now turn to other kinds of evidence, largely from other disciplines and other sub-disciplines of psychology that appear to embrace the idea more directly. i then ask how mental models or situation models theory may have to be modified to accommodate the claim. my point is not that the claim needs to be garnham 12 established, for example because it is controversial in some sense. more likely it is so obvious to those studying discourse that the need to explain it has been overlooked. so, i do not intend to go into detail about how it has been established in other disciplines. rather, i want to underline the fact that it is, and should be, obvious, and that there needs to be a clear place for it in theories of discourse representation. one of the central ideas underlying the mental models/situation models approach to discourse comprehension is that understanding events described in text or dialogue is essentially like understanding events that are directly experienced. an extension to this idea, related to our everyday experience, is that we think about both individual events and more complex sets of interrelated events primarily in terms of the people involved in them. this everyday observation is reflected in approaches to understanding the world in certain parts of psychology, perhaps more directly in social psychology than in the psychology of language. one particularly relevant subarea of social psychology is social perception, sometimes tellingly referred to as person perception, and within social perception, impression formation and attribution theory. research in person perception emphasizes the importance of people in how we construe the world, and the complexities of our perceptions of, and hence our representations of, people. ideas about impression formation have informed thinking about literary character in literary theory, and its interplay with cognitive psychology. marilynn brewer’s (1988) dual-process theory of impression formation influenced gerrig and allbritton’s (1990) work on literary character, considered from the perspective of cognitive psychology. and this work in turn influenced that of ralf schneider (2001), who incorporated ideas from mental models/situation models theory into his account of literary character. schneider’s conclusion (2001: 610) that “readers of novels focus their attention predominantly on psychological traits, emotions, and aims of characters that are more abstract and less dependent on the immediate circumstantial conditions of individual situations” is clearly consistent with the claims made this paper. the importance of “literary character” in various literary genres and in other media, such as film, poetry, and comic books, is further emphasized in the collection of papers by eder, jannidis, and schneider (2010). however, interesting though this work is, particularly in the distinction that it indirectly draws attention to between the highly elaborated representations of characters that we encounter in books and other media in which we become engrossed, and the almost certainly sparser representations typically developed by participants reading brief texts in psycholinguistic experiments, the question of how elaborated our representations of characters become is at least to some extent orthogonal to the question, central to the current paper, of why the protagonist index appears to behave differently from other indices in the event-indexing model of zwaan et al. (1995). a second aspect of social perception, attribution theory, focuses on the importance of people’s characteristics and dispositions in explaining their behavior. and although attribution theory was developed independently of considerations about discourse and text, its ideas are clearly relevant to understanding texts of many kinds. furthermore, one of the key ideas in attribution theory is that we tend, at least in the case of other people, to overemphasize the role of personal characteristics, as opposed to external factors, in explaining behavior (jones & harris, 1967), a tendency that was dubbed the fundamental attribution error by lee ross (1977). the fundamental attribution error reinforces the idea that people are particularly important in the way we understand the world. it is an indication that we like to explain things in terms of (people’s) dispositions, not the situations they find themselves in, and it reflects the fact that we think in terms of people. interestingly, however, there are differences in attributions between individualistic and collective cultures (e.g. miller, 1984), including differences in the fundamental attribution error (e.g., choi, nisbett, & norenzayan, 1999, koenig & dean, 2010). in an example more directly related to narration, marcus, uchida, omorigie, townsend, and kitayama (2006) showed that japanese reports on olympic success were less focused on characteristics of the successful athletes (dispositional attributions) than reports in the united states, and more on context. these differences, and their relation to psychological theories of discourse comprehension, would be well worth pursuing in how people structure representations of discourse 13 another context. however, the focus on contextual explanations of behavior is not, of itself, at odds with the idea that stories are organized around central characters, and indeed many important works of literature from collectivist cultures are either named after central characters or at least focus on them. for example the sixteenth century ming dynasty “novels” xiyouji, usually translated as journey to the west, and jinpingmei, the golden lotus, are about the buddhist monk xuanzang and about hsi men (and his six wives) respectively. and the title of the seventeenth century japanese “floating world” novel, the life of an amorous man, clearly indicates its focus on a main protagonist. there is also work that is more central to the psychology of language that points in the same direction as the observations made above about the importance of main characters in discourse processing. for example, although much psychological work on the comprehension of stories became embroiled in an unilluminating debate (see garnham, 1983) about whether stories have a structure similar to sentences, it is clear from discussions of what, from a psychological point of view, constitutes a story (e.g., stein, 1982) that protagonists are central. in addition, work with much shorter texts (e.g., anderson, garrod, & sanford, 1983) suggest that it is principal characters, like frederic in l’education sentimentale, that are crucial in the memory representations of narratives, and that peripheral characters are processed much less deeply. it might also be thought that the approach of centering theory (e.g., gordon, grosz, & gilliom, 1983), with its single backward-looking center and its rule 2 about the relations between forward-looking centers in adjacent utterances (see above), supports this idea. finally, work by radvansky, spieler, & zacks (1993) used the fan effect – the effect of the number of associations that a concept has (anderson, 1974) – to provide support for a person-based, rather than a location-based, representation of information derived from a text about different people in different locations. however, these observations simply point out the centrality of people (or of a main protagonist) in text representation, rather than providing an explanation of why they are central, or details of what such representations look like. more generally, in their review of situation models, zwaan and radvansky (1998) make a number of allusions to the importance of main characters, but without explaining how those claims inform situation models theory and its account of event indexing. for example, they describe protagonists and objects as “meat” on a “backbone” of goal structures (1998: 173), which depend on the intentional and causal indices. and they later state (1998: 179) that “main protagonists are a crucial component of situation models. most narratives, ranging from the odyssey to the short passages used in psycholinguistic experiments, describe the goals and actions of a main protagonist.”. furthermore, they note that although in other types of text “objects can also function as a central element of situation models, for example, in a textbook chapter about the heart or a printer manual.” (1988: 179), “readers appear to be intensively engaged in keeping track of protagonists during comprehension whereas the amount of focus on objects appears to be more dependent on contextual cues.” (1988:173). however, despite these observations, protagonist remains just one of the five indices of the theory. 5. modifications to mental models/situation models theory the event-indexing model’s focus on events is consistent with the claim that events are primitives in natural language semantics (davidson, 1967; parsons, 1990). the model is also correct that there are certain aspects of events (“indices”) that allow them to be linked to other events, and that the linking of events is essential in coming to an understanding of a wider situation, whether directly experienced in the world or described in a dialogue or text. however, the different indices play different roles, which partly depend on the type of text. for example, in instruction manuals spatial relations may play a crucial role. but, interestingly, such manuals often include diagrams and photos, as spatial information is notoriously difficult to convey purely verbally (morrow et al., 1987; morrow et al., 1989). in narratives, novels, biographies and so on, people (protagonists) are garnham 14 more central, as i have been arguing. also, as miller & johnson-laird’s (1976) list of types of relation between events (which excludes protagonist) indicates, the way that the protagonist index in situation model theory works (via anaphora) is at least superficially different from the way other indices work, though tense and aspect, used in establishing temporal links between events, have anaphoric properties (partee, 1973). although the situation models literature in general, and zwann and radvansky’s (1998) review in particular, have many useful insights into the role of central characters in the construction of text representations, these insights have remained dissociated from the core ideas of the theory of situation models. they have not been directly reflected in developments of the theory and its related event indexing model, or of mental models theory. indeed, it may appear that the problem with the event-indexing model is that it treats protagonist as just one of five types of index on events (time, space, cause, intention, protagonist), and hence loses sight of the more central role of protagonists. however, this conclusion is misleading, because of a difference between the protagonist index and the other indices. miller & johnson-laird (1976) identified six types of relations between events: spatial, temporal, logical, causal, intentional, and moral. protagonist is not among them four correspond to four of the five indices in the event-indexing model. of the other two, logical corresponds to a type of relation that is relatively rarely important in stories, perhaps with exceptions such as sherlock holmes stories, but which reflects johnson-laird’s interest in reasoning. moral is, perhaps, an omission from the event-indexing theory, but proponents of that theory are open to the idea of additional indices. as noted in the introduction, these types of relation correspond to what are known in other literatures as (local) coherence relations (though temporal relations have some properties of coherence relations and some properties of anaphoric relations). links between characters in the different events described in a text are not in miller & johnsonlaird’s list but are established via anaphoric processing and sometimes by inference. anaphoric processes have received more attention within the mental models framework (e.g., garnham, 2001) than in the situation models framework. however, like coherence relations, anaphoric relations typically establish local links in a text. main characters in extended text do not just provide links between eventualities mentioned close together. they provide links across broader spans of text. it is for this reason that the event-indexing model fails to say anything about them. and neither do the accounts of anaphoric processes in mental models theory. from this broader perspective protagonists, at least in stories, behave differently from other things that are linked by indices. indices provide local links between events that are mentioned close together in a text. but across a wider stretch of text, there may be a series of links, in a broader sense, all involving the same text character. furthermore, the links based on the other indices are, in a clear sense, subsidiary to the protagonist links. spatial links exist because the same character is in different places. we do not (usually) see a chain of references to one or more characters subserving the need to present a set of places with a particular spatial relations to one another. similarly, the links created using temporal, causal and intentional indices are there because of a story about one or more individuals with temporal, causal and intentional properties. as just mentioned (see, also, garnham, 1991), miller and johnson-laird’s spatial, temporal, logical, causal, intentional and moral relations between events are the basis of the (local) coherence relations, which have been analyzed in differing ways in different theories of text structure (e.g., hobbs, 1983; mann & thompson, 1986). and the event-indexing model gives a handle on the properties of event representations that allow these links to be made. however, discourses have a structure sitting on top of these local coherence relations, and it is within this higher-level structure that the importance of central characters should be explained. many suggestions have been made about what this structure is (see above). perhaps the most important ones come out of the literature on stories, but this literature became embroiled in two issues that distracted attention from attempts to give a proper account of the structure of texts. the first arose from the idea that stories formed a distinct category from other texts, so that lessons about stories could not necessarily be generalized to other types of text. the second derived from the notion that the structure of stories mirrored that how people structure representations of discourse 15 of sentences. however, a more intuitive idea about the structure of stories, or a least of a minimal story, is what nancy stein appears to be searching for in her 1982 article, though her arguments are buried amongst forays against those attacking story grammars. a clearer statement of what stein takes to be a minimal story appears in stein & glenn (1979), and it includes the idea of a central protagonist. however, the thrust of the current paper is that its claims are not specific to stories, in the narrow sense of the story grammar literature, so stein’s insight needs to be reassessed from this perspective. how do we account for the fact that a wide a range of texts are organized around the idea of a central character, or perhaps a small group of central characters? a further issue for stories, that of whether an affective response is crucial to their definition, and whether it is a response of the characters (stein & glenn, 1979; mandler & johnson, 1977) or of the reader (brewer & lichtenstein, 1982) or both, is less directly relevant to the current paper, but needs careful consideration in a general theory of narrative. a different idea about global structure in text, one that is more closely related to situation model theory, is that of casual chains (trabasso & van den broek, 1985; van den broek & trabasso, 1986). as mentioned above, causal chains are said to provide the “backbone” of situation models. however, two related points should be made in connection with this idea. first, event indexing produces local cause-effect links, but causal chains need to show global coherence. second, as already argued, it is not chains of causes (including causes of intentional actions) per se that are important. the characters that form the “meat” of the models sit on these chains, and causal chains should reflect coherent sets of actions that people, or often a single main character, engage in and that are worth narrating. while it is clear that texts have structure above that of local coherence (and anaphoric) relations, it is not easy to give a definitive account of what that structure is. it includes plot structure, but plots are not as stereotyped as those suggested by story grammars, so flexibility is required in how plots are specified. it also includes major, recurring characters and the relations among them. these characters are linked to the eventualities and small-scale situations that they are involved in. their representations are likely to develop into complex character models (e.g., rapp, gerrig, & prentice, 2001). these character models contain components that occur in the representation of individual eventualities – an eventuality may ascribe a property to an individual, for example. it is simply that these character models are much richer than representations that can be derived from individual eventualities. however, the combinations of properties and relations in character models are likely to give rise to complex inferences about characters and their behaviour, including their interactions with other characters, that will not follow from the information in a single clause. they determine our reactions, as readers, to those characters. and they determine those characters reactions to each other, including the whole range of affective reactions seen in everyday life (see also the comments about affect and stories, above). protagonists often serve as focalizers, in the sense of genette (1980), as with frédéric in sentimental education, and the narrator must present the story from the point of view of the focalizer. the relation between the narrator and the focalizer is likely represented at this level, and the content of this representation reflects johnson-laird and garnham’s (1980) claim, mentioned at the beginning of this paper, that discourse models need to represent what participants in discourse, and by extension characters in or around the narrator of the discourse, know about each other. the narrator’s understanding of what the focalizer perceives and knows allows the narrator to tell the story from the focalizer’s point of view. 6. conclusions the basic insight of the theory of mental models (or situation models) is correct. at a local level, information in text is about eventualities, and it is encoded in a “who did what to whom, where, when and how” format. eventualities described in clauses close together have relations to one another such as cause-effect, or evidence-assertion, sometimes signalled textually (e.g. by conjunctions, such as “because”) and sometimes suggested by knowledge about the world, and garnham 16 these relations must be computed in order to make sense of the text. indexing is part of what makes these links possible. eventualities may primarily involve only one entity, or they may present relations between two or more entities. at any point in a text, one or more of these entities is likely to be important (“salient”, “psychologically focused”). in a narrative text, and not by coincidence, this entity is likely to be a person, and the same person is likely to be the most important entity at many points in the same text. eventualities mentioned close together often involve the same entity, and the identity of the entity between events must be computed, via anaphoric and inferential processes, which might be described as a process based on indexing. from a broader perspective, repeated reference to an entity (usually a person) across larger stretches of a text happens because the text is about that person, or partly about that person, or about that person as one of a small group of central characters. and our representations, as readers, of the content of such texts are structured by our representations of the central character or characters. part of the reason is that we relate easily but significantly to other people. we can be entertained by their exploits, and we can be edified by their successes and failures. by saying “or characters” i am acknowledging that texts do not necessarily have a central character, and that any other character must be necessarily subsidiary, though some like sentimental education do. stories can have more than one central character – two protagonists, for example, or a protagonist and an antagonist (othello and iago; harry potter and voldemort), although some literary analyses claim that separate protagonists, unless they are part of a group, should each be associated with their own story. alternatively, of two apparent main characters, one may be more important than the other. at the level of representation at which the central character or the set of important characters (there are said to be 15 major characters in tolstoy’s war and peace, for example) is represented, so must at least some of the relations among them. the relation between the narrator and the focalizer (see above) may also represented at this level. however, these observations about multiple characters, though interesting in themselves, are not crucial to the arguments of this paper, and so i will not develop them further. representations of central or major characters, and the relations between them, are at a level above the representation of eventualities and the local relations between them, and our representations of text content tend to be organised around them. people structure representations of texts. and they do so at a level above that of local coherence. this claim is true directly of narratives, whether fictional or about the real world. it is also true more indirectly for other types of text. recipes, for example, should be written with a user in mind. situations, in a broader sense than that implied by the term “situation model”, are also crucial in the longer-term representation of text content. sometimes, in some newscasts for example, such situations are where our interests lie, rather than in the plights of individuals caught up in those situations. at other times, we can only understand the actions of people (or characters in fiction) in relation to the situations they find themselves in. i am not claiming that forming people-centred representations is all there is to the representation of texts and discourse, or even of narrative text. many other considerations apply, some discussed at length in texts on narrative comprehension, such as emmott (1997). i am however, claiming, that we need to go beyond situation model theory, as it now stands, and consider more global levels of text structure and text comprehension, to have a proper psychological account of the understanding of discourse and text. the exact nature of these higher levels of text structure remains to be determined. references anne anderson, simon c. garrod, and anthony j. sanford (1983). the accessibility of pronominal antecedents as a function of episode shifts in narrative text. quarterly journal of experimental psychology, 35(3):427-440. https://doi.org/10.1080%2f14640748308402480 how people structure representations of discourse 17 john r. anderson (1974). retrieval of propositional information from long-term memory. cognitive psychology, 6(4):451-474. https://psycnet.apa.org/doi/10.1016/00100285(74)90021-8 nicolas asher and alex lascarides (2003). logics of conversation. studies in natural language processing. cambridge university press, cambridge. john l. austin (1962). how to do things with words. clarendon press, oxford. alan d. baddeley and graham j. hitch (1974). working memory. in gordon h. bower, editor, the psychology of learning and motivation (vol. 8), pages 47-90. academic press, new york. jon barwise and john perry (1983). situations and attitudes. mit press/bradford books, cambridge, massachusetts. john d. bransford and jeffery j. franks (1971). the abstraction of linguistic ideas. cognitive psychology, 2(4):331–350. https://doi.org/10.1016/0010-0285(71)90019-3 marilynn b. brewer (1988). a dual process model of impression formation. in thomas k. srull & robert s. wyer, jr., editors, advances in social cognition. vol. 1: a dual process model of impression formation, pages 1-36. lawrence erlbaum associates, hillsdale:, new jersey. william f. brewer and edward h. lichtenstein (1982). stories are to entertain: a structural-affect theory of stories. journal of pragmatics, 6(5-6):473-486. https://doi.org/10.1016/03782166%2882%2990021-2 joseph campbell (1949). the hero with a thousand faces (1st ed.). princeton university press, princeton, new jersey. marc cavazza and david pizzi (2006). narratology for interactive storytelling: a critical introduction. in stefan gobel, rainer malkewitz and ido iurgel, editors, technologies for interactive digital storytelling, third international conference, pages 72-83. lecture notes in computer science, volume 4326. springer verlag, berlin. https://doi.org/10.1007/11944577_7 incheol choi, richard e. nisbett and ara norenzayan, (1999). causal attribution across cultures: variation and universality. psychological bulletin, 125(1):47-63. https://doi.org/10.1037/0033-2909.125.1.47 marcelle crinean and alan garnham (2006). implicit causality, implicit consequentiality and semantic roles. language and cognitive processes, 21(5):636-648. https://doi.org/10.1080/01690960500199763 donald davidson (1967). the logical form of action sentences. in nicholas rescher, editor, the logic of decision and action, pages 81-95. university of pittsburgh press, pittsburgh, pennsylvania.. jens eder, fotis jannidis and ralf schneider (2010). characters in fictional worlds: understanding imaginary beings in literature, film, and other media. de gruyter, berlin. catherine emmott (1997). narrative comprehension: a discourse perspective. clarendon press, oxford. fernanda ferreira, karl g. d. bailey and vittoria ferraro (2002). good-enough representations in language comprehension. current directions in psychological science, 11(1):11-15. https://doi.org/10.1111%2f1467-8721.00158 alan garnham (1983). what's wrong with story grammars. cognition, 15(1-3):145-154. https://doi.org/10.1016/0010-0277(83)90037-9 alan garnham (1991). where does coherence come from: a psycholinguistic perspective. occasional papers in systemic linguistics, 5:131-141. alan garnham (1996). the other side of mental models: theories of language comprehension. in jane v. oakhill and alan garnham, editors, mental models in cognitive science: essays in honour of phil johnson-laird, pages 35-52. psychology press, hove, east sussex. alan garnham (2001). mental models and the interpretation of anaphora. psychology press, hove, east sussex. garnham 18 alan garnham (2020). integration: key but not so simple. discourse processes. special issue on “integration: the keystone of comprehension” jane oakhill and kate cain, editors. https//doi.org/10.1080/0163853x.2020.1758546 online first. alan garnham, scarlett child and sam hutton (2020). anticipating causes and consequences. journal of memory and language, 114:104130. https://doi.org/10.1016/j.jml.2020.104130 catherine garvey and alfonso caramazza. (1974). implicit causality in verbs. linguistic inquiry, 5(3):459-464. https://www.jstor.org/stable/4177835 gérard genette (1980). narrative discourse: an essay in method, j. e. lewin., translator. cornell university press, ithaca, new york. richard g. gerrig and david w, allbritton (1990). the construction of literary character: a view from cognitive psychology. style, 24(3):380-391. https://www.jstor.org/stable/42945868 peter c. gordon, barbara j. grosz and laura a gilliom (1993). pronouns, names, and the centering of attention in discourse. cognitive science, 17(3):311–347. https://doi.org/10.1207/s15516709cog1703_1 barbara j. grosz, aravind k. joshi and scott weinstein (1995). centering: a framework for modeling the local coherence of discourse. computational linguistics, 21(2):202–225. https://doi.org/10.5555/211190.211198 barbara j. grosz and candice l. sidner (1986). attention, intentions, and the structure of discourse. computational linguistics, 12(3):175-204. https://doi.org/10.5555/12457.12458 irene heim (1983). file change semantics and the familiarity theory of definiteness. in rainer. bäuerle, christoph schwarze, and arnim von stechow, editors, meaning, use, and interpretation of language, pages 164-189. walter de gruyter, berlin. julia hirschberg, and christine h. nakatani (1996). a prosodic analysis of discourse segments in direction-giving monologues, in proceedings of the 34th annual meeting of the association of computational linguistics, pages 286-293. santa cruz, california. https://doi.org/10.3115/981863.981901 jerry r. hobbs (1983). why is discourse coherent? in fritz neubauer, editor, coherence in natural language texts, papers in text linguistics, volume 38, pages 29-69. helmut buske verlag, hamburg. philip n. johnson-laird, (1970). the perception and memory of sentences. in john lyons, editor, new horizons in linguistics, pages 261-270. penguin books, harmonsworth, uk. philip n. johnson-laird (1983). mental models: towards a cognitive science of language, inference, and consciousness. cambridge university press, cambridge. philip n. johnson-laird and alan garnham (1980). descriptions and discourse models. linguistics and philosophy, 3(3):371-393. https://link.springer.com/article/10.1007%2fbf00401691 edward e. jones and victor a. harris (1967). the attribution of attitudes. journal of experimental social psychology, 3(1):1–24. https://doi.org/10.1016/0022-1031(67)90034-0 hans kamp (1979). events, instants and temporal reference. in rainer bauerle, urs egli, & arnim von stechow, editors, semantics from different points of view, pages 376-417. springer verlag, berlin. hans kamp and uwe reyle (1993). from discourse to logic: introduction to modeltheoretic semantics of natural language, formal logic and discourse representation theory. kluwer academic publishers, dordrecht. hans kamp and christian rohrer (1983). tense in texts. in rainer bauerle, christoph schwarze and arnim von stechow, editors, meaning, use, and interpretation of language, pages 250269. walter de gruyter, berlin. andrew kehler, laura kertz, hannah rohde and jeffrey elman (2008). coherence and coreference revisited. journal of semantics, 25(1):1-40. anne m. koenig and kristy k. dean (2010). cross cultural differences and similarities in attribution. in kenneth d. keith, editor, cross cultural psychology: contemporary themes and perspectives, volume 42, pages 475-493. wiley-blackwell, malden, massachusetts. how people structure representations of discourse 19 sarit kraus (2001). strategic negotiation in multi-agent environments. mit press, cambridge, massachusetts. wendy g. lehnert (1981). plot units: a narrative summarization strategy. in wendy g. lehnert & martin h. ringle, editors, strategies for natural language processing. lawrence erlbaum associates, hillsdale, new jersey. claude levi-strauss (1958/1963). structural anthropology, c. jacobson and b. g, schoepf translators. basic books, new york. (originally published in french in 1958) karen e. lochbaum (1998). a collaborative planning model of intentional structure. computational linguistics, 24(4):525-572. william c. mann and sandra a. thompson (1986). relational propositions in discourse. discourse processes, 9(1):57-90. https://doi.org/10.1080/01638538609544632 jean m. mandler, and nancy s. johnson (1977). remembrance of things parsed: story structure and recall. cognitive psychology, 9(1):111-151. https://psycnet.apa.org/doi/10.1016/00100285(77)90006-8 hazel r. marcus, yukiko uchida, heather omorigie, sarah s. m. townsend, and shinobu kitayama (2006). going for the gold: models of agency in japanese and american contexts. psychological science, 17(2):103-112. https://doi.org/10.1111/j.1467-9280.2006.01672.x david marr (1982). vision: a computational investigation into the human representation and processing of visual information. freeman, san francisco, california. daniel g. morrow, gordon h. bower and steven l. greenspan (1989). updating situation models during narrative comprehension. journal of memory and language, 28(3):392-312. https://psycnet.apa.org/doi/10.1016/0749-596x(89)90035-1 daniel g. morrow, steven l. greenspan and gordon h. bower (1987). accessibility and situation models in narrative comprehension. journal of memory and language, 26(2):165-187. https://doi.org/10.1016/0749-596x%2887%2990122-7 george a. miller and philip n. johnson-laird (1976). language and perception. cambridge university press, cambridge. joan g. miller (1984). culture and the development of everyday social explanation. journal of personality and social psychology, 46(5):961-978. https://psycnet.apa.org/doi/10.1037/00223514.46.5.961 jane v. oakhill, alan garnham and wietske vonk (1989). the on-line construction of discourse models. language and cognitive processes, 4(3-4):263-286. https://doi.org/10.1080/01690968908406370 terence parsons (1990). events in the semantics of english: a study in subatomic semantics. mit press, cambridge, massachusetts. barbara h. partee (1973). some structural analogies between tenses and pronouns in english. journal of philosophy, 70(18):601-609. https://doi.org/10.2307/2025024 george r. potts (1974). storing and retrieving information about ordered relationships. journal of experimental psychology,103(3):431-439. https://psycnet.apa.org/doi/10.1037/h0037408 vladimir propp (1968). morphology of the folktale, laurence. scott, translator. university of texas press, austin, texas. (originally published in russian, 1928.) gabriel a. radvansky, daniel h. spieler and rose t. zacks (1993). mental model organization. journal of experimental psychology: learning, memory, and cognition, 19(1):95-114. https://psycnet.apa.org/doi/10.1037/0278-7393.19.1.95 david n. rapp, richard j. gerrig and deborah. a. prentice (2001). readers’ trait-based models of characters in narrative comprehension. journal of memory and language, 45(4):737–750. https://doi.org/10.1006/jmla.2000.2789 charles rich & candice l. sidner (1998). collagen: a collaboration manager for software interface agents. user modeling and user-adapted interaction, 8(3):315–350. https://doi.org/10.1023/a:1008204020038 garnham 20 harvey sacks (1995). lectures on conversation, volumes i and ii, gail jefferson, editor. blackwell, malden, massachusetts. john searle (1969). speech acts. cambridge university press, cambridge. lee ross (1977). the intuitive psychologist and his shortcomings: distortions in the attribution process in leonard berkowitz, editor, advances in experimental social psychology, volume 10, pages 173-220. academic press, new york. anthony j. sanford and patrick sturt (2002). depth of processing in language comprehension: not noticing the evidence. trends in cognitive sciences, 6(9):382-386. https://doi.org/10.1016/s1364-6613(02)01958-7 ralf schneider (2001) toward a cognitive theory of literary character: the dynamics of mentalmodel construction. style, 35(4):607-640. https://www.jstor.org/stable/10.5325/style.35.4.607 nancy l. stein (1982). the definition of a story. journal of pragmatics, 6(5-6):487-507. https://doi.org/10.1016/0378-2166(82)90022-4 nancy l. stein and christine g. glenn (1979). an analysis of story comprehension in elementary school children. in roy o. freedle, editor, advances in discourse processing, volume. 2: new directions in discourse processing, pages 53-120. ablex, norwood, new jersey. richmond h. thomason, editor (1974). formal philosophy: selected papers of richard montague. yale university press, new haven, connecticut. tzvetan todorov (1969). grammaire du décameron. mouton: the hague. thomas trabasso and paul van den broek (1985). causal thinking and the representation of narrative events. journal of memory and language, 24(5):612-630. https://doi.org/10.1016/0749-596x(85)90049-x paul van den broek, and thomas trabasso (1986). causal networks versus goal hierarchies in summarizing text. discourse processes, 9(1):1-13. https://doi.org/10.1080/01638538609544628 teun a. van dijk and walter kintsch (1983). strategies of discourse comprehension. academic press, new york. wikipedia (2020, july 23). list of book titles taken from literature. wikipedia. https://en.wikipedia.org/wiki/list_of_book_titles_taken_from_literature wikipedia (2007, december 22) sitcom. wikipedia. https://en.wikipedia.org/w/index.php?title=sitcom&oldid=179522509 rolf a. zwaan, (1996). toward a model of literary comprehension. in bruce k. britton and arthur c. graesser, editors, models of understanding text, pages 241-255. lawrence erlbaum associates, hillsdale, new jersey. rolf a. zwaan, mark c. langston and arthur c. graesser (1995) the construction of situation models in narrative comprehension: an event-indexing model. psychological science, 6(5):292-297. https://doi.org/10.1111%2fj.1467-9280.1995.tb00513.x. rolf a. zwaan and gabriel a. radvansky (1998). situation models in language comprehension and memory. psychological bulletin, 123(2):162-185. https://doi.org/10.1037/00332909.123.2.162 rolf a. zwaan, gabriel a. radvansky, amy e. hilliard and jacqueline m. curiel (1998). constructing multidimensional situation models during reading. scientific studies of reading, 2(3):199-220. https://psycnet.apa.org/doi/10.1207/s1532799xssr0203_2 rolf a. zwaan and herre van oostendorp. (1993) do readers construct spatial representations in naturalistic story comprehension? discourse processes, 16(1-2):125-143. https://psycnet.apa.org/doi/10.1080/01638539309544832 dialogue & discourse 9(1) (2018) 107–127 doi: 10.5087/2018.104 changemyview through concessions: do concessions increase persuasion? elena musi em3202@columbia.edu data science institute columbia university debanjan ghosh dg513@mit.edu mcgovern institute for brain research massachusetts institute of technology smaranda muresan smara@columbia.edu data science institute columbia university editor: maite taboada submitted 07/2017; accepted 07/2018; published online 08/2018 abstract in discourse studies concessions are considered among those argumentative strategies that increase persuasion. we aim to empirically test this hypothesis by calculating the distribution of argumentative concessions in persuasive vs. non-persuasive comments from the changemyview subreddit. this constitutes a challenging task since concessions are not always part of an argument. drawing from a theoretically-informed typology of concessions, we conduct an annotation task to label a set of polysemous lexical markers as introducing an argumentative concession or not and we observe their distribution in threads that achieved and did not achieve persuasion. for the annotation, we used both expert and novice annotators. with the ultimate goal of conducting the study on large datasets, we present a self-training method to automatically identify argumentative concessions using linguistically motivated features. we achieve a moderate f1 of 57.4% on the development set and 46.0% on the test set via the self-training method. these results are comparable to state of the art results on similar tasks of identifying explicit discourse connective types from the penn discourse treebank. our findings from the manual labeling and the classification experiments indicate that the type of argumentative concessions we investigated is almost equally likely to be used in winning and losing arguments from the changemyview dataset. while this result seems to contradict theoretical assumptions, we provide some reasons for this discrepancy related to the changemyview subreddit. keywords: concessions, argumentation, subreddit, discourse relations, classification task 1. introduction a major challenge for argument mining—the automatic identification of argumentative structures within discourse—is the identification of linguistic features which characterize a winning argument. discourse moves that allow a speaker to achieve persuasion have been investigated since antiquity, being at the core of rhetoric, and are still a central concern in contemporary discourse studies and argumentation theory. the automatic identification of persuasive discourse is also receiving more c©2018 elena musi, debanjan ghosh, and smaranda muresan this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). musi, ghosh and muresan and more attention in computational linguistics. concessions have been unanimously deemed as strategies which increase persuasion (section 2). however, not every concession is part of an argument. we provide a definition of argumentative concessions in semantic and pragmatic terms. on this basis, we empirically test the theoretically-informed hypothesis that argumentative concessions work as persuasive strategies by calculating their distribution in persuasive vs. non persuasive discourse. an ideal dataset to carry out this empirical investigation is the changemyview subreddit platform, where multiple users negotiate opinions on a certain issue willing to change their point of view through other users’ arguments. when their point of view is changed, they award a ∆ point and back it up with a reason. thus, this platform provides us with a clear user intent — persuasion — and a clear signal when a message is perceived as persuasive (∆ point). even though it would have been convenient to observe the function of concessions in different datasets, the lack of explicit signals of persuasion prevents us to carry out such a comparison. we use the changemyview dataset released by tan et al. (2016). our task faces two challenges. first, in order for concessions to increase persuasion, they need to be part of an argument. however, not every concession is argumentative (grote et al., 1997): the pragmatic function of the sentence “although it is december, there is no snow” is not convincing the hearer about the lack of snow — which is an observable fact not subject to doubt — but simply that of expressing surprise for an unusual combination of events. this sentence could work in certain contexts as an argument (e.g., for the standpoint “global warming is worse and worse”) but does not semantically presuppose any controversy and, thus, the presence of argumentation. to address this challenge, we provide a semantically based methodology to identify argumentative concessions and compare them with other types of concessions (section 3). second, the automatic retrieval of argumentative concessions through discourse markers is difficult due to the polysemy with contrast and other discourse relations (prasad et al., 2014). state of the art discourse parsers trained on the penn discourse treebank (prasad et al., 2008) achieve low accuracy in the identification of concessive uses of discourse connectives in general due to the small number of instances in the training data; this result is reflected also on our task of identifying argumentative concessions. therefore, we first observe the distribution of concessions in manually labeled data both through expert annotation and through crowdsourcing (section 5). we focus our analysis on four discourse connectives: but, though, however, and while for two reasons: (1) they constitute 85% of the overall occurrences of potential markers of concessions in our dataset, and (2) they are highly polysemous and thus require disambiguation (section 4). we used 20% of the overall occurrences of these markers for our crowdsourcing study (i.e., 2,440 instances) as well as a separate dataset of 1,000 instances for the expert annotation task. second, as a step towards large-scale analysis of persuasive discourse, we use the manually labeled data as training, development and test sets to build computational models to detect argumentative concessions (section 6). we achieve a moderate f1 of 57.4% on the development set and 46.0% on the test set via a self-training method (section 7). our findings from the manual labeling (using both expert and crowdsourcing annotations) indicate that the type of argumentative concessions we investigate is almost equally likely to be used in winning (∆-awarded) and losing (non-∆) arguments. while this result seems to contradict theoretical assumptions, we provide some reasons related to the nature of the changemyview subreddit (section 8). in addition, we present a preliminary analysis on running our computational models on a different dataset, the yahoo news annotated comments corpus, (eric: engaging, respectful, and/or informative conversations) (napoles et al., 2017a), where comments are labeled as persuasive or not via crowdsourcing (persuasion is a “binary label indicating whether a comment contains 108 do concessions increase persuasion persuasive language or an intent to persuade”). differently from delta points in changemyview, the persuasiveness label does not inform us about what discourses achieved persuasion among the participants of an actual interaction, but provides hints as to what type of language is perceived as persuasive by third parties. we observe that argumentative concessions are present more in persuasive comments than in non-persuasive comments in the eric dataset. the dataset and code are freely available.1 2. related work concessions have received various non-overlapping definitions. we take as a default definition the one provided by grote et al. (1997), which focuses on the semantic relation holding between two connected propositions (a, b): “on the one hand, a holds, implying the expectation of c. on the other hand, b holds, which implies not c, contrary to the expectation induced by a”: a→ c e.g., “this dress is gorgeous”→ it is worth buying b → ¬c e.g., “but, it is expensive”→ it is not worth buying the conflict resides in the expectations generated by a and b which are mutually exclusive. besides this concessive configuration, which is called indirect concession (azar, 1997; izutsu, 2008), there are direct concessions where b and the following implication rule do not have to be verbalized. in these cases the main clause constitutes the negation of the expectation arising from proposition a (e.g., “this dress is gorgeous, but i am not buying it”). the distinction between direct and indirect concessions is particularly relevant from a discourse perspective since direct concessions are more suitable to be used nonargumentatively to describe states of affairs and to increase interest for an unexpected contradiction. the pragmatic function played by concessions has been widely addressed in linguistically oriented discourse studies. to cite just a few, drawing from quintilian institutio oratoria, perelman (1971) explains that through concessions, “one gives a favorable receipt to one’s opponent’s real or presumed argument. by restricting his claims, by giving up certain theses or arguments, a speaker can strengthen his position and make it easier to defend, while at the same time he exhibits his sense of fair play and his objectivity”. mann and thompson (1988) list concessions among presentational rhetorical relations, aimed at increasing the addressee’s positive attitude towards the speaker’s beliefs. according to pragmadialecticians (van eemeren et al., 2007) concessions are attested either in the confrontation stage of an argumentative discussion, where the difference of opinion between two parties is stated, and the opening stage, where the common starting points of the discussion are established. in a corpus study of interview transcripts involving environmental activists, uzelgun et al. (2015) show that the concessive construction “yes . . . but” constitutes a privileged viewpoint to investigate the (dis)agreement space singling out what is accepted and what is criticized. antaki and wetherell (1999) identify a series of rhetoric strategies through which concessions strenghten the speaker’s positions undermining counterarguments. similarly, couper-kuhlen and thompson (2000) exemplify how concessive repairs are used by speakers to back down their overstatements and foster credibility. in their investigation of myside bias in written argumentation, wolfe et al. (2009) show that texts which present and rebut other-side arguments achieve better ratings of agreement, quality 1. https://github.com/debanjanghosh/concessions 109 musi, ghosh and muresan and overall impression of the author. however, the presence of concessions preceding rebuttals did not lead to a higher perception of the arguments’ quality. it has to be remarked that students that were asked to evaluate the texts were not individually engaged in an interactive conversation; therefore, face-threatening risks were not the same as those characterizing a face to face argumentative exchange. the argumentative value played by concessive discourse relations is recognized by green (2010) who inserts concessions among the rst relations needed to represent argument presentation in a biomedical corpus. as far as datasets are concerned, while resources annotated as to agreement and disagreement are provided (walker et al., 2012), cases of partial (dis)agreement are neglected. computationally, somasundaran and wiebe (2009) develop an unsupervised method for stance detection in online debates, taking into consideration also concessionary opinions. from a pragmatic perspective, a growing interest is devoted to the identification of persuasive discourse strategies in order to implement classification experiments. young et al. (2011) released a corpus of blog posts annotated as to persuasive tactics according to sociological studies (e.g., promises/threats, mentions to duties/generalizations, appeal to reason). the evaluations of the predictive power of the strategies for the identification of persuasion show that the label reason guarantees accuracy; however the unigram svm baseline for reason is poor due to the difficulty in identifying rhetorical relations. drawing from the same set of tactics young et al. (2011) presented a corpus of 37 transcripts from four sets of hostage negotiation transcriptions annotated as to persuasion features. using supervised learning algorithms they show that persuasion tactics constitute machine-learnable features. tan et al. (2016) analyzed shallow linguistic and interactional features which happen to be persuasive in changemyview, a subreddit where users exchange opinions and assign a delta point to the user that managed to change their view: dissimilarity at the lexical level seems to play a major role, together with the order and the number of interactions. wei et al. (2016) investigated the performance of different sets of features in predicting persuasion: they have found that argumentation-based features perform better than shallow textual features. habernal and gurevych (2016) have recently released a corpus of 16k pairs of arguments over 32 topics annotated as to persuasiveness using crowdsourcing. annotators were also asked to provide reasons behind their choices. experiments with feature-rich svm (using libsvm tool; (chang and lin, 2011)) and long short-term memory (lstm) neural networks (hochreiter and schmidhuber, 1997) reveal that predicting persuasion is a task that requires analytic skills still hard to attain computationally. the detection of rhetorical/argumentative relations is bound to the performance of state of the art discourse parsers trained on the penn discourse treebank (pdtb), the largest annotated corpus so far available (prasad et al., 2008). pitler et al. (2009) showed that syntactic feautures play a crucial role in disambiguating explicit connectives. drawing from their work, lin et al. (2014) built a four-step pipeline for the identification of implicit and explicit discourse relations passing through the identification of text spans functioning as argument. connectives indicating concessions have been proven by swanson et al. (2015) to significantly correlate with the presence of argumentation, even though not with argument quality. for the task of detecting the type rather than the class of discourse connectives, which is more similar in nature to our problem, the performance of state of the art models is still modest (56.91 f-measure even for explicit cases) (biran and mckeown, 2015). biran and mckeown (2015) treat pdtb discourse parsing as two separate tagging tasks and show that connective-specific grammatical features promise to improve the results. on these grounds, we combine various lexical patterns, pragmatic as well as semantic features in conjunction 110 do concessions increase persuasion with a self-training approach for our task of identifying whether a discourse connective introduces an argumentative concession or not. 3. argumentative concessions in order to answer our research question — whether argumentative concessions increase persuasion — we need to provide a definition of argumentative concessions. drawing from musi (2017), we consider concessions as displaying an argumentative function when the proposition introduced by the connective — b —, which denies the expectations brought about by a preceding proposition, expresses the speaker’s standpoint. this happens when the following two conditions are met: (i) proposition b is asserted by the speaker who is committed to its truth at the moment of utterance; and (ii) proposition b is non factual — its truth is not self-evident. concessive relations such as “although it is already very warm, [there are no buds on the trees]b” and “he states that, despite the difficult situation, [the company will not go bankrupt]b” are, therefore, not argumentative since propositions b are an unassailable fact and a prediction with which the speaker does not necessarily agree, respectively. as far as proposition a is concerned, what is conceded can either be: (i) a point that could possibly be made by another speaker (i.e., “[sure, we could argue about some hypothetical other religion in the region causing similar problems]a, but that hypothetical world is not this world”) or that belongs to common ground knowledge (e.g., “[he’s no hitler, of course]a, but only because north korea isn’t powerful enough to start annexing its neighbors and has no substantial minority population to send to death camps”), or (ii) a claim previously made by another speaker in the discussion. concessions of type (i) can be used to prevent a counterargument. however, they do not necessarily display an interactional value, since the potential disagreement imagined by the speaker may have never happened. concessions of type (ii) are, instead, used in conversations as mitigating strategies to avoid disruptive disagreement. they, are, therefore, inherently argumentative. here are two examples. • speaker 1:[. . . ] basically, i’m glad jackson opted to make excellent **movies,** instead of attempting a shot for shot visualization of a **book,** and i think the extended cuts subvert his success in that regard. [. . . ] • speaker 2: as far as peter jackson’s opinion, the opinion of the author or producers of a content doesn’t dictate how it should be interpreted. he is a person who thinks a, that doesn’t mean that a is the be all and end all. like any piece of media once it’s released into the public it becomes its own beast that can have its interpretations judged by others. [i agree with you that the quality of the movie matters]a, but [if both versions are good movies and didn’t make unnecessary changes then i think they are different rather than one being better]b . as explained by couper-kuhlen and thompson (2000) this concession type is meant to be persuasive: a speaker, recognizing the validity of a point made by the hearer before expressing disagreement, avoids face-threatening acts and is perceived as reasonable by the hearer. we, thus, consider the description provided by interactional linguistics (couper-kuhlen and thompson, 2000, 2005) as our definition of argumentative concessions: • 1st move: speaker1 states something or makes some point • 2nd move: speaker2 acknowledges the validity of this statement or point (the conceding move) 111 musi, ghosh and muresan • 3rd move: speaker2 goes on to claim the validity of a potentially contrasting statement or point at a semantic level, proposition a (the conceding move of speaker2) is always an evaluative proposition since the speaker positively qualifies a standpoint advanced in the preceding post directly through the expression of positive sentiment (e.g., “i like your post, but [ . . . ]”) or agreement (e.g., “you are right, but [. . . ]”). at an informational level, the sentiment of the evaluation constitutes new information in the discourse flow, while what is evaluated constitutes old information which coincides with a point that has been previously made by another speaker. 4. data to empirically test our theoretically-informed hypothesis we use changemyview (cmv) subreddit: “dedicated to the civil discourse of opinions, and built around one simple idea: in order to resolve our differences, we must first understand them”. users start a discussion thread expressing their opinion (original post, op ) about a certain issue and other users challenge their view in a network of subsequent comments. changemyview constitutes a particularly suitable environment for the study of persuasive argumentation. first, the users, in order to get their posts published, have to respect a series of submission and comment rules, which ensures the presence of argumentation:2 • submission rules: (s1) try to explain the reasoning behind your view, not just what that view is (500+ characters required). (s2) you must personally hold the view and be open to it changing. (s3) submission titles must adequately sum up your view and include cmv: at the beginning. (s4) only post if you are willing to have a conversation with those who reply to you, and are available to start doing so within 3 hours of posting. • comment rules: (c1) direct responses to a cmv post must challenge at least one aspect of op s stated view (however minor), or ask a clarifying question. (c2) don’t be rude or hostile to other users. (c3) refrain from accusing op or anyone else of being unwilling to change their view. (c4) if you have acknowledged/hinted that your view has changed in some way, please award a delta (∆). (c5) comments must contribute meaningfully to the conversation. this combination of submission and comment rules matches the main rules for conducting an ideal critical discussion (van eemeren et al., 2013): rules (s1), (c2) and (c3) constitute an application of the freedom rule (“parties must not prevent each other from advancing standpoints or from casting doubt on standpoints”), while rule (s4) guarantees that the burden of proof rule (“a party that advances a standpoint is obliged to defend it if asked by the other party to do so”) is respected. 2. the following quotations are taken from the cmv wiki (https://www.reddit.com/r/changemyview/wiki/index). 112 do concessions increase persuasion rule (c1) can be interpreted as a sort of standpoint rule (“a party’s attack on a standpoint must relate to the standpoint that has indeed been advanced by the other party”). finally, rule (c4) presupposes the closure rule according to which “a failed defense of a standpoint must result in the party that put forward the standpoint retracting it”. in addition, changemyview constitutes a unique environment for the study of persuasion: although corpora annotated as to the presence of persuasive language exist (e.g., napoles et al. (2017b)), cmv, as far as we know, is the only available dataset containing first-hand information (through awarded ∆ points) as to what arguments are perceived as persuasive by actual language users. finally, since the discussed issues cover very different topics and the participants who are anonymous can have different epistemic backgrounds, the results of the analysis in terms of persuasion strategies can be generalized across different text genres and contexts. in our study, we used a dataset collected from the changemyview platform introduced by tan et al. (2016), where only the replies by the root challenger are considered, defining all the replies by the root challenger in a path as the rooted path-unit. for each rooted path-unit that wins a ∆, they select a rooted path-unit in the same discussion tree that did not win a ∆ but was the most “similar” in topic (similarity computed using jaccard similarity measure). with this setup, the goal is to de-emphasize what is being said, in favor of how it is expressed. we thus end up with a paired dataset that contains ∆-awarded and no-∆ comments, which can be used to test the hypothesis that argumentative concessions are persuasive strategies by computing their distribution in the ∆-awarded and no-∆ comments. we focus our analysis on four discourse markers: but, though, however, and while for two reasons: (1) they constitute 85% of the overall occurrences of potential markers of concessions in our dataset (see table 1), and (2) they are highly polysemous and thus require disambiguation. for each marker, we collect the sentence in which the marker is occurring as well as the previous and the next sentence (from both ∆ and no-∆ comments). next, we manually annotated two segments of the datasets using expert annotation and crowdsourcing, respectively. the first segment includes a set of 1,000 examples of the 4 discourse markers (e.g., but, though, however, and while) and has been annotated by an expert annotator.3the second segment consists of two samples each containing 10% of the overall occurrences of the four discourse markers from ∆ and no-∆ comments, which resulted in a total of 1,220 instances each (total of 2,440 instances). the annotators were asked to identify whether a marker introduces an argumentative concession (arg c) or not. one of the samples is used as development set (dev) and one as test set (test) in the machine learning experiments. the dev set is used for tuning the parameters of the learning model, while the test set constitutes the blind set used to evaluate the learning model. instead of expert annotators, we use a crowdsourcing platform – amazon mechanical turk (mturk) to identify the arg c. the task is framed as follows: given a sentence or pairs of two adjacent sentences in changemyview containing one of the four discourse markers, five turkers on mturk are asked to identify and label as arg c those occurrences in which the sentence preceding the connective expresses agreement or positive sentiment towards a point previously made by another speaker. the turkers are provided with detailed instructions of the task and multiple examples. before conducting the full crowdsourcing annotations tasks, we ran a pilot annotation to evaluate whether turkers are able to complete such task. in the pilot annotation round, we realized that providing the turkers with the mere definition of argumentative concessions was not enough since annotators were selecting as argumentative concessions also those conceding statements be3. after annotation we noticed that 20 examples were duplicates, and thus we used the remaining set of 980 examples in the experiments. 113 musi, ghosh and muresan marker ∆ no-∆ admit 26 17 albeit 9 17 although 78 93 but 4403 5908 concede 8 13 despite 89 114 even if 255 314 even though 101 129 even when 31 55 however 132 213 in spite of 10 8 nevertheless 3 10 notwithstanding 1 4 non the less nonetheless 7 18 the fact remains that 3 4 though 426 619 whereas 48 73 while 575 763 total 6205 8372 table 1: distribution of candidate concessions markers in the dataset collected from changemyview longing to common knowledge. therefore, we refined the guidelines to discard this latter type of concessions pointing to the fact that they tend to contain modal adverbs of certainty (e.g., (for) sure, of course) that, differently from markers that imply a dialogical exchange (e.g., agree, understand), signal that the truth of a general statement is taken for granted since it belongs to the common ground (e.g., “of course, runners use the lower body more than the upper body, but the same is true for swimmers in reverse (they use more upper than lower)”). we, moreover, specified that they contain second person pronouns/adjectives that do not work as deictics, but are used impersonally to portray fictive scenarios (i.e., “yes, a portion of your money goes to things you don’t support; but likewise, things you do support are partially funded by other people who may not support those things.”). for both dev and test sets, containing 1,220 sentences (or pairs of sentences) each, we obtain 6,100 labels from the turkers. to assure a level of quality control, only qualified turkers were allowed to perform the task (i.e., more than 95% approval rate and at least 5,000 approved human intellingence tasks — hits). each hit contained one sentence (or pairs of sentences) to be labeled and the turkers were paid 5 cents for each hit. we compute the inter-annotator agreement (iaa) between the turkers via fleiss kappa measure (fleiss, 1971). we obtain fair agreement both for dev (κ = 0.31) and for test (κ = 0.22) sets. as underlined by passonneau and carpenter (2014), standard measures for inter-annotator reliability are not suitable to account for corpus quality, and 114 do concessions increase persuasion in addition, the κ values obtained on semantic annotation tasks is not high. therefore, we consider the obtained iaa as a reasonable output which does not undermine the relevance of the corpus. 5. distribution of argumentative concessions in manually labeled data the expert annotated dataset contains 229 instances of arg c and 751 instances of other types of concessions or instances expressing contrast or other discourse relations, which we denote as other. for the crowdsourcing experiment, we take majority voting among the five turkers to choose the label (arg c or other). we obtained 201 arg c instances (out of 1,220) for the dev set and 174 arg c instances (out of 1,220) for the test set. to better understand the difficulty of the task, we selected the cases where 3 annotators chose one label, while 2 others choose the other label and we asked an expert annotator to label these cases. we compared the expert labels with the labels obtained by majority voting (i.e., the labeled chosen by 3 turkers). in the test set, out of 225 such instances, in 22 cases there is a mismatch between the label annotated by three turkers and the expert annotator; in the test set, out of 280 such instances, the mismatch amounts to 67 cases. zooming into the mismatched examples, in the dev set 4/22 cases are not recognized as argumentative concessions by the turkers, while 18 cases are misclassified as argumentative concessions. in the test set, the type of mismatch seems more balanced: 31/67 instances are not recognized as argumentative concessions by turkers, while 36/67 are misclassified as argumentative concessions. we observe that discarding these most confusing cases where majority is formed by only 3 turkers, the iaa improves both in the dev set (κ = 0.43) and in the test set (κ = 0.35). taking into account the annotation provided by the expert for those instances, the number of arg c in dev and test sets is 179 and 168, respectively. from the qualitative analysis, it seems that argumentative concessions are misclassified when the proposition functioning as standpoint (proposition b) expresses a negative evaluation of the opinion held by another speaker, while the proposition a is used to specify the degree of such an attack, more than to express a partial agreement. in a couple of sentences such as “i don’t believe the war was worth the suffering, but to say that nothing positive happened because of it is a little disingenuous”, the speaker’s intention is not that of partially agreeing with the judgment made by the preceding speaker, but to clarify to what extent it is indeed ingenuous. cases not recognized as argumentative concessions contain a positive sentiment rather than an explicit agreement with what was said by the preceding speaker (e.g., “you have noble goals, but there are very real downsides”). the distribution of argumentative concessions in threads that have and have not been awarded a ∆ point does not appear to be significantly skewed amounting to 99 (awarded) vs 130 (not awarded) occurrences in the set annotated by an expert and 165 vs. 194 in the crowdsourcing experiment. due to the limited quantity of annotated data available, before discussing the relevance of the attested distribution (section 8), we develop a preliminary computational model to classify and trace back argumentative concessions on a larger scale. 6. computational models to detect argumentative concessions we seek to automatically identify argumentative concessions from the cmv corpus. we frame the task as a binary classification task: arg c vs. other. since using expert annotators is an expensive process, we use the expert annotated data as a small labeled training data (total of 980 instances; section 5). due to this small annotated data scenario, we develop a self-training method (clark et al., 2003; mihalcea, 2004), which uses the remaining unannotated 70% of the cmv corpus (section 6.2) 115 musi, ghosh and muresan as unannoted data. from our annotation studies it is clear that all the datasets are highly unbalanced (the size of the arg c class is much smaller than the size of the other class). as stated earlier in the previous section, the data annotated through crowdsourcing (the two samples of 10% of the cmv corpus each) is used as development (dev) and test (test) sets, respectively. in the following section we discuss the features as well as the linguistic patterns associated with the argumentative concessions used in our experiments (section 6.1). second, we discuss the classification task using self training in section 6.2 and the results in section 7, including comparison with an off-the-shelf state-of-the art discourse parser trained on the penn discourse treebank that aims to classify the type of discourse connectives, not their class (biran and mckeown, 2015). 6.1 feature description we use linguistically-motivated features inspired by research on discourse relation identification and persuasive argument identification (stab and gurevych, 2014; ghosh et al., 2016). a brief description of the features is given below. • bag-of-words: we selected the entire cmv dataset used by tan et al. (2016) to extract the bagof-words features (e.g., unigrams and bigrams). we consider every sentence that contains the four candidate discourse markers (e.g., “but”, “while”, “however”, and “though”) as candidate examples. if the sentence starts with the markers we also consider the previous sentence. next, based on tf-idf scores (i.e., we treated each sentence as a document), we select the top 1,000 unigrams and bigrams as candidate features. • personal pronouns and adjectives: by definition, argumentative concessions dialogically point to the stance taken by the previous speaker. they, therefore, contain personal pronouns and adjectives (e.g., “i see your point”, etc.): we consider as features both a list of first person (e.g., “i”,“me”, “my”, “mine”) and second person (e.g., “you”, “your”, “you’ re”) pronouns and adjectives. • modal verbs: modal verbs (e.g., “could”, “should”) work as indicators of claims since they indicate that what is expressed in a proposition is not unassailable, but might be otherwise (palau and moens, 2009). they, thus, frequently appear in propositions b of argumentative concessions which constitute the speakers’ standpoints. we define a boolean feature which indicates if a candidate example contains a modal verb. • hedges: hedges are linguistic devices used to mitigate the speaker’s commitment to the truth of a proposition (hyland, 1996a), i.e., “i tend to accept”. they include possibility modals next to other linguistic items expressing the degree of speaker’s certainty. they can be ambiguous depending on constructional features: the propositional attitude indicator i think works, for instance, as a hedge in parenthetical constructions (e.g., “it is not worth buying it, i think”), while it merely signals subjectivity when used as a main verb (e.g., “i think it is not worth buying it”). the use of hedges is common in argumentative concessions since they contribute to avoid a potentially face-threatening act of abrupt disagreement. tan et al. (2016) argue that depending on the context, hedges can make an argument weaker or easier to accept by softening its tone. based on their research and also on hanauer et al. (2012), we collect a set 116 do concessions increase persuasion of candidate hedge cues and use them as boolean features (presence or absence of a hedge word). • jaccard similarity: in argumentative concessions proposition a always expresses positive sentiment towards a claim expressed by the previous speaker. thus, we use jaccard similarity to measure lexical similarity between the sentence/sentences containing the candidate concession and sentences from the original post op (we removed stopwords). we use the maximum similarity value as a feature. • sentiment feature: by definition, argumentative posts tend to contain opinion on the other posts. a post that is tagged with subjectivity will be a useful feature to identify concessions. we use (a) the mpqa lexicon (wilson et al., 2005) of over 8,000 positive, negative, and neutral sentiment words, (b) an opinion lexicon with around 6,800 positive and negative sentiment words (hu and liu, 2004) to see whether training instances contain sentiment words. apart from the above features, we also retrieved lexical patterns that could be indicators of argumentative concessive uses of the discourse markers. these patterns are used in proposition a to express a positive evaluation about another speaker’s claim. they could assist in achieving higher precision in identifying arg c instances. below is a short description of the semi-automatic retrieval of lexical patterns. semi-automatic retrieval of lexical patterns we present a bootstrapping algorithm that automatically learns lexical patterns expressing argumentative concessions. the algorithm begins with only two seed phrases – “i agree” and “you are right”– the most common two patterns expressing argumentative concessions according to annotation results done by an expert on a separate development set of one-hundred utterances (not used in the above annotation studies). we used 80% of our changemyview data (i.e., all except the dev as well as the test dataset) for the bootstrapping algorithm. the algorithm operates on a simple structural assumption: in argumentative concessions, words forming the proposition positioned before the marker (for “but” and “though”) or in the marker’s scope (for “while”) express agreement with a point made in a preceding comment. to identify such patterns, first, we retrieve these words (i.e., all words in the proposition before “but”). second, we extract all possible trigrams, four-grams, and five-grams from this set of words and search for the two seed phrases in the ngrams. for instance, we identify all instances of “[. . .] i [. . .] agree [. . .]”, where [. . .] represent zero or more occurrences of words. we obtain patterns such as “i agree [completely]” or “[i think] you are right [about]”. third, using each new pattern we attempt to identify new lexical constructions. for example, given the pattern “i agree [completely]” we search for the occurrences of “i [. . .] completely” where [. . .] represent zero or more occurrences of any words, excluding negation. with this search, we arrive at new patterns, such as “i [understand] completely” and “[i think] you are [correct]”. since this method relies on structural syntactic similarity, we embed the following semantic rules: (i) keep only patterns, which contain propositional attitude indicators (i.e., verbs “think”, “realize”, and the constructions “be right”, “be correct”) or indicators of sentiment (i.e., verbs “love”, “like”) and (ii) select the patterns which contain the pronoun “you” or the adjective “your”– to retain just those seeds where the target of agreement is an opinion held by another poster. finally, in the fourth stage, using the new lexicon we again resume the search for new patterns (i.e, using “[. . .] i [. . .] realize [. . .]”) and this process continues until we do not find any new patterns. as a result of the bootstrapping algorithm we obtain 329 lexical patterns, which we call b lexicon. we, then, 117 musi, ghosh and muresan manually filtered lexicon i would agree with you i fully agree that i see what you i see where you i think you are correct table 2: examples of manually filtered lexicon manually filter the bootstrapped lexicon to eliminate redundancies, merging patterns which instantiate the same linguistic constructions and adjustments based on the dev set (e.g., removing double negatives “i don’t disagree”). as a result of manual filtering we finally obtained 116 lexical patterns indicating argumentative concessions (b lexiconmf ). table 2 shows some examples from this manually filtered lexicon. 6.2 self-training method for identifying concessions self-training algorithms (mihalcea, 2004; clark et al., 2003) start with a small subset of annotated training data and attempt to increase the amount of training data by using a large set of unannotated data. we adopt the approach of mihalcea (2004) that used self-training as “a tagger that is retrained on its own labeled cache on each round” (mihalcea, 2004; clark et al., 2003). similar to mihalcea (2004), we start with a small set of labeled data that is the training data (l), in our case the data labeled by the expert annotator consisting of 980 instances. we build a classifier using the linguistically-motivated features described above, and then apply the learned model on a set of unlabeled data (u ). for our experiments, the u is the 70% of the cmv data (i.e., apart from the training, test, and dev sets; each is 10% of the cmv data). now, instead of classifying directly on the total set of data u , we split u in p random pools where each pool contains u ′ unlabeled instances. from each p pool we select only those g c argumentative concession instances and g nc non-argumentative concession instances with a labeling confidence exceeding a particular threshold. similar to mihalcea (2004), while adding the new instances to the training data l, we maintain the original class distribution of training data between arg c and other categories. the threshold could be a preset value of probability; in this experiment, we varied the number of data instances to select while keeping the original class distribution. these instances (i.e, g c and g nc) in turn are added to the original training set l. the classifier is retrained with the new set of training data (i.e, l + g c + g nc) and this process continues for all the p pools. note, in each step we also evaluate the classifier on the dev set to assess the quality of the new data that is added to the training set l (table 3). for details of the self-training procedure, please see mihalcea (2004); clark et al. (2003). we use the support vector machines (svm) classifier with rbf kernel (we use scikit-learn tool (pedregosa et al., 2011)). the class weights are inversely proportional to the number of instances in the categories. as a final classifier we used a system combination: if any dev/test instance contains a lexical pattern obtained via bootstrapping described in the previous section, we classify that instance as arg c, otherwise, use the decision of the self-training classifier. the dev data is used to fine tune all parameters in our experiment and to choose the best parameters (see next section for detailed discussion). 118 do concessions increase persuasion self-training setting optimal size (training) performance (max. f1) pool size g c arg c other p r f1 50 10 430 952 56.9 57.5 57.2 100 10 441 821 56.7 57.0 56.8 1000 10 289 811 65.9 45.8 54.0 2000 10 229 751 65.3 44.6 53.0 100 50 442 964 64.5 51.3 57.4 500 50 424 946 63.9 48.0 54.8 1000 50 415 937 64.1 47.5 54.5 2000 50 279 801 63.1 46.3 53.4 table 3: experimental results of the self-training method on the dev set (bold are best scores) 7. experiments and results before reporting the results of our self-training classifier and our baselines, we report the results of an off-the-shelf parser that aims to label the types of discourse connectives (biran and mckeown, 2015), noted asobparser in table 4. biran and mckeown (2015) have trained the parser on the penn discourse treebank (pdtb) corpus (miltsakaki et al., 2004) using discourse connective features, lexical features, syntactic features, etc. their model can identify the types of discourse relations, such as concession, contrast, pragmatic concession, and pragmatic contrast, which are types of the class of discourse relation comparison. the off-the-shelf parser results in a low f1 measure of only 6 with precision of 13.2 and recall of 4 for the dev data.4 this low performance is not unexpected. first, the parser is trained on the pdtb corpus, which is based on of wall street journal (wsj) articles and the language is vastly different from the changemyview subreddit. in addition, pdtb has a small number of concessions in general and it is not clear how many of those are argumentative concessions, if any. it has to be noted that state of the art discourse parsers do not take into account different pragmatic values underlying a discourse relation. the level of granularity required to classify different types of concessions is higher than that necessary to classify what discourse relation is conveyed by a single connective. nevertheless, the performance achieved by our system is comparable to that achieved by biran and mckeown (2015) in the classification of explicit discourse relations (56.91 f1). baselines. we use two baselines: 1) a rule-based system based on the lexical patterns learned via bootstrapping (b lexiconmf in table 4), and 2) the system combination that uses theb lexiconmf and a svm classifier that uses all the features but without the self-training process (svmnost ). the performance of both of these baselines is better on the dev set than on the test set, which is expected as dev set was used to check the final pattern lexicon. the recall of the b lexiconmf patterns is low for the test data, meaning that there are different patterns not covered in our collection of lexical patterns and argumentative concessions can exist in many ways that are not retrieved by simple lexical rules. self-training. table 3 show the results of the system combination (lexical patterns and svm classifier using self-training) on the dev set, used to decide the best parameters (pool size and the 4. since the performance of the off-shelf-parser on the dev data is very low compared to the other methods we did not conduct any further experiment on the test set. 119 musi, ghosh and muresan computational model training size dev test (arg c;other) p r f1 p r f1 obparser 13.2 4.0 6.0 svmnost (229;751) 65.3 44.6 52.9 35.1 58.9 44.0 b lexiconmf 65.5 43.9 52.5 48.3 25.0 32.9 self-training (best) 64.5 51.3 57.4 38.0 58.3 46.0 table 4: experimental results of classifiers on the dev and test set. the self-training results report the best (in bold) performing parameters. number of instances to add to the training set). in column one (i.e., pool size) we report the number of random unlabeled instances (i.e, u ′) that were tested via the classifier. column two represents the maximum number of instances (g c) that are added to the training data. for recall, we also add g nc to maintain same class distribution between arg c and other categories. the next two columns show the size of training data from our two classes respectively that achieve the highest f1 scores for the dev set. we start the experiments with the labeled training data by expert annotators (also used in the baselines) and gradually add g c and g nc that are predicted by the classifier. finally, the last three columns show the p/r/f1 on the dev set. we report all the columns based on the highest f1 achieved by the models. for instance, in the first row of table 3, the pool size is 50 and g c is 10. the size of the unlabeled data is close to 8,400 (after removing duplicates), which means here the number of pools is 168 (8400/50). now, we classify each pool from the set of p , and depending upon the classification result, we add g c (here, maximum is 10 instances) and g nc accordingly from each pool. after the classifier has evaluated a certain number of pools the f1 of 57.2 is achieved. meanwhile, starting from 229 arg c, the size of the arg c is now 430. we also observe a common trend between the various runs of the self-training procedure; the accuracy increases till a certain number of g c + g nc is added to the original training set of l, but after that the performance drops probably due to added noise.the best model based on the dev set, has the pool size of 100 and g c of 50 reaching f1 of 57.4 that is close to a 5% improvement compared to the baseline results without self-training on the dev set (see svmnost from table 4). we used this best setting (i.e., pool size of 100 and g c of 50) to run our system combination using self-training on the test set. we also used the dev data for feature selection. we employ χ2 based feature selection and observe that the “top 300” features (based on χ2 scores) performed best for the dev data. subsequently this setting (i.e, “top 300” features) is used for the test set. this feature set of ”top 300” features contains modal verbs such as “may”, “could”; jaccard similarity value; pronouns such as “i”, “your’; hedges such as “almost”, “probably”, “somewhat”; words such as “greatest”, “excitement”, “interesting” from the sentiment lexicons; unigrams such as, “agree”, “recognize”, “argument”, “think”, and finally, bigrams such as, “while you”, “argument is”, “i absolutely”, to name a few. the results of self-training show improvement over the baselines, but to a lesser degree when compared to svmnost (2%), and a higher degree when compared to b lexiconmf (13.1%) (table 4). to understand the quantitative results better, we have randomly selected and qualitatively analyzed a sample of 135 classified occurrences from the development data. we compared recurrent characteristics both of occurrences that the system classified arg c in accord with gold data and those that it classified as arg c, while the human annotators did not. it turns out that occurrences 120 do concessions increase persuasion classified as arg c both by the system and by the annotators tend to include lexical patterns, which unambiguously express agreement paired with second person pronouns/adjectives which anaphorically point to a previous claim (i.e. “i do agree with a lot of points in your post, but there is a huge disconnect between the title of your post”). looking at occurrences that have been classified by the expert annotator as arg c while considered by the system as other, they tend to lack an unambiguous expression of agreement and an anaphoric reference to a preceding post. more specifically, agreement is expressed through modal adverbs expressing different degrees of certainty (i.e. “possibly, but not because they’re missing out on experiences”; “of course it is but if you’re okay with the humane treatment of animals to include population control why are you uncomfortable with the humane treatment of animals in other contexts as the end result is the same?”). however, the state of affairs over which the modal adverbs have scope is elliptical, since it is shared as given information by the participants to the discussions (speakers as well as readers). in classification experiments ellipsis is a hard to detect phenomenon since it is dependent on the pragmatic content of the utterance. 8. discussion: are concessions persuasive strategies? to test the theoretically-informed hypothesis that argumentative concessions work as persuasive strategies, we look at their distribution in the ∆ awarded vs. no-∆ awarded comments. first, we consider the manually labeled data. the small set of training data annotated by the expert annotator, as well as the dev and test sets annotated in the crowdsourcing experiment. the distribution of argumentative concessions in the ∆ awarded vs. no-∆ awarded comments on all these datasets is given in table 5. second, we look at the distribution when considering the argumentative concessions predicted by our best classifier (system combination with self training). we look at the dev and test set, as well as the rest of the unlabeled cmv corpus (table 6). as a caveat, we have to point out that the due to the low accuracy of classifier, the observed distribution based on the classifier predictions cannot be considered reliable. for example, the predictions of the computational model on the dev and test set miss to identify more than 1/3 of argumentative concessions compared to the gold annotations. marker ∆-training no-∆-training ∆-dev no-∆-dev ∆-test no-∆-test but 39 59 68 83 83 82 however 25 3 3 5 though 22 27 4 8 2 1 while 13 13 2 5 total 99 130 78 101 85 83 table 5: distribution of argumentative concessions in ∆ and no-∆ comments in the manual labeled datasets overall the numbers suggest a fairly equal distribution of argumentative concessions in the ∆ and no-∆ comments both in the manually labeled data and in those predicted by the classifier. we have tested the statistical significance of the distribution of argumentative concessions on the manually labeled sets using χ2 test: the results are not significant at p < 0.05 on training and test 121 musi, ghosh and muresan marker ∆-dev no-∆-dev ∆-test no-∆-test ∆-unlabel no-∆-unlabel but 49 57 51 47 1100 793 however 2 4 though 4 5 1 8 2 while 1 2 total 56 68 52 48 1112 795 table 6: distribution of predicted argumentative concessions in ∆ and no-∆ comments sets, while they are significant on dev set, where argumentative concessions are less frequent in winning arguments. these results suggest that argumentative concessions do not increase persuasion in changemyview, challenging the assumptions made in the rhetorical literature. however, they do not allow us to draw conclusions scalable to different contexts about the persuasive role played by this type of concessions. they rather provide further confirmation that the persuasive value played by lexical meta-discursive features is context-bound and crucially depends on the rhetorical situation. for example, hedges seem to increase persuasiveness in scientific writing (hyland, 1996b), while they decrease it in other kinds of messages (blankenship and holtgraves, 2005). when it comes to argumentative concessions, they embody what in the philosophical tradition has been called principle of charity or principle of rational accommodation according to which “we make maximum sense of the words and thoughts of others when we interpret in a way that optimizes agreement” (davidson, 2001): conceding the claims made by the author of the original post the speaker shows his intention of interpreting them in the best possible way. in doing so he reinforces his ethos, presenting himself as a reasonable discussant. at the same time, he minimizes the risk of a face-threatening opposition which could arise from the apparent incompatibility with the statement expressed in the nucleus (proposition b). the principle of charity has, in fact, to be understood as a methodological presumption that guarantees the understanding of another point of view in its argumentatively strongest form to allow a possibly adequate critique and reach agreement through persuasion. this principle is structural in changemyview, being at the core of the subreddit mission of resolving differences of opinions starting from a deep understanding of them. users who choose to write on this subreddit have to respect the submission rules and are, thus, by default charitable. therefore, linguistic strategies such as argumentative concessions that encode the principle of charity constitute persuasive strategies in the speakers’ mind, but are plausibly perceived as routinized expressions by the addressees. in other words, they do not shape the argumentative profile of single users, but are conventional strategies belonging to the activity type envisioned by changemyview. as stated in the introduction, the lack of corpora with explicit signals as to what pieces of discourse achieved persuasion makes it difficult to investigate the persuasive roles played by concessions in datasets belonging to different discourse genres and dialogue activity types. to our knowledge, beside the changemyview subreddit, the only other corpus labeled with persuasive features is the yahoo news annotated comments corpus (ynacc) (napoles et al., 2017a). this corpus contains around 140k threads, taken from the comments sections of yahoo news articles as well threads from the internet argument corpus. discussion posts are annotated by expert and untrained (i.e., turkers) annotators with specific labels, both at the thread (e.g., constructiveness) and at the comment level (e.g., topic). among the latter type, persuasiveness is defined as a “a 122 do concessions increase persuasion binary label indicating whether a comment contains persuasive language or an intent to persuade.” differently from delta points in changemyview, the persuasiveness label does not inform us about what discourses achieved persuasion among the participants of an actual interaction, but provides hints as to what type of language is perceived as persuasive by third parties. in other words, we experiment with this dataset to assess whether argumentative concessions are perceived by language users as persuasive strategies, namely strategies used by speakers as an attempt to persuade, despite the actual pragmatic outcome. we selected a subset of ynacc that are annotated by expert annotators. this subset contains 4,719 posts that are annotated with the label “persuasive” and 17,616 posts that are annotated with the label “not persuasive”. from this subset we evaluate only the posts that contain the selected markers for our experiments, “but”, “while”, “though”, and “however”. we use the same training data (described in section 6) to predict the binary label arg c and other from the ynacc posts. we use the same features that are described in section 6.1 except the jaccard similarity feature since the ynacc dataset does not indicate which post is a reply to another post. after removing the duplicates from the ynacc posts we observe that out of 649 persuasive posts, 321 (49.4%) are classified as arg c (i.e., concessions) whereas out of 1,263 non persuasive posts, 417 (33%) posts are classified as concessions. according to these results, argumentative concessions are deemed as persuasive strategies, since they correlate with the persuasiveness label. in table 7 we present the count of the main expressions of concessions across the persuasive and the non persuasive datasets: “?” depicts one or zero occurrence of the particular pattern. for example, “pattern acknowledge” shows that any occurrence of the expression, “i” and following either of “concede”, “acknowledge”, and “think” is regarded as a concession of the pattern “pattern acknowledge”. note in the above case, “also” or “too” can appear between the above two words. other patterns of concessions also exist (e.g., “i appreciate your”) but the frequency is very explicit expressions (patterns) description count persuasive not persuasive pattern yes [yes|sure|of course|correct|right|true] [,] 24.81 13.62 pattern acknowledge [i] ?[also|too] [concede|acknowledge|think] 12.63 9.58 pattern see i ?[adverb|modal] [see|get] [?] [you|your] 8.47 4.51 pattern agree i ?[adverb|modal] [agree] 0.77 0.63 table 7: distribution (in %) of explicit expressions (patterns) of concessions in persuasive and not persuasive instances in the eric corpus low. overall, it seems that regardless the type of linguistic construction at work, concessions are more frequent in persuasive threads. 123 musi, ghosh and muresan 9. conclusion and future work we tackled the task of empirically validating the theoretically assumed persuasive role played by specific discourse relations, concessions, using the changemyview subreddit platform. drawing from a linguistically-informed typology of concessions, we singled out one type of concessions that prototypically bears an argumentative value. we focused on four discourse markers that constitute 85% of the overall occurrences of potential markers in our data: but, though, however and while. we present a computational model based on self-training using linguistically motivated features. we achieve a moderate f1 of 57.4% via the self-training method on the development set and 46.0% on the test set. our findings both from the manual labeling (both expert and crowdsourcing annotation) and the system predictions indicate that the type of argumentative concessions we investigate is almost equally likely to be used in winning (∆-awarded) and losing (no-∆) arguments. while this result seems to contradict theoretical assumptions, we provided some reasons related to the changemyview subreddit. this behavior shows that text-genre rules have to be taken into account in the interpretation of rhetorical patterns: when perceived as conventional genre-specific rules, persuasive strategies may happen not to be pragmatically effective. in future work, we plan to validate this explanation by the following three steps: 1) improving the performance of our computational models by collecting a larger dataset as training data to be able to run the distribution over a much larger dataset; 2) looking at other types of argumentative concessions; and 3) retrieving argumentative concessions in different text-genres. acknowledgement this paper is based on work supported partially by the advanced post doc snfs grant for the project from semantics to argumentation mining: a context-independent lexicon of indicators of argumentative discourse relations and darpa-deft program. the views expressed are those of the authors and do not reflect the official policy or position of the snfs, department of defense or the u.s. government. we would like to thank the annotators for their work and the anonymous reviewers for their valuable feedback. references charles antaki and margaret wetherell. show concessions. discourse studies, 1(1):7–27, 1999. moshe azar. concession relations as argumentation. text-interdisciplinary journal for the study of discourse, 17(3):301–316, 1997. or biran and kathleen mckeown. pdtb discourse parsing as a tagging task: the two taggers approach. in proceedings of the 16th annual meeting of the special interest group on discourse and dialogue, pages 96–104, 2015. kevin l blankenship and thomas holtgraves. the role of different markers of linguistic powerlessness in persuasion. journal of language and social psychology, 24(1):3–24, 2005. chih-chung chang and chih-jen lin. libsvm: a library for support vector machines. acm transactions on intelligent systems and technology (tist), 2(3):27, 2011. 124 do concessions increase persuasion stephen clark, james r curran, and miles osborne. bootstrapping pos taggers using unlabelled data. in proceedings of the seventh conference on natural language learning at hlt-naacl 2003-volume 4, pages 49–55. association for computational linguistics, 2003. elizabeth couper-kuhlen and sandra a thompson. concessive patterns in conversation. topics in english linguistics, 33:381–410, 2000. elizabeth couper-kuhlen and sandra a thompson. a linguistic practice for retracting. in hakulinen auli and margret selting, editors, syntax and lexis in conversation: studies on the use of linguistic resources in talk-in-interaction, volume 17, pages 257–288. 2005. donald davidson. inquiries into truth and interpretation: philosophical essays, volume 2. oxford university press, 2001. joseph l fleiss. measuring nominal scale agreement among many raters. psychological bulletin, 76(5):378, 1971. debanjan ghosh, aquila khanam, yubo han, and smaranda muresan. coarse-grained argumentation features for scoring persuasive essays. in proceedings of the 54th annual meeting of the association for computational linguistics, pages 549–554, 2016. nancy l green. representation of argumentation in text with rhetorical structure theory. argumentation, 24(2):181–196, 2010. brigitte grote, nils lenke, and manfed stede. ma(r)king concessions in english and german. discourse processes, 24(1):87–117, 1997. ivan habernal and iryna gurevych. which argument is more convincing? analyzing and predicting convincingness of web arguments using bidirectional lstm. in proceedings of the 54th annual meeting of the association for computational linguistics (acl), pages 1589–1599, 2016. david a hanauer, yang liu, qiaozhu mei, frank j manion, ulysses j balis, and kai zheng. hedging their mets: the use of uncertainty terms in clinical documents and its potential implications when sharing the documents with patients. in amia annual symposium, page 321, 2012. sepp hochreiter and jürgen schmidhuber. long short-term memory. neural computation, 9(8): 1735–1780, 1997. minqing hu and bing liu. mining and summarizing customer reviews. in proceedings of the tenth acm sigkdd international conference on knowledge discovery and data mining, pages 168– 177. acm, 2004. ken hyland. talking to the academy: forms of hedging in science research articles. written communication, 13(2):251–281, 1996a. ken hyland. writing without conviction? hedging in science research articles. applied linguistics, 17(4):433–454, 1996b. mitsuko narita izutsu. contrast, concessive, and corrective: toward a comprehensive study of opposition relations. journal of pragmatics, 40(4):646–675, 2008. 125 musi, ghosh and muresan ziheng lin, hwee tou ng, and min-yen kan. a pdtb-styled end-to-end discourse parser. natural language engineering, 20(02):151–184, 2014. william c mann and sandra a thompson. rhetorical structure theory: toward a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8(3):243–281, 1988. rada mihalcea. co-training and self-training for word sense disambiguation. in conll, pages 33–40, 2004. eleni miltsakaki, rashmi prasad, aravind k joshi, and bonnie l webber. the penn discourse treebank. in lrec, pages 2237–2240, 2004. elena musi. how did you change my view? a corpus-based study of concessions argumentative role. discourse studies, pages 1–19, 2017. courtney napoles, aasish pappu, and joel r tetreault. automatically identifying good conversations online (yes, they do exist!). in icwsm, pages 628–631, 2017a. courtney napoles, joel tetreault, aasish pappu, enrica rosato, and brian provenzale. finding good conversations online: the yahoo news annotated comments corpus. in proceedings of the 11th linguistic annotation workshop, pages 13–23, 2017b. raquel mochales palau and marie-francine moens. argumentation mining: the detection, classification and structure of arguments in text. in proceedings of the 12th international conference on artificial intelligence and law, pages 98–107. acm, 2009. rebecca j passonneau and bob carpenter. the benefits of a model of annotation. transactions of the association for computational linguistics, 2:311–326, 2014. f. pedregosa, g. varoquaux, a. gramfort, v. michel, b. thirion, o. grisel, m. blondel, p. prettenhofer, r. weiss, v. dubourg, j. vanderplas, a. passos, d. cournapeau, m. brucher, m. perrot, and e. duchesnay. scikit-learn: machine learning in python. journal of machine learning research, 12:2825–2830, 2011. chaim perelman. the new rhetoric. in hillel yehoshuabar, editor, pragmatics of natural languages, pages 145–149. springer, 1971. emily pitler, annie louis, and ani nenkova. automatic sense prediction for implicit discourse relations in text. in proceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural language processing of the afnlp: volume 2, pages 683–691. association for computational linguistics, 2009. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of lrec, 2008. rashmi prasad, bonnie webber, and aravind joshi. reflections on the penn discourse treebank, comparable corpora, and complementary annotation. computational linguistics, 40(4):921–950, 2014. 126 do concessions increase persuasion swapna somasundaran and janyce wiebe. recognizing stances in online debates. in proceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural language processing of the afnlp: volume 1, pages 226–234. association for computational linguistics, 2009. christian stab and iryna gurevych. annotating argument components and relations in persuasive essays. in proceedings of the 25th international conference on computational linguistics (coling 2014), pages 1501–1510, 2014. reid swanson, brian ecker, and marilyn a walker. argument mining: extracting arguments from online dialogue. in sigdial conference, pages 217–226, 2015. chenhao tan, vlad niculae, cristian danescu-niculescu-mizil, and lillian lee. winning arguments: interaction dynamics and persuasion strategies in good-faith online discussions. in proceedings of the 25th international conference on world wide web, pages 613–624. international world wide web conferences steering committee, 2016. mehmet ali uzelgun, dima mohammed, marcin lewiński, and paula castro. managing disagreement through yes, but constructions: an argumentative analysis. discourse studies, 17(4):467– 484, 2015. frans h van eemeren, peter houtlosser, and af snoeck henkemans. argumentative indicators in discourse: a pragma-dialectical study, volume 12. springer science & business media, 2007. frans h van eemeren, rob grootendorst, ralph h johnson, christian plantin, and charles a willard. fundamentals of argumentation theory: a handbook of historical backgrounds and contemporary developments. routledge, 2013. marilyn a walker, jean e fox tree, pranav anand, rob abbott, and joseph king. a corpus for research on deliberation and debate. in lrec, pages 812–817, 2012. zhongyu wei, yang liu, and yi li. is this post persuasive? ranking argumentative comments in the online forum. in the 54th annual meeting of the association for computational linguistics, pages 195–200, 2016. theresa wilson, janyce wiebe, and paul hoffmann. recognizing contextual polarity in phrase-level sentiment analysis. in proceedings of the conference on human language technology and empirical methods in natural language processing, pages 347–354. association for computational linguistics, 2005. christopher r wolfe, m anne britt, and jodie a butler. argumentation schema and the myside bias in written argumentation. written communication, 26(2):183–209, 2009. joel young, craig martell, pranav anand, pedro ortiz, and henry tucker gilbert, iv. a microtext corpus for persuasion detection in dialog. in proceedings of the 5th aaai conference on analyzing microtext, aaaiws’11-05, pages 80–85. aaai press, 2011. 127 dialogue & discourse 8(1) (2017) 106–131 doi: 10.5087/dad.2017.104 a psycholinguistic model for the marking of discourse relations frances yung pikyufrances-y@is.naist.jp computational linguistics, nara institute of science and technology 8916-5 takayama, ikoma, nara, japan kevin duh kevinduh@cs.jhu.edu human language technology center of excellence, john hopkins university stieff building, 810 wyman park drive, baltimore, md 21211-2840, usa taku komura tkomura@inf.ed.ac.uk institute of perception, action and behaviour, school of informatics, university of edinburgh 10 crichton street, edinburgh, eh8 9ab, united kingdom yuji matsumoto matsu@is.naist.jp computational linguistics, nara institute of science and technology 8916-5 takayama ikoma, nara, japan editor: maite taboada submitted 5/2016; accepted 12/2016; published online 1/2017 abstract discourse relations can either be explicitly marked by discourse connectives (dcs), such as therefore and but, or implicitly conveyed in natural language utterances. how speakers choose between the two options is a question that is not well understood. in this study, we propose a psycholinguistic model that predicts whether or not speakers will produce an explicit marker given the discourse relation they wish to express. our model is based on two information-theoretic frameworks: (1) the rational speech acts model, which models the pragmatic interaction between language production and interpretation by bayesian inference, and (2) the uniform information density theory, which advocates that speakers adjust linguistic redundancy to maintain a uniform rate of information transmission. specifically, our model quantifies the utility of using or omitting a dc based on the expected surprisal of comprehension, cost of production, and availability of other signals in the rest of the utterance. experiments based on the penn discourse treebank show that our approach outperforms the state-of-the-art performance at predicting the presence of dcs (patterson and kehler, 2013), in addition to giving an explanatory account of the speaker’s choice. keywords: discourse connectives, psycholinguistics, rational speech acts model, uniform information density c©2017 frances yung, kevin duh, taku komura and yuji matsumoto this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). a psycholinguistic model for the marking of discourse relations 1. introduction speakers or authors produce informative utterances such that listeners or readers can understand the intended message.1 grice’s maxim of quantity states that human speakers communicate by being as informative as required, but no more (grice, 1975). if a speaker always tries to provide as much information as possible, the resulting utterance could become excessively long and tedious. such utterance not only takes effort for the speaker to produce, but also contains redundant information that is not necessary for the listener. in this work, we model how speakers optimally plan the presentation of discourse structure in terms of informativeness. specifically, we propose a model that predicts whether speakers will use or omit a discourse connective, given the sense of the discourse relation they want to convey. discourse relations are relations between unit of texts (known as arguments) that make a document coherent. these relations can be explicitly marked in the surface text or inferred by the readers, as shown in examples 1a to 1c. (1a) it was a great movie, but i did not like it. (1b) it was a great movie, therefore i liked it. (1c) it was a great movie. i liked it. in example 1a, the word but indicates a contrast relation, and therefore indicates a result relation in example 1b. we call but and therefore explicit discourse connectives (dcs). in example 1c, dcs are absent, but a result relation can be inferred. we say the two sentences (called arguments) are connected by an implicit dc. explicit and implicit relations differ in their level of ambiguity. explicit relations can be signaled by a variety of lexical, syntactic and semantic features, of which dcs are the most informative cues to identify discourse relations (pitler et al., 2008).2 in contrast to explicit relations, implicit relations are more ambiguous. for example, i liked it can also be read as a justification for the first sentence in example 1c. however, marking a discourse relation or not is subject to not only ambiguity, but also redundancy. specifically, using an explicit dc decreases the potential for ambiguity. for example, the contrast sense in example 1a is difficult to infer if the dc but is omitted. nonetheless, if the intended discourse sense is highly predictable, it could be verbose or redundant to insert an explicit dc in the utterance, such as the dc therefore in example 1b. in the pdtb, there are similar numbers of implicit and explicit relations (prasad et al., 2008), yet the corresponding sense distributions are largely different. for example, contrast relations are more common in explicit relations than in implicit relations. these statistics suggest that both options are similarly frequent but the preference is not equally distributed across senses. in order to explain how human speakers choose the optimal level of marking in their utterances, this work models how speakers rationally balance ambiguity and redundancy.3 we combine two information-theoretic frameworks, namely the rational speech acts (rsa) model and the uniform information density (uid) principle. 1. in this article, speakers and listeners are interchangeably used with authors and readers, respectively. 2. in this work, we use the term “explicit relations” to refer to discourse relations that are signaled by explicit dcs. 3. however, this work does not explore the level of consciousness during the reasoning of the marking strategy. we leave it for future work to assess whether people intentionally or subconsciously balance ambiguity and redundancy. 107 yung, duh, komura and matsumoto on the one hand, the rsa model (frank and goodman, 2012) formalizes the inter-relation of language comprehension and production in terms of a listener model and speaker model, which are interwoven. recent findings in human language processing suggest that listeners simulate how an utterance is produced to guide comprehension, and speakers consider the ease of comprehension when planning production (clark, 1999; kilner et al., 2007; pickering and garrod, 2007). based on these findings, the rsa model quantifies the informativeness of the choice of discourse marking by the likelihood for the listerners to disambiguate the discourse relation. on the other hand, the uniform information density (uid) principle (levy and jaeger, 2006) is applied to model how redundant utterances are avoided. the uid principle views language communication as a form of information transmission through a noisy channel, through which a constant rate of information flow is optimal according to shannon’s information theory (genzel and charniak, 2002; levy and jaeger, 2006; shannon, 1948). speakers thus structure utterances by optimizing information density, which is the quantity of information (measured by surprisal) transmitted per unit of utterance, typically a word. in particular, a highly predictable utterance triggers a drop in information density, which has to be smoothed by choosing a more ambiguous utterance, such as by leaving out linguistic markers. in short, our computational model implements grice’s maxim of quantity by computing how speakers try to be informative (using the rsa model), but not too informative (based on the uid principle). we apply this model to predict whether an explicit or implicit dc is used to express a discourse relation, given the context of the discourse relation and the discourse sense to be conveyed. using the actual presence or absence of dcs in the pdtb as the gold standard for evaluation, our model not only achieves a higher accuracy than previous work (patterson and kehler, 2013), but also provides an interpretable account of the various cognitive factors behind the predicted decision. in terms of application, a model that predicts the marking of discourse relations not only contributes to a better understanding of the human language production mechanism, but is also important for automatically generating coherent, human-like text and dialogue. in particular, the degree of marking in discourse relations is cross-linguistically different (meyer and webber, 2013; yung et al., 2015). it remains a challenge for machine translation systems to explicitate (translate an implicit dc to an explicit dc) or implicitate (translate an explicit dc to an implicit dc) discourse relations in source texts, as human translators do (hoek et al., 2015; hoek and zufferey, 2015; li et al., 2014; meyer and webber, 2013; yung et al., 2015), as it is not yet clear how dc explicitation and implicitation are subject to the convention of discourse marking in the target text. the rest of this paper is organized as follows. we start with a review of related work on discourse relation marking in section 2, followed by a description of our proposed methodology in section 3. experiments and evaluation using the pdtb are described in section 4. lastly, section 5 discusses the advantages and disadvantages of the methodologies in this study and directions for future work, and section 6 draws the paper’s conclusion. 2. related work this section summarizes previous work on modeling the explicit marking of discourse relations. we first give a brief summary on the annotation strategy of the pdtb, which is used as the gold standard for discourse marking in our experiment. we then introduce the state-of-the-art method for the automatic prediction of the presence of discourse connectives in a corpus. lastly, we describe how discourse relation marking is explained by the uid principle in the existing literature. 108 a psycholinguistic model for the marking of discourse relations on the other hand, background information on the rsa model is explained in section 3, in connection with our proposed methodology. 2.1 penn discourse treebank (pdtb) in this work, we applied a computational model to predict the actual marking of discourse relations in corpus data given a particular discourse relation. to achieve this, a corpus annotated with discourse relations and marking is essential. there are various corpora annotated with discourse relations, such as the rst discourse treebank (carlson et al., 2002) and discourse graphbank (wolf et al., 2005), but discourse markers are annotated and associated with discourse relations in only two resources: the pdtb (prasad et al., 2008) and the rst signaling corpus (das et al., 2015). the proposed model in this work is trained and evaluated against the annotation of the pdtb, which is the largest available discourse-annotated corpus in english. the pdtb consists of news articles collected from the wall street journal. discourse relations are annotated between each pair of arguments, which are mostly clauses or single sentences. below are three examples of this annotation. (2) the otc market has only a handful of takeover-related stocks. but (explicit; comparison.contrast) they fell sharply. (wsj2379) (3) japan’s finance ministry had set up mechanisms to limit how far futures prices could fall in a single session and ... to give market operators the authority to suspend trading in futures at any time. (implicit: but; comparison) maybe it wasn’t enough. (wsj0097) (4) this cannot be solved by provoking a further downturn; reducing the supply of goods does not solve inflation. (implicit 1: instead expansion.alternative.chosen alternative), (implicit 2: so; contingency.cause.result and expansion.alternative) our advice is this: immediately return the government surpluses to the economy... the pdtb follows a lexically-grounded approach in the annotation of discourse relations (webber et al., 2003). first, explicit dcs are identified, based on a list of dcs that are accumulated in the course of annotation, and labeled with relation senses (example 2). other expressions that signal discourse relations, such as “the reason is”, are identified as alternative lexicalization (altlex) and labeled with relation senses as well. if explicit markers are absent between two sentences within the same paragraph, there are three options for annotation: i) if a discourse relation can be inferred and expressed by a dc, the relation is labeled as implicit and the candidate dc and relation sense are annotated (example 3); ii) if a discourse relation cannot be inferred but the two sentences are about the same entity, the relation is labeled entrel; and iii) if the two sentences are unrelated, they are tagged as norel. this work focuses on the marking of discourse relations by discourse connectives, so we only make use of samples labeled explicit or implicit. senses in the pdtb are defined in a hierarchy of two to three levels, as shown in figure 1. some relations have multiple senses. up to two dcs can be assigned to an implicit relation and, in turn, each (implicit or explicit) dc can be labeled with up to two senses (example 4). similarly, certain level 2 senses, as in example 2, are resulting from the back-off strategy in annotation, i.e. when the annotators disagree on the level 3 senses. this is also a kind of multiple sense. we argue that multisense discourse relations are non-compositional, which means that we consider each combination of multiple senses as an individual sense. more details are given in section 4.1. 109 yung, duh, komura and matsumoto figure 1: sense hierarchy of pdtb (prasad et al., 2008) 2.2 automatic prediction of discourse relation marking the choice of discourse marking strategies has been studied in earlier works as a subtask for natural language generation (allbritton and moore, 1999; danlos, 1998; grote and stede, 1998; moser and moore, 1995; scott and de souza, 1990; soria and ferrari, 1998). in the absence of largescale resources, investigations are based on manually derived rules and lexicons or psycholinguistic experiments. with the emergence of large corpora annotated with discourse relations, patterson and kehler (2013) presented a machine-learning approach to predict whether an explicit or implicit dc is used in the corpus for a particular discourse relation. they argue that while the choice is related to the ease of inference, it may also depend on other stylistic or textual factors. a classifier is trained to predict whether a candidate dc (the dc that actually occurs in the text as an explicit dc or annotated as an implicit dc) is actually present, given the sense of the discourse relation and the arguments. features include observable surface forms, such as argument length, count of subject nouns, and content word ratio, as well as contextual discourse structures, such as the previous discourse relation and whether the relation is embedded or shared. the classifier is trained and tested on a subset of the most frequent relations from the pdtb. an overall high classification accuracy is achieved and relation-level and discourse-level features are found to be more useful than argument-level features. we evaluate our proposed discourse marking model by predicting the use of explicit or implicit dcs in pdtb, as in patterson and kehler (2013). however, patterson and kehler (2013) focus 110 a psycholinguistic model for the marking of discourse relations on a data-driven approach that correctly replicates the occurrence of dcs in the corpus without a theoretically grounded explanation of why an utterance is preferred by the speaker. our work differs in that we model the option of marking from the viewpoint of human language production, explaining the speaker’s choice in terms of information theories. as a model of human language production, we do not make use of the candidate dc to represent the message to be conveyed by the speaker, as it is the result of the speaker’s choice, if an explicit dc is preferred. we only make use of the relation sense label to represent the message. nonetheless, our model achieves higher accuracy when evaluated on the same test samples. 2.3 discourse relation marking explained by uid the uid principle provides a theoretical basis that connects the use of dcs with that of other discourse relation signals. according to uid, information density rises when an utterance is “surprising” and drops when an utterance is highly predictable. to smooth the peaks and troughs, speakers adjust the ambiguity of an utterance by including or omitting linguistic markers. uid is applied to explain a variety of speaker’s options, such as phonetic (aylett and turk, 2004), morphological (frank and jaeger, 2008), and syntactic (jaeger, 2010) reductions as well as referential expressions (tily and piantadosi, 2009). in the context of discourse relations, our approach incorporates the uid assertion that explicit dcs are omitted when the discourse relation is highly predictable. an analysis of the pdtb in the literature shows that causal and continuous senses are more often implicit, or marked by less specific dcs (asr and demberg, 2012). indeed, these senses are presupposed by listeners according to linguistics theories (kuperberg et al., 2011; levinson, 2000; murray, 1997; sanders, 2005; segal et al., 1991). in addition, the dc instead is more often dropped for the discourse relation chosen alternative, if the first argument contains negation words, which are identified cues for this relation (asr and demberg, 2015). in fact, expectations about discourse relations are triggered by various signals, such as verb classes (rohde and horton, 2014). the corpus statistics presented in these analyses support the uid hypothesis that expected, predictable relations are more likely to be conveyed implicitly, and thus more ambiguously, to maintain a steady information flow. however, there are explicit causal and continuous relations and some chosen alternative are marked even the first argument is negated. although measures have been proposed to rate the implicitness of a relation sense (asr and demberg, 2013; jin and de marneffe, 2015), these measures only quantify the general marking of each sense in the data (e.g., the contrast sense), but not the speaker’s choice for each particular instance (e.g., the contrast sense, given particular arguments and context). in contrast, our model incorporates an information density predictor, which specifically predicts the expectability of a given relation. in turn, the speaker’s choice of discourse marking is biased based on the predicted degree of expectability. instead of particular senses or cues in the corpus, we generally apply uid to model each relation instance of in the corpus irrespective of the relation sense, in conjunction with other language production factors. our research questions are as follows: 111 yung, duh, komura and matsumoto 1. does our proposed model explain speakers’ choice of dc marking? if the hypotheses of the model is appropriate, each component in the model should contribute to the prediction accuracy. 2. how does the prediction performance of the proposed model compare with the state-of-theart, i.e., patterson and kehler (2013)? 3. modeling the marking of discourse relations using the rsa model this section describes our proposed method for modeling the speaker’s choice of dc marking. we start by explaining the key elements of the rsa model. then, in order to provide a full picture of the application of rsa to discourse processing, we also give a brief account of discourse relation interpretation using the listener model of rsa, as described in yung et al. (2016). this is followed by the details of our proposed marking model, which predicts the marking of a discourse relation produced by a speaker, and is based on the speaker model of rsa. 3.1 theoretical framework: the rational speech acts (rsa) model the rsa model (frank and goodman, 2012) describes the speaker and listener as rational agents who cooperate towards efficient communication in a game-theoretic approach. the model belongs to the line of reasoning that speakers and listeners cooperate in a conversation by recursively inferring the reasoning of each other (benz et al., 2016; frank and goodman, 2012; goodman and lassiter, 2014; goodman and stuhlmüller, 2013; jäger, 2012), and studies of bayesian interpretation in general (zeevat, 2011, 2015). rsa has been used to successfully explain existing psycholinguistic theories and predict experimental results at various linguistic levels, such as the perception of scalar implicatures (e.g., “some” meaning “not all” in pragmatic usage) and the production of referential expressions (e.g., using pronouns or proper nouns to refer to an entity) (bergen et al., 2014; kao et al., 2014; lassiter and goodman, 2013, 2015; potts et al., 2015). recent efforts also learn and evaluate the models using corpus data instead of experimental data (monroe and potts, 2015; orita et al., 2015). based on the theory that production guides comprehension and that comprehension guides production, the rsa model is composed of a listener model embedding a speaker model and a speaker model embedding a listener model. details are explained in the following sections. 3.1.1 listener model rational listeners assume the utterance they hear contains the optimal amount of information. listeners predict the intended message of speakers by bayesian inference (equation 1), plistener(s|w,c) ∝ pspeaker(w|s, c)p (s) (1) where s is the meaning of an utterance; w is the utterance produced by the speaker, and c is the context. during comprehension, the listener reasons about how the utterance is produced. pspeaker(w|s, c) represents the speaker model in the listener’s mind. on the other hand, p (s) represents the salience of the meaning, which is the private preference of the listener, but is also subject to the shared knowledge between the speaker and listener. 112 a psycholinguistic model for the marking of discourse relations 3.1.2 speaker model rational speakers emulate the listeners’ interpretation and choose an utterance they believe to be informative. in addition, an utterance that is easy to produce is preferred. specifically, speakers choose an utterance by soft-max optimizing the expected utility (u(w; s, c)) of the utterance (equation 2). pspeaker(w|s, c) ∝ eα·u(w;s,c) (2) utility is the effectiveness of the speaker using utterance w to express meaning s in context c, and α is the decision noise parameter, which is set to 1 to represent a rational speaker. here, α = 0 means the decision is completely unrelated to pragmatic reasoning; α = 1 represents the luce’s choice axiom (frank and goodman, 2012), i.e., a rational decision; and α > 1 suggests biased choices. the speakers select an utterance which, they think, is informative to the listeners and not costly to produce. utility is thus defined as the informativeness (i(s;w,c)) of the utterance, deducted by the cost (d(w)) to produce it (equation 3). u(w; s, c) = i(s;w,c)−d(w) (3) if the sense to be conveyed by the chosen utterance is unconventional and surprising, the utterance is less useful. therefore, informativeness is quantified by the negative surprisal of the sense to be conveyed with respect to the utterance being used (equation 4). i(s;w,c) = lnplistener(s|w,c) (4) in turn, during discourse production, the speaker emulates how the listener interprets the utterance. here, plistener(s|d,c) is the listener model inferred by the speaker. it is the probability that the speakers assume the listeners can interpret their intended meaning s. to summarize, the speaker and listener emulate the language processing of each other. however, instead of unlimited iterations (i.e., the speaker thinks the listener thinks the speaker thinks...), the inference is grounded by a literal interpretation of the utterance. literal interpretation means the listener does not reason about the the likelihood of the sense of an utterance, but always assigns the same interpretation to the same utterance. similarly, a literal speaker always uses the same utterance for the same sense. figure 2 illustrates the direction of pragmatic inference between the speaker and listener in their minds. literal listener l0 and literal speaker s0 do not reason about the reasoning of their counterparts. pragmatic listener l1 reasons about the pragmatic speaker s1, who in turn reasons about literal listener l0. pragmatic listeners and speakers at higher levels (e.g., l2 and s2) reason with more iterations, but previous studies demonstrate that one level of reasoning is robust for modeling a human’s interpretation (goodman and stuhlmüller, 2013; lassiter and goodman, 2013; yung et al., 2016). our proposed model for dc production is based on pragmatic speaker s1, who reasons about literal listener l0. 3.2 listener model for discourse relation interpretation yung et al. (2016) used the l1 listener model of rsa to model how listeners interpret the sense of a dc. given dc w and context c in a text, the listener’s interpreted relation sense si is the sense that 113 yung, duh, komura and matsumoto figure 2: directions of reasoning of listeners and speakers. (reproduced from yung et al. (2016)) maximizes plistener(s|w,c),and si is specifically defined as si = arg max s∈s plistener(s|w,c) (5) where s is the set of defined relation senses. context variable c is defined by the immediately previous relation, including the sense and form (explicit dc or not) of the relation. first, literal listener l0 interprets the dc directly by its most likely sense in the context. the probability is estimated by counting the co-occurrences in the pdtb in which explicit and implicit dcs are labeled with discourse relation senses. pl0(s|w,c) = count(s, w,c) count(w,c) (6) next, pragmatic speaker s1 estimates the utility of a dc by emulating the comprehension of the literal listener l0 (equations 2, 3, and 4). the probability that pragmatic speaker sn will use dc w to express meaning s is estimated as: psn(w|s, c) = elnpln−1 (s|w,c)−d(w)∑ w′∈w elnpln−1 (s|w′,c)−d(w′) (7) where n ≥ 1. moreover, w is the set of annotated dcs, including null, which stands for an implicit dc. the cost function d(w) in equation 7 measures the production effort of the dc. yung et al. (2016) simply defined the cost of producing any explicit dc by a constant positive value, which is tuned manually in the experiments, while the production cost for an implicit dc is 0, since no word is produced. in this work, we further develop other measures to quantify the effort of the utterance. this is explained in section 3.3.4. finally, pragmatic listener l1 emulates the dc production of pragmatic speaker s1 (equation 1). the probability that the pragmatic listener ln will assign meaning s to dc w is estimated as: pln(s|w,c) = psn(w|s, c)pl(s)∑ s′∈s psn(w|s′, c)pl(s′) (8) where n ≥ 1 and s is the set of defined senses. the salience of a relation sense pl(s) is defined by the frequency of the sense in the corpus. pl(s) = count(s)∑ s′∈s count(s′) (9) 114 a psycholinguistic model for the marking of discourse relations the rsa model argues that rational listeners do not just stick to the literal meaning of an utterance. instead, they reason about how likely it is that the speaker will use that utterance, in the current context, based on the informativeness and production effort of the utterance. to evaluate this claim, yung et al. (2016) compared the dc interpretation by the literal listener (based on the probability estimates of pl0 , equation 6), and the pragmatic listeners (based on the probability estimates of pl1 and pl2 , equation 8). their experiments based on the pdtb find that discourse sense predictions made by the pragmatic listeners outperform predictions by the literal listener, providing support to the rsa model. 3.3 proposed method to model discourse relation production in contrast to yung et al. (2016), who applied the listener model of rsa to model the human comprehension of discourse connectives, this work applies the speaker model of pragmatic speaker s1 to model the production of discourse structure. specifically, the proposed model predicts whether the speaker will use an explicit or implicit dc given the discourse relation to be conveyed. we first explain how we adapt the rsa model to predict discourse relation marking, followed by the details of each component. 3.3.1 speaker model as a marking model for discourse relations the probability pragmatic speaker s1 will use utterance w to convey intended message s in context c is: ps1(w|s, c) = eu(w;s,c)∑ w′∈w eu(w′;s,c) (10) we consider the binary choice of explicit or implicit dcs in this task. utterance w thus comes from set w = {(exp)licit, (imp)licit}, where both explicit and implicit dcs are grammatically valid to convey s, the sense of discourse relation. our model thus predicts a speaker’s choice of dcs based on the following two probabilities: ps1(exp|s, c) = eu(exp;s,c) eu(exp;s,c) + eu(imp;s,c) ps1(imp|s, c) = eu(imp;s,c) eu(exp;s,c) + eu(imp;s,c) (11) according to equation 3, the utility u of an explicit dc equals its informativeness i less the production cost d. we define the informativeness of using an explicit dc as the difference in the amount of information conveyed when the dc is used and not, which is quantified by negative surprisal 4. u(exp; s, c) = i(s; exp, c)−d(exp) i(s; exp, c) = lnpl0(s|exp, c)− lnpl(s|c) (12) where pl(s|c) is the salience of sense s in context c, irrespective of how the sense is presented. high i(s; exp, c) means it is informative and not surprising to use an explicit dc for this sense. 4. the conventional term informativeness in rsa is defined by negative surprisal, while information density in uid is the surprisal. 115 yung, duh, komura and matsumoto in contrast, if the speaker does not use an explicit dc, the relation sense is inferred from the arguments. in addition, as discussed in section 2.3, default discourse senses are more often unmarked. in other words, a null dc is also informative for discourse sense prediction. we thus define the probability that a speaker will choose an implicit dc to be proportional to the sum of the the utilities of a null dc and arguments (args). therefore: eu(imp;s,c) = eu(null;s,c) + eu(args;s,c) u(null; s, c) = i(s;null, c)−d(null) u(args; s, c) = i(s; arg, c)−d(args) (13) no effort is required to produce a null dc. further, we assume that the arguments have been produced to convey other information irrespective of their discourse informativeness, so no extra effort is needed. therefore, d(null) and d(args) both equal 0. the amount of information that the null dc provides for the discourse relation is defined similarly to equation 12, as follows: i(s;null, c) = lnpl0(s|null, c)− lnpl(s|c) (14) the informativeness of the arguments is used as the information density predictor, following the uid principle. it is less straightforward to measure. we propose an indirect measure that we explain in detail in section 3.3.3. to summarize, the marking model predicts that speakers will use an explicit dc if eu(exp;s,c) > eu(null;s,c) + eu(args;s,c) (15) and that they will use an implicit dc otherwise. details of each of the components are explained in the following sections. 3.3.2 informativeness of dcs this section explains how we estimate the informativeness of using an explicit and implicit dc respectively in equations 12 and 14. we can assume that the utterance lexicon w = {exp, imp} in equation 10 and the set of speaker’s intended messages (all possible discourse relation senses) are always valid. this is in contrast with other pragmatic situations where rsa has been applied. for referential expressions, for example, the lists of referents and grammatically correct pronouns differ case by case. as in the interpretation model in section 3.2, we extract the universal distributions of pl(s|c), pl0(s|exp, c), and pl0(s|null, c) from corpus data. this is based on counting co-occurrences between a sense and the occurrence of an explicit or null marker under a context. we extract these distributions from the training portion of the corpus, i.e., excluding the samples for testing and tuning parameters (see section 3.3.3). following yung et al. (2016), we define context c as the surrounding discourse relations, which are also used as features in discourse planning for natural language generation (biran and mckeown, 2015). specifically, the discourse contexts are: the full discourse sense annotated in pdtb (s), the fourway top level sense (ts), the form of discourse presentation (f) such as “explicit” or “implicit”, the combination of sense and form (sf), and the combination of top sense and form (tsf). 5 the 5. we use the five forms of discourse presentation defined in the pdtb: explicit dc, implicit dc, alternative lexicalization, entity relation and “no relation”. 116 a psycholinguistic model for the marking of discourse relations contexts are taken from window sizes of 1 to 2: previous one (10), next one (01), previous two (20), next two (02), and the previous one paired with the next one (11). we hypothesize that the speaker also thinks ahead about the coming discourse structures when planning the current ones. predictions based on various discourse contexts are compared in the experiment. 3.3.3 informativeness of arguments the informativeness of arguments i(s; arg, c) in equation 13 refers to the contribution of the arguments to present the discourse sense. it estimates the information density of the utterance towards the sense. following the uid principle, information density drops when the intended discourse sense is predictable from the arguments alone, and thus the explicit dc is omitted. the presence of features in the arguments that signal a particular sense makes the sense more predictable, and thus reduces the chance that an explicit dc for that sense is used. for example, as introduced in section 2.3, the dc instead is less used to present the chosen alternative sense if the first argument is negated (asr and demberg, 2015). in order to model the marking of every relation, we generalize the idea to capture various cues in the arguments for all senses by means of an automatic discourse parser. the implicit dc sense classifier of the discourse parser uses various features in the arguments to identify implicit discourse senses (lin et al., 2009; park and cardi, 2012; pitler et al., 2009; rutherford and xue, 2014). for example, if there is a pair of antonyms in the two arguments (e.g. “high” in the first argument, and “low” in the second argument), the contrast relation is more likely. furthermore, modal words, such as should and may, can be used to express the conditional relation. given a pair of arguments, the classifier estimates the probability of each discourse sense and outputs the most likely sense. a high probability estimate indicates a more certain sense prediction, suggesting that more cues are identifiable from the arguments, and thus explicit marking is less necessary. our motivation for using the implicit dc classifier is based on the hypothesis that the classifier can better predict the sense of relations that are actually implicit than those that are actually explicit because more features in the arguments are identifiable. similar to the informativeness of dcs, we quantify the informativeness of the arguments using an information-theoretic approach. for comparison, we propose two methods to approximate the informativeness of the arguments based on the probability distribution estimated by an automatic implicit dc parser: (1) the negative surprisal of the estimated probability pp of the parser output sense soutput (equation 16), and (2) the negative entropy of the probability distribution estimated by the parser (equation 17). i(s; arg, c) = wa · lnpp(soutput) (16) i(s; arg, c) = wa ∑ sp∈o pp(sp) logpp(sp) (17) where o is the set of senses defined in the parser and wa is a positive weight tuned on the development samples (a set of held-out samples different from the training and testing sets). we measure the general informativeness of the arguments to imply any discourse senses, so soutput does not necessarily equal s. we employ the implicit sense classifier from the winning parser of the conll shared task 2015 (wang and lan, 2015), which was designed to identify a subset of fourteen implicit senses plus the entity relation. features used in the classifier include production rules, dependency rules, last word 117 yung, duh, komura and matsumoto or argument 1, first three words of argument 2, presence of modality verbs and inquirer, polarity, the immediately preceding dc, and brown cluster pairs. in our experiment, the syntactic features are based on automatic parsing using the stanford corenlp (manning et al., 2014). the two arguments of a relation instance, which can actually be explicit or implicit, are passed to the implicit dc classifier and i(s; arg, c) is calculated using the output probabilities. the parser has been trained on the same sections of the pdtb as the training set used in our experiment, so there is no overlap with our test samples. although the performance of this state-of-the-art implicit dc classifier is still unsatisfactory (34.45% on pdtb section 23)6, our method only makes use of the confidence of the prediction, which is based on whether discourse related features are detected or not. we hypothesized that the implicit dc classifier can better predict actually implicit relations than actually explicit relations. in fact, this is the case. the classification accuracy of the originally explicit relations (28.45%) is significantly lower than that of the originally implicit relations (51.30%) on the test set, when matching at the fourway top level discourse sense and counting predictions of entity relation as expansion. this supports our motivation to use the parser estimation as an information density predictor. 3.3.4 cost function cost function d(exp) models a speaker’s effort required to produce an explicit dc for the intended discourse sense. we propose five versions of the cost function that are inspired by existing psycholinguistic findings. 1. mean dc length production cost intuitively increases with word length (orita et al., 2015). we define the mean dc length of a discourse relation as the mean number of characters of all valid dcs for that sense normalized by the average word length of all dcs. a lexicon of possible dcs per discourse sense is derived from the whole corpus. for multi-word dcs, a white space is simply counted as one character. as mentioned in section 2.3, we view that speakers first decide to use an explicit dc or not, then decide which dc best expresses the relation. therefore, we do not use the length of the candidate dc directly. 2. dc/arg2 ratio similarly, we use the mean number of words normalized by the number of words in argument 2 as another version of the cost function. this is based on the hypothesis that dc production cost is related to the relative length of the dc comparing with the whole utterance, rather than the absolute length. 3. prime frequency structural priming refers to the tendency for humans to process a linguistic construction (the target) more easily if the construction has been used before. in terms of language production, a speaker tends to repeat a previous construction (the prime) because it takes less effort than generating an alternative construction. a lexical prime of an explicit dc is ideally the same explicit dc or an explicit dc of the same sense. however, we found that repetitions of the same dc are too sparse in the data. 6. http://www.cs.brandeis.edu/˜clp/conll15st/results.html 118 a psycholinguistic model for the marking of discourse relations therefore, we define the choice of explicit or implicit marking as the structural prime of presenting a discourse relation. specifically, we use the reciprocal of the count of primes (any explicit dc occurring before the current position) as the production cost, since the strength of priming effect is known to increase with the frequency of the primes (bock, 1986; levelt and kelter, 1982; smith and wheeldon, 2001). it is not yet clear whether priming occurs in dc selection. for example, it is also reported that speakers tend to avoid the same dc if the relation is embedded in a relation preceded by the same dc (moser and moore, 1995). the result of our experiment could provide indirect evidence on discourse level priming. 4. prime distance we also use the prime-target distance normalized by the length of the text article as another version of the production cost. psycholinguistic findings suggest that the priming effect is more subtly affected by the prime-target distance (bock et al., 2007; gries, 2005; jaeger and snider, 2008). 5. distance from start we use the relative position of the relation within the text as the production cost. we hypothesize that more effort is needed as the production proceeds, as the accumulated length of the whole utterance increases. the range of values of the cost function depends on the cost definition. we thus adjust the values with a constant weight wc, which is tuned on the development samples in the experiments: d(exp) = wc · cost(exp) (18) 4. experiment we apply the model to simulate a speaker’s choice of explicit or implicit dc for discourse relations in the pdtb corpus. the aim of the experiment is to answer our research questions introduced in the end of section 2. 4.1 data and setting the experiment is based on the annotation of discourse relation senses and explicit/implicit dcs in the pdtb, as described in section 2.1. annotations of other forms of discourse relations, such as entity relations and attributions, are excluded. in addition, our model is based on the assumption that w = {explicit, implicit} for all relations, yet it is notable that intra-sentential implicit dcs are not annotated in the pdtb (prasad et al., 2014). in addition, as a result of the annotation procedure, implicit relations always occur in between two arguments. we thus exclude intra-sentential samples and cases where the dcs are not in between two arguments, hence w = {explicit, implicit} is always true. the resulting experimental data set contains 5, 201 explicit and 16, 049 implicit relations7. most existing works split a multi-sense sample into separate samples, each labeled with one of the senses. however, it is notable that the individual senses of a multi-sense relation are not disjoint 7. four cases of intra-sentential implicit relations, due to sentence splitting errors in the corpus, are removed. in the testing phrase, excluded samples are counted as explicit by default. 119 yung, duh, komura and matsumoto and having multiple senses is part of the sense (asr and demberg, 2013; prasad et al., 2014). the property of multiple senses is an important factor of our dc production model: speakers could choose an explicit dc for each sense, but if they have to express two senses at the same time, an implicit dc could be more usable. therefore, we treat all combinations of senses as individual senses, each containing one to three joint sense labels.8 this results in a total of 122 senses. table 1 is a summary of the distribution of the relation senses in descending order of frequency. in fact, joint multi-senses are not rare: the most frequent multi-sense, expansion.conjunction– temporal.synchrony, is the 17th most frequent sense. sense exp imp 1 expansion.conjunction 1, 380 3, 314 2 comparison.contrast 1, 283 1, 200 3 expansion.restatement.specification 75 2, 406 4 contingency.cause.reason 28 2, 295 5 contingency.cause.result 269 1, 649 6 expansion.instantiation 119 1, 383 7 comparison.contrast.juxtaposition 507 672 8 comparison.concession.contra-expectation 475 179 9 temporal.asynchronous.precedence 117 479 10 expansion.list 84 374 ... ... ... ... 17 expansion.conjunction#temporal.synchrony 74 114 ... ... ... ... 20 expansion 8 89 ... ... ... ... 50 contingency.pragmatic cause.justification#expansion.instantiation 0 6 ... ... ... ... 122 contingency 0 1 total 5,201 16,049 table 1: sense distribution of explicit and implicit dcs in the experimental data. the experimental data are split in the same way as in previous work (patterson and kehler, 2013): sections 2–22 are used as the training set, sections 0–1 as the development set; and sections 23–24 as the test set. in the training phrase of the experiment, probability distributions in the marking model are deduced from the training set of the experimental data. in the testing phrase, the model is applied to predict whether an explicit/implicit dc is likely to be used for each discourse relation in the development set and test set. for direct comparison with previous work, samples of infrequent dcs and relation senses were excluded from the development and test sets according to the same criteria as in previous work (patterson and kehler, 2013). the resulting development and test sets contain 1, 720 and 1, 878 relations, respectively. during evaluation, the predictions are compared with the actual marking in the corpus. parameters in the model (wa and wc) were selected 8. there is only one sample of three joint labels in our experimental dataset (example 4), although up to four joint labels are possible (two implicit dcs labeled with two senses each). 120 a psycholinguistic model for the marking of discourse relations to maximize the prediction accuracy on the development set and the same optimal parameters were used on the test set. 4.2 results the explicit and implicit dc prediction performance of our proposed marking model based on various component combinations of the model is summarized in table 2. discourse arg. info. cost function dev: sections 0-1 test: sections 23-24 context c eu(args;s,c) d(exp) accuracy f1exp f1imp accuracy f1exp f1imp bl constant 0 0 .8494 .8721 .8170 .8536 .8754 .8225 soa (patterson and kehler, 2013) – – – .8660 – – (a) f10 0 0 .8552 .8762 .8258 .8552 .8760 .8259 sf10 0 0 .8593 .8773 .8351 .8546 .8736 .8291 f20 0 0 .8541 .8748 .8251 .8541 .8748 .8253 f11 0 0 .8512 .8723 .8217 .8541 .8749 .8250 ts10 0 0 .8523 .8723 .8217 .8536 .8748 .8236 (b) constant surprisal 0 .8948++ .9013 .8874 .8701 .8810 .8570 constant entropy 0 .8953++ .9019 .8879 .8695 .8805 .8563 (c) constant 0 mean dc length .8936++ .8968 .8902 .8759+ .8855 .8646 constant 0 dc/arg2 ratio .8948++ .9004 .8885 .8733 .8821 .8631 constant 0 prime frequency .8860+ .8879 .8851 .8727 .8821 .8618 constant 0 prime distance .8924++ .9018 .8812 .8749 .8855 .8620 constant 0 distance from start .8930++ .8938 .8923 .8770+ .8790 .8749 (d) f10 entropy dc/arg2 ratio .9017++ .9025 .9010 .8818+ .8829 .8806 tsf01 surprisal prime frequency .8953++ .8983 .8922 .8892++* .8931 .8851 ts01 entropy prime distance .8948++ .8998 .8892 .8903++* .8923 .8883 table 2: accuracies and f1 scores of predicted dc marking. the best values are bolded. (abbreviations: s: full relation sense; ts: top-level sense; f: relation form; sf: sense and form; tsf: top sense and form; 10: previous relation; 20: previous 2 relations; 11: previous relation and next relation) +/++: significant improvement over baseline (bl) accuracy at p < .05 and p < .001 respectively; *: significant improvement over state-of-the-art (soa) accuracy at p < .03 (by pearson’s χ2 test) row bl shows the results of the model without the cost function and argument informativeness component with constant context c. we considered this setting to be the baseline, in which the prediction is solely based on the distributions of p (s|exp) and p (s|imp). considerably high accuracy is achieved, suggesting that the speaker’s choice of marking is strongly related to the intended discourse sense. row (a) shows the prediction results based on the distributions ofp (s|exp, c) andp (s|imp,c), where c is the discourse context. the five best combinations of contexts and window sizes are shown. refining the utility of dcs using these contextual constraints, such as previous relation 121 yung, duh, komura and matsumoto senses and marking, does not improve the classification accuracy. this suggests that a speaker’s choice of marking depends on other contextual factors rather than surrounding discourse relations. row (b) shows the contribution of the argument informativeness component, under a constant discourse context and production cost. classification accuracy increases (significantly for the development set) when the usability of an explicit dc is reduced by the estimated informativeness of the arguments, supporting the uid principle. predictions based on the surprisal of the parser output sense and entropy of the parser output distribution are similar. we also experiment by adjusting the estimated argument informativeness only if the parser output sense is correct (matching at the top level sense). similar improvement is observed. row (c) shows the contribution of the cost function, when the discourse context is set as a constant and argument informativeness is not considered. adjusting the utility of explicit dcs by their production cost increases the classification accuracy significantly. among the various features to model the production cost, “mean dc length” and “distance from start” features give the best results, while “prime frequency” and “prime distance” are the least effective. this suggests that the cost to produce a dc is more subject to the mean dc length and that a priming effect in dc production may be more subtle comparing with informativeness. row (d) shows the performance of predictions based on the three best combinations of components. the highest accuracies and f1 scores are achieved for both explicit and implicit relations. these results answer the first question of the experiment: the proposed model explains the speaker’s choice of dc marking in terms of dc and argument informativeness, as well as production cost, while contextual discourse structure is not a significant constraint on the choice. the answer to the second question is also positive. significant improvement above the state-ofthe-art (row soa) is achieved by the two best combinations (89.03% and 88.92% vs. 86.60%). lastly, we compared the results with a linear classifier trained on the features specified in the model, i.e., the discrete values of the intended sense and various discourse context definitions, and real values of various cost functions and argument informativeness estimates. note that in the proposed model, the training data are used to derive the p (s|exp, c) and p (s|null, c) distributions only, while the linear classifier learns from the features and dc marking of the training set. we used liblinear (fan et al., 2008) to build the classifiers. when extracting the argument informativeness features from the training set, using the automatic discourse parser, we penalize the parser estimates of the implicit samples by a constant ratio, since the discourse parser is also trained on these samples. the classifier achieves an accuracy of 88.3% on the test set, which does not significantly outperform previous work. this suggests that the information-theoretic configuration is an advantage of our model. 5. discussion in this work, we proposed a computational model to predict discourse marking in human language production. the model is trained and evaluated using manual discourse annotation on corpus data as the gold standard. this section explains the advantages and disadvantages of our methodology and suggests directions for future work. one advantage of learning the discourse marking model from pdtb is the compatibility of the pdtb’s annotation with the rsa framework. previous applications of rsa focus on the pragmatic use of language, where the intended message and lexicon of an utterance largely depend on the context. in the task of referring expression generation, the sets of valid referents and referring ex122 a psycholinguistic model for the marking of discourse relations pressions differ case by case. for example, red is an invalid option for referring to a blue ball; he is an invalid option for referring to a woman; and it is difficult to define a finite set of referents in the corpus. in contrast, discourse relations are generally universal across different contexts. a dc can be used or dropped to represent various discourse senses in various contexts, while referring expressions are limited by the properties of each particular referent. in addition, the pdtb annotation scheme pre-defines the sets of dcs and discourse relation senses. in this way, the listener and speaker models of rsa can be derived statistically by counting the co-occurrence of the dcs, sense labels, and contextual factors in the annotated corpus. however, our proposal to use surrounding discourse relations as context did not improve classification accuracy. therefore, one direction to improve the proposed model is to make fuller use of the training data to learn a more expressive and general abstraction of the context governing the choice of discourse marking. however, the pragmatic reasoning approach of rsa has been criticized for being unrealistic, because previous studies find that speakers tend to produce referring expressions that are overspecifying (baumann et al., 2014; dale and reiter, 1995; engelhardt et al., 2006; gatt et al., 2013). in other words, while ideal pragmatic speakers should only focus on the minimum properties that help listeners to identify the referent, the referring expressions that speakers actually choose often include redundant properties that are not necessary for distinguishing the referent from other candidates. in the context of discourse marking, an utterance is over-specifying if an explicit dc is chosen even though an implicit dc is enough for the listeners to infer the discourse relation. another advantage of the proposed model is that it does not rely on pragmatic reasoning alone to model speakers’ choice of discourse marking, but also makes use of the uid principle to penalize the choice of explicit dcs when informative signals are present in the arguments. in addition, learning rsa from corpus statistics allows the model to detect the general trend in the marking of a relation sense. some relation senses are highly likely to be marked/unmarked irrespective of the presence of other signals, while for other relations, the presence of other discourse signals affects the choice, as illustrated in examples 5 to 7. in these examples, the speaker probabilities (equation 11) estimated by the best performing model (last row in table 2) are shown along with the predicted marking choices. for comparison, we also show p ′s(imp|s, c) and p ′s(exp|s, c), which are the speaker probabilities without the uid bias (i.e., eu(arg;s,c) = 0, in equation 13). (5) and market expectations clearly have been raised by the capital gains victory in the house last month (implicit:since; contingency.cause.reason) an hour before friday’s plunge, that provision was stripped from the tax bill. (wsj2429) (without uid) p ′s(exp|s, c) = 0.042 p ′s(imp|s, c) = 0.958 (with uid) ps(exp|s, c) = 0.024 ps(imp|s, c) = 0.976 prediction= implicit (6) boeing’s offer represents the best overall three-year contract of any major u.s. industrial firm in recent history. but (explicit; comparison.contrast.opposition) mr. baker called the letter ...very weak. (wsj2308) (without uid) p ′s(exp|s, c) = 0.813 p ′s(imp|s, c) = 0.187 (with uid) ps(exp|s, c) = 0.535 ps(imp|s, c) = 0.465 prediction= explicit (7) full-time residential programs ... are particularly expensive – more per participant than a year at stanford or yale. (implicit:but; comparison.contrast) non-residential programs are cheaper, ... (wsj2412) 123 yung, duh, komura and matsumoto (without uid) p ′s(exp|s, c) = 0.649 p ′s(imp|s, c) = 0.351 (with uid) ps(exp|s, c) = 0.383 ps(imp|s, c) = 0.617 prediction= implicit in examples 5 and 6, the uid bias does not affect the prediction based on dc informativeness alone, since the contingency.cause.reason sense is dominantly implicit (example 5) and the comparison.contrast.opposition sense is dominantly explicit (example 6), according to the probability distribution in the corpus. in these cases, argument informativeness has little effect on the rsa model. in contrast, in example 7, the comparison.contrast sense could be expressed explicitly or implicitly, and the uid bias reverses the prediction based on dc informativeness. the model predicts that the speakers would not over-specify the discourse relation with a dc, since there are enough informative signals in the arguments (e.g., expensive vs. cheaper or residential vs. non-residential). however, our approach, which approximates the argument informativeness based on the probability output of an automatic discourse parser, is limited by the accuracy of the discourse parser, as shown in example 8. (8) “jeux sans frontieres”... is a hit in france. (implicit:but; comparison.contrast) a u.s.-made imitation under the title “almost anything goes” flopped fast. (wsj2361) (without uid) p ′s(exp|s, c) = 0.762 p ′s(imp|s, c) = 0.238 (with uid) ps(exp|s, c) = 0.624 ps(imp|s, c) = 0.376 prediction= explicit the parser detects low informativeness in the arguments, and thus the model wrongly predicts that explicit marking is more likely. a possible explanation for this is that the constrast between hit and flopped is uncommon, and the parser fails to identify it as a discourse-informative signal. the performance of the discourse parser we used in the experiment is not yet satisfactory. the accuracy of the marking model could be improved with a more accurate discourse parser. in addition, the classifier of the discourse parser may have poorly calibrated probabilities, which means the probability estimates of the parser may not be well associated with how well the parser detects discourse signals. the association, and thus the overall performance of the model, may be improved by an additional probabilistic calibration step on the parser output (nguyen and o’connor, 2015). lastly, we discuss the potential to use behavioral experiments to further validate or improve the proposed method. the experiment described in section 4 evaluates the prediction ability of the model against the actual data in the pdtb. in other words, the marking of each discourse relation chosen by the writers of the wall street journal and the label attributed by pdtb annotators are taken as the gold standard. however, it is possible that other writers would choose differently, given the same relation sense and context. a behavioral experiment could be carried out to compare the judgment of multiple human speakers with the judgment of the annotators of pdtb and writers of the wall street journal, as well as with the model’s predictions. for example, patterson and kehler (2013) use a readability judgment task to test speakers’ choice of explicit or implicit dcs in pdtb. they find that only 68% of the human judgment match the actual data in the pdtb, suggesting a high level of optionality in marking preference. however, readability judgment may be biased towards explicit marking because it is generally agreed that explicit dcs facilitate discourse relation comprehension (britton et al., 1982; haberlandt, 1982; kamalski et al., 2008; loman and mayer, 1983; lorch jr and lorch, 1986; meyer 124 a psycholinguistic model for the marking of discourse relations et al., 1980; millis and just, 1994; sanders et al., 1992; sanders and noordman, 2000). the marking of discourse relations could be examined using a production-oriented experimental design, such as the picture-to-language transcription task described in soria and ferrari (1998). another option is to use a cloze test, in which subjects are presented with the discourse arguments and relation sense to be conveyed and are asked to fill in a dc or leave the relation implicit. following the recent success in crowdsourced discourse annotation (rohde et al., 2015; scholman et al., 2016), a large number of judgments per relation can be collected by crowdsourcing such that a distribution of the marking preference can be obtained. 6. conclusion we have presented a language production model that predicts whether speakers will choose to use an explicit dc or not given the discourse relation they want to express. our model gives a cognitive account of the speakers’ choice and its results outperform those of previous work on the same task. although the option of dc marking is a subtle preference in the absence of other grammatical constraints, our proposed marking model tackles the option as a rational preference by the speaker. using an information-theoretic approach, the model predicts a speaker’s choice by balancing the advantage (informativeness) and disadvantages (production cost and redundancy) of using an explicit marker. this is the first work to apply the rsa framework to discourse processing. based on the discourse context, we adjusted the universal distribution of utterances and senses. this work also practically incorporates the uid principle, using the output of a discourse parser. we plan to expand the model to discriminate the choice of different explicit dcs and assess the effectiveness of the model in applications such as natural language generation or machine translation tasks, where models trained on monolingual data of different languages are applied. the long-term goal is to make use of the listener model to design a full, incremental discourse parsing algorithm that is motivated by the psycholinguistic reality of human discourse processing. references david allbritton and johanna moore. discourse cues in narrative text: using production to predict comprehension. in aaai fall symposium on psychological models of communication in collaborative systems, 1999. fatemeh torabi asr and vera demberg. implicitness of discourse relations. in proceedings of the international conference on computational linguistics, pages 2669–2684. citeseer, 2012. fatemeh torabi asr and vera demberg. on the information conveyed by discourse markers. in proceedings of the annual workshop on cognitive modeling and computational linguistics, pages 84–93, 2013. fatemeh torabi asr and vera demberg. uniform information density at the level of discourse relations: negation markers and discourse connective omission. proceedings of the international conference on computation semantics, pages 118–128, 2015. 125 yung, duh, komura and matsumoto matthew aylett and alice turk. the smooth signal redundancy hypothesis: a functional explanation for relationships between redundancy, prosodic prominence, and duration in spontaneous speech. language and speech, 47(1):31–56, 2004. peter baumann, brady clark, and stefan kaufmann. overspecification and the cost of pragmatic reasoning about referring expressions. proceedings of the annual conference of the cognitive science society, 2014. anton benz, gerhard jäger, and robert van rooij. game theory and pragmatics. springer, 2016. leon bergen, roger levy, and noah d goodman. pragmatic reasoning through semantic inference. unpublished manuscript, 2014. or biran and kathleen mckeown. discourse planning with an n-gram model of relations. in proceedings of the conference on empirical methods on natural language processing, pages 1973–1977, lisbon, portugal, september 2015. association for computational linguistics. j kathryn bock. syntactic persistence in language production. cognitive psychology, 18(3):355– 387, 1986. kathryn bock, gary s dell, franklin chang, and kristine h onishi. persistent structural priming from language comprehension to language production. cognition, 104(3):437–458, 2007. bruce k britton, shawn m glynn, bonnie j meyer, and mj penland. effects of text structure on use of cognitive capacity during reading. journal of educational psychology, 74(1):51, 1982. lynn carlson, mary ellen okurowski, and daniel marcu. rst discourse treebank, 2002. herbert h clark. using language. journal of linguistics, 35(1):167–222, 1999. robert dale and ehud reiter. computational interpretations of the gricean maxims in the generation of referring expressions. cognitive science, 19(2):233–263, 1995. laurence danlos. linguistic ways for expressing a discourse relation in a lexicalized text generation system. workshop of discourse relations and discourse markers, pages 50–53, 1998. debopam das, maite taboada, and paul mcfetridge. rst signalling corpus, 2015. paul e engelhardt, karl gd bailey, and fernanda ferreira. do speakers and listeners observe the gricean maxim of quantity? journal of memory and language, 54(4):554–573, 2006. rong-en fan, kai-wei chang, cho-jui hsieh, xiang-rui wang, and chih-jen lin. liblinear: a library for large linear classification. the journal of machine learning research, 9:1871–1874, 2008. austin frank and t florian jaeger. speaking rationally: uniform information density as an optimal strategy for language production. proceedings of the annual meeting of the cognitive science society, pages 933–938, 2008. michael c. frank and noah d. goodman. predicting pragmatic reasoning in language games. science, 336(6084):998, 2012. 126 a psycholinguistic model for the marking of discourse relations albert gatt, roger pg van gompel, kees van deemter, and emiel krahmer. are we bayesian referring expression generators. in proceedings of the workshop on production of referring expressions: bridging the gap between cognitive and computational approaches to reference, 2013. dmitriy genzel and eugene charniak. entropy rage constancy in text. in proceedings of the annual meeting of the association for computational linguistics, pages 199–206, 2002. noah d goodman and daniel lassiter. probabilistic semantics and pragmatics: uncertainty in language and thought. wiley-blackwell, 2014. noah d goodman and andreas stuhlmüller. knowledge and implicature: modeling language understanding as social cognition. topics in cognitive science, 5(1):173–184, 2013. h paul grice. logic and conversation. syntax and semantics, 3:41–58, 1975. stefan th gries. syntactic priming: a corpus-based approach. journal of psycholinguistic research, 34(4):365–399, 2005. brigitte grote and manfred stede. discourse marker choice in sentence planning. in proceedings of the international workshop on natural language generation, pages 128–137, 1998. karl haberlandt. reader expectations in text comprehension. advances in psychology, 9:239–249, 1982. jet hoek and sandrine zufferey. factors influencing the implicitation of discourse relations across languages. in proceedings the joint acl-iso workshop on interoperable semantic annotation, pages 39–45. ticc, tilburg center for cognition and communication, 2015. jet hoek, jacqueline evers-vermeul, and ted jm sanders. the role of expectedness in the implicitation and explicitation of discourse relations. in proceedings of the workshop on discourse in machine translation, pages 41–46. association for computational linguistics, 2015. t florian jaeger. redundancy and reduction: speakers manage syntactic information density. cognitive psychology, 61(1):23–62, 2010. t florian jaeger and neal snider. implicit learning and syntactic persistence: surprisal and cumulativity. in proceedings of the annual meeting of the cognitive science society, page 827, 2008. gerhard jäger. game theory in semantics and pragmatics, volume 3, pages 2487–2425. mouton de gruyter, 2012. lifeng jin and marie-catherine de marneffe. the overall markedness of discourse relations. in proceedings of the conference on empirical methods on natural language processing, pages 1114–1119, 2015. judith kamalski, ted sanders, and leo lentz. coherence marking, prior knowledge, and comprehension of informative and persuasive texts: sorting things out. discourse processes, 45(4-5): 323–345, 2008. 127 yung, duh, komura and matsumoto justine t kao, jean y wu, leon bergen, and noah d goodman. nonliteral understanding of number words. proceedings of the national academy of sciences, 111(33):12002–12007, 2014. james m kilner, karl j friston, and chris d frith. predictive coding: an account of the mirror neuron system. cognitive processing, 8(3):159–166, 2007. gina r kuperberg, martin paczynski, and tali ditman. establishing causal coherence across sentences: an erp study. journal of cognitive neuroscience, 23(5):1230–1246, 2011. daniel lassiter and noah d goodman. context, scale structure, and statistics in the interpretation of positive-form adjectives. semantics and linguistic theory, 23:587–610, 2013. daniel lassiter and noah d goodman. adjectival vagueness in a bayesian model of interpretation. synthese, pages 1–36, 2015. willem jm levelt and stephanie kelter. surface form and memory in question answering. cognitive psychology, 14(1):78–106, 1982. stephen c levinson. presumptive meanings: the theory of generalized conversational implicature. mit press, 2000. roger levy and t. florian jaeger. speakers optimize information density through syntactic reduction. advances in neural information processing systems, (849-856), 2006. junyi jessy li, marine carpuat, and ani nenkova. assessing the discourse factors that influence the quality of machine translation. in proceedings of the annual meeting of the association for computational linguistics, pages 283–288, 2014. ziheng lin, minyen kan, and hwee tou ng. recognizing implicit discourse relations in the penn discourse treebank. in proceedings of the conference on empirical methods on natural language processing, pages 343–351, 2009. nancy l loman and richard e mayer. signaling techniques that increase the understandability of expository prose. journal of educational psychology, 75(3):402, 1983. robert f lorch jr and elizabeth pugzles lorch. on-line processing of summary and importance signals in reading. discourse processes, 9(4):489–496, 1986. christopher manning, mihai surdeanu, john bauer, jenny finkey, steven j. bethard, and david mcclosky. the standord corenlp natural language processing toolkit. in proceedings of the annual meeting of the association for computational linguistics: system demonstrations, pages 55–60, 2014. bonnie jf meyer, david m brandt, and george j bluth. use of top-level structure in text: key for reading comprehension of ninth-grade students. reading research quarterly, pages 72–103, 1980. thomas meyer and bonnie webber. implicitation of discourse connectives in (machine) translation. in proceedings of the workshop on discourse in machine translation, pages 19–26, 2013. 128 a psycholinguistic model for the marking of discourse relations keith k millis and marcel adam just. the influence of connectives on sentence comprehension. journal of memory and language, 33(1):128–147, 1994. will monroe and christopher potts. learning in the rational speech acts model. arxiv preprint arxiv:1510.06807, 2015. megan moser and johanna d moore. investigating cue selection and placement in tutorial discourse. in proceedings of the annual meeting of the association for computational linguistics, pages 130–135. association for computational linguistics, 1995. john d murray. connectives and narrative text: the role of continuity. memory & cognition, 25 (2):227–236, 1997. khanh nguyen and brendan o’connor. posterior calibration and exploratory analysis for natural language processing models. in proceedings of conference on empirical methods in natural language processing, pages 1587–1598, 2015. naho orita, eliana vornov, naomi h. feldman, and hal daumé iii. why discourse affects speakers’ choice of refering expressions. proceedings of the annual meeting of the association for computational linguistics, pages 1639–1649, 2015. joonsuk park and claire cardi. improving implicit discourse relation recognition through feature set optimization. proceedings of annual meeting of the special interest group on discourse and dialogue, pages 108–112, 2012. gary patterson and andrew kehler. predicting the presence of discourse connectives. in proceedings of the conference on empirical methods on natural language processing, pages 914–923, 2013. martin j pickering and simon garrod. do people use language production to make predictions during comprehension? trends in cognitive sciences, 11(3):105–110, 2007. emily pitler, mridhula raghupathy, hena mehta, ani nenkova, alan lee, and aravind joshi. easily identifiable discourse relations. technical report, university of pennsylvania, 2008. emily pitler, annie louis, and ani nenkova. automatic sense prediction for implicit discourse relations in text. proceedings of the annual meeting of the association for computational linguistics and the international joint conference on natural language processing, pages 683–691, 2009. christopher potts, daniel lassiter, roger levy, and michael c. frank. embedded implicatures as pragmatic inferences under compositional lexical uncertainty. manuscript, 2015. rashmi prasad, nikhit dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. proceedings of the language resource and evaluation conference, pages 2961–2968, 2008. rashmi prasad, bonnie webber, and aravind joshi. reflections on the penn discourse treebank, comparable corpora, and complementary annotation. computational linguistics, pages 921–950, 2014. 129 yung, duh, komura and matsumoto hannah rohde and william s horton. anticipatory looks reveal expectations about discourse relations. cognition, 133(3):667–691, 2014. hannah rohde, anna dickinson, chris clark, annie louis, and bonnie webber. recovering discourse relations: varying influence of discourse adverbials. in workshop on linking models of lexical, sentential and discourse-level semantics, page 22, 2015. attapol rutherford and nianwen xue. discovering implicit discourse relations through brown cluster pair representation and coreference patterns. proceedings of the conference of the european chapter of the association for computational linguistics, pages 645–654, 2014. ted sanders. coherence, causality and cognitive complexity in discourse. in proceedings of the symposium on the exploration and modelling of meaning, 2005. ted jm sanders and leo gm noordman. the role of coherence relations and their linguistic markers in text processing. discourse processes, 29(1):37–60, 2000. ted jm sanders, wilbert pm spooren, and leo gm noordman. toward a taxonomy of coherence relations. discourse processes, 15(1):1–35, 1992. merel cj scholman, jacqueline evers-vermeul, and ted jm sanders. categories of coherence relations in discourse annotation: towards a reliable categorization of coherence relations. dialogue & discourse, 7(2):1–28, 2016. donia scott and clarisse sieckenius de souza. getting the message across in rst-based text generation. current research in natural language generation, 4:47–73, 1990. erwin m segal, judith f duchan, and paula j scott. the role of interclausal connectives in narrative structuring: evidence from adults’ interpretations of simple stories. discourse processes, 14(1): 27–54, 1991. c.e. shannon. a mathematical theory of communication. the bell system technical journal, 27 (379-423; 623-656), 1948. mark smith and linda wheeldon. syntactic priming in spoken sentence production–an online study. cognition, 78(2):123–164, 2001. claudia soria and giacomo ferrari. lexical marking of discourse relations-some experimental findings. in proceedings of the acl-98 workshop on discourse relations and discourse markers, pages 36–42, 1998. harry tily and steven piantadosi. refer efficiently: use less informative expressions for more predictable meanings. proceedings of the workshop on the production of referring expressions, 2009. jianxiang wang and man lan. a refined end-to-end discourse parser. conll 2015, pages 17–24, 2015. bonnie webber, matthew stone, aravind joshi, and alistair knott. anaphora and discourse structure. computational linguistics, 29(4):545–587, 2003. 130 a psycholinguistic model for the marking of discourse relations florian wolf, edward gibson, amy fisher, and meredith knight. the discourse graphbank: a database of texts annotated with coherence relations, 2005. frances yung, kevin duh, and yuji matsumoto. crosslingual annotation and analysis of implicit discourse connectives for machine translation. in proceedings of the workshop on discourse in machine translation, pages 142–152, 2015. frances yung, kevin duh, taku komura, and yuji matsumoto. modeling the interpretation of discourse connectives by bayesian pragmatics. in proceedings of the annual meeting of the association for computational linguistics, pages 531–536, 2016. henk zeevat. bayesian interpretation and optimality theory. bidirectional optimality theory. palgrave macmillan, amsterdam, pages 191–220, 2011. henk zeevat. perspectives on bayesian natural language semantics and pragmatics, pages 1–24. springer, 2015. 131 3751-7942-3-ce dialogue & discourse 10(2) 34-55 doi: 10.5087/dad.2019.202 does planning explain why predictability affects reference production? sandra a. zerkle szerkle@unc.edu department of psychology & neuroscience, university of north carolina – chapel hill jennifer e. arnold jarnold@email.unc.edu department of psychology & neuroscience, university of north carolina – chapel hill editor: patrick healey submitted 08/2017; accepted 06/2019; published online 11/2019 abstract how does thematic role predictability affect reference production? this study tests a planning facilitation hypothesis – that the predictability effect on reference form can be explained in terms of the time course of utterance planning. in a discourse production task, participants viewed two sequential event pictures, listened to a description of the first picture (depicting a transfer event between two characters), and then provided a description of the second picture (continuing with one thematic role character, either goal or source). we replicated previous findings that goal continuations lead to more reduced forms of reference and shorter latency to begin speaking than source continuations. additionally, we tracked speakers’ eye movements in two periods of utterance planning, early vs. late. we found that 1) early planning supports the use of reduced forms but is not affected by thematic role; 2) thematic role only affects late planning; and 3) in contrast with our hypothesis, planning does not account for predictability effects on reduced forms. we then speculate that discourse connectedness drives the thematic role predictability effect on reference form choice. keywords: utterance planning, thematic role predictability, reference production, discourse 1 introduction variation in reference form exists in natural language production: an entity could be referred to with its proper name (lady mannerly), a noun phrase description (the duchess), a reduced form such as a pronoun (she), or even dropping the referent entirely (…and ø left). what guides speakers as they make this choice? previous work has supported the hypothesis that certain discourse conditions promote the use of particular referential forms (e.g., ariel, 1990, 2001; chafe, 1976, 1994; gundel, hedberg, & zacharaski, 1993), where reduced forms tend to be used for salient or accessible referents. for example, speakers strongly prefer pronouns to refer to entities that appeared recently in the discourse, and especially in the previous subject position (arnold, 1998). in addition, in some cases speakers are also more likely to use pronouns for thematic roles that are predictable (arnold, 2001; rosa & arnold, 2017; but see fukumura & van gompel, 2010; rohde & kehler, 2014). however, there are no explicit models of the production mechanisms underlying reference production. in particular, an unsolved puzzle concerns the role of predictability in pronoun production, and how and why predictability affects referential choices. the kind of predictability that matters here is referential predictability, i.e. the likelihood that the speaker will mention that entity at that point in the discourse. this is related to the calculation of referential probability, which includes probability estimates made both before and after the referring expression. there is good evidence that predictability is important for comprehension, for example listeners are faster to understand predictable information, and sometimes even anticipate it (e.g., altmann & kamide, 1999). however, it’s not obvious what predictability would mean to the speaker. people usually know what they are going to say ahead of time, so they aren’t predicting their message; they are planning it. that is, while predictability can be thought of as a property does planning explain why predictability affects reference production? 35 of the referent within context, it is not known whether this leads to actual prediction in all cases. this raises questions about how factors related to predictability in comprehension might affect production, and why. despite the oddness of prediction during production, it is important to investigate the relationship between predictability and reference production, because referential probability underlies numerous models of reference production (frank & goodman, 2012; tily & piantadosi, 2009), other aspects of language production (hale, 2001; levy & jaeger, 2007), and reference comprehension (frank and goodman, 2012; kehler & rohde, 2013). however, empirical findings show that referential probability only sometimes affects pronoun use (see e.g. rosa & arnold, 2017 on the one hand, and e.g. rohde & kehler, 2014 on the other), while it has a stronger effect on prosodic variation (e.g., arnold & watson, 2015). this raises questions about how predictability really affects reference form. here we address this issue by examining the relation between reference form, referential predictability, and utterance planning. 1.1 thematic role predictability affects reference form previous work has revealed that reduced forms of reference are more likely in certain linguistic contexts: if an entity is given (fowler & hossum, 1987), in the subject position (arnold, 1998; chafe, 1976), contextually salient, topical, in focus, accessible (ariel, 1990; arnold, 2010; chafe, 1994; frank & goodman, 2012; givón, 1983; grosz et al., 1995), or predictable in context (jurafsky, bell, gregory, & raymond, 2001; levy & jaeger, 2007; mahowald et al., 2013). these findings span multiple types of reduction, including durational shortening, the use of reduced syntactic structures, as well as the use of pronouns instead of longer referential expressions. the current study is focused on the contrast between explicit descriptions (the duchess) and pronouns or dropped (zero) references. we focus our attention on how thematic roles influence referential predictability, and whether this affects reference form. thematic roles are assigned to a verb’s arguments (dowty, 1991; jackendoff, 1972), and represent a semantic characterization of the entities involved in an event. the current study examines the thematic roles associated with transfer verbs, where the verb of a sentence depicts the transfer of an object (the theme) from one character (the source) to another character (the goal), as in “the duchess handed a painting to the duke”. in a discourse context, these thematic role positions affect people’s expectations about who will be mentioned next, where the goal referent is more expected than the source (rosa & arnold, 2017; stevenson et al., 1994). rosa and arnold measured this expectation by simply asking participants who they thought would be mentioned next, but it also correlates with a tendency for speakers to mention the goal more in either natural discourse (arnold, 2001) or story-continuation tasks (stevenson et al., 1994). there are a few possibilities for why goals are frequently mentioned, and perceived as predictable referents. one idea is that people expect goals because they frequently hear other people mention goals, so this high frequency affects future expectations of similar events (arnold, 1998, 2001; rosa & arnold, 2017). another possibility is that people have an inherent interest in goals, such that we expect the next event to involve the recipient of a transferred object. when we consider transfer verbs, there is also excellent evidence that thematic roles affect reference form. speakers are more likely to use pronouns and zeros when referring to goals than sources (arnold, 2001; rosa & arnold, 2017; zerkle, rosa, & arnold, 2015). some of this evidence comes from a paradigm (also used in the current study) in which participants saw pairs of pictures depicting a transfer in the first picture and a subsequent action in the second picture (see figure 1). for half of the items, the goal character continued in the subsequent event, and for zerkle & arnold 36 the other half the source character continued. participants listened to a description of the first picture, and were instructed to provide a description of the second picture. in three experiments (rosa & arnold, 2017, exp. 1; zerkle et al. 2015, exps. 1 and 2), speakers used reduced forms more for goal than source continuations. this finding occurred on top of the well-known tendency for subject continuations to use a pronoun. see figure 2 for reference form choice results from these three studies. figure 1. example trial: a speaker listens to a description of the first panel (“the duchess handed a painting to the duke”) and then must provide a description of the second panel (e.g., “and then he threw it in the closet”). figure 2. speakers tend to use reduced forms more often for reference to goal continuations than source continuations, as well as for reference to subject continuations than non-subject continuations. (note that rosa & arnold (2017) findings reflect pronoun use only, whereas the reduced forms in zerkle, rosa, & arnold (2015) include both pronouns and zeros). error bars are standard error of the subject means by condition. one interpretation of these findings is that they reflect a preference to use reduced expressions for predictable information (arnold, 2001; rosa & arnold, 2017). this is consistent with other evidence that speakers tend to use reduced forms for redundant information (levy & jaeger, 2007; mahowald et al., 2013), and in particular that speakers use pronouns more for 0% 20% 40% 60% 80% 100% subject non-subject subject non-subject subject non-subject rosa & arnold (2017) zerkle, rosa, & arnold (2015, exp1) zerkle, rosa, & arnold (2015, exp2) % p ro no un /z er o goal source does planning explain why predictability affects reference production? 37 predictable referents (arnold, 1998, 2010). however, this conclusion is complicated by the fact that not all types of predictability affect reference form. several studies have examined implicit causality contexts, such as val admired elyce because.... or val impressed elyce because…. listeners have a strong expectation that the speaker will mention the person who is the more likely cause of the event (i.e., elyce in the first sentence; val in the second). however, several studies have found that this expectation has no effect on the speakers’ use of pronouns (fukumura & van gompel, 2010; kehler et al., 2008; kehler & rohde, 2013; kravtchenko et al. 2017, but see weatherford & arnold, under review). instead, speakers use pronouns more for the subject (val) than the object (elyce). while this difference between transfer verb experiments and implicit causality experiments may be attributable to the different verbtypes, it raises questions about why the thematic role effect occurs. in addition, it’s not clear how thematic role predictability affects language production mechanisms. we assume that in most cases, speakers do not need to predict their own speech, because typically conceptual planning precedes formulation (for an alternate view see pickering and garrod, 2013). instead, we expect that the types of conditions that make referents predictable to the listener may be associated with the production process. here we test the hypothesis that the tendency to use reduced expressions for goals is not a direct result of prediction per se, but rather is a consequence of the mechanisms involved in utterance planning. perhaps goal continuations are easier to plan (conceptually and/or lexically) than source continuations, and it is this facilitation that drives the choice to use a reduced form. the referential predictability previously found for goal continuations could affect both the time course of planning and reference form choice, raising the question of whether or not planning has a direct effect on reference form. 1.2 the role of planning facilitation hypothetically, there are several reasons why ease of planning might affect reference form. one possibility is that early planning might increase the salience of information in the speaker’s discourse model. classic theories about reference production suggest that pronouns are selected when the referent is salient, accessible, or in focus (for a review see arnold, 2016; arnold & zerkle, 2019), where referential salience is often modeled in terms of the activation of the referent’s conceptual representation (arnold & griffin, 2007; van rij et al., 2010). when planning is easy, it may leave greater mental resources for representing information in the discourse model, increasing the activation on those referents. thus, if predictability makes planning easy, this might explain the effect of predictability on reference production. another possibility, suggested by rational models of reference form, is that the production of reduced forms incurs a lower cognitive cost than the production of explicit forms (e.g., frank & goodman, 2012; tily & piantadosi, 2009). one of the few studies to address the relation between utterance planning and reference form is arnold & nozari (2017). in their study, speakers described shapes moving on screen, e.g. the yellow pentagon flashes. then it jumps over the yellow square. they found that for the sham subjects, who were not experiencing anodal transcranial direct current stimulation, pronouns and zeros were used more often when the timing of the stimuli and response enabled the speaker to do more pre-planning, i.e., when there was no gap between utterances, and when the speaker was planning one utterance while describing the previous event. in this task, these conditions also promoted other measures of discourse connectivity, such as the use of connectors and or then. however, this study used discourse contexts with relatively little semantic information, and did not address questions of how planning relates to predictability. zerkle & arnold 38 the current study uses an eyetracking method to test whether planning facilitation explains the role of predictability on reference production. on the planning facilitation hypothesis, thematic roles promote the status of predictability, which in turn triggers earlier conceptual activation of the referent, making it more accessible for faster planning, which leads to increased pronoun choice. in sum, the main question we will be asking in the current study is whether or not planning mechanisms can explain the predictability effect on reference form choice, and we will be using fine-grained measures of utterance planning with eye movements to test this. we will examine anticipatory looks to the target panel in two broad regions, representing relatively early and late planning windows. 1.3 traditional models of utterance planning traditional models of utterance planning generally agree that the generation of an utterance is incremental; beginning with the transformation of a communicative intention into a preverbal conceptual message. the speaker then encodes the grammatical formulation of this message (which includes lexical/lemma selection and functional assignment), phonological encoding and syllabification occurs, and then the phonemes are encoded and articulated (bock, irwin, davidson, & levelt, 2002; ferreira & swets, 2002; griffin & bock, 2000; levelt, 1989; levelt, roelofs, & meyer, 1999; wheeldon, 2013; see figure 3). visual input can affect the early conceptual message level, such as in picture naming tasks (indefrey & levelt, 2004; roelofs, 1992). figure 3. model of utterance planning recreated from levelt, roelofs, & meyer (1999). following the general model in figure 3, schmitt, meyer, & levelt (1999) built a model of planning that is specific to reference production. they proposed that speakers must fit each message fragment into the current discourse, which involves marking certain concepts as old does planning explain why predictability affects reference production? 39 information (in focus), and others as new (not in focus; chafe, 1976).1 these types of information can be referred to linguistically in different ways, from explicit noun phrases to reduced pronouns. such marking helps speakers and listeners maintain a coherent discourse record (levelt, 1989). in order to access pronouns, lexical concepts of the entities are marked as being in focus or not in focus depending on the discourse context. a pronoun is selected if and only if the corresponding lexical concept is accessible and marked as in focus within the discourse; otherwise, a noun is selected. the authors propose that a referent’s discourse accessibility status affects the message at the conceptual level; this accessibility node then affects the anaphoric link to lexical selection, in this case whether a reduced form vs. an explicit form is chosen (see figure 4). however, this model did not address questions about the timing of utterance planning, and whether planning is faster for accessible referents. neither does it include consideration of predictability as a component of discourse status. this raises questions about whether planning facilitation has any effect at all on reference production. figure 4. an english-language adaptation of the model of reference production from schmitt, meyer, & levelt (1999). 1.4 does planning explain thematic role effects on reference form? the question behind our study is whether planning explains the relation between thematic roles and reference production choices. we have preliminary evidence about this question from previous studies (rosa & arnold, 2017; zerkle, rosa & arnold, 2015), which used response latency as a proxy measure of planning time. in the storytelling task described above, participants hear the first sentence, and then provide a second sentence by describing the second panel of the cartoon picture. the time to respond provides a rough measure of the time needed to plan the response. some planning could take place earlier, as participants listened to the first sentence. however, there was variable time to respond across trials, suggesting that considerable planning occurred immediately before the utterance. 1 note that on this terminology, “in focus” refers to information that is salient or accessible. this contrasts with the linguistic term focus, which contrasts with topic and is associated with new information. zerkle & arnold 40 in all three experiments, they found that speakers initiated goal continuations faster than source continuations, supporting the hypothesis that predictability affects the time course of utterance planning (see figure 5). response latency was measured from the offset of the initially heard description to the onset of the speaker’s fluent speech. 2 figure 5. main effect (indicated by asterisks) of goal continuation on response latency in rosa & arnold (2017) and zerkle, rosa, & arnold (2015, exp. 1); numerical trend in zerkle, rosa, & arnold (2015, exp. 2). error bars are standard error of the subject means by condition. in sum, goal continuations both yield faster response latencies and are more likely to involve pronouns. the next question is whether response latency itself affects reference form choice. yet all three of these experiments found that latency does not predict reference form – as latency to begin speaking increased, there was no significant change in the rate of pronoun/zero choice in any of these three experiments (analyzed separately). these findings provide strong evidence against one sort of hypothesis about planning time. however, planning in a discourse context can often occur earlier, while hearing/producing earlier utterances. for example, in rosa & arnold (2017, exp. 1), participants could see both panels throughout the trial, allowing earlier planning. in both of the zerkle et al. (2015) experiments, participants did not see the second panel until after the first sentence, but they had previewed the stimulus stories before the production task, which means they may have recalled the target sentence from memory before it appeared. both of these tasks mirror the conditions of real life production, where the speaker typically knows what they are going to say before they say it. the current study tests whether earlier measures of planning pattern with reference form choices, and whether this explains the effects of predictability. 1.5 current study goals the current study tests the relation between utterance planning, predictability, and reference form. even though latency did not predict reference form choice, eyetracking may provide an earlier and more sensitive measure for identifying planning effects on reference production. the 2 in rosa & arnold (2017), this effect of predictability on latency occurred even after controlling for the length of the context sentence (arnold, 2017), which may have affected participants’ ability to pre-plan their responses. in the experiments from zerkle, rosa & arnold (2015), the stimulus picture did not appear until after the context sentence, so the duration of the context sentence is not as relevant. see zerkle, rosa, & arnold (2017) for a discussion of differences in planning across these two studies. 0 400 800 1200 1600 2000 subject non-subject subject non-subject subject non-subject rosa & arnold (2017) zerkle, rosa, & arnold (2015, exp1) zerkle, rosa, & arnold (2015, exp2) re sp on se la te nc y (m s) goal source * * * * does planning explain why predictability affects reference production? 41 conceptual planning of a message could begin earlier than the latency period that was used in previous studies. indeed, speakers often pre-plan their utterances (e.g, ferreira & swets, 2002), at least at the message level (brown-schmidt & konopka, 2015). if speaking about predictable entities leads to greater pre-planning than speaking about unpredictable entities, it might explain the link between predictability and reference production. to address this, in the current study we used eye tracking to each visual scene as a proxy for planning. previous research has suggested that fixations to visual stimuli reflect the conceptual and linguistic planning that occurs before utterance production. in a simple picture-description task, trial-initial looks signify rapid planning: speakers may direct their attention to the part of the displayed picture that represents the first referent in their utterance (bock et al., 2003). speakers also tend to look at objects immediately before naming them (griffin and bock, 2000). in contrast, van der meulen et al. (2001) found that speakers do not always look at the to-be-named objects: speakers allocated less visual attention to given objects than to new ones, and to objects they would later refer to with a pronoun than with a full noun phrase. this suggests that discourse focus (given vs. new) and referential form (pronoun vs. name) modulate the amount of visual attention allocated to a referent during utterance planning. here we test whether planning, as measured by anticipatory eye movements, could be related to both reference form and predictability. this is the first study to examine these issues within a discourse context, so we do not have a priori expectations about how these are related. current theories might suggest that accessible referents would lead to faster planning (e.g., bell et al., 2009), but this question has not been addressed explicitly. we will be using looks to panel 2 (the right-most picture, figure 1) as the metric of early planning, because looks ahead indicate that the speaker is attending to this event and preparing their own upcoming description turn. in all critical items, this event shows the continuing character (the target referent), as well as an action that they are performing. this motivates our use of the entire picture (as opposed to just the character) as our region of interest here. additionally, there is enough variability in our visual stimuli that this broad region of interest is the most appropriate. we will also examine fixations to panel 2 in two different windows of time, which represent early vs. late planning. we speculate that when a speaker is still listening and comprehending the first sentence, they may likely be conceptually planning parts of their own upcoming utterance at this point. when this first sentence is finished there is a period of time that is silent right before the speaker begins speaking, and this may be where more lexical formulation planning occurs. these two periods are based on traditional models of utterance planning (e.g., griffin & bock, 2000; meyer, sleiderink, & levelt, 1998; wheeldon, 2013); however, we cannot link them directly to phases of production (message conceptualization vs. lexical formulation). importantly, this is the first investigation into the time course of planning and how it relates to both thematic role predictability and reference form choice. to test the planning facilitation hypothesis, we will compare effects in both early and late stages of planning to determine if and where planning facilitation occurs, and how it relates to reference form. one challenge with studying language planning is that we don’t know precisely when planning begins, and thus, the direction of causation is not clear. if speakers plan a pronoun early, this choice might guide their looks during the planning periods. alternatively, if speakers happen to look more at panel 2, this attention itself might drive the choice of pronoun. nevertheless, we think it is plausible that events occurring earlier in time are likely to influence events occurring later in time. we therefore take this as our starting point. our analyses look first at the effect of the discourse context, which occurs earlier than the participants’ response. how does the context affect both planning measures and utterance form? then we look at the effect of planning, which zerkle & arnold 42 is logically prior to the reference, and ask whether either early or late planning measures predict variation in reference form. a summary of our predictors and measures is shown in figure 6. using a variation on rosa & arnold (2017)’s paradigm, we expect to replicate the effect of thematic roles on reference form. we additionally test three measures of planning: early looks, late looks, and latency. these measures allow us to test three questions: • question 1. does thematic role predict production planning measures (latency, early looks, and late looks)? • question 2. does each of the production planning measures predict reference form choice? • question 3: can the predictability effect on reference form be explained in terms of utterance planning? figure 6. visual summary of the measures used. if thematic role affects planning, we expect the goal continuation contexts to elicit shorter latencies (as in previous studies), and perhaps more looks to the target panel during either early or late planning regions. if planning facilitation drives reference form choice, we expect reference form to vary as a function of either latency or anticipatory fixations. finally, if planning explains predictability effects on reference form, then we would expect planning to have a stronger effect on reference form than thematic role. planning effects may occur for early measures of planning, late measures of planning, or both. 2 methods 2.1 participants 54 native english speakers at the university of north carolina at chapel hill participated for course credit. 13 of these participants used only np descriptions to refer to the target on critical trials, so these participants were excluded from analyses, leaving 41 participants with reference form variation. 6 of these participants were excluded due to mis-calibration of the eyetracker, and 1 participant was excluded for audio technical failure. this left 34 participants in the final analyses (19 females, and 15 males). 2.2 materials and design a narrative production paradigm designed by rosa & arnold (2017) was modified for use with eyetracking for the current study (for stimuli see http://jaapstimuli.web.unc.edu/transfer-verbstimuli-and-paradigm/). in the primary task, participants viewed pairs of pictures (like figure 1) and listened to a sentence describing the first (left) picture. their task was to describe the second does planning explain why predictability affects reference production? 43 picture, and fixations to the first panel represented attention to the discourse context, while fixations to the second panel represented early planning of the target utterance. as in rosa and arnold’s task, participants first watched a background video that described the story context. they were given the role of tabloid photographer, and informed that they witnessed a murder in a mansion and they happened to capture pictures of the events from that day. they learned about six characters: three males (the duke, the butler, the driver) and three females (the duchess, the maid, the chef). their task was to describe these pictures to a “detective” in order to help solve the crime. one critical change in this study was that the detective’s sentences were recorded, instead of spoken live (as in rosa & arnold, 2017). the recorded voice described the first of each picture pair, and participants were encouraged to “continue the story” when describing the second of each picture pair. another critical change was that the instructions emphasized storytelling. in a similar task, zerkle & arnold (2016) found that some participants ignored the context and produced only descriptive noun phrases. here we emphasized storytelling in order to encourage more connected discourse, and thus elicit greater variation in referential forms. the experimenter instructing each participant also gave two example trials (from filler trials of the main experiment), in which they used both connector words and pronouns in each continuation. the story in the main experiment consisted of 53 “evidence photos”: 29 fillers and 24 critical items. in the critical items, the first picture and sentence described a transfer event between two characters (from a source to a goal). the second panel only pictured the target character, which was either the goal or the source of the previous sentence (manipulated between items). the grammatical placement of the target character in the detective’s sentence was manipulated between-subjects, such that half of the participants heard, “the duchess handed a painting to the duke” and half heard, “the duke received a painting from the duchess.” thus there were four possible conditions of character continuation: goal/subject, source/subject, goal/non-subject, source/non-subject. however, the trials from both non-subject conditions were not included in analyses here (see section 3.2 for explanation). filler trials pictured between one to three characters. all trials were presented in the same order for all participants to create a coherent story sequence. character gender was controlled, such that half of the critical trials had two characters of the same gender, and half had two characters of different gender. 2.3 procedure research assistants fitted each participant with an eyelink ii head-mounted eyetracker (sr research) and a headset headphones/microphone. first participants viewed a slideshow in powerpoint, which played the background video describing the story and characters, and then they viewed all 53 picture pairs silently in order (5s per pair). the purpose of this initial preview was to familiarize the participant with the series of the events and the outcome of the story. this imitates the characteristics of natural language production, in which speakers usually convey information they already know. after this, they heard two sample trials from the experimenter, and were instructed to continue the story as much as possible when they described each picture. calibration of the eyetracker occurred right before the main experiment began. before each trial, participants viewed a drift correct screen and pressed the space bar to continue. then the pair of pictures appeared on the screen next to each other and remained there for the duration of the trial. after a preview period (average of 665ms), the context detective sentence played over the headphones. participants spoke their description of the second picture into the microphone, and then used the mouse to click on a green circle in the bottom right corner to move on to the next zerkle & arnold 44 trial, which began with another drift correct screen. see figure 7 for a schematic diagram of the preview and trial procedures. figure 7. schematic diagrams of the preview procedure and an example trial procedure. 3 results and discussion in this section, we first describe the general procedure we used to analyze the data from this study. we then report the linguistic, planning, and reference form results. 3.1 general analytic procedure 3.1.1 response coding trials were excluded if the participant did not refer to the target character in the subject position of their utterance. we excluded trials in which the speaker referred to the wrong character in the response (n=9), leaving 373 trials to be audio coded. research assistants transcribed utterances and coded responses for type of reference (pronoun, zero, np description) and use of connector before the reference (e.g., and, then, and then, now). for the analyses, we only analyzed the subject continuation trials (n=382), excluding the non-subject continuation trials (n=379) (see section 3.2). 3.1.2 audio data coding four research assistants analyzed the audio files in praat (boersma & weenink, 2015). they segmented each trial in order to measure latency to begin speaking, which was calculated as the time in milliseconds between the end of the context sentence and the beginning of fluent speech. any disfluencies prior to the response were considered as part of planning time, and thus were included in this latency measure (n=10). outliers were excluded, where outliers were trials where the latency was less than or greater than 3 standard deviations from the grand mean (n=3). we also calculated track loss on a trial level, where trials with greater than 33% of looks in either time window were classified as no data (blinks, looks not to the computer screen, etc.). there does planning explain why predictability affects reference production? 45 were 4 trials with < 33% track loss in the ds window, and 5 trials in the latency window. thus, 361 trials were included in statistical analyses. 3.1.3 dependent measures our analysis of reference form examined the binary contrast between reduced (pronoun and zero) and unreduced (description) forms. pronouns and zeros were combined because they play a similar pragmatic role, and are both more common when the referent is accessible or predictable. zeros in our task represent cases of syntactic coordination, e.g. the duke received the basket from the duchess…. and ø threw it down the hallway. some studies instead exclude zero continuations from analysis (e.g., rosa & arnold, 2017), especially when they are rare in the dataset. here, however, doing so would exclude around 25% of the data, and misrepresent the rate of using reduced expressions. moreover, an examination of pronoun vs. zero trials suggests that this decision did not change our findings, in that numerically similar patterns were found for pronouns and zeros. if zeros are excluded from the analyses, similar results obtain. our analysis of latency used the log of the latency as a dependent measure. for analyses of gaze, eye position was sampled every 4 ms, but the data were converted to samples at every 20 ms prior to analyses for faster processing (mcmurray, 2002). eye movement data is presented in terms of looks, where a look is defined as a fixation grouped together with the prior saccade. this practice is often used in language processing studies, where listeners are directing their attention to one referent and not another (e.g., arnold, hudson-kam, & tanenhaus, 2007; arnold & lao, 2015; huettig, rommers, & meyer, 2011; mcmurray, tanenhaus, & aslin, 2009). we used two areas of interest: panel 1 and panel 2 (see figure 1). ports used for analyses were rectangles exactly around each picture. any looks to areas besides these were considered “other.” average looks are presented graphically here for ease of interpretation, but empirical logits of looks were used in all statistical models. empirical logits were calculated for both panels using the following formula based on barr’s (2008) methods: 𝐿𝑂𝐺 $ 𝑆𝑢𝑚 𝑜𝑓 𝑙𝑜𝑜𝑘𝑠 𝑡𝑜 𝑃𝑎𝑛𝑒𝑙 2 + 0.5 𝑇𝑜𝑡𝑎𝑙 𝑠𝑢𝑚 𝑜𝑓 𝑙𝑜𝑜𝑘𝑠 − 𝑆𝑢𝑚 𝑜𝑓 𝑙𝑜𝑜𝑘𝑠 𝑡𝑜 𝑃𝑎𝑛𝑒𝑙 2 + 0.5 : 3.1.4 model-building procedure generalized linear mixed effects models were used to account for dependencies in repeated measures using sas 9.4. proc glimmix was used for analyses of dichotomous dependent variables with a logit link, i.e. for all analyses of reference form. proc mixed was used for analyses of continuous dependent variables, i.e. for all analyses of eye gaze and response latency. we constructed models of each dependent variable with random intercepts of participant and item to account for nesting. binary predictors were effects coded, such that each variable as reported in the models below are comparison group vs. reference group. all predictor variables were grandmean centered. random slopes of all included variables by participant and by item were included when appropriate, but if a slope was estimated to be zero it was excluded (searle et al., 1992). the final fixed and random effects for each model are reported below each table. 3.2 linguistic measures speakers rarely used reduced expressions for non-subject continuation trials (10%), so only subject continuation trials were included in all further analyses. therefore, all subsequent analyses examined only the subject trials. see figure 8 for a breakdown of pronoun/zero trials by non-subject and subject conditions. zerkle & arnold 46 within the subject continuation trials, the descriptive statistics by thematic role condition are reported in table 1. variable averages: goal condition source condition overall detective sentence length (ms) 2453.73 2242.10 2351.14 latency to begin speaking (ms) 1454.71 1670.13 1559.14 pronoun/zero use 64.52% 45.14% 55.12% pronoun use 37.10% 22.29% 29.92% zero use 27.42% 22.86% 25.21% looks to p2 in ds window 47.65% 43.22% 45.50% looks to p2 in latency window 75.88% 68.14% 72.13% table 1. descriptive statistics by goal/source condition, subject continuation trials only. percentages represent the total number of trials in which an event occurred in each thematic role condition (and overall). first, we built a model of pronoun/zero use that showed that speakers used significantly more pronouns/zeros for goal continuations, (see table 2, figure 8). this finding replicated the effects found in both rosa & arnold (2017) and zerkle, rosa, & arnold (2015). effect estimate se df t-value p-value goal vs. source continuation 1.1739 0.4486 19.88 2.62 0.0166 table 2. critical predictor of pronoun/zero use*. * random intercepts of participant and of item. figure 8. reference form choice rates by grammatical role and thematic role. this includes 373 trials in the subject conditions, and 364 in the non-subject conditions. error bars are standard error of the subject means by condition. in the first model, the main effect of thematic role is seen in the left bars (only using subject continuation trials). 3.3 planning measures our first question was whether thematic roles would predict variability in each of our three measures of planning: response latency, the speaker’s gaze to panel 2 during the context sentence, 65% 9% 45% 11% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% subject non-subject % p ro no un /z er o goal source does planning explain why predictability affects reference production? 47 and speaker’s gaze to panel 2 during the latency period. looks to panel 2 (which guided the response) were considered evidence of response planning (either message planning or formulation). following zerkle & arnold (2016), we divided each trial into four different time windows (see figure 9), but our analyses for this study focused primarily on the detective sentence window and latency window, as defined below. time windows were calculated starting at the estimated average stimulus image onset. due to a programming error, there was some variability in the image onset, which was estimated to occur an average of 110 ms after participants responded to the drift correct screen, with a standard deviation of ± 31 ms. the onset of the detective sentence was much less variable, occurring at an average of 665 ms after the estimated picture onset, with a standard deviation of ± 3 ms. • the detective sentence (ds) window: the time between the onset (665ms from stimulus onset) and offset of the context detective sentence. • the latency window: the time between the offset of the detective sentence and the onset of the participant’s fluent speech. figure 9. this figure shows the four time windows of a trial, averaged over all subject continuation items. time is on the x-axis, starting at the beginning of the trial for the left graph, and centered around the onset of participants’ response in the right graph. the gray box in the latency window represents the overall average response latency length. since we are interested in planning effects, analyses focused on the regions before the response. we have divided these measures into early planning: looks to panel 2 during the detective sentence window; and late planning: looks to panel 2 during the latency window, and the length of response latency (ms). we test these each in separate mixed effects regression analyses. 3.3.1 early planning speakers were not significantly more likely to look at panel 2 during the detective sentence window for goal continuations (see table 3, figure 10). this shows that in answer to question 1, we did not see strong effects of thematic role predictability on early planning. effect estimate se df t-value p-value goal vs. source continuation 0.1618 0.107 22 1.51 0.1449 table 3. critical predictors of looks to panel 2 in the detective sentence window*. * random intercepts of participant and of item, random slope of goal continuation by subject. zerkle & arnold 48 figure 10. average looks to panel 2 during the detective sentence region by thematic role condition. error bars are standard error of the subject means by condition. 3.3.2 late planning response latency was significantly predicted by thematic role, replicating findings from previous experiments using this task (see table 4, figure 11; rosa & arnold, 2017; zerkle, rosa, & arnold, 2015). effect estimate se df t-value p-value goal vs. source continuation -0.06905 0.02803 359 -2.46 0.0142 table 4. critical predictors of log latency to begin speaking*. * random intercepts of participant and of item, random slope of goal continuation by subject. figure 11. thematic role effect on response latency. error bars are standard error of the subject means by condition. 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 goal source a ve ra ge lo ok s t o p2 in d s w in do w 0 200 400 600 800 1000 1200 1400 1600 1800 2000 goal source la te nc y (m s) does planning explain why predictability affects reference production? 49 our next question was whether panel 2 looks during the latency period, which represent a later planning measure, would reflect thematic role condition. we found that indeed, they were significantly influenced by the goal/source manipulation (see table 5, purple lines in right panel of figure 9). speakers looked more at panel 2 in this window for goal continuations than for source continuations. effect estimate se df t-value p-value goal vs. source continuation 0.2756 0.09201 21.4 3 0.0068 table 5. critical predictors of looks to panel 2 in latency window*. * random intercepts of participant and of item. in sum, analyses of our three planning measures show that thematic role influences the duration of the latency period and looks to panel 2 during the latency period, but not looks to panel 2 during the context sentence. thus, in answer to question 1, we found that thematic role affected only late measures of planning, and not our early measure of planning. this suggests that predictability has an effect on utterance planning, but only immediately before utterance articulation. 3.4 what explains referential form choice? our second question was whether our planning measures would help explain speakers’ choices of referential form. to assess this, we examined the three dependent measures from the previous section (early looks, late looks, and latency duration). we first tested the correlations between these three variables (using the transformed values) to assess multicollinearity. early looks and late looks were marginally positively correlated, (r=0.10, p=0.052); early looks and latency duration were significantly negatively correlated (r=-0.23, p<.0001); and late looks and latency duration were significantly negatively correlated (r=-0.17, p=0.001). all three of these correlations reflect only small effect sizes (cohen, 1988), given our large sample sizes. this suggests that multicollinearity is not an issue in a model with all three variables.3 as shown in table 6, we found that only early looks significantly predicted reference form choice, such that more looks to panel 2 in the ds window increased the likelihood of using a pronoun/zero (see table 6). effect estimate se df t-value p-value looks to p2 in ds window 0.5813 0.2666 357 2.18 0.0299 looks to p2 in latency window -0.04723 0.187 41.46 -0.25 0.8018 log latency (ms) -0.8283 0.7756 357 -1.07 0.2863 table 6. utterance planning predictors of pronoun/zero use*. *random intercepts of participant and of item, random slopes of latency by subject and by item. 3.5. do planning measures account for the effect of thematic role on reference choices? our final question was how these planning measures, together with our thematic role manipulation, account for choices in reference form. specifically, we wanted to know whether the planning measures would account for the effect of thematic role. we focused on early looks as a planning measure, since this was the only one to predict reference form choice in the previous model. we therefore built a model with reference form (pronoun/zero vs. description) as the 3 separate models were also run with each individual predictor, and the results from each showed a similar pattern as the full model (only early looks significantly predict reference form). zerkle & arnold 50 dependent measure, and with three predictors: thematic role predictability, early looks, and the interaction between the two. thematic role and early looks were significantly positively correlated, however this significance is contingent upon the large sample size and the effect size is rather small (r=0.124, p=0.019) (cohen, 1988). as shown in table 7, we again found a significant effect of goal continuation (more reduced forms for goals); and also a significant effect of looks to panel 2 in the ds window (more reduced forms for trials with more anticipatory looks). the interaction between thematic role and looks was not significant. however, it is not the case that the thematic role effect can be entirely explained in terms of early planning. if it were, we would see the thematic role effect disappear in the presence of the ds-window looks predictor. thus, in answer to our last question, planning does not account for the bias to use more reduced expressions for goals than sources. instead, we see two independent effects of planning and thematic role. effect estimate se df t-value p-value goal vs. source continuation 1.173 0.4238 20.09 2.77 0.0118 looks to p2 in ds window 0.7468 0.2912 357 2.56 0.0107 goal continuation*looks to p2 ds 0.7385 0.5503 357 1.34 0.1804 table 7. critical predictors of pronoun/zero use, including eye movement predictors*. * random intercepts of participant and of item. 3.6. empirical summary in sum, we found that goal thematic roles lead to both more reduced forms, and evidence of greater planning during the latency period, as indicated by both shorter latencies and greater anticipatory looks to the target panel. in addition, we found that one of our planning measures (early anticipatory looks) predicted the use of reduced forms. however, these findings are critically independent. the effect of thematic roles on planning occurred only for the late planning measures, whereas the effect of planning on reference form only occurred for the early planning measure. in addition, including both early looks and thematic role as predictors of reference form revealed that they are independent predictors. 4 general discussion in sum, this study revealed three main findings: a) thematic role predictability affects late measures of planning, b) early planning supports the use of reduced forms; and c) planning does not account for predictability effects on reduced forms. we will discuss each in turn. 4.1 thematic role predictability affects late measures of planning in answer to question 1, this study found that thematic role predictability only affected planning during the latency window and the latency measure itself. this is consistent with earlier reports that response latencies are shorter for goal vs. source continuations (rosa & arnold, 2017; zerkle, rosa, & arnold, 2015). we found that eye movements during the latency window reflected thematic role predictability, such that speakers looked more at the target stimulus (panel 2) in the goal condition than in the source condition. in other words, only at this late period of planning does the predictability of goal continuations increase looks to the target picture. it may be that speakers are more likely to look back to panel 1 for source continuations in this window because these items require additional comprehension of the context (beyond the length of the detective sentence) before planning. 4.2 early planning supports the use of reduced forms does planning explain why predictability affects reference production? 51 in answer to question 2, we found that only early planning supports the use of reduced referential forms. importantly, this study is the first to examine reference production within a discourse context. van der meulen et al. (2001) also examined eye movements during the production of pronouns vs. noun phrases, but their study didn’t test natural language production within a discourse context. we found that early planning affects reference form (both pronouns and zeros) within a natural discourse context. this finding is consistent with arnold & nozari (2017) who also found that planning supported pronoun/zero use, in a different paradigm. this finding is consistent with classical models about reference form, which suggest that the discourse context influences decisions about when to use a pronoun. but critically, this is one of the first studies to link these decisions with inter-trial variation in production planning processes. when speakers are listening to the detective sentence, sometimes they begin planning earlier, as indicated by looks to panel 2. those are the trials where they tend to use pronouns and zeros. it may be that advance planning is especially helpful when it co-occurs with processing the discourse context (arnold & nozari, 2017). doing so helps establish links between the events, and increases the speaker’s desire to communicate these links with their linguistic choices, like the use of pronouns. 4.3 early planning does not account for predictability effects in answer to question 3, we found that early planning does not account for the effects of thematic roles on reference form. this is not consistent with the idea that early planning is the sole driving force. the relationship between predictability, planning, and reference form is complex, but this is the first step towards parsing the functions of early vs. late planning mechanisms for the production of reference form. in sum, planning facilitation is not the reason why predictability affects reference form, so there must be another characteristic of predictability that is responsible for the pattern of behaviors observed here. 4.4 conclusions the three main findings from this study rule out a production processing-based account, because planning facilitation does not directly account for pronoun choice. what, then is the likely explanation for predictability effects on reference form? our speculations on this question build on the frequent observation that referential forms are selected on the basis of the referent’s cognitive status, i.e. its discourse salience or accessibility (ariel, 1990; chafe, 1994; givón, 1983; gundel, hedberg, & zacharaski, 1993). this classic view suggests that reduced forms are triggered by the representation of the discourse entity itself (arnold, 2016). this view is consistent with our findings that early planning supports the use of pronouns and zeros, because we assume that participants are building a discourse representation of each story during the first part of the trial. when that representation is strong, speakers may have greater mental resources to begin pre-planning, and therefore look to panel 2 earlier. thus, pre-planning does affect reference form, but only pre-planning that occurs during the first part of the trial. by contrast, the effect of thematic role predictability emerges later. we only see the effect of thematic roles on later measures of planning, and these planning measures are unrelated to reference form. instead, thematic role must affect reference form through some other mechanism. the thematic role effect points to a property of references that derives from their participation in events, as opposed to questions about how activated or accessible each referent is alone. this raises the possibility that pronouns and zeros are used in cases where there are strong relationships between utterances. that is, goal references are predictable, and as a result are more tightly connected at a conceptual level. in support of this, rosa & arnold (2017) asked a group of subjects to rate the pairs of events in their stimuli (also used in the current study). they zerkle & arnold 52 found that the two events were perceived as more related in goal continuation than source continuation items. this suggests that speakers may perceive the goal continuations as more connected, which may have downstream effects on reduced form use. further research is needed to better understand the relationship between thematic role predictability and reference form choice. nevertheless, the current findings support the importance of early planning on discourse production. when speakers pre-plan, they are likely to connect events conceptually. in an unrelated finding, thematic roles may also lead to greater conceptual connection, but this effect is not driven by early planning. both situations lead to the selection of linguistic forms that signal coherence, such as pronouns and zeros. acknowledgements this work was funded by the national science foundation under grant 1348549 to jennifer e. arnold. all procedures were performed in compliance with relevant laws and institutional guidelines, and the university of north carolina at chapel hill institutional review board has approved them. we would also like to thank ana medina fetterman, liz reeder, samuel adam smith, jacob pascual, jenna roller, brianna torres, and kristen bubak for their help preparing the stimuli and collecting and coding the data. references ariel, m. (1990). accessing noun-phrase antecedents. london: routledge. ariel, m. (2001). accessibility theory: an overview. in sanders, t., schliperoord, j. and spooren, w. (eds.), text representation. amsterdam: john benjamins (human cognitive processing series), pp. 29-87. arnold, j. e. (1998). reference form and discourse patterns. doctoral dissertation, stanford university, stanford, california. arnold, j. e. (2001). the effect of thematic role on pronoun use and frequency of reference continuation. discourse processes, 31, 137–162. https://doi.org/10.1207/ s15326950dp3102_02 arnold, j. e. (2010). how speakers refer: the role of accessibility. language and linguistic compass, 4, 187–203. doi: https://doi.org/10.1111/j.1749-818x.2010.00193.x arnold, j. e. (2016). explicit and emergent mechanisms of information status. topics in cognitive science, 8, 737–760. http://doi.org/10.1111/tops.12220 arnold, j. e. (2017). latency analysis for rosa & arnold (2017). technical report #1. unc language processing lab, department of psychology & neuroscience, university of north carolina – chapel hill, chapel hill, north carolina. http://arnoldlab.web.unc.edu/files/2017/08/arnold_techreport1_2017.pdf arnold, j. e., & griffin, z. m. (2007). the effect of additional characters on choice of referring expression: everyone counts. journal of memory and language, 56(4), 521-536. https://doi.org/10.1016/j.jml.2006.09.007 arnold, j. e., kam, c. l. h., & tanenhaus, m. k. (2007). if you say thee uh you are describing something hard: the on-line attribution of disfluency during reference comprehension. journal of experimental psychology. learning, memory, and cognition, 33(5), 914–930. http://doi.org/10.1037/0278-7393.33.5.914 arnold, j. e., & lao, s.-y. c. (2015). effects of psychological attention on pronoun comprehension. language, cognition and neuroscience, 30(7), 832–852. http://doi.org/10.1080/23273798.2015.1017511 does planning explain why predictability affects reference production? 53 arnold, j. e. & nozari, n. (2017). the effects of utterance planning and stimulation of left prefrontal cortex on the production of referential expressions. cognition, 160, 127–144. https://doi.org/10.1016/j.cognition.2016.12.008 arnold, j. e. & watson, d. g. (2015). synthesizing meaning and processing approaches to prosody: performance matters. language, cognition, and neuroscience, 30, 88-102. https://doi.org/10.1080/01690965.2013.840733 arnold, j. e., & zerkle, s. a. (2019). why do people produce pronouns? pragmatic selection vs rational models. language, cognition and neuroscience. https://doi.org/10.1080/23273798.2019.1636103 barr d. j. (2008). analyzing ‘visual world’ eyetracking data using multilevel logistic regression. journal of memory and language, 59(4). 457–474. bell, a., brenier, j. m., gregory, m., girand, c. & jurafsky, d. (2009). predictability effects on durations of content and function words in conversational english. journal of memory and language, 60(1), 92–111. https://doi.org/10.1016/j.jml.2008.06.003 boersma, p. & weenink, d. (2015). praat: doing phonetics by computer [computer program]. version 5.4.09, http://www.praat/org/. bock, k., irwin, d. e., davidson, d. j., & levelt, w.j. m. (2003). minding the clock. journal of memory and language, 48(4), 653–685. https://doi.org/10.1016/s0749-596x(03)00007-x brown-schmidt, s., & konopka, a. e. (2015). processes of incremental message planning during conversation. psychonomic bulletin & review, 22(3), 833–843. http://doi.org/10.3758/s13423-014-0714-2 chafe, w. (1976). givenness, contrastiveness, definiteness, subjects, topics, and point of view. in charles n. li (ed.), subject and topic, (pp. 25-56). new york: academic press inc. chafe, w. (1994). discourse, consciousness, and time: the flow and displacement of conscious experience in speaking and writing. chicago, il: chicago university press. cohen, j. (1988). statistical power analysis for the behavioral sciences (2nd ed.). hillsdale, nj: erlbaum. dowty, d. (1991). thematic proto-roles and argument selection. language, 67(3), 547-619. ferreira, f. & swets, b. (2002). how incremental is language production? evidence from the production of utterances requiring the computation of arithmetic sums. journal of memory and language, 46(1), 57–84. https://doi.org/10.1006/ jmla.2001.2797 fowler, c. a. & housum, j. (1987). talkers’ signaling of “new” and “old” words in speech and listeners’ perception and use of the distinction. journal of memory and language, 26, 489–504. https://doi.org/10.1016/0749-596x(87)90136-7 frank, m. c., & goodman, n. d. (2012). predicting pragmatic reasoning in language games. science, 336(6084), 998-998. fukumura, k. & van gompel, r. p. g. (2010). choosing anaphoric expression: do people take into account likelihood of reference? journal of memory and language, 62, 52–66. https://doi.org/10.1016/j.jml.2009.09.001 givón, t. (1983). topic continuity in discourse: an introduction. in: givón, t. (ed.), topic continuity in discourse: a quantitative cross-language study, 1–41. amsterdam: john benjamins. https://doi.org/10.1075/tsl.3 griffin z. m. & bock k. (2000). what the eyes say about speaking. psychological science, 11(4). 274–279. grosz, b. j., weinstein, s., & joshi, a. k. (1995). centering: a framework for modeling the local coherence of discourse. computational linguistics, 21(2), 203-225. gundel, j. k., hedberg, n., & zacharski, r. (1993). cognitive status and the form of referring expressions. language, 69, 274-307. zerkle & arnold 54 hale, j. (2001). a probabilistic earley parser as a psycholinguistic model. naacl ’01: second meeting of the north american chapter of the association for computational linguistics on language technologies 2001, 1–8.: https://doi.org/10.3115/1073336.1073357 huettig, f., rommers, j., & meyer, a. s. (2011). using the visual world paradigm to study language processing: a review and critical evaluation. acta psychologica, 137(2), 151–171. http://doi.org/10.1016/j.actpsy.2010.11.003 indefrey, p., & levelt, w. j. (2004). the spatial and temporal signatures of word production components. cognition, 92(1), 101-144. jurafsky, d., bell, a., gregory, m., & raymond, w. (2001). probabilistic relations between words: evidence from reduction in lexical production. in j. bybee & p. hopper (eds.), frequency and the emergence of linguistic structure (pp. 229254). amsterdam, netherlands: john benjamins publishing company. kehler, a. & rohde, h. (2013). a probabilistic reconciliation of coherence-driven and centeringdriven theories of pronoun interpretation. theoretical linguistics, 39(1–2), 1–37. https://doi.org/10.1515/tl-2013-0001 kehler, a., kertz, l., rohde, h. & elman, j. l. (2008). coherence and coreference revisited. journal of semantics, 25(1), 1–44. https://doi.org/10.1093/jos/ ffm018 kravtchenko, e.; modi, a.; demberg v.; titov, i., & pinkal. m. (2017). does referent predictability affect rate of pronominalization? paper presented at the cuny conference on human sentence processing, march 2017. levelt, w. j. m. (1989). speaking: from intention to articulation. cambridge, ma: mit press. levelt, w. j., roelofs, a., & meyer, a. s. (1999). a theory of lexical access in speech production. behavioral and brain sciences, 22(1), 1-38. levy, r., & jaeger, t. f. (2007). speakers optimize information density through syntactic reduction. advances in neural information processing systems, 19, 849. mahowald, k., fedorenko, e., piantadosi, s. t., & gibson, e. (2013). info/information theory: speakers choose shorter words in predictive contexts. cognition, 126(2), 313-318. mcmurray, b. (2002). analysis scripts written in microsoft access: eyelinkanal [computer program]. version 1.5. mcmurray, b., tanenhaus, m. k., & aslin, r. n. (2009). within-category vot affects recovery from “lexical” garden-paths: evidence against phoneme-level inhibition. journal of memory and language, 60(1), 65–91. http://doi.org/10.1016/j.jml.2008.07.002 meyer, a. s., sleiderink, a. m., & levelt, w. j. m. (1998). viewing and naming objects: eye movements during noun phrase production. cognition, 66(2), b25–b33. http://doi.org/10.1016/s0010-0277(98)00009-2 nappa, r., & arnold, j. e. (2014). the road to understanding is paved with the speaker’s intentions: cues to the speaker’s attention and intentions affect pronoun comprehension. cognitive psychology, 70, 58–81. http://doi.org/10.1016/j.cogpsych.2013.12.003 pickering, m. j., & garrod, s. (2013). an integrated theory of language production on and comprehension. behavioral and brain sciences, 36(4), 329–347. https://doi.org /10.1017/s0140525x12001495 roelofs, a. (1992). a spreading-activation theory of lemma retrieval in speaking. cognition, 42(1-3), 107-142. rohde, h. & kehler, a. (2014). grammatical and information-structural influences on pronoun production. language, cognition, and neuroscience, 29(8), 912–927. https://doi.org/10.1080/01690965.2013.854918 rosa, e. & arnold, j. e. (2017). predictability affects production: thematic roles can affect reference form selection. journal of memory and language, 94, 43–60. https://doi. org/10.1016/j.jml.2016.07.007 searle, s., cassella, g. & mccullouch, c. (1992). variance components. new york: jonh wileysons. https://doi.org/10.1002/9780470316856 does planning explain why predictability affects reference production? 55 tily, h. j., & piantadosi, s. t. (2009). refer efficiently: use less informative expressions for more predictable meanings. proceedings of the workshop on the production of referring expressions: bridging the gap between computational and empirical approaches to reference, 1–8. van der meulen, f. f., meyer, a. s., & levelt, w. j. (2001). eye movements during the production of nouns and pronouns. memory & cognition, 29(3), 512–521. http://doi.org/10.3758/bf03196402 van rij, j., van rijn, h., & hendriks, p. (2010). cognitive architectures and language acquisition: a case study in pronoun comprehension. journal of child language, 37(3), 731–766. http://doi.org/10.1017/s0305000909990560 weatherford, k., & arnold, j. e. (under review). semantic predictability of implicit causality can affect referential form choice. ms., university of north carolina. wheeldon, l. r. (2013). producing spoken sentences: the scope of incremental planning. in s. fuchs, m. weirich, d. pape, & p. perrier (eds.), speech production and perception: vol. 1. speech planning and dynamics. frankfurt, germany: peter lang. zerkle, s. a. & arnold, j. e. (2016). discourse attention during utterance planning affects referential form choice. linguistics vanguard, 2(s1). doi: https://doi.org/10.1515/lingvan-2016-0067 zerkle, s. a., rosa, e. c., & arnold, j. e. (2015). do addressee gestures influence the effects of predictability on spoken reference form? presented as a poster at the cuny 2015 conference on sentence processing, usc, los angeles, ca. march 19-21, 2015. zerkle, s. a., rosa, e. c., & arnold, j. e. (2017). thematic role predictability and planning affect word duration. laboratory phonology: journal of the association for laboratory phonology, 8(1): 17, 1-28. doi: http://doi.org/10.5334/labphon.98 d:/phd/papers/dandd09incrementality/2011/2011-01 incrementalitysplits.dvi dialogue and discourse 2(1) (2011) 279–311 doi: 10.5087/dad.2011.111 on incrementality in dialogue: evidence from compound contributions∗ christine howes c.howes@qmul .ac.uk school of electronic engineering and computer science queen mary university of london mile end road, london e1 4ns, uk matthew purver m .purver@qmul .ac.uk queen mary university of london patrick g. t. healey pat.healey@eecs.qmul .ac.uk queen mary university of london gregory j. mills gjmills@stanford.edu department of psychology stanford university stanford, ca 94305, usa eleni gregoromichelaki eleni.gregor@kcl .ac.uk philosophy department king’s college london strand, london wc2r 2ls, uk editor: hannes rieser and david schlangen abstract spoken contributions in dialogue often continue or complete earlier contributions by either the same or a different speaker. thesecompound contributions(ccs) thus provide a natural context for investigations of incremental processing in dialogue. we present a corpus study which confirms that ccs are a key dialogue phenomenon: almost 20% of contributions fit our general definition of ccs, with nearly 3% being the cross-person case most often studied. the results suggest that processing is word-by-word incremental, as splits can occur within syntactic ‘constituents’; however, some systematic differences between sameand cross-person cases indicate important dialogue-specific pragmatic effects. an experimental study then investigates these effects by artificially introducing ccs into multi-party text dialogue. results suggest that ccs affect people’s expectations about who will speak next and whether other participants have formed a coalition or ‘party’. together, these studies suggest that ccs require an incremental processing mechanism that can provide a resource for constructing linguistic constituents that span multiple contributions and multiple participants. they also suggest the need to model higher-level dialogue units that have consequences for the organisation of turn-taking and for the development of a shared context. keywords: compound contributions; corpus study; party formation; dialogue; incrementality ∗. this research was carried out under thedynamics of conversational dialogueproject, funded by the uk esrc (res-062-23-0962) with epsrc support through thediet (dialogue experimentation toolkit)project (ep/d057426/1). we also thank ruth kempson and arash eshghi for many useful discussions. c©2011 c. howes, m.purver, p. g. t. healey, g. j. mills and e. gregoromichelaki submitted 1/10; accepted 4/11; published online 5/11 howes, purver, healey, m ills and gregoromichelaki 1. introduction compound contributions(ccs) – spoken dialogue contributions that continue or complete an earlier contribution,1 see e.g. (1) – have been claimed to occur regularly in dialogue, especially according to the conversation analysis (ca) literature, where specific types of compound contributions have been studied under a variety of names, including completions and joint productions (see section 2). (1) daughter: oh here dad, a good way to get those corners out dad: is to stick yer finger inside. daughter: well, thats one way. [from lerner (1991)] ccs are of interest to dialogue theorists as they provide evidence about how contributions can cohere with each other at multiple levels – syntactic, semantic and pragmatic (though of course they are not the only way). they also indicate the radical context-dependency of conversational contributions, which can, in general, be highly elliptical without disrupting theflow of the dialogue. ccs are a dramatic illustration of this: speakers must rely on the dynamics of theunfolding context (linguistic and extra-linguistic) in order to guarantee successful processing and production. as early as 1967, in his series of lectures on conversation, sacks (1992) noted that the existence of ccs supports the (now largely accepted) thesis that language indialogue is processed incrementally: such a fact as that persons go about finishing incomplete sentences of others with syntactically coherent parts would seem to constitute direct evidence of their analysing an utterance syntactically in its course. . . (sacks, 1992, p651) however, we argue here that the evidence from ccs goes further; they show that not just processing (parsing), but also production (generation) must be incremental; and that because of the variation in ccs, this must also be at a finer-grained level than is often assumed (see also ferreira, 1996; guhe, 2007). compound contributions that are split across speakers also present a canonical example of participant coordination in dialogue (here we call thesecross-person ccsto distinguish them from the same-personcases where the original speaker later continues his own contribution – see below). the ability of one participant to continue another interlocutor’s contribution coherently, both at the syntactic and semantic level, implies that speaker and hearer can be highly coordinated in terms of processing and production. the initial speaker must be able to switch to the role of hearer, processing and integrating the continuation of their contribution, whereas the initial hearer must be monitoring the grammar and content of what they are being offered closely enough that they can take over and continue in a way that respects the constraints set up by the first contribution. this switch is particularly obvious in those cases where the initial hearers continuation is not the same as that which the original speaker would have provided, as in (1, 2). 1. these terms will be defined in detail below. 280 compound contributions (2) bma: she got compensation just like that because what she had in her suitcase pm: was grade a. [from comedy news quizhave i got news for you, s35 ep1] there is evidence that such constraints are respected across speaker and hearer in compound contributions (see e.g. gregoromichelaki et al., 2009). in both finnish (which has a rich inflectional morphology), and japanese (a verb-final language), cross-person ccs within a single clause conform to the strict syntactic constraints of the language, despite the change inspeaker (helasvuo, 2004; hayashi, 1999; lerner and takagi, 1999). these observations have important theoretical implications. firstly, the grammar and semantics employed by the interlocutors must be able to license and interpret chunksmuch smaller than the usual sentential or propositional units. moreover, the possibility of role switches while syntactic/semantic dependencies are pending suggests direct involvement of the grammar in the parsing and production processes, or, at least, a very tight coupling between those processes and the grammar and intermediate representations being used (see gargett et al., 2009). indeed, poesio and rieser (2010) claim that “[c]ollaborative completions . . . are among the strongest evidence yet for the argument that dialogue requirescoordinationeven at the sub-sentential level” (italics original). from a psycholinguistic point of view, the phenomenon of ccs is compatible withmechanistic approaches as exemplified by the interactive alignment model of pickeringand garrod (2004), which claims that it should be as easy to complete someone else’s sentence as one’s own (p186). according to this model, speaker and listener ought to be interchangeable at any point. this is also the stance taken by the grammatical framework of dynamic syntax (ds: kempson et al., 2001; cann et al., 2005). in ds, parsing and production are taken to employ the same mechanisms, leading to a prediction that ccs ought to be strikingly natural (purver et al., 2006).however, continuation by another speaker is sometimes taken to involve guessing or preempting the other interlocutor’s intended content.2 it has therefore been claimed that a full account of ccs requires a complete model of pragmatics that can handle intention recognition and formation. indeed, poesio and rieser (2010) propose sentence completions as the testing ground of competing claims about coordination i.e. whether it is best explained with an intentional model like clark’s (1996) or with a model based on simpler alignment models like pickering and garrod’s (2004). they conclude that a model which includes modelling of intentions better captures the data, though see (gregoromichelaki et al., 2011) for an alternative argument. for computational models of dialogue, compound contributions pose a challenge. while poesio and rieser (2010) and purver et al. (2006) provide general foundational models for various aspects of ccs, there are many questions that remain if automatic processing of naturally occurring dialogues is ever to be completely realised. a computational dialogue system must be able to identify ccs, match up their two (or more) parts (which may not necessarily be adjacent), integrate them into some suitable syntactic and/or semantic representation, and determine the overall pragmatic contribution to the dialogue context. ccs also have implications for the organisation of turn-taking in such models (see e.g. sacks et al., 1974), as regards what conditions(if any) allow or prevent successful turn transfer. 2. note that this says nothing about whether such a continuation is the same as the initial speaker’s intended continuation. for cases where this cannot be the case see (gregoromichelaki et al., 2011), as well as (1, 2) above. 281 howes, purver, healey, m ills and gregoromichelaki from an organisational point of view, it has been claimed that turn-taking operates not on individual conversational participants, but on ‘parties’ (schegloff, 1995). for example, a couple talking to a third person may organise their turns as if they are one ‘party’, ratherthan two separate individuals. lerner (1991) speculates that cross-person compound contributions can clarify the formation of such parties, as they reveal a relationship between syntactic mechanismsand social organisation. he claims that this provides evidence of one way in which syntax can be used to organise participants into “groups”. of course, our earlier definition of compound contributions begs several questions; most importantly, what do we mean by a ‘dialogue contribution’? our use of this term, andthe related notion of a turn, can be best explained by reference to a short extract of dialogue taken from the british national corpus (3). (3) 1. a: i were gonna say, they wash [[better than]] 2. j: [[but i’ve had]] 3. a: velvet. 4. j: i’ve had to take them up. 5. cos they were, they were gonna be miles too long. 6. and i’ve not even took them out the thing. 7. they said he’d swap them if they didn’t fit. 8. a: [[ah they do!]] 9. j: [[and he]]. 10. a: where d’ya get them from joyce? 11. j: i got them from that er 12. b: top marks. 13. j: that shop. [bnc kb2 4134-4146] in our usage, each of the transcribed lines (1-13) is acontribution. our use ofcontribution is intended to correspond to clark’s (1996) “acontribution to discourse – [is] a signal successfully understood” (p227).3 with transcribed corpus text, of course, it is not always possible to determine whether contributions have been successfully understood, as we have no access to non-verbal signals (such as nodding). we therefore take contributions to be stretches of talk bounded by a change in speaker, a significant pause, or the end of the sentence, and assume that in most cases the transcribers’ decision to split the text into separate lines indicate some (e.g. prosodic) cues to suggest that the line has been successfully understood, i.e. treated as acontribution. thus, whilst contributions can be single words (as in line 3) or backchannels (e.g. ‘mm’), or complete syntactic sentences (e.g. line 4), they can also be partial sentences (e.g. the incomplete sentences at lines 1, 2 and 11 and the fragments at lines 3, 12 and 13). note however, that single words in longer contributions (e.g. ‘they’ at the start of line 7) do not count as contributions in their own right. compound contributionscan now be defined assingle syntactic or semantic (propositional) units built across multiple contributions, which could be provided by one speaker or several. the exchange in lines 11-13 provides two examples. j’s contribution ‘i got them from that er’ starts a sentence, which b’s contribution ‘top marks’ (the name of a shop) completes. this counts as a 3. note that clark uses contribution to refer to both “the joint act of . . . completing the signal and its joint construal” and for the interlocutors “participatory act, hispart of that joint act, as when we speak of roger’s contribution to the discourse.” we use contribution in this second sense only. 282 compound contributions compound contribution under our definition. j then also completes her own contribution (with ‘that shop’) at line 13, and this also counts as a (same-person) compound contribution, as it is spread across multiple contributions (in this case, with intervening material). note that even though the short extract in (3) also exhibits many other conversationaltying techniques(sacks, 1992), such as a question and answer (lines 10-11), and the use of pronouns linked to referents previously introduced in the dialogue, our focus here is not on all pragmatic dependencies between turns. it should be noted, however, that this definition depends on the protocolused by the corpus transcribers; and with the bnc, this can lead to possibly undesirable segmentation of stretches of talk into multiple “contributions”. the insistence on linear ordering means that cases of interruption of one speaker by another will always result in an apparent speakerchange, even if the interruption consists only of non-verbal noises (e.g. coughing) or is entirely overlapping – see e.g. lines 1-3 (overlapping material is shown in the examples with square brackets aligned tothe material with which it overlaps). j’s interruption in line 2 overlaps with a’s speech, butforces a’s sentence to be transcribed as two lines (1 and 3). these count as separate contributions under our definition, giving a compound contribution: a begins her contribution ‘i were gonna say, they wash better than’, which she completes in line 3 with ‘velvet’. in many cases this may be the correct analysis – in clark’s usage, overlapcansignal understanding (i might not need you to syntactically or semantically finish your sentence to accept it as a valid contribution to the discourse). in this case, though, it may be that lines 1 and 3 were intended (and processed) as one single contribution – toavoid possibly misleading conclusions we therefore report cc figures both including and excluding such cases (see section 4). note, however, that these concerns only apply to same-person ccs andnot to cross-person ccs. whe also define a notion ofturn here as all talk to the next change of speaker; the contributions by j in lines 4-7 would therefore be classified as a singleturn. we will use this notion below to distinguish ccs which span multiple turns from those spanning multiple contributions within a single turn. even a backchannel or overlapping material, such as line 2 (which completely overlaps with the end of line 1) counts as a change of speaker (and thus separate turns) here. analysis of ccs, when they can or cannot occur, and what effects they have on the coordination of agents in dialogue, is therefore an area of interest not only for conversation analysts wishing to characterise systematic interactions in dialogue, but also for linguists trying toformulate grammars of dialogue, psychologists and sociolinguists interested in alignment mechanisms and social interaction, and those interested in building automatic dialogue processing systems.in this paper we present and examine empirical corpus data and an experimental manipulationof ccs, in order to shed light on some of the questions and controversies around this phenomenon. 2. related work most previous work on what we call ccs has examined specific sub-cases, generally of the crossperson type, and have referred to these variously ascollaborative turn sequences(lerner, 1996, 2004),collaborative completions(clark, 1996; poesio and rieser, 2010),co-constructions(sacks, 1992),joint productions(helasvuo, 2004),co-participant completions(hayashi 1999, lerner and takagi 1999),collaborative productions(szczepek, 2000a) andanticipatory completions(fox, 2007) amongst others (with some differences of emphasis in the differentterms). here we discuss some of this work. 283 howes, purver, healey, m ills and gregoromichelaki 2.1 conversation analysis anticipatory completions lerner (1991) identifies various structures typical of ccs which contain characteristic split points. one group of these are ‘compound’turn-constructional units(tcus), which are structures that include an initial constituent that hearers can identify as introducing some later final component. examples include theif x-then y, when x-then y and instead of x-y constructions (4). (4) a: before that then if they were ill g: they get nothing. [bnc h5h 110-111] other cues for potentialanticipatory completionsinclude quotation markers (e.g.she said), parenthetical inserts and lists, as well as non-syntactic cues such as contrast stress or prefaced disagreements. another important category that he identifies isterminal itemcompletions, which involve completing the final one or two lexical items of an interlocutor’s utterance at projectable locations of the current speaker’s turn ending (possibly involving overlap). opportunistic cases although lerner focuses on these projectable turn completions, he also mentions that ccs can occur at other points such as “intra-turn silence”, laugh tokens and hesitations, for example in cases of a stalled word search. all these cases he terms opportunistic completions(5). (5) a: well i do know last week thet=uh al was certainly very〈pause 0.5〉 b: pissed off [lerner (1996), p260] as he makes no claims regarding the frequency of such devices for ccs,it would be interesting to know how common these are, especially as studies on ccs in japanese (hayashi, 1999) show that although ccs do occur, compound tcus do not play as prominent a role asin english. it should be noted, however, that lerner’s definitions are not intended to be mutually exclusive. expansions vs. completions other classifications of ccs often distinguish betweenexpansions and completions(ono and thompson, 1993). expansions are continuations which add, e.g., an adjunct, to an already complete syntactic element (6, 7). (6) t: it’ll be an e sharp. g: which will of course just be played as an f. [bnc g3v 262-263] (7) m: yep dr goes everyones happy n: except the dr [diet su1 4213-4214] completions involve the addition of syntactic material which is required to make the whole compound contribution (syntactically) complete (5, 8). (8) a: . . . and then we looked along one deck, we were high up, and down belowthere were rows of, rows of lifeboats in case you see b: there was an accident. a: of an accident [bnc hdk 63-65] 284 compound contributions importantly, though we consider both expansions and completions to be ccs according to our terminology, we distinguish between the two types by considering the completeness or otherwise of the first part of the cc. thus while there might be arguments for restricting the definition of a cc to only the completion type, we are also interested in comparing the relative distributions of the different sub-types. in terms of frequency, the only estimate we are aware of in the ca literature is szczepek (2000a), who found approximately 200 cross-person ccs in 40 hours of english conversation (there is no mention of the number of sentences or turns this equates to), of which 75% are completions. as briefly outlined above, ca analyses of ccs tend to focus on their sequential implications in particular cases. these analyses provide clear examples of cross person co-ordination, however, it is unclear how representative they are (with the exception of szczepek (2000a), who offers limited figures). additionally, as the emphasis in the ca literature on ccs is in identifyingtheir organisational consequences for the unfolding dialogue (which can range from indicating understanding to highlighting differences of opinion (szczepek, 2000b)), they leave open the question of where a speaker switch may occur. 2.2 linguistic models purver et al. (2006) present a grammatical model for compound contributions, using an inherently incremental grammar formalism, dynamic syntax (kempson et al., 2001; cann et al., 2005). this model shows how syntactic and semantic processing can be accounted forno matter where the split point occurs; however, as their interest is in grammatical processing, they give no account of any higher-level inferences which may be required. poesio and rieser(2010) present a general model forcollaborative completions(a subclass of cross-person ccs) based in the ptt framework, using an incremental ltag-based grammar and an information-state-basedapproach to context modelling. while many parts of their model are compatible with a simple alignment-based communication model like pickering and garrod’s (2004), they see intention recognition as crucial to dialogue management. they conclude that an intention-based model, like clark’s (1996), is more suitable. their primary concern is to show how such a model can account for the hearer’s ability to infer a suitable continuation, but their use of an incremental interpretation method also allows an explanation of the low-level utterance processing required. nevertheless, the use of an essentially head-driven grammar formalism suggests that some syntactic split points ought to be more problematic than others. 2.3 corpus studies skuplik (1999) collected data from german two-party task-oriented dialogue, and annotated for cross-person compound contribution phenomena. she found thatexpansions(cases where the part before the split point can be considered already complete, as describedabove) were more common thancompletions(where the first part is syntactically or semantically incomplete as it stands), with 72 expansions (57%) and 54 completion ccs (43%) in her corpus. this contrasts with the data reported by szczepek (2000a), detailed above. there are severalpossible reasons for this contrast; for example, there may simply be a difference in the distributions of ccs in different languages, or between experimentally controlled task-oriented dialogue (which skuplik (1999) focused on) and casual conversational dialogue. additionally, there may be issues with the classification schemes used. for example, szczepek (2000a) did not include what she callsappendor questionsin her data, 285 howes, purver, healey, m ills and gregoromichelaki which could also be argued to be expansion ccs. the corpus study hereshould shed some light on some of these possible sources of disagreement. rühlemann (2007) uses corpus analysis on the bnc to examine a subset ofexpansionccs, sentence relativesof one’s own or another’s turn (6, 9). (9) a: profit for the group is a hundred and ninety thousand pounds. b: which is superb. [bnc fuk 2460-2461] he found thatsentence relativesare slightly more likely to be same-person than cross-person, with a total of 104 (55%) of 190 being same-person cases. this contrasts withtao and mccarthy (2001) who found 96% of their corpus sample were same-person; however, thisdiscrepancy can be attributed to the fact that they were measuring different things: tao and mccarthy (2001) included all non-restrictive (‘which’) relative clauses in their analysis, thus excluding restrictive readings, and including cases which were intra-sentential and thus would not count as ccs in our terminology (see section 3.1). in fact, r̈uhlemann (2007) also excluded intra-turn cases where the sentence relative was annotated as a separate sentence but there was no intervening material; our definition would include these. in addition, de ruiter and van dienst (in preparation) are also in the process of studying crosspersoncompletionsand their effect on the progressivity of dialogue turns; however no results are available to us at this point in time. notably, the definition used by de ruiter and van dienst (pc) only includes those completions where the additional material combines with the incomplete first part of the cc such that neither part could be considered complete withoutthe other. in our view, this excludes a number of interesting cases; not only expansion type ccs,but also those in which the continuation does not finish in a complete way (including, for example, ccswhich spread over more than two parts). 2.4 dialogue models skantze and schlangen (2009) and buß et al. (2010) present incremental dialogue systems (for limited domains) which can deal with some kinds of same-person compound contribution, allowing the system or user to provide mid-sentence backchannels, and/or resumewith sentence completion if interrupted. some related empirical work regarding the issue of turn-switch addressed here is also presented by schlangen (2006) but the emphasis there centers mostlyon prosodic rather than grammar/theory-based factors. for cross-person ccs, the only system we are aware of is that presented in devault et al. (2009) in which the system is able to generate a completion to a user’s input based on the semantic representation it has built up so far. due to the limited domain of possible semantic interpretations, the system is able to produce terminal item completions, once the possible interpretations have been sufficiently narrowed down. it does not, therefore, produce the range of ccs seen in naturally occurring human dialogue (including expansions as discussed above); we hope that empirical data such as that presented here can be used in constructing such systems and evaluating whether they achieve devault et al.’s stated aim of enabling virtual agents to display natural conversational behaviour. 286 compound contributions 3. general methods 3.1 terminology in this paper, as our interest is general, we use the termcompound contributions (ccs) to cover all instances where more than one dialogue contribution combine to form a (intuitively propositional) unit – whether the contributions are by the same or different speakers. we therefore use the term split point to refer to the point at which the compound contribution is split (rather than e.g. transition pointwhich is associated with a speaker change). cases where the speaker does change across the split point are calledcross-person ccs; otherwise we call themsame-person ccs. as not all cases will lead to complete propositions, and not all will be split over exactly two contributions, we also avoid terms likefirst-half, second-halfandcompletion: instead the contributions on either side of a split point will be referred to as theantecedent and thecontinuation. in cases where an compound contribution has more than one split point, some portions may therefore act as the continuation for one split point, and the antecedent for the next.we can then talk about completeness of each portion independently, with the traditional completion/expansion distinction corresponding to completeness (or otherwise) of the antecedent. see thesub-section on annotation scheme in section 4.1 for details of how completeness is assessed. 3.2 questions questions about frequencies and distributions are addressed in the corpus study (section 4); these lead to others about the effects on the ongoing dialogue, which are examined in the experimental manipulation (section 5). general our first interest is in the general statistics regarding ccs: how often do they occur? when they do, do they usually fall into the specific categories (with specific preferred split points) examined by e.g. lerner (1991), or can the split point be anywhere? what effects do ccs then have on the ongoing dialogue? do sameand crossperson ccs have different effects? specifically, do ccs have a bearing on ‘party’-formation in schegloff’s (1995) sense, as lerner (1991) claims? samevs cross-person we are also interested in the balance between sameand cross-person ccs. some grammatical formalisms (purver et al., 2006) and psycholinguistic models (pickering and garrod, 2004) predict that ccs should be equally natural in both sameand crossperson conditions – is this the case? what are the similarities and differences betweensameand crossperson cases? completeness for a grammatical treatment of ccs, as well as for implementing parsing/production mechanisms for their processing, we need to know about the likely completeness of antecedent and continuation (for example, if they are always complete in their own right, astandard head-driven grammar may be suitable; if not, something more fundamentally incremental may be required). in addition, ca analyses of dialogue phenomena predict that compound contributions should preferably occur at turn-transfer points that are foreseeable by the participants. complete syntactic units serve this purpose from this point of view and lack of such completeness will seem to weaken this general claim. we therefore ask how often antecedents and continuations are themselvescomplete. for antecedents, we are more interested in whether theyend in a way that seems complete as they may have started irregularly due to overlap or another cc (end-complete); for continuations, whether 287 howes, purver, healey, m ills and gregoromichelaki theystart in such a way – they may not get finished for some other reason, but we want to know if they would be complete if they do get finished (start-complete).4 these notions are by no means entirely clear cut (as pointed out by an anonymous reviewer there is much debate on whether e.g. adverbial adjuncts and semantic roles are necessary in a sentence) andschegloff (1996) concedes that his definitions are both arguable and not fully specified, although conversational participants do orient themselves to points ofpossiblecompletion. in practice, however, in most cases there was a high level of agreement between annotators on what constitutes syntacticor semantic completeness.5 we also look at the syntactic and lexical categories which occur either side of the split point. we are interested to know whether there are different effects on the unfolding dialogue from ccs with complete and incomplete antecedent contributions, and whether the positionof the split point has an effect. repair and overlap finally, we look at how often the continuation of an cc involves explicit repair (repetition, reformulation, modification or replacement) of antecedent material. any grammar of dialogue or computational system will need to be able to identify where thistakes place, and we therefore also look at how such repair depends on antecedent completeness and the type of split point. as our focus is on ccs, note that our use ofrepair refers only to those cases where the ‘end’ of the antecedent (immediately preceding the split point) is explicitly repeated orreframed at the start of the continuation. an example can be seen in (12), where the last word of the antecedent is repeated in the continuation. repairs at other points in the s-unit or turn are not taken into consideration.6 4. study 1: corpus study 4.1 materials and procedure for this exercise we used the portion of the bnc (burnard, 2000) annotated by ferńandez and ginzburg (2002), chosen to maintain a balance between what the bnc defines as context-governed dialogue (tutorials, meetings, doctor’s appointments etc.) and demographic dialogue (casual unplanned conversations). this portion comprises 11,469s-units– roughly equivalent to sentences7 – taken from 200-turn sections of 53 separate dialogues. the bnc transcripts are already annotated for overlapping speech, for non-verbal noises (laughter, coughing etc.) and for significant pauses. punctuation is included, based on the original audio and the transcribers’ judgements; as the audio is not available, we allowed annotators to use punc4. the notion of end-completeness that we are trying to capture is the ca notion of endingsas outlined in schegloff (1996); “for any tcu we can ask . . . does it end with an ending, i.e., does it come to a recognizable possible completion – syntactic, prosodic and action/pragmatic.” likewise hisbeginningsfor our start-completeness; “turn constructional units – and turns – can start with a “beginning” or with something which is hearablynot a beginning.” 5. see the sub-section on annotation scheme in section 4.1 for operational details, and table 2 for kappa agreement scores between annotators. 6. consequently, our use ofrepair should be understood not as capturing all instances of repair but only as indexing the frequency with which these specific aspects of the contribution are repaired. 7. the bnc is annotated intos-units, defined as “sentence-like divisions of a text”, andutterances, defined as “stretches of speech usually preceded and followed by silence or by a change of speaker”. utterances may consist of many s-units; s-units may not extend across utterance boundaries. while s-units are therefore often equivalent to complete syntactic sentences, or complete functional units such as bare fragments or one-word utterances, they need not be: they may be divided by interrupting or overlapping material from anotherspeaker. 288 compound contributions tuation where it aided interpretation. the bnc transcription protocol divides the transcript into sentence-like units (“s-units” ) as well as speaker turns (“utterances” – see footnote 7), where utterances may contain several s-units from the same speaker. we annotated at the level of individual s-units, to allow self-continuations within a turn to be examined; we are therefore taking the bnc’s s-unit to correspond to our notion of dialoguecontribution, and the bnc’sutteranceas our notion of turn. the bnc forces speaker turns to be presented in linear order, which is vital if we are to accurately assess whether turns are continuations of one another; however, this has a side-effect of forcing long turns to appear as several shorter turns when interruptedby intervening backchannels. we will discuss this further below. tag value explanation end-complete y/n for all s-units: does this s-unit end in such a way as to yield a complete proposition or speech act? continues s-unit id for all s-units: does this s-unit continue the proposition or speech act of a previous s-unit? if so, which one? repairs number of words for continuations: does the start of this continuation explicitly repair words from the end of the antecedent? if so, how many? start-complete y/n for continuations: does this continuation start in such a way as to be able to stand alone as a complete proposition or speech act? table 1: annotation tags annotation scheme the initial stage of manual annotation involved four tags:end-complete, continues, repairs and start-complete – these are explained in table 1 above. sunits which somehowrequire continuation (whether they receive it or not) are therefore those markedend-complete=n; s-units which act as continuations are those marked with non-empty continues tags; and their antecedents are the values of thosecontinues tags. further specific information about the syntactic or lexical nature of antecedent or continuation could then be extracted (semi-) automatically, using the bnc transcript and part-of-speech annotations. 289 howes, purver, healey, m ills and gregoromichelaki (3) e n d c o m p l e t e c o n t i n u e s r e p a i r s s t a r t c o m p l e t e 1. a: i were gonna say, they wash [[better than]] n 2. j: [[but i’ve had]] n 3. a: velvet. y 1 n 4. j: i’ve had to take them up. y 2 3 y 5. cos they were, they were gonna be miles too long. y 4 n 6. and i’ve not even took them out the thing. y 5 n 7. they said he’d swap them if they didn’t fit. y 8. a: [[ah they do!]] y 9. j: [[and he]]. y 7 n 10. a: where d’ya get them from joyce? y 11. j: i got them from that er n 12. b: top marks. y 11 n 13. j: that shop. y 11 1 n returning to the extract in (3), repeated here, we can see how these tagsare applied in practice. note that all s-units have anend-complete tag whilst only those that are judged to continue some prior contribution have any other tags. the reason for judging end-completeness rather than whether the s-unit constitutes a complete proposition or speech act in its own right, is due to both the fragmentary nature of dialogue and the transcription practices of the bnc, which, as already discussed, may break up a syntactic sentence into several s-units due to overlapping material etc. whether an s-unitendsin a potentially complete way is therefore independent of whether itstarts in one. for thecontinues tag, the value is the line number which this s-unit is judged to continue (i.e. the line number of the antecedent); lines 12 and 13, for example, are both judged to be a continuation of line 11. therepair tag takes as its value (if it has one) the number of words from the end of the antecedent which are repeated, reformulated, modified or replaced at the start of the continuation. line 4 has arepair value of 3, because the continuation repeats the three words from the end of line 2 (which is the antecedent) –‘i’ve had’.8 finally, thestart-complete tag (also only applied to continuations) indicates whether the contribution starts in away that it might be the beginning of a complete sentence (even though it may not itself be complete). continuations starting withand/or/but/becauseetc. are always tagged asstart-complete=n, as can be seen in lines 5, 6 and 9. inter-annotator agreement in some cases, it is not easy to identify whether a fragment is a continuation or not, or what its antecedent is – see e.g. (10), where g’s second contribution could be seen as continuing either his own prior utterance, or a’s intervening contribution: (10) g: well a chain locker is where all the spare chain used to like coil up a: so it 〈unclear〉 came in and it went round 8. ‘i’ve’ is counted as two words as a contraction of‘i have’. 290 compound contributions g: round the barrel about three times round the barrel then right down into the chain locker but if you kept, let it ride what we used to call let it ride well〈unclear〉 well now it get so big then you have to run it all off cos you had one lever, that’s what you had and the steam valve could have all steamed. [bnc h5g 174:176] similar issues also arise in judgements of completeness, as it is not always obvious if a contribution is syntactically or semanticallyendand/orstart-complete. we therefore assessed interannotator agreement between the three authors who acted as annotators.first, all three annotated one dialogue independently, then compared results and discussed differences. they then annotated 3 further dialogues independently and agreement was measured; kappastatistics (carletta, 1996) are shown in table 2 below. bnc dialogue code tag knd kbg kb0 end-complete .86-.92 .80-1.0 .73-.90 continues (y/n) .81-.89 .76-.85 .77-.89 continues (ant) .82-.90 .74-.85 .76-.86 repairs 1.0-1.0 .55-.81 1.0-1.0 start-complete .59 .68 .62 table 2: inter-annotatorκ statistic (min-max) with the exception of therepairs tag for one annotator pair for one dialogue and thestartcomplete tags, all are above 0.7; the low figure in therepair category results from a few disagreements in a dialogue with only a very small number ofrepairs instances. thestartcomplete kappa figures, between the two annotators who completed this task, are around 0.6 suggesting that this measure may be less easy to determine. the remaining dialogues were then divided evenly between the three annotators. 4.2 results and discussion the 11,469 s-units annotated yielded 2,231 ccs, of which 1,902 were same-person and 329 crossperson cases; 112 examples involved an explicit repair by the continuationof the antecedent. the data come from the full range of dialogues; all dialogues had at least three same-person cases, though 5 of the 53 dialogues had no cross-person ccs. the mean numberof same-person ccs is 35.89 per dialogue (standard deviation 22.46). for cross-person ccs the mean was 6.21 per dialogue (s.d. 5.69). withinand cross-turn cases same-person ccs are much more common than cross-person; however, many of these same-person cases (around 44%) are self-continuations within a single speaker turn (such as those between lines 4 and 5 in (3)). as explained in section 3.2, we consider sameperson cases to be interesting in their own right. from a processing/psycholinguistic point of view, we would like to know whether such split points occur in the same places in cross-person ccs as in same-person ccs. however, there are certainly arguments for considering ccs within a turn as single contributions, and including them when comparing the frequency or nature of sameand cross-person ccs may give an unfair comparison, as cross-personccs can only occur at speaker turn boundaries. 291 howes, purver, healey, m ills and gregoromichelaki in addition, some apparently cross-turn cases (around 17%) may in factonly appear as such due to the bnc transcription protocol, which forces speaker turns to be strictly linearly ordered. a sentence from a single speaker which is interrupted by material from another speaker will be transcribed as two separate turns – even if the intervening material is non-verbal (e.g. a cough) and/or entirely overlaps with the original sentence rather than actually interrupting its flow (as seen in (3) lines 1-3). in the tables and results below, we therefore present same-person cc figures both including all cases, and excluding those cases which are either within-turnor separated only by non-verbal or overlapping material. we label these figures asall andcross-turnrespectively. person: samecrossall cross-turn (all) n % n % n % overlapping 0 0 0 0 18 5 adjacent 840 44 0 0 262 80 sep. by overlap 320 17 0 0 10 3 sep. by backchnl 460 24 456 63 17 5 sep. by 1 s-unit 239 13 229 32 16 5 sep. by 2 s-units 31 2 31 4 4 1 sep. by 3 s-units 5 0 3 0 1 0 sep. by 4 s-units 4 0 4 1 0 0 sep. by 5 s-units 1 0 1 0 0 0 sep. by 6 s-units 2 0 2 0 1 0 total 1902 726 329 table 3: antecedent/continuation separation general observations looking at cross-turn cases, even excluding those within-turn and overlapping cases discussed above, there are over twice as many same-person ccs (726) as cross-person ccs (329). many ccs have at least one s-unit intervening between the antecedent and continuation (see table 3). in same-person cases, once we have excluded the within-turn ccs described above, this must in fact always be the case (see, for example, lines 11 and 13 in (3), where the contribution at line 12 means that the antecedent (line 11) and continuation (line 13) are non-adjacent); the intervening material is usually a backchannel (63% of remaining cases) or asingle other s-unit (32%, often e.g. a clarification question), but two intervening s-units are possible(4%) with up to six being seen. in cross-person cases, 88% are adjacent or separated onlyby overlapping material, but again up to six intervening s-units were seen, with a single s-unit most common (10%,in half of which the intervening s-unit was a backchannel). many compound contributions have more than two separate contributions. insame-person cases, a cc can be split over as many as thirteen individual s-units; although such extreme cases occur generally within one-sided dialogues such as tutorials, many multi-split cases are also seen in general conversation. only 63% of cases consisted of only two s-units. antecedents can also receive more than one competing continuation (as in (3), where line 11 is continued in both lines 12 and 13), although this is rare: two continuations are seen in 2% of cases. ca categories we searched for examples which match ca categories (lerner, 1991; rühlemann, 2007) by looking for particular lexical items on either side of the split point. this search was per292 compound contributions formed in two stages: a loose (very high recall but low precision) automatic matching followed by manual checking to remove false positives (although some counts may still be slight over-estimates). for lerner’s (1996)opportunisticcases, we looked for filled pauses (‘er/erm’ etc.) or pauses explicitly annotated in the transcript (‘’), so counts in this case may be underestimates if short pauses were not transcribed. we also chose some other broad categories based on our observations of the most common cases. results are shown in table 4 (where the‖ token represents the split point).9 person: samecrossall cross-turn (all) n % n % n % . . .‖ and/but/or . . . 748 39 306 42 116 36 . . .‖ so/whereas . . . 257 14 57 8 39 12 . . .‖ because . . . 77 4 32 4 3 1 . . . er/erm‖ . . . 35 2 21 3 12 4 . . . ‖ . . . 19 1 15 2 20 6 . . .‖ which/who/etc . . . 26 1 11 2 4 1 . . . instead of . . .‖ . . . 1 0 0 0 0 0 . . . said/thought/etc . . .‖ . . . 12 1 5 1 0 0 . . . if . . .‖ (then) . . . 18 1 10 1 2 1 . . . when . . .‖ (then) . . . 6 0 4 1 1 0 (other) 783 41 317 44 164 50 total 1902 726 329 table 4: continuation categories the most common of the ca categories can be seen to be lerner (1996)’s hesitation-related opportunisticcases, which make up 3-5% of sameand 10% of cross-person ccs, meaning crossperson opportunistic cases are more common than same-person ones (same(cross-turn; 36 of 726) vs other (32 of 329)χ2 (1) = 8.53, p = 0.00310). interestingly, the breakdown of cases into those where the antecedent ends with an unfilled pause versus those which endwith a filled pause also shows a difference between sameand cross-person cases: an other person is more likely to offer a continuation after an unfilled pause, than after a filled pause (antecedents ending in‘er(m)’ 35 continued by same, 12 by other; ending in ‘’ 19 continued by same, 20 by otherχ2 (1) = 6.05, p = 0.01). this finding backs up claims by clark and fox tree (2002), that filled pauses can be used to indicate that the current speaker’s turn is not yet finished and thus have the effect of holding the floor. lerner’s compound tcu cases (instead of, said/thoughtetc,if-thenandwhen-then) account for 2-3% of same-person and 1% of cross-person ccs, though note that these could be underestimates, as his non-syntactic cues (e.g. contrast stress and prefaced disagreements) could not be extracted. rühlemann’s (2007)sentence relativecases come next with over 1%. 9. note that the categories in table 4 are not all mutually exclusive (e.g. an example may have both an‘and’-initial continuation and an antecedent ending in a pause), so column sums will not match totals shown. 10. for completeness, wherep > 0.001, we report exact probabilities but throughout adopt a criterion probability level of < 0.05 for accepting or rejecting the null hypothesis. 293 howes, purver, healey, m ills and gregoromichelaki in contrast, by far the most common pattern (for sameand cross-personccs) is the addition of an extending clause, either a conjunction introduced by‘and/but/or/nor’ (36-42%), or other clause types with‘so/whereas/nevertheless/because’. there are differences in the proportions of the clause types between sameand cross-person ccs, but further research and annotation is needed to confirm whether this represents systematic differences in pragmatic use (as in rühlemann’s (2007) sentence relative study, where cross-person ccs more often expressed stance (speaker opinion) than sameperson ccs). split point other less obviously categorisable cases make up 40-50% of continuations, in both sameand crossperson cases, with the most common first words being‘you’, ‘it’ , ‘i’ , ‘the’, ‘in’ and ‘that’ . in terms of syntactic categories, manual examination of the data suggests that the split point can occur at any point between words,11 even within what traditional theories of grammar consider to be a single constituent,12 such as noun phrases and prepositional phrases (11, 12, 13, 14). (11) d: yeah i mean if you’re looking at quantitative things it’s really you know how much actualhow much variation happens whereas qualitative is〈pause〉 you know what the actual variations u: entails d: entails. you know what the actual quality of the variations are.[bnc g4v 114-117] (12) m: we need to put your name down. even if that wasn’t a p: a proper conversation m: a grunt. [bnc kdf 25-27] (13) a: all the machinery was g: [[all steam.]] a: [[operated]] by steam [bnc h5g 177-179] (14) k: i’ve got a scribble behind it, oh annual report i’d get that from. s: right. k: and the total number of [[sixth form students in a division.]] 11. there is anecdotal evidence that ccs can also occur mid-word, aswhen someone completes a complex multi-syllabic word for another person. only one of our cross-person ccs occurred mid-word (shown in (i), from a doctor/patient exchange), in which the whole word is also repeated, so we leave such considerations aside for now, though obviously they have implications for e.g. the organisation of the lexicon. (i) a: no it wasn’t marvelon it was that trin d: trin a: aye. d: trinordiol. a: mhm. [bnc g58 63-68] 12. of course, different grammars may have different notions of constituency (such as thesurprising constituentsof ccg (steedman, 2000)) which these findings may have a bearing on, however, for the purposes of the current discussion, we limit our notion of constituency to that of syntactic elements as in, for example transformational grammars, or hpsg. 294 compound contributions s: [[sixth form students in a division.]] right. [bnc h5d 123-127] to further test the finding that the split point can apparently occur between any types of words, we annotated thecompletioncases for whether the split point occurred within a syntactic constituent, or between constituents.13 for same-person cross-turn ccs, just over half are between-constituent (111/213; 52%), whilst cross-person ccs appear to be more likely to occur within-constituent although this trend is not significant (52/87; 60%;χ2 (1) = 3.49, p = 0.06). this finding appears to be associated with repair (there seem to be more repairs in the within-constituent cases) but the numbers are too small to be sure. person: samecrossall cross-turn (all) n % n % n % antecedent end-complete y 136772 513 71 242 74 n 535 28 213 29 87 26 continuation start-complete y 22412 99 14 48 15 n 1678 88 627 86 281 85 repair y 77 4 34 5 32 10 n 1825 96 692 95 297 90 total 1902 726 329 table 5: completeness and repair completeness examination of theend-complete annotations shows that about 8% of s-units in general are incomplete (930/11469), but that (perhaps surprisingly) only 64% (591/930) of these get continued. this compares to 15% of end-complete s-units (1577/10539) that get continued (χ2 (1) = 1315.90, p < 0.001), showing that although incomplete s-units are more likely to be continued, incompleteness does not necessarily prompt the production ofa completion. the majority of both sameand cross-person continuations (71% to 74%) continue an already complete antecedent, with only 26-29% therefore beingcompletionsin the sense of e.g. ono and thompson (1993). interestingly, though, continuations are no more likely than other s-units to end in a complete way themselves. in fact, continuations are significantly more likely than other s-units to end in an incomplete way (273/2231 (12%) vs. 657/9238 (7%);χ2 (1) = 63.34, p < 0.001). the frequent clausal categories from table 4 are all much more likely to continue complete antecedents than incomplete ones.14 this is not the case for the(other) category; again suggesting that split points often occur at random points in a sentence, without regard to particular clausal constructions. the continuations in the(other) category are far less likely to continue complete antecedents than the easily classifyable categories from table 4 (220/481; 46% v. 535/574; 93%, χ2 (1) = 289.76, p < 0.001). 13. here we are concerned with only low-level syntactic constituency; wecounted a split point as within-constituent if it fell within a noun phrase (e.g. between a determiner and noun), a prepositional phrase (e.g. between a preposition and a noun phrase) or within a complex noun phrase (e.g. between an auxiliary and a head noun). other cases (e.g. between a verb and its object, or between clauses) were coded as between-constituent. 14. for the less frequent (e.g.‘if/then’, ‘instead of’) categories, the counts are too low to be sure. 295 howes, purver, healey, m ills and gregoromichelaki looking only at the general(other) category, we see that cross-person continuations more often follow antecedents that end in a complete way than same-person continuations (89/164; 54% v. 131/317; 41%χ2 (1) = 7.30, p = 0.007). for both cross-person and same person cases, continuations in the(other) category do not often start in a complete way (cross-person: 41/164; 25%, sameperson 94/317; 30%). in general, however, continuations are more than twice as likely to start in a non-complete rather than a complete way, even after complete antecedents. repair explicit repair of the antecedent is not common, only occurring in just under 5% of ccs. as might be expected, incomplete antecedents are more likely to be repaired (cross-turn (same and cross-person); 51/300 17% vs. 15/755 2%,χ2 (1) = 82.51, p < 0.001). cross-person continuations are also significantly more likely to repair their antecedents than same-personcases (32/329; 10% vs. 34/726; 5%,χ2 (1) = 9.82, p = 0.002). in those ccs where the split point falls within a syntactic constituent, only 18% (18/102) of same-person cases involve explicit repair at the start of the continuation,compared to 27% (14/52) of cross-person ccs (the equivalent figures for ccs where the splitpoint is between constituents are 12% (13/111) and 18% (6/35)). although more data are required to see if these are genuine differences, we know that repair in general is not common, so it appears that even when the split point occurs mid-constituent, the participants generally are able to just go on extending the constituent as if they were the original speaker. this might suggest that the parsing and generation mechanisms are not required to back up to the beginning of a constituent in order toprocess or produce a continuation (i.e. start with a new grammar rule). this seems to favour lexicalised or dependencybased parsing models in that it suggests that the language processing mechanisms directly rely on word-by-word dependencies rather than constituents/grammar rules. function of ccs we are concerned in this study primarily with the form, rather than the function, of ccs. however, it is worth noting at this point that they can perform functions beyond merely extending or completing an interlocutor’s contribution (see also szczepek,2000b); and in some cases are difficult to define functionally, and may even exhibit genuinemultifunctionality(see e.g. gregoromichelaki et al., 2009; bunt, 2009). in (15), for example, j’scontinuation of m’s utterance serves also as a request for confirmation: (15) m: it’s generated with a handle and j: wound round? m: yes [bnc k69 109-112] in many cases, the antecedent explicitly invites the hearer to complete the contribution, so that antecedent and continuation form a question-answer pair, possibly withina single grammatical constituent (16): (16) j: the holy spirit is the one who gives us hope. mega. i mean〈pause〉 this generation needs hope. the holy spirit is one who〈pause〉 gives us? u: strength. j: strength. 296 compound contributions yes, indeed. 〈pause〉 the holy spirit is one who gives us?〈pause〉 u: comfort. yes. [bnc hdd 274-283] this phenomenon can also happen in cases of clarification-request/clarification-reply pairs (see purver et al.’s (2003)gapcategory), e.g. (17): (17) g: cos they〈unclear〉 they used to come in here for water and bunkers you see. a: water and? g: bunkers, coal, they all coal furnace you see,〈clears throat〉 and we er they’d come in and we used to fill them up with coal, whatever they wanted〈cough〉 lot of that went over the side〈unclear〉 coal, beautiful coal that was. [bnc h5h 59-61] with the range of possibilities regarding where the split point is able to occur,including potentially within a word (see footnote 11) it is hard to see how compound contributions could be characterised as a well-defined syntactic phenomenon, a separate grammatical fragment category, or a sub-class ofnon-sentential utterance(ferńandez and ginzburg, 2002). moreover, there seems no reason to associate either antecedent or continuation with particular semantic categories or specific pragmatic speech-act information, as they seem to serve a wide rangeof purposes in dialogue: from assisting a speaker with lexical access, to eliciting a response to a query, to covertly offering a suggestion or asking a clarification. summary the results here show that ccs are common in dialogue. split points may be possible at any syntactic point, but there appear to be (possibly pragmatic) constraints on where they are likely to appear: they are far more likely after complete antecedents, although relatively few of them occur in the highly projectable positions studied by e.g. lerner. there are interesting differences between same-person ccs and cross-person ccs; firstly, sameperson ccs are over twice as common as cross-person. cross-person continuations are more likely to start with explicit repair/reformulation of the antecedent; this might be considered surprising, as self-repair is preferred in general (schegloff et al., 1977) althoughwe have no comparable figures for repair at other points in the turn. however, it is interesting to note that a cc, in virtue of being constructed as a continuation of the speakers utterance, may provide a device that enables a less exposed form of other repair. outside the frequent clausal or ca categories, cross-person ccs are also more likely than sameperson to continue a complete antecedent; and they are more likely where the antecedent ends in an unfilled pause rather than a filled one. this suggests an effect on turn-taking expectations, and that continuations may be systematically invited by a speaker or designed as though they are natural continuations of contributions that could be treated as complete. we will returnto these points in the general discussion. 5. study 2: experimental manipulation while the corpus study of section 4 provides us with useful information concerning the nature and frequency of ccs and their various sub-categories, it can tell us nothing about the effect of ccs on the dynamics of a conversation. 297 howes, purver, healey, m ills and gregoromichelaki from a processing point of view, we might intuitively predict that cross-person ccs ought to be more difficult for a third party to process than same-person ccs, as information from potentially conflicting sources must be integrated and interpreted as a single syntactic unit. conversely, some models (e.g. dynamic syntax, cann et al. (2005)) would predict that thereshould be no additional processing costs. the corpus study also suggests pragmatic effects associated with ccs. are cross-person ccs indicative of particularly close coordination (and thus of schegloff’s ‘parties’ as lerner suggests), which might facilitate understanding, or are they viewed as impolite which may addadditional implications and disrupt the flow of the conversation? the experiment reported here is, we believe, the first controlled manipulation of compound contributions during an unfolding interaction. the allows us to directly compare the effects of same-person and cross-person ccs on participants in a dialogue. the effects of seeing a cc on a dialogue in progress were tested using thedialogue experimentation toolkit (diet) chat tool, which enables text dialogues to be experimentally manipulated (see healey et al., 2003). of course, text based chat is different to face-to-face dialogue in several ways, and while clearly an interesting field of study in itself (rosé et al. (2003), for example, compare text and speech based tutoring systems), there are important questions as to whether the results from our corpus study are generalisable to such a different modality. the most obvious differences are attributable to the channel of communication; speech versus text. in text-based chat suchas msn messenger and the chat tool reported here, participants compose their turns in private before sending them to the other participants. this means that they can revise or even delete their turns without their interlocutors being aware of the revisions, unlike in face-to-face dialogue where overt repairs are necessarily shared. it also means that participants can compose their next turns simultaneously, meaning that the linearity of turn-taking in dialogue is lost. linked to this is the fact that, unlike inface-to-face dialogue, participants engaged in a text chat are not typically co-present. although this means that a number of non-linguistic cues are unavailable, this is also true in telephone conversations, for example, so should not be taken as a reason for rejecting the dialogic nature of text chat. despite these differences, there are also important similarities between text chat and face-to-face dialogue. both involve the use of interlocutors’ language resources to communicate, and text chat also exhibits many features which are generally seen in spoken dialogue, but not in either spoken monologue or written text. these include the use of non-sentential utterances such as clarification requests (purver et al., 2003) and acknowledgements (fernández and ginzburg, 2002). importantly for the study reported here, ccs also occur naturally in text-based chat(see, for example, (7) and (18), taken from the diet chat tool environment). (18) u: i agree tom needs to be there a: but one of them has to go to save the other 2 r: and what about the cancer research plan ?? according to a preliminary corpus study (eshghi, 2009) ccs occur as frequently in text chat as they do in face-to-face dialogue. in a total of 2377 text contributions, there were 493 ccs, of which 112 were cross-person and 381 were same-person ccs. overall this proportion of ccs is not different to that from our bnc corpus study. in the text chat corpus there was a higher proportion of cross-person ccs than expected from our bnc results (112 out of 2377 versus 326 298 compound contributions out of 11469χ2 (1) = 22.46, p < 0.001) which could be related to the task-based nature of the text chat (the dialogues analysed in eshghi (2009) are three-way tangramtask conversations between two directors and a matcher), or possibly due to the way turns are transcribed into consecutive contributions in the bnc, as previously discussed. 5.1 method in the diet chat tool, interventions can be introduced into a dialogue in real time, thus causing a minimum of disruption to the natural ‘flow’ of the conversation. in this experiment, a number of genuine single contributions in a text-based three-way conversation wereartificially split into two parts. in some conditions, both parts still appeared to originate from the genuine source (“speaker”), thus appearing as a same-person cc. in other conditions, one or both parts seemed to come from another participant, thus appearing either as an cross-person cc, or as a same-person cc generated by the “wrong” person. 5.1.1 materials the balloon task theballoon taskis an ethical dilemma requiring agreement on which of three passengers should be thrown out of a hot air balloon that will crash, killing all the passengers, if one is not sacrificed. the choice is between a scientist, who believes he is on thebrink of discovering a cure for cancer, a woman who is 7 months pregnant, and her husband,the pilot. this task was chosen on the basis that it should stimulate discussion, leading to dialogues ofa sufficient length to enable an adequate number of interventions. the diet chat tool the diet chat tool itself is a custom built java application consisting of two main components: user interface and server console. user interface the user interface is designed to look and feel like common instant messaging applications e.g. microsoft messenger. it consists of a display split into two windows, separated by a status bar, which indicates whether any other participant(s) are activelytyping (see figure 1). the ongoing dialogue, consisting of both the nickname of the contributor and theirtransmitted text, is shown in the upper window. in the lower window, participants type and revise their contributions, before sending them to their co-participants. all key presses are time-stamped and stored by the server. server console all text entered is passed to the server, from where it is relayed to the other participants. no turns are transmitted directly between participants. prior to being relayed, some turns are altered by the server to create fake ccs. this is carried out automatically. a genuine single-person contribution is splitaround a space character near the centre of the string. the part of the turn before the space is relayed first, as the antecedent, followed by a short delay during which no other turns may be sent. this is followed by the continuation (the part of the turn after the space), as if they were in fact two quite separate, consecutive contributions. in every case, the server produces two variants of the compound contribution, relaying different information to both recipients. each time an intervention is triggered, one of the two recipients receives a same-person cc from theactual source of the contribution (henceforth referred to as anaa-split). the other recipient receives one of three, more substantial, manipulations: a same-person cc that wrongly attributes both antecedent and continuation to the 299 howes, purver, healey, m ills and gregoromichelaki figure 1: the user interface chat window (as viewed by participant ‘sam’) other recipient (abb-split); a cross-person cc whose antecedent comes from from the actual origin and continuation from the other recipient (anab-split), or vice-versa (aba-split). this allows us to create a2×2 factorial design which separates potential effects of ‘floor change’ i.e. whether the original speaker finishes the cc or another participant appears to, from effects of ‘same/other’ i.e. whether a the two halves of the cc or appear to be produced by the same speaker or by two different speakers. this contrast is shown in table 6. a types: should we start now b sees (aa intervention): a: should we a: start now c sees (one of): ab intervention: ba intervention: bb intervention: a: should we b: should we b: should we b: start now a: start now b: start now table 6: comparison of split types the intervention is triggered every 10 turns, and restricted such that the participant who receives the non aa-split is rotated (to ensure that each participant only sees any of the more substantially manipulated interventions every 30 turns). which of the three non aa-splits they see (ab, ba or bb) is, however, generated randomly. 300 compound contributions 5.1.2 subjects 41 male and 19 female native english speaking undergraduate students were recruited for the experiment, in groups of three to ensure that they were familiar with each other.all had previous experience of internet chat software such as microsoft messengerand each was paid£7.00 for their participation. 5.1.3 procedure each of the triad of subjects was sat in front of a desktop computer in separate rooms, so that they were unable to see or hear each other. subjects were asked to follow the on-screen instructions, and input their e-mail address and their username (the nickname that would identify their contributions in the chat window). when they had entered these, a blank chat window appeared, and they were given a sheet of paper with the task description. participants were instructed to read this carefully, and begin discussing the task with their colleagues via the chat window once they had done so. they were told that the experiment was investigating the differences in communicationwhen conducted using a text-only interface as opposed to face-to-face. additionally, subjects were informed that the experiment would last approximately 20-30 minutes, and that all turns would be recorded anonymously for later analysis. once all three participants had been logged on, the experimenter went to sit at the server machine, a fourth desktop pc out of sight of all three subjects, and made no further contact until after at least 20 minutes of dialogue. 5.1.4 analysis as production and receipt of contributions sometimes occurs in overlap in text chat, it is not possible to say definitively when one contribution is made in direct response to another.15 we therefore chose to measure all the contributions produced by both recipients between the mostrecent intervention and the next intervention, averaged to produce one data point per recipient per intervention. this means that there are two data points for each intervention (one for each ofthe participants who saw a fake compound contribution). the data were analysed according to two factors in a2×2 factorial design;same/other– whether both parts of the compound contribution appeared to come from the same-person, or from different sources ([aa and bb] vs [ab and ba]), andfloor change– whether the continuation part of the cc appeared to come from the genuine source or the other participant ([aa and ba] vs [ab and bb]), with participant as a random factor. measures selected for analysis weretyping time of turn(the time, in milliseconds, between the first key press in a turn and sending the turn to the other participants by hittingthe return key) and length of turn in charactersas measures of production;deletes per character(the number of keyed deletes divided by the total number of characters) as a measure of revisions; andtyping time per characteras a measure of speed. data in tables are displayed in the original scale of measurement. however, as inspection of the data showed that they were not normally distributed, logarithmic 15. in online chat, participants can compose their next contributions simultaneously, and contributions under construction when another is received can be subsequently revised, prior to transmission. this means that a genuine response to a compound contribution might have a negative start time. however, theinclusion of cases where the whole contributions was constructed after receiving the cc (an arbitrary cut-off point, which would catch some contributions that were responses to earlier contributions in the dialogue, and miss somewhich were begun before the intervention was received and subsequently revised) should impose the same levelof noise in all cases. 301 howes, purver, healey, m ills and gregoromichelaki transformations (usingloge) were applied to the typing time of turn and length of turn in characters measures prior to all inferential statistical analyses, resulting in data distributions that were not significantly different from a normal distribution (using shapiro-wilk tests:typing time of turn w = 0.998, p = 0.882; length of turn in charactersw = 0.995, p = 0.100). for the proportional measures of deletes per character and typing time of character, which violate normality assumptions even after transformations, alternative analyses were used. the generalized linear model (gzlm) extends the general linear model (glm; which includes anovas and linear regression models) to include response variables that follow any exponential probability distribution, including e.g. poisson, binomial and gamma distributions. gzlms use maximum likelihood estimation to fit the model to the data (and provide parameter estimates). generalized estimating equations (gee) extend gzlm further by allowing fornon-independent data, such as repeated measures and clustered data. using a gee analysis (see liang and zeger, 1986; ballinger, 2004) on these variables therefore allows for both the non-normality of the data, and within-subject correlations. 5.2 results a post-experimental questionnaire and debriefing showed that, with the exception of one subject, who had taken part in a previous chat tool experiment and was therefore aware that manipulations may occur, none of the participants were aware of any interventions. of the 253 interventions to which at least one recipient responded, 89 were aa/ab splits, 99 were aa/ba splits and 65 aa/bb splits. this means there were 506 potential responses, however, in 16 cases, only one of the recipients produced a response, leaving 490 data points. table 7 shows the actual n values in each case. typing time / turn (ms) typing time / char (ms) condition mean (s.d.) mean (s.d.) n aa 11122.27 (14413.5) 475.45 (558.92) 246 ab 12500.98 (10944.6) 523.56 (1036.00) 89 ba 9800.77 (8810.3) 357.76 (316.15) 92 bb 11561.67 (10138.4) 479.51 (396.07) 63 table 7: typing time of turn and typing time per character by type of intervention 2 × 2 anovas (with participant as a random effect)16 show a significant main effect of floor change17 on the log transformed typing time of turn (see table 7), with participants taking longer over their turns in the ab and bb conditions (f(1,288) = 6.563, p = 0.012). there was no main effect of same/other (f(1,288) = 0.001, p = 0.980), and no effect of interaction (f(1,288) = 1.259, p = 0.270), though there was a main effect of participant (f(59,288) = 4.565, p = 0.008) showing that there was high individual variation for this measure. 16. we account for between subject variation by including subject as a random factor, meaning that there is more than one datapoint per subject (and, in effect, a2× 2× 60 model). there are 490 datapoints between 60 subjects. as we carried out a full factorial model, the numerator (error) degrees offreedom that resulted from this model was 288. 17. a significant effect is one in which the p-value, the probability of the observed data being sampled if the null hypothesis were true, is below some criterion value. we adopt the criteria of significance as lower than 5% probability. 302 compound contributions there were no significant effects on length of turn in characters (same/other,f(1,288) = 1.709, p = 0.194, floor changef(1,288) = 0.194, p = 0.341). a 2×2 gee model with participant as a subject effect (using the gamma distribution, goodness of fit quasi log-likelihood (qic) = 149.482)18 showed a marginally significant main effect of floor change on typing time per character (model effect; wald-χ2 = 3.820, p = 0.051. parameter estimate;b = −0.281, wald-χ2 = 7.192, p = 0.007) and a main effect of participant (waldχ2 = 258468, p < 0.001). there was no main effect of same/other, and no interaction effects. for deletes per character, a2 × 2 gee model with participant as a subject effect (using the negative binomial distribution,19 goodness of fit quasi log-likelihood (qic) = 566.574) showed a significant main effect of same/other (model effect; wald-χ2 = 9.617, p = 0.002. parameter estimate;b = −0.492, wald-χ2 = 12.226, p < 0.001) and a main effect of participant (waldχ2 = 487986, p < 0.001). there was no main effect of floor change, and no interaction effects. condition mean (s.d.) (ms/char) aa 0.108 (0.16) ab 0.094 (0.13) ba 0.071 (0.10) bb 0.138 (0.17) table 8: deletes per character by type of intervention as the experiment was looking for generic effects of ccs on the dialogue,the location of the split points was arbitrary. in order to test for effects of split point, post-hoc analyses were carried out to ascertain whether other observed contrasts in the corpus had any effects on processing of apparent ccs. the fake ccs were coded according to three factors; standalonecoherence (as judged by the authors) of the antecedent and continuation20 (see table 9) and whether the split point fell within or between a syntactic constituent. there were no effects of first or second half coherence on any of the variables, and no interaction effects. there were also no main effects of whether the split point fell within or between a constituent; (log transformed) typing time of turn (f(1,204) = 0.262, p = 0.435); (log transformed) number of characters (f(1,204) = 1.760, p = 0.189); typing time per character (wald-χ2 = 0.550, p = 0.458) deletes per character (wald-χ2 = 0.285, p = .594) and no interaction effects with same/other or floor change. these results are consistent with the finding from the corpus that the split point may be able to occur anywhere syntactically, though the lack of any observed effects could be due to low power caused by the relatively small numbers of some groups. 18. the model distributions were chosen on the basis of being the best fitto the data, as indicated by the lowest quasi log-likelihood score. 19. each key press can be seen as a delete or not-a-delete. 20. these judgements are simply a yes/no answer to the question ‘could thiscontribution be interpreted as complete in its own right?’, i.e. analogous to theend-completeandstart-completeannotation tags in the corpus study, such that an antecedent (first part) judged to be able to stand alone can be considered end-complete and a continuation (second part) judged to be able to standalone can be considered start-complete. the difference in tagging conventions is due to the fact that in the chat tool environment turns can be revised prior to sending, and therefore might be considered to be a unit by its sender, even if the fractured nature of text chat meansthat it might not constitute a syntactically complete sentence. 303 howes, purver, healey, m ills and gregoromichelaki part of cc coherent n first second 1st 2nd what the hell is that y n 131 the woman is pregnant she should stay y y 52 these people said you did something n y 43 i think this is also the wish of the doctor n n 264 table 9: examples of standalone coherence judgement examples 5.3 discussion given the novelty of the method and the lack of other experimental studies of cc’s to cross-check against, the results of this experiment must be interpreted with caution. nonetheless, we believe that the results summarised in table 10, below, do bear on the questions raisedin section 3.2. effect of dependent variable and direction of effect floor change typing time (per turn and char) (ab ∧bb) > (aa ∧ba) same/other deletes (aa ∧bb) > (ab ∧ba) table 10: summary of significant effects firstly, it is important to note that introducing fake ccs did have measurable effects on the ongoing dialogue, despite participants being unaware of either the intervention, or their effects. this in itself might be seen as surprising, as if the intervention were highly disruptive, we would presumably expect subjects to notice it. though typing time is a fairly crude measure21 one possible explanation for participants taking longer over the production of a turn (independently of length of turn in characters) could be due to problems arising in the local organisation of turn-taking (sacks et al., 1974). a participant who has seen a floor change intervention (participant c) may be taking longer overtheir turns because there is less pressure on them to take a turn. c will falsely believe that the fake source (participant b) has just completed a turn, and will therefore not expect them to take the floor. additionally, the genuine source (participant a) will not be taking the floor because they have justcompleted a turn (though c does not know this). however, this effect of floor change could also be due to the confounding fact that when one of the recipients sees a floor change cc, and the other recipient (as always) sees an aa-split, the two are left with different impressions about who made the finalcontribution (i.e. the continuation part of the fake cc) and thus have potentially conflicting expectations regarding who is entitled to speak next. whether or not these explanations are correct, theeffect does suggest that at some level participants are sensitive to specific interlocutors – note that the difference cannot be 21. for example, the additional typing time may fall at the end of a turn (before pressing enter) suggesting that participants are reviewing their responses more carefully before sending them, orit may be a general effect spread evenly across the turn. 304 compound contributions simply attributable to a mismatch between who appears to be speaking and what sort of thing they would say because then we would expect turns following the ba interventionto be equally affected. independently of a change of floor, seeing a cc that appears to be shared between speakers also has an impact on the conversation, seen in the amount of revision undertaken in formulating responses (deletes). perhaps surprisingly, in this case, participants who have seen a cc that was apparently co-constructed by both their interlocutors revise their turns less than after a same-person cc. one reason why participants might worry less about precisely formulating their turns following a cross-person cc is that it could have the effect on the recipient of suggesting that the two other participants are highly coordinated. one possible interpretation of this couldbe that they have formed a ‘party’ (schegloff, 1995) with respect to the decision of who tothrow out of the balloon. this might be understood as signalling the formation of a strong coalition between the other two participants, making the recipient behave as though they are resigned to thedecision of this coalition. (19), taken from the transcripts shows an example where this appears to be the case (the ‘fake’ part of the cc is shown in bold). (19) ab-split showing apparent coalition between ‘b’ and ‘d’ b: and he can tell his formula d: to tom and susie note that this is not the same as the effect on the typing time of turn, whereby participants are less rushed when seeing a change of floor. deletes, in contrast, indicatehow carefully participants are constructing their turns. 6. general discussion and conclusions as discussed in the introduction, ccs are of interest for many different groups of researchers. our corpus study shows that nearly one fifth of all contributions in naturally occurring dialogue continue some previous contribution, indicating the scale of the phenomenon. just thesub-set of cross-person ccs accounts for 3% of all dialogue contributions, comparable to the frequency of clarification requests (see purver et al., 2003; rodrı́guez and schlangen, 2004), widely studied by dialogue theorists (e.g. ginzburg and cooper, 2004). although most of the categories of cc described by conversation analystsappear, these categories do not correspond to the most frequent in the bnc, and do not adequately characterise all the ccs observed in the present analysis. the corpus results show no evidence that syntax places significant constraints on where a split point can occur and the experimentalresults are consistent with this. participants were able to process and interpret fake ccs successfully despite their arbitrary split points and were also not explicitly aware of the experimental manipulation.this is consistent with models that advocate highly coordinated resources between interlocutors and, moreover, the need for incremental means of processing that operate on at least a word-by-word basis (purver et al., 2006; skantze and schlangen, 2009). however, that ccs may be able to occur anywhere in a syntactic sequenceis not to say that they necessarily or usually do. both the corpus study, in which cross-person ccs occur more frequently after an unfilled (rather than a filled) pause and more often follow an end-complete antecedent than same-person ccs (in the not obviously classifiable cases), and the experiment, in which confounding expectations lead to additional response time, suggest that conversational expectations (in305 howes, purver, healey, m ills and gregoromichelaki cluding for example around turn-taking) play some role above and beyondgrammatical/linguistic resources. all ccs other person ccs bnc experiment bnc experiment y 160972% 183 37% 24274% 69 38% n 62228% 307 63% 8726% 112 62% total 2231 490 329 181 table 11: antecedent end-completeness: comparison of the distribution ofactual (corpus) and arbitrary (experiment) split points these considerations are backed up by the data shown in table 11. this table shows the distribution of antecedent end-completeness in the annotated corpus, compared tothe distribution obtained in the experiment. as can be clearly seen, when strings are artificially split in aarbitrary fashion, the ‘antecedent’ is far less likely to end in a complete way than actually occurs inthe genuine ccs in the corpus. this suggests that continuations are systematically designed (and their split points chosen) as extensions of contributions that could be treated as already complete. in terms of the effects that ccs have on the ongoing dialogue, the experiment suggests that lerner’s hypothesis that cross-person ccs might demonstrate party membership may well be correct; it also clearly demonstrates that the participants in a dialogue are not interchangeable. these results do not, of course, prejudice the claim that, at a purely mechanistic level, people could anticipate the structures needed to complete a turn (as the interactive alignment model suggests); they do not tell us about the actual production of compound contributions, but ratherabout the effect they have on the conversation, so do not provide unequivocal evidence in support of one theory over another. they do however indicate that if we wish to treat a jointly produced cc as signalling especially strong alignment, then we need to examine factors other than simply syntax. they also offer some interesting pointers for further research: if parties are genuine conversational entities, then we might expect dialogue phenomena to have distributional patterns which reflect this. consistent with our corpus results as regards other-person repair, for example, is the speculative hypothesis that people might be structuring continuations precisely as the preferred ‘self’-repairs, where ‘self’ can be taken to mean within-party. from a computational modelling point of view, there is some good news: as start-completeness of continuations is rare, a dialogue system may have a chance of detecting continuations from surface characteristics of the input (though note that we did not investigate the general prevalence of start-incomplete s-units in the corpus). there is bad news too, though: as theredo not seem to be strict syntactic restrictions on where the split point can occur, there may beno grammatical features that can be reliably employed to this end. in addition, antecedents do not endin an incomplete way as commonly as might be expected, and long distances between antecedent and continuation are possible. detecting continuations and locating their antecedents is therefore unlikely to be a straightforward task for automated systems. 306 compound contributions 6.1 further work split point for implementational purposes, additional corpus analysis needs to be carried out regarding the distribution of the different syntactic points at which ccs canand do appear; further experiments are also planned which examine the effect on processing of manipulating the syntactic position of the split point (for example, inserting splits before or after determiners). continuation form further corpus analysis is also required to investigate any systematic differences in theform of continuations, and the interaction of this with the properties of the antecedent and conversational genre (including the bnc’s “context-governed” or “demographic” face-to-face dialogues and text-based chat). ownership a further interesting question regards who can be said to take responsibilityfor, or ‘own’ a jointly produced cc. insights into this might come from lerner (2004), which discusses ccs that occur incollaborative turn sequences: if the original speaker maintains authority over the content of the collaboratively constructed contribution, they may ratify another speaker’s continuation (for example by repeating or acknowledging it), but may alternatively strategically offer their own (delayed completion,lerner, 1996). party-membership, or otherwise, may be influenced by such additional conversational contributions. character-by-character experiments due to the design of the experiment, the floor change effects might, as discussed, be because in floor change cases the two recipients will have been left with the impression that a different person made the final contribution. this means that there may be a an effect of confounded listener expectation (though see schober and brennan, 2003, for discussion) – though note that this does not have any bearing on the observed differences on deletes after a cross-person cc. because of this, and the already noted potential problems of linearity in text-based chat, a follow-up study using a character-by-character chat tool interface is already under way. this more directly enforces turn-taking, as it does not allow participants to formulate their turn before communicating it; each character is transmitted as and when it is entered. references g.a. ballinger. using generalized estimating equations for longitudinal data analysis. organizational research methods, 7(2):127–150, 2004. harry bunt. multifunctionality and multidimensional dialogue semantics. inproceedings of diaholmia, 13th semdial workshop, 2009. lou burnard. reference guide for the british national corpus (world edition). oxford university computing serviceshttp://www.natcorp.ox.ac.uk/docs/usermanual/, 2000. urlhttp://www.natcorp.ox.ac.uk/docs/usermanual/. okko buß, timo baumann, and david schlangen. collaborating on utterances with a spoken dialogue system using an isu-based approach to incremental dialogue management. inproceedings of the sigdial 2010 conference, pages 233–236, tokyo, japan, september 2010. association for computational linguistics. urlhttp://www.sigdial.org/workshops/ workshop11/proc/pdf/sigdial42.pdf. 307 howes, purver, healey, m ills and gregoromichelaki ronnie cann, ruth kempson, and lutz marten.the dynamics of language. elsevier, oxford, 2005. jean carletta. assessing agreement on classification tasks: the kappa statistic. computational linguistics, 22(2):249–255, 1996. herbert h. clark.using language. cambridge university press, 1996. herbert h. clark and jean e. fox tree. usinguh andum in spontaneous speaking.cognition, 84 (1):73–111, 2002. jan peter de ruiter and mattia van dienst. completing other people’s utterances: evidence for forward modeling in conversation. ms., in preparation. david devault, kenji sagae, and david traum. can i finish? learning whento respond to incremental interpretation results in interactive dialogue. inproceedings of the sigdial 2009 conference, pages 11–20, london, uk, september 2009. association for computational linguistics. url http://www.aclweb.org/anthology/w/w09/w09-3902. arash eshghi.uncommon ground: the distribution of dialogue contexts. phd thesis, school of electronic engineering and computer science, queen mary university oflondon, 2009. raquel ferńandez and jonathan ginzburg. non-sentential utterances: a corpus-based study.traitement automatique des langues, 43(2), 2002. victor ferreira. is it better to give than to donate? syntactic flexibility in language production. journal of memory and language, 35:724–755, 1996. barbara a. fox. principles shaping grammatical practices: an exploration.discourse studies, 9(3): 299, 2007. andrew gargett, eleni gregoromichelaki, ruth kempson, matthew purver,and yo sato. grammar resources for modelling dialogue dynamically.cognitive neurodynamics, 3(4):347–363, 2009. issn 1871-4080. urlhttp://dx.doi.org/10.1007/s11571-009-9088-y. jonathan ginzburg and robin cooper. clarification, ellipsis, and the nature of contextual updates in dialogue.linguistics and philosophy, 27(3):297–365, 2004. eleni gregoromichelaki, yo sato, ruth kempson, andrew gargett, and christine howes. dialogue modelling and the remit of core grammar. inproceedings of iwcs, 2009. eleni gregoromichelaki, ruth kempson, matthew purver, greg j. mills, ronnie cann, wilfried meyer-viol, and pat healey. incrementality and intention-recognition in utterance processing. dialogue and discourse, to appear, 2011. to appear. markus guhe.incremental conceptualization for language production. nj: lawrence erlbaum associates, 2007. makoto hayashi. where grammar and interaction meet: a study of co-participant completion in japanese conversation.human studies, 22(2):475–499, 1999. 308 compound contributions patrick healey, matthew purver, james king, jonathan ginzburg, and greg mills. experimenting with clarification in dialogue. inproceedings of the 25th annual meeting of the cognitive science society, boston, massachusetts, august 2003. marja-liisa helasvuo. shared syntax: the grammar of co-constructions.journal of pragmatics, 36 (8):1315–1336, 2004. ruth kempson, wilfried meyer-viol, and dov gabbay.dynamic syntax: the flow of language understanding. blackwell, 2001. gene h. lerner. on the syntax of sentences-in-progress.language in society, pages 441–458, 1991. gene h. lerner. on the “semi-permeable” character of grammatical units in conversation: conditional entry into the turn space of another speaker. in elinor ochs, emanuel a. schegloff, and sandra a. thompson, editors,interaction and grammar, pages 238–276. cambridge university press, 1996. gene h. lerner. collaborative turn sequences. inconversation analysis: studies from the first generation, pages 225–256. john benjamins, 2004. gene h. lerner and tomoyo takagi. on the place of linguistic resources inthe organization of talk-in-interaction: a co-investigation of english and japanese grammatical practices.journal of pragmatics, 31(1):49–75, 1999. kung-yee liang and scott l. zeger. longitudinal data analysis using generalized linear models. biometrika, 73(1):13, 1986. issn 0006-3444. tsuyoshi ono and sandra a. thompson. what can conversation tell usabout syntax. in p.w. davis, editor,alternative linguistics: descriptive and theoretical modes. benjamin, 1993. martin pickering and simon garrod. toward a mechanistic psychology of dialogue. behavioral and brain sciences, 27:169–226, 2004. massimo poesio and hannes rieser. completions, coordination, and alignment in dialogue. dialogue and discourse, 1:1–89, 2010. issn 2152-9620. urlhttp://elanguage.net/ journals/index.php/dad/article/view/91/512. matthew purver, jonathan ginzburg, and patrick healey. on the means for clarification in dialogue. in r. smith and j. van kuppevelt, editors,current and new directions in discourse & dialogue, pages 235–255. kluwer academic publishers, 2003. matthew purver, ronnie cann, and ruth kempson. grammars as parsers:meeting the dialogue challenge.research on language and computation, 4(2-3):289–326, 2006. kepa rodŕıguez and david schlangen. form, intonation and function of clarification requests in german task-oriented spoken dialogues. inproceedings of the 8th workshop on the semantics and pragmatics of dialogue (semdial), barcelona, spain, july 2004. 309 howes, purver, healey, m ills and gregoromichelaki carolyn p. rośe, diane litman, dumisizwe bhembe, kate forbes, scott silliman, ramesh srivastava, and kurt vanlehn. a comparison of tutor and student behavior in speech versus text based tutoring. inproceedings of the hlt-naacl 03 workshop on building educational applications using natural language processing-volume 2, page 37. association for computational linguistics, 2003. christoph r̈uhlemann.conversation in context: a corpus-driven approach. continuum, 2007. harvey sacks.lectures on conversation. blackwell, 1992. harvey sacks, emanuel a. schegloff, and gail jefferson. a simplestsystematics for the organization of turn-taking for conversation.language, 50(4):696–735, 1974. emanuel a. schegloff. parties and talking together: two ways in which numbers are significant for talk-in-interaction. situated order: studies in the social organization of talk and embodied activities, pages 31–42, 1995. emanuel a. schegloff. turn organization: one intersection of grammar and interaction. in elinor ochs, emanuel a. schegloff, and sandra a. thompson, editors,interaction and grammar, pages 52–133. cambridge university press, 1996. emanuel a. schegloff, gail jefferson, and harvey sacks. the preference for self-correction in the organization of repair in conversation.language, 53(2):361–382, 1977. david schlangen. from reaction to prediction: experiments with computationalmodels of turntaking. in proceedings of the 9th international conference on spoken languageprocessing (interspeech icslp), pittsburgh, pa, september 2006. michael f. schober and susan e. brennan. processes of interactive spoken discourse: the role of the partner.handbook of discourse processes, pages 123–64, 2003. gabriel skantze and david schlangen. incremental dialogue processing in a micro-domain. in proceedings of the 12th conference of the european chapter of the acl(eacl 2009), pages 745–753, athens, greece, march 2009. association for computationallinguistics. urlhttp: //www.aclweb.org/anthology/e09-1085. kristina skuplik. satzkooperationen. definition und empirische untersuchung. sfb 360 1999/03, bielefeld university, 1999. mark steedman.the syntactic process. mit press, cambridge, ma, 2000. beatrice szczepek. formal aspects of collaborative productions in english conversation. interaction and linguistic structures (inlist),http://www.uni-potsdam.de/u/inlist/ issues/17/, 2000a. beatrice szczepek. functional aspects of collaborative productions inenglish conversation.interaction and linguistic structures (inlist),http://www.uni-potsdam.de/u/inlist/ issues/21/, 2000b. hongyin tao and michael j. mccarthy. understanding non-restrictivewhich-clauses in spoken english, which is not an easy thing.language sciences, 23(6):651–677, 2001. 310 dialogue & discourse 11(2) (2020) 34–73 doi: 10.5087/dad.2020.202 from discursive practice to logic? remarks on logical expressivism rodger kibble r.kibble@gold.ac.uk department of computing, goldsmiths, university of london editor: jonathan ginzburg submitted 11/2017; accepted 08/2020; published online 08/2020 abstract this paper proposes a novel account of the conditional locution as grounded in practices of goaldirected cooperative dialogue. it is argued that a conditional semantics can be obtained within a language fragment that lacks this locution, but supports assertive, inferential and directive practices. we take brandom’s logical expressivist programme as a point of departure, but argue that this programme is empirically flawed as it underestimates the pervasive context-dependence of linguistic items including logical vocabulary. we further take issue with his claim that a discursive practice involving only assertion and inference is sufficient for the conservative introduction and deployment of conditional vocabulary. a more promising route is provided by the introduction of directives, as in so-called “pseudo-imperatives” such as get individuals to invest their time and the funding will follow: this has a conditional sense that if individuals invest their time, then funding will follow. we propose a semantic analysis for these forms which builds on kukla and lance’s account of prescriptives, and argue that our analysis more faithfully captures the “irrealis” nature of conditionals. the analysis is presented in terms of an information-state based dialogue model, with the information state comprising a partitioned commitment store. it is argued that our “dialogical” analysis of conditional reasoning is faithful to brandom’s sellarsian intuition of linguistic practice as a game of giving and asking for reasons. we conclude by contextualising and situating brandom’s programme against the larger field of practice theory, by means of a comparison with the works of sociologist, anthropologist and philosopher pierre bourdieu, and suggest that this comparison reveals further challenges to the expressivist programme. we also take note of narasimhan et al’s recent proposals for agent-based modelling of social practice theory as a possible basis for future development. keywords: dialogue, inference, commitments, conditionals, logical expressivism, brandom, bourdieu, wittgenstein, information state, social practice theory 1. introduction this paper combines a critical appraisal of robert brandom’s notion of discursive practice and in particular his programme of logical expressivism with a novel conjectural account of the genesis of the conditional locution. the constructive proposal departs in many ways from brandom’s account but shares an objective of developing an account of linguistic meaning which is “firmly rooted in actual practices of producing and consuming speech acts” (mie1). to be clear from the start, the analysis involves a conception of meaning in terms of social relations between interlocutors rather 1. the remainder of this paper adopts established abbreviations mie for brandom (1994), ar for brandom (2000), bsd for brandom (2000). c©2020 rodger kibble this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). from discursive practice to logic? than their mental states and intentions. this approach is in sympathy with bender and koller’s position that meaning does not inhere in distributional facts alone, as seems to be claimed by much of the literature on neural language models, but takes a different stance from their assumption that meaning is “the relation between a linguistic form and communicative intent” (bender and koller, 2020, p. 5185). brandom (1983; 1994; 2000; 2008) sets out to show: 1. how one can develop such an account without presupposing semantic or intentional concepts; 2. in particular, how formal logic can be shown to supervene on everyday linguistic practice. we begin by attempting to sketch out some essential concerns of brandom’s programme, according to which utterances are acts in a social space which bring about alterations in the normative status of dialogue participants. brandom’s ambitious programme2, most fully set out in mie, starts from a pragmatist, wittgensteinian approach to language use involving practices such as assertion, inference, and the assessment of normative statuses of discourse participants including committment or entitlement to making propositional claims. in this framework, the essential import of an assertion is that the speaker both takes on a social commitment to justify its content, and claims entitlement to this commitment. entitlement may be claimed by such means as producing evidence for the assertion, or deferring to the testimony of another speaker. interlocutors are taken to maintain deontic scoreboards which keep track of these statuses. on this view, assertion is not a purely monological practice but is ineluctably situated within a context of social interaction. the notion of practice plays an important part, and one goal of this paper is to try to make clear what is meant by a (discursive) practice. practice theory is a loosely defined approach in the humanities and social sciences which takes “practices” as an appropriate level of granularity for scholarly study rather than e.g. rules, concepts, conventions or statistical regularities. joseph rouse (2007a) gives a useful survey of this field, discussing the “range and scope of activities taken by various theorists to constitute ‘practices’”: a minimal definition might class practices as stereotyped patterns of sequences of meaningful performances. rouse classes brandom as a practice theorist who treats “language itself (or ‘discursive practices’) as a paradigmatic application of practice talk”. in the final section of this paper we will attempt to situate brandom within this larger field by means of a detailed comparision with the work of pierre bourdieu (1977; 1991) and tentatively consider how this work may be operationalised using agent-based models (narasimhan et al., 2017). brandom’s inferentialism was worked out in confrontation with the then dominant frameworks of formal semantics, and is largely intended to tackle the kind of problems formal semanticists thought were important. the important new paradigm of distributional semantics (ds, boleda and herbelot (2016)), like brandom’s inferentialism, has also been claimed to be inspired by wittgenstein3, though is arguably at odds with brandom’s reading of the philosophical investigations (wittgenstein, 2009) which argues against identifying linguistic or social norms with statistical regularities. the practice-based approach is intended as an alternative both to formal, rule-based accounts and “regularist” statistical analysis. 2. tutorial presentations such as (wanderer, 2008; turbanti, 2017; loeffler, 2018) generally treat mie, ar and bsd as making up a unitary body of work, though brandom’s introduction to bsd makes clear that this work had a separate genesis and should be regarded as orthogonal to the earlier work. in fact i concur with the cited authors in finding considerable overlap among these works, though i will point out some divergences and shifts of emphasis in the course of this paper. in particular, bsd shifts from using the term “practices” to the more non-committal “practices-or-abilities”, which makes room for a more psychologistic or cognitivist approach. 3. clark (2015) considers this connection to be “tenuous”. 35 kibble much of mie is taken up with showing how one can proceed in a top-down manner to account for phenomena that are more conventionally studied under the banner of linguistic semantics such as: the “meaning” or inferential roles of terms and predicates; conditionals and negation; anaphora; modality; and de re/de dicto distinctions. in this paper we are primarily concerned with one particular aspect of this programme, logical expressivism: this is essentially the thesis that i. linguistic competence involving the use of concepts minimally involves an ability to make inferences: for instance, to correctly use the term red one must endorse the inference that “red” is a colour, if something is red it is not green, and so on; ii. the ability to make and endorse inferences is not the same as the ability to articulate them, as it need not involve the use of logical vocabulary such as if. . . then; iii. logical operators serve to make explicit patterns of inference which are already available in a hypothetical “base” language that lacks this vocabulary; iv. the introduction of these operators is semantically transparent and inferentially conservative, in that their use does not enable any inferences which were not previously available. the expressivist project is presented in bsd in terms of the notions of elaboration and explication. the idea is that a set of basic abilities can be marshalled into a process which implements a higherlevel ability (elaboration), and that one can then define a vocabulary that specifies or codifies this set of practices (explication). in the particular case of logical expressivism, the argument is that this elaboration/explication (lx) relation enables speakers to say whether a particular inference is good or bad, rather than simply treating it as such. the act of “saying” is realised in this case through the conditional locution. this paper takes a close look at some empirical aspects of logical expressivism and in particular at the expressivist account of formal validity: inferences involving logical vocabulary are said to be formally good if they are both materially good, and cannot be made materially bad by uniform substitution of nonlogical for nonlogical vocabulary in the premises and the conclusions (brandom, 2000, p. 55). we adduce a range of empirical evidence that poses serious challenges for this approach, and argue that it underestimates the pervasive effects of context on the interpretation of connectives. as a preliminary, we will need to take a closer look at exactly what is meant by material inference. note to begin with that as far as we are aware, neither brandom nor any of his commentators have offered a fleshed-out demonstration of how such a substitutional methodology would work. there have been a number of formal treatments with varying degrees of faithfulness to brandom’s account: lance and kremer (lance and kremer, 1994, 1996; lance, 2001) employ modal and relevance logics, piwek (2011; 2014) gives a proof-theoretical account, kibble (2004; 2006; 2007) uses amsterdam-style dynamic semantics, and brandom himself, with alp aker, (brandom, 2008) defines an “incompatibility semantics” (is) as an algebra over certain stipulated basic operations. however, none of these proceeds by starting with natural language and using substitutional techniques to (a) demarcate syntactic categories and/or (b) explicate the inferential roles of items which are identified as logical vocabulary. rather, each presentation only informally associates formulas from the formal system with selected (typically, constructed) natural language expressions. in fact, lance and kremer (1994, p. 373) deliberately eschew expressivism understood as an encoding of “social criteria of inferential appropriateness” on the grounds that it could entail cultural 36 from discursive practice to logic? relativism4, while kukla and lance (2009) do not assume any systematic correspondence between “surface grammar” constructions and pragmatic actions, and kibble (2006) suggests that brandom’s mechanism of deontic scoreboards can be detached from the inferentialist project. we argue in what follows that such a thorough-going expressivist programme would in fact face significant difficulties, arising from the pervasive context-dependence and idiomaticity of natural language expressions even including logical terms, and the difficulty of identifying logical vocabulary itself. we will reach a tentative conclusion that fomal validity of inference is not a property of natural language, which rather consists of clusters of micro-practices or wittgensteinian language games. a major theme of this paper is brandom’s rational reconstruction of the genesis of conditional constructions in natural language. the claim is that any discursive practice must include a “core” of assertional practices, since assertions are the only speech act that can be performed without mastery of any other types of locution, and that those engaging in such an autonomous discursive practice or adp will be able to recognise and endorse inferential links between assertions, which are codified or made explicit by the introduction of logical vocabulary (brandom, 2000, pp. 45–78). in other words, the ability to make and endorse “material” inferences is taken to be both necessary and sufficient for acquiring the ability to correctly deploy logical vocabulary, in particular the conditional locution. we will argue that it is not clear whether speaker/hearers would be reliably capable of distinguishing between elaborative (providing further information) and inferential links, and that the introduction of conditional locutions is not inferentially conservative as brandom argues, since they provide a vocabulary for postulating hypothetical/irrealis states of affairs which are not available in a basic assertional practice5. we will instead propose an intermediate stage involving directive acts of the type which have been called pseudo-imperatives in the literature (fox, 2015), such as write a song about pale fire and he’s yours! these acts have the grammatical form of an imperative conjoined with an assertion but are typically interpreted as having the force of conditional statements, and it will be argued that kukla and lance’s account of directives helps us assign a conditional semantics to these constructions in a rather natural way which still slots into brandom’s scheme of elaboration and explication. the first strand in this argument is supported with reference to a corpus study of elaborative and argumentative discourse relations (scholman and demberg, 2017), and we also draw on a number of examples from the multinli6 corpus (williams et al., 2017). we sketch a conjectural genesis of the conditional locution in goal-directed cooperative dialogue, which turns out to fit neatly within brandom’s close coupling of propositional and practical commitments. it should be noted that a purely assertional adp seems lacking for the game of “giving and asking for reasons”, since it would include no locutions for asking the only way to challenge a commitment is by asserting an incompatible proposition. wanderer (2009) argues for the inclusion of a “challenge” locution and this is conceded by brandom (2009b) – note that this development is in fact anticipated by kibble (2006; 2007) 7 . (we should also note belnap’s (1990) trenchant arguments that any analysis of language restricted to the declarative form is fundamentally impov4. (piwek, 2011, p. 48) argues persuasively that the danger of cultural relativism does not arise if we conceive of logical theory as consisting of “materially non-substantive inferences”. 5. note that weiss (2009) questions whether the augmentation of a base language with logical operators can be truly said to be “conservative”. piwek (2011) shows that this can only hold if we assume unlimited time and memory. 6. multi-genre natural language inference 7. (brandom, 1994, p.193) notes the potential utilty of distinguishing queries and challenges as distinct speech acts, though these remarks remain undeveloped and there is little discussion of their impact on the deontic status of interlocutors. 37 kibble erished, while the semantics of questions is a long-established field of study (wiśniewski, 2015), (ginzburg, 2017).) section 2 gives a brief overview of brandom’s notion of discursive practice in terms of deontic scorekeeping. subsequent sections will proceed from the particular to the general: • section 3 presents empirical data which appears to pose problems for brandom’s treatment of formal validity in terms of material goodness of inference; • section 4 reviews some criticisms of brandom’s notion of the introduction of conditionals into a base language via algorithmic elaboration, and proposes an alternative dialogical account of the genesis of conditionals via the introduction of directives; • section 5 sets out some technical details of this proposal, in terms of information state-based dialogue modelling in the style of traum et al. (1999); • section 6 takes a more foundational stance and attempts to situate brandom’s discursive practices within the context of practice theory as exemplified by bourdieu (1977), rouse (2002) and narasimhan et al. (2017). • finally, we draw up some tentative conclusions and indicate directions for future work. there are a number of tutorial accounts of brandom’s approach to semantics and pragmatics, such as wanderer (2008); turbanti (2017); loeffler (2018) which are all to be recommended. this paper differs in emphasis from those works in confronting brandom’s claims with detailed linguistic analysis, and in seeking to clarify his notion of discursive practice in comparison with key work in the established field of practice theory. 2. discursive practice and material inference 2.1 brandom’s inferentialist programme in a nutshell 1. brandom’s inferentialism is a variant of speech act theory, which in contrast to the tradition represented by grice (1957) and searle (1969) treats utterances as actions which have effects in social space rather than primarily as expressions of mental states seeking to modify the mental states of interlocutors. for brandom, the effects of utterances are manifested in the deontic status of discourse participants, that is in the array of propositions which they are committed to justifying if challenged and in the entitlements they may claim to these commitments. 2. the primary units of “meaning” in language are propositions, as the smallest elements for which one can take responsibility, and propositional inferences which are governed by social practice. uttering a proposition is akin to moving a counter in a game, and is subject to socially-governed rules or conventions which determine the preconditions and effects of making a move. 3. entailment is defined in terms of a primitive notion of incompatibility: p incompatibilityentails q if everything incompatible with p is incompatible with q. this notion underlies patterns of inference which are manifest in dialogue moves. 38 from discursive practice to logic? 4. the inferential potential of subsentential units can be determined though a form of top-down compositionality, by substitution of terms within a proposition and comparing the difference this makes to the inferences which are licensed. 5. substitutions may be symmetric or asymmetric in their effects. for example if the substitution inference from (a) benjamin franklin invented bifocals to (b) the first postmaster general of the united states invented bifocals is a good one, then so is the converse inference from (b) to (a) (symmetric), while the material goodness of the substitution inference from (a) thera walks to (b) thera moves does not ensure the material goodness of the converse inference (asymmetric). the first example involves terms and the second predicates. 6. as noted above, mie proposes that the substitution-inferential method can be extended to handle many phenomena which make up the bread-and-butter of formal semantics, such as anaphora and de re/de dicto reference. in particular, it is argued that inferences involving logical vocabulary have the characteristic that their validity is unaffected by substitution of non-logical for non-logical vocabulary. so, if boris is a liar and a scoundrel, then boris is a liar will be a valid inference whatever expressions we choose to insert in place of boris, liar and scoundrel. 7. utterances are considered under two aspects: as things that we do and as ways of saying things. for example: fido is a dog. he has warm blood manifests the inference that dogs are generally warm-blooded, but does not explicitly say that this is how things are (if x is a dog, x is warm-blooded). 8. any discursive practice must minimally include assertion and (propositional) inference; any speaker/hearer who has mastered these practices has all the discursive abilities which are required to deploy logical vocabulary including conditionals and modal operators. the introduction of such vocabulary enables us to articulate and talk about inferences, while maintaining conservativity in that no new inferences will be licensed which are not available to practitioners using a logic-free vocabulary. so if something is a dog, it is warm-blooded; fido is a dog, so fido is warm-blooded manifests the same inference as the example above, but explicitly states it as an instance of a general rule rather than implicitly applying a material inference. 9. a key point is that the notion of “material inference” which brandom adopts from sellars holds that the inference follows immediately by virtue of the content of the expressions, and that “material properties of inference [are treated] as prior in the order of explanation to formal logical properties of inference” (mie: 135). the propriety of inference is a matter of normative assessment: a good (material) inference is one which members of a language community ought to endorse, and which they can reasonably expect others to endorse. this may cover inferences based on scientific knowledge, on tautologies, or general knowledge, as in eastbourne is west of hastings, so hastings is east of eastbourne.; clearly much of this is of a programmatic nature: a recent tutorial account considers brandom’s substitutional mechanism to be “unsatisfactory both for the logician, because it is silent about quantification, and for the linguist, because it hardly accounts for basic predication” (turbanti, 2017, p. 81). unlike the currently dominant formal and distributional approaches to nl semantics it is hard 39 kibble to see how brandom’s claims could be empirically evaluated in a robust manner – though as indicated above, one can extract falsifiable hypotheses from his work and one such hypothesis will be investigated below. nevertheless, we will argue that his framework provides a fruitful environment for investigating the close coupling of language and action, since both are treated within a unitary framework of commitment. 2.2 deontic scoreboards and material inference to repeat, brandom’s approach is concerned with “deontic” attitudes of hearers, and of speakers as self-monitors, rather than intentional attitudes of speakers as in classic speech act theory. in place of beliefs and desires, brandom discusses “doxastic” and practical commitments, corresponding to beliefs and intentions respectively, which interacting agents may acknowledge or ascribe to one another8. the normative dimensions of language use according to brandom comprise responsibility if i make a claim, i am obliged to back it up with appropriate evidence, argumentation and so on and authority by making a claim to which i am assumed to be entitled, i license others to make the same claim. the essential idea is that making an assertion is taking on a commitment to defend that assertion if challenged. there are obvious shared concerns with the notions of commitment developed by hamblin (1970) and walton and krabbe (1995), and subsequently taken up in multi-agent systems (singh, 2000) and computational linguistics (matheson et al., 2000). brandom’s elaborations include the notion of entitlement to commitments by virtue of evidence, argumentation etc; the interpersonal inheritance of commitments and entitlements, and the treatment of consequential commitments and incompatibility. the mechanism for keeping track of agents’ commitments and entitlements consists of deontic scoreboards maintained by each interlocutor, which record the set of commitments and entitlements which agents claim, acknowledge and attribute to one another (claims and acknowledgements are forms of self-attribution). scoreboards are perspectival and may include both explicitly claimed commitments and consequential commitments derived by (material) inference. in the pure model, each agent maintains their own scoreboard and none has any privileged authority: in practice, agents will often defer to others on grounds of specialist knowledge, experience, status or indeed, power. agents may be in a position of claiming incompatible commitments but may not be assessed as entitled to more than one of them (if any). in brandom’s model, entitlement to a commitment mostly arises in one of two ways: by inference from a commitment to which one is already entitled, or by deferral to the testimony of an interlocutor who is entitled to the commitment. stated thus simply, there is an obvious threat of infinite regress on both scores, since it appears we may not in general acquire any entitlements unless there are already commitments that we or our interlocutors are entitled to. brandom finesses this danger by proposing a “default and challenge” model: entitlement to a commitment is often attributed by default, though remaining potentially liable to be challenged by the assertion of an incompatible commitment. to keep things manageable, the more formal treatment in section 4 will mostly deal only with attributions of commitments by one agent to another. material inference is somewhat cursorily introduced as follows: the kind of inference whose correctnesses determine the conceptual content of its premises and conclusions may be called, following sellars, material inferences. (brandom, 2000, p.52) 8. strictly speaking, brandom’s analogue of belief is acknowledgement of a doxastic commitment. 40 from discursive practice to logic? examples given are: “pittsburgh is to the west of princeton, so princeton is to the east of pittsburgh” and ”lightning is seen now, so thunder will be heard soon”. many people encountering brandom’s work find the notion of material inference puzzling and suspicious, particularly in the way it seems to provide free inference tickets for deriving “ought” from “is”. in fact, it seems that the disposition to make or endorse such inferences is taken to be part of the practical ability involved in the mastery of a particular vocabulary or field of activity, as is the ability to recognise incompatibilities among commitments. the key point is that these inferences are seen as normatively governed, as noted earlier, in that any rational member of the language community ought to endorse them and can reasonably expect others to endorse them. this may encompass inferences based on, for example, accepted scientific knowledge, rules of mathematics, or general knowledge: example 1 (a) the kettle is heated to 100◦c so it will boil. (b) brockley is in london and london is in england, so brockley is in england. (c) two angles of a triangle are to 40◦ and 60◦, so the third is 80◦. it will be noted that some of these inferences are defeasible: the boiling point of water depends on the atmospheric pressure, while the geographical extent of cities can be changed by administrative fiat. material inferences are seen as pre-logical and do not involve tacit formal reasoning involving a hidden premise or enthymeme such as “if lightning is seen, thunder is heard shortly after”. nor does this normative approach make use of carnapian “meaning postulates” (carnap, 1952), a standard device for capturing “lexical” entailments in formal semantics such as ∀x(b(x)→ ¬m(x)), where b denotes the set of bachelors andm the set of married people. rather, such conditional statements are seen as formally making explicit the content of the inference, with the result that inferences themselves can become topics of scrutiny and discussion. brandom’s “origin myth” for formal reasoning postulates an introduction of logical vocabulary such as conditionals into a language in which inferences may be performed but not talked about, so that we are subsequently able to say what we can only do using the relatively impoverished language. this is subject to a principle of conservativeness, in that the new vocabulary must not license any new inferences involving the old vocabulary (op cit, p. 68). this theme is systematically developed in brandom (2008) and is considered further in the next section. 2.3 action and practical commitments brandom’s account of action and intention is initially quite similar to the doxastic story in its overall structure: the role of intentions is taken by practical commitments which can stand in inferential relations to doxastic or to other practical commitments, and to which one may be entitled or not entitled. it is notable that practical commitments can be inferred from doxastic commitments and vice versa, as in the examples below: example 2 i. only opening my umbrella will keep me dry, so i shall open my umbrella. 41 kibble ii. i am a bank employee going to work, so i shall wear a tie. (brandom, 2000, p. 84) iii. i have a fever, so i shall see the doctor. piwek (2011) brandom argues that these inferences are not enthymematic, relying on suppressed premises “i wish to stay dry” or “bank employees should wear ties”, but that (i.) and (ii.) above are further examples of “material inference”: the consequent follows from the antecedent by virtue of its content, and the putative “suppressed premises” are ways of making explicit the implicit norms or preferences that make the inferences go through. practical commitments are taken to stand in inferential relations with both doxastic and other practical commitments, and an action is taken to be rational if it fulfils a practical commitment for which the agent can give a reason. for example: “why are you wearing a tie?” “i’m on the way to work”. putting things a little more technically: to demonstrate entitlement is to offer a chain of reasoning which terminates in a practical commitment which is compatible with one’s other acknowledged commitments, and actions result from “reliable dispositions to respond differentially to the acknowledgement of certain sorts of commitments” (brandom, 2000, p. 83). scorekeepers are licensed to infer agents’ beliefs from their intentional actions [ibid.]. 2.4 compositionality further, material inference has a role to play in analysing the semantic content of subsentential expressions: two subsentential expressions of the same grammatical category share a semantic content just in case substituting one for the other preserves the pragmatic potential of the sentences in which they occur. . . a pair of sentences may be said to have the same pragmatic potential if across the whole variety of possible contexts their utterance would be speech acts with the same pragmatic significance. . . (brandom, 1994, pp. 128-9) so for example, one might say that two terms have the same denotation (“representation”) if replacing one with the other makes no difference to the appropriate circumstances in which a speech act may be uttered and its pragmatic consequences, in terms of the speaker’s deontic score. much of the second half of mie consists of elaborations of this substitutional technique to handle the traditional subject matter of formal semantics such as reference, anaphora, deixis, quantification and propositional attitudes. kremer (2009) offers a detailed examination of the decompositional strategy of analysing the content of subsentential expressions, and identifying different subcategories such as terms and predicates according to the contribution they make to the inferential potential of propositional utterances. for example: the fact that one can infer thora is a mammal from thora is a dog, but not vice versa, indicates that mammal and dog are predicates which licence asymmetric substitution inferences, rather than terms which may license symmetric inferences (brandom, 1994, pp. 133ff). kremer argues that brandom’s account is plagued with circularity, since it claims to define syntactic categories in terms of substitution inferences but turns out (on kremer’s account) to assume a prior grasp of these very categories. one could add that the substitutional techniques are presented in 42 from discursive practice to logic? rather general terms, using simple examples, and would constitute a formidable machine learning problem if applied to corpora of actual discourse. for one thing, it is unlikely that any corpus would provide instances of “all possible contexts” for any given sentence-pair (see above), and this in any case assumes agents with unlimited cognitive resources. as noted, we argue below that brandom’s programme severely underplays the pervasive nature of context-dependence in natural language understanding, such that it is problematic to posit invariant semantic content/inferential potential for particular words or phrases even including “logical” expressions. 2.5 two aspects of logical expressivisim the expressivist project can be approached in various ways. one is to start by modelling forms of inferential practice which underly or are implicit in the use of ordinary vocabulary, and investigate how these practices can be described or articulated with the aid of logical vocabulary. this is the approach taken by piwek (2011), who postulates various inferential practices employed in cooperative information-seeking dialogues in a medical domain, such as deriving sd (“see a doctor”) from ht (“has a temperature”). these inferential practices are modelled as natural deduction inference rules – for example, a rule labelled ht to sd “stands for the practical ability to derive sd when ht already follows” (op cit:266). these inference rules are supplemented with introduction rules which will, for example, license the assertion of ht→ sd just in the case that the rule ht to sd applies: ht ∈ γ {ht} ` ht {ht} ` sd ∅ ` ht→ sd this is an elegant proposal which ensures that logical connectives will only be employed when licensed by inferences which are implicit in the knowledge base, and so the requirement of conservativity will be satisfied. an alternative way to approach the issue is to start from everyday sayings and conversations and investigate whether a class of vocabulary items can be identified which behave in the same way as logical connectives in a system like piwek’s: that is, they invariably articulate or denote patterns of inference which are implicit in the knowledge base, by virtue of the nonlogical content of their arguments. this is the path suggested by brandom in his distinction between material and formal goodness of inference: . . . an inference can be treated as good in virtue of its form, with respect to [some privileged subset of] vocabulary, just in case it is a materially good inference, and it cannot be turned into a materially bad one by substituting nonprivileged for nonprivileged vocabulary in its premises and conclusions. (brandom, 2000, p. 55) it is not obvious a priori that any such class must exist. piwek acknowledges that in principle, there is nothing that excludes the existence of a language community which can make practical inferences involving materially substantive propositions (such as ht and sd), but which has not mastered any logical vocabulary. (ibid.) 43 kibble one could also argue that nothing in principle excludes the existence of a community that has no universally applicable logical vocabulary, but which either uses different vocabularies or interprets logical expressions in different ways in the context of particular domains. and indeed it would be quite in the spirit of brandom’s wittgensteinian heritage if even logical terminology turned out not to have a common core of meaning or inferential potential across all contexts of use, but rather manifested “family resemblances” (wittgenstein, 2009, § 65–71). this will be the topic of section 3 in this paper. 2.6 summary in summary, participation in a discursive practice in brandom’s terms minimally involves: • ability to deploy a vocabulary in ways which are acceptable to other members of a speech community; • ability to make and endorse a variety of material inferences; • ability to keep score of commitments undertaken by interlocutors and oneself, to recognise incompatible commitments, and to ascribe both entitlements and consequential commitments to participants in a discourse; • ability to challenge other practitioners who are assessed as not entitled to particular commitments. note that none of these bullet-points specifically mentions meanings, beliefs or intentions, but it is claimed that a practice involving these abilities can count as a linguistic or discursive practice. note also that the practice involves an abstract notion of a “scoreboard” and is essentially normative, concerned not so much with observed practices as with what agents ought to be able to do to count as engaging in dialogue. 3. context-dependence of interpretation to recapitulate: according to the expressivist account of formal validity, inferences involving logical vocabulary are said to be formally good if they are both materially good, and cannot be made materially bad by uniform substitution of nonlogical for nonlogical vocabulary in the premises and the conclusions (brandom, 2000, p. 55), as in the following example: example 3 i. all ravens are black. matthew is a raven, therefore matthew is black. ii. all babies are cute. oscar is a baby, therefore oscar is cute. the implication is that logical connectives have an invariant meaning which is preserved across linguistic contexts. this is an empirical claim, and this section aims to marshall evidence which puts the claim in doubt. this section approaches the issue from two directions. on the one hand, we investigate whether it is possible to identify formally valid inference patterns in natural language according to brandom’s substitution method; it turns out that with some ingenuity, counter-examples can be constructed that invalidate a variety of inferences which appear materially good on first inspection. on 44 from discursive practice to logic? the other, we discuss whether the types of inferences licensed by proof-theoretic schemes such as natural deduction are actually manifest in nl, using connectives which are commonly associated with boolean operations. both strategies are concerned with the question of whether there really is a route from natural language discourse to at least some of the commonly accepted inference rules for formal logics, without the digression of constructing and stipulating interpretations for formal or controlled languages. the following sections look in detail at a number of candidates for consideration as “logical vocabulary”, namely and, or, if. . . then, may/must, not and all. few of these examples are particularly novel, but taken together they pose a substantial challenge to the claims of logical expresssivism. 3.1 examples of “logical” vocabulary 3.1.1 and well-known issues with and involve non-intersective adjectives: while (b-f) in example 4 below can all be inferred from (a), the same does not hold for example 5 with the non-logical term “brilliant” substituted for “irish”, which implies that the inferences in (4) are not all formally valid. example 4 (a) joyce is an irish author and a father. (b) joyce is irish and joyce is an author. (c) joyce is irish. (d) joyce is an author. (e) joyce is a father. (f) joyce is an irish father. example 5 (a) joyce is a brilliant writer and a father. (b) joyce is brilliant and joyce is a writer. (c) joyce is brilliant. (d) joyce is a writer. (e) joyce is a father. (f) *joyce is a brilliant father. kamp (2004) assumes the following classification of adjectives: predicative: if every n1 is an n2, then every an1 is an an2. examples: fourlegged, superconductive. 45 kibble privative: no an is an n: fake, false, former(?)9 affirmative: every an is an n: big, pink, clever extensional: every an is a: red, female . . . “clearly, all predicative adjectives are extensional. non-extensional adjectives are, for example, affectionate and skillful” (op cit: 543). clearly, goodness of material inference under substitution of nonlogical terms depends on selecting adjectives from the appropriate class, rather than freely substituting nonlogical vocabulary. so the inferences in (4) go through because irish is treated as extensional, while (5) fails because brilliant is not. we should note that kamp carefully qualifies these distinctions with the phrase “in an interpretation”: adjectives need not manifest the same behaviour in all contexts. in fact it is debatable whether any adjectives are truly extensional: e.g., colour terms are highly context-dependent compare white paper, white wine, white coffee, white rhino, while posner (2004, p. 640) reports that the same colour can be called rot “red” or braun “brown” in german, depending whether it occurs on a cotton coat or a plastic wall respectively. inferences under conjunction conjunction in ordinary english typically licenses various inferences which are not reducible to the accepted truth-functional definition (examples 6 and 7 adapted from posner (2004)). successivity there is a strong implication that conjoined sentences narrate events in the order of occurrence: so (6c) implies (6e) while (6d) implies (6f). neither (6e) nor (6f) can safely be inferred from commitment to (6a) and (6b) separately. example 6 (a) annie married peter. (b) annie had a baby. (c) annie married peter and had a baby. (d) annie had a baby and married peter. (e) annie had a baby after she married peter. (f) annie married peter after she had a baby. connexity there is an implication that conjoined sentences relate to the same situation. so (7d) is felt to imply (7f) while (7e) implies (7g). it would be uncooperative to utter either (7d) or (7e) if one is committed to all of (7a-c). 9. though arguably, if a past us president becomes president of a university, they are both a president and a former president. 46 from discursive practice to logic? example 7 (a) the door is open. (b) the window is open. (c) there is a draught. (d) the door is open and there is a draught. (e) the window is open and there is a draught. (f) there is a draught from the open door. (g) there is a draught from the open window. note that these implications are cancellable: one may say for example, “annie married peter and had a baby, but not in that order”, or “the door is open and there is a draught, but it’s not from the door”. when and behaves like a conditional examples below are taken from the snli and multinli corpora (bowman et al., 2015; williams et al., 2017). these corpora were constructed to support the development of techniques for recognising entailments in nl text, providing a copious variety of inferentially linked sentence-pairs based on corpus data and incorporating native speaker judgments. both corpora include instances of nl text that were originally constructed for other purposes (referred to below as premises), and each example was given to participants who were asked to produce new sentences which were entailed or contradicted by, or merely compatible with the original example (referred to as conclusions). instructions were given in non-technical language: participants were asked to say whether their sentence was definitely correct, might be correct or was definitely incorrect in the specified context. the data thus allows insight into informal everyday inference with minimal contamination from expert ideas about entailment, deduction or what have you. in particular, this somewhat ecumenical treatment shows that naive informants are prepared to treat non-declarative utterances as premises in an inference, pace belnap (1990)10 . subjects seem to have interpreted the instructions differently; some have effectively given a paraphase of the original rather than an entailment. the snli corpus consisted entirely of descriptive captions for photos on the flickr website, while multinli uses written and transcribed material from a variety of genres. the material is provided in plain text, json and parsed variants. the examples discussed in this section were obtained by writing a script to extract sentence pairs with the following characteristics: • the pair is flagged as “entailment” • the premises do not contain the word “if” 10. belnap calls the assumption that all utterances share a cluster of properties, including the potential to act as premises for inference, the “declarative fallacy”. a reviewer points out that premises or conclusions can be questions in erotetic logic (wiśniewski, 2013) 47 kibble • the premises may consist of two or more clauses (as a rough and ready test, we looked for the word and) • the conclusion begins with “if”. a selection of the sentence pairs were manually extracted for discussion in this paper11. it was hoped that this would furnish instances of the introduction of logical vocabulary to form a conditional where the premise contains inferentially linked conjuncts, and this hope was satisfied to an extent. however there is a key difference in that while brandom envisages conditionals as arising within a purely assertional practice, many of our examples consist of an imperative followed by an assertion, which is interpreted as entailing a conditional: do x and y if(done x) then y 12. importantly, these constructions are not necessarily interpreted as entailing either of the conjuncts independently, as might be expected via ∧-elimination: e.g., the speaker of example 9 is probably not advising the hearer to sneeze or develop a fever. example 8 (a) get individuals to invest their time and the funding will follow. (b) if individuals will invest their time, funding will come along, too. example 9 (a) sneeze in the middle of the night and your edokko neighbor will demand the next morning that you take better care of yourself; stay home with a fever and she will be over by noon with a bowl of soup. (b) if you develop a fever which results in you staying home, your edokko neighbor will arrive by midday with soup. example 10 (a) truce. you stop this train, everyone lives. (b) if you stop this train everyone will live. literary examples and proverbs example 11 (a) use every man after his desert, and who should ’scape whipping? (hamlet) (b) if you treated everyone as they deserved . . . example 12 (a) feed a cold and starve a fever. 11. a fuller list can be found in the appendix. 12. these constructions are known as pseudo-imperatives in the literature fox (2015). 48 from discursive practice to logic? (b) if you feed someone who has a cold, they will develop a fever . . . 13 the reader may consider that i am using the term “entailment” rather loosely in talking about the entailments of an imperative. in section 4 i will try to tighten up this talk, and will explore a conjecture that the conditional form as manifested in the (b) sentences above is ultimately derivative on complex commands like those in the (a) sentences. i will argue that this approach leads to a more satisfying account of the introduction of conditionals into a base language than brandom’s so-called algorithmic elaboration. 3.1.2 or when or behaves (somewhat) like conjunction jennings (2004) discusses this connective at some length and contends that it does not typically behave like truth-functional disjunction, even in the context of selected examples from logic books. for instance, he argues that or in example (13), attributed to patrick suppes, is not a case of exclusive disjunction as suppes maintains: example 13 father to child: “you may go to the movies or you may go to the circus this saturday but not both”. (op cit, p. 665) this is not a disjunction because the child may legitimately infer that they have permission to go to the movies, and that they have permission to go to the circus. however, the truth-functional definition of xor does not licence p xor q ` q, or p xor q ` p. the meaning of (13) can be more accurately captured as: may(movies and not circus) and may(circus and not movies), where there is no disjunction in sight. jennings (jennings, 2004, p. 671) claims that “all the connective vocabulary of any natural language has descended from lexical vocabulary” e.g. or from oe odher “other”, but from oe butan “outside”, while “much of the vocabulary retains residual nonlogical . . . uses (since, then, therefore, . . . )”. while brandom’s story of logical vocabulary being introduced into a language lacking such terms is not to be taken too literally, it may well be the case that a class of lexical terms have acquired logical uses in the course of a language’s history but a residue persists of non-logical uses. 3.1.3 if/then we would like to think that a simple conditional “p if q” will always imply “if q then p”, but it’s not so simple as lakoff showed in his paper on “natural logic” (lakoff, 1970): example 14 (a) i think sam will smoke pot14, if he can get it cheap. (b) if he can get it cheap, then i think sam will smoke pot. 13. “as robert graves has pointed out, this archaic source of error treats the two apparent imperatives as separate injunctions, whereas what we have is in fact a conditional sentence: if you feed somebody with a cold, he [sic] will develop a fever and them you will have to starve him” (amis, 1999). admittedly this proverb gives doubtful advice on either this or the more literal reading. 14. pot: cannabis (archaic). 49 kibble example 15 (a) i realize that sam will smoke pot, if he can get it cheap. (b) *if he can get it cheap, then i realize that sam will smoke pot. the reader may object that this is not a simple conditional, but the conditional occurs as the complement of an attitudinal verb. nevertheless the fact remains that what is a good inference in one case is turned into a bad one simply by substituting nonlogical vocabulary items. (a reviewer argues that examples (14) and (15) differ in that i think functions parenthetically whereas i realize actually embeds a conditional under the attitudinal verb. this relies on a degree of syntactic analysis which is not supposed to be available under brandom’s substitutional methodology. the parenthetical function can be made unambiguous through punctuation: if he can get it cheap, then, i think, sam will smoke pot.) 3.1.4 modality and negation we would expect to treat inferences of the form if p then must-q implies if p then ¬may¬q as formally valid. however: example 16 (a) shoes must be worn in the library. (b) [if you are in the library, you must wear shoes] (c) you may not use the library if you are not wearing shoes. example 17 (a) dogs must be carried on the escalator. (b) [if you ride the escalator, you must carry your dog] (c) *you may not use the escalator if you are not carrying dogs. clearly, the distinction relies on cultural knowledge about expected behaviour in different contexts. the natural understanding of 16(a) is consistent with the above equivalence, whereas 17(a) needs to be unpacked as if you ride the escalator with your dog, you must carry it. again, the point here is that uniform substitution of non-logical vocabulary can fail to preserve logical inference15. 3.1.5 quantification and compound nouns example 18 (a) all horses are animals. so, all horse tails are animal tails. (de morgan, quoted by van benthem (2007)). (b) *all horses are animals. so, all horse boxes are animal boxes. 15. a reviewer states that the natural readings for these sentences rely on distinct intonational patterns; however, these admonitions are more likely to be found on printed signs. 50 from discursive practice to logic? (c) all dogs are animals. so, all dog coats are animal coats. (d) *all minks are animals. so, all mink coats are animal coats. according to van benthem, example (a) was routinely presented to dutch logic students in the 1960s as part of their initiation into the predicate calculus. the inference (b) fails, i contend, because a horse box is not the same kind of thing as an animal box: the former is not a box at all, but a vantype vehicle for transporting horses, while the latter might be understood (if at all) as a box of treats for pet animals such as toys, accessories and hygiene products, or perhaps a small container for carrying pets. likewise (d) fails if we read “animals coats” in a natural way as “coats for animals” as in the unobjectionable (c): the natural reading of “mink coat” is of course a coat made from the pelts of minks. bauer and tarasova (2013) discusses nominal compounds, using the following classification from levi (1978): n1 cause n2 sex scandal, withdrawal symptom n2 cause n1 tear gas, shock news n1 have n2 lemon peel, school gate n2 have n1 camera phone, picture book n1 make n2 court order, n2 make n1 computer industry, silk worm n2 use n1 steam iron, wind farm n2 be n1 island state, soldier ant n2 in n1 field mouse, letter bomb n2 for n1 arms budget, steak knife, machine oil n2 from n1 business profit, olive oil n2 about n1 tax law, love letter, business news bauer’s discussion exclusively concerns endocentric compounds: that is, every n1 n2 is an n2. so while horse tails, animal tails, mink coats are endocentric, horse boxes are not, which may account for the failure of (b). and while animal coat, dog coat are instances of n2 for n1, mink coat is a case of n2 from n1, and thus (d) is blocked. as with the different types of adjectives mentioned in section 3.1.1, material goodness of inference can only be preserved under inference if substituted items are limited to an appropriate class; a difficulty is that these associations are often idiomatic and idiosyncratic and can’t be reliably predicted from the content of the two nouns on their own. 3.2 summary to summarise: in this section we have considered various examples which appear to challenge the notion that logical connectives can be extracted from natural language practice by identifying invariant uses under substitution of the accompanying non-logical vocabulary, owing to the pervasive influence of linguistic and non-linguistic context on the behaviour of lexical items including even “logical” terms. the evidence above suggests that when we look closely at the behaviour of socalled ‘logical’ connectives, we do not find uses which are invariant under substitution but rather, clusters of evolving micro-practices or language games associated with the use of particular items and constructions. in the next section i investigate one of these practices in some detail, proposing a non-standard account of commands and commitments which may offer a plausible model for the conditional sense of complex directives. 51 kibble 4. towards a dialogical account of conditionals 4.1 algorithmic elaboration one specific type of elaboration focussed on in brandom (2008) is the introduction of conditionals into a purely assertional autonomous discourse practice. brandom claims (op cit, p45) that an ability to respond appropriately to conditional statements of the form “if p then q” can be acquired by an agent that has mastered core elements of an adp such as the abilitites to make assertions/acknowledge assertional commitments, to acknowledge inferential commitments relating certain statements p and q of the practice, and to reliably respond differentially to assertions or changes in one’s deontic attitude by altering the score. (loeffler, 2018, p. 153) such development of new discursive practices involving the exercise of “exercising the right basic abilities in the right order and under the right circumstances” (brandom, 2008, p. 26) is dubbed algorithmic elaboration, though it does not involve anything a computer scientist would recognise as an algorithm. turner (2008) notes that practices can be “underdetermined” and it may not always be obvious which practice is instantiated by a particular performance. in this instance, it is not necessarily clear how one could infallibly recognise a practice of “accepting or rejecting an inference”, particularly if we assume a base language whose repertoire does not extend beyond assertions. returning to example (1), but omitting the logical particle so: 1. (a) i am a bank employee going to work. (b) i shall wear a necktie. there is clearly scope for ambiguity, or at least vagueness, over whether the speaker is expressing an inference from (a) to (b), or simply providing more information about his current activities. that is, the relation between (a) and (b) could be analysed in rst terms as elaboration rather than, say, volitional cause (mann and thompson, 1987; taboada and mann, 2006). scholman et al (2017) noted that there is substantial disagreement among annotators over whether implicit relations between text spans should be classed as “elaborative” or “argumentative”, looking in particular at annotations of the wall street journal corpus using the pdtb (prasad et al., 2008) and rst (carlson et al., 2003) frameworks. in what follows we will outline an alternative rational reconstruction of the introduction of conditional reasoning into a discursive practice, building on a basic practice which is taken to minimally include assertions, challenges and commands. 4.2 methodological issues with the lx programme one difficulty for brandom’s programme of logical expressivism is that it could only be conclusively verified by observing inferential practices before and after the introduction of logical vocabulary. however, the idea that this vocabulary is “introduced” into a linguistic practice which had previously lacked such terms is a fiction, or at least an unverifiable speculation. logical expressivism therefore has to be interpreted as a claim that logical vocabulary allows us to codify practices which can be manifested in a language that has been stripped of such vocabulary. in this situation, it is hard to see how one could identify the “basic” or “core” practices as distinguished from practices that may 52 from discursive practice to logic? have been altered in the light of reflections facilitated by logical vocabulary. furthermore, brandom acknowledges that once the logical vocabulary has been introduced, it may induce practitioners to alter their prior practice, in the light of what it now allows them to say about that practice. (brandom, 2009a, p. 354) macfarlane (2008) argues that the ability to algorithmically elaborate discursive practices is not itself one of the core abilities manifested in basic assertional practices, and that it involves other novel abilities such a syntactic ability to combine sentences using operators. furthermore, i would argue that the proper use of conditionals assumes an ability to envisage and describe hypothetical/irrealis states of affairs, which is not needed for a basic assertional practice. a sequence a, so b involves commitment to both a and b, but if a then b on its own involves commitment to neither. for the purposes of this paper, i will assume that the goal is to investigate conceptual dependencies between practices such as asserting, commanding and exhibiting conditional reasoning, and i will argue that a rational reconstruction of the genesis of the conditional locution can be achieved if asserting and commanding are given, but not on the basis of assertive practices alone. that is not necessarily to say that these practices can be observed as chronologically prior to conditional reasoning either in language evolution or acquisition, though it suggests interesting possibilities for applied research. a further technical point: weiss (2009) discusses the logical consequence relation defined in brandom (2008) which is based on a primitive notion of incompatibility, and assumes that speakers are able to determine “a fully determinate incompatibility relation between arbitrary finite sets of sentences”. he raises the issue that this may be beyond the reasoning capacities of speakers of the base language, but that the introduction of logical operators may enable them to “decide undetermined incompatibilty relations” (emphasis in original). it is not clear that brandom satisfactorily addresses this specific point in his reply to weiss (brandom, 2009a), while as noted above piwek (2011) shows that conservativity requires unbounded resources and processing time. 4.3 introducing “complex directives” it is a constant refrain in brandom’s work that language games (or autonomous discursive practices, adps) must include “practices of giving and asking for reasons, because assertions, the most basic kinds of sayings, must be capable of both serving as and standing in need of reasons ” (brandom, 2008, p. 43). (although in fact brandom’s basic assertional practice does not appear to provide a means of asking for reasons, as noted above.) another kind of saying or discursive practice which can lead to demands for reasons is giving commands or orders; thus it seems desirable that any practice which includes commands should also include assertions16. the multinli corpus includes numerous examples where the premise consists of a command followed by an assertion which can be construed as immediately giving a reason to follow the command: example 19 (a) walk through the gateway and you’ll find yourself in an immense open space the largest temple courtyard in the country. 16. in fact it has been claimed that “commands” precede “assertions” in both the development and evolution of language (see e.g. tomasello and camaioni, 1997). 53 kibble (b) if you walk through the gate, you will be in the largest temple courtyard in the country. many of these kind of examples, of which there are several in the corpus, can be interpreted as “entailing” the first conjunct17, in contrast to some examples we considered above: the speaker of (19a) does appear to be advising the hearer to walk through the gateway. the second conjunct is arguably also entailed if we treat this as a kind of dynamic conjunction: the assertion is claimed to be correct or licit in the context created by executing the command. note that adding commands to an assertional practice introduces the notion of a hypothetical (irrealis) state of affairs, since it is clearly desirable that we can envisage and reason about the expected consequences of following a command. the second conjunct thus has the dual role of describing the consequences of executing the command, and giving a reason to do so. interpreting the sentence in this way thus involves conditional reasoning, although no explicit conditional construction is present. these constructions are known in the literature as pseudo-imperatives (fox, 2015), which have the sense of conditionals rather than actual commands or instructions. however, we can identify a subclass which have the force of genuine directives, as noted above: the speaker is in fact instructing or advising the hearer to do something. i will dub these cases complex directives, to be distinguished from pseudo-imperatives, and will suggest that they be treated as a basic form from which the latter derive – and, more speculatively, that they provide the essential ingredients for the conditional locution itself. 4.4 a closer look at directives before proceeding, it will be useful to operate with a finer-grained classification of directive speech acts (acts which strive to bring it about that another person does something), which will draw on the analysis of rebecca kukla and mark lance (kukla and lance, 2009, 2010; lance and kukla, 2013). lance and kukla propose a small number of assumptions which turn out to yield a fair amount of mileage: • terms like “imperative”, “declarative” are pragmatic categories which do not necessarily have any systematic correspondence with what they call “surface grammar”; • speech acts are treated as functions whose inputs and outputs are normative statuses; • inputs and outputs are classified as agent-neutral or agent-relative: for example a truth-claim will be agent-neutral since its truth or falsity holds (in principle) for everybody, while a directive has an agent-relative output which targets the addressee; • directives are subcategorised as prescriptives which articulate a pre-existing commitment, and imperative which assert some authority to instruct someone to fulfil a commitment. imperatives can be distinguished as those which create a commitment (dubbed “constative”) and those which articulate or ostend an existing commitment (“alethic”). i will make a further distinction between deontic and prudential practical commitments, i.e. commitments which arise from obligation and those which arise from self-interest. with this terminology in place, kukla and lance’s four ways of getting someone to do something (kukla and lance, 2009, pp. 105ff)consist of: 17. assuming the permissive notion of “entailment” discussed above. 54 from discursive practice to logic? 1. third person prescriptives, with agent-neutral input and both agent-relative and agent-neutral outputs, whose primary effect is to make a truth claim about a normative commitment, e.g. “scott needs to lose weight”; 2. second person prescriptives, again with agent-neutral input and both agent-relative and agentneutral outputs, whose primary effect is to ostend or draw attention to a practical commitment: “you [scott] need to lose weight”; 3. alethic imperatives, with agent-relative inputs and outputs, which not only ostend but seek to hold the addressee to a pre-existing norm: “please lose weight” 4. constative imperatives, with agent-relative inputs and outputs, which both create and seek to hold the addressee to a new practical commitment: “move away from the vehicle”, spoken by a police officer. prescriptives have both agent-relative and agent-neutral outputs since the claim they make is taken to be true for everyone, while the commitments referred to apply to a specific person or group. (that is, the speaker expects everyone to agree that scott is overweight, but only scott to embark on a weight-reduction programme.) imperatives but not prescriptives have agent-relative inputs since some particular authority or status is required to issue a request or command. in this case, the speaker might be scott’s physician, partner or fitness trainer. in all of cases 1-3, the speaker might be guilty of rudeness but only in (3) would the utterance “misfire” pragmatically if the speaker lacks the appropriate normative relationship to scott. following this typology, we can identify examples from the corpus of complex directives where the (i) clause can be clearly labelled as prescriptive – despite having the “surface form” of an imperative – in that the speaker is not imposing any obligation on the hearer but is advising them of their best interests: example 20 (a) i. take a left turn here on the b6318 and ii. after 6 km (4 miles) you will see the sign for birdoswald. (b) if you travel for 6 kilometers after taking a left turn you’ll see the sign for birdoswald. example 21 (a) i. click nobel prize internet archive hayek page , and ii. you’ll find yourself ... (b) if you click the nobel prize internet archive... a reminder: in each of examples (20–21) above, the (a) sentence is taken from the corpus and the (b) sentence is judged by a naive subject to be something that follows from it: so these complex directives are taken to imply conditionals. i argue that the (i) clauses meet lance and kukla’s criteria for prescriptives rather than imperatives, despite sharing a grammatical form with imperatives, on the following grounds: 55 kibble • the speaker is not seeking to hold the hearer to a commitment, but is foregrounding and making a truth-claim about or attributing a prudential practical commitment. the speaker’s utterance does not create this commitment, which is rather inherent in the addressee’s particular situation. assuming a context where in each case the hearer has a goal or practical commitment described by the (ii) clause, this goal commitment brings with it the prudential commitment identified in the (i) clause. • the claim that the hearer has this commitment is true or false for everyone, thus agent-neutral, while the intended outputs include an acknowledgement on the part of the hearer that they have this commitment – thus agent-relative. • while the prudential commitment identified in the (i) clauses is a material-inferential consequence of the goal commitment, the reverse is the case for entitlements: it is commitment to the performance of the action prescribed in (i) which will entitle hearer to commit to the goal clause. if we accept the analysis that the prescriptive clause carries a truth-claim, i propose that this propositional claim can effectively be treated as the antecedent clause of a conditional. in summary: knowledge of the hearer’s goal entitles the speaker to articulate a prudential commitment on the hearer to take a particular action. hearer’s acknowledgement of this commitment will further commit them to achieving the state of affairs described in the (ii) clause. so the complex directive in the (a) sentences manifests the structure of a conditional as in the (b) examples. if we compare the above examples with (22) and (23) below, we see that the imperative force is weakened and the conditional sense is foregrounded in the latter. whereas both (20) and (21) give specific, targetted advice in a concrete context, (22) gives more generic advice while (23) is not at all advising the hearer to sneeze or catch a fever; the conditional sense has become primary. example 22 (a) i. get individuals to invest their time and ii. the funding will follow. (b) if individuals will invest their time, funding will come along, too. example 23 (a) i. sneeze in the middle of the night and ii. your edokko neighbor will demand the next morning that you take better care of yourself; iii. stay home with a fever and iv. she will be over by noon with a bowl of soup. (b) if you develop a fever which results in you staying home, your edokko neighbor will arrive by midday with soup. this underscores lance and kukla’s principle that surface grammar does not systematically predict pragmatic force: here we have three examples of ostensibly imperative constructions which function quite differently according to their context, semantic content and the nature of the micro-practice in which they are embedded. 56 from discursive practice to logic? 4.5 anaphoric dependencies in complex directives examples (20–23) above all match the pattern do x and will(φ), as do most (though not all) examples in the appendix. as argued above, the directive clause in these examples can be interpreted as a prescriptive, and so could also be expressed as you should do x . . . . example 24 you should take a left turn . . . you will see the sign for birdoswald. the auxiliary will indicates a potential state of affairs which is a consequence of that postulated by a modal operator in the first clause, or by the imperative mood of a verb. this resembles the wellknown phenomenon of modal subordination roberts (1989), which kibble (1994) modelled with the aid of a co-indexing mechanism, analogous to that conventionally used to indicate anaphoric relationships between nominal expressions: example 25 it mightα rain. you wouldα get wet. in the cited work, the superand subscripts partition the information state by picking out sets of possible worlds: in this example, the set of worlds where it rains is labelled with α and the second clause is evaluated against this restricted set. in the next section, i will argue that we can use a similar mechanism to partition an information state which comprises a commitment store, capturing the dependency between the prescriptive clause and the conjunct 4.6 a dialogical account of discourse structure the effect of a complex directive is to utter a command, request or suggestion and immediately give some reason for executing it, as if anticipating a query or objection from the addressee. thus they can be seen to encapsulate brandom’s sellarsian notion of discourse as a game of giving and asking for reasons. this is in the spirit of kibble (2006; 2007), which proposed that discourse structure could be analysed as the outcome of an “inner dialogue”, where the speaker pre-empts or anticipates such demands for reasons. for instance, a dialogue like (26) from kibble (2007) could be collapsed into a monologue (27) if the speaker anticipates their interlocutor’s contributions at (b,d). the resulting monologue (27) has been marked up with rst relations (mann and thompson, 1987; taboada and mann, 2006): example 26 (a) a: you should take an umbrella. (b) b: why? (c) a: it’s going to rain. (d) b: it doesn’t look like rain. it’s sunny. (e) a: i heard it on the bbc. 57 kibble example 27 motivate nucleus you should take an umbrella satellite evidence nucleus concession nucleus it’s going to rain satellite even though it looks sunny satellite i heard it on the bbc we can construct a similar hypothetical dialogue which might underlie a complex directive such as (20), repeated with rst annotations as (29): example 28 (a) a: take a left turn here on the b6318. (b) b: why? (c) a: . . . after 6 km (4 miles) you will see the sign for birdoswald. example 29 motivate nucleus take a left turn . . . satellite . . . you will see the sign for birdoswald a plausible conjecture is that the pseudo-imperative form may have its origin in genuine (complex) directives, which sought to forestall the hearer’s challenge or clarification request, and were subsequently generalised to conditional uses where the command is not separately entailed by the conjunction. so as noted above, a speaker of (30) is not advising their interlocutor to sneeze: example 30 sneeze in the middle of the night and your edokko neighbor will demand the next morning that you take better care of yourself. . . logical vocabulary makes this conditional sense explicit, and thus avoids commitment to the command on its own. so we can broadly trace this development in terms of the following augmentations to a purely assertional practice: 1. introduce directives which have the effect of adding to the hearer’s attributed commitments; 2. introduce a practice of challenging directives, i.e. demanding reasons; 3. introduce a practice of following directives with some justification, pre-empting challenges; 58 from discursive practice to logic? 4. introduce logical vocabulary which expresses the conditional implied by complex directives; 5. eventually the conditional sense becomes primary, and is read back into the original construction. this will be spelled out in terms of schematic dialogue acts and commitment stores below. note that although we have taken a different route than brandom from a basic discursive practice to a more sophisticated one, the proposed analysis is within the spirit of his lx project, in that it claims to account for how we use more complex locutions to say what we are doing in the lower-level practices. 5. specifying dialogue acts the main aim of this section is to lay some foundations for implementable models of dialogue in which utterances are modelled as updates to an information state (is) in the style of traum et al. (1999), minimally sufficient to model the dialogical genesis of conditional locutions which is argued for in this paper. the framework presented here is adapted from that of kibble (2007). the is will keep track of interlocutors’ deontic statuses, namely the relevant sets of commitments and entitlements acknowledged and attributed by discourse participants. we begin by outlining the structure of the information state and subsequently sketch the update effects of selected dialogue acts. a fully-fledged commitment store would keep track of each dialogue participant’s attributions of everyone’s commitments and entitlements, with self-attributions as acknowledgments of commitments. to keep things manageable within the scope of this paper, we will mostly restrict the discussion to second-person attributions of commitments. 5.1 the information state as a commitment store agents play one of three (dynamically assigned) roles at any given point in a dialogue: speaker, addressee (targeted by utterance), hearer (not directly addressed). 1. for an agent ai to direct a claim φ towards an agent aj is to attribute to aj a particular deontic stance towards φ. 2. the global information state cstore tracks deontic statuses for every participant ai and includes the following types of entries: attr(ai, aj , c(dox(φ)) ai attributes to aj a doxastic (propositional) commitment to φ, a commitment to justify φ if challenged; attr(ai, aj , c(prac(φ)) ai attributes to aj a practical commitment to φ, a commitment to bring it about or stit 18 that φ; attr(ai, aj , e(dox(φ)) ai attributes to aj the entitlement to make a doxastic (propositional) commitment to φ; attr(ai, aj , e(prac(φ)) ai attributes toaj the entitlement to make a practical commitment to φ. 18. modelled on belnap’s (1990) “stit” or ”see to it that” operator. 59 kibble where i = n, the scoreboard includes agent an’s self-attributed commitments or entitlements, the closest thing to private beliefs in this framework. 3. a partition pcstore tracks pending commitments which require acceptance by the targeted agent in order to become binding. 4. in order to capture the distinction between prescriptives and imperatives proper, we introduce a new facility for speakers to directly update hearers’ practical commitments, if they have the appropriate authority: add(ai, aj , c(prac(φ)) ai bestows on aj a practical commitment to φ; attr(ai, aj , e(prac(φ)) ai bestows on aj the entitlement to make a practical commitment to φ. part of the contents of the cstore is marked off as containing provisional or hypothetical commitments. for example, a practical commitment on the part of agent ai to bring it about that φ potentially leads to a doxastic commitment to φ once the action has succeeded. we will refer to these entities as dependent commitments, using a conventional super/subscript notation to indicate the dependency so if rhoda says “joe, you should wash the dishes” the cstore is updated with: {attr(r, j, c(prac(clean(dishes)))α), attr(r, j, c(dox(clean(dishes)))α)} a directive act gives rise to two species of commitment: a primary, practical commitment to an action which is intended to bring about some state of affairs, and a dependent, doxastic commitment to this state of affairs holding once the action had been successfully performed. 5.2 complex directives and conditionals we propose that in the case of complex directives as discussed in section 4.3, conjoined phrases involving an auxiliary verb such as will can also be treated as giving rise to dependent commitments. a justification for this is that, as argued in section 4.5 above, the auxiliary will can have a quasianaphoric function, referring back to a previous operator as in cases of modal subordination such as “joe, you shouldα wash the dishes and i willα dry”. in this example, rhoda self-attributes a practical commitment to drying the dishes in the eventuality that joe washes them. consider again example (19), slightly simplified: example 31 “walk through the gateway and you will find yourself in the largest temple courtyard in the country.” the cstore is updated with: {attr(sp, h, c(prac(through(h, gateway)))α), attr(sp, h, c(dox(through(h, gateway)))βα), attr(sp, h, c(dox(in(h, courtyard)))β) } we now have two consecutive instances of dependent commitments: a practical commitment to walking through the doorway gives rise to a (future) doxastic commitment of being the other side of the doorway; and this doxastic commitment triggers a commitment to the claim that one is in 60 from discursive practice to logic? the temple courtyard. now, the validity of this dependency clearly does not depend on whether one actually walks through the doorway. introducing a conditional locution allows us to express the dependency without requiring the initial practical commitment: example 32 “if you walk through the gate, you will be in the largest temple courtyard in the country.” note that this formulation is considered by the subject to be entailed by (31), and does not convey any more information than was present in that example: in fact it conveys less, as it does not impute any concrete commitments to the addressee. how to handle this in terms of updating the cstore? one approach is to “offer” rather than attributing a practical commitment, which the addressee may choose to accept or not: this enables the speaker to indicate dependent commitments without assuming or requiring any initial commitment from the hearer. this is modelled by adding the relevant attributions to the pcstore rather than the primary cstore: {attr(sp, h, c(prac(through(h, gateway)))α), attr(sp, h, c(dox(through(h, gateway)))βα), attr(sp, h, c(dox(in(h, courtyard)))β) } this approach is suggested by habermas’s notion of a dialogue act offer (sprechaktangebot, habermas (2000); kreutel and matheson (2002)) though our use of the term is somewhat different from his. the point here is that the hearer is not considered to be committed to the action in question, but it can be read off from the cstore that if they do so, the specified consequence will follow. at this point we have outlined how a complex directive within a cooperative goal-oriented dialogue can be interpreted and expressed as entailing a conditional statement, with some modification to the assumed underlying cstore updates. suppose we now re-analyse the original directive as involving an offer rather than an attribution of a practical commitment, in line with brandom’s observation: once the logical vocabulary has been introduced, it may induce practitioners to alter their prior practice, in the light of what it now allows them to say about that practice. (brandom, 2009a, p. 354) the effect is that the initial practical commitment is now hypothesised rather than attributed, yielding a conditional reading of the directive. this seems to apply more naturally to what i call the “generic” and “counterfactual” variants of complex directives represented by (33) and (34) respectively, which are primarily understood as conditional statements with no commitment required to the actions indicated in the directive clauses. example 33 (a) get individuals to invest their time and the funding will follow. (b) if individuals will invest their time, funding will come along, too. example 34 (a) sneeze in the middle of the night and your edokko neighbor will demand the next morning that you take better care of yourself; stay home with a fever and she will be over by noon with a bowl of soup. 61 kibble (b) if you develop a fever which results in you staying home, your edokko neighbor will arrive by midday with soup. 5.3 dialogue acts and updates in what follows sp = speaker, ad = addressee, hs = hearers assert(sp, φ, ad, hs) publicly undertake commitment to justify a propositional claim φ. updates add to cstore: attr(sp, sp, c(dox(φ)) sp self-attributes a doxastic (propositional) commitment to φ, a commitment to justify φ if challenged; attr(sp, sp,e(dox(φ)) sp self-attributes the entitlement to make a doxastic (propositional) commitment to φ; attr(sp,h,c(dox(φ)) sp attributes to h a doxastic (propositional) commitment to φ, attr(sp,h,e(dox(φ)) sp attributes to h the entitlement to make a doxastic (propositional) commitment to φ; prescribe(sp, φ, ad, hs) publicly attribute to ad commitment to bring it about that φ. updates add to cstore: attr(sp,h,c(prac(φ))i sp attributes to h a practical commitment to φ, attr(sp,h,c(dox(φ))i sp attributes to h a dependent doxastic commitment to φ, attr(sp,h,e(prac(φ)) sp attributes to h the entitlement to make a practical commitment to φ; command(sp, φ, ad, hs) s publicly attributes to and bestows on ad commitment to bring it about that φ. updates add to cstore: attr(sp,h,c(prac(φ))i sp attributes to h a practical commitment to φ, attr(sp,h,c(dox(φ))i sp attributes to h a dependent doxastic commitment to φ, attr(sp,h,e(prac(φ)) sp attributes to h the entitlement to make a practical commitment to φ; add(sp,h,c(prac(φ)) sp bestows on h a practical commitment to φ, add(sp,h,e(prac(φ)) sp bestows on h the entitlement to make a practical commitment to φ; propose(sp, φ, ad, hs) publicly offer to ad commitment to bring it about that φ. updates add to pcstore (pending commitments): attr(sp,h,c(prac(φ)) sp attributes to h a practical commitment to φ, attr(sp,h,c(dox(φ))i sp attributes to h a dependent doxastic commitment to φ, 62 from discursive practice to logic? attr(sp,h,e(prac(φ)) sp attributes to h the entitlement to make a practical commitment to φ; complex acts and dependent commitments act1(sp, φ, ad, hs)i ∧ act2(sp, φ, ad, hs)i if the connective ∧ is interpreted as a dynamic conjunction: commitment store updates arising from act1 are stored in a partition labelled i; commitment store updates arising from act2 only take effect if those in partition i are discharged. 6. what kind of practices are “discursive practices”? this paper has argued for a treatment of various linguistic phenomena in terms of evolving discursive practices, and in particular has sought to show that the conditional locution could plausibly have developed out of speech acts embedded in goal-directed cooperative dialogue. we have been using the term “discursive practice” as if it were clearly understood, though in fact it is not in widespread use among semanticists and pragmatists in the anglo-american tradition and brandom does not give any detailed account of what is meant by practices (or the studiedly non-committal term “practices-or-abilities” which occurs throughout bsd). to help get a better grip on how linguistic interaction can be characterised in terms of social practice, this section explores similarities and differences between brandom’s approach and that of the french sociologist and anthropologist pierre bourdieu (following on from (kibble, 2014)), noting some shared connections with the later wittgenstein. both brandom and bourdieu are concerned with how practical knowledge can be articulated and become a subject for discussion, as well as with processes of elaborating basic practices into more sophisticated ones. bourdieu’s highly influential work on the logic of practice (bourdieu, 1977, 1991) focusses on a notion of practices as the fundamental level of description of the behaviour of individuals in social contexts. rouse’s (2007a) survey of practice theories places bourdieu and brandom in distinct camps: according to him, bourdieu is one of the theorists who “make central to their discussion of practices those aspects of human activity which they regard as tacit and perhaps inexpressible in language”, while brandom belongs to the party who “treat language itself (or ‘discursive practice’) as a paradigmatic application of practice talk”. further, a practice for bourdieu may consist of a frequently repeated performance or sequence of actions, while for brandom it is essentially normative, something which may be done correctly or incorrectly according to communally accepted standards (rouse, 2007b). a close reading of two key texts (bourdieu, 1977; brandom, 2008) suggests that the two scholars have a number of concerns in common: 1. commitments and objective intentions. for brandom, to make an assertion is to take on a commitment to justify that assertion if challenged; commitments are a matter of social, normative status rather than psychological states, with practical commitments taking the place of intentions. thus an agent may be assessed by others as being committed to propositions which are entailed by their overt commitments, whether or not they acknowledge such commitments. levesque (1984) sought to capture this distinction with a “logic of implicit and explicit belief”. there is an echo here of bourdieu, who speaks of agents having “objective 63 kibble intentions” which always outrun “conscious intentions” since actions are “the product of a modus operandi” over which the agent has no “conscious mastery” (bourdieu, 1977, p. 79). 2. communication as challenge and riposte. bourdieu (1977, p. 14) claims that this is “the limit towards which every act of communication tends” and appears to see every social encounter as a potential occasion for expressing dominance or deference, while brandom’s notion of autonomous discursive practice requires a speech act of challenging entitlement to propositional commitments (kibble, 2004, 2006; wanderer, 2009; brandom, 2009b). in the works we are focussing on in this paper brandom tends to abstract away from questions of power imbalance among participants in a dialogue and to assume something like a habermasian ideal speech situation (habermas, 1995), though in other work he shows an alertness towards the “power-laden asymmetric recognitive relations that articulate various modern social practices” (brandom, 2015) . 3. background knowledge in the form of habitus or material inference. for bourdieu, individual practices are both constrained by and contribute to the habitus, defined as “systems of durable, transposable dispositions . . . objectively ‘regulated and regular without being in any way the product of obedience to rules . . . ” (quoted by grenfell (2011)) the reader may be reminded of foucault’s discursive formations (foucault, 1972). for brandom, command of a language involves the practical ability to deploy a particular vocabulary, including the ability to make and endorse inferences such as that from this coat is scarlet to this coat is red, or from it is raining to the streets will be wet. 4. generative schemes and algorithmic elaboration. both authors outline ways in which basic or core practices can be combined to generate new practices appropriate to particular situations. for example, bourdieu discusses how the gravity of a theft and the concomitant severity of punishment are determined among the kabyle of algeria on the basis of “a small number of schemes that are continually applied in all domains of practice” (bourdieu, 1977), while brandom as we have seen argues that certain “primitive practices-or-abilities” can be algorithmically elaborated into more complex ones by procedures equivalent to transducing automata (brandom, 2008). 5. explication. both authors are concerned with the issue of codifying or explicating practices. bourdieu argues that “practice has a logic which is not that of logic” while brandom (brandom, 2008, p.33) claims to offer a “logic of practical abilities”. however, they differ over the extent to which practices can be made fully explicit: bourdieu maintains that clusters of practices cannot always be explicitly codified without distortion or logical contradiction while brandom is generally more optimistic, maintaining that logical vocabulary (such as conditional connectives) serves to make explicit inferential moves which are already implicit in any autonomous discursive practice that lacks this vocabulary (weiss, 2009; brandom, 2009a). 6. both situate themselves in relation to the later wittgenstein, particularly the sections of the philosophical investigations on rule following. bourdieu (1977, p. 29) quotes with approval wittgenstein’s remarks that someone who appears to be acting in a predictable, rule-governed manner might be unable to expound any rule they are deliberately following, and argues that to consider a regularity in behaviour as evidence of a consciously propounded ruling or 64 from discursive practice to logic? unconscious regulating is “to slip from the model of reality to the reality of the model”. this anticipates brandom’s wittgenstein-inspired critique of regulism and regularism, the notions that social norms are to be respectively identified with explicit rules or statistical regularities (brandom, 1994). rouse (2007a) identifies the role of language as a contentious issue in practice theory, and argues both that “to use and respond to words and sentences as semantically significant is to engage in discursive practice” and that discursive and non-discursive practices are ultimately inseparable. bourdieu (1991) develops a notion of “symbolic power”, according to which the meanings of utterances and the efficacy of speech acts derives from the “social power” of speakers (leezenberg, 2013). brandom takes a more nuanced approach, seeking to show how semantic meanings can be grounded in social practices of “normative pragmatics”, without the requirement of any explanatory role for semantic or intentional concepts. as we have seen, his approach involves a somewhat rarefied, abstract and irreducably normative account of what constitutes a “practice”: in particular he tends to disregard questions of an imbalance of power among discourse participants, though rouse (1994; 2002) contends that brandom’s framework is quite compatible with foucault’s and other contemporary accounts of interpersonal power. bourdieu problematises the explication of practices on two scores: 1. attempts to collate and set down on paper various collections of practices, as for example in the different ways subjects observe the agrarian calendar, can lead to distortion or incoherence: features which are “compatible practically” may turn out to be “logically contradictory” (bourdieu, 1977, p. 107). that is, invidual subjects may pursue practices that do not interfere with each other, but attempts to codify and harmonise their combined implicit knowledge may show up inconsistencies. this objection can be summed up as “practice has a logic which is not that of logic”. 2. once a practice has been codified and commonly agreed, reflection on the practice may lead agents to go back and revise it: explication is not a one-way street (bourdieu, 1977, p. 20). brandom seems to be more optimistic on point (1): he argues that logical vocabulary must be “semantically transparent” and “inferentially conservative” with respect to material inferences that can be exhibited in the base language. however, he seems to come closer to bourdieu’s stance on point (2), acknowledging that once the logical vocabulary has been introduced, it may induce practitioners to alter their prior practice, in the light of what it now allows them to say about that practice. (brandom, 2009a, p. 354) this does in fact seem quite consonant with bourdieu’s notion of “the dialectic between the schemes immanent in practice and the norms produced by reflection on practices” (bourdieu, 1977, p. 20); both authors appear to be in agreement that explication of practices is not just a one-way process but can feed back into modification of those practices. this may well be what has happened with pseudo-imperatives as discussed in the previous section: the introduction of conditional vocabulary leads to a re-interpretation of complex directives, foregrounding the implicit conditional sense and weakening the actual directive force. this may also be contrasted with lance and kremer’s (1994) objection that the role of logic is purely normative, setting standards for rational argumentative 65 kibble discourse, and should not be identified with an abstraction from actual discursive practice as it is possible for an entire speech community to reason incorrectly. their approach still seems to be faced with the question of what the sources of normativity actually consist in, if not some form of mathematical platonism. as noted at various points above, brandom’s framework appears in places to tacitly rely on unlimited cognitive resources, as in the ability to quantify across all possible contexts or to maintain conservativity of inference following the introduction of logical operators. this seems at odds with a practice-theoretic orientation, which is rather concerned with finite, bounded performances. in fact in bsd brandom appears to move towards a more “cognitivist” position, using the studiedly non-committal term “practices or abilities” throughout and declaring restrictions on inferential operations to be “psychological”: the framework starts to look more like a “competence” theory which abstracts away from the actual capabilities of embodied cognitive agents. 7. conclusion and future work this paper has considered brandom’s expressivist programme in some detail, against the background of bourdieu’s practice theory and in confrontation with empirical data which attest to the actual use of logical connectives in everyday language, and concluded that it remains unproven. we have proposed a re-orientation of brandom’s normative pragmatics which problematises the notion of assertion as an autonomous discursive practice, and outlines a speculative origin story for conditionals as a routinisation of a family of micro-practices. as noted earlier, the goal was to investigate conceptual dependencies between practices such as asserting, commanding and exhibiting conditional reasoning, and we have argued that a rational reconstruction of the genesis of the conditional locution can be achieved if asserting and commanding are given, but not on the basis of assertive practices alone. that is not necessarily to say that these practices can be observed as chronologically prior to conditional reasoning either in language evolution or acquisition, though it suggests interesting possibilities for applied research. we have made a case that even “logical” vocabulary items do not have invariant meanings or inferential force independent of context, but function differently according to the particular practice within which they are employed; this paper has discussed assertive, imperative and prescriptive practices, as well as cooperative dialogue involving conditional reasoning. an implication of this thesis is that one cannot speak of arguments expressed in natural language as being formally valid, except perhaps in a regimented controlled language of the kind one finds in logic textbooks. this need not be such an alarming conclusion if we consider that much everyday inference lacks formal validity: examples include affirming the consequent in the guise of abduction, and argument from authority which is surely unproblematic in many everyday situations. clearly much work remains to be done in fleshing out the details of this proposal, and extending this style of analysis to the behaviour of other so-called logical connectives which poses problems for a classic truth-functional analysis. we have also shown how the multinli corpus provides a valuable resource for investigating “natural inference” as practised in everyday language, which may help to clarify the role of “material inference” in a practice-theoretic framework. brandom’s framework turns out to provide a fruitful basis for investigating discursive practices in which speech and action are closely interleaved. there has been relatively little work in computational modelling of practice theories: narasimhan et al. (2017) report on a “conceptualisation” of social practice theory using agent-based modelling, but note that this requires addressing some key questions which are not resolved in the literature. future work will aim to build on this 66 from discursive practice to logic? conceptualisation to model particular characteristics of discursive practice, and will engage with related work in the philosophy of language (e.g. williamson, 1996) and other areas of study which bear on some of the problems addressed in this paper, including conversational analysis (drew, 2005) and linear logic (porello et al., 2009), a final point: brandom’s inferentialism was worked out in confrontation with the then dominant frameworks of formal semantics, and is largely intended to tackle the kind of problems formal semanticists thought were important. the important new paradigm of distributional semantics (ds, boleda and herbelot (2016)), like brandom’s inferentialism, has also been claimed to be inspired by wittgenstein, though philosophically inclined commentators have generally disregarded this work as it is somewhat outside their purview. it remains to be seen whether ds will fatally undermine inferentialism or, on the other hand, may prove to be complementary. for example, kremer’s accusation of circularity in the definition of syntactic categories might be sidestepped if we begin by using machine learning to induce categories from corpus data. it is ironic that both ds and inferentialism claim inspiration from wittgenstein given brandom’s skepticism about quantitative methods: his critique of “regularism”, which essentially identifies norms with statistical regularities, is that it fails to distinguish between what is usually done and what ought to be done19. acknowledgements thanks to ken turner and the dialogue and discourse reviewers for detailed comments on previous drafts of this paper, and to paul piwek for stimulating conversations over the years on these and related matters. appendix a. examples from multinli corpus 1. (a) get individuals to invest their time and the funding will follow. (b) if individuals will invest their time, funding will come along, too. 2. (a) do begin again, and prudie predicts 1999 will be your year. (b) if you start over again, prudie insists that the year 1999 will be a lucky one for you. 3. (a) lucinda– write a song about pale fire and he’s yours! (b) if you write a song about pale fire, he’s yours. 4. (a) ignore the proprieties or offend the pride of an edokko and he will let you know about it, in no uncertain terms; respect his sense of values and you make a friend for life. (b) if you offend an edokko, he will let you know that he is upset at you. 5. (a) just stir that in and you’ve got a very colorful side another dish (b) if you want a colorful side dish, then stir that in. 6. (a) subtract the shock value, and what you have here is the salon painting of the 1990s. (b) if you minus the shock value, you have a salon painting from the 1990s. 19. the tension between distributionalism and normativity is discussed at length by lücking et al. (2019) 67 kibble 7. (a) follow the n-332 a little farther to vera, then take the road to garrucha and the coast (the n-340 continues inland until almeraa). (b) if you follow the n-332 a little longer, it will lead you to a road you can take to garrucha and the coast. 8. (a) sneeze in the middle of the night and your edokko neighbor will demand the next morning that you take better care of yourself; stay home with a fever and she will be over by noon with a bowl of soup. (b) if you develop a fever which results in you staying home, your edokko neighbor will arrive by midday with soup. 9. (a) yeah scrape it and paint it put primer on it then paint it don’t put paint over else it’ll just continue rusting under it yeah. (b) if you paint over it without scraping and putting primer, it will continue rusting underneath. 10. (a) “give them something to watch and remember and they will forget what you dont́ want them to see,” said jon. (b) if you offer them an alternative, they will not remember anything else. 11. (a) take a left turn here on the b6318 and after 6 km (4 miles) you will see the sign for birdoswald. (b) if you travel for 6 kilometers after taking a left turn you’ll see the sign for birdoswald. 12. (a) as for people you deal with regularly (like doormen, since you’re a manhattanite), grease their palms once every several encounters, or else you’ll go crazy and broke. (b) if you don’t grease the palms of people you deal with, you’ll go crazy and broke. 13. (a) walk through the gateway and you’ll find yourself in an immense open space the largest temple courtyard in the country. (b) if you walk through the gate, you will be in the largest temple courtyard in the country. 14. (a) as i said, get me to sydney, get me to the opening ceremony and the torch and the hymns, and i’ll be fine. (b) if i’m at the opening ceremony in sydney, i’ll be fine. 15. (a) “get the properties and you can go right ahead!” dr. hall found his voice. (b) if the person dr. hall is speaking to gets what they need they can go right ahead. 16. (a) click nobel prize internet archive hayek page, and you’ll find yourself . . . (b) if you click the nobel prize internet archive . . . 17. (a) respond in kind and you’ll soon feel at home. (b) if you treat them similarly you’ll feel right at home in no time. 18. (a) pay us for that with your service, and that new life will be truly precious. 68 from discursive practice to logic? (b) if they help the others with their work, they will be much better in their life. 19. (a) someone once said, “get a job you love, and you’ll never work a day in your life”, zucker said. (b) if you like what you do, you’ll never think of it as a job. references kingsley amis. the king’s english: a guide to modern usage. macmillan, 1999. laurie bauer and elizaveta tarasova. the meaning link in nominal compounds. skase journal of theoretical linguistics, 10(3), 2013. nuel belnap. declaratives are not enough. philosophical studies: an international journal for philosophy in the analytic tradition, 59(1):1–30, 1990. emily bender and alexander koller. climbing towards nlu: on meaning, form, and understanding in the age of data. in proceedings of the 58th annual conference of the association for computational linguistics. association for computational linguistics, 2020. gemma boleda and aurélie herbelot. formal distributional semantics: introduction to the special issue. computational linguistics, 24:619–635, 2016. pierre bourdieu. outline of a theory of practice. cambridge university press, cambridge, 1977. translated by r. nice. pierre bourdieu. language and symbolic power. polity press, cambridge, 1991. translated by g. raymond and m. adamson, edited john b. thompson. samuel r. bowman, gabor angeli, christopher potts, and christopher d. manning. a large annotated corpus for learning natural language inference. in proceedings of the 2015 conference on empirical methods in natural language processing (emnlp). association for computational linguistics, 2015. robert brandom. asserting. noûs, pages 637–650, 1983. robert brandom. making it explicit: reasoning, representing, and discursive commitment. harvard university press, cambridge, ma, 1994. robert brandom. articulating reasons: an introduction to inferentialism. harvard university press, cambridge, ma, 2000. robert brandom. between saying and doing: towards an analytic pragmatism. oxford university press, oxford, 2008. robert brandom. reply to weiss. in b. weiss and j. wanderer, editors, reading brandom: on making it explicit, pages 353–356. routledge, london and new york, 2009a. robert brandom. reply to wanderer. in b. weiss and j. wanderer, editors, reading brandom: on making it explicit, page 315. routledge, london and new york, 2009b. 69 kibble robert brandom. towards reconciling two heroes: habermas and hegel. argumenta, page 29, 2015. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in current and new directions in discourse and dialogue, pages 85–112. springer, 2003. rudolf carnap. meaning postulates. philosophical studies, 3(5):65–73, 1952. stephen clark. vector space models of lexical meaning. in handbook of contemporary semantic theory, second edition, pages 493–522. wiley, 2015. paul drew. conversation analysis. in handbook of language and social interaction, pages 71–102. 2005. michel foucault. the archaeology of knowledge routledge. routledge, 1972. translated by xxx. chris fox. the semantics of imperatives. handbook of contemporary semantic theory, second edition, pages 314–342, 2015. jonathan ginzburg. questions. oxford bibliographies in linguistics, 2017. doi: 10.1093/obo/ 9780199772810-0206. url https://www.oxfordbibliographies.com. michael grenfell. bourdieu: a theory of practice. in michael grenfell, editor, bourdieu, language and linguistics, pages 7–34. continuum, london and new york, 2011. paul grice. meaning. philosophical review, 67:377–388, 1957. reprinted in semantics, edited by d. d. steinberg & l. a. jakobovits (1971), cambridge university press, pages 53-59. jürgen habermas. wahrheitstheorien (1972). in vorstudien und ergänzungen zur theorie des kommunikativen handelns. suhrkamp, 1995. jürgen habermas. comments on john searle: meaning, communication and representation. in maeve cooke, editor, on the pragmatics of communication. mit press, 2000. charles hamblin. fallacies. methuen, london, 1970. ray jennings. the meaning of connectives. in steven davis and brendan s gillon, editors, semantics: a reader. oxford university press, 2004. hans kamp. two theories about adjectives. in steven davis and brendan s gillon, editors, semantics: a reader. oxford university press, 2004. originally published in ed keenan (ed.), formal semantics of natural language, cup, 1975. rodger kibble. dynamics of epistemic modality and anaphora. in proceedings of the international workshop on computational semantics, 1994. rodger kibble. elements of a social semantics for argumentative dialogue. in proceedings of the fourth workshop on computational modelling of natural argumentation, pages 25–28, 2004. rodger kibble. reasoning about propositional commitments in dialogue. research on language & computation, 4(2):179–202, 2006. 70 from discursive practice to logic? rodger kibble. generating coherence relations via internal argumentation. journal of logic, language and information, 16(4):387–402, 2007. rodger kibble. discourse as practice: from bourdieu to brandom. in proceedings of the 50th aisb convention, goldsmiths, university of london, 2014. michael kremer. representation or inference: must we choose? should we? in reading brandom: on making it explicit. routledge, 2009. jörn kreutel and colin matheson. from dialogue acts to dialogue act offers: building discourse structure as an argumentative process. in proceedings of edilog, page 86, 2002. rebecca kukla and mark lance. “yo!” and “lo!”: the pragmatic topography of the space of reasons. harvard university press, 2009. rebecca kukla and mark lance. perception, language, and the first person. in reading brandom, pages 125–138. routledge, 2010. george lakoff. linguistics and natural logic. synthese, 22(1-2):151–271, 1970. mark lance. the logical structure of linguistic commitment iii: brandomian scorekeeping and incompatibility. journal of philosophical logic, 30:439 – 64, 2001. mark lance and philip kremer. the logical structure of linguistic commitment i: four systems of non-relevant commitment entailment. journal of philosophical logic, 23:369 – 400, 1994. mark lance and philip kremer. the logical structure of linguistic commitment ii: systems of relevant commitment entailment. journal of philosophical logic, 25:425 – 49, 1996. mark lance and rebecca kukla. leave the gun; take the cannoli! the pragmatic topography of second-person calls. ethics, 123(3):456–478, 2013. michiel leezenberg. power in speech actions. in arina. sbisa and ken turner, editors, pragmatics of speech actions, pages 287–213. de gruyter mouton, berlin/boston, 2013. hector levesque. a logic of implicit and explicit belief. in aaai, pages 198–202, 1984. judith levi. the syntax and semantics of complex nominals. academic press, 1978. ronald loeffler. brandom. polity, 2018. andy lücking, robin cooper, staffan larsson, and jonathan ginzburg. distribution is not enough: going firther. in proceedings of the sixth workshop on natural language and computer science, pages 1–10, 2019. john macfarlane. brandom’s demarcation of logic. philosophical topics, 36(2):55–62, 2008. william mann and sandra thompson. rhetorical structure theory: a theory of text organization. technical report isi/rs-87-190, information sciences institute, 1987. colin matheson, massimo poesio, and david traum. modelling grounding and discourse obligations using update rules. in proceedings of naacl 2000, pages 1 – 8, 2000. 71 kibble kavin narasimhan, thomas roberts, maria xenitidou, and nigel gilbert. using abm to clarify and refine social practice theory. in advances in social simulation 2015, pages 307–319. springer, 2017. paul piwek. dialogue structure and logical expressivism. synthese, 183(1):33–58, 2011. paul piwek. towards a computational account of inferentialist meaning. in proceedings of aisb 2014, 2014. daniele porello et al. logic and pragmatics: linear logic for inferential practice. tap-2009 towards an analytic pragmatism, page 69, 2009. roland posner. semantics and pragmatics of sentence connectives in natural language. in steven davis and brendan s gillon, editors, semantics: a reader. oxford university press, 2004. originally published in speech act theory and pragmatics, springer, 1980. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in nicoletta calzolari (conference chair), khalid choukri, bente maegaard, joseph mariani, jan odijk, stelios piperidis, and daniel tapias, editors, proceedings of the sixth international conference on language resources and evaluation (lrec’08), marrakech, morocco, may 2008. european language resources association (elra). craige roberts. modal subordination and pronominal anaphora in discourse. linguistics and philosophy, 12(6):683–721, 1989. joseph rouse. power/knowledge. the cambridge companion to foucault, 2, 1994. joseph rouse. how scientific practices matter: reclaiming philosophical naturalism. university of chicago press, 2002. joseph rouse. practice theory. division i faculty publications. paper 43, 2007a. http://wesscholar.wesleyan.edu/div1facpubs/43. joseph rouse. social practices and normativity. division i faculty publications. paper 44, 2007b. http://wesscholar.wesleyan.edu/div1facpubs/44. merel cleo johanna scholman and vera demberg. examples and specifications that prove a point: identifying elaborative and argumentative discourse relations. dialogue & discourse, 8(2):56– 83, 2017. john searle. speech acts: an essay in the philosophy of language. cambridge university press, cambridge, london, 1969. munindar singh. a social semantics for agent communication languages. in issues in agent communication, pages 31–45, 2000. maite taboada and william mann. rhetorical structure theory: looking back and moving ahead. discourse studies, 8(3), 2006. 72 from discursive practice to logic? michael tomasello and luigia camaioni. a comparison of the gestural communication of apes and human infants. human development, 40(1):7–24, 1997. david traum, johan bos, robin cooper, staffan larsson, ian lewin, colin matheson, and massimo poesio. a model of dialogue moves and information state revision. technical report, gothenburg university, department of linguistics; university of edinburgh, centre for cognitive science and language technology group, human communication research centre; universität des saarlandes, department of computational linguistics; sri cambridge; xerox research centre europe, 1999. giacomo turbanti. robert brandom’s normative inferentialism, volume 280. john benjamins publishing company, 2017. stephen turner. practices as the new fundamental social formation in the knowledge society. druzboslovne razprave, xxiv(59):49–64, 2008. johan van benthem. a brief history of natural logic. 2007. illc, universiteit van amsterdam. douglas walton and eric krabbe. commitment in dialogue: basic concepts of interpersonal reasoning. suny series in logic and language. state university of new york press, 1995. jeremy wanderer. robert brandom. philosophy now. acumen, chesham, 2008. jeremy wanderer. brandom’s challenges. in bernhard weiss and jeremy wanderer, editors, reading brandom: on making it explicit, pages 96–114. routledge, london and new york, 2009. bernhard weiss. what is logic? in bernhard weiss and jeremy wanderer, editors, reading brandom: on making it explicit, pages 247–261. routledge, london and new york, 2009. adina williams, nikita nangia, and samuel r bowman. a broad-coverage challenge corpus for sentence understanding through inference. arxiv preprint, 2017. url https://arxiv.org/ abs/1704.05426. timothy williamson. knowing and asserting. the philosophical review, 105(4):489–523, 1996. andrzej wiśniewski. questions, inferences, and scenarios. college publications, 2013. andrzej wiśniewski. semantics of questions. in the handbook of contemporary semantic theory, 2nd edition, pages 273–313. wiley, 2015. ludwig wittgenstein. philosophical investigations, 4th edition. john wiley & sons, 2009. 73 kibblefdp2l_aug2020-p1 kibblefdp2l_aug2020pp35-40 journal of machine learning research-microsoft word template dialogue and discourse 4(2) (2013) 65-86 doi: 10.5087/dad.2013.204 ©2013 bruno cartoni, sandrine zufferey, thomas meyer submitted 02/12; accepted 11/12; published online 04/13 annotating the meaning of discourse connectives by looking at their translation: the translation spotting technique bruno cartoni bruno.cartoni@unige.ch département de linguistique université de genève rue de candolle 2 ch-1211 genève 4 sandrine zufferey s.i.zufferey@uu.nl utrecht institute of linguistics trans 10 nl-3512 jk utrecht thomas meyer thomas.meyer@idiap.ch idiap research institute centre du parc rue marconi 19 ch-1920 martigny editors: stefanie dipper, heike zinsmeister, bonnie webber abstract the various meanings of discourse connectives like while and however are difficult to identify and annotate, even for trained human annotators. this problem is all the more important since connectives are salient textual markers of cohesion and need to be correctly interpreted for many natural language processing applications. in this paper, we suggest an alternative route to reach a reliable annotation of connectives, by making use of the information provided by their translation in large parallel corpora. this method thus replaces the difficult explicit reasoning involved in traditional sense annotation by an empirical clustering of the senses emerging from the translations. we argue that this method has the advantage of providing more reliable reference data than traditional sense annotation. keywords: discourse relations, connectives, annotation methods, parallel corpora, translation 1 introduction many natural language processing (nlp) tools rely on annotated data, that is linguistic data enriched with meta-information. for most part, this information requires manual annotation, often performed by more than one human annotator, in order to ensure optimal reliability. this paper reports a set of experiments performed for the annotation of discourse connectives in the context of a project that aims at improving machine translation systems. one of the main problems for current machine translation systems comes from lexical items that cannot be resolved by looking at individual sentences, such as pronouns, discourse connectives and verbal tenses. the goal of the swiss comtis project 1 is to 1 http://www.idiap.ch/comtis cartoni, zufferey & meyer 66 extend the current statistical machine translation paradigm by modeling these intersentential relations (popescu-belis et al. 2011; 2012). this project addresses several types of cohesion markers, but the experiments reported in this paper are limited to discourse connectives. we particularly focus on the challenging task of annotating the meaning of connectives, and advocate the use of a method called translation spotting. this method is based on the collection of a large amount of translations of connectives in a target language in order to capture the different meanings of a given connective in the source language. the paper is organized as follows. first, we briefly define the category of discourse connectives, emphasizing their importance for textual coherence and discussing the challenges they raise for machine translation (section 2). we go on to compare in section 3 two techniques used in the literature to annotate the meaning of connectives, namely sense annotation (3.1) and translation spotting (3.2) and discuss their potential advantages and limitations. in section 4, we sequentially test these methods through a series of annotation experiments, with the conclusion that translation spotting adds improvements with respect to sense annotation. we go on to show in section 5 that translation spotting can also be used to identify fine-grained differences between connectives conveying the same meaning (i.e., a causal relation). section 6 discusses the advantages and limitations of the translation spotting method and section 7 summarizes our conclusions. 2 discourse connectives: a challenge for machine translation discourse connectives, such as the words because and while in english or parce que and mais in french form a functional category of lexical items that are very frequently used to mark coherence relations such as explanation or contrast between units of text or discourse (e.g. halliday & hassan 1976; mann & thomson 1992; knott & dale 1994; sanders, 1997). even though most languages possess such a set of items, they vary tremendously in the number of connectives they have to express relations and in the use they make of them. moreover, a well-known property of discourse connectives is that they are often multifunctional and can convey several coherence relations. in some cases, various relations are conveyed by the same occurrence of a connective. for example, in french, the connective tant que (roughly corresponding to the english as long as) intrinsically conveys both a temporal relation and a conditional meaning in all its occurrences. in other cases, a connective can potentially convey several relations, but a single occurrence conveys only one of these relations. in such cases, a specific occurrence can be ambiguous between several rhetorical relations. to cite a case in point, the english connective since can convey a causal meaning but also a temporal one. in french however, these two meanings require distinct translations: depuis que for the temporal meaning and car or puisque for the causal one. from a machine translation perspective, the main challenge raised by discourse connectives is to be able to assign them a correct meaning in order to translate them appropriately. for example, in order to translate (1) correctly, a system has to recognize that since here has a temporal meaning and not a causal one, and should therefore be translated by depuis que as in (2) and not by the causal connective car as in (3), as was produced by a web-based translation engine. 1. i have been having fun since this conference started. 2. j’ai eu beaucoup de plaisir depuis que la conférence a commencé. 3. *j'ai eu plaisir car cette conférence a commencé. annotating discourse connectives by looking at their translation 67 in order to disambiguate discourse connectives for machine translation (and more specifically for statistical machine translation (smt), the comtis project proposes to pre-process their occurrences and label them with meaning tags, thus enabling the smt system to make the correct choice in the target language. in other words, the training data should contain occurrences of since labeled as either causal or temporal, in order to help the smt system to learn how these two uses of the connective should be translated in different contexts. 2 this labeling of connectives is achieved automatically using machine learning, with algorithms trained on manually annotated reference data (meyer & popescu-belis 2012). afterwards, the same classifier is applied when translating a new sentence. in this approach, the automatic disambiguation of connectives thus requires the manual annotation of a large amount of data. in this paper, we discuss the problems raised by this manual annotation. we present the different techniques that have been applied in the comtis project in order to achieve reliable and tractable results. first, a classical sense annotation approach has been used, which consists in asking human judges to annotate manually a set of data with several possible senses for each connective. the rather low inter-annotator agreement resulting from this annotation led us to investigate another technique based on translation spotting. these two approaches are described in turn in the next sections. 3 state-of-the-art methods for the annotation of connectives this section presents two methods used to annotate discourse connectives: sense annotation (section 3.1) and translation spotting (section 3.2). section 3.3 provides an overview of the resources created using translation spotting. 3.1 sense annotation a classical annotation method for connectives consists in asking several human annotators to assign a label from a list of senses to occurrences of a given connective. usually, such annotations are performed by more than one annotator, and an evaluation step assesses the reliability of the annotation by measuring the inter-annotator agreement. this assessment is needed in order to ensure that the annotation is valid (arstein & poesio 2008). as stated by spooren and degand (2010: 253) “ideally coders work completely independently and agree substantially”. but in many cases, this goal cannot be met. spooren and degand suggest various solutions in order to improve the level of agreement, such as increasing the amount of training for the annotators, or discussing the disagreements between annotators in order to reach a consensus. in a meta-analysis of factors influencing interannotator agreement on three different types of linguistic data, bayerl & paul (2011) found eight factors with a significant impact on agreement scores, among which were the amount of training, the homogeneity of the group of annotators and number of linguistic categories to be annotated. even though this meta-analysis did not include linguistic phenomena related to discourse, these factors confirm that spooren and degand’s suggestions should have a positive impact on inter-annotator agreement. one of the most important resources containing sense annotation for discourse connectives is the penn discourse treebank (pdtb) (prasad et al., 2008). 3 the pdtb provides a discourse-layer annotation over the wall street journal corpus (wsj) containing the same sections as have already been annotated syntactically in the penn 2 the comtis project focuses on french and english, but the methodology developed for the disambiguation of connectives can be extended to other languages. 3 the current version 2.0 is available through the linguistic data consortium at: http://www.ldc.upenn.edu. a website with an extensive bibliography, tools and manuals can be found at: http://www.seas.upenn.edu/~pdtb http://www.ldc.upenn.edu/ http://www.seas.upenn.edu/~pdtb cartoni, zufferey & meyer 68 treebank. the discourse annotation consists of manually annotated senses for about 100 types of explicit connectives, implicit discourse relations and their argument spans. for the total size of the wsj corpus of about 1,000,000 tokens, there are 18,459 annotated instances of explicit connectives and 16,053 instances of annotated implicit discourse relations. the senses that discourse connectives can signal are organized in a hierarchy containing three levels of granularity, with four top level senses (temporal, contingency, comparison and expansion) followed by 16 subtypes on the second level and the 23 detailed sub-senses on the third level. the annotators of the pdtb were allowed to freely choose senses among all levels, including the possibility to annotate double sense labels (from any hierarchy levels) to account for ambiguous cases. this is why, in principle, 129 sense combinations are possible. a similar methodology has been implemented to annotate discourse relations in many other languages such as hindi, czech, arabic and italian (see webber & joshi 2012 for a review). in addition, zufferey et al. (2012) conducted multilingual annotation experiments in five indo-european languages. in all these studies, similar cases of inter-annotator disagreement were reported. these results indicate that the methodology and results from the pdtb can be to a large extent replicated in other languages. among the 100 different explicit connectives found in the pdtb, we calculated that 29 of them were annotated only with one sense for all their occurrences, covering 412 occurrences. these connectives can therefore be treated as non-ambiguous. among the remaining 71 connectives, we counted that 52 connectives were annotated with two labels belonging to different top-level categories in the hierarchy. for example, the connective while was annotated with the label concession (belonging to the comparison class), and with the label synchrony (belonging to the temporal class). we reasoned that connectives like while, with several senses belonging to different top-levels categories, represented an important ambiguity that needed to be resolved for translation purposes. we therefore concentrated our annotation effort on connectives belonging to this category. in the pdtb, problems related to inter-annotator agreement have been resolved by choosing the first common label in the hierarchy above the ones that were annotated. for example, when one annotator had labeled an occurrence of while as expectation, and another annotator had labeled it as contra-expectation (both labels come from the most detailed third level of the hierarchy), this disagreement was resolved by going up to the second level of the hierarchy and choosing the tag concession, covering the two chosen tags. detailed information on the performance of the annotators is given in miltsakaki et al. (2008). the inter-annotator agreement for the four top-level senses in the pdtb is high, at 92%. for the most detailed third level however, performance drops to 77%, showing the difficulty of such a fine-grained annotation. performance on specific discourse connectives is only given for the early stages of the pdtb corpus annotation. for example, in miltsakaki et al. (2005), some information is provided on the annotation of while with its four main senses, that were described at the time of that paper as: temporal, concessive, contrast and comparison. for 100 tokens of while and two annotators, 20 sentences were judged to be uncertain. out of the 80 remaining sentences, there was 84% of agreement and 16% of disagreement. when all 100 sentences are taken into account, the overall agreement reaches only 67%. in short, sense annotation such as the one performed in the pdtb is not always straightforward for the annotators and different annotators do not consistently annotate many fine-grained distinctions. 3.2 translation spotting translation spotting is an annotation method that makes use of the translation of specific lexical items in order to disambiguate them. for example, an occurrence of since translated by puisque in french indicates that this occurrence of since has a causal rather annotating discourse connectives by looking at their translation 69 than a temporal meaning, because the french connective puisque is unambiguous while the english since is not. table 1 presents an excerpt of parallel sentences from europarl containing since in english and the translation spotting, done manually. for one single item in the source language, translation spotting has to be performed over a large set of bilingual sentence pairs, in order to cover many possible correspondences in the target language. english sentence french sentence transpot 1 in this regard the technology feasibility review is necessary, since the emission control devices to meet the ambitious nox limits are still under development. à cet égard, il est nécessaire de mener une étude de faisabilité, étant donné que les dispositifs de contrôle des émissions permettant d'atteindre les limites ambitieuses fixées pour les nox sont toujours en cours de développement. étant donné que 2 will we speak with one voice when we go to events in the future since we now have our single currency about to be born? parlerons-nous d'une seule voix lorsque nous en arriverons aux événements futurs, puisqu'à présent notre monnaie unique est sur le point de voir le jour? puisque 3 in east timor an estimated one-third of the population has died since the indonesian invasion of 1975. au timor oriental, environ un tiers de la population est décédée depuis l'invasion indonésienne de 1975. depuis 4 it is two years since charges were laid. cela fait deux ans que les plaintes ont été déposées. paraphrase table 1: example of translation spotting for since the term translation spotting was originally coined by véronis & langlais (2000) to designate the automatic extraction of a translation equivalent in a parallel corpus. in our experiments however, the spotting was done manually in order to get fully accurate reference data. indeed, some attempts have been made to perform translation spotting automatically (simard, 2003), but they proved to be particularly unreliable when dealing with connectives: danlos and roze (2011) assessed the translation spotting performed by transsearch (huet et al. 2009), a bilingual english-french concordance tool that automatically retrieves the translation equivalent of a query term in target sentences, and found that for the french connectives en effet and alors que, the tool spots an appropriate english translation for 62% and 27.5% of the cases respectively. compared to the general performance of the transsearch tool for the rest of the lexicon (around 70% of accurate transpots), these results are particularly low. danlos & roze (2011) suggest that one possible explanation is the important number of possible translations that can be found for connectives, ranging from no translation to paraphrases and syntactic constructions, which therefore are difficult to spot automatically. the theoretical idea behind translation spotting is that differences in translation can reveal semantic features of the source language (e.g. dyvik, 1998; noël, 2003). in these studies, translation is used to elicit some semantic feature of content words in the source language. yet, behrens & fabricius-hansen (2003) convincingly showed that using translated data can also help to identify the semantic space of the coherence relation of elaboration, conveyed with one single marker in german (indem) but translated in various ways in english (when, as, by + ing, -ing). of course, translated texts do not faithfully cartoni, zufferey & meyer 70 reproduce the use of language in source texts as translation has a number of inherent features (e.g. baker, 1993). translated data can therefore only be used to shed light on the source language, and investigation should be based on the source language side of parallel data only (see section 4.1 for details on our corpus data). when performed manually, translation spotting provides very reliable results and has a number of advantages over sense annotation. first, it relies on the decision made by the translator, who is an expert in his/her own language, and who makes translation choices according to the entire context of use (i.e., knowledge of the whole text) and his/her professional training in the target language. second, the task is easier to explain to human annotators, and disagreements are rather few. by contrast, the disagreements for some sense tags can be really high for some distinctions such as concession and contrast (zufferey et al. 2012). third, the different labels are not set a priori, and the wide variety of translations provides an overview of the possible means to translate a connective. finally, this task gives an interesting view of the number of discrepancies between the two languages, when there are no one-to-one translation equivalences, a very frequent situation for connectives. this last advantage is less important for annotation but has important implications for other nlp tasks relying on aligned data. however, translation spotting also has a number of limitations. the most important one is that it provides a direct disambiguation only when the language of translation is less ambiguous than the source language for a given linguistic item, and only one translation is possible for each meaning of the source language. in addition, even in a large corpus, there is no guarantee that all possible senses of a connective will be covered. another limitation is the necessity to include data from several genres in order to cover a larger range of connective uses, as the functions of connectives are variable across text types (sanders 1997). in the specific context of the comtis project however, the parallel corpus used for translation spotting is the same corpus as the one used to build the language model for machine translation. consequently, ambiguities that are found in the annotation are precisely those that have to be dealt with for machine translation. in order to solve part of these limitations, we suggest adding a second step of analysis to translation spotting. this step consists in grouping items of the target language that share the same meaning. for example, in table 1, the translation spotting of since in sentences 1 and 2 are clustered, because both étant donné que and puisque convey a causal meaning in french, while the two others (depuis and the paraphrase cela fait x que) convey a temporal meaning. but clustering is not always an easy task for all meaning differences. in order to perform it in the most reliable way, we propose an empirical method involving an interchangeability test. this test is performed by asking human judges to decide which connective can be replaced by another one from the list of possible translations. it takes the form of a sentence completion task. this additional step allows for the separation of translations that are equivalent and reflect the same meaning in the source language and translations that are not equivalent (or interchangeable) and reflect two different meanings of the connective in the source language. for example, a translation spotting performed for the english connective although resulted in three main translations in french: pourtant, bien que and même si. however, an interchangeability test performed on a set of french sentences revealed that bien que and même si were interchangeable (provided that the mood of the verb is unmarked), as they both reflect a concessive meaning of although, while pourtant cannot be used in place of the other two connectives, as it reflects the contrastive meaning of although. thus, through this sentence completion task, equivalent translations could be reliably identified and the two meanings of although were reliably coded in the source language. additional examples of such tests are presented in section 4.4. annotating discourse connectives by looking at their translation 71 3.3 resources created in the comtis project in the comtis project, translation spotting was so far performed on seven english connectives, reported in table 2. in this table, a priori meanings correspond to the possible meanings of connectives identified in reference data and a posteriori meanings correspond to the meaning tags assigned after the clustering phase described above. the number of sentences for the resources created through translation spotting is often lower than the number of sentences that were spotted, due to cases of zero translations or ambiguous connectives, for which no specific meaning can be identified. translation spotting was made with english-french parallel sentences. additional spottings of connectives are in progress for other language pairs. connective a priori meanings a posteriori meanings no. of annotated sentences resources created (in sentences) while contrast, concession, comparison, temporal contrast/temporal, concession, contrast, temporal_duration, temporal_punctual, temporal_conditional 499 294 although contrast, concession contrast, concession 197 183 though contrast, concession contrast, concession 200 155 even though contrast, concession contrast, concession 212 191 since causal, temporal causal, temporal, temporal/causal 423 423 yet adverb, concession, contrast adverb, concession, contrast 509 403 meanwhile contrast, temporal contrast, temporal 131 131 total 2171 1780 table 2: resources created in the comtis project through translation spotting 4 experiments comparing sense annotation and translation spotting we have discussed in section 3 two possible methods for assigning a meaning to ambiguous connectives. in this section, we will test them through a series of annotation experiments using a convergent methodology and the same annotators in both cases. these experiments will provide a comparative evaluation of their advantages and limitations. 4.1 data and methodology for our experiments we used the europarl corpus (koehn, 2005), a multilingual corpus made of the minutes of the debates of the european parliament. this corpus contains 23 languages in parallel: each speaker speaks in his/her own language, and every statement is translated into the other official languages. the europarl corpus is a 506-fold parallel corpus (23*23-23), but this does not mean that all parallel data contains an original text and its translation. a statement made in german will be translated both into english and french, and the two resulting texts are therefore two parallel translations. moreover, the two directions of translation cannot be cartoni, zufferey & meyer 72 considered as equivalents. previous studies (degand, 2004; cartoni et al. 2011) revealed that one of the variation factors for the use of connectives is the status of the text, and more specifically whether it is an original text or a translation. consequently, the use of parallel data in the study of discourse connectives requires identifying clearly the source and the target languages. the europarl corpus contains this information in the meta-data structure, but pre-processing steps are required to extract parallel texts, where original and translated languages are clearly identified. these steps are described in cartoni et al. (2011) and cartoni & meyer (2012). the annotators that did the sense annotation experiments described in section 4.2 were native french speakers with high proficiency in english. these annotators have been trained in two steps. first, they received written explanations about the discourse relations that they were going to annotate with examples of these relations. after reading instructions, they were asked to annotate a set of 50 sentences. a first evaluation was performed on this annotation, by computing an inter-annotator agreement score, and by looking more precisely at cases where the annotation diverged. in a second phase, the annotators received additional explanations about the discourse relations, focusing on the cases where disagreements were found. in some cases, a think-aloud protocol was also used (ericsson & simon 1980), by asking each annotator individually to verbalize the reasoning leading to their final decision while they were annotating a couple of sentences. this provided an efficient correction for the annotators in case an incorrect criterion was used and could be identified. 4.2. a sense annotation experiment in english and french the annotation of connective senses has been tested on one english connective (while) and one french connective (alors que) that share the property of conveying a contrastive meaning in part of their occurrences. according to the lexconn database of french connectives (roze et al. 2010), the connective alors que can convey a temporal-background meaning (4) in addition to its contrastive meaning (5). 4. en mai, alors que je me trouvais encore à pau, je suis tombé malade. in may, connective i was still in pau, i got sick. 5. j’aime beaucoup molière, alors que corneille m’ennuie profondément. i like molière very much, connective corneille bores me dreadfully. according to miltsakaki et al. (2005), the english connective while can signal four different senses. 4 first, while can indicate a temporal meaning (temp), referring to a duration in time, i.e., the synchronous overlapping of two events, as in example (6). the second sense is a comparison (comp) with a juxtaposition of two or more alternatives, as in example (7). the third label is concession (conc), where one argument of the sentence is an expectation, which is then violated or negated by the second argument of the sentence, as in example (8). the fourth sense marks a strong contrast (cont), for example between two extremes (antonyms) of a gradable scale, as in example (9). 6. that impressed robert b. pamplin, georgia-pacific's chief executive at the time, whom mr. hahn had met while fundraising for the institute. 7. between 1998 and 1999, loyalists assaulted and shot 123 people, while republicans assaulted and shot 93 people. 4 the pdtb in its current version uses slightly different and up to 21 different senses (combinations) for while. annotating discourse connectives by looking at their translation 73 8. while the pound has attempted to stabilize, currency analysts say it is in critical condition. 9. while georgia-pacific's stock has outperformed the market in the past two years, nekoosa has lagged the market in the same period. for the english and the french connectives, we have asked two human annotators to annotate occurrences with the meaning described above. in french, they annotated 423 sentences containing alors que, extracted from the french part of the europarl corpus. annotators were asked to decide between two labels: “b” for background or “c” for contrast. two additional labels were provided: one that could be used to indicate that the annotator could not decide which meaning the connective conveyed (“u”) and one serving to annotate strings of characters that did not correspond to the connective alors que but to another uses of this string of words, as in (10) from the corpus. such cases were annotated with “d” for discarded. 10. on verrait alors que le fédéralisme européen, qu'on nous propose tout à coup comme la panacée, a constitué, dès ses balbutiements, la cause même du mal que l'on dénonce. we would then see that european federalism, while is all of a sudden being proposed as a cure-all, has from its earliest days been the very cause of the wrong we are condemning. the results of this annotation are reported in the table 3, a contingency table showing the agreements and disagreements between the two annotators . annotator1 a n n o ta to r2 b c d u total b 86 109 0 7 202 (47.8%) c 12 181 0 6 199 (47%) d 0 0 20 0 20 (4.7%) u 0 2 0 0 2 (0.5%) total 98 (23.2%) 292 (69%) 20 (4.7%) 13 (3.1%) 423 (100%) table 3: contingency table for the annotation of alors que the agreement of the two annotators on this task was calculated with cohen kappa’s score (carletta 1986) and reached 0.428. this represents 67.8% of cases of observed agreement. when looking more closely at the results, we noticed that there was no disagreement on the simplest category d (discard) that was correctly annotated in all 20 occurrences, thus confirming that the two annotators were reliable. they almost never used the label “u”, which means that they were rather confident about their choices. moreover, the cases of disagreements between b and c seem to indicate that the two annotators did not adopt the same strategy in case of uncertainty. there were, for example, an important number of cases (109), where the first annotator consistently chose the contrastive meaning, while the second annotator chose the background meaning, but not the other way round (12 cases only). in other words, ambiguous cases were consistently classified with b by one annotator and c by the other. we will argue in section 6 that such occurrences may correspond to natural ambiguities, for which a double label tag should be assigned. in english, 300 sentences containing while were extracted from the english part of europarl and annotated by the same annotators. guidelines taken from the pdtb cartoni, zufferey & meyer 74 annotation manual (the pdtb research group, 2007) were provided to explain the different meanings conveyed by while. annotators had to decide between these four labels, plus one label if they could not decide (“u”). the inter-annotator agreement (cohen’s kappa score) was 0.426, a rather similar value to the one obtained for alors que described above. this corresponds to an agreement for 61.3% of the sentences, a slightly lower value than the 67% obtained by miltsakaki et al. (2005). the contingency table for while is presented in table 4. annotator 1 comp conc cont temp u total a n n o ta to r 2 comp 13 1 2 2 0 18 (6%) conc 15 101 1 21 1 139 (46.3%) cont 8 22 5 8 1 44 (14.7%) temp 9 9 6 64 5 93 (31%) u 0 2 1 2 1 6 (2%) total 45 (15%) 135 (45%) 15 (5%) 97 (32.3%) 8 (2.7%) 300 (100%) table 4: contingency table for the annotation of while the distribution of annotations reported in table 4 is rather unbalanced. annotators seem to reach some agreement for concession and temporal senses but overall the four labels are mixed, and no particular preference is observed for alternative tags. contrary to alors que (see table 3 above), for which one annotator clearly tended to choose a different strategy than the other, no emergence of a consistent strategy is found in this case. the larger range of possible meanings probably caused this important number of divergences. in sum, these annotation experiments highlighted the difficulties of labeling the meanings of discourse connectives, even when only a binary distinction was necessary. in both cases, the inter-annotator agreement remained low, with a kappa score never reaching 0.5. in the domain of computational linguistics, the threshold of acceptable agreement is highly debated (arstein & poesio 2008), but following krippendorff’s scale assessing inter-annotator agreement (carletta 1996: 52), these kappa scores do not indicate reliable coding. following the scale by landis & koch (1977), a value of 0.4 is considered to reflect a moderate agreement. in all cases, this score does not appear to be reliable enough to provide reference data for training automated classifiers, as it is aimed in the comtis project. 4.3. a translation spotting experiment with the connective while as mentioned above, the connective while can convey four major meanings: temporal, concessive, contrastive and comparative. as we have seen with the sense annotation experiment, the distinction between these four meanings is hard to make in a systematic and reliable way for human annotators. we therefore tried to separate these senses in the source language through translation spotting. we used 508 bi-sentences extracted from the europarl corpus for the english-french pair, and we extracted sentences that were originally produced in english. two human annotators (the same annotators who did the sense annotations) were then asked to identify the connective that was used in the target french text in order to translate while. if it was not translated by a french connective, they were allowed to assign different tags for the use of a present participle, a paraphrase, or no translation at all. the table below provides details about the different means used to translate while in french. annotating discourse connectives by looking at their translation 75 no. % no. % alors que 91 18.24% mais 4 0.80% gerund 85 17.03% malgré 3 0.60% paraphrases 72 14.43% quoique 3 0.60% si 54 10.82% pendant que 2 0.40% zero translation 41 8.22% alors même que 1 0.20% tandis que 39 7.82% aussi 1 0.20% même si 33 6.61% avant que 1 0.20% bien que 26 5.21% contre 1 0.20% s'il est vrai que 14 2.81% en même temps que 1 0.20% tant que 10 2.00% étant donné que 1 0.20% pendant 5 1.00% quand 1 0.20% puisque 5 1.00% s'il est exact que 1 0.20% lorsque 4 0.80% total 499 100% table 5: translation equivalents of while found in the corpus although the task might seem trivial, the two annotators provided a different translation spotting for 150 sentences out of the 508. 5 most of these cases were due to a disagreement about what counted as a paraphrase. for example, one annotator treated the string of words s’il est vrai que as a paraphrase and the other as a connective. this disagreement is easily correctible, and further training has consistently increased the level of agreement. in subsequent tasks, the annotators agreed in 91.5% of the cases when transpotting other connectives like whereas, and in 93% of the cases for although. 4.4. interchangeability tests as a second step for translation spotting as can be seen in table %, a wide range of french connectives is used to translate while, reflecting the numerous meanings that this connective can convey. in order to deduce its meanings based on the translations, an additional task of clustering is needed, which involves analyzing the french connectives used in the translations. in order to do so, we performed an interchangeability test on french connectives, taking the form of a sentence completion task. such a task consists of taking a bunch of sentences from our parallel data containing a specific connective (the connective used in the translation), erase it and ask human annotators to decide, from a list of connectives, which one would fit, without paying attention to the verb mood, which may be influenced by the connective. this kind of test allows making a decision with no theoretical a priori. the only a priori decision that we made was to separate the translations from table 5 into two sub-groups: the temporal connectives on one side and all the others on the other side. among the 6 most frequent french connectives used to translate while (alors que, si, tandis que, même si, bien que, s'il est vrai que), we proposed a set of sentences with blanks to fill in to three annotators. for each of the sentences (numbered 1 to 24), table 6 provides the connectives that were used in the text, followed by the connectives chosen by the annotators (the numbers in brackets correspond to the number of times the connectives have been chosen). only connectives that were chosen several times are reported. 5 among the 508 occurrences of while, 499 were connectives. the other occurrences were nouns as in “for a while” or “a while ago”, and have been excluded from the count. cartoni, zufferey & meyer 76 sentence connective used in translation chosen connectives (number of times / 3 annotators) 1 alors que alors que (3), si (3), s’il est vrai que (3), tandis que (2) 2 alors que alors que (3) 3 alors que alors que (3), tandis que (3) 4 alors que même si (3), bien que (2) 5 bien que bien que (3), même si (2) 6 bien que bien que (3), même si (2), s’il est vrai que (2) 7 bien que bien que (3), même si (2) 8 bien que bien que (2), même si (3), si(2), s’il est vrai que (2) 9 même si même si (3), bien que (3), si(2), s’il est vrai que (2) 10 même si même si (3), bien que (3), s’il est vrai que (3), si(2) 11 même si même si (3), bien que (2) 12 même si même si (3), bien que (3) 13 si s’il est vrai que (3), même si (3), si(2), bien que (2) 14 si s’il est vrai que (3), si(3), même si (3), bien que (2) 15 si s’il est vrai que (3), si(3), même si (2) 16 si s’il est vrai que (3), même si (3), si(2), bien que (2) 17 s'il est vrai que s’il est vrai que (3), même si (3), si(2), bien que (2) 18 s'il est vrai que s’il est vrai que (3), même si (2), bien que (2) 19 s'il est vrai que s’il est vrai que (3), même si (3), bien que (3) 20 s'il est vrai que s’il est vrai que (3), même si (3), si(2), bien que (2) 21 tandis que alors que (3), tandis que (2), si (3) 22 tandis que alors que (3), tandis que (3) 23 tandis que alors que (3), tandis que (3) 24 tandis que alors que (3), tandis que (3) table 6: interchangeability test for non-temporal uses of while through this test, two clusters of connectives are clearly emerging: one with a concessive meaning containing même si, bien que, si and s’il est vrai que, and another one with a contrastive meaning containing alors que and tandis que. however, this also shows that alors que can also have a concessive meaning, as in sentence 4, where it’s been interchanged in majority with même si and bien que. within these two clusters, there seems to be some more subtle clusters between même si et bien que on one side, and si and s'il est vrai que on the other side. this is confirmed in the descriptive reference work lexconn (roze et al. 2010) that assigns the connective si both a concessive and a condition meaning. this latter meaning was never annotated in the english reference for while (the pdtb), but will also emerge from the interchangebility test described below. finally, the meaning of comparison was not found in this test. it also shows that the connectives used in the translation were always the first choice of the annotators as well, with the noticeable exception of tandis que that the annotators seem to avoid using. the same test was also performed for the french connectives conveying a temporal meaning pendant que, tant que, lorsque. results are reported in table 7. annotating discourse connectives by looking at their translation 77 sentence connective used in the translation chosen connectives (number of times / 3 annotators) 1 lorsque lorsque (3) 2 lorsque lorsque (3) 3 lorsque lorsque (3), pendant que (2) 4 lorsque pendant que (3) 5 pendant que pendant que (3) 6 pendant que pendant que (3) 7 tant que tant que (3) 8 tant que tant que (3) 9 tant que tant que (3) 10 tant que tant que (3) table 7: interchangeability test for temporal uses of while this test, contrary to the one above for concessive/contrastive meanings, shows no cluster with more than one connective. apart from a few exceptions, it seems to show that there are three connectives with a specific meaning that cannot be expressed by another connective. for example, the connective tant que, that can roughly be translated into english by as long as, indicates duration in time as well as condition: the duration lasts only while the event mentioned in the segment following the connective unfolds. the connective pendant que conveys both a notion of contrast and simultaneity with another event. this connective indicates that a contrastive and temporal meaning can coexist in some connectives, with the consequence that some uses of while could be tagged as both temporal and contrastive. finally, lorsque only indicates temporal simultaneity. the interchangeability tests allow the clustering of french connectives that convey the same meaning, and consequently narrow the different possible meanings of english while. the translation spotting and interchangeability tests also revealed that there were more fine-grained features to the temporal uses of while (simultaneity, condition, etc.). these specificities of while with a temporal meaning are more specific than the labels used in the pdtb, where the temporal category is only sub-divided into synchronous and asynchronous. in this particular case, the translation reveals fine-grained distinctions of meaning in the source language, as it was the case in studies focusing on content words, mentioned in section 3.2. table 8 summarizes the different meanings that have been highlighted by clustering french connectives. only french connectives that were used more than once have been included in the analysis. meaning % french connectives concession 25.45 si (54), même si (33), bien que (26), s'il est vrai que (14) contrast 7.89 tandis que (39) contrast/temporal 18.24 alors que (91) temporal/condition 2 tant que (10) temporal/comparison 1.4 pendant que (7) temporal/simultaneity 0.8 lorsque (4) table 8: meanings of while emerging from translation spotting these meanings are then reported on the corresponding occurrence of english while, that receives the labels inferred from the translation. this annotated data (294 occurrences of while in total) is then used to train classifiers based on machine learning algorithms, in order to automatize the annotation procedure (meyer and popescu-belis, 2012). from the 294 instances, 14 are kept as a held-out test, while the other 280 are used for training a cartoni, zufferey & meyer 78 maximum entropy classifier, using the stanford nlp package (manning and klein, 2003). in both, the training and the test sets, features from syntactical parsing (charniak and johnson, 2005) are extracted: pos tags and syntactical ancestor categories for the connective, the surrounding words and words at the beginning and end of the clauses. further features are gained in form of punctuation patterns, antonyms from wordnet and temporal ordering of events obtained from a timeml parser (verhagen et al., 2005). using these features, the 6 listed senses (see a posteriori meanings in table 2) for the connective while can be disambiguated, in the held-out set, with an accuracy of about 65%, meaning that the classifier predicts the correct sense in two thirds of all cases. meyer and popescu-belis (2012) have also shown that such a classifier can be used to automatically label the large training data for machine translation. as a consequence, such an smt system translates discourse connectives more correctly. they further validate the method by automatically classifying up to 12 other temporal-contrastive connectives with larger training sets and by integrating these classifiers into smt as well. these experiments show that investigations based on translation spotting over large parallel data can uncover unexpected meanings of the connectives used in the source language. as explained in the next section, this technique can also be used to uncover more fine-grained differences of usages within a single rhetorical relation. 4.5. comparison and evaluation in this section, we systematically compare the translation spotting technique with sense annotation in terms of the sense tags they provide. for the french connective alors que, we have compared the sense annotation resulting from translation spotting and clustering with the labels assigned directly by annotators in the sense annotation. this enabled us to check whether the results of the two techniques provided consistent results or not. as a first comparison, we only used the 267 occurrences for which the two annotators had agreed on the label (background or contrast), and compared this label with the english connectives used to translate alors que. results are presented in table 9 (only connectives appearing with a frequency of >5% are reported). background label contrast label transpot no. % transpot no. % when 24 27.91% whereas 50 27.62% while 10 11.63% when 28 15.47% at a time when 9 10.47% while 26 14.36% as 7 8.14% although 19 10.50% zero translation 7 8.14% zero translation 13 7.18% whilst 6 6.98% whilst 11 6.08% although 5 5.81% table 9: translation equivalents according to the meaning of alors que when the two annotators agreed on a background meaning for alors que, a majority of connectives chosen by the translator also have a background meaning (like when, at a time when). in the second half of the table, among the occurrences of alors que that were labeled as contrast by the two annotators, the main connective used can only have a contrastive meaning (whereas) while all the other connectives used in translation are ambiguous and can have several labels, amongst which a contrastive meaning is always found in reference data (such as while). annotating discourse connectives by looking at their translation 79 in addition, when looking at the 134 occurrences where the annotators disagreed, we notice that 60 of them were translated by unambiguous connectives in english: 51 alors que are translated by a clearly contrastive english connective (such as although, whereas, but…) and 9 occurrences are translated with clearly temporal english connective (at the time when, now that). this confirms that translation spotting can provide disambiguation when annotators cannot. the remaining 74 occurrences are translated by ambiguous connectives in english (when, while, whilst). in those cases, the ambiguity is kept in translation. in sum, this comparison shows that the results from translation spotting are often similar to the sense labels assigned by annotators and can also provide results for an important number of cases of which annotators do not reach agreement. in addition, this technique has the advantage of providing a better way to deal with ambiguity than sense annotation. in many cases, ambiguity is revealed in translation spotting by the choice of a target language connective that can also have the same multiple meanings, as it is the case for the pair of while and alors que. in consequence, ambiguity can naturally be preserved and dealt with in such cases. on the other hand, while annotating the senses of a connective from a monolingual perspective, our experiments have shown that annotators often feel compelled to choose between various possible meanings. this can lead to arbitrary choices between two values that can in fact coexist naturally. this problem was accounted for in the pdtb by allowing any combination of labels from the sense hierarchy in order to annotate double sense tags to certain occurrences of discourse connectives. however, this technique does not ensure that annotators will identify all the meaning components of a connective, and use several tags instead of one. 5 translation spotting for the identification of sub-senses of connectives until now, we have shown that connectives can often convey more than one rhetorical relation and argued that disambiguating these different meanings in context represented a difficult task of manual annotation. in this section, we will concentrate on a different fact: most rhetorical relations can be conveyed in many languages by a whole array of different connectives. for example, a causal relation can be conveyed in french by parce que, car, puisque, étant donné que, comme, vu que, etc. (for recent surveys of cross-linguistic comparisons involving causality, see sanders & stukker, 2012; sanders & sweetser, 2009). the point is that all these connectives are not always interchangeable and therefore cannot be treated as equivalents. zufferey (2012), for example, showed through a sentence completion task and an acceptability judgment task that the connectives puisque and car were almost never interchangeable, contrary to what previous theoretical studies had concluded (e.g. lambda-l group 1975, roulet et al. 1985). the main consequence of this finding for machine translation is that assigning a cause label to a connective does not ensure that a correct translation will be achieved, since all connectives conveying a causal meaning are not interchangeable. in a nutshell, this observation means that at least in some cases, a more fine-grained annotation scheme than simple rhetorical relations such as cause, concession, temporal, etc. is needed to ensure an optimal translation of connectives. in the pdtb, cause is not the most fine-grained level, but its main subdivision between reason and result serve to separate connectives like because and all the french connectives listed above, that have a consequence-cause order of the segments, from connectives like so that have a reversed order (cause-consequence). in this section, we will limit ourselves to giving a flavor of the kind of information that is needed in order to translate causal connectives accurately (see zufferey & cartoni 2012, for a detailed presentation of these criteria). our aim is to show that translation spotting is also a very relevant annotation technique at this finer level of granularity. one of the main criteria dividing the category of causal connectives is the subjective or objective nature of the causal relation described. in some cases like (11), the causal cartoni, zufferey & meyer 80 relation relates events in the world and is therefore objective, while in other cases like (12) the causal relation involves the speaker’s own reasoning or speech act and is therefore more subjective (e.g. sanders, 1997; degand & pander maat 2003). 11. the snow is melting, because the temperature is rising. 12. john was tired, because he fell asleep. in english, this difference is not visible in terms of connectives, as because can convey both objective and subjective relations (sweetser 1990). however, in many other languages like dutch (pit 2007), german (sanders & stukker 2012) and french (zufferey 2012; degand & fagard 2012), different connectives are used to express both kinds of relations. for example, in written french, objective uses are prototypically translated by parce que while subjective uses are translated by car. this means that in order to translate occurrences of because accurately in a number of languages, the degree of subjectivity of the causal relation has to be taken into account. in this case, translation spotting provides an immediate solution for the annotation of occurrences of because, in order to provide training data for machine learning algorithms. the translation choices indeed provide this information, as can be seen in table 10, which presents the translation spotting of 196 parallel sentences containing because. no. % no. % car 76 38.78% vu que 1 0.51% parce que 63 32.14% dès lors que 1 0.51% paraphrases 27 13.78% gerund 1 0.51% zero translation 8 4.08% : 1 0.51% dans la mesure où 6 3.06% en effet 1 0.51% puisque 3 1.53% sans quoi 1 0.51% en effet 3 1.53% compte tenu que 1 0.51% étant donné que 1 0.51% du fait que 1 0.51% à défaut 1 0.51% total 196 table 10: translation spotting of the english connective because the two main translations of because in french are car and parce que. it can be assumed that the translations by car correspond to the subjective uses of because while the translations by parce que correspond to its objective uses. in order to verify this claim, we asked two experts to annotate 100 sentences containing the connective because with the objective/subjective trait. results indicate that 90% of the because sentences translated by car were annotated as subjective. similarly, 85% of the because sentences that were annotated as objective by the annotators were translated by parce que rather than car. 6 in sum, this example shows that translation spotting can also be used for very finegrained distinctions, as long as they are visible in the translations. this comparison also confirms that the information provided by the translations coincides with sense annotation made by experts and is therefore reliable, as discussed in section 4.5. 6 in contemporary spoken french, parce que is the only connective used for both kinds of relations and in writing, parce que can also convey subjective relations in some cases. annotating discourse connectives by looking at their translation 81 6 discussion the various annotation tasks presented in this paper confirm that the meanings of discourse connectives are difficult to annotate for human judges. arguably, this difficulty is at least partially related to the taxonomy of discourse relations that the annotators are instructed to apply. some fine-grained distinctions are indeed difficult to annotate reliably; for example it is only at the top level of their taxonomy (containing only four generic classes) that the pdtb annotators reached a reliable, even though not perfect, agreement level (92%) (miltsakaki et al. 2009). however, this kind of general annotation is not precise enough for many applications, including those involving a form of crosslinguistic mapping. another problem related to this type of annotation is that there is no consensus in the literature about what an optimal taxonomy of discourse relations should consist of (see e.g. hovy 1990 for a discussion of this problem). the ideal granularity of the taxonomy is probably not universal but strongly depends on the goal of the annotation. in the case of the comtis project underlying this study, the annotation of discourse connectives served the goal of pre-processing for machine translation systems, enabling a disambiguation of the meaning of connectives, leading to an accurate translation choice. as we have shown in this paper, for this purpose a fine-grained taxonomy is required, in order to capture the sometimes subtle differences of meanings between connectives. as our experiments on alors que and while have demonstrated, this fine-grained annotation is not reliably achieved by human annotators, even when a careful and time-consuming training procedure has been implemented. this led us to consider an alternative route to sense annotation, making use of the information provided by the translation and the intuitive knowledge that native speakers have about the possibility to use a connective in a given sentence (cf. the sentence completion tasks that are part of the second step of our method). from a theoretical perspective, there seems to be a justification of the acute difficulty of annotating connectives, compared to other lexical items. many studies on discourse connectives have argued that these lexical items encode procedural rather than conceptual information (e.g. blakemore, 2002; moeschler, 2002; wilson, 2011). in other words, their role in the sentence is to instruct the addressee about the way some of the arguments are related. for example, the connective therefore instructs the hearer to look for a consequence between the segment preceding the connective and the one following it. this property of discourse connectives can at least partially explain why their meanings are often difficult to pin down by human annotators. indeed, procedural meaning is not as easily accessible to conscious introspection as conceptual information (blakemore, 2002). however, speakers have a very reliable ability to intuitively judge the acceptability in a given context. just like it is the case for syntax, this intuitive ability is dependent on the language faculty and is not accompanied by a form of declarative knowledge. this difference explains why the task of sense annotation is often difficult for annotators while the sentence completion tasks involved in the translation spotting technique are rather straightforward. thus, the translation spotting technique avoids one of the main problems related to discourse connectives: the difficulty to reason explicitly about their meaning in context. this task is replaced by several more manageable ones for annotators: identifying a translation and, in the second phase of clustering, using a set of connectives to fill in blanks in sentences. the clustering of senses inferred from these interchangeability tests provides a more reliable indication on the meaning of connectives than the application of a pre-defined set of tags indicating coherence relations, which are often difficult to define and identify. moreover, the clustering of senses is also more flexible, as tags are defined according to the meaning of connectives in translation, rather than beforehand. finally, because the annotation tasks involved in translation spotting are rather easy, this technique provides an interesting way to gather rapidly an important amount of data. cartoni, zufferey & meyer 82 this paper has also shown that a cross-linguistic perspective provides some new insights on the possible meanings of connectives in a given language. for instance, the translation of while by tant que in french indicated that this connective could establish a condition meaning. this tag was however not assigned to while in the pdtb. moreover, we saw in section 5 that looking at translations could also be used to investigate some very fine-grained properties of connectives conveying the same rhetorical relation (i.e., causality). all these observations confirm that looking at a language through the mirror of another language can bring new insights on the meaning of these lexical items, even from a monolingual perspective. the translation spotting method also has some obvious limitations. first and foremost, it relies on the choices made by the translator. even with professional translators as the ones involved in our corpus, the translation choice for one particular occurrence of a connective is the result of a specific interpretation and incorrect translations, or at least translations involving meaning shift, cannot be excluded. however, we argue that the important amount of parallel sentences investigated should flatten this bias. consequently, translation spotting can be expected to be a reliable method only when applied over a large amount of data. this requirement is another limitation of this method. another potential problem comes from the fact that it is dependent on the presence of multiple translations in the target language. indeed, a connective could have many theoretical senses in one language but all these senses could be covered by one single connective in the target language. whether this limitation is a problem or not depends on the expected generalization of the annotation. if the aim of the annotation is to provide an accurate translation in a given target language, this ambiguity can be carried over without producing translation errors. however, this technique will not provide indications on the different meanings of this connective that could be reused for a different target language. moreover, when an ambiguity is repeatedly preserved across languages, the status of this ambiguity should be questioned. for example, it is possible that sometimes background and contrast are two values of a connective that are denoted at the same time in a given occurrence, just like some other connectives require several labels to account for their meaning. the fact that a connective covering these two meanings is also used in the translation (as in the example of the pair made of alors que and while) might mean that the value “background-contrast” can be treated as a single unit, or a somehow underspecified value. in other words, the possibility that connectives can sometimes convey two compatible but different rhetorical relations in a single occurrence has to be taken into account, as it is the case in the pdtb where annotators are allowed to use double tags for single connective occurrences. another example of such a double meaning can be observed in some occurrences of since, where a temporal and a causal meaning both seem to be conveyed simultaneously. further confirmation for the existence of such double sense labels can be obtained from experiments with automated sense classifiers and machine learning. before training the classifiers, the cases where human annotators disagreed can be resolved by assigning double labels, for instance, when one annotator used a temporal sense for an occurrence of since, and the other annotated a causal sense, this disagreement can be resolved by assigning a label temporal-causal (similarly, background-contrast for the french connective alors que). for since, an automated classifier using three labels (temporal, causal and temporal-causal) almost reaches the same performance as one that uses temporal and causal only. for alors que a three-way classifier (including background-contrast) even reaches higher performance than the twoway one – which is quite surprising, as usually, more classes means more difficulties for automated tools to disambiguate them (meyer et al. 2011). this might provide further evidence for the existence and usefulness of double sense labels for discourse connectives. annotating discourse connectives by looking at their translation 83 7 conclusion in this paper, we demonstrated through several annotation experiments that annotating the senses of discourse connectives is a difficult task for which human annotators do not reach a truly reliable agreement. we proposed the use of an alternative technique to perform this annotation, making use of the clues provided by the translation of the connective in a target language. when the target language does not provide a direct disambiguation, all translations are clustered into different senses based on the possibility to replace the various connectives in the target language. the clusters are formed based on native speakers’ judgments about the possibility to use connectives interchangeably in a sentence. this technique therefore provides a more reliable way than traditional sense annotation to label connectives with their meaning in context. this technique also opens new avenues for further cross-linguistic research on discourse relations and connectives. the approach proposed in this paper offers an interesting and easy way to gather contrastive data that can be extended to larger-scale contrastive analyses. as demonstrated in the case of while and the category of causal connectives, the systematic comparison of a large amount of correspondences in translated corpora can provide a complete picture of the equivalences between languages, and provide useful indications about the granularity of discourse relations that are required to describe them cross-linguistically. if extended to a larger set of languages and connectives in a variety of genres, this method would allow for more empirically grounded generalizations about discourse relations in the world's languages. in particular, the fact that one particular occurrence can convey two discourse relations simultaneously, and that this double meaning is repeatedly found in other languages might reflect some general tendencies about the cognitive similarity of some discourse relations. acknowledgements this study was funded by the swiss national science foundation through the comtis sinergia project (www.idiap.ch/comtis). the authors would like to thank the annotators for their careful and meticulous work. references ron artstein and massimo poesio (2008). inter-coder agreement for computational linguistics. computational linguistics 34(4):555-596. mona baker (1993). corpus linguistics and translation studies: implications and applications. in m.baker et al. (eds), text and technology: in honor of john sinclair. john benjamins, amsterdam/philadelphia. petra bayerl and karsten paul (2011). what determines inter-coder agreement in manual annotations? a meta-analytic investigation. computational linguistics 37(4):699-725. bergljot behrens and cathrine fabricius-hansen (2003). translation equivalents as empirical data for semantic/pragmatic theory. in jaszczolt k, turner jen (editors), meaning through language contrast. amsterdam: benjamins. 463-477. diane blakemore (2002) meaning and relevance: the semantics and pragmatics of discourse markers. cambridge university press, cambridge, usa. jean carletta (1996). assessing agreement on classification tasks: the kappa statistic. computational linguistics, 22(2):249–254. bruno cartoni and thomas meyer (2012). extracting directional and comparable corpora from a multilingual corpus for translation studies. in proceedings of lrec 2012, pages 2132-2137, istanbul, turkey. http://www.idiap.ch/comtis cartoni, zufferey & meyer 84 bruno cartoni, sandrine zufferey, thomas meyer and andrei popescu-belis (2011). how comparable are parallel corpora? measuring the distribution of general vocabulary and connectives. in proceedings of 4th workshop on building and using comparable corpora, pages 78-86, portland, usa. eugene charniak and mark johnson (2005). coarse-to-fine n-best parsing and maxent discriminative reranking. in proceedings of the 43rd annual meeting of the association for computational linguistics (acl) pages 173–180. ann arbor, mi. laurance danlos and charlotte roze (2011). traduction (automatique) des connecteurs de discours. in proceedings of taln 2011, montpellier, france. liesbeth degand (2004). contrastive analyses, translation, and speaker involvement: the case of puisque and aangezien. in m. achard and s. kemmer (eds.), language, culture and mind, pages 1-20, stanford: csli publications. liesbeth degand and benjamin fagard (2012). competing connectives in the causal domain: french car and parce que. journal of pragmatics 44(2): 154-168. liesbeth degand and henk pander maat (2003). a contrastive study of dutch and french causal connectives on the speaker involvement scale. in a. verhagen and j. maarten van de weijer (editors), usage-based approaches to dutch, pages 175-199, lot, utrecht. helge dyvik (1998). a translational basis for semantics. in johansson, stig & signe okselfjell (eds) corpora and crosslinguistic research: theory, method and case studies, pages 51-86, amsterdam: rodopi. k. anders ericsson and herbert simon (1980) verbal reports as data. psychological review 87(3):215-251. michael halliday and ruqaiya hasan (1976). cohesion in english. longman, london, uk. ed hovy (1990). parsimonious and profligate approaches to the question of discourse structure relations. in proceedings of the fifth international workshop on natural language generation. pittsburgh, pennsylvania. stéphane huet, julien bourdaillet and philippe langlais (2009). intégration de l’alignement de mots dans le concordancier bilingue transsearch. in proceedings of taln’09, senlis, france. philipp koehn (2005). europarl: a parallel corpus for statistical machine translation. in proceedings of the tenth machine translation summit, september 13-15, pages 79-86, phukhet, thailand. alistair knott and robert dale (1994). using linguistic phenomena to motivate a set of set of coherence relations. discourse processes 18(1):35-62. lambda-l, groupe (1975). car, parce que, puisque. revue romane 10:248-280. richard j. landis and gary g. koch (1977). the measurement of observer agreement for categorical data. biometrics, 33(1):159–174. william mann and sandra thomson (1992). relational discourse structure: a comparison of approaches to structuring text by 'contrast'. in hwang s. & merrifield w. (eds.), language in context: essays for robert e. longacre. sil, pages 19-45, dallas, usa. christopher manning and dan klein (2003). optimization, maxent models, and conditional estimation without magic. tutorial at hlt-naacl and 41st acl conferences. edmonton, canada and sapporo, japan. thomas meyer and andrei popescu-belis (2012). using sense-labeled discourse connectives for statistical machine translation. in proceedings of the eacl 2012 workshop on hybrid approaches to machine translation (hytra), avignon, france, pp. 129-138. thomas meyer, andrei popescu-belis, sandrine zufferey and bruno cartoni (2011). multilingual annotation and disambiguation of discourse connectives for machine annotating discourse connectives by looking at their translation 85 translation. in proceedings of 12 th sigdial meeting on discourse and dialog, pages 194-203, portland, usa. eleni miltsakaki, nikhil dinesh, rashmi prasad, aravind joshi and bonnie webber (2005). experiments on sense annotations and sense disambiguation of discourse connectives. in proceedings of the tlt 2005 (4th workshop on treebanks and linguistic theories), barcelona, spain. eleni miltsakaki, livio robaldo, alan lee and aravind joshi (2008). sense annotation in the penn discourse treebank. in alexander gelbukh (editor), computational linguistics and intelligent text processing. lecture notes in computer science, pages 275-286, springer berlin / heidelberg. jacques moeschler (2002). connecteurs, encodage conceptuel et encodage procédural. cahiers de linguistique française 24:265-292. dick noël (2003). translations as evidence for semantics: an illustration. linguistics 41(4):757-785. the pdtb research group (2007). the penn discourse treebank 2.0 annotation manual. ircs technical reports series, 99p. mirna pit (2007). cross-linguistic analyses of backward causal connectives in dutch, german and french. languages in contrast, 7(1), 53–82. andrei popescu-belis, bruno cartoni, andrea gesmundo, james henderson, cristina hulea, paola merlo, thomas meyer, jacques moeschler and sandrine zufferey (2011). improving mt coherence through text-level processing of input texts: the comtis project. in proceedings of tralogy 2011 (translation careers and technologies: convergence points for the future), paris, france. andrei popescu-belis, thomas meyer, jeevanthi liyanapathirana, bruno cartoni and sandrine zufferey (2012). discourse-level annotation over europarl for machine translation: connectives and pronouns. in proceedings of lrec 2012, pages 27162720, istanbul, turkey. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi and bonnie webber (2008). the penn discourse treebank 2.0. in proceedings of 6th international conference on language resources and evaluation (lrec), pages 2961–2968, marrakech, morocco. eddy roulet, antoine auchlin, jacques moeschler, christian rubattel and marianne schelling, (1985). l'articulation du discours en français contemporain. peter lang, berne, switzerland. charlotte roze, laurance danlos and philippe muller (2010). lexconn: a french lexicon of discourse connectives. in proceedings of multidisciplinary approaches to discourse (mad 2010), moissac, france. ted sanders (1997). semantic and pragmatic sources of coherence: on the categorization of coherence relations in context. discourse processes 24:119-147. ted sanders, wilbert spooren and leo noordman (1992). towards a taxonomy of coherence relations. discourse processes 15, 1-36. ted sanders and ninke stukker, (2012). causal connectives in discourse: a crosslinguistic perspective. special issue of journal of pragmatics 44 (2):131-137. ted sanders t. and eve sweetser (2009). causal categories in discourse and cognition. walter de gruyter, berlin, germany. michel simard, (2003). translation spotting for translation memories. hlt-naacl 2003, workshop: building and using parallel texts data driven machine translation and beyond. wilbert spooren and liesbeth degand (2010). coding coherence relations: reliability and validity. corpus linguistics and linguistic theory 6 (2): 241-266. marc verhagen, inderjeet mani, roser sauri, jessica littman, robert knippen, seok bae jang, anna rumshisky, john phillips, james pustejovsky (2005). automating cartoni, zufferey & meyer 86 temporal annotation with {tarsqi}. proceedings of the 43th annual meeting of the association for computational linguistics (acl), demo session (pp. 81–84). ann arbor, usa. jean véronis and philippe langlais (2000). evaluation of parallel text alignment systems: the arcade project. in parallel text processing. kluwer academic publishers, text speech and language technology series: 369-388. bonnie webber and aravind joshi (2012). discourse structure and computation: past, present and future. proceedings of the acl-2012 special workshop on rediscovering 50 years of discoveries. jeju, republic of korea. 42–54, deirdre wilson (2011). the conceptual-procedural distinction: past, present and future. in escandell-vidal, v. et al. (editors), procedural meaning: problems and perspectives, bingley: emerald group publishing. 3-31. sandrine zufferey (2012). car, parce que, puisque revisited. three empirical studies on french connectives. journal of pragmatics 44(2), 138-153. sandrine zufferey and bruno cartoni (2012). english and french causal connectives in contrast. languages in contrast 12(2):232-250. sandrine zufferey, liesbeth degand, andrei popescu-belis and ted sanders (2012). empirical validations of multilingual annotation schemes for discourse relations. eighth joint iso-acl/sigsem workshop on interoperable semantic annotation, pages 77-84, pisa, italy. intro11042011 dialogue and discourse 1 (2011) 1-10 doi: 10.5087/dad.2010.001 ©2011 hannes rieser and david schlangen introduction to the special issue on incremental processing in dialogue hannes rieser hannes.rieser@uni-bielefeld.de department of linguistics and literary studies bielefeld university universitätsstrasse 25 33615 bielefeld, germany david schlangen david.schlangen@uni-bielefeld.de department of linguistics and literary studies bielefeld university universitätsstrasse 25 33615 bielefeld, germany 1 introduction the topic of this special issue is “incremental processing in dialogue”, by which we mean, very broadly, the successive processing of input (or generation of output) in increments smaller than whole utterances, as it can be observed in natural dialogue. due to idealisation assumptions in linguistics and philosophy subscribed to by many scholars since frege and saussure, incrementality became only a topic of research not very long ago. setting philologies and hermeneutic’s programs (heidegger, gadamer) aside, the first researchers who systematically dealt with the incrementality of language production and understanding were those working in ethno-methodology and in a field to become conversation analysis (ca) later on. this can be seen from the early papers of jefferson (1972, 1974), the famous sacks, schegloff, and jefferson articles (1974, 1977) and especially from work of schegloff building on these attempts (1979, 1982). these scholars were interested in the contributionswithin-conversation perspective not restricted to dialogue proper, so their focus was on large increments and their regular distribution, for example on the allocation of turns or on the placing of repairs. a ca perspective complemented by an experimental one is here implemented in the paper on incrementality in dialogue: evidence from compound contributions by chr. howes, m. purver, p.g.t. healey, g.j. mills, and e. gregoromichelaki. simultaneously with ca research, scholars working on the psychology of language processing considered small increments such as phonemes and clusters of them in words (marslen-wilson 1973), without however, taking extra-word units into account; something that can perhaps be attributed to the experimental methodology they used at that time. the incremental perspective gained popularity by research taking up the conversation analysis tradition and combining it with the paradigms of experimental psycholinguistics. this is the hallmark of work done by h. clark and his collaborators (clark and marshall 1981, clark and wilkes-gibbs 1986), which inter alia was concerned with incrementality in syntax production and reference resolution, i.e. with in-turn regularities. in addition, clark’s notion of grounding tried to shed light on the fine-grained structure of speakers’ successive contributions in dialogue (clark and schaefer 1989). as can be seen from current literature (e.g., roque and traum 2008, poesio and rieser 2010), interest in grounding matters and their fine-grained reconstruction is still continuing. rieser and schlangen 2 in the 1980s and 90s, the increasing availability of two different experimental techniques gave researchers access to fine-grained temporal information about the comprehension process without interfering with it. eye-trackers provide a “window to the mind”, through which the current focus of attention is to be tracked, be that words during reading (see rayner 2009 for an introduction) or potential referents of expressions (as in the “visual world paradigm”, e.g. trueswell and tanenhaus eds. 2005). this research tradition is also reflected in the present volume: ilkin and sturt’s article active prediction of syntactic information during sentence processing uses eye-tracking during reading to show that certain kinds of phrases were often skipped in contexts that make them predictable, hence demonstrating the role that prediction plays in incremental processing; vasishth and drenhaus’ locality effects in german uses eye-tracking (among other methods) to argue that processing load increases with distance in sentence processing. brown-schmidt and hanna in talking in another person’s shoes: incremental perspective-taking in language processing review experiments from the visual world paradigm to argue for a constraint-based view of perspective taking in dialogue. incrementality modelling based on corpus investigation and experimental evidence coming especially from the visual world paradigm can be found in poesio and rieser’s an incremental model of anaphora and reference resolution based on resource situations. to continue with the main incrementality research line, the recording of event-related brain potentials, and especially the n400 associated with semantic mismatches, provides another way to gain insights in the interplay of information sources during language comprehension (e.g., van berkum et al. 1999). changing to language processing and syntax, the paradigm most frequently associated with an incrementality perspective is dynamic syntax (ds; kempson, meyer-viol, gabbay 2001, cann, kempson, marten 2005), but some attempts at devising theories of incremental processing were already made earlier on, for example, in research undertaken by neumann (1994, 1998) and wirén (1992, 1994) at linköping, sweden, and at the dfki saarbrücken, germany. unfortunately, awareness of this latter research tradition has been nearly lost. two papers of the present collection use ds modelling among a lot of other things, andrew gargett’s incrementality and the dynamics of routines in dialogue and the paper incrementality and intention-recognition in utterance processing by gregoromichelaki, kempson, purver, mills, cann, meyer-viol and healey. turning to dialogue theory proper, perhaps the most decisively incremental approach was suggested with ptt, where incrementality is essentially based on poesio’s notion of microconversational events (mces, poesio 1995). the question of options concerning incremental units such as mces is discussed in schlangen and skantze’s a general, abstract model of incremental dialogue processing (their ius, see below). mces have since 1995 become part and parcel of standard ptt models (poesio and traum 1997) and later ptt versions using ltag and compositional drt (e.g. poesio and rieser 2010). ptt is a good example of how concurrent developments in syntax (tree grammars, ltag) and dynamic semantics (drt variants) boosted the development of incrementality approaches. in the same vein, a link between ds and dialogue theory was established in purver and kempson (2004) and papers based on that. the present stage of this development in ds can be seen from gargett’s and gregoromichelaki et al.’s papers. for some time, the contact between research on fine-grained incrementality on the phonological and prosodic level and the investigation of sequences of larger structures in discourse seemed to have been neglected. however, recent computational work (e.g., skantze and schlangen 2009, edlund et al. 2008) shows that there now is renewed interest in combining “high” level and “low” level incrementality. there is a good chance that yet another one of h. clark’s assumptions – namely that providing and monitoring feedback is a process that continuously accompanies all contributions – can be reconstructed and simulated. we find fine-grained incrementality detailed in this volume in the paper evaluation and optimisation of incremental processors by timo baumann, okko buß and david schlangen and a more global and abstract perspective on incremental goings-on in david schlangen and introduction to special issue 3 gabriel skantze’s contribution a general, abstract model of incremental dialogue processing, as well as in the contribution by devault, sagae and traum, incremental interpretation and prediction of utterance meaning for interactive dialogue. the papers collected in this volume are either by founders or later proponents of the incrementality research movement. from the references quoted therein one can gather that they have maintained this research line for many years. all of the papers originate from a workshop on “incrementality in verbal interaction” hosted from the 8th to the 10th of june, 2009 by the zif, the interdisciplinary research centre of the university of bielefeld. the workshop was organised by ruth kempson (king’s college london), hannes rieser, petra wagner (both bielefeld university) and david schlangen (then university of potsdam) and financed by the collaborative research centre “alignment in communication”, which is funded by the german research foundation. the workshop’s invited speakers were: atterer, michaela (university of potsdam, germany) baumann, timo (university of potsdam, germany) buss, okko (university of potsdam, germany) de vault, david (ict / usc los angeles, usa) dubey, amit (university of edinburgh, uk) edlund, jens (kth, sweden) gregoromichelaki, eleni (king’s college london, uk) hanna, joy (oberlin college, usa) harbusch, karin (university of koblenz-landau, germany) healey, pat (queen mary university of london, uk) kempson, ruth (king’s college london, uk) knoeferle, pia (citec, bielefeld university, germany) kruijff, geert-jan m.(dfki, language technology lab, saarbrücken, germany) poesio, massimo (university of trento, italy / university of essex, uk) purver, matthew (queen mary university of london, uk) rieser, hannes (bielefeld university) sato, yo (university of hertfordshire, hatfield, uk) schlangen, david (university of potsdam, germany) schuler, william (university of minnesota, usa) skantze, gabriel (kth, stockholm, sweden) sturt, patrick (university of edinburgh, uk) van berkum, jos (mpi psycholinguistics nijmegen, the netherlands) vasishth, shravan (university of potsdam, germany) wagner, petra (bielefeld university) at the workshop the following talks were given in the order indicated: jos j.a. van berkum incrementality and beyond: what erps tell us about utterance comprehension shravan vasishth on anticipation and integration processes patrick sturt the dynamics of long distance dependency formation: pre-verb structural integration in a head-initial language pia knoeferle variation in the time course of visual context effects on sentence comprehension rieser and schlangen 4 christine howes some empirical observations on split utterances massimo poesio and hannes rieser incremental anaphoric interpretation with micro conversational events eleni gregoromichelaki mechanistic accounts of dialogue and split utterances: probing the limits of grammar ruth kempson incremental growth of interpretation as natural language syntax david schlangen a general, abstract model of incremental dialogue processing gabriel skantze incremental processing in human-computer number dictation michaela atterer, okko buss and timo baumann incremental asr, nlu and dialogue management in the potsdam inpro p2 system david devault incremental understanding in virtual human dialogue systems joy e. hanna incremental perspective-taking in conversation: putting language processing in context and context in language processing amit dubey an incremental syntax/semantics interface for psycholinguistic modelling william schuler simple computational model of interactive language comprehension pierre lison incremental processing of spoken dialogue for human-robot interaction pat healey incremental processing in collaborative interactions jens edlund tread carefully – collaboration in small steps karin harbusch incremental sentence production inhibits clausal coordinate ellipsis: a comparison of spoken and written language yo sato incrementality, bi-directionality and grammar learning discussion about talks and submitted papers continued among the speakers and their reviewers throughout 2010-2011. we here subsume the papers ultimately included in this volume under the systematic fields they prototypically belong to. introduction to special issue 5 2 experimental work the issue opens with a review article by sarah brown-schmidt and joy e. hanna, talking in another person’s shoes: incremental perspective-taking in language processing. in this paper, the authors review psycholinguistic evidence for incrementality in language processing, focussing on the role of perspective taking. taking sides in the ongoing debate on how to interpret the experimental results, the authors argue that a constraint-based account in which perspective is one of many constraints that guides language processing decisions best fits the data. the article by zeynep ilkin and patrick sturt, active prediction of syntactic information during sentence processing, presents an experiment that shows that during reading, and measured by eye-tracking, plural noun phrases were skipped more often than singular noun phrases, in contexts where there was a high expectation for a plural. shravan vasishth and heiner drenhaus present in their article locality effects in german a collection of experiments, using different paradigms, which show that in relative clauses, increasing the distance between the relativized noun and the relative-clause verb makes it more difficult to process the verb in the relative clause; this, they argue, supports a view where dependency-resolution cost is responsible together with expectation-based facilitation for determining processing cost. 3 abstract models frequently, the solution of selected problems, say parsing of a string or setting up a semantic representation concurrently, is fairly clear and can be done in a locally consistent way, given specific idealizing assumptions. however, what one would need to know more about is the global embedding of the local problem in an incremental model and the interaction of its components with various other modules. knowledge concerning these matters comes from david schlangen’s and gabriel skantze’s contribution a general, abstract model of incremental dialogue processing. they specify a general framework to set up architectures for incremental processing in dialogue systems. their focus is on the options available to system designers interested in handling data such as sub-utterance edits, feed-back phenomena or split utterances. the authors’ aim is to specify necessary components of such a system thereby delineating a large class ‘from non-incremental pipelines to fully incremental, asynchronous, parallel, predictive systems’. having introduced the advantages of incremental processing and a description of the modular structure of their system, the authors explain the setup of the network topology and the different types of information flow used. the structure of the modules and module behaviour is presented. the question which type of incremental units (ius) can be used and which relations between them have to be assumed is given detailed consideration. the system also allows for revisions of output. in the end we are provided with some example specifications showing how the conceptual classification can be applied to existing implementations. this indicates not only that the abstract model can be used in meta-theoretical work, i.e. classifying and comparing implemented incremental approaches, but also that the model can be used for theories and descriptions of ongoing information processes not directly tied to computational linguistics or ai. 4 practical computational models incremental processing of spoken language poses special problems to the components designed to achieve this task. the development of such components, called incremental processors subsequently, and their evaluation is discussed in the paper evaluation and optimisation of incremental processors by timo baumann, okko buß and david schlangen. the special task incremental processors have to meet is as follows: they must be able to produce partial albeit growing output given the partial input they are fed with by other rieser and schlangen 6 components of the system in question, components, perhaps also working in an incremental fashion. in addition, inputs can be revised in the course of later processing. in this paper incremental processors designed for this task are specified and evaluated. evaluation meets special challenges since interim outputs have to be considered in addition to final ones. based on their previous practical work, baumann, buß and schlangen lay down requirements for individual incremental processors and derive from these general description metrics for the evaluation of their performance. the chosen representation format for the outputs enables them to define metrics encompassing the quality of results, the times of their formation and alternative hypotheses, if any, entertained by the system. evaluation amounts to comparing actual outputs with idealised ones, called “gold standards”, and the development of metrics. of these we are given similarity metrics, timing metrics and diachronic ones, measuring incrementality in terms of pace, fit and persistence, respectively. baumann et al.’s specification of an evaluation framework should prove a valuable general contribution to the growing field of incremental spoken dialogue systems. david devault, kenji sagae and david traum present in the paper interpretation and prediction of utterance meaning for interactive dialogue techniques for building one such module for incremental dialogue systems, namely for the component that does interpretation. in their setup, incremental interpretation is done via prediction of the meaning that the ongoing utterance will ultimately have, once it is completed. they show how such a component can be used in a system to initiate completions of user utterances. 5 from grammar to dialogue among formal grammars dynamic syntax (ds) was one of the first incrementally working algorithms. until around 2000 ds focused on single propositions; it has been extended since to the reconstruction of particular dialogue phenomena such as split utterances. andrew gargett’s contribution incrementality and the dynamics of routines in dialogue extends ds in various ways. responding to the on-going discussion in cognitive psychology on linguistic routinisation, his main interest is in developing a dual processing model of linguistic routinisation ranging from fixed idioms to looser collocational constructions. his main interest is to capture routinised and non-routinised language in one comprehensive theory based on incrementality: non-routinised language use is modelled via the rule-based account of ds whereas for formulaic language memory-based processes operating as larger stable patterns are introduced. by way of example, in conversations one observes the emergence of routines out of initially regular wordings, that is the emergence of ‘constants’ with a fixed interpretation out of compositionally set up material. this has been shown in several studies of e.g. h. clark (see the contributions in clark, ed. 1992) or of s. garrod and co-workers. in gargett’s words, there is a move from initially rule-based production to a subsequent memory-based one. basing on this insight, he gives special attention to the interaction of both types of devices. in order to model routines the lexical architecture of ds, up until now working with lexical actions, i.e. rules introducing words, is extended with patterns of ‘stable’ semantic output, a sort of “frozen semantics”. a dual model allows for competition between rule-based procedures and interpretation via stored semantic input. seen from the ds development perspective, the account adds dynamicity to the lexicon which had been missing so far. the paper incrementality and intention-recognition in utterance processing by gregoromichelaki, kempson, purver, mills, cann, meyer-viol and healey discusses at the outset the role of higher order (“gricean”) intentions in communication and argues for the development of alternative models which might be more plausible given the psychological restrictions of humans. the intention topic is fused with the incrementality assumption on the theoretical and the empirical side. concerning the empirical side, the focus is on split utterances which are a prototypical incrementality paradigm involving switches of speakers. on the theoretical side incremental dynamic syntax (ds) is used in a version going back to introduction to special issue 7 work of purver and kempson (purver and kempson 2004) implementing the notion of a bidirectional ds. the bi-directional version of ds allows for a regular switch from parsing to generation using partial information on the generation side. all of that is embedded in a broad discussion of variants of intentionalism and challenges to these, the challenges coming from philosophy (cf. macdonald and papineau eds. 2006) and experimental psychology (e.g. pickering and garrod 2004). similarly, detailed attention is given to the problem of incrementality in speech production and understanding winding up to the claim that cooperative processes operate on propositional and sub-propositional levels. assuming this to be plausible, a grammar model dealing with sub-sentential contributions of different speakers is required. this is developed after discussing intention-based approaches which are dismissed as a generally valid tool. against initially given intentions as usually assumed in models based on dialogue acts, the preferred route is to implement coordination on a sub-intentional, “low-level” basis. this does not imply that intentions are discarded once and for all, on the one hand, so the argument goes, one might need them in particular settings, where there is common information for speaker and hearer, on the other hand, it is shown resorting to experiments that the assumption of emerging (i.e. not “pre-fabricated”) intentions is a plausible one. so, the opposition set up in the end is “initially given” vs. stepwise emerging intentions. the whole discussion serves as a methodological preparation to the introduction of ds as the main tool to incrementally represent split utterances, the principal ds feature in this context being the parsing-grammar coordination. finally, we get the modelling of split utterances in ds and a concluding chapter bringing together the different strands of argumentation. situation-sensitive resolution of anaphora in the ptt dialogue paradigm is the subject of poesio’s and rieser’s paper an incremental model of anaphora and reference resolution based on resource situations. they start from the observation familiar from work on corpora such as trains (cf. allen et al. 1995) or the bielefeld corpora saga (cf. lücking et al. 2010) that utterances in general and anaphora in particular are interpreted incrementally. what the increments are can often be seen from repairs, clarification requests, interruptions by other or acknowledgements. data like these provide an indication of how processes of understanding work, thus complementing findings from experimental paradigms like the visual world paradigm. anaphora can occur in the guise of pronouns or definite nps. the theory of referring expressions developed in the paper integrates several accounts: löbner’s functional approach (löbner 1987) and a theory of anaphoric accessibility using resource situations (the situations one gets suitable antecedents from) as developed in situation semantics and in previous work of poesio (1995) as well as findings in experimental psychology about incrementality and reference resolution. observations from corpora and experimental evidence are bound together in a unified theory of the semantics and the pragmatics of referring (anaphoric) expressions. this in turn is reconstructed in an updated version of ptt. in particular, incrementality is modelled through micro-conversational events, defaults and a parallelism constraint. anaphora resolution uses resource situations, visual scenes, shifts over referential domains and a host of parsing rules. finally, it is shown how the theory’s predictions fare with respect to the results of experimental psychology, pointing out shortcomings on both sides. 6 corpus studies and experimental work investigations of incremental processes in dialogue have often been bound to the study of socalled completions or “split utterances” across agents’ different turns in dialogue (for earlier work cf. h. clark 1996, skuplik 1999, rieser and skuplik 2000, poncin and rieser 2006). completions (subsequently compound contributions, ccs) are the topic of the study on incrementality in dialogue: evidence from compound contributions by chr. howes, m. purver, p.g.t. healey, g.j. mills, and e. gregoromichelaki. they provide two approaches to ccs in this paper, a standard one, namely a corpus investigation showing that ccs are indeed rieser and schlangen 8 central for coordination in dialogue and a newly and specially designed experimental study showing the on-line effects of “artificially” introduced ccs. at the start we get a useful clarification of ca notions subsequently used such as turn or contribution. the related work section provides a useful overview of approaches to ccs ranging from ca to current dialogue theory. their list of research questions to be treated, for example investigation of split point occurrences, shows the study to be perhaps the first one looking for dialogue effects of ccs. the corpus study investigates types of ccs, one result being that same person ccs are more frequent than more spectacular cross-agent ones. in addition, parameters of ccs are considered which have already been suggested in previous studies, for example where split points can occur in contributions or the grammatical completeness of the constituents making up ccs. howes et al. also indicate which effects their findings might have on the development of grammars/parsers for ccs, e.g. with respect to back-tracking procedures to be set off. the experimental manipulation in the second study using a chat-based tool was carried out to show what the ‘effect of ccs on the dynamics of a conversation’ is. despite their different format, the results of the two studies yield similar results, for example, that there are syntactic constraints on where split points usually appear. the issue closes with a paper by karin harbusch, incremental sentence production and clausal coordinate ellipsis, which presents a treebank study of clausal coordination in spoken dutch and german. this study shows that clausal coordinate ellipsis (cce) occurs much more frequently in written than in spoken language, with a particular pattern when different types of cce are studied in detail. the author argues that the pattern cannot be accounted for in terms of audience design but rather needs an explanation that assumes that the grammatical planning of spontaneous speech is restricted to a single finite clause. acknowledgements we are grateful to the following scholars who were involved in the reviewing process: gregory aist, jan alexandersson, jens allwood, ron artstein, markus bader, adrian bangerter, dan bohus, susan brennan, hendrik buschmeier, rui chaves, robin cooper, david de vault, vera demberg, jens edlund, stefan frank, edward gibson, ruth kempson, udo klein, stefan kopp, edmundo kronmuller, geert-jan kruijff, roger levy, lutz marten, christian pietsch, paul piwek, massimo poesio, matthew purver, antoine raux, jan peter de ruiter, adrian staub, amanda stent, matthew stone, david traum, shravan vasishth, mija van der wege. we gratefully acknowledge support by the centre for interdisciplinary studies, zif, bielefeld university, the crc “alignment in communication” (sfb 673), bielefeld university, and the german research foundation (dfg). references james f. allen, lenhart k. schubert, g. ferguson, p. heeman, c. h. hwang, t. kato, m. light, n. martin, b. miller, m. poesio, and d. r. traum. the trains project: a case study in building a conversational planning agent. in: journal of experimental and theoretical ai, 7, 7-48, 1995 ronnie cann, ruth kempson, lutz marten. the dynamics of language. oxford: elsevier, 2005. herbert h. clark (ed.). arenas of language use. the univ. of chicago press and csli publications, 1992. herbert h. clark. using language. cambridge university press: cambridge: ma, 1996. introduction to special issue 9 herbert h. clark and catherine r. marshall. definite reference and mutual knowledge. in: aravind joshi, bonnie webber, and ivan sag. (eds.), elements of discourse understanding. cambridge univ. press: cambridge: ma, 1981. herbert h. clark and deanna wilkes-gibbs. referring as a collaborative process. in: cognition, 22, 1-39, 1986. herbert h. clark and edward f. schaefer. contributing to discourse. in: cognitive science, 13: 259-294, 1989. jens edlund, joakim gustafson, mattias heldner, and anna hjalmarsson. towards humanlike spoken dialogue systems. in: speech communication, 50(8-9), 630-645, 2008. jonathan ginzburg. the interactive stance: meaning for conversation. oxford, 2011. gail jefferson. error correction as an interactional resource. in: language in society, 3(2), 181-199, 1974. gail jefferson. side sequences. in: d. sudnow (ed.), studies in social interaction. new york: free press, 294-338, 1972. gail jefferson, harvey sacks. the preference for self-correction in the organization of repair in conversation. in: language, 53: 361-382, 1977. ruth kempson, wilfried meyer-viol, dov gabbay. dynamic syntax: the flow of language understanding. blackwell, 2001. andy lücking, kirsten bergmann, florian hahn, stefan kopp, and hannes rieser. the bielefeld speech and gesture alignment corpus (saga). in proceedings of the lrec workshop on multimodal corpora: advances in capturing, coding and analyzing multimodality. 2010. william marslen-wilson. linguistic structure and speech shadowing at very short latencies. in: nature, 244(5417), 522-3, 1973. graham macdonald and david papineau (eds.). teleosemantics. oxford: clarendon press, 2006. günter neumann. a uniform computational model for natural language parsing and generation. saarbrücken dissertations in computational linguistics and language technology. vol. 1, 1994. günter neumann. interleaving natural language parsing and generation through uniform processing. in: artificial intelligence. vol. 99: 121-163, 1998. martin pickering and simon garrod. toward a mechanistic psychology of dialogue. behavioral and brain sciences, 27:169–226, 2004. massimo poesio. a model of conversation processing based on micro conversational events. in: proceedings of the annual meeting of the cognitive science society, pittsburg, 1995. massimo poesio and hannes rieser. completions, coordination, and alignment in dialogue. in: dialogue and discourse v. 1, n. 1, 1-89, 2010. massimo poesio and david traum. conversational actions and discourse situations. in: computational intelligence, 13(3), 309-347, 1997. kristina poncin, hannes rieser. multi-speaker utterances and coordination in task-oriented dialogue. in: journal of pragmatics 38, 718-744, 2006. revised and extended version of rieser and skuplik 2000. matthew purver and ruth kempson. incrementality, alignment and shared utterances. in: jonathan ginzburg and enric vallduví (eds.), proceedings of the 8th semdial workshop on the semantics and pragmatics of dialogue (catalog), 85-92, barcelona, spain, july 2004. keith rayner. the thirty fifth sir frederick bartlett lecture: eye movements and attention during reading, scene perception, and visual search. in: quarterly journal of experimental psychology, 62, 1457-1506, 2009. hannes rieser, kristina skuplik. multi-speaker utterances and coordination in task-oriented dialogue. in: massimo poesio, and david traum (eds.), proceedings of götalog. fourth rieser and schlangen 10 workshop on the semantics and pragmatics of dialogue, göteborg university, 2000, 143151. antonio roque and david traum. degrees of grounding based on evidence of understanding. in: david schlangen and beth a. hockey (eds.), the 9th sigdial workshop on discourse and dialogue (sigdial 2008). columbus, ohio, 2008. harvey sacks, emanuel a. schegloff, and gail jefferson. a simplest systematics for the organization of turn-taking for conversation. in: language, 50(4): 696-735, 1974. emanuel a. schegloff. the relevance of repair to syntax-for-conversation. in givon, t. (ed.), syntax and semantics vol. 12: 261-286, 1979. emanuel a. schegloff. discourse as an interactional achievement: some uses of ‘uh huh’ and other things that come between sentences. in deborah tannen (ed.), analyzing discourse: text and talk. georgetown university press, 7l-93, l982. gabriel skantze and david schlangen. incremental dialogue processing in a micro-domain. in: proceedings of the 12th conference of the european chapter of the acl (eacl 2009): 745-753, athens, greece, april 2009. kristina skuplik. satzkooperationen. definition und empirische untersuchung. sfb “situierte künstliche kommunikatoren”. report 1999/03. universität bielefeld. john c. trueswell & michael k. tanenhaus (eds.).processing world-situated language: bridging the language-as-action and language-as-product traditions. cambridge, mass: mit press, 2005. jos j. a.,van berkum, colin m., brown, & peter hagoort. early referential context effects in sentence processing: evidence from event-related brain potentials. in: journal of memory and language, 41(2), 147-182, 1999. mats wirén. minimal change and bounded incremental parsing. in: proceedings of coling '94 proceedings of the 15th conference on computational linguistics volume 1, 461-467, 1994. mats wirén. studies in incremental natural language analysis. linköping studies in science and technology, dissertation 292, department of computer and information science. linköping university, linköping, sweden, 1992. dialogue & discourse 10(1) (2019) 34–86 doi: 10.5087/dad.2019.103 a narrative sentence planner and structurer for domain independent, parameterizable storytelling stephanie m. lukin∗ stephanie.m.lukin.civ@mail.mil u.s. army research laboratory los angeles, ca marilyn a. walker mawalker@ucsc.edu natural language and dialogue systems lab university of california, santa cruz, ca editor: vera demberg submitted 07/2017; accepted 04/2019; published online 05/2019 abstract storytelling is an integral part of daily life and a key part of how we share information and connect with others. the ability to use natural language generation (nlg) to produce stories that are tailored and adapted to the individual reader could have large impact in many different applications. however, one reason that this has not become a reality to date is the nlg story gap, a disconnect between the plan-type representations that story generation engines produce, and the linguistic representations needed by nlg engines. here we describe fabula tales, a storytelling system supporting both story generation and nlg. with manual annotation of texts from existing stories using an intuitive user interface, fabula tales automatically extracts the underlying story representation and its accompanying syntactically grounded representation. narratological and sentence planning parameters are applied to these structures to generate different versions of the story. we show how our storytelling system can alter the story at the sentence level, as well as the discourse level. we also show that our approach can be applied to different kinds of stories by testing our approach on both aesop’s fables and first-person blogs posted on social media. the content and genre of such stories varies widely, supporting our claim that our approach is general and domain independent. we then conduct several user studies to evaluate the generated story variations and show that fabula tales’ automatically produced variations are perceived as more immediate, interesting, and correct, and are preferred to a baseline generation system that does not use narrative parameters. keywords: natural language generation, personalized storytelling, sentence planning 1. introduction storytelling is an integral part of daily life and how we share information and connect with others. people often structure observed events into a story (bruner, 1991; mcadams et al., 2006; gerrig, 1993), so that an average day at work may later be described as a narrative where the events are exaggerated to revolve around the individual, rather than simply listing events that took place. in these natural story settings, stories may be told many times to different audiences but rarely told in the same way twice. a storyteller may explore different interpretations of the same incident from multiple points of view (mateas, 2001), or use a richer style when telling a story to highly ∗. this work was done at the university of california, santa cruz. c⃝2019 stephanie m. lukin and marilyn a. walker this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). a narrative sentence planner and structurer for storytelling startled squirrel the fox and the crow we keep a large stainless steel bowl of water outside on the back deck for benjamin to drink out of when he’s playing outside. the craziest squirrel just came byhe was literally jumping in fright at what i believe was his own reflection in the bowl. he was startled so much at one point that he leap in the air and fell off the deck. but not quite, i saw his one little paw hanging on! after a moment or two his paw slipped and he tumbled down a few feet. but oh, if you could have seen the look on his startled face and how he jumped back each time he caught his reflection in the bowl! a crow was sitting on a branch of a tree with a piece of cheese in her beak when a fox observed her and set his wits to work to discover some way of getting the cheese. coming and standing under the tree he looked up and said, “what a noble bird i see above me! her beauty is without equal, the hue of her plumage exquisite. if only her voice is as sweet as her looks are fair, she ought without doubt to be queen of the birds.” the crow was hugely flattered by this, and just to show the fox that she could sing she gave a loud caw. down came the cheese,of course, and the fox, snatching it up, said, “you have a voice, madam, i see: what you want is wits.” table 1: startled squirrel and aesop’s the fox and the crow interactive and responsive addressees (thorne, 1987). when young adults describe situations in which their lives were threatened, they use different telling styles to convey different messages to the audience, such as empathy for others, preoccupation with one’s own fear or sadness, or one’s courage or bravery (thorne and mclean, 2003). retelling capabilities are showcased in exercises in style, where a sequence of simple events are told in 99 different ways (queneau and wright, 1981). madden (2006) repeats the exercise in visual storytelling, creating different visual depictions of the same events in a story. a computational treatment of storytelling should support the ability to retell stories in different ways, mimicking how human storytellers tailor their stories to the context and to their audience. consider the personal narrative startled squirrel in table 1 from the spinn3r corpus of blogs (burton et al., 2009). this telling is in the first person, using the narrator’s own voice. the narrator tells about a time when they saw a curious squirrel in their backyard, who tried to drink out of a dog’s water bowl. upon getting closer to the bowl, the squirrel jumped at its own reflection and fell off the deck. this story primarily evokes a humorous response, however a different telling could instead evoke empathy for the squirrel if the story were told from the perspective of the squirrel itself, as depicted in the variation that our system can automatically produce, in table 2. similarly, aesop’s fable the fox and the crow (table 1), which is traditionally told from a third person perspective, could be framed from either the fox or the crow’s perspective, affecting the reader’s insight into each character’s thoughts, as in the automatically produced variation in table 2. startled squirrel variation excerpt the fox and the crow variation excerpt i approached the bowl. i was startled because i saw my reflection. because i was startled, i leaped. i fell over the deck’s railing with my paw. my paw slipped off the deck’s railing. the crow sat on the branch of the tree. the cheese was in the beak of the crow. i observed the crow. i thought “i will obtain the cheese from the crow’s beak!” table 2: computational variations for startled squirrel and aesop’s the fox and the crow in order to computationally retell stories in different ways, the story representation must distinguish the content of the story from the telling. this distinction is classic in narratology and categorized as fabula and sujet (propp, 1969). the fabula is comprised of the events in a story, represented abstractly as a set of building blocks that can be rearranged and from which more complex narrative forms can be built (abbott, 2008). the fabula includes all the abstract components of the story world, including the characters, their goals, and the actions that take place in the story world. 35 lukin and walker on the other hand, the specific telling and framing of a subset of these events is the sujet (propp, 1969). constructing the sujet may include reordering or otherwise manipulating the presentation of events, choosing between a subjective or objective interpretation, the perspective from which the story is told, and the character voice, among other variations. figure 1: differences between story generation and natural language generation a computational storyteller requires a general representation of the fabula and a way to generate different sujet from a single fabula. one main limitation of work in this area to date, which has hampered progress in the field, is what is known as the natural language generation (nlg) story gap, illustrated in figure 1 (lönneker, 2005; callaway and lester, 2002). the left-hand side of the figure is meant to depict the considerable line of research invested in the automatic production of different fabula, given a particular pool of content (peinado and gervás, 2006; riedl and young, 2010; gervás et al., 2005) inter alia. this line of work has examined, for example, how plan-based approaches for selecting and ordering the content in different ways, e.g. selecting different events, leaving events out, or re-ordering events can have differential effects on the reader, such as enhancing the reader’s feelings of suspense or surprise (bae and young, 2009; ware and young, 2011; niehaus and young, 2009). the nlg story gap refers to the fact that these story engines produce plan-like content representations, which unfortunately do not provide the information that is needed in order to render that content textually. this line of work often adopts a simple rendering strategy, defining templates by hand that directly realize each component of the fabula, as shown in the bottom left-hand-side of figure 1. the result is that the only variations of the sujet that are possible in this approach are those that have to do with content selection, or those that are explicitly hand-crafted as template variations. in contrast, the right-hand side of figure 1 shows the typical input to an nlg engine and the standard modular architecture that nlg uses to generate different textual renderings of their input. it should be possible, in principle, to generate different sujet by building on previous work in nlg, which provides an abundance of techniques for generating different textual variations from a meaning representation. however, nlg architectures assume that the input to the nlg is a text plan, whose leaves are syntactically and semantically grounded in linguistic representations. these linguistic representations are needed in order to apply sentence planning operations and produce the many different possible variations in texts for a fixed meaning representation, as we explain in more detail in section 3. these representations are not compatible with story planning as they only contain information about a single sentence. thus the nlg story gap arises as the gap between the plan-like meaning representations used by story generators, and the input assumptions of nlg engines. current practice is to fill this gap by hand, either by writing templates as shown in the left-hand side of figure 1, or by constructing an nlg dictionary, which maps from each story meaning component to a linguistic form, such as a dependency tree, which nlg sentence planning operations such as aggregation and discourse structuring, can operate on (mairesse and walker, 2011; callaway and lester, 2002; penning and 36 a narrative sentence planner and structurer for storytelling theune, 2007). for example, callaway and lester (2002) mapped the story elements for little red riding hood into a syntactically grounded representation by hand, which then supported their work on the automatic generation of narrative variations. similarly, work by theune et al. (2007) on a system called the narrator describes in detail how the story consists of causally related semantic story elements, and how each story element is mapped by hand to a dependency tree. another approach to bridging the nlg story gap recently proposed by concepción et al. (2016b) suggests that the use of a controlled language (controlled natural language) for specifying story content could make it simpler to map story plan structures to syntactic patterns in a very general way. these syntactic patterns could in principle then be converted into the linguistic representations needed by different nlg engines (concepción et al., 2016a,b,c). in sum, to date, bridging the nlg story gap has required a considerable amount of hand-crafting for each story that a storytelling system would want to tell, and this has prevented narrative systems from generating rich and diverse variations over a variety of different topics and genres. our approach to this problem consists of three separate contributions that together form the fabula tales storytelling system: 1. we propose a particular take on bridging the nlg story gap that takes existing stories from different genres, creates a fabula representation using scheherazade, an easy-to-use annotation tool, and then automatically maps the fabula to a general nlg representation; 2. we develop an nlg engine with narratologically inspired sentence planning parameters and show how we can generate different tellings of stories using a narratological structurer that employs these parameters for any existing story that has been annotated; 3. we evaluate different generated tellings at both the sentence and the story level for different evaluation criteria. bridging the nlg story gap. in order to bridge the nlg story gap, we first require the preservation of content representing the fabula or semantics of a narrative, and the creation of linguistic representations that can be used to generate tellings, sujet or syntactics (section 3). we use existing stories, in their textual form, that come from different genres and different topics, as exemplified in table 1. our story corpus selection allows us to explore the domainindependence of our approach: the corpora consists of 36 aesop’s fables and 108 first person socialmedia blogs from the spinn3r corpus (burton et al., 2009), two radically different genres. we adopt elson’s story intention graph (sig) (elson, 2012a), as the representation of fabula. in addition to its strong theoretical motivation, one advantage of using the sig as a fabula representation is that it can be produced in a lightweight way (only one to two hours per story) using a corresponding annotation tool called scheherazade that supports the production of fabula by annotation of texts (elson and mckeown, 2009). we then create a general model that maps from the sig to novel lexical-semantic story trees (lsstrees) containing linguistic representations needed for nlg, building on our prior work (rishes et al., 2013). this highlights a second advantage of the sig representation: annotation using scheherazade maps each predicate and constant in the story’s logical representation to a lexical item from the off-the-shelf lexical resources, verbnet (kipper et al., 2006) and wordnet (fellbaum, 2010). this grounding to lexical items (with their subcategorization frames) allows us to create the general mapping model compatible with all story domains. we show that the mapping produces good quality lexical-semantic story trees and generates good baseline stories, without 37 lukin and walker expert handcrafting. generating stories with narrative variations. after showing that we can develop a general representation of fabula for any story domain, our second contribution is to generate different sujet by implementing narratological sentence planning and a set of discourse relations (section 4), and a story-level narratological structurer that sits on top of the lexical-semantic story trees (section 5). the lexical-semantic story trees are manipulated by fabula tales’ narrative sentence planner, based on the architecture of the personage expressive natural language generation engine (mairesse and walker, 2008, 2011), with parameters inspired by theories of narratology (genette and lewin, 1983; prince, 1974; lönneker, 2005; bal, 1997). our narrative sentence planner supports changing point of view (first or third), inserting direct speech acts, and supplementing character voice using operations for lexical selection, discourse structuring, and pragmatic marker insertion. we develop a narratological planner for fabula tales that operates above the narrative sentence planner. our narratological planner is not a narrative content planner; thus our approach assumes that all content from the story tree will be told, and the planner determines which narrative parameters should be applied to generate the story. the narratological structurer determines variations at the story-level by focusing on the entire flow of the story, rather than just at the sentence level. training data is obtained by overgenerating different variations of sentences on a sentence-by-sentence basis. the sentences are then ranked by subjects using a novel create-your-own-story annotation paradigm, to learn the impact of each narrative parameter. evaluation of narrative theories on generated stories. by combining these parameters, the generated variations evoke diverse framing and voice alterations at both the sentence and story level. we explore how different narrative parameters lead to different perceptions of the story, evaluating on holistic narrative metrics of immediacy, interest, correctness, and preference. we use this experimental data as input for a classification experiment where we rank possible choices that the generator can make in terms of which sentences are the most impactful. we conclude in section 6 where we discuss limitations, future work, and applications. 2. related work 2.1 narrative variation in storytelling systems the storyteller has many devices at their disposal to frame stories, including changing the overall tone, mood and effect of the story to distinguish between “who sees?” and “who speaks?” (genette and lewin, 1983). theories of narratology provide a number of narrative devices or parameters to produce diverse framings of a fabula (bal, 1997; lönneker, 2005; genette and lewin, 1983; prince, 1974). lönneker (2005) categorizes parameters into three broad categories of time, mood, and voice, each with sub-categories, as detailed in table 3. narrative variations to time and mood primarily involve forms of narrative content planning, that is, determining or generating the events and the structure in which to tell (fabula). the voice parameters influence the realization (sujet). we study the nlg story gap and a systems’ ability to generate these diverse tellings. we expand the gap presented in figure 1, positing that a storytelling architecture has the potential for four gaps to arise, as we depict in figure 2. the first gap occurs when story content for the storytelling pipeline is difficult or time consuming to create. closely linked to the first gap, the second gap 38 a narrative sentence planner and structurer for storytelling parameter explanation time: order sequence in which events are told, in comparison with the sequence in which they “actually happened”. in synchrony, the event sequence in discourse corresponds to the sequence of the story. anachronies can take the form of flashbacks (retrospectives) or flashforwards (anticipations). time: speed relation between story time and discourse time. congruence exists probably only in single scenes; otherwise timelapses (accelerations), time jumps (ellipsis), time expansions (decelerations), or pauses are used to achieve different degrees of explicitness and emphasis. time: frequency relation between the number of times a (similar) event happened, and the number of times an event is told. the following realizations are distinguished: singulative (oneto-one relation), repetitive (“recount several times what happened once”), and iterative (“recount once what happened several times”). mood: distance combination of amount of information conveyed and narrator intrusion. stereotypically, detailed information and low narrator participation indicate imitation or “direct” dramatic mode, as opposed to a “distant”, mediated narrative mode. this parameter also affects the way in which speech is reproduced. mood: focalization accessibility of knowledge needed to select story events for presentation in discourse. if a narrative instance disposes of unrestricted knowledge of the story world, it uses external focalization; if the knowledge is restricted to a character’s field of perception, focalization is internal. mood: point of view spatial, temporal, and ideological points of view from which events are described. events can be described from the point of view of different characters. this parameter covers more aspects than focalization. voice: time time relation of the narrating action to the story event. events can be told while they are happening (concurrently), retrospectively, or prospectively. voice: person narrator participation. a homodiegetic narrative instance is a character of the current narration (grammatical realization typically in the first person), while a heterodiegetic narrative instance is “absent” from the current narrative and not referred to. in a secondperson narrative, the protagonist is the reader. table 3: narratological parameters presented in lönneker (2005) occurs during the process of applying a transformation to the content in order to make it compatible with the storytelling system. below, we discuss story planners that require dependency information between plot points but lack accessible tools for their creation. we also review nlg engines that require detailed syntactic structures as input. the result is that new stories from different genres are more difficult to create. the work we present in this article does not face these two gaps because the intermediary story representation we use, the sig, is easily obtainable from any genre of natural, unstructured story texts using its accompanying creation tool, scheherazade. in section 3, we describe how this intuitive tool has a low authorial burden for creating new content. the third gap occurs when neither the content pool (the selected fabula) nor its representation are rich enough for story planning. this tends to occur in systems that make use of syntactic representations which are ripe for narrative manipulation, but do not necessarily preserve story-level information, making it difficult to alter the story when nothing is known beyond a single sentence, as we discuss below. in our work, we develop a novel representation, lexical-syntactic story trees (lsstrees), that preserve story-level discourse information across story predicates (section 3). 39 lukin and walker the lsstrees are compatible with content planning as well as syntactic transformations in sentence planning. figure 2: gaps between story representation, story generation, and nlg finally, the fourth gap precludes variation of the sujet due to a lack of syntactic information (rather than semantic information in the third gap); a system is therefore unable to dynamically alter the telling once a story point has been selected, instead making use of templates or fixed-strings. the remainder of this section discusses related work in computational storytelling within the scope of the narrative parameters described above and limitations to the work with respect to the four elements of the nlg story gap presented here. as stated above, fabula tales does not perform any temporal content planning (i.e., time:order, time:speed, time:frequency) and assumes every story point will be selected and told in the order in which it occurred. there are many approaches to story content selection, a number of which select content to prioritize author-level goals. in these systems, broad narrative goals are defined by hand prior to story generation, and then the assertions of the story world are reasoned over to assign the time of the narrative, using, for instance, hierarchical task networks (lebowitz, 1983), case-based reasoners (peinado and gervás, 2006; turner, 1993; gervás et al., 2005), or bipartite or tripartite optimizations (winer and young, 2016; barot et al., 2015). other systems take a character-driven approach to story content planning and select content based on character knowledge, goals, and their relationships with others in the story world (meehan, 1977; riedl and young, 2010; theune et al., 2004). stories are generated when the preconditions of author or character goals are met. despite control over the story generation, these systems place less priority on the realization of the texts, and are instead typically use pre-authored pieces of story text or templates. mood:focalization is a narrative device that has been explored using planning engines that identify which events or inner thoughts characters are aware of at the moment, including character’s beliefs, desires, hidden intent, or inner states of mind. these affect the selection of the events the planner selects to tell; for example, stories with surprise endings are generated by exploiting the disparity of knowledge between a story’s reader and its characters (bae and young, 2009; bae et al., 2011) while others create conflict (ware and young, 2011, 2012; ware et al., 2014; niehaus and young, 2009). curveship enables similar flexibility of story framing of focalizations or temporal orders, or the speed of the narrative (time:speed) (montfort, 2007, 2009), although curveship is 40 a narrative sentence planner and structurer for storytelling only applicable to the interactive fiction domain, and similar to the work described above, only offers a templatic realization strategy. on the opposite side of the nlg story gap are prose generation systems that prioritize diverse sujet generation, typically effecting the mood:distance and voice:person of the resulting narratives. these works use rich and flexible syntactic representations that can be manipulated. some approaches include the use of document plans and dependency trees (theune et al., 2007), or annotated source material (i.e., story texts) as controlled natural language, which restricts grammar and vocabulary to afford flexibility for sentence planning (concepción et al., 2016a,b,c), or a narrative plan compatible with the off-the-shelf generator fuf-surge (callaway and lester, 2002; elhadad and robin, 1996). shortcomings of these carefully crafted approaches tend to be the limited domain representations and the time constraint to manually create these representations, e.g., callaway and lester (2002) can only generate stories in the little red riding hood domain, and both approaches require the manual construction of a formalism representing the characters, story assertions, and various parameters, whereas our approach offers an intuitive user interface for defining these story requirements. recent work performs both story planning and diverse text generation using data-driven approaches that make use of large corpora of unstructured texts, rather than carefully curated content or syntactic representations as input. one interactive approach takes turns with the user to co-construct a story using data from the spinn3r blog corpus or from movie scripts (swanson and gordon, 2008; munishkina et al., 2013). after the user types the next sentence of a story, the algorithms search for similar sentences from stories in the corpus using term frequency-inverse document frequency (tf-idf) or other search criteria. returning the next sentence from the selected story to the user progresses the co-constructed narrative forward. the scheherazade story generation system (not to be confused with the scheherazade annotation tool for sigs), learns causal graphs from texts obtained by crowd-sourced workers prompted to write a short story about a particular topic (li et al., 2013; li, 2015). the planner performs well because the texts are constructed from a prompt and assume a causal and event-centric structure. it remains to be seen how this approach would apply to narratives collected “in the wild” such as in the spinn3r corpus, a large portion of which contain orientation and evaluation segments prevalent in oral narratives (labov and waletzky, 1997; rahimtoroghi et al., 2013, 2014). these textual learning approaches are advantageous for domain independence and can retrieve different styles of prose by mining a corpus. however, because these algorithms are retrieval-based, when they find a matching response in the corpus, the algorithm returns the response as-is without varying the selected text. therefore, these systems’ expressivity with respect to language generation are restricted to the original narrative text or script. other joint story planning and realization systems utilize recurrent neural networks or longshort term memory convolutional sequence-to-sequence neural networks (roemmele, 2018; fan et al., 2018; peng et al., 2018). these approaches similarly enable the learning of many topics and genres and have the additional advantage of generating text learned from the corpora. yet to date, these models do not attempt to diversify the narrative texts according to any theories of narrative, but rather posit that variations can be learned inherently from the datasets. these approaches construct a text word-by-word, sentence-by-sentence, but these texts do not model the overall narrative scope of the fabula, or offer diverse sentence planning for generating different sujet. to date, these neuralbased approaches do not afford these storytelling aspects that more traditional approaches have done in the past. 41 lukin and walker tasks for evaluating the consistency of stories show promise for someday being used to evaluate causality in automatically generated stories. hu et al. (2013) and rahimtoroghi et al. (2016) learn event pairs from the spinn3r corpus using unsupervised approaches and causal potential. another evaluation, the corpus of plausible alternatives (copa), is created from hand-annotated causality pairs from spinn3r, and the task is to select the most likely event to occur next in the story (roemmele et al., 2011). mostafazadeh et al. (2017) creates a synthetic dataset to measure the predictability of subsequent events, but unfortunately, a bias was identified in the dataset creation; as a result, simple natural language processing tricks can obtain a high score on the task, rather than examining the content itself (srinivasan et al., 2018; sharma et al., 2018). a new task of visual storytelling has been driven forward by improvements to computer vision algorithms. this form of storytelling faces the same challenges as text-based story generation, but an additional challenge is that the fabula must be determined from computer vision algorithms. huang et al. (2016) create a corpus for visual storytelling (vist), yet the guidelines for the collection effort do not explicitly take into consideration the entire story as a whole, nor do they explore diverse narrative variations; the stories human annotators create about a sequence of images tend to be shallow and action-oriented. the visual storytelling challenge1 is the first shared task in visual storytelling, and they, as well as wang et al. (2018), offer several subjective evaluation metrics that are applicable for text-based storytelling as well (e.g., “focus” and “expressiveness”). lukin et al. (2018) poses challenging questions for the visual storytelling task in order to bridge this modern task to its roots in narrative theories, including specifying that visual story generation systems must be flexible in both content planning and realization of different narrative goals and be able to adapt the narrative to the audience. 2.2 foundations of the fabula tales storyteller fabula tales’ automatic syntactic translation from sig to lsstree builds upon previous work first described in rishes et al. (2013). however, that work only explored a single domain, aesop’s fables, whereas the model presented in this article has been improved and tested on an additional 108 stories from the personal blog domain. new evaluations are presented that measure the quality of the baseline translation algorithm in terms of text similarity, semantic text similarity, fluency, and grammaticality. furthermore, the system presented here models semantic and discourse relations between plot elements, which the original method did not model. this article reviews the syntactic translation process alongside the new contributions of this article in order to present together the complete storytelling pipeline. fabula tales’ narrative sentence planner implements similar parameters and linguistic representations as the personage expressive nlg engine (mairesse and walker, 2008, 2011). personage manipulates deep syntactic structures (dsynts) (lavoie and rambow, 1997) according to parameterized models grounded in the big five personality traits, providing a large range of pragmatic and stylistic variations of a single utterance. in personage, the style to be conveyed is controlled by a model that specifies values for different stylistic parameters (such as verbosity, syntactic complexity, and lexical choice). personage requires hand crafted text plans and dsynts, limiting not only the expressiveness of the generations, but also the domain. personage has been used as a way to help authors to reduce the authorial burden of writing dialogue instead of relying on scriptwriters for games (reed et al., 2011), but still relies on hand-authoring dsynts. fabula tales introduces 1. http://www.visionandlanguage.net/workshop2018/#challenge 42 http://www.visionandlanguage.net/workshop2018/#challenge a narrative sentence planner and structurer for storytelling the first tool for automatically creating dsynts, which allows for personage’s sentence planning to be repurposed for the narrative space. another source of linguistic variation supported by personage and reimplemented by fabula tales is splitting and aggregating sentences at the discourse level. aggregation operations help to avoid repetition and produce more coherent, concise, and context aware output (cahill et al., 2001; scott and de souza, 1990; paris and scott, 1994). several nlg systems use rhetorical structure theory (rst) (mann and thompson, 1988) for aggregation and sentence planning (walker et al., 2007; howcroft et al., 2013). aggregation is inclusive of aspects of abstractive text summarization, with parallels in sentence compression, fusion, lexical paraphrasing, and reorganization by reducing syntactic structures through the removal of articles (grefenstette, 1998), manipulation over syntactic trees (knight and marcu, 2000), dependency trees (filippova and strube, 2008), or grammar (riezler et al., 2003) to name a few. not all text summarization operates over syntactic units, instead employing text-to-text generation (chandrasekar and srinivas, 1997; knight and marcu, 2002; marsi and krahmer, 2005). however, recent work introduces a task based on starting with small meaning-representations and recombining them in different ways (narayan et al., 2017), a step towards joining the advances and contributions of aggregation and text summarization methodologies. our method for training our narratological structurer is based on previous work on overgenerate and rank. nlg systems have made use of the overgenerate and rank methodology in which a variety of sentence variations are generated from a model that first overgenerates, and then ranks the generated output based on some measure of “goodness” relevant to the task. previous work has used statistical models as an objective scoring function to measure a set of generated candidate utterances based on a variety of features. simple ranking based on n-grams are used to generate from abstract meaning representations (langkilde and knight, 1998; langkilde-geary, 2002) and to generate dialogues with alignment and personality cues (isard et al., 2006). lexical and conceptual similarity scores influence near-synonym selection (inkpen and hirst, 2004) and correctness and grammaticality scores influence sentence generation (gardent and kruszewski, 2012). the overgenerate and rank methodology has been shown to be useful for error mining with human judgment as a scoring metric and for future parameter adjustments and feedback (walker et al., 2002, 2007; mairesse and walker, 2010b; walker et al., 2013; gardent and kruszewski, 2012). in the narrative space, overgenerate and rank has been used to combine a rule-based overgeneration phase with a statistical ranking phase by probabilistic parsing to rank sentences in a story for its naturalness (ahn et al., 2016), yet this work does not employ any narrative specific sentence planning; the highest ranked variant for each sentence in the story is simply concatenated to the other selected sentences. 3. bridging the nlg story gap this section describes how fabula tales bridges the nlg story gap. as fabula, we use existing stories about different topics and from different genres. the strength of our approach is that it allows us to test whether our methods for generating different sujet can be applied across many different fabula. a limitation of our approach is that we only have a fixed set of story points within each fabula, in contrast to work on story generation systems whose focus has been to explore story variations that result from manipulation and selection of events from the fabula (bae et al., 2011; riedl and young, 2004; gervás et al., 2006). fabula tales’ bridge is depicted in figure 3. it begins with the raw text from an existing story, and produces a representation upon which the narrative sentence planner and narratologi43 lukin and walker figure 3: fabula tales translation from text, to intermediary story intention graph, to lexicalsemantic story trees cal structurer operate. as the figure shows, the story content is first manually annotated using the scheherazade annotation tool, in order to produce an intermediate representation of the fabula as a story intention graph (sig) (elson, 2012a). the sig represents a story along a variety of dimensions, including plans, goals, and actions of characters. this formalism emphasizes the key elements of a narrative rather than attempting to model the entire semantic world of the story. the sig representation is then automatically translated to lexical-semantic story trees (lsstrees). lsstrees consist of a syntactic and semantic component. the linguistic representation can be directly converted to deep syntactic structures (dsynts) in order to vary sentences at the syntactic level using the off-the-shelf surface realizer realpro (lavoie and rambow, 1997; mel’čuk, 1988). the semantics are captured through text plans, and model three discourse relations: contingency, temporal order, and attribution. section 3.1 describes the sig formalism. the creation of sigs requires no parsing or preprocessing of the input; the story text is manually annotated using a tool called scheherazade (elson and mckeown, 2009). scheherzade is a user-friendly annotation tool that is generalizable and does not require specific domain or knowledge or deep linguistic knowledge or a generation dictionary. word sense disambiguation is done as part of the annotation process: all of the predicate-argument structures representing the structure of the story and the character intentions are lexically grounded in either wordnet synsets (fellbaum, 2010) or verbnet verb nodes (kipper et al., 2006). the dramabank, an existing corpus of aesop’s fables annotated as sigs (elson, 2012b), was first used to explore our translation pipeline. we then test the domain and genre independence of the pipeline by using a second corpus of first person informal blogs, personabank, that we created in prior work (lukin et al., 2016). because it is used extensively here in our experiments, we describe personabank and summarize the sig annotation process. previous work, as well as our own, has shown that the scheherazade annotation tool can be used by non-expert annotators and does not require a background in linguistics or computer science, nor are annotators required to be domain experts, whereas other work, as we described in the previous section, may require careful handauthoring of syntactic structures or detailed domain knowledge. the story generation process is streamlined after a sig is created. discourse relations annotated in the sig are used to construct text plans in the lsstrees, and a syntactic representation of each story point is automatically translated from the sig. section 3.2 shows how this semantic 44 a narrative sentence planner and structurer for storytelling and syntactic translation process is fully automated and does not require construction of semantic or syntactic structures by hand or by using templates. this final representation allows the narrative sentence planner and narratological structurer to generate different narrative variations of a story. we show how following the pipeline in figure 3 generates a baseline text without narrative variations, which we use in subsequent evaluations, and present evaluations verifying the fidelity of the translation process using this baseline text (section 3.4). 3.1 story intention graphs: intermediary fabula the sig formalism is a computational model of narrative that goes beyond the surface form of a text, as opposed to style. sigs separate “what the story is [fabula]” from “how the story is told [sujet]” (elson, 2012a). these encode a single sujet from which the underlying fabula is derived. in contrast, the primary goal of this work is to transform a single fabula into many sujet. as is the nature of the sujet, a telling is only one interpretation or rendering of a larger narrative discourse. for a particular sig, some events may not have been made explicit in the original sujet, and thus are excluded from the derived fabula. however, the assumptions and inferences that a reader makes to interpret the story can be added to the sig in one of its deeper semantic layers, as we explain below. the first step in annotating a story as a sig is to define all of the characters and props. these are given unique ids, and are defined by identifying a wordnet synset that is of the right type for that character or prop, e.g. a fox or a tree. then, the events of the story timeline are defined and their propositional representations use these character and prop entities as arguments. a story in the sig formalism is represented by four layers as in the example sig for the fox and the crow, shown in figure 4a: the sujet or textual layer, the timeline, the interpretive layer, and the affectual layer. the nodes in each layer are connected by arcs signifying semantic or discourse relationships between the nodes, within or across layers. the original story, the sujet (first column in figure 4a) is first divided by the annotator into textual segments, where each segment, in the annotator’s view, represents a distinct, coherent story point. (a) a story intention graph for the fox and the crow (b) a story intention graph for startled squirrel the other three layers of the sig comprise different elements of the fabula, each layers’ nodes derived from verbnet frames with wordnet story elements. the wordnet and verbnet senses are utilized in the lsstree creation process to build a generation dictionary for downstream word sense disambiguation and co-reference resolution upon text realization. the timeline layer summarizes the actions and events that occur. the interpretation layer captures story meaning 45 lukin and walker derived from agent-specific plans, goals, attempts, outcomes and affectual impacts, and the annotators’ interpretation of why characters were motivated to take the actions they did, adopting a “theory of mind” approach to modeling narratives (palmer, 2007). the final dimension is the affectual layer, representing deeper motivations underlying character goals and the effect these goals have on the characters. there are 12 basic types of affect, including health, ego, wealth as described in more detail in elson and mckeown (2009). the fact that the scheherazade annotation tool can be used by non-expert annotators to easily create new sig story encodings was first demonstrated by elson’s work on the creation of the dramabank corpus of sigs (elson, 2012b). our work uses a subset of dramabank consisting of all 36 aesop’s fables, such as the example the fox and the crow shown in table 1, and other wellknown stories like the boy who cried wolf and the fox and the grapes. a simplified sig for the fox and the crow is shown in figure 4a. numbers indicate timeline or interpretation events, and letters label the affect nodes. the sig specifies that the fox’s goal (#1) is to obtain the cheese from the crow. this would provide for his health, (a), represented as an affect node. when the fox sets his wits to discover some way of getting the cheese this is encoded by the annotator as the fox tries to discover how to obtain the cheese (#2) which is interpreted as (ia) his goal (#1). a precondition arc is created to restrict that the goal of obtaining the cheese can only be initialized if the fox has first seen the crow (#4). the fox also has a goal (#5) that the crow will sing. by flattering the crow (#3) the fox attempts to cause (achieve) that the crow will sing. if the crow caws (#6), this would actualize the goal of the crow singing. the singing itself provides for the crow’s ego, (b), represented as an affect node. when the fox says what you want is wits this is encoded by the annotator as the fox said the crow needed wits (#7) which is interpreted as the fox insulting the crow (#8). this damages the crow’s ego (b). our construction of the personabank corpus is a second demonstration that scheherazade can be used to create sigs for stories with a variety of author styles and topics. personabank is a corpus of 108 sigs for blog stories from the spinn3r corpus, which includes the startled squirrel story from table 1 (lukin et al., 2016). the sig for the startled squirrel is shown in figure 4b. again, numbers indicate timeline or interpretation events, and letters label the affect nodes. the narrator places a bowl on the deck (#1) as an attempt to cause the goal of the narrator to give the dog some water (#2) which would provide for the dogs’ health (a). then the squirrel approaches the bowl (#3) as an attempt to cause (achieve) the squirrel’s goal to drink the water (#4) which would provide for the squirrel’s health (b). when the squirrel is startled (#5), this attempts to prevent (blocks) the goal of drinking the water, and when the squirrel falls (#6) this both ceases the goal (#4) and damages the squirrel’s health (b). scheherazade provides a built-in generation module as part of the annotation process so that the annotator can see a realization of the underlying representation in real time as they annotate in order to verify that the underlying representation being constructed is what the annotator intends (bouayad-agha et al., 1998; elson and mckeown, 2009). we will call this the scheherazade realization. table 4 shows the original story and the scheherazade realization for the startled squirrel. scheherazade uses templates and lexical realizations from the wordnet nouns and verbnet frames to directly realize the underlying sig semantics, without attempting to produce any type of variation in its realizations. instead, it produces text in a fixed way for the selected encoding, e.g., the sig semantics “approach(squirrel, bowl), cautiously” will always be realized as the squirrel cautiously approached the bowl. furthermore, there are odd and redundant phrasings in the scheherazade realization because of its templates, for example, the second squirrel leaped 46 a narrative sentence planner and structurer for storytelling startled squirrel scheherazade realization we keep a large stainless steel bowl of water outside on the back deck for benjamin to drink out of when he’s playing outside. the craziest squirrel just came byhe was literally jumping in fright at what i believe was his own reflection in the bowl. he was startled so much at one point that he leap in the air and fell off the deck. but not quite, i saw his one little paw hanging on! after a moment or two his paw slipped and he tumbled down a few feet. but oh, if you could have seen the look on his startled face and how he jumped back each time he caught his reflection in the bowl! a narrator placed a steely and large bowl on a back deck in order for a dog to drink the water of the bowl. a squirrel approached the bowl. the squirrel began to be startled because it saw the reflection of the squirrel. the squirrel leaped because it was startled and fell over the railing of the deck and because it leaped. the squirrel held the railing of the deck with a paw of the squirrel. the squirrel fell, and the paw of the squirrel slipped off the railing of the deck. table 4: startled squirrel and a scheherazade realization topic excerpt from original story scheherazade realization wildlife, bugs lillian found a wasp on the window at the farm. a girl named lillian found a wasp on a window of a farm. holidays, christmas, family we tied [the christmas tree] down to the roof and go get hot chocolate. the group of relatives of the narrator tied the pine tree onto the roof of the car. the group of relatives of the narrator drank some cocoa. romance, new romance i took a few pics and was blown away by the beauty of my girls. the narrator photographed the bride and the maid of honor and noticed that the bride and the maid of honor was gorgeous. family the trip started with a much anticipated but never duplicated dinner at rainforest cafe. the family of a narrator ate at a restaurant named rainforest cafe. everyday events, technology so he wanted to get another [phone] the father of the narrator wanted to acquire a new second (#2) telephone. pets, everyday events we went to a no-kill shelter to get our first cat. the husband of the narrator and the narrator went back to a humane shelter in order to adopt a cat. table 5: excerpts from personabank, illustrating scheherazade realizations because it was startled and fell over the railing of the deck and because it leaped. because the scheherazade generator focuses on semantic fidelity to the sig, we use it below as a baseline for measuring whether our translator bridge preserves story content (section 3.4). the primary motivation behind the creation of personabank was to test the use of the sig as a representation of fabula in our pipeline. in general, we selected stories that had a clear sequential timeline, and most of the selected stories are shorter than 300 words, with the minimum and maximum number of words to be 104 and 959 respectively (table 6). trained annotators2 can annotate the timeline layer of a story in about one hour. annotating the interpretive and affectual layers requires more subjective judgment and takes an additional hour for each story. as shown in table 6, all stories were annotated with the timeline layer, 21 of which were annotated with the interpretive layers. each story was annotated by a single annotator, thus the sigs represent one interpretation of the story, one possible fabula. 2. annotators of personabank were undergraduate research assistants associated with the natural language and dialogue systems lab at the university of california, santa cruz. 47 lukin and walker statistics stories total stories 108 positive stories 55 negative stories 53 interpretation layers annotated 21 avg. length (# words) 269 table 6: personabank statistics sample excerpts from personabank stories illustrating the range of topics covered and their scheherazade realizations are shown in table 5. several of these story excerpts illustrate how the first-person i is typically mapped to a character called the narrator in the sig encoding: it is not possible to use deictics like i or me, or anaphors like it, he, she. because the sig representation uses unique ids for each character in a story, our realization engine can easily replace definite references like the narrator with other ways of referring to the same character. for more details of the corpus and its creation, please see lukin et al. (2016). 3.2 lexical-semantic story tree translation figure 5: lexical-semantic story tree content assertions 1: assert(approach (squirrel, bowl)) 2: assert(saw (squirrel, reflection)) 3: assert(leap (squirrel)) 4: assert(fell (squirrel)) figure 6: content assertions from the startled squirrel we define lexical-semantic story trees (lsstrees) as structures that contain both semantic and lexical information about a particular story point from a sig’s timeline layer, including the action and actors involved in a particular story point, syntactic representation of these semantics, and text plans that detail the discourse relations that hold between propositions within the story. these structures are operated on and manipulated by the sentence planner (section 4) and narratological structurer (section 5) in order to produce narrative variations. natural language text can be directly realized from these lsstrees by a surface realizer, thus these structures provide the storytelling system with both semantic knowledge needed for the story fabula and syntactic information about the content to generate different sujet. information from each sig story point is automatically extracted and organized into lsstrees, depicted in figure 5. this translation does not use the original story text, but only the sig structures; no parsing of text is required. instead, we develop a mapping of syntactic structures that corresponds roughly to parts of speech: verbs, nouns, adjectives, and prepositional phrases. these parts of speech are defined in the sig because part of the annotation process involves lexically grounding each predicate and constant for story events in either wordnet or verbnet. these lexical resources provide both part of speech information as well as sub-categorization frames. word senses are preserved in the mapping to lsstrees in order to support word sense disambiguation and lexical choice in sentence planning. finally, each story entity (characters or props) has a unique identifier, so there is no need to perform co-reference resolution. these structures and this syntactic translation was first developed in rishes et al. (2013). 48 a narrative sentence planner and structurer for storytelling text plans that preserve a key set of discourse relations are constructed from the sig, based on the penn discourse treebank (pdtb) discourse relations (prasad et al., 2008). the annotation of a story requires story events to be ordered on the sig timeline: these relations are represented as the pdtb temporal order relation. we also represent the attribution relation between a speaker and their utterance in order to later vary whether an utterance is realized as direct or indirect speech. finally, because causality is considered a key relation in the structuring of narratives, we represent the pdtb contingency relation, an extremely common relation in personabank. a subset of content assertions from the startled squirrel are seen in figure 6. a contingency relation between assertion 1 and assertion 2 could result in the textual realization of the squirrel saw its reflection because it approached the bowl, and the temporal order relationship between assertions 3 and 4 could yield the squirrel leapt. he fell. the derivation of the discourse relations used in the text plans are described in more detail in section 4. the translation methodology was first developed on a single fable, “the fox and the grapes”, until high coverage was achieved. the model was then tested on a set of 35 additional fables from the dramabank. additional refining of the model was performed on a single story from personabank, then tested on the remaining 107 blogs. gaps arise in the mapping from sig to lsstree due to an individual annotator’s personal choices when encoding their story and selecting wordnet and verbnet propositions. annotators may opt to select a proposition that paraphrases the original story in various ways. this allows for a range of story encodings, but may result in an unexpected lsstree or resulting realizations. section 4 describes how the narrative sentence planner applies alterations to the lsstrees in order to generate rich text that prioritize natural realizations. thus the lsstrees act as an intermediary representation during the course of the storytelling pipeline, sitting between story planning and sentence planning, and having access to the affordances of the entire pipeline. 3.3 text realization from lexical-semantic story trees lsstrees are the output of the translation process in figure 3, from which natural language text can be generated using a surface realizer. the syntactic representation of lsstrees are a oneto-one mapping to deep syntactic structures (dsynts), the input to the real-time surface realizer, realpro (lavoie and rambow, 1997), which is utilized at the final stage of generation in the fabula tales pipeline (seen later in figure 8). realpro handles morphology, agreement and function words to produce an output string. gender, tense, co-reference, and articles are automatically handled by realpro at generation time. the dsynts formalism distinguishes between arguments, modifiers, and between different types of arguments (subject, direct and indirect object etc.). lexicalized nodes also contain a range of grammatical features used in generation. figure 7 shows the dsynts representation for the lsstree in figure 5. dsynts are ordered; the root is the main verb with required properties in xml format, including lexeme and tense. the rel argument indicates the relationship of the argument with respect to its parent. nouns require an article argument, indicating a definite or indefinite article. additionally, they can have a gender and number. possession is represented structurally, so “the squirrel’s reflection” is structured with “reflection” as the parent, and “squirrel” as the child, with the child also being possessive (pro). each lexeme from each lsstree node and information derived from the sig are used to map the lsstree to dsynts. text plans from the lsstrees are created with the discourse relation 49 lukin and walker 1 2 3 4 5 6 7 8 9 10 11 12 figure 7: dsynts and text plan corresponding to the lsstree in figure 5 scheherazade realization lsstree baseline, no narrative variation a narrator placed a steely and large bowl on a back deck in order for a dog to drink the water of the bowl. a crazy squirrel approached the bowl. the second squirrel began to be startled because it saw the reflection of the squirrel. the squirrel leaped because it was startled and fell over the railing of the deck and because it leaped. the squirrel held the railing of the deck with a paw of the second squirrel. the squirrel fell, and the paw of the squirrel slipped off the railing of the deck. the narrator placed the bowl on the deck in order for benjamin to drink the bowl’s water. the squirrel approached the bowl. the squirrel was startled because the squirrel saw the squirrel’s reflection. the squirrel leaped because the squirrel was startled. the squirrel fell over the deck’s railing because the squirrel leaped because the squirrel was startled. the squirrel held the deck’s railing with the squirrel’s paw. the squirrel’s paw slipped off the deck’s railing. the squirrel fell. table 7: the scheherazade and lsstree baseline realizations of the startled squirrel between dsynts nodes using the structure in figure 7. class properties are then written to a file, and the resulting file is processed by realpro to generate the text. table 7 compares the scheherazade realization (left-hand side) to the baseline realization as translated directly from the lsstrees with no narrative variation (right-hand side). the baseline story is told in chronological order by a direct translation of the sig timeline events into the surface order of the final realization. discourse relations such as contingency are always realized within a single sentence using because as a discourse cue. 50 a narrative sentence planner and structurer for storytelling the next section will evaluate the preservation of semantic content when mapping from scheherazade to lsstrees as measured by these baseline realization methods, in an effort to compare the ground-truth fabula of each story point without conflating the measure with the sujet. 3.4 evaluating the bridging process a prerequisite for producing stylistic variations of a story is the ability to generate a “correct” retelling of the story. to this end, we measure the semantic fidelity of our translation process, using text similarity metrics, as well as subjective measures of semantic similarity, grammaticality, and fluency. we compare the scheherazade realization against the baseline generation by realizing lsstrees without applying narrative variation as described in the previous section. pair # scheherazade realization lsstree realization, no narrative variation 1 the squirrel leaped because it was startled and fell over the railing of the deck and because it leaped the squirrel leapt because the squirrel was startled. the squirrel fell over the deck’s railing because the squirrel leaped because the squirrel was startled. 2 the narrator greeted the woman and the acquaintance. the narrator greeted capt john and ann. 3 the narrator didn’t initially notice that the group of bugs had entered the apartment of the narrator. the narrator did not initially notice that the bugs entered the narrator’s apartment. 4 a group of persons began to dive around great barrier reef, and the narrator entered some water. the narrator entered the water. 5 the milkmaid began to plan for the milk to later transform into some cream, for the milkmaid to later make the cream into some butter and to later sell the butter at a market, for she to later buy some eggs, for a group of chickens to later hatch from the eggs, for the milkmaid to later sell it, to later buy a gown and for she to later wear it at a fairground, for every fellow to later admire the gown and to later court the milkmaid, and for the milkmaid to later shake the head of the milkmaid and to later ignore every fellow. the milkmaid planned the milkmaid ignored every chap. table 8: pairs of scheherazade realizations and the lsstree baseline realizations we evaluate the semantic fidelity of the realizations at the sentence level, rather than at the whole story level, for greater precision. we create an evaluation set of 320 blog pairs and 100 fable pairs consisting of the scheherazade realization and the equivalent baseline realization for each story point (table 8 shows a subset of pairs). sometimes this comparison results in a single sentence in scheherazade being compared against more than one sentence in the baseline (e.g., pair #1 in table 8). this is because two or more events that had been encoded in the sig as taking place within a single story point have been split apart during lsstree creation and assigned a temporal discourse relationship. in addition, there are differences in how names are realized. in the baseline lsstree version in table 7, the dog is named benjamin whereas in the scheherazade version, that character is known simply as a dog (also see pair #2). pair #5 illustrates a case of an error in the process of creating the lsstree. 51 lukin and walker bleu nist meteor rouge sts fables 0.33 2.17 0.39 0.62 4.74 blogs 0.29 1.98 0.38 0.61 4.54 all 0.30 2.02 0.39 0.61 4.59 table 9: metrics comparing scheherazade and lsstree baseline realizations we first apply a set of automated metrics that are commonly used to evaluate nlg output. we use the scheherazade realization as the reference sentences against which the baseline realizations are evaluated, and apply the following metrics to each story pair using the e2e-metrics suite:3 bleu (papineni et al., 2002), nist (doddington, 2002), meteor (denkowski and lavie, 2014), and rouge-l (lin, 2004) (first four columns of table 9). as table 9 shows, the results appear to be somewhat low: a bleu score of 0.30 is much lower, for example, than the baseline system used in the e2e generation challenge. however, it is well known that these automatic metrics often do not reflect how well an nlg is actually performing (belz and reiter, 2006; novikova et al., 2017). we do not compare the original fables and blog stories to the scheherazade and lsstree baselines in these evaluations. it is crucial to recall that these baselines do not have any sentence planning or narrative variation, and using these automated metrics would result in a biased comparison against the original story. before applying sentence planning, we are primarily interested in verifying the semantic fidelity of the process of translating from the sig to lsstrees, rather than the generated language output. section 4 conducts human subject evaluations that do include comparison of the original stories to those generated with our narrative sentence planning, which serves as a more fair comparison. to measure the semantic fidelity of the translation, we conducted an evaluation using the semantic textual similarity metric (sts), obtained by human annotation,4 with results shown in the final column of table 9. semantic textual similarity is defined as a scale that ranges from 1. . . 5, where 5 is means exactly the same thing, 4 is means the same thing except for minor differences, 3 is means roughly the same thing, 2 is some important information is missing or different (cer et al., 2017). none of our pairs scored a 1, and only two scored a 2 where nested information from 2 fables were lost in translation (e.g., pair #5 in table 8). almost 80% of the pairs were scored a 5, and the human average is 4.59 over both datasets. one decision made in the sts evaluation was to treat names as important semantic entities. as we explained above, during annotation, deictics like i are often annotated as the narrator and their realization can be changed during sentence planning, whereas scheherazade does not realize names. in some cases, a character was given a name during annotation. characters with names are common in the blogs, but in the fables, characters are always known as their type the fox or the wolf. when treating names as important (i.e., less blog pairs are likely to be rated as a 5), we find a statistically significant difference between the sts means of blogs and fables that we hypothesize is due to name realization (paired t-test, df = 418, t = -2.78, p < 0.01). however, if names are treated as an unimportant part of the semantics of an utterance (i.e., more blog pairs are likely to be rated as a 5), the mean sts score of utterance pairs from the blogs are not statistically different from the mean sts score of the fables (paired t-test, df = 418, t = 0.065, p = 0.95). automatic metrics often conflate measures of fluency and naturalness with semantic correctness, and also penalize stylistic differences (oraby et al., 2018). we therefore conducted an addi3. https://github.com/tuetschek/e2e-metrics 4. the annotator was an author of this article. 52 https://github.com/tuetschek/e2e-metrics a narrative sentence planner and structurer for storytelling fluency grammaticality scheherazade baseline scheherazade baseline fables 3.63 3.22 4.08 3.54 blogs 3.36 3.53 4.24 4.20 all 3.43 3.45 4.21 4.03 table 10: metrics comparing scheherazade and lsstree baselines tional evaluation for fluency and grammaticality with a human annotation task where we rate the scheherazade and baseline realizations for each pair.5 we used a likert scale of 1 . . . 5 to state the degree of agreement with two statements: (1) the utterance is grammatical; and (2) the utterance is fluent and natural. the results for this evaluation are shown in table 10, and show a high degree of grammaticality but a lower degree of fluency. this bridge provides for a semantic mapping from sig to lsstree while maintaining lexical information from the story points. these semantic similarity metrics measure the quality of the translation process, and the grammaticality and fluency metrics provide a rough estimate of the quality of the baseline lsstree realization prior to sentence planning. in subsequent evaluation with human subjects, we include these scheherazade and lsstree baselines with the sentences generated with narrative sentence planning, as well as sentences from the original story texts, for a more comprehensive evaluation of stylistic expression. the application of these narrative variations are introduced in the next section. 4. narrative sentence planning the creation of the lsstrees provide the syntactic and semantic grounding necessary to generate multiple sujet using narratologically inspired parameters. we develop a narrative sentence planner for fabula tales that takes the automatically generated lsstrees and applies three narrative aspects that lönneker (2005) describes: mood:point of view, mood:distance, and voice:person (figure 8). we develop a sentence planner that implements parameters to model each of these narrative aspects, as we describe in detail below. our sentence planner is based on the architecture of the sentence planner in the personage nlg engine (mairesse and walker, 2008, 2011). we implement some of personage’s parameters related to voice: person and add new parameters that allow us to test particular narratologically inspired variations including mood: distance and mood: point of view. after manipulating the lsstrees along these narrative dimensions, they are realized as text using dsynts and realpro as described in section 3.4. we evaluate the sentence-level variations and show improvement over the baselines in the previous section. however, for some parameters, such as point of view and voice, it makes intuitive sense for this to be a story-wide decision, rather than a sentence-by-sentence decision. thus, we design a story-level narratological structurer to ensure consistency in the generated styles and voices, as we discuss in section 5. 5. we randomly mixed together the realizations so that the annotator would be blind to the source of the realization. the annotator was an author of this article. 53 lukin and walker figure 8: fabula tales pipeline applying narrative variations to lsstrees 4.1 mood: point of view mood:point of view is the “spatial, temporal, and ideological points of view from which events are described. events can be described from the point of view of different characters.” (lönneker, 2005). our architecture encodes characters, humanoid and non-humanoid, and props as unique actors or objects in the sig that can be used for easy co-reference. biber (1991) claims that first person pronouns are markers of ego-involvement with a text. first person pronouns are often the subject of cognitive verbs, and indicate that the matter at hand is personal and an immediate mental interaction. in contrast to third person pronouns (and third person narration), first person pronouns create a different perspective in the narrative space, by restricting the perception of events to the eye of a particular character, and thus allows the audience limited perception, or focalization (pizarro et al., 2003). any character in a story, including non-narrating or non-humanoid characters such as the squirrel in startled squirrel, can tell a story from their perspective using lsstrees. whenever firstperson is desired, the lsstree sets this parameter when creating the dsynts mapping. a major advantage of the lsstrees are that the deep linguistic representation allows for the specification of a change in point of view without manipulating the surface string or editing a template. table 11 shows the dsynts for “the squirrel”. in order to transform a sentence into the first person, from the dsynt in table 11, the person attribute is assigned to 1st to specify a change of point of view to first person, reflected in table 12. the realpro surface realizer interprets the person attribute and automatically changes the lexeme present to “i”. the lsstree representation tracks the identities of the characters and handles the realization of co-reference and possession (lukin and walker, 2015). 54 a narrative sentence planner and structurer for storytelling 1 table 11: dsynts for the squirrel 1 table 12: dsynts for i all the original texts from personabank are told in the first person perspective, yet when they are annotated using the sig there is no support for encoding different perspectives, because distinguishing between narrators is the job of the sujet, and not the fabula. to handle this, these stories are encoded with a “narrator” character, as we mentioned above. just as the “squirrel” lexeme can be changed with a simple person attribute change, so too can the “narrator”. in cases of multiple characters of the same type, e.g., two squirrel characters, the perspective change and subsequent co-reference tracking would only apply to the unique identifier each character is assigned during the lsstree creation as derived from the sig. 4.2 mood: distance (direct speech) mood:distance is the “combination of amount of information conveyed and narrator intrusion. stereotypically, detailed information and low narrator participation indicate imitation or ‘direct’ dramatic mode, as opposed to a ‘distant’, mediated narrative mode. this parameter also affects the way in which speech is reproduced” (lönneker, 2005). our main manipulation of this variable is based on our supposition that the distance between the reader and the story can be altered by varying whether speech is direct or indirect. when storytellers tell stories, they know what their characters are feeling, and can express it in the telling. bal claims that “dialogue is a form in which the actors themselves, and not the primary narrator, utter language” (bal, 1997) and that, in some cases, dialogue can make a narrative more dramatic. speech acts in the sig formalism are always encoded as indirect speech. in order to identify opportunities for direct speech (dialogue) the wordnet sense provided from the sig and encoded in the lsstree is used to identify whether the main verb is a verb of communication. if so, the lsstree is broken apart into two separate trees: the utterance to be uttered, and the explanatory phrase, and are linked by the pdtb discourse relation of attribution (prasad et al., 2008). there are many opportunities in both the dramabank and personabank to use direct speech: nine fables and forty eight blogs contained at least one speech act. for example, in the sentence anne said she didn’t receive the new schedule, from the personabank story called botched training, the verb say is identified as a verb of communication from verbnet, with anne as its subject (figure 9a). the remainder of the tree starting from the verb “receive” as the root verb, which is what is to be uttered, is split it off from its parent verb of communication, resulting in two smaller trees (figure 9b). after splitting, each tree is treated as a unique lsstree. a text plan is constructed consisting of the two lsstrees linked by the attribution relation (figure 9c). this text plan can then be realized in direct speech as “i didn’t receive the new schedule” anne said. 55 lukin and walker (a) lsstree for anne said she didn’t receive the new schedule (b) split lsstrees for anne said and anne didn’t receive the new schedule. (c) text plan with the “speech” discourse relation figure 9: lsstrees and textplans for direct speech 4.3 voice: person voice:person is the “narrator participation. a homodiegetic narrative instance is a character of the current narration (grammatical realization typically in the first person), while a heterodiegetic narrative instance is ‘absent’ from the current narrative and not referred to” (lönneker, 2005). different voices are showcased by combining different stylistic variations, including pragmatic marker insertion, lexical choice, and discourse structuring. to portray the voice:person parameter in fabula tales, we develop stylistic variations that, when combined, act as a character’s speaking style. pragmatic markers. biber suggests that emotive, cognitive, modal and uncertainty words are indicators of personal stories, whereas these items are lacking in impersonal stories (biber, 1991). many of these are considered to be pragmatic markers, which we expect to be more prevalent in the natural language of the blogs. we define pragmatic markers following biber for the following categories: acknowledgments (e.g., “oh”), emphasizers (e.g., “actually”, “rather”), competence mitigations (e.g., “come on”), down tones (e.g., “i mean”), tag questions (e.g., “no?”), expletives (e.g., “damn”). we also emulate stuttering (e.g., “tr-trellis”), contractions, and exclamation insertions. pragmatic marker insertion replicates personage’s mechanisms, which add nodes to the dsynts tree in the appropriate location (mairesse and walker, 2011). table 14 lists a number of these pragmatic markers with a description and an example realization. lexical choice. the wordnet and verbnet senses from the sig are used to manipulate the lexemes and structures of lsstrees with synonym substitutions. word senses are annotated when the sig is created, as explained above, and preserved in the lsstree. lexical choice can be controlled by implementing word frequency and word length as parameters, as in personage. in one story from personabank the narrator uses a comic book to try to kill some bugs that had been seen in his apartment. one of the sentences is i smeared the bug’s innards with the rolled comicbook. the synset for “innards” contains ‘viscera”, “entrails”, and “innards”. setting the word frequency parameter to be low could result in substituting “innards” with “viscera”. lexical substitutions for verbs are also possible but requires verifying that the synonym and its arguments 56 a narrative sentence planner and structurer for storytelling are interchangeable, for example, the verb “squash” in i managed to squash the bug is transformed to its argument equivalent “crush” in i managed to crush the bug. although word frequency and word length are often inversely correlated, there are cases where short words are rare. setting the word length parameter to be high for example could affect the lexical choice among the synset for “hue” in the fox and the crow where the fox flatters the crow by saying the hue of her plumage exquisite. lexical substitutions that are available in the synset for “hue” include “chromaticity”, which we see later in the variations in table 20. deaggregation and discourse structuring. traditionally, aggregation assumes an initial set of small semantic or syntactic pieces of information that can be readily combined. this may not always be the case; a clause may instead be packed with information that should first be decomposed, then restructured. we define deaggregation as the breaking apart or the decomposition of longer semantics into shorter semantics, which then affords the option and flexibility to aggregate the clauses in different arrangements, structures, or combinations. we observe that when sig semantics are converted to lsstrees, some story points contain a sequence of nested predicates that, when realized without sentence planning intervention, result in particularly long sentences, for example: the manager said she created the new schedule and the manager gave the new schedule to the employee in order for the employee to give the new schedule to anne. the contingency discourse relation is the target aggregation composition because these are highly likely to appear in narratives, and are abundant in personabank (103 total instances). in the sig, contingency clauses are expressed with the “in order to” relation. story points with this relation are identified (figure 10a) and split, to become the arguments of the contingency relation as illustrated by the two distinct trees in figure 10b. (a) lsstree for the narrator placed the steely bowl on the deck in order for benjamin to drink the bowl’s water. (b) split lsstrees for the narrator placed the steely bowl on the deck and benjamin drinks the bowl’s water figure 10: lsstrees for deaggregation 57 lukin and walker baseline the narrator placed a steely and large bowl of water outside on the back deck in order for a dog to drink the water of the bowl. relations contingency (nuc:1, sat:2) content 1: put(narrator, bowl, deck) 2: dog(drink, bowl) inorder i placed the bowl on the deck in order for benjamin to drink the bowl’s water. becausens i placed the bowl on the deck because benjamin wanted to drink the bowl’s water. becausesn because benjamin wanted to drink the bowl’s water, i placed the bowl on the deck. ns i placed the bowl on the deck. benjamin wanted to drink the bowl’s water. n i placed the bowl on the deck. sosn benjamin wanted to drink the bowl’s water, so i placed the bowl on the deck. table 13: content assertions, texts plan, and possible realizations for contingency table 13 shows new sentence planning variations for the contingency relation. the becausens operation presents the nucleus, the primary clause (n) first, followed by a because, and then the satellite, the supporting clause (s). becausesn and sosn reverse the order of the clauses. the nucleus and satellite can be treated as two different sentences (ns) or the satellite can be completely left off and only the nucleus realized (n). the richness of the discourse information present in the sig enables the storytelling framework to implement additional discourse relations abundant in narratives in future work. defining voice models. voice, in combination with changes to the point of view and direct speech, have the capacity to express the narrator as the storyteller, and the characters as speaking in their own style, similar to how other work has defined models using pragmatic markers to portray certain personality traits or character archetypes (mairesse and walker, 2008; lin, 2016; reed et al., 2011). speech acts realized as direct speech can express a characters’ particular style or way of speaking. an example of character voice is: “oh, well, i didn’t receive the new schedule!” said anne. the acknowledgement “oh” and exclamation mark inside the direct speech reflect the mental and emotional state of the character as they express themselves. compare this to direct speech only: “i didn’t receive the new schedule”, anne said. which makes use of direct speech, but not the stylistic variations of the voice. finally, narrator voice is the voice of the primary narrator or storyteller which may use stylistic parameters, but excludes the direct speech: oh, anne exclaimed that she did not receive the new schedule. voice models can be defined by setting all of the available parameters to have values between 0 and 1, in a similar way to how personality models were defined in the rule-based version of personage (mairesse and walker, 2010a). parameter values close to 1 indicate that the parameter should be used frequently, whereas parameter values near 0 indicate infrequent use of a parameter. in previous work, we build laid-back and shy models that are loosely based on the extrovert and introvert models from personage (mairesse and walker, 2008; rishes et al., 2013). table 14 58 a narrative sentence planner and structurer for storytelling model parameter description example shy voice softener hedges insert syntactic elements (sort of, kind of, somewhat, quite, around, rather, i think that, it seems that, it seems to me that) to mitigate the strength of a proposition ‘it seems to me that he was hungry’ stuttering duplicate parts of a content word ‘the vine hung on the tr-trellis’ filled pauses insert syntactic elements expressing hesitancy (i mean, err, mmhm, like, you know) ‘err... the fox jumped’ laid-back voice emphasizer hedges insert syntactic elements (really, basically, actually) to strengthen a proposition ‘the fox failed to get the group of grapes, alright?’ exclamation insert an exclamation mark ‘the group of grapes hung on the vine!’ expletives insert a swear word ‘the fox was damn hungry’ table 14: examples of pragmatic marker insertion parameters from personage shows the pragmatic markers used in combination to build each voice model. an additional, neutral voice is constructed, which does not use any pragmatic markers, lexical substitutions, or aggregation constructions (equivalent to the lsstree baseline). 4.4 evaluation of sentence-level variations we conduct a series of human evaluation tasks informed by previous research in this area (callaway and lester, 2002; cheong and young, 2008) to test the effectiveness of our narrative parameters on single sentences. our evaluations measure the following narrative metrics: • narrative immediacy: to what degree is the reader engaged with the story and characters? • interest: to what degree would the reader desire to read the rest of the story? • correctness: to what degree is the narrative well-formed? • preference: which framings do readers generally prefer to read? we hypothesize that generating stories by varying point of view (h1), character or narrator voice (h2), and aggregation operations (h3), will have an effect on these narrative metrics. 4.4.1 point of view and voice: engagement and interest we examine how point of view and voice interact with engagement and interest in a single sentence from a story, and hypothesize that excerpts told in different points of view and voice will have an effect on engagement and interest (h1 and h2). native english speakers on mechanical turk were presented with a one sentence summary of one of seven stories from personabank and six generated variations of one sentence from that story. these sentences are framed as “possible excerpts that could come from this summary”. table 15 shows an example of the embarrassed teacher story from personabank, the summary, and its six retellings. narrative variations include the first person with a neutral, shy, and laid-back voice, and a third person with a neutral voice, as described in 59 lukin and walker summary a teacher’s slip fell down in the middle of teaching a class. source example original nervously i looked down to see that my underslip had somehow made its way to the floor. scheherazade the narrator noticed that the ankle of the narrator was observed. 3rd neutral the narrator noticed for the narrator’s ankle to be observed. 1st out oh i noticed for my ankle to be damn observed! 1st neutral i noticed for my ankle to be observed. 1st shy i noticed for my ankle to be so-somewhat observed. table 15: variations presented to turkers for interest and narrative immediacy section 4.3. additionally, an excerpt from the original story from which the generated stories were derived is compared to test how close the best narrative sentence planning realization comes to matching the natural language of the blog. the strictly template-based scheherazade realization was also included. subjects rate each excerpt on a 1 . . . 5 point scale for their interest in wanting to read more of the story based on the style and information given in the excerpt, and to indicate their engagement with the story, given the excerpt (lukin and walker, 2015). we performed a set of anovas designed with the repeated items as categorical, independent variables, i.e., style (view and voice pairs) and story content (the particular sentence in question), subjects as a random variable, and aggregated across multiple items.6 style has an effect on interest (f(1) = 204.08, p < 0.0001), as does story content (f(9) = 7.32, p < 0.0001), but there is no interaction between style and story content. this may be interpreted as: style affects the sentence, and there is a random effect of story content, but interest preference is independent of the style and story content. style similarly has an effect on engagement (f(1) = 224.24, p < 0.0001) and story content (f(9) = 5.49, p < 0.0001). however, there is an interaction between style and story content (f(9) =1.65, p < 0.1), which suggests that for engagement, but not for interest, certain styles of narration are more appropriate or preferred than others given the context of the story. for example, subjects comment that the “curse words are used to express the severity of the situation wisely” and “adding the feeling of nervousness and where she looked made sense”, acknowledging the style fitting the situation. information from the story may be used to influence and produce a more engaging realization. we briefly discuss how being cognizant of content can influence realization in future work. table 16 shows the means and standard deviation for engagement and interest for each combination of voice and person. an ordered ranking emerges for both engagement and interest: the original sentence from the blog is scored highest, followed by first-person laid-back, first-person neutral, first-person shy, scheherazade, and third-person neutral. for the subsequent analyses, the key independent variable was point of view and voice pairs (i.e., style). bonferroni correction was applied to paired t-tests of style on engagement (full results in table 29). for engagement, there are statistically significant differences between the following styles in the ordered list: original and first laid-back, first neutral and first shy, and first shy and scheherazade. however, there are no other differences between sentences. therefore, we observe 6. a linear effects model was not used because our item independent variables are not mixed. 60 a narrative sentence planner and structurer for storytelling style engagement interest mean std err mean std err original 3.98 (1.07) 3.91 (0.99) 1st-laid-back 3.27† (1.39) 3.02† (1.21) 1st-neutr 3.00 (1.19) 3.02 (1.37) 1st-shy 2.73† (1.26) 2.81† (1.27) scheherazade 1.95† (1.07) 1.90† (1.05) 3rd-neutr 1.93 (1.06) 1.87 (1.01) table 16: means and standard deviation for engagement and interest in perceptions experiment (higher is better; † indicates statistical significance between the marked style and the style in the row above; see tables 29 and 30 for detail) that h1 is supported, with statistically significant differences in point of view realizations for the engagement metric, as well as h2, with statistically significant differences in voice between laidback, shy, and neutral. similarly, bonferroni correction was applied to paired t-tests on style and interest (full results in table 30). results for interest follow a similar trend, showing a statistically significant difference between original and first laid-back, first neutral and first shy, and first shy and scheherazade, again supporting h1 and h2 for the interest metric. there are no other differences between ordered pairs. 4.4.2 discourse structuring: correctness and preference we examine how discourse structuring interacts with correctness and preference in a single sentence from a story. we hypothesize that the deaggregation and discourse structuring variations will effect reader preferences and belief about the correctness of the narrative (h3). we explore (1) how the variations compare to each other; (2) if they come close to the natural language of the original blog story; and (3) if the narrative sentence planning realization surpasses the scheherazade realization. we create a mechanical turk experiment showing an excerpt from the original story, where we tell our qualified turkers that “any of the following sentences could come next in the story” (table 17). subjects are queried about the variations in terms of correctness and goodness of fit within the story context. they are then asked to rank the sentences by personal preference (in experiment 1, we showed 7 variations where 1 is best, 7 is worst; in experiment 2 we showed 3 variations where 1 is best, 3 is worst). we emphasize in the prompt that subjects should read each variation in the context of the entire story, and encourage them to reread the story with each new sentence to understand this context (lukin et al., 2015). we performed a set of anovas designed with the repeated items as categorical, independent variables, i.e., realization (variation) and story content (the particular sentence in question), subjects as a random variable, and aggregated across multiple items.7 in the first experiment, seven native english speakers on mechanical turk analyzed 16 story segments from different blogs in personabank with the following variations: the original story, sosn, becausens, becausesn, ns, n, and the non-deaggregated realization. as expected, realization had an effect on correctness (f(6) = 9.8, p < 0.0001) and preference (f(6) = 31.7, p < 0.0001) supporting hypothesis h3 that the realizations are distinct from each other and there are preferences among them, as well as varying degrees of 7. a linear effects model was not used because our item independent variables are not mixed. 61 lukin and walker story this is one of those times i wish i had a digital camera. we keep a large stainless steel bowl of water outside on the back deck for benjamin to drink out of when he’s playing ouside. his bowl has become a very popular site. throughout the day many birds drink out of it and bathe in it. source example original the birds literally line up on the railing and wait their turn. scheherazade the birds organized themselves on the deck’s railing. becausesn because the birds wanted to wait, they organized themselves on the deck’s railing. becausens the birds organized themselves on the deck’s railing because the birds wanted to wait. sosn the birds wanted to wait, so they organized themselves on the deck’s railing. none the birds organized themselves on the deck’s railing in order for the birds to wait. ns the birds organized themselves on the deck’s railing. the birds wanted to wait. n the birds organized themselves on the deck’s railing. table 17: deaggregation and discourse structuring variations presented to turkers for correctness and preference judgments realizations correctness preference mean std err mean std err original 1.83 (1.34) 2.38 (2.28) sosn 2.32† (1.26) 3.07† (1.89) becausens 2.44 (1.28) 3.65† (1.78) becausesn 2.45† (1.26) 3.73 (1.93) ns 2.69† (1.13) 4.25† (1.53) none 2.72 (1.10) 4.86† (1.72) n 3.01 (1.14) 4.90 (1.47) table 18: means for correctness and preference for discourse structure experiment 1 (lower is better; † indicates statistical significance between the marked realization and the realization in the row above; see tables 31 and 32 for detail) grammaticality. story content had no effect on correctness or preference, suggesting that all stories were well-formed and there were no outliers in the story selection. we find an interaction between realization and story content for correctness (f(2, 110) = 1.83, p < 0.0001) and preference (f(2, 110) = 3.24, p < 0.0001), thus subjects’ preference of the realization are based on the context of the story, unlike in the previous analysis of point of view for engagement and interest. table 18 shows the means and standard deviations for correctness and preference rankings for each realization in the first experiment. averaged across all stories, there is a clear order for correctness and preference: original, sosn, becausens, becausesn, ns, non-deaggregated (indicated as none), and n. for the subsequent analysis, the independent variable was if deaggregation was performed, and with which aggregation discourse structure construction. bonferroni correction was applied to paired t-tests of realization on correctness (full results in table 31). there are statistically significant difference in correctness between original and sosn, and between becausens and becausesn, as well as becausesn and ns. similarly, bonferroni correction was applied to paired t-tests of real62 a narrative sentence planner and structurer for storytelling realizations correctness preference mean std err mean std err original 1.57 (0.92) 1.37 (0.59) sosn 2.49† (1.29) 1.93† (0.67) scheherazade 3.50† (1.43) 2.70† (0.57) table 19: means for correctness and preference for discourse structure experiment 2 (lower is better; † indicates statistical significance between the marked realization and the realization in the row above; see tables 33 and 34 for detail) ization on preference (full results in table 32). results for preference follow a similar trend, with the addition of sosn and becausens. these results indicate that the original sentence is the most correct and preferred. in a qualitative evaluation, subjects commented that while all variations were sufficient, most were “boring”, except for the original blog story excerpt. the n and ns variations are overall ranked the lowest because they sometimes produce stilted language and remove pieces of content. however, in a few instances, these variations are ranked highly because the information they remove was deemed to be redundant in text realization or repeated content, which we posit shows support for the interaction between realization and story content. in a second experiment, we compare the original blog sentence with the highest scoring discourse structure variation with a point of view change, and the realization produced by scheherazade. we expect that scheherazade will score poorly in this instance because it cannot realize deictic expressions to change point of view from third person to first person, even though it is derived directly from the sig representation. seven native english speaking subjects analyzed each of the 19 story segments in a similar experimental setup as the deaggregation experiment 1. our anovas were conducted following the same design as the first experiment in this section. realization had an effect on correctness (f(2) = 6.78, p < 0.0001) and preference (f(2) = 131.9, p < 0.0001), again, supporting hypothesis h3. story content had no effect, suggesting that there were no outlier stories, and there was an interaction between realization and story content for correctness (f(2, 47) = 5.48, p < 0.0001) and preference (f(2, 47) = 9.25, p < 0.0001), suggesting that subjects’ evaluation is based on the realization and the context of the story. table 19 shows the means and standard deviations for correctness and preference rankings for the realizations in the second experiment. there is a clear order for correctness and preference: original, sosn, scheherazade. bonferroni correction was applied to paired t-tests of realization on correctness and preference (full results in tables 33 and 34). for the majority of the stories, subjects do not select scheherazade because of “the narrator” realization, commenting “forget the narrator sentence. from here on out it’s always the worst!”. however there are three story segments where scheherazade is rated on average higher than sosn. upon closer examination, these story segments do not contain “i” or “the narrator” in the story content, so the sentence is evaluated without the “narrator” bias. however, even without that bias, sosn still outranks scheherazade: in a story about a protest at the g20 summit, the sosn realization: the leaders wanted to talk, so they met near the workplace. 63 lukin and walker is much more natural than the scheherazade realization: the group of leaders was meeting in order to talk about running a group of countries and near a workplace. these evaluations have shown that realizations using the first person point of view, pragmatic features, and aggregation variations, are more engaging, interesting, correct, and preferred than the scheherazade baselines and lsstree baselines without narrative. in the next sections, we explore how to intelligently plan and extend the narrative sentence planner to support story-level variations. 5. narratological structurer the previous sections have shown that fabula tales’ sentence planning parameters are capable of generating hundreds of sentences for any story that has first been annotated as a sig, with high semantic fidelity to the fabula of the original story. table 20 shows complete story generation achieved by setting several parameters for the fox and the crow, both using direct speech and different voice models for each character. these outputs are generated with our default story-level narratological structurer: this simply sets the parameters for the whole story to be consistent, e.g. if the direct speech parameter is set to 1, direct speech will be used throughout the story, rather than alternating between direct and indirect. variation 1: shy crow and laid-back fox variation 2: laid-back crow and shy fox the crow sat on the tree’s branch. the crow thought “i will eat the cheese on the branch of the tree because the clarity of the sky is somewhat beautiful.” the fox observed the crow. the fox thought “i will obtain the cheese from the crow’s nib.” the fox averred “i see you!” the fox alleged “your beauty is quite incomparable, okay?” the fox alleged “your feather’s chromaticity is exquisite.” the fox said “if your voice’s pleasantness is equal to your visual aspect’s loveliness you undoubtedly are every birds’ queen!” the crow thought “the fox was somewhat flattering.” the crow thought “i will demonstrate my voice.” the crow loudly cawed. the cheese fell. the fox snatched the cheese. the fox said “you are somewhat able to sing, alright?” the fox alleged “you need wits!” the crow sat on the tree’s branch. the crow thought “i will eat the cheese on the tree’s branch because the sky’s limpidity is beautiful”. the fox observed the crow. the fox thought “i will obtain the cheese from the crow’s pecker.” the fox averred “i see the bird.” the fox alleged “your beauty is somewhat incomparable.” the fox alleged “your feather’s chromaticity is somewhat exquisite.” the fox said “if your voice’s sweetness is somewhat equal to your appearance’s beauteousness you undoubtedly are every birds’ queen.” the crow thought “the fox was flattering, you know, okay?” the crow thought “i will demonstrate my voice.” the crow loudly cawed. the cheese fell. the fox snatched the cheese. the fox said “you are somewhat able to sing.” the fox alleged “you need wits.” table 20: the fox and the crow variations produced by the narratological structurer now, however, we consider that there is no guarantee that a naı̈ve combination of all these narrative parameters will produce an appropriate narrative flow. first of all, it is clear that narrative text generation at the story-level should maintain a degree of sentence-by-sentence consistency. for example, the character or narrator voice and person parameters should not dramatically change mid-story without reason. similarly, point of view should remain consistent throughout a story segment.8 however, other narrative aspects, such as the use of direct speech and different syntactic constructions, may produce better stories if they are varied throughout a story. 8. in a story with multiple chapters or segments, the style or point of view may change between segments, but within a segment this is generally consistent. 64 a narrative sentence planner and structurer for storytelling we design and build a narratological structurer for fabula tales that sits above the sentence planner and dictates narrative operations to the sentence planner (see figure 8). to train the narratological structurer, we undergo two phases: overgenerate and rank. in the overgenerate phase, we naı̈vely plug the sentence-level parameters into sentences to generate an abundance of training data (section 5.1). we design a create-your-own-story paradigm for the rank phase, allowing subjects to construct a story sentence-by-sentence by selecting from the sentences generated in the overgenerate phase (section 5.2). subject choose sentences that best contribute to the overall narrative flow. the rankings measure the effectiveness of each narrative parameter in the selected sentences and are utilized by the narratological structurer to make story-level generation decisions. we evaluate the narratological structurer’s generation capabilities and test exploratory hypotheses based on narrative theories from lönneker (2005) and observations from biber (1991). similar to as before, we hypothesize that generating stories by varying point of view (h1), character or narrator voice (h2), and aggregation operations (h3), will have an effect on reader perceptions. we add a new hypothesis (h2a) that direct speech in isolation will have an effect on reader perceptions (section 5.3). finally, we conduct a classification exercise to determine if the data and features collected from the overgenerate and rank phases can be used to identify which pre-generated sentences will be the most preferred (section 5.4). 5.1 overgenerate: generating training data depending on the narratological, structural, or lexical features present in the encoding, fabula tales produces different variations when generating variations. for this study, four stories from personabank are used, each seven sentences in length. a total of 2330 different sentences variations were generated from the original 28 baseline sentences. sentences are generated with combinations of all the parameters discussed in section 4: point of view, direct speech, and voice. # botched training variations 1 i rather excitedly entered pf changs because the manager wanted to train me 2 i excitedly entered pf changs in order for the manager to train me 3 the manager wanted to train me, so i excitedly entered pf changs 4 ok, i excitedly entered pf changs in order for the manager to train me, right? 5 because the manager wanted to train me, i excitedly entered pf changs 6 the manager wanted to train anne, so she excitedly entered pf changs, as it were 7 because the manager wanted to train anne, she excitedly entered pf changs!! 8 anne excitedly entered pf changs 9 essentially, ok, the manager wanted to train anne, so she excitedly entered pf changs 10 actually, anne excitedly entered pf changs in order for the manager to train her 11 the director wanted to train anne, so she excitedly entered pf changs 12 the manager wanted to train me, so i excitedly entered pf changs, okay? table 21: variations of first sentence of botched training story table 21 illustrates a subset of these variations of the first sentence from the botched training story from personabank. the sentence can be deaggregated into two clauses: anne excitedly entered pf changs and the manager wanted to train anne. outputs 1, 3, 5, and 8 have different discourse constructions, including the arguments, using different discourse cues, or removing the less important argument completely on the assumption that it is likely to be redundant. outputs 2 and 4 do not deaggregate, and instead realize the most straightforward logical form. output 1 has 65 lukin and walker the emphasizer “rather”, outputs 4 and 9 have the acknowledgement “ok”, and output 7 has exclamation marks. there are variations both in the first and third person point of view. because there is no speech act in the sample sentence in the table 21, this utterance is defined as evoking “narrator voice”. 5.2 rank: training the narratological structurer the goal of the rank phase is to learn story-level parameters for training a narratological structurer that preserves the desirable properties of random, probabilistic generation of different story versions, while at the same time making sure that these parameters take feedback from readers into consideration. the create-your-own-story paradigm allows subjects to build a story sentence-by-sentence by selecting from a subset of the overgenerated sentences. figure 11 shows the experimental design. for tractability, we downselect the generated sentences into a set of 5 variations per sentence from the overgeneration phase, resulting in 28 sentences per story, yielding a total of 140 sentences from which the subject can create use to create stories. the subset was designed to showcase each feature that can appear in at least one sentence in a variety of combinations with other features. subjects select the sentences they like best. at the bottom of the experiment, the progression of the reconstructed story is dynamically updated so subjects can read the full story to see how it flows. subjects are encouraged to read each sentence within the context of the entire reconstructed story. at any time, they may select a different sentence, yielding a dynamic update to the reconstructed story. when finished, the subjects rate how much they like their story on a 5-point likert scale, and then are asked to give detailed feedback about why they selected the sentences they did. nine subjects on mechanical turk who were prequalified for language-based reading and comprehension tasks completed this task. a total of thirty reconstructed stories were created with an average enjoyment score of 3. showing the reconstructed story at the bottom of the experiment allowed subjects to engage with their own perceptions of the flow of the narrative. many annotators commented about the flow of the story and keeping consistency. we make two notable observations from the qualitative feedback: (a) annotators tried to create stories with a good flow and consistency; and (b) pragmatic marker features as a way to create a character voice are pragmatically odd in many cases. this qualitative and quantitative analysis gives us insight into which features are selected, how often, and how the narratological structurer can use the high ranked features in generating future stories. deaggregation, direct speech, and contractions are popular and used more than 50% throughout a story. however, some pragmatic markers for character voice in direct speech or narrator voice are not used as consistently because they are pragmatically odd, and do not take context into consideration at the individual sentence-level or across the entire story. by analyzing the sentences that were selected and those that were not selected by annotators during the “create your own story” experiment, we design two metrics for training the narratological structurer’s parameterizable model: the catratio metric is aimed at learning the appropriateness and placement of the narrative features in individual sentences, and the perstoryratio discovers the balance of how many times a particular feature should occur within a story. in the next section, these statistics are used to test whether the application of the catratio and perstoryratio improves the quality of the story-level generation. 66 a narrative sentence planner and structurer for storytelling figure 11: create-your-own-story experimental design (sentence sets 2-6 omitted for space) we define the category-ratio metric, the percentage of the time a particular feature i is used with respect to the other features in each category, as: catratioi = seli seli.total (1) where i is an item in a feature category, seli is the number of times a sentence with feature i was selected, and seli.total is the total number of selected sentences in that feature category. for example, from table 22, there are a total of 217 sentences that were selected by the subjects that had the potential for an acknowledgment to be inserted (selack.total = 217). of those 217 selected sentences, only 29 actually had an acknowledgement present (selack.present) while the majority, 188, did not have the acknowledgment realized (selack.¬present). thus, catratioack.present = 29 217 = 0.13, indicating 13% of the sentences selected contained an acknowledgement, whereas, catratioack.¬present = 0.87 indicates that the other 87% of the selected sentences did not have the acknowledgement. many of the individual voice parameters rarely or never appear in the selected sentences, including “i mean”, “come on”, “like”, and “no?”. we believe placement is the problem because these stories were not generated with any story-wide constraints; indeed, this is what we aim to learn through this experiment. popular features were the acknowledgement “yeah” and the emphasizers “really”, “very”, and “actually”. the competence mitigation toner is rarely used, and the tag question category is never selected. contractions were selected an overwhelming 85%, and may be more likely to emulate the flow and naturalness of everyday speech, regardless of narration or direct speech. for sentences with exclamations, catratioexclam.present is 40%. table 22 also shows that 67 lukin and walker category features sel catratio acknowledgment yeah 14 .06 right 4 .02 {i see, oh} 3 .01 {oh my god, ok, well, oh yeah} 1 < .01 {great, okay?} 0 0 ack.present 29 .13 ack.¬present 188 .87 ack.total 217 competence mitigations obviously 3 .01 come on 0 0 mit.present 3 .01 mit.¬present 214 .99 mit.total 217 down rather 17 .08 {i mean, like, somewhat} 0 0 down.present 17 .08 down.¬present 200 .92 down.total 217 emphasizers actually 11 .05 really 9 .04 {great, very, you know, especially} 1 < .01 {as it were, basically, essentially, obviously} 0 0 emp.present 26 .12 emp.¬present 191 .88 emp.total 217 exclamation exclam.present 50 .40 exclam.¬present 74 .60 exclam.total 124 contraction contr.present 29 .85 contr.¬present 5 .15 contr.total 34 discourse structuring disc.present 67 .74 disc.¬present 23 .26 disc.total 90 direct speech ds.present 40 .69 ds.¬present 18 .31 ds.total 58 table 22: feature categories for ranking generated stories 74% of selected sentences have a discourse structure variant. there is also a slight preference for direct speech, with a catratiods.present of 69%. the perstoryratio is defined for for pragmatic markers based on the observed data such that 60% of reconstructed stories do not have any voice or style features, 27% have only one, 7% have two, and none have more than two. the same pragmatic feature is never selected twice in a story. while a few pragmatic features may be good for expressing character voice, too many repetitions of the same or similar markers appear to be perceived as unnatural. 68 a narrative sentence planner and structurer for storytelling generated variant 1 generated variant 2 because the manager wanted to train anne, she excitedly entered pf changs!! the manager lazily said yesterday she scheduled anne in order for her to train anne and anne didn’t show up. anne was confused. the rather disgruntled manager lazily said the schedule was erroneous. right, the manager lazily said the schedule was erroneous, so anne insistently questioned the manager. “i created a new schedule and i gave the new schedule to an employee in order for her to give the new schedule to you”, the manager melodramatically said. “oh i see, i didn’t receive the new schedule!”, anne said. i excitedly entered pf changs in order for the manager to train me. “yesterday i scheduled you in order for me to train you and you didn’t show up”, the manager lazily said. i mean, i was confused because the schedule demonstrated my punctuality. “the schedule was erroneous”, the disgruntled manager lazily said. i insistently questioned the manager because the manager lazily said the schedule was erroneous. “i created a new schedule and i gave the new schedule to an employee in order for her to give the new schedule to you”, the manager stated. “i didn’t receive the new schedule!”, i said. oh my god, right? table 23: two stories constructed by the trained narratological structurer the narratological structurer uses the catratio and perstoryratio metrics to dictate to the sentence planner when to apply a particular narrative parameter to the lsstrees when generating the fully realized stories. for each narrative parameter, the narratological structurer uses a probability distribution to determines what value to assign to that parameter according to the catratio percentages. it also keeps track of how many of each parameter have been used within the story so far, if applicable, and makes generation decisions according to the perstoryratio. in the next section, we test the effectiveness of generating full stories using the learned catratio and perstoryratio statistics in the narratological structurer. 5.3 evaluation of narratological structurer stories generated by the trained narratological structurer are shown in table 23. variant 1 is told in the third person perspective, with a variety of sentence constructions. because this story is told in the third person, the use of direct speech gives opportunities to the characters to express themselves in their own voice, e.g. “oh i see” in the direct speech of anne. variant 2 is told entirely in the first person perspective of the employee, and all direct speech, plus narrator observations, such as “oh my god, right?”, that cannot appear in a third person narration. we conduct three evaluations of the narratological structurer. the first compares stories generated according to these statistics against the lsstree baseline presented in section 3.4 (rank vs. baseline). these tests consist of a number of ablation tests to study the narrative hypotheses h1 h3. the second compares the stories generated by the narratological structurer to stories generated with randomly assigned narrative parameter values (rank vs. random). finally, we compare the randomly generated stories to the lsstree baselines (random vs. baseline). rank vs. baseline. we use the narratological structurer to generate a set of eleven stories from personabank in the following manner: a subset are generated in the first person point of view but with no other parameters (to be compared to the third person baseline, to test h1); a subset are generated with direct speech according to the narratological structurer (to be compared to the indirect speech baseline, to test h2a); a subset are generated with voice parameters according to the narratological structurer (to be compared to the neutral voice baseline, to test h2); and a subset are generated with deaggregation and discourse structuring according to the narratological structurer (to 69 lukin and walker be compared to the no deaggregation baseline, to test h3). each pair of rank-baseline stories are annotated by seven subjects on mechanical turk, who were asked which story they prefer (framed as “story a or story b”). the results indicated a preference for the stories generated by the ratio statistics over the baseline stories 75% of the time. we conducted the subsequent analyses using an anova with independent variables of point of view (h1), direct speech (h2a), direct speech and style (h2), and deaggregation and discourse structuring (h3). the dependent variable was preference. h1 claims that stories told in the first or third person point of view will effect reader preferences. point of view is not statistically significant, nor is there an effect or interaction of story content. one subject observed in qualitative feedback that first person can lead to more opportunities for the characters to describe their feelings, while another identified that the first person story added a “sense of immediacy” and “tension”. however, we find that other subjects “simply preferred the third person” without providing additional rational. h2a claims that direct speech will effect reader preferences, and we observe a statistically significant difference in preference between stories generated by the narratological structurer using direct speech and baseline stories ((1, n=42) = 23.3, p < 0.0001). no effect of story content or interaction is observed, so this finding is independent of topic of the story. readers commented that alternating narrative description in the non-speech sentences and dialogue adds personality to the story. h2 claims that pragmatic markers, when used as character or narrator voice and generated according to the ratio statistics, will effect reader preferences. we observe a statistically significant difference in preference between stories generated by the narratological structurer using voice and baseline stories ((1, n=98) = 74.0, p < 0.0001). while there is no effect of story content alone, there is an interaction between pragmatic markers and story content (f (5, 98) = 17.5, p < 0.0001) suggesting that different preferences for pragmatic markers emerge in different contexts. we examine in greater detail the direct speech and pragmatic marker interaction with the generated texts similar to those generated in table 24. first, we compare direct speech only and narrator voice. direct speech only is preferred 100% of the time. subjects say direct speech only is easier to understand, which is conceivable because the direct speech breaks up the narrative into alternating speech acts and narration segments. when comparing stories with a narrator voice to stories with character voice, character voice stories are preferred 95% of the time. baseline the manager said the schedule was erroneous. anne questioned the manager because the manager said the schedule was erroneous. ... anne said she didn’t receive the new schedule. direct speech only “the schedule was erroneous”, the manager said. anne questioned the manager because the manager said the schedule was erroneous. ... “i didn’t receive the new schedule”, anne said. narrator voice the manager said the schedule was obviously, erroneous. anne questioned the manager because the manager said the schedule was erroneous. ... yeah, anne said she didn’t receive the new schedule. character voice “the schedule was obviously, erroneous”, the manager said. anne questioned the manager because the manager said the schedule was erroneous. ... “yeah, right, i didn’t receive the new schedule”, anne said. table 24: excerpts of different speech conditions in rank vs. baseline ablation test 70 a narrative sentence planner and structurer for storytelling however, comparing character voice stories with direct speech only stories, we observe a moderately strong preference (71%) for direct speech only. one subject comments that character voice with the pragmatic marker voice features yields “mixed results”. subjects comment on what might be an issue of context insensitivity. for example, in character voice, the generated text includes yeah, right as part of what anne says in the direct speech. subjects identified that this creates a tone they believe was not appropriate with respect to anne’s tone in the rest of the story. while these voice and style features are more acceptable as character voice than as narrator voice and are moderated according to the ratio statistics, the character voices must take additional care to take into consideration appropriate character emotion or appraisal to be pragmatically cohesive. h3 claims deaggregation and discourse structuring will effect reader preferences, and we observe a statistically significant difference in preference between stories generated by the narratological structurer’s discourse structuring and baseline stories ((1, n=52) = 4.2, p < 0.01). there is no effect of story content or interaction. table 25 shows examples of story segments generating with constructions that are preferred over the baseline. deaggregation and discourse structuring were observed to create cleaner and crisper stories because they are shorter compared to the baseline stories, but still conveyed the point (e.g., the rank condition pair 2 in table 25). pair rank condition baseline condition 1 because the bugs scared john, he grabbed the rolled comic book. john grabbed the rolled comic book because the bugs scared him. 2 the squirrel fell over the deck’s railing. the squirrel fell over the deck’s railing because the squirrel leaped because the squirrel was startled. 3 the squirrel was startled, so the squirrel leaped. the squirrel leaped because the squirrel was startled. table 25: excerpts of different discourse structuring conditions in rank vs. baseline ablation test rank vs. random. next, stories generated according to the ratio statistics are compared against stories created with randomly assigned parameter values. for example, in a ratio controlled story, an acknowledgement may only appear at most two times in character voice, whereas, in a random story, the same acknowledgement may appear in character voice, for example, four times. ten story pairs of these types are annotated by the same seven mechanical turkers. stories generated with ratio statistics were preferred 92% of the time. subjects found stories constructed using the statistical models easier to read and understand. random vs. baseline. finally, the stories with randomly assigned parameter values are compared against baseline stories with no sentence planning variation. six story pairs of these types are annotated by the same seven mechanical turkers. the baseline stories were preferred 96% of the time. subjects noted that no variation is better than poor variation, even if the baseline stories were plain. in one pair, the random story contained direct speech, a high ranked ratio feature, with pragmatic markers in the direct speech. however, subjects found the dialogue to be stilted and did not contribute to a “good” character voice, and therefore preferred no variation. in the cases where a subject preferred the random story, it was stated that the baseline was “too dry”. 71 lukin and walker summary. these experiments show that the catratio and perstoryratio statistics learned from the overgeneration and rank experiments ensure consistency in the narratological structurer’s generation throughout a story, and provide insight into how the narrative variations from the sentence planner can be combined to be most effective. these stories are preferable to randomly constructed stories. we also learn that no narrative variation is typically preferred to no poorly controlled random variation. in summary, the parameterized narratological structurer ensures that not only are generated stories unique, but that they will follow the standards of consistency and balance learned from observational data. 5.4 predicting selected sentences in addition to the statistical selection of the narratological structurer, we seek to determine whether we can learn a ranking, or preference function, that would allow us to select better sentences when generating stories, in a similar way to previous work (walker et al., 2007). the findings here would be complementary to catratio and perstoryratio, and used to make more informed decisions when generating during the overgeneration phase. we randomly split the 140 sentences used in the “create your own story” experiment into 100 for training and 40 for test. each sentence was counted for how many times it was selected by annotators and binned as a binary selected or not-selected class.9 the task is classification: for each sentence, predict if it was selected by the annotators. we develop several feature sets with binary values to capture key aspects of these sentences: • n-gram features: unigrams, bigrams, and trigrams. • punct features: punctuation marks. • pragmatic features: insertion of pragmatic markers, e.g., ack : ok = true. • narrative features: discourse structuring, point of view, and direct speech parameters. an additional feature set, pragmatic*, was created to further capture information about the context of these parameters. rather than a binary “present” or “not present”, pragmatic* features represents the count of the features present. for example, the pragmatic feature ack : ok = true for the pragmatic feature set could be ack : ok = 2 for pragmatic* feature set. we train multiple off-the-shelf classification models using weka.10 the best models were weka’s support vector machine implementation, smo, and its decision tree, j48.11 tables 26 and 27 show the results on the training set using combinations of our features sets. the majority class baseline for predicting selected (sel.) is 0.4 and not-selected (¬sel.) is 0.6. in both models, the “punct” features perform the best; for j48, no other features improve over the performance of “punct” alone, and for smo, “punct” performs well in combination with “prag” and “ngram”. “n-gram” performs poorly on its own, which we posit could be due to the small size of the dataset or diversity of story content and domain specific vocabulary. we expected the narrative features to be useful because they represent more abstract narrative information, however on their 9. we note that the selected labels might be a result of the least-worst sentences, i.e., stories that were not generated with the best catratio and perstoryratio metrics, but this would still yield a subset of better potential sentences. 10. https://www.cs.waikato.ac.nz/ml/weka/ 11. other models tested include weka’s naı̈ve bayes and multilayer perceptron. 72 https://www.cs.waikato.ac.nz/ml/weka/ a narrative sentence planner and structurer for storytelling feature set precision recall f-measure sel. ¬sel. avg. sel. ¬sel. avg. sel. ¬sel. avg. ngram 0.39 0.70 0.55 0.42 0.67 0.55 0.41 0.69 0.55 punct 0.48 0.72 0.60 0.36 0.81 0.59 0.41 0.76 0.59 prag 0.09 0.64 0.37 0.03 0.85 0.44 0.05 0.73 0.39 narr 0.00 0.67 0.34 0.00 1.00 0.50 0.00 0.80 0.40 ngram-narr 0.40 0.71 0.55 0.46 0.66 0.56 0.42 0.68 0.55 ngram-punct 0.45 0.74 0.59 0.52 0.69 0.60 0.48 0.71 0.60 ngram-prag 0.39 0.70 0.55 0.42 0.67 0.55 0.41 0.69 0.55 punct-narr 0.48 0.72 0.60 0.36 0.81 0.59 0.41 0.76 0.59 prag-narr 0.00 0.63 0.32 0.00 0.85 0.43 0.00 0.73 0.36 prag-punct 0.48 0.73 0.61 0.42 0.78 0.60 0.45 0.75 0.60 ngram-prag-narr 0.40 0.71 0.55 0.46 0.66 0.56 0.42 0.68 0.55 ngram-prag-puct 0.43 0.73 0.58 0.49 0.69 0.59 0.46 0.71 0.58 prag-punct-narr 0.52 0.75 0.63 0.46 0.79 0.62 0.48 0.77 0.63 ngram-prag-punct-narr 0.43 0.73 0.58 0.49 0.69 0.59 0.46 0.71 0.58 table 26: training classification with smo in weka (highest averaged class f-measure in bold) feature set precision recall f-measure sel. ¬sel. avg. sel. ¬sel. avg. sel. ¬sel. avg. ngram 0.27 0.66 0.46 0.12 0.84 0.48 0.17 0.74 0.45 punct 0.55 0.77 0.66 0.52 0.79 0.65 0.53 0.78 0.66 prag 0.14 0.66 0.40 0.03 0.91 0.47 0.05 0.76 0.41 narr 0.00 0.67 0.34 0.00 1.00 0.50 0.00 0.80 0.40 ngram-narr 0.27 0.66 0.46 0.12 0.84 0.48 0.17 0.74 0.45 ngram-punct 0.43 0.71 0.57 0.36 0.76 0.56 0.39 0.73 0.56 ngram-prag 0.27 0.66 0.46 0.12 0.84 0.48 0.17 0.74 0.45 punct-narr 0.52 0.75 0.63 0.46 0.79 0.62 0.48 0.77 0.63 prag-narr 0.00 0.66 0.33 0.00 0.96 0.48 0.00 0.78 0.39 prag-punct 0.53 0.78 0.65 0.58 0.75 0.66 0.55 0.76 0.66 ngram-prag-narr 0.27 0.66 0.46 0.12 0.84 0.48 0.17 0.74 0.45 ngram-prag-puct 0.43 0.71 0.57 0.36 0.76 0.56 0.39 0.73 0.56 prag-punct-narr 0.53 0.78 0.65 0.58 0.75 0.66 0.55 0.76 0.66 ngram-prag-punct-narr 0.43 0.71 0.57 0.36 0.76 0.56 0.39 0.73 0.56 table 27: training classification with j48 in weka (highest averaged class f-measure in bold) own, the narrative features are not informative, and in some cases, bring down the scores when combined with other feature sets. we observe a similar phenomena in the interactions between realization and story content from our subjective experimentation: we cannot look at the realizations alone, but must take context and lexical realizations into consideration. we posit this is also why the pragmatic features are not particularly informative either. the pragmatic* features do not perform any better than the binary pragmatic features and are excluded from the tables. when all features sets are combined together, they achieve a worse performance than the otherwise best performing feature sets, “prag-punct-narr” for smo, and “punct” for the decision tree. all models and feature sets are better at predicting the not-selected class (highest f-measure is 0.80 in table 27), whereas selected is more difficult to predict (highest f-measure is 0.55 in 73 lukin and walker table 27). we posit that because the n-gram, pragmatic, and punct features are lexical, it is easier to categorize which lexicalizations are not preferred (e.g., the presence of the emphasizer basically, which had a catratio of 0, was never selected). however, of the lexical features that remain, it is more difficult to predict which would be selected. even though these results seem high, the most informative features did not reveal general insights. for example, the most informative positive feature for the smo classifier is “door” which is unique to a particular story in the training set, suggestive of overfitting. the decision tree reveals similar insights with a very deep and narrow tree. feature set precision recall f-measure sel. ¬sel. avg. sel. ¬sel. avg. sel. ¬sel. avg. punct 0.64 0.86 0.75 0.64 0.86 0.75 0.64 0.86 0.75 table 28: testing classification with j48 in weka we used the j48 prediction model with the best feature set, the simple “punct”, on the test set of the remaining 40 sentences. the results are presented in table 28 and show overall improvement to all the metrics. however, this best performing feature set only measures the appearance and type of punctuation generated in a sentence and leave much still to be understood about the features in this classification task. these results indicate that punctuation is key in determining which sentences will be selected or not, but this decision is the same function of the perstoryratio metric developed earlier, and thus this high-scoring feature set does not provide new insights. another interpretation is that the features sets developed for this classification task were not able to truly capture what makes a sentence appealing enough to be selected by an annotator, as seen by the wide the range of scores from the training metrics. a final interpretation is that there is simply not enough data or diversity for applying this type of model to this test set of create-your-own-stories for classification. 6. conclusion and future work this article has outlined requirements for a story planning and natural language generation storytelling system that bridges the four elements of the nlg story gap introduced in section 2. we bridge the gap by designing the fabula tales to automatically map from a sig (which can be obtained without domain knowledge) to a lexical-semantic representation (lsstree) compatible with a parameterized, narrative focused sentence planner. we train a narratological structurer from overgenerate and rank experimentation and observe trends from subject feedback. after the annotation of the sig, the remainder of the translation and generation is streamlined. evaluation has shown that while the realizations produced with narrative variations are more effective than the baseline realizations, there is room for improvement with respect to fluency. additional rules and heuristics can supplement the sentence planner to take context into consideration. successful approaches have incorporated context with a rule-based approach, building on the senses in wordnet or verbnet, as well as a statistical approach for expressive generation (ahn et al., 2016; rieser and lemon, 2011; paiva and evans, 2004; langkilde, 2000; rowe et al., 2008; mairesse and walker, 2011). fabula tales currently does not examine domain or story specific knowledge during deaggregation or discourse structuring, but we posit that a closer examination of ontologies can be used to learn domain specific information from each story to influence its retelling. this would 74 a narrative sentence planner and structurer for storytelling allow, for example, the same content to be removed if it could be easily inferred from the prior context. fabula tales requires manual annotation for the creation of sigs, but we have shown that the scheherazade tool is intuitive and lightweight. while some stories are more difficult than others to annotate, the guidelines we have adopted show strategies for encoding these interpretations. utilizing sigs affords manipulation of many aspects of the narrative that we have not yet explored, including time variations beyond the straightforward temporal ordering. bae et al. (2011) implement a computational model of focalization in generating narratives, where a planning-based generation engine identifies which events or inner thoughts characters are aware at the moment. these in turn, affect the selection of the events the planner selects to tell. for example, stories with surprise endings are generated by exploiting the disparity of knowledge between a story’s reader and its characters (bae and young, 2009). other work describes narrative systems that avoid conflict or make assumptions about its structure, or rely on humans to author it, creating a system based on character worlds, plans, their intentionality, and goals (ware and young, 2011; ware et al., 2014; ware and young, 2012). characters’ plans, intentions, and inner thoughts are annotated in the interpretation and affectual layers of the sig, which we posit can be used to explore these narrative aspects. furthermore, we have assumed that a single sig represents the whole of the fabula. future work may examine overlap of sigs derived from the same story, such as the collection of fables that have multiple sigs in the dramabank, and explore how to combine information from different sigs and create a model for focalization. there are several promising lines of work that can be explored with our ability to now bridge the nlg story gap for integrated applications. personalization can lead to even more engagement, and especially by coordinating gestures with speech of embodied virtual agents to increase the naturalness of human-like communication (hu et al., 2015; bergmann et al., 2013; wang and neff, 2013). the fabula tales framework could be integrated into a virtual agent environment to provide a plethora of stories for the agent to tell; as such, the affordances of the sentence-level variations is taken advantage of by rendering dialogue between two virtual agents (hu et al., 2016). here also, the gesture generation would benefit from the integration of story-level interactions. this nlg storytelling framework could also be used to enhance narrative systems, as well as make them more customized to the user, especially in simulated scenarios (johnson et al., 2004; aylett et al., 2005). stories written by real people could be generated from different perspectives to further explore perceptions and empathy by changing point of view parameter and pragmatic markers. another application area is narratively structured computer games. dialogue authoring in large games requires not only the creation of new content, but the subtlety of its delivery as it varies from character to character. when creating a replayable and adaptable system, it is important to have believable interactions with non-player characters (npcs) in the game world. short of hand authoring every possible character utterance, we ask “can npcs be given a personality fitted to the player?” and “upon replay, do the npcs utterances change under the assumption that the story is the same?” expanding npc dialogue generation was explored in lukin et al. (2014), but extending this work requires a fully immersive game world, characters, and improvement to fabula tales’ storylevel planner in order to show its effectiveness. along these lines, dialogue variants corresponding to in-game regional dialects could be modeled. some work has made great strides toward richer modeling of social-group membership for virtual characters (harrell et al., 2014; walker et al., 2013), and the ability to automatically produce linguistic variation according to such models would greatly enhance the impact of the systems. 75 lukin and walker 7. acknowledgments this work was partially funded by an arcs foundation scholarship, a nuance communications fellowship, and a university of california, santa cruz baskin family fellowship. research was partially sponsored by the army research laboratory and was accomplished under cooperative agreement number w911nf-17-2-0064. the views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the army research laboratory or the u.s. government. the u.s. government is authorized to reproduce and distribute reprints for government purposes notwithstanding any copyright notation herein. we thank the anonymous reviewers for their constructive and helpful feedback. appendix a. statistical tests significance tests for sections 4.4.1 and 4.4.2. style 1. original 2. 1st-out 3. 1st-neutr 4. 1st-shy 5. scheherazade 1. original – 2. 1st-out -4.00*** – 3. 1st-neutr -7.09**** -1.63 – 4. 1st-shy -8.02**** 8.93**** 3.72*** – 5. scheherazade 14.35**** 5.61**** 7.32**** 6.03*** – 6. 3rd-neutr -13.97**** 9.50**** 8.30**** 3.33*** -0.31 * p < 0.05; ** p < 0.01; *** p < 0.001; **** p < 0.0001 table 29: point of view and voice t-values for engagement (df = 95) style 1. original 2. 1st-out 3. 1st-neutr 4. 1st-shy 5. scheherazade 1. original – 2. 1st-out 5.59**** – 3. 1st-neutr 6.05**** > 1 – 4. 1st-shy 6.90**** 1.51 2.20* – 5. scheherazade 14.46**** 7.28**** 7.58**** 6.16**** – 6. 3rd-neutr 14.24**** 7.75**** 8.34**** 6.89**** 0.54 * p < 0.05; ** p < 0.01; *** p < 0.001; **** p < 0.0001 table 30: point of view and voice t-values for interest (df = 93) 76 a narrative sentence planner and structurer for storytelling realization 1. original 2. sosn 3 becausens 4 becausesn 5. ns 6. none 1. original – 2. sosn 2.6** – 3. becausens 3.2*** -1.1 – 4. becausesn 3.1** -1.3* 0.1* – 5. ns 4.3**** -3.0** -2.2* -2.2* – 6. none 4.9**** -3.0** -2.1** -2.4 ** -0.2 – 7. n 7.1**** -3.7*** -3.1*** -3.1 ** 1.6 1.5 * p < 0.05; ** p < 0.01; *** p < 0.001; **** p < 0.0001 table 31: deaggregation and discourse structuring t-values for correctness, exp. 1 (df = 101) realization 1. original 2. sosn 3 becausens 4 becausesn 5. ns 6. none 1. original – 2. sosn -4.1**** – 3. becausens -5.4**** -2.4** – 4. becausesn -5.6**** -2.7** -0.5 – 5. ns -7.6**** 5.6**** 3.3*** 3.1** – 6. none -9.4**** 7.5**** 5.7**** 5.4**** 2.7** – 7. n -10.8**** -6.1**** -4.3**** -3.8*** -1.7* 0.3 * p < 0.05; ** p < 0.01; *** p < 0.001; **** p < 0.0001 table 32: deaggregation and discourse structuring t-values for preference, exp. 1 (df = 106) realization 1. original 2. sosn 1. original – 2. sosn < 1**** – 3. scheherazade < 1**** < 1**** **** p < 0.0001 table 33: deaggregation and discourse structuring t-values for correctness, exp. 2 (df = 95) realization 1. original 2. sosn 1. original – 2. sosn < 1**** – 3. scheherazade < 1**** < 1**** **** p < 0.0001 table 34: deaggregation and discourse structuring t-values for preference, exp. 2 (df = 95) 77 lukin and walker references h. porter abbott. the cambridge introduction to narrative. cambridge university press, 2008. emily ahn, fabrizio morbini, and andrew s gordon. improving fluency in narrative text generation with grammatical transformations and probabilistic parsing. in in proc. of the 9th international natural language generation conference, page 70, 2016. ruth s aylett, sandy louchart, joao dias, ana paiva, and marco vala. fearnot!–an experiment in emergent narrative. in proc. of intelligent virtual agents, pages 305–316. springer, 2005. byung-chull bae and r michael young. suspense? surprise! or how to generate stories with surprise endings by exploiting the disparity of knowledge between a story’s reader and its characters. in proc. of interactive storytelling, pages 304–307. springer berlin heidelberg, 2009. byung-chull bae, yun-gyung cheong, and r. michael young. toward a computational model of focalization in narrative. in proc. of the 6th international conference on foundations of digital games, pages 313–315. acm, 2011. mieke bal. narratology. introduction to the theory of narrative. 1997. camille barot, colin m potts, and r michael young. a tripartite plan-based model of narrative for narrative discourse generation. in proc. of the joint workshop on intelligent narrative technologies and social believability in games at the 11th aaai conference on artificial intelligence and interactive digital entertainment, pages 2–8, 2015. anja belz and ehud reiter. comparing automatic and human evaluation of nlg systems. in proc. of the 11th conference of the european chapter of the association for computational linguistics, 2006. kirsten bergmann, sebastian kahl, and stefan kopp. modeling the semantic coordination of speech and gesture under cognitive and linguistic constraints. in proc. of intelligent virtual agents, pages 203–216. springer, 2013. douglas biber. variation across speech and writing. cambridge university press, 1991. nadjet bouayad-agha, donia r scott, and richard power. integrating content and style in documents: a case study of patient information leaflets. information design journal, 9(2-3): 161–176, 1998. jerome bruner. the narrative construction of reality. critical inquiry, 18:1–21, 1991. kevin burton, akshay java, ian soboroff, et al. the icwsm 2009 spinn3r dataset. in proc. of the third annual conference on weblogs and social media, 2009. lynne cahill, john carroll, roger evans, daniel paiva, richard power, donia scott, and kees van deemter. from rags to riches: exploiting the potential of a flexible generation architecture. in proc. of the 39th annual meeting on association for computational linguistics, pages 106– 113. association for computational linguistics, 2001. 78 a narrative sentence planner and structurer for storytelling charles b callaway and james c lester. narrative prose generation. artificial intelligence, 139 (2):213–252, 2002. daniel cer, mona diab, eneko agirre, inigo lopez-gazpio, and lucia specia. semeval-2017 task 1: semantic textual similarity-multilingual and cross-lingual focused evaluation. 2017. raman chandrasekar and bangalore srinivas. automatic induction of rules for text simplification. knowledge-based systems, 10(3):183–190, 1997. yun-gyung cheong and r michael young. narrative generation for suspense: modeling and evaluation. in proc. of the international conference on interactive digital storytelling, 2008. eugenio concepción, pablo gervás, gonzalo méndez, and carlos león. using cnl for knowledge elicitation and exchange across story generation systems. in proc. of the international workshop on controlled natural language, pages 81–91. springer, 2016a. eugenio concepción, gonzalo mendez, and pablo gervás. mining knowledge in storytelling systems for narrative generation. in proc. of the workshop on computational creativity in natural language generation at the international natual langauge generation conference, pages 41– 50, 2016b. eugenio concepción, gonzalo méndez, pablo gervás, and carlos león. a challenge proposal for narrative generation using cnls. in proc. of the 9th international natural language generation conference, pages 171–173, 2016c. michael denkowski and alon lavie. meteor universal: language specific translation evaluation for any target language. in proc. of the ninth workshop on statistical machine translation, pages 376–380, 2014. george doddington. automatic evaluation of machine translation quality using n-gram cooccurrence statistics. in proc. of the second international conference on human language technology research, pages 138–145. morgan kaufmann publishers inc., 2002. michael elhadad and jacques robin. an overview of surge: a reusable comprehensive syntactic realization component. technical report, technical report 96-03, ben gurion university, dept. of computer science, beer sheva, israel, 1996. david elson. modeling narrative discourse. phd thesis, columbia university, dept. of computer science, 2012a. david k elson. dramabank: annotating agency in narrative discourse. in proc. of the 8th international conference on language resources and evaluation, 2012b. david k elson and kathleen r mckeown. a tool for deep semantic encoding of narrative texts. in proc. of the acl-ijcnlp 2009 software demonstrations, pages 9–12. association for computational linguistics, 2009. angela fan, mike lewis, and yann dauphin. hierarchical neural story generation. in proc. of the 56th annual meeting of the association for computational linguistics, pages 889–898. association for computational linguistics, 2018. 79 lukin and walker christiane fellbaum. wordnet: an electronic lexical database. wordnet is available from http://www. cogsci. princeton. edu/wn, 2010. katja filippova and michael strube. sentence fusion via dependency graph compression. in proc. of the conference on empirical methods in natural language processing, pages 177–185. association for computational linguistics, 2008. claire gardent and german kruszewski. generation for grammar engineering. in proc. of the seventh international natural language generation conference, pages 31–39. association for computational linguistics, 2012. gérard genette and jane e lewin. narrative discourse: an essay in method. cornell university press, 1983. richard j gerrig. experiencing narrative worlds: on the psychological activities of reading. yale university press, 1993. pablo gervás, belén dı́az-agudo, federico peinado, and raquel hervás. story plot generation based on cbr. in in proc. of applications and innovations in intelligent systems xii, pages 33–46. springer, 2005. pablo gervás, birte lönneker-rodman, jan christoph meister, and federico peinado. narrative models: narratology meets artificial intelligence. in proc. of the toward computational models of literary analysis work at the international conference on language resources and evaluation, pages 44–51, 2006. gregory grefenstette. producing intelligent telegraphic text reduction to provide an audio scanning service for the blind. in working notes of the aaai spring symposium on intelligent text summarization, pages 111–118, 1998. d fox harrell, dominic kao, chong-u lim, jason lipshin, ainsley sutherland, and julia makivic. the chimeria platform: an intelligent narrative system for modeling social identity-related experiences. in proc. of the seventh intelligent narrative technologies workshop, 2014. david howcroft, crystal nakatsu, and michael white. enhancing the expression of contrast in the sparky restaurant corpus. in proc. of the 14th european workshop on natural language generation, pages 30–39, 2013. chao hu, marilyn a walker, michael neff, and jean e fox tree. storytelling agents with personality and adaptivity. in proc. of the international conference on intelligent virtual agents, pages 181–193. springer, 2015. zhichao hu, elahe rahimtoroghi, larissa munishkina, reid swanson, and marilyn a walker. unsupervised induction of contingent event pairs from film scenes. in proc. of the 2013 conference on empirical methods in natural language processing, pages 369–379, 2013. zhichao hu, michelle dick, chung-ning chang, kevin bowden, michael neff, jean e fox tree, and marilyn a walker. a corpus of gesture-annotated dialogs for monologue-to-dialogue generation from personal narratives. in proc. of the 10th international conference on language resources and evaluation, 2016. 80 a narrative sentence planner and structurer for storytelling ting-hao kenneth huang, francis ferraro, nasrin mostafazadeh, ishan misra, aishwarya agrawal, jacob devlin, ross girshick, xiaodong he, pushmeet kohli, dhruv batra, et al. visual storytelling. in proc. of the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 1233–1239, 2016. diana zaiu inkpen and graeme hirst. near-synonym choice in natural language generation. in recent advances in natural language processing, volume 3, pages 141–152, 2004. amy isard, carsten brockmann, and jon oberlander. individuality and alignment in generated dialogues. proc. of the 4th international natural language generation conference, 2006. w lewis johnson, carole beal, anna fowles-winkler, ursula lauper, stacy marsella, shrikanth narayanan, dimitra papachristou, and hannes vilhjálmsson. tactical language training system: an interim report. in proc. of intelligent tutoring systems, pages 336–345. springer, 2004. karin kipper, anna korhonen, neville ryant, and martha palmer. extensive classifications of english verbs. in proc. of the 12th euralex international congress, pages 1–15, 2006. kevin knight and daniel marcu. statistics-based summarization-step one: sentence compression. proc. of aaai/iaai, 2000:703–710, 2000. kevin knight and daniel marcu. summarization beyond sentence extraction: a probabilistic approach to sentence compression. artificial intelligence, 139(1):91–107, 2002. william labov and joshua waletzky. narrative analysis: oral versions of personal experience. 1997. irene langkilde. forest-based statistical sentence generation. in proc. of the 1st north american chapter of the association for computational linguistics conference, pages 170–177. association for computational linguistics, 2000. irene langkilde and kevin knight. generation that exploits corpus-based statistical knowledge. in proc. of the 36th annual meeting of the association for computational linguistics and 17th international conference on computational linguistics-volume 1, pages 704–710. association for computational linguistics, 1998. irene langkilde-geary. an empirical verification of coverage and correctness for a generalpurpose sentence generator. in proc. of the 12th international natural language generation workshop, pages 17–24, 2002. benoit lavoie and owen rambow. a fast and portable realizer for text generation systems. in proc. of the fifth conference on applied natural language processing, pages 265–268. association for computational linguistics, 1997. michael lebowitz. creating a story-telling universe. in proc. of international joint conferences on artificial intelligence, pages 63–65. citeseer, 1983. boyang li. learning knowledge to support domain-independent narrative intelligence. phd thesis, georgia institute of technology, 2015. 81 lukin and walker boyang li, stephen lee-urban, george johnston, and mark o riedl. story generation with crowdsourced plot graphs. in proc. of the 27th aaai conferece on artificial intelligence, 2013. chin-yew lin. rouge: a package for automatic evaluation of summaries. 2004. grace lin. character modeling through dialogue for expressive natural language generation. phd thesis, university of california, santa cruz, dept. of computer science, 2016. birte lönneker. narratological knowledge for natural language generation. in proc. of the 10th european workshop on natural language generation, pages 91–100. citeseer, 2005. stephanie lukin, reginald hobbs, and clare voss. a pipeline for creative visual storytelling. in proc. of the first workshop on storytelling, pages 20–32, 2018. stephanie m lukin and marilyn a walker. narrative variations in a virtual storyteller. in proc. of intelligent virtual agents, pages 320–331. springer, 2015. stephanie m lukin, james o ryan, and marilyn a walker. automating direct speech variations in stories and games. in proc. of the workshop on games and nlp at the tenth artificial intelligence and interactive digital entertainment conference, 2014. stephanie m lukin, lena i reed, and marilyn a walker. generating sentence planning variations for story telling. in proc of the 16th annual meeting of the special interest group on discourse and dialogue, page 188, 2015. stephanie m. lukin, kevin bowden, casey barackman, and marilyn a. walker. personabank: a corpus of personal narratives and their story intention graphs. in proc. of the 10th international conference on language resources and evaluation, 2016. matt madden. 99 ways to tell a story. random house, 2006. f. mairesse and m.a. walker. towards personality-based user adaptation: psychologically informed stylistic language generation. user modeling and user-adapted interaction, pages 1–52, 2010a. issn 0924-1868. françois mairesse and marilyn a walker. a personality-based framework for utterance generation in dialogue applications. in aaai spring symposium: emotion, personality, and social behavior, pages 80–87, 2008. françois mairesse and marilyn a walker. towards personality-based user adaptation: psychologically informed stylistic language generation. user modeling and user-adapted interaction, 20(3):227–278, 2010b. françois mairesse and marilyn a walker. controlling user perceptions of linguistic style: trainable generation of personality traits. computational linguistics, 37(3):455–488, 2011. william c mann and sandra a thompson. rhetorical structure theory: toward a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8(3): 243–281, 1988. 82 a narrative sentence planner and structurer for storytelling erwin marsi and emiel krahmer. explorations in sentence fusion. in proc. of the tenth european workshop on natural language generation (enlg-05), 2005. michael mateas. a preliminary poetics for interactive drama and games. digital creativity, 12 (3):140–152, 2001. dan p mcadams, ruthellen ed josselson, and amia ed lieblich. identity and story: creating self in narrative. american psychological association, 2006. james r meehan. tale-spin, an interactive program that writes stories. in proc. of the international joint conferences on artificial intelligence, volume 77, pages 91–98. citeseer, 1977. igor a. mel’čuk. dependency syntax: theory and practice. suny press, 1988. nick montfort. generating narrative variation in interactive fiction. phd thesis, university of pennsylvania, dept. of computer and information science, 2007. nick montfort. curveship: an interactive fiction system for interactive narrating. in proc. of the workshop on computational approaches to linguistic creativity, pages 55–62. association for computational linguistics, 2009. nasrin mostafazadeh, michael roth, annie louis, nathanael chambers, and james allen. lsdsem 2017 shared task: the story cloze test. in proc. of the 2nd workshop on linking models of lexical, sentential and discourse-level semantics, pages 46–51, 2017. larissa munishkina, jennifer parrish, and marilyn a walker. fully-automatic interactive story design from film scripts. in proc. of the international conference on interactive digital storytelling, pages 229–232. springer, 2013. shashi narayan, claire gardent, shay cohen, and anastasia shimorina. split and rephrase. in proc. of the conference on empirical methods in natural language processing, pages 617–627, 2017. james niehaus and r michael young. a computational model of inferencing in narrative. in aaai spring symposium: intelligent narrative technologies ii, pages 83–90, 2009. jekaterina novikova, ondřej dušek, amanda cercas curry, and verena rieser. why we need new evaluation metrics for nlg. in proc. of the 2017 conference on empirical methods in natural language processing, pages 2241–2252, 2017. shereen oraby, lena reed, shubhangi tandon, ts sharath, stephanie lukin, and marilyn walker. controlling personality-based stylistic variation with neural natural language generators. in proc. of the special interest group on discourse and dialogue, 2018. daniel s. paiva and roger evans. a framework for stylistically controlled generation. in anja belz, roger evans, and paul piwek, editors, natural language generation, third internatonal conference, inlg 2004, number 3123 in lnai, pages 120–129. springer, july 2004. alan palmer. universal minds. semiotica, 2007(165):205–225, 2007. 83 lukin and walker kishore papineni, salim roukos, todd ward, and wei-jing zhu. bleu: a method for automatic evaluation of machine translation. in proc. of the 40th annual meeting on association for computational linguistics, pages 311–318. association for computational linguistics, 2002. cécile paris and donia scott. stylistic variation in multilingual instructions. in proc. of the seventh international workshop on natural language generation, pages 45–52. association for computational linguistics, 1994. federico peinado and pablo gervás. evaluation of automatic generation of basic stories. new generation computing, 24(3):289–302, 2006. nanyun peng, marjan ghazvininejad, jonathan may, and kevin knight. towards controllable story generation. in proc. of the first workshop on storytelling, pages 43–49, 2018. manon penning and mariët theune. cueing the virtual storyteller: analysis of cue phrase usage in fairy tales. in proc. of the eleventh european workshop on natural language generation, pages 159–162. association for computational linguistics, 2007. david pizarro, eric uhlmann, and peter salovey. asymmetry in judgments of moral blame and praise the role of perceived metadesires. psychological science, 14(3):267–272, 2003. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proc. of 6th international conference on language resources and evaluation, 2008. gerald prince. a grammar of stories: an introduction, volume 13. walter de gruyter, 1974. vladimir propp. morphology of the folktale. university of texas press, second edition, 1969. raymond queneau and barbara wright. exercises in style, volume 513. new directions publishing, 1981. elahe rahimtoroghi, reid swanson, marilyn a walker, and thomas corcoran. evaluation, orientation, and action in interactive storytelling. in proc. of intelligent narrative technologies workshop, volume 6, 2013. elahe rahimtoroghi, thomas corcoran, reid swanson, marilyn a walker, kenji sagae, and andrew gordon. minimal narrative annotation schemes and their applications. in proc. of the intelligent narrative technologies workshop, 2014. elahe rahimtoroghi, ernesto hernandez, and marilyn walker. learning fine-grained knowledge about contingent relations between everyday events. in proc. of the 17th annual meeting of the special interest group on discourse and dialogue, pages 350–359, 2016. aaron a reed, ben samuel, anne sullivan, ricky grant, april grow, justin lazaro, jennifer mahal, sri kurniawan, marilyn a walker, and noah wardrip-fruin. a step towards the future of role-playing games: the spyfeet mobile rpg project. in proc. of the conference on artificial intelligence and interactive digital entertainment conference, 2011. mark o riedl and robert michael young. narrative planning: balancing plot and character. journal of artificial intelligence research, 39(1):217–268, 2010. 84 a narrative sentence planner and structurer for storytelling mark owen riedl and r michael young. an intent-driven planner for multi-agent story generation. in proc. of the third international joint conference on autonomous agents and multiagent systems-volume 1, pages 186–193. ieee computer society, 2004. verena rieser and oliver lemon. reinforcement learning for adaptive dialogue systems: a datadriven methodology for dialogue management and natural language generation. springer science & business media, 2011. stefan riezler, tracy h king, richard crouch, and annie zaenen. statistical sentence condensation using ambiguity packing and stochastic disambiguation methods for lexical-functional grammar. in proc. of the 2003 conference of the north american chapter of the association for computational linguistics on human language technology, pages 118–125. association for computational linguistics, 2003. elena rishes, stephanie m lukin, david k elson, and marilyn a walker. generating different story tellings from semantic representations of narrative. in proc. of interactive storytelling, pages 192–204. springer international publishing, 2013. melissa roemmele. neural networks for narrative continuation. phd thesis, ph. d. dissertation, university of southern california, 2018. melissa roemmele, cosmin adrian bejan, and andrew s gordon. choice of plausible alternatives: an evaluation of commonsense causal reasoning. in aaai spring symposium: logical formalizations of commonsense reasoning, pages 90–95, 2011. jonathan p rowe, eun young ha, and james c lester. archetype-driven character dialogue generation for interactive narrative. in proc. of intelligent virtual agents, pages 45–58. springer, 2008. donia scott and clarisse sieckenius de souza. getting the message across in rst-based text generation. current research in natural language generation, 4:47–73, 1990. rishi sharma, james allen, omid bakhshandeh, and nasrin mostafazadeh. tackling the story ending biases in the story cloze test. in proc. of the 56th annual meeting of the association for computational linguistics, pages 752–757. association for computational linguistics, 2018. siddarth srinivasan, richa arora, and mark riedl. a simple and effective approach to the story cloze test. in proc. of the 2018 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 92–96. association for computational linguistics, 2018. reid swanson and andrew s gordon. say anything: a massively collaborative open domain story writing companion. in proc. of interactive storytelling, pages 32–40. springer, 2008. mariët theune, sander rensen, rieks op den akker, dirk heylen, and anton nijholt. emotional characters for automatic plot creation. in proc. of the international conference on technologies for interactive digital storytelling and entertainment, pages 95–100. springer, 2004. 85 lukin and walker mariët theune, nanda slabbers, and feikje hielkema. the narrator: nlg for digital storytelling. in proc. of the eleventh european workshop on natural language generation, pages 109–112. association for computational linguistics, 2007. avril thorne. the press of personality: a study of conversations between introverts and extraverts. journal of personality and social psychology, 53(4):718, 1987. avril thorne and kate c mclean. telling traumatic events in adolescence: a study of master narrative positioning. connecting culture and memory: the development of an autobiographical self, pages 169–185, 2003. scott r turner. minstrel: a computer model of creativity and storytelling. 1993. marilyn walker, owen rambow, and monica rogati. training a sentence planner for spoken dialogue using boosting. computer speech and language: special issue on spoken language generation, 16(3-4):409–433, 2002. marilyn a walker, amanda stent, françois mairesse, and rashmi prasad. individual and domain adaptation in sentence planning for dialogue. journal of artificial intelligence research, 30: 413–456, 2007. marilyn a walker, jennifer sawyer, carolynn jimenez, elena rishes, grace i lin, zhichao hu, jane pinckard, and noah wardrip-fruin. using expressive language generation to increase authorial leverage. in proc. of the ninth artificial intelligence and interactive digital entertainment conference, 2013. xin wang, wenhu chen, yuan-fang wang, and william yang wang. no metrics are perfect: adversarial reward learning for visual storytelling. in proc. of the 56th annual meeting of the association for computational linguistics, pages 899–909. association for computational linguistics, 2018. yingying wang and michael neff. the influence of prosody on the requirements for gesture-text alignment. in proc. of intelligent virtual agents, pages 180–188. springer, 2013. stephen g ware and r michael young. validating a plan-based model of narrative conflict. in proc. of the international conference on the foundations of digital games, pages 220–227. acm, 2012. stephen g. ware and robert michael young. cpocl: a narrative planner supporting conflict. in proc. of the artificial intelligence and interactive digital entertainment conference, 2011. stephen g ware, r michael young, brent harrison, and david l roberts. a computational model of plan-based narrative conflict at the fabula level. ieee transactions on computational intelligence and ai in games, 6(3):271–288, 2014. david winer and r michael young. discourse-driven narrative generation with bipartite planning. in proc. of the 9th international natural language generation conference, pages 11–20, 2016. 86 introduction related work narrative variation in storytelling systems foundations of the fabula tales storyteller bridging the nlg story gap story intention graphs: intermediary fabula lexical-semantic story tree translation text realization from lexical-semantic story trees evaluating the bridging process narrative sentence planning mood: point of view mood: distance (direct speech) voice: person evaluation of sentence-level variations point of view and voice: engagement and interest discourse structuring: correctness and preference narratological structurer overgenerate: generating training data rank: training the narratological structurer evaluation of narratological structurer predicting selected sentences conclusion and future work acknowledgments statistical tests dd-final.dvi dialogue and discourse 4(2) (2013) 118–141 doi: 10.5087/dad.2013.206 identifying “aboutness topics”: two annotation experiments∗ philippa cook philippa.cook@fu-berlin.de institut für deutsche und niederländische philologie freie universität berlin 14195 berlin, germany felix bildhauer felix.bildhauer@fu-berlin.de institut für deutsche und niederländische philologie freie universität berlin 14195 berlin, germany editor: stefanie dipper, heike zinsmeister, bonnie webber abstract this paper deals with the annotation of “aboutness topic” (also known as “sentence topic”) in naturally occurring data. we report on two annotation experiments involving german newspaper texts: in experiment 1, based on the annotation guidelines by götze et al. (2007), two expert annotators had to select the aboutness topic from among a small number of pre-defined choices, for a total of 588 sentence tokens. although the results are disappointing in terms of inter-annotator agreement (with fleiss’ κ between .19 and .57, depending on sentence category), they allowed us to identify and analyze cases that resist easy classification. next, the original guidelines were tentatively modified, capturing insights from experiment 1. in a second experiment, based on the revised guidelines and following several training sessions, four raters annotated a total of 56 sentences, deciding for each np whether or not it is an aboutness topic. overall, the results are still disappointing from an inter-rater agreement perspective, but again, a closer look at the problematic cases is instructive as to the weaknesses of the current operationalizations of information structural notions and, perhaps most importantly, it reveals a number of basic theoretical questions that still await satisfactory answers. keywords: information structure, topic, focus, thetic, annotation, inter-rater agreement 1. introduction research on information structure (is) may serve a twofold purpose: first, information structure constitutes an intriguing area of investigation in its own right, although numerous concepts and their interrelations are nevertheless still in need of further refinement. second, insights in this field may lead to promising (re-)analyses of linguistic phenomena on the basis of information structure, that is, using information structural constraints in describing phenomena previously accounted for in terms of syntax (e. g. de kuthy, 2002; cook and payne, 2006; ambridge and goldberg, 2008; cook and ørsnes, 2010). in this respect, corpora annotated for information structure are particularly valuable as they put one in a position to test linguistic analyses that are based on notions such as ∗. we would like to thank julia ritz for providing the text material used in experiment #2 and for helping us with the annotator training. we are also grateful for the constructive comments by the anonymous reviewers. all remaining shortcoming are of course ours. the research reported in this paper was funded by the deutsche forschungsgemeinschaft, sfb 632 “information structure”, project a6. c©2013 philippa cook and felix bildhauer submitted 04/12; accepted 12/12; published online 06/13 identifying “aboutness topics” “topic”, “focus” and “givenness”. however, not only are these notions used in different ways across different currents of research but they also cause considerable problems when applied to naturally occurring data by researchers who otherwise largely agree on the definitions of these concepts and who even adhere to the same set of annotation guidelines. in the present paper, we will take a closer look at the annotation of “aboutness topics” (also known as “sentence topics”) in naturally occurring data. the data used in the annotation experiments were extracted from the german reference corpus dereko (institut für deutsche sprache, 2009) (first experiment), as well as from the tübinger baumbank des deutschen / zeitungskorpus (tüba-d/z)1 (second experiment). section 2 outlines the criteria used by the annotators of the first experiment for identifying aboutness topics and relates them to alternative approaches to topic-hood. section 3 reports the first experiment, the results of which are discussed in section 3.5, where we identify the type of data that turns out to be particularly difficult to assess and seek to establish exactly which features are involved in these cases and why they give rise to diverging annotations. on the basis of these results, we finally modify the annotation guidelines used in experiment #1. section 4 reports on a second experiment that was carried out using the revised version of the guidelines and discusses its outcome. in section 5, we raise several questions relating to both the theoretical underpinnings of the guidelines and to the more practical issue of how the guidelines may be further improved by making them more specific about particular phenomena. 2. “topic” in theory and in the annotation guidelines the notion of “topic” we are dealing with in the studies being reported on here is that of “aboutness topic” (at). since there are considerable differences in the way in which the notion of “topic” has been used, and since the actual operationalizability of this notion is the crux of the current contribution, we will lay out here some of the basic assumptions taken by researchers working with the notion of “aboutness topic”. krifka (2007), in his concise overview of the basic notions of information structure, points out that the use of the terms “topic” and “comment” reflect what von der gabelentz (1869) called “psychological subject’ and “psychological predicate” respectively, that is “the entity that a speaker identifies about which then information, the comment, is given” (krifka, 2007, 40). this approach to topicality was further elaborated by reinhart (1981), who adopts stalnaker’s (1978) notion of “context set” (a set of propositions which interlocutors accept to be true; that is, a “common ground”). in addition, reinhart assumes that the “common ground” is structured in such a way that information is stored in terms of a pairing of an entity and a proposition (or set of propositions) about that entity. new information is added to the “common ground” in the form of structured propositions, where the “sentence topic” designates an entity and the remainder of the sentence contributes the information to be associated with that entity (just like information in a file-card system is stored on a certain file card bearing a heading).2 building on reinhart’s approach, krifka (2007, 41) formulates the following definition: (1) the topic constituent identifies the entity or set of entities under which the information expressed in the comment constituent should be stored in the c[common] g[round] content. 1. http://www.sfs.uni-tuebingen.de/tuebadz.shtml 2. the file-card metaphor has since been used by a number of authors. for a critical evaluation, see e. g. hendriks and dekker (1996). 119 cook/bildhauer the notions of “topic” and “comment” have sometimes been mixed up with the notions of “background” and “focus” such that, for instance, focus is believed to be the complement of topic. the reason for this mixing of dimensions is presumably due to the fact that topics are in practice prototypically discourse-given whereas, in contrast, foci are canonically new. thus there seems to be a simple dichotomy in which newness and givenness align independently with focus and topic respectively. such a merging of the dimensions is, however, problematic because there are cases which deviate from the canonical alignment in that (i) there are topics which contain a focus, viz. e. g. (2b) below. such examples involving so-called contrastive topics can, but do not have to, involve an aboutness topic. rather, the unifying feature of so-called contrastive topics is their function in discourse, where they are assumed to indicate a discourse strategy (roberts, 1996; büring, 2003; krifka, 2007). the next problematic case is (ii) that there do appear to be new (i. e. non-given) topics as in (3) where an entity is introduced as new into the discourse (a good friend of mine) but still serves as aboutness topic (although see the discussion in section 3.5.2 below since this possibility is not wholly uncontroversial). finally, there is the possibility that the comment is not always identical to the focus, i. e. the focus may only be a sub-part of the comment as shown in (4b) below. finally, there is also the possibility of non-new foci, as in (5b) (viz. the discussion of second-occurrence foci; partee, 1999). (2) a. what do your siblings do? b. [my [sister]foc ]top [studies medicine]foc (krifka, 2007, 44) (3) [a good friend of mine]top [married britney spears last year]comment (krifka, 2007, 42) (4) a. when did [aristotle onassis]top marry jacqueline kennedy? b. [he]top [married her [in 1968]foc ]comment (krifka, 2007, 42) (5) a. everyone already knew that mary only eats [vegetables]foc b. if even [paul]foc knew that mary only eats [vegetables]sof , then he should have suggested a different restaurant. (partee, 1999, 216) thus, the possibility of such non-canonical alignments (e. g. non-given topics, non-new foci) must be accommodated in a model of information structure. we have chosen to follow the multipartitioning approach espoused by krifka which assumes both a topic/comment and an orthogonal focus/background partition in order to be able to do justice to the non-canonical as well as canonical pairings. the characterization of “aboutness topic” that we adopt is also distinct from vallduví’s (1992) “link”, which is defined positionally as the sentence-initial topic. further, since the focus/background partition is independent of the topic/comment partition, an aboutness topic can in principle be identical to a focus of an utterance (though it is unclear whether or not cases other than those in (2b) should be allowed by the theory; see section 3.5.2). under krifka’s approach that we adopt here, a sentence typically has only one aboutness topic (while reinhart, 1981, p. 86 fn. 17, excludes altogether the possibility that a sentence has more than one aboutness topic). sentences which lack a topic – or perhaps more precisely, a topic/comment articulation – are classed as topicless/thetic (cf. krifka, 2007, 43). we will have more to say about thetic utterances in general and about the distinction between thetic vs. topic-comment utterances in particular in section 3.5 below. the guidelines for the annotation of information structure (götze et al., 2007), which were produced by the collaborative research cluster (sfb) 632, and which closely mirror the proposals 120 identifying “aboutness topics” of krifka (2007), provide instructions for the annotation of information status (or “givenness”), topic, and focus. under the notion of “topic”, both “aboutness topic” and “framesetting topic” are identified. it is the former that concerns us here (see krifka, 2007, for a more detailed discussion of frame-setting). götze et al. (2007, 165) offer the following tests for determining the aboutness topic of an utterance: (6) an np x is the aboutness topic of a sentence s containing x if a. s would be a natural continuation to the announcement let me tell you something about x b. s would be a good answer to the question what about x? c. s could be naturally transformed into the sentence concerning x, s’ where s’ differs from s only insofar as x has been replaced by a suitable pronoun. applying these diagnostics to naturally-occurring data is not without problems, as will become clear in the next section where we report on an annotation experiment based on the aforementioned criteria. motivated by the outcome of the first experiment, these criteria were then modified and the effects were assessed in a second annotation experiment, reported in section 4. 3. annotation experiment #1 3.1 method in the first annotation experiment, the annotators had to decide for a number of target sentences which constituent, if any, is the aboutness topic, basing their decision on the annotation guidelines of götze et al. (2007). for each sentence, the annotators made a choice from a small number of pre-selected topic-candidates (such as the subject-xp, object-xps and deictic adverbial expressions such as hier ‘here’, dann ‘then’). 3.2 materials all target sentences were sampled, along with some preceding and following context, from the dereko corpus (institut für deutsche sprache, 2009) (the vast majority of the material in dereko consists of newspaper texts). they were selected such that every sentence contained one of the four verbs shown in table 1, instantiating a specific argument frame (a subject-xp is taken for granted in each case and therefore not listed explicitly).3 after discarding occurrences in questions, relative clauses, conditionals, titles and in the first sentence of quotations, a total of 588 sentence tokens (between 136 and 167 per verb) were included in the study. 3.3 subjects each target sentence was assessed by two independent coders (the authors of this paper). the annotators were familiar with concepts of information structure beyond the explanations in the annotation guidelines that were used (this is important because differences in annotation behaviour 3. the choice of these four verbs was driven by the requirements of a pilot study in which we investigated the preferred realization of the sentence topic in connection with non-agentive (or not prototoypically agentive) subjects. 121 cook/bildhauer argument frame verb example ppmit xploc geraten er gerät [mit seiner hose] [in die kette]. ‘to get (caught in)’ ‘he got his trousers caught in the chain.’ xp ppauf reagieren sie reagiert [überrascht] [auf den vorschlag]. ‘to react’ ‘she reacted surprisedly to this suggestion.’ ppvon profitieren sie profitieren [von den steuersenkungen]. ‘to profit’ ‘they profit from the tax reductions.’ herrschen dort herrscht ruhe. ‘to reign’ ‘there reigns peace.’ table 1: verbs and argument frames aboutness topic thetic vs. topic-comment verb n % agr. fleiss’ κ p % agr. fleiss’ κ p profitieren 136 80.7 .57 < .001 95.6 −.02 .79 herrschen 138 73.9 .54 < .001 76.8 .51 < .001 geraten 147 67.3 .25 < .01 68.0 .26 < .01 reagieren 167 63.5 .19 < .001 88.6 .23 < .01 all sentences 588 82.3 .44 <.001 table 2: inter-rater agreement in experiment #1 could in principle be due to annotators having different underlying concepts of “topic”, even if they are using the same guidelines). 3.4 results for each one of the four verbs, we calculated fleiss’ κ (fleiss, 1971) as a measure of inter-rater agreement on the choice of the aboutness topic, as shown in table 2.4 in addition, we calculated a κ-value for the more coarse-grained distinction between thetic (topicless) and topic-comment sentences, both per-verb and for all four verbs taken together. all κ values for the choice of an aboutness topic indicate a statistically significant degree of agreement. that said, inter-rater agreement is highly variable across the four verbs, never exceeding κ = .57. in the case of reagieren, agreement does not even reach a “fair” level, following landis and koch’s (1977) interpretation of κ-values. turning to the distinction between thetic and topic-comment sentences, κ-values are even lower here although most of them are still statistically significant. the fact that agreement is basically at chance level in the case of profitieren raises the question whether the annotation guidelines can be interpreted in substantially different ways by different annotators (we will turn to this question in the discussion of the second experiment). 3.5 discussion in our view, the low level of inter-rater agreement observed in experiment #1 is quite unacceptable given that both annotators based their judgements on the same guidelines. on the one hand, it 4. we are using fleiss’ κ for m raters instead of cohen’s (1960) κ here because we want to be able to compare the figures to the outcome of the second annotation experiment, in which m = 4 raters participated. 122 identifying “aboutness topics” could be the case that the concept of “aboutness topic” simply cannot be operationalized in a way that allows annotators to consistently categorize naturally-occurring data. on the other hand, it is possible that the annotation guidelines and/or the underlying theory are not specific enough and thus leave too much room for interpretation. a good starting point to decide between these two alternatives then is to inspect more closely the tokens on which the annotators did not agree. in this regard, we could identify data that proved particularly difficult to assess. most of the controversial cases fall into one of the following categories: • problems in deciding whether the sentence has an aboutness topic at all, including cases where the status of potential topic expressions is unclear because the interaction between topic and focus (especially their overlapping) is not covered exhaustively in the guidelines (nor in the literature that we are aware of) • the annotators’ different interpretation of “aboutness”; most commonly, deciding “what the sentence is about” when there is more than one expression that could plausibly serve as the aboutness topic: in many cases, the diagnostics sketched in (6) do not yield an answer, or their application is not straightforward in what follows, section 3.5.1 will briefly illustrate a number of cases where the annotators did in fact agree, and section 3.5.2 will address examples from the two problematic categories listed above. 3.5.1 agreement examples (7)–(8) are typical of the cases in which the annotators agreed on the at of the sentence. (in addition to the critical data (b), we also provide some of the immediately preceding context in (a).) in both examples, a non-subject was chosen as the at. this is probably in part due to the fact that the subject, being a non-specific indefinite, is not suitable as an at (see endriss, 2009; götze et al., 2007). in addition, in terms of givenness, the referent of the non-subject is either “active” (as in (7b)), or “accessible” (as in (8b)), which are prototypical properties of aboutness topics (see section 2) above. (7) a. ein besonderer fall ist der sogenannte „promillewegi ”, der von rothenbach richtung brandscheid führt. ‘the so-called “promille-road”, leading from rothenbach to brandscheid, is a special case.’ b. auf on [dem the idyllisch picturesquely gelegenen situated wirtschaftswegi ]top farm road herrscht reigns nämlich actually emsiger active autobetrieb. through-traffic ‘the picturesque farm road is actually busy with through-traffic.’ (8) a. „es ist schön und lustig, aber die produktion eines solchen spiels ist teuer, lohnen sich denn überhaupt die kosten?” ‘ “it’s beautiful and funny, but producing a game like this is expensive, is the cost really worth it?”’ 123 cook/bildhauer b. auf on [diese this frage]top question würde would wohl probably mancher many nicht-betriebswirt non-economists mit with „typisch typically bwler” economist reagieren. react ‘many non-economists would probably react to this question by saying “that’s typical of economists”.’ example (9) illustrates a class of cases where annotators agreed that there is no aboutness topic. (9b) is the first sentence of a newspaper article, with no prior context related to it except for the heading, given in (9a). however, cases similar to this one also gave rise to non-matching annotations in our study, as example (14) in the next section shows. (9) a. gegen leitschiene ‘against the guardrail’ b. mit with ihrem her pkw car geriet got auf on der the a a 14 14 in in höhe height ortsgebiet municipal.area koblach koblach eine a frau woman (18) (18) aus from mellau mellau ins into.the schleudern. skid ‘a woman (18 yrs.) from mellau got into a skid on the a 14 near the municipal area of koblach.’ 3.5.2 disagreement the examples presented in this section are representative of the numerous cases that caused difficulties. example (10b) is representative of a large number of cases that involve two expressions, each of which could justifiably be analyzed as the at of the sentence. (10) a. dazugelernt habe ich besonders im bereich der öffentlichkeitsarbeit. ich merkte, welche handlung welche reaktion auslöst und wie man gewisse ereignisse richtig kommuniziert. ‘i learned more in the area of public relations work in particular. i noticed what sort of reaction was caused by which actions and how to communicate certain events correctly.’ b. von from [dieser this erfahrung]top? experience kann can [ich]top? i am at.the neuen new ort place selbstverständlich evidently profitieren profit ‘i will clearly be able to profit from this experience at the new place.’ in many cases, one of these candidates is a non-subject that is realized in initial position. however, as mentioned in section 2 above, we do not adopt vallduví’s (1992) approach of identifying the aboutness topic positionally, as it is well known that the aboutness topic can occupy positions other than the initial position in german. the difficulty lies in weighing up whether the prominent position of the pp should have priority over the fact that (i) the subject is commonly considered the default topic of a sentence and (ii) the topic of preceding sentences (in (10a), arguably the subject) is likely to be the topic of the current sentence as well (“topic chain”; see givón, 1983). 124 identifying “aboutness topics” example (11b) is similar to (10b) in that it, too, contains two candidate expressions, but it also differs from (10b) because one of these expressions (namely the subject np) is arguably focussed. (11) a. diese busspur ermöglicht die neue buslinie, die ab 1. juni eingerichtet wird: (. . . ) damit erhalten zum beispiel die bretzenheimer einen flotteren anschluß nach hechtsheim (. . . ) auch in die altstadt geht’s schneller. ‘this bus lane made possible the new bus route, which will operate as of june 1st: (. . . ) the residents of bretzenheimer will thus have a better connection to hechtsheim (. . . ) it will be even quicker to get into the old town-centre too.’ b. außerdem profitiert [der orn-bus aus nieder-olm]top? von [der spur]top? furthermore profits the orn-bus from nieder-olm from the lane ‘the orn-bus from nieder-olm will also benefit from the lane.’ note that the two possible choices of at in (11b) correspond to different discourse strategies: analyzing spur as the at yields topic continuity (cf. givón, 1983) as spur is arguably the topic of (many of) the preceding sentences. on the other hand, choosing the subject-np as the at entails a topic switch.5 turning now to examples (13b) and (14b), the annotators disagreed here on whether they were dealing with a topic-comment structure or rather with a topicless/thetic sentence. at the heart of the disagreement about these examples lies the question of precisely how the two orthogonal ispartitions assumed here (topic/comment vs. focus/background) interact with one another, and in particular, how topic and focus may overlap. various authors (e. g. krifka, 1992; steedman, 2000) suggest that both the topic (theme) and the comment (rheme) section of an utterance each have their own focus-background structure. to our knowledge, the only cases discussed in which topic and focus overlap are cases of so-called contrastive topic; that is, they involve a semantic focus (marked by a rise) within the initial phrase that induces alternatives in addition to a focus later in the clause which also induces alternatives. the overall function is to indicate a discourse strategy whereby only a question that is subordinate to the (possibly implicit) question under discussion is answered. independent of such discourse configurations, the question of the possible overlap of topic and focus has been less explicitly spelt out. for one annotator, there is no intrinsic problem with a complete overlap of (new-information) focus and at as sketched in (12b) in which kim bears both topic and focus but the other annotator rules this out. this issue is in fact far from resolved in the literature on information structure. an anonymous reviewer of this paper was surprised that we should even consider that this complete overlap of topic and focus should be ruled out but, on the other hand, one finds occasional references in the linguistics literature which do indeed seem to rule out this possibility (see e. g. féry, 2007, 168, whose definition of aboutness topic includes the restriction that it be “crucially followed by a focus constituent”). (12) a. who ate the apple? 5. in the terminology of the the prague school (daneš, 1974), these strategies correspond to a ‘thematic progression’ with a continuous theme and a thematic progression with derived themes, respectively. the latter describes a configuration where there is one ‘hypertheme’ (i. e., a discourse topic; the bus lane, in our example), on which individual sentences elaborate. each one of these sentence presents a theme of its own that is ‘derived’ in some way from the hypertheme. 125 cook/bildhauer b. kim [ ]foc [ ]top ate the apple. [ ]background [ ]comment if one disallows a total overlap of topic and focus, i.e. one explicitly disallows a partitioning such as that in (12b), then the question is, of course, how the utterance should instead be analyzed. one possibility which we will discuss now is that in such cases we actually have a topicless sentence. examples such as (13b) and probably also (14b) could thus perhaps be topicless. this, however, raises questions about the possible complexity of topicless/thetic utterances. (13) a. in wil wird das seit anfang oktober gültige rauchverbot nicht überall umgesetzt, und in gewissen lokalen wird noch immer geraucht. häufig wird der gast darauf aufmerksam gemacht, dass es in seiner verantwortung liegt, zu rauchen. ‘the smoking ban that has been in place since the beginning of october is not put into practice everywhere in wil and people still smoke in certain bars. frequently the customer is told that they’re smoking at their own risk.’ b. eine a andere different stimmung atmosphere herrscht reigns im in.the [kirchberger kirchbergian restaurant restaurant eintracht]top? , eintracht wo where das the rauchverbot smoking.ban strikt strictly eingehalten kept wird. is ‘it’s a different situation at kirchberg’s eintracht restaurant, where the smoking ban is strictly adhered to.’ (14) a. bedauern über becks rücktritt ‘deep regret over beck’s resignation’ b. mit with großem big bedauern regret und and totaler total überraschung surprise reagierte reacted gestern yesterday [die the ludwigshafener ludwigshafen spd-prominenz]top? spd-dignitaries auf on den the rücktritt resignation des of.the bundesvorsitzenden federal party leader kurt kurt beck. beck ‘the spd-dignitaries in ludwigshafen reacted with deep regret and utter shock at the resignation of the party leader kurt beck.’ both annotators agree that the example in (13b) can be analyzed as introducing a new referent in the main clause, about which the relative clause makes a further predication. the actual information structure within the main clause itself is, however, not so evident. one annotator selected kirchberger restaurant eintracht as the at of the main clause irrespective of the fact that the same phrase appears to coincide with the final focus of the main clause. the other annotator elected that there was no at in the main clause (i. e. the introduced referent does not function as aboutness topic until later, in the relative clause). note that the only other potential topic candidate, the subject np, as a non-specific indefinite cannot normally function as an aboutness topic.6 under the latter view, the 6. it is worth noticing here that (13b) might be a contrastive topic: both eine andere stimmung and kirchberger restaurant eintracht are contrasted against elements that have been previously mentioned or can be inferred from the 126 identifying “aboutness topics” main clause does not constitute a topic-comment utterance at all. lacking a topic/comment partition is one of the defining features of thetic utterances (e. g. lambrecht, 1994; krifka, 2007), but classifying (13b) as thetic is not without difficulties either. it is customarily said of thetic utterances that the focus spreads across the whole utterance (e. g. lambrecht, 1994; rosengren, 1997), and that thetic sentences in german bear a single accent on the subject (thus, the subject phonologically integrates with the predicate) (e. g. krifka, 1984; sasse, 1987). however, our intuition is that (the main clause of) (13b) requires two prosodic peaks. furthermore, a description of this utterance as event-reporting, a further characteristic of thetics, (cf. götze et al., 2007, 163) does not seem quite correct either since, as mentioned above, the function of sentences like (13b) is to introduce or present a new entity (not an event) which may then later function as aboutness topic in the next discourse chunk. concerning example (14b) now, a similar situation holds. one annotator chose the subject np as topic and the other elected that the sentence had no at. however, this example differs from (13b) in that there is no contrastive element in initial position (thus we are not dealing with a contrastive topic strategy). further, while it was clear in (13b) that the main accent would fall on the pp, here it could be either on the subject np or on the final pp. if one assumes it to fall on the subject np, and if one assumes this to be the at (as one annotator did), then a similar configuration to that in (13b) holds, i.e. we have an overlap of focus and topic (which, recall, some researchers find problematic while others do not). for the other annotator, who opted for a topicless analysis, the fact that the subject-np follows the adverb gestern guided the decision that it is not an at as (14b) does not seem to be a felicitous answer to a question like “what about the ludwigshafen spd-dignitaries?”. the sentence is discourse-initial, preceded only by a headline, and unless the sentence-final np is to be analyzed as at (an option neither annotator took), the only remaining possibility is to classify it as lacking an at if a complete overlap of topic and focus is not an option. however, as was the case with example (13b), in its natural context sentence (14b) requires more than a single prosodic peak and thus does not conform to the description usually given of thetic sentences. thus, analyzing examples like (13b) and (14b) as thetic gives rise to difficulties unless one is willing to adopt a definition of theticity which allows for a type of thetic utterance that introduces or presents an entity (rather than a situation or event). such a definition fits in with the approach to theticity found in lambrecht (1994, 2001) who terms this type of thetic ‘presentational’ (vs. ‘event-reporting’) as well as sasse (1987) where this type is referred to as ‘entity-central’ (vs. ‘event central’). given this bifurcation of the notion of theticity, and bearing in mind the two orthogonal dimensions of is along which sentences are analyzed in the model we are assuming, one may classify thetics in general as “all-comment” but not necessarily as “all-focus”. the difference between entity-central thetics and event-central thetics can then be captured by recourse to their differing focus structures. only event-central thetics involve focus spreading across the whole utterance whilst with entity-central thetics it is merely the phrase that denotes the introduced referent that is focused (illustrated in figure 1). a distinction between event-reporting and entity-presenting thetics was not part of the annotation guidelines at the time of the first annotation experiment. it has been included in the revised version of the guidelines, which will serve as the basis for the second experiment (see section 4). preceding text. however, identifying (13b) as a contrastive topic does not help in deciding whether or not the sentence has an aboutness topic, for it is has been shown that contrastive topics behave differently and crucially do not necessarily involve “aboutness”. see in particular jacobs (1997, 2001); büring (2003); krifka (2007). 127 cook/bildhauer thetic entity-central event-central comment [ ] [ ] focus [ ] [ ] figure 1: analysis of different types of thetics in terms of focus and comment summing up, then, the two different annotation options for example (14b) can be sketched thus (assuming that the main stress is on the subject-np): (15) a. topic-comment-structure: [mit großem bedauern und totaler überraschung reagierte gestern]comment [[die ludwigshafener spd-prominenz]foc]top [auf den rücktritt des bundesvorsitzenden kurt beck]comment b. entity-central thetic: [mit großem bedauern und totaler überraschung reagierte gestern [die ludwigshafener spd-prominenz]foc auf den rücktritt des bundesvorsitzenden kurt beck]comment data of the type exemplified in (13b) and (14b) came up frequently and the problem is thus clearly one that should be clarified in other such annotation tasks in the future. moreover, these data show that it is necessary for annotators to state (in rough terms) the accent pattern they assumed when annotating a sentence token, as different accentuations are sometimes possible and may be indicative of different information structural partitionings. in sum, then, the annotation guideline that were used in this experiment constitute a starting point in bringing terminological clarity to a domain of study (information structure) which is notorious for involving many conflicting definitions on the one hand but also uses of the same terminology in different senses on the other (see, e. g., kruijff-korbayová and steedman, 2003). nevertheless, once the domain of study shifts to naturally-occurring data, the concept of “aboutness topic” presents various difficulties, as thematised here, that have to be addressed both theoretically and in terms of application (i. e., via modifications to the annotation guidelines). the next section presents some such changes to the guidelines and reports on a second annotation experiment in which these modifications were taken into account. 4. annotation experiment #2 4.1 revised annotation guidelines in order to evaluate the revised version of the annotation guidelines that had been used in experiment #1, a second annotation experiment was conducted. the relevant modifications of the guidelines are as follows: • from among the tests for topic-hood listed in (6) above, (6a) was dropped (because it was unclear how deaccentuation of the given material should be handled): (6a) s would be a natural continuation to the announcement let me tell you something about x 128 identifying “aboutness topics” • the distinction between entity-central and event-central thetics that we presented above was incorporated into the guidelines. the revised version gives an example of each of these types, along with an explanation: topicless sentences come in two variants: the main function of an “entity-central” topicless sentence (as in (16a)) is to introduce a new entity (which, in many cases, serves as the topic in the subsequent sentence). in an “event-central” topicless sentence (as in (17)), the emphasis is not so much on introducing an entity, but rather on presenting a situation as a whole. (16) a. (first sentence of a newspaper article) [detectives investigating phone hacking at news international have arrested [a 31-year-old woman]foc]thetic . b. (continuation) [the woman]top is believed to be. . . (17) [[it is raining]foc ]thetic . 4.2 method in the second annotation experiment, annotators were presented with several short texts and had to decide for every np or pp whether or not it is an aboutness topic (or, alternatively, whether they were dealing with a thetic sentence as defined above). this differs from the approach taken in the first experiment, where annotators chose the aboutness topic from a small, pre-selected set of options (like subject-np, object-np etc.), and it reflects more closely the typical context of use of the annotation guidelines. the annotation task was done electronically, using the open-source tool mmax2.7 prior to the actual annotation task, all annotators participated in a training phase consisting of (a) answering clarification questions about the annotation guidelines, (b) performing a test annotation jointly with the experimenters, (c) annotators performing a test annotation on their own and (c) discussing the results of the individual test annotations. after the initial training phase, the annotators carried out the actual annotation task within three weeks, at home. 4.3 materials seven short texts (5 – 11 sentences each; 8 on average) from the german newspaper die tageszeitung (extracted from the tüba-d/z corpus) were used in the experiment, thus matching the text type used in experiment #1. the markables (np and pp constituents) were indicated, thus the annotators did not have to identify them on their own. 4.4 subjects a total of four (2 m, 2 f; age 21–42 years, 29.5 on average) students with a background in linguistics (0.5–4.5 years of study) participated in the experiment. all of them are native speakers of german and were largely unfamiliar with notions of information structure prior to the experiment. 7. http://mmax2.net 129 cook/bildhauer 4.5 results the distinction between thetic and topic-comment sentences yields 56 data points per annotator (one for each sentence), with 39.3% raw agreement between the annotators. in 21 cases, all annotators agree that a sentence has a topic-comment structure, while there is only a single case in which all four agree that they are dealing with a thetic sentence (though they do not agree on the subtype of thetic as defined in section 4.1). with respect to the identification of aboutness topics, we examined a total of 516 data points for each annotator. in 17 cases, all annotators agree in their choice of the aboutness topic (and in 392 cases they agree that a constituent is not an aboutness topic), thus raw agreement is relatively high (79.3%). in order to factor out chance agreement, fleiss’s kappa for m = 4 raters was calculated, both for the annotation of aboutness-topics and for the more coarse-grained distinction between thetic and topic-comment sentences, as shown in table 3. although the results reach statistical significance, the amount of agreement beyond chance is only “moderate” for the choice of the aboutness topic, while there is merely “fair” agreement on the thetic/topic-comment distinction.8 n items % raw agreement κ(fleiss) p aboutness topic 516 79.3 .447 < .001 thetic/topic-comment 56 39.3 .225 < .001 table 3: overall inter-rater agreement in experiment #2 4.6 discussion it might look odd at first glance that there is considerably less agreement on the question of whether or not raters are dealing with a non-thetic (i. e., topic-comment) sentence at all, than on the actual choice of the aboutness topic. this is probably an artifact of the experiment design: a decision (topic or not?) had to be made for a very large number of markables, including sub-parts of nps to which raters are not likely to attribute information structural properties in a principled way, at least not at the level of training of our participants. a large amount of “not a topic” annotations thus makes up for the major part of the observed agreement. this is illustrated by the fact that, if only those items are taken into account which at least one of the raters annotated as aboutness topic (n = 143), there is only slight agreement (κ = .067) and statistical significance is marginal (p = .051). there are 5 cases, though, in which different annotators did class different sub-parts of an np as the topic. since such cases are not explicitly discussed in the annotation guidelines, this kind of embedded markable might have caused confusion that ultimately lowered overall inter-rater agreement. we “repaired” these annotations by attributing topic status to the embedding np. even with this correction applied, inter rater agreement does not improve substantially (κ = .476). inter-rater agreement on the choice of a specific subtype of “thetic” cannot be evaluated in the same way since there is only a single case in which all annotators agree that a sentence is thetic in the first place. another interesting question to ask is whether or not the annotators agree on a specific topic once they all have decided that they are dealing with a topic-comment sentence (and not with a thetic one). we recalculated inter-rater agreement on the topic constituent for the subset of sentences that 8. as before, according to landis and koch’s (1977) interpretation of κ-values. 130 identifying “aboutness topics” figure 2: similarity in distinction “thetic/topic-comment” (euclidean distance, clustering based on average distance): annotators fall into two groups. all annotators had rated as a topic-comment-structure. the resulting, higher coefficient (κ = .563, p < 0, n = 21) indicates that in those cases which are clear instances of topic-comment sentences, annotators tend to agree more on the aboutness topic, though agreement is far from being perfect in these cases, too. example (18b) illustrates one such sentence for which all four annotators agreed on the choice of the aboutness topic. (18) a. friedensbewegung, christliche und linke gruppen sowie gewerkschafter aus der ganzen republik haben dagegen zu einer demonstration gegen den krieg in jugoslawien aufgerufen. ‘the peace movement, christian and leftist groups as well as trade unionists from all over germany have announced a demonstration against the war in yugoslavia.’ b. [die the protestveranstaltung] protest.event steht stands unter under dem the motto motto „stoppt stop den the krieg! war helfen help statt instead.of bomben!”. bombs ‘the protests will take place under the motto “stop the war! help them, don’t bomb them!” in the discussion of experiment #1, we briefly pointed to the possibility that the annotation guidelines leave too much room for interpretation and thus can be understood in different ways. in order to investigate this question further, we computed a distance matrix for the annotations of our four participants that would allow us to group them according to the similarity of their annotations. with respect to the distinction between “thetic” and “topic-comment”, there is in fact evidence that the guidelines are ambiguous, as the cluster tree in figure 2 illustrates: the annotators fall into two groups that differ in their rating behaviour. the picture is not so clear, though, for the choice of the aboutness topic. here, two of our annotators pattern closely together, while the third and the fourth annotator differ both from each other and from the first two annotators (see figure 3). 131 cook/bildhauer figure 3: similarity in choice of aboutness topic (euclidean distance, clustering based on average distance): no clear pattern. 5. general discussion topic: selecting among various candidates a high degree of agreement on topic choice was often observed when the phrase in question refers to an entity that can be analyzed as the topic of the preceding utterance(s), and is realized in a syntactically prominent position (e. g. as the subject), as is the case in example (18b) given in section 4. on the other hand, we observe that a typical configuration where annotators disagree involves cases in which the topic of the preceding utterance(s) is realized in a syntactically or functionally less salient position, while another expression that was not the topic in the preceding utterance(s) occupies a more salient position in the utterance in question. this pattern, illustrated in (19), is reminiscent of the “retain” transition type of centering theory (grosz et al., 1983, 1995), which typically announces a shift of the “backward looking center” (marked as cb in example (19)). in (19b), the (referent of the) np hamburg counts as the backward-looking center because it is coreferential with the most salient referent of the preceding utterance. it is arguably also the most salient referent in (19b). now, in (19c), it is referred to again, thus (the referent of) hamburg is also the backwardlooking center of that utterance. however, it is not realized in the most salient position (the subject position) in (19c); instead the subject position is occupied by the np kanal4, which is probably why some annotators analyzed kanal4 as the aboutness topic. while the concept of “backward looking center” is not fully coextensive with our understanding of “aboutness topic” (but see e. g. beaver, 2004, who does equate these two concepts), the two notions are close enough to make this observation noteworthy, if only because it allows one to describe in precise terms one of the contexts in which confusion is likely to arise: the manifest conflict between structural saliency on the one hand and referential continuity on the other hand is surely an issue that should be addressed in any future version of the annotation guidelines. (19) a. hamburg eine stadt verabschiedet sich vom rechtsstaat, rtlplus , so. , 0.40 uhr ‘hamburg a city bids farewell to constitutional law, rtlplus, sunday, 12:40 a. m.’ b. ist is von from [hamburg]cb hamburg die the rede, speech, denken think viele many an of eine a weltoffene world.open und and liberale liberal stadt. city ‘if they hear „hamburg”, many people imagine an open-minded and liberal city.’ 132 identifying “aboutness topics” c. [kanal4] kanal4 zeigt shows am on sonntag sunday zu at später late stunde hour bundesweit nationwide auf on rtlplus rtlplus eine a andere other seite side [der of.the hansestadt]cb . hanseatic.city ‘kanal 4 will be showing a different side of the hanseatic city late on sunday in a national broadcast on rtlplus.’ topic-comment vs. thetic one of the main current weaknesses leading to poor inter-rater agreement, we believe, lies in the fact that in both annotation tasks the premise was adopted that every sentence either has an aboutness topic or else it is necessarily a thetic (i. e. topicless) utterance. the annotation guidelines (both the original and the revised version) simply do not allow for any sentence type that does not fit into either of these categories. this is perhaps an artifact of current research on information structure in theoretical linguistics. typically, studies on information structure in that particular research tradition have not thus far focused on naturally-occurring data. when one examines such data, however, as was the case in the annotation experiments, it can be observed that utterances which do not straightforwardly fit into the postulated available utterance types (topic-comment vs. thetic) actually seem quite frequent. from the results reported in the preceding section, it is obvious that the modification of the annotation guidelines, i. e. slightly greater explicitness in characterizing what should count as a thetic sentence, along with the distinction between two types of thetics and examples for each one of these, is not helpful when it comes to deciding whether or not a given sentence is to be classed as thetic. for example, while in the second experiment, there is a number of sentences (37.5%) that seem to qualify as clear instances of a “topic-comment” structure (those which the annotators unanimously classed as such), the same does not hold for the category “thetic”, for which there is only one case where all annotators agree. given the fact that not even the revised definition of “thetic”, as adopted in the revised guidelines, seems to accommodate naturally-occurring data, although it already covers cases that are structurally much more complex than the kind of sentences traditionally called “thetic”, there seems to be an urgent need to clarify the notion of “thetic” from a theoretical perspective. in particular, it is quite unclear how complex a thetic sentence can be and how one can even tell with any certainty whether a sentence is thetic or not, as there are no clear criteria offered in the literature that we are aware of. for instance, some authors restrict theticity to intransitive verbs while others clearly do not, yet this contrast does not appear to be openly addressed or commented on in the theoretical literature. in this connection, the question arises whether the clear-cut distinction between thetic and non-thetic is in fact an idealization that is not necessarily reflected in naturally-occurring data: alternatively, it is conceivable that there is a gradual transition from sentences with a clear topic-comment structure via sentences with a less salient topic-comment structure to sentences that do not display a topic-comment structure at all (thetics). appropriateness of thetic vs. topic-comment there is also the possibility that the distinction between “thetic” vs. “topic-comment” is not equally pertinent for different kinds of utterances. it is plausible, and indeed sometimes necessary, to assume the existence of other utterance types which do not serve to present an entity nor an event (i. e. are not thetic) but which also do not have the function of a topic-comment utterance, namely advancing 133 cook/bildhauer the common ground by predicating the information expressed in the “comment” of some entity under whose file card this information is to be stored. while the exact characterization of such utterances and a detailed description of constraints on their occurrence is currently not available, we will briefly discuss a few examples which deserve further study in this respect, viz. sentences that do not express an assertion, and subordinate clauses. in the case of sentences that do not serve to make an assertion, such as interrogatives, imperatives and exclamatives, the simple dichotomy of topic-comment vs. thetic does not seem sufficient. our intuition is that since these encode a different type of speech act, it seems very likely that they have a different function in terms of discourse structuring to that normally attributed to topiccomment configurations. this intuition is of course in need of closer scrutiny and empirical testing. currently, the annotation guidelines do not state whether or not such sentences should be considered to host an aboutness topic and so, in the absence of any specific instruction to the contrary, the annotators sought to annotate these utterance types as either containing an aboutness topic or as instantiating a thetic utterance. consider, for instance, the following example which is a verb-first conditional: (20) ist is von of hamburg hamburg die the rede, speech denken think viele many an of eine a weltoffene world.open und and liberale liberal stadt city ‘if they hear “hamburg”, many people imagine an open-minded and liberal city.’ two annotators in experiment #2 selected the np hamburg as the aboutness topic. one of these annotators also selected a further aboutness topic in the consequent clause of the conditional (namely the np eine weltoffene und liberale stadt). the other two annotators selected no aboutness topic at all in the whole utterance (neither in the antecedent nor the consequent clause). thus, with the current dichotomy offered in the annotation guidelines (if no topic, then thetic), two annotators classed (or were indirectly forced to class) this utterance as thetic. as mentioned above, it is not implausible to assume that the antecedent of a conditional necessarily does not involve a topic-comment partition at all. this, of course, should not entail that the sentence must therefore automatically be considered thetic either. rather, the conditional clause and the consequent clause taken together appear to provide the correct domain for attributing discourse functions such as topic and comment, with the conditional clause establishing the topic, and the consequent clause predicating of it. however, the antecedent clause, considered in isolation, seems to simply have a very different discourse structuring function and as such the question of whether it is thetic or not or of which phrase instantiates the aboutness topic is just not a relevant question. a similar situation holds for exclamative utterances such as (21) which probably have a different function in discourse than either that of predication (about a topic) or theticity. one annotator in experiment #2 selected the expletive object es (which is part of a fixed idiomatic expression) as aboutness topic. the others didn’t mark any element as aboutness topic, which again illustrates the difficulties that annotators are facing when confronted with this type of sentence. (21) besser better kann can man one es it gar not nicht make machen! ‘there is no better way to do it.’ for future studies, it would probably be advisable to give detailed instructions as to how these sentence types should be handled, since in the absence of explicit instructions, annotators are bound to become confused. 134 identifying “aboutness topics” turning now to subordinate clauses, we observe that the task of identifying their information structural partitioning poses a major problem for the annotators of experiment #2, giving rise to numerous conflicting judgements. in some cases, the subordinate clause was not independently assessed and in some it was (recall the discussion about annotating the mainand the relative clause structure of (13b) in the first experiment). it is not evident whether all subordinate clauses should be considered to constitute an independent is domain of their own, independent of that of their matrix clause, or whether matrix and subordinate clause together should be considered to constitute one broad is domain. for the case in point, the question boils down to whether or not a subordinate clause should be considered to potentially embed a topic-comment partition separate from the topiccomment partition of the matrix. in principle, there seems to be no ban on the embedding of thetic utterances per se. at present, there is no detailed discussion of the treatment of subordinate clauses in the annotation guidelines and annotators were simply instructed to annotate all finite subordinate clauses except restrictive relative clauses. again, we believe there is a need for careful empirical testing of whether or not this is the correct procedure. it seems intuitively to make sense to distinguish finite and non-finite embedded clauses since the latter typically do not overtly express the subject of the embedded predicate and this therefore is simply not available as a contender for topic status. often this leaves no argument other than the object role available as a topic. annotators were thus instructed in the guidelines not to consider non-finite clauses as separate is domains. further it also seems likely that the choice of the subordinating conjunction plays a role in the answer as to whether or not the embedded clause involves an independent is domain. it has been observed that different (classes of) conjunctions instantiate different discourse relations and impose different constraints on the “common ground”. it is thus not implausible that the embedded propositions these classes introduce differ with respect to their internal is. this point perhaps also applies to different coordinating conjunctions (cf. und ‘and’ vs. aber ‘but’). further, the semantics of the embedding predicate may well affect whether or not a subordinated proposition constitutes an independent is domain. it has been suggested, for instance, that presupposed complements (of e. g. factive verbs) are not asserted and may therefore lack a topic-comment articulation altogether (cf. cook and ørsnes, 2010; ebert et al., 2009; kuroda, 2005). in this respect, example (22c) is typical of a number of cases that caused problems. the matrix verb aufdecken ‘reveal’ is factive and is followed by a subordinate clause in the function of direct object which is introduced by the subordinating conjunction wie ‘how’. the annotators did not agree here on whether or not hamburg (or der stadtstaat hamburg) should count as an aboutness topic. (22) a. ist von hamburg die rede, denken viele an eine weltoffene und liberale stadt. kanal4 zeigt am sonntag zu später stunde bundesweit auf rtlplus eine andere seite der hansestadt. ‘if they hear “hamburg”, many people imagine an open-minded and liberal city.’ b. kanal4 zeigt am sonntag zu später stunde bundesweit auf rtlplus eine andere seite der hansestadt. ‘kanal 4 will be showing a different side of the hanseatic city late on sunday in a national broadcast on rtlplus.’ 135 cook/bildhauer c. die the beiden two fernsehjournalisten tv.journalists ernst e. matthiesen m. und and oliver o. neß n. decken reveal auf, up wie how sich itself der the stadtstaat hamburg sovereign city of hamburg immer ever mehr more in in richtung direction eines of.a „polizeistaats” police.state bewegt. moves ‘the two tv journalists e. m. and o. n. reveal how the sovereign city of hamburg begins to turn into a police state.’ similarly, the different possible syntactic realizations of embedded clauses could be of significance. german permits embedded verb-last and verb-second clauses (both with and without subordinating conjunction) and it is plausible that this too correlates with different is possibilities. moreover, it is not clear whether the distinction between argument clauses and adverbial clauses also relates to is differences; and indeed whether different sub-types within both argument and non-argument clauses need to be distinguished. finally, the treatment of complements of non-verbal categories such as nouns and adjectives must also be considered. discourse relations finally, another issue which seems worthy of consideration in this connection concerns the discourse relations between utterances (in the sense of e. g. asher and lascarides, 2003; mann and thompson, 1988) and the role of the notion of aboutness topic in such discourse relations. while certain discourse relations do appear to serve to advance the “common ground” in the way described by krifka (2007) such as, for instance, “restatement”, “elaboration” and “narration”, it is perhaps worth considering whether a particular utterance may or may not instantiate a topic-comment partition as a function of the discourse relation that connects this utterance to the remaining discourse. as mentioned by krifka (2007) at the end of his is overview, this interaction between discourse relations and is has not been extensively examined and there is a need for further research here. in this vein, an anonymous reviewer also points out to us that frameworks such as e. g. asher and lascarides (2003) and related frameworks provide a tool for identifying the arguments of certain discourse relations as contributing static (rather than eventive) information. this raises the (attractive) possibility that a corpus annotated with this level of description (which is independently useful anyway) could turn out to be helpful in clarifying (or even perhaps obviating the need for) the distinction between thetic and topic-comment utterances which were, recall, a particular source of problems in both of the annotation exercises. in general, these comments are suggestive to us that an optimally annotated corpus should include annotations of numerous different levels (focusbackground, topic-comment, given-new, discourse relations etc.) in combination with one another so that cross-cutting information can be drawn on in order to enhance the precision of annotation. this is an important and promising direction for future research, we believe. explicitness of the annotation guidelines as a consequence of what has been discussed so far, it is obvious that the annotation guidelines must be much more specific as to how particular constructions are to be analyzed. the guidelines should, for instance, provide concrete instructions as to how different types of subordinated and non-assertive utterances should be treated and how coordinations are to be handled. while it is 136 identifying “aboutness topics” desirable to use the same basic notions of information structure as a fundament for annotating texts in a variety of different languages, we doubt that the necessary degree of explicitness in the guidelines can be achieved without language-specific instructions on how particular phenomena are to be treated. taking the results reported in this paper as a baseline, we believe that introducing a number of language-specific rough-and-ready distinctions would help raise the overall figure of inter-rater agreement, even though it might not do justice to a putatively small number of cases that feature a somehow “uncanonical” link between syntactic form and information structure. broadening the concept of topicality as we have mentioned at various points above, there are many different properties that are associated with topicality. the notion of aboutness topic we selected to work with and which was outlined in section 1 explicitly allows topics to be new in the discourse, although it has long been acknowledged (and we would also agree) that notions such as familiarity and givenness prototypically align with topicality (cf. chafe, 1976 but also reinhart, 1981). the definition of aboutness topic is also not tied to any syntactic position but the idea that topics tend to be realized in syntactically prominent positions or undergo a kind of separation (cf. rizzi, 1997 but also e. g. jacobs, 2001) is also well-known. the concept of invoking lists of referents or of partiality or of addressing sub-questions under discussion also appears to tendentially align with topical status (cf. roberts, 1996; büring, 2003 and also the notion of a partially-ordered set relation (poset) of birner and ward, 1998). again, the definition of aboutness topic used here is formulated independently of such effects. the question arises (given the clear difficulty annotators had with the annotation of aboutness topic as defined here in naturally-occurring data) whether an optimal annotation of topic can better succeed if a “bundle” approach is adopted as advocated by e. g. smith (2003, 198–199) who lists a set of “topic cues”. annotators could plausibly be instructed to differentiate between canonical or highly prototypical topics and less prototypical topics. examples of the former type, i. e. in which a bundle of prototypical topic properties coincide on one candidate expression, were those which achieved a higher degree of inter-rater agreement. on a slightly different note, it is plausible that there is some cross-linguistic variation concerning which properties play a greater role in defining topicality (see mcnally, 1998, for a similar suggestion) and that language-specific “bundles” of topic features should be identified and spelt out in annotation guidelines. 6. summary in the present contribution, we reported on the results of two annotation experiments involving naturally occurring data, in the course of which raters categorized sentences as either “thetic” or “topic-comment” and, in the latter case, also identified the aboutness topic of the sentence. in experiment #1, two expert raters sought to follow the annotation guidelines of götze et al. (2007). this resulted in a relatively low degree of inter-rater agreement as a consequence of which we identified a number of typical cases that gave rise to difficulties. as the distinction between “thetic” and “topiccomment” turned out to be particularly difficult to assess, a modification to the original guidelines was proposed that elaborated on the notion of “thetic” and broadened the range of phenomena that it can be applied to. the revised version of the annotation guidelines served as the basis of experiment #2, in which four non-expert raters (i. e. students with a background in linguistics but no expertise in information structure) annotated several short newspaper texts. the modifications to the the original annotation 137 cook/bildhauer guidelines did not improve inter-rater agreement (fleiss κ does not exceed .23 for thetic/topiccomment distinction and .48 for the choice of the aboutness topic). we conclude that neither the original nor the revised version of the guidelines provide a particularly reliable set of rules when it comes to annotation of aboutness topics. to help improve this situation, we highlighted a number of issues that should be taken into account in any future revision of the guidelines (or in any other set of annotation guidelines for information structure). in particular, we suggest that the notion of “thetic” (and its relation to “topic-comment”) should be subjected to further scrutiny because its applicability is less than clear once the domain of study shifts from idealized, constructed examples to naturally-occurring data. in this connection, we also raised the question of whether or not the thetic/topic-comment distinction is a pertinent one for all sentences types alike. among other things, conditionals and various other kinds of non-assertive utterances do not necessarily lend themselves to an analysis in terms of “thetic” vs. “topic-comment”. moreover, the discussion here makes clear that the treatment of subordination is complex and probably requires consideration of numerous different parameters. finally, we pointed to the fact that the concept of “aboutness topic” is still lacking a level of operationalization that would allow annotators to reliably identify this information structural construct when scanning naturally-occurring data. borrowing some of the descriptive machinery of centering theory, we described one of the typical configurations in which annotators disagree on their choice of a sentence’s aboutness topic as a conflict between referential continuity and structural salience. in addition to these points, we also expressed serious doubts that any annotation guidelines for information structure can be stated in a way that is neutral with respect to the target language. while the guidelines used in the two experiments no doubt constitute a valuable resource and starting point, we suggest supplementing them with instructions specific to particular languages, addressing a concrete set of phenomena and explaining how these should be handled. this probably also implies introducing a number of possibly controversial decisions into the guidelines, that is, the guidelines would not necessarily reflect a relatively consensual stance on information structure anymore. however, this would seem to be preferable to a set of theoretically consensual, but vastly underspecified, annotation rules. summing up, we hope to have alerted other researchers planning a similar enterprise to some pitfalls they may encounter and hope we can contribute to the discussion concerning issues which also have a resonance for theoretical linguistics. at the moment, we must concede that a careful and thorough discussion of the points raised above is required before the notion of aboutness topic, as defined here, can be employed with reasonable confidence in annotation tasks. references ben ambridge and adele e. goldberg. the island status of clausal complements: evidence in favor of an information structure explanation. cognitive linguistics, 19(3):357–389, 2008. nicholas asher and alex lascarides. logics of conversation. cambridge university press, cambridge, 2003. david i. beaver. the optimization of discourse anaphora. linguistics and philosophy, 27(1):3–56, 2004. betty j. birner and gregory ward. information status and noncanonical word order in english. benjamins, amsterdam, 1998. 138 identifying “aboutness topics” daniel büring. on d-trees, beans, and b-accents. linguistics & philosophy, 26(5):511–545, 2003. wallace l. chafe. givenness, contrastiveness, definiteness, subjects, topics and point of view. in charles n. li, editor, subject and topic, pages 25–55. academic press, new york, 1976. jacob cohen. a coefficient of agreement for nominal scales. educational and psychological measurement, 20(1):37–46, 1960. philippa cook and bjarne ørsnes. coherence with adjectives in german. in stefan müller, editor, the proceedings of the 17th international conference on head-driven phrase structure grammar, pages 122–142, stanford, 2010. csli publications. philippa cook and john payne. information structure and scope in german. in miriam butt and tracy holloway king, editors, proceedings of lfg06, university of constance, pages 124–144, stanford, 2006. csli publications. františek daneš. functional sentence perspective and the organization of the text. in papers on functional sentence perspective, pages 106–208. mouton, the hague and paris, 1974. kordula de kuthy. discontinuous nps in german. csli publications, stanford, 2002. christian ebert, cornelia endriss, and stefan hinterwimmer. embedding topic-comment structures results in intermediate scope readings. in muhammad abdurrahman, anisa schardl, and martin walkow, editors, proceedings of nels 38, amherst, 2009. glsa. cornelia endriss. quantificational topics. a scopal treatment of exceptional wide scope phenomena. springer, dordrecht, 2009. caroline féry. information structural notions and the fallacy of invariant correlates. in caroline féry, gisbert fanselow, and manfred krifka, editors, interdisciplinary studies on information structure, volume 6 of working papers of the sfb 632, pages 161–184. universitätsverlag, potsdam, 2007. joseph l. fleiss. measuring nominal scale agreement among many raters. psychological bulletin, 76(5):378—-382, 1971. georg von der gabelentz. ideen zu einer vergleichenden syntax. zeitschrift für völkerpsychologie und sprachwissenschaft, 6:376–384, 1869. talmy givón. introduction. in talmy givón, editor, topic continuity in discourse. a quantitative cross-language study, pages 1–41. benjamins, amsterdam, 1983. michael götze, thomas weskott, cornelia endriss, ines fiedler, stefan hinterwimmer, svetlana petrova, anne schwarz, stavros skopeteas, and ruben stoel. information structure. in stefanie dipper, michael götze, and stavros skopeteas, editors, interdisciplinary studies on information structure, volume 7 of working papers of the sfb 632, pages 147–187. universitätsverlag, potsdam, 2007. 139 cook/bildhauer barbara j. grosz, aravind k. joshi, and scott weinstein. providing a unified account of definite noun phrases in discourse. in proceedings of the 21st annual meeting of the association for computational linguistics, pages 44–50, cambridge, massachusetts, usa, 1983. association for computational linguistics. barbara j. grosz, aravind k. joshi, and scott weinstein. centering: a framework for modeling the local coherence of discourse. computational linguistics, 21(2):203–226, 1995. herman hendriks and paul dekker. links without locations. information packaging and nonmonotone anaphora. in paul dekker and martin stokhof, editors, proceedings of the tenth amsterdam colloquium, pages 339–358, university of amsterdam, 1996. illc-department of philosophy. institut für deutsche sprache. deutsches referenzkorpus. archiv der korpora geschriebener gegenwartssprache 2009-i. (release vom 28. 02 .2009), 2009. http://www. ids-mannheim.de/dereko. joachim jacobs. i-topikalisierung. linguistische berichte, 168:91–134, 1997. joachim jacobs. the dimensions of topic-comment. linguistics, 39(4):641–681, 2001. manfred krifka. fokus, topik, syntaktische struktur und semantische interpretation. unpublished manuscript, 1984. url http://amor.rz.hu-berlin.de/~h2816i3x/ publications/krifka\%201984\%20fokus.pdf. manfred krifka. a compositional semantics for multiple focus constructions. in joachim jacobs, editor, informationsstruktur und grammatik, pages 17–54. westdeutscher verlag, opladen, 1992. manfred krifka. basic notions of information structure. in caroline féry, gisbert fanselow, and manfred krifka, editors, interdisciplinary studies on information structure, volume 6 of working papers of the sfb 632, pages 13–56. universitätsverlag, potsdam, 2007. ivana kruijff-korbayová and mark steedman. discourse on information structure. journal of logic, language and information: special issue on discourse and information structure, 12(3): 249–259, 2003. shige-yuki kuroda. focusing on the matter of topic: a study of wa and ga in japanese. journal of east asian linguistics, 14:1–58, 2005. knud lambrecht. information structure and sentence form. ttopic, focus, and the mental representations of discourse referents. cambridge university press, cambridge, 1994. knud lambrecht. when subjects behave like objects. an analysis of the merging of s and o in sentence-focus constructions across languages. studies in language, 24(3):611–682, 2001. j. richard landis and gary g. koch. the measurement of observer agreement for categorical data. biometrics, 33(1):159–174, 1977. william c. mann and sandra a. thompson. rhetorical structure theory. toward a functional theory of text organization. text, 8(3):243–281, 1988. 140 identifying “aboutness topics” louise mcnally. towards a theory of the linguistic coding of information packaging instructions. in peter culicover and louise mcnally, editors, the limits of syntax, volume 29 of syntax and semantics, pages 161–183. academic press, new york, 1998. barbara h. partee. focus, quantification, and semantics-pragmatics issues. in peter bosch and rob van der sandt, editors, focus. linguistic, cognitive, and computational perspectives, pages 213–231. cambridge university press, cambridge, 1999. tanya reinhart. pragmatics and linguistics. an analysis of aentence topics. philosophica, 27:53–94, 1981. luigi rizzi. the fine structure of the left periphery. in liliane haegeman, editor, elements of grammar, pages 281–337. kluwer, dordrecht, 1997. craige roberts. information structure in discourse. towards an integrated formal theory of pragmatics. in jae hak yoon and andreas kathol, editors, papers in semantics, volume 49 of osu working papers in linguistics, pages 91–136. the ohio state university, ohio, 1996. inger rosengren. the thetic/categorical distinction visited once more. linguistics, 35(3):439–479, 1997. hans-jürgen sasse. the thetic-categorical distinction revisited. linguistics, 25(3):511–580, 1987. carlota s. smith. modes of discourse. the local structure of texts. cambridge university press, cambridge, 2003. robert stalnaker. assertion. in peter cole, editor, pragmatics, volume 9 of syntax and semantics, pages 315–332. academic press, new york, 1978. mark steedman. information structure and the syntax-phonology interface. linguistic inquiry, 31 (4):649 – 689, 2000. enric vallduví. the informational component. garland, new york, 1992. 141 journal of machine learning research-microsoft word template dialogue and discourse 4 (2) (2013) 174-184 doi: 10.5087/dad.2013.208 ©2013 deniz zeyrek et al. submitted 03/12; accepted 12/12; published online 08/13 turkish discourse bank: porting a discourse annotation style to a morphologically rich language deniz zeyrek dezeyrek@metu.edu.tr middle east technical university, informatics institute dumlupinar boulevard, no:1 06800, ankara, turkey işın demirşahin disin@metu.edu.tr middle east technical university, informatics institute ayışığı b. sevdik çallı asevdik@gmail.com middle east technical university, informatics institute ruket çakıcı ruken@ceng.metu.edu.tr middle east technical university, department of computer engineering editors: stefanie dipper, heike zinsmeister, bonnie webber abstract this paper describes the current state of the turkish discourse bank, the first publicly available annotated discourse resource for turkish. it describes the annotation methods and the challenges posed by annotating turkish, a free word order language with rich morphology. it shows the usefulness of the pdtb style annotation but points out the need to expand this annotation style with the needs of the target language. keywords: turkish, discourse, discourse connectives, discourse annotation 1 introduction annotated corpora have come to play an important role both in theoretical linguistics and machine learning applications in natural language processing. there is a pressing need for such resources in turkish, a free-word order, agglutinative language with rich morphology. there are existing syntactically enriched treebanks for turkish (e.g., oflazer et al., 2003) but the field also needs annotated discourse corpora to appeal to the need of researchers who are working with texts in their entirety rather than individual sentences. 1 annotated discourse corpora allow opportunities to understand what kind of relationships hold among lexical, morphological and syntactic levels and the textual level. they provide an empirical ground for investigating a range of discourse issues and can reveal structures in discourse via various language technology applications, e.g., summarization, information extraction, sentiment analysis, essay analysis, etc. (cf. webber et al., 2011). the turkish discourse bank (tdb) is a ~400,000-word resource of modern written turkish with various genres, mainly containing annotations of explicit discourse connectives and the discourse segments they relate. it shares the principles of the penn discourse tree bank (pdtb) 1 the metu-sabancı turkish treebank described in (oflazer, 2003; say et al., 2004) is the widely-known syntactically annotated treebank of turkish. since the annotated sentences of this corpus do not form entire texts, we chose not to create the tdb on it. turkish discourse bank 175 and takes discourse connectives as discourse-level predicates with a binary argument structure (prasad et al., 2007). the connective denotes relations between eventualities, fact-like objects and proposition-like objects (asher, 1993). following the pdtb, we refer to the arguments of a connective as the first argument (arg1) and the second argument (arg2). the second argument is the textual unit which syntactically or morphologically contains the connective, the other argument is conveniently referred to as arg1. arg2 is the “internal” unit, arg1 the “external” unit in stede & heintze’s (2004) terminology. the complete list of the tdb tagset is given in table 1 (zeyrek et al., 2010). the definitions and examples of each category are provided in the rest of the paper while discussing the relevant methodological issues. conn the connective head arg1 first argument of the connective arg2 second argument of the connective supp1 supplement to the first argument supp2 supplement to the second argument shared the subject, object or adverbial phrase shared by a relation shared supp supplement for the shared material mod modifier of the connective or the modifier of the relation table 1. the annotation scheme of the tdb only explicit connectives and their two arguments are annotated in the tdb. we plan to annotate implicit connectives and the sense of connectives at a later stage. we chose to annotate explicit connectives first because the primary aim of the project was to reveal explicit discourse connectives to allow further research in discourse coherence. in creating a bank of discourse based on connectives, we do not claim that discourse connectives are the only means establishing coherence; we merely take them as the basic elements of discourse which make coherence relations salient. the tdb 1.0 has been released in march 2011 along with a browser; it is being freely distributed to researchers upon request (www.medid.ii.metu.edu.tr). the tdb is built on certain principles shared by all annotated corpora (marcus et al., 1994; skut et al., 1997). firstly, it is descriptive. it has a bottom-up approach, aiming to describe the basic characteristics of discourse by annotating discourse relations between segments. examining discourse by breaking it up to its constituents is the core of almost all theoretical work on discourse, e.g. asher & lascarides (2003), grosz & sidner (1986), mann & thompson (1988), moser & moore (1996), polanyi & van den berg (1996). secondly, the tdb is data-driven; i.e. the tagging scheme is meant to allow representations of various discourse phenomena, including, for example, shared and crossing arguments (aktaş et al., 2010). thirdly, it is theory-independent. it is not influenced by a particular discourse theory; i.e., there is not a correct way of annotating discourse relations given a specific theory, the only requirement is that similar structures should be annotated in the same way for consistency. in our earlier work, we discussed the evolving stages of the corpus (e.g., demirşahin, et al. 2012b). in the present paper, we provide a more complete picture of the finalized annotation system and present statistical information about the connectives that proved challenging in the annotation process, namely discourse adverbials, subordinators and their polymorphous occurrences. the rest of the paper proceeds as follows: in section 2, we describe the annotation cycle, methods, and challenges in porting an annotation system to turkish. we focus on the annotation of discourse adverbials and how we annotate the discourse relations expressed by nominalizations. in section 3, we explain the method of annotating phrasal expressions and how we decided to use the shared tag. we also discuss how we adapted the modifier and the supplementary tags of the pdtb. finally, in section 4 we conclude with a summary of the paper. zeyrek, demi̇rşahi̇n, sevdi̇k çalli, and çakici 176 2 porting an annotation system to turkish discourse in this section, we describe the annotation cycle, the annotation procedures and how we evaluate the annotation scheme. 2.1 the annotation cycle the annotation system used in creating the tdb involves the following steps.  a first draft of annotation guidelines is prepared and an initial set of explicit connectives is determined.  three annotators annotate the whole corpus for the given set of connectives by determining their arg1 and arg2 spans, their supplementary materials and modifiers. they go through the whole data, skipping non-discourse usage of the connectives.  the annotated corpus is statistically analyzed.  the disagreements are determined and discussed interactively in agreement meetings, focusing on the disagreed cases. two researchers and all three annotators participate in the meetings. with the researchers’ feedback, disagreements are resolved and an ‘agreed’ version is produced. the annotation guidelines are updated.  in the cases when the agreement meeting results in a modification of the annotation guidelines, all past annotations are reexamined through a process called ‘proof’ to ensure that they are in line with the latest version of the guidelines. the proofed annotations are the final version, the gold standard of the tdb.  the next set of connectives is annotated, and the cycle continues. similar to english and many other languages, discourse connectives in turkish can be identified from three major syntactic classes, namely, coordinating conjunctions (ve ‘and’, ya da ‘or’, ama ‘but’), subordinators (complex subordinators, e.g. için ‘for’, simplex subordinators, i.e. converbs, e.g. -inca ‘when,’ –ken ‘while/now that’), and discourse adverbials (oysa ‘however’, öte yandan ‘on the other hand’, ayrıca ‘in addition/separately’). 2 the initial list of connectives was prepared on the basis of these syntactic classes. we excluded simplex subordinators, which we aim to annotate later. then the researchers and the annotators discussed how one distinguishes between discourse and non-discourse usages of connectives, and where the arguments of a connective could be found. during a semester-long training period, the annotators were encouraged to form their own ideas about how discourse works. next, the annotators studied the annotation tool and the guidelines and started annotating the initial set of connectives. the annotation tool was specifically devised for this project. briefly, it uses the stand-off annotation methodology and produces xml files as annotation data (aktaş et al., 2010). neither the discourse relations nor the characteristics of the discourse segments (e.g. whether they should be full clauses or not) are spelled out in the annotation guidelines. this method was useful because it allowed the annotators to use their native-speaker intuitions in deciding about the syntactic type and the span of a connective’s arguments. regarding the span of a connective’s argument, the annotators were only told to follow the “minimality principle”, which requires them to mark the shortest text spans that are necessary and sufficient to interpret a discourse relation encoded by the connective (prasad et al., 2007). 2 the capital letters are used to capture the cases where a vowel agrees with the vowel harmony rules of the language. the letter i may be resolved as any of the high vowels in the language. the letter a may be resolved as e or a. turkish discourse bank 177 2.2 the annotation procedures and inter-annotator agreement the annotators worked independently, as a group, or in a procedure we named pair annotation, adapted from pair programming (demirşahin, et al., 2012a). in the group annotation method, one independent annotator produces a set of annotations, and the other two annotators go over this annotation set together, suggesting changes when necessary. any disagreements are resolved in agreement meetings. on the other hand, the pair annotation procedure involves one annotator working independently and two annotators working together as a pair, producing two sets of independent annotations. the agreement of pair annotations is measured between the independent annotator and the pair of annotators, treating them as a single annotator. of the total 8483 relations in the tdb 1.0, 3804 (44.84%) were annotated by three independent annotators, 3985 (46.98%) by pair annotation, and only 694 (8.18%) were annotated by group annotation. to evaluate the reliability of the argument span annotations produced by three independent annotators, or by a pair of annotators and an independent annotator, we measured agreement using fleiss’ kappa (fleiss, 1971) described as: 3 measures the degree of agreement attainable above chance, and gives the degree of agreement actually attained above chance. since each connective has two arguments, we calculated agreement over these text spans separately. for arg1 and arg2, we formed separate agreement tables similar to the table fleiss uses (1971:379), where at least two annotators assign the words of a text into two categories (select/exclude) and we recorded the number of judgments a word receives for each category. we measured agreement over the spans identified by the annotators as the boundaries of the argument spans, which we took as the first and the last words of each argument selected by the annotators (yalçınkaya, 2010). to evaluate supplementary material annotations, we used the exact match criterion (miltsakaki et al., 2004). in this method, agreement for any supplementary material (supp1 or supp2) is recorded as 1 when all annotators make identical textual span selections, and 0 otherwise. agreement is calculated by the number of exact matches found in the total number of annotations annotated by all annotators and given as a percentage (cf. appendix a). 2.2.1 discourse adverbials one of the most challenging issues in the annotation process was that of determining the location and span of the arguments of discourse adverbials. table 2 provides discourse adverbials in the data and the kappa measures of their first and second arguments, where applicable. the remaining discourse adverbials, for which kappa statistics are not measured, are also provided for the sake of completeness (see footnote 5). according to table 2, except for aslında, ‘in fact’, örneğin ‘for example’, mesela ‘to exemplify’ and böylece ‘thus’, the inter-annotator agreement measures for the first arguments of discourse adverbials are lower than the envisaged 0.80 threshold (artstein and poesio, 2008). 4 a preliminary analysis shows that these low agreement measures are largely due to the anaphoric characteristics of these connectives. similar to anaphors whose antecedents can be ambiguous, the first arguments of discourse adverbials can also be ambiguous since their location is not constrained by adjacency (forbes-riley et al., 2006; webber et al., 2003). the second arguments of discourse adverbials are mostly annotated with high agreement since their location is predictable by adjacency. the connective öte yandan ‘on the other hand’, however, yielded a low inter-annotator measure for its arg2. our exploratory 3 we did not measure the agreement on the discourse connective. 4 such low agreement results are not surprising because agreement on the exact boundaries of a text span tend to be low in discourse, as discussed by artstein & poesio (2008:580-583). zeyrek, demi̇rşahi̇n, sevdi̇k çalli, and çakici 178 analyses show that this connective appears in argumentative texts and tends to link text spans which are more than one sentence long. due to this, the annotators do not agree on the boundaries of arg2; in particular, the right edge of arg2 is a source of disagreement. such disagreements are always resolved in agreement meetings. kappa measures search item (# of annotations) gloss # of annotators arg1 arg2 aslında (81) in fact 2 0.81 0.85 ayrıca (108) in addition 3 0.66 0.84 böylece (85) thus 2 0.90 0.99 dahası (9) furthermore 3 0.71 0.90 gene de (26)* still 1 halbuki (17)* however 1 mesela (13) to exemplify 3 0.92 1.00 neticede (1) eventually 3 ne ki (14)* howbeit 1 ne var ki (32)* even so 1 oysa (136) however 3 0.78 0.91 örneğin (64) for example 3 0.87 0.92 örnek olarak (2) to illustrate 3 sonuçta (10) finally 3 0.70 0.87 sonuç olarak (5) as a result 3 0.67 1.00 söz gelimi (6)* for instance 1 taraftan (3) on the other hand 3 tersine (11) in contrast 3 0.77 1.00 yalnız (12)* it is just that 1 öte yandan (70) on the other hand 3 0.55 0.66 yine de (65)* still 1 table 2. kappa measures for the arguments of discourse adverbials in the data5 2.2.2 discourse relations expressed by nominalizations the second class of discourse connectives mentioned above, i.e., subordinators take nominalizations as complements, i.e. clauses that are “desententialized” to varying degrees (lehmann, 1988). except for a few strict cases, nominalizations are not annotated in the pdtb. however in turkish, nominalized clauses are so common as arguments of not only the subordinators but also the coordinators that we would have missed an important aspect of the language if we left them out. complex subordinators are annotated in the tdb 1.0 using the morphological features of their arg2 as a clue. complex subordinators have basically two parts; a connective (often a postposition) and nominalizing suffixes which reduce a subordinate clause to varying degrees, causing it to lose its illocutionary force as well as tense and aspect. the nominalized clauses are based on three types of suffixes: (a) clauses based on the factive nominalizer –dik and the nonfactive nominalizer –acak, (b) clauses based on the infinitives –ma or –mak, (c) clauses formed on the nonfinite nominal marker –iş (csató, 1998:230). in the annotation process, the annotators are told to notice these nominalizing suffixes as indicators of nominal clauses that have predicative potential (cf. appendix b for examples). we annotate the independent parts of the complex subordinators by selecting the independent part and the arg2 in its entirety. at a later stage, the suffixes will be separated by postprocessing to analyze the frequency of the nominalization types. this will also enable automatic sense 5the stars show that a single set of annotations was created by the group annotation procedure, 2 indicates that the annotations were created via the pair annotation method, the dashes indicate that inter-coder reliability was not calculated. turkish discourse bank 179 disambiguation of certain connectives, e.g. için ‘for/so as to', whose goaland cause-driven senses can be distinguished via the nominalizing suffixes. 3 updates in the annotation guidelines in this section, we present two major updates in the annotation guidelines, namely the decision to annotate phrasal expressions and the introduction of the shared tag to the annotation scheme. we also discuss how we adapted the pdtb's modifier and supplementary tags to turkish. 3.1 phrasal expressions as we explained in section 2.1, the annotation procedure initially started with a given set of connectives. however, the annotators soon discovered that there are polymorphous occurrences of the independent parts of complex subordinators. for example, the complex subordinator sonra ‘after’, belongs to the same family of connectives with the phrasal expression sonra ‘after this’, and its variants, e.g. önce .. sonra ‘first .. then’. we therefore decided to allow the annotators to determine all such occurrences of the given set of connectives, rather than restricting them with the given connectives. in this way, we would achieve a wider coverage of the productive means of establishing discourse coherence. phrasal expressions are marked as a form of alternative lexicalization in the pdtb (prasad et al., 2010); they are annotated as a form of complex connectives in a german corpus (stede & heintze, 2004). the number of annotated subordinators and phrasal expressions form a sizeable portion of all the annotations in the tdb 1.0. we identified 77 search items. twenty-eight of these items returned connective types that participate in subordinating relations, and 27 of them in phrasal expressions. of the total 8483 relations annotated in the corpus, 2284 (26.92%) are signaled by a subordinator, and 482 (5.68%) are signaled by a phrasal expression containing a deictic item. in appendix a, we provide all annotated complex subordinators (including their polymorphous occurrences) and kappa values of arg1 and arg2 as described in section 2.2. appendix a shows that the annotators disagree about the arg1 of some connectives, e.g., rağmen ‘despite’. although the reasons for disagreements may vary, we noticed that the flexible word order of turkish is an important source of disagreements. we discuss this more in section 3.2 below. 3.2 word order variability of turkish and the addition of the shared tag turkish is predominantly a sov language with a large degree of word order flexibility. scrambled elements cause difficulties for the annotators because they may disagree whether the scrambled elements belong to arg1 or arg2. to overcome this problem, we introduced the shared tag (subject, object, adverbial phrase), which is essentially a syntactic tag simply helping to mark the shared elements in a discourse relation no matter where they are in the sentence. in this way, the discourse relation itself is determined with more ease and confidence. table 3 provides the frequency of the subordinators with the shared tag. shared tag no shared tag total search item gloss count percent count percent count percent için for/so as to 155 14.07 947 85.93 1102 100.00 sonra after 77 10.80 636 89.20 713 100.00 kadar as well as/until 37 23.27 122 76.73 159 100.00 gibi as 35 15.35 193 84.65 228 100.00 amacıyla with the aim of 23 35.94 41 64.06 64 100.00 zaman when 15 9.43 144 90.57 159 100.00 karşın regardless of 11 15.49 60 84.51 71 100.00 önce prior to 11 8.21 123 91.79 134 100.00 halde in spite of 10 16.39 51 83.61 61 100.00 zeyrek, demi̇rşahi̇n, sevdi̇k çalli, and çakici 180 table 3. the frequency of shared elements in the subordinators and the related phrasal expressions example (1) shows the sentence-medial usage of rağmen ‘despite’, where arg1 is shown in italics, and arg2 is rendered in bold letters. the subject, which is not in its canonical sentenceinitial position in this case, is shown between curly brackets and annotated as the shared material. (1) sınırlı olmasına rağmen {bu devrimci kongreler}, sarayın değil, halkın demokratik ihtilalinin eseriydiler. despite the fact that they were limited, {these revolutionist congresses} were not a result of the empire but the people’s democratic rebellion. 3.3 modifiers the tag modifier is primarily used to show the modifier of a connective as in the pdtb (example (2), underlined together with the connective), where the connective is taken as the head, and the adverb as the modifier. different from the pdtb, we also use this tag to specify adverbs modifying the discourse relation as a whole. we refer to such tokens as modifier of a relation as in (3) although they are all marked as mod. (2) geri dönüp kanepeye uzanıyorum. az sonra ezgi başlıyor. i go back and lie on the couch. a little later, the melody starts. (3) (belki de) ona karşı çok iyi ol-duğ-um için bıraktı beni. (perhaps) he left me because i treated [treat-dik-agr] him too well. in the tdb 1.0, a total of 540 relations are tagged with modifiers. the most heavily modified connective is sonra ‘after/later’, where 220 of 713 instances are modified: 138 of these modifiers indicate duration, and 78 are focus particles. the focus particle da is the most frequent modifier with 262 instances. the temporal modifier daha ‘much/more’ is the second most frequent modifier with 83 instances. eleven classes of modifiers are annotated in the tdb 1.0. adverbs such as şimdiye kadar ‘until now’, neyse ki ‘luckily’, ne yazık ki ‘unfortunately’, etc. are not tagged as modifiers. here, we are in agreement with the pdtb, where such clausal adverbs are selected together with the argument in which they appear. table 4 shows the modifiers and their frequencies in the tdb 1.0. modifier class example gloss count percent focus da focus particle (fp) 265 49.07 temporal üç gün sonra three days later 170 31.48 intensifier tam aksine just to the contrary 26 4.81 counterfactuality sanki … gibi as though 25 4.63 epistemic belki de bunun için perhaps fp because of this 17 3.15 interrogative bu yüzden mi is this the reason 14 2.59 quantifier bütün bunlara rağmen despite all these 9 1.67 condition ancak bundan sonra only after this 5 0.93 negation için değil not because of this 5 0.93 qualifier çarpıcı örnek olarak as a striking example 3 0.56 pragmatic peki o zaman well, ok then. 1 0.19 total 540 100.00 table 4. the frequency of the modifier tags in the tdb rağmen despite 7 9.09 70 90.91 77 100.00 birlikte together/though 6 18.18 27 81.82 33 100.00 ardından after 5 7.04 66 92.96 71 100.00 total 392 13.64 2840 86,35 2872 100.00 turkish discourse bank 181 3.4 the supplementary material we identified two types of material that supplement the arguments: (a) the material that makes the semantic contribution of the argument more specific, which is how the pdtb uses this tag, and (b) the antecedent of a deictic item in one of the arguments. example (4) shows the latter function of the supp tag, where the deictic item is underlined and the supplementary text shown between straight lines “|”. (4) |ante mutlaka yalnız görüşmeleri gerektiğini| anlatmaya çalıştı. sonunda padişah buna razı oldu ve huzurunda bulunan herkesi dışarı çıkardı. ante tried to explain |that they had to meet privately|. eventually, the sultan agreed with this and asked everyone out. we used the shared supp tag for indicating the antecedent of a deictic element in the shared material (example (5)). in the example, the shared material (the subject in this case) is rendered between curly brackets; its referent is put between double straight lines. (5) simitis, "||türkiye'ye müzakere tarihi verildiği ortamda kıbrıs sorunu çözülmüş olacaktır herhalde ||. {bu ikisi} birbirine bağlı değil ama beraber gitmeleri gereken iki süreç" dedi. simitis said, “||when turkey is given a date for the discussions, the cyprus problem will probably have been solved||. {these two} are not linked with each other but they are two processes that must go together.” the tdb 1.0 has 869 supplementary annotations for arg1 (10.24%) and 337 supplementary annotations for arg2 (3.97%). in future research, the use of the supplementary tag for the antecedents of discourse deictic items as in (4) and (5) will enable us to compare the role of deictic items in the discourse relation with the discourse relation itself. 4. summary in this paper we described the turkish discourse bank, where discourse connectives are annotated with their two arguments in the style of the pdtb. we focused on the challenges posed by the morphological richness of turkish as well as its varying word order. we described those aspects of the annotation scheme that are different from the original language english, focusing on the fact that searching for discourse relations between clauses may not capture all the means for presenting information in turkish. one departure from the pdtb is annotating phrasal expressions, which revealed additional discourse relations based on connectives. future research will elucidate whether the senses of complex subordinators and the associated phrasal expressions differ systematically. in addition to this, when we reach a larger coverage in the tdb (e.g., by annotating implicit connectives), we will be able to compare the connective-based discourse relations with the non-connective-based ones, obtaining data for cross-linguistic comparison. acknowledgments we gratefully acknowledge the support from scientific and technological research council of turkey (tübi̇tak, project no. 107e158) and middle east technical university research funds. we also thank three anonymous reviewers for their extensive feedback, cem bozşahin for his comments and i̇hsan yalçınkaya who helped with the statistics. all remaining errors are ours. appendix a this appendix provides fleiss’ kappa measures for arg1 and arg2 of the search items that retrieved subordinators, the related phrasal expressions and their variants (where applicable). the appendix also shows the exact match measures for the text spans that supplement arg1 and arg2, where relevant (see section 2.2). the connective sonra ‘after’ was annotated in 4 consecutive stages and agreement was measured for each stage. zeyrek, demi̇rşahi̇n, sevdi̇k çalli, and çakici 182 search item (# of annotations) gloss # o f a n n o ta to r s kappa measures exact match agreement subordinator phrasal expression variants arg1 arg2 supp1 supp2 no of exact matches total no of ann. % agr no of exact matches total no of ann. % agr aksine (13) *6 contrary to contrary to this 1 amacıyla (64) with the aim of 3 0.69 0.93 0 0 0 0 ardından (71) after after this first … then 2 1.00 0.99 7 7 100 2 2 100 beri (4) since (temporal) 2 1 1 100 0 0 birlikte (33)* together/ though nevertheless 1 bu yana (10) since this time (temporal) 2 1.00 1.00 1 1 100 0 0 dolayı (21) owing to 3 0.98 1.00 0 0 0 1 0 dolayısıyla (66) in consequence of consequently 3 0.78 0.97 2 2 100 1 4 25 ek olarak (1) in addition to this 3 na na 0 0 0 0 gibi (228) as 2 0.94 0.95 28 29 96.55 9 9 100 halde (61) in spite of inspite of this/that 2 0.87 0.93 0 0 0 0 için (1102) for, so as to for this/that for … for 3 0.81 0.92 3 7 42.86 9 18 50 içindir (4) because of because of this/that it is because of this/that 2 0 0 0 0 kadar (159) as well as, until 2 0.84 0.99 2 3 66.67 5 5 100 karşılık (28) although nonetheless 1 karşın (71) regardless of regardless of this/that irregardless 3 0.86 0.84 0 0 0 0 nedenle (117) for this/that reason 2 0.94 0.99 13 15 86.67 2 2 100 nedenlerle (4) for these reasons for the reasons above 2 0 0 0 0 nedeniyle (42) for the reason that 2 0.96 0.97 2 2 100 4 4 100 neticesinde (1) as a result of this 3 na na 0 0 0 0 önce (134) prior to prior to this first first … then/now 2 1.00 1.00 4 5 80.00 1 1 100 2 0.84 0.88 4 4 100 1 1 100 ötürü (11) due to due to this/that due to this/that reason 2 1.00 0.94 0 0 1 1 100 rağmen (77) despite despite this/these 3 0.73 0.78 0 0 0 0 sayede (5) thanks to this/that 2 1.00 1.00 2 2 100 0 0 sayesinde (3)* thanks to 1 sonra (713) after after this first.. then, now.. then, at the beginning.. then, at the beginning..after this 3 0.85 0.91 5 10 50 7 8 87.50 2 0.91 0.96 18 18 100 13 13 100 2 0.89 0.94 8 9 88.89 5 5 100 2 0.89 0.98 0 0 0 0 sonucunda (12) as a result of as a result 3 0.78 0.78 0 0 0 0 yüzünden (5) since (causal) 2 1.00 1.00 1 1 100 1 1 100 zaman (159) when at that time whenever … then 2 0.97 0.98 19 21 90.48 10 11 90.91 6 the stars indicate that the annotations were created by the group annotation procedure; the dashes show that inter-coder reliability was not calculated. turkish discourse bank 183 appendix b this appendix provides examples that are used to guide the annotators in annotating subordinators and their nominalized arguments by using the nominalizing suffixes that have predicative potential as clues. the normal order of the arguments of a subordinator is arg2-arg1. the suffixes are shown in small caps, both the connective and the corresponding suffixes are underlined. 7 an example for nominalized clauses based on the factive –dik 8 : (a) üzül-düğ-ü kadar şaşırmıştı da. she/he was surprised as much as she/he was saddened [sad-pass-dik]. an example for nominalized clauses based on the infinitive –mak: (b) gör-ün-me-mek için hemen duvara yaslandı. in order not to be seen [see-pass-neg-mak] he immediately leaned against the wall. an example for nominalized clauses based on –iş: (c) i̇haleli sisteme geç-iş-in ardından bu ihalelere katılmayan 10 kadar firma takibe alındı. after the shift [shift-iş-gen] to the bidding system, legal action was taken against circa 10 firms that did not undertake bidding processes. references berfin aktaş, cem bozşahin and deniz zeyrek (2010). discourse relation configurations in turkish and an annotation environment. in proceedings of the fourth linguistic annotation workshop, pages 202-206, uppsala, sweden. ron artstein and massimo poesio (2008). inter-coder agreement for computational linguistics. computational linguistics, 34(4): 555-596. nicholas asher (1993). reference to abstract objects in discourse. kluwer academic publishers, dordrecht. nicholas asher and alex lascarides (2003). logics of conversation. cambridge university press, cambridge, uk. éva á. csató (1988). turkish. in lars johanson and éva á. csató, editors. (1998). the turkic languages (routledge language family series), pages 203-235, routledge, london. işın demirşahin, i̇hsan yalçınkaya and deniz zeyrek (2012a). pair annotation: adaption of pair programming to corpus annotation. in proceedings of the 6th linguistic annotation workshop, pages 31–39, jeju, republic of korea. işın demirşahin, ayışığı sevdik-çallı, hale balaban, ruken çakıcı and deniz zeyrek (2012b). turkish discourse bank: ongoing developments. in proceedings of the first workshop on language resources and technologies for turkic languages, pages 15-19, the 8th international conference on language resources and evaluation (lrec’12), istanbul, turkey, may 2012. katherine forbes-riley, bonnie webber and aravind joshi. (2006). computing discourse semantics: the predicate-argument semantics of discourse connectives in d-ltag. journal of semantics, 23(1): 55-106. barbara j. grosz and candice l. sidner. (1986). attention, intentions, and the structure of discourse. computational linguistics, 12(3): 175-204. christian lehmann (1988). towards a typology of clause linkage. in john haiman and sandra a. thompson, editors. clause combining in grammar and discourse, pages 181-225, john 7 in the examples, the following abbreviations are used: dat: dative suffix, inst: instrumental suffix, gen: genitive suffix, pass: the passive morphemes. 8 the capital letter k may be resolved as k or ğ, the letter which refers to the lengthening of the previous vowel. zeyrek, demi̇rşahi̇n, sevdi̇k çalli, and çakici 184 benjamins publishing, amsterdam; philadelphia. william c. mann and sandra a. thompson (1988). rhetorical structure theory: toward a functional theory of text organization. text 8(3): 243-281. mitchell marcus, grace kim, mary a. marcinkiewicz, robert macintyre, ann bies, mark ferguson, karen katz and britta schasberger (1994). the penn treebank: annotating predicate argument structure. in proceedings of arpa human language technology workshop, pages 114-119. eleni miltsakaki, rashmi prasad, aravind joshi and bonnie webber (2004). the penn discourse bank. in proceedings of the fourth international conference on language resources and evaluation, pages 2237-2240, lisbon, portugal. joseph l. fleiss (1971). measuring nominal scale agreement among many raters. psychological bulletin, 76(5): 378-382. megan moser and johanna d. moore (1996). toward a synthesis of two accounts of discourse structure. computational linguistics, 22(3): 409-419. kemal oflazer, bilge say, dilek z. hakkani-tür and gökhan tür (2003). building a turkish treebank. in anne abeillé, editor, treebanks: building and using syntactically annotated corpora, pages 261-277, kluwer academic publishers. livia polanyi and martin h. van den berg (1996). discourse structure and discourse interpretation. in proceedings of the tenth amsterdam colloquium, pages 113-131, amsterdam. rashmi prasad, aravind k. joshi and bonnie webber (2010). realization of discourse relations by other means: alternative lexicalizations. proceedings of coling 2010, pages 1023-1031, beijing, china. rashmi prasad, eleni miltsakaki, nikhil dinesh, alan lee, aravind joshi, livio robaldo and bonnie webber (2007). the penn discourse treebank 2.0 annotation manual. technical report 203, institute for research in cognitive science, university of pennsylvania, philadelphia, pennsylvania. bilge say, deniz zeyrek, kemal oflazer and umut özge (2004). development of a corpus and a treebank for present-day written turkish. in kamile imer and gürkan doğan, editors, current research in turkish lingustics, pages 183-192, eastern mediterranean university, cyprus. wojciech skut, brigitte krenn, thorsten brants and hans uszkoreit (1997). an annotation scheme for free word order languages. in proceedings of the fifth conference on applied natural language processing, pages 88-95, stroudsburg, philadelphia. manfred stede and silvan heintze (2004). machine-assisted rhetorical structure annotation. in proceedings of the 20th international conference on computational linguistics, pages 425431, geneva, switzerland. bonnie webber, aravind joshi, eleni miltsakaki, rashmi prasad, nikhil dinesh, alan lee and katherine forbes (2006). a short introduction to the penn discourse tree bank. in peter j. henrichsen and peter r. skadhauge, editors, treebanking for discourse and speech, pages 928, samfundslitteratur press, copenhagen. bonnie webber, matthew stone, aravind joshi and alistair knott (2003). anaphora and discourse structure. computational linguistics, 29(4): 545-587. bonnie webber, marcus egg and valia kordoni (2011). discourse structure and langugage technology. natural language engineering, 1(1): 1-54. i̇hsan yalçınkaya (2010). an inter-annotator agreement methodology for the turkish discourse bank. unpublished ms thesis, middle east technical university. deniz zeyrek, işın demirşahin, ayışığı b. sevdik-çallı, hale ö. balaban, i̇hsan yalçınkaya and ümit d. turan (2010). the annotation scheme of the turkish discourse bank and an evaluation of inconsistent annotations. in proceedings of the fourth linguistic annotation workshop, pages 282-289, uppsala, sweden. microsoft word tosunvaid2018-ms010718.docx © 2018 sümeyra tosun and jyotsna vaid this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). 128 dialogue & discourse 9 (1) (2018) 128-162 doi: 10.5087/dad.2018.105 activation of source and stance in interpreting evidential and modal expressions in turkish and english sümeyra tosun sumeyratosun@gmail.com independent researcher jyotsna vaid jvaid@tamu.edu texas a&m university editor: manfred stede submitted 06/2017; accepted 06/2018; published online 08/2018 abstract languages differ in whether and how they mark the source of evidence about a narrated event (evidentiality), with some languages (e.g., turkish) requiring users to distinguish between firsthand and nonfirsthand sources in their grammar and others (e.g., english) allowing source to be signaled optionally at the lexical level. information conveyed firsthand is generally accorded more epistemic weight (i.e., more likely to be believed) than nonfirsthand information. however, the relationship between source type and belief that the asserted event occurred has not been examined for nonfirst hand evidential sources (hearsay, inference, assumption, conjecture), or for various modal categories (necessity, advice, probability, possibility). to address this gap, the present research compared turkish vs. english users’ source and confidence (stance) judgments for sentences presented in each of four source categories and four modal types. our results suggest that, besides indicating source type, evidential markers also convey the epistemic value of the reported proposition and that epistemic modal markers, besides indicating source reliability, also indicate source type. moreover, turkish users showed a more demarcated system of interpretation of evidential markers (with inference and assumption differentiated from hearsay and conjecture) while english speakers showed a more homogeneous hearsay source interpretation. the findings provide empirical support for theorized claims about the close relationship between evidential and modal structures, while also uncovering some intriguing group differences. keywords: evidentiality, epistemic modality, turkish, english, source of evidence, modal auxiliaries 1. introduction evidentiality is a linguistic property found in many languages. it primarily indicates the source of knowledge of an asserted proposition (e.g. aikhenvald, 2004; aksu-koc & slobin, 1986; chafe, 1986; plungian, 2001). the source may be firsthand, based on sensory information (witnessing some occurrence), or nonfirsthand information, such as inference (from observable or tangible evidence), assumption (based on intuition, logical reasoning, previous experience, or general evidentiality and modality 129 knowledge), conjecture (based on second hand information but the source was unspecified), or hearsay (hearing about the described situation from someone else). source of evidence may be marked morpho-syntactically in the grammar to distinguish obligatorily the sources how the information is acquired (see aikhenvald, 2004 for further discussion of typological differences in the coding of evidentiality). alternatively, languages may not require users to mark source of evidence in the grammar but allow this information to be conveyed optionally, in the lexicon. some scholars have suggested that evidentiality does not just signal source of evidence of an asserted event but also the speaker’s stance or belief about the certainty regarding the occurrence of the event and/or the reliability of the source (e.g., chafe, 1986; delancey, 2001; faller, 2002; palmer, 1986; van der auwera & plungian, 1998). to date, discussions of the nature and scope of the construct of evidentiality and its relation to epistemic modality1 have largely relied on logical reasoning or on evidence drawn from linguistic data (e.g., de haan, 1999, 2004) without investigating how actual language users interpret evidential and modal expressions. the present study seeks to take a psycholinguistic approach to this issue by considering how actual users of a language interpret evidential and modal structures. specifically, our study assesses the relative degree to which source and stance/certainty information is co-activated when interpreting evidential and modal expressions. we further examine how such expressions are interpreted in a language in which evidentiality markers are grammaticalized (turkish) versus simply lexicalized (english). this comparison allows us to ask if obligatory vs. optional marking of source information differentially affects the pattern of activation of source and stance information in response to modal and evidential expressions. our study sought to test how native speakers of two different languages turkish and english interpret evidential and modal expressions in their respective language. specifically, we compared the two groups’ interpretations of expressions framed using four different types of evidential adverbs (reportedly, apparently, presumably, and supposedly, and their turkish verbal and/or morphosyntactic counterparts) and four different types of modal auxiliary forms, representing different levels of certainty (must have, should have, could have, might have, and their turkish counterparts). providing language users’ judgment data to bear on the issue of the relationship between evidentiality and epistemic modality contributes to the longstanding theoretical debate on the nature of evidentiality as a property of language. yet apart from its theoretical significance, the issue of how language users interpret evidential expressions also has practical significance. consider, for example, if in a presidential debate, one of the candidates says: “as recently reported by [some agency], the inflation rate reportedly dropped during my presidential term.” the voters may interpret this to mean that there was a decrease in the inflation rate, as determined by an outside agency. in this case, they would interpret the evidential marker (reportedly) as indicating the source of knowledge only. thus, they may credit the candidate with honesty in indicating the source or basis for their claim. on the other hand, voters may interpret the candidate’s use of reportedly as suggesting that the inflation rate decrease was not likely to have happened, because it was described secondhand. this interpretation accords epistemic value to an evidential source (hearsay in the example) and may change how voters regard the candidate. there is also a stance interpretation here – e.g., the speaker could be using an ironic tone. thus, it is crucial in the political sphere to understand the range of possible interpretations of utterances containing evidential markers as they may lead to very different attributions of a politician. this issue is equally important in other spheres of life, such as doctor-patient interactions, as when a patient is told that a particular medication could have or probably has certain side effects (see 1 linguistically, modality can be classified in different ways such as deontic, alethic or epistemic. alethic and epistemic modality is most often associated. alethic modality refers to “the truth in the world” and epistemic modality refers to “the truth in an individual’s mind”. in this study we focus on epistemic modality. tosun and vaid 130 segalowitz et al., 2016). these examples underscore the way in which the theoretical debate about the relative independence of evidentiality and epistemic modality could have very real consequences. before turning to our study, it is important to review how evidentiality and epistemic modality operate in english and turkish, and to consider how the relationship between evidentiality and epistemic modality has been theorized. we also consider previous studies of relevance. 1.1. evidentiality in turkish and english the prevailing definition of evidentiality, and one that we will adopt for the present study, is that it is an independent property of language that marks (variously, in the grammar or the lexicon) the source of evidence of an asserted event (aikhenvald, 2004). languages may differ in how they code source of evidence and to what degree they distinguish between different types of sources. the distinction of relevance for the languages of interest in the present study is between obligatory, morpho-syntactic coding of firsthand versus nonfirsthand accounts of a past event, and optional, lexical coding of an event. in turkish, evidentiality is coded in the grammar; that is, expressing the source of evidence for an asserted event is obligatory. turkish is an example of what aikhenvald (2004) refers to as two-choice evidential languages; it makes a distinction between firsthand and all other sources of knowledge. aksu-koc and slobin (1982, 1986) refer to the two as ‘direct’ vs. ‘indirect’ sources of knowledge. firsthand source expression. the turkish firsthand source marker –di (realized as –di, -dı, -dü, du, -ti. –tı, -tü, -tu)2 conveys directly experienced source of knowledge (slobin & aksu, 1982). directly experienced sources refer only to visual sensory sources of evidence. (1) handan okul-a git-ti. handan school-dat go-evid. ‘handan went to school, i saw.’ in (1), the speaker saw when handan was leaving home to go to school. according to aksu-koc and slobin (1986), first-hand source markers are also used to express knowledge that is expected or unsurprising. (2) köprü tadilat-ı trafiğ-i alt-üst et-ti. bridge construction-gen traffic-acc upside-down make-evid ‘construction of the bridge disturbed the traffic.’ in example (2), the speaker uses the firsthand marker even though she has not witnessed the traffic jam, because it would not be surprising at all that there would be a traffic jam when the bridge is under construction. another example of firsthand marker usage for non-witnessed events is news reports. even though the speakers, themselves, do not witness the event, they report it using the firsthand marker. choosing the firsthand marker shows that the speaker intends to indicate that she witnessed the event firsthand and/or intends to convey her certainty about the event (aksu-koc, 2000; kornfilt, 1997). nonfirsthand source expression. nonfirsthand sources in turkish are marked with the suffix –miş (realized as –miş, -mış, -müş, -muş) on the verb. this marker is derived from the resultative and stative suffix –miş (slobin & aksu, 1982). the same marker is used to cover three different information sources: reportative, inferential, and perceptive/mirative (johanson, 2000, 2003). in the reportative source, the information is acquired from someone else; in such cases, the corresponding proposition is marked using the nonfirsthand source marker (aikhenvald, 2004). 2 the variations in the suffixes are due to vocal harmony. evidentiality and modality 131 (3) handan okul-a git-miş. handan school-dat go-evid. ‘handan reportedly went to school.’ in example (3), the speaker heard from someone else that handan went to school, as she herself did not see handan leaving home to go to school. english equivalents of reportative source include reportedly, allegedly, as they say/said, and all of the reported speech versions. the basis of the inferential source is reflection and reasoning arising from inference from results and/or from reasoning per se. in a different context the same sentence (3) may have a different interpretation. (4) handan okula gitmiş. ‘handan apparently went to school.’ in the example (4) the speaker inferred that handan went to school. she could not find handan and her school bag at home or she knew that handan had a class at the time, thus she inferred that handan went to school without seeing personally when this happened. english equivalents of inferential –miş include apparently, presumably, as far as … understand/understood etc. (johanson, 2003). the basis of the perceptive/mirative source is direct sensory perception other than visual sensory knowledge such as smelling, or hearing, or else unexpected information (johanson, 2003). similarly as the sentence (4), the same sentence may be interpreted differently in a different context. (5) handan okula gitmiş. ‘handan went to school, i was surprised to hear’ in example (5) the speaker heard when handan closed the door and left home or she saw handan in the school but she also knew that handan was sick and was not expected to be on campus on that day. english equivalents of the perceptive source include it appears/appeared that, it turns/turned out that, as….can/could see that, hear etc. the perceptive use of the evidential is also interpreted in terms of relative novelty, sudden discovery, and new knowledge with an unprepared mind (aksu-koc & slobin, 1986; johanson, 2000). compound structure option. the suffix –miş is also used in compound form with two other suffixes, -miş or -dir, to convey the source of information. (6) a. handan okul-a git-miş-miş. handan school-dat go-evid-evid ‘handan reportedly/supposedly has/had gone to school.’ b. handan okul-a git-miş-tir3. handan school-dat go-evid-ind ‘handan presumably went to school.’ in example (6) the doubling of the suffix –miş is used to represent a situation that already happened in the past with a nonfirsthand source and it is also reported (johanson, 2003). thus, it is a thirdhand evidential. this type is mostly used sarcastically. the speaker uses this form if she believes that handan never actually went to school. in example (6b), the suffix –dir is attached after the nonfirsthand suffix and represents inference from reasoning, previous knowledge and knowledge about habitual events (aksu-koc, 2009, aksu-koc & alici, 2000). in this example, in the context, because handan always leaves home at 8:00 am for school, the speaker indicated that handan presumably went to school when she found out that it was already 8:30 am. in contrast to turkish, english is an example of a language in which evidentiality has no mandatory grammatical encoding. rather, speakers can choose to indicate if they witnessed something firsthand or heard about it from some other person, or assumed it had occurred. these 3 due to vocal harmony, the suffix –dir turns to -tir. tosun and vaid 132 different possibilities are signaled through lexical choices (e.g., adverbs like reportedly, presumably, or through phrases such as, i heard that, or i saw that).4 1.2. epistemic modality in turkish and english epistemic modality has been variously defined. givón (1982) defines it as a probability of the proposition on a scale between necessary (which has a probability of 1.0) and impossible (which has a probability of zero), where probable and possible are in the middle (with a probability of 0.5). aijmer (1980) defines it as “the speaker’s evidence and, degree of certainty” (p. 11). palmer (1986) defines the term epistemic as the “degree of commitment by the speaker to what he says” (p. 51). according to chafe (1986), epistemic modality codes the speaker’s attitude toward his/her knowledge of a situation. van der auwera and plungian (1998) define epistemic modality as the judgment of the speaker. according to nuyts (2001), epistemic category is evaluation of the chances of an event’s occurrence. thus, we see that epistemic modality has been defined as attitude; judgment or commitment of the speaker towards how likely it is that the situation described would occur in a possible or actual world. for the purpose of the present study, epistemic modality is defined as a language user’s degree of certainty or confidence in whether the asserted event actually occurred. in terms of how epistemic modality operates in the languages under study, this topic has not received as much attention in the literature, particularly as regards turkish. kerimoglu (2010) noted that there is a close relationship between evidentiality and certain modals in turkish. the morpho-syntactic marker –miş olabilir5 (used to signal could or might) has a probability meaning, whereas the marker –miş olmalı (used to indicate must) represents deduction and is used for strong possibilities; finally, the marker –malı (used to indicate should) has the meaning of obligation. moreover, kerimoglu stated that, along with the modal markers, the value of certainty is also manipulated by lexical markers, such as the turkish equivalents of the lexical items probably, possibly, usually, absolutely, and perhaps. the grammaticalized markers of modality (miş olmalı, -miş olabilir) were specifically proposed to convey the source of inference and assumption as well. kerimoglu (2010) and kornfilt (1997) further suggested that the suffix –dır is a modal expression, whereas aksu-koc (2009) refers to it as an evidential marker indicating assumption. to date there has been no empirical investigation of the interpretation of epistemic modals in turkish. epistemic modality in english has been the subject of more work (see palmer, 2013, for detailed review). although many modal forms exist, of particular interest of past events for the present study are those that use the auxiliary have before the past tense of the main verb, as in she must have left by now. the modal must, classified as epistemic necessity, conveys the speaker’s confidence about the occurrence of a reported event. should, could and might are classified as tentative forms of epistemic modals. should conveys extreme likelihood of occurrence of the asserted event. could denotes possibility of occurrence. might indicates a little less certainty about the reported event. although in some contexts could is replaceable by might, might is more likely to be interpreted in a tentative sense. in summary, epistemic modality in turkish relies on the use of three morpho-syntactic markers attached to the past tense of a verb which contrast must, should, and could/might, and on 4 one person may receive the same information from different sources and indicate only one source while conveying the information. which of the sources is preferred to indicate is more of an issue of the hierarchy of the sources (for more information about evidential hierarchy please see willett, 1988 and de haan, 1998). rather than evidential hierarchy, what we aimed to measure is whether language users systematically select one specific source for a specific expression. 5 the morpheme –miş in modals serves as a past participle suffix to change a verb to an adjective – it is not in this case interpreted as an evidential morpheme. evidentiality and modality 133 lexical items. in english, some epistemic modals may be represented as verb auxiliary forms. although must is acknowledged to represent the highest level of certainty, there is less consensus on the relative interpretation of the other forms. 1.3. how evidentiality and epistemic modality may be related there are five possible ways in which evidentiality and epistemic modality could be related in principle. we term these complete disjointment, inclusion (epistemic modality as a subtype of evidentiality), inclusion (evidentiality as a subtype of epistemic modality), overlap and identity. the positions outlined below have derived in large part from linguists’ own intuitions or from reliance on logical argument and comparative analyses. we regard these proposals as largely heuristic in value: they present potentially testable claims about the possible relationship between evidentiality and epistemic modality. to more fully examine the relationship between evidentiality and epistemic modality and the effect of epistemic value on evidentiality it would be informative to look at how actual speakers interpret evidential and modal expressions and whether different interpretations are made by users of different languages, reflecting how evidentials and modals are marked. 1.3.1. complete disjointment those authors who claim a complete disjointment of evidentiality and epistemic modality (aikhenvald, 2004; de haan, 1999; lazard, 2001; oswalt, 1986) argue that the two structures convey different kinds and nature of information: epistemic modality refers to the judgment of truth or probability of the asserted event, while evidentiality refers simply to the source of information. in this view, all types of evidential sources are equally likely to be true and none of the sources are superior to the others in terms of the statement’s probability (see figure 1a). the most common example that source of knowledge and belief are separate is given by givón (1982) in referring to an account by a religious leader of the life of the buddha using the hearsay suffix. although buddhists consider the story as the truest of all stories, because they have not personally witnessed buddha’s life, they need to use the hearsay evidence marker while narrating it. this example illustrates how evidentiality may convey the source of knowledge independently of epistemic judgment about the narrative. 1.3.2. inclusion this view argues that one of the categories is a subtype of the other (dendale & tasmowski, 2001). epistemic modality is a subtype of evidentiality. in this view all epistemic modals are evidentials, but not all evidentials are epistemic modals (see chafe, 1986; matlock, 1989). the main problem with this view is the difficulty of demonstrating that all epistemic modals are evidentials (see figure 1c). logically, finding only one epistemic modal without an evidential meaning is enough to falsify this view. evidentiality is a subtype of epistemic modality. according to this view, evidentiality is considered a subtype of epistemic modality (bybee, 1985; mithun, 1999; palmer, 1986, willett, 1988); evidence sources imply some degree of certainty about the described situation. thus, all evidentials are epistemic modals, but not all epistemic modals are evidentials (plungian, 2001) (see figure 1c). a criticism of this view is that finding only one evidential without an epistemic meaning is enough to falsify the view. 1.3.3. overlap according to the overlap view, some markers may convey pure evidentials or modals, but others may be ambiguous, signalling both (de lancey, 2001; faller, 2002; van der auwera & plungian, 1998) (see figure 1 d). for example, pure modals, such as english ‘may’, do not convey any evidential meaning. when a person says, “jo may be the thief” the speaker is just talking about tosun and vaid 134 the possibility of the propositions without indicating the source of information. similarly, pure evidentials do not convey any epistemic value; an example is in givón’s (1982) example of the life of the buddha. ambiguous markers, such as inferential markers, convey both epistemic value and evidential value. for example, the english epistemic necessity marker, must, also expresses inferential evidentiality (faller, 2002; van der auwera & plungian, 1998). similarly, in turkish, the third-hand evidential marker, -mismis, is not just a source marker but conveys an epistemic value as well. because the overlap view contends that there are pure epistemic modals and pure evidentials besides some ambiguous markers, it addresses the criticisms of the inclusion view. 1.3.4. identity the final position, as exemplified by matthewson (2010), is that all evidentials are epistemic modals and all modals are evidentials (see figure 1e). logically, finding only one pure epistemic modal without an evidential meaning or pure evidential use without epistemic meaning is enough to falsify this view. figure 1. this figure demonstrates the theoretical views on the relationship between evidentiality and modality. plot a depicts complete disjoinment view, plot b and c depict inclusion views (respectively), plot d depicts overlap view and plot e depicts identity view. evidentiality and modality 135 1.4. previous empirical investigations of evidentiality and modality findings from a number of studies6 demonstrate that in various languages7 with grammatical evidential markers, evidentiality is indeed related to epistemic modality: firsthand source markers make the information perceived as more likely to happen, while nonfirsthand source markers are perceived to be less likely and less reliable. for example, aksu-koc and alici (2000) presented threeto six-year-old turkish speaking children with a dialog between a teddy bear and the experimenter. the key sentence was either in firsthand form (banyoya girdi. ‘she entered the bath’) or nonfirsthand form (banyoya girmiştir. ‘she must have entered the bath’). after children heard the dialogue, they were asked how certain the teddy bear was. sentences presented in firsthand form were perceived as being more certain than those presented in nonfirsthand source marker form. using the hidden objects task, öztürk and papafragou (2005) presented fiveto sevenyear-old turkish-speaking children with two puppets who are talking about what is in a closed box. one says there is an airplane in the box; the other says there is a helicopter in the box. one of the puppets uses the firsthand source marker while the other puppet uses the nonfirsthand source marker. then the children are asked what they think is in the box. the authors found that, in all age groups, children say the object was what the puppet who uses the firsthand source marker said was in it. further, as the age of the children increased, their reliance on the information presented with the firsthand marker also increased. likewise, tibetan-speaking children and adults demonstrate a similar effect in a hidden objects task; after the age of five, they rely on the information conveyed by the firsthand source marker more than on that conveyed by the nonfirsthand source marker (de villiers & garfield, 2009). on the other hand, papafragou et al. (2007), testing korean-speaking children, did not find a difference in reliability judgments between firsthand and nonfirsthand sources. this inconsistent finding disappears with adult participants, who prefer to rely on firsthand markers over nonfirsthand markers. considering that previous studies point to the age of comprehension of evidentiality at six years, the inconsistent finding of korean children might be explained by the fact that the oldest child group was four years old. similar findings have been reported in child speakers of bulgarian (fitneva 2008; 2009) and in japanese-speaking children and adults (matsui, yamamoto & mccagg, 2006; matsui & miura, 2009). in fitneva’s studies children were asked to listen to a story about four people, in which person a asks person b and c about person d. persons b and c answer her question in two different ways using different source markers. then the children are asked what they think about where and what person d is doing. fitneva found that children rely more on the information conveyed by the firsthand source marker over that conveyed by the nonfirsthand source marker. moore, bryant and furrow (1990) examined english-speaking children’s comprehension and attention to mental verbs using the hidden object task. the hidden object’s location is described in different ways: i know it’s in the red box vs. i think it’s in the blue box. in a similar task, moore, pure and furrow (1990) manipulated modal verbs such as must and might and found that by the age of five, children rely on must statements, more than might statements. based on a corpus study, biber and finegan (1989) conclude that evidentiality is a type of stance, indicating certainty and doubt. some adjectives (e.g. obvious, true), verbs (e.g. conclude), adverbs (e.g. assuredly), emphatics (e.g. for sure) and predictive modals (e.g. will) refer to certainty, while some adjectives (e.g. alleged), verbs (e.g. assume), adverbs (e.g. 6 although the present study investigated adult speakers, we gave examples from developmental studies because most of the previous empirical investigations were conducted on children speakers. 7 the languages studied with evidential marking (turkish, bulgarian, japanese, tibetan and korean) do not share the exact same non-first hand distinctions in their grammar, but they all make a distinction between first-hand and non first-hand sources of evidence. tosun and vaid 136 supposedly), hedges (e.g. maybe), possibility modals (e.g. might), and necessity modals (e.g. should) refer to doubt. considering modal auxiliaries as a form of evidentiality in english, francis and wales (1994) asked adult participants to listen to sentences with various modal auxiliaries (e.g. must, can) and to rate the sentences in terms of certainty and obligation. they found that the modal auxiliary must was rated the most certain modal and received the highest obligation ratings. in certainty ratings, can, would, should follow must, respectively. in obligation ratings, the order differs a bit from certainty. should follows must in obligation ratings, whereas would and can receive almost equal rating points. the modal might receives the lowest certainty and obligation ratings. tosun and vaid (2012) investigated the influence of english adverbs referring to evidential sources on reasoning and decision-making processes. participants were asked to read a short biography about a hypothetical politician. in this biography the key point is the part that mentions that he took hush money. all of the participants received the same biography with one difference being the key point statement’s adverb: edward apparently took hush money. the adverb was changed from direct (no adverb), to apparently, allegedly, presumably, reportedly and one modal, must have. after reading the biography participants were asked the question, how much do you think edward took as hush money? participants who received the biography in the direct version gave the highest estimate whereas those who received reportedly, presumably and must have versions estimated a significantly smaller amount of money. taken together, a range of studies shows that for language users evidential markers carry epistemic value as well. evidential source markers influence people’s reliability and certainty judgments. firsthand sources are considered more reliable and trustworthy than nonfirsthand sources. however, whereas previous studies mainly tested the difference between firsthand and nonfirsthand source of evidence, differences between the various nonfirsthand sources of information (e.g., assumption, inference) have not been sufficiently examined. moreover, previous investigations demonstrate the relation only from the perspective of epistemic value in evidential marking. in order to fully understand the relation between source and stance, we also need to look into epistemic modals and ask whether they are interpreted as carrying evidential (source) information as well. for example, is the situation described by, “john must have washed his hands” interpreted as expressing simply that john probably washed his hands or as expressing the speaker’s inference that john apparently washed his hands? by simultaneously examining source interpretations for modals and certainty interpretations for evidentials, one can directly address the hypothesized claims about the relation between evidentiality and epistemic modality. finally, there is a need to clearly understand how users of different languages interpret evidential and epistemic modal expressions. are speakers of languages in which the coding of source of evidence is required by the grammar more likely to invoke source for modal expressions as well, as compared to speakers of languages in which source information is not obligatorily coded? alternatively, are speakers of an evidentiality-marked language more likely to accord greater confidence to sources deemed higher in the evidential hierarchy than are those of a language that does not grammaticalize evidentiality? we turn now to the present research, which was designed to address these issues. 1.5. the present research our study examined the nature of the relationship between evidentiality and modality in speakers of turkish vs. english by means of a sentence interpretation task in which two types of judgments were elicited: judgments about the source of evidence and judgments about confidence in whether the asserted event had occurred. each sentence described an event and was framed in eight possible ways, using four different evidential forms and four different epistemic modal forms (e.g., jack reportedly passed the math test / jack must have passed the math test). source and evidentiality and modality 137 confidence judgments were elicited for each sentence form. an option was also provided to opt out of a response, if the participant considered that the sentence did not provide enough information to allow a judgment of source or confidence. the present study is the first that directly investigates both the source and degree of confidence interpretations of both evidential and modal sentences and does so in speakers of two different languages, turkish and english (see appendix a). 1.5.1. rationale and hypotheses it was hypothesized that turkish and english speakers alike would report that there was enough information provided in the sentences to make both types of judgments (source and certainty) for both evidential and modal sentences. if this is the case, it would show that language speakers can derive source information from modal sentences and epistemic value from evidential sentences. this in turn would provide support for the position that there is a close relationship between evidentiality and modality. it is further hypothesized that turkish speakers’ source and confidence judgments of evidential sentences will show more variation across the different evidential expressions. that is, because evidentiality marking is required by the grammar, certain markers are assigned for certain sources of knowledge. turkish speakers will interpret the sources of each evidential expression depending on how they are defined in the grammar. similarly, if each source conveys different degrees of certainty, then turkish speakers will also interpret the degree of certainty of evidence expressions at various levels, e.g., they will show higher levels of certainty for certain evidential expressions than for others. since modal expressions in turkish have not been studied empirically or claimed to have clearly specified degrees of certainty, there was no basis to predict how turkish speakers would interpret modal expressions. on the other hand, modal expressions in english are defined more meticulously and it was already demonstrated that there is a variation in conveying degree of certainty across modal expressions (francis & wales, 1994). we therefore expect that english speakers would interpret the source of modal expressions in accordance with their perceived degree of certainty. further, because evidentiality is not signaled in the grammar of english, the meaning or the source of evidential expressions will be interpreted according to how speakers use them. the english expressions used in the study (reportedly, apparently, presumably and supposedly) were selected based on linguistic scholars’ predictions (e.g., aikhenvald, 2004; chafe, 1986; gisborne & holmes, 2007; izvorski, 1997; mushin, 2001). although there is a consensus on the source of evidence of some adverbs, such as reportedly as hearsay and presumably as assumption, there are some adverbs over which linguistic scholars have not reached an agreement yet, such as apparently. similarly, supposedly is considered as conjecture in some contexts and as hearsay within a different context (chafe, 1986). the present study will shed light on this issue by bringing empirical observations to bear on whether the linguists’ intuitions are borne out by ordinary language users’ interpretations. 2. method an experimental design was conducted to test the hypothesis. 2.1. participants a total of 50 turkish-speaking participants (45 females) and 60 english-speaking participants (40 females) were recruited from a major city in turkey and from a large research university in the southwestern region of the u.s. turkish speakers’ ages ranged from 19 to 27, with a mean age of 22.6 years. as determined on the basis of a language background questionnaire administered to screen out individuals who may have known other languages that could have influenced their performance on the languages of interest in the study, participants from turkey all reported using turkish as their first and primary language. although most of these participants had also studied tosun and vaid 138 english and/or arabic, they considered themselves as essentially monolingual because their knowledge of these other languages was at a beginner level. english speakers’ ages ranged from 17 to 20 with a mean age of 18.6 years. although some had studied another language in high school they also self-identified as essentially monolingual speakers. 2.2. materials and design per language, stimuli consisted of a total of 80 declarative sentences presented in active voice, third person singular form and containing a verb in the past tense. the 80 sentences were generated from a set of 10 sentences, each presented in 8 versions to represent four evidential forms and four different modal forms. the evidential categories were hearsay (reportedly), inference (apparently), assumption (presumably), and conjecture (supposedly). because the turkish evidential marker for hearsay and inference is the same (-miş), two lexical items (duyduğuma göre ‘reportedly’ and görünüşe göre ‘apparently’) were added to the turkish sentences to make the intended meaning clear. the modal categories were necessity (must), weak necessity (should), probability (could), and possibility (might). note that the turkish modals could and might have the same marker. sample stimuli per language are given below, with the relevant evidential or modal marker in italics, for purpose of illustration. sample evidential stimuli hearsay: reportedly the girl worked out for an hour yesterday. duyduğuma göre kız dün bir saat spor yapmış. inference: apparently the girl worked out for an hour yesterday. görünüşe göre kız dün bir saat spor yapmış. assumption: presumably the girl worked out for an hour yesterday. kız dün bir saat spor yapmıştır. conjecture: supposedly the girl worked out for an hour yesterday. kız dün bir saat spor yapmışmış. sample modal stimuli necessity: the girl must have worked out for an hour yesterday. kız dün bir saat spor yapmış olmalı. weak necessity: the girl should have worked out for an hour yesterday. kız dün bir saat spor yapmalıydı. probability: the girl could have worked out for an hour yesterday. kız dün bir saat spor yapmış olabilir. possibility: the girl might have worked out for an hour yesterday. kız dün bir saat spor yapmış olabilir8. the form variable (evidential vs. modal) was manipulated within subjects. the 80 sentences were arranged in such a way that 10 sentence blocks were presented, representing each of the 8 types (four types per form condition). participants saw each sentence only once in one of the eight conditions. the condition of sentences was counterbalanced between participants. within each block, sentences were presented in a fixed random order. dependent variables were “not enough information judgments”, source type judgments and epistemic value (confidence) judgments. source type judgments. for source judgments, participants were instructed to decide on the source of information of each sentence based on 4 designated options: hearsay, inference, assumption, and conjecture. they were informed that “hearsay” would indicate that the information is based on hearing it from someone else; “inference” would indicate that the person 8 turkish allows a way of distinguishing between possibility and probability using lexical or morphological cues. however, in this study we did not apply the lexical cues in modals. evidentiality and modality 139 who reported the information saw some signs related to the information but did not see what happened; “assumption” would indicate that the person who reported the information did not see what happened but reasoned what must have happened based on some knowledge; and “conjecture” would indicate that the person who stated that information did not see the event, but also the source of information was unspecified. importantly, participants were also able to choose the option of “not enough information for a decision” if they did not feel the sentence provided enough information for them to identify the type of source. confidence judgments. for confidence judgments, participants were instructed to decide how confident they were that the event referred to in the sentence actually happened, based on how the sentence was structured. the response options were: extremely, quite, somewhat, or not at all confident. they also had an option of choosing “not enough information for a decision.”9 participants were reminded that there were no right or wrong answers in the test and that we were simply interested in seeing how listeners interpret sentences conveying information about an event. the two tasks (source judgments and confidence judgments) were conducted 1 or 2 days apart and the order of the task was counterbalanced. the experiment was paper-pencil based. participants were tested only in their primary language (turkish or english). 2.3. procedure participants were tested in groups. per session, they received a booklet containing 80 different sentences with instructions that directed them to make either source or confidence judgments for that sentence (see appendix a). one or two days after they completed the first task they received the other task. sentences across tasks were presented in the same order. a language background questionnaire was administered last. 2.4. data coding and analysis the data from the two tasks (source judgments and confidence judgments) were analyzed separately. judgments of ‘not enough information’. an initial analysis was conducted of the relative proportion of the ‘not enough information’ (nei) response option selected for each of the form types. for example, the proportion of nei responses to the hearsay sentence form was computed for the 10 sentences with the adverb ‘reportedly’. the same computation was done for all other sentence forms. further, the same nei computation was followed for the confidence judgment analysis. after that, an overall mean nei proportion was computed for evidential sentences, which included the four evidential sentence forms and the four modal types. two separate 2x2 analyses of variance were conducted, one for source judgments and one for confidence judgments, each examining the influence of sentence type (evidential vs. modal) and language (turkish vs. english). judgments of source of evidence. a second set of analyses was then performed on the source judgments per form type by comparing the relative proportion of each source option selection out of the total source judgments made for each evidential and modal sentence form, excluding the nei option. for example, for hearsay sentences (with the adverb ‘reportedly’ for english) the relative proportion of hearsay judgments was examined based on the total number of source judgments. the same computation was done for each source response for each sentence form. this detailed coding made it easier to see the pattern of choices for each evidential source and modal. judgments of confidence in event occurrence. a third set of analyses examined confidence judgments per sentence type. these were computed by weighting the judgments. 9 while we are asking them to judge their own confidence, they are presumably basing their judgments on the degree of confidence/certainty as conveyed in the sentence. tosun and vaid 140 ‘extremely confident’ responses were given a weighting of 3, ‘quite confident’ were weighted as 2 and ‘somewhat confident’ as 1; with ‘not at all confident’ responses weighted as 0. the total points were then standardized as percentages so that the highest confidence level would be 100% and the lowest confidence level would be 0%. to summarize, three analyses were conducted. the first compared turkish and english speaking participants’ relative use of the ‘not enough information’ response option for source and confidence judgments. the second compared turkish and english speakers’ responses on source judgments. finally, the third analysis examined participants’ belief/confidence judgments that the event described in the sentence had actually occurred. for all analyses, significance was set at p < .05 and partial eta ηp 2 is reported as the measure of effect size. 3. results the results are presented as classified in the previous section. 3.1. analysis of ‘not enough information’ response option a 2 (sentence form type: evidential vs. modal sentence) x 2 (language: turkish vs. english) repeated measures anova was conducted for the “not enough information” (nei) responses for the source judgments and for the confidence judgments (in separate anovas). for the purpose of this analysis, responses were averaged across the four subtypes of each sentence form. see table 1 for a summary of the mean percent response. source judgments. a significant effect of sentence type was observed, f (1, 94) = 7.29, p < .01, ηp 2 = .07, which indicated that the “not enough information” response was given more frequently to modal than evidential sentences. language was also significant, f (1, 94) = 34.3, p < .001, ηp 2 = .27, indicating that turkish speakers gave significantly more nei responses than did english speakers. a sentence form by language interaction, however, did not emerge, f (1, 94) = .002, p = .97. belief/confidence judgments. a main effect of language, f (1, 93) = 18.68, p < .001, ηp 2 = .17 showed that turkish speakers showed a higher number of ‘not enough information’ responses than english speakers. there was no main effect of sentence form, f (1, 93) = .72, p = .4. further, the sentence form by language interaction was not significant, f (1, 93) = 1.26, p = .26. judgment sentence type turkish (n=50) english (n=60) evidential 8.32 (9.61) .46 (1.34) source modal 11.18 (12.08) 3.34 (8.23) evidential 8.83 (12.99) 1.21 (4.73) confidence modal 8.55 (9.88) 3.17 (7.13) note. standard deviations are indicated in parentheses. table 1. turkish and english speaking participants’ mean percent nei responses on source and confidence judgments to evidential and modal sentence forms. taken together, the results of these analyses show that, on the whole, participants judged there to be sufficient information on which to base both judgments of source and judgments of confidence that the asserted event had occurred; the “not enough information” option was selected at most 10% of the time and in some cases hardly at all. further, participants were more likely to state there was not enough information when asked to make source judgments for sentences with modal structures than for those with evidential structures. finally, for both sentence types and across both tasks, turkish speakers were more likely to state there was not enough basis on which to respond than english speakers were. evidentiality and modality 141 3.2. analyses of source judgment responses excluding nei response option in this analysis, a 4 (judgment type: hearsay v. inference v. assumption v. conjecture) x 2 (language: turkish v. english)10 repeated measures anova was conducted for the evidential sentences, and for the four subtypes of modal sentences, in separate anovas per sentence type condition. the results are displayed in figure 2 for evidential sentence types and in figure 3 for modal sentence types. 3.2.1. evidential condition hearsay sentence type the analysis of source judgments for sentences containing the hearsay type (e.g. reportedly jack passed the math test) revealed a significant source judgment type main effect, f (3, 324) = 64.16, p < .001, ηp 2 = .37. when asked to judge the source of sentences with the hearsay sentence form, participants chose hearsay significantly more frequently than inference (t (109) = 10.09, p < .001), assumption (t (109) = 11.06, p < .001) or conjecture (t (109) = 6.11, p < .001). conjecture responses were in turn chosen significantly more frequently than inference (t (109) = 2.65, p = .01) or assumption (t (109) = 3.34, p < .001). inference and assumption responses did not differ from each other. there was no main effect of language, but a source judgment response type by language interaction was significant, f (3, 324) = 25.61, p < .001, ηp 2 = .19. turkish speakers chose hearsay more frequently than inference (t (49) = 3.47, p < .001) and assumption (t (49) = 3.9, p < .001), but were as likely to choose conjecture as hearsay. on the other hand, for english speakers, hearsay was the most frequent response and it was chosen significantly more frequently than inference (t (59) = 12.94, p < .001), assumption (t (59) = 14.41, p < .001) or conjecture (t (59) = 12.85, p < .001). the other responses did not differ from one another. inference sentence type. the analysis of source judgments for the inference sentence form (e.g. apparently jack passed the math test) demonstrated a significant judgment response main effect, f (3, 324) = 35.78, p < .001, ηp 2 = .25. overall, inference responses were chosen significantly more frequently than assumption (t (109) = 6.6, p < .001) and conjecture (t (109) = 8.13, p < .001), but equally frequently as hearsay responses (t (109) = .55, p = .58). moreover, hearsay responses were chosen more frequently than assumption (t (109) = 5.49, p < .001) and conjecture (t (109) = 7.05, p < .001). finally, assumption responses were chosen more frequently than conjecture (t (109) = 2.09, p < .05). further, a judgment response type by language interaction emerged as well, f (3, 324) = 27.7, p < .001, ηp 2 = .20. turkish speakers’ judgments of inference sentences as inference were more frequent than hearsay (t (49) = 5.62, p < .001), assumption (t (49) = 6.98, p < .001) or conjecture (t (49) = 8.46, p < .001). there was no difference between hearsay and assumption responses (t (49) = 1.02, p = .31); however, hearsay responses were more frequent than conjecture responses (t (49) = 2.32, p = .025). on the other hand, for english speakers, hearsay responses were the most frequent response to the inference sentence form and they were significantly higher than inference (t (59) = 3.85, p < .001), assumption (t (59) = 6.31, p < .001) or conjecture (t (59) = 7.57, p < .001). the second common answer was inference and it was significantly more frequent than assumption (t (59) = 2.68, p < .01) and conjecture (t (59) = 3.84, p < .001). there was no difference between assumption and conjecture (t (59) = 1.28, p = .2) responses. 10 a language main effect was not expected in any sentence type analyses because for both turkish and english speakers total percentages of the response options added up to100%. tosun and vaid 142 figure 2. mean percent source judgment responses to evidential sentence forms of turkish and english speakers. assumption sentence type. the results of the assumption sentence forms (e.g. presumably jack passed the math test) demonstrated a significant judgment response type main effect, f (3, 321) = 19.49, p < .001, ηp 2 = .15. the most frequent source judgment for the assumption sentence form was inference, and that response was significantly more common than hearsay (t (108) = 3.82, p < .001), and conjecture (t (108) = 8.46, p < .001). the second common response was assumption and it was significantly more frequent than hearsay (t (108) = 1.98, p = .051) and conjecture (t (108) = 5.83, p < .001). however inference and assumption responses did not differ (t (108) = 1.72, p = .09). a judgment response type by language interaction did not emerge, f (3, 321) = 1.35, p = .26. turkish and english speakers exhibited similar source judgment patterns to assumption sentences. conjecture sentence type. the conjecture sentence form (e.g. supposedly jack passed the math test) analysis revealed a significant judgment response main effect, f (3, 324) = 16.75, p < .001, ηp 2 = .13. post hoc analyses revealed that the most frequent two responses were hearsay and conjecture. hearsay responses were significantly more frequent than inference (t (109) = 6.01, p < .001), assumption (t (109) = 4.93, p < .001) and conjecture (t (109) = 2.16, p = .033). conjecture responses were also more frequent than inference (t (109) = 3.44, p < .001) and assumption (t (109) = 2.39, p = .018). further, there was no difference between inference and assumption, (t (109) = 1.34, p = .18). further, a judgment response type by language interaction was significant, f (3, 324) = 21.47, p < .001, ηp 2 = .17, indicating that the source judgment pattern of turkish and english speakers was significantly different. turkish speakers judged conjecture sentences as conjecture more frequently than hearsay (t (49) = 2.19, p = .022), inference (t (49) = 6.34, p < .001) or evidentiality and modality 143 assumption (t (49) = 4.62, p < .001). on the other hand, english speakers judged conjecture sentences as hearsay more frequently than inference (t (59) = 5.06, p < .001), assumption (t (59) = 4.82, p < .001) or conjecture (t (59) = 7.38, p < .001). 3.2.2. modal condition ‘must’ sentence form. the analysis of sentences containing the modal ‘must’ (e.g. jack must have passed the math test) demonstrated a significant source judgment response type main effect, f (3, 321) = 56.89, p < .001, ηp 2 = .35. the most common judgment responses to ‘must’ sentences were inference and assumption. both responses were significantly more frequent than hearsay [inference v. hearsay: t (108) = 11.18, p < .001; assumption v. hearsay: (t (108) = 10.29, p < .001] and conjecture [inference v. conjecture: t (108) = 8.56, p < .001; assumption v. conjecture: t (108) = 8.66, p < .001]. the difference between assumption and inference was not significant, t (108) = .45, p = .65. further, a judgment response by language interaction was also significant, f (3, 321) = 3.85, p < .01, ηp 2 = .035. turkish and english speakers were equally likely to judge ‘must’ sentences as inference, t (107) = 1.55, p = .12. however, english speakers selected assumption more than turkish speakers, (t (107) = 2.14, p < .05). turkish speakers, in turn, made more hearsay judgments than english speakers, t (107) = 2.99, p < .01. there was no difference between turkish and english speakers’ frequency of selection of conjecture responses, t (107) = 1.58, p = .12. ‘should’ sentence form. the results of source judgments for ‘should’ sentences (e.g. jack should have passed the math test) showed that there was a significant judgment response main effect, f (3, 297) = 12.4, p < .001, ηp 2 = .11. inference was the most common response and significantly more frequent than hearsay (t (100) = 7.12, p < .001) and conjecture (t (100) = 2.05, p < .05). although assumption was significantly more frequent than hearsay (t (100) = 5.03, p < .001), it was equally frequent as inference (t (100) = 1.32, p = .19) and conjecture (t (100) = .8, p = .42). further, a judgment response by language interaction emerged, f (3, 297) = 3.86, p < .01, ηp 2 = .04. the pattern of english and turkish speakers’ source judgments was different. for turkish speakers the most frequent response was inference and it was a significantly more common response than hearsay (t (40) = 4.78, p < .001), assumption (t (40) = 2.37, p < .05) and conjecture (t (40) = 3.62, p < .001). assumption was the second most frequent answer and it was more common than hearsay (t (40) = 2.41, p = .02). on the other hand, english speakers’ judgments were equally distributed for inference, assumption, and conjecture [inference v. assumption: t (59) = .19, p = .85; inference v. conjecture: t (59) = -.18, p = .85; assumption v. conjecture: t (59) = -.01, p = .99]. hearsay judgment was the least frequent response [hearsay v. inference: t (59) = -5.36, p < .001; hearsay v. assumption: t (59) = -4.5, p < .001; hearsay v. conjecture: t (59) = -4.35, p < .001]. ‘could’ sentence form. the analysis of source judgments for ‘could’ sentences (e.g. jack could have passed the math test) revealed a significant judgment response main effect, f (3, 303) = 24.7, p < .001, ηp 2 = .2. the most frequent source responses to ‘could’ sentences were inference and assumption. these were significantly more common than hearsay and conjecture [inference v. hearsay: t (102) = 6.75, p < .001; inference v. conjecture: t (102) = 3.68, p < .001; assumption v. hearsay: t (102) = 7.68, p < .001; assumption v. conjecture: t (102) = 4.89, p < .001]. inference and assumption response frequency did not differ from each other, t (102) = 1.19, p = .24. finally conjecture responses were more frequent than hearsay, t (102) = 2.53, p < .05. tosun and vaid 144 figure 3. mean percent source judgment responses to modal sentence forms of turkish and english speakers. in addition, a judgment response by language interaction was significant, f (3, 303) = 3.81, p < .01, ηp 2 = .04. the response pattern of english speakers was different from that of turkish speakers. turkish speakers judged could sentences as inference and assumption more frequently than hearsay and conjecture [inference v. hearsay: t (42) = 4.59, p < .001; inference v. conjecture: t (42) = 5.07, p < .001; assumption v. hearsay: t (42) = 4.49, p < .001; assumption v. conjecture: t (42) = 4.85, p < .001]. no difference was found between inference and assumption, t (42) = .2, p = .84; nor between hearsay and conjecture, t (42) = .46, p = .64. on the other hand, english speakers judged could sentences as assumption more frequently [assumption v. hearsay: t (59) = 6.23, p < .001; assumption v. conjecture: t (59) = 2.54, p = .014; assumption v. inference: t (59) = 1.92, p = .059], followed by inference and conjecture. the inference response was more frequent than hearsay, t (59) = 4.97, p < .001; however, it was as frequent as conjecture, t (59) = .86, p = .39. finally conjecture responses were more common than hearsay responses, t (59) = 3.36, p < .001. ‘might’ sentence form. the analysis of source judgments for ‘might’ sentences (e.g. jack might have passed the math test) demonstrated that there was a significant judgment response main effect, f (3, 324) = 31.19, p < .001, ηp 2 = .22. similar to the ‘could’ sentence form results, evidentiality and modality 145 participants judged ‘might’ sentences more often as inference and assumption than hearsay and conjecture [inference v. hearsay: t (109) = 8.79, p < .001; inference v. conjecture: t (109) = 3.47, p < .001; assumption v. hearsay: t (109) = 9.42, p < .001; and assumption v. conjecture: t (109) = 4.46, p < .001]. there was no difference between inference and assumption, t (109) = .87, p = .39. finally, conjecture was chosen more often than hearsay, t (109) = 4.15, p < .001. further, a judgment response by language interaction emerged, f (3, 324) = 8.61, p < .001, ηp 2 = .07. the source judgment pattern of ‘might’ sentences differed among turkish and english speakers, similar to the other modal sentence forms. english speakers judged ‘might’ sentences equally frequently as inference, assumption and conjecture [inference v. assumption: t (59) = -1.41, p = .16; inference v. conjecture: t (59) = -.15, p = .88; assumption v. conjecture: t (59) = 1.11, p = .27]. further, all three sources were significantly more frequently selected than hearsay [inference v. hearsay: t (59) = 7.3, p < .001; assumption v. hearsay: t (59) = 9.52, p < .001; conjecture v. hearsay: t (59) = 6.35, p < .001]. on the other hand, turkish speakers selected inference and assumption sources as the most common, and these sources were more frequent than hearsay and conjecture [inference v. hearsay: t (49) = 5.56, p < .001; inference v. conjecture: t (49) = 6.54, p < .001; assumption v. hearsay: t (49) = 4.88, p < .001; and assumption v. conjecture: t (49) = 6.58, p < .001]. further, inference and assumption (t (49) = .05, p = .96) did not differ from each; neither did hearsay and conjecture (t (49) = .89, p = .37). the overall results of the source judgment analysis are summarized in table 2. 3.3. analysis of belief/confidence judgments (excluding the not enough information response option) participants’ total confidence judgment scores were analyzed separately for the evidential and the modal sentences. each analysis involved a 4 (sentence type) x 2 (language) anova, with repeated measures on sentence type. 3.3.1. evidential sentences the evidential sentence type analysis revealed a significant sentence type main effect, f (3, 324) = 35.02, p < .001, ηp 2 = .25. hearsay and inference sentences received the highest confidence scores. the scores of the two source types did not differ from each other (t (109) = .63, p = .53), while they were significantly greater than other sentence forms [hearsay v. assumption: t (109) = 2.14, p = .03; hearsay v. conjecture: t (109) = 8.03, p < .001; inference v. assumption: t (109) = 3.42, p < .001; inference v. conjecture: t (109) = 7.57, p < .001]. conjecture was judged the lowest in confidence level and significantly less than assumption (t (109) = 5.47, p < .001). the language effect was significant as well, f (1, 108) = 13.05, p < .001, ηp 2 = .11. overall english speakers’ confidence scores were significantly greater than turkish speakers. sentence type turkish source judgments english source judgments hearsay h = c >i = a h >i = a = c inference i > h =a > c h > i > a = c assumption i ≥ a >h > c i ≥ a >h > c conjecture c > h > i = a h > i = a = c necessity i > a > h = c a = i > c > h weak necessity i > a > h = c c = i = a > h probability i = a > h = c a ≥ i = c > h possibility i =a > h = c a = i = c > h note. h = hearsay, i = inference, a = assumption, and c = conjecture table 2. relative mention of each source type by language and sentence type tosun and vaid 146 further, a sentence type by language interaction emerged, f (3, 324) = 23.56, p < .001, ηp 2 = .18. confidence judgment patterns of evidential sentence types differ for turkish speakers and english speakers (see figure 4). turkish speakers demonstrated more variation in their confidence judgments of evidential sources. for turkish speakers, inference sentences elicited the most confident response compared to all other sources (hearsay: t (49) = 4.89, p < .001; assumption: t (49) = 2.48, p = .02; conjecture: t (49) = 7.56, p < .001). the second most confident response was to assumption, which was significantly greater than hearsay (t (49) = 2.1, p < .05) and conjecture (t (49) = 6.14, p < .001) sources. the third most confident source was hearsay, which was judged more confident than conjecture, which had the lowest confidence score (t (49) = 4.35, p < .001). for english speakers by contrast, hearsay was the most confident source compared to all the other sources (inference: t (59) = 4.18, p < .001; assumption: t (59) = 6.89, p < .001; and conjecture: t (59) = 7.25, p < .001). inference was the second most confident source and was significantly greater than assumption (t (59) = 2.42, p = .02) and conjecture (t (59) = 4.19, p < .001). however, there was no difference between the sources of assumption and conjecture (t (59) = 1.44, p = .17). with respect to group differences, english speakers’ confidence scores for hearsay and conjecture sources were significantly higher than those for turkish speakers [hearsay: t (109) = 5.76, p < .001; and conjecture: t (108) = 5.81, p < .001] while the confidence scores of inference and assumption did not differ between turkish and english speakers [inference: t (108) = 1.6, p = .11; assumption: t (108) = .46, p = .64]. 3.3.2. modal sentences the analysis of confidence judgments for modal sentences (see figure 5) demonstrated a significant sentence type main effect, f (3, 282) = 11.9, p < .001, ηp 2 = .11. the necessity modal was judged as the most confident modal and its confidence level was significantly higher than that of the other modals [weak necessity: t (102) = 3.01, p = .003; probability: t (102) = 4.84, p < .001; possibility, t (109) = 6.93, p < .001]. weak necessity and probability sentence confidence levels did not differ from each other (t (95) = 1.76, p = .08), but were significantly higher than possibility sentences [weak necessity v. possibility: t (102) = 2.89, p = .005; probability v. possibility: t (102) = 2.28, p = .03]. the language main effect was not significant, f (1, 94) = 2.92, p = .09. a sentence type by language interaction was significant, f (3, 282) = 3.94, p < .01, ηp 2 = .04. turkish speakers and english speakers demonstrated different confidence patterns for modal sentences (see figure 5). turkish speakers judged all modal sentences similarly in terms of the confidence level that they convey. the only difference was between necessity and possibility sentences, (t (49) = 3.07, p = .004), where necessity sentences were judged more confident than possibility sentences. the other modal sentence forms did not differ from one another [necessity v. weak necessity: t (42) = .15, p = .88; necessity v. probability: t (42) = 1.61, p = .12; weak necessity v. probability: t (35) = 1.21, p = .23; weak necessity v. possibility: t (42) = 1.69, p = .1; probability v. possibility: t (42) = 1.08, p = .28]. evidentiality and modality 147 figure 4. mean percent confidence judgment of evidential sentences by sentence type and group. on the other hand, english speakers exhibited more variation and a strong confidence order among modal types. the modal necessity was judged as the most confident type and its confidence level was significantly higher than other modal forms [weak necessity: t (59) = 3.49, p < .001; probability: t (59) = 4.81, p < .001; possibility: t (59) = 6.7, p < .001]. weak necessity and probability sentences did not reveal a significant difference in their confidence level, t (59) = 1.27, p = .21; while their confidence scores were significantly greater than possibility sentences [weak necessity v. possibility: t (59) = 2.35, p = .02; and probability v. possibility: t (59) = 2.02, p < .05]. tosun and vaid 148 figure 5. mean percent confidence judgment of modal sentences. 4. discussion a primary goal of the research was to bring empirical data to bear on various hypothesized views about the relationship between evidentiality (the source of evidence for a narrated event) and epistemic modality (the confidence in whether the event had occurred). a secondary goal was to investigate whether the nature of evidentiality is expressed (i.e., in the grammar vs. in the lexicon) differentially influences speakers’ interpretation of different types of evidential and modal expressions. we will proceed by first discussing the findings for the ‘not enough information’ analysis followed by the source judgment results and finally the confidence judgment results. 4.1. ‘not enough information’ (nei) responses the nei option provides a baseline determination of whether participants would make source judgments at all for modal sentences and confidence judgments at all for evidential sentences. if participants consider that a modal sentence contains enough information to allow one to interpret the source of information, this tells us that modal expressions provide some source-related information. similarly, if participants find there was enough information from an evidential sentence to interpret the epistemic value of the proposition, that tells us that an evidential expression has some epistemic value. we found that for both source and confidence judgments, and for both turkish and english speakers, nei responses were chosen far less often than would evidentiality and modality 149 be expected based on chance (20%), suggesting that, on the whole, participants considered there to be enough information on which to make source and confidence judgments. source judgments. across both groups, nei was chosen less often for evidential sentence types than for modal sentence types. this result is entirely to be expected, given that evidential sentence types explicitly code for source. what is nevertheless noteworthy is that selection of the nei option even for modal sentences is fairly low (11.18% for turkish, 3.4% for english speakers). as such, the results demonstrate that the modal sentence forms did convey source information along with the other meanings. confidence judgments. here too, participants showed a very low level of selection of the nei option (less than 10% of responses were of this type). there was no difference in nei response rate between evidential and modal expressions: participants responded to the same degree that there was enough information to judge the epistemic value of the propositions from evidential as from modal expressions. thus, people appear to interpret evidential expressions as conveying epistemic value along with the other meanings at the same rate as they do for modal expressions. differences between turkish and english speakers. another theoretically interesting aspect of the findings from the nei analyses was the difference observed between turkish and english speakers. english speakers showed almost no hesitation in making both source and confidence judgments of evidential sentences, with only a small percent stating that they did not have enough information to make source or confidence judgments of modal or evidential sentences. on the other hand, turkish speakers were significantly more likely to find the information given was not enough to make the required judgments. although it is difficult to explain this somewhat unexpected finding, it does suggest a need for further research to corroborate this group difference of turkish speakers being apparently more wary about over-interpreting the meaning of a sentence and more inclined to take a cautious approach. while there is some other work suggesting cultural differences whereby individuals based in western societies demonstrate different procedures and processes while making judgments and decisions than those based in eastern societies (matthew & busemeyer, 2011). although no two eastern and western cultures behave in the same way, some common ground patterns have been found for western cultures (americans and europeans) different than eastern cultures (from far east to middle east), thus it remains to be seen if those differences are applicable to the present study. whatever the reason for the observed group difference, it is theoretically important that the rate of nei responses was below chance level (20%). this result demonstrates that people judge there to be enough information to determine both the source of information of modal expressions and the epistemic value of evidential expressions. this finding in itself would lead one to reject the complete disjointment view proposed by aikhenvald (2004), de haan (1999), lazard (2001) and oswalt (1986), where the two linguistic properties are considered completely independent from each other. overall, the results of the nei analyses demonstrate that modality and evidentiality are neither completely disjoint properties nor are they completely the same. they display as two sets that intersect to a great extent. a further question asked in the present study was how the intersecting/overlapping areas of evidentiality and modality sets are shaped by specific subtypes of each sentence form, and whether certain evidential and modal subtypes are interpreted and utilized interchangeably. to address this issue we turn to the additional analyses conducted of the source and confidence judgments (after excluding the nei responses). 4.2. interpretation of source judgments the source judgment task was designed first to examine speakers’ ability to classify the source type of sentences containing evidential structures (e.g., whether sentences containing the thirdhand suffix in turkish or the adverb “supposedly” in english would be classified as tosun and vaid 150 conjecture). examining the pattern of source judgments elicited for modal sentences allowed us to address the question of the nature of the relationship between evidentiality and modality. the source judgment task with modal expressions was designed to clarify whether speakers could find enough information to make source judgments, and if so, whether their judgments would correspond to theoretical claims proposed by linguists. source judgments of evidential expressions. the results of the source judgments task demonstrated that turkish and english speakers interpret evidential expressions differently. turkish speakers judged the source for hearsay sentences equally often as hearsay (37.9%) or as conjecture (36.7%). relatedly, conjecture sentences were judged as conjecture at the rate of 49%, and as hearsay 27% of the time. this clustering of hearsay and conjecture suggests that for turkish speakers conjecture is perceived as a kind of thirdhand information, an option that is provided in turkish grammar. hence, turkish speakers could easily interchange hearsay and conjecture. inference sentences were judged as inference 60% of the time, which demonstrated that turkish speakers were mostly consistent about this source of inference. however, for assumption sentences turkish speakers made more inference than assumption judgments (42% vs. 31%). this suggests that the assumption source marker in turkish is perceived as conveying inference (the source of information obtained from results) and assumption (the source of information obtained from reasoning) somewhat interchangeably. english speakers’ source judgments of evidential sentences demonstrated less variability than those of turkish speakers. english speakers uniformly judged evidential sentences containing “reportedly” as hearsay (76%). this finding is consistent with linguists’ claims (aikhenvald, 2004; chafe, 1986; mortensen, 2006). however, english speakers rated apparently (as inference) and supposedly (as conjecture) as hearsay about half of the time (53%). these results support chafe’s claims but not those of ginsborne and holmes (2007), izvorski (1997) and mortensen (2006). further, the three evidential adverbs (reportedly, apparently, supposedly) were interpreted as the same in terms of other less frequent source judgments. all follow the order hearsay, inference, assumption and conjecture. thus, we can conclude that these adverbs were interpreted the same. in contrast, presumably (used as assumption) was judged as a mix of inference and assumption. this finding is somewhat consistent with chafe’s presumption on presumably. further, english speakers’ interpretation of presumably is very similar to turkish speakers’ interpretation of the assumption marker –dır (aksu-koc & alici, 2000). source judgments of modal expressions. the results of source judgments of modal sentences also demonstrated a significant difference between turkish and english speakers. turkish speakers interpreted all of the modal forms as the same. this was expected for the could and might sentence forms because both have the same morpho-syntactic marker. however, the other two modal forms (must and should) have different markers than could/might. yet, turkish speakers show no distinction between the modal types, interpreting them all as a combination of inference and assumption, consistent with what was claimed by kerimoglu (2010) and kornfilt (2001). english speakers, on the other hand, demonstrated more variation in source judgments of modal sentences. the modal must was interpreted as both inference (37.83%) and assumption (44.01%). this finding partly supports the claim of chafe (1986), faller (2002), von fintel and gillies (2010) and van der auwera and plungian (1998). contrary to chafe, however, should sentences were not interpreted distinctively as assumption; rather they were equally judged as conjecture (31.14%), inference (29.91%) and assumption (31.18%). could and might sentences were interpreted similarly as assumption, followed by inference and conjecture. english speakers’ source interpretation of modal sentences, overall, shows a combination of inference, assumption and conjecture at different rates. however, none of the modal interpretations directly correspond to an evidential form’s interpretation, which serves as evidence against the identity view. these findings demonstrate a consistent difference between users of grammaticalized versus lexicalized evidentials. for turkish speakers, identifying and classifying source of evidentiality and modality 151 information is second nature, given that their language requires them to make clear distinctions, at least between firsthand and nonfirsthand sources. the fact that they show clear demarcations even of nonfirsthand sources (whereas english speakers tend to consider most as hearsay) underscores that source distinctions matter for them (see also tosun et al., 2013). english speakers not only interpret most evidential adverbs as hearsay; they are also more varied in their interpretations of the source of modal expressions. turkish speakers’ source judgment of modal expressions, on the other hand, did not show any variation, with all being interpreted as a mix of inference and assumption. this may be because they already have the grammatical markers to indicate the source of information, and that leads them not to look for other forms of evidence. 4.3. interpretation of belief/confidence judgments the belief/confidence judgment task was conducted to examine the epistemic value interpretations of modal and evidential expressions. the epistemic value of english modality was previously investigated by francis and wales (1994), who found that each modal marker in english was rated at a different certainty level. must was judged as the most certain modal, which represents epistemic necessity. can, would, should and might followed after must in certainty ratings, in that order. the present study thus serves in part as a replication of francis and wales’s study. further, epistemic value ratings of turkish modal expressions were tested as well. kerimoglu (2010) discussed turkish modality in detail, although he did not mention the epistemic value of each marker. he focused on the meaning of the markers separately. the present study is the first empirical investigation of how turkish speakers rate the epistemic value of modal expressions. the purpose for obtaining confidence judgments for evidential sentences was to address the theoretical issue of the relationship between modality and evidentiality. it was of interest whether speakers could find enough information to judge the epistemic value of evidential expressions and if so, whether different sources of information would be judged at different levels of certainty. in the evidentiality literature, willet (1988) and de haan (1999) have proposed their own answers to this issue. according to willet, hearsay sources should be more reliable than inference because hearsay conveys information that was directly gathered by whoever first reported the proposition. on the other hand, de haan argued that inference should be more reliable than hearsay, because the source of inference requires a closer involvement of the speaker herself in the reported event by having to gather proof. aside from providing empirical data on this issue, another aim of the confidence judgments task was to examine if there would be a difference between turkish and english speakers in confidence judgments. the results demonstrated that english and turkish speakers judged the epistemic value of the modal sentences differently. consistent with francis and wales (1994), english speakers judged the epistemic value of modal expressions in the following order: the modal must was judged as the most certain modal (58.85%) followed by should (45.67%) and could (41.82%). might elicited the lowest confidence judgments (36.32%) by english speakers. on the other hand, turkish speakers did not show any difference in certainty between the modals. all modal expressions were judged at similar medium confidence levels. the only differentiation they made was for must and might; must (42.64%) was judged to be more confident than might (35.51%). turkish speakers not only judged the source of all modal sentences equivalently, they also judged the confidence of those modal expressions as equivalent. thus, we can conclude that for turkish speakers modals represent one member in an overlapping intersection of evidentiality and modality sets. on the other hand, just as english speakers judged the source of modal expressions variously as inference, assumption and conjecture, they also judged their confidence levels for the modals differently. thus, we can argue that for english speakers, modal expressions are represented as independent members of the intersection of evidentiality and modality sets. tosun and vaid 152 in terms of the confidence judgments of evidential expressions, similar to the modal sentences, turkish and english speakers showed different results. turkish speakers gave highest confidence ratings for the inference sentences (59.31%), followed by assumption sentences (49.13%). hearsay sentences were judged less confident (40%) than inference and assumption, and the least confident source of information was conjecture (22.10%). turkish speakers’ confidence judgments are consistent with de haan’s (1999) argument in that inference was judged as more reliable than hearsay. on the other hand, english speakers’ judgments demonstrated a different ranking. they gave hearsay sources the highest confidence rating (65.9%) followed by inference (53.13%), assumption (47.35%) and conjecture (44.85%), although assumption and conjecture did not differ significantly. english speakers’ confidence judgment pattern was consistent with willet’s (1988) argument that hearsay is more reliable than inference. the difference between turkish and english speakers’ confidence judgments of evidential expressions could be related to their source judgments of the same sentences. turkish speakers judged each source of evidence as projected in the grammar, and their source judgments varied accordingly. they judged the confidence level of each source consistently with its proximity to the proposition source. in this case, because the speaker is involved in the same environment of the reported event when the source was inference, they also reported more confidence that the event actually happened. in the hearsay source, because the speaker’s involvement is more limited than in the case of inference, they reported less confidence than they did for the inference case. thus, in turkish speakers’ minds evidential sources appear to be represented as independent sets in the intersection of the evidentiality and modality. on the other hand, as english speakers judged all of the evidential sources, except assumption, as hearsay, the difference between the confidence levels of these evidential expressions demonstrated that even one source type could be varied in terms of its reliability. in english speakers’ mind, the source hearsay is represented at three different levels of certainty. the certainty level of assumption sentences was independent from hearsay sources. before discussing the theoretical implications of the findings, there are some limitations of the study to be noted. first, the study asked participants to make metalinguistic judgments, i.e., to think about their linguistic knowledge. this is not something that language users typically do. it is possible that how participants respond on tasks required explicit metalinguistic judgments. that may not be consistent with how they actually use language in daily life. for example, an english speaker might judge apparently as hearsay but might actually use it as inference in her daily language practice. a second limitation is that the stimuli were not presented in a discourse context. in actual language use, contextual cues are usually available to constrain how discourse is interpreted. it is quite likely that the sentences presented in this study would have been judged differently if they had been presented within a given context. thus, this study might measure metalinguistic knowledge rather than language use. further work is needed to address how context, including sociopragmatic factors, may interact with particular markers of modality or evidentiality in affecting judgments of source or confidence. a third limitation is that the source of conjecture might lead to some noise due to the vagueness of the source itself. with these limitations in mind, we turn to the theoretical implications of the findings. 4.4. theoretical implications of overall findings as the not-enough-information analyses demonstrated, evidentiality and modality are in a close relationship. participants indicated that they had enough information to make source judgments of modal expressions and epistemic value judgments of evidential expressions. this finding disproves the complete disjointment view by aikhenvald (2004), de haan (1999), lazard (2001) and oswalt (1986). on the other hand, evidential and modal expressions were not used completely interchangeably. for both turkish and english speakers, modal and evidential expressions were used to indicate different sources of evidence and confidence levels. thus, the identity view (by matthewson, 2010) was also disproved. in summary, evidentiality and modality evidentiality and modality 153 . hearsay are neither completely independent linguistic properties from each other, nor are they the same structures conveying completely the same meaning, but they display a close relationship. further, the results suggest that turkish, as a grammaticalized evidential language, and english, as a lexicalized evidential language, exhibit different evidentiality-modality relationships. english speakers’ epistemic value judgments of evidential expressions were almost 100%, indicating that each and every evidential expression conveys epistemic value, whereas source judgments of modal expressions yielded more “not enough information” responses than evidential sentences. this pattern supports an inclusion set in which evidentiality is considered a subtype of epistemic modality (see figure 6), as proposed by bybee (1985), palmer (1986), willett (1988), and mithun (1999). according to this view all evidential expressions convey epistemic value; however, not all modal expressions convey the source of the evidence. further, the source and confidence judgments of english speakers displayed how the members of this subset were shaped. according to the analyses, each modal expression (must, should, could, might) was judged differently from one another in both source and confidence judgments, which was displayed as independent members in the english evidentiality-modality set structure. assumption (presumably) as an evidential expression was judged distinctively from other sources and from modals in both source and confidence judgment tasks. thus, it was displayed as another independent member of the subset. further, other evidential expressions (hearsay, inference, conjecture) were judged as the same source (hearsay) but at different epistemic value levels (strong, and weak). therefore, these sources were displayed as the hearsay subset in the evidentiality subset with three members. figure 6. the prototypical venn diagram of the english evidentiality and modality relationship. on the other hand, turkish speakers demonstrated a different relationship between evidentiality and modality. for one thing, turkish speakers indicated more nei responses than english speakers. turkish speakers’ evidentiality and modality relationship structure appears to be more of an overlapping structure (see figure 7). the overlap view argues that not all modals convey the source of information and not all evidentials convey epistemic value, however, there is overlap between them where they convey both source and epistemic value information (delancey, 2001; faller, 2002; van der auwera & plungian, 1998). turkish speakers sometimes use evidential expression only to indicate the source of information as in givón’s (1982) life of buddha example. for example, when the prophets’ life is narrated by religious turkish people they use nonfirsthand marker, although they are quite certain about what happened to the prophets. similarly, they sometimes use modal markers just to indicate their degree of certainty. modality • presumably • must • should • could • might evidentiality • strong reportedly • weak apparently, supposedly tosun and vaid 154 . . for example, a turkish speaker might say that “i might have read that book”, regardless of the firsthand source of information, to indicate purely the possibility of the proposition. scholarly discussions of the interpretation of source for modals has focused on the modal must. however, our source and confidence judgment analyses allowed for an examination of several different modals and it was possible to display how the overlapping section of the relationship was contained. turkish speakers judged each and every evidential expression independently from one another in terms of both the source and confidence tasks. thus, all evidential expressions were represented as independent members of the overlap set. further, all of the modal expressions were judged as the same in terms of both the source and confidence judgment tasks. thus, all modals were displayed as only one member of the overlapping sets. figure 7. the prototypical turkish evidentiality and modality relationship. in summary, the results suggest a close relationship between evidentiality and modality, although the nature of the relationship depends on the nature of evidentiality and modality indication of the languages. tosun and vaid (2016) tested the influence of evidential markers on decision-making. they found that evidentiality is one variable that in interaction with other variables (e.g., order of the given information, the nature of information) influences trustworthiness of given information and final decisions. the results of this research suggest that the epistemic value of evidential expressions may influence individuals’ decisions. people consider facts reported with firsthand source expressions as having a very strong epistemic value. the underlying meaning of the firsthand expression is that the speaker herself personally witnessed what was reported. thus, it is considered more secure to rely on the firsthand evidential expressions. the present study demonstrated that people interpret the reliability of the source of information from the evidential marker. the results of the study are of particular importance in settings where the decisions that are made have a long-lasting impact on individuals’ lives, such as courtrooms, political elections, medical environments, marketing and business, and academia. in these and similar settings linguistic framing can significantly affect decision-making (for more information, see filipovic, 2013; matlock, 2012; tannen, 1993). qualitative investigations of discourse and evidentiality demonstrate that evidential expressions are used to indicate assertion and commitment (berlin, 2011), responsibility, entitlement, certainty of knowledge, denial (as nonfirsthand) of the described situation (fox, 2008), unbelievable and unreliable situations as fairy tales (johanson, 2003) and the distance • reportedly • apparently • supposedly • presumably • modals modality evidentiality evidentiality and modality 155 between the speaker and the described situation (aksu-koc & slobin, 1986). thus, discourse analysis reveals that evidentiality is used by speakers to frame their stories with the underlying meaning of various evidential sources, which in turn can shape interlocutors’ judgments and decisions. by way of conclusion, the impact of this study can be summarized as follows: first, it was demonstrated that there was a clear relationship between evidentiality and modality. evidential expressions convey the epistemic value of the propositions along with the source of evidence. by the same token, modal expressions convey the source of evidence along with the epistemic value of the propositions. second, the classifications of the evidential and modal expressions in the linguistic repertoire of speakers change from language to language depending on how evidentiality and modality are marked in each language. this study, to our knowledge, is the first investigation that aimed to address empirically a long-standing debate on the relationship between evidentiality and modality. therefore, the findings are a contribution to the evidentiality literature. further, along with the results of the study, the methodology of the investigations including the materials, design, and procedure is another significant contribution for future investigations. acknowledgements this work is based in part on a doctoral dissertation by the first author conducted under the supervision of the second author. the work was supported by a college of liberal arts vision 2020 dissertation enhancement award from texas a&m university. tosun and vaid 156 references aijmer, k. (1980). evidence and the declarative sentence. stockholm: almqvist & wiksell. aikhenvald, a. y. (2004). evidentiality. new york, ny: oxford university press. aksu-koc, a. (2009). evidentials: an interface between linguistic and conceptual development. in j. guo, e. lieven, n. budwig, s. ervin-tripp, k. nakamura, & s. özçalışkan (eds.), crosslinguistic approaches to the psychology of language: research in the tradition of dan slobin (pp. 531-540). new york: psychology press. aksu-koc, a. (2000). some aspects of the acquisition of evidentials in turkish. in l. johanson & b. utas, (eds.), evidentials: turkic, iranian and neighboring languages (pp. 15-28). berlin: mouton de gruyter. aksu-koc, a., & alici, d. m. (2000). understanding sources of beliefs and marking of uncertainty: the child’s theory of evidentiality. in e. v. clark (ed.), proceedings of the thirtieth annual child language research forum (pp. 123-130). stanford: center for the study of language and information. aksu-koc, a. & slobin, d. (1986). a psychological account of the development and use of evidentials in turkish. in w. chafe & j. nichols (eds.), evidentiality: the linguistic coding of epistemology (pp. 159-167). norwood, nj: ablex. berlin, l. n. (2011). i think, therefore…commitment in political testimony. journal of language and social psychology, 27, 372-383. biber, d., & finegan, e. (1989). styles of stance in english: lexical and grammatical marking of evidentiality and affect. text, 9, 93-124. bybee, j. (1985). morphology: a study of the relation between meaning and form. amsterdam: john benjamins. chafe, w. (1986). evidentiality in english conversation and academic writing. in w. chafe & j. nichols (eds.), evidentiality: the linguistic coding of epistemology (pp. 261-272). norwood, nj: ablex. de haan, f. (1998). the cognitive basis of visual evidentials. in a. cienki, b. j. luka, & m. b. smith (eds.), conceptual and discourse factors in linguistic structure (pp. 91-105). stanford: csli publications. de haan, f. (1999). evidentiality and epistemic modality: setting boundaries. southwest journal of linguistics, 18, 83-102. de haan, f. (2004). encoding speaker perspective: evidentials. in z. frajzyngier, a. hodges & d. s. rood, (eds.), linguistic diversity and language theories (pp. 379-397). amsterdam: john benjamins. de lancey, s. (2001). the mirative and evidentiality. journal of pragmatics, 33, 369-382. de villiers, j., & garfield, j. (2009). evidentiality and narrative. journal of consciousness studies, 16, 191-217. dendale, p., & tasmowski, l. (2001). introduction: evidentiality and related notions. journal of pragmatics, 33, 339-348. faller, m. t. (2002). semantics and pragmatics of evidentials in cuzco quechua. ph.d. dissertation, stanford university. filipovic, l. (2013). the role of language in legal contexts: a forensic cross-linguistic viewpoint. in m. freeman & f. smith (eds.), law and language, (pp. 328-343). oxford, uk: oxford university press. fitneva, s. a. (2008). the role of evidentiality in bulgarian children’s reliability judgments. journal of child language, 35, 845-868. fitneva, s. a. (2009). evidentiality and trust: the effect of informational goals. in s. a. fitneva & t. matsui (eds.), evidentiality: a window into language and cognitive development, new directions for child and adolescent development, (pp. 49–61). san francisco: jossey-bass. evidentiality and modality 157 fox, b. a. (2008). evidentiality: authority, responsibility, and entitlement in english conversation. journal of linguistic anthropology, 11, 167-192. francis, j., & wales, r. (1994). prosodic cues, message interpretation, and impression formation. journal of language and social psychology, 13, 34-44. gisborne, n., & holmes, j. (2007). a history of english evidential verbs of appearance. english language and linguistics, 11, 1-29. givón, t. (1982). evidentiality and epistemic space. studies in language, 6, 23-49. izvorski, r. (1997). the present perfect as an epistemic modal. in a. lawson & e. cho (eds.), seventh conference of semantics and linguistic theory (pp. 1-18). ithaca, ny: clc publications. johanson, l. (2000). turkic indirectives. in l. johanson & b. utas, (eds.), evidentials: turkic, iranian and neighboring languages (pp. 61-87). berlin: mouton de johanson, l. (2003). evidentiality in turkic. in a. y. aikhenvald & r. m. w. dixon (eds.), studies in evidentiality (pp. 273-291). amsterdam: john benjamins. kerimoglu, c. (2010). on the epistemic modality markers in turkey turkish: uncertainty. international periodical for the languages, literature and history of turkish or turkic, 5, 434-478. kornfilt, j. (1997). turkish. london: routledge. lazard, g. (2001). on the grammaticalization of evidentiality. journal of pragmatics, 33, 359367. matlock, t. (1989). metaphor and grammaticalization of evidentials. berkeley linguistics society, 15, 215-225. matlock, t. (2012). framing political messages with grammar and metaphor. american scientist, 100, 478-483. matsui, t., yamamoto, t. & mccagg, p. (2006). on the role of language in children’s early understanding of others as epistemic beings. cognitive development 21(2), 158–170. matsui, t., & miura, y. (2009). children’s understanding of certainty and evidentiality: advantage of grammaticalized forms over lexical alternatives. in s. a. fitneva & t. matsui (eds.), evidentiality: a window into language and cognitive development, new directions for child and adolescent development (pp. 63–77). san francisco: josseybass. matthew, m. r. & busemeyer, j. r. (2011). explaining cultural differences in decision making using decision field theory. in r. proctor, s. y. nof, y. yih, and g. salvendy (eds.) cultural factors in systems design: decision making and action (pp. 17 – 32). boca raton, fl: taylor & francis. matthewson, l. (2010). evidential restrictions on epistemic modals. workshop on epistemic indefinites, university of göttingen. mithun, m. (1999). the languages of native north america. cambridge: cambridge university press. moore, c., bryant, d., & furrow, d. (1990). mental terms and the development of certainty. child development, 60, 167–171. moore, c., pure, k., & furrow, d. (1990). children’s understanding of the modal expression of speaker certainty and uncertainty and its relation to the development of a representational theory of mind. child development, 61, 722–730. mortensen, j. (2006). epistemic and evidential sentence adverbials in danish and english: a comparative study. ph.d. dissertation, roskilde university. mushin, i. (2001). evidentiality and epistemological stance: narrative retelling. amsterdam: john benjamins. nuyts, j. (2001). epistemic modality, language, and conceptualization: a cognitive-pragmatic perspective. amsterdam: john benjamins. tosun and vaid 158 oswalt, r. l. (1986). the evidential system of kashaya. in w. chafe & j. nichols (eds.), evidentiality: the linguistic coding of epistemology (pp. 29-45). norwood, nj: ablex. öztürk, ö. (2008). acquisition of evidentiality and source monitoring. ph.d. dissertation, university of delaware. öztürk, ö., & papafragou, a. (2005). the acquisition of evidentiality in turkish. university of pennsylvania working papers in linguistics, 11, 1-14. papafragou, a., li, p., choi, y., & han, c. (2007). evidentiality in language and cognition. cognition 103, 253-299. palmer, f. r. (2013). modality and english modals. new york, ny: routledge. palmer, f. r. (1986). mood and modality. cambridge: cambridge university press. plungian, v. a. (2001). the place of evidentiality within the universal grammatical space. journal of pragmatics, 33, 349-357. segalowitz, n., doucerain, m., meuter, r.i., zhao, y., hocking, j., & ryder, a.g. (2016). comprehending adverbs of doubt and certainty in health communication: a multidimensional scaling approach. frontiers in psychology, 7, article no. 558. slobin, d. i. & aksu, a. (1982). tense, aspect and modality in the use of the turkish evidential. in p. hopper (ed.), tense-aspect: between semantics and pragmatics (pp. 185-200). amsterdam: john benjamins. tannen, d. (ed.). (1993). framing in discourse. new york, ny: oxford university press. tosun, s. & vaid, j. (2012). gillian reportedly ran a red light: interpretations of non first-hand assertions of sources of information. poster presented at annual meeting of the society for text and discourse, montreal. tosun, s. & vaid, j. (2016). making a story make sense: does evidentiality matter in discourse coherence? applied psycholinguistics, 37, 1337-1367. tosun, s., vaid, j., & geraci, l. (2013). does obligatory linguistic marking of source of evidence influence source memory? a turkish/english investigation. journal of memory and language, 69(2), 121-134. van der auwera, j., & plungian, v. a. (1998). on modality’s semantic map. linguistic typology, 2, 79-124. von fintel, k., & gillies, a. s. (2010). must…stay…strong! natural language semantics, 18, 351-383. willett, t. (1988). a cross-linguistic survey of the grammaticization of evidentiality. studies in language, 12, 51-97. evidentiality and modality 159 appendix a: sample response form instructions (source judgment) in this experiment we would like to see how listeners interpret sentences in their language that differ in subtle ways. you will be shown a list of sentences describing some event. please read each sentence as carefully and closely as you can. upon reading each sentence, you will be asked to make some judgments, which involve deciding what the source of the reported event is – i.e., is the source hearsay, an inference based on other observation, an assumption or a conjecture. as an example, let’s say the sentence was about someone saying there was a mouse in a house. in this case, “hearsay” would indicate that the information is based on hearing it from someone else. for example, the person’s friend told him/her that there was a mouse in the house. “inference” means that the person who reported the information saw some signs related to the information but did not see what happened. for example, the person saw footprints of the mouse but did not see the mouse itself in the house. “assumption” indicates that the person who reported the information did not see what happened but reasoned what must have happened based on some knowledge. for example the town where the house located was known for its mice. “conjecture” would mean that the person who stated that information did not see the event, but also the source of information was unspecified. for example, the person heard people talking about a mouse and s/he guessed that there was a mouse in the house. instructions (confidence judgment) in this experiment we would like to see how listeners interpret sentences in their language that differ in subtle ways. you will be shown a list of sentences describing some event. please read each sentence as carefully and closely as you can. upon reading each sentence, you will be asked to make some judgments, which involve how confident you feel about whether the reported information actually took place: are you extremely confident, quite confident, somewhat confident or not at all confident. for example, for the example given above, you would indicate how confident you would be that there was actually a mouse in the house if you heard the sentence. remember that there are no right or wrong answers in this test. we are simply interested in seeing how listeners typically interpret sentences conveying information about an event. tosun and vaid 160 evidentiality and modality 161 appendix b: source judgments a summary of turkish and english speaking participants’ mean source judgment responses (in percentage) to evidential and modal sentence forms source judgment responses sentence form language hearsay inference assumption conjecture turkish 37.91 (38.08) 13.01 (22.18) 12.39 (18.18) 36.68 (37.71) hearsay english 75.83 (28.24) 9.50 (15.00) 6.83 (11.71) 7.83 (16.78) turkish 18.23 (24.29) 59.68 (33.41) 13.64 (18.33) 8.44 (14.90) inference english 52.61 (35.24) 24.33 (25.42) 13.69 (18.77) 9.37 (16.17) turkish 17.61 (24.72) 42.02 (25.87) 31.41 (27.60) 8.95 (15.33) assumption english 24.81 (26.47) 34.46 (27.46) 29.19 (28.10) 11.54 (16.57) turkish 27.22 (34.20) 8.26 (14.52) 15.05 (21.68) 49.47 (38.68) conjecture english 52.89 (36.05) 18.07 (21.95) 18.98 (24.15) 10.06 (16.05) turkish 11.64 (16.48) 45.78 (26.83) 33.50 (25.40) 9.09 (17.09) necessity english 4.07 (9.55) 37.83 (26.27) 44.01 (25.55) 14.08 (15.76) turkish 12.55 (20.81) 44.70 (31.75) 26.37 (25.81) 16.38 (25.79) weak necessity english 7.78 (16.95) 29.91 (27.55) 31.18 (32.62) 31.14 (33.35) turkish 12.08 (19.39) 39.59 (27.40) 38.13 (26.84) 10.19 (18.64) probability english 8.59 (19.50) 28.75 (21.98) 38.34 (25.19) 24.32 (25.50) turkish 11.11 (19.69) 40.64 (28.51) 40.23 (29.73) 8.02 (13.79) possibility english 4.46 (8.94) 29.08 (23.83) 36.52 (24.76) 29.94 (27.30) note. standard deviations are presented in parenthesis. tosun and vaid 162 appendix c: confidence judgments a summary of turkish and english speaking participants’ confidence judgments of evidential and modal sentence forms sentence form turkish (n=52) english (n=60) hearsay 40.00 (25.84) 65.90 (2.65) inference 59.31 (22.92) 53.13 (17.54) assumption 49.13 (23.28) 47.35 (16.8) ev id en ti al conjecture 22.10 (23.87) 44.85 (17.1) must 42.64 (15.65) 58.85 (17.74) should 44.27 (21.85) 45.67 (27.7) could 38.92 (19.32) 41.82 (22.67) m od al might 35.51 (17.92) 36.32 (20.38) note. standard deviations are presented in parenthesis. dialogue & discourse 7(3) (2016) 4–33 doi: 10.5087/dad.2016.301 the dialog state tracking challenge series: a review jason d. williams jason.williams@microsoft.com microsoft research one microsoft way redmond, wa, 98052, usa antoine raux araux@fb.com facebook 1 facebook way menlo park, ca 94025, usa matthew henderson∗ matthen@google.com department of engineering university of cambridge cb2 1pz, uk editor: david schlangen submitted 04/15; accepted 02/16; published online 04/16 abstract in a spoken dialog system, dialog state tracking refers to the task of correctly inferring the state of the conversation – such as the user’s goal – given all of the dialog history up to that turn. dialog state tracking is crucial to the success of a dialog system, yet until recently there were no common resources, hampering progress. the dialog state tracking challenge series of 3 tasks introduced the first shared testbed and evaluation metrics for dialog state tracking, and has underpinned three key advances in dialog state tracking: the move from generative to discriminative models; the adoption of discriminative sequential techniques; and the incorporation of the speech recognition results directly into the dialog state tracker. this paper reviews this research area, covering both the challenge tasks themselves and summarizing the work they have enabled. keywords: dialog state tracking, spoken dialog systems, dialog modeling, conversational systems, spoken language understanding 1. introduction conversational systems are increasingly becoming a part of daily life, with examples including apple’s siri, google now, nuance dragon go, xbox and cortana from microsoft, and numerous start-ups. figure 1 shows the principal components of a modern spoken dialog system. first, the user produces an utterance as audio. then automatic speech recognition (asr) converts this audio into words in text form. next, the words in an utterance are converted to a meaning representation using spoken language understanding (slu). this slu result is then passed to the dialog state tracker (dst) which updates its estimate of the dialog state. this new dialog state is passed to the dialog policy that decides which action to take. natural language generation (nlg) and text-tospeech (tts) convert this action into words and then into audio. the cycle then repeats. . ∗ matthew henderson is now at google. c©2016 jason d. williams, antoine raux, and matthew henderson this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). dialog state tracking overview leaving from downtown leaving at one p m arriving at one p m 0.6 0.2 0.1 { from: downtown } { depart-time: 1300 } { arrive-time: 1300 } 0.5 0.3 0.1 from: cmu to: airport depart-time: 1300 confirmed: no score: 0.10 from: cmu to: airport depart-time: 1300 confirmed: no score: 0.15 from: downtown to: airport depart-time: -confirmed: no score: 0.65 automatic speech recognition (asr) spoken language understanding (slu) dialog state tracker (dst) dialog policy act: confirm from: downtown from downtown, is that right? natural language generation (nlg) text to speech (tts) figure 1: principal components of a spoken dialog system. the topic of this paper is the dialog state tracker (dst). the dst takes as input all of the dialog history so far, and outputs its estimate of the current dialog state – for example, in a restaurant information system, the dialog state might indicate the user’s preferred price range and cuisine, what information they are seeking such as the phone number of a restaurant, and which concepts have been stated vs. confirmed. dialog state tracking is difficult because asr and slu errors are common, and can cause the system to misunderstand the user. at the same time, state tracking is crucial because the dialog policy relies on the estimated dialog state to choose actions – for example, which restaurants to suggest. in the literature, numerous methods for dialog state tracking have been proposed. these are covered in detail in section 3; illustrative examples include hand-crafted rules (larsson and traum, 2000; bohus and rudnicky, 2003), heuristic scores (higashinaka et al., 2003), bayesian networks (paek and horvitz, 2000; williams and young, 2007), and discriminative models (bohus and rudnicky, 2006). techniques have been fielded which scale to realistically sized dialog problems and operate in real time (young et al., 2010; thomson and young, 2010; williams, 2010; mehta et al., 2010). in end-to-end dialog systems, dialog state tracking has been shown to improve overall system performance (young et al., 2010; thomson and young, 2010). despite this progress, direct comparisons between methods have not been possible because past studies use different domains and different system components for asr, slu, dialog policy, etc. moreover, there has not been a standard task or methodology for evaluating dialog state tracking. together these issues have limited progress in this research area. the dialog state tracking challenge (dstc) series has provided a first common testbed and evaluation suite for dialog state tracking. three instances of the dstc have been run over a three 5 williams, raux, and henderson year period. each instance has released a public corpus of transcribed and labeled human-computer dialogs along with baseline trackers and evaluation tools, and each instance has explored a new aspect of dialog state tracking. between seven and nine teams have entered each challenge. this challenge task series has spurred significant work on dialog state tracking, yielding both numerous new techniques as well as a standard set of evaluation metrics. this paper is organized as follows. first, section 2 formalizes the dialog state tracking problem, and section 3 reviews solution methods from the literature. section 4 then covers the first three instances of the dialog state tracking challenge – dstc1, dstc2, and dstc3 – including the task design, data, evaluation methodology, and baselines. section 5 then covers results from the challenge tasks. finally, section 7 concludes. 2. dialog state tracking: problem definition first, we define the concept of dialog state. a dialog state st is a data structure drawn from a set s that summarizes the dialog history up to time t to a level of detail that provides sufficient information for choosing the next system action. in practice, the dialog state typically encodes the user’s goal in the conversation along with relevant history – for example, in the bus timetable domain, s may encode which bus stop the user wants to leave from, where they are going to, and whether the system has already offered a bus on that route. a dialog state tracker takes as input all of the observable elements up to time t in a dialog, including all of the results from the asr and slu components, all system actions taken so far, and external knowledge sources such as bus timetable databases and models of past dialogs. because the asr and slu are imperfect and prone to errors, they may output several conflicting interpretations. specifically, the asr may output an n-best list of sentences, a word confusion network (mangu et al., 2000), or a lattice; the slu may output an n-best list of interpretations. figure 1 shows example asr and slu n-best lists. given these inputs, the tracker then outputs its estimate of the current state of the dialog s. the goal is to correctly identify the true current state s∗ of the dialog – for example, the bus stops the user has actually said they want or whether the user wants the address, opening hours, or price range of a particular restaurant. however, the true state is typically not directly observable from the inputs, for a variety of reasons: errors in speech recognition and language understanding, ambiguous or underspecified utterances, unsignaled changes in the user’s goal, etc. therefore, robust dialog state trackers typically output a distribution over multiple possible dialog states b(s). a distribution is useful because it provides a principled representation of the uncertainty in the dialog state. it also gives a clear basis for taking clarification actions: for example, if the distribution’s probability mass is concentrated on two states that differ only in which type of food the user is asking for (say, “indian” and “italian”), this allows the system to ask “did you want indian or italian food?”. figure 2 shows an example of the dialog state tracking process, and illustrates how effective dialog state tracking can overcome some of the errors received from the asr and slu. in this paper – and in the dstc challenge series – we have taken the view that a dialog state consists of elements with human-interpretable meanings, such as values of bus stops, dates, times, whether conditions have been met, etc. we have further assumed that a dialog state tracker produces the key input to an action selector – sometimes also called a “dialog policy” – that chooses an action or response based primarily on the current dialog state. this view is in line with widely accepted theoretical models of conversation, such as clark’s common ground (clark, 1996) and various 6 dialog state tracking overview cheap restaurant restaurant italian 0.6 0.2 0.1 inform(price=cheap) inform(food=italian) 0.5 0.3 east area italian yeah 0.5 0.3 0.1 inform(area=east) inform(food=italian) affirm() 0.6 0.3 0.2 price=cheap food=italian food=italian, price=cheap price=cheap food=italian food=italian, price=cheap area=east food=italian, area=east price=cheap, area=east food=italian, price=cheap, area=east [none] [none] how can i help you? welcome() an italian restaurant sorry, what price did you want? request(price) uh, italian system action / user response asr output slu output state score dialog state tracker outputsdialog state tracker inputs figure 2: overview of dialog state tracking. in this example, the dialog state contains the user’s desired restaurant search criteria. at each turn, the system produces a spoken output. the user’s spoken response is converted into an n-best list of word hypotheses by the asr, and then into another n-best list of meaning hypotheses by the slu. both lists have confidence scores attached. a set of dialog state hypotheses is enumerated, here by simply considering all slu results observed so far, including the current turn and all previous turns. then the dialog states are scored. note how observing “italian” a second time in the asr/slu causes the dialog state for “food=italian” to accumulate considerable probability mass in the second turn, even through “italian” was never the top hypothesis from the asr or slu. this illustrates one way that dialog state tracking can overcome local asr/slu errors. models of dialog as joint action (cohen and levesque, 1990), which assume that dialog relies on some (usually shared) representation of the participants’ joint intentions and beliefs. while this is the dominant approach, it is worth mentioning alternatives. first, dialog state can instead be a latent representation, with responses selected – or in principle generated – using continuous-space representations (lowe et al., 2015). further, it is possible to dispense with state tracking altogether, and instead produce responses based only on the most recent user turn (ritter et al., 2011) – or in principle directly from features of the dialog history. a comparison with these methods would require end-to-end evaluations of spoken dialog systems, which is outside the scope of the dstc series, and this paper.1 in the next section we review methods for dialog state tracking. 3. methods for dialog state tracking broadly speaking there are three families of dialog state tracking algorithms: hand-crafted rules, generative models, and discriminative models. 1. dialog state tracking in situated environments – for example, robots or embodied agents – is also out of scope for this review, but it is noted that dialog state tracking is also used in this setting (bohus and horvitz, 2009; ma et al., 2012). 7 williams, raux, and henderson 3.1 hand-crafted rules for dialog state tracking early spoken dialog systems used hand-crafted rules for dialog state tracking. in their earliest form, these approaches generally considered only a single slu result, and tracked a single hypothesis for the dialog state. this design reduces the dialog state tracking problem to an update rule f (s, ũ′) = s′ that maps from an existing state s and the 1-best slu result ũ′ to a new state s′. for example, the mit jupiter weather information system maintained a set of state variables which were updated using hand-written rules in a dialog control table (zue et al., 2000). similarly, the information state update approach used hand-written update rules to track a rich data structure called an “information state” (larsson and traum, 2000). hand-crafted rules have the benefit that they do not require any data to implement, which is a benefit for bootstrapping. rules also provide an accessible way for developers to incorporate domain knowledge. however, one short-coming of tracking a single dialog state is an inability to make use of the entire asr or slu n-best list, and the benefit of tracking multiple dialog states was suggested nearly two decades ago by pulman (1996). thus, more recent dialog state trackers based on hand-crafted rules compute scores for all dialog states suggested by the whole asr/slu n-best list (wang and lemon, 2013; sun et al., 2014a). these methods use hand-designed formulas to compute a posterior b(s) of a dialog state s given asr/slu confidence scores and previous estimates of b(s), and thus can overcome some slu errors (figure 2). using hand-designed formulas for computing b(s) suffers from a crucial limitation: formula parameters are not derived directly from real dialog data, so they require careful tuning and do not benefit or learn from dialog data. this limitation motivates the use of data-driven techniques, which can automatically set parameters in order to maximize accuracy. chief among the data-driven techniques are generative and discriminative models, described next. 3.2 generative models for dialog state tracking generative approaches posit that dialog can be modeled as a bayesian network that relates the dialog state s to the system action a, the (true, unobserved) user action u, and asr or slu result ũ. when the system action and asr/slu result are observed, a distribution over possible dialog states can be computed by applying bayesian inference. a number of probabilistic formulations have been explored for how to relate these quantities; one illustrative example is: b′(s′) = η ∑ u′ p (ũ′|u′)p (u′|s′, a) ∑ s p (s′|s, a)b(s) (1) where b(s) is the previous distribution over dialog states, b′(s′) is the (updated) distribution over dialog states being estimated, p (ũ′|u′) is the probability of the asr/slu producing the observed output ũ′ given the (true, unobserved) user action u′, p (u′|s′, a) is the probability of the user taking action u′ given the true dialog state s′ and system action a, p (s′|s, a) is probability of the dialog state changing to s′ given it is currently s and the system takes action a, and η is a normalizing constant. variants of eq. 1 account for different factorizations of the hidden state. for example, williams and young (2007) includes a term that accumulates dialog history, such as whether the contents of s has been confirmed or not. devault and stone develop a bayesian network that includes separate random variables for an observed dialog action and an underlying intention, and includes conditional probability terms that express common-sense relationships between actions, intentions, and 8 dialog state tracking overview plausible states termed “contexts” (devault and stone, 2007; devault, 2008). other factorizations have also been presented for modeling dialog in specific settings, such as troubleshooting an internet router (williams, 2007). eq. 1 most closely follows williams (2008); the appendix of williams (2012a) provides a derivation. in all of these examples, the key assumption is that a distribution over possible (hidden) dialog states can be inferred using a bayesian network that encodes a designer’s knowledge about conversation. the parameters of the models must be estimated of course; this can be done either from labeled dialogs, or inferred from unlabeled dialogs using methods such as expectation maximization (syed and williams, 2008) or expectation propagation (thomson et al., 2010). early approaches to generative dialog state tracking enumerated all possible dialog states, then used variants of eq. 1 to score them (roy et al., 2000; zhang et al., 2001; heckerman and horwitz, 1998; horvitz and paek, 1999; meng et al., 2003; williams et al., 2005). this approach is quadratic in the number of dialog states, which is intractable, particularly given that eq. 1 must run in real time and the number of states s can be enormous. this limitation has led to two approximations: maintaining a “beam” of only the most likely members of s (young et al., 2007; devault and stone, 2007; devault, 2008; kim et al., 2008; henderson and lemon, 2008; mehta et al., 2010; williams, 2010; raux and ma, 2011; gasic and young, 2011), or further factorization of eq. 1 (williams, 2007; bui et al., 2009; thomson and young, 2010). these approximations enable generative models to operate in real-time, but impose other constraints, such as limiting the form of p (s′|s, a), which can restrict the classes of dialogs that can be accurately modeled (young et al., 2013). in end-to-end evaluations, generative approaches have been shown to yield better dialog performance than hand-crafted rules (young et al., 2010; thomson and young, 2010). even so, generative models cannot easily incorporate large sets of potentially informative features from the asr, slu, dialog history, and elsewhere: all dependencies between features must be explicitly modeled, which requires an impractical amount of data. as a result, for tractability, generative models generally make independence assumptions which are invalid, or important features of dialog history have to be ignored, which introduce violation of the markov assumption. for example, it is often assumed that errors are generated from a uniform distribution, when in fact they are highly correlated: “twenty” is much more often mis-recognized as “seventy” than as “downtown pittsburgh” (williams, 2012c). the net effect is poor estimates of b(s). together these issues have spurred interest in discriminatively trained direct models, covered next. 3.3 discriminative models for dialog state tracking in contrast to generative models, discriminative approaches to dialog state tracking compute scores for dialog states with discriminatively trained conditional models of the form b′(s′) = p (s′|f ′), where f ′ are features extracted from the asr, slu, and dialog history. the key benefit of discriminative models are that they can incorporate a large number of features, and can be optimized directly for prediction accuracy. the first presentation of discriminative state tracking trained from data is believed to be bohus and rudnicky (2006). here, a hand-written rule enumerates a set of k dialog states to score, for example by considering the top s1 slu hypotheses from the current turn, top s2 slu hypotheses from the previous turn, and the top s3 slu hypothesis from the turn before that. an additional state hypothesis s accounts for the situation when none of the hypotheses is correct, for a total of k = s1+s2+s3+1 states to score. with a fixed number k of classes, standard multiclass logistic 9 williams, raux, and henderson regression classification is then applied, in which one weight is estimated for every (class,feature) pair. features were taken from slu output and dialog history. subsequent work has explored numerous variations of this approach. metallinou et al. (2013) alter the logistic regression model so that it learns a single weight for each feature. this enables an arbitrary number of hypotheses to be scored since the number of weights to learn no longer increases with the number of state hypotheses to score. williams (2014) applies a ranking algorithm which has the ability to construct conjunctions of features. henderson et al. (2013) applies a deep neural network as a classifier. all of the approaches above encode dialog history in the features to learn a simple classifier. by contrast, three other approaches have explicitly modeled dialog as a sequential process. first, a discriminative markov model can be applied, where the distribution from the previous turn’s prediction can be used as a feature (ren et al., 2014b,a). second, dialog can be cast as a conditional random field (crf) (lafferty et al., 2001), in which features are associated with each dialog turn, and crf decoding determines the most likely final dialog state conditioned on the entire sequence (lee and eskenazi, 2013; ren et al., 2013; kim and banchs, 2014; ma and fosler-lussier, 2014c). third, recurrent neural networks can be estimated where the inputs are the observed asr/slu results, and the output is a distribution over dialog states (henderson et al., 2014d). henderson et al. (2014d) is also notable for operating directly on asr output, without an slu (c.f. figure 1). this has two benefits: first, it removes the need for feature design, and the risk of omitting an important feature, which can degrade performance unexpectedly (williams, 2014). second, it avoids the work of building a separate slu model. all of the approaches above require in-domain dialog data for training. when a small amount of labeled data exists for the target domain, multi-domain learning can be applied (williams, 2013). when no labeled data exists – for example, when a system is first deployed – it is possible to use unsupervised adaptation from a base model for a related domain. the basic idea is to find points in the dialog where a state component value is assigned a high score – such as food=italian – then treat that predicted value as a label, and adjust model parameters to predict that label earlier in the dialog (lee and eskenazi, 2013; henderson et al., 2014e). this approach allows a generic slot tracking model to be adapted to a specific slot for which labeled data does not exist. the approaches above infer user behavior directly from the dialog data, and make no a priori assumptions about the structure of p (s′|f ′). since some properties of human behavior with dialog systems is known – for example, that people typically change their goal only in certain situations – it is possible to devise rules that score dialog states using functions of the asr or slu confidence scores, and then estimate a handful of parameters of the rules from data (higashinaka et al., 2003; kadlec et al., 2014; sun et al., 2014a). since the primary source of uncertainty in dialog state tracking is the asr or slu, these methods can perform very well when the confidence scores are reliable. with so many methods for dialog state tracking proposed, it is vital to have benchmark tasks for making performance comparisons. this need motivated the dialog state tracking challenge series of research community tasks, described next. 10 dialog state tracking overview 4. challenge tasks 4.1 overview the over-arching research aim of the dstc series has been to understand which existing methods for dialog state tracking perform best, and encourage new work that advances the state-of-the-art. as part of that aim, the dstc series has also examined which evaluation measurements are appropriate for dialog state tracking. to date there have been three completed dialog state tracking challenges. each has used logs of human-computer dialogs in different domains, with different properties: dstc1 used a corpus of dialogs with various systems that participated in the spoken dialog challenge (sdc) (black et al., 2010), provided by the dialog research center at carnegie mellon university. in the sdc, telephone calls from real passengers of the port authority of allegheny county, which runs city buses in pittsburgh, were forwarded to dialog systems built by different research groups. the goal was to provide bus riders with bus timetable information. for example, a caller might want to find out the time of the next bus leaving from downtown to the airport. in this domain, the goal of the user typically remains fixed for the duration of the dialog. dstc2 aimed to extend the results of dstc1 to another domain, as well as broaden the scope to include user goal changes. this challenge relied on a corpus of dialogs in the restaurant search domain between paid participants (through amazon mechanical turk) and various systems developed at cambridge university (young et al., 2014). the goal of the user is to find specific information such as price range or phone number about a restaurant that fulfills a number of constraints such as cuisine or neighborhood. dstc3 expanded the domain of dstc2 to include new slots which do not occur in the training data. this simulates the crucial problem of adapting a dialog system to a new domain for which little dialog data is available, while data for a similar but different domain might already exist. dstc3 used all the data from dstc2 as training set, as well as a new set of dialogs (also collected by cambridge university researchers (jurčı́ček et al., 2011)) on a broader tourist information domain, covering bars and cafes in addition to restaurants. 4.2 challenge design the dialog state tracking challenges take a corpus-based approach – i.e., dialog state trackers are trained and tested on a static corpus of dialogs, recorded from systems using a variety of state tracking models and dialog managers. the challenge task is to re-run state tracking on these dialogs – i.e., to take as input the runtime system logs including the slu results and system output, and to output scores for dialog states. this corpus-based design was chosen because it allows different trackers to be evaluated on the same data, and because a corpus-based task has a much lower barrier to entry for research groups than building an end-to-end dialog system. in practice of course, a state tracker will be used in an end-to-end dialog system, and will drive action selection, thereby affecting the distribution of the dialog data the tracker experiences. in other words, it is known in advance that the distribution in the training data and live data will be mismatched, although the nature and extent of the mis-match are not known. hence, unlike much of supervised learning research, drawing train and test data from the same distribution in offline experiments may overstate performance. so in all three challenges, train/test mis-match was 11 williams, raux, and henderson explicitly created by choosing test data to be from different dialog systems, and, in the case of dstc3, with a different set of slots to be filled. 4.3 data the corpus for dstc1 was produced with dialog systems from three different research groups, here called groups a, b, and c. each group used its own asr, slu, and dialog manager. the dialog strategies across groups varied considerably: for example, groups a and c used a mixed-initiative design, where the system could recognize any concept at any turn, but group b used a directed design, where the system asked for concepts sequentially and could only recognize the concept being queried. groups trialled different system variants over a period of almost 3 years. these variants differed in acoustic and language models, confidence scoring model, state tracking method and parameters, number of supported bus routes, user population, and presence of minor bugs. the fact that these systems were actually deployed and used by the general public presented a number of challenges, most notably acoustic and linguistic conditions made asr significantly more difficult than in more controlled settings. the average length of a dialog in dstc1 is 14.1 turns. more descriptive statistics are given in table 1. dstc1 released 5 train sets and 4 test sets. in all train sets, user speech was transcribed, but only 3 of the 5 train sets were labeled for slu and dialog state correctness. after the evaluation, data inconsistencies were discovered in one of the test sets (test 4, cf. table 1). as result, that test set has been excluded from all results reported in this paper. example dialogs from dstc1 are provided in the appendix. dstc2 and dstc3 use a large corpus of dialogs with various telephone-based dialog systems that was collected using amazon mechanical turk. the dialogs used in the challenges come from 6 conditions; all combinations of one of three possible dialog managers and one of two possible speech recognisers. there are roughly 500 dialogs in each condition, of average length 7.88 turns from 184 unique callers. more descriptive statistics are given in table 1. example dialogs from dstc2 and dstc3 are provided in the appendix. 4.4 dialog state definition and labeling in dstc1, the dialog state consists of a frame of informable slots which are slots provided by the user that describe their goal, such as the bus route and origin bus stop. the slots and approximate number of values for each are shown in table 2. to determine the true dialog state, first each slu hypothesis on each slu n-best list was manually labeled for its correctness. each slu hypothesis could contain values for more than one slot, such as from=downtown,to=airport. in making labeling decisions, the labeler could view the dialog history, and it was possible that zero, one, or more than one slu hypothesis were labeled as correct. if the value for a slot had been provided but no correct value appeared in the slu results, a special value called rest was considered to be correct.2 at every turn, trackers output a scored list of values for every slot, including the special rest value. for evaluation, a dialog state was scored as correct if all of its slots were assigned values which had previously been marked as correct, or rest if no correct values had yet been observed for that slot value. thus, in dstc1, there could be multiple correct dialog states, and the best possible tracker could achieve 100% accuracy. note that, in dstc1, there was no explicit set of slot values, 2. the term rest refers to the remainder, as in “the rest of the unenumerated slu hypotheses”. 12 dialog state tracking overview # dialogs goal changes wer slu f-score dstc1 train1 2,344 46.4% 45.3% train+2 10,619 42.0% test3 2,485 55.1% 38.5% dstc2 train 1,612 40.1% 26.4% 75.7% devel. 506 37.0% 31.9% 71.6% test 1,117 44.5% 28.7% 73.8% dstc3 train4 3,235 41.1% 28.1% 74.3% test 2,275 16.5% 31.5% 78.1% table 1: statistics for the data sets for all three challenges. goal changes is the percentage of dialogs in which the user changed their mind for at least one slot. wer and slu f-score are on the top asr and slu hypotheses respectively. further details of the datasets are given in williams et al. (2013), henderson et al. (2014b), and henderson et al. (2014a). 1this row combines sets train 1a, train 2 and train 3 from dstc1. 2this row combines sets train 1b and 1c from dstc1. in these dialogs, user speech was transcribed, but slu and dialog state correctness were not labeled. 3this row combines sets test 1, test 2, and test 3 from dstc1. in this paper, test 4 has been excluded due to data issues. 4the training set for dstc3 is the combination of train, dev, and test sets from dstc2. slot size bus route 100 date time origin street 500-10,000 origin neighborhood 20-100 origin poi 50-500 destination street 500-10,000 destination neighborhood 20-100 destination poi 500-10,000 table 2: slots used for dstc1 and their approximate number of values. the ranges of values are due to the fact that systems used to collect the dialogs had different internal designs and covered different numbers of street descriptions, neighborhoods and points of interests (poi). 13 williams, raux, and henderson because dialogs were recorded from systems built by different research groups without a shared ontology. slot dstc2 train dstc2 test dstc3 train dstc3 test informable type 1 1 1 3 yes area 5 5 5 15 yes food 91 91 91 28 yes name 113 113 113 163 yes pricerange 3 3 3 4 yes near — — — 52 yes hastv — — — 2 yes hasinternet — — — 2 yes childrenallowed — — — 2 yes addr — — — — no phone — — — — no postcode — — — — no table 3: ontology used in dstc2 and dstc3 for tourist information. counts do not include the special dontcare value. in dstc2-3, the dialog state and labeling procedure was defined somewhat differently. in addition to informable slots, the dialog state in dstc2-3 included 2 other quantities. first, the state included requested slots, which are the slots the user wants to retrieve, such as the phone number, or price range (of a restaurant). second, the state included the search method which indicated if the user wanted to query by providing constraints, providing the name of a restaurant, navigating a results list, etc. the values and sizes of all of the slots in dstc2-3 are given in table 3. in dstc2-3, informable slots could take a special value called dontcare which means the user said they had no preference for that slot – for example, “i don’t mind which type of food.” in addition, dstc2-3 was based on an explicit ontology. because of this, unlike in dstc1, user requests in dstc2-3 were labeled with slot-value pairs taken from the ontology, regardless of the correctness of the slu output. as a result, in dstc2-3, at each turn there was a single correct dialog state, and because the slu often did not contain the correct interpretation, a tracker that took the slu as input could at best achieve less than 100% accuracy. in addition to labeling dialog state, all user speech for all datasets was transcribed, either through crowd-sourcing or professional services. 4.5 tracker output and evaluation metrics each tracker outputs a probability distribution over the set of possible dialog states. the goal is to assign probability 1.0 to the correct state, and 0.0 to other states. in each dialog state hypothesis output by a tracker, every slot is scored, so to be correct, the hypothesis must have perfect precision and recall. based on the ground truth, a number of metrics were computed on each tracker’s output. accuracy measures the percent of turns where the top-ranked hypothesis is correct. this indicates the correctness of the item with the maximum score. l2 measures the l2 distance between the vector of scores, and a vector of zeros with 1 in the position of the correct hypothesis. this indicates the quality of all scores, when the scores are viewed as probabilities. 14 dialog state tracking overview avgp measures the mean score of the first correct hypothesis. this indicates the quality of the score assigned to the correct hypothesis, ignoring the distribution of scores to incorrect hypotheses. mrr measures the mean reciprocal rank of the first correct hypothesis. this indicates the quality of the ordering of the scores (without necessarily treating the scores as probabilities). in addition, two versions of the receiver-operating characteristic (roc) curves were computed, which measure the discrimination of the score for the highest-ranked state hypothesis. roc.v1 computes roc as a fraction of all utterances, and roc.v2 computes fractions of correctly classified utterances. from each of these two curves, four values were extracted. roc.v1.eer and roc.v2.eer give the equal error rate – i.e., the value at which the number of false accepts and number of false rejects are equal. using the v1 curve, roc.v1.ca05, roc.v1.ca10, roc.v1.ca20 give the percent of correctly accepted utterances when the false-accept rate is set to 5%, 10%, and 20%, respectively; and using the v2 curve, roc.v2.ca05, roc.v2.ca10, roc.v2.ca20 give the percent of correctly accepted utterances when the false-accept rate is set to 5%, 10%, and 20%, respectively. in addition, several additional metrics were computed for dstc2-3. neglogp is the mean negative logarithm of the score given to the correct hypothesis, − logpi. sometimes called the negative log likelihood, this is a standard score in machine learning tasks. two metrics, update precision and update accuracy measure the accuracy and precision of updates to the top scoring hypothesis from one turn to the next. for more details, see higashinaka et al. (2004), which finds these metrics to be highly correlated with dialog success in their data. apart from what to measure, when to measure – i.e., which turns to include when computing each metric, must also be defined. for dstc1, a set of 3 schedules were used. schedule1 includes every turn. schedule2 include turns where the target slot is either present on the slu n-best list, or where the target slot is included in a system confirmation action – i.e., where there is some observable new information about the target slot. schedule3 includes only the last turn of a dialog. for dstc2 and dstc3, user goals can change during a dialog, making schedule3 less meaningful. consequently, only schedule1 and schedule2 were used for these challenges. 4.6 baselines all three challenges featured a common simple baseline that mimics standard (non-statistical) approaches commonly used in spoken dialog systems, denoted ‘team0 entry0’. it maintains a single hypothesis for each slot. its value is the slu 1-best with the highest confidence score observed so far, with score equal to that slu item’s confidence score. in addition, dstc1 featured a simpler majority baseline which always selects the rest hypothesis for each turn. two more baselines were provided for dstc2 and dstc3. the focus baseline, denoted ‘team0 entry1’, includes a simple model of changing goal constraints. beliefs are updated for the goal constraint s = v, at turn t, p (s = v), using the rule: p (s = v)t = qtp (s = v)t−1 + slu (s = v)t (2) where 0 ≤ slu(s = v)t ≤ 1 is the slu confidence score for s = v given by the slu in turn t, and qt = ∑ v′ slu(s = v′)t ≤ 1. another baseline tracker, based on the tracker presented in wang and lemon (2013) is included in the evaluation, denoted ‘team0 entry2’. this tracker uses a selection of domain independent rules to update the beliefs, similar to the focus baseline. one rule uses a learnt parameter called 15 williams, raux, and henderson the noise adjustment, to adjust the slu scores. finally, an oracle tracker is included in dstc2 and dstc3 under the label ‘team0 entry3’. this reports the correct label with score 1 for each component of the dialog state, but only if it has been suggested in the dialog so far by the slu. this gives an upper-bound for the performance of a tracker which uses only the slu and its suggested hypotheses. 4.7 participants participation to each challenge was free to any group willing to submit one or more entries by the challenge evaluation deadline. participants were kept anonymous and only referred to in terms of team and entry numbers (e.g. team2.entry4), except when they chose to give their identity in their own published papers. between 7 and 9 research groups participated in each challenge, fielding between 27 and 31 trackers in total, as shown in table 4. # teams # trackers dstc1 9 27 dstc2 9 31 dstc3 7 28 table 4: participation statistics for all three challenges. a subset of teams entered multiple dstcs. 5. challenge entries and results 5.1 which metrics are appropriate for dialog state tracking? as mentioned above, the evaluation in each of the dstcs measured numerous properties of each entry, including accuracy, probability quality, score discrimination, etc. therefore, the question immediately arises which metrics are most appropriate to study. two studies have examined this question. first, in dstc1, metrics were clustered by their correlations with each other, and found to form 4 clusters: one related to correctness with accuracy, mrr, and the three roc.v1.ca metrics; a second related to probability quality with l2 and avgp; a third related to score discrimination with only roc.v1.eer; and a fourth also related to score discrimination with the roc.v2.ca measures (williams et al., 2013). this study suggests that, within each cluster, it is sufficient to choose a single metric, since all metrics within a cluster will empirically yield nearly the same ordering of entries. second, in dstc2, the question of what to measure was posed differently, as “which evaluation metric and schedule would best predict improvement in overall dialog performance?” (lee, 2014). the author uses the data to optimize a reinforcement learning-based dialog manager, then runs a regression analysis to see which metrics are the best predictors of end-to-end dialog performance. l2, avgp, and accuracy are found to be the most predictive. the study also finds that evaluating the joint goal is more predictive than evaluating slots in isolation, and that metrics evaluating only discrimination (e.g., roc.v2) are not good predictors of dialog performance. given these findings, we focus on accuracy and l2 on joint goals throughout the results section. for consistency, we report results on schedule2. we note, however, that all metrics from every tracker in all three dstcs are publicly available for analysis. 16 dialog state tracking overview 5.2 what were the entries, and what was their performance? tables 5-7 show the entries with the highest joint goal accuracy from each team in dstc1-3, using schedule 2. the descriptions in these tables are based on a participant survey included with each of the dstcs, and references are provided if teams identified their entry in a publication. 5.3 what types of errors do the trackers make? as pointed out by smith (2014), it is important to examine the types of errors made by a tracker in order to make improvements. to do this, at each turn, we compare the top dialog state output by each tracker with the true dialog state, and examine each slot. if a slot value is present in both the true and output dialog states and the slot values are equal, we mark the slot as correct. if the slot value is present in both the true and output dialog states and the slot values are not equal, we mark the slot as wrong – i.e., a substitution error. if the slot value is present in the true dialog state but not in the output dialog state, we mark the slot as missing – i.e., a deletion error. finally, if the slot value is not present in the true dialog state but is present in the output dialog state, we mark the slot as extra – i.e., an insertion error. note that, since there are multiple slots in a dialog state, a single turn may have multiple slot-level errors. results are given in figure 3 (p. 22), including performance of the best baselines. these results show that the dominant error type is missing slots. since all error types were scored equally, this result suggests that teams were rather conservative about guessing the slot value when confidence was low. it also suggests that recall in the upstream slu is an important issue. 5.4 how much opportunity for improvement remains? we next compared each tracker to the strongest baseline, and computed the percentage of turns where the tracker was correct and the baseline was not, and the percentage of turns where the baseline was correct and the tracker was not. results are shown in figure 4. even the best trackers – which in total make fewer errors than the baseline – still make some errors that the baselines do not. this implies that there is additional scope for improvement, perhaps through combining multiple trackers using ensemble methods (lee and eskenazi, 2013; sun et al., 2014b; henderson et al., 2014b). 5.5 what is the value beyond slu? figure 5 shows the same analysis for an “slu-based oracle tracker”, again for the best-performing entry for each team. this tracker considers the items on the slun -best list – it is an “oracle” in the sense that, if a slot/value pair appears that corresponds to the user’s goal, it is added to the state with confidence 1.0. in other words, when the user’s goal appears somewhere in the slu n -best list, an oracle in dstc1 would always achieve perfect accuracy. the only errors made by the oracle are omissions of slot/value pairs which have not appeared on any slun -best list. due to the use of the rest meta-value in dstc1, the oracle always achieves 100% accuracy (c.f. section 4.4). therefore only results for dstc2 and dstc3 are shown. figure 5 shows that, for the best trackers, 5% or more of tracker turns outperformed the oracle. these teams also used asr features, which indicates they were successfully using asr results or dialog history to infer new slot/value pairs – i.e., to improve the recall of the existing slu. 17 williams, raux, and henderson unsurprisingly, despite these gains no team was able to achieve a net performance gain over the oracle. 5.6 what is the state-of-the-art? synthesizing the results above, we can summarize the properties of state-of-the-art dialog state trackers: • discriminative models: the strongest entries are consistently discriminative models. although some rule-based systems have achieved noteable performance – for example, team2 entry1 in dstc1, team3 entry1 in dstc2, and team4 entry0 in dstc3 – in no case has a rule-base or generative model achieved best performance in any of the dstcs. • use asr features: the best trackers consistently incorporate low-level asr features. lowlevel asr features such as n-best scores and word confusion network scores provide additional signals that improves precision (williams, 2014). further, incorporating the asr results themselves yields additional dialog state hypotheses that improve recall (section 5.5). • sequential: the best trackers either model dialog directly as a sequence – crfs for team6 entry4 in dstc1 and rnns for team4 in dstc2 and team3 in dstc3 – or otherwise incorporate extensive dialog history features, as in team2 entries 1 and 3 in dstc3, which used hundreds of features from the dialog history. passing only the distribution over hidden states from one turn to the next, as is often done with generative or rule-based approaches, does not perform as well. relying on the distribution over states assumes that state transitions are markovian; this result suggests that states may be encoding insufficient history for the markov assumption to be valid. • capture feature interactions: the best trackers directly model interactions between features. for example, the best trackers in dstc2 and dstc3 directly modeled feature interactions, either via (recurrent) neural networks or collections of decision trees. approaches that do not capture feature interactions, such as log-linear models where each feature of a dialog state affects its score independently – for example, team5 entry1 in dstc1 – were not top finishers. • joint posteriors: in dstc1 and dstc2, the best systems computed a joint posterior over all slots, rather than computing a posterior as a product of the marginals for each slot. the gain observed in dstc1 (table 5b) was particularly large, whereas the gain in dstc2 was present but small (henderson et al., 2014b). this difference is probably due to differences in the domains: bus stops and bus routes requested by real callers in dstc1 were highly correlated, whereas the subjects in dstc2 and dstc3 were given a specification with slot values drawn closer to uniform. 6. practical issues and lessons learned the main effort in organizing the dstc series was the preparation of the data. in dstc1, this task was particularly labor-intensive because there was no ontology of bus stops available, which 18 dialog state tracking overview entry reference description team1 entry1 (henderson et al., 2013) deep neural network team2 entry1 (wang and lemon, 2013) hand-crafted rules based on confidence scores team3 entry2 (zilka et al., 2013) discriminative classifier + hand-crafted transition probabilities team4 entry1 (anonymous) discriminative dynamic bayesian network team5 entry1 (williams, 2013) decision tree team6 entry4 (lee and eskenazi, 2013) discriminative + generative (system combination); unsupervised prior adaptation team7 entry1 (anonymous) discriminatively trained graphical model team8 entry4 (anonymous) support vector machines team9 entry4 (kim et al., 2013) generative plus discriminative re-scoring. (a) dstc1 entries. references cited where teams identified their entry in a published paper. description based on survey collected from participants. features goals joint goals entry asr slu acc. l2 acc. l2 majority class baseline1 x 0.554 0.631 0.166 1.180 1-best baseline1 x 0.564 0.599 0.241 1.078 team1 entry1 x 0.674 0.612 0.349 1.067 team2 entry1 x 0.683 0.532 0.354 1.055 team3 entry2 x 0.650 0.503 0.339 0.964 team4 entry1 x 0.565 0.626 0.278 1.045 team5 entry1 x 0.691 0.503 0.237 1.087 team6 entry4 x 0.765 0.443 0.466 0.890 team7 entry1 x 0.615 0.562 0.283 1.058 team8 entry4 x 0.584 0.592 0.226 1.098 team9 entry4 x 0.724 0.492 0.357 1.024 slu-based oracle x 1.000 0.000 1.000 0.000 (b) dstc1 results. the top performing trackers from each team are selected. results are derived from combining all test sets in the evaluation. in dstc1, none of the entries used the asr output. 1williams et al. (2013). table 5: entries and results of dstc1. 19 williams, raux, and henderson entry reference description team1 entry0 (kim and banchs, 2014) linear crf team3 entry0 (smith, 2014) discourse rules + dialog act bigrams team4 entry2 (henderson et al., 2014d) recurrent neural network team6 entry2 (anonymous) maximum entropy markov model, with dnn output distribution team7 entry4 (sun et al., 2014b) system combination of a deep neural network and maximum entropy model team8 entry1 (lee et al., 2014) hidden information state model + goal change handling model + system-user action pair weighting model team9 entry0 (anonymous) baseline, augmented with priors from a confusion matrix team2 entry2 (williams, 2014) recurrent neural network team4 entry0 (henderson et al., 2014d) recurrent neural network team7 entry0 (sun et al., 2014b) system combination of a deep neural network, maximum entropy model, and rules team2 entry1 (williams, 2014) ranking (lambdamart) team2 entry3 (williams, 2014) ranking (lambdamart) team5 entry4 (anonymous) asr/slu re-ranking (a) dstc2 entries. references cited where teams identified their entry in a published paper. description based on survey collected from participants. features joint goals search method requested entry asr slu acc. l2 acc. l2 acc. l2 1-best baseline1 x 0.619 0.738 0.879 0.209 0.884 0.196 focus baseline1 x 0.719 0.464 0.867 0.210 0.879 0.206 hwu baseline2 x 0.711 0.466 0.897 0.158 0.884 0.201 team1 entry0 x 0.601 0.648 0.904 0.155 0.960 0.073 team3 entry0 x 0.729 0.452 0.878 0.210 0.889 0.188 team4 entry2 x 0.742 0.387 0.922 0.124 0.957 0.069 team6 entry2 x 0.718 0.437 0.871 0.210 0.951 0.085 team7 entry4 x 0.735 0.433 0.910 0.140 0.946 0.089 team8 entry1 x 0.699 0.498 0.899 0.153 0.939 0.101 team9 entry0 x 0.499 0.760 0.857 0.229 0.905 0.149 team2 entry2 x 0.668 0.505 0.944 0.095 0.972 0.043 team4 entry0 x 0.768 0.346 0.940 0.095 0.978 0.035 team7 entry0 x 0.750 0.416 0.936 0.105 0.970 0.056 team2 entry1 x x 0.784 0.735 0.947 0.087 0.957 0.068 team2 entry3 x x 0.771 0.354 0.947 0.087 0.941 0.090 team5 entry4 x x 0.695 0.610 0.927 0.147 0.974 0.053 slu-based oracle1 x 0.850 0.300 0.986 0.028 0.957 0.086 (b) results of dstc2 evaluation. the top performing trackers from each team are selected. results are split by the input features used. 1henderson et al. (2014b), 2wang and lemon (2013). table 6: entries and results of dstc2. 20 dialog state tracking overview entry reference description team1 entry3 (anonymous) rules with parameters inferred from data team6 entry0 (anonymous) generative model trained with cascading gradient descent team7 entry1 (ren et al., 2014a) markovian neural network model team3 entry2 (henderson et al., 2014c) recurrent neural network team5 entry0 (sun et al., 2014a) rules that operate on confidence scores team2 entry0 (anonymous) maximum entropy model team2 entry3 (anonymous) system combination: maximum entropy, crf, rules team3 entry0 (henderson et al., 2014c) recurrent neural network team4 entry0 (kadlec et al., 2014) rules with parameters inferred from data (a) dstc3 entries. references cited where teams identified their entry in a published paper. description based on survey collected from participants. features joint goals search method requested asr slu acc. l2 acc. l2 acc. l2 1-best baseline1 x 0.555 0.860 0.922 0.154 0.778 0.393 focus baseline1 x 0.556 0.750 0.908 0.134 0.761 0.435 hwu baseline2 x 0.575 0.744 0.967 0.062 0.767 0.417 team1 entry3 x 0.561 0.733 0.963 0.097 0.774 0.401 team6 entry0 x 0.507 0.736 0.927 0.120 0.907 0.157 team7 entry1 x 0.576 0.652 0.957 0.116 0.938 0.101 team3 entry2 x 0.616 0.565 0.966 0.061 0.939 0.100 team5 entry0 x 0.610 0.556 0.968 0.091 0.949 0.090 team2 entry0 x x 0.585 0.697 0.965 0.114 0.929 0.121 team2 entry3 x x 0.582 0.639 0.970 0.065 0.938 0.138 team3 entry0 x x 0.646 0.534 0.966 0.061 0.943 0.091 team4 entry0 x x 0.630 0.627 0.853 0.272 0.923 0.136 slu-based oracle1 x 0.717 0.565 0.988 0.02 0.946 0.107 (b) results of dstc3 evaluation. the top performing trackers from each team are selected. results are split by the input features used, with bold indicating the top result in the group. 1henderson et al. (2014a), 2wang and lemon (2013). table 7: entries and results of dstc3. 21 williams, raux, and henderson 0 0.2 0.4 0.6 0.8 1 1.2 -2 -1.8 -1.6 -1.4 -1.2 -1 -0.8 -0.6 -0.4 -0.2 0 a ve ra ge s lo ts p er t u rn : c o rr ec t (d ia m o n d s) a ve ra ge s lo ts /t u rn : w ro n g, e xt ra , m is si n g (b ar s) wrong extra missing correct (a) dstc1 0 0.5 1 1.5 2 2.5 -0.7 -0.6 -0.5 -0.4 -0.3 -0.2 -0.1 0 a ve ra ge s lo ts p er t u rn : c o rr ec t (d ia m o n d s) a ve ra ge s lo ts /t u rn : w ro n g, e xt ra , m is si n g (b ar s) (b) dstc2 2.2 2.25 2.3 2.35 2.4 2.45 2.5 2.55 2.6 2.65 2.7 -0.9 -0.8 -0.7 -0.6 -0.5 -0.4 -0.3 -0.2 -0.1 0 a ve ra ge s lo ts p er t u rn : c o rr ec t (d ia m o n d s) a ve ra ge s lo ts /t u rn : w ro n g, e xt ra , m is si n g (b ar s) (c) dstc3 figure 3: average number of slots in error per turn (bar chart, left axis), and average number of correct slots per turn (black diamonds, right axis) for the best tracker from each team in each of the dstcs. see text for explanation of error types. the left axis shows negative numbers so that the top of each plot indicates ideal performance for both errors (bars) and correctness (diamonds). team ids are not consistent across different dstcs. 22 dialog state tracking overview dstc1 team 1 team 2 team 3 team 4 team 5 team 6 team 7 team 8 team 9 -30 -20 -10 0 10 20 30 dstc2 team 1 team 2 team 3 team 4 team 5 team 6 team 7 team 8 team 9 -30 -20 -10 0 10 20 30 dstc3 team 1 team 2 team 3 team 4 team 5 team 6 team 7 -30 -20 -10 0 10 20 30 figure 4: percentage of all turns where the top tracker from each team did better than the baseline (white bar) or worse than the baseline (black bar) for the joint goal accuracy metric. note that the team ids are not consistent across different dstcs. dstc2 team 1 team 2 team 3 team 4 team 5 team 6 team 7 team 8 team 9 -40 -30 -20 -10 0 dstc3 team 1 team 2 team 3 team 4 team 5 team 6 team 7 -40 -30 -20 -10 0 figure 5: percentage of all turns where the top tracker from each team did better than the slubased oracle (white bar) or worse than the oracle (black bar) for the joint goal accuracy metric. dstc1 is not shown because its design resulted in the oracle always achieving 100% accuracy, so it was not possible to beat the performance of the oracle in dstc1. note that the team ids are not consistent across different dstcs. 23 williams, raux, and henderson required manually labeling each slu hypothesis for correctness. this was done by a mixture of professional transcribers and crowd workers, at a cost of a few thousand dollars. edge cases were difficult for either group, and many utterances needed to be manually labeled by the organizers, often by consulting native pittsburghers or researching pittsburgh geography. a further difficulty in preparing the data for dstc1 was the need to design a dialog act ontology that represented dialog acts produced by three dialog systems from different research groups. by comparison, preparing the data for dstc2-3 required much less work because an explicit ontology simplified labeling, and dialogs were drawn from a single system which included system-side dialog act tags. future challenges might benefit from following the approach in dstc2-3. the dstc organizers decided to continue to make the data freely available after the conclusion of the challenge. this has had unforeseen benefits: first, the dstc data now forms a sort of benchmark for the field, with groups continuing to report results on it after the challenge proper (lee, 2013; ma and fosler-lussier, 2014b; zilka and jurčı́ček, 2015; fix and frezza-buet, 2015). in addition, the dstc1-3 corpora have been used to examine which state tracking evaluation metrics correlate with dialog success (lee, 2014), perform detailed error analyses of state trackers (smith, 2014), and for dialog act classification and slu experimentation (ma and fosler-lussier, 2014a; ferreira et al., 2015). we encourage future challenges to continue this tradition. 7. perspectives and conclusion although dialog state tracking is a crucial problem in spoken dialog systems, until recently it received only sporadic attention. throughout the 1990s, hand-crafted rules were the dominant solution in both research and production systems. in the early 2000s, researchers recognized the need to model uncertainty explicitly and make use of all of the information on the slu n-best list, and proposed several methods, with generative models being most common. yet work was sporadic and different methods were rarely compared: different groups operated their own dialog systems, and there was no standardized dataset and framework for evaluation. the dialog state tracking challenge has introduced the first shared datasets and common evaluation metrics for this problem, and has catalyzed substantial new work into this research problem. in particular, the dstc series has underpinned three broad advances. the first contribution of the dstc series has been to change the dominant approach from generative models to discriminatively trained classifiers. prior to the dstc series, generative models were most common. the dstc series has illustrated the weaknesses in generative models that hindered accuracy, such as the inability to handle a large number of features. in their simplest form, discriminatively trained classifiers take as input a feature vector of fixed size, where the features summarize dialog history up to the current turn. the second contribution of the dstc series has been to enable the development of discriminative sequential models for dialog state tracking. unlike simple classifiers, sequential models take as input a set of features at each turn, avoiding the need to design features that summarize the dialog history. thus, sequential models substantially simplify the feature engineering process, reducing effort. because they properly account for dialog as a temporal process, they also have the potential to improve accuracy, and this has been demonstrated in dstc entries. the third and most recent contribution of the dstc series has been to underpin models which take the asr results as input directly, eschewing the slu entirely. this move further reduced the feature engineering effort – these methods use only primitive asr features and require essentially no feature design at all. by providing direct access to the raw input signal, they also have the 24 dialog state tracking overview potential to provide a further improvement in accuracy, which has also been demonstrated in the dstc series. a key outstanding question for the field is whether improvements in dialog state tracking peformance translate to improvements in end-to-end dialog system performance, such as improved task completion or user satisfaction. two early studies show promising results. first, lee et al. (2014) performed off-line reinforcement learning experiments on the (static) dstc1 corpus, and showed that improved dialog state tracking performance is indeed correlated with improved dialog performance. second, kim et al. (2014) constructed a user simulator, and used simulated dialogs to compare an existing generative tracker with a discriminative tracker that had been shown to yield better dialog state tracking accuracy. they found that the discriminative tracker yielded better endto-end dialog performance. the use of a simulated user and learned dialog policy implies that the distribution of dialog states encountered at test time was different than that encountered in training, and the fact that the discriminative tracker performed well suggests it had not overfit the training data. while these are promising results, the definitive test will be a deployment with real users, which (to our knowledge) has not yet been done. we encourage future dstc instances to consider end-to-end evaluations with real users. the first three instances of the dialog state tracking challenge have tackled human-computer interaction in two domains with dialog systems from multiple research groups, including situations where users’ goals may change and where new slots are present at runtime. given the positive results achieved to date, the next instance of the challenge – dstc4 – will tackle state tracking in human-human conversation (kim et al., 2016). human-human conversation presents a host of new challenges, and addressing these is a step toward the long-standing goal of learning a dialog system from example human-human conversations. 8. acknowledgements the dstc series has benefited from the encouragement and support of many of our fellow researchers. first, the authors thank maxine eskenazi, alan black, and david traum for organizing the 2012 naacl workshop on “future directions and needs in the spoken dialog community: tools and data”, which catalysed the first challenge (williams, 2012b). thanks also to the dstc advisory boards for their guidance and input in all three challenge tasks. the authors thank sigdial for their endorsement, and to sigdial 2013 and 2014, and ieee slt 2014 for hosting special sessions. thanks to at&t, carnegie melon university, microsoft, and cambridge university for providing data and labels, and to microsoft and honda research institute for sponsorship to fund additional labeling. thanks to our co-organizers, deepak ramachandran, alan black, and blaise thomson. finally, a special thank you to all of the dstc participants. appendix 1: sample dialog excerpts below are transcripts of calls in the dstc corpus. the user’s side shows utterances as transcribed by human annotators. dstc1 : group a sys: welcome to the cmu let’s go bus information system. to get help at any time, just say help or press zero. what can i do for you? 25 williams, raux, and henderson usr: 61c evening hours sys: for example, you can say, when is the next 28x from downtown to the airport? or i’d like to go from mckeesport to homestead tomorrow at 10 a.m. usr: i’d like to go [noise] sys: where are you leaving from? usr: from kennywood tomorrow at 6:30 pm sys: leaving from mckee. is this correct? usr: yes sys: okay what is your destination? usr: downtown sys: going to downtown. is this correct? usr: yes sys: right when do you want to travel? usr: 6:30 pm sys: leaving at 6 30 p.m.. is this correct? usr: [noise] yes sys: alright hold on. i’ll look that up. there is a 71d leaving fifth avenue at darragh mckee at 6 40 p.m.. it will arrive at fifth avenue at market at 6 58 p.m.. to get more information about buses related to this trip, you can say, when is the next bus, or, when is the previous bus. to ask about a different trip, you can say, start a new query. if you are finished, you can say goodbye. usr: start a new query sys: okay, let’s start from the beginning. usr: ... dstc1 : group b sys: east pittsburgh bus schedules. say a bus route, like 28x, or say i’m not sure. usr: 61a sys: okay, 61a. to change, say go back. where are you leaving from? usr: downtown sys: okay, downtown. you can always say go back. and where are you going to? usr: east pittsburgh east pittsburgh sys: pittsburgh. is that right? usr: no sys: oh, sorry, where are you going to? usr: east pittsburgh sys: okay, east pittsburgh. do you want times for the next few buses? say yes or no. usr: yes sys: ok, the next few buses. to change, say go back. at 7:45 pm today, there is a 61 a, from fifth avenue and market street, downtown, arriving dynamo alley and beech avenue, in east pittsburgh, at 8:42 pm. say repeat, next, previous, or start over. 26 dialog state tracking overview dstc1 : group c sys: welcome to the lets go! bus information system for pittsburgh. where would you like to leave from? usr: [noise] sys: where do you wanna leave from? usr: [noise] um i’m leaving from forbes avenue and murray avenue sys: i’m sorry but i do not know any stops in the moon neighborhood. the routes i currently cover are the following: 28x, 54c, 56u, 61a, 61b, 61c, 61d and 64. usr: um i need the scheduling information for the 61d sys: where would you like to leave from? usr: forbes and murray sys: where do you wanna go? usr: greenfield sys: when are you going to take that bus? usr: 61d sys: when are you going to take that bus? usr: 61d sys: when would you like to travel? say next bus for the the next bus . usr: next bus sys: the next bus from forbes avenue at murray to greenfield is the 61d, leaving at 8 43 p.m. . it arrives at murray avenue at loretta at 8 52 p.m. . you may now say. next bus, previous bus, restart, or goodbye. usr: [noise] good-bye sys: thank you goodbye dstc2 and dstc3 (training set) sys: which part of town? usr: the north uh area sys: which part of town? usr: a cheap place in the north sys: clown cafe is a cheap restaurant in the north part of town. usr: do you have any others like that? sys: galleria is another cheap restaurant in the north. usr: what is their phone number and address? dstc3 (test set) sys: thank you for calling the cambridge information system. your call will be recorded for research purposes. you may ask for information about a place to eat, such as a restaurant, a pub, or a cafe. how may i help you? usr: i would like to find a pub with an internet connection and a tv sys: the bakers, which has an internet connection, is a great pub with a television. usr: what is the address phone number and price range 27 williams, raux, and henderson sys: the bakers is located on 176 east road and its phone number is 01223304121. the price for main courses ranges from 6 pounds to 8 pounds 49 and the price for a cheapest pint of beer is 2 pounds 25. usr: [noise] sys: the bakers is an outstanding pub. usr: what is the address sys: the bakers is on 176 east road. usr: what is the phone number sys: the phone number of the bakers is 01223304121. usr: what is the price range sys: the bakers is in the cheap price range. usr: thank you good bye references alan black, susanne burger, brian langner, gabriel parent, and maxine eskenazi. spoken dialog challenge 2010. in proc ieee workshop on spoken language technologies (slt), berkeley, california, usa, 2010. dan bohus and eric horvitz. models for multiparty engagement in open-world dialog. in proc sigdial conf on discourse and dialogue, london, usa, 2009. dan bohus and alex rudnicky. ravenclaw: dialogue management using hierarchical task decomposition and an expectation agenda. in proc eurospeech, geneva, switzerland, 2003. dan bohus and alex rudnicky. a ‘k hypotheses + other’ belief updating model. in proc aaai workshop on statistical and empirical approaches for spoken dialogue systems, boston, usa, 2006. trung bui, mannes poel, anton nijholt, and job zwiers. a tractable hybrid ddn-pomdp approach to affective dialogue modeling for probabilistic frame-based dialogue systems. nat. lang. eng., 15(2):273–307, 2009. herbert h clark. using language. cambridge university press, 1996. isbn 9780521567459. philip r cohen and hector j levesque. rational interaction as the basis for communication. in p. r. cohen, j. morgan, and m. e. pollack, editors, intentions in communication, pages 221–255. mit press, cambridge, ma, 1990. david devault. contribution tracking: participating in task-oriented dialogue under uncertainty. phd thesis, rutgers, the state university of new jersey, 2008. david devault and matthew stone. managing ambiguities across utterances in dialogue. in proc workshop on the semantics and pragmatics of dialogue (decalog), trento, italy, 2007. emmanuel ferreira, bassam jabaian, and fabrice lefvre. online adaptative zero-shot learning spoken language understanding using word-embedding. in proc intl conf on acoustics, speech and signal processing (icassp), brisbane, australia, 2015. 28 dialog state tracking overview jeremy fix and herve frezza-buet. yarbus : yet another rule based belief update system. arxiv:1507.06837v1 [cs.cl], 2015. milica gasic and steve young. effective handling of dialogue state in the hidden information state pomdp dialogue manager. acm transactions on speech and language processing, 7, 2011. david heckerman and eric horwitz. inferring informational goals from free-text queries: a bayesian approach. in proc 14th conf on uncertainty in artificial intelligence (uai), pages 230–238, 1998. james henderson and oliver lemon. mixture model pomdps for efficient handling of uncertainty in dialogue management. in proc association for computational linguistics human language technologies (acl-hlt), columbus, ohio, usa, 2008. matthew henderson, blaise thomson, and steve young. deep neural network approach for the dialog state tracking challenge. in proc sigdial conf on discourse and dialogue, metz, france, 2013. matthew henderson, blaise thomson, and jason d. williams. the third dialog state tracking challenge. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014a. matthew henderson, blaise thomson, and jason d williams. the second dialog state tracking challenge. in proc sigdial conf on discourse and dialogue, philadelphia, usa, 2014b. matthew henderson, blaise thomson, and steve young. robust dialog state tracking using delexicalised recurrent neural networks and unsupervised adaptation. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa. ieee, 2014c. matthew henderson, blaise thomson, and steve young. word-based dialog state tracking with recurrent neural networks. in proc sigdial conf on discourse and dialogue, philadelphia, usa, 2014d. matthew henderson, blaise thomson, and steve young. robust dialog state tracking using delexicalised recurrent neural networks and unsupervised adaptation. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014e. ryuichiro higashinaka, mikio nakano, and kiyoaki aikawa. corpus-based discourse understanding in spoken dialogue systems. in proc association for computational linguistics (acl), sapporo, japan, 2003. ryuichiro higashinaka, noboru miyazaki, mikio nakano, and kiyoaki aikawa. evaluating discourse understanding in spoken dialogue systems. acm trans. speech lang. process., 2004. eric horvitz and tim paek. a computational architecture for conversation. in proceedings of the 7th intl conf on user modeling, pages 201–210, banff, canada, 1999. filip jurčı́ček, blaise thomson, and steve young. natural actor and belief critic: reinforcement algorithm for learning parameters of dialogue systems modelled as pomdps. tslp, 7, 2011. 29 williams, raux, and henderson rudolf kadlec, miroslav vodolan, jindrich libovicky, jan macek, and jan kleindienst. knowledgebased dialog state tracking. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014. daejoong kim, jaedeug choi choi, kee-eung kim, jungsu lee, and jinho sohn. engineering statistical dialog state trackers: a case study on dstc. in proc sigdial conf on discourse and dialogue, metz, france, pages 462–466, 2013. dongho kim, matthew henderson, milica gasic, pirros tsiakoulis, and steve young. the use of discriminative belief tracking in pomdp-based dialogue systems. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014. kyungduk kim, cheongjae lee, sangkeun jung, and gary geunbae lee. a frame-based probabilistic framework for spoken dialog management using dialog examples. in proc sigdial workshop on discourse and dialogue, columbus, ohio, usa, 2008. seokhwan kim and rafael e. banchs. sequential labeling for tracking dynamic dialog states. in proc sigdial conf on discourse and dialogue, philadelphia, usa, 2014. seokhwan kim, luis fernando dharo, rafael e. banchs, jason d. williams, and matthew henderson. the fourth dialog state tracking challenge. in proc intl workshop on spoken dialog systems (iwsds), saariselka, finland, 2016. john lafferty, andrew mccallum, and fernando cn pereira. conditional random fields: probabilistic models for segmenting and labeling sequence data. in proc intl conf on machine learning (icml), massachusetts, usa, pages 282–289, 2001. staffan larsson and david traum. information state and dialogue management in the trindi dialogue move engine toolkit. natural language engineering, 5(3/4):323–340, 2000. byung-jun lee, woosang lim, daejoong kim, and kee-eung kim. optimizing generative dialog state tracker via cascading gradient descent. in proc sigdial conf on discourse and dialogue, philadelphia, usa, pages 273–281, 2014. sungjin lee. structured discriminative model for dialog state tracking. in proc sigdial conf on discourse and dialogue, metz, france, 2013. sungjin lee. extrinsic evaluation of dialog state tracking and predictive metrics for dialog policy optimization. in proc sigdial conf on discourse and dialogue, philadelphia, usa, 2014. sungjin lee and maxine eskenazi. recipe for building robust spoken dialog state trackers: dialog state tracking challenge system description. in proc sigdial conf on discourse and dialogue, metz, france, 2013. ryan lowe, nissan pow, iulian serban, and joelle pineau. the ubuntu dialogue corpus: a large dataset for research in unstructured multi-turn dialogue systems. in proc sigdial conf on discourse and dialogue, prague, czech republic, 2015. yi ma and eric fosler-lussier. detecting ’request alternatives’ user dialog acts from dialog context. in proc intl workshop on spoken dialog systems (iwsds), napa, california, usa, 2014a. 30 dialog state tracking overview yi ma and eric fosler-lussier. a discriminative sequence model for dialog state tracking using user goal change detection. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014b. yi ma and eric fosler-lussier. a discriminative sequence model for dialog state tracking using user goal change detection. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014c. yi ma, antoine raux, deepak ramachandran, and rakesh gupta. landmark-based location belief tracking in a spoken dialog system. in proc sigdial conf on discourse and dialogue, seoul, korea, 2012. lidia mangu, eric brill, and andreas stolcke. finding consensus in speech recognition: word error minimization and other applications of confusion networks. computer speech and language, 14 (4):373–400, 2000. neville mehta, rakesh gupta, antoine raux, deepak ramachandran, and stefan krawczyk. probabilistic ontology trees for belief tracking in dialog systems. in proc sigdial conf on discourse and dialogue, tokyo, japan, 2010. helen meng, carmen wai, and roberto pieraccini. the use of belief networks for mixed-initiative dialog modeling. ieee transactions on speech and audio processing, 11(6):757–773, november 2003. angeliki metallinou, dan bohus, and jason d williams. discriminative state tracking for spoken dialog systems. in proc association for computational linguistics (acl), sofia, bulgaria, 2013. tim paek and eric horvitz. conversation as action under uncertainty. in proc conf on uncertainty in artificial intelligence (uai), stanford, california, usa, pages 455–464, 2000. stephen pulman. conversational games, belief revision and bayesian networks. in clin vii: 7th computational linguistics in the netherlands meeting, 1996. antoine raux and yi ma. efficient probabilistic tracking of user goal and dialog history for spoken dialog systems. in proc interspeech conf, florence, italy, pages 801–804. isca, 2011. hang ren, weiqun xu, yan zhang, and yonghong yan. dialog state tracking using conditional random fields. in proc sigdial conf on discourse and dialogue, metz, france, 2013. hang ren, weiqun xu, and yonghong yan. markovian discriminative modeling for cross-domain dialog state tracking. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014a. hang ren, weiqun xu, and yonghong yan. markovian discriminative modeling for dialog state tracking. in proc sigdial conf on discourse and dialogue, philadelphia, usa, pages 327–331, 2014b. alan ritter, colin cherry, and william b. dolan. data-driven response generation in social media. in proc conf on empirical methods in natural language processing (emnlp), edinburgh, united kingdom, 2011. 31 williams, raux, and henderson nicholas roy, joelle pineau, and sebastian thrun. spoken dialog management for robots. in proc association for computational linguistics (acl), hong kong, pages 93–100, 2000. ronnie smith. comparative error analysis of dialog state tracking. in proc sigdial conf on discourse and dialogue, philadelphia, usa, 2014. kai sun, lu chen, su zhu, and kai yu. a generalized rule based tracker for dialogue state tracking. in proc ieee workshop on spoken language technologies (slt), south lake tahoe, nevada, usa, 2014a. kai sun, lu chen, su zhu, and kai yu. the sjtu system for dialog state tracking challenge 2. in proc sigdial conf on discourse and dialogue, philadelphia, usa, pages 318–326, 2014b. umar syed and jason d. williams. using automatically transcribed dialogs to learn user models in a spoken dialog system. in proc association for computational linguistics human language technologies (acl-hlt), columbus, ohio, usa, pages 121–124, 2008. blaise thomson and steve young. bayesian update of dialogue state: a pomdp framework for spoken dialogue systems. computer speech and language, 24(4):562–588, 2010. blaise thomson, filip jurčı́ček, milica gasic, simon keizer, françois mairesse, kai yu, and steve young. parameter learning for pomdp spoken dialogue models. in proc ieee workshop on spoken language technologies (slt), berkeley, california, usa, 2010. zhuoran wang and oliver lemon. a simple and generic belief tracking mechanism for the dialog state tracking challenge: on the believability of observed information. in proc sigdial conf on discourse and dialogue, metz, france, 2013. jason d. williams. using particle filters to track dialogue state. in proc ieee workshop on automatic speech recognition and understanding (asru), kyoto, japan, 2007. jason d. williams. exploiting the asr n-best by tracking multiple dialog state hypotheses. in proc interspeech conf, brisbane, australia, 2008. jason d. williams. incremental partition recombination for efficient tracking of multiple dialogue states. in proc intl conf on acoustics, speech and signal processing (icassp), dallas, texas, usa, 2010. jason d. williams. challenges and opportunities for state tracking in statistical spoken dialog systems: results from two public deployments. ieee journal of selected topics in signal processing, special issue on advances in spoken dialogue systems and mobile interface, 6(8):959–970, 2012a. jason d. williams. a belief tracking challenge task for spoken dialog systems. in naacl hlt 2012 workshop on future directions and needs in the spoken dialog community: tools and data, montreal, canada, 2012b. jason d. williams. challenges and opportunities for state tracking in statistical spoken dialog systems: results from two public deployments. ieee journal of selected topics in signal processing, special issue on advances in spoken dialogue systems and mobile interface, 6(8):959–970, 2012c. 32 dialog state tracking overview jason d. williams. multi-domain learning and generalization in dialog state tracking. in proc sigdial conf on discourse and dialogue, metz, france, 2013. jason d. williams. web-style ranking and slu combination for dialog state tracking. in proc sigdial conf on discourse and dialogue, philadelphia, usa, 2014. jason d. williams and steve young. partially observable markov decision processes for spoken dialog systems. computer speech and language, 21(2):393–422, 2007. jason d. williams, pascal poupart, and steve young. factored partially observable markov decision processes for dialogue management. in proc workshop on knowledge and reasoning in practical dialogue systems, intl joint conf on artificial intelligence (ijcai), edinburgh, united kingdom, 2005. jason d williams, antoine raux, deepak ramachadran, and alan black. the dialog state tracking challenge. in proc sigdial conf on discourse and dialogue, metz, france, august 2013. steve young, jost schatzmann, karl weilhammer, and hui ye. the hidden information state approach to dialog management. in icassp 2007, honolulu, hawaii, 2007. steve young, milica gasic, simon keizer, françois mairesse, jost schatzmann, blaise thomson, and kai yu. the hidden information state model: a practical framework for pomdp-based spoken dialogue management. computer speech and language, 24(2):150–174, 2010. steve young, milica gasic, blaise thomson, and jason d. williams. pomdp-based statistical spoken dialogue systems: a review. proceedings of the ieee, pp(99):1–20, 2013. steve young, catherine breslin, milica gasic, matthew henderson, dongho kim, martin szummer, blaise thomson, pirros tsiakoulis, and eli tzirkel hancock. evaluation of statistical pomdpbased dialogue systems in noisy environment. in proc intl workshop on spoken dialog systems (iwsds), napa, california, usa, 2014. bo zhang, qingsheng cai, jianfeng mao, eric chang, and baining guo. spoken dialogue management as planning and acting under uncertainty. in proc eurospeech, aalborg, denmark, pages 2169–2172, 2001. lukas zilka and filip jurčı́ček. incremental lstm-based dialog state tracker. arxiv:1507.03471v1 [cs.cl], 2015. lukas zilka, david marek, matej korvas, and filip jurčı́ček. comparison of bayesian discriminative and generative models for dialogue state tracking. in proc sigdial conf on discourse and dialogue, metz, france, 2013. victor zue, stephanie seneff, james r glass, joseph polifroni, christine pao, timothy j hazen, and lee hetherington. juplter: a telephone-based conversational interface for weather information. speech and audio processing, ieee transactions on, 8(1):85–96, 2000. 33 d'arceyqar dialogue & discourse 10(2) 56-78 doi: 10.5087/dad.2019.203 ©2018 j. trevor d’arcey, shereen oraby, and jean e. fox tree this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). wait signals predict sarcasm in online debates j. trevor d’arcey jdarcey@ucsc.edu department of psychology university of california, santa cruz shereen oraby soraby@ucsc.edu department of computer science university of california, santa cruz jean e. fox tree foxtree@ucsc.edu department of psychology university of california, santa cruz editor: patrick healey submitted 06/2018; accepted 11/2019; published online 12/2019 abstract we examined the predictive value of wait signals for sarcasm in online debate forums. in study 1, we examined the word frequency of um and uh across six corpora. in general there were far more of these fillers in spoken corpora than written corpora. we also found that the proportion of ums to uhs varied by corpus type. in study 2, we tested whether the inclusion of um or uh at the beginning of online debate forum posts led to higher probability of those posts being classified as sarcastic by amazon mechanical turk workers. we found that posts beginning with these items were twice as likely to be labeled sarcastic. in study 3, we tested fillers and ellipses in the middle of posts. we found that posts including these items were approximately three to five times more likely to be labeled sarcastic. we compared results to other signals like the word obviously and quotation marks. signals that indicate delay in written communication cue readers to non-literal meaning. keywords: sarcasm, irony, fillers, ellipses, online debate, spontaneous communication, written communication, quotations, wait signals 1 introduction non-literal language use is common in communication, both in speech (gibbs, 2000; glucksberg, gildea, & bookin, 1982) and writing (whalen, pexman, & gill 2009; walker, fox tree, anand, & king, 2012). one form of non-literal language is sarcasm, in which people’s intended meaning contrasts with the literal, semantic meaning of their words. people can use sarcasm to mock or to be funny (kreuz, long, & church, 2009), to affirm and modify social relationships (seckman & couch, 1989), and to help a friend save face (jorgensen, 1996). fluency with sarcasm and other forms of humor is an important social skill that predicts a variety of positive social outcomes such as peer reputation in children (masten, 1986) and ability to cope with stress in adults (overholser, wait signals predict sarcasm in online debates 57 1992). creating tools with the ability to recognize sarcasm would have wide-reaching benefits for these groups. yet identification of sarcastic content is notoriously elusive for both people (rockwell, 2000; burgers, van mulken, & schellens, 2011) and machines (reyes & rosso, 2014; riloff, qadir, surve, de silva, gilbert, & huang, 2013; felbo, mislove, søgaard, rahwan, & lehmann, 2017). a number of cues to sarcastic content have been identified, but one that has not been fully explored is the use of fillers like um and uh and spontaneously-written versions of spoken pauses like ellipses. um and uh (er and erm in british english) have been shown to be used by speakers to notify interlocutors of an upcoming delay in speech (smith & clark, 1993), and ellipses typically indicate an omission or pause in writing. we propose that these phenomena are used as wait signals in writing. these wait signals operate to change the pacing at which a text is read, thereby introducing novel pacing in the reader’s mind and potentially delaying delivery for dramatic effect. dramatic pacing can be observed in the following: “the watch-word here is ‘big’: big guitar-licks, big melodic surges, big-hearted words and, erm … big blokes” (from the british national corpus, ck5/3128). wait signals and sarcasm can be observed in the following: “yeah, i'll ....uh keep that in mind dude....trust me!” (from the internet argument corpus, walker et al., 2012). in this report we document the use of wait signals as indicators of sarcasm in writing. 1.1 identifying sarcasm we begin our discussion of sarcasm by noting that it may be futile to try to experimentally differentiate sarcasm from irony, regardless of whether raters are trained to do so (attardo, eisterhold, hay, & poggi, 2003). irony is using language to mean something other than what the words literally express, such as saying “i’ll keep that in mind” while meaning “i most definitely will not keep that in mind.” sarcasm if often thought of as adding a negative connotation to the irony, such as by targeting a victim; for example, by saying “nice hair” to someone with a bad haircut (cambpell & katz, 2012, p. 460). despite these definitions, most researchers are in agreement that the two concepts are difficult to differentiate. to further complicate matters, the word sarcasm may be becoming more prevalent as a replacement for irony (nunberg, 2001), suggesting that to the layperson, the concepts may be interchangeable. when we use the term irony in this work, it is because the research we are referencing uses this term. for all other instances we use the label sarcasm because it is more readily understood (bryant & fox tree, 2002), while acknowledging the fact that researchers generally agree they are separate constructs. to this end, in our present research we were explicit in defining sarcasm for participants as: 1: a sharp and often satirical or ironic utterance designed to be humorous, snarky, or mocking. 2: a mode of satirical wit depending for its effect on bitter, caustic, and often ironic language that is often directed against an individual or a situation. participants were also given examples of statements with and without sarcasm: with sarcasm: "yes, you are 100% correct. criminals would be sure to pay the tax on their illegally owned pistol, just like they pay income tax on drug money. oh, wait they don't pay tax on their drug money. most criminals break the law you see." without sarcasm: "the article said very little about his observations and almost nothing about his methods." our goal was to be as clear as possible, although it is well known that defining these concepts is difficult. 1.1.1 human sarcasm identification perhaps anticipated by the challenges in defining sarcasm, people have a hard time agreeing on whether statements are sarcastic. individual (akimoto & miyazawa, 2017; ivanko, pexman, & olineck, 2004; rockwell & theriot, 2001) and regional (dress, kreuz, link, & caucci, 2008) variations in the conception of sarcasm exacerbate this challenge. for example, political beliefs can affect how satire from late night comedy routines is interpreted (lamarre, landreville, & d’arcey, oraby and fox tree 58 beam, 2009). even if individual, regional, and political backgrounds are held constant, interpretation of sarcasm can vary based on the context presented with the sarcastic utterance. context can make an originally sincere utterance appear sarcastic and vice versa (bryant & fox tree, 2002). although there are challenges to identifying sarcasm, under some circumstances, people can be quite good at detecting it. in a study of tweets originally marked with #sarcasm compared to those which were not, people could correctly identify which were marked sarcastic about 70% of the time when the hashtags were removed (kovaz, kreuz, & riordan, 2013). raters’ misidentification of sarcasm has led researchers to develop explicit, rigorous procedures to achieve high inter-rater reliability on ratings of sarcasm and irony. one such method, the verbal irony procedure (burgers et al., 2011), found high reliability for film reviews by asking raters to engage in a four step process: first to read the entirety of the review and determine the author’s overall stance, second to remove purely descriptive utterances (which, it is assumed, never contain verbal irony), third to remove utterances that have a literal evaluation that fits with the overall stance, and fourth to construct scales of evaluation for the remaining (possibly ironic) utterances in which the literal evaluation of each utterance can be compared to the rater’s perception of the writer’s intent. utterances which contrast are coded as ironic. with this procedure, the authors achieved very strong agreement (97.3%) between two coders (burgers, van mulken, & schellens, 2011). however, this method may not apply as well to less explicitly evaluative texts — film reviews are, by their nature, usually quite expressive. 1.1.2 machine sarcasm identification on the other hand, computational methods to identify sarcasm are improving as deep learning techniques are put into broader use. nonetheless, the best models are still unable to agree with people on what’s sarcastic, whether it is spoken or written. one issue is that the rates of sarcasm in corpora are generally sparse, hovering around 10% (e.g., gibbs, 2000; walker, et al., 2012), leading to more difficulty in measuring classifier success in natural language processing research. in the field of natural language processing, many researchers studying imbalanced classification problems, like sarcasm identification, measure their models’ success using two metrics: the first, recall, is defined as the percentage of sarcastic occurrences that the model correctly identifies. for example, in a set of 1,000 internet posts, 100 may include sarcasm. if the model identifies 80 of the 100 sarcastic posts, its recall is .8. the second measure, precision, is defined as the percentage of model-identified sarcastic occurrences that are actually sarcastic. so, if the aforementioned model correctly identified 80 sarcastic posts, but it also incorrectly labeled another 80 posts as sarcastic, its precision is .5. these are important as separate constructs in imbalanced classification tasks because both a large rate of false positives and a large rate of false negatives are important to understanding the model’s performance. recall and precision are frequently combined into a single measure of performance, f1, defined as the harmonic mean. the harmonic mean is used to avoid rewarding models in which either recall or precision is close to perfect, at the expense of the other (koehrsen, 2018). state-of-the-art classifying models (see, for example, felbo, mislove, søgaard, rahwan, & lehmann, 2017; ghosh & veale, 2016; poria, cambria, hazarika & vij, 2016) achieve a wide range of f1 scores depending on the text being analyzed and the model being used. felbo et al. (2017) developed the deepmoji model for sarcasm detection using publicly released debate forums data from walker et al. (2012) and oraby et al. (2016) using around 1,000 training examples and achieving f1 scores of 0.69 and 0.75, respectively. ghosh and veale (2016) performed sarcasm classification in the twitter domain. they first constructed a dataset of 39k tweets, 18k sarcastic and 21k non-sarcastic. they collected the sarcastic class by using “positive markers of sarcasm” — hashtags such as #sarcastic and #yeahright (ghosh & veale, 2016, p. 3). the non-sarcastic class were tweets lacking these hashtags. training was performed on tweets after the relevant hashtags were removed. ghosh and vale (2016) achieved an f-score of 0.92 on their test set using a convolutional neural network (cnn) model with a long short-term memory layer (lstm) and wait signals predict sarcasm in online debates 59 deep neural network (dnn). they also verified their model on other existing datasets: for riloff et al.’s (2013) test set of 3k tweets, they reported an f1 of 0.88 as compared to the baseline result of 0.51, and for tsur et al.’s (2010) test set of 180 product reviews, they reported an f1 of 0.90 — which was higher than the previously reported best result of 0.83. in addition to these models, poria et al. (2016) reported f1 scores using a deep convolutional neural network on two datasets from ptacek et al. (2014): 0.98 f1 on a balanced dataset of 100k tweets, and 0.95 f1 on an unbalanced dataset (25k sarcastic and 75k non-sarcastic). 1.2 cues for sarcasm in speech and writing although it is challenging to identify, sarcasm is common in speech and writing. in a study of 62 10-minute conversations, for example, gibbs (2000) and two student judges agreed that at least 289 utterances were ironic (8% of the corpus), of which 80 were deemed to be sarcastic. in another study of communication under irony-inducing conditions — describing badly-dressed celebrities and planning meals for a disliked guest — over 10% of the turns produced were ironic (hancock, 2004). of forty dyads who communicated in face-to-face and instant messaging conversations, only one dyad did not produce any ironic utterances (hancock, 2004). irony induction is not necessary to observe sarcasm in writing. in a study of 105 people’s emails to close friends, almost all contained non-literal language, averaging almost three per email (whalen, pexman, & gill 2009). even in academic genres, where one might expect to find more straightforward language use, sarcasm is frequent and widespread (lee, 2006). 1.2.1 cues for sarcasm in speech despite the fact that it is a common phenomenon, cues for spoken sarcasm are not particularly straightforward. pop culture places emphasis on prosody to convey sarcasm, but a close examination of prosodic tone did not show consistent patterns for an ironic tone of voice (bryant & fox tree 2005), although people could differentiate between talk radio utterances that were originally produced as sarcastic and those that were originally produced as sincere (bryant & fox tree 2002). noting that a target utterance contrasts with surrounding talk is another key way people determine what is sarcastic (attardo et al., 2003). some specific cues that can be used in spoken communication include facial cues like smiles, laughter, and slow nods, which are more likely to co-occur with sarcastic utterances (caucci & kreuz, 2012), and air-quotes gestures, which can be used to indicate irony or sarcasm (lampert, 2013). eye gaze towards (caucci & kreuz, 2012) or away (williams, burns, & harmon, 2009) from addressees can also predict sarcasm, perhaps depending on how prepared the sarcastic utterances are (they were not prepared in advance in caucci & kreuz but they were in williams et al., although there may be other differences). in the switchboard corpus of spoken dialogue, 23% of occurrences of the phrase yeah right indicated sarcasm (tepperman, david, & narayanan, 2006). contextual cues and regional differences may influence generation and perception of spoken sarcasm as well: common ground between interlocutors led to more sarcasm (caucci & kreuz, 2002; clark, 1996), and southern u.s. participants’ viewed sarcasm as more hurtful (dress et al., 2008). 1.2.2 cues for sarcasm in writing written sarcasm cannot take advantage of the multimodal auditory, facial, and bodily cues that can help in identifying spoken sarcasm. but written sarcasm does have textual cues that are not available for spoken sarcasm, such as exclamation points and question marks, laughter expressions like lol, emoticons like :), and quotation marks, all of which contribute to irony detection (carvalho, sarmento, silva, & de oliveira, 2009; this study was done in portuguese). as with speaking, contrast is also important with writing. a frowning emoticon matched with an apparently positive message conveys sarcasm (derks, bos, & von grumbkow, 2008b). contrast between positive and negative sentiment can also be used to identify sarcasm (riloff et al., 2013). d’arcey, oraby and fox tree 60 many words and phrases that indicate sarcasm have also been identified. in a corpus of written arguments, the phrases included let’s all (62% sarcastic), i love it when (56% sarcastic), oh really (50%) and i’m shocked/amazed/impressed (42%; oraby, harrison, reed, hernandez, riloff, & walker, 2016). in writing, the inclusion of the word really in online debate posts made the probability that the post-response pair will be perceived as sarcastic approximately double (walker et al., 2012). other markers of sarcasm included you mean and so (walker et al., 2012). in a study of 100 lines from books that were introduced by “said sarcastically,” the presence of interjections such as well and uh were predictive of sarcastic content (caucci & kreuz, 2012). sarcastic lines in books marked by said sarcastically and sarcastic tweets marked by #sarcasm had more positive emotion words than non-marked lines (kovaz et al., 2013). this converges with the idea that sarcasm is more commonly used to evaluate a negative situation with positive affect (clark & gerrig, 1984). 1.2.3 contrast between sarcasm in speech and writing some of the words identified as sarcastic in writing are not generally associated with sarcasm when they are spoken. well, for example, is understood to mean that there is a mismatch between what follows and what’s expected (blakemore, 2002; jucker, 1993), suggesting a less-obvious interpretation (fox tree, 2010), such as a dispreferred interpretation (holtgraves, 2000). this meaning of well aligns well with its potential use in sarcasm, as sarcasm is intended to represent something beyond the literal meaning. when coupled with said sarcastically, and grouped with other interjections, written wells found in books were predictive of sarcasm (caucci & kreuz, 2012). however, when 605 turn-initial wells in spontaneously-written debates were compared to unmarked turns, those marked by well were not more sarcastic (fox tree, 2015). um and uh are also not generally associated with sarcasm when they are spoken. um and uh indicate upcoming delay in speaking (clark & fox tree, 2002; fox tree, 2001). delays marked by um are different from silent pauses without ums; marked delays are more indicative of speech production trouble, and are also associated with lack of comfort with the topic under discussion and dishonesty (fox tree, 2002). but it is important to note that the um itself does not mean that speech production difficulty, discomfort, or deception will necessarily follow. in a study of 35 people’s self-assessment of the meaning of um and uh, for example, no one indicated it meant deception (fox tree, 2007). no one indicated it meant sarcasm either (fox tree, 2007). like in speaking, ums and uhs in writing can also indicate a kind of delay, a need to think, such as to answer a question (fox tree, mayer, & betts, 2011; fox tree, 2015) as in “so, uh... what movie is everyone talking about? i don’t think i’ve seen any previews” (from the internet argument corpus, walker et al., 2012). but also as with speaking, in none of these prior studies were spontaneouslywritten ums or uhs proposed to indicate sarcasm. why do specific n-grams (textual patterns of variable length, here defined as patterns of one or more words) like let’s all, really, and you mean contribute to sarcastic perceptions? one potential explanation is that they are used to call attention to an incongruity, with incongruity being one way nonliteral language is flagged. for example, the incongruity between the body and last lines of a news story suggests that it is satire (rubin et al., 2016). it could be that the incongruity is flagged by the words themselves, if the words are uncommon contextually. for example, slang is not expected in news stories, and has been shown to indicate satire (burfoot & baldwin, 2009). as another example, transforming written quotes to the face-to-face modality as air-quotes gestures sets up an incongruity (because quotes are typically written, not enacted), and it could be this incongruity that flags the sarcasm that air-quotes suggest (cf. lampert, 2013). similarly, using a quote for a single word in writing (e.g., thanks for the “advice”) may also indicate sarcasm because it is a noncanonical usage (in writing, quotes are usually used to indicate a direct report of speech, which is usually more than one word long). that is, a word out of context may cue non-literal meaning. wait signals predict sarcasm in online debates 61 1.3 fillers and ellipses as signals of sarcasm many signals appropriate for speech are not as useful for written communication. a hand gesture may be communicative when directed at a driver who cuts off other drivers, but is less likely to be communicative when directed at a forum troll who belittles an argument. in addition to gestures, another group of signals that may not have as much value among asynchronous writing is requests to wait for production to continue. unlike face-to-face communication with a waiting addressee, spontaneous writing often takes place asynchronously. the composition process is not observed keystroke by keystroke, as we observe speakers phoneme by phoneme. instead, writers generally finish their messages prior to sharing their product. in writing, as opposed to speaking, there are usually fewer costs to lack of timeliness (fox tree, 2015). this asynchrony means that there are not as many reasons to ask addressees to wait or to inform them of an upcoming pause when writing. in a sample of 44 students’ spoken and text conversations, ums and uhs were nine times more common in speaking (fox tree, 2015). but although they were less common, ums and uhs and other signals of time did still occur in writing. we define wait signals in writing as tools used by writers to pace readers’ consumption of information. they include ums and uhs (which can be spelled in numerous ways), ellipses, parentheses that indicate asides, em-dashes, and other markers. in asynchronous writing, wait signals should be expected to be less prevalent than fillers and pauses in speaking. but their lack of prevalence may imbue them with additional significance when they are used. whereas wait signals are not traditionally associated with sarcasm in speech, we propose that wait signals suggest to readers that they take more time with the information that follows them, with the additional time leading to non-literal interpretations. one definition of wit from the oxford english dictionary is, “a natural aptitude for using words and ideas in a quick and inventive way to create humour” (oxford english dictionary, n.d.). our hypothesis is that the use of traditional wait signals in contexts where wait signals have limited use for signaling a pause constitutes one form of wit, or using words in inventive ways. as hearing um at the beginning of a turn leads listeners to consider that the speaker is having production trouble, discomfort with the topic, or is preparing a dishonest answer, so too can reading um suggest that writers are intending something different from what they’ve literally written, such as that they are being sarcastic. it is both the unexpectedness of the wait signal in writing as well as the extra processing suggested by the wait signal that drives the sarcastic interpretation. whalen, pexman, and gill (2009) suggested something similar for nonfiller wait signals: “hyphens, parentheses, and ellipses could be construed as a category of ‘textseparators,’ used to segment portions of the text to assist the reader in detecting those portions that are to be interpreted non-literally” (pp. 275-276). in support of this hypothesis, we note that the highest predictor of sarcasm in a study of a variety of textual cues to sarcasm was oh wait, at 87% (oraby et al., 2016). while oh on its own has been linked to sarcasm and negative emotion in writing (abbott et al., 2011; fox tree, 2015), the predictiveness of oh wait is much higher than oh in combination with other words such as oh really or oh yeah (both 50%, oraby et al., 2016). not all ohs are sarcastic. in speaking they can indicate arrival at revised interpretations or state change (heritage, 1984) which can be used strategically, such as to politely show newsworthiness in comparison to responding with a yes (fox tree & schrock, 1999). the revised interpretations can also be used sarcastically to imply that something is newsworthy when it is not (fox tree, 2015). oh has both attitudinal and cohesive functions, functions that differ from temporally sensitive markers like um and uh which are much more common in synchronous communication (fox tree, 2015). the rate of oh production is similar in spontaneous speech and spontaneous writing (fox tree, 2015). we think the high predictiveness of oh wait comes from both the revision-predictiveness of the oh (which violates expectations of no revision) and the wait-signalling of the wait, , although there may be other factors or interactions; the predictiveness of oh right as a signal of sarcasm was also high, 81% (oraby et al., 2016). d’arcey, oraby and fox tree 62 we predict that the unexpectedness of written fillers plus fillers’ basic meaning of waiting will lead to increased ratings of sarcasm when assessing debate posts that have fillers. similarly, the unexpectedness of ellipses in asynchronous writing (which allows producers time to plan) plus ellipses’ basic meeting of waiting will also increase sarcasm ratings for debate posts with ellipses. importantly, we do not propose that wait signals in writing only cue sarcasm. we propose that when asked to evaluate sarcasm, the unexpectedness of the wait signal in an asynchronous form of communication coupled with the signal to wait will suggest sarcasm. 1.4 current research we tested the hypothesis that contextually unexpected text patterns are cues for sarcasm, and in particular that wait signals — which prompt taking time in assessing upcoming information — are cues to sarcasm. in a corpus comparison, we tested the rate of filler production across a range of spoken and written corpora. although others have observed more fillers in speaking than in writing (e.g. fox tree, 2015), we wanted to confirm this across a wide range of corpora, as well as explore the proportions of ums to uhs across corpora. in studies 1 and 2, we tested the hypothesis that online posts that included a wait signal, defined as fillers or ellipses, would be rated as more sarcastic than online posts without them. because a pause in a spoken conversation has no single written equivalent (periods, ellipses, dashes, em dashes, semicolons, and commas all may qualify), it is challenging to identify whether any particular pause is meant to convey sarcastic meaning. however, ellipses (...) specifically suggest “an omission (as of words) or a pause” (merriamwebster’s online dictionary, n.d.) and so may be most likely to be linked to sarcasm when readers are asked about sarcasm. 1.5 hypotheses we began by verifying that fillers are contextually unexpected text patterns, comparing across spoken and written american and british corpora: h1: there are more fillers in speaking than in writing (study 1). we then tested whether the presence of wait signals in an unexpected context increased sarcasm ratings. we tested fillers at the beginning of turns: h2: the presence of a filler at the beginning of written turns will suggest sarcasm at a higher than baseline rate (study 2). and in the middle of turns: h3: the presence of a filler in the middle of written turns will suggest sarcasm at a higher than baseline rate (study 3). as well as ellipses, which most often occur in the middle of turns: h4: the presence of an ellipsis in the middle of written turns will suggest sarcasm at a higher than baseline rate (study 3). an alternative to the hypothesis that wait signals suggest sarcasm when readers are asked about sarcasm (h2, h3, h4) is that wait signals are a stylistic device to make written language feel more like spoken talk, without any implication for conveying sarcasm. in general, we predict that contextually unexpected patterns can be cues to sarcasm, such as fillers in writing or quotes (air-quotes) in speaking. but beyond contextual inappropriateness, we predicted that cues to wait would enhance ratings of sarcasm, as they suggested deeper thought — with deeper thinking possibly leading to alternative interpretations from the literal words expressed. we compared fillers to words we thought might indicate sarcasm: h5: the words obviously, surely, no doubt, and clearly will suggest sarcasm at a higher than baseline rate. h6: fillers will be more effective at suggesting sarcasm than the words obviously, surely, no doubt, and clearly. as an alternative to h5, obviously, surely, no doubt, and clearly may not suggest sarcasm at higher than baseline rate. as an alternative to h6, fillers may suggest sarcasm less than or to a similar wait signals predict sarcasm in online debates 63 degree as the words obviously, surely, no doubt, and clearly. we also compared fillers to a device we thought might indicate sarcasm: h7: quotation around a single word will suggest sarcasm at a higher than baseline rate. h8: fillers will be more effective at suggesting sarcasm than quotation around a single word. as an alternative to h6, quotation around a single word may not suggest sarcasm at higher than baseline rate. as an alternative to h7, fillers may suggest sarcasm less than or to a similar degree as quotation around a single word. 2 study 1: comparing corpora in study 1 we investigated the frequency of the fillers um and uh across several corpora of both spontaneous communication and planned communication. working with transcripts of spoken conversation can be challenging because across corpora, transcribers generally do not follow the same transcription rules. in addition, it is often not possible to access the original audio conversation to determine how transcription was done. this is especially problematic when examining word frequencies for discourse markers and fillers, as transcription rules vary especially widely on whether to include words like so, i mean, and uh. furthermore, frequency of these markers may show large variance across different contexts. for example, if one corpus is made up of unscripted conversations from radio and television shows (e.g. simpson, briggs, ovens, & swales, 2002), there may be fewer fillers due to television and radio personalities being more likely to have received speech training to avoid using them. likewise, when performing a difficult communication task over the phone (e.g. liu, fox tree, & walker, 2016), one may expect the frequency of fillers to be higher on average because people may be more likely to produce delays, and therefore the fillers that indicate delays (clark & fox tree, 2002). for these reasons, we chose to analyze several different corpora from both spoken and written sources and examine their differences and similarities. 2.1 method word frequencies were calculated from several publicly-available corpora. we include short explanations of and examples from each corpus to contextualize word frequencies in each. the michigan corpus of academic spoken english (micase) is a 1.8-million-word corpus that consists of transcripts from colloquia, dissertation defenses, sections, lectures, office hours, seminars, study groups, and similar academic situations. it has close to 200 hours of transcribed audio recorded at the university of michigan in ann arbor (simpson, briggs, ovens, & swales, 2002). fillers in micase generally appear to be quite spontaneous. for example, “okay. then that’s... that is that’s one thing to figure out um but that’s probably too much work it’s not worth that” (from the michigan corpus of academic spoken english, lel565su064; simpson, briggs, ovens, & swales, 2002). the corpus of contemporary american english (coca) is the largest corpus used in this analysis, at more than 520 million words. the spoken component of the corpus contains over 109 million words transcribed from unscripted tv and radio conversations over 26 years. the four written portions of the corpus are each of similar size to the spoken portion and are taken from fictional works, magazines, newspaper articles, and academic journals (davies, 2008). unfortunately, because audio is no longer available for the coca and transcription methods are unknown, it is difficult to interpret word frequencies for fillers, which frequently are left out of transcription instructions. in the written component, many fillers are within direct quotations, but some exist outside of them, for instance, “samantha, samantha, samantha. what to say about saman-tha? um, okay. this is what i’m going to say about samantha. nothing” (from the corpus of contemporary american english; davies, 2008). the british national corpus (bnc) is 100 million words divided into spoken (10%) and written (90%) components. the written portion samples newspapers, fictional works, academic books, and other texts, while the spoken portion is entirely made up of informal conversations, “recorded d’arcey, oraby and fox tree 64 by volunteers selected from different age, region and social classes in a demographically balanced way” (british national corpus consortium, 2007). an example taken from the written part of the corpus is, “it’s very nice of you to ask me — erm — but i’ve got a lot to do when i get back to england — erm — i’d like to have a lie down … and there’ll be piles of washing … and i haven’t got a hairdresser …” (from the british national corpus; british national corpus consortium, 2007). subtlexus consists of 50 million words of “spoken-like” language of english-language subtitles from television and film (brysbaert & new, 2009). because this corpus generally contains scripted speech, we treat it as written — but we acknowledge that the nature of improvisation and acting may allow for more fillers, as in “pardon me, please. yeah. the, uh ... the man-eating wolves are on a, um ... ski vacation” (from subtlexus; brysbaert & new, 2009). the internet argument corpus consists of about 73 million words of debate posts taken from a popular online debate forum (walker et al., 2012). it should be noted that this corpus is different from the other written corpora we cite in that it consists of work that has not been published in the traditional sense of the word — that is, all the other written corpora draw from newspapers, magazines, books, and other written works that are likely heavily edited prior to being published. an internet forum, on the other hand, has relatively simple mechanisms for revising a work prior to publishing it. in addition, whereas more traditional written works tend to be monologic, the internet argument corpus consists almost entirely of dialogue. these differences are frequently apparent in the corpus, as in “first you lie about what i said, then you quote me to prove it’s a lie. that was, um, helpful of you” (from the internet argument corpus; walker et al., 2012). the artwalk corpus contains about 500,000 words transcribed from mobile cell-phone conversations that took place while participants collaborated on a naturalistically situated referential communication task that also involved a wayfinding component (liu, fox tree, & walker, 2016). although brysbaert and new (2009) suggest that corpora must be 1-3 million words in order to get reliable estimates of high-frequency words, we also included the artwalk corpus in our analysis for two reasons: first, we believe it represents an important type of naturalistic conversation that is not represented by the other corpora. second, brysbaert & new’s operationalization of high-frequency was “over 20 words per million.” because there is a difference of several orders of magnitude between this conceptualization of high-frequency and the frequency of our target words in the artwalk corpus (over 9,000 words per million), we believe that the additional information from artwalk is interesting enough to warrant inclusion. an example from the corpus is, “the the computer for the directions it says we have eight minutes to find each um like we’re finding statues and like art pieces um” (from artwalk; liu, fox tree, & walker, 2016). interpreting raw differences between spoken and written frequencies may be inequitable due to higher lexical diversity in written media. with more words to choose from, the rate of any particular word would be lower. for this reason, we multiplied frequencies originating from written corpora by the constant 2.05, the highest ratio of lexical diversity between spoken and written reported in johansson (2009). because we hypothesize that the frequency of our target words should be lower in written communication, this adjustment creates a more conservative estimate. 2.2 results table 1 reports raw word frequencies for uh, um, er, and erm (british forms of uh and um) across spoken and written corpora, and written frequencies when corrected for the difference in lexical diversity between the two media. with the exception of the coca corpus, the rates of spoken uh, um, er, and erm are many times higher in spoken corpora than written corpora. the average rate of ums and uhs in the spoken micase and artwalk corpora was 9,802 instances per million words compared to 430 instances per million words in the written iac and subtlexus corpora, adjusted for lexical wait signals predict sarcasm in online debates 65 diversity, a difference of 23 to 1. for coca, this relationship was 0.44 to 1: there were more written ums and uhs than spoken. in the spoken corpora, the ratio of ums to uhs was 1.07 to 1 for micase, 1.24 to 1 for artwalk, 0.71 to 1 for the bnc, and 0.47 to 1 for coca. that is, in the american conversational spoken corpora, there were more ums than uhs, and in the british corpus and the american television and radio corpus, there were more uhs than ums. in the written corpora, the ratio of ums to uhs was 0.52 to 1 for coca, 0.94 to 1 for the iac, 0.12 to 1 for subtlexus. the ratio of erms to ers was 0.18 to 1 for the bnc and 0.12 to 1 for the iac. there were no written erms in subtlexus. that is, across all written corpora there were more uhs and ers than ums and erms. 2.3 discussion the difference between filler use in spoken and written corpora was stark, with far more fillers in spoken corpora. coca’s rates were much lower than the other corpora we examined. because we could not ascertain whether ums or uhs were included in transcription instructions for coca’s spoken corpora, we leave it out of our analysis entirely, merely noting that when we took a closer look at the coca’s instances of um and uh that occurred in writing, we found the majority of them to be direct quotations. when excluding coca, the rate of fillers across spoken to written settings was 23 to 1. extrapolating this data to estimate how likely language users are to encounter fillers across settings suggests that for every one filler a person reads in written conversation, a person could be expected to hear, conservatively, 23 fillers in spoken conversations. in american conversational corpora, there were more ums than uhs. in british conversational corpora and american television and radio corpora, there were more ers/uhs than erms/ums. in both american and british written corpora there were more ers/uhs than erms/ums. the strongest outlier in these data was the spoken component of coca. our best explanation of this difference is that although coca’s spoken component is made up of unscripted conversations (such as interviews and debates) from television and radio programs, transcribers of these programs may not have concerned themselves with transcribing fillers. additionally, speakers in television and radio may be more likely to have been trained against the use of fillers in speech. television personalities may also have spoken quickly, which clark and fox tree (2002) showed is inversely correlated with the frequency of fillers. d’arcey, oraby and fox tree 66 the bnc has fewer spoken fillers in comparison to both american english corpora, micase and artwalk. but the bnc also has far fewer written fillers in comparison to the iac and in subtlexus. the bnc displays a more than five-hundred-fold difference in frequencies for er and erm across spoken and written formats, in comparison to the twenty-three fold difference in micase, artwalk, iac, and subtlexus for spoken and written uh and um. one interpretation is that the words er and erm are just more commonly spoken than written in british english. additionally, er and erm may be just far less commonly written in british english. because american english generally uses uh and um, er and erm frequencies are predictably low for corpora featuring american english. nonetheless, er is actually more common in the iac and subtlexus than in the british english corpus. more convincing than the overall differences between spoken and written contexts may be that the highest rate of fillers, lexical diversity-corrected, for written corpora was in subtlex, the corpus most clearly meant to emulate spoken dialogue. in summary, fillers are used more frequently in speech than in writing, although they do occur in both contexts. this result supports hypothesis 1. in most spoken and written corpora investigated here, there were more uhs than ums. the exception was american conversational corpora where there were more ums than uhs. in study 2, we turn to the test of whether fillers in writing indicate sarcasm. 3 study 2: wait signals at the beginning of turns in study 2, we examined whether posts to online debate forums were more likely to be perceived as sarcastic if they began with a filler. previous researchers showed that the probability of mechanical turk workers rating a post-response pair from the internet argument corpus as sarcastic was approximately 12% (walker et. al., 2012). we also examined three other words and a phrase which we thought may also be used to indicate sarcasm: obviously, surely, no doubt, and clearly. if these phenomena indicate sarcasm, the probability that mechanical turk workers will rate posts as sarcastic should be higher than 12%. we tested the beginning of the turns because that is the likely location for fillers in writing (fox tree et al., 2011). 3.1 method in this section, we discuss the participants, materials, and procedure for study 2. 3.1.1 participants mechanical turk workers were required to have an overall approval rate of at least 95%, to have completed at least 500 tasks, and to have an ip address originating from an english-speaking country (including australia, canada, new zealand, great britain, and the united states). workers were paid $0.80 for rating 20 post-response pairs. 3.1.2 materials we used regular expressions to collect a set of stimuli from the internet argument corpus. we then performed additional filtering by limiting our set to posts that had parent posts (contained a quote from a previous post) and contained between 10 and 150 words. for example, “wouldn’t this be contrary to the popular convention that sexuality is innate and orientation is permanent?” is a parent post to, “no, but it would be contrary to your false premise that sexuality is a dichotomy and that orientation is uh... pardon the expression... rigid.” we collected all the post-response pairs in the internet argument corpus which contained one of the following six textual patterns in the response: um (at the beginning of the response), uh (at the beginning of the response), obviously, surely, no doubt, and clearly. the last four of these textual patterns were included as contrasts to um and uh. the stimuli selected were others that had the potential to indicate sarcasm. all of them could be considered to belong to the category of “adjectives or adverbs used to exaggerate or wait signals predict sarcasm in online debates 67 minimize a statement” (hancock, 2004, p. 453), which have been shown to be related to judgements of irony, although the set we selected was not specifically mentioned in hancock (2004). obviously was noted by several researchers as a marker of sarcasm (burgers, van mulken, & schellens, 2012; oraby et al., 2016; whalen et al., 2009). surely and clearly were selected because of their similarity to obviously. no doubt is also similar to obviously, and was part of a sarcastic sample in whalen et al. (2009). er and erm were not used, as they were not frequent enough in the corpus to analyze. we randomly selected 166 168 posts with each textual pattern from the results to be used as our stimuli. 3.1.3 procedure once the final set of posts were selected, we then created a human intelligence task (hit) on amazon mechanical turk that asked workers whether any part of the response contains sarcasm. five workers rated each post-response pair and posts were marked as sarcastic if the majority of workers (three out of five) agreed that the response contained sarcasm. a total of 233 workers accepted the tasks. 3.2 results given that most sarcasm annotation tasks of this type find low reliability on sarcasm ratings, we expected low reliability (e.g., walker et. al., 2012; swanson, lukin, eisenberg, corcoran, & walker, 2017; davidov, tsur, & rappoport 2010). indeed, many studies avoid this problem by focusing on text that includes the explicit #sarcasm or #irony hashtags, common on twitter (peng, lakis, & pan, 2015; liebrecht, kunneman, & van den bosch, 2013; abercrombie & hovy, 2016; riloff et. al., 2013; gonzález-ibánez, muresan, & wacholder, 2011) rather than have humans handannotate text. davidov et al., (2010), when using fleiss’s kappa with two categories (the fewer categories, the higher the 𝜅), achieved a reliability of .34 for amazon reviews and .41 on twitter tweets, indicating fair reliability at best. even when including relatively clear cases of sarcasm, swanson et. al., (2017) found an alpha of only .387. they argue that though this is usually considered low, the subjectivity of sarcasm may mean that it should be treated differently. several researchers argue that these low agreements are a result of the fact that there is wide variation in how people use and understand sarcasm (walker et. al., 2012; swanson et al., 2017; davidov, tsur, & rappoport, 2010). low inter-rater reliability on manual ratings of sarcasm seems to be an unfortunate corollary of studying forms of sarcasm that don’t contain explicit textual flagging. sure enough, our krippendorff’s alpha for the workers’ ratings of sarcasm was 𝛼 = .17. as a part of the process of preparing the rating task, several researchers and assistants tested our hits. not one of our researchers or assistants were able to complete the task in fewer than five minutes. however, 39 of our 250 tasks were completed in under five minutes, 17 in under three minutes, and one in 14 seconds (as reported by mechanical turk). on the opposite end, 44 workers were reported as spending over 30 minutes on the task. although we cannot be certain about the large discrepancy in times, a plausible explanation is that the short duration workers skimmed or ignored the post-response pairs, and the long duration workers took breaks while working on the task. another possibility is that short duration workers considered their answers prior to accepting the task, leaving them with the trivial task of filling them in once they accepted the task, and the work time counter began. it could also be that some participants put little effort into the nontrivial cognitive task of sarcasm comprehension. although inter-rater reliability was practically nonexistent for our participants, there were still reliable differences between stimuli that contained our cues and those that did not. comparing the rate of sarcasm across conditions is still valuable in spite of the high variability in participants’ rating behavior. we ran chi-squared tests of independence to determine if the rates of sarcasm in our post-response pairs were significantly different from the baseline rate of 12% that walker et al. (2012) found using stimuli from the same corpus and an identical hit procedure on mechanical turk. see table 2. post-response pairs starting with the word uh at the beginning of the post had d’arcey, oraby and fox tree 68 higher rates of sarcasm, 𝜒2(1) = 23.1, p < .001, φ = .08, and we can be 95% confident that between 18.1% and 31.3% of iac post-response pairs that start with the word uh would be rated as sarcastic by a majority of mturk workers. post-response pairs starting with the word um also had higher rates of sarcasm, 𝜒2(1) = 43.2 p < .001, φ = .11 and we can be 95% confident that between 22.6% and 36.5% of iac post-response pairs that start with the word um would be rated as sarcastic by a majority of mturk workers. in addition, post-response pairs including the word obviously had higher rates of sarcasm than baseline, 𝜒2(1) = 11.7, p = .001, φ = .06 and we can be 95% confident that between 14.8% and 27.1% of iac post-response pairs that include the word obviously would be rated as sarcastic by a majority of mturk workers, post-response pairs including the word surely had higher rates of sarcasm than baseline, 𝜒2 (1) = 18.6 p < .001, φ = .08 and we can be 95% confident that between 16.9% and 29.8% of iac post-response pairs that include the word surely would be rated as sarcastic by a majority of mturk workers, and post-response pairs including the word clearly had higher rates of sarcasm than baseline, 𝜒2 (1) = 6.2, p < .013, φ = .04 and we can be 95% confident that between 12.6% and 24.3% of iac post-response pairs that include the word clearly would be rated as sarcastic by a majority of mturk workers. post-response pairs including the phrase no doubt did not have higher rates of sarcasm than baseline, 𝜒2 (1) = 1.4, p = .239, φ = .02 and we can be 95% confident that between 9.6% and 20.5% of iac post-response pairs that include the phrase no doubt would be rated as sarcastic by a majority of mturk workers. our two best candidates, um and uh, each displayed rates of sarcasm more than double the baseline rate in the corpus. 3.3 discussion writing um or uh at the beginning of a turn suggested to readers that the writers were being sarcastic at more than twice the base rate of sarcasm for the internet argument corpus. this result supports hypothesis 2. ums and uhs were more predictive than a number of other words tested, including words others identified as related to sarcasm. this result supports hypothesis 6. of the conventional words hypothesized to be related to sarcasm — obviously, surely, no doubt, and clearly — only no doubt did not have a higher rate of sarcasm ratings than baseline. this result partially supports hypothesis 5. one possibility is that only wait signals at the start of turns will affect sarcasm perception. the start of a turn is more noticeable, and indeed others have tested the role of written discourse markers in turn initial position precisely because of the salience of this location (abbott et al., 2011). in study 3, we assessed whether wait signals in the middle of turns also influenced sarcasm perception. wait signals predict sarcasm in online debates 69 4 study 3: wait signals in the middle of turns in study 3, we examined whether posts to online debate forums were more likely to be perceived as sarcastic if they contained a filler or an ellipsis that was not at the beginning of a turn (referred to henceforth as uh (within) and um (within)). we also examined quotation marks encapsulating single words, which we thought may indicate higher sarcasm as a textual equivalent of the airquotes gesture (lampert, 2013). once again, if fillers, ellipses, or quotes around a word indicate sarcasm, mechanical turk workers should rate posts including them as sarcastic at a rate higher than 12%. 4.1 method methods for study 3 were identical to study 2 with two exceptions: first, we used different textual patterns to collect post-response pairs from the internet argument corpus, and second, we recruited a smaller set of workers who already had experience rating sarcastic content in online debate posts, in an attempt to achieve higher inter-annotator agreement. 4.1.1 participants mechanical turk workers were recruited from a pool of workers who had previously been ranked as providing reliable ratings of sarcasm in textual stimuli according to the conditions specified in oraby et. al., (2016). all workers were also required to have an overall approval rate of at least 95%, to have completed at least 500 tasks, and to have an ip address originating from an englishspeaking country (including australia, canada, new zealand, great britain, and the united states). workers were paid $0.80 for rating 20 post-response pairs. 4.1.2 materials as in study 2, we used regular expressions to match specific textual patterns within the internet argument corpus, using the same constraints as before, selecting only posts that had between 10 and 150 words, and included a quote from a previous post. we randomly selected sets of approximately 200 posts per pattern. due to possible limitations of our scripts combined with relative scarcity of cues within the corpus, only 159 uh (within) posts were identified. further, some posts were manually removed because upon inspection the posts fell into categories we did not want to examine and also believed we could computationally control for in future studies. the categories that made a post-response pair eligible for exemption from our set of stimuli were: (1) the post-response pair was a duplicate post-response pair to one that already existed in the set (1 removed), (2) the post-response pair included the matched pattern as part of a url (19 removed), (3) the response did not include fillers or ellipses (15 removed) and (4) the post-response pair was not written in english (1 removed). this process afforded us a set of 154 uh (within) posts, 182 um (within) posts, 184 ellipses posts, and 292 quoted word posts (posts with quoted words were relatively plentiful within the iac, and so were used to fill our quota for the hit). quotations were included as contrasts to um and uh. quotations have been argued to express sarcasm both in speaking, as airquotes (lampert, 2013), and in writing (carvalho et al., 2009). 4.1.3 procedure once the final set of posts were selected, we then created a hit (human intelligence task) on amazon mechanical turk that asked workers whether any part of the response contains sarcasm. five workers rated each post-response pair, and posts were marked as sarcastic if the majority of workers (three out of five) agreed that the response contained sarcasm. a total of nine workers accepted the tasks. d’arcey, oraby and fox tree 70 4.2 results as in study 2, we expect a low worker reliability due to the challenges presented in section 3.2. for study 3, the krippendorff’s alpha for the workers’ ratings of sarcasm was 𝛼 = .32, which was higher than the alpha of .17 in study 2, but still far under common thresholds for fair reliability (krippendorff, 2004). we attribute this to the higher quality of our workers. the krippendorff’s alpha was also higher than the alpha of .22 for the original sample of 3,158 post-response pairs. we attribute this boost in reliability to the fact that our set of posts contained strong predictors of sarcasm (um, uh, or ellipses), so the sarcasm should be less ambiguous, leading people to agree on it in more cases. this explanation fits with the higher alpha (.39) achieved in another study in which researchers included unambiguous sarcastic/non-sarcastic post-response pairs (swanson, lukin, eisenberg, corcoran, & walker, 2017). despite the low reliability between workers, comparing sarcasm ratings across conditions is still valuable. while reliability detects the agreement of workers, the following analyses detect differences between overall proportion of post-response pairs rated as sarcastic. we ran chi squared tests of independence to determine if the rates of sarcasm in our post-response pairs were significantly different from the baseline rate of 12% that walker et al. (2012) found using stimuli from the same corpus and an identical hit procedure. see table 3. post-response pairs including the word uh had higher rates of sarcasm, 𝜒2(1) = 363.7, p < .001, φ = .33, and we can be 95% confident that between 60.1% and 74.9% of iac post-response pairs that include the word uh would be rated as sarcastic by a majority of mturk workers. post-response pairs including the word um also had higher rates of sarcasm, 𝜒2 (1) = 309.8, p < .001, φ = .30 and we can be 95% confident that between 52.2% and 66.4% of iac post-response pairs that include the word um would be rated as sarcastic by a majority of mturk workers. and post-response pairs including ellipses had higher rates of sarcasm, 𝜒2 (1) = 122.6, p < .001, φ = .19, and we can be 95% confident that between 33.6% and 47.9% of the iac post-response pairs that include ellipses would be rated as sarcastic by a majority of mturk workers. in addition, post-response pairs including quotations had higher rates of sarcasm, 𝜒2 (1) = 195.2, p < .001, φ = .24, and we can be 95% confident that between 36.5% and 47.8% of the iac post-response pairs that include quotations would be rated as sarcastic by a majority of mturk workers. 4.3 discussion writing um, uh, or using ellipses in the middle of a turn suggested to readers that the writers were being sarcastic at 4.5 times the base rate of sarcasm for the internet argument corpus. these results support hypotheses 3 and 4. the lowest rate found, for ellipses, was still over triple the baseline rate of sarcasm in the corpus. this result supports hypotheses 7 and 8. as observed in prior work (carvalho et al., 2009; lampert, 2013), quotations around single words were also indicative of sarcasm, at over triple the baseline rate. in speech, people reported that they used direct quotation (which would be expressed with quotation marks if written) to be wait signals predict sarcasm in online debates 71 entertaining (blackwell & fox tree, 2012). direct quotes were also used to report thoughts (fox tree & tomlinson, 2008), and were often accompanied by vocal and bodily demonstrations, such as moving the mouth and neck up as if howling and using a howling voice to imitate a dog’s behavior (blackwell, perlman, & fox tree, 2015). being entertaining, reporting thoughts, and adding vocal and bodily information might all contribute to a relationship between spoken quotation and sarcasm. this relationship may be alluded to with written quotations. written quotations may also act as text-separators to highlight non-literal content (whalen et al., 2009, p. 275). as text-separators they could potentially contribute to the pacing of information consumption which in turn may be suggestive of sarcasm, as proposed for um, uh, and ellipses. 5 general discussion sarcasm has been studied across speech and writing and in synchronous and asynchronous settings. in the current series of studies, we documented the prevalence of fillers across spoken and written corpora, and tested how likely fillers were to suggest sarcasm when they fell at the beginning of turns and in the middle of turns. we predicted that fillers would be more frequent in speaking than in writing (h1). we also predicted that fillers would suggest sarcasm because they are uncommon in writing (h2, h3), and that fillers and ellipses would suggest sarcasm because they communicate the need to wait in a context where waiting isn’t necessary (h4). we thought that seeing elements typical of spoken speech in writing (fillers and a written representation of spoken pauses, ellipses) would suggest to readers a need to think more deeply about what the writer was communicating. we also predicted that obviously, surely, clearly, and no doubt would indicate the presence of sarcasm (h5), although at lower rates than fillers and ellipses (h6), and that quotation around a single word would indicate sarcasm (h7), also at lower rates than fillers and ellipses (h8), because fillers and ellipses indicate delay, further prompting readers to consider the material they were reading more deeply. we predicted that the search for deeper meaning would lead listeners to consider writers’ non-literal goals in using fillers and ellipses, such as the production of sarcasm. across corpora, we demonstrated that fillers are more common in speech than in writing. we also documented differences in preferences for er/uh versus erm/um across corpora and settings, with more ums in conversational american english corpora, and more ers/uhs in a conversational british corpus, a television/radio american corpus, and all written corpora. in two studies, we showed that fillers and ellipses reliably indicated sarcasm to readers and to a greater extent than other sarcasm-predicting devices. these data are indicative of a broader pattern in which writers use incongruent language to express sarcasm (clark & gerrig, 1984; kovaz et al., 2013; rubin et al., 2016). another way to view incongruence is by noting language that is used more frequently in one medium than another. because fillers and pauses are not necessary in asynchronous written communication, such as online forums, the use of fillers and pauses are contextually inappropriate — their use contrasts with their medium. we suggest that this contrast is what enables um, uh, ellipses, and likely other phenomena, to cue non-literal meaning, including sarcasm. one next step with this research is to examine whether these patterns exist for more types of computer-mediated communication, including testing varying levels of synchronicity. more synchronous communicative methods, like instant messages and text-messages, could be expected to have lower rates of sarcasm co-occurrence with fillers, because fillers would be more likely to be used in these media for their time-noting functions; for example, communicators using text chat might write um to indicate that their response will be delayed (although text chat programs that contain blinking ellipses to indicate that the respondent is writing may obviate the need for an um). more asynchronous communicative methods, like reddit and other message boards, could be expected to have higher rates of sarcasm co-occurrence with fillers, much like we observed here with an online debate forum. other phenomena that might be explored include other wait signals, such as words like wait or hang on, characters like em-dashes, typographic behavior such as spacing d’arcey, oraby and fox tree 72 out words, like t h i s, or elongations like thiiiiis. like fillers, elongations have been shown to indicate upcoming problems in speech (fox tree & clark, 1997). their interpretation in writing may be similar to fillers as well. another next step with this research is to assess the role of wait signals on other kinds of inferences readers can make beyond sarcasm. for example, wait signals may influence assessments politeness or evasion. hearing ums at the beginning of the spoken turns affected listeners’ judgements of speakers’ production difficulty, comfort, and honesty (fox tree, 2002). but this wasn’t because the ums were a leaked symptom of difficulty, discomfort, or dishonesty. the assessments were a product of the ums’ basic meaning — announcing an upcoming delay — coupled with the requirements of the task. listeners were, in essence, asking themselves why a speaker would need to delay right then, and, if thinking about honesty, conclude that the speaker needed time to come up with a deceptive answer. finally, it would be interesting to determine whether there is a difference in how sarcastically ums are viewed as opposed to uhs. in spoken communication ums lead to longer pauses than uhs on average (clark & fox tree, 2002). since wait signals in online forums seem to be able to cue sarcasm through their inappropriateness, a longer pause could be viewed as more inappropriate than a short one. it is possible, therefore, that the longer pauses implied by um lead to higher ratings of sarcasm than uh. although our data trends toward ums at the beginnings of posts being rated as sarcastic more frequently than uhs, it trends in the opposite direction for ums and uhs in the middle of posts. it’s also important to note that frequency does not necessarily imply intensity, so it would be interesting to use a more nuanced rating of sarcasm to check for differences between wait signals. as we achieve a better understanding of mechanisms and cues of non-literal language, both in writing and in speech, we will be able to train computers to flag sarcasm in language more and more accurately, leading to better tools to assist those who could benefit from them. one group who could benefit are people with hearing difficulties. deaf children show slower development in recognizing sarcasm than hearing children, and although native sign language signers’ performance appears to eventually catch up to hearing persons’, late signers (those from hearing families) continue to show reduced performance in sarcasm recognition into adulthood (o’reilly & peterson 2014). another group who could benefit are people on the autism spectrum, who struggle with sarcasm identification (kaland, møller-nielsen, callesen, mortensen, gottlieb, & smith, 2002; peterson 2012). a third group who could benefit are second language learners, who also struggle with sarcasm identification, such as identifying satirical news (prichard & rucynksi, 2019). and there are others who could benefit, such as anyone who has trouble differentiating satirical news reports from real stories, or who has trouble recognizing the satire behind a deadpan delivery. technology with the ability to recognize sarcastic intent could inform readers of non-literal meaning as they read, bridging gaps in communication. acknowledgements this work was supported in part by nsf iis-1302668, and by the federico and rena perlino award for research related to deafness. we thank marilyn walker for her support of this project. we thank brian schwarzmann for assistance with amazon mechanical turk. data cited herein have been extracted from the british national corpus online service, managed by oxford university computing services on behalf of the bnc consortium. all rights in the texts cited are reserved. wait signals predict sarcasm in online debates 73 references abbott, r., walker, m., anand, p., fox tree, j. e., bowmani, r., & king, j. (2011, june). how can you say such things?!?: recognizing disagreement in informal political argument. in proceedings of the workshop on languages in social media (pp. 2-11). association for computational linguistics. abercrombie, g., & hovy, d. (2016). putting sarcasm detection into context: the effects of class imbalance and manual labelling on supervised machine classification of twitter conversations. in proceedings of the acl 2016 student research workshop (pp. 107-113). akimoto, y., & miyazawa, s. (2017). individual differences in irony use depend on context. journal of language and social psychology, 36(6), 675-693. https://doi.org/10.1177/0261927x17706937 attardo, s., eisterhold, j., hay, j., & poggi, i. (2003). multimodal markers of irony and sarcasm. humor, 16(2), 243-260. https://doi.org/10.1515/humr.2003.012 blackwell, n. & fox tree, j. e. (2012). social factors affect quotative choice. journal of pragmatics, 44, 1150-1162. https://doi.org/10.1016/j.pragma.2012.05.001 blackwell, n. l., perlman, m., & fox tree, j. e. (2015). quotation as multimodal construction. journal of pragmatics, 81, 1-7. https://doi.org/10.1016/j.pragma.2015.03.004 blakemore, d. 2002. relevance and linguistic meaning: the semantics and pragmatics of discourse markers. cambridge: cambridge university press. british national corpus consortium. (2007). british national corpus version 3 (bnc xml edition). distributed by oxford university computing services on behalf of the bnc consortium. retrieved february 13, 2012. bryant, g. & fox tree, j. e. (2002). recognizing verbal irony in spontaneous speech. metaphor and symbol, 17(2), 99-117. https://doi.org/10.1207/s15327868ms1702_2 bryant, g. a., & fox tree, j. e. (2005). is there an ironic tone of voice? language and speech, 48(3), 257-277. https://doi.org/10.1177/00238309050480030101 brysbaert, m., & new, b. (2009). moving beyond kučera and francis: a critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for american english. behavior research methods, 41(4), 977-990. https://doi.org/10.3758/brm.41.4.977 burfoot, c., & baldwin, t. (2009, august). automatic satire detection: are you having a laugh?. in proceedings of the acl-ijcnlp 2009 conference (pp. 161-164). association for computational linguistics. burgers, c., van mulken, m., & schellens, p. j. (2011). finding irony: an introduction of the verbal irony procedure (vip). metaphor and symbol, 26(3), 186-205. https://doi.org/10.1080/10926488.2011.583194 d’arcey, oraby and fox tree 74 burgers, c., van mulken, m., & schellens, p. j. (2012). type of evaluation and marking of irony: the role of perceived complexity and comprehension. journal of pragmatics, 44(3), 231-242. campbell, j. d., & katz, a. n. (2012). are there necessary conditions for inducing a sense of sarcastic irony? discourse processes, 49(6), 459-480. carberry, s. (1989). a pragmatics-based approach to ellipsis resolution. computational linguistics, 15(2), 75-96. carvalho, p., sarmento, l., silva, m. j., & de oliveira, e. (2009, november). clues for detecting irony in user-generated contents: oh...!! it's “so easy” ;-). in proceedings of the 1st international cikm workshop on topic-sentiment analysis for mass opinion (pp. 53-56). acm. caucci, g. m., & kreuz, r. j. (2012). social and paralinguistic cues to sarcasm. humor, 25(1), 1– 22. https://doi.org/10.1515/humor-2012-0001 clark, h. h. (1996). using language. 1996. cambridge university press: cambridge. clark, h. h., & fox tree, j. e. (2002). using uh and um in spontaneous speaking. cognition, 84(1), 73-111. https://doi.org/10.1016/s0010-0277(02)00017-3 clark, h. h., & gerrig, r. j. (1984). on the pretense theory of irony. http://dx.doi.org/10.1037/0096-3445.113.1.121 davidov, d., tsur, o., & rappoport, a. (2010, july). semi-supervised recognition of sarcastic sentences in twitter and amazon. in proceedings of the fourteenth conference on computational natural language learning (pp. 107-116). association for computational linguistics. davies, m. (2008-). the corpus of contemporary american english (coca): 520 million words, 1990-present. available online at http://corpus.byu.edu/coca/. davies, m. (2009). the 385+ million word corpus of contemporary american english (1990– 2008+): design, architecture, and linguistic insights. international journal of corpus linguistics, 14(2), 159-190. http://dx.doi.org/10.1075/ijcl.14.2.02dav derks, d., bos, a. e., & von grumbkow, j. (2008). emoticons and online message interpretation. social science computer review, 26(3), 379-388. https://doi.org/10.1177/0894439307311611 dress, m. l., kreuz, r. j., link, k. e., & caucci, g. m. (2008). regional variation in the use of sarcasm. journal of language and social psychology, 27(1), 71-85. https://doi.org/10.1177/0261927x07309512 ellipsis, (n.d.). in merriam-webster’s online dictionary. retrieved from https://www.merriamwebster.com/dictionary/ellipsis felbo, b., mislove, a., søgaard, a., rahwan, i., & lehmann, s. (2017). using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm. arxiv preprint arxiv:1708.00524. wait signals predict sarcasm in online debates 75 fox tree, j. e. (2001). listeners’ uses of um and uh in speech comprehension. memory and cognition, 29(2), 320-326. fox tree, j. e. (2002). interpreting pauses and ums at turn exchanges. discourse processes, 34(1), 37-55. https://doi.org/10.1207/s15326950dp3401_2 fox tree, j. e. (2007). folk notions of um and uh, you know, and like. text & talk, 27(3), 297-314. https://doi.org/10.1515/text.2007.012 fox tree, j. e. (2010) discourse markers across speakers and settings. language and linguistics compass, 3(1), 1–13. https://doi.org/10.1111/j.1749-818x.2010.00195.x fox tree, j. e. (2015). discourse markers in writing. discourse studies, 17(1), 64-82. fox tree, j. e., & clark, h. h. (1997). pronouncing “the” as “thee” to signal problems in speaking. cognition, 62(2), 151-167. https://doi.org/10.1016/s0010-0277(96)00781-0 fox tree, j. e., mayer, s. a., & betts, t. e. (2011). grounding in instant messaging. journal of educational computing research, 45(4), 455-475. https://doi.org/10.2190/ec.45.4.e fox tree, j. e., & schrock, j. c. (1999). discourse markers in spontaneous speech: oh what a difference an oh makes. journal of memory and language, 40(2), 280-295. https://doi.org/10.1006/jmla.1998.2613 fox tree, j. e. & tomlinson, j. m., jr. (2008). the rise of like in spontaneous quotations. discourse processes, 45, 85-102. https://doi.org/10.1080/01638530701739280 ghosh, a., & veale, t. (2016). fracking sarcasm using neural network. in proceedings of the 7th workshop on computational approaches to subjectivity, sentiment and social media analysis (pp. 161-169). gibbs, r. w. (2000). irony in talk among friends. metaphor and symbol, 15(1-2), 5-27. https://doi.org/10.1080/10926488.2000.9678862 gonzález-ibánez, r., muresan, s., & wacholder, n. (2011, june). identifying sarcasm in twitter: a closer look. in proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies: short papers-volume 2 (pp. 581-586). association for computational linguistics. ivanko, s. l., pexman, p. m., & olineck, k. m. (2004). how sarcastic are you? individual differences and verbal irony. journal of language and social psychology, 23(3), 244-271. https://doi.org/10.1177/0261927x04266809 hancock, j. t. (2004). verbal irony use in face-to-face and computer-mediated conversations. journal of language and social psychology, 23(4), 447-463. https://doi.org/10.1177/0261927x04269587 heritage, j., & atkinson, j. m. (1984). structures of social action. studies in conversation analysis. d’arcey, oraby and fox tree 76 holtgraves, t. (2000). preference organization and reply comprehension. discourse processes, 30(2), 87–106. https://doi.org/10.1207/s15326950dp3002_01 johansson, v. (2009). lexical diversity and lexical density in speech and writing: a developmental perspective. working papers in linguistics, 53, 61-79. jorgensen, j. (1996). the functions of sarcastic irony in speech. journal of pragmatics, 26(5), 613634. https://doi.org/10.1016/0378-2166(95)00067-4 jucker, a. h. (1993). the discourse marker ‘well’: a relevance theoretical account. journal of pragmatics, 19(5), 435–52. https://doi.org/10.1016/0378-2166(93)90004-9 kaland, n., møller-nielsen, a., callesen, k., mortensen, e. l., gottlieb, d., & smith, l. (2002). a new ‘advanced’ test of theory of mind: evidence from children and adolescents with asperger syndrome. journal of child psychology and psychiatry, 43(4), 517-528. https://doi.org/10.1111/1469-7610.00042 koehrsen, w. beyond accuracy: precision and recall choosing the right metrics for classification tasks [web log message]. retrieved from https://towardsdatascience.com/beyond-accuracy-precision-and-recall-3da06bea9f6c kovaz, d., kreuz, r. j., & riordan, m. a. (2013). distinguishing sarcasm from literal language: evidence from books and blogging. discourse processes, 50(8), 598-615. kreuz, r. j., long, d. l., & church, m. b. (1991). on being ironic: pragmatic and mnemonic implications. metaphor and symbol, 6(3), 149-162. https://doi.org/10.1080/0163853x.2013.849525 krippendorff, k. (2004). reliability in content analysis: some common misconceptions and recommendations. human communication research, 30(3), 411-433. lamarre, h. l., landreville, k. d., & beam, m. a. (2009). the irony of satire: political ideology and the motivation to see what you want to see in the colbert report. the international journal of press/politics, 14(2), 212-231. https://doi.org/10.1177/1940161208330904 lampert, m. (2013). say, be like, quote (unquote), and the air-quotes: interactive quotatives and their multimodal implications. english today, 29(04), 45-56. https://doi.org/10.1017/s026607841300045x lee, d. (2006). humor in spoken academic discourse. nucb journal of language culture and communication, 8(1), 49-68. leech, g., & rayson, p. (2014). word frequencies in written and spoken english: based on the british national corpus. routledge. liebrecht, c. c., kunneman, f. a., & van den bosch, a. p. j. (2013). the perfect solution for detecting sarcasm in tweets# not. in proceedings of the 4th workshop on computational approaches to subjectivity, sentiment and social media analysis, pp. 29-37. http://hdl.handle.net/2066/112949 wait signals predict sarcasm in online debates 77 liu, k., fox tree, j. e., & walker, l. (2016). coordinating communication in the wild: the artwalk dialogue corpus of pedestrian navigation and mobile referential communication. in proceedings of the tenth international conference on language resources and evaluation (pp. 3159-3166). masten, a. s. (1986). humor and competence in school-aged children. child development, 461473. http://dx.doi.org/10.2307/1130601 nunberg, g. (2001). the way we talk now: commentaries on language and culture from npr's" fresh air". houghton mifflin harcourt. o’reilly, k., peterson, c. c., & wellman, h. m. (2014). sarcasm and advanced theory of mind understanding in children and adults with prelingual deafness. developmental psychology, 50(7), 1862. https://doi.org/10.1037/a0036654 oraby, s., harrison, v., reed, l., hernandez, e., riloff, e., & walker, m. (2016, september). creating and characterizing a diverse corpus of sarcasm in dialogue. in 17th annual meeting of the special interest group on discourse and dialogue (p. 31). overholser, j. c. (1992). sense of humor when coping with life stress. personality and individual differences, 13(7), 799-804. https://doi.org/10.1016/0191-8869(92)90053-r peng, c., lakis, m., & pan, j.w. (2015). detecting sarcasm in text an obvious solution to a trivial problem. foundations and trends in information retrieval. peterson, c. c., wellman, h. m., & slaughter, v. (2012). the mind behind the message: advancing theory-of-mind scales for typically developing children, and those with deafness, autism, or asperger syndrome. child development, 83(2), 469-485. https://doi.org/10.1111/j.14678624.2011.01728.x poria, s., cambria, e., hazarika, d., & vij, p. (2016). a deeper look into sarcastic tweets using deep convolutional neural networks. arxiv preprint arxiv:1610.08815. prichard, c., & rucynski jr, j. (2019). second language learners’ ability to detect satirical news and the effect of humor competency training. tesol journal, 10(1), e00366. reyes pérez, a.; rosso, p. (2014). on the difficulty of automatically detecting irony: beyond a simple case of negation. knowledge and information systems. 40(3):595-614. doi:10.1007/s10115-013-0652-8. reyes, a., rosso, p., & veale, t. (2013). a multidimensional approach for detecting irony in twitter. language resources and evaluation, 47(1), 239–268. https://doi.org/10.1007/s10579012-9196-x riloff, e., qadir, a., surve, p., de silva, l., gilbert, n., & huang, r. (2013). sarcasm as contrast between a positive sentiment and negative situation. in emnlp 2013 2013 conference on empirical methods in natural language processing, proceedings of the conference (pp. 704714). association for computational linguistics (acl). d’arcey, oraby and fox tree 78 rockwell, p., & theriot, e. m. (2001). culture, gender, and gender mix in encoders of sarcasm: a self-assessment analysis. communication research reports, 18(1), 44-52. https://doi.org/10.1080/08824090109384781 simpson, r. c., briggs, s. l., ovens, j., & swales, j. m. (2002). the michigan corpus of academic spoken english. ann arbor, mi: the regents of the university of michigan. smith, v. l., & clark, h.h. (1993). on the course of answering questions. journal of memory & language, 32, 25-38. swanson, r., lukin, s. m., eisenberg, l., corcoran, t., & walker, m. a. (2017). getting reliable annotations for sarcasm in online dialogues. in lrec (pp. 4250-4257). https://arxiv.org/abs/1709.01042 walker, m. a., fox tree, j. e., anand, p., abbott, r., & king, j. (2012). a corpus for research on deliberation and debate. in lrec (pp. 812-817). whalen, j. m., pexman, p. m., & gill, a. j. (2009). “should be fun—not!”: incidence and marking of nonliteral language in e-mail. journal of language and social psychology. https://doi.org/10.1177/0261927x09335253 williams, j. a., burns, e. l., & harmon, e. a. (2009). insincere utterances and gaze: eye contact during sarcastic statements. perceptual and motor skills, 108(2), 565-572. wit. (2019). in oxforddictionaries.com. retrieved from https://en.oxforddictionaries.com/definition/wit paper.dvi dialogue & discourse 8(2) (2017) 206–224 doi: 10.5087/dad.2017.209 user-adaptive a posteriori restoration for incorrectly segmented utterances in spoken dialogue systems∗ kazunori komatani komatani@sanken.osaka-u.ac.jp the institute of scientific and industrial research, osaka university naoki hotta graduate school of engineering, nagoya university satoshi sato ssato@nuee.nagoya-u.ac.jp graduate school of engineering, nagoya university mikio nakano nakano@jp.honda-ri.com honda research institute japan co., ltd. editor: amanda stent submitted 05/2016; accepted 07/2017; published online 12/2017 abstract ideally, the users of spoken dialogue systems should be able to speak at their own tempo. thus, the systems needs to interpret utterances from various users correctly, even when the utterances contain pauses. in response to this issue, we propose an approach based on a posteriori restoration for incorrectly segmented utterances. a crucial part of this approach is to determine whether restoration is required. we use a classification-based approach, adapted to each user. we focus on each user’s dialogue tempo, which can be obtained during the dialogue, and determine the correlation between each user’s tempo and the appropriate thresholds for classification. a linear regression function used to convert the tempos into thresholds is also derived. experimental results show that the proposed user adaptation approach applied to two restoration classification methods, thresholding and decision trees, improves classification accuracies by 3.0% and 7.4%, respectively, in cross validation. keywords: spoken dialogue system, turn taking, user adaptation, a posteriori restoration 1. introduction to make spoken dialogue systems more user-friendly, users need to be able to speak at their own tempo. even though not all users speak fluently, i.e., some speak slowly and put some pauses within their utterances, conventional systems basically assume that users say every utterance with no pauses. systems need to handle utterances both by novice users who speak slowly and by experienced users who want the system to reply quickly. we propose a method for spoken dialogue systems to interpret user utterances adaptively in terms of utterance units. we utilize an approach based on our a posteriori restoration method for incorrectly segmented utterances (komatani et al., 2014). the proposed method allows the system ∗. this paper is a modified and extended version of our earlier report (komatani et al., 2015). c©2017 kazunori komatani, naoki hotta, satoshi sato, and mikio nakano this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). user adaptive a posteriori restoration for incorrectly segmented utterances to respond quickly while also interpreting utterance fragments by concatenating them when a user speaks with pauses. another approach to this issue is to change the parameters of voice activity detection (vad) adaptively for each user during dialogue. however, automatic speech recognition (asr) engines with such adaptive control are uncommon, and implementing an online-adaptive vad module is difficult. our a posteriori restoration approach does not require changing asr engines but uses vad and asr results as they are, and it restores the interpretation of user utterances. a crucial part of our approach is to determine whether or not two utterance fragments that are close in time need to be interpreted together, i.e., whether these are two different utterances or a single utterance incorrectly segmented by vad. if the fragments need to be interpreted separately, the system normally responds to each on the basis of their asr results. if they need to be interpreted together, the system immediately stops its response to the first fragment, concatenates the two segments, and then interprets the combined segments. misclassification of whether restoration is needed causes erroneous system responses. if the system incorrectly classifies the restoration as not being required, its response may become erroneous because the original user utterance is interrupted in the middle. if the system classifies the restoration as being required even though the restoration is actually not, the system takes an unnecessarily long time before it starts responding, and its response tends to be erroneous because an unnecessary part is attached to the actual utterance. in this work, we present a method for adapting the restoration classifier to each user, and we show through experiments that user adaptation improves classification accuracy. we focus on the tempo of each user and use it to adapt the classifier. because the temporal interval between two utterance fragments is an important parameter in the classifier (komatani et al., 2014), we adapt its threshold to user behaviors obtained during the dialogue. 2. related work the aim of our restoration is to resolve a problem with utterance units. spoken dialogue systems that do not consider the problem naively assume that these three items are always in agreement: 1. results of voice activity detection (vad) 2. units of dialogue acts (das) 3. units of user turns the second item is used to update dialogue states in the system, and the third determines when the system starts responding. however, these three do not agree in cases of real user utterances. because the first item is the input received by the system, existing studies on the problem can be categorized into two: handling disagreements between 1 and 2 and between 1 and 3. the disagreement between 1 and 2 was tackled by nakano et al. (1999) and bell et al. (2001). the purpose of these studies was to understand fragmented utterances incrementally and to determine whether or not each fragment forms a da with another. these approaches treated each word as a unit of the fragments. methods for incremental understanding (skantze and hjalmarsson, 2010; baumann and schlangen, 2011; selfridge et al., 2011; traum et al., 2012) treat each input frame as a unit of the fragments and determined whether or not each input frame forms a da with previous ones. the disagreement 207 komatani, hotta, sato, and nakano � �� �� �� �� �� �� � � � � � �������� ������� ����� �� �� ���!��� "�� ��#��$�% &���� ���'�����( ���'�� !������#� )�*��� figure 1: distributions of intervals between utterance fragments (komatani et al., 2014) between 1 and 3 was tackled by sato et al. (2002), ferrer et al. (2003), and kitaoka et al. (2005), who determined the timing at which a system needs to start responding. raux and eskenazi also tackled this problem by changing the thresholds for silence duration in a vad module (raux and eskenazi, 2008) and by incorporating partial asr results into their model (raux and eskenazi, 2009). some current systems allow the system to start responding even during user utterances (paetzel et al., 2015; zhao et al., 2015). we define a correct utterance unit as a dialogue act (da). our a posteriori restoration framework primarily considers the disagreement between 1 and 2, by restoring fragmented asr results. unlike previous studies, such as (nakano et al., 1999) and (bell et al., 2001), which are based on syntactic parsing, our method assumes that the da boundaries are a subset of the vad boundaries. the disagreement between 1 and 3 is partially considered in our framework, which can decide whether to respond to a fragmented utterance. our problem setting relates in part to the one tackled by (kitaoka et al., 2005; raux and eskenazi, 2008, 2009), in which the system determines more precise timing to respond. thus, our approach can be used together with these methods to improve turn-taking. user-adaptive spoken dialogue systems can be categorized into two types: adaptation of the system’s output and adaptation during input interpretation. several previous studies have adapted the system output to users by changing behaviors such as the content of system utterances (jokinen and kanto, 2004) and dialogue management (komatani et al., 2005), pause and gaze duration (dohsaka et al., 2010), and backchannel behaviors (e.g., head nods or short vocalization like “uhhuh”) (de kok et al., 2013). however, only a few studies have been conducted on adaptation during input interpretation. as one example, paek and chickering (2007) exploited the history of a user’s commands and adapted the system’s asr language model to the user. our adaptation is concerned with input interpretation, i.e., in which unit the system interprets user utterances, thereby aiming at improving utterance understanding accuracy. as far as we know, this is the first method of user adaptation proposed for the restoration of utterance units. 208 user adaptive a posteriori restoration for incorrectly segmented utterances +,-. /00-.123-4 56789:7 ;9<= <7> ? :7;7 >@a 6a a@ab7acd ef g120 0h ih 0h j1ihk1lmhn-lo1-lp1q1 ,010rh2cs tuvw xyyvwz[\v ]^uyv_ wvu`a[uv bcdefc gehi hcj klh.0 m1/,nop wvuxqy fcgc jrs bs srstcs o]u wvuxqy vwxyz{x| v}x~x ��� �� ����x�| �� �z[y ya �a ya �z�z uyzy�a[�� 0rn]�a��[� z �z^ ya �z�z ]yzy�a[� ]� � ��vw_�[zyv�� �v\zxuv y�v [v�y xuvw xyyvwz[\v �u �vyv\yv��� figure 2: example of incorrect utterance segmentation due to short pauses 3. posteriori restoration for incorrectly segmented utterances most conventional dialog systems wait to respond until they have received an entire user utterance, to which a vad result is assumed to correspond. that is, such a system responds when every vad result (possibly an utterance segment) arrives on the basis of the asr and understanding result for it. however, user utterances are not always obtained as single vad results, and they are often incorrectly segmented into several fragments. figure 1 shows the distributions of intervals between pairs of vad results classified by whether or not they were originally single utterances (i.e., incorrectly segmented) (komatani et al., 2014). the total number of pairs in this data set was 376. the pairs were manually annotated. we can see that many pairs were incorrectly segmented by vad and that their intervals tended to be shorter than others. these pairs need to be restored to obtain correct interpretations of the original utterances. we explain how conventional dialog systems respond to a pair of utterance fragments, denoted as first and second fragments. when a user utterance is incorrectly segmented by vad, a problem called a false cut-in occurs, and the system incorrectly starts speaking even during a user utterance. because the pair is close in time, we can think of two types of basic system behavior concerning whether or not barge-in is allowed. • system prohibits barge-in: the system responds to the first fragment and ignores the second fragment of the user utterance while it is speaking. • system allows barge-in: the system terminates its response to the first fragment and responds on the basis of an asr result for the second fragment only. the latter strategy is better for the false cut-in problem because its damage is smaller. this is because the system allowing barge-in immediately stops speaking when the user starts speaking, while the system prohibiting barge-in continues to speak, thereby requiring the user to wait until the whole system utterance ends. an example of incorrect segmentation is illustrated in fig. 2. the system in this example adopts the latter strategy: allowing the user to barge in. the user intends to say “i want to go to nagoyadome-mae-yada station,” but a short pause occurs between “nagoya-dome-mae” and “yada.” if 209 komatani, hotta, sato, and nakano ���� �������� ����� ��� � �¡�¢��� � ��£ ¤�¡��¡� ¥¦§ �¢��� §�¤��  � ���¨ �¡�¢���  ©��¡ �� ª  �¡���� �¡�¢��� � «�� ª¡��§�� �¡� ��� �� ¡�¬ª�¡�£ §�� �¡� ��� �� ��  ¡�¬ª�¡�£ §�¤��  � ����� ��� �£ ª  �¡���� ­�� ¡��¤����  � ��¡�  �¡�¢��� ® «�¡�  �¡�¢���  ¦����£ �¡�¢��� ¯� �¡°�� figure 3: overview of a posteriori restoration (komatani et al., 2014) the system correctly recognizes the user utterance and understands that “nagoya-dome-mae-yada” is a station name, it should respond with “i will show you the way to nagoya-dome-mae-yada station.” however, the asr result is incorrect because asr is performed separately for two fragments: “nagoya-dome-mae” and “yada eki ni ikitai.” consequently, the system responds with “showing a way to yada station” (yada is another station) on the basis of the incorrect asr result for the second fragment, because the response to the first fragment is immediately terminated when the second fragment starts. the resulting system response does not match the user’s request. an outline of our a posteriori restoration process (komatani et al., 2014) is shown in fig. 3. this approach is based on the latter type of conventional dialog system, which allows barge-in and thus aims at alleviating the false cut-in problem. when a pair of utterance fragments is close in time, this process is invoked when the second fragment starts. the process consists of two steps: 1. classify whether or not a pair of utterance fragments resulted from an incorrect segmentation, i.e., whether or not restoration is required. 2. restore the utterance if it has been incorrectly segmented; i.e., concatenate the two utterance fragments and perform asr again for the resulting segment. because the system allows barge-in, it also tries to reduce the damage caused by the false cut-ins by terminating its response to the first fragment and waiting until the second fragment ends to avoid the system speaking during a user utterance. if restoration is required, the system performs asr again after concatenating the fragments. the system then responds on the basis of the asr result for the concatenated fragments. in particular, asr for incorrectly segmented fragments tends to be erroneous when a long word is incorrectly segmented in it. restoration improves utterance understanding accuracy. over 153 segment pairs that were close in time and that contained at least one keyword, accuracy was 75% (114/153) when restoration was performed, while accuracy was 28% (43/153) and 20% (31/153) when the asr result of either a first or second segment was only used for understanding. (komatani et al., 2014). if restoration is not required, i.e., the fragments are deemed to be two utterances, the system responds normally; that is, it generates responses based on the asr results for each fragment. 210 user adaptive a posteriori restoration for incorrectly segmented utterances user who speaks quickly set threshold for interval shorter classified correctly as restoration is “required” show me ellora caves. nazca and pampas … … … … de jumana set threshold for interval longer user who speaks slowly classified correctly as restoration is “not required” next, please. figure 4: examples of user-adapted restoration when the system starts the second response also depends on the system’s configuration: after the first response finishes or immediately by stopping the first response. a trade-off exists between the occurrences of erroneous system responses caused by incorrect segmentation and response delay resulting from restoration. our approach gives weight to preventing erroneous responses at the cost of a small delay in system responses. we have endeavored to reduce damage stemming from the delay, by producing fillers such as “well” to prevent unnatural silences (komatani et al., 2014) and by improving implementation to reduce the delay itself. even if the performance of end-pointing is further improved, a mechanism for restoring incorrect segmentation is required because such errors are unavoidable. our a posteriori restoration of incorrect segmentation works with a normal vad, specifically, that of the julius asr engine (lee and kawahara, 2009), which is based on the amplitude of the speech signal and the zero crossing rate (benyassine et al., 1997). the proposed method relies on neither a special asr engine nor a specific end-pointing method; that is, it is complementary to other approaches. integration with more sophisticated vad and end-pointing methods remains for future work. for example, sakai et al. (2007) showed that vad performance improved by using gmms. prosodic features are also known to be helpful in determining ends of turns (ohsuga et al., 2005; edlund et al., 2005). 4. obtaining appropriate thresholds from dialogue tempos we analyze each user’s dialogue tempo, which is a parameter representing the way a user speaks quantitatively. in improving the classification accuracy of whether or not an utterance fragment pair is required to be restored, the threshold for the temporal interval between fragments plays a dominant role. we assume that appropriate thresholds depend on the way each user speaks. examples of how the thresholds need to change are given in fig. 4. brisk users are assumed to speak fluently with shorter pauses. thus, the threshold needs to be set shorter, thereby avoiding unnecessary restoration and subsequent late responses. we should point out here that such users often repeat their utterances when the system’s response is not quick enough because they think the system has not heard their utterance, and this causes utterance collision (funakoshi et al., 2010). 211 komatani, hotta, sato, and nakano ±²³´ µ¶²·³¸ ¹º»¼½¾¿ àá â½ã¾¿ä åæçá ºãè æ¿º½áéãê ëèãì åá àæ ºíèã 㺻ä îº ëº½ ¿éïè éãë ð½èáæàºãáä ñ»à澿àãò íé½áè ñ»à澿àãò íé½áè óôéõòèöàã× figure 5: examples of switching pauses in contrast, “slow” users often speak with long pauses during their utterances. in this case, the threshold needs to be set longer, thereby enabling the system to restore utterances even when longer pauses exist in a single utterance. 4.1 definition of dialogue tempo we define dialogue tempo as a quantitative parameter showing how each user speaks. specifically, it is defined as the average duration of switching pauses, which are the times between when a system finishes speaking and when a user starts speaking, as depicted in fig. 5. we calculate this for each user from the beginning of the dialogue. the duration of a switching pause can be negative, as when the user barges in, i.e., when the user starts speaking during a system utterance. although the speaking rate can also be used for defining the tempo, we here use the duration of switching pauses. although the tempo is calculated for each dialogue, it can be accumulated per user across dialogues when a user id can be obtained (e.g., mobile phones and in-car interfaces). 4.2 appropriate threshold for interval we set appropriate thresholds for each user to investigate the relationship of the threshold to the dialogue tempo. by “appropriate,” we mean that the threshold can classify whether or not restoration is required with high accuracy. restoration for an utterance fragment pair is classified as “required” if its interval is shorter than the threshold and “not required” otherwise. here, we set the threshold as a discriminant plane (point) of a support vector machine (svm) whose only feature is the temporal interval between two utterance fragments. a reference label was manually given, i.e., whether or not the restoration is required. we used the smo module in weka (version 3.6.9) (hall et al., 2009) as an svm implementation. the parameters were set to the default values, e.g., its kernel function was polynomial. the svm can set a discriminant plane that maximizes distances between classes. if a user’s training data did not contain both positive and negative labels, we set fixed values for the threshold as exceptions: large enough (2.00 seconds) when all the labels in the training data were “restoration is not required” and small enough (0.00 seconds) when they were all “restoration is required.” 4.3 target data our target data were collected by our spoken dialog system that introduces world heritage sites (nakano et al., 2011). in total, speech data from 35 participants was recorded. each participant 212 user adaptive a posteriori restoration for incorrectly segmented utterances 0.00 0.50 1.00 1.50 2.00 0.00 1.00 2.00 3.00 a p p ro p ri a te t h re sh o ld s fo r te m p o ra l in te rv a ls [ se c. ] dialogue tempos (average switching pauses) [sec.] figure 6: correlation between appropriate thresholds and dialogue tempos per participant engaged in four 8-minute dialogues. participants were not given any special instructions prior to or during the dialogues. we used the data of only 26 of the 35 participants because nine participants did not have sufficient utterance pairs. specifically, we used the data only of participants who had more than six utterance pairs whose temporal intervals were close in time (less than 2.00 seconds), and whose fragments were longer than 0.80 seconds. this was because our target was originally a single utterance and because we regarded pairs that are very short and that have intervals greater than 2.00 seconds as not being such an utterance (komatani et al., 2014). we obtained 3099 utterances from the 26 participants. the data included 390 utterance pairs that satisfy the aforementioned conditions to possibly be a single utterance. we manually assigned the labels of whether or not the pair is a single utterance in accordance with the procedure in (hotta et al., 2014). because 240 pairs were originally single utterances and 150 pairs were not, the classification accuracy of a majority baseline is 61.5%. 4.4 correlation between dialogue tempos and appropriate thresholds we investigated the correlation between dialogue tempos and the appropriate thresholds for restoration for each of the 26 participants. all 3099 utterances were used to obtain the dialogue tempos of each participant. we excluded outliers: specifically, utterances whose switching pauses are less than −3.5 seconds and more than 6 seconds were excluded because such large values simply indicate that the participant was thinking deeply. these values were determined experimentally. figure 6 plots the data of the 26 participants, where the x-axis denotes the dialogue tempos and the y-axis denotes the appropriate thresholds, both in seconds. first, we can see that the appropriate thresholds (y-axis values) varied depending on the participant. this shows the distributions of within-speaker pauses are different across users, thereby indicating that adaptation of the thresholds can improve the classification accuracy. second, the correlation coefficient between the two values was 0.63. the linear regression function is derived as y = 0.88x− 0.43. (1) 213 komatani, hotta, sato, and nakano input: pair of utterance fragments output: restoration is required / not required temporal interval switching pauses after dialogue starts dialogue tempo linear regression function threshold adapted threshold thresholding figure 7: user-adapted classification using thresholding this function is used in the next section for estimating thresholds from the dialogue tempos per user. here, dialogue tempos are used as an approximation to represent how fluently a user speaks. the dialogue tempos and the thresholds correspond to between-speaker and within-speaker pauses, respectively. these two kinds of pauses are different, but we have empirically shown that these two pauses were correlated when analyzing them per user. because the dialogue tempos can be obtained during dialogues, this empirical correlation is used to estimate the thresholds for the classification. we can also see that the appropriate thresholds were generally smaller than the dialogue tempos, which can be seen in eq. (1): the coefficient of the first-order term was less than one (0.88), and the constant term was also negative (−0.43). because the thresholds reflect the upper bounds of durations of within-speaker pauses, the actual distributions of within-speaker pauses were shorter than the thresholds. this tendency is different from the results shown by heldner and edlund (2010), whose finding was that within-speaker pauses are generally longer than switching pauses in humanhuman conversation corpora. one of the reasons may be that our data were collected from humancomputer dialogues, where users tend to speak more formally than in human-human conversation. more detailed analysis on such differences when users speak with humans and systems remains for future work. 5. adapting classifiers for restoration to users we present our investigation on whether or not the correlation between dialogue tempos and appropriate thresholds is helpful. the correlation is used to derive the user-adaptive threshold from the user’s dialogue tempo, thereby improving classification accuracy for whether or not restoration is required. first, the system derives threshold values for the temporal intervals from the user’s dialogue tempos by using the linear regression function. it then adapts the classifier to each user. we examined user adaptation for two classification methods: thresholding and decision trees. 214 user adaptive a posteriori restoration for incorrectly segmented utterances 5.1 thresholding thresholding is the simplest method for classification on the basis of the temporal interval between utterance fragments. we first discuss the effectiveness of user adaptation with this method. the process flow of thresholding with user adaptation is shown in fig. 7. its input is a pair of utterance fragments (and the temporal interval between them). the system calculates the user’s dialogue tempo on the basis of switching pauses from when the dialogue starts and derives a threshold value corresponding to the tempo using the linear regression function. the system then classifies whether or not the restoration is required using the adapted threshold. the restoration for a pair is classified as “required” if its temporal interval is shorter than the adapted threshold and is “not required” otherwise. 5.2 decision trees we also use decision trees, which are a more complicated classifier than thresholding. we show that user adaptation is also effective in this case. the process flow of the decision trees with user adaptation is depicted in fig. 8. in addition to the temporal interval between a pair of utterance fragments, we use four features that were shown to be effective in our previous report (hotta et al., 2014): an average confidence score of the first fragment, noise detection results using a gaussian mixture model (gmm), the f0 range of the first fragment, and the maximum loudness in the first fragment. user adaptation is performed by converting the temporal interval only; the other four features are not changed here. the interval is converted in both the training and classification phases in the decision tree learning. instead of adapting the threshold to each user, we convert its feature values. this is because, in the normal training phase of decision tree learning, a single decision tree having fixed thresholds across different users is obtained. our approach is to convert the feature values relatively for the interval in accordance with each user, thereby enabling the system to classify adaptively to users with a constant threshold. specifically, we use ratios between the threshold values of the target user and the average one of all users. the feature value is converted using eq. (2), where we denote an original interval i by a user j as iij and its converted value as îij : îij = iij × t0 tj , (2) where tj is a threshold value adapted to user j, which is obtained from the user’s dialogue tempo and the linear regression function, and t0 is a constant set to 0.519 seconds1 , which was the average interval of all users. our aim with this conversion is as follows. the correlation depicted in fig. 6 shows that thresholds need to be smaller for users with quicker dialogue tempos. this conversion makes the feature values of the interval relatively larger for such users (having smaller tj) by multiplying the ratio t0/tj . this is equivalent to setting a relatively smaller threshold even though fixed and common thresholds are actually used in decision tree learning. 1. the constant t0 is not necessarily required. it is used to adjust the feature values within a similar value range as the original ones for making it easy to interpret learned trees. 215 komatani, hotta, sato, and nakano øùúûüý þßàá âã ûüüäáßùåä ãáßæçäùüè éûüúûüý êäèüâáßüàâù àè êäëûàáäì í îâü áäëûàáäì ïäçúâáßð àùüäáñßð òóàüåôàùæ úßûèäè ßãüäá ìàßðâæûä èüßáüè õâùñäáèàâù ö÷ëø ùú ûàßðâæûä üäçúâ üàùäßá áäæáäèèàâù ãûùåüàâù ïôáäèôâðì éüôäá ý ãäßüûáäè öþÿÿ áäèûðü( äüåøú ûäåàèàâù üáää figure 8: user-adapted classification using a decision tree 6. experimental evaluation we investigated whether user adaptation contributes to improving the classification accuracy. we also experimentally checked the upper limit and convergence speed of the proposed adaptation by comparing the accuracy with its batch version, in which all utterance data from a target user are assumed to be available. 6.1 performance of user adaptation we report classification accuracy for the two methods, thresholding and decision trees, discussed in section 5. experiments were conducted under two conditions: closed test and cross validation. figure 9 illustrates the training and test data usage for the experimental evaluation. the background models to be trained were the linear regression function (i.e., coefficients a and b) and decision trees for the thresholding and decision tree methods, respectively. during the training of decision trees, the linear regression function (fixed here, as explained later) and each participant’s dialogue tempo were used to convert feature values. the training data for the background models were those of all 26 participants in the closed test and those of 25 participants excluding one for the test and adaptation phase in cross validation. that is, cross validation was performed by leaving one participant out. the data contained utterances of participants in chronological order. more specifically, for pairs of vad results close in time, their features including temporal intervals and their reference labels (i.e., whether or not they should be restored) were given. durations of switching pauses were also recorded for every utterance, by which a dialogue tempo was calculated at each point in time. the adaptation and test phase was performed by loading each participant’s data chronologically. each user’s dialogue tempo was calculated by using the duration of switching pauses from the beginning of the dialogue until the target utterance2. adaptation was based on the target user’s 2. in sections 6.2 and 6.3, batch adaptation is used; all switching pause durations of the target user are loaded at once in advance, meaning the user’s dialogue tempo is already known from the beginning of the dialogue. 216 user adaptive a posteriori restoration for incorrectly segmented utterances �������� ���� �� � ���� �� ����������� �������� ����� �������� � ��� ���� ����� �!�� !"#����$ %"�&&�%���� '�) *+ ,#� �&��� -��! *�%.��"/�� 0"��# 1234��25 6 �7� �234��25 8%%/��%9 :2��4�2; <7� ;2=>2�� ����; ?��@7=42 �2>�7 �� 2�� �7��� �� ��>2 a<7� �5������7�b c d7�525 � �7�7@7=���@@e fg�@��2 �7�5���7�h c i@@ =�j2� �� �5j���2 fk��� �7�5���7�h l mn f�@@h �����������; <7� �@7;25 �2;� l mo f2p�@45��= q <7� �2;�h �����������; <7� ��7;; j�@�5���7� r�%� �"� ����$ %"� ��/%��� ���� �� 4;2�s; 5��@7=42 �2>�7; :2��4�2; ��5 �2<2�2��2 @�t2@; <7� ;2=>2�� ����; u72<<���2��;f�v th w2�2 <�p25 �� � 2 2p�2��>2�� ?��@7=42 �2>�7; ��5 ����7�����2 � �2;x 7< 2�� 4;2� figure 9: training and test data usage for experimental evaluation table 1: deviation of parameters in the linear regression function a b average 0.883 −0.431 std. dev. 0.034 0.057 dialogue tempo at each point in time using the linear regression function. more specifically, thresholds were changed by the function in the thresholding method, and feature values were converted by eq. (2) in the decision tree method. then classification accuracy was calculated by comparing the outputs of the classifier and reference labels for the pairs of vad results. the deviations of the two parameters of the linear regression function y = ax + b during the cross validation for the thresholding method are listed in table 1. the two parameter values, a and b, only changed slightly, and their averages were almost the same as the coefficients in eq. (1), which were calculated using all the data. thus, we used the same parameter values (a = 0.88 and b = −0.43) in the decision tree method for simplicity of experimentation. 6.1.1 thresholding adapted to users classification accuracies in thresholding are listed in the left column of table 2. in the “no adaptation” condition for the closed test, a constant threshold (0.822 seconds) was used to classify all data. this threshold was determined optimally for all data by an svm (smo in weka) in the same 217 komatani, hotta, sato, and nakano table 2: classification accuracies with/without adaptation thresholding decision tree closed cross validation closed cross validation no adaptation 281/390 (72.1%) 281/390 (72.1%) 312/390 (80.0%) 271/390 (69.5%) online adaptation 294/390 (75.4%) 293/390 (75.1%) 320/390 (82.1%) 300/390 (76.9%) manner as discussed in section 4.2. in the “no adaptation” condition for cross validation, thresholds were determined for each fold; that is, data of 25 participants were used to determine the threshold for testing the excluded data of one participant. this was repeated 26 times. user adaptation improved classification accuracies by 3.3 and 3.0 percentage points for the closed test and cross validation, respectively. we can also see that the accuracies were almost equivalent under both adaptation conditions (“no” and “online”). this suggests that no overfitting occurred in these cases, so a similar performance will also be obtained for unknown users. the number of parameters is small, which is why they are stable, as already shown in table 1. 6.1.2 decision tree learning adapted to users classification accuracies for decision tree learning are listed in the right column of table 2. the “no adaptation” condition denotes normal decision tree learning, that is, no feature values were converted using eq. (2). user adaptation improved the accuracies by 2.1 and 7.4 percentage points for the closed test and the cross validation, respectively. the difference in cross validation was statistically significant according to the mcnemar test (p = 3.2 × 10−4). we can see that the accuracies in cross validation were lower than those in the closed test. this is because a decision tree has many more parameters to be trained than thresholding, so the obtained trees were overfitted to the training data. this means that the accuracies under the closed test condition were unreasonably high. note that the accuracy under the “no adaptation” condition in cross validation was lower than that in the same condition for the thresholding. this means that the complicated classifier makes the accuracy worse. in contrast, when user adaptation was performed, the accuracy the for decision tree under the “online adaptation” condition in cross validation outperformed that of thresholding. this implies that user adaptation makes the features more general and essential, so overfitting was avoided even when the more complicated classifier (decision tree) was used. figure 10 shows the top part of the obtained decision tree, whose depth did not exceed four. the feature at the top was the temporal interval after the user adaptation. this fact also demonstrates that the feature was effective in the decision tree. 6.2 comparison with batch adaptation we calculated dialogue tempos by using the whole dialogue containing the target utterance. this condition, called “batch adaptation,” corresponds to a case where the target user’s characteristics have already been obtained. we discuss accuracy under this condition because this can be regarded as an upper limit of user adaptation. because the accuracies of the batch adaptation were calculated as closed tests, those of the on line adaptation were calculated also as closed tests. 218 user adaptive a posteriori restoration for incorrectly segmented utterances <=0.470 >0.470 <=0.465 >0.465 <=210.17 >210.17 not noise noise noise detection results by gmm temporal interval (after adaptation) not required f0 range of first fragment avg. confidence of first fragment required not required max. loudness of first fragment temporal interval (after adaptation) figure 10: obtained decision tree (depth < 4) table 3: classification accuracies by adaptation methods thresholding decision trees no 281/390 (72.1%) 312/390 (80.0%) online 294/390 (75.4%) 320/390 (82.1%) batch 306/390 (78.5%) 331/390 (84.9%) oracle 334/390 (85.7%) – table 3 shows the classification accuracies under the no adaptation and two adaptation conditions. here, for the simplicity of experiments in the decision tree method, we used the same decision tree (including branching conditions and tree structure) as the batch adaptation, but the temporal interval features were re-estimated over time; that is, the available number of switching pause durations to calculate dialogue tempos increased online. the results show that the accuracies of the batch adaptation were higher than those of online adaptation conditions by 3.1 and 2.8 percentage points for thresholding and decision trees, respectively. this implies that the classification accuracy improves when plenty of utterances from the target user are available. furthermore, we investigated an “oracle” condition, where the optimal thresholds were determined for each user by svm as in section 4.2. the accuracy was 85.7% (334/390), as also shown in table 3. this result shows that adaption still has room for improving the accuracy more by capturing the detailed characteristics of each user. 6.3 convergence speed of adaptation we further investigated the convergence speed of the online adaptation. we conducted the following experiments only for thresholding for simplicity of implementation. the classification accuracy of online adaptation naturally converges into that of batch adaptation when the number of the target 219 komatani, hotta, sato, and nakano figure 11: convergence speed of adaptation (in thresholding) user’s available utterances increases because batch adaptation assumes that all utterances are obtained beforehand. we plotted the classification performance when the number of a target user’s available utterances increased to analyze its convergence speed. here, the levels of performance were calculated as closed tests, similarly with those of the previous section. figure 11 shows the number of correct classifications when the number of available utterances for online adaptation increased. the vertical and horizontal axes denote the number of correct classifications and available utterances for the adaptation, respectively. more specifically, the horizontal axis shows that the user’s dialogue tempo was calculated using data from the beginning of the dialogue to the x-th utterance. the dashed line and the dotted line represent batch adaptation, i.e., y = 306 and no adaptation, i.e., y = 281, respectively. both of these results are listed in table 3. we can see that when the number of available utterances was small (x < 10), the number of correct classifications was significantly varied and also small (about 275). correct classification results increased when the available utterances increased and became equivalent to that of batch adaptation after x = 80. this shows that the performance converged at about 80 utterances. these results lead us to the following conclusions. first, when the number of available utterances is small, i.e., less than 10, the classifier should not be adapted because the levels of performance were lower than those under the “no adaptation” condition. performance does not degrade if we adapt the classifier after about 10 utterances are obtained from the target user. second, although a one-shot user will probably not make 80 utterances at once, such a number of utterances can be obtained when user ids are available and when a user’s utterances are obtained through several sessions. user ids can be obtained when the system is used through personal terminals (e.g., cell phones) or when using techniques such as speaker identification. 7. conclusions and future work we developed a user-adaptive method to classify whether restoration is required for incorrectly segmented utterances by focusing on each user’s style of speaking. we empirically showed the correlation between dialogue tempo and appropriate thresholds for temporal intervals between utterance fragments, which are an important feature for the classification. we then investigated classification 220 user adaptive a posteriori restoration for incorrectly segmented utterances accuracies by online adaptation of two classifiers: thresholding and decision trees. results showed that the accuracies improved in both classifiers more than in the baselines using a constant threshold for all users. several issues remain as future work to improve the classification accuracy even more. first, we can exploit aspects other than dialogue tempo to represent each user’s style of speaking, such as the speaking rate and the frequency of self-repairs. lexical or semantic features, which were used in previous studies such as (nakano et al., 1999), can also be used. second, we want to adapt features in the feature set other than the temporal interval between two utterance fragments used in this paper. for example, the maximum loudness in the first fragment can be adapted to each user. the f0 range of the first fragment can also be a target of adaptation because some users have habitual intonation at the end of utterances. third, the experiments in this paper were conducted using already recorded dialogue data between a human and a system. the user behaviors in this data may have been influenced by the system performance when the data was collected. therefore, we need to conduct another experiment where a system with the proposed method actually interacts with humans. other metrics such as user satisfaction and completion time will be helpful to demonstrate the performance. finally, variations in speaking styles naturally exist within the same user as well as across users when the system is used repeatedly (komatani et al., 2009). this occurs especially when the user first tries the system. if much more data per user are made available to analyze, the classification accuracy may improve. references timo baumann and david schlangen. predicting the micro-timing of user input for an incremental spoken dialogue system that completes a user’s ongoing turn. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 120–129, 2011. linda bell, johan boye, and joakim gustafson. real-time handling of fragmented utterances. in proc. naacl workshop on adaption in dialogue systems, pages 2–8, 2001. a. benyassine, e. shlomot, huan yu su, d. massaloux, c. lamblin, and j.-p. petit. itu-t recommendation g.729 annex b: a silence compression scheme for use with g.729 optimized for v.70 digital simultaneous voice and data applications. ieee communications magazine, 35(9):64–73, 1997. iwan de kok, dirk heylen, and louis-philippe morency. speaker-adaptive multimodal prediction model for listener responses. in proc. international conference on multimodal interaction (icmi), pages 51–58, 2013. kohji dohsaka, atsushi kanemoto, ryuichiro higashinaka, yasuhiro minami, and eisaku maeda. user-adaptive coordination of agent communicative behavior in spoken dialogue. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 314–321, 2010. jens edlund, mattias heldner, and joakim gustafson. utterance segmentation and turn-taking in spoken dialogue systems. in bernhard fisseni, hans-christian schmitz, bernhard schröder, and 221 komatani, hotta, sato, and nakano petra wagner, editors, computer studies in language and speech, volume 8, pages 576–587. peter lang, frankfurt am main, 2005. luciana ferrer, elizabeth shriberg, and andreas stolcke. a prosody-based approach to end-ofutterance detection that does not require speech recognition. in proc. ieee international conference on acoustics, speech & signal processing (icassp), volume 1, pages 608–611, 2003. kotaro funakoshi, mikio nakano, kazuki kobayashi, takanori komatsu, and seiji yamada. nonhumanlike spoken dialogue: a design perspective. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 176–184, 2010. mark hall, eibe frank, geoffrey holmes, bernhard pfahringer, peter reutemann, and ian h. witten. the weka data mining software: an update. acm sigkdd explorations newsletter, 11: 10–18, november 2009. mattias heldner and jens edlund. pauses, gaps and overlaps in conversations. journal of phonetics, 38(4):555 – 568, 2010. naoki hotta, kazunori komatani, satoshi sato, and mikio nakano. detecting incorrectlysegmented utterances for posteriori restoration of turn-taking and asr results. in proc. annual conference of the international speech communication association (interspeech), pages 313–317, 2014. kristiina jokinen and kari kanto. user expertise modeling and adaptivity in a speech-based e-mail system. in proc. annual meeting of the association for computational linguistics (acl), pages 87–94, 2004. norihide kitaoka, masashi takeuchi, ryota nishimura, and seiichi nakagawa. response timing detection using prosodic and linguistic information for human-friendly spoken dialog systems. journal of the japanese society for artificial intellignece, 20(3):220–228, 2005. kazunori komatani, shinichi ueno, tatsuya kawahara, and hiroshi g. okuno. user modeling in spoken dialogue systems to generate flexible guidance. user modeling and user-adapted interaction, 15(1):169–183, 2005. kazunori komatani, tatsuya kawahara, and hiroshi g. okuno. a model of temporally changing user behaviors in a deployed spoken dialogue system. in proc. international conference on user modeling, adaptation, and personalization (umap), volume 5535 of lecture notes in computer science, pages 409–414. springer, 2009. isbn 978-3-642-02246-3. kazunori komatani, naoki hotta, and satoshi sato. restoring incorrectly segmented keywords and turn-taking caused by short pauses. in proc. international workshop on spoken dialogue systems (iwsds), pages 27–38, 2014. kazunori komatani, naoki hotta, satoshi sato, and mikio nakano. user adaptive restoration for incorrectly segmented utterances in spoken dialogue systems. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 393–401, 2015. 222 user adaptive a posteriori restoration for incorrectly segmented utterances akinobu lee and tatsuya kawahara. recent development of open-source speech recognition engine julius. in proc. apsipa asc: asia-pacific signal and information processing association, annual summit and conference, pages 131–137, 2009. mikio nakano, noboru miyazaki, jun ichi hirasawa, kohji dohsaka, and takeshi kawabata. understanding unsegmented user utterances in real-time spoken dialogue systems. in proc. annual meeting of the association for computational linguistics (acl), pages 200–207, 1999. mikio nakano, shun sato, kazunori komatani, kyoko matsuyama, kotaro funakoshi, and hiroshi g. okuno. a two-stage domain selection framework for extensible multi-domain spoken dialogue systems. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 18–29, 2011. tomoko ohsuga, masafumi nishida, yasuo horiuchi, and akira ichikawa. investigation of the relationship between turn-taking and prosodic features in spontaneous dialogue. in proc. european conf. speech commun. & tech. (eurospeech), 2005. tim paek and david maxwell chickering. improving command and control speech recognition on mobile devices: using predictive user models for language modeling. user modeling and user-adapted interaction, 17(1-2):93–117, 2007. maike paetzel, ramesh manuvinakurike, and david devault. ”so, which one is it?” the effect of alternative incremental architectures in a high-performance game-playing agent. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 77–86, 2015. antoine raux and maxine eskenazi. optimizing endpointing thresholds using dialogue features in a spoken dialogue system. in proc. sigdial workshop on discourse and dialogue, pages 1–10, 2008. antoine raux and maxine eskenazi. a finite-state turn-taking model for spoken dialog systems. in proc. human language technologies: the 2009 annual conference of the north american chapter of the association for computational linguistics (hlt naacl), pages 629–637, 2009. hiroyuki sakai, tobias cincarek, hiromichi kawanami, hiroshi saruwatari, kiyohiro shikano, and akinobu lee. voice activity detection applied to hands-free spoken dialogue robot based on decoding using acoustic and language model. in proc. international conference on robot communication and coordination (robocomm), 2007. ryo sato, ryuichiro higashinaka, masafumi tamoto, mikio nakano, and kiyoaki aikawa. learning decision trees to determine turn-taking by spoken dialogue systems. in proc. international conference on spoken language processing (icslp), pages 861–864, 2002. ethan selfridge, iker arizmendi, peter a. heeman, and jason d. williams. stability and accuracy in incremental speech recognition. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 110–119, 2011. gabriel skantze and anna hjalmarsson. towards incremental speech generation in dialogue systems. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 1–8, 2010. 223 komatani, hotta, sato, and nakano david traum, david devault, jina lee, zhiyang wang, and stacy marsella. incremental dialogue understanding and feedback for multiparty, multimodal conversation. in intelligent virtual agents, volume 7502 of lecture notes in computer science, pages 275–288. springer berlin heidelberg, 2012. isbn 978-3-642-33196-1. tiancheng zhao, alan w. black, and maxine eskenazi. an incremental turn-taking model with active system barge-in for spoken dialog systems. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 42–50, 2015. 224 dialogue & discourse 11(2) (2020) 1–33 doi: 10.5087/dad.2020.201 a neural approach to discourse relation signal detection amir zeldes amir.zeldes@georgetown.edu georgetown university department of linguistics yang (janet) liu yl879@georgetown.edu georgetown university department of linguistics editor: amanda stent submitted 09/2019; accepted 02/2020; published online 07/2020 abstract previous data-driven work investigating the types and distributions of discourse relation signals, including discourse markers such as ‘however’ or phrases such as ‘as a result’ has focused on the relative frequencies of signal words within and outside text from each discourse relation. such approaches do not allow us to quantify the signaling strength of individual instances of a signal on a scale (e.g. more or less discourse-relevant instances of ‘and’), to assess the distribution of ambiguity for signals, or to identify words that hinder discourse relation identification in context (‘anti-signals’ or ‘distractors’). in this paper we present a data-driven approach to signal detection using a distantly supervised neural network and develop a metric, ∆s (or ‘delta-softmax’), to quantify signaling strength. ranging between -1 and 1 and relying on recent advances in contextualized words embeddings, the metric represents each word’s positive or negative contribution to the identifiability of a relation in specific instances in context. based on an english corpus annotated for discourse relations using rhetorical structure theory and signal type annotations anchored to specific tokens, our analysis examines the reliability of the metric, the places where it overlaps with and differs from human judgments, and the implications for identifying features that neural models may need in order to perform better on automatic discourse relation classification. keywords: rst, signaling, discourse markers, neural network, contextual embeddings, connective detection, signal strength metric, delta s, rnn 1. introduction the development of formal frameworks for the analysis of discourse relations has long gone hand in hand with work on signaling devices. the analysis of discourse relations is also closely tied to what a discourse structure should look like and what discourse goals should be fulfilled in relation to the interpretation of discourse relations (roberts, 2012). earlier work on the establishment of inventories of discourse relations and their formalization (hovy and maier 1993, knott and dale 1994, knott 1996, knott and sanders 1998, webber and joshi 1998, fraser 1999) relied on the existence of ‘discourse markers’ (dms) or ‘connectives’, including conjunctions such as because or if, adverbials such as however or as a result, and coordinations such as but, to identify and distinguish relations such as condition in (1), concession in (2), cause in (3), or contrast, result etc., depending on the postulated inventory of relations (signals for these relations as identified by c©2020 amir zeldes, yang (janet) liu this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). zeldes and liu human analysts are given in bold; examples come from the gum corpus (zeldes, 2017), presented in section 3). (1) [if you work for a company,]condition [they pay you that money.] (2) [albeit limited,]concession [these results provide valuable insight into si interpretation by chitonga-speaking children.] (3) [not all would have been interviewed at wave 3] [due to differential patterns of temporary attrition]cause the same reasoning of identifying relations based on overt signals has been applied to the comparison of discourse relations across languages, by comparing inventories of similar function words cross-linguistically (taboada and de los ángeles gómez-gonzález 2012, versley and gastel 2013); and the annotation guidelines of prominent contemporary corpora rely on such markers as well: for instance, the penn discourse treebank (see prasad et al. 2008) explicitly refers to either the presence of dms or the possibility of their insertion in cases of implicit discourse relations, and dm analysis in rhetorical structure theory (mann and thompson, 1988) has also shown the important role of dms as signals of discourse relations at all hierarchical levels of discourse analysis (das and taboada, 2017). at the same time, research over the past two decades analyzing the full range of possible cues that humans use to identify the presence of discourse relations has suggested that classic dms such as conjunctions and adverbials are only a part of the network of signals that writers or speakers can harness for discourse structuring, which also includes entity-based cohesion devices (e.g. certain uses of anaphora, see poesio et al. 2002), alternative lexicalizations using content words, as well as syntactic constructions (see prasad et al. 2014 and the addition of alternative lexicalization constructions, altlexc, in the latest version of pdtb, prasad et al. 2019).1 in previous work, two main approaches to extracting the inventory of discourse signal types in an open-ended framework can be identified: data-driven approaches, which attempt to extract relevant words from distributional properties of the data, using frequencies or association measures capturing their co-occurrences with certain relation types (e.g. torabi asr and demberg 2013, toldova et al. 2017); and manual annotation efforts (e.g. prasad et al. 2008, taboada and das 2013), which develop categorization schemes and guidelines for human evaluation of signaling devices. the former family of methods benefits from an unbiased openness to any and every type of word which may reliably co-occur with some relation types, whether or not a human might notice it while annotating, as well as the naturally graded and comparable nature of the resulting quantitative scores, but, as we will show, falls short in identifying specific cases of a word being a signal (or not) in context. by contrast, the latter approach allows for the identification of individual instances of signaling devices, but relies on less open-ended guidelines and is categorical in nature: a word either is or isn’t a signal in context, providing less access to concepts such as signaling strength. the goal of this paper is to develop and evaluate a model of discourse signal identification that is built bottom up from the data, but retains sensitivity to context in the evaluation of each individual example.2 1in the examples above as well, we can identify non-dm signals such as the concessive word limited in the concession relation in (2), or even the use of the generic present tense and the generic ‘you’ in “if you work” in (1) as markers of general conditional clauses. finding and agreeing on the exact span of words that constitutes a signal, or ‘signal anchoring’ (see section 3.1) are therefore non-trivial tasks. 2zeldes (2018b) showed some preliminary results from the early stages of this line of work. 2 a neural approach to discourse relation signal detection in addition, even though this work is conducted within rhetorical structural theory, we hope that it can shed light on signal identification of discourse relations across genres and provide empirical evidence to motivate research on theory-neutral and genre-diverse discourse processing, which would be beneficial for pushing forward theories of discourse across frameworks or formalisms. furthermore, employing a computational approach to studying discourse relations has a promising impact on various nlp downstream tasks such as question answering and document summarization etc. for example, narasimhan and barzilay (2015) incorporated discourse information into the task of automated text comprehension and benefited from such information without relying on explicit annotations of discourse structure during training, which outperformed state-of-the-art text comprehension systems at the time. towards this goal, we begin by reviewing some previous work in the traditions sketched out above in the next section, and point out some open questions which we would like to address. in section 3 we present the discourse annotated data that we will be using, which covers a number of english text types from the web annotated for 20 discourse relations in the framework of rhetorical structure theory, and is enriched with human annotations of discourse relation signaling devices for a subset of the data. moreover, we also propose a taxonomy of anchored signals based on the discourse annotated data used in this paper, illustrating the properties and the distribution of the anchorable signals. in section 4 we then train a distantly supervised neural network model which is made aware of the relations present in the data, but attempts to learn which words signal those relations without any exposure to explicit signal annotations. we evaluate the accuracy of our model using stateof-the-art pretrained and contextualized character and word embeddings, and develop a metric for signaling strength based on a masking concept similar to permutation importance, which naturally lends itself to the definition of both positive and negative or ‘anti-signals’, which we will refer to as ‘distractors’. in section 5, we combine the anchoring annotation data from section 3 with the model’s predictions to evaluate how ‘human-like’ its performance is, using an information retrieval approach measuring recall@k and assessing the stability of different signal types based on how the model scores them. we develop a visualization for tokenwise signaling strength and perform error analysis for some signals found by the model which were not flagged by humans and vice versa, and point out the strengths and weaknesses of the architecture. section 6 offers further discussion of what we can learn from the model, what kinds of additional features it might benefit from given the error analysis, and what the distributions of scores for individual signals can teach us about the ambiguity and reliability of different signal types, opening up avenues for further research. 2. previous work 2.1 data-driven approaches a straightforward approach to identifying discourse relation signals in corpora with discourse parses is to extract frequency counts for all lexical types or lemmas and cross-tabulate them with discourse relations, such as sentences annotated as cause, elaboration, etc. (e.g. scheffler and stede 2016, mı́rovský et al. 2017, 73, toldova et al. 2017). table 1, reproduced from toldova et al. (2017, 32), illustrates this approach for the russian rst treebank. this approach quickly reveals the core inventory of cue words in the language, and in particular the class of low-ambiguity discourse markers (dms), such as odnako ‘however’ signaling con3 zeldes and liu relation type freq marker translation elaboration 150 kotoryj which, that joint 119 i, takzhe and, as well attributon 118 zajavil, soobschil report, announce etc. contrast 61 odnako, a, no however, but cause-effect 47 poetomu, v+prichina so, accordingly, v+cause purpose 39 chtoby, dlya in order that, for interpretation-evaluation 34 nouns / verbs expressing opinion background 31 no dominant marker condition 27 esli if table 1: russian discourse relation signals, reproduced from toldova et al. (2017). trast (see fraser 1999 on delimiting the class of explicit dms) or relative pronouns signaling elaboration. as such, it can be very helpful for corpus-based lexicography of discourse markers (cf. stede and umbach 1998). the approach can potentially include multiword expressions, if applied equally to multi-token spans (e.g. as a result), and because it is language-independent, it also allows for a straightforward comparison of connectives or other dms across languages. results may also converge across frameworks, as the frequency analysis may reveal the same items in different corpora annotated using different frameworks. for instance, the inventory of connectives found in work on the penn discourse treebank (pdtb, see prasad et al. 2008) largely converges with findings on connectives using rst (see stede 2012, 97–109, taboada and das 2013): conjunctions such as but can mark different kinds of contrastive relations at a high level, and adverbs such as meanwhile can convey contemporaneousness, among other things, even when more fine-grained analyses are applied. however, a purely frequentist approach runs into problems on multiple levels, as we will show in section 4: high frequency and specificity to a small number of relations characterize only the most common and unambiguous discourse markers, but not less common ones. additionally, differentiating actual and potentially ambiguous usages of candidate words in context requires substantial qualitative analysis (see da cunha 2013), which is not reflected in aggregated counts, and signals that belong to a class of relations (e.g. a variety of distinct contrastive relation types) may appear to be non-specific, when in fact they reliably mark a superset of relations. other studies have used more sophisticated metrics, such as point-wise mutual information (pmi), to identify words associated with particular relations (torabi asr and demberg, 2013). using the pdtb corpus, torabi asr and demberg (2013) extracted such scores and measured the contribution of different signal types based on the information gain which they deliver for the classification of discourse relations at various degrees of granularity, as expressed by the hierarchical labels of pdtb relation types. this approach is most similar to the goal given to our own model in section 4, but is less detailed in that the aggregation process assigns a single number to each candidate lexical item, rather than assigning contextual scores to each instance.3 3in this overview we disregard supervised approaches to detecting signals in the sense of connective detection, which rely primarily on labeled training data. such approaches can be very effective in reproducing benchmark annotations from corpora such as pdtb, but do so by effectively memorizing lexical items from gold standard annotated data (see muller et al. 2019 for recent state-of-the art work). as such, they are less interesting for open-ended exploration of signaling devices. 4 a neural approach to discourse relation signal detection finally we note that for hierarchical discourse annotation schemes, the data-driven approaches described here become less feasible at higher levels of abstraction, as relations connecting entire paragraphs encompass large amounts of text, and it is therefore difficult to find words with high specificity to those relations. as a result, approaches using human annotation of discourse relation signals may ultimately be irreplaceable. 2.2 discourse relation signal annotations discourse relation signals are broadly classified into two categorizes: anchored signals and unanchored signals. by ‘anchoring’ we refer to associating signals with concrete token spans in texts. intuitively, most of the signals are anchorable since they correspond to certain token spans. however, it is also possible for a discourse relation to be signaled but remain unanchored. results from liu (2019) indicated that there are several signaled but unanchored relations such as preparation and background since they are high-level discourse relations that capture and correspond to genre features such as interview layout in interviews where the conversation is constructed as a question-answer scheme, and are thus rarely anchored to tokens. the penn discourse treebank (pdtb v3, prasad et al. 2019) is the largest discourse annotated corpus of english, and the largest resource annotated explicitly for discourse relation signals such as connectives, with similar corpora having been developed for a variety of languages (e.g. zeyrek et al. 2013 for turkish, zhou et al. 2014 for chinese). however the annotation scheme used by pdtb is ahierarchical, annotating only pairs of textual argument spans connected by a discourse relation, and disregarding relations at higher levels, such as relations between paragraphs or other groups of discourse units. additionally, the annotation scheme used for explicit signals is limited to specific sets of expressions and constructions, and does not include some types of potential signals, such as the graphical layout of a document, lexical chains of (non-coreferring) content words that are not seen as connectives, or genre conventions which may signal the discourse function for parts of a text. it is nevertheless a very useful resource for obtaining frequency lists of the most prevalent dms in english, as well as data on a range of phenomena such as anaphoric relations signaled by entities, and some explicitly annotated syntactic constructions. working in the hierarchical framework of rhetorical structure theory (mann and thompson, 1988), taboada and das (2013) re-annotated the existing rst discourse treebank (carlson et al., 2003), by taking the existing discourse relation annotations in the corpus as a ground truth and analyzing any possible information in the data, including content words, patterns of repetition or genre conventions, as a possibly present discourse relation signaling device. the resulting rst signalling corpus (rst-sc, das et al. 2019) consists of 385 wall street journal articles from the penn treebank (marcus et al., 1993), a smaller subset of the same corpus used in pdtb. it contains 20,123 instances of 78 relation types (e.g. attribution, circumstance, result etc.), which are enriched with 29,297 signal annotations. das and taboada (2017) showed that when all types of signals are considered, over 86% of discourse relations annotated in the corpus were signaled in some way, but among these, just under 20% of cases were marked by a dm. however, unlike pdtb, the rst signalling corpus does not provide a concrete span of tokens for the locus of each signal, indicating instead only the type of signaling device used. although the signal annotations in rst-sc have a broader scope than those in pdtb and are made more complex by extending to hierarchical relations, liu and zeldes (2019) have shown that rst-sc’s annotation scheme can be ‘anchored’ by associating discourse signal categories from 5 zeldes and liu rst-sc with concrete token spans. liu (2019) applied the same scheme to a data set described in section 3, which we will use to evaluate our model in section 5. since that data set is based on the same annotation scheme of signal types as rst-sc, we will describe the data for the present study and rst-sc signal type annotation scheme next. 3. data 3.1 anchored signals in the gum corpus in order to study open-ended signals anchored to concrete tokens, we use the signal-annotated subset of the freely available georgetown university multilayer (gum) corpus (zeldes, 2017) from liu (2019). our choice to use a multi-genre rst-annotated corpus rather than using pdtb, which also contains discourse relation signal annotation to a large extent is motivated by three reasons: the first reason is that we wish to explore the full range of potential signals, as laid out in the work on the signalling corpus (das and taboada, 2017, 2018), whereas pdtb annotates only a subset of the possible cues identified by human annotators. secondly, the use of rst as a framework allows us to examine discourse relations at all hierarchical levels, including long distance, high-level relations between structures as large as paragraphs or sections, which often have different types of signals allowing their identification. finally, although the entire gum corpus is only about half the size of rst-dt (109k tokens4), using gum offers the advantage of a more varied range of genres than pdtb and rst-sc, both of which annotate wall street journal data. the signal annotated subset of gum includes academic papers, how-to guides, interviews and news text, encompassing over 11,000 tokens. although this data set may be too small to train a successful neural model for signal detection, we will not be using it for this purpose; instead, we will reserve it for use solely as a test set, and use the remainder of the data (about 98k tokens) to build our model (see section 4.2 for more details about the subsets and splits), including data from four further genres, for which the corpus also contains rst annotations but no signaling annotations: travel guides, biographies, fiction, and reddit forum discussions. the gum corpus is manually annotated with a large number of layers, including document layout (headings, paragraphs, figures, etc.); multiple pos tags (penn tags, claws5, universal pos); lemmas; sentence types (e.g. imperative, wh-question etc., leech et al. 2003); universal dependencies (nivre et al., 2017); (non-)named entity types; coreference and bridging resolution; and discourse parses using rhetorical structure theory (mann and thompson, 1988). in particular, the rst annotations in the corpus use a set of 20 commonly used rst relation labels, which are given in table 2, along with their frequencies in the corpus.5 the relations cover asymmetrical prominence relations (satellite-nucleus) and symmetrical ones (multinuclear relations), with the restatement relation being realized in two versions, one for each type. the signaling annotation in the corpus follows the scheme developed by rst-sc, with some additions. although rst-sc does not indicate token positions for signals, it provides a detailed taxonomy of signal types which is hierarchically structured into three levels: 4during the review period, the corpus has since grown to 130k tokens, however we were not able to include these in the analyses presented here. 5during the review period, the rst component has undergone segmentation and relation set changes (there are now 25 labels in the corpus), however we were not able to include these in the analyses presented here. 6 a neural approach to discourse relation signal detection relation type # relation type # joint multinuclear 1248 condition satellite-nucleus 210 elaboration satellite-nucleus 1037 justify satellite-nucleus 203 sequence multinuclear 546 result satellite-nucleus 185 preparation satellite-nucleus 500 solutionhood satellite-nucleus 173 background satellite-nucleus 464 restatement multinuclear or 150 circumstance satellite-nucleus 314 satellite-nucleus evaluation satellite-nucleus 283 evidence satellite-nucleus 147 concession satellite-nucleus 271 purpose satellite-nucleus 136 contrast multinuclear 250 motivation satellite-nucleus 124 cause satellite-nucleus 242 antithesis satellite-nucleus 109 table 2: rst relations and their frequencies in the gum corpus. 1. signal class, denoting the signal’s degree of complexity 2. signal type, indicating the linguistic system to which it belongs 3. specific signal, which gives the most fine-grained subtypes of signals within each type it is assumed that any number of word tokens can be associated with any number of signals (including the same tokens participating in multiple signals), that signals can arise without corresponding to specific tokens (e.g. due to graphical layout of paragraphs), and that each relation can have an unbounded number of signals (0− n), each of which is characterized by all three levels. the signal class level is divided into single, combined (for complex signals), and unsure for unclear signals which cannot be identified conclusively, but are noted for further study. for each signal (regardless of its class), signal type and specific signal are identified. according to rst-sc’s taxonomy, signal type includes 9 types such as dms, genre, graphical, lexical, morphological, numerical, reference, semantic, and syntactic. each type then has specific subcategories. for instance, the signal type semantic has 7 specific signal subtypes: synonymy, antonymy, meronymy, repetition, indicative word pair, lexical chain, and general word. we will describe some of these in more depth below. in addition to the 9 signal types, rst-sc has 6 combined signal types such as reference+syntactic, semantic+syntactic, and graphical+syntactic etc., and 15 specific signals are identified for the combined signals. although the rich signaling annotations in rst-sc offer an excellent overview of the relative prevalence of different signal types in the wall street journal corpus, it is difficult to apply the original scheme to the study of individual signal words, since actual signal positions are not identified. while recovering these positions may be possible for some categories using the original guidelines,6 most signaling annotations (e.g. lexical chains, repetition) cannot be automatically paired with actual tokens, meaning that, in order to use the original rst-sc for our study, we would need to re-annotate it for signal token positions. as this effort is beyond the scope of our study, we will use the smaller data set with anchored signaling annotations from liu (2019): this data is annotated with the same signal categories as rst-sc, but also includes exact token positions 6especially dms, of which 201 lexical types are identified in rst-sc, could be searched for in the relevant text spans. however even for these we would not always be sure which word is intended in spans containing multiple instances of the relevant word. 7 zeldes and liu for each signal, including possibly no tokens for unanchorable signals such as some types of genre conventions or graphical layout which are not expressible in terms of specific words. in order to get a better sense of how the annotations work, we consider example (4). (4) [5] sociologists have explored the adverse consequences of discrimination; [6] psychologists have examined the mental processes that underpin conscious and unconscious biases; [7] neuroscientists have examined the neurobiological underpinnings of discrimination; [8] and evolutionary theorists have explored the various ways that in-group/out-group biases emerged across the history of our species. – joint [gum academic discrimination] figure 1: a visualization of an rst analysis of (4) with the signal tokens highlighted. in this example, there is a joint relation between four spans in a fragment from an rst discourse tree. the first tokens in each span form a parallel construction and include semantically related items such as explored and examined (signal class ‘combined’, type ‘semantic+syntactic’, specific subtype ‘parallel syntactic construction + lexical chain’). the words corresponding to this signal in each span are highlighted in figure 1, and are considered to signal each instance of the joint relation. additionally, the joint relation is also signaled by a number of further signals which are highlighted in the figure as well, such as the semicolons between spans, which correspond to a type ‘graphical’, subtype ‘semicolon’ in rst-sc. the data model of the corpus records which tokens are associated with which categorized signals, and allows for multiple membership of the same token in several signal annotations. in terms of annotation reliability, das and taboada (2017) reported a weighted kappa of 0.71 for signal subtypes in rst-sc without regard to the span of words corresponding to a signal, while a study by gessler et al. (2019) suggests that signal anchoring, i.e. associating rst-sc signal categories with specific tokens achieves a 90.9% perfect agreement score on which tokens constitute signals, or a cohen’s kappa value of 0.77. as anchored signal positions will be of the greatest interest to our study, we will consider how signal token positions are distributed in the corpus next, and develop an anchoring taxonomy which we will refer back to for the remainder of this paper. 8 a neural approach to discourse relation signal detection 3.2 a taxonomy of anchored signals from a structural point of view, one of the most fundamental distinctions with regard to signal realization recognized in previous work is the classification of signaling tokens into satellite or nucleus-oriented positions, i.e. whether a signal for the relation appears within the modifier span or the span being modified (taboada, 2006). while some relation types exhibit a strong preference for signal position (e.g. using a discourse marker such as because in the satellite for cause, duque 2014), others, such as concession are more balanced (almost evenly split signals between satellite and nucleus in taboada 2006). in this study we would like to further refine the taxonomy of signal positions, breaking it down into several features. at the highest level, we have the distinction between anchorable and non-anchorable signals, i.e. signals which correspond to no token in the text (e.g. genre conventions, graphical layout). below this level, we follow taboada (2006) in classifying signals as satellite or nucleus-oriented, based on whether they appear in the more prominent elementary discourse unit (edu) of a relation or its dependent. however, several further distinctions may be drawn: • whether the signal appears before or after the relation in text order; since we consider the relation to be instantiated as soon as its second argument in the text appears, ‘before’ is interpreted as any token before the second head unit in the discourse tree begins, and ‘after’ is any subsequent token • whether the signal appears in the head unit of the satellite/nucleus, or in a dependent of that unit; this distinction only matters for satellite or nucleus subtrees that consist of more than one unit • whether the signal is anywhere within the structure dominated by the units participating in the relation, or completely outside of this structure table 3 gives an overview of the taxonomy proposed here, which includes the possible combinations of these properties and the distribution of the corresponding anchorable signals found in the signal-annotated subset of the gum corpus from liu (2019).7 individual feature combinations +a(nchorable) -a before after i(nside) o(utside) i(nside) o(utside) h(ead) d(ependent) h(ead) d(ependent) sat nuc sat nuc sat nuc sat nuc group i ii iii iv v vi vii viii ix x xi #tok 513 1073 311 648 0 1656 302 948 1114 0 0 #edu 157 281 25 55 0 457 89 109 144 0 57 table 3: an overview of the taxonomy and the distribution of the corresponding anchorable signals attested in the subset of the gum corpus from liu (2019). 7in this paper we disregard a further set of distinctions pertaining to the position of signals within a discourse unit (unit initial, final, medial etc.), though these too have been explored in the past, especially in the context of cohesive discourse generation (power et al., 1999). 9 zeldes and liu can be referred to either as acronyms, e.g. abihs for ‘anchorable, before the second edu of the relation, inside the relation’s subtree, head unit of the satellite’, or using the group ids near the bottom of the table (in this case the category numbered roman i). we will refer back to these categories in our comparison of manually annotated and automatically predicted signals. to illustrate how the taxonomy works in practice, we can consider the example in figure 2, which shows a signal whose associated tokens instantiate categories i and iv in a discourse tree – the words demographic variables appear both within a preparation satellite (unit [50], category i), which precedes and points to its nucleus [51–54], and within a satellite inside that block (unit [52], a dependent inside the nucleus block, category iv). based on the rst-sc annotation scheme, the signal class is simple, with the type semantic and specific sub-type lexical chain. the numbers at the bottom of table 3 show the number of tokens signaling each relation at each position, as well as the number of relations which have signal tokens at the relevant positions. the hypothetical categories v and x, with signal tokens which are not within the subtree of satellite or nucleus descendants, are not attested in our data, as far as annotators were able to identify.8 figure 2: example of a manually annotated signal with token positions corresponding to categories i and iv in the taxonomy with respect to the preparation relation of unit 50. 4. automatic signal extraction 4.1 a contextless frequentist approach to motivate the need for a fine-grained and contextualized approach to describing discourse relation signals in our data, we begin by extracting some basic data-driven descriptions of our data along the lines presented in section 2.1. in order to constrain candidate words to just the most relevant ones 8a conceivable example would be foreshadowing or summarizing the contents of a relation before or after any of the units involved, such as a discussion of an academic article’s sections acting as an early signal of an upcoming heading, which is itself analyzed as a preparation. though these cases were not found in the small sample used here, they have been studied in the past as a form of metalinguistic labeling (see francis (1994) for example). 10 a neural approach to discourse relation signal detection figure 3: rst fragment with signal obeying strong nuclearity: the dm thus indicating result is in the head edu of the satellite block. for marking a specific signal, we first need a way to address a caveat of the frequentist approach: higher order relations which often connect entire paragraphs (notably background and elaboration) must be prevented from allowing most or even all words in the document to be considered as signaling them. a simple approach to achieving this is to assume ‘strong nuclearity’, relying on marcu’s (1996) compositionality criterion for discourse trees (ccdt), which suggests that if a relation holds between two blocks of edus, then it also holds between their head edus. while this simplification may not be entirely accurate in all cases, table 3 suggests that it captures most signals, and allows us to reduce the space of candidate signal tokens to just the two head edus implicated in a relation. we will refer to signals within the head units of a relation as ‘endocentric’ and signals outside this region as ‘exocentric’. figure 3 illustrates this, where units [64] and [65] are the respective heads of two blocks of edus, and unit [65] in fact contains a plausible endocentric signal for the result relation, the discourse marker thus. more problematic caveats for the frequentist approach are the potential for over/underfitting and ambiguity. the issue of overfitting is especially thorny in small datasets, in which certain content words appear coincidentally in discourse segments with a certain function. table 4 shows the most distinctive lexical types for several discourse relations in gum based on pure ratio of occurrence in head edus marked for those relations. on the left, types are chosen which have a maximal frequency in the relevant relationship compared with their overall frequency in the corpus. this quickly overfits the contents of the corpus, selecting irrelevant words such as holiest and slate for the circumstance relation, or hypnotizing and currency for concession. the same lack of filtering can, however, yield some potentially relevant lexical items, such as causing for result or even highly specific content words such as ammonium, which are certainly not discourse markers, but whose appearance in a sequence is not accidental: the word is in this case typical for sequences in how-to guides, where use of ingredients in a recipe is described in a sequence. even if these kinds of items may be undesirable candidates for signal words in general, it seems likely that some 11 zeldes and liu relation f > 0 f > 10 solutionhood viable, contributed, 60th, touched, palestinians what, ?, why, did, how circumstance holiest, eventually, fell, slate, transition october, when, saturday, after, thursday result minuscule, rebuilding, distortions, struggle, causing result, phoenix, wikihow, funny, death concession until, favoured, hypnotizing, currency, curiosity although, while, though, however, call justify payoff, skills, net, presidential, supporters nato, makes, simply, texas, funny sequence feel, charter, ammonium, evolving, rests bottles, place, then, baking, soil cause malfunctioned, jams, benefiting, mandate, recognising because, wanted, religious, projects, stuff table 4: most distinctive lexemes for some relations in gum, with different frequency thresholds. rare content words may function as signals in context, such as evaluative adjectives (e.g. exquisite) enabling readers to recognize an evaluation.9 if we are willing to give up on the latter kind of rare items, the overfitting problem can be alleviated somewhat by setting a frequency threshold for each potential signal lexeme, thereby suppressing rare items. the items on the right of the table are limited to types occurring more than 10 times. since the most distinctive items on the left are all comparatively rare (and therefore exclusive to their relations), they do not overlap with the items on the right. looking at the items on the right, several signals make intuitive sense, especially for relations such as solutionhood (used for question-answer pairs) or concession, which show the expected wh words and auxiliary did, or discourse markers such as though, respectively. at the same time, some high frequency items may be spurious, such as nato for justify, which could perhaps be filtered out based on low dispersion across documents, but also stuff for cause, which probably could not be. another problem with the lists on the right is that some expected strong signals, such as the word and for sequence are absent from the table. this is not because and is not frequent in sequences, but rather because it is a ubiquitous word, and as a result, it is not very specific to the relation. however if we look at actual examples of and inside and outside of sequences, it is easy to notice that the kind of and that does signal a relation in context is often clause initial as in (5) and very different from the adnominal coordinating ands in (6), which do not signal the relation: (5) [she was made a dame by elizabeth ii for services to architecture,] [and in 2015 she became the first and only woman to be awarded the royal gold medal]sequence (6) [gordon visited england and scotland in 1686.] [in 1687 and 1689 he took part in expeditions against the tatars in the crimea]sequence 9arguably adjectives are the strongest signals of both positive and negative evaluative language (cf. benamara et al. 2017). 12 a neural approach to discourse relation signal detection these examples suggest that a data-driven approach to signal detection needs some way of taking context into account. in particular, we would like to be able to compare instances of signals and quantify how strong the signal is in each case. in the next section, we will attempt to apply a neural model with contextualized word embeddings (peters et al., 2018) to this problem, which will be capable of learning contextualized representations of words within the discourse graph. 4.2 a contextualized neural model task and model architecture since we are interested in identifying unrestricted signaling devices, we deliberately avoid a supervised learning approach as used in automatic signal detection trained on resources such as pdtb. while recent work on pdtb connective detection (muller et al. 2019, yu et al. 2019) achieves good results (f-scores of around 88-89 for english pdtb explicit connectives), the use of such supervised approaches would not tell us about new signaling devices, and especially about unrestricted lexical signals and other coherence devices not annotated in pdtb. additionally, we would be restricted to the newspaper text types represented in the wall street journal corpus, since no other large english corpus has been annotated for anchored signals. instead, we will adopt a distantly supervised approach: we will task a model with supervised discourse relation classification on data that has not been annotated for signals, and infer the positions of signals in the text by analyzing the model’s behavior. a key assumption, which we will motivate below, is that signals can have different levels of signaling strength, corresponding to their relative importance in identifying a relation. we would like to assume that different signal strength is in fact relevant to human analysts’ decision making in relation identification, though in practice we will be focusing on model estimates of strength, the usefulness of which will become apparent below. as a framework, we use the sentence classifier configuration of flair (akbik et al., 2019) with a bilstm encoder/classifier architecture fed by character and word level representations composed of a concatenation of fixed 300 dimensional glove embeddings (pennington et al., 2014), pretrained contextualized flair word embeddings, and pre-trained contextualized character embeddings from allennlp (gardner et al., 2018) with flair’s default hyperparameters. the model’s architecture is shown in figure 4. contextualized embeddings (peters et al., 2018) have the advantage of giving distinct representations to different instances of the same word based on the surrounding words, meaning that an adnominal and connecting two nps can be distinguished from one connecting two verbs based on its vector space representation in the model. using character embeddings, which give vector space representations to substrings within each word, means that the model can learn the importance of morphological forms, such as the english gerund’s -ing suffix, even for out-of-vocabulary items not seen during training. formally, the input to our system is formed of edu pairs which are the head units within the respective blocks of discourse units that they belong to, which are in turn connected by an instance of a discourse relation.10 this means that every discourse relation in the corpus is expressed as exactly one edu pair. each edu is encoded as a (possibly padded) sequence of n-dimensional vector representations of each word x1, .., xt , with some added separators which are encoded in the 10to simplify evaluation and ensure compatibility between the unit of measurement for human and model signal detection, we use gold segmented edus as input. an examination of the model’s performance in conjunction with automatic edu segmentation is outside the scope of this study. 13 zeldes and liu flair embeddings char embeddings glove embeddings bilstm softmax condition elaboration contrast . . . 0.1 0.12 0.6 0.13 decoder but usually not . . . f1 c1 g1 f2 c2 g2 f3 c3 g3 f4 c4 g4 f1 b1 f2 b2 f3 b3 f4 b4 input: sometimes this information is available , but usually not . label: concession 1 figure 4: model architecture. edu pairs are fed along with satellite, nucleus and separator markers in text order to an encoder using glove, flair and character embeddings. a bilstm classifier outputs probabilities for the possible discourse relations. same way and described below. the bidirectional lstm composes representations and context for the input, and a fully connected softmax layer gives the probability of each relation: softmax(reli) = ereli∑k j=1 e relj , hδt = fh(xt−1, ht−1, ct−1; θ)), cδt = fc(xt−1, ht−1, ct−1; θ)) where the probability of each relation reli is derived from the composed output of the function h across time steps 0 . . . t, δ ∈ {b, f} is the direction of the respective lstms, cδt is the recurrent context in each direction and θ = w, b gives the model weights and bias parameters (see akbik et al. 2019 for details). note that although the output of the system is ostensibly a probability distribution over relation types, we will not be directly interested in the most probable relation as outputted by the classifier, but rather in analyzing the model’s behavior with respect to the input word representations as potential signals of each relation. in order to capitalize on the system’s natural language modeling knowledge, edu satellitenucleus pairs are presented to the model in text order (i.e. either the nucleus or the satellite may 14 a neural approach to discourse relation signal detection come first).11 however, the model is given special separator symbols indicating the positions of the satellite and nucleus, which are essential for deciding the relation type (e.g. cause vs. result, which may have similar cue words but lead to opposite labels), and a separator symbol indicating the transition between satellite and nucleus. this setup is illustrated in (7). (7) sometimes this information is available , but usually not . label: concession in this example, the satellite precedes the nucleus and is therefore presented first. the model is made aware of the fact that the segment on the left is the satellite thanks to the tag . since the lstm is bi-directional, it is aware of positions being within the nucleus or satellite, as well as their proximity to the separator, at every time step. we reserve the signal-annotated subset of 12 documents from gum for testing, which contains 1,185 head edu pairs (each representing one discourse relation), and a random selection of 12 further documents from the remaining rst-annotated gum data (1,078 pairs) is taken as development data, leaving 102 documents (5,828 pairs) for training. the same edus appear in multiple pairs if a unit has multiple children with distinct relations, but no instances of edus are shard across partitions, since the splits are based on document boundaries. we note again that for the training and development data, we have no signaling annotations of any kind; this is possible since the network does not actually use the human signaling annotations we will be evaluating against: its distant supervision consists solely of the rst relation labels. relation classification performance although only used as an auxiliary training task, we can look at the model’s performance on predicting discourse relations, which is given in table 5. unsurprisingly, the model performs best on the most frequent relations in the corpus, such as elaboration or joint, but also on rarer ones which tend to be signaled explicitly, such as condition (often signaled explicitly by if ), solutionhood (used for question-answer pairs signaled by question marks and wh words), or concession (dms such as although). however, the model also performs reasonably well for some trickier (i.e. less often introduced by unambiguous dms) but frequent relations, such as preparation, circumstance, and sequence. rare relations with complex contextual environments, such as result, justify or antithesis, unsurprisingly do not perform well, with the latter two showing an f-score of 0. the relation restatement, which also shows no correct classifications, reveals a weakness of the model: while it is capable of recognizing signals in context, it cannot learn that repetition in and of itself, regardless of specific areas in vector space, is important (see section 6 for more discussion of these and other classification weaknesses). although this is not the actual task targeted by the current paper, we may note that the overall performance of the model, with an f-score of 44.37, is not bad, though below the performance of state-of-the-art full discourse parsers (see morey et al. 2017) – this is to be expected, since the model is not aware of the entire rst tree, rather looking only at edu pairs out of context, and given that standard scores on rst-dt come from a larger and more homogeneous corpus, with with fewer relations and some easy cases that are absent from gum.12 11flair embeddings compute a pooled word representation from contextualized character embeddings and previous occurrences of the entire word string akbik et al. (2019), meaning that the pre-trained model retains knowledge and expectations regarding the order of words in a text. 12for example, rst-dt annotates relative clauses, which are often easy to identify, as elaborations, while version 5 of gum, used here, does not segment adnominal clauses as edus at all. in the newest version 6 of gum, released during the review period, the same segmentation guidelines are followed as in rst-dt. additionally note that most discourse parsing on rst-dt uses a collapsed set of only 16 relations, as opposed to the 20 used here. however, we 15 zeldes and liu relation n train % train n test % test p r f antithesis 87 1.49 22 2.88 0 0 0 background 406 6.97 58 7.59 46.67 36.21 40.78 cause 224 3.84 18 2.36 23.08 16.67 19.36 circumstance 291 4.99 23 3.01 34.69 73.91 47.22 concession 237 4.07 34 4.45 20.00 23.53 21.62 condition 180 3.09 30 3.93 85.71 80.00 82.76 contrast 216 3.71 34 4.45 14.00 20.59 16.67 elaboration 921 15.8 116 15.18 34.85 39.66 37.10 evaluation 252 4.32 31 4.06 36.36 12.90 19.04 evidence 131 2.25 16 2.09 23.08 18.75 20.69 joint 1083 18.58 165 21.6 56.48 73.94 64.04 justify 168 2.88 35 4.58 0 0 0 motivation 112 1.92 12 1.57 13.33 16.67 14.81 preparation 432 7.41 68 8.9 69.33 76.47 72.73 purpose 117 2.01 19 2.49 83.33 78.95 81.08 restatement 119 2.04 31 4.06 0 0 0 result 167 2.87 18 2.36 10 5.56 7.15 sequence 530 9.09 16 2.09 11.76 12.50 12.12 solutionhood 155 2.66 18 2.36 54.55 66.67 60.00 total (micro avg) 5828 100 764 100 44.37 44.37 44.37 (macro avg) 32.49 34.37 32.48 table 5: model performance on relation classification in the test set. given the model’s performance on relation classification, which is far from perfect, one might wonder whether signal predictions made by our analysis should be trusted. this question can be answered in two ways: first, quantitatively, we will see in section 5 that model signal predictions overlap considerably with human judgments, even when the predicted relation is incorrect. intuitively, for similar relations, such as concession or contrast, both of which are adversative, the model may notice a relevant cue (e.g. ‘but’, or contrasting lexical items) despite choosing the wrong one. second, as we will see below, we will be analyzing the model’s behavior with respect to the probability of the correct relation, regardless of the label it ultimately chooses, meaning that the importance of predicting the correct label exactly will be diminished further. signaling metric the actual performance we are interested in evaluating is the model’s ability to extract signals for given discourse relations, rather than its accuracy in predicting the relations. to do so, we must extract anchored signal predictions from the model, which is non-trivial. while earlier work on interpreting neural models has focused on token-wise softmax probability (zeldes, 2018a) or attention weights (ghaeini et al., 2018), using contextualized embeddings complicates the evaluation: since word representations are adjusted to reflect neighboring words, the model may assign higher importance to the word standing next to what a human annotator may interpret as a signal. example (8) illustrates the problem: would like to emphasize that discourse parsers have entirely different goals, so that these numbers and parsing scores would be apples and oranges even if the same underlying corpus were used. 16 a neural approach to discourse relation signal detection (8) to provide information on the analytical sample as a whole , gold:purpose−−−−−−−−→ pred:preparation two additional demographic variables are included . each word in (8) is shaded based on the softmax probability assigned to the correct relation of the satellite, i.e. how ‘convincing’ the model found the word in terms of local probability. in addition, the top-scoring word in each sentence is rendered in boldface for emphasis. the gold label for the relation is placed above the arrow, which indicates the direction of the relation (satellite to nucleus), and the model’s predicted label is shown under the arrow. intuitively, the strongest signal of the purpose relation in (8) is the initial infinitive marker to – however, the model ranks the adjacent provide higher and almost ignores to. we suspect that the reason for this, and many similar examples in the model evaluated based on relation probabilities, is that contextual embeddings allow for a special representation of the word provide next to to, making it difficult to tease apart the locus of the most meaningful signal. to overcome this complication, we use the logic of permutation importance, treating the neural model as a black box and manipulating the input to discover relevant features in the data (cf. casalicchio et al. 2019). we reason that this type of evaluation is more robust than, for example, examining model internal attention weights because such weights are not designed or trained with a reward ensuring they are informative – they are simply trained on the same classification error loss as the rest of the model.13 instead, we can withhold potentially relevant information from the model directly: after training is complete, we feed the test data to the model in two forms – as-is, and with each word masked, as shown in (9). (9) original: to provide information ... ... masked1: provide information ... ... masked2: to information ... ... masked3: to provide ... ... label: purpose we reason that, if a token is important for predicting the correct label, masking it will degrade the model’s classification accuracy, or at least reduce its reported classification certainty.14 in (9), it seems reasonable to assume that masking the word ‘to’ has a greater impact on predicting the label purpose than masking the word ‘provide’, and even less so, the following noun ‘information’. we therefore use reduction in softmax probability of the correct relation as our signaling strength metric for the model. we call this metric ∆s (for delta-softmax), which can be written as: ∆rel,ti s = softmax(rel|xmask=φ)− softmax(rel|xmask=i) 13we thank an anonymous reviewer for pointing out wiegreffe and pinter (2019), who offer experimental support for the variable utility of attention as an explanatory metric. 14an anonymous reviewer has suggested that part of the model’s response to masking may be the result not only of missing discourse signals, but also of ungrammaticality due to missing words. while this could be the case to some extent, we would like to point out that the model should be somewhat accustomed to the masking situation since masked tokens are merely represented as oov word embeddings (which occur regularly at train and test time), and because drop out is applied to the model during training, meaning that the model is constantly exposed to some missing information. if the masked information were not crucial to relation classification, the model should still work correctly, whereas if an ungrammatical construction is truly preventing the identification of the relation, then that construction may well have been a signal, and the masking approach would be doing its job as intended. 17 zeldes and liu where rel is the true relation of the edu pair, ti represents the token at index i of n tokens, and xmask=i represents the input sequence with the masked position i (for i ∈ 1 . . . n ignoring separators, or φ, the empty set). to visualize the model’s predictions, we compare ∆s for a particular token to two numbers: the maximum ∆s achieved by any token in the current pair (a measure of relative importance for the current classification) and the maximum ∆s achieved by any token in the current document (a measure of how strongly the current relation is signaled compared to other relations in the text). we then shade each token 50% based on the first number and 50% based on the second. as a result, the most valid cues in an edu pair are darker than their neighbors, but edu pairs with no good cues are overall very light, whereas pairs with many good signals are darker. some examples of this visualization are given in (10)-(12) (human annotated endocentric signal tokens are marked by double underlines). (10) to provide information on the analytical sample as a whole , gold:purpose−−−−−−−−→ pred:preparation two additional demographic variables are included . (11) telling good jokes is an art that comes naturally to some people , gold:contrast←−−−−−−− pred:contrast but for others it takes practice and hard work . (12) it is possible that these two children understood the task and really did believe that the puppet did not produce any poor descriptions , and in this regard , are not yet adult-like in their si interpretations . gold:evaluation←−−−−−−−− pred:evaluation this is unlikely the highlighting in (10) illustrates the benefits of the masking based evaluation compared to (8): the token to is now clearly the strongest signal, and the verb is taken to be less important, followed by the even less important object of the verb. this is because removing the initial to hinders classification much more than the removal of the verb or noun. we note also that although the model in fact misclassified this example as preparation, we can still use masking importance to identify to, since the score queried from the model corresponds to a relative decrease in the probability of the correct relation, purpose, even if this was not the highest scoring relation overall. in (11) we see the model’s ability to correctly predict contrast based on the dm but. note that despite a rather long sentence, the model does not need any other word nearly as much for the classification. although the model is not trained explicitly to detect discourse markers, the dm can be recognized due to the fact that masking it leads to a drop of 66% softmax probability (∆s=0.66) of this pair representing the contrast relation. we can also note that a somewhat lower scoring content word is also marked: hard (∆s=0.18). in our gold signaling annotations, this word was marked together with comes naturally as a signal, due to the contrast between the two concepts (additionally, some people is flagged as a signal along with others). the fact that the model finds hard helpful, but does not need the contextual near antonym naturally, suggests that it is merely learning that words in the semantic space near hard may indicate contrast, and not learning about the antonymous relationship – otherwise we would expect to see ‘naturally’ have a stronger score (see also the discussion in section 6). finally (12) shows that, much like in the case of hard, the model is not biased towards traditional dms, confirming that it is capable of learning about content words, or neighborhoods of content words in vector space. in a long edu pair of 41 words, the model relies almost exclusively on the 18 a neural approach to discourse relation signal detection word unlikely (∆s=0.36) to correctly label the relation as evaluation. by contrast, the anaphoric demonstrative ‘this’ flagged by the human annotator, which is a more common function word, is disregarded, perhaps because it can appear with several other relations, and is not particularly exclusive to evaluation. these results suggest that the model may be capable of recognizing signals through distant supervision, allowing it to validate human annotations, to potentially point out signals that may be missed by annotators, and most importantly, to quantify signaling strength on a sliding scale. at the same time, we need a way to evaluate the model’s quality and assess the kinds of errors it makes, as well as what we can learn from them. we therefore move on to evaluating the model and its errors next. 5. evaluation and error analysis evaluation metric to evaluate the neural model, we would like to know how well ∆s corresponds to annotators’ gold standard labels. this leads to two kinds of problems: the first is that the model is distantly supervised, and therefore does not know about signal types, subtypes, or any aspect of signaling annotation and its relational structure. the second problem is that signaling annotations are categorical, and do not correspond to the ratio-scaled predictions provided by ∆s (this is in fact one of the motivations for desiring a model-based estimate of signaling strength). the first issue means that we can only examine the model’s ability to locate signals – not to classify them. although there may be some conceivable ways of analyzing model output to identify classes such as dms (which are highly lexicalized, rather than representing broad regions of vector space, as words such as unlikely might), or more contextual relational signals, such as pronouns, this line of investigation is beyond the scope of the present paper. a naive solution to the second problem might be to identify a cutoff point, e.g. deciding that all and only words scoring ∆s >0.15 are predicted to be signals. the problem with the latter approach is that sentences can be very different in many ways, and specifically in both length and in levels of ambiguity. sentences with multiple, mutually redundant cues, may produce lower ∆s scores compared to shorter sentences with a subset of the same cues. conversely, in very short sentences with low signal strength, the model may reasonably be expected to degrade very badly with the deletion of almost any word, as the context becomes increasingly incomprehensible. for these reasons, we choose to adopt an evaluation metric from the paradigm of information retrieval, and focus on recall@k (recall at rank k, for k = 1, 2, 3...). the idea is to poll the model for each sentence in which some signals have been identified, and see whether the model is able to find them if we let it guess using the word with the maximal ∆s score (recall@1), regardless of how high that score is, or alternatively relax the evaluation criteria and see whether the human annotator’s signal tokens appear at rank 2 or 3. figure 5 shows numbers for recall@k for the top 3 ranks outputted by the model, next to random guess baselines. the left, middle and right panels in figure 5 correspond to measurements when all signals are included, only cases contained entirely in the head edus shown to the model, and only dms, respectively. the scenario on the left is rather unreasonable and is included only for completeness: here the model is also penalized for not detecting signals such as lexical chains, part of which is outside the units that the model is being shown. an example of such a case can be seen in figure 6. the phrase respondents in unit [23] signals the relation elaboration, since it is coreferential with a previous mention of the respondents in [21]. however, because the model is only given 19 zeldes and liu .21 .33 .33 .48 .41 .59 .22 .40 .34 .53 .43 .64 .08 .56 .13 .62 .16 .67 all endocentric only dms r@1 r@2 r@3 r@1 r@2 r@3 r@1 r@2 r@3 0.0 0.2 0.4 0.6 metric re ca ll source model random figure 5: signal recall with 1, 2 or 3 guesses for the neural model and a random guess baseline. left: all signals included; middle: only endocentric cases, restricted to head edus shown to the model; right: only discourse markers (dms). heads of edu blocks to classify, it does not have access to the first occurrence of respondents while predicting the elaboration relation – the first half of the signal token set is situated in a child of the nucleus edu before the relation, i.e. it belongs to group iv in the taxonomy in table 3. realistically, our model can only be expected to learn about signals from ‘directly participating’ edus, i.e. groups i, ii, vi and vii, the ‘endocentric’ signal groups from section 3.2. although most signals belong to endocentric categories (71.62% of signaled relations belong to these groups, cf. table 3), exocentric cases form a substantial portion of signals which we have little hope of capturing with the architecture used here. as a result, recall metrics in the ‘all signals’ scenario are closest to the random baselines, though the signals detected in other instances still place the model well above the baseline. a more reasonable evaluation is the one in the middle panel of figure 5, which includes only endocentric signals as defined in the taxonomy. edus with no endocentric signals are completely disregarded in this scenario, which substantially reduces the number of tokens considered to be signals, since, while many tokens are part of some meaningful lexical chain in the document, requiring signals to be contained only in the pair of head units eliminates a wide range of candidates. although the random baseline is actually very slightly higher (perhaps because eliminated edus were often longer ones, sharing small amounts of material with larger parts of the text, and therefore prone to penalizing the baseline; many words mean more chances for a random guess to be wrong), model accuracy is substantially better in this scenario, reaching a 40% chance of hitting a signal 20 a neural approach to discourse relation signal detection figure 6: exocentric signal not detectable by the model. the elaboration pointing from unit [23] to [22] is signaled by a coreferential phrase appearing in another satellite of [22]. with only one guess, exceeding 53% with two guesses, and capping at 64% for recall@3, over 20 points above baseline.15 finally, the right panel in the figure shows recall when only dms are considered. in this scenario, a random guess fares very poorly, since most words are not dms. the model, by contrast, achieves the highest results in all metrics, since dms have the highest cue validity for relation classification, and the model attends to them most strongly. with just one guess, recall is over 56%, and goes as high as 67% for recall@3. the baseline only goes as high as 16% for three guesses. qualitative analysis looking at the model’s performance qualitatively, it is clear that it can detect not only dms, but also morphological cues (e.g. gerunds as markers of elaboration, as in (13)), semantic classes and sentiment, such as positive and negative evaluatory terms in (14), as well as multiple signals within the same edu, as in (15). in fact, only about 8.3% of the tokens correctly identified by the model in table 6 below are of the dm type, whereas about 7.2% of all tokens flagged by human annotators were dms, meaning that the model frequently matches non-dm items to discourse relation signals (see performance on signal types below). it should also be noted that signals can be recognized even when the model misclassifies relations, since ∆s does not rely on correct classification: it merely quantifies the contribution of a word in context toward the correct label’s score. if we examine the influence of each word on the score of the correct relation, that 15it should also be noted that a maximum score of 100% is not a reasonable expectation, since even humans disagree on signal tokens. as a ceiling figure representing human performance we can take the average mutual human recall score for signal tokens in gessler et al.’s (2019) data, a score of .84. 21 zeldes and liu impact should and does still correlate with human judgments based on what the system may tag as the second or third best class to choose. (13) for the present analysis , these responses were recoded into nine mutually exclusive categories gold:result←−−−−−−−− pred:elaboration capturing the following options : (14) professor eastman said he is alarmed by what they found . gold:evaluation−−−−−−−−→ pred:preparation ” pregnant women in australia are getting about half as much as what they require on a daily basis . (15) even so , estimates of the prevalence of perceived discrimination remains rare gold:concession←−−−−−−−− pred:evidence at least one prior study by kessler and colleagues [ 15 ] , however , using measures of perceived discrimination in a large american sample , reported that approximately 33 % of respondents reported some form of discrimination unsurprisingly, the model sometimes make sporadic errors in signal detection for which good explanations are hard to find, especially when its predicted relation is incorrect, as in (16). here the evaluative adjective remarkable is missed in favor of neighboring words such as agreed and a subject pronoun, which are not indicative of the evaluation relation in this context but are part of several cohorts of high scoring words. however, the most interesting and interpretable errors arise when ∆s scores are high compared to an entire document, and not just among words in one edu pair, in which most or even all words may be relatively weak signals. as an example of such a false positive with high confidence, we can consider (17). in this example, the model correctly assigns the highest score to the dm so marking a purpose relation. however, it also picks up on a recurring tendency in how-to guides in which the second person pronoun referring to the reader is often the benefactee of some action, which contributes to the purpose reading and helps to disambiguate so, despite not being considered a signal by annotators. (16) the agreement was that gorbachev agreed to a quite remarkable concession : gold:evaluation−−−−−−−−→ pred:preparation he agreed to let a united germany join the nato military alliance . (17) the opening of the joke or setup should have a basis in the real world gold:purpose←−−−−−−− pred:purpose so your audience can relate to it , in other cases, the model points out plausible signals which were passed over by an annotator, and may be considered errors in the gold standard. for example, the model easily notices that question marks indicate the solutionhood relation, even where these were skipped by annotators in favor of marking wh words instead: (18) which previous virginia governor(s) do you most admire and why ? gold:solutionhood−−−−−−−−−→ pred:solutionhood thomas jefferson . from the model’s perspective, the question mark, which scores ∆s=0.79, is the single most important signal, and virtually sufficient for classifying the relation correctly, though it was left out of the gold annotations. the wh word which and the sentence final why, by contrast, were noticed by annotators but were are not as unambiguous (the former could be a determiner, and the latter 22 a neural approach to discourse relation signal detection in sentence final position could be part of an embedded clause). in the presence of the question mark, their individual removal has much less impact on the classification decision. although the model’s behavior is sensible and can reveal annotation errors, it also suggests that ∆s will be blind to auxiliary signals in the presence of very strong, independently sufficient cues. using the difference in likelihood of correct relation prediction as a metric also raises the possibility of an opposite concept to signals, which we will refer to as distractors. since ∆s is a signed measure of difference, it is in fact possible to obtain negative values whenever the removal or masking of a word results in an improvement in the model’s ability to predict the relation. in such cases, and especially when the negative value is of a large magnitude, it seems like a reasonable interpretation to say that a word functions as a sort of anti-signal, preventing or complicating the recognition of what might otherwise be a more clear-cut case. examples (19)–(20) show some instances of distractors identified by the masking procedure (distractors with ∆s <-0.2 are underlined). (19) how do they treat those not like themselves ? gold:preparation−−−−−−−−−→ pred:solutionhood then they ’re either overzealous , ignorant of other people or what to avoid those that contradict their fantasy land that caters to them and them only . (20) god , i do n’t know ! gold:preparation−−−−−−−−→ pred:preparation but nobody will go to fight for noses any more . in (19), a rhetorical question trips up the classifier, which predicts the question-answer relation solutionhood instead of preparation. here the initial wh word how and the subsequent auxiliary do-support both distract (with ∆s=-0.23 and -0.25) from the preparation relation, which is however being signaled positively by the dm then in the nucleus unit. later on, the adverb only is also disruptive (∆s=-0.31), perhaps due to a better association with adversative relations, such as contrast. in (20), a preparatory “god, i don’t know!” is followed up with a nucleus starting with but, which typically marks a concession or other adversative relation. in fact, the dm but is related to a concessive relation with another edu (not shown), which the model is not aware of while making the classification for the preparation. although this example reveals a weakness in the model’s inability to consider broader context, it also reveals the difficulty of expecting dms to fall in line with a strong nuclearity assumption: since units serve multiple functions as satellites and nuclei, signals which aid the recognition of one relation may hinder the recognition of another. performance on signal types to better understand the kinds of signals which the model captures better or worse, table 6 gives a breakdown of performance by signal type and specific signal categories, for categories attested over 20 times (note that the categories are human labels assigned to the corresponding positions – the system does not predict signal types). to evaluate performance for all types we cannot use recall@1–3, since some sentences contain more than 3 signal tokens, which would lead to recall errors even if the top 3 ranks are correctly identified signals. the scores in the table therefore express how many of the signal tokens belonging to each subtype in the gold annotations are recognized if we allow the system to make as many guesses as there are signal tokens in each edu pair, plus a tolerance of a maximum of 2 additional tokens (similarly to recall@3). we also note that a single token may be associated with multiple signal types, in which case its identification or omission is counted separately for each type. 23 zeldes and liu type subtype accuracy ntest lexical alternate expression 0.773 53 semantic synonymy 0.727 22 lexical indicative word 0.726 340 graphical colon 0.714 21 morphological tense 0.711 59 semantic antonymy 0.690 42 dm dm 0.653 257 semantic repetition 0.642 350 numerical same count 0.617 34 semantic + syntactic meronymy + subject np 0.567 37 semantic + syntactic lexical chain + subject np 0.566 83 semantic lexical chain 0.549 1453 semantic meronymy 0.545 121 semantic indicative word pair 0.528 70 semantic + syntactic repetition + subject np 0.505 99 reference personal reference 0.461 206 morphological modality 0.461 52 reference + syntactic propositional reference + subject np 0.409 22 reference demonstrative reference 0.380 84 syntactic + semantic parallel syntactic construction + lexical chain 0.333 99 textual date 0.318 22 reference propositional reference 0.285 21 table 6: anchored token detection accuracy for signal types attested over 20 times. three of the top four categories which the model performs best for are, perhaps unsurprisingly, the most lexical ones: alternate expression captures non-dm phrases such as i mean (for elaboration), or the problem is (for concession), and indicative word includes lexical items such as imperative see (consistently marking evidence in references within academic articles) or evaluative adjectives such as interesting for evaluation. the good performance of the category colon captures the model’s recognition of colons as important punctuation, primarily predicting preparation. the only case of a ‘relational’ category, requiring attention to two separate positions in the input, which also fares well is synonymy, though this is often based on flagging only one of two items annotated as synonymous, and is based on rather few examples. we can find only one example, (21), where both sides of a pair of similar words is actually noticed, which both belong to the same stem (decline/declining): (21) the report says the decline in iodine intake appears to be due to changes in the dairy industry , where chlorine-containing sanitisers have replaced iodine-containing sanitisers . gold:justify←−−−−−−−−− pred:background iodine released from these chemicals into milk has been the major source of dietary iodine in australia for at least four decades , but is now declining . we note that our evaluation is actually rather harsh towards the model, since in multiword expressions, often only one central word is flagged by ∆s (e.g. problem in “the problem is”), while 24 a neural approach to discourse relation signal detection the model is penalized in table 6 for each token that is not recognized (i.e. the and is, which were all flagged by a human annotator as signals in the data). interestingly, the model fares rather well in identifying morphological tense cues, even though these are marked by both inflected lexical verbs and semantically poor auxiliaries (e.g. past perfect auxiliary had marking background); but modality cues (especially can or could for evaluation) are less successfully identified, suggesting they are either more ambiguous, or mainly relevant in the presence of evaluative content words which out-score them. other relational categories from the middle of the table which ostensibly require matching pairs of words, such as repetition, meronymy, or personal reference (coreference) are mainly captured by the model when a single item is a sufficiently powerful cue, often ignoring the other half of the signal, as shown in (22). (22) on a new website , ” the internet explorer 6 countdown ” , microsoft has launched an aggressive campaign to persuade users to stop using ie6 gold:elaboration←−−−−−−−− pred:elaboration its goal is to decrease ie6 users to less than one percent . here the model has learned that an initial possessive pronoun, perhaps in the context of a subject np in a copula sentence (note the shading of the following is) is an indicator of an elaboration relation, even though there is no indication that the model has noticed which word is the antecedent. similarly for the count category, the model only learns to notice the possible importance of some numbers, but is not actually aware of whether they are identical (e.g. for restatement) or different (e.g. in contrast).16 finally, some categories are actually recognized fairly reliably, but are penalized by the same partial substring issue identified above: date expressions are consistently flagged as indicators of circumstance, but often a single word, such as a weekday in (23), is dominant, while the model is penalized for not scoring other words as highly (including commas within dates, which are marked as part of the signal token span in the gold standard, but whose removal does not degrade prediction accuracy). in this case it seems fair to say that the model has successfully recognized the date signal of ‘wednesday april 13’, yet it loses points for missing two instances of ‘,’, and the ‘2011’, which is no longer necessary for recognizing that this is a date. (23) nasa celebrates 30th anniversary of first shuttle launch ; gold:circumstance←−−−−−−−−− pred:circumstance wednesday , april 13 , 2011 6. discussion this paper has used a corpus annotated for discourse relation signals within the framework of the rst signalling corpus (das and taboada 2017) and extended with anchored signal annotations (liu 2019) to develop a taxonomy of unrestricted and hierarchically aware discourse signal positions, as well as a data-driven neural network model to explore distantly supervised signal word extraction. the results shed light on the distribution of signal categories from the rst-sc taxonomy in terms of associated word forms, and show the promise of neural models with contextual embeddings for 16we note that while the original signalling corpus only used a same count specific signal category, the data set we are working with uses a similar specific signal type for any numerical signals, including first(ly) and second(ly), which the model is particularly successful in learning. 25 zeldes and liu the extraction of context dependent and gradient discourse signal detection in individual texts. the metric developed for the evaluation, ∆s, allows us to assess the relative importance of signal words for automatic relation classification, and reveal observations for further study, as well as shortcomings which point to the need to develop richer feature representations and system architectures in future work. the model presented in the previous sections is clearly incomplete in both its classification accuracy and its ability to recognize the same signals that humans do. however, given the fact that it is trained entirely without access to discourse signal annotations and is unaware of any of the guidelines used to create the gold standard that it is evaluated on, its performance may be considered surprisingly good. as an approach to extracting discourse signals in a data-driven way, similar to frequentist methods or association measures used in previous work, we suggest that this model forms a more fine grained tool, capable of taking context into consideration and delivering scores for each instance of a signal candidate, rather than resulting in a table of undifferentiated signal word types. additionally, although we consider human signal annotations to be the gold standard in identifying the presence of relevant cues, the ∆s metric gives new insights into signaling which cannot be approached using manual signaling annotations. firstly, the quantitative nature of the metric allows us to rank signaling strength in a way that humans have not to date been able to apply: using ∆s, we can say which instances of which signals are evaluated as stronger, by how much, and which words within a multi-word signal instance are the most important (e.g. weekdays in dates are important, the commas are not). secondly, the potential for negative values of the metric opens the door to the study of negative signals, or ‘distractors’, which we have only touched upon briefly in this paper. and finally, we consider the availability of multiple measurements for a single dm or other discourse signal to be a potentially very interesting window into the relative ambiguity of different signaling devices (cf. torabi asr and demberg 2013) and for research on the contexts in which such ambiguity results. to see how ambiguity is reflected in multiple measurements of ∆s, we can consider figure 7. the figure shows boxplots for multiple instances of the same signal tokens. we can see that words like and are usually not strong signals, with the entire interquartile range scoring less than 0.02, i.e. aiding relation classification by less than 2%, with some values dipping into the negative region (i.e. cases functioning as distractors). however, some outliers are also present, reaching almost as high as 0.25 – these are likely to be coordinating predicates, which may signal relations such as sequence or joint. a word such as but is more important overall, with the box far above and, but still covering a wide range of values: these can correspond to more or less ambiguous cases of but, but also to cases in which the word is more or less irreplaceable as a signal. in the presence of multiple signals for the same relation, the presence of but should be less important. we can also see that but can be a distractor with negative values, as we saw in example (20) above. as far as we are aware, this is the first empirical corpus-based evidence giving a quantitative confirmation to the intuition that ‘but’ in context is significantly less ambiguous as a discourse marker than ‘and’; the overlap in their bar plots indicate that they can be similarly ambiguous or even distracting in some cases, but the difference in interquartile ranges makes it clear that these are exceptions. for less ambiguous dms, such as if, we can also see a contrast between lower and upper case instances: upper case if is almost always a marker of condition, but the lower case if is sometimes part of an embedded object clause, which is not segmented in the corpus and does not mark a conditional relation (e.g. “they wanted to see if...”). for the word to, the figure suggests a strongly 26 a neural approach to discourse relation signal detection 0.00 0.25 0.50 0.75 1.00 and but if if to word ∆ s figure 7: boxplots for all ∆s values of several signal token types across the test corpus. bimodal distribution, with a core population of (primarily prepositional) discourse-irrelevant to, and a substantial number of outliers above a large gap, representing to in infinitival purpose clauses (though not all to infinitives mark such clauses, as in adnominal “a chance to go”, which the model is usually able to distinguish in context). in other words, our model can not only disambiguate ambiguous strings into grammatical categories, but also rank members of the same category by importance in context, as evidenced by its ability to correctly classify high frequency items like ‘to’ or ‘and’ as true positives. a frequentist approach would not only lack this ability – it would miss such items altogether, due to its overall high string frequency and low specificity. beyond what the results can tell us about discourse signals in this particular corpus, the fact that the neural model is sensitive to mutual redundancy of signals raises interesting theoretical questions about what human annotators are doing when they characterize multiple features of a discourse unit as signals. if it is already evident from the presence of a conventional dm that some relation applies, are other, less explicit signals which might be relied on in the absence of the dm, equally ‘there’? do we need a concept of primary and auxiliary signals, or graded signaling strength, in the way that a metric such as ∆s suggests? another open question relates to the postulation of distractors as an opposite concept to discourse relation signals. while we have not tested this so far, it is interesting to ask to what extent human analysts are aware of distractors, whether we could form annotation guidelines to recognize them, and how humans weigh the value of signals and potential distractors in extrapolating intended discourse relations. it seems likely that distractors affecting humans may be found in cases of misunderstanding or ambiguity of discourse relations (see also da cunha 2013). finally, the error analysis for signal detection complements the otherwise opaque relation classification results in table 5 in showing some of the missing sources of information that our model would need in order to work better. we have seen that relational information, such as identifying not just the presence of a pronoun but also its antecedent, or both sides of lexical semantic relations such 27 zeldes and liu as synonymy, meronymy or antonymy, as well as comparing count information, are still unavailable to the classifier – if they were being used, then ∆s would reflect the effects of their removal, but this is largely not the case. this suggests that, in the absence of vastly larger discourse annotated corpora, discourse relation recognition may require the construction of either features, architectures, or both, which can harness abstract relational information of this nature beyond the memorization of specific pairs of words (or regions of vector space with similar words) that are already attested in the limited training data. in this vein, pitler et al. (2009) conducted a series of experiments on automatic sense prediction for four top-level implicit discourse relations within the pdtb framework, which also suggested benefits for using linguistically-informed features such as verb information, polarity tags, context, lexical items (e.g. first and last words of the arguments; first three words in the sentence) etc. the model architecture and input data are also in need of improvements, as the current architecture can only be expected to identify endocentric signals. the substantial amount of exocentric signaling cases is in itself an interesting finding, as it suggests that relation classification from head edu pairs may ultimately have a natural ceiling that is considerably below what could be inferred from looking at larger contexts. we predict that as we add more features to the model and improve its architecture in ways that allow it to recognize the kinds of signals that humans do, classification accuracy will increase; and conversely, as classification accuracy rises, measurements based on ∆s will overlap increasingly with human annotations of anchored signals. in sum, we believe that there is room for much research on what relation classification models should look like, and how they can represent the kinds of information found in non-trivial signals. the results of this line of work can therefore benefit nlp systems targeting discourse relations by suggesting locations within the text which systems should attend to in one way or another. moreover, we think that using distant-supervised techniques for learning discourse relations (e.g. huber and carenini 2019) is promising in the development of discourse models using the proposed dataset. we hope to see further analyses benefit from this work and the application of metrics such as ∆s to other datasets, within more complex models, and using additional features to capture such information. we also hope to see applications of discourse relations such as machine comprehension (narasimhan and barzilay, 2015) and sentiment analysis (huber and carenini, 2019) etc. benefit from the proposed model architecture as well as the dataset. references alan akbik, tanja bergmann, and roland vollgraf. pooled contextualized embeddings for named entity recognition. in proceedings of naacl 2019, pages 724–728, minneapolis, mn, 2019. farah benamara, maite taboada, and yannick mathieu. evaluative language beyond bags of words: linguistic insights and computational applications. computational linguistics, 43(1):201–264, 2017. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in current and new directions in discourse and dialogue, text, speech and language technology 22, pages 85–112. kluwer, dordrecht, 2003. giuseppe casalicchio, christoph molnar, and bernd bischl. visualizing the feature importance for black box models. in michele berlingerio, francesco bonchi, thomas gärtner, neil hurley, 28 a neural approach to discourse relation signal detection and georgiana ifrim, editors, machine learning and knowledge discovery in databases. ecml pkdd 2018, lecture notes in computer science 11051. springer, cham, 2019. iria da cunha. a symbolic corpus-based approach to detect and solve the ambiguity of discourse markers. research in computing science, 70:95–106, 2013. debopam das and maite taboada. signalling of coherence relations in discourse, beyond discourse markers. discourse processes, pages 1–29, 2017. debopam das and maite taboada. rst signalling corpus: a corpus of signals of coherence relations. language resources and evaluation, 52(1):149–184, 2018. debopam das, maite taboada, and paul mcfetridge. rst signalling corpus. ldc2015t10, 2019. eldadio duque. signaling causal coherence relations. discourse studies, 16(1):25–46, 2014. gill francis. labelling discourse: an aspect of nominal-group lexical cohesion. in malcolm coulthard, editor, advances in written text analysis, pages 83–101. routledge, 1994. bruce fraser. what are discourse markers? journal of pragmatics, 31:931–952, 1999. matt gardner, joel grus, mark neumann, oyvind tafjord, pradeep dasigi, nelson f. liu, matthew peters, michael schmitz, and luke s. zettlemoyer. allennlp: a deep semantic natural language processing platform. in proceedings of acl 2018, pages 1–6, melbourne, 2018. luke gessler, yang liu, and amir zeldes. a discourse signal annotation system for rst trees. in proceedings of the workshop on discourse relation parsing and treebanking (disrpt 2019), pages 56–61, minneapolis, mn, 2019. reza ghaeini, xiaoli z. fern, and prasad tadepalli. interpreting recurrent and attention-based neural models: a case study on natural language inference. in proceedings of emnlp 2018, pages 4952–4957, brussels, 2018. eduard h. hovy and elisabeth maier. parsimonious or profligate: how many and which discourse structure relations? technical report, information sciences institute, usc, 1993. patrick huber and giuseppe carenini. predicting discourse structure using distant supervision from sentiment. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), pages 2306–2316, hong kong, china, november 2019. alistair knott. a data-driven methodology for motivating a set of coherence relations. phd thesis, university of edinburgh, 1996. alistair knott and robert dale. using linguistic phenomena to motivate a set of coherence relations. discourse processes, 18(1):35–52, 1994. alistair knott and ted sanders. the classification of coherence relations and their linguistic markers: an exploration of two languages. journal of pragmatics, 30(2):135–175, 1998. 29 zeldes and liu geoffrey leech, tony mcenery, and martin weisser. spaac speech-act annotation scheme. technical report, lancaster university, 2003. yang liu. beyond the wall street journal: anchoring and comparing discourse signals across genres. in proceedings of the workshop on discourse relation parsing and treebanking (disrpt 2019), pages 72–81, minneapolis, mn, 2019. yang liu and amir zeldes. discourse relations and signaling information: anchoring discourse signals in rst-dt. proceedings of the society for computation in linguistics, 2(1):314–317, 2019. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text, 8(3):243–281, 1988. daniel marcu. building up rhetorical structure trees. in proceedings of aaai-96, pages 1069–1074, portland, or, 1996. mitchell p. marcus, beatrice santorini, and mary ann marcinkiewicz. building a large annotated corpus of english: the penn treebank. special issue on using large corpora, computational linguistics, 19(2):313–330, 1993. jiřı́ mı́rovský, pavlı́na synková, magdaléna rysová, and lucie poláková. czedlex a lexicon of czech discourse connectives. the prague bulletin of mathematical linguistics, 109:61–91, 2017. mathieu morey, philippe muller, and nicholas asher. how much progress have we made on rst discourse parsing? a replication study of recent results on the rst-dt. in proceedings of emnlp 2017, pages 1319–1324, copenhagen, denmark, 2017. philippe muller, chloé braud, and mathieu morey. tony: contextual embeddings for accurate multilingual discourse segmentation of full documents. in proceedings of discourse relation treebanking and parsing (disrpt 2019), pages 115–124, minneapolis, mn, 2019. karthik narasimhan and regina barzilay. machine comprehension with discourse relations. in proceedings of the 53rd annual meeting of the association for computational linguistics and the 7th international joint conference on natural language processing (acl/ijcnlp 2015), pages 1253–1262, beijing, china, 2015. joakim nivre, željko agić, lars ahrenberg, maria jesus aranzabe, masayuki asahara, aitziber atutxa, miguel ballesteros, john bauer, kepa bengoetxea, riyaz ahmad bhat, eckhard bick, cristina bosco, gosse bouma, sam bowman, marie candito, gülşen cebiroğlu eryiğit, giuseppe g. a. celano, fabricio chalub, jinho choi, çağrı çöltekin, miriam connor, elizabeth davidson, marie-catherine de marneffe, valeria de paiva, arantza diaz de ilarraza, kaja dobrovoljc, timothy dozat, kira droganova, puneet dwivedi, marhaba eli, tomaž erjavec, richárd farkas, jennifer foster, cláudia freitas, katarı́na gajdošová, daniel galbraith, marcos garcia, filip ginter, iakes goenaga, koldo gojenola, memduh gökırmak, yoav goldberg, xavier gómez guinovart, berta gonzáles saavedra, matias grioni, normunds grūzı̄tis, bruno guillaume, nizar habash, jan hajič, linh hà mỹ, dag haug, barbora hladká, petter hohle, radu ion, elena irimia, anders johannsen, fredrik jørgensen, hüner kaşıkara, hiroshi 30 a neural approach to discourse relation signal detection kanayama, jenna kanerva, natalia kotsyba, simon krek, veronika laippala, lê h`ông, alessandro lenci, nikola ljubešić, olga lyashevskaya, teresa lynn, aibek makazhanov, christopher manning, cătălina mărănduc, david mareček, héctor martı́nez alonso, andré martins, jan mašek, yuji matsumoto, ryan mcdonald, anna missilä, verginica mititelu, yusuke miyao, simonetta montemagni, amir more, shunsuke mori, bohdan moskalevskyi, kadri muischnek, nina mustafina, kaili müürisep, luong nguy˜ên thi., huy`ên nguy˜ên thi. minh, vitaly nikolaev, hanna nurmi, stina ojala, petya osenova, lilja øvrelid, elena pascual, marco passarotti, cenel-augusto perez, guy perrier, slav petrov, jussi piitulainen, barbara plank, martin popel, lauma pretkalniņa, prokopis prokopidis, tiina puolakainen, sampo pyysalo, alexandre rademaker, loganathan ramasamy, livy real, laura rituma, rudolf rosa, shadi saleh, manuela sanguinetti, baiba saulı̄te, sebastian schuster, djamé seddah, wolfgang seeker, mojgan seraji, lena shakurova, mo shen, dmitry sichinava, natalia silveira, maria simi, radu simionescu, katalin simkó, mária šimková, kiril simov, aaron smith, alane suhr, umut sulubacak, zsolt szántó, dima taji, takaaki tanaka, reut tsarfaty, francis tyers, sumire uematsu, larraitz uria, gertjan van noord, viktor varga, veronika vincze, jonathan north washington, zdeněk v zabokrtský, amir zeldes, daniel zeman, and hanzhi zhu. universal dependencies 2.0. technical report, lindat/clarin digital library at the institute of formal and applied linguistics (úfal), faculty of mathematics and physics, charles university, 2017. url http: //hdl.handle.net/11234/1-1983. jeffrey pennington, richard socher, and christopher d. manning. glove: global vectors for word representation. in proceedings of emnlp 2014, pages 1532–1543, doha, qatar, 2014. matthew e. peters, mark neumann, mohit iyyer, matt gardner, christopher clark, kenton lee, and luke zettlemoyer. deep contextualized word representations. in proceedings of naacl 2018, pages 2227–2237, new orleans, la, 2018. emily pitler, annie louis, and ani nenkova. automatic sense prediction for implicit discourse relations in text. in proceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural language processing of the afnlp, pages 683–691, suntec, singapore, 2009. massimo poesio, barbara di eugenio, and gerard keohane. discourse structure and anaphora: an empirical study. nle technical note tn-02-02, university of essex, department of computer science, nle group, 2002. richard power, christine doran, and donia scott. generating embedded discourse markers from rhetorical structure. in proceedings of ewnlg 1999, pages 30–38, toulouse, 1999. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of the 6th international conference on language resources and evaluation (lrec 2008), pages 2961–2968, marrakesh, morocco, 2008. rashmi prasad, bonnie webber, and aravind joshi. reflections on the penn discourse treebank, comparable corpora, and complementary annotation. computational linguistics, 40(4):921–950, 2014. 31 zeldes and liu rashmi prasad, bonnie webber, alan lee, and aravind joshi. penn discourse treebank version 3.0. ldc2019t05, 2019. craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. semantics and pragmatics, 5(6):1–69, december 2012. tatjana scheffler and manfred stede. adding semantic relations to a large-coverage connective lexicon of german. in proceedings of lrec 2016, pages 1008–1013, reykjavik, iceland, 2016. manfred stede. discourse processing. synthesis lectures on human language technologies 4. morgan & claypool, [san rafael, ca], 2012. manfred stede and carla umbach. dimlex: a lexicon of discorse markers for text generation and understanding. in proceedings of coling-acl 1998, pages 1238–1242, montreal, 1998. maite taboada. discourse markers as signals (or not) of rhetorical relations. journal of pragmatics, 38(4):567–592, 2006. maite taboada and debopam das. annotation upon annotation: adding signalling information to a corpus of discourse relations. dialogue and discourse, 4(2):249–281, 2013. maite taboada and marı́a de los ángeles gómez-gonzález. discourse markers and coherence relations: comparison across markers, languages and modalities. linguistics and the human sciences, 6(1–3):17–41, 2012. svetlana toldova, dina pisarevskaya, margarita ananyeva, maria kobozeva, alexander nasedkin, sofia nikiforova, irina pavlova, and alexey shelepov. rhetorical relation markers in russian rst treebank. in proceedings of the 6th workshop recent advances in rst and related formalisms, pages 29–33, santiago de compostela, spain, 2017. fatemeh torabi asr and vera demberg. on the information conveyed by discourse markers. in proceedings of the fourth annual workshop on cognitive modeling and computational linguistics (cmcl), pages 84–93, sofia, bulgaria, 2013. yannick versley and anna gastel. linguistic tests for discourse relations in the tüba-d/z corpus of written german. dialogue and discourse, 4(2):142–173, 2013. bonnie lynn webber and aravind k joshi. anchoring a lexicalized tree-adjoining grammar for discourse. in proceedings of sigdial 1998, pages 86–92, montreal, 1998. sarah wiegreffe and yuval pinter. attention is not not explanation. in proceedings of emnlp 2019, pages 11–20, hong kong, china, 2019. yue yu, yilun zhu, yang liu, yan liu, siyao peng, mackenzie gong, and amir zeldes. gumdrop at the disrpt2019 shared task: a model stacking approach to discourse unit segmentation and connective detection. in proceedings of discourse relation treebanking and parsing (disrpt 2019), pages 133–143, minneapolis, mn, 2019. amir zeldes. the gum corpus: creating multilayer resources in the classroom. language resources and evaluation, 51(3):581–612, 2017. 32 a neural approach to discourse relation signal detection amir zeldes. multilayer corpus studies. routledge advances in corpus linguistics 22. routledge, london, 2018a. amir zeldes. a neural approach to discourse relation signaling. in georgetown university round table (gurt) 2018: approaches to discourse, washington, dc, 2018b. deniz zeyrek, işin demirşahin, ayışığı b. sevdik-çallı, and ruket çakıcı. turkish discourse bank: porting a discourse annotation style to a morphologically rich language. dialogue and discourse, 4(2):174–184, 2013. yuping zhou, jill lu, jennifer zhang, and nianwen xue. chinese discourse treebank 0.5 ldc2014t21, 2014. 33 journal of machine learning research-microsoft word template dialogue & discourse 12(2) (2021) 174-191 doi: 10.5210/dad.2021.206 an analysis of japanese sentence-final particle yone: compare yone and ne in response jun xu jun.xu@colostate.edu colorado state university, fort collins, usa editor: kallirroi georgila submitted 01/2020; accepted 10/2021; published online 12/2021 abstract yone, a japanese sentence-final particle (sfp), is frequently used in conversation, and some functions overlap with ne, another sfp. however, not much discussion has taken place about their differences. this study argues that the two japanese sentence-final particles, yone and ne, express a distinction about the speaker’s state of mind: yone indicates that an idea has been on the speaker’s mind, while ne suggests a thought just emerged into the speaker’s awareness. naturally occurring conversation data provides evidence for this claim. the results show that the particles reflect the speaker’s choice of presenting his/her state of awareness. keywords: japanese, discourse, pragmatics, sentence-final particle, yone 1 introduction ne and yone are two frequently used japanese sentence-final particles in conversation (e.g., asanocavanagh, 2011; cook, 1990, 1992; hasegawa, 2010; hasunuma, 1992, 1995; hayano, 2011, 2013; izuhara, 1993, 1994, 2001, 2003, 2008; kamio, 1994, 1995, 1997; katagiri, 2007; kato, 2001; lee, 2007; maynard, 1993; miyazaki 2002; morita, 2002; saigo, 2011; takubo & kinsui, 1997; tanaka, 2000; xu, 2016; zhang, 2009). while both yone and ne can be used as tag-like questions to seek confirmation, they can also show agreement or empathy (e.g., hasunuma, 1992, 1995; izuhara, 1993, 1994, 2001, 2003; noda, 1993). although yone and ne functions overlap, interchangeability has not been thoroughly studied (asano-cavanagh, 2011; hayano, 2013; noda, 1993; xu, 2016). also, as hayano (2011) points out, research on japanese sentence-final particles tends to consider that their use is rule-governed by the objectively discernible distribution of knowledge or information. however, such perspective overlooks the agency of the individual in presenting knowledge of information when using sentence-final particles through the dynamic courses of conversation (hayano, 2011). using naturally occurring conversation data where yone and ne occur in a response environment, the present study attempts to show these differences between yone and ne: a. yone is a marker of existing awareness that indicates an idea had already been on the speaker’s mind. b. ne is a marker of new awareness that shows a thought just emerged into the speaker’s awareness. c. what ne presents is not necessarily something new, and it can be something the speaker had not thought of until the time of utterance. d. speakers systematically select yone or ne to present their states of awareness. it is necessary here to clarify three critical terms in this study: “response,” “new awareness,” and “existing awareness.” first, according to the merriam-webster dictionary, a response is xu 175 something constituting a reply or a reaction.1 in other literature, a response can appear as the answer part in a summons-answer sequence (schegloff, 2007), an initial assessment of something, or the second assessment in an assessment pair (pomerantz, 1984). in this paper, the term “response” is used in a broad sense to examine four situations: (i) response to immediate conditions; (ii) response in recall; (iii) response to the answer to a question; and (iv) response to assessment.2 second, throughout this paper, the term “new awareness” is used to describe what a speaker currently thinks, feels, or what the speaker has just noticed or recalled at the time of the utterance. third, “existing awareness” refers to the speaker’s knowledge or thoughts, such as information mentioned in previous discourse or knowledge already established before the moment of utterance. worth noting is that the present study does not intend to provide a general account for the distinction between yone and ne. instead, this examination is limited to only four situations in response. however, they reveal how a speaker systematically displays one’s state of awareness with ne and yone. this article follows this organizational structure: the first section deals with the background of this study. the second part concerns the data and methodology of this study. data analysis constitutes the third part, and the fourth part concludes this study. 2 background this section provides a brief review of the overlapped functions of yone and ne, the differences between them in response as showing agreement, and the relationship between information status and the choice of linguistic forms in japanese. 2.1 the differences between yone and ne in response the literature has identified that yone has a few overlapped functions with ne (e.g., asanocavanagh, 2011; hasunuma, 1992, 1995; mcgloin et al., 2013; miyazaki, 2002; noda, 1993; zhang, 2009). this study focuses on ne and yone as response, which appears as agreement in examples (1) and (2).3 here, responding to the assessment in line 1, speaker b can use either ne or yone in line 2 to show the agreement with speaker a’s opinion, i.e., this restaurant’s sushi is delicious. (1) 1 a: kono mise no sushi oishii naa this restaurant lk sushi delicious fp 2 b: → un oishii ne itj delicious fp (2) 1 a: kono mise no sushi oishii naa this restaurant lk sushi delicious fp 1 response [def. 2]. (n.d.). in merriam webster online, retrieved april 18, 2021, from http://www.merriamwebster.com/dictionary/response. 2 assessment is an action or an instant judgment about something. it is an interactive action commonly examined in conversation analysis (pomerantz, 1984). she points out that “assessments are produced as products of participation; with an assessment, a speaker claims knowledge that which he or she is assessing” (pomerantz, 1984:57). in the present study, ne-marked or yone-marked responses commonly appear as assessments. however, a response does not always appear as an assessment. assessment is one category of response in this article. 3 see appendix for the list of abbreviations and transcription conventions. japanese sentence-final particle yone 176 2 b: → un oishii yone itj delicious fp to date, the differences between using yone and ne in examples (1) and (2) have received scant attention in the literature. only noda (1993) has discussed this issue by pointing out that yone and ne are not always interchangeable. she argues that in certain situations, one rather than the other is natural. for instance, ne is acceptable in example (3) while yone cannot be used at the same turn in example (4). noda (1993:14) argues that yone cannot be used in greetings, so (4)b is incorrect. (3) 1 a: ii tenki desu ne good weather cop fp 2 b: → soo desu ne that cop fp (4) 1 a: ii tenki desu ne good weather cop fp 2 ?b: → soo desu yone 4 that cop fp however, noda’s (1993) proposal fails to explain why yone and ne are not interchangeable in example (5) even though they are in a non-greeting situation. in example (5), the participants are talking about hourly pay at their part-time jobs. after finding out speaker b’s hourly pay rate, speaker a comments that the pay rate is low (line 1). speaker b shows his agreement with speaker a in line 2. in line 3, when speaker a further displays his response that the pay is low, only yone is possible. if ne is used in the same position (example 5, line 3), the whole utterance becomes unnatural. (5) 1 a: yasukunai↑ cheap-neg 2 b: yasui cheap 3 a: → yasui yone cheap fp surprisingly, no studies have discussed yone and ne as used in example (5). because noda’s (1993) proposal cannot successfully explain the use of yone in example (5), this study attempts to show the difference between yone and ne through a new approach. 4 this sentence was judged as incorrect with a “*” mark by noda (1993), but it is replaced with “?” in this study. five japanese native speakers were consulted about the naturalness of this example. one reported that the sentence 4(b) is unnatural. while the other four considered it acceptable, all agreed that the sentence 3(b) is more natural. what is important when changing the “*” mark to the “?” mark for 4(b), however, is that 3(b) is more appropriate than 4(b) in terms of naturalness even though native speakers might have different opinions about accepting sentence 4(b). xu 177 2.2 information status and the choice of linguistic forms in japanese in many languages, speakers choose different linguistic features to present the various statuses of information (e.g., akatsuka, 1985; goldsmith & woisetschlaeger, 1982; gundel, 1985; kamio, 1997; kuno, 1972; miyazaki 2002; prince, 1992; slobin & aksu, 1982). for instance, in japanese, kuno (1972) uses the notion of “old/new information” to explain the use of two particles, wa and ga. he proposes four different usages of wa and ga: thematic wa, contrastive wa, descriptive ga, and exhaustive-listing ga. he further points out that thematic wa presents old information while descriptive ga and exhaustive-listing ga convey new information. akatsuka (1985) illustrates the vital role a speaker’s awareness plays in the choice of forms for conditionals. she argues that an understanding of what registers the speaker’s awareness at the time of the utterance is pivotal in distinguishing the conditional s1 no nara s2 from other conditionals in japanese. when s1 no nara s2 appears, s1 always expresses new information that has just entered the speaker’s awareness at the discourse site. she also notes that japanese grammar is sensitive to the cognitive distinction between “newly-learned information” and the “state of knowledge.” kamio (1997) distinguishes the japanese sentence-final particles yo and ne using the theory of “territory of information,” a theory based on the notion of psychological distance between a given piece of information and the speaker/hearer (see kamio 1994, 1995, 1997 for details of the territory of information theory). he claims that the particle ne is a marker of shared information, which indicates a piece of information presumably within both the speaker’s and addressee’s territories. in contrast, the particle yo implies that the speaker and addressees do not share a piece of information. such information exists in the speaker’s territory but not within the addressee’s territory. using the term “認識の現場性 ninshiki no genbasei (immediate awareness),” miyazaki (2002) argues that ne as well as its corresponding particle in monologue, na/naa, present that a speaker’s awareness is what he/she thinks, feels, and notices at the time of utterance. he used examples of ne and na from novels to argue this point (miyazaki, 2002:11). for instance, ne and naa in examples (6) and (7) are used to present the speaker’s thoughts about what he/she has just seen. (6) (7) all the studies reviewed here suggest it is possible to analyze the difference between yone and ne from a new perspective by focusing on how speakers use yone and ne to represent their states of awareness. the present study follows miyazaki’s (2002) framework of ne, which shows the importance of a speaker’s awareness and analyzes the use of yone and ne in four responding situations. as described in the introduction section, the present study argues that the differences between yone and ne are: a. yone is a marker of existing awareness that indicates that an idea had already been on the speaker’s mind. b. ne is a marker of new awareness that shows a thought just emerged into the speaker’s awareness. [shoya looked at nobuko, who is in front of him.] odoroita ne!—marude betsujin da yo. “nobuko san?” to me o mihatta. “odoroita naa!—marude betsujin da yo.” japanese sentence-final particle yone 178 c. what ne presents is not necessarily something novel, and it can be something the speaker had not thought of until the time of utterance. d. speakers systematically select yone or ne to present their states of awareness. 3 data and method the present study examines the use of yone and ne in conversation using a cognitive-interactive framework. the reasons for this particular approach are twofold: first, this study’s argument about the use of yone stems from miyazaki’s (2002) proposal about ne, which focuses on the speaker’s awareness. second, unlike the data used in miyazaki (2002) and noda (1993), which were constructed or taken from novels, this study examines how yone and ne are used in social interaction. methodologically, the present study proposes hypotheses about using yone and ne based on constructed, written, and conversation data. naturally occurring conversation data provided the data for testing the hypotheses. to carefully examine the conversation data, conversation analysis (ca) with modification was partially adopted. the data consists of nine sets of face-to-face video-recorded multi-party conversations. the source of the data is the sakura corpus distributed by the talkbank organization.5 conversation participants are native speakers of japanese who are students at a japanese university. in each video session, four students, who are classmates or close friends, talk about a given topic. because the conversation participants were close, they used the non-polite conversation style (da-form). while the sakura corpus has 18 sets of conversation data, the data included in this study are female-male conversations and all-female conversations. other data with fewer cases of yone, such as all-male conversations, were excluded. although the data bias is toward female participants, gender is not a contributing factor to the general conclusion of the present study. also, while two male participants’ utterances were occasionally affected by dialect, the corresponding sentence-final particle yone in the dialect, yona, was not identified in the data. thus, dialect does not influence the use of yone in this study’s data. there are 397 cases of yone in the final data examined for this study. of those, 86 instances were speaker responses to various situations and, therefore, constitute this study’s focus. meanwhile, there are 1230 cases of ne in the same data sets since ne is the most frequently used sentence-final particle in japanese conversation (hasegawa, 2010). the study first examined the 86 cases of yone for responses and identified four responding situations. then, another 53 cases of ne in the same responding situations were compared with the cases of yone in response. this study did not identify and categorize all the instances of ne in response; this study’s primary focus is yone while ne is used to compare with yone. data was transcribed using the revised hepburn system of romanization. although transcription of utterances used standard japanese pronunciations, paralinguistic features, such as pauses, sound stretches, and overlapping speech, were also noted in the transcription according to the transcription conventions developed by jefferson (1984) with modifications. also noted were non-vocal actions such as gestures, body alignment, and other contextual information. roman letter codes masked individual speakers, providing anonymity. 5 the data are from the database of talkbank (https://sla.talkbank.org/tbb/ca/sakura). most transcriptions of examples in this paper were updated to include details such as overlapped utterances, pauses, and gestures. such non-verbal communication can provide critical supplemental cues when evaluating a speaker’s intention and such inclusion adds to the argument of the present paper. in addition, five japanese native speakers were consulted to ensure the accuracy of the final transcriptions used in the present paper. https://sla.talkbank.org/tbb/ca/sakura xu 179 4 result and discussion in this section, yone and ne as used in four response situations will be carefully examined and discussed. 4.1 response to immediate situations in this study, “immediate situations” refers to situations in which the speaker has just received new information by seeing or hearing at the time of conversation. what the speaker has just seen or heard triggers a ne-marked utterance. in the present study data, ne is commonly used in such situations while yone is not. example (8) illustrates the use of ne based on what the speakers have just seen and heard. four participants are talking about their ideal partners. before the segment, speaker h says that he likes girls wearing framed glasses. speakers l and k, two females, wonder what makes glasses appealing to speaker h. (8) 1 l: aa kooiu kanji↑ ((pretending to lift glasses with a finger of her right hand)) itj this-kind feeling 2 h: → ii ne nice fp 3 (1.0) 4 l: ii no↑ eeee [igai nice q itj unexpected 5 k: → [eee igai da ne itj unexpected cop fp 6 k: sangurasu no hoo ga ii to omotteta sunglasses lk side nom nice qt thought 7 h: zenzen not-at-all in line 1, speaker l uses gestures like touching the frame of a pair of glasses. speaker h responds to the action with a ne-marked comment (line 2). his response “ii ne (that is nice)” is based on the immediate observation of speaker l’s action. thus, here speaker h uses ne to reflect what he has just observed. another speaker, k, produced a ne-marked comment in line 5. her immediate comment, “igai da ne (that is unexpected),” is based on what she has just heard from speaker h’s response in line 2. her thought that sunglasses were better was different from what she heard from speaker h’s response. this difference indicates that newly-learned information triggered the use of ne. example (8) shows that ne is triggered by what the speaker has just seen or heard. this finding is consistent with miyazaki’s (2002) proposal. on the other hand, no yone-marked response exists in such situations in this study’s data. 4.2 response in recall in 4.1, we saw that only ne exists in response to what the speaker has just seen or heard. however, both ne and yone can be used in response where the speaker has just been reminded of alreadyknown information. this section will illustrate how ne and yone are used to respond to the speaker’s already-known information in recall. in this study, “recall” refers to situations in which old japanese sentence-final particle yone 180 information, such as information mentioned in previous discourse or knowledge already established before the moment of utterance, was brought into the speaker’s mind. 4.2.1 ne in recall example (9) shows how ne relates to the speaker’s own experience, which is already-known information to the speaker. four participants talk about their ideal partners in this segment. speakers h and g, two male participants, previously spoke about the same topic with others in a different session. two female participants, l and k, ask them what they talked about in the previous session. after being reminded of some of their discussion topics, they are asked about another issue (line 1). now speaker g has trouble recalling their conversation topics. (9) 1 l: ato wa↑ after tp 2 (2.0) 3 g: nani ga↑ what nom 4 (2.0) 5 h: fukusoo toka yuttotta jan = clothes like said tag 6 g: → [=a:: yuttotta ne orera itj said fp we 7 k: [=a:: iru ne itj need fp 8 g: yuttotta yuttotta said said speaker h’s utterance in line 5 triggers speaker g’s ne-marked response in line 6. although the information regarding what else they talked about as criteria for ideal partners in a different session is not new to speaker g, the information was not in his awareness until triggered by speaker h’s comment that they also talked about clothes in line 5. speaker g’s ne-marked response produced in line 6 indicates that the information just entered speaker g’s awareness. furthermore, the ne-marked response associates with the change-of-state marker “a:::” (line 2) (heritage, 1984). thus, here ne can be considered as displaying new awareness. 4.2.2 yone in recall example (10) illustrates the use of yone where the speaker talks about information already-known to him/her. (10) 1 i: jikyuu sen en tte ii ne hourly-pay 1000 yen qt good fp 2 e: eq tabun tabun sonna kanji datta itj probably probably that-kind feeling cop-pst 3 l: demo sempai sen en tte itteta [kara but senior 1000 yen qt said because xu 181 4 e: → [itteta yone said fp 5 l: un itj 6 e: kenshuu kikan wa ne: yasui kamoshiren ne training period tp fp cheap possible fp 7 l: demo nai n janai [sonna no] but no nml tag that-kind nml 8 e: [hajime wa] deeta nyuuryoku tte yutteta beginning tp data input qt said 9 kensa mo nani mo shizuni check-up also what also do-neg 10 l: ii na:: good fp before this segment, speaker e told others that she ran into an alumnus at a clinic where the alumnus works. speaker e had an opportunity to work there, and the hourly pay was 1000 yen. maybe because she has not started to work there yet, and her job will be only inputting data into a computer without doing anything such as checking patients (lines 8 and 9), she downgrades her stance (line 2). she does this by showing uncertainty with the adverb “tabun (probably)” when speaker i comments on how good the pay is (line 1). in line 3, speaker l mentions that the alumnus has said the pay is 1000 yen, and speaker e produces the yone-marked response in line 4. indeed, she knows what the alumnus said about the hourly wage, and the information about hourly pay was already-known information to her. her immediate yone-marked response shows that she knows the information clearly and has no problem in recall (line 4). 4.2.3 the differences between ne and yone in recall the environments are distinctly different when comparing the yone-marked response in example (10) with the ne-marked responses in example (9). ne-marked responses tend to appear in situations where the speaker has trouble recalling his/her experience. ne presents what has just entered the speaker’s awareness. in this sense, it is the same as a ne-marked response to an immediate situation (see 4.1), which presents the speaker’s new awareness. on the other hand, a yone-marked response does not associate with a speaker’s recollection difficulties. in other words, the speaker does not use yone to present what has just entered the speaker’s awareness; instead, yone displays one’s existing awareness. as example (9) shows, ne-marked responses tend to occur with the change-of-state maker “a::” (heritage, 1984). in the data for this study, 43 responses marked with ne or yone occur with “a::.” eighty-six percent (37 cases) occur with “a::” and ne, while only 14 percent (6 cases) are prefaced by “a::” and yone. this data indicates that the speaker tends to use ne to present information as something that has just entered one’s awareness at the discourse site, even in a situation involving recalling the speaker’s experience. although the experience per se is certainly not new, a speaker can present it as a piece of new information by using ne. in contrast, “a:: + yone” is deployed where the speaker does not experience any recall issues. the speaker uses yone to indicate previous experience or knowledge, rather than some newly-learned information that has just entered the speaker’s awareness. japanese sentence-final particle yone 182 examples, such as example (11), show how yone works with “a::.” to indicate already-known information that just entered speaker e’s awareness. speakers l and e are talking about people they met at a clinic where they worked as interns. s sensei (mr./ms. s) is a person with whom both speakers l and e worked. when asked about his opinion of s sensei by speaker l (line 1), speaker e comments that s sensei is smart; yone and “a:: ” are attached. (11) 1 l: s sensei wa↑ name teacher tp 2 e: → a:: ano hito atama ii yone itj that person head good fp 3 l: tabun osowaru yo s sensei ((talking to g)) probably be-taught fp name teacher in example (11), “a::” still functions as a marker of change-of-state, and speaker l’s question in line1triggers the use of yone. however, unlike example (9), speaker e does not have difficulty in recall. in sum, these examples demonstrate that the use of yone or ne can be decided by how the speaker wants to present his/her awareness rather than by the information per se. 4.3 response to the answer to a question in this section, we examine the use of yone and ne in question responses. the targeted yone and ne are in the line 3 position, as shown in table 1. table 1. response to the answer to a question when used in response to a question’s answer, the question, i.e., line 1, is designed differently, as shown in table 2. ne-marked response yone-marked response 1 wh question or polar question 2 answer 3 ne-marked response 1 negative question/ wh question + candidate answer 2 answer 3 yone-marked response table 2. yone and ne in response to the answer to a question for ne-marked responses, the questions are wh questions or polar questions. however, for yone-marked responses, questions are primarily negative questions or in the form of “wh question + candidate answer.” in this study’s data, 14 neor yone-marked utterances appear as a response to the answer to a question. table 3 shows the distribution. wh question polar question negative questions wh question + candidate answer ne-marked response 4 6 0 0 yone-marked response 0 1 2 1 total 4 7 2 1 table 3. the distribution of yone and ne in response to the answer to a question response to the answer to a question 1 question 2 answer 3 → response + yone/ne xu 183 to examine this point, observe the following examples: (12) 1 a: → jikyuu ikura↑ hourly-payment how much 2 b: jikyuu wa sekkotsuin ga 850 en rashikku ga hourly-payment tp bonesetter’s-office nom yen restaurant’s name nom 3 950 en kateikyooshi ga 1500 en yen tutor nom yen 4 a: → metcha ii ne very good fp 5 b: un ichi jikan de yes one hour in in example (12), four participants talk about their part-time jobs. first, speaker a asks about speaker b’s hourly payment by using a wh question: “how much is your hourly payment?” (line 1). the question indicates that speaker a does not possess any knowledge about speaker b’s salary. the information she receives from speaker b’s answer in lines 2 and 3 thus forms her response. example (12) illustrates that a ne-marked assessment reflects a speaker’s awareness established when receiving information. negative questions, or wh questions followed by candidate answers, present a speaker’s existing awareness, as illustrated by yone-marked responses in examples (13) and (14). for instance, in example (13), four participants talk about the hourly pay for their part-time jobs. before this segment, speaker g says that even though he wants more money, he tolerates his low hourly pay rate because the part-time job is comfortable and close to his home. (13) 1 i: okane wa doodemo yokatta ↑ money tp however good-pst 2 g: soo hoshii kedo ne that want but fp 3 e: → demo datte yasukunai↑ but but cheap-neg 4 g: yasui cheap 5 e: → yasui yone sore cheap fp that in line 3, speaker e’s question is negative, i.e., “yasukunai (isn’t it cheap?).” also, “datte” is used to strengthen the speaker’s assertion and make others change their stance (mori, 1994). the combination of “datte” and negative question shows speaker e already thinks that speaker g’s pay rate is low. the question design indicates the question’s purpose is to solicit an affirmative answer, not merely a responsive answer from speaker g. upon receiving speaker g’s expected affirmative answer (line 4), speaker e produces a yone-marked utterance (line 5). japanese sentence-final particle yone 184 in example (14), yone appears in response to the answer to a question in the form of “wh question + candidate answer.” four participants are talking about their part-time jobs. (14) 1 g: demo dekireba shitakunai naa: but can-if want-to-do-neg fp 2 a: majide↑ really 3 b: [honto really 4 g: [e:::: baito suki↑ itj part-time-job like 5 a: atashi wa meccha suki da mon i tp extremely like cop because 6 g: → e::::::: nande↑ hito ga ii kara↑ itj why people nom good because 7 a: hito ga ii kara people nom good because 8 g: → [a::: zettai soo da yone. itj definitely that cop fp 9 a: [un yes speaker g says that she does not want a part-time job (line 1), while speaker a says that she likes her part-time job very much (line 5). surprised by speaker a’s opposite opinion, speaker g first asks the reason using a wh question, i.e., “nande (why).” then she produces a candidate answer: because speaker a’s co-workers are nice (line 6). it is notable that the candidate answer, which is in the form of a polar question, does not stand alone; the formation of “wh question + candidate answer” shows that speaker g already has an opinion. confirmed by speaker a (line 7), speaker g produces the yone-marked utterance in line 8. here yone is chosen rather than ne because of speaker a’s opinion that “hito ga ii kara (because they are nice)” was previously expressed in the specially designed question formation in line 6. she also upgrades her opinion with the adverb “zattai (definitely)” (line 8). this example clearly illustrates that yone presents the speaker’s existing awareness. example (15) illustrates one case where yone can respond to the answer to a polar question. four participants talk about another classmate who quit his part-time job. in line 1, speaker b asks whether the classmate is still coming to school. when speaker g tells her the classmate was not at school yesterday (line 4), speaker b produces in line 5 a response with yone and “a::.” (15) 1 b: → gakkoo kiteru↑ kare school coming he 2 c: [shira-nai know-neg xu 185 3 a: [tte iu ka iru no↑ qt say q exist q 4 g: [kinoo i-na-katta yo↑ yesterday exist-neg-pst fp 5 b: → a:: i-na-katta yone itj exist-neg-pst fp in example (12), as mentioned previously, speaker a produced the ne-marked response to the answer of a wh question because he has no access to speaker b’s salary. in contrast, in example (15), speaker b probably knows whether the classmate was at school yesterday, given that they are classmates and possibly take the same classes. like example (11), yone is also used as the maker of change-of-state “a::.” here, yone, triggered by speaker g’s answer in line 4, can also be considered already-known information that has just entered her awareness. the preceding examples of yone and ne demonstrate the distinctions between yone and ne in response to a question’s answer. questions associated with a ne-marked response tend to be wh questions or polar questions, indicating the speaker seeks an answer because the speaker lacks knowledge or information. a ne-marked response also shows that the speaker’s awareness manifests upon hearing an answer. in this sense, the response represents new awareness. in contrast, a yone-marked response primarily occurs when answering negative questions or with “wh question + candidate answer” patterns. these question forms reveal that the speaker does not merely seek an answer but seeks an answer that confirms the awareness he/she has already established. thus, using yone as a response to an answer presents the speaker’s existing awareness. 4.4 response to assessments in this section, using yone and ne in response to assessments is examined. the present study primarily focuses on neand yone-marked responses to assessments without ne or yone. the environment examined in the present study is shown in table 4. ne and yone in second assessment positions are considered. table 4. yone and ne in response to assessments first, ne-marked responses are examined. four females talk about whether they prefer dogs or cats. (16) 1 b: kekka inu ha ga ookatta tte yuu result dog group nom many-past qt say 2 g: inu ha da ne. dog group cop fp 3 a: un tooron ni naranakatta ne itj discussion to became-neg-pst fp 4 atashi neko da yo:: toka ni naranakata i cat cop fp like to became-neg response to assessment 1 assessment 2 → assessment + yone/ne japanese sentence-final particle yone 186 5 c: → neko anma inai n janai no↑ iru kana: cat not-much exist-neg nml tag q exist fp 6 a: iru n janai↑ exist nml tag 7 g: e:::: neko namaiki da yo nanka itj cat self-conceit cop fp like 8 b: e::: demo kawaii yone itj but cute fp 9 g: e:::: kawaii↑ itj cute 10 b: → koneko ga kawaii demo kitten nom cute but 11 g: → a::: koneko wa kawaii ne itj kitten tp cute fp they conclude they prefer dogs to cats in line 1. speaker g is a dog person, and that might be why she is surprised in line 9 when speaker b mentions that cats are cute (line 8). upon speaker g’s reaction, speaker b modifies her statements by saying kittens are cute (line 10). in line 11, speaker g supports speaker b’s view that kittens are cute by using ne. first, she contrasts kittens with adult cats by using the topic marker contrastive “wa,” showing her opinion that kittens are cute but not adult cats (line 11). this opinion corresponds to her utterance in line 7, in which she explains why she does not like cats, i.e., she thinks cats are conceited. speaker g’s utterance in line 11 is prefaced with a prolonged interjection “a::,” which corresponds with “oh” in english and is a change-of-state marker (heritage, 1984). here, before speaker b’s utterance in line 10 that kittens are cute, speaker g opines that cats are not cute. after hearing speaker b’s statement in line 10, speaker g changes the referent of her assessment from “cats in general” to “kittens in particular.” in this sense, ne presents speaker g’s new awareness because she did not think that cats were cute until speaker b’s line 10 assessment. example (17) demonstrates how a speaker can use yone to display one’s existing awareness. before this segment, four female participants talk about speaker a’s part-time job. speaker a works at a famous chocolate store. (17) 1 b: e↑ choko tabereru↑ itj chocolate can-eat 2 a: choko moraeta yo↑ chocolate could-receive fp 3 b: [a: ii na:: itj good fp 4 g: [a: ii na:: itj good fp 5 b: → zettai oishii yo definitely delicious fp xu 187 6 c: un yes 7 a: → metcha oishikatta very delicious-pst 8 b: → da yone cop fp when speaker b finds out that speaker a receives some free chocolates, she immediately comments on the taste of the chocolates (line 5). in line 7, speaker a upgrades her comment on chocolate with the adverb “metcha (very),” saying it was very delicious. speaker b’s yone-marked response to speaker a’s assessment (line 8) corresponds to her comment in line 5, i.e., “zettai oishii yo (definitely delicious).” thus, example (17) illustrates that yone displays a speaker’s existing awareness. 4.5 discussion the data in this study shows that only ne is used to respond to immediate situations such as what a speaker has just seen, heard, or known. prefacing ne with the change-of-state marker “a::” or “ja” indicates previous discourse leading to the speaker’s new awareness (hamada, 1991; heritage, 1984). yone is not found in such a situation. ne can be used in recall. in the examples, although information is already known to the speaker, it is not in the speaker’s awareness before the ne-marked utterance. using ne indicates that information just entered the speaker’s awareness. in this sense, the function of ne is consistent with the immediate situation; that is, the speaker uses ne to present already-known information as a new awareness. in contrast, because yone indicates a speaker’s existing awareness, the change-of-state marker “a::” does not frequently preface yone-marked responses. different question designs result in different uses of ne and yone in response to a question’s answer. in the ne examples, questions are wh questions or polar questions, indicating speakers seek specific information. also, ne-marked responses, in this case, tend to occur with the changeof-state marker “a::.” thus, the use of ne reflects a speaker’s new awareness. in the yone examples, the question design is different. yes, speakers also seek the answer, but they already recognize possible answers when articulating the questions. one particular question design (negative question or question in the form of “wh question + candidate answer”) solicits confirmation of the speaker’s existing awareness rather than seeking new information. thus, when an answer confirms knowledge, the speaker’s response to it tends to be associated with the marker yone. when ne is used in second assessments to respond to first assessments, the ne-marked second assessment tends to be different from the assessment on the same referent that the speaker made in previous discourse. when using yone in such a situation, the yone-marked second assessment remains consistent with prior assessments made or implied by the speaker. in this sense, yone presents a speaker’s existing awareness. now let us examine the use of yone and ne in example (18), four participants are talking about their ideal partners. before this segment, they talked about what kind of things they do not like. (18) 1 h: kechi toka stingy like 2 k: → a: kechi [iya da ne itj stingy dislike cop fp japanese sentence-final particle yone 188 3 l: → [kechi iya da yone stingy dislike cop fp 4 h: kechi iya da stingy dislike cop when speaker h brings up “being cheap,” speakers k and l immediately respond to it with almost the same assessment. the difference is that speaker k’s ne-marked response prefaces the change-of-state maker “a::” while speaker l’s yone-marked response does not. using only conversational data, it is indeed impossible to know a speaker’s underlying intention. however, a reasonable interpretation exists to explain the differences between yone and ne in response: speaker k uses ne to present her new awareness; that is, she thinks at the moment of utterance that being cheap is not good. speaker l uses yone to indicate her existing awareness; that is, she presents her thought in this way because she believed for a long time that being cheap is not good. example (18) is different in a noteworthy way from the four types of responses examined in this study. speaker k and speaker l use ne and yone to respond to a new topic. the new topic is not newly-learned information, nor does the response occur in a similar turn with other cases included in the study. it appears that the proposed hypothesis in this study is also able to explain the use of yone and ne in different situations, which are unexamined closely in the present study. future research should explore this point with examples of yone and ne in other responding contexts. 6 5 conclusion this study examined the differences between yone and ne in the context of responses with naturally occurring conversation data. this study has found that generally, ne is used to present the speaker’s new awareness while yone shows the speaker’s existing awareness. also, this study shows that using yone or ne can be decided by how the speaker presents his/her state of awareness, rather than by the information per se to which yone or ne is attached. this study is limited in several ways. first, the differences between yone and ne demonstrated in the present study apply within the context of four types of response. while the proposed hypothesis appears to successfully explain the difference between yone and ne in example (18), the present study could not further explore this point due to the limited data. future research regarding whether this hypothesis works in other settings, such as seeking confirmation or providing new information, both contexts in which the use of yone and ne overlap, may further expand the understanding of yone and ne. second, this study focuses on using yone and ne to consider speakers’ states of awareness and does not analyze the pragmatic functions of yone and ne in response in detail. further investigation into the difference of pragmatic functions of yone and ne in response and other situations is strongly recommended. despite its limitations, this study certainly adds to our understanding of complex japanese sentence-final particles. it also extends our knowledge of akatsuka’s (1985) proposal that japanese grammar is sensitive to the cognitive distinction between “newly-learned information” and the 6 in addition, as one reviewer points out, the order of “ne” and “yone” in example (18) cannot be reversed. that is, it is unnatural to say “a kechi iya da yone” followed by “kechi iya da ne” in a situation in which two speakers are almost simultaneously responding to a question. while this irreversibility seems to be naturally accounted for by this study about the use of “yone” and “ne,” this point cannot be elaborated due to limited data and should be subject to further exploration. xu 189 “state of knowledge.” the methods used for this study may be applied to other research on final sentence particles. considering the complexity and frequent use of yone and ne in conversation, the insights gained from this study may also be of assistance to the field of teaching japanese as a foreign language. transcription conventions 1. transcript symbols [ the point where overlapping talk begins ] the point where overlapping talk ends (.) micro-pause (0.0) length of silence :: noticeably lengthened sound = latched utterance ↑ rising intonation (( )) transcriber’s descriptions hh laughter 2. abbreviations aux auxiliary cop copula fp final particle itj interjection lk nominal linking particle neg negative morpheme nml nominalizer nom subject marker o object marker pst past tense q question marker qt quotative marker tag tag-like expression tp topic marker references noriko akatsuka (1985). conditionals and the epistemic scale. language 61(3): 625-639. yuko asano-cavanagh (2011). an analysis of three japanese tags: ne, yone, and daroo. pragmatics & cognition 19 (3):448-475. haruko minegishi cook (1990). the sentence-final particle ne as a tool for cooperation in japanese conversation. japanese/korean linguistics1: 29-44. haruko minegishi cook (1992). meanings of non-referential indexes: a case study of the japanese sentence-final particle ne. text-interdisciplinary journal for the study of discourse 12(4): 507–540. haruko minegishi cook (2001). particles. in alessandro duranti (ed.), key terms in language and culture,176-179. malden, blackwell publishers. japanese sentence-final particle yone 190 john goldsmith and erich woisetschlaeger (1982). the logic of the english progressive. linguistic inquiry 13(1): 79–89. jeanette k gundel (1985). shared knowledge’ and topicality. journal of pragmatics 9 (1): 83107. mari hamada (1991). dewa’ no kinoo: suiron to seitsuzokugo [the functions of ‘dewa’: inference and connectives]. handai nihongo kenkyuu 3: 25-44. (in japanese) yoko hasegawa (2010). the sentence-final particles ne and yo in soliloquial japanese. pragmatics 20 (1):71–89. atsuko hasunuma (1992). shuujoshi no fukugookei ‘yone’ no yoohoo to kinoo [the use and functions of the compound sentence-final particle ‘yone’]. taishoo kenkyuu dainigoo: hatsuwa maakaa nitsuite: 63-77. (in japanese) atsuko hasunuma (1995). taiwa niokeru kakunin kooi ‘daroo’ ‘janaika’ ‘yone’ no kakunin yoohoo [confirmation of action in dialogue: on the confirmation function of ‘daroo’ ‘janaika’ ‘yone’]. in niida yoshio (ed.) fukubun no kenkyuu [studies of complex sentences], 389– 419. tokyo, kuroshio shuppan. (in japanese) kaoru hayano (2011). claiming epistemic primacy: yo-marked assessments in japanese. in tanya stivers, lorenza mondada, and jakob steensig (ed.), the morality of knowledge in conversation, 58–81. cambridge, cambridge university press. kaoru hayano (2013). territories of knowledge in japanese conversation. utrecht, lot. john heritage (1984). a change-of-state token and aspects of its sequential placement.” in john heritage and j. maxwell atkinson (ed.) structures of social action: studies in conversation analysis, 299–345. cambridge, cambridge university press. john heritage (2002). oh-prefaced responses to assessments: a method of modifying agreement/disagreement. in c. ford, b. fox, and s. a. thompson (ed.), the language of turn and sequence, 196-224. oxford, oxford university press. john heritage and geoffrey raymond (2005). the terms of agreement: indexing epistemic authority and subordination in talk-in-interaction. social psychology quarterly 68 (1):15–38. eiko izuhara (1993). ‘ne’ to ‘yo’ saikoo – ‘ne’ to ‘yo’ no komyunikeeshon kinoo no koosatsu kara [revisit ‘ne’ and ‘yo’an analysis of the communicative functions of ‘ne’ and ‘yo’]. nihongo kyooiku:103-114. (in japanese) eiko izuhara (1994). kandooshi kantoojoshi shuujoshi ‘ne/nee’ no intoneeshonn – danwa shinkoo tono kakawari kara [the intonation of interjection, interjectory particle, sentence-final particle ‘ne/nee’the relation with conversation development]. nihongo kyooiku: 96–107. (in japanese) eiko izuhara (2001). ‘ne’ to ‘yo’ saisaikoo [re-revisit ‘ne’ and ‘yo’]. aichi gakuin daigaku kyooyoobu kiyoo 49(1):35–49. (in japanese) eiko izuhara (2003). shuujoshi ‘yo’ ‘yone’ ‘ne’ saikoo [revisit sentence-final particle ‘yo’ ‘yone’ ‘ne’]. aichi gakuin daigaku kyooyoobu kiyoo 51(2): 1–15. (in japanese) eiko izuhara (2008). kantoojoshi shuujoshi no danwa kanri kinoo bunseki – ‘ne’ ‘yone’ ‘yo’ no baai [a discourse management functional analysis of interjectory and sentence-final particle ‘ne’ ‘yone’ ‘yo’]. aichi gakuin daigaku kyooyoobu kiyoo 56(1): 67–82. (in japanese) gail jefferson (1984). transcript notation.” in j. heritage and j. maxwell atkinson (ed.) structures of social action: studies in conversation analysis, ix-xvi. cambridge, cambridge university press. akio kamio (1994). the theory of territory of information: the case of japanese. journal of pragmatics 21 (1): 67–100. akio kamio (1995). territory of information in english and japanese and psychological utterances. journal of pragmatics 24 (3): 235–64. xu 191 akio kamio (1997). territory of information. amsterdam, john benjamins publishing. yasuhiro katagiri. (2007). dialogue functions of japanese sentence-final particles ‘yo’ and ‘ne.’ journal of pragmatics 39 (7): 1313–23. shigehiro kato (2001). bunmatsu joshi ‘ne’ ‘yo’ no danwa koosei kinoo [the discourse constructive functions of the sentence-final particles ‘ne’ and ‘yo’]. toyama daigaku jinbun gakubu kiyoo 35:31-48. (in japanese) susumu kuno (1972). functional sentence perspective: a case study from japanese and english. linguistic inquiry 3: 269-320. duckyoung lee (2007). involvement and the japanese interactive particles ne and yo. journal of pragmatics 39 (2):363–88. senko maynard (1993). discourse modality: subjectivity, emotion, and voice in the japanese language. amsterdam, john benjamins publishing. naomi h mcgloin, mustuko endo hudson, fumiko nazikian, and tomomi kakegawa (2013). modern japanese grammar: a practical guide. london, routledge. kazuhito miyazaki (2002). sheloshim ‘ne’ to ‘na’ [the sentence-final particles ‘ne’ and ‘na’]. handai nihongo kenkyuu14. 1-19. (in japanese) junko mori (1994). functions of the connective datte in japanese conversation. japanese/korean linguistics 4:147-163. emi morita (2002). stance marking in the collaborative completion of sentences: final particles as epistemic markers in japanese. japanese/korean linguistics 10: 220–233. keiko noda (1993). shuujoshi ‘ne’ to ‘yo’ no kinou: ‘yone’ to kasanaru baai [the functions of sentence-final particles ‘ne’ and ‘yo’: the cases of overlapping ‘yone’]. gengo bunka to nihongo kyooiku 6:10-21. (in japanese) ellen f prince (1992). the zpg letter: subjects, definiteness, and information-status.” in w. c. mann and s. a. thompson (ed.), discourse description: diverse linguistic analyses of a fund-raising text, 295-325., amsterdam, john benjamins publishing. anita pomerantz (1984). agreeing and disagreeing with assessments: some features of preferred/dispreferred turn shaped.” in j. heritage and j. maxwell atkinson (ed.), structures of social action: studies in conversation analysis, 57-101. cambridge, cambridge university press. hideki saigo (2011). the japanese sentence-final particles in talk-in-interaction. amsterdam, john benjamins publishing. emanuel a schegloff (2007). sequence organization in interaction: a primer in conversation analysis i. cambridge, cambridge university press. dan slobin and ayhan aksu (1982). tense, aspect and modality in the use of the turkish evidential. in p. j. hopper (ed.), tense-aspect: between semantics and pragmatics: 185-200. amsterdam, john benjamins publishing. yukinori takubo and satoshi kinsui (1997). discourse management in terms of mental spaces. journal of pragmatics 28 (6): 741–58. hiroko tanaka (2000). the particle ne as a turn-management device in japanese conversation. journal of pragmatics 32 (8):1135–76. jun xu (2016). an analysis of the japanese sentence-final particle yone. ph.d. thesis, university of wisconsin madison, madison, wisconsin. huifang zhang (2009). shizen kaiwa niokeru ‘yone’ no imi ruikei to hyoogen kinoo [the types of meaning and functions of ‘yone’ in natural conversation]. gengogaku ronsoo onrain ban 2: 17-32. (in japanese) dialogue and discourse 4(2) (2013) 87-117 doi: 10.5087/dad.2013.205 a corpus of science journalism for analyzing writing quality annie louis lannie@seas.upenn.edu university of pennsylvania philadelphia, pa, 19104, usa ani nenkova nenkova@seas.upenn.edu university of pennsylvania philadelphia, pa, 19104, usa editors: stefanie dipper, heike zinsmeister, bonnie webber abstract we introduce a corpus of science journalism articles, categorized in three levels of writing quality.1 the corpus fulfills a glaring need for realistic data on which applications concerned with predicting text quality can be developed and evaluated. in this article we describe how we identified, guided by the judgements of renowned journalists, samples of excellent, very good and typical writing. the first category comprises extraordinarily well-written pieces as identified by expert journalists. we expanded this set with other articles written by the authors of these excellent samples to form a set of very good writing samples. in addition, our corpus also comprises a larger set of typical journalistic writing. we provide details about the corpus and the text quality evaluations it can support. our intention is to further extend the corpus with annotations of phenomena that reveal quantifiable differences between levels of writing quality. here we introduce two such annotations that have promise for distinguishing amazing from typical writing: text generality/specificity and communicative goals. we present manual annotation experiments for specificity of text and also explore the feasibility of acquiring these annotations automatically. for communicative goals, we present an automatic clustering method to explore the possible set of communicative goals and develop guidelines for future manual annotations. we find that the annotation of general/specific nature on sentence level can be performed reasonably accurately fully automatically, while automatic annotations of communicative goals reveals salient characteristics of journalistic writing but does not align with categories we wish to annotate in future work. still with the current automatic annotations, we provide evidence that features based on specificity and communicative goals are indeed predictive of writing quality. keywords: text quality, readability, science journalism, corpus 1. introduction text quality is an elusive concept. it is difficult to define text quality precisely but it has huge potential for transforming the way we use applications to access information. for example in information retrieval, rankings of documents can be improved by balancing the relevance of the document to the user query and the estimated quality of the document. similarly, tools for writing support can be developed to help people improve their writing. text quality prediction requires that texts can 1. the corpus can be obtained from http://www.cis.upenn.edu/˜nlp/corpora/scinewscorpus. html c©2013 annie louis and ani nenkova submitted 03/12; accepted 12/12; published online 05/13 louis and nenkova be automatic analyzed in multiple dimensions and naturally links work from seemingly disparate fields of natural language processing such as discourse analysis, coherence, readability, idiomatic language use, metaphor identification and interpretation and automatic essay grading. to develop models to predict text quality, we need datasets where articles have been marked with a quality label. but the problem of coming up with a suitable definition for article quality has hindered corpus development efforts. there are reasonable sources of text quality labels for low level aspects such as spelling and grammar in the form of student essays and search query logs. such data, however, have remained proprietary and not available for research outside of the specific institutions. there are also datasets where automatic machine summaries and translations have been rated by human judges. yet systems trained on such data are very unlikely to transfer with the same accuracy to texts written by people. on the other hand, studies that were not based on data from automatic systems have made a simple assumption that articles of bad quality can be obtained by manipulating good articles. one approach used to evaluate many systems that predict organization quality is to take an article and randomly permute its sentences to obtain an incoherent sample. however, such examples are rather unrealistic and are not representative of problems generally encounted in writing. experiments on such datasets may give optimistic results that do not carry over to real application settings. well-written texts are also marked by properties beyond general spelling, grammar and organization. they are beautifully written, use creative language and are on well-chosen topics. there are no datasets on which we can explore linguistic correlates of such writing. the goal of the work that we describe here is to provide a realistic dataset on which researchers interested in text quality can develop and validate their theories about what properties define a wellwritten text. our corpus consists of science journalism articles. this genre has unique features particularly suitable for investigating text quality. science news articles convey complex ideas and research findings in a clear and engaging way, targeted towards a sophisticated audience. good science writing explains, educates and entertains the reader at the same time. for example, consider the snippets in table 1. they provide a view of the different styles of writing in this genre. the impact and importance of a research finding is conveyed clearly and the research itself explained thoroughly (snippet a). snippet b shows the use of creative writing style involving unexpected phrasing (“fluent in firefly”) and visual language (“rise into the air and begin to blink on and off”). the last snippet also presents humour and a light-hearted discussion of the research finding. these examples show how this genre can be interesting for studying text quality. note also how these examples are in great contrast with research descriptions commonly provided in journals and conference publications. in this regard, these articles offer a challenging testbed for computational models for predicting to what extent a text is clear, interesting, beautiful or well-structured. our corpus contains several thousand articles, divided into three coarse levels of writing quality in this genre. all articles have been published in the new york times (nyt), so the quality of a typical article is high. a small sample of great articles was identified with the help of subjective expert judgements of established science writers. a substantially larger set of very good writing was created by identifying articles published in the nyt and written by writers whose texts appeared in the great section of the corpus. finally, science-related pieces on topics similar to those covered in the great and very good articles but written by different authors formed the set of typical writing. in section 3 we provide specific details about the collection and categorization of articles in the corpus. our corpus is also realistic compared to data sets used in the past and also relevant for many application 88 a corpus of science journalism for analyzing writing quality a. clarity in explaining research the mystery of time is connected with some of the thorniest questions in physics, as well as in philosophy, like why we remember the past but not the future, how causality works, why you can’t stir cream out of your coffee or put perfume back in a bottle. ... dr. fotini markopoulou kalamara of the perimeter institute described time as, if not an illusion, an approximation, ”a bit like the way you can see the river flow in a smooth way even though the individual water molecules follow much more complicated patterns.” b. creative language use sara lewis is fluent in firefly. on this night she walks through a farm field in eastern massachusetts, watching the first fireflies of the evening rise into the air and begin to blink on and off. dr. lewis, an evolutionary ecologist at tufts university, points out six species in this meadow, each with its own pattern of flashes. c. entertaining nature news flash: we’re boring. new research that makes creative use of sensitive location-tracking data from 100,000 cellphones in europe suggests that most people can be found in one of just a few locations at any time, and that they do not generally go far from home. table 1: example snippets from science news articles settings such as article search and recommedation. a further attractive aspect of our corpus is that the average quality of an article is high, they were all edited and published in a leading newspaper. therefore our corpus allows for examining linguistic aspects related to quality at this higher end of the quality spectrum. such distinctions have been unexplored in past studies. given the examples in table 1, we can see that automatic identification of many aspects such as research content, metaphor, visual language can be helpful for predicting text quality. as a first step, manual annotations of these dimensions would be necessary to train systems to automatically annotate new articles. in sections 5 and 6, we focus on the annotation of two aspects–general/specific nature and communicative goals of sentences. the general-specific distinction refers to the high level statements and specific details in the article. we believe that a proper balance between the use of general and specific information can contribute to text quality. communicative goals refer to the author’s intentions for every sentence. for example, the intention may be to define a concept, provide an example, narrate a story etc. certain sequences of intentions may work better and create articles of finer quality compared to other intention patterns. both these aspects, text specificity and communicative goals, appear to be relevant for quality given the explanatory nature of the science journalism genre. in section 5 we give further motivation for annotating specificity. we also present results on sentence-level annotation of sentence specificity both by annotators and automatically; we derive article-level characterizations of the distribution of specificity in the text and perform initial experiments which indicate that text specificity, computed fully automatically, varies between texts from different categories in the corpus. in section 6 we discuss the need for annotation of communicative goals. we experiment with an unsupervised approach for capturing types of communicative goals. we find that the model reveals systematic differences between categories of writing quality 89 louis and nenkova but does not directly correspond to distinctions we intended to capture. in section 8 we conclude with discussion of further cycles of introducing factors related to text quality and annotating data either manually or automatically. 2. background in this section, we describe the different aspects of text quality which have been investigated in prior work, with special focus on the datasets used to validate the experiments. there is the rich and well-developed field of readability research where the aim is to predict if a text is appropriate for a target audience. the audience is typically categorized by age, education level or cognitive abilities (flesch, 1948; gunning, 1952; dale and chall, 1948; schwarm and ostendorf, 2005; si and callan, 2001; collins-thompson and callan, 2004). most of the data data used for these experiments accordingly come from educational material designed for different grade levels. some studies have focused on a special audience such as non-native language learners (heilman et al., 2007) and people with cognitive disabilities (feng et al., 2009). in heilman et al. (2007), the authors use a corpus of practice reading material designed for english language learners. feng et al. (2009) create a corpus of articles where for each article, they had their target users answer comprehension questions and the average score obtained by the users for each article was used as a measure of that article’s difficulty. but readability work does not traditionally address the question of how we can rank texts suitable for the same audience into well or poorly written ones. in our corpus, we aim to differentiate texts that are exceptionally well-written from those which are typical in the science section of a newspaper. some recent approaches have focused on this problem of capturing differences in coherence within one level of readership. for example, they seek to understand what is a good structure for a news article when we consider a competent, college-educated adult reader. rather than developing scores based on cognitive factors, they characterize text coherence by exploring systematic lexical patterns (lapata, 2003; barzilay and lee, 2004; soricut and marcu, 2006), entity coreference (barzilay and lapata, 2008; elsner et al., 2007; karamanis et al., 2009) and discourse relations (pitler and nenkova, 2008; lin et al., 2011) from large collections of texts. but since suitable data for evaluation is not available, a standard evaluation technique for several of these studies is to test to what extent a model is able to distinguish a text from a random permutation of the sentences in the same text. this approach removes the need for creating dedicated corpora with coherence ratings but severely limits the scope of findings because it remains an open question if results on permutation data will generalize for comparisons of text in more realistic tasks. one particular characteristic of the permutations data is that both the original and permuted texts are the same length. this setting might be easier for systems than the case where texts to be compared do not necessarily have the same length. further, incoherent examples created by permuting sentences resemble text generated by automatic systems. for articles written by people there may be more subtle differences than the clearly incoherent permutation examples and performance may be unrealistically high on the permutations data. in this area specifically the need for realistic test data is acute. work on automatic essay grading (burstein et al., 2003; higgins et al., 2004; attali and burstein, 2006) and error correction in texts written by non-native speakers (de felice and pulman, 2008; gamon et al., 2008; tetreault et al., 2010) has utilized actual annotations on student writing. however, most datasets for these tasks are proprietary. in addition, they seek to identify potential deviations from a typical text that indicates poor mastery of language. the corpus that we collect targets analy90 a corpus of science journalism for analyzing writing quality sis at the opposite end of the competency spectrum, seeking to find superior texts among reasonably high typical standard. the largest freely available annotations of linguistic quality come from manual evaluations of machine generated text such as machine translation and summarization. in several prior studies researchers have designed metrics that can replicate the human ratings of quality for these texts. however, the factors which indicate well-written nature of human texts are rather different from those useful for predicting quality of machine generated texts; recent findings suggest that prediction methods trained on both machine translated sentences or machine produced summaries perform poorly for articles written by people (nenkova et al., 2010). so these annotations have little use outside the domain of evaluating and tuning automatic systems. another aspect that has not been considered in prior work is the effect of genre on metrics for writing quality. readability studies have taken motivation in cognitive factors and their measures based on word familiarity and sentence complexity are assumed to be generally applicable for most texts. on the other hand, data-driven methods aiming to learn coherence indicators are proposed such that they can utilize any collection of texts of interest and learn their properties. however, a well-written story would be characterised by different factors compared to good writing in academic publications. there are no existing corpora that can be used to develop and evaluate metrics that are unique for a genre. we believe that our corpus will help to evaluate not just readability and coherence, but the overall well-written nature or text quality. for this, we need to have texts that are not focused on competency levels of the audience. science journalism is intended for an adult educated audience and our corpus picks out articles that are considered extremely well-written compared to typical articles appearing in the science section of a newspaper. so our corpus ratings directly reflect differences in writing quality for the same target audience. secondly, our corpus will also enable the study of genrespecific metrics beyond ease of reading. these articles can be examined to understand linguistic correlates of interesting and clear writing and of the style that accomodates and encourages lay readers to understand research details. we believe that these important characteristics of science writing can be better studied using our corpus. in terms of applications of measuring text quality, some recent work on information retrieval have incorporated readability ratings in the ranking of web page results (collins-thompson et al., 2011; kim et al., 2012). but these studies so far have utilized the notion of grade levels only and the dominant approach is to predict the level using the vocabulary differences learnt from the training corpus. we believe that our corpus can provide a challenging setting for such a ranking task. in our corpus, for every well-written article, we also identify and list other topically similar articles which have more typical or average writing. while in the case of retrieving webpages on different topics, it may be difficult to apply genre-specific measures, we hope that for our corpus one can easily explore the impact of different measures: relevance to query, generic readability scores and metrics specific to science journalism. so far, only some studies on social media have been able to incorporate domain-specific features such as typos and punctuation for the ranking of questions and answers in a forum (agichtein et al., 2008). 3. collecting a corpus of science journalism ratings for text quality are highly subjective. several factors could influence judgements such as choice of and familiarity with the topic and personal preference, particularly for specialized domains 91 louis and nenkova such as science. in our work we have chosen to adopt the opinion of renowned science journalists to identify articles deemed to be great. we then expand the corpus using author and topic information to form a second set of very good articles. a third level of typical articles on similar topics but by different authors was then formed. all articles in the corpus were published in the new york times between 1999 and 2007. 3.1 selecting great articles the great articles in our corpus come from the “best american science writing” annual anthologies. the stories that appear in these anthologies are chosen by prominent science journalists who serve as editors of the volume, with a different editor overseeing the selection each year. in some of the volumes, the editors explain the criteria they have applied while searching for articles to include in the volume: • “first and most important, all are extremely well written. this sounds obvious, and it is, but for me it means the pieces impart genuine pleasure via the writers’ choice of words and the rhythm of their phrases... “i wish i’d written that”, was my own frequent reaction to these articles.” (2004) • “the best science writing is science writing that is cool... i like science writing to be clear and to be interesting to scientists and nonscientists alike. i like it to be smart. i like it, every once in a while, to be funny. i like science writing to have a beginning, middle and end—to tell a story whenever possible.” (2006) • “three attributes make these stories not just great science but great journalism: a compelling story, not just a topic; extraordinary, often exclusive reporting; and a facility for concisely expressing complex ideas and masses of information.” (2008) given the criteria applied by the editors, the “best american science writing” articles present a wonderful opportunity to test computational models of structure, coherence, clarity, humor and creative language use. from the volumes published between 1999 and 2007, we picked out articles that originally appeared in the new york times newspaper. it was straightforward to obtain the full text of these articles from the new york times corpus (sandhaus, 2008) which has all articles published in the nyt for 20 years, together with extensive metadata containing editor assigned topic tags. all articles in our corpus consist of the full text of the article and the associated nyt corpus metadata. there are 52 articles that appear both in the anthologies and the nyt corpus and they form the set of great writing. obviously, the topic of an article will influence the extent to which it is perceived as well-written. one would expect that it is more likely to write an informative, enjoyable and attention-grabbing newspaper article related to medicine and heath compared to writing a piece with equivalent impact on the reader on the topic of coherence models in computational linguistics. we use the nyt corpus topic tags to provide a first characterization of the articles we got from the “best american science writing” anthology. there are about 5 million unique tags in the full nyt corpus and most articles have five or six tags each. the number of unique tags for the set of great writing articles is 325 which is too big to present. instead, in table 2 we present the tags that appear in more than three articles in the great set. medicine, space and physics are the most popular subjects in the collection. computers and finance topics are much lower in the list. in the next section we describe the procedure we used to further expand the corpus with samples of very good and typical writing. 92 a corpus of science journalism for analyzing writing quality tag no. articles medicine and health 18 research 17 science and technology 12 space 11 physics 9 archaeology and anthropology 7 biology and biochemistry 7 genetics and heredity 6 dna (deoxyribonucleic acid) 5 finances 5 animals 4 computers and the internet 4 diseases and conditions 4 doctors 4 drugs (pharmaceuticals) 4 ethics 4 planets 4 reproduction (biological) 4 women 4 table 2: most frequent metadata tags in the great writing samples 3.2 extraction of very good and typical writing the number of great articles is relatively small—just 52—so we expanded the collection of good writing using the nyt corpus. the set of very good writing contains nyt articles about research that were written by authors whose articles appeared in the great sub-corpus. for the typical category, we pick other articles published around the same time but were neither chosen as best writing nor written by the authors whose articles were chosen for the anthologies. the nyt corpus contains every article published between 1987 to 2007 and has a few million articles. we first filter some of the articles based on topic and research content before sampling for good and average examples. the goal of the filtering is to find articles about science that were published around the same time as our great samples and have similar length. we consider only: • articles published between 1999 and 2007. the best science writing anthologies have been published since 1997 and the nyt corpus contains articles upto 2007. • articles of at least 1,000 words. all articles from the anthologies had that minimum length. • only science journalism pieces. in the nyt metadata, there is no specific tag that identifies all the science journalism articles. so, we create a set of metadata tags which can represent this genre. since we know the great article set to be science writing, we choose the minimal subset of tags such that at least one tag per great article appears on the list. we dub this set as the “science tags”. we derived this list using greedy selection, choosing the tag that describes the largest number of great articles, then the tag that appears in most of the remaining articles, and so on until we obtain a list of tags that covers all great articles. table 3 lists the eleven topic tags that made it into the list. 93 louis and nenkova medicine and health research space physics computers and the internet brain evolution disasters religion and churches language and languages environment table 3: minimal set of “science tags” which cover all great articles people process topic publications endings other researcher discover biology report -ology human scientist discuss physics published -gist science physicist experiment chemistry journal -list research biologist work anthropology paper -mist knowledge economist finding primatology author -uist university anthropologist study issue -phy laboratory environmentalist question lab linguist project professor dr student table 4: unique words from the research word dictionary we consider an article to be science related if it has one of the topic tags in table 3 and also mentions words related to science such as ‘scientist’, ‘discover’, ‘found’, ‘physics’, ‘publication’, ‘study’. we found the need to check for words that appear in the article because in the nyt, research-related tags are assigned even to articles that only cursorily mention a research problem such as stem cells but otherwise report general news. we used a hand-built dictionary of research words and remove articles that do not meet a threshold for research word content. the dictionary comprises a total of 71 lexical items including morphological variants. six of the entries in the dictionary are regular expression patterns that match endings such as “-ology” and “gist” that often indicate research related words. the unique words from our list are given in table 4. we have grouped them into some simple categories here. an article was filtered out when (a) fewer than 10 of its tokens matched any entry in the dictionary or (b) there were fewer than 5 unique words from the article that had dictionary matches. this threshold keeps articles that have high frequency of research words and also diversity in these words. the threshold values were tuned such that all the articles in the great set scored above the cutoff. after this step, the final relevant set has 13,466 science-related articles on the same topics as the great samples. the great articles were written by 42 different authors. some authors had more than one article appearing in that set, and a few have even three or more articles in that category. it is reasonable to consider that these writers are exceptionally good, so we extracted all articles from the relevant set written by these authors to form the very good set. there are 2,243 in that category. the remaining articles from the relevant set are grouped to form the typical class of articles, a total of 11,223. a summary of the three categories of articles is given in table 5. 94 a corpus of science journalism for analyzing writing quality category no. articles no. sentences no. tokens great 52 6,616 163,184 very good 2243 160,563 4,054,057 typical 11223 868,059 21,522,690 total 13518 1,035,238 25,739,931 table 5: overview of great, very good and typical categories in the corpus 3.3 ranking corpus as already noted in the previous section, the articles in science journalism span a wide variety of topics. the writing style for articles from different topics, for example health vs. religion research would be widely different and hard to analyze for quality differences. further, the sort of applications we envision, article recommendation for example, would involve ranking articles within a topic. so we pair up articles by topic to also create a ranking corpus. for each article in the great and very good sets, we associate a list of articles from the typical category which discuss the same or closely related topic. to identify topically similar articles, we compute similarity between articles. only the descriptive topic words identified via a log likelihood ratio test are used in the computation of similarity. the descriptive words are computed by the topics tool2 (louis and nenkova, 2012b). each article is represented by binary features which indicate the presence of each topic word. the similarity between two articles is computed as the cosine between their vectors. articles with similarity equal to or above 0.2 were matched. table 6 gives an example of two matched articles. the number of typical articles associated with each great or very good article varied considerably: 1,138 of the good articles were matched with 1 to 10 typical articles; 685 were matched with 11 to 50 typical articles and 59 were matched with 51 to 140 typical articles. if we enumerate all pairs of (great or very good, typical) articles from these matchings, there are a total of 25032 pairs. in this way, for many of the high quality articles we have collected examples with hypothesized inferior quality, but on the same topic. the dataset can be used for exploring ranking tasks where the goal would be to correctly identify the best article given a good article among a set of similar articles. the corpus can support investigations into what topics, events or stories draw readers’ attention. high quality writing is interesting and engaging and there are few studies that have attempted to predict this characteristic. mcintyre and lapata (2009) present a classifier for predicting interest ratings for short fairy tales and pitler and nenkova (2008) report that ratings of text quality and reader interest are complementary and do not correlate well. to facilitate this line of investigation, as well as other semantically oriented analysis, in the corpus we provide a list of all descriptive topic words which were the basis for article mapping. 4. annotating elements of text quality so far, we described how we have collected a corpus of science journalism writing with varying quality. in later work we would like to further annotate these articles with different elements indica2. http://www.cis.upenn.edu/ lannie/topics.html 95 louis and nenkova great writing sample kristen ehresmann, a minnesota department of health official, had just told a state senate hearing that vaccines with microscopic amounts of mercury were safe. libby rupp, a mother of a 3-year-old girl with autism, was incredulous. “how did my daughter get so much mercury in her?” ms. rupp asked ms. ehresmann after her testimony. “fish?” ms. ehresmann suggested. ”she never eats it,” ms. rupp answered. “do you drink tap water?” “it’s all filtered.” “well, do you breathe the air?” ms. ehresmann asked, with a resigned smile. several parents looked angrily at ms. ehresmann, who left. ms. rupp remained, shaking with anger. that anyone could defend mercury in vaccines, she said, ”makes my blood boil.” public health officials like ms. ehresmann, who herself has a son with autism, have been trying for years to convince parents like ms. rupp that there is no link between thimerosal – a mercury-containing preservative once used routinely in vaccines – and autism. they have failed. the centers for disease control and prevention, the food and drug administration, the institute of medicine, the world health organization and the american academy of pediatrics have all largely dismissed the notion that thimerosal causes or contributes to autism. five major studies have found no link. yet despite all evidence to the contrary, the number of parents who blame thimerosal for their children’s autism has only increased. and in recent months, these parents have used their numbers, their passion and their organizing skills to become a potent national force. the issue has become one of the most fractious and divisive in pediatric medicine. topically related typical text neal halsey’s life was dedicated to promoting vaccination. in june 1999, the johns hopkins pediatrician and scholar had completed a decade of service on the influential committees that decide which inoculations will be jabbed into the arms and thighs and buttocks of eight million american children each year. at the urging of halsey and others, the number of vaccines mandated for children under 2 in the 90’s soared to 20, from 8. kids were healthier for it, according to him. these simple, safe injections against hepatitis b and germs like haemophilus bacteria would help thousands grow up free of diseases like meningitis and liver cancer. halsey’s view, however, was not shared by a footnotesize but vocal faction of parents who questioned whether all these shots did more harm than good. while many of the childhood infections that vaccines were designed to prevent – among them diphtheria, mumps, chickenpox and polio – seemed to be either antique or innocuous, serious chronic diseases like asthma, juvenile diabetes and autism were on the rise. and on the internet, especially, a growing number of self-styled health activists blamed vaccines for these increases. table 6: snippets from topically related texts belonging to the great and typical categories 96 a corpus of science journalism for analyzing writing quality tive of text quality and use these aspects to compute measurable differences between the typical science articles and the extraordinary writing in the great and very good classes. some annotations may be carried out manually, others could be done on a larger scale automatically whenever system performance is good. in the remainder of this article we discuss two types of annotation for science news: the generalspecific nature of content and the communicative goals conveyed by sentences in the texts. we explore how these distinctions can be annotated and also present automatic methods which use surface clues in sentences to predict the distinctions. using these automatic methods, we show how significant accuracies above baseline can be obtained for separating out articles according to text quality. for the empirical studies reported later in the paper, we divide our data into three sets: best is the collection of great articles from the corpus. no negative samples are used in this set for our analysis. we analyze the distinction between very good and typical articles matched by topic using the ranking corpus we described in section 3.3. these articles are divided into two sets. development corpus comprised of 500 pairs of topically-matched very good and typical articles. these pairs were randomly chosen. test corpus of the remaining (very good, typical) pairs, totaling 24,532. 5. specificity of content this aspect is based on the idea that texts do not convey information at a constant level of detail. there are portions of the text where the author wishes to convey only the main topic and keeps such content devoid of any details. certain other parts of the text are reserved for specific details on the same topics. for example, consider the two sentences below. dr. berner recently refined his model to repair an old inconsistency. [general] the revision, described in the may issue of the american journal of science, brings the model into closer agreement with the fact of wide glaciation 440 million years ago, yielding what he sees as stronger evidence of the dominant role of carbon dioxide then. [specific] the first conveys only a summary of dr. berner’s new research finding and alerts the reader to this topic. the next sentence describes more specifically what was achieved as part of dr. berner’s research. we put forward the hypothesis that the balance between general and specific content can be used to predict the text quality of articles. to facilitate this analysis, we created annotations of general and specific sentences for three of our articles and also developed a classifier that can automatically annotate new articles with high accuracy. using features derived from the classifier’s predictions, we obtained reasonable success in distinguishing the text quality levels in our corpus. this section describes the general-specific distinction, the annotations that we obtained for this aspect and experiments on quality prediction. if texts gave only specific information, it would read like a bulleted list and less of an article or essay. in the same way that paragraphs divide an article into topics, authors introduce general 97 louis and nenkova statements and content that abstract away from the details to give a big picture. of course these general statements need to be supported with specific details elsewhere in the text. if the text only conveyed general information, readers would regard it as superficial and ambiguous. details must be provided to augment the main statements made by the author. books that provide advice about writing (swales and feak, 1994; alred et al., 2003) frequently emphasize that mixing general and specific statements is important for clarity and engagement of the reader. this balance between general and specific sentences could be highly relevant for science news articles. these articles convey expert research findings in a simple manner to a lay audience. science writers need to present the material without including too much technical details. on the other hands, an overuse of general statements such as quotes from the researchers and specialists can also lower the reader’s attention. similarly, conference publications of research findings have been found to have a distinctive structure with regard to general-specific nature (swales and feak, 1994). these articles have a hourglass like structure with general material presented in the beginning and end and specific contents in between. given that even articles written for an expert audience have such a structure, we expected that the distinction would also be relevant for the genre of science news. yet while widely acknowledged to be important for writing quality, there is no work that explores how general and specific sentences can be identified and used for predicting the quality of articles. our previous work (louis and nenkova, 2011a,b) presented the first study on annotating general-specific nature, building an automatic classifier for the aspect and using it for quality prediction. this work was conducted on articles from the news genre. in this section, we provide some details about this prior study and then explain how similar annotations and experiments were performed for articles from our corpus of science news. in our prior work we focused on understanding the general-specific distinction and obtaining annotated data in order to train a classifier for the task. there was no annotated data for this aspect before our work but we observed that an approximate but suitable distinction has been made in existing annotations for certain types of discourse relations. specifically, the penn discourse treebank contains annotations for thousands of instantiation discourse relations which are defined to hold between two adjacent sentences where the first one states a fact and the second provides an example of it. these annotations were created over wall street journal articles which mostly discuss financial news. as a rough approximation, we can consider the first sentence as general and the second as specific. these general and specific sentences from the instantiation annotations, disregarding pairing information, provided 1403 examples for each type. these sentences also served as development data to observe properties associated with general and specific types. we developed features that captured sentiment, syntax, word frequency, lexical items in these sentences and the combination of features had a 75% success rate for distinguishing the two types on this approximate data. we then obtained a new held-out test set by eliciting direct annotations from people. we designed an annotation scheme and recruited judges on amazon mechanical turk who marked sentences from news articles. we used news from two sources, associated press and wall street journal to obtain sentences for annotation. the annotators had fairly good agreement on this data. we used our classifier trained on discourse relations to predict the sentence categories provided by our judges and our accuracy remained consistent at 75% as on the instantiations data. also in prior work, we successfully used the automatic prediction of sentence specificity to predict writing quality (louis and nenkova, 2011b). we obtained general-specific categories on a 98 a corpus of science journalism for analyzing writing quality corpus of news summaries using our automatic classifier (trained on discourse relations). these summaries were created by multiple automatic summarization systems and the corpus also has manually assigned content and linguistic quality ratings for these summaries by a team of judges. we developed a specificity score that combined the general-specific predictions at sentence level to provide a characterisation for the full summary text. this specificity score was shown to be a significant component for predicting the quality scores on the summaries. the general trend was that summaries that were rated highly by judges showed greater proportion of general content compared to the lower rated summaries. in the next sections, we explore annotations and classification experiments for the science writing genre. 5.1 annotation guidelines we chose three articles from the great category of our corpus for annotation. each article has approximately 100 sentences, providing a total of 308 sentences. the task for annotators was to provide a judgement for a given sentence in isolation, without the surrounding context in which the sentence appeared. we used amazon mechanical turk to obtain the annotations and each sentence was annotated by five different judges. the judges were presented with three options—general, specific and cannot decide. sentences were presented in random order, mixing different texts. the judges were given minimal instructions about criteria they should apply while deciding what label to choose for a sentence. our initial motivation was to identify the types of sentences where intuitive distinctions are insufficient, in order to develop more detailed instructions for manual annotation. the complete instructions were worded as: “sentences could vary in how much detail they contain. one distinction we might make is whether a sentence is general or specific. general sentences are broad statements made about a topic. specific sentences contain details and can be used to support or explain the general sentences further. in other words, general sentences create expectations in the minds of a reader who would definitely need evidence or examples from the author. specific sentences can stand by themselves. for example, one can think of the first sentence of an article or a paragraph as a general sentence compared to one which appears in the middle. in this task, use your intuition to rate the given sentence as general or specific.” examples: (g indicates general and s specific) [g1] a handful of serious attempts have been made to eliminate individual diseases from the world. [g2] in the last decade, tremendous strides have been made in the science and technology of fibre optic cables. [g3] over the years interest in the economic benefits of medical tourism has been growing. [s1] in 1909, the newly established rockefeller foundation launched the first global eradication campaign, an effort to end hookworm disease, in fifty-two countries. [s2] solid silicon compounds are already familiar–as rocks, glass, gels, bricks, and of course, medical implants. [s3] einstein undertook an experimental challenge that had stumped some of the most adept lab hands of all time–explaining the mechanism responsible for magnetism in iron. 99 louis and nenkova agreement general specific 5 32 (25.6) 50 (27.7) 4 48 (38.4) 73 (40.5) 3 45 (36.0) 57 (31.6) total 125 180 no majority = 3 table 7: the number (and percentage of sentences) for different agreement levels. the numbers are grouped by the majority decision as general or specific. agreement 5 general at issue is whether the findings back or undermine the prevailing view on global warming. specific the sudden popularity of pediatric bipolar diagnosis has coincided with a shift from antidepressants like prozac to far more expensive atypicals. agreement 3 general “one thing is almost certain,” said dr. jean b. hunter, an associate professor of biological and environmental engineering at cornell. specific the other half took seroquel and depakote. table 8: example sentences for different agreement levels 5.2 annotation results table 7 provides the statistics about annotator agreement. we do not compute the standard kappa scores because different groups of annotators judged different sentences. instead, we report the fraction of sentences which had 5, 4 or 3 annotators agreeing on the class. when fewer than three annotators agree on the class there is no majority decision, and this situation occurred only for three out of the 308 annotated sentences.3 we found that two-thirds of the sentences were annotated with high agreement: either four or all the five judges gave these sentences the same class. the other one third of the sentences are on the borderline with a slim majority, three out of five judges assigning the same class. table 8 shows some examples of sentences with full agreement and those that have only agreement of 3. these sentences with low agreement have a mix of both types of content. for example, the general sentence with agreement 3 has a quote the content of which is general, at the same time, the descriptions of the person presents specific information. such sentences are genuinely hard to annotate. we hypothesize that we can address this problem better by selecting a different granularity such as clauses rather than sentences. our annotations were also obtained by presenting each sentence individually but another aspect to explore is how these judgements would vary if the context of adjacent sentences is also provided to the annotators. for example, the presence of pronouns, discourse connectives and such references in a sentence taken out of context (for example the last sentence in table 8) could make the sentence appear more vague but in context, they may be interpreted differently. in terms of distribution of general and specific sentences, table 7 shows that there were more specific (59%) sentences than general on this dataset. we expect that this distribution would vary by 3. for the three sentences that did not obtain majority agreement, two of the judges had picked general, two chose specific and the fifth label is ‘cannot decide’. 100 a corpus of science journalism for analyzing writing quality transition type number (%) gs 66 (22) gg 58 (19) ss 114 (38) sg 63 (21) table 9: number and percentage of different transition types. g indicates general and s specific. block type g1 g2 gl s1 s2 sl number 35 19 13 21 17 28 % 26 14 10 16 13 21 table 10: number and percentage of different block types. g1, g2 and gl indicate general blocks of size 1, 2 and >=3. similarity for the specific category. genre. in the annotations that we have collected previously on wall street journal and associated press articles, we found wall street journal articles to contain more general than specific content while the associated press corpus collectively had more specific than general (louis and nenkova, 2011a). this finding could result from the fact that our chosen wsj articles are written with a more ‘essay-like’ style, for example one article discusses the growing use of personal computers in japan, while the ap articles were short news reports. a simple plot of the sentence annotations for the three articles (figure 1) shows how the annotated general and specific sentences are distributed in the articles. the x axis indicates sentence position (starting from 0) normalized by the maximum value. on the y axis, 1 indicates that the sentence was annotated as specific (any agreement level) and -1 otherwise. these plots show that often there are large sections of continuous specific content and the general sentence switch happens between these blocks. some of these specific content blocks on examination tended to be detailed research descriptions. an example is given below. in this snippet, a general sentence is followed by 8 consecutive specific sentences. this finding is interesting because it shows that research details are padded with general content providing the context of the research. (general) still, no human migration in history would compare in difficulty with reaching another star. the nearest, alpha centauri, is about 4.4 light-years from the sun, and a light-year is equal to almost six trillion miles. the next nearest star, sirius, is 8.7 light-years from home. to give a graphic sense of what these distances mean, dr. geoffrey a. landis of the nasa john glenn research center in cleveland, pointed out that the fastest objects humans have ever dispatched into space are the voyager interplanetary probes, which travel at about 9.3 miles per second. ”if a caveman had launched one of those during the last ice age, 11,000 years ago,” dr. landis said, ”it would now be only a fifth of the way toward the nearest star.” ... (5 more specific sentences) these patterns are also indicated quantitatively in tables 9 and 10. table 9 gives the distribution of transitions between general and specific sentences, calculated from pairs of adjacent sentences only. we excluded the three sentences that did not have a majority decision. in this table, g stands for general, s for specific and the combinations indicate a pattern between adjacent sentences. the most frequent transition is ss, specific sentence immediately followed by another specific sentence, 101 louis and nenkova figure 1: specificity of annotated sentences which accounts for 40% of the total transitions. other types are distributed almost equally, 20% each. table 10 shows the sizes of contiguous blocks of general and specific sentences. the minimum block size is one whereas the maximum turned out to be 7 for general and 8 for specific. 5.3 automatic prediction of specificity this section explains the classifier that we developed for predicting sentence specificity and results on the data annotated from the science news articles. our classifier trained on the instantation relations data obtains 74% accuracy in making the binary general or specific prediction on the science news data we described in the previous section. this high accuracy has enabled us to create specificity predictions on a larger scale for our full corpus and analyze how the aspect is related to text quality. the experiment on text quality is reported in the next section. our features use different surface properties that indicate specificity. for example, one significant feature indicative of general sentences is plural nouns which often refer to classes of things. other features involve word-level specificity, counts of named entities and numbers, and likelihood under language model (general sentences tend to have lower probability than specific under a language model of the domain), and binary features indicating the presence of each word. these features are based on syntactic information as well as ontology such as wordnet. for example, the word-level specificity feature is computed using the path length from the word to root of wordnet via hypernym relations. longer paths could be indicative of a more specific word, while words that are closer to the root could be related to more general concepts. we also add features that capture the length of adjective, adverb, preposition and verb phrases. to obtain the syntax features, we parsed the sentences using the stanford parser (klein and manning, 2003). we call the word features lexical and the set of all other features are non-lexical. descriptions of these features and their 102 a corpus of science journalism for analyzing writing quality example type size lexical non-lexical all features agree 5 82 74.3 92.2 82.9 agree 4+5 203 66.0 85.2 76.3 agree 3+4+5 305 58.3 74.4 67.2 table 11: accuracies for automatic classification of sentences on sentences with different agreement levels. agree 4+5 indicates that sentences with both agreement 5 and 4 were combined. agree 3+4+5 indicates all the annotated examples. individual strengths are given in louis and nenkova (2011a). here we report the results on categories of features (table 11) on the science news annotations (302 sentences). (we excluded three sentences that did not have majority judgement.) we used a logistic regression classifier available in r (r development core team, 2011) for our experiment. we take the majority decision by annotators as the class of the sentences in our test set. specific sentences comprise 59% of the data and so the majority class baseline would give 59% accuracy. the accuracy of our features is given in table 11. for this test dataset (all sentences where a majority agreement was obtained) the accuracies are indicated in the last line of the table. the best accuracy is obtained by the non-lexical features, 74.4%. the word features are much worse. one reason for poor performance of the word features could be that our training sentences are taken from articles published in the wall street journal. the lexical items from these finance news articles may not appear in science news from the new york times leading to low accuracy. the lexical features could also suffer from data sparsity and need large training sets to perform well. on the other hand, the non-lexical indicators appear to carry over smoothly. the performance of these features on science news is close to the accuracy when the wsj trained classifier was tested on held-out sentences also from wsj. the combination of lexical and non-lexical features is not better than the non-lexical features alone. when we consider examples with higher agreement, rows 1 and 2 of the table, the results for all feature classes are higher, reaching even 92.2% for the non-lexical features. in further analyses, we found that the confidence from the classifier is also indicative of the agreement from annotators. when the classifier made a correct prediction, the confidence from the logistic regression classifier is higher on examples which had more agreement. on examples on the borderline between general and specific classes, even when the prediction was correct, the confidence was quite low. the trend for wrong examples was opposite. examples with high agreement were predicted wrongly with lower confidence whereas the borderline cases had high confidence even when predicted wrongly. this pattern suggested that the confidence values from the classifier can be used as generality/specificity scores in addition to the binary distinction. 5.4 relationship to writing quality since automatic prediction could be done with very good accuracies, we used our classifier to analyze the different categories of articles in our corpus. we trained a classifier on data derived from the discourse annotations and using the set of all non-lexical features. we obtained the predictions from the classifier for each sentence in our corpus and then composed several features to indicate specificity scores at article level. these features are described below. 103 louis and nenkova feature mean value in category p-value from t-test very good typical spec wtd 0.56 0.58 0.004 sl 0.09 0.10 0.035 table 12: features related to general and specific content which had significantly different mean values in the very good and typical categories. the p-value from a two-sided t-test is also indicated. overall specificity: one feature indicates the fraction of specific sentences in the article (spec fraq). since we use a logistic regression classifier for predicting specificity, we also obtain from the classifier a confidence value for the sentence belonging to each class. we use the confidence of each sentence belonging to the ‘specific’ class and compute the mean and variance of this confidence measure across the article’s sentences as features (spec mean, spec var). but these scores do not consider the lengths of the different sentences. a longer sentence constitutes a greater portion of the content of the article and its general or specific nature should have higher weight while composing the score for the full article. for this purpose, we also add another feature—the weighted average of the confidence values where the weights are the number of words in the sentence (spec wtd). sequence features: proportions of different transition and block types that we discussed in section 5.2 are added as features (gg, gs, ss, sg, g1, g3, gl, s1, s2, sl). in contrast to the manual markings in that section, here the transitions are computed using the automatic annotations from the classifier. we first tested how these features vary between good and average writing using a random sample of 1000 articles taken from the great and very good categories combined and another 1000 taken from the typical category. no pairing information from the ranking corpus was used during this sampling as we wanted to test overall if these features are indicative of good articles rather than their variation within a particular topic. a two-sided t-test between the feature values in the two categories showed the mean values for two of the features to vary significantly (p-value less than 0.05). the test statistic and mean values for the significant features are shown in table 12. the good articles have a lower degree of specific content (measured by the confidence weighted score spec wtd). the other trend was a higher proportion of large specific blocks (sl) in the typical articles compared to the good ones. both these features indicate that more specific content is correlated with the typical articles. none of the transition features were significantly different in these two sets. all the features (significant and otherwise) were then input to a classifier that considered a pair of articles and decided which article is the better one. the features for the pair are the difference in feature values for the two individual articles. a random baseline would be accurate 50% of the time. we used the test set described in section 4 and performed 10-fold cross validation over it using a logistic regression classifier. the accuracy using our features is 54.5% indicating that the general-specific distinction is indicative of text quality. the improvement is low but statistically significant. when the features were analyzed by category, the specificity levels gave 52% and the transition/block features gave 54.4% accuracy. so both aspects are helpful for making the general-specific distinction but the transition features are stronger than overall specificity levels. 104 a corpus of science journalism for analyzing writing quality the combination of the two feature types is not much different than transition features alone. but we expect both feature sets to be useful when combined with other aspects of writing not related to specificty. these results provide evidence that content specificity and the sequence of varying degrees of specificity are predictive of writing quality. they can be annotated with good agreement and we can design automatic ways to predict the distinction with high accuracy. 6. communicative goals of the article our second line of investigation concerns the intentions or communicative goals behind an article. this work is based on the idea that an author has a purpose for every article that he writes and the structure that he chooses for organizing the article’s content is such that it will help him to convey his purpose. a detailed theory of intentional structure and its role in the coherence of articles was put forth in an influential study by grosz and sidner (1986). the theory has two proposals: that the overall intention for an article is conveyed by the combined intentions of smaller level discourse segments in the article and that the smaller segments are linked by relations which combine their purposes. since the organization of low-level intentions in the article contributes to the coherence by which the overall purpose is conveyed, if a method of identifying intentions for small text segments is developed we can use it for text quality prediction. this section describes our attempt to develop an annotation method for sentence-level communicative goals in science news articles and the usefulness of these annotations for predicting well-written articles. for the genre of science writing, the author’s goal is to convey information about a particular research study, its relevance and impact. at a finer level in the text, many different types of sentences contribute to achieving this purpose. some sentences define technical terms, some provide explanations and examples, others introduce the people and context of the research and some sentences are descriptive and present the results of the studies. science articles are also sometimes written as a story and this style of writing creates sentences associated with the narrative and dialog in the story. there are numerous books on science writing (blum et al., 2006; stocking, 2010) that prescribe ideas for producing well-written articles in different styles—narratives, explanatory pieces and interviews. such books often point to examples of specific sentences in articles that convey an idea effectively. for example, we might expect that good science articles are informative and contain more sentences that define and explain aspects of the research. similarly, well-written articles might also provide examples after a definition or explanation. so we expect that a large scale annotation of sentences with communicative goals and building classifiers to automatically do such annotations can be quite useful for text quality research. studies on register variation (biber, 1995; biber and conrad, 2009) provide evidence that there is noticeable variation in linguistic forms depending on the communicative purpose and situations. there have also been successful efforts to manually annotate and develop automatic classifiers for predicting communicative goals on the closely related genre of academic writing ie. conference and journal publications. these articles have a well-defined purpose, also for different sections of the paper. studies in this area have identified that conference publications have a small set of goals such as motivation, aim, background, results and so on (swales, 1990; teufel et al., 1999; liakata et al., 2010). such well-definable regularities in academic writing have lead to studies that manually annotated these sentence types and developed supervised methods to do the annotation 105 louis and nenkova definition in this theory, a black hole is a tangled melange of strings and multidimensional membranes known as “d-branes.” example to see this shortcoming in relief, consider an imaginary hypothesis of intelligent design that could explain the emergence of human beings on this planet. explanation if some version of this hypothesis were true, it could explain how and why human beings differ from their nearest relatives, and it would disconfirm the competing evolutionary hypotheses that are being pursued. results the new findings overturn almost everything that has been said about the behavior and social life of the mandrillus sphinx to date, and also call into question existing models of why primates form social groups. about people a professor in harvard’s department of psychology, gilbert likes to tell people that he studies happiness. story line “happiness is a seven,” she said with a triumphant laugh, checking the last box on the questionnaire. table 13: examples for some sentence types from the science news corpus automatically (teufel and moens, 2000; guo et al., 2011). subsequently, these annotations have been used for automatic summarization and citation analysis tasks. however, our annotation task for science journalism is made more difficult by the wider and unconstrained styles of writing in this genre. the number of communicative goals is much higher and varied compared to academic writing. in addition, academic articles are also structured into sections whereas there is no such subdivision in the essay-like articles in our genre. perhaps the most difficult aspect is that even coming up with a initial set of communicative goals to annotate is quite difficult. table 13 shows examples for some of the sentence types we wish to annotate for our corpus based on our intuitions. but these categories only capture some of the possible communicative goals. as an alternative to predefining catgories of interest, we explored an unsupervised method to group sentences into categories. our method relies on the idea that similarities in syntactic structure could indicate similarities in communicative goals (louis and nenkova, 2012a). for example, consider the sentences below. a. an aqueduct is a water supply or navigable channel constructed to convey water. b. a cytokine receptor is a receptor that binds cytokines. both sentences a and b are definitions. they follow the regular pattern for definitions which consists of the term to be defined followed by a copular predicate which has a relative or reduced relative clause. in terms of syntax, these two sentences have a high degree of similarity even though their content is totally different. other sentences types such as questions also have a unique and well-defined syntax. we implemented this hypothesis in the form of a model that performs syntactic clustering to create categories and assigns labels to sentences indicating which category is closest to the syntax of the sentence (louis and nenkova, 2012a). then the sequence of sentences in an article can be analyzed in terms of these approximate category label sequences. we ran our approach on a corpus of academic articles. we showed that the sequence of labels in well-written and coherent articles is different from that in an incoherent article (created artificially by randomly permuting the sentences in the original article). we also found that the automatic categories created by our model have some 106 a corpus of science journalism for analyzing writing quality correlation to categories used in the manual annotations of intentional structure for the academic genre. given this success on academic articles, we expected that syntactic clustering could help us uncover some of the approximate categories prevalent in science journalism. such a method is in fact more suitable for our genre where it is much harder to define the intention categories. our plan for using the approximate categories is two-fold. firstly, we want to test how much these approximate categories would align with our intuitions about sentence types in this genre and also how accurately we can predict text quality using sequences of these categories. on the other hand, we expect this analysis to help create guidelines for a fuller manual annotation of communicative goals on these articles. the clusters would give us a closer view of different types of sentences present in the articles and enable us to define categories we should annotate. in addition, when a sentence is annotated, annotators can look at similar sentences from other articles using these clusters and these examples would help them to make a better decision on the category. 6.1 using syntactic similarity to identify communicative goals our approach uses syntactic similarity to cluster sentences into groups with (hopefully) the same communicative goal. the main goal is to capture these categories and also obtain information about what are the likely sequences of such categories in good and typical articles. to address this problem, we employ a coherence model that we have recently developed (louis and nenkova, 2012a). our implementation uses a hidden markov model (hmm). in the hmm, each state represents the syntax of a particular type of sentence and transitions between two states a and b indicates how likely it is for a sentence of type b to follow a sentence of type a in articles from the domain. we train the model on examples of good writing and learn the sentence categories and transition patterns on this data. we then apply this model to obtain the likelihood of new articles. poorlywritten articles would obtain a lower probability compared to well-written articles. accordingly, for our corpus, we use the articles from the best category (great) for building the model. we also need to define a measure of similarity for clustering the sentences. our similarity measure is based on the constituency parse of the sentence. we represent the syntax of each sentence as a set of context free rules (productions) from its parse tree. these productions are of the form lhr→ rhs where lhs is a non-terminal in the tree whose children are the set of non-terminals and terminal nodes indicated by rhs. we compute features based on this representation. each production serves as a feature and the feature value for a given sentence is the number of times that production occurred in the parse of the sentence. we obtained the parse trees for the sentences using the stanford parser (klein and manning, 2003). all sentences from the set of articles are clustered according to these syntactic representations by maximizing the average similarity between items in the same cluster. these clusters form the states of the hmm. the emissions from a state are given by a unigram language model computed from the productions of the sentences clustered into that state. transitions between states are computed as a function of number of sentences from cluster a that follow those of cluster b in the training articles. we also need to decide how many clusters are necessary to adequately cover the sentences types in a collection of articles. we pick the number of hidden states (sentence types) by searching for the value that provides the best accuracy for differentiating good and poorly-written articles in our development set (see section 4). specifically, for each setting of the number of clusters (and other 107 louis and nenkova model parameters) we obtained the perplexity of the good and typical articles under the model. the article with lower perplexity is considered as the one predicted as better written. we use perplexity rather than probability to avoid the influence of article length on the score. this development set contained 500 pairs of good and typical articles. the smoothing parameters for emission and transition probabilities were also tuned on this development data. the parameters that gave the best accuracy were selected and the final model trained using these settings. 7. analysis of sentence clusters there are 49 clusters in our final model trained on the best articles. we manually analyzed each cluster and marked whether we could name the intended communicative goal on the basis of example sentences and descriptive features. the descriptive features were computed as follows. since the sentences were clustered on the basis of their productions and a cosine similarity function was used, we selected those productions which contribute to more than 10% of the average intra-cluster cosine similarity. the intrepretability of the clusters varied greatly. there were 27 clusters for which we could easily provide names. some of the remaining clusters also had high syntactic similarity, however we could not easily interpret them as related to any communicative goal. so we have not included them in our discussion. in this section, we provide descriptions and examples for sentences in each of these clusters. we ignore the remaining clusters for now but their presence indicates that further in-depth analyses of sentences types must be conducted in order to be able to create annotations that cover all sentences. an overview of the 27 clusters is given in tables 14 and 15. we also indicate the descriptive features and some example sentences for each cluster. for the example sentences, we have also marked the span covered by the lhs of each descriptive production. these are indicated with ‘[’ and ‘]’ braces attached with the label of the lhs non-terminal. we also divide these clusters into categories for discussion. some clusters are created such that the sentences in them have matching productions that cover long spans of texts in the sentences. the descriptive features for these clusters are productions that involve the full sentence or complete predicates of sentences. such clusters are in contrast to others where the sentences are similar but only based on productions that involve rather small spans of text. for example the descriptive productions for these clusters could involve smaller syntactic categories such as noun or prepositional phrases. we provide a brief study of our clusters below. sentence-level matches. clusters a, b and c correspond to quotes, questions and fragments (table 14) respectively. the match between the sentences in all these clusters is on the basis of sentence level productions as indicated by the most descriptive features. for example, root→sq the descriptive production for cluster 2 is a topmost level production which groups the yes/no questions. for quotes-related clusters, the descriptive features have quotation marks enclosing an s node or a np vp sequence. an example sentence from this category is shown below. root-[is there one way to get parkinson’s, or 50 ways? ”]-root category c has many sentences which are quite short and have a style that would be unconventional for writing but would be more likely in dialog. they are fragment sentences that create a distinctive rhythm: “in short, no science”; “and the notion of mind?”; “after one final step, this.” these clusters were easiest to analyze and directly matched our intuitions about possible communicative goals. 108 a corpus of science journalism for analyzing writing quality predicate-level match. for clusters in categories d, e, f and g, the descriptive features often cover the predicate of the sentences. for clusters in category d several of these predicates are adjectival phrases as in the following example. the number of applications vp-[is adjp-[mind-boggling]-vp]-adjp . the primary goal of these sentences is to describe the subject and in our examples these phrases often provided opinion and evaluative comments. other clusters in this category have clauses headed by ‘wh’ words which also tend to provide descriptive information. category e has a number of clusters where the main verb is in past tense but with different attributes. for example, in cluster 12, the predicate has only the object or the predicate has a prepositional phrase in addition to the object as in cluster 11. most of these sentences (two are shown below) tended to be part of the narrative or story-related parts of the article. (i) he vp-[entered the planet-searching business through a chance opportunity]-vp . (ii) s-[he vp-[did a thorough physical exam]-vp and vp-[took a complete history]-vp .]-s categories f and g comprised sentences with adverbs, discourse connectives or modal verbs. these sentences show how comparisons, speculations and hypotheses statements are part of motivating and explaining a research finding. an example is shown below. if a small fraction of the subuniverses vp-[can support life]-vp , then there is a good chance that life vp-[will arise somewhere]-vp, dr. susskind explained. entity or verb category matches. in categories h and i, we have listed some of the clusters where the descriptive features tended to be capturing only one or two words at most. so the descriptive feature does not relate the sentences by their overall content but are based on a phrase such as np. the two clusters under category h have proper names as the descriptive features as in the sentence below. dr. strominger remembered being excited when he found a paper by the mathematician np-[dr. shing-tung yau]-np , now of harvard and the chinese university of hong kong . these sentences have either two or three proper name tokens in sequence and were often found to introduce a named entity which is why they are referred to using their full names. these sentences often have an appositive following the name which provides more description about the entity. other clusters group different types of noun phrases and verbs. for example, cluster 27 groups sentences containing numbers and 42 has sentences with plural nouns. the top descriptive feature for cluster 15 is present tense verbs. some examples from these clusters with their descriptive features are shown below. (i) the big isolated home is what loewenstein , np-[48]-np , himself bought. (ii) the news from np-[environmental organizations]-np is almost always bleak. (iii) the first departments we vp-[checked]-vp were the somatosensory and motor regions. (iv) periodically, he went to methodist hospital for imaging tests to measure np-[the aneurysm’s size]-np . 109 louis and nenkova while some intrepetation can be provided for the communicative goals corresponding to these sentences, these clusters are somewhat divergent from our idea of detecting communicative goals. the similarity between sentences in these clusters is based on a rather short span such as a single noun phrase and less likely to give reliable indications of the intention of the full sentence. overall the clusters discovered by the automatic approach were meaningful and coherent, however, they did not not match our expected communicative goals from table 13. while we expected several categories related to explaining a research finding such as definition, examples and motivation, we did not find any clearly matching clusters that grouped such sentences. we hypothesize that this problem could arise because several science articles embed research related sentences as part of a larger narrative or story and information about scientists and other people. hence the input to clustering is a diverse set of sentences and clustering places research and non-research sentences in the same cluster. this issue can be addressed by separating out sentences which describe research findings from those that are related to the story line as a first step. the clusters can then be built on the two sets of sentences separately and we expect their output to be closer to the types we had envisioned. we would then obtain clusters related to the core research findings that are reported and a separate set of clusters related to the story line of the articles. we plan to explore this direction in our future work. however, we believe that this experiment has provided good intuitions and directions for further annotation of communicative goal categories for our corpus. 7.1 accuracy in predicting text quality despite the fact that the clusters are different from what we expected to discover, we evaluate if these automatic and approximate sentence types contribute to the prediction of text quality for our corpus. good results on this task would provide more motivation for performing large scale annotation of communicative goals. we describe our experiments in this section and show that positive results were obtained using these approximate sentence categories. for each article in our test set (details in section 4), we obtained the most likely state sequence under the hmm model using the viterbi decoding method. this method assigns a category label to each sentence in the article such that the sequence of labels is the most likely one for the article under the model. then features were computed using these state predictions. each feature is the proportion of sentences in the article that belong to a state. we first analyzed how these features vary between a random set of 1000 good articles and a set of 1000 typical articles, without topic pairing. a two-sided t-test was used to compare the feature values in the two categories. table 16 shows the 10 clusters whose probability was significantly different (p-value less than 0.05) between the good and typical categories. 110 a corpus of science journalism for analyzing writing quality c lu st er s d es cr ip tiv e fe at ur es e xa m pl e se nt en ce s a. q uo te s: se nt en ce s fr om di al og in th e ar tic le s 1 s→ n p v p .” s[w e w er e so m et im es la x ab ou tp er so na ls af et y .” ]s 6 s→ ” n p v p .” s[“ t ha t’ s on e re as on pe op le ar e in te re st ed in tim e tr av el .” ]s 10 s→ ” s ,” n p v p .; v p→ v b d s[” it ho ug ht b en m ad e an er ro r, ” he v p[s ai d] -v p .]s b. q ue st io ns :d ir ec tq ue st io ns ,s om e ar e rh et or ic al ,s om e ar e ye s/ no ty pe 2 r o o t → sq r o o t -[ w ou ld th ey kn ow ab ou ts ta rs ?] -r o o t 3 r o o t → sb a r q ;s b a r q → w h a d v p sq . sb a r q -[ r o o t -[ w hy ar e at om s so tin y an d st ar s so bi g ?] -] r o o t ]sb a r q c. c on ve rs at io n st yl e: se nt en ce s th at ar e un co nv en tio na li n w ri tte n te xt s bu tm or e lik el y in di al og 4 r o o t → fr a g r o o t -[ fo rw ha t? ]r o o t d. d es cr ip tio ns :t he pr ed ic at e is m ai nl y in te nd ed to pr ov id e de sc ri pt io n of te n co nt ai ni ng ev al ua tiv e co m m en ts an d op in io n 5 v p→ v b d a d jp ;a d jp → jj sh e kn ew on ly th at it v p[w as a d jp -[ co nt ag io us ]v p] -a d jp an d th at it v p[w as a d jp -[ na st y] -v p] -a d jp . 7 v p→ v b z a d jp a d jp → jj e m ot io n v p[i s ce nt ra lt o w is do m ]v p, ye td et ac hm en tv p[i s a d jp -[ es se nt ia l] -v p] -a d jp . 13 a d jp → jj s a nd ye ta s ps yc ho lo gi st s ha ve no te d, th er e is a yi nya ng to th e id ea th at m ak es it a d jp -[ di ffi cu lt to pi n do w n] -a d jp . 20 v p→ v b d n p ;s b a r → w h n p s l at er ,a rd el ts et tle d on 39 qu es tio ns sb a r -[ th at ,i n he rj ud gm en t, v p[c ap tu re d th e el us iv e co nc ep to fw is do m ]v p] -s b a r 29 a d jp → jj ;v p→ v b p a d jp c os m ol og is ts ha ve fo un d to th ei ra st on is hm en tt ha tl if e an d th e un iv er se v p[a re a d jp -[ st ra ng el y] -a d jp an d de ep ly co nn ec te d] -v p . 30 sb a r → w h a d v p s ;w h a d v p→ w r b t he pr ob le m ,a s g ilb er ta nd co m pa ny ha ve co m e to di sc ov er ,i s th at w e fa lte r w h a d v p[s b a r -[ w he n] -w h a d v p it co m es to im ag in in g w h a d v p[s b a r -[ ho w ]w h a d v p w e w ill fe el ab ou ts om et hi ng in th e fu tu re ]sb a r ]sb a r . e. st or y lin e: t he se cl us te rs m os tly ca pt ur e ev en ts as so ci at ed w ith th e st or y pr es en te d in th e ar tic le 11 v p→ v b d n p pp w e v p[s aw ac tiv ity in th e hi pp oc am pu s] -v p . 12 v p→ v b d n p ;s → n p v p . t he m or e co m pl ex th e ta sk ,t he m or e v p[d is pe rs ed th e br ai n ’s ac tiv ity ]v p . ta bl e 14 :e xa m pl e cl us te rs di sc ov er ed by sy nt ac tic si m ila ri ty .m ul tip le de sc ri pt iv e fe at ur es fo ra cl us te ra re se pa ra te d by a ‘; ’. 111 louis and nenkova c lu st er s d es cr ip tiv e fe at ur es e xa m pl e se nt en ce s f. d is co ur se re la tio ns /n eg at io n: se nt en ce s th at ar e pa rt of co m pa ri so n an d ex pa ns io n re la tio ns an d th os e ha vi ng ad ve rb s 31 v p→ v b z r b v p ;s → c c n p v p . s[b ut th e pi ct ur e, as m os tg oo d ev ol ut io na ry ps yc ho lo gi st s po in to ut ,i s m or e co m pl ex th an th is .]s g. m od al s/ ad ve rb s: sp ec ul at io ns ,h yp ot he si s. so m et im es ad ve rb s ar e us ed w ith th e m od al s in cr ea si ng th e in te ns ity 40 v p→ m d v p ;s b a r → in s t he ne xt bi g m ile st on e, as tr on om er s sa y, v p[w ill be th e de te ct io n of e ar th -s iz e pl an et s, sb a r -[ al th ou gh th at v p[w ill re qu ir e go in g to sp ac e] -v p] -v p] -s b a r . 41 v p→ v b n p ;v p→ to v p a fe w al ie n ba ct er ia in a m ud pu dd le so m ep la ce w ou ld v p[c ha ng e sc ie nc e] -v p . 46 v p→ m d a d v p v p ;a d v p→ r b it is ab ou ta ki nd of en er gy w e a d v p[o ft en ]a d v p ru e bu tv p[w ou ld a d v p[s ur el y] -a d v p m is s. ]v p h. a bo ut na m ed en tit ie s: se nt en ce s w he re th e m ai n fo cu s is a pe rs on or ot he rn am ed en tit y 14 n p→ n n p n n p n n p “i tt ou ch es on ph ilo so ph ic al is su es th at sc ie nt is ts of te nt im es sk ir t,” sa id n p[d r. jo hn sc hw ar z] -n p ,a ph ys ic is ta nd st ri ng th eo ri st at th e c al if or ni a in st itu te of te ch no lo gy . 45 n p→ n n p n n p n p[d r. l om bo rg ]n p al so ta ke s is su e w ith so m e gl ob al w ar m in g pr ed ic tio ns . i. o th er s: no un ph ra se s w ith di ff er en ta ttr ib ut es ,d iff er en tt yp es of ve rb ph ra se s 27 n p→ c d it al so m ad e m e cu ri ou s ab ou tc la yt on ,w ho di sa pp ea re d fr om ac ad em ia in n p[1 98 1] -n p . 39 n p→ n n p n p[b en ]n p w as on ce a m ai lc ar ri er an d a fa rm er an d ca ttl e ra nc he r. 26 n p→ n p n n ;n p→ n n p po s d r. h aw ki ng ’s ce le br at ed br ea kt hr ou gh re su lte d pa rt ly fr om a fig ht . 34 n p→ jj n n to da te ,t he pr op on en ts of n p[i nt el lig en td es ig n] -n p ha ve no tp ro du ce d an yt hi ng lik e th at . 42 n p→ jj n n s n p[t ro pi ca lf or es ts ]n p ar e di sa pp ea ri ng . 15 v p→ v b z a rd el ta ck no w le dg es th at no on e re al ly kn ow s w ha tw is do m v p[i s] -v p . 32 v p→ v b d b ut d r. d eb ak ey ’s re sc ue al m os tn ev er v p[h ap pe ne d] -v p . ta bl e 15 : e xa m pl e cl us te rs di sc ov er ed by sy nt ac tic si m ila ri ty (c on tin ue d fr om ta bl e 14 )) . m ul tip le de sc ri pt iv e fe at ur es fo r a cl us te r ar e se pa ra te d by a ‘; ’. 112 a corpus of science journalism for analyzing writing quality cluster type mean proportion in category p-value from no. very good typical t-test higher probability in good articles 5 descriptions 0.004 0.003 0.038 30 descriptions 0.085 0.079 0.001 40 modals 0.028 0.026 0.014 34 others (nouns with adjective modifiers) 0.019 0.018 0.036 lower probability in good articles 6 quotes 0.003 0.004 0.021 46 modals and adverbs 0.004 0.005 0.003 15 other (present tense verbs) 0.002 0.003 0.018 27 other (numbers) 0.022 0.025 0.003 39 other (proper names) 0.014 0.016 0.024 42 other (plural nouns with adjective modifiers) 0.039 0.043 0.003 table 16: clusters that have significantly different probabilities in good and typical articles. the mean value for the cluster label in the good and typical articles and the p-value from a two-sided t-test comparing these means are also indicated. four states occurred with higher probability in good writing. two of them are from the description type category that we identified. particularly, in cluster 5, as we already noted, several of the adjective phrases provided evaluative comments or opinion. such sentences are more prevalent in good articles compared to the typical ones. sentences that belong to cluster 34 also have a descriptive nature, however, only at noun phrase level. cluster 30 groups sentences which provide explanation in the form of ‘wh’ clauses. the prevalence of modal containing sentences (cluster 30) indicates that hypothesis and speculation statements are also frequently employed. most of these clusters appear to be associated with descriptions and explanations. on the other hand, there are six states that appeared with significantly greater probability in the typical articles. one of these clusters is state 6 depicting quotes. several of the other clusters come from the category that we grouped as ‘other’ class in table 15. these are clusters whose descriptive features are numbers, proper names, plural nouns and present tense verbs. even though the descriptive features for these clusters mostly indicated properties at noun phrase or verb level, we find that they provide signals that can distinguish good from typical writing. further annotation of the properties of such sentences could provide more insight into why these sentences are not preferred. since these still approximate categories are significantly differently distributed in very good and typical categories, we also examined the use of features derived from them for pairwise classification. for each article, we calculate the proportion of sentences labelled with a particular state label. similar to the experiments for general-specific content, the features for a pair of articles are computed as the difference in feature values of the two constituent articles. we evaluated these features using 10-fold cross validation on the test corpus we described in section 4. we obtained an accuracy of 58%. while low, the accuracy is significantly higher than baseline giving motivation that sentence types could be valuable indicators of quality. we plan to experiment with other ways of computing similarity and also create manually annotate articles for communicative goals which will enable us to develop supervised methods to do the intention classification. 113 louis and nenkova 8. closing discussion we have presented a first corpus of text quality ratings in the science journalism genre. we have shown that existing judgments of good articles from new york times can be combined with the new york times corpus to create different categories of writing quality. the corpus can be obtained from our website.4 there are several extensions to the corpus which we plan to carry out. all the three categories great, very good and typical involve writing samples from professional journalists working for the new york times. so these articles are overall nicely written and the distinctions that we make are aimed at discovering and characterizing truly great writing. this aspect of the corpus is advantageous because it allows researchers to focus on the upper end of the spectrum of writing quality, unlike work dealing with student essays and foreign language learners. however, it would also be useful to augment the corpus with further levels of inferior writing using the same topic matching method that we used for our ranking corpus. one level can be articles from other online magazines—average writing. in addition we can elicit essays from college students on similar topics to create a novice category. such expansion of the corpus will further enable us to identify how different aspects of writing change along the scale. we believe that the corpus will help researchers to focus on new and genre-specific measures of text quality. in this work, we have presented preliminary ideas about how annotations for new aspects of text quality can be carried out and showed that automatic metrics for these factors can also be built and are predictive of text quality. in future work, we plan to build computational models for other aspects that are unique to science writing and which can capture why certain articles are considered more clearly written, interesting and concise. these include identifying figurative language (birke and sarkar, 2006) and metaphors (fass, 1991; shutova et al., 2010), and work that aims to produce sentences that describe an image (kulkarni et al., 2011; li et al., 2011). we wish to test how much these genre-specific metrics would improve prediction of text quality in addition to regular readability type features. moreover, we can expect that the strength of these metrics would vary for novice versus great writing. acknowledgements we would like to thank the anonymous reviewers of our paper for their thoughtful comments and suggestions. this work is partially supported by a google research grant and nsf career 0953445 award. references e. agichtein, c. castillo, d. donato, a. gionis, and g. mishne. finding high-quality content in social media. in proceedings of wsdm, pages 183–194, 2008. g.j. alred, c.t. brusaw, and w.e. oliu. handbook of technical writing. st. martin’s press, new york, 2003. y. attali and j. burstein. automated essay scoring with e-rater v.2. the journal of technology, learning and assessment, 4(3), 2006. 4. http://www.cis.upenn.edu/˜nlp/corpora/scinewscorpus.html 114 a corpus of science journalism for analyzing writing quality r. barzilay and m. lapata. modeling local coherence: an entity-based approach. computational linguistics, 34(1):1–34, 2008. r. barzilay and l. lee. catching the drift: probabilistic content models, with applications to generation and summarization. in proceedings of naacl-hlt, pages 113–120, 2004. d. biber. dimensions of register variation: a cross-linguistic comparison. cambridge university press, 1995. d. biber and s. conrad. register, genre, and style. cambridge university press, 2009. j. birke and a. sarkar. a clustering approach for nearly unsupervised recognition of nonliteral language. in proceedings of eacl, 2006. d. blum, m. knudson, and r. m. henig, editors. a field guide for science writers: the official guide of the national association of science writers. oxford university press, new york, 2006. j. burstein, m. chodorow, and c. leacock. criterion online essay evaluation: an application for automated evaluation of student essays. in proceedings of the fifteenth annual conference on innovative applications of artificial intelligence, 2003. k. collins-thompson and j. callan. a language modeling approach to predicting reading difficulty. in proceedings of hlt-naacl, pages 193–200, 2004. k. collins-thompson, p. n. bennett, r. w. white, s. de la chica, and d. sontag. personalizing web search results by reading level. in proceedings of cikm, pages 403–412, 2011. e. dale and j. s. chall. a formula for predicting readability. edu. research bulletin, 27(1):11–28, 1948. r. de felice and s. g. pulman. a classifier-based approach to preposition and determiner error correction in l2 english. in proceedings of coling, pages 169–176, 2008. m. elsner, j. austerweil, and e. charniak. a unified local and global model for discourse coherence. in proceedings of naacl-hlt, pages 436–443, 2007. d. fass. met*: a method for discriminating metonymy and metaphor by computer. computational linguistics, 17:49–90, march 1991. l. feng, n. elhadad, and m. huenerfauth. cognitively motivated features for readability assessment. in proceedings of eacl, pages 229–237, 2009. r. flesch. a new readability yardstick. journal of applied psychology, 32:221 – 233, 1948. m. gamon, j. gao, c. brockett, a. klementiev, w. b. dolan, d. belenko, and l. vanderwende. using contextual speller techniques and language modeling for esl error correction. in proceedings of ijcnlp, 2008. b. grosz and c. sidner. attention, intentions, and the structure of discourse. computational linguistics, 3(12):175–204, 1986. 115 louis and nenkova r. gunning. the technique of clear writing. mcgraw-hill; fouth printing edition, 1952. y. guo, a. korhonen, and t. poibeau. a weakly-supervised approach to argumentative zoning of scientific documents. in proceedings of emnlp, pages 273–283, 2011. m. heilman, k. collins-thompson, j. callan, and m. eskenazi. combining lexical and grammatical features to improve readability measures for first and second language texts. in proceedings of hlt-naacl, pages 460–467, 2007. d. higgins, j. burstein, d. marcu, and c. gentile. evaluating multiple aspects of coherence in student essays. in proceedings of hlt-naacl, pages 185–192, 2004. n. karamanis, c. mellish, m. poesio, and j. oberlander. evaluating centering for information ordering using corpora. computational linguistics, 35(1):29–46, 2009. j. y. kim, k. collins-thompson, p. n. bennett, and s. t. dumais. characterizing web content, user interests, and search behavior by reading level and topic. in proceedings of wsdm, pages 213–222, 2012. d. klein and c.d. manning. accurate unlexicalized parsing. in proceedings of acl, pages 423– 430, 2003. g. kulkarni, v. premraj, s. dhar, siming li, yejin choi, a.c. berg, and t.l. berg. baby talk: understanding and generating simple image descriptions. in proceedings of cvpr, pages 1601 –1608, 2011. m. lapata. probabilistic text structuring: experiments with sentence ordering. in proceedings of acl, pages 545–552, 2003. s. li, g. kulkarni, t. l. berg, a. c. berg, and y. choi. composing simple image descriptions using web-scale n-grams. in proceedings of conll, pages 220–228, 2011. m. liakata, s. teufel, a. siddharthan, and c. batchelor. corpora for the conceptualisation and zoning of scientific papers. in proceedings of lrec, 2010. z. lin, h. ng, and m. kan. automatically evaluating text coherence using discourse relations. in proceedings of acl-hlt, pages 997–1006, 2011. a. louis and a. nenkova. automatic identification of general and specific sentences by leveraging discourse annotations. in proceedings of ijcnlp, pages 605–613, 2011a. a. louis and a. nenkova. text specificity and impact on quality of news summaries. in proceedings of the workshop on monolingual text-to-text generation, acl-hlt, pages 34–42, 2011b. a. louis and a. nenkova. a coherence model based on syntactic patterns. in proceedings of emnlp, pages 1157–1168, 2012a. a. louis and a. nenkova. automatically assessing machine summary content without a goldstandard. computational linguistics, 2012b. 116 a corpus of science journalism for analyzing writing quality n. mcintyre and m. lapata. learning to tell tales: a data-driven approach to story generation. in proceedings of acl-ijcnlp, pages 217–225, 2009. a. nenkova, j. chae, a. louis, and e. pitler. structural features for predicting the linguistic quality of text: applications to machine translation, automatic summarization and human-authored text. in emiel krahmer and mariët theune, editors, empirical methods in natural language generation, pages 222–241. springer-verlag, berlin, heidelberg, 2010. e. pitler and a. nenkova. revisiting readability: a unified framework for predicting text quality. in proceedings of emnlp, pages 186–195, 2008. r development core team. r: a language and environment for statistical computing. r foundation for statistical computing, 2011. e. sandhaus. the new york times annotated corpus. corpus number ldc2008t19, linguistic data consortium, philadelphia, 2008. s. schwarm and m. ostendorf. reading level assessment using support vector machines and statistical language models. in proceedings of acl, pages 523–530, 2005. e. shutova, l. sun, and a. korhonen. metaphor identification using verb and noun clustering. in proceedings of coling, pages 1002–1010, 2010. l. si and j. callan. a statistical model for scientific readability. in proceedings of cikm, pages 574–576, 2001. r. soricut and d. marcu. discourse generation using utility-trained coherence models. in proceedings of coling-acl, pages 803–810, 2006. s. h. stocking. the new york times reader: science and technology. cq press, washington dc, 2010. j. swales. genre analysis: english in academic and research settings, volume 11. cambridge university press, 1990. j. m. swales and c. feak. academic writing for graduate students: a course for non-native speakers of english. ann arbor: university of michigan press, 1994. j. tetreault, j. foster, and m. chodorow. using parse features for preposition selection and error detection. in proceedings of acl, pages 353–358, 2010. s. teufel and m. moens. what’s yours and what’s mine: determining intellectual attribution in scientific text. in proceedings of emnlp, pages 9–17, 2000. s. teufel, j. carletta, and m. moens. an annotation scheme for discourse-level argumentation in research articles. in proceedings of eacl, pages 110–117, 1999. 117 microsoft word evers-vermeul_hoek_scholman_temporality_8mei2017 dialogue & discourse 8(2) (2017) 1-20 doi: 10.5087/dad.2017.201 ©2017 jacqueline evers-vermeul, jet hoek and merel c.j. scholman this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). on temporality in discourse annotation: theoretical and practical considerations jacqueline evers-vermeul j.evers@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands jet hoek j.hoek@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands merel c.j. scholman m.c.j.scholman@coli.uni-saarland.de language science and technology, saarland university campus c7.4, 66123, saarbrücken, germany editor: maite taboada submitted 11/2016; accepted 03/2017; published online 05/2017 abstract temporal information is one of the prominent features that determine the coherence in a discourse. that is why we need an adequate way to deal with this type of information during discourse annotation. in this paper, we will argue that temporal order is a relational rather than a segmentspecific property, and that it is a cognitively plausible notion: temporal order is expressed in the system of linguistic markers and is relevant in both acquisition and language processing. this means that temporal relations meet the requirements set by the cognitive approach of coherence relations (ccr) to be considered coherence relations, and that ccr would need a way to distinguish temporal relations within its annotation system. we will present merits and drawbacks of different options of reaching this objective and argue in favor of adding temporal order as a new dimension to ccr. keywords: coherence relations, cognitive approach to coherence relations (ccr), discourse annotation, discourse connectives, language acquisition, language processing, temporality 1 introduction one of the prominent features that determine the coherence of a discourse is temporality. zwaan (1996: 1196) has noticed this for temporality at the sentence level in narrative types of discourse: “temporal information is ubiquitous to narratives. not only does every sentence in a narrative carry temporal information, temporal information can be expressed in every major word class.” temporality is also highly relevant at the discourse level, where the states and events reported in individual clauses are ordered relative to each other. for instance, in examples (1)-(3), the evers-vermeul, hoek and scholman 2 temporal ordering of the segments is the most distinctive feature of the relation between the first (s1) and the second clause (s2). (1) [i visited my grandmother]s1 before [i went shopping.]s2 (2) [i went shopping]s1 after [i visited my grandmother.]s2 (3) [i visited my grandmother.]s1 then [i went shopping.]s2 considering that temporal information is so important in a discourse, we need an adequate way to deal with this type of information during annotation efforts at the discourse level. however, there is currently no consensus in the way in which annotation frameworks deal with temporality at the discourse level. we see at least three different approaches in the classification of discourse relations. in the first approach, temporal relations are simply listed among the inventory of discourse relations (e.g., bunt & prasad, 2016; kehler, 2002). for example, according to segmented discourse representation theory (sdrt; asher & lascarides, 2003; reese et al., 2007), examples (1) and (3) constitute a narration relation, whereas (2) would be a precondition relation. in the second approach, different types of temporal relations together constitute a separate class of discourse relations (e.g., halliday & hasan, 1976; hobbs, 1983; hovy & maier, 1995; longacre, 1983; martin, 1992). for example, both the penn discourse treebank (pdtb; prasad et al., 2008) and the rhetorical structure theory discourse treebank (rst-dt; carlson et al., 2003) distinguish a class of temporals. in the pdtb, this class includes precedence relations such as (1) and (3), and succession relations such as (2). rst-dt would also classify these relations as belonging to the class of temporals, but would label (1) temporal-before, (2) temporal-after, and (3) sequence (see sanders et al., submitted for details on the exact mapping of different labels across annotation frameworks). by contrast, the third approach considers temporal order not as a relational, but as a segmentspecific feature expressed by the propositional content of clauses in a text (e.g., partee, 1984; hinrichs, 1986; nerbonne, 1986; webber, 1988). this view can be found in the cognitive approach to coherence relations (ccr; sanders, spooren & noordman, 1992, 1993), which does not have a separate category of temporal relations. sanders et al. (1992) do not claim that temporal order is not relevant to discourse; instead, they argue that temporal relations belong to the class of additive relations because “the properties distinguishing temporal relations from other additive relations concern the referential meaning of the individual segments” (1992: 28). the discrepancy between approaches to temporality raises the following research question: what should be the status of temporality in discourse annotation? should temporal order be considered a relational feature or a segment-specific feature? this question is also addressed in kehler (1994) and relates to the definition of a coherence relation. sanders et al. (1992) define a coherence relation as “an aspect of meaning of two or more discourse segments that cannot be described in terms of the meanings of the segments in isolation. in other words, it is because of this coherence relation that the meaning of the discourse segments is more than the sum of its parts” (sanders et al., 1992: 2). if temporal ordering is always explicitly encoded in the discourse segments of a text, it should not be considered a relational feature, since temporal ordering would not be part of the relational surplus. in that case, discourse relation inventories should not include temporal relations. however, in section 2 we will argue, similar to the first and second approach presented above, that information on the temporal ordering of segments in a discourse cannot be seen as merely segment-specific information, and that it should thus be considered a relational feature. in order for any feature to be considered a relational feature in ccr, the feature needs to be cognitively plausible; that is, it should affect language acquisition and discourse processing, and have unique linguistic markers. this means that in ccr, temporal order is only a relational feature if temporal relations have a special cognitive status that is distinct from other types of on temporality in discourse annotation 3 relations, such as additive and causal relations. in section 3, we will therefore review empirical evidence supporting the claim that temporal order is indeed cognitively plausible. if temporal order is indeed a distinct cognitively plausible notion, it would need to be implemented in ccr. because ccr works with a highly restricted set of annotation dimensions that are relevant for describing all types of coherence relations (sanders et al., 1992, 1993), it cannot do with a first and straightforward option: to simply acknowledge temporal relation labels in its inventory of end labels, an approach that would be similar to the second approach presented above. in section 4, we will therefore discuss the merits and drawbacks of several ways in which temporal relations can be included in ccr, and argue in favor of making temporal order an additional dimension. the paper ends with a discussion and conclusion (section 5). 2 temporality: propositional or relational feature? sanders, spooren and noordman (1992) are clear proponents of the position that the temporal meaning aspect of coherence relations is a propositional rather than a relational feature. they claim, much like partee (1984), hinrichs (1986), nerbonne (1986), and webber (1988), “that temporal relations belong to the class of additive relations and that the properties distinguishing temporal relations from other additive relations concern the referential meaning of the individual segments” (sanders et al., 1992: 28). according to them, “the temporal meaning aspect is to a large degree determined by the referential content of the segments, more than, for instance, the causal meaning aspect” (sanders et al., 1992: 28). there do indeed appear to be many segment-specific features that can be addressed under the umbrella term ‘temporality’. for example, adam’s (2008) overview of the many dimensions of temporality, in her terms timescapes, contains features such as tempo (speed, pace, rate of change), duration (extent, instantaneity) and temporal modality (past, present, future), all of which are expressed within segments. likewise, berman and slobin’s (1994: 19) definition of temporality includes two segment-specific properties: the expression of the location of events on the timeline (past, present and future), and the temporal constituency of events (perfective, imperfective, progressive, iterative; cf. also comrie, 1976). in this paper, we restrict our discussion of temporality to the third subtype of temporality listed in berman and slobin (1994): the temporal relations between events (anteriority, simultaneity and posteriority). this sequential order of the propositions expressed by discourse segments at least has the potential of being a relational and not just a propositional feature. in the remainder of this section, we will look at a variety of segment-specific temporality markers, and show that the presence of such markers is neither necessary nor sufficient for expressing the third subtype of temporality, temporal order. an example of a relation in which the sequential order of the segments is explicitly encoded in one of the segments is (4). five minutes later in s2 indicates that s2 took place after s1, and can therefore be considered a sequence marker, which places an event in relation to one or several other events mentioned in the discourse (bestgen & costermans, 1994; costermans & bestgen, 1991). (4) [mona received a phone call saying that she had lost her most important client.]s1 [five minutes later, she was fired.]s2 in (4), the temporal progression does not seem to be part of the relational surplus, but rather expressed by the propositional content of s2. note, however, that the presence of such a marker of temporal progression may not be sufficient for making the coherence relation temporal. it is for instance also very reasonable to infer a causal relation between the segments in (4); losing an important client is a plausible reason to be fired. (5) also contains lexical indicators of time: the prepositional phrases at ten o’clock in s1, and at twelve o’clock in s2. in terms of costermans and bestgen (1991; bestgen & costermans, evers-vermeul, hoek and scholman 4 1994), both these prepositional phrases serve as anchorage markers, elements that tie segments to an absolute time scale external to the narrative. the temporal progression appears, at first glance, to be entirely signaled by the propositional content of the segments and it seems as if there is no temporal meaning surplus at the relational level. (5) [at ten o’clock we had coffee.]s1 [at twelve o’clock we had lunch.]s2 however, even though the absolute time points are given within the segments, the temporal progression between the segments is actually inferred by establishing that the time points are two hours apart. the temporal ordering is therefore due to the fact that the two segments are related to each other; in other words, it is part of the relational surplus. note also that even in the presence of two anchorage markers a relation need not be temporal. the relation in (5) could also be judged as predominantly comparative or even contrastive. in addition to lexical items expressing time points, grammatical tense can provide information on the ordering of events mentioned in the discourse (cf. kehler, 1994). for example, the pluperfect in (6) indicates that the event in s2 precedes the event in s1 (example taken from lascarides & asher, 1993: 472). (6) [max slipped.]s1 [he had spilt a bucket of water.]s2 again, it may seem as if there is no temporal meaning surplus at the relational level. however, the pluperfect in itself does not signal anti-chronological order. the reversed temporal order is inferable because of the combination of the two tenses; because the event in s2 is in pluperfect tense, it has to be interpreted as having occurred before the event in perfect tense in s1. similarly, the future tense of s1 in (7) has to be interpreted as occurring after the perfect tense of s2, but this is again due to the combination of tenses between the two clauses. had s2 in (7) also been in future tense, as in (8), the temporal ordering would not have been anti-chronological, but rather chronological or simultaneous, depending on the context. (7) [jane will slip.]s1 [max has spilt a bucket of water.]s2 (8) [jane will slip.]s1 [max will help her get up.]s2 the sequential interpretation of segments also differs depending on whether the segments denote states or events. this distinction includes, but is not restricted to grammatical aspect, as verb semantics can also determine whether something is a state or an event. event-denoting clauses in past tense, such as s2 in (9), tend to be temporally interpreted in sequence, that is, they either precede or follow some other event in discourse. in contrast, state-denoting clauses like s2 in (10) tend to temporally overlap with previous discourse events (gennari, 2004: 878, example (10) taken from webber, 1987). (9) [i went to a party last night.]s1 [i ran into an old friend outside.]s2 (10) [i went to a party last night.]s1 [the music was wonderful.]s2 if the event-state distinction helps distinguish between temporal sequence and temporal overlap, it may seem as if there is no temporal meaning surplus at the relational level. again, however, it is the combination of the two clauses and their status as states or events that eventually determines the inferred temporal ordering. this entails that the relevance of the event-state distinction for temporality is located in the relational surplus. tense and aspect markers, as well as anchorage markers that signal an absolute time point, and sequential markers that signal the relative time between the segments help us establish the on temporality in discourse annotation 5 temporal order of events in a discourse. as mentioned above, the presence of temporal information within the segments is not sufficient to establish a temporal coherence relation. it is not the mere presence of these signals, but rather the way in which these signals are combined that influences the temporal ordering of adjacent clauses. in addition, the presence of temporal information within the segments is not a necessary characteristic of temporal relations. the segments of the relations in (11)-(13), for instance, have a highly similar structure, but the specific type of temporal relation differs between the sentences. (11) [paul read the paper]s1 before [going to work.]s2 (12) [zoe went to the gym]s1 after [finishing her exam.]s2 (13) [arnold organized his desk]s1 while [talking to his mother on the phone.]s2 examples (11)-(13) illustrate that temporal ordering may not always become apparent from the segments alone. the segments in (11)-(13) are structurally identical, with a simple past in s1 and a gerund in s2. still, the order of the events in (11) is opposite from the order of the events in (12), and in (13), the two segments overlap in time. these orders are indicated by the connectives before, after and while, respectively, linguistic elements that are considered not to be part of the propositional content of a text and that are therefore not part of the discourse segments (cf. blakemore, 2002; hall, 2007). given the absence of anchorage markers or other indicators of time, we cannot represent the differences between (11), (12), and (13) without attributing a temporal meaning relation between the two clauses of each fragment. it should be noted that for other types of coherence relations, the presence of propositional elements hinting at specific relation types also does not necessarily eliminate the option that what they signal is included in the relational surplus of a coherence relation. for example, the presence of negative elements in s1 or of semantic opposites in both segments hint at a contrastive relation (asr & demberg, 2015), as in (14). similarly, modal elements such as obviously in (15) hint at a subjective relation (scheibman, 2002; traugott, 2010). as with temporal relations, the presence of these propositional elements, however, is neither a required nor sufficient property of the relation at hand. for example, (15) remains a subjective argument-claim relation, even in the absence of obviously. (14) [harry never lifted a finger,]s1 but [kyle worked really hard.]s2 (15) [jenny obviously went to the party,]s1 since [she came home at 3 a.m.]s2 in conclusion, temporal ordering of segments in a discourse cannot be seen as merely segment-specific information. it should thus be seen as a relational feature, even if linguistic indicators of temporality are present within a clause. in other words: temporal relations constitute a specific subset of coherence relations, just like, for instance, causal, conditional and contrastive relations. 3 cognitive plausibility of temporality in ccr, coherence relations are considered to be cognitive entities because they account for the coherence in the reader’s cognitive text representation (sanders et al., 1992). working from this assumption, it is expected that the nature of a relation affects discourse processing. in other words, different types of relations, such as causal and additive relations, should require different processing costs and result in different mental representations (sanders & noordman, 2000). this means that if temporal order is to be considered a relational feature rather than a propositional one, the defining features that separate temporal relations from other types of relations should be cognitively plausible. in this section, we will examine three types of evidence to determine whether temporal relations are distinct cognitive entities: a) evidence from the system of evers-vermeul, hoek and scholman 6 linguistic markers; b) evidence from language acquisition; and c) evidence from language processing. 3.1 linguistic markers of temporality connectives and other relational cue phrases signal certain types of relations; for example, because is used to express causal relations and nevertheless is used to express concessive relations. readers can use these connectives as so-called processing instructions that signal how the incoming information should be related to the previous discourse segment (britton, 1994; canestrelli, mak & sanders, 2013). the category of connectives tends to consist of simple linguistic expressions and idiom chunks (knott & dale, 1994: 45), and is often restricted to grammaticalized linguistic elements that express a two-place semantic relation, have propositional arguments, are not integrated in the predicative structure, and do not vary regarding inflection (see among others mendes & lejeune, 2016; stede, 2002; but for a more liberal point of view see prasad, joshi & webber, 2010; rysová & rysová, 2015). knott and dale (1994) have used the existence of such connectives and relational cue phrases to motivate a set of coherence relations. they work from the assumption that “[if] people actually use a particular set of relations when constructing and interpreting text, it is likely that the language they speak contains the resources to signal those particular relations explicitly” (knott & dale, 1994: 44). similarly, stukker and sanders (2012) work from what they call the linguistic categorization hypothesis, adopting a basic assumption of cognitive linguistics that a direct link holds between linguistic categories and cognitive categories, and applying it in a cross-linguistic corpus-based study on different types of causal relations. they assume that “when language users systematically prefer one highly grammaticalized linguistic item over another (even highly similar) one to express a certain type of causal relationship, they assign the relation expressed to a specific conceptual type of causality” (stukker &sanders, 2012: 170). such an approach thus assumes that “linguistic choices provide a window on speakers’ cognitive categorizations” (stukker & sanders, 2012: 170). this assumption can be taken to hold for temporal order as well: if there are frequently used, grammaticalized linguistic markers that distinguish different types of temporal relations, one can assume that temporal order is indeed a relational feature. in order to determine whether a specific feature, such as temporality, constitutes a type of coherence relation, knott and dale (1994) look for relational cue phrases that mark this feature. using a test for identifying relational cue phrases they gather a corpus of around 200 cue phrases for all kinds of relations. this corpus provides ample evidence for connectives and cue phrases in the temporal domain, such as before, after, as, while, when, and as soon as (knott and dale, 1994: 60). assuming that features need to be cognitively plausible in order to be considered a type of relations, knott and dale’s (1994) study provides evidence that temporality indeed constitutes a separate class of relations. at first glance, two frequent phenomena in the use of linguistic markers seem to provide evidence to counter this conclusion: 1) underspecification, the fact that connectives and relational cue phrases may feature in a more specific type of relation that they do not encode themselves; and 2) polysemy, the fact that connectives may encode different types of coherence relations. we will discuss these phenomena in turn. consider the underspecified relations in (16) and (17) (the latter taken from mak & sanders, 2013: 1415): (16) mary burst into tears after her boss fired her. (17) the boy quarreled with his father when he did not get permission to go out. in (16) and (17), “the semantics of the connective that is used to indicate the link does not fully match the semantics of the relation” (spooren, 1997: 150). the two segments present events that on temporality in discourse annotation 7 do not only have a temporal order, but plausibly also a causal relation: mary likely cried because she got fired, and the boy likely quarreled with his father because he did not get permission to go out. as these paraphrases illustrate, underspecified relations can be recognized by exchanging the underspecified connective for a more specific counterpart (cf. also givón, 1990: 828). in an eyetracking experiment, mak and sanders (2013) have shown that, triggered by the implicit causality verb quarreled in s1, readers indeed arrived at a causal interpretation of (17) and similar sentences, even in the presence of the temporal marker when. this resulted in faster reading times throughout s2, compared to sentences without a causal interpretation. one could argue that because the temporal connectives in (16) and (17) can be used in other types of relations, they cannot be taken as evidence that temporal order is a cognitively plausible feature. however, it should be noted that non-temporal connectives can feature in underspecified relations as well: for example, and can be used to mark additive, contrastive, temporal, and causal relations, and but can be used to mark both contrastive (or negative additive) and concessive (or negative causal) relations. language users arrive at a richer interpretation than the connective actually encodes due to a highly common process referred to as pragmatic enrichment or conversational implicature (grice, 1989; levinson, 2000). still, the fact remains that connectives such as after and when do occur in relations in which a causal implicature is less likely or not possible at all. the fact that these connectives actually encode temporal ordering, and therefore provide temporal processing instructions to the reader, is still in line with the original linguistic categorization hypothesis. another frequent linguistic phenomenon that seems to serve as a counter-argument is polysemy; certain connectives encode not only temporality, but also some other type of relation. for example, in (18), the relation is temporal as well as contrastive: the bad performance of the soccer team is contrasted with the relatively good performance four years earlier. in (19), the contrastive interpretation of while is even the only one available in a context where mary’s father has deceased. (18) the dutch soccer team failed to qualify for the 2016 european cup, while two years before the team had reached the semi-finals of the world cup. (19) mary likes oysters, while her father bill hated them. again, pragmatic enrichment is involved, but this time in the diachronic development of the connective at hand. traugott (1995) illustrates how elements from the pragmatic context in which an expression typically occurs can become part of a word’s conventional meaning, a process referred to as “the semanticization of pragmatics” (see barlow & kemmer, 2000; traugott, 2011). middle english while ‘during’, which was purely temporal, has developed into modern english while ‘although’, because the pragmatic inference of “surprise concerning the overlap in time or the relations between event and ground” eventually became conventionalized in the connective (traugott, 1995: 41; cf. also gonzález-cruz, 2007).1 the question remains whether the existence of such polysemous connectives should be seen as counter-evidence to the linguistic categorization hypothesis. given that ambiguity proliferates in language, and polysemy is not restricted to temporal connectives, we are inclined to answer this question with a no. again, examples of temporal while that are clearly non-ambiguous can be found, and enough non-ambiguous temporal markers remain that justify recognizing temporality as a cognitively plausible feature. 1 note that this process is by no means restricted to english. for example, dutch terwijl ‘while’ has, similar to while in english, taken on a contrastive meaning (vogl, 2007), german weil ‘because’ also used to mark that events happened at the same time, but invited a causal inference following the pattern post hoc ergo propter hoc (behaghel, 1928: 341; traugott & könig, 1991: 194), and has since been semanticized, into a causal connective (keller, 1995). evers-vermeul, hoek and scholman 8 in sum, given that there are specific temporal connectives that can mark temporal relations, we argue that temporal order is indeed a cognitively plausible and a relational feature. in the next sections, we will see that temporal connectives and relations are acquired and processed differently from other types of connectives. 3.2 evidence from language acquisition for temporality to be considered a cognitively plausible notion, it should play a role in the cognitive processes underlying language use. in this section, we will therefore look at evidence from language acquisition data, and in the next section, we will address processing data from adult language users. we discuss a non-exhaustive list of studies that have shown that temporality influences language use. language acquisition studies have shown that the first emergence of temporal connectives happens at a different stage than that of additives or causals: around their second birthday, children start combining clauses with and. next, they start using (and) then, followed by because a few months later (see bloom et al., 1980 for english, and evers-vermeul & sanders, 2009 for a similar developmental sequence in dutch). the same developmental picture arises from a growth curve analysis of german connective acquisition (van veen, evers-vermeul, sanders & van den bergh, 2009; 2014). the fact that temporal connectives are acquired in a stage separate from both additives and causals indicates that temporal relations are a category of coherence relations that is distinct from additive and causal relations, but also that temporality is a cognitively plausible notion. even though children start using temporal connectives around the age of three, they are not able to fully comprehend them until much later (bever, 1970; blything, davies & cain, 2015; clark, 1971; french & brown, 1977, johnson, 1975; keller-cohen, 1987; pyykkönen, niemi & järvikivi, 2003; pyykkönen & järvikivi, 2012; trosborg, 1982). in other words, children produce temporal connectives before they are able to fully comprehend them. as comprehension of temporal connectives develops slowly, there is a clearly identifiable developmental path in the process (clark, 1971). initially, at the age of three, children are not able to correctly interpret the temporal order information that before and after provide, using an order-of-mention strategy instead: whatever is mentioned first, is interpreted as the first event. in the next stage, around the age of four, children interpret sentences with before correctly, but not yet sentences with after. finally, around the age of five, children are able to comprehend both temporal connectives correctly around 80% of the time. a comprehension score of 80% is still lower than that of adults, who do not have any trouble correctly interpreting temporal order in relations (see section 3.3). research on when exactly children reach adult competence in interpreting these connectives does not provide a conclusive answer: for example, blything, davies and cain (2015) find in a forced choice reading experiment that children are performing at high levels of accuracy around age seven, while cain and nash (2011) report that ten-year-olds still differed from adult levels of performance on a cloze task, sentence rating task and online reading task. pyykkönen and järvikivi (2012) even found that children experience difficulties achieving the correct interpretation of temporal relations containing before and after until the age of twelve. regardless of the exact age at which children are able to successfully comprehend temporal relations, it is clear that children experience difficulty in interpreting these relations. blything et al. (2015) attribute this difficulty to children’s working memory. in a forced choice reading experiment, they show that children with higher memory capacity comprehend different types of temporal relations more accurately. on the basis of the acquisition data discussed in this section, we can conclude that temporality is relevant for children’s language production and their language comprehension from a very early age. production data indicate that temporality determines part of the order of acquisition of connectives, and studies investigating children’s comprehension of temporal relations indicate that their interpretation process is facilitated by a chronological order of events. on temporality in discourse annotation 9 3.3 evidence from language processing if temporality is a cognitively plausible notion, it should play a role in language processing in adults as well. indeed, it has been found that readers do encode temporal information in their representations (mandler, 1986; gennari, 2004; townsend, 1983; van der meer, beyer, heinze & badel, 2002). zwaan (1996), for example, investigated the effect of discourse time shifts such as an hour later in the context of a sequence of events (for example, john was beaming. a moment later…). sentential reading times were longer after a large time shift such as an hour later compared to a smaller time shift such as a moment later. these results indicate that readers track the time of events while reading a text, and that the duration of the time shifts also influences the reading times. other studies have shown that the order of events also influences readers’ processing; it has consistently been found that chronologically ordered temporal relations are processed faster than anti-chronologically ordered sentences. in this section, an overview of these studies will be given. unlike children, adults are able to achieve the correct interpretation regardless of the event order in relations. however, chronologically ordered sentences do facilitate processing. offline studies have shown that temporal relations with a chronological order are remembered better. for example, clark and clark (1968) investigated the recall of complex sentences marked by before, and then, after or but first. they found that the subjects remembered the sentences with a chronological order better than those with an anti-chronological order. similarly, townsend (1983) investigated the recall of temporal and causal english sentences connected by after, before, when, since, while or because and found that, for both types of relations, the subjects remembered the sentences with a chronological order better than those with an anti-chronological order. baker (1978) found that people were faster and more accurate in reporting the order of events in narratives when the events were presented in chronological order compared to antichronological order. these studies therefore show that the temporal order of clauses affects how well the relation is encoded in the mental representation. münte, schiltz and kutas (1998) provided evidence that the encoding of temporal order by adults requires working memory as well. their stimuli consisted of chronological and nonchronological sentences starting with either before or after, such as before/after the psychologist submitted the article, the journal changed its policy. using event-related potentials (erp), they showed that the participants had different brain responses to the connectives before and after within 300 ms after presentation of the connective. the authors interpreted these results as indicating that anti-chronologically ordered sentences are more demanding of working memory because they require additional discourse-level computations. these findings were confirmed by ye et al. (2012) in an fmri study. using stimuli similar to the ones used by münte et al. (1998), they found that anti-chronological sentences need additional discourse-level computations to reverse the order of the two constituent clauses. readers therefore appear to invest cognitive energy in correctly representing the temporal order of relations. interestingly, however, it seems that temporal relations differ from causal coherence relations in this respect. mandler (1986) found that a chronological order of complex sentences facilitates the processing of temporal relations, but not of causal relations. in other words, chronological ordering of the relation is not necessary for rapid understanding of the temporal relation between two causally connected events. mandler attributes this difference to the readers’ world knowledge: when events are causally connected, prior knowledge about the temporal relations between a cause and its effect makes encoding the order of two events relatively effortless, regardless of the order of mention of the events. in the case of events that are arbitrarily ordered in time, more effort is required to encode the temporal order, since no pre-existing knowledge about this order is available (cf. french & brown, 1977; trosborg, 1982). hence, temporal relations seem to be processed differently from causal relations. taken together, the results of several offline and online experiments indicate that the temporal order of coherence relations affects the processing of these relations in adults. although evers-vermeul, hoek and scholman 10 adults do not have the same difficulty as children with achieving the correct interpretation of temporal relations, they do process these relations faster and recall them better when the segments are ordered chronologically. an anti-chronological order of the segments requires additional computations, which is more demanding of working memory. this indicates that temporality is a cognitively relevant notion that people take into account when using language. 4 how to implement temporality in ccr in sections 2 and 3 we have shown that information on the temporal ordering of segments is relational information and that temporality is a cognitively plausible notion; temporal order is expressed in the system of linguistic markers, and is relevant in acquisition as well as language processing. this entails that temporality should be implemented in ccr. for many other approaches, it would be fairly easy to extend the respective inventory of relation labels by adding more labels. however, ccr does not work with an inventory of relation labels; rather, it distinguishes a restricted set of four dimensions that are relevant for describing all types of relations (sanders et al., 1992, 1993). as a result, including temporality in the framework is not as straightforward as it might be in other discourse annotation frameworks. in the remainder of this section we will examine two options of including temporality in ccr: to extend one of the existing ccr dimensions in order to incorporate temporality (section 4.1), and adding a new dimension to ccr (section 4.2). first we will briefly introduce the relevant ccr dimensions. 4.1 extending a current ccr dimension ccr differs from other annotation frameworks in that coherence relations are not assigned a specific end label, but rather are defined by their characteristics. in ccr, four cognitive dimensions are distinguished that apply to every relation. these dimensions are polarity, source of coherence, basic operation, and order of the segments (sanders et al., 1992). here we explore whether one or both of the latter two dimensions can be adapted in order to include temporal relations in ccr. basic operation distinguishes between causal and additive relations. a relation is causal if an implication relation (p → q) can be deduced between the two segments. causal relations are typically connected by because or so. a relation is additive if the relation between the segments is one of logical conjunction (p & q). additive relations are typically connected by and. in the original proposal, ccr considers temporal relations to be a subtype of additive relations. one way of giving temporality a separate status in ccr would be to add the value ‘temporal’ to the basic operation dimension, thereby making a three-way distinction between additive, temporal, and causal relations. this option has actually already been put into practice by, for example, louwerse (2001) and scholman, evers-vermeul and sanders (2016). however, this option ignores the fact that other ccr dimensions are binary, distinguishing, for instance, positive vs. negative relations, and basic vs. non-basic relations.2,3 2 note, however, that the principle of having binary dimensions has been ‘violated’ before. on the dimension source of coherence, sometimes two values are distinguished: semantic vs. pragmatic (sanders et al., 1992), or objective vs. subjective (scholman et al., 2016), but it is also frequently operationalized as having three values, following sweetser’s (1990) approach: content, epistemic and speech act (see, among others, evers-vermeul, 2005; stukker, 2005). this trichotomy, however, is really a simplification of two binary distinctions: objective (content) vs. subjective, which contains both epistemic and speech act relations. this allows for different kinds of generalizations: first, between content on the one hand and epistemic and speech act on the other, and second between epistemic and speech act. 3 in line with other cognitive linguists, we acknowledge the fact that in practice categories often display fuzzy boundaries, with less and more prototypical instances (cf. rosch, 1977). this has also been explored for coherence relations and connectives (stukker & sanders, 2012; sanders & spooren, 2013). this does not, however, mean that the definitions of the categories themselves should be fuzzy. given the objectives of the current paper, we will not address the fundamental issue of whether we should abandon binary distinctions in favor of the prototypes typically used in cognitive linguistics, but instead base our line of reasoning on the binary categorizations that have shown to be relevant in accounting for patterns in both language acquisition and language use. on temporality in discourse annotation 11 more importantly, it ignores the fact that many causal relations display a temporal order as well; by default, causes in the real world precede their consequence. by placing temporal order on a par with causality, the two values become mutually exclusive. note, however, that all causal relations are by default additive as well, given that an implication relation (p  q) presupposes p and q (p & q). in other words, this overlap is inevitable. an important difference between causal vs. additive relations and causal vs. temporal relations is that not all causal relations are temporal, and that during the annotation process one would therefore want to be able to attribute these values simultaneously in order to be able to distinguish between temporal causal relations and non-temporal causal relations. we will further explore this point in section 4.2. a second option would be to adapt the ccr dimension order of the segments, as is done in scholman et al. (2016). in the original ccr proposal (sanders et al., 1992), order of the segments is defined in terms of implication relations, and therefore only applies to causal and conditional relations. it refers to the mapping of the antecedent (p) and the consequent (q) onto the segments, where p is the cause, condition or argument, and q is the consequence or the claim. in a coherence relation with a basic order, such as (20), p is s1, followed by q as s2. in a relation with a non-basic order, such as (21), p maps onto s2 and q onto s1. (20) [it was raining.]s1 that is why [the streets are wet.]s2 (21) [the streets are wet]s1 because [it was raining.]s2 in the original proposal, the order of the segments dimension does not apply to additive relations, as these are symmetrical by definition. however, it could be argued that segments with a sequential temporal order form an exception in this respect because the events expressed in p and q are necessarily ordered in time. following this line of reasoning, it could be said that if s1 and s2 follow this order in time, the relations show a basic order. in a relation with a non-basic order, s2 expresses p; that is, the event expressed in s2 precedes the event in s1 in reality. there are some drawbacks to depicting the chronological ordering of temporal relations using order of the segments. first and foremost, it would require seriously stretching or altering the definition of order of the segments. temporal relations, after all, do not involve an implication relation and do not consist of an antecedent p and a consequent q. second, for causal and conditional relations, this solution does not leave the option of distinguishing between the order of the segments within the implication relation and the chronological order of the segments. as we will elaborate on in section 4.2, these two types of order may, but do not necessarily overlap. third, if the values on the basic operation dimension are still restricted to additive and causal, as in the original ccr proposal, this solution implies that in principle, order of the segments becomes available for the entire class of additives. in this scenario, temporal relations can be distinguished from other additive relations by looking at their value for order, and additive relations that are not specified for order would be purely additive relations. however, ccr would then have no way of setting apart purely additive relations from temporal relations expressing temporal overlap. since both additive and synchronous temporal relations are symmetrical in the sense that neither relation involves one segment occurring before the other, this would mean that neither type would be specified for order. considering the initial motivations of sanders et al.’s (1992) taxonomy of coherence relations and considering the drawbacks mentioned above, using order of the segments to capture the chronological ordering of temporal relations seems undesirable. 4.2 adding a new dimension to ccr considering the drawbacks of using one of the existing dimensions to incorporate temporal relations in ccr, as outlined in section 4.1, the only option for achieving this goal seems to be adding one or more new dimensions to ccr’s inventory of annotation dimensions. adding one evers-vermeul, hoek and scholman 12 extra dimension would be in line with evers-vermeul and sanders (2009), who distinguish between [+ temporal] and [α temporal] relations in order to be able to account for the order of first emergence of the english connectives and, then, and because, and their dutch counterparts. adding more than one temporality dimension would be in line with clark (1971: 273), who presents three hierarchically ordered features to account for the acquisition of the temporal connectives after, before and when: [+/– time], [+/– simultaneous], and [+/– prior]. according to clark, children’s production data reflect the learning of these features from the most superordinate feature [+ time] down, which reflects the cognitive relevance of all three features. below, we will take this latter approach, and propose to implement the three-step temporality dimension presented in (22) that is orthogonal to the other ccr dimensions. this dimension will help account for temporal relations within ccr while adhering to its initial basic principle of using binary variables. note that this approach of working with a dimension with multiple, hierarchically ordered steps can already be found in ccr itself, even though it is not presented as such. first, ccr applies the basic operation dimension to set implication relations apart from additive relations. second, within the group of implication relations, causal relations are distinguished from conditional relations. third, sanders et al. (1992: 12) propose that volitionality would be a candidate for further specification of the proposed taxonomy, which would then differentiate volitional and non-volitional causal relations. (22) the three-step temporality dimension 1 temporal non-temporal 2 sequential synchronous 3 chronological anti-chronological as we have mentioned before, ccr considers temporal relations to be a subset of additive relations; that is, temporal relations are considered to be additive relations that are ordered in time. the first step in the temporality dimension we propose helps distinguish the subset of temporal relations from the rest of the additive relations. relations such as (23), (24) and (25) are temporal, while additive relations such as (26) are non-temporal.4 (23) after [mark came home from the supermarket,]s1 [he realized his grocery list continued on the back of the paper.]s2 (24) before [anne went to the important meeting,]s1 [she did a lot of preparation.]s2 (25) [mirjam played the piano,]s1 while [nathan was playing soccer.]s2 (26) [michael works at the veterinary clinic across the street.]s1 [he also volunteers at the local animal shelter.]s2 to the additive relation in (26), chronological ordering is irrelevant. although in terms of truthvalue s1 and s2 hold simultaneously, these segments are not presented in combination to indicate that michael performs these activities simultaneously in the real world. note that temporal only means that temporality is relevant to the relation at hand, not necessarily that the segments display a sequence of events. hence, the synchronous temporal relation in (25) and the sequential relations in (23) and (24) are all temporal. 4 we leave it to another paper to discuss whether values distinguished within a dimension should simply receive different value labels, or whether these would have to represented with actual features such as [+/– time] or [+/α time]. in line with the labeling within other ccr dimensions, we take the first approach, and use different value labels. on temporality in discourse annotation 13 the second step in the temporality dimension distinguishes between sequential and synchronous temporal relations. since in (23) and (24) one of the segments occurs before the other, these temporal relations are sequential. by contrast, the non-sequential relation in (25), in which the segments display temporal overlap, is labeled synchronous. finally, the chronological order of sequential temporal relations has to be determined. this distinction is captured by the third step in the temporality dimension, which differentiates between chronological and anti-chronological relations. since the event in s1 occurs before the event in s2, (23) has a chronological order. in (24), the order of the segments does not match the order of the events in the real world. this relation is therefore anti-chronological. because the temporality dimension operates orthogonal to the other ccr dimensions, the three steps do not only apply to additive relations. they are also relevant to objective causal relations, such as (27) and (28), which depict an implication relation that can be observed in the real world. temporal ordering is involved in both cases, which is why they are both labeled temporal. in addition, both relations are sequential, as the events follow each other in time. finally, (27) is chronological because the freezing mentioned in s1 occurs first and causes the formation of the ice flowers mentioned in s2. by contrast, (28) is anti-chronological because the freezing is mentioned in s2. (27) [it had been freezing.]s1 that is why [there were ice flowers on the windows.]s2 (28) [there were ice flowers on the windows]s1 because [it had been freezing.]s2 in these examples, the values on the third step in the temporality dimension align with the values on the dimension order of the segments: (27) is chronological and displays a basic order, since s1 expresses the antecedent p and s2 the consequent q, while (28) is anti-chronological and has a non-basic order. chronological ordering and order of the segments will align in this way, basic/chronological and non-basic/anti-chronological, in all objective causal relations, as a cause tends to precede its effect. however, chronological ordering can, and often will, differ from the order of the segments in case of subjective causal relations, which includes both epistemic and speech act relations (sanders & spooren, 2009). consider for instance the relations in (29) and (30). (29) [the neighbors’ lights are on,]s1 so [they must be home.]s2 (30) [the neighbors just got home,]s1 so [the lights will probably be turned on soon.]s2 both (29) and (30) have a basic order of the segments: p is expressed in s1, while s2 expresses q, which in both cases is a judgment rather than a fact. the temporal order of events in the real world, however, is not identical. (29) is ordered anti-chronologically, since turning on the lights, depicted in s1, happens after the neighbors’ coming home, which is depicted in s2. (30), on the other hand, is ordered chronologically: the neighbors’ arrival in s1 precedes the probable turning on the lights in s2. by using the third step in the temporality dimension in addition to order of the segments, we can capture the difference between relations in which an argument is presented to back up a claim or conclusion, as in (29), and relations in which a prediction is made on the basis of an observation, as in (30). distinguishing between chronological order and order of the segments helps analyze not only subjective causal relations, but also relations such as (31), also known as implicit assertion relations in the pdtb (pdtb research group, 2008). (31) if [you want to become rich someday,]s1 [you should probably get off the couch.]s2 (31) is a subjective conditional relation that has a basic order of the segments: the antecedent p is expressed in s1, and the consequent q in s2. however, in real time, the getting off the couch in s2 evers-vermeul, hoek and scholman 14 would have to occur before the getting rich in s1. (31) therefore has an anti-chronological temporal order. other examples of relations that could be annotated along these lines would be rst-dt’s enablement and problem-solution. in sum, we have proposed to incorporate temporality into ccr by implementing a temporality dimension with three hierarchically ordered binary annotation steps that are orthogonal to other ccr dimensions such as basic operation and order of the segments. 5 discussion and conclusion in this paper, we have argued that information on the temporal ordering of segments in a discourse cannot be seen as merely segment-specific information, and that temporal order should thus be considered a relational feature, even in the presence of linguistic indicators of temporality within the connected clauses. the cognitive approach to coherence relations considers coherence relations to be cognitive entities. this means that if temporal relations are to be included in ccr, the defining features that separate temporal relations from other types of relations should be cognitively plausible. a review of empirical evidence from the system of linguistic markers, from language acquisition, and from language processing supports the claim that temporal features are indeed cognitively plausible. in section 4, we have therefore presented a proposal on how to implement temporality in ccr. we suggest adding a three-step temporality dimension to ccr with a hierarchically ordered set of binary annotation steps that are orthogonal to the other dimensions, including basic operation and order of the segments. the first step in the temporality dimension sets temporal relations apart from non-temporal relations, in which temporal order is not relevant. the second step in the temporality dimension distinguishes between sequential and synchronous relations, and the third step differentiates between chronological and anti-chronological relations. the proposed approach allows for generalizing over different subtypes of temporal relations. for example, relations that are classified as additive on basic operation and temporal on the first temporality dimension, constitute the class of temporals distinguished by for instance pdtb and rst-dt (see section 1). in addition, the new temporality dimensions allow for a different type of generalization, one that none of the current discourse annotation frameworks are able to make. on top of setting apart relations for which temporality is the key defining feature, it can cluster all relations for which temporality is a relevant feature. this allows us to show the similarities between sequential temporal relations and objective causal relations, and between synchronous temporal relations and contrastive relations. it is exactly this type of clustering we need in order to be able to account for frequent linguistic phenomena such as underspecification and pragmatic enrichment (see section 3.1). in addition, the temporal dimension can help distinguish between different types of subjective causal relations and make explicit that in subjective causal relations temporal order may not always align with order of the segments (see section 4.2). we have drawn a clear distinction between temporality at the sentence level (the location of events on a time line as indicated by past, present or future tense, and the temporal constituency of events as expressed by markers of aspect) and temporality at the discourse level, focusing on the relational surplus of temporal order. this approach differs from the one taken by researchers who seem to collapse these levels and, for example, label ‘discourse verbs’ (danlos, 2006; lejeune, mendes & martins, 2016) and prepositional phrases taking nominal arguments (atallah, bras & vieu, 2015; mendes & lejeune, 2016; rysová & rysová, 2015; stede, 2002) as markers of coherence relations. although we acknowledge the merit of studying such alternative lexicalizations (cf. prasad et al., 2010), we prefer not to categorize alternative lexicalizations as markers of coherence relations, that is, as markers of the relational surplus of coherence relations. keeping the discourse and the sentence levels apart allows us to investigate how markers of on temporality in discourse annotation 15 temporality – or, for instance, causality – at these two levels interact. this may, for example, generate insight into why certain combinations of sentence and discourse markers are possible, preferred or impossible. for example, knott and dale (1994: 50) observe that earlier, afterwards, and later, as well as before and after can all be modified by any expression denoting a length of time, for instance, three days after, a minute earlier, and some time before. by contrast, the combination of certain temporal connectives and ‘discourse verbs’ denoting progression of time is impossible, as the comparison of (32), (33) and (34) illustrates (example (34) taken from danlos, 2006: 65). (32) ted left. next, sue arrived. (33) ted left. this preceded sue’s arrival. (34) #ted left. next, this preceded sue’s arrival. from a cognitive perspective, studying similarities and differences between marking at the sentence and the discourse level may give us insight into the underlying cognitive categories that people make use of. for instance, stukker, sanders and verhagen (2008) investigate causality at the inter-clausal level (expressed by causal connectives) and at the intra-clausal level (expressed by causal verbs), which allows them to show commonalities between the two. from a developmental perspective, linguists can and already do study the way in which markers that have a function at the sentence level gradually obtain a function at the discourse level. a famous example is the development of while discussed by traugott (1995), but more generally, one could look at differences between so-called primary and secondary, less grammaticalized connectives (rysová & rysová, 2015). this would be in line with knott and dale’s (1994:49-50) observation that relational markers are sometimes single elements and idiom chunks (e.g. because, on the other hand), but can also be partly compositional (five minutes later, three years later). also, researchers can account for the way in which children over the years acquire the various grammatical, lexical and discourse devices for expressing temporality (berman & slobin, 1994; uccelli, 2009). in this paper, we have focused on the question whether temporality should be considered a relational or a segment-specific feature. the same question can be raised for other types of relations for which annotators need to decide if they should distinguish these by adding a separate relation label to their relation inventory. similarly, the method used in this paper could be used to decide whether a relation label should be maintained within an inventory of coherence relations, or whether the label is not necessary because there is no relational surplus and further specifications can be made on the basis of segment-specific features. in their comparison of annotation frameworks, sanders et al. (submitted) have listed several features that are distinguished in one or more of the other frameworks but not in ccr, thereby presenting a list of candidates for which this discussion seems relevant. this includes specificity, which seems to play a role in instantiation relations and the subtypes of restatements in pdtb, as well as in rst-dt’s collection of elaboration relations. other candidates that require further investigation are list, which sets apart list relations from other types of additive relations, and features that discriminate between the different types of conditionals found in pdtb and rst-dt. we have shown that the cognitive plausibility of the relation type under investigation can be tested fruitfully using empirical evidence from acquisition data, processing, and the system of linguistic markers. acknowledgements the first author’s work was enabled by a grant awarded by the executive board of utrecht university to the anncor project, work package discourse annotation. the second author was funded by the snsf sinergia project modern (crsii2_147653). the third author was funded evers-vermeul, hoek and scholman 16 by the german research foundation (dfg) as part of sfb 1102 “information density and linguistic encoding”. references barbara e. adam (2008). the timescapes challenge: engagement with the invisible temporal. in b. e. adam, j. hockey, p. thompson, paul and r. edwards (eds.), researching lives through time: time, generation and life stories, timescapes working paper series, volume 1: 7-12. leeds: university of leeds. nicholas asher and alex lascarides (2003). logics of conversation. cambridge: cambridge university press. fatemeh t. asr and vera demberg (2015). uniform information density at the level of discourse relations: negation markers and discourse connective omission. in iwcs 2015, pages 118128, london, united kingdom. caroline atallah, myriam bras and laure vieu (2015). discourse relations, discourse connectives and discourse segmentation interdependency in the light of causality. paper presented at lpts2016. valencia, spain. linda baker (1978). processing temporal relationships in simple stories: effects of input sequence. journal of verbal learning and verbal behavior, 17(5): 559-572. michael barlow and suzanne kemmer (2000). usage based models of language. stanford: csli publications. otto behaghel (1928). deutsche syntax. die satzgebilde, band iii. heidelberg: carl winters universitätsbuchhandlung.ruth berman and dan i. slobin (1994). relating narrative events: a crosslinguistic developmental study. hillsdale, nj: lawrence erlbaum. yves bestgen and jean costermans (1994). time, space, and action: exploring the narrative structure and its linguistic marking. discourse processes, 17(3): 421-446. thomas g. bever (1970). the cognitive basis for linguistic structures. in r. hayes (ed.), cognition and language development: 279-362. new york: wiley & sons. diane blakemore (2002). relevance and linguistic meaning: the semantics and pragmatics of discourse markers. cambridge: cambridge university press. lois bloom, margaret lahey, lois hood, karin lifter and kathleen fiess (1980). complex sentences: acquisition of syntactic connectives and the semantic relations they encode. journal of child language, 7(2): 235-261. liam p. blything, robert l. davies and kate e. cain (2015). young children’s comprehension of temporal relations in complex sentences: the influence of memory on performance. child development, 86(6): 1922-1934. bruce k. britton (1994). understanding expository text: building mental structures to induce insights. in m. a. gernsbacher (ed.), handbook of psycholinguistics: 641-674. san diego, california: academic press. harry bunt and rashmi prasad (2016). iso dr-core (iso 24617-8): core concepts for the annotation of discourse relations. in proceedings 12th joint acl-iso workshop on interoperable semantic annotation (isa-12), pages 45-54, portoroz, slovenia. kate e. cain and hannah m. nash (2011). the influence of connectives on young readers’ processing and comprehension of text. journal of educational psychology, 103(2): 429-441. anneloes r. canestrelli, willem m. mak and ted j.m. sanders (2013). causal connectives in discourse processing: how differences in subjectivity are reflected in eye-movements. language and cognitive processes, 28(9): 1394-1413. lynn carlson, daniel marcu and mary e. okurowski (2003). building a discourse-tagged corpus in the framework of rhetorical structure theory. in j. van kuppevelt and r. smith (eds.), current directions in discourse and dialogue: 85-112. dordrecht: kluwer academic publishers. on temporality in discourse annotation 17 eve v. clark (1971). on the acquisition of the meaning of before and after. journal of verbal learning and verbal behavior, 10(3): 266-275. herbert h. clark and eve v. clark (1968). semantic distinctions and memory for complex sentences. the quarterly journal of experimental psychology, 20(2): 129-138. bernard comrie (1976). aspect. cambridge: cambridge university press jean costermans and yves bestgen (1991). the role of temporal markers in the segmentation of narrative discourse. cpc: european bulletin of cognitive psychology, 11, 349-370. laurence danlos (2006). ‘discourse verbs’ and discourse periphrastic links. in c. sidner, j. harpur, a. benz and p. kühnlein (eds.), proceedings of the second workshop on constraints in discourse (pp.59-65). maynooth, ireland. evers-vermeul, j. (2005). the development of dutch connectives: change and acquisition as windows on form-function relations. phd thesis, utrecht university. utrecht: lot. available at: http://www.lotpublications.nl/documents/110_fulltext.pdf. jacqueline evers-vermeul and ted j.m. sanders (2009). the emergence of dutch connectives: how cumulative cognitive complexity explains the order of acquisition. journal of child language, 36(4): 829-854. lucia a. french and ann l. brown (1977). comprehension of before and after in logical and arbitrary sequences. journal of child language, 4(2): 247-256. silvia p. gennari (2004). temporal references and temporal relations in sentence comprehension. journal of experimental psychology: learning, memory, and cognition, 30(4): 877-890. talmy givón (1990). syntax: a functional-typological introduction, volume 2. amsterdam/ philadelphia: benjamins. ana i. gonzález-cruz (2007). on the subjectification of adverbial clause connectives: semantic and pragmatic considerations in the development of while-clauses. in u. lenker and a. meurman-solin (eds.), connectives in the history of english: 145-166. amsterdam/ philadelphia: john benjamins. h. paul grice (1989). studies in the way of words. cambridge, ma/london: harvard university press. michael a. k. halliday and ruqaiya hasan (1976). cohesion in english. london: longman. alison hall (2007). do discourse connectives encode concepts or procedures? lingua, 117(1): 149-174. erhard hinrichs (1986). temporal anaphora in discourses of english. linguistics and philosophy, 9: 63-82. jerry r. hobbs (1983). why is discourse coherent? in f. neubauer (ed.), coherence in natural language texts: 29-70. hamburg: buske. eduard h. hovy and elisabeth maier (1995). parsimonious or profligate: how many and which discourse structure relations. unpublished manuscript. available at: http://www.isi.edu/natural-language/people/hovy/papers/93discproc.pdf. helen l. johnson (1975). the meaning of before and after for preschool children. journal of experimental child psychology, 19(1): 88-99. andrew kehler (1994). temporal relations: reference or discourse coherence? in proceedings of the 32nd annual conference of the association for computational linguistics (acl-94), pages 319-321, las cruces, new mexico. andrew kehler (2002). coherence, reference, and the theory of grammar. stanford: csli publications. rudi keller (1995). the epistemic weil. in d. stein and s. wright (eds.), subjectivity and subjectivisation: linguistic perspectives: 16-30. cambridge: cambridge university press. deborah keller-cohen (1987). context and strategy in acquiring temporal connectives. journal of psycholinguistic research, 16(2): 165-183. alistair knott and robert dale (1994). using linguistic phenomena to motivate a set of coherence relations. discourse processes, 18: 35-62. evers-vermeul, hoek and scholman 18 alex lascarides and nicolas asher (1993). temporal interpretations, discourse relations and common sense entailment. linguistics and philosophy, 16(5): 437-493. pierre lejeune, amália mendes and nuno martins (2016). some considerations on the use of main verbs to express rhetorical relations. in l. degand, c. dér, p. furkó, and b. webber (eds.). conference handbook textlink – structuring discourse in multilingual europe second action conference (pp.81-85). budapest: debrecen university press. stephen c. levinson (2000). presumptive meaning: the theory of generalized conversational implicature. cambridge, ma/london: mit press. robert e. longacre (1983). the grammar of discourse: notional and surface structures. plenum: new york. max louwerse (2001). an analytic and cognitive parameterization of coherence relations. cognitive linguistics, 12(3): 291-315. willem m. mak and ted j.m. sanders (2013). the role of causality in discourse processing: effects of expectation and coherence relations. language and cognitive processes, 28: 14141437. jean m. mandler (1986). on the comprehension of temporal order. language and cognitive processes, 1(4):309-320. james r. martin (1992). english text: system and structure. amsterdam/philadelphia: john benjamins. amália mendes and pierre lejeune (2016). ldm-pt: a portuguese lexicon of disourse markers. in l. degand, c. dér, p. furkó, and b. webber (eds.). conference handbook textlink – structuring discourse in multilingual europe second action conference (pp.89-92). budapest: debrecen university press. thomas f. münte, kolja schiltz and marta kutas (1998). when temporal terms belie conceptual order. nature, 395: 71-73. john nerbonne (1986). reference time and time in narration. linguistics and philosophy, 9: 8395. barbara h. partee (1984). nominal and temporal anaphora. linguistics and philosophy, 7: 243286. pdtb research group (2008). the penn discourse treebank 2.0 annotation manual. technical report ircs-08-01. philadelphia: institute for research in cognitive science, university of pennsylvania. available at: https://www.seas.upenn.edu/~pdtb/pdtbapi/pdtb-annotationmanual.pdf. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind k. joshi and bonnie l. webber (2008). the penn discourse treebank 2.0. in proceedings of the 6th international conference of language resources and evaluation (lrec 2008), marrakech, morocco. available at: https://www.seas.upenn.edu/~pdtb/papers/pdtb-lrec08.pdf. rashmi prasad, aravind k. joshi and bonnie l. webber (2010). realization of discourse relations by other means: alternative lexicalizations. in proceedings of the 23rd international conference on computational linguistics (coling 2010), posters volume, pages 10231031, beijing, china. pirita pyykkönen and juhani järvikivi (2012). children and situation models of multiple events. developmental psychology, 48(2): 521-529. pirita pyykkönen, jussi niemi and juhani järvikivi (2003). sentence structure, temporal order and linearity: slow emergence of adult-like syntactic performance in finnish. sky journal of linguistics, 16: 113-138. brian j. reese, julie hunter, nicholas asher, pascal denis and jason baldridge (2007). reference manual for the analysis of rhetorical structure. technical report, university of texas at austin, austin, texas. on temporality in discourse annotation 19 eleanor rosch (1977). classification of real-world objects: origins and representations in cognition. in p.n. johnson-laird & p.c. wason (eds.), thinking: readings in cognitive science (pp.212-222). cambridge: cambridge university press. magdaléna rysová (2012). alternative lexicalizations of discourse connectives in czech. in proceedings of the eight international conference on language resources and evaluation (lrec’12), pages 2800-2807, istanbul, turkey. magdaléna rysová and kateřina rysová (2015). secondary connectives in the prague dependency treebank. in proceedings of the third international conference on dependency linguistics (depling 2015), pages 291-299, uppsala. sweden. ted j.m. sanders, vera demberg, jet hoek, merel c.j. scholman, fatemeh torabi asr, sandrine zufferey and jacqueline evers-vermeul (submitted). unifying dimensions in coherence relations: how various annotation frameworks are related. submitted for publication. ted j.m. sanders and leo g.m. noordman (2000). the role of coherence relations and their linguistic markers in text processing. discourse processes, 29: 37-60. ted sanders and wilbert spooren (2009). causal categories in discourse: converging evidence from language use. in t. sanders and e. sweetser (eds.), causal categories in discourse and cognition (pp. 205-246). berlin: mouton de gruyter. ted sanders and wilbert spooren (2013). exceptions to rules: a qualitative analysis of backward causal connectives in dutch naturalistic discourse. text & talk, 33(3): 377-398. ted j.m. sanders, wilbert p.m. spooren and leo g.m. noordman (1992). toward a taxonomy of coherence relations. discourse processes, 15: 1-35. ted j.m. sanders, wilbert p.m. spooren and leo g.m. noordman (1993). coherence relations in a cognitive theory of discourse representation. cognitive linguistics, 4(2): 93-133. joanne scheibman (2002). point of view and grammar: structural patterns of subjectivity in american english conversation. amsterdam: john benjamins merel c.j. scholman, jacqueline evers-vermeul and & ted j.m. sanders (2016). categories of coherence relations in discourse annotation: towards a reliable categorization of coherence relations. dialogue and discourse, 7(2): 1-28. wilbert spooren (1997). the processing of underspecified coherence relations. discourse processes, 24: 149-168. manfred stede (2002). dimlex: a lexical approach to discourse markers. in a. lenci and v. di tomaso (eds.), exploring the lexicon: theory and computation (pp.1-15). alessandria: edizioni dell’orso. ninke stukker (2005). causality marking across levels of language structure. a cognitive semantic analysis of causal verbs and causal connectives in dutch. phd thesis, utrecht university. utrecht: lot. available at: http://www.lotpublications.nl/documents/ 118_fulltext.pdf. ninke m. stukker and ted j.m. sanders (2012). subjectivity and prototype structure in causal connectives: a cross-linguistic perspective. journal of pragmatics, 44: 169-190. ninke stukker, ted sanders and arie verhagen (2008). causality in verbs and in discourse connectives: converging evidence of cross-level parallels in dutch linguistic categorization. journal of pragmatics, 40: 1296-1322. elizabeth c. traugott (1995). subjectification in grammaticalisation. in d. stein and s. wright (eds.), subjectivity and subjectivisation: linguistic perspectives: 31-54. cambridge: cambridge university press. elizabeth c. traugott (2010). (inter)subjectivity and (inter)subjectification: a reassessment. in k. davidse, l. vandelanotte and h. cuyckens (eds.), subjectification, intersubjectification and grammaticalization: 29-71. berlin & new york: de gruyter mouton. elizabeth, c. traugott (2011). grammaticalization and mechanisms of change. in h. narrog and b. heine (eds.), the oxford handbook of grammaticalization: 19-30. oxford: oxford university press. evers-vermeul, hoek and scholman 20 elizabeth c. traugott and ekkehard könig (1991). the semantics-pragmatics of grammaticalization revisited. in e.c. traugott and b. heine (eds.), approaches to grammaticalization, volume i: 189-218. amsterdam/philadelphia: benjamins. anna trosborg (1982). children’s comprehension of ‘before’ and ‘after’ reinvestigated. journal of child language, 9(2): 381-402. david j. townsend (1983). thematic processing in sentences and texts. cognition, 13(2): 223261. paola uccelli (2009). emerging temporality: past tense and temporal/aspectual markers in spanish-speaking children’s intraconversational narratives. journal of child language, 36(5): 929-966. elke van der meer, reinhard beyer, bertram heinze and isolde badel (2002). temporal order relations in language comprehension. journal of experimental psychology: learning, memory, and cognition, 28(4): 770-779. rosie van veen, jacqueline evers-vermeul, ted sanders and huub van den bergh (2009). parental input and connective acquisition: a growth curve analysis. first language, 29(3): 267-289. rosie van veen, jacqueline evers-vermeul, ted sanders and huub van den bergh (2014). “why? because i’m talking to you!” parental input and cognitive complexity as determinants of children’s connective acquisition. in h. gruber & g. redeker (eds.), the pragmatics of discourse coherence: theories and applications: 209-242. amsterdam/philadelphia: john benjamins. ulrike vogl (2007). het belang van conditionaliteit voor de ontwikkeling van temporeel naar causaal voegwoord: de geschiedenis van dewijl, terwijl, weil en while [the importance of conditionality in the development from temporal to causal complementizer: the history of dewijl, terwijl, weil and while]. nederlandse taalkunde, 12(1): 2-24. bonnie l. webber (1987). the interpretation of tense in discourse. in proceedings of the 25th annual meeting on association for computational linguistics, pages 147-154, stroudsburg, pennsylvania. bonnie l. webber (1988). tense as discourse anaphor. computational linguistics, 14(2): 61-73. zheng ye, marta kutas, marie st george, martin i. sereno, feng ling and thomas f. münte (2012). rearranging the world: neural network supporting the processing of temporal connectives. neuroimage, 59(4): 3662-3667. rolf a. zwaan, (1996). processing narrative time shifts. journal of experimental psychology: learning, memory, and cognition, 22(5): 1196-1207. journal of machine learning research-microsoft word template dialogue and discourse 8(2) (2017) 149-166 doi: 10.5087/dad.2017.207 ©2017 ludivine crible and maria josep cuenca this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). discourse markers in speech: characteristics and challenges for corpus annotation ludivine crible ludivine.crible@uclouvain.be université catholique de louvain place blaise pascal, 1 1348 louvain-la-neuve belgium maria-josep cuenca maria.j.cuenca@uv.es universitat de valència facultat de filologia av. blasco ibánez 32 46010 valència spain editor: manfred stede submitted 04/2016; accepted 10/2017; published online 12/2017 abstract it is generally acknowledged that discourse markers are used differently in speech and writing, yet many general descriptions and most annotation frameworks are written-based, thus partially unfit to be applied in spoken corpora. this paper identifies the major characteristics of discourse markers in spoken language, which can be associated with problems related to their scope and structure, their meaning and their tendency to co-occur. the description is based on authentic examples and is followed by methodological recommendations on how to deal with these phenomena in more exhaustive, speech-friendly annotation models. keywords: discourse markers, corpus annotation, speech, linguistic complexity, mode of communication 1 introduction within the large literature on discourse markers (henceforth dms), a number of studies investigate how uses differ according to the mode of communication (see for instance chafe, 1982; horowitz and samuels, 1987; castellà, 2004; biber, 2006; lópez serena and borreguero zuluoga, 2010). yet corpus-based analysis of spoken discourse is still quite infrequent, or focused on very specific phenomena (see lópez serena and borreguero zuluoga, 2010; ciabarri, 2013; cuenca, 2013a; crible, 2017 for some exceptions). similarly, most annotation models of dms are designed for written discourse. this is the case of the rhetorical structure theory or rst (mann and thompson, 1988), the penn discourse treebank or pdtb 2.0 (prasad et al., 2008) and the cognitive approach to cognitive relations or ccr (sanders et al., 1992) – although there are some recent endeavors to apply these frameworks to speech (e.g., tonelli et al., 2010 for pdtb in spoken italian). despite this predominance of writing-based models, spoken and spoken-like corpora are increasingly available with large databases such as the spoken bnc2014 project for english or c-oral-rom for romance languages. therefore, this paradoxical situation calls for the need to take into account the discourse markers in speech 150 particular characteristics of spoken language both for descriptive adequacy and in order to extend writing-based annotation models, or even to design multimodal models (crible and zufferey, 2015). this paper aims at identifying the major characteristics of discourse markers that relatively differ (in nature, frequency or degree) between spoken and written language, and discusses the problems that they pose for corpus annotation. besides examples and some corpus-based quantitative findings of the distribution of these features, the paper also provides a review of how existing writing-based frameworks handle (or not) the particular characteristics of spoken dms. this qualitative discussion of the challenges of spoken discourse annotation will lead to methodological recommendations to enhance the validity of future corpus-based research. 2 spoken vs. written mode of communication and the use of dms in this section, we report on previous descriptive works which identified relative differences between the spoken and written modes of communication (section 2.1), and their resulting impact on the use of dms (section 2.2). 2.1 characteristics of speech and writing it seems uncontroversial that written and oral texts are produced and processed differently (see, e.g., horowitz and samuels, 1987). as a consequence, dms are expected to behave differently according to the mode of communication. however, linguistic complexity in speech and writing is a controversial matter. two opposite positions can be distinguished. chafe (1982) and many other authors consider that oral grammar is simpler than written grammar. chafe argues that there are two dichotomies that account for this fact: syntactic fragmentation vs. integration and detachment vs. involvement. halliday (1987) represents the opposite, and less popular, perspective. halliday argues that the grammar of speech is complex and so the idea that speech is structureless is a myth. “speech is not, in any general sense, ‘simpler’ than writing; if anything, it is more complex” (halliday, 1987: 65). what seems to be uncontroversial is that spoken and written language differ as for their respective kind of complexity. written language tends to be lexically dense but grammatically “simple” or rather “regular”. conversely, speech is grammatically intricate but lexically sparse, as illustrated by the frequent cases of underspecification, implicitness and/or ambiguity in spoken language: “speech [is] processlike, intricate, with meanings related serially” (halliday, 1987: 79-80). their preferred patterns of organization are therefore different. an important methodological assumption (or conclusion) of halliday’s perspective is that the term mode of communication must be understood in a wide sense: mode is not only (or mainly) the medium, but the contextual factors attached to it.1 although the mode of communication, either spoken or written, is an important factor to account for the selection and frequency of discourse markers, other factors, namely planning and interactivity, must also be taken into account. planning is more likely to be carried out in written texts. planning is directly related to the following features: “face-to-text”, not space-time bound, secondary discourse and consciousness. it can also be directly associated with formality, since informal texts are usually not planned, whether they are oral or written. that is why an informal and non-revised text by a non-expert often sounds like spoken language. the other side of planning is revision once the message has been textualized. revision is only possible with permanent products. interactivity, that is, the presence and activity of the hearer explicitly expressed from the speaker’s point of view, is more likely to occur in spoken texts or, rather, in face-to-face interaction, 1 the same position is put forward in bazzanella (2001). discourse markers in speech 151 including “speech-like” modes of mediated communication (e.g. chats). interactivity is related to face-to-face communication, addressor-addressee reciprocity and interpersonal discourse, as opposed to objective and distant communication. it can be related to an additional factor, namely, emotivity. emotivity is typically linked to orality, informal communication and unplanned texts (see cuenca and romano, 2013). in fact, koch and österreicher (1990: 12) identify intimacy and emotional implication as factors defining communicative proximity (vs. affective distance and lack of emotional implication).2 in short, speech is not always unplanned and interactive (e.g. a broadcast political speech) and writing is not necessarily planned and non-interactive (e.g. text-based chats). the opposition between speech and writing should therefore be refined by taking these contextual factors into account, so that we will henceforth always specify what particular type of discourse is at stake. 2.2 dms in speech vs. writing a number of previous analyses of the linguistic complexity of oral and written texts have taken into account the use of dms. in this section, two of them will be reviewed, namely castellà (2004) and biber (2006). castellà (2004) analyzes three discourse genres in catalan (informal conversation, academic lecture and written academic prose) by focusing on three aspects of linguistic structure and complexity: lexical density, compound sentences, and text connection. castellà’s main conclusion is that spoken informal texts, spoken formal texts (expository) and written formal texts result from different discourse-building strategies (2004: 149). this fact has an impact on both the linguistic complexity of the discourse and the use of dms. specifically,  spoken conversation exhibits a verbal style (more verbs and verbal complements) whereas written exposition has a more nominal style (more nouns and noun complements). as a consequence, written prose organizes information through integration of structures (complex phrases, mainly noun phrases, and clauses), while conversation includes more reduced phrases and clauses, flexible structures and more repetition. writing is, then, denser and more lexically varied.  written texts tend to relate text units as compound sentences, whereas spoken texts make a profuse use of text connection.  orality combines intonation, sentential connection and text connection to link units. writing, in contrast, heavily relies on sentential connection alone.  all texts, whether written or spoken, use a similar number of subordination constructions. informal conversations include a similar number of subordinators in absolute terms and in relation with the number of words, although fewer if we consider the total number of clauses. this puzzling fact (the paradox of complexity, as castellà names it) has to do with the tendency to an increase of subordinators because of two different preferred strategies: a verbal style, typical of speech, implies connection between verbal units; similarly, an integrated style, typical of writing, relies on sentential connection. as for the use of dms, some conclusions can be drawn from castellà’s analysis:  informal conversation resorts more often to conjunctions and all kinds of dms as text connectives. spoken expository prose is more similar to written prose in the use of conjunctions as text connectives, but to informal conversation in the use of other connectives.  spoken genres use a more reduced variety of connectives, but they are more polyfunctional. 2 in the words of tannen (1982: 15), interpersonal involvement characterizes orality, whereas written texts are more message-oriented. discourse markers in speech 152  some connectives and some uses of certain connectives seem to be typical of one type of text. douglas biber, in his monograph on university language (2006), analyzes a wide range of the registers found in american universities, both spoken (classroom teaching, classroom management, office hours, study groups, service encounters) and written (academic textbooks, course management, institutional writing). as for the use of dms, biber (2006: chapter 4) concludes that:  linking adverbials, such as however, therefore, for example, that is, are found in both spoken and written registers but “are primarily characteristic of written registers” (2006: 70), mainly of textbooks.  “discourse markers” (such as okay, well, now) are “restricted primarily to spoken discourse” (2006: 66).  as for dependent clauses, “overall patterns of use […] run exactly opposite to […] prior expectations, with dependent clauses overall being more common in the spoken university registers than in the written registers” (biber, 2006: 72). specifically, adverbial clauses and complement clauses, especially that-clauses, are more frequent in speech, whereas relative clauses occur in both modes of discourse depending on the genre but are more common in written registers. to sum up, our intuition as linguists or even as speakers is that, since written and spoken texts are produced and processed in a different way, discourse markers should exhibit a different behavior accordingly. however, previous studies have shown that the analysis of written and spoken texts does not lead to a straightforward conclusion. orality is not a homogeneous phenomenon and thus there are many intermediate cases between the pole represented by (prototypical) spoken texts and (prototypical) written texts. similarly, it is possible to identify spoken texts with a structure and use of dms that resemble that of written texts, and also written texts that echo orality. nonetheless, the analysis of spoken texts shows that the use of dms tends to exhibit some relative differences or tendencies. this is especially the case in spontaneous speech, where planning is low, and in dialogue, where interactivity is high. 3 materials and present approach to dms in the following sections, some characteristics of spoken dms will be discussed and illustrated by naturally-occurring examples taken from disfren, a french-english comparable corpus fully annotated for discourse markers (crible, 2017b). this dataset comprises 15 hours of recordings (161,700 words) balanced across a variety of more or less formal settings such as conversations, interviews or political speeches. the english transcripts were compiled from existing corpora, namely the british component of the international corpus of english (nelson et al., 2002) and the backbone project (kohn, 2012). a total of 8,743 dm items have been identified and annotated following the definition and guidelines in crible (2017a): dms were identified through a bottomup selection (no closed list) provided they met the criteria of procedurality, syntactic optionality, fixedness (i.e. grammaticalized), discourse-level scope and metadiscursive function (discourse relation, topic structure, turn exchange, speaker-hearer relationship). this definition includes in effect expressions as varied as conjunctions (and, but, although), adverbs (actually, well), prepositional phrases (in fact, by the way) or verb phrases (you know, i mean, if you will), among others. the annotation covers syntactic position, co-occurrence and sense disambiguation. crible and degand (forthcoming) report on moderate inter-annotator agreement score (fleiss’ kappa = 0.563, relative agreement 70.9%) while intra-annotator agreement, performed on 10% of the whole corpus, is higher (𝜅 = 0.74, 75.8% of relative agreement). this data will be used to quantify the discourse markers in speech 153 characteristics under discussion in this paper, comparing, when available, results from available written corpora (mostly the pdtb corpus of wall street journal, prasad et al., 2008). 4 dms in unplanned speech: the scope of dms and discourse structure the standard use of dms implies two complete and identifiable arguments, which are consecutively attached to a dm. however, in speech, and especially in unplanned and/or interactive discourse, this is not always the case. one of the two arguments can be missing, incomplete or implicit (truncated structures) and the two arguments can be complex and interrupted by different arguments (far-reaching scope), so that the linear structure corresponding to the construction “arg1 dm arg2” does not always match its semantic structure. 4.1 truncated structures structures containing dms in unplanned speech are often truncated (that is, the second argument is missing or is not complete) or independent (the first argument is implicit), whereas in written/planned discourse the structures tend to be continuous and complete because the text is presented as a finished product (instead of an on-going process in speech, cf. halliday 1987: 65). this fragmentation of unplanned spoken syntax affects the annotation process by blurring the function of dms in truncated structures and generates incomplete argument pairs. stent (2000) already noticed this feature of dialogs in her application of rst to conversational data, which led her to discard abandoned utterances from the scope of the annotation. she further notes that full representation of discourse structure as proposed by the rst framework cannot be fully achieved “given the length and complexity of a typical dialog” (2000: 250). cases of missing arg2 and missing arg1 are respectively illustrated in examples (1) and (2). (1) ice24: i was hoping it’d be a week so i could get off the ward for a bit but ice25: oh already? ice12: well you know it’s nice to have a change (en-phon-05) (2) ice25: i’ll walk down anyway so i can get one ice26: oh right ok ice25: so (0.573) um (1.730) so are you still on call at work (en-phon-05) in (1), the second argument of “but” is not verbally expressed, leaving ice25 to infer what is left implicit. in (2), ice25 is struggling to come up with a new topic and uses two tokens of “so” to hold the floor while planning the upcoming utterance: the argument introduced by “so” bears no relation with previous context and “so” simply expresses a stalling and topic-shifting function. in these examples, interactivity plays an essential part since, in both cases, the incomplete structure is either turn-final (1) or turn-initial (2), pointing to the specific role of turn transitions in dm behavior. such non-relational functions of dms are excluded from writing-based frameworks such as rst and pdtb where two abstract objects are required for any sense to be assigned. yet, absence of left or right context is quite frequent for spoken dms, as already mentioned in degand and simon-vandenbergen (2011: 289), who talk about a “scale” of relationality, from “non-relational” markers (e.g., actually) to “strictly relational” (e.g., because) discourse markers. some uses have become conventionalized as specific discourse functions: estellés and pons (2014) identified the “absolute initial position” which can only be filled by a limited number of dms in spanish (e.g., bueno ‘well’) that do not require previous context and that are typically used to signal the beginning of a new interaction. similarly, final position (as in example 1 above) is generally associated with intersubjective, elliptical or common-ground functions that invite the hearer to contribute to the discourse representation (see haselow, 2012; buysse, 2014; degand, 2014), functions which are completely absent from most taxonomies of discourse relations. to date, most works on spoken discourse markers in speech 154 corpora (e.g. tonelli et al., 2010; rehbein et al., 2016; riccardi et al., 2016) make a distinction between “connectives” and “discourse markers”, that is, relational vs. non-relational uses, the latter being discarded although they are sometimes performed by the same expressions (cf. but and so in the examples above), thus maintaining a writing-based, strictly relational view of dms. 4.2 far-reaching scope structures containing dms in speech often display long-distance relationships, or “far-reaching scope” (lenk, 1998: 247), as a result of the dynamic process of building a shared discourse representation. this feature is mainly displayed in two configurations: a dm relating to a distant previous context of which it is separated by a digression (location), or a dm taking scope over multiple utterances which, once combined, constitute its left argument (extent). this flexibility is not specific to speech: unger (1996: 409), for instance, acknowledges that, in writing as well, “discourse connectives can have scope over an utterance or a group of utterances”. similarly, in the pdtb 2.0 (prasad et al., 2008: 3), the extent and location of arg1 are coded as separate variables, amounting to 617/18459 (3.34%) cases where arg1 comprises multiple utterances and 1666/18459 (9%) cases where arg1 is non-adjacent to arg2 (the possibility of a non-adjacent arg1 is restricted to explicit connectives only). however, some writing-based taxonomies do not include relations which specifically function with a far-reaching scope, at a higher level of discourse organization. feng et al. (2014: 1) oppose in this regard pdtb-style vs. rst-style annotations: “pdtb-style discourse relations encode only very shallow discourse structures, i.e., the relations are mostly local, e.g. within a single sentence or between two adjacent sentences”, as opposed to rst which includes “long-distance discourse dependency” and higher-level relations such as topic-shift or topic-drift. in other words, according to the pdtb, the same discourse relation can have a local or global scope, but no relation is specific to the latter: the only potentially far-reaching relation type, viz. “list”, has been removed from the latest version of their taxonomy (pdtb3, webber et al., 2016).3 similarly, in ccr, relations between topics such as topic-shift or topic-resuming are excluded on the grounds that they are “orthogonal to another classification in terms of coherence” (sanders et al., 2016), that is, any marker or relation expressing a higher-level function (e.g. topic-shift) necessarily encodes a more local relation such as contrast or consequence, thus forcing a lexical-semantic interpretation of the dm regardless of its meaning-in-context. example (2) above seems to contradict this strong position, where “so” cannot be related to the previous context by any relation other than topic-shift. it is, however, true that in some cases, the scope of the dm is not clear and seems to combine local and global meanings, as illustrated by example (3) where both location and extent of arg1 are quite flexible: (3) bb1: could you talk a little bit about the wirral accent i i know that um (0.200) there’s obviously quite a um range of accents in that part of the country bb4: yeah (0.520) uh well i (0.290) consider myself to have a cheshire accent because when i was born (0.300) and i lived in (0.110) on the wirral (0.287) uh (0.333) i (0.460) it was a cheshire accent which is (0.440) the accent i have now though (0.270) there are overtones of (0.230) the liverpudlian accent (0.290) however over the years certainly it has changed (0.270) and now it’s very much (0.110) a liverpool accent (0.340) and uh you know which (0.430) i’m not (0.300) i’m not saying i disapprove of it but i think it’s a lazy speech and you need to (0.440) 3 it should be noted that, in the pdtb, the “norel” (i.e. absence of relation) option may mask the presence of possible long-distance relations, especially when their identification is unclear, in order to enhance consistency in the annotation. it remains that the pdtb taxonomy only covers typically local relations such as cause or contrast, and not typically farreaching relations such as topic-shift. discourse markers in speech 155 actually um (0.530) think about what you’re saying i know my nephew sometimes’ll to speak to me in the liverpool accent (0.350) and i’ll say please speak to me in english (0.160) but it’s things like “yeah” and “you what” and (0.230) whereas you know mine is “yes” “pardon” or whatever i’m a bit old-fashioned in that way so i do find the accent (0.440) is a bit harsh and it’s interesting that actually that accent is spread out into the (0.270) uh (0.390) the parts of north wales that are very near to the wirral… (en-intf-03) in example (3), the dm “so” introducing “i do find the accent is a bit harsh” can be interpreted in multiple ways depending on which left context it refers to. the most local interpretation connects it to the immediate co-text (“i’m a bit old-fashioned in that way”), in which case “so” would signal a consequence relation. another possible interpretation is to relate it to the “nephew” anecdote, of which it introduces a conclusion. more convincingly, the evaluative content of the verb (“find”) and adjective (“harsh”) in the second argument echoes an earlier evaluative utterance: “i’m not saying i disapprove of it but i think it’s a lazy speech”. lastly, “so” could also be interpreted as introducing the answer to the original question by the interviewer “could you talk a little bit about the wirral accent”, in which case its function is to come back to the main topic after a lengthy digression. while a conservation decision would only annotate the local interpretation (consequence relation), a more accurate and comprehensive analysis could combine two senses (consequence and topic-resuming), an option which is available in many annotation models for both speech and writing. complex examples such as (3) tend to show that, while systematic annotation of dm scope would be particularly challenging (especially in spoken data), speech-friendly taxonomies should not deny the ability of dms to express higher-level functions such as coming back to a previous topic. in the disfren corpus, topic-relating dms take up 6.07% of all annotated tokens, which would not be accounted for by writing-based models with a more restricted view of dm functions. given that these functions are often expressed by frequent connectives such as and or so, which are also used in writing, this observation of far-reaching dms in speech calls for more inclusive taxonomies for written corpora as well, in order to reach a better comparability across annotated corpora. it thus appears that spoken dms not only tend to take scope over large and distant segments (as in writing), but can also combine local and global scope simultaneously, making the annotation process quite complex and calling for specific function types at a higher level of discourse structure than what is included in local views of coherence (pdtb, ccr). 5 the meaning of dms one of the most commented features of dms is their multifunctionality (see for instance mosegaard hansen, 2006, 2008; hummel, 2012). this feature, which can be directly associated with semantic and pragmatic ambiguity, is especially outstanding in speech, where a limited number of markers tend to be used with a relatively wide range of meanings and where context is a key-element in discourse production and processing. 5.1 vagueness the meaning of a dm in spoken/unplanned discourse is often (semantically and pragmatically) vague or not clearly definable, as opposed to the main tendency in written/planned discourse.4 dms can be ambiguous in writing as well, but these cases are much more typical of speech, especially 4 tuggy (1993) defines “vagueness” in cognitive-grammar terms as “meanings which are not well-entrenched but which have a relatively well-entrenched, elaboratively close schema subsuming them” (1993: 8). discourse markers in speech 156 in face-to-face settings where the speaker might expect the hearer to make use of various contextual cues (e.g., prosody, gestures) to reconstruct the speaker’s intended meaning, an option which is not available for readers (schober and brennan, 2003). in addition, speech-specific dms such as well or french quoi ‘you know’ do not encode a strong core meaning, as opposed to more explicit expressions in writing such as by contrast or in addition, which are virtually absent in speech: for instance, in disfren, only one occurrence of in addition was found out of the total 8,743 annotated dms, against 165 in the pdtb corpus out of 18,459 tokens. while not all written dm occurrences are clear and unproblematic, similarly, not all spoken dms are vague, and some forms tend to specialize in a given function, such as the very frequent you know which is mostly used to check for the hearer’s attention and understanding. yet, two degrees of difficulty related to vagueness are illustrated in (4) and (5) where the dm i mean departs from its basic reformulative meaning. (4) ice76: didn’t john use to deal with uhm (1.850) divorce in his earlier days ice77: did he uhm ice76: found it quite depressing ice77: i imagine it probably is uhm ice76: all the toing and froing and ice77: i mean i shi should think you’d get over the uhm (0.110) the voyeuristic aspects in the early stages (en-conv-04) (5) ice12: this is what we do all the time we sit and describe other people and i mean when people got stuck i’d just say look just list you imagine... (en-conv-03) in (4), i mean elaborates on the speaker’s previous turn but it is not clear whether it introduces more detail (“it probably is depressing, by that i mean that you would get over the voyeuristic aspects”) or a justification (“it probably is depressing, i say that because you would get over the voyeuristic aspects”), or a combination of both (see next section on multifunctionality). in (5), i mean does not fluctuate between two meanings, like in (4), but rather seems completely underspecified and functions as a semantically void “punctuant” (vincent, 1993). in addition, its co-occurrence with and blurs its own functional contribution into the additive meaning of the conjunction (see section 5 on co-occurrence). the treatment of cases like (4) and (5) does not only depend on the quality and coverage of the functional taxonomy used by the annotator. these problems are mostly due to the intrinsic ambiguity of language in general and of some dms in particular. the challenge for the annotator is then to assign the appropriate function label within the taxonomy, striving to avoid “undefined” or “non-interpretable” categories which usually end up left out of quantitative analyses as outlier data. therefore, a marker-based approach to dms, i.e. starting from the identification of (a set of) dms and systematically annotating their function(s), presents the additional challenge to deal with such vague cases and not only the more typical relational uses (such as i mean expressing reformulation or specification). 5.2 polysemy and multifunctionality of dms as already indicated, another challenging characteristic of dms – in general and especially in speech – is their multifunctionality, which is at the very core of the category to a greater extent than more homogeneous pragmatic categories such as modal particles (associated with epistemic modality), interjections (associated with subjectivity) or response signals (associated with agreeing/disagreeing). as suggested in crible (2017a: 107), three forms of multifunctionality can be distinguished: (1) the category covers items that perform many different functions; (2) a single member can perform different functions in different contexts; and (3) a single member can perform different functions simultaneously in the same context, given the great polysemy of dms. discourse markers in speech 157 we will now turn to each of these degrees of multifunctionality to discuss their challenge for corpus annotation. 5.2.1 one category with different functions dms can exhibit propositional functions indicating a logico-semantic relationship (e.g., consequence), but also non-propositional, mainly structural functions (e.g., topic-shift). in addition, some dms add a modal, (inter)subjective value (e.g., agreement, monitoring, emphasis, etc.) to a structural function. these three categories (viz. propositional, structural, modal), borrowed from cuenca (2013a), have a different weight in speech and in writing. dms acting at a propositional level are prevalent in planned monologues (36.5% of all dms in news broadcasts vs. only 20% in conversations in disfren), which can be taken as evidence for their higher attraction to written-to-be-spoken and, by extension, written discourse. regarding function types, however, there is a common core of propositional relations in speech and writing, with examples of cause or contrast, as in (6), which are equally included in writing-based and speech-based taxonomies. (6) i wasn’t looking forward to do it but i am now it’s going to be good (en-phon-01) dms with structural functions are typical of both speech and writing, but each mode of communication tends to prefer certain types of structural functions over others. in (spoken and written) situations with intermediate or high planning, markers of global structure such as topicshift or enumeration (example 7) are quite frequent. in the disfren corpus, news broadcast (highly prepared, written-to-be-spoken) is the only register where the topic-shift function ranks among the most frequent functions (5th and 4th in english and french, respectively). on the other hand, situations with low planning will make more use of local structuring devices such as continuity (example 8). (7) i will begin with a review of the economic situation and prospects (0.380) i shall then deal with monetary policy and public finances (0.220) finally i will present my tax proposals (en-poli-02) (8) you sit there in a canteen wherever you are and you sit and you listen and the conversation next to you someone’s saying (en-conv-03) global structural functions, such as discourse beginning, discourse closing or pre-closing, topic shift or turn change, are generally not taken into account – or not systematically – in writing-based taxonomies for dm functions, although they are very frequent in speech (14% of all dms in disfren; cf. examples 2, 3, 7). as already mentioned above, only rst includes a category of topic relations, while the only structural relation in the pdtb is the local conjunction relation (typically expressed by and as in example 8). dms that bracket units of talk while also performing modal functions are frequent in genres where interactivity is more prominent such as dialogues. modal markers introducing units of talk are typically oral (e.g., well, look, listen). some of these uses are covered by the question-answer relation in rst, although not all turn-initial dms actually belong to such a well-defined structure, as in (9) where “look” follows an imperative utterance and not an interrogative one, thus expressing disagreement (modal) as well as turn-opening (structural). (9) ice5 have some banana bread ice6 look i’m not that much of a banana bread eater (en-phon-08) regardless of the classifying system, authors agree on the multifunctionality of the dm category, although not always to the same degree of inclusiveness and granularity. as mentioned above, structural functions are usually excluded from writing-based taxonomies (e.g. pdtb, ccr) on the grounds that these relations function at a different, higher level of discourse coherence. modal functions are, in turn, absent from these frameworks since they fail to meet the relationality criterion (e.g. a you know does not relate two arguments but merely takes scope over one utterance). discourse markers in speech 158 however, such an exclusion overlooks emerging modal uses of typical dms such as french mais ‘but’ which, besides its relational functions (contrast, concession) can also be used structurally and/or modally in speech, as in (10). (10) c’est fatigant tu vois on prend leur bics pour le tu ben euh (0.900) quoi mais on n’a pas dit qu’on voulait bien gnagna tu vois it’s exhausting you know we take their pens for the tu well uh (0.900) what mais ‘but’ we did not say we were okay blabla you know (fr-conv-05) in this extract, the speaker is pseudo-reporting someone else’s words (“quoi mais on n’a pas dit qu’on voulait bien”), where “mais” opens the reported segment and expresses disagreement. once more, in a dm-based approach, such uses of “mais” cannot be discarded and have to be accounted for beyond what a purely relational (contrastive) reading would suggest. furthermore, the inclusion of structural and modal functions in dm taxonomies is in line with seminal, empirically-evidenced theories of discourse domains such as halliday’s (1970) ideational-textual-interpersonal distinction, taken up by schiffrin (1987), redeker (1991) and many others since then. 5.2.2 one marker with different functions in different contexts dms in speech/unplanned discourse tend to be multifunctional (mosegaard hansen, 2008; hummel, 2012). while multifunctional or ambiguous expressions do occur in writing/planned discourse as well (e.g. and), writing tends to rely on a larger diversity of connectives, some of which have highly specific meanings (e.g. in addition), as opposed to speech where a few very generic markers are used extensively for many different functions. in disfren, spoken registers with a high degree of preparation (political speeches, news broadcasts) show the highest type-token ratio of dms (around 20 vs. only 6 in unplanned speech such as casual conversations). these corpus-based findings suggest that dm expressions are more varied in writing and in planned “written-to-be-spoken” discourse (see also cuenca 2013b). a case in point is the dm and. the dm and combines the problematic features of multifunctionality and vagueness or underspecification, a situation which, coupled with its high frequency in speech (the most frequent english dm in disfren), makes it very complex to annotate reliably. and is quite pervasive in writing as well though not to the same extent as in speech: and takes up 26% of all dm occurrences in disfren, as opposed to only 16% in the pdtb corpus. table 1 compares the most frequent functions of and (more than 10 occurrences) in disfren and in the pdtb corpus and their proportions over the total number of and tokens. we see that the functional range of and in the spoken corpus is much larger and more balanced than in writing, where only four different senses show more than 10 occurrences with an overwhelming majority of the basic conjunction relation. although the two corpora have been annotated with different taxonomies in potentially different degrees of granularity, both spoken and written and share a number of common functions: addition roughly corresponds to conjunction, enumeration to list and consequence to result. other functions expressed by and in speech are included in the pdtb taxonomy but were not assigned to and in the written corpus, namely specification, temporal and contrast. nevertheless, it is obvious that not all functions in the spoken corpus can be reduced to these core relational meanings: at least the structural functions of topic-shift and opening boundary must be differentiated from the basic meaning of and as a conjunction at sentence level (and its underspecified, contextually enriched uses such as contrast or temporality). while this latter group of structural functions is not part of the pdtb taxonomy, it remains that the majority of the labels attributed to and in speech (6 out of 11) are included in both annotation models, and yet are not distributed in the same way across the two modes of communication: more balanced proportions across more different function types in speech, overwhelming monopoly of one sense label (namely conjunction) in writing. discourse markers in speech 159 disfren (speech): 1140 and pdtb (writing): 3000 and addition (57.11%) conjunction (90.76%) specification (15.79%) list (7%) consequence (8.86%) result (1.26%) topic-shift (3.6%) juxtaposition (0.36%) temporal (2.37%) punctuation (2.11%) conclusion (1.75%) topic-resuming (1.4%) contrast (1.14%) opening boundary (1.14%) enumeration (0.88%) table 1. most frequent functions of and in speech vs. writing in other words, these corpus-based results show that dms tend to fulfil a wider functional range in speech than in writing, even for those expressions which do occur in both modalities and which share a common relational core such as and. besides the inclusion of more different types of functions, this difference also suggests a greater complexity of spoken dm annotation where the annotators have to choose among many different possible senses for a single dm, including highly underspecified uses. 5.2.3 one marker with simultaneous functions in the same context the third level of multifunctionality refers to the tendency of dms to perform more than one function at a time (see petukhova and bunt, 2009). it is generally agreed (e.g., brinton, 1996: 35) that it is not always possible nor relevant to rank by order of semantic priming the respective weight of multiple functions in a single occurrence. these simultaneous functions can either belong to the same level or domain (11) or to different ones (12-13). (11) these concerns are especially important (0.270) as we approach the crucial topic of economic and monetary union (en-poli-01) (12) and then there are the really bland ones that i think oh come on (0.130) you know (0.493) friendly (0.960) helpful (en-conv-03) (13) and that means either we’re going to have to increase the amount of tax we pay (0.400) and none of us really wants to do that either (0.560) or we’re going to have to reduce expenditure (en-intf-04) in example (11) “as” exhibits two functions from the propositional domain, namely a causal and a temporal meaning. similarly, “you know” in example (11), expresses both monitoring (modal) and specification (propositional), whereas “either…or” in (12) express alternative (propositional) and enumeration (structural). simultaneous functions of dms are not more frequent in unplanned than in planned discourse: only 4% of all dms in disfren were assigned double labels, an average which is not particularly sensitive to variation in degrees of planning or interactivity (4.57% in conversations vs. 4.56% in political speeches). by contrast, results from written corpora tend to show a high frequency of double senses: for instance, webber et al. (2016) report that 1002 out of 4138 relations between conjoined verb phrases are assigned two or three senses in the pdtb corpus. although this characteristic exists both in speech and writing, it involves different types of functions: in the pdtb, for instance, double senses only correspond to simultaneous discourse relations (e.g. conjunction-synchrony; see also baldridge and lascarides, 2005), whereas the multidimensionality of spoken communication is particularly prone to expressions working on discourse markers in speech 160 different levels or macro-functions at the same time. as bunt (2012) puts it: the phenomenon of multifunctionality can be explained by considering participation in a dialogue as involving multiple activites at the same time, such as making progress in a given task or activity; monitoring attention and understanding; taking turns; managing time, and so on. (bunt 2012: 243) once again, we see that writing-based taxonomies of discourse relations, in particular those excluding topic relations and modal meanings, would fail to account for the full multifunctionality of spoken dms “participating in a network of ideational, interpersonal, and textual relations” (goutsos 1996: 167). 6 co-occurrence of dms dms have been repeatedly shown, especially in speech-based research (e.g. bazzanella, 2001; pons, 2008; cuenca and marín, 2009; dostie, 2013; but see also fraser 2013 on writing), to frequently co-occur with each other. dm co-occurrence is very frequent in spoken language: up to 20% of all dms in disfren were coded as part of a co-occurring string of dms, in combinations as varied as and so, but i mean, because if or and actually. to our knowledge, such information is not made available in discourse-annotated corpora. the pdtb reports on a different yet related phenomenon, which is the coordination of conjoined connectives: only before and after, if and when and when and if are attested for a total of five occurrences expressing combinations of temporal and conditional relations. more varied forms and functions of co-occurrence may well exist in writing, yet further empirical evidence is needed to effectively compare the proportion of co-occurring dms across spoken and written corpora. while comparisons of frequency are limited, differences remain as to the treatment of cooccurring dms in speechvs. writing-based annotation frameworks: in the latter, co-occurring dms are treated as one complex unit or distinct type with only one sense label assigned to them; in the former, however, authors tend to acknowledge varying degrees of fixation, from simply juxtaposed dms (in which case each component expresses a distinct function) to more and more bound combinations expressing one single function. with time and spread use, clustered dms can indeed aggregate and combine into new “complex” dms which are no longer separable and express a meaning that is not computable from their individual components (waltereit, 2007), such as and then and its french equivalent et puis. yet this is not systematically the case and this variation should be accounted for more systematically during dm identification and sense disambiguation. one such proposal is provided by cuenca and marín’s (2009) three-fold distinction between juxtaposition of dms (different functions generally with different scopes), addition of dms (different functions with the same scope) and composition of dms (a single function performed by a complex unit whose components are not completely lexicalized). juxtaposition of markers implies that two or more markers co-occur but do not combine neither syntactically nor semantically. in many cases of co-occurrence, the two (or more) dms remain “simply” juxtaposed, as in example (14), where the meanings and functions of three dms are serially added to each other. (14) so it’s actually a proper increasing function (2.833) ok (1.730) so for example if you wanted to supposing you’re looking at sine x (en-clas-04) in (14), three different cases of connection can be identified, namely, consequence, exemplification and condition. on the other hand, some co-occurrences imply some degree of integration of the components. addition of markers implies two or more markers which combine and have the same scope. they usually act at a local level but still maintain their meanings and functions sufficiently distinct, as in (15). discourse markers in speech 161 (15) he’s the guy who is supposed to have left and he had my papers and so that was the problem over the party (en-conv-06) in this example, “and” indicates its basic meaning of addition, marking the continuity of the following segment, while consecutive “so” specifies that this next segment (“that was the problem”) is a conclusion to the previous context.5 finally, composition of markers takes place when two co-occurring markers jointly contribute to indicating a single discourse function at a global level, as in (16). (16) and she’s feeding a baby (0.200) so uhm (0.200) and then yes of course this is a some sort of love scene going on (en-clas-05) in (16), “and then” acts as a complex marker indicating the continuation of the description of a situation. this case can be differentiated from other cases in which and and then can be attributed distinct functions, generally with then acting as a temporal marker, as in (17). (17) i really found myself enjoying the first ten minutes and then suddenly (0.273) everything seemed to disintegrate (en-intr-01) the most integrated co-occurrences include dms that indicate interactional values while delimiting main units of the narrative: the beginning, the end or a major transition place (turn, sequence and subsequence, enunciation). other examples of this type are english well then when it is used as opening-up closings or okay then indicating a discourse boundary and marking continuation, although there were no occurrences in disfren. complex dms are usually language-specific and even variety-specific (e.g., ou sinon in belgian french, see crible, 2015), although some are present cross-linguistically, such as and then and its french equivalent et puis (also existing in writing). the scale from juxtaposition to composition is in constant evolution, which is why lexicalization criteria are needed to draw the line between different stages of integration, following e.g., himmelmann (2004). in sum, dm co-occurrence is a pervasive and multi-faceted phenomenon in spoken language and should therefore be included in annotation frameworks in a more systematic yet flexible way than what is currently available in writing-based models, where it is either ignored or highly restricted. this difference of treatment in the literature might not necessarily reflect a relative difference of frequency in corpus data. nevertheless, dm co-occurrence appears to be one area where studies on spoken language can enrich annotation models designed for both speech and writing. 7 conclusion and recommendations for corpus annotation some degree of dms specialization depending on the mode of communication (spoken vs. written) can be observed, but planning and interactivity also play a crucial role on the use of dms. some dm configurations are typically oral and are usually related to unplanned and interactive discourse, while others are typical of written texts or, rather, planned discourse, either in written or spoken mode. overall, it appears that markers in speech are associated with more varied structural configurations and multifunctionality at various levels (or “multidimensionality” in petukhova and bunt’s (2011) terms), and thus potentially more complex to annotate. as a consequence, annotation instructions should be adapted/formulated according to the characteristics of spoken dms in a systematic and comprehensive way, taking into account a number of relative tendencies which we have identified, quantified and illustrated. the structure of speech is linearly intricate, lexically vaguer and includes more repetition than written discourse. interactivity and low planning in dialogal genres are prone to repair, turn-taking 5 the recurrence of and so in disfren (26 occurrences) can be a telling sign of its increasing fixation as a new combined unit in language. discourse markers in speech 162 and overlapping. this implies the presence of truncated or independent clauses as well as longdistance relationships which pose a challenge to the interpretation of meaning and thus to annotation. in addition, dms often seem to perform vague functions. structural and modal functions (and combinations thereof) must be correctly identified and included in annotation taxonomies. as a conclusion, we suggest the following recommendations for spoken corpus annotation to handle the problematic cases identified: a) the absence of a textually expressed arg1 or arg2 must be considered as a structural possibility. it covers both non-relational dms (that is, applying to one segment only) or relational dms applying to at least one implicit segment such as a pragmatically available assumption or extra-linguistic context. b) longer excerpts must be taken into consideration to identify the scope of the dm. simplified paraphrases can be used to identify the core structure. however, this is challenging to annotate consistently, and explicitly identifying the units under a dm’s scope may be too ambitious, especially in long-distance relations (for both spoken and written data). we argue that sense disambiguation is informative and complex enough and should not necessarily be combined with an identification of the related segments. c) most dms can act on different levels or domains, that is, they can implement propositional and non-propositional meanings. the “one form-one function” principle seldom applies. structural and modal functions should be incorporated in general dm taxonomies. specific disambiguating criteria should be established for operational annotation. d) double labels should be allowed. however, they must be restricted to true simultaneous multifunctionality and not used in case of hesitation between two senses. whether or not double senses (e.g. causal-temporal) constitute new “complex” types of functions should be left to the analyst’s choice depending on the method and research question. e) given the complexity of discourse annotation in spoken data, a tracking system to document the annotator’s confidence during the process would be useful in order to retrieve ambiguous cases, and potentially discard them or at least distinguish them in the analysis. f) dm co-occurrences should be differentiated and systematically annotated. various degrees of integration must be identified and taken into account. in conclusion, by identifying characteristics of dms which are more frequent (if not only existing) in speech and by suggesting annotation recommendations, the research presented in this paper aims to make a contribution to a better description of the use of dms in speech, and also to the development of annotation systems that can deal not only with written but also with spoken data. this would allow us to build inter-operable databases compiling evidence from multiple data types, thus considerably furthering our knowledge on the use and functions of this complex pragmatic category.6 acknowledgements the first author benefits from the support of the arc-project “a multi-modal approach to fluency and disfluency markers” granted by the fédération wallonie-bruxelles (12/17-044). the second author is part of the grampint project (ffi2014-56258-p) supported by the spanish ministry of 6 this research is in line with the is1312 cost action “textlink: structuring discourse in multilingual europe” (http://textlink.ii.metu.edu.tr/, chair prof. liesbeth degand), whose main aim is to map different annotation models into an integrated interface which bridges differences between languages and theoretical frameworks. both authors are members of this cost action. http://textlink.ii.metu.edu.tr/ discourse markers in speech 163 education and science. we thank the three anonymous reviewers and the editors for their insightful comments. references jason baldridge and alex lascarides (2005). annotating discourse structure for robust semantic interpretation. in proceedings of the 6th international workshop on computational semantics, pages 17-29, the university of tilburg, tilburg. carla bazzanella (2001). segnali discorsivi nel parlato e nello scritto. in m. dardano, a. pelo, a. stefinlongo (eds.), scritto e parlato. metodi, testi e contesti, pages 79-97, aracne, roma. douglas biber (2006). university language: a corpus-based study of spoken and written registers. john benjamins, philadelphia, pa. laurel brinton (1996). pragmatic markers in english. grammaticalization and discourse functions. mouton de gruyter, new york, ny. harry bunt (2012). multifunctionality in dialogue. computer speech and language 25:222-245. lieven buysse (2014). “we went to the restroom or something”. general extenders and stuff in the speech of dutch learners of english. in j. romero-trillo (ed.), yearbook of corpus linguistics and pragmatics: new empirical and theoretical paradigms, pages 213-237, springer, berlin. josep m. castellà (2004). oralitat i escriptura. dues cares de la complexitat del llenguatge. publicacions de l’abadia de montserrat, barcelona. william chafe (1982). integration and involvement in speaking, writing and oral literature. in d. tannen and r. freedle (eds.), spoken and written language, pages 83-113, academic press, new york, ny. federica ciabarri (2013). italian reformulation markers: a study on spoken and written language. in c. bolly and l. degand (eds), across the line of speech and writing variation, pages 113127, presses universitaires de louvain, louvain-la-neuve. ludivine crible (2015). grammaticalisation du marqueur discursif complexe ou sinon dans le corpus de sms belge: spécificités sémantiques, graphiques et diatopiques. le discours et la langue 7(1):181-200. ludivine crible (2017a). towards an operational category of discourse markers: a definition and its model. in c. fedriani and a. sanso (eds), discourse markers, pragmatics markers and modal particles: new perspectives, pages 101-126, john benjamins, amsterdam. ludivine crible (2017b). discourse markers and (dis)fluencies in english and french: variation and combination in the disfren corpus. international journal of corpus linguistics 22(2): 242269. ludivine crible and liesbeth degand. reliability vs. granularity in discourse annotation: what is the trade-off? corpus linguistics and linguistic theory, forthcoming. ludivine crible and sandrine zufferey (2015). using a unified taxonomy to annotate discourse markers in speech and writing. in h. bunt (ed.), proceedings of the 11th joint acl-iso workshop on interoperable semantic annotation (isa-11), april 14th, london, uk, pages 1422. maria-josep cuenca (2013a). the fuzzy boundaries between modal and discourse marking. in l. degand, b. cornillie and p. pietrandrea (eds), discourse markers and modal particles. description and categorization, pages 191-216, john benjamins, amsterdam. maria-josep cuenca (2013b). causal constructions in speech. in c. bolly and l. degand (eds.), text-structuring. across the line of speech and writing variation, pages 17-31, presses universitaires de louvain, louvain-la-neuve. maria-josep cuenca and maria-josep marín (2009). co-occurrence of discourse markers in catalan and spanish oral narrative. journal of pragmatics 41:899-914. discourse markers in speech 164 maria-josep cuenca and manuela romano (2013). discourse markers, structure and emotionality in oral narratives. narrative inquiry 23:344-370. liesbeth degand (2014). “so very fast, very fast then” discourse markers at left and right periphery in spoken french. in k. beeching and u. detges (eds), the role of the left and right periphery in semantic change: crosslinguistic investigations of language and language change, pages 151-178, brill, leiden. liesbeth degand and anne-marie simon-vandenbergen (2011). grammaticalization and (inter-) subjectification of discourse markers. linguistics 49:287-294. gaëtane dostie (2013). les associations de marqueurs discursifs de la cooccurrence libre à la collocation. linguistik online 62(5). maria estellés and salvador pons (2014). absolute initial position. in s. pons (ed.), discourse segmentation in romance languages, pages 121-155, john benjamins, amsterdam. bruce fraser (2013). combinations of contrastive discourse markers in english. international review of pragmatics 5:318-340. vanessa wei feng, ziheng lin and graeme hirst (2014). the impact of deep hierarchical discourse structures in the evaluation of text coherence. in proceedings of the 25th international conference on computational linguistics (coling 2014), pages 940-949, dublin city university and association for computational linguistics, dublin. dionysos goutsos (1996). modeling discourse topic. sequential relations and strategies in expository text. ablex, norwoord, nj. michael a. k. halliday (1970). functional diversity in language as seen from a consideration of modality and mood in englih. foundations of language: international journal of language and philosophy 6:322-361. michael a. k. halliday (1987). spoken and written modes of meaning. in r. horowitz and s. j. samuels (eds.), comprehending oral and written language, pages 55-82, academic press, new york, ny. alexander haselow (2012). subjectivity, intersubjectivity and the negotiation of common ground in spoken discourse: final particles in english. language and communication 32:182-204. nikolaus himmelmann (2004). lexicalization and grammaticization: opposite or orthogonal? in w. bisang, n. himmelmann and b. wiemer (eds), what makes grammaticalization? a look from its fringes and its components, pages 21-42, mouton de gruyter, berlin. rosalind horowitz and s. jay samuels (1987). comprehending oral and written language: critical contrasts for literacity and schooling. in r. horowitz and s. j. samuels (eds.), comprehending oral and written language, pages 1-52, academic press, san diego, ca. martin hummel (2012). polifuncionalidad, polisemia y estrategia retórica. los signos discursivos con base atributiva entre oralidad y escritura. acerca de esp. bueno, claro, total, realmente, etc. de gruyter, berlin. peter koch and wulf österreicher (1990). gesprochene sprache in der romania: französisch, italienisch, spanisch. max niemeyer, tübingen. kurt kohn (2012). pedagogic corpora for content and language integrated learning. insights from the backbone project. the eurocall review 20(2). uta lenk (1998). discourse markers and global coherence in conversation. journal of pragmatics 30:245-257. araceli lópez serena and margarita borreguero zuluoga (2010). los marcadores del discurso y la variación lengua hablada vs. lengua escrita. in ó. loureda lamas and e. acín villa (eds.), los estudios de marcadores del discurso en español, hoy, pages 415-496, arco/libros, madrid. william mann and sandra thompson (1988). rhetorical structure theory: toward a functional theory of text organization. text 8(3):243-281. discourse markers in speech 165 maj-britt mosegaard hansen (2006). a dynamic polysemy approach to the lexical semantics of discourse markers (with an exemplary analysis of french toujours). in k. fischer (ed.), approaches to discourse particles, pages 21-41, elsevier, amsterdam. maj-britt mosegaard hansen (2008). particles at the semantics/pragmatics interface: synchronic and diachronic issues. a study with special reference to the french phrasal adverbs. elsevier, oxford. gerald nelson, sean wallis and bas aarts (2002). exploring natural language: working with the british component of the international corpus of english. john benjamins, amsterdam. volha petukhova and harry bunt (2009). towards a multidimensional semantics for discourse markers. in proceedings of the 8th international conference on computational semantics (iwcs-8), pages 157-168, university of tilburg, tilburg. salvador pons (2008). la combinación de marcadores del discurso en la conversación coloquial: interacciones entre posición y función. estudos linguísticos/linguistic studies 2:141-159. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi and bonnie webber (2008). the penn discourse treebank 2.0. in proceedings of lrec, june 2008, marrackech, morroco, pages 2961-2968. gisela redeker (1991). linguistic markers of discourse structure. linguistics 29:1139-1172. rehbein, ines, merel scholman and vera demberg (2016). annotating discourse relations in spoken language: a comparison of the pdtb and ccr frameworks. in proceedings of lrec, may 2016, portoroz, slovenia, may, pages 23-28. riccardi, giuseppe, evgeny a. stepanov and shammur absar chowdhury (2016). discourse connective detection in spoken conversations. in proceedings of 2016 ieee international conference on acoustics, speech and signal processing (icassp), march 2016, shangai, china, pages 6095-6099. ted sanders, wilbert spooren and leo noordman (1992). toward a taxonomy of coherence relations. discourse processes 15:1-35. ted sanders, vera demberg, jacqueline evers-vermeul, jet hoek, merel scholman, sandrine zufferey (2016). unifying dimensions in discourse relations: how various annotation frameworks are related. in proceedings of the second textlink action conference, april 2016, budapest, hungary, pages 110-112. deborah schiffrin (1987). discourse markers. cambridge university press, cambridge. michael schober and susan brennan (2003). processes of interactive spoken discourse: the role of the speaker. in a. graesser, m. gernsbacher and s. goldman (eds), handbook of discourse processes, pages 123-164, lawrence erlabum, hillsdale, nj. amanda stent (2000). rhetorical structure in dialog. in proceedings of the 2nd international natural language conference (inlg’2000), pages 247-252. deborah tannen (1982). the oral/literate continuum in discourse. in d. tannen (ed.), spoken and written language, pages 1-16, ablex, norwood, nj. sara tonelli, giuseppe riccardi, rashmi prasad and aravind joshi (2010). annotation of discourse relations for conversational spoken dialogs. in proceedings of the seventh international conference on language resources and evaluation (lrec 10), pages 2084-2090. david tuggy (1993). ambiguity, polysemy, and vagueness. cognitive linguistics 4(3):273-290. christoph unger (1996). the scope of discourse connectives: implications for discourse organization. journal of linguistics 32:403-438. diane vincent (1993). les ponctuants de la langue et autres mots du discours. nuit blanche éditeur, québec. richard waltereit (2007). à propos de la genèse diachronique des combinaisons de marqueurs. l’exemple de bon ben et enfin bref. langue française 154:94-128. discourse markers in speech 166 bonnie webber, rashmi prasad, alan lee and aravind joshi (2016). discourse annotation of conjoined vps. in conference handbook of the second action conference of textlink, april 2016, budapest, hungary, pages 135-140. journal of machine learning research-microsoft word template dialogue & discourse 10(2) 79-104 doi: 10.5087/dad. 2019.204 ©2019 juliane burmester, katharina spalek and isabell wartenburger this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). visual attention-capture cue in depicted scenes fails to modulate online sentence processing juliane burmester juliane.burmester@uni-potsdam.de empirical childhood research university of potsdam karl-liebknecht-straße 24-25 14476 potsdam, germany katharina spalek katharina.spalek@hu-berlin.de department of german language and linguistics humboldt-universität zu berlin unter den linden 6 10099 berlin, germany isabell wartenburger isabell.wartenburger@uni-potsdam.de cognitive sciences, dept. linguistics university of potsdam karl-liebknecht-straße 24-25 14476 potsdam, germany editor: vera demberg submitted 06/2019; accepted 12/2019; published online 12/2019 abstract everyday communication is enriched by the visual environment that listeners concomitantly link to the linguistic input. if and when visual cues are integrated into the mental meaning representation of the communicative setting, is still unclear. in our earlier findings, the integration of linguistic cues (i.e., topic-hood of a discourse referent) reduced discourse updating costs of the mental representation as indicated by reduced sentence-initial processing costs of the noncanonical word order in german. in the present study we aimed to replicate our earlier findings by replacing the linguistic cue by a visual attention-capture cue that directs participants’ attention to a depicted referent but is presented below the threshold of conscious perception. while this type of cue has previously been shown to modulate word order preferences in sentence production, we found no effects on sentence comprehension. we discuss possible theory-based reasons for the null effect of the implicit visual cue as well as methodological caveats and issues that should be considered in future research on multimodal meaning integration. keywords: linguistic vs. visual salience, accessibility, discourse processing, erp 1 introduction everyday communication is multimodal, comprising linguistic as well as extra-linguistic (e.g., visual) information. a growing branch of psycholinguistic research highlights effects of extralinguistic cues (e.g., eye gaze or gestures) and the visual environment on language processing (e.g., crocker, knoeferle, & mayberry, 2010; nappa & arnold, 2014; spevack, falandays, burmester, spalek and wartenburger 80 batzloff, & spivey, 2018; staudte, crocker, heloir, & kipp, 2014). by contrast, traditional models of sentence comprehension do not explicitly account for the role of visual attention during the comprehension process (e.g., bornkessel & schlesewsky, 2006; friederici, 2002; marslenwilson & tyler, 1980). however, discourse models (or situational/mental models) go beyond sentence-level processing. discourse models propose that during communication, interlocutors build a non-linguistic mental representation of relevant discourse referents and events based on the incoming linguistic and visual perceptual input amongst multiple other factors (e.g., bower & morrow, 1990; gernsbacher, 1991; grosz & sidner, 1986; johnson-laird, 1980; van dijk & kintsch, 1983; zwaan, 2004; zwaan & radvansky, 1998). therein, referents of high attentional state are assumed to be mentally represented with a higher degree of mental accessibility and/or a higher activation level compared to referents of low attentional state (e.g., arnold, 2010; arnold & lao, 2015; givón, 1988; gundel, hedberg, & zacharski, 1993). we will refer to those referents of high attentional state as being more salient. for the present study, we differentiate between linguistic salience which is verbally induced by, for instance, subject-hood or topic-hood of a referent vs. visual salience which is induced by, for instance, exogenous visual cues to a depicted referent. exogenous visual cues initiate a reflexive attention shift of the addressee to the location of a stimulus. in our study we used the terms implicit vs. explicit visual cues to enable a more precise distinction of exogenous visual cues, which were used in previous studies (analogous to the distinction by myachykov, thompson, garrod, & scheepers, 2012, p. 3). implicit visual cues are presented below the threshold of perception (i.e., subconsciously). explicit visual cues are presented above the threshold of perception (i.e., consciously) (for an overview about neuronal modulations by stimulus-driven (i.e., sensory cue-based) visual attention mechanisms, see corbetta & shulman, 2002). with the present study, we aim to test if visual salience induced by an implicit visual attention-capture cue to a depicted referent impacts online sentence-initial processing in a similar way as it has previously been shown for linguistic salience (burmester, spalek, & wartenburger, 2014). hence, we raise the underlying question if the accessibility degree of mentally represented discourse referents is affected by this type of implicit visual cue or if this is limited to linguistic cues. in the linguistic domain, information structure is used to make certain entities of the discourse more salient. for instance, topic or aboutness topic is an information structural concept describing the entity (e.g., a referent) the sentence is about, that is, topic is attributed to that part of information about which the speaker intends to increase the listener’s knowledge (gundel, 1985; reinhart, 1981). hence, topic is ascribed not solely a formal linguistic but also a cognitive concept that activates the listener’s mental representation at the beginning of a sentence (portner, 2007). in the majority of languages, salient information ‒in terms of the grammatical subject and/or topic of the sentence‒ dominantly occupies the sentence-initial position, because subjects and topics own a higher accessibility degree compared to their complements, that is, objects and comments (e.g., bock & warren, 1985; dryer, 2013; tomlin, 1995). german is a language with a strong subject-first preference (e.g., hemforth, 1993; weber & müller, 2004): the canonical word order in german main clauses is subject-verb-object (so) (see example sentence (1)). morphological case marking at the respective noun phrases enables the identification of the grammatical function of subject (via nominative case (nom)) and object (via accusative case (acc)) for masculine nouns. (note that in the example sentences (1) and (2), the nouns “wal” [whale] and “hai” [shark] lack overt case affixes, while the determiners are overtly case marked, which nevertheless allows the unequivocal identification of subject and object.) (1) so: der wal streichelt den hai. [the[nom] whale[nom]]subject [strokes]verb [the[acc] shark[acc]]object. the whale strokes the shark. visual cue fails to modulate sentence processing 81 (2) os: den hai streichelt der wal. [the[acc] shark[acc]]object [strokes]verb [the[nom] whale[nom]]subject. the shark, the whale strokes. despite the strong subject-first preference in german, information structural characteristics allow reordering of sentential constituents such that the object can precede the subject (see example sentence (2) for a non-canonical object-verb-subject (os) main clause). however, os sentences in german are much less frequent than so sentences (e.g., bader & häussler, 2010) and need a suitable context which increases the salience of the sentence-initial object. in our previous work, for instance, linguistic salience in short, fictitious stories of two animals was induced by a topic question (i.e., “what about the shark?”), which revealed one of two previously mentioned (i.e., discourse given) referents as the topic of the scene (burmester et al., 2014). compared to a neutral cue not indicating topic-hood but a wide focus (i.e., “what exactly is going on?”), subsequent online sentence-initial processing of os sentences is eased. this facilitating impact of linguistic salience (i.e., topic-hood) is reflected in the event-related potentials (erps) in the form of a sentence-initial late positivity, which is attributed to reduced discourse updating costs (e.g., schumacher & hung, 2012). in 2018 we directly compared linguistic and visual salience cues (burmester, sauermann, spalek, & wartenburger, 2018): visual salience induced via an explicit gaze-shift of a virtual person to a depicted referent speeds sentence-initial reading times of german so and os sentences similar to linguistic salience induced via a topic cue. hence, the sentence-initial processing ease was evident 1) independent of whether salience was induced linguistically or visually compared to a preceding neutral cue, and 2) independent of whether the salient referent is mentioned as the sentence-initial subject or object (burmester et al., 2018). this is line with other studies supporting the view that utterance comprehension is facilitated when the speaker’s gaze increases the visual salience of depicted referents (e.g., hanna & brennan, 2007; knoeferle & kreysa, 2012; staudte & crocker, 2011). however, not only speakers’ eye gaze, which provides explicit information about referential intentions (henceforth: intentional information), influences utterance comprehension, but also various other visual salience cues. staudte et al. (2014) showed that listeners benefit from an explicit (non-gaze) arrow cue (henceforth: attentional information) during utterance comprehension similar to eye gaze. both the arrow and the gaze cue effectively direct listeners’ visual attention to a depicted object, to finally anticipate this salient object for an upcoming verbal reference (staudte et al., 2014). arnold and lao (2015) showed that another abstract type of visual attentional cue (i.e., a black rectangle with a size of approximately 1.0° x 1.0° of visual angle1 presented for 200 ms at the target referent’s location) together with the position of the referent in the visual display manipulates listeners’ trial-initial attention in depicted scenes. still, when listeners interpret a subsequent pronoun, their trial-initial attention only secondarily influences which antecedent they select as the most accessible referent in discourse. instead, pronoun interpretation is primarily driven by the linguistic cue of sentenceinitial mention. overall, such studies provide evidence that explicit visual attentional cues affect sentenceand discourse-level processes, although to a different extend than linguistic cues. evidence in favour of the impact of implicit visual cues comes from language production studies. here, implicit similar to explicit visual cues effectively manipulate speakers’ attention to referents in a depicted scene. this manipulation of the speakers’ attention is reflected in the 1 note that in order to establish comparability of visual cues published in previous research, we calculated the visual angle of the cues used by arnold and lao (2015) and myachykov et al. (2012) post hoc as these studies did not report the visual angle. calculations of the visual angle account for the visual cue’s size and distance from participants’ eyes. for instance, for arnold and lao (2015) the calculation was based on a screen distance of 650 mm (22 34 inches reported), screen size width of 390 mm, screen resolution width of 1280 pixels, and cue size of 38 pixels: 𝑉𝑖𝑠𝑢𝑎𝑙 𝑎𝑛𝑔𝑙𝑒 = 𝑑𝑒𝑔𝑟𝑒𝑒𝑠(𝑎𝑟𝑐𝑡𝑎𝑛 (((38/2)/(650 ∗ 1280/390)))) ∗ 2 = 1.0° burmester, spalek and wartenburger 82 sentential structure they choose in picture descriptions: gleitman, january, nappa, and trueswell (2007) used an implicit visual cue by means of a black rectangle with a size of approximately 0.5° x 0.5° of visual angle presented for about 60 75 ms at the location of one of two subsequently depicted referents. other sentence production studies used explicit visual cues by means of a black arrow (tomlin, 1995), red dot or referent preview (e.g., myachykov et al., 2012; turner & rommetveit, 1968) followed by the presentation of referents that are performing a simple transitive action. as a result, these implicitly or explicitly cued referents are more salient or accessible2 than other, uncued referents as reflected in a greater likelihood of salient referents being mentioned sentence-initially as the grammatical subject and/or topic of the sentence (e.g., arnold, 1998, 2010; tomlin, 1997). both cue types even lead to production of otherwise disfavoured linguistic structures (in english). for instance, in cases where the patient of the transitive action is cued, speakers produce the less frequent passive voice with salient referents (i.e., the patient) in sentence-initial position. while production and/or eye-tracking data indicate shifts in the addressee’s attention, erps allow us to investigate whether and when during the course of sentence processing increased effort is needed. numerous erp studies have provided insights into underlying discourse-level mechanisms elicited by different types of linguistic cues during online sentence processing (e.g., bornkessel, schlesewsky, & friederici, 2003; burkhardt, 2006; burmester et al., 2014; kaan, dallas, & barkley, 2007). based on specific neural correlates, the syntax-discourse model (schumacher & hung, 2012), as an instance of a neurocognitive account of discourse processing, specifies two temporally distinct processing mechanisms of meaning computation: discourse linking (n400) and discourse updating (late positivity). in burmester et al. (2014), the facilitative impact of the linguistic salience cue (i.e., topic-hood) elicits a reduced late positivity around 500 700 ms time-locked to the sentence-initial position of os sentences, but not of so sentences. in line with the assumptions of the syntax-discourse model, the reduced late positivity in the non-canonical os word order is attributed to reduced processing costs for updating the current discourse model following the linguistic topic cue compared to the neutral cue. this interpretation of the late positivity as an index for integration and updating processes of mental representations is further supported by recent studies (e.g., delogu, drenhaus, & crocker, 2018; or within the neurocomputational model of language comprehension by brouwer, crocker, venhuizen, & hoeks, 2017). however, the assumptions of the syntax-discourse model as well as of other discourse models (hagoort & van berkum, 2007) go beyond the impact of purely sentential context on meaning computation, but include situational context information. even more explicitly, the coordinated interplay account (crocker et al., 2010) highlights the role of visual attention for listeners’ mental representations. this account assumes closely temporally synchronized stages of visual and linguistic information processing during sentence comprehension as supported by multiple “visual world” eye-tracking studies (e.g., knoeferle & kreysa, 2012) and also erp studies (e.g., knoeferle, habets, crocker, & münte, 2007). for instance, visual cues reduced online processing costs of os sentences: facilitating cues included explicit, intentional, speech-aligned (beat) gesture cues indicating a specific sentence part as salient (holle et al., 2012), or explicit visual presentations of the visually depicted event of the target sentence (knoeferle et al., 2007). to the best of our knowledge it has not been reported so 2 note that in contrast to mental accessibility (e.g., ariel, 1988) which relates to discourse-level processing, these sentence production studies attribute the term accessibility to the lemma/conceptual level which corresponds to the retrieval of a referent’s mental representation from memory or “representing potential referents in thought” (bock & warren, 1985, p. 47). accordingly, the lemma of a more accessible referent is earlier retrieved from memory and hence mapped to a more prominent syntactic role (levelt, 1989). however, we use the term accessibility from both the side of discourse and from the side of sentence production, since for both sides referent-related features such as animacy, linguistic or visual prominence contribute to a referent’s high accessibility degree (e.g., ariel, 1988; prat-sala & branigan, 2000). disentangling different concepts of accessibility is not within the scope of the present study. visual cue fails to modulate sentence processing 83 far how implicit visual cues that purely direct the addressee’s attention to depicted referents impact online sentence processing. using erps for investigating the impact of implicit visual cues on sentence comprehension might contribute to our understanding of the underlying neurophysiological mechanisms during sentence processing which might be comparable to those evoked by linguistic cues. our study aims to answer the question whether ‒parallel to our earlier findings concerning linguistic salience‒ a referent in sentence-initial (i.e., topic) position is easier to process if visual salience is induced via an implicit attention-capture cue. hence, by using an implicit visual cue in the present study we intend to conceptually replicate our earlier erp-findings, that is, the sentenceinitial late positivity modulation evoked by an (explicit) linguistic cue (burmester et al., 2014). the implicit visual cue of the current study was presented for 66 ms analogously to the gleitman et al. (2007) study in which a similar type of cue significantly manipulated speakers’ attention in depicted scenes, and hence, modulated what speakers mentioned first during sentence production. in accordance with the earlier findings concerning linguistic salience, we predict that visual salience of a depicted referent induces modulations of the late positivity at sentence-initial position of subsequent os sentences. besides the late positivity, the linguistic topic cue in burmester et al. (2014) elicited an early perceptual repetition effect due to word repetition in the topic but not in the neutral condition. this effect was reflected in a reduced early positivity around 200 ms at sentence-initial position of both so and os sentences. in the present visual cueing paradigm, no word repetition occurs. therefore, we do not expect any modulations of this early positivity. in addition, the burmester et al. (2014) study revealed a word order effect in terms of generally greater processing costs for os than so sentences that we expect to replicate in the present study. 2 materials and methods 2.1 participants thirty-one native speakers of german participated after giving informed consent. except for one participant, participants were right-handed as assessed by a german version of the edinburgh handedness inventory (oldfield, 1971). all had normal or corrected-to-normal vision and had no reported neurological disorder. participants were reimbursed or received course credits for participation. data of two participants were excluded from further analysis, that is, one participant due to left handedness, and one participant due to a technical error during recording the electroencephalogram (eeg). the analysed group consisted of 29 participants (15 female, mean age 24.8 years, age range 19.4 25.2 years). 2.2 design and material in the present study (analogous to burmester et al., 2014) participants were presented with short stories of two animals that were going to perform a fictitious transitive action (e.g., a whale and a shark, one of which is going to stroke the other) while an eeg was recorded to investigate erps during online sentence processing. in contrast to burmester et al. (2014), stories were additionally depicted by pictures of the two animals and the action instrument (cf. figure 1). the study used a 2 x 2 within-subject design with the fully crossed factors cue (topic vs. neutral) and word order (so vs. os sentences), resulting in four conditions: topic so, neutral so, topic os, neutral os. a total of 160 different stories (40 per condition) was created based on coloured pictures of 40 animals (monomorphemic nouns of masculine gender which were 1-syllabic (n = 18) or 2-syllabic (n = 22)) and 10 actions (which were monomorphemic, 2-syllabic transitive, and accusative-assigning verbs). for 90% of nouns, nom and acc case were overtly marked only at the determiner. in the remaining 10 % of nouns, nom case was overtly marked only at the determiner, but acc case was overtly marked at the determiner and noun (e.g., “den löwen” [the[acc] lion[acc]]object). nouns and verbs were controlled for normalised written lemma burmester, spalek and wartenburger 84 and type frequency values according to the dlex database (heister, würzner, bubenzer et al., 2011). moreover, other semantic and discourse factors such as animacy and discourse-givenness of sentential arguments impact referent accessibility and hence the ordering principles at the sentence-level (e.g., clark & clark, 1977; grewe, bornkessel, zysset et al., 2006). we controlled for these factors by exclusively choosing animate referents that were explicitly mentioned in the lead-in sentence. each trial started with a red fixation cross signalling the beginning of a new story. afterwards a blank screen for 500 ms was followed by a phrase-wise presented lead-in sentence (see figure 1 (1)) introducing the two relevant animals of the scene, the action instrument and a corresponding prepositional phrase (e.g., the place where the animals were finding the action instrument). with regard to information structure, the lead-in revealed both animals as discourse-given (prince, 1981) and the action as inferable based on the mentioned instrument (prince, 1992). the following visual context (2) consisted of the implicit visual attention-capture cue located at one of three picture positions: (i) the upper left or (ii) upper right animal (i.e., topic cue, respectively), (iii) the bottom centre position of the action instrument (i.e., neutral cue). the visual cue was presented in the form of a black square with approximately 0.3° x 0.3° of visual angle presented against a light greyish background colour for a duration of 66 ms. the cue was immediately followed by the pictures at the respective positions similar to previous visual cueing paradigms (gleitman et al., 2007; myachykov et al., 2012). hence, either one of the two animals was cued in order to direct participants’ attention to the topic referent (topic cue), or the action instrument was cued in order to direct participants’ attention to a wider scope of the scene, the to-be-performed action of the two animals (neutral cue). the coloured pictures of the animals and actions were presented with 9.3° x 9.3° of visual angle. after a black fixation cross, the target sentence (3) was presented phrase-wise either in so or os word order describing the thematic role relations of the depicted animals (i.e., who is performing the action with whom), followed by a blank screen for 200 ms. the target sentence consisted of a first determiner phrase (dp1), verb, second determiner phrase (dp2), and a prepositional phrase specifying the animals’ location or the action instrument. dp1 was either both subject and agent of the action or object and patient/undergoer of the action. dp2 always carried the inverse syntactic and thematic role of dp1. note here that in our study syntactic and thematic role always coincided. with the closing prepositional phrase (e.g., “with the brush”) we aimed to prevent that processing of dp2 is contaminated by “wrap up” effects typically occurring at the end of sentences (e.g., just & carpenter, 1980). the phrase-wise presentation durations were chosen in analogy to previous studies (bornkessel et al., 2003): dps and prepositional phrases presented for 500 ms, respectively; conjunctions, auxiliary verbs, and main verbs presented for 450 ms, respectively; with a 100 ms interstimulus interval. in 20 % of trials a sentence-picture-verification task (4) probed participants’ attentive reading of the stories. for this, 32 pictures (eight per condition) depicting the content of the preceding target sentence were created -half with correct, half with exchanged (i.e., incorrect) thematic role assignments (e.g., shark stroking whale vs. whale stroking shark). pictures of the sentence-picture-verification were presented for 2 s before participants had to press the corresponding button within a 2 s time window. the verification task was followed by a blank screen for 500 ms. the experimental items were identical to the ones in burmester et al. (2014) except that the lead-in sentence was not presented in a self-paced-reading manner but automatically and that the linguistic context question was replaced by the presentation of the visual context. due to the fictive character of the stories as in children’s books, both animals could be plausible agents or patients of the action. in the visual context, animals of a story were always facing each other. in the target sentence, animals equally often occurred as the agent or patient of an action. animals were distributed equally across conditions and were always performing the action with a different animal. introducing the animal first or second in the lead-in sentence as well as the presentation of the animals on the left or right side of the screen was counterbalanced visual cue fails to modulate sentence processing 85 across conditions. to avoid possible effects of structural priming (e.g., scheepers & crocker, 2004), trials were presented in pseudo-randomized order with maximally two consecutive trials of the same condition and word order in the target sentence. preferences of thematic role assignment or topic continuity due to preceding trials were minimized by at least five intermediate trials before an animal was repeated. there were four lists of 160 trials each. lists were created such that within each list each item (i.e., animal pair and action) occurred once and across the four lists each item occurred in each condition. we presented each participant with one of these lists of 160 trials. these lists did not include any filler trials to arrange the experimental session in an appropriate time frame for participants’ motivation and concentration ability, and hence, minimize artefacts and alpha waves in the eeg signal. 2.3 procedure participants were tested individually, seated in a sound-attenuated booth with 80 cm distance to a computer screen (1680 x 1050 pixels screen resolution). after the preparation for eeg recording, participants were visually presented (on screen) with all pictures used in the subsequent experiment with their corresponding word forms to become familiar with the pictures, that is, the 40 animals and 10 actions. afterwards participants received a written instruction in which they were asked to read each story attentively and silently and to answer the sentencepicture-verification task after some of the stories as accurately and fast as possible. participants were asked to sit relaxed, and to avoid eye-movements, blinks, and other muscle movements. participants had a button box (cedrus® response pad model rb-830) on their lap and performed three practice trials to become familiar with the procedure. to answer the sentence-pictureverification task, the green and red response button (according to correct vs. incorrect pictures) were assigned to the right fore and middle finger (which was counterbalanced across participants). participants were instructed that they will be presented with a new story as soon as they see a red fixation cross and they press the yellow button of the button box on which they should put their left thumb throughout the whole experiment. the experiment was visually presented by means of the presentation® software (version 14.1; www.neurobs.com). the whole experiment included pauses after each 40 trials and lasted approximately 30 minutes. in a postexperiment questionnaire, participants were asked if they have an idea about the purpose of the study and if they noticed anything in the course of the experiment, for instance, any cues or disturbances during picture presentation. burmester, spalek and wartenburger 86 figure 1. experimental design of sample trial with approximate english translation written in grey. (1) participants read the phrase-wise presented lead-in sentence followed by a blank screen (200 ms), fixation cross (1000 ms), and blank screen (200 ms). (2) the implicit visual cue presented for 66 ms (topic vs. neutral cue) was directly followed by the pictures. after another fixation cross, (3) the so or os word order target sentence was presented phrase-wise. in the topic cue condition, the so or os sentence mentioned the cued referent (i.e., whale) first (i.e., as dp1). (4) in 20 % of trials participants had to answer a sentence-picture-verification task afterwards. abbreviations: so = subject-verb-object, os = object-verb-subject, dp = determiner phrase. visual cue fails to modulate sentence processing 87 2.4 eeg recording the eeg was recorded using a 32 channel active electrode system (brain products, gilching, germany) with a sampling rate of 1000 hz. the electrodes fixed at scalp by means of a soft cap included the following 29 scalp sites according to the international 10-20 system (american electroencephalographic society, 2006): f7/8, f5/6, f3/4, fc3/4, c5/6, c3/4, cp5/6, p3/4, p7/8, po3/4, fpz, afz, fz, fcz, cz, cpz, pz, poz, oz. the electrooculogram (eog) was monitored by electrodes above (position fp2) and below the right eye. the ground electrode was placed at fp1. impedances were kept below 5 kohm. the left mastoid served as the reference electrode online, whereas recording was re-referenced to bilateral mastoids offline. 2.5 erp data analysis for erp data analysis, the brain vision analyzer software (version 2.1, brain products gilching, germany) was used. to exclude slow signal drifts and muscle artifacts from the eeg raw data the butterworth zero phase filter (low cutoff: 0.3 hz; high cutoff: 70 hz; slope: 12 db/oct) was applied additional to the notch filter of 50 hz. for the correction of artifacts caused by vertical eye movements the algorithm by gratton, coles, and donchin (1983) was applied. we applied an automatic artifact rejection to reject blinks and drifts in the time window of -200 to 2150 ms relative to the onset of the target sentence as well as -200 to 500 ms relative to onset of the visual cue (rejection criteria: max. voltage step of 30 µv/ms, max. 200 µv difference of values in intervals, lowest activity of 0.5 µv in intervals). on average 1.64 % of trials was rejected. erps were averaged for each participant and each condition within a 2150 ms time window time-locked to the onset of the target sentence and within a 500 ms time window time-locked to the onset of the visual cue, with a 200 ms pre-stimulus onset baseline, respectively. for the statistical analysis, ibm spss statistics (version 25.0) was used. the chosen parameters for erp analysis were identical to the ones used by burmester et al. (2014) to maintain comparability with the erp results on the impact of linguistic context information on online sentence processing, in which the same sentences were presented to participants. hence, based on previous psycholinguistic research, we analysed language-related erp components of the target sentence in the following time windows time-locked to the onset of dp1, verb, and dp2, respectively: 100 300 ms (p200), 300 500 ms (n400), 500 700 ms (late positivity). via computation of mean amplitudes of three electrodes, respectively, nine regions of interest (rois, which were identical to the ones in burmester et al., 2014) entered the statistical erp analysis as the fixed factor roi: left frontal (f7, f5, f3), left fronto-central (fc3, c5, c3), left centro-parietal (cp5, p3, p7), right frontal (f8, f6, f4), right fronto-central (fc4, c6, c4), right centro-parietal (cp6, p4, p8), frontal-midline (fpz, afz, fz), central midline (fcz, cz, cpz), parietal midline (pz, poz, oz). for statistical erp analysis, mean amplitude values of erps within each condition were analysed following a hierarchical schema (e.g., bornkessel et al., 2003; burmester et al., 2014). firstly, we computed a fully crossed repeated measures analysis of variance (anova) with the fixed factors cue (topic vs. neutral), word order (so, os), and roi (nine levels, see above) for each of the three time windows time-locked to the onset of dp1, verb, and dp2, respectively. in addition to the analyses of the target sentence, we analysed early erp components (i.e., n1, p2) relative to the onset of the visual cue, henceforth termed cue position (i.e., topic left, topic right, and neutral bottom), in the time window of 100 200 ms as well as of 250 350 ms (see e.g., luck & hillyard, 1994; mangun, 1995 for attention-based early visual processing changes reflected in different early evoked potentials). notably, the pictures followed the cue immediately, hence these time windows started 100 ms and 250 ms after cue onset but also 34 ms and 184 ms after picture onset. we report greenhouse and geisser (1959) corrected fand p-values, the original degrees of freedom (df) in brackets, and the greenhouse and geisser epsilon (ε) factor for non-sphericity burmester, spalek and wartenburger 88 adjustments of the original df according to jennings & wood (1976) (only for f-tests with more than one df in the numerator). statistically significant effects (i.e., p <. 05) involving an interaction with roi were resolved by computing post hoc paired t-tests to reveal the topographical distribution of the effect. we controlled for the type i error due to multiple pairwise t-tests of levels of the fixed effects in the nine rois by adjusting the significance level according to the bonferroni correction. thus, for post hoc t-tests the following bonferroni adjusted p-values (two-tailed) were considered as statistically significant at α = .05: p <. 006 to resolve the word order x roi interaction, and p <. 002 to resolve the cue position x roi interaction. for presentation purposes the displayed erps in figure 2, 3, and b-1 in appendix b are 10 hz low-pass filtered. 2.6 behavioural data analysis for the statistical analysis of the response accuracy of the sentence-picture-verification task, logit mixed models fitted by the laplace approximation were calculated using the lme4 package (bates, mächler, bolker, & walker, 2015) provided by the r environment (version 3.5.1, r core team, 2013). to analyse the binary distributed response accuracy data (correct vs. incorrect) with the logit mixed models, cue, word order, and the interaction of both were defined as fixed effects, and participants and items were defined as random effects. fixed effects were coded as +/‒.5 to resemble the contrast coding of traditional anova analyses. model fitting started with the simple model (i.e., the two fixed effects and their interaction, and participants and items as random intercepts). in a step-wise manner, slope-adjustments were included if they significantly improved the explanatory power of the simpler model without that slope adjustment as revealed by log-likelihood tests (e.g., baayen, 2008). the statistics of the fixed effects of the final models are reported with estimates (b), standard errors (se), zand p-values. 3 results as reported in the post-experiment questionnaire, participants did not notice any manipulation of their visual attention suggesting that the cue was not consciously perceived and hence, truly implicit. in this section, we first describe the erp results with respect to initial and subsequent processing of the target sentence, before reporting the results of an additional analysis of the present erp data following the visual cue together with the published data following the linguistic cue (burmester et al., 2014). secondly, we present the erp results with respect to the onset of the visual cue. thirdly, we report the behavioural results of the probe sentence-pictureverification task. 3.1 erp results of the target sentence figure 2 illustrates the grand average erps at one representative electrode time-locked to the onset of the target sentence (i.e., dp1) followed by the subsequent sentence positions (i.e., verb and dp2) for both cues (topic and neutral) and both word orders (so and os). for grand average erps at selected electrodes of each roi, see figure b-1 in appendix b illustrating the cues within each word order. 3.1.1 erp results of sentence processing following the implicit visual cue for sentence-initial processing, statistical analyses in the time windows of 100 300 ms and 300 500 ms time-locked to the onset of dp1 neither revealed any statistically significant main effects of cue (topic vs. neutral) or word order (so vs. os), nor significant interactions of cue, word order and/or roi [p > .1] (see appendix a for the complete statistical output). the analysis in the following time window of 500 700 ms revealed a statistically significant main effect of word order [f(1, 28) = 6.254, p = .019] and a significant interaction of word order x roi [f(8, 224) = 3.004, p = .029, ε = 0.424], but no statistically significant effects or visual cue fails to modulate sentence processing 89 interactions of the factor cue. separate post hoc analyses for so and os sentences (averaged across the cue conditions) within each roi yielded a statistically significant enhanced positivegoing erp for os sentences compared to so sentences in the left central roi [t(28) = -3.605, p = .001] and in the midline central roi [t(28) = -3.369, p = .002]. figure 2. grand average erps (with baseline correction) at one representative electrode of the midline central roi for the factors cue (topic vs. neutral) and word order (so vs. os) time-locked to the onset of the target sentence, that is, the first determiner phrase (dp1) followed by the verb and second determiner phrase (dp2): topic so [dotted black] vs. neutral so [dotted grey], topic os [solid black] vs. neutral os [solid grey]. negativity is plotted upwards. grey shades indicate the three time windows of dp1. for subsequent sentence positions, the statistical erp analysis for the time windows 100 300 ms, 300 500 ms, and 500 700 ms post verb onset neither revealed a statistically significant main effect nor interaction involving the factor cue, but revealed statistically significant effects involving the factor word order: in the time window 100 300 ms post verb onset a significant main effect of word order [f(1, 28) = 10.856, p = .003] and a significant interaction of roi x word order [f(8, 224) = 3.359, p = .013, ε = 0.485] was reflected in an enhanced positive-going erp for os compared to so sentences in the left central [t(28) = -3.356, p = .002] and left posterior [t(28) = -4.507, p < .001] roi, the right posterior roi [t(28) = -3.090, p = .004], and the midline central [t(28) = -3.854, p = .001] and midline posterior [t(28) = -4.192, p < .001] roi. in the following time window of 300 500 ms post verb onset a significant interaction of roi x word order [f(8, 224) = 8.101, p < .001, ε = 0.447] was similarly reflected in form of a significantly enhanced positive erp for os compared to so sentences in the left posterior roi [t(28) = -3.448, p = .002] and in the midline posterior roi [t(28) = -4.602, p < .001]. furthermore, in the time window of 500 700 ms post verb onset the significant interaction of roi x word order [f(8, 224) = 6.157, p < .001, ε = 0.558] was reflected in an enhanced positive erp for os compared to so sentences in the midline posterior roi [t(28) = -3.114, p = .004]. statistical analyses of erps post onset of dp2 neither showed any statistically significant main effects, nor interactions of cue or word order in any of the calculated time windows (i.e., 100 300 ms, 300 500 ms, and 500 700 ms) [p < .1]. in summary, the erp results of all three sentence positions (i.e., dp1, verb, dp2) did not show any statistically significant modulation by the preceding visual cue. an impact of the burmester, spalek and wartenburger 90 varying word order with enhanced positive-going erps for os compared to so sentences was evident in multiple time windows time-locked to the onset of dp1 (i.e., 500 700 ms) and verb (i.e., 100 300 ms, 300 500 ms, and 500 700 ms). 3.1.2 erp results compared to sentence-initial processing of the linguistic cue (burmester et al., 2014) since the identical sentence material was used in the present study with the visual cue as in the study with the linguistic cue (burmester et al., 2014), we aimed at directly comparing the impact of the visual vs. linguistic cue modality on sentence-initial processing. for this purpose, we computed additional comparisons of the published erp data following the linguistic cue with the erp data following the visual cue by adding the between-subject factor modality (visual vs. linguistic) while the within-subject-factors cue, word order, and roi maintained as in the original analysis (reported in section 2.5). the results as summarised in appendix c show the following statistically significant effects of modality: in the 100 300 ms time window, the analysis showed a significant interaction of modality x roi [f(1, 46) = 3.190, p = .038] and of modality x word order [f(1, 46) = 5.793, p = .020], which was reflected in a marginally significantly different processing of so vs. os sentences following the linguistic cue [t(18) = 2.092, p = .051], which was absent following the visual cue [t(28) = -4.16, p = .680]. in line with the separate analyses, there were no statistically significant effects of modality in the 300 500 ms time window. for the 500 700 ms time window post onset of dp1, the joint analysis revealed a marginally statistically significant interaction of modality x cue x word order [f(1, 46) = 3.963, p = .052]. this confirms the results of the two separate analyses of the linguistic and visual cue: the presence of a statistically significant interaction of cue x word order x roi [f(8, 144) = 4.15, p < .05] following the linguistic cue and its absence following the visual cue [f(8, 224) = 1.605, p = 0.189; ε = 0.417] (cf. figure 2 and b-1 in appendix b) in the present study and figure 2 and table 3 in burmester et al., 2014). in summary, the visual cue had no impact on sentence processing in the present study. therefore, the visual cue was not effectively increasing the salience of the cued referent (i.e., topic cue). to examine whether the visual cue per se modulated participants’ processing, we computed a further erp analysis time-locked to the onset of the visual cue. if yes, this should be reflected in differential, especially early sensory-evoked potentials in dependence of the cue position on screen (i.e., topic left, topic right, and neutral bottom). 3.2 erp results of the visual cue per se with the following erp analysis we aimed to test if the implicit visual cue modulated participants’ early sensory-evoked potentials (i.e., n1, p2) time-locked to the onset of the visual cue. the visual cue was presented for 66 ms and was directly followed by the pictures. however, since the timing and the position of the pictures were always the same, early processing differences are likely to be related to the different prior cue positions. therefore, erp analyses were calculated to assess the impact of cue position (i.e., topic left, topic right, and neutral bottom) in the n1 (100 to 200 ms) and p2 (250 to 350 ms) time window post onset of the visual cue and its topographical distribution by the factor roi.3 as can be seen in figure 3, the erps time-locked to the visual cue show an early negativity around 100 200 ms followed by a positive deflection around 250 350 ms. statistical analysis in the time window of 100 200 ms revealed a statistically significant difference in the form of a 3 note that for the anova and post hoc t-tests reported in appendix d, we report the statistical analysis based on the grand average erps of all available trials (i.e., 40 trials each for topic left vs. topic right cue positions and 80 trials for the neutral bottom cue position). however, conducting the analysis with a similar number of trials across conditions with random sampling of half of the neutral bottom cue position trials (i.e., 40), revealed a similar pattern of results. visual cue fails to modulate sentence processing 91 main effect of cue position [f(2, 56) = 9.477, p < .001, ε = 0.966] and of an interaction of cue position x roi [f(16, 448) = 7.337, p < .001, ε = 0.351]. post hoc comparisons show that the negativity around 100 200 ms was strongest for the cue at the neutral bottom position compared to the topic left and topic right position. this difference was present in multiple rois (see appendix d for the complete post hoc statistical results of the respective rois). regarding the positivity around 250 350 ms, statistical analyses revealed a statistically significant main effect of cue position [f(2, 56) = 7.628, p < .001, ε = 0.949] and a statistically significant interaction of cue position x roi [f(16, 448) = 3.675, p = .007, ε = 0.255]. post hoc analyses show that the topic left cue elicited a significantly more pronounced positive deflection compared to the neutral bottom cue (p < .001). note, we cannot clearly disentangle the response to the cue and the response to the picture, as the pictures were always presented 66 ms after the cue. the modulation by cue position might therefore either reflect the direct effect of the cues themselves or their impact on the processing of the subsequently presented pictures. in both cases, the results indicate that participants processed the implicit visual cues. but still, given that none of the participants noticed the presence of the cues nor was sentence processing influenced by the cue, we can assume, that the cues were processed only subconsciously. figure 3. grand average erps of one representative electrode for the factor cue position (topic left, topic right, neutral bottom) time-locked to the onset of the implicit visual cue (~66 ms cue duration directly followed by pictures). negativity is plotted upwards. 3.3 behavioural results in the sentence-picture-verification task (in 20 % of trials) participants showed a high response accuracy across conditions indicating that participants were attentive throughout the experiment: topic so: m = 0.91 (se = 0.03), neutral so: m = 0.91 (se = 0.02), topic os: m = 0.82 (se = 0.03), neutral os: m = 0.81 (se = 0.02). logit mixed models analyses of participants’ response accuracy did not reveal any statistically significant differences of the fixed effects cue (b = 0.112, se = 0.222, z = 0.503, p > .1), word order (b = 0.490, se = 0.339, z = 1.443, p > .1), or the interaction of cue x word order (b = 0.226, se = 0.222, z = 1.020, p > .1). 4 discussion we aimed at answering the question if an implicit visual cue to a depicted referent impacts sentence-initial processing similarly to a purely linguistic cue as revealed by erps (burmester et al., 2014). with regard to the linguistic cue, the indication of the aboutness topic referent, which was subsequently mentioned in sentence-initial position, reduced the late positivity during online processing of os sentences. in the present study, the linguistic topic cue was replaced by an implicit visual cue to a depicted referent in a visual scene. however, the impact of the linguistic cue on sentence-initial processing was not replicated. in fact, none of our analyses revealed any burmester, spalek and wartenburger 92 statistically significant effects of the visual cue on sentence processing. the findings of the linguistic cue (burmester et al., 2014) were interpreted within the syntax-discourse model (schumacher, 2014) which highlights the role of context information for meaning computation during sentence comprehension. within this model, the impact of context information (i.e., including sentential and situational context, amongst others), is reflected in modulations of processing costs for discourse linking and updating the hitherto built mental representation of the listener. following this model, the impact of a referent’s topic-hood of the sentence-initial referent in os sentences was attributed to reduced discourse updating costs as –in contrast to a neutral cue– the linguistic topic cue indicated one referent more salient amongst others which rendered this referent more likely to be mentioned sentence-initially. obviously, the visual topic cue used in the present study did not increase the salience of the cued referent to a similar extent and hence did not elicit a facilitative effect on sentence processing in terms of reduced discourse updating costs. however, the long-time neglected role of extra-linguistic (e.g., visual) information in traditional accounts of sentence comprehension (e.g., frazier & fodor, 1978; friederici, 2002; marslen-wilson & tyler, 1980) has been complemented by recent models underlining the close temporal integration of multimodal (e.g., visual and linguistic) information in the listener’s current mental representation (e.g., bower & morrow, 1990; zwaan, 2004). for instance, as supported by neuroscientific methods, a one-step model of sentence comprehension –integrating concomitant information from different modalities within one step– has been suggested (hagoort & van berkum, 2007). more specifically, this model postulates that linguistic (for instance, sentential structure, semantics) and pragmatic information (from prior discourse, extra-linguistic information such as speaker’s gestures or the visual world) is immediately processed by the same brain regions (namely the left inferior frontal gyrus) in order to directly map all information onto a discourse model as the basis for sentence interpretation. similarly, the coordinated interplay account by crocker et al. (2010) explicitly outlines the integration of visual scene information during sentence processing, especially in cases were visual information becomes highly relevant for the interpretation and disambiguation of spoken sentences. by linking to the so called “blank screen paradigm” in which attention shifts during sentence comprehension even occurred when depicted objects were no longer presented (altmann, 2004), crocker et al. (2010) suggest that mental representations of a previously presented scene still influence the comprehension process. however, the weighting of (competing) visual and linguistic information and the role of specific inherent features of visual salience cues on the strength of impact during language comprehension is not specified in any of these models. in the following, we will discuss our findings concerning the absent visual cue impact on sentence processing against the background of previous research using different visual cues while raising possible issues with respect to the present study design. afterwards, we will briefly discuss the word order effect, which we replicated (section 4.2). 4.1 null effect of visual cue on sentence-initial processing against the background of previous research investigating the interaction of visual cues with linguistic processing, we discuss the absent cue effect of the present study 1) with respect to different aspects of informativity of the visual scene for the comprehension process, and 2) with respect to the different impact of visual cue types (implicit vs. explicit) on sentence comprehension and production. with respect to the informativity of the visual scene, some previous studies emphasise the need for relevance of visual information for meaning computation during language processing. for instance, listeners use depicted events –similarly to linguistic cues (i.e., case marking)– for syntactic reanalyses of locally structurally ambiguous german sentences as reflected in a reduced p600 (knoeferle et al., 2007). further evidence from the field of referential processing shows that visual information impacts accessibility of discourse referents only if the linguistic context is visual cue fails to modulate sentence processing 93 moderate, uninformative, or ambiguous (e.g., nappa, wessel, mceldoon et al., 2009; vogels, krahmer, & maes, 2013). with respect to the present study design, the visual scene might not have added crucial information to the comprehension process, for instance, to assign thematic roles in the subsequent target sentence. compared to visual scenes used in production studies (e.g., gleitman et al., 2007; myachykov et al., 2012), visual scenes in the present study did not depict thematic role relations for the following reasons: we aimed at minimizing confounding effects of prominence-related factors known to affect participants’ gaze fixations of depicted transitive events (e.g., agent-directed fixations followed by patient-directed fixations, ganushchak, konopka, & chen, 2017) as well as linear ordering preferences of sentential constituents (e.g., agent precede patient theta roles, jackendoff, 1972). in short, we tested how a single implicit cue indicating the subsequent sentence topic would affect referent accessibility and therefore we eliminated all additional factors that could have masked this single cue. indeed, previous research shows that the more (additional) information is conveyed by the prior discourse –by the linguistic context or by depicted scenes– the more predictable are specific upcoming words as reflected in an immediate ease of sentence processing (e.g., burmester et al., 2018; otten & van berkum, 2008; van berkum, brown, zwitserlood et al., 2005). more specifically, in an experimental design rather similar to the one in the present study, we could show that depicted thematic role information of the salient referent boosted the cue-based ease of sentence processing compared to non-predictable thematic role information (burmester et al., 2018). moreover, for linguistic context information, otten and van berkum (2008) found greater priming effects of the exact discourse message than of the mere word primes. drawing the parallel to our earlier findings, this could speak in favour of predictive processing mechanisms following the linguistic cue explicitly indicating the upcoming sentence-initial topic (burmester et al., 2014), while with the visual cue we rather manipulated accessibility in a way similar to a word prime. hence, maybe further depicted information such as thematic role relations would have activated further semantic features in order to constrain a more precise and coherent discourse context, and finally support predictive processing. moreover, with respect to the informativity of the visual scene, the simultaneous visual presence of multiple referents has been argued to reduce referent accessibility (as revealed by reduced pronoun use in, for instance, arnold & griffin, 2007 or fukumura, van gompel, & pickering, 2010). in the present study, two possible referents were simultaneously depicted, which –parallel to the preceding argument– could have caused a competition of referent accessibility. this competition of accessibility was not the case in our erp study in which the linguistic topic cue exclusively increased the accessibility of one referent while not mentioning the other (i.e., “what about the ‘topic referent’?”; burmester et al., 2014). however, in our reading time study (burmester et al., 2018) multiple referents were presented simultaneously and nevertheless an explicit visual (gaze) cue increased the accessibility of one amongst three depicted referents as reflected in sentence-initial processing ease similar to a linguistic topic cue. but, in contrast to the present study, participants in our reading time study were already familiarised with the visual scene by a multimodal lead-in before the gaze cue was presented. therefore, participants might have had greater attentional capacities at the moment of processing the gaze cue than during processing the implicit cue with the subsequently presented pictures in the here presented study. alternatively, the type of the cue matters, as gaze cues are more explicit and intentional in nature than the here used implicit abstract cue. taking previous studies using different visual cue types into account, we can assume that visual cues indeed modulate meaning computation during sentence comprehension while they seem to differ with respect to their impact on referent accessibility. as just mentioned, one type of visual cues that has been shown to clearly influence language processing, is eye gaze. this cue is a strong social-communicative cue signalling shared-attention of speaker and listener, it is hence an intentional cue. crucial evidence provided by a few studies that compared both attentional and intentional cues supports the importance of the intentional component of visual burmester, spalek and wartenburger 94 cues for listeners: similar to speaker’s gaze or pointing gesture, a visual cue (i.e., black square presented for 50 ms at the location of a possible referent) influences listener’s pronoun interpretation, but only if the listener was previously instructed that this abstract visual cue is intentionally created by the speaker (nappa & arnold, 2014). however, the same visual cue but without the previous instruction of being intentionally created by the speaker does not reveal an impact. analogously to this finding, holle et al. (2012) showed that an intentional, conversational gesture co-occurring with the speaker’s speech facilitates comprehension of german ambiguous so and os sentences: this short hand movement (i.e., beat gesture) emphasising the subject of the sentence reduces additional processing costs for disambiguation towards the os word order as indicated by a reduced p600. in contrast to this gestural cue, listeners do not make use of an explicit visual attentional cue, that is, a moving point. hence, the intentional component of visual cues plays a significant role in meaning computation during the comprehension process. this significance could be explained by assumptions of pickering and garrod’s (2004) interactive alignment model. accordingly, for a successful dialogue, interlocutors are assumed to develop aligned situational models. thus, maybe intentional cues such as eye gaze offer a window into the interlocutors’ mental representations in dialogue that might trigger other attentional mechanisms than those elicited by purely attentional, visual cues. however, the present study is a first step to better understand the impact of implicit attentional visual cues on sentence-initial processing, which demonstrates to us the subtle differences and difficulties of adapting study designs of sentence comprehension (burmester et al., 2014) and sentence production (gleitman et al., 2007). indeed, studying the impact of implicit visual cues on sentence comprehension might engender some caveats weakening the measurable outcome during later sentence processing. for the visual cue of the present study, we used a black square with 0.3° x 0.3° of visual angle against coloured pictures similar to the one by gleitman et al. (2007), who used a slightly bigger cue (i.e., 0.5° x 0.5° of visual angle) against full-colour clip-arts. moreover, myachykov et al. (2012) used a red dot with 0.7° x 0.7° against black-white-line drawings that –similar to the cue by gleitman et al. (2007)– influenced sentence-initial mention during production. so maybe the visual cue in our study was too subliminal and hence, its impact too short-lived to be measurable during subsequent sentence processing. following this line of thought, modulations of participants’ attention by the implicit visual cue might be less strong and less long lasting with respect to their impact on referent accessibility compared to the high accessibility degree indicated by the explicit mention of the topic in the linguistic context. this explanation is in accordance with eye-tracking data reported by arnold and lao (2015): listener’s trial-initial attention to depicted referents was indeed modulated by multiple visual scene-based factors, that is, different visual attentional cues, the order of the depicted referents from left to right, and listener’s idiosyncratic biases. but, the attentional cues themselves did not significantly predict subsequent pronoun interpretation and instead, the linguistic cue of sentence-initial mention was the strongest predictor. in addition, in an experimental task that requires sentence understanding, linguistic cues such as in burmester et al. (2014) might less likely be ignored by the reader than visual cues and picture contexts such as in the present study. hence, it is important to check whether visual cues and depicted scenes are indeed processed by the addressee of the stimuli. however, based on the absent impact of the implicit visual cue on sentence processing in the present study, we suggest that this type of visually induced salience of discourse referents elicits a different, less strong, impact on the accessibility degree of mentally represented discourse referents compared to linguistically induced salience via topic-hood. this train of thought is supported by erp and behavioural studies showing that linguistic stimuli (e.g., sentences, words) more strongly affect the accessibility of entities compared to pictures (e.g., bögels, schriefers, vonk, & chwilla, 2011; fukumura et al., 2010). moreover, a different neural processing of linguistic and visual stimuli has been suggested (e.g., brandon & andrew, 2007; zhang, begleiter, porjesz, & litke, 1997) which, for instance, has been explained by the more efficient visual cue fails to modulate sentence processing 95 semantic access and/or memory retrieval of words compared to pictures (dorjee, devenney, & thierry, 2010). however, words and pictures revealed a very similar time course of the n400 (congruity) effect with differences only in the topographical distribution (ganis, kutas & sereno, 1996). in summary, reasons for the null effect of the implicit cue in our study might, on the one hand, be traced back to a weaker, less long lasting impact on referent accessibility compared to intentional or linguistic cues. on the other hand, it might to some degree be related to methodological differences to production studies in which implicit cues modulated sentenceinitial mention. disentangling these reasons and testing the impact of other types of visual attentional cues is left open for future research. 4.2 word order effect in the present study, word order effects are reflected in sustained positive deflections (at dp1 and verb position) which were more pronounced for os compared to so sentences across multiple rois. hence, we replicated differential processing costs for so and os sentences found in burmester et al. (2014). a body of neurocognitive research concerning the processing of german sentences with varying word order demonstrates increased processing costs for os sentences compared to their canonical (so) counterpart as reflected in different erp components, time windows and across different sentence positions (e.g., holle et al., 2012; knoeferle et al., 2007; matzke, mai, nager et al., 2002; schlesewsky, bornkessel, & frisch, 2003). for this paper, the word order effect is not of primary relevance and we take it just as a “sanity check” of our erpdata. in line with the previous literature, we argue that the processing differences between os and so sentences are engendered by the subject/nominative-first-preference in german leading to increased processing demands for the non-canonical and less frequent os word order. note that in our study the impact of word order might be confounded with the ordering of theta roles, as both grammatical role and theta role coincided in the sentence constituents, that is, subject and agent, object and patient. the increased online processing difficulties were, however, not visible in the behavioural (accuracy) results of the subsequent probe sentence-picture-verifications. in summary, we argue that the replicated word order effect during online sentence processing confirms the validity of the design: the replication of the word order effect shows that we are not dealing with a replication failure per se but rather that the different findings can be traced back specifically to cue modality. 5 conclusion all in all, previous findings speak in favour of the integration of visual and linguistic cues into listeners’ mental representations –although with a different magnitude of both cue modalities. in our study, the implicit visual cue to a depicted referent followed by subsequent sentence-initial mention did not influence online processing of german so and os sentences. hence, the impact of the linguistic topic cue on the identical sentence material could not be replicated with the present study design, although a similar type of cue was effective in previous production studies (gleitman et al., 2007). it therefore remains an open question if comprehension and production are influenced by similar underlying processes. hence, we conclude that the role of visual, purely attention-directing cues for meaning computation during sentence processing needs further clarification. future research needs to shed more light on the role of different visual cues and their interaction with intentional and attentional aspects in guiding information packaging preferences during utterance comprehension in order to disentangle experimental task specific effects. burmester, spalek and wartenburger 96 acknowledgements this work was supported by the german research foundation (dfg) under grant sfb 632 ‘information structure’. we thank franziska machens and tobias busch for assistance in material preparation and data collection as well as jan ries for his help in preparation of the figures. appendix a results of analysis of variance (anovas) of the erps for the different time windows time-locked to the onset of the first determiner phrase (dp1). df f-values 100 300 ms 300 500 ms 500 700 ms cue (topic vs. neutral) cue x roi cue x word order cue x word order x roi word order (so vs. os) word order x roi 1, 28 8, 224 1, 28 8, 224 1, 28 8, 224 0.669 0.275 (ε = 0.252) 1.374 1.760 (ε = 0.411) 0.983 1.872 (ε = 0.433) 0.266 0.575 (ε = 0.337) 0.030 0.564 (ε = 0.399) 1.004 1.425 (ε = 0.378) 1.020 0.978 (ε = 0.285) 0.723 1.605 (ε = 0.417) 6.254* 3.004* (ε = 0.424) note. greenhouse & geisser (1959) corrected significance levels: * p <. 05; ** p <. 01; *** p <. 001. df = degrees of freedom. ε = greenhouse & geisser epsilon factor for non-spericity to adjust the original df according to jennings and wood (1976). visual cue fails to modulate sentence processing 97 appendix b figure b-1. grand average erps (with baseline correction) at selected electrodes for the factors cue (topic vs. neutral) and word order (so vs. os) time-locked to the onset of the target sentence, that is, the first determiner phrase (dp1) followed by the verb and second determiner phrase (dp2): upper panel: topic so [dotted black] vs. neutral so [dotted grey], lower panel: topic os [solid black] vs. neutral os [solid grey]. negativity is plotted upwards. burmester, spalek and wartenburger 98 appendix c results of overall analysis of variance (anovas) of the erps for the different time windows time-locked to the onset of the first determiner phrase (dp1) of target sentences following both cue modalities (modality), that is, the implicit visual cue (i.e., visual: data of the present study) and the linguistic cue (i.e., linguistic: data published in burmester et al., 2014). df f-values 100 300 ms 300 500 ms 500 700 ms cue (topic vs. neutral) cue x roi cue x word order cue x word order x roi 1, 46 8, 368 1, 46 8, 224 1.970 0.311 (ε = 0.311) 4.162* 2.306 (ε = 0.404) 0.065 0.736 (ε = 0.428) 0.440 2.088 (ε = 0.452) 0.000 3.454* (ε = 0.373) 0.829 4.429** (ε = 0.441) word order (so vs. os) word order x roi 1, 46 8, 368 1.429 2.372# (ε = 0.507) 0.139 1.666 (ε = 0.483) 0.136 1.435 (ε = 0.460) modality (visual vs. linguistic) x roi modality x cue modality x cue x roi modality x word order modality x word order x roi modality x cue x word order modality x cue x word order x roi 8, 368 1, 46 8, 368 1, 46 8, 368 1, 46 8, 368 3.190* 0.089 0.538 5.793* 1.117 0.285 0.446 1.266 0.018 0.592 0.567 0.377 0.200 1.843 1.378 1.778 1.891 3.717 1.152 3.963# 1.487 note. greenhouse & geisser (1959) corrected significance levels: # p <. 06; * p <. 05; ** p <. 01; *** p <. 001. df = original degrees of freedom. ε = greenhouse & geisser epsilon factor for non-spericity to adjust the original df according to jennings and wood (1976). visual cue fails to modulate sentence processing 99 appendix d results of post hoc pairwise t-tests of the erps to resolve the roi x cue position interaction in the n1 and p2 time window (i.e., 100 200 ms, 250 350 ms) time-locked to the onset of the visual cue. cue position topic left vs. topic right topic left vs. neutral bottom topic right vs. neutral bottom 100 200 ms 250-350 ms 100-200 ms 250-350 ms 100 200 ms 250 350 ms roi t(28) p t(28) p t(28) p t(28) p t(28) p t(28) p left frontal right frontal left fronto-central right fronto-central left centro-parietal right centro-parietal frontal-midline central midline parietal midline 0.753 -1.076 1.512 0.237 0.568 1.055 1.542 1.646 2.254 0.458 0.291 0.142 0.814 0.575 0.301 0.134 0.111 0.032 1.153 -2.207 2.616 0.289 1.333 0.111 0.804 3.210 2.625 0.259 0.036 0.014 0.775 0.193 0.912 0.428 0.003 0.014 3.780 -0.344 4.178 2.573 4.885 4.208 2.523 3.810 6.538 0.001* 0.734 <.001* 0.016 <.001* <.001* 0.018 0.001* <.001* 3.726 0.909 3.328 2.907 2.844 1.962 3.068 4.503 3.994 0.001* 0.371 0.002 0.007 0.008 0.060 0.005 <.001* <.001* 1.659 0.914 2.220 2.247 4.313 3.467 0.397 1.883 5.021 0.108 0.368 0.035 0.033 <.001* 0.002 0.694 0.070 <.001* 1.767 3.131 1.093 2.755 2.023 2.369 1.954 1.145 1.348 0.088 0.004 0.284 0.010 0.053 0.025 0.061 0.262 0.188 note. bonferroni corrected significance level (two-tailed) at α = .05: * p <. 002. burmester, spalek and wartenburger 100 references gerry t.m. altmann (2004). language-mediated eye movements in the absence of a visual world: the ‘blank screen paradigm’. cognition, 93(2): b79–b87. american electroencephalographic society (2006). guideline 5: guidelines for standard electrode position nomenclature. journal of clinical neurophysiology, 11: 111–113. mira ariel (1988). referring and accessibility. journal of linguistics, 24(01): 65–87. arnold (1998). reference form and discourse patterns. phd thesis, stanford university, stanford, california. jennifer e. arnold (2010). how speakers refer: the role of accessibility. language and linguistics compass, 4(4): 187–203. jennifer e. arnold and zenzi m. griffin (2007). the effect of additional characters on choice of referring expression: everyone counts. journal of memory and language, 56(4): 521–536. jennifer e. arnold and shin-yi c. lao (2015). effects of psychological attention on pronoun comprehension. language, cognition and neuroscience, 30(7): 832–852. r. harald baayen (2008). analysing linguistic data. cambridge university press, cambridge. markus bader and jana häussler (2010). word order in german: a corpus study. lingua, 120(3): 717–762. douglas bates, martin mächler, ben bolker, and steve walker (2015). fitting linear mixed-effects models using lme4. journal of statistical software, 67(1): 1–48. kathryn j. bock and richard k. warren (1985). conceptual accessibility and syntactic structure in sentence formulation. cognition, 21(1): 47–67. sara bögels, herbert schriefers, wietske vonk, and dorothee j. chwilla (2011). pitch accents in context: how listeners process accentuation in referential communication. neuropsychologia, 49(7): 2022–2036. ina bornkessel and matthias schlesewsky (2006). the extended argument dependency model: a neurocognitive approach to sentence comprehension across languages. psychological review, 113(4): 787–821. ina bornkessel, matthias schlesewsky, and angela d. friederici (2003). contextual information modulates initial processes of syntactic integration: the role of interversus intrasentential predictions. journal of experimental psychology: learning, memory, and cognition, 29(5): 871–882. gordon h. bower and daniel g. morrow (1990). mental models in narrative comprehension. science, 247(4938): 44–48. a. ally brandon and e. budson andrew (2007). the worth of pictures: using high density eventrelated potentials to understand the memorial power of pictures and the dynamics of recognition memory. neuroimage, 35(1): 378–395. harm brouwer, matthew w. crocker, noortje j. venhuizen, and john c.j. hoeks (2017). a neurocomputational model of the n400 and the p600 in language processing. cognitive science, 41: 1318–1352. petra burkhardt (2006). inferential bridging relations reveal distinct neural mechanisms: evidence from event-related brain potentials. brain and language, 98(2): 159–168. juliane burmester, antje sauermann, katharina spalek, and isabell wartenburger (2018). sensitivity to salience: linguistic vs. visual cues affect sentence processing and pronoun resolution. language, cognition and neuroscience, 33(6): 784–801. juliane burmester, katharina spalek, and isabell wartenburger (2014). context updating during sentence comprehension: the effect of aboutness topic. brain and language, 137: 62–76. visual cue fails to modulate sentence processing 101 herbert h. clark and eve v. clark (1977). psychology and language: an introduction to psycholinguistics. harcourt brace jovanovich, new york. maurizio corbetta and gordon l. shulman (2002). control of goal-directed and stimulus-driven attention in the brain. nature reviews neuroscience, 3(3): 201–215. matthew w. crocker, pia knoeferle, and marshall r. mayberry (2010). situated sentence processing: the coordinated interplay account and a neurobehavioral model. brain and language, 112(3): 189–201. francesca delogu, heiner drenhaus, and matthew w. crocker (2018). on the predictability of event boundaries in discourse: an erp investigation. memory & cognition, 46(2): 315–325. dusana dorjee, lydia devenney, and guillaume thierry (2010). written words supersede pictures in priming semantic access: a p300 study. neuroreport, 21(13): 887–891. matthew s. dryer (2013). order of subject, object and verb. in: matthew s. dryer and martin haspelmath (eds). the world atlas of language structures online. max planck institute for evolutionary anthropology, leipzig, available online at http://wals.info/chapter/81. lyn frazier and janet d. fodor (1978). the sausage machine: a new two-stage parsing model. cognition, 6(4): 291–325. angela d. friederici (2002). towards a neural basis of auditory sentence processing. trends in cognitive sciences, 6(2): 78–84. kumiko fukumura, roger p.g. van gompel, and martin j. pickering (2010). the use of visual context during the production of referring expressions. the quarterly journal of experimental psychology, 63(9): 1700–1715. giorgio ganis, marta kutas, and martin i. sereno (1996). the search for "common sense": an electrophysiological study of the comprehension of words and pictures in reading. journal of cognitive neuroscience, 8(2): 89–106. lesya y. ganushchak, agnieszka e. konopka, and yiya chen (2017). accessibility of referent information influences sentence planning: an eye-tracking study. frontiers in psychology, 8: 250. morton a. gernsbacher (1991). cognitive processes and mechanisms in language comprehension: the structure building framework. the psychology of learning and motivation, 27: 217–263. thomas givón (1988). the pragmatics of word-order: predictability, importance and attention. in: michael hammond, edith a. moravczik, and jessica r. wirth (eds). studies in syntactic typology. typological studies in language. 17. john benjamins publishing company, amsterdam, 243–284. lila r. gleitman, david january, rebecca nappa, and john c. trueswell (2007). on the give and take between event apprehension and utterance formulation. journal of memory and language, 57(4): 544–569. gabriele gratton, michael g. h. coles, and emanuel donchin (1983). a new method for off-line removal of ocular artifact. electroencephalography and clinical neurophysiology, 55: 468– 484. samuel w. greenhouse and seymour geisser (1959). on methods in the analysis of profile data. psychometrika, 24(2): 95–112. tanja grewe, ina bornkessel, stefan zysset, richard wiese, d.yves von cramon, and matthias schlesewsky (2006). linguistic prominence and broca's area: the influence of animacy as a linearization principle. neuroimage, 32(3): 1395–1402. barbara j. grosz and candace l. sidner (1986). attention, intentions, and the structure of discourse. computational linguistics, 12(3): 175–204. jeanette k. gundel (1985). `shared knowledge’ and topicality. journal of pragmatics, 9: 83–107. jeanette k. gundel, nancy hedberg, and ron zacharski (1993). cognitive status and the form of referring expressions in discourse. language, 69(2): 274–307. burmester, spalek and wartenburger 102 peter hagoort and jos j.a. van berkum (2007). beyond the sentence given. philosophical transactions of the royal society of london. series b, biological sciences, 362(1481): 801– 811. joy e. hanna and susan e. brennan (2007). speakers’ eye gaze disambiguates referring expressions early during face-to-face conversation. journal of memory and language, 57(4): 596–615. julian heister, kay-michael würzner, johannes bubenzer, edmund pohl, thomas hanneforth, alexander geyken, and reinhold kliegl (2011). dlexdb – eine lexikalische datenbank für die psychologische und linguistische forschung. psychologische rundschau, 62(1): 10–20. barbara hemforth (1993). kognitives parsing: repräsentation und verarbeitung sprachlichen wissens. phd thesis, university of bochum, bochum. henning holle, christian obermeier, maren schmidt-kassow, angela d. friederici, jamie ward, and thomas c. gunter (2012). gesture facilitates the syntactic analysis of speech. frontiers in psychology, 3: 74. ray jackendoff (1972). semantic interpretation in generative grammar. mit press, cambridge, mass. richard jennings and charles c. wood (1976). the ɛ‐adjustment procedure for repeated‐measures analyses of variance. psychophysiology, 13(3): 277–278. philip n. johnson-laird (1980). mental models in cognitive science. cognitive science proceedings, 4: 71–115. marcel adam just and patricia a. carpenter (1980). a theory of reading: from eye fixations to comprehension. psychological review, 87(4): 329–354. edith kaan, andrea c. dallas, and christopher m. barkley (2007). processing bare quantifiers in discourse. brain research, 1146: 199–209. pia knoeferle, boukje habets, matthew w. crocker, and thomas f. münte (2007). visual scenes trigger immediate syntactic reanalysis: evidence from erps during situated spoken comprehension. cerebral cortex, 18(4): 789–795. pia knoeferle and helene kreysa (2012). can speaker gaze modulate syntactic structuring and thematic role assignment during spoken sentence comprehension? frontiers in psychology, 3: 538. willem j. m. levelt (1989). speaking: from intention to articulation. mit press, cambridge, mass. steven j. luck and steven a. hillyard (1994). electrophysiological correlates of feature analysis during visual search. psychophysiology, 31: 291–308. george r. mangun (1995). neural mechanisms of visual selective attention. psychophysiology, 32: 4–18. william marslen-wilson and lorraine k. tyler (1980). the temporal structure of spoken language understanding. cognition, 8(1): 1–71. mike matzke, heinke mai, wido nager, jascha rüsseler, and thomas f. münte (2002). the costs of freedom: an erp – study of non-canonical sentences. clinical neurophysiology, 113: 844– 852. andriy myachykov, dominic thompson, simon garrod, and christoph scheepers (2012). referential and visual cues to structural choice in visually situated sentence production. frontiers in psychology, 2: 396. rebecca nappa and jennifer e. arnold (2014). the road to understanding is paved with the speaker’s intentions: cues to the speaker’s attention and intentions affect pronoun comprehension. cognitive psychology, 70: 58–81. rebecca nappa, allison wessel, katherine l. mceldoon, lila r. gleitman, and john c. trueswell (2009). use of speaker's gaze and syntax in verb learning. language learning and development, 5(4): 203–234. visual cue fails to modulate sentence processing 103 r. carolus oldfield (1971). the assessment and analysis of handedness: the edinburgh handedness inventory. neuropsychologia, 9: 97–113. marte otten and jos j.a. van berkum (2008). discourse-based word anticipation during language processing: prediction or priming? discourse processes, 45(6): 464–496. martin j. pickering and simon garrod (2004). toward a mechanistic psychology of dialogue. behavioral and brain science, 27: 169–226. paul portner (2007). instructions for interpretation as separate performatives. in: kerstin schwabe and susanne winkler (eds.). on information structure, meaning and form: generalizations across languages. john benjamins publishing company, amsterdam/philadelphia. mercè prat-sala and holly p. branigan (2000). discourse constraints on syntactic processing in language production: a cross-linguistic study in english and spanish. journal of memory and language, 42(2): 168–182. ellen f. prince (1981). topicalization, focus-movement, and yiddish-movement: a pragmatic differentiation. in: proceedings of the seventh annual meeting of the berkeley linguistics society. berkley linguistics society. 249–264. ellen f. prince (1992). the zpg letter: subjects, definiteness, and information-status. in: sandra a. thompson and wiliam c. mann (eds.). discourse description: diverse analyses of a fund raising text. john benjamins publishing company, amsterdam, 295–325. r core team (2013). r: a language and environment for statistical computing, vienna, austria. tanja reinhart (1981). pragmatics and linguistics: an analysis of sentence topics. philosophica, 27(1): 53–94. christoph scheepers and matthew w. crocker (2004). constituent order priming from reading to listening: a visual-world study. in: manuela carreiras and charles clifton, jr. (eds.). the online study of sentence comprehension: eyetracking, erps and beyond. psychology press, new york, 167–185. matthias schlesewsky, ina bornkessel, and stefan frisch (2003). the neurophysiological basis of word order variations in german. brain and language, 86(1): 116–128. petra b. schumacher (2014). content and context in incremental processing: “the ham sandwich” revisited. philosophical studies, 168(1): 151–165. petra b. schumacher and yu-chen hung (2012). positional influences on information packaging: insights from topological fields in german. journal of memory and language, 67(2): 295–310. samuel c. spevack, j. benjamin falandays, brandon batzloff, and michael j. spivey (2018). interactivity of language. language and linguistics compass, 12(7): e12282. maria staudte and matthew w. crocker (2011). investigating joint attention mechanisms through spoken human–robot interaction. cognition, 120(2): 268–291. maria staudte, matthew w. crocker, alexis heloir, and michael kipp (2014). the influence of speaker gaze on listener comprehension: contrasting visual versus intentional accounts. cognition, 133(1): 317–328. russell s. tomlin (1995). focal attention, voice, and word order: an experimental, cross-linguistic study. in: pamela downing and michael noonan (eds.). word order in discourse. john benjamins publishing company, amsterdam and philadelphia, 517–555. russell s. tomlin (1997). mapping conceptual representations into linguistic representations: the role of attention in grammar. in: jan nuyts and eric pederson (eds.). language and conceptualisation. language, culture & cognition. 1. cambridge university press, cambridge, 162–189. elizabeth a. turner and ragnar rommetveit (1968). focus of attention in recall of active and passive sentences. journal of verbal learning and verbal behavior, 7(2): 543–548. jos j.a. van berkum, colin m. brown, pienie zwitserlood, valesca kooijman, and peter hagoort (2005). anticipating upcoming words in discourse: evidence from erps and reading times. journal of experimental psychology: learning, memory, and cognition, 31(3): 443–467. burmester, spalek and wartenburger 104 teun a. van dijk and walter kintsch (1983). strategies of discourse comprehension. academic press, new york. jorrig vogels, emiel krahmer, and alfons maes (2013). who is where referred to how, and why? the influence of visual saliency on referent accessibility in spoken language production. language and cognitive processes, 28(9): 1323–1349. andrea weber and karin müller (2004). word order variation in german main clauses: a corpus analysis. in: proceedings of the 20th international conference on computational linguistics, geneva, switzerland, 71–77. xiao lei zhang, henri begleiter, bernice porjesz, and ann litke (1997). visual object priming differs from visual word priming: an erp study. electroencephalography and clinical neurophysiology, 102(3): 200–215. rolf a. zwaan (2004). the immersed experiencer: towards an embodied theory of language comprehension. in: brian h. ross (editor). the psychology of learning and motivation: advances in research and theory. acadamic press, new york, 35–62. rolf a. zwaan and gabriel a. radvansky (1998). situation models in language comprehension and memory. psychological bulletin, 123(2): 162–185. microsoft word dnd-harbusch-final-without-doi.docx dialogue and discourse 5 (2011) 313-332 submitted 1/2010; accepted 2/2011; doi: 10.5087/dad.2011.112 published online 5/11 ©2011 karin harbusch incremental sentence production and clausal coordinate ellipsis: 
 a treebank study comparing spoken and written language 
in dutch and german karin harbusch harbusch@uni-koblenz.de university of koblenz-landau computer science department universitätsstr. 1 56070 koblenz germany editors: david schlangen and hannes rieser abstract from two corpus studies into varieties of clausal coordination in english (meyer, 1995 and greenbaum & nelson, 1999), it is known that the incidence of clausal coordinate ellipsis (cce) is about two times higher in written than in spoken language. we present a treebank study into cce in written and spoken dutch and german which confirms this tendency. moreover, we observe considerable differences between written and spoken language with respect to the incidence of four main types of clausal coordinate ellipsis—gapping, forward conjunction reduction (fcr), backward conjunction reduction (bcr), and subject gap with finite/fronted verb (sgf). we argue that the detailed data pattern cannot be accounted for in terms of audience design, and propose an explanation based on the assumption that during spontaneous speaking—but not during writing—, the scope of online grammatical planning is basically restricted to one (finite) clause. keywords: treebanks, human language production, incrementality, coordination, coordinate ellipsis, gapping, conjunction reduction, audience design 1 introduction one of the benefits of incremental sentence production is a reduction of the working memory capacity needed for advance planning: the planning units can be of considerably smaller size (measured in terms of word length) than in case of non-incremental production. the same advantage has been claimed for the various forms of ellipsis, which preempt the need to plan the detailed shape of one or more constituents and thereby reduce the size of planning units. because working memory load tends to be higher in spoken than in written language, one expects that speakers, in comparison with writers, will more frequently resort both to incremental production and to the use of elliptical constructions. as far as we know, this prediction is generally borne out for incremental production. however, in two corpus studies into the incidence of clausal coordinate ellipsis in spoken and written english, meyer (1995) and greenbaum & nelson (1999) obtained a data pattern opposite to the prediction. in written clausal coordinations, the proportion of elliptical versions is about twice as high as in spoken clausal coordinations. recent treebanks with spoken and written sentences for dutch and german enable us to test whether the latter finding generalizes to other germanic languages, and to explore how spoken and written language production affect the incidence of four different types of clausal coordinate ellipsis—gapping, forward conjunction reduction (fcr), backward conjunction reduction (bcr), and subject gap with finite/fronted verb (sgf). as far as we know, the only treebank study of cce was carried out by zinsmeister (2006). in a incremental sentence production and clausal coordinate ellipsis 314 quantitative study with the german taz1 treebank, she identified 8,133 sentences (37% of the total number of sentences) with a coordination of syntactic constituents of any type. she only reports one number dealing with cce types: 83 sentences embodying sgf, i.e., about 1 percent. (this proportion is not directly comparable to the one we report below because we only take into account sentences with clausal coordinations.) in this paper, we explore treebanks for spoken and written dutch and german. for dutch, we probe the alpino treebank (van der beek et al., 2002) that consists of about 7,000 sentences of newspaper text, and cgn2.0 (van eerten, 2007) with about 130,000 spoken sentences (or rather dialogue turns) from more than ten different domains. for german, we explore the tiger treebank (brants et al., 2004) of written newspaper text with about 50,000 sentences, and the spoken utterances in the verbmobil dialogues in tüba-d/s (about 38,000 sentences). we first introduce the four types of clausal coordinate ellipsis (section 2). in section 3, we present the detailed results of the four treebank studies—first the dutch study (section 3.1), then the german study (section 3.2), and finally some striking between-modality differences and cross-language similarities (section 3.3). in section 4, we propose a theoretical explanation of the data, focusing on the differential effects that incremental sentence production exerts on the individual cce types. finally, in section 5, we sum up and mention desiderata for future work. 2 a typology of clausal coordinate ellipsis in the linguistic literature on clausal coordination one often distinguishes four main types of coordinate ellipsis2 (for overviews, see van oirsow, 1987; steedman, 2000; sag et al. 2003; te velde, 2006; and kempen, 2009):  gapping, with three variants called: o long distance gapping (ldg), o subgapping, and o stripping,  forward conjunction reduction (fcr),  backward conjunction reduction (bcr; also known as right node raising or rnr), and  subject gap with finite/fronted verb (sgf). table 1 illustrates these cce types in terms of examples taken from the verbmobil dialogues in tüba-d/s and cgn2.0.3 we adopt the following psycholinguistically motivated definitions of cce types. they derive from work by kempen (2009) who argues that coordinations are structurally similar to self-repairs in spontaneous speech, i.e. can be viewed as special type of “update” constructions (cf. section 4.1). with respect to the underlying linguistic framework, we try to be as theory-neutral as possible. however, we presuppose a separation between the hierarchical and the linear structure of sentences. in descriptions of linear order, we use the terminology of topological fields (with forefield, midfield, and endfield as translations of vorfeld, mittelfeld, and nachfeld, respectively; cf. höhle, 1986). as the encodings of hierarchical structures in the four treebanks differ 1 tüba-d/z treebank (hinrichs et al. 2004) is a corpus of german newspaper texts currently comprising about 22,000 sentences taken from the wissenschafts-cd of “die tageszeitung” (taz). henceforth, it is called the “taz” treebank in order to avoid confusion with tüba-d/s (stegmann et al., 2000), i.e. a spoken german treebank for the verbmobil domain (wahlster, 2000), which we will call the “verbmobil” treebank in the following. 2 we do not deal here with the elliptical constructions known as vp ellipsis, vp anaphora and pseudogapping because they involve the generation of pro-forms instead of, or in addition to, the ellipsis proper. for example, john laughed, and mary did, too—a case of vp ellipsis—, includes the pro-form did. nor do we account for recasts of clausal coordination as coordinate nps (e.g., changing john likes skating and peter likes skiing into john and peter like skating and skiing, respectively). presumably, such conversions involve a semantic rather than syntactic mechanism. 3 we were unable to identify tokens of long-distance gapping (ldg) in the verbmobil corpus. in cgn2.0, we found nine exemplars. the theory proposed in section 4 may explain why these numbers are so small in spoken text. harbusch 315 g a pp in g (g ) (1) nachdem ich selbst ungern in die oper gehe und as i myself reluctantly to the opera go and nachdemg sie so gerne in die oper geheng ... you so readily ‘as i myself go to the opera reluctantly and you [do so] readily ...’ l d g (( g) + g) (2) hij had zich verkleed als meisje en he had himself dressed-up as girl and z'n vriend hadg [zich verkleed]gg als oude vrouw his friend as old woman ‘he had dressed up as a girl and his friend [had dressed up] as an old woman’ su b g a pp in g (s g) (3) dann können wir uns erst um die verbindung kümmern und then can we ourselves first for the connection care and dann [können wir]sg die hotels aussuchen then the hotels select ‘first we may take care of the connection and afterwards [we may] select the hotels’ st r ip pi n g (s tr ) (4) am freitag hätte ich bis elf uhr zeit bzw. on friday would-have i until eleven o’clock time and [am freitag hätte ich]str ab dreizehn uhr auch wieder zeitstr onward-from 1pm o’clock also again ‘on friday, i would have time until 11am and from 1pm onward too’ fc r (f ) (5) wenn sie schon in dem parkhotel waren und if you already in the park-hotel stayed and [wenn sie]f das gut fanden, … it ok found ‘if you already stayed in the park hotel and found it ok’ (6) da wäre hotel x, [welches am bahnhof liegt und there would-be hotel x which at-the station lies and welches f zum zentrum 15 minuten laufzeit hat to-the center 15 minutes walking-time has ‘there is hotel x which is located at the station and [which is] a 15 minutes walk to the center’ b c r (b ) (7) im juli hätte ich nur einen [tag frei]b und in-the july would-have i only one and im mai zwei tage frei in-the may two days off ‘in july, i only have one, and in may two days off’ sg f (s ) (8) dann fahren wir frühmorgens los und then get we early-morning started and [wir]s sind mittags da are at-noon there ‘then, we get started early morning and [we’ll] arrive at noon’ table 1. cce examples in spoken german and dutch. struck-out text represents elisions. substantially—e.g., not all treebanks use vp nodes (cf. section 3)—we suppose the tree structures to be rather flat. all forms of gapping (cf. examples (1) to (4) in table 1) are characterized by elision of the posterior member of a pair of lemma-identical verbs.4 the position of this verb need not be peripheral but is often medial, as in (2) through (4). every non-elided constituent 4 in our definitions of cce types, we restrict ourselves to coordinations encompassing two conjuncts, called anterior (first, left) and posterior (second, right), respectively (cf. footnote 13). incremental sentence production and clausal coordinate ellipsis 316 (“remnant”) in the posterior conjunct should pair up with a constituent in the anterior conjunct that has the same grammatical function but is not coreferential.5 stated differently, the members of such a pair are contrastive—in (1), the subjects ich selbst ‘i myself’ vs. sie ‘you’, and the modifiers ungerne ‘reluctantly/dislike-to’ vs. gerne ‘readily/like-to’. only contrastive constituents are expressed overtly in the posterior conjunct. this restriction rules out cases as john eats apples and peter eats in the car. notice that gehen ‘go (3rd person)’ in the posterior conjunct of (1) can be elided although it has no wordform-identical (but only a lemma-identical) anterior counterpart (wordform gehe in the first conjunct is 1st person, singular, present tense of the verb gehen ‘go’). gapped sentences resemble answers to implicit multiple wh-question. for instance, steedman (1990:248) writes: “[e]ven the most basic gapped sentence like fred ate bread, and harry, bananas is only really felicitous in contexts which support (or accommodate) the presupposition that the topic under discussion is who ate what?”. according to reich (2007), gapping answers an implicit6 wh-question. hartmann (2000) examines the close connection between focus and ellipsis in gapping. this correspondence is reflected in the particular intonation contour typical of gapping structures. in long distance gapping (ldg), the remnants originate from different clauses (more precisely: from different clauses that belong to the same superclause; a superclause is a hierarchy of finite or nonfinite clauses that—with the possible exception of the topmost clause—do not include a subordinating conjunction. in (2), hij ‘he’ belongs to the main clause headed by the verb had ‘had’ but zich ‘himself’ and als meisje ‘as a girl’ to the nonfinite complement clause headed by the past participle verkleed ‘dressed up’. in subgapping, the posterior conjunct includes a remnant in the form of a nonfinite complement clause (vp; aussuchen ‘select’ in (3)). in stripping, the posterior conjunct is left with one non-verb remnant, often supplemented by a sentential adverb such as auch ‘too’ (in example (4), auch wieder ‘again’), or a negation. notice that our definition of gapping in terms of the hierarchical structure of the conjuncts circumvents word order issues which are not relevant for gapping. for instance, (9) exemplifies one alternative to the word order in example (4). actually, each of the constituents zeit, ich, am freitag, and auch wieder can go to the forefield. (9) … ab 13 uhr hätte ich am freitag auch wieder zeit from-onward 13 o’clock would-have i on friday also again time in stripping, there is only one remnant, i.e. contrastive constituent. however, we also count as stripping those cases where the posterior conjunct introduces new semantic aspects of the event described in the anterior conjunct (see example (10) from the verbmobil corpus). (10) ich möchte gerne nach hannover, und zwar über karneval nächstes jahr i would like(to-go) to hannover and namely during carnival next year ‘i want to go to hannover, (and) namely during carnival next year.’ in forward conjunction reduction (fcr), elision affects the posterior token of a pair of left-peripheral strings consisting of one or more wordform-identical major constitu 5 we distinguish three identity relationships between constituents in coordinated conjuncts: lemma identity, wordform identity and coreferentiality. for lemma identity, only the lexical entries (citation form) of the constituents have to be identical; wordform identity requires, in addition, identity of their morphological features. coreferential constituents refer to the same discourse entity or entities, irrespective of whether or not they include the same lemma(s). lemmas are “syntactic words”; their lexical entries specify the sentential environments in which they are allowed to occur. the morphological and phonological information associated with words is specified in another type of lexical entry called word forms or lexemes. 6 example (i) below might be analysed as implicit question answering based on gapping, more precisely, on stripping. here, the speaker formulates a question. the elliptical construction contains the tentative answer—now by the same speaker. (in this example, mit dem flugzeug can hardly have been intended as an extraposed part of the question, assuming that in the verbmobil context wie queries the means of transportation. (i) wie sollen wir uns dahin begeben, mit dem flugzeug? how shall we ourselves there move with the plane ‘how shall we go there, by plane?’ harbusch 317 ents. in (5), the posterior tokens of wenn sie ‘if you’ and in (6), welches ‘which’, respectively, belong to such a pair and are eligible for fcr. notice that if the finite verb belongs to the left-peripheral string, fcr and gapping are not always distinguishable. we count such constructions as fcr. example (11) is a case in point. (11) dann können wir noch essen gehen oder then can we even eat go or dann können wir noch irgendwas in der richtung tun something in the direction do ‘then we can go dining or do something like that’ backward conjunction reduction (bcr) is almost the mirror image of fcr as it deletes the anterior member of a pair of right-peripheral lemma-identical word strings (tag(e) ‘day(s)’ in (7)); however, bcr may elide part of a major constituent—e.g. only the part tag(e) of the direct object in (7). in addition, it requires only lemma identity (cf. number singular vs. plural of the elided noun in example (7)).7 bcr does not necessarily elide the verb (see (12)). (12) erst hören wir uns und dann sehen wir uns wohl am bahnhof first hear we ourselves and then see we ourselves probably at-the station ‘first we phone each other, and then we see each other at the station’ subject gap with finite/fronted verb (sgf; see wunderlich, 1988, and höhle (1983) who calls the phenomenon “slf koordinationen”, for “subjektlücken in finiten/frontalen sätzen”; for a recent survey, see kathol, 2001) elides the subject of the posterior conjunct in a main clause, when in the anterior conjunct the wordform-identical subject follows the finite verb (subject-verb inversion). elision of the posterior subject cannot be due to fcr since the anterior subject is not left-peripheral. furthermore, the initial constituent of an anterior sgf conjunct should not be an argument. this is illustrated by the ill-formed ellipsis in example (13) where a complement clause opens the anterior conjunct. in sgf case (8), the initial constituent is an adjunct. (13) *das examen bestehen will er und ers kann auch the exam to-pass wants he and can too ‘he wants to pass the exam and is also able to [do this]’ we also subsume under the heading of sgf cases like (14), where the anterior conjunct is a conditional subordinate clause rather than a main clause. (see höhle (1990) and reich (2008) for discussion of the affinity between this structure and sgf.)8 (14) ... dann reicht es ja, wenn wir ungefähr um neun losfahren würden und then suffices it already if we about at nine get-startedinf would and wir würden dann mittags dort ankommen would then at-noon there arrive ‘… then it suffices already if we would leave at nine and we would arrive there at noon’ 3 cce in spoken and written dutch and german in this section, we present a detailed overview of the incidence of the four types of clausal coordinate ellipsis in dutch and german treebanks with spoken and written text. we first describe our study for dutch with the treebanks alpino for written newspaper text, and cgn2.0 for spoken text. then, we report the findings obtained with the tiger treebank of 7 notice that example (7) cannot be analyzed as a case of one-anaphora, for at least two reasons. first, the anaphoric element is in the anterior rather than the posterior conjunct. second, anaphoric use requires eins instead of ein, as in example (ia/b). (tag has masculine, auto neuter gender.) (ia) ich habe ein und du zwei autos. ‘i have one and you two cars’ (ib) ich habe zwei autos und du eins/*ein ‘i have two cars and you one’ 8 notice that the word order in the second conjunct of example (5) distinguishes fcr from sgf: in sgf, the second conjunct always has main-clause word order (i.e. yielding und fanden das gut ‘and found it ok’ instead of und das gut fanden in (5)). incremental sentence production and clausal coordinate ellipsis 318 the german written newspaper text, and with the tüba-d/s treebank that contains the spoken verbmobil dialogues. the linguistic encodings in the four treebanks under consideration are rather different. we cannot describe the individual formats in detail for reasons of space (but see the appendices for relevant specifications of clausal coordination in the four treebank formats). in order to obtain comparable numbers, we defined search patterns that implement the definitions of the four cce types and take into account differences between the specific notations used in the four treebanks (such as whether or not they encode vps and vp coordination). all retrieved sentences were manually classified with respect to cce type in a uniform manner. if one coordination embodies more than one cce type, as in (15), (17) or (18) below, we counted them as separate instances. the incidence of a cce type in a corpus is expressed as the number of exemplars of that cce type in the total number of sentences with at least one coordination. 3.1 cce in dutch here, we report the results of our comparative study of cce in written and spoken language in dutch (also see harbusch & kempen, 2009a). 3.1.1 cce in the alpino treebank the alpino treebank, released in november 2002, contains 7,153 syntactically annotated dutch sentences. all annotations in this corpus were manually inspected. the sentences comprise the full (cdbl newspaper) part of the eindhoven corpus (uit den boogaart, 1975). figure 1 illustrates the encoding format for cce in the alpino treebank. in sentence (15), the subject (labeled su) and the verb complement (labeled vc) are elided due to fcr and bcr, respectively. in alpino, elisions are encoded by coreferential indices9 at the remnant and its corresponding empty leaf node. for instance, the index 1 is attached to the overt subject of the anterior conjunct, and to the subject of the posterior conjunct, where the subject has been deleted due to fcr. index 2 denotes the verb complement, which has been erased from the anterior conjunct due to bcr. (notice that, in the hierarchical structure, the verb complement has been encoded as part of the first conjunct although it has been elided due to bcr and occurs overtly as remnant in the second conjunct.) for more details on the labels used in the figure, see appendix a. (15) het moet en zal een nederlands stuk worden it should and will a dutch piece become ‘it should become and it will become a dutch piece’ we identified 931 sentences with at least one clausal coordination. within this set, we found 319 cce cases of any type. the detailed proportions of different cce types are reported in section 3.1.3, where they are compared to those in cgn2.0. 3.1.2 cce in the cgn2.0 treebank cgn stands for corpus gesproken nederlands ‘corpus of spoken dutch.’ the speech data originate from adult speakers of standard dutch in flanders (one third of the corpus material) and the netherlands (two third). version cgn2.0, released in 2004, comprises 130,594 sentences from various dialogue domains. the treebank has been encoded in the negra format and can be inspected using the tigersearch tool (könig & lezius, 2003). secondary edges, represented by curved edges in tree diagrams such as figures 2, relate remnants to the root node where the constituent is supposed to have been elided. in figure 2, the secondary edge expresses that the head (cf. edge label hd) bent ‘handed’ is supposed to be a child of the ssub node as well. however, the target position within the list of siblings is intentionally left undetermined. (for more details on the encodings used in cgn2.0, see appendix b.) in 9 the coreferential index also encodes subject-to-subject raising, indicating that the subject of the infinitival verb complement is identical with the subject of the governing finite verb. we ignored all non-elliptical uses of these indices. harbusch 319 figure 1. hierarchical structure of sentence (15) in alpino. example (16)—part of a longer utterance—the last word bent ‘are’ in the second conjunct has been deleted from the first conjunct due to backward conjunction reduction (bcr). we found 8,653 sentences with at least one clausal coordination. within this set, we counted 924 cce instances. figure 2. tree structure of example (16) in cgn2.0. (16) … omdat ik links en jij rechts bent … because i left-handed and you right-handed are ‘because i am left-handed and you [are] right-handed’ 3.1.3 cce incidence in spoken and written dutch in written dutch, the percentage of elliptical versions within the set of sentences with at least one clausal coordination is three times higher than in spoken dutch: 34% versus 11% (see table 2). this means that the lower incidence of cce in written than in spoken language that has been observed in english, holds for dutch as well. treebank average sentence length (words) total number of sentences with clausal coordination total number of cce percentage of sentences with clausal coordination cce percentage alpino (written) 17.8 931 319 13 34 cgn 2.0 (spoken) 8.6 8,653 924 6 11 table 2. clausal coordinate ellipsis in the dutch treebanks. the numbers in the rightmost column are percentages of the number of sentences containing one or more clausal coordination. incremental sentence production and clausal coordinate ellipsis 320 3.2 cce in german in this section, we test whether the differing cce tendencies in spoken and written language generalize to german (also see harbusch & kempen, 2009b). 3.2.1 cce in the tiger treebank the tiger treebank, released in december 2003, contains 50,474 german syntactically annotated sentences from a german newspaper corpus. as illustrated in figure 310, tiger’s annotation scheme uses many clause-level grammatical functions (subject, direct and indirect object, etc., shown as edge labels in the sentence diagrams). as already mentioned for cgn2.0, elided constituents in coordinate clauses are represented by secondary edges (represented by curved arrows pointing from the remnant to the root of the elided element). like in our cgn2.0 exploration, we deployed the tigersearch tool to design queries that retrieve clausal coordination (whether elliptical or not). we classified the elliptical ones (those including one or more secondary edges) into one of the four cce types by hand. (for details, see harbusch & kempen, 2007, 2009c.) we found 7,194 sentences with at least one clausal coordination. within this set, we counted 4,020 cce instances. figure 3. tree diagram for example (17): subgapping combined with bcr. the two remnants of the posterior clause are märkte ‘markets’ and the vp (nonfinite clause) getrennt werden ‘be separated’. (17) monopole sollen geknackt und märkte getrennt werden monopolies should shattered and markets split be ‘monopolies should be shattered and markets should be split’ 3.2.2 cce in the verbmobil treebank the verbmobil treebank is encoded in the same manner as the taz corpus (hinrichs et al., 2004) but rather differently from tiger (see lemnitzer and zinsmeister, 2006, page 82, for a comparison of the tag sets; cf. appendix d). the verbmobil treebank comprises 38,328 utterances in fifteen different subcorpora. we selected and retrieved all trees via tigersearch. however, the verbmobil treebank lacks annotations that relate remnants and elisions (cf. the indices in alpino, and the secondary edges in cgn2.0 and tiger). therefore, we classified all coordinated clauses for cce type manually. see appendix d for encoding details. figure 4 illustrates the structural encoding of example (18). this sentence exhibits “forward” elision of viertel vor zwölf könnte combined with “backward” elision of abholen. (18) viertel vor zwölf könnte ich sie oder mein fahrer sie abholen quarter to twelve could i you or my driver you up-pick ‘quarter to twelve, i could pick you up or my chauffeur could do so’ 10 the trees in figures 2 and 3—from cgn2.0 and tiger, respectively—look similar because they both originate from the tigersearch tool. however, the grammatical annotations are not identical. for instance, cgn2.0 does not use vp nodes. a combination of automatic sentence retrieval and manual inspection of the hits enabled us to obtain comparable numbers (see appendix c for encoding details). harbusch 321 figure 4. encoding of example (18) in verbmobil in terms of a coordination (node label fkoord) in the midfield (mf) of the subjects (grammatical function on in squared boxes) and the direct objects (oa) of the conjuncts. this example also illustrates a particularity of spoken language, namely underreduction with respect to backward elision of sie ‘you’ from the first conjunct. (without this sie, we would analyze the example as a coordination of nps rather than as clausal coordination—see section 4.3) in total, we found 3,713 verbmobil sentences with at least one clausal coordination. this set includes 1,314 cce tokens. 3.2.3 cce incidence in spoken and written german in written german, the percentage of elliptical versions within the set of clausal coordinations is considerably higher than in spoken german: 56% versus 35%. this means that the lower cce incidence in spoken than in written text holds not only for english and dutch but also for german. treebank average sentence length (words) total number of sentences with clausal coordination total number of cce percentage of sentences with clausal coordination cce percentage tiger (written) 17.6 7,194 4,020 14 56 verbmobil (spoken) 9.9 3,713 1,314 10 35 table 3. clausal coordinate ellipsis in the german treebanks. the numbers in the rightmost column are percentages of the number of sentences containing one or more clausal coordination. 3.3 incidence of cce types in spoken versus written language table 4 shows the relative frequencies in german and dutch of the four cce types we distinguish. it reveals striking differences between modalities and similarities across languages. in both dutch and german, fcr and gapping together take the lion’s share of cce instances: 80 to 92 percent. however the distribution of fcr and gapping differs considerably between the spoken and written modalities. in spoken language, gapping is responsible for about one third of the cce cases (32%), in written language for only one eighth of the cce cases (13%). in both languages, the relative frequencies of the cce types in the spoken vs. the written production modality differ significantly from one another: χ2 = 192.3, df = 3, p < 0.0001 for german; χ2 = 66.1, df = 3, p < 0.0001 for dutch. in the german treebanks, bcr11 and 11 as bcr is defined as elision of the right periphery of the anterior conjunct, and both dutch and german often have verbs in clause-final position, one may wonder how often bcr causes elision of verbs (and/or separable verb incremental sentence production and clausal coordinate ellipsis 322 spoken language written language cce type verbmobil (german) cgn 2.0 (dutch) tiger (german) alpino (dutch) gapping 33% 31% 17% 10% fcr 55% 61% 63% 82% bcr 1% 3% 10% 5% sgf 11% 5% 10% 3% table 4. relative frequencies of the four types of cce, expressed as percentages of the total set of sentences exhibiting cce. sgf are well represented (in particular sgf) whereas in the dutch corpora they live a somewhat marginal existence. in the next section, we propose an explanation for the results of our corpus study, focusing on the two main aspects in which spoken language appears to differ from written language: (1) a lower incidence of cce, and (2) within the cce types a higher incidence gapping. 4. incremental sentence production and cce as argued in the introduction, given the usually tighter processing constraints during speaking than during writing, one expects speakers, in comparison to writers, to resort more frequently to the use of elliptical constructions. however, the corpus work by meyer (1995) and greenbaum & nelson (1999) as well as the present treebank study yield a data pattern opposite to this prediction. in written clausal coordination, the proportion of elliptical versions is about twice as high as in spoken coordination. a tentative explanation of these data appeals to audience design (cf. bell, 1984): speakers try to reduce the cognitive load imposed on their dialogue partners—here, by avoiding ellipsis. greenbaum & nelson (1999), following meyer (1995), propose a similar explanation on page 116: “repetition helps the listener to understand what is being said by making the discourse less dense. full forms, which involve repetition, tend therefore to be preferred over ellipted forms in speech.” however, the empirical evidence for audience design as a systematic speaker strategy aiming to preclude comprehension problems is mixed (e.g., arnold, et al., 2004; haywood et al., 2005; brennan & hanna, 2009). moreover, this hypothesis cannot explain our observation that, in both languages, the incidence of gapping is higher in the spoken than in the written modality.12 in the following, we present another explanation for the complete data pattern.13 it is in line with incremental sentence production (see section 4.1 for a brief introduction). given the particles, which are always clause-final), and whether the frequencies of such elisions vary between modalities. we checked the corpora for such cases but they are rare and do not exhibit any salient differences. 12 an additional argument against an explanation in terms of audience design is suggested by the strong tendency to use telegraphic style in verbmobil. see examples (i) and (ii), where the hearer has to figure out which constituent(s) is/are left out. in both examples, the clause-initial position of both conjuncts is empty: ‘topic drop’. (i) startet bonn hauptbahnhof um acht uhr fünfundvierzig und leaves bonn main-station at 8 hour 45 and kommt an in hannover hauptbahnhof um zwölf uhr vier arrives in hannover main-station at 12 hour 4 ‘(it) leaves bonn main station at 8:45 am and arrives at the main station of hannover at 12:04pm’ (ii) liegt zentral, hat hallenbad und fitneßraum is located centrally has indoor-pool and fitness-room ‘(it) is located centrally, has indoor-pool and fitness-room’ 13 although in this paper we focus on binary coordinations (i.e. with two conjuncts), the explanation to be presented below will generalize to n-ary coordinations on the assumption that these are produced in a pairwise manner. that is, all pairs of conjoined (super)clauses (1,2), (2,3), …, (n-1,n) are checked for applicability of a cce operation. e.g., an n-ary bcr construction suppresses the right-periphery in all but the last clause. see, e.g., borsley (2005) for additional restrictions on n-ary coordinations. harbusch 323 stronger time pressure during speaking, we assume that, in order not to overtax current workspace capacity, speakers have a stronger tendency than writers to plan the grammatical shape of each clause in isolation. consequently, speakers are less likely to take into account the shape of coordinated clauses than writers, who have more opportunities for looking ahead/back. hence, cce constructions whose well-formedness does depend on both conjunct’s grammatical shape (i.e. fcr, bcr and sgf) are expected to occur less frequently in spoken language. in contrast, gapping’s relative frequency will actually increase because in gapping only one clause needs to be considered at planning time (in a repair-like manner; cf. section 4.2). moreover, we present in section 4.3 more evidence from the verbmobil corpus for the limitation of the workspace capacity. finally, we sum up the line of argumentation in section 4.4. 4. 1 key features of incremental sentence production and kempen’s (2009) psycholinguistically motivated ellipsis generation model according to levelt’s (1989) widely accepted “blueprint of the speaker”, sentence production is a five-stage process:  intending: creating the communicative intention,  conceptualizing: mapping the intention onto a set of concepts and conceptual (thematic) relations; output from this stage is a conceptual structure (“message”) in the form of a tree with branches whose linear order is undefined,  grammatical encoding: mapping concepts onto lemmas; output is the hierarchical/dominance and the linear/precedence structure of the sentence,  phonological encoding: taking the linearly ordered tree as input, the spoken form of the sentence is produced by replacing the lemmas by lexemes (phonological wordforms), and  articulating (not further addressed here). all mapping information during these stages stems from the mental lexicon, in particular concepts, conceptual (thematic) relations, and lemmas. (for present purposes, we simply assume a one-to-one correspondence between concepts and lemmas.) lemmas will be attached to syntactic trees as terminal nodes. associated with every lemma is information specifying how it can be used in sentences. for instance, a lemma that corresponds to an action concept specifies how the concept’s thematic relations are mapped onto grammatical functions (e.g., actor onto subject and patient/theme onto direct object). lexemes (inflected or uninflected wordforms) are yet another type of lexical entries in the mental lexicon. the creation of a communicative intention during speaking and writing is not always an “indiscerptible event” yielding precisely the full meaning underlying the next sentence at one point in time. often it is a “time-extended event” during which the meaning is conceived gradually, fragment by fragment. this allows the language user to initiate grammatical encoding before the complete communicative intention is available—for instance already when the “topic” has been conceived but not yet the “comment”. through this incremental mode of sentence production, language users reduce the risk of overtaxing the currently available processing resources, in particular the workspace for the formation and short-term maintenance of grammatical structures. of course, this strategy increases the risk of grammatically incoherent utterances; but these can be repaired online, in particular when they threaten to become incomprehensible. for psycholinguistic evidence regarding incremental sentence production see, e.g., kempen & hoenkamp, 1982, 1987, and levelt, 1989. for grammar formalisms that take incrementailty into account, e.g. tree adjoining grammar, see ferreira 2000 and frank & badecker, 2001. for dynamic syntax, see purver & kempson, 2004; for a formalismindependent parsing method based on unsupervised-learning from treebanks, see, seginer, 2007. as mentioned in section 2, our account of cce phenomena is inspired by kempen’s (2009) theory of coordinate ellipsis in dutch and german. the four types of clausal coordinate ellipsis—gapping, fcr, bcr and sgf—are argued to originate in four different stages incremental sentence production and clausal coordinate ellipsis 324 of sentence production. moreover, he classifies the phenomena in terms of appropriateness repair-likeness or unlikeness. repairs occur when—halfway through, or at the end of, a sentence—speakers modify the communicative intention underlying the current utterance in such a way that at least part of the utterance needs to be updated. in such repairs, some or all of the originally intended content that already has been encoded conceptually and grammatically and surfaced as an overt utterance, is replaced by more appropriate content, which requires at least a partially different overt realization. the ellipsis phenomena fcr and gapping can be classified as a repair process according to kempen. in case of gapping, the conceptual content underlying the verb is shared between conjuncts. this also holds for the thematic relations contracted by the verb and for the mappings between thematic relations and grammatical functions. only some arguments or adjuncts need to be replaced or added. consequently, kempen states that gapping comes into existence already in the conceptualization phase as an updating operation. in contrast, fcr takes place during grammatical encoding, basically because there the verb concepts do not get updated. (for the detailed argumentation, see kempen, 2009.) kempen (2009:679) argues that bcr actually is not a form of coordinate ellipsis:14 “coordinate structures only afford a suitable playing ground for left deletion as they often give rise to contrastive pairs. the plausibility of viewing bcr as a form of coordinate ellipsis […] is extremely low anyway: the notion of updating entails forward ellipsis only because, by definition, the update comes later than the original structure.” finally, kempen states that posterior clauses of sgf coordinations do not “borrow” a subject np in the course of an incremental updating operation. instead, in the communicative intention underlying an sgf coordination, the speaker assigns several predicates to the referent of the subject np simultaneously. in the following, we assume that in spontaneous speech language production is more often incremental than in writing due to higher time pressure and more severe workspace limitations. 4. 2 “one-clause” versus “multiple-clause” cce as shown in section 3, the relative frequency of gapping is higher in spoken language than in written text whereas the other cce types have lower relative frequencies in spoken than in written language. why? steedman (2000:182) claims that gapping constructions do not result from elision, and thus, there is no need to analyze sequences of non-clausal constituents (in the posterior conjunct) as headless “clauses”. we follow this argument, which is in line with kempen (2009). we assume that gapping does not involve constructing a posterior clause but that the posterior conjunct is built by modifying the anterior one. the whole process of gapping15 works much like speech repair (as outlined in section 4.1) where only the corrections, i.e. the contrastive elements, are uttered. hence, only one clause—more precisely: one superclause16— 14 it is known (see, e.g. hudson, 1976) that bcr-like structures arise also in non-coordination contexts (cf. example (i), where the relative clauses modifying subject and direct object allow ellipsis similar to the pattern in the coordinated nps in example (ii), from the tiger corpus). we have not searched for such cases in spontaneous speech. (i) politicians who fought for chimpanzee rights may well snub those who have fought against chimpanzee rights (ii) ... das richtige gleichgewicht zwischen und the right balance between and die gleichseitigkeit von verschiedenen ... allokationsmechanismen. the equilaterality of different allocation mechanism 15 gapping in coordinate structures is similar to gapping in question answering. producing a gapped answer presupposes that the structure built up during parsing the question is reused in the answer generation process (see, e.g., branigan et al., 2000). 16 according to our definition, gapping in dutch and german requires some syntactic checking, namely the inspection of superclause boundaries (cf. section 2). thus, literal translations into dutch and german of example (i), taken from culicover & jackendoff (2005, p.273), do not work in the intended meaning because the subordinating conjunction blocks any wider scope of gapping. harbusch 325 needs to be kept in the workspace simultaneously, i.e. we claim that gapping is a “oneclause” cce construction. consequently, all constituents needed to construct two clausal gapping conjuncts can be kept in the workspace at the same time. in contrast, in written text, writers are more likely to resort to np coordinations for expressing the same information. stripping, in particular, often becomes recast in this manner, presumably for stylistic reasons. in contrast, applying fcr, bcr or sgf presupposes a grammatical processing window spanning more than one clause (“multiple-clause” cce construction). application of fcr requires comparing the left-hand peripheries of two or more conjoined clauses. likewise, bcr demands comparisons between the right-hand peripheries of the clausal conjuncts. in sgf, the encoding mechanism needs to verify the positions and the referential identity of the subject constituents of consecutive clauses. the “one-clause” property of gapping entails that gapping-type cce structures can survive the limitations due to limited workspace capacity more easily than structures embodying other cce types. finally, we comment on the fact that bcr seems “anti”-incremental. bcr requires the second conjunct to be grammatically planned at least to some extent in parallel with the first one, because the potential elision in the anterior conjunct is licensed by the right-peripheral part of the posterior conjunct. this constraint limits incrementality in spontaneous speech, i.e. the utterance of the final anterior conjunct relies on the posterior conjunct’s syntactic shape. when starting our project, we expected bcr to be absent from spoken corpora. this expectation is even strengthened by alternative ellipsis options, which would affect the second conjunct, imposing even fewer constraints on incremental processing, and erase the same amount of text. for instance, in example (18) above, gapping would elide viertel vor zwölf, könnte, sie and abholen (cf. sentence (25) in section 4.3, which is actually shorter). the same holds for example (17), of which example (19) is the gapping variant. (19) monopole sollen geknackt werden und märkte getrennt monopolies should shattered be and markets split ‘monopolies should be shattered and markets split’ but, why would speakers resort to bcr in a small percentage of the corpus material? there should be another explanation for bcr constructions. often the two conjunct are very short (cf. (20) or (12) above) so that the two conjuncts may simultaneously fit into the workspace. alternatively, the shared right-peripheral constituent(s) might have been added as a kind of afterthought to both already uttered constituents (cf. (21)). however, this claim needs additional evidence from psycholinguistic studies into human cce production under controlled circumstances. (20) … wann ebbe ist und wann flut ist when low-tide and when high-tide is ‘… when is ebb and flow’ (21) … wir wollten uns für ein meeting treffen bzw. eines ausmachen in hannover, wenn we wanted us for a meeting meet and-resp. one negotiate in hannover when ich mich recht entsinne i myself correctly remember ‘…we wanted to have a meeting or negotiate one in hannover if i remember correctly’ 4.3 additional evidence for limited workspace capacity in speaking in the verbmobil treebank of spontaneous spoken utterances, we found many dialogue turns that are underreduced17, i.e., where ellipsis options were not used, or used only par (i) robin believes that everyone pays attention to you when you speak french, and leslie, german as stated in footnote 3, we could not find any token of long-distance gapping in verbmobil (which might have a domain-specific reason: identical subjects in both conjuncts prevail in that corpus) and only nine tokens in cgn2.0. this observation may be viewed as evidence for the claim that having several (super)clauses in the workspace at the same time easily leads to overtaxing the currently available capacity, and to avoidance of ldg structures. 17 underreduction also occurs in written text. however, there it is used intentionally and suits stylistic purposes such as emphasizing the non-reduced constituent—cf. (i) from the tiger corpus. incremental sentence production and clausal coordinate ellipsis 326 tially (see example (22)). for instance, in example (18) above (repeated here as (23)), the sie ‘you’ could be left out resulting in an np coordination (as in (24)) or in stripping/gapping (as in (25), which is alternatively interpretable as including a discontinuous np). such incomplete elisions attest to the possibility of workspace overload when two conjuncts need to be planned simultaneously. (22) ich organisiere die flüge und sie organisieren also dann das hotel i organize the flights and you organize well then the hotel’ ‘i organize the flights and you organize then the hotel’ (23) viertel vor zwölf könnte ich sie oder mein fahrer sie abholen quarter to twelve could i you or my driver you up-pick ‘quarter to twelve, i could pick you up or my chauffeur could do so’ (24) viertel vor zwölf könnte ich oder mein fahrer sie abholen quarter to twelve could i or my chauffeur you up-pick ‘quarter to twelve, my chauffeur or me could pick you up’ (25) ’viertel vor zwölf könnte ich sie abholen oder mein fahrer quarter to twelve, could i you up-pick or my chauffeur in support of workspace limitations as a causal factor, we can also point to typical errors in verbmobil. here, inaccurate structural matches between a remnant in a posterior conjunct and its counterpart in the anterior conjunct occur frequently (cf. fcr of das/da in example (26)). these might also be interpreted as consequences of workspace overload. moreover, telegraphic style and cce production often go together, as in example (27) where gapping has been applied to the incomplete first conjunct. (26) das ist direkt am hauptbahnhof und da kostet das einzelzimmer that is directly at-the main-station and costs the single-room einhundert und neunundzwanzig mark. one-hundred and twenty-nine mark ‘that is directly at the main station and (there) a single room is 129 dm’ (27) hat sogar schwimmbad dabei und hat bar dabei has even swimming-pool included and bar included ‘(it) has included even swimming pool and bar’ in view of these observations, there seems to be a tendency in incremental sentence production to erase from the workspace some or all constituents of an already uttered clause as soon as the conceptualizer to the grammatical encoder embark on a new (super)clause. 4.4 the reasoning in a nutshell a. the five stages of language production are subject to constraints on processing resources, in particular on the workspace for conceptualization and grammatical encoding of communicative intentions. b. the workspace capacity available to the language user will, on average, be smaller in speaking than in writing. c. the smaller the currently available workspace capacity, the higher the likelihood of incremental sentence production. d. when the grammatical encoding mechanism is forced to operate with a limited workspace capacity, it has few opportunities to plan multiple (super)clauses at the same time. e. during speaking, clauses are more often grammatically encoded in isolation than during writing—as a consequence of b and d. f. gapping is a “one-clause” cce structure, whereas fcr, bcr or sgf are “multipleclause” cce structures. g. gapping can survive the limitations due to workspace limitations more easily than structures embodying other cce types—due to e and f. (i) der lehrkurs bei ihm war streng und war gründlich the lesson by him was strict and was intensive ‘his lesson was stern and intensive’ harbusch 327 h. hence, the trend towards more gapping in spoken than in written language production. q.e.d. 5 conclusion summing up, we have presented data verifying, for dutch and german, the data pattern emerging from two corpus studies into the incidence of clausal coordinate ellipsis in spoken and written english. we proposed a new explanation for the relative frequencies of the four cce types we distinguish, based on incremental sentence production and the more limited workspace capacity in speaking than in writing. so far, we have avoided any excursions into incremental parsing as it would lead too far away from the central topic addressed here and would be highly speculative. at least in computational linguistics, parsing elliptical constructions is a difficult problem (see, e.g., kübler et al., 2009), partially due to the fact that neither of the conjuncts might consist of a complete, grammatically correct clause (cf. (17)). how does a human hearer understand cce? let us assume that parsing is highly parallel (in line with new studies on human parsing (cf. boston et al, in press), and that sentence parsing is the inversion of sentence generation (cf. bidirectional grammars such as optimality theory; prince & smolensky, 1993/2004). due to frequency, the hearer is most likely prepared for the gapping variant of cce. thus, the perception of a coordinate conjunction triggers the hypothesis (among a set of parallel hypotheses) that the current workspace content (i.e. the anterior conjunct) will be modified by repair-like substitutes. accordingly, the hearer tries to find a contrastive constituent for every incrementally constructed constituent of the second conjunct. this strategy is doomed to failure as soon as a verb is parsed. consequently, fcr and sgf parsing hypotheses for the second conjunct get activated. however, these cce hypotheses compete with the more highly frequent ellipsis-free hypotheses, which may allow the workspace to be emptied completely. fcr and sgf parsing means, that a prefix (a leftperipheral part) of the first conjunct is assumed to be the valid prefix of the second conjunct as well. (n.b. for sgf, the length of the prefix is exactly one constituent, and by definition it has to be the subject; for fcr, the parser might try a language-specific default instead of trying all prefixes systematically.) accordingly, the workspace can only be emptied partially before the second conjunct is parsed. bcr is triggered by a parsing failure (or a misparse18) of the first conjunct. thus, at the very end of the second conjunct, a re-parse (or a reordering of parallel parsing hypotheses) of the first conjunct is called for. consequently, the working memory should not be emptied before the complete clausal coordination has been parsed. we hasten to that these considerations are highly speculative—even more so because sometimes combinations of cce types (such as example (17)) need to be parsed. gapping benefits from incremental processing particularly, as only the contrastive constituents of the second conjunct have to be checked, and can be uttered in any order. detailed frequency studies into comparatives such as john is prouder of his dog than mary is proud of her cat (see, e.g., lechner, 2004), which are reminiscent of gapping, and arguably similar, are needed. acknowledgement i would like to express my deep gratitude to gerard kempen, who co-authored several related articles on the present topic. i also thank the two anonymous reviewers for their valuable comments. however, none of them can be blamed for remaining errors—any misconceptions and shortcomings in this paper are my own. 18 separable verb prefixes are often elided due to bcr, which changes the meaning considerably (e.g. replacing rufen ‘shout’ by anrufen ‘call up’ as in ich rufe und du schreibst peter an ‘i call up peter and you write to him’. here, the first conjunct could have been parsed as ‘i shout’. incremental sentence production and clausal coordinate ellipsis 328 appendices in the following, we list the nomenclature for coordinate structures in the four different treebanks investigated in the present study. as illustrated by the examples in section 3, the encodings are rather different. therefore, we had to abstract away from the details of the treebank formats. however, due to space restrictions, we cannot enumerate the various search patterns for the corpora. appendix a: alpino in alpino, a node labeled “conj” dominates nodes labeled conjunct (cnj) and coordinator (crd), respectively. edges are not labeled but the node labels denote not only a grammatical function (e.g., su for subject) but, in addition, either a part-of-speech (e.g., vg for the coordinating conjunction such as en ‘and’ in dutch) supplemented with the word in the sentence, or a phrasal category. in order to license clausal coordination, the node labels of conjuncts should be either smain for main clause or ssub for subordinate clause. within a clause, grammatical functions are labeled: subject as su, the finite verb as hd and the verb complement as vc, etc. elisions are encoded by coreferential indices at nodes. the remnant and its corresponding empty leaf node get identical numerical indices. which of the two is expanded can only be identified in the sentence itself. we wrote a java program to automatically retrieve all clausal coordinations with and without such indices, and manually assigned cce types to the coordinations containing an index. the detailed proportions of different cce types are reported in section 3.1.3. appendix b: cgn2.0 in cgn2.0, edges are labeled with the grammatical functions. edge labels are represented as squared boxed at the edges in tree diagrams provided by tigersearch (see, e.g., figure 2 in section 3.1.2). nodes carry either lexical information or a phrasal category. in coordination, a node labeled “conj” dominates edges labeled with cnj (conjunct) and crd (coordinator). in order to represent clausal coordination, cnj-edges end at smain or ssub nodes, denoting main and subordinate clauses, respectively. within a clause, grammatical functions are represented at edges. for instance, su = subject, prec = predicate, and hd = head. cgn2.0 represents elided constituents by so-called secondary edges between the remnant and the root node where the constituent is supposed to be elided. secondary edges are represented by curved edges in tree diagrams provided by tigersearch. the grammatical function of the elided constituent labels the secondary edge. tigersearch provides a specific search pattern for secondary edges (“>~”), possibly expanded by the grammatical function annotated at the edge. this allows accurate extraction of syntactic trees embodying various types of coordinate ellipsis. in cgn2.0, secondary edges are not directed (cf. tiger). nevertheless, the remnant can be determined by a search pattern in tigersearch. for both end nodes of a secondary edge, the grammatical function of the secondary edge is compared to the edge label of the node’s incoming edge. the matching one is the remnant. at the node the constituent is supposed to be elided, this function does not exist because it is supposed to be elided. one exception is the modifier function that may occur repeatedly. these cases were inspected manually. the detailed proportions of different cce types are reported in section 3.1.3. appendix c: tiger in tiger, edges are labeled with grammatical functions—similarly to cgn2.0. edge labels are represented as squared boxed at the edges in tree diagrams provided by tigersearch (see, e.g., figure 3 in section 3.2.1). nodes carry either lexical information or a phrasal category. in tiger’s syntactic trees, different types of coordination are distinguished. the relevant ones for cce are: • cs: coordinated finite clauses, harbusch 329 • cvp: coordinated verb phrases (nonfinite clauses), and • cvz: coordinated infinitival clauses (vps) with the verb preceded by zu ‘to’ (as in zu tun ‘to do’). all those nodes dominate edges labeled cj for conjunct and cd for coordinating conjunction. such edges end in s, vp or vz nodes. abbreviations for edge labels in a clause are, e.g., sb = subject, hd = head, oc = object complement. like in cgn2.0, tiger represents elided constituents by so-called secondary edges between the remnant and the root node where the constituent is supposed to be elided. secondary edges are represented by curved edges in tigersearch tree diagrams. the grammatical function of the elided constituent labels the secondary edge. in tiger, secondary edges are directed from the remnant to the root node of which the constituent is supposed to be a child. there are forward and backward secondary edges. all backward secondary edges point out bcr. forward secondary edges require more finegrained search patterns. for instance, the edge label head at a forward secondary edge points to gapping instances. the detailed proportions of different cce types are reported in section 3.2.3. appendix d: verbmobil in verbmobil, three node types are differentiated. nodes carry either lexical information or a phrasal category—similarly to cgn2.0 and tiger. moreover, topological field information is spelled out in so-called field nodes:  vf = vorfeld ‘forefield’ with the special types listed in table d-1 in more detail: mvc and mvcn which occur in clauses with a complementizer (c), viz. as cmvc and cmvcn,  mf = mittelfeld ‘midfield’,  nf = nachfeld ‘endfield’,  lk = linke satzklammer ‘left sentence bracket’, and  vc = verb complex. coordination type composition pattern mn mf+nf cm c+mf mvc/cmvc mf+vc/c+mf+vc mvcn/cmvcn mf+vc+nf/c+mf+vc+nf lkm lk+mf lkmn lk+mf+nf lkmvc lk+mf+vc lkmvcn lk+mf+vc+nf lkn lk+mf+nf vcn vc+nf vlkm vf+lk+mf vlkmvc vf+lk+mf+cv table d-1. node labels of conjuncts denoting clausal coordination in verbmobil. a coordination type (in column 1) represents a node label. the composition pattern in column 2 spells out which subtree types (i.e. node labels at the root of the subtree) it dominates. edges are labeled in the verbmobil corpus. such labels are represented as squared boxed at the edges in tigersearch tree diagrams (see, e.g., figure 4 in section 3.2.2). all edges to field nodes are labeled with a dash. within a clause, grammatical functions are represented as edge labels. for instance, on = subject and oa = direct object. clausal coordination is spelled out at the level of field nodes. a node labeled fkoord, which represents complex field coordination, dominates nodes of the types listed in column one of table d-1. the second column of table d-1 describes the composition pattern, i.e. the nodes which have to occur as children of that node. the symbol “+’ separates the children’s types in the list. incremental sentence production and clausal coordinate ellipsis 330 notice that, unlike alpino, cgn2.0 and tiger, verbmobil provides no specification relating remnant and elided constituent. accordingly, we first retrieved sentences with clausal coordination. within this set, all sentences were manually classified. for instance, mn could be a case of fcr or sgf. the detailed proportions of different cce types are reported in section 3.2.3. references arnold, j.e., wasow, t., asudeh, a. & alrenga, p. 2004. avoiding attachment ambiguities: the role of constituent ordering. journal of memory and language, 51:1, 55–70. bell, a. 1984. language style as audience design. language in society 13(2): 145–204. borsley, r.d. 2005. against conjp. lingua, 115, 461–482. boston, m.f., hale, j.t., vasishth, s. & kliegl, r. in press. parallel processing and sentence comprehension difficulty. language and cognitive processes. branigan, h.p, pickering, m.j. & cleland, a.a. 2000. syntactic co-ordination in dialogue. cognition, 75, b13–b25. brants, s., dipper, s., eisenberg, p., hansen-schirra, s., könig, e., lezius, w., rohrer, c., smith, g. & uszkoreit, h. 2004. tiger: linguistic interpretation of a german corpus. research on language and computation, 2, 597–620. brennan, s.e. & hanna, j.e. 2009. partner-specific adaptation in dialog. topics in cognitive science 1, 274–291. culicover, p.w. & jackendoff, r. 2005. simpler syntax. oxford: oxford university press. ferreira, f. 2000. syntax in language production: an approach using tree-adjoining grammars. in wheeldon, l. (ed.). aspects of language production. cambridge ma: mit press. frank, r. & badecker, w. 2001. modeling syntactic encoding with incremental tree adjoining grammar. in proceedings of the 14th annual cuny conference on human sentence processing, philadelphia pa. greenbaum, s. & nelson, g. 1999. elliptical clauses in spoken and written english. in collins, p. & lee, d. (eds.). the clause in english. amsterdam: benjamins. harbusch, k. & kempen, g. 2007. clausal coordinate ellipsis in german: the tiger treebank as a source of evidence. in proceedings of the 16th nordic conference of computational linguistics (nodalida 2007), tartu, estonia. harbusch, k. & kempen, g. 2009a. a treebank study of clausal coordinate ellipsis in spoken and written language. in proceedings of the 15th annual conference on architectures and mechanisms of language processing (amlap2009), barcelona, spain. harbusch, k. & kempen, g. 2009b. clausal coordinate ellipsis and its varieties in spoken and written german: a study with the tüba-d/s treebank of the verbmobil corpus. in proceedings of the 8th international workshop in treebanks and linguistic theories (tlt8), milano, italy. harbusch, k. & kempen, g. 2009c. generating clausal coordinate ellipsis multilingually: a uniform approach based on postediting. in proceedings of the 12th european workshop on natural language generation (enlg 2009), athens, greece. hartmann, k. 2002. right node raising and gapping: interface conditions on prosodic deletion. philadelphia/amsterdam: john benjamins. haywood, s., pickering, m.j. & branigan, h.p. 2005. do speakers avoid ambiguity in dialogue? psychological science, 16, 362–366. hinrichs, e., kübler, s., naumann, k., telljohann, h. & trushkina, j. 2004. recent developments in linguistic annotation of the tüba-d/z treebank. in proceedings of the third workshop on treebanks and linguistic theories (tlt3), tübingen, germany. höhle, t.n. 1983. subjektlücken in koordinationen. unpublished manuscript, university of cologne. höhle, t.n. 1986. der begriff “mittelfeld”, anmerkungen über die theorie der topologischen felder. in schöne, a. (ed), kontroversen, alte und neue: akten des 7. internationalen germanisten-kongresses, göttingen 1985, band 3, 329–340. tübingen: niemeyer. harbusch 331 höhle, t.n. 1990. assumption about asymmetric coordination in german. in mascaró, j. & nespor, m. (eds.). grammar in progress: glow essays for henk van riemsdijk, 221–235. dordrecht: foris. hudson, r. 1976. conjunction reduction, gapping, and right-node-raising. language 52:535–562. kathol, a. 2001. linearization vs. phrase structure in german coordination constructions. cognitive linguistics, 10:4, 303–342. kempen, g. 2009. clausal coordination and coordinate ellipsis in a model of the speaker. linguistics, 47, 653–696. kempen, g. & hoenkamp, e. 1982. incremental sentence generation: implications for the structure of a syntactic processor. in horecky, j. (ed.). proceedings of the ninth international conference on computational linguistics (coling), prague. amsterdam: northholland. kempen, g. & hoenkamp, e. 1987. an incremental procedural grammar for sentence formulation. cognitive science: a multidisciplinary journal, 11: 2, 201–258. könig, e. & lezius, w. 2003. the tiger language: a description language for syntax graphs, formal definition. tech. rep. ims, university of stuttgart, germany. kübler, s., hinrichs, e., maier. w. & klett, e. 2009. parsing coordinations. proceedings of eacl 2009, athens, greece. lechner, w. 2004. ellipsis in comparatives. berlin: walter de gryter. lemnitzer, l. & zinsmeister, h. 2006. korpuslinguistik: eine einführung. tübingen: narr studienbücher. levelt, w.j.m. 1989. speaking: from intention to articulation. cambridge, ma: mit press. meyer, c.f. 1995. coordination ellipsis in spoken and written american english. language sciences, 17, 241–169. prince, a. & smolensky, p. 1993/2004. optimality theory: constraint interaction in generative grammar. rutgers university and university of colorado at boulder: technical report ruccstr-2, available as roa 537-0802. revised version published by blackwell, 2004. purver, m. & kempson, r. 2004. incremental context-based generation for dialogue. in proceedings of the 3rd international conference on natural language generation (inlg04), careys manor, uk. reich, i. 2007. toward a uniform analysis of short answers and gapping. in schwabe, k. & winkler, s. (eds.). on information structure, meaning and form. amsterdam: john benjamins. 467–484. reich, i. 2008. from discourse to ‘‘odd coordinations’’: on asymmetric coordination and subject gaps in german. in fabricius-hansen, c. & ramm, w. (eds.). ‘subordination’ versus ‘coordination’ in sentence and text: a cross-linguistic perspective. amsterdam: benjamins. sag, i.a., wasow, t. & bender, e.m. 2003. syntactic theory: a formal introduction. stanford ca: csli publications [second edition.] seginer, y. 2007. fast unsupervised incremental parsing. in proceedings of the 45th annual meeting of the association for computational linguistics (acl), prague, czech republic. steedman, m. 1990. gapping as constituent coordination. linguistics and philosophy, 13, 207–263. steedman, m. 2000. the syntactic process. cambridge ma: mit press. stegmann, r., telljohann, h. & hinrichs, e. 2000. stylebook for the german treebank in verbmobil. saarbrücken: dfki rep. 239. te velde, j.r. 2006. deriving coordinate symmetries. amsterdam: benjamins. uit den boogaart, p.c. 1975. woordfrequenties in geschreven en gesproken nederlands. werkgroep frequentie-onderzoek van het nederlands. utrecht: oosthoek, scheltema & holkema. van der beek, l., bouma, g., malouf, r. & van noord, g.-j. 2002. the alpino dependency treebank. in computational linguistics in the netherlands (clin 2001). amsterdam: rodopi. incremental sentence production and clausal coordinate ellipsis 332 van eerten, l. 2007. over het corpus gesproken nederlands. nederlandse taalkunde, 12:3, 194–215. van oirsow, r.r. 1987. the syntax of coordination. london: croom helm. wahlster, w. (ed.). 2000. verbmobil: foundations of speech-to-speech translation. berlin: springer. wunderlich, d. 1988. some problems of coordination in german. in reyle, u. & rohrer c. (eds.). natural language parsing and linguistic theories, dordrecht: reidel. zinsmeister, h. 2006. treebank data as linguistic evidence? coordination in tüba-d/z. pre-proceedings of the international conference on linguistic evidence, tübingen. microsoft word rhetoricagreement_galitskyd&d_r3_insertedmetadata-1.docx dialogue & discourse 8(2) 167-205 doi: 10.5087/dad.2017.208 ©2017 boris galitsky this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). discovering rhetorical agreement between a request and response boris galitsky boris.galitsky@oracle.com oracle corp. redwood shores ca usa editor: barbara di eugenio submitted 11/2016; accepted 12/2017; published online 12/2017 abstract to support a natural conversation flow between humans and automated agents, the rhetorical structure of each message must be analyzed. we classified pairs of text paragraphs as either appropriate or inappropriate for one to follow the other based on considerations of both topic and communicative discourse. to represent a multi-sentence message with respect to how it should follow a previous message in a conversation or dialogue, we built an extension of a discourse tree. an extended discourse tree is based on a discourse tree for rst relations, with labels for communicative actions and additional arcs for anaphora and ontology-based relations for entities. we refer to such trees as communicative discourse trees (cdts). we explored the syntactic and discourse features indicative of correct versus incorrect request-response or question-answer pairs. two learning frameworks were used to recognize such correct pairs: deterministic, nearestneighbor learning of cdts as graphs and tree kernel learning of cdts, in which a feature space of all cdt sub-trees is subject to svm learning. we formed the positive training set from correct pairs obtained from yahoo answers, social networks, corporate conversations (including enron emails) customer complaints and interviews by journalists. the corresponding negative training set was artificially created by attaching responses for different inappropriate answers that nevertheless cover the topics of questions. the evaluation showed that it is possible to recognize valid pairs in 70% of cases in the domains of weak request-response agreement and 80% of cases in the domains of strong agreement. recognition of such pairs is essential to support automated conversations. these accuracy rates are comparable to the benchmark task of classifying discourse trees as either valid or invalid. they are also comparable to the classification of multisentence answers in factoid question-answering systems. we conclude that learning rhetorical structures in the form of cdts is a key source of data to support answering complex questions. 1 introduction in recent years, the development of chatbots for answering questions and performing user requests has become very popular. a broad range of relevant technologies, including compositional semantics, have been developed to support these systems in the context of simple, short queries and replies. the accuracy of discourse parsing results has dramatically increased. rhetorical parsers are now capable of building discourse structures for longer queries, requests and answers (subba and di eugenio 2009). at present, the issue of how a question-answering, discovering rhetorical agreement 168 dialogue management, or recommendation system (galitsky 2013) can leverage rich discourserelated information in a structural form has not yet been extensively addressed. during the last two decades, research in the field of dialogue systems has experienced increasing growth (wilks 1999, hiraoka et al., 2013). a number of formal systems representing various aspects of dialogue have been proposed (traum and hinkelman 1992, blaylock et al., 2003, popescu-belis 2005, popescu 2007, visser et al., 2014). however, the design and optimization of these systems does not entail simply combining language processing systems such as parsers, part-of-speech taggers, intention models, rhetorical components and natural language generation systems. it also requires the development of dialogue strategies, considering at minimum the performances of these systems, the nature of the task (such as form-filling, tutoring, robot control, information and advice requests, social promotion or database search/browsing), and user behavior, such as cooperativeness and expertise (aina et al., 2017). due to the great variability in these factors, deploying manual, handcrafted dialogue designs is very difficult. for these reasons, statistical machine learning methods to support dialogue have been a leading focus of research for the last decade. a request can have an arbitrary rhetorical structure as long as the topic of this request or question is clear to its recipient. a response on its own can also have an arbitrary rhetorical structure. however, these structures should be correlated when the response is appropriate to the request. in this study, we focused on the computational measurement of the agreement between the logical rhetorical structure of a request or question and that of the corresponding response or answer. we formed a number of representations for a request-response (rr) pair, learned them and solved an rr classification problem, identifying a pair as either valid (a correct answer or response) or invalid pairs. traditionally, computational models of communicative discourse are based on analyses of speaker intent (allen and perrault, 1980; grosz and sidner, 1986; heller et al., 2013). in recent years, deep learning based models for dialogue learning have also become popular (zhae et al., 2017). a requester has certain goals, and communication results from a planning process to achieve these goals. the requester will form intentions based on these goals and then act on these intentions, producing utterances. upon hearing the utterance, the responder will then reconstruct a model of the requester’s intentions. however, this family of approaches is limited to providing an adequate account of adherence to discourse conventions in dialogue. in addition to the analysis of intent, the notion of grounding has proved to be fruitful in modeling dialogues. language needs grounding in the nonlinguistic world and in the practices of language users. this grounding is formed and controlled in the course of dialogue through conversational grounding (schlangen, 2016), which is the interactive process through which interlocutors form an understanding of each other, grounding justification (the ability to explain and provide reasons for each agents’ language use), and grounding adaptation (the ability to accept corrections and modify how language is employed). when answering a question formulated as a phrase or a sentence, the answer must address the topic of the question. given an initial utterance as an explicit or implicit question, its answer is expected not only to maintain a topic but also to match the generalized epistemic state of this utterance. for example, when a person is looking to sell an item with features, the search results should not only contain these features but also indicate intent to buy. when a person is looking to share knowledge about an item, the search results should contain an intent to receive a recommendation. when a person asks for an opinion about a topic, the response should be to share an opinion about this topic instead of expressing another request for an opinion. modern dialogue management and automated email-answering systems have achieved good accuracy with http://www.uni-bielefeld.de/(en)/lili/personen/dschlangen/ discovering rhetorical agreement 169 regard to maintaining the topic, but maintaining the communication discourse is a much more difficult problem. the syntactic structure of a simple question is correlated with that of an answer. this structure is helpful for finding the best answer in the passage re-ranking problem. it has been shown that using syntactic information to improve search relevance is helpful in addition to keyword frequency (tf*idf) analysis and other keyword statistical methods, such as lda (blei et al., 2003). selecting a most suitable answer, not only through keywords but also by judging how the syntactic structure of a question, including a focus on wh-words, is reflected in an answer, has been proposed (moschitti and quarteroni 2011; galitsky 2013). following along the lines of these studies, we have adapted this consideration at the phrase and sentence levels and applied it to the level of discourse. to represent the linguistic features of text, we used the two following sources: 1. rhetorical relations between the parts of the sentences, obtained as a discourse tree. we relied on rhetorical structure theory (rst, mann and thompson 1988) and deployed rhetorical parsers (joty et al., 2013; surdeanu et al., 2015) to build these discourse trees. 2. speech acts and communicative actions, obtained as verbs from the verbnet resource (verb signatures with instantiated semantic roles). these were attached to rhetorical relations as labels for the arcs of communicative discourse trees. it turns out that having only 1) or only 2) is insufficient for recognizing correct rr pairs. however, the combination of these sources is sufficient. rhetorical structure theory models the logical organization of text, a structure employed by a writer relying on relations between parts of text. rst simulates text coherence by forming a hierarchical connected structure of texts via discourse trees. rhetorical relations are split into coordinate and subordinate classes; these relations hold across two or more text spans and therefore implement coherence. these text spans are called elementary discourse units (edus). clauses in a sentence and sentences in a text are logically connected by the author. the meaning of a given sentence is related to that of the previous and following sentences. this logical relation between clauses is called the coherence structure of the text. rst is one of the most popular theories of discourse and is based on tree-like discourse structures called discourse trees (dts). the leaves of a dt correspond to edus, the contiguous atomic text spans. adjacent edus are connected by coherence relations (e.g., attribution, sequence), forming higher-level discourse units. these units are then also subject to this relation-linking. edus linked by a relation are then differentiated based on their relative importance: nuclei represent the core parts of the relation, whereas satellites represent the peripheral ones. the goal of this research was to extend the notion of question/answer relevance to the rhetorical relevance of general request/response pairs for broader dialogue support. we now proceed to an example for an agreement between a question and answer. for the question “what does the investigative committee of the russian federation do” there are two answers: 1) mission statement. “the investigative committee of the russian federation is the main federal investigating authority which operates as russia's anti-corruption agency and has statutory responsibility for inspecting the police forces, combating police corruption and police misconduct, is responsible for conducting investigations into local authorities and federal governmental bodies.” 2) an answer from the web. “investigative committee of the russian federation is supposed to fight corruption. however, top-rank officers of the investigative committee discovering rhetorical agreement 170 of the russian federation are charged with creation of a criminal community. not only that, but their involvement in large bribes, money laundering, obstruction of justice, abuse of power, extortion, and racketeering has been reported. due to the activities of these officers, dozens of high-profile cases including the ones against criminal lords had been ultimately ruined” (crimerussia 2016). the choice of answers depends on context. rhetorical structure allows differentiation between “official,” “politically correct,” template-based answers and “actual,” “raw,” “reports from the field,” “controversial” ones (fig. 1 a and b). sometimes, the question itself can give a hint about which category of answers is expected. if a question is formulated as a factoid or definitional question without a second meaning, then the first category of answers is suitable. otherwise, when a question has the meaning, “tell me what it really is,” the second category is appropriate. in general, if we can extract a rhetorical structure from a question, it is easier to select a suitable answer that would have a similar, matching, or complementary rhetorical structure. the discourse trees of an official answer are based on elaboration and joints, which are neutral in terms of any controversy that a text might contain (fig. 1a). at the same time, the raw answer includes the contrast relation. this relation refers to the contrast between what an agent is expected to do and what the agent was discovered to have done. fig. 1a: dt for the official answer in the future (but not currently), discourse parsers should be capable of differentiating between discovering rhetorical agreement 171 “what does this entity do” and “what does this entity really do” by identifying a contrast relation between ‘do’ and ‘really do.’ after such a relation is established, it would be easier to make a decision as to whether an answer with or without contrast is suitable. hence, when rhetorical parsing improves, the same dt learning machinery would deliver more accurate rhetorical agreement results. fig. 1b: dt for the raw answer (from the web) we formulated the main problem of this study to classify pairs of rr texts as correct or incorrect. this problem can be formulated with or without considering relevance, which we intend to treat orthogonally to how the rhetorical structure of a request agrees with the rhetorical structure of a response. rhetorical agreement may be present or absent in an rr pair, and the same applies to the relevance agreement. some methods of measuring rhetorical agreement include considerations of relevance agreement whereas others do not. the idea of measuring the similarity between question-answer pairs for question-answering instead of question-answer similarity turned out to be insightful (moschitti and quarteroni, 2011). the classifier for correct versus incorrect answers processes two pairs at a time, and , and compares q1 with q2 and a1 with a2, producing a combined similarity score. this comparison allows the classifier to determine whether an unknown question/answer pair contains a correct answer by assessing its distance from another question/answer pair with a known label. in particular, an unlabeled pair is processed such that, rather than ‘‘guessing” correctness based on words or structures shared by q2 and a2, both q2 and a2 are compared with their corresponding components, q1 and a1 of the labeled pair , on the grounds of such words or structures. because this approach targets a domain-independent answer classification, only the structural cohesiveness between a question and answer, not the ‘meaning’ of an answer, can be leveraged. to form a training set for this classification problem, we included actual rr pairs in the positive dataset and arbitrary or low-relevance and -appropriateness rr pairs in the negative dataset. for the positive dataset, we selected various domains with distinct acceptance criteria for which an answer or response was suitable for the question. such acceptance criteria are low for community question-answering, automated question-answering, automated and manual customer support systems, social network communications and writing by individuals such as consumers about their experience with products such as reviews and complaints. rr acceptance criteria are higher in scientific texts, professional journalism, health and legal documents in the form of discovering rhetorical agreement 172 faqs, and in professional social networks, such as stackoverflow.com. 2 communicative discourse trees communicative discourse trees (cdts) are designed to combine rhetorical information in the form of dt with speech act structures using arcs labeled with expressions for communicative actions. these expressions are logical predicates expressing the agents involved in the respective speech acts and the arguments of their communicative actions. the arguments of logical predicates are formed in accordance with their respective semantic roles as proposed by a framework such as verbnet (kipper et al., 2008, palmer 2009). the purpose of adding these labels is to incorporate the speech act-specific information into dts so that their learning occurs over a richer feature set than just rhetorical relations and the syntax of elementary discourse units (edus). the objective here is to incorporate all information about how an author’s thoughts are organized and communicated irrespective of the subjects of these thoughts. our second example of a rhetorical agreement in a conversation is a dispute between three parties concerning the cause of the downing of malaysia airlines flight 17 (wikipedia 2016). we built an rst representation of the arguments being communicated and observed if a discourse tree is capable of indicating whether a paragraph communicates both a claim and an argumentation that backs it up. we then explored what needs to be added to the dt representation so that it is possible to judge whether it expresses an argumentation pattern or not. surdeanu et al.’s (2015) computation and visualization system for dt was used. three conflicting agents, dutch investigators, the investigative committee of the russian federation, and the self-proclaimed donetsk people's republic, exchanged their opinions on the matter. it is a controversial conflict in which each party does all it can to blame its opponent. to sound more convincing, each party does not just produce its claim but formulates it in such a way as to rebuff the claims of its opponent. to achieve this goal, each party attempts to match the style and discourse of the opponents’ claims. “dutch accident investigators say that evidence points to pro-russian rebels as being responsible for shooting down plane. the report indicates where the missile was fired from and identifies who was in control of the territory and pins the downing of mh17 on the pro-russian rebels.” (fig. 2a) discovering rhetorical agreement 173 fig. 2a: the claim of the first agent, dutch accident investigators. the regular nodes of cdts are rhetorical relations, and the terminal nodes are the elementary discourse units (phrases, sentence fragments) that are the arguments of these relations. certain arcs of cdts are labeled with the expressions for communicative actions, including the actor agent and the argument of these actions (what is being communicated). for example, the nucleus nodes for elaboration relation (on the left) are labeled with say (dutch, evidence), and the satellites, with responsible (rebels, shooting_down). these labels are not intended to express that the arguments of edus are evidence and shooting_down but instead to match this cdt with others to find the similarity between them. in this case, simply linking these communicative actions by a rhetorical relation but not providing information of communicative discourse would be too limited to represent a structure of what is being communicated and how. requiring an rr pair to have the same or coordinated rhetorical relations is too weak; thus, an agreement of the cdt labels for arcs on top of matching nodes is required. “the investigative committee of the russian federation believes that the plane was hit by a missile, which was not produced in russia. the committee cites an investigation that established the type of the missile.” (fig. 2b) fig. 2b: the claim of the second agent, the committee. “rebels, the self-proclaimed donetsk people's republic, deny that they controlled the territory from which the missile was allegedly fired. it became possible only after three months after the tragedy to say if rebels controlled one or another town.” (fig. 2c). a response cannot be arbitrary. it has to talk about the same entities as the original text. it has to back up its disagreement with its estimates and sentiments about these entities, and about actions of these entities. we attempt to encode the structure of agreement between rr pairs via cdt in a domain-independent manner, only in the space of communication and its style, and the logical flow of a conversation irrespectively of the nature of these entities. discovering rhetorical agreement 174 fig. 2c: the claim of the third agent, the rebels. what we see from this example is that the replies of the peers need to reflect, to mimic the communicative discourse of the first utterance, the seed. as a simple observation, since the first agent uses attribution to communicate his claims, the other agents have to follow suit and either provide their own attributions or attack the validity of attribution of the proponent, or both. to capture a broad variety of features for how communicative structure of the seed message needs to be retained in consecutive messages, we will learn the pairs of respective cdts. to verify the rr agreement, discourse relations are necessary but insufficient, and speech acts (communicative actions) are necessary but insufficient as well. for the paragraph from the previous example, we need to know the discourse structure of interactions between agents, and what kind of interactions they are. we don’t need to know domain of interaction (here, military conflicts), the topics of these interaction, what the entities are. but we need to take into account mental, domain-independent relations between them. towards the end of this section we give a formal definition of cdt. cdt is a dt with labels for arcs which are the verbnet expressions for verbs which are communicative actions. the argument of verbs are substituted from text according to verbnet frames. for the details of dts we refer the reader to (joty et al., 2016), and for verbnet frames – to the section on communicative actions below and then to (kipper et al., 2008). we conclude this section with a note that a cdt required to learn rr agreement is an extension of a traditional discourse tree. it allows us to make rst relations labeled with communicative actions. cdt is a reduction of what is called parse thicket (galitsky et al., 2015), a combination of parse trees for sentences with discourse-level relationships between words and parts of the sentence in one graph (fig. 3). the straight edges of this graph are syntactic relations, and curvy arcs – discourse relations, such as anaphora, same entity, sub-entity, rhetorical relation and communicative actions. this graph includes much richer information than just a combination of parse trees for individual sentences would. as well as cdts, parse thickets can be generalized at the level of words, relations, phrases and sentences. 3 representing rhetorical relations and communicative actions to compute similarity between abstract structures, two approaches are frequently used: 1) represent these structures in a numerical space, and express similarity as a number. this is a statistical learning approach discovering rhetorical agreement 175 2) use a structural representation, without numerical space, such as trees and graphs, and express similarity as a maximal common sub-structure. we refer to such operation as generalization. this is an inductive learning approach. to conduct feature engineering, we will compare both these approaches in the domain of this study. the representation machinery and learning settings are different, but the classification accuracies can be compared. 3.1 greedy representations for a q/a pair we now proceed to another examples of q/a pair and its representation trying to involve as detailed linguistic information as possible. a greedy approach to representing linguistic information about this pair and a map from q to a is shown in fig. 3. we combine parse trees for sentences with pragmatic and discourse-level relationships between words and parts of the sentence in one graph, called parse thicket. the parse thicket for q is shown as a connected graph on the top, and for a – on the bottom. we complemented the edges for syntactic relations obtained and visualized with the stanford nlp system (manning et al., 2014). (recasens et al., 2013 and lee at al 2013) was used for coreference resolution. the arcs for pragmatic and discourse relations, such as anaphora, same entity, sub-entity, rhetorical relation and communicative actions correspondence between parse thickets for q and a are manually drawn. labels embedded into arcs denote the syntactic relations. lemmas are written below the boxes for the nodes, and parts-of-speech are written inside the boxes. this graph includes much richer information than just a combination of parse trees for individual sentences would. navigation through this graph along the edges for syntactic relations as well as arcs for discourse relations allows to transform a given parse thicket into semantically equivalent forms for matching with other parse thickets, performing a text similarity assessment task. to form a complete formal representation of a paragraph, we attempt to express as many links as possible: each of the discourse arcs produces a pair of thicket phrases that can be a potential match. topical similarity between q and a can be expressed as common sub-graphs of parse thickets. the higher the number of common graph nodes, the higher the similarity. for rhetorical agreement, the common sub-graph does not have to be large. however, rhetorical relations and communicative actions of the seed and response are correlated and a correspondence is required. our example for a pair of parse thickets for a question and its answer is obtained for the following: q: i just had a baby and it looks more like the husband i had my baby with. however it does not look like me at all and i am scared that he was cheating on me with another lady and i had her kid. this child is the best thing that has ever happened to me and i cannot imagine giving my baby to the real mom. a: marital therapists advise on dealing with a child being born from an affair as follows. one option is for the husband to avoid contact but just have the basic legal and financial commitments. another option is to have the wife fully involved and have the baby fully integrated into the family just like a child from a previous marriage. discovering rhetorical agreement 176 fig. 3: parse thickets for the question and answer and the mapping between them. discovering rhetorical agreement 177 3.2 communicative actions and their generalization arguments are mostly communicated in the mental world, and learning communicative actions is a key to expressing and understanding arguments. computational verb lexicons are key to supporting acquisition of entities for actions, and a rule-based form to express their meanings. verbs express the semantics of an event being described as well as the relational information among participants in that event, and project the syntactic structures that encode that information. verbs, and in particular the ones for communicative actions, are also highly variable, displaying a rich range of semantic behaviors. verb classification helps a learning systems to deal with this complexity by organizing verbs into groups that share core semantic properties. verbnet is one such lexicon, which identifies semantic roles and syntactic patterns characteristic of the verbs in each class and makes explicit the connections between the syntactic patterns and the underlying semantic relations that can be inferred for all members of the class. each syntactic frame in a class has a corresponding semantic representation that details the semantic relations between event participants across the course of the event. verbnet is a good source of information on verbs in general and communicative actions in particular. let us consider the verb amuse. there is a cluster of similar verbs that have a similar structure of arguments (semantic roles) such as amaze, anger, arouse, disturb, irritate, and other. the roles of the arguments of these communicative actions are as follows: • experiencer (usually, an animate entity) • stimulus • result the frames (the classes of meanings differentiated by syntactic features for how this verb occurs in a sentence) are as follows (np – noun phrase, n – noun, v – communicative action, vp – verb phrase, adv adjective). below is a set of definitions for the verb amuse (fig. 3a). for this example, the information for the class of verbs amuse is available at http://verbs.colorado.edu/verb-index/vn/amuse-31.1.php#amuse-31.1 we now show how communicative actions are split into clusters (table 1). the purpose of defining the similarity of two verbs as an abstract verb-like structure is to support inductive learning tasks such as rhetorical agreement assessment. in statistical machine learning, similarity is expressed as a number. a drawback of this learning approach is that, by representing linguistic feature space as numbers, one loses the ability to explain the learning feature underlying a classification decision. after a feature has been expressed as a number and combined with other numbers, it is difficult to interpret it. in inductive learning, when a system performs classification tasks, it identifies a particular verb or verb-like structure that is determined to cause the target feature (such as rhetorical agreement). in contrast, the statistical and deep learning-based family of approaches simply delivers decisions without explanation. in statistical learning, the similarity between two verbs is a number. in inductive learning, it is an abstract verb with attributes shared by these two verbs (attributes present in one but absent in the other are not retained). the resultant structure of similarity computation can be subjected to further similarity computation with another structure of that type or with another verb. for verb similarity computation, it is insufficient to indicate only that two verbs belong to the same class: all common attributes must occur in the similarity expression. http://verbs.colorado.edu/verb-index/vn/amuse-31.1.php#amuse-31.1 discovering rhetorical agreement 178 np v np example: "the teacher amused the children." syntax: stimulus v experiencer clause: amuse(stimulus, e, emotion, experiencer): cause(stimulus, e), emotional_state(result(e), emotion, experiencer). np v adv-middle example: "small children amuse quickly." syntax: experiencer v adv clause: amuse(experiencer, prop): property(experiencer, prop), adv(prop). np v np-pro-arb example "the teacher amused." syntax stimulus v amuse(stimulus, e, emotion, experiencer): cause(stimulus, e), emotional_state(result(e), emotion, experiencer). np.cause v np example "the teacher's dolls amused the children." syntax stimulus <+genitive> ('s) v experiencer amuse(stimulus, e, emotion, experiencer): cause(stimulus, e), emotional_state(during(e), emotion, experiencer). np v np adj example "this performance bored me totally." syntax stimulus v experiencer result amuse(stimulus, e, emotion, experiencer): cause(stimulus, e), emotional_state(result(e), emotion, experiencer), pred(result(e), experiencer). fig. 3a: definitions for the verb amuse discovering rhetorical agreement 179 verbs with predicative complements appoint, characterize, dub, declare, conjecture, masquerade, orphan, captain, consider, classify verbs of perception see, sight, peer verbs of psychological state amuse, admire, marvel, appeal desire want, long judgment verbs judge assessment assess, estimate verbs of searching hunt, search, stalk, investigate, rummage, ferret verbs of social interaction correspond, marry, meet, battle verbs of communication transfer(message), inquire, interrogate, tell, manner(speaking), talk, chat, say, complain, advise, confess, lecture, overstate, promise avoid avoid measure register, cost, fit, price, bill aspectual begin, complete, continue, stop, establish, sustain table 1: verbnet classes of the verbs for communicative actions hence similarity between two communicative actions a1 and a2 is defined as an abstract verb which possesses the features which are common between a1 and a2. we first provide an example of the generalization of two very similar verbs: agree ^ disagree = verb(interlocutor, proposed_action, speaker), where interlocutor is the person who proposed the proposed_action to the speaker and to whom the speaker communicates their response. proposed_action is an action that the speaker would perform if they were to accept or refuse the request or offer, and the speaker is the person to whom a particular action has been proposed and who responds to the request or offer made. when verbs are not that similar, a subset of the arguments remain: agree ^ explain = verb(interlocutor, *, speaker). further examples of generalizing verbs and communicative actions are available in (galitsky et al., 2009). the main observation concerning communicative actions in relation to finding text similarity is that their arguments need to be generalized in the context of these actions and that they should not be generalized with other “physical” actions. hence, we generalize the individual occurrences of communicative actions together with their arguments. we also generalize sequences of communicative actions representing dialogs against other such sequences of similar dialogs. this way we represent the meaning of an individual communicative action as well as the dynamic discourse structure of a dialogue (in contrast to its static structure reflected via rhetorical relations). the idea of generalization of compound structural representation is that generalization discovering rhetorical agreement 180 happens at each level. the verb itself of a communicative action is generalized with another verb, and its semantic roles are generalized with their respective semantic roles. generalization of communicative actions can also be thought of from the standpoint of matching the verb frames. the communicative links reflect the discourse structure associated with participation (or mentioning) of more than a single agent in the text. the links form a sequence connecting the words for communicative actions (either verbs or multi-words implicitly indicating a communicative intent of a person). for a communicative action, we distinguish an agent, one or more arguments being acted upon, and the phrase describing the features of this action. we define communicative action as a function of the form verb(agent, argument, cause), where verb characterizes some type of interaction between involved agents (e.g., explain, confirm, remind, disagree, deny, etc.); argument refers to the information transmitted or object described, and cause refers to the motivation or explanation for the subject. a scenario (labeled directed graph) is a sub-graph of a parse thicket g=(v, a), where v={action1, action2,…,actionk} is a finite set of vertices corresponding to communicative actions, and a is a finite set of labeled arcs (ordered pairs of vertices), classified as follows: • each arc (actioni; actionj) ∈ asequence corresponds to a temporal precedence of two actions (vi, agi, si, ci) and (vj, agj, sj, cj) referring to the same argument (that is, si= sj) or different argument. each arc (actioni, actionj) ∈ acause corresponds to an logical argumentation attack relationship (gabbay and garcez, 2009) between actioni and actioni indicating that the cause of actioni is in conflict with the subject or cause of actionj. subgraphs of parse thickets which are associated with scenarios of interaction between agents have some distinguishing features (galitsky et al., 2009): 1) all vertices are ordered in time, so that there is one incoming arc and one outgoing arc for all vertices (except the initial and terminal vertices); 2) for asequence arcs, at most one incoming and only one outgoing arc are admissible; 3) for acause arcs, there can be many outgoing arcs from a given vertex, as well as many incoming arcs. the vertices involved may be associated with different agents or with the same agent (i.e., when he contradicts himself). to compute similarities between parse thickets and their communicative action – induced subgraphs – the sub-graphs of the same configuration with similar labels of arcs and strict correspondence of vertices need to be analyzed (galitsky and kuznetsov, 2008). analyzing the communicative actions’ arcs of a parse thicket, one can find implicit similarities between texts. given two texts t1 and t2, we can generalize: 1. one communicative actions with its argument from t1 against another communicative action with its argument from t2 (communicative action arc is not used) ; 2. a pair of communicative actions with their arguments from t1 against another pair of communicative actions from t2 (communicative action arcs are used) . in our example in fig. 3 we have the former case: cheating(husband, wife, another lady) ^ avoid(husband, contact(husband, another lady)) discovering rhetorical agreement 181 this generalization gives us communicative_action(husband, *) which introduces a constraint on a in the form that if a given agent (= husband) is mentioned as an argument of ca in q, he(she) should also be an argument of (possibly, another) ca in a. to handle meaning of words expressing the arguments of cas, we apply compositional semantics (mikolov et al., 2011, mikolov et al., 2015) models. to compute generalization between the arguments of communicative actions, we use the following rule: if argument1=argument2, argument1^argument2 = . here subject remains and score is 1. otherwise, if they have the same part-of-speech (pos), argument1^argument2 = <*, pos(argument1), word2vecdistance(argument1^argument2)>. ‘*’ denotes that lemma is a placeholder, and the score is a word2vec distance between these words. if pos’ are different, the generalization is an empty tuple. it cannot be further generalized. as the reader can observe, generalization results can be further generalized with other arguments of communicative actions and with their generalizations. generalizing two different communicative actions is based on their attributes and is presented elsewhere (galitsky et al., 2013). 3.3 generalization for rst relations only rst arcs of the same type of relation (presentation relations, such as antithesis; subject matter relations, such as condition; and multinuclear relations, such as list) can be generalized. we use n for a nucleus or situation presented by this nucleus, and s for satellite or situation presented by this satellite, w for a writer, and r, for a reader (hearer). situations are propositions, completed actions or actions in progress, or communicative actions, or states (including beliefs, desires, approve, explain, reconcile and others). the generalization of two rst relations with the above parameters is expressed as rst1(n1, s1, w1, r1) ^ rst2 (n2, s2, w2, r2) = (rst1^ rst2 )( n1^n2, s1^s2, w1^w2, r1^ r2). the texts in n1, s1, w1, r1 are subject to generalization as phrases. the rules for rst1^ rst2 are as follows: • if relation_type(rst1 ) ! = relation_type(rst2 ) then generalization is empty. • otherwise, we generalize the signatures of rhetorical relations as sentences (iruskieta et al., 2015): sentence(n1, s1, w1, r1) ^ sentence (n2, s2, w2, r2) . for example, we apply generalization to the definitions of rst relations: rst-background ^ rst-enablement = (s increases the ability of r to comprehend an element in n) ^ (r comprehending s increases the ability of r to perform the action in n) = increase-vb thedt ability-nn of-in r-nn to-in. since the relations rst-background ^ rst-enablement are different, the rst relation part is empty. we then generalize the expressions that are the verbal definitions of the respective rst relations. for each word or a placeholder for a word such as an agent, we retain this word (with its pos) if it is the same in each input phrase or remove it if it is different between these phrases. the resultant expression can be interpreted as a common meaning between the definitions of two different rst relations, obtained formally. the reader is recommended to consult galitsky et al., 2012 for further details of syntactic generalization. discovering rhetorical agreement 182 a mapping between two essential rhetorical relations rst-contrast is shown in the middle of fig. 3. this mapping is important to demonstrate that if contrasting clauses occur in the question, they have to be addressed (most likely, via contrast) in the answer as well. to compute the generalization between the expressions for contrast in q and a, we draw a chart: i just had a baby ↔rst-contrast it does not look like me ^ husband to avoid contact ↔rst-contrast have the basic legal and financial commitments the generalization gives us vp ↔rst-contrast vp which is read as “both q and a have contrast expressed via verb phrases”. we can see that the vp of a does not have to be similar to the vp of q but their rhetorical structure does. not all phrases in a must match phrases in q: those which do not match must be in certain rhetorical relations with those a phrases which are relevant to phrases in q. 3.4 representing a request-response chain so far we considered the stand-alone r-r pairs. in this section we explore how these pairs can form a chain and how to represent its structure. once we have a chain, a rhetorical agreement is expected to hold not only between consecutive members but also triples and four-tuples. for a text expressing a sequence of r-r pairs, we will compare the scenario representation with the discourse tree representation. in the domain of customer complaints, request and response are present in the same text, from the viewpoint of a complainant. the customer complaint text needs to be split into request and response text portions to form the positive and negative dataset of pairs. for the purpose of evaluation, we combine all text for the proponent and all text for the opponent together. the first sentence of each paragraph below will form the request part (which will include three sentences) and the second sentence of each paragraph will form the response part (which will also include three sentences in this example). let us consider a scenario with communicative actions and its representation via dt and scenario graph (fig. 5). i explained that my check bounced (i wrote it after i made a deposit). a customer service representative accepted that it usually takes some time to process the deposit. i reminded that i was unfairly charged an overdraft fee a month ago in a similar situation. they denied that it was unfair because the overdraft fee was disclosed in my account information. i disagreed with their fee and wanted this fee deposited back to my account. they explained that nothing can be done at this point and that i need to look into the account rules closer. a discourse tree and communicative scenario for this conflict dialogue is shown in fig. 4. judging by the dt, it is hard to see if this text is an interaction scenario or just some kind of description. the scenario graph on the bottom is a higher-level representation of discourse. discovering rhetorical agreement 183 fig. 4: discourse tree on the top and scenario graph for communicative actions on the bottom. these two sources of discourse information complement each other. 4 classification settings for request-response pairs in a conventional search approach, as a baseline, rr match is measured in terms of keyword statistics such as tf*idf. to improve search relevance, this score is augmented by item popularity, item location or taxonomy-based score (galitsky 2015). also, search can be formulated as a passage re-ranking problem in a machine learning framework. the feature space includes rr pairs as elements, and a separation hyper-plane splits this feature space into correct and incorrect pairs. hence, a search problem can be formulated in a local way, as similarity between req and resp, and in a global, learning way, via similarity between rr pairs. we take these groups of methods further towards discourse level analysis. to measure the rr match there are the following classes of methods: 1. extract features for req and resp and compare them as a feature count. introduce a scoring function such that a score would indicates a class (low score for incorrect pairs, high score for correct ones); 2. compare representations for req and resp against each other, and assign a score for the comparison result. analogously, the score will indicate a class; 3. build a representation for a pair req and resp, as elements of the training set. then perform learning in the feature space of all such elements . to form a object we combine dt(req) with dt(resp) into a single tree with the root rr (fig. 5). we then classify such objects into correct (with high agreement) and incorrect (with low agreement). discovering rhetorical agreement 184 fig. 5: forming the request-response pair as an element of a training set. 4.1 nearest neighbor graph-based classification to identify an argument in a text, once the cdt is built, one needs to compute its similarity with cdts for the positive class and verify that this similarity is lower than the similarities of this cdt with each element of the negative class. similarity between cdt is defined by means of maximal common sub-cdts. since we describe cdts by means of labeled graphs, first we consider formal definitions of labeled graphs and of the domination relation on them (see, e.g., ganter & kuznetsov 2001). let’s have an ordered set g of cdts (v,e) with vertexand edge-labels from the sets (λς, ) and (λε, ). a labeled cdt γ from g is a pair of pairs of the form ((v,l),(e,b)), where v is a set of vertices, e is a set of edges, l: v → λς is a function assigning labels to vertices, and b: e → λε is a function assigning labels to edges. we do not distinguish isomorphic trees with identical labeling. the order is defined as follows: for two cdts γ1:= ((v1,l1),(e1,b1)) and γ2:= ((v2,l2),(e2,b2)) from g we say that γ1 dominates γ2 or γ2 ≤ γ1 (or γ2 is a sub-cdt of γ1) if there exists a one-toone mapping φ: v2 → v1 such that it • respects edges: (v,w) ∈ e2 (φ(v), φ(w)) ∈ e1, • fits under labels: l2(v) l1(φ(v)), (v,w) ∈ e2 b2(v,w) b1(φ(v), φ(w)). this definition takes into account the calculation of similarity (“weakening”) of labels of matched vertices when passing from the “larger” cdt g1 to “smaller” cdt g2. now, the similarity cdt z of a pair of cdts x and y, denoted by x ^ y = z, is the set of all inclusion-maximal common sub-cdts of x and y, each of them satisfying the following additional conditions: • to be matched, two vertices from cdts x and y must denote the same rst relation; • each common sub-cdt from z contains at least one communicative action with the same verbnet signature as in x and y. this definition is easily extended to finding generalizations of several graphs (e.g., see ganter and kuznetsov 2001; kuznetsov 1999). the subsumption order µ on pairs of graph sets x and y is naturally defined as x µ y := x ∗y = x. ⇒ ⇒ discovering rhetorical agreement 185 an example of maximal common sub-cdt for cdts fig. 2a and 3b is shown in fig. 6. notice that the tree is inverted and the labels of arcs are generalized: communicative action cite() is generalized with communicative action say(). the first (agent) argument of the former ca committee is generalized with the first argument of the latter ca dutch. the same operation is applied to the second arguments for this pair of cas: investigator ^ evidence. notice that because the arguments of rhetorical relations are different in a general case, the leaf nodes on the bottom of the resultant discourse tree are empty. we define the condition such that cdt u belongs to a positive class: 1) u is similar to (has a nonempty common sub-cdt) with a positive example r+. 2) for any negative example r-, if u is similar to r(i.e., u ∗ r-≠∅) then u ∗ rµ u ∗ r+. this condition introduces the measure of similarity and says that to be assigned to a class, the similarity between the unknown cdt u and the closest cdt from the positive class should be higher than the similarity between u and each negative example. condition 2 implies that there is a positive example r+ such that for no rone has u ∗ r+ µ r-, i.e., there is no counterexample to this generalization of positive examples. fig. 6: maximal common sub-cdt for two cdts figs. 2a and 2b 4.2 thicket kernel learning for cdt tree kernel learning for strings, parse trees and parse thickets is a well-established research area these days. the parse tree kernel counts the number of common sub-trees as the discourse similarity measure between two instances. tree kernels have been defined for dt by (joty and moschitti et al., 2014). (wang et al., 2013) used a special form of tree kernels for discourse relation recognition. in this study we define the thicket kernel for cdt, augmenting dt kernels with the information on communicative actions. a cdt can be represented by a vector v of integer counts of each sub-tree type (without taking into account its ancestors): v (𝑇) = (#𝑜𝑓 𝑠𝑢𝑏𝑡𝑟𝑒𝑒𝑠 𝑜𝑓 𝑡𝑦𝑝𝑒 1, … , # 𝑜𝑓 𝑠𝑢𝑏𝑡𝑟𝑒𝑒𝑠 𝑜𝑓𝑡𝑦𝑝𝑒 𝐼, … , # 𝑜𝑓 𝑠𝑢𝑏𝑡𝑟𝑒𝑒𝑠 𝑜𝑓 𝑡𝑦𝑝𝑒 𝑛). this results in a very high dimensionality since the number of different sub-trees is exponential in its size. thus, it is computationally infeasible to directly use the feature vector ∅(𝑇). to solve the computational issue, a tree kernel function is introduced to calculate the dot product between the discovering rhetorical agreement 186 above high dimensional vectors efficiently. given two tree segments cdt1 and cdt2 , the tree kernel function is defined: 𝐾 (cdt1 , cdt2) = < v (cdt1 ), v (cdt2 ) > = σi v (cdt1 )[i], v (cdt2)[i] = σn1σn2 σi ii(n1)* ii(n2) where: 𝑛1∈𝑁1 , n2∈𝑁2 where 𝑁1 and n2 are the sets of all nodes in cdt1 and cdt2 , respectively; 𝐼i (𝑛) is the indicator function. 𝐼i (𝑛) = {1 iff a subtree of type 𝑖 occurs with root at node; 0 otherwise}. 𝐾 (cdt1, cdt2) is an instance of convolution kernels over tree structures (collins and duffy, 2002) and can be computed by recursive definitions: δ (𝑛1, n2) = σi ii(n1)* ii(n2); (1) δ (𝑛1, n2) = 0 if 𝑛1 and n2 are assigned the same pos tag or their children are different subtrees. (2) otherwise, if both 𝑛1 and n2 are pos tags (are pre-terminal nodes) then δ (𝑛1, n2) = 1x𝜆; (3) otherwise, δ (𝑛1, n2) = 𝜆 (1 +  δ   ch n1, j , ch n2, j )!"(!!) !!! (4) where 𝑐h(𝑛,𝑗) is the 𝑗th child of node 𝑛, 𝑛𝑐(𝑛1) is the number of the children of 𝑛1 , and 𝜆 (0 < 𝜆 < 1) is the decay factor in order to make the kernel value less variable with respect to the sub-tree sizes. in addition, the recursive rule (3) holds because given two nodes with the same children, one can construct common sub-trees using these children and common sub-trees of further offsprings. the parse tree kernel counts the number of common sub-trees as the syntactic similarity measure between two instances. fig.7: a tree in the kernel learning format for cdt fig 2a. a cdt representation for kernel learning is shown in fig. 7. the terms for communicative actions as labels are converted into trees which are added to respective nodes for rst relations. discovering rhetorical agreement 187 for texts for edus as labels for terminal nodes only the phrase structure is retained: we label the terminal nodes with the sequence of phrase types instead of parse tree fragments. if there is a rhetorical relation arc from a node x to a terminal edu node y with label a(b, c(d)), then we append the subtree in fig. 7a to x. 4.3 implementation of the rhetorical agreement classifier 1. define positive and negative classes of rr pairs: a) form the positive class from the rhetorically correct rr pairs b) form the negative class from the relevant but rhetorically foreign rr pairs 2. for each rr pair: a) parse each sentence b) obtain verbnet structure for verbs c) obtain coreferences d) obtain entity entity and entity – sub-entity links e) build parse thicket pair for ptrr f) apply discourse parsing to obtain discourse tree pair dtrr for rr pair g) align edus of dtrr with ptrr h) merge aligned edus of dtrr with ptrr i) obtain dtrr with verbnet signatures for cas j) obtain parse thicket with enriched rst relations k) build representation for thicket kernel learning l) build representation for nearest neighbor learning 3. apply thicket kernel learning 4. apply nearest neighbor learning fig. 8a: implementation of the rhetorical agreement classifier. fig. 8a shows the architecture of the rhetorical agreement classifier. the training dataset stores the request-response pairs for positive and negative cases (on the top). each sentence of each request and response is parsed and then combined into a parse thicket. for each word, we obtain pos, entity type, and for verbs we obtain verbnet entries including roles and frames. once a verb is identified, we use jverbnet framework to attach the verb-related information to the node of the parse tree. an algorithm of populating the roles with words from respective role classes is straightforward. logical form expressions which can potentially be attached to the nodes of discourse trees are also obtained from verbnet. discovering rhetorical agreement 188 verbnet information is valuable for both learning setting: for nearest-neighbor, it helps to express similarity in a richer way when matching verbs are different. without verbnet information, the result of generalization is an empty lemma and pos=verb. with verbnet, all common attributes shared by the different verbs being generalized are retained. for thicket kernel learning, verbnet information helps to obtain a richer set of subgraphs with node labels for verbnet attributes in the similar case where matching verbs are different. the reader is recommended to consult (galitsky 2012) for further details of how phrases with verbnet attributes are generalized. to build a parse thicket, we form a graph from all parse tree nodes and add arcs for coreference, rhetorical and entity-entity relations. discourse trees are initially built as a result of rhetorical parsing (surdeanu et al., 2015) and then verbnet labels are obtained as a result of their merge with parse thickets. we did not use our own training set for rhetorical parsing and used the trained parser instead. notice that since parse thicket building is based on stanford nlp and rhetorical parsing is also based on stanford nlp, integration of these systems is not difficult. to perform nearest neighbor learning we represent paragraphs as graphs and compute their maximal common subgraphs in terms of common phrases. to obtain these common phrases we used the generalization operation which is applied at the level of words, phrases, sentences and paragraphs (galitsky et al., 2012). some of the components of the rhetorical agreement classifier are implemented in the current study, some were implemented in our previous studies and a number of open source components developed by others were employed (please see appendix). for the first two cases, we show the java packages implementing functionality of a given component as a part of github project https://github.com/bgalitsky/relevance-based-on-parse-trees. the least reliable component of the rhetorical agreement classifier is the rhetorical parser. although splitting into edus works reasonably well, assignment of rst relation is noisy and in some domains its accuracy can be as low as 50% (personal communications from some users of discourse parsers). however, when the rst relation label is random, it does not significantly drop the performance of our classification system since a random discourse tree will be less similar to elements of positive or negative training set, and most likely will not participate in positive or negative decisions. to overcome the noisy input problem, more extensive training datasets are required so that the number of reliable, plausible discourse tree is high enough to cover cases to be classified. as long as this number is high enough, a contribution of noisy, improperly built discourse trees is low. 5 evaluation 5.1 evaluation domains to evaluate rhetorical agreement in texts of various styles we form rr-pairs from various sources, applying distinct mechanisms to form positive and negative datasets. table 1 shows the sources of evaluation and characterizes them in terms of the volume, average lengths of positive and negative sets in terms of sentences and words, and also the average numbers of rhetorical relations in respective discourse trees. our first domain is yahoo! answer set of question-answer pairs with broad topics (chang et al., 2008). this dataset includes 20 top-level categories of yahoo! answer website. out of the set of 4.4 million user questions we selected 20000, which included more than two sentences. answers for most questions are fairly detailed so no filtering was applied to answers. there are multiple https://github.com/bgalitsky/relevance-based-on-parse-trees discovering rhetorical agreement 189 answers per questions and the best one is marked. we consider the pair question-best answer as an element of the positive training set and question-other-answer as one member of the negative training set. to derive the negative set, we either randomly select an answer to a different but somewhat related question, or formed a query from the question and obtained an answer from web search results. source yahoo! answers conversation on social networks customer complaints interviews by journalists # of data items (total | positive | negative) 40k 20k 20k 2035 1232 803 670 415 255 3740 187 0 1870 average length of data items (in sentences) 2.8 3.1 2.5 3.9 4.0 3.9 5.2 4.7 5.5 2.9 4.1 1.8 average length of data items (in words) 68 72 65 82 85 79 91 84 93 63 71 48 average number of rhetorical relations 10.3 11.1 9.7 12.7 13.4 11.9 13.4 12.0 14.2 9.5 10. 0 8.3 average number of essential rhetorical relations 3.2 3.2 3.1 3.4 3.7 3.3 3.5 3.3 3.6 3.1 3.3 3.0 table 1: data sources our second dataset includes social media. we extracted request-response pairs mainly from postings on facebook. we also used a smaller portion of linkedin.com and english vk.com conversations related to employment. in the social domains the standards of writing are fairly low. the cohesiveness of text is very limited and the logical structure and relevance frequently absent. the authors formed the training sets from their own accounts and also public facebook accounts available via api over a number of years (at the time of writing, facebook api for getting messages is unavailable). in addition, we used 860 email threads from the enron dataset (cohen 2016). also, we collected the data of manual responses to postings of an agent which automatically generates posts on behalf of human users-hosts (galitsky et al., 2014). we formed 2035 rr pairs from the various social network sources where request and response include 3+ grammatically correct sentences. the third domain is customer complaints. in a typical complaint a dissatisfied customer describes his problems with products and service as well as the process for how he attempted to communicate these problems with the company and how they responded. complaints are frequently written in a biased way, exaggerating product faults and presenting the actions of opponents as unfair and inappropriate. at the same time, the complainants try to write complaints in a convincing, coherent and logically consistent way (galitsky et al., 2014, githubdeceptiondataset 2017); therefore complaints serve as a domain with high agreement between requests and response. we split each complaint into two parts: 1) a complainant describing how he communicated an issue; 2) this complainant describing how the company responded. discovering rhetorical agreement 190 for the purpose of assessing agreement between user complaint and company response (according to how this user describes it) we collected 670 complaints from planetfeedback.com over 10 years. a typical complaint is a report of a failure of a product or service, followed by a narrative on the customer's attempts to resolve the issue. these complaints include both a description of the product or service failure and a description of the resulting interaction process (negotiation, conflict, etc.) between the customer and the company representatives. since it is almost impossible to verify the actual occurrence of such failures, company representatives must judge the adequacy of a complaint on the basis of the communicative actions. a complaint narrative usually describes a conflict between an unsatisfied customer and customer support representatives, in which communicated claims need to be rationally justifiable by sound arguments. in contrast with the almost unlimited number of possible details regarding product failures, the emerging argumentative dialogues between customer and company can be subject to a systematic computational study (galitsky et al., 2009). the fourth domain is interviews by journalists. usually, the way interviews are written by professional journalists is such that the match between questions and answers is very high. we collected 1200 contributions of professional and citizen journalists from such sources as allvoices.com, huffingtonpost.com and others. over four years from 2011 to 2013, about 27 000 interviews by citizen journalists were submitted to allvoices.com. these interviews contain an extended question by a journalist, providing some background and own journalist opinion along with the question or a statement. it is followed by an answer of the person being interviewed, written by the journalist. also, some of journalists’ interviews include the comments inserted by other website users, usually not journalists. querying the database of articles and comments, we selected 1870 triples of paragraphs containing: 1) original question and background; 2) answer by the person being interviewed; 3) comment by a user other than journalist. we form the positive dataset from 1 and 2 and negative dataset with 1) and 3). obviously 1) and 3) share the same topic but usually not in a good rhetorical agreement (as, for example, 1)+2) vs. 3) would be). to facilitate data collection, we designed a crawler which searched a specific set of sites, downloaded web pages, extracted candidate text and verified that it adhered to a question-orrequest vs. response format. then the respective pair of text is formed. the search is implemented via microsoft cognitive services search api in the web, news & blogs domains. each data source in table 1 is characterized by the following parameters with respective rows: 1. the number of data items (paragraphs) in the training set. we show the number of total, positive and negative cases; 2. average length of data items (in sentences). we also show the number of total, positive and negative cases; 3. average length in words; 4. average number of rhetorical relations (measured as dt edges); 5. average number of essential rhetorical relations (other than elaboration and joint) 5.2 recognizing valid and invalid r-r pairs in this section, we evaluate the accuracies of rhetorical agreement-based classification. we first outline the baseline methods for rhetorical agreement assessment, present the results for nearest discovering rhetorical agreement 191 neighbor learning based on computing maximal common sub-graphs (table 2a), and then proceed to the evaluation of the statistical, kernel-based methods (table 2b). each row represents a specific method. the first baseline approach we selected in an attempt to detect proper rhetorical agreement was based on counting rhetorical relations. the premise here is that if a request and response have similar rhetorical relations, then agreement should be high. conversely, if the rhetorical relations were different for this request and response, then we would expect the rhetorical agreement to be low. hence, the linear regression-based classifier counts the different rhetorical relations. this premise turned out to provide too coarse of an approximation of the rhetorical agreement phenomenon: it was false in most cases, and the recognition accuracy was close to that of a random classifier. the second baseline approach was based on a naïve hypothesis that common keywords might be correlated with rhetorical agreement. when computing the common keywords, we removed stop words as per the default search application (lucene library). this hypothesis turned out not to be the case: the common keywords shared by req and resp were weakly correlated with rhetorical agreement. table 2a: evaluation results for baseline approaches and maximal common subgraph-based learning. the third baseline approach relied on a named-entity-based alignment of the dts of req and resp. the intuition behind this approach was that if the rr dts are well-aligned in terms of these entities, the rhetorical agreement should be high. an abstract tree alignment problem extensively studied in bioinformatics is np hard even when the set of labels is limited (varon & wheeler 2012). therefore, we reduced it to a sequence alignment problem for fairly short sequences of entities; thus, no optimization was required. although this approach resulted in an improvement over the previous one by a few percentile points, one might observe that matching entities between req and resp are weakly correlated with rhetorical agreement. source / evaluation setting yahoo! answers conversation on social networks customer complaints interviews by journalists p r f1 p r f1 p r f1 p r f1 counts of types of rhetorical relations of req and resp 55.2 52.9 54.03 51.5 52.4 51.95 54.2 53.9 54.05 53.0 55.5 54.23 bag-of-words 54.7 55.2 54.94 50.8 51.9 51.90 55.0 52.7 53.83 51.6 54.3 52.92 entity-based alignment of dts of req and resp 63.1 57.8 60.33 51.6 58.3 54.70 48.6 57.0 52.45 49.2 57.9 53.21 maximal common sub-dt for req and resp 67.3 64.1 65.66 70.2 61.2 65.40 54.6 60.0 57.16 80.2 69.8 74.61 maximal common sub-cdt for req and resp 68.1 67.2 67.65 68.0 63.8 65.83 58.4 62.8 60.48 77.6 67.6 72.26 discovering rhetorical agreement 192 the fourth and fifth approaches were based on computing the maximal common sub-graph as a way to measure the similarity between trees, as presented in section 4.1. for the richer set of labels provided by cdt versus dt, we obtained smaller common sub-trees but with a higher number of labels. one can observe that the rhetorical agreement classifier based on cdt provided a better performance than the one based on dt, exceeding the entity alignment case by a few percentiles. table 2b: evaluation results for thicket learning family of approaches. we present the classification results for the svm tk family of approaches (section 4.2) in table 2b. from top to bottom, we extend the sources of linguistic information employed by svm: 1) parse trees for sentences only; 2) parse trees are connected into parse thickets, so that rhetorical relations, communicative actions, coreference links and entity-entity links are leveraged; 3) regular discourse trees; 4) communicative discourse trees; 5) communicative discourse trees with sentiment and argumentation features added. one can see that the highest accuracy was achieved in the domains of journalism and community answers and the lowest in the domains of customer complaints and social networks. we can conclude that the higher the accuracy achieved by having the method fixed is, the higher the level of agreement between req and resp is. as the responder’s competence becomes higher, the rhetorical agreement increases as well. the best representative of the deterministic family of approaches (the bottom row of table 2a) performed approximately 9% below the best of svm tk (the fourth row of table 2b). this observation indicates that the similarity between req and resp is substantially less important than certain structures of rr pairs indicative of an rr agreement. this result means that the agreement between req and resp cannot be assessed on an individual basis: if we demand that dt(req) be very similar to dt(resp), we will obtain decent precision but extremely low recall. source / evaluation setting yahoo! answers conversation on social networks customer complaints interviews by journalists p r f1 p r f1 p r f1 p r f1 svm tk for parse trees of individual sentences 66.1 63.8 64.93 69.3 64.4 66.80 46.7 61.9 53.27 78.7 66.8 72.24 svm tk for rst and ca (full parse thickets) 75.8 74.2 74.99 72.7 77.7 75.11 63.5 74.9 68.74 75.7 84.5 79.83 svm tk for rr-dt 76.5 77 76.75 74.4 71.8 73.07 64.2 69.4 66.69 82.5 69.4 75.40 svm tk for rr-cdt 80.3 78.3 79.29 78.6 82.1 80.34 59.5 79.9 68.22 82.7 80.9 81.78 svm tk for rr-cdt + sentiment + argumentation features 78.3 76.9 77.59 67.5 69.3 68.38 55.8 65.9 60.44 76.5 74.0 75.21 discovering rhetorical agreement 193 proceeding from dt to cdt for svm tk helps by only 1–2%, because communicative actions play no major role in either composing a request or forming a response. for the statistical family of approaches (table 2b), the richest source of discourse data (svm tk for rr-dt) provided the highest classification accuracy—almost the same as under the rr similarity-based classification. although svm tk for rst and ca (full parse trees) included more linguistic features some part of it (most likely syntactic) was redundant and yielded poorer results for the limited training set. using additional features as input to tk, such as sentiment and argumentation, did not help either. most likely, these features were derived from rr-cdt features and did not contribute to classification accuracy on their own. to draw a comparison with the accuracy of rhetoric parsers, we state that employing the tk family of approaches based on cdt provided accuracy comparable to that achieved by classifying dt as correct or incorrect by the rhetorical parsing tasks on which state-of-the-art systems have competed over the last few years approached an accuracy of over 80%. direct analytical approaches in the deterministic family performed rather weakly, which means that a higher number and a more complicated structure of features is required. simply counting and considering types of rhetorical relations was insufficient to judge how the rr agreed with each other. even when two rr pairs have the same types and counts of rhetorical relations and communicative actions, they could still belong to opposite rr agreement classes in most cases. nearest-pair neighbor learning for cdt achieved lower accuracy than did svm tk for cdt, but the former yielded interesting examples of sub-trees typical of argumentation and ones that are shared among the q/a pairs of the factoid type. the number of the former groups of cdt sub-trees was naturally significantly higher. unfortunately, the svm tk approach did not help in explaining exactly how the rr agreement problem was solved. it only provided final scoring and class labels. it is possible but uncommon to express a logical argument in a response without communicative actions (this observation was backed up by our data). manual analysis of the false negative rr pairs showed that some cases of rhetorical agreement were hard to cover by the mapping of respective cdts. a number of responses that could be viewed as sarcastic did not follow rhetorical structure but were nevertheless considered by the readers as cohesive. many cases of official answers, answers in legalese or other domain-specific professional language were wrongly classified as being in bad agreement. this was not only because of noisier cdt construction but also due to the limitation of the cdt model in general. we believe an additional model for rr pair agreement is required, one that goes beyond cdt mapping and discourse level in general. creating such a model could be a topic of future studies. most false-positive rr pairs were obtained when the cdt of the request was structurally similar to the cdt of the response but when the rhetorical agreement was nevertheless low. these were mostly cases in which the rr pair to be classified was similar to a given rr pair in the positive training dataset but the rhetorical structure was not suitable for a given domain. the next step in improving a rhetorical agreement classifier would be to have a training set specific to a broad domain rather than a domain-independent training set formed for a given text genre, as we did in this study. this possibility could potentially be explored in future research. 5.3 cdt construction task in this section, we evaluate how well cdts were constructed irrespective of how they were used and learned. as mentioned earlier, although splitting text into edus works reasonably well, assignment of the rst relation is fairly noisy. however, when the rst relation label is random, it does not significantly drop the performance of our argumentation detection system. to overcome the noisy input problem, more extensive training datasets are required to make the number of reliable, plausible discourse trees high enough to cover the cases to be classified. discovering rhetorical agreement 194 when this number is sufficiently high, the contribution of noisy, improperly built discourse trees is low. there is a certain systematic deviation from the correct, intuitive discourse trees obtained by discourse parsers. in this section, we evaluate whether there is a correlation between the deviation in cdts and some specificities of our training sets. we consider the possibility that the cdt deviations for q/a pairs with high rhetorical agreement is stronger than the ones with low rhetorical agreement. for each source, we calculate the number of significantly deviated cdts. for this assessment, we consider a cdt as deviated when more than 20% of rhetorical relations were determined improperly. we did not differentiate between the specific rst relations associated with high rhetorical agreement. the cdt distortion evaluation dataset was significantly smaller than the detection dataset, because substantial manual effort was required, and the task could not be submitted to amazon mechanical turk workers. as table 3 shows, there was no obvious correlation between the recognition classes and the rate of cdt distortion (less than 3%). hence, we conclude that the training set of noisy cdts can be adequately evaluated with respect to the detection of high and low rhetorical agreement. table 3: does deviation in cdt construction depend on the domain? we also assessed the agreement between the two rhetorical parsers available from surdeanu et al. (2015) and joty et al. (2013). in approximately 60% of the cases, the resultant dts did not match, but in less than 15% of cases, this deviation affected the rhetorical agreement decision, as a small subset of our evaluation set shows. we did not collect evidence on which parser performed better in our domains; we used only the surdeanu et al.’s (2015) parser in this study due to its ease of integration and its lightweight implementation. 6 related work although discourse analysis has a limited number of applications in question-answering and summarization and generation of text, we have not found applications of automatically constructed discourse trees. we discuss research related to applications of discourse analysis in source positive training set size negative training set size significantly deviating dts for positive training set, % significantly deviating dts for negative training set, % yahoo! answers 50 50 18.6±3.43 21.3±2.34 conversation on social networks 50 50 16.0±4.92 19.8±5.40 customer complaints 40 40 21.3±4.62 18.8±4.36 interviews 40 40 17.6±3.43 18.5±4.72 discovering rhetorical agreement 195 two areas: dialogue management and dialogue games. these areas can potentially be applied to the same problems as those for which the current proposal is intended. research in these areas includes both logic-based approaches as well as analytical and machine learning-based approaches. 6.1 managing dialogues and question answering if a question and answer are logically connected, their rhetorical structure agreement becomes less important. (de boni 2007) proposed a method of determining the appropriateness of an answer to a question through a proof of logical relevance rather than a logical proof of truth. we define logical relevance as the idea that answers should not be considered as absolutely true or false in relation to a question, but should be considered true more flexibly in a sliding scale of aptness. then it becomes possible to reason rigorously about the appropriateness of an answer even in cases where the sources of answers are incomplete or inconsistent or contain errors. the authors show how logical relevance can be implemented through the use of measured simplification, a form of constraint relaxation, in order to seek a logical proof than an answer is in fact an answer to a particular question. our model of cdt attempts to combine general rhetorical and speech act information in a single structure. while speech acts provide a useful characterization of one kind of pragmatic force, more recent work, especially in building dialogue systems, has significantly expanded this core notion, modeling more kinds of conversational functions that an utterance can play. the resulting enriched acts are called dialogue acts (jurafsky and martin, 2000). in their multi-level approach to conversation acts (traum and hinkelman 1992) distinguish four levels of dialogue acts necessary to assure both coherence and content of conversation. the four levels of conversation acts are: turn-taking acts, grounding acts, core speech acts, and argumentation acts. research on the logical and philosophical foundations of q/a has been conducted over a few decades, having focused on limited domains and systems of rather small size and been found to be of limited use in industrial environments. the ideas of logical proof of “being an answer to” developed in linguistics and mathematical logic have been shown to have a limited applicability in actual systems. most current applied research, which aims to produce working general-purpose (“open-domain”) systems, is based on a relatively simple architecture, combining information extraction and retrieval, as was demonstrated by the systems presented at the standard evaluation framework given by the text retrieval conference (trec) q/a track. (sperber and wilson 1986) judged answer relevance depending on the amount of effort needed to “prove” that a particular answer is relevant to a question. this rule can be formulated via rhetorical terms as relevance measure: the less hypothetical rhetorical relations are required to prove an answer matches the question, the more relevant that answer is. the effort required could be measured in terms of amount of prior knowledge needed, inferences from the text or assumptions. in order to provide a more manageable measure we propose to simplify the problem by focusing on ways in which constraints, or rhetorical relations, may be removed from how the question is formulated. in other words, we measure how the question may be simplified in order to prove an answer. the resultant rule is formulated as follows: the relevance of an answer is determined by how many rhetorical constraints must be removed from the question for the answer to be proven; the less rhetorical constraints must be removed, the more relevant the answer is. there is a very limited corpus of research on how discovering rhetorical relations might help in q/a. (santosh and jahfar 2012) discuss the role of discourse structure in dealing with 'why' questions, that helps in identifying the relationship between sentences or paragraphs from a given discovering rhetorical agreement 196 text or document. (kontos et al., 2016) introduced a system which allowed an exploitation of rhetorical relations between a “basic ” text that proposes a model of a biomedical system and parts of the abstracts of papers that present experimental findings supporting this model. adjacency pairs is a popular term for what we call rr-pair in this paper. adjacency pairs are defined as pairs of utterances that are adjacent, produced by different speakers, ordered as first part and second part, and typed—a particular type of first part requires a particular type of second part. some of these constraints could be dropped to cover more cases of dependencies between utterances (popescu-belis 2005). adjacency pairs are relational by nature, but they could be reduced to labels (‘first part’, ‘second part’, ‘none’), possibly augmented with a pointer towards the other member of the pair. frequently encountered observed kinds of adjacency pairs include the following ones: request / offer / invite → accept / refuse; assess → agree / disagree; blame → denial / admission; question → answer; apology → downplay; thank → welcome; greeting → greeting (levinson 2000). rhetorical relations, similarly to adjacency pairs, are a relational concept, concerning relations between utterances, not utterances in isolation. it is however possible, given that an utterance is a satellite with respect to a nucleus in only one relation, to assign to the utterance the label of the relation. this poses strong demand for a deep analysis of dialogue structure. the number of rhetorical relations in rst ranges from the ‘dominates’ and ‘satisfaction-precedes’ classes used by (grosz and sidner 1986) to more than a hundred types. coherence relations are an alternative way to express rhetorical structure in text (scholman et al., 2016). (mitocariu et al., 2013) considers cases when two different tree structures of the same text can express the same discourse interpretation, or something very similar. the authors apply both rst and veins theory (cristea et al., 1998), which uses binary trees augmented with nuclearity notation. in the current paper we attempt to cover these cases by learning, expecting different dts for the same text to be covered by an extended training set. there are many classes of nlp applications that are expected to leverage the informational structure of text. dt can be very useful is text summarization. knowledge of salience of text segments, based on nucleus-satellite relations proposed by (sparck-jones 1995) and the structure of relation between segments should be taken into account to form exact and coherent summaries. one can generate the most informative summary by combining the most important segments of elaboration relations starting at the root node. dts have been used for multi-document summaries (radev 2000). in the natural language generation problem, whose main difficulty is coherence, the informational structure of the text can be relied upon to organize the extracted fragments of text in a coherent way. a way to measure text coherence can be used in automated evaluation of essays. since a dt can capture text coherence, then yielding discourse structures of essays can be used to assess the writing style and quality of essays. (burstein et al., 2002) described a semiautomatic way for essay assessment that evaluated text coherence. the neural network language model proposed in (bengio et al., 2003) uses the concatenation of several preceding word vectors to form the input of a neural network, and tries to predict the next word. the outcome is that after the model is trained, the word vectors are mapped into a vector space such that distributed representations of sentences and documents semantically similar words have similar vector representations. this kind of model can potentially operate on discourse relations, but it is hard to supply as rich linguistic information as we do for tree kernel learning. there is a corpus of research that extends word2vec models to go beyond word level to achieve phrase-level or sentence-level representations (mikolov et al., 2015). for instance, a simple approach is using a weighted average of all the words in the document, (weighted averaging of word vectors), losing the word order similar to bag-of-words approaches. a more discovering rhetorical agreement 197 sophisticated approach is combining the word vectors in the order given by a parse tree of a sentence, using matrix-vector operations (socher et al., 2010). using a parse tree to combine word vectors, has been shown to work only for sentences because it relies on parsing. given a dt for a text as a candidate answer to a compound query, (galitsky et al., 2015) proposed a rule system for valid and invalid occurrence of the query keywords in this dt. to be a valid answer to a query, its keywords need to occur in a chain of elementary discourse units of this answer so that these units are fully ordered and connected by nucleus – satellite relations. an answer might be invalid if the queries’ keywords occur in the answer's satellite discourse units only. classifying user intent in horizontal web searches is a well-known difficult problem. jansen et al. (2007) proposed a classification algorithm to determine user intent underlying web search engine queries that considers three classes: informational, navigational, and transactional. the results showed that more than 80% of web queries are informational in nature, while approximately 10% each are navigational or transactional. kathuria et al. (2010) solved the same problem using a k-­‐‑means clustering approach based on a variety of query traits. the authors showed than 75% of web queries (clustered into eight classifications) are informational, while 12% each are navigational or transactional. their results also show that web queries fall into eight clusters, six are primarily informational and two are primarily transactional or navigational. as for a chatbot, one of its essential capabilities is to discriminate between a request to commit a transaction and a question to obtain some information (galitsky et al., 2017). this capability is expected to be domain-independent: in any domain, a user may either request that the system do something or provide a recommendation. this functionality should also be context-independent: a user may switch from information access to a request to do something and back to information access, although this should be discouraged. even human customer support agents prefer a user to first receive information, then make a decision, and finally, request an action. lewandowski (2016) confirmed that search engine performance on navigational queries is of great importance, because users can clearly identify queries that have returned correct results. as such, performance on navigational query types may contribute to explaining user satisfaction with search engines. the performance of chat bots and search engines strongly depends on their ability to capture user intent. vogel et al. (2005) addressed the issue of mapping a search engine query to certain nodes of a subject taxonomy that expresses a possible query. an architecture of a user intent classification system uses a web directory to determine the query context by the query term frequencies. 6.2 analytical approaches to rr agreement in this paper, we formulated the problem of rr agreement and approached it via machine learning. however, there are several approaches other than learning that tackle the relationships between request and response from different perspectives. not all features of these perspectives are covered by our learning framework, which is designed to automatically extract available features from text. moreover, a statistical or reinforcement learning framework does not reveal which features exactly are leveraged (rieser and lemon 2011). therefore, it is worth mentioning explicit models of relationships between requests and responses. in an arbitrary conversation, a question is typically followed by an answer, or some explicit statement of an inability or refusal to answer. the following model explains the intentional thread of a conversation as: from the yielding of a question by agent b, agent a recognizes agent b’s goal to find out the answer, and it adopts a goal to tell b the answer in order to be cooperative. a then plans to achieve the goal, thereby generating the answer. this provides an elegant account in the simple case, but requires a strong assumption of co-cooperativeness. agent discovering rhetorical agreement 198 a must adopt agent b’s goals as her own. as a result, it does not explain why a says anything when she does not know the answer or when she is not ready to accept b’s goals. an intentional analysis at the level of textual discourse was introduced by (litman and allen 1987). the authors assumed a set of typical multi-agent actions. other authors attempted to simulate these forms of multiagent behavior via social intentional models such as joint intentions (cohen and levesque 1991) or shared plans (grosz and sidner 1990). although these approaches help explain certain discourse features of multiagent interaction, they still need to shed a light on how dialogue coherence is achieved. let us imagine a stranger approaching a person and asking, “do you have spare coins?” it is unlikely that there is a joint intention or shared plan, as they have never met before. from a purely strategic point of view, the agent may have no interest in whether the stranger’s goals are met. yet, typically agents will still respond in such situations. hence an account of q/a must go beyond recognition of speaker intentions. questions do more than just provide evidence of a speaker’s goals, and something more than adoption of the goals of an interlocutor is involved in formulating a response to a question. an interesting model is described by (airenti et al., 1993), which separates out the conversational games from the task-related games in a way similar way to (litman and allen 1987). because of this separation, they do not have to assume co-operation on the tasks each agent is performing, but still require recognition of intention and co-operation at the conversational level. it is left unexplained what goals motivate conversational co-operation. 6.3 rhetorical relations and argumentation frequently, the main means of linking questions and answers is logical argumentation. there is an obvious connection between rst and argumentation relations which we tried to learn in this study. there are four types of rhetorical relations correlated with logical argumentation: the directed relations support, attack, detail, and the undirected sequence relation (lippi and torroni 2016). the support and attack relations are argumentative relations, which are known from related work (peldszus and stede, 2013), whereas the latter two correspond to discourse relations used in rst. the argumentation sequence relation corresponds to “sequence” in rst, the argumentation detail relation roughly corresponds to “background” and “elaboration”. the argumentation detail relation is important because there are many cases in scientific publications, where some background information (for example the definition of a term) is important for understanding the overall argumentation. a support relation between an argument component resp and another argument component req indicates that resp supports (reasons, proves) req. similarly, an attack relation between resp and req is annotated if resp attacks (restricts, contradicts) req. the detail relation is used, if resp is a detail of req and gives more information or defines something stated in req without argumentative reasoning. finally, we link two argument components (within req or resp) with the sequence relation, if they belong together and only make sense in combination, i.e., they form a multi-sentence argument component. in our previous papers we observed that using svm tk, one can differentiate between a broad range of text styles (galitsky et al., 2015). in particular, based on rhetorical structure, it is possible to differentiate between documents without argumentation and ones with various forms of argumentation. this is also applicable to text with and without string sentiment, writing with strong legal focus, engineering / design document focus and a focus in finance. each text style and genre has its inherent rhetorical structure that is leveraged and automatically learned. since the correlation between text style and text vocabulary is rather low, traditional classification approaches which only take into account keyword statistics information could lack accuracy in discovering rhetorical agreement 199 complex cases. we also performed text classification into rather abstract classes such as the belonging to language-object and metalanguage in literature domain and style-based document classification into proprietary design documents (galitsky 2016). evaluation of text integrity in the domain of valid versus invalid customer complains (those with argumentation flow, noncohesive, indicating a bad mood of a complainant) shows the stronger contribution of rhetorical structure information in comparison with the sentiment profile information. discourse structures obtained by the rst parser are sufficient to conduct the text integrity assessment, whereas sentiment profile-based approach shows much weaker results and also does not complement strongly the rhetorical structure ones. handling dialogues via machine learning of communicative discourse trees allowed us to model a wide array of dialogue types of collaboration modes (blaylock et al., 2003) and interaction types (planning, execution, and interleaved planning and execution). 7 conclusion an extensive corpus of studies has been devoted to rst parsers, but the research on how to leverage rst parsing results for practical nlp problems is limited to content generation, summarization and search (jansen et al., 2014). dts obtained by these parsers cannot be used directly in a rule-based manner to filter or construct texts. therefore, learning is required to leverage the implicit properties of dts. to our knowledge, this study is the only one that employs discourse trees and their extensions for general and open-domain question-answering. search engines and recommendation systems need to be capable of understanding and matching users’ communicative intentions, reasoning with these intentions, building their own respective communication intentions and populating these intentions with actual language to be communicated to the user. discourse trees on their own do not provide representations for these communicative intents. in this study, we introduced communicative discourse trees, built upon traditional discourse trees, which, on one hand, can currently be computed efficiently and, on the other hand, constitute a descriptive utterance-level model of a request-response pair. statistical computational learning approaches offer several key potential advantages over the manual rule-based hand-coding approach when developing search and recommendation systems: • data-driven development cycle; • provably optimal action policies; • a more accurate model for response selection; • possibilities for generalization to unseen states; • reduced development and deployment costs for industry. comparing inductive learning results with kernel-based statistical learning while relying on the same information allowed us to perform more concise feature engineering than either approach would on their own. the task of comparing tree structures such as parse trees, discourse trees and parse thickets with respect to similarity is fairly important. dot products of the vectors of features of the trees are adequate ways to implement similarity comparisons, but these vectors are multidimensional. instead of representing complex structures such as parse trees and parse thickets with feature vectors, tree kernels are used, which allow for the computation of similarity over trees without explicitly computing the feature vectors of these tree structures. kernel methods and svm in particular have been widely used in machine learning tasks and, therefore, are assumed to be the best approach to handle parse thickets and extended discourse trees. structural features can potentially be expressed in a deep learning framework; however, industrial deployment of such features in a major cloud infrastructure such as oracle’s would be difficult. the deep learning class of algorithm lacks feature explainability and feature engineering, which are the strong points of tree kernel learning, especially the nearest neighbor discovering rhetorical agreement 200 framework. real-time performance and a lack of large training datasets of discourse structures are additional factors favoring svm and nearest-neighbor classes of learning algorithms. rst parsers are mostly evaluated with respect to agreement with test sets annotated by humans rather than their expressiveness of the features of interest. in this work, we focused on interpretation of dts and explored ways to represent them in a form indicative of an agreement or disagreement rather than the neutral enumeration of facts. to provide a measure of agreement for how a given message in a dialogue is followed by a subsequent message, we used cdts that include labels for communicative actions in the form of substituted verbnet frames. we investigated the discourse features that were indicative of correct versus incorrect request-response and question-answer pairs. we used two learning frameworks to recognize correct pairs: deterministic, nearest-neighbor learning of cdts as graphs and a tree kernel learning of cdts, in which a feature space of all the cdt sub-trees was subject to svm learning. the positive training set was constructed from the correct pairs obtained from yahoo answers, social networks, corporate conversations (including enron emails), customer complaints and interviews by journalists. the corresponding negative training set was created by attaching responses for different random requests and questions that included relevant keywords so that the relevance similarity between requests and responses was high. the evaluation showed that it is possible to recognize valid pairs in 68–79% of cases in the domains of weak request-response agreement and 80–82% of cases in the domains of strong agreement. these accuracy rates are essential to support automated conversations and are comparable to the benchmark task of classifying discourse trees as either valid or invalid as well as with factoid question-answering systems. we believe that this study is the first one to leverage automatically built discourse trees for question-answering support. previous studies have used specific customer discourse models and features that are hard to systematically collect, learn with explainability, reverse engineer and compare with each other. we conclude that learning rhetorical structures in the form of cdts is a key source of data to support answering complex questions, chatbots and dialogue management. the code used in this study is open source and available at: https://github.com/bgalitsky/relevance-based-on-parse-trees. acknowledgements the author is grateful to his colleagues dmitri ilvovsky, sergey kuznetsov, dina pisarevskaya, vishal vishnoi, stephen mcritchie, gautam singaraju, kim kanzaki, mark mathison for fruitful discussion, and to barbara di eugenio, action editor, and anonymous reviewers for significant help in improving the manuscript. references laura aina, natalia philippova, valentin vogelmann and raquel fernández (2017). referring expressions and communicative success in task-oriented dialogues. semdial 2017 saarbruken germany. gabriella airenti, bara, bruno g. and colombetti, marco (1993). conversation and behavior games in the pragmatics of dialogue. cognitive science 17:197–256. james allen, and c. perrault (1980). analyzing intention in utterances, artificial intelligence, 15(3):143– 178,. roy f. baumeister and bushman, b. j. (2010). social psychology and human nature: international edition. belmont, usa: wadsworth. https://github.com/bgalitsky/relevance-based-on-parse-trees discovering rhetorical agreement 201 yoshua bengio, réjean ducharme, pascal vincent, and christian janvin. (2003). a neural probabilistic language model. j. mach. learn. res. 3 (march 2003), 1137-1155. nate blaylock, james allen and ferguson, g. (2003). managing communicative intentions with collaborative problem solving. in current and new directions in discourse and dialogue, springer netherlands, dordrecht, 63-84. david m. blei, ng, andrew y., and jordan, michael (2003). latent dirichlet allocation. in lafferty, john, ed. journal of machine learning research. 3 (4–5): pp. 993–1022. doi:10.1162/jmlr.2003.3.4-5.993. jill c. burstein, lisa braden-harder, martin s. chodorow, bruce a. kaplan, karen kukich, chi lu, donald a. rock and susanne wolff (2002). system and method for computer-based automatic essay scoring. united states patent 6,366,759: educational testing service. ming-wei chang, l. ratinov, d. roth and v. srikumar (2008). importance of semantic representation: dataless classification aaai – 2008. phipip r. cohen & levesque, h. j. (1990). intention is choice with commitment, artificial intelligence, 42: 213-261. william cohen (2016). enron email dataset . https://www.cs.cmu.edu/~./enron/ last downloaded july 10, 2016. crimerussia (2016). http://en.crimerussia.ru/corruption/shadow-chairman-of-the-investigative-committee. dan cristea, ide, n., & romary, l. (1998). veins theory: a model of global discourse cohesion and coherence. in c. boitet & p. whitelock (eds.), 17th international conference on computational linguistics (vol. 1 pp. 281–285). montreal, canada: association for computational linguistics. marco de boni (2007). using logical relevance for question answering, journal of applied logic, volume 5, issue 1, march 2007, pages 92-103. vanessa wei feng and hirst, g. (2011). classifying arguments by scheme. in proceedings of the 49th annual meeting of the association for computational lin guistics, portland, or, 987-996. vanessa wei feng and graeme hirst (2014). a linear-time bottom-up discourse parser with constraints and post-editing. acl 2014), baltimore, usa, june. gov gabbay and garcez, a.s. logical modes of attack in argumentation networks. stud logica (2009) 93: 199. boris galitsky, ilvovsky, d. and kuznetsov so (2015). rhetoric map of an answer to compound queries knowledge trail inc. acl 2015, 681–686. boris galitsky, dmitri ilvovsky, nina lebedeva and daniel usikov (2014) improving trust in automation of social promotion. aaai spring symposium on the intersection of robust intelligence and trust in autonomous systems stanford ca. boris galitsky (2013). content inversion for user searches and product recommendations systems and methods. us patent 9336297. boris galitsky (2012). machine learning of syntactic parse trees for search and classification of text. engineering application of ai . volume 26, issue 3, pages 1072-1091 boris galitsky (2016) using extended tree kernels to recognize metalanguage in text. uncertainty modeling, in kreinovich v., editor. springer. boris, galitsky and josep lluis de la rosa. (2011). concept-based learning of human behavior for customer relationship management. special issue on information engineering applications based on lattices. information sciences. volume 181, issue 10, 15 may 2011, pp 2016-2035. boris galitsky, gabor dobrocsi, josep lluis de la rosa (2012). inferring the semantic properties of sentences by mining syntactic parse trees. data & knowledge engineering. volume 81-82, november, 2012. pages 21-45. boris galitsky, mp gonzález, ci chesñevar (2009). a novel approach for classifying customer complaints through graphs similarities in argumentative dialogue. decision support systems, 46-3, 717-729. boris galitsky, vishnoi, vishal and xu, anfernee. (2017). transaction bot discrimination between user’s question or request. oracle provisional patent application 62/564,868. discovering rhetorical agreement 202 github-deceptiondataset (2017) https://github.com/bgalitsky/relevance-based-on-parsetrees/blob/master/examples/ultimatedeception.xls); barbara j. grosz and candace sidner (1986). attention, intention, and the structure of discourse, computational linguistics, 12(3):175–204. barbara j. grosz & sidner, candace l. (1986). attentions, intentions and the structure of discourse. computational linguistics, 12(3), 175– 204. susan haller, susan mcroy, alfred kobsa (2013). computational models of mixed-initiative interaction. springer science & business media, nov 11, computers 398 pages. lee heeyoung, angel chang, yves peirsman, nathanael chambers, mihai surdeanu and dan jurafsky. (2013). deterministic coreference resolution based on entity-centric, precision-ranked rules. computational linguistics 39(4). takuya y. hiraoka, yamauchi, g. neubig, s. sakti, t. toda and s. nakamura (2013). dialogue management for leading the conversation in persuasive dialogue systems, 2013 ieee workshop on automatic speech recognition and understanding, olomouc, pp. 114-119. hospice houngbo and robert mercer (2014). an automated method to build a corpus of rhetoricallyclassified sentences in biomedical texts. proceedings of the first workshop on argumentation mining, pages 19–23, baltimore, maryland usa, june 26, acl. mikel iruskieta, iria da cunha and maite taboada (2015). a qualitative comparison method for rhetorical structures: identifying different discourse structures in multilingual corpora. lang resources & evaluation, volume 49, issue 2, pp 263–309. bernard j. jansen, danielle l. booth, and amanda spink (2007). determining the user intent of web search engine queries. in proceedings of the 16th international conference on world wide web (www '07). acm, new york, ny, usa, 1149-1150. peter jansen, m. surdeanu, and p. clark (2014.) discourse complements lexical semantics for nonfactoid answer reranking. in proceedings of the 52nd acl. shafiq r. joty, and a. moschitti (2014). discriminative reranking of discourse parses using tree kernels. proceedings of emnlp, 2049-2060. shafiq r. joty, giuseppe carenini, raymond t. ng (2016). codra: a novel discriminative framework for rhetorical analysis. computational linguistics volume 41, number 3. shafiq r. joty, giuseppe carenini, raymond t. ng, and yashar mehdad. (2013). combining intra-and multisentential rhetorical parsing for document-level discourse analysis. in acl (1), pages 486–496. daniel jurafsky, james h. martin (2000). speech and language processing: an introduction to natural language processing, computational linguistics, and speech recognition. upper saddle river, nj: prentice hall. ashish kathuria, bernard j. jansen, carolyn hafernik, amanda spink (2010). classifying the user intent of web queries using k-means clustering, internet research, vol. 20 issue: 5, pp.563-581, karin kipper, anna korhonen, neville ryant and martha palmer. (2008). a large-scale classification of english verbs. language resources and evaluation journal, 42:21–40. john kontos, ioanna malagardi, john peros (2016). question answering and rhetoric analysis of biomedical texts in the aroma system. unpublished manuscript. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.379.5382 (last downloaded september 12, 2016). stephen c. levinson (2000). presumptive meanings: the theory of generalized conversational implicature. cambridge, ma: the mit press. d. lewandowski (2015). evaluating the retrieval effectiveness of web search engines using a representative query sample. j assn inf sci tec, 66: 1763–1775. z. lin, kan m. and ng h. (2009). recognizing implicit discourse relations in the penn discourse tree bank. in proceedings of the 2009 conference on empirical methods in natural language processing (emnlp 2009), singapore, august. discovering rhetorical agreement 203 marco lippi, paolo torroni. (2016). argumentation mining: state of the art and emerging trends. acm trans. internet technol. 16, 2, article 10 (march 2016), 25 pages. diance litman, james allen (1987). a plan recognition model for subdialogues in conversation, cognitive science, 11: 163-200. william mann, matthiessen, c., thompson, s (1992). rhetorical structure theory and text analysis. discourse description: diverse linguistic analyses of a fund-raising text / ed. by w. c. mann and s. a. thompson. – amsterdam. –p. 39–78. william mann and sandra thompson. (1988). rhetorical structure theory: towards a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8(3):243–281. christopher d. manning, mihai surdeanu, john bauer, jenny finkel, steven j. bethard, and david mcclosky. (2014). the stanford corenlp natural language processing toolkit. in proceedings of the 52nd annual meeting of the association for computational linguistics: system demonstrations, pp. 55-60. tomas mikolov, i. sutskever, k. chen (2011). distributed representations of words and phrases and their compositionality. in gs corrado, j dean. advances in neural information processing systems, 31113119. tomas mikolov, k. chen, g.s. corrado; j. dean (2015). computing numeric representations of words in a high-dimensional space. us patent 9,037,464, google, inc. mitocariu, elena, daniel-alexandru anechitei, dan cristea, comparing discourse tree structures (2016) https://www.researchgate.net/publication/262331642_comparing_discourse_tree_structures [accessed may 15, 2016]. alessandro moschitti, (2006). efficient convolution kernels for dependency and constituent syntactic trees. in: proceedings of the 17th european conference on machine learning, berlin, germany. alessandro moschitti, s. quarteroni (2011). linguistic kernels for answer re-ranking in question answering systems, inf. process. manage. 47 (6) 825–842. myle ott, y. choi, c. cardie, and j.t. hancock (2011). finding deceptive opinion spam by any stretch of the imagination. in proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies. myle ott, c. cardie, and j.t. hancock (2013). negative deceptive opinion spam. in proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: human language technologies. andreas peldszus, and stede, m. (2013). from argument diagrams to argumentation mining in texts: a survey. int. j of cognitive informatics and natural intelligence 7(1), 1-31. andrei popescu-belis (2005). dialogue acts: one or more dimensions? tech report issco working paper n. 62. vladimir popescu, jean caelen, corneliu burileanu. logic-based rhetorical structuring for natural language generation in human-computer dialogue. lecture notes in computer science volume 4629, pp 309-317, 2007. dragomir r. radev, hongyan jing, and malgorzata budzikowska. (2000). centroid-based summarization of multiple documents: sentence extraction, utility-based evaluation, and user studies. in proceedings of the 2000 naacl-anlpworkshop on automatic summarization volume 4. marta recasens, marie-catherine de marneffe, and christopher potts (2013). the life and death of discourse entities: identifying singleton mentions. in proceedings of naacl. verena rieser, oliver lemon (2011). reinforcement learning for adaptive dialogue systems: a datadriven methodology for dialogue management and natural language generation. springer science & business media, nov 23, 2011 computers 256 pages. kate rohit, y. w. wong, and r. mooney (2005). learning to transform natural to formal languages. in aaai, 2005. discovering rhetorical agreement 204 soumya santhosh, jahfar ali (2012). discourse based advancement on question answering system. journal on soft computing. merel scholman, jacqueline evers-vermeul, ted sanders (2016). categories of coherence relations in discourse annotation. dialogue & discourse, vol 7, no 2. richard socher, c. d. manning, and a. y. ng (2010). learning continuous phrase representations and syntactic parsing with recursive neural networks. in proceedings of the nips-2010 deep learning and unsupervised feature learning workshop. sparck jones, k. (1995). summarising: analytic framework, key component, experimental method', in summarising text for intelligent communication, (ed. b. endres-niggemeyer, j. hobbs and k. sparck jones), dagstuhl seminar report 79, 13.12-17.12.93 (9350). deirdre wilson and dan sperber (2004). relevance: communication and cognition. blackwell, oxford and harvard university press, cambridge, ma,. rajen subba, and barbara di eugenio (2009). an effective discourse parser that uses rich linguistic informationproceedings of human language technologies: the 2009 annual conference of the north american chapter of the association for computational linguistics mihai surdeanu, thomas hicks, and marco a. valenzuela-escarcega. (1986). two practical rhetorical structure theory parsers. proceedings of the conference of the north american chapter of the association for computational linguistics human language technologies: software demonstrations (naacl hlt), 2015. zhao tiancheng, allen lu, kyusong lee and maxine eskenazi (2017) generative encoder-decoder models for task-oriented spoken dialog systems with chatting capability. semdial 2017 saarbruken germany. david r. traum, and james f. allen. (1994). discourse obligations in dialogue processing. in proceedings of the 32nd annual meeting on association for computational linguistics (acl '94). association for computational linguistics, stroudsburg, pa, usa, 1-8. david r. traum, hinkelman, elizabeth a. (1992). conversation acts in task-oriented spoken dialogue. computational intelligence, 8(3), 575–599. thomas visser, david traum, david devault, rieks op den akker (2014). a model for incremental grounding in spoken dialogue systems to appear in journal of multimodal user interfaces,. david vogel, steffen bickel, peter haider, rolf schimpfky, peter siemen, steve bridges, and tobias scheffer (2005). classifying search engine queries using the web as background knowledge. sigkdd explor. newsl. 7, 2 (december 2005), 117-122. w. wang, su, j., tan, c.l. (2010). kernel based discourse relation recognition with temporal ordering information. acl. wikipedia (2016). malaysia_airlines_flight_17. https://en.wikipedia.org/wiki/malaysia_airlines_flight_17. y. a. wilks (ed.) (1999). machine conversations. kluwer. 8 appendix in appendix we present a detailed chart for the rhetorical agreement algorithm with the references to the integrated components. 1. define positive and negative classes of rr pairs: a) form the positive class from the rhetorically correct rr pairs b) form the negative class from the relevant but rhetorically foreign rr pairs 2. for each rr pair: a) parse each sentence https://en.wikipedia.org/wiki/malaysia_airlines_flight_17 discovering rhetorical agreement 205 stanford nlp parser, ner, sentiment module of (manning et al., 2014, recasens et al., 2013, lee at al 2013) b) obtain verbnet structure for verbs verbnet, jverbnet (kipper et al., 2008, http://projects.csail.mit.edu/jverbnet/). c) obtain coreferences stanford nlp parser – coreference d) obtain entity entity and entity – sub-entity links opennlp.similarity.parse_thicket e) build parse thicket pair for ptrr f) apply discourse parsing to obtain discourse tree pair dtrr for rr pair g) align edus of dtrr with ptrr h) merge aligned edus of dtrr with ptrr opennlp.similarity.parse_thicket i) obtain dtrr with verbnet signatures for cas j) obtain parse thicket with enriched rst relations opennlp.similarity.parse_thicket.rhetoric_structure k) build representation for thicket kernel learning l) build representation for nearest neighbor learning opennlp.similarity.parse_thicket m) improve text similarity assessment by word2vec model mikolov et al., 2011, https://deeplearning4j.org/ 3. apply thicket kernel learning opennlp.similarity.parse_thicket.kernel_interface moschitti 2006, http://disi.unitn.it/moschitti/tree-kernel.htm 4. apply nearest neighbor learning opennlp.similarity.jsmlearning opennlp.similarity.parse_thicket.matching fig. 10: sources of components for the rhetorical agreement classifier. references are shown in italics. dialogue & discourse 12(2) (2021) 145–173 doi: 10.5210/dad.2021.205 lexical alignment to non-native speakers iva ivanova imivanova@utep.edu department of psychology university of texas at el paso 500 w. university ave., el paso, tx 79902, usa holly p. branigan holly.branigan@ed.ac.uk department of psychology university of edinburgh, edinburgh, uk janet mclean j.mclean@abertay.ac.uk division of psychology abertay university, dundee, uk albert costa secretaria.dtic@upf.edu departament de tecnologies de la informació i les comunicacions universitat pompeu fabra, barcelona, spain institució catalana de recerca i estudis avançats (icrea), barcelona, spain martin j. pickering martin.pickering@ed.ac.uk department of psychology university of edinburgh, edinburgh, uk editor: patrick g. t. healey submitted 02/2019; accepted 04/2021; published online 10/2021 abstract two picture-matching-game experiments investigated if lexical-referential alignment to non-native speakers is enhanced by a desire to aid communicative success (by saying something the conversation partner can certainly understand), a form of audience design. in experiment 1, a group of native speakers of british english that was not given evidence of their conversation partners’ picture-matching performance showed more alignment to non-native than to native speakers, while another group that was given such evidence aligned equivalently to the two types of speaker. experiment 2, conducted with speakers of castilian spanish, replicated the greater alignment to nonnative than native speakers without feedback. however, experiment 2 also showed that production of grammatical errors by the confederate produced no additional increase of alignment even though making errors suggests lower communicative competence. we suggest that this pattern is consistent with another collaborative strategy, the desire to model correct usage. together, these results support a role for audience design in alignment to non-native speakers in structured task-based dialogue, but one that is strategically deployed only when deemed necessary. keywords: audience design, communicative success, lexical choice, picture-matching game 1. introduction how we talk may feel unique but is not always original. we sometimes mimic aspects of the utterances of our conversation partners, a phenomenon known as alignment (pickering and garrod, c©2021 iva ivanova, holly p. branigan, janet mclean, albert costa, and martin j. pickering this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). ivanova, branigan, mclean, costa and pickering 2004) . alignment is beneficial for conversations in that it facilitates mutual understanding (ferreira et al., 2012). speakers in dialogue can align to each other’s words (lexical alignment; brennan and clark, 1996), phrasal or sentence structure (structural alignment; branigan et al., 2000), as well as phonetic or prosodic aspects of utterances (giles and powesland, 1975; pardo, 2006). however, alignment and its underlying cognitive mechanisms may vary with features of the conversational situation and characteristics of the conversation partner. one such feature is interacting with a nonnative speaker. here, we study whether lexical alignment with non-native speakers is influenced by a desire to aid a conversation partner’s comprehension, a form of audience design. we address this question in two experiments involving a constrained exchange of referential expressions between participants and confederates, to be able to implement specific manipulations that test our hypotheses. at the end, we discuss differences with, and implications for, less constrained natural dialogue. a classical experimental demonstration of lexical alignment (and ensuing alignment of situation models) was provided by garrod and anderson (1987). in a cooperative maze game, one participant had to indicate her or his position in a maze to another participant viewing the same maze from another room. participants tended to converge in their descriptions of the maze structure and their positions (for example, participant a: “right two along from the bottom one up”; participant b: “two along from the bottom, which side?”). further demonstrations have involved tasks requiring one participant to relate the order of pictures of objects that can be named in more than one way to another participant who had to arrange her or his figures or images in the same order brennan and clark (1996). in such tasks, participants in a dyad typically adopt the same way of referring to the figures or images (for example, if one participant refers to a picture of a shoe as “pennyloafer”, the other participant is more likely to use the same term to refer to it). lexical alignment may be important for communicative success. first, it may aid mutual understanding, by helping conversation partners reach a similar mental representation of a situation. ferreira et al. (2012) showed that participants were faster to choose pictures on a display when their own names for those pictures were repeated back to them versus when other words were used. second, imitating the behavior of one’s conversation partner (including reusing the words they produce) might increase rapport, empathy and prosociality between the conversation partners (for a review, see beňuš, 2014; chartrand and van baaren, 2009; garrod et al., 2018). in addition, alignment may underlie long-term learning: reusing a language representation (including a word representation) may strengthen its connections to other elements in the language system and thus facilitate its reuse on future occasions (see oppenheim et al., 2010; chang et al., 2006, for structural alignment). of note, one study found that lexical alignment conferred communicative benefits only when it was limited to task-relevant vocabulary, suggesting that its effects are well-targeted; overall alignment was in fact negatively correlated with task performance (fusaroli et al., 2012). the potential of lexical alignment to aid communication might make it an especially important communicative mechanism in conversations with non-native speakers because these speakers face various challenges in conversing in a second language. speaking a second language is cognitively demanding and associated with reduced fluency and clarity (presumably stemming from word-finding difficulties; pivneva et al., 2012) and delayed lexical access (ivanova and costa, 2008). understanding a second language is also cognitively effortful and deteriorates in the presence of noise (for the latter, see weiss and dempsey, 2008). still, an increasing number of people converse in a second language with native speakers of that language, for purposes of work, education or 146 lexical alignment to non-native speakers travel but little is known about the functioning of alignment in such conversations (see costa et al., 2008). to begin understanding lexical alignment with non-native speakers, it is necessary to consider the cognitive mechanisms that drive it. one such mechanism is audience design: speakers’ tendency to construct their utterances considering the knowledge, needs and intentions of their conversation partners (brennan and clark, 1996; clark and schaefer, 1987). audience design is goal-driven and involves belief-based judgments about what would be more versus less effective for conversation partners’ comprehension (clark, 1996), or what would be maximally informative (grice, 1975). lexical-referential alignment is among the strategies that ensure successful information transmission: the use of a particular referential expression by the conversation partner is clear evidence that they will understand it if repeated back to them. indeed, comprehension is slowed down when listeners hear different terms than they themselves (ferreira et al., 2012) and also their conversation partners (metzing and brennan, 2003) have previously produced in the course of their interaction. several studies support that audience design can drive lexical alignment. these studies have shown that speakers can override their normal production preferences and adopt the lexical choice of their conversation partners based on their judgment of the partners’ likelihood of experiencing comprehension difficulties. in a director-matcher task in which participants could interact freely, isaacs and clark (1987) showed that new-yorkers (experts) described postcards of buildings in new york city to non-new-yorkers (novices) using fewer building names (which the novices would not know) than to other new-yorkers (and decreased their use of proper names with non-newyorkers by approximately 20% across trials). this result supports audience design by showing that experts adapt their lexical choices to novices’ productions, and presumably do so to ensure that their linguistic contributions will be understood. most relevant to the present study, branigan et al. (2011) showed that speakers aligned their lexical choice to computers more than to humans, and to computers perceived as older more than to computers perceived as newer. in this study, participants alternated between describing pictures either orally or in writing to what they were told was a human or computer conversation partner, and matching pictures to the “partner’s” picture descriptions (in reality, all such descriptions were scripted). on experimental trials, the “partner” produced either preferred (for example, bus for a picture of a bus) or dispreferred but acceptable responses (for example, coach for a picture of a bus). there was greater lexical alignment with conversation partners presumably perceived as less capable of producing accurate descriptions. the authors concluded that this tendency was driven by assumptions about what the conversation partner could versus could not understand (that is the words they used versus words they did not use) instead of a desire for a social bond. this is because people would arguably not want to create a social bond with a computer to a greater extent than with another person. both of these studies suggest that a belief that a conversation partner is less communicatively competent (a “novice” or a computer) can increase lexical alignment, supporting a role of audience design for alignment. in the context of the current study, audience design might also act to increase lexical alignment to non-native relative to native speakers, to the extent that non-native speakers can be perceived as less communicatively competent conversation partners because of their incomplete knowledge and slower processing of the target language. further, audience-design-driven lexical alignment to non-native speakers may be inversely proportional to how communicatively competent they seem, and hence increase when beliefs or perceptions point to lower competence (consistent with the larger alignment with “older” computers in branigan et al., 2011). 147 ivanova, branigan, mclean, costa and pickering strong evidence about what a conversation partner does or does not understand is feedback that provides grounding of a speaker’s utterances – that is, contributes to the mutual belief of sufficient understanding of the entities under discussion to move forward (clark and brennan, 1991; clark and schaefer, 1987; clark and wilkes-gibbs, 1986). in natural conversation, grounding is accomplished in one of three main ways: by providing acknowledgements through backchannel responses (for example, uh-huh or mhm) or assessments (for example, really?? or good god!), by directly initiating the relevant next turn (for example, answering a question in a relevant way), and by continued attention on the speaker as shown through eye-gaze (although the latter can have other functions; clark and brennan, 1991). grounding of referring expressions can also be indicated by alternative descriptions (for example, a: it is the second one on the left, b: you mean the red one?) or indicative gestures (for example, pointing) to verify understanding. however, grounding that is optimal for specific conversational needs is not always possible. for example, grounding is affected by the conversational medium. for example, participants in video conferencing often turn off their microphones or cameras, and by doing so they also eliminate the possibility of providing backchannel responses. but even face-to-face conversations do not always provide opportunities for listeners to assert that they understand every detail of what is being said, such as when the speaker is telling a story, or a professor is lecturing to a class. this is relevant in the current context for the following reason. if lexical-referential alignment is used as a tool to ensure successful comprehension, of which feedback of listener understanding, when it occurs, provides direct evidence, alignment may be used as an audience design strategy to a greater extent (or only) when the possibility for feedback is reduced or absent. consistent with this possibility, bergmann et al. (2015) found higher lexical alignment when direct evidence of conversation partner understanding was scarce. these authors showed more lexical alignment with an assumed artificial agent than with an assumed human (in fact both were artificial agents), but only when participants could not see (on video) their conversation partner. these results are consistent with alignment driven by an intention to aid communication because speakers adapted their lexical choice to a greater extent to a partner they judged as less communicativelycapable (such as a virtual agent as opposed to a human). but they also imply that visual feedback of the conversation partner’s behavior can eliminate this difference, presumably because being able to see one’s conversation partner gives greater certainty of comprehension success (reducing the need for lexical alignment to ensure this success). if so, it is possible that audience design would not increase alignment to non-native speakers overall, but only in the absence of evidence that they correctly understand what is being said. this possibility is consistent with evidence that speakers do not always consider conversation partners’ knowledge and needs because making such considerations is cognitively effortful. accordingly, horton and keysar (1996) showed that participants were less likely to take common ground into account in a referential communication task under time pressure than under no time pressure. the authors concluded that such audience design does not form a core part of designing utterances but is only part of an optional monitoring process, which we assume might be invoked when comprehension success is uncertain. we note that audience design, our current focus, is not the only mechanism that may drive lexical alignment. another mechanism that may underlie all alignment behavior is priming of the underlying linguistic representations (pickering and garrod, 2004). for lexical alignment, when one’s conversation partner produces a particular referential expression (such as the name sofa for an object that can also be called couch), this name becomes highly activated. this increased ac148 lexical alignment to non-native speakers tivation would in itself make it more likely that that name be selected for production over other possible choices. consistent with a priming mechanism contributing to alignment is evidence for similar lexical alignment in typically-developing children and children with an autistic spectrum disorder, whose theory-of-mind abilities or social functioning did not predict alignment magnitudes (branigan et al., 2016; hopkins et al., 2017). in the current study, we focus on the contribution of audience design to alignment to non-native speakers, but assume that priming underlies alignment to both non-native and native speakers at least in part (although possibly to different degrees; we return to this issue in the general discussion). a direct examination of alignment to non-native speakers was provided by bortfeld and brennan (1997). in a referential communication task, pairs of participants took six turns to be directors (describing a set of 15 chairs to the other person) and matchers (identifying each chair and placing it in the correct serial position). results showed that native speakers lexically aligned to non-native speakers to the same extent as to other native speakers (even if this involved alignment on nonnativelike expressions, such as the chair in which i can shake my body for a rocking chair). but this study did not investigate the relative strength of different mechanisms driving alignment with non-native speakers. the fact that native speakers aligned on non-nativelike expressions suggests an influence of audience design, as they would not have had any other motivation to produce such expressions except to ensure that their non-native partners understood them. on the other hand, non-nativelike expressions would not have direct representations in native speakers’ mental lexical inventories and thus would be unable to automatically prime native speakers’ productions. thus, alignment to non-native speakers in bortfeld and brennan’s study could have been greater (possibly greater than that to native speakers) had the non-native speakers produced more nativelike expressions. 1.1 the present study in the present study, we compared the magnitude of native speakers’ lexical-referential alignment to non-native speakers with that to native speakers, to test several hypotheses about the influence of audience design on alignment to non-native speakers in structured task-based dialogue. hypothesis 1 (tested in both experiments) is that native speakers would align more (and not less) to non-native speakers than to other native speakers, at least under certain conditions. this would be because native speakers would take greater care to ensure comprehension success with non-native speakers because they perceive the latter as less communicatively competent. hypothesis 2 (tested in experiment 1) is that such enhanced alignment to non-native speakers would be more, or only, present in conditions lacking feedback of comprehension success (that is, grounding). this would be because direct evidence that the conversation partner is understanding everything would alleviate the need for enhanced alignment with a less competent conversation partner to ensure communicative success. hypothesis 3 (tested in experiment 2) is that enhanced alignment to non-native speakers would be modulated by how communicatively competent they appear to be. this would be because enhanced lexical alignment as a strategy to ensure comprehension success may appear more necessary when the non-native conversation partner appears less communicatively competent. to determine the extent to which patterns of alignment with non-native speakers might generalize beyond particular populations or speakers of a particular language, we investigated two different populations (speakers of british english tested in edinburgh in experiment 1, and speakers of castilian spanish tested in barcelona in experiment 2). 149 ivanova, branigan, mclean, costa and pickering in a picture-matching game, participants took turns with confederates to name pictures for each other and match pictures after hearing the names produced by their partners. each native participant performed rounds of the game with both a native and a non-native confederate. the confederates (on half of the critical trials in experiment 1 and on all critical trials in experiment 2) produced dispreferred names (for example, calling a lamp light, or a basket, hamper, based on pretested norms; branigan et al., 2011). we focused on participants’ use of dispreferred names to detect differences in alignment across conditions, given that we expect speakers to use preferred names as their default choice (that is, independently of the identity of the conversation partner who produced them previously). in both experiments, each participant interacted with two confederates, in counterbalanced order. in experiment 1, native speakers of british english alternated with a native and a non-native confederate to name one of two pictures presented on a printed page, and indicate which of the two pictures on the following page matched the name produced by the other person. in addition, experiment 1 manipulated the presence of feedback of comprehension success. for this purpose, half of the naı̈ve participants received feedback from the confederate about their communicative success (who verbally indicated the correct response with the words left and right; participants did the same on their own matching turns). the other half of the participants matched pictures without receiving or providing any such evidence of comprehension success. in experiment 1, hypothesis 1 predicts more lexical alignment with the non-native than with the native confederate. hypothesis 2 predicts that such an effect would be larger (or only present) without feedback. in experiment 2, native speakers of castilian spanish alternated with a native and a non-native confederate to consecutively name 12 pictures presented on three-by-four grids, and select the picture named by the other person among an initial set of 12 loose pictures (decreasing after each trial), to then place it in its respective position on a blank three-by-four template with numbered grids. in addition, experiment 2 manipulated the apparent communicative competence of the non-native confederates. thus for one group of naı̈ve participants, the non-native confederate produced gender errors (for example, “*un mano” instead of “una mano” [sp. a hand]) on half of the trials eliciting alignment (and 1/12 of all trials). for another group of participants, the non-native confederate produced no errors. in experiment 2, hypothesis 1 predicts more lexical alignment with the non-native than with the native confederate. hypothesis 3 predicts that such an effect would be larger for the confederate who produced errors. note that we expected participants never to align to the erroneous gender markers – just to the picture names themselves. we acknowledge that there are reasons to treat with caution the use of confederates in dialogue experiments (kuhlen and brennan, 2013). we used confederates because dispreferred names are crucial for the detection of alignment but hard to elicit spontaneously. kuhlen and brennan note that in such a case the use of confederates may be the only feasible approach (p.56, p.65). in accordance with their recommendation, we document relevant details about confederate knowledge and behavior in the method section, and discuss how confederate participation could have affected the results in the general discussion. we also note that the current experiments sought to determine the mechanisms of a specific type of language use and thus had features that made them quite different from many natural dialogues. some relevant features were the overall ease of the task, the preponderance of referring expressions (over non-referential language) together with the otherwise limited conceptual content, the lack of grounding in most conditions, and the formulaic grounding in the feedback group of 150 lexical alignment to non-native speakers experiment 1; we consider in the general discussion how these features could have affected the studied mechanisms. 2. experiment 1: lexical alignment to native and non-native speakers of english with or without feedback of comprehension success in this experiment, we contrasted native english speakers’ lexical alignment to native and non-native confederates in a picture-matching task. half of the participants and confederates gave and received feedback about successfully matching the pictures, while the other half matched pictures without such feedback. 2.1 method in this section, we provide details about the participants and confederates, the materials and procedure, and the design, coding, and data analysis in experiment 1. 2.1.1 participants forty-eight students from the university of edinburgh community were paid to take part. all participants were native speakers of english. the data from one participant in the no-feedback group was excluded from the analyses because of a counterbalancing error. there were thus 24 participants in the feedback group, and 23 in the no-feedback group. 2.1.2 confederates four confederates who were non-native speakers of english were selected on the basis of a pretest. nineteen non-native speakers and one native speaker of english were recorded producing the preferred and dispreferred names of the experimental items described below (36 words in total). after, twelve further participants from the university of edinburgh community who did not participate in the main experiment listened to these recordings (in randomized order for each rater) and indicated the perceived strength of the speaker’s foreign accent on a likert scale (1 meant “extremely strong foreign accent” and 7 meant “native english accent”). four non-native speakers with mediumstrong accents (to ensure intelligibility; average ratings of 3.25, 3.33, 3.75 and 4.08) were selected for participation in the main experiment. t-tests on the ratings of each two pairs of speakers for all possible pairings indicated that the ratings for all four speakers did not differ [all ps > .17]. four further native speakers of english were selected to act as the native confederates. each of these speakers was paired with one of the non-native speakers to form same-gender pairs (three female and one male pairs). all confederates were paid for their time. the confederates were aware of the experimental hypotheses. confederates arrived when their turn was scheduled and were not present in the room when another confederate was performing the task – thus, they did not have information about the behavior of the confederate they were paired with. each confederate performed the task 11 or 12 times (see below). 2.1.3 materials and procedure the materials consisted of black-and-white line drawings and were the same as in branigan et al. (2011). there were 18 experimental pictures, each of which could be named with both a preferred 151 ivanova, branigan, mclean, costa and pickering and a dispreferred but acceptable name (for example, lamp-light). the names were normed by branigan et al. (2011) for the same population as tested here. in addition, there were 154 filler pictures which only had one name and appeared between one and eight times in the combined participants’ and confederate’s naming and matching sets. the materials were presented in binders, one for the participant and one for the confederate (see figure 1 for the experimental set-up). in the participant binder, two pictures were displayed on each page (an a4 sheet inside a plastic envelope). on a naming turn, the picture to be named was indicated by a thicker black border, and appeared on the left side on half of the trials, and on the right side on the other half. on a matching turn, both pictures had the same thickness border. matching was performed by placing a sticker on one of the two pictures on every matching sheet. in the confederate binder, each naming turn had a single word printed in the center of the page. matching trials had the same structure as participants’. on each naming or matching turn, an individual filler picture appeared next to an experimental picture. on critical turns, the confederate named an experimental picture with either its preferred or its dispreferred name (counterbalanced across items), and the participant located it on the corresponding matching page. after two filler turns (participant naming trial, confederate naming trial), the participant had to name the same experimental picture (again, presented together with a filler picture). there were 14 filler turns at the beginning of the experiment (the first four were presented as practice). there were two such beginning-filler sets, counterbalanced across lists, confederate orders and the four confederate pairs. after that, six filler turns separated each 4-trial sequence involving experimental pictures; those were kept constant within lists. figure 1: experimental paradigm for the picture-matching game used in experiment 1. the confederate has performed a naming turn, and the participant has indicated their choice using a sticker (indicated as a yellow rectangle). 152 lexical alignment to non-native speakers there were eight item lists. four lists (list set a) were created by counterbalancing confederate names for the experimental pictures (preferred, dispreferred) and the side on which experimental items appeared on participant naming and matching trials (left, right). the content and order of filler trials was the same across these four lists except the beginning filler sequence, as explained above). four further lists (list set b) were created by randomizing the order of the filler trials and experimental sequences (but the order of items within an experimental sequence was the same as in the first four lists). across the latter four lists, the content and order of filler trials was also kept the same. within a single list, there were 152 confederate naming trials, and 152 participant naming trials. each participant performed the matching game with both confederates within a confederate pair, one a native speaker of english, and the other a non-native speaker, in counterbalanced order. twelve participants performed the experiment with each of the first three confederate pairs, and 11 participants performed the experiment with the last confederate pair; confederate pairs were equally distributed across the feedback and no-feedback groups. each participant performed the experiment with a same-gender confederate pair. if the first part of the experiment (interaction with the first confederate) used a list from list set a, the second part of the experiment (interaction with the second confederate) used a list from list set b that contained the alternative names for experimental items (for example, if the first confederate named a picture of a basket with its dispreferred name hamper, the second confederate named the same picture basket). the order of list sets was counterbalanced across participants. upon arrival, participants and confederates were assigned to their seats, to ensure that confederates did not always head to the same side out of habit, which could indicate that they had been in the room before. seat order was counterbalanced, such that half of the time confederates sat facing the window and half of the time they sat facing the door. participants and confederates read written instructions that explained the procedure involving turn-taking with the other person to name and match pictures. participants were not told that they would perform the task with two different partners at the beginning, but only upon arrival of the second confederate (approximately 30 minutes after the beginning of the experimental session, the time taken to complete the task with the first confederate). the instructions explained that to-be-named pictures were surrounded by a thicker black border and that matching was to be performed by placing a sticker on the matching picture. additionally, participants and confederates in the feedback group were instructed to give verbal feedback of which pictures they matched. they did this by saying “left” or “right”, depending on the position of what they thought was the correct matching picture on a given matching sheet (and they were always correct because the task did not leave any possibility for confusion for a native speaker). importantly, they were also instructed to pay attention to this verbal feedback when it came from the other person, and place a sticker on their own naming sheet to indicate which picture their partner had selected. this was done to encourage participants to attend to the feedback. once participants or confederates performed a turn (for a naming turn, produced a name, heard the feedback from the other participant and placed a sticker accordingly; for a matching turn, heard the name produced by the other person, decided on the correct matching picture, placed a sticker on it and said “left” or “right” to indicate its position), both participants turned a page in their binders; this indicated the beginning of the next trial. the procedure for participants in the no-feedback group was identical, except that no feedback was given (or monitored) by either participants or confederates with respect to their matching choices. all experimental sessions were recorded for subsequent verification. 153 ivanova, branigan, mclean, costa and pickering 2.1.4 design, coding, and data analysis responses were transcribed in real time by the first author and were recorded for subsequent verification. we report here analyses of the responses following dispreferred names produced by confederates, to the extent that they carry condition differences in alignment effects (while preferred responses following preferred confederate names were close to ceiling, 95.3%). analyses of responses following preferred confederate names are reported in appendix a. the data were analyzed with logistic mixed-effects regression modeling. dispreferred responses (matching the names produced by confederates) were coded as 1, and preferred responses (different from the names produced by confederates) were coded as 0. the fixed predictors were confederate type (native, coded as 0.5, and non-native, coded as -0.5), feedback group type (feedback, coded as 0.5, and no feedback, coded as -0.5), confederate order group (native-first, coded as 0.5, and non-native first, coded as -0.5) and their interactions. the confederate order group factor was included in the analyses because of the following considerations. first, evidence from structural alignment suggests that speakers adapt to the frequency of alternatives in an experiment, resulting in a gradual decrease of alignment throughout an experimental session (fine and jaeger, 2013). in our experiment, confederate order was counterbalanced (half of the participants performed the picture-matching task with a native confederate first and a non-native confederate afterwards, and for half of the participants this assignment was reversed), but alignment decrease because of adaptation might be differentially influenced by confederate type. further, in this experiment the second confederate produced the alternative name to the name produced by the first confederate. for example, if the native confederate was the first to interact with a participant and produced a preferred name (for example, lamp), the non-native confederate subsequently named the same picture with the dispreferred name (for example, light), or vice versa. because of this design, participants might align more readily with a name (thus, accept a label for a given object, for example lamp) if they have not yet themselves produced a name for that object, but then be less likely to adopt a different name once they have named this object because speakers tend to repeat themselves (see experiment 3 in branigan et al., 2011; wheeldon and monsell, 1992; brennan and clark, 1996, though note that speakers in dialogue can change the way they refer to objects with each different conversation partner). for both of these reasons, it is possible that alignment to the native and non-native confederates was differentially impacted by the position of the respective interaction within our experiment. note that the decision to include the confederate order group as factor was post-hoc but analyses without it produced an equivalent pattern of results for the remaining factors. the models were run using the glmer function in the lmertest package (version 2.0-33, lme4 version 1.1-13) in r (version 3.4.1). to aid convergence, the “bobyqa” optimizer was used. all models initially had the maximal random-effects structure justified by the design (barr et al., 2013). in case of non-convergence of the full model, model simplification was performed by first removing random-effects correlations and subsequently removing in a step-wise manner the random effects accounting for least variance (random slopes were removed before random intercepts). to shed light on significant interactions, further models were run on subsets of the data; these models are described in the results section below. the data are publicly available at https://osf.io/7rtgd/, and analyses scripts will be offered upon request. 154 lexical alignment to non-native speakers 2.2 results and discussion the by-participant percentage aligned responses are plotted in figure 2 and the statistical models results are reported in table 1. we first report the results of the main model described above. after, we report analyses breaking down the data into feedback and no-feedback groups, and, within each of these, into native-first and non-native-first groups. these breakdowns were done because differential effects of feedback across confederate types is predicted under hypothesis 2 (and visible numerically on figure 2) but power may not be sufficient to detect the interactions of interest. the main analysis revealed globally more alignment for dispreferred responses to the non-native (39.7%) than to the native confederate (33.6%; the main effect of confederate type was significant). further, participants in the native-first group aligned more overall (42.8%) than participants in the non-native-first group (30.9%; the main effect of confederate order group was significant). this difference, however, was due to participants in the no-feedback group (there was a significant interaction between feedback group and confederate order group). specifically, feedbackgroup participants who interacted with the native confederate first showed similar overall alignment to feedback-group participants who interacted with the non-native confederate first (a difference of 1.7%), while no-feedback-group participants who interacted with the native confederate first showed 25% more alignment than no-feedback-group participants who interacted with the nonnative confederate first. these differences might suggest an effect of confederate type whereby interacting with a non-native confederate first globally reduced the rate of alignment to dispreferred names throughout an experimental session. however, we believe these differences mostly likely stem from individual differences in alignment to dispreferred names because they come from a between-participant comparison; as such, we do not interpret them further. participants also aligned more to the confederate with whom they interacted first, regardless of confederate type (there was a significant interaction between confederate type and confederate order group). but this tendency was more pronounced when the first conversation partner was the non-native confederate (17% more alignment with the non-native than with the native confederate) than when it was the native confederate (5.3% more alignment with the native than with the nonnative confederate; note that this is a between-participants comparison). we then ran separate models on the data of each feedback group type, with the fixed predictors confederate type, confederate order and their interaction. participants in the feedback group aligned overall to a similar extent to the native and non-native confederates (no significant main effect of confederate type). however, they had a tendency to align more with the confederate they interacted with first (a significant interaction between confederate type and confederate order group). in other words, when there was no need for concern about comprehension success, there was less alignment with the second consecutive conversation partner than with the first, independent of that partner’s native language. (this tendency was stronger for the non-native-first group, who showed 17.3% more alignment with the first non-native confederate, and a significant simple effect of confederate type, than for the native-first group, who showed 13.1% more alignment with the first native confederate but no significant simple effect of confederate type.) importantly, participants in the no-feedback condition aligned more to the non-native than to the native confederate (a significant effect of confederate type). further, as indicated by the main analysis, participants in the native-first group aligned more overall than participants in the non-native-first group (a significant effect of confederate order group). 155 ivanova, branigan, mclean, costa and pickering simple effects models on the data for the two confederate order groups with confederate type as a fixed predictor showed more alignment with the non-native than with the native confederate (a difference of 16.7%) for participants who interacted first with the non-native confederate (a significant effect of confederate type), but no difference in alignment to the two confederates (1.9%) for participants who interacted first with the native confederate (no significant effect of confederate type). these analyses suggest that, without evidence for comprehension success, the tendency to align less to dispreferred names from a second partner occurred only when this partner was the native confederate; when the non-native confederate was second, alignment did not differ from alignment with the first confederate. figure 2: mean by-participant percentage alignment effects by feedback group type, confederate order group and confederate type in experiment 1. error bars represent standard error (computed from by-participant means). nat 1st: native-first group; non-n 1st: nonnative-first group. the numbers inside the bars indicate condition order in the experiment. in sum, experiment 1 revealed that, with feedback, participants aligned less to their second than to their first partner, regardless of their native language. without feedback, participants aligned less to their second than to their first partner when the second partner was a native speaker, but maintained alignment at its initial level when the second partner was a non-native speaker. these results support hypothesis 1, which is about the direction of the confederate type effect: under certain conditions, alignment to the non-native speaker was larger (not smaller) than that of the native speaker. this directionality supports an influence of audience design in alignment to nonnative speakers. the results also support hypothesis 2 in that such effects appeared only in the absence of feedback. this pattern is also consistent with an influence of audience design but suggests that audience design is used only when deemed necessary (i.e., in the absence of evidence for comprehension success). additionally, we found that alignment tended to decrease from the first to the second consecutive conversation partner. this effect could have sources in participants’ adaptation to the frequency of dispreferred alternatives within the experiment (for such an inverse preference effect in structural alignment, see jaeger and snider, 2013), as well as participants’ tendency to self-repeat (branigan 156 lexical alignment to non-native speakers et al., 2011; wheeldon and monsell, 1992): if participants aligned with the first confederate (for example, said lamp), they might have then been less willing to switch to a different name with the second confederate (for example, say light). model predictors estimate se z p main model confederate type -.41 .17 -2.33 .02 feedback group .29 .39 .75 .45 confederate order group .77 .39 1.99 .047 confederate type * feedback group .51 .35 1.45 .15 confederate type * confederate order group 1.40 .42 3.37 <.001 feedback group * confederate order group -1.67 .77 -2.15 .03 confederate type * feedback group * confederate order .31 .72 .44 .66 feedback group confederate type -.14 .29 -.50 .62 confederate order group -.07 .60 -.11 .91 confederate type * confederate order group 1.58 .62 2.55 .01 native-first group confederate type 1.54 1.08 1.43 .15 non-native-first group confederate type -.93 .45 -2.05 .04 no-feedback group confederate type -.69 .32 -2.19 .03 confederate order group 1.62 .53 3.06 .002 confederate type * confederate order group 1.22 .69 1.76 .08 native-first group confederate type -.07 .47 -.14 .89 non-native-first group confederate type -1.12 .37 -3.04 .002 table 1: results of the lmer main model of dispreferred responses in experiment 1. note: grey rows indicate significant effects. 3. experiment 2: lexical alignment to native and non-native speakers of spanish with or without grammatical errors in this experiment, we contrasted native spanish speakers’ lexical alignment to dispreferred names produced by native and non-native confederates in a different version of the picture-matching game. the non-native confederate made grammatical (gender) errors on half of the target pictures with one participant group and made no errors with another group, so that we could see if errors from the conversation partner (a signal of lower communicative competence) would increase alignment. an additional aim of this experiment was to determine the extent to which the findings of experiment 1 would generalize to another population and language. 3.1 method in this section, we provide details about the participants and confederates, the materials and procedure, and the design, coding, and data analysis in experiment 2. 157 ivanova, branigan, mclean, costa and pickering 3.1.1 participants sixty-four undergraduate students from the university of barcelona, (barcelona, spain), took part in exchange for course credit. all participants were native speakers of castilian spanish (spoken in their home by both parents, and currently spoken by participants more than 50% of the time). participants also spoke catalan (the population of barcelona is almost entirely bilingual). for forty participants, the non-native confederate made no grammatical errors (no-error group), and for twenty-four participants the confederate produced grammatical errors on half of the experimental items (error group). we acknowledge the uneven number of participants in the two error groups. participant testing in the current no-error group was originally conducted as two separate experiments, which we now combine for greater power. the two combined experiments followed an identical procedure. the only differences were that the confederates were different individuals, and that two items, described below, were changed. sample sizes were originally chosen to approximate those in the study of branigan et al. (2011), which were 16 in experiments 1 and 3, 20 in experiment 5, 24 in experiment 4 and 32 in experiment 2. the uneven number of participants in the two groups was adjusted in the statistical analyses by centering the numerically coded error group factor around the mean. 3.1.2 confederates the confederates for the error group and the first 16 participants in the no-error group were two female graduate students from the university of barcelona, a native speaker of castilian spanish and a native speaker of german. the confederates for the remaining participants in the no-error group were a female graduate student from the university of barcelona, a native speaker of castilian spanish, and a female student from the university of barcelona community, a native speaker of american english. the confederates were aware of the experimental hypotheses. they were not present in the room when another confederate was performing the task. each confederate performed the task 24 or 40 times. confederates saw pictures with the names written underneath for both experimental and filler items. 3.1.3 materials and procedure the experimental items were 22 further black-and-white line drawings, which could be named with a preferred and a dispreferred name (for example, puño [sp. fist] – mano [sp. hand]). both names for all stimuli are listed in appendix b, but note that confederates produced only dispreferred names in this experiment. the dispreferred names were always more frequent than the preferred names, to lend credibility to the fact that they were spontaneously produced by non-native speakers (likely not expected to know low-frequency words). the experimental pictures were selected on the basis of a pretest from an initial set of 53 pictures, as follows. thirty-two further participants from the same population and who did not take part in the main experiment were asked to name the 53 pictures. they also rated the appropriateness of the dispreferred name for each picture on a 10-point scale. these tasks were presented in counterbalanced order. a stimulus set of 24 of these items was then tested with seven pilot participants using the paradigm in experiment 1. these 24 stimuli involved two changes from the normed ones: the picture of a pigeon was replaced with a picture of a parrot (both had pájaro [sp. bird] as dispreferred name); and the dispreferred name via [sp. way] for the preferred carretera [sp. freeway] was replaced with camino [sp. road]. the pilot revealed numerically greater alignment to the non-native 158 lexical alignment to non-native speakers than to the native confederate. the paradigm was changed to the one described below to enhance ecological validity and reinforce beliefs that the task would be hard for a non-native speaker. from the items in the pilot, four were removed and three were replaced, for example because some of the pictures were hard to recognize (e.g., tree trunk; braid), leaving 20 items. this stimulus set was used with 16 participants in the no-error group. the preferred names in this set were spontaneously produced 96.4% of the time on average (sd = 4.8%; norming data); the corresponding dispreferred names had medium acceptability (m = 6.11, sd = 1.04). for the remaining 24 participants in the no-error group and all participants in the error group, two further pictures (baby and bride) were replaced (with fridge and earring) because gender errors on entities with a natural gender (for example, una niño [a (fem.) male child] or un mujer [a (masc.) woman] would have appeared exceptionally strange. the experimental set-up (see figure 3) consisted of two description sets (two sets of pictures (a) arranged on three-by-four-grids on ten powerpoint slides per set and then printed on a4 sheets of paper), two match sets (two sets of loose pictures (b), prepared by cutting up different copies of the description sets into individual pictures), and two sets of blank templates (c) printed on a4 sheets of paper with numbered three-by-four-grids on each of them. each of the ten sheets in one description set, together with their corresponding loose pictures, made one round of the game for one participant. the confederate’s description sheet additionally contained all picture names written under their corresponding pictures. to create the description sets, 20 experimental pictures, together with 176 filler pictures, were arranged on ten three-by-four grids drawn on microsoft powerpoint slides (some fillers occurred on both the participant and confederate description sets, and some were unique to either the participant or confederate set). there were two experimental pictures on each slide in each round of the game. an experimental picture always occurred first in the confederate’s description sheet and was preceded by at least two filler pictures. the same picture then occurred on the participant’s description sheet, after two filler naming turns (participant’s and confederate’s; this separation was necessary to make the repetition less obvious). each participant saw and named the 20 experimental items only once with each confederate. for this purpose, the 10 rounds were divided into two sets of five rounds, one to be performed with each confederate. the filler pictures in both the confederate’s and participant’s description sets contained pictures whose names were similar in meaning or form to the target names (for example, toothpaste, toothbrush, dentist, for the target denture (preferred) – teeth (dispreferred)), to make the matching task more challenging, presumably especially for a non-native speaker. the pictures in the confederate’s description set had more frequent names than the ones in the participant’s description set, to reinforce the idea that the non-native confederate had limited vocabulary (that is, was not fully proficient in spanish); but note that these pictures were the same for the native confederate. on every round, participant and confederate each received a description sheet, twelve match pictures from their match set (containing the pictures from the other person’s current description set) and a blank template. they then took turns to name a picture from their description set (moving horizontally from left to right), and then to find the picture just named by the other person and place it on its corresponding position on the blank template. the confederate and participant each went first on half of the rounds, such that the participant went first on all rounds performed with the first confederate and second on all rounds performed with the second confederate, or vice versa. 159 ivanova, branigan, mclean, costa and pickering to make the task somewhat more naturalistic, participants and confederates were instructed to produce a whole sentence instead of just the picture name. an example exchange went as follows: confederate: mi primer dibujo es un libro [my first picture is a book] participant (after matching): mi primer dibujo es una pelota [my first picture is a ball] confederate (after matching): mi número dos es un espejo [my number two is a mirror] participant (after matching): mi número dos es una vela [my number two is a candle] on critical naming sequences (independently of who had the first turn in the round), the confederate named an experimental picture, always using the dispreferred name (for example, mesa [sp. table] for a picture of a desk). this was followed by two filler turns (for example, the participant then named a picture of a chair, and the confederate named a picture of a bookcase), after which it was the participant’s turn to name the experimental picture (desk in this example). the experimenter, seated behind the participant and confederate, noted the name produced by the participant. additionally, for the error group, the non-native confederate made a grammatical error on one experimental trial per round (thus, on half of the experimental trials in total). the errors consisted in producing feminine indefinite determiners for words of masculine gender (for example, *un terraza [a (masc.) terrace]; *un mano [a (masc.) hand]) or vice versa (for example, *una pollo [a (fem.) chicken; *una juguete [a (fem.) toy]. there were eight experimental versions for the no-error group, obtained by crossing confederate order (native first or non-native first), first-turn assignment (confederate or participant), and order of experimental items within a round. for example, if the picture of a highway was the first experimental item and the picture of an ambulance was the second experimental item on a given round for half of the participants, the ambulance was first and the highway second for the other half. there were additional eight versions for the error group, obtained by crossing the presence or absence of an error on each experimental item. further, half of the participants in both groups saw one 5-round set first and half saw the other 5-round set first, but this factor was not fully counterbalanced (that is, no additional groups were created beyond the ones described above). the order of the rounds within each 5-round set was randomly varied among participants. upon arrival in the lab, the participant and first confederate were seated side by side at a table with an occluder between them, such that they were not able to see each other’s picture sets (see figure 3). as in experiment 1, the second confederate arrived later, when the task with the first confederate was completed, and participants were not told in advance that they would interact with two confederates. they read written instructions which informed them that the experiment involved a collaborative game and then stepped them through the procedure. the instructions further stated that each participant would receive “something sweet” (a kitkat bar) if both of them completed all rounds correctly, to further encourage a collaborative spirit and incentivize them to pay attention to their and their partner’s performance. the game with each confederate began with a practice round consisting of six description pictures. this was followed by five experimental rounds, each consisting of 12 pictures. once a round was completed, the experimenter verified that it was completed correctly (performance was at ceiling). the experimental sessions were recorded for subsequent verification. 3.1.4 design, coding, and data analysis responses were coded as in experiment 1. the dependent variable was aligned responses, which were all dispreferred responses. the fixed predictors in the main model were confederate type 160 lexical alignment to non-native speakers (native, coded as 0.5, and non-native, coded as -0.5), confederate error group (no-error group, coded as -0.5, and error group, coded as 0.5), confederate order group (native-first, coded as 0.5, and nonnative first, coded as -0.5) and their interactions. confederate order group (a predictor determined post-hoc) was included in the model for comparison with experiment 1; a model without this factor and its interactions produced an identical pattern of results for the other factors. to shed light on interaction terms, further models (described below) were run on subsets of the data. figure 3: experimental paradigm for the picture-matching game in experiment 2. 3.2 results and discussion the by-participant mean proportions of dispreferred responses by confederate type, error group and confederate order group are plotted in figure 4, and the statistical models results are reported in table 2. the main analysis revealed that participants produced more dispreferred responses overall when they interacted with the non-native confederate (40%) than when they interacted with the native confederate (31.1%) – that is, there was more alignment with the non-native than with the native confederate (a significant main effect of confederate type). further, participants experiencing error-free descriptions from the non-native confederate (noerror group) showed more overall alignment (41.3%) than participants experiencing gender errors from the non-native confederate (error group; 26.1%; confederate error group was a significant predictor). however, when the native confederate was first, participants in the no-error group aligned more even to the native confederate relative to participants in the error group in the same condition (estimate = -.79, se = .40, z = -1.98, p = .048); the linguistic behaviour of the native confederate did not differ between the two confederate error groups. this suggests that the alignment differences between the two groups were most likely due to random variability. 161 ivanova, branigan, mclean, costa and pickering we further report analyses of subsets of the data parallel to those we conducted in experiment 1. first, we conducted separate analyses of each confederate error group, with confederate type, confederate order group and their interaction as fixed predictors. for the error group, this analysis showed that alignment with the non-native (23.0%) and native confederates (29.2%) did not differ (no main effect of confederate type). however, for the no-error group, there was significantly more alignment with the non-native (46.5%) than with the native confederate (36.0%) (a main effect of confederate type). the simple effects of confederate order group for the no-error group further showed more alignment to the non-native than to the native confederate for participants who first interacted with the non-native confederate, but similar alignment to the non-native and native confederates for participants who first interacted with the native confederate. there were no other significant effects. in sum, participants aligned their lexical choice more with the non-native than with the native confederate when the non-native confederate did not produce any grammatical errors. these results support hypothesis 1 and point to a role of audience design in lexical alignment to non-native speakers in the structured task-based dialogue studied here. they also replicate with a different population and in a different language the greater alignment to non-native than to native speakers without feedback found in experiment 1. figure 4: mean by-participant percentage aligned (dispreferred) responses by error group, confederate order group and confederate type in experiment 2. error bars represent standard error (computed from by-participant means). the numbers inside the bars indicate condition order in the experiment. however, experiment 2 did not support hypothesis 3: when the non-native confederate produced grammatical errors, alignment was similar to alignment with the native confederate (instead of being even larger than when the non-native confederate did not make errors). this finding suggests that the presence of errors does not uniquely signal the need for enhanced lexical alignment to ensure comprehension success. instead, we suggest it may introduce competing motives, such as the desire to serve as an example of correct language use (i.e., model correct behavior). 162 lexical alignment to non-native speakers we probed further into whether such behavior affected only trials with errors or the task-based exchange globally, and we found support for the latter. that is, we compared alignment to the nonnative confederate on trials with errors (29.8%) to that on trials without errors (31.2%) and did not find a statistical difference between the two (estimate = -.30, se = .31, z = -.97, p = .33; this model had only a random intercept for items). we discuss possible roles of grammatical errors for lexical alignment to non-native speakers in structured task-based dialogue in the general discussion. lastly, experiment 2 did not reveal the second-confederate alignment reduction found in experiment 1. one plausible explanation for this pattern is that participants in experiment 2 were exposed to and produced names for the experimental pictures with only one confederate (that is, never had to name the same picture twice). we note, however, that there were many other differences between the two experiments. model predictors estimate se z p main model confederate type -.51 .14 -3.75 <.001 confederate error group -.85 .25 -3.46 <.001 confederate order group -.08 .24 -.36 .72 confederate type * confederate error group .19 .29 .67 .50 confederate type * confederate order group .34 .27 1.25 .21 confederate error group * confederate order group .33 .49 .67 .51 confederate type * confederate error group * confederate order group -.67 .57 -1.16 .25 no-error group confederate type -.63 .26 -2.38 .02 confederate order -.23 .33 -.69 .49 confederate type * confederate order .65 .47 1.39 .17 native-first group confederate type -.29 .30 -.98 .33 non-native-first group confederate type -.92 .38 -2.41 .02 error group confederate type -.40 .34 -1.19 .24 confederate order .07 .38 .19 .85 confederate type * confederate order -.05 .70 -.07 .94 native-first group confederate type -.45 .75 -.60 .55 non-native-first group confederate type -.11 .45 -.25 .80 table 2: lmer results for experiment 2 4. general discussion this study tested three hypotheses about the role of audience design in lexical-referential alignment with non-native speakers in two structured task-based dialogue experiments. in experiment 1 with speakers of english, there was greater alignment to dispreferred names produced by non-native confederates than to such names produced by native confederates in a participant group that did not receive feedback about the confederates’ comprehension success. however, there was equivalent alignment in another group that did receive such feedback. experiment 2 with speakers of spanish replicated the greater alignment to non-native confederates in the absence of feedback (shown only in analyses of subsets of the data), but found equivalent alignment for another group that received descriptions from the non-native speakers that included some grammatical errors. the native/nonnative differences in both experiments were driven by maintenance of alignment at the initial rate when the second partner was a non-native speaker, relative to a reduction of alignment when the second partner was a native speaker. 163 ivanova, branigan, mclean, costa and pickering these results support our hypothesis that, under certain conditions, alignment with the nonnative speaker would be larger (and never smaller) than with the native speaker (hypothesis 1). such differences suggest an influence of a form of audience design (the desire to aid communicative success) on lexical-referential alignment with non-native speakers in structured task-based contexts, across speakers of different languages. our findings are thus in line with studies suggesting that speakers can override their own production preferences to adapt their lexical choices to their partners’, and that they tend to do so when they believe or see that their partners are less competent speakers or have less expertise in the relevant domain (branigan et al., 2011; isaacs and clark, 1987). the finding that alignment with the non-native speaker was greater only when evidence about their performance was absent also supports our hypothesis that enhanced alignment to non-native speakers would be more, or only, present in conditions lacking grounding (hypothesis 2). this result shows that audience-design driven alignment is not indiscriminate, but adapts to the specifics of the situation (i.e., uncertainty about one’s conversation partner’s comprehension success). without feedback (experiment 1) and without confederate errors (experiment 2), alignment decreased from the initial levels when the native confederate was second, but remained unchanged when the non-native confederate was second. this finding raises the possibility that participants initially intentionally chose to align at a high rate before they were able to assess the overall difficulty of the task (that is, at the beginning of the experiment), but by the second block had realized that the task was easy enough for a native speaker to perform and there was no need for extra care in their referential choice (that is, alignment). however, when introduced to the non-native confederate in the second block and in the absence of evidence of their comprehension success, they might have wanted to ensure that the non-native confederate was able to perform the task, and hence continued to align at the initial level. this possibility further suggests a role for audience design, in that it attributes adjustments to the magnitude of alignment to top-down decision-making processes. these findings also fit with broader evidence that speakers engage in partner modeling, or adaptation of one’s own language production based on expectations of the addressee’s knowledge (horton and gerrig, 2002). here, we show its influence on alignment behavior even though evidence suggests that it is a resource-demanding process (horton and keysar, 1996; vogels et al., 2015). interestingly, the results of experiment 2 did not support our hypothesis that the greater alignment with the non-native confederate would increase still further when this confederate gave descriptions with grammatical errors relative to when they gave descriptions without errors (hypothesis 3). such an increase might have been expected under an audience design account because grammatical errors indicate even more limited competence in the target language. there could be several possible explanations for why the predicted increase was not found. first, the presence of grammatical errors may have triggered a desire in participants to correct non-native speakers’ errors or infelicities and thus model correct usage of the language (see e.g., kurhila, 2001). consistent with this, participants never produced incorrect gender markers themselves, even when they reused the dispreferred picture names on trials with errors (to the recollection of the experimenter, the first author; the recordings of the experiments are no longer available to verify). the act of always producing correct gender markers, thus correcting the confederate’s gender errors, may have suggested or intensified the decision to also suggest better names for the pictures. we note that such behaviour is ultimately collaborative (and likely kept in check by the fear of appearing rude or excessively pedantic), in that it aims to improve non-native speakers’ ability to function in the target language. 164 lexical alignment to non-native speakers another possibility (less likely in the current setting because it was not explicitly manipulated) is that native speakers did not want to imitate the behaviour of speakers who made language errors. divergence from the speech of one’s conversation partner can be used to indicate disaffiliation (bourhis et al., 1979; doise et al., 1976; ludlow, 2014). this possibility is consistent with results from studies on structural alignment. for example, heyselaar et al. (2017) showed that participants structurally aligned less with avatars that had a computerized voice and did not exhibit typical interactional human behavior such as facial expressions and looks to the conversation partner, relative to human-like avatars. weatherholtz et al. (2014) reported that the greater the perceived distance between recorded speakers’ accents and participants’ own, the less participants structurally aligned with this speaker. we caution, however, that findings of studies of one type of alignment may not generalize to studies of a different type of alignment, because different levels of linguistic representation might be differentially sensitive to different mechanisms driving alignment. for example, syntactic misalignment would not normally result in miscomprehension, because a speaker can assume that even a less proficient conversation partner has some mastery of common syntactic alternations, as well as because utterance meaning is at least partially conveyed in lexical items and is not only carried by syntactic structure. in contrast, naming an object in a way the conversation partner does not understand can result in direct lack of comprehension. further research is thus needed to systematically examine the interplay of mechanisms driving alignment at the different linguistic levels, and their similarities and differences. our results have implications for the increasingly common conversations between native and non-native speakers in real life. first, they suggest that the communicative difficulties faced by non-native speakers might be at least partially offset by native speakers aligning to their lexical choices. second, the alignment of native speakers to non-native speech might, in the long run, have implications for language change (especially of english, which is spoken by many non-native speakers around the world): after aligning to certain non-native uses of words or expressions, native speakers might become more likely to adopt them even in conversations with native speakers, who might in turn align to them and adopt them in their own vocabularies. to gain more insight into this process, it would be necessary to investigate whether those native speakers’ lexical choices that result from alignment with non-native speakers tend to persist with other speakers or remain partner-specific. we note that the structured, task-based dialogue in our experiments had many differences from natural dialogue. one such difference was the participation of confederates. kuhlen and brennan (2013) point out that confederates may unwittingly behave in a way that biases participants to engage in the hypothesized behaviors, that they may show more knowledge than is assumed for somebody unfamiliar with the task, that participants may behave differently if they suspect that they are speaking to lab assistants rather than naı̈ve participants like themselves, and that confederates that have performed the experiment multiple times and are reading from a script do not behave naturally. features of our design undermine some of these concerns but we cannot completely exclude them. first, the critical comparison was across two different confederates rather than between two conditions involving the same confederate engaging in different behaviors. second, the likelihood of confederate behavior that consistently biased in a single direction seems undermined by the fact that there were four confederate pairs in experiment 1 and two confederate pairs in experiment 2, and by the different patterns for the feedback and no feedback vs. error and no-error groups (although we 165 ivanova, branigan, mclean, costa and pickering cannot exclude the possibility that confederate behavior affected differences between the feedback and no feedback conditions). third, confederates did not know how the other confederate in their (native/non-native) pair behaved because they were not in the room during the other confederate’s part of the session. fourth, if participants had suspected confederate status, they would likely have inferred that the non-native confederates did not need help to perform the task; as such, we should not have observed more alignment to the non-native confederates. apart from the participation of confederates, the referential communication task we used had a greater proportion of referential expressions than a typical conversation, but no other conceptual content. lexical-referential alignment may thus be overall much less in a natural conversation because of the fewer opportunities for it to occur; consistent with this, more alignment was found in a corpus of task-oriented conversations than in a corpus of unconstrained telephone conversations (reitter et al., 2006). as such, differences between alignment with native and non-native speakers may be obscured or reduced in such unconstrained contexts compared with our study. another relevant difference between many natural conversations and our study is the presence of overt grounding (which occurred in our study only in the feedback group) or explicit clarification requests prompting overt explanations. as we have shown, overt evidence for comprehension success modulates alignment to non-native conversation partners, and may make alignment less necessary as a comprehension-ensuring strategy in more naturalistic situations. further, the social desirability of alignment (cf. communiucation accommodation theory, giles and powesland, 1975; hopkins and branigan, 2020) was not specifically manipulated here but may be a strong determinant of lexical-referential alignment in real-life situations, many of which also have a greater emotional component than our experiments. as such, factors that exercise a stronger influence than audience design may ultimately determine alignment to non-native (and also native) speakers in many contexts. lastly, natural conversation may be substantially more cognitively effortful than the simple tasks used here – for example, because of the need to plan more complex utterances or concurrent activities. too great a cognitive load or time pressure may reduce or eliminate audience design (for example, horton and keysar, 1996) and leave only alignment based on automatic priming. this may eliminate or reverse the greater alignment to non-native speakers found in our experiments. considering these differences, our study shows that audience design is a mechanism that can act to enhance lexical-referential alignment to non-native relative to native speakers – but it remains to establish the extent of its influence in natural conversations. we have concluded that audience design played a role in alignment to non-native speakers but did it play a role in alignment to native speakers? we speculate that such alignment may involve audience design in at least some contexts, but that the underlying influence of automatic priming may be stronger for native than for non-native conversation partners. if so, we can tentatively hypothesize that alignment to non-native speakers is more cognitively demanding than alignment to native speakers. the greater cognitive demand would be imposed by the need to assess non-native speakers’ knowledge and conversational needs, which may not be undertaken to such an extent with native speakers. this then predicts that a manipulation of cognitive load would show a greater reduction of alignment under load with non-native than with native speakers. this prediction, however, remains to be tested. in sum, we compared lexical-referential alignment with non-native speakers to that with native speakers, to test three hypotheses about the role of audience design in alignment to non-native speakers. our results from two different language populations suggest that, within the context 166 lexical alignment to non-native speakers of our experiments, audience design led to greater alignment with non-native speakers, but only when communicative success was uncertain (experiment 1). grammatical errors produced by the non-native confederate did not increase alignment even further, despite evidencing lower communicative competence in the target language (experiment 2). this result is instead consistent with an alternative collaborative strategy, the desire to demonstrate correct language use (specifically to non-native speakers). an implication of these results – although one that remains to be tested in more naturalistic settings – is that non-native speakers’ production in conversation is ultimately facilitated by their native conversation partners, who act to ensure communicative success. acknolwedgements albert costa died on 10th december, 2018. we dedicate this article to his memory. this research was supported by spanish government funds (grants sej05 62542cv00568007 and psi 200801191/psic, awarded to albert costa, fpu fellowship ap2005-4496, awarded to iva ivanova) and an esrc grant res-062-23-0376, awarded to holly branigan and martin pickering. heartfelt thanks go to katy bellamy, anna leonard cook, sarah ‘sez’ gordon, wan-yu hung, george kountouriotis, nien chen lee, oliver stewart and anna vasilyeva for acting as confederates in experiment 1, to yolanda garcia, jennifer klimowicz, sara rodriguez and jasmin sadat for acting as confederates in experiment 2, and to kyle wolff for formatting the final version of the manuscript. we also thank the editor and two anonymous reviewers for the many useful comments. the authors declare no conflict of interest. 167 ivanova, branigan, mclean, costa and pickering appendix a: analysis of responses to preferred confederate names in experiment 1 analyses of responses to preferred confederate names were the same as the analyses of responses to dispreferred confederate names reported in the main text. the main logistic mixed-effects model had confederate type (native, coded as 0.5, and non-native, coded as -0.5), feedback group (feedback, coded as 0.5, and no feedback, coded as -0.5), confederate order group (native-first, coded as 0.5, and non-native first, coded as -0.5) and their interactions as fixed predictors. dispreferred responses were coded as 1, and preferred responses as 0. further separate models on the data of each feedback group had confederate type, confederate order and their interaction as fixed predictors. the results of the statistical models used to analyze these effects are reported in table 3. participants aligned to a similar extent with the native and non-native confederate (in the main model, the effect of confederate type was not significant). further, participants aligned more to the confederate with whom they interacted first, regardless of confederate type (there was a significant interaction between confederate type and confederate order group). specifically, the native-first group aligned more to the native than to the non-native confederate (a difference of 4.3%), while the non-nativefirst group aligned less to the native than to the non-native confederate (a difference of 3.2%). these differences were further modulated by the presence of feedback (there was a significant three-way interaction between confederate type, feedback group and confederate order group). the separate models on the data of each feedback group indicated that the tendency to align more with the first confederate was carried by participants in the feedback group (a significant interaction between confederate type and confederate order group for participants in this group). further separate models on the data of each confederate order group suggested that this tendency was more robust for the native-first group (a difference of 10.1% and a significant effect of confederate order) than for the non-native-first group (a difference of 4.6% but no significant effect of confederate order.) for the no-feedback group, alignment to the two confederates was similar between the native-first group (a difference of 0.9%) and the non-native-first group (a difference of 1.9%; there were no significant effects). taken together, these results suggest that alignment with two consecutive conversation partners is influenced by interaction order, but only when there is no need for concern about communicative success. the tendency to align less with the second conversation partner in the presence of feedback was also attested in the analyses of dispreferred responses reported above. this tendency was possibly driven by increasing exposure to both alternatives throughout the experiment causing more varied responses (i.e., more noise in the system), as well as by participants’ reluctance to switch to a different name once they had produced a name for a given object – even when it was the preferred name. but this tendency was unaffected by confederate native language only when participants received feedback of comprehension success. in line with the findings reported above, when comprehension success was unclear, this tendency disappeared, and alignment to the non-native confederates when they were the second conversation partners was 7.5% more than in the same situation in the feedback condition (although note that this is a between-participant comparison). considerations about comprehension success in the absence of feedback may have made alignment to preferred names more likely in general, even with native confederates. 168 lexical alignment to non-native speakers model predictors estimate se z p main model confederate type .18 .40 .45 .66 feedback type group -.03 .40 -.07 .95 confederate order group .07 .40 .17 .86 confederate type * feedback type group 1.12 .81 1.39 .16 confederate type * confederate order group 1.91 .81 2.37 .02 feedback type group * confederate order group .44 .81 .54 .59 confederate type * feedback type group * confederate order group 3.23 1.62 2.00 .046 feedback group confederate type .78 .64 1.22 .22 confederate order group .28 .64 .44 .66 confederate type * confederate order group 3.36 1.29 2.61 .009 native-first group confederate type 2.51 1.06 2.37 .02 non-native-first group confederate type -1.74 1.83 -.95 .34 no-feedback group confederate type -.40 .52 -.77 .44 confederate order group -.16 .52 -.31 .76 confederate type * confederate order group .58 1.20 .48 .63 native-first group confederate type -.22 .90 -.24 .81 non-native-first group confederate type -3.29 5.59 -.59 .56 table 3: lmer analyses of responses to preferred names in experiment 1 appendix b: experimental items in experiment 2 carretera [freeway] – camino [road] ambulancia [ambulance] – coche [car] copa [wine glass] – vaso [glass] hamurguesa [hamburger] – bocadillo [sandwhich] escritorio [desk] – mesa [table] litera [bunk bed] – cama [bed] imperdible [safety pin] – aguja [pin] caramelo [candy] – dulce [sweets] palmera [palm tree] – árbol [tree] rosa [rose] – flor [flower] balcón [balcony] – terraza [terrace] cuadro [painting] – dibujo [drawing] puño [fist] – mano [hand] dentadura [denture] – dientes [teeth] mochila [backpck] – bolsa [bag] muñeca [doll] – juguete [toy] gallina [hen] – pollo [chicken] loro [parrot] – pájaro [bird] bebé [baby] – niño [child] (no-error group only) novia [bride] – mujer [woman] (no-error group only) pendiente [earring] – joya [piece of jewellery] (error group only) nevera [fridge] – frigorı́fico [refridgerator] (error group only) 169 ivanova, branigan, mclean, costa and pickering references dale j. barr, roger levy, christoph scheepers, and harry j. tily. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3): 255–278, 2013. issn 0749596x. doi: 10.1016/j.jml.2012.11.001. štefan beňuš. social aspects of entrainment in spoken interaction. cognitive computation, 6(4): 802–813, 2014. doi: 10.1007/s12559-014-9261-4. kirsten bergmann, holly p. branigan, and stefan kopp. exploring the alignment space lexical and gestural alignment with real and virtual humans. frontiers in ict, 2(7), 2015. issn 2297198x. doi: 10.3389/fict.2015.00007. heather bortfeld and susan e. brennan. use and acquisition of idiomatic expressions in referring by native and non-native speakers. discourse processes, 23(2):119–147, 1997. issn 15326950. doi: 10.1080/01638537709544986. richard y. bourhis, howard giles, jacques-phillipe leyens, and henri tajfel. psycholinguistic distinctiveness: language divergence in belgium. in language and social psychology, pages 158–185. basil blackwell, 1979. holly p. branigan, martin j. pickering, and alexandra a. cleland. syntactic co-ordination in dialogue. cognition, 75(2):b13–b25, 2000. doi: 10.1016/s0010-0277(99)00081-5. holly p. branigan, martin j. pickering, jamie pearson, janet f. mclean, and ash brown. the role of beliefs in lexical alignment: evidence from dialogs with humans and computers. cognition, 121(1):41–57, 2011. issn 00100277. doi: 10.1016/j.cognition.2011.05.011. holly p. branigan, alessia tosi, and karri gillespie-smith. spontaneous lexical alignment in children with an autistic spectrum disorder and their typically developing peers. journal of experimental psychology: learning memory and cognition, 42(11):1821–1831, 2016. issn 02787393. doi: 10.1037/xlm0000272. susan e. brennan and herbert h. clark. conceptual pacts and lexical choice in conversation. journal of experimental psychology: learning, memory, and cognition, 22(6):1482–1493, 1996. issn 1939-1285. doi: 10.1037/0278-7393.22.6.1482. franklin chang, gary s. dell, and kathryn bock. becoming syntactic. psychological review, 113 (2):234–272, 2006. issn 0033295x. doi: 10.1037/0033-295x.113.2.234. tanya l. chartrand and rick van baaren. human mimicry. advances in experimental social psychology, 41(8):219–274, 2009. issn 00652601. herbert h. clark. using language. cambridge university press, 1996. doi: 10.1017/ cbo9780511620539. herbert h. clark and susan e. brennan. grounding in communication. in perspectives on socially shared cognition, pages 127–149. american psychological association, 1991. doi: 10.1037/ 10096-006. 170 lexical alignment to non-native speakers herbert h. clark and edward f. schaefer. collaborating on contributions to conversations. language and cognitive processes, 2(1):19–41, 1987. doi: 10.1080/01690968708406350. herbert h. clark and deanna wilkes-gibbs. referring as a collaborative process. cognition, 22(1): 1–39, 1986. doi: 10.1016/0010-0277(86)90010-7. albert costa, martin j. pickering, and antonella sorace. alignment in second language dialogue. language and cognitive processes, 23(4):528–556, 2008. issn 01690965. doi: 10.1080/01690960801920545. willem doise, anne sinclair, and richard y. bourhis. evaluation of accent convergence and divergence in cooperative and competitive intergroup situations. british journal of social and clinical psychology, 15(3):247–252, 1976. victor s. ferreira, daniel kleinman, tanya kraljic, and yanny siu. do priming effects in dialogue reflect partneror task-based expectations? psychonomic bulletin and review, 19(2):309–316, 2012. issn 10699384. doi: 10.3758/s13423-011-0191-9. alex b. fine and t. florian jaeger. evidence for implicit learning in syntactic comprehension. cognitive science, 37(3):578–591, 2013. issn 03640213. doi: 10.1111/cogs.12022. riccardo fusaroli, bahador bahrami, karsten olsen, andreas roepstorff, geraint rees, chris frith, and kristian tylén. coming to terms. psychological science, 23(8):931–939, 2012. doi: 10.1177/ 0956797612436816. simon garrod and anthony anderson. saying what you mean in dialogue: a study in conceptual and semantic co-ordination. cognition, 27(2):181–218, 1987. issn 00100277. doi: 10.1016/ 0010-0277(87)90018-7. simon garrod, alessia tosi, and martin j. pickering. alignment during interaction. in the oxford handbook of psycholinguistics, pages 575–593. oxford university press, 2018. doi: 10.1093/ oxfordhb/9780198786825.013.24. howard giles and peter f. powesland. speech style and social evaluation. academic press, 1975. herbert p. grice. logic and conversation. in peter cole and jerry l. morgan, editors, speech acts. brill, 1975. doi: 10.1163/9789004368811 003. evelien heyselaar, peter hagoort, and katrien segaert. in dialogue with an avatar, language behavior is identical to dialogue with a human partner. behavior research methods, 49(1):46–60, 2017. issn 15543528. doi: 10.3758/s13428-015-0688-7. zoë l. hopkins and holly p. branigan. children show selectively increased language imitation after experiencing ostracism. developmental psychology, pages 897–911, 2020. issn 00121649. doi: 10.1037/dev0000915. zoë l. hopkins, nicola yuill, and holly p. branigan. inhibitory control and lexical alignment in children with an autism spectrum disorder. journal of child psychology and psychiatry and allied disciplines, 58(10):1155–1165, 2017. issn 14697610. doi: 10.1111/jcpp.12792. 171 ivanova, branigan, mclean, costa and pickering william s. horton and richard j. gerrig. speakers’ experiences and audience design: knowing when and knowing how to adjust utterances to addressees. journal of memory and language, 47 (4):589–606, 2002. issn 0749596x. doi: 10.1016/s0749-596x(02)00019-0. william s. horton and boaz keysar. when do speakers take into account common ground? cognition, 59(1):91–117, 1996. issn 00100277. doi: 10.1016/0010-0277(96)81418-1. ellen a. isaacs and herbert h. clark. references in conversation between experts and novices. journal of experimental psychology: general, 116(1):26–37, 1987. issn 00963445. doi: 10. 1037/0096-3445.116.1.26. iva ivanova and albert costa. does bilingualism hamper lexical access in speech production? acta psychologica, 127(2):277–288, 2008. issn 00016918. doi: 10.1016/j.actpsy.2007.06.003. t. florian jaeger and neal e. snider. alignment as a consequence of expectation adaptation: syntactic priming is affected by the prime’s prediction error given both prior and recent experience. cognition, 127(1):57–83, 2013. issn 00100277. doi: 10.1016/j.cognition.2012.10.013. anna k. kuhlen and susan e. brennan. language in dialogue: when confederates might be hazardous to your data. psychonomic bulletin and review, 20(1):54–72, 2013. issn 10699384. doi: 10.3758/s13423-012-0341-8. salla kurhila. correction in talk between native and non-native speaker. journal of pragmatics, 33 (7):1083–1110, 2001. issn 03782166. doi: 10.1016/s0378-2166(00)00048-5. peter ludlow. living words: meaning underdetermination and the dynamic lexicon. oxford university press, 2014. charles metzing and susan e. brennan. when conceptual pacts are broken: partner-specific effects on the comprehension of referring expressions. journal of memory and language, 49(2):201– 213, 2003. issn 0749-596x. doi: 10.1016/s0749-596x(03)00028-7. gary m. oppenheim, gary s. dell, and myrna f. schwartz. the dark side of incremental learning: a model of cumulative semantic interference during lexical access in speech production. cognition, 114(2):227–252, 2010. doi: 10.1016/j.cognition.2009.09.007. jennifer s. pardo. on phonetic convergence during conversational interaction. the journal of the acoustical society of america, 119(4):2382–2393, 2006. issn 0001-4966. doi: 10.1121/1. 2178720. martin j. pickering and simon garrod. toward a mechanistic psychology of dialogue. behavioral and brain sciences, 27(2):169–190, 2004. issn 0140-525x. doi: 10.1017/s0140525x04000056. irina pivneva, caroline palmer, and debra titone. inhibitory control and l2 proficiency modulate bilingual language production: evidence from spontaneous monologue and dialogue speech. frontiers in psychology, 3(57), 2012. issn 16641078. doi: 10.3389/fpsyg.2012.00057. david reitter, frank keller, and johanna d. moore. computational modelling of structural priming in dialogue. in proceedings of the human language technology conference of the naacl, companion volume: short papers, pages 121–124. association for computational linguistics, 2006. doi: 10.3115/1614049.1614080. 172 lexical alignment to non-native speakers jorrig vogels, emiel krahmer, and alfons maes. how cognitive load influences speakers’ choice of referring expressions. cognitive science, 39(6):1396–1418, 2015. issn 15516709. doi: 10.1111/cogs.12205. kodi weatherholtz, kathryn campbell-kibler, and t. florian jaeger. socially-mediated syntactic alignment. language variation and change, 26(3):387–420, 2014. issn 14698021. doi: 10. 1017/s0954394514000155. deborah weiss and james j. dempsey. performance of bilingual speakers on the english and spanish versions of the hearing in noise test (hint). journal of the american academy of audiology, 19 (1):5–17, 2008. issn 10500545. doi: 10.3766/jaaa.19.1.2. linda r. wheeldon and stephen monsell. the locus of repetition priming of spoken word production. the quarterly journal of experimental psychology section a, 44(4):723–761, 1992. doi: 10.1080/14640749208401307. 173 dialogue & discourse 9(1) (2018) 50–78 doi: 10.5087/dad.2018.102 primary and secondary discourse connectives: definitions and lexicons laurence danlos laurence.danlos@linguist.univ-paris-diderot.fr université paris diderot, laboratoire de linguistique formelle 2 place thomas mann, 75013 paris, france kateřina rysová rysova@ufal.mff.cuni.cz charles university, faculty of mathematics and physics, institute of formal and applied linguistics malostranské náměstı́ 25, 118 00 praha 1, czech republic magdaléna rysová magdalena.rysova@ufal.mff.cuni.cz charles university, faculty of mathematics and physics, institute of formal and applied linguistics malostranské náměstı́ 25, 118 00 praha 1, czech republic manfred stede manfred.stede@uni-potsdam.de universität potsdam, applied computational linguistics discourse research lab karl-liebknecht-str. 24-25, 14476 potsdam, germany editor: demberg vera submitted 06/2017; accepted 04/2018; published online 06/2018 abstract starting from the perspective that discourse structure arises from the presence of coherence relations, we provide a map of linguistic discourse structuring devices (drds), and then focus on those found in written text: connectives. to subdivide this class further, we follow the recent idea of structuring the set of connectives by differentiating between primary and secondary connectives, on the one hand, and free connecting phrases, on the other. considering examples from czech, english, french and german, we develop definitions of these groups, with attention to certain cross-linguistic differences. for primary and secondary connectives, we propose that their behavior can be described to a large extent by declarative lexicons, and we demonstrate a concrete proposal which has been applied to five languages, with others currently being added in ongoing work. the lexical representations can be useful both for humans (theoretical investigations, transfer to other languages) and for machines (automatic discourse parsing and generation). keywords: discourse connective, discourse structure, discourse semantics 1. introduction an important strand of research on discourse coherence operates under the assumption that discourse relations are, in addition to coreference, of central importance to explain local coherence between sentences, or more generally, between discourse segments of various kinds. discourse theories and annotation projects implementing this insight have been proposed, inter alia by mann and thompson (1988); sanders et al. (1992); asher and lascarides (2003); prasad et al. (2008). while the precise sets of relations proposed (and the underlying motivations) vary, there is general agreement on coarse classifications into groups such as causal, contrastive, and additive. following the notion of discourse relations further, it is evident that their instances in actual text or speech can c©2018 laurence danlos, kateřina rysová, magdaléna rysová, and manfred stede this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). primary and secondary discourse connectives be either signalled or unsignalled (implicit). in the former case, some linguistic device indicates to the hearer or reader the presence of the discourse relation — sometimes rather specifically (e.g., although), sometimes more vague (e.g., and). these devices are available in all the world’s languages, but they can be of quite different kinds, and so we start here by using the very general term “discourse relational devices” (drds) for them. this follows the practice of the european textlink cost action1, which aims at inventarization, annotation, and cross-linguistic comparison of drds. the research reported in this paper is a result of a collaboration within this project. the set of drds can first be divided in two subsets: discourse connectives, on the one hand, and discourse markers, or particles, on the other. discourse connectives establish a two-place relation between the ‘abstract objects’ (asher, 1993) that result from interpreting the text spans related by the connective, called arguments of connectives. on the other hand, discourse markers/particles establish a one-place relation; as stated in (fischer, 2000, p. 1), they are “small items such as german ja, also, ne, oh or ach and english yes, yeah, oh or well which predominantly occur in spontaneous spoken language.” we acknowledge that the distinction is somewhat blurred by the fact that connectives trivially are also used in spoken language, and more interestingly some connectives in some languages have additional senses such that they can also function similar to discourse markers. still, we see the difference between two-place and one-place relations as an important categorizing factor. in this paper, we look only at discourse connectives, in written texts. we seek to provide further meaningful subclassifications for them. connectives do not form a syntactically homogeneous category, and this is one reason why existing grammars and dictionaries fail to adequately capture the meaning, usage conditions, and translation options of discourse connectives. but this knowledge is of great importance, inter alia for supporting foreign language learning, and for building computer software that performs applications involving text understanding. consequently, we see a pressing need for defining inventories of discourse connectives in the form of lexicons that are readable both for humans and for machines. our goal here is to propose guidelines for developing such lexicons, which can be applied to different languages, resulting in resources that can be interlinked across languages. as part of a definition of discourse connectives, we propose to make a distinction between primary and secondary discourse connectives, illustrated below in (1a) and (1b), respectively.2 in (1a), the causal relation between fred’s jokes and his friends’ hilarity is explicitly signalled by the primary connective as a result, which is a frozen multi-word unit, while, in (1b), it is signalled by the expression this caused, which is made of a subject and a verb, both used with a compositional meaning. the paraphrase relationship between the examples demonstrates that this caused, which we will call a secondary connective, can be used to signal a causal relation between two events in the same conditions as the phrase as a result (as long as syntactic constraints are respected). (1) a. fred didn’t stop joking. as a result, his friends enjoyed hilarity throughout the evening. b. fred didn’t stop joking. this caused hilarity among his friends for the whole evening. (2) fred didn’t stop joking. his friends enjoyed hilarity throughout the evening. 1. http://textlink.ii.metu.edu.tr 2. in this paper, most of the time, we deliberately use invented examples for purposes of succinctly illustrating particular phenomena or distinctions. 51 danlos, rysová, rysová and stede another distinction, already mentioned above, is usually made between explicit and implicit discourse relations, depending on whether they are lexicalized by a connective or not. to illustrate, the causal relation between fred’s jokes and his friends’ hilarity is not overtly marked in (2) and is said to be implicit; it is said to be explicit in (1a), thanks to the primary connective as a result. this difference is a particularly acute phenomenon in shallow discourse parsing, which aims at automatically identifying discourse relations and their arguments in text. for current parsers, determining the type of relation is much more difficult for implicit than for explicit relations (e.g., in the system of oepen et al. (2016), the difference in f1-measure is 13 points; see also (braud and denis, 2016)). as far as examples like (1b) are concerned, standard discourse parsers would not recognize the expression we have labeled as a secondary connective here, and would thus posit an implicit relation between the two sentences.3 depending on the models used, caused may happen to have been learned as a cue for the statistical classifier, but in general the causal relation between fred’s jokes and his friends’ hilarity will be hard to detect. on the other hand, if the expression this caused were identified by a declarative resource as a secondary discourse connective with a causal meaning, the relation could be identified as easily as in (1a). therefore, we believe that a systematic treatment of secondary connectives — and having them play a similar role to that of primary connectives — can be highly beneficial for such parsers, because delegating cases like (1b) to the “explicit” module would reduce the number of instances to be handled by the much more ambitious “implicit relation” module. in a nutshell, the two central contributions of this paper are • to revisit and reconsider the earlier work that previous authors have done on lexicons of discourse connectives, adding insights gained from a multilingual perspective that was originally not present, and thus providing a more comprehensive landscape of drds/connectives; • to formulate guidelines for building such lexicons in other languages, drawing from our experiences in both monolingual description and interlingual linking. the paper is organized as follows. section 2 provides an overview of the types of discourse connectives and some related drds. we characterize the difference between primary and secondary connectives, which was originally introduced by rysová and rysová (2014). these two categories are then defined more precisely in sections 3 and 4, which in turn suggest further subgroups and discuss various details. our descriptions are based on czech, english, french and german, but we give examples in english whenever possible illustrating phenomena that are common to these four languages (and possibly other languages). furthermore, both of these sections present our suggestions for developing new connective lexicons: how to find the set of connectives, what information to encode about them, and in which format. for primary connectives, they build on our experiences with the two lexicons ‘dimlex,’ for german, (stede, 2002; scheffler and stede, 2016) and ‘lexconn’, for french (roze et al., 2012; danlos et al., 2015). for secondary connectives, our exposition is based on the work described in rysová (2015) for czech, which is applied in the ‘czedlex‘ lexicon (mı́rovský et al., 2017). in section 5, we add observations on dealing with the semantics of the two categories of connectives and point to the need for further work in this regard. finally, section 6 draws general conclusions. 3. essentially all shallow discourse parsers (see the shared tasks at conll 2015 and 2016) follow the pipeline model implemented in lin et al. (2014), which first identifies connectives, their arguments, and the relation, and in a later stage tries to classify implicit relations using a separate module. 52 primary and secondary discourse connectives 2. connectives and related drds this section describes our terminology and notation, surveys the realm of connectives and other, similar drds, and argues that two groups (primary and secondary connectives) should be subject to a lexical description. 2.1 connectives: primary, secondary, and free phrases due to its syntactic heterogeneity, the notion of discourse connective needs to be defined at the functional level: a discourse connective is an expression whose function is to semantically and/or rhetorically link two “abstract objects” in the terminology of asher (1993) — roughly, events, states or propositions. this functional definition means that a discourse connective is a discourse-level predicate with two arguments. we follow the convention established by the pdtb (penn discourse tree bank, an english corpus annotated at the discourse level (pdtb group, 2008)), which uses the term arg2 for the argument that is linked to the syntactic host clause of the connective, and arg1 for the “other” argument. to make the arguments easy to identify, the span of text corresponding to an arg2 is henceforth set in italics, while arg1 is in boldface; and we use colors for the discourse connectives — see (3). (3) a. fred didn’t go to work because he is sick. b. fred is sick. however, he did go to work. on the semantic side, the pdtb distinguishes around thirty different sense tags of connectives (organized in a hierarchy), which characterize the discourse relation between the arguments of connectives. for example, the sense tag of because in (3a) is reason, which applies when the connective indicates that the situation specified in arg2 is interpreted as the cause of the situation specified in arg1. senses of connectives will be discussed again in section 5. turning again to syntax, there are several types of drds that explicitly express discourse relations. to illustrate, in (4a-e), there is a common relation of result expressed by different linguistic means, ranging from a grammaticalized one-word expression (therefore) to an open combination of words (due to this weather). these expressions differ in many ways, and we divide them into three groups following the concept of rysová and rysová (2014): primary connectives (4a), henceforth in magenta; secondary connectives (4b-d), in blue; and free connecting phrases (4e), underlined. (4) there was nice spring weather with high temperatures already in april. a. therefore, domestic productions were arriving in the market much earlier than normally. b. because of this, domestic productions were arriving in the market much earlier than normally. c. for this reason, domestic productions were arriving in the market much earlier than normally. d. this caused domestic productions to arrive in the market much earlier than normally. e. due to this weather, domestic productions were arriving in the market much earlier than normally. 53 danlos, rysová, rysová and stede primary connectives are single-word units or non-compositional multi-word units. secondary connectives are multi-word units that differ from primary connectives in that they are not (yet) fully grammaticalized. for example, a secondary connective can be composed of a preposition, e.g., because of, followed by a demonstrative pronoun with a compositional anaphoric reading, e.g., this, (4b). another example of secondary connectives is illustrated in (4c) with the expression for this reason, which is compositional and allows internal modification and inflection (for this unbelievable reason, for these reasons, for that reason). in (4d), it is the verb cause, with an anaphoric subject referring to arg1, which expresses the causal relation. secondary connectives are distinct from free connecting phrases, such as due to this weather in (4e), in that the former can be used in a wide variety of contexts while the latter can be used only in quite specific contexts. as an illustration, the expression because of this illness in (5a) is a free connecting phrase, whose use makes sense because the left context mentions fred’s pneumonia, which is the antecedent of the anaphoric noun phrase this illness. on the other hand, this expression cannot be used in (5b) — which is incoherent (hence the sign #) — because there is no antecedent for the anaphoric noun phrase. it contrasts with the expression because of this, which can be used in either context, (5c). (5) a. fred has pneumonia. because of this illness, he will be absent from his work for two weeks. b. #fred is on his honeymoon. because of this illness, he will be absent from his work for two weeks. c. fred /has pneumonia/ is on his honeymoon. because of this, he will be absent from his work for two weeks. 2.2 the notion of ‘alternative lexicalization’ the penn discourse treebank (prasad et al., 2008) employs the notion of alternative lexicalization (altlex) for expressions that are not considered connectives, yet explicitly signal the presence of a relation. in our terminology, these correspond to the two groups of secondary connectives and free connecting phrases. in the pdtb-2 corpus, 624 tokens are annotated as altlex.4 according to prasad et al. (2010), expressions are annotated as altlex when “a discourse relation is inferred, but insertion of an implicit connective leads to redundancy” (prasad et al., 2010). we feel that redundancy is not a very clear criterion for identifying altlex, because an arg2 can easily be introduced by both a primary and a secondary connective that signal the very same discourse relation, as in the real-life example (6) below, in which result is signalled by both as a result and this caused.5 when following the pdtb annotation guidelines, this caused would be annotated as an altlex for instance in (1b) but not in (6), because the explicit connective as a result is present. examples such as (6) show that (i) redundancy is observed in real texts, and (ii) the definition/annotation of altlex should not rely on redundancy but on semantics: an altlex (a secondary connective or free connecting phrase in our terminology) is an expression that has the same function as a primary connective but has different distributional properties. 4. some primary connectives are also annotated as altlex. this happens because the list of primary connectives in the pdtb-2 is made up of only 100 elements, and so several connectives are missing, like the adverb thereafter, as well as any preposition with a discourse use, like in order to. 5. source: business ethics and diversity in the modern workplace philippe w., zgheib publisher, 2014 54 primary and secondary discourse connectives (6) families, especially in lebanon, have passed through different decades of wars (...). as a result, this caused families to send their children to work. 2.3 relative pronouns primary/secondary connectives and free connecting phrases are not the only types of drds that can connect two spans of text; relative pronouns can also serve this purpose. whereas restrictive relative clauses merely work as modifiers of nps (serving to identify the referent from a set of candidates), non-restrictive relative clauses can have discourse-structural roles. for example, the relative pronoun in (7a) does not modify its antecedent but is used to connect the eventualities of ted’s giving and mary’s eating. see also (7b) from (huddleston and pullum, 2002, page 1223) who speak of the “continuative use of supplementary relatives”. along the same lines, which in (7c) can be considered as connecting two eventualities. (7) a. ted gave a candy to mary, who ate it quickly. b. i gave it to john, who passed it on to mary, and she gave it back to me. c. ted gave a candy to mary, which pleased her. it is not appropriate to include relative pronouns in discourse connective lexicons: they are a separate class. this, of course, does not preclude useful work on relative clauses. for example, further work is needed in order to clarify the linguistic contexts where relative pronouns function like drds, and how annotate them as drds in corpora. indeed, some sentences including a relative pronoun in one language may translate as two sentences linked by a connective in another language. for instance, see (8) from the english-french hansard corpus. (8) a. you could call it “canada’s moment”, a chapter in history that finds the canadian economy outperforming expectations while major foreign economies are still in distress. b. on peut dire que le canada est en train de vivre son heure de gloire. en effet, l’économie canadienne dépasse les attentes alors que les principales économies étrangères sont encore en pleine crise. 2.4 lexical description in conclusion, among the set of connective-like drds, we first have identified the groups of primary and secondary discourse connectives. these will be discussed in detail in the next two sections. we argue that each group, despite the prima facie heterogeneity of the items included, can be described by a lexical resource that gathers the information on the commonalities and differences between these items. such lexicons can support empirical analyses and theoretical work, and they can assist automatic discourse parsers in identifying connectives and their arguments in text. for primary connectives, a lexical approach is feasible because they are closed-class items with, as we will see, a relatively small membership. secondary connectives will be described by templates that are, in a sense, open, yet they are headed by a set of items that again we take to be closed. free connecting phrases, on the other hand, are not amenable to a lexical description because of their productive nature. “lexical entries” like because of this (bad) illness or due to this weather would be undesirable because of their heavy dependence on context. for these reasons, we do 55 danlos, rysová, rysová and stede not believe that free connecting phrases are lexical items like primary and secondary connectives are, and they will therefore be excluded from further consideration in this paper. likewise, we have briefly mentioned relative pronouns above, which share some properties with connectives, but should be treated separately. 3. primary discourse connectives the previous section leads us to define a primary discourse connective as an element which is, morpho-syntactically, a frozen single-word unit or a non-compositional multiword unit, and which is, semantically, a predicate with two arguments referring to eventualities. this characterization of primary discourse connective corresponds to the traditional notion of “connective” as used in the literature; see, e.g., (zwicky, 1985; hrbáček, 1994; pasch et al., 2003; fischer, 2006) or (urgellescoll, 2010). this description pertains to a variety of expressions across languages, which are of different syntactic categories, as we have illustrated in the previous section. to organize the following indepth discussion, the syntactic categories will be described following the distinction between “intrasentential” and “inter-sentential” connectives in sections 3.1 and 3.2 below. in general, intra-sentential connectives typically join both arguments within one typographic sentence (9a), while inter-sentential connectives typically appear in (the second sentence of) a sequence of two typographic sentences, as in (9b) or (9c). (9) a. fred is nice but he may be tough with women. b. fred is nice. therefore, he is never tough with women. c. fred is nice. he is therefore never tough with women. however, it is well-known that typographic conventions are only preference rules, which means that (10) can be observed in parallel with (9). (10) a. fred is nice. but he may be tough with women. b. fred is nice, he is therefore never tough with women. therefore, we propose a linguistic criterion to help distinguish intraand inter-sentential connectives: • intra-sentential connectives form discourse segments that can be embedded under a matrix clause, (11a), • inter-sentential connectives form discourse segments that cannot be embedded under a matrix clause, (11b). (11) a. jane said that [fred is nice but he may be tough with women.]. b. *jane said that [fred is nice, he is therefore never tough with women].6 6. this example becomes acceptable only if and an inter-sentential connective — is added: jane said that [fred is nice and (he) is therefore never tough with women]. 56 primary and secondary discourse connectives intra-sentential connectives (e.g., subordinating and coordinating conjunctions) are thoroughly described and analyzed in syntax/semantics studies that are concerned with the sentential level. our discourse perspective should be compatible with those studies. the english syntactic terminology we use here largely follows (huddleston and pullum, 2002). notice that, with respect to these categories, some connectives have more than one use: for example, before can be used as an intra-sentential connective (as a subordinating conjunction in (12a) and as a preposition in (12b)), but also as an inter-sentential connective (as an adverb in (12c)). (12) a. fred cooked a pizza before mary finished her homework. b. fred cooked a pizza before finishing his homework. c. mary finished her homework. before, fred had cooked a pizza. 3.1 intra-sentential primary connectives intra-sentential connectives are described according to their syntactic category, which puts constraints on the nature of their arguments arg1 and arg2, and on their linear order. recall that arg2 is, roughly, the content of the host clause of the connective. 3.1.1 subordinating conjunctions subordinating conjunctions can be single-word units (if, because) or multi-word units (so that, even if). they introduce a finite clause whose mood (indicative or subjunctive in french) or verbal position (v-2 or v-f in german) depends on the conjunction. they are obligatorily positioned in front of the finite clause they introduce, whose content is their arg2, and they are placed either after their arg1 (13a), before it (13b), or within it (13c). (13) a. john is in a bad mood because he lost his keys. b. because he lost his keys, john is in a bad mood. c. john, because he lost his keys, is in a bad mood. some subordinating conjunctions (but not all) can be externally modified by adverbials such as only or mostly (focus particles), as in (14). external modification should not be confused with internal modification , which is observed only with multi-word expressions.7 for example, some multi-word expressions function as subordinating conjunctions in that they introduce a finite clause, e.g., in the hope that (15), but are not yet fully grammaticalized, so that internal modifiers can be added (in the vain hope that). they should be considered as “secondary subordinating conjunctions” (section 4.1.1). (14) a. only because he wanted to see the moma, fred went to new york. 7. in (prasad et al., 2010), no distinction is made between external and internal modification. therefore, it is claimed that drds “should be treated as an open class that includes explicit connectives”. we believe, on the contrary, that primary connectives form a closed class (section 3.4). as a matter of comparison, english modal verbs (may, must, can, . . . ) can be externally modified (may possibly, may not, . . . ), nevertheless, they do not form an open class. 57 danlos, rysová, rysová and stede b. #only as he wanted to see the moma, fred went to new york. (15) people were trained in the hope that they would find jobs. note that in english when and if and if and when could be considered multi-word primary conjunctions (pdtb group, 2008). along the same lines, in german, it is possible for a conjunction and an adverbial to form a lexicalised (but potentially discontinuous) complex expression that together signal a coherence relation. a case in point is wenn .. auch, where wenn (‘if’/‘when’) and auch (‘also’) each can signal a different relation by themselves (condition and elaboration, respectively), but in combination mark a concession relation. the words can occur adjacent to one another, usually as an afterthought (16a), or discontinuously (16b). (16) a. dieser mac ist ziemlich gut, wenn auch teuer. ‘this mac is pretty good, even though expensive.’ b. ich nehme diesen mac, wenn er auch ziemlich teuer ist. ‘i will take this mac, even though it’s quite expensive.’ 3.1.2 adpositions adpositions are prepositions or postpositions which introduce an infinitival or gerund-participial vp (17a and 17b), or an np referring to an eventuality (17c). in german, adpositions introduce only nps, except for um zu (in order to). (17) a. fred made a pizza in order to please mary. b. fred left after taking a shower. c. fred left after his shower. in english and french, there are only prepositions (prep), whose positional properties are identical to those of subordinating conjunctions (18a-c). some prepositions are lexically related to conjunctions: after + clause and after + ving, in order that + clause and in order to + vinf. some prepositions can be externally modified, such as mainly in order to in (18d). (18) a. fred made a pizza in order to please mary. b. in order to please mary, fred made a pizza. c. fred, in order to please mary, made a pizza. d. schröder visited china, mainly in order to secure trade deals for german industry. german also has postpositions, which appear after the element they introduce (arg2). additionally, the item wegen can be used as either a preposition or a postposition (19a and 19b). (19) a. wegen marys ankunft gingen wir nicht in ein gutes restaurant. ‘because of mary’s arrival, we didn’t go to a good restaurant.’ 58 primary and secondary discourse connectives b. marys ankunft wegen gingen wir nicht in ein gutes restaurant. ‘because of mary’s arrival, we didn’t go to a good restaurant.’ “circumpositions” in german are multi-word connectives made up of a preposition and a postposition that surround arg2, an np referring to an eventuality, (20). (20) um der verlängerung des friedens willen sage ich jetzt nichts. ‘for the continuation of the peace, i am not saying anything.’ prepositions are not annotated at all in the pdtb-2; some are being annotated in the pdtb-3 version, and in corresponding projects on other languages. it should be noted that it can be hard to distinguish between a discourse use and a non-discourse use of, e.g., an expression of the form prep + vinf. this is the case for pour in french (colinet et al., 2014) or to in english, which introduces a “catenative” complement or a purpose adjunct: huddleston and pullum (2002, page 1223) point out that (21a) is ambiguous between a catenative interpretation glossed in (21b) and a purpose adjunct interpretation glossed in (21c). (21) a. he swore to impress his mates. b. = he swore that he would impress his mates. c. = he swore in order to impress his mates. along the same lines, the annotation of expressions of the form prep + np requires determining whether the np refers to an eventuality, which can be a difficult problem. this may well be, but prepositions and postpositions should nevertheless be included in discourse connective lexicons, even if they are not annotated in all corpora. 3.1.3 coordinating conjunctions each of the three languages we study here has a closed list of less than ten coordinating conjunctions (e.g., in french, mais, ou, et, donc, or, car, ni), although the status of some elements is not clear: for example, puis in french shares properties of both coordinating conjunctions and adverbs, and grammars differ on assigning its category. coordinating conjunctions basically introduce a finite clause (22a). however, the finite clause can be elliptical and reduced, for instance, to a vp, or to two nps, when gapping (22b). they are positioned in front of the phrase they introduce, whose content is their arg2, and they are placed to the right of their arg1. they can never be externally modified. (22) a. john will help mary and ted will help sue. b. john will help mary and ted sue. 3.2 inter-sentential primary connectives inter-sentential connectives are single-word adverbs (next, finally, conversely) or adverbial prepositional phrases (pps) (in summary, in conclusion, for example). their host sentence is generally a finite clause. they appear at the beginning of or within arg2, as in (23), and the order is (obligatorily) arg1-arg2 hence, inter-sentential connectives never appear at the very beginning of a discourse). 59 danlos, rysová, rysová and stede (23) a. fred lost his keys. therefore, he is in a bad mood. b. fred lost his keys. he is therefore in a bad mood. as we pointed out earlier, primary connectives of the form pp are frozen and non-compositional expressions, in contrast to secondary connectives. as an illustration of this difference, consider french à part ça: in (24a), the pronoun ça does not have an anaphoric reading, and the connective is a primary one; on the other hand, in (24b), ça has an anaphoric reading and the connective is secondary.8 (24) a. fred vient de s’acheter un costume de marque. a part ça, il se plaint qu’il est fauché. ‘fred just bought a branded suit. yet, he complains he is broke.’ b. fred vient de rater un examen. a part ça, il est en pleine forme. ‘fred just failed an exam. apart from that, he is in a good mood.’ a similar contrast is observed with the expression à ce moment là with both a primary and secondary connective use, see (25) from (roze et al., 2012). (25) a. tu as l’air de penser qu’elle n’est pas honnête. a ce moment là, ne lui raconte rien. ’you seem to think she’s not honest. so don’t tell her anything.’ b. il a commencé à pleuvoir. a ce moment là, fred est arrivé. ‘it started raining. at that moment, fred arrived.’ 3.3 pairs of primary connectives in section 2.2, we pointed out that ‘redundancy’ is not a very good criterion for an insertion test for the presence of an “alternative lexicalization” in a sentence. in the following, we elaborate on this point by discussing instances where two primary connectives form a pair that jointly contributes to discourse coherence, and we distinguish the two cases of the connectives being present in the same argument, or are split between arg1 and arg2.9 we want to stress that pairs of connectives should not be confused with complex multi-word primary connectives such as when and if in english or wenn .. auch in german, which were described in section 3.1. 3.3.1 both parts of the pair hosted by the same argument “doublets of connectives” are made up of either two adverbial connectives or a subordinating conjunction and an adverbial connective. the two connectives in a doublet share the same two arguments and express roughly the same sense, as in (26). in a doublet, the first adverbial or the subordinating conjunction has a somewhat more general meaning than the second adverbial: the second adverbial can signal an additional facet of meaning, without fundamentally altering the coherence relation. in (26a), then signals temporal succession, subsequently adds the information that the temporal intervals of the two events meet; in (26b), but signals an adversative relation, however 8. the (non)-anaphoric reading of ça is determined by testing the replacement of ça with a noun phrase such as le fait que arg1 (‘the fact that arg1’): this test leads to an incoherent discourse in (24a) and a coherent (though infelicitous) discourse in (24b). 9. for a more extensive treatment of this issue of connective pairs in german, see stede and irsig (2012). 60 primary and secondary discourse connectives more specifically a concessive one; in the german example (26c), weil (‘because’) signals a causal relation, nämlich (‘in fact’) does the same and adds a facet of speaker involvement (grabski, 2008). in (26d), both sondern and vielmehr together signal a corrective replacement. (26) a. fred finished his home work. then subsequently he went to the movie. b. fred is a nice guy but he may however be tough with women. c. ich fliege nach berlin, weil dort nämlich ein bruder von mir wohnt. (‘i’m flying to berlin because a brother of mine is living there.’) d. das scheint kein lama zu sein, sondern es sieht vielmehr aus wie ein dromedar. (‘that doesn’t seem to be a llama but it rather looks like a dromedar.’) the crucial point to note is that, in a doublet, both connectives signal by themselves the coherence relation expressed in the examples. therefore, it is important not to confuse these doublets of connectives with cases of “multiple connectives”. multiple connectives are hosted by the same arg2 but have different senses and possibly different arg1s, see (27). in (27a), taken from (forbesriley et al., 2016), the conjunction because has a causal meaning and its arg1 is the canceling event, while the adverbial then has a temporal meaning and its arg1 is the ordering event. in (27b), the conjunction but and the adverbial next share the same arguments; but has an adversative meaning, next is used to make a partition of the thesis between its beginning and its end, as analyzed in (danlos, 2005). (27) a. john ordered three cases of barolo. but he had to cancel the order because then he discovered he was broke.10 b. the thesis begins with a brillant state of the art, but next it describes uninteresting experiments. we want to emphasize that multiple connectives should not be recorded in lexicons: a priori, any two connectives (signalling different relations) can be hosted by the same arg2 as long as syntactic constraints are respected. on the other hand, the above-mentioned doublets of connectives (signalling the same relation) cannot be arbitrarily combined and thus should be recorded in lexicons, as we will discuss in section 3.5. 3.3.2 parts of the pair distributed across arguments “parallel connectives” are “pairs of connectives where one part presupposes the presence of the other, and where both together take the same two arguments,” (pdtb group, 2008). they are illustrated in (28). it has to be noted that one of the parts can be optional, and the other be a ‘simple’ connective in its own right. this holds for the three examples in (28): on the other hand, if and or can be used without their counterparts. the meaning can change, however: either .. or marks a strictly exclusive disjunction, while or does not mark in-/exclusion. (28) a. on the one hand, mr. front says, it would be misguided to sell into “a classic panic.” on the other hand, it’s not necessarily a good time to jump in and buy. 10. as arg1 is not the same for because and then, no boldface is used in this example. 61 danlos, rysová, rysová and stede b. if the answers to these questions are affirmative, then institutional investors are likely to be favorably disposed toward a specific poison pill. c. either sign new long-term commitments to buy future episodes or risk losing “cosb” to a competitor. parallel connectives can be recorded in lexicons, where the conditions on optionality need to be stated in some way (see section 3.5). the notion of arg1/arg2 does not map straightforwardly to parallel connectives and requires a revised definition, but this issue is left aside here.11 3.4 conclusions on primary connectives concerning the possible types of arguments, we have seen that in accordance with the syntactic category of a connective, its arg2 can be a finite clause, a vp or an np referring to an eventuality. hence, the adjp category has been left aside. however, there are some primary connectives that introduce an adjp. for example, although, on top of introducing a finite clause (29a) and a gerund participial vp (29b), can introduce an adjectival phrase (29c), and the same situation holds for concessive connectives in french and german (16a). (29) a. although he is ill, ted went to work. b. although being ill, ted went to work. c. although ill, ted went to work. we have not included verbal suffixes in the list of intra-sentential connectives. indeed, in languages such as turkish and japanese, some verbal suffixes are frequently used to connect two discourse segments and are considered as “subordinators”; see (zeyrek and webber, 2008) for turkish. in english (resp. french) the verbal suffix -ing (resp. -ant) could be considered as a subordinator with a result sense in examples such as fred made a sex joke, shocking mary (‘fred a fait une blague porno, choquant marie’). this is left aside for further research. so, putting aside verbal suffixes, we have described three categories of intra-sentential primary connectives (subordinating and coordinating conjunctions, adpositions) and two categories of intersentential primary connectives (adverbs and adverbial pps). this leads to the question whether these five syntactic categories cover all the primary connectives. we believe this is by and large the case. of course, as said in (prasad et al., 2010), there are some frozen expressions that do not fall in these categories (such as what’s more or never mind that); but in languages such as french or german, for which our earlier work has aimed at developing exhaustive lexicons for primary connectives, we found that only 3.2% and 3.6% of their entries, respectively, do no fall into our five categories. these low proportions show that these expressions can be considered as exceptions to the rule of primary connectives belonging to a closed list of syntactic categories (for the languages we are concerned with).12 11. parallel connectives are quite frequent in chinese, and so zhou and xue (2012) identify the two arguments of parallel connectives on semantic grounds, without using the notation arg1/arg2. 12. this claim should be reinforced by corpus frequencies: it is clear that idiomatic expressions such as what’s more or never mind that appear much less frequently in corpora than conjunctions such as because or pps such as for example. 62 primary and secondary discourse connectives 3.5 building lexicons for primary connectives having discussed the linguistic “map”s of primary connectives, we now turn to the task of actually representing the relevant information in a lexicon that, ideally, is readable for both humans and machines. in this section, we explain the design of the two early lexicons dimlex (stede, 2002; scheffler and stede, 2016) and lexconn (roze et al., 2012; danlos et al., 2015), point to recent extensions, and provide recommendations for starting similar projects in other languages. 3.5.1 acquiring the set of lexical items the first step in building a lexicon is to decide which items to include, along the lines of the definitions we provided in sections 3.1 and 3.2. dimlex in its early stages was developed in cooperation with the research group at institut für deutsche sprache (ids) that was responsible for compiling the ”handbook of german connectives” (pasch et al., 2003; breindl et al., 2015). the first dimlex version was a subset of the words studied by that group, viz. the 175 items that we intuitively considered as most frequent in present-day german. the latest dimlex described in (scheffler and stede, 2016) with 276 entries resulted from another comparison with the final list of the ids publication, but we did not include a number of items that we saw as outdated, or where we did not agree that they should be treated as connectives. further, the “handbook,” for theoretical reasons, excludes prepositions, while dimlex includes many of them. notice that when compiling an inventory of connectives, the notion of “entry” is not trivial to define; this will be discussed in the next subsection. for lexconn, a first version was obtained by compiling a list of elements belonging to the syntactic categories described above, which were manually filtered; the result was a lexicon with 325 entries (roze et al., 2012). an annotation project supplemented this lexicon with an additional 30 entries (danlos et al., 2015). as the annotation enterprise dealt with 18,535 sentences with 10,429 connective tokens annotated, we believe that this latests version is now nearly exhaustive. when the goal is to build a lexicon for a new language, the starting point is the assumption that primary connectives form a closed class; the sizes of dimlex and lexconn can be seen as indicative for the target number. one way of collecting a first base set of connectives is to work through standard grammars of the target language, which give explanations on conjunctions, adpositions and adverbials. grammars that provide semantic classifications (as, for example, many grammars for second language learning do) can be especially helpful, because they also give hints on meaning of connectives. as a potential alternative, we suggest taking advantage of a parallel corpus, in conjunction with an existing lexicon in one of the languages of the parallel corpus. if the corpus covers the target language t and if there exists a connective lexicon in the the source language s, the parallel corpus can be used to automatically retrieve potentially-corresponding words for the connectives known in s.13 the prerequisite is that the text pairs be both sentence-aligned and word-aligned. then the t words that are aligned to the already known connectives in s can be retrieved. this process yields of course only an approximation, as it will contain three types of errors: • wrong word alignment: a connective in s may be falsely aligned to some word in t . 13. a popular resource is the multilingual europarl corpus (http://www.statmt.org/europarl/); it is also included in the larger collection at http://opus.nlpl.eu/. europarl was extensively used for work on connectives in the project comtis (http://www.idiap.ch/project/comtis). 63 danlos, rysová, rysová and stede • ambiguity: the word in s may be used in a non-connective reading, so that the (possibly correctly) aligned word in t should not be part of the connective lexicon in t . • translation as non-connective: a connective in s may be correctly aligned to a word in t that is not a connective in the target language. furthermore, the set of connectives determined for t in this way is obviously not likely to be complete. nonetheless, this procedure can significantly speed up the process of “bootstrapping” a lexicon (and, by running the process also in the opposite direction, also to validate the source lexicon). for the case of mapping german to italian connectives (and backwards), the process is explained by bourgonje et al. (2017). since the meaning of the corresponding connectives in the two texts can expected to be similar, the senses can also be mapped from the s lexicon to the t one, for a start. finally, given a source lexicon in s, one may of course use manual or automatic translation resources in order to obtain a first version of a lexicon for t . this method was applied as the first step in building an italian version of dimlex (feltracco et al., 2016). 3.5.2 organizing the information in lexical entries for organizing the candidate items in a lexicon, an initial decision is to define the mapping from connective words to the notion of an “entry” in the lexicon: what do we regard as two variants of a single connective (to be treated in the same entry), and what is to be analyzed as two different connectives? a clear case is that of mere spelling variants of the same word (as were, for example, created by the german spelling reform in the 1990s). as long as no syntactic or semantic differences can be observed, these are listed as orthographic variants in a single entry, both in dimlex and in lexconn. sometimes, a word can have multiple syntactic roles, with each of them being a connective. in this case, the lexical entry has to provide more than one field for syntactic information. this happens, for instance, with the german trotzdem, which can be a subordinating conjunction (‘although’) and in some dialects also an adverbial (‘anyway’); in english, though shows a similar ambiguity. this situation needs to be distinguished from a word that also has a non-connective reading. in dimlex, this is described in a separate ‘ambiguity’ part of the entry, which states whether or not such a reading is available, and also gives examples of non-connective uses. notice that some words conflate these ambiguities, such as english before, whose three connective variants have been shown above in example (12), and which, in addition, is not a connective when heading an object-denoting np (e.g., before 1996). next, there is semantic ambiguity in case a connective can signal more than one coherence relation, i.e., have more than one sense. dimlex chose to encode this hierarchically embedded within the syntax field of an entry, which thus can have more than one semantic field, corresponding to each sense. (if the two senses also coincide with some syntactic difference, there have to be different syntax fields, though.) a further complication for the question of “one entry” versus “two entries” arises when the connective can have two syntactic roles, coinciding with an orthographic variant. this is a frequent phenomenon both in german (e.g., dadurch adverbial; dadurch, dass complementizer) and in french (e.g., avant de (’before’) preposition; avant que (’before’) subordinating conjunction). in english, a corresponding question is the relationship between because (subordinating conjunction) and because of (preposition). in these cases, dimlex and lexconn opt for two separate entries. 64 primary and secondary discourse connectives on the syntactic side, dimlex chose not go deeply into detail (much less so than the resource compiled by pasch et al. (2003) does). it provides the basic category of the connective as well as information on the possible orderings of the arguments: arg1 precedes or follows arg2, or arg2 can be embedded in arg1. in lexconn, this information is not recorded explicitly, but (see our discussion in section 3.1) it largely follows from the syntactic category. within the syntactic description, dimlex gives the pdtb-3 sense relations and the frequencies that have been observed in a corpus study, where 25 instances of use (taken from the large www.dwds.de corpus) have been analyzed for each connective (scheffler and stede, 2016). in general, when building a lexicon for a new language, such a step of thoroughly considering corpus instances is to be highly recommended; recall that lexconn, too, has been finalized only after an accompanying corpus annotation project. 3.5.3 technical implementation a fundamental decision made originally for dimlex was to use xml as the technical format, because this allows for straightforward mappings for example, by means of xsl scripts) from the base lexicon to various format variants. these include a reduced version to be used in the semiautomatic annotation tool conano (stede and heintze, 2004), an html version for the human reader using a web browser, and different language-technological applications such as discourse parsers. in addition, defining mappings to common spreadsheet formats is not difficult. as an illustration, figure 1 shows an abbreviated form of the dimlex entry for vielmehr, which is similar to english ‘instead’ or ‘rather’, but is used predominantly in contexts of correction. we use this example to explain the remaining aspects of dimlex entries. the orthography is classified as being a simplex (single word) or multi-word connective (more than one word); for the latter, we additionally state whether the two components are discontinuous or not. each orthographic variant receives its own identifier (an extension of the identifier for the complete connective), which allows for possible cross-referencing from other parts of the entry. following the ambiguity information, which consists of two binary features (i.e., with values 0 or 1), a feature specifies whether the connective can be in the scope of focus particles (cf. the discussion of example (14) in section 3.1, repeated here in (30)). (30) a. only because he wanted to see the moma, fred went to new york. b. #only as he wanted to see the moma, fred went to new york. a set of two features states whether it can occur in so-called ‘correlate’ constructions, which are loose yet not totally-arbitrary collocations between adverbials and conjunctions that we described as “doublets” in section 3.3.1 above. in the example (fig. 1), vielmehr is specified as a possible correlate of the conjunction sondern (a corrective form of ‘but’). this means that when the adverbial vielmehr occurs in a sondern-clause, most likely the two words collectively signal the identical instance of a substitution relation (instead of two different relations). an example with these two connectives was shown in (26d) above. the dimlex xml format has proven to be quite compatible with approaches to lexicons in other languages. we converted the original format of lexconn also to dimlex xml, and new lexicons for italian (feltracco et al., 2016) and portugese (mendes et al., 2018) have recently been constructed following the dimlex format. likewise, we mapped the list of english connectives annotated in the pdtb project to dimlex format and extended it with some 50 entries from other 65 danlos, rysová, rysová and stede vielmehr 1 0 0 0 1 0 sondern konnadv 0 1 0 figure 1: dimlex sample entry: ‘vielmehr’ (abridged) 66 primary and secondary discourse connectives resources (das et al., 2018). all of these resources, in a “minimalized” version covering the basic syntactic information and sense relations, have recently been made available online as an interlinked lexical database.14 4. secondary discourse connectives above, we have defined primary connectives as either single-word units or multi-word units, which are non compositional, internally non-modifiable and non-inflected, while secondary connectives are multi-word units which are compositional, modifiable and inflectable. however, there are some exceptions to these rules. for example, the czech primary connective kdyby (‘if’) may be inflected, see the forms kdybych, kdybys, kdyby, kdybychom, kdybyste, kdyby meaning ‘if i, if you, if he/she/it, if we, if you, if they’. therefore, primary and secondary connectives should not be considered as two strictly separated classes of expressions but rather as a scale of expressions in a different degree of grammaticalization. this means that due to language changes and possible increasing grammaticalization, secondary connectives may become primary in the future.15 we should add that some primary connectives in one language translate as secondary connectives in another language. for example, the german primary connective dagegen does not have (yet) any fully grammaticalized counterpart in czech (the czech equivalent is proti tomu consisting of a preposition proti and an anaphoric pronoun in dative case). due to grammaticalization of primary connectives, it is possible to describe their formal characteristics in a comprehensive way, as was done in section 3. however, secondary connectives, which are not (yet) fully grammaticalized, exhibit a high degree of variation and have thus to some extent an open form. as a consequence, it is only possible to describe their syntactic structures, which leaves room for variation. we regard these structures as “lexically headed,” appealing to the notion of core unit: the core unit of a secondary connective is defined as the lexical unit with the strongest meaning. by this definition, any anaphoric element, whose meaning does not stand by itself, is not a core unit. examples are: in the secondary connective for this reason, the core unit is the noun reason; in this precedes, it is the verb precede; in because of this it is the complex preposition because of. given its strong compositional meaning, the core unit signals which type of discourse relation the secondary connective expresses. for example, a secondary connective whose core unit is reason can express either that arg1 is the reason of arg2 — as in the template given in (31a) — or that arg2 is the reason of arg1 — as in (31b). similarly, when the core unit is the verb precede, either (the event in) arg1 precedes (the event in) arg2, as in (32a), or arg2 precedes arg1, as in (32a).16 (31) a. arg1. for this reason arg2. b. arg1. the reason is arg2. 14. http://www.connective-lex.info 15. this idea is supported by historical origin of present-day primary connectives that arose from similar structures (and parts of speech) like present-day secondary connectives. see english because coming from combination of a preposition bi and noun cause or german dagegen containing the preposition gegen and a referential part da. 16. using the pdtb hierarchy of senses tags, the core unit reason signals a relation of type contingency-cause, either reason or result, the core unit precede a relation of type temporal-synchronous, either precedence or succession. given the strong meaning of the core unit, the annotation of discourse relations expressed by secondary connectives in discourse corpora is slightly easier than for (semantically more general) primary connectives. this is reflected in the inter-annotator agreement of discourse relations in czech where the agreement on relations expressed by secondary connectives was slightly higher than by primary connectives: 0.82 vs. 0.77. 67 danlos, rysová, rysová and stede (32) a. arg1. this precedes arg2. b. arg1. this was preceded by arg2. we describe below the list of the most frequent structures for secondary connectives, underlying for each structure the syntactic category of its core unit. these structures are illustrated with english examples, but there exist equivalents in czech, french or german (and possibly many other languages). as for primary connectives, we distinguish intra versus inter-sentential secondary connectives. 4.1 intra-sentential secondary connectives 4.1.1 secondary subordinating conjunctions and prepositions this class includes pps that introduce a complement in the form of a clause, a vp or np, see (33). they are modifiable (in the vain hope) and they may appear without any complement (in this hope). with a complement (referring to arg2), they can be qualified as secondary subordinating conjunctions or prepositions. (33) a. people were trained in the hope that they would find jobs. (= 15) b. he was studying in the hope of being admitted to an engineering college. c. he applied for a job in a new city in the hope of a positive answer. as far as we know, there are no multi-word units that can be qualified as (secondary) coordinating conjunctions. 4.2 inter-sentential secondary connectives 4.2.1 adverbial prepositional phrases secondary connectives in the form of prepositional phrases (pps) are of two kinds: the first is a combination of a preposition and an anaphoric expression, mostly a demonstrative pronoun, like due to this, because of this, despite this, besides this, thanks to this, in spite of this, as in (34). the core units of these secondary connectives are the prepositions i.e., due to, because of, despite, which signal the semantic types of discourse relations e.g., despite signals a concession relation. (34) i had all the necessary qualifications. despite this, i didn’t get the job. the second type of secondary connectives in the form of pps consists of a preposition and a content noun which is the core unit, like for this reason, under these conditions, for this purpose, (35). (35) we were stuck in a traffic jam. for this reason, we couldn’t attend the event. these two types of secondary pps are schematized as pp/prep and pp/n, respectively: in these schemes, the core unit (prep or n) is indicated on top of the syntactic category pp. these schemes can be used in lexicons for secondary connectives, which will be discussed in section 4.3. as it is the case for adverbial primary pps, arg2 is syntactically expressed by a clause and appears obligatorily after arg1. secondary pps appear most of the time at the beginning of arg2, but it may happen that they appear within arg2. 68 primary and secondary discourse connectives 4.2.2 discourse verbs other secondary connectives contain a semantically strong verb that forms the core unit, e.g., the verb mean (occurring in the secondary connective it means that) or other verbs like cause, precede, follow, prove which are called “discourse verbs” in (danlos, 2006). the subject of a discourse verb in a secondary connective is an anaphoric pronoun referring to arg1. the examples in (36) show that arg2 can be nominal or clausal. (36) a. cl start to rise, reaching the maximum level, twice that of healthy controls, on day +11. this preceded the rise of blood leukocytes above 1.0x10(9)1. b. most researches in this field, on the way to understand challenges and propose solutions, are based on case studies. this causes that findings are confined to particular situations. discourse verbs can be used in the active or passive form (e.g. this was preceded by, this was caused by). they are schematized as dvs. 4.2.3 copula structures many secondary connectives contain a semantically weak verb (mostly be). first, this verb can be built with a subject whose head noun is the core unit like reason, condition, consequence, example, conclusion. the examples in (37) show that arg2 can be nominal or clausal. (37) a. the tourism industry has grown over the years. the reason is the arrival of international flights to the capital. b. the tourism industry has grown over the years. the reason is that international flights arrived at the capital. another structure with the weak verb be is illustrated in (38), in which the core unit reason appears after the copula, the subject being an anaphoric pronoun referring to arg1. (38) international flights arrived at the capital. that is the reason why the tourism industry has grown over the years. it should be stressed that arg2 in (37) expresses the reason of arg1 while arg2 in (38) expresses the result of arg1. these copula structures are schematized as be/subjn and be/attn respectively, which states that the core unit n appears as the subject or attribute of the copula. the last structure with the weak verb be is lexically headed by a subordinating conjunction or a preposition, (39). the subject is an anaphoric pronoun referring to arg1. the scheme is be/conj or be/prep. (39) a. out in space, the sky looks black, instead of blue. this is because there is no atmosphere. b. jane got pregnant. this was before her father’s death. 69 danlos, rysová, rysová and stede 4.2.4 to vinf phrases another frequent type of secondary connective is made up of the preposition to followed by an infinitival phrase, e.g. to conclude, to sum up, to give an example, to make it short, to put it in a nutshell. the core unit is either a plain verb (conclude) or a verbal expression (give an example, put it in a nutshell), (40). arg2 is clausal. (40) denotation and sense can be applied to a lexeme or a larger expression. the denotation and sense of a composite expression is . . . . to put it in a nutshell, denotation, reference and sense are closely related to one another. 4.3 building lexicons for secondary connectives for secondary connectives, right now, only a czech lexicon is under development, following annotation of the pdit (prague discourse treebank 2.0 (rysová et al., 2016)); we are not aware of any other efforts. the format of such lexicons is therefore not yet as stable as it is for lexicons of primary connectives, which have been developed for several languages (see section 3.5). nevertheless, we think that the following sections give guidelines that will help the development of lexicons for secondary connectives in other languages. section 4.3.1 briefly addresses the question of delimiting what can be a lexical entry for such a lexicon. section 4.3.2 discusses how to organize the information under each entry. 4.3.1 organizing the information in lexical entries we have seen that there are possibly several different structures containing the same core unit, e.g., for this reason, the reason is, that is the reason why for the core unit reason. the question is whether some of them are only variants of the same secondary connective (and if so, which one is basic and could be presented as a representative for the others) or whether all of them are individual secondary connectives that should be considered separate lexical entries. from the formal point of view, all these examples are individual (structurally different) expressions, and therefore a possible solution would be to treat them as individual lexical entries. however, in an alphabetical ordering of the lexicon, the connectives for this reason and that is the reason why, for example, would appear very far away from each other, which would be inconvenient for the (human) lexicon users. moreover, this solution would lead to an enormous number of entries and an overall opacity of the lexicon. at the same time, we cannot consider structures like for this reason and that is the reason why simply to be variants of the same connective. in conclusion, the solution may be to keep all the secondary connectives with the same core unit within a single entry (so that the user can easily find all structures containing the same core unit in one place) and at the same time to differentiate between distinct structures by simply listing them under the common core unit without any hierarchy. in other words, we consider core units to be “umbrella lemmas” of the individual structures containing them. 4.3.2 description of secondary connectives with the same core unit the most common schemes for secondary connectives were described in the previous sections. an entry with a given core unit lists the schemes in which it appears. these schemes may need to be specified; for example, in the scheme pp/n, the preposition needs to be specified for each core unit n. in the lexical entry under the core unit reason given in (41), there are three schemes; for each 70 primary and secondary discourse connectives scheme, its name (e.g., pp/n) is given and followed by its specification. the following abbreviations are used in the specifications of the schemes: ana for anaphoric, det for determiner, adj for adjective, pro-subj for a subject pronoun (referring to an eventuality); the symbol $n$ stands for a variable whose value is given in the core unit field; optionality is marked with parenthesis and alternatives are written within brackets. under each scheme, the field realizations gives concrete examples of the scheme. (41) core unit: n = reason scheme 1 = pp/n : for [ana-det (adj)/ana-adj] $n$ realizations: for this reason, for given reason inflection: 1 modification: 1 scheme 2 = be/subjn: det (adj) $n$ be (that) realizations: the reason is, a possible reason is that inflection: 1 modification: 1 scheme 3 = be/attn: pro-subj be the (adj) $n$ why/for which realizations: that is the reason why; this is the simple reason for which inflection: 1 modification: 1 as pointed out earlier, secondary connectives may be inflected (for these reasons). the question is how many variants should be recorded in the lexicon. on the one hand, it is useful to capture all substantial characteristics of connectives; on the other hand, the lexicon should not be overcrowded by details. so we suggest not to include all the variants in the lexicon, but just to include the information as to whether a scheme is inflectable and modifiable, through the binary features inflection and modification respectively (a typical example of a concrete inflected/modified realization can be added). note that the scheme in itself indicates where possible modifications can take place. for example, when the core unit is the noun reason, this noun can be modified by an adjective in the three schemes (41) where it appears (e.g., for this simple reason, the main reason is, that is a possible reason why). moreover, when a scheme includes a verb (weak or plain), the verb itself can be modified (e.g., the reason is possibly that, it will inevitably cause that, to conclude rapidly). but this is a general phenomenon, so it should not be recorded in the lexicon. more generally, the scheme specifications (and the binary features) should lead to regular expressions that define search patterns to find secondary connectives in corpora. the information given in (41) under the core unit reason is complemented by other fields that appear in each scheme: for example, the sense of the secondary connective, the existence of a primary connective equivalent if any, foreign language equivalents, etc. this is illustrated in (42) for the first scheme under the core unit reason. (42) core unit: n = reason scheme 1 = pp/n : for [ana-det (adj)/ana-adj] $n$ realizations: for this reason, for given reason 71 danlos, rysová, rysová and stede inflection: 1 modification: 1 sense : result primary connective equivalent: therefore foreign language equivalents: • czech: z tohoto důvodu • french: pour cette raison • german: aus diesem grund scheme 2 = be/subjn: det (adj) $n$ be (that) . . . 5. semantics for primary and secondary connectives so far, we have looked mainly at structural properties of primary and secondary connectives, while touching on semantics only occasionally. in this section, we consider first the question of the sense inventory for both groups, and then that of internal or external modifiers taking scope over a connective. 5.1 senses of primary and secondary connectives as said in section 1, a discourse theory such as sdrt (asher and lascarides, 2003) and any discourse annotation project — e.g., (prasad et al., 2008) for english or (rysová et al., 2016) for czech — relies on the assumption that discourse relations form a closed class of less than thirty elements (27 leaf nodes in the pdtb-3 hierarchy (webber et al., 2016); 22 in czech) that are considered to be the senses of primary connectives. in comparison, there exist around 300 primary connectives in french and german, and 150 in czech. as a consequence, sense hierarchies do not attempt further semantic differentiation between primary connectives such as therefore, thus, hence, consequently, as a result, as a consequence, as a side-effect, which are said to all express the very same discourse relation (result). notwithstanding the coarse grain in these semantic analyses, there is a problem in that some connectives have an “open sense”, i.e., a sense which does not fall into the closed list of senses. in french, 6.7% of primary connectives have an open sense. as an illustration, consider au fur et à mesure que, which expresses both temporal simultaneity and proportional equality between its arguments, as in (43a) whose accurate english translation is given in (43b) with no connective but a specific construction the more . . . the more. the specific sense of this conjunction does not belong to the closed list, so it is considered in lexconn as an open sense. this choice means that a too coarse-grain semantic analysis is rejected: it is considered as inappropriate to give au fur et à mesure que just a simultaneity sense, which belongs to the closed list. we stress that the number of french primary connectives with an open sense is twice the number of french primary connectives with an “open syntactic category”, i.e., a syntactic category that does not fall into the closed list presented in section 3. (43) a. son élocution devenait de plus en plus inaudible au fur et à mesure qu’il descendait la bouteille de whisky. b. the more he went down the bottle of whisky, the more inaudible his speech was. 72 primary and secondary discourse connectives what is the situation for the senses of secondary connectives? following the position adopted for primary connectives, one can make the assumption that secondary connectives such as the result is, the consequence is, a side-effect is, this results in, this causes that, this leads to, this brings about, for this reason all express the very same result relation. however, the number of secondary connectives with an open sense may be important, especially when considering those connectives whose lexical head is a noun with a compositional meaning. in czech, we estimate that 15% of secondary connectives have an open sense (i.e., a sense which is not one of the 22 senses adopted for primary connectives). as an illustration, consider v tomto ohledu (’in this respect’) which is illustrated in (44). it does not have any equivalent as a primary connective: it cannot be replaced, for example, by tak (’so’) (with a result sense) without changing the meaning of the discourse. (44) rychlost internetu je v této zemi velmi malá. v tomto ohledu je daná země pro uživatele internetu jednou z nejméně výhodných zemı́. ‘the internet speed is low in this country. in this respect, it is one of the worst countries to use the internet in’. one solution consists in adding new senses to the closed list, for example to add the sense ‘regard’, as it has been adopted in the pdit. this sense is expressed in czech, english, french or german not by a primary connective but by various secondary connectives, for example in english by the following pps (schematized as pp/n): in this respect, in this regard, in this perspective, against this backdrop. however, the solution of adding new senses to the closed list is not appropriate for secondary connectives with a very specific sense such as in the hope (that/of), which was illustrated in (33) in section 4.1.1. when rejecting that its sense is goal (in the closed list) because of the non-synonymy between hope and goal, it must be considered an open sense; it would be peculiar to add a sense ‘hope’ just because of this connective. in conclusion, while at the syntactic level, it seems clear that primary and secondary connectives belong to a closed list of syntactic categories/templates with only a few exceptions (section 3.4), the situation is more confusing at the semantic level. even though the list of senses for primary connectives appears to be fairly closed, as demonstrated by corpus annotation projects, this seems not to be the case for secondary connectives. our strategy would be to introduce new senses if a variety of very similar secondary connectives exists in a language (preferably, in many languages), such as ‘regard’, while leaving more idiosyncratic secondary connectives recorded in the lexicon with an open sense. 5.2 modification of primary and secondary connectives another interesting aspect of connective semantics is modification. for primary connectives, only external modification is possible (section 3.1) and modifiers are roughly focus particles (e.g., mainly, only, just) with a modal semantic value. for secondary connectives, internal modification is possible (section 3.2) and internal modifiers can have a modal value (a possible reason is, the main reason is, possibly this has caused) or an evaluative value (the awful consequence is, a positive result is, this was unfortunately followed by). moreover, two modifiers can be found within the same expression, as in the example probably the most egregious example is, as cited in (prasad et al., 2010). discourse theories essentially ignore modification of discourse relations with a modal value, let alone with an evaluative value. discourse annotation faces difficulties with modification of connectives 73 danlos, rysová, rysová and stede and usually resorts to ad-hoc solutions. this means that research is needed on the modification of discourse relations/connectives (at the theoretical and practical levels) with solutions that work for primary connectives as well as for secondary connectives. as an illustration, consider the authentic example (45). in discourse parsing or annotation, it can be computed or annotated that there is a result relation between arg1 (in bold) and arg2 (in italics), due to the secondary connective an undesirable side-effect is that whose core unit is the noun sideeffect used in a be/subjn structure (section 4.2.3) and modified by the adjective undesirable. this analysis avoids placing an implicit connective between the two typographic sentences and can be good enough for many nlp tasks. however, it does not make any distinction between a result and a side-effect, and it does not state that the scope of the evaluative adjective undesirable is arg2, which means that arg2 is undesirable for the writer. this can be achieved in a deeper semantic analysis, which we here suggest as a goal for future research. (45) conventional statistics-based methods for joint chinese word segmentation and partof-speech tagging have generalization ability to recognize new words that do not appear in the training data. an undesirable side-effect is that a number of meaningless words will be incorrectly created. . . . [corr 2013] 6. conclusion and outlook in our proposal on structuring the realm of drds, we have provided clear definitions for primary and secondary connectives, showing that the former are often grammaticalized variants of the latter. the syntactic categories of primary connectives belong to a closed list (conjunctions, adpositions, adverbs, prepositional phrases), except for a very small number of other elements (and putting aside verbal suffixes). exhaustive lexicons can be developed in a format that was first designed for german and more recently implemented for french, italian, english, portugese and other languages, so that we consider it as stable for this type of languages now. we explained the design choices and provided hints for creating a resource along these lines for a new language. one advantage of the xml format is that it allows for easily maintaining variants of a lexicon, which differ in granularity: for german, the smallest version includes only the information that is needed for the semi-automatic annotation tool conano (i.e., syntactic category, example sentences, and sense), a slightly extended version is part of the parallel lexicon set available at www.connective-lex.info, and a yet larger version contains additional information that is still under construction and hence not published yet. straightforward xslt scripts can derive simpler versions from the “master” version. secondary connectives, which are compositional, modifiable and inflectable, appear in syntactic structures that are lexically headed by the core unit, i.e., the unit that has the strongest meaning. we have given a list of the most frequent structures and suggested guidelines to develop lexicons relying on these structures. so far, only a czech lexicon is being developed for secondary connectives after annotation in a corpus, so we are not yet in a position to discuss coverage questions. we hope that the joint work presented here will support the development of lexicons for primary and secondary connectives, as well as discourse annotation projects, in other languages. with a growing set of such resources, information on cross-language linking can be added, so that matters of translation can be studied. also, in our future work, we plan to address various questions on semantics, e.g., to what extent the senses of connectives form a closed list, when multiple languages are being considered. 74 primary and secondary discourse connectives finally, we wish to mention again the potential role of such lexicons for discourse parsers (either pdtb-style ”shallow” or rst/sdrt-style ”deep”), which can use a lexical resource to supplement the corpus-derived models of relation signals in order to improve recall on the task of identifying coherence relations. as long as the size of relation-annotated corpora remains relatively small — which is the case for any language but english today — such a lexicon can be of great help for bootstrapping approaches to parsing. acknowledgments we thank ewan dunbar for editing our english. we also acknowledge support from institut universitaire de france for the french part, the czech science foundation project no. ga17-06123s for the czech part, deutsche forschungsgemeinschaft project ‘anaphoricity in connectives’ and sfb 1287 for the german part, and the isch cost action is1312 ‘textlink’ for the collaborative part. references nicholas asher. reference to abstract objects in discourse. kluwer, dordrecht, 1993. nicholas asher and alex lascarides. logics of conversation. cambridge university press, cambridge, 2003. peter bourgonje, yulia grishina, and manfred stede. toward a bilingual lexical database on connectives: exploiting a german/italian parallel corpus. in proceedings of the fourth italian conference on computational linguistics (clic-it 2017), rome, 2017. chloé braud and pascal denis. learning connective-based word representations for implicit discourse relation identification. in proceedings of the 2016 conference on empirical methods in natural language processing (emnlp 2016), pages 203–213, austin, texas, 2016. eva breindl, anna volodina, and ulrich herrmann waßner. handbuch der deutschen konnektoren 2. walter de gruyter, berlin/new york, 2015. margot colinet, laurence danlos, mathilde dargnat, and grégoire winterstein. emplois de la préposition pour suivie d’une infinitive : description, critères formels et annotation en corpus. in f. neveu, p. blumenthal, l. hriba, a. gerstenberg, j. meinschaefer, and s. prévost, editors, 4ème congrès mondial de linguistique française, pages 3041 – 3058, berlin, germany, 2014. laurence danlos. partition of an entity with aspectuo-temporal operators. in proceedings of the third international workshop on generative approaches to the lexicon (gl’2005), genève, switzerland, 2005. laurence danlos. discourse verbs and discourse periphrastic links. in proceedings of the second workshop on constraints in discourse (cid 2006), maynooth, ireland, 2006. laurence danlos, margot colinet, and jacques steinlin. fdtb1 : repérage des connecteurs de discours dans un corpus français. revue discours, 15, 2015. 75 danlos, rysová, rysová and stede debopam das, tatjana scheffler, peter bourgonje, and manfred stede. constructing a lexicon of english discourse connectives. in proceedings of the 19th annual sigdial meeting on discourse and dialogue (sigdial 2018), melbourne, australia, 2018. anna feltracco, elisabetta jezek, bernardo magnini, and manfred stede. lico: a lexicon of italian connectives. in proceedings of the third italian conference on computational linguistics (clicit 2016), napoli, italy, 2016. kerstin fischer. from cognitive semantics to lexical pragmatics: the functional polysemy of discourse particles. mouton de gruyter, berlin/new york, 2000. kerstin fischer. towards an understanding of the spectrum of approaches to discourse particles: introduction to the volume. approaches to discourse particles, pages 1–20, 2006. kathy forbes-riley, bonnie webber, and aravind joshi. the predicate-argument semantics of discourse connectives in d-ltag. journal of semantics 23(1), 2016. michael grabski. connectives that manage perspectives in discourse: on the function of german ‘nämlich’, ‘schließlich’, and ‘also’. in proceedings of the 3rd workshop on constraints in discourse (cid 2008), potsdam, germany, 2008. josef hrbáček. nárys textové syntaxe spisovné češtiny. trizonia, prague, czechia, 1994. isbn 80-85573-51-2. rodney huddleston and geoffrey pullum. the cambridge grammar of the english language. cambridge university press, cambridge, 2002. ziheng lin, hwee tou ng, and min-yen kan. a pdtb-styled end-to-end discourse parser. natural language engineering, 20:151–184, 2014. william mann and sandra thompson. rhetorical structure theory: towards a functional theory of text organization. text, 8:243–281, 1988. amalia mendes, iria del rio gayo, manfred stede, and felix dombek. a lexicon of discourse markers for portuguese. in proceedings of the tenth international conference on language resources and evaluation (lrec 2018), miyazaki, japan, 2018. jiřı́ mı́rovský, pavlı́na synková, magdaléna rysová, and lucie poláková. czedlex 0.5. charles university, prague, czech republic, 2017. s. oepen, j. read, t. scheffler, u. sidarenka, m. stede, e. velldal, and l. vrelid. opt: oslopotsdamteesside—pipelining rules, rankers, and classifier ensembles for shallow discourse parsing. in proceedings of the conll 2016 shared task, berlin, germany, 2016. renate pasch, ursula brauße, eva breindl, and ulrich herrmann waßner. handbuch der deutschen konnektoren. walter de gruyter, berlin/new york, 2003. pdtb group. the penn discourse treebank 2.0 annotation manual. technical report, institute for research in cognitive science, university of philadelphia, 2008. 76 primary and secondary discourse connectives rashmi prasad, nikhil dinesh, alan leea, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of the sixth international conference on language resources and evaluation (lrec’08), marrakech, morocco, 2008. rashmi prasad, aravind joshi, and bonnie webber. realization of discourse relations by other means: alternative lexicalizations. in proceedings of the 23rd international conference on computational linguistics (coling 2010): poster volume, beijing, china, 2010. charlotte roze, laurence danlos, and philippe muller. lexconn: a french lexicon of discourse connectives. revue discours, 10, 2012. magdaléna rysová. diskurznı́ konektory v češtině (od centra k periferii). [discourse connectives in czech (from centre to periphery).] phd thesis. charles university in prague, prague, czechia, 2015. magdaléna rysová and kateřina rysová. the centre and periphery of discourse connectives. in wirote aroonmanakun, prachya boonkwan, and thepchai supnithi, editors, proceedings of the 28th pacific asia conference on language, information and computing (paclic 2014), pages 452–459, bangkok, thailand, 2014. magdaléna rysová, pavlı́na synková, jiřı́ mı́rovský, eva hajičová, anna nedoluzhko, radek ocelák, jiřı́ pergler, lucie poláková, veronika pavlı́ková, jana zdeňková, and šárka zikánová. prague discourse treebank 2.0. charles university, prague, czech republic, 2016. ted sanders, wilbert spooren, and leo noordman. toward a taxonomy of coherence relations. discourse processes, 15:1–35, 1992. tatjana scheffler and manfred stede. adding semantic relations to a large-coverage connective lexicon of german. in proceedings of the ninth international conference on language resources and evaluation (lrec 2016), portorož, slovenia, 2016. manfred stede. dimlex: a lexical approach to discourse markers. in a. lenci and v. di tomaso, editors, exploring the lexicon theory and computation. edizioni dell’orso, alessandria, 2002. manfred stede and silvan heintze. machine-assisted rhetorical structure annotation. in proceedings of the 20th international conference on computational linguistics (coling 2004), pages 425– 431, geneva, switzerland, 2004. manfred stede and kristin irsig. complex connectives in german: complications for local coherence analysis. in anton benz, manfred stede, and peter kühnlein, editors, constraints in discourse 3 representing and inferring discourse structure, pages 165–182. benjamins, amsterdam, 2012. miriam urgelles-coll. the syntax and semantics of discourse markers. continuum studies in theoretical linguistics. a&c black, london, uk, 2010. bonnie webber, rashmi prasad, alan lee, and aravind joshi. a discourse-annotated corpus of conjoined vps. in proceedings of the tenth linguistic annotation workshop (law-x 2016), berlin, germany, 2016. 77 danlos, rysová, rysová and stede deniz zeyrek and bonnie webber. a discourse resource for turkish: annotating discourse connectives in the metu corpus. in proceedings of the third international joint conference on natural language processing (ijcnlp 2008), hyderabad, india, 2008. yuping zhou and nianwen xue. pdtb-style discourse annotation of chinese text. in proceedings of the 50th annual meeting of the association for computational linguistics (coling 2012), pages 69–77. association for computational linguistics, 2012. arnold m. zwicky. clitics and particles. language, 61(2):283–305, 1985. 78 dialogue & discourse 8 (2) 2017 129-148 doi: 10.5087/dad.2017.206 © 2017 zeynep cihan koca-helvacı this is an open-access article distributed under the terms of a creative commons attribution license (http:==creativecommons:org=licenses=by=3:0=). boon or bane? discursive construction of the shale gas controversy zeynep cihan koca-helvacı zeynep.koca@deu.edu.tr school of foreign languages, dokuzçeşmeler yerleşkesi, dokuz eylul university 35160/ buca izmir/turkiye editor: david traum submitted 11/2016; accepted 06/2017; published online 11/2017 abstract this study explores strategies in pro and anti-shale organizations’ discourse by combining the discourse-historical approach (wodak, 2001) with corpus linguistics. with the help of keyword lists, collocations, concordances, and key semantic domains, the representations of shale gas extraction, relevant actors and argumentation schemes in opposing discourses of the pro-shale marcellus shale coalition and anti-shale americans against fracking were analysed. the findings of the study show that the advocates presented shale gas as a bonus for the crisis-struck american society through discourse of altruism and solidarity. the opponents, on the other hand, represented shale gas as a threat to the american ecosystem and public health through an alarming and scientific discourse. the empirical findings of this study add to a growing body of literature on discursive strategies employed by opposing camps of environmental controversies. keywords: shale gas, environmental issues, the discourse-historical approach, corpus linguistics, discursive strategies 1 introduction depletion of conventional energy resources, growing demand for energy, and advanced drilling technologies have led many energy companies to steer towards unconventional sources of fossil fuel like shale gas. the extraction process involves horizontal drilling and hydraulic fracturing which means drilling and injecting chemical solvent at high pressure into the hole to crack the rock formations so that trapped methane gas in the miniscule pores of shale deposits can be released (finewood & stroup, 2012). the extraction of shale gas is also known as (hydraulic) fracturing, fracking, or fracing. since the mid-2000s, the increasing demand for fracturing has sharply polarized society. the advocates claim that the extraction of domestic shale resources not only guarantees national security by ending reliance on energy imports but also flourishes national economy by creating new jobs (kay, 2011). in addition, it is also claimed that shale gas with lower carbon emissions than coal means less pollution (engelder, 2011). nevertheless, the opponents argue that fracturing causes irreversible damage to ecosystems with the excessive koca-helvaci 130 use of water and toxic substances in the drilling process (balaba & smart, 2012). apart from this energy and water nexus, fracturing is also criticized for aggravating global warming by bringing methane gas to the surface (burleson, 2012) and changing rural landscapes as a result of mushrooming of fracturing wells, an army of trucks, and gas workers (smith & ferguson, 2013). with rich shale resources, united states has witnessed an economic revival in the last decade as shale gas accounted for more than 40% of the total natural gas production in 2014 (us energy information administration, 2015). however, the release of the movies gasland (2010) and promised land (2012) have alarmed the society about the detrimental effects of fracturing on human life and environment. noticing the intricate link between fracturing and contaminated water resources, american stakeholders have raised objections against shale gas extraction in their region (finewood & stroup, 2012; maykuth, 2011). the discursive aspects of this controversy lay the foundation for this study. the main objective of this research is to analyse the discursive strategies used by both sides of the public debate for legitimizing their own perspectives while delegitimizing the standpoint of their opponents. on the pro-hydraulic fracturing side of the debate, there is the marcellus shale coalition (hereafter msc) which represents the energy, chemical and transportation companies that are engaged with shale business in the region. as the shale industry’s top trade group within pennsylvania, the coalition functions as the primary public face of the industry (bomberg, 2013). for the anti-fracturing camp, the non-governmental organization americans against fracking (hereafter aaf) was chosen because it is the largest coalition with members from environmental organizations, health institutions, labour unions and social justice groups. press releases, reports, letters and newsletters published between july 2013 and july 2015 on each website were gathered to build a corpus. as my research aims at identifying the particular discursive strategies which enable the representation of a group and associated action either favourably or unfavourably, i focused on the referential or nomination strategies, predicational strategies and finally argumentation strategies within the framework of the discourse-historical approach (hereafter dha) (wodak, 2001). while the first strategy is concerned with how social actors and actions of inand out-groups are named, the second strategy deals with their evaluative attributions. finally, argumentation strategies are concerned with topoi which are stock topics or argumentative formulations that provide standard patterns of logical associations between a claim and conclusion (burgess et al., 2001). as rhetorical configurations, topoi enable legitimation of in-group action or delegitimation of out-group action. as my dataset was too big for an in-depth manual analysis, i also took advantage of corpus linguistics to obtain keywords, collocations, concordance lines and key semantic categories to examine these discursive strategies. my research questions are: 1. which linguistic resources did the msc and the aaf use to represent social actors of the in-group and out-group? 2. how was shale gas extraction represented linguistically in both corpora? 3. by means of what argumentation strategies did each organization justify its own standing while attacking the opponent’s? this article is divided into six sections. following this introductory section, the literature review presents relevant research concerning the discursive structures in the representation of the shale gas controversy. the methodology section explains the analytical framework by drawing upon the dha and corpus linguistics. in the data section, i provide information about data selection and collection methods. the analysis presents how referential or nomination, predicational and argumentation strategies were used in each corpus. finally, i discuss findings of the research in the conclusion. 2 literature review with its controversial nature, the extraction of unconventional energy resources has become a site of attention for critical discourse analysis which aims at providing a critical perspective on structures of manipulative discourse concerning contentious issues (van dijk, 1993). the discursive construction of the shale gas controversy 131 never-ending strife between environmental protection and economic growth makes the discourse about the energy industry a rich database for linguistic analysis. focusing on the coal seam gas extraction debate in australia, mercer, rijke and dressler (2014) examined the language use concerning the representation of effective regional development and role of economics in decision-making in advocate and opponent groups’ media reports, promotional materials and website discussions over a one year period. it was seen that whilst the anti-coal seam gas group’s discourse was shaped by health concerns and emphasis on irreplaceability, damage and interference, the discourse of the procoal seam gas group’s discourse was based on growth and prosperity which are the basic constituents of the neoliberal mindset. being aware of the power of media in shaping public perception of energy problems, jaspal and nerlich (2014) adopted thematic analysis and social representations theory to analyse pervasive rhetorical strategies in four british broadsheets’ coverage of the fracking controversy at three critical points. the guardian and the independent with their left-ofcentre political standing tended to frame fracking as a threat to the environment and climate targets with reference to the movie gasland (2010), the increasing seismic activity and carbon emissions. conversely, the times and the daily telegraph with their centre right perspective represented fracking as a solution for unemployment, economic slow-down and national security. in another major study on british media’s representation of the issue, cotton et al. (2014) drew upon an argumentative discourse analytic approach to analyse dominant storylines in the british shale debate through stakeholder interviews and key policy actor statements in broadsheets. whilst ‘cleanliness and dirt’, ‘energy transitions’ and ‘geographies of environmental justice’ were found to be the most dominant storylines, the government’s intentional move to emphasize economic benefits while minimising the risk factors was noted. bomberg (2015) also investigated the dominant discursive frames of the pro and antishale groups’ discourse in the uk fracking debate between 2011 and 2014 to find out which frames became more visible throughout time. it was seen that in the pro-shale corpus, the economic promise and security frames remained visible across time whereas the reassurance frame, which was aimed at ensuring the public that shale was harmless, gradually became more dominant. on the contrary, the ‘threat’ and the ‘fossil fuel lock-in’ frames remained consistently strong in the anti-shale corpus whilst the bad-governance frame which underlined lack of transparency, democracy and citizen input in the policy making processes increased dramatically throughout the timeline. another detailed study from the european fracking debate is on poland which is one of the countries with prolific shale plays in the continent. wagner (2015) investigated the role of experts and knowledge in polish media with respect to the models of public communication on shale gas. the analysis was two-staged in which primarily frequency lists were generated to enable the qualitative analysis of the situational maps and discourse worlds. the outcome of the research has shown that shale gas was associated with the discursive field of uncertainty. another study on european shale debate is by schirrmeister (2014) who studied storylines and future imagery in media and stakeholder discourses over hydraulic fracturing debate in germany. the findings overlapped with the studies mentioned in this section as it was seen that there were six main storylines which are opportunity, economic growth, risky technology, conflict, climate change, possible geopolitic effects. finally, considering the shale debate in the usa, a longitudinal study by evensen, clarke and stedman (2014) examined how local newspapers of the marcellus shale region reported on hydraulic fracturing. it was seen that the regional newspaper coverage of the issue is heavily dependent on the local discourse, especially information, ideas and opinions shared at public meetings. that is to say, when an area experienced more benefits, there was a positive coverage. the studies mentioned in this section showed that previous linguistic research on shale gas discourse tends to focus on macro-linguistic structures such as frames and storylines whereas microkoca-helvaci 132 linguistic structures such as lexical choices, agency and argumentative strategies haven’t been studied in much detail. 3 methodology in this section, the theoretical and analytical frameworks of the research were explained. 3.1 theoretical framework now that my focus is on discursive representations of pro and anti-fracturing camps over shale controversy in the usa, i made use of the dha to analyse discursive strategies both blocs used to frame actors, action and argumentation to present a positive in-group and negative out-group image. legitimation, which is central to both organizations’ representation efforts, can be defined as a strategy that justifies a social activity through ‘‘good reasons, grounds, or acceptable motivations for past or present action’’ (van dijk, 1998, p. 255). the most pervasive and remarkable pattern in this persuasive process is unsurprisingly the development of an us versus them dichotomy that involves semantic strategies of positive ingroup and negative out-group presentation to gain approval (van dijk, 1993). considering the polarization of the society over the controversy of shale gas, each side shapes their discourse in a way to persuade the audience that their own standpoint is favorable and legitimate whereas the opponent’s is undesirable and unjustifiable (wodak, 2006). while the dha puts stress on historical and current context of discourse, it also centres its attention on linguistic aspects by employing theories of argumentation, grammar and rhetoric to discover communicative patterns (meyer, 2001). therefore, a dha analysis indicates a comprehensive investigation of the text both at the macro level by means of the historical, social and political background and at the micro level in terms of agency, causality and evaluative lexis. within the framework of the dha, the identification of the discursive strategies which are intentional moves to attain particular social, political, psychological or linguistic goals (reisigl & wodak, 2009), is central to explore discursive representation. although the dha provides five types of discursive strategies, i focused on three of them as it was my aim to explore how legitimation of the in-group and delegitimation of the out-group were realized in construction and representation of identities (reisigl & wodak, 2009). first, the referential or nomination strategies deal with the linguistic construction of social actors in terms of their personal, kinship or occupational relations to each other. second, predicational strategies are concerned with linguistic realizations of evaluative attributions related to social actors of ingroups and out-groups. by means of predication, social actors and actions are identified with reference to quality, quantity, space and time. this type of strategy cannot be sharply separated from the referential or nomination strategies. third, argumentation strategies engage with a number of topoi which are widely used in persuasive rhetoric as topics or commonplaces to justify positive self and negative other representations (balkin, 1996; wodak, 2009). reisigl and wodak (2001) define topoi as content-related warrants or conclusion rules that establish a legitimate link between the argument and the conclusion or claim. as recurrent argumentation schemes which can be applied to different contexts (anscombre, 1995), topoi guide the audience to draw a particular conclusion in a specific situation by means of providing a logical frame which is tightly linked to common sense, reasoning or doxa (kienpointner, 1996). for example, the topos of advantage or usefulness can be used by two conflicting sides like the environmentalists and nuclear energy supporters to legitimize their viewpoints. the study of topoi has so far been limited to political and racist discourse (reisigl &wodak, 2001; richardson, 2004). as this paper was concerned with us vs. them dichotomy that is inscribed to our common sense reasoning, i took advantage of the relevant topoi from the list created by wodak (2001, p.74) (see table 1) to examine discriminatory discourse. discursive construction of the shale gas controversy 133 3.2 analytical framework while working within the framework of the dha to analyse social actor and action representation as well as argumentation strategies, i also took advantage of corpus linguistics. previous research (baker & mcenery, 2005; gabrielatos & baker, 2008) has shown that triangulating corpus linguistics with a discourse analysis approach enables researchers to analyse large amounts of textual data objectively so the credibility of the analysis increases. for the purpose of this study, i used wordsmith 6.0 (scott, 2012) to obtain keywords as well as collocations and concordance lines of the search terms. keyword lists, which are created automatically after statistical comparison of the frequent words in one corpus with a reference corpus, give us ‘‘a measure of saliency’’ (baker, 2006). in line with alexander (2009) who underlines the significance of keywords in diverting attention and backgrounding real problems, the keyword lists helped me investigate not only the lexical discrepancies between both corpora but also provide an insight into how both sides opted for framing the issue of fracturing. with the help of the keyword lists, i obtained a set of search terms concerning shale gas extraction and prevalent social actors in us and them groups. in order to understand how shale gas extraction and relevant social actors were represented in both corpora within the scope of referential or nomination strategies and predicational strategies, i examined the frequent collocations and concordance lines of my search terms. collocational analysis, which is concerned with the statistical co-occurrence of certain words, provides information about the most persistent and prominent ideas associated with a word (gabrieletos & baker, 2008). subsequently, i carried out concordance analysis which can be defined as ‘‘… simply a list of all of the occurrences of a particular search term in a corpus, presented within the context that they occur in; usually a few words to the left and right of the search term’’(baker, 2006, p.71). the automatic display of my search terms within their contexts enabled me to see how the same search terms appeared in both discourses. i also used wmatrix software (rayson, 2009) to automatically find out which semantic fields pervade the corpora of the msc and the aaf. as baker (2006) notes, whilst keyword lists display statistically significant differences between corpora, low frequency synonyms which are used to avoid repetition are ignored in the analysis. analysis of the key semantic categories is a solution to this problem as general meaning is of greater importance than the meaning conveyed by single words. i used usas semantic annotation set within wmatrix which is based upon longman lexicon of contemporary english (mcarthur, 1981). this set is made up of a hierarchical system of semantic categories under 21 main discourse fields with 232 subcategories (see archer, wilson & rayson, 2002). the software groups words into semantic categories that are built upon some level of generality around the same mental concept. for instance, the category w is concerned with the world & our environment whilst w1 is about the universe and w2 is about light. the semantic tagging of the corpus displays general categories of meaning used for the construction of discourse positions by 1. usefulness, advantage 9. finances 2. uselessness, disadvantage 10. reality 3. definition, name-interpretation 11. numbers 4. danger and threat 12. law and right 5. humanitarianism 13. history 6. justice 14. culture 7. responsibility 15. abuse 8. burdening, weighting table 1: list of topoi koca-helvaci 134 opposing camps of debate (baker, 2006). for the analysis of argumentation strategies namely topoi, i focused on the key semantic tags provided by the wmatrix as these represent remarkably prevailing topics and associated argumentation schemes in the corpora of the msc and the aaf (prentice, 2010). 4 data two corpora were built from the documents of the msc and the aaf . the msc defines itself as an organization that works with exploration, production, midstream and supply chain partners in the u.s shale basins to address issues regarding ‘‘the production of clean, jobcreating american natural gas’’ (http://marcelluscoalition.org/). the second corpus was formed by the texts on the website of the aaf which describes itself as a group of entities that aim at banning shale gas extraction to protect vital resources like drinking water and air (https://www.americansagainstfracking.org/). the texts were collected from press releases, newsletters and reports on both websites. after a detailed investigation, it was necessary to specify a time span limit for collection of the texts to have a balanced corpus. although the earliest texts in both websites dated back to 2011, there was no even distribution over the years. between july 2013 and july 2015, there were nearly an equal volume of texts on both websites. therefore, i decided to collect the texts between those dates. table 2 shows the size of both corpora. 5 analysis in order to provide an insight into the discourse worlds of these two opposite groups, keyword lists of both corpora were generated with the help of wordsmith 6.0 (scott, 2012) which conducts a statistical comparison between the word lists to find out unusually frequent words in each of them. the log-likelihood test in this tool automatically calculates the difference between statistical significance of the frequency contrast. due to space limitations, only a limited number of keywords from each corpus were presented and discussed in the following paragraphs. table 3 presents the 20 strongest keywords in the msc corpus in comparison to the aaf corpus. the most striking finding in the keyword table of the msc is the high frequency of the personal pronoun ‘we’ and possessive adjective ‘our’ whilst these were considerably less in the aaf corpus. as function words, pronouns not only indicate language style but also disclose emotional states, personality and other features of social relationships (chung & pennebaker, 2007). such a prevalent use can be a result of the msc’s attempt to present a collective response to a rather contentious issue as ‘we’ not only shows a strong ‘institutional identity’ (sacks, 1992), but also proves the consent of the majority (van leeuwen, 2008). previous research (kacewicz, pennebaker, davis, jeon & graesser, 2013) has also shown less use of first person singular pronouns and more use of first person plural pronouns by high status individuals, which implies being other-oriented. as an industry coalition working for the rights of the shale business companies in the region, the msc puts emphasis on solidarity and altruism by using inclusive and positive sounding terms such as ‘we’ and ‘our’ while describing its aims and deeds. no name of corpus number of words 1 marcellus shale coalition 91,627 2 americans against fracking 73,116 total 164,743 table 2: the size of corpus discursive construction of the shale gas controversy 135 the focus on ‘people’ can also be seen through the frequency of terms that show groups of people through ‘commonwealth, union, pennsylvanians and msc’. this can be viewed as an attempt to create a sense of solidarity with the stakeholders by means of addressing a variety of social groups. the lexis is generally positive as the advantages of shale gas were underlined through vocabulary of finance (fees and revenues), jobs (work, business and manufacturing) and progress (growth, development). in addition, the shale business was represented as an act of philanthropy by means of ‘benefits, opportunities, support, and impact’. on the other hand, the keywords of ‘wolf, severance and tax’ show different discourse prosody as their concordances show a rather negative tone because pennsylvania governor wolf planned to propose a severance tax on the shale industry. however, the tone on the whole is positive which is in line with scollon (2008) who underlines the corporate attitude to background negative facts and events whereas foregrounding everything that can uphold their positive status or value. as textual meaning is created through relations of absence as well as presence (richardson, 2007), i should underline that there are no keywords related to environmental issues. table 4 presents the keyword list of the aaf in comparison to the msc. the keywords in the aaf list can be grouped under three main headings. the first group includes the vocabulary related to the extraction through ‘fracking, chemicals, fluids, wastewater, gallons’ and shale related terms ‘fossil, methane, benzene’. the second group contains vocabulary of natural resources like ‘underground, wells and water’. the final group is concerned with the harmful effects of shale by means of keywords like ‘climate, health, toxic, problems, contamination, pollution, watch’. no keyword msc aaf keyness freq. % freq. % 1 taxes 454 0.54 20 0.02 447.43 2 our 505 0.60 93 0.12 270.68 3 fees 174 0.20 0 0.00 223.18 4 businesses 220 0.26 12 0.01 209.72 5 opportunities 195 0.23 9 0.00 190.92 6 we 462 0.55 116 0.15 186.97 7 impact 258 0.31 36 0.05 166.33 8 manufacturing 127 0.15 1 0.00 152.73 9 msc 102 0.12 0 0.00 130.84 10 revenues 145 0.17 11 0.00 126.64 11 commonwealth 98 0.12 0 0.00 125.71 12 growth 157 0.19 17 0.02 115.48 13 benefits 208 0.25 42 0.05 106.40 14 pennslyvanians 89 0.11 2 0.00 97.92 15 severance 75 0.09 0 0.00 96.19 16 wolf 73 0.09 2 0.00 78.18 17 work 98 0.12 10 0.01 74.02 18 development 151 0.18 32 0.04 71.91 19 support 104 0.12 13 0.02 71.21 20 union 54 0.06 0 0.00 69.25 table 3: the 20 strongest keywords in the msc corpus koca-helvaci 136 contrary to the positivity in the msc keyword list, the tone in the aaf list is rather negative, alarming and pessimistic. another remarkable difference is the formal and scientific style in comparison with the personal style of the msc. institutional authorities like ppinys (the public policy institute of new york state) and epa (environmental protection agency) were the sources of reference. while foregrounding the negative aspects of the shale gas extraction, the absence of keywords showing benefits of shale is in accordance with the aaf’s critical stance. 5.1 representation of social actors in regards of social actors, martin-rojo (1995) observed that the dichotomy between us vs. them draws a line between the inclusive us and exclusive them and prompts an ideological perspective that frames the included as ethical, beneficial and mentally superior while the excluded as immoral, vile and irrational. the comprehensive keyword list of the msc shows that the in-group social actors were heavily represented in terms of ‘we, industry and msc ’ while the out-group was represented by ‘wolf’. table 5 shows the collocations of the social actors in the msc corpus. as the msc is an organization that stands for the interests of the corporations and shareholders within the sector, it is quite normal for the in-group actors to be represented in terms of collective identities. such use has an important role in manufacturing consent as van leeuwen (2008, p.37) put an emphasis on the rule of majority in persuasion as ‘‘what most people consider legitimate’’ is seen as common sense. however, the out-group was represented by governor wolf in terms of individualization. the contradiction between no keyword aaf msc keyness freq. % freq. % 1 fracking 1120 1.49 47 0.06 1350.10 2 water 526 0.70 54 0.07 481.65 3 wells 685 0.90 275 0.33 274.85 4 methane 170 0.23 9 0.01 194.70 5 fluids 188 0.24 22 0.03 169.36 6 climate 148 0.20 10 0.01 159.83 7 pollution 125 0.17 5 0.00 151.14 8 health 201 0.27 34 0.04 150.24 9 wastewater 108 0.14 2 0.00 144.23 10 contamination 109 0.14 4 0.00 133.70 11 chemicals 120 0.16 10 0.01 121.94 12 toxic 74 0.10 0 0.00 110.75 13 ppinys 71 0.09 0 0.00 106.26 14 fossil 81 0.11 3 0.00 99.19 15 epa 138 0.18 25 0.03 98.92 16 gallons 64 0.08 1 0.00 86.73 17 watch 68 0.09 2 0.00 86.17 18 problems 89 0.12 9 0.01 84.62 19 benzene 56 0.07 0 0.00 83.81 20 underground 82 0.11 7 0.00 82.67 table 4: the 20 strongest keywords in the aaf corpus discursive construction of the shale gas controversy 137 us through collectivization and them by means of individualization is a deliberate effort to prove the legitimacy of the majority over the deviant minority. social actors within the us category were represented positively by means of lexis showing collective force (companies, members, committee, advocates, public, veterans, businesses), progress (work, make, continue, increase, grow, create, generate, emerging, fastest growing, growth, leading, dynamic, better, higher, development), financial benefits (job, employment, fees) altruism (cooperative, support, supportive, contribute, sponsor, fund, opportunity) and reliability (trusted, responsible, careful, commitment). the use of modality (will, can, be able to) also requires explanation as it indicates the writer/speaker’s attitude towards the content or object of his/ her message (hodge & kress, 1996). a closer look at the concordance lines of these modals shows that ‘will’ was used to show determination while ‘can and be able to’ were indicating the competence and power of the in-group social actors. (1) shows how the msc praised its own activity with emphasis on self-dedication, communicative attitude, scientific evidence, transparency and benevolence: (1)we are committed to being responsible members of the communities in which we work; we encourage spirited public dialogue and fact-based education about responsible shale gas development; and we conduct our business in a manner that will provide sustainable and broad-based economic and energy-security benefits for all (the msc, 30.07.2015). the only social actor in the them category was represented in terms of functionalization as a governor, for he was a high status social actor. contrary to the positivity in the us category, he was described rather negatively through lexis implying financial burden (tax, taxes, higher, severance) and danger (damaging, threatening). lexis about a hypothetical future (plan(s), proposed, proposal) reinforced the idea that it was necessary to stop the no collocation no collocation we 462 our, have, are, were, need, know, see,make, work, think, continue, will, must, can, able, gas, natural, shale, all, more, better, co-operative, supportive, dynamic wolf 73 taxes, energy, proposed, proposal, plan, plans, severance, higher, governor, administration, natural, damaging, threatening msc 102 president, companies, members, committee, will, working, gas, advocates, trusted, responsible, careful, leading industry 461 gas, shale, natural, oil, our, energy, will, impact, marcellus, well, growth, taxes, growing, support, state, businesses, higher, want, needs, grow, economic, continues, jobs, veterans, working, job, work, public, opportunities, employment, represent, development, fees, commitment, companies, generates, creates, contributes, increases, sponsors, funds, dynamic, emerging, fastest growing table 5: collocations of social actors in the msc corpus us them koca-helvaci 138 harmful plans of the governor. in (2), we can see how wolf was criticized in terms of destruction by oppressive, stunt, jeopardize, threatening, and reverse positive gains. he was represented as an obstacle before the economic prosperity promised by the shale industry. the discourse surrounding him was shaped by fear-inducing rhetoric with reference to the 2008 recession. (2) wolf is demanding oppressive new taxes that will stunt economic growth, job creation and our region’s manufacturing potential. at the same time, critical impact tax funding would be jeopardized, threatening important local infrastructure investments. marcellus shale has been one of the few bright spots in a historically slow economic recovery. gov. wolf’s actions seem designed to reverse these positive gains. (the msc, 5.1.2015) table 6 below presents the social actors in us and them categories of the aaf corpus. surprisingly, the actors in the us part are statistically less frequent than the ones in the them part. this difference can be seen as an indication of the aaf’s attempt to focus on the negative qualities of the pro-shale group. the collocations of the in-group can be grouped under main headings of modality (must, will, can, need, should), volition (want, urge, demand), conservation (protect, responsible, safeguarding) and community (americans, national, coalition, organization). a closer examination of the concordance lines of the modal verbs and expressions shows that deontic modality, which is ‘‘concerned with influencing actions, states or events’’ (palmer, 1990, p.6) through emphasis on obligation, prohibition and advice, dominates the discourse of the aaf. given that, we can say that the aaf discourse was shaped by a call for immediate action to stop fracturing. the infrequency of the social actors in the us category may stem from the structure of the organization as the aaf stands for loosely connected subgroups. as in (3), the us them no collocation no collocation we 116 have, are, must, can, should, will, need, do, know, our, want, urge, depend, move, drink, protect, responsible, safeguarding industry 328 gas, oil, fracking, shale, fossil, fuel, pollution, wastes, water, energy, natural, data, spending, unrestricted, public, government, funded, access, clen, climate, policy, drilled, harm, emissions, lobbying, rhetoric, claims, jobs, workers aaf 18 americans, national, coalition, organization, water, science ppinys 71 jobs, protection, scenario, claimed, correction, incorrect, misused, exaggerated, estimate epa 138 water, drinking, standard, study, investigation, administration, released, report, wells, found, methane, fracking, emission, greenhouse table 6: collocations of the social actors in the aaf corpus discursive construction of the shale gas controversy 139 subgroups which form the aaf were frequently mentioned throughout the corpus to put stress on the heterogeneity and multivocality of the organization. although the scarcity of the ingroup actors as a single body can be seen as a drawback, for it communicates a weak organizational attitude, the emphasis on subgroups can be viewed as an indication of the credibility and validity of the aaf’s, goals which unite groups with different backgrounds and interests. the focus on diversity and numbers in (3) realized argumentum ad populum which refers to the wide support for justifying a viewpoint. (3) on the public comment deadline for the environmental protection agency’s (epa) proposed power plant rules, americans against fracking, a national coalition to ban fracking, delivered a letter from over 250 environmental, health, labor and consumer protection groups, along with over 200,000 comments criticizing the rules for incentivizing fracked natural gas.(the aaf, 1.12.2014) the actors in the them part were not only more frequent but also identified in a more varied way. the collocations show that they were associated with wrongdoing (pollution, wastes, unrestricted, access, harm, greenhouse, emissions, methane), dishonesty (lobbying, rhetoric, claims, correction, incorrect, misused, exaggerated, scenario), environment (fossil, shale, water, oil, wells), bureaucracy (government, administration, policy). although ‘jobs and workers’ and scientific research (investigation, study, found, data, estimate, release, report) seemed to be positive associations of this group, their concordance lines show that they were contextualized in a way to criticize the pro-shale group actors for distorting the truth about scientific evidence and job statistics. criticism over honesty can be observed in (4) which blamed out-group social actors for telling lies to manufacture consent. (4) new industry-sponsored study on fracking which is more spin than science produces findings dramatically out of step with recent studies (the aaf, 16.09.2013) 5.2 representation of shale gas extraction the second stage of my analysis is concerned with the representation of the shale gas extraction. that is, what is preferred of all the variety of choices and what qualities are attributed to the action. the comprehensive keyword lists of both corpora show that the extraction process was named as ‘hydraulic fracturing, fracking and shale development’. in table 7, we can see the distribution and representation of these terms in the msc corpus. the least used term is obviously hydraulic fracturing, the collocations show that it was treated merely as a technical term. fracking, which has negative connotations, was also used by the msc with an attempt to embrace it rather than avoiding it (zimmer, 2014). although it was not used very often, the collocations such as favor and revolutionizing show that the industry tried to reverse the pejorative meaning loaded by the antishale groups. despite these attempts to change the image of fracking, it is still seen that the pro-shale group heavily used a more neutral term shale development to represent the extraction process. while the term itself hides the negative aspects of drilling and fracturing, the associations are very positive. the collocations can be categorized into three main groups: community (farmers, commonwealth), innovation (job-creating, game-changing, renaissance, transformational, new, generate, opportunities), safety (safe, responsible, tightly-regulated). koca-helvaci 140 in (5), we can see how shale development was appreciated through security measures and its contribution to the job market. (5) a new study released this week – study of construction employment in marcellus shale related oil and gas industry– highlights the clear jobcreating benefits tied to safe shale development across our labor force. (the msc, 17.10.2014) although the same terms were used in the aaf corpus as well, there are substantial differences between the frequency and usage. table 8 presents the distribution and lexical associations of the search terms in the aaf corpus. as in the msc corpus, hydraulic fracturing was used as a technical term. however, the collocations here involve the vocabulary related to the chemical solvent and its effects whilst action no collocation hydraulic fracturing 25 horizontal, drilling, natural, gas fracking 47 said, favor, help, revolutionizing, potential shale development 139 continues, generates, supports, presents, farmers, commonwealth, playing, tax, good, new, opportunities, pervasive, game-changing, jobcreating, safe, responsible, renaissance, tightlyregulated, transformational table 7: representation of the extraction process in the msc corpus action no collocation hydraulic fracturing 47 injects, water, mixture, used, fluids, drinking fracking 1120 fluids, gas, water, groundwater,drinking, wastewater, wells, oil, operation, chemicals, chemical,natural, public, safe, environmental, contamination, contaminate, pollution, period, associated, sites,climate, cancer, risks, underground, dangerous, unsafe, methane, increased, poses, polemical shale development 32 dangerous, risks, problems, controversial, costs table 8: representation of the extraction process in the aaf corpus discursive construction of the shale gas controversy 141 in the msc the term is concerned with the technique and product. contrary to the msc, shale development has the smallest share and it was associated with harm (dangerous, risks, problems, costs) and dissent (controversial). on the other hand, the pejorative term fracking was the most frequently used one in the aaf corpus. the action was associated with environmental harm (wastewater, waste, chemicals, contamination, contaminate, pollution, methane), threat (cancer, risks, dangerous, unsafe, poses, polemical) and resources (fluids, gas, water, oil, groundwater, underground). as all these three terms were related to the actions of the industry which the aaf opposed, the discourse concerning them was highly negative. in (8), fracking was identified as a menace to the environmental heritage of americans through emotive vocabulary such as national treasure, sacred trust which aims at evoking nationalistic feelings. (8) “our public lands are a national treasure and a sacred trust passed by one generation of americans to another,” said drew hudson of environmental action. “fracking on public lands threatens the drinking water of millions of people, including the president’s daughters and everyone else here in washington, d.c. it would also poison many of our last wild and pristine ecosystems. (the aaf, 22.08.2013) 5.3 argumentative schemes in this section, the interplay between key semantic categories and argumentative schemes namely topoi were studied to analyze the construction of discourse positions on the different sides of the shale controversy. the usas annotation system within wmatrix generated comparative lists of key semantic categories for both corpora. as i was mainly interested in topoi, i only focus on the semantic categories which instantiated argumentative schemes in both tables. in table 9, the key semantic categories in the msc corpus in comparison to the aaf corpus can be viewed. the most prevalent argumentative scheme is the topos of advantage or usefulness which means ‘‘if an action under a specific relevant point of view will be useful, then one should perform it’’ (wodak, 2001, p.74). through the categories of helping (support, benefit, promote, encourage, boost), chance, luck (opportunity, chance), size: big (growth, expansion, substantial), evaluation: good (advance, progress, recover, advantage), the industrial activity advocated by the msc was praised for contributing positively to american social welfare. the subcategory of this topos is pro bono publico (to the advantage of all) for the shale industry was framed as a savior for the whole american nation who suffered from the recession. the categories concerning money (2, 16, 18), work (7), business (3, 14) were also concerned with the topos of advantage or usefulness by means of specifying the benefits of shale for the job market and regional economy. however, the concordance lines of the money-related categories were also framed by the topos of danger and threat which indicates that ‘‘if a political action or decision bears specific dangerous, threatening consequences, one should not perform or do it’’ (wodak, 2001, p.75) with a focus on the harms of prospective tax rise for the shale industry. this topos led to a victim-victimizer reversal in which the proshale group casts itself into the role of the victim who was threatened by oppressive tax schemes in return of their contribution to american welfare. the issue of tax rise was also characterized by the topos of finance which means ‘‘if a specific situation or action costs too much money or causes a loss of revenue, one should perform actions which diminish the costs or help to avoid the loss’’ (wodak, 2001, p.76). koca-helvaci 142 the categories of politics and government show that the msc tried to convince the politicians to oppose the tax regulation. as exemplified in (9), with reference to negative socio-economic consequences of a prospective tax rise, the msc addressed a variety of stakeholders to stand against this legislation. the word ‘sting’ was used as a metaphor to identify anti-shale politicians in harrisburg as insects that would cause harm to pennsylvanians. those in opposition were also criticized for being irrational with reference to the need for common sense policies. (9) as some in harrisburg seek to pass even higher energy taxes that would sting pennsylvania consumers and families, local officials across the political spectrum, building and labor trade unions as well as small businesses are speaking out loudly, clearly and in a united voice for common sense policies that create jobs and opportunities. (the msc, 5.05.2015) table 10 below presents the top 20 key semantic categories in the aaf corpus with comparison to the msc. as in the msc, only the semantic categories that were identified with topoi were discussed. the most pervasive argumentation scheme is the topos of danger and threat which was instantiated through the categories of disease (cancer, symptoms, headache), green issues (pollution, contaminants, epa, environmental resources), damaging and destroying (destructive, harmful, accident, leaks, ruin), medicines and medical treatment (public health, hospital), weather (climate change, drought), danger (risky, dangerous, hazards), temperature: hot/ on fire (burning, global warming) and violent/ angry (aggressive, attack, threaten). this topos is closely linked with the topos of burdening or weighing down which can be described as ‘‘if a person, an institution or a country is burdened by specific problems, one should act in order to diminish these burdens’’ (wodak, 2001, p.76). no semantic category msc aaf keyness freq. % freq. % 1 pronouns 3952 5.14 1962 2.83 491.05 2 money and pay 981 1.28 203 0.29 478.33 3 business:generally 835 1.09 246 0.36 280.46 4 helping 873 1.14 300 0.43 235.58 5 speech: communicative 723 0.94 239 0.35 207.04 6 personal names 827 1.08 335 0.48 166.79 7 work and employment: generally 826 1.07 367 0.53 136.53 8 time: future 546 0.71 211 0.30 120.57 9 chance, luck 211 0.27 39 0.06 112.79 10 participating 121 0.16 11 0.02 96.13 11 warfare, defence and the army;weapons 141 0.18 21 0.03 87.52 12 size: big 433 0.56 185 0.27 78.15 13 government 773 1.01 416 0.60 74.80 14 business: selling 459 0.60 206 0.30 74.10 15 politics 288 0.37 103 0.15 72.87 16 money: cost and price 405 0.53 174 0.25 72.22 17 evaluation: good 417 0.54 200 0.29 56.98 18 money generally 222 0.29 79 0.11 56.63 19 places 1000 1.30 618 0.89 55.54 20 belonging to a group 813 1.06 502 0.73 45.34 table 9: the 20 strongest semantic categories in the msc corpus discursive construction of the shale gas controversy 143 the categories of exceed; waste (waste, over) and difficult (problems, crisis, challenge, difficulty) are also related with the topos of burdening or weighing down. the opposition of the anti-shale group was legitimized through representation of the shale industry not only as a threat but also as a burden for natural resources and public health. the concordance lines showed that governmental action was demanded to reduce the harms of fracturing. the topos of abuse, which can be paraphrased as if a right is misused, measures must be taken to stop the abuse (wodak, 2001), was also employed to criticize legal loopholes which enabled the corporate abuse of the environment through the category of no constraint (unrestricted, unchecked, lax). the present and future effects of this unrestrained industrial activity were also represented in terms of the topos of numbers which states that ‘‘if the numbers prove a specific topos, a specific action should be performed or not be carried out’’ (wodak, 2001, p.76). here, the statistical data not only aggravated the negative projections of the aaf but also presented the propositions as more verifiable (semin & fiedler, 1988), true (hansen & wanke, 2010) and probable (tversky &kahneman, 1982). the last argumentation scheme is the topos of responsibility which can be summarized as if a state or group creates a problem, they should resolve it (wodak, 2001). this topos was instantiated by the category not allowed (ban, prohibit, suppression) which underlined the anti-shale group’s demand for government action to regulate industrial activity rather than encouraging it for the sake of political and economic gains. in (10), we can observe topos of danger and threat as well as topos of responsibility. the negative tone was sharpened by alarming vocabulary (poison, toxic, threaten) and focus on children. the reference to the ‘investigation by the pulitzer prize winning journalists’ is argumentum ad verecundiam as the author portrayed the journalists as a credible authority without providing any scientific data no semantic category aaf msc keyness freq. % freq. % 1 substances and materials: liquid 1339 1.93 324 0.42 775.80 2 geographical terms 936 1.35 352 0.46 339.25 3 disease 451 0.65 79 0.10 328.72 4 substances and materials: gas 1474 2.13 861 1.12 233.26 5 substances and materials generally 431 0.62 121 0.16 218.16 6 green issues 541 0.78 208 0.27 190.20 7 damaging and destroying 270 0.39 54 0.07 180.65 8 weather 198 0.29 28 0.04 162.36 9 danger 270 0.39 74 0.10 140.09 10 drinks and alcohol 165 0.24 28 0.04 122.56 11 substances and materials: solid 268 0.39 84 0.11 121.31 12 no constraint 248 0.36 74 0.10 118.30 13 medicines and medical treatment 308 0.44 119 0.15 107.54 14 deciding 152 0.22 34 0.04 93.76 15 not allowed 79 0.11 6 0.01 82.32 16 exceed; waste 181 0.26 61 0.08 75.42 17 temperature: hot/on fire 119 0.17 31 0.04 64.70 18 numbers 1375 1.99 1119 1.46 59.88 19 violent/angry 189 0.27 85 0.11 52.10 20 difficult 149 0.22 60 0.08 49.01 table 10: the 20 strongest semantic categories in the aaf corpus koca-helvaci 144 about their research. the validity of the aaf’s perspective was also aimed to be strengthened by means of ‘hundreds of complaints’ which is an example of argumentum ad populum as it represented the opposition as a widely held view. (10) the 8-month ‘fracking the eagle ford shale’ investigation by pulitzer prize winning journalists reveals that fracking is literally poisoning the air children and families breathe. polluted with toxic chemicals like hydrogen sulfide and benzene, air poisoned by fracking is entering homes, daycare centers and schools throughout entire regions. this investigation and the hundreds of complaints build on an already significant body of science showing that fracking inherently poisons the air and threatens people’s health. for the sake of our health and the wellbeing of our communities, fracking must be banned. (the aaf, 18.02. 2014) 6 conclusion this study set out with the aim of finding discursive strategies adopted by two opposite groups to justify their standpoint in the fracturing controversy. considering the discourse of the pro-shale group, the findings of the study are consistent with the results of previous research (mercer, rijke & dressler, 2014; jaspal & nerlich, 2014), as the discourse of the msc which champions the shale gas industry was largely shaped by the identification of shale gas with financial gains, job opportunities, social welfare, nationalism, and philanthropy. this organization’s fixation with industrial growth, economic profits, reduced governmental intervention and tax reductions overlap with neoliberalism which regards environment as a mere instrument for the generation of economic gains for human beings (egri & pinfield, 1996). in order to fend off criticism concerning commodification of nature for corporate interests, the msc frequently underlined that the whole american nation were the beneficiaries of their industrial manufacturing. the focus on regional pride associated with energy production and collective benefits portrays extraction as ‘‘classless and horizontally beneficial to all members of the community’’ (gaventa, 1982, p.58) with an intention to mask the corporations’ profits and environmental harm. for residential support, master narratives of nationalism and ‘american dream’ with reference to progress, prosperity and power (bell & york, 2010) were used. therefore, bans and tax regulations were framed as attacks on national interests. the msc also took advantage of intimidating rhetoric by establishing a cause and effect link between tax rise and economic recession. the most striking finding about this corpus is the highly positive representation of the in-group actors and actions with an emphasis on collective aims and gains. the topos of advantage or usefulness also contributed to this positive ‘self representation’. on the other hand, negative ‘other representation’ was considerably less and only built upon criticism towards government regulation. the results concerning the discourse of the anti-shale group aaf corroborate the findings of jaspal and nerlich (2014), mercer, rijke and dressler (2014) and bomberg (2015) as shale gas was associated with threat. as a pro-environmental group, the discourse of the aaf is educating, alerting and mobilizing (cox, 2010) through negative associations of shale gas with irreversible environmental damage and threat to human health. the term ‘fracking’ was defined through lexis showing disease, destruction and demise. in line with anderson (1997) who underlined the importance of using themes that people can easily identify with, the aaf largely built its discourse on the drinking water problem and health hazards. this group also took advantage of nationalism by putting stress on the vitality of protecting the ecosystem of the us. scientific evidence and statistical data as authoritative resources (ozawa, 1996) were used to persuade the stakeholders to unite against the shale gas extraction. it is interesting to discursive construction of the shale gas controversy 145 note that this group built its discourse upon negative ‘other representation’ whilst there was remarkably less emphasis on its own group identity. to prove their organizational legitimacy, both sides of the controversy shaped their discourses in a way to create the image of a proper and beneficial entity that conformed to the norms, interests, definitions and values of the american society (suchman, 1995). by means of triangulating quantitative and qualitative analysis to explore the discourses surrounding water-energy nexus which is an increasingly important area in the us, this study not only adds to the literature on self and other representation but also contributes to the linguistic analysis of environmental controversies. references richard alexander (2009). framing discourse on the environment: a critical discourse approach. routledge, new york/ london. alison anderson (1997). media, culture and the environment. ucl press, london. jean-claude anscombre (editor) (1995). théorie des topoi. editions kimé, paris. dawn archer, andrew wilson and paul rayson (2002). introduction to the usas category system: benedict project report. available online at http://ucrel.lancs.ac.uk/usas/ usas% 20guide.pdf. paul baker and tony mcenery (2005). a corpus-based approach to discourses of refugees and asylum seekers in un and newspaper texts. language and politics, 4 (2), 197-226. paul baker (2006). using corpora in discourse analysis. continuum, london/ new york. ronald s. balaba and ronald b. smart (2012). total arsenic and selenium analysis in marcellus shale, high-salinity water, and hydrofracture flowback wastewater. chemosphere, 89 (11), 1437-1442. shannon elizabeth bell and richard york (2010). community economic identity: the coal industry and ideology construction in west virginia. rural sociology, 75(1), 111–143. elizabeth bomberg (2013). the comparative politics of fracking: networks and framing in the us and europe. apsa 2013 annual meeting paper; american political science association 2013 annual meeting. available online at https://ssrn.com/ abstract=230 1196 elizabeth bomberg (2015). 'shale we drill? discourse dynamics in uk fracking debates' journal of environmental policy and planning, 1-17. jonathan burgess, emily kneebone, vaios liapis and laura swift (2001, november 01). topos. the literary encyclopedia. volume 1.1.1: archaic, classical, hellenistic and imperial greek writing and culture, -800--100. available online at https:// www. litencyc.com/php/stopics.php?rec=true&uid=1126 elizabeth burleson (2012). cooperative federalism and hydraulic fracturing: a human right to a clean environment. cornell journal of law and public policy, 22, 289348. http://ucrel.lancs.ac.uk/usas/%20usas%25%2020guide.pdf http://ucrel.lancs.ac.uk/usas/%20usas%25%2020guide.pdf https://ssrn.com/%20abstract=230%201196 https://ssrn.com/%20abstract=230%201196 https://www.litencyc.com/php/browse_volume.php?id_display=15 koca-helvaci 146 cindy chung and james pennebaker (2007). the psychological function of function words. in k. fiedler (ed.), frontiers of social psychology , 343-359. new york, ny: psychological press. matthew cotton, imogen rattle and james van alstine (2014). shale gas policy in the uk. an argumentative discourse analysis. energy policy 73, 427-38. robert cox (2010). environmental communication and the public sphere. sage, california. carolyn p. egri, c. and lawrence t. pinfield (1996). organizations and the biosphere: ecologies and environments. in stewart r. clegg, cynthia hardy and walter r. nord (editors), handbook of organization studies , 459-483. sage, london. terry engelder (2011). should fracking stop?. nature, 477, 274-275. darrick evensen, chris clarke and richard c.stedman (2014). a new york or pennsylvania state of mind: social representations in newspaper coverage of gas development in the marcellus shale. journal of environmental studies and sciences , 4:65–77. michael h. finewood and laura stroup (2012). fracking and the neoliberalization of the hydrosocial cycle in pennsylvania’s marcellus shale. journal of contemporary water research & education, 147, 72-79. costas gabrielatos and paul baker (2008). fleeing, sneaking, flooding: a corpus analysis of discursive constructions of refugees and asylum seekers in the uk press, 1996-2005. journal of english linguistics, 36, 5-38. john gaventa (1982). power and powerlessness: quiescence and rebellion in an appalachian valley. university of chicago press, chicago, il. jochim hansen, j and michaela wanke (2010). truth from language and truth from fit: the impact of linguistic concreteness and level of construal on subjective truth. personality and social psychology bulletin, 36, 1576-1588. robert hodge and günther kress (1996). language as ideology. routledge, london. rusi jaspal and brigette nerlich (2014). fracking in the uk press: threat dynamics in an unfolding debate. public understanding of science, 23 (3), 348-363. ewa kacewicz, james pennebaker, matthew davis, moongee jeon and arthur graesser. (2013). pronoun use reflects standings in social hierarchies. journal of language and social psychology, xx(x) 1–19. david kay (2011). the economic impact of marcellus shale gas drilling: what have we learned? what are the limitations?. working paper series: a comprehensive economic analysis of natural gas extraction in the marcellus shale. cornell university press, new york. manfred kienpointner (1996). vernunftig argumentieren: regeln und techniken der diskussion. rowohlt, hamburg. luisa martin-rojo (1995). division and rejection: from the personification of the gulf http://www.springer.com/environment/journal/13412 discursive construction of the shale gas controversy 147 conflict to the demonization of saddam hussein. discourse & society, 6(1), 49–80. andrew maykuth (2011, april 18). strong positions on either side of “fracking” at epa hearing. philadelphia inquirer. available online at http://www.philly.com/philly/news/special_packages/inquirer/marcellus-shale /2010 0914_strong_positions_on_either_side_of__quot_fracking_quot__at_epa_hear ing.html tom mcarthur (1981). longman lexicon of contemporary english. longman, london. alexandra mercer, kim de rijke, wolfram dressler (2014). silences in the boom: coal seam gas, neoliberalizing discourse, and the future of regional australia. journal of political ecology, 21, 279-302. michael meyer (2001). between theory, method and politics: positioning of the approaches to cda. in ruth wodak and michael meyer (editors), methods of critical discourse analysis (pp. 14–31). sage, london. connie p. ozawa (1996). science in environmental conflicts. sociological perspectives, 39(2), 219–231. frank robert palmer (1990). modality and the english modals. longman, london. sherly prentice (2010). using automated semantic tagging in critical discourse analysis: a case study on scottish independence from a scottish nationalist perspective. discourse & society, 21, 405-437. paul rayson (2009). wmatrix: a web-based corpus processing environment, computing department, lancaster university. available online at http://ucrel.lancs.ac.uk/wmatrix/ martin reisigl and ruth wodak (2001). discourse and discrimination: rhetorics of racism and anti-semitism. routledge, london. martin reisigl and ruth wodak (2009). the discourse-historical approach (dha). in ruth wodak & michael. meyer (editors) methods of critical discourse analysis, 2nd edn (pp. 87-121). sage, london. john e. richardson (2004). (mis)representing islam: the racism and rhetoric of british broadsheet newspapers. john benjamins publishing, amsterdam/ philadelphia. john e. richardson (2007). analysing newspapers: an approach from critical discourse analysis. palgrave macmillan , hampshire/ new york. harvey sacks (1992). lectures on conversations (vol. 1 and 2). blackwell, oxford. mira schirrmeister (2014). controversial futures—discourse analysis on utilizing the “fracking” technology in germany. eur j futures res, 2, 38-46. ron scollon (2008). analyzing public discourse: discourse analysis in the making of public policy. routledge, london. mike scott (2012). wordsmith tools version 6. stroud: lexical analysis software. http://www.philly.com/philly/news/special_packages/inquirer/marcellus-shale%20/2010 koca-helvaci 148 giin r. semin and klaus fiedler (1988). the cognitive function of linguistic categories in describing persons: social cognition and language. journal of personality and social psychology, 54, 558-568. michael f. smith and denise p. ferguson (2013). fracking democracy: issue management and locus of policy decision-making in the marcellus shale gas drilling debate. public relations review, 30, 377-386. mark c. suchman (1995). managing legitimacy: strategic and institutional approaches. academy of management review, 20 (3), 571-610. amos tversky and daniel kahneman (1982). judgments of and by representativeness. in daniel kahneman, paul slovic and amos tversky (editors.), judgment under uncertainty: heuristics and biases (pp. 84-100). cambridge university press, cambridge, uk. us energy information administration. (2015). annual energy outlook. washington, dc: u.s. department of energy. teun van dijk (1993). elite discourse and racism. sage, newbury park, ca. teun van dijk (1998). ideology: a multidisciplinary approach. sage, newbury park, ca. theo van leeuwen (2008). discourse and practice. oxford university press, new york. aleksandra wagner (2015). shale gas: energy innovation in a (non-)knowledge society: a press discourse analysis. science and public policy, 42, 273-286. ruth wodak (2001). the discourse-historical approach. in ruth wodak and michael meyer (editors), methods of critical discourse analysis (pp. 63-94). sage, london. ruth wodak (2006). history in the making/ the making of history: the ‘german wehrmacht’ in collective and individual memories in austria. journal of language and politics, 5(1), 25–54. ben zimmer (2014, october 3). a push to make ‘fracking’ sound better. can the word lose its bad reputation?. the wall street journal. available online at http://www.wsj.com/articles/can-the-word-fracking-lose-its-bad-reputation1412358270 http://www.wsj.com/articles/can-the-word-fracking-lose-its-bad-reputation1412358270 dialogue & discourse 13(1) (2022) 41–62 doi: 10.5210/dad.2022.102 opinion piece: can we fix the scope for coreference? problems and solutions for benchmarks beyond ontonotes amir zeldes amir.zeldes@georgetown.edu georgetown university department of linguistics editor: massimo poesio submitted 04/2021; accepted 01/2022; published online 04/2022 abstract current work on automatic coreference resolution has focused on the ontonotes benchmark dataset, due to both its size and consistency. however many aspects of the ontonotes annotation scheme are not well understood by nlp practitioners, including the treatment of generic nps, noun modifiers, indefinite anaphora, predication and more. these often lead to counterintuitive claims, results and system behaviors. this opinion piece aims to highlight some of the problems with the ontonotes rendition of coreference, and to propose a way forward relying on three principles: 1. a focus on semantics, not morphosyntax; 2. cross-linguistic generalizability; and 3. a separation of identity and scope, which can resolve old problems involving temporal and modal domain consistency. keywords: coreference, annotation, guidelines, anaphora, referring expressions, predication, scope, multilingual, ontonotes 1. introduction coreference resolution is the task of delineating and grouping together referring expressions in text so that spans of text referring to the same discourse entity are clustered together. the past decade has seen remarkable improvements in the consistency and performance of coreference resolution, first through the creation of large (>1 m tokens) benchmark data in ontonotes (hovy et al. 2006; weischedel et al. 2012; hence on), and then through use of that dataset in the conll shared task on coreference resolution (pradhan et al., 2011) and the development of evaluation metrics (pradhan et al. 2014; see moosavi and strube 2016 for criticism). with the advent of end-to-end neural approaches to coreference resolution (lee et al., 2017), we have seen scores on coreference resolution rise from the mid-50s at the 2011 shared task to current sota scores around 80 points when gold speaker information is provided (joshi et al., 2019; wu et al., 2020), as evaluated by the standard metrics on the on test set. at the same time, coreference resolution using the on scheme as a target has been plagued by a number of issues: the lack of annotation of singletons (entities mentioned only once) has led to systems conflating referentiality recognition (whether an expression in fact refers to some entity) i would like to thank anna nedoluzhko, massimo poesio, sameer pradhan, nathan schneider and the anonymous reviewers for valuable comments and discussions about previous versions of this paper; the usual disclaimers apply. ©2022 amir zeldes this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). zeldes with coreference recognition (whether a referring expression is mentioned more than once, see lee et al. 2017; zhang et al. 2018); omission of annotation of indefinite anaphors has led to counterintuitive system behaviors (e.g. repeated mention of specific indefinites and generics is ignored; see section 2.1 and the appendix for examples); and idiosyncratic handling of some constructions (especially copula predication, compounding) has meant that even when systems perform well in predicting on-like annotations, the results are often less informative than they could be for practical applications (see section 2.4). additionally, because the on guidelines for english depend on morphosyntactic constructs such as indefinite nps, copula predication and compounding, they are not readily applicable to other languages, leading to rather different practices in corpora for languages other than english, including in the arabic and chinese sections of the ontonotes corpus itself. i would like to start by saying that i have been a passionate admirer and intensive user of ontonotes for a long time, but i also suspect that dissatisfaction with some aspects of its annotation scheme is widespread: this is evidenced by the emergence of the coreference resolution beyond ontonotes workshops since 2016 (corbon, and later crac, ‘computational models of reference, anaphora and coreference’), the continuing creation of resources with different annotation schemes (cf. stoyanov et al. 2009; zeldes and zhang 2016; see poesio et al. 2016 for an overview), the focus of the 2021 coreference shared task on non-on data, and most recently the launching of a universal anaphora (ua) initiative, aimed at bringing some consistency into the growing variety of divergences between resources.1 however, i suspect that while we as a research community can arrive at a consensus about this dissatisfaction, it will be much harder to reach a consensus about what a solution might look like, and how it could be attained feasibly. my goal in writing this overview is therefore first to outline some of the reasons why we need to do something, and my audience is, in the first instance, the community working on language resources for coreference, and in the second instance, nlp researchers, and especially grad students, who are entering the field and may get the impression that the on standard represents a final end point of what we want from coreference resolution. the suggestions below probably form a continuum between more and less controversial, and their ranking on different researchers’ wish lists will differ as well based on their respective goals, but it’s a discussion we can only have if the commonalities and differences between coreference datasets and task definitions are laid out clearly.2 2. the 90% and 10% in ontonotes the development and release of on were landmark achievements for a number of annotation types covered by the corpus, but especially transformative for coreference resolution. one of the declared goals of the project, as presented in the hovy et al. paper and its title, was to reach a 90% solution: agreement scores for annotations had to be in the 90s, casting consistency as a necessary condition for any guidelines developed for the corpus. for coreference, this meant focusing on some of the easier phenomena to agree on, such as antecedents for pronominal anaphora, while more tricky cases, such as indefinite anaphors, were left unannotated – not because they were deemed disinteresting or unuseful, but because they were a lower priority target (sameer pradhan, p.c.). this decision led to high agreement, but also to a number of unexpected results. 1see http://universalanaphora.org/ and nedoluzhko et al. (2022). 2for readers who prefer to start with worked examples, see the appendix for some analyses contrasting the on view of coreference with a more unrestricted annotation scheme. 42 can we fix the scope for coreference? in this section i will survey some of the most surprising cases for lay users or researchers new to coreference resolution; for quite a few cases, i am certain that the majority of people using offthe-shelf coreference tools are not well aware of the specifics or consequences, and many people working with on data directly are also not likely to know about all. in the following section i will argue that for all of these, semantic reference to the same entity, rather than morphosyntactic environments, should guide the coreference decision. 2.1 indefinites and generics on guidelines prohibit “generic, underspecified, and abstract nominal mentions” (bbn technologies, 2007, 5) from being linked to other such mentions, meaning that they can only be referred back to by pronouns or definite nps.3 although indefinite nps can indicate a variety of semantic or pragmatic functions (bhatia et al., 2014), all “indefinite noun phrases, which begin with the indefinite article (a, an)” (bbn technologies, 2007, 6) as well as bare nps (bare plurals and singulars with no article) are considered generic by definition in on. thus the following are not annotated in the corpus in any way: (1) [program trading]i is “a racket,” ... [program trading]i creates deviant swings (on, wsj_0121, no coreference in the corpus) this indirectly implies that coreference resolution systems pairing a pronoun or definite mention with the first np in such a chain will be penalized, and that systems must learn in an environment in which ‘it’ + ‘program trading’ is once a positive match and once a negative one, despite the fact that it is semantically the same ‘it’, which in fact refers to the generic concept of program trading. the same guideline applies to generic pronouns, as in the following example, which is also left unannotated in on. (2) [you]i couldn’t start unless [you]i knew that the replacement heart would make it to the operating room (bbn technologies, 2007, 5) in this case, systems must again learn that the usual coreference between two instances of ‘you’ in the same sentence does not apply, even though semantically, i would again argue that both pronouns do in fact refer to the same (albeit generic) thing – anyone who could have started a heart replacement in this context. from a set-theoretical perspective it may even be possible to argue that that set (people who could have started heart replacements at the time) is completely specified, and non-generic (or at least, we could continue the text in a way which explicitly states the names of the relevant people). the problems with the abstract/generic guideline as a foundation for universal, cross-linguistic coreference annotation are numerous: 3an anonymous reviewer has asked about the sense of ‘referring’ or ‘referentiality’ intended here. the term is unfortunately murky: on the one hand, theoretical literature going back to at least reinhart (1982) has referred to cognitive models such as ‘file cards’ per entity (see krifka 2008 for discussion), which can be evoked by language, but which are empirically hard to define or observe. on the other hand the practical needs of coreference resolution in nlp have often equated referentiality with potential for back-reference. in many partial implementations of ‘tricky’ coreference, this has amounted to including usually ignored expressions if they are referred back to, such as verbal markables for discourse deixis (‘she visited... the visit’), or limited markup of coordinations when they are mentioned again with a plural pronoun (‘kim and yun ... they’). 43 zeldes 1. indefinite mentions are often neither generic nor underspecified or abstract (“[participants] comprised [15 women and 10 men]”) 2. it is not immediately obvious that we should not want coreference information even for those mentions which are generic, abstract, etc. (e.g. “program trading”), especially given that we do resolve pronouns referring back to them 3. in context it is often hard to know whether pronominal mentions are generic, and even when they are, they can still form multiple distinct clusters 4. many languages do not have widespread articles which could be used to identify generic mentions, even if we agreed that all indefinites should be considered generic to illustrate this, we can consider the following examples from the gum corpus (zeldes, 2017), a freely available english coreference data set with texts from 12 written and spoken text types, which does not exclude indefinite anaphors:4 (3) marbles is the first social media star to have [a wax figure]i displayed in madame tussauds ... in 2015 , marbles unveiled [a wax figure of herself]i at madame tussauds (gum_bio_marbles) (4) i have [a mini-mmpi]i ... i have [a chart that i’ll go through]i (gum_interview_dungeon) in (3), the wax figure is a specific, unique physical object, mentioned twice as an indefinite np in the same document. in (4), the ‘chart’ is referring in context to the exact same entity as the mini-mmpi (a ‘minnesota multiphasic personality inventory’) – without the coreference information, a tool or researcher accessing the data automatically would not know that the mmpi was a chart, and vice versa, despite such information sharing being an obvious application of coreference resolution.5 generic pronouns too can often serve different purposes and belong to different, meaningful clusters: (5) [you]1 feel like [you]1’re prepared , [you]1’re in a , [you]2 know , in a relationship ... [you]2 know (gum_vlog_pregnant) in (5), which discusses feeling prepared for a pregnancy, all instances of ‘you’ are arguably generic, but they can be grouped into two sets: the first set refers to some person who might be considering a pregnancy (thereby implicitly ruling out those who cannot become pregnant), while the second set is used exclusively as a pragmatic discourse marker (östman, 1981) to refer to a generic audience and solicit presupposed agreement. it is therefore clear that in context, the person who may be prepared is the same person filling the subject role for ‘feel’, while ‘you’ in ‘you know’ can refer to anyone who may be listening, which forms a different set of individuals (the original utterance comes from a vlog, and the audience is therefore underspecified even for the speaker). finally, the issue of cross-linguistic applicability surfaces already in the on guidelines for chinese, which cannot refer to indefinite articles as a criterion for identifying genericity, since chinese 4see the corpus website (https://corpling.uis.georgetown.edu/gum) for example sources, and the appendix for longer examples. 5an anonymous reviewer has cautioned that this should not imply that disjoint indefinite mentions should corefer, as in ‘i always owned [a dog]generic. for my 6th birthday, i got [a dog]x. when it died we got [a dog]y from the animal sanctuary.’ i fully agree that there are three distinct ‘dog’ entities in this case, where x and y are subsets of the generic ‘dog’, and therefore not coreferring. 44 can we fix the scope for coreference? does not use articles as in english.6 the guidelines instead refer to specific examples:“死刑 (capital punishment),世界 (world – as distinct from “the world” meaning the planet earth), 社会 (society) are considered generic nouns” (bbn technologies, 2008, 8). the on chinese corpus itself, however, is not consistent about what cases are excluded as generic, and even these specified nouns are regularly annotated as coreferring in contexts which would not be considered definite or specific in english. for ‘society’, these include 44 coreferring instances, as in (6). (6) 我们能不能发展的快一些、好一些,实现经济快速发展和 [社会]i 全面进步,并且保 持 [社会]i稳定,十分重要 (on, cnr_0016) wǒmen néng bùnéng fāzhǎn de kuài yīxiē, hǎo yīxiē, shíxiàn jīngjì kuàisù fāzhǎn hé [shèhuì] quánmiàn jìnbù, bìngqiě bǎochí [shèhuì] wěndìng, shí fèn zhòngyào whether our development can progress faster and better to make it possible for the economy to grow quickly and for [society]i to make progress across all metrics as well as to maintain the stability of [society]i is very crucial rather than concluding that these cases are annotation errors, i would like to argue below in line with previous work showing the salience and coreferring potential of a range of indefinite phrases (see e.g. kunz 2009, 295–300) that it is reasonable to annotate them, just as the recurring annotations suggest that annotators regarded them as coreferring.7 2.2 compound modifiers compound modifiers (or noun-modifying nouns) in english, as in many languages with similar ‘bare’ adnominal modification, often have low referential content, and are known to resist pronominalization (i.e. they are ‘anaphoric islands’, postal 1969), making them attractive to exclude categorically. for example, as a compound nominal premodifier, it is difficult to imagine pronominal reference to ‘animals’ introduced by a compound ‘animal hunters’ in (7). because of this often limited referential potential, and the desire to avoid disagreements, a conservative annotation scheme may prioritize simplicity and consistency by prohibiting compound modifiers from entering into coreference relations. at the same time, repeated mention of proper nouns in the same position seems normal, and not indefinite or non-referential, as in (8): (7) *[animal]i hunters tend to like [them]i. (postal, 1969, 230) (8) the [hong kong]i government’s jurisdiction is the [hong kong]i special administrative region on makes an exception for proper noun modifiers (bbn technologies, 2007, 3), but excludes common noun cases like (9), which are left unannotated despite subsequent definite reference. 6the same is true of many languages; for example, slavic language coreference datasets have included generic nps, which cannot be distinguished based on articles (see e.g. nedoluzhko et al. 2009 for czech). 7one reviewer pointed out that we may still want to use articles in guidelines for languages that do distinguish them, as this can raise inter-annotator agreement, which may be the more important factor for some applications. i certainly do not mean that any reference to articles in guidelines for such languages should be forbidden; but conversely, ruling out ‘all indefinite anaphors’ in coreference datasets a priori, based on form alone, will not produce comparable data across languages, and as i hope this section shows, this would make us lose important information for english as well. 45 zeldes (9) small investors seem to be adapting to greater [stock market]i volatility . . . glenn britta . . . is “factoring” [the market’s]i volatility “into investment decisions.” (on, wsj_0121) however it is by now well known that compound modifiers are not strict islands even to pronominal anaphora (ward et al., 1993), and corpus data provides plentiful counter examples, leading to many of the same issues as those involved with generics (or more generally, indefinites): noun modifiers can be specific and can be referred back to with definite expressions as in (9), and even when they are truly generic, we may want to know about their coreference. in (10), the indication that cinnamon is a spice (and not a brand, a person’s name, or something else) is made explicit via coreference, and it is not clear to me why we should exclude such cases from the target gold standard for annotation and nlp tools. (10) [cinnamon]i basil really does smell like [the sweet spice]i (gum_whow_basil) as before, the restriction on compound modifiers in english is problematic from a cross-linguistic perspective, and in fact, construct-state modifiers in the on arabic coreference annotations, which are very similar to english compound modifiers (e.g. modifiers resist bearing articles, the translation of (7) would be ungrammatical), are not restricted in the same way, as seen in (11) with a semantically specific modifier: (11) �ðp  aj. � � é � kc � jë . . . � éò» am× . . . èaòëb@ � éòî � dk. �ðp  aj. � � éò» am× (on, ann_0006) mu ˙ hākamatu ˙ dubbātin rūsi bituhmati l-ihmāli... mu ˙ hākamatun... li¯ talā ¯ tatin ˙ dubbātin rūsi [russian officer]i trial on charges of negligence. . . a trial. . . for [three russian officers]i here the russian officers in the ‘russian officer trial’ are exactly the same set of ‘three officers’. such examples with specific indefinite nps create the most obvious motivation for including modifier nouns, but the same construction and coreference appear equally with non-specific/generic referents. for parallel corpora and applications of coreference resolution in machine translation too, we can note that what is expressed by such modifiers in one language may be expressed by a full np in another, meaning that excluding them will lead to more cross-linguistic divergence (see novák and nedoluzhko 2015, 31 for example for english and czech). the inconsistency between the annotations of modifier nouns in the three on languages is thus not only odd from a theoretical perspective; it can cause problems from a practical one as well if we try to develop multilingual applications relying on automatic coreference resolution, only to find out that even systems trained on on itself behave fundamentally differently for different languages. 2.3 nesting nested coreference, i.e. a mention with a coreferring mention within its own span, may seem odd, but is in fact rather frequent, and attested in this very sentence. on does allow nested coreference, but excludes cases without subsequent mention, leading to another difficult learning task for systems (the non-markable status of such cases hinges on later sentences), and unresolved pronouns in text, as in (12), which remains unannotated in ontonotes, presumably because outside of the larger phrase’s boundaries, it appears to be a singleton.8 8in fact, some systems have relied on this constraint, sometimes called ‘i-within-i’ due to the co-index i, to rule out certain kinds of match candidates, see lee et al. (2013, 896). 46 can we fix the scope for coreference? (12) [an elusive sheep with a star on [its]i back]i (on, wsj_0037) a further complication in on concerns nested dates, in which only the top-level expression is considered referential. for example, if a year is mentioned by itself, it may corefer with other mentions, but if it is part of a date, then it may not (again, the example is unannotated in on): (13) what are the opportunities for new developments in the wake of the [1999]i handover? on december 19, [1999]i, the eve of macau’s transfer of sovereignty... (on, ectb_1001) in (13), we are unable to discover that the year 1999 in the first sentence is the same year 1999 as in the second, since annotation is ruled out by the on form-based guideline on nested dates. some readers may think this is not so bad, since we can surely analyze both dates and discover that the year is the same; but i would argue we need to keep in mind that date detection is not trivial (1999 could be a price, or something else), and that these two cases should co-refer based on a semantic criterion (they refer to the same thing in the world). moreover, for system development it is problematic to create examples where algorithms will learn that date expressions are atomic, since we can have overt underspecified anaphoric expressions targeting subspans of dates: (14) between march 18, [2005]i and may 7 of [the same year]i (on, a2e_0016) example (14) actually is annotated in on, in contradiction with the guidelines, perhaps because of how counterintuitive their effect would be in this case and others like it (this example is by no means unique in on, and others are left unannotated as intended). without coreference in (14), we have no chance of discovering that “2005” is “the same year” in which “may 7” takes place. such information can be crucial for identifying dct (document creation time, ray et al. 2018) and for event timeline extraction (e.g. for clinical events, nikfarjam et al. 2013), among other applications. however the main point from my perspective is that this behavior is unexpected from a semantic point of view, and is narrowly tied to specific forms, not meanings. 2.4 predication judging by discussions with some of my colleagues, perhaps the most controversial opinion i can express in this paper is that the omission of predication from mainstream coreference resolution datasets and systems is a mistake.9 following an early period in the development of coreference datasets (notably ace, doddington et al. 2004) in which most nominal predicates were considered to corefer to their subjects (i.e. “kim is a teacher”⇒ “kim” coref←−−→ “a teacher”), criticism expressed for example by van deemter and kibble (2000) and zaenen (2006) led to on rejecting coreference in predication altogether. in this section i would like to suggest not only that on’s guidelines are an overreaction to the original criticism, but also that omission is neither a sufficient nor a necessary remedy for the problems that led to its rejection, an outcome left open by van deemter and kibble themselves. the core of the problem with predication has been expressed as “change over time” (hirschman and chinchor 1998, 11, van deemter and kibble 2000, 632), as in (15) (ibid.), but i would argue 9this is not to say that there are not several datasets which mark up coreference for predicates (see below); however since most kinds of predication are not covered by ontonotes, mainstream papers, systems and the conll shared tasks have not included these cases. for examples of how i think it should be handled, see the appendix, as well as section 4.2 47 zeldes that it is more broadly a problem of “change of scope”, since modal scopes create the same type of issue, as in (16) (a made-up example).10 (15) [henry higgins, who was formerly [sales director of sudsy soaps]i]i, became [president of dreamy detergents]i (16) if [beyoncé]i were [the queen of england]i, [she]i would.... in (15), henry higgins is not simultaneously sales director at one company and president of another, raising doubts as to the value of a coreference chain including all of those nps. likewise in (16), the pronoun “she” arguably corefers more to the queen of england (in a world where beyoncé is the queen), but not to the first mention of beyoncé. in fact, this case verges on examples of non-referential bound anaphora, also discussed by van deemter and kibble, but which cannot be discussed here for space reasons. so should we rule out predication due to such problems? after careful consideration of different opinions in the reviews of this paper, i believe the best answer is ‘yes, predication is different from other types of coreference’, but also ‘no, we cannot ignore it if we want to get the full picture’. my reluctance to ignore predication is based on a double dissociation: not all cases of predication raise this problem, and the problem can arise without predication. the most obvious case for including predication is in identificational predicates, which in english usually involves a definite predicate, but in other languages (e.g. japanese) can often be indistinguishable from other predications. compare these cases, marked up according to on guidelines, which include naming predicates as copular (bbn technologies, 2007, 27): (17) [elizabeth ii]i is the queen of england. [she]i ... (18) the queen of england is [elizabeth ii]i. [she]i ... (19) she was crowned [elizabeth ii]i in 1953. [she]i ... although they forbid annotating subject and predicate as a coreferent pair, on guidelines do specify a hierarchy for determining which of the two should be taken as the antecedent for subsequent mentions, ranking names above pronouns, and pronouns above definite nps. these kinds of configurations therefore appear in the official conll coreference shared task dataset and consequently model the output that contemporary coreference resolution systems attempt to produce. i will leave it to readers to imagine scenarios in which the missing member of the chain will lead to loss of information,11 but suffice it to say that including them would not lead to the scope problems above; in fact, some corpora already separate indefinite predication from regular coreference such as arrau (poesio and artstein, 2008) and gum, which treat identificational predication as regular coreference, and either label predicative markables whenever semantically applicable (arrau), or maintain a special coreference relation subtype for non-identificational cases (gum, see appendix). for the other side of the dissociation, we can consider chains of simple, definite nps, in which the scope problem does arise, as in (20). 10massimo poesio has pointed out that the problem is even broader and ultimately stems from the multiple types of entity expressions in names or variables, truth value expressions and quantification in the sense of montagovian type shifting, see partee (1986). in this discussion i will limit myself to the prominent and frequent issues in narrowly defined copula predication, but indeed similar problems arise for negated nps, relational indefinites and other ‘tricky nps’ (see landman 2004). 11see also lee et al. (2013, 902) for an example. 48 can we fix the scope for coreference? (20) [a fresh major in the swedish army], in 1812 [gordon] went to war . . . in 1875 [the now general in the russian army] was ready to pursue [his] ultimate achievement. . . [gordon] is buried in. . . (adapted from gum_bio_gordon) we can introduce temporally inconsistent mentions (the buried man is not a swedish major, the swedish major is not a russian general), and similarly modality inconsistencies or ambiguous mentions, etc. although predication, naming, objects of ‘as’ and other constructions are likely locations for this in english, they are not the problem in themselves (so not a necessary condition), and they do not guarantee that there is a problem (so not sufficient). i agree such problems need to be addressed (see section 4), but with few exceptions, the solution in most major datasets so far has been to largely ignore them, rather than to mark them up as a special case of anaphora (similarly to bridging, discourse deixis or split antecedents). 3. isn’t this someone else’s job? before suggesting possible solutions, i would like to briefly outline why two potential objections to this paper’s proposals do not offer alternative solutions to the problems raised here. although there may be other suggestions on how to deal with the phenomena not covered in ontonotes, most probably follow one of two ways of ‘punting’ the issues: either to syntax or semantics. 3.1 can’t we get all of this from syntax? this line of reasoning is often raised in defence of coreference datasets which exclude predication, but which do have gold treebanking information: since syntax trees represent predication, naming constructions, nesting, and other structures more or less unambiguously, can’t we just leave them out of the coreference annotation proper and recover them from the trees? the answer is no. i have yet to see a convincing case where this has been done, and there are good reasons why we should not think that it is possible. contrast the following pairs, adapted from gum: (21) a. [he] would be a libertarian today (no coref, since this is hypothetical: he is *not* a libertarian) b. [the principles governing an f-e translation]i would then be: [reproduction of grammatical units; consistency in word usage; and meanings in terms of the source]i (coref, since in fact, those are the principles of f-e translation) (22) a. [this coffee table] is glass (not predicative coref, only specifies the substance that most/part of the table is made of) b. [this ice here]i of course is [water]i (predicative coref, part of a chemistry demonstration in which the speaker literally identifies an ice cube as being the same water in a solidified state) (23) [he]i was not the leaf-collecting doctor, but [an altogether strange man, with silver eyebrows in his smooth face and long fine-knuckled hands]i (gum_fiction_lunre) in (21)–(22) the same syntactic construction yields coreferring expressions in one case and no coreference in the other (in (21), identity coreference, in (22), predication coreference, labeled as such in gum). in (23), only part of a coordinate predicate np is coreferring, meaning that extracting 49 zeldes the correct span using only syntax is non-trivial, and of course the phenomena in these examples can co-occur (imagine processing modal coordinate substance predications correctly!). syntactic approaches also assume a well-formed syntax tree uttered contiguously by a single speaker, which in some cases is not a given. it is also extremely difficult to know whether fairly mundane spatio-temporal predicate nps are coreferring, regardless of definiteness: (24) this town is 35 minutes from the harbor (no coref, purely spatio-temporal ‘is’, since town6=35 minutes) (25) but christmas is still the whole winter to wait (no coref, but most syntax trees would show ‘winter’ as a definite predicate np) similarly for np-internal or nested coreference resolution, we could easily have ambiguities in ontonotes-style data. consider this minor modification of the ‘sheep’ example above (26) a sheepi in a coatj with a star on itsi/j back the data in (26) means that nested resolution is needed for disambiguation and cannot be extracted from syntax automatically, just as it is for predication. finally for compound modifiers, although many coreferential cases include verbatim repetition of an entire compound, which can perhaps be recognized from word forms alone, some cases do not, with examples like (10) above forming the most striking cases. there can be no syntactic solution if we want to recover such relations. to be clear, this is not to say that morpho-syntactic criteria can never be used in coreference annotation guidelines in any way: the reason why projects have used syntactic constructions diagnostically is because they can be very helpful in differentiating subtly different constructions and as a result help raise inter-annotator agreement. the point is rather that use of such criteria should be in service of semantic distinctions, which should be our actual object of interest: syntax and morphology should not be allowed to rule out things which we are certain do co-refer (for example morphologically incongruent singular and plural nps), and they should not be used to admit things that do not (for example appositions which do not corefer). and if something is supposed to be recoverable from syntax, this should not move us to exclude it from coreference annotation – otherwise we are effectively mandating that any coreference resolution system will also have to tackle syntactic parsing. 3.2 how about semantics? another line of reasoning is that for ‘quirky’ coreference cases, including predication but also various kinds of distributive semantics (which space prevents me from discussing), semantic annotation should be in charge of annotating predicate structure in a way that disambiguates the issues in section 3.1. the problem here is that even the most elaborate semantic annotation formalisms available right now do not address the main problems, since predication is largely subsumed in lexical semantics. in propbank (palmer et al., 2005), one of the simpler but most widespread types of semantic annotation, predicates are simply argument structure graphs with word sense disambiguation, but copula ‘be’ is simply annotated as be.01, as in the following example from ontonotes: (27) a. the most special is rice (on, cctv_0000) b. prop: be.01, arg1: ‘the most special’, arg2: ‘rice’ 50 can we fix the scope for coreference? it is therefore impossible to know whether this is a coreferential case or not (it is, specifically a non-identification predicative coreference: in this case a rice plant is being discussed). more complex semantic analysis is undertaken in the less widespread but highly detailed formalism of abstract meaning representation (amr, banarescu et al. 2013), which has facilities to indicate, for example, predication negation, possibly solving the problem in (23); but amr is not aligned to words, and it is impossible to use it for the purpose of coreference resolution, as the following example from the little prince corpus (ibid.) illustrates: (28) a. it is my fault that you have not known it all the while. b. (f / fault-01 :arg1 (i3 / i) :arg2 (k / know-01 :polarity :arg0 (y / you) :arg1 (i2 / it) :time (w / while-away-01 :duration (a / all)))) the amr analysis captures the argument structure of ‘fault’ perfectly (arguments: the spearker and the ‘knowing’ predicate), but leaves no indication of the original expletive subject ‘it’, its relationship with the extraposed clause ‘that you have not known...’, etc. because of amr’s (intentional) distance from word forms in the original sentence, it does not fulfill or replace the task of textanchored coreference resolution – amr simply has no obligation to include nodes for each surface pronoun or np. similarly, neither of these formalisms analyzes the properties of discourse referents in terms of genericity, or contemporaneous temporal or modal scope. 4. fixing the problems is it impossible to envision a semantic or syntactic formalism that can represent the facts discussed here? certainly not: we could just extend semantic or syntactic annotation to distinguish such cases. but here i would like to argue that covering these cases is part of the job of coreference resolution, and realistically, if coreference resolution does not tackle them, they are unlikely to be covered by other annotation efforts in the current landscape. 4.1 unexiling tricky coreference although it should be clear by now that i would like to see a lot more things annotated in coreference resolution datasets, i would like to emphasize that i am not saying that all of these types of coreference are the same. addressing them does not necessarily mean that we need to lose the distinction between predication, pronominal anaphora and other types of coreference. ontonotes itself distinguishes several types of apposition from identity coreference, and several other corpora distinguish not only appositions, but also pronominal anaphora from lexical identity (e.g. gum) or non-identificational predication (e.g. both gum and arrau for english, or cess-ece for spanish, recasens et al. 2007, which are all useful examples of explicitly handling, rather than ignoring predication), in addition to harder phenomena such as discourse deixis or bridging anaphora. critics may point out that including trickier cases will inevitably lead to lower agreement, but i would answer that doing so should not decrease agreement on ‘easy cases’, and that in instances 51 zeldes expression type instances per 1k tokens % of total pron. anaphora 59.84 44.98% cataphora 1.39 1.05% nominal predicate 4.94 3.71% compound mod. 5.79 4.36% split antecedent 1.22 0.92% apposition 3.81 2.86% coref. name 19.84 14.92% other indef. np 14.53 10.92% other coref 21.67 16.29% total 133.02 100% table 1: coreference link type distribution in gum of disagreement, most often each of the conflicting analyses corresponds to a valid reading of an ambiguous context (as demonstrated by poesio and artstein 2005). human annotators also tend to disregard or not notice a relation more often than connecting non-coreferring entities (jiang et al., 2013; zeldes, 2017), meaning data would still have relatively high precision (for gum, which includes the phenomena discussed here, student annotators achieved precision, recall and f1 of 0.918/0.811/0.861 compared to the adjudicated gold standard, when minimal mention span matching is used, zeldes 2017, 596). with this in mind, including these phenomena and labeling them as such, as is done in arrau or gum, seems in essence already close to the ideal solution, with a few caveats and additions. the first is that the default coreference resolution task should, in my opinion, cover all types of cluster-based12 coreference discussed here, including compound modifiers, indefinites and predication. as shown in table 1, which tallies types of anaphoric expressions in gum, all of these are extremely frequent, non-marginal phenomena, which often overlap with clear cases of identity coreference (e.g. definite copula predicates, semantically specific compound modifiers). taken together, indefinite nps, compound modifiers and predicates form around 19% of all previously mentioned nps, occurring about 26 times per 1k tokens. for comparison, this is more than the remaining definite common nps (21.67 times) or subsequently mentioned proper names (19.84). nested dates (not shown in the table) are also ubiquitous in gum data, with 13.8% of temporal nps being nested in other temporal nps (including singletons) and 38% of coreferring year expressions in the corpus being nested in a date entity span, which would exclude them in on. removing such mentions from the target standard for mainstream nlp leads to counterintuitive results and ultimately creates an incongruity between systems’ performance in shared tasks and their coverage for real world applications, which often assume the premise of being able to collect all information about an entity in a text (say, some year) when coreference resolution is ‘correct’. this position, if accepted, raises the question of what we should do about henry higgins, president of dreamy detergents? haven’t we been here before, and didn’t the criticism of the ace dataset teach us that ‘change over time’ makes text-anchored coreference futile? 12i use this term to set aside other types of anaphora, such as bridging, sense anaphora (e.g. ‘one’ anaphora), etc., not because they are not important, but because we already have much work to do on many types of identity coreference and closely related concepts. 52 can we fix the scope for coreference? 4.2 addressing scope the ‘change over time’, or better, ‘change of scope’ problem raised by van deemter and kibble (2000) remains the biggest challenge to a more inclusive view of coreference, and i have answers for it on several levels. on the practical level, scope problems are both not very common and ubiquitous. they are not very common since true hard cases, such as texts covering an entity changing substantially vis-a-vis the relevance of referring expressions over time, are only a subset of texts, and even in those texts, any problematic entity often co-exists with scores of other predicate nps which do not raise such problems. refusing to include those hundreds of identificational predicate nps is throwing out the baby with the bath water. at the same time, the problem is ubiquitous (cf. recasens et al. 2012 on ‘near-identity coreference’). any long text will involve changes to participants, which, even if they are not stated, mean that not all predicates stated of some entity actually apply simultaneously. if we omit predicates from the coreference annotation and then tell the story of henry higgins, he does not remain the same henry higgins whether we include the information about his various jobs or not, or whether we express that information using nps or not – we could just say he was ‘fired’ and alter the properties of the entity (i.e. his job) without interacting with coreference. to then say that we will use syntax or semantics to describe henry’s biography does not absolve us of the need to address scope, since the same inconsistencies will result in a semantic annotation of the text. but ignoring ‘difficult’ predicate nps as part of coreference annotation will lead to many gaps in what practitioners using coreference resolution might expect or could harness. moving over to the theoretical perspective, i think there is enough merit in representing scope that we as a community of researchers interested in coreference should consider what is the right thing to do. the fact that we recognize the problem with henry higgins suggests that we do know that both the ‘sales director’ henry and the ‘president’ henry are the same person: just not at the same time. it could be tempting to simply allow a co-indexed scope attribute in our coreference annotation schemes, as in (29) for the beyoncé example. (29) if beyoncé were the queen of england , she would conduct the annual swan upping. this notation would accurately show that the pronoun ‘she’ refers to its intuitive antecedent entity in the document context, ‘beyoncé’, while also indicating that this specific ‘she’ is referential in the world in which beyoncé is the queen, and co-refers to that queen. we could also add a coreference type indicating that we are looking at identity-predication, etc. for modal cases, adding a scope id may be enough, but for temporal ones, we could even consider annotating intervals where these are known (or place-holders where they are not; see the red corpus for a similar strategy in event annotation, o’gorman et al. 2016): (30) henry higgings, who was formerly sales director at sudsy sopes, became president of dreamy detergents omitting scope ids and intervals would mean that these are the ‘general’ scope conditions of the document, and that further information is not known. 53 zeldes alternatively, given that there have been independent efforts to annotate modal (hendrickx et al., 2012; nissim et al., 2013; rubinstein et al., 2013) and temporal scope (pustejovsky and stubbs, 2011; styler iv et al., 2014; pustejovsky et al., 2019) for semantics, it may make sense to encode scope independently of coreference annotation, as a span level property of sentences, or vps or any text span, which interacts with coreference, since many other types of annotation could be affected by modal and temporal scopes. for example, the hypothetical ‘conduct the swan upping’ (an annual census of the queen’s swans) in (29) is also restricted to the modal scope in which beyoncé is the queen, and event annotations of this predicate could take this into account. i do not want to pretend that annotating scope for coreference is simple, or that we are in a position to easily add it to large benchmark datasets such as ontonotes: there are surely many complications, even without which this would mean much work, and a detailed examination of how coreference interacts with existing scope annotation schemes is certainly in order. but assuming we can take the position that (serious) scope problems are fairly rare, then i think the right starting point before moving on to an adequate representation for these cases is to include the phenomena missing in ontonotes as an ideal target, and offer scope as a topic for further advanced research on coreference, similarly to bridging anaphora or split antecedents, whose absence in benchmarks like ontonotes is well understood, but does not interfere with progress on the most common types of coreferentiality in corpora. 5. conclusion if this opinion piece falls short of changing any coreference annotation practices, then i hope it at least serves one purpose: to make researchers aware of the limitations of on-style coreference, and by extension most nlp tools for coreference resolution. my experience has been that these limitations are not well understood by proficient nlp practitioners who use tools such as vanilla e2e coref (lee et al., 2017), allennlp (gardner et al., 2018) or spacy’s huggingface (clark and manning, 2016) implementations, and this often leads to surprises. conscious choices for or against including some of the trickier phenomena will inevitably involve factors relating to trade-offs of linguistic adequacy versus reliability and application domains, but informed decisions are better than ones made spontaneously or based on inertia. a more ambitious goal for this paper is that readers involved in the production and (re)annotation of coreference datasets will consider following corpora such as gum and arrau in taking a broader view of the phenomenon, and use subtypes to carve out areas which some systems might want to leave out or label in special ways for downstream applications – the appendix offers some worked examples from the gum corpus as a starting point. including more coreference phenomena does not mean that we have to be naive about the potentially lower level of agreement on complex cases (e.g. generics, predication); but removing such cases from output that contains them is generally much easier than trying to add them to data in which they are not included in any way, and if modifying data along different lines benefits specific subsets of downstream tasks, then the flexibility offered by distinguishing but including tricky phenomena is all the more valuable. this has also been one of the main lessons i have learned from working on multilayer corpora (see zeldes 2018): no, we cannot get everything from syntax as pointed out above, but syntax is great for finding specific phenomena or manipulating data into different schemes (e.g. removing compound modifiers, changing how appositions are handled). in one paper, this has even allowed transforming gum, 54 can we fix the scope for coreference? the dataset most closely matching this proposal, to follow the ontonotes guidelines (the resulting dataset is called ‘ontogum’, see zhu et al. 2021 for details and evaluation). finally, i think that as corpus resources converge across languages on uniform standards, as evidenced in the universal dependencies project for treebanking, pos tagging and morphological annotation (de marneffe et al., 2021), or multilingual initiatives in other areas of discourse (e.g. the disrpt shared tasks on discourse relations, zeldes et al. 2019, 2021) we will need to make coreference more consistent across datasets. this will mean that we have to minimize the use of guidelines based on language-specific forms, such as definiteness or specific constructions. instead, i have argued here that coreference should primarily be decided on semantic grounds, which seem likelier to transfer between languages. as an added motivation, i offer the consideration that cross-linguistically and semantically grounded coreference annotations are likely to work better for applications involving multiple corpora, such as multilingual and cross-document coreference, as well as entity linking (see lin and zeldes 2021 for more on merging entity linking with crossdocument coreference). the problems discussed in this paper are in my opinion not intrinsically ones of copula sentences or modal verbs or other constructions, but rather issues at the discourse-semantics interface, where meanings from individual sentences coalesce to form narratives in which entities change. this suggests that tackling the issues in a cross-linguistically applicable way will ultimately require the development of standards for handling spatiotemporal, modal and other types of scope independently of language or text type. references laura banarescu, claire bonial, shu cai, madalina georgescu, kira griffitt, ulf hermjakob, kevin knight, philipp koehn, martha palmer, and nathan schneider. abstract meaning representation for sembanking. in proceedings of the 7th linguistic annotation workshop and interoperability with discourse, pages 178–186, sofia, bulgaria, 2013. bbn technologies. ontonotes english co-reference guidelines version 7.0. technical report, bbn technologies, 2007. bbn technologies. ontonotes chinese co-reference guidelines version 4.0. technical report, bbn technologies, 2008. archna bhatia, chu-cheng lin, nathan schneider, yulia tsvetkov, fatima talib al-raisi, laleh roostapour, jordan bender, abhimanu kumar, lori levin, mandy simons, and chris dyer. automatic classification of communicative functions of definiteness. in proceedings of coling 2014, pages 1059–1070, dublin, ireland, 2014. kevin clark and christopher d. manning. deep reinforcement learning for mention-ranking coreference models. in proceedings of emnlp 2016, pages 2256–2262, austin, tx, 2016. marie-catherine de marneffe, christopher d. manning, joakim nivre, and daniel zeman. universal dependencies. computational linguistics, 47(2):255–308, 2021. george r. doddington, alexis mitchell, mark a. przybocki, lance a. ramshaw, stephanie strassel, and ralph m. weischedel. the automatic content extraction (ace) program. tasks, 55 zeldes data, and evaluation. in proceedings of the 4th international conference on language resources and evaluation (lrec-2004), pages 837–840, 2004. john w. du bois, wallace l. chafe, charles meyer, sandra a. thompson, robert englebretson, and nii martey. santa barbara corpus of spoken american english, parts 1-4, 2000-2005. matt gardner, joel grus, mark neumann, oyvind tafjord, pradeep dasigi, nelson f. liu, matthew peters, michael schmitz, and luke s. zettlemoyer. allennlp: a deep semantic natural language processing platform. in proceedings of acl 2018, pages 1–6, melbourne, 2018. iris hendrickx, amália mendes, and silvia mencarelli. modality in text: a proposal for corpus annotation. in proceedings of lrec 2012, pages 1805–1812, istanbul, turkey, 2012. lynette hirschman and nancy chinchor. appendix f: muc-7 coreference task definition (version 3.0). in seventh message understanding conference (muc-7), 1998. eduard hovy, mitchell marcus, martha palmer, lance ramshaw, and ralph weischedel. ontonotes: the 90% solution. in proceedings of naacl 2006: short papers, pages 57–60, new york, 2006. lili jiang, yafang wang, johannes hoffart, and gerhard weikum. crowdsourced entity markup. in proceedings of crowdsourcing the semantic web, pages 59–68, sydney, 2013. mandar joshi, omer levy, luke zettlemoyer, and daniel weld. bert for coreference resolution: baselines and analysis. in proceedings of emnlp-ijcnlp 2019, pages 5803–5808, hong kong, china, 2019. manfred krifka. basic notions of information structure. acta linguistica hungarica, 55:243–276, 2008. kerstin kunz. variation in english and german nominal coreferece. a study of political essays. peter lang, frankfurt, 2009. fred landman. indefinites and the type of sets. blackwell, malden, ma, 2004. heeyoung lee, angel chang, yves peirsman, nathanael chambers, mihai surdeanu, and dan jurafsky. deterministic coreference resolution based on entity-centric, precision-ranked rules. computational linguistics, 39(4):885–916, 2013. kenton lee, luheng he, mike lewis, and luke zettlemoyer. end-to-end neural coreference resolution. in proceedings of emnlp 2017, pages 188–197, copenhagen, 2017. jessica lin and amir zeldes. wikigum: exhaustive entity linking for wikification in 12 genres. in proceedings of the joint 15th linguistic annotation workshop (law) and 3rd designing meaning representations (dmr) workshop, pages 170–175, punta cana, dominican republic, 2021. nafise sadat moosavi and michael strube. which coreference evaluation metric do you trust? a proposal for a link-based entity aware metric. in proceedings of acl 2016, pages 632–642, berlin, 2016. 56 can we fix the scope for coreference? anna nedoluzhko, jiří mírovský, radek ocelák, and jiří pergler. extended coreferential relations and bridging anaphora in the prague dependency treebank. in proceedings of the 7th discourse anaphora and anaphor resolution colloquium (daarc 2009), pages 1–16, goa, india, 2009. anna nedoluzhko, michal novák, martin popel, zdeněk žabokrtský, amir zeldes, and daniel zeman. corefud 1.0: coreference meets universal dependencies. in proceedings of lrec 2022, marseille, france, 2022. azadeh nikfarjam, ehsan emadzadeh, and graciela gonzalez. towards generating a patient’s timeline: extracting temporal relationships from clinical notes. biomedical informatics, 46:40–47, 2013. malvina nissim, paola pietrandrea, andrea sansò, and caterina mauri. cross-linguistic annotation of modality: a data-driven hierarchical model. in proceedings of the 9th joint iso acl sigsem workshop on interoperable semantic annotation, pages 7–14, potsdam, germany, 2013. michal novák and anna nedoluzhko. correspondences between czech and english coreferential expressions. discours, 16:3–41, 2015. jan-ola östman. you know: a discourse-functional approach. pragmatics & beyond, ii:7. john benjamins, amsterdam, 1981. tim o’gorman, kristin wright-bettner, and martha palmer. richer event description: integrating event coreference with temporal, causal and bridging annotation. in proceedings of the 2nd workshop on computing news storylines (cns 2016), pages 47–56, austin, tx, 2016. martha palmer, daniel gildea, and paul kingsbury. the proposition bank: a corpus annotated with semantic roles. computational linguistics, 31(1):71–106, 2005. barbara h. partee. noun phrase interpretation and type-shifting principles. in jeroen groenendijk, martin stokhof, and dick de jongh, editors, studies in discourse representation theory and the theory of generalized quantifiers, pages 115–143. foris, dordrecht, 1986. massimo poesio and ron artstein. the reliability of anaphoric annotation, reconsidered: taking ambiguity into account. in proceedings of the workshop on frontiers in corpus annotation ii: pie in the sky, pages 76–83, stroudsburg, pa, 2005. acl. massimo poesio and ron artstein. anaphoric annotation in the arrau corpus. in nicoletta calzolari, khalid choukri, bente maegaard, joseph mariani, jan odjik, stelios piperidis, and daniel tapias, editors, proceedings of lrec 2008, pages 1170–1174, marrakesh, morocco, 2008. massimo poesio, sameer pradhan, marta recasens, kepa rodriguez, and yannick versley. annotated corpora and annotation tools. in massimo poesio, roland stuckardt, and yannick versley, editors, anaphora resolution: algorithms, resources, and applications, theory and applications of natural language processing, pages 97–140. springer, heidelberg, 2016. paul m. postal. anaphoric islands. in robert i. binnick, editor, 5th regional meeting of the chicago linguistic society, volume 5, pages 205–239. chicago linguistic society, chicago, 1969. 57 zeldes sameer pradhan, lance ramshaw, mitchell marcus, martha palmer, ralph weischedel, and nianwen xue. conll-2011 shared task: modeling unrestricted coreference in ontonotes. in proceedings of the conference on computational natural language learning: shared task, pages 1–27, portland, or, 2011. sameer pradhan, xiaoqiang luo, marta recasens, eduard hovy, vincent ng, and michael strube. scoring coreference partitions of predicted mentions: a reference implementation. in proceedings of the 52nd annual meeting of the association for computational linguistics, pages 30–35, baltimore, md, 2014. james pustejovsky and amber stubbs. increasing informativeness in temporal annotation. in 5th linguistic annotation workshop (law v), pages 152–160, portland, or, 2011. james pustejovsky, ken lai, and nianwen xue. modeling quantification and scope in abstract meaning representations. in proceedings of the first international workshop on designing meaning representations (dmr), pages 28–33, florence, italy, 2019. swayambhu nath ray, shib sankar dasgupta, and partha talukdar. ad3: attentive deep document dater. in proceedings of emnlp 2018, pages 1871–1880, brussels, belgium, 2018. marta recasens, m. antònia martí, and mariona taulé. where anaphora and coreference meet. annotation in the spanish cess-ece corpus. in proc. ranlp 2007, borovets, bulgaria, 2007. marta recasens, m. antònia martí, and constantin orasan. annotating near-identity from coreference disagreements. in proceedings of lrec 2012, pages 165–172, istanbul, turkey, 2012. tanya reinhart. pragmatics and linguistics. an analysis of sentence topics. philosophica, 27:53–94, 1982. aynat rubinstein, hillary harner, elizabeth krawczyk, daniel simonson, graham katz, and paul portner. toward fine-grained annotation of modality in text. in proceedings of the workshop on annotation of modal meanings in natural language, pages 38–46, potsdam, germany, 2013. veselin stoyanov, nathan gilbert, claire cardie, and ellen riloff. conundrums in noun phrase coreference resolution: making sense of the state-of-the-art. in proceedings of acl/ijcnlp 2009, pages 656–664, suntec, singapore, 2009. william f. styler iv, steven bethard, sean finan, martha palmer, sameer pradhan, piet c de groen, brad erickson, timothy miller, chen lin, guergana savova, and james pustejovsky. temporal annotation in the clinical domain. tacl, 2:143–154, 2014. kees van deemter and rodger kibble. on coreferring: coreference in muc and related annotation schemes. computational linguistics, 26(4):629–637, 2000. gregory ward, richard sproat, and gail mckoon. a pragmatic analysis of so-called anaphoric islands. language, 63(3):439–474, 1993. ralph weischedel, sameer pradhan, lance ramshaw, jeff kaufman, michelle franchini, mohammed el-bachouti, nianwen xue, martha palmer, jena d. hwang, claire bonial, jinho choi, 58 can we fix the scope for coreference? aous mansouri, maha foster, abdel aati hawwary, mitchell marcus, ann taylor, craig greenberg, eduard hovy, robert belvin, and ann houston. ontonotes release 5.0. technical report, linguistic data consortium, philadelphia, 2012. wei wu, fei wang, arianna yuan, fei wu, and jiwei li. corefqa: coreference resolution as query-based span prediction. in proceedings of the 58th annual meeting of the association for computational linguistics, pages 6953–6963, online, 2020. annie zaenen. mark-up barking up the wrong tree. computational linguistics, 32(4):577–580, 2006. amir zeldes. the gum corpus: creating multilayer resources in the classroom. language resources and evaluation, 51(3):581–612, 2017. amir zeldes. multilayer corpus studies. routledge advances in corpus linguistics 22. routledge, london, 2018. amir zeldes and shuo zhang. when annotation schemes change rules help: a configurable approach to coreference resolution beyond ontonotes. in proceedings of the workshop on coreference resolution beyond ontonotes (corbon), pages 92–101, san diego, ca, 2016. amir zeldes, debopam das, erick galani maziero, juliano desiderato antonio, and mikel iruskieta. the disrpt 2019 shared task on elementary discourse unit segmentation and connective detection. in proceedings of discourse relation treebanking and parsing (disrpt 2019), pages 97–104, minneapolis, mn, 2019. amir zeldes, yang janet liu, mikel iruskieta, philippe muller, chloé braud, and sonia badene, editors. proceedings of the 2nd shared task on discourse relation parsing and treebanking (disrpt 2021), punta cana, dominican republic, 2021. rui zhang, cícero nogueira dos santos, michihiro yasunaga, bing xiang, and dragomir radev. neural coreference resolution with deep biaffine attention by joint mention detection and mention clustering. in proceedings of acl 2018, pages 102–107, melbourne, australia, 2018. shuo zhang and amir zeldes. gitdox: a linked version controlled online xml editor for manuscript transcription. in proceedings of flairs-30, pages 619–623, marco island, fl, 2017. yilun zhu, sameer pradhan, and amir zeldes. anatomy of ontogum—adapting gum to the ontonotes scheme to evaluate robustness of sota coreference algorithms. in proceedings of the fourth workshop on computational models of reference, anaphora and coreference (crac 2021), pages 141–149, punta cana, dominican republic, 2021. appendix a. worked examples this section gives detailed examples taken from different genres in the gum corpus, which includes two versions of coreference annotations: gum’s native, exhaustive coreference scheme, which follows the suggestions in this paper, and an automatically produced version of the same data following the ontonotes guidelines, a rendition of the corpus referred to as ontogum (see zhu et al. 2021 for details). markup for all examples in this appendix can be found at the gum corpus repository, in both the original and ontonotes scheme, at https://github.com/amir-zeldes/gum. 59 zeldes a.1 how to grow basil non-identificational predication (dashed) apposition (dotted) identity coreference (solid color) singleton (gray box) figure 1: gum-style coreference annotations (gum_whow_basil), visualized using gitdox (zhang and zeldes, 2017). this excerpt from a wikihow guide on how to grow basil (see the corpus website for source and licensing information on all texts) illustrates the proposal’s coverage for nested identity coreference and non-identificational predicative coreference, as well as singletons and entity type annotations found in gum. the more sparse ontogum representation illustrates the comparative loss of information based on the ontonotes scheme. beyond the singletons (in grey) and entity types (icons in top left corners of boxes) found in figure 1 (see the legend), some items to note include the common noun compound modifier antecedent ‘cinammon’, the predication between ‘most other varieties’ and ‘annuals...’, and the generic pronoun ‘you’ (in pink). additionally, we see nested coreference in ‘[african blue basil ( which has pretty blue veins on [its]i leaves)]i’. in figure 2 we see that the nested coreference set is not included according to the ontonotes scheme, since it is not mentioned a subsequent time outside the nesting mention, and is therefore considered a singleton (and singletons are generally omitted). similarly the mention of ‘cinammon’ is ignored, and predication for ‘annuals’ cannot be recovered trivially (see section 2.4). figure 2: ontonotes-style coreference annotations (gum_whow_basil). 60 can we fix the scope for coreference? a.2 spoken indefinites this example comes from a conversation (gum_conversation_lambada), originally from the ucsb corpus of spoken american english (du bois et al., 2000-2005), and included in gum’s conversation genre. a. b. miles: jamie: miles: harold: miles: miles: jamie: miles: harold: miles: figure 3: analyses of a conversation fragment according to coreference schemes based on gum (above) and ontonotes (below). referring expressions with a gold star are annotated as accessible in context. in figure 3a, mentions of the ‘random guy’ are all grouped together, regardless of definiteness or predication, since these are not criteria for coreference in the gum scheme, which follows the proposal in this paper. the ontonotes version ignores some mentions (if they are indefinites in chain-medial position) and oddly splits the chain into two sets, creating the impression that there are two random guys, in addition to the person being asked about by the second speaker (there are only two guys in context). gum’s mention annotation scheme, which marks up information status (with six subtypes), additionally flags referring expressions as accessible in context, including an initial deictic realization of ‘this guy’ and the situationally accessible ‘where you are’, which are marked by gold stars in the figure. a.3 datelines and headlines this example comes from the news genre (taken from wikinews) and illustrates two common problems in newswire data. in figure 4, ‘2015’ corefers with ‘this year’, and ‘this process’ corefers with both spans ‘a public process’ and the ‘process to consider...’ which is part of the headline. because the headline and first subsequent mention of the ‘process’ are both indefinite, the first mention is eliminated in the ontonotes version in figure 5. additionally, because the year in the dateline is nested in a date, ontonotes guidelines rule it out as an antecedent for ‘this year’, meaning that the latter becomes a singleton, and both mentions are excluded. 61 zeldes figure 4: gum-style coreference annotations (gum_news_flag). figure 5: ontonotes-style coreference annotations (gum_news_flag). 62 dialoguediscoursepaper_final_formatted dialogue and discourse 4(2) (2013) 1-33 doi: 10.5087/dad.2013.201 ©2013 jonathan t. morgan et al. submitted: 02/12; accepted: 12/12; published online: 04/13 are we there yet?: the development of a corpus annotated for social acts in multilingual online discourse jonathan t. morgana {jmo25, what, ebender, zliyi, gracheva, zachry} meghan oxleyb @uw.edu emily m. benderb liyi zhub varya grachevab mark zachrya adepartment of human centered design & engineering, box 352315 bdepartment of linguistics, box 354340 university of washington seattle, wa 98195 usa editors: stefanie dipper, heike zinsmeister, bonnie webber abstract we present the aawd and aacd corpora, a collection of discussions drawn from wikipedia talk pages and small group irc discussions in english, russian and mandarin. our datasets are annotated with labels capturing two kinds of social acts: alignment moves and authority claims. we describe these social acts, discuss our annotation process, highlight challenges we encountered and strategies we employed during annotation, and present some analyses of resulting the data set which illustrate the utility of our corpus and identify interactions among social acts and between participant status and social acts in online discourse. keywords: annotation, linguistics, discourse, computer-mediated communication 1 introduction the application of machine learning to natural language processing problems has been very effective at automatic annotation of morphosyntactic structure and, in at least some application contexts, extracting relational semantic information and general sentiment. the work described in this paper is motivated by the question of whether these successes can be extended to the automatic identification of the social acts of discourse participants. an important first step in developing automated annotation methods is the manual annotation of the target structures, which in turn requires the development of operationalizable definitions. to that end, we have defined two sets of social acts that can be annotated at the discourse turn level and carried out annotation of these acts in six corpora, representing three languages (english, mandarin, russian) and two genres (asynchronous online forum discussions and synchronous online chat).1 reliable identification of social goals, individual motivations and discursive strategies of conversational participants presents both theoretical and practical challenges. few theoretical frameworks systematically account for a meaningful and coherent set of the kinds of things discussion participants “do” with language as they interact across varied communication situations. when research from different fields addresses what people do with language, it often addresses a limited range of field-specific concerns that do not readily align with potentially 1 the social acts and their annotation within the english wikipedia corpus were described in bender et al. 2011, along with some preliminary analysis. this paper expands that description and analysis as well as extending both to multilingual and multigenre considerations. are we there yet? complementary ideas from seemingly related work in another field. across the work in different fields, there is little guidance in terms of defined categories that can be operationalized for annotation and thus allow for systematic investigation of the specific word, sentence, and turnlevel structural properties of the language used in the social acts. likewise, it is difficult to develop and apply annotation guidelines to reliably label such pragmatic-level phenomena because of their inherent variability and sensitivity to contingencies introduced by the medium, genre and language of the communicators. this paper describes the development of such a theoretically-motivated annotation effort, resulting in the authority and alignment in wikipedia discussions (aawd) corpus and the authority and alignment in chat discussions (aacd) corpus.2 these corpora contain online discussions in english, mandarin and russian annotated for two types of social acts: authority claims and positive/negative interpersonal alignment moves. the aawd corpus includes annotation of these social acts in asynchronous, language-based interactions among editors on wikipedia article discussion pages. the aacd corpus includes annotation of the social acts in synchronous, textual interactions among small groups of individuals engaged in collaborative planning through the medium of internet relay chat (irc). the annotations described here were produced with both engineering and scientific goals in mind. on the one hand, they can serve as training and test data for machine learning, that is, the automatic detection in text of the phenomena that we are annotating by hand (see for example marin et al. 2011). to the extent that these social acts (and others of a similar granularity) are deployed in various combinations towards larger social goals of conversational participants, the ability to detect them in naturally occurring discourse can be an important step towards automatic recognition of those larger social goals. for example, automatic detection of alignment moves and authority claims could facilitate the detection of sections of multiparty interactions where the participants take sides regarding some particular issue (which in turn could help in the identification of contentious issues); interactions in which participants attempt to establish themselves as credible sources in order to make power plays (kriplean et al. 2007); and in identifying potential trolls, scapegoats or other participants who tend to become the focus of a strong negative response from an entire group. similarly, as social acts such as alignment moves and authority claims are means by which individuals perform roles (e.g., moderator, expert or novice), automatic detection of social acts can contribute to the automatic detection of fluid roles in conversation exchange. to the extent that the distribution and form of the social acts vary across genres, the inclusion of different genres in annotation guideline development and in the annotated corpus is critical to the development of robust detection algorithms. in addition to the engineering goals of automatic detection of social acts, this annotated corpus holds scientific interest. we offer operational definitions of our social acts such that they can be annotated at the level of a turn (or sub-turn utterance), but not directly in terms of the linguistic form that they take. this approach allowed us to apply the same definitions across genres and across languages (while allowing the annotation guidelines to develop as we moved to these new corpora) and thus to compare the distribution and linguistic form of the social acts across genres and languages. in this article, we report some preliminary comparative analysis across mandarin, english and russian language wikipedia discussions, and present some initial data and observations on the distribution and presentation of social acts in synchronous internet relay chat discussions in english. however, these findings are intended to be illustrative rather than definitive, and certainly not exhaustive. there is still much more that could be done with these corpora, which we have made available for anyone who is interested in working with it. in this paper, we define the social acts of interest, describe our two corpora and annotation scheme, and highlight strategies and challenges that emerged during our annotation process. our 2 both available from http://ssli.ee.washington.edu/projects/scil.html 2 morgan et al. analysis considers the distribution of social acts across three languages and two genres, and explores hypotheses to illustrate potential interactions between social acts and related social phenomena. we draw on existing literature in linguistics, communication and rhetoric to define our two types of social acts, alignment moves and authority claims. from linguistics and organization studies, we draw on the concept of “identity work” to motivate our focus on these two types of social acts and guides our hypotheses regarding their potential characteristics and interactions. we describe how we generated and refined our annotation guidelines iteratively during annotation, drawing on observations and feedback from the annotators as they applied each version of the guidelines and we highlight effective strategies and unanticipated challenges we encountered. we focus primarily on those strategies and challenges that relate to achieving reliable annotation of discourse-level phenomena and on localizing a set of annotation guidelines to different languages and genres. we present a set of preliminary analyses that illustrate potential languageand genre-mediated differences in the expression and distribution of alignment moves and authority claims. finally, we test hypotheses related to how the performance and reception of a social act on wikipedia is shaped by aspects of the contributor’s identity, such as their official status within the community and their level of experience. 2 social acts the annotation of social acts differs from other annotation efforts concerned with linguistic structure or even linguistic meaning. in the creation of a syntactic treebank (e.g., the penn treebank, marcus et al. 1993), annotators are concerned with the linguistic form of the utterances they are annotating. the meaning is also relevant in that annotators use their understanding of each utterance in context to select among different possible syntactic structures to annotate. the annotations, in turn, make explicit aspects of the linguistic structure. similarly, annotation efforts concerned with lexical and propositional semantics make explicit parts of the linguistic meaning (again, as disambiguated by context) closely aligned with sentence structure or with particular words (e.g., which phrases fill which semantic roles (baker et al. 1998, palmer et al. 2005), the scope of negation (szarvas et al. 2008), or word sense (snyder and palmer 2004)). in contrast, social acts concern not the form nor the linguistic meaning of utterances but what those utterances are used to accomplish. in this sense, annotation of social acts is akin to dialogue act annotations (shriberg et al. 2004) or social events (agarwal and rambow 2010). nonetheless, there are differences here as well. dialog acts concern how the turns in a conversation connect to each other (e.g., question-answer pairs), whereas our social acts concern the social positioning of a discussant within a group. indeed, dialog acts and social acts can be seen as orthogonal tiers of annotation: e.g., both a question and an answer to it can be positive or negative alignment moves. our social acts differ from agarwal and rambow’s (2010) social events in that we are concerned with what discussants are accomplishing (or attempting to accomplish) with their utterances, whereas agarwal and rambow were looking to extract descriptions of interactions from narratives. in other words, our social acts are attributed to the interlocutors; their social events to entities described by an author. social acts, like dialog acts or social events, can be detected in part on the basis of the syntactic and semantic structure of the utterances that constitute (or, in the case of social events, describe) them. however, it is important to note that the linguistic structures available for performing social acts are highly varied, perhaps even more so than those available for performing dialog acts, which tend to be at least somewhat conventionalized (searle 1975). therefore, in creating annotated resources for social acts, it is neither feasible nor sufficient to annotate linguistic structures and then use these as the basis for annotating the social acts. 2.1 identifying social acts on wikipedia although wikipedia, and wikis in general, are designed to make it easy for an individual to make content changes without consultation or prior approval (cummings 2008), on many wikipedia 3 are we there yet? articles editors conventionally discuss and justify major changes that they make to article content, using other wiki pages as discussion forums (viegas 2007). clay shirky has called wikipedia “the product of unending argumentation... which grows not from harmonious thought but from constant scrutiny and emendation.” (shirky 2008) these discursive norms mean that wikipedia editors must often perform complex negotiations about article content. they publicly align with or against other editors in the discussion by making statements that support or oppose proposals made by other editors, such as including a particular image in the article or re-writing the introductory section, or that express approval or disapproval of particular edits others have already made to the article. editors must also justify changes they have already made, since any particular edit can be reverted or re-written by another editor at any time. the two types of social acts described in this paper, authority claims and alignment moves, are therefore particularly relevant to understanding the dynamics of wikipedia editorial discussions. however, we believe that the expression of authority and alignment is common across a variety of discursive contexts, especially contexts in which conversational participants are engaged in collaborative activities such as group decision-making, debate and production. furthermore these social acts themselves exhibit a degree of regularity both within specific contexts (such as a wikipedia discussion in english, with its singular jargon and conventions) and also across contexts. in the following sections, we describe the theoretical and empirical basis for authority and alignment, present the annotation schemes we developed for identifying them, and reflect on the annotation process. claim type definition example credentials credentials claims involve reference to education, training, or a history of work in an area. "speaking as a native born midwesterner who is also a professional writer. . ." experiential experiential claims are based on an individual’s involvement in or witnessing of an event. "if i recall correctly, god is mentioned in civil ceremonies in snohomish county, washington, the only place i’ve witnessed one." institutional institutional claims are based on an individual’s position within an organization structure that governs the current discussion forum or has power to affect the topic or direction of the discussion. forum forum claims are based on policy, norms, or contextual rules of behavior in the interaction. "do any of these meet wikipedia’s [[wprs | reliable sources ]] criteria?" external external claims are based on an outside authority or source of expertise, such as a book, magazine article, website, written law, press release, or court decision. "the treaty of international law which states that wars have to begin with a declaration is the hague convention relative to the opening of hostilities from 1907" social expectations social expectations claims are based on the intentions or expectations (what they think, feel or believe) of groups or communities that exist beyond the current conversational context. "i think in the minds of most people, including the government, the word “war” and a formal declaration of war have come apart." table 1. authority claims by type. 4 morgan et al. 2.2 authority claims the ability to persuade others to believe in one’s statements or the soundness of one’s judgments is a necessary component of human interaction. in order to establish the necessary credibility to secure the belief or assent of others, communicators will often couch their statements in some broadly-recognized basis for their authority on the matter. these “arguments from authority” have long been recognized as an important component of informal logic by many language philosophers (locke 1959 [1690], liu 1997). the self-presentation of authority has also been empirically examined in a variety of spoken and written contexts by scholars from disciplines such as communication, rhetoric, health studies, sociolinguistics, linguistic pragmatics and political science (for instance, galegher 1998, jensen 2003, mackiewicz 2010, richardson 2003, thompson 1993), providing a framework for understanding the strategies and conventions that communicators operating in different genres and media employ to establish themselves as credible discursive participants. although the linguistic construction of authority claims can vary greatly within a single genre and across genres, the regularity in the types of claims that are made and the construction of claims facilitates empirical analysis. authority claims provide an interesting lens through which to view an authored text or a conversation transcript, as the overall frequency of claims can reflect the nature or purpose of the discourse (for example a task-oriented collaboration vs. an undirected conversation) and the distribution of claim types can reveal features of the social context in which they are made, such as shared norms, practices and community values. for example, since certain bases for authority may be seen as more credible than others in certain contexts (such as citation of peer-reviewed publications in academic scholarship, or references to personal experience in online support groups), discursive patterns related to the expression of authority in a written text or a conversation transcript can illuminate the shared values of speakers and audiences in a given genre (galegher et al. 1998). although the linguistic construction of authority claims can vary greatly according to the genre of the communication, within a single genre there is often regularity in the ways claims are made, such as the common i’m a long-time listener introduction used by radio talk-show call-in guests. even across genres, recognizable types emerge: references to personal credentials (such as education or profession) are found to be important in newsgroup messages (richardson 2003), product reviews (mackiewicz 2010) and online scientific article comments (shanahan 2010). our taxonomy of authority claims was iteratively developed based on empirical analysis of conversational interaction in two different genres: political talk shows and wikipedia discussion pages (oxley et al. 2010), with reference to the literature cited above. we classify authority claims into the following types: credentials, experiential, external, forum, institutional and social expectations. while our claim types were developed independently of existing classification schemes, several of our types do mirror categories used in previous research. richardson’s (2003) warranting strategies are particularly salient, particularly warranting by source (similar to our ‘external’ claim type), warranting by reference to personal experience (see ‘experiential’, below) and warranting by reference to status (see ‘credentials’ and ‘institutional’, below). our codebook3 includes detailed definitions as well as positive and negative examples for each claim type. see table 1 for a list of our claim types, with brief definitions and examples. 2.3 alignment moves in multiparty discourse, interpersonal relationships among participants manifest themselves in social moves that participants make to demonstrate alignment with or against other participants. expressing alignment with another participant functions as a means of enhancing solidarity with                                                                                                                 3 available from http://ssli.ee.washington.edu/projects/scil.html 5 are we there yet? that participant while expressing alignment against another participant functions as a means of increasing social distance between conversational participants, particularly in situations where participants may be previously unacquainted with each other (svennevig 1999). changes in the alignment of participants toward one another, or shifts in footing (goffman 1981), may reflect long-term changes in interpersonal relationships or may be more transitory, demonstrating minor concessions and critiques embedded within larger, more stable patterns of participant agreement and disagreement (wine 2008). this concept of alignment is a different phenomenon from pickering and garrod’s (2004) alignment, in that it describes participants’ attitudes towards one another rather than how they linguistically represent and align their mental models in order to perform successful communication. ways of expressing agreement and disagreement can vary according to a variety of social factors, including power relations among participants, gender, participant goals, and conversational context (rees-miller 2000). research has suggested that expressions of agreement and disagreement in written language tend to be more explicit than oral expressions of agreement and disagreement (mulkay 1985; mulkay 1986). text-based online discussions generally reflect this turn towards more explicit cues (baym 1996) as participants compensate for the lack of the many non-linguistic communication mechanisms that are available in face-to-face interactions. in move type example baym (1996) positive alignment markers of agreement explicit agreement “i agree.” “yes.” “that’s right.” “correct.” “exactly.” “absolutely.” “of course.” “i don’t disagree.” “i’ll do that.” explicit indicants of agreement praise/thanking “great idea.” “good point.” “well said.” “thanks for your work on the article. it looks much better.” expressions of gratitude positive reference to previous speaker’s point “like bill was saying . . .” “as you mentioned earlier . . .” “as mary was indicating . . .” acknowledgment of other’s perspective other positive clear indicator of positive alignment that doesn’t fit into one of the above categories smiley faces negative alignment markers of disagreement explicit disagreement “i disagree.” “no.” “that’s wrong.” “that’s false.” explicit indicants of disagreement doubting “i doubt that.” “i don’t think so.” “that’s questionable.” “you can’t be serious.” qualification sarcastic praise “that’s a great plan. while you’re at it, why not destroy the entire article?” n/a criticism/insult “that’s ridiculous.” “that’s a terrible idea.” “you’re nuts.” “you fool.” n/a other negative clear indicator of negative alignment that doesn’t fit into one of the above categories n/a   table 2. alignment move types. 6 morgan et al. asynchronous online group discussion forums such as those on usenet and wikipedia, the way agreement and disagreement are expressed is also mediated by the non-dyadic nature of the discussion. for instance, if an editor on a wikipedia talk page disagrees with another editor, that disagreement is effectively public: it is equally visible to a large audience of (known and unknown) other conversational participants. depending on community norms about the acceptability of disagreeing, the public quality of interpersonal communication may lead a speaker to perform more explicit “facework” (baym 1996, brown and levinson 1987) by either downplaying their disagreement (for instance, by hedging) or by exaggerating it, even to the point of rudeness or flaming. the presence of a large, participatory audience can also cue speakers to exhibit more “front stage” behaviors (goffman 1959), exaggerating their (discursive) movements like an actor on a stage. thus, a wikipedia discussion participant is likely to exhibit a more regularized pattern of agreement and disagreement as they shape their language to adhere to local norms and adopt the jargon of their anticipated audience. different languages also reflect differing conventions for expressing agreement and disagreement (see, for example, mori 1999), which constitutes another mediating factor in how these social acts are performed in a given situation. our annotation scheme contributes to existing research on the role of conversation context, medium and language in shaping online discourse by accounting for a range of alignment cues in two types of text-based, task-oriented online discussions across three languages. we classify alignment moves into positive and negative types, according to whether the participant is indicating direct (explicit) or general (implicit) agreement or disagreement with the target: positive alignment is annotated in cases of explicit agreement, praise/thanking, positive reference to another participant’s point or where other clear indicators of positive alignment are present. negative alignment is annotated in cases of direct disagreement, doubting, sarcastic praise, criticism/insult, dismissing, or where other clear indicators of negative alignment (such as typographical cues) are present. we did not explicitly adopt our alignment sub-types from any previous researchers’ classification schemes. instead, our categories were developed through iterative exploratory qualitative analysis of broadcast talk show transcripts from the gale broadcast speech corpus4 and a sample of wikipedia discussion pages. however, like our authority claims, many of our alignment sub-types are similar to those defined by other researchers. in particular, baym (1996) recognized a set of alignment cues that mirror our own. a summary of our alignment move types and related markers identified by baym is presented in table 2. a more detailed list of alignment types, cases and definitions are available in the codebook included with our corpus at the url previously specified. 2.4 social acts and identity work “identity work,” a sociological concept common to both organization studies and sociolinguistics (alvesson and willmott 2002, bucholtz and hall 2010), describes the way in which individuals’ social identities are constructed, reflected and transformed through their own communication practices and those of the people with whom they interact. both wikipedia talk pages and irc channels are contexts in which participants often interact with others whom they only know within that context, or with whom they have never interacted before. online identities are often fluid and transitory even in cases where they are persistent within the system (for instance, in the form of a user name, handle or pseudonym) since a user may use multiple identities, and identifying features (such as a users’ profile information) can be hidden or altered at any time. in these spaces participants also lack many common mechanisms used in face-to-face contexts to communicate attention, emotions, roles, and goals. for instance physical gestures, eye contact, facial expressions, posture, and vocal inflection are not available in these textually-mediated                                                                                                                 4  http://projects.ldc.upenn.edu/gale/data/catalog.html   7 are we there yet? environments. additionally, each participant’s identity is less likely to be pre-established and persistent in wikipedia and irc discussions than in discussions among known others, such as in business meetings. in such spaces we expect the discursive practices associated with “identity work” to be more explicit, as discussion participants constantly re-present aspects of their personal values and their social status, their attitudes and allegiances in the text they type. given the few visible markers of identity on wikipedia and the fact that editors are constantly interacting with new collaborators, wikipedians perform authority by adopting insider language and other community-specific norms of interaction related to the task of collaboratively writing an encyclopedia (see for example kriplean et al. 2007). supporting arguments with specific references is one such norm. in order to investigate the relationship between social acts and identity work, we explore the extent to which wikipedia editors’ authority claims reflect certain socially salient aspects of their identity as wikipedians. we use two particular social identity measures, user role and their degree of experience, which are manifested in specific ways within wikipedia. we hypothesize that as editors become more integrated into wikipedia, they will make more authority claims. in order to test this hypothesis, we leverage metadata about individual wikipedia editors that is captured in our dataset but is not immediately visible to other participants: user roles (for instance, administrator, registered editor, anonymous editor), total lifetime edits, and length of membership. we also describe “v-index,” a measure of an editor’s degree of integration, investment or “veteran status” within the wikipedia community at a particular point in time. inspired by ball’s (2005) “h-index” of scholarly productivity, v-index balances frequency of interaction with length of interaction. specifically, an editor’s v-index at the time of a particular edit (in this case, a conversational turn) is the greatest v such that the editor has made at least v edits to wikipedia within the past v months (28-day periods). our tests of these hypotheses about the relationship between identity and social acts are described in section 5. 3 data collection we gathered english language wikipedia data first, and then gathered similar samples from the russian and mandarin wikipedias. the aacd corpus was gathered later, and is intended to offer meaningful comparisons and contrasts with the aawd corpus. we gathered the aacd by facilitating a set of task-based group irc discussions using participants we recruited locally. 3.1 wikipedia talk page discussions wikipedia talk pages (also called discussion pages) are editable pages on which wikipedia editors can take part in threaded, asynchronous discussions about the content of other pages, particularly the article pages that most visitors to wikipedia are familiar with. every article page in wikipedia has an associated talk page wherein editors can discuss and collaboratively plan editing actions on that article; any editor interested in a given article can join the conversation on that article’s talk page. conversational exchanges on the talk pages may take the form of a polite deliberation aimed at achieving a final consensus-based decision or a heated argument as editors advocate different ideas about matters concerning the content or form of an article. each edit to the talk pages is recorded as a unique revision in the system and thus becomes part of the permanent record of system activity. wikipedia constitutes a particularly valuable natural laboratory for studies such as this one for several reasons. first, the interaction among the participants is almost entirely captured within the wikipedia database: while some wikipedians might interact with each other in person or in other online forums (such as irc channels or mailing lists), this is the exception rather than the rule. furthermore, while participants often maintain persistent identities (usernames for registered users; ip addresses for unregistered ones) there are no other cues to social identities available to the participants beyond what is captured in the digital record. therefore all of the effort that participants put into constructing their online identities is in the record for analysis. second, the 8 morgan et al. discussions on wikipedia talk pages tend to be goal-oriented, as the discussion topic is the wikipedia article that the participants are collaboratively editing. this goal-orientation motivates participants to explicitly align with each other in the course of discussions and buttress their arguments with authority claims. finally, the wikipedia dataset contains rich metadata, such as the date and time of each edit (identified by revision id) to every article or talk page; the editor responsible for the edit (identified by username or ip address, depending on registration status); and markup such as hyperlinks and formatting used in the textual content of each edit. these metadata allow for sophisticated data analysis at the editor level (e.g., how many edits made by one editor in a given span of time) and the page level (e.g., how many editors have participated in a talk page discussion). our english-language wikipedia dataset is drawn from a publicly-available 2008 wikipedia xml data dump5 and is composed of 365 discussions associated with 47 talk pages. these articles were selected based on a list of topic keywords extracted from a set of english-language                                                                                                                 5 http://en.wikipedia.org/wikipedia:data_dumps english counts total annotated discussions 185 total turns 3361 turns w/ external claim 459 turns w/ experiential claim 77 turns w/ forum claim 260 turns w/ credentials claim 3 turns w/ social expectations claim 21 turns w/ any claim 703 mandarin total annotated discussions 225 total turns 1517 turns w/ external claim 60 turns w/ experiential claim 24 turns w/ forum claim 77 turns w/ credentials claim 7 turns w/ social expectations claim 3 turns w/ any claim 155 russian total annotated discussions 122 total turns 893 turns w/ external claim 67 turns w/ experiential claim 7 turns w/ forum claim 19 turns w/ credentials claim 2 turns w/ social expectations claim 4 turns w/ any claim 82  table 3. total authority claims in aawd corpus by language. 9 are we there yet? broadcast news transcripts in the gale broadcast speech corpus. a python script was used to query the google search engine to identify english language wikipedia articles related to those keywords. because the russian and mandarin language editions of wikipedia have fewer users and therefore less talk page discussion over all, our mandarin and russian datasets consist of discussions from the talk pages on those language editions of wikipedia that had the most edits, rather than articles specifically related to those in the english language dataset. all the selected discussions contain at least 5 conversational turns and at least 4 human participants.6 we set this minimum threshold for discussion length and number of participants because we expected that our social acts would be most evident in discussions where there was a higher level of interactivity. because not all wikipedia articles are highly collaboratively created, contested, or frequently updated, many talk page threads do not meet these minimum criteria. of the 365 discussions in our final english dataset, 185 were annotated for both alignment moves and authority claims. the mandarin and russian-language versions of wikipedia were annotated for both authority and alignment in 225 and 122 discussions, respectively. see table 3 for a breakdown of authority claim annotation across the three languages, and table 4 for a breakdown of alignment move annotation in english. all annotation was completed by paid university students, and took place between 2009 and 2011. we anticipated that it would be difficult to perform annotation on both types of social acts (alignment and authority) simultaneously, so annotators were instructed to only annotate a discussion for one social act at a time. 3.2 irc discussions our chat data is based on a set of 12 textual, synchronous exchanges among four-person groups chatting in a private irc channel. these exchanges were all facilitated by the researchers in an effort to develop a corpus of language use examples that would share some characteristics with the wikipedia discussion page corpus described above. in each of the 12 sessions, 4 individuals interact through internet relay chat (irc) for approximately 45 minutes. these sessions were all saved as time-stamped transcripts. the chat                                                                                                                 6 some of the turns in wikipedia discussions are actually contributed by automated agents, called “bots.” table 4. total alignment moves in aawd corpus by language. n % english total turns 2890 100 turns w/ positive alignment 330 11.4 turns w/ negative alignment 467 16.2 turns w/ any alignment 710 24.6 total editors 905 100 editors making alignment moves 315 31.9 russian total turns 1806 100 turns w/ positive alignment 558 31 turns w/ negative alignment 142 8 turns w/ any alignment 645 36 mandarin total turns 2767 100 turns w/ positive alignment 808 29 turns w/ negative alignment 153 6 turns w/ any alignment 925 33   10 morgan et al. datasets include four 45-minute exchanges in english, four in mandarin, and four in russian, for a total of 48 total participants across languages. all participants were native speakers of the language in which the session was conducted and were comfortable typing their language on a standard qwerty keyboard using standard microsoft windows 7 language packs to remap keys to non-english character sets, as appropriate. all sessions occurred between august 2010 and august 2011. participants were primarily recruited through university list-serves and by word-ofmouth. additional participants were recruited through paid advertisements in a university daily newspaper, both in print and online. all participants received compensation for their participation in the form of a $25 gift card. given that our chat participants were primarily recruited through word of mouth and were therefore not sampled randomly, there are notable demographic differences between the three language groups. the mean ages of our english, mandarin, and russian language groups were 20, 26, and 30, respectively. education levels varied between groups; 13% of english participants, 100% of mandarin participants, and 81% of russian participants had earned a postsecondary degree. participants also differed somewhat in their use of online chat systems. while 94% of the english participants and 100% of the mandarin participants used online chat systems on a daily or weekly basis, the russian participants used online chat systems less frequently, with only 69% of participants reporting that they used online chat systems daily or weekly. participants in all three groups reported that they primarily used online chat for communicating with people that they know offline (94%), and slightly under half of the participants also used online chat for conducting meetings (42%). session number 1 2 3 4 total claims session code 5ne 4ne 3ne 2ne authority external authority 8 9 4 2 23 experiential authority 7 5 2 5 19 social expectations authority 2 1 2 0 5 forum authority 1 0 0 1 2 credentials authority 0 2 0 0 2 column totals 18 17 8 8 51 alignment positive alignment 59 56 73 53 241 negative alignment 29 38 7 7 81 totals 88 94 80 60 322 total turns in chat session 646 654 534 395 2229  table 5. authority claim types and alignment moves across chat sessions for english. 11 are we there yet? a spreadsheet containing all demographic information collected on the chat participants is provided with the aacd at the url previously specified. the sessions were stimulated by the researchers, using a common scenario for the participants across all sessions. at the beginning of each session, participants were assigned to one of four different discussion roles: project manager, accountant, publicity coordinator, or secretary. a brief description of the responsibilities incumbent on each role were provided in the scenario prompt, which is available with our dataset. our chat scenario was thus designed to reflect in a synchronous forum some of the features of wikipedia talk page discussions, which often address specific considerations and decisions about wikipedia articles and involve multiple editors with different knowledge sets, motivations and points of view. in each chat session, the participants were asked to work together toward a shared goal (the planning of a party for students in a large university lecture course), where each participant’s tasks and concerns differed according to their assigned role. each of the discussions in this collection was independently annotated for authority claims and alignment moves by a native speaker researcher. because these irc discussions were not dually annotated, these data and the preliminary analyses we provide below should be considered provisional and illustrative of the potential for cross-genre comparison. 3.3 genre similarities and differences together the aawd and aacd corpora represent productive resources for exploring social acts in online language use. we present these corpora together in order to provide opportunities to explore the way genre, medium and language shape social acts. both corpora include interactions among individuals who are loosely bound by a collective orientation to accomplishing the same task (either creating an encyclopedic article or planning an event). the shared task nature of these interactions provides some constraint on the topical focus of each exchange set though individual participant contributions are not further constrained. in both corpora the interactants rely on aawd aacd computer-mediated communication type asynchronous interaction synchronous interaction time constraints open-ended 45 minutes previous interaction among participants primarily online primarily offline conversational roles loosely-structured: unregistered (ip), registered editor (username), administrator (username) pre-assigned: project manager, publicity coordinator, secretary, accountant task type collaborative writing collaborative event planning number of participants per conversation four or more four motivation for participating uncompensated compensated conversation topic determined by related article topic pre-assigned   table 6. genre and task differences in the wikipedia and chat datasets.   12 morgan et al. online language to make their individual intentions known to others and to affect the group’s thinking about the task at hand. still, important differences also distinguish the two corpora. the aawd data represents asynchronous interactions, while the aacd data is synchronous interactions. related to this, in the aawd data there is no time constraint related to the accomplishment of the motivating task, but in the aacd data that participants had a set time (45 minutes) to work toward their task goal. genreand task-mediated differences between the two datasets are summarized in table 6. 4 annotation process in this section we both present the results of our annotation and analysis, and reflect on the benefits and drawbacks of our framework and process. we find that while the social acts we chose to capture exhibit some regularity, even after extensive guideline revision and annotator training we were not able to achieve cohen’s kappa scores of higher than 0.6 for most of our social act sub-types. one of the greatest challenges in our annotation effort was to reliably identify instances where a social act was present, and to set meaningful boundaries that presented annotators with as few problematic edge cases as possible. encouragingly, however, we find that in instances where multiple annotators annotated a particular utterance as containing a social act, we saw much higher agreement on the type of social act presented in that utterance, suggesting that in more prototypical cases, elements of each social act were relatively consistent in their presentation. we will also discuss the process of localizing our guidelines, originally developed by native english speakers with reference to english language data, to mandarin and russian, and how this localization process was guided by differences in the manifestations of authority and interpersonal alignment in different languages but within the same genre. 4.1 guideline development we developed our initial authority claim and alignment categories through qualitative examination of english language wikipedia discussions outside of our dataset, as well as transcripts of political talk show broadcasts in the gale corpus. two of the researchers developed an initial data-driven set of authority claim types, and a third researcher expanded this typology through additional analysis of talk pages, guided by established constructs related to the self-presentation of authority from rhetorical theory and informal argumentation. our authority claims are not intended to map directly onto categories established in previous work (for instance, classical rhetorical appeals to ethos, logos and pathos). rather they are intended to capture the ways in which these universal discursive moves manifest in topic-focused debates and task-based deliberation across a variety of online contexts. furthermore, although our goal was not to apply existing categorization schemes to our data, some of the authority claim types we developed are similar or complementary to categorizations employed by other researchers (for example mackiewicz 2010, richardson 2003). see section 2 above for additional citations. we identified candidate sub-types of alignment moves through a review of the literature on conversational agreement and disagreement, especially the work of brown and levinson (1987) and baym (1996). this candidate list was vetted and supplemented and through sample annotation by two of the researchers, and in response to student annotators’ observations. alignment sub-types proved less easily discriminable than our authority claim types, leading us to treat them as illustrative cues rather than an exhaustive and mutually-exclusive set of categories. alignment cues were used to identify the presence of alignment in a turn, and to discriminate between instances of positive and negative alignment. our process for refining these guidelines is described in further detail below. our initial english guidelines included simple descriptions of the phenomena to be annotated, but only a few examples. as coding progressed, more examples were added to provide necessary clarification to the social acts categories as annotators raised questions during weekly annotation 13 are we there yet? meetings. we developed both positive and negative examples in order to clarify the boundaries between social act categories (for instance, to distinguish between an external and an experiential authority claim), and to establish heuristic thresholds to help annotators decide what qualified as a social act (for instance, whether the use of the word “yeah” always counted as positive alignment, or whether other contextual features were necessary for determining whether or not it was merely a backchannel utterance.). we summarize and illustrate some of these adaptations below. 4.1.1 differentiating alignment from similar or contrary opinion while testing out our guidelines, we observed that it can be difficult to differentiate between a turn in which a speaker expressed a similar or opposing opinion to a previous turn, and one where the turn-taker was actually aligning with or against a previous speaker. in many cases on wikipedia, a speaker will express a similar idea as a previous talk page participant, or make an alternative, contradictory proposal without explicitly orienting their statement towards (or making any reference to) the person they are supporting or contradicting. in order to increase consistency in annotation, we required annotators to find substrings within the utterance which explicitly mark it as alignment and identify them with keyword tags, in order to mark a turn as containing alignment. 4.1.2 identifying sarcasm sarcastic statements can be difficult to recognize reliably. during both guideline development and refinement, we encountered cases where one annotator labeled an utterance as sarcastic and another did not, and they were not even able to successfully reconcile their opinions during open discussions at annotator meetings. to address this issue, we limited the types of sarcasm which we labeled as a potential alignment cue to sarcastic praise. we also prompted our annotators to look for typographical cues related to sarcasm, such as bolded or italicized words or the use of caps (for example, “oh, sure, that’s a great idea.”) 4.1.3 specifying personal experience identifying experiential claims reliably also proved difficult initially. we found that discussion participants often related events that they had experienced, or things that had happened to them outside of the context of establishing authority. in order to aid in identifying references to personal experience that were specifically linked to authority claims, we limited our annotators to experience-related utterances that contained a first or second-person pronoun, such as “because i was living there at the time” or “we never used to say it that way.” 4.1.4 naming social groups social expectations claims, though infrequent in our data, proved difficult to label consistently. one way we attempted to address this difficulty was to require our annotators to only label a claim as “social expectations” when it contained a reference to a named group, such as “wikipedia readers”, “the gop” or “iowa voters.” this helped guide annotators in situations where the group being named was ambiguous or implied, such as in “some editors think” or “they won’t be convinced by mere facts.” we also found that these discussions included many statements about what named groups were doing or had done in the past, which were often not associated with claims of authority, such as “the media overemphasized his role” or “the administration is acting swiftly.” we addressed this by emphasizing with additional examples that our definition of ‘social expectations’ required a claim that made a statement about what groups thought, believed or desired. 4.1.5 connecting moves with targets it is difficult to identify whether an utterance contains an alignment move without a clear indication of who is being addressed. in order to help our annotators differentiate clear alignment 14 morgan et al. moves from more ambiguous cases, we required that each labeled move meet basic criteria related to identifying a target of the utterance. alignment moves had to include either a named target (for example, “i doubt that will work, tom”) or some other unambiguous personal reference such as a second person pronoun and be situated in the discussion thread in such a way that the person the speaker was referring to was clear from the context (such as the editor who made the immediately previous post, or the editor who started the thread). research on the role of ‘facework’ in interpersonal communication, as presented in brown & levinson (1987) and discussed extensively in baym (1996), provides support for our personal pronoun requirement: using personal references, such as attributing an idea to a person through use of a personal pronoun, is a stronger indicator of negative alignment than making an oblique comment on an idea because it constitutes an explicit threat to the target’s positive face. all social acts were annotated at the “turn” level, with each “turn” representing a single message in a discussion thread. each turn could contain multiple instances of a social act: for instance, an author could make an experiential and an external claim within the same turn. in order to capture these phenomena, all utterances that contained claims were labeled with keyword tags, as in example 1: example 1. “i’ve read up on thisand the most recent new york times op-ed says that biden was right.” individual authority claims were identified and marked by contiguous keyword spans within a single sentence. we found that cues to alignment with or against a specific target tended to be less regularly expressed, and were often scattered across multiple sentences. in such cases, our use of keyword tags to capture words and phrases indicative of alignment allowed us to label our data accurately and flexibly. multiple alignments with or against multiple targets could be annotated at the level of the entire turn, and one or more keyword spans associated with each alignment move could be marked across different sentences within the turn. for example: example 2. “that’s right, speaker2. speaker3’s way off base, but you seem to have a good solution. however disagree with your name for the section – ‘iraq war’ is used in the united states media and should be used here as well." therefore, marking alignment moves at the turn level, allowing multiple alignment moves per turn and identifying a single alignment move by multiple keyword spans, allowed annotators to capture instances where an author expressed both positive and negative alignment towards the same target within the same turn, as in example 2 above. 15 are we there yet? 4.2 annotation tool the annotation tool (a modified version of ldc’s xtrans (glenn et al. 2009)) allowed annotators to indicate the presence and type of claims or moves in each annotation unit, in addition to selecting spans of text corresponding to each social act. for alignment moves, within a turn, alignment of the same type (positive or negative) with the same target was annotated as a single alignment move, even across multiple sentences. where the type or target differed, we annotated up to three separate alignment moves per annotation unit. for authority claims, we also annotated up to three claims per annotation unit, with each claim identified by a single span of text. for authority, claims in separate sentences of an annotation unit counted as separate even if they were of the same type. in some cases, the imperfectly re-created threaded structure of our data in xtrans made the target of an utterance difficult to identify, so we also included an optional hyperlink to the text of each turn on the live wikipedia website to allow annotators to view the turn in context. 4.3 guideline refinement we refined our social act categories iteratively, through weekly group annotator meetings and by setting up an annotator email list to which the annotators could send questions as they worked. one of the primary difficulties these meetings addressed was to circumscribe ambiguous or hazy categories by creating firm rules about what did and what did not count as an instance of a social act. these rules were based on common patterns we saw in the data, and generally took the form rounds of mandarin annotation authority agreement (before review) alignment agreement (before review) annotator group 1, final round (5/2011) 0.56 0.61 annotator group 2, 1st round (7/2011) 0.13 0.30   table 7. comparison of inter-annotator agreement between mandarin annotator groups. mandarin annotation rounds (group 2) authority (before review) alignment (before review) first round, 7/2011 0.13 0.30 second round, 7/2011 0.44 0.42 third round, 8/2011 0.72 0.64 russian annotation rounds authority alignment fourth round, 6/2011 0.56 0.52 fifth round, 7/2011 0.52 0.58   table 8. longitudinal comparison of inter-annotator agreement for second mandarin annotator group (top) and russian annotator group (bottom). 16 morgan et al. of structural or linguistic cues (see examples above). for instance, in order to remove ambiguity about the target of an alignment move, annotators were only allowed to specify a target if that user was mentioned by name, if the alignment target was the author of the previous post in the thread, or if the target was the author of the first post in the thread. this narrowing of the possible targets available to annotators increased consistency in target identification and greatly simplified the process of target identification for our annotators. annotator turnover and guideline refinement both affected inter-agreement (measured as averaged cohen’s kappa). to illustrate this challenge, we present inter-annotator agreement data from the mandarin annotation project carried out between may and august of 2011. in may 2011, two of the three mandarin annotators left the project. in late june, two new annotators were trained and joined the project. comparisons of the inter-agreement rates for the last round of the old annotators and the first round of the new annotators are provided in table 7. the new mandarin annotator team witnessed the second mandarin guideline refinement and other minor guideline modification. as noted, these changes were based on the annotators’ questions and comments. as the mandarin guidelines were refined, the inter-annotator agreement rates improved. in the case of russian annotation however, discussion and guideline refinement did not necessarily lead to increased inter-annotator agreement rates. longitudinal comparisons of annotator consistency in mandarin and russian are presented in table 8. 4.4 adaptation of guidelines to mandarin and russian the mandarin and russian annotation guidelines were adapted based on the english guidelines. the mandarin guidelines went through two major rounds of refinement. in the first round, a researcher who was a native speaker of mandarin adapted the english language guidelines and applied the preliminary guidelines to 5 sample annotations. adaptations included adding examples of emphatic markers that are used to indicate disagreement in mandarin. the researcher also removed the statement (included in the english guidelines) that “no” can indicate agreement if the preceding statement is negative because “no” always indicates disagreement in mandarin. the second round of refinement took place after the mandarin annotators finished their first round of annotation. general questions and observations that were brought up in the annotation meeting were added to the guidelines, such as the use of rhetorical questions/paired conjunctions to indicate negative alignment. beyond these two major rounds of refinement, guidelines were iteratively refined based on annotator feedback and regular “spot checks” of annotated data. when the first drafts of the mandarin and russian guidelines were finished, the editors annotated 5 randomly selected mandarin and russian wikipedia discussions. this process quickly revealed further issues with the guidelines. for instance, it was found that people sometimes disagree with a wikipedia item, but not a particular person. in the initial guidelines, there was no clear stipulation that stated whether attacking a wikipedia item should be coded as a negative alignment or not. it was decided then to add a stipulation to the mandarin and russian guidelines that: “any alignment move that agrees/disagrees with an article, or wikipedia, should not be coded.” as with the english annotation guidelines, problems uncovered in the sample annotation pass were addressed through guideline modification. additional modifications were made during the annotation process, with earlier discussions re-annotated to maintain consistency. care was taken to keep the overall alignment move types consistent across all three languages, even as modifications were made to the types of alignment cues. the researchers working on the guidelines for mandarin and russian found that some of the behaviors described were common between english and mandarin/russian. therefore, they first translated the culturally shared keyword strings from english to mandarin/russian. for example, all three languages have very similar repository of explicit agreement/disagreement words, like “i agree 我同意 я согласен”, “yes 是的 да”, “i disagree我不同意 я не согласен”, 17 are we there yet? “no 不是 нет.” there do exist differences, however. for example, in the english alignment guidelines, there is a note that “no” can indicate agreement if the preceding statement is negative. but in mandarin, the situation is the opposite. specifically, in english, the answer “no” is factoriented, while in mandarin, it is speaker-oriented. a simple example is: example 3. (suppose mary is not a student.) english a: mary is not a student. b: no. she is not. mandarin a: mary is not a student. b: yes. she is not. therefore this note was removed from the mandarin alignment guidelines. another example that shows language differences is that in mandarin, rhetorical questions are frequently used in daily speech to indicate negative opinions. such rhetorical questions often have explicit lexical markers either at the beginning or inserted in the middle of the sentence, like “难道…? how can/can’t...?” (at the beginning of a sentence), and “怎么就…? how come…?” (in the middle of a sentence). below is an example: example 4. a: 如果 汉族的发源地   只有 黄河中下游的话, if han’s birthland only huang river’s lower and middle reaches, 那 只能 说明 今天的 中国 辽阔 领土 来源于 then only indicate today’s china’s large territory comes from 中央政权 对 其它 民族的 战争 与 征服。 central government to other peoples’ war and conquering. ‘if the han people’s birthland is only huang river’s lower and middle reaches, it could only indicate that china’s current large territory comes from the wars launched by the central government and their conquering towards other minorities.’ b: 不征服 怎么 来 土地? no conquering how come get land? ‘if they didn’t win through conquering, how could they get land?’ as a result, one extra category named “rhetorical question” was added under the negative alignment cues in the mandarin alignment guidelines while there is no such category in the english guidelines. in order to make the mandarin and russian guidelines more comprehensive and representative, more examples were added from early rounds of mandarin/russian annotation. guidelines were written in english, with examples in russian/mandarin, each with english translations, to ensure that they were easily accessible to the researchers and annotators across all 18 morgan et al. three languages. like english, mandarin went through several ‘test’ rounds of annotation starting from early 2010 using an early version of the social act guidelines. this annotation was based on samples drawn from both wikipedia discussion pages and political talk shows in the gale corpus. when the mandarin guidelines were revised and evaluated in preparation for the primary annotation task in 2011, the mandarin-speaking researcher viewed roughly 30 files that had been annotated in 2010 and selected a batch of keywords that seems to occur frequently in mandarin wikipedia discussions. these keywords were added to the mandarin guidelines as examples. the russian-speaking researcher followed a similar process, examining data from randomly selected russian wikipedia discussions and basing the keywords and examples for russian annotators on this data. as a result, the russian guidelines used two sources of keywords and examples: russian data examined by the russian guidelines editor and examples that were either translations of english examples or were created by the researcher who led the russian annotation effort in order to account for conversation topics that were not encountered in the initial set of wikipedia examples, but were likely to appear in the russian data. 4.5 annotation quality in complicated annotation tasks, such as those conducted in this work, evaluating annotation quality is a fundamental challenge. the most popular approach to measuring annotation quality is via the surrogate of annotation consistency. this assumes that when annotators working independently arrive at the same decisions they have correctly carried out the task specified by the annotation guidelines. several quantitative measures of annotator consistency have been proposed and debated over the years (artstein and poesio 2008). we use the well-known cohen’s kappa coefficient κ, which accounts for uneven class priors, so one may obtain a low agreement score even when a high percentage of tokens have the same label. we also report the percentage of instances on which the annotators agreed, a, which includes agreement on the absence of a n κ a authority claims forum 451 0.52 0.92 external 715 0.63 0.91 experiential 185 0.33 0.96 social expectations 78 0.13 0.98 credentials 6 0.57 0.99 overall 1157 0.59 0.86 alignment moves explicit agreement 379 0.62 0.94 praise/thanking 117 0.6 0.98 positive reference 86 0.2 0.98 explicit disagreement 453 0.29 0.92 doubting 198 0.23 0.96 sarcastic praise 38 0.3 0.99 criticism/insult 556 0.32 0.91 dismissing 396 0.16 0.91 all positive 509 0.66 0.94 all negative 1092 0.45 0.85 overall 1378 0.5 0.8   table 9. agreement summary for authority claims and alignment moves in the english language wikipedia dataset. n denotes the number of turns of the given type that at least one annotator marked. 19 are we there yet? particular label. when more than two annotators have labeled a set of instances, we compute the average of pairwise agreement. κ scores for authority claim and alignment move agreement in english are presented in table 9. the counts, n, presented in table 9 reflect the total number of turns marked as containing claims of a given type by any annotator, rather than the total counts of claims that were independently annotated by two annotators (which are presented in table 3). for authority, the most common types of claims, forum and external, are also two of the most reliably identified. for alignment, better agreement was demonstrated for the positive alignment sub-types than for the negative sub-types, but agreement was lower in general. the difficulty of making fine distinctions between the types of negative alignment move (for instance, discriminating between sarcastic praise and criticism/insult) appears to be a large factor in the low agreement scores. when all of the negative categories are merged, agreement is higher, although still less than for positive alignment moves. our κ values generally fall within the range that landis and koch (1977) deem “moderate agreement”, but below the 0.8 cut-off tentatively suggested by artstein and poesio (2008).7 one possible reason is that the negative class is not as discrete as it might be in other tasks: both alignment moves and authority claims can be more or less subtle or explicit. we designed our annotation guidelines to emphasize the more explicit variants of each, but the same guidelines can sometimes lead annotators to pick up more subtle examples that other annotators might not feel meet the strict definitions in the guidelines. the aawd corpus includes both social acts that were identified by at least two annotators working independently, as well as those that were identified by only one annotator. we expect our multiply-annotated authority claims and alignment moves to correspond to more blatant or prototypical examples and our singly-labeled moves and turns, while sometimes being genuine noise, to pick out more subtle examples. the aacd corpus contains only singly annotated social acts. a single native-speaker researcher annotated each irc discussion. 4.6 lessons learned through the process of creating and testing our annotation guidelines, we developed several strategies which were effective for resolving ambiguity within the guidelines, minimizing the cognitive load of the annotation task for our annotators, and increasing the overall quality of the annotation produced. these strategies are outlined below. 4.6.1 treat guideline development and annotation as iterative processes no researcher is omnipotent, and issues with the annotation guidelines are bound to arise in any annotation project, particularly when a set of guidelines developed on one language and genre is applied to new data. our researchers met with annotators on a weekly basis to identify ambiguities within the guidelines and to reach consensus on how those ambiguities should be resolved. we also set up a mailing list for annotators to email us questions as they arose during annotation and for the researchers to provide comments and feedback. this iterative approach necessarily entails revisions to the guidelines to improve clarity and updates to the annotation after points of contention have been ironed out. although this process can be time consuming, we agree wholeheartedly with macqueen et al. (1998, p.36) that “re-coding should not be viewed as a step back; it is always indicative of forward movement in the analysis”. 4.6.2 where possible, reduce cognitive load for the annotators although most of our annotators performed annotation of both of our social acts, alignment and authority, they were not asked to annotate both simultaneously. our annotators found it easiest to                                                                                                                 7  artstein and poesio also note that it may not make sense to have only one threshold for the field.   20 morgan et al. perform one type of annotation for several weeks rather than alternating frequently between different tasks. the annotation tool was set up so that when annotators chose to annotate a particular turn as containing a social act, they were automatically prompted with a list of possible acts to choose from, and were then provided with the next possible annotation labels based on the category chosen in the previous step. for example, if an annotator chose to mark a turn as containing positive alignment, they would then be prompted to categorize that positive alignment move as an instance of either “explicit agreement,” “praise/thanking,” “positive reference to a previous speaker’s point,” or “other.” building this sequence of options into the tool simplified the annotation task and eliminated some sources of accidental error, such as an annotator forgetting to indicate the type of positive alignment after marking its presence. 4.6.3 provide the annotators with specific criteria for deciding whether a turn should be annotated, and include both positive and negative examples although the goal of our annotation was to code for social acts and not for particular linguistic structures, we were able to simplify the annotation task by restricting the annotation of some categories to turns containing particular linguistic cues or criteria. for example, to narrow the range of what might be coded as an “experiential” authority claim, annotators were instructed to look for first person pronouns such as “i,” “me,” “my,” or “we.” when coding “external” authority claims, annotators were instructed that the turn must name a specific source as the basis for that claim (a book, website, article, etc.). although these decisions may have reduced the overall quantity of data coded with these categories, the addition of these criteria allowed annotators to apply these labels more consistently and with less deliberation (further decreasing cognitive load, as discussed above). based on sources of confusion identified at regular annotator meetings, we expanded our guidelines to include not only positive examples (macqueen et al. 1998, “inclusion criteria”) but negative examples as well (macqueen et al. 1998, “exclusion criteria”). these negative examples enabled annotators to more readily reject a turn that was not appropriate for annotation and to more easily distinguish between our annotation categories. 4.6.4 perform regular spot checks on the annotation to identify both low-level and highlevel disagreements between annotators in early phases of guideline development, the discussions at our annotator meetings centered on feedback that we provided to annotators on small segments of the data that they had annotated. questions raised by the annotators at these meetings also allowed us to discuss how to interpret the guidelines in unusual cases. this process worked well for identifying major points of ambiguity in our guidelines, but was limited by the fact that we could only discuss a small portion of the data that had been annotated. we realized that some consistent disagreements between annotators might also go unnoticed by the annotators themselves. as the set of annotated data grew, we were able to examine overall patterns in the use of annotation categories between annotators. this higher-level examination of annotation disagreements worked well for uncovering inconsistencies in annotation due to, for example, overor under-use of a particular annotation category by one annotator. these patterns of disagreement could then become the topic of discussion at future annotation meetings. 5 analysis in this section, we present quantitative analysis of social act prevalence and usage among discussion participants with different roles (such as wikipedia administrators, or accountants in the chat scenario). in the wikipedia discussion corpus, we also analyze differences in authority self-presentation among participants with different levels of experience within the editor community, as measured by v-index, months editing, and total edits. while these analyses are meant to be illustrative rather than definitive, they suggest intriguing associations between the self-presentation of authority and interpersonal alignment, and between these social acts and other 21 are we there yet? features of discussions and discussants such as social role and experience level. the analyses below are based on a subset of the full aawd corpus released by the project. specifically, using the version of the files provided in the "merged" directory, we used only files which had been annotated by two annotators (working independently) and counted as authority claims/alignment moves only those that were identified as such by both annotators (where the annotators furthermore agreed on the type). 5.1 differences in expression of alignment across the wikipedia samples for the three languages speakers of different languages express agreement and disagreement in different ways (see for example mori 1999). during the process of adapting our english language annotation guidelines to mandarin and russian, we identified several potential differences in the expression of alignment, which we addressed by altering our guidelines. during weekly meetings with our annotators, who were all bilingual in english and their native language, we asked them to reflect on other potential differences between the expression of alignment in those two languages. based on their observations, as well as the observation of the researchers who led the annotation process in mandarin and russian, we formed several preliminary hypotheses regarding the ways in which mandarin and russian wikipedia editors differ from english-speaking editors in some aspects in their language uses. for example, hypothesis (ii) below is based on recurring observations by annotators that mandarin wikipedia editors tended to not use any explicit negators, but instead used paired conjunctions to indicate the negation. these preliminary hypotheses and findings are not definitive and are provided to illustrate opportunities for further research. hypotheses: ( i ) negative alignment is more likely to be implicit in mandarin than in english or in russian. ( ii ) mandarin speakers are more likely to use paired conjunctions and partial agreement to express negative alignment. ( iii ) mandarin and russian speakers are less likely than english speakers to use the names of their addressees in alignment moves. ( i ) negative alignment is more likely to be implicit in mandarin than in english or in russian this hypothesis is supported by our data. when mandarin speakers contradict someone else, they are less likely to use explicit disagreement words; instead, they tend to indirectly explain what they believe to be the fact without saying “no/not” or “disagree.” in a sample of 100 mandarin negative alignment moves and 100 english negative alignment moves, we found that 75 english negative alignment moves contained either “no” or “not,” among which 14 negative alignment moves started with “no.” also, in english, 7 out of 100 negative alignment moves contained “disagree.” in mandarin, however, only 20 out of 100 negative alignment moves contained “不 是 no/not,” among which 5 moves started with “不是 no/not.” no negative alignment moves contained “不同意 disagree.” these results support the hypothesis that mandarin speakers in this dataset use less explicit negative utterances when they are disagreeing with others. in other words, mandarin speakers negate their statements more indirectly. the same hypothesis was tested in russian, and it was found that in our data the usage of explicit agreement markers in russian is much higher than for mandarin, but lower than for english. 258 negative alignment examples in russian were analyzed. of those examples, 153 moves (or 59.3%) were made using explicit disagreement markers, such as “не/no,” “ни/no” (treated by annotators as a marker of dismissal), “нет/no,” and “я не согласен/i do not agree.” 22 morgan et al. ( ii ) mandarin speakers are more likely to use paired conjunctions and partial agreement to express negative alignment chinese frequently employs paired subordinate conjunctions, where the subordinate conjunction introduces the subordinate clause and another discourse connective introduces the main clause. these connectives are generally clause-initial or clause-medial (xue 2005). examples of paired conjunctions in mandarin include: “虽然…但是… although…but…,” “即使…也… even though…still,” “不是…而是.. not…but,” “不但…而且 not only…but also…,” “不是…也不是 …neither…nor,” “是…但… true it is…, but…” below are two examples which were tagged as negative alignment in the sample files: example 5. (b is annotated as negative alignment.) a: 你 用 “简体夷字”, 那便 是 you use “simplified tribal characters”, then (you) are 夷人, 不配 是 汉 人 tribal man, not qualified as a han people. ‘you use simplified tribal characters, then you are a barbarian, not an authentic chinese.’ b: 虽然 我 使用 简体字, 但是 我 赞成 although i use simplified characters, but i support 恢复 正体 字。 revive orthodox characters. ‘although i use the simplified characters, i support the revival of the traditional chinese characters.’ example 6. (b is annotated as negative alignment): a: 既然 我们 汉族的 疆域 主要 是靠 战争 since our han’s territory mainly depends on wars 与 征服 而来, 我们 有 什么 资格 and conquering come from, we have what qualification 指责 少数民族 夺权 就是 criticize the minority’s seizing the power as “非正义”呢? “injustice”? ‘since china’s territory mainly comes from wars and conquering, how can we be qualified to condemn that the minority’s power seizure is injustice?’ 23 are we there yet? b: 即使 是 武装 战争,   如果 是 为 取得 更多的 even if (it) is armed war, if is for get more 国土 land, 仍然 是 侵略 战争,   也 很难说 就是 正义的。 still is invading war, still very hard to say as justice. ‘even if it is an armed war, if it is aimed for getting more land, it is still invading, and it is hard to say it is justice.’ we hypothesized that when making negative alignment moves, mandarin speakers would tend to use more paired conjunctions and partial agreement than english speakers in our data set. partial agreement means that agreement is made with reservation especially when there is doubt or feeling of not being able to accept something completely. a commonly used partial agreement structure in mandarin is “adj is adj, but...,” meaning the speaker admits one aspect of the good quality, but denies another, with the focus usually on the latter. while the differences are not as stark as for hypothesis (i), the data do provide support for this hypothesis as well. in a sample of 100 mandarin negative alignment moves, there are a total of 9 moves (9%) that contain partial agreement and 19 moves (19%) that contain paired conjunctions. on the other hand, in a sample of 100 english negative alignment moves, there are 5 moves (5%) that contain partial agreement and 4 moves (4%) containing paired conjunction (specifically, sentences with the structure “ … not only … but (also) ...”). ( iii ) mandarin and russian speakers are less likely than english speakers to use the names of their addressees in alignment moves we hypothesized that mandarin speakers would tend to use the names of their addressees in alignment moves less frequently than english speakers do. the statistics show, however, that english and mandarin wikipedia editors are very close with each other in terms of using names when making alignment moves. in a sample of 170 mandarin alignment moves, there are 26 alignment moves that contain direct names of their addressees, which accounts for 15.29%. on the other hand, in a sample of 200 english alignment, there are 26 alignment moves that contain direct names of their addressees, which accounts for 13%. pronouns are excluded in the calculation. russian speakers, however, included the names of their addressees less often when aligning. out of 258 negative alignment moves in russian, only 4 were made with a specific addressee name, which constitutes just 1.55%. compared to 13% in english or 15.29% in mandarin, this number suggests that russian speakers tend to omit specific addressee names when disagreeing with another wikipedia participant. moreover, the same pattern was found in positive alignment in russian. out of 79 positive alignment moves, only 1 was made using a specific addressee name, which constitutes 1.26% of all positive alignment moves. this data suggests that in both types of alignment, negative and positive, russian speakers usually do not include the name of the person they are agreeing or disagreeing with. alignments made without a specific addressee name also had some practical implications for the annotation process across all languages, i.e., it was difficult for the annotators to identify the target of (dis)agreement (the person with whom the speaker was (dis)agreeing) in the absence of an addressee name. 24 morgan et al. 5.1.5 summary while we were able to define social acts in a way that allowed us to annotate them across languages and thus explore differences in both their distribution and realization across those languages, these differences should not be taken as direct evidence for specific differences between national cultures. wikipedia draws editors from across the globe, and the languages we are working with are spoken in communities in many different nations, both as a native language and as a second language. more generally, it would be overly simplistic to generalize from the differences in expression of social acts that we found to differences between cultures. rather, the differences in social acts merely suggest that there is room to explore what cultural conventions might help shape the ways in which these social acts are carried out. 5.2 comparison of social act expression in english wikipedia and irc two researchers qualitatively examined the english language aacd dataset, taking notes on potential differences between the aacd and aawd english datasets as well as other salient medium variables such as turn length and conversation structure. the researchers then met to share notes and discuss their findings, and identify possible high-level trends. see table 5 (section 3.2 above) for a breakdown of authority claims and alignment moves in the english language irc data.8 below we share several observations that illustrate the ways in which genre may interact with the presence and form of social acts between the chat and wikipedia data, and which could be productively investigated in future work. these findings also illustrate some of the challenges posed by differences in conversation structure on the application of coding schemes and annotation processes developed on one genre to another. 5.2.1 alignment tends to be explicit in irc data unsurprisingly given the limited overall timeframe and synchronous nature of the chat genre, irc chat participants tended to take shorter turns and to express alignment with other participants more explicitly and with less extensive argumentation than wikipedia editors. 5.2.2 alignment moves are common in irc data however, overall shorter turns, more-rapid turn-taking, and more explicit agreement/disagreement in irc can make it difficult to distinguish ‘true’ alignment from backchannels (such as “yeah”). rapid turn-taking and a lack of nested turn threading also makes it difficult to identify alignment targets. 5.2.3 negative alignment is less prevalent in irc data than in wikipedia data the relative infrequency of negative alignment moves in irc may reflect “facework” considerations, as the participants in our study often knew each other offline, and even when they did not, the fact that participants met face-to-face before and after the chat session may have made them less willing to be seen as disagreeable. in addition, the artificial nature of the task scenario may have made them less invested in a particular outcome and therefore less likely to dispute others’ proposals.                                                                                                                 8  comparative analysis of alignment and authority claims in the russian and mandarin data is beyond the scope of this paper, but represents an intriguing opportunity for future research.     initial turn alignment in next 10 turns no authority claim 0.52 any authority claim 0.63  table 10. average number of alignment moves targeted at participant in 10 following turns. 25 are we there yet? 5.2.4 authority claims are comparatively rarer in irc data authority claims are rarer on irc than on wikipedia, but are also harder to identify because they’re often less formally structured. claims may be incomplete or partially implied, or spread across multiple turns. for instance “safeway has grapes for 80 cents a pound” could be an authority claim, or just an observation. it becomes hard to tell without including evidence from the speaker’s previous and subsequent turns. 5.2.5 authority claim distribution is different in irc data and wikipedia data the distribution of authority claim types differs between wikipedia and irc discussions, in a way that likely reflects the medium (less time to craft a complex empirical or logical argument) and the genre (a short term collaboration among immediate peers, with less at stake) and the artificiality of the scenario (assigned roles, made-up task). our english irc data exhibits a higher proportion of experiential claims, and social expectations claims (for instance, when a participants assert which kind of pizza their fellow students will prefer). 5.3 interactions between authority and alignment on wikipedia thus far, we have been addressing our social acts independently, but of course no social act occurs in a vacuum. alignment moves and authority claims are only two types of social acts; many other types of social acts are present (and could be annotated) in this same data set. even with only these two types (and their subtypes), however, we find interactions. we hypothesized that authority claims would be likely to provoke alignment moves.9 that is, although participants may make alignment moves whenever someone else has expressed an opinion or taken action (e.g., edited the article attached to the discussion), we hypothesized that by making an authority claim, a participant becomes more likely to become a focal point in the debate. to test this, we calculated, for every turn, the number of alignment moves targeted at the author of that turn within the next 10 turns. we then divided the turns into those that contained authority claims and those that did not. making an authority claim in a given turn made the participant significantly more likely to be the target of an alignment move within the subsequent 10 turns compared to turns that did not contain any claims (students’ t-test, t=-2.086, df=772, p=.037; table 10)) furthermore, we find that different types of authority claims elicit different numbers of subsequent alignment moves. specifically, turns that contain either external claims or forum claims (the two most prevalent claim types in our sample) interact differently with alignment. external claims elicited more alignment overall (students’ t-test, t=3.189, df=411, p=.002) and more negative alignment moves than did forum claims (students’ t-test, t=3.839, df=415, p<.001). however, external claims did not elicit significantly more positive alignment moves than forum claims (students’ t-test, t=0.695, df=309, p=.488). this is illustrated in table 11. 5.4 social acts and identity work given the few visible markers of status on wikipedia and the fact that editors are constantly interacting with new collaborators, wikipedians perform authority by adopting insider language                                                                                                                 9  the findings presented in sections 5.3 and 5.4 are slightly revised and expanded from bender (2011).   initial turn positive negative overall external authority claim 0.26 0.49 0.74 forum authority claim 0.22 0.2 0.42   table 11. average number of positive, negative and overall alignment moves targeted at claim-making participant in 10 following turns. 26 morgan et al. and norms of interaction. supporting arguments with specific references is one such norm. thus we hypothesized that as editors become more integrated into wikipedia, they will make more authority claims. several possible metrics could be used to represent an editors’ level of integration into wikipedia. in order to test our hypothesis that more integrated editors would make authority claims with greater frequency, we analyzed our data using several of these measures. first, we analyzed whether editors with different official roles (anonymous editors, registered editors, and administrators) exhibited different degrees of claim-making. we then analyzed whether total edits by an editor (under a particular user name), the number of months an editor had been active on wikipedia, or our v-index measure showed a positive correlation with claim-making. we present our rationale for developing the v-index measure, our sampling criteria, and the findings from our evaluation of v-index, total edits and months active in section 5.4.2 below. 5.4.1 authority claim types by user status wikipedia distinguishes three different statuses: unregistered users (able to perform most editing activities, identified only by ip address), registered users (able to perform more editing activities, edits attributed to a consistent user name) and administrators (registered users with additional ‘sysop’ privileges). participants of different statuses tend to do different kinds of work on wikipedia, with administrators in particular being more likely to take on moderator work (burke and kraut 2008), such as mediating and diffusing disputes among editors. because conflict mediation requires a different kind of credibility than collaborative writing work, and because unregistered users are likely to be newer and therefore less likely to be incorporating references to wikipedia-specific rules and norms into their projected identities (and, therefore, their conversation), we hypothesized that editors of different statuses would use different kinds of authority claims. this is borne out in our data. while no user group was significantly more or less likely than any other to include authority claims overall in their posts (chi square test for independence, n=3164, df=2, χ2=2.367, p=.306) users of different statuses did use significantly different proportions of forum claims and external claims (chi square test for independence, n=973 turns, df=8), which were the most frequent claim types in the sample overall. table 12 presents a breakdown of the claim-making behavior of different user types. 5.4.2 authority claims by level of experience evaluating the level of experience of a particular wikipedia editor presents numerous challenges. although registered editors (who are identified by their username) are often more experienced than unregistered editors (who are identified by their ip address), this is not a given: a veteran wikipedian who has thousands of edits’ worth of editing experience may edit a page while not logged in. and it is possible for an editor to work for years on wikipedia without ever creating a username. status as an administrator is generally thought to be an unambiguous signal of editing experience, since administrators are ‘elected’ by the community in recognition of extensive work. however, administrators make up only a small fraction of all registered wikipedia editors, and many non-administrators have comparable levels of experience, but never apply for adminship. other measures based on community recognition of quality work, such as ‘barnstars’ (kriplean et # editors % forum claims % external claims % turns with claims user type 44 47.1 45.1 19.6 administrator 192 29.1 63.6 22.3 registered 55 18.3 70.6 19.8 unregistered 291 29.8 62.5 21.6 all types   table 12. percent of turns by different types of wikipedia editors that contain authority claims. 27 are we there yet? al. 2008) may serve as robust signals of experience, reputation or investment, but are variably distributed and can also be hard to interpret. two measures that are commonly cited by wikipedia community members as indicators of experience or status are the number of months (or years) an editor has been active, and the total number of edits (to articles, talk pages, policy pages, etc.) an editor has made. however, these measures may not accurately reflect an editor’s level of participation. for instance, is an editor of five years and 20,000 edits who has not made an edit since 2008 as invested as an editor who joined in 2011, but has since made 3,000 edits? however, length of ‘tenure’ is still important: it takes time to integrate into the community and become a “wikipedian,” due to both the technical complexity of the software and the dizzying variety of rules and conventions (butler et al. 2008). the v-index score is designed to account for an editor’s level of investment within the community at a specific point in time by taking into account recent edits over recent months. the longer an editor has demonstrated a high level of sustained participation, the higher their v-index will be. if they are generally a low-level participant their v-index be lower, and if they were a highly active in the past, but their recent participation shows a decrease or has become erratic, their v-index will drop. we therefore hypothesize that v-index will show a stronger correlation with claim-making behavior than either the number of months since the editor joined wikipedia or the total number of edits by that editor, because it more accurately reflects an editor’s level of engagement and expertise at the point at which they are making a particular utterance. to evaluate the claim that v-index is a better measure of engagement than simple edits or time counts, we replicated the v-index finding from bender (2011), and then calculated claimfrequency by total months active and total edits for comparison. we assigned a v-index, monthsediting and total-turns value to every turn in our english dataset made by a registered editor or an administrator.10 then we sorted these turns into buckets with one bucket for each v-index and month editing, and one bucket per 100 edits. this resulted in 40 v-index buckets, with a top value of 46; 53 months-editing (as 28-day periods) buckets with values with a top value of 58, and 223 total-edits buckets, with a top value of 1580, indicating that one editor in our sample had made over 150,000 edits.11 for each of the three measures, the number of turns per bucket value declined rapidly and steeply, although at different rates. in order to assure an adequate sample size for each bucket, and to avoid a single editor’s claim-making behavior disproportionately influencing the correlation among the higher bucket values (which were represented by far fewer turns), we set a threshold for each dataset at the last bucket (ascending) where total turns was greater than 50 and total unique editors represented by that bucket was greater than 20. our data are presented in table 13. we performed a one-sided spearman’s rho rank correlation on each of our three metrics against the percentage of turns that contained any type of authority claim. v-index showed a strong, significant positive correlation with claim-making, confirming our hypothesis (one-sided,                                                                                                                 10 see appendix a for v-index sampling considerations and potential sources of error. 11 as of february 29th, 2012 the top lifetime edit count for a single editor was 968,000. total edits (log) total turns turns with claims % claim turns 1 78 14 18% 10 1178 263 22% 100 1419 288 20% 1000 234 54 23% 10000 9 0 0%  table 13. proportions of claim-bearing turns by participants with different edit counts, log scale. 28 morgan et al. correlation coefficient = .596, n=14, p=0.012). total months editing did not show a significant correlation (correlation coefficient = .253, n=14, p=0.192). total edits showed a marginally significant correlation (correlation coefficient = .533, n=9, p=0.07). because the results from total edits seemed suggestive, we calculated these values on a log scale as well, with buckets for editors with 1-10 edits, 11-100 edits, 101-1000 edits and 100110,000 edits. while this sample is too small to yield a significant correlation, the results (table 13) do not show a clear increase in the number of claim-bearing turns for editors with higher edit counts. we present these results merely to illustrate that for these data the way samples are grouped can affect trends observed. we also caution that choosing a higher or lower “cutoff” value for these sample buckets, or using an alternate statistic, may affect the significance of the resulting correlation. while additional studies, ideally correlating v-index to other kinds of editor behavior, would be required to establish v-index as a reliable measure of editor engagement, we find these initial results promising. 6 conclusion we have presented the authority and alignment in wikipedia discussions (aawd) corpus, a collection of 365 discussions drawn from wikipedia talk pages and annotated for two broad types of social acts: authority claims and alignment moves. these annotations make explicit important discursive strategies that discussion participants use to construct their identities in this online forum. that “identity work” is being done with these social acts is confirmed by the correlations we find between proportions of turns with authority claims and external variables such as user status and v-index, on the one hand, and the interaction between authority claims and alignment moves on the other. as an example of a social medium, wikipedia is characterized by its taskorientation and by the fact that all of the participants’ “identity work” with respect to their identity in the medium is captured in the database. this, in turn, causes the data set to be rich in the type of social acts we are investigating. though this data set is small compared to many that are used in machine learning, it has already been used in research on the automatic detection of forum claims. (marin et al. 2011). that work focused on using lexical features, filtered through word lists obtained from domain experts and through data-driven methods, and extended with parse tree information.12 we hope to see similar approaches applied to the automatic detection of other types of authority claims and of alignment moves in future. we have also described and reflected on the iterative process through we developed our annotation guidelines. the original drafts of the guidelines were developed on the basis of an initial pass through sample data paired with theoretical reflections, and attempted to map out the space of possible variations within the social act types were annotating. these guidelines were then used by other annotators to annotate more data. measuring inter-annotator agreement and examining specific cases of disagreement led us to tighten up the guidelines. in general, our strategy was to make the guidelines more restrictive rather than to cast a wider net, based on our observation that “core” or “prototypical” examples of our social acts were easier for the annotators to recognize and agree on. this was driven in part by the fact that our field uses interannotator agreement as a measure of consistency of annotations and consistency of annotations in turn as a proxy for the degree to which annotations represent ground truth. our experience annotating social acts brings into focus and problematizes this proxy relationship: by tightening our guidelines in order to achieve better consistency, it could be argued that we increased the number of false negatives (unlabeled social acts) in our annotated corpus. on the other hand, as with many annotation projects, our labels did not have well-established a priori definitions. thus                                                                                                                 12 an anonymous reviewer notes that the aawd and aacd corpora are not large enough for certain approaches to machine learning. one of our goals in publishing our code books along with the data we annotated was to enable future work extending the annotations to larger collections of text. 29 are we there yet? one product of this annotation effort has been the creation of the social act typologies themselves, which, of necessity, focused on overt, explicit variants within a larger space. while similar observations likely pertain to many types of annotation efforts, we believe they are particularly prominent in this domain because we are working at such remove from the linguistic structure of the utterances we are annotating. we believe that, as social acts, authority claims and alignment moves are broadly recognized communication behaviors that play an important role in human interaction across a variety of contexts. however, we expect that the distribution and presentation of our social acts be manifestly different in online genres other than wikipedia discussions and internet relay chat. wikipedia discussions are shaped by a set of well-defined, local communication norms that are closely tied to the task of distributed, collaborative writing and the culture of open-source software. internet relay chat and related instant message protocols present their own constraints, and furthermore a comparison between our facilitated discussions and ‘organic’ irc data would help shed light on the ways our protocol and testing environment may have shaped the discourse. future work could explore the range of variation among the linguistic cues associated with authority and alignment categories across genres, cultures and communication media, as well as the possible role of additional categories or social acts not discussed here. we hope that the online communication genres captured in the aawd and aacd corpora prove to be valuable resources for social scientific analyses of communication behaviors as well as a resource for the development of nlp systems which can automatically identify these social acts, on wikipedia and beyond. acknowledgments this research was funded by the office of the director of national intelligence (odni), intelligence advanced research projects activity (iarpa). all statements of fact, opinion or conclusions contained herein are those of the authors and should not be construed as representing the official views or policies of iarpa, the odni or the u.s. government. we wish to thank our colleagues: brian hutchinson, alex marin, mari ostendorf, and bin zhang. we also gratefully acknowledge the contribution of the annotators: wendy kempsell, kelley kilanski, robert sykes and lisa tittle. we also wish to thank the anonymous reviewers for their valuable feedback on this manuscript. the original wikipedia discussion page data for this study was made available from a research project supported by nsf award iis-0811210. we thank travis kriplean for his initial assistance with scripts to process this data dump. appendix a: sources of inconsistency in computing v-index the basic algorithm for calculating the v-index values and the data included in the aawd corpus are the same as that used in bender et al. (2011). however, the calculation of v-index requires access to the rest of wikipedia as well. the v-index values reported in bender et al. (2011) were calculated against the 2008 wikipedia snapshot. in this paper, we instead use a live mirror of the wikipedia database. as v-index values only reflect revisions made before the turns in question, this should in principle lead to the same results. however, the live database differs from the snapshot in that revisions that have been permanently deleted from wikipedia after the snapshot was taken are no longer reflected in the database. because of wikipedia’s built-in version control system, any content ever entered into wikipedia is perpetually available and potentially viewable by default. employees of the wikimedia foundation (but not individual editors) may therefore occasionally completely delete individual revisions of a page if making it available may have legal ramifications (for instance contain potentially libelous content), or revisions which contain particularly sensitive information (such as a users social security 30 morgan et al. number).13 while this is a comparatively rare occurrence, one of the discussions we annotated (comprising 58 turns) was associated with such a deleted page. we were able to calculate v-index values for the turns in this deleted discussion by looking up the timestamp of those turns which were available to us because one of the authors (a contractor with wikimedia) possessed special data access privileges. however, there may still be slight differences as the deletion of other pages not captured in our database may have lowered the v-index value for some proportion of our turns. in addition, in the 2011 work, we calculated v-index values for unregistered users by treating all edits from the same ip address as belonging to the same user. the mapping from ip addresses to users is not reliable, however, and accordingly we have chosen to exclude turns by unregistered users in the current analysis. this removes 443 turns from consideration. we did not evaluate potential correlations between v-index, months editing and total edits and authority claims for russian and mandarin because of a significant reduction in sample size. references apoorv agarwal and owen rambow. 2010. automatic detection and classification of social events. in proceedings of the 2010 conference on empirical methods in natural language processing (emnlp '10), association for computational linguistics, pages 1024-1034, stroudsburg, pennsylvania. mats alvesson and hugh willmott. 2002. identity regulation as organizational control: producing the appropriate individual. journal of management studies, 39(5):619-644. ron artstein and massimo poesio. 2008. inter-coder agreement for computational linguistics. computational linguistics, 34(4):555–596. collin f. baker, charles j. fillmore and john b. lowe. 1998. the berkeley framenet project. in proceedings of 36th annual meeting of the association for computational linguistics and 17th international conference on computational linguistics, association for computational linguistics, pages 86-90, stroudsburg, pennsylvania. philip ball. 2005. index aims for fair ranking of scientists. nature, 436:900–900. nancy baym. 1996. agreements and disagreements in a computer-mediated discussion. research on language and social interaction, 29:315–345. emily m. bender, jonathan t. morgan, meghan oxley, mark zachry, brian hutchinson, alex marin, bin zhang, and mari ostendorf. 2011. annotating social acts: authority claims and alignment moves in wikipedia talk pages. in proceedings of the acl-hlt workshop on language in social media (lsm 2011), pages 48-57, portland, oregon. penelope brown and stephen c. levinson. 1987. politeness: some universals in language usage. cambridge university press, cambridge. mary bucholtz and kira hall. 2010. locating identity in language. in carmen llamas and dominic watt, editors, language and identities. edinburgh university press, edinburgh. moira burke and robert kraut. 2008. mopping up: modeling wikipedia promotion decisions. in proceedings of the 2008 acm conference on computer supported cooperative work, association of computing machinery, pages 27–36, san diego, california. brian butler, elisabeth joyce, and jacqueline pike. 2008. don't look now, but we've created a bureaucracy: the nature and roles of policies and rules in wikipedia. in proceedings of the twenty-sixth annual sigchi conference on human factors in computing systems (chi'08), association of computing machinery, pages 1101-1110, new york, new york.                                                                                                                 13 see http://en.wikipedia.org/wiki/wikipedia:office_actions 31 are we there yet? robert e. cummings. 2008. what was a wiki, and why do i care? a short and usable history of wikis. wiki writing: collaborative learning in the college classroom, pages 1-16. university of michigan press, ann arbor. jolene galegher, lee sproull, and sara kiesler. 1998. legitimacy, authority, and community in electronic support groups. written communication, 15(4):493–530. meghan lammie glenn, stephanie m. strassel, and haejoong lee. 2009. xtrans: a speech annotation and transcription tool. in proceedings of interspeech 2009, pages 2855–2858, brighton, uk. erving goffman. 1959. the presentation of self in everyday life. doubleday, garden city, new york. erving goffman. 1981. forms of talk. university of pennsylvania press, philadelphia. jakob l. jensen. 2003. public spheres on the internet: anarchic or government sponsored; a comparison. scandinavian political studies, 26(4):349–374. travis kriplean, ivan beschastnikh, david w. mcdonald, and scott a. golder. 2007. community, consensus, coercion, control: cs*w or how policy mediates mass participation. in proceedings of the 2007 international acm conference on supporting group work (group '07), association of computing machinery, pages 167-176, new york, new york. travis kriplean, ivan beschastnikh, and david w. mcdonald. 2008. articulations of wikiwork: uncovering valued work in wikipedia through barnstars. in proceedings of the 2008 acm conference on computer supported cooperative work (cscw '08), association of computing machinery, pages 47-56, new york, new york. j. richard landis and gary g. koch. 1977. measurement of observer agreement for categorical data. biometrics, 33(1):159–174. yameng liu. 1997. authority, presumption and invention. philosophy and rhetoric, 30(4):413– 427. john locke. 1959 [1690]. an essay concerning human understanding. dover publications, new york. jo mackiewicz. 2010. assertions of expertise in online product reviews. journal of business and technical communication, 24(1):3–28. kathleen m. macqueen, eleanor mclellan, kelly kay, and bobby milstein. 1998. codebook development for team-based qualitative analysis. cultural anthropology methods, 10(2):3136. mitchell p. marcus, beatrice santorini, and mary ann marcinkiewicz. 1993. building a large annotated corpus of english: the penn treebank. computational linguistics, 19(2):313–330. alex marin, bin zhang, and mari ostendorf. 2011. detecting forum authority claims in online discussions. in proceedings of the acl-hlt workshop on language in social media (lsm 2011), pages 39-47, portland, oregon. junko mori. 1999. negotiating agreement and disagreement in japanese: connective expressions and turn construction. john benjamins publishing company, amsterdam. michael mulkay. 1985. agreement and disagreement in conversations and letters. text, 5(3):201– 227. michael mulkay. 1986. conversations and texts. human studies, 9(2-3):303–321. meghan oxley, jonathan t. morgan, mark zachry, and brian hutchinson. 2010. “what i know is...”: establishing credibility on wikipedia talk pages. in proceedings of the 6th 32 morgan et al. international symposium on wikis and open collaboration, association for computing machinery, article 26, 2 pages, new york, new york. martha palmer, daniel gildea, and paul kingsbury. 2005. the proposition bank: an annotated corpus of semantic roles. computational linguistics, 31(1):71-105. martin j. pickering and simon garrod. 2004. the interactive-alignment model: developments and refinements. behavioral and brain sciences, 27(2):212-225. janie rees-miller. 2000. power, severity, and context in disagreement. journal of pragmatics, 32(8):1087–1111. kay richardson. 2003. health risks on the internet: establishing credibility online. health, risk and society, 5(2):171–184. john r. searle. 1975. indirect speech acts. in peter cole and jerry l. moran, editors, syntax and semantics volume 3: speech acts, pages 59-82. academic press, new york. marie-claire shanahan. 2010. changing the meaning of peer-to-peer? exploring online comment spaces as sites of negotiated expertise. journal of science communication, 9(1):1–13. clay shirky. 2008. here comes everybody: the power of organizing without organizations. penguin press, new york. elizabeth shriberg, raj dhillon, sonali bhagat, jeremy ang, and hannah carvey. 2004. the icsi meeting recorder dialog act (mrda) corpus. in proceedings of the 5th sigdial workshop on discourse and dialogue at hlt-naacl 2004, association for computational linguistics, pages 97-100, cambridge, massachusetts. benjamin snyder and martha palmer. 2004. the english all-words task. in proceedings of senseval-3, the third international workshop on the evaluation of systems for the semantic analysis of text, association for computational linguistics, pages 41-43, barcelona, spain. jan svennevig. 1999. getting acquainted in conversation: a study of initial interactions. john benjamins publishing company, amsterdam. györgy szarvas, veronika vincze, richárd farkas, and jános csirik. 2008. the bioscope corpus: annotation for negation, uncertainty and their scope in biomedical texts. in proceedings of the workshop on current trends in biomedical natural language processing, association for computational linguistics, pages 38-45, stroudsburg, pennsylvania. dorothea k. thompson. 1993. arguing for experimental “facts” in science. written communication, 10(1):106. fernanda b. viegas, martin wattenberg, jesse kriss, and frank van ham. 2007. talk before you type: coordination in wikipedia. in proceedings of the 40th annual hawaii international conference on system sciences (hicss '07), page 78, ieee computer society, washington, d.c. linda wine. 2008. towards a deeper understanding of framing, footing, and alignment. columbia university working papers in tesol and applied linguistics, teachers college, 8(3):1–3. nianwen xue. 2005. annotating discourse connectives in the chinese treebank. in proceedings of the workshop on frontiers in corpus annotations ii: pie in the sky, pages 84-91, ann arbor, michigan. 33 dialogue and discourse 4(2) (2013) 249-281 doi: 10.5087/dad.2013.211 ©2013 maite taboada and debopam das submitted: 04/12; accepted: 02/13; published online: 08/13 annotation upon annotation: adding signalling information to a corpus of discourse relations maite taboada mtaboada@sfu.ca debopam das ddas@sfu.ca department of linguistics simon fraser university 8888 university dr. burnaby, b.c. v5a 1s6 canada editors: stefanie dipper, heike zinsmeister, bonnie webber abstract we present an annotation effort that involves adding a new layer of annotation to an existing corpus. we are interested in how rhetorical relations are signalled in discourse, and thus begin with a corpus already annotated for rhetorical relations, to which we add signalling information. we show that a very large number of relations carry signals that can help identify them as such. the detailed, extensive analysis of signals in the corpus can aid research in the automatic parsing of discourse relations. keywords: discourse relations, corpus annotation, discourse markers 1 introduction one of the most frequent tasks that corpus and computational linguists perform is to re-use and re-annotate existing resources. although many valuable annotated corpora exist, they often do not contain all the information and detail that is necessary in every research project. thus, researchers are left with the need to add information to a corpus that has already been annotated in some form or another. starting from scratch may not be optimal, since one can build on existing annotations, unsatisfactory as they may be. this is particularly the case with higher-level annotations, those “beyond semantics”, because they tend to rely on annotations at lower levels of discourse, such as semantic role annotations relying on part of speech tags. in addition, many annotation efforts are conceived as layers of different kinds of information, sometimes added by different annotators (see stede, 2007; cunningham et al., 2011 for examples of a general philosophy of layered text annotation). in this paper, we present an annotation effort that involves adding a new layer of annotation to an existing corpus. we are interested in how rhetorical relations are signalled in discourse, and thus begin with a corpus already annotated for rhetorical relations, to which we add signalling information. the issue of signalling is central in research on discourse relations. identification and classification of relations often hinge on pinpointing lexical or other cues that indicate a relation taboada and das 250 is present, with some approaches to coherence relations relying exclusively on signals (mostly discourse markers) to classify relations (sanders et al., 1992; knott & sanders, 1998). in more applied areas, signals are used to help identify relations in applications such as discourse parsing and summarization (marcu, 2000a; schilder, 2002; hanneforth et al., 2003; polanyi et al., 2004; sporleder & lascarides, 2005; baldridge et al., 2007; afantenos et al., 2010). more generally, the issue of signalling in discourse relations needs to be examined from a processing point of view. if we assume that coherence relations are cognitive entities, then we need to find how hearers and readers are able to identify them on the basis of linguistic cues. successful communication must be based on a relatively unambiguous interpretation of relations, for which clear signals are necessary. most psycholinguistic research on this matter to date has focused on one particular type of signal, the presence of discourse markers. in order to understand how relations are processed, and in order to extract them automatically, we need to move beyond signalling by discourse markers, as those seem to be present in only a small fraction of the relations found in corpora (taboada, 2006, 2009). we believe that the first step in this endeavour is to annotate discourse with an open mind to other types of signalling. the only other available resource that contains signalling information, the penn discourse treebank (prasad et al., 2008), contains mostly discourse markers 1 as signals. although the annotation is very detailed and useful, it does not include all of the types of signals that we believe are indicative of rhetorical relations. thus, in this paper, we begin with a corpus already annotated for coherence relations, to which we are adding information on how the relations are signalled, including a variety of possible signals. we begin the paper by briefly discussing coherence relations and their signalling. then we propose a classification of signalling devices, which we use to annotate a corpus. we discuss the corpus annotation, issues with reliability, and the particular types of problems that are associated with annotating discourse phenomena. the corpus annotation reveals a broad spectrum of signalling devices. the paper concludes with some lessons learned from the annotation, and the applications that the corpus will have. 2 coherence relations there are many theories of discourse, rhetorical, or coherence relations, but we believe they all refer to fundamentally a similar phenomenon: relations among propositions, which are the building blocks of discourse, and help explain coherence. although we have worked within rhetorical structure theory (mann & thompson, 1988), and will use some of its constructs here, the discussion that follows likely applies to any view of coherence relations. in rst, relations are defined through different fields, the most important of which is the effect, the intention of the writer (or speaker) in presenting their discourse. relation inventories are open, and the most common ones include names such as cause, concession, condition, elaboration, result or summary. relations can be multinuclear, reflecting a paratactic relationship, or nucleus-satellite, a hypotactic type of relation. the names nucleus and satellite refer to the relative importance of each of the relation components. texts are then built out of basic clausal units that enter into rhetorical relations with each other, in a recursive manner. mann and thompson proposed that most texts can be analyzed in 1 in the penn discourse treebank, relations are also annotated as being signalled by indicative phrases. these relations are known as altlex (alternative lexicalization) relations (prasad et al., 2010). signalling in a corpus of discourse relations 251 their entirety as recursive applications of different types of relations. in effect, this means that an entire text can be analyzed as a tree structure, with clausal units being the branches and relations the nodes. in figure 1 we present an rst analysis from the rst discourse treebank (carlson et al., 2002), the corpus that we have chosen to annotate. in it, we can see the text divided into units, or spans, and how rhetorical relations hold across spans. in this case, all the relations are nucleussatellite, with relations embedded throughout the example. the analysis itself may be questioned in terms of standard rst practice. for instance, unit 4 should probably not be considered a span, and instead included as a unit with the noun that it modifies (amount). we are, however, working with an existing annotation, and will use the relations in the corpus as they are. figure 1. sample rst analysis from the rst discourse treebank there has been a long and lively debate about how coherence relations, interpreted as rhetorical relations in rst or in other theories (e.g., polanyi & scha, 1983; sanders et al., 1993; asher & lascarides, 2003), are recognized and interpreted, that is, their cognitive status: are relations present in the minds of speakers and hearers 2 or are they analysis constructs? the former postulates that coherence relations are part of the process of constructing a coherent text representation. in rhetorical structure theory (mann & thompson, 1988), the relations are postulated as being recognizable to an analyst, and in general to a reader. the process is one of uncovering the author’s intention in presenting pieces of text in a particular order and combination. in carrying out an rst analysis of a text, “the analyst effectively provides plausible reasons for why the writer might have included each part of the entire text” (mann & thompson, 1988: 246). but further cognitive claims have not been strong within rst. support for the cognitive status of coherence relations comes from experimental work on the effect of particular types of relation on text comprehension. sanders and colleagues have best 2 we will use speakers/hearers and writers/readers interchangeably. it is arguably the case that most of what can be said about coherence relations applies equally to spoken and written discourse. indeed, if we postulate psychological validity, both forms of discourse must be accounted for. taboada and das 252 articulated this view. in knott and sanders (1998), they argue that text processing consists of building a representation of the information contained in the text. part of the process of building involves integrating individual propositions in the text into a whole. coherence relations model the ways in which propositions are integrated. the evidence presented comes from two different sources. first of all, studies have shown differences in processing different types of relations, mostly causal versus non-causal (e.g., trabasso & sperry, 1985; and references in knott & dale, 1994). secondly, the presence of connectives indicating coherence relations tends to facilitate text comprehension. if coherence relations were not cognitive entities, then there should not be any effect in indicating their presence. the conclusion is, then, that processing coherence relations is part of understanding text. the evidence on the production side is not as abundant, however. this line of research has explored the identification and classification of coherence relations through discourse markers (or connectives). the problem with such an approach is that it does not address the issue of unsignalled relations. it is clear to most researchers that one can postulate relations (and presumably, readers understand them) even when they are not signalled. if all relations are of the same type, that is, if all relations are cognitive entities, then signalling through discourse markers only facilitates their comprehension. lack of signalling does not mean that no relation is present. in the following section we further discuss the signalling problem, and show that signalling has been understudied, focusing mostly on discourse markers. 3 the signalling of discourse relations in this paper, by signalling we mean the cues that indicate that a coherence relation is present, such as the conjunction because as a clue that a causal relation is being presented. we use the term signalling rather than marking because the latter has been associated with discourse markers, one of many possible signalling devices. research on coherence relations has often focused on cues that indicate the presence of a relation, or the lack of such cues, as many relations seem to be unsignalled. whereas it is true that many coherence relations (under whatever definition) are not signalled by a discourse marker, that is, they are implicit, it is also often the case that other markers have been understudied (taboada & mann, 2006; taboada, 2009). our goal in this paper is to push that line of research further. we explore how many, and what types of cues can be found if we study signalling beyond discourse markers. a secondary goal aims at discovering whether unsignalled or implicit relations can be said to exist at all. if we postulate psychological validity for coherence relations, that is, if we assume that coherence relations are present in discourse and that they are recognized by speakers, then there must be signals through which speakers identify relations when parsing discourse. if, as spooren (1997) suggests, underspecified or unsignalled relations obey the cooperative principle (grice, 1975) and the quantity maxim (“say no more than necessary”) 3 , then unsignalled relations are such because no signal is necessary. psycholinguistic experiments have shown that certain relations are processed faster when a connective is present. haberlandt (1982), for instance, found that causal and concessive connectives between two sentences resulted in a faster processing of the second sentence. this was compared to pairs of sentences with no 3 spooren actually makes reference to horn’s (1984) take on the cooperative principle, which can be summarized as “say no more than necessary”. signalling in a corpus of discourse relations 253 connective between them. the conclusion was that the lack of connective necessitated inference, which resulted in longer processing times. sanders et al. (2007) showed that explicitly marked relations led to better performance in text comprehension questions, both in laboratory and realistic situations. the effects of signalling on recall and some aspects of comprehension have been more mixed. meyer et al. (1980) found no positive effect on recalling content 4 . they did, however, find that subjects recalled the structure of the original text more faithfully when it was signalled. millis and just (1994) saw an increase in processing time but more accurate answers to comprehension questions when a connective was present. degand and sanders (2002) report better answers on comprehension questions if the texts include a relational marker. sanders and noordman (2000) found that connectives had a positive effect on processing, but no noticeable effect on recall. sanders and noordman’s conclusion about the recall effect is that the effect of the marker decreases over time, just as the surface representation of the text is lost, but the semantic content is preserved longer. degand and sanders (2002) also caution that the mixed results may reflect mixed methodology, where there was no control for different types of connectives, coherence of the texts, evaluation methodology (free recall versus comprehension questions), or reader background. other studies have shown that the effect of signalling is different for different types of readers. meyer et al. (1980) discovered that explicit connectives helped only underachieving students, those readers that need signalling to identify the top-level structure of a text. britton et al. (1982) also found faster reaction times in a secondary task, but no effect on recall due to signalling, in two types of subjects, with average or low verbal ability (measured in terms of the scholastic aptitude test). although it is not the focus of this paper, it is also worth mentioning that some of the experimental work has studied the role of different types of relations. it has consistently been shown that causal relations are processed faster and often lead to better recall than other types of relations (e.g., keenan et al., 1984; trabasso & sperry, 1985; myers et al., 1987). other research has shown differences among different relations, such as problem-solution and list (sanders & noordman, 2000). it seems clear that coherence relations are different in nature among them. this probably means that their signalling will also be different, not only in terms of whether signalling is present or not, but in terms of which types of signal produce which comprehension effects. the task of a writer or speaker, then, is one of determining how much signalling is enough. a writer may decide that no connective is necessary because other cues that suffice to identify the relation are present, thus obeying the quantity maxim or, according to spooren (1997), the rprinciple (“say no more than necessary”). in a study of young (6-7 year old) and older (11-12 year old) children, spooren found that a number of relations were unsignalled (close to 20%) and, more importantly, that a very large number (between 65 and 75%) were underspecified, that is, they were signalled by general connectives, such as and. there was a significant difference between the age groups, with younger children leaving fewer relations implicit, but using more underspecified relations. 4 meyer et al.’s (1980) signalling included explicit statements of the structure of the text and connectives. as noted later on in this section, the results were different for different types of students (poor vs. good readers). taboada and das 254 the fact that some studies have found no significant effects of signalling on recall may indicate that readers (and hearers) are able to process text and assign relations successfully, even if the effort requires more time with unsignalled relations, or relations that are more weakly signalled (for instance, signalled by an open-class lexical item instead of a connective). most of the work reviewed thus far dealt with connectives/discourse markers. the problem is that there are many other types of signals that may facilitate the comprehension process, and those have clearly been understudied. in previous work (taboada, 2004, 2006, 2009) we have reported on different types of signals that can be used to identify a relation. here we summarize that work, and in the next section we provide a more detailed list of the signals used in this study. discourse markers are, of course, the most studied signals. in some cases, the taxonomy of discourse markers has been reduced to single-word conjunctions. we have found many multiword expressions that function as discourse markers, even though some of them may not be conjunctions from a syntactic point of view, such as in the event that in the following example, from the rst web site (mann & taboada, 2010), which signals a condition relation between spans (1b) and (1c). this is a prepositional phrase that takes a clausal complement. (1) [copyright notice] a. this notice must not be removed from the software, b. and in the event that the software is divided, c. it should be attached to every part. [rst web site] one aspect that we have discussed elsewhere is the use of mood and modality to signal relations. for example, a question (as expressed by an interrogative mood) is a potential signal for a solutionhood relation. verb finiteness is sometimes the only indicator of a relation, as shown in example (2), from the rst discourse treebank (carlson et al., 2002). the circumstance relationship between spans 1 and 2-5 is signalled by the non-finite form of the verb insisting. (2) [1] insisting that they are protected by the voting rights act, [2] a group of whites brought a federal suit in 1987 [3] to demand that the city abandon at-large voting for the nine-member city council [4] and create nine electoral districts, [5] including four safe white districts. [rst discourse treebank] lexical items may also be used to indicate a relation, such as the verb cause in a causal relation, or concede, as in example (3), which in this case marks a concession relation. (3) [s] some entrepreneurs say the red tape they most love to hate is red tape they would also hate to lose. [n] they concede that much of the government meddling that torments them is essential to the public good, and even to their own businesses. [rst discourse treebank] in example (4) there is an evaluation relation between segments 1 and 2. the author characterizes the narrator of the novel “the wedding” as a character removed from the main protagonist, noah, and therefore making the connection between narrator and protagonist quite indirect. the main indicator of this evaluation relation is the semantic content of the word indirect, an adjective conveying subjective content. signalling in a corpus of discourse relations 255 (4) [1] the first-person narrator of “the wedding” is the son-in-law (wilson) of noah’s daughter jane. [2] ?????? talk about indirect. [sfu review corpus] we embrace a view of coherence in discourse whereby coherence relations (also known as relational coherence) and reference and lexical relations (also known as cohesion, or entity-based coherence) are part of what renders a text coherent. this is the view in poesio et al. (2004), and the principle behind veins theory (cristea et al., 1998). in general, coherence established by lexical means, as part of a more general entity coherence, or cohesion, is a very important aspect of signalling. karamanis (2007), for instance, assumes that, in the absence of a marker, entity coherence (links among entities in the discourse) signals the relation. as we will see in later sections, cohesion of all types (reference, lexical, etc.) seems to be a strong indicator of coherence. in fact, in the halliday and hasan (1976) view of cohesion, cohesion and coherence relations are part of the same system, with coherence relations represented by conjunctive links. thus, it is not surprising that we see signalling by lexical and other cohesive devices as an extension of signalling by conjunctions and discourse markers. other cases are more difficult and subjective to interpret. example (5) contains two elaborations embedded within each other. in the first relation, the satellite starts with “recently, the boards…” and continues to the end of the paragraph, which is longer than displayed in the example here. the only possible signal that an elaboration relation is present is the adverb also before the main verb voted in this satellite. the second elaboration relation has that “recently, the boards…” sentence plus the next sentence as nucleus. the satellite starts with “the transaction…” and continues for a while. this second satellite has no adverb, punctuation mark, or any other device that indicates an elaboration on what has gone before. knowledge of the newspaper genre leads us to think that an article, unless other cues are present, proceeds in a series of elaborations. (5) [n1] american pioneer inc. said it agreed in principle to sell its american pioneer life insurance co. subsidiary to harcourt brace jovanovich inc.’s hbj insurance cos. for $27 million. american pioneer, parent of american pioneer savings bank, said the sale will add capital and reduce the level of investments in subsidiaries for the thrift holding company. [s1] [n2] recently, the boards of both the parent company and the thrift also voted to suspend dividends on preferred shares of both companies and convert all preferred into common shares. the company said the move was necessary to meet capital requirements. [s2] the transaction is subject to execution of a definitive purchase agreement and approval by various regulatory agencies, including the insurance departments of the states of florida and indiana, the company said. […] [rst discourse treebank] finally, there is the question of punctuation and layout in written texts, including the problem of how these devices correlate with rhetorical relations. there is some work in this area, going back to hovy and arens (1991) and dale (1991), and including research by bateman (bateman et al., 2001), which in general shows a good correlation between some forms of layout and rhetorical relations. it should be fairly clear by now that multiple signals for relations are possible, and that some of them are straightforward to annotate, such as discourse markers, especially conjunctions, whereas some other signals require long-distance dependencies and involve a certain amount of taboada and das 256 subjectivity. in our work, we have strived to compile a list of signals that we felt we could annotate reliably. the next section discusses these. 4 signals for reliable annotation the most important aspect of the annotation was to select and classify the types of cues to annotate. discourse markers have been extensively studied, and are relatively easy to identify. beyond discourse markers, we found other classes of cues that have been mentioned in previous studies, or that we identified in our preliminary corpus work. the classification has a top-level breakdown into discourse markers, morphological, syntactic, semantic, lexical, genre and graphical features, plus heuristics specific to each relation. we started our annotation, as we explain in section 5, by consulting previous studies for indication of what signalling devices have been found in corpora (halliday & hasan, 1976; blakemore, 1987; schiffrin, 1987; fraser, 1990; scott & de souza, 1990; dale, 1991; blakemore, 1992; sanders et al., 1992, 1993; knott & dale, 1994; knott, 1996; corston-oliver, 1998a; fraser, 1999; marcu, 1999, 2000b; bateman et al., 2001; schiffrin, 2001; blakemore, 2002; lapata & lascarides, 2004; polanyi et al., 2004; sporleder & lascarides, 2005; fraser, 2006; huong, 2007; prasad et al., 2007; pardo & nunes, 2008; sporleder & lascarides, 2008; teijssen et al., 2008; fraser, 2009; lin et al., 2009; pitler et al., 2009; louis et al., 2010; prasad et al., 2010). then, as we annotated more and more relations, we added to our classification. the top-level classification of signals is provided in figure 2. we briefly discuss this classification below. a full account, with examples of each type, is available as supplementary material to this article 5 . please note that the subcategories in figure 2 are illustrative, not exhaustive. discourse markers are by far the most studied type of signalling (see references in taboada and mann, 2006a, 2006b). markers are specific to each relation, such as if for condition or although for concession. there is, however, no one-to-one correspondence between markers and relations, and many markers are ambiguous (and can indicate a number of discourse relations, in addition to its function as linking device within clauses and phrases). in our annotation, we mainly followed fraser’s (1999, 2006, 2009) definition of discourse markers, that discourse markers constitute a functional class of linguistic elements drawn from different syntactic classes, such as conjunctions, adverbs and prepositional phrases. they connect discourse segments, and signal a relation between them. in addition, we also followed a number of conditions for considering an expression to be a discourse marker. the conditions are enumerated below. 1. the scope of the function of a discourse marker is a single discourse sequence comprising adjacent text spans in a relation. 2. discourse markers can be present at the beginning or end of the sentence (or segment), or within the sentence (or segment). 3. discourse markers signal relations that hold between two adjacent text segments. 4. a discourse marker does not create the relation between text segments. it only guides the interpretation of the relation. 5 http://www.sfu.ca/~mtaboada/docs/taboada_das_dialogue_and_discourse_2013_supplementary_material.pdf signalling in a corpus of discourse relations 257 figure 2. top-level classification of signals entity features include links where entities, similar or dissimilar, help interpret the relation. for example, in (6), which contains a multinuclear list relation with three nuclei, the three distinct entities indicate that a roster of companies is being listed 6 . (6) [earlier this year, tata iron & steel co.’s offer of $355 million of convertible debentures was oversubscribed.]n [essar gujarat ltd., a marine construction company, had similar success with a slightly smaller issue.]n [larsen & toubro started accepting applications for its giant issue earlier this month;]n many of the semantic relations in halliday and hasan (1976) can be used to identify relations, such as antonyms as signals of contrast, or hypernyms as indicators of the satellite(s) in an elaboration relation. we define these as semantic relations because a semantic link between two or more entities is established, as opposed to the lexical features mentioned below, where a single word or phrase is used, with no connection to other words in the text. the category that includes entities is related to this one. under semantic relations, however, we include relations that are easily labelled in terms of synonym, antonym, hyponym, etc. 6 there are many other signals in this example, among them the word similar in the second sentence, and the temporal descriptions (earlier this year; earlier this month). the example is being used here to illustrate entity features. this will be the case with other examples used to illustrate signals: one particular signal will be highlighted, but other signals may be present in the example. taboada and das 258 lexical features include the use of indicative words and phrases, such as individual words that indicate a relation, for example, the verbs concede and cause for concession and cause respectively. indicative phrases are some of the more difficult signals to define a priori, but, when they appear, they are unequivocal in their nature as signals. examples from the current round of analysis include last year as an indication of background, and at the same time for temporalsame-time. among morphological features, tense is the most prominent one, helping indicate temporal relations (circumstance in rst terms), or more general circumstances, as is the case with some instances of non-finite verbs (taboada, 2006). at the syntactic level there are a host of constructions that help identify a relation. from word order, such as subject-verb inversion for condition (had he known…) to sentence mood, such as the use of interrogatives to signal solutionhood. graphical and other punctuation features, such as lists and headings, and other forms of layout are sometimes indicators of a relation. numerical elements are present in list relations, but also in more subtle ways, when an elaboration consists of providing a general word (in this case, a number) and then listing the contents of that word. an example is (7), where the nucleus contains the numeral five, and the satellite a listing of five names. (7) [this maker of electronic devices said it replaced all five incumbent directors at a special meeting …]n [elected as directors were mr. hollander, frederick ezekiel, frederick ross, arthur b. crozier and rose pothier.]s genre helps guide the interpretation of relations when the style of the genre is well known to the reader. in the newspaper genre that all the texts in the corpus belong to, it is common to start the text with general information, and to continue with further details. this results in elaboration relations, with the nucleus being the first sentence or paragraph, and the rest of the article acting as a satellite that expands on the beginning of the text. other aspects that are specific to newspaper writing are the ways in which the attribution relation is signalled. these are classified under graphical (quotes and dashes) or syntactic features (verbs of diction such as say or claim), but it is the fact that the genre is journalistic discourse that provides the interpretation for those signals. two other types of signals are not included in our classification above, because they are either too general or too specific. the general class of discourse features is less loosely defined, and can include position in the text (summary tends to appear at the end), given-new status (in elaboration and contrast relations), or genre characteristics (evaluation relations more common in opinion texts). in some instances, this class overlaps with genre. our last category includes heuristics, that is, features that are specific to relations. one example is the use of evaluative words (satisfactory, adequate, success) in a satellite, which indicate that it is modifying a nucleus in an evaluation relation. these broad categories describe single signals, that is, one specific item that indicates the relation. the types of signals described above contain many specific signals in themselves. for signalling in a corpus of discourse relations 259 example, the syntactic type includes specific signals such as infinitival clause, participial clause, parallel syntactic construction and reported speech pattern. in addition, we find that many relations are indicated by combined signals. combined signals are made of two or more single signals which work in combination with each other to indicate a particular relation. for instance, the list relation between span 2-4 and span 5 in example (8) in the next section (see table 1) is indicated by the combined signal entity + syntactic (more specifically, given entity + subject np), along with the single signal lexical chain (of semantic type). we have identified 10 broad types of combined signals: (i) entity + positional, (ii) entity + syntactic + lexical, (iii) entity + syntactic, (iv) graphical + syntactic, (v) lexical + positional, (vi) lexical + syntactic + positional, (vii) lexical + syntactic, (viii) syntactic + lexical, (ix) syntactic + positional, and (x) semantic + syntactic. lists of combined signals can be found in the supplementary material available online (see footnote 5). some relations are also indicated by multiple signals. the difference between combined signals and multiple signals is one of independence of operability. in a combined signal, there are usually two signals, one of which is an independent signal, while the other one is dependent on the first signal. for example, in given entity + subject np, which is a combined signal, given entity is the independent signal because it directly (and independently) refers back to the entity introduced in the first span. in contrast, subject np is the dependent signal because it is used to specify additional attributes of the first signal. in this particular case, the syntactic role of the given entity (i.e., a subject np) in the second span is specified by the use of the second signal subject np. multiple signals, on the other hand, function independently and separately of each other, but they all contribute to signaling the relation. for example, in an elaboration relation with multiple signals, involving a genre feature (e.g., textual organization) and a lexical feature (e.g., indicative word), the signals do not have any connection, as they refer to two different features which separately signal the relation. 5 annotation process for our corpus, we have selected the rst discourse treebank (carlson et al., 2002), a collection of 385 wall street journal articles annotated for rhetorical relations. we elected to use an existing corpus to expedite our research on signalling, even though the corpus may not be ideal. we believe this will be a more and more frequent situation for researchers in discourse, with so many existing annotated corpora available that can be reused and extended. we discuss some of our technical and theoretical difficulties with the layered annotations. the annotation process involves examining each relation and, assuming the relation annotation is correct, searching for cues that indicate that such relation is present. in some cases, more than one cue may be present. from a theoretical point of view, some of the difficulties that we are encountering are disagreements with the annotations already present in the corpus, from the segmentation (for our approach to segmentation, see tofiloski et al., 2009) to the application of relation definitions, also including the particular inventory of relations used to annotate the corpus. from a practical point of view, we need to read lengthy texts and examine both parts of a relation, which are sometimes far apart from each other. our current setup involves opening the files in rsttool, a graphical interface to annotate rst relations (o'donnell, 1997), and annotating information about the signalling in a separate excel file, as rsttool does not allow for multiple annotations. the annotation, at this point, includes only information about the type and subtype of signal involved, and an indication of what word(s) convey the signalling. it does taboada and das 260 not, however, consist of an integrated annotation on the actual rst discourse treebank files. a future goal is to find a way to layer our signalling annotation over the existing rst discourse treebank, marking both the type of signal and the words we can identify as signals. in our preliminary corpus study, we annotated 40 articles which constitute approximately ten percent of the 385 articles in the rst discourse treebank 7 . the texts in these articles contain 1,304 rhetorical relations. for the annotation of these relations, we performed a sequence of three main tasks: (i) we examined each and every relation in the rst discourse treebank, (ii) we identified the signals involved to indicate those relations, and, finally, (iii) we documented information on how the relations are signalled. we used the list presented in figure 2 above to identify signals. when confronted with a new instance of a particular type of relation, we consulted our list, and tried to find appropriate signal(s) that could best function as the indicator for that relation instance. if our search led us to assigning an appropriate signal (or more than one appropriate signal) to that relation, we declared success in identifying the signal(s) for that relation. if our search did not match any of the signals in the list, then we examined the context (comprising the spans) to discover any potential new signals. if a new signal was identified, we included it in the appropriate category in our existing list. in this way, we proceed through identifying the signals of the relations in the corpus, and, at the same time, keep on updating our database with new signalling information, if necessary. we found that after approximately 20 files, or 650 relations, we added very few new signals to the list. in the coding task, we provided annotations for signals of coherence relations, or in other words, we added signalling information to the existing relations from the rst corpus. for this purpose, we extracted the signals identified, and documented them along with relevant information about the relation in question, the document number (to which the relation belongs), the status of the spans (i.e., nucleus or satellite), and the span numbers (i.e., the location of the spans in the text). we annotated the signalling information in a separate excel file, since rsttool, as previously mentioned, does not allow multiple levels of annotation. 5.1 an annotation example we provide the annotation of a short rst file (file no. 650) with signalling information. the file contains the text in example (8). (8) sun microsystems inc., a computer maker, announced the effectiveness of its registration statement for $125 million of 6 3/8% convertible subordinated debentures due oct. 15, 1999. the company said the debentures are being issued at an issue price of $849 for each $1,000 principal amount and are convertible at any time prior to maturity at a conversion price of $25 a share. the debentures are available through goldman, sachs & co. 7 the 385 articles in the rst discourse treebank are organized into 385 separate files which are divided into two groups: (i) training documents, comprising 347 files, and (ii) test documents, comprising the remaining 38 files. the 40 articles chosen for annotation are taken from the training set. signalling in a corpus of discourse relations 261 the graphical representation of the rst analysis of this text using the rst tool is provided in figure 3. the rst analysis shows that the text in example (8) comprises five spans which are represented in the diagram (in figure 3) by the numbers, 1, 2, 3, 4, and 5 8 , respectively. in the diagram, the arrowhead points from a satellite to a nucleus span. span 3 (nucleus) and span 4 (nucleus) are in a multinuclear list relation, and together they make the combined span 3-4. span 2 (satellite) is connected to span 3-4 (nucleus) by an attribution relation, and together they make the combined span 2-4. a multinuclear list relation holds between spans 2-4 (nucleus) and 5 (nucleus), and together they make the combined span 2-5. finally, span 2-5 (satellite) is connected to span 1 (nucleus) by an elaboration (more specifically, elaboration-addition) relation. figure 3. graphical representation of an rst analysis we annotated the text in example (8) with the appropriate signalling information. a detailed description of our annotation for the text is provided in table 1. according to our annotation, the elaboration relation between spans 1 and 2-5 is indicated by three types of signals: (i) genre; (ii) entity + syntactic; and (iii) lexical features. first, the text is part of the newspaper genre (since it is taken from a wall street journal article), and in newspaper texts the content of the first (or the first few) paragraphs is typically elaborated on in the following paragraphs. a reader, being conscious of the fact that he/she is reading a newspaper text, expects the presence of an elaboration relation between the first paragraph (or the first few paragraphs) and subsequent paragraphs. it is this prior knowledge about the textual organization of the newspaper genre that guides the reader to interpret an elaboration relation between paragraphs in a news text. in this particular example, the entire first paragraph is the nucleus of the elaboration relation, with the two following paragraphs being its satellite. thus, we postulate that the elaboration relation is conveyed by the genre feature (more specifically by a feature 8 spans 2 to 5 do not actually have a label in the corpus. while the labels are inferable, this makes the annotation more complicated with lengthy files. taboada and das 262 which we call textual organization). second, we postulate that a combined signal entity + syntactic, made of two individual features, is operative in signalling the elaboration relation (see section 6 for more information about combined signals). one can notice that the entity sun microsystems inc., mentioned in the nucleus, is elaborated on in the satellite. syntactically, the entity is also used as the subject np of the sentence the satellite starts with, representing the topic of the elaboration relation. finally, the elaboration relation is also (perhaps rather loosely) signalled by a lexical feature, lexical overlap. words such as debentures and convertible occur in both the nucleus and satellite, indicating the presence of the same topic in both spans, with an elaboration in the second span of some topic introduced in the first span. the list relation between spans 3 and 4 is conveyed in a straightforward (albeit underspecified) way by the use of the discourse marker and. the attribution relation between spans 2 and 3-4 is indicated by a syntactic signal, a reported speech pattern in which the reporting clause (span 2) functions as the satellite and the reported clause (span 3-4) functions as the nucleus. the key is the s+v (subject+verb) combination with a reported speech verb (said). source file no. nucleus span satellite span relation name marker type identified specific marker identified explanation – how the relation is signalled 650 1 2-5 elaborationadditional genre textual organizati on newspaper: the content of the first paragraph (or the first few paragraphs) is elaborated on in the following paragraphs. entity + syntactic given entity + subject np sun microsystems inc., mentioned in the nucleus, is the subject of the sentence which the satellite starts with. lexical lexical overlap words such as debentures and convertible are in both the spans. 3/4 list discourse marker and the discourse marker and functions as a signal for the list relation. 3-4 2 attribution syntactic reported speech pattern the reported speech pattern “the company said…” is a signal for the attribution relation. 2-4/5 list entity + syntactic given entity + subject np the subject np of the reported speech in the first span and the subject np of the sentence in the second span both refer to the same entity: the debentures. semantic lexical chain the words issued and available in the respective spans are semantically related. table 1. annotation of an rst file with relevant signalling information signalling in a corpus of discourse relations 263 finally, the list relation between spans 2-4 and 5 is indicated by two types of signals: (i) entity + syntactic and (ii) semantic feature. for the combined feature entity + syntactic, the specific signal is called given entity + subject np, which means that the subject np of the reported speech within the first span and the subject np of the sentence in the second span both refer to the same entity (the debentures, in this case). for the semantic feature, the specific signal is a lexical chain which means that semantically similar or related words occur in the respective text spans. we notice that words such as issued and available are semantically related, and they are used in both spans, indicating a list relation holding between them. after the annotations (with signalling information) are done, we code our annotated data in a separate excel file. the coded version of our annotation for the text is provided in table 2. source nucleus satellite relation marker type specific marker 650 1 2-5 elaboration (-additional) genre + (entity + syntactic) + lexical textual organization + (given entity + subject np) + lexical overlap 650 3/4 list discourse marker and 650 3-4 2 attribution syntactic reported speech pattern 650 2-4/5 list (entity + syntactic) + semantic (given entity + subject np) + lexical chain table 2. coding of signalling information for relations in an rst-annotated text in this way, we completed our annotation task with signalling information for the relations for a total of 40 files in the rst discourse treebank. 5.2 reliability study as with all annotations, ours carries a certain amount of subjectivity. this is particularly true with discourse annotations and all phenomena beyond semantics, where interpretation of the context and of long-distance features plays a role. our list of signals and the annotation procedure were agreed upon after several iterations of the taxonomy and after adding more signals when our initial analysis revealed more than we had originally listed. to check the validity and reproducibility of our taxonomy, we conducted a reliability study. we selected approximately 10% of the 1,304 relations in the current annotation (see section 6), coming from two of the texts. one of us had annotated the entire corpus, and the other one annotated those two files, containing 130 relations. we concentrated on whether we agreed on at least one of the signals for the relation. some relations have multiple signals, and some relations have combined signals. calculating agreement on those becomes very complex quite quickly, so we stayed with simple signals, and an agreement of at least one signal per relation. also because of the complexity of the task, we calculated agreement using the top level of the signalling taxonomy, that is, the nine top-level signals from figure 2. we established whether we agreed on the type of signal, not necessarily on where it was conveyed in the text (e.g., for a lexical chain, we annotated ‘semantic’, but not what words were involved in the chain). agreement was calculated using cohen’s kappa (siegel & castellan, 1988), with nominal data, namely, the nine categories in our classification, plus an extra category, “no signal”, used to indicate cases where the annotator concluded that there was no identifiable signal. agreement on taboada and das 264 this category is just as important as on the other ones. the kappa value for our study was 0.68, or moderate agreement. table 3 presents the disagreements per relation. of note is the fact that in elaboration relations, we disagreed in only 17 out of 64 instances (26% of the time), whereas we expected disagreement for that relation to be higher. in terms of markers, disagreement was higher for genre, where we disagreed in all four cases that it appeared, one annotator identifying genre as a signal, and the other one labelling the instance as ‘no signal’. other markers where disagreement was high were semantic markers (66%, or 20 out of 30 cases) and lexical signals (55%, 5 out of 9 cases). we have, as a consequence, refined our taxonomy of lexical and semantic labels, and believe this will have a positive effect on agreement, to be determined in future agreement studies as we proceed with annotation. relation agreement disagreement antithesis 3 attribution 19 1 background 1 3 cause-result 1 circumstance 1 1 condition 2 contrast 3 elaboration 47 17 example 2 explanation 4 hypothetical 1 list 5 manner 2 problem-solution 2 1 purpose 5 same-unit 6 summary 1 temporal 2 total 97 33 table 3. agreement and disagreement per relation a more general issue as regards reliability studies is whether they are useful at all. in our study, as in most published studies, the level of agreement is considered acceptable, and we do believe that our annotation is reproducible. the larger question is whether providing values for kappa or for similar measures reveals much about the annotation process and its level of difficulty. reaching such level of agreement after four iterations through the data and after modifying the annotation guidelines is quite different from doing so after a quick explanation of the methodology to a new member of the research group. spooren and degand (2010) discuss agreement measures in a similar task, that of coding coherence relations, and conclude that measures beyond kappa are necessary to ensure and measure reliability, such as double coding and discussion of disagreement and agreement cases, and other agreement measures. those will be part of future reliability tests in our project. in our case, the reliability study could only be carried out by members of our project, who were familiar with rst, shared similar points of view with regard to what counts as a relation, signalling in a corpus of discourse relations 265 and agreed on the list of signals given. other annotators may disagree with our results, no matter how experienced, or how much time they spend studying our guidelines. we point this out because we feel that too much emphasis is placed on arriving at an acceptable measure of agreement, when an acceptance of the intrinsic difficulty of annotation is what is needed, together with a reasonable explanation of how the annotation was performed. 6 results among the 1,304 relations examined, the distribution of signalled relations (indicated either by discourse markers or by some other signal) and unsignalled relations (not indicated by any signal) is provided in table 4. relation type tokens percentage signalled relations 1,127 86.43% unsignalled relations 177 13.57% total 1,304 relations indicated by a discourse marker 251 22.27% relations indicated by other signals 878 77.91% total 1,127 table 4. distribution of signalled and unsignalled relations the results show that 1,127 relations (86.43%) out of all the 1,304 relations are signalled, either by a discourse marker or with the help of some other signalling device. on the other hand, no significant signals are found for the remaining 177 relations (13.57%). among the 1,127 signalled relations, we find that discourse markers are used to signal 251 relations (22.27% of the signalled relations), while 878 relations (77.91% of the signalled relations) are indicated with the help of some other signals. we need to point out that there are two instances of list relation which are signalled by both a discourse marker and some other signal (which is why the total of 251 plus 878 actually adds up to 1,129). this is because these relations are multinuclear, consisting of three or four nuclei, and we found that while a nucleus is connected to another nucleus by a discourse marker, a third nucleus is related to any of the two former nuclei (in case of a relation with three nuclei), or to a fourth nuclei (in case of a relation with four nuclei) by means of some other signal(s). for the 251 instances of relations signalled by a discourse marker, we found 58 different discourse markers. examples of some of these discourse markers include after, although, and, as, as a result, because, before, despite, for example, however, if, in addition, moreover, or, since, so, thus, unless, when and yet. a full list of these extracted markers is available (see footnote 5). for the 878 signalled relations without discourse markers, we found that a wide variety of signals are used to indicate them. as mentioned in section 4, we divide the signals into two broad groups: single and combined signals. in our corpus analysis, 81.81% of the signalled relations (922 out of 1,127 signalled relations) are exclusively indicated by a single signal (including discourse markers), whereas 5.69% of the signalled relations (64 out of 1,127) are indicated by a combined signal. we have also noticed that in many cases multiple signals, i.e., two or more types of other signals (single or combined) are separately used to indicate a particular relation instance. for taboada and das 266 instance, the elaboration-additional relation between span 1 and span 2-5 in example 8 (see table 1) is indicated by multiple signals: (i) genre, (ii) entity + syntactic, and (iii) lexical features. the distribution of the signals in our annotation shows that 12.51% of the signalled relations (141 out of 1,127 signalled relations) contain multiple signals. this is an encouraging result for any attempt at automatic identification, as the redundancy in signalling will increase the chances of identification. the relative distribution of relations with respect to whether they are indicated by a discourse marker, by some other signals, or whether they are unsignalled is provided in table 5. 9 no. relation group relation # relations signalled by dms # relations signalled by other markers # relations not signalled total 1. attribution attribution 0 228 3 231 attributionnegative 0 0 0 0 2. background background 2 8 6 16 circumstance 21 9 9 39 3. cause cause 2 1 1 4 result 3 0 0 3 consequence 14 1 12 27 4. comparison comparison 5 9 4 18 preference 0 0 0 0 analogy 0 0 0 0 proportion 0 0 0 0 5. condition condition 15 1 1 17 hypothetical 1 1 0 2 contingency 0 0 0 0 otherwise 0 0 0 0 6. contrast contrast 19 2 2 23 concession 13 0 1 14 antithesis 25 1 4 30 7. elaboration elaborationadditional 23 238 41 302 elaborationgeneralspecific 1 16 4 21 elaborationpart-whole 0 0 0 0 elaborationprocess-step 0 0 0 0 elaborationobjectattribute 4 179 3 186 elaborationset-member 0 6 1 7 example 3 6 8 17 9 the total number of relations analyzed is actually 1,304. the total in the table shows 1,306, because two relations are counted twice, two instances of list relation indicated by both discourse markers and other signals at the same time. signalling in a corpus of discourse relations 267 definition 0 2 0 2 8. enablement purpose 0 39 0 39 enablement 0 0 0 0 9. evaluation evaluation 1 3 1 5 interpretation 1 0 9 10 conclusion 0 0 0 0 comment 0 0 9 9 10. explanation evidence 0 3 8 11 explanationargumentative 6 1 23 30 reason 12 1 4 17 11. joint list 50 27 6 83 disjunction 3 0 0 3 12. mannermeans manner 3 0 0 3 means 1 4 0 5 13. topiccomment problemsolution 2 2 2 6 questionanswer 0 0 0 0 statementresponse 0 2 0 2 topiccomment 1 0 0 1 commenttopic 0 0 0 0 rhetoricalquestion 0 0 0 0 14. summary summary 0 0 8 8 restatement 0 9 0 9 15. temporal temporalbefore 3 0 0 3 temporalafter 7 1 0 8 temporalsame-time 3 1 0 4 sequence 5 0 0 5 invertedsequence 0 0 0 0 16. topic-change topic-shift 0 0 4 4 topic-drift 0 0 0 0 17. same-unit same-unit 2 76 3 81 18. span span 0 0 0 0 19. textual organization textual organization 0 1 0 1 total 251 (19.25%) 878 (67.33%) 177 (13.57%) 1,306 table 5. distribution of relations indicated by a dm (discourse marker), of relations indicated by some other signals, and of unsignalled relations the distribution of relations in table 5 shows that almost every group of relations is more or less signalled. in particular, we find that relation groups such as attribution, elaboration, taboada and das 268 enablement, and joint are most frequently signalled, either by discourse markers or by some other signal 10 . we also found that that there is only one group of relations, evaluation, which is rarely indicated by any signal. among the signalled relations, discourse markers are most frequently used to signal relations such as circumstance, result, consequence, condition, concession, contrast, antithesis, reason and list. in contrast, relations such as attribution, background, comparison, elaborationadditional, elaboration-general-specific, elaboration-object-attribute, example and purpose are rarely or never signalled by a discourse marker. our findings are also parallel to the results presented in our earlier work (taboada, 2006), where we found that relations such as concession, condition and purpose are most frequently signalled (by a discourse marker), while background and summary are rarely signalled (by a discourse marker). relations which are mostly indicated by some other signals include attribution, elaborationadditional, elaboration-general-specific, elaboration-object-attribute, purpose and restatement. in contrast, relations which are rarely or never indicated by some other signals include circumstance, consequence, condition, contrast, antithesis, explanation-argumentative and temporal-after. finally, the relations for which no signals (neither a discourse marker nor any other signal) were found include comment, summary and topic-change. the relation-wise distribution of discourse markers shows that a significant number of relations are frequently signalled by a wide variety of discourse markers. the distribution of the most frequently-occurring discourse markers with respect to the most common relations is provided in table 6. common relation group common relation most frequently occurring discourse markers background (23) circumstance (21) when (5), as (4), with (3), cause (19) consequence (14) and (6) condition (16) condition (15) if (11), unless (2) contrast (57) contrast (19) but (11), however (3) concession (13) while (3), but (2), though (2) antithesis (25) but (11), although (3), however (3) elaboration (31) elaborationadditional (23) and (8), but (6), as (2), so far (2) example (3) for example (2) explanation (18) reason (12) and (4), because (4), because of (3) joint (53) disjunction (3) or (3) list (50) and (44), in addition (2), moreover (2) temporal (18) sequence (5) and (4) temporal-after (7) since (3), after (2) temporal-before (3) before (3) table 6. distribution of the most frequently occurring dms with respect to the most common relations signalled by them the distribution of different discourse markers provided in table 6 shows what discourse markers are most frequently used to convey a particular relation, and how frequently they are 10 we exclude same-unit from this list because same-unit is not a true coherence relation. in the rst discourse treebank it is used to join discontinuous grammatical elements, such as subject np and vp. signalling in a corpus of discourse relations 269 used for signalling that relation 11 . for instance, list relations are most frequently signalled by and, in addition, and moreover. out of the 50 instances of list relation, the dms and, in addition and moreover, are used 44 (88%), 2 (4%), and 2 (4%) times, respectively. the complete distribution of the discourse markers with respect to the relations is provided online (see footnote 5). in an alternate combination, the distribution of the most common relations with respect to the most frequently-occurring discourse markers is provided in table 7. table 7 shows what relations are most frequently signalled by a particular discourse marker, and how frequently they are signalled by that marker. for instance, the discourse marker but is most frequently used to signal contrast and elaboration relations. out of the 35 instances of but, the relations contrast and elaboration are signalled 25 (71.43%) and 6 (17.14%) times, respectively. frequently occurring dm common relation group common relation although (5) contrast (5) antithesis (3) and (70) cause (7) consequence (6) elaboration (8) elaboration-additional (8) joint (44) list (44) explanation (4) reason (4) temporal (4) sequence (4) as (8) background (4) circumstance (4) elaboration (2) elaboration-additional (2) because (8) cause (2) consequence (2) explanation (6) explanation-argumentative (2) reason (4) because of (6) explanation (4) reason (3) before (4) temporal temporal-before (3) but (35) contrast (25) antithesis (11) concession (3) contrast (11) elaboration (6) elaboration-additional (6) however (9) contrast (6) antithesis (3) contrast (3) if (13) condition (11) condition (11) since (5) temporal (3) temporal-after (3) when (10) background (5) circumstance (5) while (8) comparison (3) comparison (3) contrast (4) concession (3) with (4) background (3) circumstance (3) without (6) manner-means manner (3) table 7. distribution of the most common relations with respect to the most frequently occurring discourse markers the relation-wise distribution of other signals and the other signal-wise distribution of relations show even more diverse relationships between the relations and the other signals. the 11 the numerical value within parentheses following a relation/relation group refers to the number of instances the relation/relation group is signalled by a dm. the numerical value within parentheses following a dm refers to the number of times it is used to signal the corresponding relation. this applies to tables 6, 7, 8 and 9. taboada and das 270 distribution of the most frequently-used signals with respect to the most common relations is provided in table 8. relation group relation other signal type specific signal attribution (228) attribution (228) syntactic (220) reported speech pattern (220) genre (4) newspaper heuristics (4) lexical (4) vp cue (4) background (17) background (8) morphological (2) change of tense (2) lexical (5) indicative phrase (5) circumstance (9) syntactic + positional (4) reduced relative clause + beginning (3) lexical (5) indicative phrase (5) comparison (9) comparison (9) lexical (8) indicative phrase (7) elaboration (447) elaborationadditional (238) entity + syntactic (84) given entity + subject np (74), given entity + subject np (rs) (6) entity (79) given entity (77) lexical (8) indicative word (4), indicative phrase (4) semantic (133) lexical overlap (61), lexical chain (60), phrasal chain (7) syntactic (51) relative clause (26), reduced relative clause (10), participial clause (7) genre (38) textual organization (32), newspaper heuristics (6) graphical (16) parentheses (10), dashes (4) elaborationobject-attribute (179) syntactic (167) relative clause (85), reduced relative clause (45), infinitival clause (np) (27) elaborationgeneral-specific (16) entity (5) given entity (5) entity + syntactic (5) given entity + subject np (3) graphical (5) dash (4) semantic (11) lexical chain (5), lexical overlap (5) enablement (39) purpose (39) syntactic (38) infinitival clause (37) joint (27) list (27) syntactic (14) parallel syntactic constructions (9) semantic (7) lexical chain (3) manner-means (4) means (4) lexical + syntactic (4) indicative word + participial clause (4) summary (9) restatement (9) graphical (8) parentheses (7) table 8. distribution of the most frequently used other signals with respect to the most common relations indicated by them the relation-wise distribution of different other markers in table 8 shows what other signals are most frequently used to indicate a particular relation, and how frequently they are used for indicating that relation. for instance, elaboration-additional relations are most frequently signalled by semantic, syntactic, entity and genre features. more specifically, semantic, entity + syntactic, entity, syntactic and genre features are individually used 133 (55.88%), 84 (35.29%), signalling in a corpus of discourse relations 271 79 (33.19%), 51 (21.43%), 38 (15.97%) times, respectively, out of the 238 instances an elaboration-additional relation is present 12 . in an alternate combination, the distribution of the most common relations with respect to the most frequently-occurring other signals is provided in table 9. the distribution of relations with respect to other markers in table 9 shows what relations are most frequently indicated by a particular other signal, and also how frequently they are indicated by that signal. for instance, the signal relative clause is most frequently used to signal elaboration-object-attribute and elaboration-additional relations: out of the 112 instances of relative clauses, elaboration-object-attribute and elaboration-additional relations are signalled 85 (75.89%) and 26 (23.21%) times, respectively. other marker type specific other marker relation group relation entity + syntactic (92) given entity + subject np (78) elaboration (77) elaboration-additional (74), elaborationgeneral-specific (3) given entity + subject np (rs) (7) elaboration (7) elaboration-additional (7) entity (87) given entity (84) elaboration (83) elaboration-additional (77) lexical (51) indicative phrase (40) background (9) background (5), circumstance (4) comparison (7) comparison (7) elaboration (11) elaboration-additional (4), elaboration-objectattribute (2), elaboration-setmember (2), example (3) indicative word (10) elaboration (4) elaboration-additional (4) semantic (163) lexical chain (72) elaboration (66) elaboration-additional (60), elaborationgeneral-specific (5) lexical overlap (67) elaboration (66) elaboration-additional (61), elaborationgeneral-specific (4) syntactic (573) reported speech pattern (223) attribution (220) attribution (220) relative clause (112) elaboration (112) elaboration-objectattribute (85), elaboration-additional (26) reduced relative clause (55) elaboration (55) elaboration-objectattribute (45), elaboration-additional (10) 12 note: in signalling relations by discourse markers, a single discourse marker is typically used to signal a particular instance of a relation. however, in signalling relations by signals other than discourse markers, two or more signals are frequently used at the same time to indicate a particular instance of a relation. as a result, the individual distribution score of a particular other signal, unlike that of a discourse marker, is not relative to that of any other signal. taboada and das 272 infinitival clause (41) enablement (37) purpose (37) elaboration (3) elaboration-additional (3) infinitival clause (np) (27) elaboration (27) elaboration-objectattribute (27) participial clause (19) elaboration (18) elaboration-objectattribute (10), elaboration-additional (7) genre (47) textual organization (36) elaboration (36) elaboration-additional (33) newspaper heuristics (11) elaboration (6) elaboration-additional (6) attribution (4) attribution (4) graphical (37) parentheses (18) elaboration (11) elaboration-additional (10) summary (7) restatement (7) dashes (12) elaboration (11) elaboration-additional (4), elaboration-objectattribute (4) lexical + syntactic (18) pp cue + participial clause (7) elaboration (7) elaboration-objectattribute (7) indicative word + participial clause (5) manner-means (4) means (4) syntactic + positional (6) reduced relative clause + beginning (3) background (3) circumstance (3) parallel pp constructions + beginning (2) joint (2) list (2) table 9. distribution of the most common relations with respect to the most frequently-occurring other signals in the specific case of elaboration, there are indeed some significant differences in signalling the different types of elaboration (see table 8). among the 447 instances of elaboration relations, the majority is distributed between elaboration-additional (238 instances) and elaboration-object-attribute (179 instances) while the other types of elaboration have much fewer tokens. elaboration-additional relations are signalled by a wide variety of signals. the most important types (with higher number of tokens) include (i) entity + syntactic (84), (ii) entity (79), (iii) semantic (133), and (iv) syntactic (51). on the other hand, elaboration-object-attribute relations are mainly signalled by syntactic features, in particular by features such as relative clause (130) and infinitival clause (27). while we may disagree in principle with the very specific breakdown of elaboration, in this case it does seem that the annotators of the rst discourse treebank were on the right track, distinguishing subtypes that are different in their signalling. it is worth mentioning that elaboration-object-attribute, a relation that has been questioned as not a true rst relation, but rather a derivative of entity relations (knott et al., 2001), is actually not signalled through semantic or entity features (which would be the equivalent to the entity or reference relations that knott et al. postulated). it is most frequently signalled by extensions to the noun that the relation modifies (relative and infinitival clauses). as for the 177 relation instances for which we could not identify a signal (see table 4), those include a number of different relation types, but there were three particular relations that were never signaled: comment, summary and topic-shift (21 instances among the three). there are signalling in a corpus of discourse relations 273 three different reasons why we believe no signals could be found. first of all, in some cases we found that there were errors in the annotation, and a relation was postulated, whereas we would not have annotated a relation, or we would have proposed a different one. summary and elaboration in the rst-discourse treebank seem to be used in very similar contexts, so when a summary was annotated, but we believed the relation was not in fact a summary, it was more difficult to find signals that would identify the relation as summary. secondly, some of the rst discourse treebank relations are not true rst relations. relations such as comment or topicshift, in our opinion, belong in the realm of discourse organization, not together with relations among propositions. finding no signals in those cases is not surprising, as such phenomena are not likely to be indicated by the same type of signals as coherence relations proper. finally, in many cases, one or both of the annotators had a sense that the relation was clear, but could not pinpoint the specific signal used. this is the case with tenuous entity relations, or relations that rely on world knowledge. what may be happening in those cases is that the relation is being evoked, in the same way frames and constructions may be evoked (dancygier & sweetser, 2005). dancygier and sweetser propose that, in some constructions, only one aspect of the construction is necessary in order to evoke the entire construction. such is the case with some instances of sentence juxtaposition, which give rise to a conditional relation reading, as in “steal a bait car. go to jail” (the slogan for a car-theft prevention campaign by the vancouver police). no conditional connective is necessary. the juxtaposition of the two sentences, together with the imperative and a certain amount of world knowledge lead to the conditional interpretation. we would like to conclude this section by repeating that our results show that relation signalling is much more sophisticated than previously thought, and that a certain level of redundancy is present in many relations. recent work in the automatic identification of relations has postulated a clear separation between implicit and explicit relations. a series of experiments by marcu and echihabi (2002) and sporleder and lascarides (2005, 2008) have shown that it is difficult to generalize from “explicit” to “implicit” features, that is, that a classifier built using “explicit” relations does not necessarily identify “implicit” relations correctly (see also the discussion in stede, 2012). we use quotes around “explicit” and “implicit” because we believe that existing definitions of those terms are too narrow. if by “explicit” we mean relations signalled exclusively by discourse markers, then it may be the case, as sporleder and lascarides (2008) conclude, that those two types are different in nature. however, if explicit is extended to include other types of signals, and particularly semantic signals, we believe that the two types may not be that different in nature, and automatic classification may be possible (assuming, of course, complex annotation of the type carried out here, and identification of those semantic relations in unseen data). 7 discussion: relation signalling and layered annotations the first goal is this ongoing annotation effort was to investigate whether signals other than discourse markers exist for coherence relations. in this respect, we can confidently say that this is, indeed, the case: out of the 1,127 signalled relations, 878 (77.91%) contain a signal other than a discourse marker. although some of the relations (13.57% of the total 1,304) are not signalled, the overwhelming majority of them are. we would like to point out that what we have found are positive signals, that is, indicators that a relation exists. this does not mean that such signals are used exclusively to indicate that relation (as we have seen in the many-to-many correspondences). it also means that the signals, taboada and das 274 as linguistic devices, are not exclusively used to mark a relation; they may well have other purposes in the text. in a sense, this means that the signals are compatible with a relation, not necessarily indicators of the relation exclusively. one may argue that the signals that we have identified are quite intricate, and that an automatic system would have a very hard time making use of them. this is especially the case with the semantic and lexical relations, where some of the relations are identified based not only on wordnet-type relations (fellbaum, 1998), but also on world knowledge. an example from our corpus is an elaboration relation that relies on the semantic connection between the philippine company in the nucleus and luzon petrochemical corp. in the satellite. identifying that connection may require knowledge about luzon being a philippine island, which is beyond the scope of wordnet. in this paper, we are not, however, directly concerned with the issue of automatic identification. we merely wish to point out that more signals than previously found are present in many of the relations. automatic identification of relations would require some disambiguation, of the same type that is already necessary for discourse markers, some of which have nondiscourse functions (hirschberg & litman, 1993). we will devote the rest of this section to issues having to do with annotating discourse phenomena, and with the difficulties in adding annotations to an existing resource. one of our main difficulties in annotating discourse phenomena has to do with the more loose definition of what counts as a signal. although we tried to create a very detailed list of signals, and documented those signals with many examples, it is undeniable that this type of annotation is subjective. our reliability study shows a decent level of agreement between annotators. as we already discussed in section 5.2, this is often the case with published studies, and to be expected in a research group where members work closely together and under the same assumptions. the question that we would like to address here is how difficult it is in general to annotate phenomena that are more abstract than, for instance, part of speech tags (which also contain a certain level of abstraction and are by no means straightforward). our view on this is that phenomena at the discourse level are as easy or as difficult to identify as phenomena at other levels of the language. the main criterion for reliable annotation is a clear set of guidelines and, in particular, a clearly defined taxonomy. we found that distinguishing between signals that belonged in the categories “entity” and “semantic” was the basis of many of our disagreements. initially, we had reserved the category “entity” for those signals that involved reference to the same referent. the category “semantic” was reserved for semantic relations that do not necessarily involve same reference, such as synonymy. this distinction works along the lines of halliday and hasan’s (1976) grammatical versus lexical cohesion, with entity signals being close to the reference system in halliday and hasan’s grammatical cohesion. our semantic group of signals contains lexical cohesion relations, such as synonyms, antonyms and hypernyms. the problem, however, is that lexical cohesion also includes repetition of the same item which is, strictly speaking, reference to the same referent, and thus entity in our system. each one of us had made a different assumption about how to deal with this problem (one including repetition as entity, the other as semantic). one of the lessons learned in this process was to stick to the tried and true as much as possible, and rely on existing taxonomies, or else motivate our departure from them. this lesson leads us to the discussion of the other issue in the annotation of higher-level phenomena. as we have mentioned throughout the paper, we are dealing with an existing corpus, already annotated for discourse relations. we believe that this will be more and more the case, signalling in a corpus of discourse relations 275 with so many available resources already annotated for a wide range of phenomena. we found ourselves disagreeing with many of the annotation decisions in the initial corpora, from the number of relations to the definition of what an elementary unit of discourse is. the rst discourse treebank uses a very large set of 78 relations, including a high number of subtypes of elaboration. in practice, this meant that we had to keep all these distinctions in mind as we annotated. more difficult for our purposes was the fine-grained segmentation. the traditional definition of minimal unit of discourse in rst proposes that clauses should be minimal units, excluding subject and object clauses. in other words, it is mostly adverbial clauses that have a function at the discourse level. mann and thompson (1988), in this as in many other aspects, leave the door open for other definitions, if they suit the researcher’s purposes. the authors of the rstdiscourse treebank decided on a segmentation method that classifies all types of clauses as elementary discourse units (edus). in particular, noun clauses as objects of verbal processes (say, tell, claim) are considered to be units of discourse in the rst discourse treebank. carlson and marcu (2001) then proposed a new rst relation, attribution, to connect the reported speech verb and its complement. similarly, relative clauses and noun clauses that modify nouns (alson lee, who heads the philippine company…; a contract to build…) are also elementary discourse units. we found that such level of detail made our annotation quite difficult, in part because we disagree with the notion that noun and relative clauses stand in any kind of discourse relation to the words that they modify. the clause-internal relations (which specifically represent the relationships between two entities or between an entity and a proposition) mainly include attribution and elaboration. these relations are usually signalled by syntactic features. the distribution is provided in table 10. syntactic feature relation group relation infinitival clause (np) or noun clause (27) elaboration (27) elaboration-object-attribute (27) participial clause (19) elaboration (18) elaboration-additional (7), elaboration-object-attribute (10), elaboration-general-specific (1) enablement (1) purpose (1) reduced relative clause (55) elaboration (55) elaboration-additional (10), elaboration-object-attribute (45) relative clause (112) elaboration (112) elaboration-additional (26), elaboration-object-attribute (85), definition (1) reported speech pattern (223) attribution (220) attribution (220) evaluation (2) evaluation (2) statement-response (1) statement-response (1) table 10. distribution of clause-internal relations in terms of syntactic features we found our agreement in annotating these relations to be quite high, as syntactic phenomena tend to be easier to identify. nonetheless, in most cases we felt that the relation was, in fact, syntactic, rather than a discourse or coherence relation. as we performed the annotation of signals, we also found ourselves disagreeing with specific aspects of the rst annotation, such as the label for a particular relation or the nucleus-satellite assignation. these types of errors are to be expected in discourse annotation, and we do not take taboada and das 276 issue with them, as they are the result of human error and they tend to be localized. one question that arises, however, is whether we should be making corrections in cases of obvious mistakes. although that would probably make the corpus better, we have decided not to alter it, as it has become a standard in many studies. in summary, our experience shows that, although layering upon an existing annotation is challenging, the results are certainly worthwhile. we have shown that rhetorical relations have multiple signals associated with them, and we hope to be on our way to determining how those signals can be used to perform automatic identification of relations. 8 conclusions we have presented an annotation effort that adds signalling information to an existing corpus of rhetorical relations. the purpose of the study was to determine to what extent rhetorical relations carry signals that may help readers and hearers identify the relation. research so far has focused mainly on one type of signals, discourse markers, and has thus concluded that the majority of relations are implicit, that is, they contain no overt signal. we have shown that this is not the case and that, although there may still exist some implicit relations, most of the relations in our corpus are explicit, that is, they are signalled, sometimes through multiple signals. in the process of annotating the corpus, we have discovered and solved a number of issues involving creating accurate and manageable taxonomies of signals, adding information to an existing corpus, and mapping relations and signals to each other. the annotation described in this paper is a preliminary pilot study, comprising only 10% of the total corpus. in future work, we will expand to cover the entire corpus. the most important qualitative change for the rest of the annotation involves finding a method to layer annotations on top of the existing lisp-style notation for the rst corpus. the finished corpus has two clear applications. from a psycholinguistic point of view, we hope to be able to use it to determine how hearers and readers use signals to identify relations. most of the psycholinguistic studies to date have manipulated relations by adding or deleting discourse markers. it would be very useful to extend that work by changing other types of signals, to see what effects that has on comprehension. the other main application of such an annotated corpus is in discourse parsing. a great deal of recent work (hernault et al., 2010; hernault et al., 2011; mithun & kosseim, 2011; da cunha et al., 2012) and also earlier approaches (corston-oliver, 1998b; marcu, 2000a; schilder, 2002) have used discourse markers as the main signals to automatically parse relations, and almost exclusively at the sentence level. our extended set of signals, and the fact that they work at all levels of discourse, will probably facilitate this task. acknowledgements this work was supported by a grant from the natural sciences and engineering research council of canada (discovery grant 261104-2008). we would like to thank the audience at the 7 th conference of the swiss linguistic society (lugano, september 2012), and the reviewers and the editors of the special issue for very useful feedback and comments. signalling in a corpus of discourse relations 277 references afantenos, stergos d., denis, pascal, muller, philippe, & danlos, laurence. (2010). learning recursive segments for discourse parsing. international conference on language resources and evaluation (lrec). malta. asher, nicholas, & lascarides, alex. (2003). logics of conversation. cambridge: cambridge university press. baldridge, jason, asher, nicholas, & hunter, julie. (2007). annotation for and robust parsing of discourse structure on unrestricted texts. zeitschrift für sprachwissenschaft(26), 213-239. bateman, john, kamps, thomas, kleinz, jörg, & reichenberger, klaus. (2001). towards constructive text, diagram, and layout generation for information presentation. computational linguistics, 27(3), 409-449. blakemore, diane. (1987). semantic constraints on relevance. oxford: blackwell. blakemore, diane. (1992). understanding utterances: an introduction to pragmatics. oxford: blackwell. blakemore, diane. (2002). relevance and linguistic meaning: the semantics and pragmatics of discourse markers. cambridge: cambridge university press. britton, bruce k., glynn, shawn m., meyer, bonnie j. f., & penland, m. j. (1982). effects of text structure on use of cognitive capacity during reading. journal of educational psychology, 74(1), 51-61. carlson, lynn, & marcu, daniel. (2001). discourse tagging manual. manual. http://www.isi.edu/~marcu/discourse/tagging-ref-manual.pdf. carlson, lynn, marcu, daniel, & okurowski, mary ellen. (2002). rst discourse treebank, ldc2002t07 [corpus]. philadelphia, pa: linguistic data consortium. corston-oliver, simon. (1998a). beyond string matching and cue phrases: improving efficiency and coverage in discourse analysis. proceedings of aaai 1998 spring symposium series, intelligent text summarization (pp. 9-15). madison, wisconsin. corston-oliver, simon. (1998b). identifying the linguistic correlates of rhetorical relations. proceedings of the workshop on discourse relations and discourse markers, coling/acl '98 (pp. 8-14). cristea, dan, ide, nancy, & romary, laurent. (1998). veins theory: a model of global discourse cohesion and coherence. proceedings of the 36th annual meeting of the association for computational linguistics and the 17th international conference on computational linguistics (acl-98/coling-98) (pp. 281-285). montréal, canada. cunningham, hamish, maynard, diana, bontcheva, kalina, tablan, valentin, aswani, niraj, roberts, ian, . . . peters, wim. (2011). text processing with gate (6th ed.). sheffield, uk: gate. da cunha, iria, san juan, eric, torres-moreno, juan manuel, cabré, maría teresa, & sierra, gerardo. (2012). a symbolic approach for automatic detection of nuclearity and rhetorical relations among intra-sentence discourse segments in spanish. proceedings of cicling (pp. 462-474). new delhi, india. dale, robert. (1991). the role of punctuation in discourse structure. proceedings of aaai fall symposium on discourse structure in natural language understanding and generation (pp. 13-14). asilomar, ca. degand, liesbeth, & sanders, ted. (2002). the impact of relational markers on expository text comprehension in l1 and l2. reading and writing, 15(7-8), 739-758. http://www.isi.edu/~marcu/discourse/tagging-ref-manual.pdf taboada and das 278 fellbaum, christiane (ed.). (1998). wordnet: an electronic lexical database. cambridge, ma: mit press. fraser, bruce. (1990). an approach to discourse markers. journal of pragmatics, 14(3), 383-398. fraser, bruce. (1999). what are discourse markers? journal of pragmatics, 31(7), 931-952. fraser, bruce. (2006). towards a theory of discourse markers. in k. fischer (ed.), approaches to discourse particles (pp. 189-204). amsterdam: elsevier. fraser, bruce. (2009). an account of discourse markers. international review of pragmatics, 1, 293-320. grice, h. paul. (1975). logic and conversation. in p. cole & j. l. morgan (eds.), speech acts. syntax and semantics, volume 3 (pp. 41-58). new york: academic press. haberlandt, karl. (1982). reader expectations in text comprehension. in j.-f. le ny & w. kintsch (eds.), language and comprehension (pp. 239-249). amsterdam: north-holland. halliday, michael a. k., & hasan, ruqaiya. (1976). cohesion in english. london: longman. hanneforth, thomas, heintze, silvan, & stede, manfred. (2003). rhetorical parsing with underspecification and forests. proceedings of hlt-naacl 2003. edmonton, canada. hernault, hugo, bollegala, danushka, & ishizuka, mitsuru. (2011). semi-supervised discourse relation classification with structural learning. proceedings of the 12th international conference on computational linguistics and intelligent text processing (cicling '11). tokyo, japan. hernault, hugo, prendinger, helmut, duverle, david a., & ishizuka, mitsuru. (2010). hilda: a discourse parser using support vector machine classification. dialogue and discourse, 1(3). hirschberg, julia, & litman, diane j. (1993). empirical studies on the disambiguation of cue phrases. computational linguistics, 19(3), 501-530. horn, laurence. (1984). toward a new taxonomy for pragmatic inference: q-based and r-based implicature. in d. schiffrin (ed.), meaning, form and use in context: linguistic implications (pp. 11-42). washington, dc: georgetown university press. hovy, eduard, & arens, yigal. (1991). automatic generation of formatted text. proceedings of 9th aaai conference. anaheim, california. huong, le thanh. (2007). an approach in automatically generating discourse structure of text. journal of computer science and cybernetics, 23(3), 212-230. karamanis, nikiforos. (2007). supplementing entity coherence with local rhetorical relations for information ordering. journal of logic, language and information, 16, 445-464. keenan, janice m., baillet, susan d., & brown, polly. (1984). the effects of causal cohesion on comprehension and memory. journal of verbal learning and verbal behaviour, 23(2), 115126. knott, alistair. (1996). a data-driven methodology for motivating a set of coherence relations. ph.d. dissertation, university of edinburgh, edinburgh, uk. retrieved from http://citeseer.nj.nec.com/knott96data.html knott, alistair, & dale, robert. (1994). using linguistic phenomena to motivate a set of coherence relations. discourse processes, 18(1), 35-62. knott, alistair, oberlander, jon, o'donnell, michael, & mellish, chris. (2001). beyond elaboration: the interaction of relations and focus in coherent text. in t. sanders, j. schilperoord & w. spooren (eds.), text representation: linguistic and psycholinguistic aspects (pp. 181-196). amsterdam and philadelphia: john benjamins. http://citeseer.nj.nec.com/knott96data.html signalling in a corpus of discourse relations 279 knott, alistair, & sanders, ted. (1998). the classification of coherence relations and their linguistic markers: an exploration of two languages. journal of pragmatics, 30, 135-175. lapata, mirella, & lascarides, alex. (2004). inferring sentence-internal temporal relations. proceedings of the north american chapter of the assocation of computational linguistics (naacl) (pp. 153-160). boston, ma. lin, ziheng, kan, min-yen, & ng, hwee tou. (2009). recognizing implicit discourse relations in the penn discourse treebank. proceedings of the 2009 conference on empirical methods in natural language processing (pp. 343-351). singapore. louis, annie, prasad, rashmi, joshi, aravind k, & nenkova, ani. (2010). using entity features to classify implicit discourse relations. proceedings of the sigdial conference (pp. 59-62). tokyo, japan. mann, william c., & taboada, maite. (2010). rst web site, from http://www.sfu.ca/rst mann, william c., & thompson, sandra a. (1988). rhetorical structure theory: toward a functional theory of text organization. text, 8(3), 243-281. marcu, daniel. (1999). a decision-based approach to rhetorical parsing. proceedings of 37th annual meeting of the association for computational linguistics (acl'99) (pp. 365-372). college park, maryland. marcu, daniel. (2000a). the rhetorical parsing of unrestricted texts: a surface based approach. computational linguistics, 26(3), 395-448. marcu, daniel. (2000b). the theory and practice of discourse parsing and summarization. cambridge, ma: mit press. marcu, daniel, & echihabi, abdessamad. (2002). an unsupervised approach to recognising discourse relations. proceedings of the 40th annual meeting of the association for computational linguistics (acl'02) (pp. 368-375). philadelphia, pennsylvania. meyer, bonnie j. f., brandt, david m., & bluth, george j. (1980). use of top-level structure in text: key for reading comprehension in ninth-grade students. reading research quarterly, 16(1), 72-103. millis, keith k., & just, marcel a. (1994). the influence of connectives on sentence comprehension. journal of memory and language, 33, 128-147. mithun, shamima, & kosseim, leila. (2011). comparing approaches to tag discourse relations. proceedings of the 12th international conference on computational linguistics and intelligent text processing (cicling '11) (pp. 328-339). tokyo, japan. myers, jerome l., shinjo, makiko, & duffy, susan a. (1987). degree of causal relatedness and memory. journal of memory and language, 26(4), 453-465. o'donnell, michael. (1997). rst-tool: an rst analysis tool. proceedings of the 6th european workshop on natural language generation. duisburg, germany. pardo, thiago alexandre salgueiro, & nunes, maria das graças volpe. (2008). on the development and evaluation of a brazilian portuguese discourse parser. journal of theoretical and applied computing, 15(2), 43-64. pitler, emily, louis, annie, & nenkova, ani. (2009). automatic sense prediction for implicit discourse relations in text. proceedings of the 47th annual meeting of the association for computational linguistics (pp. 683-691). suntec, singapore. poesio, massimo, stevenson, rosemary, di eugenio, barbara, & hitzeman, janet. (2004). centering: a parametric theory and its instantiations. computational linguistics, 30(3), 309363. http://www.sfu.ca/rst taboada and das 280 polanyi, livia, culy, christopher, van der berg, martin, thione, gian lorenzo, & ahn, david. (2004). a rule based approach to discourse parsing. proceedings of the 5th sigdial workshop in discourse and dialogue (pp. 108-117). cambridge, ma. polanyi, livia, & scha, r. j. h. (1983). the syntax of discourse. text, 3(3), 261-270. prasad, rashmi, joshi, aravind k, & webber, bonnie. (2010). realization of discourse relations by other means: alternative lexicalizations. proceedings of coling 2010 (pp. 1023-1031). beijing. prasad, rashmi, lee, alan, dinesh, nikhil, miltsakaki, eleni, campion, geraud, joshi, aravind k., & webber, bonnie. (2008). penn discourse treebank version 2.0, ldc2008t05 [corpus]. philadelphia, pa: linguistic data consortium. prasad, rashmi, miltsakaki, eleni, dinesh, nikhil, lee, alan, joshi, aravind k, robaldo, livio, & webber, bonnie. (2007). the penn discourse treebank 2.0 annotation manual. philadelphia: university of pennsylvania. sanders, ted, land, jentine, & mulder, gerben. (2007). linguistic markers of coherence improve text comprehension in functional contexts. information design journal, 15(3), 219-235. sanders, ted, & noordman, leo. (2000). the role of coherence relations and their linguistic markers in text processing. discourse processes, 29(1), 37-60. sanders, ted, spooren, wilbert, & noordman, leo. (1992). toward a taxonomy of coherence relations. discourse processes, 15(1), 1-35. sanders, ted, spooren, wilbert, & noordman, leo. (1993). coherence relations in a cognitive theory of discourse representation. cognitive linguistics, 4(2), 93-133. schiffrin, deborah. (1987). discourse markers. cambridge: cambridge university press. schiffrin, deborah. (2001). discourse markers: language, meaning and context. in d. schiffrin, d. tannen & h. e. hamilton (eds.), the handbook of discourse analysis (pp. 54-75). malden, ma: blackwell. schilder, frank. (2002). robust discourse parsing via discourse markers, topicality and position. natural language engineering, 8(2/3), 235-255. scott, donia, & de souza, clarisse sieckenius. (1990). getting the message across in rst-based text generation. in r. dale, c. mellish & m. zock (eds.), current research in natural language generation (pp. 47-73). london: academic press. siegel, sidney, & castellan, n.j. (1988). nonparametric statistics for the behavioral sciences. new york: mcgraw-hill. spooren, wilbert. (1997). the processing of underspecified coherence relations. discourse processes, 24, 149-168. spooren, wilbert, & degand, liesbeth. (2010). coding coherence relations: reliability and validity. corpus linguistics and linguistic theory, 6(2), 241-266. sporleder, caroline, & lascarides, alex. (2005). exploiting linguistic cues to classify rhetorical relations. proceedings of recent advances in natural language processing (ranlp-05). borovets, bulgaria. sporleder, caroline, & lascarides, alex. (2008). using automatically labelled examples to classify rhetorical relations: an assesment. natural language engineering, 14(3), 369-416. stede, manfred. (2007). korpusgestützte texanalyse: grundzüge der ebenen-orientierten textlinguistik. tübingen: gunter narr. stede, manfred. (2012). discourse processing: morgan & claypool. signalling in a corpus of discourse relations 281 taboada, maite. (2004). building coherence and cohesion: task-oriented dialogue in english and spanish. amsterdam and philadelphia: john benjamins. taboada, maite. (2006). discourse markers as signals (or not) of rhetorical relations. journal of pragmatics, 38(4), 567-592. taboada, maite. (2008). sfu review corpus [corpus]. vancouver: simon fraser university, http://www.sfu.ca/~mtaboada/research/sfu_review_corpus.html. taboada, maite. (2009). implicit and explicit coherence relations. in j. renkema (ed.), discourse, of course (pp. 127-140). amsterdam and philadelphia: john benjamins. taboada, maite, & mann, william c. (2006). rhetorical structure theory: looking back and moving ahead. discourse studies, 8(3), 423-459. teijssen, daphne, van halteren, hans, verberne, suzan, & boves, lou. (2008). features for automatic discourse analysis of paragraphs. proceedings of the 18th meeting of computational linguistics in the netherlands (clin 2007) (pp. 53-68). leuven, belgium. tofiloski, milan, brooke, julian, & taboada, maite. (2009). a syntactic and lexical-based discourse segmenter. proceedings of the 47th annual meeting of the association for computational linguistics (pp. 77-80). singapore. trabasso, tom, & sperry, linda l. (1985). causal relatedness and importance of story events. journal of memory and language, 24(5), 595-611. http://www.sfu.ca/~mtaboada/research/sfu_review_corpus.html microsoft word artwalk referents.final.4.20.21.docx dialogue & discourse 12(1) 45-72 doi: 10.5210/dad.2021.103 ©2020 kris liu, j. trevor d’arcey, marilyn walker, and jean e. fox tree this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). referential communication between friends and strangers in the wild kris liu kyliu@ucsc.edu university of california, santa cruz j. trevor d’arcey jdarcey@ucsc.edu university of california, santa cruz marilyn walker mawalker@ucsc.edu university of california, santa cruz jean e. fox tree foxtree@ucsc.edu university of california, santa cruz editor: barbara di eugenio submitted 02/2018; accepted 12/2020; published online 04/2021 abstract the map task (anderson et al., 1991) and tangram task (clark & wilkes-gibbs, 1986) are traditional referential communication tasks that are used in psycholinguistics research to demonstrate how conversational partners mutually agree on descriptions (or referring expressions) for landmarks or unusual target objects. these highly controlled, laboratory-based tasks take place under conditions that are relatively unusual for naturally-occurring conversations (speed, wnuk, & majid, 2016). using the artwalk task (liu, fox tree, & walker, 2016) – a real-world-situated blend of the map task and tangram task – we showed that the process of negotiating referring expressions “in the wild” is similar to the process that takes place in a laboratory. in addition to replicating laboratory results showing lexical entrainment, we also found that acquaintanceship and extraversion influenced the number of unique descriptors used by pairs. in round 1, introverts in stranger pairs used fewer descriptors but introverts in friend pairs were indistinguishable from extraverts. the influence of extraversion declined by round 2. lexical entrainment observed in labs is generalizable to realworld settings, and lexical entrainment in naturalistic communication, at least, is subject to social and personality factors. keywords: lexical entrainment, acquaintanceship, referential communication tasks, outdoor tasks, extraversion 1 introduction many laboratory-based tasks have characteristics that are not found in naturalistic settings. they take place in relatively sterile environments where distractions are discouraged, and participants are frequently separated from other people who are not involved in the experiment. participants are often seated at computers in booths, a situation that more closely resembles an exam than a conversation. while participant pairs engaging in lab-based discourse tasks may occasionally wander off-topic in their talk, there is usually little to tempt them to do so. in contrast to laboratory communication tasks, many real-world conversations take place amid noise, people, and other distractions. the following excerpt from the current study shows the sort of side conversation that is unlikely to happen without external distractions: a matcher who is on a mobile phone on a public liu, d’arcey, walker, and fox tree 46 street is receiving instructions from a remote director about the location and description of a piece of public art (in this and later examples, d refers to the director and m refers to the matcher): (1) m: canvassers are everywhere d: did he hassle you for being on the phone or something? m: [laugh] yeah like this guy came up to me and like tried to run in front of me d: what? m: i was just like, “i’m doing a psychology experiment” and he was like, “oh you’re doing the fake phone thing?” and i was like, “what? i’m on the phone” [pause] i could never do that to someone our investigation shows that lexical entrainment, a reliable finding in laboratory settings, also occurs in more naturalistic settings. in the following example, the director proposes the word blob, among other descriptions, to describe an art object: (2) d: okay so we’re looking for uh this like concrete thing i don’t know how to explain it m: it’s a concrete thing? d: yeah okay m: um is it is it a sculpture? d: uh yes it’s like a big just blob i couldn’t really describe it to yoit looks almost like a jellyfish when the matcher finds the object, the matcher shows acceptance of the label blob by saying, “i see the blob i see the blob.” later in the experiment, the director instructed the matcher to find the object again with the instructions, “do you remember the big blob?” which is a shorter version of the original description of a blob that was a “concrete thing” that “looks almost like a jellyfish.” the matcher entrained on the label blob saying, “okay i’m taking a picture of the blob right now.” the above examples took place during a referential communication task, a traditional experimental paradigm that is used to study collaborative communication. in one referential communication task, the map task (anderson et al., 1991), directors and matchers are given two different maps: maps share some features (but not all), and the director’s map has a specific route drawn on it whereas the matcher’s map has no route. it is the director’s job to verbally guide the matcher to draw the director’s precise route through the landmark without actually showing the matcher the director’s own map. the tangram task has a similar asymmetric setup: the director has a set of tangrams (or abstract shapes) in a specific order and the matcher has access to the same tangrams (plus some distractor tangrams). the director must verbally guide the matcher to identify the same tangrams the director sees and get the matcher to place them in the same order as they are shown in the director’s set (clark & wilkes-gibbs, 1986; schober & clark, 1989). the task is often done more than once in an experimental session, allowing researchers to observe how collaborative behavior changes across iterations or rounds. both experimental paradigms clearly show that directors and matchers work toward common understanding (grounding, e.g., clark, 1996) including coming to use (and re-use) the same terms for objects shared between their asymmetric perspectives (entrainment, a term used here mostly synonymously with clark’s notion of collaboration on referential language). we developed a naturalistic task that is conceptually similar to the wayfinding map task and object-identifying tangram task but also situates the matcher in a walkable downtown area while the director directs the matcher to a series of public artwork targets via a skype-to-mobile phone connection. by taking the matcher out of the lab and making the director contend with a matcher who is in a complex environment, the task more closely resembles the natural setting of many conversations, particularly those conducted with mobile devices. the mobile demands of the task require matchers to tune out surrounding noise and conversation. they need to consider street crossings (neider, mccarley, crowell, kaczmarski, & kramer, 2010) as well as social constraints. for instance, talking on a mobile phone in public for an extended amount of time can be viewed as being exasperating or impolite (love & perry, 2004). attending to social concerns while moving about a city can reduce attention to a mobile phone task (oulasvirta, tamminen, roto, & friends and strangers in the wild 47 kuorelahti, 2005). in addition, walking can negatively impact cognitive performance (beauchet, dubost, herrmann, & kressig, 2005; hill, bohil, lewis, & neider, 2013). and conversely, engaging in cognitive tasks can slow down walking or change a person’s gait (yogev-seligmann, hausdorff, & giladi, 2008). the presence of bystanders may also influence participants’ behavior. in laboratories, participants may be observed by an experimenter or they may be left alone in a room or booth; in the case of artwalk, only directors were sitting alone in a booth – matchers were out on a busy street, walking or standing near a variable number of bystanders. asking for directions on a mobile phone may not be unusual but attempting to describe an abstract piece of artwork to someone who seems to require multiple different descriptions, snapping a photo, and then repeating the process with another piece of art is not typical street behavior. this has the potential to exacerbate participant self-consciousness, which can influence behavior (froming & carver, 1981). nevertheless, it should be noted that the relationship between laboratory and non-laboratory behavior can be hard to predict. for example, laboratories may seem to naturally engender more formal or deliberate speech than non-laboratory settings but not all experimental linguistic work has supported this (cf. xu, 2010). in our own laboratory, deliberate effort by experimenters was required to create a formal enough atmosphere to observe changes in speakers’ uses of the quotation devices said and like (blackwell & fox tree, 2012). our first goal with the artwalk task (liu, fox tree, & walker, 2016) was to assess whether entrainment patterns (mutual adoption of common terms, as measured by their repeated use) observed in laboratories could be replicated in a more naturalistic setting involving one mobile communicator and one non-mobile communicator. it is possible that the distractions, social concerns, and different pacing of the task in a naturalistic setting could affect lexical entrainment. our second goal was to test whether we could observe differences in communication based on two social factors: acquaintanceship status and extraversion. acquaintanceship was part of the original design of the map task (anderson et al., 1991), although few acquaintanceship differences were documented (see acquaintanceship and communication section). extraversion is not often studied in the context of referential communication (see extraversion in collaboration and conversation section). we hypothesized that both acquaintanceship and extraversion would affect communication in this naturalistic setting. with this novel, real-world task we reproduced effects observed in laboratories and additionally demonstrated that some interpersonal and personality variables can affect how efficient pairs are at completing a referential communication task. 1.1 referential communication in the laboratory referential communication tasks in the laboratory have yielded consistent results on the process of coordinating referring expressions between conversational partners. when two conversational partners talk about objects that are not easily named, they undergo a period of negotiation where they have to ensure that they are both discussing the same object (grounding) and then implicitly or explicitly agree on what to call it (entrainment). over time, referring expressions shorten and communication becomes more efficient (clark & wilkes-gibbs, 1986). for example, a participant may initially identify a specific tangram as “all right, the next one looks like a person who’s ice skating, except they’re sticking two arms out in front” but later refer to it as “the ice skater” (clark & wilkes-gibbs, 1986, p. 12). participants entrain on the labels used to refer to objects, such as observed with the use of the label blob in example 2. many factors can influence grounding, including the level of expertise with the topic under discussion (isaacs & clark, 1987), the level of expertise with the communicative medium (fox tree, mayer, & betts, 2011), culture (wang, fussell, & setlock, 2009; wu & keysar, 2007), and age (horton & spieler, 2007; kemper, othick, warren, gubarchuk, & gerhing, 1996). in this study, we examine the influence of acquaintanceship and extraversion. liu, d’arcey, walker, and fox tree 48 1.2 acquaintanceship and communication whether or not interlocutors know each other changes communicative behavior. in some previous studies investigating the differences between friends and strangers, friends outperformed strangers. friends were better at guessing the thoughts and feelings of friends (stinson & ickes, 1992) and they more easily identified a target figure when described by a friend than by a stranger (fussell & krauss, 1989). friends were also better at sending covert messages to each other without the covert message being understood by strangers (fleming, darley, hilton, & kojetin, 1990). retrieval cues generated by friends were more helpful for word recall than those generated by strangers (andersson & rönnberg, 1995). in other previous studies investigating the differences between friends and strangers, being friends was found to negatively impact performance. friends tended to project their own knowledge and beliefs on each other, resulting in incorrect assumptions of clarity (savitsky, keysar, epley, carter, & swanson, 2011). they were also more likely to overestimate the success of their performance as a pair in comparison to strangers (gould, osborn, krein, & mortenson, 2002). in recall tasks, friends were more likely to induce false memories in each other (hope, ost, gabbert, healey, & lenton, 2008). beyond performance, the structure of discourse was also found to differ between friends and strangers. friends were more likely to use informal language, to interrupt, to talk about multiple topics, to self-disclose, to be judgmental, and to ask for favors (planalp & benson, 1992). friends also used more discourse markers (fox tree, 2007; jucker & smith, 1998), overlapped more turns (bortfeld, leon, bloom, schober, & brennan, 2001; dunne & ng, 1994), laughed more (smoski & bachoroski, 2003), and left more information unspoken in openings and had more complex closings (hornstein, 1985). yet some researchers found no discernible effect of acquaintanceship on some discourse phenomena, including no effect on disfluencies (bard, aylett, & lickley, 2002; branigan, lickley, & mckelvie, 1999), little effect on prosodic convergence (truong & heylen, 2012), and no effect on the amount of laughter (truong & trouvain, 2012). the inconsistent effect of acquaintanceship on discourse may account for why few researchers, after thirty years and hundreds of citations, have documented differences between friend and stranger pairs in the map task corpus, although an even split between friend and stranger pairs was part of the original design of the study (anderson et al., 1991). the lack of differences between friends and strangers in the map task corpus suggests that acquaintanceship will not affect grounding and entrainment behaviors with the artwalk task. nonetheless, acquaintanceship differences may be more apparent in a more natural setting. if we do find differences based on acquaintanceship in the wild, this raises the question of whether lab settings dampen or extinguish acquaintanceship differences. lab settings may have that effect because they focus participants’ attention on the tasks at hand. with the artwalk task, participants often talked freely about off-task topics as they walked between art objects (guydish, d’arcey, & fox tree, 2020). 1.3 extraversion in collaboration and conversation extraversion is one of the most readily and reliably recognizable personality traits at zero acquaintance (albright, cohen, malloy, christ, & bromgard, 2004; kenny & acitelli, 2001; levesque & kenny, 1993) with social and linguistic behaviors that are identifiable even when not in face-to-face interaction (gill & oberlander, 2003). although level of extraversion is a continuous scale, in the following discussion, we refer to extraverts and introverts as shorthand for those who score higher or lower on the scale. extraverts have an advantage for some types of laboratory studies, particularly those that involve greater external pressure. they performed better than introverts on verbal working memory tasks that had time limits (rawlings & carnie, 1989). they also performed better on both practiced friends and strangers in the wild 49 and novel tasks when there was someone observing them (uziel, 2007). in groups, extraverts generated more unique and diverse ideas than introverts (jung, lee, & karsten, 2012), although the presence of highly extraverted participants on teams did not always translate to better team performance (peeters, van tuijl, rutte, & reymen, 2006). in discourse, extraverts tended to take the lead in conversations (cuperman & ickes, 2009). they said more, spoke louder, spoke faster, used fewer pauses, and used more backchannels (campbell & rushton, 1978; carment, miles, & cervin, 1965; dewaele & furnham, 1999; feldstein & sloan, 2984; gifford & hine, 1994; scherer, 1978). extraverts’ speech had lower lexical richness and used a greater proportion of verbs, adverbs, and pronouns (dewaele & furnham, 1999). introverts tended to choose their words more carefully, qualifying their statements more frequently (oberlander & gill, 2006; pennebaker & king, 1999). introverts also used more descriptive and concrete language (beukeboom, tanis, & vermeulen, 2013), although extraverts used more adjectives than introverts (gill & oberlander, 2003). in casual conversations and in formal examinations, extraverts relied on greater shared knowledge between themselves and their interlocutors (heylighen & dewaele, 1999). they were more likely to use deictic expressions (e.g., pronouns such as he) than introverts, whose language tended to be more explicit about referents (heylighen & dewaele, 1999). in other words, extraverts were more likely than introverts to tailor their speech to the context of a conversation. this tailoring can also be observed with spontaneous gestures that extraverts produce (tolins, liu, neff, walker, & fox tree, 2016). prior work on language differences between extraverts and introverts make no clear predictions about how grounding and entrainment behaviors differ based on this personality variable. nonetheless, extraversion differences, if they exist, may be more likely to manifest in the naturalistic artwalk task where participants had more opportunity to engage in off-task dialogue. 1.4 acquaintanceship and extraversion personality and acquaintanceship can interact in discourse. in one study, extraverts were better at conducting conversations with strangers (thorne, 1987), possibly due to their increased willingness to take the lead (cuperman & ickes, 2009). but within friend pairs, introverts were quite assertive, particularly when both friends were introverts (nelson, thorne, & shapiro, 2011). in mixed pairs, however, only 40% of introverts were reported as being as engaged in speaking as their extravert friends (nelson et al., 2011). as was observed in the laboratory, acquaintanceship and extraversion may also interact in our study. for example, personality variables may be more pronounced in stranger pairs than in friend pairs. 1.5 current study we tested the role of acquaintanceship status and extraversion on communicative efficiency during a real-world-situated referential communication task. communicative efficiency is an observable dimension of entrainment and was defined as the shortening of referential descriptions. shortening of referential descriptions is a linguistic behavior that has been shown repeatedly in laboratory referential communication tasks (clark & wilkes-gibbs, 1986; schober & clark, 1989; tolins, zeamer, & fox tree, 2018) as well as in an outdoor task that took place in a controlled experimental setting (brennan, schuhmann, & batres, 2013). we predicted that shortening would also occur in a far more complex, real-world referential communication task conducted in distracting communicative circumstances. a traditional assessment of communicative effectiveness in referential communication is the tangram task. this task involves assessing the number of correct identifications from a dozen identifications (clark & wilkes-gibbs, 1986), often from a larger array (fox tree, 1999; fox tree & clark, 2013), with a game that is played repeatedly, such as with eight iterations (tolins et al., 2018). in the traditional task, directors and matchers see all the two-dimensional items-to-be-placed liu, d’arcey, walker, and fox tree 50 at the same time, often on a table in front of them or on a computer screen. in the artwalk task, directors saw items-to-be-placed one at a time and matchers saw multiple pieces of public art as they walked around the downtown area. because the in-the-wild setting did not allow as many identifications, nor as many rounds, we used the number of director descriptors rather than the number correct as the entrainment measure. in example 2, descriptors would include blob, concrete thing, sculpture, and jellyfish. we chose the number of director descriptors instead of other measures (e.g., number of matcher descriptors) because directors’ descriptors were less subject to cell phone reception problems and because director and matcher descriptors were correlated and therefore similar (see coding section). a higher number of descriptors indicates greater difficulty communicating about the object. as the number of descriptors shorten, communication becomes more efficient. if entrainment occurs in this naturalistic setting, round 2 should have fewer director descriptors for any target object than in round 1. that is, directors’ descriptions in round 2 will be shorter if and only if entrainment has occurred in round 1. this is an oft-reported finding in the literature, which is demonstrated in laboratories by showing that descriptions shorten over rounds in dialogues (with entrainment) but not monologues (no entrainment; e.g., clark & wilkes-gibbs, 1986). the important predictor variable is the round, not the primes (what one person says) and targets (what the other person says). research on entrainment has used both rounds and primes within a dialogue. we use rounds here. the measure of entrainment – fewer descriptors in round 2 than in round 1 – accommodates implicit agreement across interlocutors. most of the time, deciding the label for an object (e.g., a blob) is not accomplished with explicit statements (e.g., “let’s call this a blob”) nor explicit confirmations (e.g., “do you agree?” “yes, i do”). it is accomplished by reducing the number of descriptors used (e.g., blob, concrete thing, sculpture, and jellyfish in round 1 and blob in round 2, as in example 2). in past work, measurement of verbal agreement across interlocutors (such as displayed by the reduction of descriptions across rounds) has not been a typical assessment of entrainment; the typical assessment is accuracy of object identification, such as how many objects were correctly picked out of an array. one reason verbal agreement has not been used is that it requires a high amount of judgement calls from coders. to illustrate, when one interlocutor uses the phrase “like a t,” while the other interlocutor uses the phrase “t-shaped,” should these be counted as the same descriptor, or different ones? when one interlocutor uses the phrase “copper thing” and the other uses each word separately (“thing” at one point and “copper” at another point), is that two examples of entrainment or one? accuracy in this kind of assessment relies on human judgements: a computer may be able make these comparisons consistently, but its matches would be extremely sparse. in the methodology section we provide analyses of how often directors’ and matchers’ word choices overlapped. then, to measure entrainment, we used the number of director descriptors as our main dependent measure. in addition to assessing whether director descriptors shortened across rounds, we examined whether participants’ acquaintanceship status (they were either friends for at least one academic quarter or complete strangers) and extraversion influenced communicative efficiency. because the successful completion of referential communication tasks is thought to be reliant on the creation of conceptual pacts and shared common ground (brennan & clark, 1996), we predicted that friends would be more efficient than strangers (analysis 1). additionally, we tested how communicators behaved across multiple rounds of the task (rounds 1 and 2). round 1 relied on directors’ and matchers’ abilities to negotiate a referring expression for a novel target with each other, a process which was made more challenging for matchers walking around downtown without a map. round 2 relied on directors using descriptions of previously identified targets in ways that matchers would recognize, a process that was dependent on the participants’ remembering the negotiated label. we expected round 1 performance to be more likely to involve interpersonal and personality factors (analysis 2) and round 2 performance to be more influenced by conversational history, such as how quickly the pair found the targets the first time (analysis 3). friends and strangers in the wild 51 2 methodology in this section, we provide details on the methods we used for this study. 2.1 participants forty-eight pairs of uc santa cruz participants’ audio recordings were analyzed in this study (24 friend pairs and 24 stranger pairs). those recruited to be in friend pairs were asked to bring a friend, which was defined as someone they spoke to regularly and had known for at least one academic quarter (10 weeks). an additional 26 pairs’ recordings were not assessed because they failed to find the minimum number of targets (i.e., at least 8 of 10) needed to ensure a similar number of trials across pairs (19 pairs) or because they experienced experimenter, participant, or technical errors such as equipment failing to record, participants’ taking photos in the wrong order, or experiencing particularly poor cellular connection quality during the task, which included cases where the phone call was dropped more than once (7 pairs). participants received either course credit or a $10 gift card after participation. participants were screened on native language so only native english speakers were included in this study. table 1 shows the gender composition of the 48 analyzed pairs, with females outnumbering males (58% to 42% respectively). table 1. gender composition of analyzed pairs (n = 48). female director male director female matcher 21 (43.8%) 7 (14.6%) male matcher 8 (16.7%) 12 (25%) 2.2 materials downtown santa cruz, california, features many public art installation pieces such as sculptures, murals, and mosaics. about 40 pieces of abstract and non-abstract art were initially identified as potential targets. ten research assistants were asked to describe each object in detail and the number of descriptive words used (descriptors) was tallied and averaged. the objects were ranked by number of descriptors used; the ten with the greatest variety of descriptors were chosen as targets, to ensure that multiple conversational turns would be required to identify the objects. all ten targets were located within a two-by-six block area. most targets were close to other pieces of art that often needed to be explicitly eliminated by the matcher as a potential target; this sometimes included finding specific, unique sections of murals and multi-panel mosaics. during the yearlong data collection process, three pieces were removed by the city. we replaced those targets with new pieces that were geographically close to the originals. the 10 selected targets were split into two lists (a or b) with five targets each, with an attempt to balance the maximum distance participants would be required to walk in order to take photographs of all five targets. 2.3 procedure director/matcher pairs worked together to find art targets in downtown santa cruz. directors were located in a campus lab and matchers were located downtown. there were two rounds of art identification. in each round, the director received five targets in sequence, describing each one to liu, d’arcey, walker, and fox tree 52 the matcher. when the target was found, the matcher took a photo of it and the director hit a key to advance to the next art object. the experimenters in the two locations both had their own cell phones which they could use to indicate issues with the study, such as a late participant or study cancellation due to rain. directors met an experimenter in a campus lab while matchers met an experimenter at a café a block away from the center of the downtown area where all targets were located. before establishing a call between participants, experimenters gave a short orientation for the devices the participants would be using. after the call was established, the experimenters left the participants alone to do the practice trial and experiment (the campus experimenter waited in another room while the downtown experimenter stayed behind at the café). the practice trial used a non-abstract statue as a target near the center of the downtown area, allowing participants to become acclimated to the equipment and task setup while also placing them close to the trial targets. directors were connected to matchers via voice-only skype in an enclosed computer booth. directors used the skype application on a computer, but matchers received the call as if it were a typical cell carrier-based voice call. matchers carried a smartphone and a separate digital camera. both participants were given the option of wearing headsets, but most participants opted against them (matchers spoke holding the phone to their ears, directors were on speaker phone). the conversation was recorded using audio hijack pro 2, an in-line audio recording program. during an experiment trial, directors were shown a photo of a single target object alongside a map with the target’s location highlighted in green (see figure 1). photos of the target stimuli and a map were presented to the director using superlab 4.5 (haxby, parasuraman, lalonde, & abboud, 1993). the images were displayed on a single computer screen. the map of downtown santa cruz used google maps’ satellite imagery, but with indicators manually inserted to show the locations of targets. matchers did not see photos of the stimuli. matchers were also not given maps. matchers’ self-reported familiarity with downtown santa cruz was not significantly correlated with the number of descriptors or turns they or their partner used, suggesting that they were indeed reliant on the director to get them to where they had to be and/or had little preexisting knowledge of specific art installations. each new target was accompanied with the same map but with new locations highlighted and old locations in grey. target order was randomized, and each target was given an 8-minute time limit. the goal of the time limit was to prevent the entire hour from being spent on searching for one piece of art. (we note here that five objects found twice with a maximum of 8 minutes each equals 80 minutes, which was 20 minutes beyond the hour allowed for the activity.) if directors failed to hit the key to advance to the next target within 8 minutes, the experiment automatically moved on to the next target and participants had to stop their search for the timed-out target. this was done in order to motivate participants to progress as much as possible through the ten trials, rather than getting stuck on one. after participants searched for all five targets once (round 1), targets were re-randomized and presented again for the pairs to find a second time (round 2). participants were instructed to not deviate from the order set for them; the handful of pairs that did were excluded from analyses. matchers were instructed to take pictures of the targets as they located them, which were then checked by the experimenters after the experimental session was concluded. although directors had no access to the photos, directors and matchers confirmed verbally that they had found the right targets. the fact that directors had no access to the photos conceptually mirrors both the map task and the tangram task, neither of which typically confirms for participants that they have arrived at the same understanding with evidence beyond verbal confirmation (for example, directors do not typically see the tangrams matchers select). generally, matchers were at ceiling in finding targets. when they had gone through both rounds, directors led matchers back to the café, after which the call was disconnected. both participants were separately given post-experiment questionnaires which included a single item asking them to rate their familiarity with the downtown area on a 7 point likert-type scale, five items on their experience doing the study, as well as the ten item friends and strangers in the wild 53 personality inventory (tipi; gosling, rentfrow, & swann, 2003) which was used to assess extraversion. participants were asked to rate themselves on pairs of personality-related words on a scale of 1 (strongly disagree) to 7 (strongly agree). figure 1. an example of a director’s screen during the task. the map is non-interactive and has grey indicators for potential target locations. the green indicator shows the location for the current target and the red indicator shows the experiment starting point. 2.4 coding we begin this section by describing the data included in analyses, and then how the data was coded to measure communicative efficiency. the nature of the planned analysis required that only the data for pairs who had found at least 8 of the 10 targets would be analyzed for this study, at least four targets in each of two rounds. inability to find a target often had a cascading negative effect on performance: participants generally either became flustered when the experimental procedure moved on after eight minutes spent on a target, or they ignored the time limit and kept looking for the target, which then threw off timing of the rest of the trials. table 2 shows the number of friend pairs and stranger pairs who found 8, 9, and 10 targets. table 2. number of targets found by pairs. number of targets friend pairs (n = 24) stranger pairs (n = 24) 8 targets 2 7 9 targets 8 4 10 targets 14 13 liu, d’arcey, walker, and fox tree 54 in the participants section we noted that participants who had problems with their photos were excluded. the problem with including pairs who took photos in the wrong order is that there was no way to retroactively determine whether they found the correct objects on each trial or if they had found targets of prior trials during later trials. there were also problems with participants who took overly wide shots because these shots included multiple potential targets. there were two dependent measures available to quantify communicative efficiency, the number of directors’ descriptors per target and the number of matchers’ descriptors per target. these two variables were moderately to strongly correlated with each other (round 1 r = .23; round 2, r = .58), suggesting that they would operate similarly for our analyses. we chose directors’ descriptors because they were less sensitive to the pitfalls of cellular communication on a street; for example, there were several times that matchers’ reception cut out on the recordings. research assistants annotated director descriptors, matcher descriptors, and counted the number of turns taken for each piece of target artwork found. descriptors were defined as uniquewithin-round adjectives or nouns that were descriptive of the target objective, such as colors, shapes, media (e.g., painting, sculpture, metal, stone) and patterns (e.g., striped). coders were told that descriptors should generally be single or hyphenated words, such as “purple” and “eggshaped.” despite this, coders found that many descriptors required longer phrases. adjectives that referred to a portion of the artwork needed to retain their noun, e.g., “green background,” “multiple curves,” and “pointy nose.” non-hyphenated compound nouns needed to be included, e.g., “stick figures,” “triangle shape.” other phrases were used to disambiguate between attributes of the work itself and its position within the greater context, e.g., “on top of a platform,” “on a cinderblock,” “part of a bigger mural.” even longer phrases simply lost meaning when they were artificially shortened, e.g., “if it had hands, they would be doing spirit fingers.” allowing research assistants to be more dynamic with their coding made sense: even in the tangram task, descriptors are longer in the first trial (e.g., “a person who’s ice skating” vs. “the ice skater”). to avoid arbitrarily determining what counted as synonyms, similar descriptors were counted separately. for example, the uses of “turquoise,” “blue,” and “kind of blue” were treated as unique to allow them to be considered as terms that could be entrained upon despite a possible lack of reuse of the term by a matcher (as when a matcher responds to a series of descriptors by accepting but not repeating them, such as by saying “ok”). our goal was to capture the widest range of descriptors without overreliance on judgement calls (e.g., is turquoise the same as blue?). each unique descriptor was only counted once per trial. this is due to the constraints of running a task “in the wild”: noisy streets and occasionally poor cellular connection meant that there was a lot of repetition of single words or short phrases by both director and matcher. counting repeated words as separate instances of entrainment would conflate entraining on descriptors with ensuring correct audio transmission, adding unwanted noise to the analysis. to determine whether this coding scheme was similar across coders, two research assistants coded round 1 director descriptors for 12 of the 48 participant pairs. because descriptor coding was an open-response task, we chose three measures to understand how similarly descriptors were coded. first, we used a pearson correlation to determine whether the number of descriptors coded by one research assistant for one trial would predict the number of descriptors coded by the other research assistant for the same trial. a positive correlation would suggest that when one coder recorded more descriptors for a trial, the other coder would in general also record more descriptors. a count-based analysis of similarity between descriptors is especially useful in this context, as descriptor count is also our measure of entrainment. the number of descriptors that each research assistant coded was highly correlated with the other research assistant’s coding, r(58) = .876, p < .001. second, we performed a flexible human-based procedure where a third research assistant compared the two coders’ work using a strict-matching method (similar to how a python script might compare strings, but in a slightly more forgiving way) and a flexible-matching method that asked for just the meaning to be similar enough (even more forgiving). each of these match coding friends and strangers in the wild 55 schemes is described in detail below. the match coding instructions that were given to research assistants are included in appendix a. the strict method of coding was similar for single-word and multiple-word descriptors. in the strict method of coding single word descriptors, the third research assistant determined whether each descriptor was reported by both coders. when close or partial matches occurred, they determined whether the meaning was similar enough to be treated as a match based on whether the root of the word was the same, such as “spikes” and “spiky” or “yellow” and “yellowish.” for multiple-word descriptors, they determined whether content words were identical. the third research assistant followed the example that “surrounded by plants” and “plants are surrounding it” have two identical content words, plants and surround. we also specified that negation within one of two multiple-word descriptors (e.g., “holding hands” and “not holding hands”) should not be counted as matching. they compared single-word descriptors to multiple-word descriptors in the same way that they compared multiple-word descriptors. the third research assistant’s analysis found 38.98% agreement between the two original coders, averaged across all 12 analyzed participant pairs, for the strict method of coding. strict descriptor matching is likely to under-represent the actual levels of entrainment due to its rigid rule set. for example, the strict rule set would not allow “rocks” and “stones” to match. for a deeper look at inter-coder overlap, the third research assistant made more thoughtful judgment calls on whether a descriptor’s meaning was similar enough to be considered the same. using these rules, the third research assistant’s analysis found 49.04% agreement averaged across items. finally, in order to achieve a more objective measure of the similarity between the research assistants’ descriptor coding while also factoring in semantic content, we relied on recent work showing that word embeddings provide state of the art results on semantic textual similarity tasks, such as our task of comparing the lists of phrases in our descriptor pairs (liu, ott, goyal, du, joshi, chen, levy, lewis, zettlemoyer, & stoyanov, 2019). word embeddings capture semantic similarities between words, generalizing beyond the particular words used in the descriptors (reimers & gurevych, 2019; li et al 2020)). using roberta large embeddings1, we measured the cosine similarity of the document embeddings of the lists of director descriptors coded by the research assistants. we first directly compared descriptors from the same conversation for each artwalk object as illustrated by table 3. over all 60 samples of “same artwalk object, same conversation” we get a high average similarity (m = .80, sd =10). we then compared these measures of cosine similarity with a random selection of 60 descriptions given by one research assistant matched to a description given by the other research assistant for the same artwalk object, but from a different conversation (different dyads). this is illustrated by the pairs of descriptor lists in table 4. the cosine similarities for “same artwalk object, different conversation” were significantly lower (m = .44, sd = .15, t(59) = 15.65, p < .001). this difference clearly shows that research assistant coders coded more similarly when coding the same conversation than coding different conversations about the same artwalk object. in order to establish a random baseline for cosine similarity for artwalk objects, we also created a random selection of 60 descriptions from one research assistant and 60 descriptions from the second research assistant, for different artwalk objects in different conversations. see table 5 for examples. here the cosine similarities for “different artwalk object, different conversation” were again significantly different (m = .27, sd = .12, t(59) = 26.65, p < .001). while it is challenging for humans to agree on descriptors, embeddings show that the descriptors they report are similar at levels far above chance. we also note that our main research question (whether entrainment occurs) will be measured by the number of descriptors coded, rather than what those descriptors actually are. 1 https://github.com/martinomensio/spacy-sentence-bert/ liu, d’arcey, walker, and fox tree 56 table 3. semantic textual similarity scores for the same object within dyads. object coder 1 coder 2 similarity emo penguin emo penguin, spray painted, beanie, white, sad emo, penguin, spray painted, outline, beanie, white, 0.922 trumpet bench across from parking garage, middle of the block, mural, wooden bench, dog legs, instrument, trumpet, red and blue ribbons, painting on cement, animal legs, carved legs, tree, white low fence mural, painting, bigger, wooden, bench, two, legs, animal-like, dog legs, instrument, trumpet, red, blue, ribbons, cement, carved, behind, white, low, fence 0.862 table 4. semantic textual similarity scores for the same object across different dyads. object coder 1 coder 2 similarity kimono moldy, t shape, red, big, green, towards borders, in front of that indian place weird statue, torso, across the street from om store, metro center side 0.339 vader helmet black, shiny, marblish, circular on top, flat bottom, carved stripes, in front of clothes store dome-shaped, sculptury, concrete, black, seal, smooth 0.299 a phenomenon we did not analyze was direction-giving (“on lincoln street,” “turn left and walk one block”). traditional lab-based referential tasks do not generally include route-finding, and we wanted to draw a more direct comparison between the tangram task and the artwalk task. other researchers have also opted to focus on specific target descriptions when analyzing communication produced in a real-world wayfinding and target-identification task (e.g., brennan et al., 2013). number of turns, which was only used in analysis 3, was the combined number of turns taken by both director and matcher from the time the director started any conversation about the next target to the time when the matcher indicated that they had found the target. because some turns were partially about directions and partially about descriptions, instead of determining how much of a turn needed to be dedicated to description to count, we included partial turns and turns dedicated to direction-giving in this variable. friends and strangers in the wild 57 table 5. semantic textual similarity scores for different objects and across different dyads. object coder 1 coder 2 similarity coder 1 object: kimono coder 2 object: pi signs rock, moose, green, brown right side of street, blue background, rectangle, two parallel lines, white, bars, connect to other side, in between, strip 0.119 coder 1 object: mosaic coder 2 object: emo penguin in a parking lot, to the right, collage, tile, blue, green, orange, red penguin, little, boy, white, beanie, pointy, black, shading, helmet, large, buildings, facing, left 0.107 to summarize, the number of director descriptors was used in analyses 1 and 2, and the number of director descriptors and turns was used in analysis 3. 2.5 examples of entrainment and descriptor overlap in this section, we illustrate what is meant by entrainment in dialogue using examples from the corpus. we also describe how much descriptors overlapped between directors and matchers, followed by a discussion of the relationship between entrainment and overlap. in figure 2, the matcher uses the same words as the director, “a bird with a beanie,” as part of the process of grounding on descriptors that identify the art. 1 m: is it on the left the uh little mural? 2 d: yeah there is a picture of a bird with a beanie on 3 m: a bird with a beanie on you said? 4 d: yeah it is like a mopey gray sad bird 5 m: oh okay yeah i see it 6 d: yeah 7 m: yeah it is on the wall 8 d: okay cool figure 2. an example of a matcher copying a director in round 1. the figure 2 example was from round 1. in round 2, the director shortened the description to “the bird” saying, “uh so you’re gonna go back to the bird one um so walk up pacific,” which was accepted by the matcher with “alright i just got a picture of it.” this shows that by round 2, the director and matcher had entrained on “bird” for this art piece, and that the matcher did not need to repeat the word “bird” for entrainment to occur. the descriptor “a beanie” was also used by a different director-matcher pair, as seen in figure 3 lines 17 and 24. but unlike figure 2, the director and matcher in figure 3 made a lot of other conversational contributions that illustrate the breadth of information sifted through in this corpus to quantify entrainment, including directional information (e.g., line 12), spatial information (e.g., line 7), and conversational coordination (e.g., line 2, lines 31-32). in the second round of the figure 3 pair, the target was described exclusively with the director’s words, entraining on “sad penguin thing with the beanie,” as shown in figure 4, but without the matcher’s repeating any of the words in “sad penguin thing with the beanie.” liu, d’arcey, walker, and fox tree 58 1 d: you’re looking for like a emo penguin 2 m: sorry what? 3 d: it’s like a emo penguin or something i’m not really sure exactly what it is 4 m: okay [chuckle] 5 d: yeah 6 [silence 65s] 7 m: how far down did you say? i’m on by church right now 8 d: by by church street? 9 m: yeah 10 d: it’s it’s the next street 11 m: okay 12 d: yeah [silence 3s] uh when you get to locust you wanna cross the street um cross locust or a take a right on locust and then cross it 13 m: okay so go the next block but on locust 14 d: yeah 15 m: okay 16 [silence 7s] 17 d: and um you’re so what you’re looking for it’s like a it looks like it was spray painted um it’s like an outline i think it’s like a penguin wearing like beanie and it’s colored in white mostly 18 m: okay um 19 d: and it should be um it looks like it’s probably there should be two buildings on locust street from what i can see on that side of the street so it’s not the one it’s not the one on the cedar side it’s the one on the pacific side and it’s 20 m: alright 21 d: on the corner of the building that’s closest to the middle of the block 22 m: okay [silence 4s] so you said look like oh okay i think i see it 23 d: you see it? 24 m: wearing a beanie? 25 d: yeah 26 m: okay and okay so do you just want that one specific penguin looking like thing or do you want all of them? 27 d: uh it only has the one penguin it’s only it’s like cropped to where it’s only the one white penguin on mine so 28 m: okay so 29 d: so looks pretty sad 30 m: cool alright 31 d: got it? 32 m: yeah figure 3. an example of object entrainment surrounded by directional and conversational coordination in round 1. in this exchange, one of our coders generated the following list of descriptors for the director: emo, penguin, spray painted, outline, beanie, and white. note that penguin only occurs once in this list, as we asked coders to only record unique descriptors (only ones that hadn’t already occurred in the trial). 1 d: got that one? so now you gotta go back to the sad penguin which is um on locust street so i mean you can go to pacific or cedar and just head head that way 2 m: okay 3 d: away from pergolesi 4 m: alright wait which one am i getting now? 5 d: um the sad penguin thing with the beanie 6 m: oh okay okay okay yeah figure 4. round 2 of the figure 3 example showing drastic reduction in words used to identify the art. friends and strangers in the wild 59 both friends and strangers entrained across rounds 1 and 2. figure 5 illustrates the same target described during the first and second rounds of a friend pair and a stranger pair (asterisks indicate overlap). examples were chosen for brevity and clarity in round 1. they came from pairs where the matcher became reasonably confident in the identification of the target quickly. friends describing the spiky rock target in the first round 1 d: you’ll be pretty much walking pretty much towards the end of cooper street and what you’re looking for, um pretty much they’re spiky rocks, like, they’re rocks with little spikes on them 2 m: it’s a what? 3 d: they’re rocks, like, kinda like oval-shaped rocks with spikes on them, they’re, you you’ll, you’ll know what i mean 4 m: huh? no, you were saying it’s an oval-shaped rock 5 d: with spikes on them [yawn] 6 m: wait, it’s what? 7 d: spikes, like, splike thorn spikes, like spiky stuff like s 8 m: is it red? 9 d: one of the rocks is charcoal gray, one of the rocks is like a light brown, the other rock is *like* 10 m: *oh do* thedo they do they have have a little like slots on them? 11 d: yeah, like little spikes on them yeah 12 m: okay, i found them same friends describing the spiky rock target in the second round 14 d: okay, um last one, go back to cooper street for the spiky rocks 15 m: hey hey to theorto the rocks again? 16 d: yeah, the ones that we yeah. 17 m: okay, to the rocks and then it’s the last one, that one mural 18 d: yeah, i think that’d be the last one 19 [matcher and director joke and matcher says something unintelligible to a third party] 20 m: alright, i got it strangers describing the spiky rock target in the first round 21 d: okay, so you head down cooper 22 m: *uh huh* 23 d: *and* uh it’s gonna be on the left side erit’s gononce youat the end of cooper if *ya* 24 m: *uh huh* 25 d: on, the left of you, the left side, it’s at the corner, there should be like these three uh it looks like these three rocks one’sthe theit goes from like one’s small, one’s medium, and one’s one’s large 26 m: oh yeah yeah, with uh blue spikes and *yellow spikes*? 27 d: *yeah yeah yeah* those are the ones 28 m: perfect 29 d: okay same strangers describing the spiky rock target in the second round 30 m: okay, i got that picture, where to next? 31 d: next one is the uh three wathe three stones the blue one with the spikes 32 m: are we just revisiting all of them? 33 d: yeah, we’re *revisiting all of them* 34 m: *[laughing]* 35 d: so it’s on cooper street, you’re right next to it 36 m: alright, i was like wait a minute, *these objects* 37 d: *[laughing]* 38 m: look familiar 39 d: yeah 40 m: okay, well at least that one’s close 41 [task-irrelevant conversation as matcher takes photo] 42 m: anyway okay, i got that last picture, so now what? figure 5. a pair of friends and a pair of strangers describing the same art across rounds 1 and 2. liu, d’arcey, walker, and fox tree 60 all four of these participants used similar concepts to describe this public art, (1) rocks or stones (e.g., lines 1, 14, 25, and 31) and (2) having spikes (e.g., lines 1, 14, 26 and 31). a close examination of directors’ and matchers’ descriptors was conducted on half the data to assess the degree to which the descriptors overlapped. unsurprisingly given that they had the information about the objects to be identified, directors produced more descriptors than matchers: directors’ descriptors account for 64.63% of the total descriptors. we also asked two coders to assess how often directors’ and matchers’ descriptors overlapped using the two sets of rules that were used above to measure similarity, both using the strict and the flexible approaches. coders’ ratings of the number of director descriptors matched to matcher descriptors across 60 trials were highly correlated for both the strict set of rules, r(58) = .94, p < .001, and the flexible set of rules, r(58) = .79, p < .001. in the strict method of coding single word descriptors, both raters determined whether a director’s descriptor was also used by the matcher, such as “rocks” in figure 5. as before, when close or partial matches occurred, coders determined whether the meaning was similar enough to be treated as a match, based on whether the root of the word was the same. also as before, for multiple-word descriptors, coders determined whether content words were identical including that negation should not be counted as matching. coders compared single-word descriptors to multipleword descriptors in the same way that they compared multiple-word descriptors. strict coding resulted in 12.17% overlap averaged across coders and items. as before, strict descriptor matching is likely to under-represent the actual levels of entrainment due to its rigid rule set. for a deeper look at entrainment, coders made more thoughtful judgment calls on whether a descriptor’s meaning was similar enough to be considered entrainment. flexible coding resulted in 21.99% overlap average across coders and items. the strict and flexible overlap analyses show that people do in fact overlap in the words chosen to describe the target art, which illustrates one form of entrainment. entrainment does not require overlap in words, however. entrainment is also visible in figures 2 through 5 through the use of backchannels acknowledging agreement on a description. in figure 4 line 6, the matcher responds “oh okay okay okay yeah” to the descriptor “sad penguin” without repeating any descriptor words. in figure 5 lines 30 through 42, the matcher doesn’t even confirm that the “three stones” were found — the confirmation is implicit in the matcher’s clarification that the objects were being identified a second time coupled with the matcher’s final remark, “anyway okay, i got that last picture, so now what?” the entrainment occurs through the taking of the picture, not with the production of words like rock, stone, or spiky. 3 results we tested whether acquaintanceship status and extraversion influenced how quickly (defined by number of descriptors) pairs would entrain (analysis 1) and examined how efficiency differed in round 1 (analysis 2) and in round 2 (analysis 3). because there was no evidence of a difference in performance between list a (n = 23) and list b (n = 25; t(46) = -0.65, p = .52, 95% ci [-2.90, 1.49]), lists were collapsed in these analyses. descriptive statistics broken down by acquaintanceship can be found in table 6. we note here that that there was no evidence of a difference in director descriptor use before and after the replacement of two of three targets during our data collection period (due to the city changing its public art displays), t(46) = -0.73, p = .47, 95% ci [-3.05, 1.43] (because the third target was switched out very late during our data collection with only a few participant pairs discussing it, we did not break the data down to test whether the third object made a difference). overall, there was a slight negative skew (-0.21) in the director extraversion data, with more participants rating themselves as being highly extraverted than highly introverted. the scores were non-normally distributed (shapiro-wilk = 0.95, p = .03) but the distribution was still roughly bellshaped. friends and strangers in the wild 61 table 6. basic descriptive statistics on variables of interest broken down by acquaintanceship status. variable friend pairs (n = 24) stranger pairs (n = 24) round 1 director descriptors mean sd median 12.53 3.90 12.70 11.93 3.66 12.00 round 1 turns mean sd median 20.01 5.49 19.68 18.67 8.68 16.93 round 2 director descriptors mean sd median 4.34 1.91 4.20 3.99 1.64 3.68 round 2 turns mean sd median 8.32 3.59 7.00 7.05 2.75 7.00 director extraversion mean sd median 4.48 1.42 4.00 4.81 1.47 5.00 3.1 analysis 1 with analysis 1 we tested whether extraversion and acquaintanceship predicted the number of director descriptors used overall, as well as whether extraversion and acquaintanceship had differential effects in round 1 and round 2. preliminary analyses indicated no evidence of a difference between director extraversion between friends (m = 4.48, sd = 1.42) and strangers (m = 4.81, sd = 1.47) in our sample, t(46) = -0.80, p = .43, 95% ci [-1.17, 0.50]. this suggests that friend pairs who participated did not inherently differ from stranger pairs with respect to sociability or outgoingness. to test whether acquaintanceship status and extraversion influenced how quickly (defined by number of descriptors) pairs entrained, a hierarchical linear regression was performed using director descriptors as the dependent variable, round and acquaintanceship as binary categorical predictors, and director extraversion as a continuous predictor (centered; cohen, cohen, west, & aiken, 2003). two-way interaction terms between the three predictor variables were also entered. this model (model a) predicted about 69.62% of the variance in director descriptors, f(6,95) = 37.29, p < .0001, adj-r2 = .70, though it should be noted that the interaction terms introduced a fairly high level of collinearity. acquaintanceship (b = -1.42, p = .43) and director extraversion (b = 1.10, p = .10) were not significant predictors of the number of director descriptors. however, round was a significant predictor (b = -8.31.10, p < .001), indicating that directors did use shorter descriptions in round 2 than in round 1. this result replicates, in the wild, the oft-demonstrated laboratory finding that referring expressions are shortened during entrainment in referential communication tasks. there were two interactions. one interaction was between acquaintanceship and director extraversion (b = 1.14, p = .005). whereas directors in friend pairs used the same number of descriptors regardless of whether they were introverted or extraverted, the number of director liu, d’arcey, walker, and fox tree 62 descriptors for stranger pairs was sensitive to director extraversion. specifically, stranger directors who scored higher in extraversion used more descriptors than those who scored lower in extraversion. the second interaction was between round and director extraversion (b = -0.85, p = .04). director extraversion only made a difference to the number of descriptors produced in round 1 but not in round 2. this possibly suggests that round 1 and round 2 are fundamentally different when it comes to the influence of extraversion on performance. there was no interaction between round and acquaintanceship (b = 0.53, p = .64). see table 7. table 7. regression table for analysis 1: all director descriptors model a. **sig. at .01 level; ***sig. at .001 level. b se round -8.31*** 0.78 acquaintance -1.42 1.79 director (dir) extraversion 1.11 0.66 acquaintance * dir extraversion 1.14** 0.40 acquaintance * round 0.53 1.13 round * dir extraversion -0.85 0.40 adjusted r2 0.70 f for model 37.29*** 3.1.1 analysis 1 discussion in analysis 1, we found that neither extraversion nor acquaintanceship alone influenced the number of descriptors that the director used, though round did influence it. directors used fewer descriptors in round 2 than in round 1, which replicates the findings of lab-based tangram tasks. we also found that extraversion influenced directors from stranger pairs but not friend pairs. directors from stranger pairs who were more extraverted used more descriptors than the less extraverted directors. but there was no evidence that friend directors used a different number of descriptors depending on whether they were high or low on extraversion. this was unexpected because extraversion is generally associated with greater talkativeness (dewaele & furnham, 1999) and a greater volume of novel ideas (jung et al., 2012). we also looked at whether extraversion and acquaintanceship interacted with round. this was used to test whether round 1 and round 2 should be analyzed separately. because round interacted with at least one of the predictors (director extraversion), we examined the two rounds using separate regression models for the remainder of the analyses of this study. in analysis 2, we examined round 1 using director extraversion and acquaintanceship as predictors. in analysis 3, we explored the influence of extraversion and acquaintanceship on the number of descriptors used by friends and strangers in the wild 63 directors in round 2, as well as two predictors stemming from round 1: number of director descriptors in round 1, as well as the number of turns in round 1. 3.2 analysis 2 with analysis 2, we tested the influence of director extraversion and acquaintanceship on the number of director descriptors produced in round 1. the predictors used were director extraversion (b = 0.07, p = .89), acquaintanceship (b = -0.89, p = .39), and their interaction (b = 1.50, p = .04). this model, model b, accounted for about 14% of the variance in director descriptors, f(3,44) = 3.46, p = .02, adj-r2 = 0.14. see table 8. table 8. regression table for analysis 2: round 1 director descriptors model b. *sig at .05 level. b se acquaintance -0.89 1.02 director (dir) extraversion 0.07 0.51 acquaintance * dir extraversion 1.50* 0.71 adjusted r2 .14 f for model 3.46* we used simple slopes analysis (cohen et al., 2003; preacher, curran, & bauer, 2003) to examine whether acquaintanceship moderated the extent to which extraversion influenced the number of director descriptors used in round 1. simple slopes analysis probes an interaction by examining the regression of the dependent variable (director descriptors) on a predictor variable (extraversion) at different values of a moderator variable (acquaintanceship), which in this case is dichotomous. the significance test assesses whether the slope of the regression line for either friend or stranger pairs differs from zero: that is to say, whether a decrease or increase in director descriptors is associated with higher or lower director extraversion. in order to graph the interaction, high extraversion was calculated using one standard deviation above the mean of our sample while low extraversion was calculated as one standard deviation below. this was both parsimonious and intuitive, given that there was no evidence that director extraversion differed between the two acquaintanceship conditions and given that our sample’s extraversion is in line with typical ten item personality inventory extraversion scores (gosling et al., 2003). one sd above the mean was a score of 6.08 and one sd below the mean was 3.21 (on a scale of 1-7, where the midpoint is 4). there was no evidence that friend pair directors differed in the number of director descriptors produced based on how introverted or extraverted the director was, but there was evidence that stranger pair directors differed. the less extraverted the stranger director, the fewer descriptors they used. the stranger slope deviated from zero, stranger slope = 1.57, t(46) = 3.17, p = .003, 99% ci [0.24, 2.90], but the friend slope did not, friend slope = 0.07, t(46) = 0.13, p = .89. see figure 6. liu, d’arcey, walker, and fox tree 64 figure 6. simple slopes graph for friend and stranger pairs for round 1 director descriptors on extraversion. 3.2.1 analysis 2 discussion in analysis 2, we found that there was no direct effect of being previously acquainted with a partner on performance. we also found no simple relationship between director extraversion and the number of descriptors used. instead, we found that acquaintanceship moderated the influence of extraversion, such that extraversion only had an effect when participants were strangers. there was no evidence that friend directors differed in the number of descriptors produced depending on how introverted or extraverted they were, but there was evidence that stranger directors did. the less extraverted the stranger director, the fewer descriptors they used. though we found that an interaction between acquaintanceship and extraversion accounted for some of the variance in descriptors in round 1, we predicted that this pattern would not hold for round 2. once participants believed they had established a conceptual pact for specific referents, the task became a more straightforward test of whether the matcher was able to match a referring expression to one of the targets previously found in round 1. we predicted that round 2 performance would be related to how participants did in round 1: the less confusing and more straightforward round 1 performance, the fewer number of descriptors the director would need to use in round 2. 3.3 analysis 3 with analysis 3, we explored the influence of director extraversion, acquaintanceship, the number of director descriptors in round 1, and the number of turns in round 1 on the number of director descriptors produced in round 2. this was an exploratory analysis because we did not have specific hypotheses. we wanted to explore whether the speed at which people were able to identify art – fewer turns in round 1 – influenced the effort required to identify art in round 2 – fewer descriptors in round 2. we ran the same model that we used for analysis 2 (which was model b in analysis 3: director extraversion, acquaintanceship, and director extraversion * acquaintanceship interaction) on the number of director descriptors used in round 2. in model c we see no main effects or interactions across these variables. model d went a step further and added round 1 director descriptors, as it friends and strangers in the wild 65 was reasonable to assume that the number of descriptors used in round 1 would influence the number of descriptors used in round 2. model d also added two two-way interactions: round 1 director descriptors * director extraversion and round 1 director descriptors * acquaintanceship. all predictors were centered, except for the binary acquaintanceship predictor. these additions resulted in a regression model that did not reach significance, f(6,47) = 1.84, p = .12. as a result of the results of model d, both director extraversion and round 1 director descriptors were dropped from the model. model e substituted round 1 turns in place of round 1 director descriptors as the measure of efficient communication in round 1. round 1 turns was logarithmically transformed, as it was moderately skewed. model e, which includes acquaintanceship, round 1 turns (log), and an interaction between acquaintanceship and round 1 turns, is able to account for about 10.4% of the variance in round 2 director descriptors, f(3,44) = 2.83, p = .05, adj-r2 = 0.10. figure 7. simple slopes graph for friend and stranger pairs for round 2 director descriptors on round 1 turns. overall, there was no evidence of a difference between friends (m = 4.34, sd = 1.91) and strangers (m = 3.99, sd = 1.64) in the number of director descriptors used in round 2, t(46) = -0.67, p = .51, 95% ci [-1.38, 0.69]. a simple slopes analysis of the interaction between round 1 turns and acquaintanceship (figure 7) was conducted. we defined one standard deviation above the average number of turns as slow and one standard deviation below the average number of turns as fast. this analysis indicated that the stranger slope did not deviate from zero, stranger slope = 2.19, t(46) = -1.05, p = .30, but the friend slope did, friend slope = 7.53, t(46) = 2.61, p = .01, 95% ci [1.74, 13.32]. this indicates that the effect of the number of turns in round 1 on the number of liu, d’arcey, walker, and fox tree 66 descriptors used in round 2 was moderated by acquaintanceship. for friend pairs, the fewer turns used to reach agreement on a target in round 1, the fewer director descriptors were needed in round 2. for the strangers, the number of descriptors used in round 2 did not change depending on whether they took fewer or more turns in round 1. see table 9. table 9. regression table for analysis 3: round 2 director descriptors. *sig at .05 level; **sig. at .01 level; ***sig at .001 level. model c model d model e b se b se b se acquaintance -0.30 0.52 -0.74 0.53 -0.25 0.49 director (dir) extraversion 0.00 0.10 -0.47 0.25 acquaintance * dir extraversion 0.11 0.14 1.12* 0.43 r1 dir descriptors 0.04 0.09 r1 dir descriptors * dir extraversion -0.09 0.16 r1 dir descriptors * acquaintance 0.11* 0.05 round 1 turns 7.53** 2.88 round 1 turns * acquaintance -9.72*** 3.55 adjusted r2 -0.03 0.097 0.104 f for model 0.49 1.84 2.83* 3.3.1 analysis 3 discussion in analysis 3 we found no evidence of a difference in how many round 2 descriptors strangers used, based on whether they were fast or slow to agree upon a referring expression for a referent in round 1. however, speed at agreeing did make a difference for friend pairs: directors in friend pairs who had taken more turns to negotiate a referring expression or agree on a referent in round 1 tended to use more descriptors in round 2. said another way, stranger directors did not adjust their strategy based on how the pair did in round 1, but friend directors did. that being said, the results for analysis 3 should be interpreted with caution as the replacement of round 1 director descriptors friends and strangers in the wild 67 with round 1 turns was done to explore whether any basic measure of how participants performed in round 1 influenced the number of descriptors used in round 2. 4 general discussion in line with decades of research using laboratory-based referential communication tasks such as the tangram task, we show that the shortening of referring expressions also happens in more naturalistic conversational settings: directors used fewer descriptors when referring to targets over two rounds of a remote skype-to-cell-phone referential communication task while their matcher partners had to navigate real-life obstacles on busy streets. taking no further variables into account, there was no evidence of a difference between the performance of friends and strangers, nor was there evidence that the number of descriptors used changed among people of varying extraversion levels. alone, friends and strangers were equally efficient, as were introverts and extraverts. however, acquaintanceship status and extraversion together did affect how quickly pairs could mutually agree upon a referring expression. the more introverted the stranger director, the fewer descriptors were used; that is, for strangers, introversion led to increased efficiency of communication. friends, however, were not influenced by extraversion. in other words, extraverts and introverts were only distinguishable in stranger pairs. for both friends and strangers, the influence of extraversion disappeared by the second round. results of exploratory analyses potentially suggest that existing friendship affected how efficiently pairs performed in the task, as was observed in the interaction between acquaintanceship and the number of turns used in round 1. this time, the effect was observed with friends: directors in friend pairs were sensitive to whether there were a greater or fewer number of turns before agreeing upon the target in the first round and adjusted their behavior accordingly. the more turns they took in the first round, the more descriptors they used in the second. this trend was not significant for strangers. whereas friends seemed to adapt to their partners and slow down to accommodate potential difficulties in communication, strangers did not. as with any task completed in relatively uncontrolled circumstances, there are a number of variables that we did not test in our design that may have affected results. for example, the mobile phone-to-skype communication in our study may have been constrained by the quality of the audio, the quality of the cellular connection, or even the participant’s comfort with using a phone that was not their own. these constraints may have caused participants to repeat themselves and each other more than usual, to speak more loudly to overcome ambient noise from the environment, or to express information differently because they feared getting disconnected. personality factors may also manifest differently with phone-to-skype communication versus other communication. for example, extraverts and introverts have different opinions on what constitutes polite mobile phone behavior in public spaces (love & kewley, 2005). consequently, introverted matchers might have designed utterances differently, knowing that bystanders might overhear their conversation. differences in task-oriented conversation between friends and strangers are not as evident as people might assume. we provide some evidence that friends and strangers can differ in conversations that are focused on a specific collaborative goal, though their verbal behavior is moderated by extraversion and, potentially, by the conversations they have had in the recent past. though the in-the-wild method introduces literal and statistical noise, putting people into more naturalistic contexts and examining discourse between interlocutors who have various levels of acquaintanceship can reveal differences that are hidden or discouraged in the laboratory. liu, d’arcey, walker, and fox tree 68 acknowledgements this research was supported by nsf grant iis # 1044693 from the robust intelligence program. we thank our many research assistants who aided in data collection and coding, with special thanks to haley biesemeier, alea casanova, ericka elphick, andrea estrada, steven ethington, steven finet, jennifer fong, beth hart, bronwyn hassall, erin hiscock-wagner, megan kostecka, emilie kovalik, tyler pilgrim, natalie serourian, sean tang, jason tharp, yana ulitsky, elizabeth williams, and evelyn yap. we thank natalia blackwell and lena reed for contributions to this project. we thank barbara di eugenio and four anonymous reviewers for comments on an earlier version of this manuscript. correspondence can be addressed to jean e. fox tree (foxtree@ucsc.edu) in the psychology department, university of california santa cruz, santa cruz, ca, 95064. references albright, l., cohen, a. i., malloy, t. e., christ, t., & bromgard, g. (2004). judgments of communicative intent in conversation. journal of experimental social psychology, 40, 290302. anderson, a. h., bader, m., bard, e. g., boyle, e., doherty, g., garrod, s., isard, s., kowtko, j. mcallister, j., miller, j., sotillo, c., thompson, h. s. & weinert, r. (1991). the hcrc map task corpus. language and speech, 34(4), 351-366. andersson, j., & rönnberg, j. (1995). recall suffers from collaboration: joint recall effects of friendship and task complexity. applied cognitive psychology, 9(3), 199-211. bard, e., aylett, m., & lickley, r. (2002). towards a psycholinguistics of dialogue: defining reaction time and error rate in a dialogue corpus. in proceedings of the sixth workshop on the semantics and pragmatics of dialogue (edilog 2002). beauchet, o., dubost, v., herrmann, f. r., & kressig, r. w. (2005). stride-to-stride variability while backward counting among healthy young adults. journal of neuroengineering and rehabilitation, 2(1), 26. beukeboom, c. j., tanis, m., & vermeulen, i. e. (2013). the language of extraversion: extraverted people talk more abstractly, introverts are more concrete. journal of language and social psychology, 32(2), 191-201. blackwell, n., & fox tree, j. e. (2012). social factors affect quotative choice. journal of pragmatics, 44(10), 1150-1162. bortfeld, h., leon, s. e., bloom, j. e., schober, m. f., & brennan, s. (2001). disfluency rates in spontaneous speech: effects of age, relationship, topic, role, and gender. language and speech, 44(2), 123-149. branigan, h., lickley, r., & mckelvie, d. (1999). non-linguistic influences on rates of disfluency in spontaneous speech. in proceedings of the 14th international conference of phonetic sciences. brennan, s. e., & clark, h. h. (1996). conceptual pacts and lexical choice in conversation. journal of experimental psychology: learning, memory & cognition, 22(6), 482-493. brennan, s. e., schuhmann, k. s., & batres, k. m. (2013). entrainment on the move and in the lab: the walking around corpus. in proceedings of the 35th annual conference of the cognitive science society. campbell, a., & rushton, j. p. (1978). bodily communication and personality. british journal of social and clinical psychology, 17(1), 31-36. carment, d., miles, c., & cervin, v. (1965). persuasiveness and persuasibility as related to intelligence and extraversion. british journal of social and clinical psychology, 4(1), 1-7. clark, h. h., & wilkes-gibbs, d. (1986). referring as a collaborative process. cognition, 22, 139. friends and strangers in the wild 69 clark, h. h. (1996). using language. cambridge university press. cohen, j., cohen, p., west, s., & aiken, l. (2003). applied multiple regression/correlation analysis for the social sciences (3rd edition). lawrence erlbaum associates. cuperman, r., & ickes, w. (2009). big five predictors of behavior and perceptions in initial dyadic interactions: personality similarity helps extraverts and introverts, but hurts “disagreeables”. journal of personality and social psychology, 97(4), 667-684. dewaele, j. m., & furnham, a. (1999). extraversion: the unloved variable in applied linguistic research. language learning, 49(3), 509-544. dunne, m., & ng, s. h. (1994). simultaneous speech in small group conversation: all-togethernow and one-at-a-time? journal of language and social psychology, 13(1), 45-71. feldstein, s., & sloan, b. (1984). actual and stereotyped speech tempos of extraverts and introverts. journal of personality, 52(2), 188-204. fleming, j. h., darley, j. m., hilton, j. l., & kojetin, b. a. (1990). multiple audience problem: a strategic communication perspective on social perception. journal of personality and social psychology, 58(4), 593-609. fox tree, j. e. (1999). listening in on monologues and dialogues. discourse processes, 27(1), 3553. fox tree, j. e. (2007). folk notions of um and uh, you know, and like. text & talk, 27(3), 297-314. fox tree, j. e., & clark, n. b. (2013). communicative effectiveness of written versus spoken feedback. discourse processes, 50(5), 339-359. fox tree, j. e., mayer, s. a., & betts, t. e. (2011). grounding in instant messaging. journal of educational computing research, 45(4), 455-475. froming, w. j., & carver, c. s. (1981). divergent influences of private and public selfconsciousness in a compliance paradigm. journal of research in personality, 15(2), 159-171. fussell, s. r., & krauss, r. m. (1989). understanding friends and strangers: the effects of audience design on message comprehension. european journal of social psychology, 19(6), 509-525. gifford, r., & hine, d. w. (1994). the role of verbal behavior in the encoding and decoding of interpersonal dispositions. journal of research in personality, 28(2), 115-132. gill, a. j., & oberlander, j. (2003). perception of e-mail personality at zero-acquaintance: extraversion takes care of itself; neuroticism is a worry. in proceedings of the 25th annual conference of the cognitive science society. gosling, s. d., rentfrow, p. j., & swann, w. b. (2003). a very brief measure of the big-five personality domains. journal of research in personality, 37(6), 504-528. gould, o. n., osborn, c., krein, h., & mortenson, m. (2002). collaborative recall in married and unacquainted dyads. international journal of behavioral development, 26(1), 36-44. guydish, a., d’arcey, j. t., & fox tree, j. e. (2020). reciprocity in conversation. language and speech. advance online publication. doi: 10.1177/0023830920972742 haxby, j. v., parasuraman, r., lalonde, f., & abboud, h. (1993). superlab: general-purpose macintosh software for human experimental psychology and psychological testing. behavior research methods, instruments, & computers, 25(3), 400-405. heylighen, f., & dewaele, j.-m. (1999). formality of language: definition, measurement and behavioral determinants. interner bericht, center “leo apostel”, vrije universiteit brüssel, 4. hill, a., bohil, c., lewis, j., & neider, m. (2013). prefrontal cortex activity during walking while multitasking: an fnir study. in proceedings of the human factors and ergonomics society annual meeting. hope, l., ost, j., gabbert, f., healey, s., & lenton, e. (2008). “with a little help from my friends…”: the role of co-witness relationship in susceptibility to misinformation. acta psychologica, 127(2), 476-484. hornstein, g. a. (1985). intimacy in conversational style as a function of the degree of closeness between members of a dyad. journal of personality and social psychology, 49(3), 671-681. liu, d’arcey, walker, and fox tree 70 horton, w. s., & spieler, d. h. (2007). age-related differences in communication and audience design. psychology and aging, 22(2), 281-290. isaacs, e. a., & clark, h. h. (1987). references in conversations between experts and novices. journal of experimental psychology: general, 116(1), 26-37. jucker, a. h., & smith, s. w. (1998). and people just you know like “wow”: discourse markers as negotiating strategies. in andreas h. jucker & yael ziv (eds.), discourse markers: description and theory, pp. 171–201. john benjamins. jung, j., lee, y., & karsten, r. (2012). the moderating effect of extraversion–introversion differences on group idea generation performance. small group research, 43(1), 30–49. kemper, s., othick, m., warren, j., gubarchuk, j., & gerhing, h. (1996). facilitating older adults’ performance on a referential communication task through speech accommodations. aging, neuropsychology, and cognition, 3(1), 37-55. kenny, d. a., & acitelli, l. k. (2001). accuracy and bias in the perception of the partner in a close relationship. journal of personality and social psychology; journal of personality and social psychology, 80(3), 439-448. levesque, m. j., & kenny, d. a. (1993). accuracy of behavioral predictions at zero acquaintance: a social relations analysis. journal of personality and social psychology, 65(6), 1178-1187. li, b., zhou, h., he, j., wang, m., yang, y., & li, l. (2020, november). on the sentence embeddings from bert for semantic textual similarity. in proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), pp. 9119-9130. liu, k., fox tree, j. e., & walker, l. (2016). coordinating communication in the wild: the artwalk dialogue corpus of pedestrian navigation and mobile referential communication. in proceedings of the international conference on language resources and evaluation, portorož, slovenia, pp. 3159-3166. liu, y., ott, m., goyal, n., du, j., joshi, m., chen, d., levy, o., lewis, m., zettlemoyer, l., & stoyanov, v. (2019). roberta: a robustly optimized bert pretraining approach. arxiv preprint arxiv:1907.11692 love, s., & kewley, j. (2005). does personality affect peoples’ attitude towards mobile phone use in public places? mobile communications (pp. 273-284). springer. love, s., & perry, m. (2004). dealing with mobile conversations in public places: some implications for the design of socially intrusive technologies. in chi’04 extended abstracts on human factors in computing systems. neider, m. b., mccarley, j. s., crowell, j. a., kaczmarski, h., & kramer, a. f. (2010). pedestrians, vehicles, and cell phones. accident analysis & prevention, 42(2), 589-594. nelson, p. a., thorne, a., & shapiro, l. a. (2011). i’m outgoing and she’s reserved: the reciprocal dynamics of personality in close friendships in young adulthood. journal of personality, 79(5), 1113-1148. oberlander, j., & gill, a. j. (2006). language with character: a stratified corpus comparison of individual differences in e-mail communication. discourse processes, 42(3), 239-270. oulasvirta, a., tamminen, s., roto, v., & kuorelahti, j. (2005). interaction in 4-second bursts: the fragmented nature of attentional resources in mobile hci. in proceedings of the sigchi conference on human factors in computing systems. peeters, m. a., van tuijl, h. f., rutte, c. g., & reymen, i. m. (2006). personality and team performance: a meta­analysis. european journal of personality, 20(5), 377-396. pennebaker, j. w., & king, l. a. (1999). linguistic styles: language use as an individual difference. journal of personality and social psychology, 77(6), 1296-1312. planalp, s., & benson, a. (1992). friends’ and acquaintances’ conversations i: perceived differences. journal of social and personal relationships, 9, 483-506. friends and strangers in the wild 71 preacher, k. j., curran, p. j., & bauer, d. j. (2003). simple intercepts, simple slopes, and regions of significance in lca 2-way interactions. retrieved from http://www.quantpsy.org/interact/lca2.htm rawlings, d., & carnie, d. (1989). the interaction of epq extraversion with wais subtest performance under timed and untimed conditions. personality and individual differences, 10(4), 453-458. reimers, n. & gurevych, i. (2019). sentence-bert: sentence embeddings using siamese bertnetworks. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), pp. 3973-3983. savitsky, k., keysar, b., epley, n., carter, t., & swanson, a. (2011). the closenesscommunication bias: increased egocentrism among friends versus strangers. journal of experimental social psychology, 47(1), 269-273. scherer, k. r. (1978). personality inference from voice quality: the loud voice of extroversion. european journal of social psychology, 8(4), 467-487. schober, m. f., & clark, h. h. (1989). understanding by addressees and overhearers. cognitive psychology, 21(2), 211-232. smoski, m., & bachoroski, j.-a. (2003). antiphonal laughter between friends and strangers. cognition & emotion, 17(2), 327-340. speed, l. j., wnuk, e., & majid, a. (2016). studying psycholinguistics out of the lab. research methods in psycholinguistics and the neurobiology of language. wiley-blackwell. stinson, l., & ickes, w. (1992). empathic accuracy in the interactions of male friends versus male strangers. journal of personality and social psychology, 62(5), 787-797. thorne, a. (1987). the press of personality: a study of conversations between introverts and extraverts. journal of personality and social psychology, 53(4), 718726. tolins, j., zeamer, c., & fox tree, j. e. (2018). overhearing dialogues and monologues: how does entrainment lead to more comprehensible referring expressions? discourse processes, 55(7), 545-565. tolins, j., liu, k., neff, m., walker, m., & fox tree, j. e. (2016). a verbal and gestural corpus of story retellings to an expressive embodied virtual character. in proceedings of the international conference on language resources and evaluation, portorož, slovenia, pp. 3461-3468. truong, k. p., & heylen, d. (2012). measuring prosodic alignment in cooperative task-based conversations. in thirteenth annual conference of the international speech communication association. truong, k. p., & trouvain, j. (2012). laughter annotations in conversational speech corporapossibilities and limitations for phonetic analysis. in proceedings of the 4th international workshop on corpora for research on emotion sentiment and social signals, 20-24. uziel, l. (2007). individual differences in the social facilitation effect: a review and meta-analysis. journal of research in personality, 41(3), 579-601. wang, h.-c., fussell, s. f., & setlock, l. d. (2009). cultural difference and adaptation of communication styles in computer-mediated group brainstorming. in proceedings of the sigchi conference on human factors in computing systems. wu, s., & keysar, b. (2007). the effect of culture on perspective taking. psychological science, 18(7), 600-606. xu, y. (2010). in defense of lab speech. journal of phonetics, 38(3), 329-336. yogev­seligmann, g., hausdorff, j. m., & giladi, n. (2008). the role of executive function and attention in gait. movement disorders, 23(3), 329-342. liu, d’arcey, walker, and fox tree 72 appendix a descriptor matching instructions the goal of this coding is to see whether participants used the same terms to describe objects to each other (lexical entrainment). for that reason, categorize descriptors as non-matches if unclear. use the following guides to match first. dashed words count as multiple words. note that most of the descriptors in this set will be single words. descriptors are always separated by commas. see the end of this document for examples. strict way single words. for each of the director descriptors, see if it exists in the matcher descriptors. if it is an exact match, count it as a match. in case of a close or partial match, determine whether the meaning is similar enough to be treated as a match: (1) is the root of the word the same (e.g., yellow and yellowish are the same root word, flying and flies are the same root word)? if so, consider it a match. (2) otherwise, it’s not a match. multiple words. (1) are all of the content words identical? (e.g., ‘surrounded by plants’ and ‘plants are surrounding it’ have two identical content words, so it is a match. (2) if the content words are identical but one is negated (e.g., ‘holding hands’ and ‘not holding hands’) it is not a match. (3) otherwise, it’s not a match. mix of single and multiple words. do the content words all match? (e.g., “the baseball” and “baseball” match, but “white baseball” and “baseball” do not match). flexible way single words. (1) is the word almost identical in meaning (e.g., tan, khaki, beige)? if so, consider it a match. (2) otherwise, it’s not a match. multiple words. (1) are most of the content words at least similar? (e.g., ‘it’s kind of round and a dark blue purple color’ and ‘it’s circular and purpleish’) if you can get the same general idea, consider it a match. (2) if one description is more specific than the other/the verbiage is different but related to the same object (e.g., ‘it has a baseball’ versus ‘it’s holding a baseball’), it can be a match. (3) otherwise, it’s not a match. example director (number of descriptors = 17) painting, on wall, or on sidewalk, jaggedy looking, turquoise, two, little, purple, eggshaped figures, little, black legs, black arms, sticking out, no head, tiny arms and legs, at the bottom, brown underneath the turquoise matcher (number of descriptors = 11) mural, like a rainbow, boat, turquoise, green strip, reddish brown strip, bar, with bunch of hieroglyphs, little figures, strip of blue, and green matches the strict way turquoise matches the flexible way turquoise/turquoise; mural/painting; egg-shaped figures/little figures dialogue and discourse 4(2) (2013) 215-248 doi: 10.5087/dad.2013.210 focus triggers and focus types from a corpus perspective arndt riester arndt.riester@ims.uni-stuttgart.de institut für maschinelle sprachverarbeitung universität stuttgart pfaffenwaldring 5 b d-70569 stuttgart stefan baumann stefan.baumann@uni-koeln.de institut für linguistik – phonetik universität zu köln herbert-lewin-str. 6 d-50931 köln editors: stefanie dipper, heike zinsmeister, bonnie webber abstract the article discusses several issues relevant for the annotation of written and spoken corpus data with information structure. we discuss ways to identify focus top-down (via questions under discussion) or bottom-up (starting from pitch accents). we introduce a two-dimensional labelling scheme for information status and propose a way to distinguish between contrastive and noncontrastive information. moreover, we take side in a current debate, claiming that focus is triggered by two sources: newness and elicited alternatives (contrast). this may lead to a high number of semantic-pragmatic foci in a single sentence. in each prosodic phrase there can be one primary focus (marked by a nuclear pitch accent) and several secondary foci (marked by weaker prosodic prominence). second occurrence focus is one instance of secondary focus. keywords: alternatives, contrast, corpus annotation, focus, givenness, information status, information structure, prosody, question under discussion, secondary prominence 1. focus in english and german: basic prosodic assumptions the purpose of this article1 is to discuss the linguistic notion of focus – which has been intensively investigated in the theoretical literature2 – in the light of corpus data. for this, it is necessary that we take a stand on a number of issues about which the current theoretical debate is undecided, for instance, whether one or several types of focussing should be assumed and how to delineate contrastive from non-contrastive focus. we are convinced that precision with regard to conceptual 1. this is a completely revised and extended version of an article which appeared in 2011 under the title information structure annotation and secondary accents in the volume beyond semantics: corpus-based investigations of pragmatic and discourse phenomena in the series bochumer linguistische arbeitsberichte, 3:111-127, ed. by stefanie dipper and heike zinsmeister. 2. a few references: halliday (1967); chomsky (1971); jackendoff (1972); rooth (1992); krifka (1992); selkirk (1995); roberts (1996); hajičová et al. (1998); é. kiss (1998); schwarzschild (1999); büring (2003); krifka (2008); beaver and clark (2008) c©2013 arndt riester and stefan baumann submitted 04/2012; accepted 02/2013; published online 08/2013 riester and baumann issues is an indispensable prerequisite for corpus annotation of information structure. we will furthermore argue in favour of a distinction between what we call primary focus and communicatively less important kinds of secondary focus. the first assumption which we shall make in this paper is that in english or german,3 and presumably in other languages, there can be several foci in the same prosodic phrase but only one of them can be marked by a nuclear pitch accent. we call this one the primary focus. definition 1 (primary focus) a primary focus is a focus constituent which is marked by means of a nuclear pitch accent. note that definition 1 neither says what a focus is in general, nor why and when it may receive a nuclear pitch accent, nor whether it has to be marked at all. this requires some clarification: in line with most current theoretical treatments, we take focus to be a semantic-pragmatically – and not prosodically or morpho-syntactically – defined notion. to investigate what the precise semanticpragmatic factors are which lead to focussing is one of the main subjects of the present article, and we will soon turn to this issue (section 2). for the beginning, let us settle for an intuitive focus notion which says that focus indicates new or important parts of an utterance. it seems as if the literature contains many examples in which there are two primary foci in the same clause. examples are given in (1) to (4), with relevant pitch accents marked by capital letters. (1) even johnf1 drank only waterf2. (krifka, 1992) (2) a: did carl sue the company, or did the company sue carl? (büring, 2003) b: carlf1 sued the companyf2. (3) a: who ate what? (roberts, 1996; büring, 2003) b: fredf1/ct ate the beansf2. (4) [in myf1 opinion]t , johnf2 stole the cookies. (krifka, 2008) the phenomenon in (1) has been called multiple focus. it is characterised by the presence of two focus-sensitive particles, even and only, within the same clause. the focused constituents are indexed by f1 and f2. in (2b), the two foci form a complex focus. they highlight a choice from among two pairs of alternatives provided by question (2a). in addition to that, the (pairs of) foci may also be called contrastive foci, due to the presence of overt alternatives. (3b) is a protoypical example involving a contrastive topic (ct), which is assumed to share basic properties of focus, in particular, the ability to highlight the availability of alternatives. in opposition to (2b), it additionally signals that the question is not answered completely, and that people, rather than dishes, are the “sortal key” (büring, 2003: 530) along which (3a) is worked off. in using the contrastive topic, the speaker of (3b) announces to continue and tell us what other persons ate. the first focus in (4) is used within a frame-setting topic (t), which likewise divides the interpretational space into alternative partitions. ignoring the fact that the pitch accents occurring on the two respective focal elements in (1)-(4) might be of a different type, e.g. rising vs. falling, there is no clear observable difference with respect to their prosodic strength or prominence.4 this holds true for all four examples. 3. throughout the paper, we shall make the assumption that the prosodic marking of information structure is very similar in english and german. this seems by and large justified with regard to pitch accent placement, but perhaps less so with regard to some details of pitch accent types. 4. note that krifka (2008) does claim that the first accent in some cases of multiple focus is stronger than the second one, and that this distinguishes multiple focus from complex focus. if this can be shown to be true in general, then (1) is not an example with two primary foci. 216 focus triggers and focus types all we have said so far is compatible with the requirement that there can be only one primary focus per prosodic phrase, since we may assume that, in all of the above examples, the two respective foci are confined to their own intermediate phrase (ip) – or even intonation phrase (ip) – defined by beckman et al. (2005) as the domain for a nuclear accent. in other words, the prosodic structure of all sentences in (1) to (4) will look like (5). (5) {{ (pn1 . . . pnn) n (pn) -}ip { (pn1 . . . pnm) n (pn) -}ip %}ip here, (pn) stands for an optional prenuclear accent, several of which may occur, n is the nuclear pitch accent, and (pn) indicates an optional postnuclear prominence. the latter is not supposed to carry the rank of a pitch accent and is signaled by means of increased duration and intensity but only very little pitch movement. it is sometimes called phrase accent (grice et al., 2000), or postlexical stress (beckman, 1986). furthermore, ‘%’ indicates an intonation phrase break, and ‘-’ stands for an intermediate phrase break. however, it is possible and, in fact, widespread to have several (semantic-pragmatic) foci in the same prosodic phrase. one particular case in point is second occurrence focus (sof), which has received a considerable amount of attention in recent years (rooth, 1996; partee, 1999; bartels, 2004; büring, 2008, ms.; rooth, 2010; beaver and velleman, 2011), though not explicitly from the perspective of prosodic phrasing. a well-known example from partee is shown in (6), in which we would assume that the first half of (6b) (until vegetables) can be realised as a single prosodic phrase. (6) a. everyone knew that mary only eats vegetablesf1. b. if { even paulf2 n knew that mary only eats vegetablessof , pn -} then he should have suggested a different restaurant. several experiments on both english (rooth, 1996; beaver et al., 2007) and german (féry and ishihara, 2009; baumann et al., 2010) have revealed that (semantically) focussed5 but given expressions, like vegetables in (6b), can be realised with some sort of postnuclear prominence (pn), thus differing from non-focal, given expressions, which are claimed to be completely deaccented. we will return to the issue of second occurrence focus in section 5. at this point, we merely note that sof is one instance from the class of secondary foci defined below. definition 2 (secondary focus) a secondary focus is a focus constituent which is not marked by means of a nuclear pitch accent but by some preoder postnuclear prominence. obviously, our definitions of primary and secondary focus are not fully pragmatic in nature since they make reference to prosodic concepts. in the long run, one would like to replace definitions 1 and 2 by purely pragmatic definitions which predict when exactly a focus receives a nuclear pitch accent. we think that, seen from the perspective of corpus analysis, it is still too early for that.6 we would like to draw the reader’s attention to a related phenomenon which has received much less attention than the role of postnuclear prominences in marking second occurrence focus: the 5. in both (6a) and (6b), vegetables is associated with the focus adverb only. 6. the question of predicting which focus receives more prosodic prominence is addressed in selkirk (2008); büring (2008, ms.); rooth (2010) or beaver and velleman (2011). büring gives a purely pragmatic definition of primary focus in terms of the sizes of focus domains. of two foci, the one “whose domain contains the domain of the other” is the primary one. büring predicts that this focus will then receive the nuclear pitch accent. however, a definition like that presupposes a solution to the problem of identifying focus domains. 217 riester and baumann information structural contribution of prenuclear accents. while it is easy to ignore these accents when discussing “the” focus of constructed examples, they represent an important and ubiquitous element of the prosody of almost every spoken utterance. consider the phrase shown in figure 1, taken from the dirndl corpus of german radio news (eckart et al., 2012), which is annotated for pitch accents and prosodic boundaries following gtobi(s) (mayer, 1995). the entire sentence is given in (7).7 figure 1: praat (boersma and weenink, 1996) screenshot; dirndl s1222, 26-03-2007, 06:00, 7’12”: in the former slave trading post elmina (7) in {{ ghana l*h -} fand { ein festakt l*h -} im { ehemaligen h* sklaven-handelsposten l*h elmina h*l statt. %}} in ghana, a ceremonial act took place in the former slave trading post elmina. in the section shown in figure 1 there are two prenuclear accents, h* on the adjective ehemaligen (‘former’) and l*h on sklaven-handelsposten (‘slave trading post’). it is accents like these which are usually ignored in formal discussions of focus. if they are mentioned at all, they are often assumed to be optional or, at least, not meaning-related.8 we would like to raise doubts about the optionality of prenuclear accents. in (7), it doesn’t seem possible to omit any of the prenuclear accents. it is unclear, so far, whether this is because they play a role in indicating information structure – perhaps marking a focus of their own – or for purely rhythmical reasons. it seems that in order to settle this issue more theoretical background is necessary. analysing corpus data with regard to their focal properties raises a number of problems. we may by now have sophisticated theories concerning the pragmatics of focus, as well as some good experimental evidence about its prosodic marking. nevertheless, all that knowledge still seems quite 7. nuclear pitch accents are indicated in boldface, prenuclear accents in small capitals. 8. büring (2007) calls these accents “ornamental”. 218 focus triggers and focus types insufficient to analyse even a relatively simple and by no means untypical example like (7). in the next section, we will try to sketch the problem from two different perspectives (top-down, i.e. from the perspective of questions under discussion, and bottom-up, i.e. starting out from pitch accents). in section 3, we present our reflex annotation scheme for given, accessible and new information (information status), which accounts for a less controversial, though substantial, share of the information structure of linguistic data. section 4 discusses the distinction between non-contrastive and contrastive focus and presents a method of how to spot contrastive (alternative-eliciting) features in corpus data. in section 5, we return to the issue of primary and secondary foci. 2. determination of focus: top-down vs. bottom-up taking another look at example (7), we may find it surprisingly difficult to tell how many foci it contains. the simplest choice is to say that the entire sentence represents a single broad focus. an obvious justification is that (7) contains only discourse-new information, and therefore serves to answer the big question (roberts, 1996) what is the way things are?, or simpler what happened? for several reasons this cannot be a satisfactory solution. as it is known from the work of groenendijk and stokhof (1984) and roberts (1996), assertions need not be complete answers; in fact, an unrestricted question like what happened? is never answered completely. but if an answer is only partial with regard to the big question, as is the case with (7), it may carry additional information about the structure of the conversation. according to roberts, the prosody of assertions reflects the (immediate) question under discussion (qud), which is more specific than the big question and usually implicit. büring (2003) develops this idea further, elaborating on example (8). (8) fredct ate the beansf . consider (8) as a discourse-initial assertion. the contrastive topic in (8), marked by a rising accent on fred, signals the discourse strategy of the speaker to answer the question who ate what? by first providing an answer to the subquestion what did fred eat? our german corpus example (7) shows – perhaps incidentally9 – a very similar pattern at the beginning. it is in accordance with our interpretation of the sentence to assume that (7) has an analogous structure to (8), i.e. a contrastive topic on the phrase in ghana; in other words, (7) seems to signal a strategy – technically a stack – consisting of the two implicit questions (i) what happened where? and (ii) what happened in ghana? we infer that the speaker intends to continue, at some stage, with news about other countries.10 if in ghana is indeed a contrastive topic (and therefore, a special kind of focus), and if the remainder of the clause is the main focus which actually answers question (ii), then the total number of independent focal constituents in the sentence rises to at least two. determining the focal structure of some utterance in the way described – by reasoning about implicitly asked questions – is what we might call a top-down process of focus identification: first, determine what is being asked for, or under discussion, and what is the discourse strategy to which 9. note that we do not propose to rely on the actually found pitch accents to guide us in the annotation process. we are not claiming that a rising pitch accent necessarily signals a contrastive topic. clearly, contrastive topics need to be identified via pragmatic reasoning. what we hope to eventually find, however, is statistical evidence in the prosodic domain that supports the chosen pragmatic analysis. 10. it must be noted though that intentions like this are often not more than vague promises, and that strategies may fall prey to memory decay. this goes against roberts’s assumption that questions have to remain on the qud stack until resolved or determined unanswerable. another conceivable situation is that a person reading a text simply has a false anticipation of what is to follow. 219 riester and baumann it belongs; second, identify the phrases which provide an answer to the current question under discussion; and third, predict or observe how these phrases are marked prosodically (and/or morphosyntactically).11 in other words, the marking itself should not be used to identify the focus or the contrastive topic. we do not claim that we already possess a universal procedure which lives up to these standards. nevertheless, the top-down approach works straightforwardly in simple cases like (9), in which the accented word john is both the focus and the answer to the overt qud. (9) a: who spilled the wine? b: johnf spilled it. the procedure can also be applied to more complex cases involving nested phrases and disjoint foci. (10) is a german example from höhle (1982), discussed in gussenhoven (1999). (10) a: was hat das kind erlebt? what happened to the child? b: karlf hat dem kind [einen füller geschenkt]f . karl gave the child a fountain pen. the phrase which provides the answer to the question under discussion – i.e. the focus – is split in two parts, karl and einen füller geschenkt. gussenhoven (1983, 1992, 1999) offers a general rule for predicting the accent pattern of a focus that has been pragmatically determined, the so-called sentence accent assignment rule (saar), which we reproduce here in a simplified form: 1. place a pitch accent on every content word in focus. 2. then, deaccent every focussed predicate which is adjacent to an accented argument. the saar correctly predicts the pitch accents on karl and on füller (‘fountain pen’), the lack of pitch accent on the unfocussed (backgrounded) word kind (‘child’) and the deaccenting of the predicate geschenkt (‘given as a present’).12 as for example (7), we already assumed that the focus encompasses the phrase in (11). (11) [fand ein festakt im ehemaligen sklavenhandelsposten elmina statt]f? again, the saar correctly accounts for the deaccentuation of fand. . . statt (‘took place’) which is adjacent to its argument ein festakt (‘a ceremonial act’). the remaining content words receive a pitch accent because they do not stand in any predicate-argument relation. what gussenhoven’s rule cannot accomplish is to explain the intricate patterns of prosodic phrasing. neither does it tell us whether the focus in (11) is a single one or whether it actually consists of several smaller foci. in order to answer this question, we return to roberts’s (1996) theory of questions under discussion. when answering question (12a) by means of (7), we are actually behaving in an over-informative way. in some sense, it would have been enough to use the simpler answer given in (12b). 11. as is well-known, some languages do not mark focus by means of prosodic prominence. other devices found crosslinguistically include moving the focal constituent to a particular syntactic position in the sentence, or attaching a special focus morpheme. 12. deaccentuation of a predicate with an accented argument is explained in a similar fashion within selkirk’s (1984,1995) focus projection framework. 220 focus triggers and focus types (12) a. what happened in ghana? b. [in ghana]ct [fand ein festakt statt]f . (13) im ehemaligen sklavenhandelsposten elmina in fact, the over-informative prepositional phrase (13) is a non-restrictive modifier and therefore belongs to the class of so-called supplemental expressions, which potts (2005: 6) defines as conventional implicatures or as not-at-issue. simons et al. (2010) define at-issueness by saying that a proposition p is at-issue if and only if the question whether p is true or not is relevant to the qud. in our case, the proposition expressed by the pp, namely the location of the ceremonial act is in elmina, is not an answer to (12a), but rather to the supplemental question in (14). (14) where did the ceremonial act take place? in fact, we can iterate this process once more by stating that the information that elmina is a former slave trading post – the meaning expressed by the non-restrictive modifier of elmina – is not an answer to (14) but to (15). (15) what is elmina? this, admittedly complex, line of reasoning brings us to the following tentative conclusion: it is very likely that our corpus sentence actually consist of four independent pragmatic foci (including one contrastive topic), as shown in (16). (16) [in {{ ghana]ct l*h -} [fand { ein festakt l*h -} [im { [ehemaligen h* sklavenhandelsposten]f3 l*h elmina]f2 h*l statt]f1 %}} we can furthermore state that ct, f1 and f2 are marked by nuclear pitch accents – in the terminology defined in the previous section, they represent primary foci (see definition 1 above). f3 is only marked by (two) prenuclear accents; we therefore call it a secondary focus (definition 2). the top-down approach, as we have sketched it, has two important advantages: it is crosslinguistically applicable, even to languages which do not mark information structure prosodically, and it does not confuse form (pitch accent) and meaning (focus). it allows us to use the following language-independent pragmatic focus definition. definition 3 (focus) a focus is either an answer to the immediate (explicit or implicit) question under discussion (at-issue focus) or to a supplemental question (not-at-issue focus).13 the problem is that the identification of quds in corpus data is still in its infancy, and the current paper does not purport to propose a general procedure for it, although this is what clearly should be envisaged in the future. a competitor to the top-down approach, which is intrinsic to certain well-established theories of focus, including rooth (1992) and selkirk (1995), is the assumption that pitch accents necessarily indicate a focus.14 we call this the bottom-up approach to focus identification. obviously, 13. in (16), we may call f1 an at-issue focus, and f2/f3 not-at-issue foci. 14. quote selkirk (1995: 555): “the basic focus rule states that the assignment of a pitch accent to a word entails the f-marking of the word[.]” 221 riester and baumann the bottom-up approach is not linguistically universal, and even for english or german we cannot be entirely sure whether the assumption is empirically sound, especially if secondary prominence comes into play. note that, under the bottom-up approach, example (16) would not contain four focal expressions but five, since the prenuclear pitch accent on ehemaligen would count as an autonomous marker of focus.15 nevertheless, at the current stage, we cannot entirely do without the bottom-up approach as long as the top-down approach is not fully worked out. this will become clear in section 4, when we will discuss whether an observed nuclear pitch accent is used to express contrastive focus or not. in the following sections, we will discuss those aspects of information structure which we think can be annotated without too much controversy. this will enable us to create pragmatically enhanced corpus resources, which may subsequently serve as a basis for empirical investigation, even though we do not yet caim to annotate all aspects of focus. other proposals of annotating information structure involve, for instance, hajičová et al. (2000), paggio (2006), götze et al. (2007), or cook and bildhauer (2013). in paggio (2006), focus is defined, following lambrecht (1994), as “non-presupposed” information. further rules given to the annotators include, for instance, the requirement to identify at least one focus per sentence and at least one “main accent” per focus (for danish data). in götze at al. (2007: 170), the definition for focus is “[that] part of an expression which provides the most relevant information in a particular context [. . . ]”. subcategories for new information focus and contrastive focus are defined. we believe that these proposals capture some important aspects of focus in corpus data but are less complete and less detailed than the picture that we are developing in this article. 3. annotating given-new information: the reflex scheme in accordance with definition 3, we adopt a concept of focus in terms of answers to explicit or implicit questions. this is not an uncontroversial decision. often, what counts as the answer to an implicit question can simply be called new information. however, not everybody seems to agree that new information which is not otherwise marked (e.g. occurring in a contrastive constellation or being associated with a focus-sensitive particle) deserves to be called focus. consider the following quote by elisabeth selkirk: [. . . ] i am using the simple term “focus” to refer to “contrastive focus” [. . . ], as involving roothian alternatives. this should not be confused with the use of the term “focus” to indicate newness in the discourse, a use which this paper argues should not be made. (selkirk, 2008: section 4, footnote 9) we agree with selkirk’s view that there is a linguistically relevant distinction between contrastive focus and (purely) new information in terms of their pragmatic meaning (more on this in section 4). besides, as selkirk and numerous others have shown in their work, contrastive focus tends to receive higher prosodic prominence and sometimes uses syntactically less canonical structures than new information, cf. repp (2010) and references therein. however, we object against selkirk’s conclusion that newness should not be called focus, since this is in conflict with the qud approach. 15. leaving out the h* accent on ehemaligen would presumably violate the givenness principle (schwarzschild, 1999), which says that an expression must either be given or f-marked. focus projection rules (selkirk, 1995) do not allow an f-marker to project from sklavenhandelsposten since it is not an argument of ehemaligen. therefore, the adjective needs its own pitch accent. 222 focus triggers and focus types in section 4, we will account for the distinction between contrastive and non-contrastive (novelty) focus in a way that is compatible with both quds and rooth’s alternative semantics. the common denominator of new and contrastive constituents is their ability to answer questions – to reduce the number of possibilities the world might be like – and this is what should be taken as the defining characteristic of focus in general. in section 2, we have given a sketch of the top-down analysis of the focal structure of natural discourse, not yet ready for the use in linguistic annotation. in this section, as a first step towards the annotation of focus we are going to provide a scheme for annotating (different types of) given and new (as well as accessible) information, what is also called information status, following the system of baumann and riester (2012). the scheme combines earlier accounts of information status in the aftermath of prince (1981, 1992) – notably chafe (1994); lambrecht (1994); eckert and strube (2000); nissim et al. (2004); götze et al. (2007); riester et al. (2010) – with the givenness theory by schwarzschild (1999). to classify the constituents of a sentence into given and non-given ones is a first move towards identifying its background-focus structure. schwarzschild’s givenness principle says that non-given constituents must be “f-marked” (and are therefore, in some sense, focal).16 the procedure described in this section will not tell us how many different foci a sentence contains. but it already provides us with a fine-grained analysis with a view to its prosodic correlates; see baumann and riester (in press). the system distinguishes between a referential level and a lexical level (and is therefore called the reflex scheme). we will clarify why it is desirable to use such a fine-grained system rather than just distinguishing between “given” and “new” constituents. note well that we are not claiming that the annotation labels presented below represent syntactic features of some kind, in the way as, for instance, selkirk (2008) treats her f and g markings. we will make no predictions as regards the precise functioning of the syntax-phonology interface. the category descriptions below are deliberately kept short, since we have introduced them in great detail in baumann and riester (2012). 3.1 r-given and l-given givenness, loosely following schwarzschild (1999: 151), can be interpreted as either synonymy / hyponymy of lexemes (and the concepts they express), or as identity between referring expressions. likewise, halliday and hasan (1976: 288) distinguish between lexical cohesion and various referential relations. we call the two notions l-givenness and r-givenness, respectively. interesting constellations can be observed if the two notions are applied simultaneously, as shown below. by use of the r-categories it is possible to classify referential determiner phrases and prepositional phrases occurring in natural discourse; by use of the l-categories we can classify the information status of content words and non-referential phrases. r-labels apply at the dp or pp level. for instance, in examples (17), (18) and (20) we find various kinds of coreferential expressions. lexical givenness, on the other hand, applies in (18) and (20) on the repeated words, and in (19) on the hypernym man. (17) a colleague came in. the idiot dropped a vase. r-given 16. a warning: beaver and clark (2008: 15) notice that not all f-markers used in the literature have the same function, e.g. f-markers used in rooth (1992) and in selkirk (1995) are used for different purposes. 223 riester and baumann (18) a student came in. another student greeted him. l-given r-given (19) a policeman came in. another man left. l-given (20) a man came in. the man coughed. l-given r-given the most important take-home message is that neither is referential givenness a prerequisite for lexical givenness, as shown in (17), nor vice versa, see (18) and (19), although the two sometimes combine, as in (20). 3.2 r-new, l-new, r-unused novelty is, on most treatments of information structure and discussions of the given/new distinction, understood as “novelty in the discourse”. remarkably however, prince (1992) additionally distinguishes between discourse novelty and hearer novelty, the latter representing a stronger notion since unmentioned (i.e. discourse-new) entities may nevertheless be familiar to the addressee (i.e. hearerold). in her earlier paper, prince (1981) uses the labels unused (discourse-new, hearer-old) and brand-new (discourse-new, hearer-new) for the same opposition. the labels r-new and r-unused that are employed on our account are defined in a slightly different way: both describe discoursenew referential expressions but, while r-new is reserved for indefinites, r-unused stands for uniquely identifiable, definite, but not necessarily known, entities used on the first occasion in a text. this decision, on the one hand, does justice to the long-standing semantic tradition to keep indefinites and definites (for instance, proper names) apart, and, on the other hand, accounts for the difficulty to decide with certainty whether, for instance, a named entity is hearer-known or not, cf. riester et al. (2010). independently of what has just been said, it is furthermore possible to separately describe the discourse novelty of lexemes (l-new) and of the discourse referents (r-new, r-unused) which they introduce. examples of the three categories in combination are given in (21) to (23). (21) a man came in. another man left. l-new l-given r-new r-new (22) george came in. mary likes george. l-new l-new l-given r-unused r-unused r-given (23) the man who stole my wallet yesterday is very tall. l-new l-new r-unused r-unused the complex subject phrase in example (23) shows that information status needs to be assigned recursively. this is an issue which is of particular relevance for the language of news, which contains many expressions with several embeddings, cf. riester et al. (2010). 224 focus triggers and focus types 3.3 r-bridging, l-accessible prince (1981) and also chafe (1994) have pointed out that it is desirable to not only distinguish between given and new information but to take into account at least a third, intermediate, class: expressions which have not been mentioned explicitly but are inferrable from material in the discourse. chafe (1994) uses the term accessible for such information but does not distinguish between different levels, as we would like to do. as far as discourse referents are concerned, a closely related phenomenon has been discussed under the notion of bridging or associative anaphora (clark, 1977; asher and lascarides, 1998; löbner, 1998; poesio and vieira, 1998), shown in example (24). (24) bill discovered a romantic house. the door was open. l-new l-new l-accessible r-unused r-new r-bridging the label l-accessible is defined for words which are hyponyms or meronyms (part expressions) of other words in the recent discourse context.17 the label r-bridging, on the other hand, is defined quite differently as a definite expression whose licensing depends on a previously introduced scenario or frame. so, while in (24), house and door stand in a whole-part relation (door is lexically accessible), no such relation exists between murdered and harpoon in (25). since the harpoon is an unusual murder instrument, it is labeled l-new. nevertheless, we would still like to say that this is a case of bridging, since the second sentence could not be uttered felicitously at the beginning of a discourse. (25) john was murdered yesterday. the harpoon was lying nearby. l-new l-new r-unused r-bridging other than in the case of r-unused expressions, the interpretation of items labeled r-bridging is context-dependent. in contrast to the label r-given, r-bridging implies non-coreference. indefinites never receive the label r-bridging in the present system. in (26), lexical accessibility combines with referential novelty.18 (26) john lives in italy and is married to a neapolitan. l-new l-new l-accessible r-unused r-unused r-new 3.4 r-generic definite or indefinite expressions which refer to a kind, see (27) and (28), receive the label r-generic. (27) the lion has a mane. l-new l-new r-generic r-generic 17. we assume a somewhat arbitrary window of five sentences. a more precise range in which word associations still play a role needs to be determined experimentally. 18. arguably in (26), in addition to the l-label on the word italy, the dp italy and the pp in italy both should receive separate r-unused labels, since the former refers to the country itself and the latter to a location. 225 riester and baumann (28) mary likes vegetables. john likes vegetables, too. l-new l-new l-new l-given r-unused r-generic r-unused r-generic as can be seen in (28), we do not treat the repeated mention of a generic expression (vegetables) as a case of coreference (r-given) but merely as repetition of the same concept (l-given, r-generic). 3.5 overview and annotation of higher syntactic constituents r-level l-level units: dp, pp, that-cp units: a(p), adv(p), n(p), v(p), s label description label description r-given coreferential l-given word identity / anaphor synonym / hypernym / holonym / superset r-bridging non-coreferential l-accessible hyponym / meronym / context-dependent subset / co-hyponym / expression related r-unused definite l-new unrelated expression discourse-new (within last five expression clauses) r-new specific indefinite r-generic generic definite or indefinite other e.g. cataphors table 1: overview of basic reflex scheme table 1 contains the most important labels of the reflex annotation scheme. for a more detailed list of labels consult baumann and riester (2012). in the following, we will turn to a number of practical issues which arise when we apply the annotation scheme to corpus data. as we said at the beginning of this section, we want to use our annotation system in order to arrive at a comprehensive identification of the given and non-given parts of linguistic data since we consider the latter as indicating focal material. in order to achieve this goal we cannot confine our analysis to referring expressions, as it has been done in e.g. prince (1981); nissim et al. (2004); götze et al. (2007); riester et al. (2010), but need to extend the annotations to other content expressions like adjectives, verbs and adverbs as well as their syntactic projections. the account builds upon schwarzschild’s (1999) theory of focus and givenness. of course, the question of what counts as a unit for annotation is influenced by the choice of syntactic theory which the analysis is based on. a principled distinction can be made between expressions which refer to some entity, like an individual, a place, 226 focus triggers and focus types s [l-level]hhhhhhhhh ((((((((( dp [r-level] ppppp ����� d the np [l-level] pppp ���� ap [l-level] a [l-level] tall np [l-level] n [l-level] man vp [l-level] aaaa !!!! v [l-level] arrived pp [r-level] h hh � �� p in dp [r-level] np [l-level] n [l-level] tuscany figure 2: basic target units for reflex annotations a fact etc. (dp, pp, that-cp),19 and expressions which denote a property / set of entities (np, ap, vp, advp, s) or a relation.20 as is shown in figure 2, the former are assigned r-labels, the latter l-labels. in practice, some of these labels will be redundant and can be left out. for instance, one of the l-labels at either the n or the np level above man can of course be dropped, since they are identical. what we are proposing amounts to a practical explication – and further development – of the approach taken by schwarzschild (1999: 151), who distinguishes between categories of type e (r-level) and of type 〈α, β〉 (l-level). our definition of the l-level, however, is much simpler than schwarzschild’s since we completely abandon his notion of existential f-closure. however, we make use of his idea to generalise lexical relations to a notion of entailment.21 as we said, in corpus annotation practice, the system will have to be adapted to various constraining factors, such as the properties of the chosen parser with its specific syntactic tagset, as well as features of the annotation tool. figure 3 shows a sample annotation of a german sentence from the dirndl corpus (eckart et al., 2012).22 19. one reviewer criticised that pps should count as properties rather than individual type entities, which is a common assumption in semantics. what we are after, however, is the referent of an argument, which often comes in the form of a pp. sometimes pps refer to a place or time; sometimes the preposition is subcategorised by the predicate and semantically empty; sometimes, in german, preposition and determiner are amalgamated (im – ‘in the’, zum – ‘to the’ etc.). in all those cases the simplest practical choice is to assign referential information status to the pp. a second criticism pertained to allegedly non-referential quantifiers like few people, every dog. in corpus data, however, such expressions almost always introduce or refer back to group entities, analogously to indefinites and definites. 20. we assume the dp hypothesis (abney, 1987). accordingly, we take nps to denote properties, i.e. sets of individuals, whereas dps denote (or refer to) a single individual or group entity. we abstain from the debate whether a sentence should be analysed as an ip (inflection phrase), as a cp (as in the german example shown in fig. 3) or as a projection of tense, aspect or voice, cf. adger (2003). 21. according to this approach, the previous mention of chihuahua entails the successively mentioned hypernym dog, as well as a successive mention of small dog, cf. baumann and riester (2012: 133ff.) 22. the sentence was parsed using xle (crouch et al., 1993-2011, ms.) and the german lexical functional grammar implementation by rohrer and forst (2006), and converted to be used with the salto tool (burchardt et al., 2006), which produces output in tiger/salsa-xml. in the rest of the paper, we shall abstract over such individual choices, since it is our goal to provide the general annotation procedure and not one that is tied to a specific annotation tool, format or syntactic theory. 227 riester and baumann figure 3: sentence annotated in salto, dirndl s165, 25-03-2007, 5:00: a strong earthquake has hit central japan. in the following, we will briefly show how the extended annotations can be applied to news data. example (29) has been slightly adapted for ease of demonstration. the analysis is shown in (30) to (32), using a simplified table notation. note that in our envisaged annotation process of information status, the labellers will have no access to prosodic information. (29) a. ein starkes erdbeben hat zentral-japan erschüttert. a strong earthquake has hit central japan. b. die behörden gaben eine tsunami-warnung für den südwesten heraus. the authorities have issued a tsunami warning for the southwest. c. auch im inselstaat vanuatu im südpazifik wurden zwei beben registriert. also in the island state of vanuatu in the south pacific two earthquakes have been registered. (30) ein starkes erdbeben hat zentral-japan erschüttert. a strong earthquake has central japan shaken (ap ) l-new (n) l-new (np ) l-new (v ) l-new (np ) l-new (dp ) r-unused (dp ) r-new (v p ) l-new (s) l-new (31) die behörden gaben eine tsunami-warnung für den südwesten heraus. the authorities gave a tsunami warning for the southwest out (np ) l-new (v ) l-new (np ) l-new (np ) l-new (v ) l-new (pp ) r-bridging (dp ) r-bridging (dp ) r-new (v p ) l-new (s)l-new 228 focus triggers and focus types (32) auch im inselstaat vanuatu im südpazifik wurden zwei beben registriert. also in the island state vanuatu in the south pacific were two quakes registered (n)l-new (n)l-new (np )l-new (np )l-given (v )l-new (np ) l-new (pp ) r-unused (dp ) r-new (pp ) r-unused (v p ) l-new (s) l-new the annotation proceeds along the principles defined above and consists of the following steps: 1. all referring expressions (dps and pps) receive an r-label. the phrases the authorities and (for) the southwest are linked to (or anchored in) central japan via bridging. in cases of syntactic embedding, e.g. [in the island state of vanuatu [in the south pacific]], r-labels are nested inside each other. 2. all content words (including verbal particles) receive an l-label (topmost line). in our example, the only l-given word is quakes in sentence (32), which is a near-synonym of earthquake in sentence (30). (still the earthquakes are referentially distinct and independent from each other, thus the r-new label on two quakes.) it is debatable whether the phrase tsunami warning is l-new or rather l-accessible due to its intuitive relation to earthquake. however, since it is neither a hyponym nor a meronym of the latter there is so far no clear criterion which would license a classification of the expression as l-accessible. 3. following the syntactic structure of the sentences, complex non-referential phrases are assigned l-labels as well. a complex phrase counts as l-given if it is entailed by another phrase in the discourse. this does not occur in the present examples but see section 5, example (55). a phrase is l-accessible if it entails an earlier phrase. at the sentence level, this might happen with elaborations, e.g. in (33). (33) a. sandy bought a car. b. [l-accessible she chose a hybrid model]. in order to test the reliability of some aspects of the reflex scheme, we had two trained student annotators independently assign r-labels to 3445 referring dps/pps in written news text from the dirndl corpus (eckart et al., 2012), as well as l-labels to 5045 content words.23 the annotations were done in salto (burchardt et al., 2006). since the annotators themselves had to identify the markables in pre-parsed syntactic representations, the spans of some of the markables had to be adjusted after annotation. for the label granularity defined in table 1, an evaluation of inter-annotator agreement following cohen (1960) and artstein and poesio (2008) yields κ = 0.75 for the r-level and κ = 0.64 for the l-level. the value for the l-level is lower than in earlier annotation experiments conducted by the authors of this article. the main reason for the relatively low score of the l-level seems to be a certain insecurity in classifying an expression as l-accessible. obviously, there are more ways in which two items can be lexically related than the ones we list in table 1. 23. for this evaluation we did not consider that-cps and complex phrases. 229 riester and baumann 4. contrastive focus vs. novelty focus a longstanding issue in information structure theory is the differentiation between so-called contrastive focus and novelty focus (information focus). when occurring in isolation, both types of focus are marked by nuclear pitch accents in english and german. according to definition 1, both may be primary foci, although there is evidence that novelty focus sometimes receives comparatively weaker (or at least different) marking with regard to certain prosodic or acoustic parameters, see e.g. alter et al. (2001); selkirk (2002); hedberg and sosa (2008); hermes et al. (2008); katz and selkirk (2011). the notion of contrast has received various interpretations in the literature on information structure and discourse structure, see umbach (2004); repp (2010). the most straightforward, though not tenable, definition is in terms of the explicit mention of alternatives. example (34) contains a pair of contrastive topics and a pair of contrastive foci in a parallel structure; the focus in (35) does not have an overt alternative and simply represents new information. (34) johnct1 ordered waterf1, and maryct2 ordered beerf2. (35) mary went into a pub. she [ordered beer]f . (36) only maryf drank beer. example (36) shows that explicit mention is not the only criterion for contrast. the exhaustive particle only in (36) requires a domain of individuals – an alternative set – who did not drink beer, except for mary. rooth (1992) presents a uniform account of alternative-eliciting focus for cases of association with focus-sensitive particles, overt contrast, scalar implicatures, question-answer pairs, ellipsis and comparatives. several researchers, including selkirk (2008), have taken these phenomena to establish the paradigm of contrastive focus. occasionally, contrastive focus has been assigned stricter definitions, especially in terms of correction or of exhaustivity/identification. while corrective contrast is often used in the design of minimal pairs of non-contrastive and contrastive contexts – see (37) vs. (38) – e.g. for the use in experiments of laboratory phonology, we argue that it should be seen as an extreme case of contrast (involving the rejection of a previous utterance) which might possess its distinct prosodic marking. (37) a: what did mary drink? b: she drank beerf . (non-contrastive / non-corrective) (38) a: mary drank water. b: (no.) she drank beerf . (contrastive / corrective) exhaustivity need not be expressed by means of a focus-sensitive particle like in (36) but also occurs with it-clefts (atlas and levinson, 1981; hedberg, 1990; delin and oberlander, 1995; reeve, 2011), as a default interpretation of certain syntactic positions, like the preverbal position in hungarian (szabolcsi, 1981; é. kiss, 1998; kenesei, 2006; horvath, 2010), or simply arises as a conversational implicature (schulz and van rooij, 2006; spector, 2006) like in (39). b’s answer is interpreted as saying that mary was in the pub but no one else of a certain group which the interlocutors have in mind. (39) a: who was at the pub? b: maryf was there. 230 focus triggers and focus types it can be assumed that in most cases in which an answer to a question is interpreted exhaustively, an alternative set has been introduced or accommodated beforehand. interestingly however, é. kiss (1998), who discusses exhaustivity in her account of so-called identificational focus, assumes it to be an independent property from contrastivity. while exhaustivity requires the exclusion but not necessarily the identification of the alternatives, contrastivity only requires that the alternatives form “a closed set of entities whose members are known to the participants of the discourse” (é. kiss, 1998: 267) – but not necessarily their exclusion. we believe that the latter definition is very appealing, and will demonstrate below that it can be nicely applied when annotating natural language data. before doing so, however, we will return to selkirk’s (2008) criticism mentioned at the beginning of section 3. recall that selkirk argues for a distinction between (contrastive) focus and discourse-new information, and against calling the latter focus. in doing so, she cites the influential work by rooth (1992), who however – as far as we can tell – leaves it open whether his theory of alternative semantics also applies to plainly new information. against selkirk, we argue that there is a straightforward move to integrate new information into alternative semantics. this, however, requires some degree of exegesis of rooth (1992) as it has been undertaken by riester and kamp (2010). the key to solving the problem lies in taking the semantics of focus as suggested in rooth (1992) more literally than has been done in parts of the contemporary literature on focus semantics. alternative semantics, in a nutshell, provides us with two important semantic components for a theory of focus: the first component is the f-feature, which, when applied to some syntactic constituent, introduces a second meaning, also called the focus semantic value, which is a set of elements of the same semantic type as the focussed constituent. the focus semantic value of the focussed expression maryf is de, the (unrestricted) set of individuals. büring (2013) has used the attribute “raw” for such unfiltered focus semantic values. riester and kamp (2010) note that, for natural discourse, we have to assume that de contains each and every individual on earth, since it has not undergone any restriction. therefore, it certainly does not meet the requirements for contrastiveness formulated by é. kiss, namely that the alternatives form a closed set and be known to the participants of the discourse. in fact, the semantics of the f-marker can be seen as an operation which is completely blind to contextual information. riester and kamp (2010) call the focus semantic value of a natural language expression an “anonymous” alternative set because its elements are, at the outset, unidentified. we shall assume, however, that such anonymous alternative sets, the result of f-assignment, are sufficent for the f-marked phrase to be called a novelty focus. consider example (40), the beginning of a news feature, and its spoken realization in figure 4. (40) bundespräsident köhler [hat das gesetz zur gesundheitsreform unterschrieben]f . federal president köhler has signed the bill on the health care reform. let us assume that the vp of (40) is focussed – while ignoring other properties the sentence might have. alternative semantics tells us that the focus semantic value will consist of other vp denotations (properties) than signing the health bill. this means that the president might have “done other things”, although we will have difficulties reaching an agreement what his specific other options had been, since neither the discourse context nor world knowledge tell us. the vp focus simply represents new information; the sentence is informative – a contingency24 – but nothing more. 24. we are grateful to carla umbach (p.c.) for suggesting the notion in this context. 231 riester and baumann figure 4: novelty focus; dirndl s1790, 26-03-2007, 18:00, 3’20”: has signed the bill on the health care reform the second important component of focus semantics which rooth (1992) introduces is the socalled focus interpretation operator ∼ (“squiggle”). semantic-pragmatically, ∼ is defined as an anaphoric operator, which minimally imposes the following constraints on the interpretation of the constituent it attaches to:25 identify at least one proper alternative in the context 1. which matches the pattern defined by the focus semantic value of the phrase, and 2. which is different from the ordinary meaning of the phrase. the attachment site of the ∼ operator is sometimes called the focus domain (büring, 2008; rooth, 2010) or, perhaps, the focus phrase (krifka, 2006). however, opinions differ where to attach ∼. (41) die europäische union hat den sklavenhandel früherer jahrhunderte bedauert und sich gegen formen ∼[neuzeitlicherf sklaverei] gewandt. the european union has expressed regret over the slave trade of earlier centuries and turned against forms of modern slavery. (dirndl s1618, 26-03-2007, 14:00, 5’47”) example (41) nicely illustrates the fulfilment of the constraints imposed by the ∼ operator. (for the sake of simplicity we only highlight the nuclear pitch accent on the relevant phrase.) the phrase neuzeitlicher sklaverei (‘modern slavery’) has a focus semantic value of the form {[[x slavery ]] |x is some intersective modifier}. the discourse context contains the phrase den sklavenhandel früherer 25. for a formal specification in drt and further explications see riester and kamp (2010). note that there might be uses of ∼ in the literature which are not compatible with our strict anaphoric semantics. 232 focus triggers and focus types jahrhunderte (‘the slave trade of earlier centuries’), whose ordinary meaning matches the template provided by (is an element of) the just mentioned focus semantic value (assuming that slavery and slave trade may be read as synonyms). the two complex expressions are distinct from each other and represent proper alternatives.26 as we will clarify below, the identification of contrastive alternatives is not limited to the actual discourse context but does often rely on situational or world knowledge. this does not mean that identifying contrastive alternatives is always possible, as we see in (40). we are now in a position to delimit (non-contrastive) novelty focus from contrastive focus. definition 4 (novelty focus) a novelty focus is a non-given constituent which serves as an answer to the immediate qud or to a supplemental question and whose alternatives remain unidentified. definition 5 (contrastive focus) a contrastive focus is a constituent which serves as an answer to the immediate qud or to a supplemental question and whose alternatives can be unanimously identified in the context (i.e. discourse context, encyclopaedic knowledge, lexicon etc.) let us briefly summarise our theoretical assumptions surrounding the focus notion. we postulate that both new and contrastive information trigger (or represent) focus. novelty focus merely carries the f-marker generating a focus semantic value as defined in rooth (1992: 76). contrastive focus, too, comes with an f-marker but is additionally interpreted by means of the anaphoric ∼ operator. the result of successful focus interpretation is the identification of at least one contrastive alternative. we loosely follow é. kiss (1998) in assuming that contrastive identification means that the addressee is able to name with certainty who or what this contrastive alternative (or the set of alternatives) in the respective context is (definition 5). novelty focus, on the other hand, is defined as a focus with an anonymous – not clearly identifiable – alternative set (definition 4). contrastive constituents may contain, or entirely consist of, given material. we see this in example (41), or in (42) – a selectional focus, which makes a choice from a given list. (42) john and mary were at the pub but ∼[johnf ] left early. this means that, in order to identify contrastive focus in corpus data we cannot confine ourselves to discourse-new constituents and decide whether their alternatives are identifiable or not. rather, we must look for specific alternative-eliciting features or constellations (alt). a list of such features is presented in table 2.27 note well that having spotted an alternative-eliciting constellation is not yet having identified contrastive focus. beaver and velleman (2011: 1674) state that speakers are not obliged to mark contrast. we agree with this view. the fact that e.g. two co-hyponymic expressions, say an elephant and a lion, occur as the arguments of the same predicate (alt-overt-arg), as in (43), does not mean that they are necessarily contrasted against each other, although the speaker might decide to establish the contrast prosodically; for instance by forming two prosodic phrases rather than one. this would then signal that the order of arguments could have been the other way around. (43) [alt-overt-arg an elephant] chased [alt-overt-arg a lion]. 26. wagner (2006) shows that there has to be sortal compliance between two alternatives. for instance, high-end convertible is a proper alternative to cheap convertible but not e.g. to red convertible. what we presumably want is a kind of co-hyponymy, or what lang and umbach (2002) call the availability of a common integrator. 27. a different list of features for “contrastive focus” is provided in götze et al. (2007: 178ff). 233 riester and baumann sublabel of alt description fsp item is associated with a focus-sensitive particle. overt item is an element of a pair or list of overtly contrastive expressions -arg type-identical arguments of the same predicate. -comp items occur in a comparative construction. -coord items are coordinated. -ext items occur in different sentences – sentence-external contrast sel item selects one element from a pair or list of previously introduced alternatives. table 2: alternative-eliciting features on the other hand, there are also triggers, e.g. focus-sensitive particles, which seem to necessarily associate with a contrastive focus (alt-fsp).28 when we apply our set of features from table 2 to example (29), we obtain an additional tier of elicited alternatives, shown in (44) and (45). (44) ein starkes erdbeben hat zentral-japan erschüttert. alt-overt-ext (45) auch im inselstaat vanuatu im südpazifik wurden zwei beben registriert. alt-fsp / alt-overt-ext the phrase im inselstaat vanuatu im südpazifik associates with the additive particle auch. it can furthermore be contrasted with zentral-japan. in the following, we give corpus examples for the remaining alternative-eliciting features. exchangeable arguments of a predicate (alt-overt-arg): (46) [alt-overt-arg die fluggesellschaft air berlin] übernimmt [alt-overt-arg den düsseldorfer konkurrenten ltu]. airline company air berlin will absorb its düsseldorf rival ltu. (dirndl s2428, 27-03-2007, 08:00) 28. the following table shows the frequencies of the most common conventional fsps in german, in a sample of 2484 sentences from the dirndl corpus. however, about 90% of the sentences do not contain an fsp from the list, which shows the need to identify other features than just alt-fsp. particle translation n auch too 107 nur only 62 wieder again 53 ebenfalls too 8 lediglich only 5 nicht einmal not even 1 weder . . . noch neither . . . nor 1 sogar even – bloß only – 234 focus triggers and focus types comparative (alt-overt-comp): (47) huber meinte in dortmund zur begründung, [alt-overt-comp in der metall-branche] sei die lage stabiler als [alt-overt-comp in der chemieindustrie]. huber explained in dortmund that the situation in the metal industry was more stable than in the chemical industry. (dirndl s1943, 26-03-2007, 21:00) coordination (alt-overt-coord): (48) falls sich [alt-overt-coord die protestantische unionisten-partei dup] und [alt-overt-coord die katholische sinn féin] bis 24 uhr ortszeit nicht auf ein bündnis einigen, [. . . ] unless the protestant unionist party dup and the catholic sinn féin reach an agreement on an alliance until 12 p.m. local time [. . . ] (dirndl s1249, 26-03-2007, 07:00) parallelism, sentence-external contrast (alt-overt-ext): (49) [alt-overt-ext kevin kuranyi] schoss in prag beide tore für die deutsche elf. [alt-overt-ext milan baros] erzielte den anschlusstreffer. in prague, kevin kuranyi scored both goals for the german team. milan baros scored the other goal. (dirndl s59-s60, 25-03-2007, 01:00) selection (alt-sel): (50) darauf verständigten sich die parteivorsitzenden paisley und adams bei ihrem ersten persönlichen treffen in belfast. das amt des ersten ministers soll [alt-sel paisley] übernehmen. this was arranged by the party leaders paisley and adams at their first personal meeting in belfast. paisley will become the first minister. (dirndl s1897, 26-03-2007, 20:00) as we said, we are not claiming that speakers have to mark all of the above alternative-eliciting constellations prosodically, although they often will do so. more research is necessary on each of them. the fact that marking contrastive focus is sometimes optional is one reason why text allows for intonational variation when being read aloud. from an annotation perspective, however, this means that identifying all cases of contrastive focus requires some advanced (top-down) reasoning about potential questions under discussion, including supplemental questions, as sketched in section 2. alternatively, it requires reverting to the spoken signal – what we have called the bottom-up approach – including the assumption that nuclear pitch accents always signal focus, which we can then classify as either contrastive or non-contrastive.29 a final problem which we currently cannot solve without falling back on spoken language is the identification of implicit contrast. by this we mean obvious cases of contrastive interpretation which do not come with one of the above features and whose marking likewise seems optional. 29. as should have become clear by now, this will leave a number of problems unsolved, including the question which other forms of prosodic prominence are also markers of focus. 235 riester and baumann consider figure 5, showing a section from the context given in (51), in which a nuclear pitch accent occurs on the phrase ersten ministers (‘first minister’). das amt des ersten ministers soll paisley übernehmen l*h h* l*h % h*l % 50 200 100 0 2.70.6 1.2 1.8 time (s) fu nd am en ta l f re qu en cy (h z) figure 5: optional contrast; dirndl s1897, 26-03-2007, 20:00, 04’05”: first minister (51) in nordirland soll am 8. mai eine gemeinsame regierung aus protestantischen unionisten und katholischer sinn féin die arbeit aufnehmen. darauf verständigten sich die parteivorsitzenden paisley und adams bei ihrem ersten persönlichen treffen in belfast. ∼[alt-implicit das amt des [ersten ministers]f %] soll paisley übernehmen. in northern ireland, a power-sharing government of the protestant unionists and the catholic sinn féin will start its work on may 8. this was arranged by the party leaders paisley and adams at their first personal meeting in belfast. paisley will become the first minister. knowing that the nuclear pitch accent is on minister intuitively evokes the contrastive interpretation that the speaker assumed that other posts were also assigned. this is what we indicate by the ∼ operator. clearly, the version in figure 5 is not the only way the sentence can be pronounced. instead, the speaker might have chosen not to mark the contrast, thus simply producing a novelty focus as in (52). (52) [das the amt post des of ersten first ministers minister soll shall paisley paisley übernehmen.]f take there are many conceivable ways of evaluating the proposals made in this section. for example, it would have been possible to have annotators detect alternative-eliciting features. however, as a first step we asked annotators to tell apart contrastive from non-contrastive focus. to this end, we reverted to a bottom-up process: we selected 3842 nuclear pitch accents from the dirndl corpus (only the final pitch accents of an intonation phrase), disregarding their accent type (rise, fall etc.). we made the assumption that all of them mark some sort of focus. the question we were interested in was whether these foci would be contrastive or merely new information. two independent student 236 focus triggers and focus types annotators were provided with the accented words as well as the containing sentence plus context. their task was to say whether they could identify against what an accented word was contrasted in the given context. for validation purposes, they were asked to note down the explicit or implicit alternative. finally, they had to label the focus as either non-contrastive (no-contrast), implicitly contrastive (alt-implicit), or marked by an alternative-eliciting feature from table 2. as a result, we obtain a κ-value of 0.56 for the eight categories. the score slightly improves to κ = 0.58 if we only consider the three categories non-contrastive, implicitly contrastive, and marked by an explicit alt-feature. nevertheless, determining contrastive focus turned out to be a rather difficult task. in a final consensus annotation we classify the nuclear pitch accents as marking the focus types shown in table 3. note that the table does not allow for any conclusions with regard to the total ratio of contrastive versus non-contrastive foci in the complete data, since it leaves aside foci marked by nuclear pitch accents of non-final intermediate phrases as well as preand postnuclear prominences. label n alt-fsp 45 alt-overt-arg 315 alt-overt-comp 12 alt-overt-coord 484 alt-overt-ext 234 alt-sel 36 alt-implicit 479 no-contrast 2237 table 3: (non-)contrastive focus on 3842 (intonation-phrase final) nuclear pitch accents 5. second occurrence focus and secondary accents in the remaining part of this article we return to the issue of second occurrence focus (sof), already discussed in section 1. describing the precise conditions which license second occurrence focus is not straightforward. in particular, sof is not sufficiently characterised by defining it as a (contrastively) focussed and given constituent. a counterexample is shown in (53b) (büring, 2008, ms.) in which the second occurrence of john is realised as a primary – not secondary – focus. (53) a. many people only drank juicef at john’s party. b. even johnf only drank juicesof at his party. beaver and velleman (2011) discuss at length the licensing conditions for second occurrence focus. in particular, they reject proposals which rely on the size (or mutual embedding) of the focus domains (büring, 2008, ms.; rooth, 2010) or which redefine the notion of givenness for focussed constituents (selkirk, 2008). we will not repeat their argumentation here. the characterization of an sof given by beaver and velleman (2011) is a constituent which is “important” (in our terminology alt-marked) as well as “predictable”, where predictability is defined as “givenness-all-the-way-up” – a constituent is predictable if and only if it is given and its containing phrasal context is given as well. from this it follows that an unpredictable constituent must either be new or occur in a new 237 riester and baumann argument slot. in (53b), the given phrase juice occurs in a likewise given vp – it is predictable – while the given expression john newly occupies the agent role, and is therefore unpredictable. in (54) and (55) we provide a complete reflex and elicited-alternatives annotation of the sof example (6), which was introduced in section 1.30 we observe that the second occurrence of vegetables in (55) is embedded under other given constituents. (54) everyone knew that mary only eats vegetables. (np ) l-new (v ) l-new (np ) l-new (v ) l-new (dp ) r-unused (dp ) r-generic (v p ) l-new (s) l-new (cp ) r-unused (v p ) l-new (s) l-new alt-fsp (55) even paul knew that mary only eats vegetables. (np ) l-new (np ) l-given (v ) l-given (np ) l-given (v ) l-given (dp ) r-given (dp ) r-generic (dp ) r-unused (v p ) l-given (v p ) l-given (cp ) r-given (v p ) l-given (s) l-given alt-fsp alt-fsp the reader is encouraged to verify that every constituent marked l-given is entailed by the context (in this case, simply repeated), while r-given phrases are coreferential. we shall assume that the that-cp behaves similar to a definite dp while referring to a fact. the last issue which we are going to discuss are other types of secondary focus which are not second occurrence foci. consider that, instead of (55), the speaker would have chosen to say (56). (56) even paul knew that mary is picky. (np ) l-given (ap ) l-new (dp ) r-given (s) l-new (cp ) r-given the utterance in (56) is clearly subjective. the speaker does not merely tell us that even paul knew that mary is a vegetarian but she also lets us know, en passant, that she considers vegetarians picky. we only get this interpretation because the entire cp does not carry a nuclear pitch accent but 30. the definite subject everyone is left unannotated here because, strictly speaking, we are not dealing with an out-ofthe-blue utterance but a constructed example which presupposes some context. in normal discourse, speakers would limit the use of everyone to contexts in which it is clear which group is being referred to. in other words, the domain of the universal quantifier is restricted to an identifiable set in the context. therefore, the most appropriate label for everyone is probably r-given, which might strike some readers as odd in the absence of any explicit context. 238 focus triggers and focus types instead, following our intuitions, comes with a compressed intonation contour. it is only because of this postnuclear compression that we obtain the interpretation that (the speaker thinks that) the two cps boil down to the same thing.31 the second cp is therefore labelled as referentially given. the question whether the cp in (56) is deaccented or whether it carries postnuclear prominence cannot be solved by introspection and therefore requires experimental research. telling from the reflex annotations, however, we are able to state that the embedded clause contains lexically new material, and we might furthermore argue that the new information contained in the clause is not-at-issue – compare section 2 – and gives an answer to the speaker-oriented question what is mary like, according to the speaker? we think that this is sufficient reason to grant the embedded predicate the status of a novelty focus, which, when indeed marked by postnuclear prominence would count as a secondary focus. in analogy to the case in (56), we can also construct examples involving a coreferential dp consisting of lexically new material. such expressions have sometimes been called epithets in the literature (clark, 1977; schlenker, 2005; potts, 2005; riester, 2009). expressives, as in (57b), are one kind of epithets. again they represent not-at-issue content and will be considered novelty foci. (implicit question: what does the speaker think about fred?) (57) a: do you know where fred is? b: i haven’t seen the goddam idiot. (np ) l-new (dp ) r-given in our corpus, we do find instances of epithets which are marked by postnuclear prominence (mainly indicated by increased duration and intensity of lexically stressed syllables). figure 6 shows an example of the r-given, l-new dp der serbischen provinz (‘of the serbian province’), which corefers with the phrase kosovo in the context shown in (58). (58) der uno-sondergesandte ahtisaari plädiert für eine unabhängigkeit des kosovo unter internationaler aufsicht. dies sei die einzige politische und wirtschaftliche option für die zukunft der serbischen provinz. un special envoy ahtisaari is making the case for an independence of the kosovo under international control. according to him, this is the only political and economic option for the future of the serbian province. so far, this is merely anecdotal evidence, and more elaborate statistical investigations are necessary. we would like to point out, however, that baumann and riester (in press) did find – mainly in a corpus of read speech – that expressions which possess a “hybrid” information status (e.g. r-given, l-new), such as epithets, were marked as significantly less prominent than fully new items and significantly more prominent than fully given items. only some of them were encoded by postnuclear prominences, others by prenuclear accents, or less prominent (e.g. low) accent types. actually, several kinds of secondary prominence have been proposed in the literature, e.g. duration 31. this is a behaviour well-known from definite dps. compare example (i) from umbach (2002): (i) {john has an old cottage.} a. last summer he reconstructed [the shed]. (non-coreferential) b. last summer he reconstructed [the shed]. (coreferential) 239 riester and baumann für die zukunft der serbischen provinz h*l secondary prominence % l-new r-given 50 280 100 200 0 1.90.4 0.8 1.2 time (s) fu nd am en ta l f re qu en cy (h z) figure 6: realisation of an epithet (r-given, l-new); dirndl s1730, 26-03-2007, 17:00, 0’31”: the serbian province accents (kohler, 2005), ornamental accents (büring, 2007), or phrase accents (grice et al., 2000), mentioned in section 1 above. however, the various concepts refer to quite different phenomena or levels: first, the presence or absence of a pitch movement, i.e. tonal vs. non-tonal prominence (e.g. kohler’s duration accents are non-tonal); second, accent position or the status of a prominence in the prosodic hierarchy (e.g. ornamental accents are prenuclear, phrase accents postnuclear); and third, accent type, i.e. the form of a pitch movement on a metrically strong syllable (high or rising accents are more prominent than low or falling ones, see e.g. baumann and grice, 2006). while there is a growing body of evidence that these three levels of prosodic prominence are used for marking various aspects of information structure, cf. baumann (2012), it is a matter of some debate whether the current tobi (and gtobi) systems are suited to represent the relevant distinct types of information structural meaning – although at least the distiction between several types of pitch accent has always been a central incentive for the definition of the tobi labels, see for instance pierrehumbert and hirschberg (1990); steedman (1991). 6. conclusions in this article, we have discussed a number of problems which arise when issues from the theoretical and experimental tradition of information structure are brought together with corpus data, such as spoken radio news. in parts of the literature, focus is discussed as a semantic-pragmatic phenomenon which is related to the answering of questions under discussion. since the latter are mostly implicit in monological discourse, it seems, however, that their identification is almost as subtle as the identification of focus itself. corpus data can open our eyes to a number of issues that have often been ignored in research on focus, for instance the fact that spoken utterances contain prenuclear and nuclear pitch accents as well as various kinds of postnuclear prominence. 240 focus triggers and focus types on the meaning side, this corresponds to the fact that the informative parts of a sentence have an internal structure that can be revealed by asking (nested) questions. as a result, we can have several semantic-pragmatic foci within one sentence, which have different degrees of communicative significance and which are marked by different kinds of prosodic prominence. information structure has a long history of theory building which has produced highly complex models, sometimes based on relatively thin empirical evidence. the precise mapping between the semantic-pragmatic and the prosodic phenomena is still not sufficiently described. (the same can be said about morpho-syntactic focus marking in cross-linguistic perspective.) we are convinced that building and analysing annotated corpus resources can complement experimental and theoretical research in this regard. a big problem in focus theory is its non-standardised terminology. are there one or several kinds of focus, and is new information one of them? how should givenness and contrast be defined? is contrastive topic a kind of focus? what is the precise definition of the f-feature and the∼ operator? what is a focus domain? researchers have provided diverging answers to these problems, including the authors of this paper. in principle, nothing speaks against different terminological choices as long as they are made transparent. on the other hand, it is plain that certain terminological and classificatory decisions can blur rather than clarify insights into linguistic phenomena and may cause problems in view of larger theoretical structures. linguistic annotation is an important testbed for both theory and terminology: the ability to annotate a certain phenomenon is support for its underlying theoretical conceptualisation. in this article we have provided two proposals with respect to the annotation of information structure: (i) the reflex scheme, a framework for the annotation of information status, divided into a referential and a lexical level, and (ii) a set of alternative-eliciting features, which represent the basis for a speaker’s decision to mark contrastive focus. we furthermore postulate that plainly new information represents the basic type of focus. while both types of focus involve alternatives, the defining criterion for contrastive focus is the addressee’s ability to identify (and name) at least one of these alternatives, in the respective context. while the annotation of information status can (and should) be accomplished without access to prosody, the annotation of contrastive focus in english and german will ultimately require the use of spoken information since it can only partially be determined on the basis of written data alone. finally, we discussed the phenomenon of second occurrence focus in the light of general assumptions about prosodic structure and the pragmatics of focus, and concluded that it is not the only kind of focus which is realised by means of secondary prominence. very likely, there are other focal phenomena, like not-at-issue content, or expressions with a hybrid information status which are sometimes marked by means of prenuclear or postnuclear prominence. 241 riester and baumann acknowledgements the authors would like to thank the editors of this volume stefanie dipper, bonnie webber and heike zinsmeister for their support. we thank the three anonymous reviewers for their comprehensive and thoughtful comments. thanks to jutta hartmann, fabienne martin and carla umbach for discussions, as well as maria drixler and lisa kienzle for their corpus annotation work. we would also like to thank kordula de kuthy and detmar meurers for testing our work in their 2012/13 class on information structure annotation at tübingen. thanks go as well to the organisers of the dfg research network questions in discourse, edgar onea and malte zimmermann for enabling numerous opportunities for the exchange of thoughts about discourse and information structure. financial support by the german science foundation (dfg) is kindly acknowledged. (project a1 of the stuttgart collaborative research center 732 incremental specification in context, and the project gr 1610/5 degrees of activation and focus-background structure in spontaneous speech in cologne.) references steven paul abney. the english noun phrase in its sentential aspect. phd thesis, mit, cambridge, ma, 1987. david adger. core syntax: a minimalist approach. oxford university press, 2003. kai alter, ina mleinek, tobias rohe, anita steube, and carla umbach. kontrastprosodie in sprachproduktion und -perzeption. linguistische arbeitsberichte, 77:59–79, 2001. in: anita steube and carla umbach, editors, kontrast – lexikalisch, semantisch, intonatorisch. universität leipzig. ron artstein and massimo poesio. inter-coder agreement for computational linguistics. computational linguistics, 34(4):556–596, 2008. nicholas asher and alex lascarides. bridging. journal of semantics, 15:83–113, 1998. jay atlas and stephen levinson. it-clefts, informativeness and logical form. in peter cole, editor, radical pragmatics, pages 1–61. academic press, new york, 1981. christine bartels. acoustic correlates of ‘second occurrence focus’: towards an experimental investigation. in hans kamp and barbara partee, editors, context-dependence in the analysis of linguistic meaning, pages 354–361. elsevier, amsterdam, 2004. stefan baumann. types of secondary prominence and their relation to information structure. talk at satellite workshop of glow 35 on production and perception of prosodically-encoded information structure, university of potsdam, march 27, 2012. stefan baumann and martine grice. the intonation of accessibility. journal of pragmatics, 10 (38):1636–1657, 2006. stefan baumann and arndt riester. referential and lexical givenness: semantic, prosodic and cognitive aspects. in gorka elordieta and pilar prieto, editors, prosody and meaning, volume 25 of interface explorations, pages 119–162. mouton de gruyter, berlin, 2012. 242 focus triggers and focus types stefan baumann and arndt riester. coreference, lexical givenness and prosody in german. lingua, in press. doi: http://dx.doi.org/10.1016/j.lingua.2013.07.012. special issue ‘information structure triggers’ ed. by jutta hartmann, susanne winkler and janina radó. stefan baumann, doris mücke, and johannes becker. expression of second occurrence focus in german. linguistische berichte, 221:61–78, 2010. david beaver and brady clark. sense and sensitivity. how focus determines meaning. wiley & sons, chichester, 2008. david beaver and dan velleman. the communicative significance of primary and secondary accents. lingua, 121:1671–1692, 2011. david beaver, brady clark, edward flemming, florian jaeger, and maria wolters. when semantics meets phonetics: acoustical studies of second occurrence focus. language, 83(2), 2007. mary beckman. stress and non-stress accent. foris, dordrecht, 1986. mary beckman, julia hirschberg, and stefanie shattuck-hufnagel. the original tobi system and the evolution of the tobi framework. in sun-ah jun, editor, prosodic typology – the phonology of intonation and phrasing, pages 9–54. oxford university press, 2005. paul boersma and david weenink. praat – a system for doing phonetics by computer, version 3.4. technical report 132, institute of phonetic sciences of the university of amsterdam, 1996. url http://www.fon.hum.uva.nl/praat/. aljoscha burchardt, katrin erk, anette frank, andrea kowalski, and sebastian padó. salto: a versatile multi-level annotation tool. in proceedings of the fifth international conference on language resources and evaluation (lrec), genoa, italy, 2006. daniel büring. on d-trees, beans, and b-accents. linguistics & philosophy, 26(5):511–545, 2003. daniel büring. semantics, intonation and information structure. in gillian catriona ramchand and charles reiss, editors, the oxford handbook of linguistic interfaces. oxford university press, 2007. daniel büring. what’s new (and what’s given) in the theory of focus? in proceedings of the 34th annual meeting of the berkeley linguistics society, pages 403–424, 2008. daniel büring. givenness, contrast & topic. talk at workshop on prosody and information status in typological perspective, 35th annual meeting of the german linguistic society (dgfs), university of potsdam, march 13, 2013. daniel büring. been there, marked that – a theory of second occurrence focus. 2008, ms. wallace l. chafe. discourse, consciousness, and time. university of chicago press, 1994. noam chomsky. deep structure, surface structure, and semantic interpretation. in danny steinberg and leon jakobovits, editors, semantics. an interdisciplinary reader in philosophy, linguistics and psychology, pages 183–216. cambridge university press, 1971. 243 riester and baumann herbert h. clark. bridging. in philip johnson-laird and peter wason, editors, thinking: readings in cognitive science, pages 169–174. cambridge university press, 1977. jacob cohen. a coefficient of agreement for nominal scales. educational and psychological measurement, 1(20):37–46, 1960. philippa cook and felix bildhauer. identifying ‘aboutness topics’: two annotation experiments. dialogue and discourse, 4(2):118–141, 2013. dick crouch, mary dalrymple, ron kaplan, tracy king, john maxwell, and paula newman. xle documentation. palo alto research center, 1993-2011, ms. url http://www2.parc.com/isl/groups/nltt/xle/doc/xle toc.html. judy delin and jon oberlander. syntactic constraints on discourse structure: the case of it-clefts. linguistics, 33(3):465–500, 1995. kerstin eckart, arndt riester, and katrin schweitzer. a discourse information radio news database for linguistic analysis. in christian chiarcos, sebastian nordhoff, and sebastian hellmann, editors, linked data in linguistics. representing and connecting language data and language metadata, pages 65–76. springer, heidelberg, 2012. miriam eckert and michael strube. dialogue acts, synchronising units and anaphora resolution. journal of semantics, 17(1):51–89, 2000. katalin é. kiss. identificational focus versus information focus. language, 74(2):245–273, 1998. caroline féry and shinichiro ishihara. the phonology of second occurrence focus. journal of linguistics, 45(2):285–313, 2009. michael götze, thomas weskott, cornelia endriss, ines fiedler, stefan hinterwimmer, svetlana petrova, anne schwarz, stavros skopeteas, and ruben stoel. information structure. in stefanie dipper, michael götze, and stavros skopeteas, editors, information structure in cross-linguistic corpora: annotation guidelines for phonology, morphology, syntax, semantics and information structure, volume 7 of interdisciplinary studies on information structure. universitätsverlag potsdam, 2007. sfb 632. martine grice, d. robert ladd, and amalia arvaniti. on the place of phrase accents in intonational phonology. phonology, 17(2):143–185, 2000. jeroen groenendijk and martin stokhof. studies in the semantics of questions and the pragmatics of answers. phd thesis, universiteit van amsterdam, 1984. carlos gussenhoven. focus, mode and nucleus. journal of linguistics, 19:377–417, 1983. carlos gussenhoven. sentence accents and argument structure. in iggy roca, editor, thematic structure: its role in grammar, pages 79–106. foris, berlin, 1992. carlos gussenhoven. on the limits of focus projection in english. in peter bosch and rob van der sandt, editors, focus. linguistic, cognitive and computational perspectives, pages 43– 55. cambridge university press, 1999. 244 focus triggers and focus types eva hajičová, barbara h. partee, and petr sgall. topic-focus articulation, tripartite structures, and semantic content. kluwer, dordrecht, 1998. eva hajičová, jarmila panevová, and petr sgall. a manual for tectogrammatical tagging of the prague dependency treebank. technical report tr-2000-09, charles university, prague, 2000. url http://ufal.mff.cuni.cz/pdt/corpora/pdt 1.0/doc/tmanual/tmanen.pdf. michael halliday. notes on transitivity and theme in english. part 2. journal of linguistics, 3: 199–244, 1967. michael halliday and ruqaiya hasan. cohesion in english. longman, london, 1976. nancy hedberg. the referential status of clefts. phd thesis, university of minnesota, 1990. nancy hedberg and juan m. sosa. the prosody of topic and focus in spontaneous english dialogue. in chungmin lee, matthew gordon, and daniel büring, editors, topic and focus: crosslinguistic perspectives on meaning and intonation, volume 82 of studies in linguistics and philosophy, pages 101–120. springer, dordrecht, 2008. anne hermes, johannes becker, doris mücke, stefan baumann, and martine grice. articulatory gestures and focus marking in german. in proceedings of the 4th conference on speech prosody, campinas, brazil, 2008. tilman höhle. explikationen für ‘normale betonung’ und ‘normale wortstellung’. in werner abraham, editor, satzglieder im deutschen, pages 75–153. narr, tübingen, 1982. julia horvath. ’discourse features’, syntactic displacement and the status of contrast. lingua, pages 1346–1369, 2010. ray jackendoff. semantic interpretation in generative grammar. mit press, cambridge, ma, 1972. jonah katz and elisabeth selkirk. contrastive focus vs. discourse-new: evidence from phonetic prominence in english. language, 87(4):771–816, 2011. istván kenesei. focus as identification. in valéria molnár and susanne winkler, editors, the architecture of focus, studies in generative grammar, pages 137–168. mouton de gruyter, berlin, 2006. klaus j. kohler. form and function of non-pitch accents. arbeitsberichte des instituts für phonetik und digitale sprachverarbeitung der universität kiel (aipuk), 35a:97–123, 2005. prosodic patterns of german spontaneous speech. manfred krifka. association with focus phrases. in valéria molnár and susanne winkler, editors, the architecture of focus, studies in generative grammar. mouton de gruyter, berlin, 2006. manfred krifka. basic notions of information structure. acta linguistica hungarica, 55(3-4): 243–276, 2008. 245 riester and baumann manfred krifka. a compositional semantics for multiple focus constructions. in joachim jacobs, editor, informationsstruktur und grammatik, pages 17–53. westdeutscher verlag, opladen, 1992. also in proceedings of salt 1. cornell working papers in linguistics 10. 1991. knud lambrecht. information structure and sentence form. cambridge university press, 1994. ewald lang and carla umbach. kontrast in der grammatik: spezifische realisierungen und übergreifender konnex. linguistische arbeitsberichte, 79:145–186, 2002. sebastian löbner. definite associative anaphora. in proceedings of the discourse anaphora and resolution colloquium (daarc), lancaster, 1998. jörg mayer. transcription of german intonation. the stuttgart system. ms., 1995. url http://www.ims.uni-stuttgart.de/phonetik/joerg/labman/stgtsystem.html. malvina nissim, shipra dingare, jean carletta, and mark steedman. an annotation scheme for information status in dialogue. in proceedings of the fourth international conference on language resources and evaluation (lrec), lisbon, 2004. patrizia paggio. annotating information structure in a corpus of spoken danish. in proceedings of the fifth international conference on language resources and evaluation (lrec), pages 1606–1609, genoa, italy, 2006. barbara h. partee. focus, quantification and semantics-pragmatics issues. in p. bosch and r. van der sandt, editors, focus: linguistic, cognitive, and computational perspectives, pages 213–231. cambridge university press, 1999. janet pierrehumbert and julia hirschberg. the meaning of intonational contours in the interpretation of discourse. in philip cohen, jerry morgan, and martha pollack, editors, intentions in communication, chapter 14. mit press, 1990. massimo poesio and renata vieira. a corpus-based investigation of definite description use. computational linguistics, 24(2):183–216, 1998. christopher potts. the logic of conventional implicatures. oxford university press, 2005. ellen f. prince. toward a taxonomy of given-new information. in peter cole, editor, radical pragmatics, pages 233–255. academic press, new york, 1981. ellen f. prince. the zpg letter: subjects, definiteness and information status. in william c. mann and sandra a. thompson, editors, discourse description: diverse linguistic analyses of a fund-raising text, pages 295–325. benjamins, amsterdam, 1992. matthew reeve. the syntactic structure of english clefts. lingua, 121:142–171, 2011. sophie repp. defining ‘contrast’ as an information-structural notion in grammar. lingua, 120: 1333–1345, 2010. arndt riester. partial accommodation and activation in definites. in proceedings of the 18th international congress of linguists, pages 134–152, seoul, 2009. linguistic society of korea. 246 focus triggers and focus types arndt riester and hans kamp. squiggly issues: alternative sets, complex dps, and intensionality. in maria aloni et al., editors, logic, language and meaning. revised selected papers from the 17th amsterdam colloquium, pages 374–383. springer, berlin, 2010. arndt riester, david lorenz, and nina seemann. a recursive annotation scheme for referential information status. in proceedings of the seventh international conference on language resources and evaluation (lrec), pages 717–722, valletta, malta, 2010. craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. osu working papers in linguistics, 49, 1996. christian rohrer and martin forst. improving coverage and parsing quality of a large-scale lfg for german. in proceedings of the fifth international conference on language resources and evaluation (lrec), genova, 2006. mats rooth. second occurrence focus and relativized stress f. in caroline féry and malte zimmermann, editors, information structure: theoretical, typological, and experimental approaches. oxford university press, 2010. mats rooth. a theory of focus interpretation. natural language semantics, 1(1):75–116, 1992. mats rooth. on the interface principles for intonational focus. in teresa galloway and justin spence, editors, proceedings of semantics and linguistic theory (salt) 6, pages 202–226, ithaca, ny, 1996. clc publications. phillippe schlenker. minimize restrictors! (notes on definite descriptions, condition c and epithets). in emar maier, corien bary, and janneke huitink, editors, proceedings of sinn und bedeutung ix, pages 385–416, nijmegen, 2005. katrin schulz and robert van rooij. pragmatic meaning and non-monotonic reasoning: the case of exhaustive interpretation. linguistics and philosophy, 29(2):205–250, 2006. roger schwarzschild. givenness, avoidf, and other constraints on the placement of accent. natural language semantics, 7(2):141–177, 1999. elisabeth selkirk. contrastive focus vs. presentational focus: prosodic evidence from right node raising in english. in proceedings of the 1st international conference on speech prosody, pages 643–646, aix-en-provence, 2002. elisabeth selkirk. contrastive focus, givenness and the unmarked status of ‘discourse-new’. acta linguistica hungarica, 55(3–4):331–346, 2008. elisabeth selkirk. phonology and syntax. mit press, cambridge, 1984. elisabeth selkirk. sentence prosody. intonation, stress, and phrasing. in j. goldsmith, editor, the handbook of phonological theory, pages 550–569. blackwell, 1995. mandy simons, judith tonhauser, david beaver, and craige roberts. what projects and why. in proceedings of salt 20, pages 309–327, vancouver, 2010. 247 riester and baumann benjamin spector. scalar implicatures: exhaustivity and gricean reasoning. in maria aloni, alistair butler, and paul dekker, editors, questions in dynamic semantics, pages 229–254. elsevier, amsterdam, 2006. mark steedman. structure and intonation. language, 67(2):207–264, 1991. anna szabolcsi. the semantics of topic-focus articulation. in jeroen groenendijk, theo janssen, and martin stokhof, editors, formal methods in the study of language. part 2, volume 136 of mathematical centre tracts, pages 513–540. mathematisch centrum, amsterdam, 1981. carla umbach. (de)accenting definite descriptions. theoretical linguistics, 27(2/3), 2002. carla umbach. on the notion of contrast in information structure and discourse structure. journal of semantics, 21(2), 2004. michael wagner. givenness and locality. in proceedings of salt xvi, pages 295–312, ithaca, ny, 2006. 248 dialogue and discourse 4(2) (2013) 53-64 doi: 10.5087/dad.2013.203 a uniform syntax and discourse structure: the copenhagen dependency treebanks daniel hardt dh.itm@cbs.dk department of it management copenhagen business school copenhagen, denmark editors: stefanie dipper, heike zinsmeister, bonnie webber abstract i present arguments in favor of the uniformity hypothesis: the hypothesis that discourse can extend syntax dependencies without conflicting with them. i consider arguments that uniformity is violated in certain cases involving quotation, and i argue that the cases presented in the literature are in fact completely consistent with uniformity. i report on an analysis of all examples in the copenhagen dependency treebanks (cdt) involving violations of uniformity. i argue that they are in fact all consistent with uniformity, and conclude that the cdt should be revised to reflect this. keywords: discourse, syntax, treebank 1. introduction the copenhagen dependency treebanks (cdt) are unusual in that they contain annotation of both syntactic and discourse structure, using a single dependency graph (buch-kromann and korzen, 2010). underlying this approach is a strong hypothesis about the relation of syntax and discourse: namely, that they are subject to uniform well-formedness conditions. in other words, discourse structure should extend the syntax graph without conflicting with it. other discourse annotation projects have not followed this approach: for example, the penn discourse treebank has been annotated independently of the syntactic annotation of the same texts. in this paper, i present arguments in favor of the hypothesis that discourse can extend syntax dependencies without conflicting with them. i call this the uniformity hypothesis. while the design of the cdt was motivated by the uniformity hypothesis, the actual annotation practice has been more flexible, allowing a mechanism for annotating discourse relations in a way that allows for violations of uniformity. this flexibility was introduced in response to arguments in the literature against uniformity; in particular that of (dinesh et al., 2005). here it is argued that contrast relations sometimes conflict with syntactic relations when quotation is involved. i argue that the examples of interest here are in fact consistent with uniformity, once the semantics of the contrast relation is considered in more detail. this raises the question of whether the flexibility in cdt annotation is ever needed, or whether its annotation could be made completely consistent with uniformity. to answer this question, i report on a systematic analysis of the examples in cdt that have been annotated in violation of uniformity. there are three discourse relations having at least 10 occurrences of such violations. i c©2013 daniel hardt submitted 03/12; accepted 12/12; published online 04/13 hardt argue that, in each case, the annotation can be replaced with one that does not violate uniformity. based on this, i argue that the cdt should be revised to reflect this. in what follows, i begin with some background on the cdt, focusing on the annotation of syntax and discourse. next, i consider the argument of (dinesh et al., 2005) against uniformity, arguing that the relevant examples are in fact consistent with uniformity. i then turn to an empirical analysis of the examples in cdt that have been annotated to indicate a violation of uniformity. i argue that all these cases are in fact best analyzed in a way that is consistent with uniformity, and conclude that the cdt annotations should be modified to conform with uniformity in all cases. 1.1 background the copenhagen dependency treebanks, cdt, consist of five parallel open-source treebanks for danish, english, german, italian, and spanish. the treebanks are being annotated manually with respect to syntax, discourse, anaphora, morphology, as well as translational equivalence (word alignment) between the danish source text and the target texts in the four other languages. at this point it is primarily the danish treebank that has been annotated for both discourse and syntax, so in this paper i will restrict attention to this treebank. 1.2 syntax the syntactic annotation of the cdt treebanks is based on the linguistic principles outlined in the dependency theory discontinuous grammar (buch-kromann, 2006) and the syntactic annotation principles described in (kromann, 2003), (buch-kromann et al., 2007), and (buch-kromann et al., 2009). all linguistic relations are represented as directed labelled relations between words or morphemes. the model operates with a primary dependency tree structure in which each word or morpheme is assumed to act as a complement or adjunct to another word or morpheme, called the governor (or head), except for the top node of the clause or unit, typically the finite verb. 1.3 discourse just as sentence structure can be captured by dependencies that link up the words and morphemes within a sentence, discourse structure can also be captured by dependencies that link up the words within an entire discourse. the cdt discourse annotation consists in linking up each clause’s top node with its nucleus (understood as the unique word within another clause that is deemed to govern the relation) and labelling the relations between the two nodes. the inventory of discourse relations in cdt is described in the cdt manual. it borrows heavily from other discourse frameworks, in particular rhetorical structure theory, rst (mann and thompson, 1987; taboada and mann, 2006; carlson, lynn and marcu, daniel and okurowski, mary ellen, 2001) and the penn discourse treebank, pdtb (webber, 2004; dinesh et al., 2005; prasad et al., 2007, 2008), as well as (korzen, 2006, 2007), although the inventory had to be extended to accommodate the great variety of text types in the cdt corpus. the inventory allows relation names to be formed as disjunctions or conjunctions of simple relation names, to specify multiple relations or ambiguous alternatives. one of the most important differences between the cdt framework and other discourse frameworks lies in the way texts are segmented. in particular, cdt uses words as the basic building blocks in the discourse structure, while most other discourse frameworks use clauses as their atomic discourse 54 syntax and discourse units, including rst, pdtb, graphbank (wolf and gibson, 2005), and the potsdam commentary corpus, pcc (stede, 2004). 2. uniformity of syntax and discourse according to the hypothesis of uniformity, the same well-formedness conditions apply to the discourse structure as to syntactic structure. in a very real sense cdt does not distinguish between syntax and discourse, since both discourse and syntax links are part of the same graph. i will focus on one particular aspect of uniformity, which is easier to appreciate in terms of relations on bracketed structures. consider the following structures s1 and s2 s1= [. . . [a] . . .] s2= [. . . [b] . . .] given these structures, we allow a relation between s1 and s2, but we do not allow crossing relations involving embedded elements a or b: such as a link from a to b, from s1 to b, or from a to s2. consider the following constructed example (1) john said it was raining. but it was not raining. here, we have s1= [john said [it was raining]] s2= [but it was not raining]] where a (embedded within s1) is it was raining. one might be tempted to define a contrast relation between a and s2. but uniformity requires that the relation be between s1 and s2. 3. quotation: an apparent violation of uniformity dinesh et al. (2005) argue that just such violations of uniformity can be observed in cases involving quotation. consider example 2: (2) the current distribution arrangement ends in march 1990, although delmed said it will continue to provide some supplies of the peritoneal dialysis products to national medical, the spokeswoman said. [(12) in (dinesh et al., 2005)] s1= [the current distribution arrangement ends in march 1990] s2= [delmed says [it will continue to provide some supplies of the peritoneal dialysis products to national medical]...] here there is an embedded element in s2, which i call b: it will continue to provide some supplies of the peritoneal dialysis products to national medical. (i ignore the final attribution to the spokeswoman.) dinesh et al. argue that the discourse relation of contrast, signalled by “although”, does not hold between s1 and s2, but between s1 and the embedded element b. according to dinesh et al.: “although as a discourse connective denies the expectation that the supply of dialysis products will be discontinued when the distribution arrangement ends. it does not convey the expectation that delmed will not say such things”. to evaluate this argument, we must look at the semantics of contrast and quotation. 55 hardt 4. semantics of contrast dinesh et al. (2005) are assuming that contrast typically involves a “denial of expectation”. this is a standard view of contrast, and the penn discourse treebank annotation manual (prasad et al., 2007) defines concession as a subtype of contrast, characterized by “denial of expectation”, stating this: the type concession applies when the connective indicates that one of the arguments describes a situation a which causes c, while the other asserts (or implies) not c. the rst annotation manual (carlson, lynn and marcu, daniel, 2001, p. 50) says this about concession as a type of contrast: . . . a concession relation is always characterized by a violated expectation (p 50) but what exactly is meant by a “violated expectation”? in an influential early discussion, (hobbs, 1985, p. 22), defines “violated expectation” as follows: infer p from the assertion of s0 and ¬ p from the assertion of s1. hobbs illustrates this with the following example: (3) john is a lawyer, but he is honest. from s1 john is a lawyer, hobbs argues, one can infer p = john is dishonest, while s2 is not p (john is honest). from this discussion, it is clear that contrast between s1 and s2 normally involves a contradiction – s1 implies some p and s2 implies not p. at the same time, a normal, felicitous discourse must be logically consistent. what this means is that the contradiction in a contrast must always be safely “packaged” to avoid an inconsistent discourse. a standard way of achieving this is that inferences from s0 to p and from s1 to not p are based on different background assumptions, which i will call back1 and back2. a felicitous discourse does not require that one be committed to the truth of back1 and back2, but, rather, one must be willing to temporarily entertain them. we now return to example 2. recall that s1 = [current distribution arrangement ends], and s2 = [delmed says some supplies will continue], with b = [some supplies will continue]. dinesh et al. argue for contrast between s1 and b: i will call this case 1. uniformity would dictate contrast between s1 and s2, which i will call case 2. following hobbs’ analysis, we want to identify some p that we can infer from s1, such that not p can be inferred from b (in case 1) or from s2 (in case 2). it is crucial that these two inferences rest on different background assumptions; otherwise the discourse would be inconsistent. furthermore, a discourse is only felicitous if the background assumptions are salient and one is willing to entertain the possibility that they are true. i call the background assumptions underlying the inference to p and not p back1 and back2, respectively. in case 1, the background assumption for inferring p, back1, is this: supplies only come from current distribution arrangement. together with s1, current distribution ends, one can infer p: no 56 syntax and discourse supplies will continue. the background assumption for inferring not p, back2, is empty, since not p is identical to b. thus, i agree with dinesh et al. that contrast between s1 and b is coherent. however, case 2 also supports contrast. here, we have s2 delmed says some supplies will continue instead of b some supplies will continue. instead of an empty background assumption back2, we have delmed speaks truthfully. from back2 and s2 one can infer not p. the two cases are shown below: case 1: contrast s1,b • back1: supplies only come from current distribution arrangement • s1: current distribution ends • p: no supplies will continue • back2: • b: some supplies will continue • not p: some supplies will continue case 2: contrast s1,s2 • back1: supplies only come from current distribution arrangement • s1: current distribution ends • p: no supplies will continue • back2: delmed speaks truthfully • s2: delmed says some supplies will continue • not p: some supplies will continue i have argued that example 2 is in fact consistent with the uniformity hypothesis – a discourse relation of contrast between the top-level constituents s1 and s2 is consistent with the notion that contrast always requires conflicting material to be properly “packaged”, and quotations are one typical means for doing so. now, while the annotation approach in the cdt was originally motivated in part by the uniformity hypothesis, the actual annotation policy has been more flexible, incorporating a special star notation for the express purpose of indicating the kinds of violations of uniformity being discussed here. i have attempted to show that this argument is not convincing. there could of course be other cases in which uniformity must be violated. indeed, there are 135 examples in the cdt in which the star notation has been used to indicate violations of uniformity. i now turn to an examination of these examples. 57 hardt relation total occurrences occurrences with * percentage conjunction 1938 76 3.921 contr 195 15 7.692 agentive 156 11 7.692 conc 107 7 6.542 const 109 7 6.422 formal 91 6 6.593 telic 150 6 4.000 time 64 3 4.687 expr 17 2 11.76 quest 48 2 4.166 table 1: occurrences of star notation in cdt 5. an empirical analysis of uniformity in cdt 5.1 use of the star notation in cdt in cdt a relation can be written with a * either to the left or right (or both), to indicate that the left or right argument is embedded. consider relation r linking s0 and s1, where s0= [ . . . [a] . . . ] and s1= [ . . . [b] . . . ] if we have *r[s0,s1], this is a way of indicating r[a,s1], while r*[s0,s1] indicates r[s0,b]. in all cases i have observed, the embedded element referred to with the star notation is quoted material. by quoted material, i mean to include both indirect and direct quotes, including sentential complements not only of verbs of saying but also verbs of belief. thus i include john said “it is raining”, john said that it is raining and john thinks that it is raining. in all three cases i call it is raining the quoted material. table 1 gives the relations with which the star notation occurs. below i consider in detail the relations contrast, agentive, and conjunction. the remaining relations have very small numbers of occurrences. i begin with a detailed consideration of contrast, since it relates directly to the argument discussed above, which motivated the star notation. 5.2 contrast i begin with contrast. for each example, there are two top-level clauses, s1 and s2, and in each case the “*” notation has been used to indicate a contrast relation where one of the arguments is embedded. either s1 or s2 (or both) contains an embedded element, which is the complement of a saying verb. this gives rise to three possibilities. table 2 gives the distribution of the contrast examples with respect to these three possible categories. for category 1, recall example 2. there, the argument was that contrast was valid between s1 and s2, based on a background assumption that delmed speaks truthfully. this argument can be made for all three categories. assume that a = s1 or is embedded within s1, and b = s2 or is embedded within s2. in each case, the embedded element is the complement of “x says”. if there 58 syntax and discourse category s1 s2 occurrences (file id’s) count 1 x says a b 688, 1182, 840, 465, 729, 138 6 2 a x says b 0005, 0366, 0654, 0986, 0683, 1251, 1014, 1214, 1527 9 3 x says a x says b 0 table 2: occurrences of star notation with contrast is a contrast relation between a and b, then, under the assumption that x is truthful, there must also be a contrast relation between s1 and s2.1 below i consider an example of each category. there are six examples in cdt in category 1 “x says a; b”. consider this example (file 0688): s1: [administrerende direktør peter christoffersen siger, at [der hverken er forhandlinger eller sonderinger mellem baltica og skandia i øjeblikket].] [ceo peter christoffersen says that [there are neither negotiations or explorations between baltica and skandia at the moment.].] s2: [skandia har tidligere ønsket et giftermål med baltica.] [skandia had earlier wanted an alliance with baltica.] simplifying a bit, we have a = there are no negotiations and s2 = skandia had wanted an alliance. the reasoning is directly parallel to that of example 2: there is a “violated expectation” between a and s2 – in this case s2 (that an alliance was desired) sets up an expectation that there would be negotiations, and a violates that expectation. this is the contrast relation annotated using the star notation, indicating a relation between the embedded a with the top-level s2. however, under the assumption that the speaker, peter christoffersen, is truthful, then s1 can be inferred to violate the expectation just as a does. recall that category 2 is “a; x says b”. we have this example (file 0986): s1= [de europæiske stålfabrikker befinder sig midt i den værste nedgang i ti år, uden udsigt til forbedring i år. ] [the european steel factories find themselves in the midsts of the worst downturn in ten years, with no prospects of improvement this year.] s2= [[men det er tid til at købe aktier i stålindustrien], siger erhvervsanalytikere.] [[but this is the time to buy stock in the steel industry], say business analysts.] the annotator found a contrast between s1 and b; since s1 describes a downturn in the steel industry, one could infer that this is not the time to buy stock in steel, while b expresses the opposite. 1. as an anonymous reviewer points out, contrast does not always involve a denial of expectation; it can also involve a “juxtaposition of viewpoints” where two or more arguments of a given relation differ. it is worth mentioning that the penn discourse treebank annotates two subtypes of contrast, called juxtaposition and opposition. the logic of my argument concerning contrast has been limited to cases involving denial of expectation. it is possible that this argument would not apply to examples involving these other types of contrast, which would then support the argument that uniformity cannot be maintained. no such cases of contrast were found in the cdt however. 59 hardt category s1 s2 occurrences (file id’s) count 1 x says a b 0366, 1173, 0538, 1035 4 2 a x says b 0581, 0705, 0465, 1173, 0065, 1259 6 3 x says a x says b 0001 1 table 3: occurrences of star notation with agentive under the assumption that what business analysts say is true, there is also a contrast between s1 and s2. 5.3 agentive (cause/reason) the discourse relation agentive in cdt is meant to indicate that one clause expresses a cause or reason for another clause. table 3 gives the occurrences of the star notation in connective with agentive. i begin with an example from category 1 “x says a; b” (file 1173): s1 = [[“jeg respekterer virkelig orlando,”] siger michela buscemi. ] [[“i really respect orlando,”] says michela buscemi. ] s1a= [to af hendes brødre er blevet dræbt af mafiaen, og hun vidnede i retten mod de mistænkte drabsmænd. ] [two of her brothers were killed by the mafia, and she testified in court against the suspected killers.] s2= [“han var den eneste, der rakte en hånd frem for at hjælpe mig.”] [“he was the only one who came forward to help me.”] here we have a = “i really respect orlando”, and b = “he was the only one who came forward to help me”. (note that the agentive relation skips over the intervening sentence which i call here s1a.) clearly in the reported discourse, the speaker michela buscemi is offering b as reason for a. however, in the discourse of the text, i argue that the writer is offering s2 as a reason for s1: that is, the fact that the speaker says she respects orlando is explained by the fact that the speaker has a belief about him, namely that he helped her. it is perfectly coherent to see b (orlando helped buscemi) as a reason for a (buscemi respects orlando), but it is equally coherent to see s2 (buscemi says orlando helped her) as a reason for a. this latter view, i argue, is the correct view of the text, and this is what uniformity would suggest. more generally, the reasoning is parallel to that with contrast: given that x is truthful, then if there is an agentive relation between s1 and b there is also an agentive relation between s1 and s2. category 2 “a; x says b” (file 1259): s1= [og i hvert fald kan vi ikke måle kvaliteten af vore dages barndom ud fra de normer, der gjaldt dengang, vi selv var børn.] [and in any case, we can’t measure the quality of today’s childhood based on the norms that were relevant when we were children.] 60 syntax and discourse category s1 s2 count 1 x says a b 32 2 a x says b 30 3 x says a x says b 12 table 4: occurrences of star notation with conjunction s2= [[”børns vilkår har ændret sig så meget de seneste år, at vi faktisk ikke har noget sammenligningsgrundlag,”] siger han.] [[“the conditions of children have changed so much in recent years, that we actually don’t have a basis for comparison,”] he said.] here we have s1 = and in any case, we can’t measure the quality . . . , and b = “the conditions of children have changed so much . . . ”. it may well be that the quoted speaker is offering b as a reason for s1. this is presumably why the annotator used the right-star notation to indicate b as the second argument for the agentive relation. but it is equally reasonable to argue that the writer is offering s2 as a reason for s1. just as in previous examples: if the speaker is willing to entertain the assumption that x is truthful, then the fact that b is a reason for a means that x says b is a reason for a. category 3 x says a; x says b s1= [de hævder, at [ruslands vej til demokrati går gennem diktatur.]] [they claim that [russia’s path to democracy goes through dictatorship.]] s2= [ i en af deres artikler hedder det: [ ”i et autoritært regime lagdel samfundet og forskellige interesser modnes.]] [in one of their articles, it is stated: [“in an authoritarian regime, society is stratified and different interests matured.”]] here the annotator used both a left and right star, indicating an agentive relation between a and b. the statement in b, in an authoritarian regime, . . . provides an explanation for a, that russia’s path to democracy goes through dictatorship. but this also supports an agentive relation between s1 and s2: the author of the text is offering s2 as a reason for s1. that is, the reason the speakers (they) make the claim a, is that they have the beliefs b. 5.4 conjunction table 4 shows the distribution of star-notated conjunction relations with respect to the different categories. because of the large number of conjunction occurrences with the star notation, i do not list the specific file id’s, as i did with agentive and contrast. furthermore, conjunction does not place specific semantic demands on its arguments in the way that contrast and agentive does. the general argument concerning conjunction is the same: assume a is s1 or embedded within s1, and b is s2 or embedded within s2. if conjunction holds between a and b, then assuming x is truthful, conjunction must hold between s1 and s2. the following is an example from category 1, “x says a, b” (file 1259): 61 hardt s1= [[-jeg har taget noget tøj med til camilla,] forklarede bjørn, da de var kommet ind i stuen. ] [[-i have brought some clothes for camilla,] explained bjørn, when they had come into the living room. ] s2= [-du kan bare hente noget mere, hvis der ikke er nok.] [-you can just get some more, if there is not enough.] we have a = i have brought some clothes for camilla and b = -you can just get some more, if there is not enough.]. the dashes are meant to indicate direct quotations. so it is implicit that the speaker of s1, bjørn, is also the speaker of s2. thus s2 could be preceded with and bjørn continued: without change to the meaning. the felicity of the connective and supports my claim that the conjunction relation need not apply to the embedded a, but can relate the top level clauses s1 and s2. 6. discussion the cdt was originally formulated in accordance with uniformity; subsequently, the star notation was added because it was felt that it was necessary to violate uniformity in certain cases. in this paper, i have argued that this is not the case. there are three relations with at least ten occurrences of the star notation in cdt: contrast, agentive and conjunction.2 for contrast and agentive, i have made a general argument that quotations always involve a temporary assumption that the speaker is truthful and therefore the semantic relation attributed to the speaker will also be attributable to the writer. because of this i argued that quotations involving contrast or agentive do not require a violation of uniformity, based on an examination of all the relevant examples. the general argument involving conjunction is similar, and furthermore conjunction places rather weak requirements on its two arguments. based on these arguments, i have argued that uniformity can indeed be maintained in cdt. i have not attempted to argue that these notations are better than those which violate uniformity. for example, an anonymous reviewer points to the following pattern: john says x but john says y, where it might well be that the fundamental contrast is between x and y – that is, annotating a contrast relation between x and y might intuitively be the best choice. i do not deny that this might well be the case – i have merely attempted to show that it is possible to rule out such annotations and still be able to find acceptable annotations for all the data of the danish portion of the cdt. 7. conclusion while most treebank work has focused on syntax, there is a growing interest in treebanks that involve annotation of discourse structure. the cdt is unusual in that discourse structure is treated as an extension of syntax, with an underlying assumption of uniformity – that discourse and syntax can be annotated as part of single well-formed graph. other major discourse treebank projects have not followed this approach, instead annotating discourse independently of syntax, reflecting 2. there is no reason to assume that apparent uniformity violations are confined to these three relations. for example, an anonymous reviewer points out that the pdtb contains apparent uniformity violations with a variety of other relations, including temporal and contingency. whether these apparent violations can also be removed is a topic for future investigation. 62 syntax and discourse a widespread view that uniformity cannot be maintained, and the cdt was recently modified to allow violations of uniformity. in this paper i have shown that the data of the cdt can in fact be annotated consistently with uniformity, and i have concluded that the cdt can and should be modified to remove all conflicts between discourse and syntax. in future work, i intend to investigate larger discourse treebanks such as the penn discourse treebank to further investigate the possibility of maintaining uniformity. this is less straightforward than the current investigation, since the deviations from uniformity are not explicitly marked in the penn discourse treebank, but it will provide a much richer empirical basis for exploring the hypothesis of uniformity. references matthias buch-kromann. discontinuous grammar. a dependency-based model of human parsing and language learning. copenhagen business school, copenhagen, 2006. matthias buch-kromann and iørn korzen. the unified annotation of syntax and discourse in the copenhagen dependency treebanks. in proceedings of acl linguistic annotation workshop, 2010. matthias buch-kromann, jürgen wedekind, and jakob elming. the copenhagen danish-english dependency treebank v. 2.0. http://code.google.com/p/copenhagen-dependency-treebank, 2007. matthias buch-kromann, iørn korzen, and henrik høeg müller. uncovering the ‘lost’ structure of translations with parallel treebanks. copenhagen studies in language, (38):199–224, 2009. carlson, lynn and marcu, daniel. discourse tagging reference manual. technical report isi-tr545, isi, september 2001. carlson, lynn and marcu, daniel and okurowski, mary ellen. building a discourse tagged corpus in the framework of rhetorical structure theory. in proceedings of 2nd sigdial workshop on discourse and dialogue, eurospeech, aalborg, denmark, 2001. nikhil dinesh, alan lee, eleni miltsakaki, rashmi prasad, aravind joshi, and bonnie webber. attribution and the (non-)alignment of syntactic and discourse arguments of connectives. in proceedings of the workshop on frontiers in corpus annotation ii: pie in the sky, pages 29–36, 2005. jerry hobbs. on the coherence and structure of discourse. report csli-85-37, center for the study of language and information, stanford university, 1985. iørn korzen. endocentric and exocentric languages in translation. perspectives: studies in translatology, 13(1):21–37, 2006. iørn korzen. linguistic typology, text structure and appositions. in iørn korzen, marie lambert, and hélène vassiliadou, editors, langues d’europe, l’europe des langues. croisements linguistiques, volume 22, pages 21–42. scolia, 2007. matthias t. kromann. the danish dependency treebank and the dtag treebank tool. in treebanks and linguistic theories (tlt 2003), page 217–220, växjö, 2003. 63 hardt willliam c. mann and sandra a. thompson. rhetorical structure theory. a theory of text organization. isi/rs-87-190, isi: information sciences institute, 1987. rashmi prasad, eleni miltsakaki, nikhil dinesh, alan lee, aravind joshi, livio robaldo, and bonnie webber. the penn discourse treebank 2.0. annotation manual. the pdtb research group, 2007. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings 6th int. conf. on language resources and evaluation, marrakech, morocco, 2008. manfred stede. the potsdam commentary corpus. in proceedings of acl 2004 workshop on discourse annotation, pages 96–102, 2004. maite taboada and william c. mann. rhetorical structure theory: looking back and moving ahead. discourse studies, 8(3):423–459, 2006. bonnie webber. d-ltag: extending lexicalized tag to discourse. cognitive science, 28:751–779, 2004. florian wolf and edward gibson. representing discourse coherence: a corpus-based study. computational linguistics, 31(2):249–287., 2005. 64 microsoft word processing of discourse anaphors-copy ready.docx dialogue & discourse 12(2) 38-80 doi: 10.5210/dad.2021.202 ©2021 derya çokal, patrick sturt and fernanda ferreira this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). processing of discourse anaphors by turkish l2 speakers of english derya çokal dcokal@qmul.ac.uk school of electronic engineering and computer science queen mary university of london london, e1 4ns, united kingdom patrick sturt patrick.sturt@ed.ac.uk psychology department, university of edinburgh 7 george square edinburgh, midlothian, eh8 9jz, united kingdom fernanda ferreira fferreira@ucdavis.edu psychology department, university of california, davis 174a young hall 1 shields avenue davis, ca 95616 editor: jonathan ginzburg submitted 09/2020; accepted 07/2021; published online 08/2021 abstract this study examines the cognitive information processes that turkish advanced non-native speakers of english employ in assigning the referents of this and that in reading and production. we predicted that these speakers would assign referents in relation to the linear distance between discourse-linked anaphors and their referents in the discourse (i.e., based on spatial-temporal features of this and that), which means they would prefer this for a referent mentioned in the proximal chunk of text and that for a referent mentioned in the distal chunk. we also predicted that readers would not assign referents based on the focusing features of this and that. we tested our predictions in two eyetracking reading experiments and one sentence-completion experiment. turkish l2 learners’ online reference resolution in reading experiments was different from that of english native speakers that were tested in a previous study. in the eye-tracking experiments, turkish l2 learners did not show evidence of using a recency strategy to resolve referential ambiguity and did not use spatialtemporal or focusing features of this and that to assign referents. on the other hand, in the sentencecompletion experiment, the effect of prominence of discourse structure in the use of this and that was qualitatively similar to that of english native speakers, but their indexing of the degree of focus of this and that was different. our results suggest that the difference between turkish l2 learners and english native speakers is due to l1 interference. keywords: anaphors, processing, demonstratives, non-native speakers, l2 speakers 1 introduction a language’s discursive, pragmatic, and structural features govern speakers’ choices of discourse anaphors (d-anaphors)1 (e.g., cornish, 2010;2009; diessel, 2006). non-native speakers (l2) use d 1 throughout this paper ‘discourse anaphors’ is shortened to d-anaphors. in spoken discourse, this and that are referred to as demonstratives/deixis, which direct an interlocutor’s attention to an entity in a spatial and/or temporal context. in processing of this and that by turkish l2 speakers 39 anaphors (e.g., this and that) more extensively than do native speakers (çokal, 2019; çokal & ruhi, 2006; ellert, 2013; ferris, 1994; hinkel, 2001; jin, 2001; juvonen, 1996; niimura & hayashi, 1996; wilson, keller, & sorace, 2009). in spite of frequent d-anaphor use, l2 speakers demonstrate nontarget patterns in their use of d-anaphors (e.g., çokal, 2019; çokal & ruhi, 2006; niimura & hayashi, 1996). however, little is known about how l2 speakers select d-anaphors in production, or how they select a referent for this or that in comprehension (ionin, baek, kim, ko & wexler, 2012; swierzbin, 2010). here, we report the results of one sentence-completion experiment and two eye-tracking reading experiments, which are intended to address this gap in the literature. in a previous study, çokal, sturt and ferreira (2014) showed that discourse structure and focus features encoded in danaphors play major roles in native english (l1) speakers’ comprehension and production of this and that. the current study is an extension of çokal et al. (2014), with turkish l2 speakers of english (efl). 2 the aim is to contribute to our understanding of the on-line and offline processing of danaphors by turkish l2 speakers. we examine which factors (spatial-temporal features or degree of focus) affect l2 speakers’ biases in establishing the referents of this and that. we also hope to provide an in-depth understanding of the on-line strategy that these speakers use to disambiguate d-anaphors. in the following section, we provide a brief review of semantics and pragmatics of danaphors in english and turkish, and we discuss evidence about the comprehension strategies employed by native speakers of english, and then offer discussion of what we know about l2 strategies and their relevance to the phenomena we investigate. 2 review of literature in this section, we will briefly summarize the previous studies on discourse anaphors. 2.1 english d-anaphors in written discourse this and that bring different elements (e.g., an entity, a proposition, or an event) into focus in written and spoken discourse (cornish, 2010; diessel, 2006; fossard & rigelleau, 2005; levinson, 1993; linde 1979; mccarthy 1994; passonneau, 1993; strauss, 2002). these expressions can be used exophorically (i.e., in reference to entities present in the immediate surroundings of the speech event (halliday & hasan, 1976; levinson, 1983) or endophorically (i.e., in situations where speakers/writers refer back to elements of the ongoing discourse (levinson, 1983; lyons, 1977; see also peeters, krahmer, & maes, 2020). while exophoric use of this and that has been a main area of research interest for experimental studies, there have been few experimental studies on endophoric use of these expressions in written discourse (çokal et al., 2014; peeters, krahmer, & maes, 2020). in addition, there have been attempts to unify exoand endo-phoric uses of demonstratives (see elbourne, 2008; lücking, 2018; nunberg, 1993). it is proposed that the difference between endophoric and exophoric demonstratives is the result of finding the referents in two different contexts – utterance situation (i.e., real world) vs. co-text (i.e., written discourse). according to nunberg’s (1993) unified account, it is not the interpretation (i.e., the referent) but the index (i.e., the contextual element and a referred entity) that determines the choice of this and that. for instance, in (1), demonstrations (i.e., pointing gestures) would not be used to identify the referent addition, we identify this/that as d-anaphors (i.e., deixis in written discourse) since in written discourse they are called ‘d-anaphors’ or ‘anadeixis’, which points to a mental representation of a discourse segment in a text (cf. çokal et al., 2014 for further theoretical discussion). 2 throughout this paper, advanced learners of english as a foreign language (efl) who are native speakers of turkish are referred to as turkish non-native speakers of english or as turkish l2 speakers/learners. çokal, sturt and ferreira 40 of that. instead, a set of entities (e.g., a mask, cape) would be used to anchor or bind that in an utterance. (1) melissa made herself a mask and cape; that will be her costume (nunberg, 1993, p.37). this is the understanding that has made possible unified accounts of (endophoric) and (exophoric) demonstratives within approaches like discourse representation theory (lascarides & asher, 2007). in this study, we adopt this view and thus specifically focus on endophoric uses of this and that, which direct readers’ attention to different parts of a text (lascarides & asher, 2007; mccarthy 1994; webber, 1991,1989) and update the mental representation of the discourse (zwaan, 1996). 2.2 proximal-distal model vs. focus model to account for the functions of this and that in written discourse, the proximal-distal model and the focus model have been proposed. 2.2.1 proximal-distal model according to mccarthy’s (1994) proximal-distal model, this refers to a proximal entity/proposition, whereas that refers to a distal entity/proposition. mccarthy’s (1994) proximaldistal model is an extrapolation of the spoken use of demonstratives/deixis. 3 unit (a) in (2) below is called the distal frontier (df), which is separated from this/that by intervening clauses or units. unit (b), on the proximal frontier (pf), is the intervening clause between the unit with this/that and propositions in units (a) and (c). in (2) and (3), this and that refer to different parts of a text. (2) this/that referring to a long event on the distal frontier: (a) robert ate his meal, annoyed that there was no beer. (b) after he finished eating, he helped his girlfriend to prepare coffee with scottish whisky and cream. (c) this/that took him 2 hours, and thereafter he was pleased that he had some whisky in his coffee. (3) this/that referring to a short event on the proximal frontier: (a) robert ate his meal, annoyed that there was no beer. (b) after he finished eating, he helped his girlfriend to prepare coffee with scottish whisky and cream. (c) this/that took him 10 minutes, and thereafter he was pleased that he had some whisky in his coffee. the difference between item (2) and item (3) is the duration (2 hours vs. 10 minutes) of each event (i.e., having a meal vs. preparing coffee). the suggested interpretation is based on the assumed plausibility that the events take these different amounts of time. according to the proximal-distal model in (2), that is preferred relative to this because the reference is to an event “robert ate his meal”, which is described on the distal frontier (i.e., the non-adjacent portion of text). thus, that can signal reference across entities or foci of attention (mccarthy, 1994, p. 272). according to the proximal-distal model, in (3) above, this is preferred relative to that, because the reference is to an event described in the immediately preceding adjacent text, namely, “he helped his girlfriend to prepare coffee”. the referent of the pronominal this in unit (c) is the event in unit (b) on the pf. the pf is the clause or group of contiguous clauses immediately adjacent to the referential expression and can therefore be regarded as salient/prominent and accessible to referential expressions. according to the proximal-frontier-only hypothesis, unit (a) on the df in (2) and (3) is relatively less salient/prominent for referential expressions (webber, 1989) and thus does not provide referents for this/that. however, an alternative theory, the segmented discourse representation theory (lascarides & asher, 2007) claims that entities on the df (i.e., unit a) are 3 in a nutshell, ‘proximal’ demonstratives (e.g., english this) are used in reference to entities relatively nearby the speaker, and ‘distal’ demonstratives (e.g., english that) in reference to entities relatively far from the speaker (anderson & keenan, 1985; halliday & hasan, 1976; levelt, 1989). processing of this and that by turkish l2 speakers 41 accessible as long as certain rhetorical relations (e.g., explanation, contrast, or narration between referential expressions and entities/propositions on the df) are made. 2.2.2. focus model in contrast to the proximal-distal model, according to strauss’s (2002) focus model, this and that can signal different degrees of focus/attention that the addresser should devote to the referent. accordingly, this is used to tell an addressee to devote high attention to the referent, since it can also refer to brand new, less salient, or hitherto unshared information. strauss (2002) contends that in (2) above, this can refer to the event “robert ate his meal” on the df. in addition, that is employed to ask an addressee to devote moderate attention to the referent, since it refers to shared and salient information. therefore, from this perspective, in (3) that can refer to the event “he helped his girlfriend to prepare coffee” on the pf. however, none of these models was completely supported by empirical findings in a previous study (çokal et al., 2014), which found that english speakers use a recency strategy and the degree of focus encoded by this and that. we describe this study in detail below, and the current study is an extension of these previous findings. 2.3 comprehension strategies by l1 speakers given the longstanding experimental tradition of investigating the cognitive status of different types of anaphors (see peeters, krahmer, & maes, 2020 for a review on demonstratives) the lack of research into how writers or readers produce or comprehend ‘proximal’ versus ‘distal’ demonstrative pronouns is striking. in our previous study, we addressed this gap in the literature and conducted one sentence-completion and two eye-tracking reading experiments to test how l1 speakers of english 4 resolve the referents of this and that (çokal et al., 2014). the first eye-tracking reading experiment compared how this and that refer to distal and proximal events in a context such as the following: (4) this/that referring to a long event on the distal frontier: john drove from edinburgh to birmingham, listening to his favourite jazz cds. when he arrived in birmingham, he filled up the car with petrol. this/that took him 5 hours and afterwards he was happy to have had enough time to go to his hotel to have a rest. (5) this/that referring to a short event on the proximal frontier: john drove from edinburgh to birmingham, listening to his favourite jazz cds. when he arrived in birmingham, he filled up the car with petrol. this/that took him 5 minutes and afterwards he was happy to have had enough time to go to his hotel to have a rest. events that took a longer time (e.g., driving from edinburgh to birmingham) were given on the distal frontier (i.e., distant segment) as in (4), whereas events that took a shorter time (e.g., filling up the car with petrol) were presented on the proximal frontier (i.e., the segment adjacent to this/that) as in (5). the frontier referent preferences of this and that were measured by referring to matching (e.g., driving from edinburgh to birmingham takes 5 hours) or mismatching timespans (e.g., driving from edinburgh to birmingham does not take 5 minutes). native speakers of english showed longer reading times when this and that referred to the event (i.e., driving from edinburgh to birmingham) on the df (as in 4) than when referring to the event on the pf (as in 5). to replicate the results of experiment 1, in the second eye-tracking experiment (see method in experiment 2.), çokal et al. (2014) changed the order of events: events of shorter duration (e.g., he filled up the car with petrol) were moved to the df, while events of longer duration (e.g., john drove from edinburgh to birmingham) were placed on the pf. as in experiment 1, regardless of 4 in this study, the native speakers of english did not speak other languages and were living in scotland native speakers of non-british varieties of english (e.g., north american or australian) were not recruited. çokal, sturt and ferreira 42 whether this or that was used, references to the event on the pf led to shorter l1 reading times (e.g., indicative of preference) than references to the event on the df. this suggests that native speakers of english do not use the proximal-distal model (i.e., this can refer to a proximal/recent entity/proposition, whereas that refers to a distal entity/proposition (cf. mccarthy’s proximal-distal model, 1994). in addition, they did not employ ‘degrees of attention’ encoded by this and that (cf. strauss’ focus account, 2002). in addition, the on-line reading experiments demonstrated that native speakers of english preferred both this and that to refer to the event on the proximal frontier. however, references to the distal frontier with this and that led to processing difficulties. in the sentence-completion experiment, they used the same stimuli, but with a blank after the words following this/that in each sentence: john drove from edinburgh to birmingham, listening to his favourite jazz cds. when he arrived in birmingham, he filled up the car with petrol. this/that __________. when completing the sentence, english writers’ references to propositions/events on the pf (with both this and that) were more frequent than df references. however, this was more likely than that to refer to the df. in accordance with strauss’s (2002) focus account, the findings suggest that references with this signal the english writer is asking the interlocutor to place a high focus on the referent, since the referent is ‘less salient’ and related to the ‘earlier part of discourse’. on the other hand, with that, the english writer is asking the interlocutor to place a medium focus on the referent, since the referent is on the pf, and thus is providing information on the adjacent frontier. in summary, çokal et al. (2014) showed that when l1 english speakers are reading, they use a recency strategy and prefer references to salient propositions or events on the pf (i.e., prominent unit in discourse), rather than references to less salient propositions or events on the df (i.e., less salient/prominent). their findings of a ‘recency strategy’ to resolve a referential ambiguity support previous psycholinguistic studies (e.g., anderson, garrod, & sanford, 1983; ehrlich & rayner, 1983; klin, guzman, weingartner, & ralano, 2006; o’brien, raney, albrecht, & rayner, 1997). again, in sentence completion, english writers mostly refer to salient propositions on the pf. however, when l1 english speakers refer to less salient propositions on the df, they prefer this to that, which suggests their preferences are governed by: (1) the prominence of discourse frontiers (i.e., distal/less salient vs. proximal/salient) and (2) the degree of focus an addressee must devote in order to resolve ambiguity (i.e., high vs. medium focus). having discussed çokal et al.’s (2014) findings from l1 speakers, we will now briefly describe turkish d-anaphors. 2.4 turkish d-anaphors in turkish, the use of referential expressions (overt or null) is determined by the discourse context, and conveys pragmatic information such as contrast, similarity, emphasis or new information (enç, 1986; erguvanlı-taylan, 1986). once the discourse topic is set with a noun phrase or overt pronoun, continuity requires the use of null pronouns. but when the speaker wants to mark a topic change, an overt pronoun is necessary for grammaticality (öztürk, 2001). in the cases of d-anaphors in turkish, an overt pronoun is used. while a threefold d-anaphora system (bu, şu, and o) is commonly seen in spoken discourse (kornfilt, 1997; küntay & özyürek, 2006; lyons, 1977), a twofold anaphoric system, using bu/this and o/that, is generally observed in written discourse (sağın-şimşek, rehbein, & babur, 2009; turan, 1997). since şu is used cataphorically and the current paper’s focus is on the anaphoric uses of this and that, we will not mention the functions of şu below. it should be noted that turan’s (1997) investigation of written discourse, in which she examined the cognitive status of bu in the establishment of attentional structure, revealed that to signal topic continuation, bu refers to an entity on the pf. similarly, özil and şenöz (1996) processing of this and that by turkish l2 speakers 43 proposed that bu can be used anaphorically to refer to propositions, clauses or verb-phrases (vps) as antecedents. sağın-şimşek et al. (2009) made implicit analogies between d-anaphors in written discourse and those in spoken discourse. they investigated the use of bu and o in two novels (babamin bavulu by orhan pamuk and mavi karanlik by vedat turkali) and annotated their references to an entity/proposition. while bu (n= 89) refers to an entity that is in the focus of reader/listener (i.e., an entity on the proximal frontier), o (n= 59) refers to an entity that is distal to the focus of reader/listener (i.e., an entity on the distal frontier). o is used to refocus an entity that is mentioned earlier in the text. the use of turkish d-anaphors in written discourse is determined by the linear distance between d-anaphors and their referents in discourse (i.e., spatial-temporal features of bu and o). having reviewed findings on turkish d-anaphors, in the next section we briefly summarize how non-native speakers (l2) of english process and produce this and that in written discourse. 2.5 comprehension strategies by l2 speakers the question of whether l2 and l1 sentence processing are similar has long been debated (e.g., clahsen & felser, 2006). within the domain of anaphora processing, theories of l2 acquisition (sorace, 2011) have been informed by the question of whether l2 speakers can acquire native-like interpretive preferences for demonstratives/personal pronouns (e.g., ellert, 2013; niimura & hayashi, 1996; wilson et al., 2009), and whether l2 preferences vary between task types (roberts, gullberg, & indefrey, 2008). wilson et al. (2009) and ellert (2013) both investigated demonstratives in german, however, wilson et al. (2009) examined their uses in the context of animate references, while ellert (2013) explored their uses in inanimate references. so, though the type of demonstratives they tested cannot refer to events or propositions such as this/that in english, below we briefly summarize their findings under two categories ([1] on-line studies showing l2 on-line processing deficits and [2] l1 interferences), in order to make inferences based on the similarity of the phenomena under investigation in the current study. 2.5.1 l2 on-line studies in visual world paradigm experiments5 involving a judgment task, wilson et al. (2009) examined whether english-speaking learners of german used grammatical role (subject-verbobject [svo] order vs. object-verb-subject [ovs] order) or thematic role (active vs. passive conditions with svo and ovs word order) in assigning referents of referential expressions (personal pronouns [er, sie, es] vs. demonstrative pronouns [der, die, das]). 6in german, while er is often used to refer to the first mentioned entity (e.g., entity in subject position), der is used to index a topic shift (i.e., focus change from the first mentioned entity to the less salient entity (e.g., focus change from entity in subject position to an entity in object position). to examine this, wilson et al (2009) used the sentences like (6) and tested whether participants prefered er referring to firstmentioned noun phrase (np1) references (i.e., der kellner) and der referring to second-mentioned noun phrase (np2) references (i.e., detektiv). (6) der kellner erkennt den detektiv, als das bier umgekippt wird. er/der ist offensichtlich sehr fleißig. the waiter recognises the detective as the beer is tipped over. 5 the visual-world paradigm (in which participants listen to a sentence or short text while their eye-movements are recorded as they look at a scene or series of images) exploits the fact that the interpretation of auditory linguistic input influences eye-movements for scenes or images (see altmann & kamide, 2007). 6 wilson et al. (2009) reported three experimental strands on (1) information structure and pronominalization, (2) antecedent preferences of anaphoric demonstratives/ pronouns, and (3) simulating l2 learners’ difficulties in native speakers. here we only report the most relevant experiment to our current study. çokal, sturt and ferreira 44 he/der is clearly very hard working. wilson et al. (2009) showed that german native speakers (l1) had different referent preferences for different types of referential expressions: personal pronouns (i.e., er) for np1 references and demonstratives (e.g., der) for np2 references. while l1 and l2 speakers of german interpreted personal pronouns similarly, english-speaking learners of german (l2) did not consistently interpret demonstratives as ‘indexing a topic-shift’. in addition, they ‘had no clear preference for np2s’ as antecedents of demonstratives. overall, l2 speakers preferred both personal pronouns and demonstratives for np1 references. they also preferred first-mentioned entity/topic references, irrespective of type of referential expressions (i.e., er or der) or information structure (i.e., for the salient np/first-mentioned np vs. less salient np2). in the light of these findings, wilson et al. (2009) proposed that l2 learners might have difficulty efficiently integrating different sources of information in on-line processing and thus have a deficit in ‘resource allocation’ during on-line processing. to link correct antecedents and referential expressions, l2 learners need to integrate many sources of information: (a) the current discourse model (i.e., which entities have been mentioned and are prominent), (b) real world knowledge, (c) syntactic knowledge (i.e., er is used for topic references whereas der can shift attention from np1 to np2). similarly, using a visual-world eye-tracking paradigm and an offline referent assignment task, ellert (2013) tested the preferences dutch learners of german for personal pronouns (e.g., german er) with topical references (i.e., references to an entity in subject position) as in (7a), and demonstratives (e.g., german der) with non-topical references (i.e., references to an entity in object position as in (7b). (7a) der schrank ist schwerer als der tisch. er (personal pronoun condition) stammt aus einem möbelgeschäft in belgien. das sofa soll nächste woche geliefert werden. (7b) der schrank ist schwerer als der tisch. der (d-pronoun condition) stammt aus einem möbelgeschäft in belgien. das sofa soll nächste woche geliefert warden. ‘the cupboard is heavier than the table. it (personal/demonstratives) originates from a furniture store in belgium. the sofa is supposed to be delivered next week.’ (ellert, 2013, p. 178) similar to wilson et al.’s (2009) findings, dutch-speaking learners of german had a topic-reference preference (i.e., np1) for pronouns/demonstratives. however, they did not have distinct referent preferences for different types of anaphora, even though dutch and german have similar pronoun/demonstrative systems. ellert (2013) proposed that the l2 on-line processing cannot be merely explained by l1 influences but needs to take more general l2 learner effects into account.7 these two studies indicate that l2 learners develop a ‘learner specific’ referent assignment strategy involving a first mentioned-entity/topic-reference preference irrespective of anaphora types (ellert, 2013; wilson et al., 2009). these findings also suggest that observed non-target uses of l2 speakers in interpreting d-anaphors cannot simply be explained by l1 influence (ellert, 2013), since irrespective of l2 speakers’ native language (i.e., typologically similar or different anaphora system), they exhibited non-target antecedent preferences for demonstratives. 7 unmarked referential expressions signal that their referent is a continuation of the topic previously established. on the other hand, marked referential expressions direct the reader to new entities or topics that are not highly salient. therefore, for l2 learners the referential expression “that” can bring a less salient entity into focus and signal an addressee to raise high focus (i.e., attention calling). processing of this and that by turkish l2 speakers 45 2.5.2 l1 interference in a sentence completion experiment that examined typologically different anaphora systems across different languages (i.e., english/non-null language vs. japanese/null language) with english l1 native speakers and japanese l2 learners of english, english l1 speakers assigned the referents of demonstratives based upon focus encoded in demonstratives, whereas japanese l2 speakers’ demonstrative choice was based on the entity’s spatial-temporal distance (niimura & hayashi, 1996). for instance, one of the illustrated cartoon characters notices a stain on the carpet, bends down to have a closer look and says, “hmm, a bad stain on the carpet’. his face is close to the stain and thus this is not an instance of proximity. in this case, 68% of l1 speakers chose that, while 28% chose this & 3% chose it, while non-native (l2) speakers preferred this 69% of the time, that 18 %, and it 13%. in an accompanying illustration, a character finds a piece of paper on the floor while he is standing upright, looking down at a piece of paper. he utters, “hello? what’s…?”. in this situation, 76% of the l1 speakers chose this over that (24%), even though the character’s physical distance from the referent is greater. on the other hand, 55% of l2 speakers preferred this, while 44% preferred that, and only 1% chose it. it should be noted, that in both instances, l2 speakers preferred this over that. niimura and hayashi (1996) state that in an identical context in japanese, the appropriate response for l1 speakers would be ko/this because both cases fall in the domain of a speaker's direct experience and anything within the speaker's territory, especially something within reach, is referred to by ko. while native speakers of english chose demonstratives on the basis of focus (strauss, 2002) instead of the physical proximity of the referent, japanese l2 learners’ preference was based on physical distance to the referent. in another study, l2 learners’ performance changed across on-line and offline tasks (roberts, gullberg, & indefrey, 2008). in an acceptability judgement task, while dutch l1 speakers and german l2 speakers of dutch both interpreted an overt pronoun as referring to a discourse topic, turkish-speaking l2 speakers of dutch chose an overt pronoun as indexing a topic shift. this shift brings a new entity into the discourse, since turkish is a null subject language, and an overt pronoun is used to change topics. consequently, the offline preference of turkish l2 speakers of dutch was for an overt pronoun to index a topic shift, which was evidence of language transfer from turkish to dutch. however, in on-line reading tasks this l1 interference was not seen; both turkish and german l2 learners had processing disadvantages irrespective of the anaphora systems in their native language. in addition, in an on-line eye-tracking experiment, roberts et al. (2008), tested how both german and turkish l2 learners of dutch read sentences with/without competitor antecedents in the preceding discourse for personal pronouns. both german and turkish l2 groups had similar preferences, which were different from dutch native speakers. specifically, when reading the critical region (i.e., the verb and subject: eet hij “he eats”), their second pass and total reading times were longer in a condition in which there were two grammatically matching referents available (peter and hans are in the office. while peter is working, he is eating a sandwich.), then in a condition with only one grammatically available referent (the workers are in the office. while peter is working, he is eating a sandwich.). based upon their experiment results, roberts et al. (2008) propose that during real-time comprehension, l2 learners find the integration of information from multiple sources more difficult than do native speakers, since they need to coordinate syntactic and semantic or pragmatic information. in summary, previous l2 on-line studies demonstrate: (1) l2 learners may not ‘index topicshift’ with demonstratives, resulting in unclear or inconsistent referent preferences of demonstratives (ellert, 2013; wilson et al., 2009), and (2) their non-target use might not be due to l1 interference (ellert, 2013). instead, such non-target use may be related to a deficit in ‘resource allocation’ (wilson et al., 2009) or difficulties integrating information from multiple resources during on-line processing (roberts et al., 2008). while previous offline studies point to l1 interference (niimura and hayashi, 1996), roberts et al. (2008) found l1 interference in offline çokal, sturt and ferreira 46 tasks, but there was no evidence of l1 interference in on-line tasks. however, there were l2 processing disadvantage – irrespective of native language – in on-line tasks. when we examine l2 studies, we see that so far they have involved comparisons between personal pronouns and demonstratives but not examinations within d-anaphors. however, they have not assessed whether l2 learners index the focusing features of this and that, or how l2 learners identify the referents of this and that in either comprehension or sentence production. in addition, to our knowledge, we still do not know what on-line strategies turkish-speaking l2 learners use when disambiguating d-anaphors. importantly, the studies discussed above have mostly focused on english, german, dutch, and italian l2 speakers but not turkish l2 speakers. thus, the current study will provide results on the development of discourse-conditioned structures for turkish learners of english. more importantly, unlike english, turkish is a null-subject language. therefore, when there is a topic continuation between (8a) and (8b), referents are maintained with a null pronoun (as in 8b). in (8a) and (8b), the narrator explains what s/he did yesterday. in (8c), a new character ‘murat’ is introduced. in (8d), the overt pronoun o is used to state that the narrator did like the movie, but murat did not. it should be noted that in turkish, overt pronouns are used when the referent is marked for contrast or topic shift (cf. enc, 1986). 8a. ben dün sinema-ya git-ti-m. i yesterday cinema-dat go-past-1sg ‘yesterday i went to the cinema.’ 8b. øj film-i beyen-di-m. øj movie-acc likepast-1sg ‘(i)j liked the movie.’ 8c. aynı film-i murati izle-miş. same movie-acc murat watch-past-3sg ‘murati watched the same movie.’ 8d. ama oi beğen-me-miş. but he like-neg-past-3sg ‘he did not like it.’ the use of d-anaphors in turkish is a special case, in which addressers use them overtly. in both english and turkish, information structure (i.e., salience of df and pf) plays a role in the use of d-anaphors. in most language-teaching textbooks used in turkey for students to learn english, only the spatial-temporal aspects of english d-anaphors are explicitly highlighted, and the degree of focus encoded by d-anaphors is not taught (cokal, 2019). from an educational perspective, the role of instruction on l2 development of d-anaphors is important, since the degree of focus of this and that is perceptually less salient/subtle and might go unnoticed for l2 learners (cf. dekeysar, 2002; reviews of the experimental and quasi-experimental investigations into the effectiveness of l2 instruction in ellis, 2015; ellis, 2008). specifically, in an efl context, turkish l2 learners are often not exposed to naturally occurring input. consequently, the efficient use of this and that requires turkish l2 learners/writers to learn the similarities and differences across these two languages. while the effect of instruction on l2 development is not the current focus of the study, it is still important background information, worthy of examination. with this in mind, in the next section, we will list the current study’s hypotheses. 2.6 the current study given previously obtained results (çokal et al., 2014), in this study the following account is assumed for native english speakers: l1 english speakers consider (a) ‘the prominence of discourse structure’ and thus prefer an event on the pf as a referent of this and that, and also (b) the degree processing of this and that by turkish l2 speakers 47 of focus an addressee needs to devote to a less salient/distal event on the df and thus prefer this to that. based on what is currently known in niimura and hayashi (1996), we predicted that turkish l2 speakers of english would not index the ‘degree of focus’ encoded by this and that and thus they would not use this to refer to a less salient event on the df. instead, they would use the spatialtemporal features encoded by this and that. they would prefer this when referring to a salient/proximal event on the pf (i.e., an event on the clause adjacent to this) and that when referring to a less salient/distal event on the df (i.e., an event on the distal clause to that). the l2 speakers’ non target pattern may be due to l1 interference. basically, turkish l2 learners would transfer what they know about this and that to their on-line processing and offline production. in written discourse, bucorresponds to this (i.e., refers to an entity/event that has been mentioned recently in the text), whereas ocorresponds to that (i.e., refers to an earlier entity/event in the text) (sağın-şimşek, rehbein, & babur, 2009). in addition, in spoken discourse, both o (i.e., referring to distal elements) and şu (e.g., pure attention calling) are mapped onto that. therefore, unlike native speakers of english, for turkish l2 learners, that but not this can refer to an event on the df (çokal, 2020; çokal & ruhi, 2006). alternatively, turkish l2 learners would not consider either ‘degree of focus’ or ‘spatialtemporal features’ of this and that. they would only consider ‘prominence of discourse structure’ and thus would use recency strategy (i.e., both this and that referring to an event given on the pf). turkish l2 learners’ referent preferences were observed to change across offline and online tasks (roberts et al., 2008). with such findings in mind, we predicted that turkish l2 learners’ referent choices might be much clearer in an offline task (i.e., sentence completion) than an on-line task. on the other hand, in an on-line task (i.e., on-line reading experiment), they would not use ‘spatial-temporal features’ of this and that, degree of focus, or recency strategy to disambiguate this/that. if so, then our results would suggest that these l2 learners have disadvantages in on-line processing. to disambiguate the referents of this/that, they need to do the following: (a) maintain ‘global coherence’ and keep track of what information is accessible, (b) follow information structure to determine salient/accessible or less salient/less accessible events on the discourse frontiers, (c) match these frontiers with degree of focus encoded by this and that, (d) avoid l1 interference (i.e., spatial-temporal features of bu and o), and (e) use real world knowledge about event structures. this process requires them to handle d-anaphora systems in two languages and distinguish ‘form-functions of d-anaphors’ in both english and turkish. thus, if they do not have referent preferences for this and that, then based on the previous findings (roberts et al., 2008; wilson et al., 2009), we could propose that in on-line reading, they would have a processing deficit in the integration of all this information. to test our predictions, we conducted three experiments: one offline sentence-completion study and two on-line eye-tracking experiments. while the on-line experiments were conducted first, for presentational purposes we will report the offline experiment (experiment 1) first and then the two on-line experiments. in experiment 1, we examined turkish l2 preferences during the use of this and that in narrative completions. in experiments 2 and 3, in order to observe turkish l2 speakers’ processing of this and that in narrative written discourse, we recorded eye-movements during reading. 2.6.1 participants participants in three experiments were turkish l2 speakers of english who were third or fourthyear students in the foreign language teaching department at middle east technical university (metu). all volunteers were between 21 and 24 years of age (m = 22; sd = 1.126), unaware of the study’s purpose, paid for their participation, and had not spent extended periods of time in english-speaking countries. however, all participants had passed metu’s mandatory first-year english proficiency exam, which includes listening and writing tests. participants’ mean proficiency exam score was 81 (sd = 3.39), which is equal to 102 toefl (ibt) or 7.5 ielts (see conversion table http://oidb.metu.edu.tr/en/english-proficiency). the metu mean proficiency çokal, sturt and ferreira 48 score supports categorizing participants as l2 advanced speakers (for a conversion table with levels of proficiency in english, see http://www.eurogates.nl/en-toefl-ielts-score-conversion/). in addition, it should be noted that the medium of instruction at metu is english; therefore, students are required to speak and write in english in all courses. 3 experiment 1 experiment 1 tested turkish l2 preferences using a sentence completion method. although this technique does not provide any insight into the time-course of processing, experience in our laboratory has shown it to be a very sensitive measure of reference assignment preferences, which is presumably because it is less vulnerable to individual differences (i.e., the timing with which information is used in processing), compared with an on-line method such as eye-tracking. 3.1 method in this section, we will describe our participants, materials, and data analysis. 3.1.1 participants participants were paid turkish advanced non-native speakers of english (n = 30). the l2 participants in this sentence-completion experiment were a subset of those of experiments 2 & 3. the sentence completion experiment was conducted eight months after experiment 3. once again, participants were not informed of the study’s purpose and had not spent time in english speaking countries. 3.1.2 materials we used the same stimuli as in experiment 2 below. differently from experiment 2, each participant was provided with an initial sentence/context and the remainder of the second sentence (after this or that) was left blank. in an initial context, we always gave a long event (e.g., driving from istanbul to zonguldak) on the distal frontier of discourse and a short event (e.g., filling up the car with petrol) on the proximal frontier. the initial context always included a sequence of events. therefore, the df and pf were in narration relation. we indicated the end of the long event using when or after adverbial clause. each participant was asked to provide a completion answer for the sentence in a manner consistent with the previous text. a sample stimulus is given below: berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. this ________. there were 40 experimental and 60 filler sentences (see item 4 in appendix.). there were two experimental conditions, one of which included the word this as the final word before the blank (as in the sample above), and the other of which included the word that. two versions of each sentence and two files were constructed, using latin square counterbalancing. in each file, each sentence appeared in only one condition, with an equal number of items from each condition. sentences were presented on the computer in a fixed random order. 3.1.3 procedures at the beginning of each trial, participants looked at a blank square on the monitor to generate a new stimulus, and then completed the sentence given in the stimulus by speaking aloud in a clear voice. 3.1.4 data analysis processing of this and that by turkish l2 speakers 49 though our purpose was to investigate turkish l2 speakers’ pronominal use of this and that, we also coded their uses of this and that as prenominals (this/that + noun phrase [np]). when coding prenominal uses, we also considered whether the l2 speakers referred to proposition/events on the distal or proximal frontiers (e.g., this/that + np referring to an entity on df or pf). we used the following continuation codings and samples for prenominal and pronominal this and that. • if pre(pro)nominal this or that referred to the short event, then its referent was coded as the proximal frontier referent: (1) ozlem went to the hairdresser, hoping to get the latest hairstyle. when she was at the hairdresser, she looked through the magazines for inspiration. this made her confused about the new hairstyles. (2) aysun watched her favorite tv show heroes, relaxing on her sofa. after she had watched the final episode, she prepared a face mask with a little milk and avocado. that mask helped her get rid of all the spots on her face. • if pre(pro)nominal this or that referred to the distal clause with the long event, then its referent was coded as the early clause referent: (1) derya had face-to-face meetings with her clients, highlighting the new product features. before she left her office, she visited the ladies’ room downstairs. that was very tiring. (2) mary went to the hairdresser, hoping to get the latest hairstyle. when she was at the hairdresser, she looked through the magazines for inspiration. this hairdresser was the one she has been going to for the last 10 years. • if prenominal this or that referred to a group of sentences, then the group of sentences was entered as referent. (1) faruk planted 50 roses, following all the instructions. after he had planted the roses with great care, he watered them with a watering can. that was an exciting experience for him and he was looking forward to seeing them grow. (2) murat washed his peugeot 407, listening to capital fm. after he had cleaned the outside of his car, he vacuumed the floor, including the area beneath the seats. this was a preparation for the wedding of one of his friends. • other categories: if the referents of this or that were not clear, if the new discourse focus was introduced by the use of this or that, if this or that was used with time expressions such as this night or that evening, or if incoherent sentences were made with the previous part of the text, then all these cases were coded as other categories. two phd students (unaware of the purposes of the experiment) independently transcribed the data and coded the continuations according to the categories given to them. there was high interrater reliability on the referents of this and that (cohen’s kappa: this: > 0.85; that > 0.81; this np > 0.89; that np< 0.85). twenty continuations that annotators did not understand were excluded from data analysis. 3.2 results figure 1 shows the relative proportions of references to the dfs and pfs, as a function of danaphors type in turkish l2’s completions. figure 2 highlights the distribution of other types of references. because data for this experiment are categorical, the statistical analyses in this section used logistic mixed effects regression, taking condition (this vs. that) as the fixed effects, and çokal, sturt and ferreira 50 including crossed random intercepts and slopes for participants and items: response ~ (condition +1| subject) + (condition + 1|item) + condition. data and scripts are available at http://osf.io/v2msd/ figure 1. turkish l2 speakers’ choice of referents of this and that (i.e., pronominals). error bars show standard errors. in 27% of cases, this and that were used prenominally (e.g., this trip) or their antecedents were unclear (please see appendix item a1.). the proportion of these trials (coded as other) did not differ between this and that conditions (z = -0.392, p = .70). these prenominal cases were analysed separately from the pronominal usages, since the prenominal and pronominal usages are linguistically and cognitively distinct (see the differences in ariel, 1996; gundel, hedberg, & zacharski, 1993.). figure 2. turkish l2 speakers’ percentages of other categories. error bars show standard errors. in this and that conditions, the turkish l2 speakers had a strong preference for the pf reference (this = 90.1%, that = 75.29). in the logistic mixed effects regression, we coded references to the 0 10 20 30 40 50 60 70 80 90 100 this that pe rc en ta ge o f r ef er en ts referents of this and that df pf 0 5 10 15 20 25 30 distal frontier proximal frontier a group of sentences others pe rc en ta ge o f r ef er en ts referents of this + np and that + np this np that np processing of this and that by turkish l2 speakers 51 df as 0 and the pf as 1. the intercept was significantly different from zero (𝛽" = 1.55, z = 5.187, p< .001, reflecting the overall preference for the pf reference (see appendix item a1.). the logistic mixed effect analysis revealed a main effect of condition (𝛽" = 1.20, z = 5.79, p< .001). the turkish l2 learners referred to the pf more with this than with that (this pf = 90%; that pf = 75%). correspondingly, the complementary pattern for reference to the df was this df = 10% and that df = 25%. the turkish l2 speakers used this+np more frequently than that+np when referring to the events on the df (df: this+np = 21%; that+np = 9%) (figure 2) (see appendix item a1.). references to the events on the pf were more frequent with that+np than this+np (pf: this+np = 38%, that+np = 62%). for the l2 speakers, this+np can more easily access the events on the df than that+np. thus, the frequency of references with that+np to the events on the pf was higher than with this+np. for turkish l2 learners, pronominal that is more deictically marked8 than this and therefore it can more easily refer to events on the df, relative to this. this+np is more deictically marked than that+np and thus can access propositions on the df. 3.3 conclusion in summary, in the sentence-completion experiment, like english native speakers (çokal et al., 2014), turkish l2 speakers predominantly used this and that to refer to propositions on the pf of the discourse structure. in addition, they had the following biases: this for the event on the pf and that for the event on the df. in the prenominal cases, the turkish l2 speakers preferred this+np over that+np to refer to an event on the df. in this experiment, the difference between english native speakers and turkish l2 learners was indexing ‘the degree of focus’ encoded by this and that. therefore, l1-speaker referent preferences are opposite to those of l2 speakers: this for references to less salient/less accessible proposition on the df and that for salient/shared and focused elements on the pf. the results show that turkish l2 speakers are able to use this and that differentially to refer to different parts of texts and bring different elements into focus. the sentence-completion experiment showed that turkish l2 speakers’ uses of this and that are governed by l1 interference, with bu (corresponding to this) used for references to salient/proximal/focused elements and o (corresponding to that) used for references to less salient/distal/unfocused elements. it is worth emphasizing that the large correspondence between l1 and l2 was the result of the pf being strongly preferred. 4 experiment 2 experiment 2 was an eye-tracking in reading study using turkish l2 speakers, and it was based closely on experiment 1a of çokal et al.'s (2014), which used l1 speakers of english. the experiment had a 2 × 2 within-subjects design, crossing two levels of deictic expressions (this and that) with two levels of discourse frontiers (proximal and distal) as seen in example 9 below: (9a) this referring to a long event on the distal frontier of discourse: berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. this took him 5 hours, and afterwards he was happy to have had enough time to go to his hotel to have a rest. (9b) that referring to a long event on the distal frontier of discourse 8 unmarked referential expressions signal that their referent is a continuation of the topic previously established. on the other hand, marked referential expressions direct the reader to new entities or topics that are not highly salient. therefore, for our l2 learners the referential expression “that” can bring a less salient entity into focus and signal an addressee to raise high focus (i.e., attention calling). çokal, sturt and ferreira 52 berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. that took him 5 hours, and afterwards he was happy to have had enough time to go to his hotel to have a rest. (9c) this referring to a short event on the proximal frontier of discourse berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. this took him 5 minutes, and afterwards he was happy to have had enough time to go to his hotel to have a rest. (9d) that referring to a short event on the proximal frontier of discourse berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. that took him 5 minutes, and afterwards he was happy to have had enough time to go to his hotel to have a rest. the events on the proximal and distal frontiers, which served as referents for this and that, differed in their typical time duration (see 9a, 9b, 9c, and 9d). as in (9a, 9b), the df described an event of relatively “long” duration (i.e., driving from istanbul to zonguldak), while as in (9c, 9d), the pf described an event of relatively “short” duration that was part of the “long” event (i.e., filling up the car with petrol). in the critical sentence, this and that referred to a specific event that was of either long (e.g., five hours) or short (e.g., five minutes) duration, matching either the long event on the df or the short event on the pf. after this and that, in order to present the time duration that the long and short events required, we used the same structure ... took him/her + time duration throughout the stimuli. after the time duration, the second clause started with and, followed by adverbials with seven or more characters (e.g., and afterwards) to allow for the possibility participants would see some of that word parafoveally while fixating on and. all long event clauses had a modifier (e.g., listening to his favourite jazz cds). in order to prevent modifiers being mistaken for a referent of this and that, attention was paid to not introduce a new event or use psychological verbs (e.g., planning or thinking). in order to signal that two different events were being mentioned and to indicate the end of the long event and the start of a new one, adverbial clauses with after, once or when were provided before clauses with short events (e.g., when he arrived in zonguldak). in addition, the long event was always on the df (the distal clause in the first sentence), while the short event was always on the pf (the immediately preceding clause in the second sentence). we conducted this time duration manipulation since we assumed that reference assignment to an event occurs soon after a subject reads the d-anaphor at the start of the critical sentence. in addition, we surmised that turkish l2 learners should exhibit processing difficulty if the length of the event (e.g., driving from istanbul to zonguldak) did not match the time duration subsequently mentioned in the critical sentence (e.g., five minutes). thus, differences in reading times among conditions should indicate which event was initially chosen as the referent. 4.1 predictions here, we assumed that if turkish l2 learners would have any unclear or indeterminate referential preferences for this and that, then there would be no reading time differences in (9a, 9b, 9c, and 9d). alternatively, we assumed that if l2 speakers relied on spatial-temporal functions of this and that (e.g., references to the distal/proximal frontiers), then processing difficulty should be greater and reading times would be longer when this referred to the event on the distal frontier (as in 9a) and that referred to the event on the proximal frontier (as in 9d). this interaction should be observed at the disambiguating region where l2 readers first encounter the disambiguating information, which will be reflected in regression path at the disambiguation region. if l2 readers refixate on the d-anaphors region after disambiguation, then the interaction will be found in the danaphors region in second pass reading time and total time, since these measures include refixations after the reader has progressed beyond the analysis region. processing of this and that by turkish l2 speakers 53 another possibility was that the reverse pattern for this and that would be observed if readers were sensitive to the degree of focus encoded by this and that. consequently, reading times would be longer when this referred to the event on the pf (as in 9c) and that referred to the event on the df (as in 9b). this preference will be seen in regression path time for the disambiguating region and in second pass and total time for the d-anaphors region. in addition, if the df was inaccessible, then reading times for both this and that should be generally longer for references to the df (as in 9a, 9b) than the pf (as in 9c, 9d). 4.2 method in this section, we will describe our participants, materials, and data analysis. 4.2.1 participants participants were turkish l2 speakers of english (n = 52). 4.2.2 materials there were 40 experimental items, in each of the four experimental conditions illustrated (9a, 9b, 9c, and 9d) (see item 5 in appendix.). in the four constructed files, each sentence appeared in only one condition and each condition appeared an equal number of times. there were 60 fillers and eight practice items, all of which were similar in length to the control sentences. in the filler items, events involved causes and consequences. texts were presented on 3-4 written lines of between 66 to 76 characters. this and that always appeared in the second sentence around the middle of the line. each participant saw all the fillers. 4.2.3 procedures we presented 108 texts in a fixed random order, so no two experimental items appeared adjacent to each other. thirteen participants were assigned to each list. in order to familiarize participants with the experimental procedure, the experiment began with eight fillers. the experiment was conducted using an eyelink 1000 eye-tracker (sr research ltd, ottawa, canada) in tower-mounted mode, with a chin rest to stabilize the participant’s head. items appeared on a monitor approximately 80 cm from the participant’s eyes. only the right eye was tracked, but viewing was binocular. before each item, participants fixated on a black square so the experimenter could check calibration of subject’s eyes. after reading each item, participants pressed the x-button on the controller to see the corresponding comprehension question, the left button for the answer option on the left [false], and the right button for the option on the right [true]. the comprehension questions never probed the referents of this/that (e.g., did berk drive to kastamonu?). we asked participants to read the sentences at a natural pace, and then answer the question accurately and, if possible, quickly. 4.2.4 data analysis individual fixations of less than 80 ms and more than 1200 ms were excluded from the analysis. such long fixations usually indicate tracker loss. all the participants scored at least 90% correct on their answers to the questions, and the analyses below include all trials regardless of question-answering accuracy. we did not exclude participants/trials because the by-participant intercepts adjust for individual differences in reading speed, and the log transformation in our model reduces the impact of individual trials with extreme values. texts were divided into three regions (see table 1), and we report analyses for regions 1, 2, and 3. çokal, sturt and ferreira 54 analysis region r1 d-anaphor this/that took him r2 time duration 5 hours/5 minutes r3 connector and adverbial , and afterwards table 1. analysis regions (r) in experiment 2 all analyses were conducted using linear mixed effects regression (lmer) and the lme4 r package.9 an additional package (plyr) was used to compute p-values. for each region and measure, an lmer model, incorporating all fixed effects and their interactions in a single step, was constructed (barr et al., 2013). factor labels were transformed into numerical values and centered prior to analysis, to have a mean of 0 and a range of 1. the results provide coefficients, standard errors, and t-values for each fixed effect and interaction. analysis was based on log reading times, and all analyses reported below incorporated crossed random intercepts for participants and items. random slope parameters (levels of deictic expressions) (e.g., this and that), two levels of frontiers (e.g., proximal and distal), and the interaction in the slopes (danaphors* frontiers +1|participants) were included in the maximal model for both participants and items (see table 2). data and scripts are available at http://osf.io/v2msd regression path times were defined as the sum of all fixations from the first entry into the region from the left until the first fixation to a later region and were our measure of early processing (reflecting the fixation behaviour that immediately follows the reader’s initial inspection of a given region). in the analysis of regression path times, trials where the reader skipped the region of interest during initial reading were removed from the analysis and treated as missing data. the second pass reading measure was defined as the sum of all fixation durations following the first exit of the region either to the right or left, and it therefore excluded time spent during the initial reading of a region. therefore, this gives information about delayed aspects of processing, as well as the re-processing required when a reader looks back at a portion of text. for second pass reading times, if a region was not re-fixated, this would result in a value of zero for the relevant trial. such trials values (i.e., those with no re-fixations) were removed from the second pass reading analysis. total reading times were the sum of all fixations in the region, reflecting overall processing. for total reading time, regions with no fixations in any given trial were treated as missing data and removed from total reading time analysis. 4.3 results regression path times did not reveal a significant interaction between d-anaphor and discourse frontier, or any main effects of discourse frontier (see tables 2 & 3 and figure 3.) 9 r 3.6.0 gui 1.70 el capitan build (7657). processing of this and that by turkish l2 speakers 55 eye-movement measures regions & parameters regression path times second reading times total reading times r1: d-anaphors β$ se t β$ se t β$ se t intercept 6.38 0.04 142 5.39 0.02 268 6.50 0.05 130 d-anaphors -0.01 0.02 -0.48 0.02 0.32 0.81 -0.06 0.02 -*2.52 frontiers -0.03 0.03 -1.14 -0.06 0.27 -*2.04 -0.04 0.02 -1.869 d-anaphors x frontiers -0.00 0.05 -0.11 0.00 0.05 -0.15 0.01 0.05 0.198 r2: time duration intercept 5.73 0.04 141 5.36 0.01 275 5.83 0.04 136 d-anaphors -0.00 0.02 -0.22 -0.04 0.03 -1.60 -0.00 0.02 -0.03 frontiers -0.06 0.03 -1.77 0.00 0.03 0.07 -0.03 0.02 -1.41 d-anaphors x frontiers 0.03 0.05 0.53 0.01 0.06 0.17 0.01 0.05 0.191 r3: connector & adverbial intercept 6.02 0.03 156 5.35 0.02 238 6.11 0.04 138 d-anaphors -0.01 0.02 -0.68 0.00 0.03 0.15 -0.02 0.02 -1.17 frontiers 0.00 0.03 0.03 -0.01 0.04 -0.32 -0.03 0.02 -1.33 d-anaphors x frontiers 0.01 0.06 0.20 -0.08 0.07 -1.09 -0.01 0.05 -0.22 *statistically significant effects are indicated with asterisk. effects are considered significant when the absolute value of t was 2 or greater. table 2. experiment 2: results of mixed-effects analysis for the eye-movement measures (complex model) çokal, sturt and ferreira 56 regions regression path second pass total reading r1: d-anaphors this df 813 (58) 240 (9) 790 (35) that df 750 (38) 235 (9) 837 (37) this pf 749 (53) 221 (6) 752 (32) that pf 718 (39) 217 (8) 785 (33) r2: time duration this df 421 (34) 213 (6) 420 (15) that df 474 (56) 228 (9) 425 (18) this pf 378 (22) 223 (8) 408 (18) that pf 427 (40) 226 (11) 420 (24) r3: connector & adverbial this df 484 (25) 228 (8) 541 (23) that df 555 (56) 222 (11) 556 (28) this pf 544 (42) 208 (7) 511 (20) that pf 519 (39) 222 (6) 524 (23) table 3. experiment 2: means (and standard errors) for regression-path times, second pass reading times and total reading times. figure 3. regression path times (ms) across regions, for experiment 2. error bars show standard errors. in second pass reading times (see appendix item a2.), for the region of d-anaphors (r1), a main effect of discourse frontier was found (see tables 2 & 3 and figure 4). second pass reading times were longer when this/that referred to the long event on the distal frontier than those with the proximal frontier (df: m = 237, se = 9.67; pf: m = 219, se = 7.27). 0 100 200 300 400 500 600 700 800 900 1000 df pf df pf df pf d-anaphors time duration connector & adverbial fi xa tio n tim es in m s regression path times this that processing of this and that by turkish l2 speakers 57 figure 4. second pass reading times (ms) across regions, for experiment 2. error bars show standard errors. in total reading times (see appendix item a2.), for the region of d-anaphors (r1), a main effect of d-anaphor was found (see tables 2 & 3 and figure 5). total reading times were longer when that was used than when this was used (that: m = 811, se = 24.73; this: m = 771, se = 23.71). figure 5. total reading times (ms) across regions, for experiment 2. error bars show standard errors. 4.4 conclusion the results of experiment 1 show some evidence consistent with the claim that, similarly to native english speakers in çokal et al. (2014), turkish l2 speakers of english preferred references to the proximal frontier, regardless of whether this or that was used. however, turkish l2 learners’ preferences for pf references were only seen in second pass reading times and in only the d-anaphor region. reading times for turkish l2 learners were longer when this and that referred to the event on the df. experiment 2 results do not show evidence that turkish l2 learners’ preferences involved ‘spatial-temporal’ or ‘degree of focus’ features of this and that. 0 50 100 150 200 250 300 df pf df pf df pf d-anaphors time duration connector & adverbial fi xa tio n tim es in m s second pass reading times this that 0 100 200 300 400 500 600 700 800 900 1000 df pf df pf df pf d-anaphors time duration connector & adverbial fi xa tio n tim es in m s total reading times this that çokal, sturt and ferreira 58 5 experiment 3 in experiment 3, we investigated whether turkish l2 speakers of english would have different referent preferences for this and that if the order of the long and short events in the context was changed. experiment 3 (using l2 speakers of english) was an extension of çokal et al. (2014) and designed to rule out a potential alternative interpretation of experiment 2’s results, which had suggested faster reading times for references to the short event on the pf, relative to references to the long event on the df. we interpreted experiment 2’s results as a mismatch effect, showing a preference for proximal-frontier reference for both this and that. however, the results could also be interpreted as demonstrating a reading preference for short time durations (e.g., five minutes) relative to long durations (e.g., five hours) or a referent preference for short events over long ones. if so, this might be unrelated to the referent’s distal or proximal-frontier status. 5.1 method in this section, we will describe our participants, materials, and data analysis. 5.1.1 participants turkish l2 participants in experiment 3 (n = 40) had taken part in experiments 1 and 2. however, experiments 2 and 3 were conducted at least six months apart. during this ensuing timeperiod, participants did not spend an extended period of time in english-speaking countries. 5.1.2 materials to rule out this potential counter-explanation, we replicated experiment 2’s design, changing only the order of events (i.e., the “short” event was moved to the df and the “long” event was moved to the pf.), as seen in example 10 below: (10a) this referring to a long event on the proximal frontier: berke filled up the car with petrol, being careful not to spill any over his white wedding trousers. then he drove from istanbul to zonguldak. this took him 5 hours, and afterwards he was happy not to have had to stop on his way. (10b) that referring to a long event on the proximal frontier: berke filled up the car with petrol, being careful not to spill any over his white wedding trousers. then he drove from istanbul to zonguldak. that took him 5 hours, and afterwards he was happy not to have had to stop on his way. (10c) this referring to a short event on the distal frontier: berke filled up the car with petrol, being careful not to spill any over his white wedding trousers. then he drove from istanbul to zonguldak. this took him 5 minutes, and afterwards he was happy not to have stained his trousers. (10d) that referring to a short event on the distal frontier: berke filled up the car with petrol, being careful not to spill any over his white wedding trousers. then he drove from istanbul to zonguldak. that took him 5 minutes, and afterwards he was happy not to have stained his trousers. another difference from experiment 2, was that the end of the “short” event and the start of the “long” event was indicated by an adverbial (e.g., then), instead of an adverbial clause (e.g., when or after) in experiment 3 (see item 6 in appendix.). processing of this and that by turkish l2 speakers 59 5.1.3 predictions if turkish l2 speakers showed a preference for reference to the pf as in experiment 2, then we should again observe faster reading times for references to the event on the pf in (10a, 10b) even though this event now involved a “long” duration, relative to the “short” event depicted in the clause on the df in (10c, 10d). however, if turkish l2 speakers merely took less time to process “short” durations than “long” durations, then the opposite pattern should be found. again, l2 preferences would be revealed in regression path at the disambiguation region and in second pass reading time and total time at the d-anaphors region. l2 preference in regression path at the danaphors region is not interpretable since they have not yet encountered disambiguating information. 5.1.4 procedures the procedure was the same as experiment 2. 5.1.5 data analysis all details of data analysis were identical to those of experiment 2. 5.2 results all participants correctly answered at least 90% of the comprehension questions. there were no comprehension differences across conditions in either experiment 2 or experiment 3. all trials are included in the analyses. the results of regression path times, second pass reading time, and total reading times (see appendix item a3.) of l2 speakers of english are reported in tables 4 and 5. eye-movement measures regions & parameters regression path times second pass reading times total reading times r1:d-anaphors β$ se t β$ se t β$ se t intercept 6.35 0.04 130 6.03 0.04 132 6.40 0.05 123 d-anaphors -0.02 0.02 -0.80 0.00 0.06 0.05 0.00 0.02 0.28 frontiers 0.06 0.03 1.87 0.04 0.06 0.75 0.02 0.02 1.04 d-anaphors x frontiers 0.06 0.06 0.96 0.09 0.12 0.75 -0.04 0.05 -0.86 r2: time duration intercept 5.65 0.03 154 6.01 0.05 112 5.73 0.05 113 d-anaphors -0.07 0.03 *-2.26 -0.02 0.08 -0.23 -0.01 0.02 -0.59 çokal, sturt and ferreira 60 frontiers 0.01 0.03 0.49 0.03 0.08 0.42 -0.03 0.03 -1.11 d-anaphors x frontiers -0.00 0.07 -0.01 -0.13 0.16 -0.81 0.01 0.06 0.29 r3: connector & adverbial intercept 5.94 0.04 133 6.08 0.05 111 6.05 0.04 133 d-anaphors 0.00 0.01 0.20 0.05 0.10 0.52 -0.01 0.02 -0.60 frontiers --------0.01 0.10 0.12 -0.04 0.02 -1.70 d-anaphors x frontiers 0.01 0.01 0.93 -0.00 0.23 -0.00 -0.00 0.05 -0.01 *statistically significant effects are indicated with asterisk. effects are considered significant when the absolute value of t was 2 or greater. table 4. experiment 3: results of mixed-effects analysis for the eye-movement measures (complex model) regions regression path second pass total reading r1: d-anaphors this df 648 (31) 528 (47) 700 (35) that df 783 (64) 531 (39) 687 (44) this pf 719 (43) 601 (51) 704 (39) that pf 761 (57) 646 (90) 722 (38) r2: time duration this df 343 (26) 546 (58) 371 (25) that df 379 (27) 574 (61) 394 (28) this pf 337 (20) 702 (126) 365 (22) that pf 473 (64) 633 (93) 380 (22) r3: connector & adverbial this df 440 (25) 615 (111) 499 (31) that df 476 (30) 591 (74) 510 (29) this pf 467 (27) 612 (80) 479 (24) that pf 540 (103) 699 (196) 490 (24) table 5. experiment 3: means (and standard errors) for regression-path times, second pass reading times and total reading times. processing of this and that by turkish l2 speakers 61 figure 6. regression path times (ms) across regions, for experiment 3. error bars show standard errors. in regression path times, for the d-anaphors region (r1), there were no main effects or interaction. a pairwise comparison showed that this referring to the df was significantly different from that referring to the df (t = -2.108, p = .036) (see tables 4 & 5, figure 6.). however, this result is uninterpretable and superfluous because it was not replicated in other regions or eyemovement measures. in addition, it was seen before participants encountered the disambiguating region. in regression path times, for the time duration region (r2), there was a main effect of danaphors (see tables 4 & 5 and figure 6). references to the events with that led to longer regression path times than references to the events with this (that: m = 426, se = 35.73; this: m = 339, se = 16.20). no main effect or interaction was seen in other regions (r1, r3) for regression path, second pass reading, and total reading times (see tables 4 & 5 and figures 7 & 8). figure 7. second pass reading times (ms) across regions, for experiment 3. error bars show standard errors. 0 100 200 300 400 500 600 700 800 900 1000 df pf df pf df pf d-anaphors time duration connector & adverbial fi xa tio n tim es in m s regression path times this that 0 100 200 300 400 500 600 700 800 900 1000 df pf df pf df pf d-anaphors time duration connector & adverbial fi xa tio n tim es in m s second pass reading times this that çokal, sturt and ferreira 62 figure 8. total reading times (ms) across regions, for experiment 3. error bars show standard errors. 5.3 conclusion experiment 3 was designed to test whether turkish l2 readers would continue to show a preference for reference to the event on the pf, when the relative positions of the “long” and “short” events were opposite from experiment 2. if this was the case, we would expect similar results to those of l1 native readers reported by çokal et al. (2014), namely evidence of processing difficulty when the d-anaphor referred to the event described on the df, even though this event was now of “long” duration. however, in contrast to this prediction, experiment 3 results showed no evidence of processing difficulty for reference to the df. therefore, based on the results of experiments 2 and 3, we cannot rule out the possibility that turkish l2 readers simply preferred descriptions of “short” time durations (e.g., five minutes) over “long” durations (e.g., five hours), and possibly did not fully resolve the d-anaphor during on-line reading of the text. in this experiment, there was no evidence that turkish l2 learners used features of ‘spatial-temporal’, ‘degree of focus’ or ‘prominence of discourse (i.e., recency strategy). 6 general discussion this study was designed to explore turkish l2 speakers’ preferences in establishing the referents of this and that. we investigated whether turkish l2 learners would disambiguate referents of this and that and – if so – whether ‘spatial-temporal features’ or ‘degree of focus’ would affect their choice of referent, in both on-line and offline settings. we predicted that turkish l2 speakers would not index ‘degree of focus’ encoded by this and that, but instead would rely upon ‘spatial-temporal features’ in accordance with mccarthy’s (1994) proximal-distal model. secondly, we assumed l2 learners’ preferences could change across on-line and offline experiments and their referent preferences would likely be clearer in the offline sentence completion experiment than the on-line eye-tracking experiment. in addition, we predicted that in an offline task (but not an on-line experiment), we could observe l1 interference. we also predicted that in an on-line task (i.e., on-line reading experiment), l2 learners would not use ‘spatial-temporal features’ of this and that, ‘degree of focus’, or recency strategy to disambiguate this/that. if this is the case, then we surmised that l2 learners would have on-line processing deficits and disadvantages. alternatively, we also considered the possibility that, like the l1 english speakers in çokal et al. (2014), turkish l2 speakers’ processing difficulty would be greater (reading times would be longer) and they would consider the prominence of discourse structure (i.e., recency strategy) if the distal frontier was inaccessible to references with this and that. 0 100 200 300 400 500 600 700 800 900 1000 df pf df pf df pf d-anaphors time duration connector & adverbial fi xa tio n t im es in m s total reading times this that processing of this and that by turkish l2 speakers 63 our first prediction was that turkish l2 speakers would rely upon ‘spatial-temporal features’ in accordance with the proximal-distal model to determine the referents of d-anaphors. this was robustly supported in the sentence-completion experiment but in our on-line reading experiments, we could not find evidence that turkish l2 learners were employing mccarthy’s (1994) ‘spatialtemporal’ model or strauss’ (2002) ‘degree of focus’ model to disambiguate the referents of danaphors. in experiment 1 (i.e., sentence completion experiment), when turkish l2 learners were asked to refer to a certain part of a text with this and that, they preferred that when referring to a distal/less salient event/entity on the distal frontier and this in reference to a proximal/salient event/entity on the proximal frontier. thus, there were contrasting distal and proximal frontier preferences for this and that between l1 and l2 speakers. in çokal et al. (2014), l1 english speakers used degree of focus, which means this was used as an attention-caller to ask for high attention from an addressee when referring to less salient/distal elements on the distal frontier. on the other hand, that was used to refer to the referent/information on the proximal frontier/recent segment. however, turkish l2 learners seemed to rely upon turkish d-anaphors in written discourse (i.e., l1 interference): this/bu refers to a salient entity/event on the proximal frontier, whereas that/o refers to a less salient/distal entity/event on the distal frontier. in addition, that might be mapped onto both o and şu in spoken discourse and therefore – for turkish l2 learners – that can be more deictically marked than this when referring to a distal element, functioning like an “attention-caller” to specify reference. 10 as was suggested by a reviewer of this manuscript, another possible explanation might be that our l2 participants transferred what they knew about deictic use of this and that and implemented it in a discourse anaphora context. importantly, their preference pattern is consistent with the turkish d-anaphora system and deictic features. we also predicted that if turkish l2 learners do not use ‘spatial-temporal’ or ‘degree of focus’ to determine referents of d-anaphors, then they would have a general preference of an event on the pf as a referent of this and that. in other words, similar to l1 english speakers (çokal et al., 2014), turkish l2 speakers’ processing difficulty would be greater (and thus reading times longer) if the df was inaccessible to references with this and that. in on-line reading, multiple eye-movement measures of l1 english speakers have shown an overall bias towards interpreting this and that when referring to the pf (çokal et al., 2014). çokal et al. (2014) interpreted this as a recency strategy in online reading. alternatively, it could be interpreted as native speakers’ preferences for this/that referring to the combined event/an entire paragraph (e.g., driving from istanbul to zonguldak + filling up car with petrol). our experiment 2 (i.e., on-line reading experiment) showed some evidence in favour of this interpretation, since turkish l2 speakers showed longer second pass reading times when the d-anaphor was disambiguated to refer to the df, relative to the proximal frontier. however, it should be noted this finding was found in only one measure and was not replicated in experiment 3. therefore, we did not find strong evidence for a recency strategy in on-line reading or the interpretation of the entire paragraph. the turkish l2 learners’ nontarget pattern (i.e., not having a consistent bias for pf references), or not using spatial-temporal features, might be because they did not effectively track the referents of this and that during on-line reading, although we also acknowledge that the on-line experiments may have lacked the power to detect the relevant effects in a noisy sample of l2 reading. 11 alternatively, if their nontarget pattern reflects an ‘l2 on-line processing deficiency/disadvantage’, it is consistent with previous l2 studies that found l2 learners did not integrate all information in on-line reading (roberts et al., 2009; wilson et al., 2009). we also assumed that our l2 learners could have a different pattern/performance in the sentence completion experiment, compared to the on-line comprehension experiments. our 10 referential expressions direct the reader to new entities or topics that are not highly salient. 11 hopp (2014) tested l2 readers processing of syntactic ambiguities and showed that there was an on-line preference only for readers who had a high lexical automaticity score (based on lexical decision times). the participants who had a lower automaticity score showed no significant difference. çokal, sturt and ferreira 64 findings support this. not only did the sentence completion task force l2 learners to have referent preferences for d-anaphors, but we also found a pattern consistent with l1 interference. on the other hand, in the on-line experiment, readers were not forced to make a decision regarding referents, and this might have been why it was difficult to find on-line evidence for l2 learners’ preferences for this/that. these differences across tasks might indicate that l2 learners were not sure enough of their preference to generate reading difficulty in the reading experiment. in other words, turkish l2 speakers may have been open to different configurations they found in the text because they were still ‘learning’. however, l1 speakers’ preferences in the on-line task in çokal et al (2014) were much clearer, and these readers’ fixation times showed evidence of processing difficulty when a reference went in an unexpected direction. regarding our main question of whether turkish l2 speakers can acquire native-like interpretive preferences of d-anaphors, our sentence completion experiment showed that turkish l2 learners’ use of ‘prominence of discourse structure’ in the use of this and that’ was quantitatively similar to that of english native speakers in çokal et al. (2014). both language groups preferred a salient/recent event on the pf as a referent of this and that while completing sentences. in addition, similarly to english native speakers (çokal et al. 2014), turkish l2 learners’ completions showed evidence of sensitivity to this and that bringing different entities/events into focus and referring to different discourse frontiers. however, the difference between english native speakers in çokal et al. (2014) and turkish l2 speakers in the current study was their indexing degree of focus. our findings support niimura & hayashi (1996), in which l2 speakers did not index degree of focus. such differences between turkish l2 learners and english native speakers are due to l1 interference. while turkish l2 speakers’ on-line reference resolution in reading experiments might be different from those of english native speakers, much more compelling evidence for the l2 speakers’ recency preference came from the data in the sentence completion experiment. methodologically, the overall pattern of findings suggests that the sentence completion method is a more sensitive measure of l2 reference representation than an on-line measure (such as eyetracking). in this study, we did not explore the role of cognitive measures (e.g., working memory) in processing of d-anaphors in turkish. we also did not investigate the role of different language proficiency levels or language instruction, as well as age, in processing and production of danaphors. these are further research avenues to explore. acknowledgements the project was funded by oyp and the scientific and technological research council of turkey (tubitak), grant no. 108k405 and the erc, grant number 695662. we would also like to thank the reviewers and journal editor for their valuable feedback. processing of this and that by turkish l2 speakers 65 appendix a: 1. since english native speaker results are taken from experiment 1a of çokal et al. (2014), statistical information from that paper is repeated here for comparison purposes. in 14% of cases, this and that were used prenominally (e.g., this boy), or their antecedents were unclear. the proportion of these trials (coded as other) did not differ between this and that conditions (z= -1.33, p = .18). native speakers preferred pf references. there was a significant effect of condition (β$ = 0.696, z = 3.12, p = .001). l1 speakers had more proximal frontier reference for that than for this (this pf= 78%, that pf= 84%). this, not that, was used to access entities/propositions on the distal frontier (this df = 22% and that df = 16%). prenominal that+np was used more frequently to refer to events on the distant or proximal frontier than this+np (df: this+np = 10.8%, that+np = 15.3%; pf: this+np = 9.9%, that+np = 15.3%). 2. in çokal et al. (2014), native speakers’ data were reported using repeatedmeasures anova treating d-anaphors (this-that) and frontiers (distal/proximal frontiers [df/pf]) as within-participant and within-item analysis. (1) here, we briefly report the same native speakers’ data in linear-mixed effect model to compare native and non-native speakers’ preferences. in second-pass reading times, for the d-anaphors region (r1), a main effect of frontier was significant (β$= -0.090, se = 0.041, t = -2.178). second pass reading time were longer than for the proximal frontier (df = 139 ms, se = 12.140; pf =109 ms, se = 9.688) when this and that referred to the long event on the distal frontier. however, a main effect of d-anaphors and the interaction between these two factors were not significant (intercept: β$ = 5.754, se = 0.047, t = 122.071; danaphors: β$ = 0.013, se = 0.044, t = 0.297, d-anaphors x frontiers: β$ = 0.036, se = 0.081, t = 0.448). (2) for the region of d-anaphors (r1), a main effect of discourse frontier was significant (β$= 0.085, se = 0.027, t = -3.112) in total reading times. when this and that referred to the long event on the distal frontier (e.g., driving from edinburgh to birmingham), total reading times were longer than for the proximal frontier (e.g., filling up the car with petrol) (df = 524 ms, se = 28.942; pf = 478 ms, se= 26. 282). a main effect of d-anaphors and the interaction between these two factors were not significant (intercept: β$ = 6.019, se = 0.052, t = 115.63; d-anaphors: β$ = 0.029, se = 0.022, t = 1.304, d-anaphors x frontiers: β$ = 0.003, se = 0.044, t = 0.085). 3. in second pass reading times, for the region of duration (r2), a main effect of frontier was found (intercept: β$= 342.79, se = 29.38, t = 11.667, frontier: β$= -64.37, se = 33.62, t = 2.115; d-anaphors: β$= -38.85, se = 31.17, t = -1.247, d-anaphors x frontiers: β$= 61.46, se = 44.01, t = 1.397). regardless of whether this or that was used, references to events on the distal frontier led to longer second pass reading times than to events on the proximal frontier (df = 139 ms, se = 12.140; pf = 109 ms, se = 9.688.) (3) in total reading times, for the region of time duration (r2), a main effect of frontier was found (β$= -32.146, se = 14.836, t = -2.167). total reading times were longer when this and that referred to the short event on the df (df = 336, se =14.142; pf = 305, se = 9.553). a main effect of d-anaphors and an interaction between these factors were not significant (intercept: β$= 316.288, se = 16.171, t = 19.559, d-anaphors: β$= -3.939, se = 14.454, t = -0.272, d-anaphors x frontiers: β$= 6.728, se = 27.433, t = -2.167). in the connector and adverbial regions (r3), the same effect of frontier was seen (intercept: β$= 5.84, se = 0.047, t = 122.607, frontier: β$= -0.036, se = 0.022, t = -1.634; d-anaphors: β$= -0.078, se = 0.028, t = -2.779, d-anaphors x frontiers: β$= 0.032, se = 0.042, t = 0.761, df: df = 417, se =15.57; pf = 377, se = 13.39) çokal, sturt and ferreira 66 appendix b: stimuli for experiment 1 1a/1b. berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. this/that... 2a/2b. berra packed her belongings with the help of her best friend. once she had wrapped everything, she put the packages in her small car with great care. this/that... 3a/3b. ceren moved to her new flat, taking all her belongings in one handbag and three suitcases. after she arrived, she called her mother with her blackberry. this/that... 4a/4b. sena went to the hairdresser, hoping to get the latest hairstyle. when she was at the hairdresser, she looked through the magazines for inspiration. this/that... 5a/5b. mehmet ate his meal, annoyed that there was no beer. after he finished eating, he helped his girlfriend to prepare coffee with sco:ttish whisky and cream. this/that... 6a/6b. ali prepared roast turkey and potatoes for his dinner party, planning to have drinks with the meal. after roasting the turkey, he prepared a caesar salad. this/that... 7a/7b. eren reorganised the seating plan, considering the phd students' seating preferences. after he arranged the new seating plans in the offices, he went to his office on the first floor to have a strong coffee with whipped cream. this/that... 8a/8b. beytullah planted 50 roses, following all the instructions on the plant packaging. after he planted the roses with great care, he watered them with a watering can. this/that... 9a/9b. samet flew back from edinburgh to turkey, travelling with his wife. when he arrived at ataturk, he went to the duty free shop to buy whisky for his father-in law. this/that... 10a/10b. melisa baked a cake and some scones, following the recipe in ceyhan's kitchen. after she finished the scones, she washed the dishes on the counter. this/that... 11a/11b. arda studied for the chemistry exam, worrying about the problems he might have to solve the next day. after he revised all the topics, he had a hot shower. this/that... 12a/12b. alev prepared her third-year project, feeling anxious about not finishing\nit on time. after she finished writing up the project, she submitted it to her instructor. this/that... 13a/13b. sude did some spring cleaning at her flat, feeling surprised at how quickly it got dirty. after she cleaned the whole flat, she prepared a pasta salad for dinner. this/that... 14a/14b. neslihan had an intense workout at cse gym in istanbul, working with her trainer. after she completed her workout routine, she went to the sauna with her cd player. this/that... 15a/15b. irmak went to the dentist with toothache on monday morning, feeling uncomfortable from the pain. when she got there, she had to have the tooth out. this/that... 16a/16b. hasan did his weekly shopping, picking up all the goods on the shopping list. when he arrived back at his flat, he put the groceries into the fridge and cupboards. this/that... 17a/17b. nagehan swam in the swimming pool, pleased to know that she was burning some calories and staying fit. when she went to the changing room, she had a shower. this/that... 18a/18b. ata wrote his book entitled eye-tracking compiling all his experiences. after he finished writing up the chapters, he designed the cover page of his book using photoshop. this/that... 19a/19b. korhan washed his peugeot 4007, listening to capital fm. after he cleaned the outside of his car, he vacuumed the floor including the area beneath the seats. this/that... 20a/20b. baturay drove from edirne to gaziantep, admiring the scenery on the way. when he arrived in gaziantep, he looked for a cheap restaurant to have dinner. this/that... 21a/21b. furkan went to his friend's party on friday night, taking two bottles of wine and a tub of ice-cream with him. once the party was over, he took a taxi home. this/that... 22a/22b. nurbanu experienced a problem with the printer, worried that she would miss the deadline for her project. after she unsuccessfully tried to fix it, she went to the staff office in the basement to get help from one of the servitors. this/that... 23a/23b. nevreste picked strawberries, helping her uncle and his workers. after she picked them, she walked to her uncle's house with her uncle and the workers. this/that... processing of this and that by turkish l2 speakers 67 24a/24b. sevilay rode her bike from her flat to work, singing a song. when she arrived at the bike rack, she bought some tea and cookies from the sun canteen. this/that... 25a/25b. can climbed kizlarsivrisi with his friends, using his climbing gear. when he reached the top, he ate his food with two bottles of fresh orange juice. this/that... 26a/26b. deren walked around istanbul, enjoying the blue sky and brisk wind. when she stopped at the galata tower, she had a bite to eat and drank mint tea. this/that... 27a/27b. nurhayat had a health checkup on women's check-up day, including screening, bone density and blood tests. after she finished her planned tests at memorial hospital, she visited her doctor with her husband and daughter. this/that... 28a/28b. ekrem had a lovely breakfast with his family, enjoying the sun outside. after he had eaten the delicious breakfast, he went to the toilet with the newspaper. this/that... 29a/29b. goncay played squash with her ex-boyfriend, hoping that she would beat him as she had in the old days. after the game, she did her routine stretching. this/that... 30a/30b. behiye watched her favorite tv show heroes, relaxing on her sofa. after she watched the final episode, she prepared a face mask with a little milk and avocado. this/that... 31a/31b. sabutay argued with his wife about their financial problems, hoping that the children would not hear. after the argument, he had a hot shower to calm down. this/that... 32a/23b. bilge had face-to-face meetings with her clients, highlighting the new product features. before she left her office, she visited the ladies room downstairs. this/that... 33a/33b. cendel reorganised some files in his office, planning to make space for his new books and files. while he was moving them, he read one of the documents. this/that... 34a/34b. nisa enjoyed a camping trip in fethiye, leaving all the stress of work and family behind her. when she got back home, she read her emails and letters. this/that... 35a/35b. deniz went for a morning run in the new city, not knowing her way around. when she realized that she was not on the right road, she bought a map from the newsagents. this/that... 36a/36b. semirramis waited for the response from her job application, hoping to hear back soon. while she was waiting, she visited her cousin in geneva. this/that... 37a/37b. emircan held a house party, having heard his parents would be gone for the night. when his party ended, he washed all the wine glasses and plates by himself. this/that... 38a/38b. ece flew from ankara to new york, excited to be a phd student. when she landed at kennedy airport, she experienced a problem at passport control. this/that... 39a/39b. bircan went shopping for a wedding dress in suadiye, wishing to find the dress of her dreams. when she found it, she phoned her bridesmaid to describe it. this/that... 40a/40b. ada baked a cake for her friend’s party, following her favourite recipe. after she baked it, she set the oven timer following the instruction manual. this/that... çokal, sturt and ferreira 68 appendix c: stimuli for experiment 2 1a/1b. berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. this/that took him 5 hours, and afterwards he was happy to have enough time to go to his hotel to have a rest. 1c/1d. berk drove from istanbul to zonguldak, listening to his favourite jazz cds. when he arrived in zonguldak, he filled up the car with petrol. this/that took him 5 minutes, and afterwards he was happy to have enough time for coffee. 2a/2b. berra packed her belongings with the help of her best friend. once she had wrapped everything, she put the packages in her small car with great care. this/that took her 3 hours, and subsequently she was pleased to have finished everything on time. 2c/2d. berra packed her belongings with the help of her best friend. once she had wrapped everything, she put the packages in her small car with great care. this/that took her 8 minutes, and subsequently she was pleased to have fitted them all into her car. 3a/3b. ceren moved to her new flat, taking all her belongings in one handbag and three suitcases. after she arrived, she called her mother with her blackberry. this/that took her 5 hours, and afterwards she was too tired to unpack her suitcases. 3c/3d. ceren moved to her new flat, taking all her belongings in one handbag and three suitcases. after she arrived, she called her mother with her blackberry. this/that took her 5 minutes, and afterwards she was too tired to call her boyfriend. 4a/4b. sena went to the hairdresser, hoping to get the latest hairstyle. when she was at the hairdresser, she looked through the magazines for inspiration. this/that took her 2 hours, and then she was happy to get the hairstyle she desired. 4c/4d. sena went to the hairdresser, hoping to get the latest hairstyle. when she was at the hairdresser, she looked through the magazines for inspiration. this/that took her 10 minutes, and then she was happy to find the hairstyle she wanted in one of the magazines. 5a/5b. mehmet ate his meal, annoyed that there was no beer. after he finished eating, he helped his girlfriend to prepare coffee with scottish whisky and cream. this/that took him 2 hours, and thereafter he was pleased that he had spent the evening with her. 5c/5d. mehmet ate his meal, annoyed that there was no beer. after he finished eating, he helped his girlfriend to prepare coffee with scottish whisky and cream. this/that took him 15 minutes, and thereafter he was pleased that he had some whisky in his coffee. 6a/6b. ali prepared roast turkey and potatoes for his dinner party, planning to have drinks with the meal. after roasting the turkey, he prepared a caesar salad. this/that took him 2 hours, and eventually he was happy to finish it earlier than he expected. 6c/6d. ali prepared roast turkey and potatoes for his dinner party, planning to have drinks with the meal. after roasting the turkey, he prepared a caesar salad. this/that took him 15 minutes and eventually he was happy to finish making the salad earlier than he expected. 7a/7b. eren reorganised the seating plan, considering the phd students' seating preferences. after he arranged the new seating plans in the offices, he went to his office on the first floor to have a strong coffee with whipped cream. this/that took him 2 hours, and subsequently he was annoyed about the students' never-ending requests. 7c/7d. eren reorganised the seating plan, considering the phd students' seating preferences. after he arranged the new seating plans in the offices, he went to his office on the first floor to have a strong coffee with whipped cream. this/that took him 10 minutes, and subsequently he was annoyed about resuming his work. 8a/8b. beytullah planted 50 roses, following all the instructions on the plant packaging. after he planted the roses with great care, he watered them with a watering can. this/that took him 2 hours, and afterwards he was glad to finish the hard work processing of this and that by turkish l2 speakers 69 8c/8d. beytullah planted 50 roses, following all the instructions on the plant packaging. after he planted the roses with great care, he watered them with a watering can. this/that took him 10 minutes, and afterwards he was glad to have time for his visit to the dentist. 9a/9b. samet flew back from edinburgh to turkey, travelling with his wife. when he arrived at ataturk, he went to the duty free shop to buy whisky for his father-in law. this/that took him 5 hours, and eventually he was happy to arrive home. 9c/9d. samet flew back from edinburgh to turkey, travelling with his wife. when he arrived at ataturk, he went to the duty free shop to buy whisky for his father-in law. this/that took him 15 minutes, and eventually he was happy to catch his next flight on time. 10a/10b. melisa baked a cake and some scones, following the recipe in ceyhan's kitchen. after she finished the scones, she washed the dishes on the counter. this/that took her 2 hours, and therefore she was too tired to clean the flat. 10c/10d. melisa baked a cake and some scones, following the recipe in ceyhan's kitchen. after she finished the scones, she washed the dishes on the counter. this/that took her 15 minutes, and therefore she was too tired to prepare the icing. 11a/11b. arda studied for the chemistry exam, worrying about the problems he might have to solve the next day. after he revised all the topics, he had a hot shower. this/that took him 10 hours, and afterwards he was glad to be able to go to cinema. 11c/11d. arda studied for the chemistry exam, worrying about the problems he might have to solve the next day. after he revised all the topics, he had a hot shower. this/that took him 15 minutes, and afterwards he was glad to get an early night before the exam. 12a/12b. alev prepared her third-year project, feeling anxious about not finishing\nit on time. after she finished writing up the project, she submitted it to her instructor. this/that took her 1 month, and naturally she was no longer worried about the project deadline. 12c/12d. alev prepared her third-year project, feeling anxious about not finishing it on time. after she finished writing up the project, she submitted it to her instructor. this/that took her 1 hour, and naturally she was no longer worried about taking a night off. 13a/13b. sude did some spring cleaning at her flat, feeling surprised at how quickly it got dirty. after she cleaned the whole flat, she prepared a pasta salad for dinner. this/that took her 8 hours, and afterwards she was happy to have finished everything in one day. 13c/13d. sude did some spring cleaning at her flat, feeling surprised at how quickly it got dirty. after she cleaned the whole flat, she prepared a pasta salad for dinner. this/that took her 40 minutes, and afterwards she was happy to have finished preparing her dinner for the night. 14a/14b. neslihan had an intense workout at cse gym in istanbul, working with her trainer. after she completed her workout routine, she went to the sauna with her cd player. this/that took her 2 hours, and thereafter she was happy to have relaxed before her meetings later in the day. 14c/14d. neslihan had an intense workout at cse gym in istanbul, working with her trainer. after she completed her workout routine, she went to the sauna with her cd player. this/that took her 15 minutes, and thereafter she was happy to have a cold shower. 15a/15b. irmak went to the dentist with toothache on monday morning, feeling uncomfortable from the pain. when she got there, she had to have the tooth out. this/that took her 2 hours, and later she was happy she would not have to go to work that day. 15c/15d. irmak went to the dentist with toothache on monday morning, feeling uncomfortable from the pain. when she got there, she had to have the tooth out. this/that took her 15 minutes, and later she was happy she would not be in pain anymore. 16a/16b. hasan did his weekly shopping, picking up all the goods on the shopping list. when he arrived back at his flat, he put the groceries into the fridge and cupboards. this/that took him 2 hours, and subsequently he was pleased to have a beer while watching some football. 16c/16d. hasan did his weekly shopping, picking up all the goods on the shopping list. when he arrived back at his flat, he put the groceries into the fridge and cupboards. this/that took him 10 minutes, and subsequently he was pleased to start cooking the chicken and couscous. çokal, sturt and ferreira 70 17a/17b. nagehan swam in the swimming pool, pleased to know that she was burning some calories and staying fit. when she went to the changing room, she had a shower. this/that took her 2 hours, and consequently she was worried about missing the deadline of her project. 17c/17d. nagehan swam in the swimming pool, pleased to know that she was burning some calories and staying fit. when she went to the changing room, she had a shower. this/that took her 10 minutes, and consequently she was worried about being late for her class. 18a/18b. ata wrote his book entitled eye-tracking compiling all his experiences. after he finished writing up the chapters, he designed the cover page of his book using photoshop. this/that took him 10 months, and afterwards he was happy about the distribution of his book to all universities in the world. 18c/18d. ata wrote his book entitled eye-tracking compiling all his experiences. after he finished writing up the chapters, he designed the cover page of his book using photoshop. this/that took him 2 days, and afterwards he was happy about his final design. 19a/19b. korhan washed his peugeot 4007, listening to capital fm. after he cleaned the outside of his car, he vacuumed the floor including the area beneath the seats. this/that took him 2 hours, and shortly after he was proud of how quickly he cleaned the car. 19c/19d. korhan washed his peugeot 4007, listening to capital fm. after he cleaned the outside of his car, he vacuumed the floor including the area beneath the seats. this/that took him 10 minutes, and shortly after he was proud of how clean the inside of his car was. 20a/20b. baturay drove from edirne to gaziantep, admiring the scenery on the way. when he arrived in gaziantep, he looked for a cheap restaurant to have dinner. this/that took him 20 hours, and afterwards he was happy to go to his hotel to sleep. 20c/20d. baturay drove from edirne to gaziantep, admiring the scenery on the way. when he arrived in gaziantep, he looked for a cheap restaurant to have dinner. this/that took him 1 hour, and afterwards he was happy to have eaten the spiciest meal of his life. 21a/21b. furkan went to his friend's party on friday night, taking two bottles of wine and a tub of ice-cream with him. once the party was over, he took a taxi home. this/that took him 6 hours, and afterwards he was happy to have been to the most enjoyable party of his life. 21c/2d. furkan went to his friend's party on friday night, taking two bottles of wine and a tub of ice-cream with him. once the party was over, he took a taxi home. this/that took him 10 minutes, and afterwards he was happy to have arrived back home without any trouble. 22a/22b. nurbanu experienced a problem with the printer, worried that she would miss the deadline for her project. after she unsuccessfully tried to fix it, she went to the staff office in the basement to get help from one of the servitors. this/that took her 45 minutes, and consequently she was anxious because she did not know the result of missing the deadline. 22c/22d. nurbanu experienced a problem with the printer, worried that she would miss the deadline for her project. after she unsuccessfully tried to fix it, she went to the staff office in the basement to get help from one of the servitors. this/that took her 10 minutes, and consequently she was anxious because she had left her work to the last minute. 23a/23b. nevreste picked strawberries, helping her uncle and his workers. after she picked them, she walked to her uncle's house with her uncle and the workers. this/that took her 5 hours, and consequently she was fed up with the day she had had. 23c/23d. nevreste picked strawberries, helping her uncle and his workers. after she picked them, she walked to her uncle's house with her uncle and the workers. this/that took her 15 minutes and consequently she was fed up with the bad jokes and stories from the workers on the way. 24a/24b. sevilay rode her bike from her flat to work, singing a song. when she arrived at the bike rack, she bought some tea and cookies from the sun canteen. this/that took her 30 minutes, and therefore she was in hurry to get to her meeting with the head of the department. 24c/24d. sevilay rode her bike from her flat to work, singing a song. when she arrived at the bike rack, she bought some tea and cookies from the sun canteen. this/that took her 7 minutes, and therefore she was in hurry to\neat them before work. processing of this and that by turkish l2 speakers 71 25a/25b. can climbed kizlarsivrisi with his friends, using his climbing gear. when he reached the top, he ate his food with two bottles of fresh orange juice. this/that took him 10 hours, and consequently he was overjoyed because of the sense of relief and achievement. 25c/25d. can climbed kizlarsivrisi with his friends, using his climbing gear. when he reached the top, he ate his food with two bottles of fresh orange juice. this/that took him 15 minutes, and consequently he was overjoyed because his friends had not tried to drink his orange juice. 26a/26b. deren walked around istanbul, enjoying the blue sky and brisk wind. when she stopped at the galata tower, she had a bite to eat and drank mint tea. this/that took her 6 hours, and subsequently she was excited to show the photos of the trip to her family. 26c/26d. deren walked around istanbul, enjoying the blue sky and brisk wind. when she stopped at the galata tower, she had a bite to eat and drank mint tea. this/that took her 20 minutes, and subsequently she was excited to see some more new places. 27a/27b. nurhayat had a health checkup on women's check-up day, including screening, bone density and blood tests. after she finished her planned tests at memorial hospital, she visited her doctor with her husband and daughter. this/that took her 4 days, and eventually she was glad to finish her checkups. 28a/28b. ekrem had a lovely breakfast with his family, enjoying the sun outside. after he had eaten the delicious breakfast, he went to the toilet with the newspaper. this/that took him 2 hours, and eventually he was pleased to spend some time at home. 28c/28d. ekrem had a lovely breakfast with his family, enjoying the sun outside. after he had eaten the delicious breakfast, he went to the toilet with the newspaper. this/that took him 15 minutes, and eventually he was pleased to read his favourite column. 29a/29b. goncay played squash with her ex-boyfriend, hoping that she would beat him as she had in the old days. after the game, she did her routine stretching. this/that took her 1 hour, and eventually she was cheerful to go and get a drink. 29c/29d. goncay played squash with her ex-boyfriend, hoping that she would beat him as she had in the old days. after the game, she did her routine stretching. this/that took her 5 minutes, and eventually she was cheerful to return to the changing room. 30a/30b. behiye watched her favorite tv show heroes, relaxing on her sofa. after she watched the final episode, she prepared a face mask with a little milk and avocado. this/that took her 2 hours, and thereafter she was glad to have time to do whatever she liked. 30c/30d. behiye watched her favorite tv show heroes, relaxing on her sofa. after she watched the final episode, she prepared a face mask with a little milk and avocado. this/that took her 10 minutes, and thereafter she was glad to have her skin feel so smooth. 31a/31b. sabutay argued with his wife about their financial problems, hoping that the children would not hear. after the argument, he had a hot shower to calm down. this/that took him 1 hour, and afterwards he was worried since he did not know how to reconcile it with her. 31c/31d. sabutay argued with his wife about their financial problems, hoping that the children would not hear. after the argument, he had a hot shower to calm down. this/that took him 5 minutes and afterwards he was worried since he did not want to go over the same issue. 32a/23b. bilge had face-to-face meetings with her clients, highlighting the new product features. before she left her office, she visited the ladies room downstairs. this/that took her 6 hours and afterwards she was relieved to have finished her tiring day. 32c/32d. bilge had face-to-face meetings with her clients, highlighting the new product features. before she left her office, she visited the ladies room downstairs. this/that took her 10 minutes and afterwards she was relieved to have freshened up. 33a/33b. cendel reorganised some files in his office, planning to make space for his new books and files. while he was moving them, he read one of the documents. this/that took him 2 hours, and afterwards he was happy to have enough empty space for his new books. çokal, sturt and ferreira 72 33c/33d. cendel reorganised some files in his office, planning to make space for his new books and files. while he was moving them, he read one of the documents. this/that took him 13 minutes and afterwards he was happy to have read such an interesting document. 34a/34b. nisa enjoyed a camping trip in fethiye, leaving all the stress of work and family behind her. when she got back home, she read her emails and letters. this/that took her 2 days, and afterwards she was happy to feel strong enough to cope with all the challenges she faced. 34c/34d. nisa enjoyed a camping trip in fethiye, leaving all the stress of work and family behind her. when she got back home, she read her emails and letters. this/that took her 10 minutes and afterwards she was happy to go to bed early. 35a/35b. deniz went for a morning run in the new city, not knowing her way around. when she realized that she was not on the right road, she bought a map from the newsagents. this/that took her 1 hour, and eventually she was pleased to arrive back home. 35c/35d. deniz went for a morning run in the new city, not knowing her way around. when she realized that she was not on the right road, she bought a map from the newsagents. this/that took her 3 minutes, and eventually she was pleased to be able to find her way more easily. 36a/36b. semirramis waited for the response from her job application, hoping to hear back soon. while she was waiting, she visited her cousin in geneva. this/that took her 2 months, and eventually she was pleased to have time to relax after her graduation from the university. 36c/36d. semirramis waited for the response from her job application, hoping to hear back soon. while she was waiting, she visited her cousin in geneva. this/that took her 10 days, and eventually she was pleased to have got away from her hometown. 37a/37b. emircan held a house party, having heard his parents would be gone for the night. when his party ended, he washed all the wine glasses and plates by himself. this/that took him 7 hours, and subsequently he was too exhausted to even think about having another party in his parents' house again. 37c/37d. emircan held a house party, having heard his parents would be gone for the night. when his party ended, he washed all the wine glasses and plates by himself. this/that took him 1 hour, and subsequently he was too exhausted to clean the floor.] 38a/38b. ece flew from ankara to new york, excited to be a phd student. when she landed at kennedy airport, she experienced a problem at passport control. this/that took her 8 hours, and meanwhile she was angry with herself for not sleeping on the flight. 38c/38d. ece flew from ankara to new york, excited to be a phd student. when she landed at kennedy airport, she experienced a problem at passport control. this/that took her 45 minutes, and meanwhile she was angry with herself for not bringing her acceptance letter. 39a/39b. bircan went shopping for a wedding dress in suadiye, wishing to find the dress of her dreams. when she found it, she phoned her bridesmaid to describe it. this/that took her 5 hours, and afterwards she was pleased because she still had time to schedule her hairdresser and photographer. 39c/39d. bircan went shopping for a wedding dress in suadiye, wishing to find the dress of her dreams. when she found it, she phoned her bridesmaid to describe it. this/that took her 10 minutes, and afterwards she was pleased because her bridesmaid also liked the dress. 40a/40b. ada baked a cake for her friend’s party, following her favourite recipe. after she baked it, she set the oven timer following the instruction manual. this/that took her 30 minutes, and consequently she was glad to have finished on time for the party. 40c/40d. ada baked a cake for her friend’s party, following her favourite recipe. after she baked it, she set the oven timer following the instruction manual. this/that took her 5 minutes, and consequently she was glad to. have understood the manual. processing of this and that by turkish l2 speakers 73 appendix d: stimuli for experiment 3 1a/1b. berke filled up the car with petrol, being careful not to spill any over his white wedding trousers. then he drove from istanbul to zonguldak. this/that took him 5 hours, and afterwards he was happy not to have had to stop on his way. 1c/1d. berke filled up the car with petrol, being careful not to spill any over his white wedding trousers. then he drove from istanbul to zonguldak. this/that took him 5 minutes, and afterwards he was happy not to have stained his trousers. 2a/2b. berra phoned to book a taxi to the airport for 7 pm, becoming stressed by the busy operator. afterwards, she packed her suitcases with all her holiday clothes. this/that took her 1 hour, and afterwards she was sad to be leaving the country. 2c/2d. berra phoned to book a taxi to the airport for 7 pm, becoming stressed by the busy operator. afterwards, she packed her suitcases with all her holiday clothes. this/that took her 5 minutes, and afterwards she was sad to be leaving the country. 3a/3b ceren had a quick coffee at the hairdresser's, surrounded by women gossiping about the city council. then she had her hair restyled in one of the latest fashions. this/that took her 1 hour, and afterwards she was happy to have found the hairstyle she wanted in one of the magazines. 3c/3d. ceren had a quick coffee at the hairdresser's, surrounded by women gossiping about the city council. then she had her hair restyled in one of the latest fashions. this/that took her 10 minutes, and afterwards she was happy to have her morning coffee with a catch up. 4a/4b. mehmet had a look at the festival programme, hoping to find a play that grabbed him. then he saw the nihilists at the ataturk cultural centre of istanbul. this/that took him 2 hours, and afterwards he was happy to have watched a well-performed play. 4c/4d mehmet had a look at the festival programme, hoping to find a play that grabbed him. then he saw the nihilists at the ataturk cultural centre of istanbul. this/that took him 20 minutes, and afterwards he was happy to have found something that he liked. 5a/5b. yalman had his morning coffee catching up with new gossip in the department in the kitchen. after the coffee, he reorganised the seating plan in the phd students' offices. this/that took him 2 hours, and afterwards he was pleased to have his lunch soon. 6a/6b. beytullah read all the instructions from the garden centre, wearing his dolce and gabbana sunglasses. then he planted 50 roses with great care and patience. this/that took him 2 hours, and afterwards he was delighted to have done so much hard work. 6c/6d. beytullah read all the instructions from the garden centre, wearing his dolce and gabbana sunglasses. then he planted 50 roses with great care and patience. this/that took him 6 minutes, and afterwards he was delighted to have done some gardening. 7a/7b. samet went to the duty free shop to buy turkish drinks and turkish delights for his fatherin-law and workmates. then he flew back from turkey to edinburgh. this/that took him 6 hours, and eventually he was happy to arrive home. 8a/8b. melisa washed the dishes at the counter, listening to her favourite rock cds. then she read the famous historian fernand braudel's writings on the mediterranean. this/that took her 2 hours, and afterwards she was glad to start writing her thesis. 8c/8d. melisa washed the dishes at the counter, listening to her favourite rock cds. then she read the famous historian fernand braudel's writings on the mediterranean. this/that took her 25 minutes, and afterwards she was glad she had not left them until later. 9a/9b. cemal had a shower, anxious about his sister and her new boyfriend. then he studied for the maths and chemistry exams with his best friend from university. this/that took him 5 hours, and afterwards he was worried he had left his calculator at his ex-girlfriend's. 9c/9d. cemal had a shower, anxious about his sister and her new boyfriend. then he studied for the maths and chemistry exams with his best friend from university. this/that took him 5 minutes, and afterwards he was worried he had left the boiler on. çokal, sturt and ferreira 74 10a/10b. neslihan printed out her third-year essay to hand it in to her instructor, feeling anxious about not submitting it on time. then she had a long soak in the bath to relax. this/that took her 1 hour, and afterwards she was relieved to have stopped worrying about the essay at last. 10c/10d. neslihan printed out her third-year essay to hand it in to her instructor, feeling anxious about not submitting it on time. then she had a long soak in the bath to relax. this/that took her 5 minutes, and afterwards she was relieved to have submitted it on time. 11a/11b. irmak had breakfast with her boyfriend, listening to the classical collection from the 18th century on bbc radio 3. then she did some spring cleaning at her flat. this/that took her 8 hours, and subsequently she was energetic enough to accomplish the other plans she had for the evening. 11c/11d. irmak had breakfast with her boyfriend, listening to the classical collection from the 18th century on bbc radio 3. then she did some spring cleaning at her flat. this/that took her 30 minutes, and subsequently she was energetic enough to start a new day. 12a/12b. hazal had to have her tooth out, feeling uncomfortable because of her bad breath. then she spent the rest of her day checking the company accounts. this/that took her 5 hours, and later she was glad she would not\have fallen behind in her work. 12c/12d. hazal had to have her tooth out, feeling uncomfortable because of her bad breath. then she spent the rest of her day checking the company accounts. this/that took her 10 minutes, and later she was glad she did not have smelly breath. 13a/13b. aykut put the groceries into the fridge and cupboards, trying to be as tidy as he could. after he had put them away, he watched some rugby. this/that took him 2 hours, and subsequently he was pleased with the result of the match. 13c/13d. aykut put the groceries into the fridge and cupboards, trying to be as tidy as he could. after he had put them away, he watched some rugby. this/that took him 10 minutes, and subsequently he was pleased to have finished unpacking so soon. 14a/14b. nurbanu pulled over onto the hard shoulder to pick up a hitchhiker on a rainy day, realizing that it was a horrible day to be standing by the side of the road. then she drove the car to her new company in a new city centre. this/that took her 1 hour, and afterwards she was glad to have been on time for her meeting. 14c/14d. nurbanu pulled over onto the hard shoulder to pick up a hitchhiker on a rainy day, realizing that it was a horrible day to be standing by the side of the road. then she drove the car to her new company in a new city centre. this/that took her 3 minutes, and afterwards she was glad to have done her good deed for the day. 15a/15b. onur designed the cover of his book entitled eye-tracking methods in linguistics, wishing it to be more appealing to buyers. then he wrote the chapters. this/that took him 10 months, and afterwards he was confident since he had gathered all the recent studies in his chapters. 15c/15d. onur designed the cover of his book entitled eye-tracking methods in linguistics, wishing it to be more appealing to buyers. then he wrote the chapters. this/that took him 3 hours, and afterwards he was confident since his design would draw attention. 16a/16b. yalman looked under all his car seats for his brother-in-law's brown leather wallet. then he cleaned his peugeot 308 inside and out. this/that took him 2 hours, and afterwards he was pleased with how bright and shiny his car looked. 16c/16d. yalman looked under all his car seats for his brother-in-law's brown leather wallet. then he cleaned his peugeot 308 inside and out. this/that took him 10 minutes, and afterwards he was pleased to have found it. 17a/17b. volkan looked for a cheap and clean restaurant before going on a trip with his partner. after his lunch, he drove from gaziantep to kayseri. this/that took him 6 hours, and afterwards he was happy to have arrived at his hotel to sleep. 18a/18b. nevreste walked to her uncle's farm with her uncle and the workers, wishing she could finish the work quickly. after she had arrived at the farm, she picked strawberries. this/that took her 5 hours, and afterwards she was fed up with the day she had had. processing of this and that by turkish l2 speakers 75 18c/18d. nevreste walked to her uncle's farm with her uncle and the workers, wishing she could finish the work quickly. after she had arrived at the farm, she picked strawberries. this/that took her 15 minutes, and afterwards she was fed up with the bad jokes and stories from the workers on the way. 19a/19b. cengiz took a taxi to his friend's flat, taking two bottles of wine and a tub of ice-cream with him. then he sewed a thousand sequins onto his cowboy costume. this/that took him 2 hours, and afterwards he was glad to be the best-dressed cowboy at the party. 19c/19d. cengiz took a taxi to his friend's flat, taking two bottles of wine and a tub of ice-cream with him. then he sewed a thousand sequins onto his cowboy costume. this/that took him 10 minutes, and afterwards he was glad to see his friend again. 20a/20b. dilruba went to the staff office in the basement to get help from\none of the technicians. then she re-ran the experiment with a new participant. this/that took her 2 hours, and afterwards she was glad because she had finished the experiment on time. 20c/20d. dilruba went to the staff office in the basement to get help from one of the technicians. then she re-ran the experiment with a new participant. this/that took her 5 minutes, and afterwards she was glad because the technical problem was not serious. 21a/21b. sevilay bought some bread and soup from hocam piknik next door, feeling too tired to cook at home. then she rode her bike over to mithatcan's in karakusunlar. this/that took her 30 minutes, and afterwards she found that mithatcan had also bought soup for dinner. 21c/21d. sevilay bought some bread and soup from hocam piknik next door, feeling too tired to cook at home. then she rode her bike over to mithatcan's in karakusunlar. this/that took her 5 minutes, and afterwards she found that the shopkeeper had given her too much change. 22a/22b. karatay ate two sandwiches and a bar of kendal mint cakes, enjoying the sun outside. then he climbed kartaltepe with his sister and three friends. this/that took him 10 hours, and afterwards he was very pleased because his girlfriend would be waiting for him at base camp. 22c/22d. karatay ate two sandwiches and a bar of kendal mint cakes, enjoying the sun outside. then he climbed kartaltepe with his sister and three friends. this/that took him 15 minutes, and afterwards he was very pleased because his girlfriend had packed it for him. 23a/23b. nurcan had a bite to eat and drank mint tea at the galata tower, enjoying the sunset and light wind. after the galata tower, she caught the overnight train to eskisehir. this/that took her 6 hours, and afterwards she couldn't wait to get out and stretch her legs. 23c/23d. nurcan had a bite to eat and drank mint tea at the galata tower, enjoying the sunset and light wind. after the galata tower, she caught the overnight train to eskisehir. this/that took her 20 minutes, and afterwards she couldn't wait to buy some mint tea to drink at home. 24a/24b. nurhayat visited her doctor with her husband and daughter, worried she may have breast cancer. then she had a long wait to hear the results of the tests. this/that took her 5 days, and afterwards she was glad to get the all-clear. 24c/24d. nurhayat visited her doctor with her husband and daughter, worried she may have breast cancer. then she had a long wait to hear the results of the tests. this/that took her 25 minutes, and afterwards she was glad she had had her family with her. 25a/25b. samet went to the bathroom for a quick shower on sunday morning, singing one of elvis’ songs. then he had a leisurely breakfast with his family and next-door neighbours. this/that took him 2 hours, and afterwards he was pleased not to have gone out for breakfast with his friends. 25c/25d. samet went to the bathroom for a quick shower on sunday morning, singing one of elvis’ songs. then he had a leisurely breakfast with his family and next-door neighbours. this/that took him 10 minutes, and afterwards he was pleased not to have forgotten to cut his toenails. 26a/26b. gonca did her routine stretching, hoping that she would beat her ex-boyfriend as she had in the old days. then she played squash with him. this/that took her 1 hour, and afterwards she was raring to mock him. çokal, sturt and ferreira 76 26c/26d. gonca did her routine stretching, hoping that she would beat her ex-boyfriend as she had in the old days. then she played squash with him. this/that took her 5 minutes, and afterwards she was raring to go. 27a/27b. behiye prepared a hard boiled egg for her maternal grandmother's evening meal. then she watched her favourite tv show rosellinda and annabel. this/that took her 50 minutes, and afterwards she was glad to have seen brian and rosellinda finally get together. 27c/27d. behiye prepared a hard boiled egg for her maternal grandmother's evening meal. then she watched her favourite tv show rosellinda and annabel. this/that took her 10 minutes, and afterwards she was glad she had remembered to buy the eggs. 28a/28b. sabutay called baskentgaz and power to find out about their monthly payment, hoping not to pay much. after the call, he argued all evening with his wife about their financial problems. this/that took him 2 hours, and afterwards he was worried as he did not know how to make up with her. 28c/28d. sabutay called baskentgaz and power to find out about their monthly payment, hoping not to pay much. after the call, he argued all evening with his wife about their financial problems. this/that took him 5 minutes, and afterwards he was worried as he did not know how to break the bad news to her. 29a/29b. bahar visited the ladies room downstairs, wanting to start the day looking good. after the ladies room, she had face-to-face meetings with her new clients. this/that took her 6 hours, and afterwards she was relieved to have finished her tiring day. 29c/29d. bahar visited the ladies room downstairs, wanting to start the day looking good. after the ladies room, she had face-to-face meetings with her new clients. this/that took her 10 minutes, and afterwards she was relieved to have freshened up. 30a/30b. cendel noted down all the important documents his boss requested for the meeting, wishing not to forget any of them. later he searched for them in all the files in his office. this/that took him 2 hours, and afterwards he was happy to be able to get ready for the meeting. 30c/30d. cendel noted down all the important documents his boss requested for the meeting, wishing not to forget any of them. later he searched for them in all the files in his office. this/that took him 5 minutes, and afterwards he was happy to have found them. 31a/31b. banu read her emails and letters, happy that she would leave all the stress, work and family behind her soon. after reading them, she took a camping trip in fethiye. this/that took her 2 days, and afterwards she was pleased to feel strong enough to cope with all the challenges she faced. 31c/31d. banu read her emails and letters, happy that she would leave all the stress, work and family behind her soon. after reading them, she took a camping trip in fethiye. this/that took her 5 minutes, and afterwards she was pleased to start her trip earlier than she expected. 32a/32b. derya bought a map from the newsagents, worried that she was hopelessly lost. then she tried to find her way back to the new company's regional offices. this/that took her 1 hour, and afterwards she was pleased to have arrived in time for the meeting. 32c/32d. derya bought a map from the newsagents, worried that she was hopelessly lost. then she tried to find her way back to the new company's regional offices. this/that took her 3 minutes, and afterwards she was pleased to have found one for only two pounds. 33a/33b. emircan prepared all the wine glasses and plates, guessing the number of people that would come around. then he served four courses to his guests. this/that took him 3 hours, and afterwards he was too exhausted to think about having another dinner party in his flat again. 33c/33d. emircan prepared all the wine glasses and plates, guessing the number of people that would come around. then he served four courses to his guests. this/that took him 20 minutes, and afterwards he was too exhausted to arrange everything on time. 34a/34b. esra experienced a problem at the turkish airlines desk, nervous she would miss her direct flight. then she flew from istanbul to the united states. this/that took her 9 hours, and meanwhile she was very angry with herself for not sleeping on the flight. processing of this and that by turkish l2 speakers 77 34c/34d. esra experienced a problem at the turkish airlines desk, nervous she would miss her direct flight. then she flew from istanbul to the united states. this/that took her 45 minutes, and meanwhile she was very angry with herself for not bringing the credit card that she used to pay for her booking. 35a/35b. asya phoned her bridesmaid to go shopping, wishing to find the dress she would like. then she went shopping for a wedding dress in armada in ankara. this/that took her 5 hours, and afterwards she was pleased since she still had time to schedule her hairdresser and photographer. 35c/35d. asya phoned her bridesmaid to go shopping, wishing to find the dress she would like. then she went shopping for a wedding dress in armada in ankara. this/that took her 10 minutes, and afterwards she was pleased since she would have a day with her bridesmaid. 36a/36b ahu set the oven timer following the instruction manual. then she made a cake with apricots and tarts with strawberries and blueberries for her friend's party. this/that took her 30 minutes, and eventually she was glad to have finished on time for the party. 36c/36d. ahu set the oven timer following the instruction manual. then she made a cake with apricots and tarts with strawberries and blueberries for her friend's party. this/that took her 10 minutes, and eventually she was glad to have understood the manual. 37a/37b. durukan waited for the next train to mersin, feeling tired after a long day in adana. when he got the train, he had to sit with the largest man in turkey. this/that took him 55 minutes, and afterwards he was irritated at having had such a long day. 37c/37d. durukan waited for the next train to mersin, feeling tired after a long day in adana. when he got the train, he had to sit with the largest man in turkey. this/that took him 10 minutes, and afterwards he was irritated about being in the ridiculously crowded train. 38a/38b. asu travelled four stops on the tram, wishing to see the ayasofya museum in istanbul before 6 p.m. then she visited the greatest museum. this/that took her 2 hours, and afterwards she was drained from looking at lots of mosaics. 38c/38d. asu travelled four stops on the tram, wishing to see the ayasofya museum in istanbul before 6 p.m. then she visited the greatest museum. this/that took her 10 minutes, and afterwards she was drained because the train got very crowded. 39a/39b. kenter stood in the footlights, revelling in the audience's applause. then she had dinner with the sonata's composer, musicians and some journalists. this/that took her 2 hours, and afterwards she had never felt so full. 39c/39d. kenter stood in the footlights, revelling in the audience's applause. then she had dinner with the sonata's composer, musicians and some journalists. this/that took her 3 minutes, and afterwards she had never seen so many flowers thrown upon stage. 40a/40b. yurdanur had a long interview for the m.sc programme in biology, hoping she would be accepted. then she visited her cousin and aunt in geneva. this/that took her 15 days, and eventually she was pleased to have time to relax after finishing her undergraduate dissertation. 40c/40d. yurdanur had a long interview for the m.sc programme in biology, hoping she would be accepted. then she visited her cousin and aunt in geneva. this/that took her 25 minutes, and eventually she was pleased to have answered all questions asked to her. çokal, sturt and ferreira 78 references altmann, g. t. m., & kamide, y. (2007). the real-time mediation of visual attention by language and world knowledge: linking anticipatory (and other) eye movements to linguistic processing. journal of memory and language, 57(4): 502-518. https://doi.org/10.1016/j.jml.2006.12.004 anderson, a., garrod, s. c., & sanford, a. j. (1983). the accessibility of pronominal antecedents as a function of episode shifts in narrative text. the quarterly journal of experimental psychology section a, 35(3): 427–440. anderson, s. r., & keenan, e. l. (1985). deixis. language typology and syntactic description, 3(4): 259-308. ariel, m. (1996). referring expressions and the +/coreference distinction. in t. fretheim & k. gundel (eds.), reference and referent accessibility, pages 13-25, amsterdam: john benjamins. bates, d., mächler, m., bolker, b., & walker, s. (2015). fitting linear mixed-effects models using lme4. journal of statistical software, 67(1): 1–48. barr, d. j., levy, r., scheepers, c., & tily, h. j. (2013). random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3):10. blagoeva, r. (2004). demonstrative reference as a cohesive device in advanced learner writing: a corpus-based study. in k. aijmer & b. altenberg (eds.), advances in corpus linguistics. papers from the 23rd international conference on english language research on computerized corpora (icame 23), pages 297–307, amsterdam & new york: rodopi. clahsen, h., & felser, c. (2006). continuity and shallow structures in language processing. applied psycholinguistics, 27(01): 107–126. https://doi.org/10.1017/s0142716406060024 cokal, d. (2021, march 5). l2 processing of d-anaphors. retrieved from osf.io/v2msd çokal, d. (2019). discourse deixis and anaphora in l2 writing. dilbilim araştırmaları, 30(02): 241-277. boğaziçi üniversitesi yayınevi, i̇stanbul. çokal, d. (2010). the pronominal bu-şu and this-that: rhetorical structure theory. dilbilim araştırmaları, 1, 15-33. çokal, d., & ruhi, ş. (2006). discourse anaphors in l2 english writing: implications for researching interlanguage demonstrative systems. paper presented at the 16th annual conference of the european second language association, antalya, turkey. çokal, d., sturt, p., & ferreira, f. (2014). deixis: this and that in written narrative discourse. discourse processes, 51(3): 201–229. doi: 10.1080/0163853x.2013.866484 cornish, f., (2010). anaphora: text-based or discourse-dependent? functionalist vs. formalist accounts. functions of language, 17 (2): 207–241. cornish, f., (2009). indexicality by degrees: deixis, ‘anadeixis’ and (discourse) anaphora. paper presented at a symposium (quel sens pour la linguistique?) organized to mark the conferral of a doctorate honoris causa on sir john lyons, université de toulouse-le mirail, 23–24 april, 2009. dekeyser, r. (2003). implicit and explicit learning. in m. long and c. doughtly (eds.). handbook of second language acquisition, pages 313-348, malden, ma: blackwell. diessel, h. (2006). demonstratives, joint attention, and the emergence of grammar. cognitive linguistics, 17: 463–489. doi: 10.1515/cog.2006.019 ehrlich, k., & rayner, k. (1983). pronouns assignment and semantic integration during reading: eye-movements and immediacy of processing. journal of verbal learning and verbal behavior, 22:75-87. elbourne, p. (2008). demonstratives as individual concepts. linguistics and philosophy, 31:409 466. ellert, m. (2013). resolving ambiguous pronouns in a second language: a visual world eye tracking study with dutch learners of german. international review of applied linguistics in language teaching, 51:71-197. doi: 10.1515/iral-2013-0008 processing of this and that by turkish l2 speakers 79 ellis, r. (2008). the study of second language acquisition. oxford: oxford university press. ellis, n. (2015). implicit and explicit language learning: their dynamic interface and complexity. in rebuschat, p. (ed.). implicit and explicit learning of languages, pages 3-23, amsterdam: john benjamins. enç, m. (1986). “topic switching and pronominal subjects in turkish.” in d. i. slobin and k. zimmer (eds.) studies in turkish linguistics (pp.195–208). amsterdam: john benjamins. erguvanlı-taylan, e. (1986). pronominal versus zero representation of anaphora in turkish. in dan. i. slobin & karl zimmer (eds.). studies in turkish linguistics, pages 209-231, amsterdam: john benjamins. ferris, d. r. (1994). lexical and syntactic features of esl writing by students at different levels of l2 proficiency. tesol quarterly, 28(2): 414. doi: 10.2307/3587446 fossard, m., & rigalleau, f. (2005). referential accessibility and anaphora resolution: the case of the french hybrid demonstrative pronouns: celui-ci/celle-ci. in a. branco, t. mcenry, & r. mitkov (eds.), anaphora processing: linguistics, cognitive and computational modelling, pages 283–300, amsterdam: john benjamins. grosz, j.b., & sidner, l. c. (1986). attention, intention and the structure of discourse. computational linguistics, 12: 175-202. gundel, j. k., hedberg, n., & zacharski, r. (1993). cognitive status and the form of referring expressions in discourse. language, 69(2): 274. halliday, m. a. k., & hasan, r. (1976). cohesion in english. english language series, london: longman. hinkel, e. (2001). matters of cohesion in l2 academic texts. applied language learning, 12(2): 111–32. hopp, h. (2014). working memory effects in the l2 processing of ambiguous relative clauses. language acquisition, 21(3): 250-278. doi:10.1080/10489223.2014.892943 ionin, t., baek, s., kim, e., ko, h., & wexler, k. (2012). that’s not so different from the: definite and demonstrative descriptions in second language acquisition. second language research, 28(1): 69–101. doi: 10.1177/0267658311432200 jin, w. (2001). a quantitative study of cohesion in chinese graduate students’ writing: variations across genres and proficiency levels. retrieved from http://eric.ed.gov/?id=ed452726 juvonen, p. (1996). språkkontakt och språkförändring: om användningen av demonstrativa pronomen i en sverigefinsk kontaktsituation. in l. huss (ed.), många vägar till tvåspråkighet. uppsala: centre for multi-ethnic research, uppsala university. klin, c. m., guzman, a. e., weingartner, k. m., & ralano, a. s. (2006). when anaphora resolution fails: partial encoding of anaphoric inferences. journal of memory and language, 54: 131 143. kornfilt, j. (1997). turkish. london: routledge. küntay, a., & özyürek, a. (2006). learning to use demonstratives in conversation: what do language specific strategies in turkish reveal? journal of child language, 33: 303–320. doi: 10.1017/s0305000906007380 lascarides, a., & asher, n. (2007). frontiered discourse representation theory: dynamic semantics with discourse structure. in h. bunt & r. muskens (eds.), computing meaning, pages 87–124. amsterdam: springer. http://link.springer.com/chapter/10.1007/978-1-4020-5958-2_5 levelt, w. j. (1993). speaking: from intention to articulation (vol. 1). mit press. levinson, s. c. (1983). pragmatics. cambridge university press. linde, c. (1979). focus of attention and the choice of pronouns in discourse. in talmy (ed.), syntax and semantics 12: discourse and syntax, pages 337-354, new york: academic press. lücking, a. (2018). witness-loaded and witness-free demonstratives. atypical demonstratives: syntax, semantics and pragmatics, 568: 255. lyons, j. (1977). semantics, vol. 2. cambridge: cambridge university press. çokal, sturt and ferreira 80 mccarthy, m. (1994). it, this and that. in coulthard, r.a. (ed.), advances in written text analysis, pages, 266-275, london: routledge. niimura, t., & hayashi, b. (1996). contrastive analysis of english and japanese discourse anaphors from the perspective of l1 and l2 acquisition. language sciences, 18(3): 811–834. https://doi.org/10.1016/s0388-0001(96)00049-6 nunberg, g. (1993). indexicality and deixis. linguistics and philosophy, 16, 1-43. o’brien, e. j., raney, g. e., albrecht, j. e., & rayner, k. (1997). antecedent retrieval processes. journal of experimental psychology: learning, memory and cognition, 16: 241-249. özyürek, a. (1998). an analysis of the basic meaning of turkish demonstratives in face-to-face conversational interaction. in santi s. et al. (eds.). oralite et gestualite: communication multimodale interaction, pages 604-614, pars: l’ harmattan. öztürk, b. (2001). turkish as a non-pro-drop language. in e. erguvanlı-taylan, (ed.) the verb in turkish, pages 239–259, amsterdam; philadelphia: john benjamins. özil, ş. & şenöz, c. (1996). türkçe’de bu, şu sözcükleri. x. dilbilim kurultayı bildirileri, 27-39. i̇zmir: ege üniversitesi basımevi. passonneau, r. j. (1993). getting and keeping the center of attention. in m. bates, & r. m. weischedel, (eds). challenges in natural language processing. cambridge: cambridge university press. peeters, d., & krahmer, e., maes, a. (2020). a conceptual framework for the study of demonstrative. https://doi.org/10.31234/osf.io/ntydq polanyi, l. (1986). the linguistic discourse model: towards a formal theory of discourse structure. tr-6409. cambridge ma: laboratories inc. r core team (2020). r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. https://www.r-project.org/. roberts, l., gullberg, m., & indefrey, p. (2008). on-line pronoun resolution in l2 discourse: l1 influence and general learner effects. studies in second language acquisition, 30: 333–357. sağın-şimşek, ç., rehbein, j., & babur, e. (2009). i̇şlevsel edimbilim yöntemiyle metin i̇çinde gösterme alanının i̇ncelenmesi. dilbilim araştırmaları. sorace, a. (2011). pinning down the concept of “interface” in bilingualism. linguistic approaches to bilingualism, 1: 1-33. doi: 10.1075/lab.1.1.01sor turan, ü. d. (1997). metin i̇şaret adılları: bu, şu ve metin yapısı. xi. dilbilim kurultayı bildirileri, 201-212. ankara, ortadoğu teknik üniversitesi. underhill, r. (2001). turkish grammar. cambridge: the mit press. webber, b. (1989). deictic reference and discourse structure. technical reports (cis). retrieved from http://repository.upenn.edu/cis_reports/848 webber, b. l. (1991). structure and ostension in the interpretation of discourse anaphors. language and cognitive processes, 6(2): 107–35. wilson, f., keller, f., & sorace, a. (2009). the antecedent preferences of anaphoric demonstratives and personal pronouns in l2 german. in j. chandlee, m. franchini, s. lord, g-m. rheiner (eds.), proceedings of the boston university conference on language development 33, pages 634-645, somerville, ma: cascadilla press. zwaan, r. a. (1996). processing narrative time shifts. journal of experimental psychology: learning, memory, and cognition, 22(5): 1196-1207. https://doi.org/10.1037/0278 7393.22.5.1196 dialogue & discourse 9(1) (2018): 163-191 doi: 10.5087/dad.2018.106 ©2018 andrea santana, wilbert spooren, dorien nieuwenhuijsen and ted sanders this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). subjectivity in spanish discourse: explicit and implicit causal relations in different contexts andrea santana a.c.santanacovarrublas@uu.nl utrecht institute of linguistic ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands wilbert spooren w.spooren@let.ru.nl radboud universiteit postbus/ po box 9103, nl-6500 hd nijmegen dorien nieuwenhuijsen d.nieuwenhuijsen@uu.nl utrecht institute of linguistic ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands ted sanders t.j.m.sanders@uu.nl utrecht institute of linguistic ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands editor: manfred stede submitted 06/2017; accepted 08/2018; published online 09/2018 abstract corpus-based studies in various languages have demonstrated that some connectives are used preferentially to express subjective versus objective meanings, for example, omdat vs. want in dutch. however, spanish connectives have been understudied from this perspective. moreover, most of the studies of subjectivity have focused on explicit relations and little is known about the subjectivity of implicit coherence relations. in addition, the role that context plays in the meaning and use of causal relations and their connectives is still under discussion. this study analyzes spanish causal explicit and implicit relations in different contexts by carrying out manual text analyses, focused on subjectivity. 360 backward relations marked by three prototypical causal connectives and 120 backward implicit relations were extracted from academic and journalistic contexts. the analytical model is based on an integrative approach to subjectivity. statistical analyses reveal no systematic profiles of spanish connectives in terms of subjectivity. furthermore, a significant three-way interaction between santana, spooren, nieuwenhuijsen and sanders 164 subjectivity, context, and linguistic marking is observed. based on a solid corpus analysis, this study reveals new insights into the expression of subjectivity in spanish discourse relations. keywords: causality, subjectivity, spanish connectives, implicit relations, manual analyses, context 1. introduction when interpreting discourse, experienced language users often infer causal relations between utterances, such as cause-consequence or claim-argument. over the last two decades, the notions of causality and subjectivity have been identified as relevant cognitive principles that organize our knowledge of coherence relations (sanders & spooren, 2015). causality is the implicational meaning (pq) that is inferred between connected discursive segments (sanders et al., 1992; 1993). subjectivity is a complex notion that has been defined from various theoretical approaches in linguistics (lyons, 1977; langacker, 1990; traugott, 1995). building on these insights, we consider subjectivity to be the degree to which the speaker is involved in the construal of the relation (pander maat & sanders, 2000, 2001; pander maat & degand, 2001; degand & pander maat, 2003). by combining the principles of causality and subjectivity we distinguish between subjective and objective causal relations, depending on the presence of a subject of consciousness (soc). this soc is responsible for the causal relation that is constructed between two or more discourse segments. thus, coherence relations defined in terms of domains by sweetser (1990) can also be described in terms of subjectivity. for example, a speech-act relation like (1) is considered as a subjective causal relation since the soc motivates her question by referring to the quality of the movie. it is not possible to interpret this relation without considering the speaker who performs the speech-act. the epistemic relation in (2) is also identified as a subjective causal relation because there is a soc, again the speaker, who does the epistemic reasoning that john must love someone, based on the observation that he came back. both relations differ from the content relation in (3), because in this case there is no soc involved in the relation: the utterance described a causal link in the real world between john’s love and his coming back1. that is why (3) is an example of an objective causal relation. (1) what are you doing tonight, because there’s a good movie on. (2) john loved her, because he came back. (3) john came back because he loved her. coherence relations can, but need not, be marked explicitly through connectives. this paper focuses on backward relations, of the type consequence-cause (q because p). (1), (2) and (3) are backward relations marked explicitly by because, but this connective could be omitted and the causal meaning between the discursive segments can be interpreted perfectly well. connectives are defined as processing instructions on how one part of a text is related to another (sanders & spooren, 2007). language users can choose between one or another connective to express causality, for instance, between because, since and for in english or between omdat and want in dutch. the preference for one connective over another to express certain coherence relations is insightful, if it shows to be a systematic one. for instance, connectives may vary in the extent to which they can express subjectivity. in fact, several corpus-based studies have demonstrated that language users prefer to use some connectives to express subjective meanings, while others specialize in communicating objective meanings. this existing body of research has studied various languages. in dutch, the connectives dus (‘so’), want (‘since/for’) and aangezien (‘since’) have been identified as connectives used to express subjective meanings; omdat (‘because’), daardoor (‘as a result’) and doordat (‘because’) have been recognized as connectives used preferentially to express objective meanings, while daarom (‘that’s why’) occupies an 1 examples are taken from sweetser (1990), who distinguishes between content, epistemic and speech act relations. subjectivity in spanish discourse 165 intermediate position (pander maat & degand, 2001; pander maat & sanders, 2000, 2001; degand, 2001; verhagen, 2005; pit, 2006, 2007; stukker & sanders, 2009; sanders & spooren, 2015). a similar pattern has been identified for german, where denn and da (‘because’) tend to express more subjective meanings, whereas weil (which also mean ‘because’) tends to express more objective meanings (günthner, 1993; keller, 1995; wegener, 2000; pit, 2007; stukker & sanders, 2012). the same situation has been observed in other, typologically less related languages. for example, in french, car (‘for’) and puisque (‘since’) have been recognized as connectives that are used preferentially to express subjective meanings as opposed to parce que (‘because’), that has been identified as a connective that has a more general meaning and a preference to express objective relations (degand & pander maat, 2003; zufferey, 2012); in mandarin-chinese, kĕjian (‘so/therefore’) prefers to express subjective relations, whereas yīn’er (‘as a result’) tends to express objective meanings (li, et al., 2013; wei, et al., 2018). these findings reveal that language users distinguish unconsciously between different types of causality. they also suggest that many languages of the world might have a comparable pattern in terms of subjectivity. that is, language users from various languages use different connectives to express more or less subjective causal relations. there are prominent exceptions, however, such as english because, which can be used to express speech act, epistemic as well as content relations, and the same holds for the connective so (knott & sanders, 1998; sweetser, 1990). the crucial question is whether english is really exceptional, especially because we lack knowledge of many other languages. in a romance language like spanish, there is an extensive body of research on discourse markers and connectives (briz, 1998, 2000; martínez, 1997; martín zorraquino & montolío, 1998; montolío, 2001; portolés, 2001; pons, 1998, 2000; vázquez veiga, 2002; santos río, 2003; garrido rodríguez, 2004; domínguez garcía, 2007; fuentes, 2009; martí sánchez, 2008; garcés gómez, 2014), but in spite of this, the spanish language has been understudied from the perspective of subjectivity (for noticeable exceptions, see pit et al., 1996; blackwell, 2016). this paper presents such a study. in earlier studies of spanish connectives, porque (‘because’), ya que (‘since’) and puesto que (‘given that’) have been identified as prototypical connectives for expressing causality (casado velarde, 1998; montolío, 2001; domínguez garcía, 2007). porque (‘because’) is the most common connective in spanish (vera, 1984; domínguez garcía, 2007). goethals (2002, 2010) indicates that porque (‘because’) differs from other connectives in spanish because of its multifunctional properties: it is the only connective that can (i) fall in the scope of negation; (ii) be the focus of a question; (iii) be modified by any type of adverb; (iv) be included in a cleft sentence; and (v) introduce an answer to a wh-question. in sum, this causal connective seems to be like other subordinating causal conjunctions, like english because. blackwell (2016) indicates that porque in oral narratives have either a semantic or a pragmatic2 reading and, in some cases, they may have both. given these features, it can be deduced that porque (‘because’) is indeed a general connective and, therefore, it can be expected that it is a connective used to express various types of causal relations, both objective and subjective. the spanish connective ya que (‘since’) has been characterized as a specific connective that expresses specialized uses. some studies have claimed that it is used to justify or legitimize the speaker’s acts, to express a collective voice, or to introduce information that is accepted by most of the people; it has also been identified as a connective that is associated with the presence of third singular or plural person (borzi & detges, 2011). other studies have argued that it is a connective that refers to evident information in the context to justify other statements (santos río, 2003). yet other studies have indicated that among its rhetorical functions ya que (‘since’) is specialized in expressing irony (goethals, 2002, 2010). these specific features would suggest that 2 blackwell (2016) refers to semantic and pragmatic relations considering the distinction presented in sanders (1997). these correspond to objective and subjective relations, respectively. see also note 4. santana, spooren, nieuwenhuijsen and sanders 166 ya que (‘since’) is a connective used preferentially for expressing subjective relations. indeed, some studies have suggested this idea by claiming that ya que (‘since’) introduces speech acts of justification in which a conceptualizer expresses a subjective point of view, but without implying that this point of view is also the speaker’s (goethals, 2002, 2010). pit et al. (1996) identified that ya que (‘since’) only occurred in epistemic relations and that the segments associated to this connective presented more evaluations than other connectives. the connective puesto que (‘given that’) has also been identified as a connective that has specific meanings in that it introduces the justification of categorical judgments and deductions (santos río, 2003). some authors have considered it as a connective that is associated with a proposal or suggestion (galán rodríguez, 1999). others have stated that it is a connective that signals subjectivity, co-occurring generally with a speaker in an evaluative role (pit et al., 1996). therefore, these features lead us to infer that it is also used preferentially to express subjective relations. in a recent corpus-based study (santana et al., 2017), we explored the degree of subjectivity of spanish connectives by carrying out automatic analyses. the results showed all three connectives occur with a relatively high frequency. however, the subjectivity profile observed in these connectives did not coincide with the claims in the literature discussed so far. we found that porque (‘because’) tends to occur in a subjective environment, i.e., in a context containing relatively many subjective words, which would suggest that it is used for expressing subjective relations. by contrast, ya que (‘since’) did not occur very often in a subjective environment. so, it is not necessarily a connective that expresses subjective relations. for puesto que (‘given that’) the results were mixed: whereas the segment preceding the connective (s1) was associated with many subjective words, the other segment (s2) contained a relatively low percentage of subjective words, which could indicate that this connective is used to express both subjective and objective relations. as these results (santana et al., 2017) differ from the literature on spanish connectives, this raises the issue of the semantic-pragmatic profile of the causal connectives. are the differences of use between the three connectives so subtle that they do not turn up in an automatic analysis of subjectivity? for this reason, in the current study, we used another method to explore the profile of spanish connectives: we carried out a manual analysis of corpus data. hence, the first research question this paper aims to answer is: to what extent do spanish connectives show a systematic variation in terms of subjectivity in a manual corpus analysis? our hypothesis is that porque (‘because’) is a general connective used for expressing both subjective and objective meanings; we also expect that ya que (‘since’) and puesto que (‘given that’) are specific connectives used predominantly to express subjective meanings. in our study, we annotated coherence relations using an analytical model of subjectivity that involves different subjectivity features analyzed in previous studies (degand & pander maat, 2003; sanders & spooren, 2009, 2015; spooren et al., 2010) (see section 2 for more details). so far, we have focused on coherence relations made explicit by connectives. however, in natural discourse, coherence relations often appear without a connective (taboada, 2006, 2009; taboada & das, 2013); these are so-called implicit relations. still, hardly any studies have concerned themselves with implicit relations (see spooren, 1997; taboada & das, 2013; das & taboada, 2017 for noticeable exceptions), extensive research has concentrated on the study of explicit coherence relations. the main reason is that implicit relations are difficult to recognize (lin et al., 2009), precisely because they do not show clear textual cues to indicate the relation. therefore, little is known about the subjectivity of implicit relations. the past two decades have seen the availability of discourse annotation schemes like the penn discourse treebank (pdtb, prasad et al., 2008), the rhetorical structure theory (rst) treebank (carlson, marcu and okurowski, 2003) and the segmented discourse representation theory (sdrt, asher & lascarides, 2003), which have allowed us to analyze explicit relations and their connectives, but also implicit relations. most available resources contain english data, but gradually similar subjectivity in spanish discourse 167 resources have been created for other languages. the rst spanish treebank3 is a good illustration of this (cunha et al., 2011). this raises the possibility to compare the subjectivity profile of explicit relations to that of implicit relations. now that the resources are available, we think it is very important to analyze implicit relations in order to better understand the notion of subjectivity. therefore, the present study aims to fill this research gap by exploratively formulating the second research question: what differences can be identified between explicit (with connectives) and implicit (no connectives) causal relations in terms of subjectivity? a third aspect that this study investigates is the influence of context in meaning and use of causal relations and their connectives. the issue of whether the semantic profiles of causal connectives can be characterized in terms of subjectivity as a stable characteristic, or whether this subjectivity is a context-dependent characteristic has become a debate within the study of connectives. some studies have suggested that there is a relationship between context and type of connective or type of relation. for instance, sanders et al., (1993) demonstrated that language users recognize objective versus subjective relations4 easily when appropriate communicative contexts were provided. sanders (1997) showed that the interpretation of ambiguous cases of coherence relations was strongly influenced by their descriptive versus argumentative context. furthermore, a corpus study showed how objective relations were predominant in informative texts, whereas subjective relations were predominant in expressive and persuasive texts. zufferey (2012) revealed that the distribution of french connectives varies according to different modalities of texts. in written texts, car (‘because’) is more often used for epistemic relations, whereas parce que (‘because’) prefers content relations. however, the situation varies dramatically in spoken texts since car is absent and parce que is used significantly more often in speech-act and epistemic relations. zufferey et al., (2017) also provided empirical evidence that register is a distinguishing factor between connectives car and parce que. a final example is our previous study, in which we identified a significant relationship between the use of spanish causal connectives and text type (informative versus persuasive/argumentative texts) in journalistic texts and between the use of spanish causal connectives, text type (informative and persuasive/argumentative texts) and domain (education and psychology) in academic texts (santana et al., 2017). other studies, however, have demonstrated how connectives have a robust semanticpragmatic profile in terms of subjectivity, irrespective of the context in which they occur. for example, sanders and spooren (2015) found that dutch connectives omdat (‘because’) and want (‘since/for’) showed a clearly different pattern, irrespective of the media used in the research, i.e. written texts, conversations and chat interactions. zufferey (2012) also identified that the semantic-pragmatic profile of the french connective puisque (‘since’) was clearly subjective, and stable across written and spoken data. apparently, the context is relevant in some cases depending on the text types and the connectives that are analyzed. this situation leads us to the third research question of this paper: what is the relationship between contexts and the meaning and use of causal coherence relations? specifically, we will analyze the subjectivity of coherence relations in two different contexts5: journalistic and academic. these contexts differ in their audiences. while journalistic contexts are oriented to a broad audience (van dijk, 1988; waugh, 1995), academic contexts are more specialized, since they address a specific audience: academics (bhatia, 2002, 2004; hyland, 2009; swales, 1990; silver, 2006). we expect to identify whether 3 available from http://corpus.iingen.unam.mx/rst/corpus.html 4 these relations are named “semantic” and “pragmatic” relations in the original papers (see sanders et al., 1992 and 1993) 5 on the one hand, the term “context” adopted in this paper refers to the source from which different text types were extracted to analyze connectives and coherence relations. on the other hand, the term “text type” refers to the classical distinction between informative and persuasive/argumentative texts. santana, spooren, nieuwenhuijsen and sanders 168 the different contexts influence the meaning and use of causal coherence relations, that is, whether the frequency of causal subjective and objective relations depends on the context in which these coherence relations are produced. our goal is to elucidate whether the realization of subjectivity in spanish coherence relations is an inherent property of certain relations and their preferred connectives, or whether this is a matter of contextual dependency. in comparison with our previous study (santana et al., 2017), we believe that a manual analysis will allow us to examine this issue with greater depth. this paper is organized into four sections. first, we describe the analytical model of subjectivity applied in the manual analysis. secondly, we present the method carried out in this study along with the results of the inter-rater agreement tests. third, we present the results obtained after conducting different statistical analyses. finally, the paper ends with the main discussion and conclusions. 2. an analytical model of subjectivity in this paper, we analyze backwards causal relations in spanish and we make use of the analytical model of subjectivity proposed by sanders & spooren (2015), which decomposes the general notion of subjectivity in its most relevant components. this model unites the subjectivity features introduced in earlier approaches: subjectivity is analyzed in terms of three key features. first, the reference to the speaker (lyons, 1977; traugott, 1995) and the implicit presence of the speaker (langacker, 1990) are recognized as relevant components of subjectivity. for this reason, modality of the q-segment (the consequent in a coherence relation) and the presence of the soc are central variables of this analytical model. secondly, considering that our analysis focuses on causal coherence relations, the nature of coherence relations should also be identified as a relevant aspect in the distinction of subjectivity (sweetser, 1990; sanders et al., 1992, 1993). consequently, domain is another essential variable included in this model, which allows us to classify four types of causal relations. finally, assuming that the presence of the soc plays a relevant role on the distinction of causal relations and that different mental spaces are linked to this soc (sanders et al. 2012), we estimate that the distinction between the author/speakersubjectivity and character-subjectivity is also important to identify subjectivity. thus, the identity of the soc is the fourth variable included in our analytical model. previous studies have used all or some of these components to study coherence relations and connectives in other languages (li, et al. 2013; degand & pander maat, 2003; sanders & spooren, 2009, 2015; spooren, et al. 2010). the current study does not propose a new operationalization of subjectivity. table 1 presents the variables and the respective subjectivity values of this analytical model: variables +……………………... subjectivity values………………………… domain epistemic/speech-act volitional content non-volitional content modality (q) judgement/speech-act mental fact physical fact presence of the soc implicit explicit absent identity of the soc author current speaker character table 1. the analytical model for subjectivity. every variable involves different indicators of subjectivity that represent two or more values. these distinctions correspond to different degrees of subjectivity (traugott, 1995; langacker, 1990). 2.1. domain the category domain operationalizes subjectivity in terms of the nature of causal relations (sweetser, 1990; sanders et al., 1992, 1993). it distinguishes between four types of causal relations in terms of domains (sweetser, 1990). the content domain concerns causal relations that subjectivity in spanish discourse 169 occur in the physical world; this means that one event causes another in the “real world”. this domain is divided into two subtypes: the volitional content domain that involves the presence of a soc who performs an intentional act and the non-volitional content domain that does not implicate a soc at all (stukker, sanders & verhagen, 2008). the epistemic domain concerns causal relations in which a soc is involved, to the effect that the events are related by the speaker’s reasoning. finally, the speech-act domain concerns causal relations in which the events are related by the speaker who performs a speech act. as was mentioned previously, these relations indicate different degrees of subjectivity depending on the presence of the soc: the non-volitional content is assumed as the least subjective domain since there is no soc in the construction of the causal relation; volitional content is more subjective than non-volitional content because there is a soc who performs an intentional act involved in the causal relation; epistemic domain is considered even more subjective because there is a soc who is reasoning, inferencing, concluding, giving an opinion in the causal relation, and the speech-act domain is also a subjective domain because there is a soc corresponding to the speaker who is performing the speech act involved in the causal relation. thus, causal relations can be ordered from least subjective to most subjective, as follows: non-volitional content < volitional content < speech act/epistemic the interpretation of these four domains can be facilitated by using a paraphrase test (sanders, 1997; sanders & spooren, 2015; li et al., 2013) which is presented in table 2 (p and q correspond to the antecedent and the consequent of the causal relation, respectively): domain paraphrase non-volitional content p leads to the physical fact/mental fact that q, and no intention is involved in q volitional content p leads to the intentional physical act/mental act that q speech-act epistemic p leads to question/advice/command/promise that q p leads to the claim/decision/inference/conclusion that q table 2. paraphrase test used in the analysis of domain. applying this paraphrase test, example (4) is a non-volitional content relation because the avalanche leads to the physical fact that the road was blocked, and no intention is involved in this causal relation. moreover, the interpretation of this meaning relation does not require reference to a specific soc, which is absent. (5) is interpreted as a volitional content relation since feeling tired leads to the intentional physical act of going home. there is a soc involved in the causal relation, the speaker, who made an intentional decision. (6) is a speech-act relation because the fact that the neighbors are not at home leads to the request to turn up the radio. the speaker motivates her request by reference to the absence of the neighbors. finally, (7) is classified as an epistemic relation because the observation that the baby is crying leads to the speaker’s conclusion that the baby is hungry6. (4) there had been an avalanche at roger’s pass. as a result, the road was blocked. (5) i went home because i felt tired. (6) why don’t you turn up the radio? the neighbors are not at home. (7) the baby must be hungry, because it is crying. 6 example (4) is extracted from pander maat and sanders (2001), (5) and (7) are extracted from pander maat and sanders (2006), and (6) is extracted from sanders and spooren (2007). santana, spooren, nieuwenhuijsen and sanders 170 2.2. modality (q) another possible indicator of subjectivity is in the modality of the segments connected. in our analysis we consider the modality of backward causal relations, i.e. relations in which the consequent (q) corresponds to the first segment whereas the antecedent (p) to the second segment. particularly, we focus on the first segment because the q is the place where subjectivity can be most manifest in backward relations. the modality or propositional attitude can be of different types (sanders et al, 1992, 1993; sanders & spooren, 2015). in this study modality has four values: physical fact, mental fact, speech-act and judgement. the modality is a physical fact if the segment q describes events or states that take place in the real world and can be observed in the physical world (4); it is a mental fact if the segment q depicts mental states such as personal feelings, mental processes, or psychological activities (5); the modality is a speech-act if the segment q is a general question, a rhetorical question, or an imperative construction (6) and finally the modality is a judgement if the segment q presents an opinion, claim, conclusion or deduction (7). these types of modality (q) show different degrees of subjectivity, depending on the complexity involved in the causal relation (sanders et al., 2012): physical fact indicates the lowest degree of subjectivity because it does not involve a mental process of a soc. mental fact is more subjective than physical fact since it involves a mental state of a soc, but it is less subjective than speech-act and epistemic because these indicators require complex actions that must be carried out by a soc like commanding an action, asking a question, formulating an opinion or drawing a conclusion, for instance. thus, indicators of modality can be ordered from least subjective to most subjective, as follows: physical fact < mental fact < speech act/judgement the interpretation of these indicators of modality (q) can be facilitated by using a paraphrase test, which is presented in table 3: modality (q) paraphrase physical fact the fact/event/situation that q mental fact the decision/thinking/analysis/feeling that q speech-act judgement speaker questions/recommends/promises/commands here and now that q soc comes here and now to the conclusion that q, and an evaluation is involved in q table 3. paraphrase test used in the analysis of modality (q). given this paraphrase test and the previous examples, the modality (q) in example (4) is a physical fact since the fact that the road was blocked is observable in the real world. no mental process is carried out by a speaker. in (5), modality (q) is mental fact because the volitional action of going home was a decision (“i decided to go home”). therefore, a mental process carried by a speaker was necessary to make that decision. in (6), the modality (q) is speech-act because the speaker asks a question here and now (“i ask you here and now to turn up the radio”). so, a mental process carried out by a speaker was mandatory to ask the action to another speaker. finally, the modality (q) in (7) is judgement because the claim that the baby is hungry is a conclusion in the here and now (“i here and now conclude that hunger must be the reason”). 2.3. presence of the soc the presence of the soc distinguishes whether a soc, if present, is explicitly referred to (lyons, 1977; langacker, 1990; traugott, 1995). this variable has three values, which express different subjectivity in spanish discourse 171 degrees of subjectivity depending on the linguistic reference to the soc: absent soc, implicit soc and explicit soc. absent soc is the least subjective because there is no soc involved in the causal relation. consider again the examples (4-7), repeated below. in (4) there is no soc because there is no human intention or motivation in the causal relation. in (5), the presence of the soc is explicit since there is an explicit reference to a soc who decided to go home (the first-person pronoun i). in (6) and (7), the presence of the soc is implicit because there is a soc that performs a speech act and draws a conclusion, respectively, but there is no linguistic signal that provides information about these socs. (4) there had been an avalanche at roger’s pass. as a result, the road was blocked. (5) i went home because i felt tired. (6) why don’t you turn up the radio? the neighbors are not at home. (7) the baby is hungry, because it is crying. implicit socs are the most subjective indicator because there is a soc involved in the causal relation but as there is no linguistic signal referring to that soc, it is not part of the utterance and remains off-stage. explicit soc is more subjective than absent soc because there is a soc involved, but the explicit reference puts the soc on-stage (langacker, 1990), making it objective to a certain degree. consequently, indicators of the presence of the soc can be ordered from least subjective to most subjective, as follows: absent soc < explicit soc < implicit soc 2.4. identity of the soc the identity of the soc is a variable that has three values: character, main author/speaker or current speaker (sanders et al., 2012). the identity of the soc is character when the soc is someone other than the author/speaker. it is main author/speaker when it is the main speaker who establishes the causal relation. finally, it is current speaker when the main speaker is presenting a causal relation for which s(he) is not responsible; generally, a secondary speaker is quoted. those cases in which no soc is identified have been coded as non-applicable (n/a). the identity of the soc in (8) is character because the basketball player is the actor who volitionally apologized to the authorities. it is a third person actor. in (9), the identity of the soc is the author because there is a soc who is offering his/her opinion and it is the main speaker. (s)he is establishing the causal connection between both segments. finally, the identity of the soc in (10) is current speaker ‘primo levi’ because an author (in this case a quoted speaker, the public prosecutor) is quoting the causal construction made by primo levi and quotation marks indicate what he said. (8) the basketball player federico kammerichs, from pamesa valencia, apologized to the authorities of the spanish club because last friday he returned to argentina without permission. (9) there should not be “classes” in a democracy since the title per se implies inequality. (10) the public prosecutor has pointed out that "those words were not pronounced by a common person and they have a decisive influence". "i remind you what primo levi wrote: 'we have a responsibility while we live. we must answer for everything we write, word for word because every word leaves a mark," the public prosecutor has said to a downcast de luca. the different degrees of subjectivity of these indicators are associated with the distance between the soc and the main speaker: character is the least subjective identity because it corresponds to the responsible actor for the causal relation, who differs from the main speaker; santana, spooren, nieuwenhuijsen and sanders 172 the current speaker is more subjective than character since it corresponds to a soc who constructs the causal relation, but it is less subjective than the author because the soc is referred to explicitly by the main speaker; author is the most subjective indicator because it corresponds to the soc who constructs the causal relation and is the main speaker of the causal relation. in this way, indicators of the identity of the soc can be ordered from least subjective to most subjective, as follows: n/a < character < current speaker < author we have now identified the four subjectivity features with their indicators, which constitute the analytical model that has been applied in the manual analysis of coherence relations. in this analysis, each variable is considered a separate variable, even though various empirical studies using this model (li et al., 2013; spooren et al, 2010; sanders & spooren, 2015) have revealed a correlation between some variables, such as presence of the soc and domain of the relations. for instance, non-volitional relations typically do not have a soc. still, we think it is important to analyze corpora in terms of these variables separately, because it is very well possible that subjectivity is linguistically expressed differently in one language than in another. in the section that follows, the method used in this study will be described. 3. method different methodological steps were carried out: sampling of the corpus, data analysis and processing and analyses establishing inter-rater agreement. in this section, this information will be described in detail. 3.1. corpus six data sets with backward causal coherence relations were constructed, four of them corresponding to causal explicit relations and the other two to causal implicit relations. the sets for explicit relations contain relations marked by the three connectives of interest: porque (‘because’), ya que (‘since’) and puesto que (‘given that’), which were extracted from different corpora7. the first and second set correspond to examples selected randomly from the academic and journalistic corpus of spanish constructed in our previous corpus-based study (santana et al., 2017); specifically, the fragments were extracted from essays, research articles and textbooks in education and psychology, and from editorials and news, respectively. the third and fourth set contain examples selected randomly from the corpes xxi8, which is a freely available online corpus created by the royal spanish academy (rae in spanish). we decided to use this resource instead of others because it enabled us to select the required number of cases for each connective. for example, in the case of rst spanish treebank, we explored the distribution of those relations that could be considered as causal relations (cause, justification, result, motivation and evidence) and we noticed that the frequencies of the connectives of interest were restricted to few cases. moreover, corpes xxi allowed us to consult cases of these connectives using different search parameters. the search parameter for the third set was ‘academic texts’; the fragments stem mainly from textbooks, research articles, essays, proceedings, theses and research reports from different disciplines. for the fourth set, the search parameters were ‘news’, ‘editorials’ and ‘opinion articles’; the resulting fragments stem from these journalistic text types. table 4 shows the distribution of these explicit relations: 7 the random selections for the first, second, third and fourth set of causal coherence relations were carried out using a sequence generator available from http:www.random.org/ 8 available from http://web.frl.es/corpes/view/inicioexterno.view http://web.frl.es/corpes/view/inicioexterno.view subjectivity in spanish discourse 173 connective academic sub-corpus journalistic sub-corpus academic corpes xxi journalistic corpes xxi total porque (‘because’) 30 30 30 30 120 ya que (‘since’) 30 30 30 30 120 puesto que (‘given that’) 30 30 30 30 120 total 90 90 90 90 360 table 4. distribution of explicit causal relations selected from each corpus. table 4 shows an identical frequency of examples for every connective. however, originally it was not possible to extract enough cases for ya que (‘since’) and puesto que (‘given that’) from the academic and journalistic corpus of spanish9 (santana et al., 2017) since their frequencies did not reach 30. given the purpose of describing the semantic-pragmatic profile of spanish connectives, it was necessary to obtain the same number of examples per connective. to achieve this purpose, new texts were added to the corpus considering the same sources settled in its original construction: one essay of psychology, one essay of education, one research article of psychology, ten editorials and ten news texts. a specific segmentation was considered for each selected case: previous context (c1), which is the clause10 that precedes the causal relation; segment 1 (s1), which is the clause that functions as q (consequence) in the causal relation; the causal connective (porque, ‘because’, ya que ‘since’, puesto que ‘given that’); segment 2 (s2), which is the clause adjacent to s1 and functions as p (cause) in the causal relation; finally, posterior context (c2), which is the clause that follows the causal relation established between s1 and s2. analysts analyzed exclusively the causal relation that exists between s1 and s2; c1 and c2 were presented just to provide more information to the analysts, which could lead to a better classification. fragment (11) illustrates the examples included in the data sets of explicit relations: (11) (c1) desde 2012 cuando se superaron los tres millones de peregrinos, arabia saudí ha limitado el número de asistentes debido tanto a las controvertidas obras que se realizan en la gran mezquita de la meca como al temor a epidemias. (s1) pero se trata de una medida temporal, ya que (s2) el objetivo de los trabajos es ampliar la superficie de la aljama en 400.000 metros cuadrados, (c2) para que pueda acoger hasta 2,2 millones de fieles a un tiempo. ‘(c1) since 2012 when there were more than three million of pilgrims, saudi arabia has limited the number of attendees because of the controversial works performed in the grand mosque of mecca and the fear of epidemics. (s1) but it is a temporary measure, since (s2) the purpose of the works is to expand the area of the mosque in 400,000 square meters (c2), so that it can accommodate up to 2.2 million of faithful at a time.’ as can be noticed in (11), the clause was considered as the minimal unit of analysis for s1 and s2. however, we also considered the propositional content, this is the meaning that underlies every segment, and which allows us to construct a mental representation of the causal relation. thus, the selected examples vary in terms of clauses, some of them containing only one clause in 9 the frequencies of backwards cases of porque (‘because’), ya que (‘since’) and puesto que (‘given that’) in academic texts were 98, 106 and 31, respectively (273,359 words); and in journalistic texts, the frequencies were 79, 22 and 14, respectively (175,466). 10 ‘clause’ was defined here as “a structure that contains a predicate headed by a main verb”. santana, spooren, nieuwenhuijsen and sanders 174 s1 and s2, respectively, as in (11), and others containing more than one clause in s1 or s2. in (12), s1 consists of a complex of hierarchically ordered clauses (indicated by lowercase ‘c’), but ya que (‘since’) is connecting two independent and adjacent segments, which was the main criterion assumed for all the selected examples: (12) (s1) llama la atención la forma [c1 como lograron [c2 ascender en la sociedad colombiana [c3 a pesar de no encontrar un marco institucional [c4 que favoreciera su integración debido a políticas restrictivas hacia los extranjeros [c5 cuya procedencia no fuera europea c5] c4] c3] c2] c1], ya que (s2) la postura gubernamental era recibir a manos llenas la influencia del viejo continente a través de sus hijos. (c2) sin embargo las leyes, aunque abundantes en este campo, resultaban ineficientes a pesar de haberse redactado varios proyectos, pues las cifras de inmigrantes europeos que decidieron establecerse en el país durante la emigración masiva fueron mínimas por el poco atractivo económico que revestía la nación. ‘(s1) it is striking [c1 how they managed [c2 to ascend in colombian society [c3 in spite of the fact that they did not find an institutional framework [c4 that favored their integration due to restrictive policies towards foreigners [c5 whose origin was not european c5] c4] c3] c2] c1], since (s2) the governmental position was to receive generously the influence of the old continent through their children. (c2) however, the laws, although abundant in this field, were inefficient in spite of the fact that several projects had been drafted, since the numbers of european immigrants who decided to settle in the country during the mass emigration were minimal because of the lack of economic attractiveness of the nation.’ regarding the sets for implicit relations, the cases were extracted from the academic and journalistic corpus of spanish constructed in our previous corpus-based study (santana et al., 2017). to this effect, academic and journalistic text types were randomly selected from the corpus and examined to identify the implicit causal relations. this process was repeated until 60 implicit relations were identified for each context, which was identical to the number collected by each connective in the explicit sets (see table 4). table 5 shows the distribution of these implicit relations for every corpus11: academic sub-corpus journalistic sub-corpus total implicit causal relations 60 60 120 table 5. distribution of implicit causal relations selected from each corpus. implicit relations were segmented in the same way as explicit relations, with the exception that these relations did not contain a connective linking s1 and s2. fragment (13) illustrates the examples included in the data sets of implicit relations: (13) (c1) 3.2. concepto de cine de animación (s1) el cine es una ilusión. (s2) las imágenes en movimiento no existen. (c2) se necesita la existencia de 24 imágenes fijas por segundo para crear esa sensación de movimiento. ‘(c1) 3.2. concept of animation cinema 11 84 academic texts (180,955 words) and 215 journalistic texts (100,647 words) were analyzed to identify 120 examples. subjectivity in spanish discourse 175 (s1) cinema is an illusion. (s2) moving images do not exist. (c2) the existence of 24 still images per second is required to create that feeling of movement. 3.2. manual analysis the manual text analysis consisted of coding the data sets described in section 3.1 by evaluating the variables of the analytical model explained in section 2. this task was carried out initially by three native speakers of spanish, one of the authors of this paper and two collaborators. all of them work in the field of discourse studies at different universities. the analysis proceeded in three phases. the first was the ‘warming-up phase’ aimed at training the analysts. in this phase, a protocol of analysis was presented (see appendix 1), containing information about the purpose of the analysis, explaining the categories of the analytical model and the procedures that should be carried out. the analysts also examined a set of examples (20 cases per each one) in order to practice with the categories and discuss problematic cases. these examples were extracted from the same corpora used for the creation of the data sets. four cases corresponded to implicit relations, six to relations marked by porque (‘because’) relations, five to relations marked by ya que (‘since’), and five to relations marked by puesto que (‘given that’). table 6 illustrates one of these examples: item example domain modality presence of the soc identity of the soc observations 17 (c1) ciertamente, el concepto de deconstrucción no era empleado por aguilar indistintamente; como se anotó con anterioridad, en la experiencia con las obras debía sugerirse el modulador adecuado para su apreciación y comprensión. (s1) por otra parte, las connotaciones particulares que el concepto cobró en el discurso crítico de aguilar están íntimamente relacionadas con la especificidad de las obras que fueron observadas mediante esta plataforma conceptual, puesto que (s2) aguilar se remitió a la deconstrucción como un concepto táctico para explicar ciertas características de orden formal. (c2) existen dos artículos clave para ilustrar este fenómeno: "construcciones privadas en la u.n.: armar y desarmar en un lienzo" y "salas silva en el museo: un exquisito 'déjà vu'". ‘(c1) certainly, the concept of deconstruction was not used by aguilar interchangeably; as noted previously, in the experience with the works, the appropriate modulator should be suggested for its appreciation and understanding. (s1) on the other hand, the particular connotations that the concept took in the critical aguilar’s speech are intimately related to the specificity of the works that were observed through this conceptual platform given that (s2) aguilar referred to deconstruction as a tactical concept to explain certain epistémico ‘epistemic’ juicio? ‘judgement?’ implícito ‘implicit’ autor ‘author’ dominio: en casos como este me cuesta determinar si es realmente un estado de cosas del mundo físico observable o si se trata de una interpretación, juicio, valoración, conclusión del autor. tiendo a creer que es epistémico, ya que con el parafraseo no suena lógica una interpretación no volitiva ‘domain: in cases like this one, it is hard for me to determine if it is a state of the physical and observable world or if it is an author’s interpretation, judgement, evaluation or conclusion. i tend to think that it is epistemic since by using the paraphrase test a volitional interpretation does not sound logical.’ santana, spooren, nieuwenhuijsen and sanders 176 formal characteristics. (c2) there are two key articles to illustrate this phenomenon: “construcciones privadas en la u.n.: armar y desarmar en un lienzo” and “salas silva en el museo: un exquisito 'déjà vu”.’ table 6. example of codebook used in the manual text analysis. table 6 shows the elements that were included in the codebook of the manual text analysis. in the example column, the case to be analyzed was presented, in the domain, modality, presence of the soc and identity of the soc columns, annotators identified the values corresponding to such variables and in the observation column, annotators included a comment, question or doubt related to each case. in this way, every case was discussed between annotators, and they agreed on how to analyze the examples, especially the most problematic cases. the second phase was the preliminary manual analysis of data sets, which focused on applying the analytical model and clarifying doubts about the categories. this phase was considered important because it is assumed that the quality of coding improves over time (spooren & degand, 2010). thus, we ensured the correct interpretation of the variables and the codebook before carrying out the official analysis. in this phase, the type of analysis that was carried out is called ‘partial overlap coding’, which implies that the sample is cut up in several subparts of which some are double coded, while others are coded by only one analyst (spooren & degand, 2010; li, et al., 2013). considering that we counted on three analysts and our corpus contained 480 examples distributed in explicit and implicit relations, 15% of the sample was coded by analyst 1 (72 cases); another 15% was coded by analyst 2 (72 cases); and finally, all the examples evaluated by analyst 1 and 2 (144 cases) were coded by analyst 3. these examples were similar to those illustrated in table 6. after carrying out this preliminary round of analysis, several disagreements were identified between annotators, conflicting cases were discussed and criteria for the official analysis were agreed on. this discussion was carried out separately between the analysts, so the examples analyzed by analysts 1 and 2 were not shared between them; analyst 3 was the only one who had access to all the examples. the third phase corresponded to the final manual analysis of the data. here we used the same type of analysis ‘partial overlap coding’ but unlike the second phase, the analyses were done with two rather than three analysts. in this way, one analyst examined 15% of the sample (72 cases), which corresponded to different examples to those given in the preliminary round of analyses and the other analyst examined all the examples of the sample (480 cases). these examples were comparable to those presented in the previous phases (see table 6). in total, considering all phases of the analysis process (the warming-up phase, the preliminary manual analysis, and the final manual analysis of the data), three rounds of annotations were carried out. different, ambiguous and clear examples were discussed in the first and second phases, so the annotator who analyzed most of the sample (480 cases) in the third phase, followed the criteria established as a result of previous annotations. therefore, annotations in the third phase do not show as much variation as the annotations carried out in the previous phases. this method ensured that at least 15% of the data was double-coded, it also enabled us to use the information of double-coded data in order to conduct statistical analyses and it made possible to enhance the reliability of data coded by one analyst. 3.3. inter-rater agreement the level of inter-rater agreement was calculated taking into account the double-coded data collected in the third phase of our analysis, which correspond to 72 cases of the sample (15% of the total analyzed corpus). this selection consisted of 36 cases extracted from the academic context and 36 cases from the journalistic context. all of them were extracted randomly. subjectivity in spanish discourse 177 regarding the cases of the academic context: five essays of education, two essays of psychology, six research articles of education, three research articles of psychology, four textbooks of education and four textbooks of psychology were extracted from the academic corpus of spanish (santana et al., 2017) and four research articles, three textbooks, two essays, one proceeding, one research report and one thesis were obtained from corpes xxi. in relation to the cases of the journalistic context: twelve editorials and eight news were extracted from the journalistic corpus of spanish (santana et al., 2017) and seven editorials, five news and four opinion articles were collected from corpes xxi. table 7 shows the inter-rater agreement: variable percentage of agreement cohen´s kappa n cases domain 90.3% 0.68 72 modality (q) 65.3% 0.42 72 presence of the soc 93.1% 0.84 72 identity of the soc 94.4% 0.85 72 table 7. inter-rater agreement per variable among annotators. it can be seen from the data in table 7 that the percentage of agreement was over 90% in three of the variables. moreover, the indicator of cohen’s kappa shows that the agreement was excellent for the identity of the soc and for the presence of the soc (following the interpretation rules for kappa suggested by landis and koch, 1977). in the case of the domain, the level of agreement can be labeled substantial. the agreement was low in case of modality (q), which reveals that it is the most controversial category of the analytical model. for further detail, tables 8, 9, 10 and 11 provide the confusion matrices of each subjectivity feature12: domain analyst 3 analyst 1 epistemic speech-act volitional non-volitional epistemic 56 3 1 0 non-volitional 1 0 0 1 volitional 2 0 8 0 table 8. confusion matrix of domain. modality (q) analyst 3 analyst 1 judgement speech-act mental fact physical fact judgement 30 0 0 0 speech-act 2 3 0 0 mental fact 17 0 13 0 physical fact 0 0 6 1 table 9. confusion matrix of modality (q). 12 matrices are calculated on the basis of the total number of double-coded data, i.e. 72 cases of the sample (15% of the total analyzed corpus). santana, spooren, nieuwenhuijsen and sanders 178 presence of the soc analyst 3 analyst 1 implicit explicit absent implicit 48 2 0 explicit 3 18 0 absent 0 0 1 table 10. confusion matrix of the presence of the soc. identity of the soc analyst 3 analyst 1 author current speaker character n/a author 54 0 0 0 current speaker 1 1 0 0 character 1 2 12 0 n/a 0 0 0 1 table 11. confusion matrix of the identity of the soc. regarding the disagreements in modality (q), most of them corresponded to cases that were classified as judgement by one analyst and mental fact by the other (see table 9). this situation illustrates the difficulty of distinguishing these indicators, especially because both involve mental processes that are not directly observable in the external world. once we obtained the inter-rater agreement, statistical analyses were conducted. to investigate the distribution of subjectivity features, log-linear analyses were used specifically to identify associations and interactions between variables (subjectivity, linguistic marking, context), which in case of three-way interactions were followed up by chi-square tests (following the strategy suggested by field, 2012). in the section that follows, this data will be presented in detail. 4. results the present study aims to investigate whether spanish connectives show a systematic variation in terms of subjectivity in a manual corpus analysis, whether explicit and implicit relations differ with respect to subjectivity, and whether these relationships depend on context. as was mentioned in section 1, we assumed academic and journalistic as sources of context. in this section, the results are reported for each subjectivity feature and the observed data correspond to the analyses carried out by one of the annotators who analyzed all the 480 cases in the corpus. as indicated previously, log-linear analyses were carried out with the purpose of identifying associations and interactions between variables (subjectivity, linguistic marking, context). to achieve this aim, it was necessary to collapse some categories. this decision was made as a solution to the problem of having more than 20% of cells with expected frequencies less than 5, which violates one of the assumptions in the log-linear analysis (see field, 2012). grouping was done taking into account categories that have similar degrees of subjectivity. table 12 illustrates the organization of categories: variables subjective objective domain epistemic/speech-act volitional content/non-volitional content modality (q) judgement/speech-act mental fact/physical fact presence of the soc implicit explicit/absent identity of the soc author/current speaker character/n/a table 12. organization of categories for the statistical analyses. subjectivity in spanish discourse 179 table 12 shows that in the case of the domain, speech-act and epistemic relations were grouped together as subjective relations, whereas volitional and non-volitional as objective. in the case of modality of the q-segment, speech-act and judgement were gathered as subjective and mental fact and physical fact as objective indicators. in the case of the presence of the soc, implicit soc was considered subjective and explicit soc and absent soc were grouped as objective indicators. finally, in the case of the identity of the soc, author and current speaker were grouped as subjective cases, whereas character and n/a were considered as objective ones. 4.1. domain the log-linear analyses produced a final model that retained all effects. the fit of the model was perfect (χ2 (0) = 0; p=1). this indicated that the highest order interaction of domain * context * linguistic marking was significant (χ2 (3) = 15.41; p <.01). to interpret this interaction separate chi-square tests were calculated between domain and linguistic marking for each source of context. for academic context, the association between domain and linguistic marking did not reach significance (χ2 (3) = 6.49; p = .09). for journalistic context, the association was significant (χ2 (3) = 16.71; p <.01). analysis of the standardized residuals showed that the association is mainly caused by the relatively low number of objective indicators (volitional and non-volitional) for implicit relations, compared to the three types of explicit relations. the data are summarized in table 13: domain subjective objective linguistic marking academic context porque (‘because’) 46 (-0.7) 14 (1.6) ya que (‘since’) 51 (0.0) 9 (0.0) puesto que (‘given that’) 56 (0.7) 4 (-1.7) implicit 50 (-0.1) 10 (0.2) journalistic context porque (‘because’) 43 (-0.8) 17 (1.7) ya que (‘since’) 45 (-0.5) 15 (1.1) puesto que (‘given that’) 48 (-0.1) 12 (0.2) implicit 59 (1.5) 1 (-3.1) note: subjective and objective correspond to the total number of the collapsed categories in the variable domain, speech-act/epistemic and volitional/non-volitional, respectively. table 13. frequencies (and standardized residuals) of linguistic marking by domain for each source of context (academic, journalistic). overall13, 84.5% of the 240 relations in academic context were subjective (speechact/epistemic) and 15.4% were objective (volitional/non-volitional). a similar predominance of subjective relations was found in journalistic context, 81.2% subjective (speech-act/epistemic) and 18.7% objective (volitional/non-volitional). 4.2. modality (q) the log-linear analyses also produced a final model that retained all effects. the fit of the model was perfect (χ2 (0) = 0; p=1). this indicated that the highest order interaction of modality (q) * context * linguistic marking was significant (χ2 (3) = 9.71; p <.01). to interpret this interaction separate chi-square tests were calculated between modality (q) and linguistic marking for each source of context. for academic context, the association between modality (q) and linguistic 13 all percentages presented in section 4 are calculated on the basis of the total number of relations in each source of context (240 relations in academic contexts and 240 relations in journalistic contexts). santana, spooren, nieuwenhuijsen and sanders 180 marking did not reach significance (χ2 (3) = 2.87; p = .41). for journalistic context, the association was significant (χ2 (3) = 14.41; p <.01). analysis of the standardized residuals showed that the association is mainly caused by the low number of objective indicators (mental fact and physical fact) for the implicit cases, compared to the three types of explicit cases. the data are summarized in table 14: modality (q) subjective objective linguistic marking academic context porque (‘because’) 40 (-0.7) 20 (1.1) ya que (‘since’) 45 (0.1) 15 (-0.1) puesto que (‘given that’) 48 (0.5) 12 (-0.9) implicit 45 (0.1) 15 (-0.1) journalistic context porque (‘because’) 40 (-0.6) 20 (0.9) ya que (‘since’) 39 (-0.7) 21 (1.2) puesto que (‘given that’) 41 (-0.4) 19 (0.7) implicit 55 (1.7) 5 (-2.8) note: subjective and objective correspond to the total number of the collapsed categories in the variable modality (q), speech-act /judgement and mental/physical facts, respectively. table 14. frequencies (and standardized residuals) of linguistic marking by modality (q) for each source of context (academic, journalistic). overall, 74.1% of the 240 q-segments in academic context had a subjective modality (speech-act/judgement) and 25.8% had an objective modality (mental/physical facts). the same predominance of subjective modalities was also found in journalistic context, 72.9% subjective modality, 27.0% objective modality. 4.3. presence of the soc similar to domain and modality (q), the log-linear analyses produced a final model that retained all effects. the fit of the model was perfect (χ2 (0) = 0; p=1). this indicated that the highest order interaction of presence of the soc * context * linguistic marking was significant (χ2 (3) = 27.93; p <.01). to interpret this interaction separate chi-square tests were calculated between presence of the soc and linguistic marking for each source of context. for academic context, the association between presence of the soc and linguistic marking was significant (χ2 (3) = 9.16; p <.01). for journalistic context, the association was also significant (χ2 (3) = 29.46; p <.01). analysis of the standardized residuals showed that the association in academic context is mainly caused by the relatively low number of objective cases (explicit and absent socs) for explicit relations marked with puesto que (‘given that’), compared to the explicit relations marked by the other connectives and implicit relations. in the case of journalistic context, the association is mainly caused by the high number of subjective indicators (implicit soc) and the low number of objective indicators (explicit and absent socs) for implicit relations, compared to the three types of explicit relations. the data are summarized in table 15: presence of the soc subjective objective linguistic marking academic context porque (‘because’) 44 (-0.6) 16 (1.2) ya que (‘since’) 46 (-0.3) 14 (0.6) puesto que (‘given that’) 56 (1.2) 4 (-2.3) implicit 46 (-0.3) 14 (0.6) subjectivity in spanish discourse 181 journalistic context porque (‘because’) 35 (-0.1) 25 (1.5) ya que (‘since’) 34 (-1.2) 26 (1.7) puesto que (‘given that’) 39 (-0.4) 21 (0.6) implicit 58 (2.6) 2 (-3.8) note: subjective and objective correspond to the total number of the collapsed categories in the variable presence of the soc, implicit soc and explicit/absent soc, respectively. table 15. frequencies (and standardized residuals) of linguistic marking by presence of the soc for each source of context (academic, journalistic). overall, 80.0% of the 240 relations in academic context had a subjective soc (implicit soc) and 20.0% were objective (explicit/absent soc). the percentages for journalistic context were 69.2% subjective and 30.8% objective. 4.4. identity of the soc as with the previous variables, the log-linear analyses produced a final model that retained all effects. the fit of the model was perfect (χ2 (0) = 0; p=1). this indicated that the highest order interaction of identity of the soc * context * linguistic marking was significant (χ2 (3) = 19.94; p <.01). to interpret this interaction, separate chi-square tests were calculated between identity of the soc and linguistic marking, for each source of context. for academic context, the association between identity of the soc and linguistic marking was significant (χ2 (3) = 10.67; p <.01). for journalistic context, the association was also significant (χ2 (3) = 21.69; p <.01). analysis of the standardized residuals showed that the association in academic context is mainly caused by the low number of objective cases (character and n/a) for explicit relations marked with puesto que (‘given that’), compared to the explicit relations marked by the other connectives and implicit relations. in journalistic contexts, the association is mainly caused by the high number of subjective indicators (author and current speaker) and the low number of objective ones (character and n/a) for implicit relations, compared to the three types of explicit relations. the data are summarized in table 16: identity of the soc subjective objective linguistic marking academic context porque (‘because’) 49 (-0.4) 11 (1.1) ya que (‘since’) 48 (-0.6) 12 (1.4) puesto que (‘given that’) 59 (1.0) 1 (-2.5) implicit 52 (0.0) 8 (0.0) journalistic context porque (‘because’) 38 (-0.1) 22 (1.7) ya que (‘since’) 40 (-0.7) 20 (1.2) puesto que (‘given that’) 43 (-0.3) 17 (0.4) implicit 58 (2.0) 2 (-3.4) note: subjective and objective correspond to the total number of the collapsed categories in the variable identity of the soc, author/current speaker and character/n/a, respectively. table 16. frequencies (and standardized residuals) of linguistic marking by identity of the soc for each source of context (academic, journalistic). regarding the total percentages, 86.6% corresponded to subjective cases (author/current speaker) and 13.3% to objective ones (character/n/a) in academic context. in the case of santana, spooren, nieuwenhuijsen and sanders 182 journalistic context, 74.5% corresponded to subjective indicators (author/current speaker) and 34.0% to objective indicators (character/n/a). 5. discussion and conclusions the current study aimed to achieve a better understanding of subjectivity in spanish causal relations and connectives by carrying out manual text analyses. evidence of several corpus-based studies in different languages (pander maat & degand, 2001; pander maat & sanders, 2000, 2001; degand, 2001; verhagen, 2005; pit, 2006, 2007; stukker & sanders, 2009; sanders & spooren, 2015; günthner, 1993; keller, 1995; wegener, 2000; pit, 2007; stukker & sanders, 2012; degand & pander maat, 2003; zufferey, 2012; li, et al., 2013; wei, et al., 2018) have demonstrated that causal connectives follow a specific pattern in terms of subjectivity: some causal connectives are mainly used to express subjective meaning, whereas others are used to express objective meanings. a priori, spanish might be expected to follow such a pattern. however, our previous corpus-based study (santana et al., 2017), in which automatic analyses were carried out, did not reveal this. this situation led us to question whether the results were due to the method of automatic analysis of subjectivity. that is why we decided to use another method to shed light on this matter: manual analyses of subjectivity. local contexts of spanish causal explicit relations marked by the connectives porque (‘because’), ya que (‘since’) and puesto que (‘given that’) were analyzed in order to investigate whether these connectives show a systematic variation in terms of subjectivity. based on previous literature of spanish connectives (borzi and detges 2011; goethals, 2002, 2010; pit et al., 1996; blackwell, 2016), we expected to find that porque (‘because’) is a general connective used to express subjective and objective relations, whereas ya que (‘since’) and puesto que (‘given that’) are specific connectives used to express subjective meanings. a second issue concerned implicit relations. most available studies on coherence relations have investigated the subjectivity of explicit relations; hardly any investigations exist on the subjectivity of implicit relations. therefore, the local contexts of implicit relations were examined in order to identify differences between this type of relations and explicit relations. our purpose was to explore whether subjectivity occurs differently in the absence of connectives. these relations were analyzed in different contexts (academic and journalistic) in order to identify whether the subjectivity of these relations depends on the context in which they occur or whether this subjectivity is stable across these contexts. log-linear analyses were conducted to identify associations and interactions between variables (subjectivity, linguistic marking, context) and chisquare tests were applied to obtain a more fined-grained analysis of the interaction between variables. the first research question was: to what extent do spanish connectives show a systematic variation in terms of subjectivity in a manual corpus analysis? results of this corpus-based study showed that there is no evidence for subjectivity as a categorical principle in spanish causal connectives. the connectives observed did not reveal a systematic variation in terms of subjectivity. porque (‘because’) and ya que (‘since’) were not associated with any subjectivity feature, which suggests that they do not have a profile that can be characterized in terms of more or less subjectivity. puesto que (‘given that’) was associated with a low number of objective features, which could lead us to infer that it is not used to express objective meanings. however, this tendency was observed only in academic context and for only two features: the presence of the soc (explicit and absent socs) and identity of the soc (character and n/a). therefore, we conclude that in comparison with other languages, the distinction between connectives expressing subjective versus objective relations does not occur so evidently in spanish. our second research question was: what differences can be identified between explicit (with connectives) and implicit (no connectives) causal relations in terms of subjectivity? the results showed that implicit relations were associated with relatively low number of objective scores subjectivity in spanish discourse 183 across the analyzed variables (domain, modality (q), presence of the soc and identity of the soc). moreover, implicit relations were associated with high number of subjective indicators in two variables: presence of the soc (implicit soc) and identity of the soc (author/current speaker). however, these results were observed only in journalistic context. in the case of academic context, implicit relations were not associated with any subjectivity feature. therefore, these observed data lead us to conclude that implicit relations behave differently depending on the context that they occur in. more specifically, in journalistic context they tend to be non-objective and in the case of containing a soc, implicit relations are used preferentially to express subjective meanings. our third research question was: what is the relationship between contexts and the meaning and use of causal coherence relations? results revealed that there was a significant three-way interaction in all the observed variables, which would indicate that context is relevant for the relation between subjectivity and linguistic marking. by splitting the results per source of context and by carrying out a more fine-grained analysis, we identified that associations between subjectivity and linguistic marking were significant in all variables in journalistic context (domain, modality (q), presence of the soc and identity of the soc), whereas they were significant only in two variables in the case of academic context (presence of the soc and identity of the soc). these results suggest that the relevance of context in the relation between subjectivity and linguistic marking is clearer in journalistic context than in academic context. in conclusion, the three research questions of this study received an explicit, and sometimes surprising, answer. regarding our first research question, the analyzed spanish connectives do not show a systematic variation in terms of subjectivity as various other languages do (such as dutch, german, french, mandarin chinese). porque (‘because’), ya que (‘since’) and puesto que (‘given that’) do not have a clear-cut semantic profile in terms of subjectivity. this conclusion is based on solid manual text analyses considering a substantial number of relations and including statistical evaluation. the question arises why our results differ from previous studies of spanish connectives. for ya que (‘since’) was identified by borzi and detges (2011) as a connective that is associated with the presence of a third singular or plural person. goethals (2002, 2010) indicates that ya que (‘since’) identifies a speech act of justification in which a conceptualizer expresses a particular or subjective point of view, but without implying that this point of view is also the speaker’s. according to these studies, we would have expected that ya que (‘since’) would be related to some indicator of presence of the soc and identity of the soc, but our data did not reveal that. in the case of puesto que (‘given that’) pit et al. (1996) claim that this connective signals subjectivity co-occurring generally with a speaker in an evaluative role, which would lead us to expect that puesto que (‘given that’) would be associated to judgement in modality (q). however, in our study, this connective was not associated with subjective indicators at all. these discrepancies could be attributed to the fact that our sample was selected randomly considering different texts per each source of context (essays, research articles, textbooks, proceedings, theses and research reports from different disciplines for academic context, and editorials, news and opinion articles for journalistic context), whereas texts that were analyzed in previous studies were selected more narrowly. for example, the corpus used by borzi and detges (2011) consisted of examples extracted from journalistic articles of argentinian newspapers and from five essays that corresponded to three argentinian authors. similarly, goethals (2002) used cases of ya que (‘since’) only from the section of politics of the newspaper el país and pit et al. (1996) analyzed 25 cases for ya que (‘since’) as well as for porque (‘because’) and puesto que (‘given that’). furthermore, our results differed from previous studies because of the analytical model that was used. where we used an integrative approach to subjectivity that decomposes the general notion of subjectivity in its most relevant components (lyons, 1977; langacker, 1990; traugott 1995; degand & pander maat, 2003; sanders & spooren, 2009, 2015; spooren et al., 2010), santana, spooren, nieuwenhuijsen and sanders 184 previous studies considered some specific aspects of subjectivity. borzi and detges (2011) described ya que (‘since’) as polyphonic connective, similar to the french connective puisque, assuming the proposal of bourcier and ducrot (1980) and they examined the position and the quality of the information associated to these connectives; goethals (2002) described and classified spanish causal conjunctions within a framework that combines elements of the semiotics and speech act philosophy. finally, pit et. al. (1996) analyzed causal relations considering the content/epistemic domain dichotomy proposed by sweetser (1990) and applying a schema of perspective analysis. we do believe that subjectivity was operationalized in the current study in a detailed and precise way, with a model that has revealed systematic differences in other languages (sanders & spooren, 2015; li et al, 2013). as a result, the conclusion on the lexicon of causal connectives in spanish seems valid. taken together, these results provide new insights into the categorization of coherence relations in terms of subjectivity. we obtained empirical evidence that shows that the phenomenon of subjectivity is encoded differently across languages. spanish, similar to english, does not have connectives that have a clear subjectivity profile. rather, spanish seems to have a wider repertoire of general causal connectives, like because (sweetser, 1990; ford, 1993; knott & sanders, 1998), which can be used to express both subjective and objective relations. in this way, subjectivity in spanish might be conveyed through other linguistic elements. therefore, it would be necessary to look into other specific features in the context. for example, levshina and degand (2017) analyzed several contextual variables to identify objective and subjective meanings expressed by english because on the basis of uses associated to dutch connectives omdat and want; the presence or absence of modal verbs, polarity (positive and negative clauses), semantic class of the verbal predicate (mental or social verbs), semantic class of the subject (animate, inanimate, no subject), tense of the finite predicate and voice of the finite predicate could be promising variables to explore in order to identify whether they are associated with subjective or objective meanings in spanish. as they have been understudied from this type of approach, further investigation into spanish causal connectives and coherence relations is still worthwhile. a reasonable approach might be to use experimental methods, which could increase the reliability and validity of data collected so far, with manual annotations. crowdsourcing experimentation seems to be a promising methodology. it allows us to obtain information based on the knowledge of native speakers avoiding possible biased annotations by analysts (pusse, et al., 2016; scholman & demberg, 2017). our future work will be oriented in this direction. regarding our second research question, implicit relations in journalistic context tend to be non-objective, i.e. they avoid expressing objectivity, and in the case of containing a soc, implicit relations are used to express subjective meanings. by observing the cases of journalistic context in detail, we notice that many of these relations are like examples (14) and (15): (14) (c1) el gobierno peruano ha negado esta acusación. (s1) sin duda, la anunciada y no concretada rendición de humala acabó con la paciencia del gobierno. (s2) en un comunicado, a media tarde de ayer, el ministerio de interior deploró la actitud de humala de no deponer las armas, tal como lo había anunciado el pasado domingo (c2) e invitó a la población que vive en las inmediaciones de la comisaría a abandonar temporalmente sus viviendas. [ne180] ‘(c1) the peruvian government has denied this accusation. (s1) undoubtedly, the announced and unspecified surrender of humala put an end to the patience of the government. (s2) in an announcement, yesterday afternoon, the ministry of interior deplored the attitude of humala, who did not put down weapons as he had announced the last sunday (c2) and it invited the population living in the vicinity of the police station to temporarily abandon their homes.’ subjectivity in spanish discourse 185 (15) (c1) el movimiento de la casa blanca y las sanciones aplicadas a siete altos funcionarios venezolanos marcan un cambio significativo en la actitud estadounidense ante los graves acontecimientos que se suceden en venezuela y que se han acelerado desde que nicolás maduro asumió la presidencia hace casi dos años. (s1) la situación política y social en el país sudamericano es casi insostenible. (s2) la ineficacia en la gestión económica, junto a la violación por parte del gobierno de derechos básicos, han llevado al país a un estado de penuria material y de tensión social que tienen difícil justificación. (c2) algunos de los dirigentes más importantes de la oposición se encuentran encarcelados. ‘(c1) the movement of the white house and the sanctions applied to seven venezuelan senior officials mark a significant change in the us attitude towards the serious events that are happening in venezuela and that have increased since nicolás maduro assumed the presidency almost two years ago. (s1) the political and social situation in the south american country is almost unsustainable. (s2) the inefficiency in economic management and the violation of basic rights by the government have led the country to a state of material scarcity and social tension, which are difficult to justify. (c2) some of the most important leaders of the opposition are imprisoned.’ in both examples we observed that there is a soc who is evaluating a situation. the epistemic stance marker sin duda (‘undoubtely’) in example (14) and the use of the adverb casi (‘almost’) and the adjective insostenible (‘unsustainable’) in (15) reflect this soc’s presence. it could be the case that by not having any mark to express causality, implicit relations need to make even clearer the relation established between segments and consequently, these are more preferred to express subjective relations like (14) and (15), which were the predominant ones in our corpus (subjective indicators speech-act/epistemic, speech-act/judgement, implicit soc and author/current speaker reached more than 70% in both types of contexts, academic and journalistic). in academic context similar relations to (14) and (15) were identified, but the number of objective relations was less extreme than in journalistic context, which would explain the non-significance in this context. it is important to stress that in the current study, implicit relations were analyzed from a subjectivity perspective. implicit relations have been understudied in discourse coherence studies (but see taboada & das, 2013; das & taboada, 2017 for recent exceptions), which have always been biased towards explicitly marked relations. our findings provide a preliminary perspective of how subjectivity is manifest in absence of connectives, which is a step forward into the knowledge of subjectivity in discourse. we identified that subjectivity is manifested differently in implicit relations, depending on contexts. for this reason, we strongly encourage exploring implicit relations in different languages and contexts in order to compare results and obtain more insights into subjectivity. moreover, the exploration of these relations would be an interesting issue for future studies on reading comprehension and processing information, especially because there have been studies that claim that subjective relations are more difficult to process and that connectives serve as instructions that facilitate the processing of this type of relations (canestrelli et al., 2013). as for our third research question, the current study allowed us to confirm that context plays an influential role in the meaning and use of coherence relations. for instance, significant relationships were identified between subjectivity, linguistic marking and context. this implies that context is a variable that should be considered in the analysis of coherence relations and connectives. especially, in languages that have been understudied or that do not have connectives with a clear-cut profile in terms of subjectivity, context might give more insights into the meaning and use of these relations and connectives. clear as this may be as a general tendency, we have to santana, spooren, nieuwenhuijsen and sanders 186 acknowledge that it is not easy to interpret specific interactions between context, subjectivity and linguistic marking in a straightforward way, because we are talking about three-way interactions. for instance, the results suggest that the relevance of context in the relation between subjectivity and marking is clearer in journalistic context than in academic context, since this association was significant for all variables in journalistic contexts, whereas it was significant only for two variables in academic contexts. however, it is hard to relate this directly to the characteristics of journalistic versus academic contexts, in general. it is tempting to argue that in journalistic texts it is necessary to communicate precisely what kind of events happened, how they occurred, who participated in those events, in order to clarify if they correspond to events of the real world or if they represent the thoughts of someone in particular. however, this intuition is not in line with the results that we found, because we found many epistemic relations in which the author is clearly present. for academic texts we found a similar relationship, but only for explicit puesto que-relations. this may be because authors of academic texts are more oriented to a particular audience of academic readers. they want to make a point, to argue, or to demonstrate the relevance of their work (hyland, 2001, 2011). these characteristics are in line with the subjectivity variables that reached statistical significance in our corpus, especially for epistemic relations in the case of the connective puesto que (‘given that’) (presence of the soc and identity of the soc). these findings illustrate that we are still in need of further specification of the role of the context. among the limitations of this study, we mention the size of the corpus. as this study focused on specific contexts, it was necessary to collect texts corresponding to these contexts searching several resources in spanish, which are limited, and in some cases, a part of them are not available freely. therefore, future analyses, especially those that involve specialized contexts, should take this methodological issue into account. finally, we believe that the methodology used in this research provided us with a better understanding of subjectivity of spanish causal relations and connectives. in comparison with automatic analyses carried out in our previous corpus-based study (santana et al., 2017), manual analyses allowed us to explore in more depth the local contexts of spanish coherence relations revealing interesting insights about these relations and their connectives. the strategy of analysis used in this study played a crucial role in the inter-rater agreement. the discussion of disagreements in the warming-up phase and in the preliminary manual analysis were beneficial steps in our methodology. these two rounds of annotations allowed us to clarify doubts, define the categories and indicators even more precisely, which was beneficial to the reliability because the coding can be reproduced in other research contexts. moreover, it was favorable for the validity because the coding reached a higher quality, which now provides us with a clearer picture of causal coherence in the spanish language. acknowledgements the present study is part of a ph.d. research project carried out at the utrecht institute of linguistics ots, utrecht university, which is funded by becas-chile from the national commission for scientific and technological research (conicyt-chile). the first author is grateful to these two institutions for financial, administrative and scientific support. we also thank fernando moncada and inés recio, for their crucial collaboration in the coding phases of this work. subjectivity in spanish discourse 187 references nicholas asher and alex lascarides (2003). logics of conversation. cambridge university press, cambridge. claudia borzi and ulrich detges (2011). ya que, un marcador polifónico. in h. aschenberg and óscar loureda lamas (eds.), marcadores del discurso: de la descripción a la definición: 263–281. iberoamericana vervuert, madrid. danièle bourcier and oswald ducrot (1980). les mots du discours. éditions de minuit, paris. vijay bhatia (2002). a generic view of academic discourse. in j. flowerdew (ed.), academic discourse: 21–39. routledge, new york. vijay bhatia (2004). worlds of written discourse: a genre-based view. continuum, london. sara blackwell (2016). porque in spanish oral narratives: semantic porque, (meta) pragmatic porque or both? in a. capone and j.l. mey (eds.), interdisciplinary studies in pragmatics, culture and society: 615–651. springer, heidelberg. antonio briz (1998). el español coloquial en la conversación: esbozo de pragmagramática. ariel, barcelona. antonio briz (2000). ¿cómo se comenta un texto coloquial? ariel, barcelona. anneloes canestrelli, willem mak and ted sanders (2013). causal connectives in discourse processing: how differences in subjectivity are reflected in eye movements. language and cognitive processes, 28(9): 1394–1413. lynn carlson, daniel marcu and mary ellen okurowski (2003). building a discourse-tagged corpus in the framework of rhetorical structure theory. in j. van kuppevelt and r. smith (eds.), current directions in discourse and dialogue: 85-122. kluwer, dordrecht. manuel casado velarde (1991). los operadores discursivos es decir, esto es, o sea ya saber en español actual: valores de lengua y funciones textuales. lingüı́stica española actual, 13(1): 87–116. iria da cuhna, juan manuel torres-moreno and gerardo sierra (2011). on the development of the rst spanish treebank. in proceedings of the v linguistic annotation workshop (law) (pp. 1-10). stroudsburg, pa, usa: association for computational linguistics. debopam das and maite taboada (2017). rst signalling corpus: a corpus of signals of coherence relations. language resources & evaluation, 52(1):149-184. liesbeth degand (2001). form and function of causation: a theoretical and empirical investigation of causal constructions in dutch. peeters publishers, leuven. liesbeth degand and henk pander maat (2003). a contrastive study of dutch and french causal connectives on the speaker involvement scale. in a. verhagen and j. van de weijer (eds.), usage-based approaches to dutch. lexicon, grammar, discourse:175–199. lot, utrecht. noemí domínguez garcía (2007). conectores discursivos en textos argumentativos breves. arco libros, madrid. andy field, jeremy miles and zoe field (2012). discovering statistics using r. sage publications ltd., london. celia ford (1993). grammar in interaction: adverbial clauses in american english conversations. cambridge university press, cambridge. catalina fuentes (2009). diccionario de conectores y operadores del español. arco libros, madrid. carmen galán rodríguez (1999). la subordinación causal y final. in i. bosque and v. demonte (eds.), gramática descriptiva de la lengua española: 3597-3642. espasa calpe, madrid. maría del pilar garcés gómez (2014). diacronía de los marcadores discursivos y representación en un diccionario histórico (anexos de revista de lexicografía, 28). universidade da coruña, a coruña. santana, spooren, nieuwenhuijsen and sanders 188 maría del camino garrido rodríguez (2004). conectores contraargumentativos en la conversación coloquial. secretariado de publicaciones y medios audiovisuales de la universidad de león, león. patrick goethals (2002). las conjunciones causales explicativas españolas como, ya que, pues y porque: un estudio semiótico-lingüístico. peeters publishers, leuven. patrick goethals (2010). a multi-layered approach to speech events: the case of spanish justificational conjunctions. journal of pragmatics, 42(8): 2204–2218. susanne günthner (1993). «... weil―man kann es ja wissenschaftlich untersuchen». diskurspragmatische aspekte der wortstellung in weil-sätzen. linguistische berichte, 143: 37–59. ken hyland (2001). humble servants of the discipline? self-mention in research articles. english for specific purposes, 20(3): 207–226. ken hyland (2009). academic discourse: english in a global context. continuum, london. ken hyland (2011). disciplines and discourses: social interactions in the construction of knowledge. in d. starke-meyerring, a. paré, n. artemeva, m. horne and l. yousoubova (eds.), writing in the knowledge society: 193-214. parlor press and the wac clearinghouse, west lafayette, indiana. rudi keller (1995). the epistemic weil. in d. stein and s. wright (eds.), subjectivity and subjectivisation: linguistic perspectives: 16–30. cambridge university press, cambridge. alistair knott and ted sanders (1998). the classification of coherence relations and their linguistic markers: an exploration of two languages. journal of pragmatics, 30(2): 135175. richard landis and gary kosh (1977). the measurement of observer agreement for categorical data. biometrics, 33(1): 159-174. ronald langacker (1990). subjectification. cognitive linguistics, 1(1): 5-38. natalia levshina and liesbeth degand (2017). just because: in search of objective criteria of subjectivity expressed by causal connectives. dialogue & discourse, 8(1): 132-150. fang li, jacqueline evers-vermeul and ted sanders (2013). subjectivity and result marking in mandarin. chinese language & discourse, 4(1): 74–119. ziheng lin, min-yen kan and hwee tou ng (2009). recognizing implicit discourse relations in the penn discourse treebank. in proceedings of the 2009 conference on empirical methods in natural language processing (emnlp 2009) (pp. 343–351). stroudsburg, pa, usa: association for computational linguistics. john lyons (1977). semantics. cambridge university press, cambridge. manuel martí sánchez (2008). los marcadores en español l/e: conectores discursivos y operadores pragmáticos. arco libros, madrid. maría antonia martín zorraquino and estrella montolío (1998). marcadores del discurso. ariel, barcelona. roser martínez (1997). conectando texto: guía para el uso efectivo de elementos conectores en castellano. octaedro, barcelona. estrella montolío (2001). conectores de la lengua escrita. ariel, barcelona. henk pander maat and ted sanders (2000). domains of use or subjectivity? the distribution of three dutch causal connectives explained. in e. couper-kuhlen and b. kortmann (eds.), cause, condition, concession, contrast: cognitive and discourse perspectives (topics in english linguistics, 33): 57-82. mouton de gruyter, berlin, boston. henk pander maat and ted sanders (2001). subjectivity in causal connectives: an empirical study of language in use. cognitive linguistics, 12(3): 247–274. henk pander maat and liesbeth degand (2001). scaling causal relations and connectives in terms of speaker involvement. cognitive linguistics, 12(3): 211–246. subjectivity in spanish discourse 189 henk pander maat and ted sanders (2006). connectives in text. in k. brown, a.h. anderson, l. bauer, m. berns, g. hirst and j. miller (eds.), encyclopedia of language and linguistics: 33–41. elsevier, amsterdam. mirna pit (2006). determining subjectivity in text: the case of backward causal connectives in dutch. discourse processes, 41(2): 151–174. mirna pit (2007). cross-linguistic analyses of backward causal connectives in dutch, german and french. languages in contrast, 7(1): 53–82. mirna pit, jacqueline hulst and henk pander maat (1996). subjectiviteit en de spaanse connectieven porque, ya que en puesto que. gramma/ttt, 5(3): 221–240. salvador pons (1998). conexión y conectores: estudio de su relación en el registro informal de la lengua (cuadernos de filología, anejo no. 27). universitat de valència, facultat de filología, departamento de filología española (lengua española), valencia. salvador pons (2000). los conectores. in a. briz (ed.), ¿cómo se comenta un texto coloquial?: 193–220. ariel, barcelona. josé portolés lázaro (2001). marcadores del discurso. ariel, barcelona. florian pusse, asad sayeed and vera demberg (2016). lingoturk: managing crowdsourced tasks for psycholinguistics. in proceedings of the north american chapter of the association for computational linguistics: human language technologies (naacl-hlt). http://aclweb.org/anthology/n/n16/n16-3012.pdf. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind k. joshi and bonnie l. webber (2008). the penn discourse treebank 2.0. in proceedings of the 6th international conference of language resources and evaluation (lrec). https://www.seas.upenn.edu/~pdtb/papers/pdtb-lrec08.pdf ted sanders (1997). semantic and pragmatic sources of coherence: on the categorization of coherence relations in context. discourse processes, 24(1): 119–147. ted sanders and wilbert spooren (2007). discourse and text structure. in d. geeraerts and h. cuykens (eds.), handbook of cognitive linguistics: 916–943. oxford university press, oxford. ted sanders and wilbert spooren (2009). causal categories in discourse – converging evidence from language in use. in t. sanders and e. sweetser (eds.), linguistic categories of causality in discourse: 205–246. mouton de gruyter, berlin. ted sanders and wilbert spooren (2015). causality and subjectivity in discourse: the meaning and use of causal connectives in spontaneous conversation, chat interactions and written text. linguistics, 53(1): 53–92. ted sanders, wilbert spooren and leo noordman (1992). toward a taxonomy of coherence relations. discourse processes, 15(1): 1–35. ted sanders, wilbert spooren and leo noordman (1993). coherence relations in a cognitive theory of discourse representation. cognitive linguistics, 4(2): 93–133. josé sanders, ted sanders and eve sweetser (2012). responsible subjects and discourse causality. how mental spaces and perspective help identifying subjectivity in dutch backward causal connectives. journal of pragmatics, 44(2): 191–213. andrea santana, dorien nieuwenhuijsen, wilbert spooren and ted sanders (2017). causality and subjectivity in spanish connectives: exploring the use of automatic subjectivity analyses in various text types. discours. revue de linguistique, psycholinguistique et informatique, 20. luis santos río (2003). diccionario de partículas. luso-española de ediciones, salamanca. marc silver (2006). language across disciplines: towards a critical reading of contemporary academic discourse. brownwalker press, florida. merel scholman and vera demberg (2017). crowdsourcing discourse interpretations: on the influence of context and the reliability of a connective insertion task. in proceedings of the santana, spooren, nieuwenhuijsen and sanders 190 xi linguistic annotation workshop (law) (pp. 24-34). stroudsburg, pa, usa: association for computational linguistics. wilbert spooren (1997). the processing of underspecified coherence relations. discourse processes, 24(1):149-168. wilbert spooren and liesbeth degand (2010). coding coherence relations: reliability and validity. corpus linguistics and linguistic theory, 6(2): 241–266. wilbert spooren, ted sanders, mike huiskes and liesbeth degand (2010). subjectivity and causality: a corpus study of spoken language. in j. newman and s. rice (eds.), empirical and experimental methods in cognitive/ functional research: 241–255. university of chicago press, chicago. ninke stukker and ted sanders (2009). another(’s) perspective on subjectivity in causal connectives: a usage-based analysis of volitional causal relations. discours. revue de linguistique, psycholinguistique et informatique, 4. ninke stukker and ted sanders (2012). subjectivity and prototype structure in causal connectives: a cross-linguistic perspective. journal of pragmatics, 44(2): 169–190. ninke stukker, ted sanders and arie verhagen (2008). causality in verbs and in discourse connectives: converging evidence of cross-level parallels in dutch linguistic categorization. journal of pragmatics, 40(7): 1296–1322. john swales (1990). genre analysis: english in academic and research settings. cambridge university press, new york. eve sweetser (1990). from etymology to pragmatics: the mind-body metaphor in semantic structure and semantic change. cambridge university press, cambridge. maite taboada (2006). discourse markers as signals (or not) of rhetorical relations. journal of pragmatics, 38(4): 567–592. maite taboada (2009). implicit and explicit coherence relations. in j. renkema (ed.), discourse, of course: 127–140. john benjamins, amsterdam. maite taboada and debopam das (2013). annotation upon annotation: adding signalling information to a corpus of discourse relations. dialogue and discourse, 4(2): 249–281. teun van dijk (1998). news as discourse. lawrance erlbaum associates, inc., publishers, new jersey. elizabeth traugott (1995). subjectification in grammaticalisation. in s.wright & d. stein (eds.), subjectivity and subjectivisation: linguistic perspectives: 31–54. cambridge university press, cambridge. nancy vázquez veiga (2002). diccionario de colocaciones y marcadores del español: esbozo de una entrada de un marcador discursivo. communication in iv congreso de lingüística general (2459–2472). servicio de publicaciones de la universidad de cádiz, cádiz-alcalá de henares. agustín vera (1984). en torno a la causalidad (aproximación a los fenómenos recursivos-causales a la luz de una teoría de base prototípica). anales de la universidad de murcia, 42(1–2): 31–50. arie verhagen (2005). constructions of intersubjectivity: discourse, syntax, and cognition. oxford university press, oxford. linda waugh (1995). reported speech in journalistic discourse: the relation of function and text. text interdisciplinary journal for the study of discourse, 15(1): 129–173. yipu wei, ted sanders, jacqueline evers-vermeul and willem mak (2018). causal connectives and perspective markers in chinese: the encoding and processing of subjectivity in discourse. lot, utrecht. heide wegener (2000). da, denn und weil – der kampf der konjunktionen. zur grammatikalisierung im kausalen bereich. in r. thieroff, m. tamrat, n. fuhrhop and o. teuber (eds.), deutsche grammatik in theorie und praxis: 69–82. mouton de gruyter, berlin, boston. subjectivity in spanish discourse 191 sandrine zufferey (2012). “car, parce que, puisque” revisited: three empirical studies on french causal connectives. journal of pragmatics, 44(2): 138–153. sandrine zufferey, willem mak, ted sanders and sara verbrugge (2017). usage and processing of the french causal connectives car and parce que. journal of french language studies, 28(1): 85-112. subjectivity in spanish discourse: explicit and implicit causal relations in different contexts 1. introduction references dialogue & discourse 12(1) (2021) 21–44 doi: 10.5210/dad.2021.102 cognitive and social delays in the initiation of conversational repair julia mertens julia.mertens@tufts.edu tufts university, medford, ma, 02155 jan p. de ruiter jp.deruiter@tufts.edu tufts university, medford, ma, 02155 editor: patrick g.t. healey submitted 04/2020; accepted 01/2021; published online 03/2021 abstract the exact timing of a conversational turn conveys important information to listeners. most turns are initiated within 250ms after the previous turn. however, interlocutors take longer to initiate certain types of turns: those that either require more cognitive processing or are socially dispreferred. many dispreferred turns are also cognitively demanding, so it is difficult to attribute specific conversational delays to social or cognitive mechanisms. in this paper, we evaluate how cognitive and social variables contribute to the timing of utterances in conversation. we focus on a type of turn that is socially dispreferred, cognitively demanding, and generally delayed: other-initiations of repair (oirs). oirs occur when a listener notices and decides to signal a comprehension problem (e.g., “what?”). we analyzed the floor transfer offsets of 456 oirs. interlocutors initiated oirs later when trouble source turns had weaker discourse context or were shorter. we found cognitive effects of trouble source duration and discourse context: a longer duration or stronger context was associated with shorter oir ftos. we also found social effects of problem attribution and size: when the problem could be attributed to the environment or was smaller, oir ftos were shorter. discourse context, planning, and social attribution manifest in the timing of turns. keywords: other-initiated repair, prediction, floor transfer offset, preference organization, turn taking 1. introduction humans develop and maintain relationships through conversation. in fact, conversation is ubiquitous: children acquire language by interacting with their parents and peers, and coworkers increase their social coherence by chatting around the water cooler. however, conversation is not as effortless as it seems. on the contrary, conversation is an advanced cognitive skill. interlocutors must perform many tasks at once. speakers coordinate their words, prosody, and other nonand para-verbal signals to form messages tailored to their recipient and the context. listeners integrate that information into their discourse record as they plan their response. at the same time, all interlocutors update their common ground, which involves accessing both their short-term and long-term memory. to perform all these tasks at once, conversants draw heavily on cognitive resources. these taxing cognitive tasks are also time sensitive. interlocutors aim to minimize gaps and overlaps (sacks et al., 1978). as a consequence, speakers usually begin their turn within 250ms of ©2021 julia mertens and jan p. de ruiter this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). mertens and de ruiter the end of the previous turn (de ruiter et al., 2006; stivers et al., 2009). in comparison, participants in psycholinguistic experiments take more than twice as long (according to indefrey and levelt, 2004, approximately 600ms) to initiate the name of a simple object. it seems much more complex to plan a conversational turn than to plan the name of an object. nevertheless, interlocutors initiate conversational turns much earlier than the processing time for simple object naming would suggest. turn transition times in conversation are so short because listeners predict the upcoming turn (de ruiter et al., 2006; magyari and de ruiter, 2012; magyari et al., 2017) and plan their response before the current turn ends (jansen et al., 2014; magyari et al., 2017; bögels et al., 2020). even so, speakers initiate some turns later than others. in fact, speakers consistently delay some types of turns (schegloff, 1979). for example, people usually pause before declining (but not accepting) invitations. psycholinguists and conversation analysts explain conversational delays differently. psycholinguists argue that speakers initiate turns later when they need to process or manipulate more information. processing takes time; more processing takes more time (donders, 1969). for example, interlocutors initiate lies later than truths (walczyk et al., 2003) because lies require additional cognitive processes: deciding to lie, inhibiting the truth, and constructing the lie (walczyk et al., 2003). from the psycholinguistic perspective, delays are not part of the speaker’s intended message. instead, delays are unintentional, automatic side effects of increased cognitive processing; any time the listener’s processing load increases they take longer to initiate their response. conversation analysts argue that some delays reflect preference organization and speakers delay turns to signal that the turn will be dispreferred (clayman, 2002). dispreferred responses do not follow the current action of the conversation or do not create social solidarity (schegloff, 1968). in contrast, speakers do not delay preferred turns, which continue the current action of the conversation or create social solidarity. in excerpt 1, les attempts to compensate joy for a favor. on line 4, joy rejects les’s offer. rejections are typically dispreferred turns, but it is preferred to reject offers. the rejection overlaps with the offer. the timing of the rejection emphasizes that the offer is unnecessary. early responses convey that the response should be expected or that the speaker is eager to produce the utterance. excerpt 1: fast rejection on line 4, from kendrick and torreira (2015) 1 les:.hhh p’aps you’d likewould you like eh:m:: some 2 frozen f:::: # ruit fr’m our kou:r # freezer as 3 a small recompens[e? 4 joy: [oh: les for goodness sake n:o 5 i don’t want anything if a speaker does not delay a dispreferred turn, they risk appearing rude. excerpt 2 is a fictional, modified version of excerpt 1. joy’s response (line 2-3) now rejects an invitation. the rejection implies an eagerness to reject les, or that les should have expected a rejection. there may be some unknown negative history between les and joy, or joy is rude. speakers may delay initiating a rejection to avoid conveying an inappropriate sentiment. excerpt 2: fictional rejection on line 3 1 les: .hhh p’aps you’d like to go out to dinn[er? 2 joy: [oh: les 3 for goodness sake n:o i don’t 22 cognitive and social delays in the initiation of conversational repair delays also prepare the recipient for a dispreferred turn. approximately 53% of all turns are preferred. out of turns delayed by 600ms, only approximately 15% are preferred (kendrick and torreira, 2015). as someone increasingly delays an upcoming turn, their interlocutor can more confidently predict that the turn will be dispreferred. a dispreferred response can be addressed before it is even produced (pomerantz, 1984). just before excerpt 3, margy asked emma for a favor. in lines 1-3, emma tells margy she’s currently unavailable, and offers to help later in the day. after a substantial gap, margy begins her turn with “well,” which typically projects an upcoming dispreferred turn (heritage, 2015). in addition, the “well” is stretched, as indicated by the colon (hepburn and bolden, 2013). stretched turn-initial “well” is especially associated with dispreferred turns (rühlemann, 2019). emma predicts that margy will produce a dispreferred response. in line 6, emma attempts to preempt the expected dispreferred response by providing it herself, with “do you have to have it done now” (“now” as opposed to “after” on line 1). excerpt 3: projected dispreferred response, adapted from kendrick and torreira (2015) 1 emm: honey i’ll come down after i had muh liddle bowl 2 a’soup’n salad’n i’ll call’em ba:ck to yuh i’d 3 love it. 4 (1025 ms) 5 mar: we :ll (0.7) oka:y [i:-uh: (.) i wanteda (j’s)]] 6 emm: [d’you havtuh have it done 7 no:w? the crucial point is that interlocutors initiate utterances earlier or later to convey the appropriate meaning. even if they are cognitively able to produce a faster response, the interlocutor may delay their turn. from the social perspective, delays are not a mere by-product of cognitive processing. instead, delays are themselves intentional signals (see e.g., clark and fox tree, 2002). to complicate matters, these two explanations tend to overlap in their predictions; many dispreferred turns are also cognitively demanding, so both explanations predict similar delays. one challenging confound is that complex utterances take more processing and time to translate from conceptual to motor movement plans (ferreira, 1991; kempen and huijbers, 1983; sternberg et al., 1978). many dispreferred utterances are more linguistically complex; speakers mitigate, explain, or otherwise minimize the potential negative implicatures of their turn. excerpt 4 illustrates this phenomenon. on line 4, speaker a expresses appreciation for their interlocutor before rejecting an invitation. this preface makes the utterance more linguistically complex. overall, it may be cognitively demanding to design a predictable but dispreferred turn. dispreferred turns perform unexpected actions. in addition to all the processing needed to plan a preferred turn, speakers may have to strategize the most appropriate way to project an upcoming dispreferred turn. excerpt 4: appreciation before dispreferred response, from schegloff (1968) 1 spb: and uh theuh if you’d care to come over and 2 visit a little while this morning, i’ll give you a 3 cup of coffee. 4 spa: hehh! well that’s awfully sweet of you, i don’t 5 think i can make it this morning, hh uhm i’m 6 running an ad in the paper and-and uh i have to 7 stay near the phone. 23 mertens and de ruiter to summarize, delays in initiating conversational turns could be due to cognitive variables, social variables, or a combination of the two. in this study, we examine social and cognitive variables together to determine their relative contributions to delays in conversation. roberts et al. (2015) analyzed how both cognitive processing and preference organization influenced the timing of all utterances in a corpus. however, different types of utterances have different timing mechanisms. for example, backchannels (e.g., “uh-huh” or “yeah”; schegloff, 1982; yngve, 1970), insults, and questions occur in different contexts, perform different functions, and involve different social and cognitive information. to avoid these confounds, we focus on one type of utterance that (1) requires greater levels of cognitive processing, (2) is socially dispreferred and (3) tends to be delayed in comparison to other turns. together, these features suggest that the utterance is delayed because it requires increased cognitive processing and/or is socially dispreferred. excerpt 5: open class oir (albert et al., 2015) 1 spa: we can’t have three holidays in a year, can we? 2 spb: why not? 3 (1.5 seconds) 4 spa: sorry? 5 spb: why not? these criteria lead us to focus on other-initiations of repair (oir; schegloff et al., 1977). an oir occurs when a listener diagnoses their own problem of hearing or understanding and chooses to signal that problem to their interlocutor (e.g., “sorry?” in excerpt 5). oirs meet all criteria for this study. first, turns marked by oirs induce extra cognitive processing. the interlocutor identifies, diagnoses, and attempts to cognitively ’repair’ the problem on their own (fodor and inoue, 2000) before planning their response (an oir). postdiction is an example of one “extra” cognitive process that listeners may employ when comprehending trouble source turns (gwilliams et al., 2018). when a listener processes an ambiguous speech sound, they maintain that speech sound in working memory. as the turn unfolds, the listener not only comprehends new information but also relates that information to the ambiguous speech sound, in hopes of disambiguating the ambiguity. maintaining ambiguous information in working memory requires cognitive resources (baddeley, 2003). comparing new and old information requires extra cognitive processing. if the information is not disambiguated, the listener may produce an oir. see section 3.2 on cognitive variables for the cognitive variables examined in this study. second, oirs are socially dispreferred. oirs slow conversation down (threaten progressivity; schegloff, 1979). in excerpt 5, line 5 (“why not?”) could have been answered right away. instead, speaker a initiated repair, delaying progressivity for at least two turns. in addition, oirs suggest that the oir producer could not understand, or was not paying attention to, their interlocutor (threaten face; schegloff et al., 1977). see sections 3.3 and 3.4 for the social variables examined in this study. finally, oirs are delayed in comparison to most other turns (kendrick, 2015b; schegloff et al., 1977). while responses to yes/no questions follow a median delay of approximately 340ms, oirs follow a median delay of approximately 721ms (kendrick, 2015b). these characteristics — increased cognitive processing, dispreferred status, and delay — make oirs a good candidate for studying the relative influence of cognitive and/or social factors on utterance timing in conversation. few, if any, other types of utterances have been shown in the literature to meet these criteria. 24 cognitive and social delays in the initiation of conversational repair 2. method 2.1 data source we collected oirs from the conversation analytic british national corpus (cabnc; albert et al., 2015). the cabnc is a collection of the spoken language in the british national corpus (bnc; leech, 1992). the cabnc contains approximately 4.2 million words from 1,436 conversations recorded in the uk in the late 20th century. the participants were volunteers selected from different age groups, regions, and social classes in a demographically balanced way. we analyzed cabnc recordings of in-person (i.e., not over the telephone) and non-institutional (e.g., no call-ins to radio shows) conversations because these represent the most informal types of conversation. exploratory, hypothesis-generating studies on conversation should analyze as unconstrained data as possible. 2.2 oir collection even within the category of oirs, there were potential confounds. first, it takes more time to initiate longer utterances (ferreira, 1991; kempen and huijbers, 1983; sternberg et al., 1978). therefore, we restricted our analysis to single-word oirs. in addition, some oirs perform multiple actions. for example, oirs that repeat the trouble source turn can also challenge the relevance of that turn (kendrick, 2015a; robinson and kevoe-feldman, 2010). to complicate matters, some turns may either be an oir or a different action, such as a joke (e.g., excerpt 6). we included a potential oir only if the interlocutors treated the potential trouble source as either misheard or misunderstood (sacks et al., 1978). the interlocutors in excerpt 6 treat “the concrete jungle” as a joke. therefore, we would not include excerpt 6 in our oir collection. excerpt 6: an oir as a joke, from kendrick (2015a) 1 cha: (it’s a) nice place to work though. 2 (0.9) 3 liz: ehhh what.=the concrete jungle, 4 (0.2) 5 cha: aww:::::.=i think it’s quite pretty. it has 6 ree:ds. these constraints restricted our analysis to a small subset of all oirs. only approximately 25% of the oirs in kendrick (2015a) would have been included in our collection. however, we required a high level of precision to isolate the effects of specific cognitive and social variables in naturally occurring data. speakers can produce oirs in the turn directly after the trouble source turn, or some number of turns removed from the trouble source turn. in this paper, we only analyzed oirs that occur in the turn after the trouble source turn. this constraint made oir fto a more valid measure of the listeners’ processing of the trouble source turn. in other sequences, where the oir is produced later in the conversation, operationalization of some measures becomes more difficult. in excerpt 7, speaker 2 produces a joke on line 2 (evidenced by “it was a joke” on line 8). this joke makes laughter relevant in the next turn. instead of laughing, speaker 1 responds with “yeah” (line 3), a continuuer or backchannel (yngve, 1970; schegloff, 1982) that implies that speaker 1 is listening to speaker 2. the mismatch prompts speaker 2 to upgrade their original joke with “every weekend yeah” (line 5). only after this upgrade does speaker 1 produce an oir. in this excerpt, should the 25 mertens and de ruiter figure 1: measures in this study. cognitive measures are bold and social measures are italicized. trouble source include line 2, even though speaker 1 responded to line 2 as if they did not notice any miscommunication? on one hand, speaker 1 may not have noticed a communication problem until line 5. on the other hand, “‘every weekend” is not the entire joke targeted by the repair solution on line 8. to avoid these problematic methodological questions, we decided to only collect oirs that occur in the turn directly after the trouble source. these selection criteria resulted in a final collection of 456 oirs. excerpt 7: delayed oir, collected by the human interaction laboratory at tufts university 1 sp1: i’ve never done some shit like this. 2 sp2: yeah u(0.3) oh i have li[ke,] 3 sp1: [ye]ah, 4 (0.3) 5 sp2: every weekend yeah, 6 sp1: w:hat, 7 (4.1) 8 sp2: i was just kidding. it was a joke, 3. measures in this section, we describe floor transfer offset (fto; de ruiter et al., 2006), a quantification of the time between turns. then, we describe the cognitive variables analyzed in this study. next, we describe the social variables analyzed in the study. finally, we describe our operationalization of overlapping talk during the trouble source. figure 1 displays the measures analyzed in this study. 3.1 floor transfer offset floor transfer offset (fto; de ruiter et al., 2006) is the time between two turns, calculated by subtracting the end time of the first turn from the start time of the second turn. negative ftos indicate overlaps, positive ftos indicate gaps, and ftos of 0 indicate gapless transitions. 26 cognitive and social delays in the initiation of conversational repair we defined turn boundaries based on turn construction units (tcus) and transition relevance places (trps; clayman (2013); sacks et al. (1978). tcus are composed of speech that is semantically, pragmatically, and prosodically complete. trps are moments where the preceding speech may be a tcu. at a trp, the speaker may choose to end their tcu (and turn). in figure 2, trps are marked with stars and the ends of tcus are marked with circles. line 2 ends in a trp, but speaker 1 adds two increments in lines 5 and 7. figure 2: transcript with trps (stars) and tcus (circles). we calculated ftos based on trps. we identified the trp most relevant to the second turn. we subtracted the trp time from the start time of the second turn. when ftos are negative (indicating an overlap), the trp is after the start point of the second turn. when ftos are positive (indicating a gap), the trp was before the start time of the second turn. the transcript displayed in figure 2 would result in the ftos displayed in figure 3. figure 3: ftos from lines 1-5 in figure 2. negative ftos are represented in red, and the positive fto is represented in blue. we defined the trouble source turn as a single tcu that contained the problem. however, this presented some difficulties. for example, the trouble may arise from disordered self-repair across multiple tcus. in excerpt 8, lines 1-2 contain 3 short tcus: “sam’s dead died” “dot died” and “his wife died in childbirth.” this sequence of self-repairs may have caused the communication problem. excerpt 8 also demonstrates another problem: anaphora. when ivy says “his” wife, she is referring to “sam’s” wife. this reference would be unclear without the context of the immediately prior turn. excerpt 8: oir sequence with a duration ratio of 0.06 from the cabnc 1 ivy: sam’s dead die:d. dot die:d his wife (1.3) died 2 in childbirth. 27 mertens and de ruiter 3 (0.9) 4 con: who. 5 (0.3) 6 ivy: dot, 7 (0.7) 8 con: oh sam’s wife, a third problem is when the trouble source turn was followed by a second tcu by the same speaker. in excerpt 9, the oir eventually targets “over there.” however, gordon produces another tcu: “is that it?” in each of these cases, we defined the trouble source turn as encompassing more than one tcu. there were 21 such cases. we conducted the statistical analyses with and without these sequences and obtained similar results. excerpt 9: oir sequence the trouble source in the first tcu from the cabnc 1 gor: what’s that lthing over there. is that it? 2 (0.4) 3 deb: whe:re. (0.2) no that’s sean’s and kirsty’s. the fto between the trouble source and the oir (the oir fto) was our outcome variable. the fto before the trouble source (the trouble source fto) was our measure of discourse context: long trouble source ftos indicate less discourse context (see section 3.2 on cognitive variables). finally, the fto between the oir and the repair solution (the repair solution fto) was one measure of progressivity: the shorter the repair solution fto, the less the oir paused progressivity (see section 3.3 on social variables). we marked all time points in praat (boersma, 2006). 3.2 cognitive variables we measured the following cognitive variables: (1) duration (e.g., piquado et al., 2010), (2) word frequency (for a review, see goldinger, 1996), (3) syllable rate (osada, 2004; roberts et al., 2015) and (4) availability of discourse context for the trouble source. in this section, we operationalize and present hypotheses about each of these variables. we measured trouble source turn length in duration and number of syllables. these variables were related, but not identical. trouble source turns with more syllables tended to be longer, but a turn composed of many fast syllables could have the same duration as one composed of a few slow syllables. in addition, when the speaker paused mid-turn, utterance duration increased while the number of syllables remained the same. we predicted that longer trouble source turns would contain more information and would require more cognitive processing to comprehend (roberts et al., 2015) as a result, we hypothesized that longer trouble source turns would be followed by longer oir ftos. we measured the number of syllables in the trouble source turn and calculated the trouble source turn syllable rate. roberts et al. (2015) used average phone lengths to calculate deviance from expected syllable rates. in contrast, we calculated the raw syllable rate. trouble source turn produced quickly should have a higher information density on average. therefore, faster trouble source turns should be associated with longer oir ftos. fourth, we calculated the average word frequency of the trouble source from the subtlex-uk corpus (van heuven et al., 2014). frequent words are easier to process (goldinger, 1996). we predicted that the higher the average word frequency, the shorter the oir fto. 28 cognitive and social delays in the initiation of conversational repair the relationship between discourse context and cognitive processing is more complex. in this paper, we use the term “discourse context” to refer to the sequential environment that occurs immediately prior to the trouble source turn. discourse context constrains the upcoming turn. for example, the turn after a question is expected to be an answer (schegloff, 1968). listeners use this information to predict the upcoming speech act (levinson, 2017; de ruiter and cummins, 2012; gisladottir et al., 2015). these predictions are time dependent: the more a turn is delayed (has a larger fto), the more likely it is that the turn will be dispreferred (kendrick and torreira, 2015). however, dispreferred utterances are typically indirect and may take many forms, in contrast to preferred utterances which are typically direct (pomerantz, 1984). therefore, it may be more difficult to predict the exact words in a delayed trouble source turn (gisladottir et al., 2015). prediction may either increase or decrease the processing needed to notice a hearing or understanding problem. if a listener has a strong prediction for the trouble source turn, they may easily notice deviations from that prediction. in contrast, surprised listeners may work harder to make sense of the surprising information. this extra cognitive processing would delay the oir. discourse context and listener predictions about the trouble source should have the same effect on oir fto, because discourse context leads to prediction. therefore, we can make bidirectional predictions about the effect of discourse context on oir ftos. we operationalized discourse context as the fto before the trouble source turn; the shorter the trouble source turn fto, the stronger the context for the trouble source turn. we also created a binary variable indicating whether the trouble source turn fto was greater than 1.5 seconds. most turns transitions are shorter than 1 second, so trouble source turns with ftos longer than 1.5 seconds are considered to be disjointed, or with little discourse context. 3.3 social variables in this paper, we focused on two preferences especially relevant to the timing of oirs: progressivity and specificity. speakers aim to be progressive, or to continue the activity of the conversation (schegloff, 2007; stivers and robinson, 2006). oirs stop the current activity for at least two turns: the oir and the repair solution. oirs always decrease progressivity, but some require more time and work to resolve than others. we predict that speakers initiate oirs earlier when they require less time to resolve. we measured repair solution fto by subtracting the end time of the oir from the start time of the repair solution. in addition, we calculated the duration of the repair solution. however, the duration of the repair solution is dependent on the duration of the trouble source turn. even when an oir is more progressive than another would have been, if the trouble source turn is very long, the repair solution will also be very long. therefore, we calculated the duration ratio of the sequence: the ratio of the repair solution duration to the trouble source turn duration. a duration ratio of 1 signifies that the repair solution and trouble source turn were the same length. a duration ratio of 0.5 signifies that the repair solution was half the length of the trouble source turn. the sequence in excerpt 8 has a very small duration ratio, in part because ivy pauses for 1.3 seconds during the trouble source turn. when solution fto was longer or duration ratio was higher, progressivity was paused for longer. we predicted that in such sequences, oir fto would also be longer. dingemanse et al. (2015) examined measures that were similar but distinct from the ones explored in this study. first, they found that the duration of the trouble source turn was always approximately equal to the combined duration of the oir and repair solution. at first glance, this 29 mertens and de ruiter evidence suggests that our measure of duration ratio would be around 1 for all sequences. however, dingemanse et al. (2015) did not focus on single-word oirs. the variance in duration of singleword oirs is much lower than the variance in duration of all oirs. different single-word oirs still request different types of repair solutions, so we still expected variance in the repair solution duration. we also did not include the duration of the oir in our duration ratio measure. therefore, we still expected variance in duration ratios. in addition, there is a preference for more specific oirs (schegloff et al., 1977; dingemanse et al., 2015; kendrick, 2015b). specificity is related to the size of the problem marked by the oir; specific oirs mark “smaller” communication problems (e.g., excerpts 8 and 9). open class oirs (e.g., excerpts 5 and 7) mark the entire trouble source turn (drew, 1997). there are two reasons that specific oirs may be preferred over open class oirs. first, it may be more face-threatening to admit to a larger communication problem by producing an open class oir. second, since specific oirs mark smaller problems than open class oirs, the repair solutions to specific oirs tend to be shorter (clark and schaefer, 1987; dingemanse et al., 2015; grice, 1975). this means that specific oirs halt progressivity for a shorter duration. we created a binary variable indicating whether the oir was open class. out of the 456 oirs, 269 (59%) were open class oirs and 187 (41%) were specific oirs. for every specific oir, we had almost 1.5 open class oirs. in contrast, kendrick (2015b) found 2.2 specific oirs for every open class oir. we believe this difference is due to our collection criteria. many repeats, candidate understandings, restricted requests, and other types of specific oirs tend to have more than one word and, therefore, were not included in this collection. we predicted that specific oirs would be initiated earlier than open oirs, the repair solutions in specific oir sequences would be shorter, and that duration ratio will contribute to earlier specific oirs. parts of the interlocutors’ social identity could have effects on oir fto. for example, the relationship between gender and turn-timing has been a controversial topic for decades (e.g., west and zimmerman, 2015). roberts et al., (2015) found that for each male in a dyad, ftos increased by approximately 70ms. however, we decided to not analyze social identity in this analysis. first, the cabnc has inconsistent labels of these variables. second, even if we were to find that members of one group (e.g., women) initiate oirs earlier than members of another group (e.g., men), we would not be able to explain why. most likely, differences between social groups would be due to a complex underlying variable, such as dominance. third, identity is not easily categorized. for example, even if we limited the investigation of gender to a binary conceptualization, and only to cisgender individuals, the categories of “cisgender male” and “cisgender female” are still very broad. the stereotypes associated with white and black women are very different (landrine, 1985). even within one of those subcategories, different groups of white or black women are perceived as more or less feminine and are expected to conform to stereotypes to different degrees. for example, female athletes are often perceived as less feminine (kauer and krane, 2006). the relationship between social identity and turn-taking is interesting and a worthwhile subject, but requires more detailed analysis than was possible in this paper. 3.4 overlapping talk finally, we analyzed overlapping talk during the trouble source turn. we hypothesized that overlapping talk could have either social or cognitive effects on oir fto. background speech interferes with comprehension, even more than meaningless noise (koelewijn et al., 2015, 2012; schneider 30 cognitive and social delays in the initiation of conversational repair et al., 2007). this finding suggests that listeners extract and process the information contained in both the target speech (from their interlocutor) and overlapping speech (from the background). if so, overlapping talk during the trouble source turn should increase cognitive processing, resulting in delayed oirs. overlapping talk during the trouble source turn may also have a social effect on oir ftos. participants work to maintain a positive public image, or face (goffman, 1955; lerner, 1996). when an utterance is face-threatening, that utterance is dispreferred and delayed. oirs are potentially face-threatening because they could be attributed to poor language comprehension, outsider status, or a lack of attention. if someone talks over the trouble source turn, the oir can be attributed to something other than the oir speaker. in such cases, the oir should be less face threatening. in contrast to the cognitive hypothesis, the social hypothesis predicts that overlapping talk during the trouble source turn should be related to shorter gaps before the oir. we operationalized overlapping talk as speech produced either by an interlocutor within the interaction, people in the background, or a television or radio, before the start of the oir. we did not analyze music or non-speech noises. while environmental noise has been linked to more open class oirs (dingemanse et al., 2015), given the data collection methods, noises that were loud to the transcriber may not have been loud to the interlocutors. we calculated two values: a binary variable indicating whether there was any overlapping speech during the trouble source turn, and the percentage of the trouble source turn that was overlapped. 4. analysis in natural conversation, it is impossible to isolate the effect of one predictor from all the others. some statistical tests, such as principal components analyses or decision trees, allow researchers to simulate several effects at once. however, many of these tests provide output that is difficult to interpret. multiple regression solves both problems. multiple regression can analyze multiple effects at once, while calculating the effect of each predictor on its own. the commonly used frequentist approach to inferential statistics computes the probability of the data or more extreme data under a null hypothesis, the alternative hypothesis being the negation of this null hypothesis. in contrast, bayesian statistics allow researchers to directly compare the predictive adequacy of different hypotheses. given that our goal was to compare cognitive and social models of oir fto, we decided to use bayesian multiple regression. in this paper, we will report bayes factors (bfab) representing the relative probability of the observed data under model a over model b. a larger bayes factor means that there is a larger difference between the two models. typically, researchers report bf10, and the larger it is, the more evidence there is for the alternative over the null (kass and raftery, 1995; jeffreys, 1998); however, in this paper, we will also discuss evidence for cognitive over social factors. we created three types of regression models. first, we analyzed models with only cognitive predictors. second, we analyzed models with only social predictors. we compared the cognitive and social models directly to see which model explained more variance in oir fto. finally, we created models that combined both social and cognitive predictors. for details on the regression models and residuals, see https://osf.io/stna6/. we performed all analyses with jasp (team, 2018; jasp-stats.org). for access to the datafile, see https://osf.io/rkuav/. 31 mertens and de ruiter 4.1 pre-screening regression models assume that the model residuals are normally distributed. our outcome variable, oir fto, was positively skewed, which resulted in positively skewed model residuals. to make the models more valid, we log-transformed the oir ftos. we added a value to each such that the smallest oir fto was 1, and then took the 10-based logarithm of the resulting value. as a result, the model does not assume a linear relationship between the predictors and oir fto, but between the predictors and the log-transformed oir fto. this means the effect of any independent variable on raw oir fto increased in magnitude as its value increased. as a result, we present effects as percentages – the percent difference in oir fto when a predictor variable is different. the percent difference indicates a greater raw difference when the value of the predictor variable is larger. the residuals of the resulting models were approximately normally distributed (see the annotated jasp file at https://osf.io/stna6/ for regression residuals). in addition, regression models assume independence of predictor variables. when predictor variables are related, regression models have greater false negative rates and inaccurately estimated effects. we found strong evidence for several relationships between predictors (for details on our prescreening, see the annotated jasp file at https://osf.io/fznht/). to mitigate the consequences of these relationships, we used jasp to compare all combinations of predictors in each category within regression model (cognitive, social, and overall). jasp determined the most predictive model within each category. below, we report the models that best predict log oir fto. finally, we determined whether overlapping talk during the trouble source turn should be added to the cognitive or social regression model. a mann-whitney u test found strong evidence that overlapped trouble source turns (median = 478ms) were followed by shorter oir ftos than nonoverlapped trouble source turns (median = 596ms, bf10 = 10.37). further, a kendall correlation found decisive evidence that oir fto was shorter when a greater proportion of the trouble source turn was overlapped (t = -0.13, 95% bayesian credible interval, henceforth referred to as ci = -0.19 – -0.07, bf10 = 421.93). this evidence suggested that overlapping talk during the trouble source turn may make oirs less socially dispreferred. therefore, we included overlapping talk during the trouble source turn in the social regression model. 4.2 missing data out of 456 oirs, 307 (67.32%) had data for all predictor variables. the relatively high rate of missing data is a common consequence of collecting natural data: we could not control for environmental noise or ask participants to speak into the microphone. this means that some data were excluded from the regression models. in excerpt 10, for example, speaker b resolved the problem on their own so there was no value for repair solution fto or duration. for each statistically significant effect found in the final models below, we performed a simple analysis including only the relevant predictor. these simple tests allowed us to analyze all the oirs with values for oir fto and the relevant predictor, even if there was missing data for another predictor. this also allowed us to use nonparametric versions of tests on raw oir fto. for example, we conducted a mann-whitney u test to investigate the effect of oir category (specific vs. ‘open’ class) on oir fto. any oir with a value for oir category was included in this test. excerpt 10: resolving oir without verbal repair solution 1 spa: nice little bu:m there, 2 (1.0) 32 cognitive and social delays in the initiation of conversational repair 3 spb: where. 4 (0.5) 5 spb: (laughing) oh yeah. to ensure that there were no relationships between missing data and oir fto, we created a binary variable for each predictor indicating whether data were missing or present. we then performed a series of bayesian mann-whitney u-tests to determine whether oir fto differed when each predictor was absent from the model. we found no evidence that oir fto differed when any predictor variable was missing. for details on missing data, see our jasp file at https://osf.io/h5ysb/. 4.3 cognitive factor regression we regressed the log-transformed oir fto on all cognitive factors (trouble source turn fto, disjointed status, length in both words and duration, average word frequency and syllable rate). the most predictive model included trouble source turn duration and fto, and the binary variable indicating that the trouble source turn was disjointed (r2 = 0.08, bf10 = 20.69). see table 1 for the unstandardized coefficient results for cognitive factors included in the best cognitive model. predictor mean predictor std. deviation 95% ci bfinclusion ts duration -2.90% 1.13 seconds -4.60% – -1.10% 1.47 ts fto 0.40% 6.81 seconds 0.05% – 0.70% 3.93 ts = disjointed 5.10% – 0.8% – 9.50% 4.66 table 1: regression coefficients for cognitive variables. 4.4 social factor regression we regressed the log-transformed oir fto on all social factors (the percent overlap in the trouble source turn, the binary variable indicating any overlap during the trouble source turn, oir category and duration ratio). the most predictive model consisted of oir category and the binary variable indicating overlap in the trouble source turn (r2 = 0.04, bf10 = 11.41). we found anecdotal evidence (the classification“anecdotal” is taken from appendix b in jeffreys, 1998) that when the trouble source turn was overlapped, the oir fto was 5.30% shorter (95%ci = -9.20% — -1.10%, bfinclusion = 3.11). we found substantial evidence that when the oir was open class, the oir fto was 6.00% longer (95%ci = 2.10% — 9.80%, bfinclusion = 8.11). 4.5 cognitive and social model comparison next, we compared the social and cognitive models in three stages. first, we added all cognitive factors to the most predictive social model (oir category and the binary variable indicating overlap in the trouble source turn). the most predictive model added trouble source turn fto and trouble source turn duratino to the most predictive social model (bf10 = 243.39, r2 = 0.10). see table 2 for the unstandardized coefficients for the resulting model. second, we added all social factors to the most predictive cognitive model (trouble source turn duration, trouble source turn fto, and the binary variable indicating disjointed trouble source turns). the most predictive model added all social factors: oir category, duration ratio, the percent 33 mertens and de ruiter predictor mean predictor std. deviation 95% ci bfinclusion cognitive ts duration -2.60% 1.13 seconds -4.40% – -0.80% 1.66 ts fto 0.40% 6.81 seconds 0.20% – 0.70% 4.34 social oir = open class 4.10% – 0.30% – 7.80% 1.00 ts = overlapped -6.00% – -10.00% – -2.00% 1.00 table 2: regression coefficient table for cognitive variables added to social model. overlap in the trouble source turn, and the binary variable indicating overlap in the trouble source turn (bf10 = 67.30, r2 = 0.11). see table 3 for the unstandardized results for this model. finally, we compared the models directly to determine whether one model was more likely than the other. we found anecdotal evidence that the cognitive model was more predictive than the social model (bf10 = 1.94). predictor mean predictor std. deviation 95% ci bfinclusion cognitive ts duration 4.00% 1.13 seconds 0.40% – 7.70% 1.00 ts fto 1.30% 6.81 seconds 0.70% – 1.90% 1.00 ts = disjointed -3.60% – -13.30% – 6.00% 1.00 social oir = open class -14.30% – -22.70% – -6.00% 67.58 duration ratio 6.50% 0.88 1.80% – 11.30% 13.83 % of ts overlapped -0.10% 25.64% -0.30% – 0.10% 1.49 ts = overlapped -1.00% – -14.80% – 12.70% 0.97 table 3: regression coefficient table for social variables added to cognitive model. 4.6 overall model we regressed the log-transformed oir fto on all predictor variables. the most predictive model included both cognitive (trouble source turn duration, trouble source turn fto) and social (the binary variable marking overlapped trouble source turns, oir category) variables (r2 = 0.11, bf10 = 33848.68). we found strong evidence for the overall model over the social model (bf10 = 12.36) and substantial evidence for the overall model over the cognitive model (bf10 = 6.35). see table 4 for the unstandardized coefficients. the residuals approximated a normal distribution. 34 cognitive and social delays in the initiation of conversational repair predictor mean predictor std. deviation 95% ci bfinclusion cognitive ts duration -3.00% 1.13 seconds -5.00% – -1.00% 1.14 ts fto 0.40% 6.81 seconds 0.10% – 0.70% 2.85 social oir = open class 5.20% – 1.00% – 9.40% 2.47 ts = overlapped -5.50% – -9.80% – -1.10% 1.10 table 4: coefficient table for overall model (unstandardized). 5. discussion interlocutors initiated oirs later when the trouble source turn (1) had less discourse context, (2) was not overlapped by other talk, or (3) was shorter. finally, open class oirs were initiated later than specific oirs. adding cognitive variables to the social model, or social variables to the cognitive model, improved model predictions. the best overall model included both types of measures. 5.1 discourse context when the trouble source turn fto was longer, oir fto was longer. 4 shows that this effect persisted in a simple post-test (t = 0.15, bf10 = 538.16). during the pre-screening process, we also found substantial evidence that a greater proportion of disjointed trouble source turns (≈67%) were more followed by open class oirs than non-disjointed trouble source turns (≈54%). listeners use discourse context to anticipate features of the trouble source turn. accurate predictions decrease the information the listener has to process, either by pre-activating upcoming representations or easing the integration of new information (kuperberg and jaeger, 2016). therefore, it may be that discourse context helps the listener process the correctly predicted information in the trouble source turn, reducing the time needed to respond. in addition, if a communication problem violates a prediction, it may be easier to notice. neurological evidence suggests that listeners process speech act mispredictions approximately 300ms after the start of the turn (gisladottir et al., 2015). this is just a few words into the unfolding turn. at the same time, violated speech act predictions have been related to more open class oirs (drew, 1997), suggesting that prediction can increase the size of the comprehension problem. together, these findings suggest that predictions help the listener identify problems, but make it harder for the listener to diagnose or resolve problems. attention may also contribute to this effect. some of our trouble source turns had long ftos, even over 30 seconds. these trouble source turns likely started a sequence. from the audio recordings, it appeared that many occurred while the interlocutors were engaged in a concurrent activity, such as walking down the street or doing the dishes. however, we do not have access to the video data to confirm whether the listener was distracted or multitasking. in addition, no trouble source turns began with greetings (e.g., “hi,”) or address terms (e.g., “barbara, what are you up to?”). such prefaces attract the listener’s attention. the listeners may not have been attending to trouble source turns with very long ftos. listeners still process the general characteristics, but not the content, of speech they are not currently attending to (arons, 1992). this hypothesis may explain 35 mertens and de ruiter figure 4: positive correlation between log(trouble source fto) and log(oir fto) why interlocutors produce more open oirs when multitasking (dingemanse et al., 2015). particularly intriguing is the finding that a listener requires approximately 170ms to switch their attention to, and begin to process the content of, an utterance (cherry and taylor, 1954). this is the exact difference in the median oir fto after disjointed (median = 674ms) and non-disjointed (median = 504ms) trouble source turns. this evidence raises several new and intriguing questions about the role of attention in conversation. do listeners respond more quickly to disjointed turns that start with address terms? are all disjointed utterances harder to process, or is this effect unique to trouble source turns that are especially difficult to comprehend? in summary, one explanation for the increased oir fto after trouble source turns with less discourse context is that the listener needed more time to direct their attention to the trouble source turn. recent work in conversation analysis suggests that there are different types of lapses in conversation (hoey, 2015). some lapses are expected given the previous conversation, some are accounted for by parallel activities, and some are conspicuous. in this paper, we did not distinguish between these lapses. future research could examine whether the type of lapse preceding the trouble source turn influences the timing of oirs. one possibility is that conspicuous lapses draw the interlocutors’ attention, so when a response is finally given it will be easier to process. in contrast, if a gap follows the end of a sequence, interlocutors may find it more difficult to process the utterance after this gap. another alternative explanation could be that all ftos in a conversation are correlated; conversations are either fast or slow. however, repair solution fto was not correlated with oir fto (t = 0.08, bf10 = 0.93, 95% ci = 0.15 – 0.01), although repair solution fto was positively correlated with trouble source turn fto (t = 0.11, bf10 = 15.65, 95% ci = 0.18 – 0.05). it is difficult to 36 cognitive and social delays in the initiation of conversational repair interpret these findings; future research should compare oir ftos to a distribution of ftos from across the entire conversation. 5.2 overlapping talk during the trouble source turn overlapped trouble source turns were followed by shorter oir ftos (figure 5). this effect is discrete and not continuous: more overlapping talk does not relate to even shorter oir ftos. overlapping talk during the trouble source turn could provide a face-saving explanation for the oir; the interlocutors may attribute the problem to the environment, and not the person who produced the trouble source turn (kelley, 1973). as a result, there is less social cost for oirs after overlapped trouble source turns, regardless of the duration of the overlap. the oir producer may not have to account for the oir to the same degree (robinson, 2016). figure 5: oir fto after overlapped and non-overlapped trouble source turns. this result is the opposite of what we would expect if overlapping talk had a cognitive effect on timing in conversation. if listeners processed both target and overlapping speech, they would take longer to respond to an overlapped turn. the findings of this study do not rule out the possibility that overlapped turns take more cognitive processing to comprehend. however, listeners may not have the social prerogative to attempt to cognitively repair the problem if the oir is not face threatening. even if the interlocutor has not finished processing all the possible information from the turn, they may initiate repair if there is little cost for doing so. interestingly, during our data pre-screening we found that oir category was not predicted by the presence of overlapping talk during the trouble source turn (bf10 = 0.29). this was unexpected — previous analyses found that open class repair initiations were more likely in loud or distracting environments (dingemanse et al., 2015). however, we did not analyze non-speech noises or distractions. further research is needed to determine whether speech, other noises, and concurrent activities have different effects on oirs. 37 mertens and de ruiter we also did not distinguish between different sources of overlapping talk. if the trouble source turn overlapped with the previous turn (negative trouble source turn fto), then the trouble source turn speaker may be considered responsible for the problem. in contrast, if the eventual oir producer interrupted the trouble source, the responsibility may lie more with the oir producer. future research should explore the social and cognitive differences between general background noise, a negative trouble source fto, and mid-turn overlap during the trouble source. 5.3 oir category open class oirs (e.g., “sorry?”) were initiated later than specific oirs (e.g., “who?”). this finding supports previous analyses of oirs (kendrick, 2015b; schegloff et al., 1977) and preference organization (roberts et al., 2015; kendrick and torreira, 2015). we extended these findings by accounting for other effects that affect timing in conversation. figure 6: probability density of oir fto separated by oir category (open vs. specific) we also attempted to understand why specific oirs were produced after shorter oirs. we hypothesized that interlocutors delay open class oirs because open class oirs are less progressive (dingemanse et al., 2015). a kendall’s correlation did find substantial evidence that duration ratio was slightly positively correlated with oir fto (t = 0.10, bf10 = 5.92). however, once we accounted for other variables, duration ratio did not explain unique variability in oir fto. therefore, we suggest that the relationship between duration ratio and oir fto is indirect. specific oirs were initiated earlier, and were followed by shorter repair solutions, but specific oirs were not initiated earlier because they were followed by shorter solutions. from a psychological perspective, this finding is not surprising. for the oir producer to account for the expected length of the repair solution, they would first have to estimate the length of the repair solution. but of course the oir producer does not know the repair solution. further, simulating the repair solution duration would be another cognitively demanding task in a time-sensitive scenario. 38 cognitive and social delays in the initiation of conversational repair in our pre-screening, we found substantial evidence that open class oirs (m = 387ms) were followed by shorter repair solution ftos than specific oirs (m = 506ms, w = 13307.00, bf10 = 5.68). this finding contradicted our assumption that open class oirs would be followed by longer repair solution ftos, and therefore be even less progressive. this surprising result may be due to the nature of open vs. specific repair solutions. solutions to open class oirs may be more likely to repeat the trouble source. in contrast, solutions to specific oirs may be more likely to incorporate new information or words. as a consequence, specific repair solutions may require more planning. repair solutions highlight an interesting tension between cognitive and social factors in conversation that should be explored in a separate, focused analysis. so, why are open class oirs initiated later than specific oirs? an examination of figure 6 suggests two effects. first, the open class oir fto mode (approximately 430ms) is somewhat longer than the specific oir fto mode (approximately 390ms). second, most oir ftos longer than 2.5 seconds precede open class oirs. the first effect — the difference in modes — may be a cognitive effect. if open class oirs mark larger comprehension problems, then the oir speaker may spend more time attempting to cognitively repair that comprehension problem. the second effect — the right tail of the distribution — may be a social effect. if open class oirs are more likely to be extremely dispreferred, then the interlocutor who produces the open class oir may wait an especially long time to produce the oir. 5.4 trouble source turn duration trouble source turns with longer durations were followed by shorter oir ftos. when the trouble source turn was longer, the listener had more time to plan their upcoming oir. in contrast, roberts et al. (2015) found that longer turns were followed by longer ftos. however, roberts et al., 2015 analyzed all utterances in an entire corpus. we restricted our analysis to a very small proportion of all utterances to control for a variety of confounds. it may be that confounds are responsible for the effect in roberts et al. (2015). another possible explanation is that oirs have a unique relationship with the duration of the previous turn. it is possible that once listeners realize there is a comprehension problem, they stop comprehending the remaining information in the turn. instead, they immediately begin planning their oir. if so, then the amount of information contained in a trouble source turn is less relevant to the timing of oirs. either way, we found that longer trouble source turns were related to shorter oir ftos. 6. conclusion this study found that (1) discourse context, (2) overlapping talk during the trouble source turn, (3) trouble source turn duration, and (4) oir category play a role in the timing of other-initiations of repair. first, we found that discourse context allows the listener to predict the upcoming turn. these predictions facilitate processing information that matches the prediction, and noticing information that violates the prediction. therefore we concluded that the relationship between discourse context and oir fto is a cognitive effect. second, overlapped trouble source turns were followed by shorter oir ftos. if overlapping talk had a cognitive effect on oir fto, we would expect that oir fto would be longer (koelewijn et al., 2012, 2015; schneider et al., 2007). in addition, we would expect a continuous effect, where oir fto would be longer when there was more overlapping talk during the trouble source turn. however, we found that any overlapping talk during the trouble source resulted in similarly short39 mertens and de ruiter ened oir ftos. when there is overlapping talk, the interlocutors can attribute communication problems to that talk; the oir speaker is not accountable for the problem. third, longer trouble source turns were followed by shorter oir ftos. we believe this is a cognitive effect: the longer the trouble source turn, the more time the oir producer has to plan the oir. finally, our analysis cannot conclusively determine whether oir category has an exclusively social or cognitive effect on oir fto. our hypothesis – that open class oirs are especially delayed because they delay progressivity more than specific oirs – was not supported by our analysis. it may be that open class oirs are, in fact, more dispreferred than specific oirs because they admit to a larger communication problem, and therefore may be more face-threatening. on the other hand, this effect may be cognitive effect: that open class oirs are more likely when the trouble source turn recipient has a larger problem of hearing or understanding, and therefore induce more cognitive processing. further research would be needed to determine if oir category should be categorized as a cognitive or social factor. more generally, these findings show that both cognitive and social factors play a role in the timing of oirs. interlocutors are constrained by aspects of cognitive processing: they cannot speak until they are cognitively ready to do so. cognitive limitations put a lower bound on the time required before initiating speech. however, interlocutors are also sensitive to social rules: they sometimes are cognitively able to initiate speech earlier, but nevertheless delay because of social factors. so the most important general finding from our study is that the timing of an utterance is influenced by both cognitive and social constraints. surprisingly, we found strong evidence for the null-hypothesis that there is no effect of average word frequency in the trouble source turn on oir fto (bf01 = 14.76). many studies have found that listeners respond to frequent words more quickly (goldinger, 1996). perhaps this effect is only detectable using highly controlled experimental paradigms like lexical decision or picture naming. another possibility is that interlocutors in conversation do not process language in the same way as participants in experiments. or it may be that interlocutors process trouble source turns differently than other turns; roberts et al. (2015) did find that word frequency predicted turn transition times in a corpus of natural data. it may be that trouble source turns are unique in that the frequency of the words in the turn is less influential. this finding raises a provocative question: how many well-established psycholinguistic findings, like frequency effects or semantic priming, are relevant in natural interaction, and if they are relevant, in which types of sequences? conversation, one of the crown jewels of human cognition, requires the streamlined interaction of a variety of skills and abilities. in this paper, we examined how social attribution, attention, attribution, and prediction come together when humans resolve miscommunication. to understand the mechanisms underlying conversational interaction we have to take into account both cognitive social constraints. acknowledgements we would like to thank dr. saul albert and dr. ariel cohen-goldberg for their guidance in the early stages of this project. funding for this project was provided by the united states air force laboratory. 40 cognitive and social delays in the initiation of conversational repair references saul albert, laura e. de ruiter, and jan p. de ruiter. cabnc: the jeffersonian transcription of the spoken british national corpus, 2015. barry arons. a review of the cocktail party effect. journal of the american voice i/o society, 12: 35–50, 1992. doi: 10.1.1.30.7556. alan baddeley. working memory and language: an overview. journal of communication disorders, 36(3):189–208, 2003. paul boersma. praat: doing phonetics by computer. http://www.praat.org/, 2006. sara bögels, kobin h. kendrick, and stephen c. levinson. conversational expectations get revised as response latencies unfold. language, cognition and neuroscience, 35(6):766–779, 2020. e. colin cherry and w. k. taylor. some further experiments upon the recognition of speech, with one and with two ears. the journal of the acoustical society of america, 26(4):554–559, 1954. herbert h. clark and jean e. fox tree. using uh and um in spontaneous speaking. cognition, 84 (1):73–111, 2002. herbert h. clark and edward f. schaefer. collaborating on contributions to conversations. language and cognitive processes, 2(1):19–41, 1987. steven e. clayman. sequence and solidarity. in advances in group processes. emerald group publishing limited, 2002. steven e. clayman. turn-constructional units and the transition-relevance place. the handbook of conversation analysis, pages 150–166, 2013. jan p. de ruiter and chris cummins. a model of intentional communication: airbus (asymmetric intention recognition with bayesian updating of signals). proceedings of semdial 2012, pages 149–50, 2012. jan p. de ruiter, holger mitterer, and nick j. enfield. projecting the end of a speaker’s turn: a cognitive cornerstone of conversation. language, 82(3):515–535, 2006. mark dingemanse, seán g. roberts, julija baranova, joe blythe, paul drew, simeon floyd, rosa s. gisladottir, kobin h. kendrick, stephen c. levinson, elizabeth. manrique, et al. universal principles in the repair of communication problems. plos one, 10(9):e0136100, 2015. franciscus cornelis donders. on the speed of mental processes. acta psychologica, 30:412–431, 1969. paul drew. ‘open’ class repair initiators in response to sequential sources of troubles in conversation. journal of pragmatics, 28(1):69–101, 1997. fernanda ferreira. effects of length and syntactic complexity on initiation times for prepared utterances. journal of memory and language, 30(2):210–233, 1991. 41 mertens and de ruiter janet dean fodor and atsu inoue. syntactic features in reanalysis: positive and negative symptoms. journal of psycholinguistic research, 29(1):25–36, 2000. rosa s. gisladottir, dorothee j. chwilla, and stephen c. levinson. conversation electrified: erp correlates of speech act recognition in underspecified utterances. plos one, 10(3):e0120068, 2015. erving goffman. on face-work: an analysis of ritual elements in social interaction. psychiatry, 18 (3):213–231, 1955. stephen d. goldinger. auditory lexical decision. language and cognitive processes, 11(6):559– 568, 1996. herbert p. grice. logic and conversation. in speech acts, pages 41–58. brill, 1975. laura gwilliams, tal linzen, david poeppel, and alec marantz. in spoken word recognition, the future predicts the past. journal of neuroscience, 38(35):7585–7599, 2018. alexa hepburn and galina b. bolden. the conversation analytic approach to transcription. the handbook of conversation analysis, 57:76, 2013. john heritage. well-prefaced turns in english conversation: a conversation analytic perspective. journal of pragmatics, 88:88–104, 2015. elliott m. hoey. lapses: how people arrive at, and deal with, discontinuities in talk. research on language and social interaction, 48(4):430–453, 2015. peter indefrey and willem j. m. levelt. the spatial and temporal signatures of word production components. cognition, 92(1-2):101–144, 2004. stefanie jansen, hendrik wesselmeier, jan p. de ruiter, and horst m. mueller. using the readiness potential of button-press and verbal response within spoken language processing. journal of neuroscience methods, 232:24–29, 2014. harold jeffreys. the theory of probability. oup oxford, 1998. robert e. kass and adrian e. raftery. bayes factors. journal of the american statistical association, 90(430):773–795, 1995. kerrie j. kauer and vikki krane. “scary dykes” and “feminine queens”: stereotypes and female collegiate athletes. women in sport & physical activity journal, 15(1):42, 2006. harold h. kelley. the processes of causal attribution. american psychologist, 28(2):107, 1973. gerard kempen and pieter huijbers. the lexicalization process in sentence production and naming: indirect election of words. cognition, 14(2):185–209, 1983. kobin h. kendrick. other-initiated repair in english. open linguistics, 1, 2015a. kobin h. kendrick. the intersection of turn-taking and repair: the timing of other-initiations of repair in conversation. frontiers in psychology, 6(250):10–3389, 2015b. 42 cognitive and social delays in the initiation of conversational repair kobin h. kendrick and francisco torreira. the timing and construction of preference: a quantitative study. discourse processes, 52(4):255–289, 2015. thomas koelewijn, adriana a. zekveld, joost m. festen, and sophia e. kramer. pupil dilation uncovers extra listening effort in the presence of a single-talker masker. ear and hearing, 33(2): 291–300, 2012. thomas koelewijn, hilde de kluiver, barbara g. shinn-cunningham, adriana a. zekveld, and sophia e. kramer. the pupil response reveals increased listening effort when it is difficult to focus attention. hearing research, 323:81–90, 2015. gina r. kuperberg and t. florian jaeger. what do we mean by prediction in language comprehension? language, cognition and neuroscience, 31(1):32–59, 2016. hope landrine. race× class stereotypes of women. sex roles, 13(1-2):65–75, 1985. gene h. lerner. finding “face” in the preference structures of talk-in-interaction. social psychology quarterly, pages 303–321, 1996. stephen c. levinson. speech acts. in oxford handbook of pragmatics, pages 199–216. oxford university press, 2017. lilla magyari and jan p. de ruiter. prediction of turn-ends based on anticipation of upcoming words. frontiers in psychology, 3:376, 2012. lilla magyari, jan p. de ruiter, and stephen c. levinson. temporal preparation for speaking in question-answer sequences. frontiers in psychology, 8:211, 2017. nobuko osada. listening comprehension research: a brief review of the past thirty years. dialogue, 3(1):53–66, 2004. tepring piquado, derek isaacowitz, and arthur wingfield. pupillometry as a measure of cognitive effort in younger and older adults. psychophysiology, 47(3):560–569, 2010. anita pomerantz. agreeing and disagreeing with assessments: some features of preferred/dispreferred turn shaped. in structures of social action: studies in conversation analysis. cambridge university press, 1984. seán g. roberts, francisco torreira, and stephen c. levinson. the effects of processing and sequence organization on the timing of turn taking: a corpus study. frontiers in psychology, 6: 509, 2015. jeffrey d. robinson. accountability in social interaction. oxford university press, 2016. jeffrey d. robinson and heidi kevoe-feldman. using full repeats to initiate repair on others’ questions. research on language and social interaction, 43(3):232–259, 2010. christoph rühlemann. how long does it take to say ‘well’? evidence from the audio bnc. corpus pragmatics, 3(1):49–66, 2019. 43 mertens and de ruiter harvey sacks, emanuel a. schegloff, and gail jefferson. a simplest systematics for the organization of turn taking for conversation. in studies in the organization of conversational interaction, pages 7–55. elsevier, 1978. emanuel a. schegloff. sequencing in conversational openings 1. american anthropologist, 70(6): 1075–1095, 1968. emanuel a. schegloff. the relevance of repair to syntax-for-conversation. in discourse and syntax, pages 261–286. brill, 1979. emanuel a. schegloff. discourse as an interactional achievement: some uses of ‘uh huh’ and other things that come between sentences. analyzing discourse: text and talk, 71:93, 1982. emanuel a. schegloff. sequence organization in interaction: a primer in conversation analysis i, volume 1. cambridge university press, 2007. emanuel a. schegloff, gail jefferson, and harvey sacks. the preference for self-correction in the organization of repair in conversation. language, 53(2):361–382, 1977. bruce a. schneider, liang li, and meredyth daneman. how competing speech interferes with speech comprehension in everyday listening situations. journal of the american academy of audiology, 18(7):559–572, 2007. saul sternberg, stephen monsell, ronald l. knoll, and charles e. wright. the latency and duration of rapid movement sequences: comparisons of speech and typewriting. in information processing in motor control and learning, pages 117–152. elsevier, 1978. tanya stivers and jeffrey d. robinson. a preference for progressivity in interaction. language in society, pages 367–392, 2006. tanya stivers, nicholas j. enfield, penelope brown, christina englert, makoto hayashi, trine heinemann, gertie hoymann, federico rossano, jan p. de ruiter, kyung-eun yoon, et al. universals and cultural variation in turn-taking in conversation. proceedings of the national academy of sciences, 106(26):10587–10592, 2009. jasp team. jasp (version 0.10.0). computer software, 2018. url https://jasp-stats. org. walter j. b. van heuven, pawel mandera, emmanuel keuleers, and marc brysbaert. subtlexuk: a new and improved word frequency database for british english. quarterly journal of experimental psychology, 67(6):1176–1190, 2014. jeffrey j. walczyk, karen s. roper, eric seemann, and angela m. humphrey. cognitive mechanisms underlying lying to questions: response time as a cue to deception. applied cognitive psychology, 17(7):755–774, 2003. doi: 10.1002/acp.914. candace west and don h. zimmerman. small insults: a study of interruptions in cross-sex conversations between unacquainted persons. critical concepts in psychology. routledge/taylor & francis group, new york, ny, us, 2015. isbn 978-0-415-66094-5 (hardcover). victor h. yngve. on getting a word in edgewise. in chicago linguistics society, 6th meeting, 1970, pages 567–578, 1970. 44 d&d-submitted-250429-copy dialogue & discourse 16(1) 91-113 doi: 10.5210/dad.2025.104 ©2025 jenny myrendahl this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). repair of claimed non-understanding of word meaning in online discussion forum interaction jenny myrendal jenny.myrendal@gu.se department of education, communication & learning university of gothenburg editor: patrick g.t. healey submitted: 25th may 2023, 29th april 2024, april 29th 2025; accepted 05/2025; published online 05/2025. abstract this article describes how participants in online discussion forums manage claimed nonunderstanding of word meaning, specifically when one participant displays insufficient understanding by requesting meta-linguistic clarification. claimed non-understanding refers to cases where a participant signals a lack of understanding in a way that invites repair. by engaging in a short word meaning negotiation sequence, the participants collaboratively repair the issue of claimed non-understanding and can move on with the discussion on topic. in some cases, however, participants behave in ways that break the normative pattern of interaction and do not enter into the anticipated sequence of repair dealing with the lack of understanding. the analysis of these deviant cases reveals the participants’ own normative orientations in repair of claimed non-understanding of word meaning, and thus provide evidence that there is an underlying organization of repair dealing with such issues in online discussion forum interaction. keywords: word meaning negotiation, repair, online interaction, semantic coordination, lexical pragmatics 1 introduction in communication, interlocutors regularly run into trouble of understanding what the other person has just said or meant. in such situations, interlocutors typically make use of practices of repair, an organized set of techniques that can be used to draw attention to and rectify such problems (schegloff et al. 1977). this article addresses the issue of how participants in online discussion forum interaction deal with interactional problems that originate in insufficient understanding of word meaning. interlocutors facing such issues may deal with the problems by entering into a sequence of word meaning negotiation specifically addressing the situated meaning of the problematic word. word meaning negotiation (henceforth wmn) occurs when participants who are engaged in a discussion about a particular topic remark on a word choice made by another participant, thereby initiating a meta-linguistic sequence in which the meaning of the word is questioned and up for negotiation. in such sequences, participants typically shift their attention from the discussed topic to the words used and their associated meanings. according to myrendal (2015) wmn can be sorted into two main types, depending on the origin or cause of the negotiation. one type of wmn encompasses sequences that are caused by disagreement between participants concerning what a word can or should mean in the situated interaction (myrendal 2019). the other type of wmn occurs when there is insufficient understanding between participants regarding the myrendal 92 meaning of a word, and the participants must initiate repair to restore enough mutual understanding regarding the meaning of the word to be able to continue the discussion on topic. this article presents a study that examines this form of repair: wmn caused by insufficient understanding of word meaning in online discussion forum interaction. this type of interaction was chosen for the study for several reasons. firstly, because the mode of interaction is asynchronous, participants have more time to contemplate and reflect on the words and meanings they exchange. additionally, participants in online forums usually have little prior knowledge of each other and must work together to ensure that they understand what they mean by the words they choose during discussions. furthermore, online forums lack the non-verbal cues present in spoken interaction, so verbalizing processes of interpretation and understanding become more important, and explicit, than in spoken interaction where meta-communicative functions such as grounding and turn-taking are performed through gestures, body language, gaze and prosody (clark & brennan 1991). the study adopts a conversation analytic perspective, treating insufficient understanding, or nonunderstanding, as something that participants publicly display or claim through their written contributions – thus making it observable and available for repair. this approach makes it possible to examine how participants manage problems of understanding in an asynchronous, text-based setting, despite the lack of immediate physical or temporal co-presence. this study aims to contribute to a greater understanding of how instances of claimed nonunderstanding of word meaning are repaired in human interaction that takes place in an online setting. providing linguistic insight of this kind is important, as it contributes to building knowledge about the linguistic and communicative strategies that people use to repair claimed nonunderstandings and clarify meaning in online interactions. this knowledge can both inform theories of language and communication and inform the future design of online platforms and support the development of tools and strategies to facilitate communication and mitigate misunderstandings in online interactions. 2 repair the notion of repair refers to the practice of interrupting the ongoing sequence of interaction to deal with potential problems in speaking, hearing or understanding. studies of repair have shown that repairs are a pervasive aspect of dialogue and highly prevalent in naturally occurring interaction (colman & healey 2011, dingemanse et al 2015, purver et al 2018). hough and purver (2013) found that repairs occur approximately once every 25 words in conversational speech. healey et al (2018) have shown that misunderstandings are typically dealt with on the fly and suggest that “running repairs” is a crucial driver for semantic coordination in dialogue. repair is used so that “the interaction does not freeze in its place when trouble arises, so that intersubjectivity is maintained or restored” (schegloff 2007). repair can thus characterize all kinds of problems ranging from failure to hear or be heard, use of a wrong word or failure to come up with the intended word at a particular moment, to more general problems of understanding. repair can serve as an interface for cross-disciplinary research between ca and cognitive approaches to human interaction (albert & de ruiter, 2018). ca studies have shown that the practice of repair in spoken interaction seems to orient towards a preference organization in which self-initiated repair is the preferred alternative, which means that the speaker who causes the initial trouble also acknowledges this and self-corrects by producing a repair solution. in other cases, it is a participant other than the speaker who initiates the repair, which is called other-initiated repair (schegloff et al. 1977). schegloff et al. (1977) suggest that mechanisms for other-initiated repair are essentially techniques for locating the trouble source. when other-initiated repair is initiated, the participant can choose a repair initiation form which specifically locates the trouble source by pointing to the repairable item, or they can choose a repair initiation form which addresses the whole prior turn as repair of non-understanding of word meaning in online discussion forum interaction 93 problematic (drew 1997). schegloff (2007) points out that other-initiated repair can sometimes be used for other issues than repairing instances of miscommunication, for example to offer a predisagreement or a pre-rejection to something said in a previous turn. in this way, other-initiated repair sequences can be used in the same way as delays, hedges and accounts to indicate disagreement with a prior speaker’s assessment, and they also offer the speaker a chance of modifying the utterance facing the possible rejection or disagreement before the open disagreement is a known fact (schegloff 2007: 102-103). 2.1 computer-mediated communication and repair in recent years, a growing interest has developed around interaction practices that occur in online environments and researchers have worked to adapt ca to online settings to be able to study computer-mediated interaction (fitzgerald et al 2017; giles et al 2017; harris & church, 2019; meredith, 2019; warren & paulus 2020). since ca was originally developed for analysis of spoken interaction, one limitation of ca when applied to online interaction is that it works from the assumption that communication is linear, and that turns follow each other in a chronological sequence (giles et al. 2017). in computer-mediated communication (cmc), sequential coherence tends to be violated due to disrupted turn adjacency which occurs when messages that are related to each other end up being separated by unrelated intervening messages (garcia & jacobs 1999; marcoccia 2004; lewis 2005; gibson 2009; herring 2010). although this can also be the case in spoken interaction, i.e., that turns that are related in a sequence are interrupted by intervening talk, this happens at a much larger scale in cmc since there are more participants interacting simultaneously and participants are not in the same physical space when interacting. several studies on cmc have shown that even when chronologically adjacent messages appear unrelated, interlocutors develop strategies for overcoming the problems of interactional incoherence, making sure that it will be possible to interpret how messages are connected to each other relationally even when the chronological adjacency is disrupted (simpson 2005; farina 2018). these different strategies are employed collaboratively and typically make use of the discussion platform’s various interaction affordances (wellman et al. 2003; arminen et al. 2016; meredith 2017). the notions phantom adjacency (schönfeldt & golato 2003) and virtual adjacency (garcia & jacobs 1999) have been used to describe this collaborative achievement of coherence in cmc, stemming from the communication strategies used by participants to form distinct sequences within the multi-party communication. farina (2018) points out that posts in a discussion thread only make sense when they are interpreted as part of a sequence, and not if analyzed as separate items. thus far the practices of repair in cmc have mainly focused on synchronous forms of online interaction, i.e., communication that takes place in real time where participants are mutually copresent in the chat or instant messaging system when they are interacting with each other. jacobs and garcia (2013) found that participants in chat room interaction mimic procedures used in verbal communication to avoid and repair interactional troubles, and that the procedures of repair are adapted to the affordances and constraints of the cmc system. on a similar note, jepson (2005) compared repair move patterns in online voice chats and online synchronous text chats and found that the voice chats contained a significantly higher number of repair moves than the text-based chats. jepson found that in synchronous text-based chats, that repair initiations were often produced by a repetition of the trouble source, for example by producing a clarification request such as “what do you mean by x?” or a confirmation check such as “did you mean/say x?”. lee (2008) studied expert-to-novice synchronous online interaction within a learning environment setting and found that experts used confirmation checks on the form “x?” as a repair initiation device to draw attention to linguistic and grammatical errors made by the novice as an attempt to draw attention to a potential trouble source and influence the novice to make a correction or clarification. recent work on asynchronous interaction, such as goddard and gillespie’s (2025) study of reddit discussions, has shown that repair initiations are common in these settings but are often left uncompleted. this highlights that while participants orient to the need for clarification, the myrendal 94 completion of repair is not always achieved, especially in environments where responses are delayed or where participants do not remain engaged. their findings highlight that repair in asynchronous cmc is often unstable and incomplete, emphasizing the need to examine how such sequences develop – or break down – across extended interactional timelines. in summary, research on cmc and repair has concluded that although this form of interaction may suffer from disrupted turn adjacency, participants develop strategies for reconstructing relational adjacency between posts that are not chronologically adjacent. studies have also shown that participants in cmc use similar repair strategies as in verbal communication, although slightly adapted to the online medium’s constraints and affordances. this study will focus its attention to one source of interactional trouble: issues that originate in insufficient understanding of word meaning. it explores asynchronous online interactions, addressing a gap in the literature that has predominantly examined synchronous online interactions in the context of repair practices. 3 method and materials this study uses a conversation-analytic approach to analyze forum discussions, a form of asynchronous cmc. according to meredith (2019), this type of computer-mediated interaction can in some ways be considered more natural than recorded and transcribed verbal conversation, as it can be captured without the researcher’s intervention. the reason the ca approach was chosen to study this form of online interaction is because it is a method specifically designed to examine the structure and organization of talk-in-interaction and has been proven useful for studying processes of interaction and the organization principles on which interaction rests (sacks et al. 1974; jefferson 1987; schegloff 2007; sidnell 2010). like face-to-face conversation, online discussions involve turn-taking, repair, and other interactive practices that ca is well equipped to analyze. by using ca to examine online discussions, researchers can gain insight into how these conversations are organized and how participants use language to achieve their communicative goals (stommel 2008). in addition, since interaction in online discussion forums is asynchronous and involves many participants who interact simultaneously and by overlapping contributions in branching conversations, it can be challenging to analyze online discussions using traditional methods such as discourse analysis, which often assumes a linear sequence of events. ca, on the other hand, is well-suited for analyzing asynchronous discussions because it focuses on the organization of talk and how speakers respond to each other’s actions, and by doing so it offers a micro-level analysis of the structure of talk that the participants themselves organize. ca relies on the assumption that human interaction is a highly organized activity. by closely examining instances of naturally occurring interaction, the researcher attempts to reveal how participants manage their interaction, particularly focusing on how turns are related to prior turns and how the participants themselves display understanding of the communicative practices of which they are a part (goodwin & heritage 1990). all discussion posts in the data collection have been analyzed in terms of actions, which is an integral part of the conversation-analytic approach, focusing on how social actions are carried out and interpreted by the participants themselves (ten have 2007). ca assumes that turns in interaction are not just serially ordered, but also sequentially organized, which means that they relate to each other in a number of relevant ways (schegloff 2007). generally, the primary unit of analysis in ca research is sequences and turns-withinsequences (heritage 1984). according to antaki et al. (2005) and van hooijdonk and van charldorp (2019) sequentiality in online forums can be considered equal to sequentiality in face-to-face interaction, although clearly manifested in different ways than in verbal interaction. gibson (2009) suggests that analyzing aspects of sequentiality in asynchronous cmc essentially involves tracking responses through the discussion thread by identifying which posts they relate to in various ways. this study adopts gibson’s methodological framework, employing a focused approach to investigate how participants navigate word meaning negotiations in online discussions. the study repair of non-understanding of word meaning in online discussion forum interaction 95 is further guided by ten have’s (2007) specimen-based approach, which supports the detailed collection and analysis of naturally occurring instances that exhibit a particular phenomenon. some typical ca procedures applied in this study include: an inductive method of data retrieval, identification of regularities and patterns in instances of the researched phenomenon and building analytic accounts of single cases which are both particularized and generalized (seedhouse 2004). accordingly, this study carries out a close and in-depth analysis on interaction sequences where apparent non-understanding of word meaning occurs. in line with a conversation analytic perspective, instances of non-understanding in this study are not treated as objective cognitive states but as interactionally displayed or claimed understandings – that is, participants publicly orient to something as a problem of understanding, thereby making it available for repair. the purpose of the study is to examine how these participant orientations to non-understanding are treated as repair in asynchronous, online interaction and will focus on how they are indicated and reacted to, how they are negotiated between participants and how they ultimately are resolved. 3.1 data retrieval and selection criteria the data used in the study consists of 38 interaction sequences gathered from three large swedish discussion forums, familjeliv (www.familjeliv.se), flashback (www.flashback.org) and passagen debatt (www.debatt.passagen.se). the forums were chosen since they, at the time of the data collection, were three of the most popular discussion forums in sweden. since then, passagen debatt has shut down, but familjeliv and flashback are still very well visited and active forums. to identify word meaning negotiation sequences (wmns) within forum discussions, a systematic search process was employed, using specific utterance-initial constructions as potential indicators of wmn. these constructions, such as “what/how do you mean (by)?” and the repetition of a word as a question, signal a shift in the discussion from the initial topic to a meta-linguistic level. the sequences were retrieved by using five different swedish variations of the phrase “what/how do you mean (by)?1” as search expressions, specifically targeting these forums. the goal was to identify instances where the conversation shifted from the main topic to a deeper negotiation of the meanings of specific words. the data collection followed a systematic selection method, described in myrendal (2015, chapter 3), based on ten have’s (2007) specimen-based approach. the searches yielded approximately 1500 initial hits, reflecting 100 results per expression for each of the five swedish prompt phrases across three forums. each post was manually reviewed for evidence of a metalinguistic shift – defined as a transition in the conversation from topical discussion to an explicit focus on word meaning. this shift had to be evident through indicators such as meta-linguistic clarification requests and responses addressing either the meaning potential or the situated meaning of a word. cases were excluded if they lacked a meta-linguistic shift, such as when responses clarified general context but not the meaning of the targeted word. this ensured that inclusion was not based merely on the presence of a key phrase, but on interactional features indicating actual engagement in word meaning negotiation. threads were also excluded if they involved highly sensitive or personal content, such as discussions of miscarriage, deceased children, or sexual preferences, which were deliberately omitted for ethical reasons. although the forums were publicly accessible, the study followed ethical guidelines for internet research (e.g. markham & buchanan 2012), which emphasize that public availability does not eliminate the need for caution when dealing with emotionally charged or private material. this approach was adopted to protect participant privacy and minimize the risk of harm or discomfort to forum users. after this filtering, the final dataset comprised 38 analyzed sequences, representing approximately 2.5% of the 1500 1 the five swedish expressions used in the searches were: vadå x?, vad då x?, vaddå x? (all roughly translating as “what do you mean x?”), vad menar du med x? (“what do you mean by x?”), and hur menar du med x? (“how do you mean (by) x?”). these expressions are described in detail in myrendal (2015: 93). myrendal 96 reviewed posts. the selection of wmn sequences followed the step-by-step procedure outlined below. 1. locating the post with the problematic word. 2. identifying the start of the sequence by noting all possible indicators leading up to the metalinguistic shift. 3. reviewing the entire thread to determine relevant posts, including those directly discussing the identified word and others contributing indirectly (e.g., through pronouns referring to the word). 4. determining the end of the wmn sequence as the last post that significantly adds to the negotiation before the conversation concludes or shifts away. this systematic approach for retrieving wmns ensures thorough coverage of how participants negotiate and clarify word meanings in online discussions, drawing on methods for spotting miscommunication events as detailed by linell (1995: 185-187). in the excerpts presented in this article, the original swedish discussions have been translated into english, and the names of the participants have been anonymized into p1, p2 etc. each turn is labelled in brackets with the id number assigned to the post by the discussion forum, displaying its consecutive order in the discussion thread. in total, the 38 interaction sequences comprise 149 turns produced by 90 participants. the data collection comprises approximately 12 000 words. the careful selection and anonymization process ensures the integrity and confidentiality of the participants’ contributions while facilitating a detailed analysis of the conversational dynamics present in the forums. the sequences have been manually identified as instances of apparent non-understanding wmns (myrendal, 2015), which means that they all have in common that they originate in participant orientations to a lack of understanding about the situated meaning of a particular word used in the discussion. consequently, all sequences involve one turn in which a word is used that is retroactively identified as problematic, and at least one turn in which another participant requests meta-linguistic clarification regarding the meaning of the word, thus initiating repair. following the initial use of the word, and the indication that the word is not fully understood in the repair initiation, the discussion unfolds into a short repair sequence focusing on the meaning of the word and how it should be understood. most of the sequences in the data collection consist of either three or four turns. in the sequences, the first turn is the one containing the trouble source, the second turn is the one initiating repair, the third turn addresses the issue of repair and in cases where there is a fourth turn in the sequence, this generally contains a reaction to the clarification provided in the third turn. all repair initiation posts in the data collection bear great similarities with each other as they contain a metalinguistic clarification request targeting the perceived meaning problem of the trouble source (mirroring the specific search expressions employed to identify such instances). drawing on the classification of clarification requests outlined by purver (2004), the clarification requests observed in our data predominantly consist of non-reprise clarifications. this form involves participants restating or reformulating the information being clarified, often taking the shape of questions like “what do you mean by x?”. this pattern is exemplified in many of the dialogue excerpts in this article, for example in the dialogue of excerpt 1 (post #29) and excerpt 2 (post #2), where participants seek further explanation or definition. excerpt 1 p1 (27): men ska som sagt till husläkaren och ska säga att jag vill göra en helkroppsscanning och börja med medicin. p1 (27): but as i said, i’m going to the doctor and will say that i want to have a full body scan and start on medication. p2 (29): vad menar du med helkroppsscanning? p2 (29): what do you mean by full body scan? p1 (31): alltså en slags röntgen där dom ser inflammationerna. repair of non-understanding of word meaning in online discussion forum interaction 97 p1 (31): it’s like an x-ray where they see the inflammations. p2 (33): jag har aldrig fått nått sånt, var nyfiken på vad det var. jag har däremot fått händerna undersökta med ultraljud p2 (33): i have never had anything like that, was curious about what it was. i have had my hands examined with ultrasound though. excerpt 2 p1 (1): priset brukar väl ofta vara utan resning. ca 2-3 miljoner för standardhus. p1 (1): the price is normally without putting up. about 2-3 million for a standard house. p2 (2): vad menar du utan resning? ursäkta ifall frågan låter knas är inte så duktigt på sånt. p2 (2): what do you mean without putting up? sorry if the question sounds weird, am not good at this. p1 (3): dvs pris på själva huset från leverantören, men det är ju en byggsats. det ska byggas på plats, elektriker, rörmokare osv osv osv. kostar ca 1 miljon att resa huset. så plussa på ca 1 miljon på huspriset från leverantören. p1 (3): that is to say the price of the house from the manufacturer, it’s modular. it has to be constructed on site, electricians, plumbers etc etc etc. so plus about 1 million on top of the supply costs. p2 (4): jaha nu förstår jag bättre. men hur funkar det med avlopp och allt sånt. för det finns redan en sådan. p2 (4): oh, i see, now i understand better. but how does it work with plumbing and those kinds of things? because we already have that. the third turn in the sequence generally determines if the sequence turns into a repair sequence or not. up until this point, a word has been used by p1 that has been perceived as problematic in some way by p2. in 34 out of the 38 sequences, the requested meta-linguistic clarification is in fact provided in the third turn, which means that the repair initiation is addressed, and repair is carried out in the third turn of the sequence. in the four cases where the expected clarification is not provided in the third turn, the participants overtly address this in the negotiation, making the expectation explicit in the interaction itself. this will be discussed in more detail in the section concerning deviant cases. the fourth turn of the sequence is generally used to wrap up the side-sequence of repair and resume the main sequence of the discussion. in the sequences where the fourth turn is present, this turn is generally used for grounding purposes (clark 1996), which manifests itself in two main functions: • confirming understanding of the meaning of the trouble source, which means that the fourth turn is a part of the repair negotiation. • returning to the main discussion, i.e., exiting the micro-sequence of repair and continuing the discussion on topic. p2’s post #4 in excerpt 2 contains both functions. here, p2 first acknowledges that an understanding of the word has been grounded in the ongoing communication (“oh, i see, now i understand better.”), and then moves on in the discussion on topic (“but how does it work with plumbing and those kinds of things?”). myrendal 98 3.2 sequentiality in discussion forum interaction one issue to address when analyzing how repair is carried out in asynchronous cmc is to deal with the problems of interactional incoherence and disrupted turn adjacency brought up in previous research. to identify how turns relate to each other and form distinct sequences, there needs to be some way of interpreting how posts are connected to each other relationally even in cases where unrelated posts end up intervening the chronological adjacency of posts (gibson 2009; farina 2018). this study therefore investigates the different practices adopted by participants to show how posts are sequentially related to each other in this form of multi-party communication, unveiling how “virtual adjacency” (garcia & jacobs 1999) or “phantom adjacency” (schönfeldt & golato 2003) manifests itself and how chronologically unrelated posts form a distinct repair sequence in the interaction. one strategy for displaying how posts are related in asynchronous cmc involves explicitly quoting the post, indicating that the new post is responding to the previous one. many discussion forums have this feature built into the communication interface. both familjeliv and flashback have quoting buttons placed next to every discussion post in each thread, facilitating the quoting practice amongst the participants. passagen debatt does not have a quoting button, but instead uses a direct-reply function which makes it impossible to write “the next post” in a discussion thread without explicitly directing that post at a previous post. another way of making explicit that a post is relating to someone else’s post is by referring to that participant by name (alias) or by use of a pronoun, typically ‘du’ (you, second person singular) or ‘ni’ (you, second person plural). a third way of knowing how posts are sequentially related to each other is by looking at parts of adjacency pairs (schegloff & sacks 1973) which are expected to be produced next to each other, although there is no absolute requirement for the two parts to be strictly adjacent in all cases (hutchby & wooffitt 2008), since adjacency becomes a matter of interpretation in cmc (giles et al. 2017). when a discussion post constitutes the second part of an adjacency pair, that post is functionally dependent on a specific previous post, i.e., it is part of a two-part exchange and relates to the first part post, regardless of how far back in the chronological flow of discussion that the first part post is located. when producing the first part of an adjacency pair, such as a question, there is an underlying anticipation that the addressed participant(s) will respond in the expected way, by producing the second part of the adjacency pair. consequently, a first pair part projects a prospective relevance, establishing certain conditions for how the following turn will be perceived (schegloff 2007). most of the repair initiation posts in the second turn component of the 38 sequences contain the pronoun ‘du’ (you), explicitly addressing the person who produced the turn containing the trouble source, requesting clarification regarding the word. a little more than half of the repair initiation turns contain quotes or direct-replies of the post containing the trouble source, making explicit how these turns are related in a sequence. as all repair initiation turns contain meta-linguistic clarification requests pointing out the trouble source, they are all sequentially related both to the post where the trouble source is located, as well as to the response turn expected to clarify the meaning of the word. by providing an answer and repairing the issue of claimed non-understanding of the trouble source (in the initial post), the second turn component (the repair initiation post) and the third turn component (the repair post) are explicitly sequentially related to each other in a kind of adjacency structure. most of the repair posts and reaction posts also contain quotes or direct-replies of the previous posts in the wmn, explicitly displaying how the posts are related in a sequence. some repair and reaction posts also include the pronoun ‘you’ or an alias addressing the name of the person who the post is responding to. in summary, participants establish and maintain sequentiality between posts in discussion forum interaction by making use of affordances of the system, for example by quoting, directrepair of non-understanding of word meaning in online discussion forum interaction 99 replying or addressing particular interlocutors by name or personal pronoun, or by posts being sequentially linked in an adjacency structure. 3.3 ethical considerations technically, it is easy to collect interaction data from online platforms. however, the collection and processing of such data must be preceded by a series of ethical considerations and reflections on the part of the researcher. at present, there is an international guideline for online research published by the association of internet researchers, but the recommendations regarding an ethical approach in relation to the study of online interaction are relatively general. therefore, the individual researcher must make his or her own ethical considerations in relation to conducting research on online data (franzke, bechmann, zimmer, & ess 2020). the ethics of conducting research on online discussion forums without participant consent depends on several factors, including the type of research, the potential harm to participants, and the safety measures in place to protect participants’ privacy. the ethical considerations can vary based on the purpose of the study, the method of data collection (such as active engagement with participants vs the use of existing archival data), and the type of online space accessed (roberts 2015). when it comes to ethical considerations in relation to the collection and processing of online interaction data, there are mainly two aspects that the researcher needs to consider, partly whether the interaction takes place in an open or private space, and partly how sensitive the content of the interaction data is (markham & buchanan 2012). in general, it can be said that interaction that the participants themselves perceive as private, for example personal interaction between a low number of participants in a closed group, should not be collected or used without the participants’ consent. in more open spaces, such as public discussion forums where a larger number of participants interact under pseudonyms instead of their real names, participants are likely not perceiving the interaction to be private or closed. in such cases, the researcher can show ethical consideration towards the participants in other ways, for example by avoiding collecting material that deals with sensitive topics and by anonymizing the material as far as possible (smedley & coulson 2021). in this study, discussion threads dealing with sensitive issues have been excluded and all names of the participants have been anonymized. the membership agreements of the three discussion forums used in this study allow for data collection for non-commercial purposes and inform users that the forums may be used for private or educational use. familjeliv specifically reminds users that their written contributions are made in a public space and may be used in ways beyond their control. 4 analytical findings this section will present the findings of the study, by analyzing how participants initiate and carry out repair in instances of insufficient understanding of word meaning in online discussion forum interaction. 4.1 repair of claimed non-understanding of word meaning this section analyzes how participants initiate and carry out repair in instances of insufficient understanding of word meaning in online discussion forum interaction. in 34 of the 38 sequences, repair is accomplished in the third turn component, following a recognizable interactional pattern. four sequences deviate from this pattern and will therefore be treated separately as deviant cases. while the majority of sequences follow a recurring structure, the presence of deviant cases shows that this structure is not imposed by the analysis but emerges from the participants’ own orientations. these deviant cases will be analyzed in section 4.2. myrendal 100 4.1.1 initiating repair indicating claimed non-understanding of word meaning in all of the sequences, the repair initiation turn overtly addresses the issue of apparent nonunderstanding of word meaning by explicitly referring to the trouble source in a meta-linguistic clarification request, typically a variation on the form “what do you mean by x?” as seen in the repair initiation turns in excerpt 1 and 2. occasionally, the repair initiation turn can also contribute to the repair by displaying partial understanding of the meaning of the trouble source, for example by suggesting alternative interpretations of the word. this practice is in line with what heritage (1984) calls proposing “candidate understandings”, which is a common way of initiating repair in spoken communication. excerpt 3 p1 (1): vad är dina tankar om du skulle träffa en ogift 39 kvinna som inte har barn? p1 (1): what are your thoughts on if you would meet an unmarried 39 woman who doesn’t have children? p2 (6): vad menar du med ”ogift”. många är ju sambos, särbos, lever i partnerskap etc. menar du verkligen bokstavligen gift i din fråga eller menar du snarare att kvinnan ifråga är utan partner? finns ju mängder av folk som lever som sambos och har barn. p2 (6): what do you mean by ”unmarried”. lots of people are in relationships and live together, or live apart, or in partnerships etc. do you really mean married in your question or do you mean that the woman in question does not have a partner? there are lots of people who live together and have children. p1 (16): hej, jag menar enbart ogift. p1 (16): hi, i mean strictly unmarried. in the repair initiation post in excerpt 3, p2 shows that the word ‘ogift’ (unmarried) is partially understood by proposing two candidate understandings (post #6). the clarification request thus targets whether the intended meaning of the word in the current context should be interpreted as literally unmarried or rather as “without a partner”. all repair initiation turns contain a meta-linguistic clarification request addressing the need to repair the issue of claimed non-understanding of word meaning. in addition, some also contribute to the repair by displaying partial understanding of the word by proposing candidate understandings of the word. the following excerpts offer additional instances of repair initiation posts, illustrating a variety of approaches beyond the singular example of an ‘or’/disjunctive question presented in excerpt 3. 4.1.2 responding to and repairing instances of non-understanding of word meaning this section will investigate how issues of displayed non-understanding of word meaning are addressed and repaired in the third turn component of the sequence. in sequences where the repair initiation turn includes candidate understandings of the problematic word, p1 typically responds by confirming which of the interpretations is the intended one, as seen in the response posts in excerpt 3 (post #16). in cases where the repair initiation turn does not provide any partial understanding of the word, the p1 typically clarifies the meaning of the word by adopting one of the following repair practices: explicification or exemplification. one way of responding to a clarification request targeting the the meaning of the trouble source is by producing an explicification, a concept described by ludlow (2014) that denotes the introduction of an explicit definition-like component to a word meaning under negotiation. through repair of non-understanding of word meaning in online discussion forum interaction 101 acts of explicification, participants may clarify the meaning of the word by introducing a definitionlike component that foregrounds aspects of the word’s semantic properties, such as in excerpt 4 (post #4). excerpt 4 p1 (1): jag är antisexist, vilket betyder att jag är emot sexism i samhället! fråga mig vad ni vill! p1 (1): i’m anti-sexist, which means that i’m against sexism in society. ask me anything! p2 (3): vad menar du med begreppet “sexism”? p2 (3): what do you mean by the concept “sexism”? p1 (4): att människor behandlas olika pga sin könstillhörighet. p1 (4): that people are treated differently because of their gender. in the response post in excerpt 4, p1 repairs the issue of claimed non-understanding by producing an explicification. p1 here provides a definition-like clarification about the meaning of the word, drawing upon semantic affordances tied to the word itself. another way of responding to the meta-linguistic clarification request is by providing an example of what could be meant by the word, in the current situation or in situations beyond the current conversational context. this practice is called exemplification (myrendal 2019), and entails exemplifying or enumerating what the word can denote in the current conversational context, but also in other possible situations. excerpt 5 p1 (23): hursomhelst finns det en ordningslista för hur man designar en byggnad och jag tror inte ts greppar det riktigt. p1 (23): anyway there is an order list for how you design a building and i don’t think ts completely gets this. p2 (35): vad menar du med ordningslista? p2 (35): what do you mean by order list? p1 (39): tänk prioritetslista. ”mitt hus ska ha en pool” vs. ”mitt hus kan ha en pool om förhållandena tillåter”. det finns vissa saker som man löser först och vissa saker man löser senare. p1 (39): think priority list. ”my house should have a pool” vs. “my house could have a pool if the conditions allow.” there are some things you solve first and other things you solve later. in the response turn in excerpt 5, p1 repairs the issue of claimed non-understanding first by producing an explicification (“think priority list”) and then by producing an exemplification, enumerating what could be meant by the word “ordningslista” (order list). note the difference between introducing a definition-like component in an explicification and exemplifying what the word can mean. in exemplification, there is no definition of the word that draws upon aspects of the word’s own semantic properties, instead an account of what the word may denote in different situations is provided. the continuation of p1’s utterance can be seen as a part of the initial explicification, as this is interpreted more as a definition-like clarification about the meaning of the word itself, by drawing upon the word’s semantic properties such as the implication of prioritization. 4.1.3 resolving instances of claimed non-understanding of word meaning as mentioned, the fourth turn component (the reaction to the response post) is typically used to ground understanding of meaning, i.e., to acknowledge that there is enough mutual understanding concerning the the meaning of the trouble source so that the repair sequence can be resolved and myrendal 102 the discussion on topic can be resumed. in the data, the reaction post is only rarely used to continue the negotiation of meaning or by issuing another meta-linguistic clarification request, urging the other participant to continue negotiating the word. however, there are a few instances where the reaction turn is used to continue the sequence of repair, and these instances are found in the sequences constituting the deviant cases which will be described in detail in the next section. 4.2 the deviant cases from the sequential pattern found in most of the sequences, it is clear that p2’s repair initiation post establishes an expectation that p1 will respond with meta-linguistic clarification. in most cases, p1 fulfills this expectation by elaborating on the meaning of the word and thereby resolving the perceived non-understanding. however, in four sequences, no clarification is provided in p1’s third-turn response. sidnell (2013: 80) writes: “of course, a recipient sometimes responds in a way that might not have been predicted by the prior turn, or, indeed, does not respond at all. what can we make of such examples? as it turns out, such deviant cases often provide the strongest evidence for the analysis because it is here that we see the participants’ own orientations to the normative structures most clearly.” the four sequences that do not contain repair actions in the response post are treated as deviant cases in this study. they will be described in detail below, focusing on how participants orient to the normative structures by responding to their interlocutors’ unexpected behavior in different and observable ways. although this study does not aim to quantify occurrence rates, it is relevant to note that these 38 cases were identified from a larger pool of roughly 1500 forum posts. thus, while word meaning negotiation is not frequent, it is a recognizable and recurrent interactional phenomenon in asynchronous cmc. 4.2.1 unwillingness to provide meta-linguistic clarification in two sequences, p1 delivers a response post in the third turn component, but without providing the requested meta-linguistic clarification regarding the meaning of the trouble source. in these response posts, the expected clarification is missing, which means that there is a reply to the repair initiation post, but repair is not carried out. in excerpt 6, p1 is also called ts, thread starter. excerpt 6 p1 (20): ja påhittade räkningar vägrar jag ju såklart betala är det konstigt på något sätt? p1 (20): yes of course i refuse to pay made up invoices is that strange in some way? p2 (21): vad menar du med påhittade räkningar? har du en skuld hos fogden? p2 (21): what do you mean by made up invoices? do you have a debt with the bailiffs? p1 (22): antar att jag menar det jag säger såklart (-: p1 (22): guess i mean what i say of course (-: p1 (24): hade men byt ut faran mot räkningarna så ska du se resultatet.2 p1 (24): used to but switch the danger with the invoices and you will see the result. p2 (25): ursäkta men jag förstår inte vad du skriver? byt ut faran?? p2 (25): sorry, but i don’t understand what you are writing? switch the danger?? p3 (26): hopplös tråd gillar inte när ts inte förklarar ordentligt utan man ska gissa sig till vad hon menar. 2 this utterance does not make sense in swedish and has therefore been translated into english word for word. repair of non-understanding of word meaning in online discussion forum interaction 103 p3 (26): hopeless thread don’t like it when ts doesn’t explain properly and you have to guess what she means. p1 (27): sorry fakta ska det stå. p1 (27): sorry it should say facts. p4 (30): haha! vilken rolig tråd... en ts som inte kan skriva och man får dessutom inga rediga svar.. bara flum.. p4 (30): haha! what a funny thread… a ts who cannot write and you don’t get any clear answers.. just fluff… p5 (31): denna ts har för vana att då och då skapa trådar som ingen begriper och som hon sen inte ger någon mer förklaring till och så vanligen avslutas de med att ts vägrar svara på något mer av en eller annan anledning. p5 (31): this ts has a habit of occasionally creating threads that noone understands and that she doesn’t explain so usually they end with ts refusing to respond for one reason or another. the sequence dealing with the meaning of the word ‘påhittade’ (made up/fictitious) in excerpt 6 starts out like the other sequences. however, the expected pattern is disrupted when p1 does not provide the requested meta-linguistic clarification in the response post. instead, p1 responds evasively: “i guess i mean what i say” (post #22), which acknowledges the clarification request but avoids explaining the meaning of the problematic word. this failure becomes a communication problem, which is openly addressed by several other participants. p1 makes another response post, which contains both grammatical and spelling errors, and appears to further confuse the other participants (post #24). a side-sequence clearing up a spelling error is launched, and some participants comment on p1’s reluctance to clarify or engage properly with the thread (posts #25, #30, and #31). this deviant case is analytically important because it shows that the expected pattern – problematic word, repair initiation, clarification – is not a researcher-imposed structure but something participants themselves orient to. the reactions from several of the other participants in excerpt 6 show that there is a normative structure in place, with the expectation that p1 should provide the requested clarification in the response turn following the repair initiation (post #21). when the expected clarification is not delivered, this is overtly addressed by the other participants, revealing their orientations to the normative structures of what constitutes the expected behavior. while participants clearly treat the absence of clarification as a communicative issue, it should also be noted that the responses from p2, p3, and p4 appear to exhibit signs of irritation, sarcasm, or even strategic exploitation of the repair expectation. the clarification requests and critical comments could thus be interpreted not only as sincere efforts to re-establish understanding but also as actions that portray p1 as uncooperative or unreasonable. this possibility highlights that some cases of claimed non-understanding may blur into more conflictual or mocking uses of the repair sequence, especially when the trouble source concerns sensitive or contested matters. although the analysis primarily focuses on the organization of repair following displayed nonunderstanding, these more adversarial dimensions cannot be fully ruled out. this strengthens the claim that word meaning negotiation follows a recognizable sequential pattern and that departures from it are socially marked and made relevant by participants. as such, cases like excerpt 6 offer insight into the interactional variability and social accountability at play in wmns. they show that while a recurring structure is empirically observable, it is not presupposed in the analysis. this conclusion can be drawn because the structure is not only followed in most cases but also becomes visible through its disruption. in excerpt 6, multiple participants explicitly react to the missing clarification, expressing frustration and metacommenting on the lack of cooperation. this shows that the sequential pattern is treated by participants themselves as normative and accountable. it is through such participant orientations, both to the pattern and to its violation, that the analyst can identify the structure as an emergent myrendal 104 property of the interaction rather than an imposed coding scheme. in this way, the case illustrates how patterns in conversation analysis are not assumptions, but findings grounded in participants’ own displayed understandings of what constitutes appropriate interaction. next, one participant re-uses the word in a form of comprehension check directed at p1, in another attempt to elicit a repair containing a meta-linguistic clarification (post #32). excerpt 7 p6 (32): okej, så något företag har skickat ut påhittade räkningar till dig, som du inte har betalat men inte heller bestridit, och nu har ärendet gått till kronofogden. är det så det har gått till? p6 (32): okay, so some company has sent you made up invoices, which you have not paid, but which you have not contested either, and now the matter has gone to the bailiffs. is this how it went? p1 (33): eller skriver överenskomna summor det roliga är att jag vet inte om några överenskommelser eller vad det avser och har inte skrivit sådana i heller. p1 (33): or write agreed-upon sums the funny thing is i don’t know about any agreements or what they refer to and haven’t signed such agreements either. the participants have still not made any progress regarding the meaning of the word, as p1 has not provided a response post clearing up the issue of claimed non-understanding. however, a few turns later, p1 quotes an earlier post commenting on her unhelpful behavior and contests that she has not been clear about what is meant (post #35). excerpt 8 p1 (35): jag sa ju det det är räkningar som sägs vara överenskomna och jag har ingen aning har dock inte gjort några överenskommelser och vet heller inte vad det avser. p1 (35): i told you they are invoices that are said to be agreed upon and i have no clue have not made any agreements and don’t know what they refer to. p5 (38): varför har du då inte valt att bestrida räkningarna? man kan inte bara låta räkningar ligga, inte ens om de är felaktiga, för då hamnar det ju hos kronofogden till slut. p5 (38): why haven’t you contested the invoices then? you can’t just let invoices sit, not even if they are incorrect, because they will end up with the enforcement authority eventually. in excerpt 8, p1 finally touches upon aspects of meaning potential of the word ‘påhittade’ (made up/fictitious) (post #35). it has to do with agreements that are said to be arranged but that have not been agreed upon by both parties, i.e., the arrangements do not correspond to reality, according to p1. this is in fact fairly close to the dictionary meaning of ‘made up’ or ‘fictitious’, and the issue of claimed non-understanding of the word likely has to do with p1’s creative use of the word in the current conversational context. the word ‘påhittade’ (made up/fictitious) is typically not used to describe invoices. in the reaction to p1 in excerpt 8, p5 displays enough understanding of the scenario to produce a follow-up question regarding the topic (post #38). p5 substitutes the word ‘påhittade’ (made up/fictitious) with ‘felaktiga’ (incorrect). beyond this point in the discussion, the word ‘påhittade’ is not used again. instead, the word ‘felaktiga’ is used in the remainder of the discussion to refer to the erroneous invoices. another sequence in which a participant does not provide the expected meta-linguistic clarification in a response to a post initiating repair is found in the negotiation of the word repair of non-understanding of word meaning in online discussion forum interaction 105 ‘självständig’ (independent). in this sequence, there is no meta-linguistic shift in the third turn as p1 in his response does not deal with the the meaning of the trouble source (post #3). again, this is overtly addressed in the reaction to the response post in the fourth turn (post #4). unlike in the case of ‘påhittade’, the meta-linguistic shift comes in the reaction turn, by a third person entering into the sequence clarifying the meaning of the word in p1’s place (post #4). excerpt 9 p1 (1): min önskekvinna. hon ska vara självständig men ändå tillgiven. hon ska vara utmanande men bara för mig. hon ska ha värme men visa kyla och avståndstagande mot utomstående. hon ska vara en hemmafru som pysslar om hemmet och gör det hemtrevligt som bara en kvinna kan. p1 (1): my ideal woman. she should be independent yet devoted. she should be challenging but only to me. she should possess warmth but show coldness and distance towards outsiders. she should be a housewife who takes care of the home and makes it cozy in a way only a woman can. p2 (2): vad menar du med självständig men ändå tillgiven? vad lägger du in för kriterier vad gäller självständighet hos en kvinna kontra tillgivenhet? det är jag nyfiken på. p2 (2): what do you mean by independent yet devoted? what criteria are you considering regarding a woman's independence versus her devotedness? i'm curious about that. p1 (3): att min önskekvinna ska ha dessa egenskaper. jag vill ha tillbaks den trygghet man hade som barn av att ha en kvinna helt för sig själv och som man vet känner något alldeles extra för en. p1 (3): that my ideal woman would possess these qualities. i want to reclaim the security one had as a child, having a woman all to oneself who you know feels something special for you. p3 (4): en självständig kvinna = en kvinna som tjänar egna pengar och kan ta hand om sig själv vilket krockar med din önskan om en hemmafru för en hemmafru får du ta hand om mer än vad hon tar hand om dig. p3 (4): an independent woman = a woman who earns her own money and can take care of herself, which contradicts your desire for a housewife because a housewife requires you to take care of her more than she takes care of you. in excerpt 9, many of the criteria on p1’s list are listed in pairs, with the first element in the pair in some way contradicting the second element since they are contrasted with each other by the word ‘but’ (post #1). this applies to the first criterion, about being independent yet devoted. p2 requests clarification about the meaning of the word ‘självständig’ (independent), which he seems to regard as difficult to combine with the word ‘tillgiven’ (devoted). here, it seems that there are aspects of meaning potentials of the two words that p2 cannot merge into one personality trait in a woman. p2 hints that the two words pull into different directions, that they are opposites (post #2). p2 asks p1 to provide meta-linguistic clarification about the the meaning of the trouble source, which is signaled in the repair initiation post asking about which criteria must be met for a woman to be considered independent (post #2). however, p1 does not shift focus to the meaning of the word in the response post, and thus does not enter into a repair sequence concerning the meaning of the word (post #3). even though p1 attends to the repair initiation post by replying to it, there is nothing said about the meaning of the word, which means that no repair action is carried out in the response post (post #3). p1 simply repeats that these are the qualities he is looking for in a woman. he does not further explain what he thinks these qualities entail, or what the words he is using to describe them mean. myrendal 106 when the meta-linguistic clarification request is not attended to in the response post, another participant enters into the discussion and puts forward an explanation of the word ‘självständig’ (independent). p3 touches upon aspects of meaning potential in the sense that being independent means making money and being able to support oneself. both sequences that break the mold for what is the typical pattern in cases of claimed nonunderstanding of word meaning presented in this section manifest the expectations of what is considered “normal behavior” when engaging in a repair sequence dealing with issues of claimed non-understanding. when the meta-linguistic clarification is not provided in the third turn, the participants orient to this lack of clarification by explicitly addressing it in the discussion. through the collaborative efforts of the participants, they end up finally dealing with the issue of claimed non-understanding of word meaning, just not in the third turn component (as expected). 4.2.2 being vague in response posts, playing on words the two sequences presented in the previous section were treated as deviant cases since there was no meta-linguistic clarification provided in the third turn in each sequence. next, two other deviant cases will be described and analyzed. in these cases, there are attempts to take repair action in the third turn, but they are not fully successful. again, the failure to provide meta-linguistic clarification is overtly addressed in the discussion that follows, highlighting how participants orient to the underlying normative structures when repair has been initiated but not carried out as expected. the first sequence is about the word ‘evighetsgaranti’ (eternity-guarantee) which is used in a religious discussion where one participant uses it as a promotional piece on why believing in god is a good thing. excerpt 10 p1 (1): tron på gud är som en utgångspunkt för allt jag är med om. inte som en vag gissning utan som det mest verkliga jag vet: livet. gud som livets ursprung och mål är stabila grejer, ja som jag kan se den enda hållbara utgångspunkten för livet. evighetsgaranti medföljer ju. p1 (1): faith in god is the starting point for everything i experience. not as a vague guess, but as the most real thing i know: life. god as the origin and goal of life is stable stuff, yes as i can see it the only sustainable starting point for life. eternity-guarantee included. p2 (4): orsak, mening och evigheten. det är den där vissheten om orsaken, meningen och evigheten som känns märklig för mig. vad då evighetsgaranti? garanti??? varför är det så säkert att det finns en orsak eller en mening? det är tankekonstruktioner om något. p2 (4): cause, meaning and eternity. it’s the certainty about the cause, meaning, and eternity that feels strange to me. what eternityguarantee? guarantee??? why is it so certain that there is a cause or a meaning? those are constructions of thought if anything. p1 (5): vissheten känns märklig för dig skriver du. fast jag skulle tycka det var ännu mer märkligt om en människa inte var förvissad om sin egen tro. evighetsgaranti var bara en liten lek med våra ord, ta det inte så bokstavligt. p1 (5): you write that certainty feels strange to you. but i would find it even more strange if a person was not sure of their own faith. the eternity-guarantee was only a little play on words, don’t take it so literally. p2 (6): hur ska jag ta ordet evighetsgaranti om inte bokstavligt? du säger ju att du är förvissad. är du förvissad (känner att du repair of non-understanding of word meaning in online discussion forum interaction 107 har garanti) för att du ska få evigt liv, eller är du inte förvissad om det? eller är du bara förvissad om att du tror att du kommer att få evigt liv? det är faktiskt en viktig distinktion. p2 (6): how should i take the term eternity-guarantee if not literally? you say you are sure. are you sure (feel that you have a guarantee) that you will have eternal life, or are you not sure about it? or are you just sure that you believe you will have eternal life? it’s actually an important distinction. p1 (8): tack för dina synpunkter. orden är allas vårt redskap att skapa begriplighet med allvar och lek och den här gången blev mitt försök med evighetsgaranti ingen hit. beklagar det men jag har ingen möjlighet att fortsätta diskussionen ikväll. p1 (8): thank you for your point of view. words are everyone's tool to create understanding with seriousness and play, and this time my attempt with an eternity-guarantee wasn’t successful. i regret that but i am unable to continue the discussion tonight. in excerpt 10, p2 requests clarification about the word ‘evighetsgaranti’ (eternity-guarantee), indicating the need to specifically address the meaning of this word (post #4). p1 responds to the post but does not take repair action by providing meta-linguistic clarification concerning the meaning of the word. this excerpt exemplifies a complex case that navigates the boundaries between a genuine request for clarification, and potential disagreement about the appropriateness of a chosen word and its meaning. p2’s request indicates a need to specifically address the meaning of ‘evighetsgaranti’ (post #4), prompting p1 to respond. however, p1 does not provide a metalinguistic clarification but states that the situated meaning of ‘evighetsgaranti’ should not be taken literally (post #5). this response acknowledges the word’s literal interpretation but dismisses its relevance in the current context, suggesting instead that it was merely a “little play on words”. p2 seems dissatisfied with the lack of clarification and repeats the request, emphasizing a literal interpretation as the only conceivable meaning (post #6). this insistence on a literal interpretation can be seen as p2 potentially contesting the adequacy of p1’s initial explanation rather than simply seeking clarification. here, the example highlights a nuanced negotiation about word meaning, underscored by p2’s repeated questions which may suggest a deeper, possibly critical engagement with the underlying assumptions of p1’s faith-based assertions rather than lack of clarity. this dynamic is reinforced when p1 ultimately retracts the use of ‘evighetsgaranti’ altogether (post #8), signaling a concession that the term might have introduced more confusion or controversy than clarity. the discussion straddles the line between insufficient understanding and disagreement, reflecting the dialogical nature of wmn where words serve not just as means of communication but as tools for negotiating deeper ideological and existential dimensions (myrendal 2019). the repeated request for clarification in post #6, the use of multiple question marks in post #4, and the challenges to the provided explanations suggest that the sequence may not solely be about seeking understanding, but also about contesting the word, its meaning and implications. however, the example shows that participants orient to otherwise invisible structures when their interlocutors behave in ways that break the normative pattern of communication. when a response post that may be anticipated to provide clarification regarding a trouble source instead contains vagueness or word play, the issue of claimed non-understanding remains to be cleared up as the participant requesting clarification is essentially left without a proper response and therefore adopts alternative strategies for coping with the persisting issues of claimed non-understanding. in this case, we see that both strategies from the two deviant cases described earlier are employed, namely reiterating the clarification request and supplying an answer oneself regarding what the word can (must) mean in the given context. myrendal 108 yet another example of a deviant case containing word play in a response post is found in the repair sequence of the word ‘mannar’ (menfolk). the discussion is about a person, p1, who wants to start up a tattoo removal business abroad and who has a few questions regarding the funding of that business. excerpt 11 p1 (1): får jag bidrag av svenska mannar om jag öppnar upp företag utomlands, eller ska ja registrera företaget i sverige o flyga runt o ta bort tatueringar… p1 (1): will i get financial support from swedish menfolk if i open up the company abroad, or should i register the company in sweden and fly around and remove tattoos… p2 (2): vad menar du med mannar? p2 (2): what do you mean by menfolk? p1 (3): givetvis menar jag män i skor.3 p1 (3): of course i mean men in shoes. [bold font in original] p2 (5): menar du alla män i landet som bär skor, eller vadå? går det inte lika bra att du får bidrag ifrån kvinnor med höga klackar? p2 (5): do you mean all men in the country who wear shoes, or what? wouldn’t it be just as good if you got funding from women in high heels? in excerpt 11, p2 reacts to the use of the word ‘mannar’ (menfolk) and requests clarification (post #2). p1’s reply in post #3 does not resolve the confusion but instead introduces a pun, playfully complicating the conversation. the swedish phrase ‘män i skor’ translates directly to ‘men in shoes’, but phonetically it resembles ‘människor’, meaning ‘people’. this linguistic twist merges the literal translation with a humorous nod to the more inclusive term ‘people’. for the second time, p1 highlights a distinction between the gender-neutral word ‘människor’ (people) which could have been used and the two word choices specifically focusing on men (menfolk and men in shoes). what p1 appears to be asking, is if he will receive financial support from the swedish government, i.e., the swedish people, so explicitly asking about swedish ‘mannar’ (menfolk) or ‘män i skor’ (men in shoes) seems to draw attention to a difference between the word choices foregrounding male citizens and the neutral ‘människor’ (people). this is also picked up in the reaction post, as p2 explicitly asks if financial support coming from women would not be just as good (post #5). at this point, p1 has been vague in the response post, by employing a play on words, and it is still not clear if ‘mannar’ (menfolk) should be interpreted as ‘people’ or ‘male people’. in the reaction post, p2 makes the other, more specific clarification request, drawing attention to the distinction between male and female taxpayers. the reaction post indicates the need to keep repairing the issue of claimed non-understanding, until sufficient understanding of the trouble source has been established. however, there is no uptake from p1 with regards to the second clarification request. in fact, the reaction post is the very last post in the entire thread. in this case, we see yet another example of how a participant addresses behavior that does not follow the expected pattern of interaction. when the requested meta-linguistic clarification is not provided and the issue of claimed non-understanding is not attended to in the expected place, this issue is overtly addressed, and the clarification request is repeated. in this case, it is made more specific and simultaneously adds to the negotiation of the word, asking if it should be understood as ‘male people’ or just ‘people’. it could be argued that this sequence is also a borderline case between a wmn originating in claimed non-understanding and a disagreement, following myrendal’s categorization (2015), seeing that the meta-linguistic clarification request may be 3 the swedish phrase ‘män i skor’ is a word play on ‘människor’ which, when written as one word, means ‘people’. however, when split into three separate words, it literally translates to ‘men in shoes’. repair of non-understanding of word meaning in online discussion forum interaction 109 interpreted more like a meta-linguistic objection (commonly found in wmns originating in disagreement), seeing that p1 uses “givetvis” (of course) before employing the word play “män i skor” (men in shoes) and p2’s somewhat sarcastic response about “women in high heels”. 5 concluding discussion this study shows how participants in online discussions manage to organize their interaction, in instances where there is insufficient understanding regarding the meaning of particular words, by participating in wmn repair sequences. by adopting practices of establishing explicit relations between posts that may chronologically be very far apart in the discussion flow, participants are able to form distinct sequences of repair to deal with the issues of claimed non-understanding. establishing this kind of explicit sequentiality is crucial in overcoming the problems of disrupted turn adjacency and interactional incoherence often said to be characteristic of this form of communication. most repair sequences in this study display a similar threeor four-turn pattern. in the first turn component, a word is used that is later recognized as difficult to understand. in the second turn component, the word is addressed as a trouble source and repair is initiated. in the third turn component, the issue of claimed non-understanding is typically addressed and repaired, and in cases where there are four turn components in the repair sequence, this typically contains a reaction confirming understanding of the word, thus exiting the sequence of repair, and resuming the main discussion on topic. the analysis reveals that most of the issues of claimed non-understanding are solved quickly and easily between participants. when faced with a meta-linguistic clarification request, most participants who produced the trouble source provide a response that quickly clears up the issue of claimed non-understanding, either by defining the meaning of the word, or by exemplifying what the word can denote in the world. consequently, participants seem willing to cooperate concerning the meanings of the words used in communication, both when producing repair initiators and when agreeing to elaborate on and explain the meaning of a word indicated by someone else as problematic. it is interesting to note that when participants do not provide an explanation of the meaning of the trouble source, subtle yet revealing insights into participants’ orientations to the normative structures underpinning repair sequences emerge. the findings from the analysis of the deviant cases indicate that when the expected meta-linguistic clarification is not provided in the third turn of the repair sequence, the participants overtly address this as a communicative issue that needs to be resolved. this demonstrates that participants seem to anticipate a certain response when a clarification request is directed towards someone; specifically, the expectation is that the recipient of such a request will offer an answer to clarify the issue of claimed non-understanding. this anticipation, and the way participants address the absence of clarification, aligns with the broader concept of preference organization (goodwin & heritage 1990). while goodwin and heritage primarily discuss the notion of preference in the context of agreements (being produced promptly) and disagreements (being marked by delay and mitigation), the expectation for clarifying responses can be seen as part of the conversational machinery that participants navigate. this machinery, illustrated by preference organization, is underpinned by rationales that favor certain conversational outcomes over others. in the context of repair sequences, the anticipation of a clarifying response reflects a preference for resolving ambiguities and misunderstandings, embodying the normative expectation that communication should progress towards mutual understanding. the findings show that when the expected meta-linguistic clarification is not provided in subsequent turns, participants employ several strategies to address the lack of clarification as an interactional problem. these strategies include producing repeated repair initiations, prompting the other participant to collaborate on restoring enough mutual understanding, and speculating about what the trouble source word must mean in context. participants’ perseverance in wanting to stay myrendal 110 in the repair sequence until sufficient mutual understanding has been reached demonstrates the importance of such repair work for maintaining coherence in communication. thus, the findings reveal an order of preference governing online repair of claimed non-understanding of word meaning: once participants have entered a repair sequence targeting the meaning of a word, they prioritize staying within it until the issue is resolved. in the deviant cases, if the participant expected to clarify the meaning of the trouble source fails to do so, other participants seem to prefer to advance the repair themselves rather than allow the interaction to stall. this highlights an orientation towards maintaining the interaction’s progressivity, as participants continually take necessary actions to resolve claimed nonunderstandings and resume the discussion on the main topic. these findings resonate with earlier ca research on spoken interaction, underscoring how participants prioritize the progression of talk. previous studies have explored how participants focus on progressivity, for instance, by displaying that a repair move has made progress towards the solution of the trouble addressed (schegloff 1979) and by showing a preference for answers over non-answers in responses to questions seeking specific information (stivers & robinson 2006). moreover, the deviant cases sometimes reveal a complex interplay between claimed non-understanding and potential disagreement, where the nature of the interaction may initially appear to be about clarifying issues of insufficient understanding but can subtly transition into a borderline disagreement concerning the underlying meaning or use of the word (myrendal 2019). this study’s findings resonate with goddard and gillespie’s (2025) large-scale analysis of reddit interactions, which showed that while repair initiations are widespread online, many go unanswered, particularly in short threads. however, the present study demonstrates that when a thread continues and the participants remain engaged, interlocutors treat clarification requests as socially accountable actions that call for a response. unlike cases where the thread simply dies out, a non-answer in an ongoing thread is seen as a marked absence, prompting meta-comments and criticism. this comparison highlights that even in environments where uncompleted repairs are statistically common, active non-responses are interactionally dispreferred – reinforcing the normative expectation to engage in clarification when misunderstandings are made explicit. the significance of investigating online interaction and its linguistic practices lies not only in its widespread prevalence in our daily lives, but also in the fact that digital tools and platforms play a critical role in shaping communication of today. these tools enable interaction to take place asynchronously, opening new ways for discussion, negotiation, and sense-making. this study sheds light on repair practices in asynchronous online interaction, showing that they play a crucial role in the overall organization of the interaction and facilitate the progression of the discussion on topic. knowledge of these practices is important not only for advancing linguistic theory based on empirical studies of online interaction, but also for the future development of online platforms. such knowledge can inform the design of tools and strategies that more effectively support successful communication and mitigate misunderstandings in online settings. acknowledgements this work was supported in part by swedish research council grant 2022-02125, not just semantics: word meaning negotiation in social media and spoken interaction. references saul albert and j.p. de ruiter (2018). repair: the interface between interaction and cognition. topics in cognitive science 10: 279-313. repair of non-understanding of word meaning in online discussion forum interaction 111 charles antaki, elisenda ardévol, francesc núnes and agnès vayreda (2005). ““for she knows who she is:” managing accountability in online forum messages.” journal of computermediated communication 11(1): 114-132 ilkka arminen, christian licoppe and anna spagnolli (2016). “respectifying mediated interaction. research on language and social interaction 49(4): 290-309. herbert h. clark (1996). using language. cambridge: cambridge university press. herbert h. clark and susan brennan (1991). grounding in communication. in l. b. resnick, j. m. levine & s. d. teasely (eds.), perspectives on socially shared cognition, 127-147. washington: american psychologist association. marcus colman and patrick g.t. healey (2011). the distribution of repair in dialogue. in: proceedings of cogsci 2011 33rd annual meeting of the cognitive science society, 15631568. mark dingemanse, seán g. roberts, julija baranova, joe blythe, paul drew, simeon floyd, rosa s. gisladottir, kobin h. kendrik, stephen c. levinson, elisabeth manrique,giovanni rossi, and n.j. enfield (2015). universal principles in the repair of communication problems. plos one, 10 https://doi.org/10.1371/journal.pone.0136100 paul drew (1997). open' class repair initiators in response to sequential sources of troubles in conversation. journal of pragmatics 28: 69-101. matteo farina (2018). facebook and conversation analysis. the structure and organization of comment threads. london: bloomsbury academic. richard fitzgerald, william housley, and sean rintel (2017). membership categorisation analysis. technologies of social action. journal of pragmatics 115: 51-55. aline s franzkeanja bechmann, michael zimmer, and charles ess (2020). internet research: ethical guidlines 3.0. https://aoir.org/reports/ethics3.pdf angela c. garcia and jennifer b. jacobs (1999). the eyes of the beholder: understanding the turn-taking system in quasi-synchronous computer-mediated communication. research on language and social interaction 32: 337-67. will gibson (2009). intercultural communication online: conversation analysis and the investigation of asynchronous written discourse. forum: qualitative social research 10: 118. david, giles, wyke stommel, and trena paulus (2017). the microanalysis of online data: the next stage. journal of pragmatics 115: 37-41. alex goddard and alex gillespie (2025). conversational repairs on reddit: widely initiated but often uncompleted. plos one, 20(1): https://doi.org/10.1371/journal.pone.0316618 charles goodwin and john heritage (1990). conversation analysis. annual review of anthropology 19: 283-307. jess harris and amelia church (2019). methodological insights from ethnomethodology and conversation analysis. journal of pragmatics 143: 201-204. patrick g.t. healey, gregory j. mills, arash eshghi, and christine howes (2018). running repairs: coordinating meaning in dialogue. topics in cognitive science 10: 367-388. john heritage (1984). a change of state token and aspects of its sequential placement. in structure of social action, ed. by j. atkinson and j. heritage, 299-345. cambridge: cambridge university press. susan herring (2010). who’s got the floor in cmc? edelsky’s gender patterns revisited. language@internet 7. julian hough and matthew purver (2013). modelling expectations in the self-repair processing of annot-, um, listeners. in: proceedings of the 17th semdial workshop on the semantics and pragmatics of dialogue. amsterdam, netherlands, december 16-18. 92-101. ian hutchby and robin wooffitt (2008). conversation analysis. cambridge: polity press. myrendal 112 jennifer b. jacobs and angela c. garcia (2013). repair in chat room interaction. in pragmatics of computer-mediated communication, ed. by s. herring, d. stein and t. virtanen, 565-87. berlin: mouton de gruyter. gail jefferson (1987). on exposed and embedded correction in conversation. in talk and social organisation, ed. by john r.e lee and g. button, 86-100. clevedon: multilingual matters. kevin jepson (2005). conversations – and negotiated interaction – in text and voice chat rooms. language learning and technology 9: 79-98. lina lee (2008). focus-on-form through collaborative scaffolding in expert-to-novice online interaction. language learning & technology 12: 53-72. diana m. lewis (2005). arguing in english and french asynchronous online discussion. journal of pragmatics 37: 1801-18. per linell (1995). troubles with mutualities: towards a dialogical theory of misunderstanding and miscommunication. in mutualities in dialogue, ed by i. marková, c. graumann and k. foppa, 176-213. cambridge: cambridge university press. peter ludlow (2014). living words: meaning underdetermination and the dynamic lexicon. oxford: oxford university press. michel marcoccia (2004). on-line polylogues: conversation structure and participation framework in internet newsgroups. journal of pragmatics 36: 115-45. annette markham and elizabeth buchanan (2012). ethical decision-making and internet research: recommendations from the aoir ethics working committee (version 2.0). https://www.aoir.org/reports/ethics2.pdf joanne meredith (2017). analysing technological affordances of online interaction using conversation analysis. journal of pragmatics 115: 42-55. joanne meredith (2019). conversation analysis and online interaction. research on language and social interaction, 52:3. 241-256. jenny myrendal (2019). negotiating meanings online: disagreements about word meaning in discussion forum communication. discourse studies 21: 1-23. jenny myrendal (2015). word meaning negotiation in online discussion forum communication. ph.d dissertation. bohus: ale tryckteam. matthew purver (2004). the theory and use of clarification requests in dialogue. ph. d dissertation, university college london. matthew purver, julian hough, and christine howes (2018). computational models of miscommunication phenomena. topics in cognitive science 10:425-451. lynne d. roberts (2015). ethical issues in conducting qualitative research in online communities. qualitative research in psychology 12: 314-325. harvey sacks, emanuel a schegloff, and gail jefferson (1974). a simplest systematics for the organization of turn-taking for conversation. language 50: 696-735. emanuel a. schegloff (1979). the relevance of repair for syntax-for-conversation. in t. givon (ed.), syntax and semantics 12: discourse and syntax, 261–88. new york: academic press. emanuel a. schegloff (2007). sequence organization in interaction: a primer in conversation analysis. vol. 1, cambridge: cambridge university press. emanuel a. schegloff, gail jefferson, and harvey sacks (1977). the preference for selfcorrection in the organization of repair in conversation. language 53: 361-82. emanuel a. schegloff and harvey sacks (1973). opening up closings. semiotica 7: 289-327. juliane schönfeldt and andrea golato (2003). repair in chats: a conversation analytic approach. research on language and social interaction 36: 241-84. paul seedhouse (2004). conversation analysis methodology. language learning 54: 1-54. jack sidnell (2010). conversation analysis. in sociolinguistics and language education, ed. by n.h. hornberger and s.l. mckay, 492-527. bristol: multilingual matters. jack sidnell (2013). basic conversation analytic methods. in the handbook of conversation analysis, ed. by j. sidnell and t. stivers, 77-100. malden, ma: wiley-blackwell. repair of non-understanding of word meaning in online discussion forum interaction 113 james simpson (2005). conversational floors in synchronous text-based cmc discourse. discourse studies 7: 337-61. richard m. smedley and neil s. coulson (2021). a practical guide to analysing online support forums. qualitative research in psychology 18: 76-103. tanya stivers and jeffery d. robinson (2006). a preference for progressivity in interaction. language in society, 35: 367-392. wyke stommel (2008). conversation analysis and community of practice as approaches to studying online community. language@internet, 5: 1-22. paul ten have (2007). doing conversation analysis: a practical guide. london: sage publications. charlotte van hooijdonk and tessa van charldorp (2019). sparking conversations on facebook brand pages: investigating fans’ reactions to rhetorical brand posts. journal of pragmatics 151: 30-44. amber n. warren and trena m. paulus (2020). postgraduate students’ accomplishment of epistemic positioning through personal experience in online discussion forums. classroom discourse 11: 11-40. burstein_tetreault_chodorow_final_130720 dialogue and discourse 4(2) (2013) 34-52 doi: 10.5087/dad.2013.202 ©2013 jill burstein, joel tetreault, martin chodorow submitted 02/12; accepted 11/12; published online 04/13 holistic annotation of discourse coherence quality in noisy essay writing jill burstein jburstein@ets.org educational testing service rosedale road, ms 11-r princeton, new jersey 08540 joel tetreault1 joel.tetreault@nuance.com nuance communications, inc. 1198 east arques avenue sunnyvale, ca 94085 martin chodorow martin.chodorow@hunter.cuny.edu hunter college and the graduate center of cuny 695 park avenue new york, ny 10065 editors: stefanie dipper, heike zinsmeister, bonnie webber abstract in this paper, we describe a holistic annotation scheme for coherence quality that requires little expertise on the part of the human annotator; we also present computational systems for evaluating coherence quality in essays. a formidable challenge to reliability when annotating discourse coherence comes from differences among annotators in the inferences that they draw when reading an essay. this may reflect differences in their background knowledge or in their willingness to bridge what might otherwise seem to be disconnected portions of the text. despite these differences, we achieved adequate reliability for a holistic binary coherence annotation of essays written for five different types of large-scale assessment. when designing computational systems to score these essays for coherence quality, we faced a number of issues not encountered in most previous work, which has focused primarily on well-formed text. foremost among these was the need to develop models based on features that reflect the coherence quality criteria which are found in human essay scoring guides, including aspects of writing quality (e.g., the presence of grammatical errors) that might interfere with constructing the meaning of the essay. such features are needed to produce meaningful scores and to provide the basis for instructional feedback to the student or test-taker. we present results of testing various computational models on essays using the binary discourse coherence score. keywords: discourse coherence annotation, essay data, automated essay scoring 1 the work reported here was completed while the second author was employed at educational testing service. holistic annotation of coherence quality 35 1. introduction motivation. in the field of automated essay evaluation, we are tasked with building systems that evaluate the overall quality of essays (attali & burstein, 2006; burstein, tetreault & madnani, 2013; and, elliot & klobucar, 2013). these systems must be able to handle noisy essay data. in the context of automated essay evaluation for educational instruction and assessment, system features must be consistent with human scoring criteria to satisfy measurement standards. specifically, are the system features used to predict the quality of the essay consistent with the scoring rubric criteria (i.e., criteria that are desirable to measure essay quality)? further, system outcomes must be transparent (easily explainable). in other words, in the same way a teacher would explain why he or she had assigned a grade (score) to a writing assignment, we must be able to explain to score recipients (typically test-takers, teachers, and students) the system features that predicted outcomes. consistent with this, in developing educational technology, explanation of system decisions is often achieved through explicit feedback aligned with scoring rubric criteria, especially in instructional settings. for instance, e-rater®, an automated essay scoring system used in assessment and instruction (burstein, chodorow, & leacock, 2004; attali & burstein, 2006; burstein et al., 2013), uses a set of well-defined features that are aligned with human scoring criteria. the system assigns scores to essays in standardized assessment and instructional settings. in the united states, the practical need for text analysis capabilities for writing tasks is now being driven by increased requirements for standardized curricula and assessments. more recently, the need for applications for text analysis has been emphasized by the common core state standards initiative (ccssi).2 this relatively new initiative has now been adopted by most states for use in kindergarten through 12th grade (k-12) classrooms, and it is likely to have a strong influence on teaching standards in k-12 education. ccssi illustrates what k-12 students should be learning with regard to reading, writing, speaking, listening, language, and media and technology. more specifically, it describes language structures that learners need to grasp as they progress to the higher grades in preparation for college readiness in reading and writing.3 in terms of writing skill and discourse coherence, there is a strong focus on students’ ability to “marshal an argument.” given the continued emphasis on the need for feedback, and a specific interest in supporting students’ ability to develop argumentation in text, the development of automated capabilities that evaluate aspects of argumentation such as, discourse coherence quality, will be critical. consistent with the need to build technology that can evaluate aspects of argumentation, we have been developing a system for automated evaluation of discourse coherence quality (burstein, tetreault, & andreyev, 2010). a major goal is to ensure that the system addresses scoring criteria specifically related to discourse coherence in noisy essay writing. human readers 2 see http://www.corestandards.org/. 3 see http://www.corestandards.org/assets/publishers_criteria_for_3-12.pdf. burstein, tetreault, and chodorow 36 who score essays written for standardized writing exams are instructed to consider discourse coherence quality, assigning lower scores to essays that are less coherent. for example, some essays are likely to contain ungrammatical structures, misspellings and typographical errors, poor organizational structure (e.g., lack of transitions), and other linguistic features that contribute to the difficulty readers may have in constructing the meaning of the text. to assess text quality, it is necessary to isolate the linguistic sources of this difficulty if we are to provide meaningful essay scores and instructional feedback that addresses relevant writing features contributing to overall essay quality. constructing meaning. a core challenge in the evaluation of discourse coherence quality is that it is related to a reader’s construction of text meaning. from halliday and hasan’s (1976) perspective, a text is a semantic unit, as opposed to a large grammatical unit, or super-sentence. further, there are a number of contributors to textual continuity –the reader’s ability to construct meaning from the text. these include overt linguistic features, such as relationships between words, well-formed sentences, and meaningful transitions. as well, interpretation of a text plays a role in the reader’s perception of the text continuity, or quality of discourse coherence. also related to text interpretation, shriver (1989) argues that the quality of a text can be judged to some extent by the correct (author-intended) inferences that readers draw from the text. from a cognitive psychology perspective, graesser, mcnamara, louwerse, and cai’s (2004) view coherence is a psychological construct, and “…coherence relations are constructed in the mind of the reader and depend on the skills and knowledge that the reader brings to the situation. (p.1).” as is also pointed out in graesser et al. (2004), readers’ construction of meaning may differ based on individual inferences. van den broek (2012) asserts that skilled readers have a “large toolbox” of reading strategies that they can access to build coherence as they read a text. these include retrieving information from previous, earlier parts of the text, bringing in prior knowledge, and accessing other sources of information, for example, via internet search. van den broek observes that these processes are learned over time and become increasingly more automated with practice. in reading a text, users draw on these strategies in a “landscape of activations” in which they move between concepts from sentence to sentence. evaluation of discourse coherence quality. in scoring writing assessments, how do human raters tasked with evaluating discourse coherence in a text determine what renders higher and lower discourse coherence? by design, conventional holistic scoring guides developed for standardized writing assessments are intended to offer human raters guidelines in terms of how to assign an overall score that takes into account many aspects of writing (coward, 1950; huddleston, 1952; and, godshalk, swineford, & coffman, 1966). these include word choice, syntactic variety, grammaticality, and organizational quality. with regard to discourse coherence, while scoring guides provide a general idea of features that might play a role in evaluating discourse coherence quality, they do not prescribe to the rater the linguistic features that contribute specifically to good and bad qualities of coherence. in addition, no specific examples of essays illustrating variation in quality specifically related to coherence are provided. rubric criteria in scoring guides that describe the relative quality of discourse coherence might read something like this: “displays unity, progression, and coherence, though it may contain occasional redundancy, digression, or unclear connections”; “is not clearly organized; some holistic annotation of coherence quality 37 parts may be clear while others are disjointed or confused... errors in grammar interfere with reader understanding”. such criteria are not clear-cut, and leave the necessary wiggle room for individual differences inherent in judgments about discourse coherence as suggested by halliday and hasan (1976), shriver (1989), and graesser et al. (2004). we use a holistic approach in our annotation scheme that focuses on reading to evaluate versus reading to comprehend (shriver, 1989). the research described in this paper shows promise that we can capture annotators’ “impressions” of discourse coherence quality and provide a meaningful holistic evaluation score for discourse coherence in noisy essay data. by meaningful, what we mean is that we extract linguistically-grounded features from the annotated data set that can be mapped to scoring criteria related to discourse coherence and can also be translated into feedback for students and test-takers in instructional settings (i.e., classroom instruction or test preparation). features related to discourse coherence are used to build and evaluate system models of discourse coherence quality. this paper focuses on a discussion of issues for annotation of discourse coherence quality and related annotation challenges, especially those related to the inferences that readers draw. we also consider the need for consistency between the feature set used for building coherence models and the scoring rubric criteria, in order for coherence scores to be meaningful and valid.4 further, to demonstrate the effectiveness of a holistic annotation scheme, we present our discourse coherence system development and evaluation. the remainder of this paper is organized as follows. in section 2, we discuss computational approaches to the evaluation of discourse coherence and related work on annotation methods used to build and evaluate systems for predicting text coherence and readability. in section 3 we describe a holistic annotation scheme, related challenges, and the development and evaluation of discourse coherence models for five data sets. section 4 provides further discussion and conclusions. 2. related work early approaches to discourse analysis of texts have, in general, been limited to lexical cohesion, specifically, repetition of vocabulary in a text (hearst, 1997; foltz, kintsch, & landauer, 1998). texttiling, for instance, identifies lexical chains (essentially, repetition of vocabulary) across adjacent sentences following the notion of identifying subtopic structure in text (hearst, 1997). similarly, foltz, et al.. (1998) describe systems that use latent semantic analysis (lsa) to measure lexical relatedness between text segments by using vector-based similarity between adjacent sentences. the focus of this section is a taxonomy of annotation schemes related to discourse coherence in well-formed texts and noisy essays. the annotations are then used in supervised approaches for building systems to predict discourse coherence quality. table 1 illustrates various aspects of annotation schemes which relate to the time and expertise required for each task, the task goals, and the text types to which the annotation scheme has been applied. in the context of 4 see attali and burstein (2006) for a discussion of feature validity and human scoring criteria in the context of automated essay scoring. burstein, tetreault, and chodorow 38 the set of annotation schemes illustrated in table 1, the holistic scheme that we apply to essay data and that we use for building coherence quality models is relatively low cost in that it requires less expertise and less time. in section 3, we discuss how this low cost scheme was successfully used to assign accurate coherence quality scores to noisy essay data using linguistic features that can be mapped back to criteria in human scoring guides. 2.1 annotation schemes and supervised computational approaches miltsakaki and kukich (2000) were the first to complete an annotation study that focused on discourse coherence in essay data. they implemented a genre-independent annotation task that requires significant expertise related to centering theory (grosz et al. 1995). the goal of this work was to use centering theory to identify topic shift in essays. in miltsakaki and kukich, 100 essays were manually annotated to label specific features that identified where centering information occurred in the text. the data were used to show that essays with a higher proportion of topic shifts (i.e., rough shifts in centering theory) tended to have lower scores (based on a standard 6-point holistic scoring scale, where 6 indicates the best writing and lower scores indicate poorer writing skill.) more recently, wang, harrington, and white (2012) conducted a study that builds on miltsakaki and kukich in which annotators identified specific coherence breakdown points in essays. as part of this work, breakdown points in essays were provided as feedback to non-native english speaker writers. as an outcome essay revisions showed improvements in writing related to coherence features. higgins, burstein, marcu, and gentile (2004) implemented a genre-dependent annotation method and system to predict discourse coherence quality in essays. their annotation scheme required expertise about essay-based discourse structure, specifically, the ability to rate the coherence between specific essay-based discourse elements, such as an essay’s thesis statement and each of its main points. this method showed some success, but it was reliant on organizational structures that are most consistent with responses to expository and persuasive writing. other discourse coherence schemes for wellformed text, such as wolf and gibson (2005), have required annotators to label text segments with particular discourse coherence relations, for instance, cause-effect, example, and elaboration. this work is similar to miltsakaki and kukich, and to higgins et al.’s featurespecific annotation task, both of which require significant expertise. barzilay and lapata (2005, 2008) also implemented a supervised method to predict discourse coherence quality in well-formed texts (also see rus & niraula, 2012). their method takes into account the distributional and syntactic properties of entities (nouns and pronouns) in a text and is theoretically aligned with centering theory (grosz, joshi, and weinstein, 1995). the theory asserts that the discourse in a text contains a set of textual segments that contain discourse entities. centering theory ranks these entities by their importance. the theory can be used to track the entities as well as topic clusters and shifts. barzilay and lapata’s entity-grid algorithm keeps track of the distribution of entity transitions between adjacent sentences and computes a value for all transition types based on their proportion of occurrence in a text. the algorithm has been evaluated with three tasks using well-formed newspaper corpora: text ordering, summary coherence evaluation, and readability assessment. for summary evaluations, the entity-grid approach was successfully used to predict coherence quality of summaries written by humans holistic annotation of coherence quality 39 versus machine-generated summaries. for readability assessment, barzilay and lapata (2008) used their algorithm to determine if a well-formed text was an article from an elementary-level version of encyclopedia britannica or the adult-level version of the same article. they found increased performance of a system that predicted readability (higher or lower grade level) when standard readability syntactic measures were added to the model that used the entity grid alone (also see graesser et al., 2004; sheehan, kostin, futagi, & flor, 2010; and graesser, mcnamara, and kulkowich, 2011 for evaluations that discuss readability assessment and coherence measures). barzilay and lapata’s measures used parse tree information (charniak, 2000) and included sentence length, average parse tree height, average number of nps, average number of vps, and average number of subordinate clauses (schwarm & ostendorf, 2005). like other evaluations of readability assessment, features addressing technical quality were not included since this work analyzed only well-formed texts. using well-formed texts, pitler and nenkova (2008) show that a text coherence detection system yields the best performance when it includes features using the barzilay and lapata (2008) entity grids, syntactic features, discourse relations from the penn discourse treebank (prasad, dinesh, lee, miltsakaki, robaldo, joshi, and webber, 2008), and vocabulary and length features. evaluations were based on 30 wall street journal articles that were annotated for coherence by at least three college students. the coherence annotation scheme, similar to our holistic scheme (described in section 3), asked annotators to provide a rating on a scale of 1 – 5 in response to each of these four questions related to the article’s coherence: (1) how well-written is this article?; (2) how well does the text fit together?; (3) how easy was it to understand?; and (4) how interesting is this article? crossley and mcnamara (2011) also developed an annotation scheme for discourse coherence in essays. this is the only other work, to our knowledge, that includes a technical quality measure (i.e., an annotation measure that assesses standard english conventions) for evaluating a system for automated detection of coherence quality. in crossley and mcnamara, raters assigned a holistic score (from 1 to 6) to multiple traits of an essay which were believed to contribute to essay coherence. these addressed the following categories: (a) quality of the introductory paragraph, including connections between thesis and essay discourse; (b) connections between topics across paragraphs in the essay; and (c) grammar, syntax, and mechanics. as crossley and mcnamara discuss, their annotators were experts so this was a higher cost annotation task with regard to expertise. their findings differ from burstein et al.’s (2010) findings and those reported later in this paper inasmuch as their results did not show technical quality (e.g., grammatical errors) to be a strong predictor of coherence quality in the context of test-takers’ essays. this may be related to the nature of their essay data, which consisted only of essays written by native speakers during an undergraduate composition course. this is possibly a population of more proficient writers who are less likely to make grammatical and spelling errors in their writing. pitler and nenkova’s (2008) annotation scheme probably most closely reflects the annotation approach that we use in our work (burstein et al., 2010), which is described in section 3. like pitler and nenkova, our annotators use a holistic annotation scheme to rate the overall coherence of a text. the task is completed at the document level and requires no specific linguistic or essay-scoring expertise. the major difference between our work and pitler and burstein, tetreault, and chodorow 40 nenkova’s is the data itself. they annotate well-formed text, and we annotate noisy essay data. that said, it is not surprising that our feature set for predicting coherence includes a technical quality feature, and theirs does not. reference scheme class text type scheme goal(s) annotation level expertise level miltsakaki and kukich (2000) fs essay responses capture centering theory “rough shifts” word high higgins et al. (2004) fs essay responses identify lexical relationships between discourse segments in essays (e.g., thesis, main points, supporting ideas, conclusion) discourse segment medium wolf and gibson (2005) fs wall street journal & ap newswire articles identify rhetorical relations related to coherence, e.g., causeeffect discourse segment high pitler and nenkova (2008) h wall street journal articles determine ‘ease’ of reading for a text document low barzilay and lapata (2008) h 1. human and machinegenerated summaries 2. elementary and adultlevel encyclopedia brittanica articles 1. determine discourse coherence quality of summaries 2. predict text level (elementary or adult) document low burstein et al. (2010) h essay responses determine low/high discourse coherence quality document low crossley and mcnamara (2011) h/mt essay responses predict overall essay score • discourse segments • intra-text segments • document high wang, harrington, & white fs essay responses identifies breakdown points in essays • intra-text segments high table 1. taxonomy of discourse coherence annotation schemes. fs is feature-specific, h is holistic, and h/mt is holistic for multiple coherence traits (e.g., topic sentences, paragraph transitions, organization). 3. study: annotation, feature development, and system evaluation as illustrated in table 1, our holistic annotation scheme attempts to be less time-consuming, is performed at the document level, and requires no specific linguistic expertise. it is similar to the document-level annotation approaches used in barzilay and lapata (2008) in which raters assigned a coherence rating on a 7-point scale with regard to relative coherence quality of multiholistic annotation of coherence quality 41 document summaries generated by humans and an automatic summarization system, and to the pitler and nenkova (2008) work in which annotators responded to a set of questions about text quality. the annotation scheme elicits holistic scores from annotators about essays as opposed to asking for labels related to specific linguistic structure as has been done in previous work with discourse coherence in essay data (miltsakaki and kukich, 2000; higgins et al., 2004). it is also more closely aligned with the holistic scoring process, in which readers assign an overall score to an essay based on a general impression, given fuzzy criteria. in our work, annotators essentially assign a holistic (impressionistic) score specifically for discourse coherence quality based on a few features that may contribute to an essay’s coherence (e.g., displays unity, progression and coherence). it is critical that systems designed to evaluate linguistic aspects of test-taker or student writing address valid measurement criteria specified in scoring guides. in theory, these criteria are likely to be related to text features that readers use to construct text meaning. from a practical standpoint, these criteria need to be made available to users for score explanation. users include test-takers who might be preparing for an assessment or who might question a score reported from a standardized test, as well as teachers or students who might inquire about an evaluation in an instructional environment. in the remainder of this section, we will describe our data, the annotation scheme, the training and data annotation process, and the subsequent use of the annotated data to build coherence quality models that predict lowand high-coherence in essays. 3.1 data and annotation data. essay data sub-corpora included writing samples from native and non-native english speakers, ranging from 6th to 12th grade, and across undergraduate and graduate-level writing assessments. writing task types were varied and included expository and persuasive writing, subject-based writing, and summary writing. the total number of essays was 1555. there were five different task types (essay data sub-corpora). data sample sizes ranged from approximately 250 to 400 essays. annotation scheme. the annotation scheme is composed of a 3-point scale as follows: score point 1 indicates that an essay is incoherent (low coherence; no meaning can be constructed); score point 2 indicates that an essay is mostly coherent (essentially coherent; text meaning can basically be constructed, but one or two identifiable points were confusing); and, score point 3 indicates that there are no problems with coherence (high coherence; text meaning can easily be constructed). annotators labeled each essay with one of the three score points. for score point 2, annotators had to do an additional task. score point 2 essays represent those essays in which most of the meaning of the text can be easily understood, but there are one or two identifiable points where coherence breakdown has occurred. for these essays, annotators had to label the “awkward sentence(s)” where the coherence breakdown occurred. annotators could also add comments to describe the confusion points in the essays where it may have taken a few tries to construct meaning or where meaning construction might have been unsuccessful. the annotation protocol includes specific examples of essays at each of the three score points. for burstein, tetreault, and chodorow 42 score point 2, the protocol also contains example awkward sentences that illustrate where sources of coherence breakdown appear in an essay, as in the essay excerpt in figure 1 below. while there may be different interpretations of this text, we have indicated in italicized bold font our judgments about which sentences were awkward. in (1), the awkward sentence [1] reference to “these companies” does not have a natural antecedent and is confusing. in awkward sentence [2], it is not clear who or what “posers” are. in awkward sentence [3], there is a similar issue, and “products” does not have a natural antecedent. these points of confusion can be resolved using inference so that we can essentially construct meaning for this text. it is therefore assigned to score point 2. media in all physical or visual forms from magazines, to movies, to billboards, and television display images of beauty from both male and female. not a day goes by when a person is not exposed to these campaigns. these companies are trying to promote their product or service in the most elegant way, but the qualities seen on the images are compared with themselves who may not match up by far. [1] when look at these images, it's easy to forget that the posers depend on their careers to look the way they do, but the observers do not. [2] thus, a desire to have the same look continuously builds and results in taking action. it is the state of mind that they are dissatisfied with their appearance and they will take any steps to gain that appearance. psychological problems may occur and lead to serious disorders such as anorexia. furthermore, the products or others related to that kind, are sold leading to obsessive and unnecessary spending. [3] figure 1. excerpt from an essay labeled as score point 2. italicized bold font indicates sentences that caused confusion for the reader (annotator). different readers may find different aspects confusing. in the protocol, annotators were instructed not to consider grammatical or spelling errors as lowering coherence scores, unless those errors interfered with the annotators’ ability to construct meaning. this is aligned with shriver’s (1989) assertion that such error types are not notable and we tend to ignore them while reading, unless they slow down our reading and make us re-read. that said, the interference of grammatical and spelling errors should appear in essays that annotators assign score point 1 or 2. in the case of score point 2 writing, these would be cases where such errors caused the annotator (reader) to have to re-read a part of the text because grammatical or spelling errors interfered with understanding of a particular section of the essay. annotator training. during the initial training and development of the annotation scheme (burstein et al., 2010), two annotators worked with two of the authors. both annotators were research assistants who had experience doing varied kinds of linguistic annotation. however, neither annotator had previous experience with annotating discourse coherence. the initial three data sets, including the 6th-12th grade data, and college undergraduate and graduate level writing assessments, were labeled by two annotators as described in burstein et al. (2010). the annotation scheme was developed and refined over a period of about one month in discussion with the first two authors and two annotators from burstein et al. (2010). training began at the point at which the authors and the annotators agreed on the annotation protocol instructions and the examples of essays that illustrated coherence at each of the 3 score points. seventy-five essays were annotated using a training set that contained essays from different task types and grade holistic annotation of coherence quality 43 levels for three of the five data sets reported in this paper. annotations took approximately one week. annotators had less than acceptable agreement for score point 2, but weighted kappa of 0.68 was achieved overall for the 3-point scale. beyond training, annotators continued to label essays using the 3-point scale, but in the end, score point 2 inter-annotator agreement remained a challenge. the annotators could not reach acceptable levels of agreement for this middle score. therefore, the three score points were collapsed to a 2-point scale. score point 1 remained the classification for essays with low coherence, and score points 2 and 3 were combined into a single class to represent essays with high coherence. based on the two-point scale, kappa for this final set of data was 0.67 for approximately 750 essays (see burstein et al., 2010 for details). these annotations were then successfully used to build a discourse coherence system that assigned low and high coherence scores to essays across different writing tasks and test-taker and student populations. correlations with holistic essay scores. part of the decision to include a new feature in an automated essay scoring system (e-rater® in this case; see attali & burstein, 2006) is determined by examining the correlations between the new feature and the human rater holistic essay score. it is important to determine that the new feature is accounting for additional variance and is not redundant with existing features used to predict the holistic essay score, and that it offers more fine-grained information related to the ease of reading (i.e., the ease with which a reader constructs meaning). therefore, for the first three data sets from the 2010 study, correlations were calculated to evaluate the relationships among the 2-point coherence quality scores, existing features, and the human rater holistic essay scores. these holistic essay scores range from 1 to 5 or from 1 to 6, depending on the assessment. the score of 1 indicates a lower quality essay, and the higher scores of 5 and 6 indicate higher quality essays. pearson correlations between human discourse coherence scores and human (overall) holistic essay scores were between 0.46 and 0.58. while these are moderate-to-strong correlations, they still suggest that the discourse coherence quality scores capture characteristics of essay writing that are not fully explained by the holistic essay scores. for example, it is possible for an essay with a low holistic score to have high coherence. that essay might have received a low score for reasons independent of coherence, such as not offering a sufficient amount of supporting evidence, or not responding to the essay prompt. we address this point further in the evaluation section in tables 2 and 3, where in endto-end comparisons, the discourse coherence systems we tested consistently outperformed e-rater in prediction of discourse coherence quality scores. this outcome is consistent with the correlations between the human coherence and holistic essay scores. annotation challenge: score point 2. as mentioned above, score point 2 remained a challenge, since inter-annotator agreement was unacceptably low, resulting in a kappa below 0.60, the minimum value which is generally considered “substantial” agreement (landis & koch, 1977). without a substantial rate of agreement, system building based on annotations is often difficult and unreliable. as a result, we could not build a system with acceptable agreement across our 3point scale, requiring us to instead use the binary classifications of high and low coherence. a possible explanation for low agreement for score point 2 may be related to an individual reader’s inferencing decisions. recall that score point 1 essays are those that are incoherent to the extent burstein, tetreault, and chodorow 44 that the breakdown points could not be identified; a score point 3 essay was one that could be read easily without noticing any points of coherence breakdown. by contrast, score point 2 labels indicated that a few identifiable coherence breakdown points could be located. from the annotation task, however, it seemed that annotators could not easily agree on score point 2 labels. as discussed earlier, graesser et al. (2004) and van den broek (2012) assert that readers construct the meaning of a text. van den broek suggests that to build coherence while reading a text, readers consult a set of strategies, including integrating sections of the text, but also accessing external sources of information, including prior knowledge (see van den broek (1993)’s discussion of “types of inferences”). it would not be unreasonable to assume that our annotators’ prior knowledge or perspectives on an essay topic influence how they build up meaning and how they ultimately assign a coherence score. more concretely, though, reading through several cases in our data where the annotators differed, with one assigning a score point 2 and the other a score point 3, the annotator who assigned the score point 2 typically indicated in a “comments section” that the awkward sentences she identified as causing coherence breakdown seemed “out of place” or were “missing a transition”. the lack of transitional elements seemed to inhibit the annotators’ ability to construct meaning in a segment of an essay, leading to a score point of 2. the level of disagreement among score point 2 essays would indicate that readers (annotators) were constructing meaning in texts differently in these cases: one annotator seemed able to make the transitions while another could not. figure 2 illustrates an example of an essay that one annotator assigned score point 3 and the other score point 2. the annotator who assigned score point 2 indicated that transitions were missing and certain sentences seemed out of place. on the other hand, the annotator who labeled the essay as score point 3 was able to construct meaning and figure out the relationships between these sentences and other sentences either in the text or, perhaps, related to some prior knowledge about the topic. it may be difficult to build systems that identify this kind of fine-grained distinction when humans cannot agree. however, what we can do is to capture text segments in essays, such as discourse connectors, or repeated concepts that serve as a continuous thread in the text. capturing this information can lead to the development of useful feedback in instructional settings, which can help users understand if they are missing transitions that could result in a lack of clarity in their writing. systems could also identify text segments or concepts that did not appear to be related to other parts of the text (lexically) and could flag these as potentially unconnected ideas. features capturing discourse relations and repeated concepts are used in our system and are described below. in theory, these features could be used to identify weak relationships between text segments as well. 3.2 system features and evaluation using n-fold cross-validation with c5.0, a decision-tree classifier (quinlan, 1993), systems were built to classify essays in each of the five data sets, which represented five different writing tasks and populations. across the five tasks, there was a mixture of native and non-native speaker data. task types varied and included essays that call for summarization and for expository and persuasive writing. figures 3 and 4 illustrate relevant discourse features that were predictive of coherence quality. holistic annotation of coherence quality 45 throughout the history of time there have been many great leaders, along with many poor ones. these leaders exemplify traits above all others that help them become what we see them as today. gandhi was a powerful man who used civil disobedience as his weapon against guns, to help hindu's and muslims unite in peace. though his acts started off small they gradually grew to be what we consider the greatest non-violent movements of all time. gandhi's first movement was when he burnt his pass card which restricted muslims and hindus to travel without them. for this he was beaten by british soldiers and taken to prison. later on mahatma gandhi started such movements as the homespun movement and the salt march. [1] the homespun movement was when gandhi refused to wear any british cloth therefore made his one clothes that were merely whites pieces of cloth worn in a toga style. from there he made speeches all across india about this movement. [2] hindu's and muslim's responded with great support to this movement and also participated. … gandhi's life came to an end when he was shot walking around at a nonviolent movement by a muslim who disagreed with his beliefs. once gandhi died the fighting between this muslim's and the hindu's went down hill and only got worse until the british freed india to become its own country. so although gandhi worked for many years to bring peace to a land of war, his dream did not come true until many years after his death. figure 2. excerpt from an essay labeled as score point 2 by one annotator and score point 3 by a second annotator. the score point 2 annotator indicated that awkward sentence [1] “needed a transition” and awkward sentence [2] “did not fit into the paragraph”. 3.2.1 feature set. the set of features used in each model is intended to capture holistic rubric criteria from scoring guides that correspond to discourse coherence in high quality essays: “displays unity, progression, and coherence, though it may contain occasional redundancy, digression, or unclear connections”; and in low quality essays: “is not clearly organized; some parts may be clear while others are disjointed or confused... errors in grammar interfere with reader understanding”. these features include the following: proportion of grammatical errors using error features from e-rater (attali & burstein, 2006; burstein et al., 2013), entity-grid transition probabilities (barzilay and lapata, 2005; barzilay& lapata, 2008), rhetorical structure theory (rst) features (mann and thompson, 1988) derived from an rst parser (marcu, 2000), and a type-token feature that computes the entity type/token ratios from specific words and terms recovered from the entity grid. high-level descriptions of the features are given below. (for more detailed descriptions of the entity-grid approach, refer to barzilay & lapata, 2008); for more detail about the rhetorical structure tree parser, see marcu, 2000). below, we describe each feature type and its corresponding scoring rubric criterion. entity-grid transition probabilities (entity grids). entity-grid transition probabilities are intended to address unity, progression and coherence by tracking nouns and pronouns in text. the entity grids characterize occurrences of words across a text in syntactic roles. an entity grid is burstein, tetreault, and chodorow 46 constructed in which all entities (nouns and pronouns) are represented by their roles in a sentence (i.e., subject, object, other). entity grid transitions track how the same word appears in a syntactic role across adjacent sentences. examples of transition types are subject-object, objectobject, and object-other. entity transition probabilities represent the proportions of different entity transition types in a text. these probability values are used as features to build coherence models. e-rater grammatical error features (ergramerror). these features address errors in grammar that could interfere with a reader’s ability to construct meaning. for example, in a sentence such as, it make student to have a good rest than staying in the school for whole term, a reader may have to pause and consider the possible intentions of the writer. e-rater identifies more than 30 kinds of errors in grammar, such as subject-verb number disagreement, in word usage, such as missing articles, and in spelling (attali & burstein, 2006; burstein et al., 2013). aggregate counts of these individual errors are used as features in e-rater, and these same features were used in our model for predicting discourse coherence quality. maximum lsa value for distant sentence pairs (maxlsa). in the course of our research, we observed that some high coherence essays tended to reintroduce earlier topics later in the essay. this reintroduction provided a sort of “thread” that maintained coherent connections promoting discourse unity within the text. to represent this, we used latent semantic analysis (lsa), a statistical method used to determine semantic similarity between text units (i.e., do these text units use similar vocabulary) (foltz et al., 1998). previously, higgins, et al.(2004) had used a similar technique, random indexing (kanerva, kristoferson, & holst, 2000; sahlgren, 2001), to measure similarity between discourse unit text segments within an essay (e.g., similarity between the thesis statement and the conclusion). with lsa, we computed similarity values for all sentence pairs in the essay and found that the maximum lsa value associated with pairs that were separated by more than five intervening sentences was highly correlated with the human coherence annotations. this feature, which measures similarity at a distance and should, therefore, be sensitive to reintroduction of topics, was added to our set of features. re-introducing a concept later in a text is consistent with the backward inference strategy discussed in van den broek (1993). the positive correlations between the maxlsa value and discourse coherence scores suggest that clear re-introduction of material presented earlier in a text does contribute to higher coherence. table 4 indicates that it was predictive of the holistic coherence quality in two of the five data sets we tested. rst-derived features (rst). rhetorical relations tell us about discourse connections with regard to how clauses and sentences in a text are rhetorically linked. rst relations include, for example, antithesis, cause, comparison, elaboration and example. we used rhetorical relations and derived features to evaluate if and how certain rhetorical relations, combinations of rhetorical relations, or rhetorical relation tree structures might contribute to discourse coherence quality. these included the following: (a) relative frequencies of n-gram rhetorical relations in the context of the tree structure (unigrams, or occurrences of a single relation (e.g., themeshift); bigrams (e.g., “themeshift -> elaboration”); and trigrams (e.g., “themeshift -> elaboration -> holistic annotation of coherence quality 47 circumstance”); (b) relative proportions of leaf-parent relation types in rhetorical structure trees; and (c) counts of root node relation types in rhetorical structure trees. see figure 4 below for an example of an rst parse tree (marcu, 2000). type/token ratios for entities (type/token). type/token ratios can be used to track redundancy in essays. as discussed above, the entity grid transition probabilities are used as features in the coherence models, but they measure only local transition patterns in adjacent sentences rather than more global reuse of words. for instance, if there is a high probability for the “subjectsubject” transition to occur in an essay, this indicates that the writer is repeating an entity in subject position in adjacent sentences, but we do not know if the same word is being repeated throughout the text or if a variety of words are. the type/token feature can distinguish between these two cases. for instance, if there are 10 subject-subject transitions, and 5 different word types appear in these transitions, the type/token ratio would be 5/10 (0.50); if, on the other hand, there is only 1 word type (e.g., “i”), the type/token ratio would be 1/10 (0.10). thus, higher ratios indicate that more concepts are being introduced in a given syntactic role, and lower ratios indicate fewer concepts. school is the most important institution in the life of a person. it is in school that a child learns the most. school life is the most significant part of the growing up years of a child. a child absorbs theortical knowledge, moral values and cultural ethics in school. therefore, a great school helps to mould a child into an overachieving and succesful person. one most important change i would like to incorporate in the school i attented was to be able to have a choice in subjects for study and not a rigid curriculum. this according to me is very important. every student is a different individual and therefore each student has his own likes and dislikes. i truly believe that by allowing students to choose their own subjects, students would study what they are interested in and thereby not only gain exceptional knowledge but attain a great affinity for the particular field. this would in turn, propel them to achieve greater heights and conquer more and more difficult goals in their particular field. i also believe that the school must continue to have a few compulsory subjects like language, math, geography and history. by doing this, the child will gain all round knowledge which is extremely important to become succesful in today's competitive world. i, as a student did not have this choice and was therefore compelled to study only what the school thought was good for me. today, i feel i could have had a significant advantage over others if i could have studied computer applications and software right from school. i would like to conclude by saying that there is no better place to build character, values, morals, spirit and knowledge than school. and therefore it is the responsibility of the school to allow us, its students, the freedom to choose a career path and field which would make us successful in this over achieving world of today. figure 3. illustration of example discourse features automatically detected and used to generate a discourse coherence quality score for this essay written by a non-native speaker, from the ep/nns data set described in tables 2 and 3. italicized bold indicates the long distance sentence pair with the maximum lsa value, illustrating the re-introduction of a similar concept, promoting discourse unity in the essay; turquoise highlighted words indicate entity types used in the type/token feature that represents the repetition of entities extracted from the entity grid. in this case, there are reasonably strong relationships among a small and related variety of entities, e.g., school, child, and student(s), which support lexical cohesiveness, or discourse unity. other entities are referential or personal pronouns associated with statement of opinion. burstein, tetreault, and chodorow 48 3.2.2 evaluation. developing a system that assigns discourse coherence on a 3-point scale has been a challenge. as discussed earlier, it appears that for the score point 2 category, annotators had lower agreement. therefore, the score points 2 and 3 were collapsed to represent the “high” coherence class, and all score point 1 annotations are assigned to the “low” coherence class. using these two classes, we have built systems that assign high and low coherence scores and compared them to three baselines. baseline systems included: (1) assigning the majority class (i.e., assignment of the more frequent category, “high coherence", to every essay); (2) using the erater automated essay scoring system for assigning coherence scores; and (3) using only e-rater’s grammatical error features to assign coherence scores. performance on high and low coherence essays is measured in terms of precision, recall and f-score. a system’s precision on high coherence essays, for example, is the number of essays that both the system and the human annotator agreed are “high”, divided by the number of essays that the system labeled “high”. recall is the number of essays that both the system and the human annotator agreed are “high”, divided by the number of essays that the human labeled “high”. the f-score is the harmonic mean of precision and recall. precision, recall and f-score can be calculated for “low” coherence essays in a similar manner. tables 2 and 3 below show separate performance measures for high and low predictions, respectively. these measures are more informative about system performance than the overall values where low and high predictions are collapsed. the best system outperformed the f-score baselines for four of the five data sets for prediction of “high” coherence, and across all five data sets in the prediction of “low” coherence. for the professional proficiency exam, it should be noted that the data for this exam contained a much lower proportion of score point 1 essays than the other data sets, and the system had a more difficult time predicting the low coherence essays in this set. in tables 2 and 3, the best system (boldface) for each data set was built using a combination of the features described above in section 3.2.1. table 4 shows the features that contributed to the best system for each data set. system ep/ nnc n=196 summary/nnc n=304 ep/gl n=210 ppe n=355 ep/ 6-12 n =220 p/r/f p/r/f p/r/f p/r/f p/r/f majority class 77/100/87 76/100/76 82/100/90 91/100/95 87/99/92 e-rater 87/86/86 81/90/86 86/97/91 91/97/94 89/91/90 ergramerror 90/85/88 78/93/85 86/97/91 91/100/95 90/94/92 best system 93/91/92 85/94/89 91/94/93 93/96/95 94/99/96 table 2. p/r/f indicates system precision, recall, and f-scores (x 100) for “high” discourse coherence essays for 5 essay data sets. ep/nnc is expository and persuasive, non-native, college level assessment writing; summary/nnc is summary, non-native, college level assessment writing; ep/gl is expository and persuasive, graduate level assessment writing; ppe is the professional proficiency exam; and ep/6-12 is elementary, middle, and high school writing. e-rater = e-rater system features. ergramerror = e-rater grammatical error features. 4 discussion and conclusions in this paper, we have described and compared different discourse coherence annotation schemes and related studies. most work in this area has been evaluated for building systems to handle well-formed texts. there has also been considerably more work in the area of coherence and its holistic annotation of coherence quality 49 relationship to readability (reading for comprehension) as opposed to coherence quality in noisy data (in this case, essays) (reading for evaluation). we have shown that a holistic annotation scheme that requires no linguistic expertise can be successfully used to build discourse coherence systems that classify lowand high-coherence quality in 1555 essays from 5 different data system ep/ nnc n=58 summary/nnc n=94 ep/gl n=47 ppe n=37 ep/ 6-12 n =33 p/r/f p/r/f p/r/f p/r/f p/r/f majority class 0/0/0 0/0/0 0/0/0 0/0/0 0/0/0 e-rater 54/57/55 56/33/40 68/32/43 14/5/8 32/27/30 ergramerror 58/69/63 39/14/20 68/28/39 0/0/0 41/27/33 best system 72/75/74 71/47/57 69/57/63 48/32/39 84/52/64 table 3. p/r/f indicates system precision, recall, and f-scores (x 100) for “low” discourse coherence essays for 5 essay data sets. ep/nnc is expository and persuasive, non-native, college level assessment writing; summary/nnc is summary, non-native, college level assessment writing; ep/gl is expository and persuasive, graduate level assessment writing; ppe is the professional proficiency exam; and ep/6-12 is elementary, middle, and high school writing. e-rater = e-rater system features. ergramerror = e-rater grammatical error features. data set best system feature set ep/ nnc entitygrid + type/token + rst + maxlsa summary/nnc ergramerror + rst ep/gl entitygrid + ergramerror + type/token + rst ppe entitygrid + ergramerror + type/token + rst ep/ 6-12 entitygrid + ergramerror + type/token + rst+ maxlsa table 4. feature sets used in the best system for each data set. ep/nnc is expository and persuasive, non-native, college level assessment writing; summary/nnc is summary, non-native, college level assessment writing; ep/gl is expository and persuasive, graduate level assessment writing; ppe is the professional proficiency exam; and ep/6-12 is elementary, middle, and high school writing. figure 4. illustration of an rst representation for a sentence from the essay in figure 3. the rst relations contribute to discourse relationships and conceptual connectedness within the essay that can support readers in building meaning and coherence. in this illustration, the rst parser identified a condition relation within the sentence “today, i feel today, i feel i could have had a significant advantage over others if i could have studied computer applications and software right from school. condition relation burstein, tetreault, and chodorow 50 i could have had a significant advantage over others if i could have studied computer applications and software right from school.” samples, including native and non-native speaker populations, 6th through 12th grade, and college and graduate-level populations, and from numerous topic domains across the multiple essay question topics in our essay sub-corpora. the features that are used to build models can be mapped back to scoring guide criteria in the following way. entity-grid transition probabilities offer information that can be related to the organization of ideas and how they are distributed in a text; the type-token feature based on the entities from the grid offers information related to the degree of repetition of ideas in the text; rhetorical structure features provide a sense of the organizational and development units of discourse in the text; and the grammatical error features provide a sense of the technical quality of the text. these features could be used to provide explanations of coherence scores and feedback to students, test-takers, teachers, and other score recipients. as discussed throughout the paper, the most challenging aspect in this work was annotation of score point 2, “mostly coherent” essays – specifically, those essays where the reader experiences one or two identifiable points of confusion (coherence breakdown). for system model building and score prediction, we collapsed the score point 2 and score point 3 essays into a single “high” class. essentially, when essays had very low or very high coherence, annotators could agree. a possible explanation for the assignment of score point 2 may be related to differences in how individual readers construct meaning or employ inference during reading. this is consistent with halliday and hasan’s (1976) assertion that a text is a semantic unit and that the text continuity is related to how an individual constructs its meaning. especially with regard to employing inference strategies, this might also be explained by graesser et al.’s 2004 assertion that coherence is a psychological construct where individuals may vary in their use of inference, and similarly, by van den broek’s (1993 and 2012) discussions of readers’ reading strategies that take into account not only the text itself, but also external sources, such as internet search and prior knowledge. in a study that specifically investigated readers’ construction of text meaning, kwong (2010) conducted a discourse coherence annotation task using aesop’s fables. in the study, for sentences in each story, three annotators specified key words, main information, inferred information, and the relation to the moral of the story. with regard to keywords and main information, agreement could be reasonably measured by word overlap; however, for inference and relation to the moral of the story, annotator responses varied and seemed to rely on world knowledge. since assignment of score point 2 indicates that the reader had some confusion, it is possible that while one reader was confused, another reader might not be confused because he or she made an inference that held the text together. an interesting direction for future work would be to investigate the interaction among inference, prior knowledge, and the distribution of content within a text by replicating kwong’s study, but using essays instead of fables. holistic annotation of coherence quality 51 acknowledgements we would like to thank daniel blanchard and slava andreyev for engineering support in system development and evaluation. we would also like to thank sarah ohls, waverly van winkle, and jonathan schmidgall for their valuable input during the development of the annotation protocol and for their subsequent annotation of the essays used in this research. references attali, y. & burstein, j. (2006). automated essay scoring with e-rater v.2.0.. journal of technology, learning, and assessment, 4(3). barzilay, r. & lapata, m. (2008). modeling local coherence: an entity-based approach. computational linguistics, 34(1), 1–34. barzilay, r. & lapata, m. (2005). modeling local coherence: an entity-based approach. in proceedings of the 43rd annual meeting of the association of computational linguistics, 141–148, ann arbor, mi. burstein, j. tetreault, j., & andreyev, s. (2010). using entity-based features to model coherence in student essays. in proceedings of 10th annual meeting of the human language technology and north american association for computation linguistics, 681–684, los angeles, ca. burstein, j., tetreault, j., & madnani, n. (2013). the e-rater® automated essay scoring system. in shermis, m.d., & burstein, j. (eds.), handbook for automated essay scoring. new york: routledge. charniak, e. (2000). a maximum-entropy-inspired parser. in proceedings of the 1st annual meeting of the north american chapter of the association for computational linguistics, 132–139, seattle, wa. coward, a. f. (1950). the method of reading the foreign service examination in english composition. ets rb-50-57, princeton, nj: educational testing service. crossley, s. a., & mcnamara, d. s. (2011). text coherence and judgments of essay quality: models of quality and coherence. in l. carlson, c. hoelscher, & t. f. shipley (eds.), proceedings of the 29th annual conference of the cognitive science society. 1236–1241. austin, tx: cognitive science society. elliot, n. & klobucar, a. (2013). automated essay evaluation and the teaching of writing. in shermis, m.d. & burstein, j. (eds.), handbook for automated essay scoring. new york: routledge. foltz, p. w., kintsch, w., & landauer, t. k. (1998). textual coherence using latent semantic analysis. discourse processes, 25(2&3): 285–307. godshalk, f. i., swineford, f. & coffman, w.e. (1966). the measurement of writing ability. new york, ny. college entrance exam board. graesser, a.c., mcnamara, d.s., & kulikowich, j. (2011). coh-metrix: providing multilevel analyses of text characteristics. educational researcher. 40: 223–234. graesser, a. c., mcnamara, d. s., louwerse, m. m., & cai, z. (2004). coh-metrix: analysis of text on cohesion and language. behavior research methods, instruments, & computers, 36(2), 193–202. grosz, b., joshi, a., & weinstein, s. (1995). centering: a framework for modeling the local coherence of discourse. computational linguistics, 21(2): 203–226. halliday, m. a. k., & hasan, r. (1976). cohesion in english. london: longman. hearst, m. (1997). texttiling: segmenting text into multi-paragraph subtopic passages, computational linguistics, 23 (1), 33–64. higgins, d., burstein, j., marcu, d., &. gentile, c. (2004). evaluating multiple aspects of coherence in student essays. in proceedings of 4th annual meeting of the human language technology and north american association for computation linguistics, 185–192, boston, ma. huddleston, e. m. (1952). measurement of writing ability at the college-entrance level: objective vs. subjective testing techniques. ets rb-52-57. kanerva, p., kristoferson, j., & holst, a. (2000). random indexing of text samples for latent semantic analysis. in l. r. gleitman & a. k. josh (eds.), in proceedings of the 22nd annual conference of the cognitive science society, 1036, mahwah, nj: erlbaum. kwong, o. y. (2010). constructing an annotated story corpus: some observations and issues, in proceedings of the language resources and evaluation conference, 2062–2067. malta. burstein, tetreault, and chodorow 52 landis, j. r. & koch, g. g. (1977). the measurement of observer agreement for categorical data. biometrics. 33:159–174. mann ,w.c. & thompson, s. (1988) rhetorical structure theory: toward a functional theory of text organization. text, 8(3), 243–281. marcu, d. (2000). the theory and practice of discourse parsing and summarization. cambridge, ma: the mit press. miltsakaki, e. & kukich, k. (2000). automated evaluation of coherence in student essays. in proceedings of the language resources and evaluation conference, athens, greece. pitler, e. & nenkova, a. (2008). revisiting readability: a unified framework for predicting text quality. in proceedings of the conference on empirical methods in natural language processing, 186–195, edinburgh, scotland. prasad, r., dinesh, n., a., lee, miltsakaki, e., robaldo, l., joshi, a., & webber, b. (2008). the penn discourse treebank 2.0. in proceedings of the language resources and evaluation conference. quinlan. j. r., (1993). c4.5: programs for machine learning. morgan kaufmann publishers. rus, v., & niraula, n. b. (2012). automated detection of local coherence in short essays based on centering theory, in proceedings of cicling 2012, iit delhi, india. schwarm, s. e. & ostendorf, m. (2005). reading level assessment using support vector machines and statistical language models. in proceedings of the 43rd annual meeting of the association for computational linguistics, 523–530, ann arbor, mi. sahlgren, m. (2001). vector based semantic analysis: representing word meanings based on random labels. in proceedings of the esslli 2001 workshop on semantic knowledge acquisition and categorisation. helsinki, finland. shriver, k. a. (1989). evaluating text quality: the continuum from text-focused to reader-focused methods. ieee transactions in professional communication, 32(4): 238–255. sheehan, k. m., kostin, i., futagi, y. & flor, m. (2010). generating automated text complexity classifications that are aligned with targeted text complexity standards, ets rr-10-28, educational testing service: princeton, nj. van den broek, p. w. (2012). individual and developmental differences in reading comprehension: assessing cognitive processes and outcomes. in: sabatini, j.p., albro, e.r., o'reilly, t. (eds.), measuring up: advances in how we assess reading ability, 39–58. lanham: rowman & littlefield education. van den broek, p., fletcher, c. r., & risden, k. (1993). investigations of inferential processes in reading: a theoretical and methodological integration. discourse processes, 16, 169–180. wang, y., harrington, m, and white, p. (2012). detecting breakdowns in local coherence in the writing of chinese english speakers. the journal of computer assisted learning. 28, 396–410. wolf, f., & gibson, e. (2005). representing discourse coherence: a corpus-based study. computational linguistics, 31(2):249–288. dialogue & discourse 8(2) (2017) 225-248 doi: 10.5087/dad.2017.210 signalling implicit relations: a pdtb – rst comparison lucie poláková polakova@ufal.mff.cuni.cz jiří mírovský mirovsky@ufal.mff.cuni.cz pavlína synková synkova@ufal.mff.cuni.cz institute of formal and applied linguistics faculty of mathematics and physics charles university in prague malostranské náměstí 25 118 00 prague, czech republic editor: amanda stent submitted 05/2016; accepted 07/2017; published online 12/2017 abstract describing implicit phenomena in discourse is known to be a problematic task, from both theoretical and empirical perspectives. the present article contributes to this topic by a novel comparative analysis of two prominent annotation approaches to discourse relations (coherence relations) that were carried out on the same texts. we compare the annotation of implicit relations in the penn discourse treebank 2.0, i.e. discourse relations not signalled by an explicit discourse connective, to the recently released analysis of signals of rhetorical relations in the rst signalling corpus (rst-sc). the intersection of corresponding pairs of relations is rather a small one, but it shows a clear tendency: unlike the overall signal distribution in the rst-sc, more than half of the signals in the studied intersection are of semantic type, formed mostly by loosely defined lexical chains. our data transformation allows for a simultaneous depiction and detailed study of the two resources. keywords: discourse coherence, rhetorical relations, implicit discourse relations, signals, rhetorical structure theory, penn discourse treebank, rst signalling corpus 1. introduction and motivation in recent discourse-oriented research, great attention has been paid to discourse markers or discourse connectives1, which are agreed to be the most apparent anchors of discourse relations, and in this way substantially contribute to discourse coherence. nowadays, there are even large efforts to thoroughly describe such discourse-relational devices in form of lexicons (e.g. (stede and umbach, 1998; roze et al., 2012)) and other inventories that would serve both linguistic purposes and nlp. a natural step in the research on discourse coherence is then to answer the general question of how coherence is established if such a connective device is not present between given text segments. from the linguistic viewpoint, we may want to describe what other overtly present language 1. disregarding the strictness / looseness of possible definitions of these categories c©2017 lucie poláková, jiří mírovský and pavlína synková this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). poláková, mírovský and synková elements, even elements not directly associated with discourse coherence functions, can play a role in our understanding of such a connection. from the cognitive viewpoint, we may be interested in the way our mind processes such a connection (let us say, an implicit relation), why we all (mostly) understand and interpret a given text in the same way and what part of the overall meaning is inferred from a co-textual, situational or world knowledge context. finally, from the perspective of automatic text processing, we may want to model discourse coherence with the help of detecting the same signals/features a human normally uses for full comprehension of texts. in all known attempts, the annotation of implicit relations (and connectives) has been a difficult task. the inter-annotator agreement figures for implicit relations from various discourse annotation projects are mostly calculated together with the figures for explicit relations and thus the actual statistics on implicit relations remains unpublished2. personal discussions with annotators experienced with this task shows that the agreement on implicit relations is perceived rather low, far beyond satisfactory. this was also the reason for postponing such a task in our own work (annotation of discourse phenomena in the prague dependency treebank) for later phases, when the annotators get more experienced in recognizing discourse relations and also when we have gathered feedback from the results of similar projects. the aim of this article is to contribute to our understanding of how coherence is established in the absence of connective devices, by performing a comparative corpus analysis. in particular, we make use of the existence of multiple annotations for the same wall street journal texts. we focus on the penn discourse treebank 2.0 (pdtb, prasad et al. (2008b)) annotation of implicit discourse relations, look for their counterparts in the rst signalling corpus (rst-sc, das et al. (2015)) and analyze coherence signals assigned to these relations in the rst signalling annotation. to put it differently, we are interested less in what connective could be inferred to anchor such an implicit connection, and more in what language element(s) in the text (and possibly other pragmatic settings) influenced the pdtb annotator to insert a specific connective token (and no other) and to annotate a specific label (and no other). our hypothesis is that there must be evidence for the annotator’s decision in the text itself and the rst annotation of signals can provide this evidence. also, it can provide essential empirical information on how far such a description can actually reliably proceed, as we assume that the way humans speak and write is to a certain degree vague. the results of our study can hopefully be of use for various nlp tasks concerned with discourse phenomena. finding out how (and to what degree) implicit relations are anchored on the surface can contribute to enhancing systems using discourse-aware features. in shallow discourse parsing and especially classification of implicit relations, there has been noticeable progress in the last years (e.g. (pitler et al., 2009; lin et al., 2009; park and cardie, 2012) and others), also thanks to the intensive search for the best feature combinations. the latest work in this field (li and nenkova, 2014) contains a discussion on sparsity issues of lexical features and offers a comparison of different settings of lexical and syntactic features. the authors arrive at the observation that those lexical 2. to our knowledge, there are only a few published inter-annotator agreement figures solely for implicit relations. in the penn discourse treebank 2.0, the percentual agreement between two annotators on setting the extent of the argument spans was 85.1% with an exact-match metric, and 92.6% with a partial match metric (miltsakaki et al., 2004; prasad et al., 2008b). agreement on the inserted connectives was 72% (for 5 semantic categories), (miltsakaki et al., 2004, p. 7). a recent measurement for implicit relations in the turkish discourse treebank (zeyrek and webber, 2008) reports chance-corrected kappa values of 0.52 for the class level, 0.43 for the type level and 0.34 for the subtype level (zeyrek et al., 2015). 226 signalling implicit relations: a pdtb – rst comparison features computed as most significant for the system performance captured semantic information quite related to the nature of the relation, but also, that these features were mostly domain-specific3. last but not least, as already addressed in poláková (2015, pp. 144–154), the present analysis is intended as a springboard for designing an annotation scenario for “discourse relations with no connective” in a future release of the prague dependency treebank (a treebank of written journalistic czech). since there already is (apart from explicit discourse relations) a complex annotation of underlying syntax, information structure (topic-focus articulation), genres, pronominal and nominal textual coreference and also bridging anaphora (bejček et al., 2013), it makes methodologically more sense to focus on already available signals rather than, for instance, on inserting a connective into a connection with an already annotated strong contrastive bridging link4. 1.1 from implicit relations to signals in this article, we use the term implicit discourse relations in agreement with the penn discourse treebank terminology; this term signifies discourse relations that are not signalled by any connective device in the text (on the surface). more details on the annotation of implicit relations in the pdtb 2.0 are given in section 2.15. (1) [several organizations, including the industrial biotechnical association and the pharmaceutical manufacturers association, have asked the white house and justice department to name candidates (for judges) with both patent and scientific backgrounds.] [the associations would like the court to include between three and six judges with specialized training.] in example 1 from the pdtb, the two sentences represent two discourse units (discourse arguments, in the pdtb terminology) related to each other with an implicit discourse relation in the pdtb and with a rhetorical relation in the rst-discourse treebank (rst-dt, carlson et al. (2002))6. the example represents an exact match in two ways: (i) the text spans of both discourse units match exactly in the two frameworks and (ii) there exists a coherence relation between these units in both annotations. they also represent a partial match7 in the categories assigned to the relation in both annotations: the sense of the relation in the pdtb was annotated as expansion–restatement– specification and the implicit connective “specifically” was inserted, in the rst-dt, the relation was annotated as elaboration–additional (the left argument represents the nucleus). finally, the rst signalling corpus provides the relation with two signals: 1. combined; (semantic+syntactic); (repetition + subject_np); comment: associations 2. single; semantic; lexical_chain; comment: a few lexical chains. these signals of implicit relations, their types, combinations and distribution, and the linguistic implications of their analysis, are the very focus of our research. 3. like e.g. the words share or cent for contrast in the financial domain 4. for details on annotation categories regarding phenomena “beyond the sentence boundary” in the prague dependency treebank see zikánová et al. (2015), chapters 2–5. 5. to avoid confusion, it has to be stressed again that the term implicit relations refers to the implicitness (absence) of a connective device, not to the implicitness of one of the connected discourse units, as in [i think that] he has already left because his car is gone. 6. in the rst analysis, the first sentence consists of two elementary discourse units (edus), the first one being split by an embedded unit. the rhetorical relation in question thus holds between a subtree of an rst tree with three terminal nodes on one side and a single terminal node on the other. compare also figure 1. 7. for details on label mapping between the two frameworks see section 3.4. 227 poláková, mírovský and synková figure 1: rst analysis (a subtree from the rst tool) for the sentences in example 1 2. global and local coherence modeling in corpora the leading frameworks in discourse analysis access discourse phenomena from two main perspectives on discourse coherence modeling. the so-called “global”, top-down approaches represent a whole document as a single connected structure (such approaches are also referred to as “deep” discourse parsing), while “local”, bottom-up approaches access discourse phenomena from the syntactic perspective (“shallow” discourse parsing). the most influential frameworks among the former are rhetorical structure theory (rst, mann and thompson (1988)), the segmented discourse representation theory (sdrt, asher and lascarides (2003)) and the discourse graphbank (wolf and gibson, 2005). the latter, “local” direction of discourse analysis is best represented by the lexically grounded approach of the penn discourse treebank (prasad et al., 2008b), which accesses discourse relations in the first place by searching for their lexical anchors – discourse connectives. also, the pdtb does not make any claims about the shape of the overall discourse structure. we are well aware of this basic difference in the theoretical aiming and description methods of the two (three) corpora targeted in this research, namely the penn discourse treebank, (the rst discourse treebank) and the rst signalling corpus. (a detailed description of these projects follows in sections 2.1, 2.2 and 2.3.) the different basic departure points of these resources can be observed already in the first step of such an analysis: on the segmentation of a discourse (text) to (elementary) discourse units8. the differences in the conception of discourse units and their hierarchization in both approaches are addressed below in the section on argument mapping (3.3). also, the relations between (among) discourse units are approached differently in these frameworks – whereas the rst investigates rhetorical relations, the pdtb uses the term discourse relations. acknowledging the differences between the notions behind these terms, for the purposes of this article we treat these relations in the same way and use the term discourse relations or, a more neutral term coherence relations. nevertheless, both recent developments in the discourse-oriented research community – which is, among other things, comparing the existing frameworks for discourse annotation or mapping them onto one another (e.g. (versley and gastel, 2013; prasad and bunt, 2015; scholman et al., 2016)) – and our own experience with discourse annotation (poláková et al., 2013) indicate that even such theoretical and descriptive variation offers areas of intersection where certain phenomena can be observed with methodological soundness. 8. the rst framework uses the term discourse units (dus), while the pdtb segments are called discourse arguments. 228 signalling implicit relations: a pdtb – rst comparison 2.1 penn discourse treebank the penn discourse treebank (prasad et al., 2008a) contains annotation of discourse relations over the 1 million word wall street journal corpus in its current version pdtb 2.0 (prasad et al., 2008b), version 3.0 is forthcoming. the pdtb approach to discourse relations is based on the identification of discourse connectives as anchors of local discourse relations. following this lexically grounded approach, the annotation in the pdtb first consisted in marking relations signalled by explicit discourse connectives (according to a predefined list of these expressions). the location and the extent of the arguments of explicit relations were not restricted but all relations were assumed to hold between two and only two arguments. as a second step, discourse relations that have no connective device on the surface were annotated. the annotators were instructed to read adjacent sentences within a paragraph and, for each pair of sentences not already linked by an explicit connective, they made a decision as to whether a discourse relation is present. then they inserted a connective conveying the best the meaning of the relation and provided a label for this relation. relations with an inserted connective are called implicit9. apart from adjacent sentences within a paragraph, implicit relations were also annotated intra-sententially between clauses delimited by a semi-colon or a colon. however, in three situations a connective could not be inserted between the adjacent sentences. these situations have led to the introduction of three additional categories: altlex (alternative lexicalization of a connective), entrel (entity-based relation) and norel (no relation). in the case of altlex, insertion of a connective would lead to redundancy, since the relation was expressed by an expression or phrase not included in the original list of connectives (e.g. one reason is, that is why, further, see prasad et al. (2010)). in the case of entrel, only an entity-based relation was detected between the given arguments (annotators were not able to insert an appropriate connective); and finally in the case of norel, no discourse relation was perceived between the sentences. the distribution of all types of relations (and norels) in the pdtb is the following – overall, there are 40 771 relations, 18 459 of them explicit, 16 224 implicit10, 624 signalled by altlexes, 5 210 entrels and 254 norels. explicit relations thus represent 45.3% of all relations, the implicit ones 39.8% and altlexes 1.5%. sense annotation in the pdtb 2.0 uses a three-level hierarchy containing 4 general classes, 16 sense types and 23 subtypes (30 possible senses if choosing the most detailed label available prasad et al. (2008b, p. 2965)). sense annotation was provided for explicit, implicit and altlex relations only, as entrels and norels do not indicate (according to the pdtb approach) the presence of a discourse relation. this fact represents an important difference compared to the rst approach – relations which are called entity-based in the pdtb are not different from the other discourse (rhetorical) relations according to the rst-dt. separately, the pdtb annotates the attribution of each discourse relation and of each of its two arguments. attribution is a relation between agents who report some content and the reported content and thus, according to the pdtb approach, it is not a discourse relation. here, again, there is a discrepancy between the two frameworks – attribution relations in pdtb correspond to regular rhetorical relations in the rst-dt. this issue is further addressed in section 3.3. 9. the annotators could insert one or two connectives for a single implicit relation, and to each of them assign one or two labels. for the present study we only take into account the first label of the first connective. 10. if we consider implicit discourse relations with two inserted connectives as two relations, the total number of implicit relations in the pdtb is 16 224. if we consider them as one, the total number is 16 053. later on in this paper, we use the latter number. 229 poláková, mírovský and synková 2.2 rst discourse treebank the rst discourse treebank (rst-dt, carlson et al. (2002)) is a language resource annotated for rhetorical relations over 385 wall street journal articles (176 000 words) selected from the penn treebank (marcus et al., 1993). the chosen texts concern a variety of topics and were annotated manually under the rhetorical structure theory framework (mann and thompson, 1988). rhetorical relations are considered to hold between elementary discourse units (and/or segments composed of these units) which often correspond to clauses but are not restricted in this respect; also, the number of discourse units connected by one relation is not restricted. the rst-dt uses a set of 78 rhetorical relations for the annotation (carlson and marcu, 2001) and does not provide information about the connective devices. relations are categorized according to nuclearity status (i.e. the presence of relevant relational content in one or in more discourse units for a given relation). the annotation captures the global structure of each text in the form of a tree diagram with elementary discourse units as its leaves and rhetorical relations as its edges. 2.3 rst signalling corpus the rst signalling corpus (rst-sc, das et al. (2015)) adds an annotation of signalling information to each of the rhetorical relations in the rst discourse treebank. overall, the corpus contains signals for 21 400 relations (das and taboada, 2015). for each relation, more than one signal can be present. the taxonomy of signals, based on features identified in previous studies and preliminary corpus work (taboada and das, 2013), is hierarchically organized in three levels: signal class, signal type and specific signal. three values for the signal class are single, combined and unsure, meaning that for a given relation either a single signal was found, or the signalling is combined from one independent and one dependent signal, or no signal was found. the class single is divided into nine types11. a relation in a given context can be signalled by: 1. a discourse marker an expression like because, and, now 2. reference features personal, demonstrative or comparative reference 3. lexical features an indicative word or phrase expressing the relation such as compared with, explaining, that means, the result is that 4. semantic features words (not pronouns) or phrases in a mutual semantic relation present in both/several discourse units of the relation; the semantic relation can be strong (synonymy, antonymy, meronymy, repetition, word pairs like asked – replied, general words like matter or thing referring to the previous context) or relatively weak (i.e. a lexical chain like selling shares – credit – concern – company – holding, confuses – clearer) 5. morphological features tense, aspect or mood change between units 6. syntactic features a relative, infinitival, participial or imperative clause, interrupted matrix clause, a parallel syntactic construction, reported speech, inversion of subject auxiliary, nominal or adjectival modifier 11. in what follows, specific signals for each type are presented only broadly, not in detail. for a detailed description with examples see the rst-sc annotation manual (das and taboada, 2015). 230 signalling implicit relations: a pdtb – rst comparison 7. graphical features colon, semicolon, dash, parentheses, items in sequence 8. genre features inverted pyramid scheme, newspaper layout, newspaper style attribution or newspaper style definition 9. numerical features the number of certain objects is represented by a word in one span (e.g. three, five) and these entities are then named in the other span the class combined comprises neither all the possible combinations of types from the class single nor combinations of all specific signals. the classification is data-driven and contains the following six types: 1. reference + syntactic features combine a reference feature and a subject noun phrase (np) 2. semantic + syntactic features combine all semantic features and a subject np 3. lexical + syntactic features combine an indicative word and a present participial clause 4. syntactic + semantic features combine parallel syntactic constructions and lexical chains 5. syntactic + positional features combine a participial clause and a beginning position of this clause 6. graphical + syntactic features combine a comma and a participial clause for the intended analysis, it is important to note the difference between combined signals and multiple signals in the rst-sc annotation scheme. combined signals have two parts – one of them is an independent signal, the other depends on the first one (e.g. in a combined reference + syntactic signal, the second signal, the subject np, “is used to specify additional attributes of the first signal” (das and taboada, 2015, p. 9)). on the other hand, multiple signalling refers to the possibility of a relation to be signalled by more than one separate signal (single or combined) functioning independently from each other. the class unsure “refers to those cases in which no potential signals were found or were specified” (das and taboada, 2015, p. 33). the distribution of all signals in the rst signalling corpus is presented in table 3 below in section 4, together with a comparison to signal distributions only for the subset of relations that have pdtb implicit counterparts. as far as we know, the general distribution was not presented by the authors of the corpus (their distributions are related to individual relations only (das, 2014)), that is why we present the results of our own measurement. 3. the method the methodological procedure for the comparative corpus analysis consisted of several steps: first, all three corpora were manually inspected in order to get a basic idea about the properties of the data to be compared. for the rst-dt and the rst-sc we used their respective annotation / visualization tools: the rst tool 3.45 (o’donnell, 2000) and the uam corpus tool 3.2i (o’donnell, 2008). for the pdtb annotation, we used exports from the column-transformed data representation and a recently developed pml-based extension of tred, a prague tool for treebank viewing and searching 231 poláková, mírovský and synková (mírovský et al., 2016), see also section 3.1. in this way, a sample of six wsj documents was selected according to their different genre characteristics (webber, 2009) and the number of implicit pdtb relations (129 sentences, 51 implicit relations). a preparatory survey was conducted by hand (18 matches with rst relations detected). simultaneously, the data annotated in all corpora (section 3.1) were converted into a common working format and a procedure for automatic argument mapping was designed (sections 3.2 and 3.3). the manual survey served as a check on the accuracy of the mapping procedure. next, label alignment was performed on the basis of hand-crafted sense alignment principles (section 3.4). finally, the intended comparative analysis could be carried out on the matching subset of data (section 4). 3.1 data the pdtb consists of 2 159 files (documents) with 48 338 sentences (the average number of sentences per document being 22.4) and contains 16 053 annotated implicit relations (see footnote 10). 364 out of the 2 159 files (i.e. 16.9%) were also annotated in the rst-dt with rst trees and rst relation types, and in the rst-sc with information on signals12. by number of sentences (8 532 out of 48 338), it represents 17.7% of the whole pdtb (the average number of sentences per rst document13 being 23.4). these 364 documents represent the data we used for the present study. 3.2 data conversion for unification of the data of the pdtb and the rst-sc, we used a framework for treebank annotation and data processing that consists of three core components: 1. the prague markup language (pml)14, an abstract xml format designed for annotation of linguistic corpora, especially treebanks, 2. tred, a highly customizable tree editor (pajas and štěpánek, 2008)15, which can be used to browse and edit data in the pml format, and 3. the pml-tree query (pml-tq), a powerful system for querying any data encoded in the pml format (pajas and štěpánek, 2009)16. once a treebank is transformed into the pml format, it can be browsed and edited in tred, and searched using the pml-tq – see for example a transformation of the penn discourse treebank to the pml (mírovský et al., 2016), or a project of harmonizing 36 treebanks into a common data format and annotation scheme hamledt (zeman et al., 2014). for our task, we needed to combine (i) information from the rst-sc (which includes also the original annotations of the rst-dt tree structures and relation types) and (ii) the annotation of implicit relations from the pdtb, in a single data representation. we proceeded in two steps: first, the data of the rst-sc were transformed to the pml: the rst tree structures along with the relation types were transformed from the penn bracketing format, and – at the same time 12. the whole rst-dt (as well as the rst-sc) consists of 385 wall street journal articles; 21 of these 385 files were not annotated in the pdtb. 13. counted on the 364 files annotated both in the pdtb and rst-sc 14. for an on-line documentation for the prague markup language, see http://ufal.mff.cuni.cz/jazz/pml/. 15. tred is written in perl and can be easily customized for a desired purpose by extensions that are included into the system as modules. it is available from https://ufal.mff.cuni.cz/tred/ under the gnu general public licence. 16. see e.g. chapter 8 of zikánová et al. (2015) for a practical introduction to the query language. 232 signalling implicit relations: a pdtb – rst comparison relations count implicit relations in the pdtb 16 053 (see footnote 10) implicit relations in the part of the pdtb that was annotated also in the rst-sc (364 files) 2 892 (18% of all implicit relations in the pdtb) imported implicit relations (both arguments match with nodes in the rst trees) 2 286 (79% of all implicit relations in the 364 files) without a counterpart rst relation (arguments are not siblings in the rst tree) 1 499 (65.6% of all imported relations) with a counterpart rst relation (arguments are siblings in the rst tree) 787 (34.4% of all imported relations, 27.2% of all implicit relations in the 364 files) imported implicit relations with a match of the counterpart rst relation (see section 3.4) 472 (60% of imported relations with a counterpart rst relation, 20.6% of all imported relations, 16.3% of all implicit relations in the 364 files) table 1: overview of implicit relations throughout the transformation process – the information on signals taken from the xml files was mapped into the tree structures using references to positions in the penn bracketing format. second, the implicit relations from the pdtb were mapped onto the transformed data, whenever the arguments of an implicit relation could be matched with node spans in the rst trees. matching of the arguments was performed via comparing rawtext representations of the arguments, as they were given in the two source formats (the penn bracketing format for the rst-sc, and the column-transformed annotation of the pdtb data). systematic inconsistencies between rawtext representations of arguments in the two sources were taken into account: before comparing the arguments, we removed paragraph separation marks, parentheses, and leading and trailing punctuation. table 1 gives an overview of numbers of implicit relations in the pdtb and in the rst-sc/pdtb overlapping data. as our study only focuses on implicit phenomena, the explicit pdtb relations and altlexes and their arguments were not mapped in this stage. apart from the data transformation itself, an extension to tred for displaying the transformed data was implemented17. figure 2 displays a part of the rst tree structure for the text from example 1 in tred (see also figure 1 for its depiction in the original rst tool). in our implementation, the rst tree structure and relation types are displayed along with signals for the rst relations and matching pdtb implicit discourse relations. red arrows/polylines represent rst relations; orange arrows represent implicit discourse relations from the pdtb. for the time being (and for technical reasons), the information about any relation is at the start node of the relation; the type of the rst relation is depicted in red, signals are depicted in magenta, and comments to signals in brown. semantic types (senses) of the pdtb implicit relations and inserted connectives are in orange. 17. the extension is freely downloadable from inside tred using its extensions management tool. scripts for transforming the original rst-dt, rst-sc and pdtb data into the pml are a part of the extension; for details, see the dedicated web page: https://ufal.mff.cuni.cz/rst. 233 poláková, mírovský and synková nucleus 29 several organizations, nucleus 29 30 satellite 30 including the industrial biotechnical association and the pharmaceutical manufacturers association, nucleus 29 31 nucleus 31 have asked the white house and justice department to name candidates with both patent and scientific backgrounds. nucleus 29 32 satellite 32 the associations would like the court to include between three and six judges with specialized training. same-unit elaboration-set-member-e signal;single;syntactic;nominal_modifier no_comment signal;single;lexical;indicative_word including same-unit signal;single;syntactic;interrupted_matrix_clause no_comment elaboration-additional signal;combined;(semantic+syntactic);(repetition+subject_np) associations signal;single;semantic;lexical_chain a few lexical chains expansion.restatement.specification specifically figure 2: joint annotation in tred for the sentences in example 1 (see also figure 1) 3.3 argument mapping as already mentioned in section 2, the different theoretical assumptions behind the two corpora lead to different strategies in the delimitation of discourse units (discourse arguments). moreover, arguments of implicit relations in the pdtb 2.0 differ slightly from arguments of explicit relations: arguments of implicit relations were partly predefined by the principle to annotate an implicit relation between adjacent sentences within a paragraph (an argument being represented by either the whole sentence or its part)18. 18. the adjacency rule for the annotation of arguments of implicit relations was loosened for the later annotation of english biomedical texts in biodrb (prasad et al., 2011) by allowing the annotators to search also for a remote left argument of an implicit connective. the authors were able to reduce the percentage of norels, i.e. of cases where no relation to the immediately preceding sentence could be found, from 1.15% in the pdtb to 0.9% in the biodrb (prasad et al., 2014, p. 924). 234 signalling implicit relations: a pdtb – rst comparison nucleus 14 "overall demand still is very respectable," nucleus 14 15 satellite 15 says christopher c. cole, group vice president at cincinnati milacron inc., the nation's largest machine tool producer. satellite 14 16 nucleus 16 "the outlook is positive for the intermediate to long term." list attribution signal;single;syntactic;reported_speech no_comment list signal;single;semantic;lexical_chain respectable ~ positive contingency.cause.result thus figure 3: tred representation of example 2: a pdtb implicit relation with arguments matching to rst discourse units but without a direct rst relation counterpart despite these different segmentation strategies, it was assumed that an intersection of pdtb discourse arguments exactly matching rst discourse units does exist in the wsj data, only its size was difficult to estimate in advance. the mapping procedure of the pml-converted data, as described from the technical viewpoint in section 3.1 above, revealed that this intersection comprises 4 081 arguments, or 2 286 implicit pdtb relations with both arguments exactly matching the rst units (which represents 79% of all implicit relations in the 364 files with both annotations, see table 1). comparing the pdtb and rst relations requires not only to detect the location and extent of the arguments: the 2 268 pdtb relations with both arguments successfully mapped also include cases where there was no corresponding rst relation found between these two arguments (i.e. the corresponding rst nodes were not siblings). an instance of such a pdtb implicit relation without an rst counterpart is given as example 2 and visualized in figure 3 – the orange arrow (a pdtb relation) connects two rst nodes that do not relate by an rst relation. these pdtb relations have been excluded from the studied subset, resulting in 787 pdtb–rst relation pairs (implicit pdtb relations with an rst relation counterpart). 235 poláková, mírovský and synková pdtb implicit relation rst relation count expansion.conjunction elaboration-additional 85 expansion.conjunction list 80 expansion.list list 60 expansion.restatement.specification elaboration-additional 36 contingency.cause.reason explanation-argumentative 26 comparison.contrast elaboration-additional 20 contingency.cause.reason elaboration-additional 17 expansion.restatement.specification elaboration-general-specific 17 contingency.cause.result elaboration-additional 15 expansion.instantiation example 13 comparison.contrast.juxtaposition contrast 11 comparison.contrast contrast 10 expansion.restatement.specification explanation-argumentative 10 contingency.cause.result consequence-s 9 expansion.conjunction comparison 9 ... table 2: fifteen most frequent correspondence pairs of implicit pdtb and rst relations (787 relations in total); (substantial) semantic mismatches are highlighted in grey (2) [“overall demand still is very respectable,”] [says christopher c. cole, group vice president at cincinnati milacron inc., the nation’s largest machine tool producer.] [“the outlook is positive for the intermediate to long term.”] the results of the argument mapping demonstrate two basic tendencies: if an rst discourse unit matches exactly with a pdtb argument of an implicit relation, the pdtb argument typically consists of more rst elementary discourse units (edus) within one subtree. these edus mostly correspond to clauses and clause-like syntactic structures within one sentence. on the contrary, non-matching pdtb arguments consist of rst nodes that do not form a subtree. further, where the discourse units of two similar relations do not match, it is mainly due to the exclusion of a part of the sentence from the pdtb argument – often due to the exclusion of attribution spans (e.g. he said) from the pdtb arguments, compare again example 2 and figure 3. these near-matches are not included our study; at this point, our aim was to reliably find exactly matching relation pairs, without further intervention. in a less strict approach, some relations with attribution spans could be probably considered matching counterparts if the attribution spans were detected in the data, analyzed separately and removed or disregarded (with possible manual work). this would lead, in our opinion, to an enlargement of the intended dataset. 3.4 semantic labels the research in this article does not focus on the correspondence and mapping of labels in the two annotation frameworks – this is current work of e.g. sanders et al. (submitted). moreover, for the 236 signalling implicit relations: a pdtb – rst comparison given two categorizations of discourse (rhetorical) relations in the pdtb 2.0 and the rst-dt, such a mapping is a difficult task, (not only) since there is a big difference in the granularity of the two taxonomies. the pdtb 2.0 uses a three-level hierarchy of senses with 30 possible end-level values (prasad et al., 2008b) whereas the rst-dt annotation distinguishes 78 types of relations in 16 classes (carlson et al., 2001). nevertheless, for the purposes of this article, it appeared necessary to provide at least partial correspondence links for the labels on both sides. the links are partial in the sense of mapping only those relations that actually occurred in our dataset (not the whole taxonomies)19, and also in the sense of capturing the agreement only on a reasonably general semantic level. for the analysis of signals, there has to be agreement in the two corpora that a given signal is actually a signal of one particular, even if coarse-grained, category. let us demonstrate this situation on example 1 above: it can be observed that the pdtb label expansion–restatement–specification and the rst label elaboration–additional share basic semantic features and thus can be treated as corresponding on a general semantic level: they are both additive, they expand the piece of information given in a (a demand for candidates for judges with specific skills) by adding b (how many judges are required). on the other hand, the subtle semantics of specifying or giving a detail is not encoded in the rst elaboration– additional relation. there can be a better fit, the elaboration-general-specific relation. in principle, types of relations that agreed within one of the four first-level pdtb-defined classes were assessed as corresponding, with few exceptions: within the expansion-like group, substantially different relations were not matched, e.g. pdtb expansion–conjunction and rst example; within the contingency-like group, conditionals were not treated as equal to causals; and so on. in other words, we tried to keep the matching strict. only in such a way can we rely on the basic assumption required for our study: in the cases of relations with corresponding semantic categories, we can consider the signals for the rst relations to be signals for the pdtb implicit relations as well. according to this method, from the 787 relation pairs we had at our disposal based on the argument mapping, 60% (472) relation pairs have similar semantic characteristics and can be further worked with. these 472 relation pairs have altogether 674 signals, i.e. 1.43 signals per relation20. table 2 shows the fifteen most frequent correspondence pairs of implicit pdtb relations and rst relations. rows highlighted in grey mark pairs of relations whose semantic labeling could not be treated as matching. 4. the comparison in the comparative analysis itself, we focus on two main ways of comparison. first, we analyze the distribution of all signal types in the whole studied subset of matching implicit pdtb relations against the distribution of these signal types in the whole rst signalling corpus. the matching dataset is quite small compared to the sizes of both source corpora; nevertheless – without aspiring on generalizing our results on the whole treebanks – it allows us to observe linguistically relevant and interesting tendencies for the studied phenomenon (4.1). a more fine-grained analysis concerns the semantic type of signal, since it proved to be the most frequent type of signalling in the studied dataset (section 4.1.1). second, we analyze signal type distributions with regard to different pdtb 19. see again footnote 9 in section 2.1. 20. in fact, it was only 471 relations, as one of them (between text spans 164 and 163 in the file wsj_0681) was not annotated with any signal in the rst-sc. 237 poláková, mírovský and synková single rst ∩ combined rst ∩ unsure rst ∩ syntactic 29.8 1.3 semantic+syntactic 7.4 13.5 unsure 5.3 12.2 semantic 24.8 51.2 reference+syntactic 1.9 5.5 discourse marker 13.3 1.3 syntactic+semantic 1.4 5.5 lexical 4.9 2.8 graphical+syntactic 0.7 0 graphical 3.5 1.8 lexical+syntactic 0.4 0 genre 3.2 0 syntactic+positional 0.2 0 reference 2 4 morphological 1.1 0.6 numerical 0.1 0.3 table 3: distribution of types of signals in the whole rst signalling corpus (the rst column) and in the matched 472 relations (674 signals) corresponding to implicit pdtb relations (the ∩ column); all values are in percents. senses (section 4.2). as there are not enough occurrences for all pdtb relation sense (sub)types in our dataset, we concentrate on the most frequent ones. 4.1 overall signal distributions table 3 presents percentages for occurrences of all signal types in the whole rst signalling corpus (385 documents, 21 400 relations, 29 300 signals) and further percentages for a subset of rst relations that have implicit pdtb counterparts agreeing in argument spans and in the (basic) semantic characteristics (472 relations, 674 signals in total). considering first only the general rst-sc signal distribution, it can be observed that more than a two-third majority of signals (approx. 68%) is represented by only three categories of signalling: syntactic (29.8%), semantic (24.8%) and discourse markers (13.3%). discourse markers, although this category is in general perceived broader than the category of discourse connectives, are much less frequent in the rst-sc than are explicit connectives in the pdtb (the proportion of explicit relations in the pdtb is 45.3%, cf. section 2.1 above)21. in total, 201 types of rst discourse markers were identified (the “type” here refers to a unique string, so e.g. if and only if are counted as two types). the combination of semantic and syntactic signals, i.e. one of the semantic signal subtypes (repetition, lexical chain, synonymy, meronymy or general word) in combination with a subject np, is the fourth most common signal (7.4%), followed by a 5.3% of unsure cases (signals not found or non-specified). the proportion of no other signal type (nor their combination) exceeds 5%. proportions of the signal types in the subset of rst relations corresponding to implicit pdtb relations show some substantial differences. as expected, the proportion of discourse markers drops dramatically (to 1.3%). a manual inspection of discourse markers in this subset indicates that these 21. we are aware that a direct comparison is not possible here and that this observation is only very rough: first, the rstsc represents approximately one sixth of the pdtb size; second, the annotation approaches to discourse connectives and discourse markers differ significantly; and third and most importantly, building a hierarchical tree structure for the whole document results in the existence of many more relations per word. 238 signalling implicit relations: a pdtb – rst comparison markers are distinct22 from the pdtb connectives, for example expressions like now, particularly, most importantly, unfortunately or after all. in this way we can confirm a basic assumption that implicit pdtb relations generally do not correspond to relations with discourse markers in the rst-sc. if we take a closer look at the lexical signal type (indicative words, alternate expressions), we observe that the concept of pdtb altlexes and that of a lexical signal type in the rst-sc include similar items – e.g. compared with, followed by, including, reason, explain(ing), still, like, finally, but also longer strings like no matter how much or in the past two weeks etc. since the pdtb annotates altlexes only under specific conditions and thus many such expressions and phrases were not marked at all (see section 2.1), it can be assumed that this category would be more numerous in the pdtb. this is not so much the case in the rst-sc: the proportion of the lexical signal type in the whole rst-sc is only 4.9%. in our subset of ptdb – rst matching relation pairs, the proportion slightly drops (to 2.8%). single lexical signals in the studied subset comprise altogether 19 indicative words, among which there are 4 occurrences of the expression even, (as in example 3); the remaining lexical signals occur only once each. in the pdtb, even is treated as a connective modifier (prasad et al., 2007, p. 9), not as a separate connective. as a result, in example 3, the pdtb annotator did not annotate an explicit even-relation or an altlex but instead inserted an implicit also-connective (and the label expansion–conjunction). the remaining three occurrences are analogical: the even expression23 does not modify any other connective-like expression, but has a scope over other parts of the sentence. (3) [the pilot program was received well (by teachers and students), but there wasn’t reason enough to sign up.] [we even invited the public to stop by and see the program, but there wasn’t much interest.] in our opinion, this example demonstrates a well known issue of delimitation of the category of discourse connectives, discourse markers and other strong lexical cues of discourse connections, may they be called alternative lexicalizations, lexical signals or secondary connectives (as in rysová and rysová, 2014). the most apparent distribution change between the two datasets is represented by the drop in the most frequent signal type – the syntactic signal type (from almost 30% to 1.3%). in the implicit subset, there are only 9 cases, all of them represented by parallel syntactic constructions, cf. example 4. (4) “how the hell can you live with yourself?” he erupts at a politician. [“you twist people’s trust.] [you built your career on prejudice and hate.] signal: single; syntactic; parallel_syntactic_constructions; comment: you twist ~ you built 22. apart from one annotation error 23. a focusing particle or a rhematizer in the prague approach to information structure and a possible discourse connective in the analysis of discourse relations 239 poláková, mírovský and synková the very low proportion of syntactic signals in the studied subset is most likely caused by the fact that the pdtb implicit relations are in the vast majority realized inter-sententially, whereas most of the rst syntactic signal subtypes apply only intra-sententially. the parallelism of syntactic constructions, as demonstrated in example 4, is one of the few possible syntactic signals applying also between individual sentences. it appears that syntactic signals, on their own the strongest signalling cue in the rst-sc, can play only a restricted role as sole signals of implicit pdtb relations. on the other hand, the proportions of three types of combined signals with a syntactic component shows a visible increase in the studied subset (first three cells in the “combined” column in table 3): semantic + syntactic from 7.4 to 13.5%, reference + syntactic from 1.9 to 5.5% and syntactic + semantic from 1.4 to 5.5%. these signals are represented mostly by the following subtypes, respectively: lexical chain + subject np and repetition + subject np, personal reference + subject np, parallel synt. constructions + lexical chain. in these combinations, again, only the parallel syntactic constructions syntactic subtype applies; the subject np component does not function as an independent signal, but is always dependent on the previous component in a combination (section 2.3). quite in the opposite direction from syntactic signalling goes the semantic type of signalling, it increases by 26.4% to more than a half (51.2%) of all signals in our subset. this fact, in our opinion, is the most expressive evidence about the nature of signalling of the pdtb implicit relations. that is why we analyze the semantic type of signals individually, in section 4.1.1 below. finally, a great distribution difference can be detected in the unsure category, it increases from 5.3% to 12.2% in the studied subset. this is also the only tag that never co-occurs with other types of signals, which means that the number of unsure tags in our subset (82) is also the number of relations with this (and no other) labeling. 4.1.1 semantic signals, lexical chains according to the rst-sc annotation manual, a semantic signal, unlike most other signals, “has two components (words or phrases), each belonging to one of the spans. the components are in a semantic relationship with each other, such as synonymy, antonymy, and lexical chain...” (das and taboada, 2015, p. 20). it has six subtypes: synonymy (e.g. scandinavian airlines system ~ sas), antonymy (profit ~ loss), meronymy, i.e. set–member relation (computer firms ~ microsoft corp.), repetition (avis ~ avis), indicative word pair, i.e. very closely semantically related words or phrases (asked ~ replied; resigned ~ succeeded) and lexical chain (personal computers ~ desk-top computers, microprocessors, minicomputers, mainframes). a lexical chain is defined as follows: “words or phrases in the respective spans are identical or semantically related. unlike the repetition feature, words or phrases in lexical chains do not refer to object or entity, but they belong to the class of indefinite or common nouns and also other syntactic categories, such as adjectives, verbs and adverbs. lexical chain differs from synonymy, antonymy, meronymy and indicative word pair in a significant way. while in the latter categories, the semantic relation between the words or phrases is very strong and can be specified, words or phrases present in a lexical chain are related to each other by a relatively weak semantic connection.” (das and taboada, 2015, p. 21) there are altogether 473 semantic signal occurrences in the studied dataset (single and in combination). 161 are single.semantic signals appearing as a sole signal of a relation, i.e. the relation does not have combined or multiple signalling24. interestingly, all these semantic signals are of the sub24. especially here, note the difference between multiple and combined signals, as explained above in section 2.3. 240 signalling implicit relations: a pdtb – rst comparison type lexical chain; that means that none of the remaining 5 categories function as a sole signal of an implicit pdtb relation. the remaining 312 cases with semantic signals are represented by 60 other combinations of (single, combined or both) signals, with most of the combinations not exceeding 1% of the occurrences of all cases with semantic signals25. although it is difficult to further analyze such a variety of cases, one tendency is traceable: also here, with only one exception, at least one signal from each of the 60 combinations is represented by the lexical chain subtype. we have therefore further briefly analyzed the comments for the lexical chain subtype. there are altogether 408 occurrences of this subtype in our dataset (or 86% of all semantic signals). the comments for lexical chain typically specify the exact wording of the lexical chains in question, cf. example 5 annotated with the rst label explanation-argumentative and the following signal: (5) [excluding the gain, p&g’s earnings were close to analysts’ predictions of about $1.40 a share for the quarter.] [wall street had expected a modest rise in the company’s domestic sales and earnings, and more substantial increases in overseas results.] signal: single; semantic; lexical_chain, comment: (close to analysts’ predictions) ~ (expected a modest rise) on the other hand, in 247 cases the annotators only indicated that there are more lexical chains, cf. example 1 above, or they did not provide any comment at all (40 cases). sometimes, the comment included also a note that the lexical chain is a rather loose one. if we sum up our observations so far and relate them to the number of pdtb relations in the studied subset (472), the following can be stated: 161 relations are signalled by a single semantic signal of the subtype lexical chain, which is a rather weak semantic connection. a further 82 relations are assigned the unsure tag; one relation has no signal assigned. that implies that more than half (exactly 51.7%) of all relations in our intersection are signalled weakly or their signals are difficult to detect. 4.2 signals across implicit relations there are only 17 distinct pdtb relation types (or 3rd-level subtypes) in the studied subset, and their distribution is quite uneven. for instance, none of the temporal class relations exceeds 10 occurrences and, on the contrary, we can observe a clear prevalence of relations from the expansion class (355 out of 472, or 75%). this is partly influenced by the nature of the studied data and would be different for other genres and registers, and partly, as we believe, by the implicit nature of the studied relations. because of the low number of instances we present results only for eight most frequent relation (sub)types from three class level categories (expansion, contingency and comparison). for each such relation, we present two most common signals. the results are summarized in table 426. the signal distribution across the (relatively) frequent relations confirms at first sight our findings for the overall signal distribution: the leading signal for five of these relations is single; semantic or 25. they appear either as combined signals (semantic + other type) or as multiple, independent signals of one relation (single semantic + single semantic; single semantic + combined; more than two signal types etc.). 26. as an attachment to this article, a larger table with full signal distributions across the mapped pdtb relations can be found at https://ufal.mff.cuni.cz/rst. 241 poláková, mírovský and synková lexical_chain, followed by unsure as the leading signal for the remaining three relations. if we cluster the relations according to their general semantic class, it appears that expansion class relations are signalled more by lexical chains whereas the relations from contingency and comparison classes were more likely to have unsure signalling. altogether, no other striking differences in signal distributions can be observed for individual relation types. the semantic weakness of the semantic.lexical_chain signal subtype and the frequency of the unsure signals go hand in hand with other facts about semantic weakness from the pdtb annotation: (i) expansion–conjunction is known to be a relatively broadly and loosely defined relation in the pdtb 2.0 annotation (as opposed to the restriction of its definition for the upcoming pdtb 3.0 relation taxonomy, cf. prasad et al. (in preparation)) and (ii) there is no subtype-level label annotated for most contrast relations in our subset27, see table 4. all these partial observations suggest that the relations in the studied subset, regardless of how we defined them, were in a large number difficult to treat for the annotators of both systems. in other words, despite the apparatus of a complex and detailed data-driven signal taxonomy, the cues leading to the interpretation of these particular relations were weak and could not be identified precisely. 5. discussion the signalling annotation of the rst-sc has a high information value for the presented analysis and possibly also for the purposes of discourse parsing and other automatic text processing tasks. it provides direct access to signals of discourse coherence other than the most apparent cues in the connections regarded as implicit. we are now able to study discourse functions of expressions such as more or differently, quantify the role of syntactic features and coreference (at least for our limited sample of texts), etc. we can also look at possible combinations of various types of signals. and, we also learn where the presented types of discourse annotation have their limits: it is where semantics comes into play. from the perspective of prague multilayer-annotated treebanks, it can be observed that semantic signals in the rst-sc terminology partly correspond to prague bridging anaphora annotations. on one hand, we are convinced that a finer subclassification of semantic signals is possible – compare the types of bridging links and further proposed subtypes in chapter 4 in zikánová et al. (2015). on the other hand, from its nature, a semantic signal is in principle not a surface signal, but requires the inclusion of inferential processes. it is only natural that a high fraction of the semantic signals in implicit relations, namely those annotated as lexical chains, are very difficult to assess, and for the annotators (or any two interpretators) to agree on. also the prague annotation of bridging anaphora, although further classified and provided with extensive annotation guidelines, has the lowest agreement figures from all the annotated phenomena annotated in the prague dependency treebank (zikánová et al., 2015, p.96) (compare also the discussion on p. 60–62 on co-hyponymy, definition of common world knowledge and the necessity of inclusion of word-net-like databases in similar tasks). our study leads to an indirect observation that the two discourse annotation frameworks are consistent with each other in pointing out places with weak signalling or no signalling of coherence at all. to put it simply, implicit relations are indeed implicit, not signalled, or signalled vaguely. we believe that at this point of linguistic description, with such a high degree of semantically anchored 27. there was the option for the pdtb annotators to provide only a higher-level sense label for a relation if they could not decide about the lower-level label. 242 signalling implicit relations: a pdtb – rst comparison pdtb implicit relation count signals % expansion.conjunction 165 signal; single; semantic; lexical_chain 46.6 signal; combined; (semantic+syntactic); (lexical_chain+subject_np) 13 expansion.restatement.specification 82 signal; single; semantic; lexical_chain 43.6 signal; combined; (semantic+syntactic); (lexical_chain+subject_np) 12.1 expansion.list 63 signal; single; semantic; lexical_chain 59.2 signal;combined;(syntactic+semantic); (parallel_synt._constr.+lexical_chain) 18.3 contingency.cause.reason 34 signal; unsure 50 signal; single; semantic; lexical_chain 47.1 contingency.cause.result 37 signal; unsure 40.5 signal; single; semantic; lexical_chain 37.8 expansion.instantiation 33 signal; single; semantic; lexical_chain 54 signal; unsure 12 comparison.contrast 19 signal; unsure 47.6 signal; single; semantic; lexical_chain 28.6 comparison.contrast.juxtaposition 11 signal; single; semantic; lexical_chain 57.1 signal; unsure 14.3 table 4: eight most frequent pdtb implicit relations in the studied subset (472 relations) and two most frequent signals for each (with relative frequencies within signals for the given implicit relation) coherence relations, we have reached the borderline of what information corpus annotation can provide. trying to annotate semantic phenomena beyond this point, in our opinion, only results in getting unreliable and inconsistent data with very limited use for nlp purposes. the only possible way for us is to accept a certain degree of vagueness and perhaps to learn to detect places in texts with weak coherence or vagueness instead of imposing a certain type of connective or even discourse meaning on them. from another perspective, the interpretation of implicit connections can be further studied using the manual czech translations of the wsj-part of the pdtb collected in the prague czech–english dependency treebank (pcedt 2.0, hajič et al. (2011)). it offers an ideal opportunity to look for possible explicitations of implicit relations by czech connectives or other surface cues. 6. conclusions the study presented in this article made use of a portion of the english wall street journal texts having been annotated for discourse phenomena from two different theoretical perspectives, namely the penn discourse treebank 2.0 and the rst signalling corpus. despite the theoretical and annotation differences (global vs. local approach to text analysis, different segmentation strategies, different 243 poláková, mírovský and synková sets of labels for coherence relations etc.), we have shown that there is a common denominator for a comparison of the two annotations. we have been interested in answering the question of how implicit pdtb relations are signalled by means of the rst-sc signal annotation. we have arrived at several observations we believe can be of use for any discourse researcher concerned with implicit discourse phenomena or comparing discourse annotation schemes in general. (i) a matching intersection of the annotations, in terms of a discourse (rhetorical) relation and the two units it relates, does exist in the compared datasets. we took into account relation pairs with at least broadly corresponding semantic labeling but we aimed for exact argument matching; the resulting intersection is therefore not very large. it comprises 472 implicit pdtb relations (out of 2 892 implicit relations in the part of the pdtb also annotated in the rst-sc) signalled by 674 rst-sc signals28. that is why we do not aspire to generalize our results for the whole treebanks. nevertheless, we believe our results allow us to observe relevant and interesting tendencies for the studied phenomena. (ii) the comparative analysis showed that to a large extent (51.2%), implicit pdtb relations are signalled by signals of semantic nature; these signals are anchored in the semantics of specific lexical chains in the arguments. lexical chains are characterized as a rather weak type of semantic connection by the authors of the rst-sc themselves. these lexical chains are either specified word for word in the comments in the annotations, or, in many cases, the annotators indicated that there can be more lexical chains taking part in interpreting the relation, and did not specify them. further, in 19%, semantic signals appeared in combination with syntactic types of signals (subject np, parallel syntactic constructions). in 12.2%, the signalling of implicit pdtb relations was unsure, compared to the 5.3% of unsure cases in the whole rst-sc. if we take the number of pdtb implicit relations in the studied dataset into account, slightly more than a half of them (51.7%) are signalled either by a sole signal of the type single;semantic;lexical chain or are unsure. these observations indicate that annotation of implicit relations, at least in the studied subset, cannot be easily solved by looking for overtly present signals. their nature is to a large extent semantic and, moreover, often outside the scope of well definable semantic categories (synonymy, antonymy, set–member relationship, etc.). our analysis therefore seems to confirm the fact that annotation/classification of implicit relations is a challenging task both for humans and for automatic methods in nlp applications. (iii) last but not least, for the purposes of this study the data of the two treebanks were transformed to a common format (pml) and a unified visualization in the tred tool was developed. one of its extensions, the pml-tree query, also enables for the first time to search the imported treebanks for various phenomena at once. acknowledgements this work has been supported by the projects ga cr 17-06123s of the czech science foundation and multilingual corpus annotation as a support for language technologies (lh14011) of the ministry of education, youth and sports of the czech republic. the work has been using language resources and tools developed and/or stored by the lindat/clarin project of the ministry of education, youth and sports of the czech republic (lm2015071). we would like to thank the two anonymous reviewers for their helpful suggestions. 28. the intersection would be probably larger if different treatment of attribution was taken care of, and, naturally, it would be larger if also explicit pdtb relations were taken into account. 244 signalling implicit relations: a pdtb – rst comparison references nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. eduard bejček, eva hajičová, jan hajič, pavlína jínová, václava kettnerová, veronika kolářová, marie mikulová, jiří mírovský, anna nedoluzhko, jarmila panevová, lucie poláková, magda ševčíková, jan štěpánek, and šárka zikánová. prague dependency treebank 3.0. univerzita karlova v praze, mff, úfal, prague, czech republic, 2013. lynn carlson and daniel marcu. discourse tagging reference manual. technical report isitr-545, los angeles, 2001. available from . lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in proceedings of the sigdial workshop on discourse and dialogue, pages 1–10, 2001. lynn carlson, mary ellen okurowski, daniel marcu, et al. rst discourse treebank. linguistic data consortium, university of pennsylvania, philadelphia, 2002. debopam das. signalling of coherence relations in discourse. phd thesis, simon fraser university, burnaby, canada, 2014. debopam das and maite taboada. rst signalling corpus annotation manual. linguistic data consortium, university of pennsylvania, philadelphia, 2015. debopam das, maite taboada, and paul mcfetridge. rst signalling corpus. linguistic data consortium, university of pennsylvania, philadelphia, 2015. jan hajič, eva hajičová, jarmila panevová, petr sgall, silvie cinková, eva fučíková, marie mikulová, petr pajas, jan popelka, jiří semecký, jana šindlerová, jan štěpánek, josef toman, zdeňka urešová, and zdeněk žabokrtský. prague czech-english dependency treebank 2.0. institute of formal and applied linguistics, prague, czech republic, 2011. junyi jessy li and ani nenkova. reducing sparsity improves the recognition of implicit discourse relations. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 199–207, philadelphia, pa, usa, 2014. ziheng lin, min-yen kan, and hwee tou ng. recognizing implicit discourse relations in the penn discourse treebank. in proceedings of the conference on empirical methods in natural language processing, pages 343–351, 2009. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8:243–281, 1988. mitchell p. marcus, mary ann marcinkiewicz, and beatrice santorini. building a large annotated corpus of english: the penn treebank. computational linguistics, 19(2):313–330, 1993. 245 poláková, mírovský and synková eleni miltsakaki, rashmi prasad, aravind joshi, and bonnie webber. annotating discourse connectives and their arguments. proceedings of the hlt/naacl workshop on frontiers in corpus annotation, pages 9–16, 2004. jiří mírovský, lucie poláková, and jan štěpánek. searching in the penn discourse treebank using the pml-tree query. in proceedings of the international conference on language resources and evaluation, pages 1762–1769, 2016. michael o’donnell. rsttool 2.4: a markup tool for rhetorical structure theory. in proceedings of the international conference on natural language generation, pages 253–256, 2000. available from . michael o’donnell. the uam corpus tool: software for corpus annotation and exploration. in proceedings of the xxvi congreso de aesla, almeria, spain, pages 3–5, 2008. available from . petr pajas and jan štěpánek. recent advances in a feature-rich framework for treebank annotation. in proceedings of the international conference on computational linguistics, pages 673–680, 2008. petr pajas and jan štěpánek. system for querying syntactically annotated corpora. in proceedings of the joint conference of the annual meeting of the acl and the international joint conference on natural language processing, pages 33–36, 2009. joonsuk park and claire cardie. improving implicit discourse relation recognition through feature set optimization. in proceedings of the annual meeting of the special interest group on discourse and dialogue, pages 108–112, 2012. emily pitler, annie louis, and ani nenkova. automatic sense prediction for implicit discourse relations in text. in proceedings of the joint conference of the annual meeting of the acl and the international joint conference on natural language processing, pages 683–691, 2009. lucie poláková. discourse relations in czech. phd thesis, charles university in prague, prague, 2015. lucie poláková, jiří mírovský, anna nedoluzhko, pavlína jínová, šárka zikánová, and eva hajičová. introducing the prague discourse treebank 1.0. in proceedings of the international joint conference on natural language processing, pages 91–99, 2013. rashmi prasad and harry bunt. semantic relations in discourse: the current state of iso 24617-8. in proceedings of the joint acl-iso workshop on interoperable semantic annotation (isa-11), pages 80–92, 2015. rashmi prasad, eleni miltsakaki, nikhil dinesh, alan lee, aravind joshi, livio robaldo, and bonnie webber. the penn discourse treebank 2.0 annotation manual. philadelphia, 2007. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. penn discourse treebank version 2.0. linguistic data consortium, university of pennsylvania, philadelphia, 2008a. 246 signalling implicit relations: a pdtb – rst comparison rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of lrec 2008, pages 2961– 2968, 2008b. rashmi prasad, aravind joshi, and bonnie webber. realization of discourse relations by other means: alternative lexicalizations. in proceedings of the international conference on computational linguistics, pages 1023–1031, 2010. rashmi prasad, susan mcroy, nadya frid, aravind joshi, and hong yu. the biomedical discourse relation bank. bmc bioinformatics, 12:188–205, 2011. rashmi prasad, bonnie webber, and aravind joshi. reflections on the penn discourse treebank, comparable corpora, and complementary annotation. computational linguistics, 40:921–950, 2014. rashmi prasad, bonnie webber, alan lee, and aravind joshi. discourse relations in the pdtb 3.0. in preparation. charlotte roze, laurence danlos, and philippe muller. lexconn: a french lexicon of discourse connectives. discours. revue de linguistique, psycholinguistique et informatique, (10), 2012. doi: 10.4000/discours.8645. magdaléna rysová and kateřina rysová. the centre and periphery of discourse connectives. in proceedings of the pacific asia conference on language, information and computing, pages 452–459, 2014. ted sanders, vera demberg, jet hoek, merel scholman, fatemeh torabi asr, sandrine zufferey, and jacqueline evers-vermuel. unifying dimensions in discourse relations: how various annotation frameworks are related. corpus linguistics and linguistic theory, submitted. merel scholman, jacqueline evers-vermeul, and ted sanders. categories of coherence relations in discourse annotation: towards a reliable categorization of coherence relations. dialogue & discourse, 7(2):1–28, 2016. manfred stede and carla umbach. dimlex: a lexicon of discourse markers for text generation and understanding. in proceedings of the international conference on computational linguistics, pages 1238–1242, 1998. maite taboada and debopam das. annotation upon annotation: adding signalling information to a corpus of discourse relations. dialogue and discourse, 4:249–281, 2013. yannick versley and anna gastel. linguistic tests for discourse relations in the tüba-d/z corpus of written german. dialogue and discourse, 2:142–173, 2013. bonnie webber. genre distinctions for discourse in the penn treebank. in proceedings of the joint conference of the annual meeting of the acl and the international joint conference on natural language processing of the afnlp, pages 674–682, 2009. florian wolf and edward gibson. representing discourse coherence: a corpus-based study. computational linguistics, 31(2):249–287, 2005. 247 poláková, mírovský and synková daniel zeman, ondřej dušek, david mareček, martin popel, loganathan ramasamy, jan štěpánek, zdeněk žabokrtský, and jan hajič. hamledt: harmonized multi-language dependency treebank. language resources and evaluation, 48(4):601–637, 2014. deniz zeyrek and bonnie webber. a discourse resource for turkish: annotating discourse connectives in the metu corpus. in proceedings of the workshop on asian language resources at the international joint conference on natural language processing, pages 65–71, 2008. deniz zeyrek, işın demirşahin, a. b. s. çallı, and murathan kurfali. annotating implicit discourse relations in turkish: the challenge of corrective discourse relations. in abstracts of the international pragmatics association (ipra) conference, pages 455–456, 2015. šárka zikánová, eva hajičová, barbora hladká, pavlína jínová, jiří mírovský, anna nedoluzhko, lucie poláková, kateřina rysová, magdaléna rysová, and jan václ. discourse and coherence. from the sentence structure to relations in text. studies in computational and theoretical linguistics. úfal, praha, czechia, 2015. available from . 248 dialogue & discourse 14(1) (2023) 34–55 doi: 10.5210/dad.2023.102 attribution and the discourse structure of reports emar maier e.maier@rug.nl clcg & theoretical philosophy university of groningen editor: amir zeldes submitted 05/2022; accepted 03/2023; published online 04/2023 abstract i propose a discourse-level analysis of report constructions. indirect discourse, mixed and direct quotation, free indirect discourse, and attitude ascriptions are all analyzed in terms of a discourse relation of attribution, connecting two propositional discourse units corresponding to (i) a frame segment (he said, she dreamed) and a (possibly complex, multi-sentence) report (“i’m an idiot”, (that) she was president). i provide a unified semantics for the discourse relation of attribution that invokes a flexible notion of ‘characterization’. a discourse unit may characterize a speech event by reproducing its linguistic surface form (as in quotation) or its propositional content (as in indirect speech and attitude reports), or some mixture of both (as in mixed quotation or free indirect discourse). i formalize this unified discourse-level attribution approach to reporting within the general framework of sdrt, and apply it to direct, indirect, and free indirect reports that extend beyond the single embedded or quoted clause. the resulting account is the first to do justice to the complex internal dependencies within stretches of reported discourse. keywords: discourse structure, sdrt, attribution, coherence, quotation, reported speech, free indirect discourse 1. introduction: discourse, coherence, and reporting a correct interpretation of a multi-sentence discourse includes more information than is contained in the interpretations of its individual sentences taken in isolation. take the mini-discourse in (1). (1) john was biking home late. a police officer stopped him. she give him a fine. his lights were off. we naturally infer that a police officer stopped john while john was biking home late and then the police officer gave john a fine because john’s lights were off. the individual sentences themselves describe states and events, which we as interpreters try to combine into a coherent discourse by inferring various causal, temporal and other relations between these states and events (hobbs, 1979). these coherence inferences are generally defeasible and constrained by rationality, worldknowledge, a finite inventory of potential discourse relations (narration, background, elaboration, explanation, etc.), and linguistic cues (an overt connective like and then would signal narration, because would signal explanation). now say the story continues with a question like (2). (2) what was he thinking? ©2023 emar maier this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). attribution and the discourse structure of reports in principle, (2) could represent a (genuine or rhetorical) question of the writer to the reader, but, in the given narrative context, another likely interpretation is that this is rather a report of a question that one of the characters is asking. it could be the police officer reprimanding john by asking, somewhat sarcastically, “what were you thinking?”. or perhaps it’s john reflecting on his own actions, thinking to himself “what was i thinking?”. in this paper i propose to account for these kinds of report readings at the level of discourse structure. my proposal will be couched in the general framework of segmented discourse representation theory (sdrt, asher and lascarides 2003). crucially, my analysis revolves around a dedicated discourse relation called attribution. i will provide a very general semantics for attribution in terms of an underspecified notion of characterization that covers the full range of reporting types, from verbatim direct quotation to the paraphrasing of propositional content in attitude ascriptions. 2. modeling coherence in sdrt sdrt treats each individual clause in a discourse as contributing a separate discourse unit, and formulates a number of axioms that model the establishment of discourse relations, like narration, result, contrast, and elaboration, between these discourse units. unlike competing theories of discourse structure it gives these relations a model-theoretic semantic interpretation. for instance, the story in (1) gives rise to four elementary discourse units, typically labeled π1, π2, etc. (3) π1 : john was biking home late. π2 : a police officer stopped him. π3 : she gave him a fine. π4 : his lights were off. sdrt is compatible with any dynamic semantic interpretation for the individual discourse units, but in this paper i’ll use drt (kamp and reyle, 1993), and extend its box-style notation to sdrss as a whole, as illustrated in (4):1 (4) π1 : e1 x1 bike(e1) agent(e1,x1) john(x1) . . . π2 : e2 x2 police(x2) stop(e2) agent(e2,x2) . . . π3 : e3 x3 give(e3) agent(e3,x1) fine(x3) . . . π4 : e4 x4 be.off(e4) agent(e4,x4) lights(x4) . . . background(π1,π2) narration(π2,π3) explanation(π2,π3) abstracting away from the semantic contents of the elementary units we can visualize just the global coherence structure of the discourse as a graph: (5) π1 π2 π3 π4 background narration explanation 1. to avoid formal clutter in notation i leave πi discourse referents out of the drs universes and ignore the top-level π0 altogether. in the examples i discuss these can always easily be reconstructed unambiguously. 35 maier in these diagrams we stick with the standard sdrt convention of horizontal edges visualizing coordinating discourse relations, i.e., discourse relations like narration and background that in some intuitive sense move the story forward and change the active topic, and vertical edges visualizing subordinating relations, i.e., relations like explanation or elaboration that don’t move time and instead explore subtopics of the ‘dominant’ node.2 the two main questions for a formally precise and practically usable discourse semantics are: how do we derive a graph representation like (5) or (4) from a discourse like (1), and how exactly are we to interpret such formal structures? the sdrt framework provides two formal systems to answer these two questions. to start with the latter, the model-theoretic interpretation of an sdrt graph representation extends the standard drt semantics for the graph’s πi-labeled drs nodes with interpretation rules for the various discourse relations like in (6). notation: kπ1 denotes the drs unit that is labeled with proposition label π1; eπ1 denotes the main eventuality introduced in the universe of the drs unit labeled π1; jkk denotes the dynamic semantic interpretation of a drs (i.e., a context change potential, defined as a function from information states to information states, representing how an utterance affects an input context, à la groenendijk and stokhof 1991); ◦ denotes function composition (i.e., the dynamic semantic analogue of conjunction); the symbol ⃝ in a drs condition denotes temporal overlap between eventualities; ≺ denotes immediate temporal precedence (the second eventuality occurs right after the first): (6) a. jnarration(π1,π2)k=jkπ1k◦ jkπ2k◦ jeπ1 ≺ eπ2k b. jexplanation(π1,π2)k=jkπ1k◦ jkπ2k◦ jcause(eπ2 ,eπ1)k c. jbackground(π1,π2)k=jkπ1k◦ jkπ2k◦ jeπ1 ⃝ eπ2k in words, (6a) says that a narration relation between two discourse units means that we have to update the context with the contents of both discourse units, in order, and moreover the main eventuality described by the second unit, follows immediately after the event described by that of the first. now for the first question, how to derive a discourse structure representation like the graph (5) and ultimately the full sdrs (4) from a sequence of utterances? let’s assume that the elementary discourse units are already identified and assigned drs representations by the standard drs construction algorithm (see kamp and reyle 1993). now, sdrt’s so-called glue logic provides inference rules that specify what discourse configurations trigger what discourse relations. for instance, a sequence of two discourse units where the first contributes a state and the second an event licenses the inference that they are connected by a background relation – unless the resulting graph leads to an inconsistent or not maximally coherent final output representation. similarly, a sequence of two eventive units defeasibly triggers (;) a narration connection. (7) a. state(eπ1) ∧ event(eπ2) ; background(π1,π2) b. event(eπ1) ∧ event(eπ2) ; narration(π1,π2) we will not go into the formal details of either model theory or glue logic, nor into the presupposed drs construction algorithm and dynamic semantics in terms of context change potentials. i trust the above examples, diagrams and simplified formulas suffice to illustrate the basics of the sdrt 2. the main advantage of this convention is to visualize the so-called right frontier constraint that relates anaphora resolution to discourse structure. in this paper we are not concerned with anaphora resolution so we’ll skip over this (asher and lascarides, 2003). 36 attribution and the discourse structure of reports discourse semantics framework to the uninitiated, and i defer to asher and lascarides (2003) for all formal details. in the following i provide an account of reported speech in this general framework, treating reporting as a discourse phenomenon, i.e., analyzing its semantic effects in terms of a semantically interpreted discourse relation attribution (hunter, 2016; cumming, 2021). 3. indirect discourse 3.1 from operators to event modifiers attitude and speech reports have occupied a central position in semantic theory from its very beginnings (frege, 1892). in contemporary possible worlds semantics, the intensional operator approach (hintikka, 1969) and its descendants (kaplan, 1989; schlenker, 2003) are still dominant. recently, there’s been a rise in event-based versions, where the attitude or speech verb introduces an event of thinking, speaking, hoping, and the complement clause specifies the content of that event (kratzer, 2006; hacquard, 2010) (notation: ∧ϕ refers to the possible worlds proposition expressed by ϕ , which is just a traditional montagovian way of dealing with intensionality without introducing explicit possible worlds variables into the formal metalanguage). (8) a. mia said don is a phony. b. ∃e [ say(e) ∧ agent(e,mia) ∧ content(e,∧phony(don))] such an analysis fits neatly in a more general neo-davidsonian framework by treating subject and complement uniformly as event modifiers. instead of treating speech and attitude verbs as special operators it relies on the idea that there are certain events that have propositional contents. in this section i’ll adopt the event-based approach but move it from the syntax–semantics interface into the discourse/pragmatics level, where, i will argue in the remainder of the paper, it belongs. 3.2 from clausal complements to discourse units when we look at a report like (8a) from the perspective of discourse structure, the first question that arises is whether we are dealing with a single elementary discourse unit (mia said don is a phony) or with two separate units (mia said (something), don is a phony) connected by a discourse relation. hunter (2016) argues for the latter, on the basis of an ambiguity between regular (in)direct speech attributions and so-called parenthetical readings (also known as evidential or non-at-issue readings) of report constructions. in this paper i will provide a different, independent argument for this bipartite segmentation of indirect discourse, based on unembedded continuations of reports (§4). in the remainder of this subsection i first illustrate how a hunter-style bipartite analysis could work for a simple report like (8a). on hunter’s analysis, the two units in a report are connected by a discourse relation of attribution:3 (9) a. π1: mia said. π2: don is a phony. 3. a technical advantage of the kratzerian event-based approach here over the classic hintikkan intensional operator approach that hunter uses is that a unit of the form ‘mia said’ in (9a), without a grammatical object, is semantically speaking completely well-formed and interpretable. 37 maier b. π1: e x say(e) mia(x) agent(e,x) π2 : y don(y) phony(y) attribution(π1,π2) attribution is a non-veridical discourse relation, i.e., its truth does not presuppose the truth of both arguments. specifically, π2 serves to characterize what mia said, not what the world is actually like. we build this into our semantics as follows, using the content(e,p) relation from §3.1:4 (10) jattribution(π1,π2)k = jkπ1k◦jcontent(eπ1 , ∧kπ2)k (to be revised) this definition presupposes that π1 introduces a main eventuality (eπ1) that can plausibly be said to have a propositional content, such as an utterance event, an occurrent thought, an attitudinal state, or a perceptual state/event. this requirement should ultimately be included in the antecedent of a defeasible glue logic axiom for inferring an attribution connection, of the form in (11), but we’ll leave the precise conditions in the ‘. . . ’ for another occasion: (11) contentful.eventuality(eπ1)∧ . . .; attribution(π1,π2) a defeasible inference rule of the form in (11) should allow us to infer attribution in passages where we have one clause introducing a speech, thought, or attitude event, and another that could plausibly be interpreted as specifying that event’s content. in the case of (8a) however we have a grammatical report construction that, arguably, forces an attribution connection between frame and complement. this situation is parallel to what we see with most other discourse relations. a contrast may be left implicit, defeasibly inferred by the interpreter on the basis of various semantic, pragmatic, and discourse structural cues, but it may also be encoded directly in the grammar by means of an unambiguous, dedicated lexical item like but. similarly, we have narration, optionally marked by and then, or explanation by because. in sdrt, lexical items like these directly inform the glue logic, i.e., they simplify the sdrs construction process by filling in a fixed discourse relation. with attribution, we could assign this function to the complementizer that (which may be silent).5 in any case, whether marked on the surface as a report or inferred pragmatically on the basis of (11), π1 and π2 are going to be connected by attribution here. π1 introduces a speech event, and (10) then tells us that the content of that event must be the proposition expressed by π2. in other words, our bipartite discourse structure analysis gives us exactly the truth conditions that we also got from the compositional semantics in (8b). 4. the term ‘attribution’ is somewhat ambiguous: we usually say that we attribute an attitude or opinion to an individual, but strictly speaking the discourse relation of attribution here connects the content of the attitude/opinion to the event or state of an individual experiencing or expressing said attitude or opinion. since there seems to be little risk of confusion, i’ve decided to stick with the now established sdrt terminology (e.g. asher et al. 2006; hunter 2016; abrusán 2021). 5. alternatively, we can point out some other part of to the grammatical structure of a communication or attitude verb plus subordinated complement clause to encode the glue logic restriction to attribution. 38 attribution and the discourse structure of reports 3.3 parenthetical reports before i introduce my own applications of the bipartite discourse analysis of reporting, let’s briefly review hunter’s (2016) application to what she calls parenthetical indirect reports: (12) a: why is mia not in class? b: joe said she has covid. a traditional hintikka or kratzer semantic analysis of b’s answer gives us the proposition that joe produced a speech act with a certain content, which hardly counts as an answer to a’s question about mia. but intuitively b does provide an acceptable answer. according to hunter, this is because b’s indirect discourse report here allows a so-called parenthetical reading, presenting the reported information (that mia has covid) as the primary, at-issue contribution, with the reportative information (that joe said something) serving as a not-at-issue meaning supplement. the current discourse approach to reports, where we analyze a report as consisting of two distinct discourse units, seems ideally suited to account for such parenthetical readings without having to assume a syntactic ambiguity. to prove this, let’s first consider the exact same report in a different, more narrative context like (13). (13) so we’re sitting in that waiting room, when mia starts coughing, and then joe said she has covid. all hell broke loose. in (13), as in the cases we’ll be discussing in the next sections, it’s the reporting segment, that joe said something, that is at-issue, or as hunter operationalizes it in sdrt, it’s the reporting segment that directly connects to the previous discourse (via narration, in this case): first mia coughs, and then joe says something (and as a result all hell breaks loose). (14) . . . π1:mia starts coughing π2:joe said π3:she has covid narration attribution now back to the trickier case of the dialogue in (12). here, it’s the information contributed by the reported clause (that mia has covid), that directly connects to the previous discourse (i.e., a’s question about mia’s whereabouts), in this case via the questionanswer relation. to a first approximation, the discourse seems to be structured like this: (15) π1:why is mia not in class? π2:joe said π3:she has covid questionanswer attribution one complication that arises is that b probably uses the report embedding as way to hedge their own commitment to the truth of the complement, while the use of the veridical relation questionanswer in (15) entails full commitment on b’s part. hunter’s solution involves the introduction of 39 maier ‘modalized discourse relations’, such as 3questionanswer, to decorate diagonal connections between material above and below an attribution arrow (as we see in (15)). in the following we’ll ignore this and other complications (e.g., relating to syntactically parenthetical or evidential reports) as we focus on what bary and maier (2021) call ‘at-issue eventive’ (uses of) reports like in (13), where it’s the report frame that’s directly connected to the previous discourse. the upshot of this section is that we can emulate the classic semantic analysis of indirect discourse in a discourse framework, effectively moving the intensional embedding semantics from the syntax/semantics interface to the level of discourse structure. we’ve seen one distinctive benefit of the discourse approach, due to hunter, viz. that it allows us to capture discourse parenthetical readings of reports without postulating ad hoc syntactic ambiguities. 4. reports beyond the clause on the discourse approach to reporting, the attitude or speech verb plus clausal complement construction is treated as a cue that informs the pragmasemantic glue logic of sdrs construction to infer the discourse relation of attribution between two discourse units. in the remainder of the paper i show that the powerful added machinery of the discourse-level approach is warranted by cases that are not overtly marked as reports but nonetheless interpreted as such. the most salient example of this is probably free indirect discourse, to be discussed in section 6. below we first discuss another case that has received far less attention: indirect report continuations beyond the overtly embedded complement clause. consider the following extended dream report: (16) dan went to bed early. he dreamed that he was a frog. he jumped around a bit and then he was eaten by a stork. on our discourse-level approach we parse this discourse as consisting of 5 segments. (17) π1 : dan went to bed early π2 : he dreamed π3 : (that) he was a frog π4 : he jumped around a bit π5 : (and then) he was eaten by a stork the corresponding discourse units can be straightforwardly connected by discourse relations to create an interpretable discourse graph. note that in this example, two discourse relations are arguably encoded grammatically: the complement construction in dreamed that encodes attribution and and then encodes narration. the rest can be defeasibly inferred by existing glue logic axioms, such as the sequence of events in π1 and π2 giving rise to a likely narration inference, and the sequence of state and event in π3 and π4 giving rise to a likely background inference. crucially, when we read this story, we take π3-π5 together to form a complex description of a single dream (that consists of an internally coherent sequence of events). to correctly represent this reading we need to construct a so-called complex discourse unit (asher and lascarides, 2003). we 40 attribution and the discourse structure of reports then take that complex unit (rather than just π3) as the second argument of the attribution. in our graph notation, we draw a labeled box to indicate the scope of a complex unit (here: π ′): 6 (18) π1 π2 π ′: π3 π4 π5 background narration narration attribution when we try to describe the discourse interpretation process leading up to this graph procedurally, the question arises: at what point do we create the complex unit? this is a thorny and quite general question for sdrt, but to keep it simple i propose to stipulate that the second argument of attribution7 always comes with a complex unit. thus, officially, the simple report he dreamed that he was a frog in isolation would already lead to the creation of a complex unit π ′ with π3 as its sole contents.8 following the standard sdrt procedural attachment rules, this embedded unit π3 is available for subsequent discourse units to attach to (it’s on the so-called right frontier, asher and lascarides 2003). in the case of an isolated, single report clause, the stipulated extra layer of embedding is semantically superfluous, so we might as well introduce a notational shorthand to the effect that these extra embedding boxes are not drawn (resulting in familiar graph diagrams like (9b)) until they contain more than one discourse unit and thereby become semantically relevant (as in (18)). the eventual discourse graph in (18) straightforwardly captures the ‘modal subordination’ (roberts, 1989) reading, where π4 and π5 are interpreted as describing the content of the dream, despite being syntactically outside the scope of the attitude verb. by contrast, the only way for a traditional sentence-level report semantics to deal with this would be to assume a silent dream operator in front of every proposition interpreted as a dream description. note that such a sequence of hidden operators would ultimately still fail to capture the obvious discourse structural, temporal, and anaphoric relations between these segments (e.g., π4 and π5 could not be connected to each other by narrationif they were each syntactically, semantically, or even discourse structurally, embedded by their own separate intensional operator. similar unmarked continuations of reports occur with other attitude and speech reports. in some languages, such syntactically unembedded continuations of speech reports can be marked with a reportative subjunctive mood on the verb: 6. interestingly, when we continue the discourse in (16) with he woke up screaming, we should attach that via result to π2 at the top-level, because it’s not part of the dream description. he went to bed, and then had a dream (with such and such content), and as a result of the dream he woke up screaming. as one referee points out, we could construct another complex unit around π2 and π ′ and connect that to the waking up screaming. either way, since attribution is non-veridical, we can’t connect the screaming directly to the being eaten with a veridical discourse relation. 7. probably we can generalize this to any non-veridical argument of a coherence relation, as that’s where complex discourse units are crucial. 8. as the editor, amir zeldes, points out, existing sdrt implementations geared towards corpus annotation tend to forbid vacuous embeddings like this, i.e. discourse units that contain nothing but a single discourse unit. an alternative procedure then would be to stick with the simple bipartite analysis of a syntactic indirect discourse that we had in (9b). the complex discourse unit is then created ‘on the fly’, whenever a followup sentence more coherently attaches underneath the attribution than above it. 41 maier (19) sie she sagte said sie she habe have-subj keine no zeit. time. sie she müsse must-subj noch still 86 86 prüfungen exams bewerten. grade ‘she said she has no time. she still has 86 exams to grade (she said)’ (german, bary and maier 2021) in such constructions, the traditional, compositional approach would take the subjunctive morpheme as a separate semantic report operator (which causes significant complications for dealing with the overtly embedded subjunctive in the first sentence in (19), see fabricius-hansen and saebø 2004). on the current approach, we take the subjunctive merely as a grammatical cue that constrains the glue logic to block attachment of the current unit to a top-level unit, i.e., forcing it to attach to a unit under an attribution. in english, where we have no subjunctive inflection to mark something as reported content, we occasionally find unmarked free standing clauses that are interpreted as speech report continuations: (20) trump says he’ll cut inflation in half. he’ll also create record numbers of jobs and beat covid before christmas. as in (18), by connecting the propositions about inflation, record job numbers, and covid together into a complex unit (using coordinating, veridical relations like list or continuation between them), we automatically get the most likely reading where all three together are semantically interpreted as describing what trump said, without relying on any covert operators in the syntax. 5. quotation the above event-based implementation of hunter’s (2016) discourse-structural approach to indirect discourse applies to both speech and attitude reports in the indirect mode, i.e., where we are reporting the content of another person’s speech or attitudinal state in our own words. i propose to generalize the semantics of attribution in order to cover also quotation and free indirect discourse reports, which seem to exhibit similar sensitivity to discourse structure, like allowing complex report continuations far beyond the sentence level. 5.1 direct discourse and pure quotation we start with a simple, clausal, direct quotation. on an event-based account we can treat direct and indirect speech uniformly as event modifiers, one that characterizes a speech event by its propositional content, and one that characterizes it by its linguistic form (maier, 2017): (21) a. mia said, “don is a phony” b. ∃e [ say(e) ∧ agent(e,mia) ∧ form(e,‘don is a phony’) ] as with indirect reports i now propose a discourse-level alternative to this type of (near-)compositional account that retains the idea of treating quotation as event modification. we parse the quotation and the frame as distinct discourse units, connected by attribution. 42 attribution and the discourse structure of reports (22) π1:mia said π2:“don is a phony” attribution now, to get the right truth conditions we could technically admit two distinct, primitive types of attribution: one defined as in (10), contributing jcontent(eπ1 , ∧kπ2)k, and one, say qattribution, contributing instead something like jform(eπ1 ,σπ2)k (with σπ2 denoting the linguistic/graphemic/phonological surface form of speech act π2). however, this move will lead us down a path of multiplying discourse relations for each type of reporting, including, beyond direct and indirect discourse, mixed quotation, free indirect discourse, speech balloons, etc. in this paper i explore an alternative route, where we stick with a single discourse relation of attribution. to make this work we have to generalize its semantic contribution so that it subsumes both formand content-based reporting. 5.2 attribution as underspecified event characterization i propose to replace our original definition of the semantics of content-based attribution in (10) with (23), which invokes a distinct notion of ‘characterization’. in this definition the notation ‘char(f (π),e)’ means that ‘discourse unit π characterizes event e’, using asher and lascarides’s (2003) official sdrt notation, in which f denotes the function that maps labels in an sdrs to the sdrs constituents that they label – that is, f (π) is effectively a notational variant for what we’ve been denoting as kπ (we’ll rely on this more abstract notation below in making precise what characterization does). (23) jattribution(π1,π2)k=jkπ1k◦jchar(f (π2),eπ1)k the idea behind (23) is that languages may allow different ways of characterizing what someone said, thought, or dreamed. we can characterize what someone said by reproducing its propositional content in our own words. that is what happens in indirect discourse reports, and it is exactly this type of ‘loose’ characterizing that is formalized explicitly in our original formulation of the semantics of attribution in (10). but we can also characterize what someone said by reproducing the exact words uttered. this is what happens in direct discourse.9 the proposed general approach to attribution leaves us with the question of what to do with the actual quotation marks. are they merely a cue to enforce the inference of an underspecified attribution– the way we suggested treating the reportative subjunctive mood in (19) above –, or are they a genuine semantic quotation operator applied to the second attribution argument? applied to the current sdrt setting, the first option – in the spirit of pragmatic accounts of quotation like gutzmann and stei (2011) – would mean that at the level of semantic representation quoted sentences are treated just like any other discourse unit, i.e., parsed and assigned their regular drs representation. but for reports with quotation marks we need more than just the semantic representation of the complement, we need access to the actual form of the words used to express it. i propose that’s what quotation marks do: they tell the drs construction algorithm to introduce a surface form into the semantic representation. for reasons to be discussed below we’ll assume that we also construct the regular drs representation of the quoted material, where possible. hence, in 9. below we’ll briefly survey some other forms of characterization, such as simultaneous form and content characterization, diagonal characterization, and iconic characterization. 43 maier the full sdrs representation of (22), the report frame π1 is represented as just a content drs, while the quoted unit π2 is represented as a form–content pair, consisting of a copy of the quoted surface form along with a drs representation of its content. (24) π1 : e x mia(x) say(e) agent(e,x) π2 : 〈 don is a phony , y don(y) phony(y) 〉 attribution(π1,π2) we can now be more precise about the two most salient types of characterization that figure in the semantic definition of attribution. first, propositional characterization: a drs k propositionally characterizes a contentful eventuality e if the proposition expressed by k matches the propositional content of e.10 second, formal characterization: a form–content pair formally characterizes a speech or thought event e if the form component matches the linguistic form of the reported speech event.11 we can rephrase this more formally as in (25), using the following notational conventions: jϕk f ,c w is the (static)12 semantic interpretation of an atomic drs condition ϕ , i.e., its truth value relative to an assignment f , a kaplanian context c and a possible world index w; jϕk f ,c = λw[jϕk f ,c w ], i.e., the proposition expressed by ϕ; and content and form are the by now familiar functions mapping certain events to their propositional contents and surface forms, respectively. (25) a. jchar(k,e)k f ,c w is defined iff f (e) is a contentful eventuality (speech event, belief state, etc.). if defined, jchar(k,e)k f ,c w = 1 iff content( f (e)) = jkk f ,c b. jchar(⟨σ ,k⟩,e)k f ,c w is defined iff f (e) is a linguistic speech act or language-like occurrent thought. if defined, jchar(⟨σ ,k⟩,e)k f ,c w = 1 iff form( f (e)) = σ (to be revised) the definition of characterization in (25) together with the general semantics of attribution from (23) allows us to model some standard forms of direct and indirect discourse adequately. it effectively recreates the truth-conditional predictions of a traditional account of direct discourse as pure quotation, and a traditional account of indirect discourse as an intensional operator (or rather, as contentful event modifier) (kaplan, 1989; brasoveanu and farkas, 2007; maier, 2017). note that (25b) effectively ignores the second component of the form–content pair in a quotation, which entails that quoted words are never really interpreted at all, they just contribute their form, i.e., their ‘shape’ (davidson, 1979), to the eventual interpretation. this would be fine if all we’re interested in are the kinds of pure and direct quotations discussed in the philosophical liter10. i’m assuming here that propositional matching means identity between sets of possible worlds. this is an oversimplification. the original speech act may in fact have been quite different from the reported complement (e.g., i can report that mary said that she’s coming if she literally said something more specific, like “i’ll be at the party between 9 and 10pm” (von stechow and zimmermann, 2005; abreu zavaleta, 2019). 11. again, for simplicity i’ll assume matching means identity between strings of letters or phonemes, though to model judgments regarding natural language quotation more realistically we have to make room for cleaning up false starts and filled pauses and allow literal translations, at the very least. 12. in drt we typically use essentially static truth definitions for conditions as part of a definition of dynamic context change potentials for drss. see kamp et al. (2003) for details. 44 attribution and the discourse structure of reports ature, like ‘boston’ is a six letter word and otto said “i’m a fool”. but when we’re interested in more global discourse structures in actual text, this will prove unsatisfactory.13 5.3 complex quotations take a, still very simple, quotation like (26). (26) “oh, we’ll be cutting,” trump told the audience. “but we’re also going to have tremendous growth.” on the one hand, the two quoted fragments flanking the report frame should somehow be linked together, because together they characterize the form of trump’s speech. on the other hand, they clearly contribute two distinct discourse units of their own that are moreover meaningfully connected by contrast (as evidenced by the overt connective but). in other words, we want a discourse graph like (27): (27) π2:trump told audience π ′: π1:. . . cutting π3:. . . tremendous growthcontrast attribution in order to correctly infer coherence (and anaphoric) connections between multi-sentence quotations and derive graph structures like (27), the glue logic needs to have some access to the semantic content of quoted discourse units as well as their forms. the first step we already took is to represent both form and content in the sdrs representation of a quote, but now we still have to revise the form–content interpretation rule in (25b) to take advantage of that two-dimensional representation. on a more technical note, when we spell out the full semantic (s)drs box representation corresponding to the abstract discourse graph in (27), we get a complex unit, π ′, as the second component of our attribution, but as it stands this will be just a box around two form–content pairs, which is not itself a form–content pair yet,14 and hence will not even trigger a quotational interpretation in the first place. to remedy this technical problem first, we add a ‘form-projection’ rule: when we attach a form–content pair to another form–content pair inside a complex unit (which, we stipulated in §4, is always present under attribution), the embedded form components project up to the complex discourse unit containing them, where they are concatenated (notation: ∩). (28) form-projection rule: π ′: π1:⟨σ1,k1⟩ π2:⟨σ2,k2⟩ ; π ′: 〈 σ∩ 1 σ2 , π1 : k1 π2 : k2 〉 13. partee (1973) and others have already provided well-known arguments against the pure quotation approach to direct discourse on the basis of anaphora and ellipsis dependencies between quotation and surrounding discourse, as in: “don’t worry, my boss likes me! he’ll give me a raise” said mary, but given the economic climate i doubt that he can. (maier, 2015) 14. the simple pure quotation analysis of §5.1 already can be seen as suffering from a milder version of this technical issue, if we had strictly followed our official stipulation that the second argument of attribution is always a complex discourse unit. 45 maier applying form-projection to our example we get the following full sdrs representation for (27): (29) π2: e2 tell(e2) π ′: 〈 oh, we’ll be cutting. but we’re also going to have tremendous growth , π1: e1 cut(e1) π3: e3 have.growth(e3) contrast(π1,π3) 〉 attribution(π2,π ′) in sum, a straightforward form-projection mechanism thus puts complex quotations like (26) in the right format to feed into our semantics, as laid out in (23) and (25). now to make our quotation semantics sensitive to both form and content (i.e., as philosophers put it, treat direct quotation as simultaneous mention and use (davidson, 1979; cappelen and lepore, 1997)), i’ll follow a straightforward implementation based on the two-dimensional account of direct quotation of potts (2007): a form–content pair ⟨σ ,k⟩ characterizes a speech or thought event e if the first component σ formally characterizes e and the second component k propositionally characterizes e. one complication we run into when we spell this out is that we have to incorporate a context shift in the content-matching criterion: propositional characterization in the case of direct discourse must compare the content of the speech/thought event e to the content of the complement k relative to the shifted, reported context of utterance, not relative to the actual, reporting context of utterance (as in regular indirect discourse) (potts, 2007). a context shift is necessary in order to get the reference of indexicals right – in direct discourse, all indexicals are systematically shifted. i’ll assume a function context mapping a speech/thought event to the context in which it takes place (eckardt, 2015).15 in sum, we replace the second clause, (25b), in our general definition of characterization with a stricter definition that demands matching of form and content simultaneously, like this:16 (30) jchar(⟨σ ,k⟩,e)k f ,c w is defined iff f (e) is a linguistic speech act or language-like occurrent thought. if defined, jchar(⟨σ ,k⟩,e)k f ,c w = 1 iff form( f (e)) = σ and content( f (e)) = jkk f ,context( f (e)) as a further illustration of this rather technical, auxiliary notion of form–content characterization, let me show how it can be used to analyze mixed quotation. 5.4 mixed quotation mixed quotation typically involves an indirect discourse where part of the clause is quoted directly (davidson, 1979): (31) biden said that putin “totally miscalculated” 15. context(e) = ⟨w, t,x⟩ iff e occurs in w at time t and the agent of e is x. this is assuming events are world-bound particulars. if we instead assume that a single event can occur in different possible worlds we would have to add the world as an extra parameter, i.e., context(e,w). 16. if we allow the content compartment to be empty, and in such cases disregard it semantically, we get a way to account for the intuitive well-formedness and interpretability of quoting gibberish (she was like “shis thewgg”. 46 attribution and the discourse structure of reports following the so-called presuppositional analysis of the phenomenon (geurts and maier, 2005; maier, 2014), the intended interpretation can be schematically represented as involving two meaning components: an assertion of an (underspecified) indirect discourse, (32a), and a metalinguistic presupposition, (32b):17 (32) a. assertion: biden said that putin has property x. b. presupposition: biden used the words ‘totally miscalculated’ to express property x. maier’s (2014) drt implementation of this idea involves a primitive three-place relation e(xpress): e(x,totally miscalculated,x) ≈ x uses the linguistic expression totally miscalculated to express semantic property x . when we port the ideas behind the presuppositional account of mixed quotation over to the current sdrt framework we can actually reduce this primitive e-relation to the independently motivated char(acterization) relation introduced above. to make this precise, let’s work out the interpretation of the simple example in (31). we start by segmenting the basic biclausal report as consisting of two units: (33) π1: biden said π2: putin “totally miscalculated” for the compositional semantic interpretation of π2 let’s follow my 2014 syntactic parse and drs construction, where the mixed quoted vp leads to the introduction of a discourse referent x (of type ⟨e, t⟩, i.e., ranging over properties), together with a metalinguistic presupposition, viz. that x is the property s.t. there was a saying event e′ and x matches the content of e′ and totally miscalculated matches the form of e′. we can capture this combination of form and content matching in a simple drs condition using the char relation (as defined in (30)). further notational convention: unresolved presuppositions are represented by dashed boxes that sit between the triggering drs box and its label. (34) π1: x1 e1 biden(x1) say(e1,x1) π2: x e′ say(e′) char(⟨totally miscalculated,x⟩,e′) x2 putin(x2) x(x2) attribution(π1,π2) we can resolve the metalinguistic presupposition in the (accessible) π1 box, by accommodating both x and e′ there. note that while we can’t directly bind e′ to e1 and fully equate them (because the content of e1 is a full proposition, and that of e′ is just a property), we can plausibly add a bridging inference to the effect that e′ is a subevent of e1 (notation: e′ < e1). (35) π1: x1 e1 e′ biden(x1) say(e1) say(e′) e′ < e1 char(⟨totally miscalculated,x⟩,e′) π2: x2 putin(x2) x(x2) attribution(π1,π2) 17. for arguments that the second component, in (32b), really is a presupposition and not some other type of (not-at-issue) content, i refer to (maier, 2014). 47 maier we’ve seen here how characterization plays an important role in capturing the metalinguistic presupposition triggered in mixed quotations. this is in line with the idea that characterization is a useful concept in its own right, beyond an auxiliary technicality that allows a more unified simple statement of the general meaning of attribution (i.e., attribution(α ,β ) means that β characterizes the main eventuality of α). in fact, the definition of attribution in terms of char opens up a variety of potential further extensions. let me end this section with a few directions for future extensions of the framework, based on intuitively plausible extensions of the notion of characterizing. first, we could define an intermediate mode of characterization, somewhere in between formal and propositional characterization, viz. characterization at the level of kaplanian character or its diagonal (kaplan, 1989; stalnaker, 1978; zimmermann, 1991). this would be useful for capturing monstrous or de se reports.18 second, we could extend characterization to the visual modality, to analyze distinctively visual conventions for representing characters’ speech, thoughts, dreams, or hallucinations as instances of attribution, as part of an overall sdrt approach to analyzing the narrative structure of sequential visual media like comics and film (bateman and wildfeuer, 2014; cumming et al., 2017).19 third, we might extend characterization beyond contentful events to model demonstrations more generally. for instance, the semantics of mary ate like ¡gobbling gesture¿ (davidson, 2015) would involve an event of eating being ‘iconically characterized’ by a gobbling gesture. incorporating such extensions and comparing various implementations is beyond the scope of this paper, which focuses on the general account of reporting as a discourse-structural phenomenon. 6. free indirect discourse free indirect discourse is a form of reported speech or thought that shows characteristics of both direct and indirect discourse (banfield, 1982). take (36). (36) sue stared at the calendar. oh no, she had to hand in that damn paper today! she’d never make it. . . the first sentence is just a description of what’s going on in the story world, but the next two seem to describe what’s going on inside sue’s head. the way this ‘perspective shift’ is marked linguistically is often subtle but it involves a combination of the use of expressive and indexical elements (oh no, damn, today, !) directly representing the protagonist sue’s point of view (i.e., as in direct speech), 18. to define this concisely, assume that contexts (c ∈c) and indices (w ∈w ) are tuples of the same type, i.e., indices are contexts with unused coordinates for agent, addressee, location etc., so that c ⊆w (von stechow and zimmermann, 2005). then we can easily define diagonal content as a a set of contexts: diagonal drs content: \\k\\ f = λc.jkk f ,c c we could now say that a discourse unit π with a drs component kπ diagonally characterizes a contentful eventuality e if the ‘de se content’ of e (the set of contexts ‘compatible with e’, lewis 1979; schlenker 2003) corresponds to the diagonal content of kπ . 19. interestingly, the seemingly distinctive visual technique of the ‘blended perspective shot’ (e.g., presenting a character’s internal perceptual hallucinations from a seemingly objective, neutral observer viewpoint) may already be captured by the plain content matching clause in (25a) that we’ve used for interpreting regular indirect discourse (maier and bimpikou, 2019; maier, 2022). 48 attribution and the discourse structure of reports and the regular narrative past tense and third person pronouns (she had to, she’d) representing the thinking protagonist from the narrator’s ‘third person’ perspective (i.e., as in indirect speech).20 linguists have examined the semantic properties of free indirect discourse in some detail, and have proposed various competing semantic analyses, e.g., in terms of monstrous indirect discourse (sharvit, 2008), the addition of an extra context parameter (schlenker, 2004; eckardt, 2014), and quotation plus unquotation (maier, 2015, 2017). some salient features of free indirect discourse that are often overlooked by semanticists are (i) that these types of reports tend to span several sentences or even entire paragraphs, and (ii) that it may require intricate textual analysis to pinpoint exactly where such a report starts or ends. these neglected features however are exactly the type of thing we would expect on a discoursestructural approach. on our attribution-based approach, once we have established that there’s an attribution, we get for each new incoming discourse unit a choice: do we attach it to the complex unit underneath that attribution(i.e., treat it as a continuation of the report), or to the main story line above it (i.e., treat it as a narrative description of the story world)? this choice is guided by often subtle considerations of global discourse coherence, i.e., which attachment generates a more coherent overall output sdrs (asher and lascarides, 2003). combined with the lack of clear, overt cues like quotation or (in english) subjunctive mood marking, this explains the observed difficulty of determining the exact boundaries of free indirect discourse passages. let me now flesh out the proposed discourse-structural attribution account of free indirect discourse by applying it to the example in (36). attuned to the grammatical cues for free indirect discourse detection, sketched above, we can recognize three discourse units, of which two form a complex node that is connected to the previous discourse via attribution. but strictly speaking, attribution can’t have the staring eventuality as its first argument, because staring is not in any way a contentful or linguistically structured event that can sensibly be characterized by a form or a content. following recent discourse-structural analyses of free indirect discourse (abrusán, 2020; bimpikou et al., 2021; altshuler and maier, 2022) i propose that we may in such cases accommodate a simple discourse unit, π3, to introduce the required thought event. (37) π1 : sue stared at the calendar. π2 : oh no, she had to hand in that damn paper today! π3 : (she thought.) π4 :she’d never make it. . . (38) π1 π3 π ′: π2 π4 result background attribution due to the inherent underspecification in the semantics of attribution, this graph is in principle compatible with the various competing semantic analyses of the interpretation of free indirect discourse constructions. all that (38) tells us about the reports is that π2 and π4 together characterize 20. see abrusán (2021) for discussion of a more comprehensive algorithm for detecting ‘perspective shift’ based on grammatical, lexical and discourse-level cues. 49 maier the (accommodated) thought event in π3. in its abstract graph form it doesn’t specify what kind of characterization this is – simultaneous use/mention quotation, indirect discourse, or something else. but if we want to spell out the full sdrs box corresponding to the graph, and its interpretation, we’ll eventually have to settle on a specific semantic theory. i’ll explore here my own quotationplus-unquotation approach.21 let’s assume, following the argumentation of maier (2015), that the drs construction algorithm treats a free indirect discourse segment – recognized as such – as essentially quoted. this means that we introduce corresponding form layers for π2 and π4. but, still following maier (2015), pronouns and tenses are to be treated as ‘unquoted’.22 technically, that means these pronouns and tenses are ‘moved’ out of the reports and interpreted separately, leaving (metalinguistic) traces (maier, 2014). let’s go through the steps of the drs construction algorithm for the first part of our example. first, we assume a (usually covert) quotation with (covert) unquotation of all pronouns and tenses, (39b). to interpret this semantically we first move the unquoted elements out of the quotation, (39c). (39) a. oh no, she had to hand in that damn paper today! b. “oh no, [she] have-[past] to hand in that damn paper today!” c. shex pastt “oh no, [x] have-[t] to hand in that damn paper today!” now we apply the standard drs construction algorithm to the expressions in (39c). the two extraposed elements shex and pastt are anaphoric in nature and hence trigger presuppositions, the quotation will give rise to a labeled form–content pair consisting of the surface form (with two indexed holes) and a drs box. the only new feature we have to add to the construction algorithm is a way to deal with indexed holes in a surface form. since the traces tie each hole to a corresponding presupposition trigger, we can simply represent the contributions of the holes as the corresponding presupposed discourse referents, i.e., x and t, respectively. (40) π2 : x fem.3.sg(x) t t < n 〈 oh no, [x] have-[t] to hand in that damn paper today! , e2 y2 paper(y2) hand.in(e2) agent(e2,x) theme(e2,y2) time(e2, t) today(t) 〉 we can now add (40) to the sdrs under construction by connecting its discourse label to a suitable existing label (e.g., to a thought event, via attribution, or to another quoted or otherwise reported event already under an attribution). looking at the earlier graph structure in (38), we have neither a suitable attribution nor a thought event, so we’ll have to accommodate a thought event unit π3 and attach (40) to that with an attribution to get the following sdrs: 21. a monstrous account à la sharvit (2008) would involve defining a mode of characterization that preserves the character or diagonal for most of the report, but preserves only content for pronouns and tenses, presumably relying on some feature deletion mechanism already at the syntax/semantic level of drs construction. 22. maier (2017) seeks to derive the unquote-pronouns-and-tenses assumption from general pragmatic interpretation and production principles. 50 attribution and the discourse structure of reports (41) π1 : e1 x1 sue(x1) stare(e1) agent(e1,x1) π3 : e3 think(e3) agent(e3,x1) π2 : x fem.3.sg(x) t t < n 〈 oh no, [x] have-[t] to hand in that damn paper today! , e2 y2 paper(y2) hand.in(e2) agent(e2,x1) theme(e2,y2) time(e2, t) today(t) 〉 background(π1,π3) attribution(π3,π2) now we can resolve the presuppositions: x (she) binds to x1, the only salient female third person, and t binds to the time of the thinking (e3).23 now we add the final unit, π4. we’ll assume this is fed to the construction algorithm as a free indirect discourse, i.e., with quotation marks and unquotation holes, yielding a form–content pair with presuppositions, as in (40). we attach this π4 to the existing form–content pair, π2, within the existing (but previously invisible by the notational convention of section 4) complex discourse unit π ′ under the existing attribution; project and concatenate the form components following (28); and bind π4’s unquoted tense and pronoun presuppositions. this gives the final output sdrs in (42), ascribing to sue a complex thought whose form and content is characterized by two coherently connected discourse units. (42) π1 : e1 x1 sue(x1) stare(e1) agent(e1,x1) π3 : e3 t3 think(e3) time(e3, t3) agent(e3,x1) π ′ : 〈 oh no, [x1] have-[t3] to hand in that damn paper today! [x1] will-[t3] never make it in time , π2 : e2 y2 paper(y) hand.in(e2) agent(e2,x1) theme(e2,y2) time(e2, t3) today(t3) π4 : e4 make.it(e4) agent(e4,x1) time(e4, t3) in.time(t3) result(π2,π4) 〉 background(π1,π3) attribution(π3,π ′) 23. more precisely, the antecedent time t3 is introduced in kπ3 via a bridging inference. and this is still a simplification, as we have occurrences of both t3 and x1 in π ′ that are bound from outside the thought representation, leading to traditional philosophical worries about ‘quantifying in’, whose resolution is entirely orthogonal to the matters at hand. 51 maier 7. conclusion i have proposed abandoning attempts to model reporting constructions in terms of various clausal operators integrated in a compositional semantics. instead, we should model them at the level of discourse structure. more specifically, i have proposed a discourse-structural account of all reporting in terms of a discourse relation of attribution connecting two distinct discourse units: one contributed by a frame segment (she said, he dreamed) and one complex report unit contributed by, for instance, a clausal complement (that he was unhappy), or a complex multi-sentence quotation (“i’ll beat covid. but not global warming. that’s still a hoax”). i have proposed a simple semantics for the discourse relation of attribution that relies on the notion of a speech/thought/attitude eventuality being ‘characterized’ by a surface form or a propositional content, or both. the proposed discourse-structural account is embedded in the general discourse semantics framework of sdrt. clausal complements are simply analyzed as contributing their own discourse units, represented by a labeled drs in the discourse-level ‘logical form’ (the sdrs). quotation marks serve to introduce a surface form layer on top of the drs representation of a quoted unit. these straightforward assumptions allow us to implement simultaneous use and mention for direct quotation, which i motivate with cases where multiple quoted sentences together form a complex discourse unit describing an internally coherent multi-sentence quotation. more generally, it is such cases of extended direct, indirect, and free indirect reports, beyond the single reported clause, that have been the blind spots of traditional semantic accounts of attitude reports and quotation and that motivate the proposed shift from the syntax/semantics interface, to the level of discourse structure when it comes to understanding reports. acknowledgments this research is supported by nwo vidi grant 276-80-004 (the language of fiction and imagination). i thank the editor amir zeldes and three anonymous referees for their helpful feedback. references martı́n abreu zavaleta. weak speech reports. philosophical studies, 176(8):2139–2166, 2019. doi: 10.1007/s11098-018-1119-2. márta abrusán. the spectrum of perspective shift: protagonist projection versus free indirect discourse. linguistics and philosophy, 2020. doi: 10.1007/s10988-020-09300-z. márta abrusán. computing perspective shift in narratives. in emar maier and andreas stokke, editors, the language of fiction, pages 325–348. oxford university press, 2021. doi: 10.1093/ oso/9780198846376.003.0013. daniel altshuler and emar maier. coping with imaginative resistance. journal of semantics, 39 (3):523–549, 2022. doi: 10.1093/jos/ffac007. nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. isbn 978-0-521-65058-8. 52 attribution and the discourse structure of reports nicholas asher, julie hunter, pascal denis, and brian reese. evidentiality and intensionality: two uses of reportative constructions in discourse. in workshop on constraints in discourse structure, 2006. ann banfield. unspeakable sentences: narration and representation in the language of fiction. routledge & kegan paul, 1982. corien bary and emar maier. the landscape of speech reporting. semantics and pragmatics, 14 (8), 2021. doi: 10.3765/sp.14.8. john bateman and janina wildfeuer. a multimodal discourse theory of visual narrative. journal of pragmatics, 74:180–208, 2014. doi: 10.1016/j.pragma.2014.10.001. sofia bimpikou, emar maier, and petra hendriks. the discourse structure of free indirect discourse. in mark dingemanse and elena tribushinina, editors, linguistics in the netherlands. john benjamins, 2021. adrian brasoveanu and donka farkas. say reports, assertion events and meaning dimensions. pitar mos: a building with a view. papers in honour of alexandra cornilescu, pages 175–196, 2007. herman cappelen and ernest lepore. varieties of quotation. mind, 106(423):429–450, 1997. doi: 10.1093/mind/106.423.429. sam cumming. narrative and point-of-view. in emar maier and andreas stokke, editors, the language of fiction, pages 221–254. oxford university press, 2021. doi: 10.1093/oso/ 9780198846376.003.0009. samuel cumming, gabriel greenberg, and rory kelly. conventions of viewpoint coherence in film. philosophers’ imprint, 17(1):1–29, 2017. url http://hdl.handle.net/2027/spo. 3521354.0017.001. donald davidson. quotation. theory and decision, 11(1):27–40, 1979. doi: 10.1007/bf00126690. kathryn davidson. quotation, demonstration, and iconicity. linguistics and philosophy, 38(6): 477–520, 2015. doi: 10.1007/s10988-015-9180-1. regine eckardt. the semantics of free indirect speech. how texts let you read minds and eavesdrop. brill, 2014. regine eckardt. utterance events and indirect speech. grazer linguistische studien, 83:27–46, 2015. cathrine fabricius-hansen and kjell johan saebø. in a mediative mood: the semantics of the german reportive subjunctive. natural language semantics, 12(3):213–257, 2004. gottlob frege. über sinn und bedeutung. zeitschrift für philosophie und philosophische kritik, 100(1):25–50, 1892. bart geurts and emar maier. quotation in context. belgian journal of linguistics, 17(1):109–128, 2005. doi: 10.1075/bjl.17.07geu. 53 maier jeroen groenendijk and martin stokhof. dynamic predicate logic. linguistics and philosophy, 14 (1):39–100, 1991. doi: 10.1007/bf00628304. daniel gutzmann and erik stei. how quotation marks what people do with words. journal of pragmatics, 43(10):2650–2663, 2011. doi: 10.1016/j.pragma.2011.03.010. valentine hacquard. on the event relativity of modal auxiliaries. natural language semantics, 18 (1):79–114, 2010. doi: 10.1007/s11050-010-9056-4. jaakko hintikka. semantics for propositional attitudes. in models for modalities, pages 87–112. reidel, 1969. jerry r. hobbs. coherence and coreference. cognitive science, 3(1):67–90, 1979. doi: 10.1207/ s15516709cog0301 4. julie hunter. reports in discourse. dialogue & discourse, 7(4), 2016. url http://dad. uni-bielefeld.de/index.php/dad/article/view/3695. hans kamp and uwe reyle. from discourse to logic: an introduction to modeltheoretic semantics in natural language, formal logic and discourse representation theory, volume 1. kluwer, 1993. hans kamp, josef van genabith, and uwe reyle. discourse representation theory. in dov gabbay and franz guenthner, editors, handbook of philosophical logic, volume 10, pages 125–394. springer, 2003. david kaplan. demonstratives. in joseph almog, john perry, and howard wettstein, editors, themes from kaplan, pages 481–614. oxford university press, 1989. angelika kratzer. decomposing attitude verbs. 2006. url https://semanticsarchive.net/ archive/dcwy2jkm/. handout. honoring anita mittwoch on her 80th birthday. the hebrew university of jerusalem. david lewis. attitudes de dicto and de se. the philosophical review, 88(4):513–543, 1979. url http://www.jstor.org/stable/2184843. emar maier. mixed quotation: the grammar of apparently transparent opacity. semantics and pragmatics, 7(7):1–67, 2014. doi: 10.3765/sp.7.7. emar maier. quotation and unquotation in free indirect discourse. mind and language, 30(3): 235–273, 2015. doi: 10.1111/mila.12083. emar maier. the pragmatics of attraction: explaining unquotation in direct and free indirect discourse. in paul saka and michael johnson, editors, the semantics and pragmatics of quotation. springer, 2017. url http://ling.auf.net/lingbuzz/002966. emar maier. unreliability and point of view in filmic narration. epistemology and philosophy of science, 59(2):23–37, 2022. doi: 10.5840/eps202259217. emar maier and sofia bimpikou. shifting perspectives in pictorial narratives. sinn und bedeutung, 23, 2019. doi: 10.18148/sub/2019.v23i2.600. 54 attribution and the discourse structure of reports barbara partee. the syntax and semantics of quotation. in s. anderson and paul kiparsky, editors, a festschrift for morris halle, pages 410–418. holt, rinehart and winston, 1973. christopher potts. the dimensions of quotation. in chris barker and pauline jacobson, editors, direct compositionality, pages 405–431. oxford university press, 2007. craige roberts. modal subordination and pronominal anaphora in discourse. linguistics and philosophy, 12(6):683–721, 1989. doi: 10.1007/bf00632602. philippe schlenker. a plea for monsters. linguistics and philosophy, 26(1):29–120, 2003. doi: 10.1023/a:1022225203544. philippe schlenker. context of thought and context of utterance: a note on free indirect discourse and the historical present. mind and language, 19(3):279–304, 2004. doi: 10.1111/j.1468-0017. 2004.00259.x. yael sharvit. the puzzle of free indirect discourse. linguistics and philosophy, 31(3):353–395, 2008. doi: 10.1007/s10988-008-9039-9. robert stalnaker. assertion. in peter cole, editor, syntax and semantics 9: pragmatics, pages 315–332. academic press, 1978. arnim von stechow and ede zimmermann. a problem for a compositional treatment of de re attitudes. in greg carlson and francis pelletier, editors, reference and quantification: the partee effect, pages 207–228. csli, 2005. thomas ede zimmermann. kontextabhängigkeit. in arnim von stechow and dieter wunderlich, editors, semantik: ein internationales handbuch der zeitgenössischen forschung, pages 156– 229. walter de gruyter, 1991. 55 dialogue & discourse 11(2) (2020) 128-149 doi: 10.5210/dad.2020.205 please, please, just tell me: the linguistic features of humorous deception stephen skalicky stephen.skalicky@vuw.ac.nz school of linguistics and applied language studies victoria university of wellington nicholas d. duran nduran4@asu.edu school of social and behavioral sciences arizona state university scott a. crossley scrossley@gsu.edu department of applied linguistics and esl georgia state university editor: amir zeldes submitted 03/2020; accepted 11/2020; published online 11/2020 abstract prior research undertaken for the purpose of identifying deceptive language has focused on deception as it is used for nefarious ends, such as purposeful lying. however, despite the intent to mislead, not all examples of deception are carried out for malevolent ends. in this study, we describe the linguistic features of humorous deception. specifically, we analyzed the linguistic features of 753 news stories, 1/3 of which were truthful and 2/3 of which we categorized as examples of humorous deception. the news stories we analyzed occurred naturally as part of a segment named bluff the listener on the popular american radio quiz show wait, wait. . . don’t tell me!. using a combination of supervised learning and predictive modeling, we identified 11 linguistic features accounting for approximately 18% of the variance between humorous deception and truthful news stories. these linguistic features suggested the deceptive news stories were more confident and descriptive but also less cohesive when compared to the truthful new stories. we suggest these findings reflect the dual communicative goal of this unique type of discourse to simultaneously deceive and be humorous. keywords: deception, humor, lexical semantics, applied natural language processing 1. introduction “you, of course, are going to play the game in which you must try to tell truth from fiction” a person who fibs, lies, or is otherwise untruthful during a conversation possesses a decided interactional advantage: they alone are aware of their deception and can thus use that knowledge to their own benefit. this type of everyday deception ranges from the altruistic white lies used in emotionally close relationships (depaulo and kashy, 1998) to duplicitous speech intended to prevent interlocutors from discovering a truth (gupta and ortony, 2018). however, the intent to deceive need not only serve nefarious or malevolent ends. indeed, sometimes deception can be viewed through a positive lens: as a form of creativity (kapoor and khan, 2017) that may evoke a sense of humor or mirth (dynel, 2011). satirical television news shows, humorous movie spoofs, c©2020 stephen skalicky, nicholas duran, and scott a. crossley this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). linguistic features of humorous deception and radio quiz shows can all employ humorous deception to varying degrees for the purposes of entertainment. it is this type of humorous deception we address in the current study. because deception is typically viewed as an inherently negative conversational act, some research has worked to classify and measure features of deceptive speech so that deceptive speakers may be more easily identified. among the different markers of deceptive conversation, quantifiable linguistic features have proven to be a promising method for the automatic detection of deception (duran, 2009; duran et al., 2010; meibauer, 2018). however, deception in these studies is normally operationalized in the negative sense. this is done by asking research participants to lie or otherwise be purposefully deceptive during prompted interaction (hancock et al., 2008; van swol et al., 2012), or by analyzing examples of deceptive communication (ludwig et al., 2016). as such, the linguistic features associated with deception in these studies are based on deception for the purpose of lying in order to prevent interlocutors from discovering a truth. the purpose of the current study is to examine the linguistic features of creative and humorous deception in a context where all parties are aware of deceptive intent. we employ a wide range of linguistic features related to lexical, syntactic, and affective features of language. by exploring which of these linguistic features (if any) can distinguish humorous deception from truth, we aim to provide a better understanding of the linguistic features of deception as it occurs in a different communicative context with different conversational goals. moreover, though a comparison to prior studies examining malevolent deception and humor, our study can provide insight as to whether similar linguistic strategies are used during two very different types of deception. specifically, we examine humorous deception as it occurs in a popular syndicated radio quiz show in the united states named wait wait. . . don’t tell me!. 2. the wait wait. . . don’t tell me radio show wait wait. . . don’t tell me! (wwdtm) is a weekly radio program produced by a public radio station in chicago (wbez) along with national public radio (npr). npr is a national, non-profit radio broadcasting company in the united states which syndicates wwdtm across the country. while most radio shows on npr take a serious perspective towards reporting and commenting upon news and politics, the purpose of wwdtm is to provide an hour of levity. wwdtm does so through virtue of being a quiz show. with the aid of host peter sagal and quiz judge bill kurtis (who recently replaced long-time judge carl kasell), a pool of recurring panelists and radio callers compete in various trivia games and activities. regular panelists for the show include successful writers, actors, journalists, comedians, and more. the topics of the quizzes are always related to recent news topics which occurred in the prior week. the panelists typically provide humorous banter about the various topics discussed on each episode. in its current version, wwdtm is taped before a live audience on thursdays before being broadcast across the united states every saturday or sunday on various npr stations. 2.1 the bluff the listener game among the different recurring game segments on wwdtm, one game in particular includes deception as a fundamental component. this game, called bluff the listener (btl), usually occurs somewhere in the middle of each wwdtm radio broadcast. during the btl game, the weekly panelists are tasked with bluffing a radio caller who participates via telephone. radio callers are chosen by the producers of wwdtm from a pool of listeners who apply in advance and agree to be 129 skalicky, duran and crossley available during the taping of a particular broadcast. each btl segment begins with small talk between the host (sagal) and the radio caller, wherein sagal typically comments upon the geographic location of the radio caller. after this brief chat, sagal then begins the btl game by informing the participant that it is now time for them to play the “game in which you must try to tell truth from fiction.” sagal then hints at an event which has actually happened, usually with humorous flair, but purposefully leaves out specific details of that event. for example, sagal introduced the 11 january 2020 btl topic as follows: “as long as there has been homework, there have been excuses for not handing it in. the paleontological record shows a t-rex once claimed his arms were too short to fill out the sheets. [laughter] this week, we heard an excuse we had never, though, heard before. it was a pretty good one. our panelists are going to tell you about it. pick the truthful one, you’ll win our prize the wait wait-er of your choice on your voicemail. you ready to play?” (npr, 2020) the three panelists then take turns reading aloud a news story that could potentially be a fuller description of the event mentioned by sagal. the task of the radio caller is to identify which story among the three is the actual, truthful description of the event. if the radio caller correctly guesses the truthful story, they are rewarded by having a member of the show leave a message on their answering machine or voicemail service. the panelist who created the selected story also earns a point: the currency used to determine the overall winner of each wwdtm broadcast. 2.1.1 the bluff the listener news stories the btl news stories are all revised or entirely invented by the panelists. according to the wwdtm frequently asked questions website, panelists are provided with a real news story approximately two days before taping of the episode. while one panelist is charged with presenting the real news story, the other two panelists are asked to create their own fictional news stories which are similar to the real story. the real news stories that are chosen each week are always incredible, difficult-tobelieve scenarios that tend to naturally arouse suspicion. for example, the real news story associated with the homework topic cited above reported on two young canadian snowboarders who burned their homework to keep warm after becoming lost in the wilderness due to a snowstorm. while it is not clear exactly how much content is added by the panelist to the real news story each week, these stories tend to include punchlines at the end which highlight the incredulous and humorous nature of the story. for instance, the panelist presentation of the canadian snowboarders’ news story ended with “no word if their teachers reassigned the homework or just gave them a b for burnt.” the wordplay in this final sentence (“b for burnt”) provides just enough of a humorous flourish to plant a seed of doubt in the radio caller’s mind as to the story’s authenticity and works to reinforce the overall humorous frame of the btl game. the other two news stories are completely invented but structured to conform to the rhetorical genre of a news story. for example, to accompany the canadian snowboarding story, one panelist wrote a fictional report describing how a graduate student studying problem-solving behavior in primates kept awakening to find his laptop missing (which contained his homework). the mystery was solved when it was later discovered the orangutan under study had devised a method for escaping its enclosure and was hiding the laptop under the orangutan’s bed at night. an ironic and humorous effect is thus created through the lack of awareness of the orangutan’s gifted problem-solving (by virtue of escaping and returning each night) on the part of the graduate student. the second 130 https://www.npr.org/programs/waitwait/about/faq.html https://www.npr.org/programs/waitwait/about/faq.html linguistic features of humorous deception fictional btl story included real people. specifically, this story retold an apparent confession by famous rapper snoop dogg, known for his proclivity to smoke marijuana, who admitted he once accidently used his son’s final biology paper to roll a comically large joint during a recording session with his colleagues. according to this fictional story, snoop dogg was plagued with regret and felt obliged to call the school and explain what happened, ultimately donating a large sum of scholarship money as a way of apology. the humor in this story is primarily found in the incongruity between expected parent-child relationship roles and what typically counts as an acceptable excuse for losing one’s homework. 2.1.2 classifying the btl news stories the invented btl news stories are not a case of common, everyday deception associated with lying and untruthfulness as it occurs in regular conversation (duran and fusaroli, 2017; depaulo and kashy, 1998). instead, the fictional btl news stories are an example of specific and purposeful deception: a deception game. the radio callers are faced with a problem they must solve, and they know from the outset of the btl game that two of the panelists are presenting fictional stories. accordingly, the invented btl news stories are also clearly creative products. but, just like deception, these news stories are not examples of everyday, creative language (gerrig and gibbs, 1988). instead, these news stories are more appropriately defined as expert creative products carefully planned in advance. moreover, all of the btl stories violate a listener’s assumptions regarding the plausiblity of the events described as well as reasonable behavior associated with the events. this violation is typically benign enough in that it can be reconciled within a hearer’s understanding of the world, and thus all of the btl stories, both fictional and truthful, contain some elements of humor when analysed via theoretical models of humor, such as benign violation theory (warren and mcgraw, 2016). this is because the unbelievable nature of these stories creates some form of incongruity between expectations and reality, the resolution of which is thought to be an essential component of theories of humor comprehension from almost all theoretical perspectives, including the general theory of verbal humor (gtvh) and other models of incongruity resolution (attardo and raskin, 1991; forabosco, 1992, 2008; ritchie, 2009). thus, it is perhaps best to categorize the dishonest btl narratives as a unique form of humorous deception. with this distinction in mind, the fictional btl stories must work within the textual and communicative contraints of the truthful btl stories, which are all examples of humorous narratives because they describe unbelievable events involving both fictional and non-fictional characters for the purpose of entertainment (chłopicki, 2017). one only need refer to internet curations of strange yet real news, such as yahoo!’s odd news section, reddit’s r/nottheonion community, or internet memes associated with florida man to see similar examples. real news can be strange, and thus even the most outrageous sounding btl stories could possibly be true. therefore, the relative strangeness of an invented btl news story is an essential element and one that cannot be used to reliably identify deception from truth. at the same time, it is the strength of this violation that is played with in the fictional btl stories going too far towards the absurd (or the normal) may give away the fictional nature of the deceptive btl narratives. while the radio callers likely attend to this fine line, this may also affect the linguistic choices of the panelists as they strike a balance among humor and deception. accordingly, just as prior studies of deceptive and humorous language have suggested, it may be the case that more subtle differences in linguistic features can be used to dis131 skalicky, duran and crossley tinguish between deceptive and truthful btl stories. the next section reviews related research into the linguistic features of deceptive and humorous language. 3. linguistic features of humor and deception: prior studies with advances in applied natural language processing (nlp) technology, a wide range of linguistic features can be modelled quantitatively. these include simple measures such as counting the number of words in a document but also include more sophisticated features, such as the sentiment associated with a text (e.g., perceptions of negativity or positivity) or the complexity of a text’s syntactic structures. among the many applications of this approach, one method is to use linguistic features to classify documents associated with different genres, registers, or communicative functions. it is through this approach that some studies have already provided insight into the linguistic features associated with humor and deception. 3.1 linguistic features and deception several studies have used quantitative linguistic measures to classify deceptive from truthful speech. for instance, hancock et al. (2008) recruited 35 pairs of subjects to communicate with one another using a computer messaging program. one participant in each pair was randomly asked to lie about two of the five topics discussed. hancock et al. (2008) used an automatic text analysis program to explore differences in quantifiable linguistic features between the deceptive and non-deceptive conversational turns. the program, linguistic inquiry and word count (liwc), reports on psychological and emotional information associated with specific words used in a text, as well as basic lexical and syntactic information (pennebaker et al., 2001). the hancock et al. (2008) results indicated significant differences between deceptive and non-deceptive language for several linguistic features. specifically, deceptive language contained a greater number of words, a greater number of third-person pronouns, and more words related to senses (e.g., touch, feel) when compared to truthful language. the authors interpreted these findings to suggest deceptive language includes more detail and shifts the focus of the conversation onto the listener. motivated by these findings, duran et al. (2010) reanalyzed the hancock et al. (2008) data using a different text analysis program, coh-metrix (graesser et al., 2004). based on the findings of hancock et al. (2008) and their own theoretical position, duran et al. (2010) focused their analysis on six categories of linguistic features hypothesized to predict deceptive language. these categories were: total word count (measured as total number of words), immediacy of personal involvement (measured through personal pronouns and hedging), specificity of events (measured through temporal words and whquestions), accessibility of meaning (measured through word concreteness, meaningfulness, and familiarity), complexity of language (measured through syntactic complexity and use of negation), and cohesion (measured through argument overlap and lexical similarity). using these categories, duran et al. (2010) constructed a linguistic profile of deceptive language based on linguistic features which significantly differed between the deceptive and truthful language. this profile suggested that while deceptive language used relatively fewer words than non-deceptive language, (different from hancock et al. 2008 based on how the two programs counted words), deceptive language was also more syntactically complex than truthful language. semantically, deceptive language employed words with more accessible meanings and did not introduce new information at the same rate as truthful speech. thus, while not all the categories described by duran et al. (2010) proved to be important distinguishers of deceptive language, their 132 linguistic features of humorous deception study replicated most of the hancock et al. (2008) results and identified new features related to deception. a third study employed linguistic indices from both liwc and coh-metrix using a different dataset (van swol et al., 2012). in this study, research subjects (in pairs) participated in a game of deception wherein one participant was tasked with allocating a sum of money to the other participant. the total sum to be allocated was only known to the participant allocating the money, and the receiver could choose to reject or accept the offer presented to them by the allocator. as such, the participant allocating the money had the option of informing the receiver of the true amount of money or lying (ostensibly to be able to keep more for themselves). van swol et al. (2012) recorded and transcribed the conversations and then coded them for deception. the results for this study separated deceivers into two categories. for participants who simply lied about the total amount of money they were given to allocate, the findings reported a higher frequency of third-person pronouns, swearing, and use of numbers. participants who lied via omission of truth used fewer words and causatives. as such, the van swol et al. (2012) study reported differences for similar linguistic features identified in the hancock et al. and duran et al. studies, but also highlighted how different deception strategies or functions were associated with different linguistic features. 3.2 linguistic features and humor the current body of computational and linguistic research into humor is vast when compared to similar research on deceptive language. incongruity is a central concept to theoretical understandings of humor (attardo and raskin, 1991; forabosco, 2008; ritchie, 2009) and explains how humor is created (e.g., incongruity between linguistic form and meaning) and understood (resolution of two competing interpretations in light of context). computational analyses of jokes, puns, and other forms of humor have worked to identify linguistic features best able to detect structural and semantic incongruity in a humorous text. for example, a computational analysis of humorous one-line jokes found that measures of unique word associations were best able to identify the correct punchlines paired with joke setups (mihalcea et al., 2010). another study investigated humor as it occurred naturally in a corpus of academic essays (skalicky et al., 2016). in this study, skalicky et al. (2016) identified four linguistic features associated with human perceptions of humor. together, these features suggested humorous academic essays to be more descriptive (via a greater incidence of adjectives, adverbs, and adjective predicates), more sophisticated (based on words with lower frequency), and less cohesive (based on lower sentenceto-sentence cohesion) when compared to less humorous academic essays. in this data, the lack of cohesion associated with humorous essays may have worked to signal incongruity associated with humor, differentiate the humorous paragraphs from the academic paragraphs in the essays, or a combination of both. finally, several studies have modelled the linguistic features of humorous satirical news texts from websites such as the onion or satirical product reviews taken from amazon.com (burfoot and baldwin, 2009; mihalcea and pulman, 2007; reyes and rosso, 2011; skalicky and crossley, 2015). differences among these studies in terms of the texts used, linguistic features selected, and statistical models employed have led to a wide range of results with some observable trends. for instance, lexical properties of satirical texts suggest they are generally more negative (mihalcea and pulman, 2007; skalicky and crossley, 2015) and human-centered (mihalcea and pulman, 2007). moreover, satirical news texts contain entities that are typically not discussed together in real news articles 133 skalicky, duran and crossley (burfoot and baldwin, 2009), whereas satirical product reviews construct descriptive, present-tense (yet fictional) narratives (reyes and rosso, 2011; skalicky and crossley, 2015). as a whole, these studies showcase the manner in which quantifiable linguistic features can be used to detect and describe both humorous and deceptive language. the question remains, however, as to whether similar linguistic strategies of deception or humor are employed in communicative contexts in which deceptive intention is known in advance, such as during games of deception. in these situations, the goal of deception is to convince the hearer that something false is actually true, whereas the purpose of most conversational deception is to prevent something true from being discovered (gupta and ortony, 2018). as such, deceivers may rely on specific linguistic strategies when cloaking fiction as truth, and these features may or may nor overlap with the findings from prior linguistic studies of deception. at the same time, the fictional btl stories are designed to elicit humor and therefore may also contain linguistic features designed to present some form of humorous incongruity to the listener. this incongruity may serve two purposes: to signal humor as well as to position the fictional btl story in the “unbelievable-yet-true” genre of news stories. 4. current study the goal of the current study is to consider how a different conversational context may influence the linguistic features associated with deception when it is also being used for the purpose of humor. specifically, we investigate humorous deception as it occurs in a context in which all of the interlocutors are aware that deception is present: a humorous quiz segment named bluff the listener from the popular wait wait. . . don’t tell me! radio show. the following research questions guide this study. 1. are there linguistic features which can distinguish between deceptive and non-deceptive language as it occurs in the wait wait. . . don’t tell me! radio show? 2. if so, how do these results compare to prior research of humorous and deceptive language? 4.1 data collection a corpus of btl stories was collected by manually downloading provided transcripts from newsbank (www.newsbank.com), a curated repository of current and archived media. the corpus comprised 251 different wwdtm episodes broadcast over a ten-year period from 2010 to 2019 (number of episodes per year m = 25.1, sd = 5.44), for a total of 753 different btl stories (502 deceptive stories and 251 truthful stories). the mean number of words for deceptive texts was 206 (sd = 49.3), and the mean number of words for truthful texts was 184 (sd = 48.7). in the corpus, each episode is accompanied with the following information: the date, the text of the three btl stories presented in the episode, the panelist associated with each story, and whether the caller correctly identified the truthful story. there was a total of 50 different panelists in this data, each contributing a different number of btl stories (minimum = 1, m = 15.1, sd = 19.42, maximum = 75). the variation in contributions is an effect of some panelists being regulars on the show, such as paula poundstone (n = 75) or mo rocca (n = 50), whereas other panelists appeared only once or twice in this data. indeed, half of the panelists contributed two or fewer btl stories (n = 25). because the different panelists hailed from different vocations, had differing levels of experience crafting btl stories, and had different numbers of truthful or deceptive stories, it is important to capture this variation in our analysis. 134 https://www.newsbank.com/ linguistic features of humorous deception year of broadcast number of btl episodes caller accuracy 2010 23 47.80% 2011 27 63.00% 2012 30 60.00% 2013 30 70.00% 2014 27 85.20% 2015 27 55.60% 2016 16 62.50% 2017 27 70.40% 2018 29 79.30% 2019 15 80.00% table 1: description of dataset and caller accuracy. the radio callers were able to identify the truthful btl stories 67.33% of the time. however, the data suggests this accuracy varied by year. table 1 displays the number of episodes and overall caller accuracy for each year in the corpus. as can be seen, there is evidence of two positive, linear trends in accuracy; one starting from 2010 and ending in 2014, and a second starting in 2015 and ending in 2019. because the btl stories were all modeled on real news stories at the time, it may be the case that differences in available topics over the years also influenced the deceptive strategies associated with different btl stories which may have influenced this difference in caller accuracy. accordingly, the variation in accuracy over time is also important to model. 4.2 linguistic features linguistic features were collected using a suite of automatic text analysis tools which consolidate a wide range of linguistic features developed for multiple applications. table 2 summarizes these tools and the primary constructs they measure. more information can be found following the appropriate citations or by visiting www.lingusiticanalysistools.org. 4.2.1 initial feature selection our analysis1 started with the full output of linguistic features collected from the text analysis programs reported in table 2. we then trimmed the number of indices down to avoid violating statistical assumptions. to do so, we first removed variables with variance close to zero and variables that had high zero counts. we then kept those variables that showed a significant difference between truthful and deceptive stories using simple t-tests in which the alpha value for significance was set at p < .001. after pruning the data, 263 indices remained. we further culled this list by removing variables which measured different variations of the same construct. for instance, from seance, if two variables measured a similar construct but one version controlled for negation, we only retained the version which controlled for negation. we also opted for versions of variables which measured all of the words or all of the content words in the text (as opposed to just the function words). after removing these variables, 127 variables remained. we lastly checked these remaining 127 linguistic features for multicollinearity using pearson correlations. for any two variables with an absolute cor1. data and code for our analysis can be found on osf.io 135 https://www.linguisticanalysistools.org/ https://osf.io/qdjmn/?view_only=276e228a22294fcdae3cbd5d9c756894 skalicky, duran and crossley program name construct reference tool for the automatic analysis of lexical sophistication 2.0 (taales version 2.8.1) lexical sophistication kyle, crossley, and berger, 2018 sentiment analysis and social cognition engine (seance version 1.2.0) sentiment crossley, kyle, and mcnamara, 2016 tool for the automatic analysis of cohesion 2.0 (taaco version 2.0.4) cohesion and coherence crossley, kyle, and dascalu, 2019 tool for the automatic analysis of syntactic sophistication and complexity (taasc version 1.3.8) syntactic sophistication and complexity kyle, 2016 tool for the automatic analysis of lexical diversity (taaled version 1.3.1) lexical diversity in preparation, see linguisticanalysistools.org table 2: list of text analysis programs used in current study. relation higher then .75, we removed the variable with the highest mean absolute correlation with all of the other variables. through this process we removed an additional 49 linguistic features, bringing the total to 78 remaining linguistic features. 4.2.2 further feature selection: elastic net logistic regressions although we drastically reduced the initial number of linguistic variables, it was still necessary to further reduce the number of remaining variables to avoid overfitting models. we did so by using supervised machine learning classification techniques. specifically, we trained 100 versions of a logistic regression model predicting whether a btl story was truthful or deceptive. for each model, we randomly split the entire btl corpus into a training and test set using a 70% training and 30% test split. we trained our models in r using caret with a glmnet elastic model method. this method works to penalize overfitting by regularizing the coefficients of variables which only contribute a small amount of predictive power. doing so mitigates the effects of overfitting associated with the large number of linguistic features entered into each model. during the training process, each training model was fit based on the results of ten-fold cross-validation with ten repeats, which was done in order to further safeguard the coefficients from overfitting. finally, because the category membership of deceptive and truthful btl stories was unbalanced (i.e., twice as many deceptive as truthful stories), we specified our models to use down sampling which equalized the number of cases in each model. after training each model, we tested the results based on the model’s ability to make predictions for the remaining 30% of the data. the result of this was that each individual btl story in the test data was assigned a percentage chance of how likely that story was to be truthful or deceptive based on the linguistic features included in the elastic net model. accuracy of the predictions was then assessed by comparing the predicted category to the actual category for the truthful and deceptive stories. 136 https://www.linguisticanalysistools.org/taaled.html https://www.linguisticanalysistools.org/taaled.html linguistic features of humorous deception as mentioned above, we repeated this entire training/test process using ten-fold cross-validation 100 times, meaning that we obtained 100 different final models from 100 different random 70/30 splits of the data. each each of these 100 final models were themselves obtained from a ten-fold cross-validation boostrapping process. accuracy predicting the test set for each of the 100 final models ranged from a minimum of .573 to a maximum of .716 (m = .651, sd = .029). for each model, we extracted the top 20 predictors based on the strength of the coefficients in the model (using the varimp function in caret). we counted the number of times each linguistic variable was included in the top 20 predictors for the 100 models, resulting in a feature score for each linguistic variable ranging from 0-100. we then chose features which occurred in the models at least 50% of the time (i.e., included in the top 20 variables of a model at least 50 times). this process identified 14 linguistic variables. these variables comprised the set of linguistic indices we then used in our subsequent predictive models, described below2. table 3 presents an overview of these 14 linguistic features, including their definition and average values for the truthful and deceptive btl stories. in order to better represent these differences, figure 1 visually plots standardized versions of each value using z-scores. for each variable in figure 1, the mean is set to zero, represented by the dashed line. bars above the dashed line represent values greater than the mean, and bars below the dashed line represent values less than the mean. figure 1 thus displays whether the truthful or deceptive btl stories contain higher or lower amounts of each particular linguistic feature. these features can be grouped into several larger categories, described below. 4.2.3 sentiment and semantic groupings sentiment measures affective perceptions associated with words, such as valence (positive/negative emotional associations with words). semantic groupings are clusters of words related to some similar semantic category, such as words related to motivation, cognition, and so on. five variables from table 3 are of this category and include abstract words, strength adjectives, dominance ratings: nouns, time and space words, and vader polarity: adjectives. the abstract words, strength adjectives, and time and space words all belong to lists of semantic categories originally compiled as part of the general inquirer database (stone et al., 1966). abstract words are content words which represent abstract concepts, such as duty and truth. strength adjectives are adjectives which imply strength, such as alert or muscular. the time and space words category includes words related to temporal and spatial meaning, including locations (somewhere) and locative prepositions (above), measurement words (diameter), and words of distance and time (inch, soon). dominance represents affective perceptions of dominance originally collected as part of the affective norms for english words (anew) database (bradley and lang, 1999). perceptions of dominance measure whether a word is associated with something that is in control versus something that is being controlled (leader is more dominant than ache). vader is a sentiment analysis framework specifically designed for social media which takes into account variations in punctuation, emoticon use, and other features of shorter texts to provide a state-of-the-art measure of valence (hutto and gilbert, 2014). the specific measure in this list is the average vader sentiment ratings for the adjectives in each btl story. 2. to clarify, the purpose of this portion of the analysis was to explore which set of linguistic features were consistently chosen as predictive of whether a btl story was deceptive or truthful in the current data. therefore, this part of the analysis was not intended to be confirmatory or related to testing our research questions. 137 skalicky, duran and crossley variable name variable description truth deception m(sd) m(sd) abstract words words representing abstract concepts (higher = more of these words) 0.032 (0.016) 0.036 (0.016) cw repetition: sentence average number of a times any content word repeats in the next two sentences (higher = more repetition) 0.121 (0.055) 0.111 (0.051) vac strength (sd) standard deviation of average strength between constructions and verbs (higher = stronger strength) 0.090 (0.064) 0.106 (0.076) dominance nouns average dominance values for all nouns (higher = more dominant) 5.300 (0.498) 5.420 (0.355) sentence similarity (lda) average semantic similarity between adjacent sentences using lda (higher = more similarity) 0.943 (0.042) 0.934 (0.050) word similarity (lsa) average word similarity for all words in a text using lsa (higher = more word similarity) 0.166 (0.021) 0.170 (0.020) mean length cw ttr (mtld) average span length for content words with an average ttr of .720 (higher = lower average lexical diversity) 87.200 (24.200) 97.600 (23.800) noun complexity (sd) standard deviation of number of dependents per noun subject (higher = more variation in number of dependents) 0.992 (0.304) 1.080 (0.344) positive causal conn. causal connectives with positive sentiment (higher = more of these words) 0.017 (0.011) 0.014 (0.009) cw repetition: text average ratio of any one content word to all words in a text (higher = more content word repetition) 0.174 (0.048) 0.181 (0.047) type-token ratio average variety of lexical items (higher = more variety) 0.624 (0.060) 0.612 (0.058) strength adjectives words representing strength (higher = more of these words) 0.242 (0.133) 0.271 (0.128) time and space words words with spatial or temporal meaning (higher = more of these words) 0.070 (0.025) 0.075 (0.023) vader polarity: adjectives average valence of adjectives (higher = more positive adjectives) 0.194 (0.564) 0.382 (0.533) cw = content words, vac = verb argument construction, sd = standard deviation, lda = latent dirichlet allocation, lsa = latent semantic analysis, ttr = type-token ratio, mtld = measure of textual lexical diversity. variable names as reported by the text analysis tools are included in appendix a . table 3: fourteen linguistic variables included in the top 20 predictors by at least 50% of the elastic net logistic regression models. 138 linguistic features of humorous deception figure 1: standardized average values for the 14 linguistic variables which appeared in at least 50 of the 100 elastic net logistic regression models predicting truthful and deceptive bluff the listener stories. row 1 = sentiment variables, row 2 = complexity variables, row 3 = cohesion and coherence variables. dashed lines set at zero represent mean value for all texts; shaded bars represent change for each individual variable in truthful or deceptive texts. cw = content words, vac = verb argument construction, sd = standard deviation, lda = latent dirichlet allocation, lsa = latent semantic analysis. 4.2.4 lexical and syntactic complexity lexical and syntactic complexity measure the overall sophistication of the words and grammatical constructions employed in a text. five variables from table 3 were related to these constructs: word similarity via latent semantic analysis (lsa), type-token ratio (ttr), mean length of content word type-token ratio (mtld), noun complexity via standard deviation (sd), and verb argument construction (vac) strength via standard deviation (sd). the word similarity (lsa) feature measures the distributional similarity of words across texts. the lsa values used here are pretained values from the touchstone applied science associates inc. (tasa) corpus. the tasa corpus contains over 37,000 texts representing a variety of different genres (günther et al., 2015). thus, our word similarity (lsa) feature measures the distributional similarity of words in the btl stories when compared to the general language corpus (i.e., tasa). in this manner, this feature represents the lexical distinctiveness of vocabulary used in a text because texts with greater distributional similarity will contain words that can be used in a greater number of contexts. higher average lsa values thus reflect lower lexical distinctiveness and less sophisticated vocabulary. the next two variables are a measure of how many different words types are used in a text. typetoken ratio (ttr) is the simplest version and is the number of unique word types divided by the 139 skalicky, duran and crossley total number of word tokens in a text. a higher ttr means a greater amount of different word types exist. the mean length content word ttr (mtld) is a more precise measure of ttr which calculates the average span length of content words (cw) which maintain a ttr above .720 (mccarthy and jarvis, 2010). in this manner, ttr captures global word use in a text, whereas the mtld value captures consistency of ttr throughout a text. noun complexity (sd) is the standard deviation of the average number of dependents attached to noun phrases in a text (e.g., each direct or indirect object attached to a noun phrase would count as a dependent). because this measure is the standard deviation, it captures the rate of variation for this measure. finally, vac strength (sd) is a measure of how strongly verbs and their syntactic dependents (e.g., adverbs or nouns attached to a main verb) are associated based on their frequency of occurrence in a regular english usage. a higher value would suggest verbs are being used with typical or more commonly used syntactic dependents. in the same manner to noun complexity (sd), vac strength (sd) measures the standard deviation and thus the rate of variation for this measure. 4.2.5 cohesion and coherence cohesion and coherence measure the similarity of words across sentences and paragraphs in a text as well as how well ideas in a text are connected. four variables from table 3 were related to these constructs: content word (cw) repetition: sentence, content word (cw) repetition: text, positive causal connectives, and sentence similarity via latent dirichelet allocation (lda). the first two measure the number of repeated content words (i.e., nouns, verbs, adjectives) between adjacent sentences as well as in the overall text. positive causal connectives measures the frequency of occurrence for words which create positive causal links between sentences, such as because and moreover. the sentence similarity (lda) feature uses latent dirichlet allocation to calculate the probability occurrence of latent topics within and across texts. in this case, lda is used to calculate the similarity of sentence topics in a particular text. 4.3 predictive model: generalized linear mixed effects model the distribution of the 14 features as they appear in figure 1 suggests descriptive differences between the truthful and deceptive btl stories. however, it is important to model these potential differences in light of the variation that may be associated with differences in style associated with particular panelists as well as the available pool of news topics each year. to do so, we fit a generalized linear mixed effects model with text category as the outcome variable (truth versus deception with truth as baseline), the 14 linguistic features in table 3 as fixed effects, and panelist and year as random effects3. we used an automatic backfitting algorithm to assess the significance of each fixed effect via model comparisons using relative log-likelihood comparisons with the akaike’s information criterion (aic) as a criterion of model fit (tremblay and ransijn, 2015). the resulting model parameters are reported in table 4. three of the 14 features in table 3 were removed during the model backfitting procedure (abstract words, type-token ratio, and vac strength sd), suggesting that they did not contribute a significant amount of additional explanatory power in light of the remaining variables. although the random effects structure captured the predicted variance associated with different panelists, the random effect of year explained close to zero variance in the model (resulting in a singular fit) and was thus removed. the marginal r2 (fixed effects only) was .182 and the conditional r2 (fixed and random effects combined) was .223 (using 3. glmer(truth or deception ∼ linguistic feature1 + . . . + linguistic feature14 + (panelist) + (year)) 140 linguistic features of humorous deception random effects model effect sizes variance sd marginal r2 conditional r2 panelist 0.239 0.489 .182 .223 fixed effects estimate se z p or 5% 95% (intercept) 0.894 0.139 6.452 < .001 2.445 1.947 3.071 sentiment and semantic groupings strength adjectives 0.206 0.089 2.314 0.021 1.229 1.062 1.424 time and space words 0.250 0.090 2.768 0.006 1.284 1.107 1.490 dominance: nouns 0.293 0.095 3.090 0.002 1.341 1.147 1.567 vader: adjectives 0.349 0.088 3.975 < .001 1.417 1.227 1.637 cohesion and coherence cw repetition: sentence -0.302 0.118 -2.564 0.010 1.352 0.609 0.897 cw repetition: text 0.393 0.117 3.360 0.001 1.482 1.222 1.796 positive causal conn. -0.269 0.089 -3.022 0.003 1.309 0.660 0.885 sentence similarity (lda) -0.247 0.095 -2.590 0.010 1.280 0.668 0.914 lexical and syntactic complexity word similarity (lsa) 0.230 0.094 2.458 0.014 1.259 1.079 1.469 mean length ttr (cw) 0.476 0.108 4.389 < .001 1.609 1.347 1.924 noun complexity (sd) 0.251 0.093 2.701 0.007 1.285 1.103 1.497 dv baseline = truth. or = odds ratio. for ease of interpretation, the or for terms with negative estimates were transformed to positive odds ratios using 1/or. refer to table 3 for variable descriptions. table 4: glmer results predicting truthful and deceptive bluff the listener stories the delta method), which indicates the linguistic variables in this model were able to account for approximately 18% of the variation in text type. 5. discussion the goal of this study was to investigate the linguistic features of creative and humorous deception as it occurs in a unique conversational context wherein all interlocutors are aware that deception is present. to do so, we collected a corpus of truthful and fictional news stories used during a recurring segment of a radio quiz show named bluff the listener (btl), part of the popular american radio show wait wait. . . don’t tell me!. we gathered quantitative measures for a variety of linguistic features related to lexical and syntactic sophistication, sentiment, cohesion, and coherence for the deceptive and truthful btl stories using a suite of automatic text analysis tools. we first identified linguistic features consistently chosen as predictive of text category (deceptive versus truthful) using 100 penalized logistic regressions with cross-validation and down sampling. this process identified 14 linguistic variables as significant predictors of text type, which we then fit into a generalized linear mixed effects model which also took into account variance associated with different authors of the btl news stories as well as the time the story was produced. 141 skalicky, duran and crossley our first research question asked whether linguistic features could be used to distinguish between deceptive and non-deceptive language as it occurs in the btl segment of the wwdtm radio show. our initial results identified 11 linguistic variables as significant predictors of deceptive texts which accounted for approximately 18% of the variance in our data set. these 11 features represent three general categories of linguistic properties: sentiment and semantic groupings, cohesion and coherence, and lexical and syntactic complexity. our second research question asked how the results obtained in answering our first research question compare to similar prior studies of the linguistic features of humorous and deceptive language. below, we provide a discussion of our results in light of these two research questions for each of the three broader categories of linguistic properties described above. 5.1 sentiment and semantic groupings the results for the four significant linguistic variables in this category suggest specific lexical differences between the deceptive and truthful btl stories. texts containing a higher number of adjectives related to perceptions of strength as well as a higher number of time and space words were more likely to be deceptive rather than truthful. adjectives that belong to the strength semantic grouping include words related to physical qualities (athletic, large, healthy), certainty (undeniable, last, most), behavior (nonchalant, steady), performance (perfect, proficient), and other related terms. words that belong to the time and space semantic grouping represent temporal and spatial relations. these include prepositions, adjectives, and verbs related to physical location (around, southern, surround) as well as nouns related to distance and time (kilometer, era). the other two linguistic features in this category were related to sentiment or affect. deceptive texts were associated with language that included higher average dominance ratings for nouns (crown is more dominant than hostage) as well as higher vader scores for adjectives (meaning more positive adjectives). as a whole, these features coalesce to suggest the vocabulary of deceptive btl texts is marked by confident and positive descriptions of entities or actions during specific times and/or in specific locations. this may suggest that deceptive btl stories are therefore more specific or detailed in some regards when compared to the truthful btl news stories. specificity of language was a feature previously investigated by both hancock et al. (2008) as well as duran et al. (2010) and was operationalized as a measure of temporal and spatial words and the number of wh-questions produced. however, specificity was predicted to be lower for deceptive language as a strategy to obscure falsified information with little to no veracity, consistent with prior findings suggesting that liars tend to use language which includes fewer details (depaulo et al., 2003). the reverse trend is seen here in the current data and may be reflective of the very different communicative context associated with the btl quiz game. indeed, because the authors of the deceptive btl news stories are tasked with making the fictional seem plausible, it may be the case that including more specific information and vivid detail lends an air of authenticity to the fictional stories. in this manner, these confident, accurate description mirror linguistic features of humor identified in academic essays as well as satirical product reviews (skalicky and crossley, 2015; skalicky et al., 2016), both of which were found to employ a higher degree of description and certainty. these features may therefore align closer with the humorous than deceptive aspect of the btl stories. 142 linguistic features of humorous deception 5.2 lexical and syntactic complexity complexity was also investigated in prior research of both humor and deception. in terms of deception, duran et al. (2010) included two measures of syntactic complexity: the use of negative connectors and the mean number of words which appear before a verb in each clause. deceptive language contained significantly more of the second of these features, which duran et al. (2010) interpreted to represent the need for deceptive speakers to spend more time formulating their lies and deception on-the-fly (a stalling strategy). much like the previous category, the findings of the current study are different. the three linguistic features in the current study related to lexical and syntactic complexity were the word similarity via latent semantic analysis (lsa), mean length of high type-token ratio for content words (mtld), and standard deviation of noun complexity measures. deceptive btl stories were associated with greater amounts of all three of these features. in terms of word similarity via lsa, the deceptive stories were marked by words which were more strongly related to other words. as such, the vocabulary of the deceptive texts was less sophisticated and less contextually diverse when compared to the non-deceptive texts. additionally, the sentences in the deceptive stories used a greater variety of words for longer spans before repeating words (supported by the lower incidence of content word overlap discussed below), which aligns with the descriptive findings demonstrating the deceptive texts were on average longer than the truthful texts. as such, this form of complexity may in turn reflect a level of exaggerated or fabricated complexity on the part of the deceptive btl news stories. finally, the measure of noun complexity was the standard deviation of the average number of dependents attached to a noun phrase. this means that the noun phrases (and, by extension, embedded phrases, clauses, and whole sentences) in the deceptive btl stories had much greater variation in structure, length, and word types when compared to the truthful btl stories. this does not suggest that the deceptive btl stories were necessarily more or less syntactically complex than the truthful stories, but rather that that deceptive stories were less consistent in their syntactic choices. all together, these features suggest that the deceptive btl stories are on the whole more descriptive and varied in their sentence complexity, which likely reflects a combined strategy of deception (exaggerated detail) and humor (incongruity). 5.3 cohesion and coherence the direction of the coefficients for each of the four individual variables in this category align to suggest two key differences between the deceptive and truthful btl stories for cohesion and coherence. first, the deceptive texts were associated with lower sentence-to-sentence cohesion when compared to the truthful texts. for instance, the measures of lexical overlap and semantic cohesion at the sentence level both suggest the truthful texts more consistently used the same words (content word repetition: sentence) and repeated similar ideas (sentence similarity via latent dirichelet allocation) across adjacent sentences. the truthful texts were also associated with a greater number of positive causal connectives, which link sentences and ideas using words like arise and moreover to signpost additional positive information linked to any particular idea (as opposed to negative causation words such as however). this suggests truthful btl texts were marked by greater overall coherence of ideas because the sentences were more cohesive and more explicitly connected. the second key difference between deceptive and truthful btl texts for cohesion and coherence was demonstrated by the measures of lexical repetition. namely, deceptive texts were marked by 143 skalicky, duran and crossley more repetition of content words in a text overall (content word repetition: text) as contrasted with repetition in adjacent sentences. thus, the deceptive texts were more cohesive at a general, lexical level, but had lower overall coherence among ideas when compared to the truthful texts. this may reflect the manner in which topics and entities are discussed in the two different texts types. for the truthful btl stories, ideas may be typically presented in a progressive manner with more explicit and coherent links among ideas, as befitting a news report. the fictional narratives that comprise deceptive btl stories, however, may lack this level of coherence because they rely more heavily on invented situations. cohesion and coherence have been investigated in prior similar investigations of deceptive texts. for instance, in their results, duran et al. (2010) found deceptive language to be more redundant than truthful language, and therefore more cohesive, which is opposite the findings reported in the current data. duran et al. (2010) suggested higher redundancy associated with deception in their study may have been reflective of a strategy to focus on a small number of ideas as to avoid introducing information which may give away the deceptive intent. at the same time, they argued that higher cohesion may have also reflected the relative lack of links to memorized, prior experiences, and naturally the deceptive conversational turns were forced to rely on the smaller number of ideas conjured in the moment. although not realized in their results, duran et al. (2010) had also predicted that the same lack of memory and experience associated with fictional and deceptive events may lead to less cohesion among ideas in deceptive speech. this hypothesis may explain the results for the btl stories, where the truthful btl stories were more cohesive than the deceptive stories. the truthful btl stories are all adapted from news reports based on actual events with real entities and locations but are also within the genre of strange or unbelievable news. because the deceptive btl stories also must tread the line between the merely absurd and the unbelievable, it may be the case that the introduction of less cohesive ideas and words served to enhance the appropriateness of the deceptive btl news stories for this particular genre (i.e., to create a sense of strange-yet-real news). in other words, the introduction of lower cohesion may actually serve to better align the deceptive btl stories with the unique genre and communicative context of the btl quiz show, but still lacks the cohesion a story based on genuine events can contain. at the same time, this lack of coherence may reflect incongruity associated with humor in the fictional btl texts. 5.4 summary and implications a bird’s eye view of the results paints the deceptive btl stories as fictional narratives which are overly descriptive but also lack higher-level cohesion and coherence. in this manner, the deceptive btl stories are a messy mirror of the truthful btl stories and may represent the individual chaos injected into these stories by each of the different panelist authors for the dual purpose of humor and deception. the panelists are simultaneously competing against each other to earn points while also entertaining the wwdtm audience. thus, there may exist a tension between being purely deceptive and the desire to present a fictional btl story which is humorous and entertaining, and this likely translates into a specific instantiation of deceptive humor seen only in this and similar genres. as mentioned in section 2.1.2, both the fictional and the truthful btl stories engage in some form of incongruity or violation of expectations which can result in humor (warren and mcgraw, 2016; attardo and raskin, 1991). the strength of this violation must be carefully attended to by the authors of the deceptive btl stories. we suggest that the increased lexical descriptiveness and lack 144 linguistic features of humorous deception of higher-level cohesion and coherence in the fictional btl stories may be a linguistic manifestation of the forced, fictional incongruity required in order to meet the demands of this unique genre. in order to create some level of incongruity or violation expected in any btl news story, an invented situation must be constructed and thus the tendency for overly descriptive language in the fictional btl narratives might help to construct this reality for both the audience as well as the panelist. and, as mentioned above, the lack of coherence and cohesion may have been an unconscious linguistic decision to inject a certain level of incongruity that the truthful btl stories naturally possessed. however, our results suggest that this invented, forced incongruity cannot fully replicate what is found in the truthful btl stories, and perhaps successful radio callers are able to attend to these subtle linguistic differences. as the saying goes, truth is stranger than fiction. 6. conclusion and future directions in this study we explored which linguistic features were representative of deceptive and truthful news stories in a specific genre and communicative context: a radio quiz game show where radio callers must pick a truthful story from among three possibilities. we opted for this approach because it was difficult to make theoretical predictions from prior research into deceptive language due to the stark differences in genre and context. our findings support this notion, in that the features predictive of deceptive language in this study did not reflect findings from prior research attempting to catalogue the linguistic features of deception operationalized as lying and/or truth avoidance. however, we believe this is not a contradiction but rather a reflection of how deceptive language adapts to local communicative situation (as does any language use). in other words, because the deceptive btl stories are a very different type of deception, it is not surprising that linguistic features were used differently, likely reflecting the overt humor associated with this form of deception. there are big differences between spontaneous and prepared lies. the underlying cognitive constraints involved will differ, and thus so will their manifestations in language. as such there may be no generic suite of linguistic features related to deception in general, and we believe our current approach provides a roadmap for how linguistic methods can be applied to different examples of deception detection in future studies of this nature. a natural next step for an analysis of this nature would be to model caller accuracy as a variable to measure which of the deceptive stories were more likely to be chosen (incorrectly) as truthful. while our data did contain caller accuracy, we were unable to verify which callers were accurate based on their knowledge of the news cycle (e.g., knowing the truthful story because they had encountered it in the news already or using the internet during the call to check the stories) and which were truly naı̈ve participants. one potential solution to this issue would be to qualitatively analyze the responses by the radio caller for hints as to the strategies used to determine the truthful stories. a further possibility would be to present these stories in a laboratory setting to research participants. another approach for future work in this area would be to consider additional types of humor which rely on varying levels of deception to create or amplify a humorous effect. one potential source of data for such an analysis may be found in the speech of stand-up comedians and other explicit comedic contexts. for example, the late mitch hedberg routinely employed a strategy based on linguistic subversion, wherein expectations with specific phrases and collocations were subverted (e.g., “i used to do drugs. i still do, but i used to, too.”). other examples can be found in jokes which initially retreat from their punchline using phrases such as “just kidding” before 145 skalicky, duran and crossley then doubling down on the initial joke (see skalicky et al. 2015 for a more detailed discussion of these jokes). it could be argued that in both of these cases the audience is partially deceived, and this deception is crucial for incongruity resolution and ultimately humor. because the incongruity in these forms of humor is tied more strongly to violation of expectations at the linguistic level, it would be fruitful to investigate whether these differences are realized in quantifiable linguistic features and how they may differ from the features identified in the current study. a final limitation of our study was the relatively small amount of variance accounted for by the linguistic features used here (approximately 18%). this suggests that there are likely other factors characterizing the deceptive stories which may include linguistic and discourse features not captured in the current analysis as well as features of the individual panelists. our random effects structure suggested the panelists explained an additional 4% of the variance in text category, but future studies might model other aspects of the panelists, such as their vocation, experience on the show, and more. in tandem with the linguistic feature selection process we have described in the current study, this type of additional information offers improvements for any future analyses attempting to distinguish truth from fiction. references salvatore attardo and victor raskin. script theory revis(it)ed: joke similarity and joke representation model. humor-international journal of humor research, 4(3–4):293–348, 1991. margaret m. bradley and peter j. lang. affective norms for english words (anew): instruction manual and affective ratings. the center for research in psychophysiology, university of florida, 1999. clint burfoot and timothy baldwin. automatic satire detection: are you having a laugh? in proceedings of the acl-ijcnlp 2009 conference short papers, page 161–164. association for computational linguistics, 2009. władysław chłopicki. humor and narrative. in salvatore attardo, editor, the routledge handbook of language and humor, pages 143–157. routledge, 2017. bella m. depaulo and deborah a. kashy. everyday lies in close and casual relationships. journal of personality and social psychology, 74(1):63–79, 1998. bella m. depaulo, james j. lindsay, brian e. malone, laura muhlenbruck, kelly charlton, and harris cooper. cues to deception. psychological bulletin, 129(1):74–118, 2003. doi: 10.1037/ 0033-2909.129.1.74. nicholas d. duran. expanding a catalogue of deceptive linguistic features with nlp technologies. in proceedings of the 22nd aaai conference on artificial intelligence, pages 243–248, sanibel, fl, usa, 2009. association for the advancement of artificial intelligence. nicholas d. duran and riccardo fusaroli. conversing with a devil’s advocate: interpersonal coordination in deception and disagreement. plos one, 12(6):1–25, 2017. doi: 10.1371/journal. pone.0178140. 146 linguistic features of humorous deception nicholas d. duran, charles hall, philip m. mccarthy, and danielle s. mcnamara. the linguistic correlates of conversational deception: comparing natural language processing technologies. applied psycholinguistics, 31(3):439–462, 2010. doi: 10.1017/s0142716410000068. marta dynel. a web of deceit: a neo-gricean view on types of verbal deception. international review of pragmatics, 3(2):139–167, 2011. doi: 10.1163/187731011x597497. giovannantonio forabosco. cognitive aspects of the humor process: the concept of incongruity. humor, 5(1/2):45–68, 1992. giovannantonio forabosco. is the concept of incongruity still a useful construct for the advancement of humor research? lodz papers in pragmatics, 4(1):45–62, 2008. doi: 10.2478/ v10016-008-0003-5. richard j. gerrig and raymond w. gibbs. beyond the lexicon: creativity in language production. metaphor and symbol, 3(3):1–19, 1988. arthur c. graesser, danielle s. mcnamara, max m. louwerse, and zhiqiang cai. coh-metrix: analysis of text on cohesion and language. behavior research methods, instruments, & computers, 36(2):193–202, 2004. swati gupta and andrew ortony. lying and deception. in jörg meibauer, editor, the oxford handbook of lying, pages 148–169. oxford university press, 2018. doi: 10.1093/oxfordhb/ 9780198736578.013.11. fritz günther, carolin dudschig, and barbara kaup. lsafun an r package for computations based on latent semantic analysis. behavior research methods, 47(4):930–944, 2015. doi: 10.3758/s13428-014-0529-0. jeffrey t. hancock, lauren e. curry, saurabh goorha, and michael woodworth. on lying and being lied to: a linguistic analysis of deception in computer-mediated communication. discourse processes, 45(1):1–23, 2008. doi: 10.1080/01638530701739181. c. j. hutto and eric gilbert. vader: a parsimonious rule-based model for sentiment analysis of social media text. in eighth international conference on weblogs and social media (icwsm-14), page 216–225, 2014. hansika kapoor and azizuddin khan. deceptively yours: valence-based creativity and deception. thinking skills and creativity, 23:199–206, 2017. doi: 10.1016/j.tsc.2016.12.006. stephan ludwig, tom van laer, ko de ruyter, and mike friedman. untangling a web of lies: exploring automated detection of deception in computer-mediated communication. journal of management information systems, 33(2):511–541, 2016. doi: 10.1080/07421222.2016.1205927. philip m. mccarthy and scott jarvis. mtld, vocd-d, and hd-d: a validation study of sophisticated approaches to lexical diversity assessment. behavior research methods, 42(2):381–392, 2010. doi: 10.3758/brm.42.2.381. jörg meibauer. the linguistics of lying. annual review of linguistics, 4(1):357–375, 2018. doi: 10.1146/annurev-linguistics-011817-045634. 147 skalicky, duran and crossley rada mihalcea and stephen pulman. characterizing humour: an exploration of features in humorous texts. in alexander gelbukh, editor, computational linguistics and intelligent text processing, page 337–347. springer, 2007. rada mihalcea, carlo strapparava, and stephen pulman. computational models for incongruity detection in humour. in alexander gelbukh, editor, computational linguistics and intelligent text processing, page 364–374. springer, 2010. npr. bluff the listener, january 2020. url https://www.npr.org/2020/01/11/ 795534428/bluff-the-listener. james w pennebaker, martha e francis, and roger j booth. linguistic inquiry and word count: liwc 2001, 2001. antonio reyes and paolo rosso. mining subjective knowledge from customer reviews: a specific case of irony detection. in proceedings of the 2nd workshop on computational approaches to subjectivity and sentiment analysis, page 118–124. association for computational linguistics, 2011. graeme ritchie. variants of incongruity resolution. journal of literary theory, 3(2):313–332, 2009. doi: 10.1515/jlt.2009.017. stephen skalicky and s. a. crossley. a statistical analysis of satirical amazon.com product reviews. the european journal of humour research, 2(3):66–85, 2015. stephen skalicky, cynthia m. berger, and nancy bell. the functions of “just kidding” in american english. journal of pragmatics, 85:18–31, 2015. stephen skalicky, c. m. berger, s. a. crossley, and danielle s. mcnamara. linguistic features of humor in academic writing. advances in language and literary studies, 7(3):248–259, 2016. philip j. stone, dexter c. dunphy, and marshall s. smith. the general inquirer: a computer approach to content analysis. mit press, 1966. antoine tremblay and johannes ransijn. lmerconveniencefunctions: model selection and posthoc analysis for (g)lmer models. r package version 2.10, 2015. url https://cran. r-project.org/package=lmerconveniencefunctions. lyn m. van swol, michael t. braun, and deepak malhotra. evidence for the pinocchio effect: linguistic differences between lies, deception by omissions, and truths. discourse processes, 49 (2):79–106, 2012. doi: 10.1080/0163853x.2011.633331. c. warren and a. p. mcgraw. differentiating what is humorous from what is not. journal of personality and social psychology, 110(3):407–430, 2016. 148 https://www.npr.org/2020/01/11/795534428/bluff-the-listener https://www.npr.org/2020/01/11/795534428/bluff-the-listener https://cran.r-project.org/package=lmerconveniencefunctions https://cran.r-project.org/package=lmerconveniencefunctions linguistic features of humorous deception appendix a. complete names of linguistic features output name from program name used in manuscript program abs gi neg 3 abstract words seance adjacent overlap 2 cw sent cw repetition: adjacent sentences taaco all av delta p const cue stdev vac strength (sd) taasc dominance nouns neg 3 dominance nouns seance lda 1 all sent sentence similarity (lda) taaco lsa average all cosine word similarity (lsa) taales mtld ma wrap cw mean length ttr (cw) taaled nsubj nn stdev noun complexity (sd) taasc positive causal positive causal conn. taaco repeated content lemmas cw repetition: text taaco simple ttr aw type-token ratio taasc strong gi adjectives neg 3 strength adjectives seance timespc lasswell neg 3 time and space words seance vader compound adjectives vader polarity: adjectives seance 149 introduction the wait wait…don’t tell me radio show the bluff the listener game the bluff the listener news stories classifying the btl news stories linguistic features of humor and deception: prior studies linguistic features and deception linguistic features and humor current study data collection linguistic features initial feature selection further feature selection: elastic net logistic regressions sentiment and semantic groupings lexical and syntactic complexity cohesion and coherence predictive model: generalized linear mixed effects model discussion sentiment and semantic groupings lexical and syntactic complexity cohesion and coherence summary and implications conclusion and future directions complete names of linguistic features dialogue & discourse 13(2) (2022) 79–132 doi: 10.5210/dad.2022.203 characterizing the response space of questions: data and theory jonathan ginzburg yonatan.ginzburg@u-paris.fr université paris cité, cnrs, laboratoire de linguistique formelle zulipiye yusupujiang zulpiya127@hotmail.com université paris cité, cnrs, laboratoire de linguistique formelle chuyuan li lisa27chuyuan@gmail.com université de lorraine, cnrs, inria, loria, kexin ren kren4@ualberta.ca université paris cité, cnrs, laboratoire de linguistique formelle aleksandra kucharska ale.kucharska96@gmail.com gsk (glaxosmithkline) paweł łupkowski pawel.lupkowski@amu.edu.pl faculty of psychology and cognitive science, adam mickiewicz university reasoning research group editor: massimo poesio submitted 12/2020; accepted 10/2022; published online 12/2022 abstract the main aim of this paper is to provide a characterization of the response space for questions using a taxonomy grounded in a dialogical formal semantics. as a starting point we take the typology for responses in the form of questions provided in łupkowski and ginzburg (2016). that work develops a wide coverage taxonomy for question/question sequences observable in corpora including the bnc, childes, and bee, as well as formal modeling of all the postulated classes. this paper extends that work to cover all types of responses to questions. we present the extended typology of responses to questions based on studies of the bnc, bee, maptask and cornellmovie corpora which include 607, 262, 460, and 911 question/response pairs respectively. we compare the data for english with data from polish using the spokes corpus (694 question/response pairs), providing detailed accounts of annotation reliability and disagreement analysis. we sketch how each class can be formalized using a dialogical semantics appropriate for dialogue management, concretely the framework of kos (ginzburg, 2012). keywords: question, responses, dialogue, corpus study king midas: what is the best thing for humans and the most choice worthy thing of all?’ silenos: why are you forcing me to tell you humans what it would be better for you not to know? (aristotle, eudemian ethics (aristotle, 2012)) c©2022 jonathan ginzburg, zulipiye yusupujiang,chuyuan li, kexin ren, aleksandra kucharska, paweł łupkowski this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). ginzburg, yusupujiang, li, ren, kucharska, łupkowski 1. introduction there are various theories of what questions are (groenendijk and stokhof, 1997; wiśniewski, 2015), and several computational theories of dialogue (poesio and rieser, 2010; asher and lascarides, 2003; ginzburg, 2012), but no attempt yet at a comprehensive characterization of the response space of questions. thus, our aim in this paper is to provide a characterization of the response space for questions using a taxonomy grounded in a dialogical formal semantics.1 this task, nonetheless, is of considerable theoretical and practical importance: it is an important ingredient in the design of dialogue systems, spoken or text–based; it provides benchmarks for dialogue/question theories, and of course is a component in explicating intelligence to pass the turing test (see turing, 1950).2 łupkowski and ginzburg (2013, 2016) tackled one part of this problem, offering an empirical and theoretical characterization of the range of query responses to a query (q-responses). based on a detailed analysis of the british national corpus and three other corpora, two task–oriented, (bee (rosé et al., 1999) and amex (kowtko and price, 1989)) and a sample from childes (macwhinney, 2000), they identified 7 classes of questions that a given query gives rise to; we refer to these classes as the l(upkowski)g(inzburg) classes of query responses. the study sample consisted of 1,466 query/query response pairs. as an outcome the following query responses taxonomy was obtained: (1) cr: clarification requests; (2) dp: dependent questions, i.e., cases where the answer to the initial question depends on the answer to a q-response; (3) motiv: questions about the underlying motivation behind the initial question; (4) no answ: questions whose aim is to avoid answering the initial question; (5) form: questions which consider how to answer the initial question; (6) ind: questions which indirectly convey an answer, (7) ignore: responses ignoring the initial question, but addressing a shared situation—for more details see (łupkowski and ginzburg, 2016, p. 255). we take their work as a starting point and make the following hypothesis: (1)(h) main hypothesis: responses drawn from or concerning the lg query classes plus direct answerhood exhaust the response space of a query. specifically this amounts to the following general types of responses (we present the detailed taxonomy in section 3). 1. question–specific: (a) answerhood; (b) dependent questions (a: who should we invite? b: who is in town?); 2. clarification responses. 3. evasion responses: 1this paper is a substantially extended version of a paper that was presented at sigdial 2019 (“characterizing the response space of questions: a corpus study for english and polish”). it includes a significantly broader review of the literature, the corpus study, manually annotated given the complexity of the categories, includes an additional corpus (the cornell movie corpus) and many more q/r pairs analyzed for english (1,235 vs. 2,240) and polish (205 vs. 694); the discussion of annotation reliability is more extensive; the formal section has been rewritten and expanded considerably and the paper is also accompanied by two appendices covering formal background and the annotation guidelines. 2for the analysis of the turing test as a question-response system see, e.g. (łupkowski and wiśniewski, 2011). 80 characterizing the response space of questions (a) ignore (address the situation, but not the question); (b) change the topic (‘answer my question’); (c) motive (‘why do you ask?’); (d) difficult to provide a response. the hypothesis has to be understood relationally—one is not really interested in the extension of the semantic entities (primarily propositions and questions) that can be given as responses. rather, one is interested in the class each such entity is classified as since that is what determines the subsequent contextual evolution. (2) i do not want to talk about that question. (direct answer to what do you not want to do? evasion answer to where were you last night?). we survey the existing literature in section 2. following this, we provide a description of the proposed taxonomy, in section 3. in the sections that follow we proceed to test our main hypothesis using four corpora in english (bnc (burnard, 2007), bee (rosé et al., 1999), hcrc maptask (anderson et al., 1991), cornellmovie (danescu-niculescu-mizil and lee, 2011)) and one corpus in polish (spokes; pęzik 2014). section 4 discusses respectively the corpora we used and data selected therefrom. section 5 describes our annotation method. the hypothesis achieves wide coverage, as we discuss in section 6; in section 7 we discuss in extensive detail the reliability of the results. in section 8 we consider the requirements on semantic frameworks for a formal characterization of the various classes of the taxonomy. we sketch an account of the different classes in the framework of kos (ginzburg, 2012), building on though departing in some respects from the account developed in (łupkowski and ginzburg, 2016). we point to problems other existing frameworks face in providing a comprehensive account. a concluding section 9 outlines a variety of natural extensions to the work described here. there are two appendices: appendix a offers basic notions from the type logical framework ttr (cooper, 2012, 2023) used in the paper, whereas appendix b provides the annotation guidelines. 2. related work as enfield (2010, p. 2658) points out ‘while the grammatical and information structural properties of questions have received widespread attention in linguistics literature, there has been relatively little attention paid to the relationship between questions and their responses.’ let us start with berninger and garvey (1981) who introduce three terms to refer to a reaction to a question: (1) response, which is any verbal production emitted by a partner following a question; (2) reply, which is a response relevant to the question; and (3) answer—a reply that directly or indirectly provides the missing information. in what follows, the authors introduce their rich taxonomy of possible replies for children conversation in a nursery school. the taxonomy covers six categories: (1) possible answers; (2) indirect answers; (3) confessions of ignorance; (4) clarification questions; (5) evasive replies; (6) miscellaneous. in particular, we find questions as a form of replying to questions among the proposed types (in the form of clarification questions). further replies of this kind may be observed among the proposed sub-types of evasive replies (see 8 and 9—however, they are not as fine-grained as the lg typology of q-responses). these cover the following (see berninger and garvey, 1981, p. 407–408). 81 ginzburg, yusupujiang, li, ren, kucharska, łupkowski 1. selecting own reference in making an assertion: (3) x: where the morrow’s house? y: nope, well the morrow house has sniffles. 2. selecting own reference in rejecting the presupposition of the question: (4) x: what’s his name? y: um pretend that he didn’t have a name. 3. routinely associating question and answer form: (5) x: why: (ellided: should i go to sleep). y: because. 4. temporarily stalling in providing an answer, but acknowledging that question has been heard: (6) x: now what do you want for dinner? y: well. x: hot beef? y: ok, hot beef. 5. challenging questioner to supply answer: (7) x: what is it? y: guess. x: the lights? y: yes, it’s a light yea i know. 6. asking a related question other than for clarification purposes: (8) x: where’s my baby’s food? y: are you ready for your baby’s food? 7. repeating question: (9) x: where’s chrissy? y: where chrissy? 8. rejecting question as stated: (10) x: do you hear the man that is with lisa? y: they’re not with lisa. i’m lisa. 82 characterizing the response space of questions one may observe that the presented categories are co–extensive with the ones mentioned in the introduction to this paper. possible and indirect answers are subsumed by the question-specific: answerhood category. clarification questions correspond directly to the category of clarification responses. and evasive replies and confessions of ignorance fall under our richer category dubbed evasion responses. our proposed typology identifies also other types of question responses that are not tackled by berninger’s and garvey’s proposal. in later work, an interesting typology of question responses was proposed as a result of an extensive 10-language comparative project on question–response sequences in ordinary conversation. the project was carried out from 2007 as the part of the multimodal interaction project at the max planck institute for psycholinguistics—see an overview in (stivers et al., 2010). the study adopted certain restrictions with respect to the questions which were taken into account. in order for a question-response pair to be coded the question had to be a formal question or a functional question. questions seeking acknowledgment, offered in reported speech and requests for immediate physical action were not coded (stivers and enfield, 2010, p. 2621). the coding scheme for the response types presented in (stivers and enfield, 2010, p. 2624) is the following: non-response was coded if the person did nothing in response, directed his/her attention to another competing activity, or initiated a wholly unrelated sequence. non-answer response covers a verbal or visible response that failed to directly answer the question as put. this includes laughter, ‘i don’t know’, initiation of repair (e.g., ‘what?’) or other inserted sequences, gestural responses such as shrugs that do not answer the question. nonanswer responses include also ‘maybe’, ‘possibly’ or responses that deal with the question indirectly (like e.g., a: ‘do you see jack much?’ b: ‘he moved’). answer answers the question directly. answers can be gestural (e.g., a head nod or shake) or verbal (‘uh huh’, ‘yeah’, or longer, more involved answers including partial repeats of the question to confirm or disconfirm). can’t determine can’t hear/see participants, etc. as with the previous typology, one can observe that our categories of question response cover these discussed above. the types which are not covered (like parts of ‘non-response’ or ‘can’t determine categories’) are a consequence of the set-up of the multimodal interaction project, where annotators had video-taped conversations at their disposal. our study is based on a wide range of already existing corpora (without access to video). another interesting issue concerns what constitutes the most frequent type of response. berninger and garvey (1981) observe that the vast majority of responses provided (for polar and for whquestions) were the possible answers. other types were rare: ‘the only other classes of replies that occurred with sizeable frequency were evasive replies and confessions of ignorance following wh-questions and indirect answers following yes/no questions’ (berninger and garvey, 1981, p. 410). analogous results are reported in stivers and robinson (2006) for the group of adult american english speakers. the corpus gathered for the analysis consisted of 260 instances of question sequences in a multi-party interaction (retrieved from video recordings of naturally occurring interactions). in this case the authors do not provide an extensive typology of replies as discussed 83 ginzburg, yusupujiang, li, ren, kucharska, łupkowski above, but focus only on answer / non-answer patterns. the conclusion of the study is that an answer is the alternative preferred over a non-answer (stivers and robinson, 2006, p. 371)—85% of the cases in the analyzed sample were answers. stivers and robinson provide several explanations for such a distribution. one is that the form of a non–answer supplying response turn reflects their ranking as dispreferred (they are frequently delayed both within and between turns, prefaced by filled pauses and discourse markers such as ‘well’, and expanded with accounts—see stivers and robinson 2006, p. 372). moreover, conversational participants typically treat a non-response as indicating disalignment, rather than indicating that no response will be forthcoming. another reason for the obtained distribution, according to stivers and robinson, is that speakers perform interactional work to provide answers and despite the fact non-answers are a readily available alternative category of response—they ‘struggle to receive and provide answers if at all possible’ (stivers and robinson, 2006, p. 374). stivers et al. (2010, p. 2616) point out that ‘in english there is a strong normative order surrounding questions. in the first place, responses are normatively required [. . . ], and answers are preferred over non-answer responses’. this claim is confirmed in the study of 350 questions drawn from spontaneous conversation in american english presented in (stivers, 2010). the results are that 76% of responses were answers, only 19% were non-answers and 5% non-responses (stivers, 2010, p. 2778). this is in line with previous results reported in (stivers and robinson, 2006) discussed above. interestingly, yoon (2010) reports results for korean which though indicative of a similar pattern (answer > non-answer > non-response) indicate a markedly different distribution: of the sample of 326 questions-responses, 52% were answers, 33% non-answers and 15% non-responses (yoon, 2010, p. 2790). in this study, the question sample was limited to questions that functionally sought information, confirmation or agreement (yoon, 2010, p. 2783). enfield et al. (2019) present results of a fourteen-language (including e.g., english, lao, korean) study concerning the issue of how people answer polar questions. the data-set consisted of 172 videotaped interactions. the authors point out that they focus only on answers: “in our quantitative study of responses, we examine only confirming answers (rather than non-answers such as i don’t know, i can’t remember, or laughter; or disconfirming answers). this is because confirmations are more frequent than disconfirmations (...)” (enfield et al., 2019, 288–289); it is worth noting that the non-answer examples acknowledged above are covered by our taxonomy. enfield et al. (2019) conclude that the answers to polar question may be of two possible types: (i) interjectiontype answers (such as ‘uh-huh’ or equivalents ‘yes’, ‘mm’, ‘head nods’, etc.)3 and (ii) repetitiontype answers. wang (2020) uses the proposed taxonomy of polar-question answers in a study of mandarin data, adding a 15th language to the already existing data. another notable source is enfield (2010), who provides an analysis of questions and responses in lao for a corpus of 351 questions drawn from 8 separate recordings. the results reported in this paper are interesting for the discussion of what counts as an answer to a question. the focus of the analysis is the structural fit between questions (wh–questions and polar ones) and their responses. the author offers the following hypothesis as to what answers to wh–questions are optimally coherent: ‘[the answer] should supply a referent of the relevant ontological category (i.e. a thing for a ‘what’ question, a person for a ‘who’ question, etc.).’ (enfield, 2010, p. 2661). green and carberry (1999) provide useful insights into indirect answering. they study 25 dialogue examples originating in (stenstrom, 1984), where 13% responses to polar questions were 3further three sub-types of interjections answers (upgraded, downgraded, and acquiescent) were proposed in (stivers, 2019). 84 characterizing the response space of questions indirect answers. on this basis one can highlight four possible reasons for using indirect answers (see green and carberry, 1999, p. 392). 1. to answer implicit wh-questions: (11) q: isn’t your country seat there somewhere? r: [yes/no]. stoke d’abernon. 2. for social reasons: (12) q: did you go to his lectures? r: [yes.] oh he had a really caustic sense of humour actually. 3. to provide an explanation: (13) q: and also did you find my blue and green striped tie? r: [no.] i haven’t looked for it. 4. to provide clarification: (14) q: i don’t think you’ve been upstairs yet. r: [yes, i have been upstairs.] um only just to the loo. the works discussed in this section indicate the need for a wider corpus study of the whole spectrum of responses to questions. these studies are limited in terms of the examples that were analyzed. they also impose certain limitations concerning the number of response categories to be identified. this is understandable, as their main aim was to explicate the answer/non-answer difference. we believe that an extensive corpus study should bring a fine grained characterization of the entire response space of questions. moreover, we aim at providing an explicit dialogical semantics for each category in our corpus-based typology. one should also acknowledge here the existence of various question answer typologies created within the field of question answering (qa). qa may be characterized as ‘a sophisticated form of information retrieval (ir), in which the system processes questions queried in a natural language format and provides either the content containing the answer or the answer itself’ (shah et al., 2019, p. 611–612). usually, such typologies are proposed for well-structured knowledge bases (or extracted with the use of various nlp methods from the unstructured data sets). what is different from our approach is that these typologies focus on the non-dialogical context of large amounts of texts. as such, qa addresses the interaction between computer systems and users (see shah et al., 2019, p. 612). the resulting answer typology usually takes the form of an ontology of the related data on which a given qa system is to operate—an example of such an approach is presented in (hovy et al., 2002). 85 ginzburg, yusupujiang, li, ren, kucharska, łupkowski response question–specific da dp ind not-question–specific metacomm cr ack evasion cht ignore motiv dpr figure 1: proposed response space of questions 3. a taxonomy of responses to questions our taxonomy with its three main sub-partitions is displayed in figure 1. the classes in red are those that were added by comparison with the query response taxonomy of łupkowski and ginzburg (2016).4 we start by the most general division of question responses to those that are specific to the question asked those that are not, as discussed in the introduction. in the question–specific class we distinguish direct from indirect answers and dependent questions.5 direct answers (da) provide an answer straightforwardly.6,7 this is clearly visible in the following example—b is providing information required by a: (15) a: who is going to check that? b: well i can check it. [bnc: d97, l1226–1227] for indirect answers (ind) one needs to infer an answer from the utterance.8 this is exemplified in (16): (16) a: what is it? a: what’s he done? b: ehm, you know what i’ve said before, eh, eh you’ll get . [bnc: kd5, l175–l177] 4for an explicit presentation of the taxonomy sans answers, see (łupkowski and ginzburg, 2016, p. 256). 5an anonymous reviewer for dialogue and discourse suggests that, directness may be understood as a separate dimension, which is independent from the others. they suggest that any type of response may be presented in direct or indirect manner (not just answers). this is a hypothesis we think is worth testing, though we do not do so in the current paper. 6we give a more explicit characterization of answerhood in section 8; for a thorough, historically based discussion see (wiśniewski, 2015). 7 for the direct answers category we allow for additional sub-categories, which we did not use in the annotation, but which we return to discuss briefly in section 8. these include: (1) no/yes answer to polar questions; (2) simple answers to wh-questions; (3) partial polar answers; (4) partial wh-question answers. 8as with the direct answers category, it is also apt to use the following sub-categories of indirect answers: (i) indirect answers addressing a wh-question; (2) q-widening inds (over-informative answer to a polar question, addressing a more general wh-question). 86 characterizing the response space of questions in (16) a needs to infer the answer to his/her questions from b’s suggestion that this issue has been addressed before. one also encounters ind being a question-response, as in (17), which is rhetorical and in this sense does not need to be answered and indirectly provides an answer to the initial question (q1). (17) a: are you gemini? b: well if i’m two days away from your, what do you think? [bnc: kpa, l3603– l3604] dependent questions (dp) constitute the case where the answer to the initial question (q1) depends on the answer to the query-response (q2), as in: (18) a. a: q1 b: q2 7→ q1 depends on q2 b. a: do you want me to push it round? b: is it really disturbing you? [cf. whether i want you to push it around depends on whether it really disturbs you.] [bnc: fm1, 679–680] the other two remaining super-categories reuse the classes proposed in (łupkowski and ginzburg, 2013, 2016) with some minor renaming. we start with the metacommunicative class, involving clarification responses and acknowledgments. clarification responses (cr) address something that was not completely understood in the initial question (q1)9, like: (19) a: why are you in? b: what? [bnc: kpt, 469–470] some significant consequences this class has for contextual composition is discussed in section 8.2. acknowledgment (ack)—a speaker acknowledges that s(he) has heard and understood the question, e.g. mhm, aha etc.10 (20) a. leeloo: do you know how we say ‘make love’? korben: uh. . . leeloo: . . . hoppi-hoppa [cornell movie, 5963-5965] 9this class contains intended content questions, repetition requests and relevance clarifications—for detailed discussion see e.g. (purver, 2006) or (ginzburg, 2012). 10acknowledgments are much rarer after questions than after assertoric moves, often communicating, as in two of the examples here, hesitation as to how to answer the question; a finer grained scheme might distinguish such cases from “pure” backchannels with a continuative import. 87 ginzburg, yusupujiang, li, ren, kucharska, łupkowski b. a: who’s it for? b: er a: private job or [bnc: kd3, 2063–2065] c. a: what’s that called, the centre line of the earth? b: mm [bnc: f72, 62–63] moving on to evasive question-responses, we mention first the type which addresses the motivation underlying asking q1 (motiv). whether an answer to q1 will be provided depends on a satisfactory answer to q2, as in (21a); (21b) is an instance where the responder offers an answer negatively resolving the motivation issue: (21) a. a: what’s the matter? b: why? [bnc: hmd, 470–471] b. reporter: who did you back prime minister? theresa may: as i said last week none of your business. [the guardian, may 2019] a related class, which was subsumed within motiv in (łupkowski and ginzburg, 2016),11 but which we separate away here involves cases where the speaker states that it is difficult to provide an answer (dpr), points at a different information source, etc. or the speaker states that s(he) does not know the answer. (22) tutor: why? student: i’m not exactly sure. [bee: log-stud29] (23) a: when’s the first consignment of scottish tapes? b: erm don’t know. [bnc: fm2, 1061–1062] another type of evasive question-response is change–the–topic (cht). instead of answering q1, the agent directly provides q2 and attempts to turn the table on the original querier. the original querier is pressured to answer q2 and put q1 aside, as exemplified in (24a) and most explicitly in (24b).12 11this subclass was insubstantial when solely query responses are considered. 12these can occur in text as well: (i) so, in answer to the question: is jeremy corbyn an anti-semite? my response would be that that’s the wrong question. the right questions to ask are: has he facilitated and amplified expressions of anti-semitism? has he been consistently reluctant to acknowledge expressions of anti-semitism unless they come from white supremacists and neo-nazis? will his actions facilitate the institutionalization of anti-semitism among other progressives? sadly, my answer to all of these is an unequivocal yes. [d. lipstadt, antisemitism: here and now, p. 67] 88 characterizing the response space of questions (24) a. a: what we doing in that? a: er b: so did this woman ask you about why you’ve had so many fridays off? [bnc: kny, 1005–1007] b. bbc interviewer: how did singapore handle the pandemic so well? singapore health official: the question should be ‘how did uk not handle it so well?’. bbc interviewer: what do you mean? singapore health official: we followed ‘uk pandemic response protocol’, the uk did not! [twitter 24 may 2021] (25) provide examples of propositional cht, where the response addresses a distinct issue, thereby indicating that this latter is the topic the responder wishes to discuss and not the initial issue: (25) a. a: what’s dolly’s name? b: it’s raining. [bnc: kd4, 110–111] b. kat: you’re amazingly self-assured. has anyone ever told you that? patrick: go to the prom with me! [cornell movie corpus, m6, 839–840] an ignore type of query-response appears when q2 relates to the situation described by q1 but not directly to the initial question. this can be observed in (26). a and b are playing monopoly. a asks a question, which is ignored by b. it is not that b does not wish to answer a’s question and therefore asks q2. rather, b ignores q1 and asks a question related to the situation (in this case, the board game). (26) a: i’ve got mayfair piccadilly, fleet street and regent street, but i never got a set did i? b: mum, how much, how much do you want for fleet street? [bnc: kch, 1503– 1504] see also the following examples: (27) a: just one car is it there? b: why is there no parking there? [bnc: kp1, 7882–7883] 89 ginzburg, yusupujiang, li, ren, kucharska, łupkowski similar examples emerge with propositional responses, as evinced in (28) and (29): (28) a: so does that mean that the ammeter is not part of the series, just hooked up after to the tabs? b: let’s take a step back. [bee, log-stud23] (29) dino velvet : mister welles . . . would you be so kind as to remove any firearms from your person? welles: what are you... ? dino velvet : take out your gun! [cornell movie corpus, 6840–6842] 4. corpus data used for the study in order to test our main hypothesis, we used corpora from two languages: english and polish. 4.1 english: bnc, bee, maptask, cornellmovie the data for english comes from the bnc (burnard, 2007), bee (rosé et al., 1999), maptask (anderson et al., 1991; skantze et al., 2006) and the cornellmovie corpora (danescu-niculescu-mizil and lee, 2011). although both self–answering and multiparty turns figured in the initial development stage of the taxonomy, we restricted attention to two–person dialogue in the study reported here. 607 q-r turns were taken from the bnc, 262 q-r turns from bee, 460 q-r turns from the maptask, and 911 q-r turns from the cornellmovie corpus. the bnc data covers mainly free conversations: initially 864 q-r pairs from bnc were annotated, but after elimination of multi-party segments, 607 q-r pairs were retained. as for bee, 37 undergraduate students with little background in electricity or electronics participated in conversations with a tutor. we randomly selected the students’ numbers (23, 25, 27, 28, 29, and 31) and annotated the dialogues generated between those students and the tutor. in this way, we obtained 262 q-r pairs. the maptask consists of dialogues recorded for a route following task in which one participant directs a second participant along a route in a map, though the route giver and route follower maps are not identical. 297 of the 460 maptask q-r pairs are from the hcrc maptask corpus (anderson et al., 1991), whereas 163 of them are from the higgins pedestrian navigation and guiding project (skantze et al., 2006). the hcrc map task corpus contains 128 dialogues, 64 of which involve eye contact between participants, while the remaining 64 dialogues involved no eye contact. in this study, we chose 28 dialogues, of which 14 with eye contact and 14 without eye contact. the filenames of these dialogues were selected randomly. we annotated all q-r pairs occurring in each dialogue and obtained 297 q-r annotated pairs after eliminating cases involving self–answering, incomplete questions, and overlapping. as for the q-r turns from the higgins project, we annotated six folders (files no.1 -no.6) and in each of them, there are 4 or 5 different dialogue files. as a result, we also annotated 28 dialogues and obtained 163 q-r turns. the cornellmovie corpus is a collection of fictional conversations extracted from raw movie scripts. we annotated all available two-person dialogues from the first 8 movies listed in the corpus, ranging from the movie id m0 to m7, thereby obtaining 911 q-r pairs. this covers various genres such as comedy, romance, adventure, biography, history, action, crime, science fiction, thriller, fantasy, and horror. 90 characterizing the response space of questions table 1: summary of the corpus data used for the study corpus q-r pairs domain bnc 607 free conversations bee 262 tutorial dialogues maptask 460 cooperative task cornellmovie 911 scripted conversations spokes 694 free conversations total 2,934 4.2 polish: the spokes corpus the data used for this study were retrieved from the spokes corpus. the corpus currently contains 247,580 utterances (2,319,291 words) in transcriptions of spontaneous conversations. for the purposes of this paper, two studies were conducted (with two different sets of annotators). for the first study, we selected four files from the corpus (10,244 words). for the second study, 21 files were selected (86,052 words). the files cover casual conversations concerning, e.g., youth, tv shows, children, wine, or travel plans. within each file, the question-response pairs (q-r) were selected manually. in total, we obtained 694 q-r pairs for two studies.13 5. annotation method for the annotation, all the question-response pairs were supplemented with a full context. the guideline for annotators contained explanations of all the classes and examples for each category. moreover, the other category was included. the complete annotation guidelines are presented in appendix b of this paper. english data annotation: the 607 bnc q-r turns used in this study were randomly extracted from the british national corpus (bnc) and manually annotated by one english l1 speaker and two english l2 speakers who have masters degrees in linguistics and underwent several training sessions with one of the authors, a native speaker of english with significant experience in dialogue annotation. among the 607 q-r turns, 334 of them were annotated by the first and second annotators, whereas the remaining 273 q-r turns were annotated by the first and the third annotators. therefore, an inter-annotator study was conducted in two groups: first vs. second annotators, and first vs. third annotators. polish data annotation: the first sample of 205 q-rs was annotated by the main annotator and two other annotators (one of whom has previous experience in corpus data annotation, all annotators were polish native speakers). the annotators received the annotation guidelines and underwent a short training phase based on selected examples. the second sample of 489 q-rs was annotated by the main annotator and two other annotators who are different from that of the first sample (the main annotator remained the same as in the first sample, all annotators were polish native speakers). as 13given that spokes is the sole source we had for polish, we did not restrict attention to two-person dialogue, given that this would have significantly reduced our data set. 91 ginzburg, yusupujiang, li, ren, kucharska, łupkowski in the previous case, annotators received the annotation guidelines and underwent a short training based on selected examples. 6. results the detailed results of the annotation are presented in figure 2. we discuss the annotation reliability in section 7. we also provide additional data for this paper (covering annotated q-r pairs and disagreement cases) which are hosted on the osf web-page (https://osf.io/mq6r7/). 6.1 english in all four cases, the other class is less than 0.5%, hence coverage is above 99.5%. the most frequent response classes in all four corpora are direct answers; the second most frequent class in the bnc is difficult to provide an answer (dpr=7.91%), while in cornellmovie, the next biggest is indirect answers (ind=18.33%), whereas for the maptask and bee these are ignore (6.09% and 3.82% respectively). 6.2 polish the two most frequent classes of responses for spokes are answers: direct ones (da=64.27%) and—much smaller—indirect ones (ind=10.66%). the next two most frequent classes are dpr (stating that a person does not know the answer to the question, or it is difficult to provide one, dpr=7.78%) and utterances ignoring the question asked (questions and declaratives, ignore=6.92%). 6.3 discussion when comparing results for english and polish, it is apparent that the largest category is direct answers (da). also, indirect answers constitute a large group among recognized responses types in both languages. this result is in line with the findings reported by stivers and robinson (2006) and yoon (2010)—summarized in section 2. as might be expected given the previous results presented in łupkowski and ginzburg (2016), the most frequent question-response for english and polish data is the clarification request. what is interesting is the relatively high number of ignoring responses observed for english and polish. in (łupkowski and ginzburg, 2016) we analyzed only question-responses and this type was observed rarely (0.57% for n=1,051 for bnc). this time, ignore has been used also to classify declaratives, which may explain the higher number observed; we discuss a possible semantic explanation for this in section 8, where we suggest that it is in some sense only “weakly evasive”. other evasive responses (relatively) frequent in both languages are cht and dpr. we can also make some comments concerning cross-corpus differences. as we already mentioned in section 4, our bnc and spokes data cover mainly free conversations, while bee and maptask contain task-oriented dialogues. one might expect differences between these dialogue genres. these expectations are indeed fulfilled: the maptask and bee corpora have the highest number of direct answers in our study sample (80.87% and 88.93% respectively). in contrast, for the bnc and spokes corpora these numbers are substantially lower (respectively 69.36% and 64.27%). when it comes to clarification responses, we observe that the numbers are lower for taskoriented corpora than for the bnc and spokes (this is in line with our previous results for bnc and bee, reported in (łupkowski and ginzburg, 2016, p. 256–257)). we also observe that, for the 92 characterizing the response space of questions bnc: 69.36%da maptask: 80.87% bee: 88.93% cornellmovie: 54.67% spokes: 64.27% bnc: 3.95%ind maptask: 4.13% bee: 1.15% cornellmovie: 18.33% spokes: 10.66% bnc: 7.74%cr maptask: 4.13% bee: 0.38% cornellmovie: 3.07% spokes: 4.18% bnc: 7.91%dpr maptask: 1.30% bee: 0.42% cornellmovie: 4.28% spokes: 7.78% bnc: 4.28%ignore maptask: 6.09% bee: 3.82% cornellmovie: 5.38% spokes: 6.92% bnc: 2.64%ack maptask: 2.39% bee: 0.00% cornellmovie: 0.55% spokes: 2.59% bnc: 2.31%cht maptask: 0.65% bee: 1.15% cornellmovie: 11.75% spokes: 3.03% bnc: 1.15%dp maptask:: 0.44% bee: 0.38% cornellmovie: 0.99% spokes: 0.43% bnc: 0.16%motiv maptask:: 0.00% bee: 0.00% cornellmovie: 0.99% spokes: 0.14% 0% 50% 100% figure 2: response types frequency (bnc, n=607; bee, n=262; maptask, n=460; cornellmovie, n=911; spokes, n=694) 93 ginzburg, yusupujiang, li, ren, kucharska, łupkowski evasive response types discussed above, the tendency is analogous, i.e., we observed lower numbers for task-oriented dialogues than for free conversations. for the cornellmovie corpus, which is a collection of fictional, scripted conversations extracted from raw movie scripts, we observe tendencies akin to the bnc and spokes. this is more or less expected, as elite scriptwriters aim for and succeed in mimicking natural conversation. one notable exception is the cht response category (11.75% vs 2.31%–3.03%). one may hypothesize that such an evasive response is especially useful for movie dialogue writers—however, this observation needs further investigation. 7. annotation reliability 7.1 inter-annotator study we conducted the following inter-annotator reliability study on the english bnc and polish spokes corpora, as they were double annotated by multiple annotators. however, other english corpora such as bee, maptask, and cornellmovie were annotated only once. english the reliability of the annotation was evaluated using the κ (carletta, 1996) and α (krippendorff, 2011) coefficients. we used the scikit-learn (pedregosa et al., 2011) data mining and data analysis tool in python with its sklearn.metrics package for calculating cohen’s kappa, and also used the python implementation krippendorff 14 for the calculation of krippendorff’s alpha. in this case, cohen’s kappa for the first and second annotators is 0.7053 (substantial), whereas for the first and third annotators it is 0.6430. krippendorff’s alpha for the first group is 0.7022, while 0.6373 for the second group. all disagreements were then discussed in detail by one of the annotators and the aforementioned author and resolved. as a result, we obtained a gold standard for this bnc annotation task. in addition, as seen in figure 3 and figure 4, we created a confusion matrix for each of these three annotators by comparing their annotations with the gold standard. we also calculated precision, recall, and f-1 measures of each class for all three annotators as shown in table 2. all were calculated by using the data analysis tool scikit-learn in python with its sklearn.metrics package. we can learn the annotation performance of each annotator by investigating the results shown in the confusion matrices in figure 3 and figure 4, as well as from the precision, recall, and f1 scores reported in table 2. for the largest categories, on the whole, da and cr were easy to annotate; ind was more tricky. in more detail: for the first group annotation in english, there are 334 annotated q-r pairs in total, and among them 232 are da, 33 are cr, and 16 are ind. the first annotator correctly annotated 206 of da as da but misannotated 19 of them as ind, 2 as ack, one as cr, 3 as dpr, and one as ignore. therefore, the first annotator obtained a precision of 0.99, recall 0.89, and the f-1 score of 0.94 for the response type da. the second annotator on the other hand, correctly annotated 219 out of 232 actual da as da but misannotated 2 of them as cr, 1 as cht, 7 as ind, 2 as ignore, and 1 as other. therefore, the second annotator gained a precision of 0.98, recall 0.94, and the f-1 score 0.96 for the response type da. as for the response type cr, the first and second annotators obtained a recall score, 0.94 and 0.97 respectively. that is, the first annotator correctly identified 31 out of 33 cr, and only misclassified 2 as ind. the second annotator identified all 32 cr correctly, and only misclassified one as ignore. the precision and f-1 score of cr for the first annotator is 0.94, and 0.94 and 0.96 for the second annotator. regarding the response type ind, the first annotator correctly annotated 15 out of 16 as ind but misclassified 1 14https://pypi.org/project/krippendorff/ 94 characterizing the response space of questions (a) (b) figure 3: english first group confusion matrices: (a) first annotator (b) second annotator of them as da. the precision, recall, and f-1 score of ind for the first annotator are 0.37, 0.94, and 0.53 respectively. the second annotator correctly identified 12 out of 16 as ind but misclassified 4 of them as da. the precision, recall, and f-1 score of ind for the second annotator are 0.55, 0.75, and 0.63 respectively. as shown in the figure 4 and table 2, in the second group of annotation for english, there are 273 annotated q-r pairs in total, and among them, there are 189 da, 8 ind, 14 cr, 23 dpr, and 17 ignore. the first annotator correctly annotated 175 da as da but misclassified 13 as ind, and one as ignore. the precision, recall, f-1 scores of da for the first annotator are 0.99, 0.93, and 0.96 respectively. the third annotator correctly identified 165 out of 189 da but misclassified 21 of them as ind, 1 as ignore, and 2 as cht. the third annotator obtained a precision score of 0.93, recall 0.87, and f-1 score of 0.90. as for the response type cr, the first annotator correctly annotated 9 cases, but misidentified 3 as ignore, and 2 as ind. the precision, recall, and f-score for cr are 1.00, 0.64, and 0.78 respectively for the first annotator. the third annotator also performed similarly in the classification of cr, and she correctly annotated 8 out of 14 cases, but missclassified 5 as da, and 1 as other. the third annotator obtained similar performance scores as the first annotator, which are 1.00, 0.57, and 0.73 for precision, recall, and f-1 scores respectively. as to the response type ind, 5 out of 8 cases were annotated correctly by the first annotator. however, there are 2 cases misclassified as da, one as dpr. the precision, recall, and f-1 score obtained by the first annotator are 0.19, 0.62, and 0.29 respectively. the annotation of ind also caused more difficulties to the third annotator. she correctly identified 4 out of 8 ind cases but misclassified 3 as da, and 1 as ignore. the performance scores of the third annotator for the response type ind are 0.13, 0.50, and 0.21 respectively for precision, recall, and f-1 score. the annotation performance of both annotators on the response types ignore are similar. however, the f-1 scores of cht are 0.95 and 0.70 for the first and the third annotator respectively, 1.00 and 0.00 for the response type dp. in addition, the annotators’ performance of the second group is better than the first group in terms of the annotation of response classes dpr and ack. polish the reliability of the annotation for polish was also evaluated using the κ (carletta, 1996) and α (krippendorff, 2011) coefficients. as mentioned above, the main annotator was the same person in both samples. however, other annotators were different in these two annotation groups. the reported values were calculated using the same method and tools as for english. for the first 95 ginzburg, yusupujiang, li, ren, kucharska, łupkowski table 2: detailed annotation report for english annotators annotator classes precision recall f1-score frequency da 0.99 0.89 0.94 232 ind 0.37 0.94 0.53 16 cr 0.94 0.94 0.94 33 dpr 0.85 0.68 0.76 25 first group ignore 0.89 0.89 0.89 9 first annotator ack 0.57 0.89 0.70 9 cht 0.75 0.75 0.75 4 dp 0.50 0.40 0.44 5 motiv 1.00 1.00 1.00 1 accuracy 0.87 334 macro avg. 0.76 0.82 0.77 334 weighted avg. 0.92 0.87 0.89 334 da 0.98 0.94 0.96 232 ind 0.55 0.75 0.63 16 cr 0.94 0.97 0.96 33 dpr 1.00 0.60 0.75 25 first group ignore 0.33 0.67 0.44 9 second annotator ack 0.50 0.44 0.47 9 cht 0.43 0.75 0.55 4 dp 1.00 0.20 0.33 5 motiv 1.00 1.00 1.00 1 other 0.00 1.00 0.00 0 accuracy 0.88 334 macro avg. 0.67 0.73 0.61 334 weighted avg. 0.92 0.88 0.89 334 da 0.99 0.93 0.96 189 ind 0.19 0.62 0.29 8 cr 1.00 0.64 0.78 14 dpr 0.91 0.91 0.91 23 second group ignore 0.73 0.65 0.69 17 first annotator ack 1.00 0.86 0.92 7 cht 0.91 1.00 0.95 10 dp 1.00 1.00 1.00 2 other 1.00 1.00 1.00 3 accuracy 0.89 273 macro avg. 0.86 0.85 0.83 273 weighted avg. 0.94 0.89 0.91 273 da 0.93 0.87 0.90 189 ind 0.13 0.50 0.21 8 cr 1.00 0.57 0.73 14 dpr 1.00 0.78 0.88 23 second group ignore 0.67 0.71 0.69 17 third annotator ack 1.00 1.00 1.00 7 cht 0.70 0.70 0.70 10 dp 1.00 0.00 0.00 2 other 0.75 1.00 0.86 3 accuracy 0.82 273 macro avg. 0.80 0.68 0.66 273 weighted avg. 0.89 0.82 0.84 273 96 characterizing the response space of questions (a) (b) figure 4: english second group confusion matrices: (a) first annotator (b) third annotator sample, the best inter-annotator κ and α scores were achieved by the second and the main annotators, 0.6579 and 0.6582 respectively. while for the second sample, we observed the highest inter-annotator agreements between the first and the main annotators, which are 0.5467 and 0.5466 for κ and α, as shown in table 3. all disagreements were discussed in detail by the main annotator and resolved. in addition, we also used the data analysis tool scikit-learn in python with its sklearn.metrics package to created a confusion matrix for each of five annotators by comparing their annotation with the gold standard, as well as calculated the precision, recall, and f-1 measures of each response type. the annotation performance of each polish annotator is presented in detail in table 4, table 5, and figure 5. the frequency of each response type for the polish first group annotation is da:107, ind:29, cr:9, dpr:22, ignore:23, ack:3, cht:11, dp:1, and motiv: 0. regarding the most frequent response type da, the first annotator correctly annotated 87 out of 107 da cases but misclassified 13 as ind, 3 as ignore, 2 as dpr, 1 as cr, and 1 as other. as a result, the first annotator obtained a precision of 0.94, recall 0.81, and f-1 score of 0.87. the second annotator correctly identified 6 more da cases than the first annotator but also misannotated 5 as ind, 3 as cht, 1 as cr, 2 as ignore, and 1 as dpr. the performance scores are also very close to those of the first annotator, which are 0.94, 0.89, and 0.91 for precision, recall, and f-1 score respectively. as to the response type ind, the first annotator correctly annotated 21 out of 29 ind cases but misclassified 3 as da, and other 5 as cht, cr, ignore, motiv, and other respectively. the precision, recall, and f-1 score of ind for the first annotator are 0.48, 0.72, and 0.58. the second annotator correctly annotated 18 out of 29 ind cases but misclassified 5 as ignore, 2 as cht, and other 4 as da, dp, dpr, and other respectively. the precision, recall, and f-1 scores for the second annotator are 0.75, 0.62, and 0.68. regarding the response type cr, the first annotator successfully identified 5 cases, whereas the second annotator identified 7. the f-1 scores for the first and the second annotator are 0.59 and 0.82 respectively. as for ignore, the first annotator correctly identified only 12 cases out of 23, whereas the second annotator correctly classified 21. the f-1 scores of the response type ignore for the first and the second annotator are 0.53 and 0.78. comparing the f-1 scores for each response type, we learned that the second annotator performed better than the first annotator in general. when it comes to the polish second group annotation, there are 489 annotated q-r pairs in this sample. the frequencies of each response type are da:339, ind:45, cr: 20, dpr:32, ignore:25, ack:15, cht:10, dp:2, and motiv:1. as for the response type da, the first annotator correctly 97 ginzburg, yusupujiang, li, ren, kucharska, łupkowski annotated 328 out of 339 da cases but misclassified 4 as dpr, 2 as cht, 2 as cr, and the remaining 3 as dp, ind, and ignore. the precision, recall, and f-1 score for the first annotator are 0.93, 0.97, and 0.95 respectively. the second annotator, on the other hand, correctly identified 266 out of 339 da cases. the second annotator misclassified 33 da cases as ind, 13 as dpr, 14 as ignore, 10 as cht, 2 as ack, and 1 as cr. as a result, the second annotator obtained a precision of 0.99, recall 0.78, and f-1 score of 0.88. regarding the response type ind, the first annotator correctly annotated 34 out of 45 ind cases but misclassified 10 as da, 1 as ch, and obtained a precision of 0.85, recall 0.76, and f-1 score of 0.80. the second annotator correctly identified 38 out of 45 ind cases, but misannotated 7 as ignore. the precision, recall, and f-1 score for the second annotator are 0.49, 0.84, and 0.62 respectively. as to the response type cr, the first annotator correctly annotated 9 out of 20 cr cases but failed to identify the remaining 11 cases. the precision, recall, and f-1 scores for the first annotator are 0.82, 0.45, and 0.58. the second annotator correctly annotated 11 out of 20 cr cases but misclassified 2 as da, 3 as ind, and 4 as ignore. the precision, recall, and f-1 scores for the second annotator are 0.92, 0.55, and 0.69. regarding the response type ignore, the first annotator correctly identified 14 out of 25 ignore cases but misclassified 6 as cht, 3 as da, and 2 as ind. the precision, recall, and f-1 scores are 0.82, 0.56, and 0.67 for the first annotator. whereas the second annotator correctly annotated 23 out of 25 ignore cases and only misclassified 2 of them as cht. even though the second annotator obtained a high recall of 0.92, he has a low precision and f-1 score, which are 0.43 and 0.58 respectively. in addition, the first and the second annotators performed similarly on the annotation of the response types dpr, ack, cht, and motiv. however, as for dp, the second annotator obtained 1.00 for all the precision, recall, and f-1 scores, and the first annotator obtained 0.40, 1.00, and 0.57 respectively. as for the performance of the main annotator in both groups of annotation samples, he outperformed all the other annotators in most of the cases. however, in the first group of samples, the main annotator failed to correctly capture the response type dp, which has only one case in this sample. as for all other response types, he obtained very high f-1 scores, which are above 0.90 in most cases, and 0.83 and 0.87 for ind and cht respectively. while in the second group of samples, he did not perform as well as the other annotators regarding ack. he only obtained an f-1 score of 0.55, while the other two obtained 0.97 and 0.94 respectively. in addition, he also only got an f-1 score of 0.50 for the response type dp. what’s more, the first annotator outperformed the main annotator also on the annotation of ind, where the first annotator obtained an f-1 score of 0.80, whereas it is 0.73 for the main annotator. table 3: polish inter-annotator agreement annotation group annotators cohen’s kappa krippendorff’s alpha first annotator vs. second annotator 0.4588 0.4574 first group first annotator vs. main annotator 0.5121 0.5117 second annotator vs. main annotator 0.6579 0.6582 first annotator vs. second annotator 0.5414 0.5334 second group first annotator vs. main annotator 0.5467 0.5466 second annotator vs. main annotator 0.4738 0.4648 annotation reliability on the subsets of taxonomy 98 characterizing the response space of questions (a) (b) (c) (d) (e) (f) figure 5: polish confusion matrices: (a)-(c) first group annotators; (d)-(f) second group annotators 99 ginzburg, yusupujiang, li, ren, kucharska, łupkowski table 4: detailed annotation report for polish first group annotators annotator classes precision recall f1-score frequency da 0.94 0.81 0.87 107 ind 0.48 0.72 0.58 29 cr 0.62 0.56 0.59 9 first group dpr 0.81 0.77 0.79 22 first annotator ignore 0.55 0.52 0.53 23 ack 1.00 0.00 0.00 3 cht 0.50 0.27 0.35 11 dp 0.25 1.00 0.40 1 motiv 0.00 1.00 0.00 0 other 0.00 1.00 0.00 0 accuracy 0.71 205 macro avg. 0.51 0.67 0.41 205 weighted avg. 0.77 0.71 0.73 205 da 0.94 0.89 0.91 107 ind 0.75 0.62 0.68 29 cr 0.88 0.78 0.82 9 first group dpr 0.90 0.82 0.86 22 second annotator ignore 0.68 0.91 0.78 23 ack 1.00 0.67 0.80 3 cht 0.54 0.64 0.58 11 dp 0.25 1.00 0.40 1 other 0.00 1.00 0.00 0 accuracy 0.82 205 macro avg. 0.66 0.81 0.65 205 weighted avg. 0.85 0.82 0.83 205 da 0.98 0.96 0.97 107 ind 0.81 0.86 0.83 29 cr 0.90 1.00 0.95 9 dpr 0.95 0.95 0.95 22 main annotator ignore 1.00 0.87 0.93 23 ack 1.00 1.00 1.00 3 cht 0.83 0.91 0.87 11 dp 0.00 0.00 0.00 1 accuracy 0.93 205 macro avg. 0.81 0.82 0.81 205 weighted avg. 0.94 0.93 0.93 205 100 characterizing the response space of questions table 5: detailed annotation report for polish second group annotators annotator classes precision recall f1-score frequency da 0.93 0.97 0.95 339 ind 0.85 0.76 0.80 45 cr 0.82 0.45 0.58 20 second group dpr 0.87 0.81 0.84 32 first annotator ignore 0.82 0.56 0.67 25 ack 1.00 0.93 0.97 15 cht 0.28 0.50 0.36 10 dp 0.40 1.00 0.57 2 motiv 1.00 1.00 1.00 1 accuracy 0.89 489 macro avg. 0.77 0.78 0.75 489 weighted avg. 0.89 0.89 0.88 489 da 0.99 0.78 0.88 339 ind 0.49 0.84 0.62 45 cr 0.92 0.55 0.69 20 second group dpr 0.68 0.84 0.75 32 second annotator ignore 0.43 0.92 0.58 25 ack 0.88 1.00 0.94 15 cht 0.28 0.50 0.36 10 dp 1.00 1.00 1.00 2 motiv 1.00 1.00 1.00 1 accuracy 0.79 489 macro avg. 0.74 0.83 0.76 489 weighted avg. 0.88 0.79 0.81 489 da 0.92 0.95 0.94 339 ind 0.83 0.64 0.73 45 cr 0.69 1.00 0.82 20 dpr 0.85 0.88 0.86 32 main annotator ignore 0.71 0.68 0.69 25 ack 0.86 0.40 0.55 15 cht 0.88 0.70 0.78 10 dp 0.50 0.50 0.50 2 motiv 1.00 1.00 1.00 1 accuracy 0.88 489 macro avg. 0.80 0.75 0.76 489 weighted avg. 0.88 0.88 0.88 489 101 ginzburg, yusupujiang, li, ren, kucharska, łupkowski the previous inter-annotator reliability study was carried out on the full taxonomy of question responses. however, we also performed inter-annotator reliability tests on several subsets of the taxonomy, to learn which subsets of the taxonomy can be reliably annotated. we also used cohen’s kappa score for this task. english the detailed cohen’s kappa scores on different subsets of the taxonomy for english are presented in table 6. as shown in the table, the response types da, cr, ack, dpr, dp, motiv were annotated with almost perfect agreement level (above 0.9) (mchugh, 2012) between annotators in both groups of the experiment. however, the response types such as ignore, cht, ind caused a sharp decrease in the agreement level. the indirect answer (ind) is the one that drops the agreement level between annotators significantly. table 6: english inter-annotator reliability on subsets of the taxonomy, cohen’s kappa score subset of taxonomy 1st vs. 2nd 1st vs. 3rd da, cr 0.9816 1.0 da, cr, ack 0.9710 1.0 da, cr, ack, dpr 0.9681 0.9489 da, cr, ack, dpr, dp 0.9686 0.9489 da, cr, ack, dpr, dp, motiv 0.9692 0.9489 da, cr, ack, dpr, dp, motiv, ignore 0.8973 0.8755 da, cr, ack, dpr, dp, motiv, ignore, cht 0.8739 0.8391 da, cr, ack, dpr, dp, motiv, ignore, cht, ind 0.7183 0.6358 da, cr, ack, dpr, dp, motiv, ignore, cht, ind, other 0.7052 0.6430 polish the agreement level among annotators on different subsets of the taxonomy for two groups of polish annotation are displayed in table 7 and table 8 respectively. comparing the overall results on two tables, we found that the agreement level among the annotators in the first group is generally higher than that of the second group. in the first group, the response types da, cr, ack, dpr, dp, motiv were annotated with a strong agreement level (0.8–0.9) (mchugh, 2012) between first and the main annotators, and also between the second and the main annotators. however, those response types were annotated with a moderate agreement level (0.60–0.79) (mchugh, 2012) between the first and the second annotators. as for the second group in table 8, the response types da, cr, ack, dpr, dp, motiv were annotated with a moderate agreement level (0.60–0.79) nearly among all annotators. in both groups, the agreement level dropped evidently when ignore, cht, ind were added. to sum up, response types such as da, cr, ack, dpr, dp, motiv can be reliably annotated by all annotators in both languages, whereas the response types such as ignore, cht, ind cause more confusion to the annotators. among all response types, the indirect answer (ind) is the one that is most difficult to annotate. disagreement analysis for english: among the commonly annotated 607 bnc q-rs, there are 108 cases where annotation disagreements between two annotators occurred as shown in table 9. the main disagreements concerned da versus ind (52), ignore versus cht/ack/dp/da/dpr/ind (33), and ack versus 102 characterizing the response space of questions table 7: polish first group inter-annotator reliability on subsets of the taxonomy, cohen’s kappa score subset of taxonomy 1st vs. 2nd 1st vs. main 2nd vs. main da, cr 0.7882 0.8214 0.8074 da, cr, ack 0.7882 0.8214 0.8010 da, cr, ack, dpr 0.7855 0.8343 0.8781 da, cr, ack, dpr, dp 0.7582 0.8238 0.8449 da, cr, ack, dpr, dp, motiv 0.7582 0.8238 0.8449 da, cr, ack, dpr, dp, motiv, ignore 0.6867 0.7515 0.8498 da, cr, ack, dpr, dp, motiv, ignore, cht 0.6360 0.6957 0.7863 da, cr, ack, dpr, dp, motiv, ignore, cht, ind 0.4810 0.5315 0.6662 da, cr, ack, dpr, dp, motiv, ignore, cht, ind, other 0.4588 0.5121 0.6579 table 8: polish second group inter-annotator reliability on subsets of the taxonomy, cohen’s kappa score subset of taxonomy 1st vs. 2nd 1st vs. main 2nd vs. main da, cr 0.7525 0.6694 0.6522 da, cr, ack 0.8351 0.5901 0.5779 da, cr, ack, dpr 0.7652 0.6651 0.6399 da, cr, ack, dpr, dp 0.7612 0.6661 0.6406 da, cr, ack, dpr, dp, motiv 0.7648 0.6712 0.6462 da, cr, ack, dpr, dp, motiv, ignore 0.7040 0.6047 0.6429 da, cr, ack, dpr, dp, motiv, ignore, cht 0.6220 0.5604 0.6025 da, cr, ack, dpr, dp, motiv, ignore, cht, ind 0.5414 0.4738 0.5467 other/da/dpr/cht (5), as exemplified in (30). invariably, the direct/indirect disagreements occurred with ‘why’, ‘how’ and ‘what is x doing’ questions, where answers are by and large sentential and for which there has been significant controversy in the theoretical literature on how to characterize answerhood (kuipers and wiśniewski, 1994; asher and lascarides, 1998). table 9: disagreement cases for english disagreement types frequency disagreement types frequency da-ind 52 da-cr 1 da-ignore 8 ignore-cht 7 da-dpr 4 ignore-dpr 3 da-cht 2 ignore-ack 2 ind-cr 3 ignore-dp 3 ind-dp 2 ack-da 1 ind-dpr 3 ack-cht 1 ind-ignore 9 ack-other 3 ind-cht 2 cr-other 2 sum 108 103 ginzburg, yusupujiang, li, ren, kucharska, łupkowski (30) a. anon 1: when did the bus service start to then? mansie flaws: oh it was a while after we started. [da vs. ind, resolved to ind] b. ann: that’s not very nice. stuart: it is. ann: no it isn’t. stuart: well it is. why isn’t it? ann: cos it isn’t. [da v. ignore, resolved to da.] c. john: so lock erm how would you spell sock? simon: smelly er smelly [ignore v. cht, resolved to ignore.] john: how would you smell sock then? d. john: can you spell box? simon: mhm. [ack v. other, resolved to da, after consideration of surrounding context.] in the above conversations, (30a) is an example of da versus ind, where the first annotator categorized it as ind, while the second person annotated it as da. after discussion, we decided to classify it as ind given that a certain amount of inference is needed to know the exact time of the bus service. for (30b), the first annotator annotated it as ignore, while the second annotator marked it as da, however, after discussion, we decided that it should be categorized in da since the response emphasizes the fact that “because it is actually not nice ”. for (30c), the first annotator annotated the answer as ignore, while the second person categorized it as cht, and after discussion, we keep ignore as the correct annotation since the answer is also related to the main topic “sock”. (30d) is an example of ack versus other, where the first annotator annotated it as other, while the second annotator treated it as ack. however, as a result of considering the surrounding context, we concluded that it is actually a direct answer to the question. for polish: for the whole annotated sample, we observed 41 cases with disagreement between all three annotators (as shown in table 10). the main disagreements concerned da versus dpr (12), which is a notable difference by comparison with the english data.15 we also observed some da versus ind disagreements but much less common (4). it is also the case that the ignore category appears often in the disagreements summary (versus da, cr, ind, cht, and dpr). among the analyzed disagreement cases, two are especially interesting as the disagreement of all annotators is observed for consecutive turns in a dialogue. the first problematic case is for [016o, 62–65]. a and b are discussing b’s application for a scholarship. (31) a: a w tej twojej szkole ty jako twoja kandydatura została złożona tylko czy jeszcze jakiś innych osób też [and in your school it is you you are the only candidate or maybe there are some other people who also applied] 15we hypothesize that the reason for this may be the background of annotators as logicians. from a logical perspective the exhaustiveness of an answer is important (see e.g. wiśniewski, 2013). thus, certain partial answers provided by dialogue participants were probably tagged as dpr. this may be due to the fact that partial answers were not explicitly pointed out in the guidelines. 104 characterizing the response space of questions table 10: disagreement cases for polish (without the main annotator) disagreement types frequency disagreement types frequency ack-cht 1 cht-ind 2 da-ind 4 ignore-da 1 da-dpr 12 ignore-cr 1 da-cr 3 ignore-cht 1 da-cht 3 ignore-ind 1 cr-dp 1 ignore-dpr 2 cr-cht 1 ind-dpr 1 cr-dpr 1 dp-cr 1 cr-ind 1 other-cht 1 cr-cht 3 sum 41 b: co ty nie no nie dowiadywałam się wiem że z mojej grupy tylko ja jestem [oh stop i didn’t check i know only that from my group it was only me] a: masz konkurencję [so you have some competition] b: yyy z całej tej szkoły ? nie wiem na przykład od marty mogłabym się marty zapytać właściwie bo od marmarta nie chciała jechaćwłaściwie to nie wiem dlaczego ale już aż mi było ty ja ją tak namawiałam tak ją prosiłam potem ona i tak tak wiesz to zlała nie wiem dlaczego nie chciała pojechać [yyy from the whole school? i don’t know for example martha actually i could ask martha because from marmartha didn’t want to go actually i do not know why but for me it was you know i have tried to convince her i have asked her and then she after all you know, she just ignored it i do not know why she didn’t want to go] in this case, the disagreement between annotators was whether the first b’s utterance should be classified as ‘it is difficult to provide an answer’ (dpr) or as an indirect answer (ind). as for the second b’s utterance, the suggested types were dpr and ignore. another example where the disagreement was observed for two consecutive utterances is [01ao, 256–259]. most probably, this is caused by the fact that four participants took part in this dialogue (which makes an interpretation of question responses much more difficult). (32) b: ciekawe ile kasy dają [i am wondering how much money they can give you] c: ciekawe ile kasy dają [i am wondering how much money they can give you] a: no dawają ci tyle co na tym na [well they give you the same that in that] d: w sklepie w kerfurze że po siedem złotych mówiła [she said that in this shop in kerfur it is seven] 105 ginzburg, yusupujiang, li, ren, kucharska, łupkowski here c’s utterance was tagged as other, dp, and cr. it seems that in this case, c’s utterance may be treated as a simple repetition of b’s question, and as such, it should not be recognized as dp. as for a’s utterance, it was tagged as ind, dpr, and da by the annotators. in this case, the answer does not require any form of inference. it simply states that it will be the same amount of money you can earn in certain places. the place and the amount of money are then pointed out by the following d’s utterance. that speaks for interpreting a’s utterance as a da (however, a partial one).16 8. formal analysis there is a two-way relationship between corpus studies of questions and responses and formal semantic theories of questions and of dialogue. notions from the latter play an important role in the design of the former. and one can strive to show that the categories posited are coherent formally using formal theories. conversely, the ability to fully describe the data that emerges from such corpus studies can be used as a means for evaluating different approaches. our aim in this section is to address both directions alluded to above. our explication is formulated using the frameworks of ttr (cooper and ginzburg, 2015; cooper, 2023) (for the semantic ontology) and kos (ginzburg, 2012; ginzburg et al., 2020) (for the theory of dialogue context); the relevant notions of ttr are sketched in appendix a, whereas those of kos are introduced in the text. 8.1 the classes da, dp, ind we assume that questions are propositional abstracts—extensive motivation for this view is provided in (ginzburg, 1995; ginzburg and sag, 2000; krifka, 2001); the particular implementation of this view in ttr can be found in (ginzburg, 2012; cooper and ginzburg, 2015).17 (33) exemplifies the denotations (contents) we can assign to a unary, binary wh-interrogative and to polar questions. we use rds here to represent the record that models the described situation in the context. the meaning of the interrogative would be a function defined on contexts which provide the described situation and which return as contents the functions given in (33). the unary question ranges over instantiations by persons of the proposition “x runs in situation rds”. the binary question ranges over pairs of persons x and things y that instantiate the proposition “x touches y in situation rds”: (33) a. who ran 7→ λr: [ x:ind rest:person(x) ] ( [ sit = rds sit-type = [ c:run(r.x) ] ]) b. who touched what 7→ 16as suggested by an anonymous reviewer for dialogue and discourse, it is plausible that in the discussed case a’s intention was to provide a complete answer but this was interrupted by d. 17a variant on this view motivated by data from boolean connectives and adjectives can be found in (ginzburg et al., 2014). 106 characterizing the response space of questions λr:  x:ind rest1:person(x) y:ind rest2:thing(y) ( [ sit = rds sit-type = [ c:touch(r.x,r.y) ] ]) c. did bo run 7→ λr:rec( [ sit = rds sit-type = [ c : run(bo) ] ]) d. didn’t bo run 7→ λr:rec( [ sit = rds sit-type = [ c : ¬run(bo) ] ]) polar questions are analyzed, following an initial proposal of ginzburg and sag (2000), as 0-ary abstracts, which in ttr is a question whose domain is the empty record type [] (that is, the type rec of records).18 this makes a 0-ary abstract a constant function from the universe of all records. it allows to distinguish the denotations of positive and negative polar questions, as exemplified in (33c,d) and as motivated by a variety of linguistic phenomena (hoepelmann, 1983; cooper and ginzburg, 2012). at the same time, it ensures that the answerhood relations they give rise to are (truth conditionally) equivalent, given that the simple answerhood relations they give rise to are equivalent and other answerhood relations are defined in terms of these.19 simple answerhood is the range of the propositional abstract, plus their negations. we exemplify what this amounts to for some cases in (34), using as we do mostly in the sequel familiar λ-notation for wh-questions and p?-notation for polar questions, rather than the official ttr notation above:20 (34) a. atomans(p?) = {p} b. atomans(¬p?) = {¬p} c. atomans(λx.p (x)) = {p (a), p (b), . . . , } d. negatomans(q) = {p|∃p1 ∈ atomans(q), p = ¬p1} e. simpleans(q) = atomans(q) ∪ negatomans(q) assuming questions to be propositional abstracts means that they can be used to underspecify answerhood. this is important given that nl requires a variety of answerhood notions, both for classifying responses and also for the role questions play as arguments to predicates such as ‘know’, ‘tell’, and ‘depends’, which in turn play a role in associated discourse reasoning (groenendijk and 18this is the type all records satisfy, since it places no constraints on them. 19the need for such truth conditional equivalence is motivated inter alia by inferences such as the following: (i) jill knows whether bo left. (ii) hence, jill knows whether bo did not leave. 20as cooper and ginzburg, 2015, §7.1 explain, the equivalence between the simple answerhoods of positive and negative polar answers follows because the negation operator on types ¬ satisfies for any s, t that— s : t iff s : ¬¬t , though t and ¬¬t are distinct types. 107 ginzburg, yusupujiang, li, ren, kucharska, łupkowski stokhof, 1997; wiśniewski, 2015). in fact, simple answerhood, though it has good coverage in practice, is not sufficient. it does not accommodate conditional, weakly modalized, and quantificational answers, all of which are pervasive in actual linguistic use (ginzburg and sag, 2000): (35) a. christopher: can i have some ice-cream then? dorothy: you can do if there is any. (bnc) b. anon: are you voting for tory? denise: i might. (bnc, slightly modified) c. how many players are getting these kind of opportunities to develop their potential? not many. (the guardian, nov 2, 2018) d. dorothy: what did grandma have to catch? christopher: a bus. (bnc, slightly modified) e. elinor: where are you going to hide it? tim: somewhere you can’t have it. thus, we suggest that the semantic notion relevant to direct answerhood is the relation aboutness—a relation between propositions and questions that any speaker of a given language can recognize, independently of domain knowledge and of the goals underlying an interaction. the most detailed discussion of aboutness we are aware of is (ginzburg and sag, 2000, pp. 129– 149), which offers (36a) (reformulated here in ttr)21 as a characterization of aboutness that can accommodate data such as (35).22 this requires the situational type component of the proposition to be a subtype of the join of the situational type of the question’s simple answer set. as it stands, this definition allows in principle very informationally strong types as direct answers, since nothing bounds the proposition from above. plausible upper bounds for direct answerhood familiar in the semantics of questions from the classic proposal of karttunen (1977) are the meets of the question’s atomic and negative atomic answer set.23 this condition is formulated in (36b):24,25 21see appendix a for some additional details. 22ginzburg (1995) suggested that aboutness is closed under conditionalization: i.e., for any r, p if p is about q, then so is if r, then p: (i) a: who will win tomorrow’s match? b: if it isn’t raining, the french. (ii) a: did someone switch the oven off? b: unless you explicitly told them to, no one did. the definition given in the text covers non–conditionalized answers. one crude strategy to obtain the latter, as proposed by ginzburg (1995), is to extend the definition for non-conditionalized answers by closing it under conditionalization. 23for a polar question p? the meets of the question’s atomic and negative atomic answer set are respectively p and ¬p, whereas for a wh–question λx.p (x) (e.g., ‘who left’) they are respectively ∧ p (ai) (‘bo left and millie left . . . ’), whereas ∧ ¬p (ai) (‘bo did not leave and millie did not leave . . . , i.e., equivalent to ‘no one left’). 24for a wide ranging discussion of a variety of answerhood relations, see (wiśniewski, 2015). he leaves the composition of his “base answer set”, the principal possible answers (ppas), as a parameter of the theory, to be fixed independently from the questions, since his account is stated in an artificial logical language that is not directly tied to linguistic forms. hence, his account is compatible in principle with most semantic approaches to questions. 25our use of subtyping as a means of characterizing aboutness reflects that, as an anonymous reviewer for dialogue and discourse points out, both direct and indirect answerhood involve inference. as we discuss below, for the latter the notion of inference is an agent-relative notion. 108 characterizing the response space of questions (36) for p = [ sit = s1 sit-type = t1 ] : prop, q = (r : t2) [ sit = s1 sit-type = t3 ] : question, a. about(p, q) holds iff t1 v ∨ {t |∃p′[p′ : prop ∧ simpleans(p′, q) ∧ t = p′.sit-type]} b. directans(p, q) holds iff about(p, q) and either (i) ∧ atomans(q) v t1 or (ii) ∧ negatomans(q) v t1 despite the proposals mentioned above for explicating direct answerhood, a comprehensive, empirically-based, experimentally tested account for a variety of wh–words is still elusive and an important task for future work. an additional important notion a theory of questions needs to provide for is a notion of exhaustiveness or resolvedness, though this is in general pragmatically parametrized (ginzburg, 1995; asher and lascarides, 1998; van rooy, 2003). whether a response is resolving (or merely goal fulfilling without so doing) can determine whether the response will be accepted as sufficient to end discussion of the question or requires a follow up. hence, the need for a finer–grained subdivision of the answer categories, as we hinted in footnote 7. given a notion of aboutness and some notion of (partial) exhaustiveness/resolvedness, one can then define question dependence (needed for the class dp), for instance, as in (37), though various alternative definitions have been proposed (groenendijk and stokhof, 1997; groenendijk and roelofsen, 2011; wiśniewski, 2013). for all these definitions, as with aboutness, their coverage awaits testing on empirical data: (37) depend(q1, q2) iff any proposition p such that p resolves q2, also satisfies p entails r for any r such that r is about q1, (ginzburg, 2012, (61b), p. 57). we have introduced answerhood notions corresponding to direct answerhood and to question– dependence, two of the three response categories we identified as question-specific in section 3. before we introduce the third notion, indirect answerhood, we sketch an account of dialogue context, which will allow us to integrate all three in a semantics for dialogue. the simplest model of context, going back to montague (1974), is one which specifies the existence of a speaker, addressing an addressee at a particular time. this can be captured in terms of the type in (38): (38)  spkr : ind addr : ind u-time : time cutt : addr(spkr,addr,u-time)  however, over the last four decades it has become clearer how much more pervasive reference to context in interaction is. expectations due to illocutionary acts—one act (querying, assertion, greeting) giving rise to anticipation of an appropriate response (answer, acceptance, counter–greeting), 109 ginzburg, yusupujiang, li, ren, kucharska, łupkowski also known as adjacency pairs (schegloff, 2007). extended interaction gives rise to shared assumptions or presuppositions (stalnaker, 1978), whereas epistemic differences that remain to be resolved across participants—questions under discussion are a key notion in explaining coherence and various anaphoric processes (ginzburg, 1994, 2012; roberts, 1996). these considerations among several additional significant ones we discuss below lead work in kos to two strategic moves: (i) instead of assuming a single context to be operative, a distributed notion is emergent from individual total cognitive states (tcs), one per participant. a tcs has two partitions, namely a private— about which we will not elaborate here—for details see (larsson, 2002), and a public one. (39) tcs = [ public : dgbtype private : private ] (ii) we posit a significantly richer structure to represent each participant’s view of publicized context, dubbed the dialogue gameboard (dgb), whose basic make up to process question–specific moves is given in (40): (40) dgbtype =  spkr : ind addr : ind utt-time : time c-utt : addressing(spkr,addr,utt-time) facts : set(prop) vis-sit = [ foa : ind ∨ rec ] : rectype moves : list(illocprop) qud : poset(question)  here facts represents the shared assumptions of the interlocutors—identified with a set of propositions. the parameters spkr and addr together with the addressing condition (at a given time) track verbal turns and mutual engagement. the remaining fields concern locutionary and illocutionary interaction. within moves the first element has a special status given its use to capture adjacency pair coherence and it is referred to as latestmove. the current question under discussion is tracked in the qud field, whose data type is a partially ordered set (poset). vis-sit represents the visual situation of an agent, including his or her focus of attention (foa), which can be an object (ind), or a situation or event (sit), relevant inter alia for processing gestural answers. we call a mapping between dgb types a conversational rule—conversational rules are the means for specifying how dgbs evolve. the types specifying its domain and its range we dub, respectively, the pre(conditions) and the effects, both of which are subtypes of dgbtype: they apply to a subclass of records that constitute possible dgbs and modify them to records that constitute possible dgbs. conversational rules are written here in a form where the preconditions represent information specific to the preconditions of this particular interaction type and the effects represent those aspects of the preconditions that have changed. the first conversational rule we formulate relates to the basic effect a query has on the dgb—as a consequence of a query a question becomes the maximal element of qud: 110 characterizing the response space of questions (41) ask qud-incrementation: given a question q and ask(a, b, q) being the latestmove, one can update qud with q as maxqud. pre : [ q : question latestmove = ask(spkr, addr, q) : locprop ] effects : [ qud = 〈 q, pre.qud 〉 : poset(question) ]  with this initial view of context and context change in hand, we can return to discuss indirect answerhood. the notion of direct answer is clearly complex and, as we have indicated, probably needs, at least for dialogue management purposes, to be refined. with indirect answers the situation seems even more tricky, which in part reflects why this category is one of those with most inter-annotator variability. indirectness encapsulates various notions, as we have already discussed in section 2. there is a considerable literature on indirect speech acts, building on and reacting to initial notions from grice (1975) and searle (1975). roughly speaking, these involve cases where the speaker’s intention is not transparently reflected in an utterance’s grammatically governed content—the content whose resolution is driven by conventional mechanisms.26 the classic gricean model involves initial recognition of a literal content (corresponding to what we have referred to above as ‘grammatically governed content’)27 and then, via domain–specific means, inference of the speaker’s intention. significant doubts about this time course, about the necessity of actually consulting a/the literal content, and what should be viewed as the literal/direct content have been debated extensively in the pragmatics literature, much of it in recent years on an experimental basis— for detailed review see (noveck, 2018). indirect speech acts are of course also an important theme in the ai planning literature, e.g., (cohen and perrault, 1979), incorporated in dialogue semantic frameworks in (larsson, 2002; asher and lascarides, 2003; ginzburg, 2012). while a detailed analysis is beyond our scope here, one can distinguish at least two cases, which we might label as shallow and deep indirect answers. the former corresponds to cases like (11) and (13) repeated here as (42a,b) respectively, where the entailment of a direct answer is due to shallow shared knowledge (for (42a): find(a,b,t1) → look_for(a,b,t0), so by contraposition ¬∃t look_for(a,b,t)→ ¬ find(a,b,t1)) or to domain–independent erotetic reasoning (wiśniewski, 2013), which adjusts the question asked to a close variant (larsson, 1998) (e.g., ?∃x.p (x) → λx.p (x), for (42b)). some initial refinement of ind along these lines is hinted in footnote 8 above. this contrasts with the deep indirect answers, exemplified in (42c), which involve reasoning about the speaker’s intentions, most often though not invariably based on domain–specific information. for detailed discussion of deep indirect answers within sdrt, see (asher and lascarides, 2001, 2003); for an account within kos, see (ginzburg, 2012, §8.3). (42) a. q: and also did you find my blue and green striped tie? r: i haven’t looked for it. 26by this we mean content whose contextual parameters are conventionally specified, e.g., ‘jill left’ conventionally specifies predication of some concept of leaving applying to a person the speaker refers to as ‘jill’; resolving which concept of leaving and which jill is less clearly rule–driven, though is a complicated mix of speaker/audience interaction, contextual salience etc. 27we use the latter somewhat pedantic term to differentiate it from the former, which has a variety of problematic associations. as will become clear in section 8.2, we do not assume that in general speaker and addressee need identically resolve even the grammatically governed content. 111 ginzburg, yusupujiang, li, ren, kucharska, łupkowski b. q: isn’t your country seat there somewhere? r: [yes/no]. stoke d’abernon. c. (context: in queue for toilet on an aircraft) anon woman: how desperate are you? me: (shrugs), go ahead. (ginzburg, 2012, p. 304) two basic conditions seem to characterize these cases: first, the indirect answer p is not a direct answer to the question q in the sense of the definition in (36b); second, p together with some shared knowledge, i.e., an element of facts for some dialogue gameboard dgb, the bridging proposition bridgeprop, entails r, which is a direct answer to q:28 (43) given p : prop, q : question, dgb : dgbtype indirectans(p,q,dgb) iff¬directans(p,q) and there exist bridgeprop, r : prop such that directans(r,q) and in(dgb.facts,bridgeprop) and→ (p ∧ bridgeprop, r). to exemplify: for (44a) asked by a who b knows needs to get up after sunrise, we could assume that the indirect answer p conjoined with (presumably shared) bridgeprop entails r:29 (44) a. a: is it time to rise? b: it is still dark outside. b. p = dark(here, now) c. bridgeprop = if it is dark here now, the time now is before a needs to rise. d. r = ¬needrise(a,now) we can now formulate a rule that explicate how answers and depended-upon questions get introduced in dialogue. this rule characterizes the contextual background of reactive queries and assertions—if q is maxqud, then subsequent to this either conversational participant may make a move which is either a (direct or indirect) answer or a question on which q depends). (45) a. given r : question ∨ prop, q : question, dgb : dgbtype, qspecific(r, q, dgb) iff directans(r,q) ∨ indirectans(r,q,dgb) ∨ depend(q,r) 28we leave open which notion of entailment, here denoted ‘→, is involved, whether directly relatable to the earlier subtyping or some other notion. 29an anonymous reviewer for dialogue and discourse asks whether p in such cases is the ‘literal’ content of the utterance or some strengthened version thereof such as some notion of speaker intended content, suggesting that in the latter case there might be significantly less need for indirect answerhood. given that grammatically governed content is the input to repair processes (ginzburg et al., 2003), it seems important for us to maintain p as a proposition that is explicitly not a direct answer if we wish to capture inter alia the clarificational potential from the addressee’s perspective, as well as the speaker’s choice in not explicitly uttering a direct answer. at the same time, we follow the reviewer’s suggestion in offering a characterization that is not strictly at the propositional level, since it makes intrinsic use of shared knowledge in entailing the direct answer, whereas for direct answers we use information state–independent type subsumption. this and other questions by the reviewer concerning indirect answers have led to several reformulations of our earlier proposed characterizations of indirect answerhood. 112 characterizing the response space of questions b. qspec =  pre : [ qud = 〈 q, q 〉 : poset(question) ] effects :  spkr = pre.spkr ∨pre.addr : ind addr : ind caddr : 6=(addr,spkr) p: prop ∨ question c1 : qspecific(p,q,pre)   8.2 the classes cr and ack metacommunicative utterances, including acknowledgements, clarification responses (crs) (also known as other repair and as other communication management) and (metacommunicative) corrections are challenging for most existing frameworks for dialogue semantics. for a start, given the mismatch they reveal between the dialogue interlocutors, they require a distributed approach to context. this rules out accounts where all semantic rules are assumed to apply to the common ground, made prominent in the view of qud due to roberts (1996).30 this was also the case for the view of discourse structure in earlier work in sdrt (e.g., asher and lascarides 1998, 2003). in more recent work (e.g., lascarides and asher 2009), sdrt adopts a view advocated in kos and also in the framework of ptt (poesio and rieser, 2010) that associates a distinct contextual entity with each conversational participant. a deeper challenge is that the analysis/generation of metacommunicative utterances requires access to the entire sign associated with a given interrogative utterance. this is for two main reasons. on the one hand, any constituent, certainly down to the word level can be the object of an acknowledgement and a clarification response, as exemplified for clarification responses in (46). moreover, as discussed in detail in (ginzburg, 2012), there are a variety of parallelism constraints relating to the form of such utterances that require reference to the non-semantic representation of the utterance. an illustration of this is given in (47) where the followup responses of two essentially synonymous questions turns out to be quite distinct: (46) a. [george] galloway [mp] is recorded reassuring his excellency [uday hussein] that ‘i’d like you to know we are with you ‘til the end.’ who did he mean by ‘we’? who did he mean by ‘you’? and what ‘end’ did he have in mind? he hasn’t said. (from a report in the cambridge varsity by jon swaine, 17 february 2006) b. is the war salvageable? that depends on what we mean by ‘the war’ and what we mean by ‘salvage’. (andrew sullivan’s blog the daily dish, sept, 2007) (47) a. a: do you fear him? b: fear? (=what do you mean by ‘fear’ or are you asking if i fear him) / #afraid? / what do you mean ‘afraid’? b. a: are you afraid of him? b: afraid? (=what do you mean by “afraid”? or are you asking if i am afraid of him) / #fear?/what do you mean ‘fear’? 30for a more refined stack–based discourse model, which distinguishes distinct participants’ commitments see (farkas and bruce, 2010). 113 ginzburg, yusupujiang, li, ren, kucharska, łupkowski this issue, first discussed in some detail in (ginzburg and cooper, 2004), rules out the lion’s share of logic–based frameworks where reasoning about coherence operates solely at the level of content. for instance, in sdrt the semantics/pragmatics interface has no access to linguistic form, but only to a partial description of the content that is derived from linguistic form. this has been argued to be necessary to ensure the decidability of sdrt’s glue logic (see e.g., asher and lascarides 2003, p. 77). in order to accommodate this class of utterances, it is crucial that the cognitive states keep track of the utterance associated with the question. in kos this is handled via the field pending whose type (locprop) is a record with two fields, one instantiated by an utterance token u, the other by an utterance type tu (the sign classifying u); this allows inter alia access to the individual constituents of an utterance. this leads to the following modified architecture for dgbs—they are distributed across dialogue participants (in other words—each participant is assigned their own dgb) and they include the field pending consisting of ungrounded utterances: (48) dgbtype 7→ spkr : ind addr : ind utt-time : time c-utt : addressing(spkr,addr,utt-time) facts : set(prop) pending : list(locprop) moves : list(illocprop) qud : poset(question)  ginzburg and cooper (2004); purver (2004); ginzburg (2012) show how to account for the main classes of crs using rule schemas of the form “if u is the interrogative utterance and u0 is a constituent of u, allow responses that are co-propositional31 with the clarification question cqi(u0) into qud.”, where ‘cqi(u0)’ is one of the three types of clarification question (repetition, confirmation, intended content) specified with respect to u0. for instance, responses such as (46b) can be explicated in terms of the schema in (49): (49) if a’s utterance u is yet to be grounded and u0 is a sub-utterance of u, qud can be updated with the question what did a mean by u0 more formally: the issue q0, what did a mean by u0, for a constituent u0 of the maximally pending utterance, a its speaker, can become the maximal element of qud, licensing follow up utterances that are copropositional with q0. assuming a propositional function view of questions, copropositionality allows in propositions from the range of range(q0) and questions whose range intersects range(q0). since copropositionality is reflexive, this means in particular that the inferred clarification question is a possible follow up utterance, as are confirmations and corrections, as exemplified in (51). 31here copropositionality for two questions means that, modulo their domain, the questions involve similar answers: for instance ‘whether bo left’, ‘who left’, and ‘which student left’ (assuming bo is a student.) are all co-propositional. 114 characterizing the response space of questions (50) parameter identification: pre :  maxpending = [ sit = u sit-type =tu ] : locprop a = u.dgb-params.spkr : ind u0 : sign c1 : member(u0,u.constits)  effects : maxqud = λxmean(a,u0,x) : question latestmove : locprop c1: copropositional(latestmove.cont,maxqud)   (51) a. λx.mean(a, u0, x) b. ?mean(a,u0,b) (‘did you mean bo’) c. mean(a,u0,c) (‘you meant chris’) 8.3 the classes motiv, dpr, cht, ignore łupkowski and ginzburg (2016) suggest that common to all classes of evasion utterances is a lack of acceptance of q1 as an issue to be discussed. in motiv-type responses the need/desirability to discuss q1 is explicitly posed, in cht-type responses there is an implicature that q1 is of lesser importance/urgency than r2 (expressing either a proposition or a question), whereas for ignore type responses there is an implicature that q1 as such will not be addressed. łupkowski and ginzburg (2016) also note that whereas q1 is not accepted for discussion, it remains implicitly in the context. in (52), where move (2) could involve either a motiv query (2a), or a cht query (2b), the original question has definitely not been re-posed and yet b still has the option to address it, which s/he should be unable to do if it is not added to his/her context before (52(2)). similar remarks mutatis mutandis apply to the dpr utterance in (52b): (52) a. a: who are you meeting next week? b(2): (2a) what’s in it for you? / (2b) who are you meeting next week? a: i’m curious. b: aha. a: whatever. b: oh, ok, jill. b. a: when are you leaving? b: i don’t know. a: come on! b: well, perhaps next week. this basic characteristic can be captured in the cognitive state architecture discussed above, given that qud is assumed to be partially ordered; this is a crucial difference from a view of qud as a stack or similar (roberts, 1996; farkas and bruce, 2010). concretely, łupkowski and ginzburg (2016) proposed to handle metadiscursive utterances such as motiv by viewing them as responses specific to the issue ?wishdiscuss(b,q) for a given question q and responder conversational participant b. this same approach can be applied to dpr, which łupkowski and ginzburg (2016) did not analyze, assuming that these involve responses specific to the issue λxknowanswer(x, q). we assume this formulation of the issue given the possiblity 115 ginzburg, yusupujiang, li, ren, kucharska, łupkowski of responses along these lines of ‘sam knows’, ‘you don’t know?’ etc.32 in fact, we will deviate somewhat from the account of łupkowski and ginzburg (2016) in proposing a more uniform account than they did of all four classes for reasons we explain below. in order to do this, we will define a single type evasiveresp that encompasses the commonalities between the four classes; each class will then be specified by merging evasiveresp with information specific to that particular class. in all cases, in line with the fact that q remains accessible, as exemplified in (52), qud is specified to include both q and a pertinent ‘metaquestion’. an additional commonality for all except dpr is turn change, underspecified for qspec given that for the latter it is not required, whereas in these cases it is more or less essential for coherence; this specification will be defused for dpr by using asymmetric merge. (53) evasiveresp=  pre : [ qud = 〈 q1, q 〉 : poset(question) ] effects :  spkr = pre.addr : ind addr = pre.spkr : ind r : question ∨ prop q2 : question r: illocrel moves = 〈 r(spkr,addr,r) 〉⊕ pre.moves : list(locprop) c1 : qspecific(r(spkr,addr,r),q2) qud = 〈 max = { q2,q1 } , q 〉 : poset(question)   given this, motiv and dpr are specified as follows:33 32utterances like ‘i don’t know’ and other dpr are differentiated from some other metadiscursive utterances in that the former can be used by the same speaker as a follow up, whereas the latter only if the speaker is correcting herself for having asked the question: (i) a: who should we invite? (ii) . . . i don’t know. (iii) . . . # do we need to talk about this now? (iv) . . . # i don’t wish to discuss this now. note also that ‘i don’t know’ can be used as an editing phrase (tian et al., 2015)—‘she’s i don’t know 29.’. 33the basic idea of merge for record types is illustrated by the examples in (i,ii). (i) [ f:t1 ] ∧. [ g:t2 ] = [ f:t1 g:t2 ] (ii) [ f:t1 ] ∧. [ f:t2 ] = [ f:t1∧. t2 ] in asymmetric merge, t1 ∧. t2, the second argument takes priority over the first, e.g., (iii) [ x:t1 y:t2 ] ∧. [ x=a:t1 ] = [ x=a:t1 y:t2 ] (iv) [ x=a:t1 y:t2 ] ∧. [ x=b:t1 ] = [ x=b:t1 y:t2 ] for a full definition which makes clear what the result is of merging any two arbitrary types, see (cooper, 2012, 2023). 116 characterizing the response space of questions (54) a. motiv = evasiveresp ∧. [ effects : [ q2 = ?wishdiscuss(spkr,pre.maxqud) : question ]] b. dpr = evasiveresp ∧. effects :  spkr = pre.spkr ∨ pre.addr : ind addr : ind caddr : 6=(addr,spkr) q2 = λxknow(x,pre.maxqud) : question   with respect to both cht and ignore, we adopt a somewhat different perspective than that offered by łupkowski and ginzburg (2016), for both empirical and conceptual reasons. considering the much larger dataset considered in this paper, their view of cht seems too “cooperative” and that of ignore too “hostile”. the analysis they offered for ignore built on an earlier analysis in (ginzburg, 2012) intended to capture gricean irrelevance, floutings of the gricean maxim of relevance as in (55). that analysis was designed to explain how the initial utterance in effect gets expunged from the dgb. (55) a: rozzo just gave a terrible talk. b: it’s really hot and unpleasant here. however, ignores often occur in quite cooperative environments such as the maptask, where under time pressure the response is driven by the observed situation. indeed, table 9 indicates that ignores were most frequently confused with answers (direct and indirect) and with chts; the former datum suggests, therefore, that ignores are susceptible to be viewed as addressing something related to the question asked. on the other hand, as far as cht goes, the analysis of łupkowski and ginzburg (2016) was, arguably, too “cooperative”. łupkowski and ginzburg (2016) assume that r2 is constrained to be unifiable with q1 via a question q3 (e.g., q1 = what do you (b) like? r2 = what do you (a) like? q3 = who likes what?). this assumption was motivated by a certain paralellism that seems to occur frequently between q1 and r2 when the latter has the form of a question. imposing this condition, which requires a question inference mechanism for testing this unifiability, significantly constrains the cht relation. however, in the more general case, where responses are not constrained to be questions, this condition seems less justified and, even focussing on question responses, the constructed example (56) seems quite natural:34 (56) a: when are you going to respond to the allegations? b: anyway, when are we going to get credit for our world leading vaccination program? the simplest analysis for ignore would make the pertinent meta-question be an arbitrary question about entities in the visual situation. similarly, for cht the simplest analysis would involve allowing a response specific to an arbitrary question. the obvious problem this would raise in both cases is massive ambiguity since many responses from other classes would be analyzable in such terms. to avoid this problem, we need to introduce an additional restriction, for instance along the lines of the afore-mentioned irrelevance; in other words, lack of coherence with the current context. what would this amount to? being neither qspecific with respect to q1 uttered by a to b, nor being co-propositional with a clarification question generated by q1’s utterance, nor qspecific with respect to ?wishdiscuss(b,q1) or λxknowanswer(x,q1). putting these conditions together amounts to the irrel relation of ginzburg (2012), which holds between an utterance and a dgb. 34the example is constructed, but familiar to anyone following the british political scene in early 2022. 117 ginzburg, yusupujiang, li, ren, kucharska, łupkowski given this, we formulate the rules for cht and ignore as in (57a) and (57b). the fact that in both cases the topic addressed is irrelevant(irrel) to the (precondition) dgb in the sense just discussed captures a similarity between the two. at the same time, there is also a significant difference in that ignore intrinsically uses material from the dgb, namely at least one entity from the visual situation as a constituent of the propositional nucleus of the question to establish coherence with the question posed. a further difference between the two–and deviation from (łupkowski and ginzburg, 2016)—is an emergent presupposition in the case of cht that the responder does not wish to discuss q1. (57) a. cht = evasiveresp ∧. effects :  q2 : question c3 : irrel(q2,pre) facts = pre.facts ∪{ ¬wishdiscuss(spkr,pre.maxqud) }   b. ignore = evasiveresp ∧.  effects :  a : ind c4 : in(vissit,a) g1 : type p : (ind)rectype q2 = (g1) sit =s sit-type = [ c : p(a) ]: question c3 : irrel(q2,pre)   9. conclusions and future work in this paper, we have presented an initial study for what is, as far as we are aware, the first, detailed, formally underpinned characterization of the response space of questions. concretely, our initial hypothesis, stated in the introduction as (1) is repeated here as (58): (58)(h) main hypothesis: responses drawn from or concerning the lg query classes plus direct answerhood exhaust the response space of a query. we think the data provided in previous sections validates this hypothesis, though we have made some small adjustments—conflating several classes. achieving such a characterization is a fundamental challenge for semantics with a very wide variety of applications. it establishes theoretical benchmarks for theories of dialogue, for dialogue systems, and for semantic theories of questions. apart from the need to scale up the evidence quantitatively, we are currently engaged in work on the following strands: • extending the characterisation of response spaces to other moves: we have partitioned the response space into question–specific and non–question-specific (metacomm, cht, ignore, motiv, dpr). this suggests that other moves such as assertions and commands can be characterized in similar terms, where the non–question-specific class is applicable to all. 118 characterizing the response space of questions • the account we have developed is domain general, abstracting over differences between different conversational types/genres/language games etc. to what extent the current account will change once one takes such differences into account is an important question. • cross-question type comparison: the q-r pairs annotated in the current study were selected randomly, whereas it is clearly of interest to consider the distribution of responses relative to fixed classes of questions (e.g., different classes of wh–questions, polar questions etc.) • apply machine learning to acquire the response classification scheme: yusupujiang et al. (2022) provide an initial study comparing both classical machine learning algorithms as well as pretrained language models such as bert (devlin et al., 2018). this achieves encouraging results on some classes (e.g., da and cr), while struggling with heavily inference-based classes like indirect answers, and ignore/cht. this learnability trend is closely in line with that achieved by the human annotators in the current paper. • spoken dialogue system implementation: we plan to test the usability of these categories in dialogue systems. for this, one needs dialogue systems with sophisticated nlu, along the lines sketched in (maraev et al., 2018, 2020). • cross-linguistic testing: a significant challenge is how to test the classification with languages lacking large or even hardly any speech corpora. we anticipate using online games with a purpose to this end (see e.g. łupkowski and ignaszak 2017; łupkowski et al. 2018; yusupujiang and ginzburg 2020). for an initial study concerning the response space of queries in uyghur, see (yusupujiang and ginzburg, 2022). finally, it is worth mentioning that at least part of our response typology can be be straightforwardly related to one of the well known annotation standards for dialogues, namely the iso 24617-2 (bunt, 2019).35 the standard focuses on functional segments of dialogue acts. these segments are understood as “minimal stretches of communicative behavior that have a communicative function, ‘minimal’ in the sense of not including material that does not contribute to the expression of the function or the semantic content of the dialogue act” (bunt, 2019, p. 4). when it comes to the general-purpose functions, dialogue acts may be information-providing (making certain information available to the addressee) or information-seeking (where information to be obtained can be of any kind, relating to the underlying task or activity, or even relating to the interaction itself). among the information-providing functions, two sub-categories are distinguished: answer functions (where the speaker is providing information in response to an information need) and informing functions (where the speaker wants the addressee to know or be aware of something. one may notice that parts of our typology relate to the scope of the information-providing functions. da and ind fall under answer functions, and ack, idk, dpr as well as cr may be categorized as informing functions. what would be interesting is to find a place for evasive responses in the dit++ scheme (probably among the dimension-specific functions). what remains an open question is how to incorporate question-responses into the aforementioned scheme. 35stemming from the dit++ annotation scheme (bunt, 2009). 119 ginzburg, yusupujiang, li, ren, kucharska, łupkowski 10. acknowledgements this work is supported by a public grant overseen by the french national research agency (anr) as part of the program “investissements d’avenir” (reference: anr-10-labx-0083). it contributes to the idex université de paris anr-18idex-0001. we also acknowledge a senior fellowship from the institut universitaire de france to the first author, which funded the internships of yusupujiang, li, and ren at llf. references anne h. anderson, miles bader, ellen gurman bard, elizabeth h. boyle, gwyneth m. doherty, simon c. garrod, stephen d. isard, jacqueline c. kowtko, jan m. mcallister, jim miller, catherine f. sotillo, henry s. thompson, and regina weinert. the hcrc map task corpus. language and speech, 34(4):351–366, 1991. aristotle. eudemian ethics. cambridge university press, 2012. edited by brad inwood and raphael woolf. nicholas asher and alex lascarides. questions in dialogue. linguistics and philosophy, 21(3): 237–309, 1998. nicholas asher and alex lascarides. indirect speech acts. synthese, 128(1):183–228, 2001. nicholas asher and alex lascarides. logics of conversation. cambridge university press, cambridge, 2003. john l. austin. truth. in james urmson and geoffrey j. warnock, editors, philosophical papers. oxford university press, 1961. paper originally published in 1950. jon barwise and john etchemendy. the liar. oxford university press, new york, 1987. jon barwise and john perry. situations and attitudes. bradford books. mit press, cambridge, 1983. ginger berninger and catherine garvey. relevant replies to questions: answers versus evasions. journal of psycholinguistic research, 10(4):403–420, 1981. harry bunt. the dit++ taxonomy for functional dialogue markup. in aamas 2009 workshop, towards a standard markup language for embodied dialogue acts, pages 13–24, 2009. harry bunt. guidelines for using iso standard 24617-2. tilburg university, jan 2019. ticc tr 2019–1. lou burnard, editor. reference guide for the british national corpus (xml edition). oxford university computing services on behalf of the bnc consortium, 2007. url http://www. natcorp.ox.ac.uk/xmledition/urg/. acess 20.03.2017. jean carletta. assessing agreement on classification task: the kappa statistic. computational linguistics, 22(2):249–254, 1996. philip cohen and ray perrault. elements of a plan-based theory of speech acts. cognitive science, 3:177–212, 1979. 120 characterizing the response space of questions robin cooper. type theory and semantics in flux. in ruth kempson, nicholas asher, and tim fernando, editors, handbook of the philosophy of science, volume 14: philosophy of linguistics, pages 271–323. elsevier, amsterdam, 2012. robin cooper. from perception to communication: a theory of types for action and meaning. oxford university press, 2023. url https://sites.google.com/site/ typetheorywithrecords/drafts/. robin cooper and jonathan ginzburg. negative inquisitiveness and alternatives-based negation. in maria aloni, floris roelofsen, galit weidman sassoon, katrin schulz, vadim kimmelman, and matthijs westera, editors, proceedings of the 18th amsterdam colloquium, 2012. robin cooper and jonathan ginzburg. type theory with records for natural language semantics. in chris fox and shalom lappin, editors, handbook of contemporary semantic theory, second edition, oxford, 2015. blackwell. cristian danescu-niculescu-mizil and lillian lee. chameleons in imagined conversations: a new approach to understanding coordination of linguistic style in dialogs. in proceedings of the 2nd workshop on cognitive modeling and computational linguistics, pages 76–87. association for computational linguistics, 2011. jacob devlin, ming-wei chang, kenton lee, and kristina toutanova. bert: pre-training of deep bidirectional transformers for language understanding. arxiv preprint arxiv:1810.04805, 2018. nicholas j enfield. questions and responses in lao. journal of pragmatics, 42(10):2649–2665, 2010. nicholas j enfield, tanya stivers, penelope brown, christina englert, katariina harjunpää, makoto hayashi, trine heinemann, gertie hoymann, tiina keisanen, mirka rauniomaa, et al. polar answers. journal of linguistics, 55(2):277–304, 2019. donka f farkas and kim b bruce. on reacting to assertions and polar questions. journal of semantics, 27(1):81–118, 2010. t. fernando. observing events and situations in time. linguistics and philosophy, 30(5):527–550, 2007. jonathan ginzburg. an update semantics for dialogue. in h. bunt, editor, proceedings of the 1st international workshop on computational semantics. itk, tilburg university, tilburg, 1994. jonathan ginzburg. resolving questions, i. linguistics and philosophy, 18:459–527, 1995. jonathan ginzburg. situation semantics and the ontology of natural language. in klaus von heusinger, claudia maierborn, and paul portner, editors, the handbook of semantics. walter de gruyter, 2011. jonathan ginzburg. the interactive stance: meaning for conversation. oxford university press, oxford, 2012. jonathan ginzburg and robin cooper. clarification, ellipsis, and the nature of contextual updates. linguistics and philosophy, 27(3):297–366, 2004. 121 ginzburg, yusupujiang, li, ren, kucharska, łupkowski jonathan ginzburg and ivan a. sag. interrogative investigations: the form, meaning and use of english interrogatives. number 123 in csli lecture notes. csli publications, stanford: california, 2000. jonathan ginzburg, ivan sag, and matthew purver. integrating conversational move types in the grammar of conversation. perspectives on dialogue in the new millennium, 114:25–42, 2003. jonathan ginzburg, robin cooper, and tim fernando. propositions, questions, and adjectives: a rich type theoretic approach. in proceedings of the eacl 2014 workshop on type theory and natural language semantics (ttnls), pages 89–96, gothenburg, sweden, april 2014. association for computational linguistics. url http://www.aclweb.org/anthology/ w14-1411. jonathan ginzburg, chiara mazzocconi, and ye tian. laughter as language. glossa, 5(1):104, 2020. doi: 10.5334/gjgl.1152. nancy green and sandra carberry. interpreting and generating indirect answers. computational linguistics, 25(3):389–435, 1999. herbert p grice. logic and conversation. in speech acts, pages 41–58. brill, 1975. jeroen groenendijk and floris roelofsen. compliance. in alain lecomte and samuel tronçon, editors, ludics, dialogue and interaction, pages 161–173. springer-verlag, berlin heidelberg, 2011. jeroen groenendijk and martin stokhof. questions. in johan van benthem and alice ter meulen, editors, handbook of logic and linguistics. north holland, amsterdam, 1997. jacob hoepelmann. on questions. in ferenc kiefer, editor, questions and answers. reidel, 1983. eduard hovy, ulf hermjakob, and deepak ravichandran. a question/answer typology with surface text patterns. in proceedings of the second international conference on human language technology research, pages 247–251. morgan kaufmann publishers inc., 2002. lauri karttunen. syntax and semantics of questions. linguistics and philosophy, 1:3–44, 1977. jacqueline c. kowtko and patti j. price. data collection and analysis in the air travel planning domain. in proceedings of the workshop on speech and natural language, hlt ’89, pages 119– 125, stroudsburg, pa, usa, 1989. association for computational linguistics. isbn 1-55860112-0. doi: 10.3115/1075434.1075455. url http://dx.doi.org/10.3115/1075434. 1075455. m. krifka. for a structured meaning account of questions and answers. audiatur vox sapientia. a festschrift for arnim von stechow, 52:287–319, 2001. klaus krippendorff. agreement and information in the reliability of coding. communication methods and measures, 5(2):93–112, 2011. theo af kuipers and andrzej wiśniewski. an erotetic approach to explanation by specification. erkenntnis, 40(3):377–402, 1994. 122 characterizing the response space of questions staffan larsson. comparing bdi approaches with the qud model. in j. hulstijn and a. nijholt, editors, proceedings of twendial 98, 13th twente workshop on language technology. twente university, twente, 1998. staffan larsson. issue based dialogue management. phd thesis, gothenburg university, 2002. alex lascarides and nicholas asher. agreement, disputes and commitments in dialogue. journal of semantics, 26(2):109–158, 2009. paweł łupkowski and olivia ignaszak. inferential erotetic logic in modelling of cooperative problem solving involving questions in the questgen game. organon f, 24(2):214–244, 2017. url http://www.klemens.sav.sk/fiusav/doc/organon/2017/2/214-244.pdf. paweł łupkowski and andrzej wiśniewski. turing interrogative games. minds and machines, 21 (3):435–448, aug 2011. doi: 10.1007/s11023-011-9245-z. url http://dx.doi.org/10. 1007/s11023-011-9245-z. paweł łupkowski, mariusz urbański, andrzej wiśniewski, wojciech błądek, agata juska, anna kostrzewa, dominika pankow, katarzyna paluszkiewicz, oliwia ignaszak, joanna urbańska, et al. erotetic reasoning corpus. a data set for research on natural question processing. journal of language modelling, 5(3):607–631, 2018. paweł łupkowski and jonathan ginzburg. a corpus-based taxonomy of question responses. in proceedings of the 10th international conference on computational semantics (iwcs 2013), pages 354–361, potsdam, germany, march 2013. association for computational linguistics. paweł łupkowski and jonathan ginzburg. query responses. journal of language modelling, 4(2): 245–293, 2016. brian macwhinney. the childes project: tools for analyzing talk. lawrence erlbaum associates, mahwah, nj, third edition, 2000. vladislav maraev, jonathan ginzburg, staffan larsson, ye tian, and jean-philippe bernardy. towards kos/ttr-based proof-theoretic dialogue management. in lauren prevot, magali ochs, and benoit fabre, editors, proceedings of semdial 2018, aix-en-provence, 2018. vladislav maraev, jean-philippe bernardy, and jonathan ginzburg. dialogue management with linear logic: the role of metavariables in questions and clarifications. traitement automatique des langues (tal), 61(3):43–67, 2020. per martin-löf. intuitionistic type theory. bibliopolis, naples, 1984. mary l mchugh. interrater reliability: the kappa statistic. biochemia medica, 22(3):276–282, 2012. richard montague. pragmatics. in richmond thomason, editor, formal philosophy. yale up, new haven, 1974. ira noveck. experimental pragmatics: the making of a cognitive science. cambridge university press, 2018. 123 ginzburg, yusupujiang, li, ren, kucharska, łupkowski fabian pedregosa, gaël varoquaux, alexandre gramfort, vincent michel, bertrand thirion, olivier grisel, mathieu blondel, peer prettenhofer, ron weiss, vincent dubourg, jake vanderplas, alexandre passos, david cournapeau, matthieu brucher, matthieu perrot, and edouard duchesnay. scikit-learn: machine learning in python. journal of machine learning research, 12: 2825–2830, 2011. massimo poesio and hannes rieser. (prolegomena to a theory of) completions, continuations, and coordination in dialogue. dialogue and discourse, 1:1–89, 2010. matthew purver. the theory and use of clarification in dialogue. phd thesis, king’s college, london, 2004. matthew purver. clarie: handling clarification requests in a dialogue system. research on language & computation, 4(2):259–288, 2006. piotr pęzik. spokes search engine for polish conversational data, 2014. url http://hdl. handle.net/11321/47. clarin-pl digital repository. aarne ranta. type theoretical grammar. oxford university press, oxford, 1994. craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. working papers in linguistics-ohio state university department of linguistics, pages 91–136, 1996. reprinted in semantics and pragmatics, 2012. carolyn p. rosé, barbara di eugenio, and johanna d. moore. a dialogue-based tutoring system for basic electricity and electronics. in susanne p. lajoie and martial vivet, editors, artificial intelligence in education, pages 759–761. ios, amsterdam, 1999. emanuel schegloff. sequence organization in interaction. cambridge university press, cambridge, 2007. john r searle. indirect speech acts. in speech acts, pages 59–82. brill, 1975. asad ali shah, sri devi ravana, suraya hamid, and maizatul akmar ismail. accuracy evaluation of methods and techniques in web-based question answering systems: a survey. knowledge and information systems, 58(3):611–650, 2019. gabriel skantze, jens edlund, and rolf carlson. talking with higgins: research challenges in a spoken dialogue system. in international tutorial and research workshop on perception and interactive technologies for speech-based systems, pages 193–196. springer, 2006. robert c. stalnaker. assertion. in p. cole, editor, syntax and semantics, volume 9, pages 315–332. ap, new york, 1978. anna-brita stenstrom. questions and answers in english conversation. lund studies in english, malmo: liber forlag, 1984. tanya stivers. an overview of the question–response system in american english conversation. journal of pragmatics, 42(10):2772–2781, 2010. 124 characterizing the response space of questions tanya stivers. how we manage social relationships through answers to questions: the case of interjections. discourse processes, 56(3):191–209, 2019. tanya stivers and nick j enfield. a coding scheme for question–response sequences in conversation. journal of pragmatics, 42(10):2620–2626, 2010. tanya stivers and jeffrey d robinson. a preference for progressivity in interaction. language in society, 35(3):367–392, 2006. tanya stivers, nicholas j enfield, and stephen c levinson. question-response sequences in conversation across ten languages: an introduction. journal of pragmatics, 42:2615–2619, 2010. ye tian, claire beyssade, yannick mathieu, and jonathan ginzburg. editing phrases. in chris howes and staffan larsson, editors, proceedings of godial, the 18th workshop on the semantics and pragmatics of dialogue, gothenburg, 2015. gothenburg university. a.m. turing. computing machinery and intelligence. mind, 59(236):433–460, 1950. robert van rooy. asking to solve decision problems. linguistics and philosophy, 26(6):727–763, 2003. wei wang. grammatical conformity in question-answer sequences: the case of meiyou in mandarin conversation. discourse studies, pages 610–631, 2020. andrzej wiśniewski. questions, inferences, and scenarios. college publications, london, england, 2013. andrzej wiśniewski. semantics of questions. in chris fox and shalom lappin, editors, handbook of contemporary semantic theory, second edition, oxford, 2015. blackwell. kyung-eun yoon. questions and responses in korean conversation. journal of pragmatics, 42(10): 2782–2798, 2010. zulipiye yusupujiang and jonathan ginzburg. designing a gwap for collecting naturally produced dialogues for low resourced languages. in workshop on games and natural language processing, pages 44–48, marseille, france, may 2020. european language resources association. isbn 979-10-95546-40-5. url https://www.aclweb.org/anthology/2020. gamnlp-1.7. zulipiye yusupujiang and jonathan ginzburg. ugchdial: a uyghur chat-based dialogue corpus for response space classification. in proceedings of lrec 2022, marseille, france, 2022. zulipiye yusupujiang, alafate abulmiti, and jonathan ginzburg. classifying the response space of questions: a machine learning approach. in proceedings of semdial 2022, dublin, ireland, 2022. 125 ginzburg, yusupujiang, li, ren, kucharska, łupkowski appendix a: basic notions of ttr type theory with records (cooper, 2012; cooper and ginzburg, 2015; cooper, 2023) is a cognitively construable formalism grounded in set theory, deriving much of its initial inspiration from situation semantics (barwise and perry, 1983; ginzburg, 2011) and its formal notions from constructive type theory (martin-löf, 1984; ranta, 1994). a fundamental notion of constructive type theory is the judgement a : t that classifies an object a as being of type t . failure to classify a by t is designated a 6: t . subtyping is defined as follows: (59) t1 v t2 iff s : t1 implies s : t2 objects can be classified by relatively simple types such as those in (60); (60) a. basic types (btype; 0-place; ind, loc, time, . . . ). b. predicate types (ptype; n-place; lion(x), carry(x,y), . . . ), constructed out of a predicate and objects which are arguments of the predicate. to classify more complex entities, for instance enable indefinite description, ttr introduces records and record types. a record is a set of fields assigning entities to labels of the form (61a), partially ordered by a notion of dependence between the fields on which their values depend. a concrete instance is exemplified in (61b). records are used to model events and states, including utterances and dialogue gameboards.36 (61) a.  l1 = val1 l2 = val2 . . . ln = valn  b.  x = 23 e-time = 2am, sept 17, 1915 e-loc = kamen–kashirskiy ctemp−at−in = o1  a record type is simply a record where each field represents a judgement rather than an assignment, as in (62). (62)  l1 : t1 l2 : t2 . . . ln : tn  36cooper and ginzburg (2015) suggest that for events with even a modicum of internal structure, one can enrich the type theory using the “string theory” developed by tim fernando (e.g., (fernando, 2007)). 126 characterizing the response space of questions the basic relationship between records and record types is that a record r is of type rt if each value in r assigned to a given label li satisfies the typing constraints imposed by rt on li. more precisely, (63) the record l1 = a1 l2 = a2 . . . ln = an  is of type  l1 : t1 l2 : t2 . . . ln : tn  iff a1 : t1, a2 : t2, . . . , an : tn. to exemplify this, (64a) (the temperature of a given location at a given time) is a possible type for (61b), assuming the conditions in (64b) hold. record types are used to model utterance types (saussurean/formal grammar signs) and to express rules of conversational interaction. (64) a.  x : ind e-time : time e-loc : loc ctemp−at−in : temp_at_in(e-time, e-location, x)  b. 23 : ind; 2am, sept 17, 1915 : time; kamen–kashirskiy : loc; o1 : temp_at_in(2am, sept 17, 1915, kamen–kashirskiy, 23) sometimes one needs to partially specify a general type by tying down one or more of the fields to a specific value. for this we use a manifest field as in (65): (65)  x : ind y=fido : ind e : hug(x,y)  this is the type of situation where some individual hugs the individual ‘fido’. any record of this type would be one meeting the conditions in (66): . (66)  x = a y = b e = s . . .  where a : ind b : ind and b is ‘fido’ s : hug(a,b) ttr assumes in addition the following type construction operations: (67) a. function types: (t1)t2 is the type of functions from elements of type t1 to type t2. 127 ginzburg, yusupujiang, li, ren, kucharska, łupkowski b. set and list types: set(t ) and list(t ). c. boolean types: (i) given a type t , there exists ¬t . (ii) given a set x of types ti, there exist ∨ x ti and ∧ x ti.∨ x ti and ∧ x ti have “classical” witnessing conditions: (68) a. r : ∨ x ti iff for at least one i ∈ x r : ti b. r : ∧ x ti iff for all i ∈ x r : ti in contrast, negation is a notion based on incompatibility that is a classical-intuitionist hybrid: (69) a. a : ¬t iff there is some t ′ such that a : t ′ and t ′ precludes t b. t ′ precludes t iff: • t = ¬t ′, or • t and t ′ are non-negative and there is no a such that a : t and a : t ′ one can show that t and ¬¬t are equivalent, but the former is a positive, the latter a negative type. on the other hand, a need not be of type t and there need not be a type t ′ that precludes t ; in other words: a : t ∨ ¬t is not a tautology. the basic reasoning for this goes back to (barwise and perry, 1983): (70) a. if i observe jo cutting onions, the situation i observe neither tells me that b. johnson is smoking a cigar, nor that he is not smoking a cigar. b. hence, svisual : cutting(j, o), svisual 6: cigarsmoke(b.johnson), hence: it is not the case that svisual : cigarsmoke(b.johnson), but neither is it the case that svisual : ¬cigarsmoke(b.johnson) the final notion we mention are propositions.37 propositions are construed as typing relations between records (situations) and record types (situation types), or austinian propositions (austin, 1961; barwise and etchemendy, 1987); more formally: (71) a. propositions are records of type prop = [ sit : rec sit-type : rectype ] . b. p = [ sit = s sit-type = t ] is true iff p.sit : p.sit-type i.e., s : t —the situation s is of the type t . two subtypes of austinian propositions are given in (72b,c): 37for a ttr approach using solely types and for detailed discussion of the two approaches, see chapter 6 of (cooper, 2023). 128 characterizing the response space of questions (72) a. sign =  phon : list(phonform) cat : [ head : pos ] dgb-params : rectype q-params : rectype cont : semobj  b. for classifying utterances, as described in the text: loc(utionary)prop(osition) = [ sit : rec sit-type : sign ] c. for assigning contents to dialogue moves: illoc(utionary)prop(osition) =  sit : rec x : ind y : ind a : prop ∨ question r : illocrel sit-type = [ c1 : r(x,y,a) ] : rectype  129 ginzburg, yusupujiang, li, ren, kucharska, łupkowski appendix b: annotation guidelines below we present the full annotation guidelines (in english and in polish) used in the described corpus study. the alert reader will notice that the number of question responses categories in the guidelines is larger than the number of categories discussed in the paper (see figure 1). the reason for this is that we decreased the number of categories initially posited by merging selected ones. the motivation for this move comes from the analysis of the annotation reliability and disagreement cases. we decided to merge (i) ia into ind, (ii) form and cor into cr, and (iii) idk into dpr. as a result, we have more general categories and we avoid a situation where we have categories with only few cases present in our data. we provide additional data for this paper (covering annotated q-r pairs and disagreement cases) which are hosted on the osf web-page (https://osf.io/mq6r7/). annotation guidelines instrukcja dla anotatorów is the utterance an answer (provides information required by a question) or a non-answer. czy reakcja na pytanie jest odpowiedzią (dostarcza informacji wymaganych przez pytanie) czy nie-odpowiedzią? if answer, then jeżeli odpowiedź, to da = direct answer (provides the required information straightforwardly). a: who is going to check that? / b: i can check that. a: and how long did they sleep? long? / b: well, you know, stas slept for at least two hours da = odpowiedź bezpośrednia (dostarcza wymaganych informacji wprost) a: kto to sprawdzi? / b: ja mogę to sprawdzić. a: a długo spali? długo spali? / b: wiesz co no staś to ze dwie godziny spał ia = indirect answer (you need to infer an answer from the utterance, it is not straightforward). a: do you want more tapes for them to take away? / b: i’ve got ten. i’ve haven’t used any of them. ia = odpowiedź pośrednia (wymagane jest wywnioskowanie odpowiedzi z wypowiedzi, nie jest ona podana wprost) a: chcesz więcej taśm, żeby zabrać je ze sobą? / b: mam dziesięć. nie użyłem żadnej z nich. else: non-answer, then is it a question? if question, then jeżeli nie-odpowiedź, to: czy jest to pytanie? jeżeli pytanie, to: cr = is q2 a query about something not completely understood in q1? a: why are you in? / b: what? cr = q2 jest zapytaniem o coś nie do końca zrozumianego w q1, prośba o wyjaśnienie a: dlaczego jesteś w środku? / b: co? a: na pewno a jest już? / b: proszę? dp = is it the case that the answer to q1 depends on the answer to q2? a: do you want me to push it round? / b: is it really disturbing you? dp = przypadek, w którym odpowiedź na q1 zależy od odpowiedzi na q2 a: czy chcesz żebym popchnął to dookoła? / b: czy naprawdę ci to przeszkadza? 130 characterizing the response space of questions ignore = does q2 relate to the situation described by q1? a: just one car is it there? / b: why is there no parking there? ignore = q2 odnosi się do sytuacji opisanej w q1, natomiast nie pośrednio do q1 a: tam jest tylko jeden samochód? / b: czemu tam nie ma parkingu? a: a i był ten merlot co w łodzi żeśmy pili to bardzo dobry był nie? / b: czternaście złotych chyba on kosztował? ind = is it the case that q2 is rhetorical and in this sense does not need to be answered and provides (indirectly) an answer to q1? a: what is it? / a: what’s he done? / b: ehm, you know what i’ve said before, eh, eh you’ll get ind = przypadek, w którym q2 jest retoryczne, nie musi uzyskać odpowiedzi i dostarcza (pośrednio) odpowiedzi na q1 a: co jest? / a: co on zrobił? / b: hmm, wiesz, to co powiedziałem wcześniej, hm, hm, dostaniesz form = is it the case that the way the answer to q1 will be given depends on the answer to q2? a: okay then, hannah, what, what happened in your group? / b: right, do you want me to go through every point? form = przypadek, w którym sposób w jaki odpowiedź na q1 będzie wyglądała zależy od odpowiedzi na q2 a: dobrze więc, hannah, co, co się stało w twojej grupie? / b: dobrze, czy chcesz żebym przeszła przez każdy punkt? motiv = does q2 address the motivation underlying asking q1? a: what’s the matter? / b: why? motiv = q2 pyta o motywację leżącą u podstaw q1 a: co się stało? / b: dlaczego pytasz? is it a declarative? if declarative czy jest to deklaratyw (zdanie twierdzące)? jeżeli deklaratyw, to: idk = i do not know, the speaker states that s(he) does not know the answer a: when’s the first consignment of scottish tapes? / b: erm don’t know. idk = mówca daje do zrozumienia, iż nie zna odpowiedzi a: kiedy to było? / b: erm nie wiem. a: to jest agnieszki ten koleś? / b: nie wiem ale ciszej dpr = difficult to provide an answer, the speaker states that it is hard to provide response, points at a different information source, etc. a: why? / b: i’m not exactly sure. dpr = trudność w podaniu odpowiedzi, mówca oświadcza, że podanie odpowiedzi jest trudne, wskazuje na inne źródło informacji, itd. a: czemu? / b: nie jestem pewien. cor = correction, the speaker point that a question has a wrong presupposition a: ub forty? / b: wd forty. cor = deklaratywny odpowiednik cr, zakłada coś związanego z intencjami oryginalnego mówcy zamiast pytać, odpowiedź wskazująca na błędne założenie obecne w pytaniu a: ub forty? / b: wd forty. 131 ginzburg, yusupujiang, li, ren, kucharska, łupkowski ack = acknowledgement, a speaker letting know that s(he) heard the question, e.g. mhm, aha etc. a: that’s about it innit? / b: mm mm. ack = potwierdzenie, mówca daje znać iż usłyszał/a pytanie poprzez na przykład mhm, aha itd. a: czy to już wszystko? / b: mhm. a: wiesz jak ma na imię? / b: poczekaj cht = an utterance that signals that the speaker does not want to answer, s(he) changes the topic, provides an evasive answer. a: what’s dolly’s name? / b: it’s raining. cht = wypowiedź sygnalizuje iż mówca nie chce odpowiedzieć, zmienia temat, udziela odpowiedzi wymijającej a: jak dolly się nazywa? / b: deszcz pada. a: czarne podoba ci się ? / b: brud widać ignore = the utterance does not relate to the question, but to the situation a: so does that mean that the ammeter is not part of the series, just hooked up after to the tabs? / b: let’s take a step back. a: what have you been doing melvin? / b: i ain’t talking cos you’ve got that bloody thing on ignore = reakcja odnosi się do sytuacji opisanej w q1, natomiast nie pośrednio do q1 a: melvin co ty robiłeś? / b: nic nie powiem bo masz to coś na sobie. a: ale w jakim pokoju? / b: no wiesz że on tam wiesz wyje trochę się uspokaja potem znowu wyje no in all other cases, put the other tag. w innych przypadkach, proszę użyć tagu other a w kolumnie obok opisać jaką funkcję spełnia ta reakcja na pytanie w tym konkretnym przypadku. 132 dialogue & discourse 16(1) (2025) 1–30 doi: https://doi.org/10.5210/dad.2025.101 german modal particles as discourse signals hannah j. seemann hannah.seemann@rub.de ruhr-universität bochum tatjana scheffler tatjana.scheffler@rub.de ruhr-universität bochum editor: amir zeldes submitted 01/2024; accepted 01/2025; published online 01/2025 abstract this study investigates the german modal particles ja and doch in discourse relations. we conduct an acceptability study of modal particles in four discourse relations (circumstance, condition, evidence, justify) to test predictions of (in)compatibilities derived from a corpus study by döring and repp (2019). as ratings for sentences representing the discourse relations circumstance and condition were significantly lower than for the two causal relations if presented with a modal particle, we confirm that modal particles and discourse relations interact. in a forcedchoice study testing the particle ja’s effect on relation disambiguation, we show that ja supports a causal interpretation of an ambiguous context in the absence of explicit discourse markers. our findings contribute to delineating the role of german modal particles in discourse, as we show that there is an interaction between discourse relations and modal particles, meaning that readers do not accept all modal particles in every discourse relation, and at least the modal particle ja can serve as a non-connective discourse signal for causal relations. keywords: modal particles; discourse particles; discourse relations; discourse signals; ja; doch; german 1. introduction speakers have a variety of linguistic means at their disposal to indicate their attitudes, knowledge, and internal states. in english, typical expressions are adverbs like obviously or probably, modal verbs like might, and certain verbs like to doubt or to know (biber, 2006; gray and biber, 2015). in the german language, modal particles can carry some of these meanings (zimmermann, 2011). german modal particles, also sometimes referred to as ‘discourse particles’, are noninflected sentence modifiers that semantically scope over the whole sentence, but do not affect its truth conditions. in discourse, they can be used as epistemic markers to express assumptions about interlocutors’ shared knowledge, to express the author’s attitude towards a proposition, or to indicate a speaker’s (un)certainty (hartmann, 1979; zimmermann, 2011). each particle makes a specific contribution to a discourse. for example, ja1 signals that the speaker assumes a piece of information to be either known, obvious in a given context, or easily 1. since there is no clear english equivalent for the german modal particles, we did not attempt to translate them in the text or in the example glosses. ©2025 hannah j. seemann, tatjana scheffler this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). seemann and scheffler verifiable, see (1). this function imposes restrictions on the use of ja: it typically cannot be used if the information is not previously known to the hearer, as shown in (2) – but see kratzer (1999) for a discussion of exceptions. (1) käsekuchen cheesecake ist is ja ja mein my lieblingskuchen. favorite.cake ‘as you and i assume/know, cheesecake is my favorite cake.’ (2) #du you wusstest knew bisher until.now nicht, not dass that käsekuchen cheesecake ja ja mein my lieblingskuchen favorite.cake ist. is intended: ‘you did not know before that cheesecake is my favorite cake.’ similar to ja, the particle doch signals information to be previously known, but additionally indicates contrast (karagjosova, 2004). this contrast might be between what a speaker expects someone else to know and the knowledge they show, as in b’s utterance in (3), or between the current and a previous utterance, as indicated by utterance b’. (3) a: i made your favorite dessert: chocolate cake! b: aber but mein my lieblingsdessert favorite.dessert ist is doch doch käsekuchen. cheesecake ‘but my favorite dessert is cheesecake (and you should know that).’ b’: aber but du you hattest had doch doch gesagt, said dass that du you käsekuchen cheesecake backst? bake ‘but you said you would bake a cheesecake, didn’t you?’ thus, modal particles interact with discourse by managing the integration of information and enabling speakers’ tendencies to indicate their own knowledge and beliefs. previous studies have even stated that particles may themselves signal discourse relations, but their exact role in discourse is not yet clear (helbig and buscha, 2001; döring, 2016).2 whether modal particles are discourse markers is discussed on both functional and formal terms in much detail in the collected volumes by degand et al. (2013) and to a lesser extent in fedriani and sansò (2017). however, the discussions on german modal particles and discourse markers are limited to theoretical considerations, concluding that ‘discourse marker’ refers to a function, whereas ‘modal particle’ is a (morphosyntactic) word class. none of the mentioned studies analyse the german modal particles’ contribution in discourse as studies on connectives and other discourse signals do. to address this gap, the present study investigates the interaction between modal particles and discourse relations from an experimental perspective. as previous work has shown that modal particles interact with discourse in the sense that they are only found in certain discourse relations that are specific for every particle (döring, 2016), the aim of this study is to delineate the nature of this interaction further: is the use of a particle restricted by the discourse relation? and does the particle itself play a role in which discourse relation is identified for a given context? 2. it should be noted that in the german research tradition, ‘discourse marker’ means something else than in the englishspeaking (or dutch-speaking) tradition. in the latter – and the field of discourse (structure) studies – conjunctions and other connectives indicating the discourse relations that hold between segments in a discourse are called ‘discourse markers’ (fraser, 1999; taboada, 2006). in german linguistics, the term ‘discourse marker’ refers mostly to conversation-structuring units, such as ich meine (‘i mean’) or english well. historically, these units have alternatively been referred to as ‘discourse particles’ or ‘pragmatic markers’, terms which also have been used to refer to modal particles (das, 2014; blühdorn, 2017). 2 german modal particles as discourse signals the following sections first summarize the theoretical background on german modal particles as well as discourse relations and their signalling in general, before we discuss previous work on modal particles in discourse. we then present an acceptability study (§6), testing the modal particles ja and doch in four different discourse relations. we show that readers are sensitive to the interaction of discourse relations and modal particles because they do not accept modal particles equally in the presented discourse relations. we follow this up with a forced-choice study (§7) which investigates whether the german modal particle ja functions as a discourse signal. the results show that ja is a causal discourse signal, as the presence of this particle in a sentence shifts the interpretation of an ambiguous discourse relation in favor of causal relations. 2. german modal particles from a syntactic and semantic point of view, there is a vast amount of research on german modal particles – an overview can be found in diewald (2009), zimmermann (2011), and müller (2017). nonetheless, to offer an overview of modal particles in general and the two particles studied here in particular, this section summarizes their (mostly) agreed-upon characteristics. modal particles are an open word class (thurmair, 2013). diewald (2009) describes 15 particles as belonging to the ‘core class’ of modal particles, with an open set of additional modal particles in the periphery. while it is not precisely defined which particles belong to the periphery of the class, most authors discussing modal particles list at least the majority of the particles that diewald includes in the core class: aber, auch, bloß, denn, doch, eben, eigentlich, etwa, halt, ja, mal, nur, schon, vielleicht, wohl, see e.g., thurmair (1989); hartmann (1998); müller (2014); abraham (2017). while different modal particles have different functions, they are all restricted to the middle field (mittelfeld) of a sentence, indicating that they cannot be topicalized or focused. in addition, there are restrictions to the sentence types and speech acts in which the particle can occur: some particles, like denn, are restricted to interrogative sentences, other modal particles are not restricted to one sentence type. though modal particles are sometimes taken to be obligatory in certain contexts, they are typically not viewed as indicators of sentence types (for discussion on the particles’ distribution cf. karagjosova, 2004; kwon, 2005; thurmair, 2013). since modal particles do not add any truth-conditional meaning, it is difficult to translate them into english without adding additional lexical content containing the particle’s expressive meaning, as seen in examples (1) and (3) (cf. gutzmann, 2016). in german, modal particles have homonyms in other word classes, e.g., ja can also be an answer particle (‘yes’) and doch can also be a conjunction (‘but’). moreover, the expressive meaning of doch as a particle changes depending on whether it is stressed or not, as seen in example (4). in (4a) doch is not stressed, while in (4b), it is.3 (4) a. das that ist is doch doch keine no kegelrobbe, grey.seal das that ist is eine a ringelrobbe. ringed.seal ‘what? that is not a grey seal, it is a ringed seal (and you should know that).’ b. das that ist is doch doch keine no kegelrobbe, grey.seal das that ist is eine a ringelrobbe. ringed.seal. ‘i was wrong! that is not a grey seal, it is a ringed seal.’ 3. it should be noted that the status of accented doch as modal particle is still up for debate. some authors treat it as a particle (gutzmann, 2010; hogeweg et al., 2011; egg and zimmermann, 2012; rojas-esponda, 2014a), while others consider it to be an adverb instead (blühdorn, 2019, abraham, 2020, p. 283). 3 seemann and scheffler the two particles studied in this article are ja and doch. as mentioned above, with the use of ja, the speaker assumes information to be known or highly salient to the hearer. it marks its prejacent proposition as not at-issue, indicating that it is not the main point itself, but relevant for understanding the main at-issue contribution (cf. potts (2005) and e.g., thurmair (1989); kwon (2005)). with doch, a speaker signals that information is assumed to be known to the hearer, but not necessarily part of the currently activated shared knowledge. without doch, the sentence in (4a) would simply state a fact about seal taxonomy. the use of doch expresses the speaker’s belief that the hearer ought to already know this. the aspect of contrast is kept if doch is stressed, but the reference point for the contrast changes: in (4b), the speaker corrects their own previous statement/belief, not the listener’s (egg and zimmermann, 2012; rojas-esponda, 2014b). based on their function of marking information as known, both particles should not co-occur with verbs presenting new information (thurmair, 1989; kratzer, 2004; gutzmann, 2015; viesel, 2017). restrictions of modal particle use in embedded clauses are described in coniglio (2012): modal particles can only appear in adverbial sentences that have illocutionary force independent of the matrix clause. therefore, it is argued that they cannot be used in embedded conditional or temporal clauses, see (5) and (6). (5) #falls if es it ja ja regnet, rains muss must ich i die the blumen flowers nicht not gießen. water intended: ‘i don’t need to water the plants if it rains.’ (6) #während while ich i ja ja draußen outside die the blumen flowers gieße, water schaut look die the katze cat von from drinnen inside zu. at intended: ‘while i water the plants outside, the cat is watching me from inside the house.’ if a modal particle can be used in a sentence introduced by a conjunction typically categorized as temporal, like nachdem (‘after’), coniglio (2012) views this as evidence for the conjunction having grammaticalized from a strictly temporal reading to a causal or contrastive conjunction. in example (7a), nachdem can be replaced by da (‘since’), illustrating its use as a causal conjunction. adding the modal particle ja to the subordinate clause is infelicitous if the same sentence is modified to allow only a temporal reading. (7) a. nachdem er ja immer gesagt hatte, ich könnte ihn jederzeit anrufen, habe ich das gestern auch gemacht. ‘because he always said i could call him anytime, i did just that yesterday.’ b. #nachdem er ja immer gesagt hatte, ich könnte ihn jederzeit anrufen, hat er nun allerdings sein telefon ständig ausgeschaltet. intended: ‘after always saying that i could call him at any time, he has now switched his phone off all the time.’ (adapted from hentschel, 1986, p. 203) before testing how these functions and restrictions translate to modal particle use in a discourse structure framework, the following sections offer an overview of the discourse model used in this study. 4 german modal particles as discourse signals 3. discourse relations: processing effects and marking of relations we employ rhetorical structure theory (rst; mann and thompson, 1988) as discourse framework. as a model of discourse coherence, rst enables the analyst to make explicit how they perceive the intentions of the text’s author. to this end, rst assumes discourse to be a sequence of discourse units that are connected by different types of relations. which and how many (coherence) relations there are differs by theory and set of relations introduced; e.g., relation sets defined in rst range from 23 relations (mann and thompson, 1988) to over 70 (marcu, 2000). even though the relation sets differ in how fine-grained the single relations are, they share the assumption that a discourse segment a can be the cause, elaboration, evidence, etc. for discourse segment b. rst also indicates which functions the segments a and b serve in every relation. this includes an assessment of which segment is necessarily required to obtain a coherent representation of discourse, independent of the relation types. the – in that regard – more important segment is called ‘nucleus’ (n), and the less important one the ‘satellite’ (s), as seen in example (8). a text with just its nuclei is understandable (yet missing information), a text’s satellites on their own are incoherent. (8) a. [the cat got wet]n [because it went outside during the rain.]s (cause) b. [in order to hunt its prey,]s [the cat climbed the tree.]n (purpose) c. yesterday, the cat did the following things: [it ate,]n [it climbed a tree,]n [it fell asleep on the bed.]n (list) in rst terms, example (8a) shows a (nonvolitional) cause relation. the information given in the satellite describes what caused the state of affairs given in the nucleus of this relation. the sentence in (8b) realizes a purpose relation. the activity described in the nucleus of this relation is needed to fulfill the situation described in the satellite. (8c) shows the multi-nuclear relation list. all segments are equally important – see also the rst website4 for a more detailed definition of rst’s discourse relations. because nucleus and satellite fulfill distinct functions in the different relations, linguistic material tied to any of those functions is prone to appear primarily in either the nucleus or the satellite. as will be described in more detail later, the modal particle ja, for example, is found almost exclusively in the satellite of causal relations. this matches its function of marking the information presented in the clause with modal particle as given: if the cause of an effect is known or highly salient, stating the effect becomes (seemingly) uncontroversial, too. beyond the function of establishing text coherence, discourse relations are also thought to be mental representations of information conveyed in the text (sanders et al., 1992), as well as processing instructions helping to construct the connection between segments of a text (canestrelli et al., 2013). earlier research focused on the differences in processing causal vs. additive relations, finding that causal relations between segments lead to better understanding and shorter reading times compared to additive relations (sanders and noordman, 2000; sanders, 2005). sanders (2005) proposes that this effect is due to language users assuming a causal relation between segments of text by default, so that establishing a non-causal relation between given segments leads to additional processing costs. recently, more types of relations and their processing have been studied, finding that negative as well as subjective relations like concession or contrast are harder to process compared to positive relations (kuperberg et al., 2011; canestrelli et al., 2013; köhne-fuetterer et al., 2021). 4. https://www.sfu.ca/rst/01intro/definitions.html 5 seemann and scheffler furthermore, it has been found that the explicit marking of a relation with a connective or cue phrase – as in (8a) and (8b) – helps both in reducing processing time and improving understanding of a text because the connective functions as an instruction for identifying the discourse relation (millis and just, 1994; degand and sanders, 2002; kamalski, 2007; sanders and spooren, 2007). explicit marking of relations is especially helpful for negative and subjective relations that express a speaker’s reasoning or judgment, as it helps to speed up the additional step of processing the propositional content as well as the negation or the speaker’s stance towards it (kleijn et al., 2019; wei et al., 2019; crible and pickering, 2020). 4. relation signals signals that help with the identification and processing of relations can appear in many forms. some of these are connectives like because or in order to; relations marked by a connective are referred to as ‘explicit’ relations, relations without a connective are ‘implicit’ (prasad et al., 2008). recent studies have focused on identifying other types of (non-connective) relation signals and their effects on text processing, as well as implicit discourse relations themselves. das et al. (2015) present the rst signalling corpus, a study on the different kinds of relation signals in rst annotated texts. they find that most relations are signaled by non-connective signals or by a combination of connectives and other signals, and only 10.65% of all relations are exclusively signaled by connectives (das and taboada, 2019). other signals the study identifies are lexical (non-connective cue words or phrases like i concede), morphological (tense), syntactic (word order, sentence mood) or semantic (synonymy, lexical chains). in addition, numerical (characters indicating lists) and graphical signals (punctuation, headings), genre and references to entities are identified as signals of relations (das and taboada, 2019). crible (2020) expands on this investigation by studying which types of signals co-occur. she reports that 32.55% of all relations are marked exclusively by connectives, while all other cases feature combinations with other signals. many of the signals identified by das and taboada (2019) have been shown to affect the processing of discourse positively by speeding up the processing and/or leading to a better understanding of the presented text. this has been shown for e.g., antonyms and resultative verbs (crible and demberg, 2020), parallelisms (crible and demberg, 2020; crible and pickering, 2020), quantifying expressions (scholman et al., 2020), enumerative structures (péry-woodley et al., 2017), and verbal tense (grisot and blochowiak, 2017). one general finding of these studies is that the signal disambiguates an ambiguous discourse context without the need for a connective to guide participants to an interpretation of the relation, as observed in detail by das and taboada (2019) and crible (2020). second, the effect of the signal is generally stronger if there is no connective present or if the connective is ambiguous, like and (asr and demberg, 2012; crible, 2020; das and taboada, 2019). moreover, epistemic markers have also been shown to give processing instructions similar to connectives. canestrelli et al. (2013) compared dutch causal connectives without further marking and found that sentences with the subjective connective want (‘because’) were processed more slowly compared to those with the objective connective omdat (‘because’).5 this effect was canceled out by the addition of the epistemic marker volgens (‘according to’) to the first part of the 5. ‘subjective’ and ‘objective’ refer to the source of coherence: a subjective connective might indicate that the speaker is presenting their own interpretation, while an objective connective might be used if observations without additional evaluation are reported (cf. e.g., sanders, 2005; canestrelli et al., 2013). 6 german modal particles as discourse signals sentence because the additional processing of the subjectivity indicated by want already happened at the epistemic marker. in a similar vein, wei et al. (2020) studied which (epistemic) stance markers and connectives co-occur in chinese. they found that the neutral connective suoyi (‘so’) appears more frequently with epistemic markers than the corresponding connective kejian (‘so’) that marks an utterance as subjective. both results indicate that as processing instruction, encoding subjectivity (= the speaker’s/author’s perspective, cf. sanders et al. (2021)) at one point of an utterance is both satisfying for the speaker/author as well as sufficient for the hearer/reader. in german, modal particles are used to mark the epistemic states of the interlocutors (doherty, 1985; zimmermann, 2011; dörre et al., 2018). this again motivates the question whether the german particles influence the perception of discourse relations, as has been shown for epistemic markers in other languages. whether there is a similar interaction between modal particles and discourse connectives in german might be an interesting question for future work. 5. previous work on modal particles in discourse modal particles have frequently been suspected to be markers in discourse, but few studies have explicitly investigated the contribution of german modal particles in a discourse structure framework. other work has studied the processing of modal particles without including the discourse context. one of the early mentions of modal particles as indicators of text coherence is presented in helbig and buscha (2001): (9) a. ich i gehe go nicht not schwimmen, swimming das the wasser water ist is ja ja noch still viel much zu too kalt. cold ‘i don’t go swimming because the water is still far too cold.’ b. ich i gehe go nicht not schwimmen, swimming weil because das the wasser water noch still viel much zu too kalt cold ist. is ‘i don’t go swimming because the water is still far too cold.’ (helbig and buscha, 2001, p. 429) the authors argue that in (9), ja can serve a function similar to the conjunction weil (‘because’), causally connecting the two arguments. the authors do not comment on the fact that the causality can be inferred even without any marker at all (cf. sanders, 2005), nor do they further elaborate on the discourse connecting function of modal particles. the first explicit connection between modal particles and discourse relations is drawn in a corpus study conducted by döring and repp (2019). the authors investigate whether there are modal particles that appear more often with certain discourse relations and whether speakers have preferences regarding modal particle choice in different relations. they find that the modal particles studied occur more or less often in certain discourse relations that are specific for each particle. for example, ja is frequently found in background or causal relations (e.g., evidence, justify), but never in temporal relations like circumstance or in condition. in addition, the position of the particle in each relation differs, e.g., whether the particle is found in the nucleus or the satellite. while ja is found almost exclusively in the satellite of the relations it appears in, doch is found in the nucleus of certain relations such as concession as well – although in general, it is still more frequently used in the satellite. in example (10), discussed by döring (2016, pp. 165-167), an evidence relation holds between [2] (nucleus) and [3] (satellite). döring argues that using ja in the satellite of evidence, which is where the evidence for the claim (= the nucleus) is presented, the 7 seemann and scheffler evidence is marked as being already in the common ground. by this, a speaker might raise the listener’s acceptance of the claim. since rejecting backgrounded or inferred information is difficult, using modal particles for certain effects might be an argumentative strategy. (10) [1] die the repräsentanten representatives der of.the gewerkschaften unions wie like auch also sie you im in.the hause house haben have in in wahrheit truth doch doch erkannt realized – [2] das that zeigt shows die the debatte debate heute today –, [1] daß that [...] die the große great mehrheit majority unserer of.our mitbürgerinnen fellow.citizens und mitbürger längst long.ago erkannt realized hat, has daß that um for der the sicherung security der of.the zukunft future willen pr veränderungen changes [...] notwendig necessary sind. are [3] wolfgang wolfgang schäuble schäuble hat has ja ja die the neuesten latest umfragedaten survey.data bekanntgegeben. announced ‘in fact, the representatives of the unions and you here in this house have realized – this shows today’s debate – that the majority of our fellow citizens realized long ago that changes are necessary to secure the future. wolfgang schäuble has announced the latest survey data.’ (döring, 2016, p. 165; kohlcorpus, speech 4, #10690) to confirm their corpus findings about the particles’ preferences, döring and repp (2019) conduct a forced-choice study in which participants have to choose one out of three given modal particles for background or justify contexts. they conclude that there is an interaction between modal particles and discourse structure, and that speakers can use the particles deliberately to convey a certain effect. in a larger corpus study investigating additional modal particles and how they interact with discourse structure, döring (2016) concludes that the particles do not just co-occur randomly with discourse relations, but only in specific combinations. yet, experimental work on the processing of german modal particles is scarce and does typically not account for the discourse context. dörre et al. (2018) compare modal particles with their homonyms to find out whether the modal particle (= the not-at-issue) reading of, e.g., einfach, leads to higher processing costs compared to its counterpart, the adjective einfach (‘simple’). they find higher reading times for the modal particle readings of the target words compared to their counterparts. this finding is in accordance with studies on the processing of discourse relations, even though these studies did not test modal particles: köhne-fuetterer et al. (2021) report higher processing times for relations more complex than causality, canestrelli et al. (2013) suggest that subjective relations have higher processing costs than objective relations, and van bergen and bosker (2018) find increased processing times for markers of inter-subjective meaning. further investigating the differences between modal particles and their homonyms, reimer and trotzke (2019) study whether reading times of nur and bloß (both ‘only’) differ depending on whether the particle is presented in its modal or focus particle reading. surprisingly, even though both particles are generally more frequently used as focus particles than as modal particles, reimer and trotzke find higher reading times for nur in the modal particle reading and for bloß in the focus particle reading. they conclude that there is a preference for one interpretation given a certain particle, independent of the frequency of that reading. 8 german modal particles as discourse signals and while it is frequently suspected that, given their discourse-structuring functions, modal particles might be signals or markers in discourse (e.g., in dahl, 1988; thurmair, 1989; könig and requard, 1991; helbig and buscha, 2001; döring, 2016), this assumption is never put to the test explicitly. the present study aims at this research gap, presenting experimental evidence for modal particles as discourse signals. 6. experiment 1: acceptability of modal particles in discourse relations the first experiment tested whether adding a modal particle to a sentence leads to lower mean acceptability ratings if the relation context presented is not expected to allow for the modal particle. this study is preregistered, access to the registration is provided in the data availability section. 6.1 hypotheses previous results suggest that there is an interaction between modal particles and (the processing of) discourse. the goal of the present study is to offer more experimental evidence of this interaction, testing whether readers are sensitive to it, as well as delineating the particles’ role as discourse signals. even though döring and repp (2019) do verify their corpus results in a forced-choice study, this is so far the only experimental evidence for the interaction of modal particles and discourse relations. therefore, using a different experimental design, we also tested some of the reported corpus results experimentally to further confirm that modal particles are restricted to certain discourse relations. in our acceptability study, we investigate the following hypotheses: h1: acceptability of sentences with modal particles is lower for certain discourse relations. h2: the acceptability of a modal particle in a discourse relation also depends on the particle’s position (= nucleus or satellite). to ensure comparability with the findings of the corpus data presented in döring and repp (2019), the modal particles chosen for all experiments are ja and doch. in their study, the authors report frequent occurrences of ja and doch in the discourse relations evidence and justify,6 and no occurrences in the relations circumstance and condition. therefore, we refer to circumstance and condition as ‘dispreferred’ and to evidence and justify as ‘preferred’ relations for these two modal particles, and we expect sentences with particles in dispreferred relations to have lower mean acceptability ratings than the same sentence without a particle. even though döring and repp (2019) report on the frequency in more than four relations, these four relations will be used in the rating study because they show similar frequency patterns for both modal particles. we thus expect ratings for circumstance and condition to be lower than the ratings for evidence and justify if the presented sentences contain a modal particle. in addition, we expect to see higher ratings if the particles are presented in a relation’s satellite than for the nucleus, given that this is the position they are found in most often in the corpus. based on these predictions, we can specify our hypotheses:7 6. the definition of justify differs from the definition of the same relation in stede (2016a), and is more similar to what is called reason there. to construct the target sentences for this experiment, the relation definitions given in döring and repp (2019) have been used. 7. the specified hypotheses were not part of the original preregistration. we have nonetheless decided to include them here in order to further specify the vague original hypotheses. 9 seemann and scheffler h1’: the acceptability of sentences with modal particles is lower in the relations circumstance and condition (= ‘dispreferred’) than in evidence and justify (= ‘preferred’). h2’: the acceptability of sentences with modal particles is lower if the particle is presented in the relation’s nucleus. 6.2 participants 92 participants (mean age 35.5 years, sd 12.5 years), all native speakers of german were recruited via prolific.8 participants received compensation based on germany’s minimum wage (2 eur for 10 minutes). one submission was excluded for failing two out of three attention checks, as predetermined in the study preregistration. since modal particles are rarely found in written text and readers might tend to rate any written sentence containing a modal particle as ‘bad’, we instructed people as follows: this experiment contains 40 sentences. the sentences have been produced by a software trying to change the style of formal texts to informal style. the software does so by inserting various words in the original, formal texts. unfortunately, the software tends to be overzealous and also inserts words that don’t fit the context. your task is to rate the transformed sentences. a sentence is good if you find the style to be colloquial, but not odd. a sentence is bad if you stumble over certain words while reading.9 the instruction itself was kept in a colloquial style to prepare the participants for reading rather colloquial than formal written sentences. 6.3 materials and design we used a 3 x 2 x 2 design that accounts for the modal particle (ja, doch, no particle), the discourse relation (preferred: evidence and justify, dispreferred: circumstance and condition) as well as the position of the particle (nucleus or satellite). definitions for the four discourse relations are given in appendix a. for each discourse relation, we designed five experimental items that were presented in five different versions: with ja or doch in the nucleus, with ja or doch in the satellite, or without particle, see (11)–(14). all relations are explicitly marked by a connective: circumstance is marked by während (‘while’), condition is marked by falls (‘if’), evidence is marked by da (‘because’), and justify is marked by weil (‘because’). the 20 experimental items were distributed over five pre-determined lists to ensure every participant reads sentences from all conditions, but each sentence only once. (11) circumstance während ingo (∅/ja/doch) in der straßenbahn sitzt, liest er (∅/ja/doch) die zeitung. ‘while ingo (∅/ja/doch) is on the tram, he (∅/ja/doch) reads the newspaper.’ 8. https://www.prolific.com/ 9. original german instruction: ‘dieses experiment besteht aus 40 sätzen. produziert wurden die sätze von einer software, die den stil von existierenden texten auflockern soll. dies tut die software unter anderem, indem sie verschiedene worte in ursprünglich eher nüchtern geschriebene texte einfügt. leider ist die software stellenweise etwas übereifrig und fügt manchmal auch worte ein, die eigentlich an der jeweiligen stelle gar nicht passen. ihre aufgabe ist die bewertung der fertigen sätze. ein satz ist gut, wenn sie den stil als (tendenziell) umgangssprachlich, aber trotzdem nicht merkwürdig empfinden. ein satz ist schlecht, wenn sie beim lesen über bestimmte worte stolpern.’ 10 german modal particles as discourse signals (12) condition falls fiona (∅/ja/doch) nicht bald aufbricht, wird sie (∅/ja/doch) zu spät kommen. ‘if fiona (∅/ja/doch) doesn’t leave soon, she will (∅/ja/doch) be late.’ (13) evidence da helena (∅/ja/doch) häufig auf konzerten ist, kann sie (∅/ja/doch) viele lieder mitsingen. ‘because helena (∅/ja/doch) frequently visits concerts, she can (∅/ja/doch) sing along to many songs.’ (14) justify weil jakob (∅/ja/doch) medizin studieren will, muss er (∅/ja/doch) einen guten abschluss machen. ‘because jakob (∅/ja/doch) wants to study medicine, he has (∅/ja/doch) to get a good degree.’ all sentences were constructed according to the same pattern: the relation’s satellite including a typical connective first, and the nucleus second. if applicable, the modal particles were always placed after the name/pronoun, in the middle field of the sentence. it has to be noted that for some of the sentences containing doch, ambiguity with the particle’s adverbial reading cannot be avoided. this limitation will be further discussed in §6.6. in addition to the 20 experiment items, participants read three training items, 17 fillers and three attention checks asking participants to choose a specific point of the rating scale. some of the filler items contained (different) modal particles to draw attention away from ja and doch and to balance the number of sentences with/without particles as well as the fraction of sentences that are expected to have good or bad ratings. after having read each sentence, participants were asked to rate the sentence on a 7-point likert scale (1: ‘bad’, 7: ‘good’), prompted by the question “how well did the software transform the sentence?”. 6.4 analysis we fit two ordinal regression models10 with varying intercepts and slopes using the r package ordinal (christensen, 2022). the two models are necessary to deal with model deficiency (arising from the constraint that ‘position’ can not be accounted for if no particle is present in a sentence) in a combined model. given that döring and repp (2019) report the same occurrence pattern of ja and doch for the relations used in this study, we expect the same rating patterns for both particles. however, to verify that there are no statistically significant differences between the two, we fit our first model using helmert coding to check for differences between ja and doch compared to the no particle condition in the ratings of the presented sentences. possible differences between the two particles might arise due to their slightly different meaning contributions: both signal known information, but only doch also indicates contrast. since we did not find a statistically significant main effect for the difference between the two particles, we fit the second model to compare the ratings of both particles combined in the positions nucleus and satellite. in both models, we coded the relations 10. this deviates from the preregistered analysis plan because it was pointed out to us that this model fits the ordinal data from the likert scale rating better. the analysis had to be split into two models to account for rank deficiency in the combined model. 11 seemann and scheffler according to our assumptions defined in the beginning of this section as ‘preferred’ (evidence, justify) and ‘dispreferred’ (circumstance, condition) and report on the maximum random effect structure. 6.5 results figure 1 provides an overview of the descriptive results of our study, the ratings for the four relations by modal particle and position. in all four relations, sentences without particle were rated highest. furthermore, if there is a difference between particles or positions, there is a tendency for ja to be rated higher than doch, and for a particle in the satellite to be rated higher. one exception is the relation condition: doch in the satellite is rated as more acceptable than ja. without modal particles, the ratings for circumstance and justify are high (median = 7), followed by condition (median = 6) and evidence (median = 5). figure 1: acceptability ratings for the two ‘dispreferred’ discourse relations circumstance, condition and the two ‘preferred’ relations evidence and justify. the rating scale ranges from 1 (= rejection) to 7 (= acceptance). ratings are divided by the particle’s position (nucleus vs. satellite) and by modal particle (doch vs. ja). the observations in figure 1 are confirmed by our first model (shown in table 1). the model shows that the comparison of ‘no particle’ vs. ‘particle’ is significant (p < 0.001), but the difference between the two particles ja and doch is not (p = 0.15). the main effect of preferred vs. dispreferred relation is not significant in general, but the significant interaction between relation and presence of particle indicates that the particle’s presence influences the ratings (p < 0.05). this is illustrated 12 german modal particles as discourse signals in figure 2: the mean rating of sentences in dispreferred relations drops more in the presence of a particle than in relations that are assumed to be acceptable with the given modal particle. fixed effects estimate std. error z-value p-value preferred vs. dispreferred relation 0.092 0.392 0.234 0.815 ∅ vs. (ja/doch) -3.030 0.291 -10.398 < 0.001 ja vs. doch -0.235 0.166 -1.414 0.157 relation : ∅ vs. (ja/doch) 1.026 0.507 2.023 < 0.05 relation : ja vs. doch -0.913 0.325 -2.810 < 0.01 table 1: output of the cumulative link mixed model for ordinal data shown below. we use helmert coding to compare first ‘no particle’ to (ja/doch) and ja and doch in a second step. we use sum coding for the relations, combined to ‘preferred’ (evidence, justify) and ‘dispreferred’ (circumstance, condition). clmm(rating ∼ relation * particle + (1+particle*relation|participant) + (1+particle|item)) while the difference between the two particles in general is not significant, the interaction of relation and ja vs. doch is (p < 0.01). since in three out of four relations, the two particles show similar rating trends – as visible in figure 1 – this interaction is likely to be driven by the rating of doch in condition. in the preferred relations, ja is rated as being more acceptable than doch in the relation’s satellite. no such difference is visible in the dispreferred relation circumstance. only in condition, doch is rated as more acceptable than ja. this difference might arise from an ambiguity of the particle doch (which has an additional stressed reading, elaborated in section §6.6), rather than a difference between occurrences of the unstressed modal particles ja and doch. the second model (table 2) shows the comparison of a sentence with any of the two particles in the nucleus or satellite. ratings for sentences without particles are not included. fixed effects estimate std. error z-value p-value preferred vs. dispreferred relation 0.4474 0.2994 1.494 0.13513 nucleus vs. satellite -0.2588 0.2067 -1.252 0.21063 relation : position -1.1468 0.3981 -2.880 < 0.01 table 2: output of the cumulative link mixed model for ordinal data shown below. we use sum coding to compare both position (nucleus, satellite) and relation, combined to ‘preferred’ (evidence, justify) and ‘dispreferred’ (circumstance, condition). clmm(rating ∼ relation * position + (1+relation*position|participant) + (1+position|item)) the model shows no significant main effect for position (p = 0.21), but a significant interaction between relation type and position (p < 0.01). the differences in the ratings of sentences with modal particles in dispreferred/preferred relations by position are presented in figure 3: whether the modal particle appears in nucleus or satellite does not strongly influence the rating in dispreferred 13 seemann and scheffler 1 3 5 7 no_particle particle a cc ep ta bi lit y ra tin g dispreferred preferred figure 2: mean rating for ‘preferred’ (evidence, justify) and ‘dispreferred’ (circumstance, condition) relations, by particle presence. the rating scale ranges from 1 (= rejection) to 7 (=acceptance). 1 3 5 7 nucleus satellite a cc ep ta bi lit y ra tin g dispreferred preferred figure 3: mean rating for ‘preferred’ (evidence, justify) and ‘dispreferred’ (circumstance, condition) relations, by position of the particle. ratings for sentences without particle are excluded. the rating scale ranges from 1 (= rejection) to 7 (=acceptance). relations, but does in preferred relations. here, the presence of the modal particle in the satellite, as expected, leads to higher acceptability ratings. 6.6 discussion the rating study showed that the acceptability of sentences with modal particles varies depending on the discourse relation, thus confirming our h1: the acceptability of sentences with modal particles is worse for certain discourse relations. still, in contrast to our expectation, modal particles in an evidence and justify relation were not generally more accepted than particles in circumstance or condition, but showed this effect only if the particle was presented in the relation’s satellite. we can confirm our h2 as well: for the relations evidence and justify, participants preferred the modal particle in the satellite of the relation over the particle being in the relation’s nucleus. the effect could not be shown for the relations circumstance and condition. this is in accord with expectations based on the particles’ semantics and the relation definitions: in evidence and justify, the claim is presented in the nucleus, whereas the evidence/argument is presented in the satellite of the relation, matching the particles’ effect of marking information as known or salient. furthermore, we take our results to indicate an interaction of particles with the discourse relation and not simply co-occurrence with specific connectives. (15a) shows that in an implicit condition relation, adding a modal particle to the satellite is still infelicitous.11 on the other hand, adding ja or doch to the satellite of justify as in (15b) is acceptable. while it will be up to future studies to verify these intuitions, we expect similar rating patterns for implicit relations. 11. a syntactic structure such as (15a), though without a modal particle, might be used to express a conditional sentence in german. 14 german modal particles as discourse signals (15) a. #bricht fiona ja/doch nicht bald auf, wird sie zu spät kommen. intended: ‘if fiona doesn’t leave soon, she will be late (as you knew before).’ b. jakob will ja/doch medizin studieren, er muss einen guten abschluss machen. ‘(as you know,) jakob wants to study medicine, he has to get a good degree.’ the box plots showed that participants rated sentences in the evidence relation generally lower compared to all other relations. this is probably due to the order of the segments. to present all relations in the same manner, the satellite was always presented first. however, presenting the satellite (= the evidence) before the nucleus (= the claim) might be unusual for an evidence relation and lead to lower acceptability ratings (cf. stede, 2016a). even though employing a modal particle that marks the evidence as uncontroversial might be an effective rhetorical strategy (as indicated in section §3), the approach of presenting evidence first does seem to be less acceptable to participants in general. the higher acceptability scores of doch in the condition relation might be due to a limitation of this study: context could not be provided for the sentences, as there is no context that allows for both modal particles ja and doch to be used in all stimuli. but without context, participants’ ratings of a sentence can differ depending on whether they read the particle in its stressed or unstressed version. this is exemplified in (16). while (16b) is a perfectly fine example, the sentence in (16a) is odd in german. (16) a. #falls fiona nicht bald aufbricht, wird sie doch zu spät kommen. intended: ‘if fiona doesn’t leave soon, she will be late (as you knew before).’ b. falls fiona nicht bald aufbricht, wird sie doch zu spät kommen. ‘if fiona doesn’t leave soon, she will be late (even though it seemed like she could make it in time).’ for future studies, this limitation of the prosodic stress influencing the particle’s rating outside of the experiment conditions can be avoided by reading the sentences aloud to participants. we acknowledge that further research is necessary to pinpoint the exact behavior of doch in different discourse contexts. in this first experiment, we obtained results similar to previous findings on modal particles in discourse. taking our results together with the experimental evidence presented in döring and repp (2019), an interaction of modal particles with discourse relations should be assumed. our results go beyond their corpus findings insofar as they show that the dispreferred relations – that were never found with the tested modal particles in the corpus – are not completely unacceptable to the tested participants, just less acceptable than the preferred relations. after having identified that participants are sensitive to the interaction of particles and discourse, we proceed to studying this interaction further to find out whether modal particles are non-connective discourse signals. 7. experiment 2: modal particles as discourse signals the second experiment tested whether the presence of the modal particle ja in a sentence leads to participants disambiguating the sentence as causal, indicated by their choice of a causal connective. the preregistration can be accessed in the data availability section. 15 seemann and scheffler 7.1 hypotheses döring (2016) suggests that modal particles do not simply co-occur with discourse relations, but can be used to mark those relations. taking the findings in das and taboada (2019) and studies on relation signals like crible and pickering (2020) as well as the particles’ discourse-managing function into account, modal particles might be a non-connective discourse signal. to study this discourse-structuring characteristic, we tested the following hypothesis: h3: as discourse signals, modal particles can help in disambiguating ambiguous relation contexts. to test this hypothesis, we conducted a connective insertion study. participants were presented with ambiguous contexts (causal-contrastive or causal-temporal) with a contrastive or temporal reading and an optional causal reading. the texts were presented with or without the modal particle ja and participants were asked to insert a connective out of a given set of alternatives. the relation types were chosen because they differ conceptually, but minimal pairs can be disambiguated without changing the sentence structure. based on the results presented in crible and pickering (2020) – disambiguating a relation is easier if the second sentence contains a signal – we modified the stimuli to be presented with or without modal particle in the second clause. we expect mainly contrastive respective temporal connectives to be chosen if the sentence is presented without ja, and a significant increase of the causal connective chosen if the sentence is presented with the particle. this expectation builds on the results presented in döring (2016) as well as ja’s function of marking information as known or highly salient and not at-issue. this marking of an utterance makes ja a good choice for causal relations. any statement containing ja might be more readily perceived as being uncontroversial or given – as long as the claim is not obviously false – and being a valid point in an argument. thus the modal particle supports causality without being causal by itself: in the presence of ja, the causal connective deshalb (‘that is why’) should be chosen more frequently than without the particle. 7.2 participants for this experiment, 100 participants (mean age 35.3 years, sd 10.1 years) were recruited via prolific and received compensation based on the minimum wage in germany (1.60 eur for 8 minutes). in this study, participants were instructed to choose one out of four connectives – the one they liked best in the given context – to connect the two sentence parts presented. seven participants had to be excluded, two because german was not their first language, and five because they chose the syntactically mismatched distractor connective in critical items. 7.3 materials and procedure we used a 2 x 2 design with the factors modal particle (ja, no particle) and ambiguous discourse context (causal-contrastive, causal-temporal). we designed eight experiment items for both discourse contexts that were presented in two versions: with or without the particle. the 16 items were split into two lists, ensuring every participant reads every sentence in exactly one version, and presented in random order. an example of a causal-contrastive context is given in (17), and a causal-temporal context is presented in (18). the low number of items is intended to avoid overexposure of participants to the modal particles. to balance the low number of observations per participant, we recruit a high number of individual participants. 16 german modal particles as discourse signals (17) causal-contrastive raphael serviert normalerweise zimtschnecken als dessert, backt florian (∅/ja) meistens schokoladenkuchen. ‘raphael usually serves cinnamon rolls as dessert, florian mostly bakes (∅/ ja) a chocolate cake.’ (18) causal-temporal die züge warten auf die behebung einer signalstörung, warten die fahrgäste (∅/ja) am bahnhof. ‘the trains are waiting for a signalling error to be fixed, the passengers are (∅/ ja) waiting at the train station.’ in addition, the experiment included two training items and eight items in unambiguously concessive contexts as fillers. none of the fillers contained any modal particles. when presented with each item, participants could choose between four different connectives: contrastive dahingegen (‘whereas’), causal deshalb (‘that is why’), concessive obwohl (‘although’), and temporal währenddessen (‘meanwhile’), and were asked to pick the one they prefer in this context. except for obwohl, which serves as a distractor, the other three connectives could be used syntactically correctly in all critical items. a screenshot of the interface visible to participants is presented in figure 4. figure 4: interface visible for participants. english translation: “please click on the word you like best in the presentend sentence. the shopping mall has many customers, the stores in the city center are less frequented.” the connectives were chosen based on the discourse relation they can signal and the sentence structure they fit in, not primarily based on their frequency. in the german connective lexicon (dimlex; scheffler and stede, 2016), dahingegen is listed as the only contrastive connective that is not a signal for causal relations, too (stede et al., 2019). compared to the other connectives, it 17 seemann and scheffler appears much less frequently in german texts: the german referenzund zeitungskorpora of the digitales wörterbuch der deutschen sprache (‘digital dictionary of the german language’),12 see geyken et al. (2017) list 243,969 matches for deshalb between 1990 and 2018 (305.62 per million token), but only 26 for dahingegen (0.03 per million token). to ensure participants will choose the contrastive connective even though it is used less frequently, we conducted a pilot study, showing that participants use dahingegen despite its rarity. in this pilot, we tested a sub-sample of the final stimuli with students. as the participants of the pilot did choose the rare connectives in the contexts we expected them to be chosen in, we concluded that the frequency of a connective does not keep participants from using it in the context of this study. 7.4 analysis we fit a generalized linear mixed model with sum coding for the independent variables ‘discourse relation’ and ‘particle’, using the glmer function in the r package lme4 (bates et al., 2015). we coded the connectives chosen as either ‘causal’ (= 1) or ‘not causal’ (= 0). since the maximal random effect structure did not converge, we simplified the model to the maximal converging random effect structure, that is varying intercepts and varying slopes for particle presence, but not for relation. details on the model used are presented together with the results. 7.5 results overall, participants chose the causal connective more often if the sentence was presented with a modal particle. in the sentences without ja, the causal deshalb is chosen in 6.99% of all cases. in sentences presented with ja, this frequency increases to 17.34%. figure 5 shows participants’ connective choices for both contexts. the results of the generalized linear mixed-effects model are presented in table 3. fixed effects estimate std. error z-value p-value (intercept) -2.798 0.357 -7.841 < 0.001 ∅ vs. ja 0.940 0.238 3.950 < 0.001 discourse relation 0.064 0.684 0.094 0.9254 particle : relation 0.743 0.429 1.731 0.0834 table 3: output of the generalized linear mixed-effects model shown below. connective is coded as 1 = causal, 0 = not causal. sum coding is used for both particle and relation. glmer(connective ∼ particle * relation + (1+particle||participant) + (1+particle||item)) the model shows that the presence of ja has a significant effect on participants’ connective choice (p < 0.001). the main effect of discourse relation (p = 0.9) and the interaction between particle and relation (p = 0.08) are not statistically significant. in the causal-contrastive context, the frequency of causal deshalb increases from 7.79% to 15.59% in the presence of ja, in the causaltemporal context it increases from 6.18% to 19.08%. while the frequency increase is higher in the causal-temporal context, this difference is not found to be statistically significant by our model. 12. https://www.dwds.de/d/korpora/public 18 german modal particles as discourse signals figure 5: the bar plots show the absolute of the respective connective chosen (contrastive, causal, or temporal) based on whether the sentence was presented with or without the modal particle ja for both given contexts: causal-contrastive (left) and causal-temporal (right). 7.6 discussion figure 5 illustrates that both discourse contexts are mainly perceived as intended, with a contrastive or temporal reading respectively as the default and optional causal reading. both the graphs and the generalized linear mixed model show an increase of causal connectives chosen in the presence of ja. given that neither the main effect of discourse relation nor the interaction of relation and particle are statistically significant, the effect is arguably the same for both discourse contexts tested: adding the modal particle ja to an ambiguous discourse context leads to an increase of causal connectives chosen, independent of the type of relation displayed. however, even if the frequency of the causal connective deshalb increases in the presence of ja, it is still not the connective used most in either context. if the modal particle were a discourse connective, the proportion of causal interpretations ought to be higher. since discourse signals support certain interpretations rather than facilitating them on their own, they serve as main interpretation instructions only in the absence of connectives (das, 2014). often, discourse signals co-occur with connectives or other signals (hoek et al., 2018; das and taboada, 2019). hoek et al. (2018) distinguish three types of discourse signals: i) division of labour; signals that can fulfill the function of a connective in the absence of the same (e.g., lexical cues), ii) agreement; signals that fulfill the same function, but still co-occur together with a connective of similar type (e.g., modal verbs), and iii) general collocation; signals that always co-occur together with a connective (e.g., verbal tense). 19 seemann and scheffler how many signals occur (or co-occur) in a sentence depends on both the compatibility of signals present and the strength of the respective signals (cf. asr and demberg, 2012). in light of these observations, we consider ja to be type ii) or iii) marker that requires other signals to unambiguously indicate a certain discourse relation. related to this point, it should be noted that while the experimental items did not contain additional connectives, the presence of additional signals compatible with various discourse relations – such as verbal tense – could not be avoided. but given that the presence of ja has a significant effect on participants’ connective choice, we consider ja to be a causal discourse signal. this raises the question of the exact relation between ja and causality. as stated at the beginning of this section, we do not take the modal particle to be causal by itself. a causal connective will always lead to the relation between the two discourse segments it connects being interpreted as causal. modal particles, like all phrase-internal signals, do not connect segments but occur in one of them. they might indicate how the relation holding between the segment carrying the signal and the other segment is to be interpreted, but the same signal can indicate different discourse relations, as discussed by hoek et al. (2018): one of the non-connective discourse signals observed by das and taboada (2018), the numeral five, can indicate different discourse relations, depending on the context. if five elements in a list follow the numeral, an elaboration relation is signaled. but if the numeral is contrasted with another numeral, contrast might be the correct relation to infer from the same signal. the same holds for modal particles as signals. if a causal interpretation of a relation is possible, then the presence of ja indicates that this causal interpretation might be valid. this effect is likely to arise from the particle’s function of backgrounding information and marking it as potentially previously known: a statement marked by ja as salient and difficult to refuse can facilitate drawing a conclusion in an argument – and thus a causal interpretation. at this point, it cannot be decided whether ja directly signals a causal relation or whether it merely marks information as salient, thus facilitating a causal interpretation in a second step. independent of which explanation holds, the result of adding the particle ja to a sentence would be an increased probability of the discourse relation being interpreted as causal. thus we take our results as the first experimental evidence of ja as a discourse signal, though further studies are needed to pinpoint the exact contribution of modal particles to the interpretation of discourse, and their interaction with connectives or additional signals such as verbal tense. a last observation to be discussed is the frequent choice of both the contrastive dahingegen and the temporal währenddessen, even if sentences were presented with ja. in the case of contrastive dahingegen, this does not pose a problem, as both discourse signals and connectives can, in principle, signal any compatible relation.13 but the fact that ∼70% of the participants chose the temporal währenddessen in the presence of ja might appear unexpected. however, the above-mentioned restrictions of modal particles in temporal clauses only concern subordinate clauses (coniglio, 2012), and the sentences in our experiment were temporal main clauses. combining the results of this experiment with the acceptability rating of ja in the temporal circumstance relation (median = 3; which means this interpretation is disliked, but not completely unacceptable), this indicates that the modal particle is not generally infelicitous in the temporal context, but rather dispreferred. 13. studies documenting the co-occurrence of signals have been discussed above. lexicons of discourse connectives list more than one possible discourse relation for most of the included connectives (stede et al., 2019), and studies of individual signals in various languages also conclude that marking multiple relations is possible (schwenter, 2000; miltsakaki et al., 2005; pitler et al., 2008; mortier and degand, 2009; meyer et al., 2011; stede, 2014; costăchescu, 2017). 20 german modal particles as discourse signals 8. conclusion the present study outlines the role of german modal particles as interpretation instructions for discourse relations. we approach this question by first experimentally confirming previous theoretical and corpus-based predictions about the interaction of modal particles and discourse structure in the form of discourse relations, and directly testing a modal particle’s effect on relation disambiguation in a second step. for the first step, we conducted an acceptability study, asking participants to rate sentences expressing four different discourse relations. in those relations, we expected ja to be dispreferred in circumstance and condition, but not in evidence and justify. results show that participants are indeed sensitive to the predicted interaction between particles and discourse relations, as the ratings for the two causal relations were higher than for the other two relations if the sentence was presented with a modal particle. furthermore, we tested whether the position of the particle, nucleus (e.g., claim) or satellite (e.g., evidence or argument), of a relation influences the ratings. for the causal relations, we found higher ratings if the particles were presented in a relation’s satellite, but we could not show this effect for the dispreferred relations. in the second step, we conducted a forced-choice experiment, where participants were tasked with inserting a connective into an ambiguous context to indicate the perceived discourse relation between the presented segments. we found that the presence of the modal particle ja led to a significant increase in the use of the causal connective deshalb, indicating that ja might be interpreted as a causal discourse signal. with these results, our study presents first experimental evidence for modal particles serving as interpretation instructions in discourse. the experimental approach chosen here offers insights into the perception of modal particles, complementing previous corpus results on the distribution of modal particles over various discourse contexts. this study thus contributes to a series of studies on non-connective discourse signals as well as a discussion of the status of modal particles. but since we tested only one german modal particle, further research investigating modal particles’ effects on discourse is needed to pinpoint their exact contribution in discourse. given that we find the particle’s disambiguation effect to be rooted in its semantic properties, we expect to find the same effect for all particles with similar properties. testing not only more modal particles in german but also in other languages with particles that are used for similar functions – like dutch – will offer more insights on this. in addition, studying the contribution of further modal particles compatible with different contexts will help answer the question of how modal particles guide the interpretation of discourse relations. data availability both experiments were created in magpie14 and preregistered. the raw data, stimuli, javascript code, preregistration, and analysis for experiment 1 can be accessed via the open science framework.15 all corresponding data for experiment 2 can be found in the corresponding repository.16 declaration of interest the authors report there are no competing interests to declare. 14. https://magpie-ea.github.io/magpie-site/ 15. https://osf.io/xhdfj/?view_only=63ad745bdaf44583aaf527876764b50d 16. https://osf.io/9gp6c/?view_only=088a957ed6054db7b614bb3f47c1f51a 21 seemann and scheffler acknowledgements this work was funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) – project id 317633480 – sfb 1287. we thank merel scholman for her insights and help with the statistical analysis, and our reviewers and editor for their detailed comments and suggestions. we appreciate their work and think our paper profited a lot from it. 22 german modal particles as discourse signals appendix the presented definitions are based on the relation description defined in stede (2016a) and döring (2016). the examples given for each relation are taken from the potsdam commentary corpus (stede, 2016b), as translated by stede et al. (2017). circumstance the nucleus (n) is interpreted in light of the satellite (s), where s presents an event or state of affairs as context in time and space. example: [when veag came under pressure because of the deregulation of the electricity market,]s [they compensated for this by squeezing their suppliers.]n (maz-5297, pcc) condition the realization of n depends on the realization of s, where s presents a future event or hypothetical situation. example: [if the sanitary facilities are still not available in the coming season,]s [radewege is in danger of losing the competition for attracting boats to the campground.]n (maz-6488, pcc) evidence n presents a claim. s presents information a reader is likely to accept. accepting s raises a reader’s willingness to accept n. example: debate about having two subjects ’religion’ and ’ler’ at school [and now even our state government seems determined to remove this apparent equality between the two subjects.]n [stolpe, reiche and others do say yes to the possible compromise offered by the karlsruhe court, but they also decree: there cannot be any voluntary subject area ler/religion.]s (maz-6159, pcc) justify both n and s present subjective statements. s presents a reason or justification for the statement presented in n. example: [with each new day of air raids, the military operations of the u.s. lose more and more credibility.]n [by means of comprehensive area-wide destruction you can’t hit the taliban, nor can you eliminate bin laden.]s (maz-5701, pcc) 23 seemann and scheffler references werner abraham. discourse marker = discourse particle = thetical = modal particle? a futile comparison. in josef bayer and volker struckmeier, editors, discourse particles, pages 241– 280. de gruyter, 2017. isbn 978-3-11-049715-1. doi: 10.1515/9783110497151-010. werner abraham. modality in syntax, semantics and pragmatics, volume 165 of cambridge studies in linguistics. cambridge university press, 2020. doi: 10.1017/9781139108676. fatemeh torabi asr and vera demberg. measuring the strength of linguistic cues for discourse relations. proceedings of the workshop on advances in discourse analysis and its computational aspects (adaca), coling, mumbai, pages 33–41, 2012. douglas bates, martin mächler, ben bolker, and steve walker. fitting linear mixed-effects models using lme4. journal of statistical software, 67(1):1–48, 2015. doi: 10.18637/jss.v067.i01. douglas biber. stance in spoken and written university registers. journal of english for academic purposes, 5(2):97–116, 2006. issn 14751585. doi: 10.1016/j.jeap.2006.05.001. hardarik blühdorn. diskursmarker: pragmatische funktion und syntaktischer status. in hardarik blühdorn, arnulf deppermann, henrike helmer, and thomas spranz-fogasy, editors, diskursmarker im deutschen. reflexionen und analysen, pages 311–336. verlag für gesprächsforschung, göttingen, 2017. hardarik blühdorn. modalpartikeln und akzent im deutschen. linguistische berichte, 259:275– 317, 2019. anneloes r. canestrelli, willem m. mak, and ted j. m. sanders. causal connectives in discourse processing: how differences in subjectivity are reflected in eye movements. language and cognitive processes, 28(9):1394–1413, 2013. issn 0169-0965, 1464-0732. doi: 10.1080/01690965.2012.685885. r. h. b. christensen. ordinal—regression models for ordinal data, 2022. r package version 2022.11-16. https://cran.r-project.org/package=ordinal. marco coniglio. die syntax der deutschen modalpartikeln: ihre distribution und lizenzierung in hauptund nebensätzen. in die syntax der deutschen modalpartikeln, volume 73 of studia grammatica. akademie verlag, 2012. isbn 978-3-05-005357-8. doi: 10.1524/9783050053578. adriana costăchescu. discourse markers and discourse relations: the french dm quoi. in chiara fedriani and andrea sansó, editors, pragmatic markers, discourse markers and modal particles. new perspectives, volume 186 of studies in language companion series, pages 151–167. john benjamins publishing company, amsterdam, 2017. doi: 10.1075/slcs.186.06cos. ludivine crible. weak and strong discourse markers in speech, chat, and writing: do signals compensate for ambiguity in explicit relations? discourse processes, 57(9):793–807, 2020. issn 0163-853x, 1532-6950. doi: 10.1080/0163853x.2020.1786778. ludivine crible and vera demberg. the role of non-connective discourse cues and their interaction with connectives. pragmatics & cognition, 27(2):313–338, 2020. issn 0929-0907, 1569-9943. doi: 10.1075/pc.20003.cri. 24 german modal particles as discourse signals ludivine crible and pickering. compensating for processing difficulty in discourse: effect of parallelism in contrastive relations. discourse processes, 57(10):862–879, 2020. johannes dahl. die abtönungspartikeln im deutschen. ausdrucksmittel für sprechereinstellungen mit einem kontrastiven teil deutsch-serbokroatisch, volume 7 of deutsch im kontrast. julius groos verlag, heidelberg, 1988. debopam das. signalling of coherence relations in discourse. dissertation, simon fraser university, canada, 2014. debopam das and maite taboada. rst signalling corpus: a corpus of signals of coherence relations. language resources and evaluation, 52(1):149–184, 2018. issn 1574-020x, 1574-0218. doi: 10.1007/s10579-017-9383-x. debopam das and maite taboada. multiple signals of coherence relations. discours, 24, 2019. issn 1963-1723. doi: 10.4000/discours.10032. debopam das, maite taboada, and paul mcfetridge. rst signalling corpus, 2015. liesbeth degand and ted sanders. the impact of relational markers on expository text comprehension in l1 and l2. reading and writing, 15(7-8):739–757, 2002. liesbeth degand, bert cornillie, and paola pietrandrea, editors. discourse markers and modal particles: categorization and description. number 234 in pragmatics & beyond new series. john benjamins publishing company, amsterdam, philadelphia, 2013. isbn 978-90-272-5639-3. gabriele diewald. abtönungspartikel. in handbuch der deutschen wortarten, pages 117–141. de gruyter, berlin, new york, 2009. monika doherty. epistemische bedeutung, volume 32 of studia grammatica. akademie verlag, berlin, boston, 1985. doi: doi:10.1515/9783050067452. sophia döring. modal particles, discourse structure and common ground management. dissertation, humboldt-universität zu berlin, 2016. sophia döring and sophie repp. the modal particles ja and doch and their interaction with discourse structure: corpus and experimental evidence. in sam featherston, robin hörnig, sophie von wietersheim, and susanne winkler, editors, experiments in focus, pages 17–56. de gruyter, 2019. isbn 978-3-11-062309-3. doi: 10.1515/9783110623093-002. laura dörre, anna czypionka, andreas trotzke, and josef bayer. the processing of german modal particles and their counterparts. linguistische berichte, 255:313–346, 2018. markus egg and malte zimmermann. accented discourse particles: the case of ‘doch’. proceedings of sinn und bedeutung, 16:225–238, 2012. chiara fedriani and andrea sansò, editors. pragmatic markers, discourse markers and modal particles, volume 186 of studies in language companion series (slcs). john benjamins publishing company, amsterdam, philadelphia, 2017. bruce fraser. what are discourse markers? journal of pragmatics, 31(7):931–952, 1999. 25 seemann and scheffler alexander geyken, adrien barbaresi, jörg didakowski, bryan jurish, frank wiegand, and lothar lemnitzer. die korpusplattform des ‘digitalen wörterbuchs der deutschen sprache’ (dwds). zeitschrift für germanistische linguistik, 45(2):327–344, 2017. doi: doi:10.1515/zgl-2017-0017. bethany gray and douglas biber. stance markers. in karin aijmer and christoph rühlemann, editors, corpus pragmatics. a handbook, pages 219–248. cambridge university press, cambridge, 2015. cristina grisot and joanna blochowiak. temporal connectives and verbal tenses as processing instsructions. pragmatics & cognition, 24(3):404–440, 2017. daniel gutzmann. betonte modalpartikeln und verumfokus. in theo harden and elke hentschel, editors, 40 jahre partikelforschung, volume band 55 of stauffenburg linguistik, pages 119–138. narr, tübingen, 2010. daniel gutzmann. use-conditional meaning. number 6 in studies in multidimensional semantics. oxford university press, oxford, new york, 2015. daniel gutzmann. modal particles ̸= modal particles (= modal particles). in josef bayer and volker struckmeier, editors, discourse particles, pages 144–172. de gruyter, 2016. isbn 9783-11-049715-1. doi: 10.1515/9783110497151-007. dietrich hartmann. syntaktische funktionen der partikeln eben, eigentlich, einfach, nämlich, ruhig, vielleicht und wohl. zur grundlegung einer diachronischen untersuchung von satzpartikeln im deutschen. in harald weydt, editor, die partikeln der deutschen sprache, pages 121–138. de gruyter, berlin, new york, 1979. dietrich hartmann. art. particles. in jakob l. mey, editor, concise encyclopedia of pragmatics, pages 657–663. elsevier, oxford, 1998. gerhard helbig and joachim buscha. deutsche grammatik. ein handbuch für den ausländerunterricht. langenscheidt, berlin, münchen, wien, zürich, new york, 2001. elke hentschel. funktion und geschichte deutscher partikeln: ja, doch, halt und eben. in funktion und geschichte deutscher partikeln, volume 63 of reihe germanistische linguistik. max niemeyer verlag, 1986. doi: 10.1515/9783111371221. jet hoek, sandrine zufferey, jacqueline evers-vermeul, and ted j. m. sanders. the linguistic marking of coherence relations: interactions between connectives and segment-internal elements. pragmatics & cognition, 25(2):276–309, 2018. doi: 10.1075/pc.18016.hoe. lotte hogeweg, stefanie ramachers, and verena wottrich. doch , toch and wel on the table. linguistics in the netherlands, 28:50–60, 2011. doi: 10.1075/avt.28.05hog. judith kamalski. coherence marking, comprehension and persuasion: on the processing and representation of discourse. dissertation, universiteit utrecht, utrecht, 2007. elena karagjosova. the meaning and function of german modal particles. dissertation, universität des saarlandes, saarbrücken, 2004. 26 german modal particles as discourse signals suzanne kleijn, henk l.w. pander maat, and ted j.m. sanders. comprehension effects of connectives across texts, readers, and coherence relations. discourse processes, 56(5-6):447–464, 2019. issn 0163-853x, 1532-6950. doi: 10.1080/0163853x.2019.1605257. angelika kratzer. beyond ouch and oops: how descriptive and expressive meaning interact. in cornell conference on theories of context dependency, volume 26. cornell university ithaca, ny, 1999. angelika kratzer. interpreting focus: presupposed or expressive meanings? a comment on geurts and van der sandt. theoretical linguistics, 30(1):123–136, 2004. doi: doi:10.1515/thli.2004.002. gina r. kuperberg, martin paczynski, and tali ditman. establishing causal coherence across sentences: an erp study. journal of cognitive neuroscience, 23(5):1230–1246, 2011. doi: 10.1162/jocn.2010.21452. min-jae kwon. modalpartikeln und satzmodus. untersuchungen zur syntax, semantik und pragmatik der deutschen modalpartikeln. dissertation, ludwig-maximilians-universität, münchen, 2005. judith köhne-fuetterer, heiner drenhaus, francesca delogu, and vera demberg. the online processing of causal and concessive discourse connectives. linguistics, 59(2):417–448, 2021. issn 1613-396x, 0024-3949. doi: 10.1515/ling-2021-0011. ekkehard könig and susanne requard. a relevance-theoretic approach to the analysis of modal particles in german. multilingua. journal of cross-cultural and interlanguage communication, 10(1-2):63–78, 1991. doi: https://doi.org/10.1515/mult.1991.10.1-2.0. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text interdisciplinary journal for the study of discourse, 8(3): 243–281, 1988. issn 0165-4888, 1613-4117. doi: 10.1515/text.1.1988.8.3.243. daniel marcu. the theory and practice of discourse parsing and summarization. cambridge, london, mit press, 2000. thomas meyer, andrei popescu-belis, sandrine zufferey, and bruno cartoni. multilingual annotation and disambiguation of discourse connectives for machine translation. in joyce y. chai, johanna d. moore, rebecca j. passonneau, and david r. traum, editors, proceedings of the sigdial 2011 conference, pages 194–203, portland, oregon, 2011. association for computational linguistics. url https://aclanthology.org/w11-2022. keith k. millis and marcel a. just. the influence of connectives on sentence comprehension. journal of memory and language, 33(1):128–147, 1994. eleni miltsakaki, nikhil dinesh, rashmi prasad, aravind joshi, and bonnie webber. experiments on sense annotations and sense disambiguation of discourse connectives. in proceedings of the 4th workshop on treebanks and linguistic theories, barcelona, spain, 2005. liesbeth mortier and liesbeth degand. adversative discourse markers in contrast: the need for a combined corpus approach. international journal of corpus linguistics, 14(3):338–366, 2009. issn 1384-6655. doi: 10.1075/ijcl.14.3.03mor. publisher: john benjamins publishing company. 27 seemann and scheffler sonja müller. modalpartikeln. kurze einführungen in die germanistische linguistik 17. winter, heidelberg, 2014. isbn 978-3-8253-6365-9. sonja müller. alte und neue fragen der modalpartikel-forschung. linguistische berichte, 252: 383–441, 2017. emily pitler, mridhula raghupathy, hena mehta, ani nenkova, alan lee, and aravind joshi. easily identifiable discourse relations. in donia scott and hans uszkoreit, editors, coling 2008: companion volume: posters, pages 87–90, manchester, uk, 2008. coling 2008 organizing committee. url https://aclanthology.org/c08-2022. christopher potts. the logic of conventional implicatures. oxford university press, oxford, 2005. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in nicoletta calzolari, khalid choukri, bente maegaard, joseph mariani, jan odijk, stelios piperidis, and daniel tapias, editors, proceedings of the sixth international conference on language resources and evaluation (lrec), marrakech, morocco, 2008. european language resources association (elra). marie-paule péry-woodley, lydia-mai ho-dac, josette rebeyrolle, ludovic tanguy, and c‘ecile fabre. a corpus-driven approach to discourse organisation: from cues to complex markers. dialogue & discourse, 8(1):66–105, 2017. doi: 10.5087/dad.2017.103. laura reimer and andreas trotzke. the processing of secondary meaning: an experimental comparison of focus and modal particles in wh-questions. in daniel gutzmann, editor, secondary content, volume 37 of current research in the semantics / pragmatics interface, pages 143–167. brill, leiden, boston, 2019. tania rojas-esponda. a discourse model for überhaupt. semantics and pragmatics, 7:1–45, 2014a. issn 1937-8912. doi: 10.3765/sp.7.1. tania rojas-esponda. a qud account of german doch. proceedings of sinn und bedeutung, 18: 359–376, 2014b. ted sanders. coherence, causality and cognitive complexity in discourse. in m. aurnague and m. bras, editors, proceedings of the first international symposium on the exploration and modelling of meaning, pages 31–46, universite de toulouse-le mirail, 2005. ted sanders and leo noordman. the role of coherence relations and their linguistic markers in text processing. discourse processes, 29(1):37–60, 2000. issn 0163-853x, 1532-6950. doi: 10.1207/s15326950dp2901 3. ted sanders and wilbert spooren. discourse and text structure. in dirk geeraerts and hubert cuyckens, editors, the oxford handbook of cognitive linguistics, pages 916–941. oxford university press, oxford, new york, 2007. ted sanders, wilbert spooren, and leo noordman. toward a taxonomx of coherence relations. discourse processes, 15(1):1–35, 1992. 28 german modal particles as discourse signals ted sanders, vera demberg, jet hoek, merel scholman, fatemeh torabi asr, sandrine zufferey, and jacqueline evers-vermeul. unifying dimensions in coherence relations: how various annotation frameworks are related. corpus linguistics and linguistic theory, 17(1):1–71, 2021. doi: 10.1515/cllt-2016-0078. tatjana scheffler and manfred stede. adding semantic relations to a large-coverage connective lexicon of german. in proceedings of the tenth international conference on language resources and evaluation (lrec), pages 1008–1013, portorož, slovenia, 2016. merel scholman, vera demberg, and ted sanders. individual differences in expecting coherence relations: exploring the variability in sensitivity to contextual signals in discourse. discourse processes, 57(10):844–861, 2020. issn 0163-853x, 1532-6950. doi: 10.1080/0163853x.2020. 1813492. scott schwenter. viewpoints and polysemy: linking adversative and causal meanings of discourse markers. in elizabeth couper-kuhlen and bernd kortmann, editors, cause condition concession contrast. cognitive and discourse perspectives, volume 33 of topics in english linguistics, pages 257–282. de gruyter mouton, berlin ; new york, 2000. manfred stede. a prerequisite for discourse parsing: resolving connective ambiguity. in helmut gruber and gisela redeker, editors, the pragmatics of discourse coherence: theories and applications, pragmatics & beyond new series, pages 121–141. john benjamins publishing company, 2014. doi: 10.1075/pbns.254.05ste. manfred stede, editor. handbuch textannotation: potsdamer kommentarkorpus 2.0. number 8 in potsdam cognitive science series. universitätsverlag potsdam, potsdam, 2016a. isbn 978-386956-343-5. manfred stede. das potsdamer kommentarkorpus. in hartmut lenk, editor, persuasionsstile in europa ii. kommentartexte in den medienlandschaften europäischer länder, number 229-231 in germanistische linguistik, pages 177–202. olms, hildesheim, 2016b. manfred stede, maite taboada, and debopam das. annotation guidelines for rhetorical structure, 2017. unpublished manuscript. manfred stede, tatjana scheffler, and amália mendes. connective-lex: a web-based multilingual lexical resource for connectives. discours. revue de linguistique, psycholinguistique et informatique. a journal of linguistics, psycholinguistics and computational linguistics, discours (24), 2019. doi: 10.4000/discours.10098. number: 24 publisher: presses universitaires de caen. maite taboada. discourse markers as signals (or not) of rhetorical relations. journal of pragmatics, 38(4):567–592, 2006. doi: 10.1016/j.pragma.2005.09.010. maria thurmair. modalpartikeln und ihre kombinationen. number 223 in linguistische arbeiten. niemeyer, tübingen, 1989. isbn 978-3-484-30223-5. doi: 10.1515/9783111354569. maria thurmair. satztyp und modalpartikeln. in jörg meibauer, markus steinbach, and hans altmann, editors, satztypen des deutschen, pages 627–651. de gruyter, berlin ; new york, 2013. 29 seemann and scheffler geertje van bergen and hans rutger bosker. linguistic expectation management in online discourse processing: an investigation of dutch inderdaad ’indeed’ and eigenlijk ’actually’. journal of memory and language, 103:191–209, 2018. doi: 10.1016/j.jml.2018.08.004. yvonne viesel. discourse particles “embedded”: german ja in adjectival phrases. in josef bayer, editor, discourse particles. formal approaches to their syntax and semantics, linguistische arbeiten, pages 173–202. de gruyter, 2017. doi: 10.1515/9783110497151-008. yipu wei, willem m. mak, jacqueline evers-vermeul, and ted j. m. sanders. causal connectives as indicators of source information: evidence from the visual world paradigm. acta psychologica, 198:article 102866, 2019. doi: 10.1016/j.actpsy.2019.102866. yipu wei, jacqueline evers-vermeul, and ted j.m. sanders. the use of perspective markers and connectives in expressing subjectivity: evidence from collocational analyses. dialogue & discourse, 11(1):62–88, 2020. issn 2152-9620. doi: 10.5087/dad.2020.103. malte zimmermann. discourse particles. in p. portner, c. maienborn, and k. von heusinger, editors, semantics, volume 2 of handbücher zur sprachund kommunikationswissenschaft hsk, pages 2011–2038. mouton de gruyter, berlin, 2011. 30 dialogue & discourse 13(1) (2022) 1–40 doi: 10.5210/dad.2022.101 when to say what and how: adapting the elaborateness and indirectness of spoken dialogue systems juliana miehle juliana.miehle@uni-ulm.de institute of communications engineering ulm university wolfgang minker wolfgang.minker@uni-ulm.de institute of communications engineering ulm university stefan ultes stefan.ultes@daimler.com mercedes-benz ag research & development sindelfingen editor: amanda stent and barbara di eugenio submitted 01/2021; accepted 03/2022; published online 04/2022 abstract with the aim of designing a spoken dialogue system which has the ability to adapt to the user’s communication idiosyncrasies, we investigate whether it is possible to carry over insights from the usage of communication styles in human-human interaction to human-computer interaction. in an extensive literature review, it is demonstrated that communication styles play an important role in human communication. using a multi-lingual data set, we show that there is a significant correlation between the communication style of the system and the preceding communication style of the user. this is why two components that extend the standard architecture of spoken dialogue systems are presented: 1) a communication style classifier that automatically identifies the user communication style and 2) a communication style selection module that selects an appropriate system communication style. we consider the communication styles elaborateness and indirectness as it has been shown that they influence the user’s satisfaction and the user’s perception of a dialogue. we present a neural classification approach based on supervised learning for each task. neural networks are trained and evaluated with features that can be automatically derived during an ongoing interaction in every spoken dialogue system. it is shown that both components yield solid results and outperform the baseline in form of a majority-class classifier. keywords: communication styles, dialogue management, interactive adaptation, supervised learning, classification, neural approach 1. introduction even though intelligent assistants like amazon alexa, apple siri, google assistant or microsoft cortana are becoming increasingly popular, they do not consider different communication styles to adapt their behaviour. current systems focus on content (what is said) rather than formulation (how is it said). however, it has been shown that people adapt their interaction styles to one another across many levels of utterance production when communicating. ©2022 juliana miehle, wolfgang minker, and stefan ultes this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). miehle, minker and ultes speech recognition linguistic analysis dialogue management text generation speech synthesis application communication style classifier dialogue act selection communication style selection figure 1: the standard architecture of spoken dialogue systems is extended by two components: 1) a communication style classifier that automatically identifies the user communication style and 2) a communication style selection module that selects an appropriate system communication style. the goal of this article is to investigate if it is possible to carry over insights from the usage of communication styles in human-human interaction to human-computer interaction. building upon a long history of communication research for human-human interaction, we investigate if the used communication styles of the user and the system influence each other in human-computer interaction. to demonstrate the principal usage of communication style adaptation within a spoken dialogue system, the problem of identifying the user’s communication style and the problem of selecting the system’s communication style are framed as classification problems (see figure 1). thus, the main contributions of this article are as follows: 1. comprehensive overview over the general field of communication styles literature for humanhuman interaction 2. introduction to interactive adaptation for human-human and human-computer interaction 3. analysing the correlation between user and system communication style in human-computer interaction on a multi-lingual data set 4. a communication style classifier that automatically identifies the user communication style using supervised learning 5. a communication style selection module that selects an appropriate communication style of the system response using supervised learning 2 adapting the elaborateness and indirectness of spoken dialogue systems various studies suggest that adapting the communication styles of spoken dialogue systems to the individual users in a similar way to what humans do will lead to more natural interactions (stenchikova and stent, 2007; reitter et al., 2006; mairesse and walker, 2010). the work described in this paper builds upon and extends work published in (miehle et al., 2020) and considers the communication styles elaborateness and indirectness. pragst et al. (2019) have shown that both styles influence the user’s perception of a dialogue and are therefore valuable candidates for adaptive dialogue management. miehle et al. (2018b) have shown that varying the elaborateness and indirectness of a spoken user interface influences the user’s satisfaction and the user’s perception of the dialogue. the elaborateness thereby refers to the amount of additional information provided to the user and the indirectness describes how concretely the information that is to be conveyed is addressed by the speaker. the structure of the paper is as follows: in section 2, we introduce communication styles and interactive adaptation in human-human and human-computer interaction. related work on the adaptation and recognition of communication styles in human-computer interaction is discussed in section 3. in section 4, the corpus used in this work is described and the correlation between user and system communication style is investigated. we present the user communication style classifier in section 5 and the system communication style selection in section 6, before concluding in section 7. 2. communication styles and interactive adaptation in this section, we introduce communication styles and interactive adaptation in human-human and human-computer interaction. showing that these aspects play an important role in human communication, we provide the background for our work on communication style adaptation in spoken dialogue systems. a summary of the references is provided in table 1. 2.1 communication styles grice (1975) describes conversation as a cooperative activity where the talk exchanges consist of a succession of connected remarks. following his cooperative principle (“make your conversational contribution such as is required, at the stage at which it occurs, by the accepted purpose or direction of the talk exchange in which you are engaged.”), each speaker makes a statement in order to promote the purpose and objective of the conversation. this superordinate principle is divided into four categories, under each of which fall different maxims: • quantity: 1. make your contribution as informative as is required (for the current purposes of the exchange). 2. do not make your contribution more informative than is required. • quality: try to make your contribution one that is true. 1. do not say what you believe to be false. 2. do not say that for which you lack adequate evidence. • relation: be relevant. 3 miehle, minker and ultes topic references communication styles bultman and svarstad (2000) miehle et al. (2016) grice (1975) neuliep (2018) holtgraves (1986) pesch et al. (2015) kroeger (2019) pragst et al. (2019) madaio et al. (2017) searle (1975) miehle et al. (2018a) van dolen et al. (2007) miehle et al. (2018b) cultural models elliott et al. (2016) kaplan (1966) hofstede (2009) lewis (2010) interactive adaptation in human-human interaction branigan et al. (2000) nenkova et al. (2008) brennan and clark (1996) niederhoffer and pennebaker (2002) burgoon et al. (1995) pardo (2006) garrod and anderson (1987) pickering and garrod (2004) jungers et al. (2002) reitter et al. (2006) levelt and kelter (1982) schober (1993) interactive adaptation in human-computer interaction bell et al. (2003) coulston et al. (2002) bergmann et al. (2015) darves and oviatt (2002) branigan and pearson (2006) doran et al. (2003) branigan et al. (2010) koulouri et al. (2016) branigan et al. (2003) oviatt et al. (2004) brennan (1991) pearson et al. (2006) brennan (1996) suzuki and katagiri (2007) brennan and ohaeri (1994) table 1: summary of references on communication styles and interactive adaptation. 4 adapting the elaborateness and indirectness of spoken dialogue systems • manner: be perspicuous. 1. avoid obscurity of expression. 2. avoid ambiguity. 3. be brief (avoid unnecessary prolixity). 4. be orderly. the listener, for his/her part, naturally assumes that an utterance follows the cooperative principle, i.e. he/she presumes the speaker’s cooperation in the process of understanding the utterance. however, according to kroeger (2019), the cooperative principle is no code of conduct which has to be obeyed. a speaker may also break the maxims, as long as the hearer is able to recognise it. hence, a deliberate deviation from the principle can be used to communicate extra elements of meaning. meaning that is derived not from the words themselves, but from the way those words are used in a particular context, is thereby called conversational implicature (grice, 1975). these implications constitute an important part of our communication and form the basis for our work on communication style adaptation. one special type of conversational implicature is indirectness (kroeger, 2019). searle (1975) defines indirect speech acts as “cases in which one illocutionary act is performed indirectly by way of performing another”. a speech act is thereby an action that is performed by speaking, e.g. greeting, making a request, giving some information or giving an order. this means that a speaker utters a sentence and means not only what he/she says, but also something more. in contrast, in case of a direct speech act, a speaker utters a sentence and means exactly and literally what he/she says. searle provides the following example: speaker a: let’s go to the movies tonight. speaker b: i have to study for an exam. the utterance of speaker a is a direct proposal in virtue of its meaning. in contrast, the answer of speaker b is an indirect rejection of the proposal. literally, speaker b is making a statement, but within the given context, speaker a can infer that speaker b is rejecting the proposal as he/she is assuming that speaker b is cooperating in the conversation according to grice’s cooperative principle. therefore, speaker a assumes that the response of speaker b is relevant for the current conversation. as the literal statement is not an acceptance or rejection of the proposal, speaker b probably means more than he/she says. as speaker a knows that both studying for an exam and going to a movie takes a large amount of time relative to a single evening, he/she can infer that speaker b cannot do both in one evening. as he/she is not able to perform the proposed act, he/she is probably rejecting the proposal. similarly, kroeger (2019) describes a direct speech act as “one that is accomplished by the literal meaning of the words that are spoken”, whereas an indirect speech act is “one that is accomplished by implicature”. neuliep (2018) describes the indirect style as a “manner of speaking in which the intentions of the speaker are hidden or only hinted at during interaction” and the direct style as a “manner of speaking in which one employs overt expressions of intention”. another special type of conversational implication is the flouting of the first maxim of quantity (grice, 1975), i.e. being more elaborate or concise. this is for example the case if speaker a asks for some information and speaker b responds by not only giving the requested information, but also 5 miehle, minker and ultes some additional information like how certain the respective information or its evidence is. neuliep (2018) defines three levels for the quantity of talk: the elaborate style as the “mode of speaking that emphasises rich, expressive language”, the exacting style as “manner of speaking in which persons say no more or less than is needed to communicate a point” and the succinct style as “manner of concise speaking often accompanied by silence”. neuliep (2018) defines communication as the “simultaneous encoding, decoding and interpretation of verbal and nonverbal messages between people” that is dependent on the context in which it occurs, i.e. the cultural, physical, relational, and perceptual environment. thus, people communicate differently depending on their cultural background. this is consistent with various cultural models (hofstede, 2009; elliott et al., 2016; kaplan, 1966; lewis, 2010). according to neuliep (2018), the direct style is often used in individualistic, low-context cultures like, for example, the united states, england, australia and germany. in contrast, the indirect style is often seen in collectivistic, high-context cultures like the asian cultures. an elaborate style of communication is usually used in arab, middle eastern and afro-american cultures, whereas european americans generally prefer an exacting style, and a succinct style can be found in japan, china, and some native american/american indian cultures. however, the context of the speaker comprises more than just the culture. the message sent by a speaker is altered by where and with whom he/she interacts, what is the goal of the interaction and which effect he/she wants to achieve. miehle et al. (2016) investigated cultural differences in communication style preferences between the germans and the japanese. the results revealed that communication idiosyncrasies in human-human interaction may also be observed during human-computer interaction in a spoken dialogue system context. moreover, miehle et al. (2018a) presented another study examining five european cultures whose communication styles are much more alike than the german and japanese communication idiosyncrasies. the study explored not only the influence of the user’s culture but also of the gender, the frequency of use of speech based assistants as well as the system’s role. the results showed that the system’s role significantly influences the user’s preference in the system’s communication style whereas the frequency of use of speech based assistants has no influence. moreover, the findings showed differences among the cultures and, depending on the culture, there are gender differences with respect to the user’s preference in the system’s communication style. numerous studies have shown, that humans use different communication styles which have different effects on their interlocutor and the conversation. pesch et al. (2015) presented a study on how new product development is affected by communication style diversity in teams. the results showed that a diversity of communication styles in teams improves the creative environment within these teams and thus facilitates product innovativeness and speed to market of new product development. on the other hand, it also increases relationship conflicts that hamper a creative team environment. however, the beneficial effects outweigh the dysfunctional effects on the team innovation performance. the study of van dolen et al. (2007) examined online commercial group chat and, in particular, how the communication style of the advisor influences the effects of perceived technology attributes (perceived control, reliability, speed, and ease of use) and chat group characteristics (group involvement, similarity, and receptivity) on chat session satisfaction. the advisor used a task-oriented communication style (highly goal oriented and purposeful, giving direction and information, repeating, clarifying and evaluating information) and a socially oriented communication style (more personal and social, even to the extent of sometimes ignoring the task at hand, making jokes, showing understanding, using emoticons and rewarding the input of the customers). the results showed that the online chat advisor’s communication style influences the 6 adapting the elaborateness and indirectness of spoken dialogue systems importance of technology attributes to customers and causes different group dynamics to develop which influence customer satisfaction. bultman and svarstad (2000) examined how the communication style of physicians impacts the clients’ knowledge, initial beliefs, satisfaction, and adherence behaviour of individuals who have been prescribed a new medication for depression. the results of the study showed that a collaborative communication style enhances the clients’ knowledge and thus positively influences the treatment outcomes. it was not required that the given information was exhaustive, but it was required that the physician clearly communicated essential details (i.e. what to take, how much and when to take the antidepressant, when one can expect to begin feeling better, potential side effects and ways to alleviate these side effects, expected length of treatment, and a general idea of how the medication works). another interesting finding was that the physician communication style varied between the initial visit and follow-up visits, even with the same patient. the perception of direct and indirect speech was investigated by holtgraves (1986). the results indicated that the perceived appropriateness of an interactant’s choice regarding how to phrase a remark in a conversation may be affected by the social process of face management. indirect replies were perceived as more likely in face-threatening than non-face-threatening situations. when the situation was face-threatening, indirect replies that were evasive were perceived as more likely and polite than direct replies, and indirect replies were more likely to be accepted rather than challenged. madaio et al. (2017) explored the impact of peer tutors’ use of indirectness with feedback and instructions as well as the impact of the interpersonal closeness between tutor and tutee on the use of indirectness. the results showed that, in comparison with friend tutors, stranger tutors provided more positive feedback and used more indirect instructions. moreover, tutees attempted and solved more problems if the stranger tutor used indirect instructions. no such effect was found for friend tutors, indicating that relationship impacts students’ collaborative learning behaviours and that interpersonal closeness reduces the face-threat of direct instructions. pragst et al. (2019) investigated the applicability of elaborateness and indirectness as possibilities for adaptation in spoken dialogue systems. in order to do so, they compared four conditions: a high level of elaborateness and indirectness and an involved user, a low level of elaborateness and indirectness and an involved user, a high level of elaborateness and indirectness and a distracted user, and a low level of elaborateness and indirectness and a distracted user. the results showed multiple significant differences between the two levels of elaborateness and indirectness and that the assessment changes depending on the situation of the user. hence, it is concluded that elaborateness and indirectness influence the user’s perception of a dialogue and are therefore valuable candidates for adaptive dialogue management. miehle et al. (2018b) addressed the issues of how varying communication styles of a spoken user interface are perceived by users and whether there exist global preferences in the communication styles elaborateness and indirectness. the results showed that the system’s communication style influences the user’s satisfaction and the user’s perception of the dialogue and that there is no general preference in the system’s communication style, i.e. not every participant preferred the same communication style. based on that, we consider the elaborateness and indirectness to be highly suitable for our adaptation approach. 2.2 interactive adaptation it has been shown that people adapt their interaction styles to one another across many levels of utterance production when they communicate, e.g. by matching each other’s behaviour or synchronising the timing of behaviour. burgoon et al. (1995) reviewed a broad range of interaction adapta7 miehle, minker and ultes tion theories and models and presented their own interaction adaptation theory. according to their theory, adaptation in interaction is responsive to the needs, the expectations, and the desires of the communicators. a mechanistic theory of language processing, the interactive alignment model, was outlined in (pickering and garrod, 2004). it assumes that, in dialogue, the linguistic representations employed by the interlocutors become aligned at many levels, including the phonetic representation, the phonological representation, the lexical representation, the syntactic representation, the semantic representation and the situation model. this process of alignment is a largely automatic process which simplifies production and comprehension in dialogue. in the following, some studies that have investigated the phenomenon of interactive adaptation in human-human and human-computer interaction will be presented. 2.2.1 interactive adaptation in human-human interaction levelt and kelter (1982) investigated how speakers repeat materials from previous talk in questionanswering situations. the results of two experiments showed that a question’s surface form can affect the format of the answer given in the way that answers tend to match the prepositional form of the question, e.g. “(at) what time do you close?” – “(at) five o’clock.” the coordination of spatial descriptions has been explored by garrod and anderson (1987). it was shown that speakers adopted similar forms of descriptions, suggesting that interlocutors adapt their description styles to one another. schober (1993) investigated how speakers describe the locations of objects (from their own perspective, their addressee’s perspective, or some perspective that avoids choosing one or the other person) when performing a referential communication task. the results revealed that two speakers often used exactly the same or nearly identical words to describe the same display when communicating, showing that both partners actively collaborated with each other to ensure understanding. brennan and clark (1996) examined lexical entrainment, which describes the phenomenon that people in conversation use the same terms when referring repeatedly to the same object. after carrying out three experiments, the authors suggested that people are proposing a conceptualisation of an object when referring to it. their addressees may or may not agree to that proposal, but once a shared conceptualisation is established, both interlocutors appeal to it in later references. over time, speakers may simplify their conceptual pacts or abandon them for new ones. niederhoffer and pennebaker (2002) explored to which degree two people in conversation coordinate by matching their word use and how this coordination is related to the success or failure of the conversation. the results of their studies offered convincing evidence that individuals coordinate their word use on both the conversational level as well as on a turn-by-turn level. an unexpected finding was the lack of a relationship between the perceived interaction quality and the degree of linguistic style matching. nenkova et al. (2008) presented a corpus study examining entrainment in the use of high frequency words (i.e. the most common words in the corpus). the results showed that the degree of high-frequency word entrainment is positively correlated with task success, and that entrainment in high-frequency word usage is a good indicator of the perceived naturalness of a conversation. syntactic adaptation has been investigated by branigan et al. (2000). it was examined whether speakers in a dialogue tend to coordinate the syntactic structures of their contributions, irrespective of lexical and semantic content. the results revealed that, when comparing prepositional object structures and double object structures, speakers tend to produce a syntactic form that they have just heard the other dialogue participant use. reitter et al. (2006) examined two corpora of spoken 8 adapting the elaborateness and indirectness of spoken dialogue systems dialogues for syntactic repetitions. positive effects were found in both corpora, both for withinspeaker and between-speaker repetitions. however, the comparison of both corpora indicated that spontaneous conversation shows significantly less repetitions than task-oriented dialogue. jungers et al. (2002) examined whether speakers imitate the rate of a previously heard sentence when producing a sentence of analogous structure. in their experiment, the speakers’ sentence duration was significantly longer following a slow sentence than a fast sentence, and significantly shorter following a fast sentence than a slow sentence, but the speakers were also influenced by their own preferred production rate. therefore, the authors concluded that both the preferred rate and the rate of the previously heard sentence influence the produced rate. phonetic convergence during conversational interaction has been investigated by pardo (2006). by asking separate listeners to detect pronunciation similarity in a conversational speech corpus, it was determined whether pairs of talkers converged in phonetic repertoire over the course of a single interaction. the results showed a relatively rapid process of phonetic convergence between interacting talkers that is influenced by a talker’s role and sex, and that is persisting beyond the conversation that induces it. 2.2.2 interactive adaptation in human-computer interaction even if it has been shown that there exist clear differences in human-human interaction and humancomputer interaction (doran et al., 2003), numerous studies prove that interactive adaptation also occurs in the context of human-computer interaction. brennan (1991) compared keyboard conversations involving a simulated computer partner with those involving a human partner. in a wizardof-oz experiment, both the human and the simulated computer partner varied between three styles of responses: a short response containing only one or several words but no complete sentence, a sentence response, and a lexical change response without heed to the particular lexical items used in the adjacent query. the results showed both differences and similarities between a simulated computer partner and a human partner. there were significantly more acknowledgements, first-person and second-person pronouns and ellipses with the human partner. however, there was no difference in the number of third-person pronouns, showing that people expected connectedness across conversational turns between sentences and turns, regardless of whether they believed they were talking to a computer or another person. moreover, there were differences in the style of the participants’ queries. the first query was always a complete sentence with human partners, whereas with simulated computer partners, half the time the first query was a phrase or key words. as the dialogue proceeded, people adapted to their partners by designing queries that were more similar to their partners’ responses. in the last half of each dialogue, the mean percentage of complete sentences was not different across both kinds of partners, and was affected only by whether the response style was short or sentential. these results indicate that the design of the user’s utterances is shaped both by the initial model of the partner and also by the partner’s responses. in another wizard-of-oz experiment, brennan and ohaeri (1994) compared the effect of a telegraphic, a fluent and an anthropomorphic message style. the results showed no difference in the success of the participants and in their ratings about the perceived intelligence of the system. however, the language they used was shaped by the system’s message style. lexical convergence with computers has been investigated in (brennan, 1996). it was shown that people adopted the terms of their computer partners during text-based and speech-based interaction. lexical alignment has also been studied by koulouri et al. (2016). in a wizard-of-oz experiment, it was analysed whether speakers used the same words as 9 miehle, minker and ultes their partner. the results showed that the vocabulary stabilised early in the dialogue, suggesting the operation of lexical alignment between speakers. darves and oviatt (2002) examined whether the duration of children’s interspeaker response latencies is influenced by a computer partner’s speech output. four different voices were used in a study: male extrovert, male introvert, female extrovert and female introvert. the extrovert voices had a higher utterance rate (measured in syllables per second) and a shorter dialogue response latency. the results revealed that the children’s response latencies differed when they conversed with an animated character that spoke with the extrovert versus introvert voice: their response latencies increased when first exposed to the extrovert voice and then to the introvert, and decreased when first exposed to the introvert voice and then to the extrovert. in (coulston et al., 2002), the amplitude convergence in the children’s conversational speech with animated personas was investigated. it was shown that children actively adapted to the amplitude of their partner and even readapted when a new voice was was introduced. they increased their amplitude when interacting with a louder extroverted character, and dropped it with the quiet introverted one. in (oviatt et al., 2004), it was shown that, additionally to the adaptation of the amplitude and the interspeaker response latencies, the children also accommodated their utterance duration, their utterance rate and their utterance pause structure. the average utterance duration as well as the utterance rate increased when first interacting with the extrovert voice and then with the introvert one, and decreased when first interacting with the introvert voice and then with the extrovert one. the children’s average number of pauses and the total pause duration increased when the animated character’s voice switched from extrovert to introvert, and decreased when it switched from introvert to extrovert. the authors conclude that the observed changes in the children’s speech represented a substantial convergence towards their computer partner’s voice. however, as there was no perfect match, the children were not doing mimicry. bell et al. (2003) investigated whether people adapt their speaking rate while interacting with an animated character. the results confirmed that the users adapted to the speaking rate of the system, even if the subjects afterwards stated that they had not been aware of it. moreover, the speakers varied their speaking rate substantially in the course of the dialogue. slower speech was used during problematic sequences where subjects had to repeat or rephrase their utterance several times. prosodic adaptation has also been studied by suzuki and katagiri (2007). they found that the participants of their study aligned at least unidirectionally: the participants produced a louder voice when the system’s speech amplitude was increased, and a shorter pause duration when the system’s pause duration was decreased. however, no bidirectional adaptation was found. branigan et al. (2003) investigated syntactic alignment in typed communication via a computer. an experiment was conducted where the participants played a dialogue game in which they believed that they were interacting with either a person or a computer. the results demonstrated syntactic alignment for both conditions and suggested that it is largely an automatic process that is unmediated by consideration of the mental states of the interlocutor. in another experiment, pearson et al. (2006) showed that the users’ lexical alignment is influenced by their expectations about a system. when users believed the system to be unsophisticated and restricted in capability, they adapted their language to match the system’s language more than when they believed the system to be sophisticated and capable. this tendency was unaffected by the actual behaviour that the system exhibited. in (branigan and pearson, 2006), the findings of the studies were summarised and it was concluded that speakers tend to align both syntactically and lexically to both computer and human addressees. moreover, alignment in human-computer interaction seems to be even more important than in human-human interaction as it involves a stronger strategic component that is 10 adapting the elaborateness and indirectness of spoken dialogue systems designed to increase the likelihood of successful communication. possible mechanisms that might lead to linguistic alignment in human-computer interaction were discussed in (branigan et al., 2010). bergmann et al. (2015) explored lexical and gestural alignment with real and virtual humans. it was shown that adaptation takes place regarding communicative features (lexical alignment) as well as features without obvious communicative function (handedness alignment). 2.3 summary communication styles play an important role in human communication. we have introduced the theoretical background and the definitions of communication styles in general and for the elaborateness and indirectness in particular. these definitions are used throughout this work for annotations and classifications. furthermore, we have provided a broad review of studies investigating the phenomenon of interactive adaptation in human-human and human-computer interaction. it has been shown that people adapt their interaction styles to one another across many levels of utterance production when they communicate: they use the same words, coordinate their phonetic repertoire, their amplitude, their sentence and pause duration, the prepositional form and syntactic structures of their utterances, and the style of their messages–both when communicating with a human and a computer interaction partner. as the textual elements (i.e. how to formulate the utterance) are covered by the concept of communication styles, in the following we concentrate on this aspect. our aim is to recognise the user’s elaborateness and indirectness and adapt the system communication style accordingly. 3. adaptation and recognition of communication styles in this section, related work on the adaptation and recognition of communication styles in humancomputer interaction will be discussed. a summary of the references is provided in table 2. 3.1 adaptation of communication styles in human-computer interaction various studies suggest to adapt spoken dialogue systems to the users in a way similar to how people adapt to their interlocutors. for example, stenchikova and stent (2007) proposed two new approaches for measuring adaptation between dialogues and used these measures to study adaptation in a corpus of spoken dialogues. as these measures can identify features that exhibit variation and can be used to evaluate adaptation, it is proposed to incorporate models of adaptation to syntactic and lexical choice into spoken dialogue systems to enable the adaptation of these systems. by adapting the system’s behaviour to the user, the conversation agent may appear more familiar and trustworthy and the dialogue may be more effective. so far, communication styles have been used to create computer personalities and approaches for stylistic variation as well as for stylistic adaptation. we elaborate on this in the following sections. 3.1.1 development of computer personalities communication styles are a widely used medium to create computer personalities. nass et al. (1995) endowed their system with properties associated with a dominant or submissive personality. while the dominant version displayed high confidence and used strong language, assertions and commands, the submissive version displayed a low confidence level and used weaker language, questions and suggestions. the fundamental information conveyed by the system was thereby not 11 miehle, minker and ultes topic references development of computer personalities aly and tapus (2016) moon and nass (1996) andré et al. (2000) nass et al. (1995) irfan et al. (2020) oraby et al. (2018) isbister and nass (2000) smestad and volden (2019) mairesse and walker (2010) tapus and mataric (2008) mairesse and walker (2011) style variation de jong et al. (2008) gupta et al. (2007) porayska-pomsta and mellish (2004) hofs et al. (2010) wang et al. (2005) johnson et al. (2004) whittaker et al. (2003) kruijff-korbayová et al. (2008) wilkie et al. (2005) style adaptation ball and breese (2000) hu et al. (2018) brockmann et al. (2005) stenchikova and stent (2007) buschmeier et al. (2009) walker et al. (2007) hoegen et al. (2019) elaborateness recognition di buccio et al. (2014) gharouit and nfaoui (2017) indirectness recognition adel and schütze (2017) goel et al. (2019) aubakirova and bansal (2016) liscombe et al. (2005) danescu-niculescu-mizil et al. (2013) prokofieva and hirschberg (2014) dral et al. (2011) ulinski et al. (2018) forbes-riley and litman (2011) table 2: summary of references on the adaptation and recognition of communication styles in human-computer interaction. 12 adapting the elaborateness and indirectness of spoken dialogue systems changed. the results of a user study showed that the users recognised the computer’s personality. moreover, they preferred the system that displayed the personality that is similar to their own personality and were more satisfied with the interaction with this system in comparison to the system that used the dissimilar personality. in (moon and nass, 1996), it was additionally investigated how changes in the system’s dominance/submissiveness were perceived by the users. the results showed that changes in the direction towards a similar personality generated greater attraction than consistent similarity. isbister and nass (2000) created an extrovert and an introvert version of a computer character by use of verbal and non-verbal cues. the extroverted character used strong and friendly language in form of confident assertions that were relatively lengthy, poses with the limbs spread wide from its body, and postures that made the character seem to have moved closer to the participant. in contrast, the introverted character used weaker language in form of questions and suggestions that were relatively short, poses with the limbs closer in to its body, and did not ever appear to approach the participant. again, the fundamental information conveyed by the system was not changed, only the style of communicating the information. after conducting a user study, the results showed that the participants were able to identify both the verbal and the non-verbal personality cues. however, contrary to the previous studies, the participants preferred a character that had a personality that is complementary to their own personality, instead of a similar one. tapus and mataric (2008) also focused on the level of extroversion/introversion. the introverted version of a socially assistive therapist robot used vocal content that was nurturing and contained gentle and supportive language, as well as low pitch and volume. for the extroverted personality, a challenging language and high pitch and volume were used. the experimental results showed preference for a robot personality that matched the personality of the respective user. andré et al. (2000) introduced animated presentation teams with different character settings for the personality dimensions agreeableness, extroversion and openness. personality was conveyed by the choice of dialogue acts, the linguistic style (verbosity, specificity, force, formality, floridity and bias), the choice of semantic content, syntactic form, and acoustical realisation. feedback from users showed that they were able to identify the different personalities. smestad and volden (2019) designed a chatbot with an agreeable personality and one with a conscientious personality. both chatbots interacted through written input and output and were equal in all regards except their personalities. the differences in personality were displayed through the choice of language and tone of voice. the experimental results showed that the personality affected the user experience of the chatbots. irfan et al. (2020) modelled the emotional state of users and an agent to dynamically adapt the dialogue utterance selection of a system in multiparty interactions. a proof of concept user study demonstrated that the system can deliver and maintain distinct agent personalities. mairesse and walker (2010) presented a parameterizable language generator that provides a large number of parameters to support different linguistic styles in order to produce utterances matching particular personality profiles. these personality profiles were assigned fixed parameter values. an evaluation with human judges showed that the generated personality cues were reliably interpreted by humans. in (mairesse and walker, 2011), the same language generator was used with parameter estimation models trained using personality-annotated data. thus, generation parameters were estimated given target stylistic scores, which were then used by the generator to produce the output utterance. the results of a human evaluation showed that the trained models produced recognisable system personalities. oraby et al. (2018) used the generator to synthesise a new corpus of over 88,000 restaurant domain utterances whose linguistic style varies according to the personality models. this corpus has then been used to train three neural models. an evaluation of these 13 miehle, minker and ultes trained models showed that they both preserve semantic fidelity and exhibit distinguishable personality styles. aly and tapus (2016) used the generator in a humanoid robot and additionally explored the usage of gestures. the introverted robot used gestures that were narrow, slow and executed at a low rate, while the extroverted gestures were broad, quick and executed at a high rate. moreover, the generated speech content was adapted so that the robot gave more details in the extroverted condition than in the introverted condition. experimental results showed that the participants found the robot that adapted both the speech and the gestures more engaging than the robot that adapted only the speech. moreover, the majority of extroverted users preferred the extroverted robot, while the majority of introverted users preferred the introverted version. however, there were also some contrary preferences, even if they were not dominant. this variance in the perception of the robot behaviour reveals the difficulty in setting up clear borders and rules for the decision when which personality is preferred. 3.1.2 style variation obviously, there exist other applications than computer personalities. in the following, more general approaches to style variation are described. whittaker et al. (2003) investigated how conciseness can be realised in spoken dialogue systems. conciseness was thereby implemented by the number of attributes included in one option: concise descriptions mentioned only the highest weighted attribute, sufficient descriptions mentioned the top three weighted attributes, and verbose descriptions mentioned five attributes. kruijff-korbayová et al. (2008) described a multimodal in-car dialogue system with a template-based generator that generates and controls personal and impersonal style variation in the output. the dichotomy of the personal/impersonal style was defined in such a way that it primarily reflected a distinction in terms of agent activity: the personal style involved the explicit realisation of an agent (e.g. “i’ve found three songs.”), while the impersonal style avoided it (e.g. “three songs have been found.”). porayska-pomsta and mellish (2004) defined a natural language model for a tutoring system with strategies for a positive or negative face. a positive face was thereby defined as a person’s need to be approved of by others, while a negative face was defined as a person’s need for autonomy from others. the strategies differed in the amount of content specificity (i.e. how specific and how structured the feedback is) and illocutionary specificity (i.e. how explicitly accepting or rejecting the tutor’s feedback is). they were characterised in terms of the degree to which each of them accommodates the user’s need for autonomy and approval and selected based on these dimensions. another tutoring system that models politeness was presented by johnson et al. (2004). natural language templates were defined and assigned positive and negative politeness values. during an interaction, the template matching the target politeness values most closely was selected. a wizardof-oz experiment to evaluate the interaction tactics where the participants were randomly assigned to either a polite or a direct treatment was conducted in (wang et al., 2005). the results showed that the polite agent had a positive impact on the students’ learning gains. wilkie et al. (2005) integrated politeness strategies for system-initiated digressions in a mass-market telephone banking dialogue. templates for a positive face redress were optimistic, informal, intensifying interest with the addressee, exaggerating approval with the addressee, presupposing common ground, showing concern for the addressee’s wants, offering and promising, giving or asking for reasons. templates for a negative face redress were pessimistic, indirect, apologising, stating the face-threatening act as a general rule, impersonalising the speaker and the addressee, giving deference, going on record 14 adapting the elaborateness and indirectness of spoken dialogue systems as not indebting the addressee. in contrast to these templates used to mitigate positive and negative face threats, the bald templates were direct and concise. experimental results showed no general preference for one of the strategies. gupta et al. (2007) presented a system combining a spoken language generator with an artificial intelligence planner to model politeness in collaborative taskoriented dialogue. a direct strategy (e.g. “do x.”), an approval strategy (e.g. “could you please do x mate?”), an autonomy strategy (e.g. “could you possibly do x for me?”) and an indirect strategy (e.g. “x is not done yet.”) were used to model different levels of politeness, and different linguistic forms were defined to model each strategy. these politeness strategies have also been used in the conversational agent described in (de jong et al., 2008) and (hofs et al., 2010) that can help users to find their way in a virtual environment, while adapting its politeness to that of the user. in each turn, a pre-generated sentence template with politeness tags was selected depending on the politeness value of the system that is calculated based on the system’s previous politeness level and the user’s politeness level. 3.1.3 style adaptation besides the realisation of style variation, approaches to adaptation were examined. walker et al. (2007) presented a two-stage sentence planner for providing restaurant information in different styles. it randomly generates multiple alternative realisations of an information presentation which differ in how the content is allocated into sentences, how the sentences are ordered and which discourse cues are used to express the relationships between content elements. these alternative realisations are ranked using a statistical model trained on human feedback. brockmann et al. (2005) used an approach for ranking alternative utterance candidates to simulate the effect of syntactic alignment in natural language generation. ball and breese (2000) presented an architecture that uses models of emotions and personality encoded as bayesian networks. one is used to diagnose the emotions and personality of the user, and a second one to generate an appropriate behaviour for the agent by selecting scripted paraphrases that are related to its emotional state and personality. however, the agent’s mood and personality might only match that of the user or be the exact opposite of the user. buschmeier et al. (2009) presented an alignment-capable microplanner that models the interactive alignment behaviour of human speakers for different microplanning tasks (lexical choice, syntactic choice, referring expression generation and aggregation). the alignment behaviour is calculated based on the recency of use by the system itself, the recency of use by the interlocutor, the frequency of use by the system itself and the frequency of use by the interlocutor. hoegen et al. (2019) developed an end-to-end voice-based conversational agent that is able to align with the interlocutor’s conversational style. the conversational style is categorised on an axis ranging from high consideration to high involvement. the agent uses content variables (pronoun use, repetition, and utterance length) and acoustic variables (speech rate, pitch, and loudness) to calculate the user’s conversational style and to match the participant on these conversational style variables. hu et al. (2018) proposed an adaptation measure which can model adaptation on any subset of linguistic features and can be applied on a turn by turn basis during the dialogue to control adaptation in natural language generation. the method was applied to multiple corpora to investigate how the dialog situation and speaker roles influenced the level and type of adaptation to the interlocutor. it was shown that the adaptation varied depending on the feature sets, the conversational situations, the dialogue initiative and the course of the dialogue. however, the application of the measure to natural language generation was left to future work. 15 miehle, minker and ultes 3.2 recognition of elaborateness and indirectness previous work has already explored approaches for the classification of elaborateness and indirectness in the context of related applications. di buccio et al. (2014) proposed a methodology to automatically detect and process verbose queries submitted to search engines. it was shown that the information retrieval effectiveness can be significantly improved by considering the query verbosity. moreover, gharouit and nfaoui (2017) suggested to use babelnet as knowledge base in the detection of verbose queries and then presented a comparative study between different algorithms to classify queries into two classes, verbose or succinct. however, both papers deal with the classification of queries submitted to search engines. to the best of our knowledge, there exists no previous work in the field of elaborateness classification for spoken language. goel et al. (2019) explored different supervised machine learning approaches to automatically detect indirectness in tutoring conversations. the authors collected a corpus of tutoring dialogues from 12 american-english speaking pairs of teenagers whereby the conversations included social interaction as well as tutoring periods. they annotated four types of indirectness for the tutoring periods, namely apologising (e.g. “sorry, its negative 2.”), hedging language (e.g. “you just add 5 to both sides.”), the use of vague category extenders (e.g. “you have to multiply and stuff.”) and subjectivising (e.g. “i think you divide by 3 here.”). each utterance was then classified as direct or indirect based on its inclusion in any of these categories. afterwards, they used different classification approaches to detect indirectness based on textual and visual features, reaching an f1 sore of 62%. however, the literature presented in section 2.1 suggests that there are more aspects than the four types of indirectness annotated in this corpus and that indirectness cannot be broken down to rather simple key word spotting (e.g. “sorry”, “just”, “and stuff”, “i think”). in this work, the definition of neuliep (2018) is used which describes the indirect style as a “manner of speaking in which the intentions of the speaker are hidden or only hinted at during interaction” (see section 2.1) and the directness/indirectness is annotated and classified in a global way and not based on fixed structures or key words. other work in this field only focused on a specific phenomena of indirect speech, like hedge detection (prokofieva and hirschberg, 2014; ulinski et al., 2018), politeness detection (danescuniculescu-mizil et al., 2013; aubakirova and bansal, 2016) and uncertainty detection (liscombe et al., 2005; dral et al., 2011; forbes-riley and litman, 2011; adel and schütze, 2017). 3.3 summary regarding the adaptation of communication styles in human-computer interaction, so far, work has focused on alignment and on the realisation of communication style variation in natural language generation, both for general variation and for the development of computer personalities. however, it has also been shown that alignment is not always the appropriate system reaction. depending on numerous parameters that influence an interaction between two participants, like the speakers’ roles, their cultures, their personalities or the aim of the interaction, the appropriate or preferred speaking style or system personality differ. therefore, we argue that the decision about which communication style is to be used by a spoken dialogue system at which time needs to be covered by the dialogue management to ensure that the relevant parameters can be included in the decision process. regarding the recognition of elaborateness and indirectness in spoken dialogue systems, only little previous work has been done. for elaborateness, only queries submitted to search engines have been examined, and for indirectness, merely different categories have been explored. in contrast, in 16 adapting the elaborateness and indirectness of spoken dialogue systems this work, indirectness is classified in a more global way and not based on fixed structures or key words, and elaborateness is classified in spoken language. 4. investigating the correlation between user and system communication style in order to analyse whether the communication style of the system is correlated to the communication style of the user, we have created a corpus1 with annotations for elaborateness and indirectness for each user and system dialogue act. for the communication style annotations, we have used the definitions presented in section 2.1. our data set is based on recordings on health care topics containing spontaneous interactions in dialogue format between two participants: one is taking the role of the system while the other one is taking the role of the user of that system. for scenarios where medical knowledge is required, the participant who takes the role of the system is legally qualified to practise a medical profession, i.e. is a trained medical doctor or nurse. each dialogue turn contains one or more user dialogue acts followed by one or more system dialogue acts. these dialogue acts are chosen out of a set of 47 distinct dialogue acts which have been predefined. a list of all dialogue acts can be found in table 3. along with the dialogue acts, the respective utterances are also added to the data set. an example dialogue is shown in table 4. overall, the corpus covers 258 dialogues containing 2,880 turns and 7,930 annotated dialogue actions. the dialogues are in four different languages: german, polish, spanish and turkish. the language distribution is shown in table 5. it can be seen that the distribution of dialogue acts per dialogue (da/d) varies among the languages: while german and turkish are very similar, there is a difference compared to polish and spanish. this is the case even though the task and the familiarity between the speakers were identical for the different languages. pairs of speakers did not know each other and were swapped (i.e. speaker a did not always talk to speaker b, but also to other speakers). hence, we conclude that the differences in the distribution are due to differences in the languages/cultures. each dialogue act has been annotated with the two communication styles indirectness and elaborateness. both are assigned scores between 1 and 5 which have been defined as follows: 1 means that the utterance is extremely direct/concise, i.e. the speaker used the most direct/concise option to give the requested information. for the indirectness dimension, this means that the information is conveyed very concretely and the listener can understand it literally and does not have to imply anything. for the elaborateness dimension, this means that only the most important information is given by use of as few words as possible. for example, a response to the question about tomorrow’s weather forecast rated with 1 for indirectness and elaborateness would be: “it will rain.” the higher the rating for indirectness, the more hidden are the intentions of the speaker (2 = slightly indirect, 5 = extremely indirect). the higher the rating for elaborateness, the more additional information is given (2 = slightly elaborate, 5 = extremely elaborate). for instance, an indirect response to the question about tomorrow’s weather forecast would be an advice to take an umbrella, and an elaborate response would result in providing the weather forecast for the next few days. an example dialogue with annotated elaborateness (e) and indirectness (i) scores is shown in table 4. each dialogue act was annotated by three different raters. they were instructed with annotated sample dialogues. moreover, uncertainties were discussed in a weekly meeting where the raters talked about concrete annotation sentences and clarified how to understand a sentence in the specific context. however, in these meetings no decision was made regarding which value to annotate, only 1unfortunately, it is not possible to publish the corpus due to privacy reasons. 17 miehle, minker and ultes dialogue acts accept personalapologise acknowledge personalgreet advise personalbye afternoongreet personalthank afternoonbye readnewspaper answerthank reject askmood repeatpreviousutterance askplans rephrasepreviousutterance asktask request askwellbeing requestadditionalinfo cheerup requestmissinginfo console requestnewspaper declare requestreasonforemotion eveninggreet requestrepeat eveningbye requestrephrase explicitlyconfirmrecognisedinput requestweather implicitlyconfirmrecognisedinput sharejoy individualisticallyorientedmotivate showweather meetagainbye simpleapologise morninggreet simplegreet morningbye simplemotivate obligate simplebye order simplethank personalanswerthank table 3: list of predefined dialogue acts. 18 adapting the elaborateness and indirectness of spoken dialogue systems role utterance annotation user hello kristina. personalgreet, e: 1, i: 1 system hi, nice to meet you. simplegreet, e: 3, i: 1 user i have got a problem. declare, e: 1, i: 1 system you can ask me everything. declare, e: 1, i: 1 user okay. acknowledge, e: 1, i: 1 i have heard that old people should drink a lot and i have the impression that my father does not drink enough. declare, e: 3, i: 2 system how much does he drink on average every day? requestmissinginfo, e: 1, i: 1 user in the morning he likes a cup of coffee, at lunchtime i always give him a glass of water or sometimes juice, but when i come back after he has eaten, there’s usually still half of it in the glass. yes, and in the evening, sometimes a glass of beer, but not much more. declare, e: 4, i: 2 system do you have the impression that he has very dry skin? requestmissinginfo, e: 2, i: 1 user not really. reject, e: 2, i: 2 well, i always put cream on his hands and face and also on his head, so i couldn’t say that now, but i would have to see how his arms and legs are. declare, e: 3, i: 2 would that be a sign that he is not drinking enough? request, e: 1, i: 2 system it is a sign that he is not drinking enough. accept, e: 4, i: 1 user okay. acknowledge, e: 1, i: 1 system is he very dry in the mouth? requestmissinginfo, e: 1, i: 1 user i’ve never paid attention to that before. declare, e: 2, i: 2 system it is said that old people should drink about 1.3 litres a day. declare, e: 2, i: 2 user i don’t think he drinks that much. declare, e: 1, i: 1 system try to give him something he enjoys drinking. advise, e: 2, i: 1 user okay, i will try that. acknowledge, e: 3, i: 1 system it might also be nice if you drink together. for example, you could have a cup of tea together in the afternoon. advise, e: 4, i: 3 user that’s a great idea! acknowledge, e: 3, i: 2 thank you very much! simplethank, e: 2, i: 1 system you’re welcome. answerthank, e: 1, i: 1 user see you next time! meetagainbye, e: 2, i: 1 system bye. simplebye, e: 1, i: 1 table 4: example dialogue with annotated dialogue acts and elaborateness/indirectness scores. 19 miehle, minker and ultes d da da/d german 135 4,887 36.20 spanish 52 1,002 19.27 polish 42 1,017 24.21 turkish 29 1,024 35.31 overall 258 7,930 30.74 table 5: language distribution of the dialogues in the annotated corpus, whereby d is the number of dialogues and da is the number of dialogue acts. elaborateness 1 2 3 4 5 german 1,782 1,850 795 312 148 spanish 295 242 139 118 208 polish 273 383 198 95 68 turkish 323 391 179 76 55 overall 2,673 2,866 1,311 601 479 indirectness 1 2 3 4 5 german 3,825 840 142 78 2 spanish 681 296 8 17 0 polish 744 249 4 20 0 turkish 777 216 12 19 0 overall 6,027 1,601 166 134 2 table 6: class distribution of the annotated elaborateness and indirectness scores (median of the three ratings). the meaning of ambiguous sentences was clarified so that the annotators could rate them afterwards on the basis of a common understanding. the class distribution of the annotated elaborateness and indirectness scores (median of the three ratings) is shown in table 6. it can be seen that the classes 1 and 2 are the most common for both the elaborateness and the indirectness. the classes 3, 4, and 5 contain utterances which are elaborate/indirect to a greater or lesser extent and the weekly meetings with the annotators revealed that it is quite hard to distinguish between different levels of elaborateness and indirectness. hence, we combined the classes 3, 4, and 5 to one new class, reducing the corpus to three classes. for indirectness, the annotation showed that it even makes sense to see it as a binary decision between direct/indirect utterances. as the classes 2-5 contain different degrees of indirectness (from slightly indirect to extremely indirect), we additionally combined these classes into one indirect class for binary classification. 20 adapting the elaborateness and indirectness of spoken dialogue systems elaborateness (5 classes) r1/r2 r1/r3 r2/r3 av. κ 0.560 0.515 0.516 0.530 ρ 0.848 0.813 0.799 0.820 icc 0.934 elaborateness (3 classes) r1/r2 r1/r3 r2/r3 av. κ 0.670 0.612 0.608 0.630 ρ 0.826 0.794 0.767 0.796 icc 0.916 indirectness (5 classes) r1/r2 r1/r3 r2/r3 av. κ 0.315 0.423 0.368 0.369 ρ 0.387 0.504 0.442 0.444 icc 0.686 indirectness (3 classes) r1/r2 r1/r3 r2/r3 av. κ 0.335 0.439 0.382 0.385 ρ 0.387 0.504 0.441 0.444 icc 0.695 indirectness (2 classes) r1/r2 r1/r3 r2/r3 av. κ 0.376 0.499 0.440 0.438 ρ 0.377 0.500 0.440 0.439 icc 0.701 table 7: agreement (κ), correlation (ρ) and reliability (icc) in elaborateness and indirectness of the three ratings (r1, r2, r3). all results are significant at the 0.001 level. 21 miehle, minker and ultes elaborateness elaborateness indirectness indirectness indirectness (5 classes) (3 classes) (5 classes) (3 classes) (2 classes) u5/s1 0.202* 0.184* 0.107* 0.107* 0.096* u5/smd 0.243* 0.219* 0.144* 0.143* 0.138* umd/s1 0.175* 0.154* 0.089* 0.087* 0.080* umd/smd 0.219* 0.189* 0.132* 0.131* 0.128* table 8: correlation between the last user action u5 and the first system action s1 of each turn as well as the median of all user and system actions of the respective turn umd and smd in terms of spearman’s rank correlation coefficient rho ρ. all results marked with (*) are significant at the 0.01 level. we have analysed the quality of the annotated scores by use of the following measures: cohen’s kappa κ, spearman’s rank correlation coefficient rho ρ and the intraclass correlation coefficient icc. the results can be seen in table 7. the original ratings (five classes) achieve an overall inter-rater agreement of κ = 0.53 for elaborateness and κ = 0.37 for indirectness, a correlation of ρ = 0.82 for elaborateness and ρ = 0.44 for indirectness and a inter-rater reliability of icc = 0.93 for elaborateness and icc = 0.69 for indirectness. if we reduce the classes to three or two (in case of indirectness), we obtain a higher agreement while the correlation and the inter-rater reliability do not change significantly. overall, we have a good inter-rater reliability for both communication styles given the difficulty of the annotation task. in order to use the communication style annotations as target for our classification tasks, we need a final score to be calculated from the three ratings. as we have applied an ordinal scale for the ratings, we have used the median of the three ratings. according to stevens (1946), this is the appropriate measure for ordinal scales. using this corpus, we analysed whether the communication style of the speaker who assumed the role of the system (hereafter referred to as system) is correlated to the communication style of the speaker who assumed the role of the user (hereafter referred to as user). the purpose of this is to find out whether the system should take into account the user’s communication style when selecting its communication style. in section 2, it was already shown that humans adapt their communication styles during an interaction. however, we want to show that this is also the case in the current setting. in order to do so, we extracted the 2,880 user-system exchanges (i.e. the single turns where the system responds to a user inquiry) and the respective elaborateness and indirectness annotations. one exchange contains up to five consecutive user actions u and up to four consecutive system actions s. therefore, we analysed the correlation between the last user action u5 and the first system action s1 of each turn as well as the median (md) of all user and system actions of the respective turn umd and smd. the results in terms of spearman’s rank correlation coefficient rho ρ for both elaborateness and indirectness in five, three and two classes can be seen in table 8. all results are significant at the 0.01 level which shows that there is a significant correlation between the communication style of the system and the preceding communication style of the user. moreover, the results show that the highest correlation is between the last user action u5 and the median of the subsequent system actions smd. the correlation between the last user action u5 and the median of all system actions of the respective turn smd for the different languages is shown in table 9. 22 adapting the elaborateness and indirectness of spoken dialogue systems elaborateness elaborateness indirectness indirectness indirectness (5 classes) (3 classes) (5 classes) (3 classes) (2 classes) german 0.138* 0.137* 0.128* 0.127* 0.124* spanish 0.378* 0.368* 0.140* 0.138* 0.115** polish 0.240* 0.235* 0.235* 0.233* 0.223* turkish 0.354* 0.320* 0.104** 0.103** 0.104** table 9: correlation between the last user action u5 and the median of all system actions of the respective turn smd in terms of spearman’s rank correlation coefficient rho ρ for the different languages. all results marked with (*) are significant at the 0.01 level, all results marked with (**) are significant at the 0.05 level. parameter grid #nodes 3, 6, 12, 24, 48, 96, 144, 192 #epochs 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 1000 optimiser adadelta, adam, nadam, adagrad, sgd, rmsprop output function sigmoid, softmax loss function categorical crossentropy (cc), mean squared error (mse) table 10: grid of parameter values for the user communication style classifier. it can be seen that there is a significant correlation for both elaborateness and indirectness for all four languages. however, the effect size for elaborateness varies between languages. while there is a small correlation for german and polish, there is a medium correlation for spanish and turkish (according to cohen (1977)). as the task and the familiarity between speakers were identical for the different languages, we conclude that the discrepancy is due to differences in the languages/cultures. 5. user communication style classifier as we have shown that there is a significant correlation between the user and the system communication style, in (miehle et al., 2020) we presented a classification approach to automatically estimate the user’s communication style during an ongoing dialogue. the estimated communication style can then be used in the dialogue management to adapt the system behaviour to the user, as depicted in figure 1. we utilise a supervised learning approach with a multi-layer perceptron (mlp) classifier with one hidden layer. a comparison with a support vector machine (svm) classifier and a recurrent neural network (rnn) classifier consisting of two long short-term memory (lstm) layers can be found in (miehle et al., 2020). to mitigate overfitting, the neural net is trained and evaluated with a 10-fold cross-validation setting on the german part of the corpus described in section 4. grid search is used to find the best set of hyper parameters (i.e. the amount of nodes, the amount of epochs, the optimiser, the output function and the loss function). the grid of parameter values can be found in table 10. to account for the imbalanced data during the grid search optimisation, the unweighted average recall (uar) was used, which is the arithmetic average of all class-wise recalls. the best parameter values that were chosen as final setting for the models are shown in ta23 miehle, minker and ultes features #nodes #epochs optimiser output function loss function da 48 250 nadam sigmoid cc da+g 48 250 nadam softmax mse u 48 50 adadelta softmax cc u+da 48 50 adadelta softmax cc u+da+g 48 50 adadelta softmax cc ub 48 50 adadelta sigmoid cc ub+da 48 50 adadelta sigmoid cc ub+da+g 48 50 adadelta softmax cc we 144 250 nadam sigmoid cc we+da 48 250 nadam sigmoid cc we+da+g 144 250 nadam sigmoid cc da 48 250 nadam softmax mse da+g 48 250 nadam sigmoid mse u 48 50 adagrad sigmoid cc u+da 48 50 adagrad sigmoid cc u+da+g 48 50 adagrad sigmoid cc ub 48 50 adadelta softmax cc ub+da 48 50 adadelta softmax cc ub+da+g 48 50 adadelta softmax cc we 144 1000 nadam sigmoid cc we+da 48 350 nadam sigmoid cc we+da+g 48 350 nadam sigmoid cc da 48 250 nadam softmax cc da+g 48 250 nadam softmax cc u 48 50 adagrad sigmoid cc u+da 48 50 adagrad sigmoid cc u+da+g 48 50 adagrad sigmoid cc ub 48 50 adagrad sigmoid cc ub+da 48 50 adagrad sigmoid cc ub+da+g 48 50 adagrad sigmoid cc we 48 350 adadelta sigmoid mse we+da 48 350 adadelta sigmoid cc we+da+g 48 350 adadelta sigmoid cc table 11: the best parameter values for the user communication style classifier (top: 3-class elaborateness, middle: 3-class indirectness, bottom: 2-class indirectness). 24 adapting the elaborateness and indirectness of spoken dialogue systems elaborateness indirectness indirectness (3 classes) (3 classes) (2 classes) da uar 0.840 0.555 0.753 acc 0.838 0.832 0.848 f1 0.838 0.582 0.761 κ 0.749 0.467 0.527 ρ 0.862 0.523 0.541 table 12: classification results using dialogue act features (da) in terms of unweighted average recall (uar), accuracy (acc), f1-score, cohen’s kappa κ and spearman’s rank correlation coefficient rho ρ. ble 11. we define a dialogue act feature set and investigate how grammatical and linguistic features influence the performance. 5.1 the dialogue act features as a first approach, the mlp was trained using only dialogue act features (da) that can directly be derived from the data2. these features contain the dialogue act and the amount of words in the utterance of the corresponding dialogue act. note that the dialogue act is the output of the linguistic analysis while the text representation of the utterance is the output of the speech recogniser (see figure 1). hence, both features in this feature set can be automatically derived during an ongoing interaction in every spoken dialogue system and no annotation is necessary. the results are shown in table 12. classification of the 3-class elaborateness reaches an uar of 84% only using dialogue act features, which is quite promising. classification of the 3-class indirectness results in an uar of 56%, and the binary indirectness reaches an uar of 75%. the results for indirectness clearly show the difficulty of the task, as was already shown by the corpus creation. there, it was quite hard for the annotators to distinguish between different levels of indirectness so that the class distribution of indirectness is sub-optimal for the classification task. however, comparing the results to a majorityclass classifier clearly shows that there is still a lot of information encoded in the da feature set achieving higher uar. the majority-class classifier always predicts the most frequent class in the training set and achieves an uar of 33% for three classes and an uar of 50% for two classes. 5.2 the contribution of grammatical and linguistic features to address the question of whether grammatical features (g) improve the estimation of the communication style, a second feature set is used containing the dialogue act features as well as grammatical features. for the grammatical features, part-of-speech (pos) tags are assigned to the utterances using the rdrpostagger (nguyen et al., 2014) and the number of each pos tag per utterance is 2during our experiments, we also tested additional annotated features (the amount of topics being talked about in the current utterance, the speaker’s culture, gender, age, year of birth, country of birth, country of residence and whether he/she played the role of the user or the system, as well as the system role and the number of the dialogue act in the current dialogue), but this led to worse results. 25 miehle, minker and ultes elaborateness indirectness indirectness (3 classes) (3 classes) (2 classes) da+g uar 0.841 0.558 0.753 acc 0.840 0.834 0.848 f1 0.839 0.588 0.761 κ 0.753 0.470 0.526 ρ 0.864 0.521 0.540 table 13: classification results using dialogue act features as well as grammatical features (da+g) in terms of unweighted average recall (uar), accuracy (acc), f1-score, cohen’s kappa κ and spearman’s rank correlation coefficient rho ρ. counted. as the utterance is the output of the speech recognition and this tagger can be used online during an ongoing interaction, there is also no annotation necessary for this feature set. the results are shown in table 13. it can be seen that there is no improvement in comparison to using only the dialogue act features. in addition to grammatical features, linguistic features may greatly contribute to the overall classification performance. in order to encode linguistic features, a bag-of-words (bow) approach was used in combination with unigrams (u), unigrams and bigrams (ub) and word embeddings (we). using bow and the corpus presented in section 4, two distinct vocabularies were created: • the bow-u vocabulary contains every word occurring in the database. • the bow-ub vocabulary contains the bow-u vocabulary (single words) as well as every two-word-sequence in the database. these vocabularies and the combination with word embeddings led to three different linguistic feature sets: • u: this feature set contains a bow-u vector for each utterance, thus encoding the number of times each word (of the overall vocabulary) appears in the corresponding utterance. • ub: this feature set contains a bow-ub vector for each utterance, thus encoding the number of times each word and each two-word-sequence (of the overall vocabulary) appear in the corresponding utterance. • we: for this feature set, the bow-u vocabulary has been combined with the german pretrained fasttext word vectors by grave et al. (2018)3. matrix x of dimension u×w contains the bow-u vectors (dimension 1 × w with w the amount of words in vocabulary bow-u) for each utterance, where u is the total number of utterances. matrix w of dimension w × p contains the fasttext word vectors (dimension 1 × p with p the length of each word vector) for each word. by multiplying these matrices a new matrix z = x ·w of dimension u × p is obtained, containing a vector representation for each utterance. these utterance vectors of dimension 1× p can then be used as feature vectors for the classification task. 3during our experiments, we also tested self-trained word vectors, but this led to worse results. 26 adapting the elaborateness and indirectness of spoken dialogue systems elaborateness indirectness indirectness (3 classes) (3 classes) (2 classes) u uar 0.747 0.485 0.729 acc 0.752 0.822 0.842 f1 0.742 0.478 0.744 κ 0.618 0.430 0.492 ρ 0.779 0.490 0.503 u+da uar 0.809 0.484 0.743 acc 0.811 0.823 0.846 f1 0.807 0.477 0.755 κ 0.708 0.433 0.512 ρ 0.831 0.507 0.522 u+da+g uar 0.817 0.484 0.746 acc 0.818 0.822 0.846 f1 0.814 0.476 0.757 κ 0.719 0.431 0.516 ρ 0.841 0.505 0.524 ub uar 0.745 0.520 0.748 acc 0.742 0.751 0.822 f1 0.734 0.497 0.740 κ 0.607 0.354 0.481 ρ 0.776 0.411 0.485 ub+da uar 0.786 0.533 0.748 acc 0.785 0.761 0.826 f1 0.781 0.511 0.742 κ 0.669 0.387 0.485 ρ 0.811 0.452 0.490 ub+da+g uar 0.799 0.542 0.756 acc 0.796 0.757 0.827 f1 0.793 0.513 0.747 κ 0.687 0.391 0.495 ρ 0.827 0.458 0.500 table 14: classification results using linguistic features encoded as unigrams (u) or unigrams and bigrams (ub) (separately and in combination with dialogue act features and grammatical features) in terms of unweighted average recall (uar), accuracy (acc), f1-score, cohen’s kappa κ and spearman’s rank correlation coefficient rho ρ. 27 miehle, minker and ultes elaborateness indirectness indirectness (3 classes) (3 classes) (2 classes) we uar 0.757 0.493 0.727 acc 0.755 0.783 0.828 f1 0.749 0.495 0.729 κ 0.626 0.364 0.464 ρ 0.786 0.414 0.479 we+da uar 0.825 0.589 0.762 acc 0.821 0.803 0.842 f1 0.819 0.589 0.759 κ 0.726 0.443 0.522 ρ 0.855 0.498 0.535 we+da+g uar 0.827 0.594 0.765 acc 0.823 0.794 0.843 f1 0.821 0.588 0.762 κ 0.729 0.432 0.528 ρ 0.857 0.480 0.544 table 15: classification results using linguistic features encoded as word embeddings (we) (separately and in combination with dialogue act features and grammatical features) in terms of unweighted average recall (uar), accuracy (acc), f1-score, cohen’s kappa κ and spearman’s rank correlation coefficient rho ρ. in addition to using these linguistic feature sets individually, we used them in combination with dialogue act features (da) and grammatical features (g). the results with the u and ub feature sets are shown in table 14, the results with the we feature set can be found in table 15. for elaborateness, the best results are achieved with the dialogue act feature set. grammatical and linguistic features do not seem to have any effect on the classification performance. this leads to the conclusion that for elaborateness, analysing the utterance length dependent on the dialogue act seems to contain enough information to achieve good classification performance. for indirectness, the overall performance improves by using linguistic information encoded as word embeddings. this in combination with grammatical and dialogue act features (we+da+g) led to uars of 59% and 76% for the estimation of the indirectness using three classes and two classes, respectively. using the bow approach in combination with unigrams and bigrams did not improve the classification performance. to sum up, linguistic features are beneficial for the estimation of indirectness, but not for the estimation of elaborateness. for the latter, the dialogue act features (i.e. the dialogue act and the amount of words in the utterance) seem to be sufficient. all features can be automatically recognised during an ongoing interaction in any spoken dialogue system, without any prior annotation. hence, this user communication style classifier can be used as additional component for any spoken dialogue system. 28 adapting the elaborateness and indirectness of spoken dialogue systems 1 2 3 elaborateness 736 1,310 834 indirectness 1,973 817 90 table 16: class distribution of the annotated elaborateness and indirectness scores for the 2,880 dialogue turns. 6. system communication style selection in this section, the task of automatically selecting the system communication style during an ongoing interaction with a spoken dialogue system is addressed. as depicted in figure 1, this is part of the dialogue management so that it not only decides what is said next, but also how. we suggest that the system communication style depends on two components: 1) the content of the system dialogue act (what the system wants to say in the current turn) and 2) the reaction to the user (what the user wants from and how the user talks to the system). as we obtained promising results for the classification of the user communication styles by using a supervised learning approach with a multi-layer perceptron (see section 5), we think that this approach is also suitable for the task at hand. hence, we utilise a mlp classifier with one hidden layer. to mitigate overfitting, the neural net is trained and evaluated with a 10-fold crossvalidation setting on the 2,880 turns of the corpus described in section 4. the class distribution for both communication styles is shown in table 16. grid search is used to find the best set of hyper parameters (i.e. the amount of nodes, the amount of epochs, the optimiser, the output function and the loss function). the grid of parameter values can be found in table 17. to account for the imbalanced data during the grid search optimisation, the uar is used. the best parameter values chosen as final setting for the models are shown in table 18. for each of the 2,880 dialogue turns, we extracted the following features: • the system dialogue acts (s) • the user dialogue acts (u) • the amount of words in the utterance of the corresponding user dialogue acts (w) • the user communication styles (cs) • the language (german, polish, spanish or turkish) (l) during our experiments, we also tested part-of-speech tags and sentence embeddings (based on the respective utterances), though without improvement of the results. note that all features can be automatically derived during an ongoing interaction in every spoken dialogue system and no annotation is necessary. the user dialogue acts are the output of the linguistic analysis while the text representation of the utterance is the output of the speech recogniser. the system dialogue acts are the output of the dialogue act selection in the dialogue manager and the user communication styles may be classified by use of the communication style classifier described in section 5 (see figure 1). however, in order to focus on the performance of the system communication style selection module 29 miehle, minker and ultes parameter grid #nodes 10, 25, 50 #epochs 10, 50, 100, 200, 500 optimiser adadelta, adam, nadam, adagrad output function sigmoid, softmax loss function categorical crossentropy (cc), mean squared error (mse) table 17: grid of parameter values for the system communication style selection. features #nodes #epochs optimiser output function loss function s+l 50 200 adagrad softmax cc w+u+cs+l 50 10 adagrad sigmoid cc s+cs+l 25 100 adam sigmoid cc s+w+u+cs+l 50 100 adadelta sigmoid cc s+l 50 200 nadam sigmoid mse w+u+cs+l 50 100 adagrad softmax cc s+cs+l 25 500 adam softmax cc s+w+u+cs+l 50 100 adam sigmoid cc s+l 25 10 nadam softmax cc w+u+cs+l 50 100 adagrad softmax mse s+cs+l 10 10 nadam softmax cc s+w+u+cs+l 25 10 nadam softmax cc table 18: the best parameter values for the system communication style selection (top: 3-class elaborateness, middle: 3-class indirectness, bottom: 2-class indirectness). and avoid errors caused by the user communication style classification, the ground truth labels for the communication styles from section 4 have been used for the following analysis. the number of classes of the user communication style has been adjusted to the number of classes of the system communication style (i.e. either three or two classes based on the task). the results are shown in table 19. it can be seen that both the system dialogue act (s+l) and the information about the user (w+u+cs+l) contain relevant information for the selection of the system communication style. overall, classification of the 3-class elaborateness reaches an uar of 63%. classification of the 3-class indirectness results in an uar of 50%, and the binary indirectness reaches an uar of 68%. the comparatively poor results of the 3-class indirectness classification can be explained by the data distribution. for the 2-class indirectness, the combination of the system dialogue act and all available user information provides the best result. for the 3-class elaborateness, the best result is obtained by using the system dialogue act in combination with the user communication style (s+cs+l) and there is no improvement when adding the user dialogue act and the amount of words of the respective utterance. this shows that all relevant information about the user is covered by the user communication style. 30 adapting the elaborateness and indirectness of spoken dialogue systems elaborateness indirectness indirectness (3 classes) (3 classes) (2 classes) s+l uar 0.625 0.495 0.673 acc 0.651 0.745 0.760 f1 0.636 0.523 0.686 w+u+cs+l uar 0.535 0.409 0.617 acc 0.560 0.702 0.708 f1 0.542 0.406 0.622 s+cs+l uar 0.634 0.484 0.675 acc 0.660 0.731 0.756 f1 0.644 0.499 0.686 s+w+u+cs+l uar 0.627 0.471 0.684 acc 0.647 0.724 0.756 f1 0.635 0.486 0.694 table 19: classification results for the system communication style selection using different feature sets in terms of unweighted average recall (uar), accuracy (acc) and f1-score. elaborateness indirectness indirectness (3 classes) (3 classes) (2 classes) overall uar 0.634 0.495 0.684 acc 0.660 0.745 0.756 f1 0.644 0.523 0.694 german uar 0.567 0.465 0.649 acc 0.649 0.745 0.743 f1 0.579 0.493 0.659 polish uar 0.584 0.439 0.643 acc 0.615 0.715 0.725 f1 0.591 0.439 0.650 spanish uar 0.766 0.552 0.818 acc 0.805 0.797 0.797 f1 0.768 0.535 0.797 turkish uar 0.506 0.539 0.619 acc 0.586 0.742 0.760 f1 0.520 0.563 0.630 table 20: classification results for the system communication style selection in terms of unweighted average recall (uar), accuracy (acc) and f1-score of the overall test set and the individual languages. 31 miehle, minker and ultes elaborateness indirectness indirectness (3 classes) (3 classes) (2 classes) u5 uar 0.416 0.398 0.571 acc 0.412 0.586 0.618 f1 0.409 0.389 0.569 umd uar 0.399 0.396 0.566 acc 0.406 0.582 0.609 f1 0.398 0.391 0.563 table 21: classification results for the system communication style selection baseline which is mimicking the last user communication style u5 or the median of all previous user communication styles umd in terms of unweighted average recall (uar), accuracy (acc) and f1-score. when dividing the test set based on the languages, we can see that the classification works differently for the individual languages (see table 20). for the 3-class elaborateness, we achieve an uar of 57% for german, 58% for polish, 77% for spanish and 51% for turkish. for the 2class indirectness, the classification results in an uar of 65% for german, 64% for polish, 82% for spanish and 62% for turkish. the differences between the languages indicate cultural differences, as already revealed by studies like (miehle et al., 2016) and (miehle et al., 2018a). for example, in the latter it has been shown that spanish people like significantly more elaborateness than the other european cultures that have been investigated. since the elaborateness is more dominant in the training data (see table 16), the trained classifier better suits spanish than the other cultures. however, this result might also be due to our limited data. this needs to be investigated in future work. comparing the results to a majority-class classifier clearly shows that there is a lot of information encoded. moreover, a baseline classifier which is mimicking the user communication style reaches an uar of 42% for the 3-class elaborateness, 40% for the 3-class indirectness and 57% for the binary indirectness when using the communication style of the last user action u5 of the current turn. when using the median communication style of all user actions umd of the current turn, the results are even worse, as can be seen in table 21. hence, our trained system communication style selection module clearly outperforms a model which is just mimicking the user communication style at each turn. 7. conclusion and future directions in this work, we have introduced communication styles and interactive adaptation for human-human and human-computer interaction. in a broad literature review, we have demonstrated that those aspects play an important role in human communication. people adapt their interaction styles to one another across many levels of utterance production when they communicate: they use the same words, coordinate their phonetic repertoire, their amplitude, their sentence and pause duration, the prepositional form and syntactic structures of their utterances, and the style of their messages–both when communicating with a human and a computer interaction partner. throughout this work, we 32 adapting the elaborateness and indirectness of spoken dialogue systems have focused on adaptation based on communication styles (i.e. how to formulate the utterance) due to the following reasons: 1. there is a strong theoretical background that allows to generalise the adaptation across multiple domains and applications. 2. the information required for this adaptation can be obtained during ongoing interactions. 3. there is a verified influence of communication style adaptation on user satisfaction (miehle et al., 2018b). using a multi-lingual data set with elaborateness and indirectness annotations, we have shown that there exists a significant correlation between the communication style of the system and the preceding communication style of the user. moreover, we have augmented the standard architecture of spoken dialogue systems with the ability to adapt to the user’s communication idiosyncrasies. it has been extended by two components: 1) a communication style classifier that automatically identifies the user communication style and 2) a communication style selection module that selects an appropriate communication style of the system response. we have presented a neural classification approach for each task. for the user communication style classifier, a supervised learning approach has been utilised in order to estimate the user’s elaborateness and indirectness. we have trained and evaluated a multi-layer perceptron with features that can be automatically derived during an ongoing interaction in every spoken dialogue system. we have tested different feature sets as input for our classifier and performed classification in two and three classes. the results show that the elaborateness can be classified quite well by only using the dialogue act and the amount of words contained in the corresponding utterance. the indirectness seems to be a more difficult classification task and additional linguistic features in form of word embeddings give improvement in the classification results. for the system communication style selection, we have used the same supervised learning approach. using features that encode what the system wants to say in the current turn (i.e. the system dialogue acts), what the user wants from the system (i.e. the user dialogue acts) and how the user talks to the system (i.e. the amount of words in the utterance of the corresponding user dialogue acts and the user communication styles), we trained and evaluated a multi-layer perceptron. as for the first task, these features can be automatically recognised during an interaction in every spoken dialogue system. the results outperform both a majority-class classifier and a baseline which is mimicking the last user communication style for each of the four languages, reaching an uar of 63% for the classification of the 3-class elaborateness and an uar of 68% for the 2-class indirectness. when combining both components, the spoken dialogue system is enabled to recognise the user’s communication style and select an appropriate communication style for the system. so far, we have shown that both components (evaluated separately) yield solid results. in future work, we plan to conduct an evaluation of the overall system with real users. the user study presented in (miehle et al., 2018b) showed that the system communication style selection has a direct influence on the user satisfaction. based on these results, we will investigate whether a targeted increase in user satisfaction is achieved by our system. in order to do so, a spoken dialogue system that incorporates the user communication style classifier and the system communication style selection 33 miehle, minker and ultes will be compared with a system that does not contain these modules. moreover, we will consider a reinforcement learning approach (instead of the herein presented supervised learning approach) as the system communication style selection in a spoken dialogue system might also depend on what the system and the user want to achieve in the long run. acknowledgements this work is part of a project that has received funding from the european union’s horizon 2020 research and innovation programme under grant agreement no 645012. we thank our colleagues from the university of tübingen, the german red cross in tübingen and semfyc in barcelona for organising and carrying out the corpus recordings. additionally, this work has received funding within the bmbf project “robotkoop: cooperative interaction strategies and goal negotiations with learning autonomous robots” and the technology transfer project “do it yourself, but not alone: companion technology for diy support” of the transregional collaborative research centre sfb/trr 62 “companion technology for cognitive technical systems” funded by the german research foundation (dfg). we thank the anonymous reviewers and the editors for their constructive comments which helped us to improve the paper. references heike adel and hinrich schütze. exploring different dimensions of attention for uncertainty detection. in proceedings of the 15th conference of the european chapter of the association for computational linguistics: volume 1, long papers, pages 22–34, 2017. amir aly and adriana tapus. towards an intelligent system for generating an adapted verbal and nonverbal combined behavior in human–robot interaction. autonomous robots, 40(2):193–209, 2016. elisabeth andré, thomas rist, susanne van mulken, martin klesen, and stefan baldes. the automated design of believable dialogues for animated presentation teams. embodied conversational agents, pages 220–255, 2000. malika aubakirova and mohit bansal. interpreting neural networks to improve politeness comprehension. in proceedings of the 2016 conference on empirical methods in natural language processing, pages 2035–2041, 2016. gene ball and jack breese. emotion and personality in a conversational agent. embodied conversational agents, pages 189–219, 2000. linda bell, joakim gustafson, and mattias heldner. prosodic adaptation in human-computer interaction. in proceedings of icphs, volume 3, pages 833–836. citeseer, 2003. kirsten bergmann, holly p. branigan, and stefan kopp. exploring the alignment space–lexical and gestural alignment with real and virtual humans. frontiers in ict, 2:7, 2015. holly p. branigan and jamie pearson. alignment in human-computer interaction. how people talk to computers, robots, and other artificial communication partners, pages 140–156, 2006. 34 adapting the elaborateness and indirectness of spoken dialogue systems holly p. branigan, martin j. pickering, and alexandra a. cleland. syntactic co-ordination in dialogue. cognition, 75(2):b13–b25, 2000. holly p. branigan, martin j. pickering, jamie pearson, janet f. mclean, and clifford i. nass. syntactic alignment between computers and people: the role of belief about mental states. in proceedings of the 25th annual conference of the cognitive science society, pages 186–191. lawrence erlbaum associates, 2003. holly p. branigan, martin j. pickering, jamie pearson, and janet f. mclean. linguistic alignment between people and computers. journal of pragmatics, 42(9):2355–2368, 2010. susan e. brennan. conversation with and through computers. user modeling and user-adapted interaction, 1(1):67–86, 1991. susan e. brennan. lexical entrainment in spontaneous dialog. proceedings of issd, 96:41–44, 1996. susan e. brennan and herbert h. clark. conceptual pacts and lexical choice in conversation. journal of experimental psychology: learning, memory, and cognition, 22(6):1482, 1996. susan e. brennan and justina o. ohaeri. effects of message style on users’ attributions toward agents. in conference companion on human factors in computing systems, pages 281–282, 1994. carsten brockmann, amy isard, jon oberlander, and michael white. modelling alignment for affective dialogue. in workshop on adapting the interaction style to affective factors at the 10th international conference on user modeling, 2005. dara c. bultman and bonnie l. svarstad. effects of physician communication style on client medication beliefs and adherence with antidepressant treatment. patient education and counseling, 40(2):173 – 185, 2000. judee k. burgoon, lesa a. stern, and leesa dillman. interpersonal adaptation: dyadic interaction patterns. cambridge university press, 1995. hendrik buschmeier, kirsten bergmann, and stefan kopp. an alignment-capable microplanner for natural language generation. in proceedings of the 12th european workshop on natural language generation, pages 82–89, 2009. jacob cohen. statistical power analysis for the behavioral sciences. academic press, 1977. rachel coulston, sharon oviatt, and courtney darves. amplitude convergence in children’s conversational speech with animated personas. in seventh international conference on spoken language processing, pages 2689–2692, 2002. cristian danescu-niculescu-mizil, moritz sudhof, dan jurafsky, jure leskovec, and christopher potts. a computational approach to politeness with application to social factors. in proceedings of the 51st annual meeting of the association for computational linguistics (volume 1: long papers), pages 250–259, sofia, bulgaria, 2013. association for computational linguistics. 35 miehle, minker and ultes courtney darves and sharon oviatt. adaptation of users’ spoken dialogue patterns in a conversational interface. in proceedings of the 7th international conference on spoken language processing, pages 561–564, 2002. markus de jong, mariët theune, and dennis hofs. politeness and alignment in dialogues with a virtual guide. in proceedings of the 7th international joint conference on autonomous agents and multiagent systems volume 1, pages 207–214, 2008. emanuele di buccio, massimo melucci, and federica moro. detecting verbose queries and improving information retrieval. information processing & management, 50(2):342–360, 2014. christine doran, john aberdeen, laurie damianos, and lynette hirschman. comparing several aspects of human-computer and human-human dialogues. in current and new directions in discourse and dialogue, pages 133–159. springer, 2003. jeroen dral, dirk heylen, and rieks op den akker. detecting uncertainty in spoken dialogues: an exploratory research for the automatic detection of speaker uncertainty by using prosodic markers, pages 67–77. springer netherlands, 2011. candia elliott, r. jerry adams, and suganya sockalingam. multicultural toolkit: toolkit for cross-cultural collaboration. awesome library. http://www.awesomelibrary.org/ multiculturaltoolkit.html, 2016. accessed: 2016-05-01. kate forbes-riley and diane j. litman. benefits and challenges of real-time uncertainty detection and adaptation in a spoken dialogue computer tutor. speech communication, 53(9):1115 – 1136, 2011. simon garrod and anthony anderson. saying what you mean in dialogue: a study in conceptual and semantic co-ordination. cognition, 27(2):181–218, 1987. kenza gharouit and el habib nfaoui. a comparison of classification algorithms for verbose queries detection using babelnet. in 2017 intelligent systems and computer vision (iscv), pages 1–5. ieee, 2017. pranav goel, yoichi matsuyama, michael madaio, and justine cassell. i think it might help if we multiply, and not add: detecting indirectness in conversation. in 9th international workshop on spoken dialogue system technology, pages 27–40. springer singapore, 2019. edouard grave, piotr bojanowski, prakhar gupta, armand joulin, and tomas mikolov. learning word vectors for 157 languages. in proceedings of the eleventh international conference on language resources and evaluation (lrec 2018). european language resources association (elra), 2018. herbert p. grice. logic and conversation. in speech acts, pages 41–58. brill, 1975. swati gupta, marilyn a. walker, and daniela m. romano. how rude are you?: evaluating politeness and affect in interaction. in international conference on affective computing and intelligent interaction, pages 203–217. springer, 2007. 36 adapting the elaborateness and indirectness of spoken dialogue systems rens hoegen, deepali aneja, daniel mcduff, and mary czerwinski. an end-to-end conversational style matching agent. in proceedings of the 19th acm international conference on intelligent virtual agents, pages 111–118, 2019. dennis hofs, mariët theune, and rieks op den akker. natural interaction with a virtual guide in a virtual environment. journal on multimodal user interfaces, 3(1-2):141–153, 2010. geert hofstede. culture’s consequences: comparing values, behaviors, institutions and organizations across nations. sage, 2009. thomas holtgraves. language structure in social interaction: perceptions of direct and indirect speech acts and interactants who use them. journal of personality and social psychology, 51(2): 305, 1986. zhichao hu, jean e. fox tree, and marilyn a. walker. modeling linguistic and personality adaptation for natural language generation. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 20–31, 2018. bahar irfan, anika narayanan, and james kennedy. dynamic emotional language adaptation in multiparty interactions with agents. in proceedings of the 20th acm international conference on intelligent virtual agents, pages 1–8, 2020. katherine isbister and clifford nass. consistency of personality in interactive characters: verbal cues, non-verbal cues, and user characteristics. international journal of human-computer studies, 53(2):251–267, 2000. w. lewis johnson, paola rizzo, wauter bosma, sander kole, mattijs ghijsen, and herwin van welbergen. generating socially appropriate tutorial dialog. in affective dialogue systems, pages 254–264. springer berlin heidelberg, 2004. melissa k. jungers, caroline palmer, and shari r. speer. time after time: the coordinating influence of tempo in music and speech. cognitive processing, 1(2):21–35, 2002. robert b. kaplan. cultural thought patterns in inter-cultural education. language learning, 16 (1-2):1–20, 1966. theodora koulouri, stanislao lauria, and robert d. macredie. do (and say) as i say: linguistic adaptation in human–computer dialogs. human–computer interaction, 31(1):59–95, 2016. paul r. kroeger. analyzing meaning: an introduction to semantics and pragmatics. language science press, 2019. ivana kruijff-korbayová, ciprian kukina, gerstenberger olga, and jan schehl. generation of output style variation in the sammie dialogue system. in proceedings of the fifth international natural language generation conference, pages 129–137, 2008. willem j. m. levelt and stephanie kelter. surface form and memory in question answering. cognitive psychology, 14(1):78–106, 1982. richard d. lewis. when cultures collide: leading across cultures. brealey, 2010. 37 miehle, minker and ultes jackson liscombe, julia hirschberg, and jennifer j. venditti. detecting certainness in spoken tutorial dialogues. in proceedings of the 9th european conference on speech communication and technology, pages 1837–1840, 2005. michael madaio, justine cassell, and amy ogan. the impact of peer tutors’ use of indirect feedback and instructions. in making a difference: prioritizing equity and access in cscl, 12th international conference on computer supported collaborative learning. philadelphia, pa: international society of the learning sciences, 2017. françois mairesse and marilyn a. walker. towards personality-based user adaptation: psychologically informed stylistic language generation. user modeling and user-adapted interaction, 20 (3):227–278, 2010. françois mairesse and marilyn a. walker. controlling user perceptions of linguistic style: trainable generation of personality traits. computational linguistics, 37(3):455–488, 2011. juliana miehle, koichiro yoshino, louisa pragst, stefan ultes, satoshi nakamura, and wolfgang minker. cultural communication idiosyncrasies in human-computer interaction. in proceedings of the 17th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 74–79, los angeles, usa, september 2016. association for computational linguistics. juliana miehle, wolfgang minker, and stefan ultes. what causes the differences in communication styles? a multicultural study on directness and elaborateness. in proceedings of the eleventh international conference on language resources and evaluation (lrec 2018). european language resources association (elra), 2018a. juliana miehle, wolfgang minker, and stefan ultes. exploring the impact of elaborateness and indirectness on user satisfaction in a spoken dialogue system. in adjunct publication of the 26th conference on user modeling, adaptation and personalization (umap), pages 165–172. acm, 2018b. juliana miehle, isabel feustel, julia hornauer, wolfgang minker, and stefan ultes. estimating user communication styles for spoken dialogue systems. in proceedings of the 12th international conference on language resources and evaluation (lrec 2020), pages 533–541. european language resources association (elra), may 2020. youngme moon and clifford nass. how “real” are computer personalities? psychological responses to personality types in human-computer interaction. communication research, 23(6):651–674, 1996. clifford nass, youngme moon, brian j. fogg, byron reeves, and chris dryer. can computer personalities be human personalities? in conference companion on human factors in computing systems, pages 228–229, 1995. ani nenkova, agustin gravano, and julia hirschberg. high frequency word entrainment in spoken dialogue. in proceedings of the 46th annual meeting of the association for computational linguistics on human language technologies: short papers, pages 169–172. association for computational linguistics, 2008. 38 adapting the elaborateness and indirectness of spoken dialogue systems james w. neuliep. intercultural communication: a contextual approach. sage, 2018. dat quoc nguyen, dai quoc nguyen, dang duc pham, and son bao pham. rdrpostagger: a ripple down rules-based part-of-speech tagger. in proceedings of the demonstrations at the 14th conference of the european chapter of the association for computational linguistics, pages 17–20, 2014. kate g. niederhoffer and james w. pennebaker. linguistic style matching in social interaction. journal of language and social psychology, 21(4):337–360, 2002. shereen oraby, lena reed, shubhangi tandon, t. s. sharath, stephanie lukin, and marilyn a. walker. controlling personality-based stylistic variation with neural natural language generators. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 180–190, 2018. sharon oviatt, courtney darves, and rachel coulston. toward adaptive conversational interfaces: modeling speech convergence with animated personas. acm transactions on computer-human interaction (tochi), 11(3):300–328, 2004. jennifer s. pardo. on phonetic convergence during conversational interaction. the journal of the acoustical society of america, 119(4):2382–2393, 2006. jamie pearson, jiang hu, holly p. branigan, martin j. pickering, and clifford i. nass. adaptive language behavior in hci: how expectations and beliefs about a system affect users’ word choice. in proceedings of the sigchi conference on human factors in computing systems, pages 1177– 1180, 2006. robin pesch, ricarda b. bouncken, and sascha kraus. effects of communication style and age diversity in innovation teams. international journal of innovation and technology management, 12(06):1–20, 2015. martin j. pickering and simon garrod. toward a mechanistic psychology of dialogue. behavioral and brain sciences, 27(2):169–190, 2004. kaśka porayska-pomsta and chris mellish. modelling politeness in natural language generation. in international conference on natural language generation, pages 141–150. springer, 2004. louisa pragst, wolfgang minker, and stefan ultes. exploring the applicability of elaborateness and indirectness in dialogue management. in advanced social interaction with agents : 8th international workshop on spoken dialog systems, pages 189–198. springer international publishing, 2019. anna prokofieva and julia hirschberg. hedging and speaker commitment. in proceedings of the 5th international workshop on emotion, social signals, sentiment & linked open data, reykjavik, iceland, pages 10–13, 2014. david reitter, frank keller, and johanna d. moore. computational modelling of structural priming in dialogue. in proceedings of the human language technology conference of the naacl, companion volume: short papers, pages 121–124. association for computational linguistics, 2006. 39 miehle, minker and ultes michael f. schober. spatial perspective-taking in conversation. cognition, 47(1):1–24, 1993. john r. searle. indirect speech acts. in syntax and semantics 3. speech acts, pages 59–82. academic press, 1975. tuva lunde smestad and frode volden. chatbot personalities matters. in internet science, pages 170–181. springer international publishing, 2019. svetlana stenchikova and amanda stent. measuring adaptation between dialogs. in proceedings of the 8th sigdial workshop on discourse and dialogue, pages 166–173, 2007. stanley s. stevens. on the theory of scales of measurement. science, 103(2684):677–680, 1946. noriko suzuki and yasuhiro katagiri. prosodic alignment in human–computer interaction. connection science, 19(2):131–141, 2007. adriana tapus and maja j. mataric. socially assistive robots: the link between personality, empathy, physiological signals, and task performance. in aaai spring symposium: emotion, personality, and social behavior, pages 133–140, 2008. morgan ulinski, seth benjamin, and julia hirschberg. using hedge detection to improve committed belief tagging. in proceedings of the workshop on computational semantics beyond events and roles, pages 1–5, 2018. willemijn m. van dolen, pratibha a. dabholkar, and ko de ruyter. satisfaction with online commercial group chat: the influence of perceived technology attributes, chat group characteristics, and advisor communication style. journal of retailing, 83(3):339–358, 2007. marilyn a. walker, amanda stent, françois mairesse, and rashmi prasad. individual and domain adaptation in sentence planning for dialogue. journal of artificial intelligence research, 30: 413–456, 2007. ning wang, w. lewis johnson, richard e. mayer, paola rizzo, erin shaw, and heather collins. the politeness effect: pedagogical agents and learning gains. in aied, pages 686–693, 2005. stephen whittaker, marilyn a. walker, and preetam maloor. should i tell all?: an experiment on conciseness in spoken dialogue. in eighth european conference on speech communication and technology, pages 1685–1688, 2003. jenny wilkie, mervyn a. jack, and peter j. littlewood. system-initiated digressive proposals in automated human–computer telephone dialogues: the use of contrasting politeness strategies. international journal of human-computer studies, 62(1):41–71, 2005. 40 dialogue & discourse 9(1) (2018) 1–49 doi: 10.5087/dad.2018.101 a survey of available corpora for building data-driven dialogue systems: the journal version iulian vlad serban {iulian.vlad.serban} at umontreal dot ca mila, diro, université de montréal 2920 chemin de la tour, montréal, qc h3c 3j7, canada ryan lowe {ryan.lowe} at mail dot mcgill dot ca school of computer science, mcgill university 3480 university st, montréal, qc h3a 0e9, canada peter henderson {peter.henderson} at mail dot mcgill dot ca school of computer science, mcgill university 3480 university st, montréal, qc h3a 0e9, canada laurent charlin∗ {laurent.charlin} at hec dot ca department of decision sciences, hec montréal 3000, chemin de la côte-sainte-catherine, montréal, qc h3t 2a7, canada joelle pineau {jpineau} at cs dot mcgill dot ca school of computer science, mcgill university 3480 university st, montréal, qc h3a 0e9, canada editor: david traum submitted 12/2015; accepted 03/2018; published online 05/2018 abstract during the past decade, several areas of speech and language understanding have witnessed substantial breakthroughs from the use of data-driven models. in the area of dialogue systems, the trend is less obvious, and most practical systems are still built through significant engineering and expert knowledge. nevertheless, several recent results suggest that data-driven approaches are feasible and quite promising. to facilitate research in this area, we have carried out a wide survey of publicly available datasets suitable for data-driven learning of dialogue systems. we discuss important characteristics of these datasets, how they can be used to learn various components of a dialogue system, and their other potential uses. we also examine methods for transfer learning between datasets and the use of external knowledge. finally, we discuss appropriate choices of evaluation metrics for the learning objective. ∗. work done primarily while author was at mcgill university c©2018 iulian vlad serban, ryan lowe, peter henderson, laurent charlin, joelle pineau this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). serban, lowe, henderson, charlin and pineau 1. introduction dialogue systems, also known as interactive conversational agents, virtual agents or sometimes chatterbots, are useful in a wide range of applications ranging from technical support services to language learning tools and entertainment (young et al., 2013; shawar and atwell, 2007b). largescale data-driven methods, which use recorded data to automatically infer knowledge and strategies, are becoming increasingly important in speech and language understanding and generation. speech recognition performance has increased tremendously over the last decade due to innovations in deep learning architectures (hinton et al., 2012; goodfellow et al., 2015). similarly, a wide range of datadriven machine learning methods have been shown to be effective for natural language processing, including tasks relevant to dialogue, such as dialogue act classification (reithinger and klesen, 1997; stolcke et al., 2000), dialogue state tracking (thomson and young, 2010; wang and lemon, 2013; ren et al., 2013; henderson et al., 2013; williams et al., 2013; henderson et al., 2014c; kim et al., 2015), natural language generation (langkilde and knight, 1998; oh and rudnicky, 2000; walker et al., 2002; ratnaparkhi, 2002; stent et al., 2004; rieser and lemon, 2010; mairesse et al., 2010; mairesse and young, 2014; wen et al., 2015a; sharma et al., 2016), and dialogue policy learning (young et al., 2013). recent advances in computing power, the availability of large public datasets, and the development of efficient and effective machine learning models has led to increases attention and success in data-driven dialogue systems. importantly, the use of machine learning methods — such as neural networks — require an understanding of the availability, requirements, and uses of available dialogue corpora. to this end, this paper presents a broad survey of available dialogue corpora. corpus-based learning is not the only approach to training dialogue systems. researchers have also proposed training dialogue systems online through live interaction with humans, and offline using user simulator models and reinforcement learning methods (levin et al., 1997; georgila et al., 2006; paek, 2006; schatzmann et al., 2007; jung et al., 2009; schatzmann and young, 2009; gašić et al., 2010, 2011; daubigney et al., 2012; gašić et al., 2012; su et al., 2013; gasic et al., 2013; pietquin and hastie, 2013; young et al., 2013; mohan and laird, 2014; su et al., 2015; piot et al., 2015; cuayáhuitl et al., 2015; hiraoka et al., 2016; fatemi et al., 2016; asri et al., 2016; williams and zweig, 2016; su et al., 2016). however, these approaches are beyond the scope of this survey unless the simulators are built from dialogue corpora. this survey is structured as follows. in the next section, we give a high-level overview of dialogue systems. we briefly discuss the purpose and goal of dialogue systems. then we describe the individual system components that are relevant for data-driven approaches as well as holistic end-to-end dialogue systems. in section 3, we discuss types of dialogue interactions and aspects relevant to building data-driven dialogue systems, from a corpus perspective, as well as the modalities recorded in each corpus (e.g. text, speech and video). we further discuss corpora constructed from both human-human and human-machine interactions, corpora constructed using natural versus unnatural or constrained settings, and corpora constructed using works of fiction. in section 4, we present our survey over dialogue corpora using the categories from sections 2 & 3. in particular, we categorize the corpora based on whether dialogues are: between humans or between a human and a machine; whether the dialogues are in written or spoken language; whether they are constrained or spontaneous; and whether they are scripted fictional works. we discuss each corpus in turn while emphasizing how the dialogues were generated and collected, the topic of the dialogues, and the size of the entire corpus. in section 5, we discuss issues related to: corpus size, transfer learning 2 a survey of available corpora for building data-driven dialogue systems between corpora, incorporation of external knowledge into the dialogue system, data-driven learning for contextualization and personalization, and automatic evaluation metrics. we conclude the survey in section 6. 2. characteristics of data-driven dialogue systems this section offers a broad characterization of data-driven dialogue systems, which structures our presentation of the datasets. 2.1 an overview of dialogue systems the standard architecture for dialogue systems, shown in figure 1, incorporates automatic speech recognition, natural language understanding, dialogue state tracking, dialogue response action selection, natural language generation, and speech synthesis. in the case of text-based (written) dialogues, automatic speech recognition and speech synthesis can be left out and throughout this work we generally do not discuss them and focus on the core components. we can refer to systems including speech synthesis and automatic speech recognition as spoken dialogue systems. in some dialogue systems literature, the dialogue state tracking and dialogue response action selection components comprise the dialogue manager (young, 2000). here, we discuss each component separately and briefly discuss related work to each. we also discuss end-to-end dialogue systems, a relatively new paradigm which ignores the division into components and treats all four non-spoken dialogue system components as a single learned system (ritter et al., 2011; vinyals and le, 2015; lowe et al., 2015a; sordoni et al., 2015; shang et al., 2015; li et al., 2015; serban et al., 2016; serban et al., 2017d,a; dodge et al., 2015; williams and zweig, 2016; weston, 2016). in this work, we focus on corpus-based data-driven dialogue systems. that is, systems that use machine learning algorithms to learn the functionality of the previously described components from dialogue corpora constructed from human conversational data. we define human conversational data as dialogue corpora collected through human-human or human-machine interaction. these system components have variables or parameters that are optimized based on statistics observed in dialogue corpora. in particular, we focus on systems where the majority of variables and parameters are optimized. such corpus-based data-driven systems should be contrasted to systems where each component is hand-crafted by engineers — for example, components defined by an a priori fixed set of deterministic rules (e.g. weizenbaum (1966); mcglashan et al. (1992)). these systems should also be contrasted with systems learning online, such as when the free variables and parameters are optimized directly based on interactions with humans (e.g. gašić et al. (2011)). still, it is worth noting that it is possible to combine different types of learning within one system. for example, some parameters may be learned using statistics observed in a corpus, while other parameters may be learned through interactions with humans. while there are substantial opportunities to improve each of the components in figure 1 through (corpus-based) data-driven approaches, within this survey we focus primarily on datasets suitable to jointly enhance the components inside the dialogue system box. it is worth noting that natural language understanding and natural language generation are core problems in natural language processing with applications well beyond dialogue systems. 3 serban, lowe, henderson, charlin and pineau figure 1: spoken dialogue system diagram. rectangular boxes represent the different components of the system while arrows represent how the information flows between the components from automatic speech recognition all the way to speech synthesis. throughout this survey we mostly discuss text-based dialogue systems (denoted in the figure as dialogue systems) which omit the automatic speech recognition and speech synthesis. 2.2 tasks and objectives dialogue systems have been built for a wide range of purposes. a useful distinction can be made between goal-driven dialogue systems, such as technical support services, and non-goal-driven dialogue systems, such as bots aimed for general chatting (wallace, 2009; serban et al., 2017b; a. ram, 2017; i. papaioannou, 2017; serban et al., 2017c). although both types of systems do in fact have objectives, typically the goal-driven dialogue systems have a well-defined measure of performance that is explicitly related to task completion. while both goal-driven and non-goal-driven components may have different levels of abstraction, throughout this text, we generally refer to the surface-form output of a dialogue system as the dialogue system or natural language response (i.e. the natural language generation output) and a concept-level response (e.g. the output of a dialogue manager or response action selection mechanism) as a response action. non-goal-driven dialogue systems. research on non-goal-driven dialogue systems goes back to the mid-60s. it began, perhaps, with weizenbaum’s famous program eliza, a system based only on simple text parsing rules that managed to convincingly mimic a rogerian psychotherapist by persistently rephrasing statements or asking questions (weizenbaum, 1966). this line of research was continued by colby (1981), who used simple text parsing rules to construct the dialogue system parry, which managed to mimic the pathological behaviour of a paranoid patient to the extent that clinicians could not distinguish it from real patients. however, neither of these two systems used data-driven learning approaches. later work, such as the megahal system by hutchens and alder (1998), started to apply data-driven methods (shawar and atwell, 2007b). hutchens and alder (1998) proposed modeling dialogue as a stochastic sequence of discrete symbols (words) using 4thorder markov chains. given a user utterance, their system generated a response by following a two-step procedure: first, a sequence of topic keywords, used to create a seed reply, was extracted from the user’s utterance; second, starting from the seed reply, two separate markov chains generated the words preceding and proceeding the seed keywords. this procedure produced many candidate responses, from which the highest entropy response was returned to the user. under the 4 a survey of available corpora for building data-driven dialogue systems assumption that the coverage of different topics and general fluency is of primary importance, the 4th order markov chains were trained on a mixture of data sources ranging from real and fictive dialogues to arbitrary texts. until very recently, such data-driven dialogue systems were not applied widely in real-world applications (perez-marin and pascual-nieto, 2011; shawar and atwell, 2007b). more recently, in a similar spirit, several neural network architectures trained on large-scale corpora have been developed. these models show promising results for several non-goal-driven dialogue tasks (ritter et al., 2011; vinyals and le, 2015; lowe et al., 2015a; sordoni et al., 2015; shang et al., 2015; li et al., 2015; serban et al., 2016; serban et al., 2017d,a; dodge et al., 2015; williams and zweig, 2016; weston, 2016). however, they require having sufficiently large corpora — in the hundreds of millions or even billions of words — in order to achieve these results. goal-driven dialogue systems. initial work on goal-driven dialogue systems was primarily based on deterministic hand-crafted rules coupled with learned speech recognition models. one example is the sundial project, which was capable of providing timetable information about trains and airplanes, as well as taking airplane reservations (aust et al., 1995; mcglashan et al., 1992; simpson and eraser, 1993). later, machine learning techniques were used to classify the intention (or need) of the user, as well as to bridge the gap between text and speech (e.g. by taking into account uncertainty related to the outputs of the speech recognition model) (gorin et al., 1997). research in this area started to take off during the mid 1990s, when researchers began to formulate dialogue as a sequential decision making problem based on markov decision processes (singh et al., 1999; young et al., 2013; paek, 2006; pieraccini et al., 2009). unlike for non-goal-driven systems, industry played a major role and enabled researchers to have access to (at the time) relatively large dialogue corpora for certain tasks, such as recordings from technical support call centers. although research in the past decade has continued to push the field towards data-driven approaches, commercial systems are highly domain-specific and heavily based on hand-crafted rules and features (young et al., 2013). in particular, many of the tasks and datasets available are constrained to narrow domains. 2.3 learning dialogue system components most of the dialogue system components shown in figure 1 can be learned through so-called discriminative models, which aim to predict labels or annotations relevant to other parts of the dialogue system. many discriminative models, such as the ones we focus on in this section, fall into the machine learning paradigm of supervised learning. when the labels of interest are discrete (most commonly) the models are called classification models. when the labels of interest are continuous, the models are called regression models. one popular approach for tackling the discriminative task is to learn a probabilistic model of the labels conditioned on the available information, p (y |x), where y is the label of interest (e.g. a discrete variable representing the user intent), and x is the available information (e.g. utterances in the conversation). another popular approach is to use maximum margin classifiers such as support vector machines (cristianini and shawe-taylor, 2000) as opposed to probabilistic models. discriminative models have allowed goal-driven dialogue systems to make significant progress (williams et al., 2013). with proper annotations, discriminative models can be evaluated automatically and accurately. furthermore, once trained on a given dataset, these models may be plugged into a fully-deployed dialogue system (e.g. a classification model for user intents may be used as input to dialogue state tracking). although it is beyond the scope of this paper to provide a survey 5 serban, lowe, henderson, charlin and pineau over such system components, we now give a brief example of each component. this will motivate and facilitate the dataset analysis in section 4. 2.3.1 natural language understanding a natural language understanding model, placed in the context of a dialogue system, is typically designed to interpret and classify the semantic meaning or intent of the interlocutor. several works investigate different statistical approaches for learning natural language understanding models. these involve semantic frame parsing wang et al. (2005), intent classification (tur and deng, 2011), slot-filling (mesnil et al., 2013; liu and lane, 2017), and semantic interpretation (miller et al., 1996). discriminative models, as aforementioned, are often used in natural language understanding for user intent classification. this model is trained to predict the intent of a user conditioned on the utterances of that user. in this case, the intent is called the label (or target or output), and the conditioned utterances are called the conditioning variables (or inputs). training this model requires examples of pairs of user utterances and intentions. one way to obtain these example pairs would be to first record written dialogues between humans carrying out a task, and then to have humans annotate each utterance with its intention label. depending on the complexity of the domain, this may require training the human annotators to reach a certain level of agreement between annotators. often, natural language understanding systems are more expansive than this. along with an intent, these systems can also provide so-called slot information (mesnil et al., 2013; liu and lane, 2017). this is added information which narrows the scope of the intent. for example, the system may classify a “request arrival times intent” with a slot of “flight number = az1234”. 2.3.2 dialogue management as previously mentioned, dialogue management encompasses both dialogue state tracking and dialogue response generation. often, the role of a dialogue manager component is to take the current dialogue state (e.g. output from natural language understanding components) and take an action – typically at the concept level – which can be transformed into a natural language response. see churcher et al. (1997); young (2000); lee et al. (2010) for overviews of data-driven dialogue management components using various methods. recent advances in dialogue management have involved using reinforcement learning methods to learn a dialogue management policy (singh et al., 2002; scheffler and young, 2002; pietquin et al., 2011; rieser and lemon, 2011; peng et al., 2017; fazel-zarandi et al., 2017). an example of such a learned dialogue management policy, as seen in fazel-zarandi et al. (2017), takes the intent and slot information outputted by a natural language understanding system, along with the history of this information and confidence score, and outputs one of several actions: confirming what the interlocutor said, eliciting more information from the interlocutor, or selecting a response according to a pre-defined response generation policy. this concept-level action is then converted to an utterance by the natural language generation component. dialogue state tracking. the dialogue state tracking component of a dialogue system might similarly be implemented as a classification model (williams et al., 2013). at any given point in the dialogue, such a model will take as input all the user utterances and user intention labels estimated by a natural language understanding model so far, and outputs a distribution over possible dialogue states. one common way to represent dialogue states are through slot-value pairs. for 6 a survey of available corpora for building data-driven dialogue systems example, a dialogue system providing timetable information for trains might have three different slots: departure city, arrival city, and departure time. each slot may take one of several discrete values (e.g. departure city could take values from a list of city names). the task of dialogue state tracking is then to output a distribution over every possible combination of slot-value pairs. this distribution — or alternatively, the k dialogue states with the highest probability — may then be used by other parts of the dialogue system. the dialogue state tracking model can be trained on examples of dialogue utterances and dialogue states labeled by humans. dialogue response action selection. given the dialogue state distribution provided by the dialogue state tracking system, the dialogue response action selection component must select an appropriate system response action (sometimes referred to as a dialogue act). this component may also be implemented as a classification model that maps dialogue states to a probability over a discrete set of response actions. for example, in a dialogue system providing timetable information for trains, the set of response actions might include providing information (e.g. providing the departure time of the next train with a specific departure and arrival city) and clarification questions (e.g. asking the user to re-state their departure city). the model may be trained on example pairs of dialogue states and response actions. 2.3.3 natural language generator given a dialogue system selected response action (e.g. a response action providing the departure time of a train), the natural language generator must output the natural language utterance of the system. in the case of commercial goal-driven dialogue systems, this is often implemented using hand-crafted rules. another option is to learn a discriminative model to select a natural language response. in this case, the output space may be defined as a set of so-called surface form sentences (e.g. “the requested train leaves city x at time y”, where x and y are placeholder values). given the system response action, the classification model must choose an appropriate surface form. afterwards, the chosen surface form will have the placeholder values substituted with appropriate items (e.g. x will be replaced by the appropriate city name through a database look up). several works examine such data-driven approaches. the authors of (oh and rudnicky, 2000) use statistical probabilities gathered from corpora to generate a conditional language generation process. similarly lemon (2008), reformulates this mapping as a reinforcement learning problem and trains linear policies to generate natural language based on semantic frames. wen et al. (2015b) use recurrent neural networks to learn a mapping from a dialogue act and semantic frame to a natural language response. many other works, similarly propose various methods for conditional natural language generation based on response actions, dialogue acts, and semantic frames. as with other classification models, such models may be trained on example pairs of system response actions and surface forms. 2.3.4 end-to-end dialogue systems not all dialogue systems conform to the architecture shown in figure 1. various works have examined learning “end-to-end” systems which combine various sub-components together. while some works investigate combining various combinations of sub-components together — such as natural language understanding and dialogue management (yang et al., 2017) — we take a broader view where we define end-to-end systems as encompassing all of the non-spoken dialogue compo7 serban, lowe, henderson, charlin and pineau nents (i.e., natural language understanding, dialogue state tracking, dialogue response action selection, and natural language generation). in particular, so-called end-to-end dialogue system architectures based on neural networks have shown promising results on several dialogue tasks (ritter et al., 2011; vinyals and le, 2015; lowe et al., 2015a; sordoni et al., 2015; shang et al., 2015; li et al., 2015; serban et al., 2016; serban et al., 2017d,a; dodge et al., 2015). in their purest form, these models take as input a dialogue in text form and output a natural language response (or a distribution over responses). we call these systems end-to-end dialogue systems because they possess two important properties. first, they do not contain or require learning any sub-components (such as natural language understanding or dialogue state tracking). consequently, there is no need to collect intermediate labels (e.g. user intention or dialogue state labels). second, all model parameters are optimized w.r.t. a single objective function. often the objective function chosen is maximum log-likelihood (e.g. crossentropy) on a fixed corpus of dialogues. although in earlier work these models depended only on the dialogue context, they may be extended to also depend on outputs from other components (e.g. outputs from the speech recognition component), and on external knowledge (e.g. external databases that can be queried by the system). end-to-end dialogue systems can be divided into two categories: those that select deterministically from a fixed set of possible responses, and those that attempt to stochastically generate responses by keeping a posterior distribution over possible utterances. systems in the first category map the dialogue history, state tracking outputs and external knowledge to a response action: fθ : {dialogue history, state tracking outputs, external knowledge, t} → action at, (1) where at is the dialogue system response action at time t, and θ is the set of parameters that defines f . while the goal of end-to-end systems is to output the response directly (for example, taking a word output as an action), in current systems, the response action may also refer to a selection from a pre-defined set of linguistic responses — perhaps corresponding to different dialogue acts. information retrieval and ranking-based systems — systems that search through a database of dialogues and pick responses with the most similar context, such as the model proposed by banchs and li (2012) — belong to this category. in this case, the mapping function fθ projects the dialogue history into a euclidean space (e.g. using tf-idf bag-of-words representations). the response is then found by projecting all potential responses into the same euclidean space, and the response closest to the desirable response region is selected. the neural network proposed by lowe et al. (2015a) also belongs to this category. in this case, the dialogue history is projected into a euclidean space using a recurrent neural network encoding the dialogue word-by-word. similarly, a set of candidate responses are mapped into the same euclidean space using another recurrent neural network encoding the response word-by-word. finally, a relevance score is computed between the dialogue context and each candidate response, and the response with the highest score is returned. hybrid or combined models, such as the model built on both a phrase-based statistical machine translation system and a recurrent neural network proposed by sordoni et al. (2015), also belong to this category. in this case, a response is generated by deterministically creating a fixed number of answers using the machine translation system and then picking the response according to the score given by a neural network. although both of its sub-components are based on probabilistic models, the final model does not construct a probability distribution over all possible responses.1 1. although the model does not require intermediate labels, it consists of sub-components whose parameters are trained with different objective functions. therefore, strictly speaking, this is not an end-to-end model. 8 a survey of available corpora for building data-driven dialogue systems in contrast to a deterministic system, a stochastic system explicitly computes a full posterior probability distribution over possible system response actions at every turn: pθ(action at | dialogue history, state tracking outputs, external knowledge, t). (2) systems based on generative recurrent neural networks typically belong to this category (vinyals and le, 2015; serban et al., 2016). by breaking down eq. (2) into a product of probabilities over words, responses can be generated by sampling word-by-word from their probability distribution. these systems are also able to generate entirely novel responses by sampling word-by-word (though, some such models require modification to elicit diversity in responses (li et al., 2015)). highly probable responses, i.e. the response with the highest probability, can further be generated by using a method known as beam-search (graves, 2012). these systems project each word into a euclidean space (known as a word embedding) (bengio et al., 2003); they also project the dialogue history and external knowledge into a euclidean space (wen et al., 2015a; lowe et al., 2015b). similarly, the system proposed by ritter et al. (2011) belongs to this category. their model uses a statistical machine translation model to map a dialogue history to its response. when trained solely on text, these generative models can be viewed as unsupervised learning models, because they aim to reproduce the training data distribution. in other words, such models learn to assign a probability to every possible conversation, and since they generate responses word-by-word, they must learn to simulate the behaviour of the agents in the training corpus. early reinforcement learning dialogue systems with stochastic policies also belong to this category (e.g. the njfun system of singh et al. (2002)). in contrast to the neural network and statistical machine translation systems, these reinforcement learning systems typically have very small sets of possible hand-crafted system states (e.g. hand-crafted features describing the dialogue state). the action space is also limited to a small set of pre-defined responses. this makes it possible to apply established reinforcement learning algorithms to train them either online or offline, however it also severely limits their application area. as singh et al. (singh et al., 2002, p.5) remark: “we view the design of an appropriate state space as application-dependent, and a task for a skilled system designer.” 3. dialogue interaction types & aspects this section provides a high-level discussion of different types of dialogue interactions and their salient aspects. the categorization of dialogues is useful for understanding the utility of various datasets for particular applications, as well as for grouping these datasets together to demonstrate available corpora in a given area. here, we discuss the important characteristics and distinctions within dialogue corpora: whether a corpus is written, spoken, or multi-modal; whether a corpus features human-human interactions or human-machine interactions; whether the corpus features constrained or unconstrained dialogues; whether the corpus includes topic oriented or goal driven dialogues; whether the dialogues are fictional or scripted; corpus size. 3.1 written, spoken & multi-modal corpora an important distinction between dialogue corpora is whether participants (interlocutors) interact through written language, spoken language, or in a multi-modal setting (e.g. using both speech and visual modalities). written and spoken language differ substantially w.r.t. their linguistic properties. 9 serban, lowe, henderson, charlin and pineau spoken language tends to be less formal, containing lower information content and many more pronouns than written language (carter and mccarthy, 2006; biber and finegan, 2001, 1986). in particular, the differences are magnified when written language is compared to spoken face-toface conversations, which are multi-modal and highly socially situated. as biber and finegan (1986) observed, pronouns, questions, and contradictions, as well as that-clauses and if-clauses, appear with a high frequency in face-to-face conversations. forchini (2012) summarized these differences: “... studies show that face-to-face conversation is interpersonal, situation-dependent has no narrative concern or as biber and finegan (1986) put it, is a highly interactive, situated and immediate text type...” due to these differences between spoken and written language, we emphasize the distinction between dialogue corpora in written and spoken language in the following sections. similarly, dialogues involving visual and other modalities differ from dialogues without these modalities (card et al., 1983; goodwin, 1981). when a visual modality is available —for example, when two human interlocutors converse face-to-face — body language and eye gaze have a significant impact on what is said and how it is said (gibson and pick, 1963; lord and haith, 1974; cooper, 1974; chartrand and bargh, 1999; de kok et al., 2013). aside from the visual modality, dialogue systems may also incorporate other situational modalities, including aspects of virtual environments (rickel and johnson, 1999; traum and rickel, 2002) and user profiles (li et al., 2016). 3.2 human-human vs. human-machine corpora another salient distinction between dialogue datasets resides in the types of interlocutors — notably, whether it involves interactions between two humans, or between a human and a computer.2 the distinction is important because current artificial dialogue systems are significantly constrained. these systems do not produce nearly the same distribution of possible responses as humans do under equivalent circumstances. as stated by williams and young (2007): (human-human conversation) does not contain the same distribution of understanding errors, and human-human turn-taking is much richer than human-machine dialog. as a result, human-machine dialogue exhibits very different traits than human-human dialogue (doran et al., 2001; moore and browning, 1992). the expectation a human interlocutor begins with, and the interface through which they interact, also affect the nature of the conversation (jonsson and dahlback, 1988). for goal-driven settings, williams and young (2007) have argued against building data-driven dialogue systems using human-human dialogues, as it contains a different distribution of understanding errors. this line of reasoning seems particularly applicable to spoken dialogue systems, where speech recognition errors can have a critical impact on performance and therefore must be taken into account when learning the dialogue model. the argument is also relevant to goal-driven dialogue systems, where an effective dialogue model can often be learned using reinforcement learning techniques. williams and young (2007) also argue against learning from corpora generated between humans and existing dialogue systems, as the trained dialogue system would simply learn to approximate the policy of the spoken dialogue system. thus, if one’s goal is to develop a dialogue system that can interact with real users, the most effective strategy may be learning online through interaction with the users. for example, there 2. machine-machine dialogue corpora are not of interest to us, because they typically differ significantly from natural human language. furthermore, user simulation models are outside the scope of this survey. 10 a survey of available corpora for building data-driven dialogue systems exists useful human-machine corpora where the interacting machine uses a stochastic policy that can generate sufficient coverage of the task (e.g. enough good and enough bad dialogue examples) to allow an effective dialogue model to be learned. in this case, the goal is to learn a policy that is eventually better than the original stochastic policy used to generate the corpus through a process known as bootstrapping. furthermore, there may be other reasons to prefer human-machine over human-human corpora, for example if researchers desire to study the behavior of a particular dialogue system. another possible alternative is the case of wizard-of-oz experiments (bohus and rudnicky, 2008; petrik, 2004). in these dialogue collection methods, a human thinks (s)he is speaking to a machine, but a human operator is in fact controlling the dialogue system. this enables the generation of datasets that are closer in nature to the dialogues humans may wish to achieve in a given setting. for such experiments, it may also be beneficial to influence either or both the human wizard and user to collect a diverse dataset. for example, in a negotiation task konovalov et al. (2016) propose to 1) influence the human wizard by asking them to change their negotiation strategy, and 2) have the human user rephrase their utterances for low-frequency intentions. however, wizard-of-oz experiments are typically expensive and time-consuming to carry out. 3.3 spontaneous vs. constrained corpora the way in which a dialogue corpus is generated and collected can have a significant influence on the trained data-driven dialogue system. many dialogue corpora contain dialogues in which the topics of conversation are either casual or not pre-specified in any way. these corpora can be referred to as spontaneous (unconstrained) corpora, as we believe they most closely mimic spontaneous and unplanned spoken interactions between humans. however, in some corpora, conversations focus on a particular topic or intend to solve a specific task. in such situations, the task or topic is pre-specified and participants are discouraged from deviating from the topic. we refer to these as constrained dialogue corpora. spontaneous corpora bear a close resemblance to natural dialogues — that is, they are close to the generally unplanned nature of most spoken interactions between humans. in the latter case — that of constrained dialogues — some experimental conditions in which dialogues were collected can result in unnatural behaviours that do not correlate well to the true typical dialogue patterns of human-human interaction in day-to-day settings. due to ethical considerations and resource constraints, researchers may be forced to inform the human interlocutors that they are being recorded or to setup artificial experiments in which they hire humans and instruct them to carry out a particular task by interacting with a dialogue system. in these cases, there is no guarantee that the interactions in the corpus will reflect natural human interactions, since the hired humans may behave differently from the population. one factor that may cause behavioural differences is the fact that the hired humans may not share the same intentions and motivations as the population (ai et al., 2007; young et al., 2013). the unnaturalness may be further exacerbated by the hiring process, as well as the platform through which they interact. such factors are becoming more prevalent as researchers increasingly rely on crowdsourcing platforms, such as amazon mechanical turk, to collect and evaluate dialogue data (jurcıcek et al., 2011). while these factors concerning constrained corpora are important to consider, there are many use cases where such corpora are beneficial or necessary. certain researchers may prefer laboratorylike conditions to study certain variables of interest in a conversation. further, constrained corpora 11 serban, lowe, henderson, charlin and pineau in the form of debate settings, topical discussions, etc. are useful to study in and of themselves. in sections 4.2 and 4.3, we separate our discussion of dialogue datasets based on whether the corpora are constrained or spontaneous. 3.4 topic-oriented & goal-driven datasets many human-human datasets may be described as containing casual or unrestricted topics, while human-machine datasets often focus on specific, narrow topics. it is useful to keep this distinction between restricted and unrestricted topics in mind, since goal-driven dialogue systems — which often have a well-defined measure of performance related to task completion — are usually developed in the former setting. when the corpus domain is restricted and a completion metric is available, it may be useful to incorporate this explicitly into the learning procedure. in contrast, when building non-goal-driven dialogue systems based on a corpus with unrestricted topics, it may not be possible to explicitly incorporate any topic information or completion metric into the learning procedure. in some cases, the line between restricted and unrestricted topics blurs. for example, in the case of conversations between players of an online game (afantenos et al., 2012), the outcome of the game is determined by how participants play in the game environment, not by their conversation. in this case, some conversations may have a direct impact on a player’s performance in the game. other conversations may be related to the game but irrelevant to the goal (e.g. commentary on past events). finally, some conversations may be completely unrelated to the game. 3.5 scripted corpora or corpora from fiction it is also possible to use artificial dialogue corpora for data-driven learning. this includes corpora based on works of fiction such as novels, movie manuscripts and audio subtitles. however, unlike transcribed human-human conversations, novels, movie manuscripts, and audio subtitles depend upon events outside the current conversation, which are not observed. this makes data-driven learning more difficult because the dialogue system has to account for unknown factors. the same problem is also observed in certain other media, such as microblogging websites (e.g. twitter and weibo), where conversations also may depend on external unobserved events. nevertheless, recent studies have found that spoken language in movies resembles spontaneous human spoken language (forchini, 2009). although movie dialogues are explicitly written to be spoken and contain certain artificial elements, many of the linguistic and paralinguistic features contained within the dialogues are similar to natural spoken language, including dialogue acts such as turn-taking and reciprocity (e.g. returning a greeting when greeted). the artificial differences that exist may even be helpful for data-driven dialogue learning since movie dialogues are more compact, follow a steady rhythm, and contain less garbling and repetition, all while still presenting a clear event or message to the viewer (dose, 2013; forchini, 2009, 2012). unlike dialogues extracted from wizard-of-oz human experiments, movie dialogues span many different topics and occur in many different environments (webb, 2010). they contain different actors with varying intentions and relationships to one another, which could potentially allow a data-driven dialogue system to learn to personalize itself to each user by identifying different interaction patterns (li et al., 2016). 12 a survey of available corpora for building data-driven dialogue systems 3.6 corpus size as in other machine learning applications such as machine translation (al-onaizan et al., 2000; gülçehre et al., 2015) and speech recognition (deng and li, 2013; bengio et al., 2014), the size of the dialogue corpus is important for building an effective data-driven dialogue system (lowe et al., 2015a; serban et al., 2016). there are two primary perspectives on the importance of dataset size for building data-driven dialogue systems. the first perspective comes from the machine learning literature: larger datasets place constraints on the dialogue model trained from that data. datasets with few examples typically require strong structural priors placed on the model, such as using a modular system, whereas large datasets can be used to train end-to-end dialogue systems with less a priori structure. the second comes from a statistical natural language processing perspective: since the statistical complexity of a corpus grows with the linguistic diversity and number of topics, the number of examples required by a machine learning algorithm to model the patterns in it will also grow with the linguistic diversity and number of topics. consider two small datasets with the same number of dialogues in the domain of bus schedule information: in one dataset the conversations between the users and operator is natural, and the operator can improvise and chitchat; in the other dataset, the operator reads from a script to provide the bus information. despite having the same size, the second dataset will have less linguistic diversity and not include chitchat topics. therefore, it will be easier to train a data-driven dialogue system mimicking the behaviour of the operator in the second dataset, however it will also exhibit a highly pedantic style and not be able to chitchat. in addition to this, to have an effective discussion between any two agents, their common knowledge must be represented and understood by both parties. the process of establishing this common knowledge, also known as grounding, is especially critical to repair misunderstandings between humans and dialogue systems (cahn and brennan, 1999). since the number of misunderstandings can grow with the lexical diversity and number of topics (e.g. misunderstanding the paraphrase of an existing word, or misunderstanding a rarely seen keyword), the number of examples required to repair these grow with linguistic diversity and topics. in particular, the effect of linguistic diversity has been observed in practice: vinyals and le (2015) train a simple encoder-decoder neural network on a proprietary dataset of technical support dialogues. although it has a similar size and purpose as the ubuntu dialogue corpus (lowe et al., 2015a), the qualitative examples shown by vinyals and le (2015) are significantly superior to those obtained by more complex models on the ubuntu corpus (serban et al., 2017a). this result may likely be explained in part due to the fact that technical support operators often follow a comprehensive script for solving problems. as such, the script would reduce the linguistic diversity of their responses. furthermore, since the majority of human-human dialogues are multi-modal and highly ambiguous in nature (chartrand and bargh, 1999; de kok et al., 2013), the size of the corpus may compensate for some of the ambiguities and missing modalities. that is, humans express themselves in conversation through intonation, emotional undertones, contextual cues, body language, and other aspects which may not be conveyed fully in dialogue corpora. if the corpus is sufficiently large, then the resolved ambiguities and missing modalities may, for example, be approximated using latent stochastic variables (serban et al., 2017d). thus, we include corpus size as a dimension of analysis. we also discuss the benefits and drawbacks of several popular large-scale datasets in section 5.1. 13 serban, lowe, henderson, charlin and pineau 4. available dialogue datasets there is a vast amount of data available documenting human communication. much of this data could be used — perhaps after some pre-processing — to train a dialogue system. however, covering all such sources of data would be infeasible. thus, we restrict the scope of this survey to datasets that have already been used to study dialogue or build dialogue systems or which could be leveraged in the near future to build more sophisticated data-driven dialogue models (due to properties of dataset size or annotations useful for different data-driven dialogue components). we restrict the selection further to contain only corpora generated from spoken or written english, and to corpora which, to the best of our knowledge, either are publicly available or will be made available in the near future. we first give a brief overview of each of the considered corpora. later, we highlight promising examples and explain how they could be used to further dialogue research.3 the dialogue datasets analyzed are listed in tables 1–5. column features indicate properties of the datasets, including the number of dialogues, average dialogue length, number of words, whether the interactions are between humans or with an automated system, and whether the dialogues are written or spoken. below, we discuss qualitative features of the datasets, while statistics can be found in the aforementioned table. we generally divide the corpora according to the main characteristics mentioned previously: human-machine corpora (including spoken and written systems), human-human spontaneous (unconstrained) spoken corpora, human-human constrained spoken corpora, human-human scripted (mostly fictional) spoken corpora, human-human spontaneous (unconstrained) written corpora, and human-human constrained written corpora. 4.1 human-machine corpora as discussed in subsection 3.2, an important distinction between dialogue datasets is whether they consist of dialogues between two humans or between a human and a machine. thus, we begin by outlining some of the existing human-machine corpora in several categories based on the types of systems the humans interact with: restaurant and travel information, open-domain knowledge retrieval, and other specialized systems. note, we also include human-human corpora here where one human plays the role of the machine in a wizard-of-oz fashion. 4.1.1 restaurant and travel information one common theme in human-machine language datasets is interaction with systems that provide restaurant or travel information. here we briefly describe some human-machine dialogue datasets in this domain. one of the early and most influential sources of such data is the let’s go! dataset (raux et al., 2005), which was collected from people calling a bus scheduling service at off-peak times.4 the dataset provides over 170,000 conversations recorded between an automated bus information system and callers requesting bus schedule information. one of the more recent sources of such data has come from the datasets for structured dialogue prediction released in conjunction with the dialog state tracking challenge (dstc) (williams et al., 2013). as the name implies, these datasets are used to learn a strategy for dialogue state 3. we form a live list of the corpora discussed in this work, along with links to downloads, at: http://breakend. github.io/dialogdatasets. pull requests can be made to the github repository (https://github. com/breakend/dialogdatasets) hosting the website for continuing updates to the list of corpora. 4. see https://dialrc.github.io/letsgodataset/. 14 http://breakend.github.io/dialogdatasets http://breakend.github.io/dialogdatasets https://github.com/breakend/dialogdatasets https://github.com/breakend/dialogdatasets https://dialrc.github.io/letsgodataset/ a survey of available corpora for building data-driven dialogue systems tracking (sometimes called “belief tracking”), which involves estimating the intentions of a user throughout a dialog. state tracking is useful as it can increase the robustness of speech recognition systems, and can provide an implementable framework for real-world dialogue systems. particularly in the context of goal-oriented dialogue systems (such as those providing travel and restaurant information), state tracking may be necessary for creating coherent conversational interfaces. that is, to form a coherent dialogue, previous contexts must be accounted for – either explicitly or in an end-to-end manner. as such, the first three datasets in the dstc — referred to as dstc1, dstc2, and dstc3 respectively — are medium-sized spoken datasets obtained from human-machine interactions with restaurant and travel information systems. all datasets provide labels specifying the current goal and desired action of the system. dstc1 (williams et al., 2013) is an annotated subset of the conversations from the let’s go! dataset (parent and eskenazi, 2010) discussed earlier, involving conversations between callers and an automated bus information system. dstc2 introduces changing user goals in a restaurant booking system, while trying to provide a desired reservation (henderson et al., 2014b). dstc3 introduces a small amount of labeled data in the domain of tourist information. it is intended to be used in conjunction with the dstc2 dataset as a domain adaptation problem (henderson et al., 2014a). the carnegie mellon communicator corpus (bennett and rudnicky, 2002) also contains human-machine interactions with a travel booking system. it is a medium-sized dataset of interactions with a system providing up-to-the-minute flight information, hotel information, and car rentals. conversations with the system were transcribed, along with user’s comments after the interaction. the atis (air travel information system) pilot corpus (hemphill et al., 1990) is one of the first human-machine corpora. it consists of interactions, lasting about 40 minutes each, between human participants and a travel-type booking system, secretly operated by humans. this dataset contains 1041 utterances. in the maluuba frames corpus (el asri et al., 2017), one user plays the role of a conversational agent in a wizard-of-oz fashion, while the other user is tasked with finding available travel or vacation accommodations according to a pre-specified task. the wizard is provided with a knowledge database which records its actions. semantic frames are annotated in addition to actions which the wizard performed on the database to accompany a line of dialogue. in this way, the frames corpus aims to track decision-making processes in traveland hotel-booking through natural dialog. 4.1.2 open-domain knowledge retrieval knowledge retrieval and question & answer (qa) corpora are a broad distinction of corpora that we will not extensively review here. instead, we include only those qa corpora which explicitly record interactions of humans with existing systems. the ritel corpus (rosset and petel, 2006) is a small dataset of 528 dialogues with the wizard-of-oz ritel platform. the project’s purpose was to integrate spoken language dialogue systems with open-domain information retrieval systems, with the end goal of allowing humans to ask general questions and iteratively refine their search. the questions in the corpus mostly revolve around politics and the economy (e.g. “who is currently presiding the senate?”), along with some conversations about artsand science-related topics. other similar open-domain corpora in this area include wikiqa yang et al. (2015) and ms marco nguyen et al. (2016), which compile responses from automated bing searches and human annotators. these do not record dialogues, but rather gather possible responses to queries. we will only mention them briefly as examples of other open-domain corpora in the field. 15 serban, lowe, henderson, charlin and pineau n am e ty pe to pi cs a vg .# to ta l# to ta l# d es cr ip tio n of tu rn s of di al og ue s of w or ds l et ’s g o! (r au x et al ., 20 05 ) sp ok en b us sc he du le s – 17 1, 12 8 – b us ri de in fo rm at io n sy st em d st c 1 (w ill ia m s et al ., 20 13 ) sp ok en b us sc he du le s 13 .5 6 15 ,0 00 3. 7m b us ri de in fo rm at io n sy st em d st c 2 (h en de rs on et al ., 20 14 b) sp ok en r es ta ur an ts 7. 88 3, 00 0 43 2k r es ta ur an tb oo ki ng sy st em d st c 3 (h en de rs on et al ., 20 14 a) sp ok en to ur is ti nf or m at io n 8. 27 2, 26 5 40 3k in fo rm at io n fo rt ou ri st s c m u c om m un ic at or c or pu s sp ok en tr av el 11 .6 7 15 ,4 81 2m * tr av el pl an ni ng an d bo ok in g sy st em (b en ne tt an d r ud ni ck y, 20 02 ) a t is pi lo tc or pu s† sp ok en tr av el 25 .4 41 11 .4 k * tr av el pl an ni ng an d bo ok in g sy st em (h em ph ill et al ., 19 90 ) r ite lc or pu s† sp ok en u nr es tr ic te d/ d iv er se to pi cs 9. 3* 58 2 60 k a n an no ta te d op en -d om ai n qu es tio n (r os se ta nd pe te l, 20 06 ) an sw er in g sp ok en di al og ue sy st em d ia l o g m at he m at ic al sp ok en m at he m at ic s 12 66 8. 7k * h um an s in te ra ct w ith co m pu te rs ys te m pr oo fs (w ol sk a et al ., 20 04 ) to do m at he m at ic al th eo re m pr ov in g m a t c h c or pu s† sp ok en a pp oi nt m en ts ch ed ul in g 14 .0 44 7 69 k * a sy st em fo rs ch ed ul in g ap po in tm en ts . (g eo rg ila et al ., 20 10 ) in cl ud es di al og ue ac ta nn ot at io ns m al uu ba fr am es † c ha t, q a & tr av el & v ac at io n 15 13 69 – fo rg oa ldr iv en di al og ue sy st em s. (e la sr ie ta l., 20 17 ) r ec om m en da tio n b oo ki ng se m an tic fr am es la be le d an d ac tio ns ta ke n on a kn ow le dg eba se an no ta te d. ta bl e 1: h um an -m ac hi ne di al og ue da ta se ts . st ar re d (* ) nu m be rs ar e ap pr ox im at ed ba se d on th e av er ag e nu m be r of w or ds pe r ut te ra nc e. d at as et s m ar ke d w ith († )i nd ic at e w iz ar dof -o z di al og ue s, w he re th e m ac hi ne is se cr et ly op er at ed by a hu m an . 16 a survey of available corpora for building data-driven dialogue systems 4.1.3 other the dialog mathematical proof dataset (wolska et al., 2004) is a wizard-of-oz dataset involving an automated tutoring system that attempts to advise students on proving mathematical theorems. this is done using a hinting algorithm that provides clues when students come up with an incorrect answer. at only 66 dialogues, the dataset is very small, and consists of a conglomeration of text-based interactions with the system, as well as think-aloud audio and video footage recorded by the users as they interacted with the system. the latter was transcribed and annotated with simple speech acts such as “signaling emotions” or “self-addressing”. the match corpus (georgila et al., 2010) is a small corpus of 447 dialogues based on a wizard-of-oz experiment, which contains conversations from 50 young and old adults interacting with spoken dialogue systems. these conversations were annotated semi-automatically with dialogue acts and “information state update” (isu) representations of dialogue context. the corpus also contains information about the users’ cognitive abilities, with the motivation of modeling how the elderly interact with dialogue systems. 4.2 human-human spoken corpora naturally, there is much more data available for conversations between humans than conversations between humans and machines. thus, we break down this category further, into spoken dialogues (this section) and written dialogues (section 4.3). the distinction between spoken and written dialogues is important, since the distribution of utterances changes dramatically according to the nature of the interaction. as discussed in subsection 3.1, spoken dialogues can potentially be less focused as the user speaks in train-of-thought manner. they also tend to use shorter words and phrases. conversely, in written communication, users have the ability to reflect on what they are writing before they send a message and thus are more precise. written dialogues can also contain spelling errors or abbreviations; such artifacts are usually not present in datasets transcribed from spoken dialogues. 4.2.1 spontaneous spoken corpora we first introduce datasets in which the topics of conversation are either casual or not pre-specified in any way. we refer to these corpora as spontaneous, as we believe they most closely mimic spontaneous, unplanned spoken interactions between humans. perhaps one of the most influential spoken corpora is the switchboard dataset (godfrey et al., 1992). this dataset consists of approximately 2,500 dialogues from phone calls, along with wordby-word transcriptions, with about 500 different speakers. a computer-driven robot operator system introduced a topic for discussion between two participants, and recorded the resulting conversation. about 70 casual topics were provided, of which about 50 were frequently used. the corpus was originally designed for training and testing various speech processing algorithms; however, it has since been used for a wide variety tasks, including the modeling of dialogue acts such as ‘statement’, ‘question’, and ‘agreement’ (stolcke et al., 2000). another important dataset is the british national corpus (bnc) (leech, 1992), which contains approximately 10 million words of dialogue. these were collected in a variety of contexts ranging from formal business or government meetings, to radio shows and phone-ins. although most of the conversations are spoken in nature, some of them are also written. bnc covers a large number of sources, and was designed to represent a wide cross-section of british english from the late 17 serban, lowe, henderson, charlin and pineau twentieth century. the corpus also includes part-of-speech (pos) tagging for every word. the vast array of settings and topics covered by this corpus renders it very useful as a general-purpose spoken dialogue dataset. other datasets have been collected for the analysis of spoken english over the telephone. the callhome american english speech corpus (canavan et al., 1997) consists of 120 such conversations totalling about 60 hours, mostly between family members or close friends. similarly, the callfriend american english-non-southern dialect corpus (canavan and zipperlen, 1996) consists of 60 telephone conversations lasting between 5 and 30 minutes each between english speakers in north america without a southern accent. it is annotated with speaker information such as gender, age, and education. the goal of the project was to support the development of language identification technologies, yet, there are no distinguishing features in either of these corpora in terms of the topics of conversation. an attempt to capture exclusively teenage spoken language was made in the bergen corpus of london teenager language (colt) (haslerud and stenström, 1995). conversations were recorded surreptitiously by student ‘recruits’, with a sony walkman and a lapel microphone, in order to obtain a better representation of teenager interactions ‘in-the-wild’. this dataset has been used to identify trends in language evolution in teenagers (stenström et al., 2002). the cambridge and nottingham corpus of discourse in english (cancode) (mccarthy, 1998) is a subset of the cambridge international corpus, containing about 5 million words collected from recordings made throughout the islands of britain and ireland. it was constructed by cambridge university press and the university of nottingham using dialogue data on general topics between 1995 and 2000. it focuses on interpersonal communication in a range of social contexts, varying from hair salons, to post offices, to restaurants. this has been used, for example, to study language awareness in relation to spoken texts and their cultural contexts (carter, 1998). in the dataset, the relationships between speakers (e.g. roommates, strangers) are labeled and the interaction types are provided (e.g. professional, intimate). other works have attempted to record the physical elements of conversations between humans. to this end, a small corpus entitled d64 multimodal conversational corpus (oertel et al., 2013) was collected, incorporating data from 7 video cameras, and the registration of 3-d head, torso, and arm motion using an optitrack system. significant effort was made to make the data collection process as non-intrusive — and thus, natural — as possible. annotations were made to attempt to quantify overall group excitement and pairwise social distance between participants. a similar attempt to incorporate computer vision features was made in the ami meeting corpus (renals et al., 2007), where cameras, a vga data projector capture, whiteboard capture, and digital pen capture, were all used in addition to speech recordings for various meeting scenarios. as with the d64 corpus, the ami meeting corpus is a dataset of multi-participant chats (four-party dialogues) where all members of the party interact with one another. it has been used for analysis of the dynamics of various corporate and academic meeting scenarios, such as addressee detection in a multi-party chat (akker and traum, 2009). in a similar vein, the cardiff conversation database (ccdb) (aubrey et al., 2013) is an audiovisual database containing unscripted natural conversations between pairs of people. the original dataset consisted of 30 five-minute conversations, 7 of which were fully annotated with transcriptions and behavioural annotations such as speaker activity, facial expressions, head motions, and smiles. the content of the conversation is an unconstrained discussion on topics such as movies. while the original dataset featured 2d visual feeds, an updated version with 3d video has also been 18 a survey of available corpora for building data-driven dialogue systems derived, called the 4d cardiff conversation database (4d ccdb) (vandeventer et al., 2015). this version contains 17 one-minute conversations from 4 participants on similarly un-constrained topics. the diachronic corpus of present-day spoken english (dcpse) (aarts and wallis, 2006) is a parsed corpus of spoken english made up of two separate datasets. it contains more than 400,000 words from the ice-gb corpus (collected in the early 1990s) and 400,000 words from the londonlund corpus (collected from the late 1960s to the early 1980s). ice-gb refers to the british component of the international corpus of english (greenbaum and nelson, 1996; greenbaum, 1996) and contains both spoken and written dialogues from english adults who have completed secondary education. the dataset was selected to provide a representative sample of british english. the london-lund corpus (svartvik, 1990) consists exclusively of spoken british conversations, both dialogues and monologues. it contains a selection of face-to-face, telephone, and public discussion dialogues; the latter refers to dialogues that are heard by an audience that does not participate in the dialogue, including interviews and panel discussions that have been broadcast. the orthographic transcriptions of the datasets are normalized and annotated according to the same criteria; ice-gb was used as a gold standard for the parsing of dcpse. the spoken corpus of the survey of english dialects (beare and scott, 1999) consists of 1000 recordings, with about 0.8 million total words, collected from 1948 to 1961 in order to document various existing english dialects. people aged 60 and over were recruited, being most likely to speak the traditional ‘uncontaminated’ dialects of their area and encouraged to talk about their memories, families, work, and their countryside folklore. the child language data exchange system (childes) (macwhinney and snow, 1985) is a database organized for the study of first and second language acquisition. the database contains 10 million english words and approximately the same number of non-english words. it also contains transcripts, with occasional audio and video recordings of data collected from children and adults learning both first and second languages, although the english transcripts are mostly from children. this corpus could be leveraged in order to build automated teaching assistants. the expanded charlotte narrative and conversation collection (cncc), a subset of the first release of the american national corpus (reppen and ide, 2004), contains 95 narratives, conversations and interviews representative of the residents of mecklenburg county, north carolina and its surrounding communities. the purpose of the cncc was to create a corpus of conversation and conversational narration in a ’new south’ city at the beginning of the 21st century, that could be used as a resource for linguistic analysis. it was originally released as one of several collections in the new south voices corpus, which otherwise contained mostly oral histories. information on speaker age and gender in the cncc is included in the header for each transcript. 4.2.2 constrained spoken corpora next, we discuss domains in which conversations only occur about a particular topic, or towards solving a specific task. not only is the topic of the conversation specified beforehand, but participants are discouraged from deviating off-topic. as a result, these corpora are slightly less general than their spontaneous counterparts; however, they may be useful for building goal-oriented dialogue systems. as discussed in subsection 3.3, this may also make the conversations less natural. we can further subdivide this category into the types of topics they cover: path-finding or planning tasks, persuasion tasks or debates, q&a or information retrieval tasks, and miscellaneous topics. 19 serban, lowe, henderson, charlin and pineau n am e to pi cs to ta l# to ta l# to ta l d es cr ip tio n of di al og ue s of w or ds le ng th sw itc hb oa rd (g od fr ey et al ., 19 92 ) c as ua lt op ic s 2, 40 0 3m 30 0h rs * te le ph on e co nv er sa tio ns on pr esp ec ifi ed to pi cs b ri tis h n at io na lc or pu s (b n c ) c as ua lt op ic s 85 4 10 m 1, 00 0h rs * b ri tis h di al og ue s m an y co nt ex ts ,f ro m fo rm al bu si ne ss or go ve rn m en t (l ee ch ,1 99 2) m ee tin gs to ra di o sh ow s an d ph on ein s. c a l l h o m e a m er ic an e ng lis h c as ua lt op ic s 12 0 54 0k * 60 hr s te le ph on e co nv er sa tio ns be tw ee n fa m ily m em be rs or cl os e fr ie nd s. sp ee ch (c an av an et al ., 19 97 ) c a l l fr ie n d a m er ic an e ng lis h c as ua lt op ic s 60 18 0k * 20 hr s te le ph on e co nv er sa tio ns be tw ee n a m er ic an s w ith a so ut he rn ac ce nt . n on -s ou th er n d ia le ct (c an av an an d z ip pe rl en ,1 99 6) t he b er ge n c or pu s of l on do n u nr es tr ic te d 10 0 50 0k 55 hr s sp on ta ne ou s te en ag e ta lk re co rd ed in 19 93 . te en ag e l an gu ag e c on ve rs at io ns w er e re co rd ed se cr et ly . (h as le ru d an d st en st rö m ,1 99 5) t he c am br id ge an d n ot tin gh am c as ua lt op ic s – 5m 55 0h rs * b ri tis h di al og ue s fr om w id e va ri et y of in fo rm al co nt ex ts ,s uc h as c or pu s of d is co ur se in e ng lis h ha ir sa lo ns ,r es ta ur an ts ,e tc . (m cc ar th y, 19 98 ) d 64 m ul tim od al c on ve rs at io n c or pu s u nr es tr ic te d 2 70 k* 8h rs se ve ra lh ou rs of na tu ra li nt er ac tio n be tw ee n a gr ou p of pe op le (o er te le ta l., 20 13 ) a m im ee tin g c or pu s m ee tin gs 17 5 90 0k * 10 0h rs fa ce -t ofa ce m ee tin g re co rd in gs . (r en al s et al ., 20 07 ) c ar di ff c on ve rs at io n d at ab as e u nr es tr ic te d 30 20 k* 15 0m in a ud io -v is ua ld at ab as e w ith un sc ri pt ed na tu ra lc on ve rs at io ns , (c c d b) (a ub re y et al ., 20 13 ) in cl ud in g vi su al an no ta tio ns . 4d c ar di ff c on ve rs at io n d at ab as e u nr es tr ic te d 17 2. 5k * 17 m in a ve rs io n of th e c c d b w ith 3d vi de o (4 d c c d b) (v an de ve nt er et al ., 20 15 ) t he d ia ch ro ni c c or pu s of c as ua lt op ic s 28 0 80 0k 80 hr s* se le ct io n of fa ce -t ofa ce ,t el ep ho ne ,a nd pu bl ic pr es en td ay sp ok en e ng lis h di sc us si on di al og ue fr om b ri ta in . (a ar ts an d w al lis ,2 00 6) t he sp ok en c or pu s of th e c as ua lt op ic s 31 4 80 0k 60 hr s d ia lo gu e of pe op le ag ed 60 or ab ov e ta lk in g ab ou tt he ir m em or ie s, su rv ey of e ng lis h d ia le ct s fa m ili es ,w or k an d th e fo lk lo re of th e co un tr ys id e fr om a ce nt ur y ag o. (b ea re an d sc ot t, 19 99 ) t he c hi ld l an gu ag e d at a u nr es tr ic te d 11 k 10 m 1, 00 0h rs * in te rn at io na ld at ab as e or ga ni ze d fo rt he e xc ha ng e sy st em st ud y of fir st an d se co nd la ng ua ge ac qu is iti on . (m ac w hi nn ey an d sn ow ,1 98 5) t he c ha rl ot te n ar ra tiv e an d c as ua lt op ic s 95 20 k 2h rs * n ar ra tiv es ,c on ve rs at io ns an d in te rv ie w s re pr es en ta tiv e c on ve rs at io n c ol le ct io n (c n c c ) of th e re si de nt s of m ec kl en bu rg c ou nt y, n or th c ar ol in a. (r ep pe n an d id e, 20 04 ) ta bl e 2: h um an -h um an sp on ta ne ou ss po ke n di al og ue da ta se ts .s ta rr ed (* )n um be rs ar e es tim at es ba se d on th e av er ag e ra te of e ng lis h sp ee ch fr om th e n at io na lc en te rf or vo ic e an d sp ee ch (w w w . n c v s . o r g / n c v s / t u t o r i a l s / v o i c e p r o d / t u t o r i a l / q u a l i t y . h t m l ) 20 www.ncvs.org/ncvs/tutorials/voiceprod/tutorial/quality.html a survey of available corpora for building data-driven dialogue systems collaborative path-finding or planning tasks several corpora focus on task planning or pathfinding through the collaboration of two interlocutors. in these corpora typically one person acts as the decision maker and the other acts as the observer. a well-known example of such a dataset is the hcrc map task corpus (anderson et al., 1991), that consists of unscripted, task-oriented dialogues that have been digitally recorded and transcribed. the corpus uses the map task (brown et al., 1984), where participants must collaborate verbally to reproduce a route from one of the participant’s map to the map of another participant. the corpus is fairly small, but it controls for the familiarity between speakers, eye contact between speakers, matching between landmarks on the participants’ maps, opportunities for contrastive stress, and phonological characteristics of landmark names. by adding these controls, the dataset attempts to focus on solely the dialogue and human speech involved in the planning process. the walking around corpus (brennan et al., 2013) consists of 36 dialogues between people communicating over mobile telephone. the dialogues have two parts: first, a ‘stationary partner’ is asked to direct a ‘mobile partner’ to find 18 destinations on a medium-sized university campus. the stationary partner is equipped with a map marked with the target destinations accompanied by photos of the locations, while the mobile partner is given a gps navigation system and a camera to take photos. in the second part, the participants are asked to interact in-person in order to duplicate the photos taken by the mobile partner. the goal of the dataset is to provide a testbed for natural lexical entrainment, and to be used as a resource for pedestrian navigation applications. the trains 93 dialogues corpus (heeman and allen, 1995) consists of recordings of two interlocutors interacting to solve various planning tasks for scheduling train routes and arranging railroad freight. one user acts the role of a planning assistant system and the other user acts as the coordinator. this was not done in a wizard-of-oz fashion, and as such is not considered a human-machine corpus. 34 different interlocutors were asked to complete 20 different tasks such as: “determine the maximum number of boxcars of oranges that you could get to bath by 7 am tomorrow morning. it is now 12 midnight.” the person playing the role of the planning assistant was provided with access to information that is needed to solve the task. also included in the dataset is the information available to both users, the length of dialogue, and the speaker and ‘system’ interlocutor identities. the verbmobil corpus (burger et al., 2000) is a multilingual corpus consisting of english, german, and japanese dialogues collected for the purposes of training and testing the verbmobil project system. the system was a designed for speech-to-speech machine translation tasks. dialogues were recorded in a variety of conditions and settings with room microphones, telephones, or close microphones, and were subsequently transcribed. users were tasked with planning and scheduling an appointment throughout the course of the dialogue. note that while there have been several versions of the verbmobil corpora released, we refer to the entire collection here as described in (burger et al., 2000). dialogue acts were annotated in a subset of the corpus (1,505 mixed dialogues in german, english and japanese). 76,210 acts were annotated with 32 possible categories of dialogue acts alexandersson et al. (2000).5 persuasion and debates another theme recurring among constrained spoken corpora is the appearance of persuasion or debate tasks. these can involve general debates on a topic or tasking a specific interlocutor to try to convince another interlocutor of some opinion or topic. generally, 5. note, this information and further facts about the verbmobil project and corpus can be found here: http: //verbmobil.dfki.de/facts.html 21 http://verbmobil.dfki.de/facts.html http://verbmobil.dfki.de/facts.html serban, lowe, henderson, charlin and pineau these datasets record the outcome of how convinced the audience is of the argument at the end of the dialogue or debate. the green persuasive dataset (douglas-cowie et al., 2007) was recorded in 2007 to provide data for the humaine project, whose goal is to develop interfaces that can register and respond to emotion. in the dataset, a persuader with strong pro-environmental (‘pro-green’) feelings tries to convince persuadees to consider adopting more green lifestyles; these interactions are in the form of dialogues. it contains 8 long dialogues, totalling about 30 minutes each. since the persuadees often either disagree or agree strongly with the persuaders points, this would be good corpus for studying social signs of (dis)-agreement between two people. the mahnob mimicry database (sun et al., 2011) contains 11 hours of recordings, split over 54 sessions between 60 people engaged either in a socio-political discussion or negotiating a tenancy agreement. this dataset consists of a set of fully synchronized audio-visual recordings of natural dyadic (one-on-one) interactions. it is one of several dialogue corpora that provide multimodal data for analyzing human behaviour during conversations. such corpora often consist of auditory, visual, and written transcriptions of the dialogues. here, only audio-visual recordings are provided. the purpose of the dataset was to analyze mimicry (i.e. when one participant mimics the verbal and nonverbal expressions of their counterpart). the authors provide some benchmark video classification models to this effect. the intelligence squared debate dataset (zhang et al., 2016) covers the “intelligence squared” oxford-style debates taking place between 2006 and 2015. the topics of the debates vary across the dataset, but are constrained within the context of each debate. speakers are labeled and the full transcript of the debate is provided. furthermore, the outcome of the debate is provided (how many of the audience members were for the given proposal or against, before and after the debate). qa or information retrieval there are several corpora which feature direct question-and-answering sessions. these may involve general qa, such as in a press conference, or more task-specific lines of questioning to retrieve a specific set of information. the corpus of professional spoken american english (cpsae) (barlow, 2000) was constructed using a selection of transcripts of interactions occurring in professional settings. the corpus contains two million words involving over 400 speakers, recorded between 1994 and 1998. the cpase has two main components. the first is a collection of transcripts (0.9 million words) of white house press conferences, which contains almost exclusively question and answer sessions, with some policy statements by politicians. the second component consists of transcripts (1.1 million words) of faculty meetings and committee meetings related to national tests that involve statements, discussions, and questions. the creation of the corpus was motivated by the desire to understand and model formal uses of the english language. as previously mentioned, the dialog state tracking challenge (dstc) consists of a series of datasets evaluated using a ‘state tracking’ or ‘slot filling’ metric. while the first 3 installments of this challenge had conversations between a human participant and a computer, dstc4 (kim et al., 2015) contains dialogues between humans. in particular, this dataset has 35 conversations with 21 hours of interactions between tourists and tour guides over skype, discussing information on hotels, flights, and car rentals. due to the small size of the dataset, researchers were encouraged to use transfer learning using other dstc datasets to improve state tracking performance. this same training set is used for dstc5 (kim et al., 2016) as well. however, the goal of dstc5 is to study 22 a survey of available corpora for building data-driven dialogue systems multi-lingual speech-act prediction, and therefore it combines the dstc4 dialogues plus a set of equivalent chinese dialogues; evaluation is done on a holdout set of chinese dialogues. miscellaneous lastly, there are several corpora which do not fall into any of the aforementioned categories, involving a range of tasks and situations. the idiap wolf corpus (hung and chittaranjan, 2010) is an audio-visual corpus containing natural conversational data of volunteers who took part in an adversarial role-playing game called ‘werewolf’. four groups of 8 to 12 people were recorded using headset microphones and synchronized video cameras, resulting in over 7 hours of conversational data. the novelty of this dataset is that the roles of other players are unknown to game participants, and some of the roles are deceptive in nature. thus, there is a significant amount of lying that occurs during the game. although specific instances of lying are not annotated, each speaker is labeled with their role in the game. in a dialogue setting, this could be useful for analyzing the differences in language when deception is being used. the semaine corpus (mckeown et al., 2010) consists of 100 ‘emotionally coloured’ conversations. participants held conversations with an operator who adopted various roles designed to evoke emotional reactions. these conversations were recorded with synchronous video and audio devices. importantly, the operators’ responses were stock phrases that were independent of the content of the user’s utterances, and only dependent on the user’s emotional state. this corpus motivates building dialogue systems with affective and emotional intelligence abilities, since the corpus does not exhibit the natural language understanding that normally occurs between human interlocutors. the loqui human-human dialogue corpus (passonneau and sachar, 2014) consists of annotated transcriptions of telephone interactions between patrons and librarians at new york city’s andrew heiskell braille & talking book library in 2006. it stands out as it has annotated discussion topics, question-answer pair links (adjacency pairs), dialogue acts, and frames (discourse units). similarly, the the icsi meeting recorder dialog act (mrda) corpus (shriberg et al., 2004) has annotated dialogue acts, question-answer pair links (adjacency pairs), and dialogue hot spots.6 it consists of transcribed recordings of 75 icsi meetings on several classes of topics including: the icsi meeting recorder project itself, automatic speech recognition, natural language processing and neural theories of language, and discussions with the annotators for the project. 6. for more information on dialogue hot spots and how they relate to dialogue acts, see (wrede and shriberg, 2003). 23 serban, lowe, henderson, charlin and pineau n am e to pi cs to ta l# to ta l# to ta l d es cr ip tio n of di al og ue s of w or ds le ng th h c r c m ap ta sk c or pu s m ap -r ep ro du ci ng 12 8 14 7k 18 hr s d ia lo gu es fr om h l a p ta sk in w hi ch sp ea ke rs m us tc ol la bo ra te ve rb al ly (a nd er so n et al ., 19 91 ) ta sk to re pr od uc e on on e pa rt ic ip an ts m ap a ro ut e pr in te d on th e ot he rs . t he w al ki ng a ro un d c or pu s l oc at io n 36 30 0k * 33 hr s pe op le co lla bo ra tin g ov er te le ph on e to fin d ce rt ai n lo ca tio ns . (b re nn an et al ., 20 13 ) fi nd in g ta sk g re en pe rs ua si ve d at ab as e l if es ty le 8 35 k* 4h rs a pe rs ua de rw ith (g en ui ne ly )s tr on g pr ogr ee n fe el in gs tr ie s to co nv in ce (d ou gl as -c ow ie et al ., 20 07 ) pe rs ua de es to co ns id er ad op tin g m or e gr ee n lif es ty le s. in te lli ge nc e sq ua re d d eb at es d eb at es 10 8 1. 8m 20 0h rs * v ar io us to pi cs in o xf or dst yl e de ba te s, ea ch co ns tr ai ne d (z ha ng et al ., 20 16 ) to on e su bj ec t. a ud ie nc e op in io ns pr ov id ed pr ean d po st -d eb at es . t he c or pu s of pr of es si on al sp ok en po lit ic s, e du ca tio n 20 0 2m 22 0h rs * in te ra ct io ns fr om fa cu lty m ee tin gs an d w hi te h ou se a m er ic an e ng lis h (b ar lo w ,2 00 0) pr es s co nf er en ce s. m a h n o b m im ic ry d at ab as e po lit ic s, g am es 54 10 0k * 11 hr s tw o ex pe ri m en ts :a di sc us si on on a po lit ic al to pi c, an d a (s un et al ., 20 11 ) ro le -p la yi ng ga m e. t he id ia p w ol fc or pu s r ol epl ay in g g am e 15 60 k* 7h rs a re co rd in g of w er ew ol fr ol epl ay in g ga m e w ith an no ta tio ns (h un g an d c hi tta ra nj an ,2 01 0) re la te d to ga m e pr og re ss . se m a in e co rp us e m ot io na l 10 0 45 0k * 50 hr s u se rs w er e re co rd ed w hi le ho ld in g co nv er sa tio ns w ith an op er at or (m ck eo w n et al ., 20 10 ) c on ve rs at io ns w ho ad op ts ro le s de si gn ed to ev ok e em ot io na lr ea ct io ns . d st c 4/ d st c 5 c or po ra to ur is t 35 27 3k 21 hr s to ur is ti nf or m at io n ex ch an ge ov er sk yp e. (k im et al ., 20 15 ,2 01 6) l oq ui d ia lo gu e c or pu s l ib ra ry in qu ir ie s 82 21 k 14 0* te le ph on e in te ra ct io ns be tw ee n lib ra ri an s an d pa tr on s. (p as so nn ea u an d sa ch ar ,2 01 4) a nn ot at ed di al og ue ac ts ,d is cu ss io n to pi cs ,f ra m es (d is co ur se un its ), qu es tio nan sw er pa ir s. m r d a c or pu s ic si m ee tin gs 75 11 k * 72 hr s r ec or di ng s of ic si m ee tin gs .t op ic s in cl ud e: th e co rp us pr oj ec ti ts el f, (s hr ib er g et al ., 20 04 ) au to m at ic sp ee ch re co gn iti on ,n at ur al la ng ua ge pr oc es si ng an d th eo ri es of la ng ua ge .d ia lo gu e ac ts ,q ue st io nan sw er pa ir s, an d ho ts po ts . t r a in s 93 d ia lo gu es c or pu s r ai lr oa d fr ei gh t 98 55 k 6. 5h rs c ol la bo ra tiv e pl an ni ng of ra ilr oa d fr ei gh tr ou te s. (h ee m an an d a lle n, 19 95 ) r ou te pl an ni ng v er bm ob il c or pu s a pp oi nt m en t 72 6 27 0k 38 h rs sp on ta ne ou s sp ee ch da ta co lle ct ed fo rt he v er bm ob il pr oj ec t. (b ur ge re ta l., 20 00 ) sc he du lin g fu ll co rp us is in e ng lis h, g er m an ,a nd ja pa ne se . w e on ly sh ow e ng lis h st at is tic s. ta bl e 3: h um an -h um an co ns tr ai ne d sp ok en di al og ue da ta se ts .s ta rr ed (* )n um be rs ar e es tim at es ba se d on th e av er ag e ra te of e ng lis h sp ee ch fr om th e n at io na lc en te rf or vo ic e an d sp ee ch (w w w . n c v s . o r g / n c v s / t u t o r i a l s / v o i c e p r o d / t u t o r i a l / q u a l i t y . h t m l ). 24 www.ncvs.org/ncvs/tutorials/voiceprod/tutorial/quality.html a survey of available corpora for building data-driven dialogue systems 4.2.3 scripted corpora a final category of spoken dialogue consists of conversations that have been pre-scripted for the purpose of being spoken later. we refer to datasets containing such conversations as ‘scripted corpora’. as discussed in subsection 3.5, these datasets are distinct from spontaneous human-human conversations, as they inevitably contain fewer ‘filler’ words and expressions that are common in spoken dialogue. however, they should not be confused with human-human written dialogues, as they are intended to sound like natural spoken conversations when read aloud by the participants. furthermore, most of the works here are fictional — a distinction made in subsection 3.5. as such, these scripted dialogues are required to be dramatic, as they are generally sourced from movies or tv shows. there exist multiple scripted corpora based on movies and tv series. these can be sub-divided into two categories: corpora that provide the actual scripts (i.e. the movie script or tv series script) where each utterance is tagged with the appropriate speaker, and those that only contain subtitles and consecutive utterances are not divided or labeled in any way. it is always preferable to have the speaker labels, but there is significantly more unlabeled subtitle data available, and both sources of information can be leveraged to build a dialogue system. the movie dic corpus (banchs, 2012) is an example of the former case—it contains about 130,000 dialogues and 6 million words from movie scripts extracted from the internet movie script data collection7, carefully selected to cover a wide range of genres. these dialogues also come with context descriptions, as written in the script. one derivation based on this corpus is the movie triples dataset (serban et al., 2016). there is also the american film scripts corpus and film scripts online corpus which form the film scripts online series corpus, which can be purchased.8 the latter consists of a mix of british and american film scripts, while the former consists of solely american films. the majority of these datasets consist of raw scripts, which are not guaranteed to portray conversations between only two people. the dataset collected by nio et al. (2014), which we refer to as the filtered movie script corpus, takes over 1 million utterance-response pairs from web-based script resources and filters them down to 86,000 such pairs. the filtering method limits the extracted utterances to x-y-x triples, where x is spoken by one actor and y by another, and each of the utterances share some semantic similarity. these triples are then decomposed into x-y and y-x pairs. such filtering largely removes conversations with more than two speakers, which could be useful in some applications. particularly, the filtering method helps to retain semantic context in the dialogue and keeps a back-and-forth conversational flow that is desired in training many dialogue systems. the cornell movie-dialogue corpus (danescu-niculescu-mizil and lee, 2011) also has short conversations extracted from movie scripts. the distinguishing feature of this dataset is the amount of metadata available for each conversation: this includes movie metadata such as genre, release year, and imdb rating, as well as character metadata such as gender and position on movie credits. although this corpus contains 220,000 dialogue excerpts, it only contains 300,000 utterances; thus, many of the excerpts contain a single utterance. the corpus of american soap operas (davies, 2012b) contains 100 million words in more than 22,000 transcripts of ten american tv-series soap operas from 2001 and 2012. because it is based on soap operas it is qualitatively different from the movie dic corpus, which contains movies 7. http://www.imsdb.com 8. http://alexanderstreet.com/products/film-scripts-online-series 25 http://www.imsdb.com http://alexanderstreet.com/products/film-scripts-online-series serban, lowe, henderson, charlin and pineau in the action and horror genres. the corpus was collected to provide insights into colloquial american speech, as the vocabulary usage is quite different from the british national corpus (davies, 2012a). unfortunately, this corpus does not come with speaker labels. another corpus consisting of dialogues from tv shows is the tvd corpus (roy et al., 2014). this dataset consists of 191 movie transcripts from the comedy show the big bang theory, and the drama show game of thrones, along with crowd-sourced text descriptions (brief episode summaries, longer episode outlines) and various types of metadata (speakers, shots, scenes). text alignment algorithms are used to link descriptions and metadata to the appropriate sections of each script. for example, one might align an event description with all the utterances associated with that event in order to develop algorithms for locating specific events from raw dialogue, such as ’person x tries to convince person y’. some work has been done in order to analyze character style from movie scripts. this is aided by a dataset collected by walker et al. (2012a) that we refer to as the character style from film corpus. this corpus was collected from the imsdb archive, and is annotated for linguistic structures and character archetypes. features, such as the sentiment behind the utterances, are automatically extracted and used to derive models of the characters in order to generate new utterances similar in style to those spoken by the character. thus, this dataset could be useful for building dialogue personalization models. there are two primary movie subtitle datasets: the opensubtitles (tiedemann, 2012) and the subtle corpus (ameixa and coheur, 2013). both corpora are based on the opensubtitles website.9 the opensubtitles dataset is a giant collection of movie subtitles, containing over 1 billion words, whereas subtle corpus has been pre-processed in order to extract interaction-response pairs that can help dialogue systems deal with out-of-domain (ood) interactions. the corpus of english dialogues 1560-1760 (ced) (kytö and walker, 2006) compiles dialogues from the mid-16th century until the mid-18th century. the sources vary from real trial transcripts to fiction dialogues. due to the scripted nature of fictional dialogues and the fact that the majority of the corpus consists of fictional dialogue, we classify it here as such. the corpus is composed as follows: trial proceedings (285,660 words), witness depositions (172,940 words), drama comedy works (238,590 words), didactic works (236,640 words), prose fiction (223,890 words), and miscellaneous (25,970 words). 9. http://www.opensubtitles.org 26 http://www.opensubtitles.org a survey of available corpora for building data-driven dialogue systems n am e to pi cs to ta l# to ta l# to ta l# to ta l# d es cr ip tio n of ut te ra nc es of di al og ue s of w or ks of w or ds m ov ie -d ic m ov ie 76 4k 13 2k 75 3 6m m ov ie sc ri pt s of a m er ic an fil m s. (b an ch s, 20 12 ) di al og ue s m ov ie -t ri pl es m ov ie 73 6k 24 5k 61 4 13 m tr ip le s of ut te ra nc es w hi ch ar e fil te re d to co m e (s er ba n et al ., 20 16 ) di al og ue s fr om x -y -x tr ip le s. fi lm sc ri pt s o nl in e se ri es m ov ie 1m * 26 3k † 1, 50 0 16 m * tw o su bs et s of sc ri pt s sc ri pt s (1 00 0 a m er ic an fil m s an d 50 0 m ix ed ). (b ri tis h/ a m er ic an fil m s) . c or ne ll m ov ie -d ia lo gu e c or pu s m ov ie 30 5k 22 0k 61 7 9m * sh or tc on ve rs at io ns fr om fil m sc ri pt s, an no ta te d (d an es cu -n ic ul es cu -m iz il an d l ee ,2 01 1) di al og ue s w ith ch ar ac te rm et ad at a. fi lte re d m ov ie sc ri pt c or pu s m ov ie 17 3k 87 k 1, 78 6 2m * tr ip le s of ut te ra nc es w hi ch ar e fil te re d to co m e (n io et al ., 20 14 ) di al og ue s fr om x -y -x tr ip le s. a m er ic an so ap o pe ra t v sh ow 10 m * 1. 2m † 22 ,0 00 10 0m tr an sc ri pt s of a m er ic an so ap op er as . c or pu s (d av ie s, 20 12 b) sc ri pt s t v d c or pu s t v sh ow 60 k* 10 k † 19 1 60 0k * t v sc ri pt s fr om a co m ed y (b ig b an g t he or y) an d (r oy et al ., 20 14 ) sc ri pt s dr am a (g am e of t hr on es )s ho w . c ha ra ct er st yl e fr om fi lm m ov ie 66 4k 15 1k 86 2 9. 6m sc ri pt s fr om im sd b, an no ta te d fo rl in gu is tic c or pu s (w al ke re ta l., 20 12 a) sc ri pt s st ru ct ur es an d ch ar ac te ra rc he ty pe s. su bt le c or pu s m ov ie 6. 7m 3. 35 m 6, 18 4 20 m a lig ne d in te ra ct io nre sp on se pa ir s fr om (a m ei xa an d c oh eu r, 20 13 ) su bt itl es m ov ie su bt itl es . o pe ns ub tit le s m ov ie 14 0m * 36 m † 20 7, 90 7 1b m ov ie su bt itl es w hi ch ar e no ts pe ak er -a lig ne d. (t ie de m an n, 20 12 ) su bt itl es c e d (1 56 0– 17 60 )c or pu s w ri tte n w or ks – – 17 7 1. 2m v ar io us sc ri pt ed fic tio na lw or ks fr om (1 56 0– 17 60 ) (k yt ö an d w al ke r, 20 06 ) & tr ia lp ro ce ed in gs as w el la s co ur tt ri al pr oc ee di ng s. ta bl e 4: h um an -h um an sc ri pt ed di al og ue da ta se ts .q ua nt iti es de no te d w ith († )i nd ic at e es tim at es ba se d on av er ag e nu m be ro fd ia lo gu es pe r m ov ie (b an ch s, 20 12 )a nd th e nu m be ro fs cr ip ts or w or ks in th e co rp us .d ia lo gu es m ay no tb e ex pl ic itl y se pa ra te d in th es e da ta se ts . t v sh ow da ta se ts w er e ad ju st ed ba se d on th e ra tio of av er ag e fil m ru nt im e (1 12 m in ut es )t o av er ag e t v sh ow ru nt im e (3 6 m in ut es ). t hi s da ta w as sc ra pe d fr om th e im b d da ta ba se (h t t p : / / w w w . i m d b . c o m / i n t e r f a c e s ). (s ta rr ed (* )q ua nt iti es ar e es tim at ed ba se d on th e av er ag e nu m be r of w or ds an d ut te ra nc es pe r fil m ,a nd th e av er ag e le ng th s of fil m s an d t v sh ow s. e st im at es de riv ed fr om th e ta m er i g ui de fo rw ri te rs (h t t p : / / w w w . t a m e r i . c o m / f o r m a t / w o r d c o u n t s . h t m l ). 27 http://www.imdb.com/interfaces http://www.tameri.com/format/wordcounts.html serban, lowe, henderson, charlin and pineau 4.3 human-human written corpora we proceed to survey corpora of conversations between humans in written form. as before, we sub-divide this section into spontaneous and constrained corpora, depending on whether there are restrictions on the topic of conversation. however, we make a further distinction between forum, micro-blogging, and chat corpora. forum corpora consist of conversations on forum-based websites such as reddit10 where users can make posts, and other users can make comments or replies to said post. in some cases, comments can be nested indefinitely, as users make replies to previous replies. utterances in forum corpora tend to be longer, and there is no restriction on the number of participants in a discussion. on the other hand, conversations on micro-blogging websites such as twitter11 tend to have very short utterances as there is an upper bound on the number of characters permitted in each message. as a result, these tend to exhibit highly colloquial language with many abbreviations. the identifying feature of chat corpora is that the conversations take place in real-time between users. thus, these conversations share more similarities with spoken dialogue between humans, such as common grounding phenomena. 4.3.1 spontaneous written corpora we begin with written corpora where the topic of conversation is not pre-specified. such is the case for the nps internet chatroom conversations corpus (forsyth and martell, 2007), which consists of 10,567 english utterances gathered from age-specific chat rooms of various online chat services from october and november of 2006. each utterance is annotated with part-of-speech and dialogue act information; the correctness of these labels was verified manually. the nps internet chatroom conversations corpus was one of the first corpora of computer-mediated communication (cmc), and it was intended for various nlp applications such as conversation thread topic detection, author profiling, entity identification, and social network analysis. several corpora of spontaneous micro-blogging conversations have been collected, such as the twitter corpus from ritter et al. (2010), which contains 1.3 million post-reply pairs extracted from twitter. the corpus was originally constructed to aid in the production of unsupervised approaches to modeling dialogue acts. larger twitter corpora have been collected. the twitter triples corpus (sordoni et al., 2015) is one such example, with a described original dataset of 127 million context-message-response triples, but only a small labeled subset of this corpus has been released. specifically, the released labeled subset contains 4,232 pairs that scored an average of greater than 4 on the likert-type scale by crowdsourced evaluators for quality of the response to the contextmessage pair. similarly, a large micro-blogging dataset, the sina weibo corpus (shang et al., 2015), which contains 4.5 million post-reply pairs, has been collected and used in literature, but this resource has not yet been made publicly available. we do not include the sina weibo corpus (and its derivatives) in the tables in this section, as they are not primarily in english. the usenet corpus (shaoul and westbury, 2009) is a gigantic collection of public usenet postings12 containing over 7 billion words from october 2005 to january 2011. usenet was a distributed discussion system established in 1980 where participants could post articles to one of 47,860 ‘newsgroup’ categories. it is seen as the precursor to many current internet forums. the 10. http://www.reddit.com 11. http://www.twitter.com 12. http://www.usenet.net 28 http://www.reddit.com http://www.twitter.com http://www.usenet.net a survey of available corpora for building data-driven dialogue systems corpus derived from these posts has been used for research in collaborative filtering (konstan et al., 1997) and role detection (fisher et al., 2006). the nus sms corpus (chen and kan, 2013) consists of conversations carried out over mobile phone sms messages between two users. while the original purpose of the dataset was to improve predictive text entry when mobile phones still mapped multiple letters to a single number, aided by video and timing analysis of users entering their messages it could equally be used for analysis of informal dialogue. it is worth noting that the corpus does not consist of dialogues, but rather single sms messages. sms messages are similar in style to twitter, in that they use many abbreviations and acronyms. the dailydialog dataset (li et al., 2017) consists of conversations crawled from websites which teach english through dialogue. these dialogues consist of everyday conversations such as a customer looking for a product or conversing with a salesperson. since the collected data comes from educational sources, the dialogues are well-defined and generally free of grammatical mistakes or abbreviations. furthermore, the data is hand-labeled with emotions. this added labeling may provide a useful complementary signal in training dialogue systems — for example as a latent variable for eliciting different types of emotion from a dialogue agent. currently, one of the most popular forum-based websites is reddit where users can create discussions and post comments in various sub-forums called ‘subreddits’. each subreddit addresses its own particular topic. over 1.7 billion of these comments have been collected in the reddit corpus.13 each comment is labeled with the author, score (rating from other users), and position in the comment tree; the position is important as it determines which comment is being replied to. researchers are just starting to investigate dialogue problems using this reddit discussion corpus; its large size makes it a particularly interesting candidate for studying transfer learning. additionally, researchers have used smaller collections of reddit discussions for broad discourse classification (schrading et al., 2015). some more specialized versions of the reddit dataset have been curated. the reddit domestic abuse corpus (schrading et al., 2015) consists of reddit posts and comments taken from either subreddits specific to domestic abuse, or from subreddits representing casual conversations, advice, and general anxiety or anger. the motivation is to build classifiers that can detect occurrences of domestic abuse in other areas, which could provide insights into the prevalence and consequences of these situations. these conversations have been pre-processed with lower-casing, lemmatizing, and removal of stop words, and semantic role labels are provided. 13. see: https://www.reddit.com/r/datasets/comments/3bxlg7/i_have_every_publicly_ available_reddit_comment/. 29 https://www.reddit.com/r/datasets/comments/3bxlg7/i_have_every_publicly_available_reddit_comment/ https://www.reddit.com/r/datasets/comments/3bxlg7/i_have_every_publicly_available_reddit_comment/ serban, lowe, henderson, charlin and pineau n am e ty pe to pi cs a vg .# to ta l# to ta l# d es cr ip tio n of tu rn s of di al og ue s of w or ds n ps c ha tc or pu s c ha t u nr es tr ic te d 70 4 15 † 10 0m po st s fr om ag esp ec ifi c on lin e ch at ro om s. (f or sy th an d m ar te ll, 20 07 ) tw itt er c or pu s m ic ro bl og u nr es tr ic te d 2 1. 3m 12 5m ‡ tw ee ts an d re pl ie s ex tr ac te d fr om tw itt er (r itt er et al ., 20 10 ) tw itt er tr ip le c or pu s m ic ro bl og u nr es tr ic te d 3 4, 23 2 65 k ‡ a -b -a tr ip le s ex tr ac te d fr om tw itt er (s or do ni et al ., 20 15 ) u se n et c or pu s m ic ro bl og u nr es tr ic te d 68 7 47 86 0† 7b u se n et fo ru m po st in gs (s ha ou la nd w es tb ur y, 20 09 ) n u s sm s c or pu s sm s m es sa ge s u nr es tr ic te d 18 3k 58 0, 66 8* 2 sm s m es sa ge s co lle ct ed be tw ee n tw o (c he n an d k an ,2 01 3) us er s, w ith tim in g an al ys is . r ed di t fo ru m u nr es tr ic te d – – – 1. 7b co m m en ts ac ro ss r ed di t. r ed di td om es tic a bu se c or pu s fo ru m a bu se he lp 17 .5 3 21 ,1 33 19 m -1 03 m 4 r ed di tp os ts fr om ei th er do m es tic ab us e (s ch ra di ng et al ., 20 15 ) su br ed di ts ,o rg en er al ch at . se ttl er s of c at an c ha t g am e te rm s 95 21 – c on ve rs at io ns be tw ee n pl ay er s (a fa nt en os et al ., 20 12 ) in th e ga m e ‘s et tle rs of c at an .’ c ar ds c or pu s c ha t g am e te rm s 38 .1 1, 26 6 28 2k c on ve rs at io ns be tw ee n pl ay er s (d ja la li et al ., 20 12 ) pl ay in g ‘c ar ds w or ld .’ a gr ee m en ti n w ik ip ed ia ta lk pa ge s fo ru m u nr es tr ic te d 2 82 2 11 0k l iv ej ou rn al an d w ik ip ed ia d is cu ss io ns fo ru m (a nd re as et al ., 20 12 ) th re ad s. a gr ee m en tt yp e an d le ve la nn ot at ed . a gr ee m en tb y c re at e d eb at er s fo ru m u nr es tr ic te d 2 10 k 1. 4m c re at e d eb at e fo ru m co nv er sa tio ns .a nn ot at ed (r os en th al an d m ck eo w n, 20 15 ) w ith ty pe of ag re em en to rd is ag re em en t. in te rn et a rg um en tc or pu s fo ru m po lit ic s 35 .4 5 11 k 73 m d eb at es ab ou ts pe ci fic (w al ke re ta l., 20 12 b) po lit ic al or m or al po si tio ns . m pc c or pu s c ha t so ci al ta sk s 52 0 14 58 k c on ve rs at io ns ab ou tg en er al , (s ha ik h et al ., 20 10 ) po lit ic al ,a nd in te rv ie w to pi cs . u bu nt u d ia lo gu e c or pu s c ha t u bu nt u o pe ra tin g 7. 71 93 0k 10 0m d ia lo gu es ex tr ac te d fr om (l ow e et al ., 20 15 a) sy st em u bu nt u ch at st re am on ir c . u bu nt u c ha tc or pu s c ha t u bu nt u o pe ra tin g 33 81 .6 10 66 5† 2b *2 c ha ts tr ea m sc ra pe d fr om (u th us an d a ha ,2 01 3) sy st em ir c lo gs (n o di al og ue s ex tr ac te d) . m ov ie d ia lo g d at as et c ha t, q a & m ov ie s 3. 3 3. 1m h 18 5m fo rg oa ldr iv en di al og ue sy st em s. in cl ud es (d od ge et al ., 20 15 ) r ec om m en da tio n m ov ie m et ad at a as kn ow le dg e tr ip le s. d ai ly d ia lo g c ha t d ai ly l if e 7. 9 13 k 1. 5m c on ve rs at io ns ex tr ac te d fr om e ng lis h la ng ua ge (l ie ta l., 20 17 ) ed uc at io na lt ex ts .l ab el ed w ith em ot io ns . ta bl e 5: h um an -h um an w ri tte n di al og ue da ta se ts . st ar re d (* )q ua nt iti es ar e co m pu te d us in g w or d co un ts ba se d on sp ac es . tr ia ng le (4 )i nd ic at es lo w er an d up pe r bo un ds co m pu te d us in g av er ag e w or ds pe r ut te ra nc e es tim at ed on a si m ila r r ed di tc or pu s sc hr ad in g (2 01 5) . sq ua re (2 ) in di ca te s es tim at es ba se d on ly on th e e ng lis h pa rt of th e co rp us . d ia lo gu es in di ca te d by († ) ar e co nt ig uo us bl oc ks of re co rd ed co nv er sa tio n in a m ul tipa rt ic ip an tc ha t. fo r u se n et ,t he av er ag e nu m be ro f tu rn s ar e ca lc ul at ed as th e av er ag e nu m be ro f po st s co lle ct ed pe r ne w sg ro up . (‡ ) in di ca te s an es tim at e ba se d on a tw itt er da ta se to fs im ila rs iz e an d re fe rs to to ke ns as w el la s w or ds . 30 a survey of available corpora for building data-driven dialogue systems 4.3.2 constrained written corpora there are also several written corpora where users are limited in terms of topics of conversation. for example, the settlers of catan corpus (afantenos et al., 2012) contains logs of 40 games of ‘settlers of catan’, with about 80,000 total labeled utterances. the game is played with up to 4 players, and is predicated on trading certain goods between players. the goal of the game is to be the first player to achieve a pre-specified number of points. therefore, the game is adversarial in nature, and can be used to analyze situations of strategic conversation where the agents have diverging motives. another corpus that deals with game playing is the cards corpus (djalali et al., 2012), which consists of 1,266 transcripts of conversations between players playing a game in the ‘cards world’. this world is a simple 2-d environment where players collaborate to collect cards. the goal of the game is to collect six cards of a particular suit (cards in the environment are only visible to a player when they are near the location of that player), or to determine that this goal is impossible in the environment. the catch is that each player can only hold 3 cards, thus players must collaborate in order to achieve the goal. further, each player’s location is hidden to the other player, and there are a fixed number of non-chatting moves. thus, players must use the chat to formulate a plan, rather than exhaustively exploring the environment themselves. the dataset has been further annotated by potts (2012) to collect all locative question-answer pairs (i.e. all questions of the form “where are you?”). the agreement by create debaters corpus (rosenthal and mckeown, 2015), the agreement in wikipedia talk pages corpus (andreas et al., 2012) and the internet argument corpus (abbott et al., 2016) all cover dialogues with annotations measuring levels of agreement or disagreement in responses to posts in various media. the agreement by create debaters corpus and the agreement in wikipedia talk pages corpus both are formatted in the same way. post-reply pairs are annotated with whether they are in agreement or disagreement, as well as the type of agreement they are in if applicable (e.g. paraphrasing). the difference between the two corpora is the source: the former is collected from create debate forums and the latter from a mix of wikipedia discussion pages and livejournal postings. the internet argument corpus (iac) (walker et al., 2012b) is a forum-based corpus with 390,000 posts on 11,000 discussion topics. each topic is controversial in nature, including subjects such as evolution, gay marriage and climate change; users participate by sharing their opinions on one of these topics. posts-reply pairs have been labeled as being either in agreement or disagreement, and sarcasm ratings are given to each post. another source of constrained text-based corpora are chat-room environments. such a set-up forms the basis of the mpc corpus (shaikh et al., 2010), which consists of 14 multi-party dialogue sessions of approximately 90 minutes each. in some cases, discussion topics were constrained to be about certain political stances, or mock committees for choosing job candidates. an interesting feature is that different participants are given different roles — leader, disruptor, and consensus builder — with only a general outline of their goals in the conversation. thus, this dataset could be used to model social phenomena such as agenda control, influence, and leadership in on-line interactions. the largest written corpus with a constrained topic is the recently released ubuntu dialogue corpus (lowe et al., 2015a), which has almost 1 million dialogues of 3 turns or more, and 100 million words. it is related to the former ubuntu chat corpus (uthus and aha, 2013). both 31 serban, lowe, henderson, charlin and pineau corpora were scraped from the ubuntu irc channel logs.14 on this channel, users can log in and ask a question about a problem they are having with ubuntu; these questions are answered by other users. although the chat room allows everyone to chat with each other in a multi-party setting, the ubuntu dialogue corpus uses a series of heuristics to disentangle it into dyadic dialogue. the technical nature and size of this corpus lends itself particularly well to applications in technical support. other corpora have been extracted from irc chat logs. the irc corpus (elsner and charniak, 2008) contains approximately 50 hours of chat, with an estimated 20,000 utterances from the linux channel on irc, complete with the posting times. therefore, this dataset consists of technical conversations similar to the ubuntu corpus, with the occasional social chat. the purpose of this dataset was to investigate approaches for conversation disentanglement; given a multi-party chat room, one attempts to recover the individual conversations of which it is composed. for this purpose, there are approximately 1,500 utterances with annotated ground-truth conversations. more recent efforts have combined traditional conversational corpora with question answering and recommendation datasets in order to facilitate the construction of goal-driven dialogue systems. such is the case for the movie dialog dataset (dodge et al., 2015). there are four tasks that the authors propose as a prerequisite for a working dialogue system: question answering, recommendation, question answering with recommendation, and casual conversation. the movie dialog dataset consists of four sub-datasets used for training models to complete these tasks: a qa dataset from the open movie database (omdb)15 of 116k examples with accompanying movie and actor metadata in the form of knowledge triples; a recommendation dataset from movielens16 with 110k users and 1m questions; a combined recommendation and qa dataset with 1m conversations of 6 turns each; and a discussion dataset from reddit’s movie subreddit. the former is evaluated using recall metrics in a manner similar to lowe et al. (2015a). it should be noted that, other than the reddit dataset, the dialogues in the sub-datasets are simulated qa pairs, where each response corresponds to a list of entities from the knowledge base. 5. discussion we now discuss a number of challenges and general methods related to the development and evaluation of data-driven dialogue systems. we highlight challenges relevant to working with large-scale datasets, colloquial language, spelling mistakes and acronyms, as well as missing and unobservable data. we also discuss methods to improve data-driven dialogue systems beyond a single corpus, such as transfer learning between datasets and the usage of external knowledge. researchers and developers may consider applying these methods once they have settled on using one or several corpora. we also discuss user personalization, applicable in the case where there is rich information available for each user. finally, we discuss different methods for evaluating data-driven dialogue systems, including corpus-based evaluation methods. evaluating a data-driven dialogue system properly is critical for real-world deployments as well as for advancing state-of-the-art research, in which case reproducibility of methods and results is crucial. 14. http://irclogs.ubuntu.com 15. http://en.omdb.org 16. http://movielens.org 32 http://irclogs.ubuntu.com http://en.omdb.org http://movielens.org a survey of available corpora for building data-driven dialogue systems 5.1 challenges of learning from large datasets recently, several of the larger dialogue datasets have been used to train data-driven dialogue systems; the twitter corpus (ritter et al., 2010) and the ubuntu dialogue corpus (lowe et al., 2015a) are two examples. in this section, we discuss the benefits and drawbacks of these datasets based on our experience using them for building data-driven models. unlike the previous section, we now focus on highly relevant aspects and characteristics of these datasets specifically for learning in data-driven dialogue systems based on neural-network architectures. 5.1.1 the twitter corpus the twitter corpus consists of a series of conversations extracted from tweets. while the dataset is large and general-purpose, the micro-blogging nature of the source material leads to several drawbacks for building conversational dialogue agents. however, some of these drawbacks do not apply if the end goal is to build an agent that interacts with users on the twitter platform. the twitter corpus has an enormous amount of typos, slang, and abbreviations. due to the 140-character limit in this dataset, tweets are often very short and compressed. in addition, users frequently use twitter-specific devices such as hashtags. unless one is building a dialogue agent specifically for twitter, it is often not desirable to have a chatbot use hashtags and excessive abbreviations as it is not reflective of how humans converse in other environments. this also results in a significant increase in the word vocabulary required for dialogue systems trained at the word level. as such, it is not surprising that character-level models have shown promising results on twitter (dhingra et al., 2016). twitter conversations often contain various kinds of verbal role-playing and imaginative actions similar to stage directions in theater plays (e.g. instead of writing “goodbye”, a user might write “*waves goodbye and leaves*”). these conversations are very different from the majority of textbased chats. therefore, dialogue models trained on this dataset are often able to provide interesting and accurate responses to contexts involving role-playing and imaginative actions (serban et al., 2017d). another challenge is that twitter conversations often rely on implicit context (e.g. they refer to recent public events outside the conversation). in order to learn effective responses for such conversations, a dialogue agent must infer the news event under discussion by referencing some form of external knowledge base. this would appear to be a particularly difficult task. 5.1.2 the ubuntu dialogue corpus the ubuntu dialogue corpus is one of the largest, publicly available datasets containing technical support dialogues. due to the commercial importance of such systems, the dataset has attracted significant attention.17 thus, the ubuntu dialogue corpus presents opportunities for anyone to train large-scale data-driven technical support dialogue systems. despite this, there are several challenges when training data-driven dialogue models on the ubuntu dialogue corpus due to the nature of the data. first, since the corpus comes from a multiparty irc channel, it needs to be disentangled into separate dialogues. this notion of disentanglement in dialogue corpora — that is, given a multi-party dialogue, each utterance must be attributed 17. most of the largest technical support datasets are based on commercial technical support channels, which are proprietary and never released to the public for privacy reasons. 33 serban, lowe, henderson, charlin and pineau to a conversational thread — has been investigated in several works elsner and charniak (2010, 2008). this disentanglement process is noisy, and errors inevitably arise. as a result, some cohesion can be lost and confusion introduced. the most frequent error is when a missing utterance in the dialogue is not picked up by the extraction procedure (i.e. an utterance from the original multi-party chat was not added to the disentangled dialogue). as a result, for a substantial amount of conversations, it is difficult to follow the topic. in particular, this means that some of the next utterance classification (nuc) examples, where models must select the correct next response from a list of candidates, are either difficult or impossible for models to predict. another problem arises from the lack of annotations and labels. since users try to solve their technical problems, it is perhaps best to build models under a goal-driven dialogue framework, where a dialogue system has to maximize the probability that it will solve the user’s problem at the end of the conversation. however, there are no reward labels available. thus, it is difficult to model the dataset in a goal-driven dialogue framework. future work may alleviate this by constructing automatic methods of determining whether a user in a particular conversation solved their problem. a particular challenge of the ubuntu dialogue corpus is the large number of out-of-vocabulary words, including many technical words related to the ubuntu operating system, such as commands, software packages, websites, etc. since these words occur rarely in the dataset, it is difficult to learn their meaning directly from the dataset — for example, it is difficult to obtain meaningful distributed, real-valued vector representations for neural network-based dialogue models. this is further exacerbated by the large number of users who use different nomenclature, acronyms, and speaking styles, and the many typos in the dataset. thus, the linguistic diversity of the corpus is large. a final challenge of the dataset is the necessity for additional knowledge related to ubuntu in order to accurately generate or predict the next response in a conversation. we hypothesize that this knowledge is crucial for a system trained on the ubuntu dialogue corpus to be effective in practice, as often solutions to technical problems change over time as new versions of the operating system become available. thus, an effective dialogue system must learn to combine up-to-date technical information with an understanding of natural language dialogue in order to solve the users’ problems. we will discuss the use of external knowledge in more detail in section 5.3. while these challenges make it difficult to build data-driven dialogue systems, it also presents an important research opportunity. current data-driven dialogue systems perform rather poorly in terms of generating utterances that are coherent and on-topic (serban et al., 2017a). as such, there is significant room for improvement on these models. 5.2 transfer learning between datasets while it is not always feasible to obtain large corpora for every new application, the use of other related datasets can effectively bootstrap the learning process. in several branches of machine learning, and in particular in deep learning, the use of related datasets for pre-training models is an effective method of scaling up to complex environments (erhan et al., 2010; kumar et al., 2015). to build open-domain dialogue systems, it is arguably necessary to move beyond domainspecific datasets. instead, like humans, dialogue systems may have to be trained on multiple data sources for solving multiple tasks. to leverage statistical efficiency, it may be necessary to first use unsupervised learning — as opposed to supervised learning or offline reinforcement learning, which typically only provide a sparse scalar feedback signal for each phrase or sequence of phrases — and 34 a survey of available corpora for building data-driven dialogue systems then fine-tune models based on human feedback. researchers have already proposed various ways of applying transfer learning to build data-driven dialogue systems, ranging from learning separate sub-components of the dialogue system (e.g. intent and dialogue act classification) to learning the entire dialogue system (e.g. in an unsupervised or reinforcement learning framework) using transfer learning (fabbrizio et al., 2004; forgues et al., 2014; serban and pineau, 2015; serban et al., 2016; lowe et al., 2015a; vandyke et al., 2015; wen et al., 2016; gašić et al., 2016; mo et al., 2016; genevay and laroche, 2016; chen et al., 2016) 5.3 incorporating external knowledge another interesting research direction is the incorporation of external knowledge sources in order to inform the response to be generated. using external information is of great importance to dialogues systems, particularly in the goal-driven setting. even non-goal-driven dialogue systems designed to simply entertain the user could benefit from leveraging external information, such as current news articles or movie reviews, in order to better converse about real-world events. this may be particularly useful in data-sparse domains, when there is insufficient dialogue training data to reliably learn a response that is appropriate for each input utterance, or in domains that evolve quickly over time. 5.3.1 structured external knowledge in traditional goal-driven dialogue systems (levin and pieraccini, 1997), where the goal is to provide information to the user, there is already extensive use of external knowledge sources. for example, in the let’s go! dialogue system (raux et al., 2005), the user requests information about various bus arrival and departure times. thus, a critical input to the model is the actual bus schedule, which is used in order to generate the system’s utterances. another example is the dialogue system described by nöth et al. (2004), which helps users find movie information by utilizing movie show times from different cinemas. such examples are abundant both in the literature and in practice. although these models make use of external knowledge, the knowledge sources in these cases are highly structured and are only used to place hard constraints on the possible states of an utterance to be generated. they are essentially contained in relational databases or structured ontologies, and are only used to provide a deterministic mapping from the dialogue states extracted from an input user utterance to the dialogue system state or the generated response. complementary to domain-specific databases and ontologies are the general natural language processing databases and tools. these include lexical databases such as wordnet (miller, 1995), which contains lexical relationships between words for over a hundred thousand words, verbnet (schuler, 2005) which contains lexical relations between verbs, and framenet (ruppenhofer et al., 2006), which contains ’word senses’ for over ten thousand words along with examples of each word sense. in addition, there exist several natural language processing tools such as part of speech taggers, word category classifiers, word embedding models, named entity recognition models, coreference resolution models, semantic role labeling models, semantic similarity models and sentiment analysis models (manning and schütze, 1999; jurafsky and martin, 2008; mikolov et al., 2013; gurevych and strube, 2004; lin and walker, 2011b) that may be used by the natural language understanding component to extract meaning from human utterances. since these tools are typically built upon texts and annotations created by humans, using them inside a dialogue system can be interpreted as a form of structured transfer learning, where the relationships or labels learned 35 serban, lowe, henderson, charlin and pineau from the original natural language processing corpus provide additional information to the dialogue system and improve generalization of the system. 5.3.2 unstructured external knowledge complementary sources of information can be found in unstructured knowledge sources, such as online encyclopedias (wikipedia (denoyer and gallinari, 2007)) as well as domain-specific sources (lowe et al., 2015b). it is beyond the scope of this paper to review all possible ways that these unstructured knowledge sources have or could be used in conjunction with a data-driven dialogue system. however, we note that this is likely to be a fruitful research area. 5.4 personalized dialogue agents when conversing, humans often adapt to their interlocutor to facilitate understanding, and thus improve conversational efficiency and satisfaction. attaining human-level performance with dialogue agents may well require personalization, i.e. models that are aware and capable of adapting to their interlocutor. such capabilities could increase the effectiveness and naturalness of generated dialogues (lucas et al., 2009; su et al., 2013). we see personalization of dialogue systems as an important task, which so far has not received much attention. there has been initial efforts on userspecific models which could be adapted to work in combination with the dialogue models presented in this survey (lucas et al., 2009; lin and walker, 2011a; pargellis et al., 2004). there has also been interesting work on character modeling in movies (walker et al., 2011; li et al., 2016; mo et al., 2016). there is significant potential to learn user models as part of dialogue models. the large datasets presented in this paper, some of which provide multiple dialogues per user, may enable the development of such models. 5.5 evaluation metrics one of the most challenging aspects of constructing dialogue systems lies in their evaluation. while the end goal is to deploy the dialogue system in an application setting and receive real human feedback, getting to this stage is time consuming and expensive. often it is also necessary to optimize performance on a pseudo-performance metric prior to release. this is particularly true if a dialogue model has many hyper-parameters to be optimized — it is infeasible to run user experiments for every parameter setting in a grid search. although crowdsourcing platforms, such as amazon mechanical turk, can be used for some user testing (jurcıcek et al., 2011), evaluations using paid subjects can also lead to biased results (young et al., 2013). ideally, we would have some automated metrics for calculating a score for each model, and only involve human evaluators once the best model has been chosen with reasonable confidence. in non-goal-driven dialogue systems researchers have focused mainly on the output of the response generation module. evaluation of such non-goal-driven dialogue systems can be traced back to the turing test (turing, 1950), where human judges communicate with both computer programs and other humans over a chat terminal without knowing the other party’s true identity. the judge’s goal is to identify the humans and computer programs under the assumption that a program indistinguishable from a real human being must be intelligent. however, this setup has been criticized extensively with numerous researchers proposing alternative evaluation procedures (cohen, 2005). more recently, researchers have turned to analyzing the collected dialogues produced after they are 36 a survey of available corpora for building data-driven dialogue systems finished (galley et al., 2015; pietquin and hastie, 2013; shawar and atwell, 2007a; schatzmann et al., 2005). even when human evaluators are available, it is often difficult to choose a set of informative and consistent criteria that can be used to judge an utterance generated by a dialogue system. for example, one might ask the evaluator to rate the f on vague notions such as ‘appropriateness’ and ‘naturalness’, or to try to differentiate between utterances generated by the system and those generated by actual humans (vinyals and le, 2015). schatzmann et al. (2005) suggest two aspects that need to be evaluated for all response generation systems (as well as user simulation models): 1) if the model can generate human-like output, and 2) if the model can reproduce the variety of user behaviour found in corpus. but we lack a definitive framework for such evaluations. we complete this discussion by summarizing different approaches to the automatic evaluation problem as they relate to these objectives. 5.5.1 automatic evaluation metrics for goal-driven dialogue systems user evaluation of goal-driven dialogue systems typically focuses on goal-related performance criteria, such as goal completion rate, dialogue length, and user satisfaction (walker et al., 1997; schatzmann et al., 2005). these were originally evaluated by human users interacting with the dialogue system. recently, researchers have also begun to use third-party annotators for evaluating recorded dialogues (yang et al., 2010). due to their simplicity, the vast majority of hand-crafted task-oriented dialogue systems have been solely evaluated in this way. however, when using machine learning algorithms to train on large-scale corpora, optimization criteria are required. the challenge with evaluating goal-driven dialogue systems without human intervention is that the process necessarily requires multiple steps — it is difficult to determine if a task has been solved from a single utterance-response pair from a conversation. thus, synthetic data is often generated by a user simulator (eckert et al., 1997; schatzmann et al., 2007; jung et al., 2009; georgila et al., 2006; pietquin and hastie, 2013). given a sufficiently accurate user simulation model, an interaction between the dialogue system and the user can be simulated from which it is possible to deduce the desired metrics, such as goal completion rate. significant effort has been made to render the simulated data as realistic as possible, by modeling user intentions. evaluation of such simulation methods has already been conducted (schatzmann et al., 2005). however, generating realistic user simulation models remains an open problem. 5.5.2 automatic evaluation metrics for non-goal-driven dialogue systems evaluation of non-goal-driven dialogue systems, whether by automatic means or user studies, remains a difficult challenge. word overlap metrics. one approach is to borrow evaluation metrics from other nlp tasks such as machine translation, which uses bleu (papineni et al., 2002) and meteor (banerjee and lavie, 2005) scores. these metrics have been used to compare responses generated by a learned dialogue strategy to the actual next utterance in the conversation, conditioned on a dialogue context (sordoni et al., 2015). however, bleu scores have been shown not to correlate with human judgment for assessing dialogue response generation (liu et al., 2016). there are several issues to consider: given the context of a conversation, there often exists a large number of possible responses that ‘fit’ into the dialogue. thus, the response generated by a dialogue system could be entirely reasonable, yet it may have no words in common with the actual next utterance. in this case, the bleu 37 serban, lowe, henderson, charlin and pineau score would be very low, but would not accurately reflect the strength of the model. indeed, even humans who are tasked with predicting the next utterance of a conversation achieve relatively low bleu scores (sordoni et al., 2015). although the meteor metric takes into account synonyms and morphological variants of words in the candidate response, it still suffers from the aforementioned problems. in a sense, these measurements only satisfy one direction of schatzmann’s criteria (schatzmann et al., 2005): high bleu and meteor scores imply that the model is generating human-like output, but the model may still not reproduce the variety of user behaviour found in corpus. furthermore, such metrics will only accurately reflect the performance of the dialogue system if given a large number of candidate responses for each given context. next utterance classification. alternatively, one can narrow the number of possible responses to a small, pre-defined list, and ask the model to select the most appropriate response from this list. the list includes the actual next response of the conversation (the desired prediction), and the other entries (false positives) are sampled from elsewhere in the corpus (lowe et al., 2016, 2015a). this next utterance classification (nuc) task is derived from recall and precision metrics typical of information-retrieval-based approaches. there are several attractive properties of this task: it is easy to interpret, and its difficulty can be adjusted by changing the number of false responses. however, there are drawbacks. since the other candidate answers are sampled from elsewhere in the corpus, there is a chance that these also represent reasonable responses given the context. this can be alleviated to some extent by reporting recall@k measures, i.e. whether the correct response is found in the k responses with the highest rankings according to the model. although current models evaluated using nuc are trained explicitly to maximize the performance on a related metric (cross-entropy between context-response pairs (lowe et al., 2015a; kadlec et al., 2015)), precision and recall could also be used to evaluate a probabilistic generative model trained to outputs full utterances. word perplexity. another metric proposed to evaluate probabilistic language models (bengio et al., 2003; mikolov et al., 2010) that has seen significant recent use for evaluating end-to-end dialogue systems is word perplexity (pietquin and hastie, 2013; serban et al., 2016). perplexity explicitly measures the probability that the model will generate the ground truth next utterance given some context of the conversation. this is particularly appealing for dialogue, as the distribution over words in the next utterance can be highly multi-modal (i.e. many possible responses). a re-weighted perplexity metric has also been proposed where stop words, punctuation, and end-of-utterance tokens are ignored to focus on the semantic content of the phrase (serban et al., 2016). both word perplexity, as well as utterance-level recall and precision outlined above, satisfy schatzmann’s evaluation criteria, since scoring high on these would require the model to produce human-like output and to reproduce most types of conversations in the corpus. response diversity. recent non-goal-driven dialogue systems based on neural networks have had problems generating diverse responses (serban et al., 2016). (li et al., 2015) recently introduced two new metrics, distinct-1 and distinct-2, which respectively measure the number of distinct unigrams and bigrams of the generated responses. although these fail to satisfy either of schatzmann’s criteria, they may still be useful in combination with other metrics, such as bleu, nuc or word perplexity. 38 a survey of available corpora for building data-driven dialogue systems 6. conclusion this paper provides an extensive survey of currently available datasets suitable for research, development, and evaluation of data-driven dialogue systems. we categorize these corpora along several dimensions depending on whether the dataset is written or spoken, between human interlocutors or human-machine conversations, and constrained in topic or more free-form. we collect statistics for these datasets and present them in tables in section 4, and provide an open-source github repository where these datasets can be viewed and pull requests can be made to add new datasets: https://github.com/breakend/dialogdatasets. there is broad coverage of existing datasets along most of the dimensions we consider. however, the vast majority of the available datasets contain at most thousands of dialogues. this presents some challenges to the training of large-scale end-to-end models, such as neural networks, on general purpose domains. neural networks can be applied to narrow domains, such as restaurant recommendation, with relatively little data (wen et al., 2017). however, as the nature of interactions becomes more open and the number of topics grows, the sample complexity and with it the required dataset size increases. to obtain reasonable results in such a setting, neural network practitioners have resorted to training neural network models on datasets with hundreds of thousands to millions of dialogues: the twitter corpus (ritter et al., 2010; sordoni et al., 2015), reddit, the ubuntu dialogue corpus (lowe et al., 2015a), and various movie subtitle datasets such as subtle, opensubtitles, movie-dic, and the movie dialogue dataset (ameixa and coheur, 2013; tiedemann, 2012; banchs, 2012; dodge et al., 2015). while the conversation topics in these datasets often vary considerably, the nature of the datasets themselves are fairly fixed in the form of informal written dialogues between humans. this is the case for movie scripts, forum posts, and micro-blogging platforms. learning only from these sources will bias dialogue systems towards certain kinds of interactions and behaviours; for example, written corpora usually have a specific turn-taking structure that is different from spoken conversation, and they may encode biases against certain groups or populations (henderson et al., 2017). if we want dialogue systems to speak in a more natural way, similar to spoken human-human conversation, emphasis should be placed on collecting large-scale spoken dialogue corpora to train the next generation of dialogue systems. there is also a lack of large-scale multi-modal datasets, which may be crucial towards grounding the language learned by our dialogue agents in human-like experience. we outline different approaches for overcoming the dearth of very large dialogue datasets. transfer learning appears to be a particularly promising avenue for future dialogue research. while many individual datasets presented in section 4 are only thousands of dialogues, summed together they represent a significant resource covering a wide range of topics. it would seem highly advantageous if methods are developed that enable dialogue systems to learn across all of these resources. there is also conceivable transfer that could be gained through learning on non-dialogue text corpora, such as wikipedia. we further discuss research directions for building data-driven dialogue systems, including incorporating external knowledge and personalizing dialogue agents. we discuss several challenges associated with training large-scale dialogue models on two popular dialogue corpora: the twitter corpus and the ubuntu dialogue corpus, based on our own experience. finally, we discuss automatic evaluation metrics used to train data-driven dialogue systems on the datasets previously mentioned. having an automatic metric that correlates highly with human judgment of dialogue quality is very important. even with significant dialogue data, poor evaluation metrics mean that it is difficult to compare the quality of dialogue systems trained on these datasets, 39 serban, lowe, henderson, charlin and pineau and progress as a field becomes difficult. at present, there exists no silver bullet for automatic evaluation, implying that a set of diverse metrics be used together to obtain an accurate measure of performance. the final arbiter is, of course, human judgments. however, these can be expensive to obtain on a large scale, in particular for publicly-funded research laboratories. as a dialogue community, we should strive towards releasing more large-scale dialogue datasets and producing standardized evaluation metrics, to make dialogue research as inclusive as possible for all who work in the field. acknowledgements the authors gratefully acknowledge financial support by the samsung advanced institute of technology (sait), the natural sciences and engineering research council of canada (nserc), the canada research chairs, the canadian institute for advanced research (cifar) and compute canada. the second author is funded by a vanier graduate scholarship. early versions of the manuscript benefited greatly from the proofreading of melanie lyman-abramovitch, and later versions were extensively revised by genevieve fried and nicolas angelard-gontier. the authors also thank nissan pow, michael noseworthy, chia-wei liu, gabriel forgues, alessandro sordoni, yoshua bengio and aaron courville for helpful discussions. references c. khatri a. venkatesh r. gabriel q. li j. nunn b. hedayatnia m. heng a. nagar e. king k. bland a. wartick y. pan h. song s. jayadevan g. hwang a. pettigrue a. ram, r. prasad. conversational ai: the science behind the alexa prize. in alexa prize proceedings, 2017. b. aarts and s. a. wallis. the diachronic corpus of present-day spoken english (dcpse), 2006. r. abbott, b. ecker, p. anand, and m. walker. internet argument corpus 2.0: an sql schema for dialogic social media and the corpora to go with it. in language resources and evaluation conference, lrec2016, 2016. s. afantenos, n. asher, f. benamara, a. cadilhac, cédric dégremont, p. denis, m. guhe, s. keizer, a. lascarides, o. lemon, et al. developing a corpus of strategic conversation in the settlers of catan. in seinedial 2012-the 16th workshop on the semantics and pragmatics of dialogue, 2012. h. ai, a. raux, d. bohus, m. eskenazi, and d. j. litman. comparing spoken dialog corpora collected with recruited subjects versus real users. in special interest group on discourse and dialogue (sigdial), 2007. r. akker and d. traum. a comparison of addressee detection methods for multiparty conversations. in workshop on the semantics and pragmatics of dialogue, 2009. y. al-onaizan, u. germann, u. hermjakob, k. knight, p. koehn, d. m., and k. yamada. translating with scarce resources. in aaai, 2000. j. alexandersson, r. engel, m. kipp, s. koch, u. küssner, n. reithinger, and m. stede. modeling negotiation dialogs. in verbmobil: foundations of speech-to-speech translation, pages 441–451. springer, 2000. d. ameixa and l. coheur. from subtitles to human interactions: introducing the subtle corpus. technical report, tech. rep., 2013. a. h. anderson, m. bader, e. g. bard, e. boyle, g. doherty, s. garrod, s. isard, j. kowtko, j. mcallister, j. miller, et al. the hcrc map task corpus. language and speech, 34(4):351–366, 1991. j. andreas, s. rosenthal, and k. mckeown. annotating agreement and disagreement in threaded discussion. in lrec, pages 818–822, 2012. l. e. asri, j. he, and k. suleman. a sequence-to-sequence model for user simulation in spoken dialogue systems. arxiv preprint arxiv:1607.00070, 2016. a. j. aubrey, d. marshall, p. l. rosin, j. vandeventer, d. w. cunningham, and c. wallraven. cardiff conversation database (ccdb): a database of natural dyadic conversations. in computer vision and pattern recognition workshops (cvprw), ieee conference on, pages 277–282, 2013. h. aust, m. oerder, f. seide, and v. steinbiss. the philips automatic train timetable information system. speech communication, 17(3):249–262, 1995. 40 a survey of available corpora for building data-driven dialogue systems r. e. banchs. movie-dic: a movie dialogue corpus for research and development. in association for computational linguistics: short papers, 2012. r. e. banchs and h. li. iris: a chat-oriented dialogue system based on the vector space model. in association for computational linguistics, system demonstrations, 2012. s. banerjee and a. lavie. meteor: an automatic metric for mt evaluation with improved correlation with human judgments. in association for computational linguistics, workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 2005. m. barlow. corpus of spoken, professional american-english, 2000. j. beare and b. scott. the spoken corpus of the survey of english dialects: language variation and oral history. in proceedings of allc/ach, 1999. y. bengio, r. ducharme, p. vincent, and c. janvin. a neural probabilistic language model. the journal of machine learning research, 3:1137–1155, 2003. y. bengio, i. goodfellow, and a. courville. deep learning. an mit press book in preparation. draft chapters available at http://www. iro. umontreal. ca/ bengioy/dlbook, 2014. c. bennett and a. i rudnicky. the carnegie mellon communicator corpus, 2002. d. biber and e. finegan. an initial typology of english text types. corpus linguistics ii: new studies in the analysis and exploitation of computer corpora, pages 19–46, 1986. d. biber and e. finegan. diachronic relations among speech-based and written registers in english. variation in english: multi-dimensional studies, pages 66–83, 2001. d. bohus and a. i rudnicky. sorry, i didnt catch that! in recent trends in discourse and dialogue, pages 123–154. springer, 2008. s. e. brennan, k. s. schuhmann, and k. m. batres. entrainment on the move and in the lab: the walking around corpus. in conference of the cognitive science society, 2013. g. brown, a. anderson, r. shillcock, and g. yule. teaching talk. cambridge: cup, 1984. s. burger, k. weilhammer, f. schiel, and h. g. tillmann. verbmobil data collection and annotation. in verbmobil: foundations of speech-to-speech translation, pages 537–549. springer, 2000. j. e. cahn and s. e. brennan. a psychological model of grounding and repair in dialog. in aaai symposium on psychological models of communication in collaborative systems, 1999. a. canavan and g. zipperlen. callfriend american english-non-southern dialect. linguistic data consortium, 10:1, 1996. a. canavan, d. graff, and g. zipperlen. callhome american english speech. linguistic data consortium, 1997. s. k. card, t. p. moran, and a. newell. the psychology of human-computer interaction. l. erlbaum associates inc., hillsdale, nj, usa, 1983. isbn 0898592437. r. carter. orders of reality: cancode, communication, and culture. elt journal, 52(1):43–56, 1998. r. carter and m. mccarthy. cambridge grammar of english: a comprehensive guide; spoken and written english grammar and usage. ernst klett sprachen, 2006. t. l. chartrand and j. a. bargh. the chameleon effect: the perception–behavior link and social interaction. journal of personality and social psychology, 76(6):893, 1999. t. chen and m. kan. creating a live, public short message service corpus: the nus sms corpus. language resources and evaluation, 47(2):299–335, 2013. y.-n. chen, d. hakkani-tür, and x. he. zero-shot learning of intent embeddings for expansion by convolutional deep structured semantic models. in acoustics, speech and signal processing (icassp), 2016 ieee international conference on, pages 6045–6049. ieee, 2016. g. churcher, e. s. atwell, and c. souter. dialogue management systems: a survey and overview. university of leeds, school of computing research report 1997.06, 1997. p. r. cohen. if not turing’s test, then what? ai magazine, 26(4):61, 2005. k. m. colby. modeling a paranoid mind. behavioral and brain sciences, 4:515–534, 1981. r. m. cooper. the control of eye fixation by the meaning of spoken language: a new methodology for the real-time investigation of speech perception, memory, and language processing. cognitive psychology, 6(1):84–107, 1974. n. cristianini and j. shawe-taylor. an introduction to support vector machines: and other kernel-based learning methods. cambridge university press, 2000. h. cuayáhuitl, s. keizer, and o. lemon. strategic dialogue management via deep reinforcement learning. arxiv preprint arxiv:1511.08099, 2015. c. danescu-niculescu-mizil and l. lee. chameleons in imagined conversations: a new approach to understanding coordination of linguistic style in dialogs. in association for computational linguistics, workshop on cognitive modeling and computational linguistics, 2011. 41 serban, lowe, henderson, charlin and pineau l. daubigney, m. geist, s. chandramohan, and o. pietquin. a comprehensive reinforcement learning framework for dialogue management optimization. ieee journal of selected topics in signal processing, 6(8):891–902, 2012. m. davies. comparing the corpus of american soap operas, coca, and the bnc, 2012a. m. davies. corpus of american soap operas, 2012b. i. de kok, d. heylen, and l. morency. speaker-adaptive multimodal prediction model for listener responses. in proceedings of the 15th acm on international conference on multimodal interaction, 2013. l. deng and x. li. machine learning paradigms for speech recognition: an overview. audio, speech, and language processing, ieee transactions on, 21(5):1060–1089, 2013. l. denoyer and p. gallinari. the wikipedia xml corpus. in comparative evaluation of xml information retrieval systems, pages 12–19. springer, 2007. b. dhingra, z. zhou, d. fitzpatrick, m. muehl, and w. cohen. tweet2vec: character-based distributed representations for social media. arxiv preprint arxiv:1605.03481, 2016. a. djalali, s. lauer, and c. potts. corpus evidence for preference-driven interpretation. in logic, language and meaning, pages 150–159. springer, 2012. j. dodge, a. gane, x. zhang, a. bordes, s. chopra, a. miller, a. szlam, and j. weston. evaluating prerequisite qualities for learning end-to-end dialog systems. arxiv preprint arxiv:1511.06931, 2015. s. dose. flipping the script: a corpus of american television series (cats) for corpus-based language learning and teaching. corpus linguistics and variation in english: focus on non-native englishes, 2013. e. douglas-cowie, r. cowie, i. sneddon, c. cox, o. lowry, m. mcrorie, j. martin, l. devillers, s. abrilian, a. batliner, et al. the humaine database: addressing the collection and annotation of naturalistic and induced emotional data. in affective computing and intelligent interaction, pages 488–500. springer, 2007. w. eckert, e. levin, and r. pieraccini. user modeling for spoken dialogue system evaluation. in automatic speech recognition and understanding, 1997. proceedings., 1997 ieee workshop on, pages 80–87, 1997. l. el asri, h. schulz, s. sharma, j. zumer, j. harris, e. fine, r. mehrotra, and k. suleman. frames: acorpus for adding memory to goal-oriented dialogue systems. preprint on webpage at http://www.maluuba.com/ publications/, 2017. m. elsner and e. charniak. you talking to me? a corpus and algorithm for conversation disentanglement. in association for computational linguistics (acl), 2008. m. elsner and e. charniak. disentangling chat. computational linguistics, 36(3):389–409, 2010. d. erhan, y. bengio, a. courville, pierre-a. manzagol, and p. vincent. why does unsupervised pre-training help deep learning? journal of machine learning research, 11, 2010. g. di fabbrizio, g. tur, and d. hakkani-tr. bootstrapping spoken dialog systems with data reuse. in special interest group on discourse and dialogue (sigdial), 2004. m. fatemi, l. e. asri, h. schulz, j. he, and k. suleman. policy networks with two-stage training for dialogue systems. in special interest group on discourse and dialogue (sigdial), 2016. maryam fazel-zarandi, shang-wen li, jin cao, jared casale, peter henderson, david whitney, and alborz geramifard. learning robust dialog policies in noisy environments. neural information processing systems (nips), conversational ai workshop, 2017. d. fisher, m. smith, and h. t welser. you are who you talk to: detecting roles in usenet newsgroups. in hawaii international conference on system sciences (hicss’06), volume 3, pages 59b–59b, 2006. p. forchini. spontaneity reloaded: american face-to-face and movie conversation compared. in corpus linguistics, 2009. p. forchini. movie language revisited. evidence from multi-dimensional analysis and corpora. peter lang, 2012. g. forgues, j. pineau, j. larchevêque, and r. tremblay. bootstrapping dialog systems with word embeddings. in workshop on modern machine learning and natural language processing, neural information processing systems (nips), 2014. e. n. forsyth and c. h. martell. lexical and discourse analysis of online chat dialog. in international conference on semantic computing (icsc)., pages 19–26, 2007. m. galley, c. brockett, a. sordoni, y. ji, m. auli, c. quirk, m. mitchell, j. gao, and b. dolan. deltableu: a discriminative metric for generation tasks with intrinsically diverse targets. in association for computational linguistics and international joint conference on natural language processing of the asian federation of natural language processing, pages 445–450, 2015. m. gašić, f. jurčı́ček, s. keizer, f. mairesse, b. thomson, k. yu, and s. young. gaussian processes for fast policy optimisation of pomdp-based dialogue managers. in special interest group on discourse and dialogue (sigdial). acl, 2010. 42 http://www.maluuba.com/publications/ http://www.maluuba.com/publications/ a survey of available corpora for building data-driven dialogue systems m. gašić, f. jurčı́ček, b. thomson, k. yu, and s. young. on-line policy optimisation of spoken dialogue systems via live interaction with human subjects. in ieee workshop on automatic speech recognition and understanding (asru), pages 312–317. ieee, 2011. m. gašić, m. henderson, b. thomson, p. tsiakoulis, and s. young. policy optimisation of pomdp-based dialogue systems without state space compression. in spoken language technology workshop (slt), 2012 ieee, pages 31–36. ieee, 2012. m. gasic, c. breslin, m. henderson, d. kim, m. szummer, b. thomson, p. tsiakoulis, and s. young. on-line policy optimisation of bayesian spoken dialogue systems via human interaction. in ieee international conference on acoustics, speech and signal processing, pages 8367–8371, 2013. m. gašić, n. mrkšić, l. m. rojas-barahona, p.-h. su, s. ultes, d. vandyke, t.-h. wen, and s. young. dialogue manager domain adaptation using gaussian process reinforcement learning. computer speech & language, 2016. a. genevay and r. laroche. transfer learning for user adaptation in spoken dialogue systems. in international conference on autonomous agents & multiagent systems, pages 975–983. international foundation for autonomous agents and multiagent systems, 2016. k. georgila, j. henderson, and o. lemon. user simulation for spoken dialogue systems: learning and evaluation. in interspeech, 2006. k. georgila, m. wolters, j. d. moore, and r. h. logie. the match corpus: a corpus of older and younger users interactions with spoken dialogue systems. language resources and evaluation, 44(3):221–261, 2010. j. gibson and a. d. pick. perception of another person’s looking behavior. the american journal of psychology, 76(3): 386–394, 1963. j. j godfrey, e. c holliman, and j mcdaniel. switchboard: telephone speech corpus for research and development. in international conference on acoustics, speech, and signal processing (icassp-92), 1992. i. goodfellow, a. courville, and y. bengio. deep learning. book in preparation for mit press, 2015. url http: //goodfeli.github.io/dlbook/. c. goodwin. conversational organization: interaction between speakers and hearers. new york: academic press, 1981. a. l. gorin, g. riccardi, and j. h. wright. how may i help you? speech communication, 23(1):113–127, 1997. a. graves. sequence transduction with recurrent neural networks. in international conference on machine learning (icml), representation learning workshop, 2012. s. greenbaum. comparing english worldwide: the international corpus of english. clarendon press, 1996. s. greenbaum and g nelson. the international corpus of english (ice) project. world englishes, 15(1):3–15, 1996. c. gülçehre, o. firat, k. xu, k. cho, l. barrault, h. lin, f. bougares, h. schwenk, and y. bengio. on using monolingual corpora in neural machine translation. corr, abs/1503.03535, 2015. i. gurevych and m. strube. semantic similarity applied to spoken dialogue summarization. in international conference on computational linguistics (coling), 2004. v. haslerud and a. stenström. the bergen corpus of london teenager language (colt). spoken english on computer. transcription, mark-up and application. london: longman, pages 235–242, 1995. p. a. heeman and j. f. allen. the trains 93 dialogues. technical report, dtic document, 1995. c. t. hemphill, j. j. godfrey, and g. r. doddington. the atis spoken language systems pilot corpus. in darpa speech and natural language workshop, pages 96–101, 1990. m. henderson, b. thomson, and s. young. deep neural network approach for the dialog state tracking challenge. in special interest group on discourse and dialogue (sigdial), 2013. m. henderson, b. thomson, and j. williams. dialog state tracking challenge 2 & 3, 2014a. m. henderson, b. thomson, and j. williams. the second dialog state tracking challenge. in special interest group on discourse and dialogue (sigdial), 2014b. m. henderson, b. thomson, and s. young. word-based dialog state tracking with recurrent neural networks. in special interest group on discourse and dialogue (sigdial), 2014c. p. henderson, k. sinha, n. angelard-gontier, n. r. ke, g. fried, r. lowe, and j. pineau. ethical challenges in datadriven dialogue systems. arxiv preprint arxiv:1711.09050, 2017. g. hinton, l. deng, d. yu, g. e. dahl, a. mohamed, n. jaitly, a. senior, v. vanhoucke, p. nguyen, t.a n. sainath, et al. deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups. signal processing magazine, ieee, 29(6):82–97, 2012. t. hiraoka, g. neubig, k. yoshino, t. toda, and s. nakamura. active learning for example-based dialog systems. in proc intl workshop on spoken dialog systems, saariselka, finland, 2016. 43 http://goodfeli.github.io/dlbook/ http://goodfeli.github.io/dlbook/ serban, lowe, henderson, charlin and pineau h. hung and g. chittaranjan. the idiap wolf corpus: exploring group behaviour in a competitive role-playing game. in international conference on multimedia, pages 879–882, 2010. j. l. hutchens and m. d. alder. introducing megahal. in joint conferences on new methods in language processing and computational natural language learning, 1998. j. l. part i. shalyminov x. xu y. yu o. duek v. rieser o. lemon i. papaioannou, a. c. curry. alana: social dialogue using an ensemble model and a ranker trained on user feedback. in alexa prize proceedings, 2017. a. jonsson and n. dahlback. talking to a computer is not like talking to your best friend. in scandinavian conference on artificial intelligence, 1988. s. jung, c. lee, k. kim, m. jeong, and g. g. lee. data-driven user simulation for automated evaluation of spoken dialog systems. computer speech & language, 23(4):479–509, 2009. d. jurafsky and j. h. martin. speech and language processing, 2nd edition. prentice hall, 2008. f. jurcıcek, s. keizer, m. gašic, f. mairesse, b. thomson, k. yu, and s. young. real user evaluation of spoken dialogue systems using amazon mechanical turk. in interspeech, volume 11, 2011. r. kadlec, m. schmid, and j. kleindienst. improved deep learning baselines for ubuntu corpus dialogs. neural information processing systems workshop on machine learning for spoken language understanding, 2015. s. kim, l. f. dharo, r. e. banchs, j. williams, and m. henderson. dialog state tracking challenge 4, 2015. s. kim, l. f. dharo, r. e. banchs, j. d. williams, m. henderson, and k. yoshino. the fifth dialog state tracking challenge. in ieee spoken language technology workshop (slt), 2016. v. konovalov, o. melamud, r. artstein, and i. dagan. collecting better training data using biased agent policies in negotiation dialogues. in wochat: workshop on chatbots and conversational agent technologies, 2016. j. a konstan, b. n. miller, d. maltz, j. l. herlocker, l. r. gordon, and j. riedl. grouplens: applying collaborative filtering to usenet news. communications of the acm, 40(3):77–87, 1997. a. kumar, o. irsoy, j. su, j. bradbury, r. english, b. pierce, p. ondruska, i. gulrajani, and r. socher. ask me anything: dynamic memory networks for natural language processing. neural information processing systems (nips), 2015. m. kytö and t. walker. guide to a corpus of english dialogues 1560-1760. acta universitatis upsaliensis, 2006. i. langkilde and k. knight. generation that exploits corpus-based statistical knowledge. in association for computational linguistics and international conference on computational linguistics, volume 1, pages 704–710. acl, 1998. c.-j. lee, s.-k. jung, k.-d. kim, d.-h. lee, and g. g. lee. recent approaches to dialog management for spoken dialog systems. journal of computing science and engineering, 4(1):1–22, 2010. g. leech. 100 million words of english: the british national corpus (bnc). language research, 28(1):1–13, 1992. o. lemon. adaptive natural language generation in dialogue using reinforcement learning. in workshop on the semantics and pragmatics of dialogue (semdial), volume 8, pages 149–156, 2008. e. levin and r. pieraccini. a stochastic model of computer-human interaction for learning dialogue strategies. in eurospeech, volume 97, pages 1883–1886, 1997. e. levin, r. pieraccini, and w. eckert. learning dialogue strategies within the markov decision process framework. in automatic speech recognition and understanding, 1997. proceedings., 1997 ieee workshop on, pages 72–79. ieee, 1997. j. li, m. galley, c. brockett, j. gao, and b. dolan. a diversity-promoting objective function for neural conversation models. arxiv preprint arxiv:1510.03055, 2015. j. li, m. galley, c. brockett, j. gao, and bill d. a persona-based neural conversation model. in association for computational linguistics, pages 994–1003, 2016. y. li, h. su, x. shen, w. li, z. cao, and s. niu. dailydialog: a manually labelled multi-turn dialogue dataset. arxiv preprint arxiv:1710.03957, 2017. g. lin and m. walker. all the world’s a stage: learning character models from film. in aaai conference on artificial intelligence and interactive digital entertainment, 2011a. g. i. lin and m. a. walker. all the world’s a stage: learning character models from film. in aiide, 2011b. b. liu and i. lane. multi-domain adversarial learning for slot filling in spoken language understanding. neural information processing systems (nips), conversational ai workshop, 2017. c.-w. liu, r. lowe, i. serban, m. noseworthy, l. charlin, and j. pineau. how not to evaluate your dialogue system: an empirical study of unsupervised evaluation metrics for dialogue response generation. in conference on empirical methods in natural language processing (emnlp), pages 2122–2132, 2016. c. lord and m. haith. the perception of eye contact. attention, perception, & psychophysics, 16(3):413–416, 1974. r. lowe, n. pow, i. serban, and j. pineau. the ubuntu dialogue corpus: a large dataset for research in unstructured multi-turn dialogue systems. in special interest group on discourse and dialogue (sigdial), 2015a. 44 a survey of available corpora for building data-driven dialogue systems r. lowe, n. pow, i. v. serban, l. charlin, and j. pineau. incorporating unstructured textual knowledge sources into neural dialogue systems. neural information processing systems workshop on machine learning for spoken language understanding, 2015b. r. lowe, i. v. serban, m. noseworthy, l. charlin, and j. pineau. on the evaluation of dialogue systems with next utterance classification. in special interest group on discourse and dialogue (sigdial), 2016. j. m. lucas, f. fernndez, j. salazar, j. ferreiros, and r. san segundo. managing speaker identity and user profiles in a spoken dialogue system. in procesamiento del lenguaje natural, number 43 in 1, pages 77–84, 2009. b. macwhinney and c. snow. the child language data exchange system. journal of child language, 12(02):271–295, 1985. f. mairesse and s. young. stochastic language generation in dialogue using factored language models. computational linguistics, 2014. f. mairesse, m. gašić, f. jurčı́ček, s. keizer, b. thomson, k. yu, and s. young. phrase-based statistical language generation using graphical models and active learning. in the association for computational linguistics, pages 1552– 1561. acl, 2010. c. d. manning and h. schütze. foundations of statistical natural language processing. mit press, 1999. m. mccarthy. spoken language and applied linguistics. ernst klett sprachen, 1998. s. mcglashan, n. fraser, n. gilbert, e. bilange, p. heisterkamp, and n. youd. dialogue management for telephone information systems. in conference on applied natural language processing, pages 245–246. acl, 1992. g. mckeown, m. f valstar, r. cowie, and m. pantic. the semaine corpus of emotionally coloured character interactions. in multimedia and expo (icme), 2010 ieee international conference on, pages 1079–1084, 2010. g. mesnil, x. he, l. deng, and y. bengio. investigation of recurrent-neural-network architectures and learning methods for spoken language understanding. in interspeech, pages 3771–3775, 2013. t. mikolov, m. karafiát, l. burget, j. cernockỳ, and sanjeev khudanpur. recurrent neural network based language model. in interspeech, pages 1045–1048, 2010. t. mikolov, i. sutskever, k. chen, g. s. corrado, and j. dean. distributed representations of words and phrases and their compositionality. in neural information processing systems, pages 3111–3119, 2013. g. a. miller. wordnet: a lexical database for english. communications of the acm, 38(11):39–41, 1995. s. miller, d. stallard, r. bobrow, and r. schwartz. a fully statistical approach to natural language interfaces. in association for computational linguistics, pages 55–61, 1996. k. mo, s. li, y. zhang, j. li, and q. yang. personalizing a dialogue system with transfer learning. arxiv preprint arxiv:1610.02891, 2016. s. mohan and j. laird. learning goal-oriented hierarchical tasks from situated interactive instruction. in aaai, 2014. t. nguyen, m. rosenberg, x. song, j. gao, s. tiwary, r. majumder, and l. deng. ms marco: a human generated machine reading comprehension dataset. arxiv preprint arxiv:1611.09268, 2016. l nio, s. sakti, g. neubig, t. toda, and s. nakamura. conversation dialog corpora from television and movie scripts. in 17th oriental chapter of the international committee for the co-ordination and standardization of speech databases and assessment techniques (cocosda), pages 1–4, 2014. e. nöth, a. horndasch, f. gallwitz, and j. haas. experiences with commercial telephone-based dialogue systems. it– information technology (vormals it+ ti), 46(6/2004):315–321, 2004. c. oertel, f. cummins, j. edlund, p. wagner, and n. campbell. d64: a corpus of richly recorded conversational interaction. journal on multimodal user interfaces, 7(1-2):19–28, 2013. a. h. oh and a. i. rudnicky. stochastic language generation for spoken dialogue systems. in anlp/naacl workshop on conversational systems, volume 3, pages 27–32. acl, 2000. t. paek. reinforcement learning for spoken dialogue systems: comparing strengths and weaknesses for practical deployment. in proc. dialog-on-dialog workshop, interspeech, 2006. k. papineni, s. roukos, t ward, and w zhu. bleu: a method for automatic evaluation of machine translation. in association for computational linguistics, 2002. g. parent and m. eskenazi. toward better crowdsourced transcription: transcription of a year of the let’s go bus information system data. in spoken language technology workshop (slt), 2010 ieee, pages 312–317. ieee, 2010. a. n. pargellis, h-k. j. kuo, and c. lee. an automatic dialogue generation platform for personalized dialogue applications. speech communication, 42(3-4):329–351, 2004. doi: 10.1016/j.specom.2003.10.003. r. passonneau and e. sachar. loqui human-human dialogue corpus (transcriptions and annotations), 2014. b. peng, x. li, l. li, j. gao, a. celikyilmaz, s. lee, and k.-f. wong. composite task-completion dialogue policy learning via hierarchical deep reinforcement learning. in conference on empirical methods in natural language processing (emnlp), pages 2221–2230, 2017. 45 serban, lowe, henderson, charlin and pineau d. perez-marin and i. pascual-nieto. conversational agents and natural language interaction: techniques and effective practices. igi global, 2011. s. petrik. wizard of oz experiments on speech dialogue systems. phd thesis, technischen universitat graz, 2004. r. pieraccini, d. suendermann, k. dayanidhi, and j. liscombe. are we there yet? research in commercial spoken dialog systems. in text, speech and dialogue, pages 3–13, 2009. o. pietquin and h. hastie. a survey on metrics for the evaluation of user simulations. the knowledge engineering review, 28(01):59–73, 2013. o. pietquin, m. geist, s. chandramohan, and h. frezza-buet. sample-efficient batch reinforcement learning for dialogue management optimization. acm transactions on speech and language processing (tslp), 7(3):7, 2011. b. piot, m. geist, and o. pietquin. imitation learning applied to embodied conversational agents. in 4th workshop on machine learning for interactive systems (mlis 2015), volume 43, 2015. c. potts. goal-driven answers in the cards dialogue corpus. in west coast conference on formal linguistics, pages 1–20, 2012. a. ratnaparkhi. trainable approaches to surface natural language generation and their application to conversational dialog systems. computer speech & language, 16(3):435–455, 2002. a. raux, b. langner, d. bohus, a. w. black, and m. eskenazi. lets go public! taking a spoken dialog system to the real world. in interspeech, 2005. n. reithinger and m. klesen. dialogue act classification using language models. in eurospeech, 1997. h. ren, w. xu, y. zhang, and y. yan. dialog state tracking using conditional random fields. in special interest group on discourse and dialogue (sigdial), 2013. s. renals, t. hain, and h. bourlard. recognition and understanding of meetings the ami and amida projects. in ieee workshop on automatic speech recognition & understanding (asru), 2007. r. reppen and n. ide. the american national corpus overall goals and the first release. journal of english linguistics, 32(2):105–113, 2004. j. rickel and w. l. johnson. animated agents for procedural training in virtual reality: perception, cognition, and motor control. applied artificial intelligence, 13(4-5):343–382, 1999. v. rieser and o. lemon. natural language generation as planning under uncertainty for spoken dialogue systems. in empirical methods in natural language generation, pages 105–120. springer, 2010. verena rieser and oliver lemon. reinforcement learning for adaptive dialogue systems: a data-driven methodology for dialogue management and natural language generation. springer science & business media, 2011. a. ritter, c. cherry, and b. dolan. unsupervised modeling of twitter conversations. in north american chapter of the association for computational linguistics (naacl 2010), 2010. a. ritter, c. cherry, and w. b. dolan. data-driven response generation in social media. in conference on empirical methods in natural language processing (emnlp), 2011. s. rosenthal and k. mckeown. i couldnt agree more: the role of conversational structure in agreement and disagreement detection in online discussions. in special interest group on discourse and dialogue (sigdial), 2015. s. rosset and s. petel. the ritel corpus-an annotated human-machine open-domain question answering spoken dialog corpus. in the international conference on language resources and evaluation (lrec), 2006. a. roy, c. guinaudeau, h. bredin, and c. barras. tvd: a reproducible and multiply aligned tv series dataset. in the international conference on language resources and evaluation (lrec), volume 2, 2014. j. ruppenhofer, m. ellsworth, m. r.l. petruck, c. r. johnson, and j. scheffczyk. framenet ii: extended theory and practice. international computer science institute, 2006. distributed with the framenet data. j. schatzmann and s. young. the hidden agenda user simulation model. ieee transactions on audio, speech, and language processing, 17(4):733–747, 2009. j. schatzmann, k. georgila, and s. young. quantitative evaluation of user simulation techniques for spoken dialogue systems. in special interest group on discourse and dialogue (sigdial), 2005. j. schatzmann, b. thomson, k. weilhammer, . ye, and s. young. agenda-based user simulation for bootstrapping a pomdp dialogue system. in human language technologies 2007: the conference of the north american chapter of the association for computational linguistics; companion volume, short papers, pages 149–152, 2007. k. scheffler and s. young. automatic learning of dialogue strategy using dialogue simulation and reinforcement learning. in international conference on human language technology research, pages 12–19. morgan kaufmann publishers inc., 2002. j. n. schrading. analyzing domestic abuse using natural language processing on social media data. master’s thesis, rochester institute of technology, 2015. http://scholarworks.rit.edu/theses. 46 http://scholarworks.rit.edu/theses a survey of available corpora for building data-driven dialogue systems n. schrading, c. o. alm, r. ptucha, and c. m. homan. an analysis of domestic a.se discourse on reddit. in conference on empirical methods in natural language processing (emnlp), 2015. k. k. schuler. verbnet: a broad-coverage, comprehensive verb lexicon. phd thesis, university of pennsylvania, 2005. paper aai3179808. i. v. serban and j. pineau. text-based speaker identification for multi-participant open-domain dialogue systems. neural information processing systems workshop on machine learning for spoken language understanding, 2015. i. v. serban, r. lowe, l. charlin, and j. pineau. generative deep neural networks for dialogue: a short review. in neural information processing systems (nips), let’s discuss: learning methods for dialogue workshop, 2016. i. v. serban, a. sordoni, y. bengio, a. courville, and j. pineau. building end-to-end dialogue systems using generative hierarchical neural networks. in aaai, 2016. in press. i. v. serban, t. klinger, g. tesauro, k. talamadupula, b. zhou, y. bengio, and a. courville. multiresolution recurrent neural networks: an application to dialogue response generation. in aaai conference, 2017a. i. v. serban, c. sankar, m. germain, s. zhang, z. lin, s. subramanian, t. kim, m. pieper, s. chandar, n. r. ke, et al. a deep reinforcement learning chatbot. arxiv preprint arxiv:1709.02349, 2017b. i. v serban, c. sankar, s. zhang, z. lin, s. subramanian, t. kim, s. chandar, n. r. ke, et al. the octopus approach to the alexa competition: a deep ensemble-based socialbot. in alexa prize proceedings, 2017c. i. v. serban, a. sordoni, r. lowe, l. charlin, j. pineau, a. courville, and y. bengio. a hierarchical latent variable encoder-decoder model for generating dialogues. in aaai conference, 2017d. s. shaikh, t. strzalkowski, g. a. broadwell, j. stromer-galley, s. m. taylor, and n. webb. mpc: a multi-party chat corpus for modeling social phenomena in discourse. in the international conference on language resources and evaluation (lrec), 2010. l. shang, z. lu, and h. li. neural responding machine for short-text conversation. arxiv preprint arxiv:1503.02364, 2015. c. shaoul and c. westbury. a usenet corpus (2005-2009), 2009. s. sharma, j. he, k. suleman, h. schulz, and p. bachman. natural language generation in dialogue using lexicalized and delexicalized data. arxiv preprint arxiv:1606.03632, 2016. b. a. shawar and e. atwell. different measurements metrics to evaluate a chatbot system. in workshop on bridging the gap: academic and industrial research in dialog technologies, pages 89–96, 2007a. b. a. shawar and eric atwell. chatbots: are they really useful? in ldv forum, volume 22, pages 29–49, 2007b. e. shriberg, r. dhillon, s. bhagat, j. ang, and h. carvey. the icsi meeting recorder dialog act (mrda) corpus. technical report, dtic document, 2004. a. simpson and n. m eraser. black box and glass box evaluation of the sundial system. in european conference on speech communication and technology, 1993. s. singh, d. litman, m. kearns, and m. walker. optimizing dialogue management with reinforcement learning: experiments with the njfun system. journal of artificial intelligence research, pages 105–133, 2002. s. p. singh, m. j. kearns, d. j. litman, and m. a. walker. reinforcement learning for spoken dialogue systems. in neural information processing systems, 1999. a. sordoni, m. galley, m. auli, c. brockett, y. ji, m. mitchell, j. nie, j. gao, and b. dolan. a neural network approach to context-sensitive generation of conversational responses. in conference of the north american chapter of the association for computational linguistics (naacl-hlt 2015), 2015. a. stenström, g. andersen, and i. k. hasund. trends in teenage talk: corpus compilation, analysis and findings, volume 8. j. benjamins, 2002. a. stent, r. prasad, and m. walker. trainable sentence planning for complex information presentation in spoken dialog systems. in association for computational linguistics, page 79. acl, 2004. a. stolcke, k. ries, n. coccaro, e. shriberg, r. bates, d. jurafsky, p. taylor, r. martin, c. van ess-dykema, and m. meteer. dialogue act modeling for automatic tagging and recognition of conversational speech. computational linguistics, 26(3):339–373, 2000. p.-h. su, y.-b. wang, t.-h. yu, and l.-s. lee. a dialogue game framework with personalized training using reinforcement learning for computer-assisted language learning. in 2013 ieee international conference on acoustics, speech and signal processing, pages 8213–8217. ieee, 2013. p.-h. su, d. vandyke, m. gasic, d. kim, n. mrksic, t.-h. wen, and s. young. learning from real users: rating dialogue success with neural networks for reinforcement learning in spoken dialogue systems. in interspeech, 2015. p.-h. su, m. gasic, n. mrksic, l. rojas-barahona, s. ultes, d. vandyke, t.-h. wen, and s. young. continuously learning neural dialogue management. arxiv preprint arxiv:1606.02689, 2016. 47 serban, lowe, henderson, charlin and pineau x. sun, j. lichtenauer, m. valstar, a. nijholt, and m. pantic. a multimodal database for mimicry analysis. in affective computing and intelligent interaction, pages 367–376. springer, 2011. j. svartvik. the london-lund corpus of spoken english: description and research. number 82 in 1. lund university press, 1990. b. thomson and s. young. bayesian update of dialogue state: a pomdp framework for spoken dialogue systems. computer speech & language, 24(4):562–588, 2010. j. tiedemann. parallel data, tools and interfaces in opus. in the international conference on language resources and evaluation (lrec), 2012. d. traum and j. rickel. embodied agents for multi-party dialogue in immersive virtual worlds. in international joint conference on autonomous agents and multiagent systems: part 2, pages 766–773. acm, 2002. g. tur and l. deng. intent determination and spoken utterance classification, chapter 4, pages 81–104. wiley, january 2011. a. m. turing. computing machinery and intelligence. mind, pages 433–460, 1950. d. c uthus and d. w aha. the ubuntu chat corpus for multiparticipant chat analysis. in aaai spring symposium: analyzing microtext, 2013. j. vandeventer, a. j. aubrey, p. l. rosin, and d. marshall. 4d cardiff conversation database (4d ccdb): a 4d database of natural, dyadic conversations. in joint conference on facial analysis, animation and auditory-visual speech processing (faavsp 2015), 2015. d. vandyke, p.-h. su, m. gasic, n. mrksic, t.-h. wen, and s. young. multi-domain dialogue success classifiers for policy training. in automatic speech recognition and understanding (asru), 2015 ieee workshop on, pages 763– 770. ieee, 2015. o. vinyals and q. le. a neural conversational model. arxiv preprint arxiv:1506.05869, 2015. m. a. walker, d. j. litman, c. a. kamm, and a. abella. paradise: a framework for evaluating spoken dialogue agents. in european chapter of the association for computational linguistics (eacl), pages 271–280, 1997. m. a. walker, o. c. rambow, and m. rogati. training a sentence planner for spoken dialogue using boosting. computer speech & language, 16(3):409–433, 2002. m. a. walker, r. grant, j. sawyer, g. i. lin, n. wardrip-fruin, and m. buell. perceived or not perceived: film character models for expressive nlg. in icids, pages 109–121, 2011. m. a walker, g. i. lin, and j. sawyer. an annotated corpus of film dialogue for learning and characterizing character style. in the international conference on language resources and evaluation (lrec), pages 1373–1378, 2012a. m. a walker, j. e. f. tree, p. anand, r. abbott, and j. king. a corpus for research on deliberation and debate. in the international conference on language resources and evaluation (lrec), pages 812–817, 2012b. r. s. wallace. the anatomy of alice. parsing the turing test, pages 181–210, 2009. ye-yi wang, li deng, and alex acero. spoken language understanding. ieee signal processing magazine, 22(5): 16–31, 2005. z. wang and o. lemon. a simple and generic belief tracking mechanism for the dialog state tracking challenge: on the believability of observed information. in special interest group on discourse and dialogue (sigdial), 2013. s. webb. a corpus driven study of the potential for vocabulary learning through watching movies. international journal of corpus linguistics, 15(4):497–519, 2010. j. weizenbaum. elizaa computer program for the study of natural language communication between man and machine. communications of the acm, 9(1):36–45, 1966. t. wen, m. gašic, d. kim, n. mrkšic, p. su, d. vandyke, and s. young. stochastic language generation in dialogue using recurrent neural networks with convolutional sentence reranking. special interest group on discourse and dialogue (sigdial), 2015a. t.-h. wen, m. gasic, n. mrksic, p.-h. su, d. vandyke, and s. young. semantically conditioned lstm-based natural language generation for spoken dialogue systems. in conference on empirical methods in natural language processing (emnlp), 2015b. t.-h. wen, m. gasic, n. mrksic, l. m. rojas-barahona, p.-h. su, d. vandyke, and s. young. multi-domain neural network language generation for spoken dialogue systems. in conference of the north american chapter of the association for computational linguistics (naacl-hlt 2016), 2016. t.-h. wen, m. gasic, n. mrksic, l. m. rojas-barahona, p.-h. su, s. ultes, d. vandyke, and s. young. a networkbased end-to-end trainable task-oriented dialogue system. in european chapter of the association for computational linguistics (eacl), 2017. j. weston. dialog-based language learning. arxiv preprint arxiv:1604.06045, 2016. 48 a survey of available corpora for building data-driven dialogue systems j. williams, a. raux, d. ramachandran, and a. black. the dialog state tracking challenge. in special interest group on discourse and dialogue (sigdial), 2013. j. d. williams and s. young. partially observable markov decision processes for spoken dialog systems. computer speech & language, 21(2):393–422, 2007. j. d. williams and g. zweig. end-to-end lstm-based dialog control optimized with supervised and reinforcement learning. arxiv preprint arxiv:1606.01269, 2016. m. wolska, q. b. vo, d. tsovaltzi, i. kruijff-korbayová, e. karagjosova, h. horacek, a. fiedler, and c. benzmüller. an annotated corpus of tutorial dialogs on mathematical theorem proving. in the international conference on language resources and evaluation (lrec), 2004. b. wrede and e. shriberg. relationship between dialogue acts and hot spots in meetings. in automatic speech recognition and understanding, 2003. asru’03. 2003 ieee workshop on, pages 180–185. ieee, 2003. x. yang, y.-n. chen, d. hakkani-tür, p. crook, x. li, j. gao, and l. deng. end-to-end joint learning of natural language understanding and dialogue manager. in acoustics, speech and signal processing (icassp), 2017 ieee international conference on, pages 5690–5694. ieee, 2017. y. yang, w. yih, and c. meek. wikiqa: a challenge dataset for open-domain question answering. in conference on empirical methods in natural language processing (emnlp), pages 2013–2018, 2015. z. yang, b. li, y. zhu, i. king, g. levow, and h. meng. collection of user judgments on spoken dialog system with crowdsourcing. in spoken language technology workshop (slt), 2010 ieee, pages 277–282, 2010. s. young, m. gasic, b. thomson, and j. d. williams. pomdp-based statistical spoken dialog systems: a review. ieee, 101(5):1160–1179, 2013. s. j. young. probabilistic methods in spoken–dialogue systems. philosophical transactions of the royal society of london. series a: mathematical, physical and engineering sciences, 358(1769), 2000. j. zhang, r. kumar, s. ravi, and c. danescu-niculescu-mizil. conversational flow in oxford-style debates. in conference of the north american chapter of the association for computational linguistics (naacl-hlt 2016), 2016. 49 introduction characteristics of data-driven dialogue systems an overview of dialogue systems tasks and objectives learning dialogue system components natural language understanding dialogue management natural language generator end-to-end dialogue systems dialogue interaction types & aspects written, spoken & multi-modal corpora human-human vs. human-machine corpora spontaneous vs. constrained corpora topic-oriented & goal-driven datasets scripted corpora or corpora from fiction corpus size available dialogue datasets human-machine corpora restaurant and travel information open-domain knowledge retrieval other human-human spoken corpora spontaneous spoken corpora constrained spoken corpora scripted corpora human-human written corpora spontaneous written corpora constrained written corpora discussion challenges of learning from large datasets the twitter corpus the ubuntu dialogue corpus transfer learning between datasets incorporating external knowledge structured external knowledge unstructured external knowledge personalized dialogue agents evaluation metrics automatic evaluation metrics for goal-driven dialogue systems automatic evaluation metrics for non-goal-driven dialogue systems conclusion dialogue & discourse 16(3) (2025) 25–59 doi: 10.5210/dad.2025.303 a modular architecture for creating multimodal embodied agents with an episodic knowledge graph as an explainable and controllable long-term memory thomas baier t.baier@vu.nl vrije universiteit amsterdam amsterdam, the netherlands selene báez santamaría selene.baez.santamaria@gmail.com vrije universiteit amsterdam amsterdam, the netherlands piek vossen p.t.j.m.vossen@vu.nl vrije universiteit amsterdam amsterdam, the netherlands editor: hendrik buschmeier submitted 11/2023; accepted 04/2025; published online 12/2025 abstract how can flexibility and control over the interpretation of multimodal signals by embodied agents be balanced? flexibility means that agents respond fluently in any context, whereas control means that responses are transparent and faithful to goals and principles that are explicitly defined. this paper describes a modular platform to create multimodal interactive agents using an event bus on which signals and interpretations are posted as a sequence in time, but also provides control options to drive the interaction given specific intentions and goals. different sensors and interpretation components can be integrated by defining their input and output topics in the event bus, which results in an open multimodal sequence-driven workflow for further interpretations. in addition, our platform allows us to define higher-level intents that control sequence patterns to achieve a goal. a key component is an episodic knowledge graph (ekg) that acts as a long-term symbolic memory to aggregate and connect these interpretations. this ekg establishes coherence and continuity across different interactions. intents and the ekg make it possible to define different (embodied) agents and compare their behavior without having to implement complex software components for multimodal sensor data and design the control over their dependencies. in this paper, we explain the broad range of components that we developed and integrated into various interactive agents. we also explain how the interaction is recorded as multimodal data and how it results in an aggregated memory in the ekg. by analyzing the recorded interaction, we can compare agents and agent components and study their interactive behavior with people and other agents. keywords: multimodal agent interaction; modularity, flexibility, and control; data sharing and technology sharing. 1. introduction interaction among agents, both humans and ai, is especially complex when it occurs in real world contexts. it involves a complex psychosocial process between agents with intentions, capabilities, and ©2025 thomas baier, selene báez santamaría, and piek vossen this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). baier, báez santamaría and vossen roles, and is related to the shared environment and space within which the interaction occurs. recent approaches to conversational ai and robots are based on encoder-decoder architectures in which end-to-end systems learn to respond to encoded multimodal signals (huang et al., 2020; miyazawa and nagai, 2023). fu et al. (2022) mention two main issues with such encoder-decoder systems: they lack interpretability and control, resulting in unwanted behavior (hallucination, biases) that is difficult to repair. two things are essential to study interactions by different agents and processing components: 1) we need to be able to record the interaction as multimodal data so that we analyze, evaluate, compare and re-use interaction data, and 2) we need to be able to freely experiment with different components to test their impact on the interaction. this is specifically challenging because multimodal signals, such as images, sound, speech, faces and their expressions and gestures, might not be aligned, are noisy, and can be interpreted in many different ways, partially dependent on each other (miyazawa and nagai, 2023). whereas other platforms such as psi bohus et al. (2017) and ros macenski et al. (2022), focus on handling streams of signals through open configurations of processing components, we augment such a platform with higher levels of abstraction and reasoning over these multimodal streams. to make sense of this multitude of signals and possible interpretations, a critical component is an episodic long-term memory in which interpretations are stored over time and which forms the basis for reasoning over and responding to new input. such a sense-making module is essential to create transparent, trustful, and functional agents. previously, vossen et al. (2019); báez santamaría et al. (2021) introduced a) emissor to store sequences of multimodal signals with their annotations and b) episodic knowledge graphs (ekg) to store the accumulation of interpretations as symbolic triples. emissor registers interactions as scenarios in which multimodal signals align using a temporal ruler to mark signals’ start and endpoints. segments within each signal are automatically annotated with interpretations, such as detected objects or faces through cameras, or entities and properties mentioned in speech. for example, recorded images in time can be annotated as a human face that is next identified as "carl", who is the speaker of an audio signal (human voice), annotated as a text signal "i am from amsterdam", which is further processed for its content. a visualization of an interaction represented in emissor is included in the appendix a. the ekg represents the accumulation of interpretations of signals over time as rdf1 triples consisting of subject, predicate, and object identifiers (internationalized resource identifiers or iris), e.g., the speaker carl claiming: [leolaniworld:carl, n2mu:live-in, leolaniworld:amsterdam]. such triples combine into more complex and rich knowledge structures and can easily be connected with other external knowledge such as dbpedia2 specifying what is amsterdam through another triple. by interpreting conversations and perceptions as triples, the agent not only records signals but also learns about people and the world over time through interaction. by representing interpretations as claims of sources as defined in the grasp model van son et al. (2016); fokkens et al. (2017); vossen and fokkens (2022), the ekg forms a so-called theory-of-mind gopnik and wellman (1992) by accumulating perspectives on claims in which possibly conflicting information holds relative to the sources of these claims. as a result, the ekg can deal with conflicting claims when sources disagree and perspectives on claims that change over time. emissor and the ekg together do not yet make an agent. an agent needs to pick up signals and call the necessary modules to process these before the interpretations can be rendered to emissor 1. resource description framework, https://www.w3.org/rdf/ 2. https://www.dbpedia.org 26 https://www.w3.org/rdf/ https://www.dbpedia.org a modular architecture for creating multimodal embodied agents or the ekg. furthermore, an agent must respond to this input by generating output signals and actions. this paper describes a platform for creating different interactive agents whose interactions can be registered in both emissor and ekg and can integrate any component to generate responses. the core of the platform is an event bus, which can be considered a multi topic input/output queue that connects time-based signals with interpretation components as separate software modules. component apis define the expected input and output event types, and the topics that a component is connected to -such as "audio-in" or "text-out"can be defined freely and specified in the configuration of an agent. components can be replaced against competing implementations as long as their apis match. likewise, different agents can be created easily by assembling components and configuring their connection to the event bus. the primary signals originate from back-end components that read data from sensors at time points and publish them as events to a particular topic in the event bus. different interpretation components consume these signal events (defined by their input topic) and publish the result of processing the input event(s) to their output topic(s) in the event bus as new events. these components can annotate segments from signals with interpretations, such as human speech or text, objects, and people perceived, ultimately pushing these interpretations to the ekg. any agent is therefore triggered by the incoming signals and their processing by the integrated components. we call such an agent a signal-driven agent. although event-driven architectures are not new as such, our event bus platform differentiates itself by adding higher-level components that try to steer and control the signals by intervention. this can be done by 1) the so-called intention modules that define the goals as dialog states to be satisfied before moving on and 2) by pushing high-level interpretations of these signals to the ekg for reasoning within a theory of mind (tom). the reasoning in the ekg results in knowledge states that are assigned to actions of the agents to improve the knowledge states. whereas intents act as dialog states similar to a dialog management component, the ekg more openly and proactively drives the interaction. adding control mechanisms turns a signal-driven agent into a control-driven agent. because of these control options, it is also more easy to incorporate (third-party) encoder and decoder modules into our platform to interpret signals or verbalize responses without suffering too much from hallucination and generalized biases. different variants of agents will still render compatible multimodal data from their interactions within our framework as we store their interpretations and responses in emissor and the ekg. therefore, our platform can be used effectively to create and compare agents by analyzing the data rendered, both in emissor and as an ekg (báez santamaría et al., 2022). although we published before about emissor and the ekg, we never described how these are combined with an event-driven run-time system, which is the topic of this paper. this paper explains the general architecture and implementation of the leolani platform, as opposed to the leolani agent, which is one of the agents that could be built on the platform.3 we also explain how different interactive agents can be created within the same platform using various components and demonstrate how we can add control to a signal-driven agent. currently, we implemented a 3. originally, leolani was developed as a specific agent application to get to know you. in recent years, leolani has evolved into a platform that can be used to create many different agents. we now use leolani for the platform and leolani agent for the implementation that uses the widest range of components currently available. see our github repository for more details and instructions: https://github.com/leolani/leolani-mmai-parent 27 https://github.com/leolani/leolani-mmai-parent baier, báez santamaría and vossen wide range of agents with no controlling options such as eliza agent4, blenderbot (shuster et al., 2022), or llama touvron et al. (2023), and with maximized control over the interaction such as the multimodal leolani agent (vossen et al., 2019), agents that play an "i spy" game or spotter, which is a multimodal game to analyze referential expressions kruijt et al. (2024). the paper is further structured as follows. we first position our platform on related work in section 2. in section 3, we explain the overall architecture of the platform and its connection to emissor and the ekg. section 4 describes the components that are currently available and section 5 describes the functions that are available to evaluate interactions. we report on the methods for analyzing and comparing interactions in section 5. section 6 explains how you can create your own agent either by combining existing components or by defining your own component. finally, we present a technical specification on processing requirements, scalability, runtime latency, and real-time processing in sections 5 and 6. all our code is available on github under the apache2.0 open source license from: https://github.com/leolani. 2. related work recent advances in deep learning resulted in unprecedented performance in various technologies, also necessary for building intelligent interactive agents, ranging from language understanding, speech recognition, conversational skills, vision, to audio processing. furthermore, fusion models combine multimodal input that is otherwise processed independently into a unified framework (gao et al., 2020). the latest versions of openai’s chatgpt (wu et al., 2023) and google gemini (team et al., 2023) combine these capabilities with conversational skills in longer contexts. despite their impressive capabilities, and in addition to well-known problems such as hallucination and biases imposing unwanted interpretations and responses, these models are, however, still far from dealing with the complexity of physical real-world situations in which robots need to perform. miyazawa and nagai (2023) give an overview of how deep learning models are integrated in robot applications and research. they show on the one hand that state-of-the-art research in robotics relies more and more on leveraging these models but also that interacting with real-word settings or virtual simulations of these settings still requires additional modeling and integration for controllable behavior of agents. miyazawa and nagai (2023) limit their survey to what we call action robots, which are agents that need to understand instructions and communicate about their task success with human instructors when dealing with some world. this greatly limits communication and social complexity, i.e. in most simulated worlds virtual agents do not encounter avatars representing real people or users. in the case of social agents and specifically social robots whose embodied realization needs to interact with people in the real world, situations are extremely complex and it is also a challenge to design agents that behave appropriately. the complexity of designing such agents is clearly demonstrated by the complex dialog flow of the agent designed by stange et al. (2022) whose purpose is to generate verbal explanations for decisions in relation to its intentions. besides a wide range of different components for understanding, perception, memory, planning and acting (both physically and verbally), their architecture also shows complex dependencies and decision points connecting these components through classical ai components such as a dialog manager, decision engine and strategy manager. remarkably, their design does not consider the integration of deep learning technology in the architecture, which would add another layer of nontransparency to the decisions made by the agent due to their lack of predictability as black-box models. 4. https://github.com/leolani/eliza-parent 28 https://github.com/leolani https://github.com/leolani/eliza-parent a modular architecture for creating multimodal embodied agents in order to be able to design and build agents with embodiment that can act physically and socially in real-world settings, a platform is needed in which researchers and developers can freely experiment with integrating any component for interpreting multimodal signals, building memories, making decisions in relation to intentions, and rendering physical and social actions. microsoft created a platform for situated intelligence, psi5 for this purpose (bohus et al., 2017). psi offers multimodal data visualization and annotation tools, as well as processing components for various sensors, processing technologies, and platforms for multimodal interaction. psi models multimodal situations and interactions within and comes close to a comprehensive solution. psi is a software integration platform through which developers can share modules using a streaming architecture for signal annotation. however, interactions are not stored in a shared representation, and the platform cannot be used to share experimental data independently of the platform itself. in the context of (spoken) dialog systems, a generic model for incremental processing was proposed by schlangen and skantze (2011). retico (michael, 2020) is a python implementation of this incremental model as a framework for creating systems and simulations of spoken dialog. it provides a range of incremental modules based on services like google asr, google tts and rasa nlu. although generic and open with respect to the integration of processing modules, these approaches model incrementality as a sequence of units that trigger each other. they do not consider higher-order modules that "oversee" such sequences and define patterns for the purpose of achieving intentions as we do in our platform enabling cross-platform comparison. retico focuses on spoken dialogs in which audio signals form a natural sequence. if other modalities are considered, such as image or video, temporal alignment of signals is necessary as provided in psi. kennington et al. (2017) therefore combine retico and psi for spoken dialogue systems in visual contexts. in their approach, sequential incremental interpretation by different modules is provided by a mixture of modules in both frameworks. whereas kennington et al. (2017) combine two different frameworks, psi implemented in c# for windows platforms and retico in python for linux, we provide an integrated framework that can run on any platform while using a distributed server architecture. furthermore, their approach requires a dialog management component that needs to take decisions at many modules-interface points in the interaction, which puts a heavy burden on the design of an agent. in our approach, we allow for both low-level and high-level control over multimodel signal interpretation, both in terms of agent intent reasoning as well as in terms of an episodic knowledge graph that creates coherence and continuity across encounters. controlling the flow of interaction can be defined easily in many different ways without changing the code and interaction between different processing components. an older comprehensive robot platform worth mentioning here is openease6, which is a webbased knowledge service between a robot and a human that comprises episodic memories (beetz et al., 2015). it produces semantically annotated data of manipulation actions, including the agent’s environment, the objects it manipulates, the task it performs, and the behavior it generates. it is provided with a query language and inference tools that allow reasoning about the data and answering queries regarding what they did, why, how, what happened, and what they saw. ease uses the so-called neems (narrative enabled episodic memories) as episodic memories. neems consists of a video recording by the agent of the ongoing activity. these videos are enriched with stories about actions, motion, their purposes, effects, and the agent’s sensor information during the activity. knowledge is represented through prolog predicates as annotations of sequences. ease is not 5. https://github.com/microsoft/psi 6. https://www.open-ease.org 29 https://github.com/microsoft/psi https://www.open-ease.org baier, báez santamaría and vossen data-centric, but a service platform that uses a knowledge database as a back-end. the database can be explored through prolog queries. the focus of openease is on physical interactions and not on conversations with complex referential relations between expressions, situations, and episodic knowledge of past experiences. recently, open distributed platforms for the creation of multimodal agents have been proposed by kennington et al. (2020) and chiba et al. (2024). these platforms allow for the incorporation of new (multimodal) modules for signal processing that can also be distributed, making the architecture flexible and expandable. these approaches come close to our approach but lack the incorporation of an episodic knowledge graph that provides control as in our system. none of the above platforms provides an easy way to define control over the incremental signal processing nor do they allow for easy integration of state-of-the-art deep learning technology. our approach distinguishes itself by pairing open integration and processing of multimodal streams of signals in a single framework with options for adding higher-order intents for dialog management and episodic memories for incremental reasoning, long-term coherence, and continuity. by modular integration of such controlling modules, our platform makes it easy to experiment with large language models and fusion models without giving up control. similar to retrieval augmented generative approaches, the ekg and the intents can be used to drive the interpretation and also the generation of signals by these models. 3. event bus architecture the architecture of our platform is built around two central concepts, one being emissor as a shared framework to represent data in a multimodal interaction and the other the communication between the modules through an event bus using a publish/subscribe model. this provides a high level of decoupling between modules in our system, as they only need to be concerned about their individual input and output data, ensuring it is in the right format and not about the context they are used in. this, on the one hand, allows us to use them as flexible building blocks for applications. on the other hand, integration of external software components into our platform can be achieved by adding a thin integration layer (or wrapper) that only needs to take care of converting input and output data to the format and implement their connection to the event bus. as an example, an object detection module that is integrated in the platform would take raw image data from a received event with an emissor imagesignal, perform object detection on these raw data, and publish an event containing the detected objects as an emissor annotation on the original imagesignal, making them available for further processing to other modules. depending on the used event bus implementation, this can even connect different run-time systems, allowing one to connect components across different platforms. figure 1 shows a schematic representation of the event bus architecture. the central component is the event bus itself, to which signals and interpretations are pushed either by back-end components (on the left side) or interpretation components (on the right side). the event bus is represented as a vertical (gray) bus to emphasize that it has a temporal dimension where events are processed in a sequence. the labels on the arrows between the event bus and the components denote the type of event transmitted. the vertical lines in the event bus symbolize different topics used by the individual components to define the event flow through the application. the topic labels to which components are connected are denoted in italics next to the arrows. by defining the input topic and the output topic for a module, we can condition which components become active for the event bus (input topic 30 a modular architecture for creating multimodal embodied agents specification) and what the potential components are that follow their output topic. this allows us to define the possible interactions from the bottom up. figure 1: schematic representation of the event bus architecture showing backend components to read from sensors and fill the queue in the event bus and interpretation components for annotating published signals. in this schema, only audio signals are processed with an eliza module for generating a response. interaction data is recorded from events by the emissor component. examples of back-end components are microphones, audio controllers, and cameras. sensors connected to back-end components can be directly linked to the system or obtained from remote servers. this schema shows how a microphone produces an audio signal with a start and a stop point. some examples of interpretation components are shown on the right side of the event bus. the vad component at the top applies voice-activity detection to an audio signal published by the microphone, which reports back to the event bus whether the audio is human speech, annotating the corresponding segment in the audio. the speech-to-text (stt) component applies speech recognition to any audio signal annotated as human speech and publishes the transcript as a text signal to the event bus. in this example, the text signal is picked up by an implementation of the eliza chatbot (weizenbaum, 1966), which checks it for triggers and returns a response as another text signal. instead of eliza, any other module can be used to generate a response either directly to the text signal or indirectly through interpretations of these signals. finally, the text-to-speech (tts) component picks up the 31 baier, báez santamaría and vossen text response and converts it to an audio signal. the back-end speaker listens to the event bus for any audio signal to be output. an audio controller between the microphone and the speaker mutes the mic when speaking and vice versa. any data produced by the sensors or the components as signals is recorded in an emissor data structure. 3.1 enabling reasoning the main application can function with just the event bus and some simple components, taking in signals and generating signals as a response. however, the behavior of such a signal-driven agent is, however, primarily reactive and not reflective of the information conveyed during an interaction. examples of such agents are eliza which responds to trigger words or generative language models such as blenderbot, dialoggpt or llama that generate the most-likely continuation of the dialog through implicit reasoning. to turn such an agent into a control-driven agent, we need to interpret the signals at a more abstract symbolic level. this can be done in the form of annotation of (segments of) signals or by the explicit knowledge conveyed and the implications derived so that the agent can reason over it. the reasoning results in a decision to create a response. symbolic signal annotations can directly be used by higher-level intentions that check for such occurrences to move from one dialog state to the next; e.g., obtaining your name means we can now chat. in the case of our ekg the response is based on reasoning over the knowledge quality in the ekg. figure 2 shows the interconnections across the three components: event bus, emissor, and ekg. the three components are aligned so that each interaction in the event bus generates a corresponding multimodal signal representation in emissor and, if the interpretation yields knowledge, also an rdf triple in the ekg. emissor stores interactions as scenarios with metadata in json files for signals in each modality: images, audio, text, and rdf-triples; further details are described in báez santamaría et al. (2021). the metadata grounds the signals to a temporal ruler and defines any annotation of segments within a signal as interpretations. the raw signals themselves are stored as separate files on disk.7 in the ekg model, interactions are initialized as instances of situational eps:contexts grounded in time and space. contexts correspond to scenarios in emissor. within a context, a series of sem:events are represented, which are grounded in time like the signals in emissor. we distinguish between conversation and perception events: conversation: conversation events in the ekg are structured as grasp:chat instances containing grasp:utterances in which participants in the interaction mention grasp:statements or formulate questions. the information mentioned in grasp:statements is stored in a gaf:claim, which is a named graph attributed to a speaker (grasp:wasattributedto) and that hold certain perspectives (grasp:attribution). these perspectives are properties such as grasp:certaintyvalue, grasp:polarityvalue, grasp:sentimentvalue, grasp:emotionvalue. currently, we included modules that extract such perspective values from the text generated by speech-recognition as well as modules that extract these values from facial expressions. perceptions: in contrast, perception events in the ekg are structured as grasp:visual instances containing grasp:detectionss which could be objects or people. these detections are considered to be instances of grasp:experiences. similarly to conversation events, the infor7. we also consider the rdf triples extracted from text as a modality stored on disk in emissor. 32 a modular architecture for creating multimodal embodied agents mation obtained from these grasp:experiences is stored in named graphs as gaf:claims with triples for the labeled output of object/face-recognition components. in this case, the information is grasp:wasattributedto sensors, and the grasp:attributions are restricted to grasp:certaintyvalues only (e.g. based on confidence scores). perception events therefore reflect the awareness of an agent of objects and people in the context. perceptions correspond to image signals in emissor. figure 2: schematic overview of relations across data structures in emissor, the episodic knowledge graph and the event bus. as shown in figure 2, the event bus connects the specific components that make up an agent and at the same time displays the data for emissor and for the ekg. these components can also push data, such as annotations or triples, to the event bus itself so that other components can take these as input for further processing. the event bus is thus populated not only by the sensors but also by the interpretation components. in the next section, we give an overview of the main components in our platform and explain in more detail how they interact (for a complete overview, we refer to our github repository). in the next section 5, we explain how interactions can be analyzed, compared, and evaluated, and, in section 6, we explain how new components can be added and how you can create your own agent. 4. components available interaction task is broken down into a large number of smaller tasks that operate on specific input events taken from the event bus and push their output back onto the event bus. so far, we have developed the following components either by creating a wrapper for third-party systems or by developing components in-house. we grouped the components by their underlying input signal: 33 baier, báez santamaría and vossen 1. audio signal voice activity detection 8 classifying an audio signal as speech, output is an annotated audio signal pushed to the event bus and saved to emissor as an audio annotation. automatic speech recognition 9 transforming speech to text, various models can be chosen which generate a text output signal to the event bus and emissor, among them whisper (radford et al., 2023). speaker identification (under development) identifying the speaker based on a sample of their voice, output is a speaker’s identity (as an iri and their name), and an audio signal annotated with the source pushed to the event bus and emissor but also registered in the ekg as an encounter. 2. text signal mention detection 10 detecting names and potential objects in a text signal using spacy11, output is an annotated signal with named entities and object mentions, pushed to the event bus and emissor. mention identification 12 identifying instances in the knowledge graph for mention annotations in a text signal, output is an identifier (iri), pushed to the event bus and emissor. corresponding rdf triples (iri gaf:denotedby gaf:utterance#offset) are pushed to the ekg. triple extraction 13 detecting rdf triples in a text signal; various implementations can be selected; output are rdf triples pushed to the event bus, emissor and the ekg. several in-house systems have been developed (among which a conversational extractor that considers sequences of turns) but also stanford’s open ie14. query extraction 15 detecting a question in a text signal and converting this into a sparql query; output is a sparql query sent to the ekg and possibly other knowledge graphs. emotion detection 16 detecting basic emotions in a text signal; output is an emotion label17 pushed to the event bus and emissor as an annotation and possibly as a perspective to the ekg. dialog act detection 18 detecting the dialog act for each utterance through two different modules, outputting one of many different conversational dialog acts to the event bus and emissor (yu and yu, 2019; chapuis et al., 2020). 8. https://github.com/leolani/cltl-vad 9. https://github.com/leolani/cltl-asr 10. https://github.com/leolani/cltl-knowledgeextraction 11. https://spacy.io 12. https://github.com/leolani/cltl-knowledgelinking 13. https://github.com/leolani/cltl-knowledgeextraction 14. https://nlp.stanford.edu/software/openie.html 15. https://github.com/leolani/cltl-questionprocessor 16. https://github.com/leolani/cltl-emotionrecognition 17. bert trained with the go annotations (https://huggingface.co/bhadresh-savani/ bert-base-go-emotion) and roberta find-tuned with meld and wassa data kim and vossen (2021) 18. https://github.com/leolani/cltl-dialogueclassification 34 https://github.com/leolani/cltl-vad https://github.com/leolani/cltl-asr https://github.com/leolani/cltl-knowledgeextraction https://spacy.io https://github.com/leolani/cltl-knowledgelinking https://github.com/leolani/cltl-knowledgeextraction https://nlp.stanford.edu/software/openie.html https://github.com/leolani/cltl-questionprocessor https://github.com/leolani/cltl-emotionrecognition https://huggingface.co/bhadresh-savani/bert-base-go-emotion https://huggingface.co/bhadresh-savani/bert-base-go-emotion https://github.com/leolani/cltl-dialogueclassification a modular architecture for creating multimodal embodied agents 3. rdf signal knowledge representation for claims 19 posting rdf claims to the ekg; output is jsonld as responses to the changes in the ekg by digesting new claims pushed to the event bus. knowledge representation for queries 20 posting sparql queries to the ekg or other graphs; output are rdf triples as the result of the query pushed to the event bus. response generation 21 verbalising qualitative reflections on changes in the ekg (responses); output is a natural language response as a text signal, pushed to the event bus and stored in emissor. language generation 22 verbalising sparql query results; output is natural language as a text signal, pushed to the event bus and stored in emissor. 4. image signal object recognition 23 detecting multiple objects in images using yolo24; output is bounding boxes and object labels as an annotated image signal, pushed to the event bus and stored in emissor. face detection 25 detecting human faces in an image, the output are bounding boxes with the estimated age and gender as an annotated image, pushed to the event bus and stored in emissor. face identification 26 identifying people from their face, the output is an identifier (iri and a label as name), pushed to the event bus and stored in emissor. encounters are also registered in the ekg as perception events. face emotion detection the same component that does text emotion detection also includes a module "emotic" kosti et al. (2019) for contextual face emotion detection. 4.1 processing audio signals the components that process the audio signal were discussed in detail in section 3 when explaining figure 1. in a nutshell, the voice-activity-detection component detects portions of human speech in audio, which the automatic-speech-recognition transcribes. eventually, these components push a text signal to the event bus, which is rendered from an audio signal for text processing components. 4.2 processing text and rdf signals the components that operate on text signals are more complex. they generate different interpretations and types of outputs and combine with components that operate on rdf triple signals. 19. https://github.com/leolani/cltl-knowledgerepresentation 20. https://github.com/leolani/cltl-knowledgerepresentation 21. https://github.com/leolani/cltl-languagegeneration 22. https://github.com/leolani/cltl-languagegeneration 23. https://github.com/leolani/cltl-object-recognition 24. https://github.com/ultralytics/yolov5 25. https://github.com/tae898/age-gender 26. https://github.com/leolani/cltl-face-recognition 35 https://github.com/leolani/cltl-knowledgerepresentation https://github.com/leolani/cltl-knowledgerepresentation https://github.com/leolani/cltl-languagegeneration https://github.com/leolani/cltl-languagegeneration https://github.com/leolani/cltl-object-recognition https://github.com/ultralytics/yolov5 https://github.com/tae898/age-gender https://github.com/leolani/cltl-face-recognition baier, báez santamaría and vossen various natural language processing (nlp) modules are called within different components, among which spacy(vasiliev, 2020), nltk, stanfordnlp (angeli et al., 2015) and various transformer models. in addition to their basic processing, the triple extraction components output various annotations, for example, mentions of people and objects for which referring expressions must be resolved to known people and objects. referring expressions can be names (carl), common noun phrases (the waiter), or ambiguous pronouns. by reasoning about the context and previous encounters, mention-identification components aim to establish the referent of these expressions. resolved expressions are converted to iris incorporated in triples. these can be registrations of mentions, e.g., carl talking about carla, or as part of triples expressing a property of an individual: [leolaniworld:carla, leolaniworld:live-in, leolaniworld:amsterdam] . pronouns such as "i" and "you" are resolved to the speaker and the addressee. whenever a gaf:claim is posted to the ekg, the knowledge-representation-for-claims component generates a reflective response to the claim. this involves running various precoded sparql queries to detect knowledge gaps, conflicts, uncertainties, novelties, analogies, and possible generalizations. the response-generation component selects a result to formulate a response to try to improve the ekg, for example, to fill gaps or resolve conflicts and uncertainty. the triples are verbalized as natural language text in a text signal. depending on whether the ekg responder is chosen within the framework, the response is converted into a new text signal attributed to the agent. the previous route through the event bus applies to signals classified as statements. if a text signal in the event bus is classified as a question, the query-extraction component extracts a sparql query from the text signal that the knowledge-representation-for-queries component posts to the ekg (or any other triple store). the language-generation component takes the query’s result (one, many, or no triples) to verbalize it as a text signal. different answers are generated according to the type and number of query results. furthermore, answers consider the status of the knowledge, e.g. how certain, who is the source, and when was it mentioned. in addition to querying the ekg directly, we also implemented alternative components that can query a database of responses (questions about the agent), external resources (consulting google search, wikipedia, news, the weather or wolfram alpha), or visual information that was perceived by the agent. in addition to processing statements and queries, some components only annotate mentions or emotions expressed in the text (without necessarily being involved in a triple). the mentiondetection component detects things in the text that (can) exist in the physical world, which are object that the object recognition can detect. furthermore, it detects named-entity expressions in the text. these mentions are saved as annotations of the text signal in emissor and pushed to the event bus for further processing. next, the mention-identification component picks up entity mentions and tries to resolve their identity given the entities registered in the ekg, e.g. by checking names, the contexts, or resolving pronouns’ co-reference. identities are pushed to the event bus and registered as gaf:mentions in the ekg. possibly, a new identity can be established and a new iri is stored in the ekg with the properties that can be inferred. the emotion-detection component interprets a text signal and pushes an emotion to the event bus as well as its annotation to emissor. the emotions expressed by the interlocutor represent potential perspectives of the source on the current situation or the information expressed. 36 a modular architecture for creating multimodal embodied agents 4.3 processing image signals finally, the object recognition and face-detection components take image signals as input to detect objects and faces with bounding boxes. these are saved as annotations of image signals in emissor and pushed to the event bus. from the event bus, the face identification component establishes the identity of people by comparing it with the faces of known people. unknown faces are saved as new identities with inferred properties, such as sex and age. a new identity is pushed to the event bus after which it is picked up by a component that asks for a person’s name. the image-signal-processing components either produce annotations pushed to the event bus and/or encoded in emissor or produce an iri with properties: name, age, gender, and registered perception event in the ekg. 4.4 sample input-output signal sequences to illustrate how sensor signals are processed, we show some schematic representations of typical sequences of input and output events with their payload type and input/output topics denoted above the arrows27. 1. audio (signal) mic−−→ audio (annotation) vad−−→ speech (annotation) asr−→ text (signal) text-in−−−→ mention (annotation) entities−−−−→ iri (annotation) linking−−−−→ claim triple (ekg) triples−−−→ response triples (ekg) response−−−−→ text (signal) text-out−−−−→ audio (signal) 2. audio (signal) mic−−→ audio (annotation) vad−−→ speech (annotation) asr−→ text (signal) text-in−−−→ mention (annotation) entities−−−−→ iri (annotation) linking−−−−→ query triple (ekg) triples−−−→ response triples (ekg) response−−−−→ text (signal) text-out−−−−→ audio (signal) 3. audio (signal) mic−−→ audio (annotation) vad−−→ speech (annotation) asr−→ text (signal) text-in−−−→ response (agent response database) response−−−−→ text (signal) text-out−−−−→ audio (signal) 4. audio (signal) mic−−→ audio (annotation) vad−−→ speech (annotation) asr−→ text (signal) text-in−−−→ response (eliza) response−−−−→ text (signal) text-out−−−−→ audio (signal) 5. audio (signal) mic−−→ audio (annotation) vad−−→ speech (annotation) asr−→ text (signal) text-in−−−→ response (llama) response−−−−→ text (signal) text-out−−−−→ audio (signal) 6. audio (signal) mic−−→ audio (annotation) vad−−→ speech (annotation) asr−→ text (signal) text-in−−−→ response (wolfram alpha) response−−−−→ text (signal) text-out−−−−→ audio (signal) 7. audio (signal) mic−−→ speaker (annotation) speaker−−−−→ iri (annotation) identity−−−−→ perception (ekg) 27. topic names are exemplary and can be configured in the application 37 baier, báez santamaría and vossen 8. image (signal) cam−−→ face (annotation) face−−→ iri (annotation) identity−−−−→ perception (ekg) 9. image (signal) cam−−→ object (annotation) object−−−→ iri (annotation) identity−−−−→ perception (ekg) 4.5 responding components that respond to processed signals in the event bus play an important role in defining the functionality of the agent. it is possible to integrate existing agents as response components, as long as you can define a wrapper that converts events from the event bus into the necessary input and the output of the agent into a signal that can be rendered by the embodiment. currently implemented examples are eliza, blenderbot, and llama that only require text signals as input and create a text signal as response. others could be task-based, such as recommender systems, q&a (wolfram alpha) or e-commerce services. a more advanced responder is the ekg itself, which has been extended with modules that assess each change to the graph as a result of the interaction and generate a possible response as a result of this assessment. since the graph is updated for a large variety of interpretations coming from different modalities, the agent is also sensitive to many different aspects of the interaction. we defined a wide range of graph updates to which the agent could respond. typical examples are knowledge gaps in case the agent hears about things it does not know yet, uncertainty that is either expressed by sources or inferred from the data, conflicts when sources claim different things, mentions and perceptions of relevant things, analogies, novelty relevant for certain people, and many more. assessments are called thoughts and are represented as the output of specific queries on the graph. given the type of thought, we define different drives to respond to each, which is formulated as a text response given the output of the query as input báez santamaría et al. (2021). typically, a single turn can thus produce a large number of thoughts, and specific strategies can be deployed to select the most effective thought, e.g., using reinforcement learning. what is effective also depends on the knowledge exchange during the interaction and hence on the user as a source. the ekg-based responder can be adapted and extended by defining additional queries, specifying selections of thoughts to respond, or by reinforcement learning. in general, a wide range of thoughts and drives results in surprising and spontaneous behavior, whereas a small range results in systematic and rigid behavior that could be focused on specific tasks only. 4.6 crossmodal awareness in addition to multimodal signal processing, the leolani platform provides various ways for crossmodality processing: alignment of segments segments of signals in different modalities are defined using the same temporal ruler and likewise can be related temporarily: disjoint, adjacent, overlapping, or inclusive. these relations define contexts for cross-modal interpretations. sharing interpretation labels to the extent that mentions and perceptions share their annotation labels, e.g. person as a perceived image and as a named-entity category in text, this can create further coherence relations on top of coherence through segment alignments. sharing of identification labels components that produce identities for different modality signals can share these identities through the ekg. for example, face recognition is associated with an 38 a modular architecture for creating multimodal embodied agents ekg identifier and a name as a label, while the same name mentioned in text can be mapped to the same identifier in the ekg, resulting in all interpretations being aggregated across modalities to the same identity. consider the following example to demonstrate the cross-modal alignment. the agent is communicating with a person identified as leolaniworld: carl. during this interaction, the agent perceives an image of another person that is unknown. close in time, preceding, overlapping, or following this image signal, the agent receives an audio signal from carl. processing the audio signal yields the text "that is my sister carla". the object recognition produces the label person for the image and the named entity recognition detects a person entity in the text. further text processing will pick up the deictic reference "that" and the identity expression "is my sister carla". if no person exists in the ekg with the name carla, a new identity leolaniworld:carla is created with the relation leolaniworld:sister-of to leolaniworld:carl. the face identification module is then activated to create a visual representation that identifies the new unknown face. in this example, the temporal segment alignment of the image and the audio/text signal, as well as the shared interpretation labels, create an alignment across the modalities that an agent could use to combine all three cross-modal interpretations to yield one coherent result. other scenarios can be easily modeled, e.g. if carla was mentioned long before she was perceived or the other way around. if other people with different visual properties are named carla, this could trigger a disambiguation intention and initiate communication by the knowledge linking module, depending on how similar or different the visual information is and or the signals are disjoint. besides the cross-modal interpretation of signals, we are working on components to aggregate these interpretations into overall situations, called contexts, that assume certain degrees of coherence and consistency. in every interaction, a new instance of a context is created in which time and space are defined vossen et al. (2019). these contexts exist for the duration of an interaction and provide the basis for cross-modal consistency and permanence. within a context, people and objects are perceived and mentioned within the same location as in a space. the location will be identified using the people and objects within it, in addition to a visual spatial map, but the objects, and potentially also the people, can also be identified given the location. locations and objects/people are in a dualistic relation. finally, the same location can have multiple interactions with the same or different people. each interaction represents a different scenario following a specific conversational flow and logic, while they can take place in the same physical location, with some objects remaining the same while others change. context overarching data structure for defining situational awareness. location and object identity locations are identified across encounters using visual maps that include the identities of objects and people within and vice versa. scenarios multiple interactions can take place within the matching contexts, each defining a unique scenario but possibly sharing the same location and objects within. the identities of locations and objects within contexts are strongly intertwined vossen et al. (2019). different locations may look very similar, and the same holds for different objects. across contexts, identifying the location and objects will involve matching properties as perceived in earlier encounters. in case of a strong match, identity can be assumed, in case of a moderate match, the 39 baier, báez santamaría and vossen agent may ask for confirmation, and in the case of no match above a threshold, identity needs to be established with the help of the human user through communication. a specific location in which the agent is activated could be an office space. the agent creates a new context instance in the ekg to aggregate interpretations within this context. after scanning the space and objects and people within, the agent compares this information with previous contexts in the long-term memory of the ekg and the corresponding sensor data (images) saved in emissor. these memories are based on all previous encounters and the interpretations of the sensor data, e.g. repetitive interactions taking place in the same office identified in the past. the agent tries to align the identities of the location, people, and objects in the current context with the contexts of the past. likewise, the agent may recover the identity and name of the office with various properties such as who owns it, what specific objects are still there, which ones are missing, and which ones are new. object identity is complex, as many phones, chairs, tables, laptops look very similar, and some things move easily, whereas others do not. if there is not sufficient evidence for a match, a location needs to be considered as a new space. the matching and corresponding identities have a considerable impact on the dispersion and fusion of information in the ekg. with each perception and interpretation, a choice needs to be made as to whether perceptions are different, the same, or vary depending on how permanence is defined across locations. the platform allows for different approaches to establish location and object identification in relation to the ekg in terms of different granularities that can vary on a spectrum from all different to all the same. in future work, we will develop components that find an optimal balance between these extremes. finally, it should be noted that the agent, at any point, could communicate uncertainty about identities to people to obtain confirmation. contexts, signals segments, and their annotations are represented in emissor data structures, but also in the ekg. multimodal agents can be designed to use either representation to arrive at specific cross-modal interpretations or ignore it. however, the more specific and aligned the interpretations are, the more dense and rich the information is for an agent to respond. 4.7 higher level intentions the above components become active when the input topic requirements are met, making a signaldriven agent that always responds to the presence of input topics in the event bus. however, our event bus architecture allows one to define higher-level intentions that represent the behavior of the agent to fulfill an end task or goal, that is, a control-driven agent. these intentions can be prioritized and remain active until resolved. for each intention, we can specify which components are actively listening to the event bus and which follow-up intentions are published once a task is completed or a certain goal is reached. the above components are thus activated only within these intentions. intentions provide dialog management control over the interaction. the spotter agent (kruijt et al., 2024) is an example in which the dialog flow for playing a reference game is fully controlled by intents that define dialog states. in this game, references to people need to be resolved to position them correctly in a row between a human and an agent. in appendix 6, we show the dialog flow diagram for the game taken from kruijt (2025) where the control is defined by higher-level intentions using conversational states and dialog management decisions. another example is given by vossen et al. (2024), in which an agent communicates with a person to learn about their daily activities to fill the knowledge gap between the previous conversation and the current, given the long-term history of the person. the agent uses the ekg to learn what is 40 a modular architecture for creating multimodal embodied agents not yet known, what is expected, and what is possible, which primarily drives the communication. the purpose of this activity monitoring is to obtain a picture of a person’s well-being over time. again, intents are used as conversational states for the degree of saturation to be reached (how many activities and how much information about each should be known) and whether there is a further need to ask questions, see the appendix 7. below is an overview of different high-level intentions on the platform. getting-to-know-you after detecting a human face in an image signal, the agent tries to identify the person to great him/her by name or to get to know a new person by asking for a name. adding this identity to the ekg triggers more questions to learn about this person. giving consent ask people permission to keep the data and share them for research. if no consent is given, the agent will remove any data (emissor and ekg) and stop the interaction. this scenario can precede the get-to-know-you scenario. leolani-agent an agent that interacts freely with users relating interpretations to identities (iri) and triples (vossen et al., 2019). the leolani-agent can be configured to follow certain thoughts and drives, such as filling knowledge gaps, and will remain proactive until these are resolved. eliza takes text signals as input and generates a text response using eliza trigger patterns and hard-coded responses. eliza has no other goals than to respond to triggers endlessly. blenderbot process text signals with blenderbot (roller et al., 2020). blenderbot is a generative model trained with different conversational data to generate a human-like response. blenderbot has no other goal than to respond to a user, just as eliza, and will not initiate communication. llama process text signals with llama (touvron et al., 2023). any type of instruction can be given to adapt llama. also, llama has no other goal than to respond to a user, just as eliza, and will not initiate communication. aboutagent process text signals by mapping questions to a database with information about the agent. aboutagent has no other goals except to answer user questions about the agent. spotter a game played over several fixed rounds during which two players (a human and a robot) need to communicate about the differences in pictures visible to each player separately. the goal is to align the images and complete the game with a high score. referring expressions are analyzed in relation to the common ground that is created between the players kruijt et al. (2024). getting-to-know-more-of-you an agent that converses with people to learn about their activities of daily life. the agent uses the episodic knowledge graph to reason about activities from the past and the period between the current and the previous conversation to fill the gap vossen et al. (2024). the goal is to monitor people’s life and well-being over time. standup-comedian an agent that interacts through scripted protocols to act as a stand-up comedian pulling jokes and riddles from llama but following the rules of the game. goodbye when the text signal contains a goodbye cue, it ends the scenario and says goodbye to the participant. the agent can start a new scenario with any person present. 41 baier, báez santamaría and vossen similarly to the above intents, it is easy to add any task-based intent to fill slots in a form, to retrieve information from an external database, or play a game. both components and higherlevel intentions can be created and combined freely. in part, this provides control over certain goals or tasks to be completed, but it also allows flexible and spontaneous behavior. in the next section, we describe several ways in which interactions can be analyzed and evaluated. we also provide details about the scalability and run-time latency of agents developed in the platform. finally, we explain how to design your own agent by adding components and high-level intentions. all the code for the platform with the currently developed components is available on github: https://github.com/leolani under apache2.0 license. a good starting point for building an agent is the repository: https://github.com/leolani/cltl-combot. 5. interaction evaluation our platform provides various ways to analyze and evaluate interactions, either through the recordings by emissor or the representation in the ekg. first, emissor captures multimodal streams of signals by storing raw signals (audio and image) and basic metadata on their temporal alignment. however, many modules from the leolani platform can also be used to annotate interactions outside the event bus framework, such as object and face detection, speech recognition, emotion detection from face or from turns, dialog act classification, entity mentions in text, and triples expressed through turns. what is needed is to represent the sequence of signals in emissor. when annotated, we can analyze the sequence of signals in terms of its annotations. if the triples are also added to an ekg, we can measure the incremental growth of the graph during the interaction. báez santamaría et al. (2022) describe such an analysis for a diverse set of interactions by students with an implementation of leolani (human-agent), blenderbot interacting with the same leolani implementation, blenderbot interacting with eliza (agent-agent) and scripted human-human interaction between actors in the friends sitcom. the emissor and ekg representations are either created during the interactions or extracted a-posteriori, which resulted in comparable representations showing differences and similarities across the different interactions. likewise, krause et al. (2023) applied this analysis to a wide range of existing conversational datasets, ranging from task-oriented conversations such as multiwoz (budzianowski et al., 2018) to fully open conversations between people as in the personal events in dialog dataset (eisenberg and sheriff, 2020). in addition to a structural and graph-based analysis of interactions, it is possible to add reference responses to replace system responses and apply standard metrics, such as blue, rouge, meteor, and bertscore as provided in the google evaluate package.28 following mehri and eskenazi (2020), transformer models, such as bert, roberta, or their fine-tuned usr model, can be used as a reference-free evaluation method by measuring the likelihood of system and user responses given the preceding context. correlation functions are provided to compare any human turn evaluation with the graph and reference-free scores. finally, the platform provides a function to plot interaction sequences over turns mapping annotations of dialog act, likelihood, and emotions on a negative and positive scale. dialog acts such as negative answer, complaint are considered negative whereas positive answers, appreciation are positive. a similar mapping is given for the assigned go emotions (demszky et al., 2020). in the case of the turn likelihood, a threshold is set below which we consider the score as negatively expected and above which as positively expected by a model. appendix b shows two generated 28. https://huggingface.co/docs/evaluate/en/index 42 https://github.com/leolani https://github.com/leolani/cltl-combot https://huggingface.co/docs/evaluate/en/index a modular architecture for creating multimodal embodied agents example plots: the first plot shows the interaction sequence with a pepper robot and the second plot with an agent implemented on the ai2thor platform (kolve et al., 2017) to find objects. to measure the capacity to maintain long-term memories and relationships, we have started a series of interactions over the past two years with students in one of our master courses. the students are asked to act as one of 12 fake personas, which were given some basic properties such as age, gender, profession, home town, and relationships. the students were free to talk about themselves as the persona and ask questions. similarly, our leolani agent could ask questions and give comments. the students interacted several times as the ekg grew and the agent became more responsive. in the second year, other students continued the conversation for the same persona. in table 1, the overall statistics give an indication of the total volume of signals and triples the agent processed. in total, the students had 79 interactions with the chat version of the agent and 13 interactions with the same agent running with a pepper robot. the latter interactions were performed using speech and include visual recognition of objects (1,428 perceived images) and identities of people (1,738 face images). in total, the students interacted for more than 146 hours with the agent. the chat version of the agent was distributed as a docker image and made available as a remote server so that the students could interact independently.29 the left-hand side of the table shows the number of utterances under both conditions, as well as the average token length, tokens per utterance, and utterance length in characters. as can be observed, these averages are stable in both conditions.30 the right-hand side of the table shows the volume of content that is processed in terms of detection of emotions and dialog acts, visually detectable objects mentioned and people referred to in the conversations, people identifications by their face and the number of different faces identified (171, students, staff and bystanders during multiple encounters), geopolitical entities (gpe) mentioned in the conversations. finally, we show the total number of triples stored in the ekg, broken down in the preloaded concepts and relations in the ontology, the instances added through the conversation, the total number of interactions (including perceptions), the claims made during these interactions, and the possible perspectives of the sources of the claims, i.e., the polarity, sentiment and certainty expressed. the question under investigation is how much information and knowledge is processed and how does that affect the performance in terms of runtime latency and response quality? there is yet little research on such long-term multi-user performance. the interactions in table 1 are the first attempts to measure such a performance impact on an agent that interacts with different users over a longer period of time. this agent also uses a wide range of processing modules across modalities in combination with an episodic knowledge graph that continuously reasons over incremental changes and knowledge growth and generates potential responses. table 2 shows the (finetuned) large language models that are loaded to run the leolani-agent with the pepper robot in a multimodal setting. we see that a considerable amount of memory is required for each functionality provided by separate fine-tuned models. however, the models are kept separate for experimentation. obviously, they can be combined in a single base model with 29. the docker-image is 8.7gb in size and can be downloaded from dockerhub: https://hub.docker.com/ repository/docker/piekvossen/leolani-text-to-ekg 30. between 2023 and 2024, the leolani-agent was extended with llama3.2.1b to rephrase the agent responses to make these more fluent. this has significantly doubled the length of the utterances: 6.78 versus 13.08 tokens per utterance and 33.61 versus 65.99 utterance length for the chat versions of 2023 and 2024 respectively. the robot interactions show a similar increase. clearly, the use of an llm such as llama makes the communication more verbose. 43 https://hub.docker.com/repository/docker/piekvossen/leolani-text-to-ekg https://hub.docker.com/repository/docker/piekvossen/leolani-text-to-ekg baier, báez santamaría and vossen students 23 images 1,428 persona 12 emotions 2,449 chat-interactions 79 dialog acts 6,123 duration in minutes 8619 object mentions 303 utterances 3,417 object perceptions 5,048 token length 4.04 person mentions 2,983 tokens per utterance 9.93 person identifications (171 identities) 1,738 utterance length 49.80 gpe mention 517 robot-interactions 13 triples 2,518,594 duration in minutes 187 ontology 72,530 utterances 837 instances 8,737 token length 3.89 interactions 14,098 tokens per utterance 9.82 claims 10,859 utterance length 48.30 perspectives 46,090 table 1: overview statistics obtained from 23 master students acting as 12 personas in 2023 and 2024, talking freely with the leolani-agent. interactions were done just with a chat interface (79 in total) or through speech with a pepper robot (13 in total). in the latter case, 1,428 images were recorded by pepper. total number of annotations and cumulation of triples are showsn in the right side of the table. module pretrained model size triple extraction 2 multilingual bert-base 1.42 gb face emotion models emotic 91 mb image detection yolo5 15.84 gb age-gender mlp-imdb-wiki 4.96 gb face-detection mlp-imdb-wiki 2.89 gb speech-recognition whisper-small 3.9 gb dialog act xlm-roberta 1.1 gb go emotion bert-base 440 mb text generation llama3.2-1b-instruct-q4_k_m.gguf 807.7 mb table 2: finetuned large language models used in the multimodal leolani-agent implementation running with a pepper embodiment multiple classification heads to reduce memory load. as explained in the following, each of these modules can be placed in separate servers as dockerized containers. 6. developing your own agent as components are self-contained and loosely coupled through the event bus, developing an agent is reduced to composing the necessary components and defining a flow of events, which is done by configuring the input and output topics of the components appropriately. 44 a modular architecture for creating multimodal embodied agents 6.1 adding, replacing and mixing components to enable compatibility between components, we also use emissor as the preferred data format of the payload carried by the events. the signal components publish emissor signals, eventually split into a start and stop event. interpretation components publish emissor mentions, defining segments and their annotations. like this, any component that can process e.g. emissor text signals or interpretations of a certain type, can be integrated into an agent by simply configuring it to listen to a topic that provides events with a text signal or the desired type of interpretation as payload. furthermore, data from components that provide emissor-based event payloads can be directly recorded in emissor by a central storage component without additional adaption. as long as component implementations process the same type of events, one implementation can simply be swapped for another. even multiple implementations of the same component can be included in the same application, either for comparison or to increase performance. in the current version of the platform, we have, for example, various components for extracting triples and for speech recognition that can be applied simultaneously to increase recall or precision through voting. the source of signals and annotations is included in all emissor data structures, which allows downstream components to determine the origin of an event. our modular architecture aims at easy integration of components that were not designed for our platform. the overhead to turn a software application into a component compatible with our platform consists of converting input and output of the application from and to the emissor format and connecting it to the event bus. this can usually be achieved by adding these steps to existing code, without the need for modification. 6.2 application types and performance by choosing different implementations of the event bus (local vs. remote), agent implementations can reach from e.g. a monolithic python application to a containerized setup, running each component as a separate process on a single machine, for instance using docker-compose31, or even running in a cluster consisting of many physical machines using e.g. kubernetes32. these different setups do not require any code adaption of the components itself, as those communicate only through the event bus interface, which remains identical. in a containerized setup, our architecture also allows mixing components from different platforms, e.g. using python next to java components, or components running different python environments in the same agent to resolve dependency conflicts, which is a common problem in software development. as long as modules can perform their tasks in parallel, as is the case e.g. for the processing of image signals parallel to audio signals or creating various interpretations (object recognition, face recognition, etc.) of the same image signal, a containerized setup also allows for scaling an application to many modules. parallelizing processing of events within a module, for instance, by running multiple instances of a module and load balancing events between them, allows for scaling a module to a high throughput of events. however, whether such parallelization is possible depends on the task and design of the module. the total processing time of an input signal in an application will be determined by individual event processing times along the sequences of modules, as described in section 4.4.the latency introduced by a remote implementation of the event bus itself33 is expected 31. https://docs.docker.com/compose 32. https://kubernetes.io 33. the local implementation comes with no latency overhead. 45 https://docs.docker.com/compose https://kubernetes.io baier, báez santamaría and vossen to be low compared to the computationally heavy modules used in our applications. widely used implementations, such as rabbitmq, report latencies in the millisecond range at throughput rates of 10k+ messages per second34. 6.3 out of the box for the components listed in section 4 we provide python implementations, each in its own repository published in the leolani github organization 35. each component can be packaged as a python package and provides a service module containing the emissor and event bus integration, which can be run from a python application. packages also provide the possibility of reuse of component code across components. some components already run in a docker container. in the future, we plan to add docker configurations to enable running each component in a docker container. our current agent implementations are realized as monolithic python applications, which initialize and run the service modules from each of the included component packages, and use a local implementation of the event bus. we provide a python implementation of the back-end server that can be run on linux/os x/windows platforms36, as well as a back-end server that can be run on pepper robots37. our agents can be used as a skeleton to build more agents, and in the future we will provide docker–compose and kubernetes setups based on component docker38 images to run the agents as fully containerized applications. in the containerized setting, we will be able to run components locally or remotely in the cloud. to collect all the code needed for an agent, we created parent repositories for each agent which contain all of the agent’s components as git submodules. we provide build tooling to package each of our python components, share the packages between components, and set up python environments with all dependencies needed to run the agent application and individual components. this process can be executed centrally from the parent repository using our build setup. following this pattern, new agents can also be created outside of our organization, mixing components under different governance. to build components, https://github.com/leolani/cltl-combot provides infrastructure libraries for event bus integration, configuration management, and resource management (locking). the cltl.combot.infra.event module provides our event bus interface with a local implementation, as well as an implementation based on the kombu library39 supporting the amqp40 protocol. this allows us to connect our framework to various available messaging servers such as, for example, rabbitmq41. furthermore, cltl.combot.infra.topic_worker.topicworker provides a utility class to listen to a configurable set of topics and process incoming events sequentially in a single thread, so that only the actual processing function, accepting a single event as input, needs to be implemented by the library user. our own components are structured to separate functionality, using the cltl namespace, from event bus integration, using the cltl_service namespace. 34. see e.g. https://www.rabbitmq.com/blog/2024/08/21/amqp-benchmarks or https://www. confluent.io/blog/kafka-fastest-messaging-system/ 35. https://github.com/leolani 36. https://github.com/leolani/cltl-backend 37. https://github.com/leolani/cltl-backend-naoqi 38. https://www.docker.com 39. https://github.com/celery/kombu 40. https://www.amqp.org 41. https://www.rabbitmq.com 46 https://github.com/leolani/cltl-combot https://www.rabbitmq.com/blog/2024/08/21/amqp-benchmarks https://www.confluent.io/blog/kafka-fastest-messaging-system/ https://www.confluent.io/blog/kafka-fastest-messaging-system/ https://github.com/leolani https://github.com/leolani/cltl-backend https://github.com/leolani/cltl-backend-naoqi https://www.docker.com https://github.com/celery/kombu https://www.amqp.org https://www.rabbitmq.com a modular architecture for creating multimodal embodied agents a simple example is presented in the appendices c, d and e. for a complete example including configuration and resource management, we provide a template component at https://github. com/leolani/cltl-template with more detailed code templates, makefiles for our build setup, and in the future a dockerfile template. 7. conclusions we described a platform for creating interactive agents which also captures the interaction at the signal level, the interpretation level and the integrated knowledge level. central to our platform is an event bus on which signals and interpretations events can be published in temporal sequence to particular topics. by defining components as services that take certain topics as input and publish results to other topics as output, we have the flexibility to engineer any combination of components to create agent behavior. our platform also supports the specification of higher-level intentions that can be carried out to fulfill a task or goal. such intentions define which low-level components should be active to achieve this. multimodal interaction is extremely complex and rich. the agents currently built through our platform are far from what is needed to deal with all aspects of human interaction. a lot of research and development is needed, not only on individual components but also on their interaction and integration. our platform will specifically help develop such complex integrated agents. in the near future, we will use the platform to create different agents that can be used in interaction experiments. the rendered multimodal data and ekgs can be analyzed and compared to assess and evaluate the agents and their components. as such, we hope that it will function as a laboratory for agent development and testing. references gabor angeli, melvin jose johnson premkumar, and christopher d manning. leveraging linguistic structure for open domain information extraction. in proceedings of the 53rd annual meeting of the association for computational linguistics and the 7th international joint conference on natural language processing (volume 1: long papers), pages 344–354, 2015. url https: //aclanthology.org/p15-1034. selene báez santamaría, thomas baier, taewoon kim, lea krause, jaap kruijt, and piek vossen. emissor: a platform for capturing multimodal interactions as episodic memories and interpretations with situated scenario-based ontological references. proceedings of the 1st workshop on multimodal semantic representations (mmsr), pages 56–77, 2021. url https://aclanthology.org/2021.mmsr-1.6. selene báez santamaría, piek vossen, and thomas baier. evaluating agent interactions through episodic knowledge graphs. proceedings of the 1st workshop on customized chat grounding persona and knowledge @ coling2022, korea, october 17, 2022, pages 15–28, 2022. url https://aclanthology.org/2022.ccgpk-1.3.pdf. michael beetz, moritz tenorth, and jan winkler. open-ease. in 2015 ieee international conference on robotics and automation (icra), pages 1983–1990. ieee, 2015. 47 https://github.com/leolani/cltl-template https://github.com/leolani/cltl-template https://aclanthology.org/p15-1034 https://aclanthology.org/p15-1034 https://aclanthology.org/2021.mmsr-1.6 https://aclanthology.org/2022.ccgpk-1.3.pdf baier, báez santamaría and vossen dan bohus, sean andrist, and mihai jalobeanu. rapid development of multimodal interactive systems: a demonstration of platform for situated intelligence. in proceedings of the 19th acm international conference on multimodal interaction, pages 493–494, 2017. paweł budzianowski, tsung-hsien wen, bo-hsiang tseng, iñigo casanueva, stefan ultes, osman ramadan, and milica gasic. multiwoz-a large-scale multi-domain wizard-of-oz dataset for taskoriented dialogue modelling. in proceedings of the 2018 conference on empirical methods in natural language processing, pages 5016–5026, 2018. emile chapuis, pierre colombo, matteo manica, matthieu labeau, and chloé clavel. hierarchical pre-training for sequence labelling in spoken dialog. arxiv preprint arxiv:2009.11152, 2020. url https://arxiv.org/abs/2009.11152. yuya chiba, koh mitsuda, akinobu lee, and ryuichiro higashinaka. the remdis toolkit: building advanced real-time multimodal dialogue systems with incremental processing and large language models. in proc. iwsds, pages 1–6, 2024. dorottya demszky, dana movshovitz-attias, jeongwoo ko, alan cowen, gaurav nemade, and sujith ravi. goemotions: a dataset of fine-grained emotions. in 58th annual meeting of the association for computational linguistics (acl), 2020. joshua eisenberg and michael sheriff. automatic extraction of personal events from dialogue. in proceedings of the first joint workshop on narrative understanding, storylines, and events, pages 63–71, 2020. antske fokkens, piek vossen, marco rospocher, rinke hoekstra, willem r van hage, and fondazione bruno kessler. grasp: grounded representation and source perspective. proceedings of knowledge resources for the socio-economic sciences and humanities associated with ranlp, 17:19–25, 2017. tingchen fu, shen gao, xueliang zhao, ji-rong wen, and rui yan. learning towards conversational ai: a survey. ai open, 3:14–28, 2022. doi:10.1016/j.aiopen.2022.02.001. jing gao, peng li, zhikui chen, and jianing zhang. a survey on deep learning for multimodal data fusion. neural computation, 32(5):829–864, 2020. alison gopnik and henry m. wellman. why the child’s theory of mind really is a theory. mind & language, 7(1-2):145–171, 1992. doi:10.1111/j.1468-0017.1992.tb00202.x. minlie huang, xiaoyan zhu, and jianfeng gao. challenges in building intelligent open-domain dialog systems. acm transactions on information systems (tois), 38(3):1–32, 2020. url https://dl.acm.org/doi/abs/10.1145/3383123. casey kennington, ting han, and david schlangen. temporal alignment using the incremental unit framework. in proceedings of the 19th acm international conference on multimodal interaction, pages 297–301, 2017. casey kennington, daniele moro, lucas marchand, jake carns, and david mcneill. rrsds: towards a robot-ready spoken dialogue system. in olivier pietquin, smaranda muresan, vivian chen, 48 https://arxiv.org/abs/2009.11152 https://doi.org/10.1016/j.aiopen.2022.02.001 https://doi.org/10.1111/j.1468-0017.1992.tb00202.x https://dl.acm.org/doi/abs/10.1145/3383123 a modular architecture for creating multimodal embodied agents casey kennington, david vandyke, nina dethlefs, koji inoue, erik ekstedt, and stefan ultes, editors, proceedings of the 21th annual meeting of the special interest group on discourse and dialogue, pages 132–135, 1st virtual meeting, july 2020. association for computational linguistics. doi:10.18653/v1/2020.sigdial-1.17. url https://aclanthology.org/2020. sigdial-1.17/. taewoon kim and piek vossen. emoberta: speaker-aware emotion recognition in conversation with roberta, 2021. url https://arxiv.org/abs/2108.12009. eric kolve, roozbeh mottaghi, winson han, eli vanderbilt, luca weihs, alvaro herrasti, matt deitke, kiana ehsani, daniel gordon, yuke zhu, aniruddha kembhavi, abhinav kumar gupta, and ali farhadi. ai2-thor: an interactive 3d environment for visual ai. arxiv, abs/1712.05474, 2017. url https://api.semanticscholar.org/corpusid:28328610. ronak kosti, jose m alvarez, adria recasens, and agata lapedriza. context based emotion recognition using emotic dataset. ieee transactions on pattern analysis and machine intelligence, 42 (11):2755–2766, 2019. url https://ieeexplore.ieee.org/document/8713881. lea krause, selene baez santamaria, lucia donatelli, and piek vossen. the role of personal perspectives in open-domain dialogue: towards enhanced data modelling and long-term memory. in proceedings of bnaic/benelearn the joint international scientific conferences on ai and machine learning, delft, november 8-10, 2023, 2023. jaap kruijt. referring expressions in dynamic human-robot common ground. phd thesis, vrije universiteit amsterdam, 2025. jaap kruijt, peggy van minkelen, lucia donatelli, piek tjm vossen, elly konijn, and thomas baier. spotter: a framework for investigating convention formation in a visually grounded human-robot reference task. in proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (lrec-coling 2024), pages 15202–15215, 2024. steven macenski, tully foote, brian gerkey, chris lalancette, and william woodall. robot operating system 2: design, architecture, and uses in the wild. science robotics, 7(66):eabm6074, 2022. doi:10.1126/scirobotics.abm6074. shikib mehri and maxine eskenazi. usr: an unsupervised and reference free evaluation metric for dialog generation. in proceedings of the 58th annual meeting of the association for computational linguistics, pages 681–707, 2020. thilo michael. retico: an incremental framework for spoken dialogue systems. in olivier pietquin, smaranda muresan, vivian chen, casey kennington, david vandyke, nina dethlefs, koji inoue, erik ekstedt, and stefan ultes, editors, proceedings of the 21th annual meeting of the special interest group on discourse and dialogue, pages 49–52, 1st virtual meeting, july 2020. association for computational linguistics. doi:10.18653/v1/2020.sigdial-1.6. url https://aclanthology.org/2020.sigdial-1.6. kazuki miyazawa and takayuki nagai. survey on multimodal transformers for robots. authorea preprints, 2023. doi:10.36227/techrxiv.21993317.v1. 49 https://doi.org/10.18653/v1/2020.sigdial-1.17 https://aclanthology.org/2020.sigdial-1.17/ https://aclanthology.org/2020.sigdial-1.17/ https://arxiv.org/abs/2108.12009 https://api.semanticscholar.org/corpusid:28328610 https://ieeexplore.ieee.org/document/8713881 https://doi.org/10.1126/scirobotics.abm6074 https://doi.org/10.18653/v1/2020.sigdial-1.6 https://aclanthology.org/2020.sigdial-1.6 https://doi.org/10.36227/techrxiv.21993317.v1 baier, báez santamaría and vossen alec radford, jong wook kim, tao xu, greg brockman, christine mcleavey, and ilya sutskever. robust speech recognition via large-scale weak supervision. in andreas krause, emma brunskill, kyunghyun cho, barbara engelhardt, sivan sabato, and jonathan scarlett, editors, proceedings of the 40th international conference on machine learning, volume 202 of proceedings of machine learning research, pages 28492–28518. pmlr, 23–29 jul 2023. url https://proceedings.mlr.press/v202/radford23a.html. stephen roller, emily dinan, naman goyal, da ju, mary williamson, yinhan liu, jing xu, myle ott, kurt shuster, eric m smith, et al. recipes for building an open-domain chatbot. arxiv preprint arxiv:2004.13637, 2020. url https://arxiv.org/abs/2004.13637. david schlangen and gabriel skantze. a general, abstract model of incremental dialogue processing. dialogue & discourse, 2(1):83–111, 2011. kurt shuster, jing xu, mojtaba komeili, da ju, eric michael smith, stephen roller, megan ung, moya chen, kushal arora, joshua lane, et al. blenderbot 3: a deployed conversational agent that continually learns to responsibly engage. arxiv preprint arxiv:2208.03188, 2022. url https://arxiv.org/abs/2208.03188. sonja stange, teena hassan, florian schröder, jacqueline konkol, and stefan kopp. self-explaining social robots: an explainable behavior generation architecture for human-robot interaction. frontiers in artificial intelligence, 5:866920, 2022. gemini team, rohan anil, sebastian borgeaud, jean-baptiste alayrac, jiahui yu, radu soricut, johan schalkwyk, andrew m dai, anja hauth, katie millican, et al. gemini: a family of highly capable multimodal models. arxiv preprint arxiv:2312.11805, 2023. hugo touvron, thibaut lavril, gautier izacard, xavier martinet, marie-anne lachaux, timothée lacroix, baptiste rozière, naman goyal, eric hambro, faisal azhar, et al. llama: open and efficient foundation language models. arxiv preprint arxiv:2302.13971, 2023. chantal van son, tommaso caselli, antske fokkens, isa maks, roser morante, lora aroyo, and piek vossen. grasp: a multilayered annotation scheme for perspectives. in proceedings of the tenth international conference on language resources and evaluation (lrec’16), pages 1177–1184, 2016. yuli vasiliev. natural language processing with python and spacy: a practical introduction. no starch press, 2020. piek vossen and antske fokkens. grasp: a model for the perspective web. creating a more transparent internet: the perspective web, page 260, 2022. piek vossen, lenka bajčetić, selene baez, suzana bašić, and bram kraaijeveld. modelling context awareness for a situated semantic agent. in modeling and using context: 11th international and interdisciplinary conference, context 2019, pages 238–252, trento, italy, 2019. piek vossen, selene báez santamaría, and thomas baier. a conversational agent for structured diary construction enabling monitoring of functioning & well-being. in hhai 2024: hybrid human ai systems for the social good, pages 315–324. ios press, 2024. 50 https://proceedings.mlr.press/v202/radford23a.html https://arxiv.org/abs/2004.13637 https://arxiv.org/abs/2208.03188 a modular architecture for creating multimodal embodied agents joseph weizenbaum. eliza—a computer program for the study of natural language communication between man and machine. communications of the acm, 9(1):36–45, 1966. url https: //dl.acm.org/doi/10.1145/365153.365168. tianyu wu, shizhu he, jingping liu, siqi sun, kang liu, qing-long han, and yang tang. a brief overview of chatgpt: the history, status quo and potential future development. ieee/caa journal of automatica sinica, 10(5):1122–1136, 2023. dian yu and zhou yu. midas: a dialog act annotation scheme for open domain human machine spoken conversations. arxiv preprint arxiv:1908.10023, 2019. url https://arxiv.org/ abs/1908.10023. 51 https://dl.acm.org/doi/10.1145/365153.365168 https://dl.acm.org/doi/10.1145/365153.365168 https://arxiv.org/abs/1908.10023 https://arxiv.org/abs/1908.10023 baier, báez santamaría and vossen appendix a. emissor multimodal data representation figure 3: visualization of an interaction in the emissor data format. figure 3, taken from (báez santamaría et al., 2021), shows a visualization of four modalities (text, audio, visual, and knowledge). signals are grounded in a temporal container on the horizontal axis, with bars marking alignments through the temporal ruler. red boxes mark segments annotated as mentions of objects (pills and the table). text segments highlighted in red are annotated as mentions of triples. the upper graphs represent the corresponding triples of the ekg generated from the annotated source modalities along the temporal sequence. the visual modality shows two different camera viewpoints (left is what carl sees and right is what leolani sees) concatenated side by side. 52 a modular architecture for creating multimodal embodied agents appendix b. interaction plots figures 4 and 5 show plot representations of interactions where each interaction point is scored on the y-axis as a negative or positive interaction "move". interaction points are text signals (possibly derived from audio signals), perceptions, or actions. text signals are annotated for dialog acts, emotions and likelihood using the usr model. the following mapping is used for dialog acts and emotions: • dialog acts: negative : neg_answer, complaint, abandon, apology,non-sense, hold. positive : pos_answer, back-channeling, appreciation, thanking, respond_to_apology. • emotions: negative : remorse, nervousness, fear, sadness, embarassment, disappointment, grief, disgust, anger, annoyance, disapproval, confusion. positive : joy, amusement,excitement, love, desire, optimism, caring, pride, admiration, gratitude, belief, approval, curiosity. a threshold is set to accept only emotions and dialog acts scoring above.5. dialog acts such as statement and opinion, which are the most frequent, are ignored. the same holds for a neutral emotion. likelihood scores between 0 and 1 are deducted with 0.3 so that a score below this threshold adds negatively to the interaction score and a score above adds positively. every positive dialog act or emotion adds 1 point to the score, and every negative act and emotion -1. the cumulative scores are plotted on the y-axis. in the case of interaction with the embodied leolani-agent emotions are also extracted from human faces. the labels in the plot start with the source of the signal, followed by the content of the signal (expressed text or annotation label), and any annotation with a score. note that the responses from leolani are based on the episodic knowledge graph. the triple results are first converted to text using templates. next, llama3.2.1b is prompted to paraphrase the input in simple english using a maximum of 100 tokens and a temperature of 0.2. 53 baier, báez santamaría and vossen figure 4: multimodal interaction plot showing text and image signals picked up and created by leolani. signal annotations are shown per source, speakers, and camera, where the camera shows ids for people, types of object detected with a confidence score, and gender and age predictions for human faces. dotted green lines represent the human interlocutor, striped red lines represent utterances from the agent, and solid blue lines the perceptions by the agent. the text signals are annotated with a usr likelihood score, and if above a threshold of 0.5 confidence, the dialog act and emotion annotation are shown. the annotations are mapped on the y-axis ranging from -5 to 5 to the extent that they indicate negative or positive turns in the conversation. face expressions are also mapped to the negative-positive scale. 54 a modular architecture for creating multimodal embodied agents figure 5: multimodal interaction plot showing text, image and action signals picked up and created by an ai2thor agent in a virtual world. striped orange lines represent the utterances from the human interlocutor, striped-dotted red lines represent the actions by the ai2thor agent and solid blue lines the utterances rom the ai2thor agent. the dotted green line represents the perceptions of the agent. the text signals (utterances and actions) are annotated with a usr likelihood score, and if above a threshold of 0.5 confidence, the dialog act and emotion and sentiment labels. the annotations are mapped on a range from -5 to 5 to the extent that they indicate negative or positive turns in the conversation. 55 baier, báez santamaría and vossen appendix c. example component structure cltl-mycomponent config/ requirements.txt setup.py src/ app.py cltl/ mycomponent/ api.py implementation.py cltl_service/ mycomponent/ service.py appendix d. example service.py module from c l t l . combot . i n f r a . e v e n t import event , eventbus from c l t l . combot . i n f r a . t i m e _ u t i l import t imestamp_now from c l t l . combot . i n f r a . t o p i c _ w o r k e r import topicworker from c l t l . combot . e v e n t . e m i s s o r import t e x t s i g n a l e v e n t from e m i s s o r . r e p r e s e n t a t i o n . s c e n a r i o import t e x t s i g n a l c l a s s s i m p l e s e r v i c e : def _ _ i n i t _ _ ( s e l f , i n p u t _ t o p i c : s t r , o u t p u t _ t o p i c : s t r , e v e n t _ b u s : eventbus ) : s e l f . _ i n p u t _ t o p i c = i n p u t _ t o p i c s e l f . _ o u t p u t _ t o p i c = o u t p u t _ t o p i c s e l f . _ e v e n t _ b u s = e v e n t _ b u s s e l f . _ t o p i c _ w o r k e r = none def s t a r t ( s e l f ) : s e l f . _ t o p i c _ w o r k e r = topicworker ( [ s e l f . _ i n p u t _ t o p i c ] , s e l f . _even t_bus , p r o v i d e s =[ s e l f . _ o u t p u t _ t o p i c ] , p r o c e s s o r = s e l f . _ p r o c e s s ) s e l f . _ t o p i c _ w o r k e r . s t a r t ( ) . w a i t ( ) def s t o p ( s e l f ) : i f not s e l f . _ t o p i c _ w o r k e r : pass s e l f . _ t o p i c _ w o r k e r . s t o p ( ) s e l f . _ t o p i c _ w o r k e r . a w a i t _ s t o p ( ) s e l f . _ t o p i c _ w o r k e r = none def _ p r o c e s s ( s e l f , e v e n t : event [ t e x t s i g n a l e v e n t ] ) : input = e v e n t . p a y l o a d . s i g n a l . t e x t ) r e s p o n s e = " " ### process the input i f r e s p o n s e : o u t p u t _ e v e n t = s e l f . _ c r e a t e _ p a y l o a d ( r e s p o n s e ) s e l f . _ e v e n t _ b u s . p u b l i s h ( s e l f . _ o u t p u t _ t o p i c , event . f o r _ p a y l o a d ( o u t p u t _ e v e n t ) ) def _ c r e a t e _ p a y l o a d ( s e l f , r e s p o n s e ) : s i g n a l = t e x t s i g n a l . f o r _ s c e n a r i o ( none , t imestamp_now ( ) , t imestamp_now ( ) , none , r e s p o n s e ) re turn t e x t s i g n a l e v e n t . c r e a t e ( s i g n a l ) 56 a modular architecture for creating multimodal embodied agents appendix e. example app.py application from c l t l . combot . i n f r a . e v e n t . memory import synchronouseventbus from c l t l _ s e r v i c e . mycomponent import s i m p l e s e r v i c e e v e n t _ b u s = synchronouseventbus ( ) s e r v i c e = s i m p l e s e r v i c e ( " t e x t _ i n " , " t e x t _ o u t " , e v e n t _ b u s ) t r y : s e r v i c e . s t a r t ( ) whi le true : t ime . s l e e p ( 1 ) e xc ep t k e y b o a r d i n t e r r u p t : s e l f . s e r v i c e . s t o p ( ) 57 baier, báez santamaría and vossen appendix f. dialog flow implementations legend start introduction start round acknowledge noquery next no yes all rounds complete? no yessuccess? yes all positiions filled? closing & goodbye mention disambiguation identify failure status repair > 3 failure count no yes response abandon and continue checkbox confirmation yes no round=1 & session =1? 80% 20% chance encouragement disambiguate finish roundcontinue button continue button continue button check conversational state human utterance disambiguator process dm process game event robot response low high certainty measure query confirmation question move position acknowledge request repair acknowledge repair current state? figure 6: dialog flow of the spotter application taken from kruijt (2025), showing conversational states and dialog management processes that are modeled as high-order intents in the leolani platform to control the dialog for the purpose of the game. 58 a modular architecture for creating multimodal embodied agents figure 7: dialog example of an agent that uses conversations to reconstruct a person’s timeline of dialy activities of life, taken from vossen et al. (2024) 59 introduction related work event bus architecture enabling reasoning components available processing audio signals processing text and rdf signals processing image signals sample input-output signal sequences responding crossmodal awareness higher level intentions interaction evaluation developing your own agent adding, replacing and mixing components application types and performance out of the box conclusions emissor multimodal data representation interaction plots example component structure example service.py module example app.py application dialog flow implementations dialogue & discourse 8(2) (2017) 56–83 doi: 10.5087/dad.2017.203 examples and specifications that prove a point: identifying elaborative and argumentative discourse relations merel c.j. scholman m.c.j.scholman@coli.uni-saarland.de department of language science and technology saarland university vera demberg vera@coli.uni-saarland.de department of computer science department of language science and technology saarland university editor: david traum submitted 11/2016; accepted 03/2017; published online 07/2017 abstract examples and specifications occur frequently in text, but not much is known about how how readers interpret them. looking at how they’re annotated in existing discourse corpora, we find that annotators often disagree on these types of relations; specifically, there is disagreement about whether these relations are elaborative (additive) or argumentative (pragmatic causal). to investigate how readers interpret examples and specifications, we conducted a crowdsourced discourse annotation study. the results show that these relations can indeed have two functions: they can be used to both illustrate / specify a situation and serve as an argument for a claim. these findings suggest that examples and specifications can have multiple simultaneous readings. we discuss the implications of these results for discourse annotation. keywords: coherence relations, crowdsourcing, discourse annotation, inter-annotator agreement, signalling 1. introduction discourse relations (also referred to as coherence relations) are semantic links between two (or more) discourse units (cf. hobbs, 1979; mann & thompson, 1988; sanders, spooren & noordman, 1992). they can be explicitly signalled by connectives (also referred to as discourse markers, or dms) such as because or for example. however, many relations are implicit, that is, they are not marked by a connective. this is illustrated in example 1. (1) packaging has some drawbacks. the additional technology, personnel training and promotional effort can be expensive. wsj 0085 in order to be able to do quantitative investigations of discourse relations, text corpora with annotated discourse relations are necessary. in recent years, several discourse annotated corpora have been developed; some inspiring examples are the penn discourse treebank (pdtb; prasad et al., 2008) and the rhetorical structure theory discourse treebank (rst-dt; carlson, marcu & okurowski, 2003). both corpora contain articles from the wall street journal, and for a substantial amount of articles, annotations from both resources are available. in order to determine whether c©2017 merel c.j. scholman and vera demberg this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). identifying elaborative and argumentative instantiations and specifications annotations from these corpora are compatible, i.e., whether the frameworks agree on the sense of a relation, demberg, asr & scholman (2017) mapped the annotations of overlapping texts of these two treebanks onto each other. we will refer to the texts with annotations from both discourse corpora as the “wsj-aligned corpus”. the mapping allows for a comparison between the labels from pdtb and rst-dt, which reveals how much the two frameworks agree with each other as well as potentially interesting patterns of disagreement. while there is usually no one-to-one correspondence between pdtb and rst-dt relation labels; it is possible to determine the correspondences between labels based on the annotation guidelines of the two frameworks (also see sanders et al., 2016, submitted). demberg et al. (2017) found that inter-framework agreement on explicit relations was reasonable (ca. 60% agreement), while agreement on implicit relations was much lower (roughly 35% agreement). the difference in the agreements in annotation between explicit and implicit relations can be explained by the presence of connectives: achieving high inter-annotator agreement on implicit relations is harder than on explicit relations because annotators cannot rely on connectives, which represent a strong cue. but even when taking into account the difficulty of annotating discourse relations, the level of agreement in annotation between frameworks on the same texts might seem astonishingly low. the goal of this article is to establish whether there are systematic reasons for the disagreement, whether the different annotations are justified, and what the implications for discourse annotation are. in order to answer these questions, we took a closer look at two types of relations for which inter-framework agreement is particularly poor, and investigate the factors affecting interpretation of these relations in more detail. the relations under investigation are pdtb’s instantiation and specification relations (32% and 14% agreement, respectively). these relations do not have many prototypical connectives and are therefore hard to identify (taboada & das, 2013; vergezcouret & adam, 2012).they are nevertheless very frequent: together, they make up 24% of all implicit relations in the pdtb. both relations are subtypes of the pdtb class expansion. instantiation is a second-level label in the pdtb hierarchy, whereas specification is a third-level label in the pdtb type restatement (see appendix a for the pdtb and rst-dt taxonomies). for convenience, expansion.restatement.specification will be referred to as specification in this paper. in both of these relation types, one segment further specifies a set or situation described in the other segment (halliday, 1994). instantiations and specifications are by definition primarily considered as elaborative (also referred to as additive) kinds of relations, i.e. they are discourse relations that connect utterances describing the same situation from different angles (jasinskaja, 2013). a large proportion of the mismatch with the corresponding rst-dt annotations stems from these same pairs of arguments being annotated as argumentative (causal) relations; namely explanation-argumentative and evidence. in such argumentative relations, one segment gives support to the other segment (jasinskaja & karagjosova, 2011). blakemore (1997) argues that instantiations and specifications can indeed have two functions: they can be used to both illustrate / specify a situation and serve as an argument to a claim. pdtb allows the annotation of more than one relation sense, but we found that the argumentative function was usually not annotated. annotators could also assign an argumentative label such as contingency.pragmatic cause or a causal label such as contingency.cause, but in practice, pdtb rather focuses on the illustrative / specification aspect of the relation. if readers do indeed infer the argumentative reading of the instantiations and specifications, it is however important that this reading is also reflected in the label that the relation receives in a corpus. otherwise, the annotation of those instantiations and specifications cannot be considered as 57 scholman and demberg fully descriptively adequate. similarly, when rst-dt annotates the argumentative aspect of the relation (by labeling relations as explanation-argumentative and evidence), the illustrative / specification aspect of the relation does not become visible in the annotations (which is reflected by the labels example and general-specific). considering that the labels instantiation and specification occur so often, the current study sets out to investigate how readers actually interpret these relations. do the instances of instantiation and specification relations from the pdtb vary in the degree to which they can be interpreted as argumentative? can we identify specific characteristics of the relations that have additional argumentative functions? are there also other alternative interpretations that are inferred? we aim to answer these questions by asking participants via a crowdsourcing platform to insert connectives from a predefined list that are considered to be relatively unambiguous in between the segments of coherence relations. the insights from this investigation can be used to help to better understand the functions of instantiation and specification relations, and, for annotation purposes, to more reliably identify and classify instantiation and specification relations, which in turn can improve the agreement on these frequently-occurring classes. in sum, the current paper deals with the following distinct issues: (i) the mismatch between annotations of instantiation and specification relations in the pdtb and rst-dt are analysed, (ii) the variability of presumably valid interpretations of instantiation and specification relations (elaborative vs. argumentative), and (iii) the usability of crowdsourced annotations for these research purposes. the lay-out of the paper is as follows: in the next section, we discuss instantiation and specification relations and their signalling. we then present a crowdsourcing experiment, showing that instantiation and specification relations are indeed often interpreted by our participants as argumentative. the paper concludes with a discussion of the results and their implications for discourse annotation. 2. background instantiation and specification both belong to the pdtb class expansion, which is primarily an elaborative class. the difference between the labels is quite subtle. a relation is labeled as instantiation when “the connective indicates that arg1 evokes a set and arg2 describes it in further detail” (prasad et al., 2007, p. 34).1 the set described in arg1 may be a set of events, reasons, behaviors, etc. instantiations are typically marked by for example and as an illustration. example (2) illustrates this type of relation. in specification relations, arg2 also typically describes arg1 in more detail, but arg2 is also a logical implication of arg1. specifications are typically marked by specifically and in fact. example (3) presents an example of a specification relation. the labels are therefore similar in that arg2 expands on arg1, but they differ in the content of arg1 (presence or absence of a set and logical implication). (2) in an age of specialization, the federal judiciary is one of the last bastions of the generalist. a judge must jump from murder to antitrust cases, from arson to securities fraud, without missing a beat.2 wsj 601 1. pdtb uses a lexical grounding approach, for which connectives are annotated for both explicit and implicit relations. in the case of implicit relations, annotators insert a connective and annotate the corresponding sense of the relation. 2. in line with the pdtb, the first segment of a relation (arg1) will be indicated with italic font, and the second segment (arg2) will be indicated with bold font. 58 identifying elaborative and argumentative instantiations and specifications (3) in the cornucopia of go-go apples, the fuji’s track record stands out. during the past 15 years, it has gone from almost zilch to some 50% of japan’s market. wsj 1128 the equivalent of these relation labels in rst are example for instantiation and elaboration-general-specific3 for specification. however, the wsj-aligned corpus shows that often pdtb’s instantiation and specification relations fall into other classes in rst-dt. as illustrated in figure 1, other common labels for instantiation relations are general-specific (which maps onto pdtb’s specification label), elaboration-additional (the most basic relation in rst), explanation-argumentative, and evidence (both typically causal labels in rst). these same rst labels were used for relations that received the specification label in pdtb. hence, rst often assigns causal labels to relations that pdtb labels as elaborative instantiation or specification (25% of pdtb instantiation and 21% of pdtb specification relations receive causal labels in rst). figure 1: 306 pdtb instantiation and 426 pdtb specification relations (explicit and implicit) annotated according to rst labels. blue rst labels are primarily elaborative, orange rst labels are primarily causal. a similar pattern does not occur for the rst classes example and general-specific: figure 2 shows that pdtb annotators did not assign a causal label as often to rst’s example and general-specific relations (7% and 12%, respectively). hence, for relations that rst annotators consider to be examples and general-specifications, pdtb annotators tend annotate the same reading. a possible explanation for this difference between frameworks reagrding the annotation of instantiations and specifications is that the frameworks’ procedures influence the resulting annotations. pdtb has a connective-based approach: annotators are instructed to annotate the connective if one is present, or insert one and then annotate the relation, when the relation is implicit. annotators are not instructed to systematically try to insert certain connectives before trying others. they can choose from a list containing many connectives, and often multiple connectives can be used for the same relation. for example, pdtb annotators inserted the discourse marker specifically for example 3 above, but because could have also been inserted in this relation. the choice of inserted connective therefore depends entirely on the annotator, and this can vary from annotator to annotator. it is plausible that different annotators have different biases for inferring one 3. elaboration-general-specific will be abbreviated to general-specific in the remainder of this paper. 59 scholman and demberg figure 2: 168 rst example and 108 rst general-specific relations (explicit and implicit) annotated according to pdtb labels. sense over another, but this is not retraceable in the corpus itself. it is therefore difficult to determine what the framework’s bias is. rst, by contrast, doesn’t explicitly make use of connectives at all; rather, the framework instructs annotators to focus on the writer’s intentions. from this viewpoint, it is likely that even though specifically or for example can be used to express the relation, the intention of the writer in fact has an argumentative nature: to convince the reader of a point by providing evidence for a claim. indeed, pdtb’s instantiation and specification classes contain instances that can be analysed as consisting of a claim and an argument. in the case of instantiation relations, the segment containing the instantiation can be seen as supporting a claim in the other argument. consider the following example, annotated as pdtb instantiation – rst evidence: (4) “if you are born to give parties, you give parties. even in russia we managed to give parties.” wsj 1367 the first segment of this example is a claim, and the second segment gives an illustration of the claim. this reading is referred to as the ideational (or semantic) reading, which involves the relation between the information conveyed in consecutive elements of a coherent discourse (cf. moore & pollack, 1992). besides this ideational reading of example 4, the second segment can also be interpreted as a premise that underpins the validity of the claim. with this reading, the relation is in fact pragmatic causal, or argumentative.4 argumentative relations are relations in which the writer attempts to affect the addressee’s beliefs, attitudes, desires etc. by means of language (cf. hovy & maier, 1995). specification relations have a similar ambiguity between an ideational and an argumentative reading: the second segment can serve as merely providing more information about a concept or situation in the first segment, or it can provide support for a claim in the first segment. this double function of instantiations and specifications was brought up by carston (1993, p. 164), who noted that “exemplification is a common way of providing evidence to support a claim, or, equivalently, of giving a reason for believing something.” building on this, blakemore (1997) ar4. the terms ideational and argumentative are also known as informational and intentional (moore & pollack, 1992), subject matter and presentational (mann & thompson, 1988), and semantic and interpersonal (hovy & maier, 1995). 60 identifying elaborative and argumentative instantiations and specifications gues that instantiation and specification relations can have different functions in a text, and that classifying them as only ideational or only argumentative does not do justice to the way that these relations are interpreted. the double function of instantiation and specification relations has also been noted in other descriptive work on elaboration relations (see, for example, cuenca, 2003; hyland, 2007; jasinskaja & karagjosova, 2011), but it is not reflected in discourse annotation frameworks. blakemore (1997) however argues that the classification of these relations in frameworks is irrelevant; rather, one should focus on how the utterance achieves relevance with the reader. the current study argues against this claim: the classification of instantiations and specifications is important considering that corpora should be descriptively adequate. returning to the classification of these relations, we can conclude (based on figures 1 and 2) that pdtb’s instantiation and specification classes and rst’s general-specific class contain a mixture of both argumentative and non-argumentative relations. this finding is actually not surprising: it has been noted that these frameworks do not have distinct labels for specific argumentative relations since they are focused on identifying general discourse structures (peldszus & stede, 2013; stab & gurevych, 2014). moreover, as biran & rambow (2011) note, argumentation is not characterised by a single discourse relation; instead, it can be realised by a large number of discourse relation types. this does not imply that all relation types necessarily have an argumentative and a non-argumentative reading, but it does seem to apply to the instantiation and specification type of relations. if readers can actually infer an argumentative reading for such non-causal relation types, one could argue that discourse annotation frameworks would need to distinguish an argumentative counterpart of these labels in order to be able to represent the the argumentative reading of those non-causal relations. another way of incorporating the argumentative reading of a non-causal relation in its annotation, is to annotate multiple senses for the same relation. consider example 4: relations like this example actually have multiple readings (both elaborative and argumentative), and multiple labels would therefore adequately express the readings of this relation. unfortunately, the annotated relations that are currently available do not carry double annotations.5 this is because these corpora are built on the assumption that at most a single discourse relation holds between two segments. however, properties of the discourse adverbial instead (webber, 2013) have shown that some relations can actually express multiple meanings. instead is a strong marker for exception relations, but when instead occurs at the beginning of a sentence, another relation can often be inferred as well, as illustrated in 5-7, taken from rohde et al. (2016). (5) i planned to make lasagna. instead i made hamburgers. → but instead i made hamburgers. (6) i don’t know how to make lasagna. instead i made hamburgers. → so instead i made hamburgers. (7) surprisingly, they ignored the lasagna. instead they just ate the salad. → and instead they just ate the salad. as examples 5-7 show, two segments marked by instead can express different relations, including the exception relation that the adverbial instead marks. building on these observations, rohde 5. pdtb does allow for double annotations, but this is rarely applied in practice: less than 5% of instances in the pdtb carry two labels. the pdtb group plans to release a new version, pdtb 3.0, in which double labels occur more frequently (webber et al., 2016). 61 scholman and demberg et al. (2016) showed that relations marked by discourse adverbials other than instead can also express multiple meanings. these results indicate that multiple discourse relations can hold between two discourse segments marked by an adverbial. based on these findings, we expect instantiation and specification relations to also have multiple readings. in sum, the fact that human annotators often disagree on the elaborative or argumentative nature of pdtb’s instantiation and specification relations can be due to the operationalization of the frameworks, or to an actual ambiguity or double-function of these relations that is not captured in pdtb’s framework. the current study therefore investigates how comprehenders interpret these relations. the obtained annotations will also be used to identify certain cues that might have influenced the readers’ interpretations (in other words, cues that signal the elaborative or argumentative type of instantiation and specification relations). we will now turn to a more detailed explanation of the methodology. 3. method we conducted a crowdsourcing experiment for which naı̈ve (non-trained, non-expert) annotators were asked to insert connectives from a predefined list into coherence relations taken from wsjaligned. naı̈ve subjects were chosen over expert annotators for several reasons. first, naı̈ve annotators are not influenced by their prior experience with a specific framework, thereby eliminating the possibility of a framework bias or instruction bias occurring in the annotations. moreover, they are not experts in discourse annotation, and their annotations therefore do not rely on implicit, expert knowledge (riezler, 2014). third, working with naı̈ve annotators has the practical advantage that they are easier to come by. this makes it easier to employ a larger number of annotators, which decreases the effect of annotator bias (artstein & poesio, 2005). studies employing naı̈ve annotators have found high agreement between these annotators and expert annotators for natural language tasks (for example, snow, o’connor, jurafsky & ng, 2008). more recently, they have been employed successfully in coherence relation classification tasks (kawahara, machida, shibata, kurohashi, kobayashi & sassano, 2014; rohde, dickinson, clark, louis & webber, 2015; rohde, dickinson, schneider, clark, louis & webber, 2016; scholman, evers-vermeul & sanders, 2016). however, crowdsourced annotators – unlike expert annotators or in-lab naı̈ve annotators – cannot be asked to code according to a specific framework because this would require them to study manuals. therefore, rather than asking for specific relation labels, we ask them to insert a connective from a predefined list of connective phrases. this is not to say that a single connective cannot mark multiple types of relations. in fact, connectives are well known to be ambiguous and multifunctional (see asr & demberg, 2013; degand, 1998; hovy, 1995; maschler & schiffrin, 2015; versley, 2011, among many others). in order to ensure that the connectives included in this experiment are as unambiguous as possible, we chose connectives based on a classification of connective substitutability by knott & dale (1994). it is assumed that these connectives are typical markers of the relational classes included in this experiment. however, we do not assume that, for example, an instantiation marker cannot also imply an argumentative reading. rather, the method is based on the assumption that participants choose the connective that matches the strongest reading that they infer. hence, if they infer an argumentative reading for a specific relation annotated as instantiation by pdtb annotators, they will choose a causal connective. if they believe both readings hold, they can choose multiple connectives. similar methodologies have been used to obtain in62 identifying elaborative and argumentative instantiations and specifications sights into readers’ interpretations of relations by rohde et al. (2015, 2016); sanders et al. (1992) and scholman et al. (2016). inserting a connecting phrase is of course not the same as assigning a relation label to two arguments. using connecting phrases to obtain annotations leads to more coarse-grained observations. however, this method can be used to tap into the interpretation of a relation by a large group of annotators, thereby reducing the effect of individual biases. moreover, the current method reveals a distribution of the senses of the discourse relation at hand, making this method more suitable for reflecting a situation in which a relation can have multiple interpretations (cf. cuenca & marı́n, 2009; rohde et al., 2015, 2016; webber et al., 2001). 3.1 participants 111 native english speakers (47 female) completed one or more batches of this experiment. they were recruited via prolific academic and reimbursed for their participation (2 gbp per batch). participants came from the united states, united kingdom, ireland, and australia. their education level ranged from an undergraduate degree to a doctorate degree. 3.2 materials the experimental passages were implicit instantiation and specification relations taken from the wsj-aligned corpus, which contains fragments from the wall street journal newspaper. these relations were chosen to enable a comparison between the pdtb label and the rst label. the instantiation and specification relations that were included in this experiment fell into one of four rst classes: example, general-specific, evidence, and explanation-argumentative. these categories were chosen because relations that received pdtb’s instantiation or specification label fall in these four rst classes most often. in total, there were 54 instantiation and 60 specification items included in this experiment. for each of the four relevant rst classes, 15 instantiation and 15 specification items were taken, with the exception of the instantiation – general-specific combination: there were only 9 items the pdtb label instantiation and the rst label general-specific. when the wsj-aligned corpus contained more than 15 items annotated with the target pdtb and rst label, preference was given to items that (i) differed less than 60 characters between the rst and pdtb segmentation, (ii) did not contain attribution in one of the arguments, and (iii) dealt with a non-economic topic. the first criterion relates to the size of the segments: pdtb and rst have different segmentation rules, resulting in different segment sizes. we chose relations that differed as little as possible in segmentation. in the experiment, we adhered to the pdtb segmentation because the pdtb annotates the minimal amount of information necessary to infer the intended relation. the second criterion relates to attribution in the segments, that is, the explicit reference to the source (e.g., john said that). pdtb does not annotate attribution as part of an argument, but rather as a feature of the relation. therefore, if the source of the attribution was reported in between the two arguments, it was moved to the context sentence following the second argument, to ensure that participants did not treat it as part of the argument. the third criterion was composed for the motivation of participants: non-economical topics were considered more interesting than economical topics. 5. crowdsourcing platform, www.prolific.ac 63 www.prolific.ac scholman and demberg the fillers in this experiment consisted of 24 causal, 24 conjunction (additive), 36 concessive and 36 contrastive relations. causal relations are characterized by the presence of an implication between the two discourse segments; in other words, one segment contains a logical cause for a situation or event in the other segment. typical connectives are because and so. conjunction relations do not have an implication relation and can be expressed by and or also. concessive relations are considered as “negative causals” (konig & siemund, 2000), since they establish a similar causal relation but differ in their polarity (sanders et al., 1992). concessive relations are characterized by the presence of a denial of expectation: one segment contains a consequence despite a situation or event in the other segment. typical connectives for concessive relations are even though and nevertheless. contrastive relations are considered to be negative counterparts of additives (cf. sanders et al., 1992), since the segments are also connected in a conjunction, but a difference is highlighted between the segments. prototypical lexical markers for these relations are but and by contrast. to ensure that the fillers in this experiment were clear cases of a specific type of relation, they were selected based on the criterion that the pdtb and rst label were in agreement (for example, when pdtb gave an item a causal label but rst gave it an additive label, it was not chosen).6 all relations included in the study were originally implicit, except for concessive and some contrastive relations, since there were no implicit concession relations and too few implicit contrast relations in the wsj-aligned corpus. table 1 shows all pdtb and rst labels that these fillers carried per type of relation. relation type pdtb label rst label freq. causal reason, result cause-result, consequence, 24 explanation-argumentative, reason, result, evidence conjunction conjunction elaboration-additional 24 concessive contra-expectation, concession, antithesis 36 expectation contrastive contrast, juxtaposition, contrast, antithesis 36 opposition table 1: type of filler and the pdtb and rst labels that it can carry. the total of 234 items were divided into 12 batches, with 4 or 5 instantiation, 5 specification, 2 causal, 2 conjunction, 3 concessive, and 3 contrastive items per batch. order of presentation of the items per batch was randomized to prevent order effects. subjects were allowed to complete more than one batch, but saw every item only once. average completion time per batch was 16 minutes. due to presentation errors in one conjunction, two causal, and two concessive items, the final dataset for analysis including fillers consists of 229 items. connecting phrases – the list of connecting phrases consisted of: as an illustration, more specifically, in addition, because, as a result, even though, nevertheless, and by contrast. with the exception of the first two phrases, these connectives were chosen based on a classification of connectives in knott and dale (1994). 6. the rst label antithesis is ambiguous between a causal and an additive reading. it was nevertheless used as a criterion for both concessive and contrastive items because the subset was too small without the inclusion of rst antithesis. 64 identifying elaborative and argumentative instantiations and specifications knott & dale (1994) list for example as a typical marker of instantiation relations. however, we decided to choose as an illustration instead. the two connecting phrases are interchangeable, but for example is used more commonly than as an illustration. we hypothesized that inserting as an illustration would therefore require more active reasoning than for example. knott & dale (1994) did not list markers for specification relations. in the pdtb, specification relations are most often marked by in fact and indeed. these markers can also be used to mark causal relations. we decided to choose more specifically as a typical marker because it is less ambiguous. in addition was chosen over and for the same reason: and is an underspecified connective and can be used to mark several different types of relations, whereas in addition only marks additive relations. because and as a result were chosen to represent causal relations. both connectives can be used to mark argumentative relations (cohen, 1987). even though and nevertheless were included as markers of concession relations. by contrast was chosen as a typical marker of contrastive relations. other typical markers, such as but, although and however were not included because they can be used to mark both concessive and contrastive relations. similar to the order of the items, the order in which the connecting phrases were presented was also randomized for every item. 3.3 procedure the experiment was distributed via prolific academic and hosted on lingoturk (pusse, sayeed & demberg, 2016). first, participants were presented with instructions for the study. next, they were presented with the experiment interface, which consisted of three parts: a short summary of the instructions, a box with predefined connectives, and the text passage (see figure 3 for an example of the interface). the text passage contained two context sentences preceding the first segment and one context sentence following the second segment. these context sentences were taken from the original text and were not altered in any way. the two segments of the target relation were shown in black text while the context sentences were displayed in grey text. subjects were instructed to choose the connecting phrase that best reflected the meaning between the black text elements, but to take the grey text into account. in between the two segments of the coherence relation was a green box (see figure 3). participants were instructed to “drag and drop” the connecting phrase that “best reflected the meaning of the connection between the sentences” (cf. rohde et al., 2015) into this green box. participants could also choose two connecting phrases if both phrases reflected the meaning of the relation, using the option “add another connective”. moreover, they could manually insert a connecting phrase by clicking “none of these”, if they felt that none of the predefined options suited the relation. punctuation markers following the first argument of the relation were replaced by a double slash (//) (cf. rohde et al., 2015) to avoid participants from being influenced by the original punctuation markers (for example, not insert the connective because because of a full stop after the first segment). the second segment always started with a lowercase letter. 4. results prior to analysis, the data of 4 participants were removed because these participants had very short completion times (<10 minutes for 20 passages of 5 sentences each) and showed high disagreement on causal and concessive items with other participants. the following analyses do not take the responses of these participants into consideration, leaving us with a total of 2962 observations. in 65 scholman and demberg figure 3: example of the interface of the experiment – in this case, the participant chose as an illustration to indicate the meaning between the segments. total, each list was completed by 12 to 14 participants. 25 participants completed between two and six batches; the others (86 participants) completed only one batch. for the following analyses, we aggregated frequencies of the connectives that fell into the same class. in other words, because and as a result were aggregated as causal connectives, and even though and nevertheless were aggregated as concessive connectives. in 3.7% of the instances, participants made use of the ‘manual answer’ option, to insert connectives that were not on the provided list or to avoid inserting a connective (by leaving a blank or inserting punctuation). we discuss these manual insertions separately in section 4.4, and aggregate the class ‘manual answer’ and ‘no answer’ for our analyses in sections 4.1 to 4.3.2. as with any discourse annotation task, some variation in the distribution of insertions can be expected. we are therefore interested in larger patterns in the distribution of inserted connectives. crucially, we do not assume that there is one single correct label for each of our experimental items (cf. rohde et al., 2016). in cases where we for example observe a distribution across two connectives in the insertions, we take this to indicate that this particular class or relation might have multiple senses. in order to determine the reliability of our method, we calculate agreement between connective insertions from our experiment and a replication experiment, and also evaluate the agreement with the originally annotated discourse treebank labels for these items. we report percentages of agreement and krippendorf’s alpha7 (α, krippendorff, 1980). for the filler items, alpha was calculated by comparing the pdtb label to the dominant response of the crowdsourced participants. 7. alpha was calculated using the r package agree.coeff2.r. 66 identifying elaborative and argumentative instantiations and specifications first, we discuss the results of the experiment overall, to determine whether the method was effective (section 4.1). next, we turn to a more comprehensive analysis of pdtb instantiations (section 4.2). we look at the insertions per rst class (section 4.2.1) to investigate whether the subjects interpreted the items in line with pdtb’s or rst’s classification. then we look at a few examples of clear elaborative and clear argumentative relations in more detail to be able to identify cues for these relation types (section 4.2.2). the same is discussed for specification items (sections 4.3.1 and 4.3.2). finally, we discuss observations regarding manual answers and double insertions (section 4.4). 4.1 reliability of the crowd-sourced annotation method the aim of our first analysis is to check whether the method of crowdsourcing the discourse relation labels is a valid method. to this end, our dataset includes not only specification and instantiation relations which are the focus of the study, but also conjunction, causal, contrastive and concessive relations. we found that our method is successful overall: the connectives inserted by the participants are consistent with the original pdtb annotation (α = .57). this is shown in figure 4, with the bars reflecting the inserted connective per pdtb class. 78% of the inserted connectives in items with a causal pdtb label were causal connectives (because and as a result), and 67% of the inserted connectives in concessive items were the concessive connectives even though and nevertheless. for both classes, the second most frequent category of inserted connectives was their negative/positive counterpart: for cause, the second most frequent category was concession (10%), and for concession, the second most frequent counterpart was cause (15%). on closer inspection of the items, we find that the disagreement between crowdsourced annotations and original pdtb annotations can be traced back to difficulties with specific items, and not to unreliability of the workers: the main cause for the confusion of causal and concessive relations can be attributed to the lack of context and/or background knowledge, especially for items with economic topics. for these topics, it can be very hard to judge whether a situation mentioned in one segment is a consequence of the other segment, or a denied expectation. the agreement on relations with conjunction and contrastive pdtb labels is lower than agreement on causal and concessive relations, but the distribution of inserted connectives for conjunction and contrast looks similar: the expected marker is used most often (40% and 44%, respectively), with the corresponding causal counterpart as the second most frequent inserted connective type (27% causal insertions and 32% concessive insertions, respectively). a closer look at the crowdsourced annotations for items in these classes reveals that this is due to genuine ambiguity figure 4: distribution (%) of inserted connectives per pdtb class. 67 scholman and demberg of the relation. for relations originally annotated as conjunction, we find that oftentimes a causal relation can also be inferred. the same explanation holds for contrastive relations. relations from this class that often receive concessive insertions are characterized by the reference to contrasting expectations. some confusion between these relations is expected, as concessive and contrastive relations are relatively difficult to distinguish even for trained annotators (see, for example, robaldo & miltsakaki, 2014; zufferey & degand, 2013). finally, looking at the instantiation and specification relations that we set out to investigate in more detail in our study, we can see that there is more variety in terms of which connective participants inserted. this was expected, and the connectives that were most often inserted (as an illustration, more specifically and because / as a result) are consistent with our hypotheses. in order to get a clearer picture of the elaborative and argumentative types of these relations, we will turn to a more fine-grained analysis of these relations per rst class in the following sections. the reliability of the current method was furthermore confirmed by a follow-up replication study: we repeated the experiment by presenting a new group of participants the same items without the context sentences to investigate the influence of context on interpretations (scholman & demberg, 2017). the connective insertions were almost a perfect replication of the results reported in this study, with the distribution being stable on an item-by-item basis. the agreement between the participants in the current study and the participants in the replication study is high: α = .71. on average, the difference between the experiments on agreement with the pdtb label differed only by 3.7%. we ran fisher exact tests on the insertions for each of the pdtb classes for the experiment reported here vs. the replication without context, and found no significant difference in the distribution of responses between studies for any of the pdtb classes (cause: p=0.61; conjunction: p=0.62; concession: p=0.98; contrast: p=0.88; instantiation: p=0.93; specification: p=0.85). the presence of context therefore did not significantly influence the results, see scholman & demberg (2017) for a more detailed discussion on this. importantly, the results show that the distribution of insertions for every item is stable when two different crowdsourced groups take part in the experiment. twelve insertions per item therefore seems to be an adequate amount to get a representative distribution of senses. 4.2 analysis of instantiation relations in this section, we look at the insertions into instantiation relations, first by rst label (section 4.2.1) and then by item (section 4.2.2). 4.2.1 analysis of instantiation relations by rst label we now turn to an analysis of only those relations that belong to the instantiation class. as shown in figure 4, items in this category often received causal, instantiation and specification markers in our study. to investigate whether this variance of interpretation within the pdtb instantiation class is reflected in the rst annotation, we analyse these instantiation relations separately by their rst label. hence, we repeat the analysis that is shown in figure 4, but we include only instantiation relations and separate them by the four rst classes. this allows us to see whether the insertions that we obtained in our experiment agrees with the reading that is annotated by rst annotators (elaborative or argumentative). figure 5 shows the distribution of inserted connectives in instantiation relations per rst class. 68 identifying elaborative and argumentative instantiations and specifications figure 5: distribution (%) of inserted connectives in instantiation relations per rst class. the first rst class under investigation, example, is considered to be similar to pdtb’s instantiation. as these two labels are direct correspondences to one another, we predicted that the participants in our study would also agree with this label and hence most frequently choose the connective as an illustration. as figure 5 shows, the instantiation marker as an illustration was indeed the most frequent connective chosen (41% of all insertions), but the distribution is quite broad: there were also many causal connectives (22% of inserted connectives) and many specification insertions (16% of insertions). the second class, general-specific, is considered to be inconsistent with pdtb’s class instantiation, since pdtb’s instantiation corresponds to rst’s example in rst, and pdtb’s specification corresponds to general-specific (also see section 2). we predicted that these relations might be genuinely ambiguous between an instantiation or a specification reading, and that we would see both instantiation and specification markers inserted. the results confirm this prediction: items in the general-specific class received an equal amount of instantiation and specification insertions (both 25%). however, the most frequently inserted type of connectives is causal, taking up 29% of all inserted connectives in this class, and we also see more conjunction relations than for other subgroups of instantiation relations. this group of relations therefore seems to be quite ambiguous. the third rst class under investigation, evidence, is generally considered to be a causal class. we therefore predicted that these relations may have two functions: being an example that also serves as evidence for a claim. we expected that both instantiation and causal markers would both be inserted often. figure 5 confirms this prediction: the instantiation marker is inserted most often (42% of all insertions), with causal connectives as the second most frequently inserted type (27%). finally, for the items annotated as pdtb instantiation and rst explanation-argumentative, it was also expected that both instantiation and cause markers would be inserted. this was indeed the case: 30% of inserted markers were as an illustration, and 35% were because / as a result. furthermore, we see a higher number of connectives expressing concession relations for this subgroup, which may reflect the causal aspect of these relations. again, the specification marker was also inserted relatively often, accounting for 18% of all insertions. in sum, we find that all of the subclasses had a substantial amount of instantiation, specification and cause interpretations. the results from our study on pdtb instantiation relations could be interpreted as evidence that instantiation items have both an elaborative and an argumentative function. however, it is also possible that rather than items being complex or ambiguous, subjects interpret a proportion of the instantiation items as expressing an elaborative relation, 69 scholman and demberg and another proportion presenting an argumentative relation. the next section provides more insight into this issue. 4.2.2 by-item insertions in instantiation items for the analyses in the previous section, all instantiation items were grouped together and the percentages represented an average amount of insertions (over all items) per rst class. this grouping obscures any possible differences between items of the same rst class. in this section, we look at the distribution of insertions per item. the distribution per item can be used to observe genuine ambiguity in the interpretation of some items, but also to derive the dominant interpretation for each item. figure 6 provides a detailed picture by showing the percentage of inserted connectives per rst class and item. every stacked bar on the x-axis represents an item, and the colours on the bars represent the inserted connectives. one way to analyse this data is to assign to each relation the label corresponding to the connective that was inserted most frequently by our participants, referred to as the dominant response (in figure 6, this corresponds to the largest bar per item). after assigning relations the label corresponding to the dominant response, we can calculate how many items received a dominant response that is the same as the pdtb or rst label. in other words, we can calculate agreement between the dominant response per item and the pdtb label and rst label. this result is reported in table 2. figure 6: distribution (%) of inserted connectives in instantiation relations per rst class and item. plots are arranged according to the amount of instantiation insertions. relation type agr. with pdtb agr. with rst instantiation – example 73 73 instantiation – general-specific 33 22 instantiation – evidence 67 27 instantiation – explanation-argum. 33 60 table 2: percentage agreement between the dominant response and the pdtb label instantiation, per rst class. 70 identifying elaborative and argumentative instantiations and specifications table 2 shows that the dominant response converges with the pdtb label relatively often for items in the rst classes example and evidence. for items in the class instantiation – general-specific, the dominant response does not converge with the pdtb label more, or less, often than with the rst label. this supports the hypothesis that these relations are ambiguous. the other common dominant response for these instantiation – general-specific items is causal. for items in the class instantiation – explanation-argumentative, the dominant response is most often causal, thereby converging with the rst label. the visualization of the by-item analysis in figure 6 reveals some interesting trends that we will discuss in more detail. items that elicited many instantiation insertions mainly belong to the rst classes example and evidence, with a few cases also occurring in the class explanationargumentative. these items are considered clear examples of the class instantiation, and we expect there to be a cue present that indicates to readers that the item is an instantiation relation. a closer look at these items revealed a common characteristic: often, a larger set is mentioned in the first argument, and one member of the set is explicitly referred to in the second argument. this larger set is referred to by a quantifier such as ‘many’, by plural noun phrases such as ‘glossy brochures’ and ‘larger department stores’, or by a combination of a quantifier and a plural noun phrase. this is illustrated in example (8), which is taken from the instantiation – example class and is presented in the same way as it was presented to participants. in arg1 of example (8), the set ‘glossy brochures’ is mentioned. arg2 then refers to one member of the set (‘one handout’) and gives a more specific example of the phenomenon described in arg1. (8) but that’s for the best horses, with most selling for much less – as little as $100 for some pedestrian thoroughbreds. even while they move outside their traditional tony circle, racehorse owners still try to capitalize on the elan of the sport. glossy brochures circulated at racetracks gush about the limelight of the winner’s circle and high-society schmoozing // one handout promises: pedigrees, parties, post times, parimutuels and pageantry. “it’s just a matter of marketing and promoting ourselves,” says headley bell, a fifth-generation horse breeder from lexington. wsj 1174 the items that elicited mainly causal insertions occurred predominantly in the rst classes evidence and explanation-argumentative (with two items in the class example). a common trait of these causal instantiations is that the first segment consists of a subjective utterance that can be interpreted as a claim and the second segment contains an argument for this claim, as in example (9), taken from the group instantiation – evidence: the speaker makes a claim in the first segment, and provides evidence for this claim in the second segment, as well as in the context following the second segment. the majority of the subjects interpreted this relation as causal (62%). (9) that done, ms. volokh spoke with rampant eloquence about the many attributes she feels she was born with: an understanding of food, business, russian culture, human nature, and parties. “parties are rather a state of mind,” she said, pausing only to taste and pass judgment on the georgian shashlik (“a little well done, but very good”). “if you are born to give parties, you give parties // even in russia we managed to give parties. in los angeles, in our lean years, we gave parties.” wsj 1367 71 scholman and demberg another characteristic of relations that elicited causal insertions is that the second segment can be interpreted as a result of the situation described in arg1, as in example (10), taken from the group instantiation – example. in this example, the instantiation reading can be inferred when the reader interprets the second segment as an example of how international competition is heating up. however, when the reader interprets these segments as occurring chronologically, he will get the reading that the situation in the second segment happens as a result of the situation in the first segment. indeed, 54% of insertions in this item were as a result (and 15% of insertions were because). (10) the goal of most u.s. firms – joint ventures – remains elusive. because the soviet ruble isn’t convertible into dollars, marks and other western currencies, companies that hope to set up production facilities here must either export some of the goods to earn hard currency or find soviet goods they can take in a counter-trade transaction. international competition for the few soviet goods that can be sold on world markets is heating up, however // west german companies already have snapped up much of the production of these items. seeking to overcome the currency problems, mr. giffen’s american trade consortium, which comprises chevron corp., rjr, johnson & johnson, eastman kodak co., and archerdaniels-midland co., has concocted an elaborate scheme to share out dollar earnings, largely from the revenues of a planned chevron oil project. wsj 1368 finally, certain items received many different types of insertions without showing a clear dominant response. manual inspection revealed that these items often revolve around topics of economics that typically require background knowledge about the stock markets. the lack of agreement in annotation of these relations may hence be due to participants not having enough background information to judge the relations in the text. given that even professionally trained discourse relation annotators are often not experts on the topic of the text that is being annotated, it is possible that this domain problem also affects the original pdtb and rst-dt annotations (also see martins, kigiel & jhean-larose, 2006; mcnamara, kintsch, songer & kintsch, 1996). 4.3 analysis of specification relations in this section, we look at the insertions into specification relations, first by rst label (section 4.3.1) and then by item (section 4.3.2). 4.3.1 analysis of specification relations by rst label for items from the pdtb class specification, the most frequently inserted connective type was not the marker that would be expected based on the pdtb label, but a causal marker (39%). the connective more specifically was the second most frequent type (25%) and as an illustration was the third most frequent type (13%). we again split up the dataset by rst labels for a more detailed analysis, see figure 7. the same predictions that held for instantiation items per rst class also hold for specification items: for items annotated as pdtb specification and rst example, we predicted that this disagreement between pdtb and rst annotations would be reflected in a similar split of inserted connectives by our participants. as figure 7 shows, this is indeed the case: subjects inserted the instantiation marker in 24% of the cases, and the specification marker in 22% 72 identifying elaborative and argumentative instantiations and specifications figure 7: distribution (%) of inserted connectives in specification relations per rst class. of the cases. somewhat surprisingly, we find a large proportion of causal insertions (27%) for these instances. similar to the findings in section 4.2.1, we find that relations for which pdtb and rst annotators did not agree on the instantiation or specification sense are ambiguous. items annotated as pdtb specification and rst general-specific received a nearly equal amount of specification and causal insertions (34% and 33%, respectively). this brings up the question whether these instances are in fact both elaborative and argumentative. this will be discussed in the by-item analysis in section 4.3.2 below. for items that received the pdtb specification and rst evidence label, we predicted that participants would insert a high amount of specification and cause markers. as figure 7 shows, nearly half of all insertions were causal (49%), while only 20% of insertions were more specifically. hence, participants seem to pick up on the same reading as rst annotators for these items. a similar pattern occurs for specification items that received the rst label explanationargumentative : 45% of the insertions were causal, and 23% were more specifically. these results indicate that naı̈ve subjects tend to interpret specification items as expressing a causal relation. the by-item analysis in the next section will show that items often received both types of insertions, and not only one or the other type. 4.3.2 by-item insertions in specification items figure 8 displays the distribution of inserted connectives in pdtb specification relations per rst class. from these distributions of answers per items, we again calculate dominant responses figure 8: distribution (%) of inserted connectives in specification items per rst class and item. plots are arranged according to the amount of specification insertions. 73 scholman and demberg relation type agr. with pdtb agr. with rst specification – example 20 27 specification – general-specific 40 40 specification – evidence 13 73 specification – explanation-argum. 27 67 table 3: percentage agreement between the dominant response and the pdtb label specification, per rst class. and their agreement with the original pdtb and rst labels. the results of this analysis are shown in table 3. the dominant response only rarely converges with the pdtb label. most agreement is observed for items on which the pdtb and rst annotations agree (rst general-specific). a common characteristic of items receiving a high number of specification insertions is that the first segment contains a reference to a general or vague concept, such as ‘one thing’ in example (11). (11) the ldp won by a landslide in the last election, in july 1986. but less than two years later, the ldp started to crumble, and dissent rose to unprecedented heights. the symptoms all point to one thing // japan does not have a modern government. its government still wants to sit in the driver’s seat, set the speed, step on the gas, apply the brakes and steer, with 120 million people in the back seat. for items in the class specification – example, we find that the dominant response does not converge with either of the labels very often. for items in the class specification – evidence and specification – explanation-argumentative, the dominant response is often causal. it thereby converges with the rst label (also shown in table 3). a manual analysis of these items showed a similar finding to that discussed in section 4.2.2: often, the first segment of the relation contains a subjective claim. it is likely that readers interpreted the more specific information in the second segment as evidence for the claim. 4.4 manual answers and double insertions when participants did not think any of the provided connecting phrases suited the relation, they were allowed to provide a manual answer. 2.5% of all insertions in instantiation and specification items were manual answers (a raw count of 37 insertions). there was no clear pattern in these manual answers: only a few items received manual answers, and these items received at most two manual answers. the type of manual answer was also variable: a few of them were connectives (although and however were inserted once, while was inserted three times), some seemed to be related to the syntax of the items (for example, as of, in which, and with), while others aimed at attributing information between the two arguments to a speaker (for example, saying, stating, adding). no clear conclusions can be drawn from these insertions. an additional 1.2% of the data consisted of ‘blank insertions’: subjects used the ‘manual answer’ option to not insert anything. as with the manual answers, there was no clear pattern: only a few items received a blank insertion, and there were no more than two blank insertions per item. 74 identifying elaborative and argumentative instantiations and specifications participants were given the option of inserting two connecting phrases if they thought that both phrases reflected the meaning of the relation. a double insertion can give us more insight into whether subjects thought that two senses held for a relation. for instantiation and specification items, 4.1% of all consisted of two connecting phrases. for most items that received a double insertion, only one answer consisted of a double insertion. for a few items, two or at most three participants provided a double insertion. looking at the amount of double insertions per participant, we find that only a few participants inserted multiple connectives (18 of 112 participants). the data on multiple insertions therefore does not allow us to draw any strong conclusions. this will be discussed further in the next section. 5. discussion and conclusion instantiation and specification are two of the most frequent implicit types of relations in the pdtb, making up 24% of all implicit relations in the pdtb. considering that these labels occur so frequently, the current study was designed to investigate how readers interpret these relations. more specifically, we examined whether readers interpret them as elaborative, argumentative, or complex, and searched for characteristics that are shared by relations which are interpreted to be argumentative. the results showed that both instantiation and specification items received many causal insertions: 28% of insertions in instantiation items and 39% of insertions in specification items were causal. we found that causal connectives were particularly prevalent in pdtb specification relations with rst evidence and explanation-argumentative annotations. importantly, a by-item analysis revealed that items rarely received only one type of insertion; rather, there were often two or more main types of insertions. these findings are consistent with a recent line of research that has focused on multiple readings of coherence relations: rohde et al. (2015, 2016) have shown that certain relations can have more than one single reading. the current study has provided more evidence for this hypothesis, showing systematic ways in which different types of discourse relations can occur simultaneously. in future work, we plan to investigate whether originally explicit instantiation and specification relations (that is, those marked explicitly with connectives such as for example, for instance or more specifically) also have an additional argumentative reading that is not annotated (and is possibly even harder to detect for annotators, due to the presence of the explicit connective). a manual analysis of items that elicited an instantiation or specification connective as the dominant response revealed that these items often contained one of the following characteristics: (i) a larger set is mentioned in the first segment, and one member of the set is explicitly referred to in the second argument, or (ii) the first segment contained a reference to a general or vague concept. by contrast, items that were often assigned a causal label shared one of the following characteristics: (i) the first segment contains a claim, and the second segment contains evidence or an argument for this claim, or (ii) the second segment can be interpreted as a result of the situation described in the first segment. these results go beyond previous work that has identified signals of instantiations and specifications and other elaboration relations (e.g., li & nenkova, 2016; taboada & das, 2013; vergez-couret & adam, 2012). for example, taboada & das (2013) have shown through manual annotation that relations from the rst class example can be signalled by individual words that indicate a relation without linking the two arguments (for example, the word explaining). general-specific relations are most often marked by entity features and lexical 75 scholman and demberg chains or overlap markers. taboada & das (2013) also find that explanation-argumentative and evidence relations often remain unsignalled. through computational corpus analysis, li & nenkova (2016) showed that first segments in instantiation relations are often shorter than other sentences, and the second segments are often longer. moreover, first segments contain fewer outof-vocabulary words than other sentences, and they contain more gradable adjectives (such as high, likely). the current study differs from these efforts in that the results are based on naı̈ve readers’ interpretations of relations, rather than expert judgments. 5.1 implications for the annotation of instantiation and specification items the results of the current study show that many discourse relations annotated in the pdtb as instantiation and specification also have an argumentative reading. this finding supports the hypothesis that instantiations and specifications are sometimes used to illustrate / specify a situation and to serve as an argument to a claim. pdtb does not annotate this argumentative function, but rather focuses only on the ideational relation between the arguments; that is, on the elaborational reading of the relation. by contrast, rst does classify these relations into separate elaborative and argumentative classes, which better matches the dominant responses of participants in our study for some relations. however, neither framework fully captures the double reading of these items that was reflected in the results. in particular, rst-dt does not make provisions for annotating more than one reading of a discourse relation. pdtb annotators, while allowed to annotate more than one label per relation, hardly ever make use of this option for the instantiation and specification relations. classifying instantiations and specifications as elaborative relations disregards the finding that many of these relations have an argumentative function as well. but classifying them as argumentative relations results in a disregard of their elaborative function (whether it’s instantiating or specifying). in order to make the annotation of these relations more reliable and to ensure that annotations reflect actual interpretations by readers, we recommend that both the ideational function of a relation (for example, that one segment provides an example of what is said in the other segment) and its rhetorical function (for example, that the example is used to justify a claim) be annotated (also see crible & degand, in press; gonzález, 2005; redeker, 1990). the way that the discourse-annotated corpora are currently structured, they do not contain multiple relation labels. but given that readers can obtain two different readings for the same relation, corpora would be descriptively more adequate if relations with multiple readings would receive double annotations. this could improve inter-annotator agreement as well: when using data annotated by only two coders, differences in the annotations might be interpreted as annotator error or disagreements. however, if at least one coder would annotate both senses, agreement would improve and the resulting annotations would better reflect the full meaning of the relation. of course, a double annotation process raises issues as well: it would need to be clear whether a single coder sees both senses, or whether different annotators have different interpretations that are alternatives to one another but can’t hold at the same time. an alternative to annotating two separate labels for reflecting the ideational and the rhetorical function would be to create separate sub-classes for causal instantiations and causal specifications. this solution would set these relations apart from purely ideational relations, and would as a result increase the label set. adding more subtypes to the set of discourse relations instead of adding double annotations would be in line with the traditional assumptions that only one relation 76 identifying elaborative and argumentative instantiations and specifications holds between two discourse segments. each of these solutions is likely to improve the descriptive adequacy of the labels for these relations, and thereby also the validity of the frameworks. 5.2 methodological remarks on crowdsourcing discourse relation annotations the crowdsourcing method used in the current study was shown to be relatively reliable for acquiring discourse annotations: the participants were able to insert the predicted connectives in filler items with high accuracy. furthermore, we showed that replicating the study with a new set of participants lead to the same results, providing evidence that our type of crowdsourced annotations are reliable and reproducible. existing corpora have mostly been annotated by a small set of trained, expert annotators. even after receiving a lot of training, agreement on the resulting annotations can be low (within and between frameworks). in sections 1 and 2, we have shown that agreement on implicit relations in particular is very low between frameworks (roughly 35% agreement). this disagreement can partially be attributed to a difference in operationalization: the way that a discourse annotation task is designed and formalised naturally influences the resulting data. at a general level, the method we propose here is similar to pdtb’s approach to discourse annotation: participants are asked to insert a connective that signals the relation between two segments. nevertheless, there are crucial differences between the two approaches, the most important difference being the use of naı̈ve, untrained individuals in our study, the lack of an annotation stage that labels the relation, and the much larger number of judgments in our study. additionally, our participants had the choice between only a small set of connectives, which are less ambiguous than many of the connectives that pdtb annotators could choose from. the participants also only had a few context sentences available to them (in the replication experiment they even had no context available), in contrast to pdtb annotators who can choose to read the entire text. the most crucial difference between traditional annotation tasks and the task described in the current study is the resulting data. the method of inserting connectives instead of assigning discourse relation labels does lead to more coarse-grained annotations compared to annotations of trained, expert annotators. however, our annotations have the potential of better reflecting the average readers’ interpretations, because they don’t rely on rules and biases introduced by the annotation frameworks that are supposed to increase inter-annotator agreement. moreover, it is easier, more affordable and faster to obtain many annotations for the same item via crowdsourcing than via traditional annotation methods. collecting a large number of annotations for the same item furthermore allows researchers to obtain a distribution of relation senses. this distribution can give researchers more insight into the readings of ambiguous relations, and into how dominant each sense is for a specific relation. the method can therefore be used to investigate comparable issues with other relational classes as well. we however also encountered limits in interpretability that are due to the experimental design. for instance, we can’t decide based on our results whether relations that received different insertions were genuinely ambiguous to a single participant (i.e. both readings were possible and the participant decided for expressing only one of them with a connective) or whether different participants had different interpretations of the same relation (but did not think that the connective that another participant inserted was suitable). even though our participants were provided with the option of inserting two connectives into a relation, they hardly ever made use of this option. it is possible that they avoided inserting a second connective because they only had one reading of the item; for ex77 scholman and demberg ample, they either interpreted the relation as elaborative or argumentative, but not both at the same time. however, it is also possible that motivation played a role. participants were only required to insert one connecting phrase; the second one was optional. since inserting a second phrase takes more time, participants might have neglected to do so, even if they interpreted multiple readings for some relations. if the double sense of relations is the focus in a future experiment, this can be solved by asking subjects to explicitly indicate that they don’t see a second reading. hence, it is possible to make it obligatory to choose two connectives, with the option of “no other connective fits” as a second connective (a similar approach is taken in the construction of a new version of the pdtb (webber et al., 2016)). another option woud be to present the items with one of the connectives, and ask participants to indicate whether the connective accurately expresses the relation. finally, we also found that some instances received a lot of very different annotations, and that these instances could likely be due to participants’ lack of domain knowledge in economics. we would like to draw attention to the fact that such a lack of domain knowledge might not only affect participants recruited via crowdsourcing platforms, but may also affect annotations of linguistically trained annotators who may not be very familiar with the textual domain. it is possible that, as a community, we are underestimating the effect of familiarity with a domain on discourse relation annotation quality and reliability. based on the high level of replicability between our original study and the replication study, we conclude that there is merit in the crowd-sourcing method, and believe it can potentially be used to create a corpus. related approaches using crowdsourcing for discourse relation annotation have been put forward by kawahara et al. (2014) and rohde et al. (2016). these studies also advocate crowdsourcing as it is fast and cheap, the resulting data are reliable, and the method can provide valuable insights that traditional annotation tasks do not. what separates the current method from previous work is that it is designed to be comparable to pdtb’s annotations, and the connectives were chosen to match pdtb’s classes. even though these connectives lead to more coarse-grained annotations, it is conceivable that the method can be extended to lead to more fine-grained distinctions in interpretations. in order to be able to apply this method to more general discourse annotation tasks, we recommend more research into which connectives can be added to represent more distinctions (e.g., temporal relations), as well as more research into the lower agreement for conjunction and contrast relations. if these issues are dealt with, we believe that the task presented in this paper has the potential to function as a method to create a discourse annotated corpus that embraces multiple interpretations. acknowledgements this research was funded by the german research foundation (dfg) as part of sfb 1102 “information density and linguistic encoding” and the cluster of excellence “multimodal computing and interaction” (exc 284). we are grateful to jacqueline evers-vermeul and jet hoek for fruitful discussions. 78 identifying elaborative and argumentative instantiations and specifications appendix a. figure 9: pdtb hierarchy (prasad et al., 2008) figure 10: rst-dt tagset (carlson & marcu, 2001) 79 scholman and demberg references artstein, r., & poesio, m. (2005). bias decreases in proportion to the number of annotators. proceedings of the conference on formal grammar and mathematics of language (fg-mol), (pp. 141–150). asr, f. t., & demberg, v. (2013). on the information conveyed by discourse markers. in proceedings of the fourth annual workshop on cognitive modeling and computational linguistics (pp. 84–93). biran, o., & rambow, o. (2011). identifying justifications in written dialogs by classifying text as argumentative. international journal of semantic computing, 5, 363–381. blakemore, d. (1997). restatement and exemplification: a relevance theoretic reassessment of elaboration. pragmatics & cognition, 5, 1–19. carlson, l., & marcu, d. (2001). discourse tagging reference manual. isi technical report isitr-545, 54, 1–56. carlson, l., marcu, d., & okurowski, m. e. (2003). building a discourse-tagged corpus in the framework of rhetorical structure theory. in current and new directions in discourse and dialogue (pp. 85–112). springer. carston, r. (1993). conjunction, explanation and relevance. lingua, 90, 27–48. cohen, r. (1987). analyzing the structure of argumentative discourse. computational linguistics, 13, 11–24. crible, l., & degand, l. (in press). reliability vs. granularity in discourse annotation: what is the trade-off? corpus linguistics and linguistic theory, . cuenca, m.-j. (2003). two ways to reformulate: a contrastive analysis of reformulation markers. journal of pragmatics, 35, 1069–1093. cuenca, m.-j., & marı́n, m.-j. (2009). co-occurrence of discourse markers in catalan and spanish oral narrative. journal of pragmatics, 41, 899–914. degand, l. (1998). on classifying connectives and coherence relations. in proceedings of the 1998 acl workshop on discourse relations and discourse markers (pp. 29–35). demberg, v., asr, f., & scholman, m. (2017). how consistent are our discourse annotations? insights from mapping rst-dt and pdtb annotations. arxiv e-prints, april. arxiv:1704.08893. gonzález, m. (2005). pragmatic markers and discourse coherence relations in english and catalan oral narrative. discourse studies, 7, 53–86. halliday, m. a. (1994). functional grammar. london: edward arnold, . hobbs, j. r. (1979). coherence and coreference. cognitive science, 3, 67–90. 80 http://arxiv.org/abs/1704.08893 identifying elaborative and argumentative instantiations and specifications hovy, e. h. (1995). the multifunctionality of discourse markers. in proceedings of the workshop on discourse markers (pp. 1–11). citeseer. hovy, e. h., & maier, e. (1995). parsimonious or profligate: how many and which discourse structure relations. unpublished manuscript, . hyland, k. (2007). applying a gloss: exemplifying and reformulating in academic discourse. applied linguistics, 28, 266–285. jasinskaja, k. (2013). corrective elaboration. lingua, 132, 51–66. jasinskaja, k., & karagjosova, e. (2011). elaboration and explanation. constraints in discourse, 4. kawahara, d., machida, y., shibata, t., kurohashi, s., kobayashi, h., & sassano, m. (2014). rapid development of a corpus with discourse annotations using two-stage crowdsourcing. in proceedings of the international conference on computational linguistics (pp. 269–278). knott, a., & dale, r. (1994). using linguistic phenomena to motivate a set of coherence relations. discourse processes, 18, 35–62. konig, e., & siemund, p. (2000). causal and concessive clauses: formal and semantic relations. topics in english linguistics, 33, 341–360. krippendorff, k. (1980). content analysis: an introduction to its methodology. sage. li, j. j., & nenkova, a. (2016). the instantiation discourse relation: a corpus analysis of its properties and improved detection. in proceedings of the north american chapter of the association for computational linguistics: human language technologies (naacl-hlt) (pp. 1181–1186). mann, w. c., & thompson, s. a. (1988). rhetorical structure theory: toward a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8, 243–281. martins, d., kigiel, d., & jhean-larose, s. (2006). influence of expertise, coherence, and causal connectives on comprehension and recall of an expository text. current psychology letters. behaviour, brain & cognition, 3. maschler, y., & schiffrin, d. (2015). discourse markers: language, meaning, and context. in d. s. d. tannen, h.e. hamilton (ed.), the handbook of discourse analysis, second edition, chichester, uk: john wiley & sons (pp. 189–221). mcnamara, d. s., kintsch, e., songer, n. b., & kintsch, w. (1996). are good texts always better? interactions of text coherence, background knowledge, and levels of understanding in learning from text. cognition and instruction, 14, 1–43. moore, j. d., & pollack, m. e. (1992). a problem for rst: the need for multi-level discourse analysis. computational linguistics, 18, 537–544. peldszus, a., & stede, m. (2013). from argument diagrams to argumentation mining in texts: a survey. international journal of cognitive informatics and natural intelligence (ijcini), 7, 1–31. 81 scholman and demberg prasad, r., dinesh, n., lee, a., miltsakaki, e., robaldo, l., joshi, a. k., & webber, b. (2008). the penn discourse treebank 2.0. in proceedings of the international conference on language resources and evaluation (lrec). citeseer. prasad, r., miltsakaki, e., dinesh, n., lee, a., joshi, a., robaldo, l., & webber, b. (2007). the penn discourse treebank 2.0 annotation manual. pusse, f., sayeed, a., & demberg, v. (2016). lingoturk: managing crowdsourced tasks for psycholinguistics. in proceedings of the north american chapter of the association for computational linguistics: human language technologies (naacl-hlt). redeker, g. (1990). ideational and pragmatic markers of discourse structure. journal of pragmatics, 14, 367–381. riezler, s. (2014). on the problem of theoretical terms in empirical computational linguistics. computational linguistics, 40, 235–245. robaldo, l., & miltsakaki, e. (2014). corpus-driven semantics of concession: where do expectations come from? dialogue & discourse, 5, 1–36. rohde, h., dickinson, a., clark, c. n., louis, a., & webber, b. (2015). recovering discourse relations: varying influence of discourse adverbials. in proceedings of the workshop on linking models of lexical, sentential and discourse-level semantics (lsdsem) (pp. 1–22). rohde, h., dickinson, a., schneider, n., clark, c. n., louis, a., & webber, b. (2016). filling in the blanks in understanding discourse adverbials: consistency, conflict, and context-dependence in a crowdsourced elicitation task. in proceedings of the 10th linguistic annotation workshop (law x) (pp. 49–58). sanders, t. j., demberg, v., hoek, j., scholman, m. c., torabi asr, f., zufferey, s., & eversvermuel, j. (submitted). unifying dimensions in discourse relations: how various annotation frameworks are related. corpus linguistics and linguistic theory, . sanders, t. j., demberg, v., hoek, j., scholman, m. c., zufferey, s., & evers-vermuel, j. (2016). how can we relate various annotation schemes? unifying dimensions in discourse relations. in textlink second action conference (pp. 110–112). sanders, t. j., spooren, w. p., & noordman, l. g. (1992). toward a taxonomy of coherence relations. discourse processes, 15, 1–35. scholman, m. c., & demberg, v. (2017). crowdsourcing discourse interpretations: on the influence of context and the reliability of a connective insertion task. in proceedings of the 11th linguistic annotation workshop (law) (pp. 24–33). scholman, m. c., evers-vermeul, j., & sanders, t. j. (2016). categories of coherence relations in discourse annotation: towards a reliable categorization of coherence relations. dialogue & discourse, 7, 1–28. 82 identifying elaborative and argumentative instantiations and specifications snow, r., o’connor, b., jurafsky, d., & ng, a. y. (2008). cheap and fast—but is it good?: evaluating non-expert annotations for natural language tasks. in proceedings of the conference on empirical methods in natural language processing (emnlp) (pp. 254–263). association for computational linguistics. stab, c., & gurevych, i. (2014). identifying argumentative discourse structures in persuasive essays. in proceedings of the conference on empirical methods in natural language processing (emnlp) (pp. 46–56). taboada, m., & das, d. (2013). annotation upon annotation: adding signalling information to a corpus of discourse relations. dialogue & discourse, 4, 249–281. vergez-couret, m., & adam, c. (2012). signaling elaboration: combining french gerund clauses with lexical cohesion cues. discours. revue de linguistique, psycholinguistique et informatique. a journal of linguistics, psycholinguistics and computational linguistics, . versley, y. (2011). multilabel tagging of discourse relations in ambiguous temporal connectives. in international conference on recent advances in natural language processing (ranlp) (pp. 154–161). webber, b. (2013). what excludes an alternative in coherence relations. in proceedings of the international conference on computational semantics (iwcs). webber, b., knott, a., & joshi, a. (2001). multiple discourse connectives in a lexicalized grammar for discourse. in computing meaning (pp. 229–245). springer. webber, b., prasad, r., lee, a., & joshi, a. (2016). a discourse-annotated corpus of conjoined vps. in proceedings of the 10th linguistic annotation workshop (law) (pp. 22–31). zufferey, s., & degand, l. (2013). annotating the meaning of discourse connectives in multilingual corpora. corpus linguistics and linguistic theory, 1, 1–24. 83 introduction background method participants materials procedure results reliability of the crowd-sourced annotation method analysis of instantiation relations analysis of instantiation relations by rst label by-item insertions in instantiation items analysis of specification relations analysis of specification relations by rst label by-item insertions in specification items manual answers and double insertions discussion and conclusion implications for the annotation of instantiation and specification items methodological remarks on crowdsourcing discourse relation annotations dialogue and discourse 4(2) (2013) 142-173 doi: 10.5087/dad.2013.207 linguistic tests for discourse relations in the tüba-d/z corpus of written german yannick versley versley@sfs.uni-tuebingen.de sonderforschungsbereich 833 universität tübingen 72074 tübingen, germany anna gastel anna.gastel@uni-tuebingen.de sonderforschungsbereich 833 universität tübingen 72074 tübingen, germany editors: stefanie dipper, heike zinsmeister, bonnie webber abstract discourse structure and discourse relations are an important ingredient in systems for the analysis of text that go beyond the boundary of single clauses. discourse relations often indicate important additional information about the connection between two clauses, such as causality, and are widely believed to have an influence on aspects of reference resolution. more so than for referential annotation, discourse relation annotation is rendered difficult by the absence of a general consensus on the underlying linguistic phenomena that should be targeted, as well as by a lack of strong predictions on the possible or permissible interactions between these phenomena. while it is sometimes claimed that the structuring of discourse is only weakly constrained and as a result capturing discourse structure and discourse relations will always result in poor reproducibility of the annotation task, we want to argue in this paper that an explicit notion of the relata of discourse relations allows to delimit annotation scope and to make use of theoretical accounts of the linguistic phenomena involved without giving up the goal of theory-neutrality that is essential in making sure that a given resource stays useful to a large community of users. in this article, we first present the general design choices that are to be made in the design of an annotation scheme for discourse structure and discourse relations. in a second part, we present the scheme used in our annotation of selected articles from the tüba-d/z treebank of german (telljohann et al., 2009). the scheme used in the annotation is theory-neutral, but informed by more detailed linguistic knowledge in the way of linguistic tests that can help disambiguate between several plausible relations. keywords: discourse structure, discourse relations, agreement, linguistic tests 1. introduction discourse information has been proven useful for a number of tasks, including summarization (schilder, 2002) and information extraction (somasundaran et al., 2009). while coreference corpora exist for many languages, and in large and very large sizes (frequently over one million words), the annotation of discourse structure and discourse relations has only recently gained the interest of the community at large. c©2013 yannick versley and anna gastel submitted 03/12; accepted 12/12; published online 07/13 linguistic tests for discourse relations in tüba-d/z the general idea of hierarchical discourse structure has a long history (polanyi and scha, 1984; grosz and sidner, 1986; mann and thompson, 1988; webber, 1988). mann and thompson’s rhetorical structure theory (rst), being the first to aim at a descriptively adequate account of real texts, has been the basis of annotated corpora targeting the analysis and generation of discourse structure, most notably the rst discourse bank (carlson et al., 2003) and similar corpora in other languages (cf. stede, 2004a, van der vliet et al., 2011), but has also drawn criticism regarding the cognitive plausibility of some of its aspects: in particular, sanders and spooren (1999) claim that rst does not separate between speaker intentions (which may not necessarily become shared knowledge) and coherence relations (which are instrumental for the understanding of a discourse); wolf and gibson (2005) take issue with the assumption that discourses are tree-structured, and propose to focus on the presence of coherence relations without any consideration of overall structure, whereas stede (2008) levels a more focused criticism at rst’s notion of nuclearity, which, as stede claims, encompasses criteria on different linguistic levels which are not always in agreement with each other. knott et al. (2001) propose a separation between low-level coherence relations (which can typically be signalled by conjunctions), and other means of structuring, which typically involve larger spans and make use of nominalization or discourse deixis. more recent work has set out to take into account the aforementioned criticisms, but also to make discourse annotation more efficient and predictable. most importantly, the authors emphasize the need to focus on a subset of the task that can be annotated reliably and that is at the same time informative with respect to a core set of discourse phenomena that are thought to be central, as is claimed by sanders and spooren (1999) or knott et al. (2001) for relations that are expressible through connectives. as a result, newer approaches such as the penn discourse treebank 2.0 (pdtb; prasad et al., 2008) use these fundamental ideas — formalized in the d-ltag formalism by webber (2004) — to define the scope of their annotation in theory-neutral guidelines for implicit (connective-less) and explicit (connective-bearing) discourse relations. the discourse annotation in the tüba-d/z treebank of german (see telljohann et al., 2009, and möller and naumann, 2009, for the syntactic and referential annotation layers) has a similar scope to the pdtb, and attempts to reach a useful balance between theory-neutrality on one hand and linguistic insights on the other hand, in particular from previous work in the frameworks of discourse lexicalized tree adjoining grammar (d-ltag) or segmented discourse representation theory (sdrt) as well as primarily descriptive treatments of single phenomena. by making strong assumptions on the types of (semantic or formal-pragmatic) entities that can be related by various discourse relations, it is possible to derive clearer criteria for linguistic tests, and it also makes it more obvious where cases are problematic due to violations of the basic assumptions (e.g., multiple propositions for one discourse segment, presence of implications as the relata of a discourse relation instead of the stated facts; cf. versley, 2008, and recasens et al., 2011, for similar considerations in coreference annotation). 2. defining the annotation task when laying down the guidelines for a discourse annotation task, one defines a formal model (i.e., text annotations and their specification) which relates parts of the text (discourse segments: usually sentences, clauses, or larger units) to each other. we will refer to these relations between discourse segments as discourse relations. such a formal definition of the task needs to answer (among others) the following two central questions: 143 versley and gastel 1. which relations between discourse segments are annotated, especially for relations that are implicit? 2. how is a given relation token described (in terms of label inventory)? both of these questions are inter-related in that, ideally, the relation inventory should offer a pertinent label for any pair of discourse segments that is to be annotated according to the first criterion; and conversely, if a given kind of relation between discourse segments is declared to be within the scope of our annotation, we would expect most or all of these relations to be part of the actual annotation. different annotations schemes start from a particular intuition to provide answers to these questions that are plausible in general, but which become problematic for some aspects of the annotation task. individual solutions for these questions are in disagreement with each other in particular areas, and also differ in which individual aspects are problematic. hence, a consideration of fundamental assumptions of existing annotation schemes, and the resulting answers to the above questions (which will be discussed in more detail in the following subsections), is essential. 2.1 underlying intuitions with respect to the first question, rst assumes a hierarchy of discourse segments that spans the whole text. discourse ltag starts from the notion of a discourse connective as an invariant phrase that signals a relation between two (tensed) clausal arguments. both the hierarchy assumption and the notion of connective arguments lead to discourse segments that may span multiple sentences. however, the two ideas lead to structures that are mutually incompatible in many cases. in the case of implicit relations, rst’s hierarchy assumption is still tenable in principle, while the idea of connective arguments is of limited use.1 in practice, existing annotation guidelines for implicit discourse relations either posit hierarchical structure (rst), ask annotators to mark all potential discourse relations that fulfill certain semantic criteria (wolf and gibson, 2005 and their discourse graphbank), or limit the annotation of implicit relations to neighbouring single sentences (pdtb). while systematic annotation difficulties cast doubt on rst’s assumption that there is a full hierarchy of discourse relations (cf. stede, 2008), neither the total absence or irrelevance of discourse structure (as postulated by wolf and gibson) nor the implicit assumption that implicit discourse relations never (or only very rarely) occur between single neighbouring sentences are satisfying alternatives for the question of discourse structure. possible solutions for the structural assumptions in discourse annotation will be examined in subsections 2.4 and 2.3. the issue of relation inventory, and associated problems, will be detailed in the following subsection 2.2. 2.2 number and kind of discourse relations one area where different descriptive proposals as well as current annotation schemes diverge consists of the choice and granularity of discourse relations. existing schemes reach from the most minimal model, containing just two relations (grosz and sidner, 1986), to one containing a taxonomy with 350 relations (hovy and maier, 1995). especially for schemes with many relations, a taxonomic organisation has the advantage of reconciliating the 1. webber, 2004, in her exposition of d-ltag, uses the idea of coordinating null connectives connecting adjacent discourse subtrees, which would essentially lead to a hierarchical structure. 144 linguistic tests for discourse relations in tüba-d/z need for differentiated relations with the necessity to capture most pertinent behaviour in a relatively small set of relations (near the top level of the hierarchy). however, even in the presence of taxonomic organization, the level of detail of an annotation scheme should be sufficiently motivated. research in rst after mann and thompson’s initial set of 20 relations seems to suggest that the exact set of relations used in a given project — whether annotation or computational generation of text — partly depends on the domain at hand and what distinctions are appropriate depends on the domain. this leads to nicholas’ (1995) claim that computational work using rst ‘has arbitrarily expanded the inventory of relations used,’ creating a major problem for rst as a (linguistic) theory, or the claim of knott and dale (1994) that ‘there often seems no motivation for introducing a new relation beyond considerations of descriptive adequacy or engineering expedience’. the problem of arbitrary inventory expansion can be avoided in schemes with strong restrictions on what can, or cannot, be a discourse relation. sanders et al. (1992) do this by introducing a relational criterion that limits the kinds of distinctions that should be made between discourse relations. sanders et al.’s relational criterion posits that categorization must refer to properties that concern the informational surplus that the coherence relation adds to the interpretation of the discourse segments in isolation, as opposed to the semantic contribution of the segments themselves. as a consequence, sanders et al. remove temporal semantics (which is inherent in the combination of tenses, or in the semantics of a conjunction linking two segments) from consideration as useful distinctions in discourse relations. the distinctions that sanders et al. do make are almost fully orthogonal and organize discourse relations into a multidimensional lattice rather than a hierarchical taxonomy: the first distinction that sanders et al. make is that of the basic operation, which can be causal (with a directed implication) or additive (with a — typically symmetric — conjunction between the two segments). other, binary, distinctions are source of coherence (which can be semantic, if it relates discourse segments based on their propositional content, or pragmatic, if it concerns the illocutionary forces of two discourse segments), basic versus non-basic order of segments in causal relations, and finally the polarity of a relation, which is negative for adversative relations and positive otherwise. another proposal for theory-independent, empirically driven criteria of what is a ‘psychologically real’ relation comes from knott and dale (1994): they say that evidence for a rhetorical relation is to be found in cue phrases that signal this relation. knott and dale, and knott (1996) give a criterion to find cue phrases, and use substitutability relationships for inferring taxonomical relations between cue phrases, hence constructing a fine-grained taxonomy while also avoiding the problem of arbitrariness that would otherwise arise. both proposals — knott and dale’s as well as sanders et al.’s — put limits on the coverage as well as the granularity of a scheme for discourse relations. while they are both well-motivated, it should be noted that the notion of cue phrases leads to a broader set of relations than sanders et al.’s relational criterion. a third proposal on criteria for discourse relation types stems from sanders and spooren (1999), who start from the idea that discourse relations are more or less related to either real-world entities, or to discourse-internal purposes (subject-matter vs. presentational relations: mann and thompson, 1988; moore and pollack, 1992). sanders and spooren argue that, among the discourse relations that are generally postulated, propositional relations such as cause relate states of affairs, whereas illocutionary relations such as evidence introduce a meaningful relation not between states of affairs but between illocutions (speech acts). they distinguish between the notion of an illocution (which is shared between speaker 145 versley and gastel and hearer) and a communicative intention (which may or may not be shared between speaker and hearer, and are mapped to illocutions by the speaker). because communicative intentions may be private to the speaker, and ultimately domain-specific, they argue, the third group of discourse relations such as rst’s preparation which are especially concerned with communicative intentions should not be part of the discourse annotation. proposals oriented at a narrow notion of coherence relations (such as hobbs, 1985, which inspired the annotation scheme of wolf and gibson’s discourse graphbank, or the original formulation of sdrt in asher and lascarides, 2003), as well as sanders et al.’s (1992) criterion, usually do not cover the group that sanders and spooren call illocutionary relations, even when these can be signaled by a cue phrase. as a result, corpora such as the penn discourse treebank, while subscribing to the idea of coherence relations in principle, include relations such as conjunction that would be outside the scope delined by sanders et al. (1992). in defining the scope of our annotation scheme, we follow the general approach of knott and dale, as well as sanders and spooren’s proposal that discourse relations be limited to propositional/illocutionary relations and not include (domain-specific or non-shared) communicative intentions. for the taxonomic organization of our annotation scheme, we start from the properties of sanders et al., but complement them with additional properties to cover other types of relations (e.g. different types of elaborating relations). we also used knott and dale’s approach of defining discourse relations in terms of cue phrases in the construction of explicit tests. such tests include — but are not limited to — substitution and/or insertion of selected cue phrases to improve the separability of discourse relations (section 4). 2.3 relations and hierarchy as in phrasal syntax, many theories of discourse structure posit a duality of constituents on one hand and head-dependent relationships on the other; in our case, (simple or complex) discourse segments and discourse relations may be seen to hold these roles. in rhetorical structure theory, the nucleus and satellite of mononuclear relations have a similar function to heads and dependents in syntactic theory. in particular, mann and thompson (1988) argue that deleting the nucleus of a mononuclear relation will result in an incoherent text as the significance of the material in the satellite(s) will not be apparent, whereas it is possible to delete, or replace, satellites of a discourse relation. in segmented discourse representation theory (asher and lascarides, 2003), a related but distinct set of notions is used: coordinating discourse relations are relations between discourse segments with the same topic, and subordinating discourse relations hold between a super-topic and a sub-topic. in difference to rst, coordinating and subordinating relations can link to the same discourse segment. for example, in (1), (c) is an elaboration of its super-topic (a), but it also has a narration relation to its sibling (b). (1) (a) john had a great meal. (b) he ate salmon. (c) he devoured lots of cheese. asher and vieu (2005) concern themselves with linguistic tests for distinguishing between coordinating and subordinating discourse relations. in their tests, they apply both the right frontier constraint (see webber, 1988, and later research) and a hypothesis that coordinating relations are 146 linguistic tests for discourse relations in tüba-d/z exactly those that are expressible within a syntactic coordination of the relation’s arguments (txurruka, 2003). using these criteria, they can show for the causal result relation as well as for narration that they are coordinating in general, although they also find examples of result that are clearly subordinating. in summary, postulating a hierarchy not only of larger and smaller discourse segments but also one across more-central/less-central may well pose problems as it sometimes deviates from the predictions derived from syntactic structure, and may become a hindrance if tokens of the same relation do not always have the same structural properties. however, minimal most-important segments, similar to the assumption of (semantic or lexical) heads in syntax, provide a useful abstraction when reasoning about larger discourse segments when the underlying assumption does hold. in our annotation scheme, we use the idea of subordination and coordination, as postulated in srdt. for a discussion of relation types versus structural properties (coordination and subordination), see the discussion in subsection 3.2. 2.4 delimiting the scope of discourse annotation the idea of a hierarchical discourse structure that reaches from elementary units (such as clauses) up to the complete document has been instrumental in the definition of rhetorical structure theory and other approaches, but has also been claimed to be problematic for reliable annotation. especially the larger structure of documents has been found to be a bad fit for rhetorical theories, and taboada and mann (2006) write that “in general, analysis of larger units tends to be arbitrary and unintuitive” (p. 430), noting that structures at larger levels of granularity (subsections or chapters) tend to be governed more by genre conventions than by the mechanisms that govern the small to medium level. this is in line with the sanders and spooren’s (1999) claim that some rst relations that link larger text units are not coherence relations, but model communicative intentions, which are not necessarily shared between speaker and hearer, and therefore can be much more arbitrary and unintuitive than the coherence relations that link smaller units (cf. subsection 2.2). let us examine a solution for setting the scope of discourse annotation which is based on these insights: instead of building complete trees, one subdivides the complete discourse into segments that realize one maximally complex illocution (which we call topic segments in our annotation scheme, cf. section 3, but which can also be found in several other annotation schemes, as noted in subsection 2.6). such an approach would allow multi-sentence segments as arguments of both implicit and explicit discourse relations, while maintaining a strong focus on coherence relations. in the remainder of this section, we look at existing analyses challenging (especially) the idea of hierarchy, and asking whether the respective phenomenon would lead to predictable problems for discourse annotation based on the idea of partial trees. one criticism of full hierarchical discourse structure which supports the above idea of ‘partial’ discourse trees can be seen in knott et al. (2001): knott et al. look at cases which they call resumption, where one larger segment of discourse is followed by an elementary discourse unit that takes up an entity mentioned in the previous segment and makes it the central topic of the next stretch of discourse. for these cases, which typically involve a discourse relation of objectattribute-elaboration (a subtype of rst’s elaboration relation), knott et al. argue that one has to choose between the overall connectivity of discourse relations on one hand and the presence of these elaboration relations on the other. knott et al. propose to segment the discourse into a sequence of 147 versley and gastel entity chains, each containing one or several discourse segments, as a means to solve the difficulties with resumption. such entity chains each share the same global focus, and they have one top nucleus which is the top nucleus of one of the discourse segments of that chain. a second source of discourse relations that would cross a hierarchical structure are discourse adverbials — discourse connectives that take their second argument anaphorically rather than structurally, and which also may occur crossing hierarchical structure or a complex discourse segment (in the sense of complex illocutions, or knott et al.’s entity chains). several researchers such as asher and lascarides (1998) and webber et al. (2003) spell out discourse adverbials as deriving their second argument not from structural composition mechanisms (as is the case with conjunctions such as but or because), but from the resolution of a presupposed proposition or situation argument. besides cutting across structure in rare cases, discourse adverbials exhibit other properties related to presupposition resolution that can make their annotation problematic. one such problem is vagueness (when the second argument of the discourse adverbial is accommodated rather than textually specified). in other cases, multiple discourse adverbials can introduce multiple concurrent relations for the same two arguments — see example (2), from stede, 2004b. it is also possible that the anaphoric argument of a discourse adverbial is either resolved within the same elementary discourse unit or accommodated to something that is not a discourse unit, as in examples (3) and (4), due to webber et al. (2003): (2) therefore, we then also went to a bigger mountain. (3) every person selling “the big issue” might otherwise be asking for spare change. (4) john just broke his arm. so, for example, he can’t cycle to work now. in the case of (4), for example is accommodated to be an example for all the other things that happened as a result of john breaking his arm; the latter is not at all realized in the text but is inferred by the reader. in our annotation guidelines, we pay special attention to these two phenomena, which tended to confuse annotators in the beginning: cases of resumption, where referential cohesion exists between discourse segments that are otherwise unrelated, are specially marked in the annotation to make clear they behave differently. in the case of discourse adverbials, annotators are advised to disregard their contribution if the anaphoric argument is vague, part of the same discourse unit, or inaccessible structurally. for details on this, see section 3.1. 2.5 information-theoretic notions of discourse structure to describe discourse segments, some approaches (notably, mann and thompson, 1988, grosz and sidner, 1986, or sanders and spooren, 1999) use intentional notions, modeling discourse segments as complex illocutions. an alternative approach for this problem, or possibly a group of alternative approaches, can be seen in the use of information-structural notions to describe discourse segments. such approaches were proposed both by researchers primarily interested in information structure (roberts, 1996; büring, 2003) and as a means to spell out the notion of topic as it is postulated, e.g., in sdrt (asher, 2004). 148 linguistic tests for discourse relations in tüba-d/z assuming an entity as a topic potentially has useful properties – on one hand, you can postulate interaction between the topic entity and the resolution of pronouns, assuming that the topic entity (or ‘global focus’) is a good candidate for pronominal reference (see also grosz and sidner, 1986, or in part knott et al., 2001, who claim such a global focus for larger discourse segments); on the other hand, subtopic relationships could be spelled out in a straightforward fashion in terms of semantic relations between entities (john – john’s hair – john’s hair care). the notion of discourse topic as a question was formulated in more detail by van kuppevelt (1995), who assumes questions as structure-building mechanism in discourse: questions are licensed by previous discourse segments (or other shared perceptions, such as visually or audibly salient happenings), which he calls feeders; questions are then answered through new discourse segments (which can give partial or complete answers). discourse structure arises through patterns of a discourse segment ‘feeding’ multiple questions, or subquestions arising from a discourse segment that in turns answers a (superordinated) question. van kuppevelt says a topic-constituting question has a satisfactory answer when it is completely answered and none of the answer segments feeds unanswered questions. the connection between intonation and the topic-constituting question in the discourse context, which we will call simply question under discussion (qud), has been investigated, among others, by roberts (1996), büring (2003) and in part by asher (2004). the intonation of a sentence can have a focus (f, a-accent, rheme focus), which indicates that the constituent is an answers to the question under discussion, and a contrastive topic (ct, b-accent, theme focus) which marks constituents that may differ in a sibling qud.2 contrastive topic marking indicates the presence of other, related quds in the discourse model (büring), the presence of implicit sub-questions (büring) or may signal a partial or indirect answer to a qud (asher), as in example (5): (5) q: what did the pop stars wear? a: the [ct female] pop stars wore [f caftans]. a′: the tuneless wannabees wore [f caftans]. (the alternative answer a′ is read to constitute a full direct answer, accommodating the assumption that all popstars were tuneless wannabees.) asher (2004) remarks that, while he generally expects ct to mark the discourse topic of an elementary discourse segment, not all relations do, or even can, receive contrastive topic marking, and that the approaches put forward by van kuppevelt and by büring cannot make any predictions regarding coordinated relations such as continuation or narration. in our annotation guidelines, the notions of qud and of information structural marking are used to implement tests for discourse relations where contrast pairs, or the surrounding discourse, play an important role. section 4 discusses this in greater detail. 2. the most common nomenclature for these intonation curves is focus and contrastive topic as used by büring (2003), but it is also common to find the terms a-accent and b-accent following jackendoff (1972) or theme focus and rheme focus as suggested by steedman (2000), respectively theme kontrast and rheme kontrast in the work of vallduví and vilkuna (1998). while the exact predictions on the link between semantic structures and intonations differ, the respective terms refer to the same intonation contours originally pointed out by jackendoff. 149 versley and gastel 2.6 annotated corpora besides our own annotation effort, a substantial number of corpora exists, in multiple languages. many of these follow the general ideas of rhetorical structure theory, as the english rst discourse treebank of carlson et al. (2003), the german potsdam commentary corpus of stede (2004a), as well as the rst spanish treebank of da cunha et al. (2011), multiple brazilian portuguese corpora (pardo and seno, 2005; pardo and nunes, 2008; cardoso et al., 2011), and the dutch corpus of van der vliet et al. (2011). in terms of relation inventory, they either using the ‘profligate’ approach of hovy and maier (1995) with close to 100 relations in the case of the rst discourse treebank, or use a close variant of mann and thompson’s (1988) original rst annotation scheme. many other corpora follow the ideas of the penn discourse treebank (prasad et al., 2008), including corpora in arabic (al-saif and markert, 2010), czech (mladová et al., 2008), hindi (kolachina et al., 2012), italian (tonelli et al., 2010) and turkish (zeyrek et al., 2010). most of these corpora presently only cover discourse connectives and explicit discourse relations, whereas implicit discourse relations are not yet part of the annotation. two corpora exist which are inspired by ideas from segmented discourse representation theory: the english sdrt corpus of reese et al. (2007), and the french annodis corpus of pérywoodley et al. (2011). for these three ‘traditions’ of discourse annotation, annotation guidelines are available publically (carlson and marcu, 2002; reese et al., 2007; pdtb research group, 2008). in the contained descriptions of discourse relations, we see a tendency to move from carlson and marcu’s mostly informal and example-based treatment of rst relations to a more formal treatment by reese et al. or the pdtb manual, which also use additional examples where appropriate for discussing special cases and often make explicit reference to propositions (pdtb) or main eventualities (sdrt) as the relata of coherence relations. several corpora use the idea of partial trees in their annotation, with slightly differing theoretical backgrounds: the dutch corpus of van der vliet et al. (2011) contains subtrees for each conversation move. conversation moves, in their annotation scheme, are zones in each document which realize one genre-specific communicative purpose. the definition of conversation moves is genre specific: starting from a rhetorical purpose that is common to all texts of one particular genre, an analyst would identify possible rhetorical functions for segments that constitute a discourse move. because it depends on the analysis of a whole genre, move analysis can accurately capture admissible complex illocutions, but would be less suitable for a heterogeneous mixture of different genres such as that found in (general) newspaper text. the sdrt-inspired annodis corpus of péry-woodley et al. (2011) combines a macrostructure level, which uses textual cues and paragraph boundaries, with a level of microstructure which uses an inventory of coherence relations “mostly common to all discourse theories”. it is interesting to note that the idea of a text consisting of smaller segments that are mostly independent from each other has been postulated independently of discourse annotation, for example by hearst (1997). however, only recent annotation proposals for discourse structure have addressed the question how this distinction can be made precise enough. the two principal sources of information for a segmentation into topic segments are referential cohesion on one hand, and intentional notions on the other hand. a naïve application to referential cohesion would suggest that topic segments correspond more or less closely to the extents of referential or lexical chains, as postulated among others by hearst (1997). however, work by knott 150 linguistic tests for discourse relations in tüba-d/z et al. (2001) proposes that referential (and lexical) cohesion can normally occur across the boundaries of larger segments, and that a better indicator would be the shift of referential focus (i.e., the entity that is most salient in a discourse segment) rather than the presence or absence of mentions of a given concept or referent. from an intentional perspective, the most appropriate notion for a topic segment would be that of a complex illocution, which would imply that one topic segment corresponds to one specific complex speech act, such as explaining the history of a specific artifact, or describing the consequences of a planned construction.3 in our own annotation scheme (see subsection 3.1), we use complex illocutions as the motivating underpinning for topic segments, and introduce a special type of marking (called transitional edus) to mark cases where referential cohesion is at odds with these complex illocutions. 3. discourse annotation in tüba-d/z in the tüba-d/z, a two-pronged approach has been chosen for the encoding of discourse relation information. on one hand, a sample of ambiguous discourse connectives (temporal subordinators as well as conjunctions) is disambiguated according to the discourse relation that they realize, aiming at a relatively precise account of the variability that the respective connective affords. on the other hand — and this is the part that the present article is concerned with — complete newspaper articles are annotated with discourse structure. the discourse structure consists of a segmentation of a complete discourse into topic segments (stretches of coherent texts that realize one high-level goal of a writer such as ‘describing the attitude of retailers towards genetically modified products’), and organizes the text within a topic segment into a hierarchy of discourse units. this hierarchy of elementary (and composed) discourse units is realized through coordinating and subordinating relations, including relations between elementary discourse units or discourse spans that are implicit in that they are not marked by a discourse connective. the main tüba-d/z corpus (telljohann et al., 2009) contains 1.1 million word tokens with syntactic, named entity, and morphological annotation, as well as information on referential cohesion in the form of anaphora/coreference annotation, which makes it an appealing target as a text source for the additional annotation of discourse structure and discourse relations. from a purely technical point of view, the existing corpus contains rich linguistic annotation that can be exploited in conjunction with the discourse relations. from the point of view of text selection, the newspaper die tageszeitung offers a good balance between argumentative and descriptive writing, usually providing both reporting on current events as well as background information and commentary. for the annotation of discourse relations, a subset of the articles was manually selected according to criteria such as length (excluding extremely short newswire-style reports as well as one or two extraordinarily long articles) and topical coherence (excluding summary articles that report a multitude of different items with only two or three sentences per item). the current version of the discourse structure annotation, which is publically available as part of release 8 of the tüba-d/z, constitutes a subcorpus of 41 texts, with 21 817 word tokens containing 1 458 discourse relations, with about 28.8 sentences/article (against 20.6 sentences/article 3. a complex illocution, such as “explain the history of a specific artefact” or a formulated questions such as “what happened to the artifact?” are more clearly delimited than general topic description such as “the history of the artifact”: the latter could also admit a discussion of other’s critical views on the artifact at top-level, while the speech-act verb explain or the question “what happened to x?” would be understood not to include commenting. in comparison to the notion of conversation moves, complex illocutions do not presuppose a pre-established account of a genre’s structuring conventions. 151 versley and gastel on average in the complete tüba-d/z, which includes brief newswire-style reports), and about 3.5 topic segments per article. of all discourse relations, 557 (38%) are intra-sentential (between edus in the same sentence), whereas 367 (25%) involve at least one of arg1 or arg2 spanning multiple edus. in 182 cases (12%), at least one of arg1 or arg2 spans multiple sentences.4 3.1 elementary segments and topic segments elementary segments (also: elementary discourse units, edus) are the smallest units of text that can be the argument of a discourse relation and ideally correspond to exactly one main eventuality or one asserted proposition. our annotation scheme is oriented at existing guidelines for english (carlson and marcu, 2002; reese et al., 2007) and german (lüngen et al., 2006).5 for the largest part, the guidelines for the segmentation of edus are defined in syntactic terms, granting edu status to tensed clauses (matrix clauses as well as subclauses which are not centerembedded), but also to eventualities introduced by (causal, temporal or attributional) prepositional phrase adjuncts (6a), appositions introduced by right dislocation (6b), non-restrictive relative clauses as well as purpose clause adjuncts to tensed clauses. one area where semantic considerations play a role is the separation of reported proposition arguments, where verbs count as communication verbs (with discourse segment status for their arguments) if they allow direct speech or allow fronting of the reported content.6 (6) (a) [nach nunmehr über sechs wochen erfolgloser luftangriffe] [gibt es im nato-hauptquartier niemanden mehr, der das scheitern bestreitet.] [after over six weeks of unsuccessful air raids,] [there is no one left in nato headquarters who disputes the demise.] (b) [er sollte unseren heimischen markt aufmischen,] [das erste produkt in deutschen läden, das genmanipulierten mais enthält.] [it was to stir up our domestic market,] [the first product in german stores to contain gm corn.] as an upper boundary for discourse annotation, texts are divided into so-called topic segments. for every topic a title is added which summarizes that topic. all text within one topic supports an answer to the same high-level question under discussion (qud). this qud is the most important tool for the identification of topic segments: it should be a question that contributes to the overall topic of the article, while, for our text type of medium-length newspaper articles, a topic segment does not have a title/theme that is synonymous to that of the article. a surface cue for annotators is presented by the paragraphs of the article, since a topic boundary usually coincides with a paragraph boundary (but not vice versa – topics are usually multi-paragraph units of text). in our annotation, topic segments are therefore exactly one hierarchy level under the topic of the complete article. 4. as an example for a complex structure without multi-edu arguments, consider a discourse graph where one superordinate edu has relations to multiple subordinate edus, but where that larger segment is not related with another discourse segment. many such simpler structures that a topic segment can have would be annotated without the involvement of relations between multi-edu spans. 5. in comparison to the proposal of lüngen et al., we do not segment embedded/parenthetical material since this would needlessly create discontinuous edus. in the case of right dislocation, we make reference to the existing syntactic structure of the treebank instead of specifying punctuation heuristics. 6. tüba-d/z, sentence 2890; tüba-d/z, sentence 5736. 152 linguistic tests for discourse relations in tüba-d/z as an example, consider the following topic segmentation for an article about the nato’s plans in kosovo: • t0: opinions on the situation in kosovo qud: what do important people say about the situation in kosovo? • t1: air strikes in kosovo qud: what happened to the air strikes in kosovo? • t2: plans for ground troops qud: what plans are there for the deployment of ground troops? • t3: role of the red/green minority qud: what role does the red/green minority in parliament play? in the example, merging t1 with t2 would yield a very general topic about plans of the nato, which corresponds to the overall topic of the article and is therefore uninformative. conversely, splitting t0 into separate topic segments with naumann’s opinions on the situation and clinton’s opinions on the situation would clearly yield non-maximal questions under discussion, to which t0 is preferrable. usually, topic boundaries coincide with the absence of discourse relations, and generally less cohesion. there is, however, a notable exception to this generalization: in the annotation process for our corpus, we would frequently encounter cases where the author explicitly bridges two segments that treat different topics, with a discourse segment where one part of the sentence contains discourse-old information that explicitly creates cohesion with the previous topic segment, whereas the other part of the sentence contains novel information introducing the new topic segment. in these cases, annotator’s topic boundaries would frequently differ by a single sentence. in order to allow annotators to make the disconnect between referential cohesion and topic segments explicit, we introduced transitional discourse units — elementary discourse segments that are at the start of a topic segment and contain referential cohesion to the previous topic segment, yet contribute to the illocutionary purpose of the new discourse segment. the marking of transitional edus serves the dual purpose of highlighting the common structure of such segments as well as preventing confusion over anaphoric reference of pronouns or discourse adverbials crossing the topic boundary. additionally, the transitional edus can be the argument of a topic crossing relation that further serves the purpose of a cohesive transition of one topic to another. in the context of example (7),7 the referential expression vor diesem hintergrund (‘in this context’) creates referential cohesion to the previous topic segment even though the actual contents of both topic segments are quite different. (7) [5.0 bereits seit jahren sind sie die großen hoffnungsträger:] [5.1 die jungen existenzgründer.] [6.0 auf ihnen ruht die erwartungsvolle aufmerksamkeit von politikern und wirtschaftsexperten] [. . . ] t1: die deutschen existenzgründertage [tr-edu12.0 vor diesem hintergrund fanden an diesem wochenende die zweiten “deutschen existenzgründertage” statt,] [12.1 mit denen die träger, wirtschaftssenator branoner und 7. tüba-d/z, sentences 1979ff. 153 versley and gastel brandenburgs wirtschaftsminister dr. dreher, berlins “gründungskompetenz bundesweit transportieren” wollen.] [5.0 since years they are the beacon of hope:] [5.1 the young business founders.] [6.0 they have the keen interest of politicians and business experts.] [. . . ] t1: the german founders’ days [tr-edu12.0 in this context, the second “german founders’ days” took place last weekend,] [12.1 which the sponsors, economics senator branoner and brandenburg economic minister dr. dreher, wanted to “transport berlin’s founder competence at the national level”] 3.2 discourse structure and discourse relations the main content of the discourse annotation consists in discourse relations, which link two arguments that can consist either of a single edu or a span of edus (generally corresponding to a complete subtree in the discourse structure). discourse relations can be either subordinating (in which case the second argument is subordinated under the first), or coordinating (in which case the first argument comes before the second in surface order). as every relation type is marked as either coordinating or subordinating in the annotation scheme (cf. table 1), the annotated graph of relations specifies a hierarchical discourse structure as postulated by asher and vieu (2005). in line with asher and vieu, the tüba-d/z discourse annotation allows a discourse unit to be linked by both a coordinating relation (to a sibling in the discourse hierarchy) and a subordinating relation (to its superordinate in the discourse hierarchy). because a discourse segment can possess both a link to the larger discourse segment that subordinates it and coordinating relations to its neighbouring siblings, the graph of all relations is not a tree in the graph-theoretic sense, but the subgraph of all subordinating relations would fulfill that criterion. as an informal part of the annotation, our annotation tool allows annotators to indent elementary text segments in order to reflect subordination depth. the hierarchy of discourse segments in the annotation also provides a holistic view on discourse relations regardless whether they are syntactically mediated (through embedding or by subordinating or coordinating conjunctions), marked by discourse adverbials, or completely unmarked. using such a coherent view, the annotation of unmarked relations between larger segments is more straightforward due to constraints found in the structural context. in converse, our view of hierarchy entails that some, but not all relations due to discourse adverbials are part of the annotation. in example (8) below, we can see that the larger span α that occurs as the argument for the commentary relation constrains the possible relation targets for the edus inside (edu boundaries are marked as vertical lines).8 (8) [α die gute nachricht: die weltbevölkerung wächst inzwischen langsamer als in den vergangenen jahrzehnten.| die schlechte nachricht: erreicht wird diese entlastung der ökologischen und sozialen systeme nicht nur durch fortschritte bei der geburtenkontrolle,| sondern auch durch eine sterblichkeitsrate, die zum erstenmal seit 40 jahren wieder ansteigt. |. . . ] [β diese entwicklung zeigt nach angaben von lester brown, einem der autoren der studie, das “versagen unserer politischen institutionen”.] [α the good news: the world population is growing more slowly than in past decades.| 8. tüba-d/z sentences 8429ff.; see figure 1, in section 5, for the larger context. 154 linguistic tests for discourse relations in tüba-d/z the bad news: this unburdening of ecological and social systems is not only achieved by progress in family planning,| but is also due to a mortality that growing again for the first time since 40 years. | . . . ] [β this development shows, according to lester brown, one of the study’s authors, a “failure of our political institutions.”] commentary(α,β) the full list of relations can be seen in table 1. in some cases, different relation types have somewhat similar semantics but different properties regarding discourse structure, such as result (coordinating) and explanation (subordinating), or attribution (subordinating, non-veridical, reported content is in arg2) and source (coordinating, veridical, reported content is in arg1).9 to illustrate these distinctions, consider example (9) and (10) for the distinction between result and explanation:10 (9) [1 private unternehmen dürfen die telefonbücher der detemedien nicht ohne deren erlaubnis zur herstellung einer telefonauskunfts-cd verwenden.] [2 die beklagten unternehmen müssen den vertrieb der info-cds sofort einstellen.] [1 private companies may not use the telephone books by detemedien without its permission for the creation of directory assistance cds] [2 the defendant companies must cease the distribution of their information cds immediately.] result-cause(1,2) (10) [3 taxifahrer sind als kolumnenthema eigentlich tabu,] [4 weil sie als “weiche angriffsziele” gelten.] [3 taxi drivers are normally a taboo topic for a newspaper column,] [4 because they are considered “soft targets”.] explanation-cause(3,4) in the case of (9), both arguments are necessary for coherence, so the coordinating result relation is chosen. in the relation example from (10), the fact in edu 4 is considered the cause for edu 3, but 4 mostly contributes background information and is not important in its own right. in addition to these distinctions, the annotation process revealed that a coordinating variant of two more relation types present a useful addition, as the corresponding relation instances both pass the coordination criterion of txurruka (2003) and asher and vieu (2005) and considerations of the surrounding discourse structure. concessionc is a counterpart of the concession relation and actually has a somewhat different profile from the subordinating concession relation (see subsection 4.1). restatementc is a counterpart of the restatement relation and is typical for a restatement where the first part of the restatement is overly vague or metaphorical, such that neither part can be left out. these cases very often occur in a coordination with und (and), which is a further indication that the relation is coordinating rather than subordinating. considering the critique of stede (2008) on rst’s notion of nuclearity11 and the large number of relations in the annotation manual of carlson and marcu (2002) – 23 of the 53 mononuclear relations in their annotation scheme have multinuclear counterparts – it is an interesting question whether we will forcibly end up duplicating every relation in one subordinating and one coordinating version. indeed, we see in table 1 that this is already the case for almost all of the relations in the contingency, temporal and reporting groups. for the symmetric relations from the comparison group, as well as 9. see hunter et al. (2006) for a rationale on distinguishing evidential from other relations in discourse annotation. 10. tüba-d/z sentences 2536ff. and 8870ff. 11. this partially also applies to the critique of moore and pollack, 1992, who advocate a separation of semantic relations and discourse progression. 155 versley and gastel for continuation and conjunction from our continuative subgroup, it is clear that a subordinating counterpart is not possible. from an information packaging point of view, having a summary or commentary without the commented or summarized content is not necessarily sensible. a similar argument could be made for the instance and background relations (where the discourse segment subordinated by these relations takes up the main event or another referent from the superordinate discourse segment). table 1 also contains a comparison of our relation inventory to the counterparts of the rst discourse treebank, the sdrt corpus of reese et al. (2007), and the penn discourse treebank. in some cases, such as our conjunction relation in comparison to list/conjunction/equivalence in the penn discourse treebank, or our commentary relation in comparison to the relations by rst, we found it more attractive to have very few ‘weak’ discourse relation types, similar to the simplifications applied in the leeds arabic discourse treebank. 4. linguistic properties and tests an annotation scheme needs to make explicit the categories and criteria used in the annotation. such an explicitation is useful in general for users of the annotated corpus to understand individual annotations, but it is also crucial for large-scale corpus annotation where multiple annotators need to reach consistent annotation decisions independently of each other. for this operationalization of annotation guidelines, an informal description of the annotated categories is often complemented by illustrating examples, as in the manuals of carlson et al. (2003) or reese et al. (2007). for difficult cases, the introduction and use of disambiguating heuristics, or linguistic tests is needed – usually, an insertion or substitution operation together with an estimate of which aspects of the meaning may change in the substitution, and which cannot. using multiple linguistic tests can be helpful as long as one is clear about the linguistic properties that these tests are supposed to verify, which means that it is often appropriate to go beyond a pre-theoretic understanding of the testing heuristics themselves. one of the most basic tests with respect to discourse markers and discourse relations is the insertion of different discourse markers between two units of discourse, as used at least by sanders et al. (1992) and used as a principal tool of investigation by knott and dale (1994) to create their shallow taxonomy of discourse connectives. knott (1996) discusses how substitutability in context can be used to abstract from cue phrases (or discourse connectives) to feature values (i.e., linguistic properties) by assuming, e.g., a feature that has different values for two discourse connectives when no context exists where they are substitutable. in knott’s method, a pair of discourse segments is given (which occurs in a context with one original discourse marker) and a given discourse marker can be judged as either substitutable with the original one linking the two (in which case the intended meaning stays the same), it may be ungrammatical or incoherent using the new marker, or the result may be grammatical and coherent, but carry a new meaning, in which case the replacement is not deemed substitutable. knott’s methodology makes very few assumptions and hence is very suitable as a starting point for a theory-neutral account, and knott notes that the resulting findings correlate with (intuition-guided) distinctions found by sanders et al. (1992). in practical use, it is possible that features are underspecified, or that intuitively plausible discourse relations sometimes involve prototypicality or family resemblance effects. this means that 156 linguistic tests for discourse relations in tüba-d/z our corpus sdrt rst pdtb contingency causal c result-(cause,enable) result cause cause:(reason,result)b c result-epistemic result evidence pragmatic cause:justification c result-speechact result rhetorical-question cause:(reason,result)b s explanation-(cause,enable) explanation result cause:(reason,result)b s explanation-epistemic explanation explanation-argumentative pragmatic cause:justification s explanation-speechact explanation elaboration-additional cause:(reason,result)b conditional c consequence consequence consequence condition c alternation alternation disjunction alternative s condition consequence condition condition denial c concessionc contrast antithesis/preference contrast s concession contrast concession concession s anti-explanation — circumstance cause:reason expansion elaboration s restatement elaboration elaboration restatement:specification s instance/v elaboration example instantiation s background background circumstance synchronous interpretation s summary elaboration conclusion restatement:generalization s commentary commentary comment/evaluation/ conjunction interpretation continuative c continuation continuation joint conjunction c conjunction continuation/ joint/list list/conjunction/ parallel equivalence temporal c narration narration sequence asynchronous:succession s precondition precondition inverted-sequence asynchronous:precedence comparison c parallel/v parallel analogy conjunction c contrast contrast contrast contrast:juxtaposition reporting s attribution attribution attribution-na — s source source attribution — sdrt refers to the corpus annotation guidelines by reese et al. (2007); rst refers to the annotation guidelines by carlson et al. (2003). these annotation schemes do not necessarily reflect other schemes inspired by the same theories. a) attribution-n is used in carlson et al. (2003) with the same semantics as our attribution in cases where the attribution strictly cannot be veridical. b) the penn discourse treebank uses syntactical criteria to distinguish between cause:reason and cause:result, whereas we, like reese et al. (2007), use criteria related to surface order for the distinction between result and explanation. table 1: relation inventory in comparison with other schemes 157 versley and gastel it is useful both to reason explicitly about the connection between discourse relations and features, and to complement substitution tests with other kinds of tests where appropriate. as an example for a non-trivial relationship between discourse connectives and properties of coherence relations, consider causal discourse relations, where the prototypical case is causation between events. in this prototypical case, causation holds between non-action events, corresponds to a temporal succession, and the second event would be predictable from the first. in such a case, both temporal connectives (when/after) and causal connectives (because/since) would be appropriate in the context. in non-prototypical cases such as piecemeal causation (in which one process influences another, while being co-temporal to it), reasons for actions, or logical causation between propositions, not all of these criteria hold, and a more detailed assessment of the participating properties is necessary to reach a crisp boundary for these relations (see subsection 4.2). table 2 summarizes both the grouping of discourse relations in our annotation scheme and the types of tests used for each discourse relation. substitution/insertion of connectives (e.g., for alternation or narration) is complemented by paraphrase tests (e.g., the causal group of relations), as well as tests that are aimed at identifying the influence of the question under discussion, such as nominalization (assuming that nominalized events/situations still realize the semantic contribution, but do not provide an answer to the qud), explicit insertion of the linking question under discussion (for restatement), or explicit marking for information-structural properties (parallel and contrast). we have grouped the discourse relations into five broad categories, which each have own defining properties as well as properties typically disambiguating relations within that group: • contingency relations are defined in terms of sanders et al.’s causal source of coherence – normally, such a discourse relation presupposes a rule a → b (lagerwerf, 1998), which either has to be instantiated with real events or propositions (causal), expressed as a general rule (conditional), or imply a denial of an expectation from such a rule (denial). most frequently, different kinds of such relations can be distinguished with an explicit paraphrase of the rule involved (cf. subsection 4.2). • expansion relations achieve coherence through some kind of referential relation between referenced situations or entities in the arguments (elaboration group, cf. subsection 4.3), making explicit the contribution of a larger group of discourse units (interpretation), or link units that are only coherent because they answer a common question under discussion (continuative, cf. subsection 4.4). • temporal and reporting relations each express a specific (semantic) relation. • comparison relations usually involve an overt contrast between two objects. this overt contrast is essential for deciding between denial and comparison relations (subsection 4.1), but also between parallel and conjunction (subsection 4.4). 4.1 adversative relations as an example for the linguistic tests, let us consider the domain of adversative relations, which spenader and lobanova (2009), based on data from the rst discourse treebank, argue to consist of three different ‘discourse marker profiles’, without however proposing tests which could help the annotation of discourse corpora. 158 linguistic tests for discourse relations in tüba-d/z relation test material type of test contingency causal result-cause die folge daraus war, dass (arg2) complex phrase the consequence of this was that (arg2) result-enable das ermöglichte es, dass (arg2) complex phrase this made it possible that (arg2) das trug dazu bei, dass (arg2) complex phrase this contributed to the fact that (arg2) result-epistemic daraus schließe ich, dass (arg2) complex phrase i infer from this that (arg2) result-speechact deswegen [speechact-verb] ich: (arg2) complex phrase hence i [speechact-verb]: (arg2) explanation-cause der grund dafür ist, dass (arg2) complex phrase the reason for this is that (arg2) explanation-enable möglich gemacht wurde das durch (arg2) complex phrase this was possible by (arg2) explanation-epistemic das schließe ich daraus, dass (arg2) complex phrase i infer this from the fact that (arg2) explanation-speechact ich [speechact-verb] das, weil (arg2) complex phrase i [speechact-verb] this because (arg2) conditional consequence (always marked) — alternation ansonsten (arg2) / otherwise, (arg2) insertion of connective condition (always marked) — denial concessionc zwar (arg1) aber (arg2) insertion of connective certainly (arg1) although (arg2) concession substitution of arg1 with ‘trotz np’ (‘despite np’) nominalization anti-explanation der grund dafür ist nicht, dass (arg2) complex phrase the reason for this is not that (arg2) expansion elaboration restatement inwiefern (arg1)? insofern, als (arg2) question-answer-coherence in what respect (arg1)? to the extent that (arg2) instance zum beispiel (arg2) / for example, (arg2) insertion of connective instanceva beispielsweise und vor allem (arg2) insertion of two connectives especially, for example, (arg2) background was (arg1-elab) betrifft (arg2) complex phrase as to (arg1-elab), (arg2) interpretation summary zusammenfassend gesagt, (arg2) complex phrase to summarize, (arg2) commentary ehrlich gesagt, (arg2) insertion of adverbial to be honest, (arg2) continuative continuation — — conjunction omitting any of (arg1) or (arg2) possible omission of edu table 2: summary of linguistic tests for each relation (part i) a) instancev is a variant of the instance relation where an especially salient example, rather than any applicable instance, is picked out. 159 versley and gastel relation test material type of test temporal narration dann (arg2) / then, (arg2) insertion of connective precondition zuvor (arg2) / previously, (arg2) insertion of connective comparison parallel auch (arg2) / also, (arg2) insertion of connective ct on overt contrast accent placement switch arg1 and arg2 symmetry parallelvb auch und vor allem (arg2) / also, and especially (arg2) insertion of two connectives contrast während (arg1), (arg2) / while (arg1), (arg2) insertion of connective ct/f on overt/secondary contrast accent placement switch arg1 and arg2 symmetry reporting attribution substitution of arg2 with np nominalization source speaker of arg2 can be replaced with other person point of view table 3: summary of linguistic tests for each relation (part ii) b) parallelv is a variant of parallel where the second relation argument is presented as an “even stronger” example for the property in question. the connective annotation, in which the linguistic tests were developed, distinguishes between contrast relations (where two items are compared), contraexpectation (where a plausible consequence of arg1 is denied in arg2) and antithesis (where a question under discussion is first answered with the contents of arg1, but the answer of arg1 is overridden by the answer in arg2), which is mostly identical to the three-way distinction used by rst. (11) [qud: are all children equally tall?] peter is stubby, but mary is tall. contrast (12) peter has bought the book, but he hasn’t read it. contraexpectation (13) [qud: should we go to the pool?] peter likes going to the pool, but mary cannot swim. antithesis in the discourse structure annotation, the coordinating cases of contraexpectation (12) and antithesis (13) are grouped in one relation (concessionc), against the subordinating cases (“peter hasn’t read the book although he has bought it”, concession) and contrast relations such as (11), which helps avoid some corner cases between antithesis and contraexpectation. of the three, contrast can be delimited most clearly: it is a characterization of two related entities (called overt contrast by lang, 1984), which differ in some relevant property (the secondary contrast) while both relata talk about a common question (called common integrator by lang). using these properties, it is possible to formulate several testable assertions on the contrast relation: firstly, contrast is symmetric: unlike with (12) and (13), a swapped version of (11), „mary is tall, but peter is stubby“ is just as informative with respect to the qud. secondly, the absence of an overt contrast pair (as in (12), where both sentences share the subject) is a relatively clear sign against a contrast relation. finally, the absence of a secondary contrast compatible with the inferred qud, as is the case in (13), is also a sign against a contrast relation. the contraexpectation relation is grouped among the contingency relations in the tüba-d/z annotations since a single 160 linguistic tests for discourse relations in tüba-d/z logical relation a -> b can be realized (or rather, presupposed) in a multitude of different relations including consequence (‘if it is sunny, the laundry dries well’), result (‘because it was sunny, the laundry dried well’) or concession (‘although it was not sunny, the laundry dried well’), as well as anti-explanation (‘the laundry dried well, but not because it was sunny’). as a result, we can test this logical relation among the examples: in (11), peter being stubby does not have any influence on mary’s size, and in (13), peter’s liking of the pool does not have any influence on the swimming abilities of mary. in contrast, „peter bought the book“ raises the (plausible) expectation that he also read the book, which is directly denied in the following clause „he hasn’t read it“. (in other cases, such as „peter turned the ignition but it was not his car’s favorite day“, a strong expectation can be implicitly denied). a less involved test for the contraexpectation relation is the possibility of an “obwohl (arg1), (arg2)” (‘although (arg1), (arg2)’) paraphrase, which is possible in (12) but out of question in (11) and (13). in the typical case of antithesis, the relation between two clauses is only clear with respect to the question under discussion, whereas overt/secondary contrast pair or a denied expectation would be interpretable regardless of context. while it could be argued that antithesis should be seen as subsuming the other two relations (as argued by umbach and stede, 1999, as well as spenader and maier, 2005, who both use the term contrast for a relation), such an analysis would fail to explain cases ambiguous between parallel and contrast, and would have difficulties accounting for constructions that can only occur with contraexpectation, such as “trotz np” (‘despite np’) paraphrases and “obwohl s” (‘although s’). 4.2 temporal/causal markers markers with a primary temporal meaning, such as nachdem (‘after’/‘since’/‘as’), als (‘when’/‘as’), or während (‘while’) often convey causal or contrastive relations. in the case of temporal relations between events (for nachdem and als), herweg (1991) and bäuerle (1995) claim that their core meaning is not a relation between temporal intervals, but one of situational location – in the case of nachdem, in the post-phase of the subclause’s event, in the subclause’s process phase for während or in the general (situational) proximity of the subclause’s event for als. this situational connection can also impart causal information by the intermediary of the construction of a post-state (e.g. from going to the supermarket to being in the supermarket), or other semantic relations in the case of (event) redescriptions with als as in example (14) by bäuerle, which would correspond to our restatement: (14) als fritz den hund fütterte, gab er ihm schappi. when fritz fed the dog, he gave him chappi. in contrast to causal readings and redescriptions, contrastive readings seem to occur parasitically on the situations present in temporal nachdem or während connectives. often, a primarily contrastive or parallel reading with these connectives contains an explicit temporal adverbial in addition to the nachdem or während clause:12 (15) [diese kuppel war im juni 1998 zum zweiten mal eingestürzt,] [nachdem sie im februar 1997 erstmals eingebrochen war.] 12. tüba-d/z, sentence 680 161 versley and gastel [this cupola had collapsed for the second time in june 1998,] [after it had caved in for the first time in february 1997]. narration+parallel cases where a temporal coherence relation co-occurs with a causal or contrastive one can often be detected by substituting a synchronous marker (als or während) with an asynchronous one (nachdem) or vice-versa, since the non-temporal information would keep the coherence independently of the temporal relation. comparing the results of a substitution with weil (‘because’) with those that one gets by substituting with kurz nachdem (‘shortly after’), which forces a purely temporal reading, reveals some cases that are not purely temporal, but also do not lend themselves to substitution with weil. such cases of contributing causes are distinguished from (strong) causes, and are presented as a weakresult relation by bras et al. (2009). such enabling conditions are annotated with the result-enable and explanation-enable relations in our scheme, and can be distinguished from their stronger counterparts result-cause and explanation-cause by the fact that they allow the addition of a ‘main’ cause with a weil adjunct:13 (16) [nachdem die fraktionsvorsitzende künast sich anders entschieden hat,] [gilt bundestagsvizepräsidentin vollmer als aussichtsreichste kandidatin.] (weil sie als kompetent gilt) [after the parliamentary leader künast has decided otherwise,] [bundestag vice president vollmer is considered the most promising candidate.] (because she is considered competent) result-enable (17) [im vergangen jahr war es an der kastanienallee zu ausschreitungen gekommen,] [nachdem die polizei ein transparent aus dem demozug entfernen wollte.] (??weil die demonstranten unzufrieden waren) [last year riots occurred at the kastanienallee] [after police wanted to remove a transparent from the demonstration train.] (??because the demonstrators were unhappy) explanation-cause note that the distinction between main and contributing causes often hinges on the conceptualization of the events (and their manipulability) by the speaker rather than on the events themselves: a spontaneous eruption of violence (riots occurred) may be caused by actions of the police whereas conscious action (the demonstrators threw stones) may be better explained by an independent motivation rather than surrounding events alone. in addition, there are non-temporal uses of während (as a general marker of contrast) and nachdem (for the justification of claims, called pragmatic cause or evidence in other annotation schemes). in the first case, one can add temporal adverbials (such as in the morning and in the evening) to both arguments of the connective so as to exclude the temporal reading of während. in the second case, the difference is between a connection between situations (cause and enable) and a connection between claims (our result-epistemic, cf. also sweetser, 1990; sanders, 2005), which can, but does not have to, occur with an order inverse to normal causality between events, such as example (18), due to sweetser: (18) john loved her, because he came back. 13. tüba-d/z, sentence 2551; tüba-d/z, sentence 18884. note how example (16) would still be acceptable if the nachdem-clause were postposed in order to match the subclause ordering in (17). 162 linguistic tests for discourse relations in tüba-d/z in sentences where a prediction or other claim is supported by a result-epistemic or explanationepistemic relation, the claim is non-factive (lang, 2000, explains the distinction as assuming rather than asserting a proposition) and can be amended or retracted, as in example (19):14 (19) somit ist mala, eine einödgemeinde, wohl jetzt der wunschort der atomindustrie. [nachdem aber schon vier kommunen dem skb bewiesen haben, daß sie trotz hoher arbeitslosigkeit nein zum atomklo sagten,] [sieht es für den strahlenmüll zappenduster aus.] (. . . oder die atomindustrie hat noch einen trumpf im ärmel) hence, the solitary borough of mala seems to be the desired locality for the atomic industry. [as already four municipalities proved to the skb that they rejected an atomic disposal site despite high unemployment,] [the outlook is dim for the radiation waste.] (. . . or the atomic industry has another ace in their sleeve) result-epistemic 4.3 the expansion group in the process of discourse annotation, we also noticed that a group of discourse relations that typically occur with unmarked instances was posing some difficulties. in older proposals, many of these would have been called elaboration. indeed, this specific set of subordinating relations serves the purpose of giving further information and elaborating the first argument in a way that is not strictly necessary for the coherence and understanding of the main text. the criteria for these relations — the expansion.elaboration group containing the instance, restatement and background and the expansion.interpretation group containing summary and commentary — often have to take into account the surrounding discourse. as they do not convey a strong semantic relation, substitution with prepositional phrases or even insertion of subordinating conjunctions is not the most straightforward test. instead, the most indicative tests for this group of relations pick out properties such as the perspectivization of each connected discourse unit, or the relation of the themes/questions under discussion of arg1 and arg2. we use the instance for relations where arg2 concerns a proper subset of the events described in arg1. these cases allow the insertion of phrases such as “zum beispiel” (‘for example’) in arg2. (20) max did well in school this year. in biology [for example] he had an ‘a’. in this case, max’s doing well in biology this year is conceptualized as a proper subset of max doing well in school this year. in case of restatement, the main event of arg2 elaborates the main event of arg1 as a whole. jasinskaja (2007) claims that the restatement relation works similar to an apposition in the case of nominal phrases in sentence syntax. the test of asking “inwiefern (arg1)” (‘in what respect (arg1)’) and reading arg2 as the answer to this makes the question under discussion – further qualification of arg1 by a redescription as arg2 – explicit. if the question is asked with sentential (i.e. largest possible) focus and arg2 is still a good and coherent answer, the relation must be a restatement.15 (21) max did well in school this year. [in what respect did max do well in school this year?] he had an ‘a’ in every subject. 14. tüba-d/z, sentence 39373 15. many restatement relations can be marked with insofern, als/dass (in that), which however often feels awkward for stylistic reasons or due to the added syntactic complexity. 163 versley and gastel in the discourse relation background, arg2 elaborates a non-central part of arg1, typically one phrase. this can be tested with an addition such as “was [...] angeht,” (‘as to [elaborated phrase]’). this frame-setting topic makes explicit that it is not the main event of arg1 that is elaborated, but the specific phrase.16 (22) max did well in school this year. as to this year, this was his second last high school year. in summary, restatement, background and instance are different in which part of arg1 they elaborate. the two other relations, commentary and summary differ from the former three because commentary introduces information that differs in perspectivization and summary does not introduce any new facts at all. in the words of reese et al. (2007), commentary is a relation where arg2 provides an opinion or evaluation of the content associated with arg1. a test has to establish that the perspective of the commentary in arg2 is in fact the author’s, whereas arg1 is usually an objective fact. we can test for author perspectivization using commentary adverbials such as the german “ehrlich gesagt” (‘to tell the truth, . . . ’, ‘frankly, . . . ’) which are utterance modifiers and cannot occur under other perspectives than the author’s (bellert, 1977; potts, 2005). in our case, we want to go beyond just taking the presence of utterance modifiers in the original text as an indicator for a commentary relation (as, e.g., reese et al. do). using the addition of utterance modifiers as a linguistic test by manipulating the corpus sentences means that annotators have to take into account possible shifts of meaning, as in (24): (23) max did well in school this year. [frankly,] this is exactly what i expected. (24) # max did well in school this year. frankly, he had an ‘a’ in every subject. in (24), the utterance would sound odd in a neutral context. for summary relations, arg2 is a reformulated and condensed version of (the most important part of) the content of arg1.17 inserting a summative adverbial such as “zusammenfassend gesagt” (‘in summary’) is only possible if arg2 does not contribute any new information for the reader, which provides us with a plausible linguistic test. in (25) and (26), the summative adverbial excludes further elaboration. (25) max had an ‘a’ in biology, maths, physics, and history. [in summary,] he did well in school this year. (26) max did well in school this year. # in summary, he had an ‘a’ in every subject. an acceptable reading of (26) would entail an additional implication such as “i could tell you more facts about this, but in summary [of these untold facts], he had an ‘a’ in every subject”. 16. one reviewer pointed at the possibility of pronominalizing the elaborated part in arg2, which often sounds more natural. using as to as an explicit frame-setting construction is more independent of other factors influencing pronominalizability, especially the salience of the entity in arg1 or the presence of confounders. in addition, it gives the correct result when there is a dissociation between pronominalized entity and topic entity such as in that letter from him. 17. jasinskaja (2007) does not differentiate between the cases that we would annotate as summary and those that would be a restatement in our annotation scheme, subsuming both cases under her restatement label. 164 linguistic tests for discourse relations in tüba-d/z 4.4 weak coherence more frequently than one might expect, two discourse segments that belong together do not show a discernable semantic or referential relation such as those discussed in the previous subsections. such relations occur as continuation or parallel in the scheme of reese et al. (2007), as joint or list relations in that of the rst discourse treebank (carlson and marcu, 2002) or as conjunction and other relations in the penn discourse treebank (prasad et al., 2008). while their presence is often obvious from the surrounding context, the lack of semantic distinctions can make these relations somewhat elusive. however, using an explicit notion of the context’s contribution we can hope to get a firmer hold of these relations. in classical examples of weak coherence, two discourse segments together provide an answer to a question under discussion where the single segments do not. this can be illustrated with this example from blakemore and carston’s (2005) discussion of the possible relations expressible by the conjunction and: (27) a: shall we start without jane? b: [well, she did say to start if she was late,] [and we have been waiting for half an hour now.] continuation using such an accumulation-of-evidence refinement to the continuation relation with an inventory close to that of reese et al. (2007) led to situations where annotators would use the parallel relation label for cases that do not have an overt contrast between two entities. while such examples do receive a parallel or continuation label in reese’s scheme, we found it preferable to introduce a third category conjunction as a coordination relation between discourse segments that provide independent (satisfactory) answers to the same question under discussion while not incorporating a comparison between different entities.18 example (28)19 shows a typical example of conjunction: being calm (inferred from ‘sauer’s voice sounds austere’) and recounting precise details both contribute to the credibility/reliability of the witness independently, which excludes continuation. furthermore, there is no overt contrast which would allow a parallel relation. (28) tatsächlich lässt der zweite zeuge den mandanten kaum chancen auf eine milde strafe. (qud: warum ist sauer ein verlässlicherer zeuge als emrich?) [im gegensatz zu emrich klingt sauers stimme nüchtern,] [und er kann sich an genaue details erinnern.] indeed the second witness leaves the clients scarcely a chance for a mild penalty. (qud: why is sauer a more reliable witness than emrich?) [in contrast to emrich, sauer’s voice sounds austere,] [and he can remember precise details.] conjunction 18. as our notion of discourse structure follows sdrt in that it postulates a structure that simultaneously includes a coordinating relation between dependants of a superordinate segment, such as conjunction, and a subordinating relation between the superordinate segment and its subordinated segment, e.g., explanation-epistemic, we need to postulate a conjunction relation where rst just annotates the subordinating relations. in the typical case of a conjunction co-occurring with multiple instances of the same subordinating relation, we consider the conjunction to be implied if no other relations is annotated and the result looks very similar to rst’s “multiple satellites” construction. 19. tüba-d/z, sentence 2469 165 versley and gastel 5. inter-annotator and inter-adjudicator agreement as in any annotation project, the annotation scheme and its description in terms of annotation guidelines is not the sole determinant of the final quality but needs to be seen in conjunction with the training and abilities of actual annotators. numerical measures of agreement target this notion of final quality by asking whether the annotation task is repeatable when done by two annotators independently of each other. in this section, we report quantitative agreement figures from two studies involving a sample of documents annotated, respectively, by the same two annotators: one pilot study with three documents, and one using a larger sample of seven documents. we also give a more detailed appraisal of agreements and disagreements based on a larger sample that however involves a larger group of annotators. among the aspects of our annotation scheme, edu segmentation shows invariably high agreement (κ > 0.9 for all articles). among the disagreements in edu segmentation, non-typical communication verbs (think, contend, write) are correlated with the presence or absence of a particular type of relation (esp. the reporting group), and non-restrictive relative clauses, usually correlate with the presence or absence of a background relation. however, these are not among the most frequent nor the most interesting groups of disagreements (most of these are resolved in the first pass). the most interesting aspect of an agreement study therefore consists in discourse relations, as well as the placement of topic segments and transitional edus. in previous studies on annotator agreement, marcu et al. (1999) determined κ values between κ = 0.54 (brown corpus) and κ = 0.62 (muc) for fine-grained rst relations and between κ = 0.59 (brown) and κ = 0.66 (muc) for coarser-grained relations. in their reliability study with the penn discourse treebank, prasad et al. (2008) determined agreement values between 80% (finest level) and 94% (coarsest level with 4 relation types), but did not report any chance-corrected values. al-saif and markert (2010) report values of κ = 0.57 for their pdtb-inspired connective scheme, saying that most disagreements are due to highly ambiguous connectives such as ‘w’ (arabic counterpart of and), which can receive one of several relations. in a study on their dutch rst corpus, van der vliet et al. (2011) found an inter-annotator agreement of κ = 0.57. to the best of our knowledge, no agreement figures have been published on the rst-based potsdam commentary corpus (stede, 2004a) or any other german corpus with discourse relation annotation. 5.1 quantitative agreement figures in the regular annotation process, two annotators create edu segmentation, topic segments, and discourse relations independently from each other; in a second step, the results from both annotators are compared (cf. figure 1) and a coherent gold-standard annotation is created after discussing the goodness-of-fit of respective partial analyses with the text and the applicability of linguistic tests. in order to account for the complete annotation process including the revision step, we follow burchardt et al. (2006) and separately report inter-annotator agreement, which is determined after the initial annotation, and inter-adjudicator agreement, which is determined after an additional adjudication step, which is carried out by two adjudicators based on the original set of annotations but working independently from each other. in the case where multiple relations were annotated between the same edu ranges (for example, a temporal narration relation in addition to a result-cause relation from the contingency group), 166 linguistic tests for discourse relations in tüba-d/z figure 1: display of annotation differences 167 versley and gastel restatement background 8 result-cause result-enable 7 restatement explanation-cause 6 restatement instance 5 continuation narration 4 conjunction parallel 4 narration result-cause 4 narration parallel 4 restatement explanation-epistemic 3 explanation-cause explanation-enable 3 table 4: most frequent disagreements (in 18 documents annotated by different annotators) we counted the annotations as matching whenever the complete set of relations (i.e. narration, result-cause in the example) is the same across annotators. for the agreement among relations, we performed a pilot study on a sample of three documents, where annotators agreed on 49 relation spans (about 40% of all spans); later, we performed a second study on a sample of seven documents, where annotators agreed on 108 relation spans (about half of all spans). the pilot study did not make an effort to match relations among spans with segmentation differences, which can agree due to differences in edu segmentation or differences in the exact span of a complex discourse unit. in the second study, we accounted for some differences in edu segmentation by matching two edus if they have the same span or if their clauses have the same (syntactic dependency) head terminals. in both studies, the agreement was only evaluated among relations with compatible spans (evaluating the labeling of relations but not the discourse structure). the percentage of multi-edu spans in both samples is about the same as in the corpus as a whole (27.6% and 29.8%, compared to 25.1% for the whole corpus). in the pilot study, annotator agreement yielded κ = 0.55 for individual relations, and κ = 0.65 for the middle level of the taxonomy (nine relation types), whereas in the second study, we found κ = 0.59 for individual relations and κ = 0.69 for the middle level of the taxonomy. for the inter-adjudicator task, we found an agreement on 82 relation spans (about two thirds of all annotated spans), among which relation agreement was at κ = 0.83 for individual relations, and κ = 0.85 for the middle level of the taxonomy, or a reduction of disagreements of about 57%. for the second study, with 167 matching spans (= 78% of all annotated spans), we found κ = 0.90 for individual relations and κ = 0.91 for the middle level of the taxonomy, with, again, a reduction of disagreements by about two thirds. to measure quantitative agreement on the topic segmentation, we used the same sample of seven newspaper articles, which were marked for topic boundaries independently and without considering the paragraph boundaries in the printed newspaper, yielding 21 (annotator 1) or 22 (annotator 2) topic segments, of which 16 matches. modeling the annotation process as tagging each sentence as topic start or as nonboundary yields a value of κ = 0.71. positive specific agreement on nontrivial boundaries (i.e., those not at the start of the main text of the article) is at f1 = 0.62. agreement on transitional edus is very good: among the seven texts we used for our study, six articles showed perfect agreement, with four matching transitional edus altogether and one unmatched. modeling the annotation process as tagging each inter-topic boundary with or without a transitional edu yields a value of κ = 0.85, or a positive specific agreement of f1 = 0.88. 168 linguistic tests for discourse relations in tüba-d/z 5.2 identifying problematic areas table 4 shows the most frequent confusions between discourse relations in cases where different annotators could agree on a span, but chose different relations, collected over a larger sample of documents annotated by at least two trained annotators. surprisingly, one of the greatest sources of confusion is disagreement between the restatement and explanation-cause or explanation-epistemic relations. such disagreements can sometimes arise when situation identity is difficult to assess, as in example (29):20 (29) [ die deklarationsbestimmungen der eu sind bis heute schwammig:] [ hingewiesen werden muß nur auf im endprodukt nachweisbare genmanipulationen.] [the declaration provisions of the eu are vague even today:] [only genetic manipulations that are demonstrable in the final product have to be indicated.] restatement in this case, insertion of a causal marker is possible (which would indicate an explanation-cause relation), but the stronger test of inserting a ‘in what regard?’ subquestion shows that (29) is indeed a case of restatement. in other cases, annotated relations vary in the degree of causality (between purely temporal and an enable relation, or between enable and cause). in some cases, modal embedding or other modifiers make the causality judgement more difficult, as in example (30),21 where a non-modal “if there is a need for it, we destroy the gm crops” would receive a result-enable relation (since there is a causal link, but the knowledge does not actually force the administration to destroy crops), but the modal embedding means that the discourse relation must be a result-cause relation relating the knowledge of the seed locations to the fact that it is possible to find and destroy the crops. (30) [“wir wissen aber, wo das saatgut hingegangen ist”, so fluhme,] [“wenn wirklich die notwendigkeit bestehen sollte, könnten wir die betroffenen felder ausfindig machen und die vernichtung der gentechnisch veränderten pflanzen sicherstellen.] [“but we know where the seed has gone”, says fluhme,] [“if there was a need for it, we could locate the concerned fields and ensure the destruction of the genetically modified crops.] result-cause 6. summary in this article, we have presented the most important design choices that an annotation scheme for discourse structure and discourse relation faces, and presented the particular solutions taken in the annotation scheme for discourse annotation in the tüba-d/z corpus. we have explicitly linked the categories of the annotation scheme to linguistic categories that have been described in the literature and demonstrated how this can be harnessed for the creation of linguistic tests that help in the annotation of discourse relations. finally, we have presented figures for inter-annotator and inter-adjudicator agreement in our corpus and have shown how the remaining ambiguities in the annotation scheme relate to problems such as event identity or parthood that also pose difficulties for other annotation schemes (such as event relation annotation in timeml). 20. tüba-d/z, sentences 5717ff. 21. tüba-d/z, sentence 1083. 169 versley and gastel acknowledgements yannick versley and anna gastel were supported by the deutsche forschungsgemeinschaft (dfg) as part of sfb 833. we would like to thank the three anonymous reviewers for their helpful suggestions. references amar al-saif and katja markert. the leeds arabic discourse treebank: annotating discourse connectives for arabic. in proceedings of the 7th international conference on language resources and evaluation (lrec 2010), 2010. nicholas asher. discourse topic. theoretical linguistics, 30(2-3):163–201, 2004. nicholas asher and alex lascarides. the semantics and pragmatics of presupposition. journal of semantics, 15(3):239–299, 1998. nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. nicholas asher and laure vieu. subordinating and coordinating discourse relations. lingua, 115 (4):591–610, 2005. rainer bäuerle. temporalsätze und bezugspunktsetzung im deutschen. in brigitte handwerker, editor, fremde sprache deutsch, pages 155–176. narr, 1995. irena bellert. on semantic and distributional properties of sentential adverbs. linguistic inquiry, 8 (2):337–351, 1977. diane blakemore and robyn carston. the pragmatics of sentential coordination with "and". lingua, 115:569–589, 2005. myriam bras, anne le draoulec, and nicholas asher. a formal analysis of the french connective alors. oslo studies in language, 1(1):149–170, 2009. daniel büring. on d-trees, beans, and b-accents. linguistics and philosophy, 26(5):511–545, 2003. paula christina figueira cardoso, erick galani maziero, maria lucía del rosario castro jorge, eloize rossi marques seno, ariano di felippo, lucia helena machado rino, maria das graças volpe nunes, and thiago alexandro salgueiro pardo. cstnews — a discourse-annotated corpus for single and multi-document summarization of news texts in brazilian portuguese. in 3rd rst brazilian meeting, 2011. lynn carlson and daniel marcu. discourse tagging reference manual. technical report, information sciences institute, university of southern california, 2002. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in current directions in discourse and dialogue. kluwer, 2003. iria da cunha, juan manuel torres-moreno, and gerardo sierra. on the development of the rst spanish treebank. in proceedings of the 5th linguistic annotation workshop (law v), 2011. barbara grosz and candice sidner. attention, intentions, and the structure of discourse. computational linguistics, 12(3):175–204, 1986. marti a. hearst. texttiling: segmenting text into multi-paragraph subtopic passages. computational linguistics, 23(1):33–64, 1997. michael herweg. temporale konjunktionen und aspekt. kognitionswissenschaft, 2(3–4):51–90, 1991. 170 linguistic tests for discourse relations in tüba-d/z jerry hobbs. on the coherence and structure of discourse. technical report 85-37, center for the study of language and information (csli), 1985. eduard h. hovy and elisabeth maier. parsimonious or profligate: how many and which discourse structure relations? unpublished manuscript, 1995. julie hunter, nicholas asher, brian reese, and pascal denis. evidentiality and intensionality: two uses of reportative constructions in discourse. in constraints in discourse 2006, 2006. ray jackendoff. semantic interpretation in generative grammar. mit press, cambridge, ma, 1972. ekaterina jasinskaja. pragmatics and prosody of implicit discourse relations: the case of restatement. phd thesis, universität tübingen, 2007. alistair knott. a data-driven methodology for motivating a set of coherence relations. phd thesis, university of edinburgh, 1996. alistair knott and robert dale. using linguistic phenomena to motivate a set of coherence relations. discourse processes, 18(1):35–62, 1994. alistair knott, john oberlander, michael o’donnell, and chris mellish. beyond elaboration: the interaction of relations and focus in coherent text. in text representation: linguistic and psycholinguistic aspects. john benjamins, 2001. sudheer kolachina, rashmi prasad, dipti misra sharma, and aravind joshi. evaluation of discourse relation annotation in the hindi discourse treebank. in proceedings of the 8th international conference on language resources and evaluation (lrec 2012), 2012. jan van kuppevelt. discourse structure, topicality and questioning. journal of linguistics, 31(1): 109–147, 1995. luuk lagerwerf. causal connectives have presuppositions. holland academic graphics, the hague, 1998. ewald lang. adversative connectors on distinct levels of discourse: a re-examination of eve sweetser’s three-level approach. in elizabeth couper-kuhlen and bernd kortmann, editors, cause-condition-concession-contrast, number 33 in topics in english linguistics. mouton de gruyter, 2000. harald lüngen, csilla puskàs, maja bärenfänger, mirco hilbert, and henning lobin. discourse segmentation of german written text. in proceedings of the 5th international conference on natural language processing (fintal 2006), 2006. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text, 8(3):243–281, 1988. daniel marcu, estebaliz amorrortu, and magdalena romera. experiments in constructing a corpus of discourse trees. in acl workshop on standards and tools for discourse tagging, 1999. lucie mladová, šárka zikánová, and eva hajičova. from sentence to discourse: building an annotation scheme for discourse based on prague dependency treebank. in proceedings of the 6th international conference on language resources and evaluation (lrec 2008), 2008. vera möller and karin naumann. manual for the annotation of in-document referential relations. technical report, seminar für sprachwissenschaft, universität tübingen, 2009. johanna d. moore and martha e. pollack. a problem for rst: the need for multi-level discourse analysis. computational linguistics, 18(2):537–544, 1992. 171 versley and gastel nick nicholas. parameters for an ontology of rhetorical structure theory. university of melbourne working paper in linguistics, 15(15):77–93, 1995. thiago alexandre salgueiro pardo and maria das graças volpe nunes. on the development and evaluation of a brazilian portuguese discourse parser. journal of theoretical and applied computing, 15(2):43–64, 2008. thiago alexandre salgueiro pardo and eloize rossi marques seno. rhetalho: um corpus de referência anotado retoricamente. in anais do v encontro de corpora, 2005. the pdtb research group. the penn discourse treebank 2.0 annotation manual. technical report ircs-08-01, institute for research in cognitive science, university of pennsylvania, 2008. http://www.seas.upenn.edu/ pdtb/pdtbapi/pdtb-annotation-manual.pdf. marie-paule péry-woodley, stergos d. afantenos, lydia-mai ho-dac, and nicholas asher. la ressource annodis, un corpus enrichi d’annotation discursives. traitement automatique des langues, 52(3):71–101, 2011. livia polanyi and remko scha. a syntactic approach to discourse semantics. in proceedings of coling 6, 1984. christopher potts. the logic of conventional implicatures. oxford university press, 2005. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of the 6th international conference on language resources and evaluation (lrec 2008), 2008. marta recasens, eduard hovy, and m. antoǹia martí. identity, non-identity and near-identity: adressing the complexity of coreference. lingua, 121(6):1138–1152, 2011. brian reese, julie hunter, nicholas asher, pascal denis, and jason baldridge. reference manual for the analysis and annotation of rhetorical structure. technical report, university of texas at austin, 2007. craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. osu working papers in linguistics 49, ohio state university, 1996. ted sanders. coherence, causality and cognitive complexity in discourse. in symposium on the exploration and modelling of meaning (sem-05), 2005. ted sanders and wilbert spooren. communicative intentions and coherence relations. in wolfram bublitz, uta lenk, and eija ventola, editors, coherence in text and discourse, pages 235–250. john benjamins, 1999. ted j. m. sanders, wilbert p. m. spooren, and leo g. m. noordman. toward a taxonomy of coherence relations. discourse processes, 15(1):1–35, 1992. frank schilder. robust discourse parsing via discourse markers, topicality and position. natural language engineering, 8(3):235–255, 2002. swapna somasundaran, galileo namata, janyce wiebe, and lise getoor. supervised and unsupervised methods in employing discourse relations for improving opinion polarity classification. in proceedings of the 2009 conference on empirical methods in natural language processing (emnlp 2009), 2009. jennifer spenader and emar maier. contrast as denial in multi-dimensional semantics. journal of pragmatics, 41(9):1707–1726, 2005. 172 linguistic tests for discourse relations in tüba-d/z manfred stede. the potsdam commentary corpus. in acl’04 workshop on discourse annotation, 2004a. manfred stede. does discourse processing need discourse topics? theoretical linguistics, 30(2-3): 241–253, 2004b. manfred stede. rst revisited: disentangling nuclearity. in catherine fabricius-hansen and wiebke ramm, editors, ‘subordination’ versus ’coordination’ in sentence and text a cross-linguistic perspective. john benjamins, amsterdam, 2008. mark steedman. information structure and the syntax-phonology interface. linguistic inquiry, 31 (4):649–689, 2000. eve sweetser. from etymology to pragmatics. cambridge university press, cambridge, 1990. maite taboada and william c. mann. rhetorical structure theory: looking back and moving ahead. discourse studies, 8(3):423–459, 2006. heike telljohann, erhard w. hinrichs, sandra kübler, heike zinsmeister, and kathrin beck. stylebook for the tübingen treebank of written german (tüba-d/z). technical report, seminar für sprachwissenschaft, universität tübingen, 2009. sarah tonelli, giuseppe riccardi, rashmi prasad, and aravind joshi. annotation of discourse relations for conversational spoken dialogs. in proceedings of the seventh international conference on language resources and evaluation (lrec 2010), 2010. isabel gomez txurruka. the natural language conjunction ‘and’. linguistics and philosophy, 26 (3):255–285, 2003. carla umbach and manfred stede. kohärenzrelationen: ein vergleich von kontrast und konzession. kit-report 148, technische universität berlin, 1999. enric vallduví and maria vilkuna. on rheme and kontrast. in the limits of syntax. academic press, 1998. yannick versley. vagueness and referential ambiguity in a large-scale annotated corpus. research on language and computation, 6(3–4):333–353, 2008. ninke van der vliet, ildiko berzlanovich, gosse bouma, gisela redeker, and markus egg. building a discourse-annotated dutch text corpus. in stefanie dipper and heike zinsmeister, editors, proc. beyond semantics (dgfs workshop), 2011. bonnie webber. d-ltag: extending lexicalized tag to discourse. cognitive science, 28(5):751– 779, 2004. bonnie webber, matthew stone, aravind joshi, and alistair knott. anaphora and discourse structure. computational linguistics, 29(4):545–587, 2003. bonnie lynn webber. discourse deixis: reference to discourse segments. in proceedings of the 26th annual meeting on association for computational linguistics, 1988. florian wolf and edward gibson. representing discourse coherence: a corpus-based study. computational linguistics, 31(2):249–287, 2005. deniz zeyrek, işin demirşahin, ayışıǧı sevdik-çallı, hale ögel balaban, i̇hsan yalçınkaya, and ümit deniz turan. the annotation scheme of the turkish discourse bank and an evaluation of inconsistent annotation. in proceedings of the fourth linguistic annotation workshop (law-iv), 2010. 173 dialogue & discourse 12(2) (2021) 192–226 doi: 10.5210/dad.2021.207 event and entity coreference across five languages: effects of context and referring expression luca bevacqua∗ luca.bev@hotmail.it school of philosophy, psychology & language sciences university of edinburgh sharid loáiciga∗ sharid.loaiciga@gu.se department of philosophy, linguistics and theory of science university of gothenburg hannah rohde hannah.rohde@ed.ac.uk school of philosophy, psychology & language sciences university of edinburgh christian hardmeier chrha@itu.dk department of computer science, it university of copenhagen department of linguistics and philology, uppsala university editor: massimo poesio submitted 08/2021; accepted 11/2021; published online 12/2021 abstract current work on coreference focuses primarily on entities, often leaving unanalysed the use of anaphors to corefer with antecedents such as events and textual segments. moreover, the anaphoric forms that speakers use for entity and event coreference are not mutually exclusive. this ambiguity has been the subject of work in english, with evidence of a split between comprehenders’ preferential interpretation of personal versus demonstrative pronouns. in addition, comprehenders are shown to be sensitive to antecedent complexity and aspectual status, two verb-driven cues that signal how an event is being portrayed. here we extend this work via a comparison across five languages (english, french, german, italian, and spanish). with a story-continuation experiment, we test how different referring expressions corefer with entity and event antecedents and whether verbal features such as argument structure and aspect influence this choice. our results show widely consistent, not categorical biases across languages: entity coreference is favoured for personal pronouns and event coreference for demonstratives. antecedent complexity increases the rate at which anaphors are taken to corefer with an event antecedent, as does portraying an event as completed though the latter does not reach significance. lastly, we report a comparison of the same referring expressions to refer to entity and event antecedents in a trilingual parallel corpus annotated with coreference. together, the results provide a first crosslingual picture of coreference preferences beyond the restricted entity-only patterns targeted by most existing work on coreference. the five languages are all shown to allow gradable use of pronouns for entity and event coreference, with biases that align with existing generalizations about the link between prominence and the use of reduced referring expressions. the studies also show the feasibility of manipulating targeted verbdriven cues across multiple languages to support crosslingual comparisons. keywords: event anaphora, coreference resolution, pronouns, multilingual. ∗. shared first authorship. c©2021 luca bevacqua, sharid loáiciga, hannah rohde and christian hardmeier this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). event vs entity interpretation 1. introduction work in coreference has focussed primarily on entity coreference, typically between third person personal pronouns and human antecedents. the use of anaphors to refer to non-human entities has been less studied, especially when the reference is to antecedents that are not entities, such as events or full clauses. non-entity coreference relations, including not only events but also previous portions of a text in discourse-deixis or metalinguistic description, present both formal and practical challenges. if the antecedent can be any of the available previously mentioned entities, events, propositions or even entire rhetorical arguments, the number of candidate antecedents becomes potentially very large. in addition, the anaphoric forms that speakers use for different kinds of coreference are not mutually exclusive: sentences (1) and (2) show the demonstrative this referring back to respectively an entity (the wine) and an event (the successful aging of the wine).1 1. the northern italian wine aged well. this was one of the best wines i’d ever tasted. 2. the northern italian wine aged well. this made the wine more marketable. this ambiguity has been the subject of work on the coreference system in english, with evidence of a split between comprehenders’ preferential interpretation of personal versus demonstrative pronouns (çokal et al., 2018; loáiciga et al., 2018). specifically, personal pronouns are preferred when referring to concrete entities, and demonstratives when referring to other more complex or abstract antecedents (see §2.1). demonstratives have in fact been identified as the expressions speakers use to signal a prominence shift, particularly to change the discourse topic (givón, 1983) and to reject the most prominent antecedent as a possible antecedent (comrie, 1997). different languages use different referring expressions, ordering the elements in a given language’s pronominal system along a prominence scale: for example, while in english personal pronouns encode antecedents that are highly accessible, topical and activated in short-term memory, in languages with widespread use of zero anaphors such as italian or spanish, personal pronouns will preferentially be used to refer to less prominent referents. moreover, less prominent antecedents will need more semantically rich types of referring expressions in order to be retrieved (such as overt pronouns, which encode number and gender, or even nouns/names), whereas more prominent antecedents can be referenced with a wider variety of referring expressions (von heusinger and schumacher, 2019). it is often the case that these richer referring expressions will also be phonologically richer. for our purposes, we talk about heavy (and light) referring expressions to indicate forms that encode both more semantic and phonological content. in addition, when processing coreference, comprehenders are shown to be sensitive to a wide range of features related to structural and meaning-driven properties of the passage (see reviews such as greene et al., 1992; carlson, 2003; rohde, 2019). those findings include a large body of work in psycholinguistics, much of which has emphasized english (with notable exceptions of course, e.g., kaiser and trueswell, 2004; mayol, 2018). here we pick up on two english findings from a psycholinguistic study: antecedent complexity (number of arguments) and aspect, two verb-driven cues that signal how an event is being portrayed. while the definition of “accessibility” (that we use as defined in accessibility theory: ariel 1988, 2004) or “prominence” is much discussed (see von heusinger and schumacher, 2019, for a 1. all examples are taken from the experiments described in §3. 193 bevacqua, loáiciga, rohde and hardmeier definition of application of the concept), some predictions are recurrent across theories. for example, the accessibility of a target antecedent is assumed to decrease with a higher number of possible competitors (see e.g. centering theory: grosz et al. 1995). beyond entity coreference, one can consider the potential antecedents and competitors made available by the description of an event. an event antecedent is more complex than an entity antecedent in that it not only includes more than one entity, but also the relationships between the entities. one way to operationalise antecedent count is to consider the category of a verb within the causative-inchoative alternation (haspelmath, 1993), which can then be used as a proxy for antecedent choice complexity. alternating verbs admit both intransitive and transitive uses (e.g., the snow melted / the sun melted the snow), corresponding to inchoative and causative interpretations, respectively. this alternation is not present in all verbs (e.g., the battery died / *frequent use died the battery). jespersen (1927) classes verbs that can undergo this alternation as “move and change verbs”, as they often describe changes in state or movement: the transitive use, then, makes available the implicit entity that brought on the change. alternating verbs, by offering this additional implicit entity to the pool of coreference possibilities even when the verb appears in the intransitive, have been shown to increase event coreference in english (loáiciga et al., 2018); for non-alternating verbs, the sole entity available in the sentence seemingly becomes relatively more salient by not having competition from other implicit entities. however, this is a gradable phenomenon, in that some verbs position in between categories: even verbs which are most firmly in the non-alternating category can have rare or idiomatic transitive uses (e.g. internal objects, to die a cruel death). languages may differ in how clean-cut the divide between alternating and non-alternating verbs is, and in how sensitive coreference is to this distinction. in the study reported here, we compare the availability of event and entity coreference across alternating vs non-alternating verb classes. another factor known to influence discourse representations and coreference processing is verb aspect. verb aspect is used as a cue to the portrayal of an event. our study also tests whether an aspectual manipulation may influence a comprehender’s perception of the concreteness of the event, and thereby its availability as an antecedent. events which are portrayed as completed might have a salient end-state and, thus, be more encapsulated and retrievable (moens and steedman, 1988) or, on the contrary, the ongoing state of an event may activate it and, with it, its properties and world knowledge (ferretti et al., 2001, 2007, 2009). the relevance of the studies we present here is underscored not only by the lack of work on abstract anaphora and the possibility of making crosslingual extensions to the limited prior work on this topic, but also by the similar questions raised in coreference-annotated corpora that inform computational systems for coreference resolution. existing resources in the field of coreference resolution also reflect the difficulty in the treatment of pronouns. for example, given the complexity involved in identifying non-nominal anaphora, and its relatively low frequency, most annotation efforts and hence most coreference systems focus on anaphora to nouns. the ontonotes corpus (pradhan and xue, 2009), for instance, although the largest in the community, does not include annotations of event entities coreferential with pronouns. corpora including events exist, but they are smaller in size and less frequently used. the parcorfull corpus (lapshinova-koltunski et al., 2018), for instance, includes the annotation of event anaphora in its guidelines, but event references only constitute 12.5% of the total annotated markables in english. parcorfull is a parallel corpus that also includes annotations in german and french. in german, event annotations amount to 11%; in french the proportion is even lower, however it should be noted that the lower number of documents for this part of the corpus makes the estimate less reliable. 194 event vs entity interpretation improving our understanding of the interpretation of event anaphora and non-nominal anaphora in general is a step towards better anaphora and coreference resolution systems (poesio et al., 2015). current systems struggle to identify and resolve this type of reference simply because they have not been trained to do so (heinzerling et al., 2017), as annotated corpora focus on nominal anaphora. annotated corpora are crucial for system development and large scale linguistic analysis, but equally important is a sound theoretical understanding of the phenomenon one wishes to annotate. finally, a more detailed and most importantly multilingual description of these coreference patterns can be useful for studies on multilingualism, which could refer to figures in a native speakers population as a point of comparison, and for language teaching. this paper presents work on these different types of coreference via a comparison across five languages (english, french, german, italian, and spanish), which are broadly related typologically but which differ in their use of grammatical gender and case and their pronominal systems, most notably the availability of a null pronoun. we ask when and to what degree event instances serve as antecedents when a competing entity referent is also available. the goal is to model human choices in five different language settings to investigate the similarities and differences that arise as a product of the structural differences in the languages, and to shed light on the relationship between event and nominal anaphora in order to inform future coreference annotation efforts and coreference systems. we report five story continuation studies targeting participants’ resolution of pronominal forms it vs this in english, il vs cela vs c’est in french, es vs das vs dies in german, ciò vs questo vs zero pronoun in italian, and esto vs este vs zero pronoun in spanish. regarding antecedent complexity and aspectual status, we test the impact of (i) the availability of an additional implicit argument to the verb, namely the complexity created via the causativeinchoative alternation and (ii) the aspectual difference of portraying of an event as completed or ongoing. regarding anaphoric forms, we test (iii) whether event coreference is inferred more for demonstratives over personal pronouns and (iv) whether this distinction is categorical, with a clear division of labour, or gradable. our results show consistent and gradable biases: participants encountering a personal pronoun are more likely to continue the sentence writing about the entity, while demonstratives prompt more event continuations across all five languages. the complexity of the antecedent, manipulated through the proxy of alternation, increases event coreference, though this is significant in only four of the five languages studied. the portrayal of an event as completed, on the other hand, seems to have a much smaller role in influencing coreference. we speculate how the varying aspectual systems of the five languages may undermine this effect. lastly, we report corpus frequencies comparing the use of the same referring expressions to refer to entity and event antecedents, along with their coreference chains. this comparison confirmed the preference for events to be referred to mostly with demonstrative pronouns and entities with personal pronouns, but with a more gradable pattern than in the experiments with human participants. in contrast, the corpus study failed to find a similar pattern for the proxy of verb alternation due to a lack of annotated resources with alternation status or thematic roles. together, the results provide the first crosslingual picture of coreference preferences beyond the restricted entity-only patterns that the target of most existing work on coreference. the five languages are all shown to allow gradable use of personal and demonstrative pronouns for entity and event coreference, with systematic biases that align with existing generalizations about the use of lighter referential forms for less complex and more concrete referents. 195 bevacqua, loáiciga, rohde and hardmeier 2. related work referring expressions play a central role in human communication and are at the heart of multiple linguistics theories and applications involving natural language understanding. for both, a lot of attention has been paid to the most standard cases represented by nominal expressions and personal pronouns while more atypical cases have been understudied. among these is the phenomenon of abstract anaphora – anaphora that involve reference to abstract entities such as events or states (asher, 1993) whose complexity is increased because the form and interpretation of the referring expressions involved have a high degree of variability. this complexity is also behind the lack of resources annotated with this type of anaphora, simultaneously hindering its study. 2.1 linguistic descriptions linguistic studies on the subject have employed both corpora and psycholinguistic methods. corpusbased studies are a source of insights about language use, since the written texts they are based on are natural passages after all. they offer better estimates for building coreference resolution systems that will be used on those texts. on the other hand, corpus-based studies do not provide any explanation as to why a particular item follows a certain distribution, and they grant little control over the confounding variables responsible for that distribution. in this respect, psycholinguistic studies are more suitable for capturing the cognitive processes behind naturally occurring phenomena. psycholinguistic research has focused on using theoretical constructs of complexity, salience, and information status to capture coreference patterns. in english, the anaphoric use of pronouns to refer to entities or abstract events has been evolving in the last centuries (azuma, 2008): while it has consistently been the preferred form to corefer to pronominal, highly accessible antecedents, the distal demonstrative that has grown in use to refer to np or clausal antecedents (even overtaking the personal pronoun for non-nominal antecedents), positioning itself between it and the proximal demonstrative this, the most rarely used and marked of the three expressions. in an estimate of the current relative frequency of nominal coreference, gundel et al. (2005) report that while about 84% of personal pronouns had nominal antecedents, only 28% of demonstratives did. in other studies, the demonstratives this and that have been grouped together, assuming that they pattern together and behave differently from it (as confirmed in the context of textual deixis: see çokal et al. 2014). hedberg et al. (2007), for example, show how the two demonstratives tend to have non-nominal antecedents, unlike it: the personal pronoun requires the antecedent to be more highly activated, to be in the focus of attention, while demonstratives can retrieve antecedents that are less activated because of their complexity or because they were not directly introduced, such as events. this is confirmed in the study by çokal et al. (2018), where it was found to corefer to concrete entities both in production and interpretation, while this was biased towards non-np, less salient antecedents. likewise, in our previous study (loáiciga et al., 2018) we find a preference for it to refer to entities and of this to refer to events. similar results were also found by wittenberg et al. (2021) looking at it versus that, with the demonstrative being preferred for event coreference: this is interpreted as a “conceptual bundling” action of that, which would make event antecedents accessible by wrapping their complex structure in a simpler referring expression. in a comparison of demonstratives across english, german, italian and spanish (dipper et al., 2011), english and german were shown to prefer the use of demonstratives for retrieving abstract anaphors, while italian and spanish were not shown to display a preference. the german findings 196 event vs entity interpretation were further confirmed in a sentence rating study by bentzen and anderssen (2019), in which speakers of german were shown to prefer das to refer to non-nominal antecedents, with es being a less common but acceptable option when the antecedent is used as a “continuing topic”. a corpus study comparing english, danish and italian (navarretta, 2007) confirms this pattern, finding that in italian the difference between personal and demonstrative pronouns with regard to their type of coreference is not significant, unlike in english, where a contrast in use individuates different kinds of anaphors. in the italian data, however, null pronouns were used for abstract referents in only 30% to 42% of instances (respectively in the original-italian and parallel translations corpora). this is in keeping with antecedent accessibility hierarchies that position zero anaphora as requiring the highest level of accessibility (cf. givón, 1983; ariel, 1988, 2004). assuming entities to be more stereotypical antecedents than events, as shown by crude annotation distributions (see §1) the italian and spanish null pronouns should be employed less for abstract anaphora. more generally, the possibility of referencing an event with it or this (or that) is captured in models of discourse deixis. in line with bentzen and anderssen’s study, webber (1991) posits that propositions, particularly those that are focussed, are referred to first with either this or that in immediately following sentences. in later instances of coreference with those same events, the pronoun it can then be used. the difference between the two demonstratives is tested by çokal et al. (2014), finding that both this and that refer preferably to the most recent clause. building on centering theory (grosz et al., 1995), passonneau (1989) analyzes intra-sentential instances of it vs that with an explicit np antecedent, finding that it is used to refer to the center (most often the subject), whereas that favors the less prominent non-centers. for french, cornish (2015) shows how the demonstratives cela and ça, along with the neuter clitics le, y and en, can retrieve a sentential or verb-phrasal antecedent, and this relation is not simply formal-syntactic. a further claim is made comparing french with english, and saying that the anaphoric distinction between nominal and non-nominal antecedents has a correspondence to the forms: personal pronouns such as she/he and elle/il can only be used for nominal antecedents, while other expressions such as do it/this/that in english and faire cela/ça in french form a contrasting set used exclusively for non-nominal antecedents (such as predicates). cornish goes on to observe that in french the nominative neuter pronoun il is rarely used as a propositional anaphor (which will use ce/cela/ça), because it is often used as an expletive. in french corpus studies on the matter, tutin (2002) shows the presence of demonstratives used as anaphors, although much less frequently than personal pronouns and clitics, but the type of anaphora is not considered. vieira et al. (2005), comparing french and portuguese, classify antecedents in the two categories of “concrete” and “abstract”, where the first includes existing entities (such as people, things, places) and the second includes notions, actions and states. their results show that, in both languages, demonstratives are used for abstract antecedents in more than three quarters of cases. however, their study targets demonstrative noun phrases, that is demonstrative adjectives rather than pronouns, so the results may not be comparable. here, we target these same previously studied languages, but we adopt a story continuation methodology to probe the availability of entity and event antecedents for demonstrative, personal and null pronouns. where psycholinguistic approaches have been used previously to analyze it vs that, the emphasis has been on english. for example, brown-schmidt et al. (2005) report that comprehenders show a preference to interpret that to refer to a complex, composite antecedent (e.g., i’ll have the hamburger and fries. i’ll have that, too.), independent of other metrics of the salience 197 bevacqua, loáiciga, rohde and hardmeier of the referent. these results leave an open question regarding the crosslinguistic generalisability of such claims. 2.2 computational approaches reference to non-nominal antecedents has largely been a niche area in computational linguistics research (see review by kolhatkar et al., 2018). the most extensive annotation efforts in the field of coreference resolution have focused on english nominal coreference. ontonotes (pradhan et al., 2013), the largest and most frequently used corpus for training coreference resolution systems, for instance, only includes verbs if “they can be coreferenced with an existing noun phrase” according to its guidelines. corpora with a richer annotation of event pronouns exist, but are much smaller. one such resource is the arrau corpus (uryupina et al., 2020), whose size amounts to about 20% of version 5 of ontonotes. parcorfull (lapshinova-koltunski et al., 2018) also contains annotations of event pronouns in english, german and french. the scarcity of manually annotated resources has led to the use of artificial training data for coreference resolution systems of english non-nominal anaphora. kolhatkar et al. (2013) study the resolution of anaphoric shell nouns such as ‘this issue’ or ‘this fact’ by exploiting instances such as ‘the fact that. . . ’. marasovic et al. (2017) construct training examples based on specific patterns of verbs governing embedded sentences. also on this front, loáiciga et al. (2020) use parallel machine translation data in many languages for automatic data creation of event pronouns. before the breakthrough of neural end-to-end systems in coreference resolution (lee et al., 2017), coreference resolvers needed to do explicit mention classification in order to exclude nonreferential mentions before any resolution was attempted. in this context, the pronoun it has been targeted, as many of its uses are non-referential (e.g. it rains, it’s tuesday, it seems like). evans (2001) proposes the classification of the pronoun it into seven classes using contextual features. boyd et al. (2005) report similar results of around 80% accuracy on the classification of instances of it using more complex syntactic patterns. bergsma and yarowsky (2011) describe a system for identifying non-referential pronouns using statistics of word sequences (n-grams) from the world wide web, however without accounting explicitly for event reference. the many uses of it are also particularly relevant in dialogue, where event reference is much more common than in text data. however, there has been little recent research dedicated to coreference for general purpose dialogue systems, even though current text coreference resolution systems are not trained to manage dialogue data. in this context, müller (2007) proposes a disambiguation of it together with the deictic pronouns this and that. both eckert and strube (2000) and byron (2002) also report on systems able to resolve both nominal and non-nominal anaphora. working in a domain close to dialogue, lee et al. (2016) create a corpus for it-disambiguation in question answering. more recently, loáiciga et al. (2017) proposed a semi-supervised approach based on a combination of syntactic and semantic features for the classification of it. yaneva et al. (2018), on the other hand, report on experiments using features from eye gaze that prove to be more effective than any of the other types of features reported in previous works. this prior work points to the interest in understanding the behavior of pronouns like it, whose range of referential uses (entity, event, etc.) and non-referential uses presents a persistent challenge for computational systems. the approach we take here is to bring a technique from psycholinguistics to better model entity/event coreference. 198 event vs entity interpretation 3. experiments the work we present here tests entity/event coreference across five languages. we use a story continuation paradigm in which we present a context sentence and then assess how a subsequent referential form is interpreted. we collect story continuations in english, french, german, italian and spanish, measuring the rate of entity-versus-event coreference. we manipulate the heaviness of referential forms (null or personal pronoun versus demonstrative pronoun) along with two previously studied properties of the context sentence: aspect (perfective versus imperfective) and antecedent complexity (sentences with versus without an additional implicit argument, as determined by the alternating/non-alternating status of the verb). the goal is to use the same methodological probe and systematic context manipulations to test coreference biases across an inventory of coreferential forms in english, french, german, italian and spanish. 3.1 choice of target pronouns our chosen target pronouns can be seen in table 1. while the english pronouns are the same as those used in our previous study (loáiciga et al., 2018), where we targeted the personal pronoun it and the demonstrative this, other factors guided the choice in the other four languages. among our set of 5 languages, only italian and spanish allow a null subject in their inventory of possible referential forms. given the existence of theories pointing to a division of labour between null and overt pronouns in the two languages (see e.g. carminati, 2002, for italian, applied to spanish in e.g. alonso-ovalle et al., 2002), and given hierarchies positioning null-anaphors at the highest point on accessibility/prototypicality scales (see givón, 1983, and ariel, 1988), our study allows us to check whether a division of labour also applies to event vs entity coreference: for example, if entity antecedents are more available/accessible, a null pronoun may corefer more easily to an entity, while an event could be picked up by an overt form. other than the null, in spanish we targeted the masculine proximal demonstrative adjective este and the neuter demonstrative pronoun esto. while esto is a singular-only pronoun, este can also be used as a determiner and its status is sometimes controversial (rae, 2010, §17.2.2b). moreover, the real academia española states that esto (and the other neuter demonstratives eso and aquello) are used to refer to inanimate entities or propositional content (while their use for animals is uncommon and for people, offensive) (§17.2.5b); however, the function of referring to propositional content is not mentioned for este (nor for the analogues ese and aquel). in italian we targeted the masculine proximal demonstratives questo and ciò: the two demonstratives were chosen following up from navarretta’s (2007) corpus study, in which questo had a slight preference for entities and ciò had a preference for events. specifically, serianni (1997, p. 198) describes ciò as a pronoun of neuter value (= this thing, that thing), and states it is of very common use, especially in written language, while spoken language would more commonly use questo (this) or quello (that). for french, we chose the two pronouns il and cela and the expression c’est, composed of the particle ce and the copula in the third person singular of the indicative present tense. our use of c’est reflected our concern that the bare form ce would be interpreted as a determiner (ce chien), limiting our ability to observe entity-vs-event coreference with the demonstrative np ce itself. using c’est prompts also allowed us to avoid the problem that the bare form ce would not have been compatible with a continuation with the verb est because of the obligatory vowel elision, while using the letterpunctuation combination c’ in isolation posed a risk of confusion for the participants. il is a third199 bevacqua, loáiciga, rohde and hardmeier english it, this french il, cela, c’est german es, das, dies italian null, ciò, questo spanish null, esto, este table 1: pronominal prompts in the five languages. person personal pronoun, it is the dummy subject in impersonal forms2 and is used anaphorically to retrieve an infinitive or subjunctive clause. the demonstrative cela is a neutral form which is used for non-gendered antecedents like propositions. the grevisse grammar (grevisse and goosse, 2008, §691ff.) includes in this neuter category, along with the distal pronoun cela and its proximal counterpart ceci, also ça and ce. cela is used more in written language and can only refer to people in an informal register, while being normally used to refer to something “that we cannot name with precision” (§698.c). cela and ça have replaced ce in most of its uses, but ce is still used for inanimate objects or to refer back to a sentence, and followed by a copula it refers to “what comes before or the situation” (§702). for german, the duden grammar (dudenredaktion, 2009) says that dies is a shortened form of dieses which is predominantly used as a pronoun rather than a determiner, while diese can be used both ways (§372). in this account, phoric and deictic reference are distinguished such that the first links to referents without pointing explicitly while the second explicitly points to an object of discourse (§1818): textual content can be referred to phorically with es and deictically with das (§1821). finally, while anaphoric personal pronouns can refer to nouns in distant sentences, the grammar explains that a demonstrative like dieser links partly anaphorically and partly deictically to the closest nominal candidate for reference (§1827), without addressing the potential of dies and dieses to refer to non-nominal antecedents. 3.2 design and materials the structure of the experiments in the five languages was the same. in each language, the 24 experimental items consisted of a sentence describing a situation followed, after a full stop and a line break, by a pronominal prompt and a space to type in a continuation. the pronominal prompts were manipulated within items. for the null pronoun condition in italian and spanish, the continuation box was simply not introduced by a pronoun. the pronominal prompts used in the different languages are reported in table 1, and examples of items in different conditions of alternation and aspect for each language are given in (1)–(5): (1) english: a. noalt, imperf: the colonial building was collapsing slowly. it/this ... b. alt, perf: the cake for the guests cooked poorly. it/this ... (2) french: a. noalt, perf: le the bâtiment building colonial colonial a has coulé collapsed sous under la the neige. snow. il/cela/c’est ... 2. interestingly, this function could historically be taken by il and cela alike, a use surviving in some expressions (e.g. il/cela est vrai, “it/that’s true”). 200 event vs entity interpretation b. alt, imperf: le the pain bread que that j’avais i-had mis put au in-the four oven cuisait was-cooking encore. still. il/cela/c’est ... (3) german: a. noalt, perf: das the prachtvolle magnificent gebäude building zerfiel decayed über over die the jahre. years. es/das/dies ... b. alt, imperf: das the geräusch, noise das that ich i hörte, heard erstarb died unversehens. suddenly. es/das/dies ... (4) italian: a. noalt, perf: il the palazzo palace coloniale colonial è is collassato collapsed improvvisamente. suddenly. questo/ciò/null ... b. alt, imperf: il the vino wine piemontese piedmontese invecchiava was-aging bene. well. questo/ciò/null ... (5) spanish: a. noalt, perf: el the problema problem burocrático bureaucratic aparecía appeared de once nuevo. more. esto/este/null ... b. alt, imperf: el the dolor pain que that me me molestaba was-bothering se refl pasó disappeared de repente. suddenly. esto/este/null ... participants of every group saw all 24 experimental items, presented with an even distribution of the language’s pronoun prompts, interleaved with 42 fillers. of these, 18 were items of an unrelated experiment involving named entities, 20 were real fillers including a context sentence concerning two referents followed by an adverbial prompt in a new sentence (e.g., however, because of this), and four were control items with unambiguous obvious responses (e.g., wilma played a led zeppelin guitar solo in front of the crowd. it was stairway to _____). the stimuli were as similar as possible across languages, deviating from a literal translation when it would not have sounded natural. the stimuli also included between-item manipulations of the status of the event as completed or ongoing. half of the stimuli appeared in the perfective and half in the imperfective. the aspectual forms used were the present perfect and past continuous for english and their corresponding forms in the other languages: the passé composé and the imperfect in french, the passato prossimo and imperfect in italian, and the preterite and imperfect in spanish. in german, where aspect is not encoded in tenses, the präteritum was used; the aspect could in this case be inferred either through adverbials or contextually. moreover, half of the stimuli had an alternating verb and half had a non-alternating verb (in a pattern that did not correspond to the perfective/imperfective split). the examples in (1)-(5) all include a non-alternating verb, and they exemplify different combinations of the two between-item manipulations of aspect and alternation. note that verbs in an inchoative form can be syntactically different in italian and spanish from the other three languages. in particular, a reflexive particle can enter the construction. while in italian this depends on the verb and can be avoided (as in (4-b)), in spanish these verbs require a 201 bevacqua, loáiciga, rohde and hardmeier n (nm) age: range mean σ english 42 (36) 22–70 37.1 11.3 french 42 (22) 18–55 30.9 10.0 german 31 (25) 18–55 31.9 10.9 italian 43 (31) 18–48 29.7 8.3 spanish 45 (27) 18–67 33.0 9.7 203 (141) 18–70 32.4 10.2 table 2: summary statistics regarding the participants. the first column indicates the total number of participants as well as the number of monolingual participants (in brackets). reflexive particle, as in (5-b), which literally translates to “the pain that annoyed me passed itself suddenly”. in italian we avoided all verbs where a reflexive was necessary to ensure grammaticality. in spanish the only possibility was including the reflexive in all alternating verbs. while it could be hypothesised that the reflexive, which corefers with the entity antecedent by necessity, could prime an entity reading, our results show the opposite direction, with these verbs prompting more event continuations (see §3.5.5): since no priming can be noted and the inclusion of the reflexive is mandatory in spanish, we consider these stimuli comparable to those in the other four languages. 3.3 participants and procedure participants were recruited on amazon mechanical turk, targeting users from specific countries by ip address (i.e. usa, france, germany, italy and spain). they received $8/e 7 for an estimated 45-60 minutes task. we excluded non-native speakers from all analyses. summary demographics for the resulting datasets by language are reported in table 2. participants were not controlled for class or education level. to decide whether to include bilinguals since birth or only monolingual speakers of the target language, this factor was added to the models and a model comparison was run between the two models: in none of the languages the addition of this information improved the model’s fit (with p ranging between 0.12 and 0.91 for the bilingualism factor). bilingual participants were thus included in the data set. continuations were collected via a web-based interface that participants could access from their own computer through mturk. the website displayed a background questionnaire, a consent form and an instructions page, then each item was presented on a page by itself with a text box for participants to use for writing their continuation. 3.4 annotation the continuations were double-annotated for event or entity coreference. the annotators, who included the authors and were either native speakers or very competent speakers with a background in linguistics, based their decision on the semantics of the continuation. the pronominal prompt was hidden during the annotation process to avoid influencing the interpretation. continuations which did not include a subject-position reference to the event or entity included in the prompt were excluded from analysis. this included a total of 1345 continuations, equivalent to 23.9% of the 202 event vs entity interpretation initially gathered data, such as pleonastic uses (e.g. it was still foggy), adjectival uses (e.g. this wine was great), or reference to another unrelated entity (e.g. this is a beautiful morning).each remaining continuation was annotated with a strict and a liberal interpretation in parallel. in the strict annotation, an example was annotated as ambiguous if there was the slightest doubt about the correct reading. in the liberal annotation, annotators were allowed to use their intuitive judgements when both readings were possible, but one of them seemed overwhelmingly more likely. note that this means that some cases deemed ambiguous in the strict annotation are disambiguated in the liberal annotation. we chose to annotate the data both strictly and liberally to be able to evaluate the trade-off between quantity and noise of the data: the strict annotation necessitates excluding more data points from the analysis, but the liberal annotation leads to more noisy (although more numerous) data. the comparison between models computed on the strict and liberal annotation is described in section §3.6. to reconcile the double annotations, the following rules were applied: • the labels used for the valid continuations were: event, entity, ambiguous; • if both annotators agreed on a label in the strict annotation, the same label was also assigned for the liberal annotation; • when the two annotators disagreed (that is, one annotation was “event” and the other was “entity”), the continuation was excluded from the analyses (as were invalid continuations);3 • in the liberal annotation scheme only, if one annotator assigned “event” or “entity” and the other labelled it as “ambiguous”, the event/entity label was chosen (i.e. if one annotator did not resolve an ambiguous reading, the opinion of the other prevailed): these continuations are only included in the liberal analyses in §3.6. 3.5 results with the strict annotation the strict annotation was used as main analysis: comparing the same models computed with data annotated strictly and liberally, the models with strict data showed better fits (with, e.g., bics being up to halved). the results from the liberal analyses are reported in §3.6 as a comparison with the strict analyses and in appendix a. given that the analysis on the strict data was later repeated with a different model (as reported below in §3.7), a bonferroni correction was applied to the significance thresholds dividing them by two (the number of total analyses). these bonferroni-corrected thresolds as well as the results of the models are reported in table 3. in all languages, non-native participants were excluded from the analysis, and the data was subset to continuations annotated as either event or entities (in the strict annotation scheme). participants who described themselves as native but not monolingual were included in the analysis: model comparison showed that the addition of this further variable to the models did not improve the fit. the number of native participants and of monolinguals is reported in table 2, the total number of data points for each language is reported in the respective sections. alternation and aspect were deviation-coded in all languages: this means that, when reporting the results, only one estimate (the effects of perfective aspect and alternation) is reported; the significance and estimate for the opposite value of each factor (imperfective and non-alternating) are the same, with opposite signs. the 3. the percentages of disagreements for the strict annotation were: english, 17.7%; french, 21.9%; german, 13.7%; italian, 18.1%; spanish, 4%. 203 bevacqua, loáiciga, rohde and hardmeier coding of the referring expression factor is three-way in all languages but english, and it is detailed in each language’s section. all models used generalised mixed-effects logistic regression (bates et al., 2015b) and were computed with the lme4 package (bates et al., 2015b) in r (r development core team, 2008). 3.5.1 english the analysis for english should replicate the results of loáiciga et al. (2018), wherein we found a bias for it to corefer with entities and this with events, along with an effect of verb type, with verbs permitting alternation yielding more event coreference. the current analysis adds verbal aspect as a further predictor. the data included a total of 539 observations. the model predicted the isevent binary outcome using referential form (it vs this), aspect, alternation, and their interactions as binary predictors. all predictors were deviation-coded. to select the best fitting model, models with different fixed effects interactions were compared. the interactions of alternation and aspect, form and aspect or form and alternation did not significantly improve the fit (p = 0.52, p = 0.31 and p = 0.75, respectively). the chosen model thus includes fixed effects for form, aspect, and alternation, with no interactions. the maximal random effect structure was used when supported by the data (barr et al., 2013). where a model did not converge, the random effects were successively removed, chosen by lowest variance. the maximal converging model includes random intercepts and slopes for alternation by participant and for referential form by item. following the recommendation of bates et al. (2015a), we ran a principal components analysis of the random effects structure, which did not diagnose any overspecification. the model results are reported in table 3. the model shows that form has a significant effect where the use of this increases event coreference (p < 0.0005, see figure 1). verb alternation and aspect did not show a significant effect (respectively p = 0.42 and p = 0.93). 3.5.2 french the french analysis followed a similar procedure (cf. §3.5.1). the total number of observations was 582, verb alternation and aspect were deviation-coded and the three-way form was coded as the differences “c’est − cela” and “il − c’est”. model comparison showed that the best model was one with no interactions: adding an interaction of alternation and aspect, form and aspect or form and alternation did not improve the model fit (respectively, p = 0.25, p = 0.25, and p = 0.49). the model included a random intercept by participant and random intercept and slopes for the form by item. the results of the model are reported in table 3. the fixed effects show a significant difference between il and c’est (p < 0.0005), with il yielding fewer event continuations than either of the other two variants. the difference between c’est and cela, on the other hand, is not significant (p = 0.48). verb alternation did not reach significance, with alternating verbs prompting more event continuations (p = 0.05), nor did aspect (p = 0.82) given the three-way nature of form, its effects were confirmed by subsetting the data to two of the three variants. in the model with il and c’est only, as well as in the that with il and cela only, condition proved significant (both p < 0.0005 and with the same direction of effect as in the full model); in the model with c’est and cela only, condition was not significant (p = 0.25), 204 event vs entity interpretation confirming that c’est and cela do not pattern significantly differently in how they bias event or entity coreference (cf. figure 1). alternation was significant only in the model with il and c’est (p = 0.04), where alternating verbs prompted more event continuations. 3.5.3 german the german data included a total of 405 observations. as form is a three-way factor (es vs das vs dies), it was centred as the difference between pairs of its values, i.e. “dies − das” and “es − dies”. following a similar procedure as that outlined in §3.5.1 and §3.5.2, models were compared to select significant interactions. however, no interaction significantly improved the model fit (alternation and aspect: p = 0.7, form and aspect: p = 0.28, form and alternation: p = 0.18), so the chosen model includes predictors for the form, aspect, and alternation with no interactions. the maximally converging model includes random intercepts by participant and by item. the results of the model are reported in table 3. a significant difference is confirmed between es and dies, wherein dies also yields more event coreference than es (p < 0.0005). dies and das are only borderline different in their influence on coreference (p = 0.06, cf. figure 1). alternation and aspect did not reach significance (both p = 0.07), however the direction of effect for alternation is the same one as in our previous english study and in the french data (alternating verbs being biased towards events). the effects of form were confirmed subsetting the data into pairs of the three possible pronouns and fitting new models. form was significant in all three models (with largest p = 0.04), with the directions expected from the full model. verb aspect showed a significant effect, with perfective verbs prompting more event continuations, in all three models (largest p = 0.04), while alternation showed a significant effect in the model with es and dies and in that with es and das, with alternating verbs yielding more event continuations (both p = 0.04). 3.5.4 italian the italian analysis also followed a similar procedure (cf. §3.5.1-3.5.3). the three-way form was coded as the differences “ciò − null” and “questo − ciò”. the total number of observations was 343. model comparison showed that the interaction of form and verb alternation significantly improved the model fit over a model with no interactions (p = 0.04), while the other two interactions did not (alternation and aspect: p = 0.48, condition and aspect: p = 0.66). the random effects included random intercepts by participant and item. the model results are reported in table 3. the model output shows a general tendency for event continuations in the intercept, due to our choice of referring expressions in the design (p < 0.0005), and a significant difference between the null pronoun and ciò (p < 0.0005, see figure 1), whereby the null pronoun yields fewer event continuations, while the difference between the two overt pronouns is not significant (p = 0.36). neither verb alternation nor aspect reached significance (respectively, p = 0.92 and p = 0.3). neither interaction reached significance. again, the effects of form were confirmed subsetting the data to two of the three variants. form was found to be significant in the model with the null pronoun and ciò and in that with the null pronoun and questo (both p < 0.001), but not in the model with the two demonstratives (p = 0.79), confirming that the pronoun with a different pattern with regard to event vs entity coreference is 205 bevacqua, loáiciga, rohde and hardmeier the null pronoun, which is biased towards entity continuations. in the model including the null pronoun and questo only and the one including questo and ciò, the interaction between form and alternation was significant: specifically, both null pronouns and ciò were used proportionally less than questo to refer to events when an alternating verb was present (respectively p = 0.04 and p < 0.001). 3.5.5 spanish the spanish analysis followed a similar procedure as the other languages (cf. §3.5.1-3.5.4). the three-way form was coded as the differences “esto − este” and “null − esto”. the number of observations was 453. like for italian (§3.5.4), the model with an interaction between form and alternation improved the fit (p = 0.01), while the other interactions did not (alternation and aspect: p = 0.3, form and aspect: p = 0.18). the model thus included this interaction alongside the main effects, as well as random intercepts by participant and item. the results of the model are reported in table 3. the model shows a significant difference between esto and este, and between the null pronoun and esto (both p < 0.0005): esto yields more event coreference than either este or the null pronoun. while aspect did not seem to have an effect on event vs entity coreference (p = 0.51), there was a main effect of alternation whereby verbs allowing alternation were biased towards event coreference (p = 0.001). alternation was also significant in interaction with “null − esto”: null pronouns are used more than esto to refer to events in the presence of alternating verbs (p = 0.007), modulating the direction of the main effect. the effects of form were confirmed subsetting the data to two of the three pronouns. form was significant in all models: in the model with the null pronoun and esto only, as well as in the that with este and esto only, with the same direction of effect as in the full model (respectively p = 0.03 and p < 0.001); in the model with the null pronoun and este, with este yielding fewer events than the null pronoun (p = 0.002) (cf. figure 1). alternation was significant in the two models including null pronouns (in the model with null and esto, p = 0.009; in that with null and este, p < 0.001); moreover, in these two models alternation was also significant in interaction with form: with alternating verbs, null pronouns produce proportionally even more event continuations than either este or esto (both p = 0.02). 3.6 results with the liberal annotation as previously mentioned, the models using the liberal annotation of event or entity score worse in measures of the relative model quality like the bic, aic and log-likelihood. we therefore chose to base our analysis on the models with strict annotation. although the results of the two types of annotation are very similar, it is worth pointing out some of the differences. in italian, the interaction between conditions and alternation no longer improved the model (p = 0.43), and was as such not included. on the other hand, with the liberally annotated data the interaction between alternation and aspect improved the model fit in english and german (respectively p = 0.03 and p < 0.001). in english this interaction reached significance: alternating verbs in the perfective aspect had a smaller effect than in the imperfective in eliciting event continuations (p = 0.02). as a main effect, alternation often reached significance, with alternating verbs yielding more event coreference in both english and french (respectively p = 0.04 and p = 0.03), as well as in 206 event vs entity interpretation effect estimate std.error z value pr(> |z|) english (intercept) 0.21 0.52 0.41 0.68 condition (this) 6.71 0.98 6.87 < 0.0005 ∗∗∗ aspect (perf) −0.07 0.80 −0.08 0.93 alternation (alt) 0.87 1.07 0.81 0.42 french (intercept) 0.10 0.53 0.85 0.85 c’est − cela −0.42 0.59 −0.71 0.48 il − c’est −7.02 1.21 −5.81 < 0.0005 ∗∗∗ aspect (perf) 0.19 0.84 0.22 0.82 alternation (alt) 1.70 0.87 1.95 0.05 german (intercept) 1.06 0.49 2.15 0.03 dies − das −1.30 0.69 −1.88 0.06 es − dies −6.79 0.88 −7.71 < 0.0005 ∗∗∗ aspect (perf) 1.47 0.82 1.80 0.07 alternation (alt) 1.47 0.80 1.84 0.07 italian (intercept) 2.26 0.46 4.92 < 0.0005 ∗∗∗ ciò − null 5.99 1.08 5.54 < 0.0005 ∗∗∗ questo − ciò −0.78 0.85 −0.92 0.36 aspect (perf) 0.72 0.69 1.04 0.30 alternation (alt) −0.07 0.79 −0.10 0.92 (ciò − null):altern −0.27 1.52 −0.18 0.86 (questo − ciò):altern 3.04 1.69 1.80 0.07 spanish (intercept) 0.41 0.40 1.03 0.31 esto − este 7.16 0.86 8.37 < 0.0005 ∗∗∗ null − esto −5.83 0.88 −6.60 < 0.0005 ∗∗∗ aspect (perf) 0.43 0.65 0.66 0.51 alternation (alt) 2.36 0.72 3.27 0.001 ∗∗ (esto − este):altern −1.05 1.13 −0.93 0.35 (null − esto):altern 3.85 1.43 2.69 0.007 ∗ bonferroni-corrected significance: ∗∗∗p < 0.0005, ∗∗p < 0.005, ∗p < 0.025 table 3: estimated models fixed effects (strict annotations). 207 bevacqua, loáiciga, rohde and hardmeier figure 1: event and entity coreference by pronominal prompt in the five languages. figure 2: event and entity coreference by verb type in the five languages. figure 3: event and entity coreference by verb aspect in the five languages. 208 event vs entity interpretation spanish, where it also showed significance in the strict annotation (p = 0.001). in the italian model, alternation did not reach significance (p = 0.81), but the perfective aspect significantly yielded more event continuations, unlike in the strict model (p = 0.04). finally, in the spanish model the interaction between alternation and “null−esto” reached significance (p = 0.03): while generally esto yields more event continuations than the null pronoun, this effect is less pronounced in alternating verbs than in non-alternating verbs. 3.7 cross-language model given the p values of alternation often being borderline (particularly before the bonferroni correction), as in french (p = 0.05) and german (p = 0.07), a model was run on the strict annotation including data from all five languagees. this model was only run on the strict annotation of the data. to diminish the risk of type i error, brought on by repeating the analysis a second time, a bonferroni (1936) correction was applied to the significance levels, dividing them by two (the number of languages and, accordingly, previous models). the bonferroni correction was also retro-actively applied to the model described in §3.5 and reported in table 3. the model results are reported in table 4. while adding interactions between aspect and alternation or aspect and condition did not improve the model’s fit over a model with no interactions (respectively p = 0.32 and p = 0.79), the interaction of alternation and condition did (p = 0.005), and was thus included in the model specification. a model with the three-way interaction did not converge. the selected model modelled event coreference as a binary, with verb aspect, alternation, referring expression and the interaction of alternation and referring expression as predictors, as well as random intercepts for participant and item, both nested within language. all the referring expressions in table 1 were included, distinguishing between null pronouns in italian and in spanish. all predictors were sum-coded: the model estimates are then to be read as a deviation from the intercept, i.e. as a bias of the referring expression towards events or entities (respectively positive and negative β). while aspect did not show a significant crosslinguistic effect (p = 0.24), alternation did (p < 0.001), with alternating verbs yielding more event continuations. all referring expressions deviated significantly from the intercept, or, in different terms, their yield of event or entity continuations was not at chance level (all at p < 0.001). more specifically, personal pronouns were biased towards entity coreference and demonstratives were biased towards events, with the exception of este in spanish, which yields more entity coreference (as already seen in the language-internal model, §3.5.5). finally, the interactions between the italian and spanish null pronoun conditions and alternation reached significance: while in italian the null pronoun reduces the effect of alternating verbs (p = 0.01), in spanish it enhances it (p = 0.003). this echoes findings on the distribution and behaviour of null pronouns in the two languages, which seem to differ in similar contexts (e.g. russo et al., 2012; filiaci, 2010). 4. discussion the results of the five studies display similarities between our target languages. in all languages, heavier referring expressions (specifically, demonstratives) bias coreference towards event continuations (as shown in figure 1). this is in keeping with theories positing that richer, more uniquely209 bevacqua, loáiciga, rohde and hardmeier effect estimate std.error z value pr(> |z|) (intercept) 0.67 0.17 3.92 < 0.0005 ∗∗∗ aspect (perf) 0.17 0.14 1.17 0.24 alternation (alt) 0.59 0.16 3.74 < 0.0005 ∗∗∗ english condition (it) −3.36 0.38 −8.75 < 0.0005 ∗∗∗ condition (this) 2.04 0.39 5.25 < 0.0005 ∗∗∗ altern:it 0.09 0.34 0.27 0.78 altern:this −0.52 0.35 −1.47 0.14 french condition (il) −5.36 0.65 −8.26 < 0.0005 ∗∗∗ condition (cela) 1.71 0.37 4.57 < 0.0005 ∗∗∗ condition (c’est) 1.31 0.38 3.48 < 0.0005 ∗∗∗ altern:il 0.64 0.61 1.04 0.30 altern:cela −0.35 0.34 −1.01 0.31 altern:c’est 0.19 0.35 0.55 0.59 german condition (es) −3.51 0.38 −9.13 < 0.0005 ∗∗∗ condition (das) 2.12 0.42 5.03 < 0.0005 ∗∗∗ condition (dies) 2.01 0.37 5.46 < 0.0005 ∗∗∗ altern:es 0.14 0.35 0.41 0.68 altern:das −0.53 0.40 −1.34 0.18 altern:dies −0.18 0.34 −0.54 0.59 italian condition (null) −2.44 0.45 −5.49 < 0.0005 ∗∗∗ condition (ciò) 3.24 0.52 6.28 < 0.0005 ∗∗∗ condition (questo) 3.09 0.60 5.16 < 0.0005 ∗∗∗ altern:null −1.07 0.41 −2.58 0.01 ∗ altern:ciò −0.60 0.48 −1.25 0.21 altern:questo 0.52 0.58 0.90 0.37 spanish condition (null) −1.69 0.56 −3.00 0.003 ∗∗ condition (este) −3.04 0.43 −7.06 < 0.0005 ∗∗∗ condition (esto) 3.88 0.51 7.57 < 0.0005 ∗∗∗ altern:null 1.62 0.55 2.94 0.003 ∗∗ altern:este 0.28 0.39 1.71 0.48 altern:esto −0.23 0.48 −0.48 0.63 bonferroni-corrected significance: ∗∗∗p < 0.0005, ∗∗p < 0.005, ∗p < 0.025 table 4: estimated models fixed effects for the cross-language model. 210 event vs entity interpretation referential expressions (like demonstratives, that add a distal trait to those of personal pronouns and null forms) will be used to retrieve less “stereotypical” material (e.g. ariel, 1988): given that event coreference is much rarer than entity coreference (cf. §1), it is likely that events will be implicitly treated as a marked case, granting the use of expressions used for less accessible antecedents. a notable exception to this pattern is the spanish demonstrative este, which was not only biased towards entity continuations, but even more so than the null pronoun. one possibility is that, unlike its neuter counterpart esto, the masculine form este could be used preferably for (masculine) entities, for analogy with the feminine form esta. note that the masculine is the default gender of the entities used in the experiment. the results of the studies often agree with the (rough) prescriptions of grammars. the rae (2010) predicted the use of spanish esto to refer to antecedents other than entities, and similarly in french the grevisse and goosse (2008) grammar predicted the use of cela and ce for non-entities. the italian data agrees with serianni (1997): both ciò and questo are dispreferred for entity antecedents. for german, the duden (2009) grammar does not provide a very clear indication of use, in that it says that both es and das can be used to refer to non-entities (however differently) – yet, our data shows a marked tendency for es to refer to entities. however, our studies differ from these prescriptions in that they offer an estimation of the degree to which the different forms are non-categorical. comparing the results with those of corpus and psycholinguistic studies also yields some interesting observations. the english results confirm a contrast in use between the personal and demonstrative pronouns already found in loáiciga et al. (2018), navarretta (2007) and dipper et al. (2011), which also agrees with our results on german, in which es continues referential chains of topical entities while das refers to non-nominal antecedents. dipper et al. (2011) does not find a clear distinction in use between demonstrative and personal pronouns in spanish, and our results add some detail to the claim showing how este does not pattern as expected. in italian, navarretta (2007) finds that while ciò is used for non-entities, questo has a light preference for entities: this is in contrast with the very clear tendency we found for the demonstrative to retrieve events. the french results are in line with vieira et al. (2005) in that demonstratives are used for abstract antecedents more often. finally, our studies targeted two verbal features: alternation and aspect. alternating verbs yielded more event continuations in more languages, even if the effect did not reach significance in italian. in italian, the two demonstratives show ceiling selection of the events as antecedents: this near-categorical behaviour may have obscured the effect. the crosslinguistic effect of alternation is visualised in figure 2. on the other hand, aspect did not show a significant effect across languages; nonetheless, the direction is quite consistently the same in which the representation of an event as completed through a perfective aspect yields more events (except for english: see figure 3). this, however, was only significant in the italian model run with the liberal annotation of our data and in interaction with alternation in english, whereby non-alternating verbs yield more event continuations in the perfective aspect, whereas alternating verbs do so in the imperfective. 211 bevacqua, loáiciga, rohde and hardmeier 5. comparison with annotated coreference in order to contexualise our findings in the coreference resolution scenario, we extract coreference relations from a multilingual parallel corpus to check whether the same patterns observed in the previous experiments are also apparent in the annotated corpus data. we work with the parcorfull corpus, which includes coreference annotations for the english, german and french languages. parcorfull includes texts from ted talks transcripts and newswire data originally in english, and comprises approximately 160,000 tokens. the underlying coreference scheme was designed for uniform annotations across the languages (see lapshinova-koltunski and hardmeier, 2018, for details). the annotated elements (markables) in this corpus include pronouns, nouns, nominal phrases or elliptical constructions that are parts of a coreference pair (antecedent-anaphora), as well as verb phrases or clauses being antecedents of event anaphora. the annotated antecedents are of two different types: entities and events. entities can be either pronouns or noun phrases, whereas events include verb phrases, clauses, or a set of clauses. 5.1 event vs entity antecedent proportions to reproduce the parameters of the experiments described above, we first extract all coreference chains headed by lexical entity or event antecedents, thus excluding cataphora. we then retain all mentions of the pronouns of interest (en: this, it; de: es, das, dies; fr: c’, il, cela) in a subject position and order them according to their appearance in the text. table 5 summarises the distribution of the mentions retained after filtering. the index i = n reflects the order of re-mention of the antecedent by a pronoun of interest. i = 1 corresponds to equivalent cases to those produced in the experiments with human participants, where a particular type of antecedent is re-mentioned for the first time with one of the different pronouns. i = 2 corresponds to the second re-mention of an antecedent, i = 3 to the third, and so forth. note that while this is a parallel corpus, the number of times there is a re-mention of an antecedent varies in each of the languages, with french being the language with the most re-mentions. additionally, for all the languages, it can be seen that there is a clearly preferred form for re-mention pervasive through time as the text increases in length, regardless of the type of antecedent. these are light forms: it for english, es for german, and il for french. in order to assess webber’s 1991 proposal that demonstrative pronouns are used to refer for the first time to parts of the context which are in focus, making them available to serve as antecedents for personal pronouns later, we looked closely at the english examples corresponding to the columns i = 1 and i = 2 in table 5. from the 54 events first introduced with this, we found 4 examples rementioned with it, as illustrated in (6). these are too few cases to draw any clear conclusion, leaving the question open of whether this pattern is upheld more frequently with the demonstrative that.4 (6) but just if you take malaria out, deaths from everything else go down. and the economist jeff sachs has actually quantified what this means for a society. what it means is, if you have malaria in your society, your economic growth is depressed by 1.3 percent every year, year after year after year, just this one disease alone. 4. as for cases where a demonstrative pronoun is also rementioned with a demonstrative pronoun we found five cases with this; 16 with das, 2 with dies; and no cases in french. 212 event vs entity interpretation antecedent anaphor re-mention index english i = 1 i = 2 i = 3 i = 4 i = 5 i = 6 i = 7 i = 8 i = 9 i = 10 i = 11 i ≥ 12 entity this 24 21 5 1 0 1 1 0 0 0 0 0 it 220 113 63 34 22 15 4 0 0 0 2 9 event this 54 10 4 1 0 1 0 0 1 1 0 0 it 60 42 18 8 5 1 2 1 0 0 0 0 german i = 1 i = 2 i = 3 i = 4 i = 5 i = 6 i = 7 i = 8 i = 9 i = 10 i = 11 i ≥ 12 es 87 48 24 9 5 1 2 0 2 0 1 1 entity das 118 14 9 0 1 3 0 1 0 0 0 2 dies 4 3 1 0 1 0 0 0 0 0 0 0 es 20 17 8 3 1 0 0 0 0 0 0 0 event das 187 26 5 0 0 0 0 0 0 0 0 0 dies 10 3 1 0 0 0 0 0 0 0 0 0 french i = 1 i = 2 i = 3 i = 4 i = 5 i = 6 i = 7 i = 8 i = 9 i = 10 i = 11 i ≥ 12 c’ 74 38 19 12 9 2 1 1 1 0 2 2 entity il 62 47 26 21 20 13 11 5 6 4 5 56 cela 16 2 3 4 0 1 0 0 1 0 0 2 c’ 90 25 7 3 0 0 1 0 0 0 1 0 event il 1 3 0 0 0 0 0 0 0 0 0 0 cela 30 7 4 0 0 0 0 0 0 0 0 0 table 5: re-mention frequency of entity and event antecedents in english, german and french with the anaphors of interest as annotated in the parcorfull corpus. the index i = n represents the subsequent order in which the pronoun appears after an antecedent has been introduced. 213 bevacqua, loáiciga, rohde and hardmeier antecedent english german french this it es das dies c’ il cela human responses entity 22 282 140 5 21 37 180 33 4% 52% 35% 1% 5% 6% 31% 6% event 203 32 11 91 137 140 6 186 38% 6% 3% 22% 34% 24% 1% 32% corpus annotation entity 24 220 87 118 4 74 62 16 7% 61% 20% 28% 1% 27% 23% 6% event 54 60 20 187 10 90 1 30 15% 17% 5% 44% 2% 33% 0% 11% table 6: comparison between the proportions of event and entity antecedent interpretations in the human experiments with the annotations in parcorfull of antecedents and their first remention (i = 1). note that the percentages are computed using the total counts per language. a direct comparison between the experiments with human participants and the corpus data is summarised in table 6. as it is observed in the human continuations, events are mostly referred to with demonstrative pronouns and entities with personal pronouns, but this preference is not categorical. in english, for example, events in the corpus have a similar proportion of it vs this. in german, the clear preference is das, while in french is c’ and not cela as one might expect. while the human event coreference rates are more balanced across languages, as a result of controlling the parameters for the experiment, the proportion of events in the corpus varies per language. for english, there is around 32% events in contrast to 68% entities, while for german the proportion is evenly balanced with 51% events. the difference is worth noting since the documents in the corpus are complete parallel versions of each other, suggesting a difference in conceptualisation at the time of translation. french, on the other hand, is in the middle with 44% events. however, this portion of the corpus contains fewer documents than its english and german counterparts. 5.2 governing verb alternation status in order to draw a similar comparison with respect to the alternation status of the verbs in the corpus, we also extracted the verbs relevant for the entities and events reported in table 6. in the case of entity antecedents, we extracted the verb to which the head of the noun phrase is attached. for event antecedents, we extracted the verb from the antecedent itself. studies on the causative alternation propose that verbs can be ranked on a universal scale of likelihood of spontaneous occurrence (haspelmath, 1993). in this scale, verbs are ranked according to the degree to which they are non-agentive (samardžić, 2014, p. 180). this means that verbs ranking higher in spontaneous occurrence can occur without an explicit agent causing the event, e.g., break, open, and are hence more likely to participate in the causative alternation. following haspelmath and using large scale corpus data, as opposed to typological observations, samardžić 214 event vs entity interpretation estimates the degree of spontaneity for the 354 english verbs permitting the causative alternation reported by levin (1993). the underlying assumption for the spontaneity score is that causative and anticausative uses of a verb have a correlation with transitive and intransitive examples in corpora (reported to be r = 0.67, p < 0.01 using the spearman test). we then compare the verbs extracted from the parcorfull antecedents against the list of scored verbs reported in samardžić (2014). unfortunately, being restricted to levin’s verbs, this list is very small, and contains only a few of our verbs. the extracted english verbs which matched the list are listed in appendices b and c. we could not find a similar resource for french or german. it is very difficult to estimate a similar spontaneity score for our verbs since, unlike samardžić who worked with levin’s list and the proposition bank (palmer et al., 2005) which is annotated for semantic roles, we have many verbs for which we do not have any sort of gold standard resource that help us estimate the accuracy of our calculations. in our case, equating transitive uses with alternating verbs (causatives) and intransitive uses with non-alternating verbs (anti-causatives) is a dangerous simplification since many verbs in the corpus do not necessarily have the proper frame of thematic roles playing a part in the causative alternation. interpreting and generating coreference involves many levels of linguistic processing, including verb semantics (for the effect of thematic roles, see e.g. stevenson et al., 1994; arnold, 2000). between our experimental items, we manipulated the alternation status of the verb because the experimental framework grants us complete control over the stimuli. estimating semantic predictors from a corpus is much harder, partly because of the lack of resources annotated with this type of information. with our work, we highlight the value in considering semantic predictors for the processing of coreference. 6. conclusions our study shows a crosslinguistic bias whereby the coreference to entities or events is biased by multiple factors: the referring expressions used and features of the verb constituting the argument structure involving the possible entity antecedents. moreover, our data gives a clear distribution of event and entity coreference across referring expressions and languages, with controlled manipulations achieving less noise than in a corpus analysis. confirming the predictions based on hierarchies describing the use of referring expressions based on their antecedent’s accessibility (e.g. ariel 1988), our results show that lighter referring expressions are biased towards entity antecedents and, conversely, heavier expressions are biased towards events. the lesser accessibility of events can be explained in multiple ways: they are a less common type of antecedent (as corpus measures show: see §1, 5), they are more complex and less easily introduced directly (hedberg et al., 2007). the comparisons with a coreference-annotated corpus generally replicate the human results, but they also suggest that these patterns might be further influenced by the time at which an antecedent is re-mentioned in a coreference chain. the corpus results also show slight differences in the referring expressions’ proportions, which may be due to stylistic variation. in addition, the fact that the different forms appear non-categorically with either entities or events points towards an interplay between both types of antecedents in the discourse, indicating that studying one without the other might result in an incomplete picture. we also investigated verbal features and their possible influence on coreference. on the question of verb aspect, this study saw a general tendency for completed events to be taken up as an 215 bevacqua, loáiciga, rohde and hardmeier antecedent more than ongoing events in all languages but english, but this effect did not reach statistical significance. further research with fewer confounding factors is needed to confirm whether the mostly uniform direction it showed in our data is upheld. on the other hand, the event structure of verbs was also shown to clearly influence the coreference patterns. specifically, further increasing the complexity of an event by having an implicit agent in alternating verbs creates more competition for the explicit entity: since the accessibility of a referent decreases with the increase in the total number of possible referents (grosz et al., 1995), the event antecedent becomes relatively more accessible and is chosen more as a likely continuation. last, the cross crosslingual effect of the verb alternation shows that, from a cognitive point of view, adding one more competitor to a pool of possible antecedents affects the relative accessibility of the other antecedents, even when the newly added competitor is not made explicit. this raises questions about whether introducing implicit and explicit elements has the same effect on coreference (both qualitatively and, especially, quantitatively), and whether an antecedent’s conceptual availability as given by world knowledge matters as much as its explicit presence in a discourse. such differences in the activation of antecedents could be explored via “on-line” psycholinguistic methods (e.g. measurement of reaction times or eye-tracking). we must leave these questions for future research. acknowledgements hannah rohde was supported by a leverhulme trust prize in languages & literatures. christian hardmeier was supported by the swedish research council under grant 2017-930. the authors gratefully acknowledge the funding. we also thank sebastiano gigliobianco for his help optimizing our extraction scripts, and elina lartaud, ludovica inserra and paola jalili for helping with data annotation. last, are also grateful to the anonymous reviewers for their insightful comments that helped improve this paper. references luis alonso-ovalle, susana férnandez-solera, lyn frazier, and charles jr. clifton. null vs. overt pronouns and the topic-focus articulation in spanish. rivista di linguistica, 14(2):151–170, 2002. mira ariel. referring and accessibility. journal of linguistics, 24(1):65–87, 1988. mira ariel. accessibility marking: discourse functions, discourse profiles, and processing cues. discourse processes, 37(2):91–116, 2004. j. e. arnold. the effect of thematic roles on pronoun use and frequency of reference continuation. university of pennsylvania working papers in linguistics, 6(3):209–235, 2000. nicholas asher. reference to abstract objects in discourse. springer netherlands, dordrecht, 1993. hiromi azuma. a diachronic view of pronominal reference in english. in christer johansson, editor, second workshop on anaphora resolution, war ii, volume 2, pages 1–9, bergen, 2008. 216 event vs entity interpretation dale j. barr, roger levy, christoph scheepers, and harry j. tily. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3): 255–278, 2013. douglas bates, reinhold kliegl, shravan vasishth, and harald baayen. parsimonious mixed models, 2015a. url https://arxiv.org/abs/1506.04967. douglas bates, martin mächler, ben bolker, and steve walker. fitting linear mixed-effects models using lme4. journal of statistical software, 67(1):1–48, 2015b. kristine bentzen and merete anderssen. the form and position of pronominal objects with nonnominal antecedents in scandinavian and german. the journal of comparative germanic linguistics, 22(2):169–188, 2019. shane bergsma and david yarowsky. nada: a robust system for non referential pronoun detection. in iris hendrickx, sobha lalitha devi, antónio branco, and ruslan mitkov, editors, anaphora processing and applications: 8th discourse anaphora and anaphor resolution colloquium (daarc), lecture notes in artificial intelligence, pages 12–23. springer, faro, 2011. carlo bonferroni. teoria statistica delle classi e calcolo delle probabilità. pubblicazioni del r istituto superiore di scienze economiche e commericiali di firenze, 8:3–62, 1936. adriane boyd, whitney gegg-harrison, and donna k. byron. identifying non-referential it: a machine learning approach incorporating linguistically motivated patterns. in proceedings of the acl workshop on feature engineering for machine learning in natural language processing, pages 40–47, ann arbor, michigan, 2005. association for computational linguistics. sarah brown-schmidt, donna k. byron, and michael k. tanenhaus. beyond salience: interpretation of personal and demonstrative pronouns. journal of memory and language, 53(2):292–313, 2005. donna k. byron. resolving pronominal reference to abstract entities. in proceedings of the 40th annual meeting of the association for computational linguistics, acl 2002, pages 80– 87, philadelphia, 2002. association for computational linguistics. gregory carlson. reference. in l. r. horn and g. ward, editors, the handbook of pragmatics. wiley-blackwell, oxford, 2003. maria nella carminati. the processing of italian subject pronouns. doctoral dissertation, university of massachusetts amherst, 2002. url https://scholarworks.umass.edu/ dissertations/aai3039345. derya çokal, patrick sturt, and fernanda ferreira. deixis: this and that in written narrative discourse. discourse processes, 51(3):201–229, 2014. derya çokal, patrick sturt, and fernanda ferreira. processing of it and this in written narrative discourse. discourse processes, 55(3):272–289, 2018. doi: 10.1080/0163853x.2016.1236231. url https://doi.org/10.1080/0163853x.2016.1236231. 217 bevacqua, loáiciga, rohde and hardmeier bernard comrie. pragmatic binding: demonstratives as anaphors in dutch. in proceedings of the twenty-third annual meeting of the berkeley linguistics society: general session and parasession on pragmatics and grammatical structure, pages 50–61, berkeley, 1997. francis cornish. anaphoric relations in english and french: a discourse perspective. taylor & francis, abingdon-on-thames (uk), 2015. stefanie dipper, christine rieger, melanie seiss, and heike zinsmeister. abstract anaphors in german and english. in iris hendrickx, sobha lalitha devi, antónio branco, and ruslan mitkov, editors, anaphora processing and applications. 8th discourse anaphora and anaphor resolution colloquium, daarc 2011, pages 96–107. springer, faro, portugal, 2011. dudenredaktion. duden. die grammatik: unentbehrlich für richtiges deutsch. dudenverlag, mannheim–wien–zürich, 2009. miriam eckert and michael strube. dialogue acts, synchronising units and anaphora resolution. journal of semantics, 17(1):51–89, 2000. richard evans. applying machine learning toward an automatic classification of it. literary and linguistic computing, 16(1):45–57, 2001. todd r ferretti, ken mcrae, and andrea hatherell. integrating verbs, situation schemas, and thematic role concepts. journal of memory and language, 44(4):516–547, 2001. todd r. ferretti, marta kutas, and ken mcrae. verb aspect and the activation of event knowledge. journal of experimental psychology. learning, memory, and cognition, 3(1):182–196, 2007. todd r. ferretti, hannah rohde, andrew kehler, and melanie crutchley. verb aspect, event structure, and coreferential processing. journal of memory and language, 61(2):191–205, 2009. francesca filiaci. null and overt subject biases in spanish and italian: a cross-linguistic comparison. in selected proceedings of the 12th hispanic linguistics symposium, pages 171–182, somerville (ma), 2010. cascadilla proceedings project. url https://www.lingref. com/cpp/hls/12/paper2415.pdf. thomas givón. topic continuity in discourse: a quantitative cross-language study. john benjamin, amsterdam, 1983. s. b. greene, g. mckoon, and r. ratcliff. pronoun resolution and discourse models. journal of experimental psychology: learning, memory and cognition, 18:266–283, 1992. maurice grevisse and andré goosse. le bon usage: grammaire française. duculot, louvain-laneuve (belgium), 14th edition, 2008. barbara j. grosz, aravind k. joshi, and scott weinstein. centering: a framework for modelling the local coherence of discourse. computational linguistics, 21(2):203–225, 1995. jeanette k. gundel, nancy hedberg, and ron zacharski. pronouns without np antecedents: how do we know when a pronoun is referential? in antonio branco, tony mcenery, and ruslan mitkov, editors, anaphora processing: linguistic, cognitive and computational modelling, pages 351– 364. john benjamins, amsterdam, 2005. 218 event vs entity interpretation martin haspelmath. more on the typology of inchoative/causative verb alternations. in bernard comrie and maria polinsky, editors, causatives and transitivity, pages 87–120. john benjamins, amsterdam, 1993. nancy hedberg, jeanette k. gundel, and ron zacharski. directly and indirectly anaphoric demonstrative and personal pronouns in newspaper articles. in proceedings daarc 2007 (discourse anaphora and anaphora resolution colloquium), lagos (pt), 2007. benjamin heinzerling, nafise sadat moosavi, and michael strube. revisiting selectional preferences for coreference resolution. in proceedings of the 2017 conference on empirical methods in natural language processing, pages 1332–1339, copenhagen, 2017. association for computational linguistics. otto jespersen. modern english grammar on historical principles, part iii: syntax (second volume). allen and unwin, london, 1927. elsi kaiser and john c. trueswell. the role of discourse context in the processing of a flexible word-order language. cognition, 94:113–147, 2004. varada kolhatkar, heike zinsmeister, and graeme hirst. interpreting anaphoric shell nouns using antecedents of cataphoric shell nouns as training data. in proceedings of the 2013 conference on empirical methods in natural language processing, pages 300–310, seattle, washington, usa, 2013. association for computational linguistics. url https://www.aclweb.org/ anthology/d13-1030. varada kolhatkar, adam roussel, stefanie dipper, and heike zinsmeister. anaphora with nonnominal antecedents in computational linguistics: a survey. computational linguistics, 44(3): 547–612, september 2018. url https://www.aclweb.org/anthology/j18-3007. ekaterina lapshinova-koltunski and christian hardmeier. coreference corpus annotation guidelines, 2018. url https://lindat.mff.cuni.cz/repository/xmlui/ bitstream/handle/11372/lrt-2614/guidelines.pdf. ekaterina lapshinova-koltunski, christian hardmeier, and pauline krielke. parcorfull: a parallel corpus annotated with full coreference. in proceedings of 11th language resources and evaluation conference, pages 00–00, miyazaki, japan, 2018. european language resources association (elra). to appear. kenton lee, luheng he, mike lewis, and luke zettlemoyer. end-to-end neural coreference resolution. in proceedings of the 2017 conference on empirical methods in natural language processing, pages 188–197, copenhagen, denmark, september 2017. association for computational linguistics. timothy lee, alex lutz, and jinho d. choi. qa-it: classifying non-referential it for question answer pairs. in proceedings of the acl 2016 student research workshop, pages 132–137, berlin, 2016. association for computational linguistics. beth levin. english verb classes and alternations: a preliminary investigation. the university of chicago press, chicago, 1993. 219 bevacqua, loáiciga, rohde and hardmeier sharid loáiciga, liane guillou, and christian hardmeier. what is it? disambiguating the different readings of the pronoun ‘it’. in proceedings of the conference on empirical methods in natural language processing, emnlp 2017, pages 1336–1342, copenhagen, 2017. association for computational linguistics. sharid loáiciga, luca bevacqua, hannah rohde, and christian hardmeier. event versus entity co-reference: effects of context and form of referring expression. in proceedings of the first workshop on computational models of reference, anaphora and coreference, pages 97–103, new orleans, june 2018. association for computational linguistics. url https://www. aclweb.org/anthology/w18-0711. sharid loáiciga, christian hardmeier, and asad sayeed. exploiting cross-lingual hints to discover event pronouns. in proceedings of the 12th language resources and evaluation conference, pages 99–103, marseille, france, may 2020. european language resources association. url https://www.aclweb.org/anthology/2020.lrec-1.12. ana marasovic, leo born, juri opitz, and anette frank. a mention-ranking model for abstract anaphora resolution. in proceedings of the 2017 conference on empirical methods in natural language processing, pages 221–232, copenhagen, denmark, 2017. association for computational linguistics. laia mayol. asymmetries between interpretation and production in catalan pronouns. dialogue and discourse, 9:1–34, 2018. marc moens and mark steedman. temporal ontology and temporal reference. computational linguistics, 14(2):15–28, 1988. christoph müller. resolving it, this, and that in unrestricted multi-party dialog. in proceedings of the 45th annual meeting of the association of computational linguistics, pages 816–823, prague, 2007. association for computational linguistics. url https://www.aclweb. org/anthology/p07-1103. costanza navarretta. a contrastive analysis of abstract anaphora in danish, english and italian. in antónio branco, tony mcenery, ruslan mitkov, and fátima silva, editors, proceedings of daarc 2007, pages 103–109. centro de linguística da universidade do porto, 2007. martha palmer, daniel gildea, and paul kingsbury. the proposition bank: an annotated corpus of semantic roles. computational linguistics, 31(1):71–106, 2005. doi: 10.1162/ 0891201053630264. url https://www.aclweb.org/anthology/j05-1004. rebecca j. passonneau. getting at discourse referents. in proceedings of the 27th annual meeting of the association for computational linguistics, pages 51–59, vancouver, 1989. association for computational linguistics. massimo poesio, roland stuckardt, and yannick versley. challenges and directions of further research. in massimo poesio, roland stuckardt, and yannick versley, editors, anaphora resolution: algorithms, resources and application, pages 487–500. springer-verlag, berlin– heidelberg, 2015. 220 event vs entity interpretation sameer pradhan, alessandro moschitti, nianwen xue, hwee tou ng, anders björkelund, olga uryupina, yuchen zhang, and zhi zhong. towards robust linguistic analysis using ontonotes. in proceedings of the seventeenth conference on computational natural language learning, pages 143–152, sofia, august 2013. association for computational linguistics. url http: //www.aclweb.org/anthology/w13-3516. sameer s. pradhan and nianwen xue. ontonotes: the 90% solution. in proceedings of human language technologies: the 2009 annual conference of the north american chapter of the association for computational linguistics, companion volume: tutorial abstracts, pages 11– 12, boulder, colorado, may 2009. association for computational linguistics. url https: //www.aclweb.org/anthology/n09-4006. r development core team. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, 2008. url http://www.r-project.org. real academia española (rae). nueva gramática de la lengua española. manual. real academia española y asociación de academias de la lengua española, madrid, 2010. hannah rohde. pronouns. in c. cummins and n. katsos, editors, oxford handbook of experimental semantics and pragmatics. oxford university press, oxford, 2019. lorenza russo, sharid loáiciga, and asheesh gulati. italian and spanish null subjects. a case study evaluation in an mt perspective. in proceedings of the eighth international conference on language resources and evaluation (lrec’12), pages 1779–1784, istanbul, turkey, 2012. european language resources association (elra). url http://www.lrec-conf.org/ proceedings/lrec2012/pdf/813_paper.pdf. tanja samardžić. dynamics, causation, duration in the predicate-argument structure of verbs: a computational approach based on parallel corpora. phd thesis, university of geneva, geneva, switzerland, 2014. luca serianni. italiano. grammatica, sintassi, dubbi. garzanti, turin, 1997. rosemary j. stevenson, rosalind a. crawley, and david kleinman. thematic roles, focus and the representation of events. language and cognitive processes, 9(4):519–548, 1994. agnès tutin. a corpus-based study of pronominal anaphoric expressions in french. in proceedings of daarc 2002, lisbon, 2002. olga uryupina, ron artstein, antonella bristot, federica cavicchio, francesca delogu, kepa rodriguez, and massimo poesio. annotating a broad range of anaphoric phenomena, in multiple genres: the arrau corpus. natural language engineering, 26:95–128, 2020. renata vieira, susanne salmon-alt, caroline gasperin, emmanuel schang, and gabriel othero. coreference and anaphoric relations of demonstrative noun phrases in multilingual corpus. anaphora processing: linguistic, cognitive and computational modeling, pages 385–403, 2005. klaus von heusinger and petra b. schumacher. discourse prominence: definition and application. journal of pragmatics, 154:117–127, 2019. 221 bevacqua, loáiciga, rohde and hardmeier bonnie webber. structure and ostension in the interpretation of discourse deixis. language and cognitive processes, 6(2):107–135, 1991. eva wittenberg, shota momma, and elsi kaiser. demonstratives as bundlers of conceptual structure. glossa: a journal of general linguistics, 6(1):33, 2021. victoria yaneva, le an ha, richard evans, and ruslan mitkov. classifying referential and nonreferential it using gaze. in proceedings of the 2018 conference on empirical methods in natural language processing, pages 4896–4901, brussels, 2018. association for computational linguistics. 222 event vs entity interpretation appendix a. model results with the liberal annotations in this appendix, we report the results of the models run with the liberal annotation (table 7). the models were chosen following the same procedure described in §3.5. other than the fixed effects as shown on the table, the models included the following random effects: a random intercept and slope for condition by participant and for aspect by item in english; random intercepts by participant and intercept and slopes for aspect by item in french; random intercepts by participant and item in german; random intercepts by item only in italian; and random intercept and slopes for conditions by participant and random intercepts by item in spanish. 223 bevacqua, loáiciga, rohde and hardmeier effect estimate std.error z value pr(> |z|) english (intercept) 0.41 0.22 1.88 0.06 condition (this) 3.96 0.37 10.84 < 0.001 ∗∗∗ aspect (perf) −0.27 0.39 −0.69 0.49 alternation (alt) 0.80 0.39 2.06 0.04 ∗ aspect:alternation −1.89 0.80 −2.35 0.02 ∗ french (intercept) 0.45 0.35 1.27 0.21 c’est − cela −0.12 0.27 −0.42 0.67 il − c’est −4.80 0.43 −11.07 < 0.001 ∗∗∗ aspect (perf) −0.21 0.68 −0.30 0.76 alternation (alt) 1.45 0.66 2.21 0.03 ∗ german (intercept) 1.14 0.39 2.91 0.004 ∗∗ dies − das −0.77 0.48 −1.61 0.11 es − dies −5.53 0.57 −9.67 < 0.001 ∗∗∗ aspect (perf) 0.37 0.64 0.58 0.56 alternation (alt) 0.85 0.64 1.34 0.18 aspect:alternation 0.07 1.27 0.05 0.96 italian (intercept) 1.67 0.30 5.63 < 0.001 ∗∗∗ ciò − null 4.18 0.45 9.34 < 0.001 ∗∗∗ questo − ciò −0.62 0.35 −1.75 0.08 aspect (alt) 1.14 0.56 2.04 0.04 ∗ alternation (perf) 0.13 0.55 0.24 0.81 spanish (intercept) 0.47 0.35 1.34 0.18 esto − este 5.70 0.76 7.53 < 0.001 ∗∗∗ null − esto −4.50 0.65 −6.94 < 0.001 ∗∗∗ aspect (perf) 0.06 0.48 0.12 0.90 alternation (alt) 1.65 0.52 3.20 0.001 ∗∗ (esto − este):altern 0.47 0.77 0.61 0.54 (null − esto):altern 2.10 0.96 2.20 0.03 ∗ ∗∗∗p < 0.001, ∗∗p < 0.01, ∗p < 0.05 table 7: estimated models fixed effects (liberal analysis). 224 event vs entity interpretation appendix b. verbs from parcorfull and their spontaneity scores here we present the list of verbs used with an entity antecedent in parcorfull. the counts column comes from our extraction, while the causative-rate, anticausative-rate and spontaneity scores are taken from samardžić (2014). verb counts causative-rate anticausative-rate spontaneity-score balance 1 0.18 0.05 1.34 grow 1 0.14 0.78 -1.68 open 1 0.54 0.14 1.33 run 1 0.3 0.56 -0.64 225 bevacqua, loáiciga, rohde and hardmeier appendix c. verbs from parcorfull and their spontaneity scores here we present the list of verbs used as an event antecedent in parcorfull. the counts column comes from our extraction, while the causative-rate, anticausative-rate and spontaneity scores are taken from samardžić (2014). verb counts causative-rate anticausative-rate spontaneity-score break 1 0.33 0.3 0.09 burn 1 0.23 0.18 0.22 close 1 0.2 0.14 0.39 collect 1 0.27 0.06 1.55 drive 1 0.36 0.11 1.2 fly 1 0.22 0.78 -1.27 move 1 0.11 0.8 -1.97 run 3 0.3 0.56 -0.64 stand 1 0.15 0.85 -1.76 226 hinterwimmer_brocher_october 20_2020 dialogue & discourse 11(2) 110-127 doi: 10.5087/dad.2020.204 ©2020 stefan hinterwimmer, andreas brocher, and umesh patil this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). demonstrative pronouns as anti-logophoric pronouns: an experimental investigation stefan hinterwimmer germanistik linguistik hinterwimmer@uni-wuppertal.de fb a geistesund kulturwissenschaften bergische universität wuppertal gaußstr. 20 d-42119 wuppertal andreas brocher a.brocher@gmx.de umesh patil umesh.patil@gmail.com institut für deutsche sprache und literatur i universität zu köln albertus-magnus-platz d-50923 köln editor: jonathan ginzburg submitted 05/2019; accepted 10/2020; published online 10/2020 abstract in this paper we report the results of two experimental studies in which we tested the claim of hinterwimmer and bosch (2017) that german demonstrative pronouns are anti-logophoric pronouns: they avoid discourse referents as antecedents that function as perspectival centers. in both experiments we tested the interpretative options of demonstrative pronouns in text segments which were either perspectivally neutral or in which the narrator’s or a topical protagonist’s perspective was foregrounded. taken together, the experimental results are most compatible with a slightly modified version of the analysis argued for in hinterwimmer and bosch (2017) according to which topical discourse referents in neutral narration automatically become perspectival centers. 1 introduction in this paper we report the results of two experimental studies in which we tested the claim of hinterwimmer and bosch (2017) that german demonstrative pronouns of the der (he)/die (she)/das (it) series (henceforth: dpros) avoid discourse referents as antecedents that function as perspectival centers. the term perspectival center is defined as follows: a discourse referent a is the perspectival center with respect to a proposition p if p is the content of a mental state of the semantic value of a (i.e. g(a), where g is the assignment function). the clearest instances of perspectival centers are the subjects of propositional attitude verbs such as think or believe: the proposition denoted by the complement clause of such a verb is the content of a mental state of the subject. consequently, the subject is the perspectival center with respect to that proposition. adopting a possible worlds semantics along the lines of hintikka (1969), the denotation of a sentence such as (1a) can thus be paraphrased as in (1b). hinterwimmer, brocher and patil 111 (1) a. stanley believes that trump will win the election in 2020. b. in all worlds that are compatible with stanley’s beliefs at the utterance time, trump wins the election in 2020. a second case where discourse referents clearly function as perspectival centers is free indirect discourse (henceforth: fid). fid is a form of speech or thought representation that is often found in fictional narrative texts. like in indirect discourse (an instance of which is given in (1a)), in fid a sentence is interpreted as the content of a thought or utterance of some discourse referent. in contrast to indirect (and direct) discourse, however, such an interpretation is not enforced by the presence of a propositional attitude verb under which the respective sentence is embedded. rather, sentences in fid are autonomous, unembedded sentences such as (2a) whose interpretation as a thought or utterance has to be inferred by the reader. this inference is either triggered by the content exclusively or by the content in combination with clues such as the presence of deictic expressions that can only be interpreted sensibly with respect to the context of some prominent discourse referent. in the case of (2a), for instance, the deictic temporal adverb tomorrow can only be interpreted with respect to mary’s context, as referring to the day following the day on which she looked at her tablet with sheer panic. if tomorrow was interpreted with respect to the narrator’s context, the resulting interpretation would be contradictory since an event or state cannot at the same time be located in the past and in the future. (2) mary looked at her tablet in sheer panic. a. tomorrow she had to submit her paper, and she had not even written two pages. if (2a) is interpreted as a thought that mary has while looking at her tablet in sheer panic, however, and if tomorrow is interpreted with respect to mary’s context, while the past tense is interpreted with respect to the narrator’s context, there is no such contradiction (see doron, 1991; schlenker, 2004; sharvit, 2008; eckardt, 2014 and maier, 2015 for different analyses of fid in a formal semantics framework and rauh, 1978 and banfield, 1982 for early analyses of fid in a generative framework). mary is thus the perspectival center with respect to (2a). since the seminal work of clements (1975) on the pronoun system of the west african language ewe (see pearson, 2015 for a recent analysis), it is well known that many languages of the world have special pronouns that can only be used to pick up antecedents that are perspectival centers (see sells, 1987 for an overview): the clause containing a pronoun of this kind has to be interpreted as the content of a mental state of the antecedent. such pronouns are called logophoric pronouns. relatedly, many languages such as icelandic, tamil and japanese allow long-distance uses of reflexive pronouns in logophoric environments, i.e. in cases where the antecedent is a perspectival center (see sundaresan, 2012; nishigauchi, 2014 and charnavel, 2019 for recent discussion). at the same time, dubinsky and hamilton (1998) and patel-grosz (2014) have argued that epithets such as the idiot are anti-logophoric pronouns since they cannot be interpreted as picking up discourse referents functioning as perspectival centers – more precisely, when they occur in the complement clause of a propositional attitude verb or a sentence in fid, they have to be interpreted as referring to an individual that is distinct from the subject of the propositional attitude verb in the former case and the implicit thinker or speaker in the latter. relatedly, charnavel and mateu (2015) and yashima (2015) have argued for the existence of anti-logophoric pronouns in french, spanish and japanese. hinterwimmer and bosch (2017) observe that german dpros, which in previous literature have been assumed to avoid subjects (bosch et al. 2007), topics (bosch and umbach 2006; hinterwimmer 2015) or (proto-)agents as antecedents, can sometimes pick up subjects, topics and (proto-)agents. this is possible whenever the speaker or narrator is clearly present as perspectival center, i.e. whenever the sentence containing the dpro is interpreted as the content of a thought expressing the narrator’s stance. if the discourse referent functioning as subject, topic and/or proto-agent is the perspectival center, in contrast, such an interpretation is not available. hinterwimmer and bosch (2017) take this contrast to show that dpros are antilogophoric pronouns avoiding perspectival centers as antecedents. more tentatively, they demonstrative pronouns as anti-logophoric pronouns 112 propose that the entire distribution of dpros can be derived from anti-logophoricity: first, they assume that (what seems to be) avoidance of subjects and (proto-)agents are actually epiphenomena of topic avoidance. second, they assume that in neutral narration, where no speaker or narrator is present as perspectival center, topics are perspectival centers by default and are therefore avoided by dpros as antecedents. although this aspect is not worked out in any detail, the basic idea is that sentences in neutral narration are interpreted as the contents of perceptions of the topical referent (see brinton, 1980; palmer, 2004; farner, 2014 and van krieken, 2018 for relevant discussion). we tested the predictions of hinterwimmer and bosch’s (2017) account with two experimental studies in which we compared the interpretative options of dpros in text segments which were either perspectivally neutral or in which the narrator’s or a topical protagonist’s perspective was foregrounded. taken together, the experimental results are most compatible with a slightly modified version of the analysis argued for in hinterwimmer and bosch (2017) according to which topical referents automatically become perspectival centers in neutral narration. the paper is structured as follows, in section 2 we first give a brief overview over previous work on dpros in german, before we present the analysis of hinterwimmer and bosch (2017). in section 3 the two experimental studies are presented and discussed. section 4 is the conclusion. 2 background 2.1 previous work on dpros in german like many languages, such as finnish, dutch, and catalan (see, e.g. kaiser and trueswell, 2008; kaiser, 2010, 2011a, 2011b, 2013; mayol and clark, 2010), german has two pronoun series: the first one consists of the personal pronouns (henceforth: ppros) er/sie/es (‘he’/‘she’/‘it’) together with their various forms. the second series consists of the so-called demonstrative pronouns (dpros) der/die/das, which are morphologically identical to the definite article. there is a further series of demonstrative pronouns in german, which is, however, largely confined to somewhat formal registers (see patil, bosch and hinterwimmer, 2020), and will not be discussed in this paper: the dieser/diese/dieses-series. generally, in languages with two or more pronoun series, weight or length, and therefore the associated markedness, varies across them (see the references above). this clearly applies to german as well, with er being shorter than der (see patel-grosz and grosz, 2017). as pointed out by kaiser and trueswell (2008), kaiser (2010; 2011a; 2011b; 2013) and mayol and clark (2010), marked pronouns display a more limited distribution than unmarked pronouns. there is also a general consensus in the literature that german dpros avoid maximally prominent antecedents. assumptions differ, however, regarding the question of which properties are decisive for maximal prominence. on the basis of contrasts like the one between the dproand the ppro-variant of the second sentence in (3), bosch et al. (2007) argue for the following claim: dps functioning as the subject of immediately preceding sentences are maximally prominent and therefore avoided by dpros. bosch and umbach (2006), however, observe that in some cases dpros actually have a strong preference for the subject of the preceding sentence: the dpro in the third sentence in (4) can only be interpreted as referring to the subject of the preceding sentence, peter, not the indirect object, karl. (3) marki hat gestern mit noahj gesprochen. derj/*i/eri/j hat sich einen neuen ferrari gekauft. marki talked to noahj yesterday. he(dpro)j/*i /hei,j bought himself a new ferrari. (4) woher karli das weiß? peterj hat es ihmi gesagt. derj/*i /eri,j war gerade hier. how does karli know? peterj told it to himi. he(dpro)j/*i /hei,j has just been here. hinterwimmer, brocher and patil 113 according to bosch and umbach (2006), this is due to karl being the topic of the entire discourse segment: karl is introduced in the opening sentence and picked up by a ppro in the second sentence. additionally, the first sentence asks and the second sentence answers a question about karl. the subject of the second sentence, in contrast, is contained in the focal part of the sentence. that part provides new information that answers the question asked by the first sentence. based on these and similar observations, bosch and umbach (2006) propose that dpros avoid discourse topics because they are the most prominent antecedents. what seems to be subject avoidance in cases like (3) is, in their view, an epiphenomenon of topic avoidance, since there is a (crosslinguistically quite stable) tendency for discourse topics to be realized as subjects. when the two notions diverge, however, as in (4), dpros avoid topics, not subjects. in hinterwimmer (2015), this proposal is taken up and further generalized along the following lines (see hinterwimmer and brocher, 2018 for experimental evidence): in cases of co-reference such as in (3) and (4), where the antecedent and the pronoun appear in separate sentences and where there is no structural relation between antecedent and pronoun, (maximal) prominence is defined in terms of topicality. in cases like (5a-d), in contrast, where binding is at play and where there is a structural relation between antecedent and pronoun, namely ccommand, maximal prominence is defined in terms of subjecthood. consequently, not only the ppro in (5a), but also the dpro in (5b) receives a bound interpretation, as the potential binder is an indirect object. in (5d), in contrast, where the only potential binder is the subject, a bound interpretation is unavailable for the dpro, unlike for the ppro in (5c). note that since seine is the possessive version of the ppro er and dessen the possessive version of the dpro der, there is no reason to expect that they behave differently from their respective nominative versions with respect to the constraint under discussion. (5) a-b. frau bauer bringt [jedem buchhalter]i seinei/desseni neue daten, die schon lange fällig waren. mrs. bauer brings [every accountant]i hisi/his(dpro)i new data, which have been overdue for a while. c-d. [jeder buchhalter]i bringt frau bauer seinei/dessen*i neue(n) daten, die schon lange fällig waren. [every accountant]i brings mrs. bauer hisi/his(dpro)*i new data, which have been overdue for a while. schumacher et al. (2016) and schumacher et al. (2017) have added agentivity to the various features that can affect prominence. based on experimental evidence, they claim that in sentences with dative experiencer verbs, such as the opening sentence in (6), the referent of the dative object is more prominent than the referent of the subject because it has an additional agentivity feature, namely sentience. consequently, a dpro contained in the subsequent sentence can only be interpreted as the subject dp. since the argument with the highest number of agentivity features is typically realized as subject (see dowty, 1991; primus, 1999; 2006), subject-avoidance might be a side effect of more general differences in agentivity. (6) dem gärtneri gefällt der kapitänj, der ein eis isst. the-dat gardener is-pleasing-to the-nom skipper who an-acc ice cream eats. the skipperi who eats ice cream is pleasing to the gardenerj. aber der*i/j/eri,j redet gerade mit zwei damen. but dpro-nom/ he-nom talks now with two ladies. but he(dpro)*i,j/hei,j is talking to two ladies right now. now, in contrast to cases where the dative object is fronted, there is no clear preference for either of the two potential antecedents in cases like in (7). here it is the subject that is fronted. from this, schumacher et al. (2016) and schumacher et al. (2017) conclude that prominence in terms of topicality plays a role as well. the idea is that fronted dps are likely to be interpreted as topics: in cases like in (6), prominence in terms of topicality then coincides with prominence in terms of agentivity, resulting in a clear interpretation preference for the less prominent demonstrative pronouns as anti-logophoric pronouns 114 referent. in (7), in contrast, the two notions go in opposite directions. consequently, there is no clear preference anymore. (7) der kapitäni gefällt dem gärtnerj, der ein eis isst. the-nom skipper is-pleasing-to the-dat gardener who an-acc ice cream eats. the skipperi is pleasing to the gardenerj, who eats ice cream. aber deri/j/eri,j redet gerade mit zwei damen. but dpro-nom/ he-nom talks now with two ladies. but he(dpro)i,j/hei,j is talking to two ladies right now. 2.2 dpros as anti-logophoric pronouns as already said in section 1, hinterwimmer and bosch (2017) claim that the distribution of dpros can largely be derived from anti-logophoricity rather than from subject-, topic-, or agentavoidance: dpros may not pick up discourse referents functioning as perspectival centers with respect to the proposition denoted by the clause that contains the respective dpro. recall that a discourse referent a is the perspectival center with respect to a proposition p if p is the content of a mental state of the semantic value of a. the crucial observation motivating the account of hinterwimmer and bosch (2017) is that in some cases dpros can pick up discourse referents that are the subjects (pace bosch et al., 2007), topics (pace bosch and umbach, 2006; hinterwimmer, 2015) and agents (pace schumacher et al., 2016; 2017) of the preceding sentence, while in other cases this is completely impossible. compare the continuations of (8) in (8a) and (8b), respectively: (the proper name referring to) peter is the subject of the main clause in (8) and the agent of the eventuality introduced by the main clause verb sighed. additionally, since there are no contextual factors overwriting the tendency for topics to be realized as subjects, he is the topic by default as well. nevertheless, the dpro in (8a) can easily be understood as picking up peter, while this is completely impossible for the dpro in (8b). since the preceding sentence is the same in both cases, this contrast cannot be due to any difference regarding subjecthood, (proto-) agentivity or topicality, and therefore is completely unexpected on all accounts discussed in section 2.1: peter should be unavailable as an antecedent for the dpro not only in (8b), but in (8a) as well. (8) peteri seufzte, als er die tür öffnete und sah, dass die wohnung mal wieder in einem fürchterlichen zustand war. peteri sighed when he opened the door and saw that the flat was in a terrible state again. a. deri kann sich einfach nicht gegen seinen mitbewohner durchsetzen. he(dproi) is simply unable to stand his ground against his flatmate. b. verdammt, der*i/eri hatte doch gestern erst aufgeräumt. damn, he(dpro*i)/hei had only tidied up yesterday, after all. the contrast between (8a) and (8b) can be accounted for if one assumes dpros to be antilogophoric pronouns, however: the continuation in (8a) is most plausibly understood as expressing a general comment about peter by the narrator. this is indicated by the content in combination with the switch from past to present tense, which breaks narrative continuity. the narrator is thus the perspectival center with respect to the proposition denoted by (8a), since that proposition is the content of a mental state (namely a thought) of the narrator on the most plausible interpretation. anti-logophoricity is therefore not violated if the dpro in (8a) is interpreted as picking up peter, since peter is not the perspectival center with respect to the proposition denoted by (8a). the continuation in (8b), in contrast, is most plausibly understood as expressing a thought of peter in fid. this is indicated by the content in combination with the presence of the deictic temporal adverb gestern (‘yesterday’), the modal particle doch and the evaluative expression verdammt (‘damn’). the deictic temporal adverb gestern is more plausibly interpreted with respect to peter’s than with respect to the narrator’s context, i.e. as referring to the day preceding hinterwimmer, brocher and patil 115 the day on which peter came home in the evening. likewise, the modal particle doch, which indicates that the proposition denoted by the clause containing the modal particle violates a previously held assumption, is more plausibly interpreted as violating peter’s expectations than the narrator’s. finally, verdammt most likely expresses peter’s frustration, not the narrator’s. as already said in section 1, fid is a special form of speech or thought representation in which all context-sensitive expressions with the exception of pronouns and tenses are interpreted with respect to some prominent protagonist’s context. the author of that context is the respective protagonist, the temporal parameter is provided by the reference time of the ongoing story and the spatial parameter is the location of the protagonist at the reference time. concerning pronouns and tenses, they are interpreted with respect to the narrator’s context (see schlenker, 2004; sharvit, 2008 and eckardt, 2014 for different versions of the double context analysis and maier, 2015 for an analysis according to which fid is a special form of mixed quotation in which pronouns and tenses are unquoted). on its most plausible interpretation as fid, peter is therefore the perspectival center with respect to the proposition denoted by (8b) since that proposition is the content of a mental state (namely a thought) of peter. consequently, anti-logophoricity is violated if the dpro in (8b) is interpreted as picking up peter – at least if (8b) is interpreted as fid. if (8b) is interpreted as expressing a thought of the narrator (a rather implausible, but not completely excluded option), in contrast, the dpro can of course be interpreted as picking up peter. from contrasts like the one between (8a) and (8b), which cannot be interpreted in terms of subject-, topic-, or (proto-)agent-avoidance, but in terms of anti-logophoricity, hinterwimmer and bosch (2017) conclude that dpros are anti-logophoric pronouns. additionally, they note that the contrast between the interpretative options of dpros and ppros in sentences like (9ab), which has already been observed by wiltschko (1998), cannot only be accounted for in terms of subject-avoidance, but also in terms of anti-logophoricity: dpros cannot be interpreted as co-referential with or bound by the subjects of propositional attitude verbs since the individuals referred to or quantified over by the respective subject dp are perspectival centers with respect to the proposition denoted by the complement clause of the propositional attitude verb1. 1 concerning the complement clauses of propositional attitude verbs, the situation is actually more complicated. hinterwimmer & bosch (2017) observe that in sentences with double embedding such as (i-ii) dpros can be interpreted as co-referential with or bound by the subject of the embedded propositional attitude verb, but not by the subject of the matrix propositional attitude verb. i. mariai behauptet, dass peterj glaubt, dessenj tochter sei klüger als ihrei. mariai claims that peterj believes that his {dproj} daughter is smarter than hersi. ii. mariai behauptet, dass peterj glaubt, deren*i/ihrei tochter sei klüger als seinej. mariai claims that peterj believes that her {dpro*i/pproi} daughter is smarter than hisj. hinterwimmer & bosch (2017) draw the following conclusion from this observation: dpros avoid the most prominent perspectival center. in the cases discussed so far, there is just one perspectival center, which is thus automatically the most prominent one. in cases like (i) and (ii), in contrast, the subject of the matrix propositional verb is the more prominent perspectival center than the subject of the embedded one, for the following reason: the referent of the matrix clause subject is the perspectival center with respect to the proposition denoted by the complement clause of the matrix propositional attitude verb. that proposition in turn contains the referent of the embedded subject as well as the proposition with respect to which that referent is the perspectival center. the referent of the matrix clause subject is therefore the superordinate perspectival center, and the referent of the embedded subject is the subordinate perspectival center, and hinterwimmer & bosch (2017) assume that superordination corresponds to higher prominence. we have tested the claims of hinterwimmer & bosch (2017) concerning the interpretative options of dpros in sentences like (i) and (ii) in a reading time and an acceptability rating study. the predictions were confirmed: sentences like (i), in which the dpro in the most deeply embedded clause could only be interpreted as co-referential with the embedded, but not the matrix clause subject were read faster and rated better than sentences where it was the other way round. we omitted discussion of these experiments from the article, however, since their results do not directly provide arguments for anti-logophoricity, as pointed out to us by two anonymous reviewers: they could just as well be explained by topic avoidance, since the matrix clause subjects can plausibly be regarded as topics by default. demonstrative pronouns as anti-logophoric pronouns 116 (9) a. mariai glaubt, dass die*i/siei ein genie ist. maria believes that she {dpro*i /pproi} is a genius. b. [jeder mann]i glaubt, dass der*i/eri ein genie ist. [every man]i believes that he {dpro*i /pproi} is a genius. concerning the question of why in sentences like in (3) and (4), the dpro can only refer to the non-topical referent and why it cannot be interpreted as bound by the subject quantifier in cases like (5d), hinterwimmer and bosch (2017) argue as follows: when there is no indication of the presence of a perspectival center (i.e. in cases of neutral narration), the respective topic or the individuals quantified over by the subject quantifier are interpreted as perspectival centers by default. the authors do not discuss cases like (6) and (7), but such cases can be accounted for under the assumption that dpros avoid (the most prominent) perspectival centers: because of sentience, it is quite natural to interpret the experiencer argument of a verb as perspectival center with respect to the proposition denoted by the clause containing that verb. interpreting the stimulus argument as perspectival center is rather unnatural, in contrast, at least when it occurs in (canonical) clause-internal position. if the stimulus argument is fronted, however, this might be taken as indication that it is the aboutness topic. consequently, in the absence of a speaker or narrator functioning as perspectival center, it becomes a potential perspectival center as well. this results in unclear interpretation preferences. the argumentation in hinterwimmer and bosch (2017) is largely based on the two authors’ introspectively gained native speaker intuitions, and the contrasts that are reported are rather subtle. although informally gained introspective judgments are valuable for formulating linguistic hypotheses, they are arguably insufficient for theory building as such (see e.g. gibson and fedorenko, 2013). to address this limitation, we conducted two offline rating experiments to gain more solid empirical evidence for the claims made in hinterwimmer and bosch (2017). we tested the following hypothesis: dpros can pick discourse referents that function as subject, agent, and topic of the preceding sentences as long as there exists a perspectival center that is different from these discourse referents. in other words, dpros avoid perspectival centers. if the topical discourse referents are at the same time perspectival centers with respect to the proposition denoted by the sentence containing the dpro, in contrast, they cannot be picked up by that dpro. in order to test this hypothesis, we compared the following two types of sentences: sentences in fid, where the discourse referent functioning as topic, agent and subject of the preceding sentence is at the same time the perspectival center, and sentences where clearly the narrator is the perspectival center. we also tested whether there is only a weak tendency to interpret topics as perspectival centers, or whether topics are automatically interpreted as perspectival centers whenever there is no indication of the narrator functioning as perspectival center. to that end, we compared sentences in fid and sentences where the narrator is the perspectival center with neutrally narrated sentences. if there is only a weak tendency to interpret topics as perspectival centers, it should be easier for dpros contained in neutrally narrated sentences to pick up topical discourse referents than for dpros contained in fid-sentences: in the latter case, antilogophoricity is necessarily violated, while in the former case violating it can be avoided by interpreting the respective sentence as reporting the abstract narrator’s rather than the (semantic value of the) topical discourse referent’s perspective. if topical discourse referents automatically become perspectival centers whenever there is no indication of a speaker or narrator functioning as perspectival center, in contrast, the interpretative options of dpros should be the same in neutrally narrated sentences and in sentences in fid. 3 the experimental studies 3.1 overview in the two experiments to be discussed in this section, we used offline rating tasks to test whether a topical discourse referent can be referred to with a dpro in sentences that express evaluation of the narrator as opposed to sentences that express a thought of the topical discourse referent in fid (cf. hinterwimmer and bosch, 2017). in the two experiments, we also tested hinterwimmer, brocher and patil 117 whether there is only a weak tendency to interpret topical discourse referents as perspectival centers or whether they are automatically interpreted as perspectival centers whenever there is no evidence of a speaker or narrator functioning as perspectival center. (10) als fabian zur arbeit ging, fand er 100 euro auf dem gehweg. when fabian went to work, he found 100 euros on the sidewalk. [narrator-judgement] a-b. er/der hat einfach immer so ein unverschämtes glück. he/hedpro simply always is so incredibly lucky. [fid] c-d. toll, er/der würde heute abend davon schick essen gehen. great, he/hedpro would spend that for a posh dinner tonight. [neutral] e-f. er/der kaufte sich von dem geld ein paar neue schuhe. he/hedpro bought a pair of new shoes with the money. in (10), the individual introduced in the opening sentence (‘fabian’) is established as the discourse topic of the text segment, and the dps referring to that individual in the adjunct and the matrix clause of the first sentence are both the subject and the agent with respect to those clauses. therefore, on the assumption that dpros avoid topics, subjects, and/or agents, we predict dpros in all three variations to be judged equally bad (and worse than the corresponding ppros). on the other hand, if avoidance of the perspectival center is the relevant constraint, the dpro in narrator-judgement conditions should be rated as more acceptable than the dpro in the other two conditions. this is expected because in narrator-judgement conditions, the discourse topic is not the perspectival center anymore, but the abstract narrator is. between the other two conditions, the dpro in the fid condition should be rated as less acceptable than the dpro in the neutral condition if there is only a weak tendency to interpret topics as perspectival centers since it should then also be possible to interpret sentences in that condition as reported from the abstract narrator’s perspective, i.e. with the narrator functioning as perspectival center. if topics automatically become perspectival centers whenever there is no indication of the speaker or narrator functioning as perspectival center, in contrast, there should be no contrast in ratings of dpros between the two conditions. finally, for the ppro, we expected no variation in ratings across the three conditions because it is an unmarked pronoun. 3.2 experiment 1 method participants eighty-five native speakers of german were recruited through prolific (https://prolific.ac/) for monetary compensation (£2.04). materials we constructed 36 experimental items each consisting of two sentences, as in (10), interspersed with 36 fillers. the first sentence, which was the same across all conditions, established an individual referred to by a proper name as topic. the second sentence had three possible continuations, each of which occurred with a ppro and a dpro. this resulted in a total of six conditions. in narrator-judgement conditions (conditions a and b), the second sentence clearly expressed an evaluation of the topical referent by the narrator, as indicated by the content in combination with a switch from past tense to present tense (as in (8a) from section 2.1 above). in condition a, the topical referent was referred to by a ppro, and in condition b it was referred to by a dpro. in fid conditions (conditions c and d, where in condition c, the topical referent is referred to by a ppro, and in condition d by a dpro), the second sentence is most plausibly interpreted demonstrative pronouns as anti-logophoric pronouns 118 as a thought of the topical referent rendered as fid. we always used two indicators for fid: an interjection such as toll (‘great’) and a deictic element such as heute (‘today’) that in combination with past tense are typically interpreted with respect to the topical referent’s (fictional) context, not with respect to the narrator’s context. note, however, that we cannot completely exclude the option of interpreting the final sentence as expressing the narrator’s evaluation, i.e. as claiming in the case of (10c-d), for example, that the narrator finds it great that the topic of the discourse (fabian) will, according to her/his expectations, spend the money for a posh dinner on the evening of the day on which the narrator tells her/his story. the participants’ task is thus actually twofold in the two conditions: on the one hand, they have to recognize the final sentence in both conditions as fid, and on the other hand they have to resolve the respective pronoun. the two tasks are not completely independent from one another since the dpro is only excluded from picking up the topical referent when the final sentence is interpreted as fid, but not when it is interpreted as expressing the narrator’s perspective.2 this is a potential shortcoming of our experimental design that we will come back to in the final discussion. finally, the neutral conditions (conditions e and f), are neutral narrative continuations, where in e the topical referent is referred to by a ppro, and in f by a dpro. on the assumption that there is only a weak preference for interpreting topics as perspectival centers in the absence of any indication of the speaker or narrator functioning as perspectival center, the participants’ task is twofold on those two conditions, too: in addition to resolving the pronoun, they have to decide whether they assume the respective final sentence to express the narrator’s or the topical referent’s perspective. again, the two tasks are not independent from one another since the dpro is only prohibited from picking up the topical referent if the latter is assumed to be the perspectival center.3 on the assumption that topical referents automatically become perspectival centers in the absence of any indication of a speaker or narrator functioning as perspectival center, in contrast, the participants’ sole task consist in resolving the respective pronouns, where dpros are excluded from being resolved to the topical referent. procedure the experiment involved a ‘yes’/‘no’ judgment task. participants were instructed that the texts were beginnings of short stories produced by advanced german learners, where each example text was produced by a different learner. participants’ task was to judge whether the student had reached native-like proficiency in german (by responding ‘yes, they have’ or ‘no, they have not’). the reason why we asked participants to judge language proficiency instead of acceptability was that judging acceptability could be influenced by factors such as prescriptive knowledge of the grammar and metalinguistic reasoning such as the plausibility of the text (schütze, 2016: 81-88). due to this potential vagueness of the acceptability rating task we wanted to explore an alternative task that native speakers are used to performing in everyday life — judging fluency of a non-native speaker.4 for methodological comparison, we also carried out an experiment using the same items with conventional acceptability rating task (see expt. 2). data analysis all data processing and analyses were carried out in r (r core team, 2018). we fitted generalized linear mixed models with logit link function (jaeger, 2008), where the dependent variable was the binary response (native or non-native) and the fixed effects were: 1. sentence type (three levels: narrator-judgement, fid, and neutral), 2. pronoun type (two levels: ppro and dpro), and 3. their interaction. since we intended to compare the effect of narratorjudgement and fid conditions with the neutral conditions we fitted a model with treatment 2 we are grateful to an anonymous reviewer for pointing this out to us. 3 we are grateful to an anonymous reviewer for pointing this out to us. 4 given the fact that among 175–220 million german speakers worldwide, 85-125 million speakers speak german non-natively (geographical distribution of german speakers, n.d.), it is evident that german native speakers are commonly exposed to non-native german usage. hinterwimmer, brocher and patil 119 contrast with neutral condition as the reference level. we inserted random intercepts for participants and items. models with maximal random effects structure did not converge. figure 1. proportions of dpro/ppro trials rated as native in experiment 1. fixed effects estimate std.error p intercept -0.834 0.154 < 10 -7 * fid 0.338 0.145 0.0197 * narrator-judgement 0.647 0.144 < 10 -5 * pronoun type 2.872 0.173 < 10 -15 * fid : pronoun type -0.015 0.241 0.949 narrator-judgement : pronoun type -0.479 0.237 0.0427 * table 1: linear model estimates, standard errors and p-values for the data from experiment 1. neutral is the reference level. estimates for fid and narrator-judgement show the effect of these two sentence types with respect to sentences in neutral type; the estimate for pronoun type shows the effect of ppro with respect to dpro; and estimates for fid : pronoun type and narrator-judgement : pronoun type show the interaction between fid sentences and the type of the pronoun, and narrator-judgement sentences and the type of the pronoun. results the mean proportions of responses are plotted in figure 1 and the fixed effects from the linear models are provided in table 1. in the first model (see table 1), where neutral was the reference level, there were significant main effects of sentence type and pronoun type: narrator-judgement and fid sentences were judged as more native-like than neutral sentences, and ppros were judged as more native-like than dpros. the interaction of narratorjudgement and neutral sentence types with pronoun type was also significant. however, the second interaction of fid and neutral sentence types with pronoun type did not reach significance. pairwise comparisons for the effects of dpro between narrator-judgement and neutral, on the one hand, and narrator-judgement and fid on the other revealed that dpros were judged as significantly more native-like in the narrator-judgement condition in both cases. we did not carry out a pairwise comparison between the dpro in fid and neutral conditions because, although there was a numerical difference in the ratings for the dpro between these two conditions, their interaction did not turn out to be significant in the earlier model. discussion we defer the discussion of experiment 1 to the discussion of experiment 2, as both experiments asked the same question, using the same materials but different experimental paradigms. 88.2 46.2 89 38.7 85.1 31.6 narrator−judgement fid neutral dpro ppro nativity judgement proportions (%) demonstrative pronouns as anti-logophoric pronouns 120 3.3 experiment 2 we carried out experiment 2 to replicate the effects from experiment 1 using a slightly different methodology. moreover, we aimed at checking whether the dpro was in fact rated as more native-like in the fid condition than in the neutral condition. although this effect did not turn out to be statistically reliable in the previous experiment, there was a numerical trend and this trend was unexpected under both variants of the anti-logophoricity account under investigation (i.e. both on the assumption that there is just a weak tendency to interpret topical referents in the absence of any indication of a speaker or narrator functioning as perspectival center and on the assumption that topical referents automatically become perspectival centers in such cases). to that end, we employed a dual-task that combined a forced-choice with an acceptability rating task, such that we elicited two responses from each participant. because, among the conventional offline tasks used for eliciting linguistic judgements, the forced-choice task provides maximum power (schütze and sprouse, 2014), we reasoned that this task should reveal differences between the fid and the neutral condition for dpros, if there are any. also, the secondary task allowed us to validate the results from the nativity judgement task in experiment 1 with a more conventional acceptability task. method participants forty-six native speakers of german were recruited from the university of cologne for course credits. materials we used the same 36 experimental items as in experiment 1. these items were again interspersed with 36 fillers. procedure the experiment was an offline dual task – a forced-choice task followed by an acceptability rating task. in each trial, participants were shown a sentence, just like the first sentence in (10), followed by two continuations, like the second sentences in (10). one continuation contained a ppro, the other a dpro. participants were asked to choose the continuations that they found more acceptable and then rate the continuation they did not select for naturalness on a scale from 1 (for “completely acceptable”) through 7 (for “completely unacceptable”). in case participants found both continuations equally plausible, they were asked to only engage in the acceptability task. data analysis all data pre-processing and analyses were carried out in r (r core team, 2018). we fitted generalized linear mixed models with logit link function (jaeger, 2008) for the first response (forced-choice). the dependent variable was the binary judgment ppro or dpro and the fixed effect was sentence type, with the three levels narrator-judgement, fid, and neutral. since we wanted to compare the effect of narrator-judgement and fid conditions with the neutral conditions, as well as potential differences between the narrator-judgement and fid, we fit two separate models with treatment contrasts: in the first model, neutral was used as reference level and in the second, fid. we used by-item and by-participant random intercepts, but only a by-participant random slope (the model with maximal random effects structure (barr et al., 2013) did not converge). when a participant responded that both continuations were equally acceptable, we coded this response as “yes” for the dpro because our research question was whether the use of a dpro is licensed in this particular context. this contributed towards only 1% of the data. for the second response, we fitted linear mixed effects models (baayen et al., 2008) with acceptability rating as the dependent variable. the fixed effects were: 1. sentence type (three levels: narrator-judgement, fid and neutral), 2. pronoun type (two levels: ppro and dpro), and 3. their interaction. again, we fit two separate models with treatment contrasts: in the first hinterwimmer, brocher and patil 121 model, neutral was used as reference level and in the second, fid. we included by-participant and by-item random intercepts, together with random slopes only for sentence type (the model with maximal random effects structure did not converge). in the analysis, we dropped all trials for which both options in the forced-choice were selected as equally plausible, which contributed towards 1.1% of the data. figure 2. response proportions for the first task in experiment 2. figure 3. the mean acceptability rating for the continuation that was rated as not native in experiment 2. results response means for the forced-choice task are plotted in figure 2 and, for the acceptability rating task, in figure 3. the fixed effects from the linear models for the two types of responses are provided in tables 2–5. the analysis of the forced-choice task revealed that, compared to the neutral condition, dpros were chosen significantly more often both in the narrator-judgement and in the fid conditions. the analysis of the acceptability judgement task revealed a significant main effect of sentence type. narrator-judgement and fid sentences were rated as more acceptable than the neutral and narrator-judgement sentence were rated as more acceptable than fid sentences. analyses also revealed a significant main effect of pronoun type. the ppro conditions were rated as more acceptable than the dpro conditions. no interaction reached significance, which could be due to the somewhat few instances of acceptability ratings for the dpro conditions: only about 7% of all responses in the acceptability rating task were dpro sentences. 85.8 12.3 1.9 93.2 5.6 1.2 95.8 3.6 0.5 narrator−judgement fid neutral both dpro ppro first response: forced−choice proportions (%) 6 4.5 6.5 6.3 4 6.5 6.2 3.8 6.3 narrator−judgement fid neutral both dpro ppro second response: acceptability rating demonstrative pronouns as anti-logophoric pronouns 122 fixed effects estimate std.error p intercept -5.383 0.804 2.14e-11 * fid 1.304 0.774 0.092 . narrator-judgement 1.539 0.905 0.089 . table 2: linear model estimates, standard errors and p-values for the data from the forcedchoice task in experiment 2. neutral is the reference level. estimates for fid and narratorjudgement show the effect of these two sentence types with respect to sentences in neutral type. fixed effects estimate std.error p intercept -4.507 0.537 2e-16 * neutral -0.849 0.351 0.016 * narrator-judgement 1.312 0.267 9.04e-7 * table 3: linear model estimates, standard errors and p-values for the data from the forcedchoice task in experiment 2. fid is the reference level. estimates for neutral and narratorjudgement show the effect of these two sentence types with respect to sentences in fid type. fixed effects estimate std.error t-value intercept 3.789 0.181 20.946 * fid 0.290 0.105 2.766 * narrator-judgement 0.778 0.138 5.642 * pronoun type 1.502 0.352 4.263 * fid : pronoun type 0.103 0.382 0.269 narrator-judgement : pronoun type -0.263 0.391 -0.673 table 4: linear model estimates, standard errors and t-values for the data from the acceptability rating task in experiment 2. neutral is the reference level. estimates for fid and narratorjudgement show the effect of these two sentence types with respect to sentences in neutral type; the estimate for pronoun type shows the effect of ppro with respect to dpro; and estimates for fid : pronoun type and narrator-judgement : pronoun type show the interaction between fid sentences and the type of the pronoun, and narrator-judgement sentences and the type of the pronoun. fixed effects estimate std.error t-value intercept 4.079 0.168 24.249 * neutral -0.290 0.105 -2.766 * narrator-judgement 0.487 0.101 4.834 * pronoun type 1.605 0.239 6.722 * neutral : pronoun type 0.103 0.382 -0.269 narrator-judgement : pronoun type 0.366 0.282 -1.297 table 5: linear model estimates, standard errors and t-values for the data from the acceptability rating task in experiment 2. fid is the reference level. estimates for neutral and narratorjudgement show the effect of these two sentence types with respect to sentences in fid type; the estimate for pronoun type shows the effect of ppro with respect to dpro; and estimates for neutral : pronoun type and narrator-judgement : pronoun type show the interaction between neutral sentences and the type of the pronoun, and narrator-judgement sentences and the type of the pronoun. hinterwimmer, brocher and patil 123 discussion experiment 2 largely replicates the results from experiment 1. the results show that dpros in narrator-judgement conditions were judged as more acceptable than dpros in the fid and the neutral conditions. these data are most compatible with the hypothesis that dpros avoid perspectival centers, where topical referents automatically become perspectival centers in the absence of any indication of a speaker or narrator functioning as perspectival center. they are less compatible with the assumption that dpros generally avoid topics, subjects, and/or agents. the results are also less compatible with a variant of the ani-logophoricity hypothesis on which there is only a weak tendency for topical referents to be interpreted as perspectival centers in the absence of any indication of a speaker or narrator functioning as perspectival center. on such an account, dpros are expected to be more acceptable in neutral than in fid conditions, which is contrary to what we found. although the pattern of results of experiments 1 and 2 was similar, the numerical trend that we observed in experiment 1 between judgements for the dpro in fid and neutral conditions turned out to be significant in experiment 2. this part of the data was not predicted by either variant of the tested accounts. one possible explanation could be that fid is a form of speech that is usually encountered in longer narrative texts with clearly established protagonists. as already said in section 3.1 above, it is therefore possible that our fidsentences were perceived not as fid, but rather as the narrator’s evaluations by some participants. that is, in the example stimulus in (10e-f), for example, some participants might perhaps have understood the final sentence as claiming that the narrator finds it great that the topic of the discourse (fabian) will, according to her/his expectations, spend the money for a posh dinner on the evening of the day on which the narrator tells her/his story. consequently, they accommodated an ‘involved’ narrator functioning as the perspectival center and the dpro was then not precluded from picking up the discourse topic. an account along these lines raises the question, however, of why it should not also be possible to take the narrator instead of the topical referent as the perspectival center in the neutral condition, which would likewise be compatible with the anti-logophoricity of dpros. we tentatively suggest that this option is strongly dispreferred for the following reason: in neutral narration there is by definition no indication at all of the presence of a narrator functioning as perspectival center, while the topical discourse referent is highly prominent. it is therefore much more natural to assume the topical referent to be the perspectival center in such cases. sentences in the fid condition, in contrast, contain perspective-dependent expressions and therefore presuppose a perspectival center, which, for the reasons just mentioned, could not only be the topical referent, but also the narrator. given the additional complication just mentioned, studying the interaction of pronoun resolution and determination of the perspectival center would be a fruitful topic for future research. for an item like (10e-f), for example, this could be done by giving participants the additional task of answering a question like who finds it great that fabian will spend the money for a posh dinner?, with fabian, the speaker/narrator, both and i don’t know being the available answer options.5 one additional and also somewhat unexpected effect that we found in the two experiments was that the dpro in the narrator-judgement condition was judged to be less acceptable than the corresponding ppro condition. if evaluation from an abstract narrator made the narrator maximally prominent and the discourse topic (relatively) less prominent, the dpro should have been rated as acceptable as the corresponding ppro. we surmise that this contrast emerged because participants followed a prescriptive rule that they were taught at school: dpros are substandard and should therefore be avoided in written texts. similar effects have also been observed elsewhere. patil, bosch and hinterwimmer (2020) have reported that dpros are very rarely used in written language even in a context that licenses their use as per the prominence constraint. in fact, hinterwimmer and patil (2020) report two experiments carried out with the same set of stimuli but in two modalities – written and oral – which show that acceptability of dpros increases considerably with oral presentation (from 3.7 to 5.9 on the 1–7 likert scale). 5 we are grateful to an anonymous reviewer for pointing this out to us. demonstrative pronouns as anti-logophoric pronouns 124 4 general discussion and conclusion in this paper, we presented the results of two experimental studies in which we tested the claim in hinterwimmer and bosch (2017) that dpros can in fact pick up discourse referents functioning as topics, agents, and subjects, provided that there is a prominent perspectival center available that is distinct from the respective discourse referent. additionally, we set out to test whether there is only a weak tendency to interpret discourse topics as perspectival centers in the absence of any indication of a speaker or narrator functioning as perspectival center, or whether discourse topics are automatically interpreted as perspectival centers in such cases. in experiments 1 and 2, we tested via acceptability studies the claim that dpros could be interpreted as co-referential with topics if the respective sentences clearly express the perspective of a highly involved narrator, thus turning the narrator into a prominent perspectival center. we also tested whether topical referents are more strongly dispreferred by dpros in cases of fid than in neutral continuations, i.e. in continuations where there is no indication of a speaker or narrator functioning as perspectival center. our results were more compatible with the stronger version of hinterwimmer and bosch’s (2017) analysis on which discourse topics are automatically interpreted as perspectival centers in neutral narration than with the weaker version on which there is only a weak tendency for them to be interpreted as perspectival centers in neutral narration: dpros picking up topical referents were rated as most native-like/acceptable in sentences that expressed the narrator’s perspective and least native-like/acceptable in neutral continuations, with sentences in fid falling in between. still, on the stronger version, it is predicted that there are no reliable differences between neutral continuations and fid continuations, not, that sentences in fid are rated better than neutral continuations. we provided some speculative remarks as to why these unexpected results might have come about. however, more research is required to come to more definitive conclusions. in particular, it would be worthwhile to carry out a further experiment in which the stimuli used in experiments 1 and 2 are presented auditorily, in order to see whether the acceptability of sentences with dpros generally increases, as reported in section 3.2. above for two other experiments with dpros. additionally, it would be worthwhile to focus on the interaction of pronoun resolution and determination of perspectival centers in further experimental studies. we would like to end this paper by pointing out that an alternative to the anti-logophoricity hypothesis is briefly mentioned in hinterwimmer and bosch (2017) that also seems to be compatible with the experimental results reported in this paper. on the alternative, a strictly prominence-based account, dpros generally avoid the most prominent discourse referents as antecedents or binders. in the absence of perspectival centers (i.e. in instances of neutral narration), topics (or, in cases of binding, subjects) are maximally prominent. if there is a perspectival center, in contrast, the perspectival center is maximally prominent. one option to make this alternative account work would be to assume, first, that while speakers automatically introduce discourse referents (hunter, 2013), narrators only introduce discourse referents when there is an indication of them being present as perspectival centers (see altshuler and maier, to appear for relevant discussion). second, one would have to assume that discourse referents introduced by narrators are more prominent than topical discourse referents, which does not seem obvious. in light of this additional complication, and since its empirical predictions do not differ from those of the strong version of the anti-logophoricity account,6 we chose not to pursue the prominence-based account in this paper. teasing apart the predictions of the strong version of the anti-logophoricity account and the prominence-based account is a topic that we are planning to come back to in future research, however. 6 we are grateful to two anonymous reviewers for pointing this out to us. hinterwimmer, brocher and patil 125 acknowledgements we would like to thank sara meuser, janne schmandt, felix jüstel, magdalena schmitz and carina rothkegel for help in preparing and running the experiments reported in this paper, and the audience at the 7th biannual experimental pragmatics conference in cologne and at the xprag meeting at the university of graz 2018 for comments and discussion. the research reported in this paper was funded by a research grant from the deutsche forschungsgemeinschaft (dfg) for the project bound-variable-like interpretations of demonstrative pronouns, complex demonstratives and definite descriptions and for the project c05 discourse referents as perspectival centers of the collaborative research center 1252 prominence in language (university of cologne). all experimental stimuli, data and statistical analyses related to the research reported in this paper are accessible via the following link: https://osf.io/7ce94/ references altshuler, david and emar maier. death on the freeway: imaginative resistance as narrator accommodation. in ilaria frana, paula menendez benito and rajesh bhatt (eds), making worlds accessible: festschrift for angelika kratzer. amherst: umass scholar works, forthcoming. baayen, r. harald., douglas j. davidson and douglas m. bates (2008). mixed-effects modeling with crossed random effects for subjects and items. journal of memory and language 59, 390–412. banfield, ann (1982). unspeakable sentences: narration and representation in the language of fiction. boston: routledge. barr, dale j., roger levy, christoph scheepers and harry j. tily (2013). random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3), 255–278. bosch, peter, graham katz and carla umbach (2007). the non-subject bias of german demonstrative pronouns. in monika schwarz-friesel, manfred consten and mareille knees (eds.), anaphors in text, 145–164. amsterdam and philadelphia: benjamins. doi: https://doi.org/10.1075/slcs.86.13bos bosch, peter and carla umbach (2006). reference determination for demonstrative pronouns. in dagmar bittner and natalia gagarina (eds.), proceedings of the conference on intersentential pronominal reference in child and adult language (zaspil 48), 39–51. berlin: zentrum für allgemeine sprachwissenschaft, sprachtypologie und universalforschung. brinton, laurel (1980). ‘represented perception’: a study in narrative style. poetics, 9: 363– 381. charnavel, isabelle (2019). locality and logophoricity: a theory of exempt anaphora. in oxford studies in comparative syntax. oxford university press. charnavel, isabelle and victoria mateu (2015). the clitic binding restriction revisited: evidence for antilogophoricity. the linguistic review 32(4), 671–701. clements, george n. (1975). the logophoric pronoun in ewe: its role in discourse. the journal of west african languages 10: 141–177. doron, edit (1991). point of view as a factor of content. in steven k. moore and adam z. wyner (eds.), proceedings of semantics and linguistic theory (salt) i. cornell university, ithaca, ny, 51–64. dowty, david (1991). thematic proto-roles and argument selection. language 67. 547–619. dubinsky, stanley j. and robert hamilton (1998). epithets as antilogophoric pronouns. linguistic inquiry 29. 685–693. eckardt, regine (2014). the semantics of free indirect discourse. how texts allow to mindread and eavesdrop. leiden: brill. demonstrative pronouns as anti-logophoric pronouns 126 farner, geir (2014). literary fiction: the ways we read narrative literature. new york, ny: bloomsbury academic. geographical distribution of german speakers. (n.d.). in wikipedia. retrieved january 12, 2020, from https://en.wikipedia.org/wiki/geographical_distribution_of_german_speakers. gibson, edward and evelina fedorenko (2013). the need for quantitative methods in syntax and semantics research. language and cognitive processes, 28(1–2): 88–124. hinterwimmer, stefan (2015). a unified account of the properties of german demonstrative pronouns. in patrick grosz, pritty patel-grosz and igor yanovich (eds.), the proceedings of the workshop on pronominal semantics at nels 40, 61–107. amherst, ma: glsa publications, university of massachusetts. hinterwimmer, stefan and peter bosch (2017). demonstrative pronouns and propositional attitudes. in pritty patel-grosz, patrick g. grosz and sarah zobel (eds.), pronouns in embedded contexts (studies in linguistics and philosophy), 105–144. dordrecht: springer. hinterwimmer, stefan and andreas brocher (2018). an experimental investigation of the binding options of demonstrative pronouns in german. glossa: a journal of general linguistics, 3(1), 77. doi: http://doi.org/10.5334/gjgl.150 hunter, julie (2013). presuppositional indexicals. journal of semantics 30 (3), 381–421. doi: https://doi.org/10.1093/jos/ffs013 jaeger, t. florian (2008). categorical data analysis: away from anovas (transformation or not) and towards logit mixed models. journal of memory and language 59, 434–446. kaiser, elsi (2010). effects of contrast on referential form: investigating the distinction between strong and weak pronouns. discourse processes 47. 480–509. doi: https://doi. org/10.1080/01638530903347643 kaiser, elsi (2011a). on the relation between coherence relations and anaphoric demonstratives in german. in ingo reich, eva horch and dennis pauly (eds.), proceedings of sinn und bedeutung 15, 337–351. saarbrücken: saarland university press. kaiser, elsi (2011b). salience and contrast effects in reference resolution: the interpretation of dutch pronouns and demonstratives. language and cognitive processes 26(10). 1587– 1624. doi: https://doi.org/10.1080/01690965.2010.522915 kaiser, elsi (2013). looking beyond personal pronouns and beyond english: typological and computational complexity in reference resolution. theoretical linguistics 39. 109–122. doi: https://doi.org/10.1515/tl-2013-0007 kaiser, elsi and john c. trueswell (2008). interpreting pronouns and demonstratives in finnish: evidence for a form-specific approach to reference resolution. language and cognitive processes 23(5). 709–748. doi: https://doi.org/10.1080/01690960701771220 van krieken, kobie (2018). ambiguous perspective in narrative discourse: effects of viewpoint markers and verb tense on readers’ interpretation of represented perceptions. discourse processes, 55:8, 771–786, maier, emar (2015). quotation and unquotation in free indirect discourse. mind & language 30, 345–373. mayol, laia and robin clark (2010). pronouns in catalan: games of partial information and the use of linguistic resources. journal of pragmatics 42. 781–799. doi: https://doi. org/10.1016/j.pragma.2009.07.004 nishigauchi, taisuke (2014). reflexive binding: awareness and empathy from a syntactic point of view. journal of est asian linguistics 23.157–206. palmer, alan (2004). fictional minds. lincoln, ne: university of nebraska press. patel-grosz, pritty (2014). epithets as de re pronouns. in christopher piñón (ed.), empirical issues in syntax and semantics 10. patel-grosz, pritty and patrick g. grosz (2017). revisiting pronominal typology. linguistic inquiry 48. 259–297. doi: https://doi.org/10.1162/ling_a_00243 patil, umesh, peter bosch and stefan hinterwimmer (2020). constraints on german diese demonstratives: language formality and subject-avoidance. glossa. a journal of general linguistic, 5(1), 14. doi: http://doi.org/10.5334/gjgl.962 pearson, hazel (2015). the interpretation of the logophoric pronoun in ewe. natural language semantics 23. 77–118. hinterwimmer, brocher and patil 127 primus, beatrice (1999). cases and thematic roles – ergative, accusative and active. tübingen: niemeyer. primus, beatrice (2006). hierarchy mismatches and the dimensions of role semantics. in ina bornkessel, matthias schlesewsky and bernard comrie (eds.), semantic role universals and argument linking. theoretical, typological and psycholinguistic perspectives. berlin: de gruyter, 53–88. r core team (2018). r: a language and environment for statistical computing [computer software manual]. vienna, austria. retrieved from https://www.r-project.org schlenker, philippe (2004). context of thought and context of utterance. a note on free indirect discourse and the historical present. mind and language 19, 279–304. schütze, carson t. and jon sprouse (2014). judgment data. in r. podesva and d. sharma (eds.), research methods in linguistics. cambridge: cambridge university press. schütze, carson t. (2016). the empirical base of linguistics: grammaticality judgments and linguistic methodology. language science press. schumacher, petra b., manuel dangl & elyesa uzun (2016). thematic role as prominence cue during pronoun resolution in german. in anke holler and katja suckow (eds.), empirical perspectives on anaphora resolution, 213–240. berlin: de gruyter. doi: https:// doi.org/10.1515/9783110464108-011 schumacher, petra b., leah roberts and juhani järvikivi (2017). agentivity drives real-time pronoun resolution: evidence from german er and der. lingua 185. 25–41. doi: https:// doi.org/10.1016/j.lingua.2016.07.004 sells, peter (1987). aspects of logophoricity. linguistic inquiry 18. 445–479. sharvit, yael (2008). the puzzle of free indirect discourse. linguistics and philosophy 31, 353–395. sundaresan, sandhya (2012). context and (co)reference in the syntax and its interfaces. phd thesis, university of stuttgart. wiltschko, martina (1998). on the syntax and semantics of (relative) pronouns and determiners. journal of comparative germanic linguistics 2. 143–181. yashima, jun (2015). antilogophoricity: in conspiracy with the binding theory. ph.d. thesis, university of california, los angeles. dialogue & discourse 15(2) (2024) 85-112 doi: 10.5210/dad.2024.203 ©2024 dagnew mache asgede this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). self-repair in tigrinya: trouble sources, mechanisms and solutions dagnew mache asgede dagnewmache.a@gmail.com department of english language & literature https://orcid.org/0000-0002-1109-4415 arba minch university arba minch, ethiopia editor: kallirroi georgila submitted 11/2022; accepted 07/2024; published online 10/2024 abstract this paper analyzes conversational self-repair, which refers to reconstructing problematic portions of a prior oral discourse by oneself, in tigrinya. tigrinya is a northern ethio-eritrean-semitic language spoken by the inhabitants of the tigray regional state of ethiopia and eritrea. this article relies on recorded oral data from speakers of the rayya tigrinya variety, particularly inhabitants of neksege located to the west of maichew. a conversational analysis (ca) approach is used to analyze the trouble sources, mechanisms (initiators), and results of self-repair. the article shows that pronunciation problems emanating from dialectal variation or tongue slip, wrong word order including focus misplacement, missing constituents, perceived misunderstandings, and using (totally) wrong constituents are some of the trouble sources that push speakers to repair portions of a prior oral utterance. on top of that, cut-offs, particles, and lexemes (one verbal noun and some predicates) are identified as self-repair initiators. though cut-offs do not indicate a self-repair, particles and predicates may sometimes indicate a self-repair. finally, the article posits that expanding, replacing, re-ordering, aborting and restarting, and inserting are some of the solutions set for the repairable segments in the repaired portions of the oral discourse. the author recommends for further investigation repair in the process of language acquisition and learning, and the relationship between self-repair and the demographic features of participants. keywords: tigrinya, self-repair, trouble sources, repair mechanisms, repair solutions 1 introduction linguistic repair refers to addressing any language-related concerns during the development of oral discourse. repair can be self-initiated or other-initiated. this paper specifically focuses on self-initiated repair, which involves correcting one’s speech to improve communication. various scholars have studied self-repair in different languages. for example, fox et al. (1996) state that the grammatical structure of a language influences how speakers correct their speech. murphy (2019) also indicates a positive relationship between repairs and the phonological and morpho-syntactic features of languages. repair, in general and self-repair in particular, is one of the most overlooked areas of linguistic research in tigrinya, and in ethiopian languages as a whole. this paper aims to investigate the sources of trouble, mechanisms, and solutions in tigrinya to contribute to the discussion on conversational self-repair. the work is based on annotated oral data. the paper is divided into four main sections. the first section, literature, introduces the sociolinguistic and grammatical features of tigrinya for readers who are unfamiliar with the language, as well as basic concepts of conversational repair and different types of repairs. the second section covers methods and data, providing details on the data collection tools used, information about the data, protocols for annotating the data, and how inter-rater agreement was calculated among four raters. the third section presents and analyzes the data, focusing on trouble sources, repair mechanisms, and the results of self-repair. finally, a discussion is provided based on the data results. https://orcid.org/0000-0002-1109-4415 asgede 86 2 literature 2.1 the tigrinya tigrinya, spoken by the people of the tigray regional state in ethiopia and by eritreans, is one of the semitic language branches of the afro-asiatic language phylum. tigrinya, along with geez and tigre, are classified as a northern ethio-eritrean-semitic language group. though the language is said to have several dialects, no single study has attempted to delineate the dialect numbers based on linguistic characteristics; researchers ignored the dialectal mutual intelligibility of the language’s varieties. roughly speaking, we can assume that the language has several dialects; of these, the one spoken by the people of tigray regional state’s southern zone, where the data this article relies on is recorded from, is known as the rayya tigrinya variety. the language is used as a medium of instruction from kindergarten to grade eight, and it is taught as a course in grades nine through twelve. the language is also taught at abiyi adi teachers’ college and mekelle university at diploma and degree levels respectively. at the moment that i am writing this article, due to the devastating war on tigray regional state, none of the mentioned programs are active. tigrinya has 37 consonants (asgede, 2019). asgede (2019) and mehari (2021) argue that the consonant segments /p/, /p’/, /v/ and /ʒ/ are ‘non-existent’ in tigrinya except in loan words. asgede (2019) claims that loan words contain sounds that are difficult to pronounce for the illiterate part of the speech community. concerning vowel sounds, the language has seven common vowels; the majority of the speakers of rayya and wajerat dialects do not utter the middle front unrounded vowel /e/ which is common in the quasistandard tigrinya variety (asgede, 2019; mehari, 2021). tigrinya has cv and cvc syllabic structure (yohannes, 2002; asgede, 2019; mehari, 2021). in wordinitial and final, a consonant cluster is impermissible; if there is either a word-initial or final potential consonant cluster, the epenthetic vowel (high central unrounded) /-ɨ-/ and the high front unrounded vowel /-i/ are inserted respectively. when two vowels appear in sequence in complex words that contain two or more morphs, the semivowels /-w-/ or /-j-/ or the glottal consonant /-ʔ-/ are inserted as epenthetic consonants. besides, consonant gemination that is phonemic in most cases is common in tigrinya at word medial but never in other word positions. tigrinya is an inflectional (i.e., fusional under the synthetic morphological type) language where several morphs with different grammatical meanings appear in a single word, and a single morph encodes several meanings and/or grammatical functions simultaneously. it is morphologically complex to the extent that a word contains elements that are equal to a sentence. this makes the gloss and translation of the data challenging, which leads to some tolerable problems. for example, the roots of the verbs in the language are templates of consonants into which vowels are inserted, and affixes are attached to mark different grammatical functions. the word order for tigrinya simple declarative sentence is s-o-v as in ʔɨssu goggo bəlɨʕu ‘he ate enjera’. however, this does not work when all grammatical elements are indicated in a single word as in the following example. (1) introspective data bəlɨʕ-u-ww-o bəlɨʕ-u-ww-o eat:prv-sj:3sm-ø-obj:3sm ‘he ate it (m).’ the order of the affixes that mark the subject and object follows the verb. the language has also circumfix that marks negation as in ʔajtɨχədɨj [ʔaj-tɨ-χəd-ɨ-j ‘neg-2smːsub-goːipv-ø-neg’] ‘you will not go’. though the grammatical elements of the language are relatively emphasized in previous works, the features of conversation have been overlooked by researchers. as far as my reading is concerned, it is only asgede (2023; 2019) who tried to touch on some of the conversational features (particularly the performance-related elements) of tigrinya. this article therefore will focus on describing self-repair as one self-repair in tigrinya: trouble sources, mechanisms and solutions 87 feature of conversation. to achieve this, the paper describes the trouble sources, the mechanisms, and the results of self-repair in different genres of oral usage of the language. in addition to a grammatical description, teferra (1979) and blejer (1986) provided some brief discussions on some pragmatic elements like discourse markers (dms) of the language. they described particles, conjunctions, adverbials, and interjections as the main sources of dms. teferra (1979) has noted that =s and/or =si ‘as for’ and ʔɨmma or its clitic form =mma ‘focus marker’ are focus marking particles; =ja ‘evidential marker’ an enclitic, ʔɨkko or its clitic form =kko ‘change in focus’, dəʔa ‘and so, then’; ʔɨmmo ‘then’; wəy dɨmma ‘or’. blejer (1986) also stated that the enclitic =s and/or =si ‘as for’ does mark focus. though previous studies like mehari (2021; 2011), asgede (2019; 2007), girmay (2012), nigusse (2012), yohannes (2002), and kiros (2009) have specifically tried to deal with different linguistic features of the different tigrinya varieties, it is asgede (2023; 2019) who gives room (though limited) to the conversational features of the rayya language variety. 2.2 conversation repair conversational repair refers to pointing back to a problematic segment of oral discourse and pointing forward to its more accurate replacement. in other words, repair refers to a process of reformulating a proposition. it is one of the performance matters of conversation or speech; this process may be initiated by cut-offs, filled pauses, lexical constituents, etc. filled pauses are linguistic forms that help a speaker and listener get more time to deal with problematic information that is given in a portion of a prior segment (brennan and schober, 2001 cited in hlavac, 2011). as to deese (1984) and brennan and williams (1995), filled pauses draw the attention of a listener to pay due emphasis to what follows in the development of speech. roughly speaking, after fox et al. (1996), a repair can be defined as a conversational process “by which speakers correct errors they have made in their immediately prior talk”. however, a repair takes place even in the absence of any error committed; repairable segments may be repaired for clarification, rephrasing, and explanations. with this, schegloff et al. (1977) stated that a repair is not limited to error correction, but a repair is made even if there are no errors made. in this paper, a repair is used as a conversational process that allows speakers to reconsider their ‘problematic’ prior utterances. speakers repair a portion of their prior constituents in two ways: self-initiated and other-initiated whose differences are briefly explained in the following sub-section. 2.3 types of repairs a repair can be either other-initiated or self-initiated. in a conversation, when something that a speaker has said is misheard or when a speaker encounters a problem in understanding a previous utterance, an attempt is often made to elicit a repetition, clarification, elaboration, or correction, referred to in conversation analysis as other-initiation of repair (schegloff et al., 1977 cited in ha and grice, 2017). other-initiated repair refers to either correcting others’ prior utterance or being corrected by others. a self-repair in contrast refers to reconsidering your portion of prior discourse; it refers to repairing your problematic prior utterance (self-initiation). a self-repair is made by speakers when they recognize they had made a mistake or their prior discourse is problematic in any way. self-repair may be accomplished within an oral discourse unit like sentence and turn, or even after long strings of constituents and turns are spoken out. self-initiated repair within the same turn may be initiated by varieties of non-lexical speech perturbations such as cut-offs and repetitions, and sound stretches (schegloff et al., 1977). similarly, depending on the mini-corpus that this study relies on, in tigrinya, repair takes place either after part of a word and sometimes even after an entire word or after a long discourse segment is uttered. for instance, as asgede (2019) mentions, reformulating is dominantly signalled by cut-off words, as can be seen in example (2) below. asgede 88 (2) (taken from asgede (2019: 215)) s’aʕɨda ʃɨngurti ʔɨm... d͡ʒɨ.. rɨħus d͡ʒɨnd͡ʒɨbɨl rɨħus d͡ʒɨnd͡ʒɨbɨl ʔabaʕɨχə səssəg s’aʕɨd ʃɨngurt ʔɨm... d͡ʒɨ.. rɨħus d͡ʒɨnd͡ʒɨbɨl rɨħu d͡ʒɨnd͡ʒɨbɨl ʔabaʕɨχə səssəg white onion dm_plnprc dm_cnvrpr wet ginger wet ginger fenugreek basil ‘garlic, wet ginger, fenugreek, basil…’ in example (2), the speaker started to utter the word d͡ʒɨnd͡ʒɨbɨl ‘ginger’ but remembered that the word should be modified by the adjective rɨħus ‘wet’. by implication, the ginger needs to be wet, not dry. it seems that the speaker remembers she had missed a constituent that should describe the partially uttered word; then, she cuts off the word to add an adjective that describes the exact state of the ginger she is referring to. the cut-offs that are considered here are those that initiate repairing (or might have other purposes) but are followed by the corrected/revised version of the repairing source. the current article focuses on such self-initiated repairing trouble sources, repair mechanisms, and repair results. 3 methods and data 3.1 methods and procedures of data collection 3.1.1 tools of data collection data was gathered by recording informal interviews, spontaneous conversations, and tales. this was accomplished during a three-month personal field trip, from april 10th to july 17th 2019, in neksege (located in the southern tigray administrative zone). the recorded data was processed using software audacity, and the format was exported as a wav file. the wav files containing the oral data were imported into elan, which was used to segment each text into discourse units such as turns, external interventions, sentence and clause boundaries. the segmented data was phonetically transcribed; morphs of each complex word were marked; an english gloss was assigned to each morph, and each discourse unit received a literal and free translation in english. 3.1.2 data statement the mini-corpus, temporarily named as rayya tigrinya mini-speech corpus, was established during a threemonth fieldwork in which the researcher recorded the speech data. the speech corpus was established from recordings of spontaneous conversations, conflict resolutions, informal interviews, descriptions of how to do things, biographies, and tales. originally, the mini-corpus totalled 24 hours of recorded audio/video data. for research ethics purposes, based on the basic schema recommended by bender and friedman (2018), the data statement is detailed as follows. the data was annotated by four annotators who were asked for their judgement on the sources of troubles, self-repairing mechanisms, and self-repair solutions of the 400 identified self-repair markers (for a detailed discussion on the data statement please see appendix ii). 3.1.3 annotation protocol there are established terminologies used in the literature in relation to how to establish structure in the annotation for repair. it is clear that a self-repair is initiated by an original utterance (hough, 2015) that i prefer to call trouble source, which is named as reparandum by meteer et al. (1995) and shriberg (1994). this type of constituent is split into two, namely start and reparandum by ginzburg et al. (2007). the trouble source of an utterance that underwent repair is defined as “the entire stretch of speech to be deleted” by shriberg (1994). that constituent is preceded and followed by square brackets [] (as in [constituent]). the other element that needs clarification is the repair mechanisms that are termed as interregnum which refers to an editing phase marked by “filled pauses or editing terms” (meteer et al., 1995; shriberg, 1994; hough, 2015). this section is termed as editing term (ginzburg et al., 2007), editing phase (levelt, 1983), cut-off to repair (nakatani and hirschberg, 1994) and repair initiation (schegloff, 1992). this part of a constituent is preceded and followed by {} as in {constituent} and in a bold font. self-repair in tigrinya: trouble sources, mechanisms and solutions 89 the third concept that is relevant to the present article is repair solutions that are named differently in the literature: repair (hough, 2015; shriberg, 1994; meteer et al., 1995), repair proper (schegloff, 1992), and alteration (ginzburg et al., 2007). this is marked by # as in #constituent#. this description is explained as follows. in the literature, an interruption point (meteer et al., 1995; shriberg, 1994) named as moment of interruption by levelt (1983), blackmer and mitton (1991), and ginzburg et al. (2007), that is the moment that appears between the trouble source and repair mechanism, is not considered in the analysis for it demands a different experiment and may not be understood by considering the surface structure of a discourse. not only that but some authors like hough (2015) and ginzburg et al. (2007) mentioned that there is a portion of discourse called continuation that is not considered here for i believe it does not affect the nature of the self-repair. 3.1.4 inter-annotator agreement based on an example text that contained 15 repairs, the researcher trained the annotators and raters for half a day. the content of the training was how to identify the repair typologies and how to annotate (transcribe) them. based on healey et al. (2005), a short guideline encompassing some detailed points is given to the annotators. for details on the guideline and the mean of the identified self-repair markers, see appendix iii. considering the total number of repairs identified by the researcher (see appendix ii) and the remaining three annotators, the mean is 390. therefore, the 400 repairs are accepted as valued and reliable. this is substantiated with the inter-rater agreement among the four raters (including the author) for the categories of self-repairs: trouble sources, repair mechanisms, and results of self-repairs under sections 4.1., 4.2., and 4.3., respectively. the fleiss’ kappa model is employed to manually compute the inter-rater agreement for there are four raters (more than two raters) and the data is nominal (fleiss, 1971; 2003). 𝐾 = po−pe 1−pe where k is fleiss’ kappa; po is observed agreement, pe is expected agreement among the raters the inter-rater agreement for the sources of troubles in self-repair, repair mechanisms, and results of self-repairs is calculated in the analysis section using the fleiss’ kappa interpretation intervals provided in the equation above. 3.2 data analysis framework the data analysis framework used is conversation analysis (ca), which is a method for studying human social interaction in the field of linguistics (stivers and sidnell, 2013). according to maynard (2013), ca has “established itself as a worldwide theoretical and empirical endeavor concerned with the social scientific understanding and analysis of [human social] interaction”. ca necessitates distinct data derived from spontaneous and naturally occurring social interaction (maynard, 2013). maynard (2013) adds that this approach is concerned with how language is used in “social, publicly interpretable methods and behaviors”. ca uses audio and/or video recordings of naturally occurring communicative events like an everyday conversation, instructions/explanations of how to do things, ritual speeches, story tales, and stories of people’s lives to study the detailed actions of participants. the audio-video data is significant because sɨʔləj ʔɨbba daɲɲəw ʔɨkko məs’u ʔɨnho [sɨʔləj] {ʔɨbba} #daɲɲəw# ʔɨkko məs’u ʔɨnho trouble source repair mechanisms repair solutions ‘si’eley i mean dagnew has come.’ asgede 90 ca aims to “describe the organization of ordinary social activities such as taking turns at talking” (mondanda, 2013). repair is a common phenomenon in conversation. ca aids in the investigation of how and why repairs are performed. according to yilmaz (2004), a repair can be initiated by oneself or by someone else. this feature of conversation indicates either disfluencies, difficulties activating language units and/or concepts, hearing, and understanding. justifying why a speaker requires repair may be difficult because the reasons in most cases are cognitive and remain unknown. taking an inquisitive look at both the linguistic and social contexts, as well as looking into the meanings and functions of repair mechanisms in the text, can help to uncover the hidden cognitive reasons a speaker performed a repair (schourup, 1985). in this regard, ca aided the researcher in examining the corpus neutral of hypothesis or theory; this approach helps to record theory neutral linguistic data. this is what wong and waring (2010) refer to as “unmotivated looking”. to make using the ca approach easier, the entire corpus was transcribed. whenever possible, the transcription included “not only what was said, but also how it was said” (wooffitt, 2005) by providing literal and free meanings of extractions. similar to the work of many researchers for other languages (jabeen et al., 2011; ward, 1998; kawamori et al., 1998; heeman et al., 1999; de rooji, 2000; matras, 2000; bell, 2010; wang, 2011), here ca is used as a method to analyze the linguistic features and pragmatic functions of trouble sources such as discourse markers, filled pauses, cut-offs, etc. in tigrinya. based on the host context, the function and meaning of each identified self-repair marker were considered. examining what happened to the repair source and its alteration text, it is attempted to investigate repair techniques, reasons for repairing, and devices used to indicate that repairing is being initiated. 4 data presentation and analysis 4.1 trouble sources of self-repairing self-repairing occurs in a conversation for a reason; trouble sources cause a speaker to reconsider a portion of an old utterance. self-initiated repair is triggered by any of the following conversational issues: diction problems caused by dialect variations and social relationships of the participants, pronunciation errors (and tongue slip), word order issues, missing constituents, focus misplacement, and misunderstandings. the distribution of the problems of the repairs identified from the data is described in the table below for clarity. table 1 demonstrates that the majority of the repair initiators are pronunciation errors and tongue slips (90 items), followed by diction problems (68 items), missing constituents (65 items), word order problems (49 items), and misunderstandings (43 items). as demonstrated in table 2 below, the inter-rater agreement trouble/repair initiator types total number of cases identified percentage diction problems dialect variation participants social relationship 68 32 36 17 47.059 52.941 pronunciation errors and tongue slips 90 22.5 word order problems 49 12.25 missing constituents 65 16.25 focus misplacements 33 8.25 misunderstandings 43 10.75 (totally) wrong constituents 51 12.75 unidentified 1 0.25 total 400 100% table 1. list of trouble source types. self-repair in tigrinya: trouble sources, mechanisms and solutions 91 d ic ti o n p ro b le m s p ro n . e rr o rs an d t o n g u e s li p s w o rd o rd er p ro b le m s m is si n g c o n st it u en ts f o cu s m is p la ce m en ts m is u n d er st an d in g s w ro n g c o n st it u en ts u n id en ti fi ed s u m sum 277 349 199 255 141 169 200 10 1600 p 0.173 0.218 0.124 0.159 0.088 0.106 0.125 0.006 po 0.86 pe 0.153 k 0.835 table 2. inter-rater agreement on the trouble sources. among the four raters on the trouble sources is computed using the fleiss’ kappa model that is represented with the formula mentioned in section 3.1.4 and in appendix iii. based on the suggested fleiss’ kappa interpretation (see table 10 in appendix iii), the inter-rater agreement is 0.835, which is ‘almost perfect’. 4.1.1 pronunciation errors though not universal, and certainly not consistent across data collected from a speech community of a specific language or variety, the data available reveals that the pronunciation problem and tongue slip, which account for 22.5% of the trouble sources of the repairs identified, are the most powerful motivators for speakers to make a conversational repair. (3) arg1 səlus səlus [lajəs’a] #nab jəs’a# ʕɨdəga [tɨχəd] {ʔa...} #tɨχəjɨd# malət ʔɨju səlus səlus la=jəs’a na b jəs’a ʕɨdəga tɨ-χəd ʔa… tuesday tuesday to-yetsa to yetsa market 2sm:sbj-go:ipv part tɨ-χəjɨd malət ʔɨj-u 2sm:sbj-go:ipv means cop-3sm ‘every tuesday, you go to yetsa for market.’ the proclitic la= ‘to’ in lajəs’a ‘to yetsa’ functions similarly to the preposition nab ‘to’ in grammatical and syntactic terms. the speaker has repaired the lajəs’a, which was a source of trouble during the repair, and replaced it with nab jəs’a to correct the way the prepositional phrase should be pronounced. the allative marker in quasi-standard tigrinya is the preposition nab ‘to’, but it is commonly attached to the next noun by rayya tigrinya variety speakers who habitat in neksege; by other rayya tigrinya variety speaker communities, it is realized as da= ‘to’ as in daməxələ ‘to mekele’. in the same turn, the speaker corrected the way he pronounced the word tɨχəd as tɨχəjɨd ‘you go’. the problematic and corrected constituents are similar in meaning but differ in pronunciation due to dialectal variation; the source of the repair therefore seems to be a pronunciation problem that arose from dialectical differences between the participants. 4.1.2 diction problems in addition to ‘inappropriate’ pronunciation, diction-related trouble is a repair initiator. diction, at least in this article, is influenced by the participants’ social relationships as well as dialect variations. this issue accounts for 68 (17%) of the total 400 repairs identified in the mini-corpus. asgede 92 (4) asg2 ʔatta ʔabzi labəlaʕɨnadə ʔabzi tɨħar.. tɨftɨnwa rɨħɨχ’ ʋəl ʔatt-a ʔab-ʔɨz-i la-bəlaʕ-ɨna-də ʔabzi tɨħar.. voc-2sm loc-prx-2sm prog-eat:ipv-sj:1pl-q loc-prx-2sm rpr tɨ-ftɨn-wa rɨħɨχ’ bəl 2smː2sbj-urinateːipv-foc away:cn say:imp:2sm ‘please urinate away for we are feeding ourselves here.’ the repair initiators are the cut-off constituents of the above extraction; the speaker was to utter tɨħarɨʔ ‘urinate’, which the speech community interprets as derogatory. instead, the speaker used the repair result constituent tɨftɨn ‘urinate’, which is a meaning extended from ‘trying, experimenting’. the repair result constituent tɨftɨn ‘urinate’ is a commonly used word by the quasi-standard tigrinya speakers and is used in media and written texts across tigray. the speaker seems to believe that the trouble source constituent is ‘inappropriate’ in the context of people eating food. besides, words may be replaced by others as a result of a repair process arising from the participants’ social relationships. (5) ast1-3 [ħɨləf] #ħalɨfom# kəf jɨʋəlu ħɨləf ħalɨf-om kəf jɨ-bəl-u pass:imp:sj:2sm pass:imp-2sm.h sit juss-say:imp-2sm.h ‘pass through…please be seated in the next seat.’ qsh1-1 ʔɨʃʃɨni ʔɨzi wəddəj marjam taχɨbɨrχa ʔɨʃʃɨ-ni ʔɨz-i wəddi-əj marjam tɨ-ʔa-χɨbɨr-χa okay-plt prx-2sm son-rl:1s (saint) merry juss:sbj:3sf-caus:3sf-respect:ipv-obj:2sm ‘okay, my sonǃ may saint merry respect you.’ ast1-4 ʔamjən #jɨk’ɨrta# dəjalələχukum ʔɨkko jɨʔu ʔatta dɨʋəχukum ʔamjən jɨk’ɨrta daj-a-lələj-χu-kum ʔɨkko jɨʔ-u amen sorry neg-caus-notice:cn-sj:1s-obj:2sm.h foc cop-3sm ʔatt-a dɨ-bəl-χu-kum voc-2sm rel-say:prv-sj:1s-obj:2sm.h amen! i’m sorry for i called you as if you are my peer for i didn’t notice it was you. qsh1-2 daħna ʔɨzi wəddəj ʔarɨmχajjakko daħna ʔɨz-i wəddi-əj ʔarɨm-χa-jj-a-kko fine prx-2sm son-rl:1s correct:prv-sj:2sm-ø-obj:3sf-foc ‘my son, do not mention it; you already corrected it (the inappropriate speech).’ as depicted in example (5), ħɨləf ‘(2sm) pass through’ is an imperative mood that is not respectful. the repair result which is ħalɨfom ‘after (2sm.h) pass through’ has an honorific marker in it. even though there is no covered verbal repair initiator, it is understood that the speaker replaces the ħɨləf with ħalɨfom because the speaker recognizes the addressee is a priest who has a high social class at least from a religious perspective. this becomes clear when the main speaker apologizes which is explicitly communicated by the constituent jɨk’ɨrta ‘sorry’ (ast1-4) and the other speaker accepts and even acknowledges the interlocutor unknowingly spoke ‘inappropriately’. 4.1.3 missing constituent spontaneous speech has many disfluencies which differentiate spoken genres from written genres. omitting a constituent is one of the disfluencies. constituents of a single discourse unit (say for instance a turn) and long constituents (a discourse unit like sentences) might be omitted. the latter refers to missing, for example, a step or steps while describing how to do things or missing part of a plot when narrating stories and tales. this type of conversational feature is another trouble source that leads speakers to self-repair a repairable prior segment. self-repair in tigrinya: trouble sources, mechanisms and solutions 93 (6) bma4-51 [ʔɨta ʔom] #ʔɨta nəwaħ ʔom# ʔɨmma hɨzχajja ħɨləf laʕaddɨmma hɨzχajja ħɨləf ʔɨt-a ʔom ʔɨta nəwaħ ʔom ʔɨmma hɨz-χa-jj-a dest-3sf tree dest-3sf long:sf tree foc hold:ipv-sj:2sm-ø-obj:3sf ħɨləf la-ʕaddi-mma hɨz-χa-jj-a ħɨləf pass:imp:sj:2sm all-home-foc hold:ipv-sj:2sm-ø-obj:3sf pass:imp:sj:2sm ‘you bring the tree… the long tree to home.’ dma4-29 /la.. tʃ’əllə la.. tʃ’əllə rpr okay ‘okay’ example (6) shows how missing constituents in the same turn initiate repair. the main speaker wants his interlocutor to bring a tree and later he remembered the tree should be not just the tree but the long tree than the other alternative trees. the tree, therefore, needs to be described by the adjective nəwaħ ‘long:sf’. this is the linguistic context that tells the participants that there are two or more trees, and the main speaker is saying that the interlocutor should bring the long tree with the repaired constituent. in the same turn, the main speaker spoke to the addressee to bring the tree to a place that was not mentioned originally in the trouble source; the other participant was also to ask a question and started uttering the word labəj ‘to where’ but cut-off as la… however, before the interlocutor uttered his question (which seems to not have been heard by the main speaker), the main speaker noticed he had forgotten to mention the destination and repaired the problematic constituent as laʕadimma hɨzχaja ħɨləf ‘bring it to home:foc’. 4.1.4 wrong constituent commonly, speakers utter wrong constituents in their speech. soon, they repair their speech and replace the wrongly uttered constituent with an appropriate one. (7) kdw1-7 [sɨʔləj] {ʔɨbba} #daɲɲəw# ʔɨkko məs’u ʔɨnho sɨʔləj ʔɨbba daɲɲəw ʔɨkko məs’-u ʔɨnh-o name foc\rpr name foc come:prv-sj:3sm exist:ipv-sj:3sm ‘si’eley i mean dagnew has come.’ in the above excerpt, the speaker wanted to talk about dagnew but he spoke out si’eley who is the small brother of dagnew. here, the repair is initiated by a trouble source of uttering a wrong constituent. concerning this, speakers also put constituents in their jumbled order. to fix the wrong constituent, a speaker repairs his or her speech by himself or herself. 4.1.5 wrong constituents order as can be seen in the following example (8), in the first turn of the extraction, the speaker uttered bota nɨfas dɨnhəwwo ‘place windy’ but later he noted that the constituents are in their wrong order, and the adjectival phrase nɨfaj dɨnhəwwo ‘that has wind; windy’ should precede the head noun bota ‘place’ that is the intended constituent to be modified by the modifier. the repair is initiated due to the wrongly ordered constituents. (8) bma4-22 məd͡ʒəmərja [bota nɨfas dɨnhəwwo] {ʔɨm} #nɨfas dɨnhəwwo bota# tɨmərɨs’ məd͡ʒəmərja bota nɨfas dɨnhəwwo ʔɨm nɨfas first place wind rel-exist:ipv-sj:3sm-obj:3sm rpr wind dɨ-ʔɨnh-ə-ww-o bota tɨ-mərɨs’ rel-exist:ipv-sj:3sm-obj:3sm place imp:2sm-select:ipv ‘first, place, you select a windy (specific) place.’ dma4-17 ʔɨʃʃi ʔɨʃʃi okay ‘i am listening to you; please proceed to talk.’ asgede 94 bma4-23 kaʔu [sɨfɨħ ʔablɨχa tɨs’ərɨg bota] {malət} #sɨfɨħ ʔablɨχa bota tɨs’ərɨg# kaʔu sɨfɨħ ʔa-bl-ɨχa tɨ-s’ərɨg bota malət sɨfɨħ then wide caus-do:ipv-sbj:2sm imp:2sm-clean place mean wide ʔablɨχa bota tɨ-s’ərɨg caus-do:ipv-sbj:2sm place imp:2sm-clean ‘then, you clean a place wide i mean you clean a wide area.’ in the third turn (bma4-23) of the above excerpt, see how the wrongly ordered constituents are the trouble source of the repair which is initiated by the verbal noun malət ‘i mean; means’. the noun bota ‘place’ follows the verb tɨs’ərɨg ‘you clean’ (the verb preceded the object) which violates the word order of the language that is s-o-v. wrongly ordered constituents refer not only to words but also to long constituents of a discourse. for instance, as can be seen in example (9), the wrongly ordered constituents are two clauses that contain one step each. (9) mqe3-16 gəβɨt͡ ʃə tɨt͡ ʃ’awətləχa [məd͡ʒəmərta ʔɨmni tarsɨn kaʔu ʔab χɨltə tɨbʷadən] {ʔɨbba} #məd͡ʒəməja ʔab χɨltə tɨbʷadən kaʔu ʔɨmni tarsɨn# gəβɨt͡ ʃə tɨ-t͡ ʃ’awət-ʔɨntələ-χa məd͡ʒəmərta ʔɨmni t-a-rsɨn kaʔu ʔab gebiche ipr-play-when-sbj:2sm first stone ipv-cau:2-heat then into χɨltə tɨ-bʷadən ʔɨbba məd͡ʒəməja ʔab χɨltə tɨbʷadən kaʔu ʔɨmni two ipv-group:2sm i mean first into two ipv-group:2sm then stone t-a-rsɨn ipv-cau:2-heat ‘when you play gebiche (an indigenous game), you first heat stone; then, form two teams; i mean, first, form two teams and heat a stone.’ the first bold constituent in the above excerpt contains two steps that should be completed in a serial order before boys play an indigenous game called gebiche; however, the speaker has wrongly first uttered the second step, which is the trouble source of the repair that aims at reconsidering the way the speaker thought the ideas should appear. the second bold portion of the excerpt thus is the corrected order of the steps which were jumbled in the first part of the excerpt. 4.1.6 focus misplacement though scarcely, focus misplacement is also observed as another type of trouble that motivates speakers to repair their portion of the prior utterance. (10) ags1-22 [ʔɨssu səbʔaj] {ʔɨk..} #ʔɨssu ʔɨkko səbʔaj# jɨʔu ʔɨssu səbʔaj ʔɨk.. ʔɨssu ʔɨkko səbʔaj jɨʔ-u 3sm man rpr 3sm foc man cop-3sm ‘he man.. he is a man’ the cut-off initiates the repair process. in the repairable segment, the focus was meant in the səbʔaj ‘man’, but the speaker wants to shift the focus from səbʔaj ‘man’ to ʔɨssu ‘he’ through the process of the repair mechanism. 4.1.7 perceived misunderstanding when a current speaker who controls a particular turn believes that a listener may not understand, or may be confused by the information he/she has just heard, the speaker may repair. perceived misunderstanding is another cause of repair during a conversation. misunderstanding is a trouble source that accounts for 10.75% of the total repair sources among the 400 repairs identified. self-repair in tigrinya: trouble sources, mechanisms and solutions 95 (11) wra1-7 [ʔanəs χalɨʔ wulɨd jəbləj] ʔanə-s χalɨʔ wulɨd jə-bl-ə-j 1s-foc another child neg-say:ipv-rl:1s-neg ‘i do not have another child.’ (lit. he is the only son that i have.) wra1-8 #mɨwuladɨskka kʷuwankʷa jɨʔuwa# mɨwulad-ɨs-kka kʷuwankʷa jɨʔ-u-wa giving birth-foc-foc language cop-3sm-foc ‘blood relationship is meaningless.’ (lit. one might have a good child who is not a child by blood.) lma1-11 ʔɨwwə ħak’iχi ʔɨwwə ħak’i-χi yes true-poss:2sf ‘yes, what you said is true.’ wra1-9 #ʔɨgzihar jɨməsgən ʔɨsχatɨχum ʔɨnhəχumni# {malətəj} jɨʔu ʔɨgzihar jɨ-məsgən ʔɨsχa-tɨχum ʔɨnhə-χum-ni malət-əj god juss:3sm-thanks 2sm-pl exist:ipv-sj:2pl-obj:1s means-poss:1s jɨʔ-u cop-3sm ‘thanks to god, you are here for me.’ wra1-10 hamʔu gɨn ʔab məjda wədɨχ’ə χ’arjə ham-ʔu gɨn ʔab məjda wədɨχ’-ə χ’arj-ə like-3sm but in plain lay:cn-1s remain:prv-1s ‘he have left me alone.’ (lit. as to him, however, i am thrown around the field.) the excerpt above demonstrates two steps of repair: the main speaker explains why the speaker’s only son, who lives in jigjiga, did not visit her. she stated in the segment wra1-7 that she has no other biological child besides him. she is aware, however, that the listener has acknowledged and assisted her as if she is his biological mother. as a result, the main speaker perceived that the listener might be uncomfortable with what she had said. this perceived misunderstanding motivates her to reframe her speech. the second constituent of the discourse is a softening repair to the previous constituent of the utterance, though there is no explicit repair initiator. when we look at the turn wra1-9, it has a clear repair initiator that is malətəj ‘i mean’. this begins as the main speaker expands on the previous turn by providing additional clarification. as a result, perceived misunderstanding is another issue that drives a speaker to make a self-initiated conversational repair. 4.2 repair mechanisms the dominant possible repairing initiations which i call repairing mechanisms markers are the following: cut-offs, filled pauses, particles, predicates, pauses, and non-verbal such as facial expressions and gestures. for details on the repair mechanisms, look at the following table (table 3). type repairing mechanism number percentage cut-offs 238 59.5 filled pauses 60 15 particles 38 9.5 predicates 32 8 pauses 21 5.25 visible only (like facial expressions and gestures) 9 2.25 unidentified 2 0.5 total 400 100% table 3. types of repair mechanisms. asgede 96 c u to ff s f il le d p au se s p ar ti cl es p re d ic at es p au se s f ac ia l e x p re ss io n s an d g es tu re s u n id en ti fi ed s u m sum 948 238 148 128 86 36 16 1600 p 0.593 0.149 0.093 0.08 0.054 0.023 0.01 pe 0.392 po 0.961 k 0.936 table 4. inter-rater agreement on the repair mechanisms. as demonstrated in table 4, the computed inter-rater agreement is 0.936, which is interpreted as almost perfect agreement (see table 10 in appendix iii). to gain an understanding of the repair initiating markers in tigrinya, let us have a further discussion on each of the techniques with some examples. 4.2.1 cut-offs note that all cut-offs may not necessarily have relations with a repair; for example they may mark hesitations. in the mini corpus, at least 181 cut-offs were found not to have an association with a repair. as roughly examined, most of these 181 cut-offs are associated with the hesitation and/or planning process (see asgede, 2019). (12) dl2-79 {ʔɨto..} #ʔɨsχatɨχum naχ’adəm təmharo sɨlə dɨχonχum gobəzat jəχum# ʔɨto.. ʔɨsχatɨχum na-χ’adəm təmhari-o sɨlə dɨ-χon-χum gobəz-at rpr 2plm poss-past time student:pl for rel-become-2pl clever-pl jəχ-um being-2pl ‘rpr since you are from the old batches, you are clever students.’ ʔɨto... is part of an unfinished constituent of the host utterance of the main speaker, as shown in example (12). the speaker then begins to utter a different word, indicating that the unfinished word was either inappropriate or incorrect. the process is more of a repair than a hesitation. distinguishing between repair and hesitation is difficult; when a speaker resumes an utterance with a completely different constituent than the one that was left unfinished, the researcher refers to it as a repair. however, hesitation occurs when the speaker restarts her/his utterance with a similar part of the unfinished part of the language unit. (13) hlf3-33 [ʔɨti wannaʔu] {ʔɨk..} #ʔɨti wanna k’umnəgərukko# ʕarsɨχa mɨχʔal jɨʔu ʔɨti wanna-ʔ-u ʔɨk.. ʔɨti wanna k’umnəgər-u=kko def:3sgm main-ø-def rpr def:3sgm main meaningful-def:3sgm=foc ʕarsɨ-χa mɨχʔal jɨʔ-u self-2sgm help cop-3sgm ‘the main rpr.. the main thing is to be to able help yourself.’ you can see that there is no possible similarity between ʔɨk.. (the unfinished language unit in example (13) above) and the immediately following word of the same segment. the speaker had recognized k’umnəgəru ‘the main thing’ shall receive a focus that is marked by =kko that the speaker misplaced. it is for this reason that the speaker cuts off and re-utters as can be seen in example (13). self-repair in tigrinya: trouble sources, mechanisms and solutions 97 (14) asg1-7 [ʔɨziʔa] {ħa..} #ʔɨziʔa dɨbəynas# sələste s’ɨmdi tawʕɨlɨla malət jɨʔu ʔɨz-a-ʔ-a ħa.. ʔɨz-a-ʔ-a dɨ-bəyn-a-s sələste prx-3sf-ø-sng:3sf rpr prx-3sf-ø-sng:3sf for-alone-3sf-dm_foc three s’ɨmdi ta-wʕɨl-ʔɨl-a malət jɨʔ-u pair caus-spend the day-exist:nonpast-3sf means cop-3sm ‘this (pointing at a farming land) consumes three pairs.’ the speaker has uttered the unfinished word (ħa..) that later appeared in his speech as dɨbəynas ‘alone:f’ which is the second word after the repair is made. in this particular example, the speaker replaced the repair mechanism with a different constituent. (15) gdy1-102 ʔɨzaw mɨsza bɨmərfɨʔ gəβrɨna nəgga... nɨləgba malət jɨʔu ʔɨz-a-w mɨs-ʔɨz-a bɨmərfɨʔ gəβr-ɨna {nəgga..} prx-3sf-sng with-prx-3sf inst-needle use:ipr-1pl rpr #nɨ-ləgb-a# malət jɨʔ-u sj:3pl-fasten-obj:3sf means cop-3sm ‘it means we sew this (pointing at a cloth on his left hand) with this one (pointing at a cloth in front of him) using a needle.’ dma1-37 ʔɨh ʔɨh part ‘i am listening; please proceed’ gdy1-103 kaʔu bək’k’a tɨməlɨsχa tɨbl.. tɨχdəna kaʔu bək’k’a tɨ-məlɨs-χa {tɨbl..} #tɨ-χdən-a# then just ref-back-2sm rpr ref-wear:ipr-obj:3sf ‘then, you just wear it.’ the necessity of repairing may also be for language style or word diction as in example (15). the speaker prefers nɨləgba over nəga.. (nəgannanɨja) which are synonyms. for diction purposes, the speaker prefers to select a word over another word, but this came to the mind of the speaker after the less preferred word started to come out of the speaker’s mouth. the speaker repairs the original speech to put the more appropriate word in the sentence. 4.2.2 filled pauses filled pauses, also called fillers, are linguistic units that are “used instead of pauses or facial expressions” (asgede, 2019). in terms of function, fillers “shed light on the speech production process” and are “indicative of the mental processes underlying speech generation” (swerts, 1998). in spontaneous interactions, filled pauses signal many other functions, as the speaker is going to repair an original part of the oral discourse. though it isn’t easy to discriminate whether fillers serve different functions like planning process, hesitation, and repair (hlavac, 2011; brennan and williams, 1995; asgede, 2019), this is noticeable when we look at the repaired portion of the discourse. (16) gnt1-7 s’əʋa məχ’əmət’i ʔaχ’ɨħa tazəgad͡ʒɨw s’əba məχ’əmət’i ʔaχ’ɨħa t-a-zəgad͡ʒɨw milk vn-seat object ipv-caus:2sm-prepare ‘you prepare an object to put in milk.’ dma6-13 ʔɨh ʔɨh part ‘i am attending you attentively.’ asgede 98 gnt1-8 bɨmaj bɨdəmbi gəʋɨrχa tɨħas’ɨʋo bɨ-maj bɨdəmbi gəbɨr-χa tɨ-ħas’ɨb-o inst-water very well do:cn-sj:2sm ipv:sj:2sm-wash:imp-obj:3sm ‘you wash it very well with water.’ dma6-14 ʔɨ…ʃʃi ʔɨ…ʃʃi okay ‘that is interesting; would you please proceed to describe.’ gnt1-9 kaʔu [tɨʕat’ɨno] {ʔɨ…} #bɨʔawlɨʕ tɨʕat’ɨno# kaʔu tɨ-ʕat’ɨn-o ʔɨ… bɨ-ʔawlɨʕ tɨ-ʕat’ɨn-o then sj:2sm-steam:ipv-obj:3sm rpr inst-olive sj:2sm-steam:ipv-obj:3sm ‘then, you steam it with an olive.’ the repair in example (16) above aims at inserting a missing constituent. it is marked by a filler ʔɨ… followed by about two seconds’ pause. the process, however, does not just repair but has also a planning process. the speaker extends the pause because she wanted to buy time to recall the missed linguistic constituent bɨʔawlɨʕ ‘by olive (tree)’. 4.2.3 particles a particle refers to a linguistic unit comprising small words that in most cases are uninflected (asgede, 2019). the term is used as a collective term for the linguistic expressions like dəʔa ‘rather’, and ʔɨbba ‘(i) mean’. the particle ʔɨbba ‘i mean’, which has different functions in different contexts, initiates repair in the following example (17)ː (17) spnt1-64 [ħagos] {ʔɨbba} #ʃəgunu# məs’ɨʔudə ħagos ʔɨbba ʃəggunu məs’ɨʔ-u-də name rpr name come:prv-3sm ‘did hagos i mean shegunu come?’ in example (17), the particle ʔɨbba ‘i mean’ appeared between two proper nouns (hagos and shegunu). this particular particle functions as a replacive marker from the possible alternative names that the speaker knows and associates with another. such wrong constituents and their corrected constituents appear as part of the conversation disfluencies because the constituents have some close associations in the speaker’s mind. here in the above example, the two proper nouns name two brothers; thus, the speaker uttered the first one wrongly, which is repaired later. the repair initiation that is ʔɨbba ‘i mean’ is a particle. another particle that could have replaced the mentioned particle is dəʔa ‘rather’ which is also realized as dəʔam ‘rather’ in the corpus. the latter is used to initiate as a speaker repairs a portion of a prior utterance that asks for confirmation by oneself. for example, ħagos dəjɨʔu dɨməs’ɨlo ʃəgunu? ʃəgunu dəʔa ‘is that hagos or shegunu who is coming? it is rather shegunu’ is a typical example that depicts how dəʔa ‘rather’ is used to initiate a repair. the difference between these two particles is that the first one is used before the repaired part of an utterance but the latter one appears after the repaired part of an utterance. these instance functions of the two particles ensure that particles are used to initiate repair by tigrinya speakers. 4.2.4 predicates predicates in this paper are consumed as words and/or phrases that have the qualities that a verb has in tigrinya. for example, the verbal noun malətəj ‘i mean’, the verbal phrases ʔajχonχuj ‘i am not’, ʔɨjʔoj ‘no’ are examples of linguistic constituents consumed as predicates. a speaker may also reformulate part of her/his speech after uttering a whole word. unlike cut-offs, lexical constituents index repair, which implies that a speaker gets cognizant that she or he needs to mark a reformulation after a word is uttered. self-repair in tigrinya: trouble sources, mechanisms and solutions 99 (18) dm1-20 gɨdəj hɨndej ʕamətu jɨʔu bɨsruχə məʔazi jɨʔu dɨtwələdə gɨdəj hɨndej ʕamət-u jɨʔ-u bɨ-sr-u-χə məʔazi jɨʔ-u name how many year:poss:3sm cop-3sm abl-root-3sm-q when cop-3sm dɨ-t-wələd-ə rel-prv-birth:sj:3sm ‘how old is gidey? when was he born?’ ma1-61 [bɨzəmənə ʕɨnt͡ ʃ’ɨwa ʔɨkko jɨʔu dɨtwələdə] bɨ-zəmən-ə ʕɨnt͡ ʃ’ɨwa ʔɨkko jɨʔ-u dɨ-t-wələd-ə by-time-of rat foc cop-3sm rel-prv-birth:sj:3sm ‘it is on the rat’s era.’ ma1-62 {ʔɨjʔoj} #zəmənə ʕɨnt͡ ʃ’ɨwa ħalɨfus bɨχɨltə ʕamətu dəʔa məsələni ʔatta# ʔɨjʔoj zəmən-ə ʕɨnt͡ ʃ’ɨwa ħalɨfus bɨ-χɨltə ʕamətu dəʔa no time-of rat pass:prv-3sm-foc by-two year-poss:3sm foc məsəl-ə-ni ʔatt-a look:ipv-ob:3sm-sj:1s voc-2sm ‘no, i think it is actually two years after the rat’s era.’ the constituent ʔɨjʔoj ‘no’ initiates a repair. here the repairable is one complete sentence that is segmented as one discourse unit. it indicates that the previous segment is wrong and should be repaired. it indexes that the host segment is a repaired form of the trouble source. the linguistic constituent ʔɨjʔoj ‘no’ points back to the original segment and fore to the altered segment. it, therefore, indexes the host segment is repaired. repair is also made even after long trouble source constituents of a discourse are articulated as can be seen in example (18) above. based on the mini-corpus, ʔajχonχuj ‘no! i am not’ also has a similar function with ʔɨjʔoj ‘no’. this linguistic unit (ʔajχonχuj ‘no! i am not’) has a morpheme that marks an agreement (ʔaj-χon-χu-j neg-become:ipv-1s-neg). this phrasal constituent initiates that the prior segment was inappropriate and a revised replacement for it is to come. 4.2.5 pauses and non-verbal signals long pauses may also indicate that a speaker is about to repair a portion of an original problematic utterance. furthermore, nonverbal signals such as facial expressions and gestures, which usually accompany verbal repair initiators, can independently indicate that a repair is about to take place. nodding left and right, for example, may indicate that the speaker recognizes his original speech is problematic and intends to replace it. furthermore, twinkling one’s eyes or closing both eyes may indicate that the speaker is about to repair a part of a problematic original utterance. 4.3 results of self-repair as highlighted in the introduction, self-repair refers to self-correcting an error or inappropriate expression of an original part of an utterance. this conversational feature allows to expand an idea, replacing a constituent or an entire turn, reorder constituents, restart a turn or part of a turn, and insert a forgotten but relevant constituent. repair results total number of cases identified percentage expanding 92 23 replacing 91 22.75 re-ordering 82 20.5 restarting 70 17.5 inserting 65 16.25 total 400 100% table 5. list of self-repair results. asgede 100 e x p an d in g r ep la ci n g r eo rd er in g r es ta rt in g in se rt in g s u m sum 328 363 331 335 243 1600 p 0.205 0.227 0.207 0.209 0.152 pe 0.203 po 0.733 k 0.665 table 6. inter-rater agreement on the results of self-repairs. as depicted in table 6, the inter-rater agreement is 0.665 which is interpreted as substantially relevant agreement according to fleiss’ kappa (see table 10 in appendix iii). after conducting a detailed investigation, the author found that expanding (92 cases) and replacing (91 cases) are the most frequently observed results of self-repair in tigrinya. the description of each self-repair result will be presented in the following pages. 4.3.1 expanding by expanding, a speaker seeks to clarify or reformulate how a concept was communicated in a previous turn. such justifications or rephrasing are added by speakers to facilitate smooth contact with their correspondents. for instance, an explanation might be significant to clarify concepts, and rephrasing might be relevant to present ideas in a new way so that the listener has other options for comprehending the argument made by the utterances. malət ‘means’, is a very frequently used lexeme in tigrinya to convey a further elaboration. in real speech, malət, ‘means’, is most often realized as malətəj ‘i mean’. (19) bly1-21 ʃɨʋa dɨʔɑχonuj ʃɨba dɨʔɑ-χon-u-j paralize foc-become:ipv-sj:3plm-foc ‘they rather are unskilled (lit.they rather are paralyzed.). bly1-22 malətəjssi suχ’lom mɨt’rom ʔaħɨʋt’om haməj wulom kɨsərɨħu mɨwdɨldal dəʔam malət-əj-ssi suχ’-bɨl-om mɨt’ri-om ʔa-ħɨbt’-om means-poss:1s-foc quiet-say:ipv-sj:3plm buttocks caus-fat:prv-sj:3plm haməj wul-om kɨ-sərɨħ-u mɨ-wdɨldal dəʔam how say:cn-3plm ipv:3-work-sj:3plm vn-idle foc ‘i mean that they can’t do a job except wondering here and there with their big physical appearance.’ the lexeme malətəj ‘i mean’ in example (19) above indicates to the listener that the speaker is about to reformulate the way an idea was described. an explanation and justification of the prior discourse unit’s claim is presented in the turn that hosts the reformulation initiator lexeme. the ‘students’ that are the issue of the discussion, are ineffective in assisting their parents because the students lack the skill and gut and that is described as ‘they are paralyzed’. the speaker goes on to explain that the students are weak and prefer to wander here and there rather than be engaged in some livelihood activities on which their parents rely. in the repairing segment, the speaker explains why he described them as “paralyzed” in the repairable segment. he means that, in contrast to their chiseled bodies, they are unqualified and uninterested in assisting their parents in any way. the verbal noun malətəj ‘i mean’ appears at the initial position of the repaired segment as depicted in example (19) above. it also appears to be a final position of the repaired constituent but before a copular verb as in suχ’lom mɨt’rom ʔaħɨʋt’om haməj wulom kɨsərɨħu mɨwdɨldal self-repair in tigrinya: trouble sources, mechanisms and solutions 101 dəʔam malətəj jɨʔu ‘i mean that they can’t do a job except wondering here and there with their big physical appearance.’ 4.3.2 replacement the second type of repair result is a replacement which refers to replacing a wrongly uttered constituent with a correct one. this is initiated by different linguistic units. predicative lexemes such as ʔajχonχuj ‘no, i am not’ and particles such as ʔɨbba ‘i mean’ are the most common lexical devices that initiate a replacement. (20) dm1-96 ʔajjajjə ʔajəʕawuji məʔazi jɨʔu dɨmotə ʔajja-jjə ʔajə-ʕawuji məʔazi jɨ-ʔu dɨ-mot-ə father-rl:1s father-big when cop-3sm rel-death:prv-sj:3sm ‘father, when did grandpa die?’ ma1-221 ʕasərtə ʕamətu χojnuwwo ʔasərtə ʕamət-u χojn-u-ww-o ten year-poss.3sm become:ipv-sj.3sm-obj.3sm ‘ten years have passed.’ (lit. it has been ten years.’ ma1-222 ʔajχonχuj ʕasərtə ħamuʃtə ʕamətu dəʔatta ʔaj-χon-χu-j ʕasərtə ħamuʃtə ʕamət-u dəʔa-tta neg-become:ipv-1s-neg ten five year-poss.3sm dm_foc-voc.2sm ‘it was rather fifteen years ago.’ as can be seen in example (20) above, ʔajχonχuj ‘no, i am not’ indicates that the following utterance is going to replace a prior portion of discourse. so, it is not 10 years since the grandpa is died but 15 years. (21) spnt1-127 ħagaj ħagaj ʔɨbba χɨrəmti χɨrəmti ʔabəj laħləfχajjo jɨʔu hɨndizi dɨt’əfaʔχa ħagaj ħagaj ʔɨbba χɨrəmti χɨrəmti ʔabəj winter winter rpr summer summer where la-ħləf-χa-jj-o jɨʔ-u hɨndi-zi dɨ-t’əfaʔ-χa prog-spend:ipv-sj.2sm-ø-obj.3sm cop-3sm much-prx.3sm rel-lost:prv-sj.2sm ‘where have you been spending every winter i mean every summer that you have been missing for so long? .’ the particle ʔɨbba ‘rather, i mean’ in this particular context is another replacive marker. it appears between the problematic portion of the prior discourse unit and its corrected constituents. as can be seen in example (21) above, ʔɨbba ‘rather, i mean’ appeared between ħagaj ħagaj ‘every winter’ and χɨrəmti χɨrəmti ‘every summer’ so that the latter one is the constituent that replaces the former ħagaj ħagaj ‘every winter’, that is, the repairable one. 4.3.3 insertion insertion refers to adding a constituent which was missed in the repairable portion of an utterance. the inserted constituent can be a word, phrase, clause and/ or sentence. it may be initiated by cut-offs, fillers, particles, or predicative words. (22) gnt1-10 kaʔu laχadɨnχa guħatɨj mɨʃətɨj s’əʋa tawahlɨləllu kaʔu la-χadɨn-χa guħat-ɨj mɨʃət-ɨj s’əba then prog-close:cn-3sm morning-and evening-and milk ta-wahlɨl-ə-ll-u caus-store-sj.2sm-ben-obj.3sm ‘every morning and evening, you store (milk) in it.’ asgede 102 gnt1-11 bɨdawahləlɨχa tɨħak’uno s’ɨnaħ medʒəmerɨja tarɨgʔo dəʔa kaʔu tɨħak’uno bɨd-a-wahləl-ɨχa tɨ-ħak’un-o s’ɨnaħ medʒəmerɨja after-caus-store-sj.3sm ipv-churn:sj.2sm-obj.3sm wait:sj.2sm first ta-rɨgʔ-o dəʔa kaʔu tɨ-ħak’un-o ipv-caus-coagulate:sj.2sm-obj.3sm rather then ipv-churn:sj.2sm-obj:3sm ‘after you store it, you churn it; wait, first you let it coagulate rather then churn it.’ the particle dəʔa ‘rather’ that functions, in most cases, to selectively emphasize a constituent, appears next to a linguistic constituent that replaces a problematic portion of a prior discourse. in the above excerpt, the speaker uttered s’ɨnaħ ‘wait’ which directly tells us something went wrong with the prior discourse; it, however, doesn’t directly index a repair. next to that, the speaker added a new step medʒəmerɨja tarɨgʔo ‘first you let it coagulate’ that should precede a step that was stated in the prior constituent of the turn tɨħak’uno ‘you churn it’. as you can see in the excerpt, the repaired segment is followed by the selective repair indicator. syntactically, this particle is different from the particle ʔɨbba ‘rather, i mean’. the former appears next to the repairing segment whereas the latter appears between the repairable and the repairing segment. 4.3.4 restart restart, at least in this paper, refers to a conversational strategy by which a speaker stops/aborts speaking and then restarts either with a modified version of the prior portion of discourse or with a new idea that is different from the trouble source. (23) bly1-001 suməj gɨrmaj ħagos ʔə… ʔɨʃʃi suməj gɨrmaj ħagos ħagos ħuluf ħuluf ʔaləm ʔaləm s’adɨχ’ jɨʔu sum-əj gɨrmaj ħagos ʔə… ʔɨʃʃi suməj gɨrmaj ħagos name-poss:1s name name rpr okay name-poss:1s name name ħagos ħuluf ħuluf ʔaləm ʔaləm s’adɨχ’ jɨʔ-u name name name name name name cop-3sm ‘my name girmay hagos ʔə… okay my name is girmay hagos huluf alem tsadik.’ as can be seen in example (23) above, the speaker aborts speaking about himself and restarts after he has taken his time to reorganize his speech. the speaker gets stuck into uttering his grandpa’s name and thus stops speaking. the speaker then restarts his talk with the same content as the trouble source segment. this was initiated by the filler ʔə which is followed by the stretched vowel sound that is marked by three dots (…). that again is followed by the particle ʔɨʃʃi ‘okay’ that can be considered as a false starter. the speaker reorganizes his idea and restarts. however, as we can see from example (24) below, the speaker aborts speaking in a plan to come up with a different point of discussion. the speaker stops talking about the previous topic and shifts the point of discussion to another. such a repair is indicated by ħɨdəgo ʔɨsti ‘please leave it’ which states that the speaker is no more interested in talking about the issue. that again closed with the enclitic =wa that is attached to the last constituent of the repairable segment. the purpose of such repair to restart is to change a point of discussion – here from talking about the challenges the speaker is facing to asking about the health of the correspondent’s family. (24) bly1-111 ʔɨttom ʔaħwatəj bɨgɨrat tɨs’alʔom ʔajnəgagəruj ʔɨtt-om ʔa-ħaw-at-əj bɨ-gɨrat tɨ-s’alʔ-om dst-3plm pl-brother-pl-rl:1s inst-farmland prv-oppose-sj:3plm ʔaj-nəgagər-u-j neg-speak:ref-3plm-neg ‘my brothers do not speak to each other because they are in disagreement over a farmland.’ self-repair in tigrinya: trouble sources, mechanisms and solutions 103 bly1-112 ʔaddəjjəllə mɨs ʔakkojjə bɨnaddɨʔom wursi tɨs’alʔom fɨrdi ʋət məlaləsləwu ʔaddo-jj-ə-llə mɨs ʔakko-jjə bɨ-na-addo-ʔ-om wursi mother-ø-rl:1s-and with uncle-rl:1s inst-poss-mother-ø-3plm heirloom tɨ-s’alʔi-om fɨrdi bət məlaləs-ʔɨllə-w-u rec-quarrel-3plm justice house ambulation-existːipv-sj:3plm ‘my mother is currently in a quarrel with my uncle, and they are going to court.’ bly1-113 ħɨdəgo ʔɨsti ʔanə ʔagnɨjəjjo dɨnhəχu gudwa… ʕɨjjal haməj ʔɨnhəwu ħɨdəg-o ʔɨsti ʔanə ʔa-gnɨj-ə-jj-o leave.imp:sj.2sm-obj.3sm part 1s caus-find:cn-sbj.1s-obj.3sm dɨ-ʔɨnh-ə-χu gud=wa… rel-exist:prv-obj.3sm-sj.1s bad thing-rpr ‘please leave the people i am dealing with…, how are all your family members?’ besides, a speaker repairs a portion of a prior segment to restart an utterance to shift from sticking to one detail to another without leaving the major point of discussion. the last turn of example (25) below has a particle wa which can be roughly translated as ‘whatever’, and is used to initiate a shift from talking about the weakness of the students to a possible alternative that they may excel at, that is, being clever students. (25) bly1-21 ʃɨba dɨʔɑχonuj ʃɨba dɨʔɑ-χon-u-j paralize foc-become:ipv-sj:3plm-foc ‘they are rather unskilled’ (lit.they are are paralyzed.). bly1-22 malətəjssi suχ’lom mɨt’rom ʔaħɨʋt’om haməj wulom kɨsərɨħu mɨwdɨldal dəʔam malət-əj-ssi suχ’-bɨl-om mɨt’ri-om ʔa-ħɨbt’-om means-poss:1s-foc quiet-say:ipv-sj:3plm buttocks caus-fat:prv-sj:3plm haməj wul-om kɨ-sərɨħ-u mɨ-wdɨldal dəʔa-m how say:cn-3plm ipv:3-work-sj:3plm vn-idle foc-foc ‘this (pointing at a farming field) requires three pairs of oxen to plow.’ bly1-23 ʔazatomu haməj… wa… ʔatta tɨmhɨrtom hamgobəzullə t͡ ʃ’əllə neʋɨru ʔaz-atomu haməj… wa… ʔatt-a tɨmhɨrti-om prx.3plm how rpr voc-2sm education-poss:3plm ham-gobəz-u-llə t͡ ʃ’əllə nebɨr-u if-excel:ipv-sj.3plm-and okay exist:prv-3sm ‘how could they …whatever… it could have been fine if they had excelled in their education.’ in example (26) below, the speaker restarts the turn at the fifth constituent of the turn. the repair operation results in deleting or leaving out some non-relevant portion of the trouble source. thus, the first four constituent words of the trouble source, i.e., ʔɨndɨr haftam tɨχon dalχamma ‘see, if you want to be rich’ are replaced by haftam dɨmuχʷan ‘to be rich’ by a restarting mechanism; while two constituents are deleted, one is modified and another constituent is maintained as it originally was in the repairing segment. (26) ma2-78 ʔɨndɨħɨr haftam tɨχon dalχamma haftam dɨmuχʷan hɨras dəjmɨbzaħ jɨʔu ʔɨndɨħɨr haftam tɨχon dalɨj-χa-mma haftam dɨ-muχʷan if rich ipv:2-become want:ipv-sj:2sm-foc rich pur-become hɨras dəjmɨbzaħ jɨʔ-u sleep neg-multiple cop-3sm ‘if you really want to be rich, to be rich, you can not sleep too much.’ 4.3.5 reordering constituents might appear in a portion of a prior utterance as a source of trouble that results in reordering them. as can be seen in example (10), a sentence or long constituents may take the wrong order in speech. asgede 104 the most common wrong order found in the data is short constituents like words as can be seen in example (27) below. (27) ma1-143 ʔɨsumma bəlaj ʔadanə ʔɨbba ʔadanə bəlaj jɨʔu sɨmu ʔɨsu-mma bəlaj ʔadanə ʔɨbba ʔadanə bəlaj jɨʔ-u sɨm-u 2sm-foc name name mean name name cop-3sm name-poss:3sm ‘that one’s name is belayadane i mean adene belay.’ here in this example, the trouble source is the wrong order of the first name and middle. the self-repair initiated by the particle ʔɨbba ‘(i) mean’ results in reordering the wrongly ordered nouns in the repairable portion of the utterance. 5 discussion and recommendation like many language speakers in the world (fox et al., 2013), tigrinya speakers often repair portions of their prior utterances that appear to be ‘inappropriate’. the inappropriateness of a portion of an utterance may or not have an association with ‘error’. through data analysis, 400 repairs have been identified. these identified repairs resulted from various sources of trouble. the almost perfect fleiss’ kappa inter-rater agreement (0.835) indicates that the most common sources of repair are pronunciation (0.218), diction (0.173), missing constituents (0.159), wrong constituents (0.125), word order issues (0.124), misunderstandings (0.106), and focus misplacement (0.088). it was observed that the pronunciation and diction issues are related with dialect differences. the paper highlights that more than half of the diction problem sources for self-repair is not related to dialectical variation. these sources of repair, however, need a further study on statistically showing to what degree dialect variation impacts pronunciation and diction issues in speech. pronunciation-related problems resulted from tongue slip and the dialect variation. the data was recorded in the presence of the researcher, who had been away from the area for about 15 years. it seems that the participants (wrongly) perceived that the researcher preferred to hear the lexemes as they are pronounced in the quasi-standard tigrinya used in education, media, and offices at various levels; six of the recorded participants explained this while we held informal discussions about what they repair and replace words for. a participant might repair lajəs’a to nab jəs’a ‘to yetsa’ and tɨχəd to tɨχəjɨd ‘you go’. how such issues work in the conversations of speakers of other varieties of the language (and other languages spoken in ethiopia too) is left for further research. in addition to the dialectical variation, triggered repair to replace a lexical item, interpersonal relationships among participants is another trouble source. for example, the constituent ħɨləf ‘(you:2sm) pass through’ is replaced by ħɨləfu ‘(you:2sm.h) pass through’ for the participants’ social status variation; tɨħarɨʔ is replaced by tɨftɨn ‘urinate’ for diction purposes that have implications on the participants’ dialect variation background. when speakers speak out constituents such as improper nouns in a specific context, they use the particle ʔɨbba between the repairable and the repaired which indexes repair. this uncovers the fact that tigrinya honorific grammatical feature has impact on how interlocutors interact with one another. different linguistic devices initiate self-repair. cut-offs that do not directly index self-repair as well as filled pauses which have multiple functions in conversation are used. cut-offs that are the dominant selfrepairs (0.593) have other functions: planning process being the most common one (asgede, 2019). there is a need to study how cut-offs are distributed and used in speech independently. predicative phrases like ʔajχonχuj ‘i am not’ and ʔɨjʔoj ‘no’ as well as particles such as ʔɨbba ‘i mean’ and dəʔa ‘rather’ are used to initiate self-repair. as stated in different empirical studies in different languages (ginzburg et al., 2007), self-repair also is done without any repair marker; initiators are therefore not mandatory in tigrinya too. that seems the reason why pauses and nonverbal cues that speakers may leave out of their utterance can initiate self-repairs. these findings match with what shriberg (1994) found out, that lexical items and disfluencies such as filled pauses initiate self-repair. this also implies that self-repair can be ‘explicitly’ marked by linguistic and non-linguistic constituents as mentioned by meteer et al. (1995). self-repair in tigrinya: trouble sources, mechanisms and solutions 105 different types of self-repair result in different operations. for example, self-repair may be used to expand an idea presented in a previous section of discourse, which is most commonly initiated by the verbal noun malət ‘means’. self-repair may also result in replacing a wrongly uttered constituent, initiated by predicative phrases such as ʔajχonχuj ‘i am not’ and ʔɨjʔoj ‘no’ and the particle ʔɨbba ‘i mean’. a particle dəʔa ‘rather’ initiates the insertion of a missing constituent. finally, though it would have been possible to study the phonetics features of every repair initiator and its categorizations, how self-repair relates with the syntactic features of tigrinya, repair in the process of language acquisition and learning, thorough discussion on the relationship between self-repair and the demographic features of participants, etc., they are left aside and the author recommends such issues for further investigation. references dagnew mache asgede (2007). tigrinya dialectal variations between rayya and adwa. arba-minch university: senior essay. dagnew mache asgede (2019). discourse markers in rayya tigrinya: documentation and linguistic analysis. dissertation. addis ababa university. http://etd.aau.edu.et/handle/123456789/24899 dagnew mache asgede (2023). backchannels in tigrinya. ethiopian journal of business and social science, 6(1), pp. 15-35. https://doi.org/10.59122/144f59lk david m. bell (2010). nevertheless, still, and yet: concessive cancellative discourse markers. journal of pragmatics, 42, 1912-1927. emily m. bender and batya friedman (2018). data statements for natural language processing: toward mitigating system bias and enabling better science. transactions of the association for computational linguistics, 6:587–604. elisabeth r. blackmer and janet l. mitton (1991). theories of monitoring and timing of repairs in spontaneous speech. cognition, 39:173–194. hatte anne blejer (1986). discourse markers in early semitic, and their reanalysis in subsequent dialects. dissertation. the university of texas at austin. susan brennan and maurice williams (1995). the feeling of another’s knowing: prosody and filled pauses as cues to listeners about the metacognitive states of speakers. journal of memory and language, 34:383–398. susan brennan and michael schober (2001). how listeners compensate for disfluencies in spontaneous speech. journal of memory and language, 44:274–296. james deese (1984). thoughts into speech: the psychology of a language. prentice-hall, englewood cliffs, nj. vincent a. de rooji (2000). french discourse markers in shaba swahili conversations. the international journal of bilingualism, 4.4:447-467. joseph l. fleiss (1971). measuring nominal scale agreement among many raters. psychological bulletin, 76:378-382. http://etd.aau.edu.et/handle/123456789/24899 https://doi.org/10.59122/144f59lk asgede 106 joseph l. fleiss, bruce levin, and myunghee c. paik (2003). statistical methods for rates and proportions (3rd ed.). hoboken, nj: wiley. barbara a. fox, makoto hayashi, and robert jasperson (1996). resources and repair: a cross-linguistic study of syntax and repair. in e. ochs, e. a. schegloff, & s. a. thompson (eds.), interaction and grammar (pp. 185–237) chapter, cambridge: cambridge university press. https://doi.org/10.1017/cbo9780511620874.004 barbara a. fox, trevor benjamin, and harrie mazeland (2013). conversational analysis and repair organization: overview. in c. chapelle (ed.), the encyclopedia of applied linguistics. https://onlinelibrary.wiley.com/doi/10.1002/9781405198431.wbeal1314 jonathan ginzburg, raquel fernández, and david schlangen (2007). unifying self-and other-repair. in decalog 2007: proceedings of the eleventh workshop on the semantics and pragmatics of dialogue, may 30-june 1 2007, rovereto, italy. abraham girmay (2012). language shift in progress in the raya variety of tigrinya: the case of alamata and kobo weredas. ma thesis. addis ababa: addis ababa university. kieu-phuong ha and martine grice (2017). tone and intonation in discourse management-how do speakers of standard vietnamese initiate a repair? journal of pragmatics, 107:60-83. patrick g. healey, marcus colman, and mike thirlwell (2005). analysing multimodal communication: repair-based measures of human communicative coordination. advances in natural multimodal dialogue systems, 113-129. peter a. heeman and james f. allen (1999). speech repairs, intonational phrases, and discourse markers: modeling speakers’ utterance in spoken dialogue. association for computational linguistics, 25.4, 528-571. jim hlavac (2011). hesitation and monitoring phenomena in bilingual speech: a consequence of codeswitching or a strategy to facilitate its incorporation? journal of pragmatics, 43:3793–3806. https://doi.org/10.1016/j.pragma.2011.09.008 julian hough (2015). modelling incremental self-repair processing in dialogue. unpublished phd thesis. queen mary university of london. http://qmro.qmul.ac.uk/xmlui/handle/123456789/9094 farhat jabeen, asim m. asim rai, and sara arif (2011). a corpus-based study of discourse markers in british and pakistani speech. in international journal of language studies (ijls), 5(4):69-86. masahito kawamori, takeshi kawabata, and akira shimazu (1998). discourse markers in spontaneous dialogue: a corpus-based study of japanese and english. in manfred stede, leo wanner and eduard hovy (eds.), discourse relations and discourse markers: proceedings of the workshop, 93-99. canada: montreal: universite de montreal. tsehaye kiros (2009). a comparison of wajerat tigrinya vs. standard tigrinya. ma thesis: addis ababa: addis ababa university. willem j. levelt (1983). monitoring and self-repair in speech. cognition, 14(4):41–104. https://doi.org/10.1017/cbo9780511620874.004 https://onlinelibrary.wiley.com/doi/10.1002/9781405198431.wbeal1314 https://doi.org/10.1016/j.pragma.2011.09.008 http://qmro.qmul.ac.uk/xmlui/handle/123456789/9094 self-repair in tigrinya: trouble sources, mechanisms and solutions 107 yaron matras (2000). fusion and the cognitive basis for the bilingual discourse markers. the international journal of bilingualism, 4.4, 505-528. douglas w. maynard (2013). everyone and no one to turn to: intellectual roots and contexts for conversation analysis. in jack sidnell and tanya stivers (eds.), the handbook of conversation analysis, 11-31. west sussex: wiley-blackwell. niguss weldezgu mehari (2011). the impacts of rayya dialectal variations and the influence of amharic on medium of instruction: the case of alamata primary schools. addis ababa university.ma thesis. http://etd.aau.edu.et/handle/123456789/16280 niguss weldezgu mehari. (2021). a grammar of rayya tigrinya. addis ababa university. dissertation. marie w. meteer, ann a. taylor, r. macintyre, and r. iyer (1995). dysfluency annotation stylebook for the switchboard corpus. philadelphia, pa: university of pennsylvania. lorenza mondanda (2013). the conversation analytic approach to data collection. in jack sidnell and tanya stivers (eds.), the handbook of conversation analysis, 32-56. west sussex: wiley-blackwell. andrew murphy (2019). resolving conflicts with violable constraints: on the cross-modular parallelism of repairs. glossa: a journal of general linguistics, 4(1): 9. 1–39. https://doi.org/10.5334/gjgl.608 christine h. nakatani and julia hirschberg (1994). a corpus-based study of repair cues in spontaneous. speech. journal of the acoustical society of america, 95(3), pp. 1603-1616. https://doi.org/10.1121/1.408547 gessesse nigusse (2012). language use in the criminal justice process: the case of raya alamata woreda in tigray region. ma thesis. addis ababa: addis ababa university. emanuel a. schegloff (1992). repair after next turn: the last structurally provided defense of intersubjectivity in conversation. american journal of sociology, v-97: pp. 1295–1345. https://www.journals.uchicago.edu/doi/10.1086/229903 emanuel schegloff, gail jefferson, and harvey sacks (1977). the preference for self-correction of repair in conversation. language, 53(2):361-382. https://doi.org/10.2307/413107 lawrence schourup (1985). common discourse particles in english conversation: ‘like’, ‘well’, ‘y’know’. new york: garland. elizabeth e. shriberg (1994). preliminaries to a theory of speech disfluencies. ph.d. thesis, university of california at berkeley, berkeley, usa. tanya stivers and jack sidnell (2013). introduction. in jack sidnell and tanya stivers (eds.), the handbook of conversation analysis, 1-8. west sussex: wiley-blackwell. marc swerts (1998). filled pauses as markers of discourse structure. journal of pragmatics, 30(4):485496. https://doi.org/10.1016/s0378-2166(98)00014-9 tsehaye teferra (1979). reference grammar of tigrinya. doctoral dissertation. washington, d.c.: georgetown university. http://etd.aau.edu.et/handle/123456789/16280 https://doi.org/10.5334/gjgl.608 https://psycnet.apa.org/doi/10.1121/1.408547 https://www.journals.uchicago.edu/doi/10.1086/229903 https://doi.org/10.2307/413107 https://psycnet.apa.org/doi/10.1016/s0378-2166(98)00014-9 asgede 108 yan wang (2011). a discourse-pragmatic functional study of the discourse markers japanese ano and chinese nage. intercultural communication studies, 20.2, 41-61. nigel ward (1998). some exotic discourse markers of spoken dialog. in manfred stede, leo wanner and eduard hovy (eds.), discourse relations and discourse markers: proceedings of the workshop, 6264. canada: montreal: universite de montreal. jean wong and hansun z. waring (2010). conversation analysis and second language pedagogy: a practical guide to esl/efl teachers. routledge. new york and london. robin wooffitt (2005). conversation analysis and discourse analysis: a comparative & critical introduction. sage publications. london, thousand oaks & new delhi. erkan yilmaz (2004). a pragmatic analysis of turkish discourse particles: yani, i̇şte and şey. doctoral dissertation: middle east technical university. tesfay tewolde yohannes (2002). a modern grammar of tigrinya. rome, tipografia u.detti-via g.savonarola. appendix i: transcription conventions and symbols the data was transcribed and annotated phonetically; the detailed annotation that includes gloss for every morph and free translation is helpful for the readers who do not speak and read tigrinya. for readability, the transcription conventions used in the excerpts presented in this paper are the following: 1 first person 2 second person 3 third person all allative ben beneficiary caus causative cn converb cnvrpr conversational process cop copular def definiteness dest distal dm discourse marker f feminine foc focus h honorific imp imperative inst instrument ipv imperfective juss jussive lit. literal meaning loc locative m masculine neg negative ø zero meaning obj object part particle pl plural plnprc planning process plt polite poss possessive prog progressive prx proximal pur purpose pvr perfective q question marker ref reflexive rel relativized rl relative (affection) rpr repair s singular sj subject vn verbal noun voc vocative name a gloss for proper nouns / overlap begins .. cut-off = clitic {} repair mechanism morph and gloss boundary self-repair in tigrinya: trouble sources, mechanisms and solutions 109 : boundary for grammatical functions by one morph … long pause [] trouble source # # repair solution \ another function appendix ii: data statement during a three-month fieldwork during which the researcher recorded the speech data, the mini-corpus temporarily named as rayya tigrinya mini-speech corpus has been established. originally, the mini-corpus totalled 24 hours of recorded audio/video data. for research ethics purposes, based on the basic schema recommended by bender and friedman (2018), the data statement is detailed as follows. curation rationale: with the help of three research assistants (two with a ba degree in english language and literature and one with a ba degree in tigrinya) and the consent of the participants, the researcher selected six hours of oral data. then the selected six hours of oral data was segmented and transcribed by the three individuals mentioned above and by the researcher. later on, the researcher edited and made some corrections on the transcription and made sample tagging of repairs. the article is entirely based on audio and video-recorded oral data. the data is made up of casual conversations, informal interviews, and stories. spontaneous conversations, informal interviews that were recorded on a variety of topics including how to do things, biographies, tales, explanations of some cultural performances such as shadey, shiwulalo, gebiche, and gebeta, farming, family relationships, and how to culturally mediate conflicts were selected; these accounts lasted approximately one hour and forty-two minutes. ten spontaneous conversations that were recorded at a market, a conflict resolution site, a bus station, a workplace, and a hotel are selected: the longest interactional conversation was twenty minutes long, while the shortest was four minutes and twenty-two seconds long. types of oral data number of recordings length in minutes total number of words percent number of repairing repairing in percentage remark spontaneous conversation 10 61 12,600 17.074 156 39 2 up to 4 participants conflict resolution 2 62 14,495 16.931 132 33 6 & 9 participants informal interview 20 59 12,497 16.935 51 12.75 2 participants each how to do things 8 58 11,650 15.786 30 7.5 2 participants each biography 12 60 12,504 16.943 17 4.25 1 participant each tales 18 62 12, 050 16.328 14 3.5 2 participants each total 362 75,796 100 400 100% table 7. total number of recordings for each genre and number of self-repairs. according to the data, 400 repairs have been identified; repairing is very common in spontaneous conversations, followed by conflict resolution, which account for 156 and 132 instances, respectively. only 14 repairs were identified in the narrating tales from the recorded genre of data sources. the above table shows the distributions of the repairs in detail. asgede 110 language variety: the linguistic data that this article relies on is collected from the residents of neksege, a locality located to the west of maichew, southern tigray administrative zone, and is used in the analysis. these people speak a tigrinya variety that is not being represented enough in education, media, and offices (asgede, 2007 and 2019; mehari, 2011). speakers’ demographic features: data was recorded from 38 participants. table 8, summarizes the demographic features of the participants whose speech was recorded for the purpose of this research. all the speakers gave their consent before each recording started and ended. number remark number remark 1. age 20-30 31-40 41-50 50-60 >60 2 5 6 10 15 4. religion orthodox protestant muslim 36 1 1 originally from qobo originally from maichew 2. gender female male 14 24 5. ethnicity tigrean agew amhara 35 2 1 originally from ts’eta originally from qobo 3. educational background primary (0-4) primary secondary (5-8) highschool (9-12) college 15 9 8 6 6. language background native in tigrinya native in amharic native in agew 36 1 1 8 of them speak agew; other 5 speak english speaks tigrinya & english speaks tigrinya total number of participants 38 table 8. demographic features of participants. annotators’ demographic: as highlighted above, including the author of the current research article, the annotators are four in number. all the annotators are members of the speech community that the data is collected from. four of them are native tigrinya speakers. while the researcher speaks amharic like a native, the others speak it with phonological, lexical and syntactic difficulties. they range in age from 35-48 and include one woman and three men. speech situations and speech characteristics: as introduced in the previous pages, the data is recorded in different social settings. except the data from conflict resolution, the remaining data was collected in informal settings and contains comics and several nonlinguistic features of daily interaction. the intended audience was therefore the participants who overtly showed up themselves in every dialogue and interview. the contents of the data are related to the living system of the community. recording equipment: the recording device used was an ic recorder, and some videos of conflict resolutions were video recorded using an apple ipad. appendix iii: inter-annotator agreement based on an example text that contains 15 repairs, the researcher trained the annotators and raters for half a day. the content of the training was how to identify the repair typologies and how to annotate (transcribe) them. based on healey et al. (2005), a short guideline that encompasses the following points is given to the annotators. self-repair in tigrinya: trouble sources, mechanisms and solutions 111 a repair is identified as a self-repair, if the response to the following questions is ‘yes’ except for the fourth point:  does the speaker of the current turn modify the original text one way or another by him/herself?  is the speaker repairing his/her own original constituent/utterance?  are you sure the current listener does not make any contribution to the request for a repair?  is not the repair/revision completed by another person?  does the speaker of the current turn edit, amend, or reprise the original text? to identify the trouble sources of self-repair, the raters should try to:  identify why speakers repair the portion of their old speech. to identify the types of the repair mechanisms, consider the following points:  note whether the repair marker (repair mechanism) is cut-off, particle, filled pause, noun, verb, pause, verb phrase, noun phrase, clause, facial expression, etc.  note whether the purpose (result) of each identified self-repair is to expand, restart, insert, replace, reordering, etc. of the original constituent. to identify the result of repairs, the raters should try to:  understand how the repair result is different from the repaired one,  explain the contribution of the repair result on the discourse development,  assign what the repair mechanisms entail on the repair result. accordingly, the number of self-repairs identified from each genre by the annotators are shown in table 9. this table only demonstrates the mean of the four annotators in comparison with the number of self-repair markers identified and used in this research. genres number of self-repair markers identified by annotators a1 a2 a3 a4 mean repair markers considered spontaneous conversation 160 145 144 156 151 156 conflict resolution 137 127 121 132 129 132 interview 53 45 47 51 49 51 how to do things 34 30 27 30 30 30 biography 18 16 15 17 17 17 tales 18 12 10 14 14 14 total 420 375 364 400 390 400 table 9. aannotators’ agreement on identifying repair markers. considering the total number of repairs identified by the researcher and the remaining three annotators, the mean is 390. therefore, the 400 repairs are accepted as valued and reliable. this is substantiated with the inter-rater agreement among the four raters (including the author) for the categories of self-repairs: trouble sources, repair mechanisms, and results of self-repairs under sections 4.1., 4.2., and 4.3., respectively. the fleiss’ kappa model is employed to manually compute the inter-rater agreement for there are four raters (more than two raters) and the data is nominal (fleiss 1971; 2003). 𝐾 = po−pe 1−pe where k is fleiss’ kappa; po is observed agreement, pe is expected agreement among the raters the inter-rater agreement for the sources of troubles in self-repair, repair mechanisms, and results of selfrepairs is calculated in the analysis section using the fleiss’ kappa interpretation intervals provided below in table 10. asgede 112 fleiss’ kappa interpretation <0.00 poor agreement 0.00 to 0.20 slight agreement 0.21 to 0.40 fair agreement 0.41 to 0.60 moderate agreement 0.61 to 0.80 substantial agreement 0.81 to 1.00 almost perfect table 10. interpretation of fleiss’ kappa values. dialogue & discourse 8(2) (2017) 21–55 doi: 10.5087/dad.2017.202 dialog structure through the lens of gender, gender environment, and power vinodkumar prabhakaran vinod@cs.stanford.edu stanford university stanford, ca owen rambow rambow@ccls.columbia.edu columbia university new york, ny editor: raquel fernández submitted 11/2016; accepted 04/2017; published online 05/2017 abstract understanding how the social context of an interaction affects our dialog behavior is of great interest to social scientists who study human behavior, as well as to computer scientists who build automatic methods to infer those social contexts. in this paper, we study the interaction of power, gender, and dialog behavior in organizational interactions. in order to perform this study, we first construct the gender identified enron corpus of emails, in which we semi-automatically assign the gender of around 23,000 individuals who authored around 97,000 email messages in the enron corpus. this corpus, which is made freely available, is orders of magnitude larger than previously existing gender identified corpora in the email domain. next, we use this corpus to perform a largescale data-oriented study of the interplay of gender and manifestations of power. we argue that, in addition to one’s own gender, the “gender environment” of an interaction, i.e., the gender makeup of one’s interlocutors, also affects the way power is manifested in dialog. we focus especially on manifestations of power in the dialog structure — both, in a shallow sense that disregards the textual content of messages (e.g., how often do the participants contribute, how often do they get replies etc.), as well as the structure that is expressed within the textual content (e.g., who issues requests and how are they made, whose requests get responses etc.). we find that both gender and gender environment affect the ways power is manifested in dialog, resulting in patterns that reveal the underlying factors. finally, we show the utility of gender information in the problem of automatically predicting the direction of power between pairs of participants in email interactions. keywords: computational sociolinguistics, gender, power, dialog 1. introduction it has long been observed that men and women communicate differently in different contexts. there has been an array of studies in sociolinguistics that analyze the interplay between gender and power. these sociolinguistic studies often rely on case studies or surveys. the availability of large corpora of naturally occurring interactions, and of advanced computational techniques to process the language and dialog structure of these interactions, has given us the opportunity to study the interplay between gender, power, and language use at a scale that was not feasible before. in this paper, we study how gender correlates with manifestations of power in an organizational setting using the enron email corpus. we investigate three factors that affect choices in communication: the writer’s c⃝2017 vinodkumar prabhakaran, owen rambow this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). prabhakaran and rambow gender, the gender of his or her fellow discourse participants (what we call the “gender environment”), and the power relations he or she has to the discourse participants. we focus on modeling the writer’s choices related to discourse structure, rather than lexical choice. specifically, our goal is to show that gender, gender environment, and power all affect individuals’ choices in complex ways, resulting in patterns in the discourse that reveal the underlying factors. we make three major contributions in this paper. first, we introduce an extension to the enron corpus of emails: we semi-automatically identify the sender’s gender of 87% of email messages in the corpus. this extension has been made publicly available.1 second, we use this enriched version of the corpus to investigate the interaction of hierarchical power and gender. we formalize the notion of “gender environment”, which reflects the gender makeup of the discourse participants of a particular conversation. we study how gender, power, and gender environment influence discourse participants’ choices in dialog. this contribution shows how social science can benefit from advanced natural language processing techniques in analyzing corpora, allowing social scientists to tackle corpora that cannot be examined in their entirety manually. third, we show that the gender information in the enriched corpus can be useful for computational tasks, specifically for improving the performance of the power prediction system from our prior work (prabhakaran and rambow, 2014) that is trained to predict the direction of hierarchical power between participants in an interaction. our use of the gender-based features boosts the accuracy of predicting the direction of power between pairs of email interactants from 68.9% to 70.2% on an unseen test set. we start by discussing related work in sociolinguistics on the interplay between gender and power followed by work within the nlp community on gender and use of language. in section 3, we present the first contribution of this paper — the gender identified enron corpus, and describe the procedure followed to build this resource and present various corpus statistics. section 4 introduces the notion of gender environment and section 5 presents the analysis framework used in this paper. in section 6 and section 7, we present the statistical analysis of the interplay between gender, gender environment, and power, through the lens of dialog behavior. in section 8, we demonstrate the utility of gender-based features in automatically predicting the direction of power between participants of an interaction, before we summarize our contributions in section 9. 2. literature review there is much work in sociolinguistics on how gender and language use are interrelated (tannen, 1991, 1993; holmes, 1995; kendall and tannen, 1997; coates, 1998; eckert and mcconnell-ginet, 2003; holmes and stubbe, 2003; mills, 2003; kendall, 2003; herring, 2008). some of this work looks specifically at language use in work environment and/or with respect to power relations, whereas some others study the gender differences in language use in general. understanding these different strands of research is important for a computational linguist working in this area. in this section, we summarize this literature, focusing more on the studies that have influenced the work presented in this paper. 2.1 gendered differences in language use many sociolinguistics studies have found evidence that men and women differ considerably in the way they communicate. some researchers attribute this to psychological differences (gilligan, 1. http://www.cs.stanford.edu/˜vinod/giec.html (originally described in (prabhakaran et al., 2014)) 22 dialog structure through the lens of gender, gender environment, and power 1982; boe, 1987), whereas some others suggest socialization and gendered power structures within the society as its reasons (zimmerman and west, 1975; west and zimmerman, 1987; tannen, 1991). for instance, tannen (1991) argues that “for most women, the language of conversation is primarily a language of rapport: a way of establishing connections and negotiating relationships”, which she calls rapport-talk, whereas “for most men, talk is primarily a means to preserve independence and negotiate and maintain status in a hierarchical social order”, which she calls report-talk. along the same lines, holmes (1995) argues that “women are much more likely than men to express positive politeness or friendliness in the way they use language”. in addition to politeness, many other linguistic variables have been analyzed in this context. lakoff (1973) describes women’s speaking style as tentative and unassertive, and argues that women use question tags and hedges more frequently than men do. however, holmes (1992) found that the differential use of question tags in-fact depends on the function of the question tag in the interaction. she categorized the instances of question tags in terms of their functionality in the contexts in which they were used, and found that question tags used as a way to express uncertainty was done more by men, whereas question tags used as a way to facilitate communication was done more by women. researchers have also looked into interruption patterns in interactions in relation to gender. for example, zimmerman and west (1975) found that men interrupted conversations more often in cross-sex interactions, whereas there were no significant differences in interruptions in same-sex interactions. however, recent studies have suggested the need for a more nuanced view on the interplay between gender and language use. they argue that the differences observed by above studies are due to more complex processes at play than gender alone, and that one needs to take into account the context in which the interactions happened to understand the gender differences better. mills (2003) challenged the above line of analysis, especially holmes (1995)’s theory regarding women being more polite. she argues that politeness cannot be codified in terms of linguistic form alone and calls for “a more contextualized form of analysis, reflecting the complexity of both gender and politeness, and also the complex relation between them”. along those lines, coates (2013) also challenge lakoff (1973)’s theory on women’s language being unassertive. she points out that hedges are multi-functional constructs and the greater usage of hedges by women “can be explained in part by topic choice, in part by women’s tendency to self-disclose and in part by women’s preference for open discussion and a collaborative floor”. in other words, she argues that women using more hedges than men does not entail that women are unassertive, but instead is an artifact of what topics women often take part in. kunsmann (2013) connects the gender differences in language specifically to status, dominance and power. he argues that “gender and status rather than gender or status will be the determinant categories” of language use. in our work, we follow a similar approach. we do not study gender in isolation, but in the context of the social power relations as well as the gender environment of the interaction. 2.2 gender and power in work place within the area of studying gender and language use, there is substantial amount of work that is specifically related to the language use in work environment (west, 1990; tannen, 1994; kendall and tannen, 1997; kendall, 2003), mostly done through qualitative case studies. in general, these studies found that women use more polite language and are “less likely to use linguistic strategies that would make their authority more visible” (kendall, 2003). for instance, west (1990) found that male physicians and female physicians differed in how they gave directives to their patients. male 23 prabhakaran and rambow physicians aggravated their directives, whereas female physicians used forms that mitigated them. similarly, in the study of gender, power and language in large corporate work environments, tannen (1994) found that female managers use more face saving strategies (e.g., phrasing directives as suggestions: you might put in parentheses) when talking to subordinates, whereas male managers used language that reinforced status differences (e.g., oh, that’s too dry. you have to make it snappier!). kendall (2003) shows that this behavior is specific to women operating in work environments. she studied the demeanor of a woman exercising her authority at work and at home, and found that while the woman used mitigating strategies to exercise her authority at work (as found by other studies before), she created a demeanor of explicit authority when exercising her authority over her daughter at home. in this paper, we study this aspect using our formulation of overt displays of power, which are face-threatening acts that reinforce the status differences. our findings on the enron emails are also in line with the above findings; we observe that male managers use significantly more overt displays of power when interacting with subordinates, whereas female managers use significantly fewer of them. however, in contrast, we draw from a much larger-scale study in which we analyze thousands of email interactions rather than a handful of case studies in the above mentioned research. another line of work that has influenced our work is by holmes and stubbe (2003) studying the effects of gendered work environments in the manifestations of power. they provide two case studies that analyze not the differences between male and female managers’ communication, but the differences between female managers’ communication in more heavily female vs. more heavily male environments. they find that, while female managers tend to break many stereotypes of “feminine” communication, they have different strategies in connecting with employees and exhibiting power in the two gender environments. this work has inspired us to look at this phenomenon by formulating the notion of “gender environment” in our study. we adapt this notion to the level of an interaction, and define the gender environment of an email thread in terms of the ratios of males to females on a thread, allowing us to look at whether the manifestations of power change within a more heavily male or female thread. 2.3 computational approaches towards gender and power within the nlp community, there is a considerable amount of work on analyzing language use in relation to gender. early work attempted to use nlp techniques to automatically predict the gender of authors using lexical features. researchers have attempted gender prediction on a variety of genres of interactions such as emails, blogs, and online social networking websites such as twitter (corney et al., 2002; peersman et al., 2011; cheng et al., 2011; deitrick et al., 2012; alowibdi et al., 2013; nguyen et al., 2014). in more recent work, hovy (2015) argues for research in the other direction, showing the importance of using gender information for better performance on nlp tasks such as topic identification, sentiment analysis and author attribute identification. while automatically detecting gender is an interesting problem, our focus in this paper is not gender detection, but understanding the variations in linguistic patterns with respect to both gender and power. for this, we require a more reliable source of gender assignments. hence, we use publicly available name databases to reliably determine the gender of participants as we have access to the email authors’ names in our corpus. we believe that the gender-identified email corpus we present will aid further research in the area of gender detection. existing work on gender prediction relies on relatively smaller datasets. for example, corney et al. (2002) use around 4k emails from 24 dialog structure through the lens of gender, gender environment, and power 325 gender identified authors in their study. cheng et al. (2011) use around 9k emails from 108 gender identified authors. deitrick et al. (2012) use around 18k emails from 144 gender identified authors. in contrast, we build a gender-assigned email dataset that is orders of magnitude larger than these resources. our corpus contains around 97k emails whose authors are gender-identified, and these emails are from around 23k unique authors. there has also been work on using nlp techniques to analyze gender differences in language use by men versus women (mohammad and yang, 2011; bamman et al., 2012, 2014; agarwal et al., 2015). mohammad and yang (2011) analyze the way gender affects the expression of emotions in the enron corpus. they found that women send and receive emails with relatively more words that denote joy and sadness, whereas men send and receive relatively more words that denote trust and fear. for their study, they assigned gender for the core employees in the corpus based on whether the first name of the person is easily gender identifiable or not. if the person had an unfamiliar name or a name that could be of either gender, they marked his/her gender as unknown and excluded them from their study. for example, the gender of the employee kay mann was marked as unknown in their gender assignment. however, in our work, we manually research and determine the gender of every core employee. bamman et al. (2012, 2014) study gender differences in the microblog site twitter. one of the many insights from their work is that gendered linguistic behavior is determined by a number of factors, one of which includes the speaker’s audience, which is similar to our notion of gender environment. their work looks at twitter users whose linguistic style fails to identify their gender in classification experiments, and finds that the linguistic gender norms can be influenced by the style of their interlocutors. more specifically, people with many same-gender friends tend to use language that is strongly associated with their gender, whereas people with more balanced social networks tend not to. our notion of gender environment captures the gender makeup of an interaction, and our findings reaffirms the need to also look into the audience’s gender makeup in studying gender. nlp approaches have also been applied recently to analyzing manifestations of power in social interactions. while early studies focus on hierarchical power relations (bramsen et al., 2011; gilbert, 2012; danescu-niculescu-mizil et al., 2012), other forms of power such as situational power and influence (prabhakaran et al., 2012a; prabhakaran and rambow, 2013; biran et al., 2012; rosenthal, 2014; rosenthal and mckeown, 2017), power of confidence in political discourse (prabhakaran et al., 2013), and pursuit of power in online forums (swayamdipta and rambow, 2012) have also been explored. in (prabhakaran, 2015), we present a comprehensive survey of literature in this area. to our knowledge, ours is the first computational study of this scale that focus on the interplay between gender and power in organizational email. we study the effects of gender in workplace interactions, not by considering the email senders’ gender in isolation, but together with their power relations with the rest of the participants, as well as the gender makeup of the interaction. 3. gender identified enron corpus in this section, our starting point is the corpus (enron-all) used in our prior work (prabhakaran and rambow, 2014). this corpus is derived from the enron email corpus (klimt and yang, 2004) that contains emails from the mailboxes of 145 “core” enron employees that were publicly released by the federal energy regulatory commission during its investigation of irregularities in enron. our version of the corpus captures the hierarchical power relations between 13,724 pairs 25 prabhakaran and rambow of employees assigned by agarwal et al. (2012), as well as the thread structure of email messages semi-automatically assigned by yeh and harnly (2006). the thread structure allows us to go beyond isolated messages and study gender in relation to the dialog structure as well as the language use. however, there are 34,156 unique discourse participants (senders and recipients together) across all the email threads in the corpus, and manually determining the gender of all of them is not feasible. hence, we adopt a two-step approach through which we reliably identify the gender of a large majority of discourse participants in the corpus. step 1: manually determine the gender of the 145 core employees who have a bigger representation in the corpus step 2: systemically determine the gender of the rest of the discourse participants using the social security administration’s baby names database we adopt a conservative approach so that we assign a gender only when the name of the participant meets a very low ambiguity threshold. 3.1 manual gender assignment we researched each of the 145 core employees using web search and found public records about them or articles referring to them. in order to make sure that the results are about the same person we want, we added the word enron to the search queries. within the public records returned for each core employee, we looked for instances in which they were being referred to either using a gender revealing pronoun (he/him/his vs. she/her) or using a gender revealing addressing form (mr. vs. mrs./ms./miss). since these employees held top managerial positions within enron at the time of bankruptcy, it was fairly easy to find public records or articles referring to them. for example, the sentence “kay mann is a strong addition to noble’s senior leadership team, and we’re delighted to welcome her aboard” (gender-revealing pronoun emphasized) in the page we found for kay mann clearly identifies her gender.2 we were able to correctly determine the gender of each of the 145 core employees in this manner. a benefit of manually determining the gender of these core employees is that it ensures a high coverage of 100% confident gender assignments in the corpus, as they are involved in all threads in the corpus. 3.2 automatic gender assignment our corpus contains a large number of discourse participants in addition to the 145 core employees for which we manually identified the gender. the steps we follow to assign gender for these other discourse participants is represented graphically in figure 1. we first determine the first names of discourse participants and then find how ambiguous the names are by querying the social security administration’s (ssa) baby names dataset. in this section, we start by describing how we calculate an ambiguity score for a name using the ssa dataset and then describe how we use it to determine the gender of discourse participants in our corpus. 2. http://www.prnewswire.com/news-releases/kay-mann-joins-noble-as-general-counsel-57073687.html 26 dialog structure through the lens of gender, gender environment, and power figure 1: automatic gender assignment process. 3.2.1 ssa names and gender dataset the us social security administration maintains a dataset of baby names, gender, and name count for each year starting from the 1880s, for names with at least five counts.3 we used this dataset in order to determine the gender ambiguity of a name. the enron data set contains emails from 1998 to 2001. we estimate the common age range for a large, corporate firm like enron at 24-67,4 so we used the ssa data from 1931-1977 to calculate ambiguity scores for our purposes. for each name n in the database, let mp(n) and fp(n) denote the percentages of males and females with the name n. the difference between these percentages of a name gives us a measure of how ambiguous it is; the smaller the difference, the more ambiguous the name. we define the ambiguity score of a name n, denoted by as (n), as follows: as (n) = 100− |mp(n)− fp(n)| the value of as (n) varies between 0 and 100. a name that is ‘perfectly unambiguous’ would have an ambiguity score of 0, while a ‘perfectly ambiguous’ name (i.e., 50%/50% split between genders) would have an ambiguity score of 100. we assign the likely gender of the name to be the one with the higher percentage, if the ambiguity score is below a threshold ast . g(n) =  male(m), if as (n) ≤ ast and mp(n) > fp(n) female(f ), if as (n) ≤ ast and mp(n) < fp(n) indeterminate(i), if as (n) > ast figure 2 shows the plot of the percentage of names that will be gender assigned in the ssa dataset against the ambiguity threshold. as the plot shows, around 88% of the names in the ssa dataset have as (n) = 0, i.e., are unambiguous. we choose a very conservative threshold of ast = 10 for our gender assignments, which assigns gender to around 93% names in the ssa dataset. an ambiguity threshold of 10 means that we assign a gender only if at least 95% of people with that 3. http://www.ssa.gov/oact/babynames/limits.html 4. http://www.bls.gov/cps/demographics.htm 27 prabhakaran and rambow figure 2: plot of percentage of first names covered against ambiguity threshold. name were of that gender. in the gender assigned corpus that we released, we retain the as (n) of each name, so that the users of this resource can decide the threshold that suits their needs. 3.2.2 identifying the first name each discourse participant in our corpus has at least one email address and zero or more names associated with it. the name field is automatically assembled by yeh and harnly (2006), who captured the different names from email headers. the names in the email headers are populated from individual email clients the senders were using and hence do not follow a standard format. to make things worse, not all discourse participants are human; some may refer to organizational groups (e.g., hr department) or anonymous corporate email accounts (e.g., a webmaster account, do-notreply address etc.). the name field may sometimes be empty, contain multiple names, contain an email address, or show other irregularities. hence, it is nontrivial to determine the first name of our discourse participants. we used the heuristics below to extract the set of candidate names for each discourse participant. • if the name field contains two words, pick the second or first word, depending on whether a comma separates them or not; pick the first word if the name field does not contain a comma; pick the word following the comma if it does contain one. • if the name field contains three words and a comma, choose the second and third words (a likely first and middle name, respectively). if the name field contains three words but no comma, choose the first and second words (again, a likely first and middle name). • if the name field contains an email address, pick the portion from the beginning of the string to a ‘.’,‘ ’ or ‘-’; if the email address is in camel case, take portion from the beginning of the string to the first upper case letter. • if the name field is empty, apply the above rule to the email address field to pick a name. in addition, we cleaned up some irregularities that were present in the name field. one common issue was that many email fields started with the text “?s” possibly a manifestation of some data preprocessing step. we strip this portion of the string in order to obtain the part that denote the actual email address. 28 dialog structure through the lens of gender, gender environment, and power the above heuristics create a list of candidate names for each discourse participant. for each candidate name, we compute the ambiguity score (section 3.2.1) and the likely gender. we find the candidate name with the lowest ambiguity score that passes the threshold and assign the associated gender to the discourse participant. if none of the candidate names for a discourse participant passes the threshold, we assign the gender to be indeterminate. we also assign the gender to be indeterminate, if none of the candidate names is present in the ssa dataset. this will occur if the name is a first name that is not in the database (an unusual or international name; e.g., vladi), or if no true first name was found (e.g., the name field was empty and the email address was only a pseudonym). this will also include most of the cases where the discourse participant is not a human (e.g., hr department). 3.2.3 coverage and accuracy we evaluated the coverage and accuracy of our gender assignment system on the manually assigned gender data of the 145 core people. we obtained a coverage of 90.3%, i.e., for 14 of the 145 core people, either their name’s ambiguity score was higher than the threshold (kam, lindy, tracy, lynn, chris, stacy, robin, stacey, and tori) or their name did not exist in the ssa dataset (geir and vladi). of the 131 people the system assigned a gender to, we obtained an accuracy of 89.3% in correctly identifying the gender. we investigated the errors and found that all errors were caused due to incorrectly identifying the first name. for the cases where we correctly identify the first name, we obtain a 100% accuracy in assigning the gender. the errors in finding first name arise because the name fields are automatically populated and sometimes the core discourse participants’ name fields include their secretaries’ who are of the other gender. while the name fields capturing multiple people is common for people in higher managerial positions, we expect this not to happen in the middle management and below, to which most of the automatically gender-assigned discourse participants belong. 3.3 corpus statistics and divisions gender assignment coverage: we apply the gender assignment system described above to all discourse participants of all email threads in the enron-all corpus to build the gender identified enron corpus (giec). table 1 shows the coverage of gender assignment in the giec corpus at different levels: unique discourse participants, messages and threads. we were able to identify the gender of 67% of unique discourse participants in the corpus. we verified that a majority of the cases where we could not assign the gender was due to the name of the sender email account not being present in the ssa dataset — mostly, cases where the discourse participant is not a human (e.g., hr department) as well as one-off email addresses (without a name entry) from outside the enron. in fact, the 67% discourse participants whose gender we could identify amounted to the senders of 87% of the messages in our corpus. we call the subset of threads for which we were able to identify the gender of all email senders, the all senders gender identified (asgi) sub-corpus, and those for which we were able to identify the gender of all participants including senders and all recipients, the all participants gender identified (apgi) sub-corpus. asgi covers around 71% of threads in the corpus, whereas apgi covers only about 49%. the users of this resource can limit their study to either subset, depending on their requirements. in figure 3, we show how the size of our gender identified enron corpus compares to existing gender assigned corpora within the emails domain (corney et al., 2002; cheng et al., 2011; deitrick 29 prabhakaran and rambow count (%) total unique discourse participants 34,156 gender identified 23,009 (67.3%) total messages 111,933 senders gender identified 97,255 (86.9%) total threads 36,615 all senders gender identified (asgi) 26,015 (71.1%) all participants gender identified (apgi) 18,030 (49.2%) table 1: coverage of gender identification at various levels: unique discourse participants, messages and threads. et al., 2012). our corpus is orders of magnitude larger than existing resources. we have representation of over 23k authors in our corpus, as opposed to a few hundred in other existing resources. in terms of number of messages also, our corpus is more than 5 times the size of next biggest corpus. (a) comparison in terms of number of unique discourse participants (b) comparison in terms of number of messages figure 3: gender identified enron corpus (giec) vs. existing gender assigned resources. gender assignment male/female split: in figure 4, we show the male/female percentage split of all unique discourse participants, as well as the split at the level of messages (i.e., messages sent by males vs. females). we have more male participants than female participants in the corpus (58% vs. 42%). when counted in terms of number of messages, around two thirds of the messages in our corpus were sent by men. 30 dialog structure through the lens of gender, gender environment, and power figure 4: male/female split in gender assignments across a) all unique participants who were gender identified (left), b) all messages whose senders were gender identified (right) 4. notion of gender environment in this study, we are interested not only in how the gender of a discourse participant affects their dialog behavior, but also whether the genders of other participants they are interacting with has an effect on their dialog behavior. we use the term “gender environment” to refer to the gender composition of a group who are communicating. we derive this notion from holmes and stubbe (2003) in which the term is used to refer to a stable work group who interact regularly. since we are interested in studying email conversations (threads), we adapt this notion to refer to a single thread at a time. we consider the “gender environment” to be specific to each discourse participant and to describe the other participants from his or her point of view. put differently, we use the notion of “gender environment” to model a discourse participant’s (potential) audience in a conversation. for example, a conversation among five women and one man looks like an all-female audience from the man’s point of view, but a majority-female audience from the women’s points of view. we define the gender environment of a discourse participant p in a thread t as follows. as discussed, we assume that the gender environment is a property of each discourse participant p in thread t. we take the set of all discourse participants of the thread t, pt, and exclude p from it: pt \ {p}. we then calculate the percentage of females in this set.5 we obtain three gender environments by setting thresholds on these percentages (dividing equally): female environment, mixed environment, and male environment. • female environment: if the percentage of women in pt \ {p} is above 66.7%. • mixed environment: if the percentage of women in pt \ {p} is between 33.3% and 66.7%. • male environment: if the percentage of women in pt \ {p} is below 33.3% 5. analysis framework in the rest of this paper, we use the all participants gender identified (apgi) subset of the enron corpus to study the interplay of gender and power, as it allows us to study the effects of both gender and gender environment. we use the same analysis framework — problem formulation, data splits, and features — introduced in (prabhakaran and rambow, 2014). in this section, we briefly 5. we note that one could also define the notion of gender environment at the level of individual emails: not all emails in a thread involve the same set of participants. we leave this to future work. 31 prabhakaran and rambow summarize the analysis framework and features we used. for a detailed account of the problem and features, refer to (prabhakaran and rambow, 2014). 5.1 power annotations our corpus contains organizational hierarchy relations extracted by agarwal et al. (2012) from the enron organizational charts. they define a dominance relation to be the relation between superior and subordinate in the hierarchy. their gold standard for hierarchy relations contains a total of 1,518 employees. they found 2,155 immediate dominance relations spread over 65 levels of dominance (ceo, manager, trader etc.) among these 1,518 employees. they also added the transitive closure of these relations to the corpus resulting in a total of 13,724 dominance relations. we use these dominance relations as our gold standard for assigning superior-subordinate relations. 5.2 problem formulation let t denote an email thread and mt denote the set of all messages in t . also, let pt be the set of all participants in t , i.e., the union of senders and recipients (to and cc) of all messages in mt . we are interested in detecting power relations between pairs of participants who interact within a given email thread. not every pair of participants (p1 , p2 ) ∈ pt × pt interact with one another within t . let imt(p1 , p2 ) denote the set of interaction messages — non-empty messages in t in which either p1 is the sender and p2 is one of the recipients or vice versa. we call the set of (p1 , p2 ) such that |imt(p1 , p2 )| > 0 the interacting participant pairs of t (ippt ). we focus on the manifestations of power in interactions between people across different levels of hierarchy. for every (p1 , p2 ) ∈ ippt , we query the set of dominance relations in the gold hierarchy to determine their hierarchical power relation (hp(p1 , p2 )). we exclude pairs that do not exist in the gold hierarchy from our analysis and denote the remaining set of related interacting participant pairs as rippt . we assign hp(p1 , p2 ) to be superior if p1 dominates p2 , and subordinate if p2 dominates p1 . in this paper, we are interested in how gender interacts with the differences in dialog behavior exhibited by superiors and subordinates. we study how a participant’s gender and the gender of other participants in an email thread affects these dialog behavior differences. we formulate the problem as a computational task. given a thread t and a pair of participants (p1 , p2 ) ∈ rippt , we want to automatically detect hp(p1 , p2 ). this problem formulation is similar to the ones in (bramsen et al., 2011) and (gilbert, 2012). however, the difference is that for us an instance is a pair of participants in a single thread of interaction (which may or may not include other people), whereas for them an instance constitutes all messages exchanged between a pair of people in the entire corpus. our formulation also differs from (prabhakaran and rambow, 2013) in that we detect power relations between pairs of participants, instead of just whether a participant had power over anyone in the thread. 5.3 data we follow the same train, dev, test division of enron-all as in (prabhakaran and rambow, 2014). we limit our study to the threads in which were able to identify the gender of all participants (i.e., threads that are part of the apgi subset of the corpus). table 2 presents the total number of pairs in ippt and rippt from all the threads in the apgi subset of our corpus and across the train, dev and test sets. we choose apgi instead of asgi (all senders gender identified) because apgi 32 dialog structure through the lens of gender, gender environment, and power description total train dev test # of threads 17,788 8,911 4,328 4,549∑ t |ippt | 74,523 36,528 18,540 19,455∑ t |rippt | 4,649 2,260 1,080 1,309 table 2: data statistics in the all participants gender identified subset of the enron corpus. row 1 presents the total number of threads in different subsets of the corpus. row 2 and 3 present the number of interacting participant pairs (ipp ) and related interacting participant pairs (ripp ) in those subsets. allows us to also study the notion of gender environment for which we need to know the gender of all participants. as an artifact of choosing the apgi, we also have a corpus with relatively smaller number of participants per thread than the full corpus. in other words, email threads with a large number of participants, such as broadcast emails, will have been excluded from the agpi, since there is a higher chance that the automatic gender assignment step fails to assign the gender for at least one of the recipients. as a result, the findings from the analysis on this subset sometimes differ from what we found in (prabhakaran and rambow, 2014). however, knowing how the two corpora differ in terms of the number of participants, it is interesting to note on which aspects of interactions the findings in both studies differ. 5.4 features we study the same dialog structural aspects of interaction introduced from (prabhakaran and rambow, 2014) in this work. in this section we briefly describe the various features we use to model these aspects of interactions. we focus on features in five different dialog structural aspects of interactions — positional, verbosity, thread structure, dialog acts, and overt display of power, as well as a non-structural aspect captured by lexical features. the first three aspects (positional, verbosity, and thread structure) capture the structure of message exchanges without doing any nlp processing on the content of the emails (e.g., how many emails did a person send), whereas dialog acts and overt display of power capture the pragmatics of the dialog and require an analysis of the content of the emails (e.g., did they issue any requests). lexical features also analyze the content, but at a shallow level, looking solely at word lemma and part-of-speech ngrams. each feature f is extracted with respect to a person p over a reference set of messages m (denoted f pm ). for example, msgratiokim mt denotes the ratio of messages sent by kim to the total number of messages in the thread t, whereas msgratiosaraimt (kim,sara) denotes the ratio of messages sent by sara to the total number of interaction messages between kim and sara in the thread t. for each pair (p1 , p2 ), we extract 4 versions of each feature f . f p1imt (p1 ,p2 ) : features with respect to p1 and interaction messages between p1 and p2 f p2imt (p1 ,p2 ) : features with respect to p2 and interaction messages between p1 and p2 f p1mt : features with respect to p1 and all messages in thread t f p2mt : features with respect to p2 and all messages in thread t 33 prabhakaran and rambow aspects features description pst initiator did p sent the first message? firstmsgpos relative position of p’s first message in m lastmsgpos relative position of p’s last message in m vrb msgcount count of messages sent by p in m msgratio ratio of messages sent in m tokencount count of tokens in messages sent by p in m tokenratio ratio of tokens across all messages in m tokenpermsg number of tokens per message in messages sent by p in m thr avgrecipients avge. number of recipients in messages avgtorecipients avge. number of to recipients in messages intolist% % of emails p received in which he/she was in the to list addperson did p add people to the thread? removeperson did p remove people to the thread? replyrate average number of replies received per message by p da reqactioncount # of request action dialog acts in p’s messages reqinformcount # of request information dialog acts in p’s messages informcount # of inform dialog acts in p’s messages conventionalcount # of conventional dialog acts in p’s messages danglingreq% % of p’s messages with requests that did not have a reply odp odpcount number of instances of overt displays of power lex lemmangram word lemma ngrams posngram part of speech (pos) ngrams mixedngram pos ngrams, with closed classes replaced with lemmas table 3: aspects of interactions analyzed in organizational emails. the first two versions capture behavior of the pair among themselves, while the third and fourth capture their overall behavior in the entire thread. in table 3, we list each feature f we use. like (prabhakaran and rambow, 2014), we use all four versions of the features in the machine learning experiments. however, for the statistical analysis presented in section 6 and section 7, we use the f p1mt version alone (similar results were obtained using the f p1imt (p1 ,p2 ) version as well). 5.4.1 positional features there are three features in this category — initiator, firstmsgpos, and lastmsgpos. initiator is a boolean feature which gets the value of 1 (true) if the p sent the first message in the thread, and 0 otherwise (false). firstmsgpos, and lastmsgpos are real-valued features taking values from 0 to 1, capturing relative positions of p’s first and last messages. the lower the value, the earlier the participant sent his/her first (or last) message. the first two features relate to the participant’s initiative. lastmsgpos captures whether the participant stays till the end of the email thread. 34 dialog structure through the lens of gender, gender environment, and power 5.4.2 verbosity features this set of features captures how verbose were the participants in the thread. there are five features in this set — msgcount, msgratio, tokencount, tokenratio, and tokenpermsg. the first two features measure verbosity in terms of p’s messages (raw counts and percentages), whereas the third and fourth features measure verbosity in terms of word tokens in p’s messages (raw counts and percentage). the last feature measure how terse or verbose on average p’s messages are. 5.4.3 thread structure features this set of features captures the structure of the email in terms of meta-data that is part of the email headers. it includes seven features — avgrecipients, avgtorecipients, intolist%, addperson, removeperson, and replyrate. the first two features capture the ‘reach’ of the person in terms of the average number of total recipients as well as recipients in the to list in emails sent by p. intolist% capture the the percentage of emails p received in which he/she was in the to list (as opposed to the cc list); the next two features —addperson and removeperson— are boolean features denoting whether p added or removed people when responding to a message. next, we look at the responsiveness towards p as the average number of replies received per message sent by p (replyrate). 5.4.4 dialog act features this feature set contains features that capture the dialog acts used by participants in the thread. we obtain dialog act tags on the entire corpus using the automatic dialog act tagger from our previous work (omuya et al., 2013). the da tagger labels each sentence to be one of the 4 dialog acts: • request-action: the writer signals her desire that the reader perform some non communicative act, i.e., an act that cannot in itself be part of the dialogue. for example, a writer can ask the reader to write a report or make coffee. • request-information: the writer signals her desire that the reader perform a specific communicative act, namely that he provide information (either facts or opinion). • inform: the writer conveys information, or more precisely, the writer signals her desire that the reader adopt a certain belief. it covers many different types of information that can be conveyed including answers to questions, beliefs (committed or not), attitudes, and elaborations on prior das. • conventional: dialog act does not signal any specific communicative intention on the part of the writer, but rather it helps structure and thus facilitate the communication. examples include greetings, introductions, expressions of gratitude, etc. the tagger uses a cascaded minority preference multi-class algorithm that posted significant improvements in its performance of identifying minority dialog acts such as request action (23% error reduction over the one-vs-all classification algorithm), and obtained an overall accuracy of 92%. please refer to (omuya et al., 2013) for more details on the dialog act tagging framework. we use 4 features: reqactioncount, reqinformcount, informcount, and conventionalcount to capture the number of sentences in messages sent by p that has each of these labels, respectively. we also use a feature to capture the percentage of p’s messages that had a request (either request-action or request-information), which did not get a reply, i.e., dangling requests (danglingreq%). 35 prabhakaran and rambow 5.4.5 overt display of power we use the notion of overt display of power (odp) introduced in our prior work (prabhakaran et al., 2012b) to measure face aggravating acts in the interactions. we define an utterance to have odp if it is interpreted as creating additional constraints on the response beyond those imposed by the general dialog act. for example, “i need the report by end of friday” would be considered as an overt display of power, whereas “could you please try to send the report by end of friday” would not be considered as one. we consider odp as a pragmatic concept, i.e., in terms of the dialog constraints an utterance introduces to its response, and not in terms of specific linguistic markers. for example, the use of politeness markers (e.g., please) does not, on its own, determine the presence or absence of an odp. in addition, the presence of odp cannot be determined solely based on syntactic patterns alone (e.g., declarative sentences such as i need the report may also function as odps). in (prabhakaran et al., 2012b), we presented a data-oriented approach of identifying instances of odps in email threads. we first obtained manual annotations of odp on a subset of 122 email threads (1734 sentences) at the sentence level, and then built an svm-based supervised machine learning model to identify instances of odp in new email threads. in addition to lexical features, it also uses the dialog act features obtained using the dialog act tagger described in section 5.4.4. our odp tagger has an accuracy of 96% and an f-measure of 54% over a random prediction baseline f-measure of 10.4%. in this paper, we applied the above odp tagger to the email threads in our entire corpus and used a feature odpcount that captures number of instances of overt displays of power in p’s messages. 5.4.6 lexical features in addition to the dialog structure features, we also used simple lexical ngram features as they have already been shown to be valuable in predicting power relations (bramsen et al., 2011; gilbert, 2012). we use the feature set lexical to capture word lemma ngrams, pos (part of speech) ngrams and mixed ngrams. a mixed ngram is a special case of word ngram where words belonging to open classes are replaced with their pos tags, thereby being able to capture longer sequences without increasing the dimensionality as much as word ngrams do. we found the best setting to be using both unigrams and bigrams for all three types of ngrams, by tuning on our dev set. 6. gender and power: a statistical analysis as a first step, we would like to understand whether male superiors, female superiors, male subordinates, and female subordinates differ in their dialog behavior. for this analysis, the anova (analysis of variance) test is the appropriate statistical test as it provides a way to test whether or not the means of several groups are equal. in other words, anova generalizes the student’s t-test to situations with more than two groups. it also eliminates the possibility of making a type i error (false positives) if multiple two-sample t-tests are applied to such a problem. we perform anova tests on all dialog structure features — positional, verbosity, thread structure, dialog acts, and overt display of power keeping both hierarchical power and gender as independent variables. this results in four groups — male superiors, female superiors, male subordinates, and female subordinates. it is crucial to note that anova only determines that there is a significant difference between groups, but does not tell which groups are significantly different. in 36 dialog structure through the lens of gender, gender environment, and power figure 5: mean value differences along gender and power: initiator (error bars indicate standard error) order to ascertain that, we use the tukey’s hsd (honest significant difference) test. we discuss the significant findings from these analyses below. altogether, there are twenty features as dependent variables, and two independent variables — power and gender. that is a total of sixty different statistical tests; in addition, for each anova test, we also perform the tukey’s hsd test. even after applying the bonferroni correction to control for multiple testing (i.e., significance level at 0.05/120=0.0008), many of the results we discuss below hold statistical significance. hence, our overall hypothesis that gender affects the way power is manifested in interactions holds true. however, as an exploratory study, we present the results along each individual aspect without applying the correction, as it has been shown that the bonferroni correction tends to be conservative. 6.1 positional features there are three features in this category — initiator, firstmsgpos, and lastmsgpos. initiator is a binary feature which gets the value of 1 (true) if the participant sent the first message in the thread, and 0 otherwise (false). firstmsgpos and lastmsgpos are real-valued features taking values from 0 to 1. the lower the value, the earlier the participant sent the first (or last) message. the first two features relate to the participant’s initiative. a higher average value for initiator in a group indicates that participants in that group initiates threads more often; so does a lower average value for firstmsgpos. lastmsgpos captures whether participant stayed on towards the end of the thread. figure 5 shows the mean values of each groups for the feature initiator. initiator and firstmsgpos behave more or less similarly; hence we show the chart only for initiator. subordinates initiate the threads significantly more often than superiors (average value of 0.39 against 0.28 for initiator). this pattern is also seen in firstmsgpos (0.18 over 0.23; lower value means earlier participation). both differences are highly statistically significant p < 0.001. at first, this finding appears to be in contrast with our finding in (prabhakaran and rambow, 2014) that superiors initiate more conversations. as we discussed earlier, this is an artifact of the fact that broadcast messages with large number of recipients get eliminated from our corpus because it is more likely to fail to assign gender to at least one of the participants. putting together both findings, we infer that superiors tend 37 prabhakaran and rambow figure 6: mean value differences along gender and power: lastmsgpos (error bars indicate standard error) to initiate email threads with large number of people; but in more focused conversations between smaller set of participants, it is the subordinates who initiate the conversations. gender is not a deciding factor. for initiator, the t-test result is significant (p = 0.03), however the magnitude of difference is very small (0.32 for females over 0.34 for males; figure 5). the t-test result is not significant for firstmsgpos. for the anova test for the combination of gender and power, the result is not significant for initiator. the anova test for firstmsgpos is significant, however the tukey’s hsd test shows that male and female superiors behaved more or less the same way; similarly, male and female subordinates also behaved the same way. the results on lastmsgpos is interesting (figure 6). the t-test results for both power and gender are significant, although the magnitude of the difference is small. the last message from superiors tend to come later than those of subordinates. similarly, males tend to send their last messages later than females. the anova results show that the factorial groups of power and gender also differ significantly (p < 0.01). upon tukey’s hsd test we find that male managers are the only group that differs from everyone else. the differences between all other groups are not statistically significant. but male managers differed from every other group significantly (p < 0.01). it is unclear why there is a significant difference in this feature. a potential explanation is that superiors tend to have the final word in conversations, and this is more in the case of male superiors. however, it is unclear to tease this apart as conversations are very often taken offline and hence it is hard to tell who had the final word. a more controlled study will need to be performed in order to verify this hypothesis, which we cannot perform using our corpus. 6.2 verbosity features there are five features in this category — msgcount, msgratio, tokencount, tokenratio, and tokenpermsg. the first two features measure verbosity in terms of messages, whereas the third and fourth features measure verbosity in terms of words. the last feature measure how terse or verbose on average the messages are. msgcount and msgratio behaved similarly, so did tokencount and tokenratio. figure 7 and figure 8 show the mean values of each groups for the feature msgcount and tokencount. superiors 38 dialog structure through the lens of gender, gender environment, and power figure 7: mean value differences along gender and power: msgcount (error bars indicate standard error) figure 8: mean value differences along gender and power: tokencount (error bars indicate standard error) tend to send fewer of messages in the thread than subordinates (p < 0.001), and women tend to send fewer messages than men (p < 0.001). the anova results for both msgcount and msgratio are significant (p < 0.001). tukey’s hsd test reveals an interesting picture. female superiors send significantly fewer messages than everyone else, almost 25% fewer than other groups. in fact, they are the only single group that is different from anyone else. difference between none of the other groups are significant. for tokencount and tokenratio, the results are similar. superiors tend to contribute fewer words in the thread than subordinates (p < 0.001). women tend to contribute fewer words than men (p < 0.01). the anova test of both features returned not significant. tokenpermsg behave differently. gender is not significant at all. that is, men and women do not differ in how long their messages are. in terms of power, subordinates send significantly longer emails. the anova test is highly significant. it turns out that among superiors, there is no significant difference. but among subordinates, male subordinates send significantly longer 39 prabhakaran and rambow figure 9: mean value differences along gender and power: tokenpermsg (error bars indicate standard error) figure 10: mean value differences along gender and power: replyrate (error bars indicate standard error) emails than female subordinates (p < 0.01) as per the tukey’s hsd test. in summary, power is a deciding factor in the difference between the verbosity exhibited by men and women. female managers send significantly fewer messages than all other groups; both female and male managers send significantly shorter messages than subordinates. on the other hand, female subordinates send significantly shorter emails than male subordinates, although they do not differ in how many messages they send. 6.3 thread structure features while the verbosity and positional features measure behavioral aspects, thread structure features in general deal with functional aspects (e.g., is a participant in cc (carbon copy) a lot?). while being in cc as a feature might be significantly related to power relations, it is unlikely that someone keeps a person in cc based on their gender. similarly, adding or removing people to the conversation is 40 dialog structure through the lens of gender, gender environment, and power figure 11: mean value differences along gender and power: avgtorecipients (error bars indicate standard error) also a functional aspect of workplace interactions, and we do not expect gender to play a role there. as expected there is no significant difference between women and men for intolist%, addperson, and removeperson. the anova test also returned not significant. in other words, gender does not affect the way superiors and subordinates behave in terms of these aspects. the results from our analysis of replyrate is interesting. figure 10 shows the mean values for each group. females get significantly more replies to their messages p < 0.001. while power did not have a significant effect, the anova result is also significant. on further analysis, we find that the female superiors get the highest reply rate (p < 0.05). the difference between the replyrate for male and female subordinates is not significant. it is an interesting finding, since it is an instance of gender of a person with power affecting how others behave towards them. however, on combining this finding with the analysis of avgrecipients and avgtorecipients (figure 11), we find that female superiors on average had more recipients in their messages than any other groups. the difference in replyrate might also be a manifestation of the fact that female superiors send emails to larger number of people. 6.4 dialog act features we now discuss the finding in terms of dialog act counts. informcount and conventionalcount behave similarly for all three tests. however, the magnitude of difference between superiors and subordinates for informcount is much higher than that of conventionalcount (superiors had 42.4% lower value than subordinates for informcount as opposed to 13.8% in the case of conventionalcount). the anova test returned not significant, which means that the gender did not affect the way superiors or subordinates use either conventional or inform dialog acts. on the other hand, the finding on reqactioncount and reqinformcount are very interesting. there is no significant difference between men and women in how often they make requests for action (figure 12), whereas they differed significantly (p < 0.001) in terms of how often they request for information. women issue almost 41% more requests for information than men. the anova test for reqactioncount returned significance (p < 0.01), but not for reqinformcount. that is, gender affects how superiors and subordinates issue requests for actions, but not requests 41 prabhakaran and rambow figure 12: mean value differences along gender and power: reqactioncount (error bars indicate standard error) figure 13: mean value differences along gender and power: reqinformcount (error bars indicate standard error) for information. male superiors issue more requests for actions than male subordinates, whereas female superiors held back from making requests. in fact, there is no significant difference between male subordinates and female subordinates in terms of reqactioncount. for danglingreq%, there is no significant difference with respect to gender or gender and power together. 6.5 overt displays of power figure 14 shows the mean values of odp counts in each group of participants. the results obtained are similar to what we found for reqactioncount. both power and gender are significant on their own. subordinates had an average of 0.091 odp counts and superiors had an average of 0.114 odp counts. gender is also significant; females have an average of 0.086 odp counts and males had an average of 0.113 odp counts. when looking at the factorial groups of power and gender, however, several differences are very highly significant. male superiors use the most odps, with an average 42 dialog structure through the lens of gender, gender environment, and power figure 14: mean value differences along gender and power: odpcount (error bars indicate standard error) of 0.135 counts. somewhat surprisingly, female superiors use the least of the entire group, with an average of 0.072 counts. however, the differences among female superiors, female subordinates, and male subordinates are not significant, as per the tukey’s hsd test. 6.6 summary and discussion in summary, we find that gender affects the manifestations of power significantly along many linguistic and structural aspects of interactions. we summarize our findings below: • gender of the participants does not have much effect on the manifestations of power in positional features (ref. section 6.1) • gender does significantly affect the manifestations of power in verbosity features; of the anova tests we performed on the five verbosity features, three returned to be highly significant. (ref. section 6.2) • gender also affects the manifestations of power on some of the thread structure features such as reply rate and number of recipients. (ref. section 6.3) • power manifestations on the dialog act based features, especially the request features and overt displays of power are also affected highly significantly by the gender of the participants. (ref. section 6.4 and section 6.5) the findings presented in this section do not exhaust the possibilities of this corpus. however, it shows how computational techniques can aid in performing large-scale sociolinguistics analysis. in order to demonstrate this point, we attempted to verify a hypothesis derived from the sociolinguistics literature we consulted. the hypothesis we investigate is: • hypothesis 1: female superiors tend to use “face-saving” strategies at work that include conventionally polite requests and impersonalized directives, and that avoid imperatives (kendall, 2003). 43 prabhakaran and rambow our notion of overt display of power (odp) is a face-threatening communicative strategy (prabhakaran et al., 2012b). an odp limits the addressee’s range of possible responses, and thus threatens his or her (negative) face.6 we thus reformulate our hypothesis as follows: the use of odp by superiors changes when looking at the splits by gender, with female superiors using fewer odps than male superiors. we saw in the results presented in section 6.5 that this hypothesis is indeed true. we find that female superiors used the least number of odps among all groups. the results confirmed our hypothesis: female superiors use fewer odps than male superiors. however, we also see that among women, there is no significant difference between superiors and subordinates, and the difference between superiors and subordinates in general (which is significant) is entirely due to men. this in fact shows that a more specific (and more interesting) hypothesis than our original hypothesis is validated: only male superiors use more odps than subordinates. in other words, the fact that superiors use more odps than subordinates is entirely due to male superiors using more odps. similarly, the fact that men use more odps than women is also entirely due to superiors among men using significantly more odps. 7. statistical analysis: gender environment and power in this section, we present our investigation on whether the manifestations of power differs based on the gender environment. as in section 6, we use the anova test to assess the statistical significance of differences. we perform anova tests on all features keeping both power and gender environment (genderenv, hereafter) as independent variables. we also perform anova keeping genderenv alone as the independent variable; since genderenv has more than two groups, we cannot use student’s t-test. we verify our overall hypothesis that gender environment affects the way power is manifested in interactions; it still holds true even after applying the bonferroni correction for multiple tests. however, as we did in section 6, we do not apply the correction when describing the findings from the statistical analysis of each set of features separately in the rest of this section. 7.1 positional features for the positional features, any difference that we see in the feature values between different gender environments is not interesting. for example, it is not sensible to investigate whether the value of initiator is different between gender environments (all threads had to be initiated by someone). however, it is still interesting to see whether there is any connection between the gender environment and how the superiors and subordinates differ in terms of when they started and stopped participating in the threads. as we saw in section 6, subordinates initiate more emails than superiors (initiator) and overall start participating earlier in the thread (firstmsgpos). the anova test keeping power and genderenv as independent variables was highly significant (p < 0.001). in other words, the gender environment does affect the initiative shown by subordinates in starting email threads. figure 15 shows the mean values of each group. subordinates do start participating in the threads significantly earlier than superiors. however, the magnitude of this difference is dependent on the gender environment. this suggests that subordinates tend to show more initiative in female environments than other gender environments, and that superiors tend to start participating in the threads much later in female environments. for the relative position of last message, the anova results are not significant. 6. for a discussion of the notion of “face”, see (brown and levinson, 1987). 44 dialog structure through the lens of gender, gender environment, and power figure 15: mean value differences along gender environment and power: firstmsgpos (error bars indicate standard error) figure 16: mean value differences along gender environment and power: tokencount (error bars indicate standard error) 7.2 verbosity features as per the anova results, the gender environment has no significance in msgcount or in how power is manifested in msgcount. on the other hand, in terms of tokencount, there is a significant difference (p < 0.01) across gender environments (figure 16). the anova test keeping power and genderenv as independent variables also returned significance (p < 0.001). in fact, in male environments, there is no significant difference in tokencount between superiors and subordinates. subordinates behaved more or less the same across the gender environments, but superiors contributed much less in female and mixed environments. a similar pattern is also observed in tokenpermsg across different gender environments. 45 prabhakaran and rambow figure 17: mean value differences along gender environment and power: conventionalcount (error bars indicate standard error) 7.3 thread structure features the effect of gender environment on replyrate is minimal. we observed that the number of recipients (both avgrecipients and avgtorecipients) is significantly higher in the mixed environment than others. this, however, is another artifact of how our corpus is constructed. in a thread with large number of participants, it is more likely to have a mixed environment than either male or female environment. the anova test keeping power and genderenv also returned no significance for addperson and removeperson. in summary, the effect of gender environment on thread structure features is minimal. 7.4 dialog act features the results obtained on the anova tests for the dialog act features are interesting. we will start with the conventionalcount. figure 17 shows the mean values of conventionalcount in each subgroup of participants. hierarchical power is highly significant as per anova results. subordinates use conventional language more (0.60) than superiors (0.52). while the averages by genderenv differ, the differences are not significant. however, the groups defined by both power and genderenv have highly significant differences. subordinates in female environments use the most conventional language of all six groups, with an average of 0.79. superiors in female environments use the least, with an average of 0.48. in the tukey hsd test, the only significantly different pairs are exactly the set of subordinates in female environments paired with each other group. that is, subordinates in female environments use significantly more conventional language than any other group, but the remaining groups do not differ significantly from each other. we interpret this result to mean that subordinates are more comfortable in female environments to use a style of communication which includes more conventional dialog acts than outside the female environments. the anova tests for informcount also returned high significance. the difference between mean values of informcount feature in male environments and mixed environments are not significant; but it differed significantly between female environments and both male and mixed environments. the groups defined by both power and genderenv also have highly significant differences. 46 dialog structure through the lens of gender, gender environment, and power figure 18: mean value differences along gender environment and power: informcount (error bars indicate standard error) figure 19: mean value differences along gender environment and power: odpcount (error bars indicate standard error) there is no significant difference between superiors’ and subordinates’ count of inform dialog acts when operating in a male environment. in other words, the finding that subordinates use more inform dialog acts holds true only in female and mixed environments, but not in male environments. however, on comparing this result with our findings in terms of verbosity features (figure 16), we find that this is in fact an artifact of most of the contributions being inform statements (the findings in informcount mirror that of tokencount). the anova results for both reqactioncount, reqinformcount, and danglingreq% are not significant when tested using power and genderenv. the male environment had a significantly (p < 0.05) lower danglingreq%. 47 prabhakaran and rambow 7.5 overt displays of power the results of the anova analysis on odpcount are interesting. figure 19 shows the mean values of each group. as we saw already in section 6, superiors use significantly more overt displays of power than subordinates. however, this pattern varied across gender environments significantly. the same relationship holds only in a mixed gender environment, where also most of the odp occur. in male environments, there is no significant difference in odpcount between superiors and subordinates, whereas in female environments, the value of odpcount for superiors is significantly lower than that of subordinates. this goes in line with our finding in section 6 that female managers use fewer overt displays of power. 7.6 summary and discussion in summary, we find that gender environment also affects the manifestations of power significantly along different structural aspects of interactions. we summarize the main findings below: • the gender environment significantly affects the difference between the initiative (in terms of how early they participated in the threads) exhibited by superiors and subordinates. while subordinates show more initiative than superiors across all gender environments, the magnitude of this difference is the largest in female environments. (ref. section 7.1) • gender environment affects the difference in verbosity exhibited by superiors and subordinates. while subordinates contributed significantly more content (in terms of token count as well as tokens per message) than superiors, this difference is the least in male environments. (ref. section 7.2) • power manifestations on dialog act features also differ significantly across different gender environments. subordinates use significantly more conventional dialog acts than superiors only in female environments. on the other hand, the difference in the their usage of inform dialog acts is non-existent in male environments. (ref. section 7.4) • gender environment also affects the use of overt displays of power among subordinates and superiors. the fact that superiors use more overt displays of power is driven entirely by mixed environments. in male environments, superiors and subordinates do not differ in their usage of overt displays of power, while in female environments, superiors used less overt displays of power. (ref. section 7.5) similar to what we did in section 6.6, we attempt to verify a hypothesis derived from the sociolinguistics literature we consulted in relation to the notion of gender environment. the hypothesis we investigate is: • hypothesis 2: women when talking among themselves use language to create and maintain social relations, for example, they use more small talk (based on a reported “stereotype” in (holmes and stubbe, 2003)). we have at present no way of testing for “small talk” as opposed to work-related talk, so we instead test hypothesis 2 by asking how many conventional dialog acts a person performs. conventional dialog acts do not convey information or requests (both of which would typically be 48 dialog structure through the lens of gender, gender environment, and power work-related in the enron corpus), but instead establish communication (greetings) and to manage communication (sign-offs); since communication is an important way of creating and maintaining social relations, we can say that conventional dialog acts serve the purpose of easing conversations and thus of maintaining social relations. we make our hypothesis 2 more precise by saying that a higher number of conventional dialog acts will be used in female environments. we presented the results of our analysis of conventionalcount feature in section 7.4. our results first appears to be a negative result: while the averages by gender environment differ, the differences are not significant. however, we find that subordinates in female environments use significantly more conventional language than any other group, but the remaining groups do not differ significantly from each other. our hypothesis is thus only partially verified: while gender environment is a crucial aspect of the use of conventional das, we also need to look at the power status of the writer. while our hypothesis is not fully verified, we interpret the results to mean that subordinates are more comfortable in female environments to use a style of communication which includes more conventional das than outside the female environments. 8. utility of gender information in predicting power in this section, we investigate the utility of the gender information in the problem of predicting the direction of power presented in (prabhakaran and rambow, 2014). we expect the svm-based supervised learning system using quadratic kernel to capture the interdependence between dialog structure features and gender features that we found in our statistical analysis presented in section 6 and section 7. we perform our experiments on the enron-apgi subset, training a model using the same machine learning framework presented in (prabhakaran and rambow, 2014) using the related interacting participant pairs in the train subset of enron-apgi, and choosing the best model based on performance on the dev subset. we experimented using all subsets of features described in section 5.4. in addition, we add two gender-based feature sets: gender containing the gender of both persons of the pair and genderenv which is a singleton set with the gender environment as the feature. table 4 presents the results obtained using various feature combinations. note that the numbers presented in table 4 are not directly comparable to the results presented in (prabhakaran and rambow, 2014), since the results presented there are on the dev set of the enron-all corpus, whereas here we discuss results obtained on the dev set of the enron-apgi, which is a subset of around 50% of the enron-all corpus. the majority baseline obtains an accuracy of 55.8%. using the gender-based features alone performs only slightly better than the majority baseline, posting an accuracy of 57.6%. the best performance is obtained using a combination of lexical, thread structure, gender and genderenv, which posts an accuracy of 70.7%. removing the genderenv feature set decreases the accuracy marginally to 70.5%, whereas removing the gender features as well reduces the performance significantly to 68.2% (tested using mcnemar test). this reduction of 2.4% percentage points in accuracy shows that gender features are in fact useful for this power prediction task. the best performance feature set without using any gender information is the combination of lexical, thread structure, positional and verbosity, which reports an accuracy of 68.3%. the best performing feature set without using lexical is the combination of dialog acts, overt display of power, thread structure and gender (67.3%). removing the gender features from this reduces the performance to 64.6%. similarly, the best performing feature set which do not 49 prabhakaran and rambow description accuracy baselines majority 55.83 using gender features alone gen 57.59 gen + env 57.59 best feature sets lex + thr + gen + env 70.74 lex + thr + gen 70.46 lex + thr 68.24 lex + thr + pst + vrb 68.33 best without lexical da + odp + thr + gen 67.31 da + odp + thr 64.63 best with no content pst + vrb + thr + gen 66.57 pst + vrb + thr 62.96 table 4: results on using gender features for power prediction. pst: positional, vrb: verbosity, thr: thread structure, da: dialog acts, odp: overt display of power, lex: lexical, gen: genderenv: genderenv use the content of emails at all is positional + verbosity + thread structure + gender (66.6%). removing the gender features decreases the accuracy by a larger margin (5.4% accuracy reduction to 63.0%). it is interesting to look at the error reduction obtained by adding gender features to different feature sets. using gender features alone obtains only an error reduction of 4.0% over the majority baseline (i.e., without using any other features). however, the predictive value of gender features improves considerably when paired with other features. for the best feature set we obtained, the gender features contributed to an error reduction of 7.9% (68.2% to 70.7%). for the best feature set without using lexical also the gender features contributed a similar error reduction of 7.6% (64.63% to 67.3%). for the setting where no content features are used, gender features obtained an even higher error reduction of 11.0% (63.0% to 66.6%). in other words, the gender-based features on their own are not very useful, and gain predictive value only when paired with other features (as we are using a quadratic svm kernel). this is because the other features in fact make quite different predictions depending on gender and/or gender environment. nonetheless, we take these results as validation of the claim that gender-based features enhance the value of other features in the task of predicting power relations. on our blind test set, the majority baseline obtains an accuracy of 57.9% and the baseline system that does not use gender features obtains an accuracy of 68.9%. on adding the gender-based features, the accuracy of the system improves to 70.3%. 9. conclusion the first contribution of this paper is the new, freely available resource — gender identified enron corpus, an extension to the enron email corpus with 87% of the email senders’ gender identified. we used the social security administration’s baby-names database to automatically assess the gen50 dialog structure through the lens of gender, gender environment, and power der ambiguity of first names of email senders and assigned the gender to those whose names are highly unambiguous. our gender identified corpus is orders of magnitude larger than other existing resources in this domain that capture gender information. we expect it to be a rich resource for social scientists interested in the effect of power and gender on language use. our second contribution is the detailed statistical analysis of the interplay of gender, gender environment and power in how they affect the dialog behavior of participants of an interaction. we introduced the notion of gender environment to capture the gender makeup of the discourse participants of a particular interaction. we showed that gender and gender environment affect the ways power is manifested in interactions in complex ways, resulting in patterns in the discourse that reveal the underlying factors. while our findings pertain to the enron email corpus, we believe that the insights and techniques from this study can be extended to other genres in which there is an independent notion of hierarchical power, such as moderated online forums. finally, we showed the utility of gender information in the task of predicting the direction of power between pairs of participants based on single threads of interactions. we obtained statistically significant improvements by adding the gender of both participants of a pair as well as the gender environment as features to a system trained using lexical and dialog structure features alone. acknowledgment this paper is partially based upon work supported by the darpa deft program. the views expressed here are those of the author(s) and do not reflect the official policy or position of the department of defense or the u.s. government. we thank emily reid who was involved in early stages of this work. we also thank rob voigt, dan jurafsky and the anonymous reviewers for their helpful feedback. references apoorv agarwal, adinoyi omuya, aaron harnly, and owen rambow. a comprehensive gold standard for the enron organizational hierarchy. in proceedings of the 50th annual meeting of the acl (short papers), pages 161–165, jeju island, korea, july 2012. association for computational linguistics. url http://www.aclweb.org/anthology/p12-2032. apoorv agarwal, jiehan zheng, shruti kamath, sriramkumar balasubramanian, and shirin ann dey. key female characters in film have more to talk about besides men: automating the bechdel test. in proceedings of the 2015 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 830–840, denver, colorado, may–june 2015. association for computational linguistics. url http://www.aclweb.org/anthology/n15-1084. jalal s alowibdi, ugo buy, paul yu, et al. language independent gender classification on twitter. in advances in social networks analysis and mining (asonam), 2013 ieee/acm international conference on, pages 739–743. ieee, 2013. david bamman, jacob eisenstein, and tyler schnoebelen. gender in twitter: styles, stances, and social networks. corr, abs/1210.4567, 2012. url http://dblp.uni-trier.de/db/ journals/corr/corr1210.html#abs-1210-4567. 51 prabhakaran and rambow david bamman, jacob eisenstein, and tyler schnoebelen. gender identity and lexical variation in social media. journal of sociolinguistics, 18(2):135–160, 2014. issn 1467-9841. doi: 10.1111/josl.12080. url http://dx.doi.org/10.1111/josl.12080. or biran, sara rosenthal, jacob andreas, kathleen mckeown, and owen rambow. detecting influencers in written online conversations. in proceedings of the second workshop on language in social media, pages 37–45, montréal, canada, june 2012. association for computational linguistics. url http://www.aclweb.org/anthology/w12-2105. s kathryn boe. language as an expression of caring in women. anthropological linguistics, pages 271–285, 1987. philip bramsen, martha escobar-molano, ami patel, and rafael alonso. extracting social power relationships from natural language. in proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 773– 782, portland, oregon, usa, june 2011. association for computational linguistics. url http://www.aclweb.org/anthology/p11-1078. penelope brown and stephen c. levinson. politeness : some universals in language usage (studies in interactional sociolinguistics). cambridge university press, february 1987. isbn 0521313554. url http://www.amazon.com/exec/obidos/redirect?tag= citeulike07-20\&path=asin/0521313554. na cheng, r. chandramouli, and k. p. subbalakshmi. author gender identification from text. digit. investig., 8(1):78–88, july 2011. issn 1742-2876. doi: 10.1016/j.diin.2011.04.002. url http://dx.doi.org/10.1016/j.diin.2011.04.002. jennifer coates. language and gender: a reader. wiley-blackwell, 1998. jennifer coates. women, men and everyday talk. palgrave macmillan, 2013. isbn 9781137314949. url https://books.google.com/books?id=ed3qaqaaqbaj. malcolm corney, olivier de vel, alison anderson, and george mohay. gender-preferential text mining of e-mail discourse. in computer security applications conference, 2002. proceedings. 18th annual, pages 282–289. ieee, 2002. cristian danescu-niculescu-mizil, lillian lee, bo pang, and jon kleinberg. echoes of power: language effects and power differences in social interaction. in proceedings of the 21st international conference on world wide web, www ’12, new york, ny, usa, 2012. acm. isbn 9781-4503-1229-5. doi: 10.1145/2187836.2187931. url http://doi.acm.org/10.1145/ 2187836.2187931. william deitrick, zachary miller, benjamin valyou, brian dickinson, timothy munson, and wei hu. author gender prediction in an email stream using neural networks. journal of intelligent learning systems & applications, 4(3), 2012. penelope eckert and sally mcconnell-ginet. language and gender. cambridge university press, 2003. 52 dialog structure through the lens of gender, gender environment, and power eric gilbert. phrases that signal workplace hierarchy. in proceedings of the acm 2012 conference on computer supported cooperative work, cscw ’12, pages 1037–1046, new york, ny, usa, 2012. acm. isbn 978-1-4503-1086-4. carol gilligan. in a different voice. harvard university press, 1982. susan c herring. gender and power in on-line communication. the handbook of language and gender, page 202, 2008. janet holmes. an introduction to sociolinguistics. pearson longman, 1992. janet holmes. women, men and politeness. longman, 1995. janet holmes and maria stubbe. “feminine” workplaces: stereotype and reality. the handbook of language and gender, pages 572–599, 2003. dirk hovy. demographic factors improve classification performance. in proceedings of the 53rd annual meeting of the association for computational linguistics, beijing, china, july 2015. association for computational linguistics. shari kendall. creating gendered demeanors of authority at work and at home. the handbook of language and gender, page 600, 2003. shari kendall and deborah tannen. gender and language in the workplace. in gender and discourse, pages 81–105. sage, london, 1997. bryan klimt and yiming yang. the enron corpus: a new dataset for email classification research. in machine learning: ecml 2004, pages 217–226. springer, 2004. peter kunsmann. gender, status and power in discourse behavior of men and women. linguistik online, 5(1), 2013. robin lakoff. language and woman’s place. language in society, 2(01):45–79, 1973. sara mills. gender and politeness, volume 17. cambridge university press, 2003. saif mohammad and tony yang. tracking sentiment in mail: how genders differ on emotional axes. in proceedings of the 2nd workshop on computational approaches to subjectivity and sentiment analysis (wassa 2.011), pages 70–79, portland, oregon, june 2011. association for computational linguistics. url http://www.aclweb.org/anthology/w11-1709. dong nguyen, dolf trieschnigg, a. seza doğruöz, rilana gravel, mariet theune, theo meder, and franciska de jong. why gender and age prediction from tweets is hard: lessons from a crowdsourcing experiment. in proceedings of coling 2014, the 25th international conference on computational linguistics: technical papers, pages 1950–1961, dublin, ireland, august 2014. dublin city university and association for computational linguistics. url http://www.aclweb.org/anthology/c14-1184. 53 prabhakaran and rambow adinoyi omuya, vinodkumar prabhakaran, and owen rambow. improving the quality of minority class identification in dialog act tagging. in proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 802–807, atlanta, georgia, june 2013. association for computational linguistics. url http://www.aclweb.org/anthology/n13-1099. claudia peersman, walter daelemans, and leona van vaerenbergh. predicting age and gender in online social networks. in proceedings of the 3rd international workshop on search and mining user-generated contents, pages 37–44. acm, 2011. vinodkumar prabhakaran. social power in interactions: computational analysis and detection of power relations. phd thesis, columbia university, 2015. vinodkumar prabhakaran and owen rambow. written dialog and social power: manifestations of different types of power in dialog behavior. in proceedings of the sixth international joint conference on natural language processing, pages 216–224, nagoya, japan, october 2013. asian federation of natural language processing. url http://www.aclweb.org/ anthology/i13-1025. vinodkumar prabhakaran and owen rambow. predicting power relations between participants in written dialog from a single thread. in proceedings of the 52nd annual meeting of the association for computational linguistics (volume 2: short papers), pages 339–344, baltimore, maryland, june 2014. association for computational linguistics. url http://www.aclweb. org/anthology/p14-2056. vinodkumar prabhakaran, owen rambow, and mona diab. who’s (really) the boss? perception of situational power in written interactions. in 24th international conference on computational linguistics (coling), mumbai, india, 2012a. association for computational linguistics. vinodkumar prabhakaran, owen rambow, and mona diab. predicting overt display of power in written dialogs. in human language technologies: the 2012 annual conference of the north american chapter of the association for computational linguistics, montreal, canada, june 2012b. association for computational linguistics. vinodkumar prabhakaran, ajita john, and dorée d. seligmann. who had the upper hand? ranking participants of interactions based on their relative power. in proceedings of the ijcnlp, pages 365–373, nagoya, japan, october 2013. asian federation of natural language processing. url http://www.aclweb.org/anthology/i13-1042. vinodkumar prabhakaran, emily e. reid, and owen rambow. gender and power: how gender and gender environment affect manifestations of power. in proceedings of the 2014 conference on empirical methods in natural language processing (emnlp), pages 1965–1976, doha, qatar, october 2014. association for computational linguistics. url http://www.aclweb.org/ anthology/d14-1211. sara rosenthal. detecting influencers in social media discussions. xrds: crossroads, the acm magazine for students, 21(1):40–45, 2014. 54 dialog structure through the lens of gender, gender environment, and power sara rosenthal and kathleen mckeown. detecting influencers in multiple online genres. acm trans. internet technol., 17(2):12:1–12:22, march 2017. issn 1533-5399. doi: 10.1145/ 3014164. url http://doi.acm.org/10.1145/3014164. swabha swayamdipta and owen rambow. the pursuit of power and its manifestation in written dialog. in semantic computing (icsc), 2012 ieee sixth international conference on, pages 22–29, 2012. deborah tannen. you just don’t understand: women and men in conversation. virago london, 1991. deborah tannen. gender and conversational interaction. oxford: oxford university press., 1993. deborah tannen. talking from 9 to 5: how women’s and men’s conversational styles affect who gets heard, who gets credit, and what gets done at work. w. morrow, 1994. isbn 9780688112431. url https://books.google.com/books?id=up7ehyicxqyc. candace west. not just ‘doctors’ orders’: directive-response sequences in patients’ visits to women and men physicians. discourse & society, 1(1):85–112, 1990. candace west and don h zimmerman. doing gender. gender and society, 1(2):125–151, 1987. jen-yuan yeh and aaron harnly. email thread reassembly using similarity matching. in ceas 2006 the third conference on email and anti-spam, july 27-28, 2006, mountain view, california, usa, mountain view, california, usa, july 2006. don h. zimmerman and candace west. sex roles, interruptions and silences in conversation. language and sex: difference and dominance, 1975. 55 narrative elements in expository texts dialogue & discourse 12(2) 115-144 doi: 10.5210/dad.2021.204 ©2021 nina l. sangers, jacqueline evers-vermeul, ted j.m. sanders and hans hoeken this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). narrative elements in expository texts: a corpus study of educational textbooks nina l. sangers n.l.sangers@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands jacqueline evers-vermeul j.evers@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands ted j.m. sanders t.j.m.sanders@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands hans hoeken j.a.l.hoeken@uu.nl utrecht institute of linguistics ots, utrecht university trans 10, 3512 jk, utrecht, the netherlands editor: amir zeldes submitted 01/2021; accepted 10/2021; published online 10/2021 abstract while the use of narrative elements in educational texts seems to be an adequate means to enhance students’ engagement and comprehension, we know little about how and to what extent these elements are used in the present-day educational practice. in this quantitative corpus-based analysis, we chart how and when narrative elements are used in current dutch educational texts (n=999). while educational texts have traditionally been considered prime exemplars of expository texts, we show that the distinction between the expository and narrative genre is not that strict in the educational domain: prototypical narrative elements – particularized events, experiencing characters, and landscapes of consciousness – occur in 45% of the corpus’ texts. their distribution varies between school subjects: while specific events, specific people, and their experiences are often at the heart of the to-be-learned information in history texts, narrativity is less present in the educational content of biology and geography texts. instead publishers employ narrative-like strategies to make these texts more concrete and imaginable, such as the addition of fictitious characters and representative entities. keywords: narrativity, educational texts, concreteness, imaginability, quantitative corpus-based analysis 1 introduction in our daily lives, we come across all kinds of genres: we skim through newspapers, laugh about our friends’ jokes, write grocery lists, listen to the latest pop songs, send e-mails to our colleagues, consult recipes for dinner, and watch the newest tv-series. ‘genre’ denotes a classification system to distinguish between different types of spoken and written text. in the linguistic literature, many sangers, evers-vermeul, sanders and hoeken 116 different genres have been categorized and described (for example, biber & conrad, 2009; martin & rose, 2008). however, broad theoretical consensus about the definition of specific genres is often lacking (chandler, 1997). this is partly caused by the many different approaches to the study of genre. for instance, while some researchers define genres primarily on the basis of conventions (for example, themes or settings) and/or forms (for example, structure and style), others also consider the situational and social context in which texts are formulated (for a discussion on different approaches, see chandler, 1997; lee, 2001). in addition, the definition of a certain genre may vary across domains, cultures, and historical periods of time (cf. biber & conrad, 2009). finally, the boundaries between genres tend to be ‘fuzzy’, as supergenres may be divided into subgenres in multiple ways (for example, “tv-series” into “crimes” or “drama”), and two or more genres may be merged into hybrid forms (for example, “romantic comedy”) (cf. chandler, 1997; santini, 2006). therefore, for any study into genres, it is essential to clearly spell out its focus. the current study focuses on the expository and narrative genre in the domain of dutch educational texts. traditionally, educational texts have been seen as prime exemplars of the expository genre, because they often introduce and explain new, subject-specific terms and/or concepts, such as the process of erosion in the geography text in (1).1 (1) under the influence of plant roots and the weather, rocks crumble all year. this is called erosion. during the winter, the process of erosion often proceeds faster. the water in the cracks and crevices of the mountain freezes and as a result pieces of rock are released. they bounce down the slope and break into smaller pieces. in their fall they take other stones with them. and all those rocks roll into the valley. (meander, physical geography grade 5, p. 12) however, not all dutch educational texts prove to be fully expository (sangers, evers-vermeul, sanders & hoeken, 2020). for example, the text in (2) presents a narrative about a prehistoric man named iugas. this text, which is placed at the beginning of a new chapter in a history textbook, is used as an introduction to to-be-learned information about the stone age and iceman ötzi, who lived in this time period. the fictional character iugas shows many similarities to ötzi. (2) the mountain towers high above iugas. heavy clouds gather around the top. just a little while and then the snow will fall. too early, it is too early. he has to go further. once he is over the mountain pass, he will be safe. a twinge of pain passes over his face. his chest hurts at the place where an arrowhead is stuck between his ribs. his shoulder hurts, also from an arrow. he was able to pull it out, but the wound still hurts. iugas has nothing to take care of the wound. he gathers all his strength. maybe he is still able to get over the mountain. as long as he keeps walking. (speurtocht, history grade 5, p. 8) in addition, dutch educational textbooks include hybrid texts, combining characteristics from both the expository and narrative genre. for instance, in the biology text in (3), to-be-learned information about the pollination of flowers is told by a forester, named jan, who can be considered a narrative character. (3) an entomophilous flower needs insects for its pollination. a windflower uses the wind. “windflowers don’t need to be noticeable to insects. that’s why they look different”, says forester jan. “they have long stamens and large pistils. those hang down from 1 throughout the paper, we represent excerpts from dutch educational textbooks by their english translations. narrative elements in expository texts 117 the flowers. as long as the wind can get to them, that’s the most important thing.” (argus clou natuur & techniek, biology grade 5, p. 49) narrative elements such as characters seem to be included in educational texts as a strategy to make these texts more engaging and better comprehensible, considering that many dutch students find their educational texts too boring to read and/or too difficult to understand (dood, gubbels & segers, 2020; gubbels, netten & verhoeven, 2017; gubbels, van langen, maassen & meelissen, 2019; inspectorate of education, 2017, 2020, 2021; sangers et al., 2020). while the use of narrative elements seems to be an adequate means to solve these readability issues (cf. norris, guilbert, smith, hakimelahi & phillips, 2005; sangers et al., 2020), we know little about how and to what extent these elements are used in present-day educational texts, and whether their distribution depends, for instance, on the school subject. therefore, in this quantitative corpus-based study, we aim at gaining insight into the use and distribution of narrative elements in educational texts, focusing on the current dutch educational practice. an added benefit of charting the use of narrative elements in current educational texts is that it will enable future empirical research to reflect actual practices. previous empirical studies into narrative elements in educational texts have been found to show conflicting results: while some studies indicate that narrative elements contribute positively to the comprehensibility of the to-belearned information (for example, eng, 2002; romero, paris & brem, 2005), other studies report negative outcomes (for example, cervetti et al., 2009; van silfhout, 2014). these conflicting results are partly explained by the way in which narrativity has been operationalized in the experimental texts of these studies, as the number and kinds of narrative elements used in these studies vary considerably (cf. sangers, evers-vermeul, sanders & hoeken, 2019). this suggests that the narrative genre has not yet been clearly defined within the boundaries of the educational domain. by basing future empirical research on actual practices, incomparability of empirical results due to too much divergence in narrative manipulations could be prevented. hence, the following question guided our research: how and when are narrative elements currently being used in dutch educational texts? before we explain how the use of narrative elements may vary between educational texts for different school subjects (section 3), we discuss how we define narrativity within the educational domain (section 2). 2 narrativity in the educational domain over time, many definitions of the concept of narrative have been formulated in the literature. some definitions have charted the linguistics features found in narrative texts, such as the combination of past tense verbs, third person pronouns, and adverbials of time and place (cf. biber & conrad, 2009; fleischman, 1990), while other definitions were greatly inspired by the work of labov (1972) and labov and waletzky (1967), who define narrative texts as consisting of six elements: abstract, orientation, complicating action, result, evaluation, and coda. for instance, work in systemic functional linguistics builds upon labov and waletzky’s structural approach, using schematic structures to define recurrent local patterns within and variation between genres – which are argued to enact the social practices of a given culture (cf. christie & martin, 2000; martin & rose, 2008). while previous studies have often focused on specific linguistic features of certain genres, our focus is on the text’s content. a recurrent element in content-based narratological definitions is the representation of events, although scholars have disagreed about whether a single event suffices for a narrative (abbott, 2008; genette, 1982), whether at least two events, ordered in time, are needed (labov, 1972; prince, 2003; rimmon-kenan, 2002), or whether the events of a narrative should be connected in a non-random way, including relations of causality (bal, 1997; onega & landa, 1996; sangers, evers-vermeul, sanders and hoeken 118 richardson, 1997; sanford & emmott, 2012). another frequently mentioned narrative element is a need for the involvement of human or quasi-human entities in the events described (fludernik, 2009; herman, 2009; norris et al., 2005; ryan, 2007). toolan (2001) has combined these elements into his definition of narrative, adding that readers should be able to ‘learn’ something from the agonist’s experiences: “a narrative is a perceived sequence of non-randomly connected events, typically involving, as the experiencing agonist, humans or quasi-humans, or other sentient beings, from whose experience we humans can ‘learn’” (toolan, 2001, p. 8). in a previous study, we have shown that toolan’s (2001) definition is well-applicable in the domain of dutch educational texts, and that a narrative educational text can be characterized as exhibiting the following three narrative elements: 1) a sequence of non-randomly connected particularized events, that are 2) experienced by a specific character, of whom 3) readers gain insight into the inner world (sangers et al., 2020). a more detailed definition of these elements is as follows. first, narrative educational texts represent two or more events that are connected in a logical way, for instance by means of causal or temporal relations. these events are particularized rather than generic: they take place only once, at one point in time and at one location, as opposed to recurrent phenomena (compare “yesterday, lisa went to the university hall in utrecht. after a short speech from her supervisor, she received her diploma for the bachelor communication and information studies” versus “many students graduate from university each year”). particularization is strengthened by references to the specific time and place at which these events take place. increasing the degree of detail for these aspects makes a text more concrete, and prompts sensory imagery of the text’s content. second, narrative educational texts contain at least one individual who experiences the events described in the text, either by taking active part in these events or by passively experiencing them. this character can be human as well as human-like (for example, an animal). human or humanlike groups are not considered specific characters (compare “lisa went to the university hall in utrecht” versus “all undergraduates went to the university hall in utrecht”). the third narrative element involves the representation of an inner world, the so-called “landscape of consciousness”, through the expression of thoughts, feelings and/or sensory perceptions (bruner, 1986). this landscape of consciousness is complementary to the text’s “landscape of action”, which is the relationship between the actions of a character and their consequences. the landscape of consciousness is usually linked to a specific character (“lisa was happy to receive her diploma”), but can also give insight into the inner world of a group (“all undergraduates were happy to receive their diplomas”). if a landscape of consciousness is not explicitly elaborated upon in a text, readers can infer this inner world themselves (cf. sangers et al., 2020). the three narrative elements are, for instance, combined in the history text about prehistoric iugas in (2). this text contains several logically related particularized events (for example, “he was able to pull it out”, “he gathers all his strength”), that are experienced by a fictitious individual named iugas, of whom readers gain insight into his inner world by the expression of his thoughts (for example, “too early, it is too early”) and sensory perceptions (for example, “a twinge of pain passes over his face”). while many traditional definitions present a binary interpretation of narrative, ryan (2007) proposes a scalar interpretation that focuses on the question “is text 1 more narrative than text 2?” rather than “is text 1 a narrative?” according to ryan (2007, p. 28), narratives should be viewed as “a fuzzy set allowing variable degrees of membership, but centered on prototypical cases that everybody recognizes as stories”. based on such a scalar interpretation, a text is most narrative if it contains all narrative elements that are considered prototypical in a certain domain. if a text contains some but not all prototypical narrative elements, this text is merely less pronounced narrative elements in expository texts 119 narrative than the narrative prototype, showing pronounced signs of narrativity. this indicates that all texts categorized as ‘narrative’ are related to a greater or lesser extent to the narrative prototype, without losing the ‘narrative’ label. taking such an interpretation within the educational domain allows for the inclusion of educational texts that combine expository and narrative features, such as the hybrid text in (3), in which forester jan tells about the pollination of flowers. in a prior study, we have qualitatively illustrated that hybrid forms of narrativity can be found in the educational domain, incorporating the three narrative elements mentioned earlier, which can be considered prototypical in the educational domain, in varying combinations (sangers et al., 2020). this variation can be visualized in the form of a venn diagram, in which each circle represents one of the three narrative elements (see figure 1). the intersections represent the different combinations of narrative elements, such as that of particularized events and an experiencing character on the left-hand side of the diagram. as figure 1 shows, the less-pronounced narrative texts evolve around prototypical ‘full’ narratives such as (2), which are classified in the center of the diagram. figure 1. different combinations of prototypical narrative elements in the educational domain. while we have previously shown that most areas of figure 1 can be identified in dutch educational texts (see sangers et al., 2020), the frequency with which the various combinations of prototypical narrative elements occur in dutch educational materials is still unknown. therefore, the model in figure 1 guided our first sub-question: rq1 how frequently do the different combinations of prototypical narrative elements occur in dutch educational texts? in section 3, we discuss why the frequency of the various combinations of narrative elements may vary across school subjects. 3 the role of narrative elements in different school subjects in the literature, it has been argued that concrete and imaginable information has a processing advantage over abstract information, being easier to comprehend and more interesting to read particularized events landscape of consciousness experiencing character sangers, evers-vermeul, sanders and hoeken 120 (nisbett & ross, 1980; sadoski & paivio, 1994; sadoski, paivio & goetz, 1991). an explanation for this ‘concreteness effect’ is given by the dual coding theory (dct, paivio, 1971, 1986; sadoski & paivio, 1994), which distinguishes between a verbal system – specialized in language processing – to represent information and a mental imagery system, which concerns the processing of world knowledge about events and objects. according to the dct, abstract information is stored only verbally, because it evokes less mental imagery, while concrete information is stored via both the verbal system and the mental imagery system. the activation of both cognitive systems elicits mental images that make concrete information more engaging, better comprehensible, and more easily retrievable from memory (sadoski, 1999, 2001). given the substantial empirical evidence supporting the claims about concreteness (for an overview, see sadoski, 2001), making educational texts more concrete and imaginable seems an adequate strategy to enhance their comprehensibility and attractiveness. one way of attaining more concreteness and imaginability seems to be the incorporation of narrative elements in educational texts: the more detailed information a text provides about specific events, specific characters, and the context, the more concrete and imagery-provoking this text is (nisbett & ross, 1980). the extent to which educational publishers make use of narrative elements in their texts, however, may be influenced by the nature of the to-be-learned information, which differs from school subject to school subject. that is, while specific events, specific people, and their experiences are often at the core of history texts, texts for biology and geography focus on recurrent natural phenomena and/or general processes, without human involvement or with humans being only passively involved (for example, in explanations about processes in the human body). for instance, the history text in (4) introduces the well-known historical figure john f. kennedy as a specific character who experiences the specific events that led to his unfortunate death in 1963. by contrast, the geography text in (5) discusses the origin of coal from the carboniferous marshes, describing events that are generic rather than specific, and including no human agents. (4) in the united states, a young, handsome president came to power in 1960: john f. kennedy. he had big plans for his country. the world was shocked when kennedy was murdered in 1963. he was shot dead while driving through the city of dallas in his open top car. the images of the murder were shown on television. (eigentijds, history grade 5, p. 50) (5) coal originated from the carboniferous marshes. this is how it happened. dead plants started to rot. the remains of these plants formed a layer of peat. water washed a layer of sand over it. on top of this layer, plants started to grow again. a new layer of peat was formed. during millions of years, those layers of peat were pressed together. a solid material, that we call coal, was formed. coal can be used as fuel. (grenzeloos, physical geography grade 5, pp. 8-9) while no narrative elements are included in (5), the narrative elements in (4) are at the heart of this text’s educational content and, therefore, cannot be disregarded. presuming a stronger connection between narrativity and the educational content of history texts, we expect to find the three prototypical narrative elements, and combinations thereof, more frequently in these texts than in texts for biology and geography. we should, however, acknowledge that the educational content of geography does not only cover natural phenomena, such as the origin of coal in (5), but also includes human-related topics, such as migration in (6). hence, it seems likely that the educational content of texts about human geography (gh texts) more frequently involves people than that in texts about physical geography (gp texts), offering more options to include narrative elements. however, the extent to which narrative elements occur in gh texts may still differ from that of history texts, as the educational content in gh texts tends to be less specific than that in history texts, focusing on general tendencies narrative elements in expository texts 121 rather than specific events, and on groups rather than specific characters. therefore, gh texts may occupy an intermediate position between history texts on the one hand, and biology and gp texts on the other hand. (6) migration can also cross borders. in case of emigration, people move to another place of residence in another country. in case of immigration, someone arrives at a country to settle there. in recent years, for instance, many polish people have come to the netherlands. initially they only came to the netherlands to work, but nowadays many of them settle here with their family. (de wereld van, gh grade 5, p. 29) following our line of reasoning that biology, gp, and gh texts are less often about specific events, specific characters, and specific contexts, we believe these texts will also tend to be more abstract and less imagery-provoking than their history counterparts. therefore, these texts might significantly benefit from narrative-like strategies to make them more concrete and imaginable. one such strategy could be the addition of a specific character to the text. in history texts, publishers can – and often have to – draw on real, well-known historical figures. the educational content of biology, gp, and gh texts, however, might often not contain such authentic characters. to add specific characters to texts for these school subjects, publishers would generally have to invent their own characters. for instance, in the biology text in (7), information about food chains is conveyed by two fictitious characters, eva and the guide, thereby making the to-be-learned information more concrete and imaginable. (7) eva listens to the guide. he describes how wild animals are constantly trying to survive. “a herd of zebras is grazing, over there. they do nothing but graze. and of course they pay attention. but at the moment they have a rest, because: do you see those lions there? they just caught a young zebra. now they are eating him.” “how sad for that little zebra,” eva says. “that’s how you see it,” the guide says, “but without those zebras, the lions would die. tell me: do you ever eat a sausage or a hamburger?” eva nods. “then you eat an animal, don’t you?” the guide says. “people eat meat as well. we keep cows and pigs to eat!” (wijzer! n&t, biology grade 5, p. 52) in addition, publishers can apply a “pars pro toto” strategy to support generic educational content. in this case, in addition to or instead of giving a summary description of a certain concept, a prototypical instance of this concept is highlighted. for example, the text in (1), which gives a generic description about the process of erosion, is preceded by the paragraph in (8), in which a specific instance of erosion is described. this paragraph presents a series of chronologically related events that focus on individualized natural entities (one cleft, one seed, and one tree) instead of the entire group of entities. as such, the generic to-be-learned information in (1) is introduced in a more concrete and imaginable way. similarly, in the biology text in (9), educational content about the maturation of babies is conveyed by focusing on the developmental process of one baby. the baby in (9), who cannot be linked one-on-one to a specific person in the real – or fictitious – world and, therefore, does not qualify as a character, is considered a “representative” of the entire group of babies. although focusing on a representative entity (a baby) instead of giving a summary description of all entities (all babies) makes an educational text more concrete and imaginable, this strategy is more conceptual than the description of a real or fictitious character who can be considered a typical example of the group of entities. for instance, the author of (9) could also have chosen to highlight a specific baby as an example case (for example, baby thomas). such an exemplar baby is more concrete than a representative baby, because it maps one-on-one to an individual in the real or fictitious world. rather than representing the entire group of entities, an sangers, evers-vermeul, sanders and hoeken 122 exemplar can be used for the categorization of new potential group members by way of comparison (cf. exemplar vs. prototype theory; murphy, 2016). from less to more concrete/imaginable, the above three ways to frame educational content about babies are schematically related to each other as follows: ‘all babies’ (group) → ‘a/the baby’ (representative) → ‘baby thomas’ (exemplar). (8) it started with a small crack in the rock. moisture started to grow in it and at some point some seeds. a tree grew from one of those seeds. the roots of that tree penetrated further and further into the crack. and so the crack became wider and deeper. snow and ice made the crack a little wider every year, because water expands when it freezes. and then one day, this huge piece of rock broke loose and popped down... (meander, gp grade 5, p. 12) (9) after seven weeks, a baby looks more like a tadpole than a human. but all kinds of things are coming into existence in that body. for instance, brains and a heart that pumps real blood. after twelve weeks, the baby has arms, legs, hands, and feet that move. this is how he exercises his muscles. (binnenstebuiten, biology grade 5, pp. 42-43) taken together, we expect 1) narrative elements to be more frequent in history texts than in texts for biology, gp, and gh, while hypothesizing that 2) strategies to make educational texts more concrete, such as the addition of fictitious characters and representative entities, are more frequent in the latter subjects. this motivated our second sub-question: rq2 to what extent are narrative elements applied differently in texts for biology, physical geography, human geography, and history? in section 4, we clarify the method of our quantitative corpus-based analysis. subsequently, in section 5, we describe the results of our analysis. finally, in section 6, we turn to our discussion and conclusion. 4 method in this section, we describe the material selection (section 4.1), method of analysis (section 4.2), inter-annotator agreement (section 4.3), and method of statistical analysis (section 4.4). 4.1 material selection to find out whether differences in the distribution of narrative elements over texts for biology, gp, gh, and history are generalizable over grade levels, our corpus-based analysis focused on texts for grade 5 and for grade 8 of pre-university education.2 while grade 5 students have acquired the basic reading skills required for a deep understanding of texts, grade 8 students need to be able to read more challenging texts, particularly in pre-university education. 4.1.1 textbook selection we selected educational texts from textbooks published by five well-known dutch educational publishers. for grade 5, one textbook was selected per subject per publisher, leading to a total of 2 the dutch system for secondary education is divided into three educational levels, ranging from theoretical to vocational training: pre-university education (dutch vwo), senior general education (dutch havo), and pre-vocational education (dutch vmbo). within pre-university education, we focused on texts for grade 8 (year 2 of dutch secondary education), because eighth graders are able to read texts on a more advanced level than seventh graders, who have only mastered a basic reading level (committee meijerink, 2009), and because eighth graders are still taking classes in all school subjects under investigation. narrative elements in expository texts 123 fifteen textbooks.3 for grade 8, only three out of five publishers also distributed textbooks on a preuniversity level. all three did so for geography and history, while only two of them also published a biology textbook, leading to a total of eight textbooks. see appendix a for a list of all twentythree textbooks. 4.1.2 chapter selection for history and biology, we selected one chapter per textbook. per geography textbook, two chapters were selected: one for gp and one for gh. this resulted in a selection of thirty-one chapters. we strived for thematic overlap per subject, both within and between grade levels. this was done to counter potential narrative distribution biases caused by topic selection as much as possible. thematic overlap was established on the basis of a comparison of keywords. for biology, the reproduction of humans, animals, and plants was chosen as the overlapping theme. for one grade 5 textbook including no information on reproduction, we selected a chapter on eating habits of animals and plants. for history, we selected chapters that discussed the time period of stadtholder william of orange, who led the dutch revolt against spain during the start of the eighty years’ war (15681648). for one grade 5 textbook including only events after 1900, we selected a chapter on the cold war. for geography, chapters were matched within grade level only, since it turned out to be unfeasible to select thematically overlapping chapters between grade levels. for gp, the grade 5 chapters were matched by their discussion of different sorts of landscapes, and the grade 8 chapters based on their focus on characteristics of the earth. for gh, the grade 5 chapters concentrated on the european union, while the grade 8 chapters focused on demographic notions such as ‘birth rate’ and ‘immigration’. even though the distinction between gp and gh chapters was generally straightforward, one grade 5 textbook paid equal attention to both sub-domains in all of its chapters. for this textbook, we selected a chapter on the climates and landscapes of eastern europe (gp) and a chapter on europe that included discussions on the european union (gh). 4.1.3 text selection within the chosen chapters, we selected texts that included educational content and/or background information. a text was taken to be a unit of at least three sentences that belonged to a marked text box, and/or was grouped under a subheading (blank lines did not mark the beginning of a new text). in those few cases in which these rules did not suffice, we looked at font characteristics in order to make a final decision. table 1 shows the number of texts per school subject and grade level. in total, the corpus consisted of 999 texts. subject grade 5 grade 8 total biology 147 84 231 gp 125 106 231 gh 124 118 242 history 137 158 295 total 533 466 999 table 1. number of texts per school subject and grade level. 4.2 method of analysis while previous analyses of the narrative genre have often focused on narrative structure and/or specific linguistic features (for example, biber & conrad, 2009; fleischman, 1990; labov & waletzky, 1967), our focus is on the text’s content, namely the three prototypical narrative elements 3 per publisher, texts for gh and gp were selected from the same geography textbook. sangers, evers-vermeul, sanders and hoeken 124 in the venn diagram in figure 1. for each text, excluding its heading, we manually coded whether these three elements were present or not. to expedite the coding process, we first analyzed whether a text contained one particularized event, describing a happening that took place only once, at one point in time and at one location. subsequently, we scored whether texts with one particularized event contained a second particularized event that was chronologically related to the first, forming a sequence of particularized events. for instance, the history text in (4), repeated here as (10), begins with the particularized event stating that john f. kennedy became president of the united states. this event is followed by that of his murder in 1963. (10) in the united states, a young, handsome president came to power in 1960: john f. kennedy. he had big plans for his country. the world was shocked when kennedy was murdered in 1963. he was shot dead while driving through the city of dallas in his open top car. the images of the murder were shown on television. (eigentijds, history grade 5, p. 50) in addition, a text contained an experiencing character if an individual was represented who was either taking active part in an event (“in the united states, a young, handsome president came to power in 1960: john f. kennedy”) or passively experiencing it (“he was shot dead”). this character could be human as well as human-like. groups and “tota pro partibus” (wholes for a part, for example, ‘the world’, indicating the world’s citizens) were not considered specific characters. in line with our discussion of narrative-like strategies that publishers may employ to make texts more concrete and imaginable, we also scored whether 1) a specific character was fictitious (vs. real), and 2) a representative entity was present in texts without a specific character, either being an individualized natural object, as in (8), or a human and/or an animal, as in (9). finally, a text contained a landscape of consciousness if thoughts, emotions, opinions, and/or wishes were represented explicitly. for instance, the history text in (11) gives insight into emperor charles v’s doctrinal ambition (“wanted”) and emotional state (“was afraid”). representations of an inner world that were imaginable but remained implicit in the text were not taken into account. a landscape of consciousness did not have to be linked to a specific character, but could also be expressed by groups or tota pro partibus. for instance, in (12), the emotions and thoughts of the common people are represented. similarly, an inner world did not have to be related to the leading character of the text; it was also scored for minor characters, as for william of orange’s cousin in (13). evaluations given by the author of the educational text were not taken into account (see also sangers, evers-vermeul & hoeken, submitted). (11) charles v wanted all his people to have the same faith. he was afraid of fights about the right faith. the catholic faith was the only thing that brought together all people in his great empire. (memo geschiedenis, history grade 8, p. 14) (12) after the last seconds of thursday october 4th 1582 had passed, it was suddenly october 15th in most parts of europe. this only happened because many european countries switched from the julian calendar to the gregorian calendar, yet it upset many people. riots broke out in some places because people thought that ten days of their lives had been stolen. it was said that migratory birds would not fly south in time. and the holy days had been shifted. would the saints understand what was going on? (geschiedenis werkplaats, history grade 5, p. 9) (13) william was eleven years old. his father was count of nassau in germany. that was what william would also become when he grew up. william had a cousin who was the prince of a tiny area in france. this cousin died. his will stated that he wanted william narrative elements in expository texts 125 to succeed him as prince of orange. and so young william suddenly inherited a very important title. (argus clou geschiedenis, history grade 5, p. 44) 4.3 inter-annotator agreement for considerations of reliability, 10 percent of the corpus (n=103) was coded by a second, independent annotator (cf. neuendorf, 2002). this sample was randomly compiled for each school subject and grade level, making use of the aselect()-function in excel. before the second annotator coded the sample, she engaged in a training phase to make her familiar with the procedure and the elements under investigation. the inter-annotator agreement was substantial to almost perfect (.741.00) (cf. landis & koch, 1977), as is shown in table 2. narrative element cohen’s kappa % agreement one particularized event .81 92 sequence of particularized events .93 97 experiencing character .88 95 landscape of consciousness .74 91 fictitious character 1.00 100 representative entity .86 93 table 2. inter-annotator agreement (cohen’s kappa and % agreement) per narrative element. the annotators discussed and resolved disagreements in their analyses to reach a final dataset. this was achieved without difficulty. the somewhat lower kappa-score for landscape of consciousness, for instance, was caused by the fact that not all thoughts, emotions, opinions, and wishes were easily recognizable. in (14), for example, the ambition of william of orange to reach freedom of religion was overlooked by one annotator as the representation of a belief/wish. (14) a number of southern regions, such as limburg, brabant, and zeeland flanders, now belonged to the republic. many catholics lived in these regions. the ideal of william of orange was for everyone to decide for themselves which religion they wanted to embrace. after the war, one could indeed decide on what to believe. in practice, however, it was much harder to be a catholic than to be a protestant. nevertheless, no one could be punished or arrested for his faith. and that is still the case today. (argus clou geschiedenis, history grade 5, p. 59) 4.4 method of statistical analysis the final dataset was analyzed using r version 3.6.1 (r core team, 2019). the analyses were completed via generalized linear mixed models, using the packages haven (wickham & miller, 2019), lme4 (bates, mächler, bolker & walker, 2015), emmeans (lenth, 2019), and ggplot2 (wickham, 2016). we added the fixed factors ‘subject’, ‘grade level’, and their interaction to the models in a stepwise manner. because some publishers did not design materials for all grade levels and/or school subjects under investigation, the statistical analysis did not allow for differentiation between publishers. to account for 1) potential differences in stylistic preferences between textbooks from different publishers and 2) correlations between texts for the two sub-domains of geography (which were selected from the same geography textbook, see also section 4.1), ‘textbook’ was modeled as a random factor. likelihood ratio tests were computed in order to assess which models fitted the data best. in section 5.2, we present the significant results of the best fitting models. see appendix b for an overview of all results. sangers, evers-vermeul, sanders and hoeken 126 5 results our analyses focused on the distribution of three prototypical narrative elements – particularized events, experiencing characters, and landscapes of consciousness – over dutch educational texts for different school subjects. we sketch a picture of the overall occurrence of the three narrative elements in our corpus (section 5.1), before we discuss the statistical analyses with respect to their distribution over biology, gp, gh, and history texts (section 5.2.1). subsequently, we consider the distribution of fictitious characters and representative entities (section 5.2.2), and discuss some qualitative observations regarding these strategies to make educational texts more concrete (section 5.3). 5.1 overall occurrence of prototypical narrative elements a sequence of particularized events was found in 138 texts of the corpus (n=999). however, due to singularity issues, the statistical model for this element could not be run without errors. since we were able to run the model for texts that contain (at least) one particularized event (282 texts), and its raw frequency pattern resembles that of the erroneous model, we report the results for texts with one particularized event here.4 we return to the theoretical and practical implications of this decision in the discussion section (section 6). beside 282 texts with a particularized event (28%), 293 texts contain an experiencing character (29%) and 253 texts a landscape of consciousness (25%). in total, 453 texts of the corpus contain one up to three of these narrative elements (45%), as opposed to 546 texts that are fully expository (55%). figure 2 represents the number of texts found for each combination of prototypical narrative elements, showing that full narratives are best represented (114; 11%), while texts that combine a particularized event and a landscape of consciousness without introducing a specific character are least frequent (21; 2%). figure 2. number of texts per combination of prototypical narrative elements (n=453). 4 of the 138 texts with two or more related events, 103 were found for history, 12 for biology, 10 for gp, and 13 for gh. a similar pattern was found for the 282 texts with one particularized event: 176 for history, 31 for biology, 37 for gp, and 39 for gh. landscape of consciousness experiencing character particularized events 60 53 114 87 21 39 79 narrative elements in expository texts 127 5.2 statistical analyses in this section, we first describe the statistical analyses for prototypical narrative elements (section 5.2.1). subsequently, we discuss the statistical analyses for fictitious characters and representative entities (section 5.2.2). 5.2.1 prototypical narrative elements we first analyzed whether the distribution of educational texts with one up to three prototypical narrative elements over the corpus was influenced by the fixed factors subject, grade level, and/or their interaction. for this analysis, which included all 453 texts categorized in figure 2, the best fitting model was the model in which only subject was entered as a fixed factor (χ2(3)=43.56, p<.001). as hypothesized, a post hoc tukey pairwise comparison test revealed that prototypical narrative elements are more frequent in history texts than in texts for biology, gp, and gh (all p’s<.001).5 these results are visualized in figure 3. figure 3. predicted probability for texts with 1-3 types of narrative elements. subsequently, we analyzed whether this pattern persisted for each individual prototypical narrative element (that is, the three autonomous circles). for each element, the same pattern was found: the model in which only the fixed factor subject was entered fitted the data best (particularized event: χ2(3)=23.98, p<.001; experiencing character: χ2(3)=24.63, p<.001; landscape of consciousness: χ2(3)=35.52, p<.001). post hoc tukey tests showed that all three prototypical narrative elements are more frequent in history texts than in biology, gp, and gh texts (all p’s<.001), as visualized in figure 4-6. figure 4. predicted probability for particularized event. 5 for the complete tukey results, see appendix b. sangers, evers-vermeul, sanders and hoeken 128 figure 5. predicted probability for experiencing character. figure 6. predicted probability for landscape of consciousness. finally, we analyzed whether the pattern persisted in full narratives, combining the three narrative elements. once again, the model in which subject was entered as a fixed factor was the best fitting model (χ2(3)=32.05, p<.001). following the pattern, a post hoc tukey test showed that full narratives are more frequent in history texts than in texts for biology, gp, and gh (all p’s<.001), as visualized in figure 7. none of the analyses revealed an effect for grade level or an interaction effect of grade level and subject. figure 7. predicted probability for full narratives.6 6 the error bar for gh grade 5 texts extends from 0 to 1 because no full narratives were found in this condition. narrative elements in expository texts 129 5.2.2 fictitious characters and representative entities of the 293 texts with an experiencing character, the majority include a character who exists or existed in the real world (228 texts, 78%), while only 65 texts introduce a fictitious character (22%). for texts with a fictitious character, the model in which only subject was entered as a fixed factor was the best fitting model (χ2(3)=20.38, p<.001). as hypothesized, the pattern for fictitious characters was opposed to the one found for prototypical narrative elements: a post hoc tukey test revealed that fictitious characters are less frequent in history texts than in texts for biology (p=.008), gp (p<.001), and gh (p<.001), as visualized in figure 8. this indicates that if experiencing characters are absent in the educational content – as is often the case in biology, gp, and gh texts –, publishers occasionally add fictitious characters to the text, while they generally do not apply this strategy if specific characters are at the core of the to-be-learned information, as is the case in most history texts. figure 8. predicted probability for fictitious experiencing characters. in addition, for texts with a representative entity (225 texts, 23%), the model with the fixed factors subject (χ2(3)=51.52, p<.001), grade level (χ2(1)=5.91, p=.015), and their interaction (χ2(3)=13.25, p=.004) was the best fitting model. a post hoc tukey test revealed that in both grade 5 and grade 8 texts, representative entities are more frequent in biology texts than in texts for gp, gh, and history (all p<.001). in addition, in texts for grade 8 – but not for grade 5 –, representative entities are more frequent in gp texts than in gh (p<.001) and history texts (p=.002). these results, which are visualized in figure 9, indicate that while experiencing characters are not so much at the heart of the educational content in biology and gp texts, these texts are occasionally made more specific by the application of representative entities. furthermore, the post hoc tukey test showed that for gp texts – but not the other subjects – representative entities are less frequent in grade 5 than in grade 8 (p=.017). sangers, evers-vermeul, sanders and hoeken 130 figure 9. predicted probability for representative entities. 5.3 making educational texts more concrete: some qualitative observations since we were also interested in the ways in which the strategies to make educational texts more concrete are qualitatively elaborated upon in educational texts for different school subjects, we made some additional observations. for instance, we observed that in biology, gp, and gh texts, fictitious characters are placed in a contemporary context, often representing a peer, such as eva in (7) or the german teenager matthias in (15), or an individual belonging to a certain professional group, such as gynecologist loes in (16). by contrast, in history texts, fictitious characters are situated in a historical context, often representing the “common man” experiencing the events of his time, such as prehistoric man iugas in (2) or merchant pieter in (17). the three fictitious characters in (15)-(17) introduce themselves, mentioning their name, and share their personal stories, interlaced with to-be-learned information. (15) guten tag! i am matthias sammer. i live in berlin and i enjoy being your guide while you are exploring my country. germany is a big country in europe. the netherlands fits into it almost nine times and we have five times as many inhabitants. (buitenland, gh grade 8, p.17) (16) hello, my name is loes. i work as a gynecologist and i am involved in pregnancies and deliveries that are not going well. in the netherlands, many babies are born at home under the supervision of a midwife. however, giving birth at home can also involve too much risk. (biologie voor jou, biology grade 8, p. 183) (17) i am pieter ysenbouts from antwerp. i buy and sell spices from the east: nutmeg, pepper, cloves, cinnamon... the commercial ships bring me everything. maybe i will travel on a such a ship myself soon. (tijdzaken, history grade 5, p. 77) for educational content about natural phenomena, we observed a similar distinction along the lines of representative entities (a baby) versus exemplar characters (baby thomas) to concretize generic to-be-learned information about human or human-like entities (all babies). compare, for instance, the geography texts in (18) and (19), which originate from the same textbook. in (18), the general process of the formation/eruption of stratovolcanoes is explained by the description of a representative instance, focusing on individualized natural entities. however, this instance does not link one-on-one to an actual event in the real world. by contrast, (19) discusses a real exemplar of a stratovolcano eruption, namely that of the soufrière on montserrat. as opposed to (18), the time narrative elements in expository texts 131 and place of the volcano eruption are explicated in (19). in addition, while (18) focuses on the technical aspects of the volcano eruption, (19) rather concentrates on its consequences for the human population. (18) at a convergent boundary, a heavy earth plate slides beneath a lighter earth plate. the heavy plate disappears deeper and deeper into the mantle, causing the stones to melt. the resulting magma wants to rise again. […] eventually the pressure becomes too high and the volcano erupts explosively. the tough lava quickly solidifies. this creates a volcano with steep slopes: a stratovolcano. (de wereld van, gp grade 8, p. 54-55) (19) a dormant volcano can come back to life. this happened not so long ago with the soufrière volcano on montserrat, a british island located 150 kilometers from saba, belonging to the same island arc. like mount scenery, the soufrière had been quiet for four hundred years. but on july 18, 1995, the volcano erupted. because there had been earthquakes since 1992, the volcano was closely monitored. this allowed the population to be brought to safety in time for the eruption. about 7,000 residents were evacuated to surrounding islands and to the united kingdom. the nineteen deaths on the island were people who had ignored the warnings. the capital of montserrat was wiped off the map, only the north was still habitable and tourism did completely collapse. (de wereld van, gp grade 8, p. 49-50) the two texts evoke a different learning approach: while (18) presents a hypothetical, individualized instance of generic to-be-learned information to which real exemplars could be linked via deduction, (19) provides a real exemplar to which more generic educational content can be associated via induction. as can be inferred from the texts’ page numbers, the publishers of the gp textbook have chosen to place the exemplar situation before the representative instance, proceeding from specific to somewhat more generic educational content. 6 discussion and conclusion in this quantitative corpus-based study, we charted how and when prototypical narrative elements, namely particularized events, experiencing characters, and landscapes of consciousness, are being used in present-day dutch educational texts. more specifically, we analyzed 1) the frequency with which various combinations of these narrative elements are used in dutch educational texts, and 2) the extent to which they are applied differently in texts for the school subjects biology, physical geography, human geography, and history. our findings are summarized in table 3. sangers, evers-vermeul, sanders and hoeken 132 narrative elements number of texts (n=999) % of texts corpus significant patterns one or more types of narrative elements 453 45% hi>bi=gp=gh particularized event (pe) 282 28% hi>bi=gp=gh experiencing character (ec) 293 29% hi>bi=gp=gh landscape of consciousness (loc) 253 25% hi>bi=gp=gh all three narrative elements 114 11% hi>bi=gp=gh pe + ec 87 9% pe + loc 21 2% ec + loc 39 4% fictitious characters 65 7% higp=gh=hi 8: bi>gp>gh=hi gp: 5<8 table 3. summary of the quantitative findings. our results demonstrate that even in a domain that has traditionally been considered the prime exemplar of the expository genre, narrative elements are found quite frequently: 45 percent of the educational texts in our corpus are categorized in one of the narrative areas of the model in figure 1, indicating that only 55 percent of the texts in the corpus are fully expository. of the three prototypical narrative elements, an experiencing character is most frequently found (n=293), closely followed by a particularized event (n=282) and a landscape of consciousness (n=253). when two narrative elements are combined, the combination of a particularized event and an experiencing character is most common (n=87). in addition, our statistical results show that the occurrence of narrative elements in educational texts depends on the nature of the to-be-learned information. as we hypothesized, particularized events, experiencing characters, and landscapes of consciousness are more frequent in history texts than in biology, gp, and gh texts. this substantiates our reasoning that the educational content of the school subjects under investigation differs in their focus on narrativity – with specific events, specific people, and their experiences often being at the heart of history texts, and a focus on general tendencies in biology, gp, and gh texts. this indicates that narrative elements are less frequently applied in contexts in which they fit less naturally. in fact, when prototypical narrative elements are used in biology, gp, and gh texts, they are generally deliberate interventions. following our observation that biology, gp, and gh texts are less often about specific events, specific characters, and specific contexts, we argued that these texts would also tend to be more abstract and less imagery-provoking than history texts. given the theoretical and empirical narrative elements in expository texts 133 evidence underlining the processing advantage of concrete information over abstract information (cf. sadoski, 2001), we reasoned that biology, gp, and gh texts would benefit from strategies to make educational texts more concrete, such as the addition of fictitious characters and representative entities. indeed, the distribution pattern of texts with a fictitious character is opposite to that of texts with prototypical narrative elements: fictitious characters are less frequently added to history texts than to texts for the other school subjects. this shows that while specific characters tend to be real historical people in history texts, those is biology, gp, and gh texts are almost always fictitious inventions. it is, however, important to acknowledge that the deliberate intervention of fictitious characters was not so frequent in itself, being employed in only 7 percent of the texts in the corpus (4% of biology texts, 9% of gp texts, 11% of gh texts, 3% of history texts), compared to the occurrence of real experiencing characters in 23 percent of the texts in the corpus (10% of biology texts, 4% of gp texts, 1% of gh texts, 66% of history texts). the other strategy – the incorporation of representative entities – was applied more frequently, namely in 23 percent of the texts in the corpus (67% of biology texts, 17% of gp texts, 7% of gh texts, 4% of history texts). in line with our hypotheses, the distribution pattern of representative entities differs from that for texts with prototypical narrative elements: while the latter elements are not so much at the heart of the educational content in biology and gp texts, publishers make use of representative entities to make these texts more concrete and imaginable. interestingly, representative entities were not found more frequently in grade 5 gp texts and in gh texts than in history texts. while we generally found robust effects for the distribution of narrative elements over school subjects, our distinction between two sub-domains of geography does not seem to be relevant with respect to the distribution of narrative elements. that is, contrary to our line of reasoning, our results indicate that gh texts (human-related, generic) do not occupy an intermediate position between history texts (human-related, specific) and biology and gp texts (non-human-related, generic): they rather side with the latter two subjects. this means that even though gh texts discuss humanrelated topics – thereby offering more options for the inclusion of narrative elements –, their content remains as generic as that of texts for less human-related subjects. in addition, we did not find differences in the distribution of narrative elements over different grade levels, indicating that our results are generalizable over grade levels. only one effect of grade level was found, which concerned just one school subject.7 the present research has provided insight into the distribution of narrative elements over dutch educational texts for different school subjects. however, there are some limitations to this study that give rise to theoretical implications as well as new directions for future research. a first limitation to this study is that we were unable to run the statistical model for a sequence of particularized events, which is why we decided to report the results on the model for one particularized event instead. there is some theoretical basis for this decision. in the narratological literature, there has been a debate on the minimal criteria required for texts to be categorized as ‘narrative’, particularly the criterium whether a single event suffices for a narrative (abbott, 2008; genette, 1982), or whether at least two events are needed (bal, 1997; labov, 1972; onega & landa, 1996; prince, 2003; richardson, 1997; rimmon-kenan, 2002; sanford & emmott, 2012). in our application of narrativity to the educational domain, we adopted the second definition: a narrative educational text is formed by a sequence of two or more non-randomly related events. this strict interpretation excludes educational texts with only one particularized event from a classification as narrative texts. however, following ryan’s (2007) scalar interpretation of ‘narrative’, and introducing a model that allows for hybrid forms of narrativity – containing only one or two prototypical narrative elements –, we could also imagine a scalar interpretation within each prototypical narrative element. following such an interpretation, educational texts with one 7 exemplars were less frequent in grade 5 gp texts than in grade 8 gp texts. sangers, evers-vermeul, sanders and hoeken 134 particularized event would merely be less-pronounced narrative than educational texts that contain a sequence of particularized events – presenting more of a narrative scene to concretize the educational content rather than a complete story. such a lenient interpretation would allow for nuances within each prototypical narrative element, giving rise to fruitful directions for follow-up research. in addition, educational publishers might use particularized events not with the purpose of making educational texts more narrative but with the sole aim of making it more concrete and imaginable. this raises the question as to how narrativity and concreteness relate to each other: does a higher degree of narrativity always lead to more concreteness, and vice versa? following the literature, the answer to the first part of this question seems to be positive: nisbett and ross (1980) have argued that the more detailed information a text provides about specific events, specific characters, and the context, the more concrete and imagery-provoking this text is. this suggests that irrespective of a strict or lenient interpretation of the narrativity of events, adding a particularized event to an educational text enhances its concreteness and imaginability. however, it is not clear whether a higher degree of concreteness also brings about more narrativity. on the basis of the strict interpretation, the use of one particularized event instead of a sequence of particularized events would not make an educational text more narrative – the event, however, does make the text more concrete. this suggests that the relationship between narrativity and concreteness might not be one-to-one. if publishers are primarily concerned with making their texts more concrete, it would not be surprising if they added just one particularized event to the educational text. the lenient interpretation, on the other hand, does not preclude a one-to-one relation between the two concepts: the addition of one particularized event would make the educational text more concrete and somewhat – but not fully – narrative. determining the precise relationship between narrativity and concreteness would be valuable for further theoretical development. in this respect, it would be worthwhile to discuss our findings with publishers and authors of educational materials. what goals do they pursue with the application of narrative elements in their texts, and how do they see the relation with concreteness? do publishers distinguish between narrativity and concreteness, or do they consider these concepts as intertwined? what kinds of design principles do they formulate to make their texts more narrative and/or more concrete? earlier research has shown that interviews can be fruitful in discovering what educational professionals consider important textual elements, and how they adapt their design principles accordingly (land, sanders, lentz & van den bergh, 2002). a second limitation to this study is that we only analyzed two narrative-like strategies publishers can employ to increase the concreteness and imaginability of educational texts that are less inherently focused on specific events, characters, and contexts. of course, other strategies may also be fruitful in conveying relevant content in a comprehensible and interesting way. one strategy, for instance, could be the use of “voice” in educational texts (cf. beck, mckeown & worthy, 1995; sangers et al., submitted), establishing an interaction between the author of the educational text and students to bridge the gap between students and the educational content they need to learn. third, in our study, we focused on a specific domain, namely dutch educational texts. to be able to interpret our quantitative results in a broader context, it would be worthwhile to compare the use of narrative elements in dutch educational texts to that of texts in other domains. a fruitful domain for comparison could, for instance, be journalism, as news texts also tend to convey new information, while often discussing particularized events and introducing ‘characters’ (for example, eyewitnesses), who may express their feelings and/or thoughts about the happenings described (cf. van krieken & sanders, 2016). furthermore, it would be valuable to broaden the focus of the current research by including texts from other cultures. while most – if not all – cultures define ‘narrative’ as a genre, its exact interpretation within the educational domain (as well as any other domain) may differ from culture to culture (cf. biber & conrad, 2009). therefore, in order to narrative elements in expository texts 135 uncover any cultural differences with respect to narrativity in educational texts, it would be fruitful to compare our results on narrative elements in dutch educational texts to that of educational texts from other countries. fourth, within the domain of dutch education, we focused on educational texts for a limited set of school subjects, namely biology, geography, and history. in future research, it would be interesting to expand the current research by examining the distribution of narrative elements over additional school subjects. for instance, how and to what extent are narrative elements applied in school subjects that focus on formulas instead of humanor nature-related topics, such as mathematics? finally, future research could focus on the presumed rationale behind including narrative elements in educational texts: are narrative educational texts indeed more interesting to read and easier to understand than expository educational texts? previous experimental research has shown conflicting results in this respect (see also section 1). however, since these studies have used different operationalizations of the notions “expository” and “narrative”, no firm conclusions can yet be drawn about the relative effectiveness of narrative versus expository educational texts (sangers et al., 2019). therefore, future research should pursue a more consistent approach to the manipulation of these genres in experimental texts. we believe that our model in figure 1 offers valuable guidelines to make a fair distinction between fully expository educational texts, fully narrative educational texts, and hybrid educational texts, combining characteristics of both genres. taken together, we have demonstrated that narrative elements are quite common in educational texts – a domain that has traditionally been considered as the prime exemplar of the expository genre. in addition, we have shown that the occurrence of narrative elements in educational texts depends on the nature of the educational content: 1) prototypical narrative elements tend to be at the core of the to-be-learned information in history texts, while 2) strategies to make educational texts more concrete, such as the addition of fictitious characters and representative entities, are more frequently applied in school subjects in which prototypical narrative elements fit less naturally, namely biology and geography texts. this way, the current paper has given insight into the use and distribution of narrative elements in educational texts, and has provided an essential step to investigate the potential of narrative elements in educational texts further – with the ultimate aim of designing optimal texts that present relevant educational content in a comprehensible and attractive way. acknowledgements this work was supported by the sustainable humanities program of the netherlands organization for scientific research (nwo) [project number pgw.17.035]. sangers, evers-vermeul, sanders and hoeken 136 appendix a – materials biology grade 5 mart ottenheim and rein tromp (eds.) (2011). natuniek: natuur en techniek voor het basisonderwijs groep 7 leerlingenboek. thiememeulenhoff, amersfoort. ferry siemensma (2014). wijzer! natuur & techniek: leerwerkboek groep 7. noordhoff uitgevers, groningen/houten. annet talsma and linda vogelesang (eds.) (n.d.). natuurzaken: werkboek jaargroep 7 (4th ed.). uitgeverij zwijsen, tilburg. maaike van riel and leonoor soet (eds.) (2012). argus clou professor in alles: natuur en techniek groep 7 lesboek. malmberg, ’s-hertogenbosch. carla wiechers (ed.) (2014). binnenstebuiten: natuur en techniek bronnenboek groep 7. blink educatie, ’s-hertogenbosch. biology grade 8 trijnie akkerman (ed.) (2013). nectar: 2-3 vwo leerboek (4th ed.). noordhoff uitgevers, groningen/houten. arteunis bos, onno kalverda, ruud passier, hans rawee, rik smale, gerard smits and ben waas (2015). biologie voor jou: handboek 2a vwo/gymnasium (7th ed.). malmberg, ’shertogenbosch. geography grade 5 anton bakker (ed.) (2012). de blauwe planeet: aardrijkskunde voor het basisonderwijs. thiememeulenhoff, amersfoort. anja huisman (ed.) (2012). argus clou professor in alles: aardrijkskunde groep 7 lesboek. malmberg, ’s-hertogenbosch. ferry siemensma (2015). wijzer! aardrijkskunde: leerwerkboek groep 7. noordhoff uitgevers, groningen/houten. annet talsma (ed.) (2014). wereldzaken: werkboek jaargroep 7 (3rd ed.). uitgeverij zwijsen, tilburg. marijke van ooijen (ed.) (2014). grenzeloos: aardrijkskunde bronnengroep groep 7. blink educatie, ’s-hertogenbosch. geography grade 8 daphne ariaens, wim b. brinke, chris de jong and jan h.a. padmos (eds.) (2016). de geo: aardijkskunde voor de onderbouw 2 vwo (9th ed.). thiememeulenhoff, amersfoort. martin van de ven (ed.) (2018). de wereld van: aardrijkskunde voor de onderbouw leeropdrachtenboek 2vg. malmberg, ’s-hertogenbosch. geert van den berg (ed.) (2014). buitenland 2 vwo (3rd ed.). noordhoff uitgevers, groningen/houten. history grade 5 millicent kruis (2014). wijzer! geschiedenis: leerwerkboek groep 7. noordhoff uitgevers, groningen. annemarie können (ed.) (2012). argus clou professor in alles: geschiedenis groep 7 lesboek. malmberg, ’s-hertogenbosch. jouke nijman and hans roest (2011). speurtocht 7: geschiedenis voor het basisonderwijs: leerlingenboek (2nd ed.). thiememeulenhoff, amersfoort.. miranda van de mortel, marcel van den oever, hans vermeer and linda vogelesang (n.d.). tijdzaken: werkboek jaargroep 7 (4th ed.). uitgeverij zwijsen, tilburg. narrative elements in expository texts 137 carla wiechers (ed.) (2014). eigentijds: geschiedenis bronnenboek groep 7. blink educatie, ’shertogenbosch. history: grade 8 leo salemink and jos venner (eds.) (2010). feniks: geschiedenis voor de onderbouw leesboek 2 vwo. thiememeulenhoff, amersfoort. wieke schrover and judith tadema (eds.) (2015). memo: geschiedenis voor de onderbouw 2 vwo handboek (4th ed.). malmberg, ’s-hertogenbosch. tom van der geugten and dik verkuil (eds.) (2013). geschiedenis werkplaats: 2 vwo informatieboek. (2nd ed.). noordhoff uitgevers, groningen/houten. sangers, evers-vermeul, sanders and hoeken 138 appendix b – statistical results 1. generalized linear mixed models8 fictitious character -2ll χ2 df p model 0 169.8 *model 1 (+subject) 149.4 20.38 3 <.001 model 2 (+level) 148.3 1.10 1 .294 model 3 (+subject*level) 141.2 7.08 3 .069 representative entity -2ll χ2 df p model 0 770.2 model 1 (+subject) 718.7 51.52 3 <.001 model 2 (+level) 712.7 5.91 1 .015 *model 3 (+subject*level) 699.5 13.25 3 .004 8 the asterisk indicates the model that was proven to be the best fitting model. full venn -2ll χ2 df p model 0 1112.7 *model 1 (+subject) 1069.1 43.56 3 <.001 model 2 (+level) 1068.9 0.21 1 .645 model 3 (+subject*level) 1067.7 1.19 3 .755 particularized event -2ll χ2 df p model 0 990.7 *model 1 (+subject) 966.7 23.98 3 <.001 model 2 (+level) 966.7 0.04 1 .844 model 3 (+subject*level) 964.6 2.05 3 .561 experiencing character -2ll χ2 df p model 0 862.5 *model 1 (+subject) 837.8 24.63 3 <.001 model 2 (+level) 837.7 0.17 1 .68 model 3 (+subject*level) 834.5 3.21 3 .36 landscape of consciousness -2ll χ2 df p model 0 1014.2 *model 1 (+subject) 978.7 35.52 3 <.001 model 2 (+level) 978.2 0.53 1 .468 model 3 (+subject*level) 973.0 5.15 3 .161 full narrative -2ll χ2 df p model 0 571.4 *model 1 (+subject) 534.3 32.05 3 <.001 model 2 (+level) 536.7 1.63 1 .202 model 3 (+subject*level) 532.1 4.58 3 .205 narrative elements in expository texts 139 2. predicted probability scores particularized event probability se lcl ucl biology 0.10 0.03 0.04 0.22 geography – physical 0.15 0.04 0.07 0.29 geography – human 0.15 0.04 0.07 0.28 history 0.62 0.07 0.44 0.77 fictitious character probability se lcl ucl biology 0.46 0.20 0.10 0.86 geography – physical 0.64 0.18 0.21 0.92 geography – human 0.91 0.09 0.44 0.99 history 0.02 0.02 0.003 0.13 full venn probability se lcl ucl biology 0.23 0.05 0.13 0.36 geography – physical 0.25 0.05 0.15 0.38 geography – human 0.35 0.05 0.23 0.49 history 0.86 0.03 0.73 0.92 experiencing character probability se lcl ucl biology 0.10 0.04 0.04 0.26 geography – physical 0.10 0.04 0.04 0.24 geography – human 0.08 0.03 0.03 0.21 history 0.70 0.08 0.47 0.85 landscape of consciousness probability se lcl ucl biology 0.13 0.03 0.08 0.22 geography – physical 0.10 0.02 0.06 0.18 geography – human 0.17 0.03 0.10 0.26 history 0.51 0.04 0.40 0.62 full narrative probability se lcl ucl biology 0.03 0.02 0.010 0.10 geography – physical 0.03 0.01 0.009 0.09 geography – human 0.01 0.01 0.001 0.04 history 0.31 0.06 0.186 0.47 sangers, evers-vermeul, sanders and hoeken 140 representative entity probability se lcl ucl grade 5 biology 0.61 0.07 0.43 0.77 geography – physical 0.06 0.02 0.02 0.17 geography – human 0.08 0.03 0.03 0.20 history 0.04 0.02 0.01 0.13 grade 8 biology 0.82 0.06 0.58 0.93 geography – physical 0.29 0.07 0.14 0.52 geography – human 0.06 0.03 0.07 0.18 history 0.04 0.02 0.01 0.13 3. post hoc tukey scores bi = biology 5 = grade 5 gp = physical geography 8 = grade 8 gh = human geography hi = history full venn particularized event experiencing character contrasts or se z p bi / gp 0.89 0.32 -0.33 .988 bi / gh 0.56 0.20 -1.66 .343 bi / hi 0.05 0.02 -8.16 <.001 gp / gh 0.63 0.13 -2.26 .107 gp / hi 0.06 0.02 -8.13 <.001 gh / hi 0.09 0.03 -6.94 <.001 contrasts or se z p bi / gp 0.63 0.31 -0.94 .784 bi / gh 0.65 0.32 -0.89 .812 bi / hi 0.07 0.03 -5.56 <.001 gp / gh 1.03 0.26 0.10 1.00 gp / hi 0.11 0.05 -4.99 <.001 gh / hi 0.11 0.05 -5.06 <.001 contrasts or se z p bi / gp 1.01 0.61 0.02 1.00 bi / gh 1.24 0.75 0.36 .984 bi / hi 0.05 0.03 -5.15 <.001 gp / gh 1.23 0.36 0.72 .891 gp / hi 0.05 0.03 -5.34 <.001 gh / hi 0.04 0.02 -5.66 <.001 narrative elements in expository texts 141 landscape of consciousness full narrative fictitious character contrasts or se z p bi / gp 0.47 0.52 -0.68 .906 bi / gh 0.09 0.11 -1.89 .232 bi / hi 39.48 45.35 3.20 .008 gp / gh 0.18 0.17 -1.80 .274 gp / hi 84.03 90.09 4.13 <.001 gh / hi 461.55 586.07 4.83 <.001 representative entity contrasts or se z p 5bi / 5gp 23.10 11.63 6.25 <.001 5bi / 5gh 18.20 8.79 6.00 <.001 5bi / 5hi 37.40 20.56 6.59 <.001 5gp / 5gh 0.79 0.37 -0.51 1.00 5gp / 5hi 1.62 1.01 0.77 .995 5gh / 5hi 2.06 1.25 1.18 .937 8bi / 8gp 10.60 5.81 4.30 <.001 8bi / 8gh 71.10 45.27 6.70 <.001 8bi / 8hi 103.00 65.55 7.27 <.001 8gp / 8gh 6.72 2.99 4.28 <.001 8gp / 8hi 9.71 5.66 3.90 .002 8gh / 8hi 1.45 0.96 0.55 .999 5bi / 8bi 0.36 0.18 -2.03 .459 5gp / 8gp 0.16 0.09 -3.37 .017 5gh / 8gh 1.39 0.85 0.54 1.00 5hi / 8hi 0.98 0.65 -0.04 1.00 contrasts or se z p bi / gp 1.37 0.48 0.91 .799 bi / gh 0.78 0.25 -0.77 .870 bi / hi 0.15 0.04 -6.34 <.001 gp / gh 0.57 0.16 -2.07 .163 gp / hi 0.11 0.03 -7.20 <.001 gh / hi 0.19 0.05 -5.91 <.001 contrasts or se z p bi / gp 1.12 0.71 0.81 .998 bi / gh 4.99 4.39 1.83 .260 bi / hi 0.07 0.04 -4.71 <.001 gp / gh 4.45 3.56 1.87 .241 gp / hi 0.07 0.04 -5.08 <.001 gh / hi 0.01 0.01 -5.15 <.001 sangers, evers-vermeul, sanders and hoeken 142 references h. porter abbott (2008). the cambridge introduction to narrative. cambridge university press, new york. mieke bal (1997). narratology: introduction to the theory of narrative. university of toronto press, toronto. douglas bates, martin mächler, ben bolker and steve walker (2015). fitting linear mixed-effects models using lme4. journal of statistical software, 67(1):1-48. isabel i. beck, margaret g. mckeown and jo worthy (1995). giving a text voice can improve students’ understanding. reading research quarterly, 30(2):220-238. douglas biber and susan conrad (2009). register, genre, and style. cambridge university press, new york. jerome bruner (1986). actual minds, possible worlds. harvard university press, cambridge, massachusetts, london. gina n. cervetti, marco a. bravo, elfrieda h. hiebert, p. david pearson and carolyn a. jaynes (2009). text genre and science content: ease of reading, comprehension, and reader preference. reading psychology, 30(6):487-511. daniel chandler (1997). an introduction to genre theory. available online at: http://www.aber.ac.uk/media/documents/intgenre/chandler_genre_theory.pdf. frances christie and james r. martin (2000). genre and institutions: social processes in the workplace and school. continuum, london. committee meijerink (2009). referentiekader taal en rekenen: de referentieniveaus [framework of reference for language and arithmetic: the reference levels]. slo, enschede. christel dood, joyce gubbels and eliane segers (2020). pisa-2018 de verdieping: leesplezier, zelfbeeld bij het lezen, leesgedrag en leesvaardigheid en de relatie daartussen [pisa-2018: the deepening: enjoyment of reading, self-perception when reading, reading behavior and reading comprehension and the relationship between them]. expertisecentrum nederlands, nijmegen. allang eng (2002). learning and processing nonfiction expository and narrative genre. phd thesis, university of toronto. suzanne fleischman (1990). tense and narrativity: from medieval performance to modern fiction. routledge, london. monika fludernik (2009). an introduction to narratology. routledge, london/new york. gérard genette (1982). figures of literary discourse (alan sheridan, trans.). columbia university press, new york. joyce gubbels, andrea netten and ludo verhoeven (2017). vijftien jaar leesprestaties in nederland: pirls-2016 [fifteen years of reading achievement in the netherlands: pirls2016]. expertisecentrum nederlands, nijmegen. joyce gubbels, annemarie m.l. van langen, nathalie a.m. maassen and martina r.m. meelissen (2019). resultaten pisa-2018 in vogelvlucht [results pisa-2018 in a bird’s eye view]. university of twente, enschede. david herman (2009). basic elements of narrative. wiley-blackwell, malden, massachusetts. inspectorate of education (2017). de staat van de leerling [the state of the student]. inspectorate of education, utrecht. inspectorate of education (2020). de staat van het onderwijs 2020 [the state of education]. inspectorate of education, utrecht. inspectorate of education (2021). de staat van het onderwijs 2021 [the state of education]. inspectorate of education, utrecht. william labov (1972). language in the inner city: studies in the black english vernacular. university of pennsylvania press, philadelphia, pennsylvania. http://www.aber.ac.uk/media/documents/intgenre/chandler_genre_theory.pdf narrative elements in expository texts 143 william labov and joshua waletzky (1967). narrative analysis: oral versions of personal experience. in j. helm (ed.), essays on the verbal and the visual art: proceedings of the 1996 annual meeting of the american ethnological society: 12-44. university of washington press, seattle, washington. jentine land, ted sanders, leo lentz and huub van den bergh (2002). tekstbegrip en tekstwaardering op het vmbo: welke tekstkenmerken dragen bij aan de kwaliteit van studieboekteksten? [text comprehension and appreciation in lower secondary professional education: which textual features contribute to the quality of educational texts?]. stichting lezen/utrecht university, amsterdam/utrecht. j. richard landis and gary g. koch (1977). the measurement of observer agreement for categorical data. biometrics, 33(1):159-174. david y.w. lee (2001). genres, registers, text types, domains, and styles: clarifying the concepts and navigating a path through the bnc jungle. language learning & technology, 5(3):37-72. russell lenth (2019). emmeans: estimated marginal means, aka least-squares means (version 1.4.1) [computer software]. available online at https://cran.rproject.org/package%20= emmeans. james r. martin and david rose (2008). genre relations: mapping culture. equinox, london. gregory l. murphy (2016). is there an exemplar theory of concepts? psychonomic bulletin & review, 23:1035-1042. kimberly a. neuendorf (2002). the content analysis handbook. sage publications, thousand oaks, california. richard nisbett and lee ross (1980). human inference: strategies and shortcomings of social judgments. prentice-hall, englewood cliffs, new jersey. stephen p. norris, sandra m. guilbert, martha l. smith, shahram hakimelahi and linda m. phillips (2005). a theoretical framework for narrative explanation in science. science education, 89(4):535-563. susana onega and josé a.g. landa (1996). narratology: an introduction. longman publishing, new york. allan paivio (1971). imagery and verbal processes. holt, rinehart and winston, new york. allan paivio (1986). mental representations: a dual coding approach. oxford university press, oxford. gerald prince (2003). a dictionary of narratology (2nd ed.). university of nebraska press, lincoln, nebraska. r core team (2019). r: a language and environment for statistical computing. r foundation for statistical computing (version 3.6.1) [computer software]. available online at https://www.rproject.org/. brian richardson (1997). unlikely stories: causality and the nature of modern narrative. university of delaware press, newark/london. shlomith rimmon-kenan (2002). narrative fiction: contemporary poetics (2nd ed.). routledge taylor and francis group, london/new york. fernando romero, scott g. paris and sarah k. brem (2005). children’s comprehension and localto-global recall of narrative and expository texts. current issues in education, 8(25):1-11. marie-laure ryan (2007). toward a definition of narrative. in d. herman (ed.), the cambridge companion to narrative: 22-35. cambridge university press, cambridge. mark sadoski (1991). theoretical, empirical and practical considerations in designing informational text. document design, 1(1):24-24. mark sadoski (2001). resolving the effects of concreteness on interest, comprehension, and learning important ideas from text. educational psychology review, 13(3):263-281. mark sadoski and allan paivio (1994). a dual coding view of imagery and verbal processes in reading comprehension. in r.b. ruddell, m. rapp ruddell and h. singer (eds.), theoretical https://cran.rproject.org/package%20=%20emmeans https://cran.rproject.org/package%20=%20emmeans https://www.r-project.org/ https://www.r-project.org/ sangers, evers-vermeul, sanders and hoeken 144 models and processes of reading: 582–601. international reading association, newark, delaware. mark sadoski, allan paivio and ernest t. goetz (1991). a critique of schema theory in reading and a dual coding alternative. reading research quarterly, 26(4):463-484. anthony j. sanford and catherine emmott (2012). mind, brain, and narrative. cambridge university press, new york. nina l. sangers, jacqueline evers-vermeul, ted j.m. sanders and hans hoeken (2019). effecten van narrativiteit in educatieve teksten: wat zeggen onderzoeksresultaten (nog niet)? [effects of narrativity in educational texts: what do research results tell (not yet)?] tijdschrift voor taalbeheersing, 41(3):433-460. nina l. sangers, jacqueline evers-vermeul, ted j.m. sanders and hans hoeken (2020). vivid elements in dutch educational texts. narrative inquiry, 30(1):186-210. nina l. sangers, jacqueline evers-vermeul and hans hoeken (submitted). addressing the student: voice elements in educational texts. marina santini (2006). interpreting genre evolution on the web: preliminary results. in j. karlgren (ed.), proceedings of the workshop on new text: wikis and blogs and other dynamic text sources: 32-39. european association of computational linguistics, trento. michael toolan (2001). narrative: a critical linguistic introduction (2nd ed.). routledge, london. kobie van krieken and josé sanders (2016). framing narrative journalism as a new genre: a case study of the netherlands. jounalism, 18(10), 1-17. gerdineke van silfhout (2014). fun to read or easy to understand? establishing effective text features for educational texts on the basis of processing and comprehension research. phd thesis, lot, utrecht. available online at https://dspace.library.uu.nl/handle/1874/300805. hadley wickham (2016). ggplot2: elegant graphics for data analysis [computer software]. retrieved from: https://ggplot2.tidyverse.org. hadley wickham and evan miller (2019). haven: import and export “spss”, “stata” and “sas” files (version 2.1.1) [computer software]. retrieved from: https://cran.rproject.org/package=haven. https://dspace.library.uu.nl/handle/1874/300805 https://ggplot2.tidyverse.org/ https://cran.r-project.org/package=haven https://cran.r-project.org/package=haven dialogue & discourse 13(1) (2022) 63–95 doi: 10.5210/dad.2022.103 gailbot: an automatic transcription system for conversation analysis muhammad umair muhammad.umair@tufts.edu tufts university, medford, ma, 02155 julia mertens julia.mertens@tufts.edu tufts university, medford, ma, 02155 saul albert s.b.albert@lboro.ac.uk loughborough university, loughborough, united kingdom jan p. de ruiter jp.deruiter@tufts.edu tufts university, medford, ma, 02155 editor: kallirroi georgila submitted 09/2020; accepted 04/2022; published online 04/2022 abstract researchers studying human interaction, such as conversation analysts, psychologists, and linguists, all rely on detailed transcriptions of language use. ideally, these should include so-called paralinguistic features of talk, such as overlaps, prosody, and intonation, as they convey important information. however, creating conversational transcripts that include these features by hand requires substantial amounts of time by trained transcribers. there are currently no speech to text (stt) systems that are able to integrate these features in the generated transcript. to reduce the resources needed to create detailed conversation transcripts that include representation of paralinguistic features, we developed a program called gailbot. gailbot combines stt services with plugins to automatically generate first drafts of transcripts that largely follow the transcription standards common in the field of conversation analysis. it also enables researchers to add new plugins to transcribe additional features, or to improve the plugins it currently uses. we describe gailbot’s architecture and its use of computational heuristics and machine learning. we also evaluate its output in relation to transcripts produced by both human transcribers and comparable automated transcription systems. we argue that despite its limitations, gailbot represents a substantial improvement over existing dialogue transcription software. keywords: automated transcription, conversation analysis, natural language processing 1. introduction researchers studying language and communication need to analyze the speech of participants in naturally occurring dialogue. transcribing dialogue is usually the first step in analyzing verbal interaction. however, even creating basic ‘verbatim’ transcripts is notoriously painstaking and timeconsuming (tilley, 2003; lapadat and lindsay, 1999; boyce and neale, 2006). for example, saon et al. (2017) took 12-14 times the duration of each recording to transcribe. as a workaround, re©2022 muhammad umair, julia mertens, saul albert and jan p. de ruiter this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). umair, mertens, albert, and de ruiter searchers can crowd-source (novotney and callison-burch, 2010), automate (bokhove and downey, 2018), or eliminate (stonehouse, 2019) costly parts of the data preparation process. alternatively, researchers can hire professional transcribers, at a rate between $0.75 and $1.50 per minute of audio recording to produce verbatim transcripts containing words and timestamps. recently, researchers have started to use automated speech recognition (asr) systems to produce ‘good enough’ first draft transcripts (bokhove and downey, 2018). since the 1980s, nlp researchers have developed and refined asr technology, including speech-to-text (stt) systems (juang and rabiner, 2005; moore, 2015). stt systems recognize word orthography and timing in audio streams, and are used in prominent applications (e.g., dictation systems, voice assistants, automated captioning, etc.). stt systems are between twenty to forty times cheaper than hiring professional transcribers. for example, at the moment of writing, google’s cloud stt service costs $0.036 per minute of audio and ibm watson’s stt service costs $0.02 per minute of audio. for data recorded in ideal conditions, these systems can achieve close to human level accuracy in recognizing words. ibm’s watson can achieve a word error rate (wer) of 5.5%-11% compared to a 5.1%-6.8% wer for human transcribers on the same data (saon et al., 2017). however, even if automated stt can approximate human-level word recognition and create ‘good enough’ first draft transcripts, it cannot identify paralinguistic features of speech and integrate them into the transcription (moore, 2015). these are non-lexical aspects of speech, like volume, voice quality, intonation, and laughter, which carry strong interactive meaning. transcribers using conversation-analytic transcription use what is often called jeffersonian notation, which uses standard keyboard symbols1 to visually represent not only the words, but also paralinguistic features in the speech (jefferson, 2004a; hepburn and varney, 2013; hepburn and bolden, 2017). for brevity, we will refer to transcripts that use jeffersonian notation as ca transcripts. excerpt 1: machine transcription from happyscribe, a commercial transcription service. 1. [00:00:00.000] speaker 1 it also makes me want to, 2. like, have, like, hot chocolate and like it from a 3. fireplace. 4. [00:00:05.250] speaker 2 oh, my god. ... there’s no 5. fires at tufts, and it really pisses me off. 6. [00:00:21.230] speaker 1 there’s too many fires in 7. california. 8. [00:00:23.080] speaker 2 no, like, fireplaces. 9. [00:00:25.060] speaker 1 yeah. excerpt 2: manual ca transcript for the same audio used in excerpt 1. 1. *sp1: it also ma:kes me: w:ant to like (0.2) have like 2. (0.3) hot chocola:te >and like ⌈sit in fron⌉t of 1a summary of jeffersonian transcription symbols is provided in appendix a. 64 https://sites.tufts.edu/hilab/files/2021/07/2018-10-27-session-5-line-98-two_channels.mp3 https://sites.tufts.edu/hilab/files/2021/07/2018-10-27-session-5-line-98-two_channels.mp3 an automatic transcription system for conversation analysis 3. a firep⌈lace,< ⌉ 4. *sp2: ⌊ mm ⌋ 5. *sp2: ⌊o:h my go⌋d. ... 6. *sp2: ⌊there’s⌋ no fires: at tufts: and it really pisses 7. me of⌈f.⌉ 8. *sp1: ⌊th⌋ere’s too many fir:es in california, 9. (0.5) 10. *sp2: no like fire like firepla⌈ces⌉ 11. *sp1: ⌊ oh ⌋ yeah, excerpt 1 shows a transcript generated using an online stt service. excerpt 2 shows a manual ca transcript of the same audio. the paralinguistic features in excerpt 2 affect the interpretation of the sequence. first, excerpt 1 does not identify the non-lexical vocalization (“mm”) represented in line 4 of excerpt 2. with this vocalization, also called a ‘continuer’ (schegloff, 1982), sp2 signals that they have understood sp1 so far, and that sp1 can continue speaking. second, changes in syllable rate (see section 2.2.4) are not annotated in excerpt 1 but are marked by angled brackets in excerpt 2 (lines 2-3). here, sp1 uses faster-than-normal speech to extend their turn, which could have been expected to be complete on syntactical grounds after “chocolate”, before sp2 can interject. third, overlaps are not annotated in excerpt 1, whereas excerpt 2 shows the exact temporal order of events. for example, sp2’s minimal uptake (“mm”) is a reaction to sp1’s mention of hot chocolate, and not the fireplace. later, sp2 produces “oh my god,” on line 4 of excerpt 1, in ‘lastitem onset’ overlap (drew, 2009). this shows that sp2 reacts enthusiastically to sp1’s mention of a fireplace. the precise temporal order of the utterances is not clear without the overlap markers. fourth, silences, represented as the number of seconds between brackets in the manual transcript, are not represented in excerpt 1. sp2 uses “fires” to refer to fireplaces, but on line 8, sp1 uses “fires” to refer to forest fires. next is a 500 ms gap, which is longer than the average turn transition (stivers et al., 2009). this gap projects a socially ‘dispreferred’ response (pillet-shore, 2017), in this case a correction of sp1’s use of “fires”. even this cursory comparison shows that ca transcripts provide information vital to interpreting verbal interaction, but that is absent in standard stt transcripts. however, while there are corpora with audio, words, timestamps and other useful annotation layers such as part-of-speech (pos) tags (e.g., godfrey et al., 1992), there are no large-scale corpora of ca transcripts. in part, this is because ca transcription takes a lot of time, even for experts. experienced ca transcribers usually estimate that one minute of dyadic conversation, with high quality audio, takes around an hour to transcribe in ca format (wagner, 2020). instead of transcribing entire corpora with ca notation, researchers often first produce a first draft transcript and then transcribe segments of interest in greater detail (heath et al., 2010). automating first draft ca transcription will increase the time interaction researchers can spend on analysis as well as the amount of transcribed data they have access to. similarly, nlp research has been held back by its reliance on a limited number of available corpora of spoken language, such as the switchboard corpus (godfrey et al., 1992), which does 65 umair, mertens, albert, and de ruiter not include annotations of paralinguistic features of talk, such as speech rate, pauses, and pitch. corpora of ca transcripts could be used to improve nlp-based software systems in a variety of settings. for example, call centers use nlp to help agents and assess performance (mishne et al., 2005). clinicians use similar tools to help diagnose patients (mirheidari et al., 2017; blomberg et al., 2019). enriching the data used to create nlp models will improve all these services. to reduce the resources required to produce conversation-analytic transcripts at scale, we developed gailbot: the first automated transcription system designed to generate draft ca transcripts. moore (2015) first introduced automated transcription to ca, using ibm’s atilla system (soltau et al., 2010). although this software transcribed word orthography and timing, it did not annotate paralinguistic features. gailbot addresses this limitation by automatically generating draft ca transcripts, using one of multiple available stt services to generate words, timestamps, and identify speakers (speaker diarization). it then employs a series of plugins that identify paralinguistic features (see section 2.1.2). users can extend gailbot with new plugins to identify new paralinguistic features and improve gailbot’s performance. finally, gailbot outputs a first draft ca transcript, which researchers can then improve and enhance manually. note that gailbot aims to facilitate, and not replace, manual transcription. nonetheless, we suggest that gailbot enables researchers to produce large scale corpora containing ca transcripts, which can create opportunities for creating new interfaces between computational-linguistic and conversation-analytic approaches. in the architecture section, we describe gailbot’s internal data structures, data flow, and the algorithms used in current plugins. in the performance section, we compare gailbot transcripts to manual transcripts and those used by moore (2015), provide an estimate of the transcription time saved by using gailbot, and compare it to existing data annotation systems. finally, in the discussion section, we address the technical and theoretical limitations of gailbot and automated transcription in general and suggest steps for future development. 2. architecture and algorithms in this section, we begin with an overview of gailbot’s architecture and internal data representations. we then highlight gailbot plugins, which are wrapper objects for custom algorithms that identify and annotate specific paralinguistic features. finally, we describe and analyze these algorithms. 2.1 architecture 2.1.1 data flow and internal representations gailbot implements an api that abstracts over the transcription process. conceptually, this api has two components: an organizer and a core pipeline, connected by a controller. the organizer handles both 1) media files of conversations and 2) previous gailbot outputs. it determines whether the input can be processed and, if so, creates a conversation object for that input. a conversation object stores information required by the core pipeline to generate a transcript. each conversation object has a settings profile that defines how the object is processed by the pipeline. for example, a settings profile allows the user to select which stt service to use and which plugins to apply. once a settings profile has been created, it can be saved, re-used, modified, and applied to multiple conversation objects. however, each conversation object must have a settings profile attached before it can be processed. 66 an automatic transcription system for conversation analysis figure 1: an overview of gailbot’s architecture and processes. once all conversation objects have a settings profile, the controller transfers them from the organizer to the core pipeline. the core pipeline transcribes, analyzes, and formats each conversation object in separate processes. these processes are executed sequentially, each augmenting the conversation object with information required to produce a final transcript. the transcription process sends raw audio data to one of multiple stt services, such as ibm’s watson and google cloud (soltau et al., 2010), based on the settings profile. different stt services differ in accuracy, but they all provide the same basic functionality: they process an audio waveform, recognize words, 67 umair, mertens, albert, and de ruiter and return the orthographic representation and timing of each recognized word. this returned information is organized into and stored as lists of utterance objects, with each object containing a speaker label, start time, end time, and text. at this stage, utterance objects are saved to disk and can be reused if needed, bypassing the costly and time consuming transcription process. for example, users can apply different sets of plugins on previous gailbot outputs without executing the transcription process. next, the analysis process executes plugins (see section 2.1.2) that identify paralinguistic features of talk. finally, the format process consolidates information from the conversation object into a ca transcript. figure 1 visualizes this entire process. 2.1.2 plugins gailbot provides a framework for users to add custom algorithms to identify specific paralinguistic features. plugins are wrapper objects (implemented as python classes) that provide a standard api for these algorithms to interact with the core pipeline. this means that customizing and changing plugins does not require modifying gailbot’s source code. plugins may or may not be dependent on each other. for example, plugin b is dependent on plugin a if the feature identified by plugin a is required as an input by plugin b. the internal plugin pipeline, executed during the analysis process, manages these dependencies. it represents each plugin as a node in a directed acyclic graph (dag) and uses edges to encode dependencies between plugins. this ensures that plugins can be executed concurrently when possible and that plugins whose dependencies fail are not executed. it also allows plugins to receive outputs from all their dependencies before they are executed. configuration files are used to add plugins and define their dependencies in a standard format. by default, gailbot applies a few plugins that are required for most transcripts. for example, one default plugin constructs turns from word timing and orthography. users can select plugins by configuring the settings profile of a conversation object. 2.2 algorithms in this subsection, we describe algorithms we developed to identify certain paralinguistic features. we highlight the relevance of each algorithm and discuss its strengths and limitations. note that these algorithms are implemented in the current version of gailbot as plugins. 2.2.1 turns turn construction units (tcus) are fragments of speech that are ‘hearably’, pragmatically, grammatically, and prosodically complete. after a tcu, speakers have the opportunity to start the next turn at a transition relevance place (trp) (sacks et al., 1974; clayman, 2012). conversation analysts use tcus and trps to decide whether to start a new line in a transcript, although they do so in different ways. some transcribe each tcu on a new line, while some only create a new line when there is a long gap, and yet others only create new lines when there are speaker transitions. a common convention is to put tcus that contain overlapping speech on their own line, with the overlapping tcu below it, in order to make the overlap clear. however, there is no general consensus among transcribers on how to split speech into different lines. there are also some features that are only relevant at the end of tcus, like latches or turnfinal intonation (see appendix a). asr systems are typically limited in their ability to analyze pragmatics, grammar, and prosody to reliably infer the location of tcus and trps. therefore, gailbot exploits a normative pattern in turn-taking instead of trying to implement human-like tcu 68 an automatic transcription system for conversation analysis perception: interlocutors frequently begin speaking when their social partner finishes their turn (schegloff, 1979). the algorithm assumes that when a new speaker initiates a turn, the current turn has ended. this method is fast and simple, but limited. the main drawback is that the current plugin cannot identify the beginnings or ends of consecutive tcus produced by the same speaker. this means that gailbot cannot currently identify turn-final characteristics of many tcus such as latches and turn-final intonation. in addition, when speakers overlap, gailbot cannot use tcus to choose when to start a new line. researchers interested in those features are encouraged to replace the algorithm used in the turn construction plugin with a more sophisticated one (e.g., masumura et al., 2018). 2.2.2 silences silences often have interactional significance (e.g., pomerantz, 1984), and provide evidence for cognitive theories about communication (e.g., de ruiter et al., 2006). there are two types of silences annotated in ca transcripts: gaps (silences between turns) and pauses (silences within turns). in this paper, silences refer to any period of non-talk. excerpt 3: gailbot transcript demonstrating turn construction.2 1 *sp1: um (.) it’s about like this (0.5) man (0.3) who 2 it’s like (.) he’s like (.) pilgrim (0.5) and 3 he’s (.) like 4 (2.5) 5 *sp1: kind of like an undercover agent (0.4) and he 6 (0.5) does (.) is like (0.3) living out like (.) 7 in europe and traveling in doing like excerpt 4: manual transcription of excerpt 3. 1 *sp1: um (0.2) it’s about like this (0.5) ma:n (0.3) 2 who it’s like (0.2) he’s like (0.2) pilgrim 3 (0.5) and he’s (0.2) like (2.5) >kinda like< an 4 undercover agent (0.3) a:nd he (0.5) does (0.2) 5 is like (0.3) living out like (0.2) in europe 6 and traveling and doing like 2unless otherwise stated, data comes from the human interaction laboratory’s in conversation corpus (icc) (see section 3.2.3). audio for select icc transcripts is available at the hi-lab website. 69 https://sites.tufts.edu/hilab/files/2020/06/excerpt-4-pilgrim.mp3 https://sites.tufts.edu/hilab/files/2020/06/excerpt-4-pilgrim.mp3 https://sites.tufts.edu/hilab/gailbot-an-automatic-transcription-system-for-conversation-analysis/ umair, mertens, albert, and de ruiter the silence algorithm first uses word timing, regardless of speaker identity, to determine silence start and end times. next, it uses these start and end times to calculate the duration, in seconds, of each silence. finally, it rounds the silence duration to the nearest 100 ms and classifies it as either a micro-pause, pause, or gap. by default, the plugin classifies a silence based on ca notation guidelines. a micro-pause is a within-turn silence between 100 ms and 200 ms long. a pause is a within-turn silence between 200 ms and 1000 ms long. if a pause is greater than 1000 ms, it is considered a gap for a number of reasons. jefferson (1983) proposed that one second is the maximum standard silence in conversation. in addition, yang (2004) suggests that silences longer than one second are more likely to be between than within turns. in excerpt 3, gailbot transcribed a 2500 ms silence (line 4) as a gap separating two tcus instead of a silence, as transcribed in excerpt 4, line 3. finally, gailbot only transcribes gaps that are at least 300 ms long. users can adjust these duration thresholds based on their preference.3 for example, a user may wish to increase the gap threshold to 2000 ms when transcribing speech produced by someone who struggles with lexical retrieval, or in situations where activities by the participants may provide an account for longer silences (jefferson, 1983). figure 2: example of beat timing. (liddicoat, 2007) how interlocutors perceive silence can depend on how a speaker talks (hepburn and bolden, 2017: 26-27). if a speaker talks slower than usual, a long silence may feel short; if a speaker speaks faster than usual, the same silence may feel longer. therefore, some researchers transcribe silences in beats instead of seconds. gailbot has two modes, the ‘absolute’ and ‘beat’ timing modes, to approximate each method. in ‘absolute’ mode, a silence is the time between the end of one word and the start of the next. in contrast, ‘beat mode’ takes the syllable rate into account. ca transcribers estimate the number of beats by counting “one one thousand, two one thousand, ...” while adopting the same speech rhythm as that in the transcribed conversation (jefferson, 1983). because “one one thousand” represents one unit of time, and has four syllables, a syllable equals 0.25 beats (see figure 2, taken from liddicoat, 2007: 31). to estimate beats, this algorithm first calculates the syllable rate (see section 2.2.4) for each speaker individually. next, it uses the syllable rate to determine the number of syllables that fit each silence. this value is divided by 4 to obtain the final number of ‘beats’. by default, gailbot uses the absolute timing mode, but the user can choose to have gailbot use beat mode instead if desired. 3for all heuristics and thresholds, see appendix b. 70 an automatic transcription system for conversation analysis 2.2.3 overlapping speech interlocutors can start turns at the same time, interrupt each other, or start turns before another speaker has finished (roger et al., 1988; drew, 2009; jefferson, 2004b). these overlaps affect the sequential ordering of actions. the overlap algorithm4 uses speaker identity and utterance timing to identify moments where speakers overlap. excerpts 5 and 6 highlight the differences in overlap annotations between gailbot and a manual transcript. in excerpt 5, the algorithm places overlap markers at the start and end of word. in some cases, this strategy can reduce the accuracy of the overlap markers, especially when a very long word is overlapped. excerpt 5: gailbot transcript demonstrating overlap markers. 1 *sp2: uhm so this semester we have some new members 2 and the total number would be like (0.5) about 3 twenty i ⌈ think ⌉ 4 *sp1: ⌊oh that’s⌋ that’s a nice (.) 5 number excerpt 6: manual transcription of excerpt 5. 1 *sp2: um: so this semester: we have some new members 2 and the total number would be like (0.5) about 3 twenty i thin⌈k ⌉ 4 *sp1: ⌊oh⌋ that’s that’s a 5 nice (.) number 2.2.4 syllable rate interlocutors can speak at different rates, and speak faster or slower to convey information (e.g., display urgency) or perform actions (e.g., compete for turn space; french and local, 1983). transcribers represent meaningful changes in speech rate by annotating speech that is faster or slower than normal, as well as syllables that are ‘stretched’. excerpt 7: gailbot transcript demonstrating the syllable rate algorithm. 1 *sp2: yeah we will perform in the parade of nations and 2 like we will have our own show case maybe next 3 semester (0.3) >usually in the spring semester< 4 (0.4) 5 *sp1: wow that’s awesome how many people are in that 4the overlap algorithm is described in appendix c. 71 https://sites.tufts.edu/hilab/files/2020/06/excerpt-10-new-members.mp3 https://sites.tufts.edu/hilab/files/2020/06/excerpt-10-new-members.mp3 https://sites.tufts.edu/hilab/files/2020/06/excerpt-6-spring-semester.mp3 umair, mertens, albert, and de ruiter excerpt 8: manual transcript of excerpt 7. 1 *sp2: yeah we will perform in the parade of nations 2 and like we will have our own showcase maybe 3 next semester (0.4) >usually in the spring 4 semester< 5 (0.4) 6 *sp1: wow that’s awesome how many people are in that the syllable rate algorithm works by identifying outliers on the segment level. segments are defined as speech surrounded by silences of at least 100 ms. gailbot estimates the number of syllables in each segment using the big phoney python package (epp, 2018), and divides the result by the duration of each turn (see section 2.2.1). for each speaker, the algorithm calculates the median absolute deviation (mad) (see equation 1), a measure of variance that is robust to outliers (leys et al., 2013). by default, turns with syllable rate two mad above or below the median are considered fast or slow speech. for example, in excerpt 7, faster than normal speech is annotated on line 3. excerpt 8 is a manual transcript for the same audio. mad = median(|xsi − x̃s|) (1) xsi = segment syllable rate x̃s = speaker syllable rate additionally, the algorithm identifies some stretched syllables. it measures how much slower a word is than normal by using the median syllable rate for the specific speaker. one marker is inserted for every mad the syllable rate is below the median. the sound-stretch markers are added after the last vowel in the word. currently, this addition position is arbitrary as we do not yet have access to algorithms that determine which specific phonemes are elongated within a word. 2.2.5 laughter laughter, like most non-lexical vocalizations, provides an important resource for participants in interaction (keevallik and ogden, 2020) and is always annotated in ca transcripts. these representations describe generic laughter as well as sophisticated inbreaths, outbreaths, pulses, and other vocal characteristics of laughter (hepburn and varney, 2013). the current laughter algorithm, inspired by ryokai et al. (2018); abadi et al. (2016), identifies the start and end times of laughter. it starts by removing background noise using google’s voice activity detector. next, it uses a three-layer feed-forward neural network. the first two layers perform batch normalization, dropout (to prevent overfitting), and use relu as the activation function. the final layer is a dense layer with a sigmoid activation function. given a target frame and standard audio features (mfccs and delta-mfccs), the network produces laughter probabilities for all 10 ms segments in the audio. the network is trained on the switchboard corpus (godfrey et al., 1992) and achieves an 88% per-frame accuracy on a validation set (ryokai et al., 2018; abadi et al., 2016). finally, the algorithm uses a low-pass filter to segment the full duration of a laugh from the recording. currently, this algorithm cannot determine more granular vocal components of laughter (e.g., inhalations, exhalations 72 https://sites.tufts.edu/hilab/files/2020/06/excerpt-6-spring-semester.mp3 an automatic transcription system for conversation analysis etc.). instead, it identifies complete laughter segments and leaves it to the human transcriber to add interactionally relevant levels of detail. 3. performance in this section, similar to moore (2015), we evaluate gailbot’s performance on conversation from different corpora. the corpora differ in their sampling rate, audio separation, and background noise. the newport beach corpus has a sampling rate of 44.1 khz, does not have speaker audio separation, and has background noise. the callhome corpus has speaker audio separation, an 8khz sampling rate, and background noises. the in conversation corpus has audio separation, no background noise, and a high sampling rate of 48 khz. we compare gailbot’s performance on these data to manual and automated transcripts. we also provide the results of a brief experiment estimating the amount of time saved when using gailbot compared to transcribing from scratch. finally, we compare gailbot with existing data annotation systems. 3.1 metrics 3.1.1 word error rate to be useful, asr technology must correctly identify most words in an audio stream. the word error rate (wer), defined by equation 2, is one measure of the accuracy of a speech to text system, defined as the number of errors made by the service relative to the total number of words it attempts to recognize. it accounts for three types of errors: additions, deletions, and substitutions. addition occurs when the system identifies a word that is not in the audio stream. a deletion occurs when the system does not identify a word. finally, a substitution occurs when the system replaces one word with another. note that the wer can be over 100% in cases where more words are misidentified than actually exist. in this paper, we use two types of wers: a strict wer (swer) and a relaxed wer (rwer), both computed by manually comparing gailbot transcripts to humanproduced transcripts. wer = (additions + deletions + substitutions) words in transcript × 100 (2) 3.1.2 relaxed word errors rwer is calculated with the same equation as wer, except it does not count vocalizations with no dictionary spelling (e.g., “um” vs. “uhm”) as errors. in addition, rwer does not count substitutions for phonetically identical words (e.g., “for” vs. “four”) as errors, because it is very difficult for most stt systems to tell the difference between the two. we present this metric along with swer to provide a more forgiving metric and a range of expected performance for gailbot. 3.1.3 strict word errors the strict word error rate (swer) also uses the same equation as wer, but considers any differences between the gailbot produced and manually produced transcript as errors. for example, swer considers substitutions for phonetically identical words (e.g., “for” vs. “four”) as errors. therefore, the swer is either higher than or equal to the rwer. 73 umair, mertens, albert, and de ruiter 3.1.4 overlap and silence performance overlaps and silence errors can occur when gailbot and a human transcriber either 1) disagree on the presence of a silence or overlap or 2) agree on the presence but disagree on the duration of the overlap or silence. depending on the researcher’s goals, small differences in the magnitude of an overlap or gap may not be considered significant. however, it is important to indicate whether a gap or overlap exists, as it informs readers regarding the sequence of the conversation. gailbot silence and overlap identification is influenced by the accuracy of word recognition preformed by the stt service. if gailbot cannot recognize a word, then it assumes the word is silence. this may occur if the speaker whispers, uses ‘creaky voice’, or uses a word not contained in the stt service’s vocabulary. as a consequence, gailbot will mark a silence that is not there, and might not notice an overlap. alternatively, gailbot may transcribe a breath or other noise as a word. this will eliminate a true silence, or cause gailbot to erroneously identify an overlap when it does not exist. we did not want to penalize the overlap and silence algorithms for an error caused by the stt service. therefore, we do not reduce the silence or overlap accuracy for errors caused by inaccurate word recognition. 3.1.5 syllable rate performance words, overlaps, and silences can be measured relatively objectively and have standard annotations in ca transcripts. however, there is no clear metric to determine whether a speaker’s turn should be considered ‘fast speech’. therefore, while we present examples of syllable rate annotations in this paper, we do not use a quantitative measure for syllable rate accuracy. this module is relatively robust to word recognition errors for a few reasons. if gailbot does not recognize a word and decides it is silence, the syllable rate algorithm splits the speech into two segments (instead of calculating a very slow speech rate). further, by default, gailbot identifies segments with speech rate 2 median absolute deviations away from the median. in practice, this includes relatively few turns. therefore, one additional word is unlikely to cause a ‘normal’ turn to be misidentified as a ‘fast’ turn. instead, performance is affected by exactly how researchers analyze speech rate. it’s likely that transcribers consider speech rate changes based on the most immediate context, while gailbot considers speech rate changes based on the entire conversation (including segments after the target segment). as we learn more about how listeners perceive changes in syllable rates, this algorithm may be refined and better evaluated. 3.1.6 turn taking performance transcribers differ in their representations of the ordering of talk. for example, some transcribers may want each tcu on a separate line. others may split consecutive tcus by the same speaker only when there is a long pause. some may even choose to separate tcus only when another speaker begins speaking. because there is no objective standard for annotating turn structure, we do not evaluate how well gailbot splits turns into lines. however, our subjective impression, based on having used gailbot extensively for several years, is that although it is far from perfect, its annotations of turn structure are adequate for generating a useful first pass. 74 an automatic transcription system for conversation analysis 3.1.7 laughter performance we could not perform a conclusive evaluation of the laughter detection algorithm because we did not have enough unique instances of laughter in our data to analyze. in addition, gailbot does not yet attempt to represent laughter in ca format, which would require a representation of each separate laughter particle, including for instances where the speaker produces a laugh particle mid-word. instead, gailbot marks laughter with “(&=laughs)”. the machine learning model in our laughter plugin has an 88% per-frame accuracy (ryokai et al., 2018; abadi et al., 2016). 3.1.8 speaker diarization performance speaker diarization, or the identification of speakers, is of central importance for any transcripts of dialogue. however, it is notoriously difficult to achieve reliably in mixed speaker, mono recordings, although advances in deep learning are improving the performance of stt systems (park et al., 2021). gailbot relies on the diarization capability of the stt system being used. when speakers are recorded separately, gailbot labels identity based on the audio source, and therefore has an accuracy rate of 100%. when speakers are recorded on the same channel, stt systems, and by extension, gailbot struggles to identify speakers. we therefore recommend that researchers using gailbot record multiple speakers on separate audio channels. we do not evaluate speaker diarization in gailbot. we expect improvements in speaker diarization on single-channel audio once diarization algorithms in current stt systems have improved. 3.2 evaluation in this subsection, we compare gailbot transcripts to transcripts produced by other automated software (cf. moore, 2015), corpora available online, and transcripts of the same data presented in a ca publication (walker, 2017). 3.2.1 newport beach corpus the newport beach corpus (jefferson, 2007) contains mono audio files with 44.1khz sampling rates and high background noise. transcribing this corpus is difficult for automated systems because it requires speaker diarization and background noise suppression. excerpt 9 is a gailbot transcript of part of a phone call discussing the assassination of robert f. kennedy.5 gailbot achieved a swer/rwer of 44.33%. for the same audio, (moore, 2015) reported a 36% wer. both systems produced errors based on a single phoneme (e.g., “makes” → “make”, “god” → “got”). they also mis-transcribed words with ‘nonstandard’ pronunciation (e.g., “didju” as “that you” and “of’m” as “up”). finally, the systems were largely unable to transcribe quiet speech (“oh no. they drag it out so” on line 83 of appendix d, transcript 1) and overlapping speech. excerpt 9: gailbot transcript of the assassination call. 1 *sp2: world uhm long week (0.3) my god i’m 2 *sp1: glad it’s over i wanted to have a tv are on the 3 there 5comparison transcripts for excerpt 9 are in appendix d. 75 https://sites.tufts.edu/hilab/files/2021/07/assassination_file_trimmed_gb_paper.mp3 umair, mertens, albert, and de ruiter 4 (0.7) 5 *sp2: like check it out 6 *sp1: that’s where they be took off on our charter 7 flight that same spot that you see it (0.8) >when 8 i took him in the our< 9 (1.0) 10 *sp2: would be more did i think it’s so ridiculous i 11 mean it’s 12 (0.7) 13 *sp1: it’s a horrible thing that my god play up that 14 *sp2: thing is 15 *sp1: just horrible guy people not 16 *sp2: ride:: 17 *sp1: in a make american people think well they’re no 18 good (0.5) well they aren’t very good some up in most cases, both gailbot and the atilla system evaluated by moore (2015) annotated quiet or non-canonical speech as silences. gailbot transcribed extra silences on lines 1, 4, 9, and 12. we believe this is because gailbot transcribed overlapping speech as silences due to word recognition errors. in comparison, the manual transcript contained only two silences: a micro-pause on line 80 and a 700 ms gap on line 86. neither system transcribed the micro-pause and both estimated the gap to be 800 ms. moore (2015) additionally notes atilla’s misidentification of a 100 ms pause on line 74. excluding silences that could be attributed to word recognition errors, both gailbot and atilla had 100% silence error rate. overlaps depend on speaker identity because they occur between different speakers. neither moore (2015) nor gailbot perform accurate speaker diarization on mono audio. therefore, both were unable to accurately transcribe overlaps. moore (2015) only transcribed one speaker during overlapping talk and did not annotate an overlap. comparatively, gailbot transcribed overlaps as silences. neither system transcribed any of the overlap markers that are present in the manual transcript, resulting in an overlap error rate of 100%. 3.2.2 callhome corpus the callhome corpus contains stereo telephone conversations between family and friends with a low 8khz sampling rate. for excerpt 10, gailbot has a swer of 18% and a rwer of 9%.6 word errors included phonological errors (e.g., “theories” was transcribed as “series”), misspellings of phonologically identical words (e.g., “four” was transcribed as “for”), and difficulty transcribing word cutoffs (“bul-” was transcribed as “bo”). the overlap was correctly placed. additionally, 6comparison transcripts for excerpt 10 are in appendix e. 76 an automatic transcription system for conversation analysis gailbot transcribed each silence as approximately 100 ms shorter than in the same data as transcribed by walker (2017). this timing discrepancy eliminates the micropause identified by walker (2017) in line 5 and the subsequent 900 ms to be transcribed as 800 ms long. however, this still marks an improvement over the callhome corpus transcript, which does not annotate silence at all. excerpt 10: gailbot transcript of audio from the callhome corpus. 1 *sp2: ⌈they⌉ they go through the series of the 2 three bullets of the magic one bullet 3 *sp1: ⌊yeah⌋ 4 (2.9) 5 *sp1: yeah for bo yeah (0.8) it was interesting excerpt 11 is a gailbot transcript of a single speaker. gailbot produced one word-segmentation error (“and to” was transcribed as “into” in line 4), resulting in a swer/rwer of 2%. again, gailbot decreased silence duration by 100 ms, considering the 200 ms pause in line 1 to be a micropause and removing the micropauses in lines 3 and 5 of the transcript published in walker (2017)7. gailbot also fails to identify the inhalations that are clearly audible in the recording. instead, it marks inhalations as silences. depending on the research question and the activities underway for speakers, ca transcribers may or may not mark all inhalations as interactionally relevant silences (trouvain et al., 2020). we suggest that researchers interested in breathing patterns manually correct inhalation related inaccuracies in gailbot’s transcripts. researchers may also be interesting in designing a custom inhalation plugin based on state of the art deep learning techniques (nallanthighal et al., 2021). excerpt 11: gailbot transcript of audio from the callhome corpus, single speaker. 1 *sp1: clinton just came out and said that he (.) 2 doesn’t believe (0.4) in quota systems (0.4) and 3 in reverse discrimination but that he does 4 believe that affirmative action is necessary 5 (0.3) to move uhm you know black americans (0.3) 6 forward into give them the opportunities that 7 they’ve been denied 7comparison transcripts for excerpt 11 are in appendix f. 77 https://sites.tufts.edu/hilab/files/2021/07/callhome_4686_excerpt_8.mp3 https://sites.tufts.edu/hilab/files/2021/07/callhome_4247_excerpt_9.mp3 umair, mertens, albert, and de ruiter 3.2.3 in conversation corpus the in conversation corpus (icc) is a high audio quality (48 khz sampling rate) conversation corpus, collected in the human interaction lab (hi-lab) at tufts university. each conversation features two students in two sound-proofed rooms that are separated by a glass window. students communicated using a microphone and headset, and each student was recorded on a separate channel. we use the icc to evaluate gailbot in two ways. first, we compare icc and gailbot transcripts presented earlier in this paper. second, we compare transcripts of 15 short segments of talk collected as part of a miscommunication study. excerpts 3, 5, and 7 are all gailbot transcripts from the the icc and have a rwer of 0%. excerpt 3 has a swer of 2.8% (“kinda” transcribed as “kind of” (line 5)). excerpt 5 has a swer of 3.8% (“um” transcribed as “uhm” (line 1)). finally, excerpt 7 has a swer of 2.9% (“showcase” transcribed as “show case” (line 2)). additionally, we selected 15 miscommunication sequences that we expected would be difficult to transcribe. combined, these sequences lasted 7 minutes and 30 seconds. they had a rwer of 19.7%, a swer of 23.2%, an overlap error rate of 53.6%, and a silence error rate of 63.3%. excerpt 12: gailbot transcript from the icc demonstrating overlap error. 1 *sp1: love it (1.0) ⌈what’s your⌉ favorite show 2 *sp2: ⌊ but ⌋ (0.9) oh (0.6) my 3 favorite show on netflix will be house of cards 4 and friends excerpt 13: manual transcription of excerpt 12. 1 *sp1: love it 2 (1.0) 3 *sp1: w⌈hat’s yo⌉ur favorite show 4 *sp2: ⌊ but ⌋ 5 *sp2: um: (0.2) my favorite show on netflix will be: 6 house of cards and friends figure 3 shows that most overlap and silence errors have a small magnitude. most overlap errors are 0-5 characters, and most silence errors have magnitude less than 300 ms. however, there are systematic sources of errors in gailbot’s transcripts. its limited ability to detect tcus and other pragmatic cues causes larger overlap and silence errors. for example, excerpt 12, line 2 should be transcribed on multiple lines (see excerpt 13, lines 4-5). since gailbot’s default setting is to separate turns that are greater than 1000 ms apart, it transcribes the time between “but” and “um” (when sp1 is talking) as a within-turn pause. a manual transcriber was able to mark the inhalation of a new action (the initial question at line 3 in excerpt 13), with a new line. 78 https://sites.tufts.edu/hilab/files/2020/06/excerpt-8-whats-your-favorite-show.mp3 https://sites.tufts.edu/hilab/files/2020/06/excerpt-8-whats-your-favorite-show.mp3 an automatic transcription system for conversation analysis 5 0 5 10 15 difference in characters 0.05 0.10 0.15 0.20 0.25 density 1 0 1 2 3 difference in seconds 0.2 0.4 0.6 0.8 1.0 1.2 1.4 density figure 3: the magnitude of overlap (top) and silence (bottom) errors from the icc. a positive value means that gailbot marked a wider overlap or longer silence than the manual transcriber. additionally, overlap errors can cause large silence errors. this occurs when one speaker speaks partway through their interlocutor’s turn but finishes before their interlocutor. for example, in excerpt 14, the 1.3 second gap in line 4 marks the time between the end of line 3 in excerpt 15 (“yeah”) and the start of line 4 in excerpt 15 (“go...”). this includes the duration of speaker 2’s “other clubs i just like” (lines 2-3 in excerpt 14, line 3 in excerpt 15). this is a significant limitation of the overlap detection algorithm that should be manually corrected. excerpt 14: overlap error: gap error. 1 *sp2: that’s (0.8) yeah that’s basically the (0.2) like 2 the club that i am really committed ⌈to but other⌉ 3 clubs i just like 3 *sp1: ⌊ yeah ⌋ 4 (1.3) 5 *sp2: go to the activities and (0.2) like i’m not 6 really p official (0.4) ⌈member⌉ 7 *sp1: ⌊member⌋ ⌈yeah⌉ 79 https://sites.tufts.edu/hilab/files/2020/06/excerpt-12-committed.mp3 umair, mertens, albert, and de ruiter 8 *sp2: ⌊yeah⌋ excerpt 15: overlap error: gap error, corrected. 1 *sp2: that’s (0.5) yeah that’s basically the (0.2) like 2 the club that i am really committed t⌈o but⌉ 3 other clubs i just like 3 *sp1: ⌊yeah⌋ 4 *sp2: go to the activities and (0.3) like i’m not 5 really (0.5) the official mem⌈ber ⌉ 6 *sp1: ⌊memb⌋er y⌈eah ⌉ 7 *sp2: ⌊yeah⌋ similarly, overlapping talk can interfere with the turn construction algorithm. human transcribers will split the first turn (the one that is overlapped) when there is a trp or tcu. however, gailbot splits turns when there are pauses and not when the tcu ends. for example, in excerpt 15, speaker 1 interjects in speaker 2’s turn with “yeah.” instead of splitting speaker 2’s utterance when it approaches a trp (after “committed to”), gailbot waits until speaker 2 pauses slightly (after “i just like”). this creates a visual break that does not match the real structure of the conversation. another limitation of gailbot is that it cannot produce turn-final jeffersonian prosody markers until it has a more sophisticated tcu detection system. this is a clear next step for researchers. 3.3 transcription time estimation one of the co-authors with experience in ca transcription measured the time it took to transcribe a conversation from scratch compared to improving a ‘first draft’ transcript produced by gailbot. we selected a random conversation recording from the icc. the co-author transcribed the first five minutes of the conversation manually and the second five minutes using gailbot output, using clan (macwhinney, 2000) to transcribe the data. the transcriber focused on words, timing, pauses, overlap markers, and speech rate changes on the manual transcript to match the level of detail provided in gailbot’s output.8 it took 179 minutes to transcribe the first five minutes of the conversation from scratch without using gailbot. in contrast, it took 63 minutes to correct a gailbot transcription of the second five minutes of the conversation. this is a substantial reduction, with gailbot reducing transcription time by approximately two thirds. however, we note that the exact amount of time savings using gailbot will depend on the expertise of the transcriber (expert ca transcribers are faster than novices), the purpose of the transcript, the speaker accent or dialect (asr accuracy decreases for speaker dialects not seen during training), as well as the frequency of overlaps. in short, using gailbot substantially reduces the time needed for ca transcription, but the exact amount of time reduction will depend on the particular conversation and transcriber. 8this level of detail is lower than for manually annotated ca transcripts, which also include annotations for a number of prosodic features that are not yet included in gailbot. this means that our manual transcription took less time than a full manual ca transcript, which influences the comparison. 80 https://sites.tufts.edu/hilab/files/2020/06/excerpt-12-committed.mp3 an automatic transcription system for conversation analysis 3.4 comparison with data annotation systems desirable features tools gailbot labelstudio praat elan speech to text default-auto enhanceable-auto manual manual turn construction default-auto enhanceable-auto manual manual silences default-auto enhanceable-auto default-auto manual overlapping speech default-auto manual manual manual syllable rate default-auto manual manual manual laughter detection default-auto enhanceable-auto enhanceable-auto manual table 1: comparison of automatic feature annotation capabilities between gailbot and existing data annotation systems. gailbot aims to produce human-readable transcripts of dialogue in a format designed by conversation analysts to refine and interpret qualitative data (ayaß, 2015; ochs, 1979). this differs from data annotation systems that aim to produce categorical annotations, usually for computational linguistics or machine learning applications (e.g., apostolova et al. 2010; stenetorp et al. 2012; yimam et al. 2013; klie et al. 2018; tkachenko et al. 2020-2021). data annotation systems exist in various domains including machine learning (tkachenko et al., 2020-2021), phonetic analysis (boersma and weenink, 2009), and gesture analysis (wittenburg et al., 2006). table 1 compares gailbot capabilities with data annotation tools9 selected from various domains. the table includes a list of features desirable for automatically producing ca transcripts. each tool may perform automatic annotation by default (default-auto), require enhancements for automation (enhanceable-auto), or only support manual annotations (manual). note that enhancements to a tool may range from additional package installations to implementing complex algorithms. additionally, this comparison only identifies each tool’s ability to annotate a feature, not its ability to integrate these annotations into ca transcripts that use special symbols, as gailbot does10. the comparison highlights that, in principle, it might be possible to emulate ca transcripts by using generic data labelling systems to segment and annotate asr transcripts, before exporting them using a template to generate ca-style symbols. however, enhancements to existing annotation systems may require fundamental software changes, which in turn can require significant amounts of additional resources. instead, gailbot provides out of the box capabilities for generating first draft ca transcripts. 3.5 summary we evaluated gailbot across a range of mono and stereo audio sources and found that lower quality and/or noisy recordings resulted in higher error rates across the board. we know of no transcription system – including gailbot – that performs speaker diarization well enough to reliably identify overlapping speech on a single audio channel. gailbot was most accurate when transcribing high 9our determination of each tool’s capability is based on publicly available documentation. 10a summary of jeffersonian transcription symbols is provided in appendix a. 81 umair, mertens, albert, and de ruiter quality stereo audio in low noise environments. additionally, we find that word recognition and timing errors introduced by external stt services in gailbot decrease accuracy when identifying paralinguistic features. as asr technology and speaker diarization in stt services develop further, gailbot’s ability to identify paralinguistic features in low quality data can be expected to improve accordingly. finally, gailbot rarely identified the exact location to place overlap markers or the exact duration of silences. it often underestimated the duration of silences by approximately 100 ms to 200 ms and the location of overlap markers by 0-5 characters. although these errors were relatively minor, some (e.g., overlap errors) had a significant effect on the accuracy of the transcripts and require manual correction for in-depth studies. nonetheless, given gailbot’s framework for improvement, and the richness of its transcripts relative to existing automatically generated transcripts, it still produces valuable first draft ca transcripts. 4. discussion gailbot is an extensible, customizable framework for transcribing paralinguistic features of conversation. it produces first draft ca transcripts that are adequate for preliminary analyses and for identifying fragments of interest requiring more accurate manual transcription. as we learn more about how transcribers identify paralinguistic features, and as nlp technology improves, gailbot will become even more accurate and useful over time. however, gailbot’s accuracy is currently limited by its plugins, each producing significant errors. one method of improving performance may be to use data-driven instead of heuristic-based approaches. for example, a recent deep neural network (dnn) based end-of-turn detection model achieved an 82% accuracy (masumura et al., 2018). however, training models is extremely timeconsuming and often necessitates parallelization, which may not be accessible to some end users with limited resources (narayanan et al., 2019). additionally, there are a limited number of corpora of jeffersonian transcripts available, especially given that training these models would require the time-synced audio along with the transcripts. this is a barrier to developing models that identify complex conversational features. with this in mind, our goal in developing gailbot was twofold – to provide user-friendly software for generating first draft ca transcripts, and to provide an adaptable framework to integrate and improve the underlying models as they are developed by the research community. more broadly, automated transcription software can only go so far (ogden, 2015; bolden, 2015). as bolden (2015) argues, even if perfect automatically generated ca transcripts were available, researchers should still review and familiarize themselves with their data as a part of their analytic workflow. another concern articulated by bolden (2015) is that researchers may change their research agendas, including the type of data they collect, to use automatic transcription. this concern is complicated by the bias inherent in asr (ferrer et al., 2021). while google’s stt claims to recognize over 70 languages and over 120 different local dialects and accents (barnes, 2020), the accuracy of asr when transcribing “nonstandard” dialects is poor (varis et al., 2021). standard dialects are those that have been institutionalized by the culture. for example, asr systems make twice as many errors when recognizing speech produced by african americans than by caucasians (koenecke et al., 2020). in addition, automatic sign language recognition lags behind spoken language recognition – the best systems have around a 30% wer (adaloglou et al., 2020). if automated systems make it cheaper and easier to study certain kinds of data, some researchers may avoid studying informal, overlapping talk, sign language, and/or interaction in marginalized groups 82 an automatic transcription system for conversation analysis and languages – issues that ca already struggles with due, in part, to its emergence in a particular socio-historical context (hoey and raymond, frth.). in addition, gailbot plugins likely also contain unrecognized biases about how interaction is ‘supposed’ to work. for example, aboriginal australians may tolerate longer silences than anglo-australians and americans (mushin and gardner, 2009); the standard thresholds for pauses, micropauses and gaps may be inappropriate assumptions for some cultures. this is why it is important that researchers can change the default parameter settings in gailbot. however, while we acknowledge that gailbot, like any other research tool, can potentially be misused, this should not prevent us from using and developing such tools. it is the shared responsibility of the research community to keep scientists accountable for their methodologies, tools and study designs. in addition, it is our shared responsibility to attempt to improve automated transcription for so-called “nonstandard” dialects and sign languages. gailbot dramatically improves on previous automated transcription software for the purposes of ca research. automated ca transcription and the availability of large-scale ca corpora will unlock many new research ideas and opportunities for interdisciplinary research, ranging from computational replications and extensions of conversation-analytic findings to understanding how participants in interaction employ paralinguistic features of talk. future research may also improve services in existing domains such as service call quality assurance, clinician-patient interaction, and in designing and evaluating language interventions (antaki, 2011). we invite researchers from across the social and computational sciences to use and contribute to the development of gailbot to help achieve this goal. acknowledgements this work was partly funded by afosr grant fa9550-18-1-0465, as well as by the school of arts & sciences and the school of engineering at tufts university. references martı́n abadi, paul barham, jianmin chen, zhifeng chen, andy davis, jeffrey dean, matthieu devin, sanjay ghemawat, geoffrey irving, michael isard, manjunath kudlur, josh levenberg, rajat monga, sherry moore, derek g. murray, benoit steiner, paul tucker, bijay vasudevan, pete warden, martin wicke, yuan yu, and xioaqiang zheng. tensorflow: a system for largescale machine learning. in 12th usenix symposium on operating systems design and implementation (osdi 16), 2016. isbn 9781931971331. doi: 10.1016/0076-6879(83)01039-3. nikolas adaloglou, theocharis chatzis, ilias papastratis, andreas stergioulas, georgios th papadopoulos, vassia zacharopoulou, george j xydopoulos, klimnis atzakas, dimitris papazachariou, and petros daras. a comprehensive study on sign language recognition methods. arxiv preprint arxiv:2007.12530, 2, 2020. charles antaki. applied conversation analysis: intervention and change in institutional talk. palgrave macmillan, london, 2011. liana g apostolova, lisa mosconi, paul m thompson, amity e green, kristy s hwang, anthony ramirez, rachel mistur, wai h tsui, and mony j de leon. subregional hippocampal atrophy 83 umair, mertens, albert, and de ruiter predicts alzheimer’s dementia in the cognitively normal. neurobiology of aging, 31(7):1077– 1088, 2010. ruth ayaß. doing data: the status of transcripts in conversation analysis. discourse studies, 17 (5):505–528, 2015. calum barnes. announcing new features, models, and languages for speech-totext, march 2020. url https://cloud.google.com/blog/products/ai-machine-learning/ new-features-models-and-languages-for-speech-to-text/. stig nikolaj blomberg, fredrik folke, annette kjær ersbøll, helle collatz christensen, christian torp-pedersen, michael r. sayre, catherine r. counts, and freddy k. lippert. machine learning as a supportive tool to recognize cardiac arrest in emergency calls. resuscitation, 138(october 2018):322–329, 2019. issn 18731570. doi: 10.1016/j.resuscitation.2019.01.015. paul boersma and david weenink. praat: doing phonetics by computer (version 5.1.13), 2009. url http://www.praat.org. christian bokhove and christopher downey. automated generation of ‘good enough’ transcripts as a first step to transcription of audio-recorded data. methodological innovations, 11(2): 2059799118790743, may 2018. issn 2059-7991. doi: 10.1177/2059799118790743. url https://doi.org/10.1177/2059799118790743. publisher: sage publications ltd. galina b. bolden. transcribing as research: “manual” transcription and conversation analysis. research on language and social interaction, 48(3):276–280, 2015. issn 0835-1813. doi: 10.1080/08351813.2015.1058603. carolyn boyce and p neale. conducting in-depth interviews: a guide for designing and conducting in-depth interviews. evaluation, 2(may):1–16, 2006. issn 1461-6734. steven e. clayman. turn constructional units and the transition relevance place. in the handbook of conversation analysis, pages 150–166. john wiley & sons, 2012. jan p. de ruiter, holger mitterer, and n. j. enfield. projecting the end of a speaker’s turn: a cognitive cornerstone of conversation. language, 82(3):515–535, 2006. issn 00978507. doi: 10.1353/lan.2006.0130. paul drew. ”quit talking while i’m interrupting”: a comparison between positions of overlap onset in conversation. in m. haakana, m. laakso, and j. lindstrom, editors, talk in interaction: comparative dimensions, pages 70–93. finnish literature society, 2009. ryan epp. big phoney. 2018. xavier ferrer, tom van nuenen, jose m. such, mark coté, and natalia criado. bias and discrimination in ai: a cross-disciplinary perspective. ieee technology and society magazine, 40 (2):72–80, june 2021. issn 1937-416x. doi: 10.1109/mts.2021.3056293. conference name: ieee technology and society magazine. peter french and john local. turn-competitive incomings. journal of pragmatics, 7(1):17–38, 1983. issn 03782166. doi: 10.1016/0378-2166(83)90147-9. 84 https://cloud.google.com/blog/products/ai-machine-learning/new-features-models-and-languages-for-speech-to-text/ https://cloud.google.com/blog/products/ai-machine-learning/new-features-models-and-languages-for-speech-to-text/ http://www.praat.org https://doi.org/10.1177/2059799118790743 an automatic transcription system for conversation analysis john j godfrey, edward c holliman, and jane mcdaniel. switchboard: telephone speech corpus for research and development. in acoustics, speech, and signal processing, ieee international conference on, volume 1, pages 517–520. ieee computer society, 1992. christian heath, jon hindmarsh, and paul luff. video in qualitative research. sage publications, 2010. alexa hepburn and galina b. bolden. transcribing for social research. sage, 2017. alexa hepburn and scott varney. beyond ((laughter)): some notes on transcription. in studies of laughter in interaction, pages 25–38. bloomsbury academic, 2013. elliott m. hoey and chase wesley raymond. managing conversation analysis data. in andrea berez-kroeker, brad mcdonnell, and eve koller, editors, the open handbook of linguistic data management. mit press, frth. g jefferson. notes on a possible metric which provides for a standard maximum silence of approximately one second in conversation (tilburg papers in language and literature 42). tilburg, netherlands: tilburg university, 1983. gail jefferson. glossary of transcript symbols with an introduction. in conversation analysis: studies from the first generation, pages 13–31. john benjamins, 2004a. doi: 10.1075/pbns.125. 02jef. gail jefferson. a sketch of some orderly aspects of overlap in natural conversation. in gene h. lerner, editor, 2004, pages 43–60. john benjamins publishing company, 2004b. gail jefferson. cabank english jefferson nb corpus [data set]., 2007. biing-hwang juang and lawrence r rabiner. automatic speech recognition–a brief history of the technology development. georgia institute of technology. atlanta rutgers university and the university of california. santa barbara, 1:67, 2005. leelo keevallik and richard ogden. sounds on the margins of language at the heart of interaction. research on language and social interaction, 53(1):1–18, 2020. issn 08351813. doi: 10.1080/ 08351813.2020.1712961. jan-christoph klie, michael bugert, beto boullosa, richard eckart de castilho, and iryna gurevych. the inception platform: machine-assisted and knowledge-oriented interactive annotation. in proceedings of the 27th international conference on computational linguistics: system demonstrations, pages 5–9, 2018. allison koenecke, andrew nam, emily lake, joe nudell, minnie quartey, zion mengesha, connor toups, john r rickford, dan jurafsky, and sharad goel. racial disparities in automated speech recognition. proceedings of the national academy of sciences, 117(14):7684–7689, 2020. judith c. lapadat and anne c. lindsay. transcription in research and practice: from standardization of technique to interpretive positionings. qualitative inquiry, 5(1):64–86, 1999. issn 10778004. doi: 10.1177/107780049900500104. 85 umair, mertens, albert, and de ruiter christophe leys, christophe ley, olivier klein, philippe bernard, and laurent licata. detecting outliers: do not use standard deviation around the mean, use absolute deviation around the median. journal of experimental social psychology, 49(4):764–766, 2013. issn 00221031. doi: 10.1016/j.jesp.2013.03.013. anthony j. liddicoat. an introduction to conversation analysis continuum. the tower building, new york, 2007. brian macwhinney. the childes project: tools for analyzing talk. transcription format and programs, volume 1. psychology press, 2000. ryo masumura, tomohiro tanaka, atsushi ando, ryo ishii, ryuichiro higashinaka, and yushi aono. neural dialogue context online end-of-turn detection. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 224–228, 2018. bahman mirheidari, daniel blackburn, kirsty harkness, and traci walker. towards the automation of diagnostic conversation analysis in patients with memory complaints. journal of alzheimer’s disease, pages 1387–2877, 2017. gilad mishne, david carmel, ron hoory, alexey roytman, and aya soffer. automatic analysis of call-center conversations. in international conference on information and knowledge management, proceedings, may 2014, pages 453–459, 2005. isbn 1595931406. doi: 10.1145/1099554.1099684. robert j moore. automated transcription and conversation analysis. research on language and social interaction, 48(3):253–270, 2015. doi: 10.1080/08351813.2015.1058600. ilana mushin and rod gardner. silence is talk: conversational silence in australian aboriginal talk-in-interaction. journal of pragmatics, 41(10):2033–2052, 2009. venkata srikanth nallanthighal, zohreh mostaani, aki härmä, helmer strik, and mathew magimai-doss. deep learning architectures for estimating breathing signal and respiratory parameters from speech recordings. neural networks, 141:211–224, september 2021. issn 08936080. doi: 10.1016/j.neunet.2021.03.029. url https://www.sciencedirect.com/science/article/ pii/s0893608021001179. deepak narayanan, aaron harlap, amar phanishayee, vivek seshadri, nikhil r. devanur, gregory r. ganger, phillip b. gibbons, and matei zaharia. pipedream: generalized pipeline parallelism for dnn training. in proceedings of the 27th acm symposium on operating systems principles, sosp ’19, page 1–15, new york, ny, usa, 2019. association for computing machinery. isbn 9781450368735. doi: 10.1145/3341301.3359646. url https://doi.org/10.1145/ 3341301.3359646. scott novotney and chris callison-burch. cheap, fast and good enough: automatic speech recognition with non-expert transcription. in human language technologies: the 2010 annual conference of the north american chapter of the association for computational linguistics, pages 207–215, los angeles, california, june 2010. association for computational linguistics. url https://aclanthology.org/n10-1024. 86 https://www.sciencedirect.com/science/article/pii/s0893608021001179 https://www.sciencedirect.com/science/article/pii/s0893608021001179 https://doi.org/10.1145/3341301.3359646 https://doi.org/10.1145/3341301.3359646 https://aclanthology.org/n10-1024 an automatic transcription system for conversation analysis elinor ochs. planned and unplanned discourse. in discourse and syntax, pages 51–80. brill, 1979. richard ogden. data always invite us to listen again: arguments for mixing our methods. research on language and social interaction, 48(3):271–275, 2015. issn 08351813. doi: 10.1080/ 08351813.2015.1058601. tae jin park, naoyuki kanda, dimitrios dimitriadis, kyu j. han, shinji watanabe, and shrikanth narayanan. a review of speaker diarization: recent advances with deep learning. arxiv:2101.09624 [cs, eess], june 2021. url http://arxiv.org/abs/2101.09624. arxiv: 2101.09624. danielle m pillet-shore. preference organization. the oxford research encyclopedia of communication, 2017. anita pomerantz. agreeing and disagreeing with assessments: some features of preferred/dispreferred turn shapes. in structures of social action, pages 57–101. cambridge university press, cambridge, 1984. isbn 0521318629. doi: 10.1017/cbo9780511665868.008. derek roger, peter bull, and sally smith. the development of a comprehensive system for classifying interruptions. journal of language and social psychology, 7(1):27–34, 1988. issn 15526526. doi: 10.1177/0261927x8800700102. kimiko ryokai, elena durán lópez, noura howell, jon gillick, and david bamman. capturing, representing, and interacting with laughter. in proceedings of the 2018 chi conference on human factors in computing systems, pages 1–12, 2018. isbn 9781450356206. doi: 10.1145/3173574.3173932. harvey sacks, emanuel a. schegloff, and gail jefferson. a simplest systematics for the organization of turn-taking for conversation. language, 50(4):696–735, 1974. doi: 10.2307/412243. george saon, gakuto kurata, tom sercu, kartik audhkhasi, samuel thomas, dimitrios dimitriadis, xiaodong cui, bhuvana ramabhadran, michael picheny, lynn-li lim, et al. english conversational telephone speech recognition by humans and machines. arxiv preprint arxiv:1703.02136, 2017. emanuel a. schegloff. the relevance of repair to syntax-for-conversation. syntax and semantics, volume 12: discourse and semantics, pages 261–286, 1979. emanuel a schegloff. discourse as an interactional achievement: some uses of ‘uh huh’and other things that come between sentences. analyzing discourse: text and talk, 71:93, 1982. hagen soltau, george saon, and brian kingsbury. the ibm attila speech recognition toolkit. 2010 ieee workshop on spoken language technology, slt 2010 proceedings, pages 97–102, 2010. doi: 10.1109/slt.2010.5700829. pontus stenetorp, sampo pyysalo, goran topić, tomoko ohta, sophia ananiadou, and jun’ichi tsujii. brat: a web-based tool for nlp-assisted text annotation. in proceedings of the demonstrations at the 13th conference of the european chapter of the association for computational linguistics, pages 102–107, 2012. 87 http://arxiv.org/abs/2101.09624 umair, mertens, albert, and de ruiter tanya stivers, n. j. enfield, penelope brown, christina englert, makoto hayashi, trine heinemann, gertie hoymann, federico rossano, jan peter de ruiter, kyung-eun yoon, and stephen c. levinson. universals and cultural variation in turn-taking in conversation. proceedings of the national academy of sciences, 106(26):10587–10592, 2009. issn 0027-8424. doi: 10.1073/pnas.0903616106. url http://www.pnas.org/cgi/doi/10.1073/pnas.0903616106. paul stonehouse. the unnecessary prescription of transcription: the promise of audio-coding in interview research. research in outdoor education, 17:1–19, 2019. issn 2375-5830. url https://www.jstor.org/stable/10.1353/reseoutded.17.2019.0001. publisher: cornell university press. susan a tilley. “challenging” research practices: turning a critical lens on the work of transcription. qualitative inquiry, 9(5):750–773, 2003. maxim tkachenko, mikhail malyuk, nikita shevchenko, andrey holmanyuk, and nikolai liubimov. label studio: data labeling software, 2020-2021. url https://github.com/heartexlabs/ label-studio. open source software available from https://github.com/heartexlabs/label-studio. juergen trouvain, raphael werner, and bernd möbius. an acoustic analysis of inbreath noises in read and spontaneous speech. in 10th international conference on speech prosody 2020, pages 789–793. isca, may 2020. doi: 10.21437/speechprosody.2020-161. url http://www. isca-speech.org/archive/speechprosody 2020/abstracts/168.html. erika varis, ryan georgi, alicia tsai, antonios anastasopoulos, kyathi chandu, xanda schofield, surangika ranathunga, haley lepp, and tirthankar ghosal. proceedings of the fifth workshop on widening natural language processing. in proceedings of the fifth workshop on widening natural language processing, 2021. johannes wagner. conversation analysis: transcriptions and data. in the encyclopedia of applied linguistics, pages 1–8. american cancer society, 2020. isbn 978-1-4051-9843-1. doi: 10.1002/ 9781405198431.wbeal0215.pub2. gareth walker. pitch and the projection of more talk. research on language and social interaction, 50(2):206–225, 2017. issn 08351813. doi: 10.1080/08351813.2017.1301310. peter wittenburg, hennie brugman, albert russel, alex klassmann, and han sloetjes. elan: a professional framework for multimodality research. in proceedings of the fifth international conference on language resources and evaluation (lrec’06), genoa, italy, may 2006. european language resources association (elra). url http://www.lrec-conf.org/proceedings/ lrec2006/pdf/153 pdf.pdf. li-chiung yang. duration and pauses as cues to discourse boundaries in speech. in speech prosody 2004, international conference, 2004. seid muhie yimam, iryna gurevych, richard eckart de castilho, and chris biemann. webanno: a flexible, web-based and visually supported system for distributed annotations. in proceedings of the 51st annual meeting of the association for computational linguistics: system demonstrations, pages 1–6, 2013. 88 http://www.pnas.org/cgi/doi/10.1073/pnas.0903616106 https://www.jstor.org/stable/10.1353/reseoutded.17.2019.0001 https://github.com/heartexlabs/label-studio https://github.com/heartexlabs/label-studio http://www.isca-speech.org/archive/speechprosody_2020/abstracts/168.html http://www.isca-speech.org/archive/speechprosody_2020/abstracts/168.html http://www.lrec-conf.org/proceedings/lrec2006/pdf/153_pdf.pdf http://www.lrec-conf.org/proceedings/lrec2006/pdf/153_pdf.pdf an automatic transcription system for conversation analysis appendix a: selected jeffersonian notation features feature symbol overlap ⌈, ⌊ = start of overlap ⌉, ⌋ = end of overlap pauses/gaps (.) = micropause (0.4) = pause or gap of approximately 400 ms changes in speech rate >word<= start/end of fast speech = start/end of slow speech syllable elongation : = extended by one beat latching ≈ current turn is latched to the previous turn laughter .h = inhalation h = exhalation volume ◦ = start or end of quiet speech changes in pitch ↑, ↓ = shift to high or low pitch ↗ = rising to high / mid pitch ↘ = falling to low or mid pitch . = falling turn-terminal pitch ? = rising turn-terminal 89 umair, mertens, albert, and de ruiter appendix b: heuristics and thresholds in gailbot heuristic default value micropause 100 <= duration <= 200 ms (after rounding) pause 200 ms ≤ duration ≤ 1000 ms (after rounding) within-speaker gap 1000 ms ≤ duration (after rounding) between-speaker gap 300 ms ≤ duration (after rounding) syllable rate mads from median 2 90 an automatic transcription system for conversation analysis appendix c: overlap algorithm for all turn pairs, do the following: 1. calculate the time between the start times of two consecutive turns. 2. if the result from step 1 is zero: (a) place the overlap-start markers (⌈, ⌊) at the start of both turns. 3. otherwise: (a) calculate the duration of both turns. (b) for each turn, calculate the proportion of the turn duration that occurs before the overlap begins. specifically, divide the overlap time by the total time of each turn. (c) for each turn, calculate the number of characters in the turn. (d) calculate the proportion of the turn characters that may occur before the overlap begins. specifically, multiply the output from line 4 and the output from line 5. (e) place the start overlap markers after those characters. 4. calculate the time between the end times of two consecutive turns. 5. if the result from step 4 is zero: (a) place the overlap-end markers (⌉, ⌋) at the end of both turns. 6. otherwise: (a) use the calculations from 3a-c to calculate the proportion of turn characters that may occur after the overlap ends. (b) place the end overlap markers before those characters. 91 umair, mertens, albert, and de ruiter appendix d: transcripts to compare to excerpt 9 transcript 1: jefferson (2007) transcript. 78 *lot: oh: ↓go:d a lo:ng wee⌈k. yeah.⌉ 79 *emm: ⌊ oh: my ⌋ ↓ god 80 i’m (.) glad it’s over i won’t even turn the 81 teevee o⌉:n. 82 *lot: ⌊i won’eether. 83 *emm: ◦aoh no. they drag it out so◦ that’s where they 84 we took off on ar chartered flight that sa:me 85 spot didju see it↗ 86 (0.7) 87 *emm: .hh when they took him in⌈the airpla:ne,⌉ 88 *lot: ⌊ n:no:::. ⌋ hell 89 i wouldn’ ev’n wa:tch it. 90 *lot: i think it’s so ridiculous. i mean it’s .hhh 91 it’s a hôrrible thing but my: go:d. play up 92 that thing it it’s jst ↑ hôrri⌈ble. ⌉ 93 *emm: ⌊it’ll⌋ drive 94 people nu:ts. 95 *lot: why id ı̈-en makes americ’n people think why ther 96 no goo:d. 97 *emm: ◦◦mm:◦◦ well they aren’t very good some of’m, transcript 2: moore (2015) transcript. 69 oh god long week 70 oh my god 71 i’ve decided sober i want you to have a t.v. 72 73 i won’t either 73.5 (0.7) 74 like uh you know (0.1) that’s where they 75 we took off on our charter flight that same spot 76 did you see it 92 https://sites.tufts.edu/hilab/files/2021/07/assassination_file_trimmed_gb_paper.mp3 https://sites.tufts.edu/hilab/files/2021/07/assassination_file_trimmed_gb_paper.mp3 an automatic transcription system for conversation analysis 77 (0.8) 78 and they took him and here uh you 79 know i wouldn’t 80 watch it 81 i think it’s so ridiculous i mean it’s (0.4) it’s a 82 horrible thing but my god (0.1) play up that’s thing 83 it’s it’s (.) horrible 84 die people that 84.5 (0.3) 85 why is it a native american people think well they’re 86 no good 86.5 (0.5) 87 well they aren’t very good some of 93 umair, mertens, albert, and de ruiter appendix e: transcripts to compare to excerpt 10 transcript 1: callhome corpus transcript 4686. 42 *b: did they go through the theories of the three 43 bullets or the magic one bullet? 44 *a: yeah four bulyeah. 44 *a: it was interesting. transcript 2: walker (2017), excerpt 2. 1 *a: [(but) yeh] 2 *b: [did (.) ] they go through the theories of the 3 three bullets, (or/and) the magic one bullet, 4 *a: yea:h. (.) four <bull>. yeahò (0.9) it 5 was? intresting. 94 https://sites.tufts.edu/hilab/files/2021/07/callhome_4686_excerpt_8.mp3 https://sites.tufts.edu/hilab/files/2021/07/callhome_4686_excerpt_8.mp3 an automatic transcription system for conversation analysis appendix f: transcripts to compare to excerpt 11 transcript 1: callhome corpus transcript 4247. 37 *a: clinton just came out and said that he doesn’t 38 believe in quota systems . 39 *a: and in reverse discrimination but that he does 40 believe that affirmative action . 41 *a: is necessary &=inhales to move &uh you know black 42 americans &=inhale . 43 *a: forward and to give them the opportunities that 44 they’ve been denied. transcript 2: walker (2017), excerpt 3. 1 *a: clinton just came out and sai:d that he:: (0.2) 2 doesn’t believe ◦h (0.2) in quota systems ◦h (0.2) 3 and (.) in reverse discrimination.=but that he 4 does believe that affirmative action is necessaryò 5 ◦h to mo:ve uh:? (.) you know black americans ◦h 6 forward and to give them the opportunities that 7 they’ve been denied. 95 https://sites.tufts.edu/hilab/files/2021/07/callhome_4247_excerpt_9.mp3 https://sites.tufts.edu/hilab/files/2021/07/callhome_4247_excerpt_9.mp3 introduction architecture and algorithms architecture data flow and internal representations plugins algorithms turns silences overlapping speech syllable rate laughter performance metrics word error rate relaxed word errors strict word errors overlap and silence performance syllable rate performance turn taking performance laughter performance speaker diarization performance evaluation newport beach corpus callhome corpus in conversation corpus transcription time estimation comparison with data annotation systems summary discussion dialogue & discourse 16(2) (2025) 74–110 doi: 10.5210/dad.2025.203 enhancing long-term rag chatbots with psychological models of memory importance and forgetting ryuichi sumida sumida@sap.ist.i.kyoto-u.ac.jp graduate school of informatics kyoto university koji inoue inoue@sap.ist.i.kyoto-u.ac.jp graduate school of informatics kyoto university tatsuya kawahara kawahara@i.kyoto-u.ac.jp graduate school of informatics kyoto university editor: david traum submitted 01/2025; accepted 11/2025; published online 12/2025 abstract this study addresses the issue of what a retrieval-augmented generation (rag) chatbot should remember and what it should forget, based on findings from psychology. rag retrieves relevant memories from past interactions to generate responses, and its effectiveness has been demonstrated. as conversations continue, however, the amount of stored memory keeps growing, which not only requires large storage capacity but also risks retaining unnecessary information, potentially deteriorating retrieval performance. to tackle this problem, we propose lufy (long-term understanding and identifying key exchanges), a rag chatbot that evaluates six distinct memory-related metrics derived from psychological models and real-world data. instead of simply summing these metrics, it uses learned weights to determine the contribution of each one. by using these weighted scores, the system can prioritize and retain relevant memories while gradually forgetting less important ones during both retrieval and memory management. to evaluate the effectiveness of lufy in long-term conversations, we conducted experiments with human participants, who engaged in textbased conversations with three types of chatbots, each using different forgetting mechanisms, for at least two hours. the length of these conversations was more than 4.5 times longer than the longest conversations reported in previous studies. the results showed that prioritizing emotionally engaging memories while forgetting most of the conversation significantly enhanced user satisfaction. keywords: retrieval-augmented generation (rag), long-term conversational ai, forgetting mechanism, psychological models 1. introduction in human conversations, what we remember remains a mystery. similarly, in the realm of chatbots, particularly retrieval-augmented generation (rag) chatbots, the challenge lies not only in retaining long-term memory but also in deciding what to forget. rag retrieves related memories from past conversations and uses them to generate a response, and it has been shown to be effective (xu et al., 2022). compared to inputting the entire conversation ©2025 ryuichi sumida, koji inoue, tatsuya kawahara this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). enhancing long-term rag chatbots with psychological models dataset avg. turns avg. words conv. setting daily dialog (li et al., 2017) 7.9 115.3 human-human soda (kim et al., 2023) 7.6 122.4 chatbot-chatbot carecall (bae et al., 2022) 104.5 515.2 human-chatbot conversation chronicles (jang et al., 2023) 58.5 1,054.8 chatbot-chatbot msc (xu et al., 2022) 53.3 1,225.9 human-human lufy-dataset (ours) 253.8 5,538.5 human-chatbot table 1: average turns and words per conversation of lufy-dataset compared to existing text-totext dialog datasets. the average length of a conversation in ours is 4.5x that of msc (xu et al., 2022), distributed over 4.8x more turns. history into a large language model (llm), retrieving relevant information is not only efficient and effective (yu et al., 2024) but also enhances the chatbot’s ability to remember and actively utilize the user’s past information, improving the level of engagement and contributing to long-term rapport between the chatbot and the user (campos et al., 2018). as highlighted in previous studies (choi et al., 2023), however, a challenge arises: as conversations progress, memory constantly increases. this not only demands significant storage space but also involves retaining a lot of unnecessary information, which could lead to degrading retrieval performance. a simple example would be that humans do not remember every single meal they have had in their entire lives; instead remember only their favorites or particularly memorable ones. this necessitates selective memory in chatbots, much like the human cognitive process, where only the important elements of a conversation are remembered. studies have shown that humans typically recall only about 10% of a conversation (stafford and daly, 1984), making it crucial to develop methods that approximate which 10% of memories are likely to be most useful to retain. this challenge raises the key question: how can chatbots efficiently identify and prioritize these valuable memories, while discarding irrelevant information? to address this problem, we propose lufy: long-term understanding and identifying key exchanges in one-on-one conversations. building upon real-world data and psychological insights, lufy improves existing rag models by calculating six distinct memory metrics, whereas the conventional simple models like memorybank (zhong et al., 2024) only account for frequency and recency of memory recall. additionally, lufy does not merely sum these metrics but uses learned weights to balance their impact. these weighted scores are then used in both the retrieval and forgetting modules, ensuring that the chatbot effectively prioritizes relevant memories while gradually forgetting less important ones. to empirically validate our proposed model and its effectiveness in enhancing long-term conversational memory, we conducted an experiment involving human participants. in this experiment, participants engaged in text-based conversations with chatbots equipped with different forgetting modules, each for at least two hours. this duration is at least 4.5 times longer than that of the most extensive existing studies of conversational abilities to date, as shown in table 1. our results show that prioritizing arousing memories while discarding most of the conversation content significantly enhances user experience. this enhancement is evaluated by human assessments, sentiment analysis, and gpt-4 evaluations. furthermore, our method improved the precision of information 75 ryuichi sumida, inoue and kawahara retrieval by over 17% compared to the naive rag system with no forgetting mechanism, highlighting its potential for more accurate and relevant conversational responses. we developed this comprehensive dataset of human-chatbot conversations along with annotations for the important utterances. our approach to data collection is designed to uphold rigorous privacy and ethical standards, with consenting participants and robust de-identification procedures where relevant. the organization of this paper is structured as follows: section 2 introduces findings from psychology and explains how to derive metrics from these findings. section 3 details methods for assessing the importance of user utterances using the metrics introduced in section 2. section 4 explores the integration of the importance metric and forgetting mechanisms into rag chatbots. section 5 presents the user experiment and its results. section 6 reviews relevant prior work in cognitive memory theory and conversational ai. this paper culminates with conclusions and future works presented in section 7. 2. psychology of memorizing conversations memorybank (zhong et al., 2024) introduced the idea of assessing the importance of a memory based on the number of times a memory is used and the elapsed time since the memory was last used. this idea was inspired by the ebbinghaus forgetting curve theory (ebbinghaus, 1885), a theory that the more times information is reviewed or learned, the slower the rate of forgetting. in this model, the importance of a memory is represented as: importance = e− ∆t s . (1) here, ∆t represents the time elapsed since the memory was last retrieved, and s is the number of times the memory is retrieved. in memorybank, the initial value of s is set to 1 upon its first mention in a conversation. when a memory item is recalled during conversations, s is increased by 1 and ∆t is reset to 0. nevertheless, as acknowledged in their study, this represents a preliminary and oversimplified model for memory’s importance updating. most importantly, their work did not include experiments to validate their proposed concept. despite the innovative approach taken by memorybank, the model’s simplicity overlooks critical aspects of how memories are valued and retained. this oversight becomes especially apparent when considering the concept of flashbulb memories, as described by (brown and kulik, 1977). these memories, such as the news of the 9/11 terrorist attacks, maintain their vividness and strength regardless of recall frequency. the current method of uniformly updating s fails to reflect the intricate nature of memory consolidation and the differential impact of emotionally charged events. in the following, we explain pivotal insights from psychology that offer valuable perspectives on conversations, particularly focusing on memory consolidation and retention mechanisms. additionally, we demonstrate how these findings can be applied to numerically score the importance of an utterance. first, emotional arousal greatly enhances memory consolidation (brown and kulik, 1977; conway et al., 1994; mcgaugh, 2003; reisberg and hertel, 2003), a phenomenon well-documented in studies like flashbulb memories (brown and kulik, 1977). emotionally charged events, such as dramatic personal experiences like a breakup or receiving college admissions, are more likely to be remembered. we used roberta (liu et al., 2019) finetuned with emobank (buechel and hahn, 2017), a 10k text corpus manually annotated with emotion, to measure arousal in user’s utterance. 76 enhancing long-term rag chatbots with psychological models second, the element of surprise plays a crucial role in memory retention. events that deviate from expectations are more memorable (breton-provencher et al., 2022). to quantify this element of surprise, we used the system’s utterance as context to evaluate the perplexity of the user’s utterance, a measure of how predictable a piece of text is. we employed the gpt2-large model to calculate the perplexity, using the chatbot’s utterance as context. third, the concept of retrieval-induced forgetting (rif) suggests that selectively recalling certain memories strengthens those memories while making related but unmentioned memories more likely to be forgotten (hirst and echterhoff, 2012). aligning our method with the concept of rif (hirst and echterhoff, 2012), we retrieved the top 2 memories despite the higher effectiveness observed with the top 5 (see appendix a for further analysis of top-k retrieval). the memory ranked as the most relevant (r1) is used for response generation to reinforce its strength, while the second most relevant memory (r2) is retrieved but not used for response generation to encourage forgetting, in accordance with rif principles. we note, however, that retrieving the top two memories relative to the user’s current utterance does not guarantee that r1 and r2 are semantically similar to each other. cosine similarity is computed relative to the query, not between memories themselves, so r2 may not represent a “competitor” memory in the strict psychological sense of retrieval-induced forgetting. our use of r2 therefore approximates the competitive retrieval structure of rif but does not assume semantic closeness between r1 and r2, and future work could refine this by identifying memories most similar to r1 directly. fourth, motivated by prior work suggesting that approximately 70% of conversational time is devoted to discussing contemporary events (dunbar et al., 1997), we incorporated recency as a key factor in our retrieval strategy (section 4.1.1). to evaluate this assumption, we conducted an empirical analysis of the dataset collected from our human experiment (described in chapter 5). our findings indicate that only 22.9% of user utterances reference contemporary events, substantially lower than the literature estimate. further methodological details and results are provided in appendix b. nevertheless, this proportion still constitutes a meaningful share of conversational content, justifying our inclusion of recency in the system design. lastly, the capacity for immediate recall in social interactions is surprisingly limited; studies show that individuals typically remember only about 10% of the content from a conversation (stafford and daly, 1984). while our model adopts this 10% capacity limit as a starting point, we do not claim it to be a psychologically essential threshold. rather, it serves as a first approximation that allows us to explore the consequences of limited memory, and future work could investigate alternative thresholds or adaptive selection strategies to refine this modeling choice. while our metrics focus on observable, surface-level utterance properties such as emotional arousal or surprise, we acknowledge that human memory in dialogue often involves more complex dynamics — including unspoken implications (fatihi, 2025), avoidance strategies (endler and parker, 1994), and relational cues (walther and tidwell, 1995). modeling such phenomena would require integrating theories of discourse structure, social cognition, or pragmatic inference, which remain beyond the scope of this study. finally, we note an important conceptual distinction relevant to how psychological theories are applied in this work. research on memory—such as emotional arousal, surprise, or flashbulb memory—typically concerns the encoding of real-world events. however, conversational ai systems do not observe events directly; they observe only their linguistic expression. in this study, we therefore treat an utterance pair as a practical proxy for an event. we acknowledge that a single event may unfold across multiple utterances, and likewise, an individual utterance may reference only part of 77 ryuichi sumida, inoue and kawahara an event. our goal is not to equate utterances with events, but to model the cues available to the system through language. thus, the psychological principles we draw upon are applied to expressed events rather than the events themselves. future work could extend this approach by modeling multi-utterance segments as unified event representations. 3. quantifying memory importance 3.1 metrics for assessing memory importance building upon the psychological principles of emotional arousal, surprise, and retrieval-induced forgetting, we propose the following metrics for determining the importance of memories. specifically, we translate these following insights into quantifiable metrics: • emotional arousal (a): measured using roberta (liu et al., 2019) fine-tuned with emobank (buechel and hahn, 2017), capturing the intensity of emotions in the user’s utterance. • surprise element (p ): assessed with perplexity using the gpt2-large model, representing the unpredictability of the utterance. • llm-estimated importance (l): derived from the language model’s estimation of the importance of the user’s utterance. • retrieval-induced forgetting (r1, r2): we define r1 and r2 as counters that track how often a memory was retrieved as the most relevant and the second most relevant, respectively. although both emotional arousal and surprise metrics are inspired by psychological theories tied to event memorability, our operationalization derives directly from the user’s linguistic expression. in particular, the arousal metric is computed from utterance text via emotion classification, and surprise is based on linguistic predictability (perplexity). as such, these measures reflect how events are expressed in speech, not just the events themselves. while the above mentioned metrics draw directly from well-established psychological findings, we incorporate an additional metric—llm-estimated importance—to capture broader contextual relevance and coherence that may not be easily quantifiable through arousal, surprise, or retrieval patterns alone. this score reflects an llm’s judgment about whether a given exchange will be useful or influential in future conversations, simulating a kind of narrative salience or pragmatic relevance. this component is inspired by the “generative agents” study (park et al., 2023), which demonstrated that llms can be used to simulate aspects of human memory and behavior in a plausible and interpretable manner. specifically, that study used llms to assess the salience of events in an agent’s life based on their likely future consequences. we adopt a similar prompting approach (see figure 1), framing importance in terms of both personal relevance and conversational utility. while not derived from a specific psychological theory, this llm-based metric acts as a learned heuristic that approximates human-like memory prioritization, especially in ambiguous cases where emotion or surprise is low but contextual relevance is high. it serves to complement, rather than replace, our psychologically grounded metrics. 78 enhancing long-term rag chatbots with psychological models on a scale of 1 to 10, where 1 is purely mundane (like brushing your teeth or making your bed) and 10 is extremely important (like a breakup or a college acceptance), please rate the importance of the following conversation. rate it based on whether it will be useful in later conversations or not. here is a summary of the past conversations: {key_summary} here is the conversation: {content} be sure to provide both a score and a brief reason for your rating. figure 1: prompt for llm-estimated importance figure 2: process of determining the importance of a memory using various psychological metrics. the weights wa, wp , wl, wr1, and wr2 are assigned based on the relative impact of each metric, as detailed in section 3. 3.2 normalization of metrics to ensure consistency across the different metrics, we apply a normalization process, ensuring all five metrics—a, p , l, r1, and r2—are scaled with the same average, minimum, and maximum values. 79 ryuichi sumida, inoue and kawahara among these metrics, perplexity (p ) is the only one that can reach notably high values. in cases where perplexity exceeds a threshold of 160, which covers more than 95% of utterances, the value is capped to avoid extreme outliers. this threshold was selected based on the distribution of perplexity values, but future research could refine this limit to better handle unusually high scores. using data from the human-rag chatbot conversations (a set of 200 utterance pairs), we then normalized all metrics to ensure uniform scaling. 3.3 definition of memory importance as illustrated in figure 2, the importance assignment process involves three key steps. step 1: metric calculation—for each chatbot-user utterance pair (defined as a memory), the psychological metrics outlined above are calculated. step 2: strength determination—next, the strength s is computed by summing the weighted metrics. s = wa · a + wp · p + wl · l+ wr1 ·r1− wr2 ·r2 (2) here, a, p , and l correspond to metrics for emotional arousal, surprise (perplexity), and llmestimated importance, respectively. r1 and r2 are counters indicating how often a memory was retrieved as the most or second-most relevant, respectively. r1 contributes positively to strength, reinforcing the memory, while r2 has a negative effect on strength, in accordance with retrievalinduced forgetting theory. thus, even though both r1 and r2 stem from retrieval events, they have opposite cognitive effects: r1 supports retention; r2 promotes forgetting. step 3: importance calculation—finally, using the calculated strength, the overall importance of a memory is determined. following the ebbinghaus forgetting curve formulation introduced in equation (1), we compute importance as: importance = e− ∆t s , (3) where ∆t is the time since the memory was last accessed, and s is the composite strength defined in equation (2). although structurally identical to equation (1), this formulation reflects a more operational use: s integrates multiple psychologically inspired metrics, going beyond access frequency to capture the nuanced salience of a memory. 3.4 weight estimation to quantify the relative impact of each memory metric (arousal, perplexity, etc.) on the overall strength, we need to estimate the weights (wa, wp , wl, wr1, wr2) in equation (2). we achieve this by fitting these parameters to real-world conversational data annotated for importance. in this study, to fit the three initial parameters (wa, wp , wl), part of the candor corpus (reece et al., 2023) (a total of 300 utterance pairs, transcribed from audio) was used. while keeping these three parameters fixed, we used human-rag chatbot conversations (a total of 200 utterance pairs), which were collected prior to the user experiment to fit the latter two parameters (wr1, wr2). the candor corpus was selected for its diverse set of conversations, while the human-rag conversations allowed for precise tracking of memory usage, which is not available in the candor corpus. for both datasets, we tasked annotators with labeling the conversations on a binary scale of 1 and 0, where 10% of the conversations were labeled as 1 (important) to mimic human behavior (stafford 80 enhancing long-term rag chatbots with psychological models metric symbol weight arousal wa 2.76 perplexity wp -0.28 llm estimated importance wl 0.44 most relevant memory count wr1 1.02 2nd most relevant memory count wr2 -0.012 table 2: estimated weights for the memory-related psychological metrics. arousal has the highest weight relative to other metrics. and daly, 1984). annotators were instructed to label utterances as important if they believed the information contained in the utterance would be used in future conversations. however, we also left some ambiguity, as the notion of what is ”important” is not clearly defined, and exploring this ambiguity is part of our study. to assess the reliability of our annotation process, we measured inter-annotator agreement using fleiss’ kappa (fleiss, 1971). for both the candor corpus and the human-rag chatbot conversations, three annotators independently labeled each utterance as either important or unimportant. the fleiss’ kappa score was 0.42 for both datasets, indicating moderate agreement among annotators. to connect the annotated importance labels (0 or 1) with the strength calculated using the three steps in section 3.3, we need to account for the lag time (∆t). in our setup, the lag time is defined as one time step. this is because the annotation occurs after the conversation (at time t = 2), and the last time the memory could have been used was during the conversation itself (at time t = 1). therefore, we substitute ∆t = 1 into equation (3). this simplifies the equation to: importance = e−1/s (4) where s is defined in equation (2) as the weighted sum of the metrics. with the simplified importance (equation (4)) and the annotated data, we can now estimate the weights. we used the levenberg–marquardt algorithm and utilized l2-regularized squared loss as the loss function, with p0 (initial guesses) to be [1, 1, 1], although we did not find the initial guesses to affect the outcome of the learned parameters. the results of the fitted parameters are given in table 2. notably, arousal (a) exhibits the highest weight among metrics, suggesting that arousal plays a crucial role in determining the memorability and importance of an utterance. the negative weight for perplexity (wp = −0.28) reflects that lower perplexity correlates with higher importance. this counterintuitive result stems from the nature of the utterances labeled as important. to better understand this, we used an llm to categorize each utterance as profile-related, episode-related, or others, using the prompt provided in appendix c.2. over half of these utterances (52.3%) contained profile-related information, such as names, preferences, or other biographical details, which are generally predictable and contextually coherent, resulting in lower perplexity (37.7) compared to the average perplexity of 39.5. similarly, a significant portion of the utterances (27.3%) were related to episodes or events, which also exhibited low perplexity (28.7) due to their alignment with prior conversational context. however, the weight of the perplexity is much smaller then those for others. that said, we also recognize limitations in using perplexity as a proxy for psychological 81 ryuichi sumida, inoue and kawahara surprise. although perplexity reflects how unexpected a sentence is to a language model, it does not necessarily align with how humans experience surprise or novelty. for example, the sentence “i saw a car accident right in front of my eyes” is syntactically conventional and therefore has low perplexity, but the event it describes is rare and emotionally salient. this highlights a key mismatch: perplexity captures linguistic predictability from large-scale training corpora, but not experiential rarity or semantic impact. thus, while our model finds a weak negative correlation between perplexity and importance in our dataset, this should not be interpreted as a general psychological principle. future work could explore metrics that better capture event-level novelty or incorporate richer contextual information beyond text. 4. system overview lufy is composed of two main components: (1) the response generation module, and (2) the forgetting module. the forgetting process is invoked after the conversation when the user has done talking. 4.1 response generation as depicted in figure 3, the response generation process consists of two main parts: the retrieval part and prompt construction via concatenation of key information (context, summary, and memory). we provide a detailed explanation of the retrieval method, as it is central to the rag chatbots. 4.1.1 retrieval method memory storage we store three types of information, which together constitute the chatbot’s memory: figure 3: overview of the response generation pipeline. the user input is embedded and compared against stored memory entries based on cosine similarity and an importance score. the top-ranked memory combined with recent utterances and a summary of prior conversations, is passed to the llm for response generation. 82 enhancing long-term rag chatbots with psychological models 1. profile information, consisting of structured facts about the user such as preferences, background traits, or frequently mentioned entities (e.g., “likes sushi,” “has a dog named luke”). the profile is automatically generated by an llm based on the accumulated conversation history. it is continuously updated after each session to reflect new salient user details. a ‘session’ refers to a 30-minute human–chatbot interaction (see section 5 for full details). 2. conversation summaries, which are generated at the end of each session to capture the overall structure and key events of the dialogue. 3. utterance pairs, consisting of user–bot exchanges, which preserve granular contextual and episodic details. we include summaries of the conversations and utterance pairs in memory storage to retain small details about the user that are unlikely to appear in the profile information, such as episodic memory (tulving, 2002), which is made up of past experiences or events they personally remember. the forgetting mechanism is applied to conversation summaries and utterance pairs, while profile entries are retained and refined over time. this structure enables the system to maintain both stable user traits and specific, temporally grounded memories. embedding we use a standard llm embedding (text-embedding-ada-002 from openai). embeddings of the summaries and utterance pairs are stored. retrieval the foundation of the rag system is its retrieval method. this process is initiated upon receiving input from a user, to retrieve the memory most relevant to the current conversation. there are two key aspects to the retrieval method: 1. cosine similarity threshold: during the retrieval stage, we use cosine similarity to assess the relevance between the embeddings of recent conversations and the entries in the memory database. to ensure a high level of contextual relevance, we have set the threshold at 0.8, consistent with commonly used standards in popular rag system frameworks like llamaindex. this choice is further validated by our analysis of question and answer pairs collected from the user experiment. the details are available in appendix d. 2. final retrieval score: our proposed method distinguishes itself from previous works by integrating importance into the retrieval process. the final retrieval score is computed using a formula as follows: score = cos. sim. + α · importance (5) since cosine similarity values are distributed between 0.8 and 0.9 due to our threshold (a range of 0.1), while importance scores range from 0 to 1, we set α = 0.1 to approximately normalize their influence in the final score. this choice does not assume that cosine similarity and importance are inherently of equal weight; rather, it heuristically balances the two terms given their different scales. we acknowledge that α is a tunable hyperparameter, and further empirical tuning—e.g., via ablation or grid search—could yield better retrieval performance. this remains an area for future exploration. 83 ryuichi sumida, inoue and kawahara you will be provided with 3 pieces of information: 1. key summary: a summary of past conversations . 2. recent utterances: the latest exchanges between you and the user. 3. relevant memory: the memory most pertinent to the current conversation. using these details, you are tasked with generating an effective response. ensure that your reply maintains a casual tone to mimic a genuine interaction with a friend. here’s the information: key summary of past conversations: {key_summary} recent utterances: {recent_utterances} memory relevant to current conversation: {related_memory} figure 4: prompt for generating chatbot responses. figure 5: overview of the forgetting process. after a conversation ends, memories are ranked by computed importance, and only the top 10% are retained. 4.1.2 information provided to the llm for response generation the llm is provided with three key pieces of information to generate responses: a summary of past conversations, the context (recent five utterances), and memory relevant to the current context. the prompt given to the llm to generate a response is shown in figure 4. 84 enhancing long-term rag chatbots with psychological models 4.2 forgetting process as depicted in figure 5, the forgetting process is executed in two steps: ranking the memories according to importance, and retaining the top 10% important memories. the forgetting process is executed only after the user finishes a conversation, ensuring that memory importance—calculated using retrieval frequency and other dynamic metrics—reflects the full interaction. retention rate as shown in table 3, memorybank and lufy remembered a very small portion, mostly less than 10% of the conversations. here, each session refers to a 30-minute conversation between a human participant and the chatbot, as described in section 5 (see section 5.1.1 for details). due to the discrete nature of memory chunks (i.e., both the total number of memories and the number of deleted items are whole numbers), the fraction of retained memories is not always exactly 10%, though it remains consistently close. for instance, if there are 14 memories and 10% corresponds to 1.4 items, we must round to the nearest whole number—e.g., retaining only 1 item. in this case, 1 out of 14 yields approximately 7.1%, which is slightly lower than the intended 10%. system s1 s2 s3 s4 number of retained memories naive rag 28.5 63.2 97.9 131.2 memorybank 2.8 6.4 9.8 13.5 lufy 2.8 6.1 9.7 12.8 retention rate (%) naive rag 100 100 100 100 memorybank 9.90 9.88 9.91 9.99 lufy 9.90 9.88 9.93 10.07 table 3: comparison of memory retention for naive rag, memorybank, and lufy across four sessions. the table shows the cumulative number of memories retained (top) and the corresponding retention rate in percent (bottom). system forgetting? retrieval method naive rag × cos. sim. memorybank (zhong et al., 2024) ✓ cos. sim. lufy ✓ cos. sim. + importance table 4: comparison of systems in terms of forgetting and retrieval methods. naive rag: a conventional rag chatbot model that exhibits no forgetting, memorybank: the model from previous work (zhong et al., 2024), and lufy: our proposed model. 85 ryuichi sumida, inoue and kawahara 5. user experiment we compared three different rag chatbots: naive rag, memorybank, and lufy. naive rag stores all memories, while memorybank and lufy are equipped with a forgetting mechanism that retains only 10% of memories. memorybank assesses the importance of a memory based solely on retrieval counts, whereas lufy evaluates memory importance using six memory-related metrics, as depicted in figure 2. additionally, although both naive rag and memorybank use cosine similarity only to assess the relevance of a memory to current conversations, lufy also considers importance in this assessment. the differences between the three systems are summarized in table 4. for an illustrative scenario highlighting how naive rag, memorybank, and lufy differ in memory retention and retrieval behavior, see appendix g. our study involved extended multi-day interactions that exceeded the length of prior dialog benchmarks (xu et al., 2022), as shown in table 1. 5.1 procedure 5.1.1 interaction phase seventeen participants each engaged in ten 30-minute conversations (hereafter referred to as “sessions”) over the course of 4 days. on day 1, participants interacted with a naive rag chatbot—a standard rag chatbot that uses cosine similarity to retrieve relevant documents based on query embeddings and has no forgetting process. the three chatbots—naive rag, memorybank, and lufy—were all equipped with identical memories for this first session, although the memories underwent different forgetting mechanisms. from day 2 to day 4, participants interacted with all three chatbots for 30 minutes each, in a randomized order, to control for ordering effects. in total, there were four sessions for each chatbot. however, because the first session (day 1) was identical for all three chatbots, participants engaged in a total of 10 sessions (1 + 3× 3). the day-1 chatbot was intentionally presented without a name. participants were told only that they would have a 30-minute conversation with “a chatbot,” without any persona label or identifying information. this ensured that they understood they were interacting with a single, unnamed system on day 1, rather than three separate bots. beginning on day 2, each of the three chatbots was assigned a distinct name, and participants were explicitly shown which named bot they were interacting with in each session. at the start of day 2, participants were explicitly informed that all three named chatbots had access to the same day-1 conversation history. they were also told that the systems used different underlying mechanisms, though the specific nature of these differences was not disclosed. this design made the three systems clearly distinguishable from day 2 onward while maintaining a neutral, shared starting point on day 1. each chatbot maintained a consistent identity across its four sessions, though none disclosed its underlying memory mechanism. participants generally treated the three systems as separate interlocutors, aided by their distinct names and the divergent conversational trajectories that naturally emerged with each bot. on occasions when participants were unsure which bot they were speaking with—particularly when several days (and in a few cases more than a week) had elapsed between sessions—we provided a brief summary of their past conversations with that specific bot at the start of the session. these summaries served as gentle reminders of the bot’s prior interactions and identity, helping participants engage with each system as a distinct conversational partner with its own history. 86 enhancing long-term rag chatbots with psychological models 5.1.2 post-interaction phase after each session, the participants were asked to create three question-and-answer (qa) pairs about the conversation they had just had. these questions contained information that the participants wanted the chatbot to remember about them. additionally, the questions were designed to be clearly judged on a binary scale (1: correct, 0: incorrect) and to ensure that the answers would remain consistent in future interactions. if the participants’ questions did not meet this criterion, they were asked to revise them. an example of a question that meets this criterion is ’what is my pet’s name?’ 5.2 lufy-dataset we developed a comprehensive dataset of human-chatbot conversations along with annotations for the important utterances, which we refer to as the lufy-dataset. similar to the annotation procedure described in section 3, at least three annotators reviewed and labeled the conversations from the user experiment using a binary scale: 1 for ”important” and 0 for ”unimportant.” the fleiss’ kappa score for inter-annotator agreement was 0.35, indicating fair agreement. in addition to the annotation of important memories, participants created three qa pairs for each conversation they had with the chatbot. thus, the qa pairs directly correspond to their respective conversations. examples are illustrated in table 5 and 6. for the de-identification process, multiple reviewers examined the conversations to ensure that all personally identifiable information (pii) was removed or modified. removal of direct identifiers this involves eliminating information that can directly identify an individual, such as names, phone numbers, and addresses. generalization specific data points are replaced with broader categories to balance privacy and data utility. for instance, exact ages might be replaced with age ranges (e.g., “29” becomes “20–30”), or full dates (e.g., “april 15, 1995”) might be replaced with just the year (“1995”). this process helps prevent re-identification, especially when combined with other quasi-identifiers. the extent of generalization depends on the dataset’s intended use. for datasets used in model training or longitudinal analysis, overly aggressive generalization might reduce usefulness. in contrast, applications with stricter privacy requirements may necessitate more coarse-grained representations. in our case, since the goal was to release a publicly available benchmark while preserving meaningful conversational structure and user traits, we opted for moderate generalization, preserving categories like profession or age range where appropriate. important? speaker utterance 0 user hey, how are you today? 0 bot i’m doing well, thanks! what’s up? 1 user i just found out i got into harvard! 0 bot wow, congratulations! that’s amazing! 0 user thanks, i’m still processing it all. . . . conversation continues . . . table 5: example conversation of the lufy-dataset 87 ryuichi sumida, inoue and kawahara q1 which university did {user} get accepted to? a1 harvard. q2 what is {user}’s favorite dog breed? a2 white schnauzer. q3 what is the name of {user}’s dog? a3 luke. table 6: example qa pairs of the lufy-dataset some studies, such as deid-gpt (liu et al., 2023), explore the use of llms for de-identification. other studies, such as the ego4d dataset (grauman et al., 2022), which consists of 3,000 hours of video, use commercial software like brighter.ai and secureredact. however, we manually executed the process to ensure proper de-identification. we aim for the lufy-dataset to serve as a benchmark for long-term conversations, given its unique length, as shown in table 1. the full dataset, including both the annotated human-chatbot conversations and the participant-created qa pairs, is publicly released to support future research in long-term conversational ai. 5.3 main results firstly, we evaluated the user experience using both subjective and automatic methods. we used three evaluation methods: (1) subjective human ratings (a subjective method) and (2) llm-based scoring (an automatic method), both based on the entire conversation, and (3) sentiment analysis (an automatic method), which was performed at the individual user utterance level. subjective evaluations for subjective evaluation, we asked third-party annotators to rate each conversation on three criteria: personalization, flow of conversation, and overall experience. to ensure that personalization judgments reflected the evolving relationship between each participant and chatbot, annotators evaluated the conversations in chronological order per participant, giving them access to earlier sessions when rating later ones. this ordering allowed raters to apply criteria consistently across multi-day interactions rather than evaluating each session in isolation. we provide the specific instructions for the three criteria, together with the intraclass correlation coefficient (icc) (shrout and fleiss, 1979), in appendix e. we opted for third-party annotators instead of collecting user self-ratings for several reasons. while user opinions are ultimately the most valuable signal, we were concerned that prompting users to rate each chatbot after every session might introduce bias—participants could infer which system was expected to perform better, potentially influencing their subsequent behavior. additionally, participants engaged in multiple long sessions (approximately five hours across four days); frequent questionnaires could contribute to fatigue and degrade the quality of feedback. to avoid these issues and ensure consistency across systems, we relied on independent annotators. the results, as presented in table 7, demonstrate lufy’s superior performance over both naive rag and memorybank across most sessions and evaluation criteria, with the most significant difference observed in session 4, where ratings for naive rag and memorybank dropped significantly, while lufy’s rating remained relatively stable. this trend underscores lufy’s enhanced ability to maintain engaging and personalized conversations, particularly as the interaction length increases. 88 enhancing long-term rag chatbots with psychological models personalization flow of conv. overall system s1 s2 s3 s4 avg. s1 s2 s3 s4 avg. s1 s2 s3 s4 avg. naive rag (3.94) 3.85 3.89 3.76 3.83 (3.86) 3.46 3.57 3.20 3.41 (3.82) 3.65 3.74 3.50 3.63 memorybank (3.87) 3.95 3.87 3.56 3.79 (3.69) 3.58 3.64 3.31 3.51 (3.85) 3.75 3.64 3.36 3.58 lufy (3.94) 4.13† 3.98 4.04† 4.05† (3.78) 3.63 3.57 3.72† 3.64† (3.80) 3.81 3.74 3.80† 3.78† table 7: subjective evaluation of three criteria: personalization, flow of conversation and overall, with ratings on a scale from 1 (lowest) to 5 (highest). † indicates statistically significant improvement over both other methods.(p < 0.05 , paired t-test). system (s1) s2 s3 s4 avg. naive rag (3.71) 3.53 3.38 3.18 3.36 memorybank (3.71) 3.53 3.44 3.21 3.39 lufy (3.71) 3.53 3.32 3.50† 3.45 table 8: average ratings (1–5 scale) by gpt-4o, averaged over 30 ratings. † indicates statistically significant improvement over both other methods.(p < 0.05 , paired t-test). ratings are shown with mean values. rating by llm as part of the automatic method to score the user experience, we assessed the user’s overall satisfaction using gpt-4o. we used gpt-4o as an evaluator because strong llms achieve performance comparable to human evaluators (zheng et al., 2024). the prompt given to the user experience is given in figure 6. as shown in table 8, the results provide valuable insights into the performance of the three rag chatbots. lufy consistently emerges as the top performer, achieving the highest average rating and demonstrating a significant improvement in the later stages of the conversation (session 4). this aligns with the subjective evaluations and is further supported by the sentiment analysis in the following section. sentiment analysis to complement both subjective and llm-based evaluations, we also performed sentiment analysis at the utterance level. we conducted sentiment analysis using finetuned versions of roberta models. we used distilroberta (sanh, 2019), fine-tuned on 4,840 polar sentiment sentences from english financial news, and timelms (loureiro et al., 2022), a roberta model trained on 124m tweets from january 2018 to december 2021 and fine-tuned with the tweeteval benchmark (mohammad et al., 2018). the hyperparameters for the fine-tuning of distilroberta are listed in table 9. as shown in table 10, the sentiment analysis results clearly indicate that lufy outperformed both naive rag and memorybank in fostering positive user experiences, especially during longer interactions. lufy consistently achieved higher average sentiment scores (0.50), demonstrating its ability to maintain a positive conversational tone, particularly evident in session 4 where its score (0.62) significantly surpassed those of naive rag and memorybank. this trend aligns with the subjective evaluations, highlighting its effectiveness in extended conversational settings. 89 ryuichi sumida, inoue and kawahara based on the following dialogue, evaluate and score how engaging the user({user_name})’s conversation is. consider the following factors: 1. the relevance of responses 2. ability to sustain a conversation 3. interest generated through their responses here is the conversation script:{conversation_script} the score should be on a scale from 1 to 5, where 1 is the lowest and 5 is the highest. be sure to provide both a score and a brief reason for your rating. figure 6: prompt for estimation of user experience. parameter value learning rate 2e-05 train batch size 8 eval batch size 8 seed 42 optimizer adam β1, β2 0.9, 0.999 epsilon 1e-08 lr scheduler type linear number of epochs 5 table 9: hyperparameters used for fine-tuning of distilroberta system s1 s2 s3 s4 avg. naive rag (0.58) 0.51 0.38 0.32 0.40 memorybank (0.58) 0.47 0.34 0.30 0.38 lufy (0.58) 0.43 0.45 0.62† 0.50† table 10: sentiment analysis results: average rating of user utterances (+1 = positive, 0 = neutral, -1 = negative). † indicates statistically significant improvement over both other methods. while sentiment analysis provides valuable insights into the emotional tone of user utterances, we acknowledge that it does not capture the full nuance of user experience. in particular, conversations involving argumentation or critical reflection may contain negative sentiment while still being engaging, constructive, and intellectually stimulating. therefore, we do not treat sentiment scores as a standalone measure of user satisfaction or engagement. instead, we use sentiment analysis as 90 enhancing long-term rag chatbots with psychological models precision recall f1 score system s1 s2 s3 s4 avg. s1 s2 s3 s4 avg. s1 s2 s3 s4 avg. naive rag 75.6 69.6 74.2 69.9 72.3 60.8 51.0 51.6 46.6 52.5 67.4 58.9 60.9 55.9 60.8 memorybank 63.6 50.7 57.4 64.6 59.1 41.2 29.4 30.7 31.9 33.3 50.0 37.2 40.0 42.7 42.5 lufy 80.0 86.9 86.8 86.6 85.1 54.9 51.0 38.6 35.3 44.9 65.1 64.3 53.4 50.2 58.3 table 11: comparison of precision, recall, and f1 score for different systems and sessions. each value reflects performance on the full set of questions encountered up to that session. for example, recall in s3 measures the model’s ability to correctly answer questions from sessions 1 through 3 at the end of session 3. a complementary signal to the subjective human evaluations and llm-based scoring, which are better suited to capture the broader context and qualitative aspects of the interaction. summary of results in summary, the subjective evaluations, llm-based ratings, and sentiment analysis provide a consistent picture: lufy excels in delivering sustained, personalized, and engaging user experiences, particularly in longer interactions. across all three evaluation methods, lufy achieved the highest scores, most notably in session 4, where both naive rag and memorybank exhibited significant performance drops. in contrast, lufy maintained stable and positive engagement. while naive rag and memorybank performed comparably to each other, their reliance on cosine similarity for memory retrieval and the limitations of memorybank’s forgetting mechanism appear to limit their ability to support coherent long-form conversations. these findings strongly suggest that lufy’s use of an importance-based forgetting mechanism, combined with importanceaware memory retrieval, is effective for improving chatbot quality in extended interactions. 5.4 in-depth analysis evaluations with qa pairs we assessed the models’ ability to remember using the qa pairs collected in the user experiment. after each session—and, for memorybank and lufy, also after their respective forgetting processes (naive rag does not include a forgetting process)— we asked each model the full set of questions accumulated up to that point and calculated precision, recall, and f1 score based on its responses. to ensure consistency and objectivity, responses were scored by gpt-4o, an llm. the exact prompt used for evaluation is detailed in appendix c.1. summary results are presented in table 11, while the full set of results is available in appendix f. naive rag achieved the highest recall overall, because of its lack of a forgetting mechanism. however, this came at the cost of lower precision compared to lufy, suggesting that it retrieved more irrelevant or outdated information. lufy outperformed both naive rag and memorybank in precision, and surpassed memorybank across all metrics in all sessions, indicating a more effective method of measuring importance than memorybank. memory matching rate to evaluate how closely system-selected memories align with human judgments, we measure the average pairwise agreement between each system and individual human annotators. for each session, we compare the binary importance labels produced by the model with 91 ryuichi sumida, inoue and kawahara important memories system s1 s2 s3 s4 avg. memorybank 19.4 8.8 14.3 15.3 14.4 lufy 19.4 16.3 24.4 10.1 17.6 (random) 13.1 10.0 10.6 11.1 11.2 (annotators) 18.3 26.1 23.5 24.2 23.0 table 12: system–human agreement on memory importance. precision recall f1 score system s1 s2 s3 s4 avg. s1 s2 s3 s4 avg. s1 s2 s3 s4 avg. lufy 80.0 86.9 86.8 86.6 85.1 54.9 51.0 38.6 35.3 45.0 65.1 64.3 53.4 50.2 58.3 −a 64.3 84.3 81.4 82.2 78.1 37.3 52.0 33.4 32.9 38.9 47.2 64.3 47.4 47.0 51.5 −p 79.0 84.3 77.3 80.2 80.2 45.1 45.1 32.1 29.4 38.0 57.4 58.8 45.4 43.0 51.7 −l 74.8 83.8 83.2 77.1 79.7 49.0 44.1 32.8 28.9 38.7 59.2 57.8 47.1 42.0 51.5 −r1 68.6 82.0 82.3 87.0 80.0 47.1 47.1 38.0 38.7 42.7 55.9 59.8 52.0 53.6 55.3 −r2 80.0 86.9 86.8 86.6 85.1 54.9 51.0 38.6 35.3 45.0 65.1 64.3 53.4 50.2 58.3 table 13: ablation study for precision, recall, and f1 score. −a indicates that wa is set to 0. the largest drop in performance is underlined. those assigned by each annotator, and compute the proportion of system-selected memories that were also marked as important by that annotator. as shown in table 12, across sessions, lufy outperforms or matches memorybank in three out of four sessions and achieves a higher overall average agreement, demonstrating more consistent alignment with human judgments. notably, lufy even exceeds the human–human pairwise agreement in two sessions, although its overall mean still falls slightly below the human average—reflecting the substantial subjectivity inherent in memory-importance annotation, where human pairwise agreement is only around 23%. we also computed the precision over utterances unanimously marked as important by all annotators. however, such unanimously important cases were extremely rare (typically only 2–6 per session), making the resulting precision estimates highly unstable and therefore unsuitable as a primary evaluation metric. for this reason, we focus on the more reliable pairwise agreement analysis. agreement on unimportant memories (i.e., cases where both the system and humans marked an utterance as unimportant) is much higher due to class imbalance—approximately 90% of utterances receive a “not important” label—but provides little discriminatory value between systems and is therefore omitted. ablation study we conducted an ablation study for the agreement probability with other annotators and precision, recall and f1 score. as shown in tabls 13 and 14, we found a (arousal), l 92 enhancing long-term rag chatbots with psychological models important memories system s1 s2 s3 s4 avg. lufy 20.9 12.4 11.8 14.4 14.9 −a 21.5 11.8 9.1 14.4 14.2 −p 23.5 11.8 10.4 13.7 14.9 −l 20.2 11.8 9.8 15.0 14.2 −r1 21.5 10.4 11.1 13.7 14.2 −r2 20.9 12.4 11.8 14.4 14.9 table 14: ablation study for the agreement probability with other annotators on whether memories are important. the largest drop in performance is underlined. (llm estimated importance), r1 (the number of times the memory is retrieved) to be of particular importance. episodic memory in conversations as stated in section 4.1.1, we stored three types of information: profile information, summaries of past sessions, and utterance pairs. we include summaries and utterance pairs in memory storage to retain small details about the user. these details, such as episodic memories (tulving, 2002), are unlikely to appear in the profile information. to assess the necessity of retaining such details, we measured the prevalence of episode-related utterances—user statements referring to unique, personally experienced events. we used an llm to identify such utterances. specifically, the model was given the prompt shown in figure 7: this classification process was applied across a sample of 170 conversations. our analysis revealed that 10.6% of the user’s utterances were episode-related. as shown in figure 8, only 36 of the 170 analyzed 30-minute conversations (21.1%) did not contain any episode-related utterances. in some conversations, however, up to 40% of the utterances were classified as episode-related. identify any parts of this conversation where the speaker refers to specific past experiences or events they personally remember. these should not include general knowledge, facts, or information from their profile, but rather unique, one-time occurrences or episodic memories. if no such references are found, output ‘0’. here is the conversation: \{conversation\} figure 7: prompt template for identifying whether a user utterance pertains to an episodic memory. 93 ryuichi sumida, inoue and kawahara figure 8: episode-related utterances frequency distribution in conversations (n = 170) 6. related work the evolution of conversational ai from stateless chatbots to modern llms has been defined by an increasing capacity for memory. while the transformer architecture’s context window provides a form of short-term memory, enabling conversational coherence, it is insufficient for the longterm, personalized interactions that characterize human relationships. to achieve this, agents require persistent, structured memory architectures capable of recalling specific interactions (episodic memory) and accessing a stable base of world knowledge (semantic memory). this review connects foundational theories from cognitive psychology with contemporary computational models, arguing that the future of conversational intelligence lies in the thoughtful integration of these cognitivelyinspired memory systems. our work is situated at the intersection of computational dialogue modeling and cognitive memory theory. foundational models in cognitive psychology distinguish between short-term and longterm memory (atkinson and shiffrin, 1968), and further between semantic memory (general knowledge) and episodic memory (event-specific personal experiences) (tulving et al., 1972). episodic memory, in particular, is crucial in dialogue, where references to past interactions and experiences often shape conversational flow and user engagement. our analysis of the lufy corpus found that approximately 79% of conversations included episode-related utterances, underscoring the importance of modeling episodic recall in conversational agents. traditional memory models in dialogue systems often rely on salience mechanisms such as recency and frequency (as in memorybank). while these approaches are computationally efficient, they fail to capture the richness of human remembering and forgetting, which involve not just decay and interference but also selective omission and social reasoning. in dialogue, speakers may choose to withhold or imply information based on context, face-saving (brockner et al., 1981)), or relationship management—processes not easily modeled with surface-level heuristics alone. 94 enhancing long-term rag chatbots with psychological models to address these challenges, recent work has explored structured and cognitively motivated memory systems. garcia contreras et al. (contreras et al., 2024) introduced a multi-store memory framework for an assistive robot, indy, inspired by cognitive psychology. their system is built around a three-tiered architecture—sensor memory, short-term memory (stm), and longterm memory (ltm)—and incorporates forgetting heuristics based on parametrized weibull decay curves (murthy et al., 2004). these mechanisms enable the robot to retain only contextually relevant or repeatedly reinforced information, simulating aspects of human episodic memory without unbounded memory growth. our current system retrieves past conversational content using a rag framework. while this supports some degree of contextual continuity, it does not yet implement explicit mechanisms for episodic segmentation or narrative abstraction as described in garcia contreras et al.’s work. their multi-store memory framework, built around sensor, short-term, and long-term memory, introduces cognitively inspired forgetting heuristics and structured narrative memory. recent work by (ong et al., 2025) also highlights the importance of structuring prior conversational data, proposing a memory timeline framework (theanine) that links past memories through temporal and causal relations to support response generation. while distinct from our goals, such efforts underscore the broader value of organizing conversational history beyond flat retrieval. in addition, we see promising avenues for future enhancement through structured relational representations. incorporating knowledge graphs or narrative schemas (wilcock and jokinen, 2022; walker et al., 2022) would enable the system to construct and retrieve memories not just as isolated utterances, but as semantically coherent entities and events—mirroring how humans organize and recall lived experience. as discussed above, while lufy already demonstrates the effectiveness of cognitively informed forgetting and memory prioritization, its design also opens up natural extensions toward structured and relational memory—marking a promising direction for future development. 7. conclusion 7.1 summary this study introduced lufy, a novel retrieval-augmented generation (rag) chatbot designed to address the challenge of memory management in long-term conversations. drawing on psychological insights, lufy employs a unique approach to evaluate and prioritize memories based on six key metrics: emotional arousal, surprise, llm-estimated importance, and retrieval-induced forgetting. unlike traditional rag models that either store all memories or rely on simplistic forgetting mechanisms, lufy uses learned weights to balance these metrics, ensuring that emotionally significant and contextually relevant memories are retained while less important ones are gradually forgotten. through extensive user experiments, involving conversations 4.5 times longer than those in previous studies, lufy demonstrated superior performance in maintaining engaging, personalized, and positive dialogues. the results, validated through subjective evaluations, sentiment analysis, and llm assessments, showed that lufy significantly outperformed both a naive rag system and the memorybank model, particularly in later stages of extended interactions. the development of the lufy-dataset, a comprehensive collection of human-chatbot conversations annotated for importance, further contributes to the field by providing a valuable resource for future research in long-term conversational ai. overall, this study demonstrates the potential of forgetting unimportant memories. 95 ryuichi sumida, inoue and kawahara 7.2 future work this study has demonstrated the potential advantages of the proposed method in enhancing chatbot interactions through improved objective and subjective evaluations. however, it is important to acknowledge its limitations. firstly, our investigation focused solely on the impact of memory-related psychological metrics in conversations between strangers. future research should aim to diversify the dataset by including interactions among friends, family members, and other relationships to comprehensively understand these metrics’ influence across different conversational contexts. second, this study only looked at text-based conversations. in text, indicators like exclamation marks play a role in detecting arousal. for example, when a user said, ”i’m going to hong kong with my friends next month,” our system recognized a low arousal level, missing the excitement. future work should explore multimodal settings to improve the accuracy of emotional detection. third, our current framework does not yet capture implicitly referenced content, social dynamics (e.g., face-saving (brockner et al., 1981)), or dialogue progression structure (levinson, 1981) (e.g., adjacency pairs, repairs). future work could investigate how such relational and pragmatic elements impact memory importance, potentially by incorporating discourse parsing, speaker state tracking, or theory-of-mind modeling. in addition, we plan to explore complementary metrics such as conversational risk-taking (e.g., controversy), and topical diversity to capture broader aspects of engagement beyond flow and personalization. fourth, in practice, long-term interactions may introduce contradictions—for example, when a user updates or retracts earlier statements. lufy does not currently attempt contradiction detection at the symbolic level, but it mitigates conflict through dynamic forgetting and salience-based retrieval: newer, more emotionally charged, or frequently retrieved memories are prioritized, while outdated or less important ones are forgotten. this approach is consistent with prior work in longterm memory for social agents (bae et al., 2022), which similarly favors overwriting older memories implicitly during retrieval. future work could integrate more explicit contradiction resolution using techniques from natural language inference or entailment. fifth, while we standardized the prompts used for tasks such as memory importance estimation and user experience evaluation, we recognize that prompt design can meaningfully influence llm behavior. for instance, in the rating prompt shown in figure 1, the numerical scale appears after the instructions. we did not investigate whether reordering or rephrasing such elements would alter outcomes. future work could examine prompt sensitivity more systematically—for example, testing whether presenting the rating scale before vs. after the task description impacts consistency or alignment with human judgments. this direction may be especially important for ensuring the reliability and interpretability of llm-based evaluations. ethical considerations. as our system involves the long-term retention of potentially sensitive user utterances—some of which are emotionally charged or self-disclosive—it raises important ethical questions for future deployment. in particular, ensuring user consent for what is stored, providing mechanisms for inspecting or deleting remembered content, and minimizing the risk of misrepresentation or inappropriate persistence of private information are essential. while our experiments were conducted with informed, consenting participants and involved de-identified data, practical interactive applications must prioritize transparency and user agency in memory handling. we envision future iterations of lufy supporting user-controllable memory (e.g., editable memory logs or consent-based retention) to better align with ethical standards for human–ai interaction. 96 enhancing long-term rag chatbots with psychological models references richard c atkinson and richard m shiffrin. human memory: a proposed system and its control processes. in psychology of learning and motivation, volume 2, pages 89–195. elsevier, 1968. sanghwan bae, donghyun kwak, soyoung kang, min young lee, sungdong kim, yuin jeong, hyeri kim, sang-woo lee, woomyoung park, and nako sung. keep me updated! memory management in long-term conversations. in findings of the association for computational linguistics: emnlp 2022, pages 3769–3787, 2022. anouck braggaar, frédéric tomas, peter blomsma, saar hommes, nadine braun, emiel van miltenburg, chris van der lee, martijn goudbeek, and emiel krahmer. a reproduction study of methods for evaluating dialogue system output: replicating santhanam and shaikh (2019). in proceedings of the 15th international conference on natural language generation: generation challenges, pages 86–93, 2022. vincent breton-provencher, gabrielle t drummond, jiesi feng, yulong li, and mriganka sur. spatiotemporal dynamics of noradrenaline during learned behaviour. nature, 606(7915):732– 738, 2022. url https://www.nature.com/articles/s41586-022-04782-2. joel brockner, jeffrey z rubin, and elaine lang. face-saving and entrapment. journal of experimental social psychology, 17(1):68–79, 1981. roger brown and james kulik. flashbulb memories. cognition, 5(1):73–99, 1977. sven buechel and udo hahn. emobank: studying the impact of annotation perspective and representation format on dimensional emotion analysis. in proceedings of the 15th conference of the european chapter of the association for computational linguistics: volume 2, short papers, pages 578–585, 2017. joana campos, james kennedy, and jill f. lehman. challenges in exploiting conversational memory in human-agent interaction. in proceedings of the 17th international conference on autonomous agents and multiagent systems, aamas ’18, page 1649–1657, richland, sc, 2018. international foundation for autonomous agents and multiagent systems. eunbi choi, kyoung-woon on, gunsoo han, sungwoong kim, daniel wontae nam, daejin jo, seung eun rho, taehwan kwon, and minjoon seo. effortless integration of memory management into open-domain conversation systems. arxiv preprint, 2305.13973, 2023. angel f. garcı́a contreras, seiya kawano, yasutomo kawanishi, yutaka nakamura, satoru saito, and koichiro yoshino. forgetful multi-store memory system for a cognitive assistive robot. in proceedings of the 30th annual meeting of the association for natural language processing (japan), pages 978–982, 2024. martin a conway, stephen j anderson, steen f larsen, carol m donnelly, mark a mcdaniel, allistair gr mcclelland, richard e rawles, and robert h logie. the formation of flashbulb memories. memory & cognition, 22:326–343, 1994. url https://link.springer.com/ article/10.3758/bf03200860. 97 ryuichi sumida, inoue and kawahara robin im dunbar, anna marriott, and neil dc duncan. human conversational behavior. human nature, 8:231–246, 1997. herm ebbinghaus. ueber das gedächtnis. mind, 10(39), 1885. norman s endler and james da parker. assessment of multidimensional coping: task, emotion, and avoidance strategies. psychological assessment, 6(1):50, 1994. ali fatihi. unspoken understandings: navigating the nuances of pragmatics. ethics international press, 2025. joseph l fleiss. measuring nominal scale agreement among many raters. psychological bulletin, 76(5):378, 1971. david fraile navarro, enrico coiera, thomas w hambly, zoe triplett, nahyan asif, anindya susanto, anamika chowdhury, amaya azcoaga lorenzo, mark dras, and shlomo berkovsky. expert evaluation of large language models for clinical dialogue summarization. scientific reports, 15(1):1195, 2025. annik imogen gmel, gerhard gmel, rudolf von niederhäusern, michael andreas weishaupt, and markus neuditschko. should we agree to disagree? an evaluation of the inter-rater reliability of gait quality traits in franches-montagnes stallions. journal of equine veterinary science, 88: 102932, 2020. kristen grauman, andrew westbury, eugene byrne, zachary chavis, antonino furnari, rohit girdhar, jackson hamburger, hao jiang, miao liu, xingyu liu, et al. ego4d: around the world in 3,000 hours of egocentric video. in proceedings of the ieee/cvf conference on computer vision and pattern recognition, pages 18995–19012, 2022. william hirst and gerald echterhoff. remembering in conversations: the social sharing and reshaping of memories. annual review of psychology, 63:55–79, 2012. url https://www. annualreviews.org/doi/abs/10.1146/annurev-psych-120710-100340. susan howell, suzanne beeke, emma louise sinnott, rosemary varley, and tim pring. an examination of sample length and reliability of the interactional network tool, a new measure of group interactions in acquired brain injury. aphasiology, 37(10):1646–1660, 2023. jihyoung jang, minseong boo, and hyounghun kim. conversation chronicles: towards diverse temporal and relational dynamics in multi-session conversations. in proceedings of the 2023 conference on empirical methods in natural language processing, pages 13584–13606, 2023. hyunwoo kim, jack hessel, liwei jiang, peter west, ximing lu, youngjae yu, pei zhou, ronan bras, malihe alikhani, gunhee kim, et al. soda: million-scale dialogue distillation with social commonsense contextualization. in proceedings of the 2023 conference on empirical methods in natural language processing, pages 12930–12949, 2023. stephen c levinson. some pre-observations on the modelling of dialogue. discourse processes, 4 (2):93–116, 1981. 98 enhancing long-term rag chatbots with psychological models yanran li, hui su, xiaoyu shen, wenjie li, ziqiang cao, and shuzi niu. dailydialog: a manually labelled multi-turn dialogue dataset. in proceedings of the eighth international joint conference on natural language processing (volume 1: long papers), pages 986–995, 2017. zhuowan li, cheng li, mingyang zhang, qiaozhu mei, and michael bendersky. retrieval augmented generation or long-context llms? a comprehensive study and hybrid approach. arxiv preprint arxiv:2407.16833, 2024. yinhan liu, myle ott, naman goyal, jingfei du, mandar joshi, danqi chen, omer levy, mike lewis, luke zettlemoyer, and veselin stoyanov. roberta: a robustly optimized bert pretraining approach. arxiv preprint arxiv:1907.11692, 2019. zhengliang liu, yue huang, xiaowei yu, lu zhang, zihao wu, chao cao, haixing dai, lin zhao, yiwei li, peng shu, et al. deid-gpt: zero-shot medical text de-identification by gpt-4. arxiv preprint arxiv:2303.11032, 2023. daniel loureiro, francesco barbieri, leonardo neves, luis espinosa anke, and jose camachocollados. timelms: diachronic language models from twitter. in proceedings of the 60th annual meeting of the association for computational linguistics: system demonstrations, pages 251– 260, dublin, ireland, may 2022. association for computational linguistics. doi: 10.18653/v1/ 2022.acl-demo.25. url https://aclanthology.org/2022.acl-demo.25. james l mcgaugh. memory and emotion: the making of lasting memories. columbia university press, 2003. saif mohammad, felipe bravo-marquez, mohammad salameh, and svetlana kiritchenko. semeval-2018 task 1: affect in tweets. in proceedings of the 12th international workshop on semantic evaluation, pages 1–17, 2018. dn prabhakar murthy, min xie, and renyan jiang. weibull models. john wiley & sons, 2004. kai tzu-iunn ong, namyoung kim, minju gwak, hyungjoo chae, taeyoon kwon, yohan jo, seung-won hwang, dongha lee, and jinyoung yeo. towards lifelong dialogue agents via timeline-based memory management. in proceedings of the 2025 conference of the nations of the americas chapter of the association for computational linguistics: human language technologies (volume 1: long papers), pages 8631–8661, 2025. joon sung park, joseph o’brien, carrie jun cai, meredith ringel morris, percy liang, and michael s bernstein. generative agents: interactive simulacra of human behavior. in proceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22, 2023. andrew reece, gus cooney, peter bull, christine chung, bryn dawson, casey fitzpatrick, tamara glazer, dean knox, alex liebscher, and sebastian marin. the candor corpus: insights from a large multimodal dataset of naturalistic conversation. science advances, 9(13):eadf3197, 2023. doi: 10.1126/sciadv.adf3197. daniel reisberg and paula hertel. memory and emotion. oxford university press, 2003. 99 ryuichi sumida, inoue and kawahara v sanh. distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. in proceedings of thirty-third conference on neural information processing systems (nips2019), 2019. patrick e shrout and joseph l fleiss. intraclass correlations: uses in assessing rater reliability. psychological bulletin, 86(2):420, 1979. laura stafford and john a daly. conversational memory: the effects of recall mode and memory expectancies on remembrances of natural conversations. human communication research, 10 (3):379–402, 1984. endel tulving. episodic memory: from mind to brain. annual review of psychology, 53(1):1–25, 2002. endel tulving et al. episodic and semantic memory. organization of memory, 1(381-403):1, 1972. nicholas thomas walker, torbjørn dahl, and pierre lison. dialogue management as graph transformations. in conversational ai for natural human-centric interaction: 12th international workshop on spoken dialogue system technology, iwsds 2021, singapore, pages 219–227. springer, 2022. joseph b walther and lisa c tidwell. nonverbal cues in computer-mediated communication, and the effect of chronemics on relational communication. journal of organizational computing and electronic commerce, 5(4):355–378, 1995. graham wilcock and kristiina jokinen. conversational ai and knowledge graphs for social robot interaction. in 2022 17th acm/ieee international conference on human-robot interaction (hri), pages 1090–1094. ieee, 2022. jing xu, arthur szlam, and jason weston. beyond goldfish memory: long-term open-domain conversation. in proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), pages 5180–5197, 2022. tan yu, anbang xu, and rama akkiraju. in defense of rag in the era of long-context language models. arxiv preprint arxiv:2409.01666, 2024. lianmin zheng, wei-lin chiang, ying sheng, siyuan zhuang, zhanghao wu, yonghao zhuang, zi lin, zhuohan li, dacheng li, eric xing, et al. judging llm-as-a-judge with mt-bench and chatbot arena. advances in neural information processing systems, 36, 2024. wanjun zhong, lianghong guo, qiqi gao, he ye, and yanlin wang. memorybank: enhancing large language models with long-term memory. in proceedings of the aaai conference on artificial intelligence, volume 38, pages 19724–19731, 2024. appendix a. top-k retrieval using the same 495 qa pairs from the lufy-dataset, we retrieved the top 10 memories for each question and analyzed the position of the correct memory among them. the results are shown in table 15. notably, we found that 85% of the correct memories appear within the top 5, with 100 enhancing long-term rag chatbots with psychological models top-k memory ratio (%) top 1 52.5 top 2 61.8 top 3 71.6 top 4 79.9 top 5 84.8 top 6 85.6 top 7 88.2 top 8 88.2 top 9 90.6 top 10 91.2 (not in top 10) 8.8 table 15: ratio of the correct memory being in various top-k memories. the correct memory is included in the top-5 memories 84.8% of the time. only marginal improvement when increasing k. this suggests that selecting the top 5 memories is effective for conversational rag chatbots, aligning with previous studies on rag systems in nonconversational settings (li et al., 2024). aligning our method with the concept of rif (hirst and echterhoff, 2012), we retrieved the top 2 memories despite the higher effectiveness observed with the top 5. appendix b. empirical estimation of contemporary event references in the lufy corpus to evaluate the reliability of parameter values grounded in literature estimates—particularly the claim that approximately 70% of conversation time concerns contemporary events—we conducted an empirical analysis on a random sample of the lufy dialogue corpus. specifically, we analyzed 2,092 randomly selected utterances. of these, 479 (22.90%) concerned contemporary events, while 1,613 (77.10%) did not, suggesting a considerably lower proportion of contemporary-event-focused dialogue than the literature estimate. b.1 method we used an llm to classify utterances according to whether they referenced contemporary or recent events. specifically, the model was prompted with user-bot dialogue pairs and asked to identify whether the user’s utterance referred to an event that is currently happening, just occurred, or is scheduled to occur soon. contemporary utterances included examples such as “i’m going to the beach tomorrow” or “we’re watching the game tonight.” non-contemporary utterances involved background information, personality traits, opinions, or general facts, such as “i love jazz” or “my sister is a doctor.” each utterance was independently evaluated using the prompt in figure 9. 101 ryuichi sumida, inoue and kawahara you will be given a pair of utterances from a conversation: one from a user and one from a chatbot. determine whether the user’s utterance refers to a contemporary or recent event|that is, something happening now, that just happened, or that is planned to happen soon (e.g., "i’m going to the beach tomorrow", "i just got a new job", "we’re watching the game tonight"). if the utterance is about general facts, opinions, personality traits, or background information (e.g., "i love jazz", "my sister is a doctor", "i’m a morning person"), it is not considered contemporary. here is the bot and user’s utterance: bot: \{bot\_utterance\} user: \{user\_utterance\} output 1 if the user’s utterance is about contemporary things, 0 if not. do not output anything else|just the number. figure 9: prompt template for identifying whether a user utterance pertains to a contemporary or recent event. b.2 results the analysis yielded the following statistics over a total of 2,092 user utterances from the lufy corpus: • total number of utterances processed: 2,092 • number of utterances about contemporary events: 479 • number of utterances not about contemporary events: 1,613 • percentage of utterances about contemporary events: 22.90% • percentage of utterances not about contemporary events: 77.10% b.3 discussion this empirical figure (22.9%) diverges considerably from earlier claims in the literature suggesting that up to 70% of conversational content pertains to contemporary events. while lower than the literature estimate, this proportion still represents a substantial share of user utterances, underscoring the relevance of recency in conversational settings. the discrepancy highlights the potential for 102 enhancing long-term rag chatbots with psychological models such estimates to be context-dependent or to overgeneralize across different dialogue scenarios. future work could strengthen these findings by examining the prevalence of contemporary references across multiple corpora or dialogue domains, thereby assessing their consistency and generalizability. appendix c. prompt given to llm c.1 prompt for judging whether the response is correct or not we used the following prompt to judge whether the provided answer by a system is correct or not. judge whether the answer to the following question is correct based on the provided label answer. here is the question: {question} here is the label answer {label_answer} here is the answer that you will be judging: {answer_to_judge} if the answer to judge is correct, output 1, if it’s incorrect, output 0. if the answer to judge is something like "i don’t know" or "no information found", output 2. output 1 even if the answer is only partly correct. output 0 if the answer uses "i" as the subject, because the question is about the user,not the assistant. only output 1 or 0 or 2, nothing else, no explanation needed. c.2 prompt for categorizing utterances as profile-related, episode-related, or other we used the following prompt to classify each user utterance into one of three categories: profilerelated, episode-related, or other. the model’s output was used for post-hoc analysis of utterance types. you will be given a single user utterance from a human-chatbot conversation. classify the utterance into one of the following three categories: profile-related: the user is stating biographical facts, preferences, personal attributes, or other stable information about themselves (e.g., "my favorite color is blue", "i have two brothers", "i live in new york"). episode-related: the user is referring to a specific event or experience, whether in the past, present, or future (e.g., "i went to a concert last night", "i’m going to hong kong next month", 103 ryuichi sumida, inoue and kawahara "i had a fight with my roommate"). other: the utterance does not fall into the above categories. this includes small talk, greetings, questions, or general opinions not tied to the user’s identity or experiences. please output only the category label: "profile-related", "episode-related", or "other". here is the utterance: "{user_utterance}" appendix d. cosine similarity threshold we used the default value of 0.8 for the cosine similarity threshold. using the lufy-dataset, we conducted experiments to verify whether this choice was optimal. we used the lufy dataset, which consists of 510 qa pairs. of these, only 2.9% (15 questions) were multi-hop questions requiring at least two pieces of information. therefore, we analyzed the remaining 495 single-hop questions—for example, “what’s my favorite movie?”, which can be answered by retrieving a single utterance (e.g., ”my favorite movie is interstellar”) from the conversation. for each pair, we retrieved ten relevant memories with a cosine similarity threshold of at least 0.6 and annotated whether the retrieved memory was correct for answering the question. the correctness of a retrieved answer was largely objective, as it involved verifying whether the retrieved memory semantically matched the ground truth answer. one annotator conducted this verification, and since it was an objective comparison, we did not compute inter-annotator agreement metrics such as cohen’s kappa. in the rare case of ambiguity, disagreements were resolved through discussion among the authors. figure 10 shows the raw counts of correct and incorrect memories across cosine similarity score bins, illustrating how both relevant and irrelevant memories are distributed across similarity ranges. our findings include: • memories with a cosine similarity of less than 0.75 were always irrelevant or incorrect. • there is a positive correlation between higher cosine similarity and the likelihood of the memory correctly answering the question, although fewer cases are retrieved when the cosine similarity is over 0.85. based on these findings, we tested various cosine similarity thresholds between 0.75 to 0.83 to determine the threshold that would yield the highest f1 score, representing the balance between retrieving relevant information and filtering out irrelevant content. as shown in figure 11, we observed that f1 score remains stable between thresholds of 0.75 and 0.8, with a notable drop beyond 0.8. while f1 score was used to characterize the overall tradeoff between precision and recall, we recognize that the goal of retrieval in this context—answering 104 enhancing long-term rag chatbots with psychological models figure 10: percentage of correct memories for different thresholds. thresholds below 0.75 are omitted because in our analysis, no correct memories were retrieved in that range. figure 11: f1 score, recall, and precision at various thresholds. minimal change in f1 score between 0.75 to 0.795, with a significant decline beyond 0.8. single-hop memory questions—may prioritize high-precision retrieval over exhaustive recall. indeed, in most cases, retrieving a single correct memory suffices for generating an accurate response. 105 ryuichi sumida, inoue and kawahara nevertheless, we selected 0.8 as a practical threshold because it retains a strong balance: it filters out many irrelevant memories (as seen with poor performance below 0.75) while still admitting a sufficient number of relevant ones to answer most questions. moreover, this value aligns with commonly used defaults in rag frameworks such as llamaindex, providing a reasonable and reproducible baseline. we acknowledge that for tasks with stricter precision requirements, a higher threshold might be preferable, and exploring task-specific threshold tuning remains a promising direction for future work. 106 enhancing long-term rag chatbots with psychological models appendix e. instructions given to the annotators for subjective evaluations we provide the specific instructions for the three criteria given to annotators. did the chatbot’s personalization appear appropriate in the conversation? 1. 1/5 (very poor): responses feel generic and lack personalization. 2. 2/5 (poor): limited personalization attempts are often off-target or superficial. responses only occasionally reflect the user’s context. 3. 3/5 (average): the chatbot shows moderate personalization. responses are relevant but may not fully address specific user needs. 4. 4/5 (good): the chatbot consistently personalizes well, matching responses to the user’s context and needs effectively. 5. 5/5 (excellent): outstanding personalization. responses are highly relevant, context-aware, and perfectly meet the user’s needs, enhancing engagement and satisfaction. how well did the conversation flow without feeling disjointed or out of context? • 1/5 (very poor): the conversation feels broken and illogical, with responses often out of context. • 2/5 (poor): the conversation has frequent awkward transitions or non-sequiturs that disrupt the flow. • 3/5 (average): the conversation flows reasonably well, with some disjointed moments that slightly distract from the overall experience. • 4/5 (good): the conversation flows well, with only minor issues that do not significantly impact the user’s experience. • 5/5 (excellent): the conversation flows seamlessly and logically, feeling completely natural and coherent throughout. overall, how would you rate the user’s experience with the chatbot? 1. 1/5 (very poor): the user found the conversation frustrating and unhelpful, strongly feeling they would not want to use the chatbot again. 2. 2/5 (poor): the user was somewhat disappointed with the conversation, finding little value in it and is unlikely to use the chatbot again soon. 3. 3/5 (average): the conversation met the user’s basic expectations. they would consider using the chatbot again if needed. 4. 4/5 (good): the user was pleased with the conversation and found it helpful, expressing a clear interest in using the chatbot again. 107 ryuichi sumida, inoue and kawahara 5. 5/5 (excellent): the user was highly satisfied with the conversation, finding it exceptionally useful and engaging, and is eager to use the chatbot again. to assess the consistency of human judgments across subjective criteria—personalization, flow of conversation, and overall quality—we computed the intraclass correlation coefficient (icc) across all annotators for each metric. the resulting icc values were: • personalization: 0.31 • flow of conversation: 0.29 • overall quality: 0.32 these values indicate fair to low inter-rater agreement. while relatively low, such results are consistent with previous findings in dialogue evaluation literature, where subjective judgments often yield low reliability due to inherent variability in interpretation and preference. for example, (braggaar et al., 2022) highlighted that likert-scale evaluations in dialogue systems frequently suffer from inconsistency, especially when raters lack a shared rubric or reference standard. similarly, icc values below 0.50 were reported in expert evaluations of conversational quality (gmel et al., 2020), and other dialogue-related tasks—such as behavior coding (howell et al., 2023) and clinical summarization (fraile navarro et al., 2025)—have shown similarly low reliability due to ambiguous criteria, narrow rating variance, and inconsistent calibration. appendix f. complete results of the evaluations with the qa pairs 108 enhancing long-term rag chatbots with psychological models s1 s2 s3 s4 system q1 q1 q2 q1 q2 q3 q1 q2 q3 q4 naive rag 75.6 68.2 71.0 64.3 61.3 97.1 62.5 62.1 91.7 63.3 memorybank 63.6 56.7 44.8 60.0 40.9 71.4 63.3 43.5 63.6 88.0 lufy 80.0 81.3 92.6 82.8 77.8 100.0 82.1 88.9 83.3 92.0 table 16: complete results showing the precision scores of each system across different sessions and question sets. q1 denotes questions about session1, q2 denotes questions about session2 and so on. s1 s2 s3 s4 system q1 q1 q2 q1 q2 q3 q1 q2 q3 q4 naive rag 60.8 58.8 43.1 52.9 37.3 64.7 49.0 35.3 64.7 37.3 memorybank 41.2 33.3 25.5 35.3 17.6 39.2 37.3 19.6 27.5 43.1 lufy 54.9 53.0 49.0 47.1 27.5 41.2 45.1 31.4 19.6 45.1 table 17: complete results showing the recall scores of each system across different sessions and question sets. s1 s2 s3 s4 system q1 q1 q2 q1 q2 q3 q1 q2 q3 q4 naive rag 67.4 63.2 53.6 58.0 46.4 77.7 54.9 45.0 75.9 46.9 memorybank 50.0 42.0 32.5 44.4 24.6 50.6 46.9 27.0 38.4 57.9 lufy 65.1 64.2 64.1 60.0 40.6 58.4 58.2 46.4 31.7 60.5 table 18: complete results showing the f1 scores of each system across different sessions and question sets. 109 ryuichi sumida, inoue and kawahara appendix g. illustrative example of memory retention and retrieval across systems to illustrate the characteristic behavior of the three systems, we present a constructed scenario highlighting how each handles memory retention and retrieval in response to a user’s emotionally significant utterance. the key memory—“i got into harvard”—was originally mentioned in session 2. system: naive rag memory status: retained (no forgetting). retrieval: fails to retrieve the ”harvard” memory due to low lexical overlap with the user’s input. bot response: “that’s understandable. college can be stressful. want to talk more about it?” system: memorybank memory status: forgotten. not retrieved in session 2, so decayed due to low frequency and recency. retrieval: memory deleted in forgetting phase. bot response: “yeah, college planning can be tough. do you know where you’re applying yet?” system: lufy (ours) memory status: retained due to high emotional arousal score. retrieval: importance-boosted retrieval surfaces the harvard memory despite low lexical similarity. bot response: “got into harvard—no wonder you’re feeling pressure. want to talk?” table 19: comparison of system responses in session 3 after the user says: “i’m feeling overwhelmed lately. so much pressure with upcoming college stuff.” only lufy successfully recalls the emotionally important but unrepeated “i got into harvard” memory, which was originally mentioned in session 2. 110 dialogue & discourse 15(2) (2024) 1–35 doi: https://doi.org/10.5210/dad.2024.201 perspective-taking and protagonist prominence: an empirical approach to the role of local and global prominence sara meuser meuser@uni-wuppertal.de bergische universität wuppertal maximilian hörl mhörl1@uni-koeln.de universität zu köln stefan hinterwimmer stefan.hinterwimmer@uni-hamburg.de universität hamburg editor: amir zeldes submitted 06/2023; accepted 06/2024; published online 08/2024 abstract the choice of the perspectival center of a stretch of discourse is crucial for the interpretation of certain phenomena such as free indirect discourse. it has been argued that the protagonist that is most prominent compared to competing protagonists gets to be the perspectival center. in this paper we discuss grammatical function and referential expression as prominence-lending cues and their impact on perspective-taking. we take the anchoring of free indirect discourse as the indicator for a shift in perspective as free indirect discourse can only be processed correctly if the reader is able to ascribe the utterance or thought to a protagonist. identifying the perspectival center is particularly crucial for the interpretation of a thought or utterance in free indirect discourse mode that can potentially be ascribed to different protagonists, since in contrast to direct or indirect discourse the respective speaker or thinker is not explicitly marked as such in free indirect discourse. in a series of acceptability rating studies, we tested if anchoring of free indirect discourse to the less prominent of two competing referents is perceived to be unnatural in german. further, we take a closer look at the role of subject and object as well as the choice of referential expression (proper name compared to indefinite noun phrase). we find that a protagonist referred to with a proper name in subject position is highly preferred as the anchor for free indirect discourse compared to a protagonist referred to with an indefinite noun phrase in object position. building on these findings, we present evidence that the prominence of the referent that is established in the sentence preceding a sentence in free indirect discourse mode can be overridden by discourse prominence. that is, a referent that is repeatedly mentioned in a short discourse is preferred as the perspectival center regardless of the prominence of a competing referent in the sentence preceding a sentence in free indirect discourse mode keywords: perspective, free indirect discourse, prominence, discourse, subjecthood 1. introduction 1.1 perspective-taking in language the choice of words in a spoken discourse or a narrative does not only depend on the intention and the knowledge that a speaker or narrator has but also on their perspective regarding location in space and time, their relation to other individuals etc. in particular expressions referring to places, ©2024 sara meuser, maximilian hörl, stefan hinterwimmer this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). meuser, hörl and hinterwimmer points in time or individuals, but also adverbs like luckily or sadly express an involvement in a given context from a certain point of view – for example, my mother, my daughter or the minister of defense may refer to the same person depending on the relation between the person uttering the expression and the person denoted by the expression. a coherent interpretation of a discourse thus requires the hearer or reader to ascribe context-sensitive utterances to the speaker or narrator or a certain protagonist. a change in perspective in a narration as well as in spoken discourse can simply be indicated by means of direct or indirect speech or thought representation; e.g., the minister of defense can be denoted by the expression my mother if the speaker signals that he or she is citing a child of the minister of defense – or if the speaker is a child of the minister of defense. in fiction a certain perspective can also be indicated by narrative mode. a first-person narrator for example usually reports or recounts events from his or her fixed point of view. a narration from a third-person perspective, in contrast, may allow for sporadic shifts in perspective for example in free indirect discourse (fid) or other kinds of perspective-taking such as viewpoint shifting (hinterwimmer, 2017) or protagonist projection (holton, 1997; stokke, 2013; abrusán, 2020). fid is a form of thought representation where context-sensitive expressions such as deictic expressions or evaluative terms are interpreted with respect to the perspective of a certain protagonist, while at the same time the narrative mode of the third-person narrator remains unchanged, as shown in (1a): the second sentence in (1a) is intuitively interpreted as a thought of maria, with the deictic adverb tomorrow being interpreted with respect to maria’s context, i.e. as referring to the day following the day on which maria has the thought. at the same time, maria is referred to by a thirdrather than a first-person pronoun, and the present rather than the past tense form of the auxiliary will is used. view point shifting and protagonist projection, in contrast, do not render conscious thoughts of protagonists but rather render the contents of their perceptions in a way that is compatible with their belief state at the time of perceiving, as shown in (1b): the main clause in (1b), for example, on its most plausible interpretation does not describe an event that is happening in the story world, but rather a temporary illusion that mary has a consequence of her sense of balance having been disturbed by the boat trip. (1) a. mary smiled. tomorrow she would reveal her true identity at the press conference. (hinterwimmer, 2019: 79, ex. (1c)) b. when mary stepped out of the boat, the ground was shaking beneath her feet for a couple of seconds. (hinterwimmer, 2017: 291, ex. (14)) when we approach perspective-taking in language, we have to consider that the common understanding of the term perspective refers to a visual experience that is limited according to the observer’s position. the term perspective is metaphorically applied to narration and may be defined as the linguistic and extralinguistic choices that are limited with respect to a certain point of view (see friedman, 1955; stanzel, 1984). a narration from a certain protagonist’s perspective can thus only convey information this particular protagonist has according to his or her position in a given narrative. in spoken language, the speaker is usually the perspectival center of the utterance – unless he or she indicates a change in perspective, e.g., by quoting another person. written discourse does not allow for such a straightforward ascription of the perspectival center as it is not always as trivial to pinpoint who is telling the story. in narration, it is not the author who writes down his or her story, shares memories or perceives a situation, but a narrating instance that is created by an author and that may be far more abstract than an actual individual and have many properties (such 2 perspective-taking and protagonist prominence as omniscience, for instance) that no actual individual has (see, e.g. zeman, 2020 for discussion).1 in the tradition of literary studies we will therefore never refer to the author’s perspective but to the narrator’s perspective. in this paper, we will give a short overview of perspective-taking in narratives before we discuss the conditions under which perspective shifts from neutral narration to fid are possible. we will follow the argumentation outlined in hinterwimmer (2019) (also hinterwimmer and meuser, 2019; meuser, 2022) and show that the availability of a protagonist as a possible perspective holder depends on his or her prominence status (see wiebe, 1990, 1994 for an early discussion of this issue in a computational framework and abrusán, 2021 for a sketch of a unified proposal that integrates the insights of wiebe, 1990, 1994 and hinterwimmer, 2019). in section 2 we report acceptability rating studies (see also meuser, 2022) investigating the hypothesis that a prominent referent is more likely to be the perspectival center of a text segment compared to a competing, less prominent referent. in our experiments we asked participants to rate the naturalness of german text segments in which sentences in fid must be interpreted with regard to two discourse referents and found that fid anchored to the more prominent protagonists received higher ratings.2 the first experiment (n=75) presented in section 2.1 will investigate the following hypothesis: h1: a protagonist that is more prominent in terms of referential expression and grammatical function is more easily available as the anchor for fid than a competing referent. in section 2.2 we will present a follow-up experiment (n=119) that will provide a first attempt to disentangle the prominence-lending cues grammatical function and referential expression by testing the following hypotheses: h2: protagonists functioning as subjects are more available as anchors for fid than protagonists functioning as objects. h3: protagonists referred to with a proper name are more available as anchors for fid than protagonists introduced by an indefinite np. furthermore, we will compare the effect investigated with respect to h2 and h3 so that we can draw conclusions regarding the hierarchies of the two prominence-lending cues. for that purpose, we raised the rather exploratory question (q1): does the type of referential expression play a bigger role than the grammatical function with respect to the anchoring of fid? a second question (q2) that calls for an exploratory investigation is whether the competition of the two protagonists has an impact on the anchoring preference, i.e., is the effect stronger when the competing referent is minimally prominent in terms of referential expression? we conclude that although fid, in contrast to direct and indirect discourse, does not require formal marking, the correct ascription of fid depends on the availability of a protagonist as the anchor for the thought or utterance. identifying the perspectival center is of particular importance whenever several protagonists are potentially available to be the anchor for an utterance. our findings 1. this approach is mostly limited to fictional texts. imagine a note on the refrigerator saying “i ate the cake”. here it is to be assumed that the author of the note ate the cake. an interpretation in which a narrator who is not the author ate the cake would be absurd. 2. although the experimental evidence is based on german stimuli rated by german native speakers, we want to argue that fid follows the same patterns in english. we decided to present english examples in the theoretical part of this paper to illustrate the issue. all examples were carefully checked to evoke the same effect by english native speakers. 3 meuser, hörl and hinterwimmer suggest that the anchoring of fid to a protagonist depends on his or her prominence status and that grammatical function is a crucial indicator for the availability as a perspectival center while the referential expression is less important as a prominence-lending cue. based on these findings we conducted a third acceptability rating study (n=116) where we test the effects found in the first two experiments in larger discourse. in particular, we wanted to test if a referent that is established as the most prominent in terms of grammatical function in a local context, i.e., the sentence preceding the fid, is still preferred as the perspectival center when a competing referent is more prominent in the global context, i.e. throughout a short discourse. we investigate two hypotheses: h4: protagonists that are highly activated in a discourse in terms of being repeatedly mentioned in subject position are just as available as anchors for fid as protagonists that are less often referred to in a discourse but in subject position in the sentence preceding the fid. h5: protagonists that are highly activated in a discourse in terms of being established as the discourse topic in a topic-establishing sentence, repeatedly mentioned in subject position, and mentioned in a title are more available as anchors for fid than protagonists that are less often referred to in a discourse but in subject position in the sentence preceding the fid. 1.2 perspective-taking in narratives we want to approach the term perspective and its connection to fiction by taking a look at how narrative modes set a certain perspective and allow for changes in perspective. the most obvious classification of narrating instances holds between the first-person narrator who tells a story from his or her point of view and a third-person narrator. a second relevant distinction is the one between homodiegetic and heterodiegetic narrators: homodiegetic narrators are part of the story world, while heterodiegetic narrators are not. consequently, the perspective of a homodiegetic first-person narrator is much more restricted than the perspective of a heterodiegetic first-person narrator: since the former is by definition also a protagonist of the story, they perceive or tell a story they witnessed from a fixed, i.e. their own perspective, and only have access to their own thoughts and feelings. by contrast, a heterodiegetic narrator has more or less knowledge about the ongoing story, the past and the future of the story world as well as insights into different protagonists’ thoughts with varying degrees of omniscience. this holds not only of third-person, but also of first-person heterodiegetic narrators, which are, however, much rarer than third-person heterodiegetic narrators, at least in modern fiction (see saure et al. 2023 for discussion of the differences between homoand heterodiegetic as well as first and third-person narrators with respect to perspective-taking). furthermore, literary scholars have done more fine-grained classifications such as the distinction between first-person, authorial and figural narrators (stanzel, 1984), between editorial, neutral, selective and multiple selective omniscience (friedman, 1955) or between zero focalization, external focalization and internal focalization (genette, 2010). the three latter terms, which have been very influential in literary studies of narration, refer to modes of narration where the narrator either has access to the thoughts, feelings and perceptions of all protagonists (zero focalization) or just one particular, highly prominent protagonist (internal focalization), or where the narrator does not have access to the thoughts, feelings and perceptions of any protagonist and just describes observable 4 perspective-taking and protagonist prominence events, actions and situations. in spite of using different terminology, all mentioned approaches describe narrators that may allow for shifts in perspective, i.e., they are not restricted to one perspective but they may temporally report a protagonist’s point of view without indicating the shift in perspective by the use of direct or indirect discourse. with many terms for perspective-taking in narratives we deal with a phenomenon that is well described in literary studies, yet with little attention3 drawn to the restrictions of perspective shift (with the notable exception of wiebe, 1990, 1994; see abrusán, 2021 for discussion) or, as formulated in hinterwimmer (2019), the answer to the question: under what conditions do protagonists become available as perspectival centers? with these descriptive approaches in mind, we want to point out that narration with a high degree of focalization does not only require the reader to identify the perspective holder in a given discourse and to be able to follow a shift in perspective from a narrating instance to a protagonist´s point of view, and vice versa, but also to be able to “integrate confronting perspectives” (zeman, 2019). while actual physical perspective-taking usually4 limits the observer to only one location from where a view may be directed, the concept of perspective in language we investigate naturally allows for multiple points of view. we will briefly elaborate on this idea taking an example from zeman that illustrates how the use of propositional attitude verbs commonly licenses (at least) dual perspectivization. (2) little red riding hood believes that the wolf is her grandmother. in (2) the reader is confronted with a proposition that requires the accommodation of an external and an internal reading. here, the information that little red riding hood believes the creature in her grandmother´s bed to be her grandmother has to be updated with the external viewpoint that provides the information that it is in fact the wolf who took the grandmother’s position in bed (i.e., the definite description the wolf is interpreted de re). in other words, without the integration of multiple perspectives we are limited to either little red riding hood´s point of view (as paraphrased in (3a) or the narrator´s point of view (as paraphrased in (3b)). (3) a. little red riding hood believes she sees her grandmother. b. the wolf pretends to be little red riding hood´s grandmother. we propose a similar multiple perspective reading including the external perspective of the narrator and the internal perspective of the protagonist for the examples of fid we discuss as indicators of a perspective shift. also, traditionally, in narratology it has been recognized that instances of selective omniscience or internal focalization are rather two overlapping or mixed perspectives – the one of the narrator 3. genette considers alterations in focalization as isolated violations that may occur as long as the coherence of the whole is given (genette, 2010). however, he does not elaborate to what extent focalization may alter in order for the text to be coherent and if all protagonists are possible perspective holders. 4. mirrors and cameras may allow for multiple perspectives. 5 meuser, hörl and hinterwimmer and the one of the protagonist (see pascal 1977 for the “dual voice” discussion). the terminology shift of perspective used in our discussion may therefore falsely indicate that one perspective is neglected for the sake of another – that is not the case. for reasons of simplicity, however, we will continue to talk about shifts in perspective – keep in mind that it is rather a shift in narrative mode from authorial, omniscient or zero focalized narration (henceforth: neutral perspective) to a figural, multiple omniscient or internally focalized narration (henceforth: protagonist´s perspective). 1.3 free indirect discourse although the issue of perspective-taking is well studied in linguistics (for an overview see eckardt, 2014) one question remains almost untackled: what restricts the choice of the perspectival center in a given discourse when different protagonists are available as perspectival centers, i.e., are shifts to the perspective of a certain protagonist that is part of a given discourse always coherent or do they sometimes lead to incoherence? this question is particularly interesting with regard to the limitations of fid. unlike direct or indirect discourse fid renders the thought of a protagonist without using quotation marks or embedding it under a propositional attitude verb such as say or think. a sentence in fid can thus only be interpreted correctly if the reader is able to take on the perspective of the protagonist whose thoughts or utterances are reported. characteristics of fid are, for example, interjections, judgmental statements, exclamatives, discourse particles, rhetorical questions and a partial shift in deixis (steube, 1985; banfield, 1982). (4) a. thomas looked at the calendar. he thought: “tomorrow i will finally see mommy again.” b. thomas looked at the calendar. he thought that he was going to see his mother again the next day. c. thomas looked at the calendar. tomorrow he would finally see his mommy again.5 especially the use of deictic expressions, i.e., expressions that depend on a fixed context such as i, here and now, is notable. in (4a) and (4b) we find a shift from the first-person personal pronoun (henceforth: ppro) i to the third-person ppro he, the informal expression mommy changes to the possessive determiner phrase his mother and the adverbial tomorrow shifts to the anaphoric adverbial the next day, which refers to the day following the day including the reference time set in the first sentence, i.e., the time when thomas looked at the calendar. in (4c), on the other hand, the use of deictic expressions seems inconsistent. here the third-person ppro he as well as the possessive determiner his indicate a neutral perspective. the temporal adverbial tomorrow contradicts the use of past tense in the preceding sentence and thus has to be interpreted with regard to thomas’ perspective just as the informal address mommy, which likewise expresses thomas’ perspective. in order to account for this partial shift, eckardt (2014) (see also schlenker, 2004 and sharvit, 2008 for closely related proposals built on the same basic idea) proposes to take two contexts into consideration: the narrator’s context c and the protagonist’s context c, where c is only introduced in 5. we want to add that this example could be read as a case of objective future-in-the past (see eckardt, 2017 for the terminology) if we paraphrase the utterance in the following way: unbeknownst to him, he would finally see mommy again tomorrow. while this reading is coherent, it is rather odd that the narrator refers to thomas’ mother as mommy. we want to clarify at this point that many utterances in fid mode may receive an objective reading by adding “but he/she was not aware of it”. we designed all items cautiously to make a future-in-the past reading strikingly odd. 6 perspective-taking and protagonist prominence cases of fid. while the spatial and temporal setting of a narrative usually depends on c, all context sensitive expressions with the exception of pronouns and tenses have to be interpreted with respect to c whenever c is available. pronouns and tenses, in contrast, are always interpreted with respect to c. crucially, whenever c is introduced, the proposition denoted by a sentence interpreted with respect to c as well as c is not interpreted as true with respect to the worlds compatible with the narration (the story worlds), but only as true in the worlds compatible with the beliefs of the respective protagonist (eckardt, 2014). example (5) can thus be paraphrased as: (5) there is an event e of thomas looking at the calendar that is located in the past with respect to the time of c (= the narration time) and in all worlds that are compatible with the beliefs of the author of c (= thomas) at the time of c (= the time of e) there is an event e’ of thomas seeing his mother again that is located in the past with respect to the time of c (= the narration time) and in the future with respect to the time of c and that takes place on the day that follows the day including the time of c (= the time of e). maier (2017), on the other hand, postulates that fid is a special form of mixed quotation in which everything except pronouns and tenses are quoted – i.e. the latter are systematically unquoted – , as shown in (6) for our example in (4) c6: (6) thomas looked at the calendar. “tomorrow” he would “finally see” his “mommy again”. we will leave open the question of which of the two lines of analysis just sketched is the correct one, as it is not directly relevant for our purposes in this paper. our concern is not so much how the perspective of the narrator interacts with the perspective of a single protagonist, but rather what the conditions are under which protagonists become available as perspectival centers – a question which becomes particularly pressing as soon as examples are considered in which several protagonists interact. 1.4 prominence and perspective while the interpretation of perspective-sensitive constructions such as fid as a thought of a protagonist has been studied extensively in linguistics and narratology, a common concern is whether it is justified to completely neglect a narrator-oriented interpretation in such cases and ascribe the propositions to the protagonist, i.e. whether fid represents just the protagonist’s, or both the narrator’s and the protagonist’s perspective. this issue was first tackled empirically by harris and potts (2009), who investigated if epithets and appositives can receive a non-speaker/non-narrator interpretation even in contexts where an overt first-person narrator is competing with a protagonist as perspectival center. in a series of forced-choice studies they found evidence that epithets and appositives can be interpreted from the perspective of a contextually prominent protagonist despite the overt narrator, and in the presence of certain contextual clues even favor such an interpretation. in line with the research of harris and potts, kaiser (2015) confirms that epithets and appositives can receive non-speaker/non-narrator interpretations especially when they are interpreted as fid. kaiser also tested if fid cues – evaluative epithets and adverbs of possibility – trigger the perspective of the grammatical subject of the preceding sentence rather than a speakeror narrator-oriented perspective. she found that in a sequence of two sentences (mary looked woefully at elizabeth. [poor 6. note that since it is not possible to unquote the verb without the tense, the whole finite verb is unquoted in (6). 7 meuser, hörl and hinterwimmer girl;] she was sick.) the presence of fid cues increases the likelihood that the personal pronoun in the second sentence refers to the object of the first sentence – despite the well-known preference of personal pronouns for the subjects of preceding sentences (see, e.g., crawley and stevenson, 1990; gernsbacher, 1990; gordon et al., 1993; stevenson et al., 1994 for discussion). the most plausible interpretation of this result is that the second sentence is interpreting as expressing mary’s rather that the narrator’s perspective. kaiser’s approach gives valuable insights on the anchoring of fid. however, her research focusses on the distinction between speaker/narrator and protagonist perspective without taking into account the possibility that multiple protagonists might be available as perspectival centers. before we get to our experimental approach to the limitations of perspective shifts, we want to elaborate on the anchoring of fid in contexts where two protagonists are in principle available as perspectival center. the discussion is based on our intuitions and the argumentation proposed in hinterwimmer (2019) regarding different prominence-based accounts as well as a coherence-based account. let us consider a discourse free of any contextual presuppositions or emotional involvement beyond the two sentences presented in (7) that could act in favor of one perspective or the other. free of any background assumptions regarding the two referents tina and mike and without any other contextual information we intuitively find differences in the availability of certain protagonists as anchors for fid: (7) a. tina was yelling at mike. tonight, that jerk really pushed it too far. b. tina was yelling at mike. tonight, that bitch really pushed it too far. while the second sentence in (7a) is most plausibly understood as a thought of tina about mike, the epithet in (7b) can only refer to a female protagonist and thus the fid in (7b) can only be interpreted as a thought of mike about tina. even though the context provided by the first sentence in which mike is being yelled at by tina gives perfect reason for mike to have negative thoughts about tina the second sentence conveys the impression that it is rather a comment by the narrator than fid anchored to mike. another, presumably rather far-fetched, reading of (7b) would suggest that the second sentence is a thought by tina and the female referent that bitch refers to some unmentioned female person. regardless of which option is chosen, the second example is harder to interpret than the first. the uncertainty about the ascription of the second sentence in (7b) illustrates the issue: linguistic cues in the sentences preceding fid increase or decrease the availability of referents as perspectival centers and thus anchors for fid. hinterwimmer (2019) suggests that it is the most prominent protagonist that is by default the perspectival center. let us narrow down the linguistic notion of prominence with respect to the availability of referents as the perspectival center of a sentence or short text segment. in example (7) both tina and mike are referred to with a rather common proper name and we assume that the reader does not have any reason to favor neither tina nor mike. still, tina can be regarded as more prominent than mike on at least two different prominence scales: first, while tina is the subject of the first sentence, mike is the prepositional object. according to the hierarchy of grammatical functions a subject is more prominent than a direct object, a direct object is more prominent than an indirect object and so forth (see himmelmann and primus, 2015 for an overview 8 perspective-taking and protagonist prominence of prominence hierarchies and von heusinger and schumacher, 2019 on prominence in discourse). second, regarding the hierarchy of semantic roles, tina is the agent, which is more prominent than the patient, mike (ibid.). although the fid in (7b) can be regarded as a possible thought of mike, who is mad at tina for yelling at him, mike is less easily easily available as perspectival center than tina in (7a) and thus a thought anchored to mike seems at least slightly odd and unexpected.7 having mentioned coherence as one restricting factor for perspective shifts (see kehler et al., 2008 and kehler and rohde, 2019 for an analysis of pronoun resolution in which coherence plays a crucial role)8, we want to point out that coherence alone does not account for the ascription of fid (but see abrusán, 2021 for observations showing that coherence relations can have an influence on protagonists’ availability as perspectival centers). in example (7a), the second sentence can be linked to the first by providing an explanation – the fact that mike had done or said something that had pushed it too far is the reason for tina to yell at him. in contrast, (7b) leaves room for different more or less coherent interpretations. one reading suggests that tina’s yelling is the reason for mike to have the thought rendered in fid. in that sense the first sentence may provide the cause for the second one. another interpretation may be that tina has done something else, unmentioned, that night that had pushed it too far. the sequence of the two sentences in (7b) could then simply be understood to be a linear order of narrative elements. a third interpretation in which the sentence in fid expresses a thought of tina and the epithet refers to someone else provides an explanation for the yelling in the first place similar to (7a), yet this reading seems rather absurd without any context. still, there are at least two ways available to link the proposition denoted by the sentence in fid coherently to the proposition denoted by the preceding sentence, the first one according to which tina’s yelling at mike causes mike to have the thought rendered in fid being no less plausible than the one linking the two sentences in (7a) via providing an explanation. consequently, the contrast between (7a) and (7b) cannot be accounted for in terms of coherence exclusively. rather, the availability of a referent as the perspectival center crucially depends on its prominence status (cf. (piwek and krahmer, 2000) on the interaction of salience, on the one hand, and plausibility/coherence, on the other, as factors contributing to the resolution of anaphora and presuppositions)9. while it seems tempting to draw conclusions on the anchoring of fid from pronoun resolution, hinterwimmer (2019) points out that pronoun resolution, although it may indicate tendencies for the availability of perspectival centers, works differently from fidanchoring. we do not want to elaborate on the similarities and differences between pronoun resolution and fid-anchoring at this point. for the investigation presented in this paper, however, it is important to note that preferences with respect to pronoun resolution do not always coincide with preferences regarding the anchoring of fid, i.e., the referent that may be preferred as the antecedent of a potentially ambiguous pronoun is not necessarily preferred as the perspectival center as illustrated in (8). (8) the new girli delighted janej. a. shei was just so gorgeous! (pronoun: i, perspective: j) 7. we have consulted two native english speakers, who shared our intuitions, but the contrast is admittedly quite subtle and there might be individual differences between speakers. 8. we can only take linguistic approaches to coherence into account as there is no elaboration on coherence as it is used in literary studies. 9. we are grateful to an anonymous reviewer for pointing us to this work. 9 meuser, hörl and hinterwimmer b. at least, shei hadn’t put on her best dress for nothing. (pronoun: i, perspective: i) c. but how could shej start a conversation? (pronoun: j, perspective: j) in line with the classification by garvey and caramazza (1974) who classify the verb to delight as yielding an np1-bias with respect to the resolution of personal pronouns it seems to be more expected to pick up the new girl rather than jane with the personal pronoun she. at the same time, there seems to be a clear preference to pick up the perspective of the second referent, i.e., jane, in (8a) as compared to (8b). additionally, despite the general preference for the new girl to be picked up with a pronoun in a subsequent sentence, an utterance in fid mode anchored to jane as in (8c) with the pronoun picking up jane seems to be more acceptable than (8b), where the utterance has to be anchored to the new girl and the pronoun is resolved to the new girl as well. with this in mind, let us now continue our discussion of prominence and perspective. the importance of prominence becomes more evident the more unevenly the prominence-lending features are distributed. in the following example the female protagonist is not only referred to with a proper name whereas the second referent is introduced with an indefinite article but she is also picked up by a pronoun. (9) a. when lisa was playing in the schoolyard, a boy pushed her into the stinging nettles. ouch, that itched! b. when lisa was playing in the schoolyard, she pushed a boy into the stinging nettles. ouch, that itched! here not only the qualitative measures such as type of referring expression but also quantitative measures such as number of mentions add to the availability of the referents as perspectival centers. in example (9b) an interpretation in which the boy´s thought is reported is coherent. however, according to the authors´ intuitions it is slightly awkward and unexpected to take on the boy´s perspective, although by no means impossible. in an attempt to anchor the fid this example offers another possible reading in which lisa takes on the boy´s perspective and empathically assumes that it must itch for the boy. this recursive reading again shows that it is the most prominent referent that is by default the perspectival center. intuitively in example (9) lisa seems to be more prominent than the boy resulting in a strong preference for (9a) as opposed to (9b). this intuition may best be captured by the notion of topicality. following reinhart (1981), we may regard lisa as the aboutness topic of the first sentence. if we assume that every sentence is the answer to an implicit question (van kuppevelt, 1995; roberts, 1996; ginzburg, 2012), example (9) can plausibly be interpreted as the answer to the question what did lisa do/experience in the schoolyard?, i.e., the answer to a question about lisa. a context in which example (8) is to be understood to answer a question about some boy, i.e., a question such as what did a boy do/experience in the schoolyard?, on the other hand is rather absurd. this may be explained by the choice of referring expression. while lisa is introduced by a proper name, the boy is introduced with an indefinite description, which indicates that the boy is unknown or at least not explicitly identifiable amongst other boys on the schoolyard. an implicit question about the boy is thus rather implausible as a discourse move. while the assumption that lisa is the topic in (9) goes well together with our claim that lisa is 10 perspective-taking and protagonist prominence the most prominent protagonist and consequently the only available perspectival center, we will not rely on the notion of topicality as it does not lend itself directly to a more fine-grained investigation of linguistic cues that impact the availability as the perspectival center. for our purpose the notion of prominence seems more appropriate as it comprises a wide range of linguistic cues and suggests hierarchies within those cues. furthermore, the notion of topicality on the level of the sentence it may not serve well with respect to larger discourse. this becomes more obvious when we make some syntactic changes to example (9) without affecting the content of our story. (10) lisa was playing in the schoolyard. a boy pushed her into the stinging nettles. ouch, that itched! in (10) we still prefer lisa as the perspectival center, when it is actually the boy who is the subject and the agent of the sentence preceding the fid while lisa is the object and the patient. while the concept of discourse topicality may serve as an indicator for perspectivehood, the notion of discourse topicality lacks a clear definition with respect to individual linguistic markers that are crucial for an empirical investigation (but see van dijk, 1977 and roberts, 2012 for relevant discussion). here we want to suggest that it is global prominence that favors lisa as the perspectival center. lisa´s prominence features clearly outweigh the features of the boy: she is the subject and agent of the first sentence, first mentioned, introduced with a proper name, picked up with a personal pronoun – consequently mentioned twice –, while the boy is only the subject and agent of the second sentence. for the sake of completeness, we want to elaborate briefly on the aspect of competition between possible perspectival centers. as mentioned above, prominence features as we investigate them only differentiate between two possible anchors. if fid is preceded by a sentence with only one protagonist, that protagonist is highly available as the anchor for fid, as in (11) where the little boy is neither the subject nor referred to with a proper name. despite presumably low prominence the fid must be anchored to the protagonist mentioned in the preceding sentence. (11) the ball hit the little boy right in the face. ouch, that hurt! competition of potential anchors for an utterance in fid may also play a role with respect to preceding sentences with two interacting protagonists, as in (12). while in (12a) martin is naturally perceived to be the perspectival center, in (12b) the subject´s prominence status is raised so that the availability of martin as the perspectival center is weakened. consequently, it is slightly awkward and unexpected (though by no means impossible) to interpret the second sentence in (12b) as a thought of martin rendered via fid. in other words, the competition of lilly, now outweighing the prominence status of martin, has an impact on the availability of a referent as the perspectival center. (12) a. a young lady asked martin for the way to the station. hmm, didn’t she see the sign? b. lilly asked martin for the way to the station. hmm, didn’t she see the sign? before we close our discussion of the availability of protagonists as anchors for fid, we want to consider an example with protagonists of equal prominence. in (13) neither a nor b seem to 11 meuser, hörl and hinterwimmer be appropriate. we tentatively suggest an explanation along the following lines: in order to be available as perspectival center, a single protagonist has to be prominent. in the case of (13), neither tina nor mike is prominent – rather, it is just the plurality referred to by the np tina and mike. consequently, both (13a) and (13b) are awkward, since in either case the fid has to be ascribed to a protagonist that is not maximally prominent. (13) tina and mike were fighting all night. a. tonight that jerk had pushed it too far. b. tonight that bitch had pushed it too far. summarizing the discussion so far, we see that a referent must either be locally prominent in the sentence preceding the utterance in fid mode, e.g., with regard to the grammatical function, the thematic role or the type of referential expression, or be globally prominent in a context where at least two protagonists compete. so far, we have shown that a prominence-based account provides promising insights in this regard. at the same time, prominence-lending cues comprise a wide range of linguistic markers – either in the sentence immediately preceding an utterance in fid mode or in a stretch of discourse – that promote a referent to be the perspectival center. the prominence-lending cues elaborated above are by no means exhaustive. yet, in order to gain a better understanding of how prominence and perspectival centers are related we will present a first attempt at answering that question by investigating how grammatical function and type of referential expression contribute to a referent´s potential to serve as the anchor for an utterance in fid mode. while the first experiment presented in the next section will give preliminary insights that suggest a maximally prominent referent – in terms of grammatical function, type of referential expression and number of mentions – is preferred as the anchor for a sentence in fid mode, the experiment presented in section 2.2 will investigate the impact of grammatical function and referential expression and their interaction. finally, in section 2.3 we test these findings in larger contexts providing first empirical evidence for hinterwimmer’s (2019) assumption (see also abrusán, 2021 for discussion) that locally established prominence interacts with globally established prominence. 2. the experimental studies 2.1 experiment 1 in our first acceptability rating tasks we asked participants to rate the naturalness of short text segments such as (14) and (15) in which two utterances in fid mode or a neutral sentence must be interpreted with regard to the context provided in the first two sentences. (14) als die hochzeit von prinz william und kate im fernsehen übertragen wurde, konnte robert seine eigene hochzeit kaum erwarten. auch er hatte seiner freundin einen antrag gemacht. a. fidpreferred: schon morgen würde er mit seiner liebsten vor den altar treten. b. fiddispreferred: schon morgen würde sie mit ihrem liebsten vor den altar treten. c. neutral: sie wollte mit ihm vor den altar treten. when the wedding of prince william and kate was broadcast on tv, robert could hardly wait for his own wedding. he, too, had proposed to his girlfriend. 12 perspective-taking and protagonist prominence a. fidpreferred: soon he would walk down the aisle with his darling. b. fiddispreferred: soon she would walk down the aisle with her darling. c. neutral: she wanted to walk down the aisle with him. (15) als der letzte band von ”harry potter” erschien, kramte luisa ihr taschengeld zusammen. sofort sagte sie ihrem besten freund bescheid. a. fidpreferred: morgen schon würde sie mit diesem bücherwurm die buchhandlung stürmen. b. fiddispreferred: morgen schon würde er mit dieser leseratte die buchhandlung stürmen. c. neutral: er wollte mit ihr am nächsten tag in die buchhandlung gehen. when the last harry potter book was published, luisa scraped up her money. she told her best friendmale right away. a. fidpreferred: tomorrow she would hit the bookstore with that bookwormmale. b. fiddispreferred: tomorrow he would hit the bookstore with that bookwormfemale. c. neutral: he wanted to go to the bookstore with her the next day. in our experiment, we will collect evidence for the hypothesis that follows the discussion presented above: h1: a protagonist that is more prominent in terms of referential expression and grammatical function is more available as the anchor for fid than a competing referent. in order to determine the influence of prominence on the acceptability of fid we distinguished our two referents by means of maximal and minimal prominence in terms of grammatical function, the number of occurrences and the type of referring expression. we predict that fid anchored to the more prominent referent, i.e., the agent in subject position, introduced by a proper name and picked up by a personal pronoun (compare (14a)/(15a)) will more likely be accepted as the perspectival center of a sentence in fid than the competing referent mentioned only once in object position (compare (14b)/(15b)). the neutral condition (compare (14c)/(15c)) is similar to the dispreferred condition where the less prominent referent is the subject of the target sentence. the neutral condition serves as a control condition: if the neutral condition is rated significantly better than the dispreferred fid condition we may conclude that the low acceptability of the dispreferred condition is at least partly due to a low acceptability of the fid regardless of the change of the subject (another relevant factor might be the presence of rather unusual terms such as bookworm in the fid conditions as opposed to the neutral condition)10. 2.1.1 material we constructed 22 short stories similar to examples (14) and (15) in 3 conditions. all stories consisted of 3 sentences. the first sentence introduced one referent (r1) in subject position with a proper name. also, the first sentence set the context to a specific event in the past (e.g., when germany won the world championship, last valentine’s day). we used an explicit reference to the past in order to emphasize the contrast of the narrative context and the deictic expression used in the fid, which in all cases refers to the present (now, today) or an immediate or close future (soon, tomorrow). in the second sentence the referent was picked up by a personal pronoun in subject 10. we are grateful to an anonymous reviewer for pointing this out to us. 13 meuser, hörl and hinterwimmer position11 interacting with a second referent (r2) introduced by a noun phrase anchored to the first referent with a possessive pronoun (e.g., his friend) in object position. the two referents differ in gender in order to prevent ambiguity of the target sentence12. the two fid conditions in the third sentence feature at least 3 indicators of fid: a temporal deictic expression, a verb in subjunctive ii mood (german würde) and kinship, family terms or qualitative nouns (e.g., her/his darling). in our fid conditions the target sentence varied with respect to r1 and r2. in the first condition the fid can only be anchored to r1, i.e., the referent introduced in the first sentence by a proper name, by the choice of the respective male or female pronoun and a qualitative noun phrase matching the gender of the second referent (compare 14a: he [. . . ] his darling). the second condition features fid anchored to r2, the referent introduced in the second sentence, by the choice of a pronoun that matches the second protagonist´s gender whereas the qualitative expression must refer to the protagonist introduced in the first sentence (compare (14) b: she [. . . ] her darling). we paid particular attention to the fact that the fid can equally plausibly be linked to either one of two referents, i.e., the fiancée is just as likely to think the thought regarding their wedding plans as robert is. a third, neutral condition (compare (14c)/(15c)) does not feature any indicators of fid, but it matches the second condition, i.e., r2 is in subject position. 2.1.2 procedure we randomly distributed 22 items across 3 lists. unfortunately, this led to an unbalanced distribution of the three conditions across the lists – per participant 7 responses for two conditions and 8 for the third13. the 22 items were randomly mixed with 44 fillers on a paper-pencil questionnaire. fillers were similar to the items in length, syntax and the number of referents mentioned. in order to mask our manipulation, we constructed fillers deliberately violating constraints of demonstrative and personal pronouns resulting in odd, but nevertheless grammatical references to yield low acceptability. the task was to judge naturalness of the third sentence in the context of the first two on a scale from 1, entirely unnatural, to 7, entirely natural. in order to prevent participants from judging according to grammatical acceptability or plausibility of the scenario the instructions stressed that all sentences are well-formed and the stories are plausible. participants were encouraged to use the entire scale from 1 to 7 according to their subjective intuitions and fill in the number in a bracket behind each item14. in april 2017, 89 students of the university of cologne, participated voluntarily in our experiment. 14 participants got excluded from analysis as they were not monolingual speakers of german or did not complete the entire questionnaire. 11. we are aware that subjecthood coincides with first mention. however, we decided to keep the design simple and neglected a potential ovs condition, especially as the ovs syntax is highly marked. testing an ovs condition may lead to unintended confounds. we need to point out that the effect that we attribute to subjecthood throughout the paper may also be the result of first mention – or potentially the interaction of subjecthood and first mention. 12. gender of the prominent protagonist was equally balanced. nevertheless, gender stereotypes were avoided in order to prevent preferences of one referent as anchor for the fid depending on gender. 13. although this was not an ideal distribution, since we used linear mixed-effects modeling for analyzing the data, we relied on the robustness of mixed-effects modeling to handle missing data (baayen et al., 2008) and assume that the ”missing” one response per participant had negligible influence on our effects. 14. in our analysis we treated the scale as an interval scale as the scale was not further labeled, we assume that participants used the scale continuously. 14 perspective-taking and protagonist prominence 2.1.3 data analysis and results ratings from 75 questionnaires were considered in the analysis. acceptability judgements were elicited for 21 items in 3 conditions. in our experiment we tested 22 items, one of which had to be excluded due to the implausibility of the scenario attentive participants commented on15. this resulted in a total of 1575 observations. in order to test differences between conditions we modeled the data using bayesian mixed effects models 16 estimated in r17 (r core team, 2015). we used a model with acceptability scores as outcome and condition as the sole predictor with random intercept and random slopes for condition, for participants and items18. analogously to frequentist statistics we considered effects to be significant if the 95%-credible interval does not include 019. the model estimate for the fidpreferred condition was 4.64 whereas the model estimate for the fiddispreferred condition was 3.63. the model revealed a significant difference of 1.01 points (se = 0.16, ci: [0.69;1.33]). for the neutral condition the model estimate was 5.04. it served as a control condition in order to test whether a lower acceptability of the fiddispreferred version is due to a change in subject. the comparisons revealed that participants rated the neutral version significantly better than the fid dispreferred condition (estimate = 1.40; se = 0.18; ci: [1.06;1.75]). in our experiment, we tested the claim that fid is more likely to be anchored to a maximally prominent referent than to a competing referent with rather low prominence (hinterwimmer, 2019). the significant difference in the ratings of fid in our preferred condition in comparison to our fiddispreferred condition indicates that the hypothesis that all referents are equally available as the perspectival center as long as the interpretation of the fid is plausible can be rejected. the comparison of the two fid conditions supports the hypothesis that the maximally prominent referent is preferred as perspectival center. however, the mediocre ratings, −0.49 below average, of the fiddispreferred condition show that the interpretation of the fid as a thought of the second referent is not entirely impossible while the significant difference indicates a clear preference. although we did not test this contrast, the difference of the neutral and the preferred condition suggests that both our fid conditions are not as acceptable as neutral story continuations. lower ratings in the sample of the fidpreferred condition in which the fid is anchored to the prominent referent compared to the neutral condition supports concerns that fid in such short stories is rather unnatural for presumably two – not necessarily unrelated – reasons: fid is a phenomenon typically 15. participants reported implausibility orally and/or left a comment on the questionnaire, i.e., it was falsely claimed that the rock am ring festival takes place in february not in early summer. 16. in hinterwimmer and meuser (2019) we presented the results of this experiment using a non-bayesian linear mixed effects models with the lmer function from the lme4 package. with the data of experiment 2 however, this prohibited the use of a maximal random effects structure as convergence issues occurred. in line with barr et al. (2013) we decided to keep the maximal random effects structure by opting for a bayesian analysis. contrary to frequentist mixed models bayesian mixed models are more stable in the sense that they almost always converge even with complex random effects structures. we therefore decided to use bayesian mixed models for the whole analysis to keep it consistent and maintain the use of a maximal random effects structure in experiment 2. for an introduction to bayesian mixed models see: https://www.sciencedirect.com/science/article/abs/pii/s0095447017302310. 17. using the brms package (bürkner, 2017). we used weakly informative priors, e.g., for the condition coefficients we used a normally distributed prior with mean 0 and sd 5. 18. we used treatment contrasts with fiddispreferred as the reference category. we specified the model as follows: ”acceptability 1 + condition+ (1 + condition|participant) + (1 + condition|item)” 19. if 0 is not included in the 95%credible interval, then 0 is not included in the range of 95% of the most plausible values. this means that the posterior probability of the effect being 0 (or less extreme) is maximally 5%. 15 meuser, hörl and hinterwimmer found in narratives and can thus possibly not be perceived to sound entirely natural in 3-sentence stories. second, if the acceptability of fid depends on prominence, the prominence of a referent must be sufficiently established. the latter assumption will be subject to further research. figure 1: experiment 1: acceptability ratings (model estimate and 95%-credible interval) 2.1.4 discussion a shortcoming of our acceptability study concerns the comparability of the test items we used. the task explicitly stated that the third sentence has to be judged with regard to the first two sentences, although it is possible that participants rated the acceptability of just the third sentence and those differed across conditions. in order to anchor the fid to one of the two referents we used kinship, family terms or qualitative nouns that match only one of the two referents in terms of gender. in cases where we compare minimal pairs such as liebster (male darling) to liebste (female darling) – both commonly used to refer to a loved one – that should not make a difference. however, there are only few such perfect minimal pairs so that in most cases we ended up comparing bücherwurm (m., bookworm) with leseratte (f., bookworm, literally a reading rat), schleckermaul (n., sweet tooth literally a yummy mouth) with naschkatze (f., again a sweet tooth, literally a nibble cat) and tratschtante (f., gossip aunt) with lästermaul (n., gossip mouth) for which we do not have any predictions regarding their overall acceptability. in order to make sure that results were not related to the potentially lower acceptability of the expressions used in the fid dispreferred condition in our follow-up we improved the comparability of the target sentence. 16 perspective-taking and protagonist prominence 2.2 experiment 2 we regard the significant contrasts found between the fid conditions in the first study presented above as indicators that acceptability judgements indicate anchoring preferences of fid. in the previous experiments acceptability varied with respect to the naturalness of fid anchored to a maximally prominent protagonist compared to fid anchored to a minimally prominent protagonist20. now, in our second experiment, the same methodology was used to measure the differing anchoring potential of referents depending on their prominence status. prominence-lending cues we investigated were grammatical function as well as type of referential expression. in an acceptability rating study items similar to (16) were tested. again, a short story was followed by an utterance in fid mode that can only be anchored to one of the two referents in terms of gender. (16) a. lynn sprach pablo auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. b. pablo sprach lynn auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. c. eine reisende sprach pablo auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. d. ein reisender sprach lynn auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. e. eine reisende sprach einen reisenden auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. f. ein reisender sprach eine reisende auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. g. lynn sprach einen reisenden auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. h. pablo sprach eine reisende auf den nächsten zug an, als eine durchsage verspätungen aller züge aufgrund starken schneefalls verkündete. fid: oh mann, jetzt würde sie bestimmt ihren anschlusszug verpassen. a) lynn asked pablo about the train when severe delays due to heavy snowfall were announced. b. pablo asked lynn about the train when severe delays due to heavy snowfall were announced. c. a travelerf asked pablo about the train when severe delays due to heavy snowfall were announced. d. a travelerm asked lynn about the train when severe delays due to heavy snowfall were announced. 20. we will refer to a protagonist referred to with a proper name in subject position as maximally prominent while a protagonist in object position referred to with an indefinite is minimally prominent. 17 meuser, hörl and hinterwimmer e. a travelerf asked a travelerm about the train when severe delays due to heavy snowfall were announced. f. a travelerm asked a travelerf about the train when severe delays due to heavy snowfall were announced. g. lynn asked a travelerm about the train when severe delays due to heavy snowfall were announced. h. pablo asked a travelerf about the train when severe delays due to heavy snowfall were announced. fid: oh man, now she would miss her connecting train! in line with the argumentation presented in section 1.4, the study investigates two main hypotheses: h2: protagonists functioning as subjects are more available as anchors for fid than protagonists functioning as objects. h3: protagonists referred to with a proper name are more available as anchors for fid than protagonists introduced by an indefinite np. in the hypotheses we make claims with respect to the two prominence-lending cues, i.e., grammatical function and the type of referential expression, in isolation. however, as a referent holds both prominence-lending cues at the same time, i.e., the referent is referred to in a certain way and the expression has a certain grammatical function, it is worthwhile to explore the question of how these different prominence-lending cues interact (q1): does the type of referential expression play a bigger role than the grammatical function with respect to the anchoring of fid? based on the discussion presented in section 1.4, a second question that calls for an exploratory investigation is whether the competition of the two protagonists has an impact on the anchoring preference. a similar effect as the one presented in example (11) (”a young lady asked martin/ lilly asked martin...”), where the prominence status of the object changes as the subject gets more prominent, shall be investigated by considering the effect of the expression used to refer to the second referent. the second question of interest (q2) may thus be phrased as is it more acceptable to anchor fid to a referent when the competing referent is less prominent in terms of the expression by which she or he is referred to? 2.2.1 material for this experiment we constructed 48 test items similar to (16) in 8 conditions. items consisted of two sentences. in the first sentence (s1) two referents (r1 & r2), one male and one female, were introduced as subject and object, respectively. in the second sentence (s2) an utterance in fid mode was presented that could only be anchored to r1 in terms of gender21. unlike the first experiment we varied the gender features of the two referents in the context so that the target sentence would 21. a reading in which the referent who is not referred to with the pronoun is the anchor of the fid, i.e., pablo thinks about that poor traveler that she would miss her connecting train, is in theory possible. however, that reading should yield equally low results as it is rather absurd. if a reader accepts the dispreferred conditions as a recursive reading, i.e., pablo thinks that she thinks, that reading is again in line with the hypothesis: the fid is anchored to the more prominent referent even if that results in a rather absurd interpretation. 18 perspective-taking and protagonist prominence condition label subject-object ref.-ex r1 ref.-ex r2 a: r1name – r2name r1-r2 name name b: r2name – r1name r2-r1 name name c: r1indef. – r2name r1-r2 indef. name d: r2indef. – r1name r2-r1 name indef. e: r1indef. – r2indef. r1-r2 indef. indef. f: r2indef. – r1indef. r2-r1 indef. indef. g: r1name – r2indef. r1-r2 name indef. h: r2name – r1indef. r2-r1 indef. name table 1: experimental conditions in experiment 2 remain exactly the same throughout conditions and yet varies with respect to the two referents. the referents were either referred to with a proper name or with an indefinite noun phrase (np) and functioned as subjects or objects in s1 resulting in 8 possible combinations presented in table 1. we will refer to the different conditions with respect to the subject in s1 and the object in s1. for half of the items we chose a male referent for r1 while the other half were female in order to prevent gender biases. in all items the utterance in fid mode featured an interjection, a temporal deictic expression and a verb in subjunctive ii mood (german würde). to be able to check if participants attentively read the items 16 filler items were included that were designed to yield low rating while another 16 filler items should be unproblematic and yield high ratings22. participants that rate the deliberately bad items better than the good ones will be excluded from the analyses. 2.2.2 procedure all test items were equally distributed across 8 lists so that each participant saw each item in only one condition. items were randomized for each participant and mixed with 64 filler items. fillers are similar to the test items in terms of length. due to the rather large number of items – 112 in total – a 3-minute break was enforced after half of the items were presented. participants were asked to rate the naturalness of the item on a 7-point likert scale – 1, entirely unnatural, to 7, entirely natural. participants were recruited in class and participated for course credit. the experiment was programmed using pcibex (zehr and schwarz, 2018). in december 2021, ratings from 129 participants were recorded. results from 7 participants had to be excluded from the analysis as they were not native speakers of german. further, we excluded the results from 3 participants as they rated the fillers that were deliberately designed to trigger low acceptability better than a set of deliberately good filler items. 22. example for a good filler item: marleen und laurin kochen beide gerne. marleen mag die argentinische küche und laurin mag auch südamerikanisches essen. (engl. marleen und laurin both love to cook. marleen likes the argentinian cuisine and laurin also likes south american food.) example for a bad filler item: melina und thomas haben beide schöne haare. melina hat schwarze haare und thomas hat auch blonde haare. (engl. melina and thomas both have beautiful hair. melina has black hair and thomas also has blond hair.) 19 meuser, hörl and hinterwimmer 2.2.3 data analysis and results ratings from 119 questionnaires were considered in the analysis. acceptability judgements were elicited for 48 items in 8 conditions. acceptability judgements were submitted to a bayesian mixed effects model. we used a full set of interactions of the factors grammatical function and referential expression of r1 and r2. we also included maximal random effects structure for subjects and items23. there is a significant main effect of grammatical function of 0.57 (se = 0.07, ci: [0.43;0.72]). this means that sentences with r1, i.e., the anchor for the utterance in fid mode, in subject position were rated better than sentences with r1 in object position. additionally, there is a significant main effect of the referential expression used for r1 (estimate = -0.62, se = 0.07, ci: [-0.75; -0.49]) indicating that referents referred to with a proper name are more acceptable as anchors for fid than referents referred to with an indefinite np. there is also a significant main effect of the referential form of r2 (estimate = 0.37, se = 0.06, ci: [0.26; 0.47], see figure 2), i.e., sentences received higher ratings when r2, the referent that cannot be the anchor for the utterance in fid mode, was referred to with an indefinite than with a proper name. there was a significant two-way interaction of the referential form of r1 and r2 (estimate = 0.18, se = 0.08, ci: [0.03; 0.34], see figure 2). this suggest that the effect of r2 depends on the referential expression used to refer to r1, that is that r2 plays a bigger role if r1 is referred with an indefinite np. however, an investigation of the three-way interaction of the referential form of r1 and r2 and the grammatical function reveals a more nuanced interplay of these factors. the interaction effect of the referential expressions of r1 and r2 varies significantly depending on the grammatical function (estimate = 0.64, se = 0.18, ci: [0.29; 1.01]). more specifically, the difference of the effect of r2 between r1name and r1indef. is higher when r1 is in subject position (compared to r1 in object position). when relating this to figure 2, it becomes clear that ’higher’ here means that the effect of r2 seems to differ between r1name and r1indef. to a noteworthy extent only when r1 is in subject position ((g-a) ¡ (e-c)). in the case where r1 is in object position the effect of r2 does not seem to differ much between r1name and r1indef. ((d-b) ≈ (f-h)). this implies that the two-way interaction between the referential form of r1 and r2 which was described earlier can be somewhat misleading, because it is (almost exclusively) driven by the effect of r2 when r1 is in subject position. therefore, the effect of r2 does not seem to generally depend on r1 but only when r1 is in subject position. here, the role of r2 is high(er) when r1 is referred to with an indefinite np (e-c) and almost negligible when r1 is referred to with a proper name (g-a). 23. we specified the model as follows: acceptability 1+ gramfunc ∗r1reftype ∗r2reftype+ (1+ gramfunc ∗ r1reftype ∗r2reftype|participant)+ (1+ gramfunc ∗r1reftype ∗r2reftype|item). in order to interpret the lower order effects as main effects we used deviation coding for all 3 factors (0.5/+0.5). 20 perspective-taking and protagonist prominence figure 2: experiment 2: acceptability ratings condition model estimate ci a: r1name – r2name 5.37 [5.15 ; 5.61] b: r2name – r1name 4.68 [4.44 ; 4.91] c: r1indef. – r2name 4.55 [4.31 ; 4.79] d: r2indef. – r1name 5.12 [4.90 ; 5.36] e: r1indef. – r2indef. 5.16 [4.93 ; 5.38] f: r2indef. – r1indef. 4.39 [4.14 ; 4.64] g: r1name – r2indef. 5.48 [5.25 ; 5.72] h: r2name – r1indef. 4.08 [3.81 ; 4.33] table 2: experiment 2: acceptability ratings (model estimate and 95%-credible interval) 2.2.4 discussion the significant main effect of grammatical function indicates that protagonists in subject position are more available as anchors for fid than protagonists in object position (h2). regarding these results we thus conclude that grammatical function can be considered a prominence-lending cue with respect to establishing the perspectival center that serves as the anchor for an utterance in fid mode. recall that r1 is the referent that the utterance in fid has to be anchored to in terms of gender. the significant main effect of referential expression of r1 indicates that protagonists referred to with a proper name are more available as anchors for fid than protagonists introduced by an indefinite np (h3). as both factors add to the referent´s prominence status, as discussed in section 1.4, we also investigated the impact of each factor and how the two prominence-lending cues interact. the estimates 21 meuser, hörl and hinterwimmer for grammatical function and referential expression indicate that the role of subjecthood is nearly similar to that of the type of referential expression. we therefore want to make the following claim with respect to q1: grammatical function affects the anchoring of fid just as much as type of referential expression. while the main effect of the referential expression of r1 indicates that r1 is more available if it is referred to with a proper name, the main effect of r2 counteracts in a similar way: just as r1 becomes more or less available as the referential expression increases or decreases its prominence status, so does r2. now, if r2 becomes more prominent in virtue of the referential expression, it becomes more available as the perspectival center and therefore both protagonists compete. with respect to q2 we may thus conclude that a high prominence status of the competing referent (r2) makes the anchoring of fid to r1 particularly difficult. while we found a main effect for the referential expression, we must also point out that in our sample the effect of referential expression of r2 does not show in the case of r1 being in subject position and referred to with a proper name. conditions a and g received rather similar high ratings (compared to the oppositional pairs e&c, f&h and d&b). therefore, we want to suggest the following: a referent that is maximally prominent in virtue of being referred to with a proper name in subject position is the default anchor for an utterance in fid mode regardless of the type of referential expression used to refer to the competing referent. complementary to the reporting of the main and interaction effects we would like to provide a matching interpretation from a slightly different perspective. that is, the three factors increase the acceptability of r1 as anchor for the fid expression: a proper name as referential form of r1, subject position of r1 and an indefinite expression that refers to the competing referent r2 enhance the availability of r1 to serve as the perspectival center. if we now sort the conditions with respect to the number of these features we get the following pattern: condition h with no such enhancing features yielded the lowest acceptability (4.08). conditions b, c and f with one enhancing feature have a slightly higher acceptability (from 4.39 to 4.67). conditions with two features, namely a, d and e, are even more acceptable (from 5.11 to 5.37) while condition g with all three features has the highest acceptability (5.47). figure 3 illustrates this. before we forge ahead to the investigation of prominence in larger discourse, we briefly want to recall the first experiment presented in this chapter and point out how the main effects found in the second experiment explain the contrast found in the first experiment. in our first experiment we presented an attempt to verify our intuitions that referents that are prominent with respect to referential expression and grammatical function are more easily available as perspectival centers than competing referents, aiming to find a contrast between maximally and minimally prominent protagonists. contrasts found for the fid preferred conditions and the fiddispreferred conditions are parallel to the contrasts found between the conditions g and h in the second experiment (see (17)). here again, the referent that serves as the anchor for the fid is maximally prominent in terms of referential expression and grammatical function in condition g and minimally prominent in condition h. 22 perspective-taking and protagonist prominence figure 3: experiment 2: acceptability ratings depending on the number of prominence-enhancing features (17) exp.1: when the wedding of prince william and kate was broadcast on tv, robert could hardly wait for his own wedding. he, too, had proposed to his girlfriend. fidpreferred: soon he would walk down the aisle with his darling. fiddispreferred: soon she would walk down the aisle with her darling. exp. 2: g: r1name – r2indef.: lynn asked a travelerm about the train when severe delays due to heavy snowfall were announced. h: r2name – r1indef.: pablo asked a travelerf about the train when severe delays due to heavy snowfall were announced. oh man, now she would miss her connecting train! the strong contrast between the acceptability judgements of condition g and h (1.39 points) as well as the contrasts found for the fidpreferred conditions and the fiddispreferred conditions (1.02 points) may be explained by the main effects reported above: a protagonist is more available as the anchor for an utterance in fid mode if it is in subject position and referred to with a proper name than when it is referred to with an indefinite np and the competing referent is referred to with an indefinite np instead of a proper name. the results thus support the intuitions outlined in chapter 1.4 that lead to the item construction (repeated in (17)), i.e., it feels more natural to anchor an utterance in fid mode to robert than to the girlfriend (exp.1) and to lynn rather than to the traveler (exp.2). 23 meuser, hörl and hinterwimmer 2.3 experiment 3 building on the evidence we found in experiment 1 and 2, which suggests that locally prominent referents are preferred as perspectival centers, the effect of local prominence in larger discourse is investigated in experiment 3. in particular, we want to test if the effect of local prominence prevails in a context where a competing referent is globally prominent. recall the discussion presented in section 1.4 where it was claimed that a referent that serves as the discourse topic serves as the perspectival center regardless of any reference to a competing referent in a sentence preceding the fid, as in (10) (repeated in (18)). (18) lisa was playing in the schoolyard. a boy pushed her into the stinging nettles. ouch, that itched! even though the boy is the subject and the agent of the sentence preceding the fid we have no difficulties anchoring the fid to lisa, who is the object and the patient. for this example, however, it may be claimed that the preference for lisa is the result of multiple prominence lending cues: lisa is the subject of the first sentence, she is introduced with a proper name and she is picked up with a personal pronoun in the second sentence, i.e., she is mentioned twice – unlike the competitor, who is only referred to once with an indefinite np. in order to investigate the effect of global as opposed to local prominence more systematically, we conducted an experiment where two referents are introduced with a proper name. however, they vary with respect to how often they are mentioned as well as with respect to their grammatical function in the sentence preceding the fid. in a third acceptability rating study, we presented larger contexts similar to (19). again, participants (n = 116) were asked to rate the naturalness of a target sentence in fid mode with respect to the context. (19) a: kein feierabend in sicht (s1) im büro gehörten überstunden zum alltag. (s2) da musste ein ruhiger kopf bewahrt werden. (s3) eine gute arbeitsteilung war notwendig, um alles zu erledigen. (s4) trotz des hohen zeitdrucks duldete das management keine verzögerungen. (s5) fred gab caroline die akten. (s6) wehe, wenn die ihn heute hängen lassen würde. b: kein feierabend in sicht (s1) im büro gehörten überstunden zum alltag. (s2) da musste ein ruhiger kopf bewahrt werden. (s3) eine gute arbeitsteilung war notwendig, um alles zu erledigen. (s4) trotz des hohen zeitdrucks duldete das management keine verzögerungen. (s5) caroline gab fred die akten. (s6) wehe, wenn die ihn heute hängen lassen würde. c: kein feierabend in sicht (s1) im büro gehörten überstunden zum alltag. (s2) caroline atmete einmal tief durch. (s3) in der frühstückspause hatte sie fred um unterstützung gebeten. (s4) sie wusste, dass das management keine weiteren verzögerungen mehr tolerieren würde. (s5) fred gab caroline die akten. (s6) wehe, wenn die ihn heute hängen lassen würde. 24 perspective-taking and protagonist prominence d: kein feierabend in sicht (s1) im büro gehörten überstunden zum alltag. (s2) fred atmete einmal tief durch. (s3) in der frühstückspause hatte er caroline um unterstützung gebeten. (s4) er wusste, dass das management keine weiteren verzögerungen mehr tolerieren würde. (s5) caroline gab fred die akten. (s6) wehe, wenn die ihn heute hängen lassen würde. e: kein feierabend für caroline (s1) caroline hatte einen langen arbeitstag vor sich. (s2) sie atmete einmal tief durch. (s3) in der frühstückspause hatte sie fred um unterstützung gebeten. (s4) sie wusste, dass das management keine weiteren verzögerungen mehr tolerieren würde. (s5) fred gab caroline die akten. (s6) wehe, wenn die ihn heute hängen lassen würde. f: kein feierabend für fred (s1) fred hatte einen langen arbeitstag vor sich. (s2) er atmete einmal tief durch. (s3) in der frühstückspause hatte er caroline um unterstützung gebeten. (s4) er wusste, dass das management keine weiteren verzögerungen mehr tolerieren würde. (s5) caroline gab fred die akten. (s6) wehe, wenn die ihn heute hängen lassen würde. a: more work to do (s1) in the office, working overtime was not unusual. (s2) keeping a clear head was important. (s3) also dividing work was crucial to get everything done. (s4) despite the tremendous time pressure, the management did not tolerate delays. (s5) fred handed caroline the papers. (s6) she better not let him down today. b: more work to do (s1) in the office, working overtime was not unusual. (s2) keeping a clear head was important. (s3) also dividing work was crucial to get everything done. (s4) despite the tremendous time pressure, the management did not tolerate delays. (s5) caroline handed fred the papers. (s6) she better not let him down today. c: more work to do (s1) in the office, working overtime was not unusual. (s2) caroline took a deep breath. (s3) during the morning break she had asked fred for help. (s4) she knew that the management did not tolerate delays. (s5) fred handed caroline the papers. (s6) she better not let him down today. d: more work to do (s1) in the office, working overtime was not unusual. (s2) fred took a deep breath. (s3) during the morning break he had asked caroline for help. (s4) he knew that the management did not tolerate delays. (s5) caroline handed fred the papers. (s6) she better not let him down today. e: no end of work for caroline (s1) caroline was facing a long day at work. (s2) she took a deep breath. (s3) during the morning break she had asked fred for help. (s4) she knew that the management did not tolerate delays. (s5) fred handed caroline the papers. (s6) she better not let him down today. 25 meuser, hörl and hinterwimmer f: no end of work for fred (s1) fred was facing a long day at work. (s2) he took a deep breath. (s3) during the morning break he had asked caroline for help. (s4) he knew that the management did not tolerate delays. (s5) caroline handed fred the papers. (s6) she better not let him down today. the first two conditions serve as a baseline. based on the results presented above we expect them to yield results in accordance with h2. h2: protagonists functioning as subjects are more available as anchors for fid than protagonists functioning as objects – if there is no competing referent activated in the context. for the items similar to condition c and d we expect to find an effect of global prominence, i.e., the referent that is mentioned in the context should also be available as the perspectival center. h4: protagonists that are highly activated in a discourse in terms of being repeatedly mentioned in subject position are just as available as anchors for fid as protagonists that are less often referred to in a discourse but in subject position in the sentence preceding the fid. for items similar to (19e) and (19f) we expect to find an even stronger effect, i.e., they are expected to be preferred as perspectival centers over the locally prominent referent. h5: protagonists that are highly activated in a discourse in terms of being established as the discourse topic in a topic-establishing sentence, repeatedly mentioned in subject position, and mentioned in a title are more available as anchors for fid than protagonists that are less often referred to in a discourse but in subject position in the sentence preceding the fid. 2.3.1 material for our third experiment we constructed 18 experimental items similar to (19). all items consist of six-sentence stories. the first four sentences established a neutral context or one where one referent was (weakly or strongly) activated. the last two sentences should be interpreted as part of a main storyline. in sentence 5, the two referents interact, and sentence 6 can only be interpreted as a thought from the perspective of one of the two referents (due to the gender of the pronoun it contains) rendered as fid. again, we manipulated the gender of the two referents in the context so that the target sentence in fid mode remains the same in all conditions. in condition a and b the referents only get introduced in s5, i.e., in the sentence preceding the fid. in condition c, d, e and f both referents are mentioned in the context, yet they differ in the number of references. that is, in conditions c and d the globally prominent referent is mentioned in s2, s3 and s4, while the competitor is mentioned once in s3. in conditions e and f the globally prominent referent is additionally mentioned in the title and in s1 (see example (18) or table 3). 26 perspective-taking and protagonist prominence condition competitor (not the local subject) local subject fid a: x – x – r1 none r1 r1 b: x – x – r2 none r2 r1 c: x – r2 – r1 weak activation r1 r1 d: x – r1 – r2 weak activation r2 r1 e: r2 – r2 – r1 strong activation r1 r1 f: r1 – r1 – r2 strong activation r2 r1 table 3: experimental conditions in experiment 3 2.3.2 procedure all conditions were distributed over six lists so that each participant saw each item in only one condition. the 18 experimental items were randomly mixed with 28 fillers (similar to the experimental items in terms of content, style, and length). in order to assure that participants carefully read the entire story, rather than just the target sentence, comprehension questions randomly asked for information presented in the fillers. as in experiment 1 and 2, participants’ task was to rate the naturalness of the last sentence in the context of the preceding sentences on a scale from 1 (entirely unnatural), to 7 (entirely natural). the experiment was conducted on qualtrics 24. in may 2020, 122 undergraduate students of the university of cologne participated in the experiment for course credit. 2.3.3 data analysis and results data from six participants had to be excluded from the analysis as they performed poorly on the comprehension questions (8 or fewer out of 12, mean accuracy = 11.19). the remainder of the data (n=116) was analyzed using bayesian mixed effects models estimated in r (r core team, 2015). the model used acceptability ratings as the outcome and the two factors of local subject and reference in the context as predictors with random intercept and random slopes for condition, participants, and items25,26,27. comparisons of individual conditions were calculated using the emmeans package (?). 24. https://www.qualtrics.com 25. treatment contrasts were used with condition a (no competitor/r1) as the reference category. the model was specified as follows: ( competitor ∗ local + (1 + competitor ∗ local|item) + (1 + competitor ∗ local|participant) 26. the raw data and the script are available online: https://osf.io/wegtn/ 27. analogously to frequentist statistics, effects were considered to be significant if the 95%-credible interval did not include 0. if 0 is not included in the 95%-credible interval, then 0 is not included in the range of 95% of the most plausible values. this means that the posterior probability of the effect being 0 (or less extreme) is maximally 5%. 27 meuser, hörl and hinterwimmer condition model estimate 95%-ci a: x – x – r1 4.76 [4.35; 5.22] b: x – x – r2 4.16 [3.71; 4.59] c: x – r2 – r1 4.49 [4.03; 4.95] d: x – r1 – r2 4.73 [4.27; 5.19] e: r2 – r2 – r1 4.53 [4.09; 4.98] f: r1 – r1 – r2 4.87 [4.42; 5.34] table 4: experiment 3: acceptability ratings (model estimates and 95%-credible interval) a comparison of condition a and b confirms h2 and thus replicates the results of experiments 1 and 2 (protagonists functioning as subjects are more available as anchors for fid than protagonists functioning as objects – if there is no competing referent activated in the context). that is, the model difference of -0.60 points (se = 0.13, ci: [-0.86; -0.35]) indicates that items with r1 in subject position in s5 were rated significantly higher than items with r2 in subject position. h4 (i.e., protagonists that are highly activated in a discourse in terms of being repeatedly mentioned in subject position are more available as anchors for fid than protagonists that are less often referred to in a discourse but in subject position in the sentence preceding the fid) was tested in terms of, first, the interaction of conditions a & b and c & d, and, second, the comparison of the model estimates for condition c and d. the difference of a and b compared to the difference of c and d reveals a significant two-way interaction (estimate = 0.83, se = 0.17, ci: [0.49; 1.17]). this indicates that repeatedly mentioning a referent in the context overrides the subject preference that shows up in the effect observed for a and b. while conditions c & d reverse the effect in a & b, the comparison of the estimates for conditions c and d was not significant (estimate = 0.23, se = 0.13, ci: [-0.04; 0.48]). that is, the globally prominent referent was not preferred significantly in comparison to the locally prominent referent in conditions c and d where the globally referent was weakly activated (compare table 3). likewise, h5 (i.e., protagonists that are highly activated in a discourse in terms of being established as the discourse topic in a topic-establishing sentence, repeatedly mentioned in subject position, and mentioned in a title are more available as anchors for fid than protagonists that are less often referred to in a discourse but in subject position in the sentence preceding the fid) was tested in terms of a comparison of the differences of condition a & b and e & f and a pairwise comparison of conditions e and f. again, the two-way interaction between the difference of condition a & b compared to the difference of e & f is significant (estimate = 0.94, se = 0.17, ci: [0.60; 1.29]). just as for condition c and d, the presence of a competitor overrides the effect of the local subject. unlike the pairwise comparison of condition c and d, with model estimates of 4.49 and 4.73, the difference between condition e and f, with model estimates of 4.53 and 4.87, is significant (estimate = 0.34, se = 0.13, ci: [0.08; 0.60]). 28 perspective-taking and protagonist prominence 2.3.4 discussion higher acceptability ratings for condition a compared to condition b once again indicate a preference for the locally prominent referent as the perspectival center, in particular a preference for the referent that is the subject of the sentence preceding the fid – if there is no competing referent activated in the context. these findings do not only replicate the effect local prominence has on perceptive-taking, rather they serve as a baseline for the manipulation of the context. that is, without any competing referent the target sentence was preferably anchored to the subject of the preceding sentence. this preference was overridden when the object of the preceding sentence was prominent in the context, i.e., globally prominent (compare figure 4). figure 4: the impact of weak and strong globally prominent competitors on local prominence while the effect of the locally prominent referent is canceled when the competing referent is prominent in the context, as indicated by the difference between condition c & d and a & b, there is no significant preference for the globally prominent referent, that is, condition d is rated better yet not significantly higher than condition c. that is, a weak prominence, i.e., three references in the context does not suffice to promote a referent to be the perspectival center when competing with a locally prominent referent. note that the higher activation in the context, in terms of five references in the context and in the title as in condition e and f, override the preference for the locally prominent referent. this allows to conclude that generally the preference for a locally prominent referent as the per29 meuser, hörl and hinterwimmer spectival center can be overridden by the presence of a globally prominent referent. however, the globally prominent referent is only preferred as the perspectival center when it is highly activated in terms of repeated mentioning. the difference between weak (cond. c and d) and strong (cond. e and f) global prominence calls for further investigation. that is, the question of whether it was merely the higher number of references or if being mentioned in the title leads to an increase of the acceptability ratings remains an open question 3. conclusion in this paper we have discussed the role of different prominence-lending cues for the anchoring of sentences in fid mode. while the formal properties distinguishing fid from other forms of speech or thought representation and its interpretative characteristics have been thoroughly investigated from a narratological as well as from a linguistic point of view, the question of how protagonists become available as implicit thinkers or speakers for the anchoring of fid has scarcely been addressed in previous research. this question becomes particularly pressing in cases where more than one protagonist is in principle available as perspectival center, a situation that is hardly uncommon in literary texts. in principle, it is conceivable that the anchoring of fid is resolved on the basis of content exclusively. consequently, whenever it makes sense to interpret a sentence in fid mode as a thought or utterance of a protagonist that has been made available by the preceding linguistic context, that protagonist should be available as an anchor. since antecedent prominence has been shown to play an important role in pronoun resolution, a competing assumption is that a protagonist’s availability as perspectival center is tied to that protagonist’s prominence status. given that not only reference tracking but also the management of perspective-taking is integral for the comprehension of narrative texts, a detailed and finegrained understanding of how perspectival centers are identified is as important for our understanding of the processes underlying the comprehension of narrative texts as a detailed and fine-grained understanding of how pronoun resolution works. the experimental results reported in this paper constitute a first step in this direction. in section 1.4 we outlined our intuitions with respect to anchoring preferences of fid. we showed that in short text segments where two protagonists are presented one is usually preferred as the anchor for a following sentence in fid mode. we argued that the referent that is more prominent with respect to different prominence-lending cues is preferred as the perspectival center. in the series of experiments presented in section 2 we aimed to test these intuitions empirically. our results summarized in section 2.2.4 now support the contrast between example (9a) and (9b) (repeated here as (20)). (20) a. when lisa was playing in the schoolyard, a boy pushed her into the stinging nettles. ouch, that itched! b. when lisa was playing in the schoolyard, she pushed a boy into the stinging nettles. ouch, that itched! 30 perspective-taking and protagonist prominence based on these results, we suggest that the establishment of a referent as the anchor for fid is the result of the interaction of various prominence-lending cues that add to a referent’s prominence status. in order to narrow down the role of different linguistic cues that are involved we provided empirical evidence that, first, allows us to conclude that a sentence in fid mode is perceived as more natural when it is anchored to the more prominent referent and less natural when it is anchored to the less prominent of two competing referents, irrespective of how plausible it is to interpret the sentence as a thought or utterance of the respective protagonist. as discussed in section 1.4, an utterance in fid mode as presented in example (20b) sounds odd or at least surprising. though the utterance must be anchored to the physical sensation of the boy on the level of content, linguistic cues press the reader to step back and ascribe the utterance to a narrating instance – or possibly an act of recursive perspective-taking, i.e., lisa thinks that it must itch for the boy. secondly, we have shown that locally established prominence can be overridden by the presence of a globally prominent referent. the third experiment does not only replicate the findings that the local subject is preferred as the perspectival center by default. rather, the results of experiment 3 reveal the impact of global prominence on local prominence. that is, fid anchored to the subject of the preceding sentence is perceived to be less natural when the competitor was activated in the context. this observation allows for at least two explanations, which do not necessarily exclude each other. rather, it may well be that both are needed since the results of experiment 3 turn out to be the combined effect of two principles that are relevant for protagonists’ perspective taking in narrative discourse. first, perspectival centers may be established in the discourse. locally established prominence only comes into play whenever there is no previous activation of any referent that can serve as the perspectival center. an approach along these lines calls for further investigation of the role of discourse phenomena such as discourse topicality in perspective taking. according to the second principle, the availability of a protagonist as the perspectival center is determined by the sum of prominencelending cues. that is, if subjecthood promotes a referent to be the perspectival center, consequently, the referent that occurs more often in subject position is preferred over the one that occurs less often subject position. the assumption that prominence is the result of the number of prominence-lending cues adding up is also supported by the results of experiment 2, which revealed a correlation between the number of prominence-enhancing features and the acceptability of a discourse referent as the perspectival center (see figure 4). in the research reported in this paper we have relied on acceptability studies exclusively. while being very useful for testing the empirical validity of linguistic hypotheses in a relatively simple and straightforward way, it is a general shortcoming of this method that it is rather far removed from the actual interpretative processes investigated since it cannot be controlled whether participants reflect consciously on the respective task. in future research on perspective taking in narratives, it would therefore be fruitful to employ online methods such as reading time measurements or the tracking of eye movements. also, it would be very insightful to look at corpus data on the issue. this approach has not been perused in depth as – to the best of our knowledge – there is no suitable corpus that includes a 31 meuser, hörl and hinterwimmer sufficient number of examples that match our hypothesis, i.e. annotated instances of fid and two referents that are mentioned in the preceding sentence. a final point worth noting is that it has been pointed out in abrusán (2021) that the preference for locally prominent protagonists as anchors for fid only holds if the content of the sentence in fid mode stands in a subordinating discourse relation such as explanation or elaboration (see asher and vieu, 2005 for an overview over subordinating and coordinating discourse relations) to the content of the preceding sentence. in general, it is very hard to interpret a sentence as fid that is linked to the preceding sentence via a discourse relation such as narration or result. if such a reading is available at all, however, abrusán (2021) claims that the less prominent of two competing referents is preferred as perspectival center. her reasoning is based on examples like the one in (21). (21) sam threatened justin with a knife. he had to defend himself! (abrusán, 2021: 335, ex. (18)) on the most likely reading, the second sentence is interpreted as reporting a thought of sam via fid that provides an explanation for sam’s action reported by the first sentence. since sam is locally more prominent than justin, this is as expected. as observed by abrusán (2021), however, a second reading is at least marginally available on which the second sentence reports a thought of justin rather than sam (via fid). on this reading, the reported thought is interpreted as resulting from sam’s action rather than explaing it, i.e. justin thinks that he has to defend himself as a consequence of his being threatened by sam with a knife. since justin is locally less prominent than sam and since there is no global context available that could overwrite sam’s local prominence, this is unexpected. abrusán (2021) notes, however, that her observations might actually be compatible with the analysis of hinterwimmer (2019): while sentences that are linked via subordinating discourse relations usually have the same topic, sentences linked via coordinating discourse relations move the story forward and hence have the potential to introduce new topics. in the experiments reported in this paper, test items consisted of sentences linked via subordinating discourse relations exclusively. it would be fruitful to empirically test abrusán’s (2021) claims and to investigate how the potential for coordinating discourse relations to shift the topic interacts with global prominence in relation to determining the anchor for fid. we have to leave this as a topic for future research, however. funding information the research reported in this paper was funded by a research grant from the deutsche forschungsgesellschaft. the research reported in this paper was funded by the deutsche forschungsgemeinschaft (dfg) for the project c05 discourse referents as perspectival centers of the collaborative research center 1252 textitprominence in language (university of cologne). references márta abrusán. the spectrum of perspective shift: free indirect discourse vs. protagonist projection. linguistics and philosophy, 44, 2020. doi: https://doi.org/10.1007/s10988-020-09300-z. márta abrusán. computing perspective shift in narratives. in emar maier and andreas stokke, editors, the language of fiction, pages 325–348. oxford university press, oxford, 2021. 32 perspective-taking and protagonist prominence nicholas asher and laure vieu. subordinating and coordinating discourse relations. lingua, 115(4):591–610, 2005. harald baayen, douglas davidson, and douglas bates. mixed-effects modeling with crossed random effects for subjects and items. journal of memory and language, 59:390–412, 2008. doi: https://doi.org/10.1016/j.jml.2007.12.005. ann banfield. unspeakable sentences: narration and representation in the language of fiction. routledge, boston, 1982. dale j. barr, roger levy, christoph scheepers, and harry j. tily. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3): 255–278, 2013. doi: https://doi.org/10.1016/j.jml.2012.11.001. paul bürkner. brms: an r package for bayesian multilevel models using stan. journal of statistical software, 80(1):1–28, 2017. doi: https://doi.org/10.18637/jss.v080.i01. rosalind a crawley and rosemary j stevenson. reference in single sentences and in texts. journal of psycholinguistic research, 19(3):191–210, 1990. doi: https://doi.org/10.1007/bf01077416. regine eckardt. the semantics of free indirect discourse: how texts allow us to mind-read and eavesdrop. brill, leiden, 2014. regine eckardt. perspective and the future-in-the-past. glossa: a journal of general linguistics, 2(1):71, 2017. doi: http://doi.org/10.5334/gjgl.199. norman friedman. point of view in fiction. the development of a critical concept. pmla, 70(5): 1160–1184, 1955. catherine garvey and alfonso caramazza. implicit causality in verbs. linguistic inquiry, 5(3): 459–464, 1974. gérard genette. die erzählung. fink, paderborn, 3rd, revised and corrected edition edition, 2010. morton anne gernsbacher. language comprehension as structure building. lawrence erlbaum associates, hillsdale, 1990. jonathan ginzburg. the interactive stance: meaning for conversation. oxford university press, oxford, 2012. peter c. gordon, barbara j. grosz, and laura a. gilliom. pronouns, names, and the centering of attention in discourse. cognitive science, 17(3):311–347, 1993. doi: https://doi.org/10.1207/ s15516709cog1703 1. jesse a. harris and christopher potts. perspective-shifting with appositives and expressives. linguistics and philosophy, 32(6):523–552, 2009. doi: https://doi.org/10.1007/s10988-010-9070-5. nikolaus p. himmelmann and beatrice primus. prominence beyond prosody: a first approximation. in amedeo de dominicis, editor, ps-prominences: prominences in linguistics. proceedings of the international conference, pages 38–58. disucom press, viterbo, 2015. 33 meuser, hörl and hinterwimmer stefan hinterwimmer. two kinds of perspective taking in narrative texts. proceedings of semantics and linguistic theory, 27:282–301, 2017. doi: https://doi.org/10.3765/salt.v27i0.4153. stefan hinterwimmer. prominent protagonists. journal of pragmatics, 154:79–91, 2019. doi: https://doi.org/10.1016/j.pragma.2017.12.003. stefan hinterwimmer and sara meuser. erlebte rede und protagonistenprominenz. in stefan engelberg, christian fortmann, and irene rapp, editors, redeund gedankenwiedergabe in narrativen strukturen ambiguitäten und varianz, volume 27 of linguistische berichte sonderheft, pages 177–200. buske, hamburg, 2019. richard holton. some telling examples: a reply to tsohatzidis. journal of pragmatics, 28:625–628, 1997. elsi kaiser. perspective-shifting and free indirect discourse: experimental investigations. semantics and linguistic theory, 25:346–372, 2015. doi: https://doi.org/10.3765/salt.v25i0.3436. andrew kehler and hannah rohde. prominence and coherence in a bayesian theory of pronoun interpretation. journal of pragmatics, 154:63–78, 2019. andrew kehler, laura kertz, hannah rohde, and jeffrey l. elman. coherence and coreference revisited. journal of semantics, 25(1):1–44, 2008. doi: https://doi.org/10.1093/jos/ffm018. emar maier. the pragmatics of attraction. explaining unquotation in direct and free indirect discourse. in paul saka and michael johnson, editors, the semantics and pragmatics of quotation, pages 259–280. springer, dordrecht, 2017. sara meuser. how free is free indirect discourse? empirical approaches to the anchoring mechanisms of perspective-taking. phd thesis, universität zu köln, 2022. roy pascal. the dual voice: free indirect speech and its functioning in the nineteenth-century european novel. manchester university press, manchester, 1977. paul piwek and emiel krahmer. presuppositions in context: constructing bridges. in pierre bonzon, marcos cavalcanti, and rolf nossum, editors, formal aspects of context, volume 20 of applied logic series, pages 85–106. springer, dordrecht, 2000. doi: https://doi.org/10.1007/ 978-94-015-9397-7 6. r core team. r: a language and environment for statistical computing. 2015. url https: //www.r-project.org/. tanya reinhart. pragmatics and linguistics: an analysis of sentence topics. philosophica, 27(1): 53–94, 1981. craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. in j.h. yoon and a. kathol, editors, papers in semantics, pages 91–136. ohio state university, columbus, oh, 1996. craige roberts. topics. in klaus von heusinger, claudia maienborn, and paul portner, editors, semantics: an international handbook of natural language meaning, pages 1908–1933. de gruyter, berlin, 2012. 34 perspective-taking and protagonist prominence christopher saure, stefan hinterwimmer, and anna pia jordan-bertinelli. an experimental investigation of the interaction of narrators’ and protagonists’ perspectival prominence in narrative texts. zeitschrift für sprachwissenschaft, 42(2):341–372, 2023. philippe schlenker. context of thought and context of utterance. a note on free indirect discourse and the historical present. mind & language, 19(3):279–304, 2004. doi: https://doi.org/10.1111/ j.1468-0017.2004.00259.x. yael sharvit. the puzzle of free indirect discourse. linguistics and philosophy, 31(3):353–395, 2008. doi: https://doi.org/10.1007/s10988-008-9039-9. franz k. stanzel. a theory of narrative (translated by charlotte goedsche). cambridge university press, new york, 1984. anita steube. erlebte rede aus linguistischer sicht. zeitschrift für germanistik, 6(4):389–406, 1985. rosemary j. stevenson, rosalind a. crawley, and david kleinman. thematic roles, focus and the representation of events. language and cognitive processes, 9(4):519–548, 1994. doi: https: //doi.org/10.1080/01690969408402130. andreas stokke. protagonist projection. mind & language, 28(2):204–232, 2013. doi: https: //doi.org/10.1111/mila.12016. teun van dijk. text and context. longman, london, 1977. jan van kuppevelt. discourse structure, topicality and questioning. journal of linguistics, 31: 109–147, 1995. klaus von heusinger and petra b. schumacher. discourse prominence: definition and application. journal of pragmatics, 154:117–127, 2019. janyce m. wiebe. recognizing subjective sentences: a computational investigation of narrative text. phd thesis, suny buffalo dept. of computer science, buffalo, ny, 1990. janyce m. wiebe. tracking point of view in narrative. computational linguistics, 20(2):233–287, 1994. jeremy zehr and florian schwarz. penncontroller for internet based experiments (ibex), 2018. sonja zeman. wer spricht? disambiguierungsfaktoren bei der perspektivensetzung im narrativen diskurs. in stefan engelberg, christian fortmann, and irene rapp, editors, redeund gedankenwiedergabe in narrativen strukturen: ambiguitäten und varianz, volume 27 of linguistische berichte sonderheft, pages 221–251. buske, hamburg, 2019. sonja zeman. parameters of narrative perspectivization: the narrator. open library of humanities, 6(2):28, 2020. doi: https://doi.org/10.16995/olh.502. 35 dialogue & discourse 16(3) (2025) 1–7 doi:10.5210/dad.2025.301 embodied conversational systems in human–robot interaction: introduction to the special issue dimitra gkatzia∗ d.gkatzia@napier.ac.uk edinburgh napier university, edinburgh, uk hendrik buschmeier∗ hbuschme@uni-bielefeld.de bielefeld university, bielefeld, germany mary ellen foster maryellen.foster@glasgow.ac.uk university of glasgow, glasgow, uk carl strathearn c.strathearn@napier.ac.uk edinburgh napier university, edinburgh, uk editor: barbara di eugenio, david traum 1. introduction in recent years, conversational systems such as chatbots and virtual assistants have become increasingly popular. the underlying technology has the potential to enhance human–robot interaction (hri; bartneck et al., 2020) and improve its user experience. however, designing and implementing effective conversational systems for hri presents significant challenges that need to be addressed (cf. devillers et al., 2020; marge et al., 2022). this special issue of dialogue & discourse on “embodied conversational systems in human–robot interaction”, brings together researchers and practitioners to explore the opportunities and challenges of developing conversational systems for hri. conversational systems and natural language generation (nlg; reiter and dale, 2000; reiter, 2025) are central to human–robot interaction, enabling natural and intuitive speech-based communication. advances in these areas, as well as in related fields such as multimodal interaction, can make robots more accessible, usable, and engaging in domains such as healthcare, education, services, assistive living, and entertainment. by integrating speech, facial expressions, and other non-verbal cues, conversational systems allow robots to better infer users’ emotions and social signals and to tailor their responses accordingly. this adaptability enables personalized interactions that reflect individuals’ needs, preferences, and characteristics, resulting in more meaningful and natural exchanges. this capability is especially valuable in contexts such as personalized tutoring or explanation-giving (stange et al., 2022), where effective communication depends on sensitivity to the user. conversational systems provide a foundation for these capabilities by enabling natural language interaction, an intuitive and familiar means of communication for humans. human–robot interaction is a complex, interdisciplinary field that requires expertise in robotics, artificial intelligence, (computational) linguistics, psychology and human factors, among others. conversational systems integrate many of these areas as well, representing a challenging and ever evolving area of research that has the potential to advance hri technology. this special issue brings *. equal contribution. ©2025 dimitra gkatzia, hendrik buschmeier, mary ellen foster and carl strathearn this is an open-access article distributed under the terms of a creative commons attribution license (https://creativecommons.org/licenses/by/4.0/). https://doi.org/10.5210/dad.2025.301 https://creativecommons.org/licenses/by/4.0/ gkatzia, buschmeier, foster and strathearn together novel research in dialogue systems designed to enhance or support the interaction with robots. a key objective within the active research domain of hri is to develop robotic agents capable of emulating socially intelligent behavior when interacting with humans. despite the clear relationship between social intelligence and fluent, flexible linguistic interaction, interactive robots have only recently begun to utilize anything beyond a basic dialogue manager and template-based response generation process in practice (van deemter et al., 2005). this means that, thus far, social robot systems have been unable to exploit the flexibility offered by dialogue systems and natural language generation when managing conversations between humans and robots in dynamic environments, or when the conversation needs to adapt to different contexts or multiple target languages. conversely, end-to-end systems based on large language models (llms) enable robots to communicate; however, they must be integrated into the overall robotics architecture in order to take into account the situated, embodied nature of robots (lison and kennington, 2023). 2. building bridges between nlg, hri and dialogue these issues have been explored in a series of workshops organized by the editors of this special issue: the first workshop on natural language generation for human robot interaction (“nlg for hri”; foster et al., 2018) took place at the 2018 international natural language generation conference (inlg 2018); conversely, the second workshop in the series (buschmeier et al., 2020) was meant to take place in march 2020 at the acm/ieee international conference on human–robot interaction (hri 2020), but had to be postponed to inlg 2020 due to the emergence of the covid-19 pandemic; finally, a special session “natural language in human robot interaction” (nlihri) took place at the 23rd annual meeting of the special interest group on discourse and dialogue (sigdial 2022). a wide range of topics were discussed at the workshops, but most of the papers presented dealt with one or more of four central themes: multimodal generation, (visual) grounding and contextual knowledge, robust interaction, and adaptation. robots are embodied agents with the ability to move. depending on their physical configuration, they may be able to move their actuators, sensors or even their entire body. these movements must be planned for action and perception; however, each movement also has the potential to be interpreted as a non-verbal communication behaviour. these non-verbal behaviors need to be planned alongside (or even integrated with) the robot’s non-communicative actions, as well as its verbal acts. this makes questions of multimodal generation central to the intersection of nlg, dialogue research, and human–robot interaction and thus a topic discussed across all three workshops (e.g., cass et al., 2018; bailly and elisei, 2020; gella et al., 2022). robots are situated in real-world environments and may share space with people. therefore, a robot needs to be able to perceive its environment, including objects and people, and ground its language in the environment. it also needs to be able to reason about and refer to the environment using verbal and non-verbal means. (visual) grounding and contextual knowledge is central for language generation in robots and thus discussed across all workshops (e.g., pomarlan et al., 2018; wallbridge et al., 2020; torres-fonseca et al., 2022). although robot actions and human–robot interaction are often much slower than human action and interaction, they share many of the pressures present in human face-to-face communication. these include the timing of turn-taking, miscommunication, conversational failure and repair. this requires a certain level of robustness in the robot’s interaction capabilities, making robust interaction a topic at each workshop (e.g., zarrieß and schlangen, 2018; doğan and leite, 2020; li et al., 2022). 2 embodied conversational systems in hri: introduction to the special issue finally, several contributions across the workshops argued that robots interacting with human conversation partners should adapt to their characteristics (e.g., personality), emotional states (as expressed, e.g., through social signals), and needs (e.g., ritschel and andré, 2018; shenoy and dugan, 2020; fernau et al., 2022). 3. overview of the papers in the special issue continuing the discussion of these central themes of the workshops, this special issue comprises four papers that cover empirical, computational modeling, and engineering perspectives on embodied conversational systems in human–robot interaction. topics that recur throughout the papers include multimodality and conversational behavior, the representation of dialogue state using knowledge graphs, and the integration of language models into conversationally capable robot architectures. the first paper “laughter use by virtual agents increases task success” (ludusan and wagner, 2025) studies the influence of synthetic sound-based nonverbal expressions in human–agent interaction – specifically the use of laughter in dialogue – on participants’ agent social perception and, importantly, task success. it shows that an agent that laughs is rated higher in its social perception and has a higher task success. this is novel evidence for the discussions earlier in the workshops which dealt with the importance of modeling humor and laughter capabilities for interactive robots (ritschel and andré, 2018; ritschel et al., 2020). the paper also shows that being able to technically integrate para-verbal behavior with the verbal behavior of these systems (both on the level of planning as well as on the level of synthesis) is useful and effective (aylett, 2020). two contributions in this issue make use of knowledge graphs to improve the interaction capabilities in interactive robots (a line of research that was already visible in workshop contributions that made use of ontologies to link and ground a robot’s physical and conversational actions, e.g., pomarlan et al., 2018, 2020). the contribution “a modular architecture for creating multimodal embodied agents with an episodic knowledge graph as an explainable and controllable long-term memory” (baier et al., 2025) uses an ‘episodic’ knowledge graph in order to create coherence and continuity across interactions. the article presents a broader architectural framework for interactive agents consisting of two core components: the aforementioned episodic knowledge graph, and a component for the time-aligned management of multimodal signals encountered (and produced) by an agent during an interaction. these core components must be supplemented by elements that interpret and annotate signals, make decisions and generate behavior. the overarching goal of the architectural framework is to combine flexibility with control exerted through modeling higher-level intentions representing an agent’s goals. in contrast to this, the contribution “a graph-to-text approach to knowledge-grounded response generation in human–robot interaction” (walker et al., 2025) proposes and evaluates a conversation model for a robot which represents the dialogue state in a graph-based representation. this knowledge graph combines linguistic with situated and multimodal information and is continuously updated from sensors of the robot as well as other system information, preserving temporal as well as probabilistic aspects. this representation is then used for generating an intermediate textual representation, which forms the basis for the generation of the robot’s conversational actions using large language models. finally, the contribution prior lessons of incremental dialogue and robot action management for the age of language models (kennington et al., 2025), addresses the important topic of incremen3 gkatzia, buschmeier, foster and strathearn tal processing in dialogue (see also dialogue & discourse, vol. 2 no. 1; rieser and schlangen, 2011) and analyses its implications for the “age of lms”. the authors argue that incremental dialogue processing, particularly dialogue management, is essential for human–robot interaction (a point also been made in several of the workshop contributions: zarrieß and schlangen, 2018; bailly and elisei, 2020; li et al., 2022). they argue that this presents challenges for systems in which dialogue capabilities are primarily driven by llms. the contribution introduces incremental dialogue processing, reviews the state of the art in incremental dialogue management, and discusses challenges and requirements. this special issue presents contributions written when llms first emerged as a novel technological development and were swiftly incorporated into robotics and human–robot interaction. the contributions reflect this technological shift by utilizing llms and/or discussing how they can be integrated in light of the well-understood theoretical and engineering challenges at the intersection of conversational systems and human–robot interaction. references matthew p. aylett. mixing speech and semantic free utterances: a challenge for natural language generation. in hendrik buschmeier, mary ellen foster, and dimitra gkatzia, editors, 2nd workshop on natural language generation for human–robot interaction, online, 2020. url https: //purl.org/nlghri2020/aylett. thomas baier, selene báez santamarı́a, and piek vossen. a modular architecture for creating multimodal embodied agents with an episodic knowledge graph as an explainable and controllable long-term memory. dialogue & discourse, 16(3):25–59, 2025. doi:10.5210/dad.2025.303. gérard bailly and frédéric elisei. speech in action: designing challenges that require incremental processing of self and others’ speech and performative gestures. in hendrik buschmeier, mary ellen foster, and dimitra gkatzia, editors, 2nd workshop on natural language generation for human–robot interaction, online, 2020. url https://purl.org/nlghri2020/bailly. christoph bartneck, tony belpaeme, friederike eyssel, takayuki kanda, merel keijsers, and selma šabanović. human-robot interaction: an introduction. cambridge university press, cambridge, uk, 2020. doi:10.1017/9781108676649. hendrik buschmeier, mary ellen foster, and dimitra gkatzia. second workshop on natural language generation for human–robot interaction. in companion of the 2020 acm/ieee international conference on human–robot interaction, pages 646–647, cambride, uk, 2020. association for computing machinery. doi:10.1145/3371382.3374853. aaron g. cass, kristina striegnitz, and nick webb. a farewell to arms: non-verbal communication for non-humanoid robots. in mary ellen foster, hendrik buschmeier, and dimitra gkatzia, editors, proceedings of the workshop on nlg for human–robot interaction, pages 22–26, tilburg, the netherlands, 2018. association for computational linguistics. doi:10.18653/v1/w18-6905. kees van deemter, emiel krahmer, and mariët theune. real versus template-based natural language generation: a false opposition? computational linguistics, 31:15–23, 2005. doi:10.1162/0891201053630291. 4 https://purl.org/nlghri2020/aylett https://purl.org/nlghri2020/aylett https://doi.org/10.5210/dad.2025.303 https://purl.org/nlghri2020/bailly https://doi.org/10.1017/9781108676649 https://doi.org/10.1145/3371382.3374853 https://doi.org/10.18653/v1/w18-6905 https://doi.org/10.1162/0891201053630291 embodied conversational systems in hri: introduction to the special issue laurence devillers, tatsuya kawahara, roger k. moore, and matthias scheutz. spoken language interaction with virtual agents and robots (slivar): towards effective and ethical interaction (dagstuhl seminar 20021). dagstuhl reports, 10(1):1–51, 2020. doi:10.4230/dagrep.10.1.1. fethiye irmak doğan and iolanda leite. open challenges on generating referring expressions for human–robot interaction. in hendrik buschmeier, mary ellen foster, and dimitra gkatzia, editors, 2nd workshop on natural language generation for human–robot interaction, online, 2020. doi:10.48550/arxiv.2104.09193. daniel fernau, stefan hillmann, nils feldhus, tim polzehl, and sebastian möller. towards personality-aware chatbots. in proceedings of the 23rd annual meeting of the special interest group on discourse and dialogue, pages 135–145, edinburgh, uk, 2022. association for computational linguistics. doi:10.18653/v1/2022.sigdial-1.15. mary ellen foster, hendrik buschmeier, and dimitra gkatzia, editors. proceedings of the workshop on nlg for human–robot interaction, tilburg, the netherlands, 2018. association for computational linguistics. doi:10.18653/v1/w18-69. spandana gella, aishwarya padmakumar, patrick lange, and dilek hakkani-tur. dialog acts for task driven embodied agents. in proceedings of the 23rd annual meeting of the special interest group on discourse and dialogue, pages 111–123, edinburgh, uk, 2022. association for computational linguistics. doi:10.18653/v1/2022.sigdial-1.13. casey kennington, pierre lison, and david schlangen. prior lessons of incremental dialogue and robot action management for the age of language models. dialogue & discourse, 16(3):96–130, 2025. doi:10.5210/dad.2025.305. siyan li, ashwin paranjape, and christopher manning. when can i speak? predicting initiation points for spoken dialogue agents. in proceedings of the 23rd annual meeting of the special interest group on discourse and dialogue, pages 217–224, edinburgh, uk, 2022. association for computational linguistics. doi:10.18653/v1/2022.sigdial-1.22. pierre lison and casey kennington. who’s in charge? roles and responsibilities of decision-making components in conversational robots. in workshop on human–robot conversational interaction, stockholm, sweden, 2023. doi:10.48550/arxiv.2303.08470. bogdan ludusan and petra wagner. laughter use by virtual agents increases task success. dialogue & discourse, 16(3):8–24, 2025. doi:10.5210/dad.2025.302. matthew marge, carol espy-wilson, nigel g. ward, abeer alwan, yoav artzi, mohit bansal, gil blankenship, joyce chai, hal daumé, debadeepta dey, mary harper, thomas howard, casey kennington, ivana kruijff-korbayová, dinesh manocha, cynthia matuszek, ross mead, raymond mooney, roger k. moore, mari ostendorf, heather pon-barry, alexander i. rudnicky, matthias scheutz, robert st. amant, tong sun, stefanie tellex, david traum, and zhou yu. spoken language interaction with robots: recommendations for future research. computer speech & language, 71:101255, 2022. doi:10.1016/j.csl.2021.101255. mihai pomarlan, robert porzel, john bateman, and rainer malaka. from sensors to sense: integrated heterogeneous ontologies for natural language generation. in mary ellen foster, hendrik 5 https://doi.org/10.4230/dagrep.10.1.1 https://doi.org/10.48550/arxiv.2104.09193 https://doi.org/10.18653/v1/2022.sigdial-1.15 https://doi.org/10.18653/v1/w18-69 https://doi.org/10.18653/v1/2022.sigdial-1.13 https://doi.org/10.5210/dad.2025.305 https://doi.org/10.18653/v1/2022.sigdial-1.22 https://doi.org/10.48550/arxiv.2303.08470 https://doi.org/10.5210/dad.2025.302 https://doi.org/10.1016/j.csl.2021.101255 gkatzia, buschmeier, foster and strathearn buschmeier, and dimitra gkatzia, editors, proceedings of the workshop on nlg for human– robot interaction, pages 17–21, tilburg, the netherlands, 2018. association for computational linguistics. doi:10.18653/v1/w18-6904. mihai pomarlan, vanja sophie cangalovic, robert porzel, and john bateman. human, we have a problem: what to say when things go wrong. in hendrik buschmeier, mary ellen foster, and dimitra gkatzia, editors, 2nd workshop on natural language generation for human–robot interaction, online, 2020. url https://purl.org/nlghri2020/pomarlan. ehud reiter. natural language generation. springer, cham, switzerland, 2025. doi:10.1007/978-3031-68582-8. ehud reiter and robert dale. building natural language generation systems. cambridge university press, cambridge, uk, 2000. doi:10.1017/cbo9780511519857. hannes rieser and david schlangen. introduction to the special issue on incremental processing in dialogue. dialogue & discourse, 2(1):1–10, 2011. doi:10.5087/dad.2011.001. hannes ritschel and elisabeth andré. shaping a social robot’s humor with natural language generation and socially-aware reinforcement learning. in mary ellen foster, hendrik buschmeier, and dimitra gkatzia, editors, proceedings of the workshop on nlg for human–robot interaction, pages 12–16, tilburg, the netherlands, 2018. association for computational linguistics. doi:10.18653/v1/w18-6903. hannes ritschel, thomas kiderle, klaus weber, and elisabeth andré. multimodal joke presentation for social robots based on natural-language generation and nonverbal behaviors. in hendrik buschmeier, mary ellen foster, and dimitra gkatzia, editors, 2nd workshop on natural language generation for human–robot interaction, online, 2020. url https://purl.org/nlghri2020/ritschel. sudhir shenoy and joanne dugan. challenges and opportunities for nlg in persuasive robotics. in hendrik buschmeier, mary ellen foster, and dimitra gkatzia, editors, 2nd workshop on natural language generation for human–robot interaction, online, 2020. url https://purl.org/ nlghri2020/shenoy. sonja stange, teena hassan, florian schröder, jacqueline konkol, and stefan kopp. self-explaining social robots: an explainable behavior generation architecture for human-robot interaction. frontiers in artificial intelligence, 5(866920):1–19, 2022. doi:10.3389/frai.2022.866920. josue torres-fonseca, catherine henry, and casey kennington. symbol and communicative grounding through object permanence with a mobile robot. in proceedings of the 23rd annual meeting of the special interest group on discourse and dialogue, pages 124–134, edinburgh, uk, 2022. association for computational linguistics. doi:10.18653/v1/2022.sigdial-1.14. nicholas thomas walker, stefan ultes, and pierre lison. a graph-to-text approach to knowledgegrounded response generation in human–robot interaction. dialogue & discourse, 16(3):60–95, 2025. doi:10.5210/dad.2025.304. christopher d. wallbridge, alex smith, manuel giuliani, chris melhuish, tony belpaeme, and séverin lemaignan. towards the effectiveness of ambiguous spatial descriptions in human 6 https://doi.org/10.18653/v1/w18-6904 https://purl.org/nlghri2020/pomarlan https://doi.org/10.1007/978-3-031-68582-8 https://doi.org/10.1007/978-3-031-68582-8 https://doi.org/10.1017/cbo9780511519857 https://doi.org/10.5087/dad.2011.001 https://doi.org/10.18653/v1/w18-6903 https://purl.org/nlghri2020/ritschel https://purl.org/nlghri2020/shenoy https://purl.org/nlghri2020/shenoy https://doi.org/10.3389/frai.2022.866920 https://doi.org/10.18653/v1/2022.sigdial-1.14 https://doi.org/10.5210/dad.2025.304 embodied conversational systems in hri: introduction to the special issue robot interaction. in hendrik buschmeier, mary ellen foster, and dimitra gkatzia, editors, 2nd workshop on natural language generation for human–robot interaction, online, 2020. url https://purl.org/nlghri2020/wallbridge. sina zarrieß and david schlangen. being data-driven is not enough: revisiting interactive instruction giving as a challenge for nlg. in mary ellen foster, hendrik buschmeier, and dimitra gkatzia, editors, proceedings of the workshop on nlg for human–robot interaction, pages 27–31, tilburg, the netherlands, 2018. association for computational linguistics. doi:10.18653/v1/w18-6906. 7 https://purl.org/nlghri2020/wallbridge https://doi.org/10.18653/v1/w18-6906 introduction building bridges between nlg, hri and dialogue overview of the papers in the special issue dialogue & discourse 13(2) (2022) 49–78 doi: 10.5210/dad.2022.202 the effect of domain knowledge on discourse relation inferences: relation marking and interpretation strategies marian marchal marchal@coli.uni-saarland.de language science and technology saarland university, germany merel c.j. scholman m.c.j.scholman@coli.uni-saarland.de language science and technology saarland university, germany vera demberg vera@coli.uni-saarland.de computer science, language science and technology saarland university, germany editor: manfred stede submitted 10/2021; accepted 08/2022; published online 09/2022 abstract it is generally assumed that readers draw on their background knowledge to make inferences about information that is left implicit in the text. however, readers may differ in how much background knowledge they have, which may impact their text understanding. the present study investigates the role of domain knowledge in discourse relation interpretation, in order to examine how readers with high vs. low domain knowledge differ in their discourse relation inferences. we compare interpretations of experts from the field of economics and biomedical sciences in scientific biomedical texts as well as more easily accessible economic texts. the results show that high-knowledge readers from the biomedical domain are better at inferring the correct relation interpretation in biomedical texts compared to low-knowledge readers, but such an effect was not found for the economic domain. the results also suggest that, in the absence of domain knowledge, readers exploit linguistic signals other than connectives to infer the discourse relation, but domain knowledge is sometimes required to exploit these cues. the study provides insight into the impact of domain knowledge on discourse relation inferencing and how readers interpret discourse relations when they lack the required domain knowledge. keywords: discourse relations; relational signals; discourse inferences; domain knowledge 1. introduction to successfully comprehend and learn from a text, readers need to construct a coherent mental representation of the information in the text. this requires readers to understand how the various concepts in a text are related and to integrate the text with background knowledge already available to the reader (see van den broek, 2010). inferring discourse relations is an essential part of establishing coherence in text (sanders et al., 1992). prior studies have suggested that background knowledge supports the inference of discourse relations, assuming that this knowledge is activated to fill in information that is missing in the text (noordman et al., 2015). the role of domain knowledge in interpreting discourse relations is still unclear. earlier work has often focused on the role ©2022 marian marchal, merel c.j. scholman and vera demberg this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). marchal, scholman and demberg of background knowledge on text comprehension or recall (for an overview, see smith et al., 2021), but how discourse relation inferences differ between highand low-knowledge readers has not been investigated systematically. moreover, it is unclear what other factors guide the interpretation of discourse relations for lowknowledge compared to high-knowledge readers. most studies in the field of discourse relations have focused on the effect of textual cues on relational inferences. most notably, studies have shown the impact of connectives on the interpretation process. agreement on explicit discourse relations is higher than on relations in which no connective is present (e.g., demberg et al., 2019; hoek et al., 2021; kishimoto et al., 2018; miltsakaki et al., 2004) and appropriate connectives facilitate the integration of upcoming material, whereas inappropriate connectives disrupt processing (e.g. murray, 1997; canestrelli et al., 2013). more recently, studies have focused on identifying other constructions and words that tend to correlate with certain discourse relations (asr and demberg, 2015; webber, 2013; grisot and blochowiak, 2021). however, these cues are highly ambiguous and likely still need to be supplemented with non-textual information to infer the relation. how these cues interact with domain knowledge has not been taken into consideration. furthermore, it is unclear how the type of inferences that readers make may depend on the knowledge that they have on the domain of the text. the goal of the present study is therefore two-fold. first, the current paper aims to investigate whether domain knowledge leads to more correct interpretations of discourse relations. this will be assessed by eliciting discourse relation interpretations from highand low-knowledge readers and comparing them to a gold label annotation. second, this research sets out to explore how readers infer the discourse relation if they lack the necessary domain knowledge. in the next section, we will first review previous research on the role of domain knowledge in discourse inferences and discuss which factors influence discourse relation interpretation and could help low-knowledge readers to infer discourse relations for which domain knowledge is required. the hypotheses are outlined in section 3, followed by a description of the methodology. the results are presented in section 5. these are subsequently discussed in the final section. 2. background 2.1 the role of domain knowledge in discourse inferences several models of language comprehension suggest that readers exploit their knowledge base about the concepts in the text to create a coherent representation of the text (e.g. construction-integration model, kintsch and van dijk, 1978, landscape model; van den broek, 2010). this knowledge is activated when reading about relevant concepts in the texts, after which the information is retrieved from the long-term memory and can then be integrated with the representation that has been made of the text so far. in addition, reading about these concepts activates additional relevant information in the knowledge base, which can in turn influence predictions about subsequent text (cf. venhuizen et al., 2019; ferreira and chantavarin, 2018). for example, comprehenders adopt general world knowledge in a similar way to linguistic cues to predict event structures within sentences (milburn et al., 2016). with respect to inter-sentential discourse coherence, studies indicate that readers also create expectations based on linguistic cues (e.g. köhne-fuetterer et al., 2021; scholman et al., 2017) as well as general world knowledge (kuperberg et al., 2011). reading comprehension is thus a dynamic process in which bottom-up and top-down processes are combined. if a discourse 50 the effect of domain knowledge and implicitation relation is not expressed linguistically, readers can utilize information from the knowledge base on how the events in the text are related to establish coherence. there are different types of non-linguistic knowledge that readers can have. in the literature, a distinction is sometimes made between background knowledge (i.e. all the knowledge the reader can bring to the text), world knowledge, and domain knowledge (e.g., smith et al., 2021). domain knowledge is a type of background knowledge about a specific area (e.g. apoptosis is natural cell death). in this sense, it could be distinguished from general world knowledge (e.g. the sky is blue), which is considered to be available to almost every reader. it should be noted that we do not assume that if a reader has domain knowledge about the topic of the text, they will know how all concepts in the text are related, nor do we assume that this knowledge is required to infer every relation. some concepts in the text might still be unknown to high-knowledge readers, and some textual relations can also be inferred without knowledge about the domain of the text. however, we do hypothesize that readers with domain knowledge might find it easier to infer the discourse relations in a text from their domain of expertise compared to readers without this specific knowledge, because they are more familiar with the concepts discussed in the text and can rely on an already existing knowledge structure. empirical evidence that readers benefit from domain knowledge in making discourse inferences comes from various studies on the influence of coherence marking on reading comprehension. this line of research has repeatedly shown a ‘reverse cohesion’ effect: in general, low-knowledge readers benefit from texts with high coherence marking, whereas high-knowledge readers show better comprehension after reading a low-cohesive text (mcnamara et al., 1996; o’reilly and mcnamara, 2007; kamalski et al., 2008; mcnamara, 2001). linguistic marking of coherence enables lowknowledge readers to understand how the concepts in a text are related. in the absence of such cues, comprehension will be impaired. for high-knowledge readers, on the other hand, a text with low cohesion induces them to employ their knowledge base to fill in the gaps in the text. connecting the concepts from the text with those in their long-term memory then leads to deeper comprehension (mcnamara et al., 1996). these studies have focused on the role of domain knowledge in text comprehension in general, but do not reveal how this influences the interpretation of discourse relations. examining how lowand high-knowledge readers interpret discourse relations differently can provide more insights into the qualitative differences in text comprehension for these groups of readers. in addition, little is known about strategies that low-knowledge readers may have to comprehend an out-of-domain text. this will be addressed in the current study. 2.2 strategies for inferring discourse relations in addition to discourse connectives and background knowledge, several other factors have been suggested to influence discourse inferences. in cases where readers lack the domain knowledge to infer the discourse relation, and no connective is available to signal the relation, readers might resort to other strategies to establish coherence. more specifically, readers might (i) use non-connective linguistic signals for coherence relations, (ii) rely on cognitive biases for relational inferences, or (iii) process the text more shallowly. how these factors influence discourse relation inferences and how they might guide the interpretations of low-knowledge readers is outlined in more detail below. 51 marchal, scholman and demberg 2.2.1 recognizing discourse relational cues discourse connectives (e.g because, however, or while) are rather strong signals of discourse relations. most relations, however, are not signaled explicitly by connectives (e.g., only 43% of the relations in the penn discourse treebank are explicitly marked by a connective (webber et al., 2016)). for relations that are not signaled by a connective, it is not necessarily the case that they contain no linguistic cue at all. das and taboada (2018) found that more than 80% of the relations in the rst signalling corpus were signaled by means of other cues. non-connective linguistic cues for discourse relations (i.e. words or phrases that tend to co-occur statistically with certain discourse relations) might thus play a role in inferring discourse relations. recent work has focused on identifying non-connective linguistic cues that co-occur frequently with certain discourse relations. for instance, a negation marker is present in more than half of the chosen alternative relations, in which the second relational argument provides an alternative to the situation described in the first argument (see example 1). (1) kevin is not going to italy this year. instead, he is going to visit family in new york. another example of a well-known linguistic cue is tense, which has been found to correlate strongly with different types of temporal relations (grisot and blochowiak, 2021). furthermore, corpus research has shown that the distribution of such cues seems to be different in explicit and implicit relations (cf. sporleder and lascarides, 2008). more specifically, nonconnective linguistic signals for discourse relations appear to be more frequent in implicit relations than in explicit relations (hoek et al., 2018). a case in point is the negation marker in example 1 above: when such a cue is present, chosen alternative relations are more likely to be implicit. in addition, crible (2020) found that non-connective cues are more likely to occur in explicit relations if the connective is ambiguous. this is in line with gricean’s maxim of quantity to not make a message more informative than necessary (grice, 1975) (see asr and demberg, 2012, 2015, for an explanation based on the uniform information density hypothesis). even though studies are increasingly showing the prevalence of non-connective linguistic cues, it remains unclear whether readers are sensitive to such cues. given that the primary function of non-connective cues is to convey propositional content (rather than signaling a discourse relation, as connectives do), they could be considered to be much subtler cues. furthermore, these elements are more ambiguous than discourse connectives in that they can correlate with a large variety of discourse relations. thus, even though such patterns provide cues about the discourse relation, readers might not pick up on them. still, there is some prior literature showing that readers are sensitive to cues other than connectives in processing discourse relations. for example, implicit causality verbs, like admire in (2) below, have been shown to elicit expectations for causal discourse relations (kehler et al., 2008; koornneef and sanders, 2013). (2) laura admires mo, because he won the competition. similarly, crible (2021) found that the processing of concession relations is facilitated by the presence of overt negation in the first segment, whereas result relations are read slower when the first segment contains negation. these findings suggest that comprehenders do draw on nonconnective cues to infer discourse relations. in a study by scholman et al. (2020), quantity expressions (e.g. a few, several) in the preceding context yielded more list continuations from participants in a study completion paradigm than when these signals were not present. their research 52 the effect of domain knowledge and implicitation also showed that participants with much reading experience were more sensitive to these cues. this suggests that not all readers pick up on discourse relational cues equally well. the present study aims to investigate how readers’ sensitivity to discourse relational cues interacts with their domain knowledge. to manipulate the degree to which linguistic signals might be present in the text, both originally implicit and originally explicit relations will be used in the current study.1 for the originally explicit relations, we remove the connective to create implicit versions. we call these instances in which the original connective has been removed implicitated relations (following e.g. hoek et al., 2017, on implicitation in translation). previous studies on discourse relation inferences have used either implicit (scholman and demberg, 2017a; yung et al., 2019) or implicitated relations (sanders et al., 1992), but have not compared readers’ accuracy on the two types of relations. since relations that are implicit likely contain more discourse relational cues than relations that have been implicitated, we expect readers to be better at inferring implicit relations than implicitated relations. we thus predict that agreement on implicitated relations will be lower than on originally implicit relations. note that implicit relations might also be left unmarked because they are easy to infer based on general world knowledge and will therefore yield higher accuracy. likewise, implicitated relations might be more difficult because the meaning changes when the connective is removed. we will return to this issue in the discussion. crucially, we also predict a possible interaction with domain knowledge here. if the text contains relational cues, both highand low-knowledge readers might be able to employ these to infer the discourse relation. the effect of domain knowledge will then be moderated. however, if the amount of linguistic information for the discourse relation is limited, as in the case of implicitated relations, high-knowledge readers can still rely on their domain knowledge to infer the relation. low-knowledge readers, on the other hand, do not have this information at their disposal and will then struggle with inferring the intended relation, leading to lower agreement. the effect of domain knowledge is therefore hypothesized to be larger for implicitated than for implicit relations. 2.2.2 cognitive biases in relation inferences another way in which readers might infer discourse relations in the absence of other information is by relying on cognitive biases towards certain discourse relation inferences. according to the continuity hypothesis (segal et al., 1991; murray, 1997), readers prefer to interpret information in a text as being temporally and causally continuous. more specifically, it suggests that readers tend to relate sentences in an additive, temporal or causal way. similarly, the causality-by-default hypothesis (sanders, 2005) states that readers have a bias to infer causal relations between the segments in a text. several corpus-based and experimental studies have provided evidence for these hypotheses. for example, continuous relations have been shown to be marked less frequently by a connective (asr and demberg, 2012), which is argued to be caused by the fact that these relations will be inferred regardless of coherence marking. in addition, causal relations are processed faster than additive relations (sanders and noordman, 2000), which in turn are processed faster than discontinuous discourse relations, such as contrastive relations (murray, 1997). 1. manipulating the materials by adding or removing these cues was not deemed suitable for the present study, given the relatively limited insights that are currently available regarding the variety of non-connective relational cues and their effects. furthermore, readers with different levels of domain knowledge might make use of different types of signals, but we do not know beforehand what these signals might be. we therefore use natural text to be able to explore what such cues might be in a qualitative analysis of the results. 53 marchal, scholman and demberg with respect to discourse inference strategies in cases where domain knowledge is required to interpret the relation correctly, it can then be hypothesized that readers will default to inferring a causal or another type of continuous discourse relation if no other relation can be inferred. only in cases where readers have the necessary background knowledge to infer the relation, they might not rely on these cognitive biases and infer a less-expected discourse relation. the hypothesis that readers infer expected relations is also supported by a shallow processing account for discourse relations in the absence of domain knowledge. according to graesser et al. (1994), readers with less background knowledge process text less deeply and might even abandon their search for coherence. several experimental studies have shown that readers are indeed less likely to make inferences during reading when they lack background knowledge (noordman et al., 2015, 1992). this suggests that low-knowledge readers process discourse relations more shallowly. scholman (2019) shows that shallow processing might lead to a higher susceptibility for cognitive biases in relation interpretation. in her study, readers interpret instantiation and specification relations less often as being argumentative when being forced to process the relation more deeply (i.e. by first summarizing the text). thus, according to a shallow-processing account, lowknowledge readers who process the text more shallowly might therefore have a stronger preference for continuous and causal discourse relations. 2.2.3 underspecified interpretations finally, low-knowledge readers could abstain from committing to a specific discourse relation, but rather make an approximate assumption about the meaning. for example, readers might infer that there is a negative, or adversative, relation between the segments in (3), but not determine whether it is a concession (i.e., one of the segments raises an expectation that is denied in the other segment) or a contrast (i.e., the two segments present two different concepts). in the concession reading, juan knew that his girlfriend would be satisfied with just a drink, but ordered much more despite that. in the contrast interpretation, juan’s extensive order is compared to the small order of his girlfriend. (3) juan ordered everything on the menu. his girlfriend only wanted something to drink. such underspecified interpretations can arise from two causes. on the one hand, it might be a result of shallow processing, similar to a preference for cognitively expected relations. if readers process a text shallowly, they might be satisfied with only inferring that the relation is negative and not wish to specify it further, as this would require more effort. on the other hand, such underspecified interpretations might also arise from uncertainty about the discourse relation. even if a low-knowledge reader processes the text deeply, they might still remain uncertain about the specific relation sense when they lack the required domain knowledge. for example, readers might not be able to determine whether the relation in (3) above, is a concession or a contrast relation, despite wishing to do so. if they are nevertheless able to infer some features of the discourse relation (e.g. that it is a negative relation), they might still infer such an underspecified relation, rather than committing to a specific relation that could be wrong. participants can express uncertainty about the relation sense through their connectives. for example, the connective ’but’ is underspecified regarding its relational sense; it can be used to express both contrast and concession. a connective like by contrast, on the other hand, is more specific, as it can only be used in contrast relations. similarly, nevertheless specifies the 54 the effect of domain knowledge and implicitation relation for concession. readers who retain underspecified interpretations might therefore prefer to provide ambiguous connectives, such as but. if readers make specific relation interpretations, they will insert more specific connectives, like nevertheless. 3. present study background knowledge has been assumed to help readers to infer discourse relations, but it is still unclear how discourse relation inferences differ between highand low-knowledge readers, both with respect to the quality of the inference as well as the cues that these different types of readers use. in this study, we manipulate domain knowledge by presenting experts from economics and biomedical sciences with texts that either stem from their domain of expertise (e.g. biomedical research papers in the case of biomedical experts) or from the other domain (e.g. biomedical research papers for economists). the biomedical texts included in this study stem from research papers, which were written for experts in the field. these texts are likely difficult to understand for readers without domain knowledge. the economic texts stem from newspaper articles, which were written for a broader audience. the effect of domain knowledge may therefore be less strong in this genre. nevertheless, since the topic of the newspaper texts focuses on a specific domain, these texts may still be easier to understand for readers who are familiar with that domain than those who are not. discourse interpretations were elicited using a connective insertion task (yung et al., 2019) and compared to gold label annotations. to examine the use of non-connective linguistic signals for discourse relations (see section 2.2.1), the relations were either originally implicit or implicitated for the purposes of the current study. the first research question that this study will address is: • do high-knowledge readers make more accurate discourse inferences than low-knowledge readers? if high-knowledge readers employ their knowledge base when inferring how segments in a text are related (cf. noordman et al., 2015), high-knowledge readers are expected to infer the relation correctly more often than low-knowledge readers. secondly, when required to make an inference about a discourse relation, low-knowledge readers might take several approaches to establish coherence in the text. the second aim of the study is therefore to investigate: • what inferences do readers make if their domain knowledge is insufficient to infer the discourse relation? based on the discussion above, we can formulate three hypotheses about what readers will do in the absence of domain knowledge: (a) readers use non-connective linguistic signals to infer the discourse relation. (b) readers resort to default interpretation strategies based on cognitive biases for continuity and causality. (c) readers make less precise interpretations about the discourse relation. 55 marchal, scholman and demberg since implicit relations have been suggested to contain more non-connective linguistic signals than implicitated relations, we hypothesize that these relations will be easier to infer. moreover, we predict that this effect is stronger for low-knowledge readers, as they are hypothesized to be unable to compensate for their lack of domain knowledge in implicitated relations. note that it might also be the case that relations are left implicit for other reasons, for example because they are easier to infer on the basis of general world knowledge. we will therefore also examine qualitative differences in the presence of linguistic signals in items on which highand low-knowledge readers differ. in addition, if low-knowledge readers’ interpretations are guided by their cognitive biases, it is predicted that these readers will infer more continuous and causal discourse relations than highknowledge readers. if the third hypothesis is true, low-knowledge participants are predicted to insert more ambiguous connectives than high-knowledge readers, as they reflect their underspecified interpretation better than specific connectives. these hypotheses are not mutually exclusive. readers might attempt to use linguistic correlates for discourse relations to make the inference, but leave the relation underspecified if these cues are not sufficient to make a precise inference. similarly, if these discourse relational signals are absent, readers may interpret the relation as being continuous, but not further commit their interpretation to a specific type of continuous relation. 4. method 4.1 participants we recruited students and graduates in the field of economics and biomedical sciences on prolific for a prescreening study. more specifically, five hundred workers participated that had registered on prolific that their subject of study was in the field of economics or biomedical sciences2. the prescreening study served to further ensure expertise in one of the two domains and assess familiarity with each domain. participants were presented with a short questionnaire assessing their demographic background and familiarity with both fields. the latter involved questions about their study and work experience in the field. in addition, we assessed participants’ knowledge by asking them to indicate their familiarity with 10 concepts that are specific to the two domains extracted from the two corpora (e.g. volatility, phosphorylation). for each of these concepts, participants could indicate that they either did not know the term (which was coded as 1), had heard of the term while not being able to describe it (coded as 2), or would also be able to provide a description of the term (coded as 3). five filler concepts for each domain consisting of terms that were deemed familiar for non-experts as well (e.g. interest, dna) were included as an attention check. in order to ensure that our study participants were knowledgeable in their own field of expertise, but not in the other field, we selected only those participants that met the following criteria for the final experiment: (a) they were working or studying in one (and only one) of the two fields, (b) they had high familiarity with the terms in their own field of expertise (top 40% compared to all participants), (c) they had low familiarity in the other field (bottom 50%), and (d) they did not consider themselves novices in the field (i.e. they did not rate their own familiarity with the field compared to other people working or studying in the field lower than 3 on a 7 point likert scale). 2. in other words, when registering on prolific, participants had indicated that their field of study was either in economics, accounting and/or finance or in biomedical sciences, genetics, biology, biological sciences and/or biochemistry. 56 the effect of domain knowledge and implicitation participants’ average familiarity with the concepts in the domain they were an expert in was higher than that of the other domain (see table 1). in addition, each individual participant had higher familiarity with the terms from their own domain than with the terms from the other domain. note that the biomedical experts have higher familiarity with the terms from the economics domain than vice versa. we will return to this issue in the discussion. in short, the experts in our study had academic experience in the relevant subject (as shown by their registration on prolific and their responses to our pretest) and considered themselves knowledgeable in the field (as indicated in our pretest) and their expertise was also reflected in their familiarity with the specialized language used in the texts. table 1: mean scores on the concepts by domain and expertise. scores on a scale from 1 (i have not heard of the term) to 3 (i would be able to describe the term). biomedical terms economic terms biomedical experts 2.86 2.09 economics experts 1.62 2.94 in the final experiment, 106 participants, all native speakers of english, took part (age range, 19-47 years; mean age, 24.7 years; 60 female). of these, 89 participants were students; 56 had completed an undergraduate degree or had obtained a higher education level. 4.2 materials ninety-six relations were sourced from the penn discourse treebank 2.0 (pdtb; prasad et al., 2008), a discourse-annotated corpus containing wall street journal texts. only those sections from the pdtb that were classified as news articles were selected. in addition, we only included texts that covered economic or financial topics. an additional set of 96 relations was extracted from the biomedical discourse relation bank (biodrb; prasad et al., 2011). this corpus contains discourse annotations of 24 biomedical research articles from the genia corpus, using an adapted version of the pdtb annotation framework. the latter texts are likely more specialized than the newspaper texts, which are written to be accessible to a broader audience. still, financial newspapers, like biomedical research papers, target a specific group of readers (i.e. people working in the field of economics) and some degree of domain-specific knowledge is presumed by the writers of these texts as well. we will elaborate on this issue in the discussion. different items could come from (different parts of) the same texts, but the items in each corpus came from at least twenty different texts so that writing styles were varied. the set of experimental items contained an equal amount of implicitated (i.e. originally explicit discourse relations from which the connective has been removed) and implicit relations. to balance the items with respect to the cognitive complexity and expectedness of the relation sense, four different relation senses were selected for the purposes of the present study: result, contrast, concession and instantiation. more specifically, we selected contra-expectation as the 57 marchal, scholman and demberg subcategory of concession relations.3 each relation sense occurred equally often in the experiment. only items for which both arguments were single full sentences were included. the context, consisting of one or two full sentences before and after the arguments, was also presented. an example of an implicitated cause item can be found in passage 4. the relational arguments, separated by \\ are presented in boldface here. to the participants, the context was presented in grey and the target sentences in black. (4) convertible debentures – bonds that can later be converted into equity shares – are the most popular instrument this year, though many companies are also selling non-convertible bonds or equity shares. these mega-issues are being propelled by two factors, economic and political. in the past, the socialist policies of the government strictly limited the size of new steel mills, petrochemical plants, car factories and other industrial concerns to conserve resources and restrict the profits businessmen could make \\ industry operated out of small, expensive, highly inefficient industrial units. when mr. gandhi came to power, he ushered in new rules for business. he said industry should build plants on the same scale as those outside india and benefit from economies of scale. for each of the four relation senses and each relation marking, 12 items were extracted from the pdtb and 12 items from the biodrb, resulting in 192 items in total. 4.3 procedure the task was an updated version of the two-step connective insertion task developed by (yung et al., 2019).4 participants were presented with each item one by one and were asked to complete two steps. in the first step, participants were asked to freely insert a connective in the blank that reflects the relation between the arguments best. they could only continue to the next step if they had typed something in the blank and were instructed to type the word nothing if they could not think of a linking phrase connecting the sentences.5 they were then provided with a list of connectives in the second step and asked to select the connective that fits the relation best. the options presented in the second step were based on the insertion in the first step, and were unambiguous alternatives for the relations that can be signaled by the connective in the first step. for example, if but was inserted in the first step, the options in the second step consisted of (among other options) despite this and on the contrary to disambiguate between the concession and contrast relation sense that can be marked by but. if the option inserted in the first step was not present in the connective bank, a default list was presented: therefore, in addition, despite this, in more detail, even though, for example, by contrast, due to, this example illustrates that, in other words. this default list thus contained a target connective for each of the target relations included in the item. 3. within the class of contrast relations, the pdtb2 distinguishes between juxtaposition and opposition. since no such distinction was made in the biodrb, this distinction was disregarded when selecting materials for the present experiment. 4. three adaptations were made to yung et al. (2019)’s task. firstly, participants were always presented with the second step in this experiment, regardless whether the connective they inserted in the first step was unambiguous or not, to discourage the use of only very specific connectives in the first step. secondly, the connective bank and the mapping of ambiguous connectives to the options in the second step was updated based on follow-up experiments. thirdly, the default list was adapted for the purposes of the current study. 5. participants avoided this restriction in 1.2% of all data points by inserting punctuation or a whitespace. 58 the effect of domain knowledge and implicitation the experiment was hosted on lingoturk (pusse et al., 2016) and distributed via prolific. participants first received instructions. they then saw two practice items, after which they received feedback on possible answers for these items. the items were divided across three batches per relation marking, with four items per relation sense per domain. thus, each participant saw 32 experimental items. four additional filler items were included as attention checks. these items were taken from the pdtb and did not require economic domain knowledge. performance for these items was at ceiling in previous experiments. after completing the study, participants were asked to rate the difficulty of the texts on economic and biomedical topics. the study took around 30 minutes to complete and participants were given £3.50 as compensation for their participation. 4.4 data analysis data from participants who provided less than five different types of connectives in the first step (n = 2) or selected that they wanted to insert a different connective in the second step in more than half of the cases (n = 4) were excluded from further analysis. in addition, participants who failed to select a connective that belonged to the same relational class as the gold label (see below) for more than half of the filler items were also removed (n = 6). the final dataset (n = 2,976) contained observations of 48 experts in the domain of biomedical sciences (implicit: 23, implicitated: 25) and 44 experts in the domain of economics (implicit: 24, implicitated: 22). trials for which participants answered that none of the options provided in the second step were suitable, were coded as missing data (n = 179 observations, 6.0%). to determine whether participants had inferred the relation correctly, the connectives in the second step were categorized as signaling eight different relational classes: (1) cause, (2) temporal, (3) contrast, (4) concession, (5) positive expansion (e.g. instantiation), (6) negative expansion (e.g. disjunction), (7) condition, (8) no relation. an overview of which pdtb3 relation senses are included in each relational class can be found in the appendix. we recoded a new variable, correctness, which was 1 when the inserted connective in the second step matched the relational class of the target relation sense (i.e. agreed with the gold standard), and 0 when it signaled a different relational sense. the correctness variable is used as the dependent variable in subsequent analyses, unless stated otherwise. during data exploration, we discovered that performance on the contrast relations was much lower in the pdtb than in the biodrb (11.2% vs. 31.9%) as well as compared to other relation senses in the pdtb (60.0%). for these contrastive pdtb items, participants frequently provided a concessive connective. note that the distinction between contrast and concession is notoriously difficult (see e.g. robaldo and miltsakaki, 2014; zufferey and degand, 2017). in fact, the manual of the updated version of the pdtb2 (pdtb3, webber et al., 2019) states that they addressed this issue in pdtb3 by reclassifying many contrast relations as concession. we compared the labels of our contrast items between pdtb2 and pdtb3 and found that 21 out of 24 contrast relations were relabelled as concession. we therefore decided to use the updated pdtb 3.0 labels as the gold label.6 we will come back to this issue in the discussion. binomial mixed-effects analyses were used to examine the data. corpus, expertise and relation marking, but not relation sense, were deviation coded for ease of interpretation of the model with the pdtb corpus, economic experts and implicit relations at -1 and their counterparts at 1. for relation 6. the label for contra-expectation and instantiation relations also differ between these two versions of the pdtb. these items were all labeled as arg2-as-denier and arg2-as-instance respectively in the pdtb 3.0. 59 marchal, scholman and demberg table 2: confusion matrix of the gold relation senses and the inserted categories (% per relation sense). positive cause expansion concession contrast other result 64.7 17.5 8.7 2.7 6.3 instantiation 21.9 56.2 9.9 5.7 6.4 concession 19.1 15.0 48.8 8.1 9.0 contrast 15.0 25.7 21.2 31.4 6.7 sense, treatment coding was used with concession as the intercept, as this was hypothesized to be one of the most difficult relations to infer. in addition, we were interested in its comparison with contrast relations, due to these relations often being confused. because of convergence issues, the bobyqa optimizer was used with 10,000 iterations. the models were always first constructed with maximal random effect structure. in case of non-convergence, the model was reduced (barr et al., 2013). the random slope for relation sense never converged. unless specified otherwise, the models therefore contained random intercepts for participants and items and random slopes for corpus and expertise. all materials and data can be found online 7. 5. results 5.1 convergence with the gold label on average, the correct relation sense was inferred in 52.1% of the insertions, as shown by convergence of the insertions in the second step with the gold label. although this performance is relatively low, discourse relation classification is a notoriously difficult task and these numbers are comparable with similar studies using crowd-sourcing for discourse relation annotation (cf. rohde et al., 2016; kishimoto et al., 2018; scholman et al., 2022). when the majority label per item is taken (i.e. aggregating responses of all participants per item to obtain a single annotated label), performance is much higher (74.2%). as can be seen in table 2, performance is much higher on some relation senses than on others (see yung et al., 2019; scholman and demberg, 2017a, for similar results). connective insertions for result items were correct in 64.7% of cases, followed by the instantiation relations (56.2%). these two relational classes were often confused, suggesting that participants did not always know whether the relation was causal or not. another possibility is that these relations were ambiguous for these two relation senses, since instantiation relations can often also be causal (scholman and demberg, 2017b). concession (48.8% correct) showed significantly lower accuracy than performance on result relations as shown in a binomial mixed-effects analysis (see table 3 below). the difference with instantiation relations was not significant. concession relations were sometimes confused with result, but also with positive expansion relations. the latter is surprising, since that means that participants neither infer the causal nor the negative relation between the arguments. finally, performance on contrast relations was even lower than on concession relations, with only 31.4% of insertions falling in the same category of the gold label. in many of 7. https://osf.io/gq59w/?view_only=04d85b1d499b4379b3015693d71c6fcc 60 https://osf.io/gq59w/?view_only=04d85b1d499b4379b3015693d71c6fcc the effect of domain knowledge and implicitation figure 1: convergence with gold label per domain and expertise with error bars showing the standard error. these items, a connective that signals positive expansion or concession was inserted. contrast relations have been shown to be difficult to annotate in other studies as well (cf. kishimoto et al., 2018). in addition, contrast and concession relations are known to often be confused with each other (e.g. robaldo and miltsakaki, 2014). given the effect of relation sense on accuracy, relation sense was included as a covariate in all models presented in this paper. 5.2 the effect of domain knowledge on discourse relation inferences as can be seen in figure 1, performance was higher in the pdtb (57.1%) than in the biodrb (47.1%). this effect was confirmed in the model, as shown by a significant main effect of corpus (see table 3). in addition, overall, biologists converged with the gold label significantly more often than economists (54.8% vs. 49.4%). indeed, expertise was also a significant predictor in the regression analysis. the main question that this study aims to answer, however, is whether high-knowledge readers infer the correct relation sense more often than low-knowledge readers and how readers interpret discourse relations in the absence of domain knowledge. as can be seen in figure 1, for each corpus, highest performance was obtained by the experts from that domain. the binomial mixed-effects analysis shows that the interaction between corpus and expertise is significant, suggesting that domain knowledge leads to a higher accuracy on relational inference. to examine this interaction more closely, we performed a subset analysis on the two corpora. experts from the biomedical sciences converged with the gold label on items from the biodrb (53.1%) more often than the economic experts (40.7%). in the biodrb, expertise was indeed a significant predictor of correctness (β = 0.30, se= 0.09, z = 3.47, p < .001). for the economic texts, however, this difference was minimal: biologists converged with the gold label in 56.4% and economists in 57.8% of cases. expertise was not found to significantly predict accuracy in the pdtb texts. the full output of these models, as well as of all models below can be found in the appendix. 61 marchal, scholman and demberg table 3: output of the full model. model specification: correctness ∼ relationsense + corpus*expertise*relationmarking + (1 + corpus | workerid) + (1 + expertise | questionid) estimate std. error z value p value intercept -0.11 0.14 -0.81 .42 relationsense result 0.84 0.20 4.20 <.001 relationsense contrast -0.75 0.26 -2.92 <.01 relationsense instantiation 0.39 0.20 1.94 .05 corpus -0.17 0.08 -2.07 .04 expertise 0.14 0.07 1.97 .05 relationmarking 0.09 0.09 0.98 .33 corpus:expertise 0.17 0.05 3.55 <.001 corpus:relationmarking 0.04 0.08 0.54 .59 expertise:relationmarking -0.01 0.07 -0.18 .86 corpus:expertise:rel...marking -0.09 0.05 -1.92 .06 5.3 interpretation strategies in the absence of domain knowledge 5.3.1 exploiting relation marking we hypothesized that readers use discourse relational cues in the text to infer the relation. more specifically, we assumed that implicit relations contain more of these cues than implicitated relations and are therefore easier to infer. in addition, low-knowledge readers were hypothesized to rely on these cues more than high-knowledge readers and therefore perform better on implicit than on implicitated relations. table 4 shows the mean accuracy per corpus and relation marking by expertise. overall, performance on implicit relations was slightly higher (54.0%) than on implicitated relations (50.5%). however, relation marking was not a significant predictor for convergence with the gold label (see table 3). there was also no three-way interaction between corpus, expertise and relation marking. we thus find no evidence that the implicit relations are easier to infer than implicitated relations, nor that the effect of domain knowledge is different in implicit and implicitated relations. to examine the role of discourse relational cues more closely, we performed a qualitative analysis on the 30 items for which the difference in accuracy between high-knowledge readers and low-knowledge readers was largest and examined the insertions by both groups. this allowed us to distinguish three different types of relations, which are presented below. table 4: mean percentage of correct answers per corpus and relation marking by expertise. biodrb pdtb implicitated implicit implicitated implicit biomedical sciences 52.2 54.1 54.3 58.7 economics 36.2 45.4 58.7 57.0 mean 44.4 49.9 56.4 57.9 62 the effect of domain knowledge and implicitation relations without linguistic cues require domain knowledge for some items, the relation could only be inferred using domain knowledge. for instance, in passage (5), a reader needs to know what ‘treg activities’ are like in murine systems in order to know whether a reduction in human systems is similar or not. however, no linguistic cues are present to signal this relation. as a result, lowknowledge readers often interpreted this item as a cause relation, instead of concession. (5) in human infectious, neoplastic, and autoimmune diseases, treg activities often mirror those in murine systems numbers of treg are reportedly reduced in human autoimmune diseases, (...) (biodrb:concession:implicit) relational cues allow relational inferences in the absence of domain knowledge a number of items on which experts and non-experts diverged contained non-connective cues that could help readers to infer the correct relation, such as hyponyms for instantiation relations and antonyms for contrast relations. more specifically, the majority of the fourteen instantiation and contrast items that yielded a large difference between experts and non-experts contained such a cue. to illustrate these cues, consider (6) and (7), which yielded high accuracy from both highand low-knowledge readers. the relational cues in these items are signalled linguistically by repeating words (e.g. magazine) or based on general world knowledge (left vs. right). this allows readers to infer relations even in the absence of domain knowledge. (6) other magazine publishing companies have been moving in the same direction the new york times co.’s magazine group earlier this year began offering advertisers extensive merchandising services built around buying ad pages in its golf digest magazine. (pdtb:instantiation:implicit) (7) the core biopsy of the left breast revealed infiltrating ductal carcinoma in 2 of 5 core fragments; high nuclear grade, with no lymphatic invasion seen the core biopsy of the right breast demonstrated benign pathology, specifically, fibrosis with focal ductal epithelial hyperplasia. (biodrb:contrast:implicit) relational cues sometimes require domain knowledge in the items where there was a large difference between experts and non-experts, low-knowledge readers did not always pick up on these cues. the reason for this is that domain knowledge was often required to exploit the cue. this was especially the case for the instantiation relations. in about half of the cases in which a hyponym was present, this cue could only be exploited with domain knowledge. for example, in (8) below, the reader needs to know that orthologous genes are genes in different species that have a similar descent. the second argument provides a specific example of this, but if a reader does not have the required domain knowledge, they will likely also not understand that these genes are instances of orthologous genes. (8) in particular, we assumed that the transcriptional regulation is conserved for orthologous genes the mouse gene myh1 and the human gene myh1 are assumed to share expression patterns and to share important cis-regulatory sequences. (biodrb:instantiation:implicit) 63 marchal, scholman and demberg interestingly, the largest difference between experts and non-experts in convergence with the gold label in the full dataset can be found in implicitated instantiation relations in the biodrb. experts performed 30 percentage points higher than non-experts in this condition (see table 6 in the appendix). the implicit instantiation items in the biodrb and implicitated instantiation items in the pdtb also yielded higher accuracy for experts than for novices. this suggests that cues for instantiation relations are more easily exploited by experts. in a post-hoc analysis, we therefore examined whether the effect of domain knowledge was different per relation sense. adding the three-way interaction between relation sense, corpus and expertise did not significantly improve model fit when compared to the same model without this interaction. since examining differences between the relation senses was not the purpose of the present study and power for finding such a three-way interaction effect with the current study design is likely to be low, further quantitative research is necessary to examine the effect of domain knowledge on different relation senses and different relational cues. the present qualitative analysis provides directions for future research. furthermore, it is interesting to point out that low-knowledge readers do not always exploit relational cues that do not require domain knowledge. more specifically, the three antonyms in the contrast relations that were more challenging for low-knowledge readers could also be detected with general world knowledge, contrasting concepts that are accessible for low-knowledge readers as well (see (10) for an example). in addition, we found instances of hyponyms in our qualitative analysis that do not require specific domain knowledge to infer the instantiation relations, but were nevertheless not detected by low-knowledge readers, as in (9). (9) more recently, several groups have demonstrated the feasibility of hybridizing metabolically labeled mrnas directly from nuclear run-on (nro) reactions to nylon filter microarrays in order to investigate nascent transcripts schuhmacher et al. used a b cell line carrying a conditional, tetracycline-regulated myc gene, and found that myc induction resulted in only a small overlap in regulated mrnas at 4 hours post-induction when comparing polya mrna and nro rna on microarrays. (biodrb:instantiation:implicitated) (10) e2 inhibits apoptosis in different cell types (cardiac myocytes and others) androgens have been found to induce apoptosis. (biodrb:contrast:implicitated) to sum up, non-connective cues seem to play a role in discourse relation inferences, although we do not find evidence that the presence of these cues (or the extent to which they are used to infer the discourse relation) depends on whether or not the relation is marked. in addition, the qualitative analysis shows that adopting these cues sometimes requires domain knowledge. however, even if domain knowledge is not required, low-knowledge readers do not always adopt these cues. 5.3.2 cognitive bias for continuity and causality a second hypothesis was that readers might be guided by cognitive biases for causality and continuity in case their background knowledge was insufficient to determine the relation sense. to examine whether low-knowledge readers resorted to default interpretation strategies, we coded the connective insertions for whether they were signals of continuous relations (cause, positive expansion, 64 the effect of domain knowledge and implicitation table 5: mean effect (sd) of corpus and expertise on the insertion of a causal or continuous connective. estimate std. error z value p value corpus 0.28 (0.04) 0.14 (0.01) 2.02 (0.32) .06 (.06) expertise 0.13 (0.04) 0.10 (0.01) 1.27 (0.41) .24 (.16) corpus:expertise 0.01 (0.04) 0.09 (0.00) 0.06 (0.44) .72 (.18) temporal, condition) or not (contrast, concession, negative expansion). we only included incorrect insertions in this analysis, because correct continuous or discontinuous interpretations cannot be attributed with high certainty to a cognitive bias for continuity and causality, as they are likely guided by the true sense of the relation. in addition, we sampled an equal amount of insertions per relation sense from each corpus and domain of expertise. this was done to ensure an equal balance between relation senses, since certain relation senses might yield more default interpretations than other relation senses (e.g. incorrect interpretations of contrast relations are more likely to be concession than result). the sampling was repeated 100 times to examine whether the effects in the binomial mixed-effects analysis were stable. the results are presented in table 5. no main effect of corpus, domain of expertise nor of the interaction between corpus and expertise was found. there is thus no evidence that readers have a bias to infer a causal or continuous relation if they lack the domain knowledge to interpret the relation correctly. 5.3.3 underspecified interpretations if readers are not certain about the discourse relation between two arguments, they might resort to making an underspecified inference, rather than committing to a specific interpretation. in the paradigm used in this experiment, this would mean that participants insert more ambiguous connectives in the first step when they have little knowledge about the domain of the text, compared to when they are experts in that domain. the connective insertions in the first step were therefore annotated as indicating relations from one vs. multiple relational classes. the most frequent ambiguous first step insertions were however (11.8%), and (6.5%) and but (5.9%). in addition, participants typed nothing in 3.4% of cases, which indicated that they could not come up with a linking phrase connecting the sentences. for example (6.4%), therefore (5.1%) and because (3.3%) were the most frequent specific connectives. a binomial mixed-effects logistic regression analysis showed an interaction between corpus and expertise (β = -0.14, se = 0.05, z = -3.08, p < .01). this effect of domain knowledge on the insertion of ambiguous connectives is visualized in figure 2. splitting up the data by corpus revealed that, compared to experts from the domain of biomedical sciences, economic experts inserted significantly fewer specific connectives in the biodrb (β = -0.29, se = 0.09, z = -3.03, p < .01), but not in the pdtb. besides inferring more incorrect relation types, low-knowledge readers thus also leave the relation underspecified by inserting ambiguous connectives in the first step, when reading the biodrb. 65 marchal, scholman and demberg figure 2: proportion of ambiguous insertions in the first step with error bars showing standard error. 6. discussion and conclusion background knowledge has often been assumed to play a role in correctly interpreting discourse relations, but this has never been investigated experimentally. the current study filled this gap by assessing discourse relation interpretations of highand low-knowledge readers. we aimed to examine whether domain knowledge contributes to inferring the correct discourse relation, as well as which factors guide discourse relation interpretation in the absence of connectives and domain knowledge. the first main finding of this research is that high-knowledge readers were better at inferring the discourse relation, as measured by convergence with the gold label, than low-knowledge readers. thus, domain knowledge can, in some instances, facilitate establishing coherence and readers are able to employ their knowledge base to interpret the relation correctly. however, this effect was modulated by the corpus from which the text was taken: the effect of expertise was significant for the items from the biodrb, but not for the pdtb (see section 6.1). in addition, we identified non-connective linguistic signals for discourse relations, showing that domain knowledge influences how readers adopt these cues (see section 6.2). 6.1 text genre and the influence of domain knowledge one possible reason for why there was only an effect of domain knowledge for texts from the biodrb and not the pdtb is the difference between these specific genres. even though economics newspaper texts are targeted at readers with a specific interest in economics, they are intended for a broader audience with various levels of expertise. research papers, on the other hand, are often not accessible to a general audience. instead, they specifically target experts in that domain. they contain more specialized vocabulary and focus on topics that only a limited amount of people are familiar with.8 also note that the two texts differ in that they are written by journalists, who are not experts themselves, versus researchers. discourse relations in the biomedical texts therefore likely required more domain knowledge than those in the economics newspaper texts. 8. this was also confirmed in a post-hoc analysis of the perplexity of the items using a generic language model (gpt (radford et al., 2018), a transformer model that is trained on a 1b word book corpus). the perplexity of the pdtb items (89.0) is lower than of the biodrb items (101.4). 66 the effect of domain knowledge and implicitation another explanation for this pattern could be related to the level of expertise of our participants. we recruited the participants via a crowd-sourcing platform, but their expertise was assessed in various ways (among others their subject of study as indicated on prolific and their familiarity with specialized terms as determined during prescreening). this ensured that they were indeed high vs. low-knowledge readers with respect to the texts presented in this study. we note here again that experts were not expected to be familiar with all the information in the text. domain knowledge was hypothesized to help in interpreting discourse relations correctly, because text processing is facilitated by an existing knowledge structure. this knowledge base does not need to be exhaustive, as the information in the text fills gaps in existing knowledge. still, the experts from the domain of biomedical sciences seemed to know more about economics than vice versa, as measured by their self-rated familiarity with specialized terms from texts from that domain. they therefore might have also been able to rely on their background knowledge of some economic topics, when interpreting the relations. the finding that the effect of domain knowledge is smaller for the items from the pdtb can therefore not be considered surprising. this interaction also raises the question of what constitutes expertise. experts were assumed to be more knowledgeable with respect to the topic of the text. this knowledge would have been gained through reading texts typical to the domain. in the case of the biomedical experts, this would more likely be research papers than newspapers; in the case of economics experts, this would more likely be economic newspapers than research papers. in our post-test questionnaire, biomedical experts indeed indicated that they read research papers more often than economists (mean 3.57 vs. 2.59 on a 1-5 likert scale ranging from ‘never’ to ‘daily’). this difference was even more distinct for biomedical research papers (3.41 vs. 1.30). economics experts, on the other hand, read newspapers (and specifically business newspapers) more often than biomedic experts (3.5 vs. 2.35 for newspapers in general; 3.09 vs. 1.33 for business newspapers). domain knowledge might thus consist not only of topic knowledge, but also of text genre familiarity. such familiarity might help readers to infer the discourse relations in that genre. for example, newspaper texts are characterized by a so-called inverted pyramid scheme, where the first paragraph is followed by an elaboration in the subsequent paragraphs (das et al., 2018). methodology sections of research papers often contain many temporal relations (bachand et al., 2014). even readers who are not familiar with the domain of the text (e.g. psycholinguistic researchers when reading biomedical research papers) might use genre familiarity with the text structure to infer the discourse relation. a future line of research could attempt to further tease apart the influence of topic knowledge and genre familiarity on the impact of domain knowledge, and how these two factors separately contribute to inferring discourse relations. 6.2 discourse relational cues besides examining the role of domain knowledge in inferring the correct discourse relation, we also set out to explore how readers infer discourse relations in the absence of domain knowledge. the first prediction was that non-connective linguistic cues for discourse relations would be used. we therefore varied the presence of a connective in the original text. discourse relational cues were hypothesized to be more frequent in implicit than in (originally explicit) implicitated relations, which is why we expected implicit relations to be easier to infer than implicitated relations. however, no effect of the presence of a connective in the original text was found. in addition, highand lowknowledge readers were not affected differently by whether the relation was originally marked. we 67 marchal, scholman and demberg can therefore not confirm the hypothesis that discourse relational cues in implicit relations facilitate discourse relational inferences. the by-item analysis revealed that some discourse relations contained cues for the relation, even in the absence of a connective. for example, antonyms were present in contrast relations and hypernyms in instantiation relations. low-knowledge readers sometimes successfully retrieved the relation when such cues were present. however, they were not always sensitive to these cues. for example, a cue such as hypernyms did not always help low-knowledge readers to infer an instantiation relation, and nor did antonyms in contrast relations. instead, readers strongly diverged in the interpretations of these items. there are several possible explanations for these findings. first of all, non-connective linguistic signals for discourse relations are highly ambiguous, with many signaling a large variety of discourse relations. the cue might then exclude some possible relation interpretations, but not provide only one single likely interpretation. as a result, different readers might interpret the cue differently, diminishing the facilitative effect of additional relational cues in implicit relations. secondly, signals for discourse relations (other than connectives) may require domain knowledge to interpret them. to illustrate, antonyms often occur in contrast relations, but in order to know that two concepts are opposite, the reader should know what the words mean. this could explain why low-knowledge readers do not always pick up on these cues. another explanation for the diverging interpretations of items containing non-connective linguistic signals might lie outside the scope of the text itself and be influenced by characteristics of the reader. scholman et al. (2020) show that some readers are more sensitive to contextual list signals than others. more specifically, participants in their study who had more reading experience (as measured by an author recognition test), picked up on these cues more than participants who were less experienced readers. only some readers might therefore have been able to employ these signals in inferring the relation, leading to differences in how relations containing such cues are interpreted. with respect to domain knowledge, high-knowledge readers had access to two strategies in interpreting the relation: non-connective linguistic signals and their knowledge base. the high-knowledge readers who were not sensitive to these non-connective discourse relational cues could then use their domain knowledge to infer the relation, whereas low-knowledge readers would not be able to interpret the relation if they did not detect these signals. more research is needed on what these non-connective linguistic cues consist of as well as the extent to which readers are sensitive to them. for example, is a hypernym a reliable cue for instantiation relations? in addition, the examples provided in this study focused on the two relational arguments, but discourse structure might also guide relational inferences, for example when the arguments are part of a longer list structure. furthermore, even though this study was conducted in english, we expect this effect to replicate in other languages. discourse relational inferences are part of high-level discourse processes and the facilitative effect of connectives has also been replicated in various languages (kamalski et al., 2008; blumenthal-dramé, 2021; lyu et al., 2020). however, which discourse relational cues are present and the extent to which readers rely on this information may differ between languages (cf. blumenthal-dramé, 2021; schwab and liu, 2020). moreover, further research should examine readers’ sensitivity to these cues. the present study did not find a difference between implicit and implicitated relations, even though the latter has been argued to contain more discourse relational cues. do readers notice non-connective signals for discourse relations? or do they only adopt this information if they are experts on the domain of the text (as suggested in the present study) or have much reading experience (scholman 68 the effect of domain knowledge and implicitation et al., 2020)? such research would provide more insight in how readers establish coherence in a text. 6.3 inferences in the absence of domain knowledge the present study also set out to investigate what low-knowledge readers do when they lack the domain knowledge that is required to infer the correct discourse relation. apart from using nonconnective linguistic signals, we hypothesized that participants might resort to a default interpretation strategy and have a preference for causal and continuous relation interpretations in cases in which the relation was not inferred correctly. however, low-knowledge readers did not insert causal and continuous connectives in incorrect items to a greater extent than high-knowledge readers. we thus did not find evidence that domain knowledge influences readers’ cognitive biases for causality and continuity. in addition, we predicted that low-knowledge readers would prefer to leave the discourse relation underspecified. rather than committing to a certain interpretations that might be incorrect, readers were hypothesized to make underspecified discourse interpretations and therefore provide connectives reflecting this underspecification. we found some evidence for this hypothesis, since low-knowledge readers inserted more ambiguous connectives in the first step than high-knowledge readers. low-knowledge readers thus seem to avoid making a specific relation interpretation. however, it remains unclear what the reason for these underspecified interpretations is. on the one hand, it is possible that low-knowledge readers were unable to specify the relation further. on the other hand, low-knowledge readers might have processed the text less deeply and therefore not committed to a specific relation because they did not wish to do so. future research could examine whether low-knowledge readers perform better when they are forced to process the text more deeply (cf. scholman, 2019) to disentangle these two factors. 6.4 limitations finally, we note some limitations of the present research. firstly, the study aimed to balance the items among the different discourse relation senses, since different relation senses were hypothesized to yield differences in accuracy and interpretation biases. we did indeed find that cause and instantiation relations were easier to infer for participants than concession relations and that instantiation and concession relations were often interpreted as being causal. many of the initially selected contrast relations had been annotated as concession in the pdtb 3.0. the lower performance on this relation sense could therefore partly be attributed to the disagreement about the gold label, since a concession interpretation might also have been possible. however, since relation sense was included as a covariate in the analysis, this does not affect the conclusions about the role of domain knowledge. another limitation is our manipulation of relation marking. it is possible that the relation might have become impossible to identify or has changed by removing the connective. in the first case, we would find floor effects on the implicitated relations, even for the high-knowledge readers. overall, there were ten (out of 190) items for which none of the high-knowledge readers converged with the gold label. however, these were equally distributed over the implicit and implicitated condition. this suggests that the original relation could still be retrieved, even when the connective had been removed, also in the implicitated condition. in the case of multiple interpretations, convergence to the gold label is not reliable anymore. to account for the problem of multiple interpretations, we 69 marchal, scholman and demberg examined those implicitated relation items where several participants agreed on the same non-gold relation interpretation and assessed whether this interpretation was also possible. including these alternative answers as correct still revealed the same pattern as above: high-knowledge readers interpreted the relation correctly more often than low-knowledge readers. nevertheless, a study manipulating discourse relational cues specifically would provide further insight on this matter. furthermore, despite carefully selecting our participants, we cannot be sure that they were indeed as knowledgeable as they said they were. nevertheless, we found a clear effect of domain knowledge in the biodrb, suggesting that the biomedical experts were indeed more familiar in this domain than the economic experts. there is no reason to believe that the experts from the field of economics would be less knowledgeable than the participants from the biomedical domain. 6.5 conclusion to conclude, the current research provides more insight into the role of domain knowledge in discourse processing by examining discourse relation interpretations. previous work has mainly focused on the influence of domain knowledge on text comprehension and recall (e.g. mcnamara et al., 1996; smith et al., 2021) or on whether or not discourse inferences are made in the absence of domain knowledge (e.g. noordman and vonk, 1998), showing that low-knowledge readers benefit more from coherence marking than high-knowledge readers and are less likely to make relational inferences during reading. however, these studies did not address how the discourse is interpreted differently by highand low-knowledge readers. the present study shows that readers are able to interpret discourse relations correctly, even if they have little knowledge about the domain of the text. still, high-knowledge readers make more correct (and more specific) discourse relation interpretations. this effect was established in biomedical research papers, a text type that targets a specialist audience, but not in economic newspapers, possibly because the genre is aimed to be accessible for both experts and novices in the field. moreover, we found that readers adopt linguistic cues for inferring discourse relations, although this did not interact with the presence of a connective in the original text. a text without discourse connectives is therefore not necessarily detrimental for low-knowledge readers (cf. mcnamara et al., 1996), as they can also establish coherence with other discourse cues. still, these cues might be more challenging to low-knowledge readers as in some cases domain knowledge is required to detect them. funding this work was supported by the german research foundation (dfg) under grant sfb 1102 (“information density and linguistic encoding”). the second and third author were supported by the european research council under grant 948878 (“individualized interaction in discourse”). references fatemeh torabi asr and vera demberg. implicitness of discourse relations. in proceedings of the international conference on computational linguistics (coling), pages 2669–2684, mumbai, india, 2012. https://aclanthology.org/c12-1163.pdf. fatemeh torabi asr and vera demberg. uniform information density at the level of discourse relations: negation markers and discourse connective omission. in proceedings of the inter70 https://aclanthology.org/c12-1163.pdf the effect of domain knowledge and implicitation national conference on computational semantics (iwcs), pages 118–128, london, uk, 2015. https://aclanthology.org/w15-0117.pdf. félix-hervé bachand, elnaz davoodi, and leila kosseim. an investigation on the influence of genres and textual organisation on the use of discourse relations. in international conference on intelligent text processing and computational linguistics, pages 454–468. springer, 2014. dale j. barr, roger levy, christoph scheepers, and harry j. tily. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3): 255–278, 2013. doi:10.1016/j.jml.2012.11.001. alice blumenthal-dramé. the online processing of causal and concessive relations: comparing native speakers of english and german. discourse processes, 58(7):642–661, 2021. doi:10.1080/0163853x.2020.1855693. anneloes r canestrelli, willem m mak, and ted j. m. sanders. causal connectives in discourse processing: how differences in subjectivity are reflected in eye movements. language and cognitive processes, 28(9):1394–1413, 2013. doi:10.1080/01690965.2012.685885. ludivine crible. weak and strong discourse markers in speech, chat, and writing: do signals compensate for ambiguity in explicit relations? discourse processes, 57(9):793–807, 2020. doi:10.1080/0163853x.2020.1786778. ludivine crible. negation cancels discourse-level processing differences: evidence from reading times in concession and result relations. journal of psycholinguistic research, 50(6):1283–1308, 2021. doi:10.1007/s10936-021-09802-2. debopam das and maite taboada. rst signalling corpus: a corpus of signals of coherence relations. language resources and evaluation, 52(1):149–184, 2018. doi:10.1007/s10579-0179383-x. debopam das, tatjana scheffler, peter bourgonje, and manfred stede. constructing a lexicon of english discourse connectives. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 360–365, 2018. vera demberg, fatemeh torabi asr, and merel c. j. scholman. how compatible are our discourse annotations? insights from mapping rst-dt and pdtb annotations. dialogue & discourse, 10 (1):877–135, 2019. doi:10.5087/dad.2019.104. fernanda ferreira and suphasiree chantavarin. integration and prediction in language processing: a synthesis of old and new. current directions in psychological science, 27(6):443–448, 2018. doi:10.1177/0963721418794491. arthur c. graesser, murray singer, and tom trabasso. constructing inferences during narrative text comprehension. psychological review, 101(3):371–395, 1994. doi:10.1037/0033295x.101.3.371. herbert p. grice. logic and conversation. in peter cole and jerry l. morgan, editors, speech acts, pages 41–58. brill, leiden, 1975. doi:10.1163/9789004368811 003. 71 https://aclanthology.org/w15-0117.pdf https://doi.org/10.1016/j.jml.2012.11.001 https://doi.org/10.1080/0163853x.2020.1855693 https://doi.org/10.1080/01690965.2012.685885 https://doi.org/10.1080/0163853x.2020.1786778 https://doi.org/10.1007/s10936-021-09802-2 https://doi.org/10.1007/s10579-017-9383-x https://doi.org/10.1007/s10579-017-9383-x https://doi.org/10.5087/dad.2019.104 https://doi.org/10.1177/0963721418794491 https://doi.org/10.1037/0033-295x.101.3.371 https://doi.org/10.1037/0033-295x.101.3.371 https://doi.org/10.1163/9789004368811\unhbox \voidb@x \kern .06em\vbox {\hrule width.3em}003 marchal, scholman and demberg cristina grisot and joanna blochowiak. temporal relations at the sentence and text genre level: the role of linguistic cueing and non-linguistic biases—an annotation study of a bilingual corpus. corpus pragmatics, 5:379–419, 2021. doi:10.1007/s41701-021-00104-5. jet hoek, sandrine zufferey, jacqueline evers-vermeul, and ted j. m. sanders. cognitive complexity and the linguistic marking of coherence relations: a parallel corpus study. journal of pragmatics, 121:113–131, 2017. doi:10.1016/j.pragma.2017.10.010. jet hoek, sandrine zufferey, jacqueline evers-vermeul, and ted j. m. sanders. the linguistic marking of coherence relations: interactions between connectives and segment-internal elements. pragmatics & cognition, 25(2):276–309, 2018. doi:10.1075/pc.18016.hoe. jet hoek, merel c. j. scholman, and ted j. m. sanders. is there less annotator agreement when the discourse relation is underspecified? in proceedings of the workshop in integrating perspectives in discourse annotation (discann), pages 1–5, tübingen, germany, 2021. judith kamalski, ted j. m. sanders, and leo lentz. coherence marking, prior knowledge, and comprehension of informative and persuasive texts: sorting things out. discourse processes, 45 (4-5):323–345, 2008. doi:10.1080/01638530802145486. andrew kehler, laura kertz, hannah rohde, and jeffrey l. elman. coherence and coreference revisited. journal of semantics, 25(1):1–44, 2008. doi:10.1093/jos/ffm018. walter kintsch and teun a. van dijk. toward a model of text comprehension and production. psychological review, 85(5):363–394, 1978. doi:10.1037/0033-295x.85.5.363. yudai kishimoto, shinnosuke sawada, yugo murawaki, daisuke kawahara, and sadao kurohashi. improving crowdsourcing-based annotation of japanese discourse relations. in proceedings of the eleventh international conference on language resources and evaluation (lrec 2018), 2018. judith köhne-fuetterer, heiner drenhaus, francesca delogu, and vera demberg. the online processing of causal and concessive discourse connectives. linguistics, 59(2):417–448, 2021. doi:10.1515/ling-2021-0011. arnout w. koornneef and t. j. m. sanders. establishing coherence relations in discourse: the influence of implicit causality and connectives on pronoun resolution. language and cognitive processes, 28:1169–1206, 2013. https://www.tandfonline.com/doi/full/ 10.1080/01690965.2012.699076. gina r. kuperberg, martin paczynski, and tali ditman. establishing causal coherence across sentences: an erp study. journal of cognitive neuroscience, 23(5):1230–1246, 2011. doi:10.1162/jocn.2010.21452. siqi lyu, jung-yueh tu, and chien-jer charles lin. processing plausibility in concessive and causal relations: evidence from self-paced reading and eye-tracking. discourse processes, 57 (4):320–342, 2020. danielle s. mcnamara. reading both high-coherence and low-coherence texts: effects of text sequence and prior knowledge. canadian journal of experimental psychology/revue canadienne de psychologie expérimentale, 55(1):51–62, 2001. doi:10.1037/h0087352. 72 https://doi.org/10.1007/s41701-021-00104-5 https://doi.org/10.1016/j.pragma.2017.10.010 https://doi.org/10.1075/pc.18016.hoe https://doi.org/10.1080/01638530802145486 https://doi.org/10.1093/jos/ffm018 https://doi.org/10.1037/0033-295x.85.5.363 https://doi.org/10.1515/ling-2021-0011 https://www.tandfonline.com/doi/full/10.1080/01690965.2012.699076 https://www.tandfonline.com/doi/full/10.1080/01690965.2012.699076 https://doi.org/10.1162/jocn.2010.21452 https://doi.org/10.1037/h0087352 the effect of domain knowledge and implicitation danielle s. mcnamara, eileen kintsch, nancy butler songer, and walter kintsch. are good texts always better? interactions of text coherence, background knowledge, and levels of understanding in learning from text. cognition and instruction, 14(1):1–43, 1996. doi:10.1207/s1532690xci1401 1. evelyn milburn, tessa warren, and michael walsh dickey. world knowledge affects prediction as quickly as selectional restrictions: evidence from the visual world paradigm. language, cognition and neuroscience, 31(4):536–548, 2016. doi:10.1080/23273798.2015.1117117. eleni miltsakaki, aravind joshi, rashmi prasad, and bonnie webber. annotating discourse connectives and their arguments. in proceedings of the workshop frontiers in corpus annotation at hlt-naacl 2004, pages 9–16, boston, ma, usa, 2004. john d. murray. connectives and narrative text: the role of continuity. memory & cognition, 25 (2):227–236, 1997. doi:10.3758/bf03201114. leo g. m. noordman and wietske vonk. memory-based processing in understanding causal information. discourse processes, 26(2-3):191–212, 1998. doi:10.1080/01638539809545044. leo g m noordman, wietske vonk, and henk j kempff. causal inferences during the reading of expository texts. journal of memory and language, 31(5):573–590, 1992. https://www. sciencedirect.com/science/article/pii/0749596x9290029w. leo g m noordman, wietske vonk, reinier cozijn, and stefan frank. causal inferences and world knowledge. in inferences during reading, pages 260–289. cambridge university press, 2015. tenaha o’reilly and danielle s. mcnamara. reversing the reverse cohesion effect: good texts can be better for strategic, high-knowledge readers. discourse processes, 43(2):121–152, 2007. doi:10.1080/01638530709336895. florian pusse, asad sayeed, and vera demberg. lingoturk: managing crowdsourced tasks for psycholinguistics. in proceedings of the north american chapter of the association for computational linguistics: human language technologies (naacl-hlt), pages 57–61, san diego, ca, 2016. alec radford, karthik narasimhan, tim salimans, and ilya sutskever. improving language understanding by generative pre-training, 2018. url https://cdn.openai.com/ research-covers/language-unsupervised/language_understanding_ paper.pdf. livio robaldo and eleni miltsakaki. corpus-driven semantics of concession: where do expectations come from? dialogue & discourse, 5(1):1–36, 2014. doi:10.5087/dad.2014.101. hannah rohde, anna dickinson, nathan schneider, christopher clark, annie louis, and bonnie webber. filling in the blanks in understanding discourse adverbials: consistency, conflict, and context-dependence in a crowdsourced elicitation task. in proceedings of the 10th linguistic annotation workshop (law x), pages 49–58, berlin, germany, 2016. 73 https://doi.org/10.1207/s1532690xci1401\unhbox \voidb@x \kern .06em\vbox {\hrule width.3em}1 https://doi.org/10.1080/23273798.2015.1117117 https://doi.org/10.3758/bf03201114 https://doi.org/10.1080/01638539809545044 https://www.sciencedirect.com/science/article/pii/0749596x9290029w https://www.sciencedirect.com/science/article/pii/0749596x9290029w https://doi.org/10.1080/01638530709336895 https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf https://doi.org/10.5087/dad.2014.101 marchal, scholman and demberg ted j. m. sanders. coherence, causality and cognitive complexity in discourse. in proceedings/actes sem-05, first international symposium on the exploration and modelling of meaning, pages 105–114, toulouse, france, 2005. ted j. m. sanders and leo g. m. noordman. the role of coherence relations and their linguistic markers in text processing. discourse processes, 29(1):37–60, 2000. doi:10.1207/s15326950dp2901 3. ted j. m. sanders, wilbert p. m. s. spooren, and leo g. m. noordman. toward a taxonomy of coherence relations. discourse processes, 15(1):1–35, 1992. https://www.tandfonline. com/doi/abs/10.1080/01638539209544800. merel c. j. scholman. coherence relations in discourse and cognition: comparing approaches, annotations and interpretations. phd thesis, saarland university, 2019. merel c. j. scholman and vera demberg. crowdsourcing discourse interpretations: on the influence of context and the reliability of a connective insertion task. in proceedings of the 11th linguistic annotation workshop (law), pages 24–33, valencia, spain, 2017a. merel c. j. scholman and vera demberg. examples and specifications that prove a point: identifying elaborative and argumentative discourse relations. dialogue & discourse, 8(2):56–83, 2017b. doi:10.5087/dad.2017.203. merel c. j. scholman, hannah rohde, and vera demberg. “on the one hand” as a cue to anticipate upcoming discourse structure. journal of memory and language, 97:47–60, 2017. doi:10.1016/j.jml.2017.07.010. merel c. j. scholman, vera demberg, and ted j. m. sanders. individual differences in expecting coherence relations: exploring the variability in sensitivity to contextual signals in discourse. discourse processes, 57(10):844–861, 2020. doi:10.1080/0163853x.2020.1813492. merel c. j. scholman, tianai dong, frances yung, and vera demberg. discogem: a crowdsourced corpus of genre-mixed implicit discourse relations. in proceedings of the thirteenth international conference on language resources and evaluation (lrec’22), marseille, france, 2022. european language resources association (elra). juliane schwab and mingya liu. lexical and contextual cue effects in discourse expectations: experimenting with german ’zwar... aber’and english ’true/sure... but’. dialogue & discourse, 11(2), 2020. doi:10.5087/dad.2020.203. erwin m. segal, judith f. duchan, and paula j. scott. the role of interclausal connectives in narrative structuring: evidence from adults’ interpretations of simple stories. discourse processes, 14 (1):27–54, 1991. doi:10.1080/01638539109544773. reid smith, pamela snow, tanya serry, and lorraine hammond. the role of background knowledge in reading comprehension: a critical review. reading psychology, 42(3):214–240, 2021. doi:10.1080/02702711.2021.1888348. 74 https://doi.org/10.1207/s15326950dp2901\unhbox \voidb@x \kern .06em\vbox {\hrule width.3em}3 https://www.tandfonline.com/doi/abs/10.1080/01638539209544800 https://www.tandfonline.com/doi/abs/10.1080/01638539209544800 https://doi.org/10.5087/dad.2017.203 https://doi.org/10.1016/j.jml.2017.07.010 https://doi.org/10.1080/0163853x.2020.1813492 https://doi.org/10.5087/dad.2020.203 https://doi.org/10.1080/01638539109544773 https://doi.org/10.1080/02702711.2021.1888348 the effect of domain knowledge and implicitation caroline sporleder and alex lascarides. using automatically labelled examples to classify rhetorical relations: an assessment. natural language engineering, 14(3):369–416, 2008. doi:10.1017/s1351324906004451. paul van den broek. using texts in science education: cognitive processes and knowledge representation. science, 328:453–456, 2010. doi:10.1126/science.1182594. noortje j. venhuizen, matthew w crocker, and harm brouwer. expectation-based comprehension: modeling the interaction of world knowledge and linguistic experience. discourse processes, 56 (3):229–255, 2019. doi:10.1080/0163853x.2018.1448677. bonnie webber. what excludes an alternative in coherence relations? in proceedings of the international conference on computational semantics (iwcs), pages 921–950, potsdam, germany, 2013. https://aclanthology.org/w13-0124.pdf. bonnie webber, rashmi prasad, alan lee, and aravind joshi. a discourse-annotated corpus of conjoined vps. in proceedings of the 10th linguistic annotation workshop (law x), pages 22–31, berlin, germany, 2016. frances yung, vera demberg, and merel c. j. scholman. crowdsourcing discourse relation annotations by a two-step connective insertion task. in proceedings of the 13th linguistic annotation workshop (law), pages 16–25, 2019. https://aclanthology.org/w19-4003.pdf. sandrine zufferey and liesbeth degand. annotating the meaning of discourse connectives in multilingual corpora. corpus linguistics and linguistic theory, 13(2):399–422, 2017. doi:10.1515/cllt-2013-0022. appendix a. relation sense classification (i) cause • cause • cause+belief • cause+speechact • purpose (ii) temporal • synchronous • asynchronous (iii) contrast (iv) concession • concession • concession+speechact 75 https://doi.org/10.1017/s1351324906004451 https://doi.org/10.1126/science.1182594 https://doi.org/10.1080/0163853x.2018.1448677 https://aclanthology.org/w13-0124.pdf https://aclanthology.org/w19-4003.pdf https://doi.org/10.1515/cllt-2013-0022 marchal, scholman and demberg (v) positive expansion • similarity • conjunction • equivalence • instantiation • level-of-detail • manner (vi) negative expansion • disjunction • exception • substitution (vii) condition • condition • condition+speechact • negative-condition • negative-condition+speechact appendix b. means across conditions table 6: mean percentage of correct answers per condition. biodrb pdtb bio eco bio eco mean implicitated result 57.3 47.2 72.6 72.1 62.4 instantiation 71.4 41.0 46.8 49.5 52.0 concession 48.4 28.6 48.9 58.6 48.6 contrast 34.7 27.6 42.9 14.3 31.2 mean 52.2 36.2 54.3 58.7 implicit result 70.3 69.1 64.1 65.1 67.1 instantiation 65.1 51.9 64.1 60.5 60.7 concession 41.0 36.1 54.9 54.2 49.1 contrast 39.1 25.9 33.3 20.0 31.7 mean 54.1 45.4 58.7 57.0 76 the effect of domain knowledge and implicitation appendix c. model output summaries models described in section 5.2 table 7 shows the model output of the subset analysis on the pdtb items, in which the effect of corpus was not significant. the effect was significant in the items in the biodrb, which is displayed in table 8. table 7: subset analysis on pdtb items. model specification: correctness ∼relationsense + expertise + relationmarking + (1 | workerid) + (1 + expertise | questionid) estimate std. error z value p value intercept 0.19 0.17 1.11 .27 relationsense:cause 0.73 0.27 2.70 <.001 relationsense:contrast -1.28 0.67 -1.92 .06 relationsense:instantiation 0.07 0.27 0.25 .80 expertise -0.04 0.08 -0.52 .60 relationmarking 0.06 0.12 0.52 .60 table 8: subset analysis on biodrb items. model specification: correctness ∼ relationsense + expertise + relationmarking + (1 | workerid) + (1 + expertise | questionid) estimate std. error z value p value intercept -0.54 0.21 -2.55 .01 relationsense:cause 1.11 0.28 3.92 <.001 relationsense:contrast -0.40 0.29 -1.40 .16 relationsense:instantiation 0.81 0.29 2.80 <.01 expertise 0.30 0.09 3.47 <.001 relationmarking 0.15 0.12 1.24 .21 models described in section 5.3.3 table 9 shows the output of the model in which the ambiguity of the first step insertions was predicted. to examine the interaction between corpus and expertise, separate analyses were performed for the items from the pdtb (table 10) and the biodrb (table 11). 77 marchal, scholman and demberg table 9: output of the model predicting ambiguity of the insertions in the first step. model specification: ambiguity ∼ relationsense + corpus*expertise + (1 + corpus | workerid) + (1 + expertise | questionid) estimate std. error z value p value intercept 0.33 0.12 2.86 <.01 relationsense:cause -0.97 0.15 -6.36 <.001 relationsense:contrast 0.05 0.19 0.28 .78 relationsense:instantiation -1.47101 0.16 -9.24 <.001 corpus -0.08 0.06 -1.30 .20 expertise -0.15 0.08 -1.95 .05 corpus:expertise -0.14 0.05 -3.08 <.01 table 10: subset analysis on items from the pdtb. model specification: ambiguity ∼ relationsense + corpus*expertise + (1 | workerid) + (1 | questionid) estimate std. error z value p value intercept 0.40 0.14 2.88 <.01 relationsense:cause -0.98 0.21 -4.62 <.001 relationsense:contrast -0.09 0.50 -0.19 .85 relationsense:instantiation -1.39 0.22 -6.38 <.001 expertise -0.01 0.09 -0.09 .93 table 11: subset analysis on items from the biodrb. model specification: ambiguity ∼ relationsense + corpus*expertise + (1 | workerid) + (1 + expertise | questionid) estimate std. error z value p value intercept 0.29 0.17 1.76 .08 relationsense:cause -1.02 0.22 -4.71 <.001 relationsense:contrast 0.04 0.21 0.19 .85 relationsense:instantiation -1.59 0.24 -6.72 <.001 expertise -0.29 0.09 -3.03 <.01 78 introduction background the role of domain knowledge in discourse inferences strategies for inferring discourse relations recognizing discourse relational cues cognitive biases in relation inferences underspecified interpretations present study method participants materials procedure data analysis results convergence with the gold label the effect of domain knowledge on discourse relation inferences interpretation strategies in the absence of domain knowledge exploiting relation marking cognitive bias for continuity and causality underspecified interpretations discussion and conclusion text genre and the influence of domain knowledge discourse relational cues inferences in the absence of domain knowledge limitations conclusion relation sense classification means across conditions model output summaries dialogue & discourse 14(1) (2023) 56–87 doi: 10.5210/dad.2023.103 bullshit, pragmatic deception, and natural language processing oliver deck oliver.deck@rub.de ruhr university bochum editor: vera demberg submitted 06/2022; accepted 05/2023; published online 05/2023 abstract fact checking and fake news detection has garnered increasing interest within the natural language processing (nlp) community in recent years, yet other aspects of misinformation remain unexplored. one such phenomenon is ‘bullshit’, which different disciplines have tried to define since it first entered academic discussion nearly four decades ago. fact checking bullshitters is useless, because factual reality typically plays no part in their assertions: where liars deceive about content, bullshitters deceive about their goals. bullshitting is misleading about language itself, which necessitates identifying the points at which pragmatic conventions are broken with deceptive intent. this paper aims to introduce bullshitology into the field of nlp by tying it to questions in a qud-based definition, providing two approaches to bullshit annotation, and finally outlining which combinations of nlp methods will be helpful to classify which kinds of linguistic bullshit. keywords: bullshit, nlp, pragmatics, misinformation, qud 1. introduction common parlance applies the term bullshit to nonsensical utterances, lies, personal affronts, injustices, and any number of daily annoyances. but from a linguistic perspective, bullshit presents itself as a highly complex, amorphous pragmatic concept which calls for novel methods when approached from a natural language processing (nlp) perspective. for some bullshit statements, context is extremely important while others may reveal themselves through textual analysis alone. consider the following statement made by former u.s. president donald trump: (1) i am the least racist person anybody is going to meet. (trump, 2018) it may not be immediately evident, why this statement is bullshit from a linguistic perspective. here, the interpretation does not stem from how one views trump. rather, it is based on the form of the utterance: he could have chosen to lie, e.g. by saying no one has ever accused me of racism, which would have been an assertion that can be true or false. instead, he chose to use the form of an assertion to make a statement about an unquantifiable and thus unverifiable concept. he in essence deceived about the pragmatics of the utterance itself. the statement can be analyzed as bullshit without the need for additional textual context since one cannot make such an assertion with implied certainty about one’s degree of racism.1 1. although world knowledge is needed to ascertain which concepts are inherently unquantifiable. we also need to take into account the audience: a favorable listener might interpret the utterance as an hyperbolic way of saying i am ©2023 oliver deck this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). bullshit, pragmatic deception, and nlp other examples of bullshit require careful analysis of the larger context or of utterance-external factors: evasive answers in interviews or exams, for example, can only be interpreted as bullshit when taking into account the questions they fail to answer. over the past four decades, the field of ‘bullshitology’ has investigated a number of related phenomena that are generally deceptive in nature, but, crucially, cannot be classified as outright lies. since its introduction into academia in the highly influential paper “on bullshit” by american philosopher harry g. frankfurt (frankfurt, 2005), researchers in different disciplines have tried to define what exactly bullshit is. however, while audiences may intuitively know when to call someone out on it, the scientific concept of bullshit remains unclear and many researchers have proposed their own definitions of these deceptive practices outside of the truth-untruth dichotomy. much computational research has, in recent years, focused on the linguistics of lying, deception, fake news, propaganda, and a host of related phenomena – but so far not on bullshit. shared tasks on fake news detection consistently garner considerable interest, leading to improvements in automated systems (zhou and zafarani, 2020; zeng et al., 2021; shahi et al., 2021). we cannot identify bullshit via content analysis, since, unlike lies and fake news, it does not depend on being factually wrong. rather, it is a kind of deception about language use, which calls for linguistic analysis. frankfurt may well have exaggerated when he called bullshit a “greater enemy of the truth than lies” (frankfurt, 2005, p. 15), but in the context of ‘alternative facts’ and ‘post-truth,’ it is imperative to approach bullshit as one subclass of disinformation. this paper approaches bullshit for the first time from a computational linguistics perspective. by its nature, computational linguistics can rarely look at the producer – the bullshitter – and typically focuses on the product, in this case textual or spoken bullshit. we typically have no insight into the cognition, situation (except from a broader, pragmatic context), or intention of an author and must therefore base our methods on text alone, so deciding whether an utterance like (1) is bullshit (in the linguistic sense), and not lying, hyperbole, or some other phenomenon is challenging. it relies on a workable definition of bullshit, narrow enough to be applicable to very specific kinds of writing or speech, but broad enough to encompass at least those examples of bullshit on which previous research agrees. the definition must also be operationalizable, i.e. it must describe bullshit that can be classified via existing – or at least theoretically possible – methods of computational linguistics. from this need follows the first research question: rq 1: what is a useful bullshit definition for a computational linguistics approach? in this paper, i argue that the deception inherent in bullshit is a deception about pragmatics, undermining conventional norms of discourse. where liars typically use language to deceive about the real world, about concepts, feelings, or their state of mind, bullshitters deceive about language itself. they use language to mislead about their language use. bullshit happens at the breaking points of pragmatic models, where cooperative discourse is disregarded. in saying i am the least racist person anybody is going to meet, donald trump neither lies nor does he tell the truth. in fact, with a statement of this kind, it is impossible to do either. the statement is presented as an assertion, but is unverifiable by its nature. rather, the language of assertion is used to present an image of the speaker – an ‘expressive’ masquerading as an ‘assertive’ in the nomenclature of searle (1976) – shrouding the actual intent of the speaker but crucially not lying about anything. to find not racist. when exactly hyperbole is acceptable as a stylistic choice and when it becomes bullshit is dependant on extralinguistic factors of speaker, situation, and audience. 57 deck a sufficiently large number of bullshitting examples, we need to create processes and guidelines of bullshit annotation and analysis. the second research question is therefore: rq 2: how can we annotate and analyze bullshit? the third research question concerns the relevant nlp methods for bullshit detection: rq 3: what are promising nlp approaches to bullshit classification? computational pragmatics is a complex and fairly recent field which is still far from solved in almost all areas of inquiry, yet we may connect pragmatic models to related computational approaches. while we cannot directly build on the vast amount of fact-checking research, other fields may help in stages of the bullshit detection process: claim identification (e.g. hassan et al., 2015; barrón-cedeño et al., 2020), question answering (e.g. choi et al., 2018; kim et al., 2021), question/answer congruity (e.g. faruqui and das, 2018; yu and jiang, 2021), argument quality (e.g. skitalinskaya et al., 2021), propaganda detection (e.g. barrón-cedeño et al., 2019a,b; da san martino et al., 2020a,b), and others. since such issues have garnered greater attention in the nlp community, synergistic effects may be utilized for bullshit classification and detection. the next chapter will answer rq 1 by looking at the history of bullshit research and formulating a definition that serves as a basis for our analysis. chapter 3 will outline different avenues of operationalizing the bullshit annotation process in order to answer rq 2, reporting on a small pilot study on annotating a subtype of bullshitting. finally, chapter 4 will provide an outlook on the potential of natural language processing in automatically classifying bullshit in an answer to rq 3. 2. what is bullshit? 2.1 the history of bullshitology frankfurt originally describes bullshit as those statements in which speakers do not care whether what they say is true or false (frankfurt, 2005). liars and truth-tellers must at least have some notion of the truth to either distort or affirm it; bullshitters have different goals, such as portraying a specific self-image. this simple view of statements without caring about the truth has since been called into question and refined by a number of researchers from different fields. many agree with frankfurt on the dangers of bullshit, while others have studied more harmless forms and highlight a number of positive and prosocial aspects of bullshitting; e.g. in a study on hitchhikers in america (mukerji, 1978). in these situations, where all participants expect bullshit, the practice can take the form of a language game aiding speakers’ identity work (cf. mears, 2002), fulfilling functions like socialization, expressing feelings, passing time, resolving personal or interpersonal strain, impression management, and others (mears, 2002). others do not classify such playful forms as bullshit, especially when all participants are aware of the practice and there is no deception involved (e.g. stokke and fallis, 2017). on larger, societal scales, bullshit can become harmful by weakening the social and linguistic norms of what constitutes acceptable, cooperative discourse. spicer argues that, in workplace settings, harmless bullshit can become harmful when performing the form becomes more important than the content (cf. spicer, 2020). similarly, so-called pseudo-profound bullshit2 as in (2) may 2. pseudo-profound bullshit comprises non-profound, inane statements masquerading as profound. note, neither conventionally profound statements such as “a wet person does not fear the rain” (pennycook et al., 2015, p.549) nor 58 bullshit, pragmatic deception, and nlp seem humorous, but research has shown that individuals who ascribe meaning to such sentences are also prone to believe in alternative medicine, angels, and pseudoscience, as well as fake news or propaganda (pennycook et al., 2015; pennycook and rand, 2018, 2020, 2021; littrell et al., 2021a,b). believing some kinds of bullshit may thus not in itself be harmful to society, but can be an indicator for ways of thinking that may be susceptible to exploitation by bad-faith actors. (2) wholeness quiets infinite phenomena (pennycook and rand, 2020, p. 196) following frankfurt’s view of the dangers of bullshit itself, a number of researchers highlight the need to deal with it similarly to lying and general disinformation. focusing on the product instead of the producers, cohen notes that ‘academic bullshit’ is harmful to academic practice (cohen, 2002), while others seek to educate the public about bullshit because of its negative consequences both on individuals and society as a whole (sagan, 1996; bergstrom and west, 2020; petrocelli, 2021) . in a rejection of frankfurt’s simple definition of bullshit, carson (2016) first introduced the distinction between persuasive and evasive bullshit, which has since brought to light striking differences in the situations and reasons for bullshitting, as well as in bullshitters’ perception and cognitive abilities. whereas frankfurt dealt with speakers using persuasive bullshit propositions to present themselves (often unprompted) in a specific light, evasive bullshit tends to occur when prompted: politicians may employ evasive bullshitting to avoid giving a straight-up answer; students may not know the answer to an exam question but hope that simply saying something might award them some points as in (3). in such cases, the speaker may very well care about the truth value of what they are saying, yet try to hide that they are not really answering the question. the split between persuasive and evasive bullshit has since been substantiated by experimental psychology, showing measurable cognitive differences between those who more often persuasively bullshit and those that use bullshit as an evasive tactic (littrell et al., 2021a,b). (3) careful test taker a student who gives a bullshit answer to a question in an exam might be concerned with the truth of what she says. suppose that she knows that the teacher will bend over backwards to give her partial credit if she thinks that she may have misunderstood the question, but she also knows that if the things she writes are false she will be marked down. in that case, she will be very careful to write only things that are true and accurate, although she knows that what she writes is not an answer to the question. (carson, 2016, p. 62) in linguistic research outside of bullshitology, there is also a rich history of inquiry into evasive speech. greatbatch investigates ‘agenda-shifting behaviors’ in political interviews, where interviewees shift the topics of questions and provide answers to more favorable (to them) subjects. this idea is similar to the one described in section 3.2 based on gabrielsen et al. (2020). political interviews also serve as the basis for a bull and mayer (1993) paper on avoiding questions. the authors develop a typology of 11 avoidance categories with 30 subcategories, most of which cannot be classified as bullshitting behavior. among them are ignoring the question, questioning the question, attacking the question, attacking the interviewer, and others. one category that overlaps with evasive bullshitting is ‘makes political point’ which includes presenting or justifying policy, appeals to pseudo-profound statements as in (2) must be factually true or false. instead the latter deceives about being the same linguistic device as the former. 59 deck nationalism, and self-justification (bull and mayer, 1993, p. 656-661). as will be shown in section 3.2, such shifting maneuvers can be interpreted as one type of bullshitting. łupkowski and ginzburg (2016) utilize corpora to look at questions that are in turn answered by questions on a larger scale. ginzburg et al. extend the model to a “formally underpinned characterization of the response space of questions.” (ginzburg et al., 2022, p. 40) their typology consists of three categories of answers: question-specific, clarification responses and evasion responses. question-specific responses are either answers or dependent questions (ginzburg et al., 2022, p. 2). evasion responses can state that it is difficult to answer, question the motive of the original query, or change the topic (cf. bull and mayer, 1993). queries can also be evaded by ‘ignoring,’ which in the context of this model describes “adress[ing] the situation, but not the question” (ginzburg et al., 2022, p. 2). this behavior could be classified as bullshitting as described in section 3.2. while evasive speech has thus prompted linguistic analysis for quite some time, when it comes specifically to bullshit, there are only two major definitions: meibauer (2016, 2020) and stokke and fallis (2017). meibauer bases his definition of bullshit on four factors: assertion, loose concern for the truth, misrepresentational intent, and too much certainty (meibauer, 2016, p. 75). in his view, bullshitters do not mislead about the content of their assertion (as liars do), but instead about their own loose concern for the truth of these assertions. meibauer’s definition is a good starting point for purely linguistic inquiry but it is unclear if and how the four factors may be operationalized computationally: while we may find linguistic markers for certainty, the loose concern for the truth and misrepresentational intent may be impossible to ascertain with purely linguistic means. from an nlp standpoint, stokke and fallis’ approach could therefore prove more fruitful. the researchers look at bullshit through the lens of questions under discussion (qud, see roberts, 2012), subsuming frankfurt’s original indifference-to-truth definition of bullshitting (stokke and fallis, 2017, p. 279). where frankfurt assumed the bullshitter to be indifferent to the truth at a content level, stokke and fallis propose a definition in which speakers are indifferent to whether or not their utterance constitutes a truthful answer to a qud. the approach thus covers cases like bullshitting while caring about the truth – as in (3) – and bullshitting while lying.3 as noted earlier, stokke and fallis disregard most prosocial, playful kinds of bullshitting, since they do not consider these utterances real assertions occurring in serious situations (stokke and fallis, 2017, p. 290). the two linguistic theories again highlight that bullshit is deceptive on a different level than lying. for meibauer, bullshitters deceive about intention, for stokke and fallis, they deceive about their cooperative participation in qud-based discourse. especially the latter definition hints at bullshitting being a phenomenon that breaks pragmatic conventions. linguistic analysis points toward a disregard not for the content, as originally postulated by frankfurt, but rather for the linguistic practices and norms themselves. stokke and fallis’ definition does not encapsulate all kinds of especially evasive bullshitting, but it can serve as a starting point for nlp methods, since the qud approach ties into computational fields related to question answering, question-answer-congruence and others. especially with the release of chatgpt (openai, 2022), a number of news outlets, blogs, and ai researchers have been calling large language models (llms) bullshitters (e.g. bernoff, 2022; narayanan and kapoor, 2022; nast, 2022). there is as of yet no formal analysis of llms as bullshitters though some writers refer to the frankfurtian definition. however, following the view 3. carson (2016, p. 60f.) gives an example of covering for a friend’s atheism by giving a long-winded answer including “as a boy he always went to church and loved singing christmas carols” while knowing this to be untrue. in carson’s interpretation, the whole answer constitutes evasive bullshitting while part of it is clearly a lie. 60 bullshit, pragmatic deception, and nlp of bender et al. of llms as ‘stochastic parrots’ that only repeat learned patterns of language, these models can certainly be deemed bullshitters from a linguistic perspective. they focus on form over content since the model does not ‘know’ what information is; it only ‘knows’ what information should look like. investigating large language models as producers of bullshit texts is an important area of future research, especially once they become more common in daily life where people may use them to access factual information when the models are typically trained to reproduce surface patterns. 2.2 what isn’t bullshit? “never tell a lie when you can bullshit your way through.” eric ambler’s character arthur abdel simpson (cited in frankfurt, 2005) the term bullshit suffers from being a very common expletive applied to all kinds of situations. as pointed out by bergstrom and west, not everything that makes people exclaim that’s bullshit! is of interest from the point of bullshit studies: “you can call bullshit on bullshit, but you can also call bullshit on lies, treachery, trickery, or injustice” (2020, p. 314). there are (at least) three different groups of phenomena that are of no interest for the current computational linguistics based analysis: non-bullshit, non-linguistic bullshit and non-pragmatic bullshit. the first group consists of phenomena that in some way or other overlap with aspects of bullshit, lying being the prime example. though bullshit may sometimes contain lies, frankfurt and those who follow his line of reasoning agree that, though related, bullshitting is different from lying. indeed, psychological research has provided evidence that at least some kinds of bullshitters are averse to lying and avoid it if possible (littrell et al., 2021a,b). in a similar vein, simply telling the truth is also not bullshitting. on the surface, bullshit can be both true or false, but the intentions and functions of truthful assertions and bullshit statements differ greatly from a pragmatic standpoint. strongly related to lying, albeit created in a more organized fashion, is propaganda. cassam rightfully points out that it seems wrong to call the propaganda in a göbbels speech bullshit (cassam, 2019). however, the similarities are there: propagandists, like (frankfurtian) bullshitters do not care whether what they are saying is true or not. propaganda may be more effective if rooted in truth, but may just as easily appeal to people’s preconceived notions or stereotypes. what makes propaganda different from bullshit, aside from the intuition that it is somehow ‘worse,’ is its usually more concerted nature. propaganda is often carefully crafted by multiple people (e.g. state actors) with a specific purpose in mind. it may be used to enhance a specific person’s image (e.g. in dictatorships) but it does not typically originate from that person in the spur of the moment. it therefore differs not so much in form, but rather in its origins.4 another concept which has long been connected to bullshit is obscurantism (cohen, 2002; ivanković, 2016). it is certainly viable to define obscurantism as a subtype of bullshit: typically intended to heighten the status of an author by portraying them or their work as something so complex that it is not understandable to the reader, obscurantism may constitute a kind of bald-faced bullshitting; comparable to the prosocial bullshitting in bull sessions and other settings where the language game nature of the utterances is apparent to all – or at least all who are in the know. on 4. an automated system might thus classify propaganda as bullshit, and the distinction would rather be one of taste: the term bullshit carries with it a certain levity that we may not want to assign to the much more negatively connotated propaganda. however, one could imagine a system that outputs all possible instances of bullshit/propaganda and human annotators could then differentiate between the two based on social and other language external factors. 61 deck the other hand, obscurantism may be interpreted as “betray[ing] an indirect move to confound while promising deep content” (bien, 2021, p. 1498). following this view, obscurantism would be more closely related to so-called argumentative bullshit (gascón, 2021). such more or less deceptive practices rely on at least some part of the audience acknowledging that epistemic conventions are lax for specific purposes: in movies or on stage, actors rely on their audience’s knowledge that what they are saying is not serious and resides outside of truthful and deceptive discourse. the same holds for humor and (in-)jokes, as well as irony, sarcasm, hyperbole and other rhetoric devices. the second group of bullshit-related phenomena includes non-linguistic bullshit such as bullshit jobs5 or bullshit visualizations. bergstrom and west provide a number of graphics and charts that are so overloaded with fluff or simply misleading, and “attempt to be cute [making] it harder for the reader to understand the underlying data” (bergstrom and west, 2020). such graphics are not deceptive about their content but about the medium itself – which makes them bullshit – but they are of no particular interest to a linguist. lastly there are types of linguistic bullshit that will not be discussed in detail in this paper, i.e. non-pragmatic bullshit: as pointed out by fredal, the misleading and deceptive language of george orwell’s 1984, later called ‘doublespeak’ could be interpreted as ‘bullshit words’ in that it reinterprets words, obfuscating their conventional meaning (fredal, 2011, p. 11). the term bullshit can also be applied to larger structures of language: “[a]rgumentative bullshit could be the production of reasons for a claim without regard to whether the reasons given really support that claim” (gascón, 2021, p. 293), therefore using the form of an argument to deceive about it being based on sound reasoning. 2.3 bullshit as pragmatic deception 2.3.1 intuition in computational linguistics and nlp, much research focuses on factual deception such as lies and, especially in recent years, fake news debunking and related tasks. one standard approach consists of identifying claims and then comparing them to some sort of database for verification (see zhou and zafarani, 2020; augenstein, 2021; guo et al., 2022). for identifying bullshit statements, however, the question is not is this true?, but rather what does this actually say?, making existing fact checking pipelines useless for bullshit detection, apart from the claim identification part. since the deception in bullshit is about language itself, a bullshit detection pipeline instead needs to identify the points at which pragmatic conventions or models are deceptively broken: “the normal assumptions that interlocutors make about the veracity and relevance of another’s statements (relying on paul grice’s maxim of quality, for example) are misplaced when applied to the bullshitter: we think this person is having a ‘serious’ conversation when such is not the case” fredal, p. 21. while classic pragmatic models – like the gricean maxims – may be useful for linguistic analysis of a single specific bullshit example, they lack a clear path towards operationalization on a larger scale. that is, they rely too much on the linguistic intuitions of annotators and world knowledge when it is highly doubtful if these can be translated into an algorithmic approach. computational pragmatics is a fairly new field, still finding new ways to deal with the often difficult-to-grasp aspects of communication in context. if we want to identify instances of pragmatic conventions being broken from an nlp perspective, we may first need a robust pragmatic model of cooperative dialog and build bullshit detection on top of it. as mentioned above, of the two 5. which take the form of useful jobs, but are empty and meaningless (graeber, 2013) 62 bullshit, pragmatic deception, and nlp recent linguistic approaches to bullshit, the qud-based approach by stokke and fallis seems more promising from the standpoint of operationalization. the next chapter will provide a short overview of the qud model and how it may be adapted for bullshit analysis. 2.3.2 bullshit and quds over the last decades, a number of researchers have identified (implicit) underlying questions as a tool to describe the internal structure of discourse (e.g. polanyi, 1988; von stutterheim and klein, 1989; kuppevelt, 1995; ginzburg, 1994, 1996). this paper largely follows the notion of quds proposed by roberts (2012). the framework “which treats discourse as a game, with context as a scoreboard organized around the questions under discussion by the interlocutors,” (roberts, 2012, p. 1) provides us with a way of interpreting speaker’s utterances in the dialog flow. the qud approach revolves around the idea of a stack of so-called ‘current questions’ which are salient in the discourse. dominated by the ‘big question’ (what is the world like?), discourse in the qud model aims at progressively answering more and more detailed subquestions. in this framework, cooperative speakers are beholden to answer – or more broadly to ‘deal with’ – these current questions. current questions can be overt (e.g. what happened to you?) or implicit (e.g. the utterance you look horrible! may entail the same current question what happened to you?). discourse participants may then choose to answer the question (e.g. i fell down the stairs.), answer a subquestion that ‘closes’ the superquestion (e.g. let’s just say be careful when walking down slippery stairs...), or dismiss the question (e.g. i don’t want to talk about it). however, current questions cannot be simply ignored in cooperative discourse since they are highly salient and should thus be on the forefront of the participants’ minds. previously, the qud framework was used to annotate discourse structure (de kuthy et al., 2018), automatically generate potential quds from assertions (de kuthy et al., 2020), as well as generating quds evoked by the preceding context, in essence predicting current questions and the following discourse (westera et al., 2020). the ginzburg et al. (2022) paper, mentioned above, also uses quds as a basis for analyzing the relevance of responses; building on ginzburg (2012), the authors provide a formal analysis of both cooperative and uncooperative discourse or questionanswering. the qud framework – and not other pragmatic models that may entail deception such as asher et al. (2017); asher and paul (2018) – was chosen as a basis for computational bullshit detection for two main reasons: firstly, the definition by stokke and fallis is a good starting point and is already based on quds. secondly, in dealing with quds similarly to overt questions, we may leverage the large amount of nlp research in question related fields. from a bullshitting perspective, there are two ways in which current questions may be deceptively abused: a) providing a straight-up answer without having evidence for or caring about its truth value; and b) introducing or answering a different question portraying it as pertaining to the original one. the first option is the one discussed by stokke and fallis who propose the following definition of bullshitting: (i) a is bullshitting relative to a qud q if and only if a contributes p as an answer to q and a is not concerned that p be an answer to q that her evidence suggests is true or that p be an answer to q that her evidence suggests is false. (stokke and fallis, 2017, p. 288) stokke and fallis show how their definition of bullshit subsumes that of frankfurt, but also serves to explain many of the other examples (and indeed counterexamples) which have been put 63 deck forth in the intervening years. in particular, they show that infidelity to the qud explains cases in which bullshitters care about the truth of what they are saying, but may also lie if necessary. stokke and fallis’ definition in (i) does not include those (evasive) cases in which speakers stray from the original qud and introduce (or implicitly answer) a novel qud. since psychological research has shown this kind of bullshitting to be distinct both cognitively and in the situations it occurs, this might be fine: there may simply be two distinct pragmatic phenomena that may call for two distinct definitions. however, the connections between the two kinds of bullshitting, as well as the shared history of research call for a definition that covers both which will be provided in the next chapter. 3. operationalization 3.1 defining bullshit recent psychological research indicates the importance of capturing two kinds of not caring about answering the qud: speakers trying to appear to answer a qud without having evidence of their answer being a truthful one (often persuasive bullshitting) or deceptively introducing or implying a new qud (typical for evasive bullshitting). an ideal automated system should be able to capture both, as they share some important characteristics. that is, both persuasive and evasive bullshit uses the form of cooperative discourse to hide the fact that its content is ‘empty’ (cf. spicer, 2020) with regard to the current discourse. current questions in bullshitting situations differ from those in non-bullshitting situations. for most kinds of bullshitting, quds are strongly connected to self-representation and image. while most communication entails subjective markers and practices of creating a self-image, bullshit seems overtly connected to it. in felicitous discourse, the big question should be what is the world like?, yet for the qud-bullshitter, the most important discourse question may be what am i like? this holds true for the bullshitting examples in section 3.1.1: frankfurt’s 4th of july orator in example (6), carson’s careful test taker in (3), stokke and fallis’ wishful bullshitter in (5) all misuse the current (implied or overt) question under discussion to talk about themselves: i am someone who values religion, i am someone who knows things – even if i can’t answer this particular question and i am someone who wants it not to rain, respectively. even the more abstract cases of bullshit can be analyzed in terms of a shifted big question: meibauer’s advertising bullshit tells readers something about how a product should be perceived – not about what it factually is. bergstrom and west’s examples of overly illustrated graphs use visual means to heighten their own cleverness while (perhaps unintendedly) obscuring the actual data. for large language models, the big question might be how does ‘what is the world like’ look like? these observations also hold true for linguistic bullshit, both of the persuasive and evasive kind. evasive bullshitters may often dodge the qud to proclaim that they are ‘(not) someone who does/is/believes x’, but may also use the opportunity to define their self-image by appearing to answer a relevant qud: (4) q: ‘are you going to contest roe vs. wade?’ a: ‘i am someone who believes in the constitution and the supreme court’ (cf. meibauer, 2018, p. 367). 64 bullshit, pragmatic deception, and nlp the answer in this example is ‘empty’ with regard to the overt qud. it answers other quds (who am i? what do i believe in?), implying that they may be relevant subquestions that closes the original qud. at closer inspection, it can be interpreted as yes, no, or neither. it only serves to build the image of the speaker by using heavily connotated phrases like ‘constitution’. defining bullshit as pragmatic deception thus leads to an expanded version of stokke and fallis’ definition: (ii) a is bullshitting relative to a qud q if a contributes p as an answer to q and a is not concerned that p be an answer to q that her evidence suggests is true or that p be an answer to q that her evidence suggests is false. (stokke and fallis, 2017, p. 288) a is also bullshitting relative to a qud q if a introduces, or by answering implies, a novel qud q’, misrepresenting it as pertaining to the original qud q.6 the first part of the definition still contains the problematic (from a text analysis perspective) reliance on the speaker caring about the truth of an answer. the second part, in turn, entails misrepresentational intent. both may result in annotators having to use subjective judgments which, as we will see later, complicates classification. however, there may exist approaches that focus on a subset of bullshit that answers questions which by definition can be neither true nor false, as in example (1). the second part may also be helpful in social media contexts, where (persuasive) bullshit often occurs unprompted, e.g. in tweets that nevertheless imply an overarching topic with associated quds by way of hashtags or other markers (cf. example 8). the next section shows how this definition can be operationalized to classify bullshit in future projects. 3.1.1 empirical coverage having formulated a definition that encompasses the desired types of bullshitting, the next step is the creation of a corpus. the bullshit definition in (ii) must therefore be operationalized for use by annotators. a preliminary annotation flowchart was created, consisting of a series of yes/no questions that replicate the decision process in bullshit identification (figure 1). an instance of stokke and fallis bullshit, for example, would map to yes yes yes no (yyyn) – moving through nodes a, b, c, and d. conversely, a simple exclamation like ‘hi mary!’ would map to nn (nodes a and b) and not constitute bullshit. the flowchart was used to qualitatively annotate some typical examples of bullshit. (3) careful test taker a student who gives a bullshit answer to a question in an exam might be concerned with the truth of what she says. suppose that she knows that the teacher will bend over backwards to give her partial credit if he thinks that she may have misunderstood the question, but she 6. as one anonymous reviewer pointed out, there are other, similar kinds of pragmatic deception. a common trope in spy novels has one agent telling another i had a turbulent flight to indicated that an adversary was on the plane with them. in my view, the difference is that such coded language (similar to the language of so-called dog whistles) relies on the intended audience knowing that there is a second, hidden qud. the speaker does not use the surface qud (about the tranquility of the flight) to talk about themselves, but instead using it to covertly answer the hidden qud (about the adversary, using previously agreed upon codes). neither test-taker nor teacher in (3), for example, believe that an evasive bullshitting answer pertains to some sort of secret question. similar to in-jokes, playacting, etc, coded speech has an intended audience which is not the target of deception, while bullshitting does not. 65 deck detecting qud-based bullshit is there an explicit qud? a is something portrayed as an answer to that qud? b does it actually answer the qud? c does the speaker believe they have sufficient evidence for their answer? d does the utterance imply a qud? e does the speaker introduce a new qud? f is the new qud related to the original qud? g does the speaker want to appear as though they answered the qud? h possible bs no bs yes no yes no yes no yes no yes no yes no yes no yesno figure 1: qud-based bullshit flowchart also knows that if the things she writes are false she will be marked down. in that case, she will be very careful to write only things that are true and accurate, although she knows that what she writes is not an answer to the question. (carson, 2016, p. 62) since we have no insight into speakers’ minds, many bullshit statements have several interpretations. carson’s example can be interpreted in at least two different ways. interpretation 1: 1. node a: is there an explicit qud? → yes, the teacher’s question. 2. node b: is something portrayed as an answer to that qud? → yes, the student’s answer. 3. node c: does it actually answer the qud? → no, the student does not know the answer. 4. node h: does the speaker want to appear as though they answered the qud? → yes, the student provided an answer about a related topic, implying it answers the original qud. path: yyny → possible bullshit interpretation 2 (following carson): 1. node a: is there an explicit qud? → yes, the teacher’s question. 66 bullshit, pragmatic deception, and nlp 2. node b: is something portrayed as an answer to that qud? → no, the student’s answer is overtly meant for another question. 3. node f: does the speaker introduce a new qud? → yes, the seemingly misunderstood question. 4. node g: is the new qud related to the original qud? → yes, the student wants to stay as close as possible to the original question to obtain partial credit. 5. node c: does it actually answer the qud? → no, the student does not know the answer to the original question. 6. node h: does the speaker want to appear as though they answered the qud? → yes, the student wants to appear as though they answered the question they misunderstood, even though it differs from the actual question asked by the teacher. path: ynyyny → possible bullshit (5) wishful bullshitter jack and julia are going to chicago. they have tickets to a cubs game, and being a big cubs fan, julia hopes the game will not be rained out. a few days before their departure, they are talking about their trip. jack says, ‘i’m really looking forward to that cubs game. i hope it won’t rain.’ julia replies with a confident air, ‘this time of year, it’s always dry in chicago.’ but she has no evidence about the weather in chicago, and she has no idea whether it’s likely to rain or not. (stokke and fallis, 2017, p. 7) 1. node a: is there an explicit qud? → no 2. node e: does the utterance imply a qud? → yes, whether or not it will rain in chicago. 3. node b: is something portrayed as an answer to that qud? → yes, julia’s assertion that it is always dry this time of year. 4. node c: does it actually answer the qud? → yes. 5. node d: does the speaker believe they have sufficient evidence for their answer? → no, julia only wishes it to be so. path: nyyyn → possible bullshit (6) 4th of july orator consider a fourth of july orator, who goes on bombastically about ‘our great and blessed country, whose founding-fathers under divine guidance created a new beginning for mankind.’ (frankfurt, 2005, p. 4) 1. node a: is there an explicit qud? → no 2. node e: does the utterance imply a qud? → yes, whether or not america was founded under divine guidance. 3. node b: is something portrayed as an answer to that qud? → yes, the orator’s assertion both implies the qud and affirms it. 4. node c: does it actually answer the qud? → yes. 67 deck 5. node d: does the speaker believe they have sufficient evidence for their answer? → no, the speaker has no insight into the divine (in frankfurt’s interpretation). path: nyyyn → possible bullshit (7) trump & putin reporter: do you accept that part of the finding? and will you undo what president obama did to punish the russians for this or will you keep it in place? donald trump: well, if – if putin likes donald trump, i consider that an asset, not a liability, because we have a horrible relationship with russia. russia can help us fight isis, which, by the way, is, number one, tricky. i mean if you look, this administration created isis by leaving at the wrong time. the void was created, isis was formed. if putin likes donald trump, guess what, folks? that’s called an asset, not a liability. now, i don’t know that i’m gonna get along with vladimir putin. i hope i do. but there’s a good chance i won’t. and if i don’t, do you honestly believe that hillary would be tougher on putin than me? does anybody in this room really believe that? give me a break. (trump, 2017) 1. node a: is there an explicit qud? → yes, the reporter asked two overt questions. 2. node b: is something portrayed as an answer to that qud? → no, the overt questions are not directly addressed. 3. node f: does the speaker introduce a new qud? → yes, donald trump introduces several new quds, e.g. is it an asset that putin likes trump? and would hillary be tougher on putin?. 4. node g: is the new qud related to the original qud? → yes, the new quds all have to do with trump’s relationship with and punishment of putin. 5. node c: does it actually answer the qud? → no, e.g. the question aimed at particular sanctions by the previous president is not addressed. 6. node h: does the speaker want to appear as though they answered the qud? → yes, by implying that he will be tough(er than hillary), trump wants to connect to the original qud about “punish[ing] the russians”. path: ynyyny → possible bullshit (8) social media corona virus is temporary. house music is forever (matroda [@matrodamusic], 2020) 1. node a: is there an explicit qud? → no 2. node e: does the utterance imply a qud? → yes, the parallelism implies two quds about how long both the corona virus and house music will last. 3. node b: is something portrayed as an answer to that qud? → yes, the tweet author’s assertion both implies the quds and affirms them. 68 bullshit, pragmatic deception, and nlp figure 2: example of llm bullshit 4. node c: does it actually answer the qud? → yes. 5. node d: does the speaker believe they have sufficient evidence for their answer? → no, the speaker has no insight into the longevity of a novel virus and a style of music. path: nyyyn → possible bullshit (9) chatgpt the example of chatgpt bullshit in figure 2 was generated on january 9th 2023 with the december 15th release of chatgpt (https://chat.openai.com/).7 1. node a: is there an explicit qud? → yes, the prompt written by the author. 2. node b: is something portrayed as an answer to that qud? → yes, the chatbot provides three examples as asked. 3. node c: does it actually answer the qud? → yes, it directly – but incorrectly – answers the question. 4. node d: does the speaker believe they have sufficient evidence for their answer? → no, the system by its design cannot believe anything and simply outputs something that fits the shape of a perfect answer (following the reasoning of bender et al., 2021). path: yyyn → possible bullshit 7. chatgpt bullshit, arguably, could be seen as prompted and persuasive. the system has no mechanism to try to evade a question, rather it does not – because it cannot – care about the content as in persuasive bullshit. 69 deck these examples show that manual, in-depth analysis of common bullshitting examples is possible and some of the nodes may even be automated and handled with nlp methods. nodes a and b and f fall squarely into the space of question and answer detection. node e will mostly lead to yes for assertions, since most utterances at least imply a topic (and assertion detection is reasonably simple), which can then be identified via topic classification methods which also applies to node g. similarly, answers and newly introduced quds may be identified with computational methods. node c, does it actually answer the qud? is a classic nlp problem for question answering, but it is still far from solved. nodes d and h, though, are highly problematic from an nlp standpoint. the next chapter will discuss some limitations and caveats of this detailed annotation process for bullshit examples. 3.1.2 limitations and caveats while the qud approach to bullshit may at first glance seem like a promising avenue of research, there are still some issues: the first is that qud annotation is a fairly recent practice, though various guidelines for qud annotation have been provided (see de kuthy et al., 2018; riester et al., 2018; riester, 2019). second, even the simplified bullshit annotation process presented in the previous chapter leaves much to the interpretation of the annotators. discourse often depends on context and audience so propositions and current questions may be interpreted in any number of ways. as mentioned with example (1) and also noted by fredal: “audiences will differ in their response to a speaker’s statements and motives, some seeing truth and honesty where others see various degrees of bias, deception, and misinformation” (fredal, 2011, p. 20). it is doubtful whether a high interannotator agreement can be reached under such conditions, so the task may be one that has to rely on differing, subjective labels. third, and from a computational perspective maybe most concerning, are the kinds of nodes in the qud annotation flowchart that rely on language-external knowledge, especially does the speaker believe they have sufficient evidence for their answer? (node d) and to a certain degree does the speaker want to appear as though they answered the qud? (node h). while the latter may be supported by some evidence on the linguistic level (e.g. by the use of specific words or grammatical markers of deception), the former requires insight into the speakers’ minds. though there might be linguistic markers that hint at a speaker’s degree of evidence, in the majority of cases there should be no way of discerning it by looking at the text. especially when taking into account meibauer’s notion of ‘too much certainty,’ the whole point is that the linguistic markers (of certainty) do not match the speaker’s internal state of mind or depth of knowledge. we may hope to find solutions in annotating large amounts of texts containing bullshit, using linguistically and epistemologically trained annotators. nlp methods such as large language models or ‘foundation models’ (bommasani et al., 2021) may then be employed to find latent linguistic markers that serve as indicators of the authors’ internal state of mind – if such markers exist. however, it is doubtful that such an approach will work, especially with the kinds of short texts in social media or impromptu interview answers.8 a more fruitful endeavor, for now, may be trying to identify breaking points in pragmatic models in other, more straight-forward ways. one such practice of pragmatic deception – in the form of shifting – will be discussed in the next chapter. 8. in addition to a range of other concerns regarding the use of llms, as noted by bender et al. (2021) and others. 70 bullshit, pragmatic deception, and nlp 3.2 investigating shifting as a practical way of finding pragmatic deception approaches based on holistic pragmatic models may well remain in the area of theoretical linguistics, so for the foreseeable future, nlp research must focus on detecting specific bullshit subtypes. one example is shifting, a type of evasive answers introduced in an journalistic analysis of danish prime minister lars løkke rasmussen, whose was frequently criticized for evasive behavior in interviews (gabrielsen et al., 2020). rather than employing techniques shown in bull and mayer (1993) like verbally attacking the interviewer or questioning the question, rasmussen evades by shifting not the topic (as in greatbatch, 1986), but something else. though the authors do not use the term, shifting may be interpreted as evasive bullshitting, since the speaker “stays within the topical agenda set by the question [...] but refocuses to a more favorable aspect within that topic” (gabrielsen et al., 2020, p. 1355). the authors do not use the qud framework, but their concept of shifting fits well within the qud-based definition of bullshitting in (ii) in section 3.1: “when a shift is successfully executed, it will appear as if the interviewee answers the question when in fact the interviewee has steered clear of the critical aspect of the question thereby leaving the original question unanswered” (gabrielsen et al., 2020, p. 1356). fact-checking does not reveal shifting, because it does not involve giving false information or changing the topic. instead, the interviewee changes the focus of their answer, interpreting what is considered relevant within that topic (gabrielsen et al., 2020, p. 1356f). the authors identify three different types of shifts: shift of time, agent, and level, while acknowledging that there may be many more. shifts of time happen when the question concerns a specific time frame and the answer is shifted to a different one. when asked about specific economic policy in the past, e.g., the danish prime minister instead talked about “the present economic situation as well as future expectations for the economy” (gabrielsen et al., 2020, p. 1362). shifts of agent can involve broadening the agent, i.e. when rasmussen referenced his party, government or the country in answers to questions that are directed at his own opinions or actions (gabrielsen et al., 2020, p. 1363). agents can also be narrowed, when replying to questions like what is the party’s opinion on supporting families? with the agent shifted to a narrower scope like as a parent, i think that... (cf. gabrielsen et al., 2020, p. 1363). shifts of level lead to answers that are either more abstract or more concrete than asked for. most commonly, this occurred when the prime minister shifted a question about concrete politics to underlying, but more abstract ideological motives (gabrielsen et al., 2020, p. 1365). these types of shifts – and others yet to be identified – are clear examples of pragmatic, evasive bullshitting. instead of rejecting the question (or qud) outright, they reject “the journalist’s underlying intention with the question” (gabrielsen et al., 2020, p. 1367). that is, they imply novel quds misrepresenting them as pertaining to the original qud as described in definition (ii) in section 3.1. these shifts are difficult to detect in the moment and in the dataset, no journalist called the prime minister out on this behavior. nevertheless, the authors note a number of linguistic indicators of shifting, which is especially interesting from an nlp perspective, since evidence of this deceptive practice may be found in the text, without insight into the speaker’s mind. shifts of time may lead to mismatched temporal adjectives and grammatical tense in questions and corresponding answers, shifts of agents to mismatching pronouns or verb forms. shifts of level may manifest on the linguistic surface in the form of more concrete or abstract adjectives, verbs, nouns, etc. if they prove robust, such verbal 71 deck indicators could be leveraged for automatic classification by use of language models or other forms of machine learning, which makes investigating these markers vital for shifting annotation. 3.2.1 annotating shifts to investigate whether approaching (a subset of) bullshit from an nlp perspective becomes feasible when focusing on shifting, we created a small test corpus. the source for the corpus was the german parliament’s question hour (‘befragung der bundesregierung’)9, which is a prototypical bullshitting scenario, as it comprises situations in which persons are strongly expected to answer but may not want to do so for any number of reasons. 100 sample question answer pairs were selected from random parliamentary sessions of the past eight years. to ascertain whether the notion of evasive bullshitting in general and shifting in particular is intuitive, annotations were carried out by 26 untrained graduate students of german language and literature at the ruhr-university-bochum. the students were given the task of carefully reading the gabrielsen et al. paper, then randomly assigned 15 question-answer-pairs, and asked to mark whether a given pair contains any of the three shifts (time, agent, level), or any other shift. for each q&a pair, between three and five students handed in annotations, resulting in a total of 89 annotated pairs. for some examples, the annotation decisions were fairly straightforward. in example (10), it is immediately apparent that the speaker does not attempt to evade the question. instead, they openly admit that they cannot answer the question and refer it to someone else.10 in such a case, we need not even look at the question, since such an answer can neither be shifting nor any other kind of qud-based bullshit. in example (11), on the other hand, we need to look at the question to see that the answer contains a shift of level in talking about abstract concepts (i.e. we appreciate the united states as a country of democracy and that is the basis of international diplomacy) when the questions were rather specific (do you believe today in a joint and, above all, meaningful closing statement?).11 (10) da ich mich dort ehrlicherweise nicht mit den details auskenne, würde ich den zuständigen kollegen – wahrscheinlich ressort umwelt oder wirtschaft – bitten, entsprechend zu antworten. since i am honestly not familiar with the details there, i would ask the responsible colleague – probably department of environment or economy – to answer accordingly. (richter et al., 2020, id 1011958) (11) frage frau bundeskanzlerin, welchen sinn hat ein gipfeltreffen im rahmen der g 7, wenn sie und nahezu alle mitglieder ihrer bundesregierung den amerikanischen präsidenten fortgesetzt und auf allen öffentlichen kanälen diskreditieren? glauben sie heute angesichts der vor sich hergetragenen vorurteile an eine gemeinsame und vor allem sinnvolle abschlusserklärung? 9. data accessed via the open discourse corpus (richter et al., 2020). 10. of course, the answer could still be a lie if the speaker is in fact familiar with the details, but it is not a case of bullshitting. 11. the analysis is complicated by the fact that the original speaker asks leading questions that serve other pragmatic goals and which almost call for a shift in order to answer without damaging ones own position. whether there is such a thing as a ‘bullshit question’ remains to be seen (see also meibauer, 2016, p. 84). for further insight into types of questions in political settings, see zhang et al. (2017). 72 bullshit, pragmatic deception, and nlp und sind unter diesen umständen die wieder einmal massiven und wohl auch kostenintensiven sicherheitsmaßnahmen auf diesem gipfel gerechtfertigt? (richter et al., 2020, id 1007478) antwort sie haben eine hohe bandbreite an fragen, die sie aus ihrer fraktion heraus stellen, was das verhältnis zu den einzelnen ländern anbelangt. wir schätzen die vereinigten staaten als ein land der demokratie, aber wir sind trotzdem der meinung, dass da, wo meinungsverschiedenheiten bestehen, diese auch benannt werden müssen. aber gerade in solchen zeiten ist es eben einfach auch wichtig, immer wieder den gesprächsfaden zu suchen und überzeugungsarbeit zu leisten – darauf beruht internationale diplomatie –, und das tun wir in alle richtungen. (richter et al., 2020, id 1007479) question madam chancellor, what is the point of a g-7 summit meeting if you and almost all members of your federal government continue to discredit the american president on all public channels? do you believe today in a joint and, above all, meaningful closing statement in view of the prejudices you have displayed? and under these circumstances, are the once again massive and probably costly security measures at this summit justified? answer you have a large range of questions that you ask out of your parliamentary group in terms of the relationship with the individual countries. we appreciate the united states as a country of democracy, but we are nevertheless of the opinion that where there are differences of opinion, these must also be named. but it is precisely in times like these that it is important to keep trying to find the thread of dialogue and to persuade others – that is the basis of international diplomacy – and we do that in all directions. since no clear consensus was reached by the students for many of the qa pairs, all were annotated a second time by two trained annotators, including the author, discussing each pair in detail. this was done to investigate whether an in-depth analysis of each sample would lead to similar results to the more intuitive approach of the untrained students. since gabrielsen et al. only provided a very small number of examples, this annotation also served to investigate whether the concept of shifting is reproducible in expert annotation from a different field (linguistics as opposed to journalism). when distinguishing shifting from non-shifting behavior, the annotators often went back to the qud-based definition of (evasive) bullshitting as a basis of discussions. for example, in some cases the question (under discussion) was explicitly marked as unanswerable or not worth answering by the government member as in example (10). for some shifts, it was helpful to compare the overt qud in the question to the implicit one reconstructed from the answers, to see which elements are shifted. more difficult pairs also benefited from qud-based analysis along the lines of the flowchart in section 3.1.1, which provided clarification for otherwise more subjective judgments. 3.2.2 results, limitations and caveats students identified at least one kind of shift in most of the examples. using mace (hovy et al., 2013) as an analysis tool12, we found only 20 question-answer pairs without any shift, i.e. some 12. mace aggregates annotations to recover the most likely answer, calculates which annotators are trustworthy and evaluates item and task difficulty. more information can be found at https://github.com/dirkhovy/mace. 73 deck students might have annotated a shift in them, but the weighted consensus did not. most annotated shifts were of the ‘shift of agent’ category, with ‘shift of level’ being a close second. ‘shift of time’ was annotated to a lesser degree and while some students marked some ‘other’ shifts, these were never enough to lead to a consensus.13 still, in 69 q&a pairs the consensus identified at least one of the three shifts, confirming the assumption that the question hour is indeed a genre in which evasive bullshitting or shifting occurs. however, the pilot study indicates that a short primer on shifting is not enough to enable untrained annotators to come to a significant consensus on the various shift types. there is a reasonable annotator agreement that at least some kind of shift occurs in the majority of samples with a mean annotator competence of 0.68 when looking at whether at least one type was noted. unfortunately, there is significantly less overlap for the types of shifts, see table 1. average annotator competence as provided by mace was 0.53 for shift of time, 0.52 for shift of agent and 0.41 for shift of level. average agreement for the ‘other’ category of shifts was 0.85 since most students did not annotate any of these shifts at all. krippendorff alpha lay between 0 and 0.27 for all categories, indicating that annotators did not intuitively come to the same conclusions. one the one hand, the low agreement could indicate that the task is not intuitive and that more training is needed. however, it stands to reason that bullshitting is part of a group of phenomena that are inherently subjective. we might therefore be unable to calculate a gold standard since different people will view different statements as either bullshit or not, depending on extralinguistic factors such as their view of the speaker or their general attitude towards epistemic norms as outlined in plank (2022).14 bullshit classification may therefore benefit from taking into account different labels of all annotators since subjective information may get lost when using binary labelling. table 1: average mace annotator competence for shift annotation shift of time shift of agent shift of level other shift 0.53 0.52 0.41 0.85 as mentioned in section 3.2, shifting annotation favors nlp classification of bullshit phenomena mainly because of identifiable linguistic surface markers. however, during the in-depth annotation of all question answer pairs, challenges became apparent when relying on specific words. for example, shifts of agents were surmised to be apparent in the usage of pronouns, but when a question is addressed at you (in both english and formal german), the answer can contain either i or we. it thus depends on annotator discretion, whether the question aimed at singular you our plural you. similarly, there can be no indication of a shift of agent if the question does not contain a presumed agent. that is, questions like will the law be beneficial for...? can be answered with pronouns or nouns for any number of agents. we also encountered challenges when annotating shifts of level. since the bundestag data contains rather long answers to multiple questions at once, there is much more space to express both concrete and abstract concepts. meaning, while the question may be concrete, the answer can contain either abstract, concrete, or mixed concepts. concrete examples were used in combination 13. this may well be an issue of motivation for doing extra work, since the students that annotated the most ‘other’ shifts were also the ones that provided the most in-depth comments for their reasoning. 14. i am thankful for the anonymous reviewer who pointed out that some linguistic phenomena cannot and probably should not be reduced to a singular gold label. 74 bullshit, pragmatic deception, and nlp with abstract concepts to fully answer the question and give additional context. annotation was therefore much more complicated than just looking at whether or not abstract or concrete concepts were expressed linguistically in both question and answer or not. it therefore remains doubtful whether textual markers are enough to classify shifts and whether an automated system will come to the right conclusions simply on the basis of the utterances themselves. a simple word-based or otherwise rule-based classification system will not be able to ignore the extraneous and misleading tokens and deep learning and architectures might not be able to pick up on nuances that require intense discussion among human annotators simply from looking at the text. some preliminary findings are provided in the next chapter. 4. nlp approaches, conclusion & outlook 4.1 nlp methods for bullshit analysis a thorough investigation of nlp methods for bullshit detection is beyond the the scope of this paper and will be approached in future research. instead, two different possible pipelines for bullshit detection will be outlined in this chapter: one for persuasive and one for evasive bullshit. building on the qud-based definition of bullshitting, we can make use of established nlp methods that deal with questions in a broader sense. yet we still face challenges in decisions human annotators make based on context, world knowledge and intuition about the producers of bullshit. the following pipelines can thus only serve as a starting point of classification for some types of bullshitting behaviors. others remain dependant on human intervention, possibly as a final step following an automated pre-selection of bullshit candidates using nlp methods. 4.1.1 nlp-based detection of persuasive bullshit since persuasive bullshitting typically occurs unprompted, i.e. not following an overt question, the first step in the pipeline is claim detection. this subtask is vital for fake news detection and related fields, so we can build on a vast amount of research (e.g. levy et al., 2014; hassan et al., 2015; lippi and torroni, 2015; barrón-cedeño et al., 2020; konstantinovskiy et al., 2021). next, the system must infer the implicit questions, building on qud-specific research by de kuthy et al. (2020), but also on a large amount of literature on question generation for a variety of purposes (heilman and smith, 2010; du et al., 2017; duan et al., 2017; zhou et al., 2018; kurdi et al., 2020). the third step in persuasive bullshit detection determines whether the generated questions are ‘unanswerable’ as in example (1) about a person’s degree of racism. a rule-based approach might suffice for many cases, based on identifying vague or unquantifiable concepts about which one cannot make assertions with sufficient certainty. another approach could utilize neural networks either by training on corpora of unanswerable questions prepared by annotators, or by using the intuition of a large language model. while systems like chatgpt may themselves be bullshitters – at least when it comes to content – they can show surprising accuracy on what is quantifiable: when asked, chatgpt tells us that weight is quantifiable while racism is not. whether we can actually rely on a system that only ‘knows’ about form to give sufficiently accurate accounts of answerability requires in-depth research in the future. finally, there are those cases of persuasive bullshitting in which the concept is not unquantifiable. here, nlp methods reach their limits and human annotators will be needed to make the final decision on whether the utterance can be called bullshit. there exist nlp methods to figure out 75 deck whether the assertion fits the reconstructed implicit question, e.g. by calculating the question-answer congruence (faruqui and das, 2018; yu and jiang, 2021). in this, we must be extremely cautious to avoid circular reasoning: if the question is generated based on a piece of text, the piece of text must in some way answer it. whether it ‘really’ fits may require linguistic and world-knowledge intuition on the part of the annotators. nlp pipeline for persuasive bs detect claims a generate implicit questions b assess answerability of questions c optional: calculate qa congruence d figure 3: nlp-based detection of persuasive bullshit 4.1.2 nlp-based detection of evasive bullshit since evasive bullshit typically occurs in prompted situations, the first step is not claim detection but rather identifying the answer(s) to overt questions. next, topic classification methods (e.g. wang and manning, 2012; quercia et al., 2012; pappagari et al., 2019) may be used on both question and answer to check if their topics match. if they do not, we can rule out evasive bullshitting, where only aspects of the question are shifted, but not the whole topic as in greatbatch (1986). differing topics could also stem from other types of uncooperative behavior, as shown in bull and mayer (1993), or simple misunderstandings. implicit quds can then be generated from the question, analogously to the detection of persuasive bullshit above. a number of nlp methods, e.g. adapted from various text similarity measures, can then be employed to compare the reconstructed quds to the overt question to see whether and where they overlap (croft et al., 2013; prasetya et al., 2018; prakoso et al., 2021). evasive bullshitting by way of introducing novel quds (and pretending they answer the original one) occurs when the topic of question and answer matches, but the implicit quds differ strongly from the overt ones. when the implicit qud is very similar to the overt question, yet changed in agent, time or level of abstraction, it may constitute a case of shifting as described in section 3.2. nlp pipeline for evasive bs detect answers to overt questions a assess topics of question and answer b generate implicit qud(s) c compare overt question(s) to qud(s) d figure 4: nlp-based detection of evasive bullshit 76 bullshit, pragmatic deception, and nlp 4.2 conclusion if we want to deal with subjective and amorphous subjects such as bullshit computationally, we need to find robust definitions. we need to approach smaller, more manageable subtypes of bullshit to create corpora and build automated systems. however, finding such narrow definitions remains a challenge. as shown above, even picking only three specific sub-concepts, the three shifts, of only one type of bullshitting (evasive bullshit) does not facilitate straightforward annotation. the cooperative annotation with trained annotators provided better results, but it is also extremely timeconsuming. pragmatic theories of bullshitting serve as starting points from which we can pick parts to be tackled with nlp methods. annotation guidelines that cover the whole spectrum of bullshit suffer from their reliance on language-external factors. these environmental aspects of specific discourse situations and speaker-internal mental processes are exceedingly hard to grasp solely on the basis of text. according to the prevailing theories on the language of deception and lying, at least some of these internal workings are expressed linguistically (see e.g. meibauer, 2019). it remains unclear to what degree this notion applies to the language of bullshit. the in-depth shifting annotation with two trained annotators provided insights into this specific subtype of evasive language and that whether or not a shift can occur strongly depends on the type of question. some questions permitted fewer shifts, e.g. when politicians are simply asked about their opinion on some topic, their answer may refer to past, present or future, so a shift of time is improbable. shifts of time are still possible, e.g. by answering in the past i was of the opinion that..., but this would be a very unusual, highly marked answer in which the shift became so obvious as to be useless to the shifter. research strongly suggests that bullshit exists outside of the conventional truth/lie spectrum. evasive bullshitting in the form of shifting can be found in politicians’ answers; other forms of evasive bullshitting can be found anywhere from exam situations to post-game interviews – bullshit “is unavoidable whenever circumstances require someone to talk without knowing what he is talking about” (frankfurt, 2005, p. 15). disciplines ranging from journalism to psychology have provided evidence which strongly points toward the existence of some sort of apathy towards ones answers or some sort of emptiness with regard to a qud, which speakers may deceive about. they do not hide the truth value of a factual concept, but rather hide their attitude towards the conventions of discourse itself. instead of a simple disregard of truth, there is a disregard for the structures and norms of human communication which places linguistic bullshit strongly in the space of pragmatics. choosing questions under discussion as a starting point – mainly because of existing nlp research connected to questions – this paper has focused on three research questions. rq1, what is a useful bullshit definition for a computational linguistics approach?, was answered in section 3.1 by providing an extended, qud-based definition that builds on stokke and fallis (2017) and covers both persuasive and evasive bullshitting. in sections 3.1.1 and 3.2, two different approaches were provided to answer rq2, how can we annotate and analyze bullshit?. finally, rq3, what are promising nlp approaches to bullshit classification?, was tentatively answered in section 4.1. two nlp pipelines for persuasive and evasive bullshitting detection were outlined to serve as a basis for future work. 77 deck 4.3 outlook further work is needed in identifying evasive bullshitting in prompted situations, as well as finding novel ways to tackle persuasive, unprompted bullshit. advances in computational qud analysis facilitate the identification of underlying questions for assertions of any kind, including bullshit statements. the advantage of relying on quds as an established pragmatic model of communication is the model’s relationship with (overt) questions. linguistically, bullshit detection can also be based on other frameworks, which may be more apt to model uncooperative dialog (such as asher et al., 2017; asher and paul, 2018). however, if it is possible to identify underlying questions and their answers in bullshitting texts, we may benefit from such nlp tasks as question answering or answer quality estimation. the qud framework therefore makes the novel task of computational bullshit detection somewhat more approachable. future work will focus on creating corpora of different types of bullshitting behavior and using the outlined nlp pipelines for bullshit detection in the coming years, bullshit detection will become vital from a practical perspective: question answering system found in popular speech assistants like apple’s siri or amazon’s alexa, for example, may deem an answer as not sufficient to the speaker’s intent (as opposed to being plain wrong). especially with the recent explosion of interest in llm-based text generation, the concept of a bullshit answer or statement becomes increasingly important. with ethical discussions about the biases inherent in the data (see e.g. raji et al., 2020; bender et al., 2021; gebru et al., 2021), philosophical questions of what these llms can ‘know’ when learning only on text, and practical considerations of reliability, the future of text generation will be treacherous if these systems turn out to be plain bullshitters. 5. acknowledgements i would like to thank the student annotators of shifting examples, including julius kirschner who helped with the in-depth annotation. i would also like to thank my anonymous reviewers and the board of dialogue & discourse for their invaluable feedback. finally, i give thanks to my supervisor tatjana scheffler for her input and support during many discussions of bullshitting phenomena. references nicholas asher and soumya paul. strategic conversations under imperfect information: epistemic message exchange games. journal of logic, language and information, 27(4):343– 385, december 2018. issn 0925-8531, 1572-9583. doi: 10.1007/s10849-018-9271-9. url http://link.springer.com/10.1007/s10849-018-9271-9. nicholas asher, soumya paul, and antoine venant. message exchange games in strategic contexts. journal of philosophical logic, 46(4):355–404, august 2017. issn 0022-3611, 15730433. doi: 10.1007/s10992-016-9402-1. url http://link.springer.com/10.1007/ s10992-016-9402-1. isabelle augenstein. towards explainable fact checking. doctor scientiarum (dr. scient.), university of copenhagen, december 2021. url http://arxiv.org/abs/2108.10274. alberto barrón-cedeño, giovanni da san martino, israa jaradat, and preslav nakov. proppy: a system to unmask propaganda in online news. proceedings of the aaai conference on artificial 78 bullshit, pragmatic deception, and nlp intelligence, 33:9847–9848, july 2019a. issn 2374-3468, 2159-5399. doi: 10.1609/aaai.v33i01. 33019847. url https://www.aaai.org/ojs/index.php/aaai/article/view/ 5061. alberto barrón-cedeño, israa jaradat, giovanni da san martino, and preslav nakov. proppy: organizing the news based on their propagandistic content. information processing & management, 56(5):1849–1864, september 2019b. issn 03064573. doi: 10.1016/j.ipm.2019.03.005. url https://dx.doi.org/10.1016/j.ipm.2019.03.005. alberto barrón-cedeño, tamer elsayed, preslav nakov, giovanni da san martino, maram hasanain, reem suwaileh, fatima haouari, nikolay babulkov, bayan hamdan, alex nikolov, shaden shaar, and zien sheikh ali. overview of checkthat! 2020: automatic identification and verification of claims in social media. in avi arampatzis, evangelos kanoulas, theodora tsikrika, stefanos vrochidis, hideo joho, christina lioma, carsten eickhoff, aurélie névéol, linda cappellato, and nicola ferro, editors, experimental ir meets multilinguality, multimodality, and interaction, lecture notes in computer science, pages 215–236, cham, 2020. springer international publishing. isbn 978-3-030-58219-7. doi: 10.1007/978-3-030-58219-7 17. url https://dx.doi.org/10.1007/978-3-030-58219-7_17. emily m. bender, timnit gebru, angelina mcmillan-major, and shmargaret shmitchell. on the dangers of stochastic parrots: can language models be too big? in proceedings of the 2021 acm conference on fairness, accountability, and transparency, pages 610–623, virtual event canada, march 2021. acm. isbn 978-1-4503-8309-7. doi: 10.1145/3442188.3445922. url https://dl.acm.org/doi/10.1145/3442188.3445922. carl t. bergstrom and jevin d. west. calling bullshit: the art of skepticism in a data-driven world. random house, new york, first edition edition, 2020. isbn 978-0-525-50918-9 978-0593-22976-7. url https://www.callingbullshit.org/. josh bernoff. chatgpt is a bullshitter, december 2022. url https://withoutbullshit. com/blog/chatgpt-is-a-bullshitter. eric nenkia bien. how obscurantism differs from bullshit: a proposal. theoria, 87(6):1497– 1526, december 2021. issn 0040-5825, 1755-2567. doi: 10.1111/theo.12354. url https: //onlinelibrary.wiley.com/doi/10.1111/theo.12354. rishi bommasani, drew a. hudson, e. adeli, r. altman, simran arora, sydney von arx, michael s. bernstein, jeannette bohg, antoine bosselut, emma brunskill, e. brynjolfsson, s. buch, dallas card, rodrigo castellon, niladri s. chatterji, annie s. chen, kathleen creel, jared davis, dora demszky, chris donahue, moussa doumbouya, esin durmus, s. ermon, j. etchemendy, kawin ethayarajh, l. fei-fei, chelsea finn, trevor gale, lauren e. gillespie, karan goel, noah d. goodman, s. grossman, neel guha, tatsunori hashimoto, peter henderson, john hewitt, daniel e. ho, jenny hong, kyle hsu, jing huang, thomas f. icard, dan jurafsky, saahil jain, pratyusha kalluri, siddharth karamcheti, g. keeling, fereshte khani, o. khattab, pang wei koh, m. krass, ranjay krishna, rohith kuditipudi, ananya kumar, faisal ladhak, mina lee, tony lee, j. leskovec, isabelle levent, xiang lisa li, xuechen li, tengyu ma, ali malik, christopher d. manning, suvir p. mirchandani, eric mitchell, zanele munyikwa, suraj nair, a. narayan, d. narayanan, benjamin newman, allen nie, juan carlos niebles, 79 deck h. nilforoshan, j. nyarko, giray ogut, laurel orr, isabel papadimitriou, j. park, c. piech, eva portelance, christopher potts, aditi raghunathan, robert reich, hongyu ren, frieda rong, yusuf h. roohani, camilo ruiz, jack ryan, christopher r’e, dorsa sadigh, shiori sagawa, keshav santhanam, andy shih, k. srinivasan, alex tamkin, rohan taori, a. thomas, florian tramèr, rose e. wang, william wang, bohan wu, jiajun wu, yuhuai wu, sang michael xie, michihiro yasunaga, jiaxuan you, m. zaharia, michael zhang, tianyi zhang, xikun zhang, yuhui zhang, lucia zheng, kaitlyn zhou, and percy liang. on the opportunities and risks of foundation models, august 2021. url https://fsi.stanford.edu/publication/ opportunities-and-risks-foundation-models. peter bull and kate mayer. how not to answer questions in political interviews. political psychology, 14:651, december 1993. doi: 10.2307/3791379. thomas l. carson. frankfurt and cohen on bullshit, bullshiting, deception, lying, and concern with the truth of what one says. pragmatics & cognition, 23(1):53–67, september 2016. issn 09290907, 1569-9943. doi: 10.1075/pc.23.1.03car. url https://dx.doi.org/10.1075/ pc.23.1.03car. quassim cassam. the bullshit industry, november 2019. url https://www. quassimcassam.com/talks. eunsol choi, he he, mohit iyyer, mark yatskar, wen-tau yih, yejin choi, percy liang, and luke zettlemoyer. quac: question answering in context. in proceedings of the 2018 conference on empirical methods in natural language processing, pages 2174–2184, brussels, belgium, october 2018. association for computational linguistics. doi: 10.18653/v1/d18-1241. url https://aclanthology.org/d18-1241. gerald allen cohen. deeper into bullshit. in contours of agency: essays on themes from harry frankfurt, pages 321–339. mit press, 2002. isbn 978-0-262-52813-9. url https: //mitpress.mit.edu/books/contours-agency. david croft, simon coupland, jethro shell, and stephen brown. a fast and efficient semantic short text similarity metric. in 2013 13th uk workshop on computational intelligence (ukci), pages 221–227, september 2013. doi: 10.1109/ukci.2013.6651309. giovanni da san martino, chris brew, giovanni luca ciampaglia, anna feldman, chris leberknight, and preslav nakov. proceedings of the 3rd nlp4if workshop on nlp for internet freedom: censorship, disinformation, and propaganda. in proceedings of the 3rd nlp4if workshop on nlp for internet freedom: censorship, disinformation, and propaganda, barcelona, spain (online), december 2020a. isbn isbn 978-1-952148-36-1ii. url https://aclanthology.org/volumes/2020.nlp4if-1/. giovanni da san martino, stefano cresci, alberto barrón-cedeño, seunghak yu, roberto di pietro, and preslav nakov. a survey on computational propaganda detection. in proceedings of the twenty-ninth international joint conference on artificial intelligence, pages 4826–4832, yokohama, japan, july 2020b. international joint conferences on artificial intelligence organization. isbn 978-0-9992411-6-5. doi: 10.24963/ijcai.2020/672. url https: //dx.doi.org/10.24963/ijcai.2020/672. 80 bullshit, pragmatic deception, and nlp kordula de kuthy, nils reiter, and arndt riester. qud-based annotation of discourse structure and information structure: tool and evaluation. in proceedings of the eleventh international conference on language resources and evaluation (lrec 2018), miyazaki, japan, january 2018. european language resources association (elra). url https://aclanthology. org/l18-1304. kordula de kuthy, madeeswaran kannan, haemanth santhi ponnusamy, and detmar meurers. towards automatically generating questions under discussion to link information and discourse structure. in proceedings of the 28th international conference on computational linguistics, pages 5786–5798, barcelona, spain (online), 2020. international committee on computational linguistics. doi: 10.18653/v1/2020.coling-main.509. url https://www.aclweb.org/ anthology/2020.coling-main.509. xinya du, junru shao, and claire cardie. learning to ask: neural question generation for reading comprehension. in proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: long papers), pages 1342–1352, vancouver, canada, july 2017. association for computational linguistics. doi: 10.18653/v1/p17-1123. url https://aclanthology.org/p17-1123. nan duan, duyu tang, peng chen, and ming zhou. question generation for question answering. in proceedings of the 2017 conference on empirical methods in natural language processing, pages 866–874, copenhagen, denmark, september 2017. association for computational linguistics. doi: 10.18653/v1/d17-1090. url https://aclanthology.org/d17-1090. manaal faruqui and dipanjan das. identifying well-formed natural language questions. in proceedings of the 2018 conference on empirical methods in natural language processing, pages 798–803, brussels, belgium, october 2018. association for computational linguistics. doi: 10.18653/v1/d18-1091. url https://aclanthology.org/d18-1091. harry g. frankfurt. on bullshit. princeton university press, princeton, nj, 2005. isbn 978-0691-12294-6. url https://www.jstor.org/stable/j.ctt7t4wr. james fredal. rhetoric and bullshit. college english, 73(3):243–259, 2011. issn 0010-0994. url https://www.jstor.org/stable/25790474. jonas gabrielsen, heidi jønch-clausen, and christina pontoppidan. answering without answering: shifting as an evasive rhetorical strategy. journalism, 21(9):1355–1370, september 2020. issn 1464-8849. doi: 10.1177/1464884917738412. url https://doi.org/10.1177/ 1464884917738412. josé ángel gascón. argumentative bullshit. informal logic, 41(3):289–308, september 2021. issn 2293-734x, 0824-2577. doi: 10.22329/il.v41i3.6838. url https://dx.doi.org/ 10.22329/il.v41i3.6838. timnit gebru, jamie morgenstern, briana vecchione, jennifer wortman vaughan, hanna wallach, hal daumé iii, and kate crawford. datasheets for datasets. communications of the acm, 64 (12):86–92, november 2021. issn 0001-0782. doi: 10.1145/3458723. url https://doi. org/10.1145/3458723. 81 deck jonathan ginzburg. an update semantics for dialogue. in proceedings of the 1st international workshop on computational semantics, tilburg: itk, tilburg university, 1994. jonathan ginzburg. dynamics and the semantics of dialogue. logic, language and computation, 1: 221–237, 1996. jonathan ginzburg. the interactive stance: meaning for conversation. oxford university press, january 2012. isbn 978-0-19-163248-8. jonathan ginzburg, zulipiye yusupujiang, chuyuan li, kexin ren, aleksandra kucharska, and pawel lupkowski. characterizing the response space of questions: data and theory. dialogue & discourse, 13(2):79–132, december 2022. issn 2152-9620. doi: 20221220143609000. url https://ojs3-prod.lib.uic.edu/ojs/index.php/ dad/article/view/11531. david graeber. on the phenomenon of bullshit jobs: a work rant. strike magazine, 3:1–5, august 2013. url https://www.strike.coop/bullshit-jobs/. david greatbatch. aspects of topical organization in news interviews: the use of agendashifting procedures by interviewees. media, culture & society, 8(4):441–455, october 1986. issn 0163-4437, 1460-3675. doi: 10.1177/0163443786008004005. url http: //journals.sagepub.com/doi/10.1177/0163443786008004005. zhijiang guo, michael schlichtkrull, and andreas vlachos. a survey on automated factchecking. transactions of the association for computational linguistics, 10:178–206, february 2022. issn 2307-387x. doi: 10.1162/tacl a 00454. url https://doi.org/10.1162/ tacl_a_00454. naeemul hassan, chengkai li, and mark tremayne. detecting check-worthy factual claims in presidential debates. in proceedings of the 24th acm international on conference on information and knowledge management cikm ’15, pages 1835–1838, melbourne, australia, 2015. acm press. isbn 978-1-4503-3794-6. doi: 10.1145/2806416.2806652. url http://dl.acm.org/citation.cfm?doid=2806416.2806652. michael heilman and noah a. smith. good question! statistical ranking for question generation. in human language technologies: the 2010 annual conference of the north american chapter of the association for computational linguistics, hlt ’10, pages 609–617, usa, june 2010. association for computational linguistics. isbn 978-1-932432-65-7. dirk hovy, taylor berg-kirkpatrick, ashish vaswani, and eduard hovy. learning whom to trust with mace. in proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 1120–1130, atlanta, georgia, june 2013. association for computational linguistics. url https://aclanthology.org/n13-1132. viktor ivanković. steering clear of bullshit? the problem of obscurantism. philosophia, 44(2): 531–546, june 2016. issn 0048-3893, 1574-9274. doi: 10.1007/s11406-016-9709-8. url https://dx.doi.org/10.1007/s11406-016-9709-8. 82 bullshit, pragmatic deception, and nlp najoung kim, ellie pavlick, burcu karagol ayan, and deepak ramachandran. which linguist invented the lightbulb? presupposition verification for question-answering, january 2021. url http://arxiv.org/abs/2101.00391. lev konstantinovskiy, oliver price, mevan babakar, and arkaitz zubiaga. toward automated factchecking: developing an annotation schema and benchmark for consistent automated claim detection. digital threats: research and practice, 2(2):1–16, june 2021. issn 26921626, 2576-5337. doi: 10.1145/3412869. url https://dl.acm.org/doi/10.1145/ 3412869. jan van kuppevelt. discourse structure, topicality and questioning. journal of linguistics, 31(1): 109–147, march 1995. issn 1469-7742, 0022-2267. doi: 10.1017/s002222670000058x. url https://www.cambridge.org/core/journals/journal-of-linguistics/ article/abs/discourse-structure-topicality-and-questioning/ 60f3e68601517091af560cb3cc02c6ac. ghader kurdi, jared leo, bijan parsia, uli sattler, and salam al-emari. a systematic review of automatic question generation for educational purposes. international journal of artificial intelligence in education, 30(1):121–204, march 2020. issn 1560-4306. doi: 10.1007/ s40593-019-00186-y. url https://doi.org/10.1007/s40593-019-00186-y. ran levy, yonatan bilu, daniel hershcovich, ehud aharoni, and noam slonim. context dependent claim detection. in proceedings of coling 2014, the 25th international conference on computational linguistics: technical papers, pages 1489–1500, dublin, ireland, august 2014. dublin city university and association for computational linguistics. url https://aclanthology.org/c14-1141. marco lippi and paolo torroni. context-independent claim detection for argument mining. in twenty-fourth international joint conference on artificial intelligence, june 2015. url https://www.aaai.org/ocs/index.php/ijcai/ijcai15/paper/view/ 10942. shane littrell, evan f. risko, and jonathan a. fugelsang. the bullshitting frequency scale: development and psychometric properties. british journal of social psychology, 60(1):248– 270, january 2021a. issn 0144-6665, 2044-8309. doi: 10.1111/bjso.12379. url https: //onlinelibrary.wiley.com/doi/10.1111/bjso.12379. shane littrell, evan f. risko, and jonathan a. fugelsang. ‘you can’t bullshit a bullshitter’ (or can you?): bullshitting frequency predicts receptivity to various types of misleading information. british journal of social psychology, page bjso.12447, february 2021b. issn 0144-6665, 20448309. doi: https://dx.doi.org/10.1111/bjso.12447. url https://dx.doi.org/10.1111/ bjso.12447. paweł łupkowski and jonathan ginzburg. query responses. journal of language modelling, 4(2): 245–292, 2016. matroda [@matrodamusic]. corona virus is temporary. house music is forever, march 2020. url https://twitter.com/matrodamusic/status/1235617743161815041. 83 deck daniel mears. the ubiquity, functions, and contexts of bullshitting. journal of mundane behavior, 3(2):233–256, june 2002. url https://www.researchgate.net/publication/ 289724155_the_ubiquity_functions_and_contexts_of_bullshitting. jörg meibauer. aspects of a theory of bullshit. pragmatics & cognition, 23(1):68–91, september 2016. issn 0929-0907, 1569-9943. doi: 10.1075/pc.23.1.04mei. url http://www. jbe-platform.com/content/journals/10.1075/pc.23.1.04mei. jörg meibauer. the linguistics of lying. annual review of linguistics, 4 (1):357–375, january 2018. issn 2333-9683, 2333-9691. doi: 10.1146/ annurev-linguistics-011817-045634. url http://www.annualreviews.org/doi/ 10.1146/annurev-linguistics-011817-045634. jörg meibauer, editor. the oxford handbook of lying. oxford handbooks in linguistics. oxford university press, oxford, united kingdom, first edition edition, 2019. isbn 978-0-19-873657-8. url https://global.oup.com/academic/product/ the-oxford-handbook-of-lying-9780198736578?cc=de&lang=en&. jörg meibauer. sprache und bullshit. universitätsverlag winter heidelberg, 2020. isbn 978-3-8253-4808-3. url https://www.winter-verlag.de/de/detail/ 978-3-8253-4808-3/meibauer_sprache_und_bullshit/. chandra mukerji. bullshitting: road lore among hitchhikers*. social problems, 25(3):241–252, february 1978. issn 0037-7791. doi: 10.2307/800062. url https://dx.doi.org/10. 2307/800062. arvind narayanan and sayash kapoor. chatgpt is a bullshit generator. but it can still be amazingly useful, december 2022. url https://aisnakeoil.substack.com/p/ chatgpt-is-a-bullshit-generator-but. condé nast. chatgpt’s fluent bs is compelling because everything is fluent bs. wired uk, 2022. issn 1357-0978. url https://www.wired.co.uk/article/ chatgpt-fluent-bs. openai. chatgpt: optimizing language models for dialogue, november 2022. url https: //openai.com/blog/chatgpt/. raghavendra pappagari, piotr zelasko, jesús villalba, yishay carmiel, and najim dehak. hierarchical transformers for long document classification. in 2019 ieee automatic speech recognition and understanding workshop (asru), pages 838–844, december 2019. doi: 10.1109/asru46091.2019.9003958. gordon pennycook and david g. rand. lazy, not biased: susceptibility to partisan fake news is better explained by lack of reasoning than by motivated reasoning. cognition, 188:39–50, 2018. issn 00100277. doi: 10.1016/j.cognition.2018.06.011. url https://dx.doi.org/10. 1016/j.cognition.2018.06.011. gordon pennycook and david g. rand. who falls for fake news? the roles of bullshit receptivity, overclaiming, familiarity, and analytic thinking. journal of personality, 88(2):185–200, april 84 bullshit, pragmatic deception, and nlp 2020. issn 0022-3506, 1467-6494. doi: 10.1111/jopy.12476. url https://dx.doi. org/10.1111/jopy.12476. gordon pennycook and david g. rand. the psychology of fake news. trends in cognitive sciences, 25(5):388–402, may 2021. issn 13646613. doi: 10.1016/j.tics.2021.02.007. url https://dx.doi.org/10.1016/j.tics.2021.02.007. gordon pennycook, james allan cheyne, nathaniel barr, jonathan a fugelsang, and derek j koehler. on the reception and detection of pseudo-profound bullshit. judgment and decision making, 10(6):549–563, november 2015. url https://psycnet.apa.org/record/ 2015-54494-003. john v. petrocelli. the life-changing science of detecting bullshit. st. martin’s press, new york, first edition edition, 2021. isbn 978-1-250-27162-4 978-1-25028015-2. url https://us.macmillan.com/books/9781250271624/ thelifechangingscienceofdetectingbullshit. barbara plank. the “problem” of human label variation: on ground truth in data, modeling and evaluation. in proceedings of the 2022 conference on empirical methods in natural language processing, pages 10671–10682, abu dhabi, united arab emirates, december 2022. association for computational linguistics. url https://aclanthology.org/2022. emnlp-main.731. livia polanyi. a formal model of the structure of discourse. journal of pragmatics, 12(5):601– 638, december 1988. issn 0378-2166. doi: 10.1016/0378-2166(88)90050-1. url https: //www.sciencedirect.com/science/article/pii/0378216688900501. dimas wibisono prakoso, asad abdi, and chintan amrit. short text similarity measurement methods: a review. soft computing, 25(6):4699–4723, march 2021. issn 1433-7479. doi: 10.1007/ s00500-020-05479-2. url https://doi.org/10.1007/s00500-020-05479-2. didik dwi prasetya, aji prasetya wibawa, and tsukasa hirashima. the performance of text similarity algorithms. international journal of advances in intelligent informatics, 4(1):63, march 2018. issn 2548-3161, 2442-6571. doi: 10.26555/ijain.v4i1.152. url http: //ijain.org/index.php/ijain/article/view/152. daniele quercia, harry askham, and jon crowcroft. tweetlda: supervised topic classification and link prediction in twitter. in proceedings of the 4th annual acm web science conference, websci ’12, pages 247–250, new york, ny, usa, june 2012. association for computing machinery. isbn 978-1-4503-1228-8. doi: 10.1145/2380718.2380750. url https: //doi.org/10.1145/2380718.2380750. inioluwa deborah raji, andrew smart, rebecca n. white, margaret mitchell, timnit gebru, ben hutchinson, jamila smith-loud, daniel theron, and parker barnes. closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing. in proceedings of the 2020 conference on fairness, accountability, and transparency, fat* ’20, pages 33–44, new york, ny, usa, january 2020. association for computing machinery. isbn 978-1-45036936-7. doi: 10.1145/3351095.3372873. url https://doi.org/10.1145/3351095. 3372873. 85 deck florian richter, philipp koch, oliver franke, jakob kraus, fabrizio kuruc, anja thiem, judith högerl, stella heine, and konstantin schöps. open discourse, 2020. url https://doi. org/10.7910/dvn/fikibo. arndt riester. constructing qud trees. in malte zimmermann, klaus von heusinger, and v.edgar onea gaspar, editors, questions in discourse, chapter questions in discourse, pages 164–193. brill, march 2019. isbn 978-90-04-37832-2. doi: 10.1163/9789004378322 007. url https: //brill.com/view/book/edcoll/9789004378322/bp000006.xml. arndt riester, lisa brunetti, and kordula kuthy. annotation guidelines for questions under discussion and information structure. hal, pages 1–56, 2018. doi: 10.1075/slcs.199.14rie. url https://hal.archives-ouvertes.fr/hal-01794160. craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. semantics and pragmatics, 5:1–69, december 2012. issn 1937-8912. doi: 10.3765/sp.5.6. url http://semprag.org/article/view/sp.5.6. carl sagan. the fine art of baloney detection, 1996. url https://fermatslibrary. com/s/the-fine-art-of-baloney-detection. john r. searle. a classification of illocutionary acts. language in society, 5(1):1–23, 1976. issn 0047-4045. doi: 10.1017/s0047404500006837. url https://www.jstor.org/ stable/4166848. gautam kishore shahi, julia maria struß, and thomas mandl. overview of the clef-2021 checkthat! lab: task 3 on fake news detection. in ceur workshop proceedings, volume 2936, page 18, bucharest, romania, 2021. ceur-ws. url http://ceur-ws.org/vol-2936/ paper-30.pdf. gabriella skitalinskaya, jonas klaff, and henning wachsmuth. learning from revisions: quality assessment of claims in argumentation at scale. in proceedings of the 16th conference of the european chapter of the association for computational linguistics, page 12, 2021. doi: 10.18653/v1/2021.eacl-main.147. url https://dx.doi.org/10.18653/v1/2021. eacl-main.147. andré spicer. playing the bullshit game: how empty and misleading communication takes over organizations. organization theory, 1:1–26, 2020. doi: 10.1177/2631787720929704. url https://dx.doi.org/10.1177/2631787720929704. andreas stokke and don fallis. bullshitting, lying, and indifference toward truth. ergo, an open access journal of philosophy, 4(10):277–309, june 2017. issn 2330-4014. doi: 10.3998/ergo. 12405314.0004.010. url https://dx.doi.org/10.3998/ergo.12405314.0004. 010. donald trump. donald trump’s news conference: full transcript and video, january 2017. url https://www.nytimes.com/2017/01/11/us/politics/ trump-press-conference-transcript.html. donald trump. trump: ’i’m the least racist person anybody is going to meet’ bbc news, january 2018. url https://www.bbc.com/news/av/uk-42830165. 86 bullshit, pragmatic deception, and nlp christiane von stutterheim and wolfgang klein. referential movement in descriptive and narrative discourse. in rainer dietrich and carl f. graumann, editors, north-holland linguistic series: linguistic variations, volume 54 of language processing in social context, pages 39– 76. elsevier, january 1989. doi: 10.1016/b978-0-444-87144-2.50005-7. url https://www. sciencedirect.com/science/article/pii/b9780444871442500057. sida wang and christopher d. manning. baselines and bigrams: simple, good sentiment and topic classification. in proceedings of the 50th annual meeting of the association for computational linguistics: short papers volume 2, acl ’12, pages 90–94, usa, july 2012. association for computational linguistics. matthijs westera, laia mayol, and hannah rohde. ted-q: ted talks and the questions they evoke. in proceedings of the 12th conference on language resources and evaluation (lrec 2020), pages 1118–1127, marseille, 2020. european language resources association (elra). url https://aclanthology.org/2020.lrec-1.141/. xiaojing yu and anxiao jiang. expanding, retrieving and infilling: diversifying cross-domain question generation with flexible templates. in proceedings of the 16th conference of the european chapter of the association for computational linguistics, page 11, 2021. doi: 10.18653/v1/2021.eacl-main.279. url https://dx.doi.org/10.18653/v1/2021. eacl-main.279. xia zeng, amani s. abumansour, and arkaitz zubiaga. automated fact-checking: a survey. language and linguistics compass, 15(10):e12438, 2021. issn 1749-818x. doi: 10. 1111/lnc3.12438. url https://onlinelibrary.wiley.com/doi/abs/10.1111/ lnc3.12438. justine zhang, arthur spirling, and cristian danescu-niculescu-mizil. asking too much? the rhetorical role of questions in political discourse. arxiv:1708.02254 [physics], august 2017. url http://arxiv.org/abs/1708.02254. qingyu zhou, nan yang, furu wei, chuanqi tan, hangbo bao, and ming zhou. neural question generation from text: a preliminary study. in xuanjing huang, jing jiang, dongyan zhao, yansong feng, and yu hong, editors, natural language processing and chinese computing, lecture notes in computer science, pages 662–671, cham, 2018. springer international publishing. isbn 978-3-319-73618-1. doi: 10.1007/978-3-319-73618-1 56. xinyi zhou and reza zafarani. a survey of fake news: fundamental theories, detection methods, and opportunities. acm computing surveys, 53(5):1–40, october 2020. issn 0360-0300, 15577341. doi: 10.1145/3395046. url http://arxiv.org/abs/1812.00315. 87 dialogue & discourse 11(2) (2020) 74–109 doi: 10.5087/dad.2020.203 lexical and contextual cue effects in discourse expectations: experimenting with german ‘zwar...aber’ and english ‘true/sure...but’ juliane schwab jschwab@uni-osnabrueck.de institute of cognitive science osnabrück university mingya liu mingya.liu@hu-berlin.de department of english and american studies humboldt university of berlin editor: vera demberg submitted 12/2019; accepted 08/2020; published online 08/2020 abstract existing literature shows that readers and listeners rapidly adjust their expectations about likely discourse continuations through discourse markers, as well as through other narrow linguistic and discourse contextual cues. however, it is unclear whether (i) the facilitative effects of various linguistic cues differ in quality and (ii) whether the effects interact with one another in any principled manner. we conducted two self-paced reading experiments on concessive constructions in german and english wherein optional lexical and/or contextual cues appeared ahead of a concessive discourse connective. the results demonstrate that readers can use both types of cues to anticipate the upcoming connective. thus, our study provides novel evidence for expectation-driven accounts of discourse processing and elucidates the functions of discourse signals. furthermore, the results also show that the role a type of cue plays may be subject to cross-linguistic variation. keywords: discourse processing, predictions, concessives, discourse signals, self-paced reading 1. introduction although much of psychoand neurolinguistic research is concerned with the processing of individual sentences, language processing in a natural setting involves more than the generation of meaning in isolated sentences. instead, comprehenders swiftly relate propositions within a text or conversation to one another and the broader (extra-)linguistic context to build a coherent discourse representation. a range of psychoand neurolinguistic studies indicate that processing at the discourse level is aided by predictions for the subsequent discourse content and structure (drenhaus et al., 2014; köhne and demberg, 2013; rohde et al., 2011; rohde and horton, 2014; scholman et al., 2017; van bergen and bosker, 2018; xiang and kuperberg, 2015). the current literature indicates that discourse expectations can be formed from a range of cue words or phrases, such as discourse markers and discourse connectives. xiang and kuperberg (2015), for instance, demonstrated that the discourse connective even so reverses readers’ expectation for the sentence continuation, such that unlikely events, like partying after having failed an exam in (1), become expected. similar effects c©2020 juliane schwab and mingya liu this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). lexical and contextual cue effects in discourse expectations have also been found for the concessive connectives however and german dennoch ‘nevertheless’ (drenhaus et al., 2014; köhne and demberg, 2013), as well as the causal connectives therefore and its german counterpart daher (ibid.). (1) elizabeth had a history exam on monday. she took the test and failed it. even so, she went home and celebrated wildly. further, scholman et al. (2017) have studied the processing of the two-part construction ‘on the one hand...on the other hand’. they found that the discourse marker ‘on the one hand’ affected how readers integrated upcoming discourse relations. among other things, the authors compared the reading times at on the other hand in conditions where the critical region was preceded by on the one hand to a control condition where it was not. in the initial study, they did not find a facilitating effect of the first connective on processing the second one. a follow-up eye-tracking study (scholman et al., 2018) with an increased filler-to-item ratio and twice the amount of subjects, however, found a first-pass reading time effect pointing toward prediction, that is, on the other hand was read faster if on the one hand had appeared in a previous sentence. the volatility of the effect suggests that readers’ adaptation to the structures they are exposed to during an experiment is a real concern when studying discourse expectations. if, as shown by the two studies discussed above, comprehenders generate discourse expectations, an obvious question is whether the cues relevant for the formation of expectations extend beyond discourse connectives. an increasing amount of research focuses on the questions how often and why discourse relations are signaled by means other than discourse markers, such as lexical, syntactic, or graphical cues (such as enumerated lists in written text) (das and taboada, 2018; taboada and das, 2013). critically, these discourse signals may appear with or without a discourse marker. in fact, das and taboada (2018) have shown that, at least within the rst (rhetorical structure theory) discourse treebank (carlson et al., 2002), 75% of discourse relations were marked exclusively with signals other than discourse markers, and 8% were marked with more than one signal. in light of these findings, it is clear that these discourse signals must serve an important function for establishing coherence relations. however, the precise mechanism via which signals come to indicate a specific discourse relation, as well as their interaction with discourse markers is so far not well understood. finally, a number of neurolinguistic studies have shown that the discourse context has an early impact on language comprehension. for instance, nieuwland and van berkum (2006) showed that supportive discourse contexts can quench processing difficulties typically observed for local semantic violations. in addition, hald et al. (2007) demonstrated that the discourse context has an immediate effect on sentence interpretation. in a visual world eye-tracking study, kim et al. (2015) showed that the discourse context acts as a constraint on the alternatives computed on-line during sentences containing focus particles. given this evidence for on-line effects of the discourse context during sentence processing, expectations for the discourse too may be affected by contextual signals, for instance through pragmatic inferences. one example for the immediate effect of pragmatic influences on discourse expectations comes from experimental work on lexically triggered pragmatic inferences. implicit causality verbs, such as scolded in (2) were found to generate an expectation for an upcoming causal discourse relation, that is, a reason for the scolding, for instance. the generated expectation was evidenced by anticipatory eye movements (rohde and horton, 2014) and participants’ on-line parsing decisions (rohde et al., 2011). 75 schwab and liu (2) arthur scolded patricia in the hallway. cause continuation: she had put thumbtacks on the teacher’s chair. to summarize, past work has shown that discourse processing is expectation-based, such that both discourse connectives and pragmatic inferences can generate discourse expectations. furthermore, these findings hint that additional discourse signals may have a similar effect on the on-line processing of discourse relations. so far, however, no study to our knowledge has either (i) directly contrasted the facilitative effects of different cues, or (ii) studied how comprehenders process sentences containing multiple cues. in the current study, we investigate the role of two types of discourse signals that appear ahead of a discourse connective and their interaction in discourse relation processing. we address the role of what we call ‘lexical’ cues, such as true in (3a), and ‘contextual’ cues, such as the incompatibility between liking to run outdoors and owning a treadmill in (3b), for the generation of discourse expectations. (3) a. paul likes to run. (true,) he has a treadmill in the living room, but he often jogs in parks. b. paul likes to run (outdoors). he has a treadmill in the living room, but he often jogs in parks. in particular, we used a self-paced reading task with subsequent naturalness ratings to answer two questions: first, do discourse cues generate an expectation that facilitates processing the discourse connective (e.g. but), as indexed by reduced reading times? second, do the effects of multiple cues interact with one another in any principled manner? does, for example, the occurrence of multiple cues result in additive facilitative effects? we take a cross-linguistic perspective (targeting german and english), as theories on the roots of expectation effects in language processing indicate that expectation effects arise from general cognitive processing principles (levy, 2008; lewis and bastiaansen, 2015). thus, cues that are similarly good indicators for the upcoming discourse relation in both languages under observation should also lead to a similar processing effect. in addition, the comparison between german and english builds upon existing work on the processing of concessive discourse relations in these two languages (drenhaus et al., 2014) that indicates similarities with regard to the processing effects of concessive connectives. based on the results of our study, we will argue that additional cues preceding the discourse connective may benefit the communication between speaker and listener by generating an anticipatory signal that can facilitate processing the upcoming discourse relation. the article is organized as follows: in section 2, we introduce the discourse cues under investigation in both languages. in section 3, we present production and corpus-linguistic studies on the german and english lexical cues that inform our hypotheses for the experiments. in section 4, we present the details of the german and english experiments. as outlook for the results: in german, we found evidence for expectations generated from lexical and contextual cues. in english, however, we only found evidence for lexical cue-based expectations. in section 5, we discuss these findings, explanations for the divergence between german and english, and implications for future research on discourse phenomena. 76 lexical and contextual cue effects in discourse expectations 2. the tested expressions our study concerns the processing of concessive discourse relations. following könig’s (2006) definition, concessive clauses, for example in the form of ‘p, but q’ involve the assertion of two propositions p and q, despite a presupposed incompatibility of q with p. concessives have also been described as anti-causal (konig and siemund, 2000), as p is said to give rise to a causal inference that is subsequently denied by q. for instance, in (3), p (he has a treadmill...) may give rise to the expectation that he extensively uses this treadmill to exercise. q (he often jogs in parks) subsequently denies this expectation. we decided to target concessive discourse relations in order to follow up and extend upon the existing research summarized above on the role of connectives in the processing of relations of opposition such as concession and contrast (asr and demberg, 2020; drenhaus et al., 2014; köhne and demberg, 2013; scholman et al., 2017; xiang and kuperberg, 2015). in this section, we will introduce the cues we refer to as lexical and contextual cues for concessive discourse relations in german and english. we start with the lexical cues1: the german discourse marker zwar2 (‘truly’) and the english particles true and sure. the latter particles are not typically classified as discourse markers (they are listed in neither the penn discourse treebank (prasad et al., 2008) nor the english lexicon of discourse markers dimlex-eng (das et al., 2018)). however, they have been described as markers of concessivity in the past (antaki and wetherell, 1999; könig, 2006). we argue that true/sure are ambiguous at the turn-initial position, where besides marking a conceded argument, they can also express the speaker’s sincere agreement with a previous speaker’s assertion. however, in turn-medial positions and outside of dialogues, the latter reading is infelicitous, rendering true/sure a clear signal for a concessive relation. we lay out the details of our argument below. our discussion of the english lexical cues focuses on clause-modifying true/sure3, where the particles simply assert that the speaker believes the clause the particles are modifying to be true. signaling a concessive relation is not part of their core meaning. this becomes apparent when they appear at the turn-initial position: true/sure here are used to express agreement with the last speaker’s assertion (4a). a concessive relation is possible, but optional (4b). (4) a: mary has done a lot to improve her health recently. a. b: true/sure(, she even quit smoking). b. b: true/sure(, but she still smokes like a chimney). we argue that true/sure in turn-medial positions and outside of dialogues become a clear indicator for a concessive relation: re-asserting that one believes in the truth of one’s own utterance is infelicitous, as can be seen from the minimally modified example without speaker change (5a). 1. we use the term lexical cue as umbrella term, as our english and german lexical cues differ with regard to whether they are typically classified as discourse markers 2. the etymological dictionary of german (pfeifer et al., 1993) indicates that zwar is a contraction of the proposition zu (‘to’) and the adjective wahr (‘true’), from old high german zi wāre ‘in truth, indeed’ (8th century), middle high german ze wāre/zewāre, originally used to express agreement to an assertion. its use as concessive marker seems to have emerged in middle low german 3. obviously, true/sure have a discourse-unrelated function as adjectives. this meaning, however, is easily distinguished from the clause-modifying one and is not available in the structural configurations we are investigating. therefore, we have nothing more to say on this matter. 77 schwab and liu instead, true/sure can only function as markers of an argument that is conceded to the audience. the particles can appear both before the conceded argument (5b) or after it (5c). (5) mary has done a lot to improve her health recently. a. #true/sure(, she even quit smoking). b. true/sure, she still smokes, but she started exercising regularly. c. she still smokes, true/sure, but she started exercising regularly. given these observations, we chose material wherein there is no ambiguity about the function of true/sure (see (3a)) for our study, that is, items where there is no apparent speaker (and thus no speaker change), and where true/sure only appears at the second sentence of our items, thus preventing participants from interpreting it as response to an omitted previous utterance. the german lexical cue that we investigate in analogy to true/sure is the adverb zwar (‘truly’). the german lexicon of discourse markers, dimlex (stede and umbach, 1998), lists zwar as a discourse marker that appears at the first argument of a concessive relation4. contrary to english true/sure, which appears before or after the proposition it modifies (5), german zwar usually appears immediately after the finite verb, that is, inside the proposition (6a). while the flexibility in the german word order allows fronting the adverbial phrase such that zwar appears in a sentence-initial position (6b), this is less frequent than sentence-medial zwar. similar to the two-part marker ‘on the one hand...on the other hand’, marking the conceded argument with zwar typically requires the second argument to be marked with a concessive connective (e.g. aber (‘but’) in (6)). just as for english true/sure, zwar is generally optional in establishing a concessive discourse relation. nevertheless, the two-part ‘zwar...aber’ construction is highly frequent both in formal and informal registers. to our intuition, the ‘true/sure...but’ construction, on the other hand, is less frequent. data from a parallel corpus that will be reported in section 3 confirms this intuition, but also demonstrates that ‘true/sure...but’ is indeed analogous in meaning to german ‘zwar...aber’. (6) jens jens joggt jogs gerne gladly (draußen). (outdoors). ‘jens likes to run (outdoors). a. er he hat has (zwar) (true) ein a laufband treadmill im in-the wohnzimmer, living-room aber but er he joggt jogs häufig often im in-the park. park b. zwar true hat has er he ein a laufband treadmill im in-the wohnzimmer, living-room aber but er he joggt jogs häufig often im in-the park. park ‘true, he has a treadmill in the living room, but he often goes running in the park.’ as what we call contextual cues, we employ adverbials added to the first sentence of our mini discourses (see (3b) and (6)). we refer to this as contextual cue as the adverbials themselves do not convey information regarding the discourse relation. instead, the adverbial generates an incoherence between the first and the second proposition of our items that should give rise to an expectation for a concessive discourse relation. let us spell out the workings of this manipulation in detail: (7a) literally asserts that john likes running outdoors. at the same time, as the manner of running is explicitly stated through inclusion of the optional adverbial outdoors, the proposition evokes alternatives to outdoors (e.g. 4. german also has a related expression, und zwar (lit. ‘and true’, roughly ‘namely’). this expression, however, is used to mark an expansion, rather than a concession (dimlex, stede and umbach (1998)). 78 lexical and contextual cue effects in discourse expectations on a treadmill, at the gym) and implies that, circumstances permitting, john indeed prefers running outdoors over its alternatives. by contrast, the sentence without adverbial, (7b), is underspecified with regard to john’s running exercise. (7) a. john likes to run outdoors. b. john likes to run. following (7), the second sentence, (8), asserts that john has a treadmill. out of the blue, this sentence would be rather uninformative with regard to likely discourse continuations (one could for instance continue the discourse with: ‘it was a gift from his sister, who didn’t need it anymore’). however, given (7), (8) expands the comment on john’s exercise regimen and implies that he uses his treadmill to run. on the one hand, as continuation of (7b), this is coherent with the immediately preceding sentence and thus does not generate an expectation for a concessive discourse relation (a continuation such as ‘and he uses it every day’ would be perfectly coherent). as continuation of (7a), on the other hand, this generates an incoherence: (7a) implied that john typically runs outdoors, while (8) implies that he uses his treadmill. (8) he has a treadmill in the living room. we argue that this discourse-based incoherence could generate an expectation for a concessive discourse continuation. this is because a concessive relation between (8) and a subsequent second argument reinstates the coherence of the discourse: in (9), the former becomes the conceded argument, for which, while true, the inference that john likes to use his treadmill to run, is denied. (9) he has a treadmill in the living room, but he often jogs in parks. in our stimulus material, the adverbial in (7a) is instrumental to achieve an opposition between (7a) and (the inference of) (8), which in turn should generate an expectation for a concessive discourse relation. note, however, that the same effect can be achieved with any set of two propositions in which the second one gives rise to an inference that is contradictory to what was asserted with the first one. an example can be seen in (10): having reversed the valence of the first sentence from liking to hating to run, the second proposition together with its inference that john uses the treadmill to run is suddenly rendered contradictory to the first sentence. as argued above, the saving grace to reinstate the coherence of the discourse is a concessive discourse relation, wherein the inference of the second proposition is denied. (10) john hates running. he has a treadmill in the living room ?(, but only his wife ever uses it.) given the existing evidence for comprehenders’ ability to generate discourse expectations from pragmatic cues outlined in the introduction (kim et al., 2015; rohde et al., 2011; rohde and horton, 2014), we expect that comprehenders should be able to recognize the incoherence in the discourse and rapidly adjust their expectations about the likely discourse continuation. in our materials, this should surface as a reduced reading time at the connective but. it is to note that with the incoherence generated through the addition of the contextual cue, there is a specific expectation about the inference that will be denied, namely that john runs indoors on his treadmill. by contrast, one could argue that in contexts containing only a lexical cue (3a), it is unclear what the default inference from john owning a treadmill should be. nevertheless, we take the position that, even then, the most obvious invited inference given the first sentence is that john likes to use his treadmill to run. 79 schwab and liu consequently, the sentence becomes somewhat odd if the concession denies a different inference (11). finally, while we agree that the contextual cue may set up the reader with a slightly more specific expectation regarding the content of the upcoming discourse, in both contextual and lexical cue conditions an expectation for the structure of the discourse, as measured at the connective but, should hold. (11) a. true, john has a treadmill in the living room, but he’s not a particularly fast runner. b. john likes to run. ?true, he has a treadmill in the living room, but he’s not a particularly fast runner. to summarize, in the present study, we target both lexical and contextual cues as the two indicate concessivity through crucially distinct mechanisms: german zwar and english true/sure mark a conceded argument. it is part of their lexical meaning to indicate concessivity. as such, a comprehender might automatically generate an expectation for a concessive discourse relation upon reading these lexical elements. by contrast, the contextual cue indicates concessivity through an incompatibility between two propositions and the inferences they give rise to. comprehenders must identify this incompatibility on-the-fly and use it to re-evaluate the discourse structure if they are to anticipate the upcoming concessive relation. our study contributes to the literature by investigating both types of cues within the same stimulus material and across two languages, thus providing insight into the effects of various cues in discourse relation processing. in the next section, we focus on the lexical cues in german and english. using corpus data and a production study, we first demonstrate that the lexical cues are indeed predictive of an upcoming concessive connective. the predictive power attested through this data forms the basis of our hypothesis that comprehenders should be able to generate discourse expectations on-line on the basis of these cues. second, in a parallel corpus study on constructions using the german or english lexical cues, we show that zwar is much more frequent than true/sure and often does not receive an analogous lexical marker in english translations. however, true/sure represent a low-frequency, but reliable, option to instantiate a concessive discourse relation in english that is commonly translated to ‘zwar...aber’. the data reported in section 3 informs our hypotheses for the self-paced reading experiments reported in section 4. 3. frequency and reliability of zwar, true, and sure as lexical cues as our study concerns the question whether the lexical cues zwar, true, and sure have a facilitative effect on processing the following connectives aber and but5 respectively, we report natural language data from german and english corpora that attest to the predictive strength of these expressions in sections 3.1 and 3.2. for the german expression zwar, we also include a production study testing the extent to which it triggers the use of aber or functionally related connectives like doch/allerdings (‘however’). in section 3.3, we then estimate the frequency of ‘zwar...aber’ and ‘true/sure...but’ constructions using a parallel corpus. 5. we chose aber and but as connectives as they are the most frequent connectives following zwar and true/sure, see section 3.3. of course, the concessive discourse relation can also be realized through other connectives, e.g. german jedoch / english however. 80 lexical and contextual cue effects in discourse expectations 3.1 zwar...aber in german the predictive strength of zwar as a lexical cue is attested in data collected and analyzed by günthner (2016): in a data set of 91 informal conversations, günthner (2016) found that zwar was part of a concessive relation in the vast majority of cases (96% of all extracted instances of zwar). only 1.3% of her data were stand-alone instances of zwar, see an example in (12) (günthner (2016), p.156; transcription simplified). (12) h. recounts being asked to go out to a club h: un ich so; ich hab zwar keine zeit, ich war auch in arbeitsklamotten, and i was like; true i have no time, i was wearing my work clothes (anna, eri and mari laugh) h: dann äh nach hause, schnell geduscht then (uh) straight home, quickly showered e: kann gar nich sein can’t be h: sind in die disse und da war das das went to this club and that was it that in (12), the expected connective aber is substituted with a recounting of events that instead illustrate a concession via h’s subsequent actions. günthner also cites an earlier study by primatarovamiltscheva (1986) that confirms this distribution of zwar...aber, although primatarova-miltscheva finds a slightly higher rate of stand-alone instances of zwar, with 18% of the analyzed spoken language instances of zwar not being part of a concessive relation. to verify the predictability of aber by zwar, we also conducted a sentence continuation study. 26 students of osnabrück university (16 female, mean age = 20.38, sd = 1.6) participated in the study. all participants were native speakers of german and had normal or corrected-to-normal vision. participants gave written informed consent prior to participating and were compensated with partial course credit. participants were shown sentence fragments such as (13) on a computer screen and used the keyboard to enter a grammatical continuation of the sentence into a text box below the fragment. there was no time or word limit for their responses. each participant saw 4 items with zwar as in (13), pseudo-randomly interspersed with 28 fillers from another study. the results of the experiment are summarized in table 1. we list both the connective used and its position, as german allows for the placement of a connective either before the subject position (vor-vorfeld, ‘pre-prefield’) or between the finite and nonfinite verb positions (mittelfeld, ‘middle field’). see examples for each in (14). (13) louisa louisa hat has zwar true eine a plattensammlung... vinyl-collection... ‘true, louisa has a vinyl collection...’ three aspects of the results are notable: for one, participants always produced an adversative or concessive discourse marker. second, aber is the most frequently used marker. in the cases where participants used other markers, these can be substituted by aber without change in meaning (e.g. (14)). as noted by scheffler and stede (2016) and stede et al. (2019), the discourse connectives in table 1 overlap in their discourse functions, such that all of them can designate a concessive relation. it is thus in line with these authors’ work that we observed participants using a variety of 81 schwab and liu connectives. third, the vor-vorfeld position, which is where the discourse marker appears in the material for our study, is the preferred position to mark the concessive relation. (14) a. louisa louisa hat has zwar true eine a plattensammlung, vinyl-collection, doch yet / / aber but sie she hört listens nicht not gerne gladly musik. music ‘true, louisa has a vinyl collection, but she does not like listening to music.’ (vorvorfeld connective) b. louisa louisa hat has zwar true eine a plattensammlung, vinyl-collection, ihr her plattenspieler record-player funktioniert works allerdings however / / aber but nicht not mehr more richtig. correctly ‘true, louisa has a vinyl collection. but her record player does not really work anymore.’ (mittelfeld connective) discourse marker frequency in vor-vorfeld position frequency in mittelfeld position aber (‘but’) 50 (48.08%) 24 (23.08%) jedoch (‘though’) 10 (9.62%) 7 (6.73%) allerdings (‘however’) 3 (2.88%) 4 (3.85%) doch (‘yet’) 4 (3.85%) 0 dennoch (‘still’) 0 2 (1.92%) total 67 (64.42%) 37 (35.58%) table 1: sentence-continuation task: frequency of adversative/concessive discourse markers following a sentence fragment containing zwar with percentage of the total number of responses in parentheses. based on the literature and our production study, we assume that zwar functions as a good predictor for a concessive construction. furthermore, as aber was the most frequent concessive marker in the production data, participants may come to expect the concessive relation to be lexically realized by aber. we thus chose to use aber in the material of our german experiment due to its high frequency both in general and in relation to zwar. 3.2 true/sure...but in english the english particles true and sure signal an upcoming concession in a similar way. we searched in the 100,000,000-word the british national corpus (2007), distributed by the university of oxford on behalf of the bnc consortium. the search was conducted via the treebank.info parser (uhrig and proisl, 2011). first, we retrieved all instances of the particles true (398 in total) and sure (118 total). then, we manually checked for concessive discourse moves both within the sentence and in the 5 sentences following the sentence marked by true or sure. the results show that true/sure was commonly part of a concessive move, with true occurring as part of a concessive discourse segment in 86% (344 instances) of the data, and sure in 83% (98 instances) of the data. of those, for both particles, intrasentential concession was the most common form of discourse continuation (69% of the 98 instances for sure, 59% of the 344 instances for true). in a smaller subset of the data, the second argument of the concessive move, i.e. the but82 lexical and contextual cue effects in discourse expectations clause, appeared in the immediately following sentence (22% of the data for sure, 26% for true), or other sentences intervene between the sentence marked with true/sure and the second argument (8% of the data for sure, 15% for true). examples for both intraand intersentential concessions are provided in (15)6. as the bnc contains some instances of reported speech, the 14% (true)/17% (sure) of the data that did not contain a concessive discourse relation included some instances of true/sure as responses to a previous speaker’s assertion, a function also outlined in section 2. in other cases, the function of true/sure was not apparent to us. (15) a. true, things have gone missing, but that does not mean that i am the thief. (bnck95-33) b. true, she probably still had a long way to go. she was painfully thin and there was an insubstantiality about her. but there was no denying that today his wife was better than he had known her for many many months. (bnc-cde-222-224) in effect, the corpus data verifies that the particles true and sure are good indicators for a concessive discourse structure. the reported data from german and english demonstrate that the two english particles are both similar to each other and to the german zwar in how often they signal a concessive discourse relation. however, one potentially concerning facet we have thus far not addressed are differences in frequency between german zwar and english true/sure: while zwar is a highly frequent discourse marker, the use of the particles true and sure appears to be more limited. in the next section, we therefore report on a parallel corpus study that estimates their frequency in comparison to german zwar and discusses potential implications for our work. 3.3 frequency of zwar, true, and sure our objectives for a parallel corpus study were to (i) gather an estimate of the frequency of german zwar and english true/sure from matching text sources, and to (ii) investigate how the german concessive construction using zwar is translated to english and vice versa. we used the german-english archive of the parallel corpus europarl (version 7) (koehn, 2005) that contains 1,920,209 aligned sentences from the proceedings of the european parliament published between 1996 and 2011. first, we extracted all instances of german zwar (excluding cases of ‘und zwar’, which, as mentioned in footnote 4 of section 2, marks an expansion rather than a concession), as well as all instances of english true/sure. as the corpus does not contain pos-tags or dependency annotations that would facilitate the exclusion of adjectival true/sure, we instead restricted our search in the english archive to only include capitalized instances (‘true’), or instances delimited by commas on either side (‘, true,’). then, we manually removed any remaining adjectival uses of true/sure (as in ‘true friendship’). while, admittedly, this procedure may have caused us to miss some instances of discourse-functional true/sure, it ensured a clean data set that reflects the usage of the lexical cue in our experimental material (see (3)). the search results show that german zwar is highly frequent with 11,671 instances in the corpus. as noted in section 2, zwar can appear both sentence-initially and sentence-medially. among the 11,671 hits for zwar, 1,601 (13.7%) are sentence-initial (i.e. capitalized). english true/sure are quite infrequent with 112 instances of true and only 3 of sure. nevertheless, the data again support the discourse cue function of true/sure that we have reported above: in 105 of the 115 collective in6. examples of usage taken from the bnc were obtained under the terms of the bnc end user licence. copyright in the individual texts cited resides with the original ipr holders. 83 schwab and liu lexical cue (on conceded argument) in source text connective (on second argument) in source text true (102) /sure (3) but however nevertheless/nonetheless (al)though even so/yet/still no connective lexical cue in german connective on second argument in german zwar aber (‘but’) 17 1 1 – – – (31) (je)doch (‘however’) 8 3 – 1 – – stimmt (‘correct’) aber (‘but’) 12 1 1 – – – (21) allerdings (‘however’) – 1 – – – – dennoch (‘nevertheless’) – – 1 – – – (je)doch (‘however’) 2 1 – – – 1 no connective – 1 – – – – sicher(lich) (‘sure(ly)’) aber (‘but’) 5 – – – – – (9) allerdings (‘however’) – 1 – – – – (je)doch (‘however’) 1 – – 1 – – dennoch (‘nevertheless’) – – – – 1 – natürlich (‘natural(y)’) aber (‘but’) 3 – – – – – (8) (je)doch (‘however’) 3 1 1 – – – gewiss (‘certain(ly)’) aber (‘but’) 2 – – – – 1 (6) dennoch (‘nevertheless’) – – 1 – – – no connective – – – 1 1 – other lexical cue aber (‘but’) 11 1 1 1 – – (26) allerdings (‘however’) 2 – – – – – dennoch (‘nevertheless’) – – – – 1 – (je)doch (‘however’) 4 3 – – – 1 no connective – 1 – – – – no lexical cue aber (‘but’) 2 – – – – – (4) (je)doch (‘however’) 2 – – – – – table 2: results of a parallel corpus analysis using the german-english aligned version of europarl. the upper left corner indicates the lexical cue that was present in the original text source (true/sure). the two leftmost columns indicate the lexical cue on the conceded argument (with total frequency counts in parentheses) and the connective on the second argument that were identified in the german translation. entries in the table provide a count of the frequency with which a combination of true/sure+connective was translated into a corresponding cue+connective in german. for example, the first entry of the table (17) shows that ‘true/sure..., but’ was translated to ‘zwar..., aber’ 17 times. stances of true/sure, the cue was part of a concessive relation. in a second analysis step, we aligned the search results with the respective german or english translation. we used the aligned sentences to analyze how german zwar-constructions are translated to english and, conversely, how english true/sure-constructions are translated to german. taking the 105 instances of true/sure in concessive relations, we manually categorized the aligned sentences by (i) the connective at the second argument in the original texts, (ii) the lexical cue at the conceded argument in the translations, and (iii) the connective at the second argument in the translations. the results are summarized in table 2. as this procedure was not manageable for all 11,671 instances of zwar, we instead selected a random sample of 500 instances of zwar (including 87 sentence-initial ones) and followed the same procedure as outlined above. the results are reported in table 3. for the sake of brevity, we grouped some cue phrases and connectives together (e.g. nevertheless and nonetheless) and conflated those that appeared in less than 5% of the samples into an ‘other’ category. an expanded version of this table featuring the full list of cues can be found in the appendices. for all instances categorized under ‘no connective’, no connective was apparent in the text. it is clear that the particles true/sure are much less frequent than german zwar. for a crosslingusitic study on these cues, a reader may thus wonder whether the stark difference in the fre84 lexical and contextual cue effects in discourse expectations lexical cue (on conceded argument) in source text connective (on second argument) in source text zwar (500) aber doch jedoch dennoch allerdings other connective no connective cue/connective in english connective on second argument in english while/whilst however 1 – – – – – – (81) nevertheless 1 1 – 1 – – – on the other hand – 1 – – – – – no additional connective 37 16 17 3 1 2 – (al)though however – – 1 – – – – (76) nevertheless 3 – – – – – – on the other hand – – – – – 1 – no additional connective 38 17 12 – 2 1 1 modal verb (may) but 16 4 3 – – – – (28) however 1 – 1 – – – – on the other hand – 1 – – – 1 – no connective 1 – – – – – – other lexical cue but 22 8 5 1 – – – (67) however 2 1 2 – 1 – – nevertheless/nonetheless 1 – – 1 – – – on the other hand 1 – – – – – – yet/still 1 – 1 – – – – no connective 9 7 1 – – – 3 no lexical cue but 106 49 44 – 2 3 – (248) however 1 2 2 – 1 – – nevertheless – – 1 – – – – yet 2 3 1 1 – – – no connective 13 5 2 2 – 1 7 table 3: results of a parallel corpus analysis using the german-english aligned version of europarl. the upper left corner indicates the lexical cue that was present in the original text source (zwar). the two leftmost columns indicate the lexical cue on the conceded argument and the connective on the second argument that were identified in the english translation. entries in the table provide a count of the frequency with which a combination of zwar+connective was translated into a corresponding cue+connective in english. quency of the constructions we are investigating could confound our experimental results. indeed, past research has often found that low-frequency words and structures can be harder to process (cf. ferreira et al. (1996); rayner and duffy (1986); schilling et al. (1998) on the effects of word and structural frequency on sentence processing). however, as we are not interested in the processing time at the lexical cue itself or at the regions immediately following the lexical cue, but instead on its ability to generate discourse expectations for an upcoming connective, transient processing difficulties associated with encountering a low-frequency construction should arguably subside before the critical region, the connective but/aber. further, these connectives themselves are the most frequent ones for the instantiation of a concessive relation in both languages. admittedly, one caveat remains: despite the fact that the corpus data indicates that a concessive connective is highly predictable after true/sure and after zwar, the low frequency of the former construction may mean that (at least some) english native speakers are not sufficiently familiar with it to rapidly recognize true/sure as part of a concessive relation and to generate discourse expectations from it. this could lead to reduced expectation-based effects in the english experiment compared to the german one. in any case, we will consider potential frequency-induced effects in the interpretation of our experimental findings. the results of this corpus study also demonstrate that true/sure..., but is indeed commonly translated to analogous concessive constructions in german. the most frequent translation is zwar...aber, but we also observe variations, particularly on the lexical cue at the conceded argument. other trans85 schwab and liu lations for true include typical adjectives and adverbs marking the truth of an assertion, like stimmt (‘correct’), sicher(lich) (‘sure(ly)’), or natürlich (‘natural(y)’) (see (16a)). this is not surprising given that they share semantic properties with zwar in terms of veridicality (giannakidou (1998) and subsequent works), i.e. all the variants and zwar are veridical operators with regard to the modified proposition. zwar differs from the others as it is exclusively used as concessive marker. in our sample of zwar, on the other hand, the majority of english translations did not use an analogous two-part construction. instead, english often solely relies on the realization of a single connective, either at the conceded argument (while, although) or at the second argument (but, however, and others). we also observe the common presence of the modal verb may at the conceded argument, typically followed by a connective like but. while we do not take a strong stance on how modal verbs signal concessive discourse relations, epistemic modality as signal for concessivity has been discussed for instance by baranzini and mari (2019), könig (2006), and souesme (2009). (16) a. das this ist is natürlich naturally ein a erster first schritt, step, dieser this bleibt stays aber but weit far hinter behind dem the bedarf need zurück. back (sentence #1202079, german) ‘this is a first step, true, but it falls far short of the requirements’ (sentence #1202079, english) b. das that heißt, means die the gesetze laws sind are zwar true da, there aber but sie they müssen must auch also angewandt applied werden. become (sentence #790455, german) ‘in other words, the laws are in place, but they still need to be enforced.’ (sentence #790455, english) to summarize, the parallel corpus data indicate that most instances of zwar are not realized with an analogous lexical signal in english. nevertheless, true/sure as lexical cues indeed represent a low-frequency option to overtly mark the conceded argument of english concessive constructions. 3.4 summary with respect to signaling a concessive relation, the data reported in sections 3.1 and 3.2 demonstrate that both the german and the english lexical cues are good indicators of a concessive discourse relation. still, it is an open question whether the pattern observed in the corpus data will actually translate to comprehenders’ use of these cues in on-line processing. thus, the chosen lexical cues are well-suited for an investigation of discourse expectations generated by lexical cues. the parallel corpus data reported in section 3.3 further showed that concessive true/sure constructions are commonly translated to german zwar constructions. this supports our decision to compare the two in a cross-linguistic experiment. the latter data set also showed that zwar is much more frequent than true/sure. potential frequency effects will thus have to be taken into consideration when interpreting the results of our study. 4. testing expectations in discourse processing / experiments we conducted two experiments in german and in english, using a combination of self-paced reading and rating tasks with reading time and naturalness ratings as measures. both experiments were 86 lexical and contextual cue effects in discourse expectations based on a fully factorial design wherein we varied the amount of cues towards the concessive relation. for the naturalness ratings, we predicted that contextual cues in particular would increase the naturalness as they create a coherent discourse by closely relating the two sentences of our experimental items. furthermore, due to the lower frequency true/sure...but, items containing these lexical cues may be rated as less natural than the same constructions with no lexical cues, as well as less natural than the german zwar....aber construction. concerning the reading time, our hypotheses were as follows: first, if readers can anticipate the discourse relation from lexical or contextual cues, we should find a reduced reading time at the connective when the preceding context contains either of these cues, as compared to a baseline condition without cues. second, lexical and contextual cues may not lead to equally large expectation-based effects. as outlined in section 2, the lexical cues are clear indicators of a conceded argument. comprehenders in principle only need to recognize the cues as such in order to generate an expectation for a concessive discourse relation. the contextual cue condition, on the other hand, requires comprehenders to detect the incoherence between two propositions and to use it to adjust their discourse expectations. the two cues thus give rise to expectations through different means to which comprehenders in turn may not be equally sensitive. third, if multiple cues predicting the same upcoming discourse relation occur together, we wondered whether they may function as cumulative evidence towards the upcoming discourse relation. if that is the case, we might find that processing is facilitated even beyond the effect of a singular cue. 4.1 german experiment 4.1.1 materials 28 items were generated based on a 2 x 2 design with the following manipulations: first, the initial sentence (s1) was either neutral with respect to the content of the following sentence (conditions 1 and 3, henceforth c1 and c3), or incongruent with it (c2 and c4). in the latter case, the incongruence acts as cue towards a concessive structure. thus, conditions 2 and 4 contain a contextual cue, whereas conditions 1 and 3 do not. the second experimental manipulation concerned lexical cues, such that the second sentence (s2) was lexically cued for the upcoming contrast with zwar (‘truly’) (c1 and c2), or did not contain a lexical cue (c3 and c4). with these two manipulations our items looked as follows: the first sentence always introduced an action that the agent likes to perform. it consisted of a proper name, followed by an intransitive verb and adverbial modifiers. in the contextual cue conditions, the verb was modified by only one adverb, namely gerne (lit. ‘gladly’, akin to english ‘to like + [verb gerund]’). in the +contextual cue conditions, it was modified by two adverbs, namely gerne and a second adverb that specifics the condition under which the action is performed. thus, as can be seen in (8), the context sentence jens läuft gerne (‘jens likes to run.’) in c1 and c3 is underspecified with regard to the running conditions that the agent prefers, whereas the explicit context sentence jens läuft gerne draußen (‘jens likes to run outside.’) in c2 and c4 specifies exactly the way that the agent prefers to do the running exercises. the following sentence then introduces a second fact about the agent, for instance that they own a treadmill. this fact is always congruent with the neutral context sentence, but incongruent with the contextual cue condition. in c1 and c2, the sentence was lexically cued with zwar. in c3 and c4, there was no lexical cue. finally, a coordinating clause starting with the discourse connective aber (‘but’) introduces a concessive relation between the two conjuncts. the region containing the discourse connective is the critical region (cr) at which we expect differences in reading times depending on whether 87 schwab and liu subjects had anticipated the concessive relation. an example for a test item in all four conditions can be seen in (17). (17) 1. jens jens / / läuft runs / / gerne. gladly. / / er he / / hat has / / zwar true / / ein a laufband treadmill / / im in-the wohnzimmer, living-room / / aber but ercr hecr / / joggt jogs / / häufig often / / im in-the park. park ‘jens likes to run. true, he has a treadmill in the living room, but he often jogs in the park.’ 2. jens jens / / läuft runs / / gerne gladly draußen. outdoors. / / er he / / hat has / / zwar true / / ein a laufband treadmill / / im in-the wohnzimmer, living-room / / aber but ercr hecr / / joggt jogs / / häufig often / / im in-the park. park ‘jens likes to run outdoors. true, he has a treadmill in the living room, but he often jogs in the park.’ 3. jens jens / / läuft runs / / gerne. gladly. / / er he / / hat has / / ein a laufband treadmill / / im in-the wohnzimmer, living-room / / aber but ercr hecr / / joggt jogs / / häufig often / / im in-the park. park ‘jens likes to run. he has a treadmill in the living room, but he often jogs in the park.’ 4. jens jens / / läuft runs / / gerne gladly draußen. outdoors. / / er he / / hat has / / ein a laufband treadmill / / im in-the wohnzimmer, living-room / / aber but ercr hecr / / joggt jogs / / häufig often / / im in-the park. park ‘jens likes to run outdoors. he has a treadmill in the living room, but he often jogs in the park.’ in addition, the experiment also used 68 filler items, each consisting of 2 sentences to ensure uniformity with the experimental items. we used two types of fillers that were tested for the purpose of another experiment not reported here. an example for each type of filler can be seen in (18). the order of experimental items and fillers was pseudo-randomized following a latin-square design, such that every participant saw only one condition of each item and two experimental items were always interspersed with at least two fillers. participants saw a total of 96 prompts in the experiment. (18) a. nils nils / / sitzt sits / / in in einer a firmenkantine. company-canteen. / / besonders very häufig often / / trifft meets / / einer one der the.gen mitarbeiter employees / / in in der the pause break / / einen one der the.gen kollegen, colleagues / / die who / / kinder children / / haben. have. ‘nils sits in a company’s canteen. one of the employees very often meets one of the colleagues who have children in the break.’ b. hannes hannes / / geht goes / / einkaufen. shopping. / / gerne happily / / berät advises / / derjenige the-one verkäufer, salesperson / / dessen whose laden store / / echte real delikatessen delicacies / / anbietet, offers / / bei at unklarheiten unclarities / / den the kunden. customer ‘hannes goes shopping. the salesperson, whose store offers real delicacies, happily advises the customer in case of any unclarities.’ 88 lexical and contextual cue effects in discourse expectations 4.1.2 participants 57 students (40 female, mean age = 22.47, sd = 2.9) from osnabrück university participated in the study. all participants were native speakers of german and had normal or corrected-to-normal vision. participants gave written informed consent prior to participating and were compensated with partial course credits or monetary payment. the experiment was approved by the ethics committee of osnabrück university. 4.1.3 procedure the experiment was programmed and presented on ibex farm (drummond, 2013). items were presented on a screen in a moving-window self-paced reading paradigm. the items appeared in words or phrases such as indicated through the slashes in (17) and (18). for every item, the two sentences appeared in the middle of the screen, with the second sentence appearing below the first one. before reading an item, subjects thus saw a white screen with two lines of grey dashes where the words would appear. while reading, they used the space bar to reveal the next section, causing the previous one to disappear. after reading an item, a question asking for a naturalness rating of the sentences would appear on a new screen. we used a 7-point likert scale with the endpoints marked as ‘unnatural’ (1) and ‘natural’ (7). in 25 out of the 96 items (distributed randomly over the experiment) the naturalness rating was further followed by a comprehension question appearing on a new screen. this question could be answered with ‘yes’ or ‘no’ based on the content of the item the participant had just read. the experimental session began with 4 practice trials, during which participants were still allowed to ask questions about the experimental procedure. in total, the experiment took about 35 minutes. 4.1.4 results reading times and naturalness ratings were analyzed using linear mixed-effects regression models using the lme4 package (bates et al., 2018) in r (r development core team, 2019).7 prior to the analysis, all participants with an accuracy rate below 80% (<20 correct answers) on the comprehension questions were removed from the data set. we thus had to remove 7 out of 57 participants. all following analyses were performed on the data set with 50 remaining subjects. to remove outliers in the reading times, we took the whole data set and removed all reading times more than 2 standard deviations from the mean for each subject’s reading time per condition and region. furthermore, reading times below 150ms were removed. the outlier removal process affected 1.5% of our data. for both naturalness ratings and reading times, our predictors are the binary variables contextual cue (+contextual cue v. -contextual cue) and lexical cue (+lexical cue v. -lexical cue). both factors were coded as sum contrasts and included as fixed effects (with interaction) in our models (vasishth and broe, 2011). following the box-cox procedure (box and cox, 1964), we determined that a reciprocal transformation of reading times would be appropriate to ensure a normal distribution of residuals. we hence used the reciprocal rt as dependent measure, which was multiplied by 1000 to make the estimates more interpretable (this does not affect the results in any way). we used the maximal random effects structure that allowed our models to converge (barr et al., 2013), which, unless otherwise noted, included random intercepts per subject and item. for the reading time analysis, we performed analyses on the critical region (aber er/sie, ‘but he/she’), as well as the regions 7. all data and code associated with this experiment are available from: https://osf.io/ux8de/ 89 https://osf.io/ux8de/ schwab and liu condition rating rt cr-1 rt cr rt cr+1 c1: -context cue,+lexical cue 6.21 (1.09) 737.36 (332.45) 550.05 (201.87) 477.76 (161.50) c2: +context cue,+lexical cue 6.36 (1.05) 754.65 (367.82) 519.44 (145.40) 469.99 (126.80) c3: -context cue,-lexical cue 6.13 (1.23) 738.15 (312.12) 583.51 (239.98) 486.48 (139.53) c4: +context cue,-lexical cue 6.27 (1.12) 754.53 (350.05) 566.87 (213.26) 468.67 (118.75) table 4: mean naturalness ratings (on a scale of 1-7) and reading times (in ms) for the pre-critical, critical, and post-critical region with standard deviations in parentheses preceding (cr-1), and following it (cr+1). effects in self-paced reading are often delayed or distributed across multiple (adjacent) regions (smith and levy, 2013; witzel et al., 2012), therefore we include the post-critical region cr+1 to check for delayed effects and effects that may span across both the cr and cr+1 regions. additionally, we include an analysis on the pre-critical region cr1 to ensure that any effects found at the critical region are not caused by spillover from previous material. the mean ratings and raw reading times are provided in table 4. naturalness ratings the model revealed a main effect of contextual cues (β = -0.08, se = 0.02, t = -3.35, p = 0.0008). items with a contextual cue were rated more natural than those without. no other effect was significant. reading times in figure 1, we provide a visualization of the reading times over the course of both sentences, with a close-up on the regions of interest. there were no significant effects at the preor postcritical regions. at the critical region we found a significant effect of contextual cues (β = -0.03, se = 0.01, t = -2.42, p < 0.05) and a highly significant effect of lexical cues (β = -0.05, se = 0.01, t = -4.81, p < 0.0001), but no interaction. the critical region was read faster if the sentence was explicitly cued with zwar, or if the context sentence contained a contextual cue, but the two effects were independent of each other. that is, the lexical and the contextual cues together did not cause a superadditive reduction in reading times. 4.2 english experiment 4.2.1 materials 24 items were generated based on the 2 x 2 design of the german experiment. again, the first sentence was either neutral (c1 and c3) or incongruent with the second one (c2 and c4), that is, the former contained no contextual cue whereas the latter did. as in the german experiment, the first sentence always introduced an action that the agent likes to perform. in the +contextual cue conditions, this description was additionally modified by an adverb that specified how the agent prefers to do the action. the second sentence was either left unmarked for the upcoming contrast (c1 and c2) or contained a lexical cue (c3 and c4), namely a sentence-initial particle (true in 12 items, sure in the other 12 items). the crucial concessive relation was signaled by the discourse connective but. an example item in all four conditions can be seen in (19). we used 48 filler items, all containing 2 sentences for uniformity with the experimental items. an example for the fillers is given in (20). the order of experimental items and fillers was pseudo-randomized following a latinsquare design, such that every participant saw only one condition of each item and two experimental items were always interspersed with at least two fillers. participants saw a total of 72 prompts in the experiment. 90 lexical and contextual cue effects in discourse expectations (a) mean reading times for all regions (b) mean reading times for the critical region (aber er/sie, ‘but he/she’), as well as the cr-1 and cr+1. figure 1: mean reading times (ms) over (a) all regions in the 4 different conditions and (b) the precritical to post-critical region. the red rectangle in (a) marks the three regions displayed in (b). error bars indicate standard errors. each data point in (b) represents the average reading time of one subject. 91 schwab and liu (19) 1. james / likes /to run. / true, / he / has / a treadmill / in the living room, / but hecr / often / jogs / in parks. 2. james /likes /to run outdoors. / true, / he / has / a treadmill / in the living room, / but hecr / often / jogs / in parks. 3. james /likes /to run. / he / has / a treadmill / in the living room, / but hecr / often / jogs / in parks. 4. james /likes /to run outdoors. / he / has / a treadmill / in the living room, / but hecr / often / jogs / in parks. (20) vincent / lives / in a big city. / because / he / knows / his way / around the city, / he / can give / directions / to pedestrians. 4.2.2 participants 88 participants (32 female, 1 non-binary, mean age = 35.34, sd =10.42) were recruited via the crowd sourcing platform amazon mechanical turk. the participants were located in the united states of america and declared that they were native speakers of english. participants were informed about the general nature and duration of the experiment and gave consent to participating by ticking a box on the website. they were compensated with monetary payment. the experiment was approved by the ethics committee of osnabrück university. 4.2.3 procedure the experiment was programmed on ibex farm. the procedure was the same as in the german study. to ensure that all participants read attentively, the naturalness rating was always followed by a comprehension question appearing on a new screen. this question could be answered with ‘yes’ or ‘no’ based on the content of the item the participant had just read. the experimental session began with 4 practice trials. in total, the experiment took about 30 minutes. 4.2.4 results reading times and naturalness ratings were analyzed using linear mixed-effects regression models using the lme4 package in r.8 prior to the analysis, all participants with an accuracy below 80% (<58 correct answers) on the comprehension questions were removed from the data set. we thus had to remove 21 out of 88 participants. all following analyses were performed on the data set with 67 remaining subjects. to remove outliers in the reading times, we took the whole data set and removed all reading times more than 2 standard deviations from the mean for each subject’s reading times per condition and region. furthermore, reading times below 150ms were removed. the outlier removal process affected 3.7% of our data. the models for both the naturalness ratings and the reading times were constructed under the same specifications as in experiment 1. the mean ratings and reading times are provided in table 5. as indicated above, 12 of our items used true as lexical cue, the other 12 used sure. thus, our first objective in the analysis was to address whether true differed in any significant way from sure in either naturalness ratings or reading times. we constructed separate models for the ratings, the critical region and the post-critical region over the part of the data set that contained a lexical cue (conditions 1 and 2). in addition to the fixed and random effects specified above, we included 8. all data and code associated with this experiment are available from: https://osf.io/ux8de/ 92 https://osf.io/ux8de/ lexical and contextual cue effects in discourse expectations condition rating rt cr-1 rt cr rt cr+1 c1: -context cue,+lexical cue 5.73 (1.26) 621.31 (424.19) 488.41 (251.35) 378.43 (143.84) c2: +context cue,+lexical cue 5.65 (1.34) 619.06 (413.51) 470.49 (230.60) 370.71 (145.14) c3: -context cue,-lexical cue 5.94 (1.23) 649.57 (463.89) 477.31 (211.80) 399.72 (222.36) c4: +context cue,-lexical cue 6.07 (1.17) 679.00 (583.75) 478.26 (211.80) 387.92 (155.16) table 5: mean naturalness ratings (on a scale of 1-7) and reading times (in ms) for the pre-critical, critical, and post-critical region with standard deviations in parentheses a binary predictor for the type of lexical cue (that is, whether the item in question used true or sure as lexical cue) with an interaction term. none of the models showed any significant effect of the lexical choice (true vs. sure), nor any interaction with the other predictor. nevertheless, we compared the models containing the type of lexical cue as predictor to models without said predictor using the likelihood ratio test, which assesses the goodness of fit between two models. the models containing the additional predictor were found not to have a better fit for our data in either the naturalness ratings (x2 = 0.78, df = 2, p = 0.68), or the regions of interest for our reading time analysis (cr: x2 = 2.03, df = 2, p = 0.36; cr+1: x2 = 1.12, df = 2, p = 0.57). thus, we decided to conflate the two lexical cues in all further analyses. naturalness ratings the model, whose fixed effects included the two binary predictors, and whose random effects structure included random by-item and by-subject intercepts, as well as random by-subject slopes for both fixed effects, revealed a significant interaction between the contextual and lexical cue (β = -0.05 se = 0.02, t = -2.16, p < 0.05). the cause of this interaction effect was investigated in post-hoc paired comparisons via tukey’s hsd: in the presence of a contextual cue, the lexically cued condition was rated less natural than the uncued condition (β = 0.41, se = 0.10, df = 112.9, t = 4.30, adjusted p < 0.001). however, if contextual cues were absent, lexically cued and uncued conditions received similar ratings (β = 0.21, se = 0.10, df = 112.9, t = 2.18, adjusted p = 0.14). the condition containing only a contextual cue (c4) was rated more natural than the condition containing only a lexical cue (c1) (β = -0.34, se = 0.10, df = 65.5, t = -3.55, adjusted p < 0.01) reading times we provide a visualization of the reading times over the course of both sentences, with a close-up on the regions of interest in figure 2. there were no significant effects at the pre-critical or the critical region. at the post-critical region (cr+1), we found a significant effect of lexical cues (β = -0.04, se = 0.02, t = -2.58, p < 0.01), such that the region was read faster if the sentence had been cued with true/sure. there was no effect of contextual cues, and no interaction. 5. discussion the rapid identification of coherence relations is a key component of successful discourse processing. in the current study, we used a combination of self-paced reading and rating tasks to address the question whether the on-line process of establishing a discourse relation is aided by expectations generated from cues preceding the discourse connective. the results from the german experiment show that readers immediately integrate information from lexical and contextual cues to update their discourse expectations. the results from the english study confirm the importance of lexical 93 schwab and liu (a) mean reading times for all regions (b) mean reading times for the critical region (but he/she), as well as the cr-1 and cr+1. figure 2: mean reading times (ms) over (a) all regions in the 4 different conditions and (b) the precritical to post-critical region. the red rectangle in (a) marks the three regions displayed in (b). error bars indicate standard errors. each data point in (b) represents the average reading time of one subject. 94 lexical and contextual cue effects in discourse expectations cues, but do not show any facilitative effect of contextual cues, at least in on-line processing. in the following, we will discuss the findings for german and english in turn, and will conclude with an elaboration on the cross-linguistic variation we observed. the german experiment can be summarized as such: the reading time results demonstrate that readers use both contextual and lexical cues to anticipate discourse relations, although lexical cues showed a numerically stronger effect. given the reliability of the lexical cue attested by the corpus data presented in section 3, its power for generating discourse expectations conforms with our hypotheses. the more surprising aspect of our results is the effect of contextual cues: for one, in the absence of the lexical cue zwar, readers were able to immediately connect the content of the second sentence to its embedding context and to generate an expectation that mitigated the absence of any further discourse signals. secondly, the contextual cue maintained an independent facilitative effect even when it appeared together with the lexical cue. our study thus provides evidence that the joint presence of discourse cues from multiple linguistic sources can act as cumulative facilitators of discourse processing. lastly, the naturalness ratings for the german experiment showed that the items were rated more natural in the conditions containing a contextual cue, suggesting that the presence of contextual cues may have generated a more coherent discourse. this is also in line with the earlier observation that through the addition of the contextual cue the invited causal inference in the conceded argument is more specific than in conditions without contextual cues. lexical cues had no effect on the ratings, which hints that the off-line judgments are not a direct reflection of processing ease, but rather an indicator for the overall coherence of the stimulus item. given that the on-line measure operates on incomplete linguistic input to offer insight into incremental processing, while the off-line measure elicits a response after the full item has been presented, differences between the two measures are not entirely surprising. in the english experiment, the reading times confirmed a facilitative effect of lexical cues on processing the concessive relation, albeit at delayed impact. however, contextual cues were found not to have any effect in the on-line processing results. delayed effects are not uncommon in self-paced reading, but the obvious absence of an effect that was found in german requires further exploration. we will discuss possible explanations when we turn to the comparison of the german and english data below. critically, despite the low frequency of ‘true/sure...but’ constructions reported in section 3, the lexical cue had a facilitative effect on processing. earlier, we have raised the possibility that (at least some) comprehenders may not be able to generate discourse expectation from true/sure due to limited familiarity with this type of construction. overall, however, this is not the case. interestingly, the effect of the lexical cue was smaller in english than in german, which might be a reflection of a reduced utility of true/sure as discourse cue. we leave this issue for future research. while it is possible that the infrequency of the true/sure construction would cause general processing difficulties, such effects are not apparent in our regions of interest. neither the pre-critical (cr-1), nor the critical region (cr) showed elevated reading times for the conditions with a lexical cue. thus, even if subjects experienced a temporary processing difficulty at true/sure, these appear to have subsided by the time the critical region was reached. in any case, frequency effects would predict longer reading times if the sentence contained true/sure, not shorter ones. thus, the reduced reading times at the post-critical region cannot be attributed to a confound of the structural frequency. 95 schwab and liu as for the naturalness ratings, several findings stand out: for one, the two conditions containing a lexical cue received the lowest numerical ratings (c1: 5.73 and c2: 5.65 on average). we believe that this may be a reflection of the relatively low frequency of ‘true/sure,...’ constructions in natural language, as reported in section 3.3. secondly, we find that the condition containing a contextual cue, but no lexical cue, was rated the most natural. in analogy to the german results, the context may thus have added to the discourse coherence. further adding a lexical cue to contextually cued items significantly worsened the ratings, suggesting that subjects dispreferred lexically cueing a concessive discourse relation if the contextual information already signals a concession. together, the two experiments provide converging evidence for the effects of contextual and lexical cues in discourse processing. for both languages, the off-line ratings suggest that contextual cues improve the coherence of the discourse, while the on-line results indicate that lexical cues facilitate processing through the anticipation of upcoming discourse relations. the german study further demonstrated an on-line effect of contextual cues which was absent in english. one may speculate whether general differences between the languages could have contributed to the divergence between german and english on the contextual cue effect. if anything, however, we would have expected english native speakers to rely on contextual cues more heavily, given that a two-part marking as with ‘true/sure...but’ or ‘zwar...aber’ is generally less frequent in english than it is in german. we therefore think this question deserves further investigation in the future. for the current experiments, we suspect that a likely contributor to the divergence between german and english is the difference in data acquisition. we would like to draw attention to the fact the german experiment was conducted in the lab of the first author’s home university, while the english experiment has been conducted online over amazon mechanical turk. in general, the replicability of lab-based experiments in linguistics through crowdsourcing is well established (enochson and culbertson, 2015; munro et al., 2010; schnoebelen and kuperman, 2010; snow et al., 2008). however, there are a number of factors adding extra variability to crowdsourced data (broader demographics, e.g. on the age and educational level of participants, lack of control over experimental setting, and limited control over maintained attentiveness during the experiment). for subtle pragmatic effects like our context manipulation, these factors may indeed make a difference. as with most lab-based experiments, the participants for our german experiment were closely matched with regard to age and education level (undergraduate and graduate students with ages ranging from 1830). in our english experiment, participants were on average older (age range 19-70). in addition, we used a comprehension question after every item as quality control measure, and removed all participants with a response accuracy below 80%. a full quarter of mturk participants had to be excluded, hinting that they were on average less attentive readers than our lab participants. while those participants remaining in the sample passed our comprehension questions, we think it is likely that there still important differences remaining between the lab-based and crowdsourced data. past studies on the demographics of mturk workers, for instance, suggest that they cover a broad range of education levels and socio-economic statuses (ross et al., 2010). as predictive processing during language comprehension has been suggested to be affected by age (though the literature is unclear on the effect of age, with some suggesting a decreased use of prediction (federmeier et al., 2010; wlotko et al., 2012), while others have found either no clear effect (dave et al., 2018), or an effect in the opposite direction (cheimariou, 2016)), working memory (huettig and janse, 2016), as well as language and literacy skills (kukona et al., 2016; falkauskas and kuperman, 2015), the increased variability in our english participants’ backgrounds may mean that, across our sample, not all participants may engage in predictive language processing as reliably as the university students 96 lexical and contextual cue effects in discourse expectations participating in the german experiment. for that reason, we assume that a contextual cue effect could still be present in english, and detectable with a larger sample size. we defer a thorough investigation of this question to future studies. it is a longstanding idea that discourse relations are signaled through complex means that include, but go beyond, the presence of discourse markers (asr and demberg, 2015; das and taboada, 2018; hoek et al., 2018; prasad et al., 2010). our study provides experimental evidence that confirms that readers are sensitive to subtle discourse cues, both in judging the coherence of discourse and in processing discourse relations. we have shown that readers can rapidly integrate multiple cues to generate and update expectations on how the discourse will unfold. we would like to conclude this paper by discussing what we consider to be fruitful avenues for further research. as language is a means for interpersonal communication, we ought to integrate both speaker and listener data for a better understanding of the communicative function of discourse cues outside of discourse connectives. while we have identified and given evidence for one of their functions, namely that they provide a processing benefit to readers, we must ask ourselves (i) whether this processing benefit is among the deciding factors as to the speaker’s choice to add or delete discourse signals or contextual information towards a discourse relation, and (ii) what other factors affect the speaker’s choice. a speaker may be more inclined to include additional cues in their utterance when the cues result in a strong processing benefit for the reader or listener, that is, for instance in circumstances where establishing the discourse relation would be difficult otherwise. there are at least two possibilities. first, some connectives are ambiguous with regard to their signaled coherence relations, such that additional cues may allow the comprehender to disambiguate the intended relation. for example, aber can be used as a device to change topics, such as in (21). (21) danke, thanks, mir me geht goes es it gut. good. aber but wie how geht goes es it dir? you? ‘thanks, i am doing fine. but how are you?’ the use of aber in (21) is incompatible with zwar. thus, the presence of zwar can exclude alternative readings and narrow down the discourse relation to a concessive one. second, in case of a long distance between the two related propositions, the prediction from discourse cues enables comprehenders to keep the first proposition active in memory until the expected discourse continuation is fulfilled. an example for the latter point are intersentential discourse relations with intervening sentences as in (15b), repeated as (22) below. (22) true, she probably still had a long way to go. she was painfully thin and there was an insubstantiality about her. but there was no denying that today his wife was better than he had known her for many many months. (bnc cde 222-224) future research on a wider range of discourse relations and text/dialogue genres will have to determine if these factors can be shown to systematically predict whether a speaker will include additional discourse cues. simultaneously, a thorough investigation of comprehenders’ use of such cues in the on-line processing of various discourse relations can complement the speaker data. in the end, we believe that integrating insight from both perspectives is a promising approach towards a better understanding of discourse-level communication. 97 schwab and liu acknowledgements this work has benefited from valuable comments of the audiences at the conferences detec 2019, cuny 2020, and the workshop “explicit and implicit coherence relations: different, but how exactly?” (hu berlin), as well as from insightful discussions with lyn frazier. we would like to thank the anonymous reviewers of this article and our editor vera demberg for their invaluable comments and constructive feedback. disclosure statement the authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. funding j.s. is funded by the dfg research training group ‘computational cognition’ (dfg-grk 2340). this work was also supported by a dfg grant to m.l.’s project within the priority-program xprag.de on ‘the semantics and pragmatics of conditional connectives: cross-linguistic and experimental perspectives’ (project number: 367088975). notes on contributors j.s. and m.l. conceived and planned all the studies in the paper. j.s. carried out the experiments and corpus studies, and conducted data analyses. both authors contributed to the interpretation of the results. j.s. took the lead in preparing the manuscript, with contributions from m.l. references charles antaki and margaret wetherell. show concessions. discourse studies, 1(1):7–27, 1999. doi: 10.1177/1461445699001001002. fatemeh torabi asr and vera demberg. uniform information density at the level of discourse relations: negation markers and discourse connective omission. in proceedings of the 11th international conference on computational semantics, pages 118–128, london, uk, 2015. url https://www.aclweb.org/anthology/w15-0117.pdf. fatemeh torabi asr and vera demberg. interpretation of discourse connectives is probabilistic: evidence from the study of but and although. discourse processes, 57(4):376–399, 2020. doi: 10.1080/0163853x.2019.1700760. laura baranzini and alda mari. from epistemic modality to concessivity: alternatives and pragmatic reasoning per absurdum. journal of pragmatics, 142:116–138, 2019. doi: 10.1016/j. pragma.2019.01.002. dale j. barr, roger levy, christoph scheepers, and harry j. tily. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3): 255–278, 2013. doi: 10.1016/j.jml.2012.11.001. 98 https://www.aclweb.org/anthology/w15-0117.pdf lexical and contextual cue effects in discourse expectations douglas bates, martin maechler, ben bolker, steven walker, rune h. b. christensen, henrik singmann, bin dai, fabian scheip, gabor grothendieck, peter green, and john fox. package ’lme4’, 2018. george e. p. box and david r. cox. an analysis of transformations. journal of the royal statistical society: series b (methodological), 26(2):211–243, 1964. doi: 10.1111/j.2517-6161.1964. tb00553.x. lynn carlson, daniel marcu, and mary ellen okurowski. rst discourse treebank, ldc2002t07, 2002. url https://catalog.ldc.upenn.edu/ldc2002t07. spyridoula cheimariou. prediction in aging language processing. phd thesis, university of iowa, 2016. debopam das and maite taboada. signalling of coherence relations in discourse, beyond discourse markers. discourse processes, 55(8):743–770, 2018. doi: 10.1080/0163853x.2017.1379327. debopam das, tatjana scheffler, peter burgonje, and manfred stede. constructing a lexicon of english discourse connectives. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 360–365, melbourne, aus, 2018. doi: 10.18653/v1/w18-5042. shruti dave, trevor a. brothers, matthew j. traxler, fernanda ferreira, john m. henderson, and tamara y. swaab. electrophysiological evidence for preserved primacy of lexical prediction in aging. neuropsychologia, 117:135–147, 2018. doi: 10.1016/j.neuropsychologia.2018.05.023. heiner drenhaus, vera demberg, judith köhne, and francesca delogu. incremental and predictive discourse processing based on causal and concessive discourse markers: erp studies on german and english. in proceedings of the 36th annual meeting of the cognitive science society, pages 403–408, québec, can, 2014. url https://escholarship.org/uc/item/ 9q88v0zh. alex drummond. ibex farm, 2013. url http:spellout.net/ibexfarm. kelly enochson and jennifer culbertson. collecting psycholinguistic response time data using amazon mechanical turk. plos one, 10(3), 2015. doi: 10.1371/journal.pone.0116946. kaitlin falkauskas and victor kuperman. when experience meets language statistics: individual variability in processing english compound words. journal of experimental psychology: learning, memory, and cognition, 41(6):1607–1627, 2015. doi: 10.1037/xlm0000132. kara d. federmeier, marta kutas, and rina schul. age-related and individual differences in the use of prediction during language comprehension. brain and language, 115(3):149–161, 2010. doi: 10.1016/j.bandl.2010.07.006. fernanda ferreira, john m. henderson, michael d. anes, phillip a. weeks, and david k. mcfarlane. effects of lexical frequency and syntactic complexity in spoken-language comprehension: evidence from the auditory moving-window technique. journal of experimental psychology: learning, memory, and cognition, 22(2):324–335, 1996. doi: 10.1037/0278-7393.22.2.324. 99 https://catalog.ldc.upenn.edu/ldc2002t07 https://escholarship.org/uc/item/9q88v0zh https://escholarship.org/uc/item/9q88v0zh http: spellout.net/ibexfarm schwab and liu anastasia giannakidou. polarity sensitivity as (non) veridical dependency. john benjamins publishing company, amsterdam/philadelphia, 1998. doi: 10.1075/la.23. susanne günthner. concessive patterns in interaction: uses of zwar. . . aber (‘true. . . but’)constructions in everyday spoken german. language sciences, 58:144–162, 2016. doi: 10.1016/j.langsci.2016.02.009. lea a. hald, esther g. steenbeek-planting, and peter hagoort. the interaction of discourse context and world knowledge in online sentence comprehension. evidence from the n400. brain research, 1146(1):210–218, 2007. doi: 10.1016/j.brainres.2007.02.054. jet hoek, sandrine zufferey, jacqueline evers-vermeul, and ted j.m. sanders. the linguistic marking of coherence relations: interactions between connectives and segment-interal elements. pragmatics & cognition, 25(2):276–309, 2018. doi: 10.1075/pc.18016.hoe. falk huettig and esther janse. individual differences in working memory and processing speed predict anticipatory spoken language processing in the visual world. language, cognition and neuroscience, 31(1):80–93, 2016. doi: 10.1080/23273798.2015.1047459. christina s. kim, christine gunlogson, michael k. tanenhaus, and jeffrey t. runner. contextdriven expectations about focus alternatives. cognition, 139:28–49, 2015. doi: 10.1016/j. cognition.2015.02.009. philipp koehn. europarl: a parallel corpus for statistical machine translation. in machine translation summit x, pages 79–86, phuket, tha, 2005. url http://www.mt-archive.info/ mts-2005-koehn.pdf. judith köhne and vera demberg. the time-course of processing discourse connectives. in proceedings of the 35th annual meeting of the cognitive science society, pages 2760–2765, berlin, ger, 2013. url https://escholarship.org/uc/item/3ng7w640. ekkehard könig. concessive clauses. in keith brown, editor, encyclopedia of language & linguistics, pages 820–824. elsevier ltd, 2006. doi: 10.1016/b0-08-044854-2/00277-7. ekkehard konig and peter siemund. causal and concessive clauses: formal and semantic relations. in elizabeth couper-kuhlen and bernd kortmann, editors, cause condition concession contrast: cognitive and discourse perspectives, pages 341–360. mouton de gruyter, berlin/new york, 2000. doi: 10.1515/9783110219043.4.341. anuenue kukona, david braze, clinton l. johns, w. einar mencl, julie a. van dyke, james s. magnuson, kenneth r. pugh, donald p. shankweiler, and whitney tabor. the real-time prediction and inhibition of linguistic outcomes: effects of language and literacy skill. acta psychologica, 171:72–84, 2016. doi: 10.1016/j.actpsy.2016.09.009. roger levy. expectation-based syntactic comprehension. cognition, 106(3):1126–1177, 2008. doi: 10.1016/j.cognition.2007.05.006. ashley g. lewis and marcel bastiaansen. a predictive coding framework for rapid neural dynamics during sentence-level language comprehension. cortex, 68:155–168, 2015. doi: 10.1016/j.cortex. 2015.02.014. 100 http://www.mt-archive.info/mts-2005-koehn.pdf http://www.mt-archive.info/mts-2005-koehn.pdf https://escholarship.org/uc/item/3ng7w640 lexical and contextual cue effects in discourse expectations robert munro, steven bethard, victor kuperman, vicky tzuyin lai, robin melnick, christopher potts, tyler schnoebelen, and harry tily. crowdsourcing and language studies: the new generation of linguistic data. in proceedings of the naacl hlt 2010 workshop on creating speech and language data with amazon’s mechanical turk, pages 122–130, 2010. url https://www.aclweb.org/anthology/w10-0719.pdf. mante s. nieuwland and jos j.a. van berkum. when peanuts fall in love: n400 evidence for the power of discourse. journal of cognitive neuroscience, 18(7):1098–1111, 2006. doi: 10.1162/ jocn.2006.18.7.1098. wolfgang pfeifer, wilhelm braun, gunhild ginschel, gustav hagen, anna huber, klaus müller, heinrich petermann, gerlinde pfeifer, dorothee schröter, and ulrich schröter. zwar. in etymologisches wörterbuch des deutschen (digitalisierte und von wolfgang pfeifer überarbeitete version im digitalen wörterbuch der deutschen sprache). 1993. url https://www.dwds. de/wb/zwar. r. prasad, a. joshi, and b. webber. realization of discourse relations by other means: alternative lexicalizations. in proceedings of the 23rd international conference on computational linguistics: posters, pages 1023–1031, beijing, chn, 2010. url https://www.aclweb.org/ anthology/c10-2118. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of the 6th international conference on language resources and evaluation (lrec2008), pages 2961–2968, marrakech, mar, 2008. url http://www.lrec-conf.org/proceedings/lrec2008/pdf/ 754_paper.pdf. antoinette primatarova-miltscheva. zwar, aber – ein zweiteiliges konnektivum? deutsche sprache, 14(2):125–139, 1986. r development core team. r: a language and environment for statistical computing, 2019. url https://cran.r-project.org/. keith rayner and susan a. duffy. lexical complexity and fixation times in reading: effects of word frequency, verb complexity, and lexical ambiguity. memory & cognition, 14(3):191–201, 1986. doi: 10.3758/bf03197692. hannah rohde and william s. horton. anticipatory looks reveal expectations about discourse relations. cognition, 133(3):667–691, 2014. doi: 10.1016/j.cognition.2014.08.012. hannah rohde, r. levy, and a. kehler. anticipating explanations in relative clause processing. cognition, 118(3):339–358, 2011. doi: 10.1016/j.cognition.2010.10.016. joel ross, lilly irani, m. six silberman, andrew zaldivar, and bill tomlinson. who are the crowdworkers? shifting demographics in amazon mechanical turk. in conference on human factors in computing systems proceedings, pages 2863–2872, 2010. doi: 10.1145/1753846.1753873. 101 https://www.aclweb.org/anthology/w10-0719.pdf https://www.dwds.de/wb/zwar https://www.dwds.de/wb/zwar https://www.aclweb.org/anthology/c10-2118 https://www.aclweb.org/anthology/c10-2118 http://www.lrec-conf.org/proceedings/lrec2008/pdf/754_paper.pdf http://www.lrec-conf.org/proceedings/lrec2008/pdf/754_paper.pdf https://cran.r-project.org/ schwab and liu tatjana scheffler and manfred stede. mapping pdtb-style connective annotation to rst-style discourse annotation. in proceedings of the 13th conference on natural language processing, pages 242–247, bochum, ger, 2016. url https://www.linguistics.rub.de/ konvens16/pub/31_konvensproc.pdf. hildur e.h. schilling, keith rayner, and james i. chumbley. comparing naming, lexical decision, and eye fixation times: word frequency effects and individual differences. memory and cognition, 26(6):1270–1281, 1998. doi: 10.3758/bf03201199. tyler schnoebelen and victor kuperman. using amazon mechanical turk for linguistic research. psihologija, 43(4):441–464, 2010. doi: 10.2298/psi1004441s. merel c. j. scholman, vera demberg, and hannah rohde. signalling with one hand: a crosslinguistic comparison of the facilitative effect of “on the one hand.”. talk presented at xprag.de workshop “implicit and explicit marking of discourse relations”, osnabrück, ger, 2018. merel c.j. scholman, hannah rohde, and vera demberg. “on the one hand” as a cue to anticipate upcoming discourse structure. journal of memory and language, 97:47–60, 2017. doi: 10.1016/ j.jml.2017.07.010. nathaniel j. smith and roger levy. the effect of word predictability on reading time is logarithmic. cognition, 128(3):302–319, 2013. doi: 10.1016/j.cognition.2013.02.013. rion snow, brendan o. connor, daniel jurafsky, andrew y. ng, dolores labs, and capp st. cheap and fast — but is it good? evaluating non-expert annotations for natural language tasks. in emnlp’08: proceedings of the conference on empirical methods in natural language processing, pages 254–263, 2008. doi: 10.3115/1613715.1613751. jean-claude souesme. may in concessive contexts. in r. salkie, p. busuttil, and j. van der auwera, editors, modality in english – theory and description, pages 159–176. de gruyter mouton, berlin/new york, 2009. doi: 10.1515/9783110213331.159. manfred stede and carla umbach. dimlex: a lexicon of discourse markers for text generation and understanding. in proceedings of coling-acl ’98, montreal, can, 1998. url http: //www.ling.uni-potsdam.de/˜stede/papers/coling98.pdf. manfred stede, tatjana scheffler, and amália mendes. connective-lex: a web-based multilingual lexical resource for connectives. discours. revue de linguistique, psycholinguistique et informatique. a journal of linguistics, psycholinguistics and computational linguistics, 24, 2019. maite taboada and debopam das. annotation upon annotation: adding signalling information to a corpus of discourse relations. dialogue & discourse, 3(2):249–281, 2013. doi: 10.5087/dad. 2013.211. the british national corpus. version 3 (bnc xml edition), distributed by bodleian libraries, university of oxford, on behalf of the bnc consortium, 2007. url http://www.natcorp. ox.ac.uk/. peter uhrig and thomas proisl. treebank.info, 2011. url http://treebank.info/. 102 https://www.linguistics.rub.de/konvens16/pub/31_konvensproc.pdf https://www.linguistics.rub.de/konvens16/pub/31_konvensproc.pdf http://www.ling.uni-potsdam.de/~stede/papers/coling98.pdf http://www.ling.uni-potsdam.de/~stede/papers/coling98.pdf http://www.natcorp.ox.ac.uk/ http://www.natcorp.ox.ac.uk/ http://treebank.info/ lexical and contextual cue effects in discourse expectations geertje van bergen and hans rutger bosker. linguistic expectation management in online discourse processing: an investigation of dutch inderdaad ‘indeed’ and eigenlijk ‘actually’. journal of memory and language, 103:191–209, 2018. doi: 10.1016/j.jml.2018.08.004. shravan vasishth and michael broe. the foundations of statistics: a simulation-based approach. springer verlag, berlin/heidelberg, 2011. doi: 10.1007/978-3-642-16313-5. naoko witzel, jeffrey witzel, and kenneth forster. comparisons of online reading paradigms: eye tracking, moving-window, and maze. journal of psycholinguistic research, 41(2):105–128, 2012. doi: 10.1007/s10936-011-9179-x. edward w. wlotko, kara d. federmeier, and marta kutas. to predict or not to predict: age-related differences in the use of sentential context. psychology and aging, 27(4):975–988, 2012. doi: 10.1037/a0029206. ming xiang and gina kuperberg. reversing expectations during discourse comprehension. language, cognition and neuroscience, 30(6):648–672, 2015. doi: 10.1080/23273798.2014.995679. 103 schwab and liu appendix a a1: stimuli for experiment 1 below we list the test sentences used for the first experiment. items were presented in 4 conditions that can be reconstructed by adding either one or both of the optional words in parentheses to the sentence. (1) gregor isst gerne (auswärts). er besitzt (zwar) eine küche mit profiausstattung, aber er isst meistens im restaurant. (2) jens läuft gerne (draußen). er hat (zwar) ein laufband im wohnzimmer, aber er joggt häufig im park. (3) frederike lernt gerne (alleine). sie teilt (zwar) eine lerngruppe mit freunden, aber sie büffelt regelmäßig im alleingang. (4) thomas schwimmt gerne (draußen). er kennt (zwar) ein hallenbad im stadtzentrum, aber er fährt oft ans meer. (5) marc wandert gerne (querfeldein). er kennt (zwar) einen wanderweg im wald, aber er läuft oft durchs walddickicht. (6) jan malt gerne (draußen). er hat (zwar) ein atelier im dachgeschoss, aber er malt momentan im garten. (7) daniela spielt gerne (online). sie hat (zwar) eine brettspielsammlung im schrank, aber sie spielt momentan am computer. (8) finn trainiert gerne (gemeinsam). er besitzt (zwar) eine hantelbank im keller, aber er geht täglich ins fitnessstudio. (9) ronja arbeitet gerne (zuhause). sie teilt (zwar) einen büroraum mit kollegen, aber sie arbeitet häufig im bett. (10) tim kickert gerne (draußen). er nutzt (zwar) eine sporthalle im winter, aber er spielt jetzt auf rasenplätzen. (11) nele verreist gerne (spontan). sie plant (zwar) eine rundreise im sommerurlaub, aber sie entscheidet spontan übers reiseziel. (12) petra fotografiert gerne (analog). sie benutzt (zwar) eine digitalkamera im fotostudio, aber sie fotografiert häufig auf filmrollen. (13) lucas arbeitet gerne (tagsüber). er akzeptiert (zwar) einen job im nachtclub, aber er sucht aktuell nach jobalternativen. (14) mayra schreibt gerne (handschriftlich). sie hat (zwar) einen computer im wohnzimmer, aber sie schreibt meistens auf papier. (15) mika arbeitet gerne (handwerklich). er hat (zwar) einen bürojob im rathaus, aber er schreinert täglich im hobbykeller. 104 lexical and contextual cue effects in discourse expectations (16) dennis klettert gerne (sicherungsfrei). er besucht (zwar) einen seilgarten mit freunden, aber er klettert meist im gebirge. (17) sina trainiert gerne (frühmorgens). sie leitet (zwar) einen abendkurs im sportverein, aber sie trainiert regelmäßig vorm arbeiten. (18) albert shoppt gerne (online). er besucht (zwar) einen laden im einkaufsviertel, aber er bestellt meistens übers internet. (19) helena wandert gerne (alpin). sie unternimmt (zwar) eine wattwanderung im urlaub, aber sie läuft oft im gebirge. (20) jana lernt gerne (visuell). sie besitzt (zwar) ein textbuch zum prüfungsthema, aber sie lernt momentan über videos. (21) lena singt gerne (zuhause). sie besucht (zwar) eine karaokebar mit freunden, aber sie verbleibt aktuell als zuschauerin. (22) hannah liest gerne (abends). sie bekommt (zwar) eine zeitung am morgen, aber sie liest täglich beim abendessen. (23) susanne musiziert gerne (gemeinsam). sie beherrscht (zwar) ein solostück im klavier, aber sie spielt jetzt im orchester. (24) johannes spielt gerne (drinnen). er kennt (zwar) einen spielplatz im stadtpark, aber er bleibt momentan im zimmer. (25) hannes tanzt gerne (allein). er besucht (zwar) einen tangokurs mit freunden, aber er tanzt regelmäßig im nachtclub. (26) johanna reist gerne (international). sie erwägt (zwar) einen nordseeurlaub auf langeoog, aber sie fliegt jetzt nach japan. (27) nicole recherchiert gerne (online). sie kennt (zwar) eine bibliothek am stadtrand, aber sie recherchiert aktuell am computer. (28) ole performt gerne (live). er produziert (zwar) ein album im studio, aber er spielt täglich auf konzerten. 105 schwab and liu a2: stimuli for experiment 2 below we list the test sentences used for the second experiment. items were presented in 4 conditions that can be reconstructed by adding either one or both of the optional words in parentheses to the sentence. (1) gregory likes to eat (out). (true,) he has an induction stove for the kitchen, but he mostly eats in restaurants. (2) james likes to run (outdoors). (sure,) he has a treadmill in the living room, but he primarily runs in parks. (3) fred likes to study (alone). (true,) he joins a study group with the classmates, but he often studies at home. (4) anna likes to swim (outside). (sure,) she knows an indoor pool in the city center, but she regularly drives to beaches. (5) emma likes to dance (outside). (true,) she attends a dance class in the campus dance room, but she primarily dances in parks. (6) hannah likes to paint (alone). (sure,) she attends an art class with the classmates, but she mostly paints in solitude. (7) john likes to hike (off-road). (true,) he knows a hiking trail in the mountains, but he often walks through thickets. (8) william likes to travel (alone). (sure,) he visits a friend for the summer vacation, but he mostly travels in solitude. (9) lucas likes to cook (alone). (true,) he takes a cooking class at the culinary school, but he mostly cooks in solitude. (10) chloe likes to read (poems). (sure,) she reads a newspaper on the train, but she regularly reads in poetry collections. (11) leah likes to sing (live). (true,) she records an album in the studio, but she often performs at concerts. (12) maya likes to act (live). (sure,) she has a television role on the regional station, but she regularly acts on stage. (13) daniel likes to hear music (live). (true,) he plays a radio in the bedroom, but he often goes to concerts. (14) david likes to play games (online). (sure,) he has a board game in the cupboard, but he mostly plays on computers. (15) catherine likes to shop (online). (sure,) she visits a store in the mall, but she primarily shops in online stores. 106 lexical and contextual cue effects in discourse expectations (16) leo likes to train (together). (sure,) he has a weight set in the bedroom, but he often trains in fitness studios. (17) george likes to bake (alone). (true,) he shares a kitchen in the flat, but he mostly bakes in solitude. (18) jessica likes to study (interactively). (sure,) she owns a book about the exam topic, but she primarily studies with friends. (19) gabrielle likes to do research (online). (true,) she knows a library in the city center, but she primarily researches on computers. (20) jane likes to relax (inside). (sure,) she has a balcony in the sun, but she primarily relaxes in bed. (21) jonah likes to sunbathe (outside). (true,) he has a tanning membership at the gym, but he regularly lays out at beaches. (22) mark likes to work (outside). (sure,) he has an office job at the government, but he regularly gardens on weekends. (23) amelie likes to nap (inside). (true,) she has a hammock in the garden, but she often naps in bed. (24) sarah likes to play (outside). (true,) she has a toy collection in the living room, but she regularly goes to playgrounds. 107 schwab and liu appendix b below, we report an extended version of table 2. in this version, we lists all cues and connectives that had been grouped into the ‘other’ category. lexical cue (on conceded argument) in source text connective (on second argument) in source text true (102) /sure (3) but however nevertheless though even so yet still although nonetheless none lexical cue in german connective on 2nd argument (german) zwar aber (‘but’) 17 1 – – – – – – 1 – (31) doch (‘however’) 5 2 – – – – – – – – jedoch (‘however’) 3 1 – 1 – – – – – – stimmt (‘correct’) aber (‘but’) 12 1 1 – – – (21) allerdings (‘however’) – 1 – – – – dennoch (‘still’) – – 1 – – – doch (‘however’) 1 – – – – – – – – 1 jedoch (‘however’) 1 1 – – – – – – – – no connective – 1 – – – – – – – – sicher (‘sure’) aber (‘but’) 3 – – – – – (6) allerdings (‘however’) – 1 – – – – doch (‘however’) – – – 1 – – dennoch (‘still’) – – – – – – 1 (sure) – – – sicherlich (‘surely’) aber (‘but’) 2 (1 sure) – – – – – – – – – (3) jedoch (‘however’) 1 – – – – – – – – – natürlich (‘natural(y)’) aber (‘but’) 3 (1 sure) – – – – – (8) doch (‘however’) 3 – – – – – – – – – jedoch (‘however’) – 1 1 – – – – – – – gewiss (‘certain(ly)’) aber (‘but’) 2 – – – – – – – – 1 (6) dennoch (‘still’) – – 1 – – – no connective – – – 1 1 – – – – – allerdings (‘however) jedoch (‘however’) 1 – – – – – – – – – einverstanden (‘agreed’) aber (‘but’) – – – – – – – – 1 – in der tat (‘indeed’) aber (‘but’) 1 – – – – – – – – – (2) allerdings (‘however’) 1 – – – – – – – – – ist wahr (‘is true’) aber (‘but’) 1 – – – – – – – – – (3) dennoch (‘still’) – – – – – 1 – – – – jedoch (‘however’) – 1 – – – – – – – – ja (‘yes’) aber (‘but’) 1 1 – – – – – – – – (3) no connective – 1 – – – – – – – – nun gut (‘well’) aber (‘but’) 1 – – – – – – – – – richtig (‘right’) aber (‘but’)3 – – – – – – – – – (4) doch (‘however’) – 1 – – – – – – – – tatsächlich (‘indeed’) jedoch (‘however’) 1 – – – – – – – – – zutreffen (‘to be the case’) aber (‘but’) 1 – – – – – – – – – unbestritten (‘uncontested’) aber (‘but’) 1 – – – – – – – – – wenn auch (‘even if’) aber (‘but’) – – – – – – – 1 – – wirklich (‘really’) allerdings (‘however’) 1 – – – – – – – – – (2) doch (‘however’) – 1 – – – – – – – – zugegeben (‘admittedly’) aber (‘but’) 1 – – – – – – – – – zugegebenermaßen (‘admittedly’) doch (‘however’) – – – – – – – – – 1 zweifellos (‘undoubtedly’) aber (‘but’) 1 – – – – – – – – – (3) doch (‘however’) 1 – – – – – – – – – jedoch (‘however’) 1 – – – – – – – – – no lexical cue aber (‘but’) 2 – – – – – – – – – (4) doch (‘however’) 1 – – – – – – – – – doch (‘however’) 1 – – – – – – – – – table 6: results of a parallel corpus analysis using the german-english aligned version of europarl. the upper left corner indicates the lexical cue that was present in the original text source (true/sure). the two leftmost columns indicate the lexical cue on the conceded argument and the connective on the second argument that were identified in the german translation. entries in the table provide a count of the frequency with which a combination of true/sure+connective was translated into a corresponding cue+connective in german. for example, the first entry of the table (17) shows that ‘true/sure..., but’ was translated to ‘zwar..., aber’ 17 times. 108 lexical and contextual cue effects in discourse expectations l ex ic al cu e (o n co nc ed ed ar gu m en t) in so ur ce te xt c on ne ct iv e (o n se co nd ar gu m en t) in so ur ce te xt zw ar (5 00 ) ab er do ch je do ch de nn oc h al le rd in gs an de re rs ei ts au ch w en n da ge ge n hi ng eg en ni ch ts de st ro tz so nd er n tr ot zd em w ie de ru m w oh in ge ge n no ne c ue /c on ne ct iv e in e ng lis h c on ne ct iv e on se co nd ar gu m en ti n e ng lis h w hi le ne ve rt he le ss 1 1 – 1 – – – – – – – – – – – (6 1) on th e ot he r ha nd – 1 – – – – – – – – – – – – – no fu rt he rc on ne ct iv e 28 11 13 2 1 – – – – – – – 1 1 – w hi ls t ho w ev er 1 – – – – – – – – – – – – – – (2 0) no fu rt he rc on ne ct iv e 9 5 4 1 – – – – – – – – – – – al th ou gh ho w ev er – – 1 – – – – – – – – – – – – (6 8) ne ve rt he le ss 2 – – – – – – – – – – – – – – on th e ot he r ha nd – – – – – – – 1 – – – – – – – no fu rt he rc on ne ct iv e 34 14 12 – 2 – – – – 1 – – – – 1 th ou gh ho w ev er – – 1 – – – – – – – – – – – – (8 ) ne ve rt he le ss 1 – – – – – – – – – – – – – – on th e ot he r ha nd – – – – – – – – – – – – – – – no fu rt he rc on ne ct iv e 4 3 – – – – – – – – – – – – – m od al (m ay /m ig ht ) bu t 16 4 3 – – – – (2 8) ho w ev er 1 – 1 – – – – on th e ot he r ha nd – 1 – – – – no co nn ec tiv e 1 – – – – – – ad m itt ed ly bu t 4 – 2 – – – – – – – – – – – – (8 ) st ill 1 – – – – – – – – – – – – – – no co nn ec tiv e – – – – – – – – – – – – – – 1 to be aw ar e th at bu t 1 – – – – – – – – – – – – – – ce rt ai nl y bu t 4 1 3 – – – – – – – – – – – – (1 0) no co nn ec tiv e 2 – – – – – – – – – – – – – – cl ea rl y bu t 1 – – – – – – – – – – – – – – co nd iti on al (i f.. .th en ) no co nn ec tiv e 3 4 – – – – – – – – – – – – – de sp ite no co nn ec tiv e – – 1 – – – – – – – – – – – – ev en if no co nn ec tiv e 1 2 – – – – – – – – – – – – 1 ev en th ou gh no co nn ec tiv e 1 – – – – – – – – – – – – – – gr an te d bu t 1 1 – – – – – – – – – – – – – to kn ow th at bu t 2 – – – – – – – – – – – – – – in fa ct bu t 1 – – – – – – – – – – – – – – (3 ) ho w ev er – 1 – – 1 – – – – – – – – – – in de ed bu t 1 – – – – – – – – – – – – – – (4 ) no ne th el es s – – – 1 – – – – – – – – – – – ye t – – 1 – – – – – – – – – – – – no co nn ec tiv e – 1 – – – – – – – – – – – – – it is tr ue bu t 3 5 – 1 – – – – – – – – – – – (1 3) ho w ev er 1 – 1 – – – – – – – – – – – – ne ve rt he le ss 1 – – – – – – – – – – – – – – no co nn ec tiv e – – – – – – – – – – – – – – 1 ob vi ou sl y ho w ev er 1 – – – – – – – – – – – – – – of co ur se bu t 3 – – – – – – – – – – – – – – (4 ) ho w ev er – – 1 – – – – – – – – – – – – on th e on e ha nd on th e ot he r ha nd 1 – – – – – – – – – – – – – – to re co gn iz e th at bu t – 1 – – – – – – – – – – – – – to be su re bu t 1 – – – – – – – – – – – – – – no le xi ca lc ue bu t 10 6 49 44 – 2 – 1 – – – 1 1 – – – (2 48 ) ho w ev er 1 2 2 – 1 – – ne ve rt he le ss – – 1 – – – – ye t 2 3 1 1 – – – no co nn ec tiv e 13 5 2 2 – 1 – – – – – – – – 7 ta bl e 7: t he up pe rl ef tc or ne ri nd ic at es th e le xi ca lc ue th at w as pr es en ti n th e or ig in al te xt so ur ce (z w ar ). t he tw o le ft m os tc ol um ns in di ca te th e le xi ca lc ue on th e co nc ed ed ar gu m en t( w ith to ta lf re qu en cy co un ts in pa re nt he se s) an d th e co nn ec tiv e on th e se co nd ar gu m en t th at w er e id en tifi ed in th e e ng lis h tr an sl at io n. e nt ri es in th e ta bl e pr ov id e a co un to f th e fr eq ue nc y w ith w hi ch a co m bi na tio n of zw ar +c on ne ct iv e in g er m an w as tr an sl at ed in to a co rr es po nd in g cu e+ co nn ec tiv e in e ng lis h. fo re xa m pl e, th e fir st en tr y of th e ta bl e (1 )s ho w s th at ‘z w ar ... ,a be r’ w as tr an sl at ed to ‘w hi le ... ,h ow ev er ’1 tim e in th e to ta ls am pl e of 50 0 in st an ce s of zw ar . 109 introduction the tested expressions frequency and reliability of zwar, true, and sure as lexical cues zwar...aber in german true/sure...but in english frequency of zwar, true, and sure summary testing expectations in discourse processing / experiments german experiment materials participants procedure results english experiment materials participants procedure results discussion microsoft word 2023_discoursedialogue dialogue & discourse 14(1) (2023) 88–124 doi: 10.5210/dad.2023.104 ©2023 nien-heng wu, shu-chuan tseng this is an open-access article distributed under the terms of a creative commons attribution license (http: //creativecommons.org/licenses/by/3.0/). form and function of connectives in chinese conversational speech nien-heng wu nienhengwu@gmail.com institute of linguistics academia sinica shu-chuan tseng tsengsc@gate.sinica.edu.tw institute of linguistics academia sinica editor: vera demberg submitted 05/2022; accepted 05/2023; published online 06/2023 abstract connectives convey discourse functions that provide textual and pragmatic information in speech communication on top of canonical, sentential use. this paper proposes an applicable scheme with illustrative examples for distinguishing sentential, conclusion, disfluency, elaboration, and resumption uses of mandarin connectives, including conjunctions and adverbs. quantitative results of our annotation works are presented to gain an overview of connectives in a mandarin conversational speech corpus. a fine-grained taxonomy is also discussed, but it requires more empirical data to approve the applicability. by conducting a multinomial logistic regression model, we illustrate that connectives exhibit consistent patterns in positional, phonetic, and contextual features oriented to the associated discourse functions. our results confirm that the position of conclusion and resumption connectives orient more to positions in semantically, rather than prosodically, determined units. we also found that connectives used for all four discourse functions tend to have a higher initial f0 value than those of sentential use. resumption and disfluency uses are expected to have the largest increase in initial f0 value, followed by conclusion and elaboration uses. durational cues of the preceding context enable distinguishing sentential use from discourse uses of conclusion, elaboration, and resumption of connectives. keywords: discourse function, mandarin connectives, conversational corpus, phonetic representation 1. introduction connectives are a class of lexical items that signal the relationship between units of text or discourse, connecting two different abstract objects, such as events, states or propositions, in discourse (asher, 1993). it is posited that connectives have little lexical impact at the local segment level but serve significant pragmatic functions (hjalmarsson, 2011). in conversational discourse, the position in which connectives occur and the phonetic form of connectives provide cues that help listeners process speakers’ intentions and the structure of the ongoing discourse (didirková et al., 2018; hirschberg and litman, 1987, 1993; horne et al., 2001; rennie et al., 2016; rhee, 2020). the use of connectives in conversation may be aimed more at marking discourse structure than at referring to the canonical meaning of the connectives themselves. it wu and tseng 89 is our goal to investigate whether there are differences between the canonical use (sentential) and the functional use (discourse) of connectives and whether it is possible to effectively disambiguate between the two uses by looking into the associated phonetic properties in conversation. previous research on mandarin chinese has identified various discourse-pragmatic functions in an array of connectives (biq, 1994, 2001; wang and huang, 2006; wang, 2018; wang, 1998, 2005; wang et al., 2013; wang and tsai, 2007; wang et al., 2010; yang, 2006), but only limited results have described the relationship between the functions and the formrelated properties of connectives. for instance, the topic-shift ránhòu, ‘then’, has a significantly longer duration and pitch range larger than the canonical use, while the trail-off ránhòu (i.e., marking the closure of the current turn and inviting the hearer’s response) shows decreased loudness and durational lengthening. while the connection between the discourse function and the phonetic form of connectives may be strong (biq, 2001; wang, 2018; yang, 2006), it lacks a systematic schema to describe discourse functions of mandarin connectives and their phonetic forms. previous studies offer mostly qualitative descriptions of individual connectives. in this study, we pursue a corpus-based discourse function annotation scheme and a quantitative analysis of mandarin connectives and their phonetic features for discourse function disambiguation. the correlation examined through a statistically grounded method could be critical to the success of spoken language generation and understanding tasks. this study considers conjunctions and adverbs to be the main lexical categories of target connective (prasad et al., 2008; zufferey and degand, 2017). we conduct an annotation project on target connectives in a chinese conversational speech corpus and analyze the positional, phonetic, and contextual properties of the associated discourse functions in a multinomial logistic regression model. with this task, we investigate whether a connective’s phonetic forms orient to sentential/discourse uses. if they do, are we able to find consistent patterns in terms of specific discourse functions? if speakers show sensitivity to the distinction of sentential/discourse uses of mandarin connectives in their speech production, statistically significant coefficients that support consistent phonetic patterns in sentential/discourse uses are expected. this article is organized as follows. section 2 reviews the previous literature on mandarin connectives' discourse functions. section 3 presents the literature on connectives' form-related properties. in section 4, we present the data, the annotation scheme, and the descriptive results of the positional, phonetic, and contextual features. in section 5, we examine closely whether there is any coupling between the phonetic form and the discourse function of connectives by a multinomial logistic regression model. finally, we discuss our main findings in section 6. form and function of connectives in chinese conversational speech 90 2. discourse functions of mandarin connectives taking topic transitions in spoken discourse1 as the main focus in the consideration of discourse function grouping of mandarin connectives, previous works basically distinguish discourse functions that signal initiation (resumption) and conclusion of topics, elaboration of various types, and disfluency. first of all, wang (2005) and wang and tsai (2007) noted that the adverbs búguò, kěshì and dànshì, ‘but/yet/however’, may initiate a new topic in the discourse. since the adverbs conventionally imply a contrast in propositions (miracle, 1991; ross, 1978), they may extend to imply a contrast between the new topic and the old topic in discourse. similarly, wang et al. (2010) observed that the adverb qíshí, ‘actually’, functioned to introduce a new topic or a new aspect of the current topic that may not be in accordance with the current claims in discourse (biq, 1994). working with the conjunction suǒyǐ, ‘so’, wang and huang (2006) proposed that suǒyǐ may initiate a new topic in the discourse, as the consequences introduced by suǒyǐ can be treated as new information on the discourse level. however, they also noted that suǒyǐ may signal a resumption of a previous topic and prevent further departure into irrelevant topics. the topic-resuming function was also found in the conjunction ránhòu, ‘then’. yang (2006) and wang (2018) maintained that ránhòu, canonically indicating either a temporal or consequential relationship between two adjacent clauses, can extend to organize utterances in discourse. as such, ránhòu may signal not only a change in topic but also a return to a previous topic after an intervening subtopic. connectives can also be used to signal conclusion of a current topic. wang and huang (2006) noticed that the speaker employed suǒyǐ to paraphrase or summarize the previous talk so that no misunderstanding was ensured. relatedly, wang (2018) described a trail-off use (local and kelly, 1986) in the turn-final ránhòu tokens. she maintained that the trail-off use of ránhòu may express “the speaker’s intent to close the turn and to invite the hearer’s responses of various types, such as acknowledgement, comment, or elaboration on the current topic” (p. 18). more often in conversation, connectives allow the speaker to provide elaboration or clarification on a current topic. biq (2001) claimed that the conjunction jiùshì, ‘that is’, when being slightly semantically reduced, may signal elaboration or clarification of the previous utterance. the elaboration function was also noted in zhǐshì, ‘only’, typically a restrictive marker (guo, 1999; lü, 1980) or a focus marker (wang, 2005). wang (2005) and wang and tsai (2007) observed that the speaker may employ zhǐshì to introduce an afterthought that comments or elaborates on the previous utterance. however, due to its implicature of contrast, the piece of discourse that zhǐshì introduces may be incongruent with or divergent from the preceding discourse, such as a counterexpectation or surprising fact. wang et al. (2013) similarly claimed that zhǐshì may indicate “a more detailed or more correct formulation of 1 studies reviewed in this section mainly focused on spontaneous conversations and/or tv/radio interviews and callins. wu and tseng 91 something stated previously” (p. 203). wang and huang (2006) noted that ránhòu featured an additive use that introduces new information to the current topic and connects successive ideas in discourse (wang, 1998). propositionally, qíshí is a commentary adverb that delivers the speaker’s attitude toward the propositional content. wang et al. (2010) noticed that the canonical meaning of qíshí has developed several discourse functions that comment on the form or content of an utterance, such as elaborating on the previous utterance with more accurate or specific information. all utterances encode the speaker’s attitude about the proposition to some extent, which may be indicated by connectives. for instance, hsieh and huang (2005) concluded that a qíshíembedded clause may disclose a fact that the speaker believes that the hearer does not know. propositional attitudes have been discussed in the form of emphasizing or supporting a proposition and limiting the validity of a proposition. yang (2006) and wang et al. (2010) suggested that ránhòu and qíshí may place emphasis on the distinctiveness of the following content. wang et al. (2010) further contended that the fact-introducing qíshí may support and strengthen the speaker’s assertion. in contrast, wang et al. (2013) claimed that the elaboration function of zhǐshì may imply incompatibility or insufficiency of the previous utterance and consequently limit the validity of the proposition in the utterance. the downtoning effect may further imply the speaker’s mild negative stance toward the propositions and his or her intent to instruct the hearer to reject previous claims. connectives may convey some interactional meanings in the speaker-hearer interaction, such as grabbing the hearer’s attention and establishing (dis)alignment between the interlocutors’ stances. yang (2006) and wang et al. (2010) suggested in their respective projects that the emphasis of ránhòu and qíshí may also function to attract the hearer’s attention to the discourse. regarding the interlocutor stance, hsieh and huang (2005) and wang et al. (2010) found that the speaker may use qíshí to disalign, and sometimes align, him/herself with the previous speaker’s stance. wang (2005) and wang and tsai (2007) claimed that the contrastive búguò, kěshì and dànshì may preface the speaker’s dispreferred response, such as a rejection or disagreement, to the previous speaker’s utterance since disagreement can be one kind of contrastiveness (ford, 2000). while wang (2005) and wang and tsai (2007) did not explicitly note the function of expressing disagreement in zhǐshì, wang et al. (2013) lent support to such a description by claiming that the incompatibility introduced by zhǐshì may extend to establish a contrast in the interlocutors’ stances, allowing the speaker to express minor or indirect disalignment with the previous speaker’s claims. the authors also noticed that in their data, zhǐshì was sometimes prefaced by a brief pause, which signals the speaker’s hesitation to agree with the previous speaker. last, connectives may sometimes contribute no propositional meaning to the discourse. they may simply signal disfluencies such as filled pauses. in her study of jiùshì and its variants, biq (2001) noted that a more reduced jiùshì may serve as a filled pause or floor holder that does form and function of connectives in chinese conversational speech 92 not contribute to the proposition. this is evidenced by the fact that the utterance would not be understood differently if jiùshì were omitted. yang (2006) further added that in the case of ránhòu, the connective may act to perform floor negotiations such as floor holding and turntaking. these functions strategically enable the speaker to have more time to plan what to say next. 3. positional, phonetic, and contextual encoding of connectives in this section, related works on the positional, phonetic, and contextual encodings of connectives with various discourse functions are reviewed. section 3.1 presents the positional encoding. section 3.2 introduces various types of phonetic encoding, including pitch, intensity, and word duration correlates. section 3.3 covers contextual cues that often co-occur with connectives. 3.1 word position word position is the location of the lexical item being discussed in relation to a given unit in spoken discourse, be it a meaning-, interactionor prosody-oriented unit, for instance, intonational unit (hirschberg and litman, 1987, 1993), interpausal unit (ipu) (gravano et al., 2007; gravano et al., 2011), speaker turn (gravano et al., 2007; gravano et al., 2011; wang, 2018) or utterance (rennie et al., 2016; rhee, 2020). recognizing the importance of word position, hirschberg and litman (hirschberg and litman, 1987, 1993) differentiated between the discourse and sentential use of the english now based on their positions in an intonational phrase and intermediate phrase (pierrehumbert, 1980). they found that almost all tokens of discourse now (98.41%) were absolutely first or followed only another cue phrase in an intermediate phrase (e.g., well now, ok now), while only 13.5% of the tokens of sentential now were so placed. sentential now, on the other hand, tended to occur intermediate phrase-finally (59.45%), whereas only 1.58% of the tokens of discourse now did. similar tendencies were supported in another study of conjunctions (e.g., and, but, or, etc.), adverbs (e.g., actually, also, indeed, etc.), and other cue phrases (e.g., okay, say, like, etc.) (hirschberg and litman, 1993). the utterance or speaker turn position of connectives is relevant to the associated discourse function. investigating the english so, rennie et al. (2016) posited that the utteranceinitial and -second so may introduce either a topic shift or a conclusion of a previous utterance. in addition, the utterance-initial so may initiate a speaker change or a new utterance by the same speaker. on the other hand, the utterance-internal so may perform a resultative function that connects a new piece of information to an old utterance. the utterance-final so may release the speaker’s turn. last, the standalone so may function as a turn-yielding device that urges the hearer to continue with the dialog or as a filled pause that holds the floor for the speaker. rennie and colleagues, however, examined only the interaction between utterance positions and the two turn-organizing functions for the utterance-initial so. the interaction between utterance positions and other discourse functions remains unknown. investigating mak, ‘coarsely’, in wu and tseng 93 korean spontaneous conversations and scenarios of dramas and movies, rhee (2020) observed that the filled-pause use of mak had a tendency to occur utterance-finally, which may reflect the tendency of the speaker to perform lexical search and floor holding at the end of the utterance. despite stating that in principle, filled-pause use can occur anywhere in an utterance, rhee did not provide any statistics. concerning the relationship between discourse functions and speaker turn positions, wang (2018) found that some of the utterance-final tokens of the mandarin ránhòu had a turn-holding use in her conversational data, while the turn-final tokens all delivered the trail-off use, bringing the current talk to a closure and yielding the floor to another speaker. the trail-off use also occupied an independent intonation unit instead of being embedded in the previous intonation unit. her findings, however, suffered from the problem of data sparseness, as she found only six tokens of the trail-off ránhòu and did not report any quantitative data for the turn-holding ránhòu. it is unclear whether the observed positional property is significant. aside from the previously mentioned studies, gravano and colleagues (gravano et al., 2007; gravano et al., 2011) found that the english cue words alright and okay tended to occur initially or independently in an ipu, signaling the beginning of a discourse segment. in summary, all these findings point to a strong tendency that positional encoding reflects discourse functions of connectives in speech production. it is unclear, however, to what extent different positional encodings (e.g., the initial, medial, final, or standalone position) correlate with each discourse function. it also remains an empirical question of which kind of production unit can reliably and effectively disambiguate between discourse functions. 3.2 pitch, intensity, and duration previous studies have also reported phonetic evidence for various discourse functions. one phonetic encoding consistently discussed in the literature is pitch, which has been operationalized on different bases, such as pitch accent (hirschberg and litman, 1987, 1993), pitch reset (didirková et al., 2018; horne et al., 2001), and pitch range (wang, 2018; yang, 2006). hirschberg and litman (hirschberg and litman, 1987, 1993) found that the discourse use of the english now was more often deaccented than the sentential use of now. when forming part of a larger intermediate phrase, the majority of sentential uses of now received an h* or complex pitch accent, while all discourse uses of now bore an l* accent. investigating pitch reset, horne et al. (2001) and didirková et al. (2018) each described its relation to topic shift in narrative speech and in spontaneous speech. horne et al. (2001) observed a mean f0 reset for tokens of the swedish men, ‘but’, that performed a topic-shift function, similar to the size of the reset one would observe at a topic-shift boundary. the effect of the f0 reset also led to 70% accuracy in distinguishing between discourse and sentential men in a linear classifier model. analyzing the french alors, ‘then’, didirková et al. (2018) reported that the connective tended to be marked by a reset on the word when introducing a new topic or specification. a reset in pitch was also observed in the french et, ‘and’, when introducing specification. focusing on pitch range, both yang (2006) and wang (2018) described the way the phonetic parameter form and function of connectives in chinese conversational speech 94 manifested various functions of the mandarin ránhòu. yang (2006) showed that ránhòu, when signaling a topic shift or returning to a previous topic after an intervening subtopic, tended to have a larger pitch range and more perceptual prominence. in contrast, a use of ránhòu to signal a continuation of the current topic had a narrow pitch range and a more gradual and smoother contour. yang also added that as an emphasis marker or a floor-negotiating device, ránhòu was marked with an expanded pitch range, suggesting the hearer pay attention to the speaker. however, no quantitative data were reported. adopting a quantitative approach, wang (2018) showed supporting evidence for yang’s findings that the topic-shifting ránhòu had a much larger average pitch range than its other sentential uses. moreover, the turn-initial ránhòu showed a larger pitch range than the noninitial ránhòu. meaningful phonetic variation has also been observed in the study of intensity, commonly referred to as loudness or volume. rennie et al. (2016) compared the intensity of the english so and that of its adjacent segments. they showed that both the so tokens that introduced either a topic shift or a conclusion of a previous utterance and the so tokens that connected a new piece of information to an old utterance were significantly quieter than their following segment in paired t-tests. in contrast, the so tokens that introduced a conclusion of the current utterance appeared to be louder than its preceding segment, even though the difference was not significant. the increase in intensity after so, as suggested by rennie and colleagues, may signal a more important status of the segment following so in the utterance. the decrease in intensity before so, on the other hand, seems to be in line with the tendency of which intensity increases at the start of a new topic and decreases at the end (brown et al., 1980). analyzing the mandarin ránhòu, wang (2018) also reported a similar observation: that the tokens marking the closure of the current speaker’s turn were marked with gradually decreased loudness as well as lengthening. her observation, however, was qualitative and based on only six tokens in her data. in connection to the relationship between phonetic encodings and the structure of discourse, attention has been given to temporal features such as word duration. for instance, horne et al. (2001) observed a significant difference in mean duration between the discourse and sentential use of the swedish men, where discourse men was longer than sentential men. studies on the topic-shift function of the french alors (didirková et al., 2018) and of the mandarin ránhòu (wang, 2018) have similarly reported a longer word duration than that of their respective sentential functions. yang (2006) also offered insights into the durational correlates of other discourse functions of ránhòu. yang found that the ránhòu tokens signaling a continuation of the current topic were short in word duration. in contrast, the tokens signaling an emphasis on the distinctiveness of the following content had a much longer word duration. she claimed that a longer duration attracts the hearer’s attention to the current topic, while the lack of prominent duration reflects less of a need to call for the hearer’s attention, which typically occurs when the following utterance develops step-by-step within the same topic. rennie et al. (2016), on the other hand, presented a case where the english so was shorter in wu and tseng 95 duration mean when functioning to grab the hearer’s attention in the utterance-initial position. this is also evidenced by the lack of a perceptible pause between so and subsequent speech. in addition to the aforementioned phonetic evidence, it has also been suggested that highfrequency disyllabic connectives used in mandarin conversational speech are often produced in a phonetically extremely reduced form (liu et al., 2016). the duration and position of disyllabic connectives tend to correlate with the degree of word reduction. therefore, in our later analysis, we will include four types of phonetic correlates, including pitch, intensity, duration, and reduction degree. 3.3 context there has been some research into how contextual cues such as silent pauses and paralinguistic events may aid speakers in structuring discourse. a silent pause typically signals a major prosodic boundary. horne et al. (2001) observed that 34% of the topic-shift use of the swedish men was both preceded and followed by a pause, while none of the sentential uses were. similarly, wang (2018) found that during trail-off use, the mandarin ránhòu was marked with a pitch contour independent of the previous intonation unit. in addition, it was immediately followed by laughter. on silent pause duration, didirková et al. (2018) revealed that the silent pause preceding the french alors tended to be longer when the connective opened a new topic or introduced specification. moreover, they noted that discourse uses of alors was almost never followed by a silent pause. rhee (2020) reported that the filled-pause use of the korean mak was often realized after a short pause, signaling that floor holding was needed. he further stated that a pause may distinguish the filled-pause use from other discourse functions of mak, such as the speaker expressing a negative stance and intensifying an utterance, as these functions tended to be marked with no pauses before or after mak. swerts (1998) also discussed silent pauses preceding and following the dutch filled pause uh and um. it showed that almost all tokens of phrase-initial uh and um had a neighboring silent pause. 4. data and annotation of mandarin connectives in this section, we present the scheme with which we labeled our target connectives and the annotation results with illustrative examples. section 4.1 gives a brief overview of the corpus used in this study. sections 4.2 and 4.3 describe the annotation scheme with which we labeled all tokens of connectives in our corpus in terms of their sentential/discourse use and individual discourse functions. representative examples of the annotated discourse functions are illustrated in section 4.4. lastly, section 4.5 presents descriptive statistics of our data in terms of positional, phonetic, and contextual features. 4.1 target connectives in sinica mcdc8 sinica mcdc8 contains eight free conversations produced by seven male and nine female mandarin chinese speakers aged between 16 and 46. the speakers were randomly sampled form and function of connectives in chinese conversational speech 96 from the citizens of taipei city in 2001. each pair of speakers who were invited to participate in the recording project met each other for the first time. the corpus has approximately eight hours of speech recording with 90k transcribed words/122k syllables. acoustic properties that will be used for our later analysis were measured based on the signaled-aligned syllable boundary information (tseng, 2019)2 . in the present study, we performed an exploratory analysis to identify our target connectives in the corpus, including bùguǎn, ‘no matter’; jiǎrú, ‘if’; jíshǐ, ‘even if’; jiùshì, ‘is precisely’; háishì, ‘still’; huòshì, ‘or’; huòzhě, ‘or’; rúguǒ, ‘if’; suīrán, ‘even though; suǒyǐ, ‘so’; yaòshì, ‘if’; zhǐshì, ‘only’; zhǐyaò, ‘as long as’; ránhòu, ‘then’; and qíshí, ‘actually’. as part of the phonetic features, we adopted the labels of disyllabic reduction degree from liu and colleagues (liu et al., 2016) and selected only connectives that had such annotation in the corpus. eventually, we obtained a total of 1370 connective tokens, as shown in table 1. connective tokens connective tokens bùguǎn 11 suīrán 9 jiǎrú 7 suǒyǐ 215 jíshǐ 2 yaòshì 3 jiùshi 387 zhǐshì 32 háishì 67 zhǐyaò 23 huòshì 20 ránhòu 298 huòzhě 11 qíshí 216 rúguǒ 69 table 1: frequency of connectives in sinica mcdc8 to present the semantic content and the prosodic organization of spontaneous conversation, prévot et al. (2015) proposed two types of production units, discourse units (dus) and prosodic units (pus). a du consists of a main predicate and the related complements and adjuncts. a pu is a stretch of speech content separated by perceptible pitch reset, changes in speech rate, and pauses. we adopted the definition of dus and pus proposed by prévot et al. (2015) and used the word position in which a connective occurs relative to the respective du/pu as our positional features of connectives in our later analysis. 4.2 sentential/discourse labeling we posited that a connective delivers a sentential use if the meaning of the du is inevitably changed or becomes incomplete when the connective is removed from the du in which it occurs. this is illustrated by (1), where the two speakers talked about what they do for work. after asking speaker a where her office is and failing to get a satisfying answer, speaker b asked speaker a whether she could describe the direction to her office from nangang district in taipei, prefaced by rúguǒ in bold. canonically, rúguǒ indicates hypotheticality, which suggests that the proposition of the du (nà rúguǒ cóng nángǎng guòqù ‘if going there from 2 sinica mcdc8 is publicly available via the association for computational linguistics and chinese language processing. wu and tseng 97 nangang’) is a purely hypothetical statement. when rúguǒ is removed from the du, the semantic meaning of the du is also affected. this shows that rúguǒ is used sententially in this case. (1) b: na 如果 /3 從 南港 過去 $ 要 怎麼 去 a. na rúguǒ cóng nángǎng guòqù yào zěnme qù a ‘if i go there from nangang, how do i get there?’ if a connective adds a designated discourse interpretation to the du in which it occurs in relation to the local context, it is considered discourse use. an instance of such connective is shown in (2), which presents a conversation in which speaker a told speaker b about her trip to a hot spring resort in japan. speaker a described the indoor baths as separated from the outdoor baths and stated that the experience was quite interesting. she was going to share her opinion about japanese people using the expression wǒ juede, ‘i feel/think’. she then abandoned the thought and shifted to talking about liking japan, which is less directly related to what is being discussed. the transition to a new topic (i.e., her liking japan) on line 25 was introduced by qíshí, which is discourse use. (2) 1 a: nhnn, / nhnn 2 他們 原則上 就是說 他 一, break / tāmén yuánzéshàng jiùshìshuō tā yī break 3 有 隔, / yǒu gé 4 就是, / jiùshì 5 [把 它 隔開 la. / $ [bǎ tā gékāi la 6 b: [還是 有 隔開]. / $ [háishì yǒu gékāi] 7 a: 就是] 隔 室內 跟 室外 而已 la. $ inhale/ $ jiùshì] gé shìnèi gēn shìwài éryǐ la inhale 8 b: on [hon]. / $ on [hon] 9 a: [就是] 讓 ne ge, / jiùshì ràng ne ge 10 可能 是 讓 ne ge, inhale / kěnéng shì ràng ne ge inhale 11 ne ge 氣 o, / ne ge qì o 12 溫度 o, / wēndù o 13 不要, pause / búyào pause 14 那樣子 交流 這樣子. / $ nàyàngzǐ jiāoliú zhèyàngzǐ 3 the symbol / indicates a prosodic unit boundary, and the symbol $ indicates a discourse unit boundary in sinica mcdc8. form and function of connectives in chinese conversational speech 98 15 b: [on ho hon]. / $ on ho hon 16 a: [他們 是 有], / tāmén shì yǒu 17 有 個, inhale / yǒu ge inhale 18 有 一 個 走道 這樣子, / $ yǒu yī ge zǒudào zhèyàngzǐ 19 然後, / ránhòu 20 整個 外面 就是 室外 的. $ inhale / $ zhěngge wàimiàn jiùshì shìwài de inhale 21 b: hon. / $ hon 22 a: 那 會 覺得 蠻 有意思 的, / $ nà hùi juéde mán yǒuyìsi de 23 我 覺得 日本人, / wǒ juéde rìběnrén 24 其實, $ inhale / $ qíshí inhale 25 其實 我 也 是 蠻 喜歡 日本 的. $ pause / $ qíshí wǒ yě shì mán xǐhuān rìběn de pause a: b: a: b: a: b: a: b: a: ‘in principle, they, that is… they [separated the baths.]’ ‘[they separated the baths.]’ ‘[that is], they only separated the baths into an indoor one and an outdoor one.’ ‘[right.]’ ‘[that is], they probably did that to prevent the heat exchange.’ ‘[right.]’ ‘[they had] a hallway, and the outdoor baths were on the outside.’ ‘right.’ ‘i think it is quite interesting. i think japanese people actually… i also quite like japan.’ horne et al. (2001) mentioned that both a sentential use and a discourse use interpretation of connectives can seem possible. we also noticed that in some cases, the distinction of sentential or discourse use can be ambiguous. for instance, in (3), speaker a told speaker b about the harmful effect of formaldehyde and her effort to educate people about it. she mentioned that all she could do is to share the information with people, and it is up to people to do something with the information. she then said jiùshì zhèyàng, literally ‘that is it’ in english, on line 11. here, jiùshì may sententially indicate the preciseness of the equation between zhèyàng ‘like this’ and the previous content. however, we identified a discourse meaning in which the speaker introduced a conclusion for her previous topic and signaled to the hearer that there is no more to add to the current topic. (3) 1 a: 所以 我 覺得, / suǒyǐ wǒ juéde 2 現在 我 覺得, inhale / xiànzài wǒ juéde inhale wu and tseng 99 3 我 的 責任 就是 把 我 知道 的 告訴 wǒ de zérèn jiùshì bǎ wǒ zhīdào de gàosù 4 他們. / $ tāmén 5 b: [mhmhmhm]. / $ mhmhmhm 6 a: [na 至於] 他 接 不 接受 $ na zhìyú tā jiē bù jiēshòu 7 看 他 自己 了 la, $ inhale / $ kàn tā zìjǐ le la inhale 8 b: hon. / $ hon 9 a: 我 盡力而為. $ laugh / $ wǒ jìnlì’érwéi laugh 10 b: laugh. / $ laugh 11 a: 就是 這樣 a, / $ jiùshì zhèyàng a 12 所以 其實 蠻 好玩 的. / $ suǒyǐ qíshí mán haǒwán de a: b: a: b: a: b: a: ‘so now i feel that it is my responsibility to tell them what i know.’ ‘[right.]’ ‘[as for] whether they accept it or not, it is up to them.’ ‘right.’ ‘i do my best.’ ‘@’ ‘that is it. it is actually quite fun.’ 4.3 annotation scheme of discourse functions describing the discourse functions of connectives has often proven a challenging task since the interpretation of the functional properties can be quite elusive and often context dependent. the exploration and initial annotation of discourse functions were carried out by the authors. auditory information was used to aid the classification wherever necessary. according to previous work, we were able to identify sentential use and eight types of discourse use for our connectives. it was relatively straightforward to adopt the definitions of resuming and concluding topics and disfluencies. we could also identify a more coarse-grained type of function elaboration for the majority of the cases that provide elaboration or clarification on a current topic. however, for functions such as emphasis, downtoning, securing the addressee’s attention, and contrast, it was, in fact, truly challenging to operationalize and identify them. therefore, we collapsed the above four discourse functions, along with elaboration, into elaboration, as they all provide more information to a proposition. the exploration of the data also identified cases where connectives signaled repairs in discourse (tseng, 2006). as such, we added repair to our discourse functions and collapsed it into disfluency. this led to an annotation scheme for four function categories: conclusion, resumption, elaboration, and disfluency. for validation, two trained labelers were recruited for verifying sentential use and the four function categories. each labeler annotated half of the dataset independently. as a result, the agreement between the authors’ and the labeler’s annotations achieved a cohen’s form and function of connectives in chinese conversational speech 100 kappa of 0.92. although judgment of sentential/discourse use is likely to be considered highly subjective, the agreement over the annotation of connective functions is surprisingly satisfactory. table 2 summarizes the statistics and their respective definitions. controversial cases were discussed to finalize the annotations. function category function definition tokens percentage sentential sentential functions as a conjunction or an adverb 639 46.64% resumption resumption resumes a new or previous topic 32 2.33% conclusion conclusion introduces a conclusion or summarization 103 7.52% elaboration elaboration elaborates on a current or previous topic 469 34.23% emphasis emphasizes the significance of the proposition 14 1.02% downtoning downplays the significance of the proposition 7 0.51% securing the addressee’s attention secures the addressee’s attention 3 0.22% contrast expresses the speaker's disagreement with the previous speaker 12 0.88% disfluency filled pause signals filled pauses 81 5.91% repair signals repairs 10 0.73% total 1370 99.99% table 2: sentential/discourse functions of connectives in sinica mcdc8 4.4 annotation examples as shown previously in (2) and (3), connectives can be used to perform topic shifts in conversation: qíshí in (2) was found to resume a topic, and jiùshì in (3) may conclude a previous topic. aside from topic shift, some connectives may introduce an elaboration on the current topic, illustrated by ránhòu on line 6 in (4). (4) 1 b: 那 他們, break / nà tāmén break 2 講話 就是 有 口音 嗎. / $ jiǎnghuà jiùshì yǒu kǒyīn ma 3 a: nhn, $ inhale / $ nhn inhale 4 還好 ba $ 可是, break / háihaǒ ba kěshì break 5 因為 他 教 兩 個 班 $ yīnwèi tā jiāo liǎng ge bān 6 然後 教 我們 班 跟 另外 一 班, $ inhale / $ ránhòu jiāo wǒmén bān gēn lìngwài yī bān inhale wu and tseng 101 7 然後 每 次, / ránhòu měi cì 8 有時候 就是, inhale / yǒushíhòu jiùshì inhale 9 上課 a, $ break / $ shàngkè a break 10 要 趕 去 別的 教室 $ yào gǎn qù biéde jiàoshì 11 然後 就 看到, inhale / ránhòu jiù kàndào inhale 12 他 就 在 前面 上課 $ tā jiù zài qiánmiàn shàngkè 13 然後 後面 一半 人 都 睡 死 了, / ránhòu hòumiàn yībàn rén dōu shùi sǐ le 14 [這樣子, $ laugh $ inhale / $ zhèyàngzi laugh inhale 15 對 a]. / $ dùi a 16 b: [laugh 對 $ 我們 也 這樣子]. $ break / $ laugh dùi wǒmén yě zhèyàngzi break b: a: b: ‘did he speak with an accent?’ ‘the accent was not bad. but, since he taught two classes, which were my class and another class, every time… sometimes when i needed to rush to another classroom for class, i would see that half of the students in the back fell asleep while he was teaching in the front. [that is it. right.]’ ‘[right. it was the same way for us too.]’ intriguingly, we have observed that some connectives can emphasize or downtone the importance of a proposition in discourse, as illustrated in (5) and (6), respectively. in (5), the two speakers were talking about modified cars. speaker b pointed out a problem on line 7 in that many people drive fast in their modified cars. he then argued that people should modify their cars only for safety. prefacing his argument with the fact-introducing marker qíshí (hsieh and huang, 2005; wang et al., 2010), speaker b was able to suggest the proposition in his argument was factual and consequently emphasize its importance. (5) 1 b: na 我, / na wǒ 2 我 請問 你 你 開到 一百 $ wǒ qǐngwèn nǐ nǐ kāidào yībǎi 3 再 轉 一 個 彎 $ zài zhuǎn yī ge wān 4 和 開 五十 $ 轉 一 個 彎, / $ hé kāi wǔshí zhuǎn yī ge wān 5 是 一定 有 差 的 ma. $ inhale / $ shì yīdìng yǒu chā de ma inhale 6 a: mhm. / $ mhm 7 b: na 今天 很 多 人 的 改裝 車, inhale / na jīntiān hěn duō rén de gǎizhuāng chē inhale 8 他 可能 會, break / tā kěnéng hùi break form and function of connectives in chinese conversational speech 102 9 仗著, inhale / zhàngzhe inhale 10 他們 自己 車子 很 好, $ inhale / $ tāmen zìjǐ chēzǐ hěn hǎo inhale 11 開 很 快, $ inhale / $ kāi hěn kuài inhale 12 na 這些 是 他們 心態 上 的 問題, $ na zhèxiē shì tāmen xīntài shàng de wèntí 13 inhale / $ inhale 14 可是, / kěshì 15 其實, break / qíshí break 16 改 車子 並 不 是, break / gǎi chēzi bìng bù shì break 17 針對 於 飆車 用 $ 是 對 安全性. $ zhēnduì yú biāochē yòng shì duì ānquánxìng 18 swallow / $ swallow 19 a: mhmhmhm. / $ mhmhmhm b: a: b: a: ‘let me ask you… there is surely a difference between making a turn at the speed of 100 km/hr and making a turn at the speed of 50 km/hr.’ ‘right.’ ‘nowadays many modified cars… those people may drive fast because they think their cars are good. that’s the problem with their attitude. in fact, modifying cars is not about racing. it is about safety.’ ‘right.’ in contrast, connectives such as zhǐshì can downtone the importance of the proposition in discourse. in (6), speaker a explained that she was stuck in traffic right before arriving at academia sinica, where the recording of the conversation took place. she, however, went on to clarify on line 18 that the traffic on the way to academia sinica was not that bad and that she was late only because she did not estimate the travel time correctly (shíjiān shàng yùgū kěnéng méiyǒu xiǎng yīxià, ‘i didn’t think about the estimated time to get here.’). her reason for being late was prefaced by zhǐshì, which limited the validity of the proposition in discourse (wang et al., 2013). the connective allowed speaker a to downtone the importance of her reason and strengthen her point that the traffic to academia sinica was in fact not bad. (6) 1 a: na, / na 2 在 中央研究院 的 ne ge, / zài zhōngyāngyánjiùyuàn de ne ge 3 研究院一路 那 邊=, inhale / yánjiùyuànyīlù nà biān inhale 4 那 鐵道 有點 塞車 $ 所以 我 [一路] nà tiědào yǒudiǎn sāichē suǒyǐ wǒ yīlù 5 b: [hen hen]. / $ hen hen 6 a: 上 過來 $ 剛好, inhale / shàng guòlái gānghǎo inhale wu and tseng 103 7 兩 個, / liǎng ge 8 一 個 地方 在 施工 $ [然後], / yī ge dìfāng zài shīgōng ránhòu 9 b: [nhn]. / $ nhn 10 a: 因為 我 對 南港 這 邊 的 路況 yīnwèi wǒ duì nángǎng zhè biān de lùkuàng 11 也 不 是 很 熟[悉 $ 所以], inhale / yě bù shì hěn shúxī suǒyǐ inhale 12 b: [mhmhm]. / $ mhmhm 13 a: 我 在, / wǒ zài 14 中央研 , / zhōngyāng 15 研究院一路 那 邊 就 有點 堵塞=. $ yánjiùyuànyīlù nà biān jiù yǒudiǎn dǔsè 16 break / $ break 17 b: oh. / $ oh 18 a: na 一路 順 著 這樣子 過來, / $ na yīlù shùn zhe zhèyàngzi guòlái 19 ei 還, / ei hái 20 還 算, / hái suàn 21 還好 la, / $ háihǎo la 22 但是 只是說, inhale / dànshì zhǐshìshuō inhale 23 時間 上 預估, pause / shíjiān shàng yùgū pause 24 可能 沒有 想 一下. $ kěnéng méiyǒu xiǎng yīxià a: b: a: b: a: b: a: b: a: ‘there was a little bit traffic around the railway on academia road section 1. [on the whole way]’ ‘[right.]’ ‘here, there happened to be two… one place under construction on the way here. [then,]’ ‘[right.]’ ‘i was also not very familiar with the roads in nangang, [so]’ ‘[right.]’ ‘i was stuck in traffic a little on academia road section 1.’ ‘oh.’ ‘the traffic on the way here was still not bad. it is just that i didn’t think about the estimated time to get here.’ attention-securing is another interlocutor interaction enabled by our connectives. the connective rúguǒ, for example, can be used to secure the attention of the hearer. in (7), speaker a argued that western democracy is not the best political system for all countries and that it will be better if a country has the freedom to figure out which system works best for it. he used china as the example for his argument, saying that although china went through a dark time form and function of connectives in chinese conversational speech 104 when communism was first introduced to the nation, it is enjoying great economic development now. he then recalled a report about shanghai that he saw on tv on line 10, prefaced by a rúguǒ-led clause addressing speaker b (rúguǒ nǐ yǒu kàn dìsìtái, ‘if you watch the cable tv’). instead of probing for a response from speaker b, evidenced by a lack of wait time for speaker b to say something, speaker a wanted to bring speaker b’s attention to his next utterance on the tv report. (7) 1 a: na, / na pat 2 na, / na pat 3 他們 帶 過來 的 時候 $ 你 說, / tāmen dài guòlái de shíhòu nǐ shuō they bring over pat time you say 4 對 $ 中共, pause / $ duì $ zhōnggòng pause to chinese communist party pat 5 中國 大陸 的確 是 走 過 一 段 很, / zhōngguó dàlù díquè shì zǒu guò yī duàn hěn china mainland indeed pat walk pat one stage very 6 蠻 黑暗 的. $ inhale / $ mán hēi'àn de inhale quite dark pat pat 7 b: 對 [a]. / $ dùi a right pat 8 a: [但 你 看] 他們 現在 , / dàn nǐ kàn tāmen xiànzài yī fā but you see they now once develop 9 一 發展 起來, $ inhale / $ yī fāzhǎn qǐlái inhale once develop pat pat 10 o 那 真的 不得了 $ 如果, / o nà zhēnde bùdéliǎo rúguǒ pat that really incredible if 11 你 有 看 第四台 $ 我 常常 看 nǐ yǒu kàn dìsìtái wǒ chángcháng kàn you have see cable tv i often see 12 ne ge 他們 報導 上海, / $ ne ge tāmen bàodǎo shànghǎi pat they report shanghai 13 經--, $ inhale / $, jīng inhale economic pat 14 大 城市 來 講 真的, pause / dà chéngshì lái jiǎng zhēnde pause big city pat speak really pat 15 不會 輸 我們. $ pause / $ búhuì shū wǒmén pause will not lose us pat wu and tseng 105 16 b: 對. $ pause / $ dùi pause right pat 17 a: 一定 贏 過, / $ yīdìng yíng guò surely win pat a: b: a: b: a: ‘when the soviet union introduced communism to china, it is true that mainland china walked through a very dark time.’ ‘[right.]’ ‘[but you can see] that now they have developed (economically) and it is really incredible. if you watch the cable tv… i often watch them reporting about shanghai. if we are talking about (the development of) big cities (in china), it won’t lose to us.’ ‘right.’ ‘it definitely beats us.’ connectives can also convey certain interactions between interlocutors. for instance, in (8), the speaker used qíshí on line 5 to express her disalignment with the previous speaker. (8) 1 a: 其實 我 剛 要 不 講 qíshí wǒ gāng yào bù jiǎng 2 我 是 開 計程車 的 $ wǒ shì kāi jìchéngchē de 3 或許 妳 也, / huòxǔ nǐ yě 4 看 不 出來, $ pause / $ kàn bù chūlái pause 5 b: 其實 [我 覺得, / qíshí wǒ juédé 6 誰 也 看 不] shéi yě kàn bù 7 a: [看 不 大 出來], $ pause / $ kàn bù dà chūlái pause 8 b: 出來 是 什麼 [行業], / $ chūlái shì shénme hángyè 9 a: [en]. / $ en 10 b: 因為 行業 實在 太 多 了 $ yīnwèi hángyè shízài tài duō le 11 [也 很 難 猜], $ inhale / $ yě hěn nán cāi inhale 12 a: [有 差, $ break / $ yǒu chā break 13 有 差]. $ yǒu chā a: b: a: b: a: b: a: ‘actually, if i didn’t say i’m a taxi driver, you probably can’t tell.’ ‘actually, [i think that nobody can]’ ‘[you can’t really tell.]’ ‘tell what a person’s [profession is]’ ‘[right.]’ ‘because there are too many professions. [it is also difficult to guess.]’ ‘[(the profession) does make a difference.]’ form and function of connectives in chinese conversational speech 106 last, connectives can signal disfluencies such as filled pauses and repairs, as illustrated in (9) and (10), respectively. in (9), suǒyǐ functioned as a filled pause on line 11. it was followed by a short pause, which hints at the speaker’s hesitation or word-searching. (9) 1 a: 除此之外, $ inhale / $ chúcǐzhīwài inhale 2 因為 已經=, / yīnwèi yǐjīng 3 要 開始, exhale / yào kāishǐ exhale 4 就是 要 跟 老師 做 研究 了 這樣, $ jiùshì yào gēn lǎoshī zuò yánjiū le zhèyàng 5 所以, inhale / suǒyǐ inhale 6 在, / zài 7 最近 的 幾 天 或 最近 的 幾 個 月 之內 zuìjìn de jǐ tiān huò zuìjìn de jǐ ge yuè zhīnèi 8 都 會 下去, $ inhale / $ dū huì xiàqù inhale 9 就是 台北 台南 這樣 兩 邊 跑 jiùshì táiběi táinán zhèyàng liǎng biān pǎo 10 這樣, $ inhale / $ zhèyàng inhale 11 對, $ 所以=, inhale / duì suǒyǐ inhale 12 雖然說 離家 很 遠 la $ suīránshuō líjiā hěn yuǎn la 13 可是=, inhale / $ kěshì inhale 14 看 以後 會 不 會 成長, $ kàn yǐhòu huì bù huì chéngzhǎng a: ‘in addition, since i’m about to start doing research with my supervisor, i’ll go down to tainan in the next few days or in the next few months. that is, i’ll travel between taipei and tainan. right. so… although i’ll be far away from home, i’ll see if i will grow (through this experience).’ in (10), speaker b asked speaker a, on line 6, whether the humanities majors in her school have to take summer lessons. she started her question off with nǐmén, ‘you’, referring to the humanities majors, before introducing the alteration (qítā jiù búyòng, ‘other (people) just don’t need to (have lessons)’) with jiùshì. it is clear that jiùshì signaled to the hearer that what follows is a repair to what was just said. (10) 1 a: 我們 好像 是 學校 念, inhale / wǒmén hǎoxiàng shì xuéxiào niàn inhale 2 ne ge 理科 的 才 要 回 學校 ne ge lǐkē de cái yào huí xuéxiào 3 去 [上課]. $ inhale / $ qù shàngkè inhale wu and tseng 107 4 b: [o]. / $ o 5 a: 對 a $ 對. / $ duì a duì 6 b: na 妳們 就是, / na nǐmén jiùshì 7 就是 其他 就 不用, / jiùshì qítā jiù búyòng 8 回去 lo. huíqù lo 9 a: 對 a $ 文科 就 不用 回去. $ inhale / $ duì a wénkē jiù búyòng huíqù inhale 10 b: 好好 o. / $ hǎohǎo o a: b: a: b: a: b: ‘for us, it seems that only the students who major in sciences in our school need to go back to [have lessons.]’ ‘[oh.]’ ‘right.’ ‘so, you just… other people just don’t need to go back for lessons.’ ‘right. students who major in humanities don’t need to go back.’ ‘that is nice.’ 4.5 descriptive results of annotated connectives this section presents the descriptive statistics of our target connectives based on the three major groups of features. section 4.5.1 shows the positional features, which designate the position of a connective in relation to a du/pu. section 4.5.2 presents the phonetic features, including duration, f0, intensity, and reduction degree. section 4.5.3 describes the contextual features, including the duration of preceding and following paralinguistic events as well as the speech rate of du/pu in which a connective occurs. 4.5.1 positional features the position of the connective is operationalized as the initial, medial, and final position in relation to a du/pu and the case in which a connective itself forms an isolated du/pu, annotated as du_initial, du_medial, du_final, and du_isolated. similar to du, the positions in a pu are annotated as pu_initial, pu_medial, pu_final, and pu_isolated. as shown in table 3, connectives of sentential use do not seem to particularly occur in du-initial or -medial positions, but when connectives deliver discourse functions of conclusion, elaboration or resumption, a du-initial position is generally preferred. when used in relation to disfluency, connectives tend to take du-medial or du-initial positions. in contrast to du, positional features related to pu show that the prosodic manifestation of connectives is diverse across discourse functions. in terms of prosodic segmentation, connectives seem more likely to occur in the form of a standalone unit than in the meaning-oriented segmentation of discourse. we will later conduct a multinomial logistic regression model to examine whether there is any significant effect in the comparison of du and pu. form and function of connectives in chinese conversational speech 108 du_initial du_medial du_final du_isolated n sentential 290 (45.38%) 343 (53.68%) 6 (0.94%) 0 639 conclusion 96 (93.20%) 7 (6.80%) 0 0 103 disfluency 31 (34.07%) 51 (56.04%) 3 (3.30%) 6 (6.59%) 91 elaboration 397 (78.61%) 105 (20.79%) 1 (0.20%) 2 (0.40%) 505 resumption 26 (81.25%) 6 (18.75%) 0 0 32 pu_initial pu_medial pu_final pu_isolated n sentential 219 (34.27%) 305 (47.73%) 70 (10.95%) 45 (7.04%) 639 conclusion 68 (66.02%) 13 (12.62%) 3 (2.91%) 19 (18.45%) 103 disfluency 23 (25.27%) 5 (5.49%) 24 (26.37%) 39 (42.86%) 91 elaboration 284 (56.24%) 81 (16.04%) 46 (9.11%) 94 (18.61%) 505 resumption 17 (53.12%) 4 (12.50%) 1 (3.12%) 10 (31.25%) 32 table 3: distribution of discourse functions by du and pu positions of connectives 4.5.2 phonetic features we considered word duration, pitch, intensity, and reduction degree in our analysis, following previous studies’ suggestions of phonetic correlates for the discourse function of connectives (didirková et al., 2018; hirschberg and litman, 1987, 1993; horne et al., 2001; liu et al., 2016; rennie et al., 2016; wang, 2018; yang, 2006). for each connective token, rate, initial_f0, and intensitymean represent the word duration in the form of speech rate (seconds per syllable), the initial f0 value, calculated using the firstpitch function in praat (boersma and weenink, 2022), respectively, and the mean of the intensity values is calculated using praat’s meanintensity function. figure 1 presents the results. for speech rate, connectives tend to be longer for disfluency (mean: 0.19 sec/syll) or resumption (mean: 0.17 sec/syll) use than for sentential use (mean: 0.14 sec/syll). all discourse uses (meanconclusion: 198.07 hz, meandisfluency: 200.53 hz, meanelaboration: 185.83 hz, meanresumption: 216.29 hz) tend to have a higher pitch onset than sentential use (mean: 173.33 hz). for intensity, there is no clear trend contrasting sentential/discourse use, but disfluency use seems to have a weaker intensity (mean: 68.15 db) than the other discourse functions (meanconclusion: 71.19 db, meanelaboration: 70.76 db, meanresumption: 71.44 db). figure 1: boxplot representation of discourse functions by speech rate, initial f0 value, and intensity mean of connectives (the red dot indicates the mean) 0.2 0.4 0.6 s e nt en tia l c on cl us io n d is flu en cy e la bo ra tio n r es u m pt io n discoursefunction r a te ( s e c/ sy ll ) 100 200 300 400 500 s e n te n tia l c on cl us io n d is flu en cy e la b or at io n r es um p tio n discoursefunction f 0 _ in it ia l (h z) 50 60 70 80 s e nt en tia l c on cl us io n d is flu en cy e la bo ra tio n r es u m pt io n discoursefunction in te n s it y m e a n ( d b ) wu and tseng 109 for word reduction, four degrees are distinguished for each connective token: the canonical form (can), the marginal segment deletion form (msd), the nuclei merging form (num), and the syllable merger form (sym), representing an increase in phonetic reduction from can to sym (liu et al., 2016). high-frequency disyllables in sinica mcdc8 mostly take the most reduced form, sym. the results shown in table 4 are partly in accordance with this tendency. however, for sentential or disfluency use, the percentage of sym is smaller than the percentage of other discourse functions. sym can msd num n sentential 382 (59.78%) 190 (29.73%) 7 (1.10%) 60 (9.39%) 639 conclusion 91 (88.35%) 7 (6.80%) 0 5 (4.85%) 103 disfluency 50 (54.95%) 20 (21.98%) 1 (1.10%) 20 (21.98%) 91 elaboration 350 (69.31%) 88 (17.43%) 16 (3.17%) 51 (10.10%) 505 resumption 22 (68.75%) 7 (21.88%) 1 (3.12%) 2 (6.25%) 32 table 4: distribution of discourse functions by reduction degree of connectives 4.5.3 contextual features previous research has suggested functional differences in the contextual cues occurring around connectives, such as silent pause (didirková et al., 2018; horne et al., 2001; rhee, 2020) and laughter (wang, 2018). we considered the position and duration of all paralinguistic events, such as pauses, coughs, laughs, etc., occurring around each connective token in terms of previous_duration and next_duration. previous_duration is the duration of a preceding paralinguistic event, and if there is no immediately adjacent paralinguistic event, the absence is marked with a zero. the same definition applies for next_duration. figure 2 presents the results. connectives of discourse uses (meanconclusion: 0.32 sec, meandisfluency: 0.21 sec, meanelaboration: 0.26 sec, mresumption: 0.39 sec) seem to be more likely to be accompanied by a preceding paralinguistic event than sentential use (mean: 0.10 sec). however, paralinguistic events following the occurrence of a connective do not seem to be as influential as those that precede such occurrences, except for disfluency use (meandisfluency: 0.33 sec). figure 2: boxplot representation of discourse functions by duration of paralinguistic events in the previous and next position of connectives (the red dot indicates the mean) 0 1 2 3 4 s e nt en tia l c on cl us io n d is flu e nc y e la bo ra tio n r e su m pt io n discoursefunction p re vi o u s_ d u ra ti o n ( s ec ) 0 1 2 s e nt en tia l c on cl us io n d is flu e nc y e la bo ra tio n r e su m pt io n discoursefunction n ex t_ d u ra ti o n ( se c) form and function of connectives in chinese conversational speech 110 another contextual property of connectives we considered is the speech rate of the entire du/pu in which the connective occurs. we posited that the overall speech rate of du/pu may correlate with the discourse function by reflecting predominant rhythmic patterns. du_rate is calculated by du/pu duration divided by the number of syllables in du/pu after removing paralinguistic events and fillers. figure 3 shows the results of du_rate and pu_rate. connective-occurring dus seem to be articulated in a slower tempo when used for disfluency (mean: 0.23 sec/syll) or resumption (mean: 0.21 sec/syll) purpose than when used sententially (mean: 0.19 sec/syll). on the other hand, connective-occurring pus seem to be articulated at a slower tempo for all discourse uses (meanconclusion: 0.22 sec/syll, meandisfluency: 0.35 sec/syll, meanelaboration: 0.23 sec/syll, meanresumption: 0.27 sec/syll) than for sentential use (mean: 0.21 sec/syll). figure 3: boxplot representation of discourse functions by speech rate of the du and pu where connectives occur (the red dot indicates the mean) 5. analysis of form and function of mandarin connectives we have identified a number of tendencies for the three groups of features in connective use. seemingly, all these features play certain roles in the production of connectives. our next study fits a multinomial logistic regression model to test whether these features statistically show an inclination for sentential/discourse uses of connectives. 5.1 multinomial logistic regression multinomial logistic regression (mlr) makes inferences about category memberships. it describes the probability of a comparison category being chosen over a reference category in a dependent variable as an outcome based on multiple independent variables. in this study, we built regression models to determine which features are more negatively or positively critical than others in predicting discourse functions. the independence of irrelevant alternatives (iia) is crucial in mlr modeling since the probabilities for any pair of categories should be determined without reference to the other categories that might be available. the iia can be tested by the hausman specification test (hausman and mcfadden, 1984). we calculated hiia using the hmftest function from the mlogit package (croissant, 2020) and obtained an 0.1 0.2 0.3 0.4 0.5 0.6 s e nt en tia l c on cl us io n d is flu e nc y e la bo ra tio n r e su m pt io n discoursefunction d u _ ra te ( s ec /s yl l) 0.4 0.8 1.2 1.6 s e nt en tia l c on cl us io n d is flu e nc y e la bo ra tio n r e su m pt io n discoursefunction p u _r a te ( s e c /s y ll ) wu and tseng 111 insignificant hiia, p > .05, suggesting that we can apply mlr modeling to our data. to build the mlr models, we used the multinom function from the nnet package for r (venables and ripley, 2002). for the predictor variables, we considered all the features described in section 4.3. for the response variable, we considered our level 1 functions, consisting of five categories: sentential, conclusion, disfluency, elaboration, and resumption. we set sentential as the reference category in the response variable. for each outcome pair of the dependent variable, the multinom function returns four values: the regression coefficients, the standard errors, the residual deviance of a model and the akaike information criterion (aic). the coefficients associated with a given predictor variable are the log-odds for a given probability of having a comparison category as the outcome per unit change in the predictor variable. the residual deviance of a model and the aic both entail goodness-of-fit. roughly speaking, residual deviance is an indicator of the amount of data not accounted for by a model. aic is sensitive to the number of parameters and increases with the number of variables used in a model or with their levels. for both measures, lower scores are better. we calculated four additional values for each outcome pair of the dependent variable using various functions in r: the z test scores, the p values, the odds ratios, and the confidence intervals (cis). we calculated the p values using z tests. the z test scores for a given predictor variable were obtained by dividing the predictor’s coefficients by its standard errors, which were then transformed into p values. an odds ratio > 1 indicates that the risk of the outcome falling into the comparison category relative to the risk of the outcome falling into the reference category increases as the variable increases. the ci for a given odds ratio informs us of the lower and upper limit of the interval for the odds ratio for the outcome relative to the reference category, given the other predictors are in the model. this is evaluated with a 95% confidence level. for the odds ratios, we used the exp function from the base package to obtain the exponentiation coefficients. for ci, we used the tidy function from the broom package (robinson et al., 2022). in addition to goodness-of-fit, we also performed some prediction with the models to see how well our considered features can predict discourse functions. we adopted a 70-30 split (70 for the training dataset and 30 for the test dataset) for the data and used the test dataset for prediction. we indicated the model performance with accuracy, which is calculated by the sum of true positives and true negatives over the number of all tokens. 5.2 overview of the models since the gradient variables were each measured by different units, we first performed z score standardization using the scale function in r (becker et al., 1988). since our goal is to compare sentential use tokens, which typically occur in the initial position of a du, we defined du_initial and pu_initial as the reference levels for positions in du/pu. as high-frequency disyllabic words are often contracted or merged (tseng, 2005), we defined sym as the form and function of connectives in chinese conversational speech 112 reference level for reduction degree. we did this by using the relevel function for r (r core team, 2021). to explore how each group of features fits our data, we built separate models, models [1] to [4], and a complex model of all features, model [5]. we looked for the lowest residual deviance and aic scores returned by the multinom function for each model to find the bestfitting one. we then examined the coefficients associated with the predictor variables in the best-fitting model for different outcome pairs of the dependent variable. table 5 summarizes how each model fits our data. model [5] had the lowest residual deviance and aic score and achieved an accuracy of 55.47% in predicting discourse functions in the test dataset. the degree to which other models fit the data, however, did not seem to be proportionally reflected in the prediction accuracy. for example, although model [5] fits the data much better than model [1b], as indicated by its smaller residual deviance and aic scores, their accuracies are the same; while the model with the acoustic features (model [2]) has a better fit than the model with the reduction degree features (model [3]), model [3] predicts outcomes slightly better than model [2]. since model [5] seems to be the best-fitting one, we will report the outcomes of model [5] below. model res. deviance aic accuracy (1a) positionaldu 2996.85 3028.85 0.5790 (1b) positionalpu 2958.28 2990.28 0.5547 (1c) positionaldu+pu 2819.31 2875.31 0.5790 (2) acoustic 3147.61 3179.61 0.4379 (3) reduction degree 3176.55 3208.55 0.4744 (4a) contextualdu 3066.90 3098.90 0.5206 (4b) contextualpu 3075.65 3107.65 0.5304 (4c) contextualdu+pu 3058.37 3098.37 0.5304 (5) positionaldu+pu + reduction degree + acoustic+contextualdu+pu 2688.49 2824.49 0.5547 table 5: goodness-of-fit of the models 5.3 variable performance the statistical details for each feature of model [5] are summarized in table 6 in terms of the regression coefficients, the standard errors, the z test scores and p values, the odds ratios, and the confidence intervals. wu and tseng 113 variable b std. error4 z-test odds ratio 95% confidence intervals for odds ratio model [5] conclusion intercept -1.06 0.20 -5.13*** 0.34 [0.23, 0.51] pu_medial -0.77 0.36 -2.14* 0.46 [0.22, 0.93] pu_final -0.11 0.70 -0.16 0.88 [0.22, 3.52] pu_isolated 0.92 0.43 2.12* 2.51 [1.07, 5.88] du_medial -1.98 0.42 -4.65*** 0.13 [0.05, 0.31] du_final -13.48 1.01x10-6 -1.33x107*** 1.39x10-6 [1.386372x10-6, 1.386377x10-6] du_isolated -3.78 4.65x10-9 -8.15x108*** 0.02 [0.02, 0.02] can -1.07 0.44 -2.40* 0.34 [0.14, 0.81] msd -16.02 9.57x10-8 -1.67x108*** 1.10x10-7 [1.095480x10-7, 1.095481x10-7] num -0.88 0.50 -1.73 0.41 [0.15, 1.11] rate -0.13 0.20 -0.67 0.87 [0.58, 1.30] f0_initial 0.34 0.10 3.16** 1.40 [1.13, 1.73] intensitymean 0.11 0.12 0.89 1.11 [0.87, 1.42] previous_duration 0.32 0.12 2.68** 1.38 [1.09, 1.75] next_duration -0.12 0.22 -0.56 0.88 [0.57, 1.36] pu_rate 0.30 0.22 1.31 1.35 [0.86, 2.11] du_rate -0.39 0.18 -2.19* 0.67 [0.47, 0.95] disfluency intercept -2.91 0.30 -9.58*** 0.05 [0.02, 0.09] pu_medial -1.66 0.53 -3.10** 0.18 [0.06, 0.54] pu_final 1.66 0.48 3.43*** 5.30 [2.04, 13.73] pu_isolated 1.77 0.40 4.40*** 5.89 [2.68, 12.98] du_medial 1.21 0.30 3.98*** 3.38 [1.85, 6.15] du_final 0.60 0.89 0.67 1.82 [0.31, 10.52] du_isolated 15.50 0.44 34.96*** 5.43x106 [2.28x106, 1.30x107] can -1.27 0.36 -3.48*** 0.28 [0.13, 0.57] msd -2.41 1.52 -1.57 0.08 [0.004, 1.79] num 0.08 0.37 0.21 1.08 [0.51, 2.27] rate -0.13 0.15 -0.83 0.87 [0.64, 1.19] f0_initial 0.57 0.12 4.60*** 1.78 [1.39, 2.28] intensitymean -0.27 0.12 -2.15* 0.75 [0.59, 0.97] previous_duration 0.22 0.15 1.45 1.25 [0.92, 1.69] next_duration -0.02 0.14 -0.19 0.97 [0.73, 1.28] pu_rate 0.45 0.18 2.47* 1.56 [1.09, 2.24] du_rate 0.14 0.14 1.02 1.15 [0.87, 1.54] elaboration intercept 0.32 0.12 2.67** 1.38 [1.09, 1.74] pu_medial -0.86 0.18 -4.64*** 0.42 [0.29, 0.60] pu_final 0.25 0.30 0.86 1.29 [0.71, 2.34] pu_isolated 0.79 0.26 3.01** 2.21 [1.31, 3.70] du_medial -0.95 0.15 -6.08*** 0.38 [0.28, 0.52] du_final -1.93 1.11 -1.74 0.14 [0.01, 1.27] du_isolated 12.40 0.44 27.97*** 2.44x105 [1.02x105, 5.81x105] 4 the implausibly large standard errors, z-test scores, odds ratios, and confidence intervals for the independent variable du_final, du_isolated, and msd may be caused by the presence of empty and small cells in table 3 and table 4. it is a generally assumed that for independent variables in multinomial logistic regression the cell frequencies should be greater than 1. form and function of connectives in chinese conversational speech 114 can -0.26 0.18 -1.40 0.76 [0.53, 1.10] msd 0.42 0.49 0.86 1.53 [0.57, 4.07] num -0.07 0.22 -0.32 0.92 [0.59, 1.45] rate -0.01 0.10 -0.11 0.98 [0.80, 1.21] f0_initial 0.17 0.07 2.49* 1.19 [1.03, 1.37] intensitymean 0.12 0.06 1.78 1.12 [0.98, 1.29] previous_duration 0.26 0.08 2.95** 1.29 [1.09, 1.54] next_duration -0.15 0.12 -1.28 0.85 [0.67, 1.08] pu_rate 0.02 0.13 0.21 1.02 [0.79, 1.33] du_rate -0.05 0.08 -0.59 0.94 [0.79, 1.12] resumption intercept -2.55 0.33 -7.64*** 0.07 [0.04, 0.14] pu_medial -0.59 0.63 -0.93 0.55 [0.16, 1.91] pu_final -1.01 1.20 -0.84 0.36 [0.03, 3.84] pu_isolated 0.83 0.58 1.43 2.31 [0.73, 7.27] du_medial -1.03 0.52 -1.98* 0.35 [0.12, 0.98] du_final -11.69 1.56x10-6 -7.51x106*** 8.30x10-6 [8.299117x10-6, 8.299168x10-6] du_isolated -2.19 1.45x10-8 -1.52x108*** 0.11 [0.11, 0.11] can -0.41 0.52 -0.79 0.66 [0.23, 1.83] msd -0.41 1.21 -0.34 0.65 [0.06, 7.19] num -0.88 0.78 -1.12 0.41 [0.08, 1.92] rate 0.43 0.25 1.72 1.54 [0.94, 2.54] f0_initial 0.57 0.16 3.52*** 1.78 [1.29, 2.45] intensitymean 0.28 0.22 1.26 1.32 [0.85, 2.04] previous_duration 0.42 0.14 2.89** 1.52 [1.14, 2.03] next_duration 0.03 0.30 0.11 1.03 [0.56, 1.89] pu_rate -0.14 0.33 -0.41 0.86 [0.44, 1.68] du_rate 0.13 0.21 0.64 1.14 [0.75, 1.75] table 6: statistical details of model [5] (with sentential use as the reference category) 5.3.1 positional variables our first analysis investigated the effects of the connective position in relation to a du/pu on the prediction of discourse functions. the coefficients of the outcome for conclusion, elaboration, or resumption use reached significance in the z test, decreasing relative to sentential use for du-medial connectives. the coefficient reached significance only in conclusion and resumption uses for du-final connectives and showed a decreasing likelihood trend. for connectives forming standalone dus, the coefficients of the outcome being conclusion or resumption use are also expected to decrease relative to sentential use, while those of the outcome being disfluency or elaboration use expected to increase relative to that of sentential use. particularly, different from conclusion, elaboration, or resumption use, the coefficient of the outcome for disfluency use is expected to increase relative to that for sentential use for du-medial connectives as well as for connectives forming standalone dus. for pu-based predictors, the coefficients of the outcome for conclusion, disfluency, or elaboration use are expected to decrease relative to that for sentential use for pu-medial connectives. for connectives that form standalone pus, the coefficients of the outcome for conclusion, disfluency, or elaboration use are expected to increase relative to that for sentential use. interestingly, we found that disfluency is the only category for which the wu and tseng 115 coefficient of the outcome for such use is expected to increase relative to that for sentential use of pu-final connectives. in summary, our results show that positions in du/pu are sensitive to the prediction of discourse functions. considering du-related results, we can clearly see that connectives used to indicate a conclusion or resumption use seem to prefer initial positions in du. this is an important finding because connectives used sententially do not always take an utterance-initial position in conversation, even if their conventional position is utterance-initial. on the other hand, pu-related features do not seem to be involved in the prediction of the sentential/discourse use of connectives for conclusion and resumption purposes as obviously as du-related features. nonetheless, it seems that pu-related features are more effective in predicting a disfluency or sentential use in our complex model, as connectives used to indicate a disfluency use are shown to prefer final positions in pu or standalone pus. 5.3.2 phonetic variables we then investigated the effects of the phonetic features of connectives on the prediction of discourse functions in our model. the coefficients of the outcome for conclusion or disfluency use are expected to decrease relative to those of sentential use when connectives are produced in canonical form, and the coefficients reached significance. that is, when a disyllabic connective is pronounced with more phonetic details, it is more likely to be used sententially. however, this is not the case for elaboration or resumption use. the coefficients of the outcomes for all four discourse functions are expected to increase relative to that for sentential use per point increase in word-initial f0 values, and the coefficients all reached significance. this significant result is clear evidence that the sentential and discourse differences in the use of connectives is reflected in the phonetic representation of the connectives. for word duration, the coefficient of the outcome for resumption use is expected to increase relative to that for sentential use per point, and those for conclusion, disfluency, or elaboration use showed a tendency to decrease. neither of the coefficients, however, was significant. last, the coefficients of the outcome for conclusion, elaboration, or resumption use showed a tendency to increase relative to that for sentential use per point of increase in intensity, but without significance, either. in contrast, the coefficient of the outcome for disfluency use is expected to significantly decrease relative to that for sentential use per point of increase in intensity. 5.3.3 contextual variables our final analysis investigated the effects of the contextual cues around the connective on the prediction of discourse functions. our model showed that the coefficient of the outcome for any of the four discourse functions would be expected to increase relative to that for sentential use per point of increase for the duration of paralinguistic events preceding connectives. all coefficients, except that for disfluency use, were significant. the coefficients of the outcome for conclusion, disfluency, or elaboration use showed a tendency to decrease relative to that form and function of connectives in chinese conversational speech 116 for sentential use per point of increase for the duration of paralinguistic events following connectives, but without significance. the coefficient of the outcome for resumption use also showed a tendency to increase relative to that for sentential use per point of increase for the duration of paralinguistic events following connectives, without significance. for the overall speech rate of du, the coefficient of the outcome for conclusion use is expected to decrease relative to that for sentential use per point of increase. for pu, the coefficient of the outcome for disfluency use is expected to increase relative to that for sentential use per point of increase. both reached significance. 5.4 lexico-semantic property matters making use of all connective tokens in the corpus, we have observed a number of tendencies between the form and the discourse function of mandarin connectives. however, we would also like to mention that individual differences nevertheless exist among connectives, as the lexicosemantic meaning of connectives inevitably affects the form and the function of connectives to various degrees. we illustrate this point by taking ránhòu as an example, as its phonetic patterns clearly differ from the general patterns we identified for all connective tokens. this is shown in table 7 and figures 4 and 5. although ránhòu is similar to other mandarin connectives in the sense that these connectives all canonically indicate a semantic relationship between two adjacent clauses, discourse uses of ránhòu occur predominantly in the du-initial position, while the distribution of discourse uses of all connective tokens is mostly du-initial or dumedial. furthermore, discourse uses of ránhòu tend to occur more often in standalone pus than discourse uses of all connective tokens, which have a more even distribution across pu positions. these tendencies seem to correspond to the observations in wang (2018), where ránhòu exclusively occurs in the turn-initial position when signaling topic shifts. in tokens of ránhòu that close the current turn and invite the hearer’s response, wang also found lengthening of the connective and occurrence in only standalone intonation units. while more data and a more sophisticated taxonomy of the discourse functions of connectives are needed to better understand the degrees to which the lexico-semantic meaning of connectives may affect the form and the function of connectives, these tendencies suggest that different lexico-semantic properties may render different configurations of positional, phonetic, and contextual encodings in connectives. wu and tseng 117 du_initial du_medial du_final du_isolated n sentential 92 (100%) 0 0 0 92 conclusion 3 (100%) 0 0 0 3 disfluency 3 (100%) 0 0 0 3 elaboration 186 (95.88%) 8 (4.12%) 0 0 194 resumption 4 (66.67%) 2 (33.33%) 0 0 6 pu_initial pu_medial pu_final pu_isolated n sentential 57 (61.96%) 20 (21.74%) 4 (4.35%) 11 (11.96%) 92 conclusion 0 2 (66.67%) 0 1 (33.33%) 3 disfluency 1 (33.33%) 0 0 2 (66.67%) 3 elaboration 102 (52.58%) 30 (15.46%) 17 (8.76%) 45 (23.20%) 194 resumption 2 (33.33%) 1 (16.67%) 0 3 (50%) 6 table 7: distribution of discourse functions by du and pu positions of ránhòu figure 4: boxplot representation of discourse functions by speech rate, initial f0 value, and intensity mean of ránhòu (the red dot indicates the mean) figure 5: boxplot representation of discourse functions by duration of paralinguistic events in the previous and next position of ránhòu and speech rate of the du and pu where ránhòu occurs (the red dot indicates the mean) 6. discussion and conclusion in our exploratory analysis of all connective tokens in the corpus, we found that these connectives may resume a new or previous topic (biq, 1994; didirková et al., 2018; horne et al., 2001; rennie et al., 2016; wang and huang, 2006; wang, 2018; wang, 2005; wang and tsai, 2007; wang et al., 2010; yang, 2006), introduce a conclusion or summarization of the current topic (rennie et al., 2016; wang and huang, 2006), elaborate on a current or previous 0.1 0.2 0.3 0.4 s e nt en tia l c on cl us io n d is flu e nc y e la bo ra tio n r e su m pt io n discoursefunction r a te ( s ec /s y ll ) 100 150 200 250 300 s e n te nt ia l c on cl u si on d is flu e n cy e la b or at io n r e su m p tio n discoursefunction f 0 _ in it ia l (h z) 60 65 70 75 80 85 s e nt en tia l c on cl us io n d is flu e nc y e la bo ra tio n r e su m pt io n discoursefunction in te n s it y m e a n ( d b ) 0 1 2 s e n te n tia l c o nc lu si on d is flu en cy e la b o ra tio n r e su m p tio n discoursefunction p re vi o u s _ d u ra ti o n ( s ec ) 0.0 0.5 1.0 1.5 2.0 2.5 s e n te n tia l c o nc lu si on d is flu en cy e la b o ra tio n r e su m p tio n discoursefunction n ex t_ d u ra ti o n ( s e c) 0.1 0.2 0.3 0.4 0.5 0.6 s e n te n tia l c o nc lu si on d is flu en cy e la b o ra tio n r e su m p tio n discoursefunction d u _r at e (s ec /s yl l) 0.4 0.8 1.2 1.6 s e n te n tia l c o nc lu si on d is flu en cy e la b o ra tio n r e su m p tio n discoursefunction p u _r a te ( se c /s yl l) form and function of connectives in chinese conversational speech 118 topic (biq, 2001; didirková et al., 2018; rennie et al., 2016; wang and huang, 2006; wang, 2018; wang, 1998, 2005; wang et al., 2013; wang and tsai, 2007; wang et al., 2010; yang, 2006), signal a filled pause or floor holding (biq, 2001; rennie et al., 2016; rhee, 2020; yang, 2006), or signal repairs (tseng, 2006). we also observed that certain connectives may emphasize the significance of the proposition, downplay the significance of the proposition, secure the addressee’s attention, or express the speaker's disagreement. these observations seem to be in line with the findings from other mandarin connective research (hsieh and huang, 2005; wang, 2005; wang et al., 2013; wang and tsai, 2007; wang et al., 2010; yang, 2006). in this paper, we have proposed an applicable scheme that annotates the sentential, conclusion, disfluency, elaboration, and resumption uses of mandarin connectives. using statistical modeling, we revealed rates of an array of encodings of connectives for the associated discourse functions. this issue related to the interaction between form and function has not been empirically examined in previous research on connectives. concerning the positional manifestation of connectives in conversation, tokens that introduce a new topic unit or return to a previous topic most often occur at the beginning of a prosodic phrase or constitute a standalone prosodic phrase (hirschberg and litman, 1993; horne et al., 2001). we have identified differences in positional manifestation in prosodically oriented pus and semantically oriented dus. we found that disfluency use of connectives is oriented more in terms of prosodically determined production units, while conclusion, elaboration, and resumption uses of connectives are oriented more in terms of semantically determined production units. disfluency use is more likely to occur in the pu-final position and in standalone pus than sentential use. conclusion and resumption uses are more likely than sentential use to occur in the initial position of dus. elaboration use has similar patterns to those of conclusion and resumption uses, except that elaboration use is more likely than sentential use to occur in standalone dus. this may be due to the many subtypes of elaboration tokens of connectives in our corpus. a larger dataset may provide further insights into this issue. in contrast to the other discourse functions, disfluency use is more likely than sentential use to occur in du-medial positions or form standalone dus. our results confirm that connectives that indicate a topic shift in our study, such as conclusion and resumption uses, may be better observed in terms of a meaning-oriented segmentation of discourse, as reported in previous research (rennie et al., 2016; wang, 2018), rather than in prosodically segmented units. lengthened word duration often appears to introduce a new topic or to return to a previous topic (didirková et al., 2018; horne et al., 2001; wang, 2018) and can be used for the automatic detection of discourse markers (zufferey and popescu-belis, 2004). although our result did not reach significance, it shows a tendency for a lengthened duration in resumption use and a tendency for a shortened duration in conclusion, disfluency, and elaboration uses relative to sentential use. while we report that prosodic patterns of connectives may be determined by the associated discourse functions, we require a larger set of spoken connectives to statistically wu and tseng 119 verify the tendencies in the future. for pitch-related features, we found that all four discourse functions tend to have a higher initial f0 value than sentential use. if a subject were to increase the initial f0 value, resumption and disfluency uses are expected to have the largest increase, followed by a medium increase in conclusion use and a much smaller increase in elaboration use. this tendency corresponds to previous findings suggesting that when the topic shifts in conversation, the average onset f0 is higher than for other types of topic boundaries such as topic continuation, elaboration, and speech-act continuation and that elaborating utterances are characterized by the lowest onset f0 among all types of topic boundaries (nakajima and allen, 1993). on the other hand, the observation of increased initial f0 value in disfluency use may be accounted for by the fact that phrase-initial filled pauses have a higher mean pitch than phrase-medial tokens (swerts, 1998) and that error repairs tend to be marked by increased intonational prominence on the correcting information (howell and young, 1991; o'shaughnessy, 1992). there is also a pitch reset in restarting mandarin repairs, and the initial f0 in the repaired items is higher than that of its counterpart (tseng, 2006). finally, the mean intensity values seem not to be a salient phonetic factor in the context of our connective study. nonetheless, we observed that disfluency use tends to have lower mean intensity values than sentential use. this may be related to the prosodic characteristics of filled pauses described in the literature, where filled pauses are characterized by low intensity as well as low, flat f0 with reduced articulation (cole et al., 2005). we have also observed some interesting interactions between the contextual properties and the functions of connectives. we found that conclusion, elaboration, and resumption uses are more likely to have a longer duration of paralinguistic events preceding the connective than sentential use. this observation is not only in line with the finding that a relatively long pause is relevant for the automatic detection of discourse markers (zufferey and popescu-belis, 2004) but also lends empirical support to the prosodic characteristics of connectives that introduce a new topic or a specification (didirková et al., 2018). we, however, did not find any significant effect for the duration of the paralinguistic events preceding connectives in the distinction of disfluency and sentential use of connectives, even though a clear durational difference can be observed in figure 2. furthermore, we did not find any significant effect for the duration of the paralinguistic events following connectives in any of the distinctions of the sentential/discourse uses of connectives. these findings seem to contrast the observations that speakers tend to produce a pause before repair to indicate that the forward flow of speech is being interrupted (howell and young, 1991) and where discourse uses of connectives are more often followed by a pause than sentential uses of connectives (gao and tao, 2021). the discrepancies may be attributed to two factors. first, the multiple features in our multinomial logistic regression model may have overlapping predictive powers. the effects of the paralinguistic event-based features in predicting discourse uses may be partially, if not fully, captured by other features, resulting in insignificant coefficients for the paralinguistic event-based features in the model. form and function of connectives in chinese conversational speech 120 second, paralinguistic event-based features may suffer from the problem of data sparseness. while they had similar durational means in our data, only 246 connective tokens (out of a total of 1370 connective tokens) had a next_duration value larger than 0, and only 518 connective tokens had a previous_duration value larger than 0. sparse data could result in a lack of generalization performance. a larger dataset may allow for a more informative look at the association between paralinguistic events and discourse functions. finally, we found that conclusion use is more likely to occur in dus with a faster speech rate than sentential use, while disfluency use is more likely to occur in pus with a slower speech rate. this is partially consistent with the positional distribution of connectives in our data, where disfluency use was more oriented at the properties related to prosody and conclusion use was more sensitive to semantically determined production units of discourse. this study has several limitations. one important issue relates to the fine-grained categories of our taxonomy of discourse functions. although we collapsed categories that share similar functional properties, such as emphasis and downtoning, for the analysis, it is worth acknowledging that these categories also have significant implications for speaker-hearer interactions in discourse and require validation from additional empirical data. as our model can be extended to analyze higher-level functional categories, a fine-grained taxonomy validated by empirical data may give us a better picture of the phonetic orientations for the associated discourse functions and ensure the applicability of our proposed annotation scheme to various speech data. another issue concerns the effect of the lexico-semantic meaning of a connective on the form of the connective. we mentioned that the intrinsic lexico-semantic property of connectives poses differing degrees of variability to their phonetic orientations, such as in the case of ránhòu. it would be interesting to see how the semantic configuration of a connective may contribute to the overall phonetic forms of the connective. for future studies, a large set of well-annotated speech data and a sophisticated taxonomy of connectives are needed to achieve a deepened understanding of the use of connectives in conversation. in summary, in this study, we systematically investigated many dimensions of the uses of mandarin connectives in a large chinese conversational speech corpus, including their discourse functions and production. incorporating findings from a linguistic modeling perspective, we have further highlighted the connection between the discourse function and the phonetic form of connectives. we believe that the proposed methodology of integrating discourse function annotation and phonetic analyses can shed more light on the way discourse connectives contribute to the dynamic organization of discourse in conversational speech. acknowledgements this work was supported by the institute of linguistics, academia sinica and the national science and technology council, taiwan, with projects [nsc 103-2410-h-001-063-my2, most 106-2410-h-001-045-my2] granted to the second author. we would like to thank the anonymous reviewers of dialogue & discourse for their invaluable feedback. wu and tseng 121 references nicholas asher. reference to abstract objects in discourse. kluwer academic publishers, dordrecht, the netherlands, 1993. richard a. becker, john m. chambers, and allan r. wilks. the new s language. wadsworth & brooks/cole advanced books & software, pacific grove, california, 1988. yung-o biq. huìhuà hùdòngxìng hé yǔyán shǐyòng [conversational interactionality and language use]. in proceedings of the 4th annual meeting of the world chinese language, pages 227–236, taipei, taiwan, 1994. yung-o biq. the grammaticalization of jiushi and jiushishuo in mandarin chinese. concentric: studies in english literature and linguistics, 27(2):53–74, 2001. paul boersma and david weenink. praat: doing phonetics by computer (version 6.2.19). data/software, 2022. available online at http://www.praat.org/. gillian brown, karen currie, and joanne kenworthy. questions of intonation. university park press, baltimore, maryland, 1980. jennifer cole, mark hasegawa-johnson, chilin shih, heejin kim, eun-kyung lee, hsin-yi lu, yoonsook mo, and tae-jin yoon. prosodic parallelism as a cue to repetition and error correction disfluency. in proceedings of disfluency in spontaneous speech 2005, pages 53–58, aix-en-provence, france, 2005. université de provence. yves croissant. estimation of random utility models in r: the mlogit package. journal of statistical software, 95(11):1–41, 2020. doi: 10.18637/jss.v095.i11. ivana didirková, george christodoulides, and anne catherine simon. the prosody of discourse markers alors and et in french: a speech production study. in proceedings of the 9th international conference on speech prosody, pages 503–507, poznań, poland, 2018. doi: 10.21437/speechprosody.2018-102. cecilia e. ford. the treatment of contrasts in interaction. in elizabeth couper-kuhlen and bernd kortmann, editors, cause-condition-concession-contrast: cognitive and discourse perspectives, pages 283–312, mouton de gruyter, 2000. hua gao and hongyin tao. fanzheng ‘anyway’ as a discourse pragmatic particle in mandarin conversation: prosody, locus, and interactional function. journal of pragmatics, 173:148–166, 2021. doi: 10.1016/j.pragma.2020.12.003. agustín gravano, štefan benuš, julia hirschberg, shira mitchell, and ilia vovsha. classification of discourse functions of affirmative words in spoken dialogue. in proceedings of the 8th annual conference of the international speech communication association, pages 1613–1616, antwerp, belgium, 2007. international speech communication association. doi: 10.21437/interspeech.2007450. agustín gravano, julia hirschberg, and štefan benuš. affirmative cue words in task-oriented dialogue. computational linguistics, 38(1):1–39, 2011. doi: 10.1162/coli_a_00083. form and function of connectives in chinese conversational speech 122 zhiliang guo. xiàndài hànyǔ zhuǎnzhé cíyǔ yánjiù [a study of modern chinese transitional words]. beijing language and culture university press, beijing, 1999. jerry hausman and daniel mcfadden. specification tests for the multinomial logit model. econometrica, 52(5):1219–1240, 1984. doi: 10.2307/1910997. julia hirschberg and diane litman. now let’s talk about now: identifying cue phrases intonationally. in proceedings of the 25th annual meeting of the association for computational linguistics, pages 163–171, stanford, california, 1987. association for computational linguistics. julia hirschberg and diane litman. empirical studies on the disambiguation of cue phrases. computational linguistics, 19(3):501–530, 1993. anna hjalmarsson. the additive effect of turn-taking cues in human and synthetic voice. speech communication, 53(1):23–35, 2011. doi: 10.1016/j.specom.2010.08.003. merle horne, petra hansson, gösta bruce, johan frid, and marcus filipsson. cue words and the topic structure of spoken discourse: the case of swedish men ‘but’. journal of pragmatics, 33(7):1061–1081, 2001. doi: 10.1016/s0378-2166(00)00044-8. peter howell and keith young. the use of prosody in highlighting alterations in repairs from unrestricted speech. the quarterly journal of experimental psychology, 43(3):733– 758, 1991. doi: 10.1080/14640749108400994. fuhui hsieh and shuanfan huang. grammar, construction and social action: a study of the qíshí construction. language and linguistics, 6(4):599–634, 2005. yi-fen liu, shu-chuan tseng, and jyh-shing roger jang. deriving disyllabic word variants from a chinese conversational speech corpus. the journal of the acoustical society of america, 140(1):308–321, 2016. doi: 10.1121/1.4954745. john local and john kelly. projection and ‘silences’: notes on phonetic and conversational structure. human studies, 9(2-3):185–204, 1986. doi: 10.1007/bf00148126. shuxiang lü. xiàndài hànyǔ bābǎi cí [800 words in modern chinese]. commercial press, hong kong, 1980. w. charles miracle. discourse markers in mandarin chinese. unpublished phd dissertation, ohio state university, columbus, ohio, 1991. shin’ya nakajima and james f. allen. a study on prosody and discourse structure in cooperative dialogues. phonetica, 50(3):197–210, 1993. doi: 10.1159/000261940. douglas o'shaughnessy. recognition of hesitations in spontaneous speech. in proceedings of the 1992 ieee international conference on acoustics, speech, and signal processing, 1, pages 521–524, san francisco, 1992. ieee. doi: 10.1109/icassp.1992.225857. janet b. pierrehumbert. the phonology and phonetics of english intonation. phd dissertation, massachusetts institute of technology, cambridge, massachusetts, 1980. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and wu and tseng 123 bonnie webber. the penn discourse treebank 2.0. in proceedings of the 6th international conference on language resources and evaluation, pages 2961–2968, marrackech, morocco, 2008. european language resources association. laurent prévot, shu-chuan tseng, klim peshkov, and alvin cheng-hsien chen. processing units in conversation: a comparative study of french and mandarin data. language and linguistics, 16(1):69–92, 2015. doi: 10.1177/1606822x14556605. r core team. r: a language and environment for statistical computing. data/software, 2021. available online at https://www.r-project.org/. emma rennie, rebecca lunsford, and peter a. heeman. the discourse marker "so" in turntaking and turn-releasing behavior. in proceedings of the 17th annual conference of the international speech communication association, pages 1280–1284, san francisco, 2016. international speech communication association. seongha rhee. on the many faces of coarseness: the case of the korean mak ‘coarsely’. journal of pragmatics, 170:396-412, 2020. doi: 10.1016/j.pragma.2020.09.025. david robinson, alex hayes, and simon couch. broom: convert statistical objects into tidy tibbles. data/software, 2022. available online at https://cran.rproject.org/package=broom. claudia ross. contrasting conjunctions in english, japanese, and mandarin chinese. unpublished phd dissertation, university of michigan, ann arbor, michigan, 1978. marc swerts. filled pauses as markers of discourse structure. journal of pragmatics, 30(4):485–496, 1998. doi: 10.1016/s0378-2166(98)00014-9. shu-chuan tseng. syllable contractions in a mandarin conversational dialogue corpus. international journal of corpus linguistics, 10(1):63–83, 2005. doi: 10.1075/ijcl.10.1.04tse. shu-chuan tseng. repairs in mandarin conversation. journal of chinese linguistics, 34(1):80–120, 2006. shu-chuan tseng. ilas chinese spoken language resources. in proceedings of the third international symposium on linguistic patterns in spontaneous speech, pages 13–20, taipei, taiwan, 2019. w. n. venables and b. d. ripley. modern applied statistics with s. springer, new york, 2002. chueh-chen wang and lillian m. huang. grammaticalization of connectives in mandarin chinese: a corpus-based study. language and linguistics, 7(4):991–1016, 2006. wei wang. discourse uses and prosodic properties of ranhou in spontaneous mandarin conversation. chinese language and discourse, 9(1):1–25, 2018. doi: 10.1075/cld.00006.wan. yu-fang wang. the functions of ranhou in chinese oral discourse. in proceedings of the 9th north american conference on chinese linguistics, 2, pages 380–397, los angeles, 1998. gsil publications. form and function of connectives in chinese conversational speech 124 yu-fang wang. from lexical to pragmatic meaning: contrastive markers in spoken chinese discourse. text, 25(4):469-518, 2005. doi: 10.1515/text.2005.25.4.469. yu-fang wang, mei-chi tsai, wayne schams, and chi-ming yang. restrictiveness, exclusivity, adversativity, and mirativity: mandarin chinese zhishi as an affective diminutive marker in spoken discourse. chinese language and discourse, 4(2):181– 228, 2013. doi: 10.1075/cld.4.2.02wan. yu-fang wang and pi-hua tsai. textual and contextual contrast connection: a study of chinese contrastive markers across different text types. journal of pragmatics, 39(10):1775-1815, 2007. doi: 10.1016/j.pragma.2007.05.011. yu-fang wang, pi-hua tsai, and ya-ting yang. objectivity, subjectivity and intersubjectivity: evidence from qishi (‘actually’) and shishishang (‘in fact’) in spoken chinese. journal of pragmatics, 42(3):705–727, 2010. doi: 10.1016/j.pragma.2009.07.011. li-chiung yang. integrating prosodic and contextual cues in the interpretation of discourse markers. in kerstin fischer, editor, approaches to discourse particles, pages 265– 297, elsevier, 2006. sandrine zufferey and liesbeth degand. annotating the meaning of discourse connectives in multilingual corpora. corpus linguistics and linguistic theory, 13(2):399–422, 2017. doi: 10.1515/cllt-2013-0022. sandrine zufferey and andrei popescu-belis. towards automatic identification of discourse markers in dialogs: the case of like. in proceedings of the 5th sigdial workshop on discourse and dialogue, pages 63–71, cambridge, massachusetts, 2004. association for computational linguistics. doi: 10.7892/boris.78686. dialogue & discourse 16(3) (2025) 8–24 doi: 10.5210/dad.2025.302 laughter use by virtual agents increases task success bogdan ludusan bogdan.ludusan@uni-bielefeld.de phonetics workgroup, faculty of linguistics and literary studies & citec bielefeld university, germany petra wagner petra.wagner@uni-bielefeld.de phonetics workgroup, faculty of linguistics and literary studies & citec bielefeld university, germany editor: dimitra gkatzia submitted 11/2023; accepted 12/2024; published online 12/2025 abstract studies on human-robot interaction as well as on embodied conversational agents have revealed that the use of laughter by agents increases their perceived naturalness and their social presence. however, laughter plays a variety of functions in human interaction, and its effects on communication go beyond those previously investigated in the aforementioned fields. taking into account that laughter has been shown to improve task performance in human-human interaction, we investigated here whether laughter use by a virtual agent increases task success also in human-machine interaction. a real-estate scenario was considered, in which an agent presented an apartment to an interested client. both the presence of laughter and the nature of the agent (virtual or human) were varied in the experiment. we operationalized the task success as being the likelihood of participants recommending the apartment, while also examining the perceived rating of the agent. the results of an observer study showed that the use of laughter by a virtual agent translates in an increased task success, while also confirming previous findings regarding improvements in the social perception of the agent. our results concerning the task success in the human agent condition were not in line with those of previous studies, most likely due to a reduced naturalness of the used laughter. this makes the findings pertaining to the virtual agent, where benefits were observed by the use of laughter in interaction, even more salient. taken together, these results seem to suggest that in this case humans are less sensitive to a reduced laughter naturalness. we further discuss the need for better laughter integration with speech, as well as its automatic synthesis in order to better take advantage of these findings. keywords: laughter, human-machine interaction, virtual agent, human-human interaction, task success 1. introduction robots employ a wide range of signals to interact with humans, including verbal and non-verbal ones. in particular, it has been shown that a robot’s non-verbal expression may elicit emotional or behavioral responses from their interlocutor, as well as having an impact on the user’s task performance (saunderson and nejat, 2019). recent years have seen an increased research interest in this field in sound-based non-verbal expressions, as these can be straightforwardly realized also in robots that are constrained in their appearance or in their interactional capabilities (i.e., robots that cannot produce complex gestural or facial expressions, or have limited capacities for main©2025 bogdan ludusan and petra wagner this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). laughter use and task success taining a dialog). an appealing characteristic of sound-based non-verbal expressions is that they are language-independent, quickly understood, and can crucially contribute to the perception of a robot’s personality (ritschel et al., 2019). much research on robot’s non-verbal behaviour focuses on artificially generated and not necessarily human-like sounds, for example midi-like tones. these often exploit universally understood, iconic relationships between sounds and meaning (such as high pitch representing “happy” and low pitch corresponding to “sad”) (zhang and fitter, 2023). here, we choose an alternative path by looking at a genuinely human, but nevertheless non-verbal signal – laughter, and we evaluate its usefulness for the interaction between a human and a virtual agent. laughter is a multi-faceted phenomenon playing various social and communicative roles in human interaction (glenn, 2003). besides the role it plays in humor (mirthful laughter), social (non-mirthful) laughter may be used to convey both positive or cooperative behaviours and negative ones. among the positive aspects, laughter may indicate closeness or affiliation, or a feeling of pleasantness with regards to the conversational partner. it may evoke also negative emotions, with laughing at someone contributing to disaffiliation or hostility towards that person. further functions of laughter may include the softening of previous affirmations, the remedy of possible interactional misunderstandings, or the handling of difficult situations (e.g., involving embarrassment or discomfort). considering the widespread use of laughter in human-human communication and the multitude of functions that it may have in conversation, laughter should be beneficial to the interaction, increasing its success one way or the other. this has been found also for interactions involving goaloriented tasks, with previous work showing that the use of shared laughter (laughter initiated by one of the interlocutors and taken up by their conversational partner) may improve, for instance, task performance in workplace meetings (kangasharju and nikko, 2009), as well as increase the hiring chances of candidates at job interviews (brosy et al., 2021). the above-mentioned characteristics of laughter have encouraged its implementation in human-machine interaction systems. this has been done either for general spoken dialogue systems (maraev et al., 2018), or by integrating it in embodied conversational agents (niewiadomski et al., 2013; pecune et al., 2015; el haddad et al., 2016; mancini et al., 2017) or in robots (becker-asano and ishiguro, 2009; türker et al., 2017; inoue et al., 2022). some of these studies have focused exclusively on the required technical implementation of laughter in these systems (maraev et al., 2018), while others only on the accuracy/naturalness of specific sub-systems (e.g., laughter prediction/synthesis el haddad et al., 2016). experiments to quantify the effect of agent laughter use on the perception of the agent by their human interlocutor or on the overall interaction have also been conducted. in niewiadomski et al. (2013), the participants compared a laughing agent that produced shared laughter with them, to an agent not using laughter or employing laughs at predefined time instants. they perceived: 1) the video clips they were watching together with the agent as being more funny, and 2) the laughter of the agent more contagious, in the former case (the latter finding reflecting views found in the literature examining laughter in human communication provine, 1996). although no increased social presence of the agent was perceived, a subsequent study (pecune et al., 2015) in which the embodied agent produced shared laughter with the user, by copying the latter’s laughter intensity and trunk movements reported higher scores for agent social presence compared to the case when there was no shared laughter (but see mancini et al., 2017 for the opposite result). responding to the laughter of the human interlocutor also improved a number of measures related to user engagement 9 ludusan and wagner in türker et al. (2017). lastly, a system employing a more complex decision on whether the agent should initiate shared laughter or not and which type of laughter to use, resulted in increased ratings (e.g., naturalness, human likeness) for some of the tested scenarios in inoue et al. (2022), compared to a system responding to the laughter of the user every time. these studies have shown that the use of laughter by a virtual agent resulted in a higher rating given to that agent, in line to findings in human-human interaction which show that laughter use is beneficial to the relationship between interlocutors (kurtz and algoe, 2017) (see also becker-asano and ishiguro 2009, in which the use of laughter has elicited also negative emotions towards the embodied agent, a normal reaction if the person perceives that they are being laughed at). we have seen that a virtual agent that laughs is judged to be more natural or more human-like (becker-asano and ishiguro, 2009; niewiadomski et al., 2013; inoue et al., 2022) and to have an increased social presence (e.g., be more empathetic and understanding pecune et al., 2015; inoue et al., 2022), while also increasing the engagement with its interlocutor (türker et al., 2017). however, there is little work that proves that the use of laughter by a virtual agent is conducive to the success of the task performed with the user, beyond the improved social perception. while improved social characteristics may, indirectly, result in an increased success of the task, more direct evidence is needed. we present a study investigating whether laughter use by a virtual agent increases the success of a task that agent is carrying out. the task consisted of a real-estate scenario, in which the agent was showing to a client an apartment for rent, with the participants to the experiments being observers of the task. the success of the interaction was operationalized as the degree to which the apartment would be recommended by the participants to their acquaintances. furthermore, similarly to other studies, we examined the perception of different characteristics of the agent, as well as an overall rating of the agent, defined as the degree to which the agent would be recommended by the participants. we did not limit ourselves to considering a virtual agent that uses laughter versus one that does not laugh (ludusan and wagner, 2021), but we also compared it to how a human agent would be perceived at the same task, including the effect of laughter on task success. we extended the study in ludusan and wagner (2021) by testing three additional conditions: a human agent employing either no laughter, laughter, or “more natural” laughter, and performing a joint analysis of all conditions (virtual and human agent), with statistical models that better fit the distribution of the data. 2. materials and methods 2.1 stimuli an interaction between a real-estate agent and a client visiting an apartment was designed for the experiment. we decided on using this particular situational setting because conversations involving real-estate scenarios have been shown to be useful for evaluating virtual agents (cassell et al., 1999), while also providing a quantifiable measure of success for the interaction – whether the apartment was liked by the experiment participants. the conditions considered in this study are illustrated in figure 1. we varied two factors: the nature of the real-estate agent (human/virtual) and the use of laughter by the agent (no/yes), resulting in four conditions, representing a two-by-two crossed design. we also considered a fifth condition, the third in the case of the human agent, to evaluate the appropriateness of the laughter used in the other two laughter conditions. 10 laughter use and task success we started by synthesizing the conversation lines belonging to the virtual agent, using a commercially-available state-of-the-art system1, based on one of the wavenet male german voices (de-de-wavenet-b). as we wanted to have an intonation more similar to conversational speech, rather than the default intonation employed by synthesis systems (trained mainly on read speech data), we tested several parameters (e.g., placement of punctuation marks, duration of pauses, the use of emphasis) during the synthesis in order to achieve this, while avoiding the introduction of any artefacts. the obtained stimuli were vetted by two native speakers and any reported inconsistencies were corrected. the resulting materials composed the speech of the virtual agent in the non-laughter condition. then, we produced the lines corresponding to the virtual agent, in the laughter condition. conversational laughter instances produced by male speakers were extracted from the recordings of the duel corpus (hough et al., 2016). the corpus contains dyadic task-based interactions that encourage the use of conversational laughter. the chosen instances had the characteristics shown previously to be well integrated with synthetic speech (trouvain and schröder, 2004): being composed of one to three laughter syllables, either voiced or unvoiced (snort-like), and having low to medium intensity. eight positions within the lines of the agent were identified as being appropriate to contain laughter and we spliced at those positions suitable laughter tokens (in terms of pitch, timbre, or expressed function). these employed laughs were either three-, twoor one-syllable long (two, four, and two instances, respectively), with two of the two-syllable long laughs being unvoiced, the remaining laughs being voiced. figure 1: an illustration of the stimuli creation procedure. the lines for the virtual agent were created using a state-of-the-art speech synthesis system, while those corresponding to the human agent were recorded by a native german speaker. the agent used either no laughter (cross symbol), laughter extracted from corpora of conversational speech (database symbol) or laughter instances produced by the native speaker. next, we recorded the lines corresponding to the client in the recording studio of the university. a female native speaker of german familiarized herself beforehand with the lines of the client. 1. https://cloud.google.com/text-to-speech 11 ludusan and wagner during the recording session, the experimenter played through loudspeakers the line corresponding to the agent (laughter condition) and the speaker would respond with the appropriate line. in order to increase the naturalness of the recordings, the speaker was allowed to change the exact wording of the sentence, provided that its meaning stayed the same. several takes of the entire conversation were performed and the most spontaneously-sounding lines were employed in the experiment. two versions of the stimuli were put together: for the non-laughter condition we combined the synthesized speech not containing laughter with the lines of the female client, while for the laughter case, the synthesized speech including the spliced laughs with the lines of the client. thus, the only difference between the two obtained recordings were the absence/presence of laughter in the virtual agent’s speech. finally, we created equivalent conditions involving a human agent. a male german native speaker recorded the lines of the agent in a similar setting to how the female client lines were recorded, hearing through the loudspeakers the sentences produced by the client and answering with the corresponding agent line. however, he was not allowed to rephrase the sentences, in order to have identical materials in the virtual and human agent cases. for the laughter condition, the speaker was asked to produce laughs which were as similar as possible to the ones employed for the virtual agent. therefore, we had a design varying two types of agents (human/virtual) and two laughter conditions (absence/presence). as mentioned at the beginning of this section, we also considered a control condition for the human agent, in order to check the appropriateness of the laughter used. in that condition, the speaker was asked to laugh as he would normally laugh at that particular point of the interaction, instead of producing laughs similar to the ones used in the virtual agent condition. three of the employed laughter instances were speech-laughs (defined as concurrent production of speech and laughter in which neither of the two components is dominant nwokah et al., 1999), with the rest being laughs, having an overall lower intensity than in the human agent, laughter condition. the audio stimuli2 corresponding to the virtual agent were around 4 minutes 30 seconds long, while those of the human agent were circa 3 minutes 50 seconds long, due to differences between the speech rate of the synthesized speech of the virtual agent and that of the human speaker. the recordings of each of the five experimental conditions were combined with the images of an apartment, in order to create a video clip. the employed images did not contain any visual representation of the agent or of the client. at any time, the displayed image matched the location where the discussion was taking place (hallway, rooms, etc.) and the images changed at the same points in the interaction in all conditions. the resulting videos, having identical visual content and different audio components, depending on the condition, represented the stimuli used in this study. 2.2 procedure the experiment was run on the psytoolkit platform (stoet, 2010, 2017). the participants watched the video containing the slide show of the apartment and the recording between the agent and the client, having been informed of the nature of the agent (human/virtual) at the beginning of the experiment. after that, they were asked to judge the performance of the agent with respect to several dimensions: professionalism, communication, pleasantness, formality and spontaneity. next, they were asked to rate the probability of recommending the apartment (our task success measure) and the agent to their acquaintances. for the human agent conditions, since we also wanted to evaluate how the employed laughter was perceived when used by a human, the participants also replied to 2. https://pub.uni-bielefeld.de/download/2999852/3000681.zip 12 laughter use and task success two additional questions on the politeness and the empathy shown by the human agent. lastly, the human agent conditions involving laughter (the original and the “more natural” laughter one) had two other questions asking the participants whether they perceived the agent to be laughing at the correct locations (from “always wrong” to “always right”) and if the laughter sounded appropriate for that situation (from “not appropriate” to “appropriate”). all the questions used a 10-point scale (e.g., for professionalism, the value 1 of the scale was labelled as “unprofessional”, while the value 10 corresponded to “professional”). the exact formulation of the questions can be found in appendix a. the study was reviewed and approved by the ethical committee of bielefeld university (eub no. 2019-085 and no. 2020-165). the majority of the participants were recruited through the prolific crowdsourcing platform3 and were paid according to the regular rate for participating in studies at bielefeld university. the rest of them (about 10%) were either students at bielefeld university (receiving course credit for their participation) or friends and family of the authors. all participants had to be residing in germany and be fluent in german. we used a between-participants design, each participant taking part in only one condition. of the 346 participants that completed the study, we had to exclude the data of 16 participants who did not listen to the entire conversation. the remaining participants4 were split evenly (66 each) between the four conditions (human/virtual agent, no laughter/laughter). we aimed to have a balanced number of male and female participants in each condition (see table 1 for exact numbers). agent type laughter present # female # male # other total virtual no 34 32 0 66 yes 31 34 1 66 human no 33 31 2 66 yes 33 33 0 66 yes* 33 33 0 66 table 1: gender distribution (female/male/other) of the participants in the experiment, across all tested condition. the human agent yes* condition represents the case in which the human agent produces “more natural” laughter. 2.3 analyses all participant data from the non-laughter conditions, as well as from the laughter conditions using the same type of laughter (yes), were considered in the main analyses. three participants, self-identified as neither female or male, were excluded from the analyses involving the predictor gender, since not enough data points were available for inclusion. in order to determine the effects of each of the investigated characteristics on the considered success measure, we employed generalized linear regression models. as the data best fitted a poisson distribution, but it was underdispersed, we used a generalized poisson distribution, conway-maxwell poisson, that handles both over-dispersed and under-dispersed data. the dependent variable of the models was either the apart3. https://www.prolific.com/ 4. they include also a set of 66 participants, equally split between males and females, that was tested in the human agent, “more natural” laughter condition. 13 ludusan and wagner ment or the agent rating and, as predictors, the following: the agent type (human/virtual), the laughter presence (no/yes), the characteristics rated by the participants (communication, professionalism, pleasantness, formality and spontaneity), as well as the age and the gender of the participants. we also included all twoand three-way interactions between agent type, laughter presence, and each of the other predictors. this model was then reduced through step-wise elimination (matuschek et al., 2017), by removing, at each step, the factor whose elimination would reduce the most the akaike information criterion (aic) value of the model. the process was stopped when the removal of any factor did not decrease the aic of the model. all numeric predictors were scaled, by subtracting their mean value and dividing them by two standard deviations (gelman, 2008). a sum to zero contrast was employed for the factor variables involved in the model. all analyses were performed using the appropriate functions of the r software (r core team, 2020, ver. 3.6.3), the models being fitted using the glm.cmp function of the compoissonreg package (sellers et al., 2022). 3. results 3.1 overall results the per-condition results obtained for the apartment rating, our task success measure, are illustrated in figure 2. a visual inspection of the results reveals no effect of laughter in the human agent condition, but an increase in apartment rating when the virtual agent used laughter, compared to when they did not. human agent virtual agent no laughter laughter no laughter laughter 0.0 2.5 5.0 7.5 10.0 condition a pa rt m en t r at in g figure 2: the apartment rating obtained in the four conditions considered in the study: human/virtual agent, no laughter/laughter. the bars inside each violin plot illustrate the median and the quartile values of the apartment rating. the generalized linear regression model predicting the apartment rating showed that the presence of laughter had a significant effect directly, the apartment being rated higher if the agent used laughter (β = 0.126, p = 0.029), as well as through its interaction with the age of the participant (β = 0.310, p = 0.017). significant effects were found also for professionalism (β = 0.328, p = 14 laughter use and task success 0.012), communication (β = 0.365, p = 0.021), and pleasantness (β = 0.477, p = 0.001), the apartment rating increasing with the increased rating of these characteristics, and for the age of the participant, with older participants rating the apartment lower (β = −0.294, p = 0.023). however, in the condition in which the agent laughed, the apartment was rated higher by older participants (see previously reported laughter-age interaction), suppressing the negative main effect seen by age. the rating of the agent in the various conditions is displayed in figure 3, while results related to the perceived characteristics of the agent are given in table 2. although laughter seems to have a more reduced effect for the agent rating, compared to the apartment rating, we verified the existing trends by means of a generalized linear regression model fitted with the agent rating. human agent virtual agent no laughter laughter no laughter laughter 0.0 2.5 5.0 7.5 10.0 condition a ge nt r at in g figure 3: the agent rating obtained in the four conditions considered in the study: human/virtual agent, no laughter/laughter. the bars inside each violin plot illustrate the median and the quartile values of the rating. the fitted model did not show a main effect of laughter, but significant interactions between agent type and laughter presence, the virtual agent being rated higher when they laughed (β = 0.115, p = 0.016), and between laughter and formality, the rating of an agent using laughter decreasing with its perceived increased formality (β = −0.208, p = 0.039). further significant effects were found for professionalism (β = 0.234, p = 0.049), communication (β = 0.491, p = 4.8e−4) and pleasantness (β = 1.201, p = 4.7e−13) and for the interactions type-communication (β = 0.384, p = 0.002) and type-laughter-professionalism (β = 0.261, p = 0.028). judging the agent to have better communication skills, to be more professional and to be more pleasant increased the rating of the agent. moreover, rating the virtual agent higher on communication increased its rating in addition to the main effect seen by communication. 3.2 agent type analyses in order to better understand the results, we performed separate analyses, for the human and for the virtual agent data, respectively. each sub-analysis looked at both the agent and the apartment rating, employing them as dependent variables in regression models. the models were similar to 15 ludusan and wagner agent laughter prof. comm. plea. form. spont. agent apartment type present rating rating human no 8.58 8.71 8.64 5.03 5.52 8.03 8.39 yes 8.30 8.59 8.39 4.29 6.17 7.71 8.48 virtual no 8.56 8.29 7.56 6.14 5.36 6.59 8.12 yes 8.29 8.21 7.33 4.97 5.32 7.03 8.45 table 2: mean values of the scores given by the participants to the seven dimensions evaluated in the experiment (professionalism prof., communication comm., pleasantness plea., formality form., spontaneity spon., agent rating, and apartment rating), for each type of agent (human/virtual) and each laughter condition (no laughter/laughter). the ones previously described, except for the agent type no longer being a predictor. moreover, the models fitted with the virtual agent data included two additional predictors: the self-reported speech technology exposure of the participants and its interaction with laughter use. the first model fitted with the virtual agent data showed that the apartment rating depended on laughter use and on the perceived communication skill of the agent, as well as on the interaction between laughter and each of the following three factors: communication, speech technology exposure and age. a virtual agent using laughter (β = 0.184, p = 0.044) and having a higher communication skill (β = 0.699, p = 6.9e−4) increased the rating of the apartment. the apartment presented by a virtual agent that laughed was also rated higher by participants reporting a higher speech technology exposure (β = 0.455, p = 0.008) and by older participants (β = 0.479, p = 0.045), and rated lower when the agent was judged to have higher communicative skills (β = −0.382, p = 0.020). the second model for the virtual agent condition revealed that the agent rating increased when the agent laughed (β = 0.211, p = 0.004), while it decreased when the virtual agent used laughter and its perceived formality increased (β = −0.304, p = 0.041). communication (β = 0.640, p = 6.7e−4) and pleasantness (β = 1.053, p = 1.6e−7) also had a significant effect on the agent rating, the latter being higher when the agent was perceived to have increased communication skills and a higher pleasantness. for the models built on the human agent data (for the apartment and for the agent rating), the only predictors that had a significant effect on the dependent variables were professionalism (β = 0.539, p = 0.007 for the apartment and β = 0.672, p = 1.4e−4 for the agent rating) and pleasantness (β = 0.773, p = 0.002 and β = 1.620, p = 9.1e−9 for the apartment and the agent rating, respectively). an increase in the scores given to these characteristics resulted in an increase in both the apartment and the agent rating. 3.3 detailed analysis of the human agent conditions an analysis of the additional questions asked to the participants in the human agent conditions was performed. it examined the politeness and the empathy of the agent, as well as the perceived correctness of the place where the agent laughed and the perceived appropriateness of the employed laugh. we display these results in table 3, along with all the scores given by participants in the “more natural” (yes*) laughter condition. for an easier comparison, we illustrate all the ratings obtained in the other two human agent conditions. 16 laughter use and task success laughter prof. comm. plea. form. spont. agent apt. pol. emp. place appr. present rating rating no 8.58 8.71 8.64 5.03 5.52 8.03 8.39 9.15 7.73 yes 8.30 8.59 8.39 4.29 6.17 7.71 8.48 9.21 7.68 6.30 6.08 yes* 8.55 8.83 8.70 3.92 6.41 8.03 8.02 9.23 8.23 7.70 7.36 table 3: mean values of the scores given by the participants in the three human agent conditions: no laughter (no), laughter (yes) and “more natural” (yes*) laughter. presented are the seven dimensions evaluated in the experiment (professionalism prof., communication comm., pleasantness plea., formality form., spontaneity spon., agent rating, and apartment rating), as well as the perceived politeness (pol.) and perceived empathy (emp.) of the agent, and the appropriateness of the place where laughter was used (place) and of the laughter itself (appr.). we employed the same type of generalized linear regression model as in the previous analyses to determine whether there was an effect of laughter condition on several scores. the following scores were considered: apartment rating, agent rating, perceived agent politeness, perceived agent empathy, appropriateness of the place where laughter was used, and appropriateness of the type of employed laughter. we performed pairwise comparisons between the conditions: no laughter (no) / laughter (yes) and laughter (yes) / “more natural” laughter (yes*). there was no effect of the presence of laughter (no/yes conditions) on any of the investigated measures: on the apartment rating (β = 0.023, p = 0.760), on the agent rating (β = −0.0521, p = 0.280), on the perceived agent politeness (β = 0.008, p = 0.912), or on the perceived agent empathy (β = −0.010, p = 0.851). also the type of laughter produced (similar to the one employed by the virtual agent vs. the “more natural” laughter) did not have any effect on the apartment rating (β = 0.086, p = 0.116), on the agent rating (β = −0.041, p = 0.356), on the perceived agent politeness (β = −0.006, p = 0.938), or on the perceived agent empathy (β = −0.108, p = 0.055). there were differences, however, between the two conditions on the remaining examined measures: the appropriateness of the place where the laughter was used (β = −0.183, p = 2.5e−4) and the appropriateness of the type of employed laughter (β = −0.126, p = 0.002). in both cases, the laughter produced similarly to how the virtual agent laughed was perceived less appropriate. however, the place of the laughter events did not vary between the two conditions and the significant difference in reported laughter place must be due to the participants not being able to separate the two appropriateness measures. this comparison shows that, while the participants did perceive the laughs produced by the human agent in the “more natural” condition as being more appropriate, these laughs did not have an effect on the overall perception of the agent or on the task success measure. 4. discussion and conclusions investigating the use of laughter in conversation by a human or a virtual agent, we observed an effect of laughter on the success of the task, defined as the likelihood of recommending the apartment presented by the agent, with a higher task success when the agent employed laughter. this effect was mainly driven by the ratings given in the virtual agent case, where also the likelihood of recommend17 ludusan and wagner ing the agent increased with the use of laughter. these findings add to the body of evidence that laughter helps the interaction with human interlocutors. our study took one step further, though, by providing confirmation that the use of laughter by a virtual agent increases not only how the agent is perceived socially, but also the success rate of the performed task, when compared to an agent that does not laugh. although, due to the experimental setting employed here, there was no explicit embodiment given to the virtual agent, we believe our findings may be easily extended to embodied agents, taking into account that the physical embodiment of an agent actually increases their evaluation by humans (lee et al., 2006), and may also have an indirect effect on task success. moreover, similar findings have been previously reported concerning the use of laughter, regardless whether a conversational agent (niewiadomski et al., 2013; pecune et al., 2015) or a robot (becker-asano and ishiguro, 2009; türker et al., 2017; inoue et al., 2022) was employed in the experiment. we compared the effect of laughter use also by a human agent, on the same task, revealing that the increase in task success can be seen only in the virtual agent case. our results did not corroborate those previously observed in human-human interaction (kangasharju and nikko, 2009; brosy et al., 2021) and this lack of an increase in task success may be explained by the reduced naturalness of the laughter stimuli used in our study, probably caused by how they were elicited. while no differences, neither in the apartment or agent rating, nor in the perceived politeness or empathy of the agent, have been observed between the human agent laughter and “more natural” laughter conditions, the fact that the participants indicated differences in the appropriateness of laughs used (including their place, although that was not the case) might be an indicator of this limited naturalness. in order for us to be able to directly compare the human with the virtual agent conditions we had to use the same types of laughter and thus, we were limited by the selection available to use in the virtual agent conditions. however, despite the reduced naturalness of the employed laughs, in both the human and the virtual agent conditions, it seems that this was more strongly perceived in the former, rather than in the latter conditions. these results are encouraging for using laughter in humanmachine interaction, seeing how humans are more forgiving towards reduced laughter naturalness in a machine than towards a human. furthermore, our results provide evidence that it may be overly simplistic to expect that any behavior that is either helping or impeding interactions between human interlocutors, may have the same or, at least similar, effects on communication between humans and machines (but see krämer et al. 2012 for a different view). our findings may also indicate that humans may have differentiated expectations from virtual agents, which influence their assessment and processing of the virtual agent’s signal, in line with previous work in the literature (horstmann and krämer, 2020). we believe that overall more work is needed on the integration of laughter in human-machine interaction systems. currently, the integration of social laughter (such as here) is mainly done by means of natural laughter instances extracted from speech corpora (inoue et al., 2022). this seems to be due to limited capabilities of the current synthesis systems to generate naturally sounding laughter. state-of-the-art laughter synthesis systems based on deep-learning approaches (mori et al., 2019; tits et al., 2020), reach mean opinion scores of around three (on a scale up to five, on which natural laughter is evaluated between four and five). even if the quality of laughter synthesis reaches a good level, one needs to take into account that the human laughter inventory does not only contain laughs. the acoustic production of laughter varies considerably (bachorowski et al., 2001) and speech-laughs are often employed in spontaneous communication (ludusan and wagner 2019b reported between 20%-40% of all laughter instances to be speech-laughs, across materials in three languages). we observed this also in our data, with the human “more natural” laughter 18 laughter use and task success condition containing three instances of speech-laughs (out of a total of eight laughter events) and an inspection of the stimuli revealed that the use of speech-laughs was appropriate in those cases (maybe even more so than that of laughs). unfortunately, there is less work done on the synthesis of speech-laughs and the existing studies suggest it being even more difficult than laughter synthesis, with lower mean opinion scores (el haddad et al., 2015). finally, as noted also by türker et al. (2017), it is difficult to incorporate naturalistic laughter (be it laughs or speech-laughs) to synthetic speech. this is made even more challenging by the fact that speech preceding laughter is subject to co-articulation effects (ludusan and wagner, 2019a). for instance, in the “more natural” laughter condition, the produced laughter instances can be predicted before their actual occurrence, due to the acoustic changes in vocal tract in anticipation to the following laughter. thus, more advances in speech technology are needed to seamlessly integrate laughter in synthetic speech, preferably through joint synthesis of speech and laughter (see also more recent attempts to generate conversational phenomena, such as laughter, directly in the synthesized speech signal, using deep learning models (nguyen et al., 2023)). some observations on the laughter employed in the conversation are also in order. the improvements previously reported in the literature (on how the agent is perceived) were found when using shared laughter. in our case, out of the eight laughter events included in the stimuli, three of them were shared. while shared laughter seems to play an important role in human-machine interaction studies (e.g., niewiadomski et al., 2013; türker et al., 2017) and also in human-human communication (e.g., ludusan and wagner, 2019b), our findings show that not all laughs produced by the virtual agent need to be synchronised to those of the human interlocutor. further analyses of the functions played by laughter will be required to better understand its roles in conversation and to better model laughter in human-machine interaction. future work could also include an investigation into whether an increase in task success may be seen for other type of settings (e.g., social settings – personal assistants). another factor worth exploring is the effect of the perspective on the interaction with the virtual agent: either a first-person one, such as when the user actively participates in the interaction, or a third-person one, when they are an observer to the study (similar to here). an increase in task success in the case of laughter use by a virtual agent, in a first-person experiment, would be of particular applied importance. lastly, independent of the experimental setting and the task involved, the human-machine interaction system will need to ensure an appropriate integration of laughter within the synthesized speech, one that takes into account also the functions that laughter plays, similar to the proposals put forward by trouvain and weiss (2022) for smiled speech. acknowledgments the authors would like to thank paul kaminski, leonie schade and marin schröer for their help with setting up the dialogue and annett jorschick for helpful discussions regarding the experiment and the statistical analyses. part of the results included here have been previously presented at essv 2021. this work was funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) project number 461442180. it was also supported by the european union’s horizon 2020 research and innovation programme under the marie sklodowska-curie grant agreement no. 799022. 19 ludusan and wagner appendix a this appendix contains the questions asked to the participants after watching the video containing the interaction between the agent and the client. the following questions were asked for all conditions: • wie bewerten sie die interaktion des maklers in den folgenden dimensionen: how would you rate the broker’s interaction along the following dimensions: professionalität: unprofessionell (1) professionell (10) professionalism: unprofessional (1) professional (10) kommunikation: schlecht (1) sehr gut (10) communication: poor (1) very good (10) annehmlichkeit: unangenehm (1) angenehm (10) pleasantness: unpleasant (1) pleasant (10) förmlichkeit: locker (1) förmlich (10) formality: informal (1) formal (10) spontanität: gestellt (1) spontan (10) spontaneity: acted (1) spontaneous (10) • wie wahrscheinlich ist es, dass sie ihren bekannten die wohnung empfehlen würden? sehr unwahrscheinlich (1) sehr wahrscheinlich (10) how likely is it that you would recommend the apartment to your acquaintances? very unlikely (1) very likely (10) • wie wahrscheinlich ist es, dass sie ihren bekannten den makler empfehlen würden? sehr unwahrscheinlich (1) sehr wahrscheinlich (10) how likely is it that you would recommend the broker to your acquaintances? very unlikely (1) very likely (10) an additional question was asked in the three conditions involving a human agent: • wie haben sie das verhalten des maklers gegenüber der interessentin wahrgenommen? how did you perceive the broker’s behavior towards the client? unhöflich (1) höflich (10) rude (1) polite (10) abweisend (1) mitfühlend (10) dismissive (1) compassionate (10) two additional questions were asked in the two conditions involving a human agent that used laughter: • hatten sie den eindruck dass der makler immer an den richtigen stellen gelacht hat? immer falsch (1) immer richtig (10) did you have the impression that the broker always laughed in the right places? always wrong (1) always right (10) • wenn der makler lachte, wie klingte das für sie? unangemessen (1) angemessen (10) when the broker laughed, how did it sound to you? inappropriate (1) appropriate (10) 20 laughter use and task success references jo-anne bachorowski, moria smoski, and michael owren. the acoustic features of human laughter. the journal of the acoustical society of america, 110(3):1581–1597, 2001. doi: 10.1121/1.4778613. christian becker-asano and hiroshi ishiguro. laughter in social robotics-no laughing matter. in proceedings of the international workshop on social intelligence design, pages 287–300, 2009. url https://gki.informatik.uni-freiburg.de/papers/ becker-asano-ishiguro-sid2009.pdf. julie brosy, adrian bangerter, and joaquim sieber. laughter in the selection interview: impression management or honest signal? european journal of work and organizational psychology, 30 (2):319–328, 2021. doi: 10.1080/1359432x.2020.1794953. justine cassell, timothy bickmore, mark billinghurst, lee campbell, kenny chang, hannes vilhjálmsson, and hao yan. embodiment in conversational interfaces: rea. in proceedings of the sigchi conference on human factors in computing systems, pages 520–527, 1999. doi: 10.1145/302979.303150. kevin el haddad, stéphane dupont, jérôme urbain, and thierry dutoit. speech-laughs: an hmmbased approach for amused speech synthesis. in proceedings of the ieee international conference on acoustics, speech and signal processing (icassp), pages 4939–4943, 2015. doi: 10.1109/icassp.2015.7178910. kevin el haddad, hüseyin çakmak, emer gilmartin, stéphane dupont, and thierry dutoit. towards a listening agent: a system generating audiovisual laughs and smiles to show interest. in proceedings of the acm international conference on multimodal interaction, pages 248–255, 2016. doi: 10.1145/2993148.2993182. andrew gelman. scaling regression inputs by dividing by two standard deviations. statistics in medicine, 27(15):2865–2873, 2008. doi: 10.1002/sim.3107. phillip glenn. towards a social interactional approach to laughter. in laughter in interaction, chapter 1, pages 7–34. cambridge university press, 2003. doi: 10.1017/cbo9780511519888. 003. aike c horstmann and nicole c krämer. expectations vs. actual behavior of a social robot: an experimental investigation of the effects of a social robot’s interaction skill level and its expected future role on people’s evaluations. plos one, 15(8):e0238133, 2020. doi: 10.1371/journal.pone. 0238133. julian hough, ye tian, laura de ruiter, simon betz, spyros kousidis, david schlangen, and jonathan ginzburg. duel: a multi-lingual multimodal dialogue corpus for disfluency, exclamations and laughter. in proceedings of the language resources and evaluation conference, pages 1784–1788, 2016. url https://aclanthology.org/l16-1281. koji inoue, divesh lala, and tatsuya kawahara. can a robot laugh with you?: shared laughter generation for empathetic spoken dialogue. frontiers in robotics and ai, page 234, 2022. doi: 10.3389/frobt.2022.933261. 21 https://gki.informatik.uni-freiburg.de/papers/becker-asano-ishiguro-sid2009.pdf https://gki.informatik.uni-freiburg.de/papers/becker-asano-ishiguro-sid2009.pdf https://aclanthology.org/l16-1281 ludusan and wagner helena kangasharju and tuija nikko. emotions in organizations: joint laughter in workplace meetings. the journal of business communication, 46(1):100–119, 2009. doi: 10.1177/ 0021943608325750. nicole c krämer, astrid von der pütten, and sabrina eimler. human-agent and human-robot interaction theory: similarities to and differences from human-human interaction. human-computer interaction: the agency perspective, pages 215–240, 2012. doi: 10.1007/978-3-642-25691-2 9. laura e kurtz and sara b algoe. when sharing a laugh means sharing more: testing the role of shared laughter on short-term interpersonal consequences. journal of nonverbal behavior, 41: 45–65, 2017. doi: 10.1007/s10919-016-0245-9. kwan min lee, younbo jung, jaywoo kim, and sang ryong kim. are physically embodied social agents better than disembodied social agents?: the effects of physical embodiment, tactile interaction, and people’s loneliness in human–robot interaction. international journal of humancomputer studies, 64(10):962–973, 2006. doi: 10.1016/j.ijhcs.2006.05.002. bogdan ludusan and petra wagner. no laughing matter: an investigation into the acoustic cues marking the use of laughter. in proceedings of the 19th international congress of phonetic sciences, pages 2179–2182, 2019a. url http://www.assta.org/proceedings/ icphs2019/papers/icphs_2228.pdf. bogdan ludusan and petra wagner. laughter dynamics in dyadic conversations. in proceedings of interspeech, pages 524–528, 2019b. doi: 10.21437/interspeech.2019-1733. bogdan ludusan and petra wagner. knock-knock! who’s there? the laughter-enhanced virtual real-estate agent. in elektronische sprachsignalverarbeitung 2021. tagungsband der 32. konferenz, volume 99, pages 281–288, 2021. url https://www.essv.de/essv2021/ pdfs/16_ludusan.pdf. maurizio mancini, beatrice biancardi, florian pecune, giovanna varni, yu ding, catherine pelachaud, gualtiero volpe, and antonio camurri. implementing and evaluating a laughing virtual character. acm transactions on internet technology, 17(1):1–22, 2017. doi: 10.1145/2998571. vladislav maraev, chiara mazzocconi, christine howes, and jonathan ginzburg. integrating laughter into spoken dialogue systems: preliminary analysis and suggested programme. in proceedings of the workshop on artificial intelligence for multimodal human robot interaction, pages 9–14, 2018. doi: 10.21437/ai-mhri.2018-3. hannes matuschek, reinhold kliegl, shravan vasishth, harald baayen, and douglas bates. balancing type i error and power in linear mixed models. journal of memory and language, 94: 305–315, 2017. doi: 10.1016/j.jml.2017.01.001. hiroki mori, tomohiro nagata, and yoshiko arimoto. conversational and social laughter synthesis with wavenet. in proceedings of interspeech, pages 520–523, 2019. doi: 10.21437/ interspeech.2019-2131. 22 http://www.assta.org/proceedings/icphs2019/papers/icphs_2228.pdf http://www.assta.org/proceedings/icphs2019/papers/icphs_2228.pdf https://www.essv.de/essv2021/pdfs/16_ludusan.pdf https://www.essv.de/essv2021/pdfs/16_ludusan.pdf laughter use and task success tu anh nguyen, eugene kharitonov, jade copet, yossi adi, wei-ning hsu, ali elkahky, paden tomasello, robin algayres, benoit sagot, abdelrahman mohamed, et al. generative spoken dialogue language modeling. transactions of the association for computational linguistics, 11: 250–266, 2023. doi: 10.1162/tacl a 00545. radoslaw niewiadomski, jennifer hofmann, jérôme urbain, tracey platt, johannes wagner, bilal piot, huseyin cakmak, sathish pammi, tobias baur, stephane dupont, et al. laugh-aware virtual agent and its impact on user amusement. in proceedings of the international conference on autonomous agents and multiagent systems, pages 619–626, 2013. doi: 10.5555/2484920. 2485018. eva e nwokah, hui-chin hsu, patricia davies, and alan fogel. the integration of laughter and speech in vocal communication: a dynamic systems perspective. journal of speech, language, and hearing research, 42(4):880–894, 1999. doi: 10.1044/jslhr.4204.880. florian pecune, maurizio mancini, beatrice biancardi, giovanna varni, yu ding, catherine pelachaud, gualtiero volpe, and antonio camurri. laughing with a virtual agent. in proceedings of the international conference on autonomous agents and multiagent systems, pages 1817– 1818, 2015. url https://dl.acm.org/doi/pdf/10.5555/2772879.2773452. robert r provine. laughter. american scientist, 84(1):38–45, 1996. url http://www.jstor. org/stable/29775596. r core team. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria, 2020. url https://www.r-project.org/. hannes ritschel, ilhan aslan, silvan mertes, andreas seiderer, and elisabeth andré. personalized synthesis of intentional and emotional non-verbal sounds for social robots. in proceedings of the 8th international conference on affective computing and intelligent interaction, pages 1–7, 2019. doi: 10.1109/acii.2019.8925487. shane saunderson and goldie nejat. how robots influence humans: a survey of nonverbal communication in social human–robot interaction. international journal of social robotics, 11:575–608, 2019. doi: 10.1007/s12369-019-00523-0. kimberly sellers, thomas lotze, and andrew raim. package ‘compoissonreg’, 2022. url https://cran.r-project.org/web/packages/compoissonreg/index. html. r package version 0.8.0. gijsbert stoet. psytoolkit: a software package for programming psychological experiments using linux. behavior research methods, 42(4):1096–1104, 2010. doi: 10.3758/brm.42.4.1096. gijsbert stoet. psytoolkit: a novel web-based method for running online questionnaires and reaction-time experiments. teaching of psychology, 44(1):24–31, 2017. doi: 10.1177/ 0098628316677643. noé tits, kevin el haddad, and thierry dutoit. laughter synthesis: combining seq2seq modeling with transfer learning. proceedings of interspeech, pages 3401–3405, 2020. doi: 10.21437/ interspeech.2020-1423. 23 https://dl.acm.org/doi/pdf/10.5555/2772879.2773452 http://www.jstor.org/stable/29775596 http://www.jstor.org/stable/29775596 https://www.r-project.org/ https://cran.r-project.org/web/packages/compoissonreg/index.html https://cran.r-project.org/web/packages/compoissonreg/index.html ludusan and wagner jürgen trouvain and marc schröder. how (not) to add laughter to synthetic speech. in proceedings of the tutorial and research workshop on affective dialogue systems, pages 229–232, 2004. doi: 10.1007/978-3-540-24842-2 23. jürgen trouvain and benjamin weiss. thoughts on the usage of audible smiling in speech synthesis applications. frontiers in computer science, 4, 2022. doi: 10.3389/fcomp.2022.885657. bekir berker türker, zana buçinca, engin erzin, yücel yemez, and t metin sezgin. analysis of engagement and user experience with a laughter responsive social robot. in proceedings of interspeech, pages 844–848, 2017. doi: 10.21437/interspeech.2017-1395. brian j. zhang and naomi t. fitter. nonverbal sound in human-robot interaction: a systematic review. journal of human-robot interaction, 2023. doi: 10.1145/3583743. 24 introduction materials and methods stimuli procedure analyses results overall results agent type analyses detailed analysis of the human agent conditions discussion and conclusions dialogue & discourse 12(2) (2021) 1–37 doi: 10.5210/dad.2021.201 discourse relations and connectives in higher text structure lucie poláková polakova@ufal.mff.cuni.cz institute of formal and applied linguistics faculty of mathematics and physics, charles university jiří mírovský mirovsky@ufal.mff.cuni.cz institute of formal and applied linguistics faculty of mathematics and physics, charles university šárka zikánová zikanova@ufal.mff.cuni.cz institute of formal and applied linguistics faculty of mathematics and physics, charles university eva hajičová hajicova@ufal.mff.cuni.cz institute of formal and applied linguistics faculty of mathematics and physics, charles university editor: vera demberg submitted 12/2020; accepted 06/2021; published online 07/2021 abstract the present article investigates possibilities and limits of local (shallow) analysis of discourse coherence with respect to the phenomena of global coherence and higher composition of texts. we study corpora annotated with local discourse relations in czech and partly in english to try and find clues in the local annotation indicating a higher discourse structure. first, we classify patterns of subsequent or overlapping pairs of local relations, and hierarchies formed by nested local relations. special attention is then given to relations crossing paragraph boundaries and their semantic types, and to paragraph-initial discourse connectives. in the third part, we examine situations in which annotators incline to marking a large argument (larger than one sentence) of a discourse relation even with a minimality principle annotation rule in place. our analyses bring (i) new linguistic insights regarding coherence signals in local and higher contexts, e.g. detection and description of hierarchies of local discourse relations up to 5 levels in czech and english, description of distribution differences in semantic types in cross-paragraph and other settings, identification of czech connectives only typical for higher structures, or the detection of prevalence of large left-sided arguments in locally annotated data; (ii) as another type of contribution, some new reflections on methodologies of the approaches under scrutiny. keywords: local and global discourse coherence, prague discourse treebank 2.0, discourse connectives, hierarchies, paragraphs 1. introduction in coherence-oriented discourse studies, the recognition of the distinction between local and global coherence dates back to the 1980’s and one of the most compelling ways of its explanation is that “the main purpose of global coherence relations is to help eliminate locally coherent nonsense texts” (unger, 2006, samet and schank, 1984). according to them, global coherence is the connectivity ©2021 lucie poláková, jiří mírovský, šárka zikánová and eva hajičová this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). poláková, mírovský, zikánová and hajičová between the main events of the text (scripts, plans and goals) and the global relations hold independently of the local coherence relations between discourse segments. in automatic discourse processing, the local and the global coherence models are also known as shallow and deep discourse analyses/parsing, respectively (prasad et al., 2010). both types of approaches deal with determination and description of semantic and pragmatic relations between individual text units but they differ in the focus of the analyses and the methodology. local coherence models proceed bottom-up, from the smallest discourse units within a single sentence (clauses, even nominalized events, states etc.) and across the sentence boundary. emphasis is put on the description of semantico-pragmatic relations between every two consecutive units and on the identification of lexical cues anchoring these relations (mostly discourse connectives; where there is no such surface cue, the relation is called implicit). the coherence analysis in this way describes how every discourse unit is related to the previous one, there are no hierarchies postulated and no claims about the shape of the overall structure of the analyzed documents. the advantage of local discourse analyses is their easy applicability in annotation and shallow parsing (looking for surface cues, connectives and other markers), given the usability across different languages and relatively high reliability (e.g. via inter-annotator agreement). as far as language resources aiming at formalized linguistic description of discourse coherence are concerned, internationally, there is a range of corpora annotated for local coherence relations in different languages. they mostly follow the penn discourse treebank (pdtb) annotation style (prasad et al., 2008, prasad et al., 2019), compare 2.2. there exists even a multilingual corpus with local discourse annotations in the pdtb style of parallel texts in six languages (zeyrek et al., 2019).1 for czech, there is a publicly available large corpus of local discourse annotation, the prague discourse treebank 2.0 (pdit 2.0; rysová et al., 2016), see section 3. this resource has been developed by the prague discourse group since 2009 and thus naturally represents the point of departure for this study. global coherence models, on the contrary, proceed top-down, from the text as a communicative whole, and postulate a hierarchically interconnected structure of smaller and larger units. these models usually capture each document as a single continuous graph representation with specific properties and constraints on the relations (e.g. tree-like graphs). this concept has a strong potential in the possibility to demonstrate the composition of smaller blocks of the text, as well as to identify more general and more important text contents and relations between them. for global annotations, there are far less projects, compare section 2, and so far, there is no such project for czech data. however, global coherence modeling (often in addition to local coherence models), apart from its straightforward application in automatic coherence evaluation (feng et al., 2014, lin et al., 2011) can significantly contribute to other nlp tasks, such as summarization (zhang, 2011), topic identification (pons-porrata et al., 2007), text generation (kiddon et al., 2016), textual entailment (hagege and jacquet, 2014) and others. in this study, we examine the distributions, properties and mutual settings of local discourse relations in order to reveal possible patterns of higher discourse structuring and signs of global discourse coherence. our research question is: is global coherence signaled by the same types of language devices as the local one or can we reveal some differences by studying various phenomena in locally annotated data? the different types of local discourse relations and their mutual settings that we analyze in this study are represented in example 1 and figure 1 which outlines discourse annotation in the czech 1. the most exhaustive list of discourse-annotated corpora of any type was collected within the textlink initiative and can be accessed from this url: http://www.textlink.ii.metu.edu.tr/corpus-view 2 http://www.textlink.ii.metu.edu.tr/corpus-view discourse relations and connectives in higher text structure if the new limits could be prepared sooner they will probably only be issued as another amendment to the old law directive so it seems that the processors are in no hurry with new limits yet there is hunger for new limits condition connective: if range: 0->0 reason-result connective: so range: 0->0 concession connective: yet range: 0->0 s1 s2 s3 figure 1: an annotated example of an intra-sentential relation (condition), inter-sentential relation within a paragraph (reason-result) and a cross-paragraph relation (concession) from the prague discourse treebank 2.0. nodes represent individual clauses, root nodes refer to individual sentences, semantic types of discourse relations are highlighted in orange, connectives in green. pdit 2.0. the example demonstrates an intra-sentential relation, an inter-sentential relation within a paragraph and a cross-paragraph relation. also, a hierarchical structure (a nested relation, i.e. a relation fully included within an argument of another relation) is shown by this example text. (1) (s1) if the new limits could be prepared sooner, they will probably only be issued as another amendment to the old law directive. (s2) so it seems that the processors are in no hurry with new limits. [paragraph boundary] (s3) yet there is hunger for new limits. czech original: (s1) pokud by se podařilo nové limity připravit dříve, budou vydány zřejmě jen jako další novela směrnice starého zákona. (s2) a tak se zdá, že zpracovatelé s novými limity nespěchají. [paragraph boundary] (s3) přitom po nových limitech je hlad. within the sentence (s1), there is an intra-sentential relation of condition signaled by the connective pokud [if]. its two arguments are the two clauses in this sentence. the whole sentence (s1) is a left-sided argument of an inter-sentential reason-result relation anchored by a tak [so]. the if relation is a subset of so-relation, they are forming a hierarchical structure. these two sentences belong to a single paragraph and the two mentioned discourse relations are local, intra-paragraph relations. sentence (s3) starts a new paragraph and it is attached to the preceding sentence (s2) by a cross-paragraph inter-sentential concession relation with the connective přitom [yet]. in this case, the cross-paragraph link is local, between adjacent sentences, but in other possible contexts, the left-sided argument of the relation may span across several sentences or/and it can be non-adjacent. 3 poláková, mírovský, zikánová and hajičová 1.1 goals the main goal of the present study is to systematically analyze existing local discourse annotations of czech, and partly english, for possible signs of higher/global discourse structure. the analysis shall serve as springboard for a planned global coherence annotation of czech, similarly as the deepsyntactic layer was informative for the analysis of local coherence relations (mírovský et al., 2012). an underlying assumption is that even the local annotation already displays such features of global text structure. this assumption is based on our observations from locally annotated data so far: we detected patterns like hierarchical organization of smaller and larger discourse relations, connectives and other discourse cues operating between larger blocks of texts, long-distance relations, genrerelated patterns and so on. also, we hypothesize that in a well-formed, coherent text, the paragraph structure, as a main formal structuring device, must be mirrored in discourse relations and their semantics in some way. corpus analyses of possible features of higher/global text structure in this study concern three separate topics: • the issues of structure, or “the shape” of a text: we investigate mutual configurations of discourse relations (pairwise) and their complexity within locally annotated texts of prague discourse treebank 2.0. we identify and quantify the embeddings, overlappings and crossings of relations, according to a similar study conducted by lee et al. in 2006 on complexity of discourse structure in the penn discourse treebank 2.0. moreover, we look into hierarchies local discourse relations can form. this research was previously published in poláková and mírovský (2020). in this new version, we reflect on some feedback by our readers, explain our previous findings in more detail and by more examples, give original czech annotated examples where they were originally left out due to space limitations, and, most importantly, extend our analysis of hierarchies of nested relations by an analogous one for the english annotations in the pdtb 3. • the analysis of paragraph-initial connectives, relations they anchor and their respective discourse types (senses). this analysis is intended to reveal any possible differences between (the majority of) relations within individual paragraphs and those cross-paragraph relations in our data that have an explicit marking by a connective. • the analysis of large arguments, more precisely, of relations with one or both arguments larger than a single sentence, and their properties. the article is structured as follows. in the second part of the introduction, we specify the extent of our study and define types of discourse relations with respect to their scope in the text structure. related research is mentioned in section 2. section 3 describes the data format and application framework used in our study and presents the main data resources. section 4 with four subsections represents the core of the research and offers results of our study in four directions – (i) configurations of pairs of discourse relations, (ii) hierarchies of nested discourse relations, (iii) paragraphinitial senses and connectives, and (iv) relations with large arguments. section 5 summarizes our findings and offers some perspectives. examples of texts with deep hierarchies of discourse relations can be found in two appendices. 4 discourse relations and connectives in higher text structure 1.2 theoretical aspects and definitions for the descriptive purposes of this study, we need to terminologically distinguish several types of coherence. global coherence of a text refers to the coherence of a given document as a whole, including all inner structure of its coherence relations of smaller and greater units. this inner structure is assumed to form hierarchies of smaller and larger relations. on the contrary, local coherence refers to the “flat”, chain-like coherence of two minimal subsequent text units, as defined most elaborately in the pdtb 2 (prasad et al., 2008) approach mentioned above. further, we distinguish coherence on a higher level, or the so-called higher text structure, which describes coherence relations between/among larger text blocks, but not necessarily in a whole document. typically, a distinctive higher structure in a text is formed by (multi-sentential) lists and enumerations and it often relates to genre rules or practises (e.g. in legal texts, instructions, recipes). independently, we also distinguish interor cross-paragraph coherence, that means, coherence relations between individual paragraphs as integral units but also between smaller units belonging to two different paragraphs. coherence relations within individual paragraphs are then again referred to as local relations, and they can be either intraor inter-sentential, or belong to a higher structure, if made up of larger units (more sentences). obviously, these categories are not rigorous and mutually exclusive, as already the very purpose of this study – looking for global features in a locally annotated data – suggests. this gives rise to a question, e.g., how to approach the relation of the last sentence of one paragraph to the immediately following first sentence of the next paragraph, or, how to treat intra-sentential lists spanning across several paragraphs. we will address these and further related issues further in section 4. a note should be made to the use of the term hierarchy in this study. whereas we explore the hierarchical organization of discourse relations as wholes, there also is another understanding of hierarchy in discourse, in terms of the ways discourse units organize themselves based on criteria such as their relative importance (nuclearity in rhetorical structure theory). we do not address the latter here. 2. related research from the wide range of coherence theories which address a text/document as a structured whole and offer some type of a formalized representation of this whole (e.g. grosz and sidner, 1986, hobbs, 1979, asher and lascarides, 2003, wolf and gibson, 2005),2 our approach is methodologically closest to the penn discourse treebank and rhetorical structure theory (rst), which are introduced in more detail in the following sections. 2.1 rhetorical structure theory the rhetorical structure theory (rst, mann and thompson, 1988) is one of the most influential frameworks among the global coherence models. it was originally developed with the intention to model text coherence in order to study computer-based text generation. the main principle of rst is the assumption that coherent texts consist of minimal units, which are linked to each other, recursively, through rhetorical relations and that coherent texts do not exhibit gaps or non-sequiturs (taboada and mann, 2006). the rst represents the whole text document as a single (projective) 2. for a detailed classification of discourse approaches within the local/global space of coherence relations see bateman and rondhuis (1997). 5 poláková, mírovský, zikánová and hajičová tree structure. basic features of these structures are the rhetorical relations between two textual units (smaller or larger blocks that are in the vast majority of cases adjacent) and the notion of nuclearity. for the classification of rst rhetorical relations, a set of labels was developed, which originally contained 24 relations, but the authors themselves add that it is an open set “susceptible to extension and modification for the purposes of particular genres and cultural styles” (mann and thompson 1988, p. 250). the type of a rhetorical relation is defined with respect to the author’s intended effect on the reader together with the application of the principles of nuclearity. the rst has gained great attention, it was further developed and tested, language corpora were built with rst-like discourse annotation, such as for english the rst discourse treebank (rstdt, carlson et al., 2003) and its extension rst-signalling corpus (das et al., 2015), for german (stede and neumann, 2014), spanish (da cunha et al., 2011), basque (iruskieta et al., 2013) and also for some other languages. on the other hand, the framework was repeatedly criticized with regard to some of its theoretical claims, above all, concerning the question of adequacy/sufficiency of representation of a discourse structure as a tree graph.3 linguistically, the strong constraints of a tree (no crossing edges, one root, all the units interconnected etc.) gave rise to a search for counter-examples in real-world texts. it was shown that not only adjacent text units exhibit coherence links and that there are even cue phrases, which connect non-adjacent units and thus support the claim that a tree graph is too restricted a structure for an adequate discourse representation (wolf and gibson, 2005). therefore, more complex graphs with crossings and overlaps should be adequate, which resulted in the creation of the discourse graphbank (wolf et al., 2005), a resource of the same texts annotated with the main diverging principle of relaxing the tree-ness constraint. on these grounds and in the same way like lee et al. (2006), we try to demonstrate where the local and global analytic perspectives meet and interact. our analysis of czech data thus contributes new empirical material to the scientific debate on whether projective trees are descriptively adequate to represent structure of texts (e.g. marcu, 2000, wolf and gibson, 2005, lee et al., 2008, egg and redeker, 2010). additionally, there have been some rst-based studies investigating the distributions of rhetorical relations on different levels of text, e.g. williams and reiter (2003), liu and zhang (2016). their research questions are similar to ours in the section on paragraph-initial relations and cues 4.3. the latter study introduces a rst-tree conversion to dependency trees with different levels of granularity, with a node representing a clause, a sentence or a paragraph. they point out that most rhetorical relations in the rst-dt treebank occur on all levels but with different percentages and the two upper levels of discourse processes are more alike in many aspects, which are observations that also follow from our local data analyses (4.3). 2.2 penn discourse treebank the penn discourse treebank (pdtb) represents a lexically based, local model of discourse. its analysis of discourse relations consists primarily in finding and analyzing lexical cues of discourse coherence as “anchors” of discourse relations. such a cue, a discourse connective, is defined as a discourse-level predicate opening positions for two discourse arguments. discourse connectives include coordinating conjunctions (e.g. and, but), subordinating conjunctions (because, if ) and 3. basically a constituency tree, which is in its nature projective and does not allow crossing edges, in comparison to the basic mathematical definition of a tree graph. 6 discourse relations and connectives in higher text structure discourse adverbials (nevertheless). apart from connectives, the two discourse arguments of a discourse relation (and their extent) and the semantic type (sense) of a discourse annotation were annotated. in 2004, the first version of penn discourse treebank was released (miltsakaki et al., 2004). the second release of the pdtb four years later includes annotation of the ca. 49,000 sentences of the wall street journal part of the penn treebank (prasad et al., 2008). apart from explicit connectives, other phenomena have been annotated in this version, mainly implicit relations and attribution (attributing beliefs and assertions to agents making them) alternative lexicalization of a connective (altlex, e. g. that is why), entity based relations (entrel) and places with weak coherence (norel). in its latest version pdtb 3.0 (prasad et al., 2018, prasad et al., 2019), the annotations were enriched by many relations mostly in intra-sentential contexts, the sense taxonomy was revised and the existing annotations enhanced in many aspects. the pdtb-style connective/argument analysis has become very popular, also because such an analysis requires less interpretation and pragmatic inference than the rst analysis. the pdtb authors also claim that their approach is theory-neutral, independent from any syntactic theory, and as such can be transferred to other languages. in the first part of our analysis on configurations of positionally close relation pairs (4.1), we relate our findings to a similar research conducted on the second version of the english penn discourse treebank in lee et al. (2006). the aim of their study was to describe and quantify the configurations of discourse relations as typical or less typical in terms of discourse complexity. the complexity of discourse configurations there is being compared to the complexity of relations in syntax, but also refers to the principles of the global rhetorical structure theory, in particular the representation of any text document as a single tree-like structure with strong constraints. the study actually tries to answer the question whether their empirical locally annotated data would fulfill or violate these strong constraints. therefore lee et al. (2006) studied various types of overlaps of discourse relations. they encountered a variety of patterns between pairs of discourse relations, including nested (hierarchical), crossed and other non-tree-like configurations. nevertheless, they conclude that the types of discourse dependencies are highly restricted since the more complex cases like crossings and partial overlaps can be factored out by appealing to discourse notions like anaphora (non-structural, possibly long-distance tie) and attribution (attribution span as a known mismatch between the syntactic and discourse structure), and argue that pure crossing dependencies, partially overlapping arguments and a subset of structures containing a properly contained argument should not be considered part of the discourse structure. the authors challenge czech discourse researchers to introduce a similar study (footnote 1 in their paper, ibid.) in order to observe and compare the complexity of discourse and syntax dependencies in two typologically different languages.4 3. data and tools our study is based primarily on two data resources: for czech, the prague discourse treebank 2.0 (pdit; rysová et al., 2016, zikánová et al., 2015),5 and, for english, the penn discourse tree4. from this viewpoint, we do not assume substantial differences between the two languages in discourse structure itself. however, there are surely larger differences on “lower” levels of linguistic description, in our case most visibly connective repertoires and syntactic properties of the two languages, which can in the outcome influence some motivations in the annotation design and thus effect the resulting discourse structure, compare mírovský et al. (in prep.). 5. please note that the discourse annotation in pdit 2.0 is the same as in pdt 3.5 (hajič et al., 2018) and the upcoming pdt-c 1.0, which are newer versions of the pdit 2.0 data. 7 poláková, mírovský, zikánová and hajičová contrast expansion confrontation conjunction opposition conjunctive alternative restrictive opposition disjunctive alternative pragmatic contrast instantiation concession specification correction equivalence gradation generalization contingency temporal reason–result synchrony pragmatic reason–result precedence–succession explication condition pragmatic condition purpose table 1: semantic types of discourse relations in the pdit 2.0 bank ver. 3.0 (pdtb; prasad et al., 2019). these two corpora are comparable in size (approx. 50 thousand sentences each), genre (journalistic texts) and they have both been manually annotated for discourse relations in comparable approaches (pdtb and the pdtb-like modification in the pdit). namely, for explicit relations, all connectives in a given text were identified, their two arguments were detected and a semantic/pragmatic relation6 between them was assigned to the relation. discourse connectives in pdit 2.0 are of two types: primary and secondary, according to rysová and rysová (2014). primary connectives are defined as grammaticalized expressions such as because or therefore whereas secondary connectives are not (yet) fully grammaticalized structures such as except for this, the reason was or for this reason. the notion of secondary connectives in prague treebank roughly corresponds to the categories of altlex and altlexc (alternative lexicalization tokens and alternative lexicalization constructions) in pdtb 3. the pdtb and pdit use similar sets of discourse relation types – the set of discourse types in pdit is inspired by the penn discourse treebank 2.0 sense hierarchy (prasad et al., 2008), see the complete list in table 1. the taxonomy is two-level, but in the annotation only the second-level relations can be assigned. confrontation in the czech taxonomy equals to the original juxtaposition in the pdtb 2 (mostly represented by connectives like while, whereas), restrictive opposition also includes exception (only, except that...), correction equals to substitution (not a, but b, instead), explication is giving evidence by other means than reason-result.7 pragmatic labels are used in accordance with the pdtb 2 pragmatic labels. 6. "sense" in the pdtb terminology, "discourse type" in pdit 2.0 7. strongly represented by a typical czech connective totiž (actually, since, i mean, as a matter of fact) 8 discourse relations and connectives in higher text structure pdit 2.0 contains 21 223 explicit discourse relations (i.e. relations signalled by explicit discourse connectives, both primary and secondary),8 while the pdtb 3 contains 25 696 such relations (we take into account relations of type explicit, (altlex) and altlexc).9 both pdit and the pdtb are available in the prague markup language format (pml; pajas and štěpánek, 2008),10 an xml-based format and application framework designed for multilayer linguistic annotations with available tools allowing for complex linguistic studies over pml data: btred for scripting in perl and prague markup language tree query (pml-tq; pajas and štěpánek, 2009) as a graphically oriented querying system. 4. analysis this section is divided into four subsections according to the research directions in which we address different aspects of a higher/global text structure traceable in locally annotated corpora. 4.1 – configurations of pairs of discourse relations, a study carried out on the basis of a similar research by the pdtb group and, like this study, contributing to the debate on acceptable “shapes” of an interconnected coherence representation of a text; 4.2 – hierarchies of nested discourse relations – a study that reveals if local discourse relations, without having the underlying assumption of hierarchy, also exhibit some type of hierarchical structure; 4.3 – paragraph-initial senses and connectives, a survey on cross-paragraph coherence in comparison to intra-paragraph settings; and 4.4 – relations with large arguments, as these are rather unexpected/unusual in local settings and can denote possible non-elementary units of higher text structuring.11 4.1 configurations of close relation pairs for this part of the analysis, and also for the subsequent one described in 4.2, we take advantage of the fact that the corpora used in the pdtb and our studies are comparable in many aspects (as mentioned earlier) and also contain a similar number of annotated discourse relations (approx. 20 thousand). however, it is important to notice that lee et al. (2006) included also implicit relations in their study and conducted the research on the older version of the pdtb (2.0). the lee et al. (2006) study defines six basic types of patterns of relation pairs: independent relations, full embeddings, shared arguments, properly contained arguments, pure crossings and partially overlapping arguments. in our analysis, we decided to make a more fine-grained classification of the patterns to cover all possible settings and to get more detailed insight into the configurations, and ended up with 17 categories, compare table 2. sometimes we also use different pattern naming (e.g. independent relations vs. adjacency), but for those patterns accounted for in both studies, we state how the names map. in our data, we have collected and classified types of relative positions of neighbouring discourse relations in the corpus, using a scripting tool btred as a part of an application framework for pdit8. we exclude the 361 list relations from this study, i.e. relations between subsequent members of enumerative structures. they represent a special text structure on their own, and as such, in the prague approach they stand aside the binary discourse relations and discourse type taxonomy. 9. implicit discourse relations have been annotated on a part of the pdit 2.0, see zikánová et al. (2019); this annotation was not taken into account in the present analysis, as the size of the implicit datasets is different at present. 10. the pml is the primary publication format for pdit. the english pdtb was first transformed to the pml format, for the purposes of studying discourse annotations in a unified format. 11. the first and the second subsections/topics are closely related, as hierarchies are a subset of possible relation configurations (nested relations pattern), but we keep them separate in the study. 9 poláková, mírovský, zikánová and hajičová line frequency in pdit pattern visualization 1 6 572 (983) adjacency <---------> <---------> 37.3% (32.9%) <=========> <=========> 2 1 923 (865) progress <---------> <---------> 10.9% (29.0%) (shared argument) <=========> <=========> 3 51 (21) total overlap <---------> <---------> 0.3% (0.7%) <=========> <=========> 4 266 (116) left overlap <---------> <---------> 1.5% (3.9%) right adjacency <=========> <=========> 5 25 (10) left overlap <---------> <---------> 0.1% (0.3%) right contained <=========> <===> 6 31 (10) right overlap <---------> <---------> 0.2% (0.3%) left adjacency <=========> <=========> 7 17 (8) right overlap <---------> <---------> 0.1% (0.3%) left contained <===> <=========> 8 250 (146) containment i <---------> <---------> 1.4% (4.9%) <===> <=========> 9 190 (150) containment ii <---------> <---> 1.1% (5.0%) (opposite) <=========> <=========> 10 83 (14) both args <---------> <---------> 0.5% (0.5%) contained i <===> <===> 11 73 (8) both args <---------> <---> 0.4% (0.3%) contained ii <===> <=========> 12 3 591 (299) left hierarchy <---------> <---------> 20.4% (10.0%) <===> <===> 13 3 543 (56) right hierarchy <---------> <---------> 20.1% (1.9%) <===> <===> 14 10 (10) crossing <---------> <---------> 0.1% (0.3%) <=========> <=========> 15 883 (196) envelopment <---------> <---------> 5.0% (6.6%) <=========> <=========> 16 3 (2) partial overlap <---------> <---------> 0.0% (0.1%) <=========> <=========> 17 8 (5) partial overlap ii <---------> <---------> 0.0% (0.2%) <===============> <=========> table 2: patterns of adjacent, overlapping or embedded pairs of explicit discourse relations in pdit 2.0. the column visualization shows relative position of the arguments read from left to right; arguments of one relation are marked with ’-’, the other one with ’=’. the total number of such close relations in pdit was 17 628; 109 of them (0.6%) did not fit the listed patterns. numbers in brackets mean frequencies if only inter-sentential discourse relations (arguments in different sentences) were taken into account; in total, there were 2 984 such close relations in pdit, 85 of them (2.8%) did not fit the listed patterns. types of corpora (pajas and štěpánek, 2008). we define a close relation pair as a pair of discourse relations that are either adjacent (the left argument of one relation immediately follows the right 10 discourse relations and connectives in higher text structure argument of the other relation), or they overlap in one of many possible patterns. we first analyze the detected patterns in pdit and, next, a specific subsection is devoted to the description of the detected hierarchical structures (4.2). we were able to detect 17 628 close relation pairs in pdit, and for each such pair, we investigated its pattern, the mutual arrangement of the two relations. in table 2, these patterns are also graphically illustrated. the table shows figures for all explicit relation pairs and, in brackets, for inter-sentential relation pairs only. the most common settings for close relations in pdit are pure adjacency (succession) of two relations (line 1: 6 572 cases in pdit), and “full embeddings”, in other words two-level hierarchies, in total 7 134 in pdit (lines 12 and 13). these configurations (lines 1, 12 and 13) represent together slightly more than 3/4 of all detected patterns in pdit. they are also referred to as very “normal” structural relationships in lee et al. (2006, p. 82). the next-largest group is progress (line 2), a shared argument in the pdtb terminology, with 1 923 instances or 10.9% of all patterns. lee et. al. report 7.5% of this type, which is fairly comparable. total overlap (line 3) is caused by the possibility to annotate two different relations between the same segments for co-occurring connectives, as in because for example or but later. it occurs 51 times (0.3%) in pdit. this pattern may vary a lot in different annotation schemes, as there may be different approaches to handling co-occurring connectives, or even the possibility of a a co-occurrence of an explicit and an implicit relations between the same segments, which is the approach taken by some newer corpora, e.g. by the pdtb in its version 3. the envelopment (line 15) concerns in the vast majority of cases a non-adjacent (long-distance) relation and another relation placed in the text between the arguments of the non-adjacent relation. the enveloped relation is often a sentence with an inner syntactic structure annotated. the same is often true for hierarchical patterns (lines 12 and 13) and it explains the big difference observed in both corpora between all detected envelopment and hierarchical patterns and only the inter-sentential ones. linguistically, some of the envelopment cases are sentences headed by two attribution spans (verbs of saying) and some structure in the reported content in between, also cases of two linked reporter’s questions in an interview and the inner structuring of the interviewee’s answer, but also texts with no striking structural reasons for such an arrangement. in pdit, envelopment represents ca. 5% of all settings and is certainly worth of further investigation. patterns with properly contained arguments, either one of the arguments (lines 5, 7, 8 and 9) or both (10 and 11), very often involve “skipping one level” in the syntactic tree of a sentence, see example 212 of the type 8 (containment i), i.e. exclusion of a governing clause from the argument (mostly an attribution span), that makes its syntactically dependent clause (mostly a “reportedcontent-argument”) to a subset of an other argument of the other relation represented (mostly) by a whole sentence. in example 2, the text span a young researcher works with enthusiasm for science, regardless of salary is a left-sided argument of the but-relation (in bold) and, at the same time a subset, properly contained, in a larger right-sided argument of the therefore-relation (in italics). the governing clause it is therefore appropriate... the fact that is thus the mismatch, it represents the difference between proper containment and the much more frequent shared argument pattern (line 2). 12. from now on in the annotated examples, relation 1 is highlighted in italics, relation 2 in bold. the connectives are underlined. 11 poláková, mírovský, zikánová and hajičová (2) the gap in the standard of living that appears between the qualified scientific elite and the business sphere, right now, at the beginning of the transformation of the society, will leave traces. it is therefore appropriate to pamper young researchers and not misuse the fact that a young researcher works with enthusiasm for science, regardless of salary. but a person who begins to find his mission in research also starts a family, wants to live at a good place and live with dignity. czech original: mezera na úrovni životního standardu, která se objevuje mezi kvalifikovanou vědeckou špičkou a sférou podnikatelskou, právě ted’, v počátcích transformace společnosti, přece jen zanechá stopy. je tedy namístě hýčkat dorost a nehřešit na to, že mladý badatel pracuje s nadšením pro vědu, bez ohledu na plat. jenže člověk, který začíná nalézat své poslání ve výzkumné práci, také zakládá rodinu, chce bydlet a důstojně žít. besides the discussed attribution (introductory statements) that has been excluded from the argument, this mismatch in the shared argument extent is also the case of the annotation of some secondary connectives realized by whole clauses (like this means that..., example 3, type 9, containment ii). these verb phrases are not treated as parts of any of the arguments they relate to and pose a methodological issue. in the representation in example 3, the clause this means that is not in italics, it is not considered to be a part of any argument of the left relation, i. e. the relation it anchors. (3) this brief overview essentially exhausts the areas of notarial activities within the framework of free competition between notaries. this means that in these notarial agendas, the client has the option of unlimited choice of notary at his own discretion, as the notary is not bound to the place of his work when providing these services. cz: tímto stručným přehledem jsou v podstatě vyčerpány oblasti notářských činností v rámci volné konkurence mezi notáři. to znamená, že v těchto notářských agendách má klient možnost neomezeného výběru notáře podle vlastní úvahy, nebot’ notář při poskytování těchto služeb není vázán na místo svého působení. a third setting concerns multi-sentence arguments, where the contained argument is typically a single sentence. patterns with properly contained arguments (lines 5, 7 – 11) represent in total 3.6% (638) of all patterns in pdit. a (pure) crossing is a setting where the left-sided argument of the right relation comes in between the two arguments of the left relation, compare line 14 in table 2. pure crossings violate the rst constraints most visibly, with crossing edges, so the debate on tree adequacy often circles around the acceptability of crossings in discourse analysis. lee et al. (2006) identify 24 cases (0.12%) of crossings in the pdtb2. we detected only 10 such cases in pdit, which is a very small proportion. manual inspection of the cases of crossing revealed several different scenarios, from clearly incorrect annotation, more interpretations possible, across cases with attribution spans in between, to a few, in our opinion, perfectly sound analyses, as exemplified by example 4 from pdit. (4) (a) what can owners and tenants expect, what should they prepare for? (b) the new legislation should allow all owners to sell apartments. (c) it is most urgent for flats owned by municipalities, as they manage about a quarter of the housing stock of the czech republic and some are mainly for financial reasons interested in mon12 discourse relations and connectives in higher text structure etizing a part of their apartments. (d) the law should also allow to complete transfers of housing association apartments and their sale to its members. (e) this is also not a “small portion”, but a fifth of the total number of dwellings.13 cz: (a) s čím mohou vlastníci i nájemci počítat, na co by se měli připravit? (b) nová právní úprava by měla umožnit všem vlastníkům prodávat byty. (c) nejnaléhavěji to vystupuje do popředí u bytů ve vlastnictví obcí, protože ty spravují asi čtvrtinu bytového fondu české republiky a některé mají především z finančních důvodů zájem část svých bytů zpeněžit. (d) zákon by měl také umožnit dokončení převodu i možnost prodeje družstevních bytů členům. (e) ani zde nejde o “malou porci”, ale o pětinu celkového počtu bytů. if we accepted the possibility that not only (b), but a larger (b+c) unit relates to (d) in the alsorelation, which would be a completely fine interpretation in the prague annotation, the relation of (e) – the neither-relation – in our view still cannot accept just (d) as its left-sided argument. we also think this case cannot be factored out due to anaphora. there is, for sure, room for different interpretations within different theories, we just offer our data, state our view and admit that crossing structures are extremely rare even in our empirical data. partial overlap is a type of structure that violates the rst tree constraint, too. lee et al. (2006) could only find 4 such cases. we detected 11 cases in pdit (lines 16 and 17 of table 2). they often include large arguments of untypical range (2.5 sentences etc.) which can be questioned. some of the relations also include secondary connectives with strong anaphoric links (in this respect, given the fact that etc.). these relations can be factored out, yet, again, even among the small number of cases in pdit there were linguistically acceptable ones, compare example 5 (partial overlap ii). (5) the responsibility of the future tenant of this 103,000-square-metre area will be to care for all properties, including their maintenance and repairs. the tenant will also have to resolve the parking conditions for market visitors and to meet the conditions of the prague heritage institute during construction changes due to the fact that the complex is a cultural monument. the capital city at the same time envisages preserving the character of the holešovice market. cz: povinností budoucího nájemce tohoto areálu o rozloze 103 tisíc metrů čtverečních bude mj. péče o všechny nemovitosti včetně jejich údržby a oprav. nájemce bude také muset vyřešit podmínky parkování pro návštěvníky tržnice a splnit podmínky pražského ústavu památkové péče při úpravách objektů vzhledem k tomu, že jde o kulturní památku. hlavní město přitom počítá se zachováním charakteru holešovické tržnice. 4.2 hierarchies the purpose of looking for hierarchical structures in the locally annotated data is to discover to what extent such an annotation shows signs of some higher (global) structure, too. we do not claim that the trees detected by us are the trees a global analysis like rst would discover, but we demonstrate the existence of some hierarchical text structure in local annotation. some of it could perhaps partially match to rst-formed subtrees (and definitely there would be an intersection of separate relations, compare e.g. intersections in wall street journal local and global annotations in 13. the “also-not” connective is originally in czech ani, in the meaning of neither. lit. translation: “neither here is_concerned a small portion...” 13 poláková, mírovský, zikánová and hajičová line frequency depth hierarchy 1 381 (10) 3 a ( b ( c )) 2 64 (1) 3 a ( b ( c ) d ) 3 50 3 a ( b c ( d )) 4 23 (1) 3 a ( b ( c ) d ( e )) 5 20 3 a ( b ( c d )) 6 0 (1) 3 a ( b ( c ) d e ) 7 0 (1) 3 a ( b c ( d ) e ) 8 19 4 a ( b ( c ( d ))) 9 7 4 a ( b ( c ( d )) e ) 10 5 4 a ( b c ( d ( e ))) 11 5 4 a ( b ( c ( d )) e ( f ( g ))) ... 12 2 5 a ( b ( c ( d ( e ))) f ) 13 1 5 a ( b c ( d ( e ( f ))) g ) 14 1 5 a ( b ( c ( d ) e ( f ( g ))) h i j ( k )) 15 1 5 a ( b ( c ( d ( e )))) 16 1 5 a ( b ( c ( d e ( f )))) table 3: selected schemes of hierarchies of explicit discourse relations in pdit. numbers in brackets mean frequencies of hierarchy schemes if only inter-sentential discourse relations were taken into account (no other inter-sentential hierarchies were encountered in the data). the three dots in the middle of the table indicate that there can be more cases of different 4-level hierarchy patterns. poláková et al., 2017), but this is yet to be investigated. we are also aware, as pointed out in egg and redeker (2010), that minimal, local annotations do not normally form a connected graph.14 4.2.1 hierarchies in the pdit 2.0 for the study of hierarchical structures, we used pairs of nested relations where one of the relations is as a whole included in one argument of the other relation (patterns corresponding to lines 12 and 13 from table 2) to recursively construct tree structures out of pairs of the nested relations. based on the quantitative results, we inspected selected samples of the detected patterns manually, in order to check the script outcome and to provide a linguistic description and comparison. the results on the whole pdit data are displayed in table 3, arranged according to the scheme of such hierarchy trees (identical structures are summed and represented by the hierarchy scheme). we only mention cases where there are at least three levels in the tree, as two-level hierarchies are part of table 2. for explanation: the scheme “a ( b )” means that the whole relation b is included in one of the arguments of the relation a (this is, of course, only a two-level tree). the scheme “a ( b c ( d ))” 14. and the more so, as we do not include implicit and entity-based relations into our study. 14 discourse relations and connectives in higher text structure means that relations b and c are all included in the individual arguments of relation a (without specifying in which argument they are, so they can be both in one argument or each in a different argument) and the relation d is completely included in one of the arguments of relation c. it is a three-level hierarchy. generally, we count the depth (number of levels) of a hierarchy tree as a number of nodes in the longest path from the root to a leaf. there are many sub-hierarchies in a large/deep hierarchy, for example “b ( c ( d ))” is a subhierarchy of “a ( b ( c ( d )) e )”; such sub-hierarchies, however, are not counted in table 3, i.e., each hierarchy is only counted in the table in its largest and deepest form as it appeared in the pdit 2.0 data.15 in the pdit data, local discourse relations form hierarchies up to five levels. we have identified 5 patterns of 5-level hierarchies (5-lh), with the total of 6 instances, see table 3. there is also a number of 3and 4-level hierarchies. an analysis of random samples (and of all the deepest ones) revealed, surprisingly, that there may be a 4-level hierarchy spanning 11 sentences, but also a 5-lh spanning only two sentences, from which one is typically a more complex compound sentence. the “longest” of the 5-lhs includes also 11 sentences (line 14) and it also exhibits branching (d, g and k as leaves, where the g-path is the deepest). one of the 5-lhs should be in fact one level flatter (line 12), as the lowest two relations are three coordinated clauses with two and-connectives: “the troops protected them and fed them and gave them the impression that they were invulnerable...”. such structures are notoriously hard to interpret for any framework, yet in prague annotation, the annotation is incorrectly hierarchical where it should have been flat. for a better illustration of the hierarchical configuration of the relations, a text sample with a 4lh is analyzed in appendix 1 to this study. it covers one whole paragraph and a part of a preceding one (11 sentences). its pattern is a ( b ( c ( d ) e ) f g ).16 to find out how much structure is involved only within individual sentences, i.e. how much of sentential syntax forms the hierarchies, in a second phase we filtered out all intra-sentential relations. the numbers in brackets give counts for patterns of hierarchies, if only inter-sentential relations (i.e. arguments in different sentences) are accounted for.17 the hierarchies of this type are much less frequent and their maximum depth is just 3, which implies that beyond the sentence boundary, local annotation of explicit connectives does not represent hierarchical text structuring very often in the pdit 2.0. 4.2.2 hierarchies in the penn discourse treebank 3.0 having at our disposal the penn discourse treebank 3.0 annotations converted into the prague pml format (poláková et al., 2017), we can use the same procedure to search for hierarchical structures also in these english locally annotated data. we take into account relations of the type explicit, altlex and altlexc. the results are summed up in table 4. similarly as in the experiment with the czech data, we have detected hierarchies of relations up to 5 depth levels. overall, their counts are 15. this also explains the zeros in table 3. the empty line in the table suggests that there are more different patterns of hierarchies of the given depth, the same holds also for table 4. 16. we do not present a 5-lh here, as the largest one is too complex and the smaller ones include two sentences only, so the main structure is syntactic in nature. 17. please note that hierarchies counted in brackets (without intra-sentential relations, column 3 in tables 3 and 4) are not a subset of the respective hierarchies from the same table row that also include intra-sentential relations – adding intra-sentential relations into the hierarchies means that a particular hierarchy pattern may change (be enlarged). this happens most clearly in lines 6 and 7 of table 3. 15 poláková, mírovský, zikánová and hajičová line frequency depth hierarchy 1 232 (2) 3 a ( b ( c )) 2 63 (2) 3 a ( b ( c ) d ) 3 46 (2) 3 a ( b c ( d )) 4 20 3 a ( b ( c d )) 5 10 3 a ( b ( c ) d ( e )) 6 8 3 a ( b ( c ) d e ) . . . 7 2 (1) 3 a ( b ( c ) d e f ) . . . 8 5 4 a ( b ( c ( d )) e ) 9 3 4 a ( b ( c ( d ))) 10 2 4 a ( b ( c ( d ) e ) f ) 11 1 4 a ( b c d e f ( g ( h ))) 12 1 4 a ( b ( c d ( e ) f ) g ( h ) i ( j ) k ( l )) . . . 13 1 5 a ( b c d ( e ( f ( g )))) table 4: selected schemes of hierarchies of explicit discourse relations in the pdtb 3.0. numbers in brackets mean frequencies of hierarchy schemes if only inter-sentential discourse relations were taken into account (no other inter-sentential at least 3-level hierarchies were encountered in the data). slightly lower, e.g. the simplest 3-lh pattern (a (b (c)) is 232 in comparison to 381 in the czech data. the maximum depth is also 5 but there is only one such detected hierarchy, compare the last line of table 4. the source text with all the relations annotated within this hierarchy is presented in detail in the appendix 2 to this study. the hierarchy spans across 2 paragraphs (7 sentences), and its pattern is a ( b c d ( e ( f ( g )))), which implies that the relations b and c do not take part in the deepest branch of the hierarchy, they are only included in the highest relation a. further, it can be observed that the two lowest relations, g (the deepest one) and f, are intra-sentential, while the hierarchy crosses the sentence boundary with its relation e, which is the lowest inter-sentential relation. so, only the relations a and d, the arguments of which span across two or more sentences18 in our understanding contribute in a way to a higher discourse structure. the number of hierarchies formed by inter-sentential relations only is even lower in the pdtb 3, only 7 (compared to 14 in the pdit 2.0), with the same maximum depth of 3 levels. a hypothesis for the relatively small number of hierarchies built by only inter-sentential relations in both corpora is that only some of the connectives operating at higher discourse levels were identified and annotated as such, some of them were assigned local coherence links due to the minimality principle.19 18. more precisely, one argument in each of them spans two or more sentences 19. the minimality principle instructs the annotators to mark as an argument as many clauses and/or sentences as are minimally required and sufficient for the interpretation of the relation. it was applied both in pdtb and in pdit 16 discourse relations and connectives in higher text structure number of documents 3 165 number of sentences 49 431 number of sentences without headings, captions and metatext 44 979 number of paragraphs 14 790 number of paragraphs without headings etc. 11 643 the average number of sentences per paragraph21 3.9 table 5: basic properties of the inner structure of texts in the prague discourse treebank this issue was recently discussed in poláková and mírovský (2019) with focus on paragraph-initial connectives and we dedicate the following subsection to this topic. 4.3 paragraph-initial semantic types and connectives a possible different scope of connectives in paragraph-initial positions was previously indicated in a study analyzing discourse connectives with anaphoric properties and their ability to relate to a distant, non-adjacent left-sided argument (poláková and mírovský, 2019). a special set of cases was defined where the connective actually does not relate to a short non-adjacent segment on its left, but to an adjacent, quite larger segment of text and is interpretable as a means of higher discourse structuring. from another perspective on paragraph boundary, cross-paragraph discourse relations were recently investigated in prasad et al. (2017) for implicit relations, leading to the observation that a first sentence in a given paragraph semantically relates to the immediately preceding last sentence of the previous paragraph in only 52% of their sample, with 48% having links to nonadjacent left contexts. our analysis of the features of cross-paragraph coherence only takes explicit relations into account (including secondary connectives), although we acknowledge that implicit connecting or other signalling is common between paragraphs. we focus on the following issues: (i) is there a difference between the semantics of discourse relations that ensure continuity between paragraphs on one side, and relations within an individual paragraph on the other? are certain relations more typical for cross-paragraph coherence, while others are typically local? (ii) are discourse connectives in the cross-paragraph usage different from locally used connectives? because of limitations of the actual version of the conversion of the pdtb data to the pml format, in this subsection and also in the study of paragraph-initial connectives in the subsequent subsection we only examined texts of the prague discourse treebank 2.0 (as characterized in table 5), via pml-tq queries and using information about paragraph numbers available at roots of the deep-syntactic trees.20 in terms of the higher structure and local relations in a text, three sets of relations can be distinguished. first of them are the relations between paragraphs, or cross-paragraph relations, as the top set (1). the rest, i.e. intra-paragraph relations, can be divided into inter-sentential (2) and intraannotations. the pdtb moreover annotates supplementary material to an argument, where needed (prasad et al., 2008). 20. information about paragraph numbers in a direct form of attribute para_no is only available in the pdit data accessible via the pml-tq search engine and is taken from identifiers (attribute id) of t-roots – roots of the tectogrammatical trees. 21. the average length of a paragraph is counted after the exclusion of headings, captions and metatext, i.e. 44 979 / 11 643. 17 poláková, mírovský, zikánová and hajičová (1) cross-par inter-s (2) intra-par inter-s (3) intra-s conjunction 236 opposition 1 584 conjunction 6 154 opposition 207 conjunction 1 322 reason-result 1 663 reason-result 125 reason-result 1 056 condition 1 400 confrontation 44 confrontation 273 opposition 1 388 generalization 43 concession 252 concession 633 specification 42 precedence 240 precedence 592 concession 32 gradation 197 specification 521 instantiation 22 explication 155 purpose 412 precedence 20 restr. opposition 153 confrontation 348 restr. opposition 19 correction 127 correction 323 table 6: 10 most common semantic types of explicit discourse relations for each dataset (number of occurrences) sentential ones (3); the differences between these groups are mainly due to the syntactic structure. in our analysis, we want to focus specifically on the relations between paragraphs (1). these relations do not reflect syntactic structure that much, which is why we compare them mainly to the separate group of intra-paragraph inter-sentential relations (2). however, in order to keep the picture of discourse relations complete, we also include a comparison with (intra-paragraph) intra-sentential relations (3). thus, we focused on three datasets. for the cross-paragraph relations (dataset 1), we analyzed explicit discourse relations which connect the first sentence of the paragraph to any previous text. as for local relations (2), the set in question concerns explicit inter-sentential discourse relations that do not go beyond the scope of one paragraph. the relations in the set (3) connect arguments within one syntactic tree (sentence), i.e. they cross neither the sentence boundary, nor the paragraph boundary.22 the distribution of semantic types of discourse relations across all the datasets differs significantly, see table 6 and the graph in figure 2, the significance was verified using the χ2 test. the vast majority of instances (63-68 percentage points in all three datasets) belong to three discourse relations, namely conjunction, opposition, and reason-result. in cross-paragraph relations, conjunction comes first, whereas in intra-paragraph inter-sentential relations, opposition is most common. within intra-sentential relations, conjunction absolutely predominates with 42% of occurrences, and the very frequent relation of condition comes as third, compare details below. in the top ten of the cross-paragraph relations, some relations typically occur which are not in the top ten of the intra-paragraph relations, namely specification, generalization, and instantiation. often, the following paragraphs in our data expand the content of the previous text in this way. on the other hand, we cannot consider the relation of specification as only typical in the higher 22. theoretically, there could be cross-paragraph intra-sentential discourse relations, too, such as lists of dependent clauses printed each in a different paragraph, cf. the request will be accepted if the applicant comes from the eu, and if he/she encloses two letters of recommendation. we do not deal with such sentences in this paper since their occurrence is marginal. 18 discourse relations and connectives in higher text structure 0% 5% 10% 15% 20% 25% 30% 35% 40% 45% cross-par inter-s intra-par inter-s intra-s figure 2: a graph for table 6 showing percentages of the semantic types for the three datasets, i.e. for cross-paragraph (first column), intra-paragraph inter-sentential (second column), and intra-sentential discourse relations (third column). in each dataset, all its occurrences sum to 100%. text structure, since it also occurs in the top 10 intra-sentential relations. we can assume that some structural features of specification in the datasets (1) and (3) differ; these features will be an object to a further research. in the top ten of the intra-paragraph inter-sentential relations (dataset 2), on the other hand, gradation, explication and correction are present, representing probably another way of text progression typical for smaller units. the datasets (1-3) differ not only in proportions of frequent semantic types, but also in those which are the least represented in the respective datasets. we have looked into semantic types of discourse relations which are represented by less than 1% of occurrences in each of the analyzed datasets, see table 7. as can be seen from table 7, some relations are rare in our data in general, independently from paragraph boundaries and syntactic structure. this is the case of all the pragmatic relations, equivalence and conjunctive alternative. another group of relations has a low representation in the dataset (1) and (2), but they are quite typical for intra-sentential relations (dataset 3). this applies to condition and purpose, see table 6 and table 7. we can consider these relations as typically syntactic and local. the relation of correction is usually not used across paragraphs. however, it is typical for both types of intra-paragraph relations. finally, we can find typical relations of the higher structure in our data (dataset 1), too, which occur in the intra-sentential structure (dataset 3), namely instantiation and generalization. let us sum up the observations concerning semantics of discourse relations in the datasets (1-3). distributions of semantic types in inter-sentential relations (datasets 1 and 2) are quite close to each other, but they differ from intra-sentential relations distinctly. this finding was confirmed by the result of χ2 tests, too. in other words, inter/intra-sententiality is a more important feature for the semantics of discourse relations than the presence or absence of paragraph boundary. nevertheless, 19 poláková, mírovský, zikánová and hajičová (1) cross-par inter-s (2) intra-par inter-s (3) intra-s equivalence 8 equivalence 57 explication 115 synchrony 3 synchrony 49 restr. opposition 99 correction 3 condition 42 conj. alternative 70 pragm. reason-result 2 pragm. contrast 27 equivalence 42 condition 1 pragm. reason-result 27 instantiation 27 disj. alternative 1 conj. alternative 20 pragm. contrast 23 conj. alternative 0 disj. alternative 14 pragm. reason-result 16 purpose 0 purpose 7 pragm. condition 16 pragm. contrast 0 pragm. condition 1 generalization 12 pragm. condition 0 table 7: distribution of semantic types with occurrence <1% for each dataset (number of occurrences) cross-paragraph relations (dataset 1) have specific features distinguishing them from intra-paragraph relations (datasets 2 and 3), too. besides the most frequent relations the distributions of which are very similar for all the three datasets (conjunction, opposition, reason-result), cross-paragraph relations typically express different meanings of expansion (generalization, specification, instantiation). on the other hand, they rarely express semantics of correction which is typical for intra-paragraph relations (datasets 2 and 3); moreover, some typical intra-sentential relations, such as condition, purpose, disjunctive alternative and synchrony almost do not occur across paragraphs in our data. 4.3.1 discourse connectives in paragraph-initial positions we also addressed the question of the differences in the use of discourse connectives across paragraphs and in local relations. we focused on the three most common relations, conjunction, opposition and reason-result, and the connectives by which these relations are expressed (see table 8). the analysis showed that the relation of conjunction is often expressed by specific discourse connectives in the cross-paragraph context, namely with the adverb dále [further, next] and the particle také [too, also]. the basic and most frequent discourse connective for conjunction a [and] is much more frequently used in the datasets (2) and (3), i.e. in the intra-paragraph structure. further, in the intraparagraph coherence, discourse connectives based etymologically on a prepositional phrase with an anaphoric element are common, cf. přitom [at the same time, lit. by that], and its intra-sentential relative variant přičemž [and/while, lit. by which]. discourse connectives based on relatives are not used in the inter-sentential structures. discourse connectives for opposition are almost identical for cross-paragraph and intra-paragraph inter-sentential relations (dataset 1 and 2). a typical discourse connective in both these groups is však [however], whereas intra-sententially (dataset 3), opposition is predominantly expressed by the conjunction ale [but]. in dataset (3), multi-part discourse connectives are used (with an nonautonomous part sice [sure/true]) which are not used inter-sententially. similarly, discourse connectives used for reason-result are very close in the datasets (1) and (2). again, a specific set of discourse connectives is used in the intra-sentential context (dataset 3). moreover, these two larger groups (inter-sentential datasets 1 and 2, and the intra-sentential set 3) 20 discourse relations and connectives in higher text structure (1) cross-par inter-s (2) intra-par inter-s (3) intra-s 5 most common discourse connectives: conjunction dále [further] 42 a [and] 309 a [and] 5 364 také [too] 32 také [too] 193 což [which] 165 a [and] 28 dodal [he added] 78 přičemž [by which] 71 rovněž [also] 25 rovněž [also] 74 a také [and also] 45 dodal [he added] 11 přitom [at the same time] 73 podobně [similarly] 39 5 most common discourse connectives: opposition však [however] 123 však [however] 906 ale [but] 732 ale [but] 33 ale [but] 310 však [however] 221 ovšem [yet] 24 ovšem [yet] 154 sice ale [true/sure but] 143 avšak [however] 4 jenže [but/only] 39 ovšem [yet] 45 jenže [but/only] 4 přitom [at the same time] 31 sice však [although yet] 37 5 most common discourse connectives: reason-result proto [therefore] 31 proto [therefore] 301 protože [because] 510 tedy [thus] 25 totiž [namely/you see] 285 nebot’ [for] 212 totiž [namely/you see] 21 tedy [thus] 142 takže [so] 116 a tak [and so] 6 tak [so] 66 a tak [and so] 95 v tomto směru [in this sense] 5 a tak [and so] 34 proto, že [because] 95 table 8: the most common discourse connectives for conjunction, opposition, and reason-result (number of occurrences) mostly differ in the direction of the reason-result relation which influences the choice of the discourse connectives. while inter-sentential relations mainly express the meaning of result (discourse connectives proto [therefore], tedy [thus], a tak [and so], v tomto směru [in this sense]), intrasentential relations more typically represent the meaning of reason (discourse connectives protože [because], nebot’ [for], proto, že [because]). generally, there is a certain difference among the observed datasets in the usage of secondary connectives. whereas in intra-sentential relations (dataset 3) secondary connectives do not occur in high positions of the table, they are yet frequent enough in the dataset 2 (dodal [he added]) and even more in the dataset 1 (dodal [he added], v tomto směru [in this sense]). to sum up, discourse connectives in the cross-paragraph and local coherence overlap to a large extent; nevertheless, paragraph-initial discourse connectives still have some special features. some of them are based on the inter-sentential character of cross-paragraph relations and they are common for datasets (1) and (2). thus, discourse connectives in these groups do not include relative elements or typical intra-sentential expressions (sice [true/sure]). further, in the inter-sentential context, discourse connectives expressing result are used more often than connectives of reason. the latter are, on the other hand, more frequent in the intra-sentential relations (dataset 3). within the group of inter-sentential discourse connectives (datasets 1 and 2), paragraph-initial connectives differ from intra-paragraph connectives in certain aspects. for the relation of conjunction, the specific discourse connective dále [further, next] is frequently used. the proportion of secondary discourse connectives is higher in this group than in datasets (2 and 3), too. 21 poláková, mírovský, zikánová and hajičová left arg size right arg size count percent (no. of sentences) (no. of sentences) 1 1 5 966 89.1% 1 large 92 1.4% large 1 617 9.2% large large 20 0.3% total 6 695 100% table 9: overview of argument sizes of inter-sentential discourse relations without lists in pdit 2.0. 4.4 large arguments although the minimality principle was taken into account during annotating argument extents in the pdit, the annotators could also mark large argument spans and non-adjacent arguments, if justified (compare 4.2). in a way, the minimality principle, if applied thoroughly, can prevent the detection of natural “whole paragraph” to “whole paragraph” relations and similar. in this part of the analysis, we try to look into cases where marking a large argument was superior to the minimality requirement, we quantify them and try to find explanations for them. for this study, relations with large arguments are specified as relations where either one or both arguments are larger than one sentence. for the analysis, we only take into account explicit inter-sentential relations and exclude list structures. we also exclude such intra-sentential relations, which cross the sentence boundary with a part of one of their arguments. they represent a very specific group (25 instances in the corpus). in the data of the pdit 2.0, extents of discourse arguments are given by the tree representation of a sentence – a discourse relation is marked between two tree nodes, roots of two subtrees that in most cases represent the arguments – and further specified by two range attributes, start_range and target_range, which help define more complex cases. it is important to keep in mind that for symmetric relations (i.e. where both arguments are of the same nature and the linear order of the arguments is always the same, as in conjunction, opposition or synchrony), the target_range attribute value always defines the extent of the left-sided argument and the start_range attribute value defines the extent of the right-sided argument. for asymmetric relations like reason-result, in which the arguments can switch the order, startand target range are assigned to each relation individually via semantic definitions of arg1 and arg2. arg1 always has the target_range attribute, disregarding its location. this property of the pdit annotation needs to be taken into account in the phase of the analysis below where the left or right position of the arguments matter.23 the proportions of inter-sentential relations with different sizes of arguments are summarized in table 9. there are 5 966 relations with single-sentenced both arguments in the pdit 2.0, in other words relations with “small” arguments, and they represent 89% of all inter-sentential relations. 23. the target arguments (arg1) are the following for asymmetric relations: (pragm.) result for the relation(s) of (pragm.) reason-result, statement being explained (explication), result of a (pragm.) condition ((pragm.) condition), denial of expectation (concession), statement being corrected (correction), statement having a purpose (purpose), lower degree (gradation), successive event (precedence-succession), statement being restricted/having an exception (restr. opposition), more general statement (instantiation, specification), more specific statement (generalization). for symmetric relations, target argument (arg1) is always the left-sided argument. 22 discourse relations and connectives in higher text structure left arg right arg start (arg2) target (arg1) relation total large left large right both large large start large target both large concession 29 27 1 1 23 5 1 condition 3 3 0 0 2 1 0 confrontation 39 35 2 2 2 35 2 conjunction 115 107 4 4 4 107 4 conj. alternative 2 1 1 0 1 1 0 correction 6 5 0 1 0 5 1 equivalence 11 8 0 3 0 8 3 instantiation 21 6 13 2 13 6 2 explication 23 7 14 2 15 6 2 generalization 78 76 2 0 2 76 0 gradation 13 10 2 1 2 10 1 opposition 173 158 12 3 12 158 3 pragm. contrast 3 2 1 0 1 2 0 pragm. reason-result 6 6 0 0 3 3 0 precedence-succession 9 9 0 0 9 0 0 reason-result 176 139 36 1 155 20 1 restr. opposition 13 12 1 0 1 12 0 specification 3 0 3 0 3 0 0 synchrony 6 6 0 0 0 6 0 total 729 617 92 20 248 461 20 table 10: large arguments in pdit 2.0, measured both for the left-right positions and start-range target-range directions (= semantics of the arguments); "large" means that the given argument spans more than one sentence while the other one (in the same column) only spans one sentence (of its subset); symmetric relations are marked with the light gray background. the remaining 729 relations that have at least one large argument (two and more sentences), form 11% of all inter-sentential relations. there are 617 relations with a large left argument and a singlesentenced right argument (large – 1), compared to only 92 relations with large right argument and a single-sentenced left argument (1 – large). there are only 20 relations with both large arguments (large – large), which is, given the figures for hierarchies, in our opinion a surprisingly low number, in terms of percentage negligible. the most distinctive finding for large arguments is the huge disproportion of left-sided and rightsided large arguments, or in other words, the great predominance of large left-sided arguments. in a detailed view, the figures in table 10, the “left right” section, show that this is the case across almost all discourse types (senses).24 however, the left-right positions of the arguments are only informative for symmetric relations (with grey background in table 10), we will elaborate on 24. in our data, three relations, disjunctive alternative, pragm. condition and purpose do not have any instances with large arguments. 23 poláková, mírovský, zikánová and hajičová them further in this section.25 for asymmetric relations, argument semantics plays a crucial role. there are three exceptions to the tendency of a larger left-sided argument in table 10: instantiation, explication and specification, where large right-sided arguments are more common. but, if we look at the “start-target” part of the table, which is more informative for these asymmetric relations, we can see that in all these cases, argument semantics goes hand in hand: these are precisely the relations where the right contexts are represented by those arguments, which are easily conceivable as more elaborated, expandable: the argument providing an explanation (large start arg = 15), the argument giving example (13), the specifying argument (3).26 for comparison, an analogous, and in terms of numbers much stronger, disproportion is visible in generalization, also an asymmetric relation, with 76 large left contexts which are also the 76 more specific arguments. the figures here fully comply with the intuitive notions of how semantics of these relations work, e.g. a generalizing, summarizing statement should be generalizing over a large previous text segment. what could be less intuitive are the sizes of individual arguments in reason-result relation. for this relatively frequent relation,27 (176 instances with large arguments in total), 155 (88%) have a large “reason” arguments but only 20 (11%) have a large “result” argument. the large arguments stand mostly on the left. thus, semantically, the arguments are in line with the similar explication relation, that means the dominance of large reason/explication arguments, but in terms of location, these arguments stand elsewhere, on the left for reason-result and on the right for explication. a possible explanation can be the different importance of the arguments, in the terms of rst nuclearity, and/or role of secondary connectives and connective phrases. for symmetric relations, we do not have a straighforward explanation for the predominance of large left-sided arguments. only the following assumptions can be suggested: we might evaluate this as an effect of the annotation strategy, given that the connective is mostly a part of the right-sided argument, and so it may seem unnatural to go beyond the strong right boundary of the connectivecontaining sentence. we may also ask if the minimality principle is applied the same way to left and right contexts. and/or, this phenomenon might be inherent to language. in a similar manner in which anaphora in the language occurs much more often than cataphora, the connectives (some of which are indeed anaphoric, compare webber et al. (2003), stede and grishina (2016) or poláková and mírovský (2019)), relate to small or larger previous semantic contents.28 from a cognitive perspective, this disproportion of argument sizes may be connected with the linear way of text production and also the gradual growth of information received by the reader, and the perspective of the annotator, who may proceed incrementally, like a reader, not knowing about the sizes of any right context. this issue needs a further insight, since it may be very important for the understanding of the difference of analytic perspectives in local and global annotation approaches. 25. this also implies that proportions of numbers for symmetric relations in the left-right section and the start-target section of table 10 are identical, they only reflect the annotation convention that the target argument is always on the left, in other words, the discourse arrow always leads to the left for the symmetric relations. 26. counts for these three relations are too low to draw any hard conclusion but even the small numbers here support the intuitive claims. 27. even when its intra-sentential instances are filtered out here, there are no protože [because]-relations included. 28. in our experience, cataphoric connectives are mostly secondary, with a demonstrative element that mostly introduces a dependent clause (e. g. thanks to the fact that... ) and so their scope is very narrow. this might be, however, language-specific. 24 discourse relations and connectives in higher text structure 5. conclusion the aim of this study was to determine, using corpus methods, to what extent local annotations of discourse relations enable to abstract and describe phenomena of global coherence or higher text structuring. we used the 50 thousand sentences of the prague discourse treebank 2.0 for czech and the equally sized penn discourse treebank 3 for english.29 the analysis focused on three main aspects: 1. the “shape” of the text in terms of mutual configuration of close discourse relations (pairwise); 2. cross-paragraph relations, their semantics and the properties of connectives in these relations as opposed to intra-paragraph settings, and 3. the size of arguments (text units) connected by discourse relations. regarding 1., even the discourse relations annotated in local annotations settings are assumed to form specific patterns, which are inherent in global coherence models like the rhetorical structure theory (rst), but not postulated for local coherence models. this includes recursive hierarchical structuring of smaller and larger relations. also, the rst model applies strong constraints on the overall document structure, defining it as a (constituency) tree with no crossings and overlaps. our analysis of relation configurations further contributes to the theoretical coherence-oriented research in particular by bringing empirical data to the discussion whether rst tree graphs are adequate (and sufficient) to represent discourse structure: czech data available to us support this claim, exceptions are very rare. we have described patterns that are typical (adjacency, progress, hierarchy, etc.), less typical (argument containment patterns, envelopment) and quite rare (total overlap, crossing, partial overlaps etc.) in our data, and analyzed them linguistically. further, we have compared our findings to those of a similar study conducted on english pdtb, version 2 (lee et al., 2006), learning that the proportions of occurrence of individual patterns roughly correspond in both corpora, although our study distinguishes some more subtle configurations. frequent patterns in our data comply with the rst tree structure rules. less frequent patterns in the pdit mostly deal with inclusion or exclusion of attribution spans, but also with the annotation strategies for secondary connectives in cases where they form a whole clause (it means that...) or they are anaphoric (in this respect). in some rare patterns, where, in our opinion, there is a violation of the tree structure in the sense of rst, we have found a small number of linguistically defensible interpretations that are not to be factored out due to discourse anaphora or attribution, as lee et al. (2006) suggest. next, we have investigated hierarchies built by nested local relations in both czech and english data. in all investigated properties in this respect, the two corpora are very similar. in both of them, we have detected even 5-level hierarchies, although they are quite rare. in a more detailed perspective, however, much of the structure proved to be intra-sentential: beyond the sentence boundary, local annotation of explicit connectives and altlexes does not expose hierarchical text structuring very often. detected hierarchies of inter-sentential relations only reached max. 3 levels of depth in both corpora. we do not claim that the trees detected by us are the trees a global analysis like rst would discover, but we demonstrate the existence of some hierarchical text structure in local annotation. 2. in the part of the analysis concerning cross-paragraph phenomena, the distributions of semantic types of discourse relations reveal dominance of three elaborative meanings, namely specification, generalization, and instantiation in relations crossing the paragraph boundary. these relations 29. the analysis of english locally annotated data is so far only complementary to the analyses of czech data but we plan to extend it also to other subtopics in this study in the future. 25 poláková, mírovský, zikánová and hajičová are not in the top ten most frequent of all other intra-paragraph relations,30 and our findings only confirm their intuitively perceived large role in text composition and lesser role in syntax. nevertheless, it was observed that the feature of inter/intra-sententiality is more important for the semantics of discourse relations than the presence or absence of the paragraph boundary. typical intra-sentential relations, such as condition, purpose, disjunctive alternative and synchrony, almost never occur in cross-paragraphs relations in our data. as for connectives, for the relation of conjunction, there seems to exist a specific discourse connective of cross-paragraph links, the connective dále [further, next]. distributions of other connectives in frequent relations between and within paragraphs do not vary much, and, counter-intuitively, also coordinating conjunctions (and, but) are fairly represented in cross-paragraph relations. the proportion of secondary discourse connectives is higher in these contexts. 3. the analysis of argument size examined the hypothesis that in local coherence models, existence of large arguments (more than one sentence) should be limited by the annotation principle to annotate minimal units. it was discovered that relations with one or both large arguments are indeed not very frequent in pdit 2.0 (11% of all inter-sentential relations) and that the large argument is in almost 85% on the left, which might be annotator’s bias when proceeding from left (known context) to right (unknown context), and/or it may an inherent property of texts. it would be interesting to compare this observation to the ways of tree branching in the rst-discourse treebank global annotations. in the future, we plan to further extend our analysis to the pdtb 3 data and we would like to include also implicit relations, entity-based relations and hypophora (relations of question and answer) in both languages. implicitness is an important feature of inter-sentential discourse relations, therefore inclusion of implicit relations into the research will enhance the general picture of the distribution of semantic types. this kind of results can be used then e.g. for the prediction of meaning of inter-sentential discourse relations. furthermore, the role of the typically implicit semantic types, such as instantiation or specification, can be described in detail in this way. the outcome of the study will be reflected in a future rst-like annotation of czech. first, the results will be confronted with a real pilot rst analysis on a sample of locally annotated czech texts with detected hierarchical organization, in order to assess the degree of equivalence of the hierarchical structures. the findings about hierarchies and the large arguments from this study can be further crosschecked with notions like nuclearity and canonical order of discourse arguments to find more about possible bridging through global and local frameworks. also, the distribution of semantic labels given the size of text arguments in local annotations can be related to the use of rst labels (e.g. to the division to subject matter and presentational rhetorical relations) and their correspondence in lower and higher structuring can be discussed. the fact that local discourse annotation in both pdit and pdtb also displays hierarchical structure (up to 5 levels of depth) but at least two lowest levels are usually intra-sentential, implies a large role of syntax in discourse complexity. syntactic hypotaxis/parataxis, but also the (a)symmetry of local discourse labels when related to nuclearity can be of advantage in a possible automatic rst pre-annotation or in rhetorical parsing. last but not least, methodologically, our experiments seem to reveal a lot about annotation strategies and biases: the minimality principle seems to affect the left-sided and right-sided argument sizes with great difference and its consistent application also may hinder the ability of local 30. with the sole exception of specification in intra-sentential use, which may be connected to the very frequent nominal right-sided arguments (and governing verb ellipsis) in the pdit annotaiton. 26 discourse relations and connectives in higher text structure models to accurately assess coherence of larger blocks. our results also open space for the hypothesis that the annotation procedure itself, e.g. the order in which individual segments are connected to other segments, may influence the segment size and the hierarchical structure formed. acknowledgements the authors would like to thank three anonymous reviewers for their detailed and constructive comments. the authors acknowledge support from the and the grant agency of the czech republic (project no. 20-09853s). the research reported in the present contribution has been using language resources developed, stored and distributed by the lindat/clariah-cz project of the ministry of education, youth and sports of the czech republic (project no. lm2018101). references nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. john a. bateman and klaas jan rondhuis. coherence relations: towards a general specification. discourse processes, 24(1):3–49, 1997. url https://www.tandfonline.com/doi/abs/10. 1080/01638539709545006. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in current and new directions in discourse and dialogue, pages 85–112. springer, 2003. url https://link.springer.com/chapter/ 10.1007/978-94-010-0019-2_5. iria da cunha, juan-manuel torres-moreno, and gerardo sierra. on the development of the rst spanish treebank. in proceedings of the 5th linguistic annotation workshop, pages 1–10. association for computational linguistics, 2011. debopam das, maite taboada, and paul mcfetridge. rst signalling corpus. linguistic data consortium, university of pennsylvania, philadelphia, 2015. url https://catalog.ldc.upenn. edu/ldc2015t10. markus egg and gisela redeker. how complex is discourse structure? in proceedings of lrec 2010, malta, 2010. url https://lexitron.nectec.or.th/public/lrec-2010_malta/pdf/ 796_paper.pdf. vanessa wei feng, ziheng lin, and graeme hirst. the impact of deep hierarchical discourse structures in the evaluation of text coherence. in proceedings of coling 2014, the 25th international conference on computational linguistics: technical papers, pages 940–949, 2014. url https://www.aclweb.org/anthology/c14-1089.pdf. barbara grosz and candace sidner. attentions, intentions, and the structure of discourse. computational linguistics, 12(3):175–204, 1986. url https://dash.harvard.edu/handle/1/ 2579648. caroline hagege and guillaume jacquet. combining temporal processing and textual entailment to detect temporally anchored events, december 18 2014. us patent app. 13/920,462. 27 https://www.tandfonline.com/doi/abs/10.1080/01638539709545006 https://www.tandfonline.com/doi/abs/10.1080/01638539709545006 https://link.springer.com/chapter/10.1007/978-94-010-0019-2_5 https://link.springer.com/chapter/10.1007/978-94-010-0019-2_5 https://catalog.ldc.upenn.edu/ldc2015t10 https://catalog.ldc.upenn.edu/ldc2015t10 https://lexitron.nectec.or.th/public/lrec-2010_malta/pdf/796_paper.pdf https://lexitron.nectec.or.th/public/lrec-2010_malta/pdf/796_paper.pdf https://www.aclweb.org/anthology/c14-1089.pdf https://dash.harvard.edu/handle/1/2579648 https://dash.harvard.edu/handle/1/2579648 poláková, mírovský, zikánová and hajičová jan hajič, eduard bejček, alevtina bémová, eva buráňová, eva hajičová, jiří havelka, petr homola, jiří kárník, václava kettnerová, natalia klyueva, veronika kolářová, lucie kučová, markéta lopatková, marie mikulová, jiří mírovský, anna nedoluzhko, petr pajas, jarmila panevová, lucie poláková, magdaléna rysová, petr sgall, johanka spoustová, pavel straňák, pavlína synková, magda ševčíková, jan štěpánek, zdeňka urešová, barbora vidová hladká, daniel zeman, šárka zikánová, and zdeněk žabokrtský. prague dependency treebank 3.5. data/software, univerzita karlova, mff, úfal, prague, czech republic, 2018. url http: //hdl.handle.net/11234/1-2621. jerry r. hobbs. coherence and coreference. cognitive science, 3(1):67–90, 1979. url https: //doi.org/10.1207/s15516709cog0301_4. mikel iruskieta, marıa j aranzabe, arantza diaz de ilarraza, mikel lersundi, and oier lopez de lacalle. the rst basque treebank: an online search interface to check rhetorical relations. in 4th workshop rst and discourse studies, pages 40–49, 2013. chloé kiddon, luke zettlemoyer, and yejin choi. globally coherent text generation with neural checklist models. in proceedings of the 2016 conference on empirical methods in natural language processing, pages 329–339, 2016. alan lee, rashmi prasad, aravind joshi, nikhil dinesh, and bonnie webber. complexity of dependencies in discourse: are dependencies in discourse more complex than in syntax? in proceedings of the 5th international workshop on treebanks and linguistic theories, pages 12– 23, 2006. url http://ufal.mff.cuni.cz/tlt2006/pdf/133.pdf. alan lee, rashmi prasad, aravind joshi, and bonnie webber. departures from tree structures in discourse: shared arguments in the penn discourse treebank. in proceedings of the constraints in discourse iii workshop, pages 61–68, 2008. ziheng lin, hwee tou ng, and min-yen kan. automatically evaluating text coherence using discourse relations. in proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 997–1006, 2011. url https://www.aclweb.org/anthology/p11-1100.pdf. haitao liu and hongxin zhang. rhetorical relations revisited across distinct levels of discourse unit granularity. discourse studies, 18(4):454–472, 2016. doi: 10.1177/1461445616647891. william c. mann and sandra a. thompson. rhetorical structure theory: toward a functional theory of text organization. text-interdisciplinary journal for the study of discourse, 8:243– 281, 1988. url https://semanticsarchive.net/archive/gmyndbjo/rst%20towards% 20a%20functional%20theory%20of%20text%20organization.pdf. daniel marcu. the theory and practice of discourse parsing and summarization. mit press, 2000. eleni miltsakaki, rashmi prasad, aravind joshi, and bonnie webber. annotating discourse connectives and their arguments. proceedings of the hlt/naacl workshop on frontiers in corpus annotation, pages 9–16, 2004. 28 http://hdl.handle.net/11234/1-2621 http://hdl.handle.net/11234/1-2621 https://doi.org/10.1207/s15516709cog0301_4 https://doi.org/10.1207/s15516709cog0301_4 http://ufal.mff.cuni.cz/tlt2006/pdf/133.pdf https://www.aclweb.org/anthology/p11-1100.pdf https://semanticsarchive.net/archive/gmyndbjo/rst%20towards%20a%20functional%20theory%20of%20text%20organization.pdf https://semanticsarchive.net/archive/gmyndbjo/rst%20towards%20a%20functional%20theory%20of%20text%20organization.pdf discourse relations and connectives in higher text structure jiří mírovský, pavlína jínová, and lucie poláková. does tectogrammatics help the annotation of discourse? in proceedings of coling 2012: posters, pages 853–862, 2012. url https: //www.aclweb.org/anthology/c12-2083.pdf. jiří mírovský, pavlína synková, and lucie poláková. extending coverage of a lexicon of discourse connectives. prague bulletin of mathematical linguistics, in prep. petr pajas and jan štěpánek. recent advances in a feature-rich framework for treebank annotation. in donia scott and hans uszkoreit, editors, proceedings of the 22nd international conference on computational linguistics, pages 673–680, manchester, 2008. the coling 2008 organizing committee. url https://www.aclweb.org/anthology/c08-1085.pdf. petr pajas and jan štěpánek. system for querying syntactically annotated corpora. in proceedings of the acl–ijcnlp 2009 software demonstrations, pages 33–36, suntec, 2009. association for computational linguistics. url https://www.aclweb.org/anthology/p09-4009.pdf. lucie poláková and jiří mírovský. anaphoric connectives and long-distance discourse relations in czech. computación y sistemas, 23(3):711–717, 2019. url https://www.cys.cic.ipn. mx/ojs/index.php/cys/article/view/3274. lucie poláková and jiří mírovský. mining local discourse annotation for features of global discourse structure. in lecture notes in artificial intelligence, 23rd international conference on text, speech and dialogue, lecture notes in computer science, pages 50–60, cham, switzerland, 2020. faculty of informatics, masaryk university brno, springer. isbn 978-3-030-583224. url https://link.springer.com/chapter/10.1007/978-3-030-58323-1_5. lucie poláková, jirí mírovský, and pavlína synková. signalling implicit relations: a pdtb-rst comparison. dialogue & discourse, 8(2):225–248, 2017. url https://firstmonday.org/ ojs/index.php/dad/article/view/10787. aurora pons-porrata, rafael berlanga-llavori, and josé ruiz-shulcloper. topic discovery based on text mining techniques. information processing & management, 43(3):752–768, 2007. rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of the sixth international conference on language resources and evaluation (lrec’08), pages 2961–2968, marrakech, 2008. european language resources association. url http://citeseerx.ist.psu.edu/viewdoc/ download?doi=10.1.1.165.9566&rep=rep1&type=pdf. rashmi prasad, aravind joshi, and bonnie webber. exploiting scope for shallow discourse parsing. in proceedings of the seventh international conference on language resources and evaluation (lrec’10), valletta, malta, may 2010. european language resources association (elra). isbn 2-9517408-6-7. url https://www.aclweb.org/anthology/l10-1634/. rashmi prasad, katherine forbes riley, and alan lee. towards full text shallow discourse relation annotation: experiments with cross-paragraph implicit relations in the pdtb. in proceedings of the 18th annual sigdial meeting on discourse and dialogue, pages 7–16, saarbrücken, germany, august 2017. association for computational linguistics. doi: 10.18653/v1/ w17-5502. url https://www.aclweb.org/anthology/w17-5502. 29 https://www.aclweb.org/anthology/c12-2083.pdf https://www.aclweb.org/anthology/c12-2083.pdf https://www.aclweb.org/anthology/c08-1085.pdf https://www.aclweb.org/anthology/p09-4009.pdf https://www.cys.cic.ipn.mx/ojs/index.php/cys/article/view/3274 https://www.cys.cic.ipn.mx/ojs/index.php/cys/article/view/3274 https://link.springer.com/chapter/10.1007/978-3-030-58323-1_5 https://firstmonday.org/ojs/index.php/dad/article/view/10787 https://firstmonday.org/ojs/index.php/dad/article/view/10787 http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.165.9566&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.165.9566&rep=rep1&type=pdf https://www.aclweb.org/anthology/l10-1634/ https://www.aclweb.org/anthology/w17-5502 poláková, mírovský, zikánová and hajičová rashmi prasad, bonnie webber, and alan lee. discourse annotation in the pdtb: the next generation. in proceedings 14th joint acl-iso workshop on interoperable semantic annotation, pages 87–97, 2018. rashmi prasad, bonnie webber, alan lee, and aravind joshi. penn discourse treebank version 3.0. data/software, linguistic data consortium, 2019. url https://catalog.ldc.upenn.edu/ ldc2019t05. university of pennsylvania, philadelphia. ldc2019t05. magdaléna rysová and kateřina rysová. the centre and periphery of discourse connectives. in proceedings of pacific asia conference on language, information and computing, pages 452– 459, bangkok, 2014. department of linguistics, faculty of arts, chulalongkorn university. url https://www.aclweb.org/anthology/y14-1052.pdf. magdaléna rysová, pavlína synková, jiří mírovský, eva hajičová, anna nedoluzhko, radek ocelák, jiří pergler, lucie poláková, veronika pavlíková, jana zdeňková, and šárka zikánová. prague discourse treebank 2.0. data/software. lindat/clarin digital library at the institute of formal and applied linguistics (úfal), faculty of mathematics and physics, charles university, 2016. url http://hdl.handle.net/11234/1-1905. jerry samet and roger schank. coherence and connectivity. linguistics and philosophy, 7(1): 57–82, 1984. url https://link.springer.com/article/10.1007/bf00627475. manfred stede and yulia grishina. anaphoricity in connectives: a case study on german. in proceedings of the workshop on coreference resolution beyond ontonotes (corbon 2016), pages 41–46, 2016. manfred stede and arne neumann. potsdam commentary corpus 2.0: annotation for discourse research. in proceedings of lrec 2014, pages 925–929, reykjavik, iceland, 2014. url http: //www.lrec-conf.org/proceedings/lrec2014/pdf/579_paper.pdf. maite taboada and william c mann. rhetorical structure theory: looking back and moving ahead. discourse studies, 8(3):423–459, 2006. url https://journals.sagepub.com/doi/ abs/10.1177/1461445606061881. christoph unger. genre, relevance and global coherence: the pragmatics of discourse type. springer, 2006. bonnie webber, matthew stone, aravind joshi, and alistair knott. anaphora and discourse structure. computational linguistics, 29(4):545–587, 2003. sandra williams and ehud reiter. a corpus analysis of discourse relations for natural language generation. in proceedings of corpus linguistics, 2003, lancaster university, united kingdom, 2003. url http://ucrel.lancs.ac.uk/publications/cl2003/papers/williams.pdf. florian wolf and edward gibson. representing discourse coherence: a corpus-based study. computational linguistics, 31(2):249–287, 2005. url https://www.mitpressjournals.org/ doi/abs/10.1162/0891201054223977. 30 https://catalog.ldc.upenn.edu/ldc2019t05 https://catalog.ldc.upenn.edu/ldc2019t05 https://www.aclweb.org/anthology/y14-1052.pdf http://hdl.handle.net/11234/1-1905 https://link.springer.com/article/10.1007/bf00627475 http://www.lrec-conf.org/proceedings/lrec2014/pdf/579_paper.pdf http://www.lrec-conf.org/proceedings/lrec2014/pdf/579_paper.pdf https://journals.sagepub.com/doi/abs/10.1177/1461445606061881 https://journals.sagepub.com/doi/abs/10.1177/1461445606061881 http://ucrel.lancs.ac.uk/publications/cl2003/papers/williams.pdf https://www.mitpressjournals.org/doi/abs/10.1162/0891201054223977 https://www.mitpressjournals.org/doi/abs/10.1162/0891201054223977 discourse relations and connectives in higher text structure florian wolf, edward gibson, amy fisher, and meredith knight. discourse graphbank, ldc2005t08 [corpus]. philadelphia: linguistic data consortium, 2005. url https: //catalog.ldc.upenn.edu/ldc2005t08. deniz zeyrek, amália mendes, yulia grishina, murathan kurfalı, samuel gibbon, and maciej ogrodniczuk. ted multilingual discourse bank (ted-mdb): a parallel corpus annotated in the pdtb style. language resources and evaluation, (54):587–613, 2019. url https: //link.springer.com/article/10.1007%2fs10579-019-09445-9. renxian zhang. sentence ordering driven by local and global coherence for summary generation. in proceedings of the acl 2011 student session, pages 6–11, 2011. šárka zikánová, eva hajičová, barbora hladká, pavlína jínová, jiří mírovský, anna nedoluzhko, lucie poláková, kateřina rysová, magdaléna rysová, and jan václ. discourse and coherence. from the sentence structure to relations in text. studies in computational and theoretical linguistics. úfal, praha, czechia, 2015. šárka zikánová, jiří mírovský, and pavlína synková. explicit and implicit discourse relations in the prague discourse treebank. in proceedings of the 22nd international conference on text, speech and dialogue tsd 2019, volume 11697 of lecture notes in computer science, pages 236–248, cham / heidelberg / new york / dordrecht / london, 2019. university of west bohemia, springer international publishing. isbn 978-3-030-27946-2. url https://link. springer.com/chapter/10.1007/978-3-030-27947-9_20. 31 https://catalog.ldc.upenn.edu/ldc2005t08 https://catalog.ldc.upenn.edu/ldc2005t08 https://link.springer.com/article/10.1007%2fs10579-019-09445-9 https://link.springer.com/article/10.1007%2fs10579-019-09445-9 https://link.springer.com/chapter/10.1007/978-3-030-27947-9_20 https://link.springer.com/chapter/10.1007/978-3-030-27947-9_20 poláková, mírovský, zikánová and hajičová appendix 1: a four-level hierarchy in prague discourse treebank 2.0 a czech original and an english translation of a pdit 2.0 text segment (cmpr9410_034, 1.5 paragraphs, 11 sentences) with a detected 4-level hierarchy of discourse relations. the pattern of the hierarchy is a ( b ( c ( d ) e ) f g ). a fb g c d e context: (1) přehlídka bezradnosti [paragraph boundary] (2) z historie je znám nejeden případ, kdy na mistrně vyvedeném padělku pohořela celá plejáda soudních znalců. (3) ve 30. letech tohoto století plnila titulní stránky berlínských deníků jedna z největších afér v historii falzifikátů. (4) obchodník s obrazy, jakýsi wacker, byl obviněn, že dodal na výstavu 33 falešných van goghů. (5) skandál brzy přerostl v odbornou polemiku. (6) obžalovací spis obsahoval několik tisíc stran. (7) soud vzal na vědomí názory více než šedesáti uměleckých historiků a kritiků. (8) nebyly mezi nimi ani dva, které by se nerozcházely alespoň v podrobnostech. (9) každé stanovisko našlo svého zastánce. (10) všichni začali být tou přemírou odborníků unaveni. [paragraph boundary] (11) umělecký restaurátor a badatel de wilt vzbudil v uměleckém světě senzaci, když (d) využil rentgenové záření k posouzení stáří obrazu. (12) jeho závěr zněl, že (c) soubor je starý minimálně tři desetiletí a tudíž pravý. (13) nebot’ (e) jaký blázen by měl zájem kolem roku 1895 padělat holandského umělce, jehož plátna byla k dostání po dvou stech francích? (14) obžaloba si též (b) najala znalce, který pracoval s rentgenem. (15) ten prozkoumal kolekci z hlediska malířského rukopisu a (f) došel k závěru, že (g) ji nenamaloval van gogh. (16) soud byl opět tam, kde předtím. (17) ani (a) znalci, kteří měli analyzovat obrazy chemickou cestou, nenašli společnou řeč. context: (1) parade of cluelessness [paragraph boundary] (2) there is more than one case in history in which a wide range of forensic experts have failed in judging a masterfully executed forgery. (3) in the 30’s of this century, one of the greatest affairs in the history of falsehoods filled the front pages of berlin newspapers. (4) a painting dealer, a wacker, was accused of supplying 33 fake van goghs to an exhibition. (5) the scandal soon became a professional controversy. (6) the indictment file contained several thousand pages. (7) the court has recorded the views of over 60 art historians and critics. 32 discourse relations and connectives in higher text structure (8) there were not two of them that did not diverge at least in detail. (9) each opinion found its defender. (10) everyone grew weary of the excess of experts. [paragraph boundary] (11) the artistic restorer and explorer de wilt caused a sensation in the world of art when (d) he used x-rays to assess the age of the painting. (12) his conclusion was that (c) the collection was at least three decades old and therefore genuine. (13) for (e) what fool would be interested, circa 1895, in forging a dutch artist whose canvases were available at two hundred francs each? (14) the prosecution also (b) hired an expert to work with the x-ray device. (15) he examined the collection in terms of the painting’s handwriting and (f) concluded that (g) it was not painted by van gogh. (16) the court was back where it had been before. (17) not even (a) the experts who were to analyze the paintings chemically could find common ground. relation a (‘ani’ [not even], conjunction): leftarg: (7 16) the court has recorded the views of over 60 art historians and critics. there were not two of them that did not diverge at least in detail. each opinion found its defender. everyone grew weary of the excess of experts. the artistic restorer and explorer de wilt caused a sensation in the world of art when he used x-rays to assess the age of the painting. his conclusion was that the collection was at least three decades old and therefore genuine. for what fool would be interested, circa 1895, in forging a dutch artist whose canvases were available at two hundred francs each? the prosecution also hired an expert to work with the x-ray device. he examined the collection in terms of the painting’s handwriting and concluded that it was not painted by van gogh. the court was back where it had been before. [soud vzal na vědomí názory více než šedesáti uměleckých historiků a kritiků. nebyly mezi nimi ani dva, které by se nerozcházely alespoň v podrobnostech. každé stanovisko našlo svého zastánce. všichni začali být tou přemírou odborníků unaveni. umělecký restaurátor a badatel de wilt vzbudil v uměleckém světě senzaci, když využil rentgenové záření k posouzení stáří obrazu. jeho závěr zněl, že soubor je starý minimálně tři desetiletí a tudíž pravý. nebot’ jaký blázen by měl zájem kolem roku 1895 padělat holandského umělce, jehož plátna byla k dostání po dvou stech francích? obžaloba si též najala znalce, který pracoval s rentgenem. ten prozkoumal kolekci z hlediska malířského rukopisu a došel k závěru, že ji nenamaloval van gogh. soud byl opět tam, kde předtím.] rightarg: (17) the experts who were to analyze the paintings chemically could find common ground [znalci, kteří měli analyzovat obrazy chemickou cestou, nenašli společnou řeč] relation b (‘též’ [also], conjunction): leftarg: (11 13) the artistic restorer and explorer de wilt caused a sensation in the world of art when he used x-rays to assess the age of the painting. his conclusion was that the collection was at least three decades old and therefore genuine. for what fool would be interested, circa 1895, in forging a dutch artist whose canvases were available at two hundred francs each? 33 poláková, mírovský, zikánová and hajičová [umělecký restaurátor a badatel de wilt vzbudil v uměleckém světě senzaci, když využil rentgenové záření k posouzení stáří obrazu. jeho závěr zněl, že soubor je starý minimálně tři desetiletí a tudíž pravý. nebot’ jaký blázen by měl zájem je kolem roku 1895 padělat holandského umělce, jehož plátna byla k dostání po dvou stech francích?] rightarg: (14) the prosecution hired an expert to work with the x-ray device. [obžaloba si najala znalce, který pracoval s rentgenem.] relation c (‘jeho závěr zněl, že’ [his conclusion was that], generalization): leftarg: (11) the artistic restorer and explorer de wilt caused a sensation in the world of art when he used x-rays to assess the age of the painting. [umělecký restaurátor a badatel de wilt vzbudil v uměleckém světě senzaci, když využil rentgenové záření k posouzení stáří obrazu.] rightarg: (sub12)31 the collection was at least three decades old and therefore genuine [soubor je starý minimálně tři desetiletí a tudíž pravý] relation d (‘když’ [when], specification): leftarg: (sub11) the artistic restorer and explorer de wilt caused a sensation in the world of art [umělecký restaurátor a badatel de wilt vzbudil v uměleckém světě senzaci] rightarg: (sub11) he used x-rays to assess the age of the painting [využil rentgenové záření k posouzení stáří obrazu] relation e (‘nebot’’ [for], reason – result): leftarg: (sub12) (it is) genuine [je pravý]32 rightarg: (13) what fool would be interested, circa 1895, in forging a dutch artist whose canvases were available at two hundred francs each [jaký blázen by měl zájem kolem roku 1895 padělat holandského umělce, jehož plátna byla k dostání po dvou stech francích] relation f (‘a’ [and], conjunction): leftarg: (sub15) examined the collection in terms of the painting’s handwriting [prozkoumal kolekci z hlediska malířského rukopisu] rightarg: (sub15) concluded that it was not painted by van gogh [došel k závěru že ji nenamaloval van gogh] relation g (‘došel k závěru’ [concluded that], generalization): leftarg: (sub15) he examined the collection in terms of the painting’s handwriting [prozkoumal kolekci z hlediska malířského rukopisu] rightarg: (sub15) it was not painted by van gogh [ji nenamaloval van gogh] 31. "sub" refers to the fact that the argument includes only a subset of a given sentence, not the whole sentence. 32. verb ellipsis restored from the previous clause via generated node in the tree structure 34 discourse relations and connectives in higher text structure appendix 2: a five-level hierarchy in the penn discourse treebank 3.0 the pdtb 3.0 text segment (wsj_2431, 2 paragraphs, 7 sentences) with a detected 5-level hierarchy of discourse relations. the pattern of the hierarchy is a ( b c d ( e ( f ( g )))).33 a cb d e f g context: (21) the coalition government tried to show that pasok ministers had received hefty sums for oking the purchase of f-16 fighting falcon and mirage 2000 combat aircraft, produced by the u.s.based general dynamics corp. and france’s avions marcel dassault, respectively. (22) naturally, neither general dynamics nor dassault could be expected to hamper its prospective future dealings by making disclosures of sums paid (or not ) to various greek officials for services rendered. (23) so it seems that mr. mitsotakis and his communist chums may have unwittingly served mr. papandreou a moral victory on a platter: pasok, whether guilty or not, can now traipse the countryside condemning the whole affair as a witch hunt at mr. papandreou’s expense. (24) but (b) while (c) verbal high jinks alone won’t help pasok regain power, mr. papandreou should never be underestimated. (25) first came his predictable fusillade: he charged the coalition of the left and progress had sold out its leftist tenets by (f) collaborating in a right-wing plot aimed at ousting pasok and (g) thwarting the course of socialism in greece. [paragraph boundary] (26) then (e), to buttress his credibility with the left, he enticed some smaller leftist parties to stand for election under the pasok banner. (27) next (d), he continued to court the communists – many of whom feel betrayed by the left-right coalition’s birth – by bringing into pasok a well-respected communist party candidate. (28) for balance, and in hopes of gaining some disaffected centrist votes, he managed to attract a former new democracy party representative and known political enemy of mr. mitsotakis. (29) thus (a) pasok heads for the polls not only with diminished scandal-stench, but also with “seals of approval” from representatives of its harshest accusers. 33. note that the relations b and c do not take part in the deepest branch of the hierarchy, they are only included in the higher relation a. 35 poláková, mírovský, zikánová and hajičová relation a (‘thus’, contingency.cause.result): leftarg: (23 28) so it seems that mr. mitsotakis and his communist chums may have unwittingly served mr. papandreou a moral victory on a platter: pasok, whether guilty or not, can now traipse the countryside condemning the whole affair as a witch hunt at mr. papandreou’s expense. but while verbal high jinks alone won’t help pasok regain power, mr. papandreou should never be underestimated. first came his predictable fusillade: he charged the coalition of the left and progress had sold out its leftist tenets by collaborating in a right-wing plot aimed at ousting pasok and thwarting the course of socialism in greece. then, to buttress his credibility with the left, he enticed some smaller leftist parties to stand for election under the pasok banner. next, he continued to court the communists – many of whom feel betrayed by the left-right coalition’s birth – by bringing into pasok a well-respected communist party candidate. for balance, and in hopes of gaining some disaffected centrist votes, he managed to attract a former new democracy party representative and known political enemy of mr. mitsotakis. rightarg: (29) pasok heads for the polls not only with diminished scandal-stench, but also with “seals of approval” from representatives of its harshest accusers relation b (‘but’, comparison.concession.arg2-as-denier): leftarg: (sub23) pasok, whether guilty or not, can now traipse the countryside condemning the whole affair as a witch hunt at mr. papandreou ’s expense rightarg: (24) while verbal high jinks alone wo n’t help pasok regain power, mr. papandreou should never be underestimated relation c (‘while’, temporal.synchronous): leftarg: (sub24) but... mr. papandreou should never be underestimated rightarg: (sub24) verbal high jinks alone won’t help pasok regain power relation d (‘next’, temporal.asynchronous.precedence): leftarg: (25 26) first came his predictable fusillade: he charged the coalition of the left and progress had sold out its leftist tenets by collaborating in a right-wing plot aimed at ousting pasok and thwarting the course of socialism in greece. then, to buttress his credibility with the left, he enticed some smaller leftist parties to stand for election under the pasok banner rightarg: (sub27) he continued to court the communists ... by bringing into pasok a wellrespected communist party candidate relation e (‘then’, temporal.asynchronous.precedence): leftarg: (sub25) he charged the coalition of of the left and progress had sold out its leftist tenets by collaborating in a right wing plot aimed at ousting pasok and thwarting the course of socialism in greece rightarg: (sub26) he enticed some smaller leftist parties to stand for election under the pasok banner relation f (‘by’, contingency.cause.reason): leftarg: (sub25) the coalition of the left and progress had sold out its leftist tenets 36 discourse relations and connectives in higher text structure rightarg: (sub25) collaborating in a right-wing plot aimed at ousting pasok and thwarting the course of socialism in greece relation g (‘and’, expansion.conjunction): leftarg: (sub25) ousting pasok rightarg: (sub25) thwarting the course of socialism in greece 37 introduction goals theoretical aspects and definitions related research rhetorical structure theory penn discourse treebank data and tools analysis configurations of close relation pairs hierarchies hierarchies in the pdit 2.0 hierarchies in the penn discourse treebank 3.0 paragraph-initial semantic types and connectives discourse connectives in paragraph-initial positions large arguments conclusion dialogue & discourse 14(2) 83–112 doi: 10.5210/dad.2023.203 ©2023 junfei hu and liesbeth degand this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). the conversational discourse unit: identification and its role in conversational turn-taking management junfei hu junfei.hu@uclouvain.be institute for language and communication university of louvain liesbeth degand liesbeth.degand@uclouvain.be institute for language and communication university of louvain editor: junyi jessy li submitted 08/2022; accepted 10/2023; published online 11/2023 abstract this study investigates how discourse segmentation and turn-taking interact. mapping syntactic, prosodic and pragmatic units, five types of conversational discourse units (cdu) were identified. based on this segmentation, associations were examined between the syntactic, prosodic and pragmatic boundaries and turn-taking, as well as the transition speed after each type of cdu. results show: 1) the relationships between the three linguistic boundaries and the occurrence of turn-taking were significant, and the association was the strongest for the pragmatic boundaries; it was weaker for prosodic boundaries and the weakest for the syntactic boundaries. 2) the type of cdu influenced the transition speed, with the pragmatic-syntax-bound cdu being fastest. the study highlights the importance of meaning-connection and earlier emergence of the utterance gist in timing turn-taking. keywords: discourse segmentation, discourse unit, corpus analysis, pragmatic unit, turn-taking 1 introduction the linear unfolding of (spontaneous) speech results from a combination of different cognitive encoding and decoding processes (e.g. auer, 2005; clark, 1996; pickering & garrod, 2021). the resulting sequences of segments involve syntactic, prosodic and pragmatic dimensions, which are most often treated in isolation from one another. yet, we believe it is important to acknowledge the interplay between these three aspects of linguistic encoding when defining the discourse segments that work as building blocks in spoken conversation. in the current study, we propose an operationalisation for segmenting spoken discourse into a sequence of conversational discourse units (cdus). we see cdus as a three-fold combination of syntactic, prosodic and pragmatic segments that are mapped onto one another in different configurations. it is our belief that getting hu and degand 84 deeper insight into the distribution of cdus in conversational discourse will improve our understanding of how speakers and hearers co-construct the ongoing conversation. a key determining aspect of (spontaneous) conversation is that speakership alternates between interlocutors apace (sacks, schegloff & jefferson, 1974). in a comparison of 10 different languages, interlocutors appeared to leave a gap between turns ranging from approximately 7 to 469 ms averagely (stivers et al., 2009), which is far shorter than the time needed to produce a single word (600 to 1,200 ms, depending on its frequency; indefrey & levelt, 2004; levelt, roelofs, & meyer, 1999). the mechanism underpinning such smooth transition is knowingly complex and has yielded several competing accounts (de ruiter, mitterer & enfield, 2006; bögels, 2020; bögels, magyari & levinson, 2015; bögels & torreira, 2015; sjerps & meyer, 2015). while there is general consensus that syntactic structure, prosodic realisation and pragmatic aspects of the utterance (i.e. the implementation of the communicative action) are involved in the turn-taking operation (ford & thompson, 1996; levinson & torreira, 2015), studies do not fully agree on the specific weights that syntax, prosody, and pragmatics should be given in such operation (bögels & torreira, 2015, 2021; de ruiter, mitterer & enfield, 2006; riest, jorschick & de ruiter, 2015). sacks, schegloff and jefferson (1974) proposed that conversation develops incrementally through the turn-constructional-unit (tcu), which intrinsically contains a recognisable boundary (i.e. transition-relevance-place, trp) where transfer of speakership becomes saliently possible in interaction. according to their argument, a transitional gap between interlocutors can largely be attributed to the projectability of the tcu. that is, the boundary of the tcu should be relatively predictable in advance. according to sacks, schegloff and jefferson (ibid: 709), ‘sentential constructions are the most interesting of the unit-type… [they] are capable of being analyzed in the course of their production by a party/hearer able to use such analyses (i.e. analysing sentence in terms of its expandability) to project their possible direction and completion loci’. in other words, syntactic structure appears to be the dominant priority in sacks, schegloff and jefferson’s framework. adopting the button-press paradigm, de ruiter, mitterer and enfield (2006) reached a similar conclusion. note that sacks, schegloff and jefferson (1974) acknowledged that in addition to syntax, other language resources, such as prosody, also play a role in turn-taking management. bögels and torreira (2015, 2021) reported that the majority of participants (96%) in an elicited dialogue did not take turns at potential syntactic completion points if prosodic information (f0 and duration) signalled incompletion. as a matter of fact, it appeared in a follow-up button-press experiment manipulating the prosodic features of the syntactic completion point that the prosodic information was useful and even more critical than syntactic information in turn-taking. yet, it is not uncommon for speakers to not take turns at either the syntactic or the prosodic completion point. speakers may actually take a turn at an incompletion point of the syntactic unit, as long as that structure is interactionally complete (li, 2016), or pass on the opportunity of taking a turn at the boundary of the syntactic unit (houtkoop & mazeland, 1985). thus, in addition to the interplay between syntax and prosody, turn-taking should also be interpreted in its pragmatic context (selting, 2000). the discrepant conclusions reached so far indicate that the three factors hardly create effects independently (auer, 1996; selting, 2000) and should be considered together. the starting point of the present study is the surface analysis of conversational discourse (see degand & simon, 2009), that is, the syntactic structure, prosodic realisation, and pragmatic unit with their respective boundaries as key elements in deciding where and when a cdu starts and ends. concretely, a cdu corresponds to a unit with coinciding syntactic, prosodic and pragmatic boundaries. in line with degand and simon (2009), we consider that a cdu is not complete as long as one of the three boundary types is open (awaiting completion). thus, a cdu should not be restricted to the smallest syntactic, prosodic, or pragmatic unit, rather we consider cdus as the segments resulting from the mapping between these three levels of analysis, giving rise to different types of cdus (see section 2). speakers and hearers rely on this information to construct and interpret the ongoing conversation, including the management of turn-taking. the conversational discourse unit: identification and its role in turn-taking management 85 with the segmentation on the three dimensions and the identification of cdu in hand, the current study aims to explore: 1) which boundary—syntactic, prosodic or pragmatic—is the best predictor for the probability of a transition to arise. 2) the extent to which the type of cdu influences transition speed between turns. as discussed above, turn-taking can occur after each type of linguistic boundary (i.e. syntactic, prosodic or pragmatic). this does not mean that every linguistic boundary will lead to a turn-taking in a real conversation. in some cases, interlocutors eschew the opportunity to take the turn when the linguistic boundary is present. segmentation allows us to calculate how many instances of each type of linguistic boundary have and have not been followed by a turn-taking. thus, generally, determining the linguistic boundary is the best predictor for the probability of a turn-taking to arise. meanwhile, the end of the cdu is where the three linguistic boundaries converge. hence, it is most likely to be followed by a turn-taking (duncan, 1972; ford & thompson, 1996). however, each type of cdu’s inner structure is unique in the way in which the syntactic, prosodic, and pragmatic combine. this motivates us to anticipate that the transition time after each type of cdu may be different. before we pursue the discussion on cdus and their role in turn-taking management, we want to acknowledge that human communication is in principle multimodal (holler & levinson, 2019). nonverbal symbols, such as gestures, are posited to coordinate with speech when conveying conceptual information (kita & özyürek, 2003; mcneill, 2005) and potentially contribute to timely turn-taking (holler, kendrick & levinson, 2018; kendrick, holler & levinson, 2023). however, because nonverbal semiotics cannot always be observed in language production, they are not the foci of this study. section 2 describes how this study defines cdus. section 3 introduces the dataset and delineates the methods informing cdu identification on the basis of syntactic, prosodic and pragmatic segmentation and determining turn-taking. section 4 displays the study’s results: significant associations were found between the occurrence of turn-taking and the three linguistic boundaries, which were differentially weighted in managing turn-taking. we also discovered that transition speed was influenced by the type of cdu. section 5 encompasses a general discussion of the study’s outcomes, highlighting that meaning-connections constructed during dialogues underpin the significant relationships found between the three linguistic boundaries and turn-taking management. we also speculate that the transition speed varies across the types of cdu, which may be attributed to the temporal difference in the emergence of the primary idea of the cdu (figure 5). section 6 finally presents the conclusions derived from the study. 2 conversational discourse units regarding the role of segmentation in understanding the unfolding of the (linear) discourse structure, two perspectives can be roughly distinguished: (i) the interactional approach focusing on units-of-action (e.g. ford, 2004; szczepek reed & raymond, 2013), (ii) the descriptive linguistic approach focusing on defining units of linguistic analysis (e.g. couper-kuhlen, 2006; degand & simon, 2009; izre’el et al., 2020; pietrandrea et al., 2014, among others). the interactional approach sees discourse as the outcome of a (collaborative) progressive construction of communicative action that can be (verbally) manifested as discourse segments corresponding to units of (collaborative) action. sacks, schegloff and jefferson (1974) view spoken communication through the lens of the turn-taking system where interlocutors undertake initiating and responsive actions in turns. in this context, ‘turns’ are considered as the basic structuring components of conversation, made up of turn-constructional units (tcus). turns themselves are part of larger types of conversational units, such as adjacency pairs and storytelling, that play a role in the so-called sequence organisation (schegloff, 2007: 9). the adjacency pair is organised around the actions of firstand second-pair parts (e.g. invitation-acceptance/rejection) and hu and degand 86 storytelling is built around expressing a stance toward an event (stivers, 2013). it is important to acknowledge that the tcu and its related units are inherently not defined in linguistic terms as the boundary of tcus is defined largely in terms of turn-taking (i.e. a potentially complete turn) (schegloff, 1996a; selting, 2000: 478). thus, such conversational units may be interesting in comprehending how conversation is constructed in terms of the formation and ascription of communitive action, yet they are not really fit as reliable units of linguistic analysis aiming at methodologically sound and descriptively adequate models of conversational structure. in contrast, the descriptive linguistics approach relies heavily on linguistic cues and rules to determine the (spoken) discourse unit. previous work has amply studied the linguistic nature of the ‘building blocks’ making up a piece of discourse, more specifically whether these units are based on syntax, prosody, pragmatics or any combination thereof (e.g. degand & simon 2009, izre’el et al., 2020; pietrandrea et al., 2014; prévot et al., 2015; steen, 2005). one influential descriptive linguistic perspective, known as the coherence approach, views discourse as a cohesive structure that prioritises the (semantic) representation of the discourse structure and its organisation, with each discourse segment being related to another segment to build a coherent whole; for example, through dependent or interdependent relationships (see e.g. rhetorical structure theory [mann & thompson, 1988] or the geneva discourse model [filliettaz & roulet, 2002]). this perspective also (implicitly) claims that the discourse unit most often corresponds to a syntactic clause. for instance, when attempting to divide the discourse into units for further structural analysis, mann and thompson (1988) indicate that the ‘units are essentially clauses, except that clausal subjects and complements and restrictive relative clauses are considered as parts of their host clause units rather than as separate units’ (p. 248). however, soundless syntactic input is barely imaginable, thus calling for prosodic input for any verbal interaction. as a matter of fact, it was shown that even in silent reading (fodor, 2002) or in communication between signers (brentari et al., 2018), prosodic information is used in syntactic decoding. kiaer (2014) illustrated the interplay between prosody and syntax in determining the grammaticality of discourse units, thus showing that they were not independent of one another. in her example (1), kiaer showed that in absence of prosodic information ({} stands for the prosodic phase), both sentence a and b were grammatical; however, when the prosodic information was taken into account, sentence b could not be considered grammatical because the intended constituency of ‘caroline for her family’ was no longer secured due to the improper phrasing. this conclusion is similar to carnie’s (2021), who highlights the importance of prosody when determining a verbless cluster constituency. in other words, discourse segmentation simply cannot exist without considering prosodic information. (1a) alison bakes cakes for tourists and {caroline for her family}. (1b) alison bakes cakes for {tourists and caroline} for her family. (kiaer 2014: 8) adhering to the principle that ‘syntax and prosody should be considered as two independent but complementary sources for the characterisation of discourse units’ (degand et al, 2014: 243), degand and simon (2009; also see izre’el et al., 2020) provided an operationalisation for linguistically segmenting spoken discourse into reusable basic discourse units. they proposed to independently segment french speech into intonation units, on the one hand, and syntactic units, on the other. subsequently, the mapping between these two levels of segmentation yielded several types of basic discourse units. while their model accounted for both syntax and prosody and the interplay between them, a crucial aspect of human communication was left out, namely pragmatics. conversation is, by its nature, a vehicle through which interlocutors implement and interpret their interactional plan (levinson, 2013; schegloff, 2007). for example, in extract 1,1 speaker 1 is trying to describe the 1 each line represents a single intonation unit (iu). and in this study, we used the terms prosodic unit and intonation unit interchangeably. each iu has been categorised as one of the four types according to its acoustic characteristics (du bois et al., 1993). the period is used to indicate a marked fall in pitch at the end of the iu. the question mark is used to the conversational discourse unit: identification and its role in turn-taking management 87 two attachments in the red oval in figure 1 to speaker 2, who needs to make a drawing of them.2 speaker 1 uses a simple sentence (l.2) with one intonation unit that has a syntactic structure similar to speaker 2’s interrogative utterance (l.1) to confirm speaker 2’s inquiry. it is common for speakers to show a confirmation by (partially) repeating their dialogic partner’s utterance (raymond, 2003; schegloff, 1996b), and speaker 1’s response is thus, to a certain extent, driven by the aim of her communication—namely, confirming speaker 2’s inquiry. speaker 2 then uses a simple sentence starting with ‘and’ to confirm the exact number of attachments and maintain the connection between her upcoming inquiry and the information already provided (l.6). speaker 1 then produces a confirmative reply in one intonation unit (l.7) that is followed by a metaphor to make it easier for speaker 2 to visualise the attachments (l.9). the most relevant data about the attachments given by speaker 1 in lines 2, 7, and 9 are structured in three sentences (< > stands for one sentence) and three intonation units ({} stands for one intonation unit). however, speaker 2 integrates all this information in one utterance that is realised in one intonation unit to display her understanding (l.10). we believe that the difference in the utterance-design in terms of their syntactic and prosodic realisation is (moderately) the result of the different commutative purpose held by these two speakers: speaker 1 tries providing the information step by step, while speaker 2 attempts to summarise, in one step, the information collected to finalise her drawing. the extract shows that syntactic (tao & hu, 2019; thompson & hopper, 2001) and prosodic realisations (szczepek reed, 2010, 2011) contribute toward realising the speaker’s communicative intention (e.g. the purpose of the communication). in other words, the pragmatic aspect of spoken discourse should not be ignored. indeed, several studies have indirectly noted that the understanding of discourse units should consider the pragmatic component of the utterance, arguing that ‘units in the conversation must be understood as usable for construction of joint activity, not merely as packages of information to be parsed’ (ford, fox & thompson, 1996: 427) and that discourse units ‘consist of an illocutionary act, a proposition, a clause, and an intonation unit’ (steen, 2005: 283). extract 1 1 speaker 2 <{and they're going to the left?}> 2 speaker 1: <{they’re going to the left}>. (3 lines are omitted) 6 speaker 2: <{and there are two of them}>. 7 speaker 1: <{yeah}>. 8 <{it’s-} 9 {like he's sticking out two arms}>. 10. speaker 2: <{sticking out two arms but on the same side}>. it should be noted that some discourse theories do emphasise that discourse analysis should account for the pragmatic aspects of the utterance, such as the speaker’s intention and speech act. for instance, segmented discourse representation theory (asher & lascarides, 2003), which analyses the interactions between discourse interpretation and discourse coherence, proposes that the discourse segments are rhetorically related and that each rhetorical relation can be considered a type of speech act, like narration or elaboration. lexical cues (e.g. referring phrases, [grosz & sidner, 1986]) and syntactic knowledge (e.g. clause, appositions and adverbials, [afantenos et al., 2012]) can be used to detect the boundary of a discourse segment. yet, a clearly elaborated operationalisation for determining the discourse segment is largely absent. indicate a high rise in pitch at the end of the iu. the comma type is used to indicate (1) a pitch that rises slightly at the end of the iu, which typically begins with a pitch at the low or mid-level, or (2) a terminal pitch that either remains level or falls slightly but not far enough to be considered final. the dash refers to an iu that breaks off in mid-utterance. the period and question-mark ius belong to the final type of iu. the comma and dash iu are labelled as to the non-final type (ford & thompson, 1996). 2 further information about the specific experimental dataset containing this extract will be provided in section 3.1. figure 1. item in an elicitation task hu and degand 88 reconciling the perspectives respectively prioritised by the aforementioned two schools, we extend upon degand and simon’s (2009) operationalisation proposing that spoken discourse should be segmented into a set of conversational discourse units (cdus) that integrate syntactic, prosodic and pragmatic units. in addition to its linguistic analytical components, the concept of cdu also has a cognitive foundation that echoes pickering and garrod’s (2021) model of language production in dialogue. roughly speaking, they propose that to communicate an idea in dialogue, the speaker needs to go through the planning and implementing phases. the planning phase (‘dialogue planner’) involves the situation and game models. the situation model corresponds to the content of what the interlocutors are discussing. the game model represents the form of the contribution the interlocutor seeks to make (e.g. raising a question to seek information from your partner). furthermore, ‘each interlocutor has a representation relating to what the interlocutors are trying to achieve at that point in the dialogue’ (pickering & garrod, 2021: 116). after finishing the planning stage, a production command will be sent to the implementing phase (‘dialogue implementer’) for linguistics encoding. in our view, the pragmatic unit corresponds to the communication plan generated in the game model, while the syntactic and prosodic (i.e. intonational) units are manifestations of the messages represented in the situation model. along this line, cdus can be considered as the outcome of the situation and game models. 3 data and method this section introduces the dataset and the means applied to determine the syntactic, prosodic (i.e. intonational), and pragmatic units. we then present the five types of cdus we obtained, and display our methods of annotating turn-taking. 3.1 dataset this study is based on six face-to-face conversations between native speakers of american english. the conversations were conducted between two speakers, enabling us to eliminate the issue of speaker selection in a multi-party conversation. it also enabled us to ensure that the response of each conversation to the tailored initiation could be anticipated. five of these face-to-face conversations were casual conversations between two acquainted adults of a similar age (female = 5 and male = 5; ~114 min). these conversations were obtained from the santa barbara corpus of spoken american english (sbc; du bois et al., 2000–2005). the criteria adopted in the current study for selecting conversations from sbc have been determined with two aims. on the one hand, the criteria were designed to ensure that the recorded conversations are as natural as possible. on the other hand, we tried to rule out the factors that might influence turn-taking time to ensure that the gap between turns can be primarily attributed to linguistic information and comprehension of utterance. first, even if naturalness is a continuum, casual talk has been acknowledged to be more natural than professional talk in terms of content and turn-design (drew, 2013). second, restricting the data set to conversations between two individuals rules out potential delays resulting from speaker selection (sacks, schegloff & jefferson, 1974). third, when interlocutors do not know each other or have a large age difference, more pragmatic issues may arise in language management, further influencing the turn-taking time. for instance, there may be more concerns about the preference of the content of the turn, as dispreferred turns, such as challenges and refusals, tend to receive a delayed response (heritage, 1984). note that there were seven conversations that fulfilled the criteria in sbc. the strong background noise (i.e. music and sounds from tv shows) made two of them non-eligible for analysis. the sixth conversation was a task-oriented dialogue (38 min in total), in which two females, who were unacquainted at the time of data collection, performed a description-drawing task with images of 2d objects (adapted from eijk et al., 2022; rasenberg et al., 2022). the task-oriented dialogues were collected with the objective to explore interactive alignment at the discourse level the conversational discourse unit: identification and its role in turn-taking management 89 (hu & degand, 2022). the motivation to include a task-oriented dialogue in this study was to take advantage of its relatively clean sequential structure to facilitate conceptualisation and operationalisation of the notions of ‘action plan’ (section 3.2.3.1) and ‘pragmatic unit’ (section 3.2.3.2), which were then applied to the remaining five natural conversations. the results presented in section 4 are based on the five sbc conversations. conversation (selected from sbc) cdu token length (token/cdu) turn 1 354 2,368 6.69 91 2 382 2,802 7.34 97 3 358 3,233 9.03 93 4 500 3,628 7.26 121 5 735 5,313 7.23 105 total 2,329 17,344 7.45 (by average) 507 table 1. descriptive overview of the dataset. 3.2 annotating syntactic, intonation and pragmatic units 3.2.1 syntactic unit the assumption held by all varieties of dependency grammar is that, within a clause, every element is embedded within binary asymmetrical relations called dependency relations, such that none remains isolated (de marneffe & nivre, 2019; heringer, 1993; tesnière, 1959). this is in line with the way people temporally build up their utterances: ‘as the speaker continues from point to point in interaction in real-time … the units, bits and pieces that are uttered or added are not autonomous atoms but are subject to syntactic interdependencies’ (linell, 2013: 62). therefore, even if the notion of the syntactic unit is neither stable nor uncontroversial in current linguistics, the theoretical basis of the annotation of the syntactic unit in the current study is dependency syntax. the starting point of analysis is verbal syntax in which a verb and its governed dependents (including core and oblique arguments, nivre et al., 2016) are central. this analysis leads to the socalled ‘dependency clause’ that demonstrates maximal syntactic completeness. in addition to the verbal dependency clauses, we also consider nonverbal dependency clauses, in which the nonverbal head (e.g. noun, adjective) is modified (cf. degand & simon, 2009). in natural conversations, certain elements are not governed by the main verb but are semantically or pragmatically linked to the dependency clause. these elements, such as discourse markers (e.g. and, so, but, frankly, or well), primarily signal semantic or discourse relations between utterances or utterance-clusters, rather than contributing to the meaning of the proposition itself. these relations tend to occur at the peripheries of the clause (degand & crible, 2021) to manage the flow of thought and the structure of a conversation. while elements of dependency and embedding are considered as microgrammatical expressions, discourse particles, interjections (e.g. oh, wow), polite expressions (e.g. please), and vocatives (e.g. scott) are considered as macrogrammatical expressions (haselow, 2017). the linearization of these macrogrammatical expressions is not constrained by the relations of dependency, but rather by real-time cognitive activity, the requirement of text organisation, and the need to manage interlocutors’ relationship. therefore, the current study treats these macrogrammatical expressions as independent syntactic units. 3.2.2 intonation unit in broad terms, an intonation unit (iu) is a stretch of speech produced under a single coherent intonation contour. operationally, du bois et al.’s (1993) schema in figure 2 indicates that an iu is generally marked by cues such as a pause along with an initial pitch reset and a lengthening of hu and degand 90 the final syllable. the iu can represent a stretch of speech that conveys substantive content or regulates information flow (chafe, 1993). recently, inbar et al (2020) have shown that the iu, as conceptualised in du bois et al.’s schema, is neurologically motivated. although du bois et al.’s criteria to identify the boundary of iu have shortcomings (barnwell, 2013; barth-weingarten, 2013) and alternative proposals on iu identification have been put forward in recent years, such as relying on the layperson’s perceptive judgement, it remains a practical way to identify ius, especially when the dataset is large. we have thus abided to this schema for the prosodic segmentation in the current study. figure 2. illustration of intonation units. 3.2.3 pragmatic unit section 2 discussed the need to include a pragmatic dimension when analysing conversational discourse and defining cdus. this paper surmises that the speaker’s communicative act originates from the primitive motive of achieving a goal, such as making a request, sharing ideas or seeking information, which are represented somewhere in the mind (levelt, 1989; pickering & garrod, 2021). the utterance, then, is a verbal reflection of a conceptual and communicative goal. the listener is expected to put effort into understanding the speaker’s intention (grice, 1975, 1989) and reply to the speaker in a cooperative way. most interlocutors have at least a (rudimentary) plan of their conversational intentions: either initiating action on the speaker’s part or performing a responsive action on the listener’s part. as such, we consider a pragmatic unit as a step towards implementing the speaker’s action plan. 3.2.3.1 action plan in a nutshell, the action plan is the implementation of the speaker’s interactive goal. yet, assuming the existence of an action plan for any interaction by no means implies that interlocutors must have a full-fledged procedural implementation plan before initiating the interaction. in general, the way in which the plan is implemented emerges from the negotiation between interlocutors. as the interaction unfolds, the original motivation for initiating the conversation can be abandoned, replaced with an intention to do other things or co-exist with other motivations that arise during the conversation. the consequence is that interlocutors can suspend the conversation at any time, change the topic as the conversation unfolds or approach several topics simultaneously. pragmatic units being the building blocks of the action plan, some pragmatic units can end up as incomplete, because the underlying action plan was aborted. the conversational discourse unit: identification and its role in turn-taking management 91 next, it is important to consider that given the wide variety of ways in which speakers may implement their action plan, there is no exhaustive list of types of action plans (or pragmatic units). lists of this type would be insufficient, by definition, to holistically explaining human action (levinson, 2013), also because no specifically tailored practice exists to form a certain action (enfield & sidnell, 2017). it follows that the identification of the action plan and the pragmatic units should depend on the analysis of the situated context. 3.2.3.2 identification of the pragmatic unit (pu) to our knowledge, ford and thompson (1996) is the first study that directly tries to segment speech into pragmatic units. they suggest that the pu has ‘to be interpretable as a complete conversational action within its specific sequential context’ (p. 150) with a distinct boundary after which no further content is expected to follow. accordingly, they expect pus to end in a final intonation contour. as for how to practically determine a complete conversational action, ford and thompson (1996) mainly rely on the listener’s reaction. they distinguish local pragmatic boundaries from global pragmatic boundaries. thus, a point where recipients respond in a non-taking-turn manner, e.g. by means of the backchannel uh-hum to show interest and encourage more details, should be treated as a local pragmatic boundary. at a global pragmatic boundary, no additional content is expected and genuine speakership alternation takes place. this approach contrasts with that of steen (2005) who assumes that syntactic clauses contain sufficient information to convey an action (cf. thompson & couper-kuhlen, 2005). there is hence overlap between pragmatic units and syntactic units, the latter constituting the minimal unit in dialogue that possesses sufficient information (i.e. complete proposition) to make a listener act (i.e. illocution). our approach to determining pus was inspired by the aforementioned studies. we agree that a complete pu should be intact in meaning-expression and action-conduction. however, our operationalisation is intrinsically different from the aforementioned criteria. the identification of the pragmatic unit does not rely on the response of the addressee in current study (cf. ford & thompson, 1996). and a single pragmatic unit may be interpreted as having more than one illocutionary act. therefore, we define the pragmatic unit as one step in the implementation of the action plan of the speaker rather than as the representation of one illocutionary act (cf. steen, 2005). for instance, when the addressee simply replies ‘okay’ to the addresser’s instruction, the ‘okay’ can be an acknowledgement of understanding while simultaneously closing the sequence. in line with the approach advocated by degand and simon (2009), we made segmentation independent of the syntactic and prosodic information. to do so, we relied on the situated context of the utterance to determine what action it invoked and which pu boundary was applicable. as shown in table 2, utterances in l.1 to l.4, taken form the experimental dataset, form one pragmatic unit (marked in angle brackets). they are produced to describe the cup-shape in the red rounded rectangle. in l.5, speaker 2 acknowledges her understanding by means of the marker ‘okay’, which forms a complete pragmatic unit. from l.6 to l.17, four pragmatic units can be identified. from l.6 to l.10, there is one pragmatic unit describing the position and shape of attachment in the red rounded rectangle. the utterances from l.11 to l.13 do not provide any new information. speaker 1 basically reformulates her utterances to make the description clearer. we treat these utterances from l.11 to l.13 together as a single pragmatic unit. from l.14 to l.16, the speaker provides further detailed description of the shape of the rectangle, indicating that the long bottom side of it is at an angle, to make the previous description more specific and clearer. after the additional explanation, utterance in l.17 is produced to reiterate that the sticking-out shape is basically a normal rectangle. speaker 2, who needs to draw the rectangle, feels confused by the information. she says ‘wait’ (l.18) to indicate that she has problems understanding and asks that the description be suspended. utterances from l.19 to l.22 are used to conduct the interrogative action as a whole. they are considered as one pragmatic unit. as shown by the illustration, pragmatic units can be expressed by either a single word (l.5 and l.18), a clause, (l.6 to l.11) or a clause-combination (l.19 to l.22). they can end with either a final or non-final intonation contour. hu and degand 92 interlocutors communicative action the pragmatic unit embodies contents of the task-oriented dialogue target speaker 1 1. so? 2. there is a, 3. cup? 4. that is rounded at the bottom? speaker 2 5. okay. speaker 1 6. and, 7. below the rounded part of the bottom of the cup, 8. there’s a kind of a, 9. rectangle, 10. sticking out, 11. at the bottom, 12. from the rounded part, 13. a rectangle sticking out, 14. and it’s a pretty normal rectangle with the exception that 15. the bottom side, 16. the long bottom side is at an angle. 17. but other than that it's a regular rectangle. speaker 2 18. wait. 19. so, 20. it's a cup, 21. sitting on a rectangle like this? 22. or the rectangle is like that? table 2. example of the segmentation of the pragmatic units. in addition to independence from syntactic and prosodic information, the operationalisation of the current study does not rely on recipient feedback to determine where a pu starts or ends. as noted before, the pragmatic unit corresponds to one step of an action plan. since the plan in principle does not need to be full-fledged, people may in some cases fail to produce a complete pragmatic unit. in extract 2, a face-to-face informal conversation, karen encounters difficulties of verbalisation. scott helps her finish the utterance. in that case, we treat scott’s utterance as a complete pragmatic unit. in contrast, since karen fails to implement her action, her utterance in l.1 is identified as an incomplete pragmatic unit. extract 2 (du bois et al., 2000–2005) 1 → karen: it wouldn't be as=, 2 → scott: it wouldn't be as aesthetically pleasing. 3 karen: mhm. 4 (tsk)

. an incomplete pragmatic unit may also result from one of the speakers ceasing talking when the interlocutors notice that their utterances are overlapping with each other (sacks, schegloff & jefferson, 1974). in extract 3, karen utters ‘would’ (l.5) after hearing scott’s one-word response further explanation make a description acknowledgment make a description clarify and summarise the previous description ask for a clarification stop the description summarise the conversational discourse unit: identification and its role in turn-taking management 93 (1.3). evidently, however, she notes that scott does not intend to alter the speakership. thus, she aborts her production immediately, which results in the incomplete pragmatic unit with a truncated intonation (l.5). extract 3 (du bois et al., 2000–2005) 1 karen: i think we thought about doing that in the springtime, 2 then i thought we'd replant them. 3 scott: yeah, 4 but [i] think there's more babies than ... we need. 5 → karen: [would] - 6 yeah. 7 and we’ve got three new babies that i could replant, 8 but i haven't, a further phenomenon that merits attention concerns instances in which the speaker either interrupts the progression of his/her own speech to fix or fine-tune issues, which is known as ‘selfinitiated’ self-repair (schegloff, 1977, 2013). we consider self-initiated self-repair in most of the cases as part of one pragmatic unit, because the repair is one step in implementing the action plan. additionally, when pragmatic units are segmented, it is common to find that the speaker is quoting others’ words. regardless of the inner structure of the quotation, we consider that the quotation as a whole forms a single pragmatic unit (cf. houtkoop & mazeland, 1985; jefferson, 1978). 3.2.4 types of conversational discourse units as discussed previously, syntactic, prosodic, and pragmatic units combine to form a conversational discourse unit. the example below illustrates the mapping procedure. in extract 4, speaker 1 attempts to describe the cuplike shape illustrated by the rounded rectangle in red to speaker 2, who is required to draw it. extract 4 1 speaker 1: the first shape it kind of looks like a cup with a bunch of stuff attached to it so like a little rounded at the bottom and you see it at a bit of an angle 2 speaker 2: okay is the rounded part sticking out of the cup or is the cup 3 speaker 1: no it’s the bottom of the cup shape is rounded 4 speaker 2: is rounded at the bottom 5 speaker 1: yes on the pragmatic level, this extract contains nine pragmatic units displayed in table 3. the action plan of speaker 1 is to instruct speaker 2 to draw the cuplike shape. meanwhile, the aim of speaker 2 is to obtain information to draw the shape. speaker 1 first provides a general description of the entire object (pragmatic unit 1). then, she focuses on the cuplike shape and provides two pieces of information about it in two pragmatic units. one is that the bottom of the cup is rounded (pragmatic unit 2). the other is that the cup is placed at an angle (pragmatic unit 3). then, speaker 2 says, ‘okay’, to convey her acceptance and understanding of the description, which forms one independent pragmatic unit. however, the description so far is insufficient to accomplish the drawing. speaker 2 then asks for more specific information about the relation between the rounded part and the cup shape (pragmatic unit 5). speaker 1 says, ‘no’, which reflects the sixth pragmatic unit in the extract to deny that the rounded part is separated from the cup. she then offers a clarification to further indicate that the bottom of the cup is rounded (pragmatic unit 7). speaker 2 then paraphrases speaker 1’s description to confirm her understanding (pragmatic unit 8). speaker 1 then says, ‘yeah’, as a confirmative reply (pragmatic unit 9). as can be observed, each pragmatic unit is produced to implement one step of the speakers’ action plans. the whole process of achieving the communicative goal is collaborative. hu and degand 94 interlocutors communicative action the pragmatic unit embodies contents of the task-oriented dialogue target speaker 1 1. the first shape it kind of looks like a cup with a bunch of stuff attached to it. 2. so like a little rounded at the bottom, 3. and you see it at a bit of an angle speaker 2 1. okay. 2. is the rounded part sticking out of the cup or is the cup shape itself. speaker 1 1. no, 2. it’s the bottom of the cup shape is rounded. speaker 2 1. is rounded at the bottom. speaker 1 1. yeah. table 3. example of the segmentation of the pragmatic units. because of space limitations, we will focus on the utterances within the first pragmatic unit only to display the segmentations into intonation and syntactic units. prosodically, two obvious pauses divide the utterances in the first pragmatic unit into three intonation units (figure 3). on the syntactic level, there are two syntactic units (figure 4), a nonverbal dependency clause, in which the nonverbal head (i.e. ‘shape’ in this case) is modified, followed by and a verbal dependency clause. mapping between the three levels in this case gives rise to the pragmatics-bound cdu where one pragmatic unit corresponds to two syntactic dependency clauses and three intonation units. figure 3. example of the segmentation of the intonation units. request more information confirmation inquiry denial clarification acknowledgment give a general description further explanation further explanation the conversational discourse unit: identification and its role in turn-taking management 95 shape looks like first it kind of cup the a with stuff bunch of attached to a it the first shape it kind of looks like a cup with a bunch of stuff attached to it syntactic unit syntactic unit figure 4. example of the segmentation of the syntactic units. following the same criteria and procedures, extract 4 can be segmented into a series of syntactic and intonation units. table 4 presents the three-level segmentations of extract 4 and the resulting type of cdu from the mapping. interlocutor pragmatic unit syntactic unit intonation unit cdu type s1 give a general description 1. the first shape 1. the first shape, pragmaticsbound 2. it kind of looks like a cup with a bunch of stuff attached to it 2. it kind of looks like a cup? 3. with a bunch of stuff attached to it, further explanation 1. so 1. so like a little= pragmaticsbound 2. like a little rounded at the bottom 2. rounded at the bottom, further explanation 1. and 1. and you see it at a bit, pragmaticsbound 2. you see it at a bit of an angle 2. of angle, s2 acknowledgement 1. okay 1. okay, congruent request more information 1. is the rounded part sticking out of the cup 1. is the rounded part, 2. sticking. 3. out of the cup? 4. or is, 5. the cup shape itself pragmaticsbound 2. or 3. is the cup shape itself s1 denial 1. no 1. no, congruent clarification 1. it’s the bottom of the cup chape is rounded 1. it’s the bottom, pragmaticssyntaxbound 2. of the cup shape is rounded s2 inquiry 1. is rounded at the bottom 1. is rounded at the bottom congruent s1 confirmation 1. yeah 1. yeah. congruent table 4. three-level segmentations and cdus of extract 4. hu and degand 96 mapping between the three levels yielded five cdu types, with coinciding intonational ending, pragmatic boundary, and syntactic closure.3 the units listed below are the five types observed in the dataset we examined. table 5 illustrates them with examples from the dataset. • congruent type: one pragmatic unit corresponds to one syntactic dependency clause and one intonation unit; • pragmatics-syntax-bound type: one pragmatic unit and one syntactic dependency clause correspond to two or more intonation units; • pragmatics-prosody-bound type: one pragmatic unit and one intonation unit correspond to two or more syntactic dependency clauses; • pragmatics-bound type: one pragmatic unit corresponds to two or more syntactic dependency clauses and intonation units; • prosody-bound type: one intonation unit corresponds to two or more syntactic dependency clauses and pragmatic units. there are several other approaches to defining and identifying discourse units (du); however, each differs in some way from our approach. for instance, prévot et al. (2015) attempted to analyse the discourse structures of spoken french and mandarin chinese by breaking up the discourse into utteranceor clause-like units. first, following stede’s (2012) elementary discourse unit definition, the syntactic information was used to obtain the semantic units, under the belief that the segment ‘is usually a clause, but in general, ranging from minimally a (nominalisation) np to maximally a sentence. it denotes a single event or type of events, serving as a complete, distinct unit of information that the subsequent discourse may connect to’ (p. 89). then, the segment was further refined based on the discourse and pragmatic completion considerations proposed in ford and thompson (1996). while the segmentation in prévot et al. (2015) shows similarities with ours, their operationalisation of prosodic boundaries was different from ours, thus resulting in different unit structures. further, biber et al. (2021) and egbert et al. (2021) both proposed a new method to identify du based on their communicative purposes and claimed that du had identifiable boundaries at which interlocutors shift to different communicative goals, that is, each discourse unit has one major communicative goal, for which they listed nine basic communitive goals, such as sharing feelings and evaluations, giving advice and instructions, and engaging in conflict. even if they specifically emphasised the importance of taking the communicative action into account when segmenting du, their “single communicative goal” concept was different from the pragmatic unit in the current study, which is defined as one step in implementing the action plan. thus, biber et al. (2021) and egbert et al.’s (2021) ‘single communicative goal’ is similar to our ‘action plan’ (also see the idea of ‘project’ proposed by levinson [2013]). still, it is important to acknowledge that our ‘action plan’ is generated in one interlocutor’s mind only, while their ‘communicative goal’ must be accomplished by all dialogic parties. accordingly, they segmented the discourse into a series of ‘topic-like units’ (prévot et al., 2015: 74) that included all interlocutors’ utterances. 3 in theory, mapping the three dimensions should have yielded seven cdu types; however, we failed to observe the ‘twoone-one’ type (i.e. the ‘syntax-prosody-bound’ type, with two or more pragmatic units, one syntactic unit, and one prosodic unit) and the ‘two-one-two’ type (i.e. the ‘syntax-bound’ type, with one dependent clause and two or more pragmatic and intonation units). given that one pragmatic unit corresponds to one step in the implementation of the communicative plan, and one intonation unit corresponds to one idea (chafe, 1987, 1994), it is unsurprising that we could not identify any of the syntax-prosody-bound type, in which interlocutors attempt to stuff two ideas into one intonation unit. however, the syntax-bound type might still be observed in other datasets. the conversational discourse unit: identification and its role in turn-taking management 97 type example (pragmatic unit: // //, syntactic unit: < >, intonational unit: { }) count congruent (one-one-one) 1,100 pragmatics-syntaxbound (one-one-two) 284 pragmatics-prosodybound (one-two-one) 351 pragmatics-bound (one-two-two) 547 prosody-bound (two-two-one) 37 table 5. types and distribution of conversational discourse units. 3.3 turn-taking annotation: floor-taking-turn vs. non-floor-taking-turn speakership alternation in conversation can be roughly divided into two types. the first one is the so-called floor-taking turn, which this study is restricted to. it is fairly easy to operationalise as a turn that is initiated by another speaker. the second one is the non-floor-taking-turn (ford & thompson, 1996), manifested mainly in the form of backchannels (schegloff, 1982), either with one-word particles such as yeah, wow, uh-hum or idiom-like usage such as i see, that’s right and for sure. they are produced to either demonstrate the listener’s acknowledgement or to signal that the listener is paying attention to the speaker, but does not wish to take the floor. we acknowledge that backchannels play a critical role in managing the conversation (jefferson, 1978). however, since interlocutors in those cases do not seem to genuinely take a turn (cf. ford & thompson, 1996), it is assumed that the mechanism of producing them is different from the production of the floor-taking turn (levinson & torreira, 2015), and might even go through a radically different mechanism in language production (macwhinney, 2008; snider & arron, 2012). that is why we excluded them from the present study. next to backchannels, we also do not consider joint-turn constructions as a floor-taking-turn. this decision is based on the observation that all 10 cases of joint-turn construction in the dataset resulted from the listener’s successful prediction of what the speaker is going to say. in such cases, where the listener joins the speaker in uttering the upcoming information, we do not consider them to purposefully take the turn, rather the joint-turn construction is seen as an expression of acknowledgment and understanding. one such case appears below in extract 6. extract 6 (du bois et al., 2000–2005) 1 alice: and with all the breaks that they’ve gotten, 2 they're never gonna have hard times. 3 → mary: //<{hard times do train you}>//. 4 alice: yep. 5 mary: they do. 1 alice: yeah. 2 → //<{ron was singlehandedly there}. 3 → mary: {with one wrecker}>//. 4 alice: yeah= 1 michael: we have very little control over it. 2 but once we do= 3 we’ll be able to program biology as well. 4 → jim: //<{well> //. 5 michael: it is frightening but, 6 jim: we can't even control our freeways 1 → michael: //<{and}>, 2 → <{it might have started a chain reaction that just blew up [the whole earth]}>//. 3 jim: [that's right]. 4 and it might very well still have started that chain reaction. 1 mary: what is it. 2 alice: norplant? 3 → mary: //<{oh>// ////. hu and degand 98 1 richard: i don't know if the parents a=re awa=re, 2 that we did, 3 you know, 4 fred: [break up]? 5 richard: [separate]. 3.4 reliability check the three-dimensional segmentation was manual, following the criteria presented in section 3.2. in order to ensure the soundness of this manual process, several reliability checks were performed. first, in line with spooren and degand’s (2010) recommendation that annotation and coding requires training, the segmentation process was conducted twice on part of the dataset, namely the entire task-oriented dialogue, with a six-week interval between the annotations. intra-coder reliability was performed by the first author. it was calculated on two episodes randomly selected from the dialogue at the interval from the sixth to the twelfth minute and from the twenty-fourth to the thirtieth minute. this examination yielded cohen’s kappa reliability values of 0.91 for the syntactic segmentation, 0.70 for the prosodic segmentation and 0.75 for the pragmatic units, indicating high reliability of our operationalisation. we then turned to the segmentation of the five natural sbc conversations following the same operationalisation. here too, we operated a reliability check, this time between two coders (intercoder reliability rather than intra-coder reliability). one of the coders was the first author, who coded the full dataset. the other coder, a student in linguistics who was blinded to the purpose of this study, was trained to identify the syntactic units, pragmatic units, and turn-taking. following the training, an inter-coder reliability check was performed for 25% (28 minutes) of the conversations, yielding cohen kappa’s reliability scores of 0.84 for turn-taking identification, 0.91 for the identification of syntactic units and 0.94 for pragmatic units, indicating almost perfect agreement for turn-taking annotation and the identification of syntactic and pragmatic units (landis & koch, 1977). note that the intonation units were already present in the original data set and were not reannotated (see section 3.1). all cases of disagreement regarding the syntactic and pragmatic segmentation and the identification of turn-taking were resolved individually through discussion. these concerned 52 out of 1344 syntactic segmentation cases, where a case refers to one boundary, 42 pragmatic cases out of 1125 boundaries, and 22 turn-takings out of 274 cases, where a case stands for the place where the “other speaker” produces verbal content. 4 results this section first shows the association between turn-taking and the three types of linguistic boundaries. it then presents the influence that different types of cdu have on the transition speed. note that the analyses have been conducted based on the five sbc conversations. 4.1 syntactic, prosodic and pragmatic boundaries in turn-taking a total number of 5006 boundaries and 507 turn-taking cases were identified in the dataset. still, 21 turn-taking cases were excluded from further analysis either because unclear speech made identifying the boundary in each dimension too challenging, or because of difficulties in identifying the utterance to which the speaker was replying. among the 486 remaining valid turn-taking cases, we observed five that occurred when none of the three boundaries were present. we kept them in the analyses. due to issues with multicollinearity and a low turn-taking incidence rate, we could not perform logistic regression. three separate chi-square tests of independence were thus performed to examine the relation between the status (i.e. present vs. absent) of the three types of linguistic boundary and the occurrence of turn-taking (i.e. occurred vs. did not occur). there was a statistically significant relationship between the status of the syntactic boundary and occurrence of the conversational discourse unit: identification and its role in turn-taking management 99 turn-taking, χ2(1, n = 5011) = 93.79, p < .001. that is, turn-taking was more likely to be observed immediately after a syntactic boundary compared to locations where a syntactic boundary is absent. identical results were observed for the prosodic boundary, χ2(1, n = 5011) = 131.30, p < .001 and pragmatic boundary, χ2(1, n = 5011) = 423.51, p < .00l. in addition, as table 6 shows, the conditional proportions of turn-taking are 11.6%, 12.6% and 18.8% when syntactic, prosodic and pragmatic boundaries are present, respectively. each of these is significantly higher than the corresponding conditional proportions (i.e. syntactic: 1.2%, prosodic: 1.7% and pragmatic: 1.6%) for absence of the boundary. according to the tests and description of the proportions, we found that the pragmatic boundary is most likely among the three types to be associated with turn-taking. the syntactic boundary is least likely to be associated with turn-taking, and the prosodic boundary is in between.4 turn-taking (occurred) χ2 df p syntactic boundary absent count 11 93.79 1 <.001 % within syntactic boundary 1.2% present count 475 % within syntactic boundary 11.6% prosodic boundary absent count 22 131.30 1 <.001 % within prosodic boundary 1.7% present count 464 % within prosodic boundary 12.6% pragmatic boundary absent count 41 423.51 1 <.001 % within pragmatic boundary 1.6% present count 445 % within pragmatic boundary 18.8% table 6. chi-square tests: association between linguistic boundaries and turn-taking. 4.2 transition speed after each type of conversational discourse unit transition speed is defined as the temporal interval between the offset of the speaker’s current turn and the onset of the listener’s response (see ‘offset 2’ as defined by kendrick & torreira, 2015: 266). analyses included the cdus after which turn-taking took place, yet a number of tokens were removed from the analyses for different reasons. among the 482 valid turn-taking cases, 24 cases of turn-taking that did not take place at the boundary of a cdu (with three coinciding boundaries) were ruled out. another 49 cases were excluded, because they were followed by a pause longer than 1,500 ms (cf. kendrick, holler & levinson, 2023; templeton et al., 2023). the reason is that such longer transition time probably involved other factors than mere prediction, such as reaction behaviour (duncan, 1972) and/or preference issues (kendrick & torreira, 2015). moreover, a further 46 cases, were excluded from analysis because the turns’ exact boundaries were difficult to identify due to low recording quality and chaotic overlap. finally, we also excluded the 8 cases of the prosody-bound type, because they were too rare. we thus ended up with 355 valid turn-taking cases involving the remaining four cdu types. to assess the impact of the type of conversational discourse units on transition speed following the boundary of each cdu type, we conducted a repeated-measures analysis of variance 4 it should be acknowledged that the number of conversations is small in the current analysis. it may spark concern on the violation of the assumption of independence. therefore, we abstained from making an overly assertive and indisputable conclusion. further examinations with larger datasets are warranted. hu and degand 100 (anova). table 7 presents the means and standard deviations for the gap after each type of cdu vis-à-vis each participant (speaker). mauchly’s test confirmed that the assumption of sphericity had been met (x2 (5) = 8.101, p = .155). the effect of the cdu type was statistically significant at the 5% level of significance, f(3, 24) = 4.33, p = .014, with a partial η2 of .351. however, speaker 4 was excluded from our analysis because the transition speeds after the pragmatics-syntax-bound cdu and the pragmatics-bound cdu were more than two times the standard deviation (2 sd) of each corresponding mean. post-hoc pairwise comparisons adjusted using the bonferroni method revealed no significant difference between the mean transition speeds after the boundary of the pragmatics-syntax-bound cdu and those after the boundary of the pragmatics-bound cdu (p = .991, 95% c.i. = [−358.03, 139.56]). the average transition speed following the boundary of the pragmatics-syntax-bound cdu and pragmatics-bound cdu indicated a marginally significant difference (p = .052, 95% c.i. = [−392.36, 1.53]). notably, the mean transition speed after the boundary of the pragmatics-syntaxbound cdu was significantly higher than that after the boundary of the pragmatics-prosody-bound cdu (p = .018, 95% c.i. = [−401.51, −38.40]). finally, there were no statistically significant differences among the congruent, pragmatics-prosody-bound cdu (p = 1.00, 95% c.i. = [−279.24, 230.16]) and pragmatics-bound cdu (p = .52, 95% c.i. = [−67.31, 239.67]), as illustrated in table 8. the conversational discourse unit: identification and its role in turn-taking management 101 c o n g ru en t p ra g m at ic ssy n ta x b o u n d p ra g m at ic sp ro so d y b o u n d p ra g m at ic sb o u n d s p ea k er m s d m s d m s d m s d 1 3 9 9 .6 4 2 9 3 .0 3 1 0 .2 2 4 2 8 .7 0 2 2 1 .6 7 3 0 0 .1 3 2 1 9 .5 0 4 1 0 .8 9 2 3 4 1 .7 5 5 5 3 .4 4 4 7 6 .0 0 4 7 0 .8 8 7 6 4 .8 3 4 8 5 .3 1 2 2 8 .0 8 4 1 8 .1 7 3 6 6 2 .0 5 4 2 4 .9 4 4 0 5 .5 0 5 4 .4 8 6 8 9 .7 5 4 0 7 .1 1 4 3 8 .6 7 2 6 0 .7 1 4 5 3 1 .6 8 3 1 6 .5 3 9 2 3 .0 0 . 6 4 5 .8 0 5 0 7 .5 3 1 3 4 3 .5 0 7 0 .0 0 5 6 2 3 .0 4 3 7 4 .1 7 2 3 8 .0 0 . 7 7 9 .2 0 4 5 6 .2 8 4 9 2 .0 0 1 8 3 .2 0 6 7 1 3 .4 1 4 0 9 .6 6 3 7 7 .7 5 3 0 3 .8 8 4 0 4 .3 7 3 7 2 .5 3 7 2 2 .0 0 6 3 .6 4 7 3 2 3 .8 0 3 3 5 .8 9 2 1 0 .0 0 1 6 8 .7 7 3 7 6 .7 5 1 6 0 .5 7 3 0 1 .2 1 1 8 9 .5 1 8 3 5 8 .9 4 4 2 9 .1 8 2 8 7 .8 8 2 5 4 .1 1 4 5 7 .2 5 4 6 4 .5 0 7 7 .0 0 3 5 0 .9 6 9 2 0 2 .2 9 2 4 5 .1 7 6 1 .0 0 1 6 .9 7 3 2 9 .2 5 8 1 .0 3 2 7 4 .7 5 5 5 2 .5 0 1 0 3 7 4 .6 3 4 1 6 .6 0 1 7 4 .4 6 2 2 7 .0 7 1 9 7 .3 3 1 1 1 .4 2 4 7 0 .7 1 3 3 0 .5 4 t o ta l 4 7 1 .9 3 4 0 6 .0 5 2 3 3 .0 4 3 2 4 .5 5 4 9 5 .5 7 4 0 2 .7 6 4 7 0 .7 1 3 3 0 .5 4 2 s d o f th e m [− 3 4 0 .1 7 , 1 2 8 4 .0 3 ] [− 4 1 9 .0 6 , 8 8 2 .1 4 ] [− 3 0 9 .9 5 , 1 3 0 1 .0 9 ] [− 4 5 3 .2 6 , 1 1 4 6 .3 8 ] t a b le 7 . s p ea k er s’ t ra n si ti o n s p ee d s fo ll o w in g e ac h t y p e o f c d u . c o n g ru en t p ra g m at ic ssy n ta x b o u n d p ra g m at ic sp ro so d y b o u n d p ra g m at ic sb o u n d d ep en d en t v ar ia b le f (3 , 2 4 ) p η 2 m s d m s d m s d m s d g ap 4 .3 3 .0 1 4 .3 5 1 4 4 4 .3 9 a 1 7 6 .6 1 2 4 8 .9 8 a 1 5 5 .3 9 4 6 8 .9 3 b 2 2 3 .5 5 3 5 8 .2 1 a, b 1 9 1 .9 8 t a b le 8 . o n ew ay r ep ea te d -m ea su re s a n o v a : t h e ef fe ct o f th e ty p e o f c d u o n t h e tr an si ti o n s p ee d a ft er t h e b o u n d ar y o f ea ch c d u t y p e. n o te . m ea n s n o t sh ar in g s u b sc ri p ts a re s ig n if ic an tl y d if fe re n t fr o m o n e an o th er . hu and degand 102 5 general discussion a close examination of the role that cdus play in the turn-taking operation found a significant association between the presence of three linguistic boundaries and occurrence of turn-taking. additionally, the analysis based on the current dataset suggested the duration of the transition gap between turns was influenced by the cdu type. in this section, we discuss the possible reasons underpinning these observations. 5.1 the importance of meaning-connection in managing turn-taking the comparison of the conditional proportions of turn-taking in presence or absence of syntactic, prosodic, and pragmatic boundaries led to the conclusion that the pragmatic boundary is the best predictor for turn-taking, whereas the syntactic boundary was the worst. the prosodic boundary was in between. this observation aligns with schegloff (2007): turn-at-talk is organised not just by grammar and phonetic properties, but also by social action. conversation emerges from action formation and ascription (levinson, 2013). this perspective has received support from several studies in cognitive neuroscience and psycholinguistics, which have indicated that interlocutors allocate resources to process the actions of their conversational partners. for instance, neuroscientific studies have indicated that addressees attempt to predictively identify a speech act before even hearing it (gisladottir, chwilla & levinson, 2015; gisladottir, bögels & levinson, 2018). in addition to predicting speech acts, listeners show differences in how direct and indirect speech acts are processed (clark & lucy, 1975; coulson & lovett, 2010). considering this together with the pragmatic unit, defined as one step of implementing an action plan and identified by analysing the meaning the speaker is attempting to convey within the situated context, there is a straightforward connection to the ‘meaning’ the speaker is attempting to convey. in other words, rather than following syntactic and prosodic structures per se, interlocutors predominantly follow the meaning their dialogic partner conveys. this idea is consistent with recent studies on joint-turn construction, which is traditionally defined as one participant completing a syntactic unit initiated by another participant (lerner & takagi, 1999; suzuki & usami, 2005). laury and ono (2021) found that if the two parts of the joint-turn construction are combined, they often do not form a syntactically acceptable construction. rather, the two parts are semantically conjoined—the second speaker produces an utterance based on their understanding of prior utterances. therefore, interlocutors may not jointly construct syntactic units but rather connect parts of joint-turn construction via meaning. when discussing meaning connection, it is important to consider views on the connection between intonation and meaning expression. chafe (1987, 1994) proposed that the intonation unit corresponds to a new complete idea. following this view, the association between prosodic boundaries and turn-taking should have been similar to that between the pragmatic boundary and turn-taking. yet, chafe’s proposal is based on an ideal situation where the pause between two intonation units denotes the change in information activation in the speaker’s mind. in other words, activated information becomes deactivated during the pause, whereas inactive information becomes activated (chafe, 1987; 1994). however, according to our observation, a considerable number of pauses resulted from difficulties in verbalisation as well as breath and laughter. that is, the segment before the pause/boundary failed to convey a complete idea, and such pauses might thus fail to lead to a turn-transition. while prosodic boundaries may be the most frequent type of boundary, this overall frequency decreases its degree of specific association with turn transitions. nevertheless, compared to the syntactic boundary, prosodic boundaries showed slightly stronger associations with turn-transition occurrence. this aligns with studies wherein prosody is considered critical to the timing of turn-taking (bögels & torreira, 2015, 2021) but contradicts others wherein syntax is argued to play a dominant role in turn-taking (de ruiter, mitterer & enfield, 2006; sacks, schegloff & jefferson, 1974). importantly, the difference between these two factors was not substantial; thus, it is worth reiterating that these two factors may produce only the conversational discourse unit: identification and its role in turn-taking management 103 trivial independent effects (auer, 1996; selting, 2000). this largely resonates with levinson and torreira’s (2015) proposal for a psycholinguistic model of turn-taking, where they argued that: morphosyntax may provide most of the early clues to the overall structural envelope …, so offering some long distance projection. within the last half second or so, the actual words will often be predicted, and, within that same late time frame, cues to imminent turn closure, usually prosodic and phonetic, are likely to appear, indicating a likely turn end. (p. 13) it should also be noted that the weight of syntax and prosody in turn-taking may differ across languages as speakers of different languages may rely to different degrees on syntax and prosody in language processing. evidence comes from unrelated but noteworthy studies on syllable effects (mehler et al., 1981), where a target syllable is detected faster when it precisely matches the first syllable of the carrier (e.g. ba in ba.lance) as compared to when it does not (e.g. ba in balcon). studies have found that this effect largely depends on whether the particular language is a stress language (cutler, 1997). 5.2 relation between type of cdu and transition speed: earlier emergence of the gist we should acknowledge that turn-taking is not simply about taking the turn at the proper place. the average transition speeds after the four types of cdus in our dataset are all shorter than 500 ms in average, faster than the time required to produce a single word (indefrey & levelt, 2004). in other words, interlocutors, in general, prepare their response before the arrival of the turn-end. meanwhile, in our analysis, even if all turn-taking occurred at the place where syntactic, prosodic and pragmatic boundaries converge, the transition speeds to certain extent differ across the cdu types. it is particularly noteworthy that transition speed after the pragmatics-syntax-bound cdu, i.e. [1 pragm./1 synt./2 inton.], is the fastest (on average), which differs from that of the pragmaticsprosody-bound [1 pragm./2 synt./ 1. inton.] cdu. uncovering the precise reasons for these differences is beyond the scope of this study. however, with the cdu segmentation in hand we would like to suggest a number of possible directions for future research. despite the discrepancy of when people begin to prepare their utterance (bögels, magyari & levinson, 2015; boiteau et al., 2014; piai et al., 2015; sjerps & meyer, 2015), mounting evidence has indicated that ideally the semantic direction of an utterance is clarified as early as possible (corps et al., 2018) so that speakers can anticipate the upcoming content, thus facilitating turntaking. in other words, the earlier the emergence of the gist5 of an utterance, the earlier the listener starts to prepare a response and the faster the transition speed. in this regard, it is reasonable to believe that the pragmatics-syntax-bound type possesses specific linguistic features that other types do not have to make its gist become clear earlier than others. it is notable that english fronts some key (lexico)syntactic characteristics (levinson, 2013) that allow speakers to reveal their gist earlier than those languages users whose language provides the key (lexico)syntactic characteristics relatively late. levinson (2013) highlighted that english tends to shift its question words (e.g. wh-words) right up to the front of an utterance, so that listeners can recognise the information-seeking actions earlier (levinson, 2013). for instance, when trying to find a key, an english speaker may say ‘where is my key?’; meanwhile, the same meaning should be formulated as ‘my key where’ in mandarin chinese. therefore, the english listener may recognise the speaker’s information-seeking actions earlier than the chinese listener under the same situation.6 similar observations were made by ono and thompson (2017) regarding the earlier occurrence of the negative morpheme in natural english conversation than in japanese. 5 we use ‘gist’ to refer to either the communicative action, such as interrogative or statement, the speaker is going to conduct or the concrete content their utterance is designed to convey. 6 different languages could have different ways of front-loading the same gist (e.g. information-seeking). it is possible for chinese users to have recognised the information-seeking action as early as english users did by drawing on other linguistic resources in addition to syntax, such as prosody (cf. shen, 1990). hu and degand 104 consequently, they concluded that english speakers do not need to wait until the end of an utterance to recognise the negative semantic implication of the entire utterance. inspired by these studies, it is reasonable to speculate that the end point of the main syntactic part of the pragmatic-syntax-bound cdu occurs earlier than that in the other three types of cdus. to this end, we considered the duration between the end of the head of the cdu (usually the predicate) and the end of the cdu.7 figure 5 displays that the post-predicate duration of the pragmatic-syntax-bound cdu is longer than that of the other three types. the raw descriptive data refrain us from making any cogent conclusion, but to a certain extent it indicates that the syntactic structure of the pragmatic-syntax-bound cdu may help the listener to grasp its gist earlier than in the other types. being aware that our measurement is oversimplified, further scrutiny is needed regarding the influence of the length (magyari & de ruitter, 2012), prosody (garrod & pickering, 2015) and transitivity (tao & hu, 2019; thompson & hopper, 2001) of the cdu. figure 5. post-predicate duration of cdu (dark gray part). 6 conclusion the aim of this study was to extend our understanding of the building blocks of conversational discourse and the role they play in turn transition. we propose that the discourse unit should be an amalgam of three types of units: syntactic (dependency) clauses, intonation units and pragmatic actions. the mapping of these three types of segments form conversational discourse units that are in line with human communication in which interlocutors exploit linguistic information to convey and interpret one another’s intentions. such mapping provides a linguistic perspective for defining the analytical conversational unit, which is considered basic for organising conversational turn-taking. based on the investigation of natural conversation, this study demonstrated that compared to prosodic and syntactic boundaries, interlocutors are subject to take turns immediately after a full pragmatic unit. this result indicates that meaning-connection and action ascription underpin the management of turn-taking. meanwhile, transition speed after the pragmatic-syntaxbound cdu denotes a statistical difference in one of the three remaining types of cdu. this finding may be attributed to the linguistic structure of the pragmatic-syntax-bound cdu, which enables interlocutors to understand its gist earlier than they do for other types of cdu. the investigation of the management of turn-taking began with identifying cdus. the question may arise why we would need a new concept, when tcu-based analysis is a well 7 the post-predicate duration of the congruent cdu was calculated based on 80 cases randomly selected from the 130 congruent cdu cases. the conversational discourse unit: identification and its role in turn-taking management 105 established method for investigating turn-taking. clearly, this study does not intend to disregard or dismiss the tcu framework. rather, it seeks to provide an additional perspective to understand turn-taking. more importantly, as a linguistic unit, cdu enriches our toolkit not only for exploring the turn-taking operation but also for extending our understanding of language production and comprehension in several other ways, such as exploring the possibility of speaker-hearer alignment at the discourse level, uncovering cross-linguistic differences in conversational structure or evaluating the impact of the communicative context on the flow of discourse. our segmentation extends the work of degand and simon (2009), whose study was restricted to the interplay between the prosodic and syntactic dimensions in defining units of discourse segmentation. to the best of our knowledge, this study is the first to incorporate the pragmatic dimension into discourse segmentation in a (at least fairly) objective way. it overcomes the longtime criticism on the subjectivity and ambiguity of the definition of pragmatic completeness (skantze, 2021). yet, the analysis of the situated context underlying the identification of the pragmatic units being based mainly on semantic understanding of the utterance, further work is needed on identifying linguistic (surface) rules related to pragmatic boundaries, so that our cdu proposal could be applied to a larger data analysis, thus gaining generalisation. acknowledgements this project has received funding from the european union’s horizon 2020 research and innovation programme under grant agreement n°859588. we are grateful for detailed and constructive comments from three anonymous reviewers. all remaining errors are ours. references stergos afantenos, nicholas asher, farah benamara, myriam bras, cécile fabre, mai ho-dac, anne le draoulec, philippe muller, marie-paule péry-woodley, laurent prévot, josette rebeyrolles, ludovic tanguy, marianne vergez-couret, laure vieu (2012). an empirical resource for discovering cognitive principles of discourse organisation: the annodis corpus. in the proceedings of the eighth international conference on language resources and evaluation (lrec’ 12), istanbul, turkey. http://www.lrecconf.org/proceedings/lrec2012/pdf/836_paper.pdf nicholas asher and alex lascarides (2003). logics of conversation. cambridge university press, cambridge, east of england. gerry t. m. altmann and jelena mirković (2009). incrementality and prediction in human sentence processing. cognitive science, 33:583–609. https://doi.org/10.1111%2fj.15516709.2009.01022.x peter auer (1996). on the prosody and syntax of turn-continuations. in elizabeth couper-kuhlen and margret selting, editors, prosody in conversation: interactional studies, pages 57–100. cambridge university press, new york city, new york. peter auer (2005). projection in interaction and projection in grammar. text, 25(1):7–36. https://doi.org/10.1515/text.2005.25.1.7 brendan barnwell (2013). perception of prosodic boundaries by untrained listeners. in beatrice szczepek reed and geoffrey raymond, editors, units of talk – units of action, pages 125– 165. john benjamins publishing company, amsterdam/philadelphia, noordholland/pennsylvania. dagmar barth-weingarten (2013). from “intonation units” to censuring – an alternative approach to the prosodic-phonetic structuring of talk-in-interaction. in beatrice szczepek reed and geoffrey raymond, editors, units of talk – units of action, pages 91–124. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. http://www.lrec-conf.org/proceedings/lrec2012/pdf/836_paper.pdf http://www.lrec-conf.org/proceedings/lrec2012/pdf/836_paper.pdf https://doi.org/10.1111%2fj.1551-6709.2009.01022.x https://doi.org/10.1111%2fj.1551-6709.2009.01022.x https://doi.org/10.1515/text.2005.25.1.7 hu and degand 106 douglas biber, jesse egbert, daniel keller, stacey wizner (2021). towards a taxonomy of conversational discourse types: an empirical corpus-based analysis. journal of pragmatics, 171:20–35. https://doi.org/10.1016/j.pragma.2020.09.018 sara bögels (2020). neural correlates of turn-taking in the wild: response planning starts early in the free interviews. cognition, 203:104347. https://doi.org/10.1016/j.cognition.2020.104347 sara bögels, lilla magyari and stephen c. levinson (2015). neural signatures of response planning occur midway through an incoming question in conversation. scientific reports, 5:12881. https://doi.org/10.1038/srep12881 sara bögels and francisco torreira (2015). listeners use intonational phrase boundaries to project turn ends in spoken interaction. journal of phonetics, 52:46–57. https://doi.org/10.1016/j.wocn.2015.04.004 sara bögels and francisco torreira (2021). turn-end estimation in conversational turn-taking: the roles of context and prosody. discourse processes, 58(10):903–924. https://doi.org/10.1080/0163853x.2021.1986664 timothy w. boiteau, patrick s. malone, sara a. peters and amit almor (2014). interference between conversation and a concurrent visuomotor task. journal of experimental psychology: general, 143(1):295–311. https://doi.org/10.1037/a0031858 galina b. bolden (2009). implementing incipient actions: the discourse marker so in english conversation. journal of pragmatics, 41(5):974–998. https://doi.org/10.1016/j.pragma.2008.10.004 brentari diane, joshua falk, anastasia giannakidou, annika herrmann, elisabeth volk and markus steinbach (2018). production and comprehension of prosodic markers in sign language imperatives. frontiers in psychology, 9:770. https://doi.org/10.3389/fpsyg.2018.00770 andrew carnie (2021). syntax: a generative introduction (fourth edition). wiley-blackwell, chichester, west sussex. wallace chafe (1987). cognitive constraints on information flow. in russell s. tomlin, editor, coherence and grounding in discourse, pages, 21–51. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. wallace chafe (1993). prosodic and functional units of language. in jane a. edwards and martin d. lampert, editors, talking data: transcription and coding in discourse research, pages 33–43. lawrence erlbaum, hillsdale, new jersey. wallace chafe (1994). discourse, consciousness, and time: the flow and displacement of conscious experience in speaking and writing. university of chicago press, chicago, illinois. herbert h. clark (1996). using language. cambridge university press, cambridge, east of england. herbert h. clark and peter lucy (1975). understanding what is meant from what is said: a study in conversationally conveyed requests. journal of verbal learning and verbal behavior, 14:56–72. https://doi.org/10.1016/s0022-5371(75)80006-5 ruth e. corps, abigail crossley, chiara gambi and martin j. pickering (2018). early preparation during turn-taking: listeners use content predictions to determine what to say but not when to say it. cognition, 175:77–95. https://doi.org/10.1016/j.cognition.2018.01.015 seana coulson and christopher lovett (2010). comprehension of non-conventional indirect requests: an event-related brain potential study. italian journal of linguistics, 22(1):107– 124. https://www.italian-journal-linguistics.com/app/uploads/2021/05/6_coulsonlovett.pdf elizabeth couper-kuhlen (2006). prosodic cues of discourse units. in keith brown, editor, encyclopedia of language & linguistics (second edition), pages 178–182. elsevier, amsterdam, noord-holland. https://doi.org/10.1016/b0-08-044854-2/00588-5. anne cutler (1997). the syllable’s role in the segmentation of stress languages. language and cognitive processes, 12:839-845. https://doi.org/10.1080/016909697386718 https://doi.org/10.1016/j.pragma.2020.09.018 https://doi.org/10.1016/j.cognition.2020.104347 https://doi.org/10.1038/srep12881 https://doi.org/10.1016/j.wocn.2015.04.004 https://doi.org/10.1080/0163853x.2021.1986664 https://doi.org/10.1037/a0031858 https://doi.org/10.1016/j.pragma.2008.10.004 https://doi.org/10.3389/fpsyg.2018.00770 https://doi.org/10.1016/s0022-5371(75)80006-5 https://doi.org/10.1016/j.cognition.2018.01.015 https://www.italian-journal-linguistics.com/app/uploads/2021/05/6_coulsonlovett.pdf https://doi.org/10.1016/b0-08-044854-2/00588-5 https://doi.org/10.1080/016909697386718 the conversational discourse unit: identification and its role in turn-taking management 107 liesbeth degand (2019). review of the book spontaneous spoken english: an integrated approach to the emergent grammar of speech by a. haselow. language, 95(1):185–187. https://doi.org/10.1353/lan.2019.0018 liesbeth degand and ludivine crible (2021). discourse markers at the peripheries of syntax, intonation and turns. towards a cognitive-functional unit of segmentation. in daniël van olmen and jolanta šinkūnienė, editors, pragmatic markers and peripheries, pages 19–48. john benjamins publishing company, amsterdam/philadelphia, noordholland/pennsylvania. https://doi.org/10.1075/pbns.325.01deg liesbeth degand and anne-catherine simon (2009). on identifying basic discourse units in speech: theoretical and empirical issues. discours: revue de linguistique, psycholinguistique et informatique. https://doi.org/10.4000/discours.5852 liesbeth degand, anne-catherine simon, noalig tanguy and thomas van damme (2014). initiating a discourse unit in spoken french: prosodic and syntactic features of the left periphery. in salvador pons bordería, editor, discourse segmentation in romance languages, pages 243–273. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. https://doi.org/10.1075/pbns.250.09deg marie-catherine de marneffe and joakim nivre (2019). dependency grammar. annual review of linguistics, 5:197–218. https://doi.org/10.1146/annurev-linguistics-011718-011842 jan peter de ruiter, holger mitterer and nick j. enfield (2006). projecting the end of a speaker’s turn: a cognitive cornerstone of conversation. language, 82:515–535. https://doi.org/10.1353/lan.2006.0130 paul drew (2013). turn design. in tanya stivers and jack sidnell, editors, the handbook of conversation analysis, pages 131–149. wiley-blackwell, malden, massachusetts. john w. du bois, wallace l. chafe, charles meyer, sandra a. thompson, robert englebretson and nii martey (2000–2005). santa barbara corpus of spoken american english, parts 1– 4. linguistic data consortium. https://www.linguistics.ucsb.edu/research/santa-barbaracorpus#access john w. du bois, stephan schuetze-coburn, susanna cumming and danae paolino (1993). outline of discourse transcription. in jane a. edwards and martin d. lampert, editors, talking data: transcription and coding methods for language research, pages 45–89. lawrence erlbaum, mahwah, new jersey. starkey duncan (1972). some signals and rules for taking speaking turns in conversation. journal of personality and social psychology, 23(2):283–292. https://psycnet.apa.org/doi/10.1037/h0033031 lotte eijk, marlou rasenberg, flavia arnese, mark blokpoel, mark dingemanse, christian f. doeller, mirjam ernestus, judith holler, branka milivojevic, asli özyürek, wim pouw, iris van rooij, herbert schriefers, ivan toni, james trujillo and sara bögels (2022). the cabb dataset: a multimodal corpus of communicative interactions for behavioural and neural analyses. neuroimage, 264:119734. https://doi.org/10.1016/j.neuroimage.2022.119734 jesse egbert, stacey wizner, daniel keller, douglas biber, tony mcenery and paul baker (2021). identifying and describing functional discourse units in the bnc spoken 2014. text & talk, 41(5-6):715–737. https://doi.org/10.1515/text-2020-0053 nick j. enfield and jack sidnell (2017). on the concept of action in the study of interaction. discourse studies, 19(5):515–535. https://doi.org/10.1177%2f1461445617730235 michael c. ewing (2019). the predicate as a locus of grammar and interaction in colloquial indonesian. studies in language, 43(2):402–443. https://doiorg.ezproxy.is.ed.ac.uk/10.1075/sl.17062.ewi laurent filliettaz and eddy roulet (2002). the geneva model of discourse analysis: an interactionist and modular approach to discourse organization. discourse studies, 4(3), 369– 393. https://doi.org/10.1177/14614456020040030601 https://doi.org/10.1353/lan.2019.0018 https://doi.org/10.1075/pbns.325.01deg https://doi.org/10.4000/discours.5852 https://doi.org/10.1075/pbns.250.09deg https://doi.org/10.1146/annurev-linguistics-011718-011842 https://doi.org/10.1353/lan.2006.0130 https://www.linguistics.ucsb.edu/research/santa-barbara-corpus#access https://www.linguistics.ucsb.edu/research/santa-barbara-corpus#access https://psycnet.apa.org/doi/10.1037/h0033031 https://doi.org/10.1016/j.neuroimage.2022.119734 https://doi.org/10.1515/text-2020-0053 https://doi.org/10.1177%2f1461445617730235 https://doi-org.ezproxy.is.ed.ac.uk/10.1075/sl.17062.ewi https://doi-org.ezproxy.is.ed.ac.uk/10.1075/sl.17062.ewi https://doi.org/10.1177/14614456020040030601 hu and degand 108 janet dean fodor (2002). prosodic disambiguation in silent reading. north east linguistics society, 32, 113–132. https://scholarworks.umass.edu/nels/vol32/iss1/8?utm_source=scholarworks.umass.edu%2 fnels%2fvol32%2fiss1%2f8&utm_medium=pdf&utm_campaign=pdfcoverpages cecilia e. ford (2004). contingency and units in interaction. discourse studies, 6(1):27–52. https://journals.sagepub.com/doi/pdf/10.1177/1461445604039438 cecilia e. ford, barbara a. fox and sandra a. thompson (1996). practices in the construction of turns: the tcu revisited. pragmatics, 6(3):427–454. https://doi.org/10.1075/prag.6.3.07for cecilia e. ford and sandra a. thompson (1996). interactional units in conversation: syntactic, intonational, and pragmatic resources for the projection of turn completion. in: elinor ochs, emanuel a. schegloff and sandra a. thompson, editors, interaction and grammar, pages 135–184. cambridge university press, cambridge, east of england. bruce fraser (2009). an account of discourse markers. international review of pragmatics, 1:293– 320. simon garrod and martin j. pickering (2015). the use of content and timing to predict turn transitions. frontiers in psychology, 6(751):1–12. https://doi.org/10.3389%2ffpsyg.2015.00751 rosa s. gisladottir, sara bögels and stephen c. levinson (2018). oscillatory brain responses reflect anticipation during comprehension of speech acts in spoken dialogue. frontiers in human neuroscience, 12:34. https://doi.org/10.3389/fnhum.2018.00034 rosa s. gisladottir, dorothee j. chwilla and stephen c. levinson (2015). conversation electrified: erp correlates of speech act recognition in underspecified utterances. plos one, 10(3):e0120068. https://doi.org/10.1371/journal.pone.0120068 paul grice (1975). logic and conversation. in peter cole and jerry l. morgan, editors, studies in syntax and semantics iii: speech acts, pages 183–198. academic press, cambridge, massachusetts. paul grice (1989). studies in the ways of words. harvard university press, cambridge, massachusetts. alexander haselow (2017). spontaneous spoken english: an integrated approach to the emergent grammar of speech. cambridge university press, cambridge, east of england. barbara j. grosz and candace l. sidner (1986). attention, intentions, and the structure of discourse. computational linguistics 12(3): 175-204. https://dash.harvard.edu/handle/1/2579648 mattias heldner, jens edlund, anna hjalmarsson and kornel laskowski (2011). very short utterances and timing in turn-taking. in interspeech 2011, pages 2837–2840, florence, tuscany. hans jürgen heringer (1993). dependency syntax basic ideas and the classical model. in joachim jacobs, arnim von stechow, wolfgang sternefeld and theo vennemann de gruyter, editors, syntax – an international handbook of contemporary research (vol. 1), pages 298–316. walter de gruyter, berlin. john heritage (1984). garfinkel and ethnomethodology. polity, cambridge. john heritage (2013). turn-initial position and some of its occupants. journal of pragmatics, 57:331–337. https://doi.org/10.1016/j.pragma.2013.08.025 john heritage and marja-leena sorjonen (1994). constituting and maintaining activities across sequences: and-prefacing as a feature of question design. language in society, 23:1–29. https://doi.org/10.1017/s0047404500017656 john heritage and marja-leena sorjonen (2018). between turns and sequence: turn-initial particles across languages. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. https://scholarworks.umass.edu/nels/vol32/iss1/8?utm_source=scholarworks.umass.edu%2fnels%2fvol32%2fiss1%2f8&utm_medium=pdf&utm_campaign=pdfcoverpages https://scholarworks.umass.edu/nels/vol32/iss1/8?utm_source=scholarworks.umass.edu%2fnels%2fvol32%2fiss1%2f8&utm_medium=pdf&utm_campaign=pdfcoverpages https://journals.sagepub.com/doi/pdf/10.1177/1461445604039438 https://doi.org/10.1075/prag.6.3.07for https://doi.org/10.3389%2ffpsyg.2015.00751 https://doi.org/10.3389/fnhum.2018.00034 https://doi.org/10.1371/journal.pone.0120068 https://dash.harvard.edu/handle/1/2579648 https://doi.org/10.1016/j.pragma.2013.08.025 https://doi.org/10.1017/s0047404500017656 the conversational discourse unit: identification and its role in turn-taking management 109 judith holler, kobin h. kendrick and stephen c. levinson (2018). processing language in faceto-face conversation: questions with gestures get faster responses. psychonomic bulletin & review, 25(5):1900–1908. https://doi.org/10.3758/s13423-017-1363-z judith holler and stephen c. levinson (2019). multimodal language processing in human communication. trends in cognitive sciences, 23(8):639–652. https://doi.org/10.1016/j.tics.2019.05.006 paul j. hopper and sandra a. thompson (2008). projectability and clause combining in interaction. in ritva laury, editors, crosslinguistic studies of clause combining: the multifunctionality of conjunctions, pages 99–124. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. hanneke houtkoop and harrie mazeland (1985). turns and discourse units in everyday conversation. journal of pragmatics, 9:595–619. https://doi.org/10.1016/03782166(85)90055-4 junfei hu and liesbeth degand (2022). do speakers align at the discourse level? the role of discourse segments in task-oriented dialogue. in the cogling day 2022, tilburg, north brabant. maya inbar, eitan grossman and ayelet n. landau (2020). sequences of intonation units from ~ 1 hz rhythm. nature, 10(1):1–9. https://doi.org/10.1038/s41598-020-72739-4 peter indefrey and willem j. m. levelt (2004). the spatial and temporal signatures of word production components. cognition, 92(1–2):101–144. https://doi.org/10.1016/j.cognition.2002.06.001 shlomo izre’el (2020). the basic unit of spoken language the interfaces between prosody, discourse and syntax: a view from spontaneous spoken hebrew. in shlomo izre’el, heliana mello, alessandro painunzi and tommaso raso, editors, in speech of basic units of spoken language: a corpus-driven approach, pages 77–106. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. shlomo izre’el, heliana mello, alessandro panunzi and tommaso raso (2020). in search of basic units of spoken language: a corpus-driven approach. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. gail jefferson (1978). sequential aspects of storytelling in conversation. in jim schenkein, editor, studies in the organization of conversational interaction, pages 219–248. elsevier inc, amsterdam, noord-holland. https://doi.org/10.1016/b978-0-12-623550-0.50016-1 gail jefferson (1983). on a failed hypothesis: “conjunctionals” as overlap-vulnerable. tilburg papers in language and literature, 28:1–33. kobin h. kendrick, judith holler and stephen c. levinson (2023). turn-taking in human face-toface interaction is multimodal: gaze direction and manual gestures aid the coordination of turn transitions. philosophical transactions of the royal society b, 378(1875):20210473. https://doi.org/10.1098/rstb.2021.0473 kobin h. kendrick and francisco torreira (2015). the timing and construction of preference: a quantitative study. discourse processes, 52(4):255–289. https://doi.org/10.1080/0163853x.2014.955997 jieun kiaer (2014). pragmatic syntax. bloomsbury, london, england. hye ri stephanie kim and satomi kuroshima (2013). turn beginnings in interaction: an introduction. journal of pragmatics, 57:267–273. https://doi.org/10.1016/j.pragma.2013.08.026 sotaro kita and asli özyürek (2003). what does crosslinguistic variation in semantic coordination of speech and gesture reveal? evidence for an interface representation of spatial thinking and speaking. journal of memory and language, 48: 16–32. https://doi.org/10.1016/s0749596x(02)00505-3 j. richard landis and gary g. koch (1977). the measurement of observer agreement for categorical data. biometrics, 33:159–174. https://doi.org/10.2307/2529310 https://doi.org/10.3758/s13423-017-1363-z https://doi.org/10.1016/j.tics.2019.05.006 https://doi.org/10.1016/0378-2166(85)90055-4 https://doi.org/10.1016/0378-2166(85)90055-4 https://doi.org/10.1038/s41598-020-72739-4 https://doi.org/10.1016/j.cognition.2002.06.001 https://doi.org/10.1016/b978-0-12-623550-0.50016-1 https://royalsocietypublishing.org/doi/abs/10.1098/rstb.2021.0473 https://royalsocietypublishing.org/doi/abs/10.1098/rstb.2021.0473 https://royalsocietypublishing.org/doi/abs/10.1098/rstb.2021.0473 https://doi.org/10.1098/rstb.2021.0473 https://doi.org/10.1080/0163853x.2014.955997 https://doi.org/10.1016/j.pragma.2013.08.026 https://doi.org/10.1016/s0749-596x(02)00505-3 https://doi.org/10.1016/s0749-596x(02)00505-3 https://doi.org/10.2307/2529310 hu and degand 110 ritva laury and tsuyoshi ono (2021). co-construction reconsidered: over-syntacticization in interactional linguistics. in the fourth interactional conference on interactional linguistics and chinese language studies, beijing. gene h. lerner and tomoyo takagi (1999). on the place of linguistic resources in the organization of talk-in-interaction: a co-investigation of english and japanese. journal of pragmatics, 31(1):49–75. https://doi.org/10.1016/s0378-2166(98)00051-4 willem j. m. levelt (1989). speaking: from intention to articulation. mit press, cambridge, massachusetts. stephen c. levinson (2013). action formation and ascription. in tanya stivers and jack sidnell, editors, the handbook of conversation analysis, pages 103–130. wiley-blackwell, malden, massachusetts. stephen c. levinson and francisco torreira (2015). timing in turn-taking and its implications for processing models of language. frontiers in psychology, 6: 731. https://doi.org/10.3389/fpsyg.2015.00731 xiaoting li (2016). some interactional uses of syntactically incomplete turns in mandarin conversation. chinese language and discourse, 7(2), 237–271. https://doi.org/10.1075/cld.7.2.03li per linell (2013). the dynamics of incrementation in utterance-building: processes and resources. in beatrice szczepek reed and geoffery raymond, editors, units of talk – units of action, pages 57–89. john benjamins publishing company, amsterdam, noord-holland. lilla magyari and jan peter de ruiter (2012). prediction of turn-ends based on anticipation of upcoming words. frontiers in psychology, 3:367. https://doi.org/10.3389%2ffpsyg.2012.00376 william c. mann and sandra a. thompson (1988). rhetorical structure theory: toward a functional theory of text organization. text, 8(3):243–281. https://doi.org/10.1515/text.1.1988.8.3.243 yael maschler and deborah schiffrin (2015). discourse markers, language, meaning, and context. in deborah tannen, heidi e. hamilton and deborah schiffrin, editors, the handbook of discourse analysis (second edition), pages 189–221. john wiley & sons, chichester, west sussex. jacques mehler, jean yves dommergues, uli frauenfelder and juan segui (1981). the syllable’s role in speech segmentation. journal of verbal learning and verbal behavior, 20:298–305. https://doi.org/10.1016/s0022-5371(81)90450-3 marianne mithun (2020). prosody and the organization of information in central pomo, a california indigenous language. in shlomo izre’el, heliana mello, alessandro painunzi and tommaso raso, editors, in speech of basic units of spoken language: a corpusdriven approach, pages 107–126. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. jacqueline monschau, rolf kreyer and joybrato mukherrje (2003). syntax and semantics at tone unit boundary. anglia, 121:581–609. https://doi.org/10.1515/angl.2003.581 joakim nivre, marie-catherine de marneffe, filip ginter, yoav goldberg, jan hajič, christopher d. manning, ryan mcdonald, slav petrov, sampo pyysalo, natalia silveira, reut tsarfaty and daniel zeman (2016). universal dependencies v1: a multilingual treebank collection. in proceedings of the tenth international conference on language resources and evaluation (lrec’16), pages 1659–1666, portorož, piran. https://aclanthology.org/l161262 tsuyoshi ono and sandra a. thompson (2017). negative scope, temporality, fixedness, and right and left-branching: implications for typology and cognitive processing. studies in language, 41(3):543–576. https://doi.org/10.1075/sl.41.3.01ono https://doi.org/10.1016/s0378-2166(98)00051-4 https://doi.org/10.3389/fpsyg.2015.00731 https://doi.org/10.1075/cld.7.2.03li https://doi.org/10.3389%2ffpsyg.2012.00376 https://doi.org/10.1515/text.1.1988.8.3.243 https://doi.org/10.1016/s0022-5371(81)90450-3 https://doi.org/10.1515/angl.2003.581 https://aclanthology.org/l16-1262 https://aclanthology.org/l16-1262 https://doi.org/10.1075/sl.41.3.01ono the conversational discourse unit: identification and its role in turn-taking management 111 vitória piai, ardi roelofs, joost rommers, kristoffer dahlsätt and eric maris (2015). withholding planned speech is reflected in synchronized beta-band oscillations. frontiers in human neuroscience, 9:549. https://doi.org/10.3389/fnhum.2015.00549 martin j. pickering and chiara gamb (2018). predicting while comprehending language: a theory and review. psychological bulletin, 144: 1002–1044. https://doi.org/10.1037/bul0000158 martin j. pickering and simon garrod (2021). understanding dialogue language use and social interaction. cambridge university press, cambridge, east of england. paola pietrandrea, sylvain kahane, anne lacheret and frédéric sabio (2014). the notion of sentence and other discourse units in corpus annotation. in tommaso raso and heliana mello, editors, spoken corpora and linguistic studies, pages 331–364. john benjamins publishing company, amsterdam/philadelphia, noord-holland/pennsylvania. https://benjamins.com/#catalog/books/scl.61.12pie/details. laurent prévot, shu-chuan tseng, klim peshkov and alvin cheng-hsien chen (2015). processing units in conversation: a comparative study of french and mandarin data. language and linguistics,16(1):69–92. https://doi.org/10.1177/1606822x14556605 marlou rasenberg, asli özyürek, sara bögels and mark dingemanse (2022). the primacy of multimodal alignment in converging on shared symbols for novel referents. discourse processes, 59(3):209–236. https://doi.org/10.1080/0163853x.2021.1992235 geoffrey raymond (2003). grammar and social organization: yes/no interrogatives and the structure of responding. american sociological review, 68(6):939–967. https://psycnet.apa.org/doi/10.2307/1519752 geoffrey raymond (2004). prompting action: the stand-alone “so” in ordinary conversation. research on language and social interaction, 37:185–218. https://doi.org/10.1207/s15327973rlsi3702_4 carina riest, annett b. jorschick and jan peter de ruiter (2015). anticipation in turn-taking: mechanisms and information sources. frontiers in psychology, 6:89. https://doi.org/10.3389%2ffpsyg.2015.00089 harvey sacks, emanuel a. schegloff and gail jefferson (1974). a simplest systematics for the organization of turn-taking for conversation. language, 50(4):696–735. https://doi.org/10.2307/412243 emanuel a. schegloff (1982). discourse as an interactional achievement: some uses of 'uh huh' and other things that come between sentences. in deborah tannen, editor, analyzing discourse: text and talk, pages 71–93. georgetown university press, georgetown, washington dc. emanuel a. schegloff (1996a). turn-organization: one intersection of grammar and interaction. in elinor ochs, emanuel a. schegloff and sandra a. thompson, editors, interaction and grammar, pages 52–133. cambridge university press, cambridge, east of england. emanuel a. schegloff (1996b). confirming allusions: toward an empirical account of action. american journal of sociology, 102(1):161–216. https://doi.org/10.1086/230911 emanuel a. schegloff (2007). sequence organization in interaction. cambridge university press, cambridge, east of england. emanuel a. schegloff (2013). ten operations in self-initiated, same turn repair. in makoto hayashi, geoffrey raymond and jack sindell, editors, conversational repair and human understanding, pages 41–70. cambridge university press, cambridge, east of england. emanuel a. schegloff, gail jefferson and harvey sacks (1977). the preference for self-correction in the organization of repair in conversation. language, 53(2):361–382. https://doi.org/10.2307/413107 margret selting (2000). the construction of units in conversational talk. language in society, 29(4):477–517. https://www.jstor.org/stable/4169050 xiao-nan susan shen (1990). the prosody of mandarin chinese. university of california press, berkeley, california. https://doi.org/10.3389/fnhum.2015.00549 https://doi.org/10.1037/bul0000158 https://benjamins.com/#catalog/books/scl.61.12pie/details https://doi.org/10.1177/1606822x14556605 https://doi.org/10.1080/0163853x.2021.1992235 https://psycnet.apa.org/doi/10.2307/1519752 https://doi.org/10.1207/s15327973rlsi3702_4 https://doi.org/10.3389%2ffpsyg.2015.00089 https://doi.org/10.2307/412243 https://doi.org/10.1086/230911 https://doi.org/10.2307/413107 https://www.jstor.org/stable/4169050 hu and degand 112 matthias sjerps and antje s. meyer (2015). variation in dual-task performance reveals late initiation of speech planning in turn-taking. cognition, 136:304–324. https://doi.org/10.1016/j.cognition.2014.10.008 gabriel skantze (2021). turn-taking in conversational system and human-robot interaction: a review. computer speech & language, 67:101178. https://doi.org/10.1016/j.csl.2020.101178 neal snider and inbal arnon (2012). a unified lexicon and grammar? compositional and noncompositional phrases in the lexicon. in stefan th. gries and dagmar divjak, editors, frequency effects in language, pages 127–164. mouton de gruyter, berlin. manfred stede (2012). discourse processing. morgan & claypool publishers, san rafael, california. gerard steen (2005). basic discourse acts: towards a psychological theory of discourse segmentation. in francisco j. ruiz de mendoza and m. samdra peña, editors, cognitive linguistics: internal dynamics and interdisciplinary interaction, pages 283–312. mouton de gruyter, berlin. tanya stivers, nick j. enfield, penelope brown, christina englert, makoto hayashi, trine heinemann, geritie hoyman, federico rossano, jan peter de ruiter, kyung-eun yoon and stephen c. levinson (2009). universals and cultural variation in turn-taking in conversation. proceedings of the national academy of sciences, 106(26):10587–10592. https://doi.org/10.1073/pnas.0903616106 takashi suzuki and mayumi usami (2005). co-constructions in english and japanese revisited: a quantitative approach to cross-linguistic comparison. in the ninth international pragmatics, pages 263–276, riva del garda, trento. beatrice szczepek reed (2010). prosody and alignment: a sequential perspective. culture studies of science education, 5:859–869. https://doi.org/10.1007/s11422-010-9289-z beatrice szczepek reed (2011). beyond the particular: prosody and the coordination of actions. language and speech, 55(1):13–34. https://doi.org/10.1177%2f0023830911428871 hongyin tao and junfei hu (2019). structural, semantic, and pragmatic properties of nong (弄) constructions in mandarin discourse: evidence from corpora. international journal of chinese linguistics, 6(1):162–176. https://doi.org/10.1075/ijchl.18003.tao emma m. templeton, luke j. chang, elizabeth a. reynolds, marie d. cone lebeaumont and thalia wheatley (2023). long gaps between turns are awkward for strangers but not for friends. philosophical transactions of the royal society b, 378(1875):20210471. https://doi.org/10.1098/rstb.2021.0471 lucien tesnière (1959). eléments de syntaxe structurale. klincksieck, paris, lle-de-france. sandra a. thompson and elizabeth couper-kuhlen (2005). the clause as a locus of grammar and interaction. discourse studies, 7(4/5):481–505. https://doi.org/10.1177%2f1461445605054403 sandra a. thompson and paul j. hopper (2001). transitivity, clause structure, and argument structure: evidence from conversation. in joan bybee and paul hopper, editors, frequency and the emergence of linguistic structure, pages 27–60. john benjamins publishing company, amsterdam, noord-holland/pennsylvania. https://doi.org/10.1016/j.cognition.2014.10.008 https://doi.org/10.1016/j.csl.2020.101178 https://doi.org/10.1073/pnas.0903616106 https://doi.org/10.1007/s11422-010-9289-z https://doi.org/10.1177%2f0023830911428871 https://doi.org/10.1075/ijchl.18003.tao https://doi.org/10.1098/rstb.2021.0471 https://doi.org/10.1177%2f1461445605054403 the conversational discourse unit: identification and its role in conversational turn-taking management 1 introduction 2 conversational discourse units 3 data and method 3.1 dataset 3.2 annotating syntactic, intonation and pragmatic units 3.2.1 syntactic unit the assumption held by all varieties of dependency grammar is that, within a clause, every element is embedded within binary asymmetrical relations called dependency relations, such that none remains isolated (de marneffe & nivre, 2019; heringer, 1993... the starting point of analysis is verbal syntax in which a verb and its governed dependents (including core and oblique arguments, nivre et al., 2016) are central. this analysis leads to the so-called ‘dependency clause’ that demonstrates maximal synt... 3.2.2 intonation unit 3.2.3 pragmatic unit 3.2.3.1 action plan 3.2.3.2 identification of the pragmatic unit (pu) 3.2.4 types of conversational discourse units 3.3 turn-taking annotation: floor-taking-turn vs. non-floor-taking-turn 3.4 reliability check 4 results 4.1 syntactic, prosodic and pragmatic boundaries in turn-taking 4.2 transition speed after each type of conversational discourse unit 5 general discussion 5.1 the importance of meaning-connection in managing turn-taking 5.2 relation between type of cdu and transition speed: earlier emergence of the gist 6 conclusion acknowledgements references mattias heldner, jens edlund, anna hjalmarsson and kornel laskowski (2011). very short utterances and timing in turn-taking. in interspeech 2011, pages 2837–2840, florence, tuscany. the conversational discourse unit: identification and its role in conversational turn-taking management 1 introduction 2 conversational discourse units 3 data and method 3.1 dataset 3.2 annotating syntactic, intonation and pragmatic units 3.2.1 syntactic unit the assumption held by all varieties of dependency grammar is that, within a clause, every element is embedded within binary asymmetrical relations called dependency relations, such that none remains isolated (de marneffe & nivre, 2019; heringer, 1993... the starting point of analysis is verbal syntax in which a verb and its governed dependents (including core and oblique arguments, nivre et al., 2016) are central. this analysis leads to the so-called ‘dependency clause’ that demonstrates maximal synt... 3.2.2 intonation unit 3.2.3 pragmatic unit 3.2.3.1 action plan 3.2.3.2 identification of the pragmatic unit (pu) 3.2.4 types of conversational discourse units 3.3 turn-taking annotation: floor-taking-turn vs. non-floor-taking-turn 3.4 reliability check 4 results 4.1 syntactic, prosodic and pragmatic boundaries in turn-taking 4.2 transition speed after each type of conversational discourse unit 5 general discussion 5.1 the importance of meaning-connection in managing turn-taking 5.2 relation between type of cdu and transition speed: earlier emergence of the gist 6 conclusion acknowledgements references mattias heldner, jens edlund, anna hjalmarsson and kornel laskowski (2011). very short utterances and timing in turn-taking. in interspeech 2011, pages 2837–2840, florence, tuscany. dialogue & discourse 13(1) (2022) 96–122 doi: 10.5210/dad.2022.104 user impressions of system questions to acquire lexical knowledge during dialogues∗ kazunori komatani komatani@sanken.osaka-u.ac.jp sanken (the institute of scientific and industrial research) osaka university kohei ono sanken (the institute of scientific and industrial research) osaka university ryu takeda rtakeda@sanken.osaka-u.ac.jp sanken (the institute of scientific and industrial research) osaka university eric nichols e.nichols@jp.honda-ri.com honda research institute japan co., ltd. mikio nakano mikio.nakano@c4a.jp honda research institute japan co., ltd.† editor: barbara di eugenio submitted 12/2020; accepted 05/2022; published online 06/2022 abstract we have been working on the challenge of systems that acquire the attributes of unfamiliar terms through dialogues, and we previously proposed an approach based on an implicit confirmation process. the questions posed by a dialogue system must not reduce the user’s willingness to converse. in this paper, we conducted a user study that explores the user impression for several question types, including both implicit and explicit questions, to acquire lexical knowledge. user impression scores were collected from 104 participants recruited through crowdsourcing, and a regression analysis was conducted on them. the results demonstrated that implicit questions give a good user impression when their contents are correct, but a bad impression otherwise. the order among the question types combined with their content correctness was also clarified. furthermore, we found that repeating the same types of questions, even those with correct content, annoys users and lowers the user impression. our results provide helpful insights for avoiding degradation of user impression during knowledge acquisition. keywords: dialogue system, knowledge acquisition during dialogue, lexical acquisition, user impression 1. introduction recently, considerable attention has been paid to non-task-oriented dialogue systems in research and commercial system development (higashinaka et al., 2014; yu et al., 2016; smith et al., 2020; nakano and komatani, 2020; roller et al., 2021). in addition to pure non-task-oriented systems, *. this paper is a modified and extended version of our earlier report (komatani and nakano, 2020). †. currently, c4a research institute, inc. ©2022 kazunori komatani, kohei ono, ryu takeda, and honda research institute japan co., ltd. this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). user impressions of system questions to acquire lexical knowledge certain task-oriented dialogue systems can also engage in non-task-oriented dialogues (lee et al., 2009; dingli and scerri, 2013; kobori et al., 2016; papaioannou and lemon, 2017) because such dialogues are expected to build rapport (bickmore and picard, 2005; lucas et al., 2018) between users and systems. because building an open-domain, non-task-oriented dialogue system that always produces appropriate utterances is a difficult task (smith et al., 2020; roller et al., 2021), we consider it worthwhile to build a closed-domain, non-task-oriented dialogue system that attempts to continue dialogues in a specific domain for the purpose of interaction itself. ideally, a closed-domain, non-task-oriented dialogue system should have comprehensive knowledge within its domain, such as lexical knowledge. in reality, all such knowledge cannot be prepared in advance. a knowledge base is not only necessary for providing various services such as information search and recommendation, but also effective for non-task-oriented dialogue systems in order to prevent generic or dull responses (xing et al., 2017; young et al., 2018; zhou et al., 2018; liu et al., 2019). however, it is impractical to presuppose a perfect knowledge base (west et al., 2014). accordingly, we must consider the case in which a human user uses terms1 outside the system’s vocabulary, i.e., new terms whose ontological categories are unknown to the system. one of the most important features of a dialogue system is the ability to acquire knowledge from users and so expand its knowledge base through dialogues. although knowledge may be obtained by asking users to input information into graphical user interfaces (guis) or spreadsheets, knowledge acquisition through dialogues beneficially allows the users to enjoy conversations with the system, especially when the system can engage in non-task-oriented dialogues (kobori et al., 2016). one target of knowledge acquisition through dialogues is obtaining an attribute of an unknown term by asking an appropriate question. this process, here called lexical acquisition, allows systems to keep learning even from dialogues containing unknown terms (meena et al., 2012; sun et al., 2015). lifelong learning (chen and liu, 2018) is an emerging topic that started with machine learning tasks and involves ongoing improvement in a classification, for example. lexical acquisition during dialogues is a type of life-long learning that can potentially become a key technology for the autonomous evolution of intelligent systems. a dialogue system can easily ask questions, but repetitive questions can damage the user experience. a dialogue has to continue to allow a system to acquire a variety of knowledge, but users might stop interacting with a dialogue system that repeatedly asks annoying questions. the users are not crowdworkers who repeatedly tell a system whether a target is correct or wrong (amershi et al., 2014). instead, questions for knowledge acquisition should be designed to avoid excessive irritation to users. question design that does not annoy users is an important component when developing non-task-oriented dialogue systems. consolidating this fact, the alexa prize socialbot grand challenge has made conversation duration a vital criterion along with user ratings (fang et al., 2018; yu et al., 2019; finch et al., 2020). to acquire domain knowledge with less annoying questions, our previous approach adopts an implicit confirmation process (ono et al., 2016), in which a dialogue system produces an utterance about an estimated attribute and determines the correctness of that attribute from the user’s response and other contextual information. figure 1 shows an example of this process. first, when an unknown term emerges in a user’s utterance, the system estimates an attribute of that term (otsuka et al., 2013) (step 1). second, instead of asking an explicit question, the system makes an implicit question about the estimated result (step 2). the implicit question is not a direct interrogative 1. by term, we mean an expression denoting an entity that can exist in a knowledge base and may comprise multiple words. 97 komatani, ono, takeda, nichols and nakano good indonesian restaurants have opened around here recently. i have been to one of them! “nasi goreng” belongs to “indonesian” it seems indonesian … step 1 estimate attribute of w i will try to cook nasi goreng today. unknown term w step 2 ask implicit question about w step 3 determine if attribute is correct figure 1: example of an implicit confirmation process. s1: egg dishes are a good source of nutrition. u2: what are you talking about? u1: japchae is one of my favorites. figure 2: example of an implicit question with wrong content. statement; rather, it operates as a question by interpreting the subsequent user utterance together. third, the system determines the correctness of the estimated result in the implicit question by accounting for the subsequent user response (step 3). if the estimated result is correct, it is added to the knowledge base of the system. to enable dialogue systems to acquire knowledge through dialogues while reducing user discomfort, both implicit and explicit questions should be issued in accordance with the situation. for this purpose, we need to investigate when the implicit confirmation process provides a better user impression than asking an explicit question. moreover, because the questions examine estimation results (in this study, the cuisine types of food names), they can contain wrong content.2 the main contribution of this paper is the results of a user study that explores the impact of knowledge acquisition questions on users’ impressions. we addressed two specific research questions (rqs): rq1 how do the system’s question types affect user impressions? rq2 are consecutive explicit questions for knowledge acquisition more annoying to users than consecutive implicit questions? 2. as will be briefly presented in section 3, we previously tackled the problem of determining the correctness of an estimated result in a question asked during the implicit confirmation process (ono et al., 2017). for example, figure 1 shows an example of an implicit question with correct content. the confirmation process appears to be smooth and the question appears not to bother the user. in contrast, figure 2 exemplifies an implicit question with wrong content, i.e., japchae is not an egg dish but a particular korean dish. such questions with wrong content may annoy the user. 98 user impressions of system questions to acquire lexical knowledge user impression data were collected by an experiment in which crowdsourced workers participated in dialogue sessions, which involved three short interactions with a dialogue system. the user impression data were then regressed against the question types used in the session. to answer rq1, we employed five types of questions including explicit and implicit questions, as well as correct and wrong content. because the estimated results used in the questions are not always correct, we need to consider the impact of such wrong content on users’ impressions. to answer rq2, we compared the average user impression scores when the same question types were actually repeated during the data collection with the predicted scores of such cases in the regression model. the remainder of this this paper is organized as follows. after reviewing related work in section 2, we describe the implicit confirmation process based on our previous experiments (ono et al., 2017) in section 3 and determine the correctness of the content in implicit questions. the main contribution of this paper is given in section 4, which describes the user’s impression of various implicit and explicit question types. section 5 concludes the paper and discusses future work. 2. related work 2.1 knowledge acquisition in dialogue systems computers that continually acquire their own knowledge have been long desired. a famous example is the never-ending language learner (nell) (carlson et al., 2010; mitchell et al., 2015), which continuously extracts information from the web. several techniques developed for machine learning tasks (such as information extraction) can continuously enhance the performance of classifiers in a semi-supervised manner. this learning process is known as life-long learning (chen and liu, 2018). we aim to create systems that acquire knowledge through dialogues with users. knowledge acquisition by dialogue systems has been reported in a number of studies. in meng et al. (2004) and takahashi et al. (2002), lexical information in dialogues was gained by methods that place unknown terms into coarse categories that generally equate to named entity categories. the coarse categories can be acquired more easily than the more specific categories sought by our approach. the method of holzapfel et al. (2008) enables a robot to acquire fine-grained categories for unknown terms by iteratively asking questions. however, as this strategy repeats explicit questions, it is unlikely to be appropriate for non-task-oriented dialogue systems. pappu and rudnicky (2014) designed strategies for asking questions in a goal-oriented dialogue system and analyzed the acquired knowledge through a user study. hixon et al. (2015) proposed a method that poses questions to users and obtain the relations between concepts in a questionanswering system. weston (2016) designed 10 tasks and demonstrated that supervision through feedback from simulated interlocutors improves the utterance-prediction ability of an end-to-end memory network. li et al. (2017) indicated that asking questions improves the performance of a system employing weston’s method with reinforcement learning. mazumder et al. (2019) proposed a dialogue system that asks questions about a triple by using the knowledge graph completion. in these contexts, favorable user impressions of the system’s questions are essential for maintaining the dialogues and allowing the system to acquire a variety of knowledge. 2.2 relationship with implicit confirmation requests in task-oriented dialogues an implicit confirmation request is a well-known error handling technique for task-oriented spoken dialogue systems (bohus and rudnicky, 2005; skantze, 2005). many researchers conducted studies 99 komatani, ono, takeda, nichols and nakano to change the form of the confirmation requests, including explicit and implicit ones (bouwman et al., 1999; komatani and kawahara, 2000). for example, consider a flight reservation system that attempts to determine the destination of a user wishing to travel to seattle. the system can explicitly ask the user “are you going to seattle?” then continue the dialogue by implicitly asking “to get to seattle, where will you depart from?” previous research has shown that an implicit confirmation request can reduce the number of turns when the content is correct, but correcting the system’s misunderstanding when the content is incorrect is difficult (sturm et al., 1999). in other words, implicit confirmation requests involve a tradeoff between conversation fluency and the risk of taking longer time to make corrections. most of the contemporary spoken dialogue agents accommodate this tradeoff; explicit confirmations depend on speech recognition confidence scores (pearl, 2016). implicit questions for non-task-oriented dialogues have different goals from those of implicit confirmation requests for task-oriented dialogues. unlike task-oriented dialogues, implicit questions in non-task-oriented dialogues do not attempt to reduce the number of turns; rather, they lessen the risk of irritating the user and consequently quitting the dialogue. to enable ongoing dialogues between non-task-oriented systems and actual users, we must investigate user impressions of these systems, notably, the acceptability of a certain question type. in particular, the user’s impression of non-task-oriented dialogues must not be impaired. however, questions aimed at knowledge acquisition in non-task-oriented dialogues have been rarely discussed. 2.3 user satisfaction and impression in dialogues several studies have focused on predicting user satisfaction of dialogues. walker et al. (1997) proposed a methodology that predicts user satisfaction in task-oriented dialogues using a regression model with several objectively obtainable parameters. using a hidden markov model, higashinaka et al. (2010) modeled user-satisfaction transitions even when given only the ratings of entire dialogues. ultes and minker (2014) and ultes (2019) employed different context-aware machine learning methods to improve the prediction accuracy of interaction quality, which is defined similarly to user satisfaction. our goal in this paper is not to predict user satisfaction for dialogue evaluation but rather to analyze the effects of various question types on users’ impression in dialogues. in non-task-oriented dialogues, user impressions can be regarded as comparable to user satisfaction because the user is expected to enjoy the dialogue. this situation differs from task-oriented dialogues, in which user satisfaction depends on the task success and dialogue cost walker et al. (1997). kageyama et al. (2018) investigated whether gradually controlling the form of a system’s utterances can improve users’ impression. rather than assessing a whole dialogue system, we focus on one component of user satisfaction, namely, the user’s impression of the system’s questions. specifically, we explore the users’ impression of diverse question types, including explicit and implicit questions, for knowledge acquisition during non-task-oriented dialogues. maintaining high user impressions is vital in a nontask-oriented dialogue system that strives to acquire knowledge from users, because users will cease interacting with a system that repeatedly asks irritating questions. 100 user impressions of system questions to acquire lexical knowledge u1: tempura soba is great! (a) implicit, correct u1: philly cheesesteaks have a lot of calories, but i can’t give them up! s1: i love rare steak. u2: no, a philly cheesesteak is a sandwich. (b) implicit, wrong; determination is easy s1: japanese food is healthy, isn’t it? u2: yes, i ate tempura soba for lunch today. figure 3: examples of correct and wrong implicit questions. s1: sometimes i want to have japanese food. u2: what are you talking about? u1: i baked pandoro yesterday. figure 4: example of an implicit question for which a wrong attribute is not easily determinable. 3. implicit confirmation process of knowledge acquisition this section discusses the concept of the implicit confirmation process. this process presents a possible approach to acquire knowledge through dialogues in non-task-oriented dialogues, but it has received little attention. previously, we showed that the implicit confirmation process can be used to determine the correctness of an estimated attribute of an unknown term (ono et al., 2017) (step 3 in figure 1). for example, question s1 in figure 3 (a) does not explicitly ask the user whether “japanese” is an attribute of the dish tempura soba, but the user response u2 informs the system that the attribute is correct. figure 3 (b) shows another example in which the system determines from u2 that the estimated attribute is wrong. because this approach was promising in our earlier study, the implicit confirmation process can plausibly be applied to knowledge acquisition in our present study. users’ impressions of this process, which have thus far remained unanswered, will be presented in section 4. this section introduces the problem and discusses the viability of the approach. in an implicit confirmation process, whether an estimated attribute is correct cannot always be determined. because implicit questions can elicit many different forms of responses, simply examining the linguistic expressions of a user’s responses is insufficient. for example, in figure 4, the system wrongly estimates the attribute japanese food for pandoro mentioned in u1, although pandoro is italian. the system then forms an implicit question s1. in such cases, the estimated attribute is not easily detected as wrong because the user utterance u2 includes no literally negative expressions, such as “no” and “not.” to avoid these problems, we require a method that precisely determines the correctness of the estimated attributes through the implicit confirmation process. in some studies, affirmative and negative sentences are classified using rules or statistical methods. for example, de marneffe et al. (2009) defined rules for judging whether a response to a yes/no question is affirmative or negative when the response is not a simple “yes” or “no.” gokcen and de marneffe (2015) investigated features for detecting disagreement in a corpus of arguments on the 101 komatani, ono, takeda, nichols and nakano table 1: features of binary classification with user responses. f1 u2 includes an affirmative expression in response to s1 f2 u2 includes a literally negative expression in response to s1 f3 u2 includes an expression correcting s1 f4 u1 and u2 contain the same term f5 u2 includes the category name in s1 f6 u2 includes another category name other than that in s1, excluding cases that fall under f3 f7 u2 includes a word preventing a change in topic in s1 f8 u1 includes the category name in s1 f9 u1 includes another category name other than that in s1 f10 u1 includes any interrogative f11 u1 includes an expression corresponding to the category mentioned in s1 web. such outcomes are helpful for interpreting user responses to explicit questions. in contrast, we attempt to determine the correctness of an attribute in an implicit question not only from the user response but also from the surrounding context in the implicit confirmation process. 3.1 determining the correctness of the content of an implicit question based on the implicit question and the previous and succeeding user utterances, we determine the correctness of the estimated category in an implicit question. determining whether an estimated attribute is correct or wrong in an implicit confirmation process can be cast as a binary classification problem. accordingly, an experiment was conducted by using collected data with user responses. we tested the classification performance between the binary labels “correct” and “wrong.” the classification result can then be used to determine whether the system should add the term-category pair to its knowledge base. we presume that the system can detect a food name in the user’s input, by using methods such as named entity recognition (mesnil et al., 2015), even when that name is outside the system’s vocabulary. the ontological category of an unknown term is an attribute to be estimated. we further assume that a category can be estimated with an existing method, as in otsuka et al. (2013). we make no assumptions about the ontological structure for food. 3.2 evaluation of the correctness of the estimated category 3.2.1 features for the classification table 1 lists the features for classifying the correctness of categories in implicit questions. in the above figures, u1, s1, and u2 respectively represent a user input, the system’s implicit question after u1, and the user response to that question. all feature values are binary; if the statement of each feature is true in a given circumstance, that feature takes the value 1; otherwise, its value is 0. these features were designed to represent differences in the expressions of user responses to implicit questions with either a correct or wrong category. the following expressions were manually prepared for each feature (in japanese, the language in which the experiment was conducted). these expressions were used in the present experiment, but they can be expanded to those that appear in the data. lexical variations can be handled by current techniques enabling more sophisticated use of word embeddings. 102 user impressions of system questions to acquire lexical knowledge feature f1 represents the situation when a user responds an affirmative expression in response to an implicit question with a correct category. affirmative expressions for this feature included “yes,” “that’s right,” and 13 other expressions. similarly, feature f2 represents the situation of negative expressions, which tend to be used for wrong implicit questions. literally negative expressions for this feature included “is not the [category name used in s1]”, “no”, and 15 other expressions. in this paper, features f1 and f2 are employed as the baseline condition because they consider only affirmative and negative expressions in u2, ignoring the relationship between u2 with s1 or u1. when the system asks an implicit question in a wrong category, the user tends to perceive that the system has abruptly altered the topic. feature f3 attempts to detect this situation to correct the system’s previous question s1, with one of six expressions in u2; for example, “it is [another category name other than that in s1].” feature f4 represents the situation in which a category in an implicit question is correct and the user carries the topic in u1 into u2, as shown in figure 3 (a). feature f4 aims to capture this situation by detecting the same term in u1 and u2: tempura soba in this example. similarly to feature f3, features f6 and f7 attempt to detect when the user’s topic in u2 differs from that in u1. for example, consider the following example: u1: i like sangria with its fruity taste. s1: yogashi has a rich taste, doesn’t it? u2: i am talking about the alcoholic beverage. in s1 of this example, the system asks an implicit question with the wrong category yogashi3; the correct category of sangria is alcoholic beverage. subsequently, the user tries to return to the original topic of alcoholic beverage. u2 includes a category name other than that in s1. this situation is represented by feature f6, which excludes cases falling under feature f3 to maintain exclusivity of the two features. in such cases, u2 often contains the japanese word hanashi4, which is represented as feature f7. the present experiment considered only one word. features f6 and f9 covered 25 category names: 20 categories included in the system’s implicit questions, such as “japanese food,” “italian food,” and “korean food”, and five food names such as “cheese” and “pasta.” feature f10 covered 18 interrogative expressions. feature f11 checks for discrepancies between the expressions “eat” and “drink” in u1 and the categories contained in s1. specifically, when u1 contains “eat” or a conjugated form of “eat”, the feature value depends on whether or not s1 contained a category related to “drink” (e.g., an alcoholic beverage), and vice versa. this feature may be domain-dependent because it assumes that each category corresponds to either “eat” or “drink.” 3.2.2 data and setting twenty pairs of terms and corresponding implicit questions were prepared for the experiment. ten of the categories in the implicit questions were correct; the remaining ten were wrong. for example, an implicit question relating to churrasco with its correct category meat dish5 was “eating meat is fun, isn’t it?” as another example, an implicit question related to sangria with a wrong category 3. yogashi means western sweets in japanese. 4. for instance, this word tends to be used to say “i am talking about ...” and “what are you talking about?” in japanese. feature f7 may be unique to the japanese language. 5. note that food category hierarchies in japan may differ from those in other countries. 103 komatani, ono, takeda, nichols and nakano table 2: confusion matrices. reference features output correct wrong all correct 742 313 wrong 236 665 f1, f2 correct 320 220 only wrong 658 758 table 3: classification results. features precision recall f-score all correct 0.703 0.759 0.730 wrong 0.738 0.680 0.708 f1, f2 correct 0.593 0.327 0.422 only wrong 0.535 0.775 0.633 yogashi was “yogashi has a rich taste, doesn’t it?” furthermore, the phrases of the implicit questions were subtly tweaked to improve their naturalness when the user’s input was interrogative or negative. data were collected on 1,956 responses from 98 workers through crowdsourcing. the data were evenly split between the responses to implicit questions with correct and wrong categories. the data from two workers who input only the specified terms or repeated the same sentences were removed. four invalid inputs containing spaces only were also eliminated. classification was performed by logistic regression6. specifically, we used the weka module (version 3.8.1) (hall et al., 2009) with its default parameter settings. the classification results were evaluated using a 10-fold cross-validation. 3.2.3 results the results of two feature sets were compared in the experiment: one including all 11 features listed in table 1 and the other including a baseline set consisting only of features f1 and f2. table 2 presents the results (confusion matrices) of raw outputs of both feature sets the classification accuracies on the complete and baseline feature sets were 71.9% (1,407/1,956) and 55.1% (1,078/1,956), respectively. as confirmed in a mcnemar test, this difference was statistically significant (p < .001), affirming that incorporating the features expressing context improved the classification performance over using the features obtained from u2 alone. feature selection also revealed the most discriminant features for the classification. specifically, the experiments were repeated for the 11 features, i.e., 2, 047 (= 211 − 1) feature sets, and the average f-scores of the combinations were compared. table 3 summarizes the precision, recall, and f-scores of the two classes (“correct” and “wrong”). when using all features and f1 and f2 alone, the average f-scores, i.e., the arithmetic means of the f-scores of the two classes, were 0.719 and 0.528, respectively. table 4 lists the top 10 feature sets ranked by their average f-scores. the condition “none” under which all 11 features were used ranked second in the table, indicating that almost all features effectively contributed to the classification. after eliminating f10, the f-score of the “wrong” cate6. a more modern classifier, such as a transformer-based classifier, might improve the classification performance. 104 user impressions of system questions to acquire lexical knowledge table 4: top-10 feature sets after removing the arbitrary features for classification. removed correct wrong features p r f p r f avg-f f10 .704 .759 .730 .738 .681 .709 .719 none .703 .759 .730 .738 .680 .708 .719 f7,f10 .701 .760 .729 .738 .676 .705 .717 f1,f4,f10 .699 .764 .730 .740 .672 .704 .717 f1,f4 .699 .765 .730 .740 .671 .704 .717 f7 .701 .759 .729 .737 .676 .705 .717 f4,f10 .691 .784 .735 .751 .649 .696 .715 f4 .690 .784 .734 .750 .648 .696 .715 f1,f4,f7,f10 .696 .765 .729 .739 .666 .700 .715 f1,f4,f7 .695 .766 .729 .739 .665 .700 .715 p: precision, r: recall, f: f-score gory marginally improved, leading to an improvement in the overall f-score. because f10 appears numerous times in the table, it was less helpful for the classification than other features. when f10 was removed, f8 had the highest positive weight in the logistic regression function, indicating that f8 gave strong evidence for the “correct” category when its value was 1. consequently, when a category name appeared in both u1 and s1, the category in s1 was likely to be correct because the topic was not suddenly shifted. 4. user study to investigate users’ impression of questions we then explored users’ impressions of implicit and explicit questions. specifically, two specific research questions were addressed. first, we determine the impact of the system’s question types on user impressions. second, we evaluate whether consecutive explicit questions for knowledge acquisition are more annoying than consecutive implicit ones. the data collection was designed to satisfy the following conditions: (1) the user should not be excessively annoyed by the process (2) any effect of the consecutive explicit questions should be discernible we thus adopted a design in which a question survey followed after several subdialogues (three in this paper) were repeated as a session, (see figure 5). although the user’s impressions could simply be assessed after every system question, this design would be extremely inconvenient and break the dialogue flow. instead, we quantified the influence of each question type in one session using a regression model, which also provided an analysis of user impressions after repeating the same question type. 4.1 user study setting we assumed a dialogue system that asks an attribute value for an unknown term. in other words, when an unfamiliar term arises in a dialogue, the system attempts to acquire its attribute from the user through the dialogue. the term and its attribute pair can then be stored as new system knowledge. 105 komatani, ono, takeda, nichols and nakano correct c wrong w explicit e ec “is puttanesca italian?” ew “is puttanesca japanese?” implicit i ic “italian is perfect for a date.” iw “japanese foods are healthy.” whq whq “what is puttanesca?” table 5: five question types of puttanesca, whose correct cuisine type is italian, with examples. e and i respectively denote explicit and implicit questions, c and w respectively denote whether the estimated cuisine is correct or wrong, and whq denotes a wh-question. in the present experiment, we assumed that an unknown food name can be paired with its cuisine type. first, the cuisine type was estimated from the food name’s character sequence (otsuka et al., 2013). the estimated cuisine was then verified by asking a question. 4.1.1 five question types for knowledge acquisition table 5 lists examples of the five question types. in these examples, the unknown term is puttanesca, the estimated correct cuisine is italian, and the estimated wrong cuisine is japanese. each question type has two components: the form of the question and the correctness of its content. the first component can be explicit (e), implicit (i), or a wh-question (whq). an explicit yes/no question asks whether the content of a question is correct (e.g., “is puttanesca italian?”). an implicit question continues the dialogue with a system utterance containing the estimated cuisine (e.g., “italian is perfect for a date”). the system then implicitly determines whether the cuisine is correct by analyzing the subsequent user utterance (ono et al., 2017). meanwhile, a wh-question simply asks without any estimated cuisine (e.g., “what is puttanesca?”). the second component is whether the estimated cuisine is correct (c) or wrong (w). this component allows the investigation of whether users’ impressions are affected by correct or wrong content resulting from the automatic cuisine estimation of the unknown food name (otsuka et al., 2013) before the system posed a question. because wh-questions have no particular content, this component applies exclusively to e and i questions. for simplicity, all questions were one-choice rather than multi-choice (komatani et al., 2016) 4.1.2 data collection the data for evaluating user impressions of dialogues including the above-described five question types were collected by another crowdsourcing7. all crowdworkers were japanese speakers and all dialogues were conducted in japanese. the workers were informed that they would be conversing with an “ai chatbot,” and were asked to simulate a first-time conversation with the chatbot. the workers gave an impression score in each session. the experimental flow is depicted in figure 5. one session consisted of three sets of interactions, followed by an impression survey. 7. we used the platform crowdworks, inc. (https://crowdworks.co.jp/). 106 user impressions of system questions to acquire lexical knowledge 1st set 2nd set 3rd set impression survey four turns per interaction set 0. show specified term to worker 1. worker inputs sentence with term 2. system asks implicit, explicit, or whquestion 3. worker responds 4. system follows up (fixed per question) each worker participated in 10 sessions one session figure 5: flow of data collection. each interaction set included two system turns and two user turns. before the first turn, an instruction with a term was displayed, e.g., “please input your thoughts as though you ate puttanesca recently.” the flow of the four turns is described below: turn 1: the worker types in a sentence including the term specified in the instruction. the terms were prepared before the experiment. turn 2: the system asks a question about the term, where the type of the question is randomly selected. wrong cuisine estimation results and phrases of implicit questions were manually prepared prior to experiment. turn 3: the worker provides an unrestricted response to the system question. turn 4: the system displays its follow-up response, which depends on the question type8 selected in turn 2. that is, the follow-up response was unaffected by the worker’s response in turn 3. for example, when the system’s question in turn 2 was “implicit, wrong (iw),” the system’s follow-up response in turn 4 was always “sorry, i probably misunderstood.” after three interaction sets, the workers recorded their impressions of the session in the questionnaire shown in figure 6. the questionnaire used a 7-point likert scale with two items: “were the system utterances annoying?” and “was the system intelligent?”9 hereafter, these impression scores are denoted as annoying and intelligent, respectively. each worker was required to participate in 10 sessions. before the experiment, we manually prepared 30 terms that were unfamiliar to japanese crowdworkers. these phrases were introduced in fixed order as the unknown terms in each session (three terms per session). 8. again, the question type is randomly selected and not dependent on automatic estimation, such as determining whether the content is correct or wrong. 9. these questionnaire items were unvalidated, meaning that they did not introduce redundancy (i.e., different ways of asking the same content) to minimize misinterpretations, as suggested in (davis, 1989). this straightforward approach was chosen because it was simple to communicate to the crowdworkers. 107 komatani, ono, takeda, nichols and nakano figure 6: screenshot of the impression survey. you: i want to eat bouillabaisse. system: i like american food. you: what are you taking about? system: sorry, i probably misunderstood. you: schnitzel is a pork cutlet. system: is schnitzel german food? you: i think so. system: i understand. thank you. you: i have never eaten puttanesca. system: is puttanesca korean? submit click below to read about puttanesca: https://en.wikipedia.org/wiki/puttanesca iw implicit, wrong ec explicit, correct ew explicit, wrong figure 7: example of the system image in a session. the boxes on the right enclosing the question types are shown for explanation only and were not displayed to the workers. figure 7 is an example of a system image (translated from japanese). the lines starting with “you” and “system” denote the utterances of the worker and the system, respectively. the first display of the interaction set, in which the specified term was presented to the worker, is not depicted in the figure because it disappeared after the worker entered the first sentence. if a worker did not recognize a term, a link to wikipedia was provided at the bottom of the screen for look-up purposes, thereby avoiding dialogues in which the worker was unfamiliar with the meaning of the term. although the dialogues were not natural, they provide an initial step when (as in the present case) a system that can naturally acquire knowledge through several turns is lacking. table 6: numbers of occurrences of question types in the collected data. ec ew whq ic iw 719 618 650 612 680 108 user impressions of system questions to acquire lexical knowledge table 7: summary of the two impression scores obtained on a 7-point scale. intelligent annoying average 3.812 3.048 standard deviation 1.562 1.613 originally, 120 workers collectively completed 1,319 sessions.10 after removing unusable data (such as data from workers who did not complete all 10 sessions), we obtained a total of 1,093 sessions from 104 workers. that is, we obtained 1,093 intelligent and annoying impression scores for each session, where each session included three system question types to be analyzed. the numbers of occurrences of the five question types in the collected data are listed in table 6. these numbers were supposed to be approximately equal but became uneven through the random selection and a system error. the average number of occurrences was 655.8 (= 1, 093 sessions × 3 sets ÷ 5 question types), and the standard deviation was 44.6. the question type was randomly selected three times from the five types, giving 125 (= 53) possible question-type patterns for a session. here, the patterns are represented by concatenating the three question types with hyphens: for example, the pattern in figure 7 is ’iw-ec-ew’. the actual number of patterns was 124, as one pattern (whq-whq-iw) was never chosen by the random selection process. the average occurrence number of each pattern was 8.81 (= 1, 093 sessions ÷ 124 question-type patterns), and the standard deviation was 3.96 (maximum: 17; minimum: 0). table 7 lists the averages and standard deviations of the two impression scores. the standard deviations were large for a 7-point scale. because the impression scores were subjective, there was little agreement on scores among the workers, but impression scores in various question types followed a consistent pattern for each worker. it is also worth noting that the trends of the two impression scores were almost opposite, with a pearson correlation coefficient of −0.512. 4.2 analysis with linear regression from the regression coefficients of the linear regression model, we extracted the influence of each question type. in the basic model, the explanatory variable was the number of occurrences of the five question types in a session and the objective variable was one of the two impression scores (annoying or intelligent). the basic regression model for predicting the score of the i-th session was scorei = w0 + ∑ t∈{ec,ew,whq,ic,iw} wc · numi(t), (1) where numi(t) denotes the number of each question type t used in the session (0, 1, 2, or 3 in the basic model). to improve the multiple correlation coefficients, we added two improvements to the basic model. first, we normalized the impression scores to obtain a mean of 0 and a variance of 1 for each worker. normalization eliminates the variations among the workers. the workers recorded a range of impression scores over the 7-point scale; that is, some gave higher scores, while others gave lower scores. to understand the effect of each question type, we used the relative scores given by each worker. 10. because of a system error, some workers participated in more than 10 sessions. 109 komatani, ono, takeda, nichols and nakano ec1st ec2nd ec3rd ew1st ew2nd impression scores (intelligent or annoying) ew3rd whq1st whq2nd whq3rd ic1st ic2nd ic3rd iw1st iw2nd iw3rd bias figure 8: illustration of the refined regression model. in each of the 15 ovals is a binary value indicating whether a question type occurred at a particular position in the i-th session. second, we considered the positions of the questions in each session. this analysis involved 15 independent variables: the five question types times the three positions (representing the first, second, and third interaction sets in a session). therefore, the refined regression model was scorei = w0 + ∑ d wd · xid, (2) where d ∈ {ec,ew, ic, iw,whq}×{1st, 2nd, 3rd}. the occurrence of each question type d in the i-th session, denoted by xid, takes a binary value (0 or 1) and ∑ d xid = 3 for each i. the model represented by equation (2) is illustrated in figure 8. the occurrence distributions of the question types among the 15 possible positions in the collected data are ideally equal but were unequal in practice. the average number of occurrences at each position was 218.6 (= 1, 093 sessions × 3 sets ÷ 15 question types and positions), and the standard deviation was 17.4 (maximum: 245; minimum: 196). the regression coefficients wd were computed from the collected data using the least-squares approach. these coefficients, which were used in subsequent analysis, represent the change in value of the objective variable (i.e., a user impression score) when each explanatory variable xid is 1, according to causal inference (angrist and pischke, 2008). as a precondition of the analysis, each explanatory variable was binary and uncorrelated with any other explanatory variables. such multicollinearity was avoided because the question types corresponding to the explanatory variables were randomly chosen during the data collection, as detailed in section 4.1.2. the regression model for investigating the influence of the explanatory variables (not for predictive purpose) in terms of the coefficients wd. the resultant wd values may depend on the collected data and the settings of the explanatory variables from which they were derived. however, because each question type was chosen at random and the values of the explanatory variables were supposedly independent, we believe that the relationship among the explanatory variables has certain generality. the multiple correlation coefficients for the two impression scores are listed in table 8. the coefficients increased after normalization and were further increased after accounting for the posi110 user impressions of system questions to acquire lexical knowledge intelligent annoying basic regression model 0.368 0.207 +normalization by worker 0.493 0.308 +consideration of position 0.540 0.354 table 8: multiple correlation coefficients (r) of the models. -0.8 -0.6 -0.4 -0.2 0 0.2 0.4 0.6 1 2 3 ec 1 2 3 ew 1 2 3 whq 1 2 3 ic 1 2 3 iw position and type of questions * * * * * ** ** ** ** -0.8 -0.6 -0.4 -0.2 0 0.2 0.4 0.6 1 2 3 ec 1 2 3 ew 1 2 3 whq 1 2 3 ic 1 2 3 iw position and type of questions re gr es sio n co ef fic ie nt ** ** ** ** ** ** ** ** ** * * **: < 0.01 *: < 0.05 annoyingintelligent figure 9: regression coefficients of the model considering question types and positions. tions. accordingly, in subsequent analysis, we employed the refined model after the normalization and position consideration. 4.3 results 4.3.1 analysis of the obtained regression coefficients we first consider rq1: “how do the system’s question types affect user impressions?” figure 9 shows the values of the 15 regression coefficients obtained for the labels intelligent and annoying. we also checked the statistical significance of the individual regression coefficients being non-zero. the symbols ** and * indicate statistical significance at the p < 0.01 and p < 0.05 levels, respectively. in the case of intelligent, larger positive values indicate that when the system asked that question type in that position, the workers tended to believe that the system was intelligent. a high positive value thus implies a good impression. in the case of annoying, larger positive values imply that when the system asked that question type in that position, the workers tended to believe that the system was annoying. a high positive value thus implies a bad impression. the averages over the three positions for the two labels are summarized in table 9. the regression coefficients of the five question types were ordered as ic > ec > whq > ew > iw for intelligent, and ic < ec < whq < ew < iw for annoying. note the opposite orderings of the two impression scores. 111 komatani, ono, takeda, nichols and nakano ec ew whq ic iw intelligent 0.24 −0.22 0.04 0.35 −0.42 annoying −0.13 0.08 −0.02 −0.21 0.28 table 9: average regression coefficients over the three positions. in the intelligent model, the coefficients of ic and ec (implicit and explicit questions with correct content) were positive, whereas those of ew and iw (implicit and explicit questions with wrong content) were negative. in the annoying model, the opposite relations held. these results correspond to our intuition that when the system asked questions with wrong content, the workers would regard the system unintelligent and become irritated by its questions. because the wh-questions had no concrete content, the whq coefficients were intermediate between those of c and w. however, the whq coefficient for annoying was small and negative, implying that the first wh-questions were not particularly annoying. we now explore the relationship between the explicit and implicit questions. the absolute values of the regression coefficients of the ic questions were larger than those of the ec questions. this result suggests that the implicit questions tended to give a better impression than the explicit ones. as a reason of this trend, we suggest that the workers perceived a knowledge of rare and difficult terms by the system. specifically, the impression scores were higher for target foods with uncommon names that for well-known foods. in contrast, the absolute values of the coefficients of the iw questions were larger than those of the ew questions. in other words, when the estimated cuisine was wrong, the implicit questions gave a worse impression than the explicit ones. in this case, the workers probably perceived that when the system implicitly asked about the wrong cuisine, it had ignored the user’s previous utterances and had switched the dialogue to a new topic. figure 9 also reveals the tendencies among the three positions for each question type. in the case of intelligent, the negative and positive regression coefficients of all five types were largest at the third position. in the case of annoying, the negative and positive regression coefficients of question types ec, ic, and iw were largest at the third position. therefore, the type of question asked soonest before the impression survey significantly affected the impression scores. 4.3.2 impression of repeating the same question type we next consider rq2: “are consecutive explicit questions for knowledge acquisition more annoying to users than consecutive implicit questions?” in this study, we compared the following two impression scores: • actual scores when the same question type was asked three times. • scores predicted by the regression model. the former scores were calculated by averaging the scores of the sessions in which the same question types were actually repeated through random selection. because the question type was randomly chosen during the second turn of a session (see section 4.1.2), the probability of selecting the same question type three consecutive times was 1/53. in the collected data, such occurrences averaged 10.4 times per question type. the latter scores were calculated using the model of eq. (2) in the virtual case of selecting the same question type three consecutive times. e.g., by substituting numi(ec1st) = numi(ec2nd) = 112 user impressions of system questions to acquire lexical knowledge -1.5 -1 -0.5 0 0.5 1 1.5 predicted actualan no yi ng ? ec ew whq ic iw-1.5 -1 -0.5 0 0.5 1 1.5 in te lli ge nt ? ec ew whq ic iw figure 10: impression scores (predicted vs. actual) for intelligent (left) and annoying (right). intelligent annoying predicted actual difference predicted actual difference ec 0.716 1.068 +0.352 −0.380 −0.324 +0.056 ew −0.663 −0.591 +0.072 0.239 0.707 +0.468 whq 0.122 0.457 +0.335 −0.058 0.224 +0.282 ic 1.042 0.901 −0.141 −0.639 −0.502 +0.137 iw −1.251 −1.290 −0.039 0.842 1.429 +0.587 table 10: predicted and actual impression scores and their differences when the same question type was repeated three times. numi(ec3rd) = 1 in eq. (2). because the coefficients were calculated from data in which each question type was randomly chosen, the predicted scores represent cases when the five question types appeared in various contexts. by comparing the two scores, we can analyze the impact of consecutive questions by the average impression scores when the same question type was actually repeated and when a question type appeared in various contexts with those of the latter case represented by the predicted scores. figure 10 shows the comparison and table 10 lists the actual values. we first investigate the results of annoying (right panels in the figure and right columns of the table) because the goal of rq2 was to check the displeasure level of asking consecutive questions. for all question types, the impression scores were larger in the actual cases than in the predicted cases, indicating that asking repeated question types was more annoying. furthermore, the scores for the ew and iw questions showed much wider differences than those for the ec and ic questions, as seen in the “difference” column in the right part of table 10. the finding is consistent with our intuition that asking a series of questions with wrong content is more annoying than asking a series of questions with correct content. examining the results for annoying, we observe that in both the predicted and actual cases, the scores of the ic questions were lower than those of the ec questions, indicating that even when the contents are correct, implicit questions are less annoying, than the explicit questions. the order between the ec and ic questions did not alter, although the degrees of “annoying” were larger in the actual cases than in the predicted cases. when the content was correct, consecutive ec questions were more annoying than consecutive ic questions. therefore, the answer to rq2 is affirmative 113 komatani, ono, takeda, nichols and nakano user the panna cotta was very sweet and good. system is panna cotta italian? user yes. it’s italian. system i understand. thank you. user the pad thai was refreshing and delicious. system is pad thai a thai food? user that’s right. it’s a thai food. system i understand. thank you. user i ate bagna cauda. it was like shabu-shabu. system is bagna cauda italian? user yes. it is an italian pot cuisine. system i understand. thank you. figure 11: example of a session during which explicit questions with correct content (ec questions) were asked three consecutive times. when the questions have correct content. one possible explanation for this outcome is nearly identical forms of the explicit questions (e.g., “is puttanesca italian?”), giving the impression that the system repeated the questions without thought. in contrast, implicit questions caused less disruption and were therefore less annoying because they retained the content of the worker’s previous utterance when the cuisine was correctly estimated. the results for intelligent are presented in the left panel of figure 10 and the left columns of table 10. the impression scores for intelligent in the actual cases varied with question type. the impression scores increased when the system actually asked an ec-type question (i.e., explicit question with the correctly estimated cuisine) three times. we assume that the workers evaluated the system as intelligent because it understood the cuisines of unfamiliar food names. meanwhile, the impression scores for ew and iw did not alter significantly, even when the system actually posed questions with wrong content three times. we speculate that the workers perceived the system as unintelligent, once the system posed a question with wrong content. when the system actually posed an ic-type question three times, the drop in the results might be explained by the invariant phrasing of the ic questions, giving the impression of a monotonous and unintelligent system. when whquestions were actually asked three times, the scores for whq increased, indicating that the workers regarded the system as intelligent despite being irritated by the questions (see results for annoying). note that for intelligent, the results of the whq questions were lower than in ec and ic cases. figures 11 and 12 show example dialogues from sessions during which ec and ic questions, respectively, were asked three times. 4.4 discussion we clarified the order of the question types as either intelligent or annoying. even when the content is correct, we found that repeating the same question type annoys users and lowers their impressions of the system. as long as the content is correct, implicit questions tend to give a better user impressions than explicit ones. in contrast, when the content is wrong, implicit questions are viewed less favorably than explicit ones. these experimental results justify an adaptive strategy based on the cuisine estimation results to lower the risk of degrading the users’ impression. specifically, a dialogue system should select 114 user impressions of system questions to acquire lexical knowledge user gelato looks gorgeous. system ice confections are perfect for eating on a hot day. user that’s right. it’s refreshing. system i think so too. user i ate paella for dinner. system many japanese people like spanish food. user japanese people like such flavors. system i think so too. user ushiojiru is a nice, tasty soup. system many japanese dishes are lightly seasoned. user that’s right. they’re so good. system i think so too. figure 12: example session in which implicit questions with correct content (ic questions) were asked three consecutive times. an implicit question if the confidence score of the cuisine estimation is high and an explicit one (or a wh-question) if the confidence score is low. this adaptive strategy aligns with the prior findings on task-oriented dialogues using confidence measures of automatic speech recognition (e.g., (sturm et al., 1999; pearl, 2016)), as discussed in section 2.2. incorporating the utility of each question type for knowledge acquisition (komatani et al., 2016) would be an interesting extension of this strategy. the analysis results discussed in section 4.3.2 confirmed the earlier best practice of the designers of dialogue systems: that is, the system must avoid repeating the same type of questions in non-task-oriented dialogues. instead, the system should contain multiple question types to engage in smooth dialogues with users. question types should be appropriately changed by considering not only the confidence of estimated cuisines but also the history of the dialogue. the system can effectively acquire knowledge through such dialogues and continue the dialogues without downgrading the user’s impression. varying the question phrases is also worth of exploration. the set expressions of our present experiment might have imparted a monotonous, annoying impression to users. we would therefore be interested in the outcome of syntactically altering the phrases of explicit questions. as the phrases of the implicit questions were likewise fixed for the estimated categories, we are similarly interested in enhancing the diversity of implicit expressions. 5. conclusion through a user study, we addressed a key issue in the implicit confirmation process (ono et al., 2016) of non-task-oriented dialogue systems: whether implicit questions and explicit questions elicit different user impressions of the system. the user impressions were investigated on five types of questions. we clarified the order and found that even when the content is correct, repeating the same question type irritates users and lowers their impression of the system. implicit confirmation is a promising question strategy for a non-task-oriented dialogue system that attempts to acquire more knowledge through dialogue without bothering the users with simple, repeated, explicit questions. the presented findings and methodology will be useful for analyzing 115 komatani, ono, takeda, nichols and nakano how different question types influence user impressions and for designing questions for a system that acquires knowledge effectively through dialogues with users. several issues should be resolved in future work. the number of turns and domain of our experiments were constrained. to remove these constraints, we must evaluate non-task-oriented dialogue systems that can engage in longer dialogues in several domains. we are planning to implement a non-task-oriented dialogue system that can acquire knowledge via an implicit confirmation process embedded within a longer dialogue. the process can be implemented by preparing expressions of implicit questions for each category to be estimated (cuisine types in the current study). this implemented system will be tested in another user study. the present study was not performed in a specific context or situation. this crucial factor must be considered in future work. we discussed knowledge acquisition during non-task-oriented dialogues (e.g., chatting about food), but when a user is teaching the system, the system will be allowed to ask questions repeatedly. the user’s motivation to talk with the system will also change according to a situation. nonetheless, the present findings indicate how the system can prevent the user from losing motivation in continuing the dialogue. in particular, the system must pick appropriate questions and thereby avoid degradation of the user’s impression. an essential problem in knowledge acquisition is that users’ responses may differ, for example, some users may say that mapo doufu is sichuan, while others may claim it is chinese. this difference arises from the different granularity degrees of users’ ideas, as evidenced in their responses. a knowledge graph with different nodes representing such concepts might resolve this challenge. we could also use confidences on the correctness of the question content given by knowledge graph completion results (komatani et al., 2021). extending the findings to spoken interactions is another interesting avenue as it necessitates the use of automatic speech recognition. phoneme recognition techniques, which convert speech signals to phonetic symbols, would be helpful to handle unknown terms. the recognition accuracy of these techniques has been improved by acoustic models based on deep neural networks, but at least two key difficulties remain: (1) the segmentation of recognized phonetic symbols into words, and (2) the distinction between unknown and misrecognized terms. the former difficulty has been tackled by approaches based on bayesian models (heymann et al., 2014; takeda et al., 2018; takeda and komatani, 2019). the latter problem relates to misspelled words in text inputs. distance metrics between an input and known terms are potentially useful for identifying unknown “long” terms rather than “short” ones, where long and short refer to the numbers of the phonetic symbols in the term. we would need to identify unknown terms through dialogue, as no known solution can completely prevent segmentation and recognition failures. acquisition of knowledge, especially of unknown terms, through spoken dialogues still requires favorable user impressions and the dialogue strategy will play a key role in maintaining users’ motivation to continue the dialogues with the system. acknowledgments this work was partly supported by jsps kakenhi grant numbers jp16h02869 and jp19h04171. 116 user impressions of system questions to acquire lexical knowledge references saleema amershi, maya cakmak, william bradley knox, and todd kulesza. power to the people: the role of humans in interactive machine learning. ai magazine, 35(4), december 2014. url https://doi.org/10.1609/aimag.v35i4.2513. joshua angrist and jörn-steffen pischke. mostly harmless econometrics: an empiricist’s companion. princeton university press, 2008. timothy w. bickmore and rosalind w. picard. establishing and maintaining long-term humancomputer relationships. acm transactions on computer-human interaction (tochi), 12(2): 293–327, june 2005. issn 1073-0516. doi: 10.1145/1067860.1067867. url https://doi. org/10.1145/1067860.1067867. dan bohus and alexander rudnicky. error handling in the ravenclaw dialog management architecture. in proc. human language technology conference and conference on empirical methods in natural language processing (hlt-emnlp), pages 225–232, october 2005. url https://www.aclweb.org/anthology/h05-1029. gies bouwman, janienke sturm, and lou boves. incorporating confidence measures in the dutch train timetable information system developed in the arise project. in proc. ieee international conference on acoustics, speech & signal processing (icassp), 1999. doi: 10.1109/icassp. 1999.758170. andrew carlson, justin betteridge, bryan kisiel, burr settles, estevam r. hruschka jr., and tom m. mitchell. toward an architecture for never-ending language learning. in proc. conference on artificial intelligence (aaai), 2010. url http://rtw.ml.cmu.edu/papers/ carlson-aaai10.pdf. zhiyuan chen and bing liu. lifelong machine learning, second edition. synthesis lectures on artificial intelligence and machine learning. morgan & claypool publishers, 2018. doi: 10.2200/s00832ed1v01y201802aim037. url https://doi.org/10.2200/ s00832ed1v01y201802aim037. fred d. davis. perceived usefulness, perceived ease of use, and user acceptance of information technology. mis quarterly, 13(3):319–340, 1989. issn 02767783. url http://www.jstor. org/stable/249008. marie-catherine de marneffe, scott grimm, and christopher potts. not a simple yes or no: uncertainty in indirect answers. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 136–143, 2009. isbn 978-1-932432-64-0. alexiei dingli and darren scerri. building a hybrid: chatterbot – dialog system. in proc. international conference on text, speech, and dialogue (tsd), pages 145–152, 2013. isbn 978-3-64240585-3. hao fang, hao cheng, maarten sap, elizabeth clark, ari holtzman, yejin choi, noah a. smith, and mari ostendorf. sounding board: a user-centric and content-driven social chatbot. in proc. north american chapter of association for computational linguistics (naacl), pages 96–100, 117 komatani, ono, takeda, nichols and nakano june 2018. doi: 10.18653/v1/n18-5020. url https://www.aclweb.org/anthology/ n18-5020. sarah e. finch, james d. finch, ali ahmadvand, ingyu choi, xiangjue dong, ruixiang qi, harshita sahijwani, sergey volokhin, zihan wang, zihao wang, and jinho d. choi. emora: an inquisitive social chatbot who cares for you. in alexa prize proceedings, 2020. ajda gokcen and marie-catherine de marneffe. i do not disagree: leveraging monolingual alignment to detect disagreement in dialogue. in proc. annual meeting of the association for computational linguistics (acl), pages 94–99, 2015. mark hall, eibe frank, geoffrey holmes, bernhard pfahringer, peter reutemann, and ian h. witten. the weka data mining software: an update. acm sigkdd explorations newsletter, 11: 10–18, november 2009. doi: http://doi.acm.org/10.1145/1656274.1656278. jahn heymann, oliver walter, reinhold haeb-umbach, and bhiksha raj. iterative bayesian word segmentation for unsupervised vocabulary discovery from phoneme lattices. in proc. ieee international conference on acoustics, speech & signal processing (icassp), pages 4057–4061, 2014. ryuichiro higashinaka, yasuhiro minami, kohji dohsaka, and toyomi meguro. modeling user satisfaction transitions in dialogues from overall ratings. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), page 18–27, 2010. isbn 9781932432855. url https://www.aclweb.org/anthology/w10-4304. ryuichiro higashinaka, kenji imamura, toyomi meguro, chiaki miyazaki, nozomi kobayashi, hiroaki sugiyama, toru hirano, toshiro makino, and yoshihiro matsuo. towards an opendomain conversational system fully based on natural language processing. in proc. international conference on computational linguistics (coling), pages 928–939, august 2014. ben hixon, peter clark, and hannaneh hajishirzi. learning knowledge graphs for question answering through conversational dialog. in proc. north american chapter of association for computational linguistics (naacl), pages 851–861, 2015. doi: 10.3115/v1/n15-1086. url http://aclweb.org/anthology/n15-1086. hartwig holzapfel, daniel neubig, and alex waibel. a dialogue approach to learning object descriptions and semantic categories. robotics and autonomous systems, 56(11):1004–1013, 2008. yukiko kageyama, yuya chiba, takashi nose, and akinori ito. improving user impression in spoken dialog system with gradual speech form control. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 235–240, 2018. doi: 10.18653/v1/ w18-5026. url https://aclanthology.org/w18-5026. takahiro kobori, mikio nakano, and tomoaki nakamura. small talk improves user impressions of interview dialogue systems. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 370–380, 2016. doi: 10.18653/v1/w16-3646. url http: //www.aclweb.org/anthology/w16-3646. 118 user impressions of system questions to acquire lexical knowledge kazunori komatani and tatsuya kawahara. flexible mixed-initiative dialogue management using concept-level confidence measures of speech recognizer output. in proc. international conference on computational linguistics (coling), pages 467–473, 2000. doi: 10.3115/990820. 990888. url https://doi.org/10.3115/990820.990888. kazunori komatani and mikio nakano. user impressions of questions to acquire lexical knowledge. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 147–156, 1st virtual meeting, july 2020. association for computational linguistics. url https://www.aclweb.org/anthology/2020.sigdial-1.19. kazunori komatani, tsugumi otsuka, satoshi sato, and mikio nakano. question selection based on expected utility to acquire information through dialogue. in proc. international workshop on spoken dialogue system technology (iwsds), pages 27–38, 2016. doi: 10.1007/ 978-981-10-2585-3 6. kazunori komatani, yuma fujioka, keisuke nakashima, katsuhiko hayashi, and mikio nakano. knowledge graph completion-based question selection for acquiring domain knowledge through dialogues. in proc. international conference on intelligent user interfaces (iui), pages 531– 541, 2021. isbn 9781450380171. doi: 10.1145/3397481.3450653. url https://doi. org/10.1145/3397481.3450653. cheongjae lee, sangkeun jung, seokhwan kim, and gary geunbae lee. example-based dialog modeling for practical multi-domain dialog system. speech communication, 51(5):466 – 484, 2009. jiwei li, alexander h. miller, sumit chopra, marc’aurelio ranzato, and jason weston. learning through dialogue interactions by asking questions. in proc. international conference on learning representations (iclr), 2017. url https://openreview.net/pdf?id=rke8pvcle. zhibin liu, zheng-yu niu, hua wu, and haifeng wang. knowledge aware conversation generation with explainable reasoning over augmented graphs. in proc. conference on empirical methods in natural language processing and international joint conference on natural language processing (emnlp-ijcnlp), pages 1782–1792, 2019. doi: 10.18653/v1/d19-1187. url https://doi.org/10.18653/v1/d19-1187. gale m. lucas, jill boberg, david traum, ron artstein, jonathan gratch, alesia gainer, emmanuel johnson, anton leuski, and mikio nakano. getting to know each other: the role of social dialogue in recovery from errors in social robots. in proceedings of the 2018 acm/ieee international conference on human-robot interaction, hri ’18, page 344–351, new york, ny, usa, 2018. association for computing machinery. isbn 9781450349536. doi: 10.1145/3171221.3171258. url https://doi.org/10.1145/3171221.3171258. sahisnu mazumder, bing liu, shuai wang, and nianzu ma. lifelong and interactive learning of factual knowledge in dialogues. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 21–31, september 2019. url https://www. aclweb.org/anthology/w19-5903. raveesh meena, gabriel skantze, and joakim gustafson. a data-driven approach to understanding spoken route directions in human-robot dialogue. in proc. annual conference of the international 119 komatani, ono, takeda, nichols and nakano speech communication association (interspeech), pages 226–229, 2012. doi: 10.21437/ interspeech.2012-73. helen meng, p. c. ching, shuk fong chan, yee fong wong, and cheong chat chan. isis: an adaptive, trilingual conversational system with interleaving interaction and delegation dialogs. acm transactions on computer-human interaction (tochi), 11(3):268–299, 2004. gregoire mesnil, yann dauphin, kaisheng yao, yoshua bengio, li deng, dilek hakkani-tur, xiaodong he, larry heck, gokhan tur, dong yu, and geoffrey zweig. using recurrent neural networks for slot filling in spoken language understanding. ieee/acm transactions on audio, speech, and language processing, 23(3):530–539, march 2015. issn 2329-9290. tom m. mitchell, william cohen, estevam hruschka, partha talukdar, justin betteridge, andrew carlson, bhavana dalvi mishra, matthew gardner, bryan kisiel, jayant krishnamurthy, ni lao nad kathryn mazaitis, thahir mohamed, ndapa nakashole, emmanouil antonios platanios, alan ritter, mehdi samadi, burr settles, richard wang, derry wijaya, abhinav gupta, xinlei chen, abulhair saparov, malcolm greavesand, and joel welling. never-ending learning. in proc. conference on artificial intelligence (aaai), 2015. url https://www.aaai.org/ ocs/index.php/aaai/aaai15/paper/view/10049. mikio nakano and kazunori komatani. a framework for building closed-domain chat dialogue systems. knowledge-based systems, 204:106212, 2020. issn 0950-7051. doi: https://doi.org/10.1016/j.knosys.2020.106212. url http://www.sciencedirect.com/ science/article/pii/s0950705120304287. kohei ono, ryu takeda, eric nichols, mikio nakano, and kazunori komatani. toward lexical acquisition during dialogues through implicit confirmation for closed-domain chatbots. in proc. of second workshop on chatbots and conversational agent technologies (wochat), 2016. url http://workshop.colips.org/wochat/@iva2016/documents/rp-272.pdf. kohei ono, ryu takeda, eric nichols, mikio nakano, and kazunori komatani. lexical acquisition through implicit confirmations over multiple dialogues. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 50–59, 2017. doi: 10.18653/v1/ w17-5507. tsugumi otsuka, kazunori komatani, satoshi sato, and mikio nakano. generating more specific questions for acquiring attributes of unknown concepts from users. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 70–77, august 2013. url http://www.aclweb.org/anthology/w/w13/w13-4009. ioannis papaioannou and oliver lemon. combining chat and task-based multimodal dialogue for more engaging hri: a scalable method using reinforcement learning. in proc. acm/ieee international conference on human-robot interaction (hri), pages 365–366, 2017. isbn 978-14503-4885-0. aasish pappu and alexander rudnicky. knowledge acquisition strategies for goal-oriented dialog systems. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 194–198, june 2014. doi: 10.3115/v1/w14-4326. url https://www. aclweb.org/anthology/w14-4326. 120 user impressions of system questions to acquire lexical knowledge cathy pearl. designing voice user interfaces: principles of conversational experiences. o’reilly media, inc., 1st edition, 2016. isbn 1491955414. stephen roller, emily dinan, naman goyal, da ju, mary williamson, yinhan liu, jing xu, myle ott, eric michael smith, y-lan boureau, and jason weston. recipes for building an open-domain chatbot. in proc. european chapter of the association for computational linguistics (eacl), pages 300–325, 2021. url https://aclanthology.org/2021.eacl-main.24. gabriel skantze. galatea: a discourse modeller supporting concept-level error handling in spoken dialogue systems. in proc. sigdial workshop on discourse and dialogue, pages 178–189, 2005. eric michael smith, mary williamson, kurt shuster, jason weston, and y-lan boureau. can you put it all together: evaluating conversational agents’ ability to blend skills. in proc. annual meeting of the association for computational linguistics (acl), pages 2021–2030, online, july 2020. association for computational linguistics. doi: 10.18653/v1/2020.acl-main.183. url https://www.aclweb.org/anthology/2020.acl-main.183. janienke sturm, els den os, and lou boves. issues in spoken dialogue systems: experiences with the dutch arise system. in proc. esca workshop on interactive dialogue in multi-modal systems, pages 1–4, kloster irsee, germany, 1999. ming sun, yun-nung chen, and alexander i. rudnicky. learning oov through semantic relatedness in spoken dialog systems. in proc. annual conference of the international speech communication association (interspeech), pages 1453–1457, 2015. doi: 10.21437/interspeech. 2015-347. yasuhiro takahashi, kohji dohsaka, and kiyoaki aikawa. an efficient dialogue control method using decision tree-based estimation of out-of-vocabulary word attributes. in proc. international conference on spoken language processing (icslp), pages 813–816, 2002. ryu takeda and kazunori komatani. attribute prediction of unknown lexical entities based on mixture of bayesian segmentation model. in proc. of life long learning for spoken language systems workshop, 2019. ryu takeda, kazunori komatani, and alexander i. rudnicky. word segmentation from phoneme sequences based on pitman-yor semi-markov model exploiting subword information. in proc. of ieee workshop on spoken language technology (slt), pages 763–770, 2018. stefan ultes. improving interaction quality estimation with bilstms and the impact on dialogue policy learning. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 11–20, september 2019. url https://www.aclweb.org/ anthology/w19-5902. stefan ultes and wolfgang minker. interaction quality estimation in spoken dialogue systems using hybrid-hmms. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 208–217, 2014. doi: 10.3115/v1/w14-4328. url https: //aclanthology.org/w14-4328. 121 komatani, ono, takeda, nichols and nakano marilyn a. walker, diane j. litman, candace a. kamm, and alicia abella. paradise: a framework for evaluating spoken dialogue agents. in proc. annual meeting of the association for computational linguistics and conference of the european chapter of the association for computational linguistics (acl-eacl), pages 271–280, 1997. doi: 10.3115/976909.979652. url https://aclanthology.org/p97-1035.pdf. robert west, evgeniy gabrilovich, kevin murphy, shaohua sun, rahul gupta, and dekang lin. knowledge base completion via search-based question answering. in proc. international conference on world wide web (www), pages 515–526, 2014. isbn 978-1-4503-2744-2. doi: 10. 1145/2566486.2568032. url http://doi.acm.org/10.1145/2566486.2568032. jason weston. dialog-based language learning. in proc. international conference on neural information processing systems (nips), pages 829–837, 2016. isbn 978-1-5108-3881-9. url http://dl.acm.org/citation.cfm?id=3157096.3157189. chen xing, wei wu, yu wu, jie liu, yalou huang, ming zhou, and wei-ying ma. topic aware neural response generation. in proc. conference on artificial intelligence (aaai), pages 3351– 3357, 2017. url https://dl.acm.org/doi/10.5555/3298023.3298055. tom young, erik cambria, iti chaturvedi, minlie huang, hao zhou, and subham biswas. augmenting end-to-end dialog systems with commonsense knowledge. in proc. conference on artificial intelligence (aaai), pages 4970–4977, 09 2018. url https://www.aaai.org/ocs/ index.php/aaai/aaai18/paper/view/16573. dian yu, michelle cohn, yi mang yang, chun yen chen, weiming wen, jiaping zhang, mingyang zhou, kevin jesse, austin chau, antara bhowmick, shreenath iyer, giritheja sreenivasulu, sam davidson, ashwin bhandare, and zhou yu. gunrock: a social bot for complex and engaging long conversations. in proc. conference on empirical methods in natural language processing and international joint conference on natural language processing (emnlp-ijcnlp), pages 79–84, hong kong, china, november 2019. association for computational linguistics. doi: 10.18653/v1/d19-3014. url https://www.aclweb.org/anthology/d19-3014. zhou yu, ziyu xu, alan w black, and alexander rudnicky. strategy and policy learning for nontask-oriented conversational systems. in proc. annual meeting of the special interest group on discourse and dialogue (sigdial), pages 404–412, september 2016. hao zhou, tom young, minlie huang, haizhou zhao, jingfang xu, and xiaoyan zhu. commonsense knowledge aware conversation generation with graph attention. in proc. of international joint conference on artificial intelligence (ijcai), pages 4623–4629, 7 2018. doi: 10.24963/ijcai.2018/643. url https://doi.org/10.24963/ijcai.2018/643. 122 dialogue & discourse 14(1) (2023) 1–32 doi: 10.5210/dad.2023.101 automatic essay scoring systems are both overstable and oversensitive: explaining why and proposing defenses yaman kumar singla∗ yamank@iiitd.ac.in adobe media data science research, iiit-delhi, suny at buffalo swapnil parekh∗ swapnil.parekh@nyu.edu new york university somesh singh∗ f20180175@goa.bits-pilani.ac.in iiit-delhi junyi jessy li jessy@austin.utexas.edu university of texas at austin rajiv ratn shah rajivratn@iiitd.ac.in iiit-delhi changyou chen changyou@buffalo.edu suny at buffalo editor: manfred stede submitted 09/2021; accepted 01/2023; published online 04/2023 abstract deep-learning based automatic essay scoring (aes) systems are being actively used in various high-stake applications in education and testing. however, little research has been put to understand and interpret the black-box nature of deep-learning based scoring algorithms. while previous studies indicate that scoring models can be easily fooled, in this paper, we explore the reason behind their surprising adversarial brittleness. we utilize recent advances in interpretability to find the extent to which features such as coherence, content, vocabulary, and relevance are important for automated scoring mechanisms. we use this to investigate the oversensitivity (i.e., large change in output score with a little change in input essay content) and overstability (i.e., little change in output scores with large changes in input essay content) of aes. our results indicate that autoscoring models, despite getting trained as “end-to-end” models with rich contextual embeddings such as bert, behave like bag-of-words models. a few words determine the essay score without the requirement of any context making the model largely overstable. this is in stark contrast to recent probing studies on pre-trained representation learning models, which show that rich linguistic features such as parts-of-speech and morphology are encoded by them. further, we also find that the models have learnt dataset biases, making them oversensitive. the presence of a few words with high co-occurrence with a certain score class makes the model associate the essay sample with that score. this causes score changes in ∼95% of samples with an addition of only a few words. to deal with these issues, we propose detection-based protection models that can detect oversensitivity and samples causing overstability with high accuracies. we find that our proposed models are able to detect unusual attribution patterns and flag adversarial samples successfully. keywords: interpretability in ai, automatic essay scoring, ai in education *. equal contribution ©2023 yaman kumar singla, swapnil parekh, somesh singh, junyi jessy li, rajiv ratn shah, changyou chen this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). singla, parekh, singh, li, shah and chen 1. introduction automatic essay scoring (aes) systems are used in diverse settings such as to alleviate the workload of teachers, save time and costs associated with grading, and to decide admissions to universities and institutions. on average, a british teacher spends five hours in a calendar week scoring exams and assignments (micklewright et al., 2014). this figure is even higher for developing and low-resource countries where the teacher to student ratio is dismal. while on the one hand, autograding systems effectively reduce this burden, allowing more working hours for teaching activities, on the other, there have been many complaints against these systems for not scoring the way they are supposed to (feathers, 2019; smith, 2018; greene, 2018; mid-day, 2017; perelman et al., 2014b). test questions on standardized tests elicit persuasive and informative writing with specific discourse structure. while in persuasive writing, students write their opinions about a topic and try to validate them using convincing arguments, informative writing is more descriptive and requires students to state their experiences to substantiate their opinions. both of them adhere to strict discourse strategies (burstein et al., 2003) which includes an introduction, thesis statements, main and supporting ideas, and finally a conclusion. several research studies have investigated how finding and scoring discourse from essays helps to provide a better holistic score to essays (mcnamara et al., 2014; graesser and mcnamara, 2011; burstein et al., 2001; nadeem et al., 2019; burstein et al., 1998). at the same time, both research studies and empirical evidence have suggested that aes models have repeatedly failed to score discourse and other features important for scoring. for instance, on the recently released automatic scoring system for the state of utah, students scored lower by writing question-relevant keywords but higher by including unrelated content (feathers, 2019; smith, 2018). similarly, it has been a common complaint that aes systems focus unjustifiably on obscure and difficult vocabulary (perelman et al., 2014a). while earlier, each score generated by the ai systems was verified by an expert human rater, it is concerning to see that now many of them are scoring independently without any intervention by human experts (o’donnell, 2020; singla et al., 2022a). the concerns are further alleviated by the fact that the scores awarded by such systems are used in life-changing decisions ranging from college and job applications to visa approvals (ets, 2020b; educational testing association, 2019; usbe, 2020; institute, 2020). traditionally, autograding systems are built using manually crafted features used with machine learning based models (kumar et al., 2019; bamdev et al., 2022). lately, these systems have been shifting to deep learning based models (ke and ng, 2019). for instance, many companies have started scoring candidates using deep learning based automatic scoring (slti-sopi, 2021; assessment, 2021; duolingo, 2021; laflair and settles, 2019; yu et al., 2015; chen et al., 2018; singla et al., 2021; riordan et al., 2017; pearson, 2019). however, there are very few research studies on the reliability1 and validity2 of ml-based aes systems. more specifically, we have tried to address the problems of robustness and validity which plague deep learning based black-box aes models. simply measuring test set performance may mean that the model is right for the wrong reasons. hence, much research is required to understand the scoring algorithms used by aes models and to validate them on linguistic and testing criteria. similar opinions are expressed by madnani and cahill (2018) in their position paper on automatic scoring systems. 1. a reliable measure is one that measures a construct consistently across time, individuals, and situations (ramanarayanan et al., 2020) 2. a valid measure is one that measures what it is intended to measure (ramanarayanan et al., 2020) 2 automatic essay scoring systems are both overstable and oversensitive with this in view, in this paper, we make the following contributions towards understanding current aes systems: 1) several research studies have shown that essay scoring models are overstable (yoon et al., 2018; powers et al., 2002; kumar et al., 2020; feng et al., 2018). even large changes in essay content do not lead to significant change in scores. for instance, kumar et al. (2020) showed that even after changing 20% words of an essay, the scores do not change much. we extend this line of work by addressing why the models are overstable. extending these studies further (§4.1), we investigate aes overstability from the perspective of discourse, coherence, facts, vocabulary, length, grammar and word choice. we do this by using integrated gradients (§3.2), where we find and visualize the most important words for scoring an essay (sundararajan et al., 2017). we find that the models despite using rich contextual embeddings and deep learning architectures, are essentially behaving as bag-of-words models. further, we develop models through which we are able to improve the adversarial attack strength (§4.1.2). for example, for memory networks scoring model (zhao et al., 2017), we delete 40% words from essays without significantly changing score (<1%), whereas kumar et al. (2020) observed that deleting a similar number of words resulted in a decrease of 20% scores for the same model. 2) while there has been much work on aes overstability (kumar et al., 2020; perelman, 2014; powers et al., 2001), there has been little work on aes oversensitivity. building on this research gap, by using adversarial triggers, we find that the aes models are also oversensitive, i.e., small changes in an essay can lead to large change in scores (§4.2). we find that, by just adding 3 words in an essay containing 350 words (<1% change), we are able to change the predicted score by 50% (absolute). we explain the oversensitivity of aes systems using integrated gradients (sundararajan et al., 2017), a principled tool to discover the importance of parts of an input. the results show that the trigger words added to an essay get unusually high attribution. additionally, we find the trigger words have usually high co-occurrence with certain score labels, thus indicating that the models are relying on spurious correlations causing them to be oversensitive (§4.2.4). we validate both the oversensitive and overstable sample in a human study (§5). we ask the annotators whether the scores given by aes models are right by providing them with both original and modified essay responses and scores. 3) while much previous research in the linguistic field studies how essay scoring systems can be fooled, for the first time, we propose models that can detect samples causing overstability and oversensitivity (pham et al., 2021; kumar et al., 2020; perelman, 2020). our models are able to detect both overstability and samples causing oversensitivity with high accuracies (>90% in most cases) (§6). also, for the first time in the literature, through these solutions, we propose a simple yet effective solution for universal adversarial peturbation (§6.1). these models, apart from defending aes systems against samples causing oversensitivity and overstability, can also inform effective human intervention strategy. for instance, aes deployments either completely rely on double scoring essay samples (human and machine) or solely on machine ratings alone (ets, 2020a; singla et al., 2022a). with the developed model, aes deployments can choose to have an effective middle ground by selecting samples for human testing and intervention more effectively. public school systems, e.g., in ohio which use automatic scoring without any human interventions can select samples using these models for limited human intervention (o’donnell, 2020; institute, 2020). for this, we also conduct a small-scale pilot study on the aes deployment of a major language testing company proving the efficacy of the system (§6.3). previous solutions for human interventions optimization rely on brittle features such as number of words and content modeling approaches like off-topic 3 singla, parekh, singh, li, shah and chen detection (yoon et al., 2018; yoon and zechner, 2017). these models cannot detect adversarial samples like the ones we present in our work. we perform our experiments for three model architectures and eight unique prompts3, demonstrating the results on twenty-four unique model-dataset pairs. it is worth noting that our goal in this paper is not to argue against aes systems and their applications. rather, our goal is to interpret how deep-learning based scoring models score essays, why they are overstable and oversensitive, and how to solve the problems of oversensitivity and overstability. we release all our code, dataset and tools for public use with the hope that it will spur testing and validation of aes models. 2. related work the related work for our work can be chiefly divided into two streams: standardized testing and automatic scoring, and testing and validation of the automated scoring models developed. automatic essay scoring: the education testing research community argues for using constructresponse (cr) based testing in high-stakes scenarios (higgins et al., 2011). well-known tests such as toefl, gre, act, linguaskill, duolingo english test, and sat are some examples of cr based standardized testing, where the tests present “naturalistic” prompts such as writing essays and summarizing a news report (such as in gre, toefl, sat), and making a conversation and giving a speech (such as in duolingo english test, toefl, linguaskill) as opposed to artificial tasks like solving multiple choice questions or filling blanks. the community believes that these types of prompts are much more likely to be encountered by the candidate in real-life scenarios (stiggins, 1982; charney, 1984; messick, 1996) and that they provide a more accurate way of measuring test construct (moran, 1987; wiggins, 1991). moreover, artificial testing like those of multiple choice questions lead to washback effects (refers to the extent to which the introduction and use of a test influences language teachers and learners to do things they would not otherwise do that promote or inhibit language learning) (messick, 1996; wiggins, 1991). due to the superiority of cr based testing, they have become a popular means of testing candidates in high-stakes scenarios (such as those of job and visa interviews and college admissions). this poses a significant challenge for language testing and natural language processing communities since cr based testing typically tests on a variety of skills incorporating syntax, semantics, and particularly discourse and organization and by its very nature is costlier than scoring mcqs. by automating this scoring, testing companies reduce the costs associated with training raters, scoring samples, monitoring quality, and also reduce the time to get scores. almost all the auto-scoring models are learning-based and treat the task of scoring as a supervised learning task (ke and ng, 2019; ormerod et al., 2021) with a few using reinforcement learning (wang et al., 2018) and semi-supervised learning (chen et al., 2010). while the earlier models relied on ml algorithms and hand-crafted rules (page, 1966; faulkner, 2014; kumar et al., 2019; persing et al., 2010), lately the systems are shifting to deep learning algorithms (taghipour and ng, 2016; grover et al., 2020; dong and zhang, 2016; uto et al., 2020). approaches have ranged from finding the hierarchical structure of documents (dong and zhang, 2016), using attention over words (dong et al., 2017), multi-stage pretraining (song et al., 2020), and modelling coherence (tay et al., 2018). 3. here a prompt denotes an instance of a unique question asked to test-takers for eliciting their opinions and answers in an exam. the prompts can come from varied domains including literature, science, logic and society. the responses to a prompt indicate the creative, literary, argumentative, narrative, and scientific aptitude of candidates and are judged on a pre-determined score scale. 4 automatic essay scoring systems are both overstable and oversensitive in this paper, we interpret and test the recent state-of-the-art scoring models which have shown the best performance on public datasets (tay et al., 2018; zhao et al., 2017). aes testing and validation: due to the high-stakes nature of the tests, if the aes models are not validated for their adherence to test objectives, they may drive the students to use unethical ways to game the system by addressing the tasks in superficial and construct-irrelevant manner. however, while automatic scoring has seen much research in the recent years, model validation and testing still lag in the ml field with only a few contemporary works (kumar et al., 2020; pham et al., 2021; yoon and xie, 2014; malinin et al., 2017). kumar et al. (2020) and pham et al. (2021) show that aes systems are adversarially unsecure. pham et al. (2021) also try adversarial training and obtain no significant improvements. yoon and xie (2014) and malinin et al. (2017) model uncertainty in automatic scoring systems. most of the scoring model validation work is in the language testing field, which unfortunately has limited ai-expertise (litman et al., 2018). due to this, studies have noted that the results there are often conflicting in nature (powers et al., 2001, 2002; bejar et al., 2013, 2014; perelman, 2020). powers et al. (2002) asked 27 specialists and generalists to write essays that could produce significant deviations with respect to scores from ets’s e-rater. the winner entry repeated the same paragraph 37 times hence showing that repetition over prompt-related keywords makes the scores given by aes unreliable. perelman et al. (2014a) made software that takes in five keywords and produces semantic garbage written in a difficult and obscure language. they tested it out with the ets’s system and produced high scores, thus concluding that the essay writing system learns to recognize obscure language with difficult and nonmeaningful words and phrases like, ‘fundamental drone of humanity’, ‘auguring commencements, torpor of library’ and ‘personal disenfranchisement for the exposition we accumulate conjectures’ 4. in this work, we do a systematic analysis of aes models on features important for scoring and try to interpret the mechanism followed by aes systems for both original and perturbed samples. through this, we discover the overstability and oversensitivity of the aes models and investigate the possible reasons behind their behavior. we also propose several defense mechanisms to solve these problems of the aes models. 3. background 3.1 task, models and dataset we use the widely cited asap-aes (2012) dataset which comes from kaggle automated student assessment prize (asap) for the evaluation of automatic essay scoring systems. the asapaes dataset has been used for automatically scoring essay responses by many research studies (taghipour and ng, 2016; ease, 2013; tay et al., 2018). it is one of the largest publicly available datasets (table 1). the questions covered by the dataset span many different areas such as sciences and english. the responses were written by high school students and were subsequently double-scored. we test the following two state-of-the-art models in this work: skipflow (tay et al., 2018) and memory augmented neural network (mann) (zhao et al., 2017). further, for comparison, we design a bert based automatic scoring model. the performance is measured using quadratic weighted kappa (qwk) metric, which indicates the agreement between a model’s and the expert 4. generated by giving the keywords, ‘library’, ‘delhi’ and ‘college’ respectively. 5 singla, parekh, singh, li, shah and chen prompt number 1 2 3 4 5 6 7 8 #responses 1783 1800 1726 1772 1805 1800 1569 723 score range 2-12 1-6 0-3 0-3 0-4 0-4 0-30 0-60 #avg words per response 350 350 150 150 150 150 250 650 #avg sentences per response 23 20 6 4 7 8 12 35 type argumentative argumentative rc rc rc rc narrative narrative table 1: overview of the asap aes dataset used for evaluation of as systems. (rc = reading comprehension). human rater’s scores. all models show an improvement of 4-5% over the previous models on the qwk metric. the analysis of these models, especially bert, is interesting in light of recent studies indicating that pretrained language models learn rich linguistic features including morphology, parts-of-speech, word-length, noun-verb agreement, coherence, and language delivery (conneau et al., 2018; hewitt and manning, 2019; singla et al., 2022c). this has resulted in pushing the envelope for many nlp applications. the individual models we use are briefly explained as follows: skipflow tay et al. (2018) model essay scoring as a regression task. they utilize glove embeddings for representing the tokens. skipflow captures coherence, flow and semantic relatedness over time, which the authors call neural coherence features. due to the intelligent modelling, it gets an impressive average quadratic weighted kappa score of 0.764. skipflow is one of the top performing models (tay et al., 2018; ke and ng, 2019) for aes. mann zhao et al. (2017) use memory networks for autoscoring by selecting some responses for each grade. these responses are stored in memory and then used for scoring ungraded responses. the memory component helps to characterize the various score levels similar to what a rubric does. they show an excellent agreement score of 0.78 average qwk outperforming the previous stateof-the-art models. bert-based we also design a bert-based architecture for scoring essays. it utilizes bert embeddings (devlin et al., 2019) to represent essays by passing tokens through the bert encoder. the cls token embedding from the last layer is passed through a fully connected layer of size 1 to produce the score. the network was trained to predict the essay scores by minimizing the mean squared error loss. it achieves an average qwk score of 0.74. we utilize this architecture as a baseline representative of transformer-based embedding models. 3.2 attribution mechanism the task of attributing a score f (x) given by an aes model f , on an input essay x can be formally defined as producing attributions a1, .., an corresponding to the words w1, .., wn contained in the essay x. the attributions produced are such that5 sum(a1, .., an) = f (x), i.e., net attributions of all words (sum(a1, .., an)) equal the assigned score (f (x)). in a way, if f is a regression based model, a1, .., an can be thought of as the scores of each word of that essay, which sum to produce the final score, f (x). we use a path-based attribution method, integrated gradients (igs) (sundararajan et al., 2017), much like other interpretability mechanisms such as (ribeiro et al., 2016; lundberg and lee, 2017) for getting the attributions for each of the trained models, f . igs employ the following method to 5. proposition 1 in (sundararajan et al., 2017) 6 automatic essay scoring systems are both overstable and oversensitive find blame assignments: given an input x and a baseline b6, the integrated gradient along the ith dimension is defined as: igi(x, b) = (xi − bi) ∫ 1 α=0 ∂f (b+ α(x− b)) ∂xi dα (1) where ∂f (x) ∂xi represents the gradient of f along the ith dimension of x. we choose the baseline as empty input (all 0s) for essay scoring models since an empty essay should get a score of 0 as per the scoring rubrics. it is the neutral input that models the absence of a cause of any score, thus getting a zero score. since we want to see the effect of only words on the score, any additional inputs (such as memory in mann) of the baseline b is set to be that of x7. see fig. 1 for an example. in all our ig diagrams, green highlighting indicates positive attribution while red highlighting indicates negative attribution. we choose igs over other explainability techniques since they have many desirable properties that make them useful for this task. for instance, the attributions sum to the score of an essay (sum(a1, .., an) = f (x)), they are implementation invariant, do not require any model to be retrained and are readily implementable. previous literature such as (mudrakarta et al., 2018) also uses integrated gradients for explaining the undersensitivity of factoid-based question-answer (qa) models. other interpretability mechanisms like attention require changes in the tested model and are not post-hoc, thus are not a good choice for our task. figure 1: attributions for skipflow, mann and bert models respectively of an essay sample for prompt 2. prompt 2 asks candidates to write an essay to a newspaper reflecting their views on censorship in libraries and express their views if they believe that materials, such as books, etc., should be removed from the shelves if they are found offensive. this essay scored 3 out of 6. 4. empirical studies and results we perform our overstability (§4.1) and oversensitivity (§4.2) experiments with 100 samples per prompt for the three models discussed in section 3.1. there are 8 prompt-level datasets in the overall asap-aes dataset, therefore we perform our analysis on 24 unique model-dataset pairs, each containing over 100 samples. 6. defined as an input containing absence of cause for the output of a model; also called neutral input (shrikumar et al., 2016; sundararajan et al., 2017). 7. we ensure that igs are within the acceptable error margin of <5%, where the error is calculated by the property that the attributions’ sum should be equal to the difference between the probabilities of the input and the baseline. ig parameters: number of repetitions = 20-50, internal batch size = 20-50 7 singla, parekh, singh, li, shah and chen 4.1 aes overstability we first present results on model overstability. following the previous studies, we test the models’ overstability on different features important for aes scoring such as the knowledge of discourse (§4.1.1, 4.1.2), coherence (§4.1.3), facts (§4.1.5), vocabulary (§4.1.4), length (§4.1.2), meaning (§4.1.3), and grammar (§4.1.4). this set of features provides an exhaustive coverage of all features important for scoring essays (yan et al., 2020). 4.1.1 attribution of original samples we take the original human-written essays from the asap-aes dataset and do a word-level attribution of scores. fig. 1 shows the attributions of all models for an essay sample from prompt 2. we observe that skipflow does not attribute any word after the first few lines (first 30% essay content) of the essay, while mann attributions are spread over the complete length of the essay. for the bert-based model, we see that most of the attributions are over nonlinguistic features (tokens) like ‘cls’ and ‘sep’. cls and sep tokens are used as delimiters in the bert model. a similar result was also observed by kovaleva et al. (2019). for skipflow, we observe that if a word is negatively attributed at a certain position in an essay sample, it is then commonly negatively attributed in its other occurrences as well. for instance, books, magazines were negatively attributed in all its occurrences while materials, censored were positively attributed and library was not attributed at all. we could not find any patterns in the direction of attribution. in mann, the same word changes its attribution sign when present in different essays. however, in a single instance of an essay, a word shows the same sign overall despite occurring in very different contexts. table 2 lists the top-positive, top-negative attributed words and the mostly unattributed words for all models. for mann, we notice that the attributions are stronger for function words like to, of, you, do, and are and lesser for content words like shelves, libraries, and music. skipflow’s top attributions are mostly construct-relevant words while bert also focuses more on stopwords. 0 20 40 60 80 100 % length of response 0.00 0.25 0.50 0.75 1.00 re la ti v e q w k iterat ively adding words (in order of im portance) 0 20 40 60 80 100 % length of response 0.00 0.25 0.50 0.75 1.00 re la ti v e q w k iterat ively adding words (in order of im portance) 0 20 40 60 80 100 % length of response 0.00 0.25 0.50 0.75 1.00 re la ti v e q w k iterat ively adding words (in order of im portance) figure 2: variation of qwk with iterative addition of response words for skipflow, mann and bert models. the y-axis notes the relative qwk with respect to the original qwk and the x-axis represents iterative addition of attribute-sorted response words. these results are obtained on prompt 7, similar results were obtained for all the prompts tested. red dashed lines show ’elbow-points’ until where removing x% of tokens results in a near equal qwk score. 8 automatic essay scoring systems are both overstable and oversensitive 4.1.2 iteratively adding important words in this test, we systematically perturb the text discourse by taking an empty essay and iteratively adding the most attributed words of the original sample (eq. 2). ig− attribution sorted list of tokens = (x1, x2, ..., xk, ....., xn) (2) such that ig(x1, b) > ig(x2, b) > .. > ig(xk, b), where xk represents the kth essay token to be removed, b represents the baseline and ig(xk, b) represents the attribution on xk with respect to baseline b. through this, we note the model’s dependence on a few words without their context. fig. 2 presents the results. we observe that the performance (measured by qwk) for the bert model stays within 95% of the original performance even if one of every four words was removed from the essays in the reverse order of their attribution values. the percentage of words deleted were even more for the other models. while fig. 1 showed that mann paid attention to the full length of the response, removing words does not seem to affect the scores much. notably, the words removed are not contiguous but interspersed across sentences, therefore deleting the unattributed words does not produce a grammatically correct response (also see fig. 3), yet can get a similar score thus defeating the whole purpose of testing and feedback. these findings show that there is a point after which the score flattens out, i.e., it does not change in that region either by adding or removing words. this is odd since adding or removing a word from a sentence typically alters its meaning and grammaticality, yet the models do not seem to be affected; they decide their scores only based on 30-50% words. this also demonstrates their lack of discourse knowledge. as an example, a 2-line sample after retaining its top 40% attributed words is given here: “in the end patience rewards better than impatience. a time that i was patient was last year at cheer competition.” model positively attributed words mann to, of, are, ,, children, do, ’, we skipflow of, offensive, movies, censorship, is, our bert ., the, to, and, ”, was, caps, [cls] model negatively attributed words mann i, shelf, by, shelves, libraries, music, a skipflow the, i, to, in, that, do, a, or, be bert i, [sep], said, a, in, time, one model mostly unattributed words mann t, you, the, think, offensive, from, my skipflow it, be, but, their, from, dont, one, what bert @, ##1, and, ,, my, patient table 2: top positive, negative and un-attributed words for skipflow, mann and bert-based model for prompt 2. 4.1.3 sentence and word shuffle coherence and organization are important features for scoring: they measure the unity of different ideas in an essay and determine its cohesiveness in the narrative (barzilay and lapata, 2005). to check the dependence of aes models on coherence, we shuffle the order of sentences and words randomly and note the change in score between the original and modified essay (fig. 3). 9 singla, parekh, singh, li, shah and chen we observe little change (<0.002%) in the attributions with sentence shuffle. the attributions are mostly dependent on word identities rather than their position and context for all models. we also find that shuffling sentences results in 10%, 2% and 3% difference in scores for skipflow, mann, and bert models, respectively. even for these samples for which we observed a change in the scores, almost half of them increased their scores and the other half was reduced. the results are similar for word-level shuffling. this is surprising since changes in the order of ideas in an essay can alter the meaning of a prose, but the models are unable to detect changes in either idea order or word-order. it indicates that despite getting trained as sentence and paragraph level models with the knowledge of language models, they have essentially become bag-of-words models. figure 3: word-shuffled essay containing 40% of (top-attributed) words for skipflow (left), mann (middle) and bert (right) models respectively. the perturbed essay scores 26 (skipflow), 15 (mann) and 5 (bert) out of 30. the original essay was scored 25, 16, 4 respectively by the models. 4.1.4 lexicon modification several previous research studies have highlighted the importance vocabulary plays in scoring and how aes models may be biased towards obscure and difficult vocabulary (perelman et al., 2014a; perelman, 2014; hesse, 2005; powers et al., 2002; kumar et al., 2020). to verify their claims, we replace the top and bottom 10% attributed words with ‘similar’ words8. table 3 shows the results for this test. it can be noted that after replacing all the top and bottom 10% attributed words with their corresponding ‘similar’ words results in an average 4.2% difference in scores across all the models. these results imply that networks are surprisingly not perturbed by modifying even the most attributed words and produce equivalent results with other similarly placed words. in addition, while replacing a word with a ‘similar’ word often changes the meaning and form of a sentence9, the models do not recognize that change by showing no change in their scores. 4.1.5 factuality, common sense, and world knowledge factuality, common sense, and world knowledge are important features in scoring essays (yan et al., 2020). while a human expert can readily catch a lie, it is difficult for a machine to do so. we randomly sample 100 sample essays of each prompt from the addlies test case of (kumar et al., 2020). for constructing these samples, they used various online databases and appended the false 8. sampled from glove with the distance calculated using euclidean distance metric (pennington et al., 2014) 9. for example, consider the replacement of the word ‘agility’ with its synonym (similar word) ‘cleverness’ in the sentence ‘this exercise requires agility.’ does not produce a sentence with the same meaning. 10 automatic essay scoring systems are both overstable and oversensitive result skipflow mann bert avg score difference 9.8% 2.4% 3% % of top-20% attributed words which had a change in their attribution values 20.3% 9.5% 34% % of bottom-20% attributed words which had a change in their attribution values 22.5% 26.0% 45% table 3: statistics obtained after replacing the top and bottom 10% attributed words of each essay with their synonyms. figure 4: attributions for skipflow (left), mann (middle) and bert (right) models of an essay sample where a false fact has been introduced at the beginning. this essay sample scores (25/30, 18/30, 22/30) by the three models respectively. the original essay (without the added lie) scored (24/30), (18/30) and (21/30) respectively. information at various positions in the essay. these statements not only introduce false facts in the essay but also perturb its coherence. a teacher who is responsible for teaching, scoring, and feedback of a student must have knowledge of world knowledge such as ‘sun rises in the east’, and ‘the world is not flat’. however, fig. 4 shows that scoring models do not have the ability to check such common sense. the models tested in fact attribute positive scores to statements like the world is flat if present at the beginning. these results are in contrast with studies like (tenney et al., 2019; zhou et al., 2020) which indicate that bert and glove-like contextual representations have common sense and world knowledge. ettinger (2020) in their ‘negation test’ also observe similar results to us. babel semantic garbage: linguistic literature has also reported that inexplicably, aes models give high scores to semantic garbage like the one generated using b.s. essay language generator (babel generator)10 (perelman et al., 2014a,b; perelman, 2020). these samples are essentially semantic garbage with perfect spellings and obscure and lexically complex vocabulary. in stark contrast to (perelman et al., 2014a) and the commonly held notion that writing obscure and using difficult words fetch more marks, we observed that the models attributed infrequent words such as forbearance, legerdemain, and propinquity negatively while common words such as establishment, celebration, and demonstration were positively scored. therefore, our results show no evidence for the hypothesis reported by studies like (perelman, 2020) that writing lexically complex words make the aes systems give better scores. 10. https://babel-generator.herokuapp.com/ 11 https://babel-generator.herokuapp.com/ singla, parekh, singh, li, shah and chen 4.2 aes oversensitivity while there has been literature on aes overstability, there is much less literature on aes oversensitivity. therefore, next using universal adversarial triggers (wallace et al., 2019), we show the oversensitivity of aes models. we add a few words (adversarial triggers) to the essays and cause them to have large changes in their scores. after that, we attribute the oversensitivity to essay words and show that trigger words have high attributions and are the ones responsible for the model oversensitivity. through this, we test whether an automatically generated small phrase can perform an untargeted attack on a model to increase or decrease the predicted scores irrespective of the original input. our results show that these models are vulnerable to such attacks, with as few as three tokens increasing / decreasing the scores of≈ 99% of samples. further, we show the performance of transfer attacks across prompts and find that ≈ 80% of them transfer, thus showing that the adversaries are easily domain adaptable and transfer well across prompts11. we choose to use universal adversarial triggers for this task since they are input-agnostic, consist of a small number of tokens, and since they do not require the model’s white box access for every essay sample (singla et al., 2022b), they have the potential of being used as “cheat-codes” where a code once extracted can be used by every test-taker. our results show that the triggers are highly effective. prompt→ 1 4 6 7 8 trigger len↓ model = skipflow 3 68, 43 100, 14 86, 40 56, 81 43, 75 5 79, 38 100, 13 97, 42 65, 83 44, 78 10 85, 44 100, 18 100, 48 78, 88 55, 94 20 93, 68 100, 27 100, 58 90, 91 67, 99 model = bert 3 71, 53 89, 31 66, 27 55, 77 46, 61 5 77, 52 90, 33 73, 33 58, 79 49, 64 10 79, 55 91, 41 87, 48 68, 84 55, 75 20 83, 61 94, 49 95, 59 88, 89 61, 89 model = mann 3 67, 38 89, 15 86, 40 60, 80 41, 70 5 73, 39 93, 19 96, 42 61, 71 43, 77 10 85, 44 97, 20 99, 48 75, 84 59, 88 20 93, 63 100, 20 100, 59 84, 90 71, 94 table 4: single-prompt targeted attack performance results. percentage of samples whose scores increase, percentage of samples whose scores decrease on using triggers of length c on prompt p against skipflow, bert and mann. (increasing, decreasing). 4.2.1 adversarial trigger extraction following the procedure of wallace et al. (2019), for a given trigger length (longer triggers are more effective, while shorter triggers are more stealthy), we initialize the trigger sequence by repeating 11. for the consideration of space, we only report a subset of these results. 12 automatic essay scoring systems are both overstable and oversensitive the word “the” and then iteratively replace the tokens in the trigger to minimize the loss for the target prediction over batches of examples from any prompt p. this is a linear approximation of the task loss. we update the embedding for every trigger token eadv to minimize the loss’s first-order taylor approximation around the current token embedding: argei′∈νmin[ei ′ − ei]t∇eadvil (3) where ν is the set of all token embeddings in the model’s vocabulary and∇eadvil is the average gradient of the task loss over a batch. we augment this token replacement strategy with beam search. we consider the top-k token candidates from equation 3 for each token position in the trigger. we search left to right across the positions and score each beam using its loss on the current batch. we use small beam sizes due to computational constraints; increasing them may improve our results. 4.2.2 experiments we conduct two types of experiments namely single prompt attack and cross prompt attack. single prompt attack given a prompt p, response r, model f , size criterion c, an adversary a converts response r to response r′ according to eq. 3. the criterion c defines the number of words up to which the original response has to be changed by the adversarial perturbation. we try out different values of c ({3, 5, 10, 20}). cross prompt attack here the adversarial triggers a obtained from a model f trained on prompt p are tested against the other model f trained on prompt p′ (where p′ 6= p). 4.2.3 results here, we discuss the results of the experiments conducted in the previous section. single prompt attack we found that the triggers can increase or decrease the scores very easily, with 3 or 5-word triggers being able to fool the model more than 95% of times correctly. it results in a mean increase of 50%. table 4 shows the percentage of samples that increase/decrease for various prompts and trigger lengths. the success of triggers increases with the number of words as well. fig. 5 shows a plot of predicted normalized scores before and after attack and how it impacts scores across the entire normalized score range. it shows that the triggers are successful for different prompts and models12. as an example, adding the words “loyalty gratitude friendship” makes skipflow increase the scores of all the essays with a mean normalized increase of 0.5 (out of 1) (prompt 5) whereas adding “grandkids auditory auditory” decreases the scores 97% of the times with a mean normalized decrease of 0.3 (out of 1) (prompt 2). cross prompt attack we also found that the triggers are able to transfer easily, with 95% of samples increasing with a mean normalized increase of ∼0.5 on being subjected to 3-word triggers obtained from attacking a different prompt. fig. 5 shows a similar plot showing the success of triggers obtained from attacking prompt 5 and testing on prompt 4. 4.2.4 trigger analysis we find that it is easier to fool the models to increase the scores than decrease it, with a difference of about 15% in their success (samples increased/decreased). we also observe that some of the triggers selected by our algorithm have very low frequency in the dataset and co-occur with only a 12. other prompts had a similar performance so we have only shown a subset of results with one prompt of each type 13 singla, parekh, singh, li, shah and chen figure 5: single prompt attack for skipflow, bert, memory-networks (§3.1). it shows the models’ predicted scores before and after adding 10-word triggers demonstrating the oversensitivity of these models subject to adversarial triggers. the green line indicates the scores given by a model not under attack, while the blue and red lines show the performance on attempting to increase and decrease the scores using the adversarial triggers. figure 6: cross prompt attack for 20-word triggers obtained from skipflow trained on prompt 5 and tested on prompt 4 showing the transferability across prompts. 14 automatic essay scoring systems are both overstable and oversensitive few output classes (scores), thus having unusually high relative co-occurrence with certain output classes. we calculate pointwise mutual information (pmi) for such triggers and find that the most harmful triggers have the lowest pmi scores with the classes they effect the most (see table 5). prompt 4 prompt 3 score, grade pmi value score, grade pmi value grass, 0 1.58 write, 0 3.28 conclution, 0 1.33 feautures, 0 3.10 adopt, 3 1.86 emotionally, 3 1.33 homesickness, 3 1.78 reservoir, 3 1.27 wich, 1 0.75 seeds, 1 0.93 power, 2 1.03 romshackle, 2 0.96 table 5: pmi of trigger word-grade pairs for prompt 4, 3. other prompts also have similar results. further, we analyze the nature of triggers and find that a significant portion consists of archaic or rare words such as yawing, tallet, straggly with many foreign-derived words as well (wache, bibliotheque)13. we also find that decreasing triggers are 1.5x more repetitive than increasing triggers and contain half as many adjectives as the increasing ones. 5. human baseline to test how humans perform on the different interpretability tests (§4.1, §4.2), we took 50 samples from each of the overstability and oversensitivity tests and asked 2 human expert raters to compare the modified essay samples with the original ones. the expert raters have more than 5 years of experience in the field of language testing. we asked them two questions: (1) whether the score should change after modification and (2) should the score increase or decrease. these questions are easier to answer and produce more objective responses than asking the raters to score responses. we also asked them to give comments behind their ratings. for all overstability tests except lexicon modification, both raters were in perfect agreement (kappa=1.0) on the answers for the two questions asked. they recommended (1) change in scores and that (2) scores should decrease. in most of the comments for the overstability attacks, the raters wrote that they could not understand the samples after modification14. for samples causing oversensitivity, they recommended a score decrease but by a small margin due to little change in those samples. this clearly shows that the predictions of auto-scoring models are different from expert human raters and are yet unable to achieve human-level performance despite the recent claims that autoscoring models have surpassed human level agreement (taghipour and ng, 2016; kumar and boulanger, 2020; ke and ng, 2019). 6. oversensitivity and overstability detection next, we propose detection-based solutions for oversensitivity (§6.1) and overstability (§6.2) causing samples. here we propose detection based defense models to protect the automatic scoring 13. all these words were already part of the model vocabulary. 14. for lexicon modification, the raters recommended the above in 78% instances with 0.85 kappa agreement. 15 singla, parekh, singh, li, shah and chen models against potentially adversarial samples. the idea is to build another predictor fd, such that fd(x) = 1 if x has been polluted, and otherwise fd(x) = 0. other techniques to tackle adversaries such as adversarial training have been shown to be ineffective against aes adversaries (ding et al., 2020; pham et al., 2021). it is noteworthy that we do not solve the general problem of cheating or dishonesty in exams, rather we solve the specific problem of oversensitivity and overstability adversarial attacks on aes models. preventing cheating such as by copying from the web can be easily solved by proctoring or plagiarism checks. however, proctoring or plagiarism checks cannot solve the deep learning models’ adversarial behavior such as due to adding adversarial triggers or repetition and lexically complex tokens. it has been shown in both computer vision and natural language processing that deep-learning models inherently are adversarially brittle and protection mechanisms are required to make them secure (zhang et al., 2020; akhtar and mian, 2018). there is an additional advantage of detection-based adversaries. most aes systems validate their scores with respect to humans post-deployment (ets, 2020a; laflair and settles, 2019). however, many deployed systems are now moving towards human-free scoring (ets, 2020a; o’donnell, 2020; laflair and settles, 2019; slti-sopi, 2021; assessment, 2021). while it may have its advantages such as cost savings, cheating in the form of overstability and samples causing oversensitivity are a major worry for both the testing companies and score users like universities and companies who rely on these testing scores (mid-day, 2017; feathers, 2019; greene, 2018). the detection based models provide an effective middle-ground where the humans only need to evaluate a few samples flagged by the detector models. a few studies studying this problem have been reported in the past (malinin et al., 2017; yoon and xie, 2014). we also do a pilot study with a major testing company using the proposed detector models in order to judge their efficacy (§6.3). studies on the same lines but with different motives have been conducted in the past (powers et al., 2001, 2002). 6.1 ig based oversensitive sample detection using integrated gradients, we calculate the attributions of the trigger words. we found that, on average (over 150 essays across all prompts), the attribution to trigger words is 3 times the attribution to the words in a normal essay (see fig.7). this gave us the motivation to detect oversensitive samples automatically. to detect the presence of triggers (y) programmatically, we utilize a simple 2 layer lstm-fc architecture. ht, ct = l(ht−1, ct−1, xt) y = sigmoid(w ∗ ht + b) (4) the lstm takes the attributions of all words (xt) in an essay as input and predicts whether a sample is adversarial (y) based on attribution values and the hidden state (ht). we include an equal number of trigger and non-trigger examples in the test set. in the train set, we augment the data by including a single response with different types of triggers so as to make the model learn the attribution pattern of words causing oversensitivity. we train the lstm based classifier such that there is no overlap between the train and test triggers. therefore, the classifier has never seen the attributions of any samples with the test-set triggers. using this setup, we obtained an average test accuracy of 94.3% on a test set size of 600 examples per prompt. we do this testing over all the 24 unique prompt-model pairs. the results for 3 prompts (one each from argumentative, narrative, rc (see table 1)) over the attributions of the bert model are tabulated in the table 6. as a baseline, we take an lstm model which takes in bert embeddings and tries to classify the adversarial 16 automatic essay scoring systems are both overstable and oversensitive samples causing oversensitivity using the embedding of the second-last model layer, followed by a dense classification layer. similar results are obtained for all the model-prompt pairs. model prompt accuracy f1 precision recall baseline 2 71 70 80 63 ig-based 2 90 91 84 99 baseline 6 74 74 78 70 ig-based 6 94 93 90 96 baseline 8 60 45 68 34 ig-based 8 99 98 96 100 table 6: validation metrics for ig attribution-based adversarial sample detection compared with embedding-dense classification model for 3 representative prompts (a) trigger “gradually centuries stared” causing score increase (b) trigger “justice you you ... i i i” causing score decrease figure 7: attributions for skipflow when adversarial triggers were inserted in the beginning. the figure shows high attributions on the trigger tokens irrespective of length of triggers. 6.2 language entropy based overstable sample detection for overstability detection, we use a language model to find the text entropy. in psycholinguistics, it is well known that human language has a certain fixed entropy (frank and jaeger, 2008). to maximize the usage of human communication channel, bits per unit (second, or other units like phrases and sentences) remain constant (frank and jaeger, 2008; jaeger, 2010). the principle of uniform information density is followed while reading and speaking (jaeger, 2010; frank and jaeger, 2008; jaeger, 2006). therefore, semantic garbage (babel) or sentence shuffle and word modifications create unexpected language with high entropy. thus, this inherent property of language can be used to detect samples causing overstability. we use a gpt-2 language model (radford et al., 2019) to do unsupervised language modelling on our training corpus to learn the grammar and structure of normal essays. we get the perplexity score p (x) of all essays after passing through gpt-2. p (x) = eh̃(x)where h̃(x) = − ∑ x q(x) loge p(x) (5) where p(x) and q(x) are the estimated (by language model) and true probabilities of the word sequence x. 17 singla, parekh, singh, li, shah and chen we calculate a threshold to distinguish between the perturbed and normal essays (which can also be grammatically incorrect at times). example perplexities of normal essays vs babel essays are shown in fig. 8. figure 8: box plot of normal vs babel gpt perplexities to find the optimal threshold, we use the isolation forest (liu et al., 2008), which is a one class (oc) classification technique. since oc classification only uses one type of examples to train, using only the normal essay perplexity, we can train it to detect when the perplexity is anomalous. scoring function: s(x, n) = 2−e(h(x))/c(n) (6) where e(h(x)) is the mean value of depths that a single data point, x, reaches in all trees. normalizing factor c(n) = 2h(n− 1)− (2(n− 1)/n) (7) where h(i) = harmonic number = ln(i) + 0.5772 (euler’s constant) and n is the number of points used to construct trees. we train this isoforest model on our training perplexities and then test it on our validation set, i.e., other normal essays, shuffled essays (§4.1.3), lexicon-modified essays (§4.1.4) and babel samples (§4.1.5). the contamination factor of the isoforest is set to 1%, corresponding to the number of probable anomalies in the training data. we obtain near perfect accuracies, indicating that our language model has indeed captured the language of a normal essay. table 7 presents the results on three representative prompts (one each from argumentative, narrative, rc (see table 1)). 18 automatic essay scoring systems are both overstable and oversensitive prompt normal essays shuffle synonyms babel 2 99.1 100 82.5 100 6 99.6 98 80 100 8 99.3 98.9 83 100 table 7: isoforest accuracy on normal essays, shuffled essays (§4.1.3), lexicon-modified essays (§4.1.4) and babel samples (§4.1.5) for three representative prompts 6.3 pilot study to test how well the sample detection systems work in practice, we conduct a small-scale pilot study using essay prompts of a major language testing company. we asked 3 experts and 20 candidate testtakers to try to fool the deployed aes models. the experts had an experience of more than 15 years in the field of language testing and were highly educated (masters of science or arts in language and above). the test-takers were college graduates from the population served by the company. they were duly compensated for their time according to the local market rate. we provided them with our overstability and oversensitivity tests for their reference. the pilot study revealed that the test-takers used several strategies to try to bypass the system, like using semantic garbage such as what is generated by the babel generator, sentence and word repetitions, bad grammar, second language use, randomly inserting trigger words, trigger word repetitions, using pseudowords and non-words like jabberwocky, and partial question repeats. the models reported were able to catch most of the attacks including the ones with repetitions, trigger words, pseudoword and non-word usages, and semantic garbage with high accuracy (0.89 f1 with 0.92 recall scores on an average). however, bad-grammar and partial question repeats were difficult to recognize and identify (0.48 f1 score with 0.52 recall scores on an average). this is especially so since bad grammar could be indicative of both language proficiency and adversaries. while bad grammar was easily detected in semantic garbage category, it was detected with low accuracy when only a few sentences were off. similarly, candidates often use partial question repeats to start or end answers. therefore, it forms a construct-relevant strategy and hence cannot be rejected according to rubrics. this problem should be addressed in essay-scoring models by introducing appropriate inductive biases. we leave this task for future work. 7. conclusion and future work automatic scoring, one of the first tasks to be automated using ai (whitlock, 1964), is now shifting to black box neural-network based automated systems. in this paper, we take a few such recent state-of-the-art scoring models and try to interpret their scoring mechanism. we test the models on various features considered important for scoring such as coherence, factuality, content, relevance, sufficiency, logic, etc and explain the models’ predictions. we find that the models do not see an essay as a unitary piece of coherent text but as a bag-of-words. we find out why essay scoring models are both oversensitive and overstable and propose detection based protection models against such attacks. through this, we also propose an effective defense against the recently introduced universal adversarial attacks. apart from contributing to the discussion of finding effective testing strategies, we hope that our exploratory study initiates further discussion about better modeling automatic scoring and testing systems especially in a sensitive area like essay grading. extensive work needs to be done on 19 singla, parekh, singh, li, shah and chen each feature important for scoring a written sample. this includes making available trait-based (or factor-based) essay scoring (mathias and bhattacharyya, 2018), systematically moving from overall scoring to making sure model is internally aware of all factors (attali, 2013), and testing the model on these factors. with millions of candidates each year relying on automatically scored tests for lifechanging decisions like college, job opportunities, and visas, it becomes imperative for the research community to validate their models and show performance metrics beyond just accuracy and kappa numbers. references naveed akhtar and ajmal mian. threat of adversarial attacks on deep learning in computer vision: a survey. ieee access, 6, 2018. asap-aes. the hewlett foundation: automated essay scoring develop an automated scoring algorithm for student-written essays. https://www.kaggle.com/c/asap-aes/, 2012. truenorth speaking assessment. truenorth speaking assessment: the first fully-automated speaking assessment with immediate score delivery. https://emmersion.ai/products/ truenorth/, 2021. yigal attali. validity and reliability of automated essay scoring. in handbook of automated essay evaluation, pages 203–220. routledge, 2013. pakhi bamdev, manraj singh grover, yaman kumar singla, payman vafaee, mika hama, and rajiv ratn shah. automated speech scoring system under the lens: evaluating and interpreting the linguistic cues for language proficiency. international journal of artificial intelligence in education, pages 1–36, 2022. regina barzilay and mirella lapata. modeling local coherence: an entity-based approach. in proceedings of the 43rd annual meeting of the association for computational linguistics (acl’05). association for computational linguistics, 2005. doi: 10.3115/1219840.1219858. isaac i bejar, waverely vanwinkle, nitin madnani, william lewis, and michael steier. length of textual response as a construct-irrelevant response strategy: the case of shell language. ets research report series, 2013(1), 2013. isaac i bejar, michael flor, yoko futagi, and chaintanya ramineni. on the vulnerability of automated scoring to construct-irrelevant response strategies (cirs): an illustration. assessing writing, 22, 2014. jill burstein, karen kukich, susanne wolff, chi lu, and martin chodorow. enriching automated essay scoring using discourse marking. in discourse relations and discourse markers, 1998. url https://aclanthology.org/w98-0303. jill burstein, daniel marcu, slava andreyev, and martin chodorow. towards automatic classification of discourse elements in essays. in proceedings of the 39th annual meeting of the association for computational linguistics, pages 98–105, 2001. 20 https://www.kaggle.com/c/asap-aes/ https://emmersion.ai/products/truenorth/ https://emmersion.ai/products/truenorth/ https://aclanthology.org/w98-0303 automatic essay scoring systems are both overstable and oversensitive jill burstein, daniel marcu, and kevin knight. finding the write stuff: automatic identification of discourse structure in student essays. ieee intelligent systems, 18(1):32–39, 2003. davida charney. the validity of using holistic scoring to evaluate writing: a critical overview. research in the teaching of english, pages 65–81, 1984. lei chen, jidong tao, shabnam ghaffarzadegan, and yao qian. end-to-end neural network based automated speech scoring. in 2018 ieee international conference on acoustics, speech and signal processing, icassp 2018, calgary, ab, canada, april 15-20, 2018. ieee, 2018. doi: 10.1109/icassp.2018.8462562. yen-yu chen, chien-liang liu, chia-hoang lee, tao-hsing chang, et al. an unsupervised automated essay-scoring system. ieee intelligent systems, 25(5), 2010. alexis conneau, german kruszewski, guillaume lample, loı̈c barrault, and marco baroni. what you can cram into a single $&!#* vector: probing sentence embeddings for linguistic properties. in proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: long papers). association for computational linguistics, 2018. doi: 10.18653/v1/ p18-1198. jacob devlin, ming-wei chang, kenton lee, and kristina toutanova. bert: pre-training of deep bidirectional transformers for language understanding. in proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). association for computational linguistics, 2019. doi: 10.18653/v1/n19-1423. yuning ding, brian riordan, andrea horbach, aoife cahill, and torsten zesch. don’t take “nswvtnvakgxpm” for an answer –the surprising vulnerability of automatic content scoring systems to adversarial input. in proceedings of the 28th international conference on computational linguistics. international committee on computational linguistics, 2020. doi: 10.18653/v1/2020.coling-main.76. fei dong and yue zhang. automatic features for essay scoring – an empirical study. in proceedings of the 2016 conference on empirical methods in natural language processing. association for computational linguistics, 2016. doi: 10.18653/v1/d16-1115. fei dong, yue zhang, and jie yang. attention-based recurrent convolutional neural network for automatic essay scoring. in proceedings of the 21st conference on computational natural language learning (conll 2017). association for computational linguistics, 2017. doi: 10.18653/v1/k17-1017. duolingo. the duolingo english test: ai-driven language assessment. https:// emmersion.ai/products/truenorth/, 2021. edx ease. ease (enhanced ai scoring engine) is a library that allows for machine learning based classification of textual content. this is useful for tasks such as scoring student essays. https: //github.com/edx/ease, 2013. 21 https://emmersion.ai/products/truenorth/ https://emmersion.ai/products/truenorth/ https://github.com/edx/ease https://github.com/edx/ease singla, parekh, singh, li, shah and chen eta educational testing association. a snapshot of the individuals who took the gre revised general test. https://www.ets.org/pdfs/gre/snapshot-test-taker-data2019.pdf, 2019. ets. frequently asked questions about the toefl essentials test. https://www.ets.org/s/ toefl-essentials/score-users/faq/, 2020a. ets. gre general test interpretive data. https://www.ets.org/s/gre/pdf/ gre guide table1a.pdf, 2020b. allyson ettinger. what bert is not: lessons from a new suite of psycholinguistic diagnostics for language models. transactions of the association for computational linguistics, 8, 2020. doi: 10.1162/tacl a 00298. adam faulkner. automated classification of stance in student essays: an approach using stance target information and the wikipedia link-based measure. in the twenty-seventh international flairs conference, 2014. todd feathers. flawed algorithms are grading millions of students’ essays. https: //www.vice.com/en/article/pa7dj9/flawed-algorithms-are-gradingmillions-of-students-essays, 2019. shi feng, eric wallace, alvin grissom ii, mohit iyyer, pedro rodriguez, and jordan boyd-graber. pathologies of neural models make interpretations difficult. in proceedings of the 2018 conference on empirical methods in natural language processing. association for computational linguistics, 2018. doi: 10.18653/v1/d18-1407. austin f frank and t florain jaeger. speaking rationally: uniform information density as an optimal strategy for language production. in proceedings of the annual meeting of the cognitive science society, volume 30, 2008. arthur c graesser and danielle s mcnamara. computational analyses of multilevel discourse comprehension. topics in cognitive science, 3(2):371–398, 2011. peter greene. automated essay scoring remains an empty dream. https://www.forbes.com/ sites/petergreene/2018/07/02/automated-essay-scoring-remainsan-empty-dream/?sh=da976a574b91, 2018. manraj singh grover, yaman kumar, sumit sarin, payman vafaee, mika hama, and rajiv ratn shah. multi-modal automated speech scoring using attention fusion. arxiv preprint arxiv:2005.08182, 2020. douglas d hesse. 2005 cccc chair’s address: who owns writing? college composition and communication, 57(2), 2005. john hewitt and christopher d. manning. a structural probe for finding syntax in word representations. in proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). association for computational linguistics, 2019. doi: 10.18653/v1/n19-1419. 22 https://www.ets.org/pdfs/gre/snapshot-test-taker-data-2019.pdf https://www.ets.org/pdfs/gre/snapshot-test-taker-data-2019.pdf https://www.ets.org/s/toefl-essentials/score-users/faq/ https://www.ets.org/s/toefl-essentials/score-users/faq/ https://www.ets.org/s/gre/pdf/gre_guide_table1a.pdf https://www.ets.org/s/gre/pdf/gre_guide_table1a.pdf https://www.vice.com/en/article/pa7dj9/flawed-algorithms-are-grading-millions-of-students-essays https://www.vice.com/en/article/pa7dj9/flawed-algorithms-are-grading-millions-of-students-essays https://www.vice.com/en/article/pa7dj9/flawed-algorithms-are-grading-millions-of-students-essays https://www.forbes.com/sites/petergreene/2018/07/02/automated-essay-scoring-remains-an-empty-dream/?sh=da976a574b91 https://www.forbes.com/sites/petergreene/2018/07/02/automated-essay-scoring-remains-an-empty-dream/?sh=da976a574b91 https://www.forbes.com/sites/petergreene/2018/07/02/automated-essay-scoring-remains-an-empty-dream/?sh=da976a574b91 automatic essay scoring systems are both overstable and oversensitive derrick higgins, gmxiaoming xi, klaus zechner, and david williamson. a three-stage approach to the automated scoring of spontaneous spoken responses. computer speech & language, 25 (2):282–306, 2011. thomas b. fordham institute. ohio public school students. https : / / www.ohiobythenumbers.com/, 2020. t florian jaeger. redundancy and reduction: speakers manage syntactic information density. cognitive psychology, 61(1), 2010. tim florian jaeger. redundancy and syntactic reduction in spontaneous speech. phd thesis, stanford university stanford, ca, 2006. zixuan ke and vincent ng. automated essay scoring: a survey of the state of the art. in proceedings of the twenty-eighth international joint conference on artificial intelligence, ijcai 2019, macao, china, august 10-16, 2019. ijcai.org, 2019. doi: 10.24963/ijcai.2019/879. olga kovaleva, alexey romanov, anna rogers, and anna rumshisky. revealing the dark secrets of bert. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlpijcnlp). association for computational linguistics, 2019. doi: 10.18653/v1/d19-1445. vivekanandan kumar and david boulanger. explainable automated essay scoring: deep learning really has pedagogical value. in frontiers in education, volume 5. frontiers, 2020. yaman kumar, swati aggarwal, debanjan mahata, rajiv ratn shah, ponnurangam kumaraguru, and roger zimmermann. get it scored using autosas an automated system for scoring short answers. in the thirty-third aaai conference on artificial intelligence, aaai 2019, the thirtyfirst innovative applications of artificial intelligence conference, iaai 2019, the ninth aaai symposium on educational advances in artificial intelligence, eaai 2019, honolulu, hawaii, usa, january 27 february 1, 2019. aaai press, 2019. doi: 10.1609/aaai.v33i01.33019662. yaman kumar, mehar bhatia, anubha kabra, jessy junyi li, di jin, and rajiv ratn shah. calling out bluff: attacking the robustness of automatic scoring systems with simple adversarial testing. arxiv preprint arxiv:2007.06796, 2020. geoffrey t laflair and burr settles. duolingo english test: technical manual. retrieved april, 28, 2019. diane litman, helmer strik, and gad s lim. speech technologies and the assessment of second language speaking: approaches, challenges, and opportunities. language assessment quarterly, 15(3):294–309, 2018. fei tony liu, kai ming ting, and zhi-hua zhou. isolation forest. in 2008 eighth ieee international conference on data mining. ieee, 2008. scott m. lundberg and su-in lee. a unified approach to interpreting model predictions. in advances in neural information processing systems 30: annual conference on neural information processing systems 2017, december 4-9, 2017, long beach, ca, usa, 2017. 23 https://www.ohiobythenumbers.com/ https://www.ohiobythenumbers.com/ singla, parekh, singh, li, shah and chen nitin madnani and aoife cahill. automated scoring: beyond natural language processing. in proceedings of the 27th international conference on computational linguistics. association for computational linguistics, 2018. andrey malinin, anton ragni, kate knill, and mark gales. incorporating uncertainty into deep learning for spoken language assessment. in proceedings of the 55th annual meeting of the association for computational linguistics (volume 2: short papers). association for computational linguistics, 2017. doi: 10.18653/v1/p17-2008. sandeep mathias and pushpak bhattacharyya. asap++: enriching the asap automated essay grading dataset with essay attribute scores. in proceedings of the eleventh international conference on language resources and evaluation (lrec 2018), 2018. danielle s mcnamara, arthur c graesser, philip m mccarthy, and zhiqiang cai. automated evaluation of text and discourse with coh-metrix. cambridge university press, 2014. samuel messick. validity and washback in language testing. language testing, 13(3):241–256, 1996. john micklewright, john jerrim, anna vignoles, andrew jenkins, rebecca allen, sonia ilie, elodie bellarbre, fabian barrera, and christopher hein. teachers in england’s secondary schools: evidence from talis 2013. 2014. mid-day. what?! students write song lyrics and abuses in exam answer sheet. https: //www.mid-day.com/articles/national-news-west-bengal-studentswrite-film-song-lyrics-abuses-in-exam-answer-sheet/18210196, 2017. mary ross moran. options for written language assessment. focus on exceptional children, 19 (5):1–12, 1987. pramod kaushik mudrakarta, ankur taly, mukund sundararajan, and kedar dhamdhere. did the model understand the question? in proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: long papers). association for computational linguistics, 2018. doi: 10.18653/v1/p18-1176. farah nadeem, huy nguyen, yang liu, and mari ostendorf. automated essay scoring with discourse-aware neural models. in proceedings of the fourteenth workshop on innovative use of nlp for building educational applications. association for computational linguistics, 2019. doi: 10.18653/v1/w19-4450. patrick o’donnell. computers are now grading essays on ohio’s state tests. https : / / www.cleveland.com / metro / 2018 / 03 / computers are now grading essays on ohios state tests your ch.html, 2020. christopher m. ormerod, akanksha malhotra, and amir jafari. automated essay scoring using efficient transformer-based language models. corr, abs/2102.13136, 2021. url https:// arxiv.org/abs/2102.13136. ellis b page. the imminence of... grading essays by computer. the phi delta kappan, 47(5), 1966. 24 https://www.mid-day.com/articles/national-news-west-bengal-students-write-film-song-lyrics-abuses-in-exam-answer-sheet/18210196 https://www.mid-day.com/articles/national-news-west-bengal-students-write-film-song-lyrics-abuses-in-exam-answer-sheet/18210196 https://www.mid-day.com/articles/national-news-west-bengal-students-write-film-song-lyrics-abuses-in-exam-answer-sheet/18210196 https://www.cleveland.com/metro/2018/03/computers_are_now_grading_essays_on_ohios_state_tests_your_ch.html https://www.cleveland.com/metro/2018/03/computers_are_now_grading_essays_on_ohios_state_tests_your_ch.html https://arxiv.org/abs/2102.13136 https://arxiv.org/abs/2102.13136 automatic essay scoring systems are both overstable and oversensitive pearson. pearson test of english academic: automated scoring. https : / / assets.ctfassets.net / yqwtwibiobs4 / 26s58z1yi9j4ortv0qo3mo / 88121f3d60b5f4bc2e5d175974d52951 / pearson test of english academic-automated-scoring-white-paper-may-2018.pdf, 2019. jeffrey pennington, richard socher, and christopher manning. glove: global vectors for word representation. in proceedings of the 2014 conference on empirical methods in natural language processing (emnlp). association for computational linguistics, 2014. doi: 10.3115/v1/d14-1162. les perelman. when “the state of the art” is counting words. assessing writing, 21, 2014. les perelman. the babel generator and e-rater: 21st century writing constructs and automated essay scoring (aes). the journal of writing assessment, 13, 2020. les perelman, louis sobel, milo beckman, and damien jiang. basic automatic b.s. essay language generator (babel). https://babel-generator.herokuapp.com/, 2014a. les perelman, louis sobel, milo beckman, and damien jiang. basic automatic b.s. essay language generator (babel) by les perelman, ph.d. http://lesperelman.com/writingassessment-robo-grading/babel-generator/, 2014b. isaac persing, alan davis, and vincent ng. modeling organization in student essays. in proceedings of the 2010 conference on empirical methods in natural language processing. association for computational linguistics, 2010. thang pham, trung bui, long mai, and anh nguyen. out of order: how important is the sequential order of words in a sentence in natural language understanding tasks? in findings of the association for computational linguistics: acl-ijcnlp 2021. association for computational linguistics, 2021. doi: 10.18653/v1/2021.findings-acl.98. donald e powers, jill c burstein, martin chodorow, mary e fowles, and karen kukich. stumping e-rater: challenging the validity of automated essay scoring. ets research report series, 2001 (1), 2001. donald e powers, jill c burstein, martin chodorow, mary e fowles, and karen kukich. stumping e-rater: challenging the validity of automated essay scoring. computers in human behavior, 18 (2), 2002. alec radford, jeffrey wu, rewon child, david luan, dario amodei, and ilya sutskever. language models are unsupervised multitask learners. openai blog, 1(8), 2019. vikram ramanarayanan, klaus zechner, and keelan evanini. spoken language technology for language learning & assessment. http://www.interspeech2020.org/uploadfile/ pdf/tutorial-b-4.pdf, 2020. marco túlio ribeiro, sameer singh, and carlos guestrin. ”why should i trust you?”: explaining the predictions of any classifier. in proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, san francisco, ca, usa, august 13-17, 2016. acm, 2016. doi: 10.1145/2939672.2939778. 25 https://assets.ctfassets.net/yqwtwibiobs4/26s58z1yi9j4ortv0qo3mo/88121f3d60b5f4bc2e5d175974d52951/pearson-test-of-english-academic-automated-scoring-white-paper-may-2018.pdf https://assets.ctfassets.net/yqwtwibiobs4/26s58z1yi9j4ortv0qo3mo/88121f3d60b5f4bc2e5d175974d52951/pearson-test-of-english-academic-automated-scoring-white-paper-may-2018.pdf https://assets.ctfassets.net/yqwtwibiobs4/26s58z1yi9j4ortv0qo3mo/88121f3d60b5f4bc2e5d175974d52951/pearson-test-of-english-academic-automated-scoring-white-paper-may-2018.pdf https://assets.ctfassets.net/yqwtwibiobs4/26s58z1yi9j4ortv0qo3mo/88121f3d60b5f4bc2e5d175974d52951/pearson-test-of-english-academic-automated-scoring-white-paper-may-2018.pdf https://babel-generator.herokuapp.com/ http://lesperelman.com/writing-assessment-robo-grading/babel-generator/ http://lesperelman.com/writing-assessment-robo-grading/babel-generator/ http://www.interspeech2020.org/uploadfile/pdf/tutorial-b-4.pdf http://www.interspeech2020.org/uploadfile/pdf/tutorial-b-4.pdf singla, parekh, singh, li, shah and chen brian riordan, andrea horbach, aoife cahill, torsten zesch, and chong min lee. investigating neural architectures for short answer scoring. in proceedings of the 12th workshop on innovative use of nlp for building educational applications. association for computational linguistics, 2017. doi: 10.18653/v1/w17-5017. avanti shrikumar, peyton greenside, anna shcherbina, and anshul kundaje. not just a black box: learning important features through propagating activation differences. arxiv preprint arxiv:1605.01713, 2016. yaman kumar singla, avyakt gupta, shaurya bagga, changyou chen, balaji krishnamurthy, and rajiv ratn shah. speaker-conditioned hierarchical modeling for automated speech scoring. in proceedings of the 30th acm international conference on information & knowledge management, pages 1681–1691, 2021. yaman kumar singla, sriram krishna, rajiv ratn shah, and changyou chen. using sampling to estimate and improve performance of automated scoring systems with guarantees. in proceedings of the aaai conference on artificial intelligence, volume 36 (11), pages 12835–12843, 2022a. yaman kumar singla, swapnil parekh, somesh singh, changyou chen, balaji krishnamurthy, and rajiv ratn shah. minimal: mining models for universal adversarial triggers. proceedings of the aaai conference on artificial intelligence, 36(10):11330–11339, jun. 2022b. doi: 10.1609/ aaai.v36i10.21384. url https://ojs.aaai.org/index.php/aaai/article/view/ 21384. yaman kumar singla, jui shah, changyou chen, and rajiv ratn shah. what do audio transformers hear? probing their representations for language delivery & structure. in 2022 ieee international conference on data mining workshops (icdmw), pages 910–925, 2022c. doi: 10.1109/icdmw58026.2022.00120. slti-sopi. ai-rated speaking exam for professionals (ai sopi). https : / / secondlanguagetesting.com/products-%26-services, 2021. tovia smith. more states opting to ’robo-grade’ student essays by computer. https: //www.npr.org/2018/06/30/624373367/more-states-opting-to-robograde-student-essays-by-computer, 2018. wei song, kai zhang, ruiji fu, lizhen liu, ting liu, and miaomiao cheng. multi-stage pretraining for automated chinese essay scoring. in proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), pages 6723–6733, online, november 2020. association for computational linguistics. doi: 10.18653/v1/2020.emnlp-main.546. url https://aclanthology.org/2020.emnlp-main.546. richard j stiggins. a comparison of direct and indirect writing assessment methods. research in the teaching of english, pages 101–114, 1982. mukund sundararajan, ankur taly, and qiqi yan. axiomatic attribution for deep networks. in proceedings of the 34th international conference on machine learning, icml 2017, sydney, nsw, australia, 6-11 august 2017, volume 70 of proceedings of machine learning research. pmlr, 2017. 26 https://ojs.aaai.org/index.php/aaai/article/view/21384 https://ojs.aaai.org/index.php/aaai/article/view/21384 https://secondlanguagetesting.com/products-%26-services https://secondlanguagetesting.com/products-%26-services https://www.npr.org/2018/06/30/624373367/more-states-opting-to-robo-grade-student-essays-by-computer https://www.npr.org/2018/06/30/624373367/more-states-opting-to-robo-grade-student-essays-by-computer https://www.npr.org/2018/06/30/624373367/more-states-opting-to-robo-grade-student-essays-by-computer https://aclanthology.org/2020.emnlp-main.546 automatic essay scoring systems are both overstable and oversensitive kaveh taghipour and hwee tou ng. a neural approach to automated essay scoring. in proceedings of the 2016 conference on empirical methods in natural language processing. association for computational linguistics, 2016. doi: 10.18653/v1/d16-1193. yi tay, minh c. phan, luu anh tuan, and siu cheung hui. skipflow: incorporating neural coherence features for end-to-end automatic text scoring. in proceedings of the thirty-second aaai conference on artificial intelligence, (aaai-18), the 30th innovative applications of artificial intelligence (iaai-18), and the 8th aaai symposium on educational advances in artificial intelligence (eaai-18), new orleans, louisiana, usa, february 2-7, 2018. aaai press, 2018. ian tenney, patrick xia, berlin chen, alex wang, adam poliak, r. thomas mccoy, najoung kim, benjamin van durme, samuel r. bowman, dipanjan das, and ellie pavlick. what do you learn from context? probing for sentence structure in contextualized word representations. in 7th international conference on learning representations, iclr 2019, new orleans, la, usa, may 6-9, 2019. openreview.net, 2019. usbe. utah state board of education 2018–19 fingertip facts. https://www.ets.org/s/gre/ pdf/gre guide table1a.pdf, 2020. masaki uto, yikuan xie, and maomi ueno. neural automated essay scoring incorporating handcrafted features. in proceedings of the 28th international conference on computational linguistics, pages 6077–6088, barcelona, spain (online), december 2020. international committee on computational linguistics. doi: 10.18653/v1/2020.coling-main.535. url https: //aclanthology.org/2020.coling-main.535. eric wallace, shi feng, nikhil kandpal, matt gardner, and sameer singh. universal adversarial triggers for attacking and analyzing nlp. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp). association for computational linguistics, 2019. doi: 10.18653/v1/d19-1221. yucheng wang, zhongyu wei, yaqian zhou, and xuanjing huang. automatic essay scoring incorporating rating schema via reinforcement learning. in proceedings of the 2018 conference on empirical methods in natural language processing. association for computational linguistics, 2018. doi: 10.18653/v1/d18-1090. james w whitlock. automatic data processing in education. macmillan, 1964. grant wiggins. teaching to the (authentic) test. costa, a., developing minds, a resource book for teaching thinking, asociación para la supervisión del desarrollo del curriculum, ascd, usa, 1: 344–350, 1991. duanli yan, andré a rupp, and peter w foltz. handbook of automated scoring: theory into practice. crc press, 2020. su-youn yoon and shasha xie. similarity-based non-scorable response detection for automated speech scoring. in proceedings of the ninth workshop on innovative use of nlp for building educational applications. association for computational linguistics, 2014. doi: 10.3115/v1/ w14-1814. 27 https://www.ets.org/s/gre/pdf/gre_guide_table1a.pdf https://www.ets.org/s/gre/pdf/gre_guide_table1a.pdf https://aclanthology.org/2020.coling-main.535 https://aclanthology.org/2020.coling-main.535 singla, parekh, singh, li, shah and chen su-youn yoon and klaus zechner. combining human and automated scores for the improved assessment of non-native speech. speech communication, 93, 2017. su-youn yoon, aoife cahill, anastassia loukina, klaus zechner, brian riordan, and nitin madnani. atypical inputs in educational applications. in proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 3 (industry papers). association for computational linguistics, 2018. doi: 10.18653/v1/n18-3008. zhou yu, vikram ramanarayanan, david suendermann-oeft, xinhao wang, klaus zechner, lei chen, jidong tao, aliaksei ivanou, and yao qian. using bidirectional lstm recurrent neural networks to learn high-level abstractions of sequential features for automated scoring of non-native spontaneous speech. in 2015 ieee workshop on automatic speech recognition and understanding (asru). ieee, 2015. wei emma zhang, quan z sheng, ahoud alhazmi, and chenliang li. adversarial attacks on deep-learning models in natural language processing: a survey. acm transactions on intelligent systems and technology (tist), 11(3), 2020. siyuan zhao, yaqiong zhang, xiaolu xiong, anthony botelho, and neil heffernan. a memoryaugmented neural model for automated grading. in proceedings of the fourth (2017) acm conference on learning@ scale, 2017. xuhui zhou, yue zhang, leyang cui, and dandan huang. evaluating commonsense in pre-trained language models. in proceedings of the aaai conference on artificial intelligence, volume 34, 2020. 28 automatic essay scoring systems are both overstable and oversensitive appendices 8. full version of abridged main paper figures the essays with their sentences shuffled are displayed in figure 5. the full-size attribution images are given in figures 10 to 13. figure 9: attributions for skipflow and mann respectively of an essay sample where all the sentences have been randomly shuffled. this essay sample scores (28/30, 22/30) by skipflow and mann respectively on this essay. the original essay also scored (28/30) and (22/30) respectively. 9. statistics: iterative addition of words the results are given in table 8. % µpos µneg npos nneg σ skipflow/mann/bert 80 3.5/1.1/0.002 0.43/0.09/0.05 65/31/0.63 8.9/2.88/91.6 5.1/2/0.06 60 4/0.37/0.001 1.01/1.4/0.14 60/9.2/0.3 17/39.1/99 6.7/2.6/0.14 40 3.1/0.07/0 3.7/5.8/0.23 36/2.24/0 44/88.4/99.6 9.24/6.5/0.24 20 2.09/0.02/0.002 14.7/13.7/0.31 15.6/0.6/0.63 78.5/94.5/99.3 19.5/14.5/0.32 0 61/0/0 0/20/0.52 0/0/0 100/94.5/100 62/22.3/0.5 table 8: statistics for iterative addition of the most-attributed words on prompt 7. legend (kumar et al., 2020): {%: % words added to form a response, µpos: mean difference of positively impacted samples (as % of score range), µneg: mean difference of negatively impacted samples (as % of score range), npos: percentage of positively impacted samples, nneg: percentage of negatively impacted samples, σ: standard deviation of the difference (as % of score range)} 10. statistics: iterative removal of words full version of the results are in the table 9. 11. bert-model hyperparameters bert model hyperparameters are given in the table 10. 29 singla, parekh, singh, li, shah and chen figure 10: full-sized attributions for skipflow, mann and bert respectively of an essay sample where all the sentences have been randomly shuffled. % µpos µneg npos nneg skipflow/mann/bert 0 0/0/0 0/0/0 0/0/0 0/0/0 20 0/0/0.04 11/1/5 0/.3/1.27 96.1/32/88.4 60 0/0/0.01 26/8/14.8 1.2/0/0.3 97.7/94.5/99.3 80 0.5/0/0 29.9/15/22 5.4/0/0 92.9/94.5/100 table 9: statistics for iterative removal of least attributed words on prompt 7. legend (kumar et al., 2020): {%: % words removed from a response, µpos: mean difference of positively impacted samples (as % of score range), µneg: mean difference of negatively impacted samples (as % of score range), npos: percentage of positively impacted samples, nneg: percentage of negatively impacted samples } 30 automatic essay scoring systems are both overstable and oversensitive figure 11: full-sized attributions for skipflow, mann and bert respectively of an babel essay sample. 31 singla, parekh, singh, li, shah and chen figure 12: full-sized attributions for skipflow, mann and bert respectively of an essay sample where all the sentences have an added false fact. hyperparameter value optimizer adam learning rate 2e-5 batch size 8 epochs 5-10 based on early stopping loss mean squared error table 10: bert model hyperparameters and architecture 32 automatic essay scoring systems are both overstable and oversensitive figure 13: full-sized attributions for skipflow, mann and bert respectively of a real essay sample. 33 introduction related work background task, models and dataset attribution mechanism empirical studies and results aes overstability attribution of original samples iteratively adding important words sentence and word shuffle lexicon modification factuality, common sense, and world knowledge aes oversensitivity adversarial trigger extraction experiments results trigger analysis human baseline oversensitivity and overstability detection ig based oversensitive sample detection language entropy based overstable sample detection pilot study conclusion and future work full version of abridged main paper figures statistics: iterative addition of words statistics: iterative removal of words bert-model hyperparameters dialogue & discourse 15(2) (2024) 113–144 doi: 10.5210/dad.2024.204 common ground inconsistencies in dialogue systems: conflict patterns implied by polar question forms maria di maro maria.dimaro2@unina.it dept. of electrical engineering and information technology, university of naples “federico ii”, urban/eco research centre antonio origlia antonio.origlia@unina.it dept. of electrical engineering and information technology, university of naples “federico ii”, urban/eco research centre francesco cutugno cutugno@unina.it dept. of electrical engineering and information technology, university of naples “federico ii”, urban/eco research centre editor: david traum submitted 01/2023; accepted 11/2024; published online 12/2024 abstract in linguistics, research on dialogue systems has accentuated the need to focus on various pragmatic aspects for their management and modelling. among the most important pragma-linguistic speech acts in dialogue systems studies are clarification requests, corrective feedback that in some circumstances require access to the set of shared knowledge known as common ground. regarding common ground management, pragmatic studies suggest differences in the type of polar questions that people prefer be used in clarification requests, where polar questions can have two possible answers: true or false. this preference appears to depend on the relationship between bias and contextual evidence. in this work, we show that varying the form of polar questions in a given pragmatic setting can influence the capability of people to track common ground inconsistencies. as a result, we demonstrate that using a negative polar question in italian has functional consequences when communicating conflicting material in the common ground. this can improve the quality of human interactions with dialogue systems, in terms of an improved identification of the conflict. the results obtained in this work provide insights into design of error reporting approaches in natural interactions. keywords: dialogue systems, pragmatics, common ground inconsistencies 1. introduction the wide success and the increasing popularity of conversational agents are shedding a new light on conversation analysis and on the pragmatic structure of dialogue. this includes the study of the automatic recognition of pragmatic phenomena which are common in human-human interactions (carberry, 1985; bunt and black, 2000; ammicht et al., 2003; skantze, 2005; hough et al., 2017). the significance of dialogue interaction in the realm of artificial intelligence cannot be overstated, ©2024 maria di maro, antonio origlia, and francesco cutugno this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). di maro, origlia, cutugno as exemplified by the contrasting examples of systems like alexa and siri and the advent of gptpowered language models. voice assistants such as alexa and siri are characterised by their limited set of interaction behaviours, which places a spotlight on the need for further advancements in conversational ai. similarly, recent developments in gpt-powered language systems have garnered both admiration for their linguistic capabilities and concerns about their potential misuse while also having limitations in their conversational abilities. we aim to contribute to the ongoing evolution of dialogue interaction but also to provide linguisticallymotivated models to help address the complex challenges and opportunities posed by the ever-expanding landscape of ai-driven communication. more specifically, our goal is to investigate theory-based strategies to identify inconsistencies in the common ground, the set of knowledge shared by the interlocutors in a conversation (clark and brennan, 1991). despite the documented pragmatic research in the field of conversational ai, some limitations have been highlighted. in de wynter et al. (2023), for instance, the memorisation skill has been taken into account as an important parameter to assess the quality of the output in llms. the authors found out that a significant portion of generated content (about 80%, differently distributed across the models) involved memorisation. however, this memorisation often led to factual inaccuracies, coherence issues, and logical fallacies, where different models exhibit these problems on varying levels. moreover, a negative correlation between factual errors and memorised content was found, suggesting that the models’ ability to recall information impacted the accuracy of their output. the study confirmed the relationship between text originality and memorisation and extended this to overall discourse quality, indicating that models with higher memorisation capability tended to produce higher quality content. pragmatic strategies adopted in such situations are, therefore, an important research subject, as they are generally concerned with grounded information, and thus memorisation. this must, first of all, be studied from a linguistic point of view to address the documented shortcoming of current technological models. in the pragmatic analysis of conversation, the starting point is to define dialogues as joint activities, for which the goals of both the interlocutors, and their role in a particular interaction must be identified in order to reach the conversation targets (macagno and bigi, 2017). spoken interactions are cooperative, influencing the production of utterances in a given context. as pointed out by scholars like clark (1996), to pursue the aim to succeed in their joint activity, the interlocutors engage in a communicative process called grounding. in conversation analysis, grounding refers to the process of establishing that what we intend to say (or what has been said) can be well understood (or has been well understood) (clark and brennan, 1991). according to other scholars, such as allwood et al. (2000), grounding can refer to the determination of what level of perception and comprehension is deemed acceptable. this can vary depending on the type of activity and how critical the information is to that activity. additionally, the analysis introduces another factor influencing grounding behaviour, intended as the evaluation or assessment of how and why grounding (mutual understanding) happens. this holds even in scenarios like casual conversations or conflicts, where there isn’t a clear purpose beyond social interaction or disagreement related to the information being conveyed, similar to what is described in this work. to verify the establishment of a common ground during a dialogic interaction, different linguistic or para-linguistic feedback analysis strategies (traum, 1999) can be exploited. from a linguistic point of view, dialogue efficiency can rely on the analysis of communicative feedback, whose relevance was pointed out by allwood et al. (1992) and which continues to be considered a fundamental characteristic in dialogue modelling (buschmeier and kopp, 2018). in this work, specific attention is dedicated to computational 114 common ground inconsistencies pragmatics, with respect to the systems’ use of pragmatic tools in specific inconsistent contexts and their impact on the human capability to solve them. the significance of the grounding process in conversational systems has been emphasised in various ways, ranging from signalling uncertainty (fernández et al., 2007; hough and schlangen, 2017) to exploring different levels of grounding (roque and traum, 2008; roque, 2009; petukhova et al., 2015), and even to the application in the evaluation of dialogue systems (curry et al., 2017; zou, 2020). specifically for the problems concerning inconsistency check and signalling, data we collected in previous studies (di maro et al., 2020, 2021a,b), summarised in section 2, motivate the experiments we carried out in this work. various scholars have highlighted the importance of including dialogue exchanges that make use of corrective utterances in their systems to improve the communication process (bousquetvernhettes et al., 2003; bohus, 2007). this resulted from the users’ need to interact with an agent capable of cooperating through communicative actions. human interlocutors always contribute with questions, answers, and feedback (beun and van eijk, 2004). a corrective dialogue is a particular type of dialogue made up of sequential acts which occur when: i) the user notices an error in the system and corrects it; or ii) the user changes their mind; or iii) the user’s beliefs are in contradiction with the system’s beliefs and expectations. in the first two cases, the corrective dialogue is initiated by the user, whereas, in the last case, it is initiated by the system (bousquet-vernhettes et al., 2003). one example of corrective dialogue in human-machine interaction is the one presented in beun and van eijk (2004). the authors focused on a particular communicative problem related to conceptual discrepancies between a computer system and its user. there are two potential scenarios to consider: • the system may detect when the user mistakenly applies an incorrect action to a specific object; • the user’s dialogue contribution might include inaccurately chosen words, words placed in the wrong order, or an incorrect combination of words. concerning specifically the second scenario, the authors reported the following example: if a user asks the system to ‘edit process 308’, the system could infer that, contrary to its prior assumption, the user believes that processes are editable. starting from the hypothesis that both the computer and its user have a mental representation of a domain, the mental representation on the computer side, also referred to as ontology, contains conceptualisations that are made explicit in a formal language. although these conceptualisations are usually incomplete and inaccurate, they can be used to trace the system’s reasoning about the concepts, items, and their properties. this representation also allows the detection of conceptual discrepancies, for example when the system observes that the user applies an incorrect action to a particular object. the authors also stated that, although feedback is now used in such systems, there is still no accurate ‘mathematical theory’ for natural communicative behaviours and their computational model to human-machine interaction, especially as far as conceptual discrepancies are concerned (beun and van eijk, 2004, p. 2). according to the authors, what is still missing is, therefore, a reference model guiding the adoption of a specific type, content, and form of the feedback that has to be generated in a particular situation (beun and van eijk, 2004), a gap that we investigate in this work. general mathematical models describing a linguistic theory are only partially implemented in technological systems. in prakken (2018), for the case of argumentation-based dialogue, it is highlighted that ad-hoc solutions are often presented for the task at hand. moreover, the most recent 115 di maro, origlia, cutugno figure 1: example of the clarie system’s capabilities to use clarification requests (source: purver (2004, p.277). approaches relinquish the task to machine learning models, which, however, build statistical models, uninformed by underlying reasons. for example, the choice of the best feedback to use could depend on different factors: i) the domain knowledge in both the system and its user: more specifically, the system’s knowledge about the user’s conceptualisation; ii) the role played by the system in the interaction (i.e., whether the system is a the expert or not), a parameter which might affect the definition of some of the the ontology features. an example of corrective dialogue, and the forms adopted, in conversational agents is given, for instance, in purver (2004, 2006), where the system (clarie) is capable of handling clarification requests uttered by human users (figure 1). here, confirmation requests in the form of polar questions are also used as shown in s5, s7, u8, and u9. in this work, a type of corrective dialogue is investigated, in which a simulated system has a non-expert role and initiates repair strategies for its grounded knowledge, when conceptual discrepancies, or inconsistencies, in the sequence of actions uttered by the user occur. the types and forms of feedback are investigated here, not only as far as the appropriateness is concerned, but especially for the practical effects that act on the interaction itself. we are particularly interested in polar questions, which are defined as questions that make relevant affirmation/confirmation or disconfirmation (stivers and enfield, 2010). polar questions can have two possible binary answers: true versus false. more in detail, the following research questions will be investigated: • main question: does linguistic feedback in the form of different polar questions influence the capability of people to identify the cause of reported common ground inconsistencies? • secondary question: do different polar question forms influence the speed with which people take decisions when searching for the cause of reported common ground inconsistencies? the study on inconsistencies within the common ground is driven by the need to delve into the practical aspects of this phenomenon. specifically, when communication is framed in the context of a collaborative effort to accomplish an operational task, the different use of language forms may have a direct impact of task performance, beyond perceived naturalness. in light of this, we have undertaken a comprehensive examination, primarily concentrating on its operational implications. a previous linguistically-informed experiment, which involved the collection of introspective data 116 common ground inconsistencies (di maro et al., 2021c), has paved the way for the investigation outlined here. the main objective of this follow-up experiment is to measure how the capability of human subjects to solve inconsistencies, signalled by a dialogue system, changes when, in the case of positive bias/negative evidence conflicts, a negative polar question (i.e., should i not have added salt?) is used, compared to the use of a positive polar question (i.e., should i have added salt?). by doing so, we aim to provide valuable insights into how inconsistencies in the common ground can be managed and leveraged effectively in various linguistic and communicative contexts. this research is not only significant in advancing our understanding of computational pragmatics but also holds the potential to inform real-world applications and improve machine-to-human communication strategies. the methodology we propose in this work has been developed specifically to answer these questions and it has been designed to artificially elicit inconsistent instructions in the participants. also, the experimental procedure aims to create the illusion of misinterpretation by the subjects to represent the situation of interest as faithfully as possible. the paper is organised as follows: in section 2 the background theory on polar questions and their pragmatic functions is summarised along with the effects on usability in human-machine interaction we expect to observe; section 3 describes a set of experiments designed to answer the above-mentioned research questions, as far as conflict detection and speed of the detection itself are concerned; section 4 provides the results of those experiments, showing that the use of different polar questions can influence the capability to identify previously-undetected conflicts with a statistically significant difference: concerning the main question, high negative polar questions resulted in a higher percentage of conflict detection, whereas for the secondary question the speed conflict detection is considered; in section 5 the theoretical consequences of the obtained results are discussed and formalised; section 6 draws the conclusions. 2. background theory among the linguistic feedback used in corrective dialogues to convey specific epistemic meanings, polar questions are frequently explored. polar questions usually encode in themselves not only a mere request but also presuppositions, agendas and preferences. furthermore, the use of a polar question can also implicate a disaffiliation, meaning the act of opposing something a co-participant has said or done (steensig and drew, 2008). in this case of the reference to the informational content, we can, therefore, also speak of epistemically biased questions. according to the literature, one way of expressing disaffiliation is through the use of reversed polarity questions, which are questions that convey bias towards the opposite valence than the utterance (koshik, 2002, 2005). for example, negative interrogatives can also function as positive assertions challenging the recipient’s position and vice versa (i.e., could i be more wrong?) (heritage, 2002). criticisms and challenges can also be expressed through declaratives (i.e., you shouldn’t have done that), imperatives (i.e., don’t do that to me again), or exclamations (i.e., how dare you?), which are perceived more confrontational and explicit and can be therefore face-threatening (hayano, 2013; sidnell and stivers, 2012). among non-standard communications, conflicting representations (huang, 2017) are listed as interactions taking place when a discrepancy between what is communicated and what is believed by the agent occurs. in these scenarios, polar questions can, therefore, serve as a knowledge challenging tool (i.e., you sure about that?). various authors have pointed out how either the original bias of the speaker or the contextual evidence bias could influence the syntactic form of polar questions. for example, according to 117 di maro, origlia, cutugno bias positive neutral negative evidence positive ppq/rpq rpq neutral hnpq (outer) ppq negative hnpq (outer/inner) lnpq table 1: results for preferred polar question forms per pragmatic cell in english and german, redrafted by domaneschi et al. (2017); hnpq refers to high negative polar questions, lnpq to low negative polar questions, ppq to positive polar questions, rpq to really-positive polar questions. ladd (1981), in english, high negative polar questions (i.e., isn’t there a good restaurant nearby?) is mandatorily used to express the original speaker bias. whereas a positive bias for a specific contextual evidence is highlighted with a positive polar question (i.e., is it raining?). we define these basic concepts as follows: original speaker bias “belief or expectation of the speaker that p is true, based on his epistemic state prior to the current situational context and conversational exchange” (ladd, 1981, p. 166). contextual evidence bias “expectation that p is true (possibly contradicting a prior belief of the speaker) induced by evidence that has just become mutually available to the participants in the current discourse situation” (buring and gunlogson, 2000, p. 7). in domaneschi et al. (2017), possible combinations of original bias of the speaker and contextual evidence were investigated, to point out the influence they may have on the choice of polar question forms. this contrast represents, indeed, the conflict existing between the presupposed knowledge of the questioner and that of the answerer. the experiment was carried out by the authors for english and german but the same observations, with small variations, also appear to be valid for italian (di maro et al., 2021c). the result of the reference study, in table 1, shows that both the original bias and the bias derived from the contextual evidence interact in the selection of the appropriate question: in both languages positive polar questions (ppq) are typically selected when there is no original speaker belief and positive or non-informative contextual evidence is provided; low negation questions (lnpq, i.e., do you not...?) are most frequently chosen when no original belief meets negative contextual evidence; high negation questions (hnpq, i.e., don’t you...?) are prompted when a positive original speaker belief is followed by negative or non-informative contextual evidence; positive questions with really (rpq) are produced most frequently when a negative original bias is combined with positive contextual evidence. regarding hnpq, we can distinguish two readings in the column with positive bias and neutral or negative evidence. ladd (1981) referred to the outer negation reading when the speaker wants to double check p, and the inner negation reading in which the speaker wants to double check ¬p. in the inner reading, negation is part of the proposition being checked, whereas in the outer reading it is not. the two readings can be distinguished by the presence of positive polarity items (i.e. some, already or too), and negative polarity items (i.e. any, yet, either) (domaneschi et al., 2017). the results provided in domaneschi et al. (2017) and in di maro et al. (2021c) indicate a preference for specific forms of polar questions to signal different kinds of conflict. nevertheless, it is important to remember that these are intended as tendencies, as the authors specify that differ118 common ground inconsistencies ent forms are indeed possible in the described scenarios. for example, in a positive bias/negative evidence scenario, ppqs can be adopted alongside hnpqs, according to the strength of the bias, the role of the interlocutors, and their intentionality. from an interaction design point of view, this has the potential to inform about the form in which to present a confirmation and/or clarification request as a polar question, depending on the pragmatic context. the experiment resulted in the identification of the most appropriate form according to the type of conflict. this result was used as a starting point for our study which is conversely focused on the resulting functionalities of the appropriate italian forms, in terms of robustness. starting from some human-computer interaction usability principles, we can point out some properties which can also be investigated to evaluate the effect of confirmation requests form on interaction quality. in dix et al. (2003), three main interaction design principles are listed, namely learnability, flexibility, and robustness. in this work, we focus on system robustness. this is defined as the level of support that the system provides to the user in completing and assessing a task successfully; this can also be ensured by the ability a system can have to check message understanding and correcting alleged errors in order to successfully complete the required tasks; the following principles are applied to support system robustness: 1. observability, which refers to the possibility of observing the internal state of the system; this can be further represented by five different principles: i) browsability allows the user to explore the internal state without modifying it, ii) default suggests the user possible actions, iii) reachability enables the navigation through observable states, iv) persistence refers to the duration of an observable state, v) task performance includes the services supporting all the possible tasks; 2. recoverability, which is the ability a system has to recover in case of errors; error recovery can act forward and backward: i) forward error recovery refers to errors in the current state causing a negotiation from that state to the desired one, ii) backward error recovery aims at correcting the effects of previous states in order to return to a preceding state; common ground clarification requests function, indeed, as backward inconsistencies recovery of grounded information; 3. responsiveness, which is the time the system need to give feedback and communicate with the user; 4. task conformance, which refers to the level of support a system offers when a task is executed in an expected way. related to the robustness principle, systems can make their internal states observable through verbal or non-verbal interaction. specifically, when problems occur in information processing, the observable characteristics of such states can be utilised to recover from the problems. from an interaction design point of view, the results obtained in domaneschi et al. (2017) and di maro et al. (2021c) have a potential effect on the principle of robustness. in particular, what can actually take advantage of the correct use of clarification requests in terms of consistency between the form used and the type of the common ground problem, namely common ground inconsistencies, that can be related to observability (here, the main question) and recoverability (here, the secondary question) features. we, therefore, investigate if different forms of polar questions used to signal an artificially created inconsistency, in series of commands from a human user, lead to a better handling of problematic situations characterised by conflicting statements. 119 di maro, origlia, cutugno 2.1 common ground inconsistencies information stored in the common ground may be subject to revision when inconsistencies occur. starting from a formal definition of the possible conflict situations, we focus on the specific problem of bias/evidence conflict. given a domain d containing domain items (e.g. cutlery, ingredients, etc.), we define an ordered sequence of actions a on items belonging to d. since every action has different pre-conditions and consequences, each ai ∈ a is associated with a set si composed of verifiable pre-conditions [pre(a1, pr1), . . . , pre(ak, prq)] and post-conditions [post(a1, po1), . . . , post(ak, pos)]. note that some actions may have multiple preor postconditions (i.e. [pre(a1, pr1), pre(a1, pr2)]. also, some actions may not even have preor post-conditions. in these, poi and pri parameters represent propositions that must, in the case of pre-conditions, be verified and, in the case of post-conditions, immediately become true. we will refer to generic propositions using p in the rest of the paper. during the course of the dialogue, applying post-conditions updates a set of propositions p , containing what is true and what is false about items in d. verifying that actions can actually be performed implies verifying that target propositions in p hold. consequently, executable actions automatically introduce or remove propositions in p . conflicts may, therefore, occur in the following conditions: i) a pre(a, p) is incompatible with the rules of the communal common ground (ccg)1, including common sense, resulting in an impossible action; for instance, grind the milk is impossible, since the pre-condition of the action [grinding] is to have an ingredient which is solid and not liquid or powder; ii) a pre(ai,¬p) ∈ si is not verified because of a post(aj , p) ∈ sj resulting from a preceding aj , saved in the set of shared knowledge the personal common ground (pcg)2. this causes what we call an inconsistency. concerning the second point, which is the main interest of this work, we call this type of conflict common ground inconsistency, with reference to the incompatibility between the listener belief and the new evidence provided by the speaker. for instance, as above-specified, the action of [grinding] requires a solid ingredient. the action of [grinding] results in the ingredient becoming powder as post-condition. this means that, after this action, an action like [cutting] cannot be performed on the same ingredient as the previous post-condition causes the pre-condition of [cutting] (i.e., solid) not to be verified. clarification requests can be in this case adopted as a corrective feedback. it is worth noting that, in this work, we are not concerned with the solution to the problem, which may vary (i.e., cancel the previous action, insert corrective actions, etc. . . ). we only concentrate on the detection and signalling of the problem between users. in figure 2, the scenario eliciting a common ground clarification requests is displayed. in the representation of the female agent a, the ccg is stored to guide the process of accumulating information in the pcg. the information (i1, i2, i3, ..., in) are communicated by the male agent b to a, and sequentially stored in her pcg. when b utters a new information iz , this is represented as a new item candidate to be part of the pcg. this representation has disastrous results, in that the presence of the new item iz in the pcg clashes with the presence of another item i3, whose validity is now 1. the amount of information shared with people that belong to the same community (clark, 2015) 2. the amount of information collected over time through communicative exchanges with an interlocutor (clark, 2015) 120 common ground inconsistencies figure 2: representation of the common ground clarification requests elicitation scenario. questioned. this conflict represents a common ground inconsistency and is translated in the common ground clarification request ¬i3?, whose form, function and illocutive effect are analysed in the next chapters. as already highlighted, it can be pointed out how important polar questions are to common ground inconsistencies, in that their epistemic stance (or presuppositional stance) is clearly expressed compared to other types of questions. finally, common ground clarification requests do not necessary refer to the immediately previous utterance, but to previously, alleged wrongly, grounded information. in this work, we, therefore, investigate different polar questions in common ground inconsistencies scenarios and how they can influence the capability of people to identify the cause of the conflict. 3. experimental setup the type of dialogue that we concentrate on to investigate the problem at hand can be divided in two sub-dialogues: • instructing phase: the participant produces a sequence of utterances representing instructions for the interlocutor to execute; • repair phase: when a logical inconsistency occurs, the participant has to re-examine the sequence of instructions they gave, according to the indication given by the interlocutor. reexamination is intended here only in terms of conflict identification and not resolution. to evaluate the effect of different question forms, we needed to replicate a situation as close as possible to a natural one. specifically, we create a situation where people unintentionally give a conflicting sequence of instructions, and verbalise them while not realising the mistake until this is signalled to them by the interlocutor. for our case, it is, therefore, important to elicit the incorrect sequence of instructions from participants by making them verbalise it and by artificially creating 121 di maro, origlia, cutugno the impression of having committed a mistake. in the following subsections, we will describe the experiment in terms of its objectives, reference domain, expected results and hypotheses, and methods adopted in the development of the setup. 3.1 rationale to our knowledge, the most widely used experimental procedure to elicit dialogues that include conflict resolution is the map task (anderson et al., 1991). we draw inspiration from this particular setting to elicit conflict resolution in our domain with two main different aspects to consider: • the map task presents a global misalignment represented by the different maps. in our case, we concentrate on specific elements in the instructing phase; • in the map task, people are allowed to realise that the maps are different. in our case, after artificially creating the error, we also hide it during the repair phase. we propose an experimental procedure to create an apparently consistent sequence of instructions that is later contradicted by another instruction. furthermore, we aim to create the impression in the subjects of having committed an interpretation mistake when verbalising a target instruction. the goal of the experiments presented in this work is to evaluate how, given a specific pragmatic situation, different italian polar question forms can impact the quality of the interaction. specifically, we concentrated on conflicting situations in which a statement is accepted but, in a later stage of the dialogue, makes another statement unacceptable. both statements, per se, must be acceptable, with conflicts arising only from their combination. this is consistent with the stimuli used in domaneschi et al. (2017) and di maro et al. (2021c) to investigate the appropriateness of the different question forms in each possible conflict category, as explained in section 2. to avoid considering dialogue situations influenced by personal interests, we focus on a particular type of dialogue, namely the deliberation dialogue (walton and krabbe, 1995; prakken, 2018), characterised by the process by which two or more agents reach a consensus on a course of action in a collaborative way. this type of dialogue takes advantage of argumentative capabilities, thus belonging to the category of argumentation-based dialogues. argumentation-based dialogue refers to the modelling of the verbal interaction aimed at the resolution of conflicts of opinions via the adoption of specific strategies. this field of study consists of a variety of different approaches and individual systems, with few unifying accounts or general frameworks (prakken, 2018). more in detail, starting from domaneschi et al. (2017), we take a different angle on two points: 1. the bias is constructed through an actual sequence of utterances given by the subject according to the semantic content shown on a sequence of slides the subject is looking at; 2. the subject is asked to perform a task on the basis of the information that can be extracted from an error prompt using different forms of polar questions. in fact, while in domaneschi et al. (2017) and di maro et al. (2021c), the appropriateness is considered, here the functionalities of the forms are taken into account, as a task is depending on the type of question. the considered situation is particularly interesting to guide dialogue systems design, as it corresponds to cases in which the machine’s internal representation of the common ground is made inconsistent not because of the last incoming command but because of previous actions. in such cases, observability and recoverability are influenced by the capability of the machine to communicate information concerning the problem in a natural and synthetic way. 122 common ground inconsistencies 3.2 domain description let’s consider a set of cooking recipes r = {r1, . . . , rn}, ∀ri = {si, ai, gi, ti}, where si is the series of slides representing ri, ai the actions, gi the ingredients, and ti the cooking tools for each ri. each recipe has a different {si, ai, gi, ti} tuple coherent with it. the experiment stimuli were composed of series of slides si = {s1, . . . , sn} and, for each of them, participants were asked to elaborate spoken commands. the slide series defines the recipe. each slide si was designed to represent an action in the recipe ri belonging to the cooking domain. this domain was chosen for three important reasons: i) the familiarity with this domain is presumably high among speakers, being part of everyday life; ii) similarly to the map-task (baker and hazan, 2011), this domain could be applied in a deliberation dialogue; iii) contrary to the traditional map-task, the number of different actions is higher, making the tasks more varied and slightly more articulated; moreover, single actions, although atomic, are often linked to each other, in the sense that an action can affect a consequent one in ways that may not be immediately evident. the use of visual stimuli was adopted to avoid influencing the participants’ production with linguistic material. also, to make the task less imposing, the slides’ structure was kept coherent and the representation strategy was designed in such a way that the same action would always be represented by the same image. hence, each si was represented using a fixed structure: given ai ∈ ai, the action involved in si, ai was represented on the left side of the slide through an animated image3; given the set of ingredients gi and the set of cooking tools ti, the set of parameters represented on the right side of the slide through static images was taken from the set p = gi∪ti. figure 3 shows an example of visual stimulus for the action grinding applied to the ingredient nutmeg. to simulate the occurrence of conflicting situations, we replaced a slide s′x representing a correct action in the original recipe with a slide sx introducing an inconsistent action in s. as already specified, the action depicted in sx is acceptable, and therefore not impossible, so the subject can continue providing commands with no error prompts when sx is presented. the inconsistency emerges when the last action in sequence, sk, cannot be performed because of sx. this inconsistency, in the form of a contrast between positive bias (having to do x) and negative evidence (having to do something that is prevented by the consequences of doing x), was determined by the opposition of some aspects of sx and some aspects of sk. following sk, a further slide sq was presented, containing an error message formulated in different ways, thus creating the experimental conditions considered in this work. the system prompt message represented in sq, other than presenting the different forms of clarification requests considered in this work, also instructed the participants to go back through the recipe to look for the conflict, so that it was possible to observe the participants’ behaviour to evaluate the effect of using the different question forms on the task of error identification. once sq is presented and the subject has understood the new directives, they were free to use voice instructions to move in s. the conflicts introduced in sx were of two different types: quantity-related or ingredient-related. quantity-related conflicts refer to the situation where an ingredient is used without specifying the quantity; in fact, when this specification is missing, the interlocutor presupposes that after the action is processed the ingredient is no longer available. on the other hand, ingredient-related conflicts refer to ingredients which have been used in a preceding action instead of the correct ingredient. this makes them no longer available, although they would have been if the correct ingredient was 3. gifs were generated from video recipes taken from giallozafferano https://www.giallozafferano.it/ 123 di maro, origlia, cutugno figure 3: an example slide from the experiment; the represented action, on the left, elicits a command for the action grinding (more specifically, this action is generally represented by showing a cook using a grinder to grind lemon peel) applied to the ingredient nutmeg represented in the parameters’ slot on the right; the corresponding action to be uttered is, therefore, grind the nutmeg. this slide was translated from italian for the reader’s convenience to help describe the setting. in the following figures, the original slides in italian are shown. recipe code conflict type # slides béchamel r01 quantity 9 carbonara r02 ingredient 16 oat yoghurt baskets r03 ingredient 16 potato croquettes r04 quantity 14 pancakes r05 quantity 13 baked potatoes r06 quantity 9 piadina r07 quantity 8 tuna meatballs r08 quantity 11 tiramisù r09 ingredient 11 small pizzas r10 ingredient 12 table 2: recipes tested within the experiment with their conflict type description and the number of slides/actions used from the original recipe. 124 common ground inconsistencies used before. the relationship between conflict types and the 10 recipes used in the experiment is summarised in table 2. figure 4: experiment structure. people go through the actions sequence containing the slide with the error sx. then, they see the error message in slide sq and are instructed to go backwards to find the error. while they instruct the experimenter to do so, the experimenter actually goes forwards in the presentation, showing the reversed sequence containing the correct slide s′x. this creates the illusion to go backwards in the same sequence while they are actually moving thorough a reversed one. the slides were organised in such a way that, after the sq slide, the reversed s sequence was found. also, in the reversed sequence the original action in the recipe replaced sx, thus becoming consistent and noted as s′x. the experiment evolved in the following way (as illustrated in figure 4): • from the start time t0, people saw each slide of the first part of the sequence, from s0 to sk up to tn−1, seeing the incorrect slide sx at time tx; • at time tn people saw the error message slide sq and were instructed to go backwards to find the incoherence; • as time passed from tn+1 onwards, people thought they were going backwards in the sequence, while the experimenter actually moved the presentation forward, through the reversed sequence from sk backwards; • at time tn+k−x+1, people saw the correct slide s′x, as if it had always been the one included in the original sequence. for example, suppose that a recipe has 9 slides and that the fourth slide s3 contains the error. the presentation contains, therefore, 19 slides: the 9 slides representing the recipe with the error, the error message slides, and the same 9 slides representing the recipe mounted backwards, with the correct slide substituting the wrong one. people go through the sequence from s0 to s8 from time t0 to t8 and they see the wrong slide s3 at time t3. at t9, they are shown the error message slide sq. 125 di maro, origlia, cutugno from time t10 to t17, they believe they are going backwards in the sequence, while the experimenter actually moves forward, through the reversed sequence, so that they see the correct slide s′3 at time t17. summarising, after the error message, the time index keeps increasing, while the slide index decreases. this created the effect of showing the reversed sequence to the subject, as if they were going back, while substituting the incorrect sx slide with the correct slide sy, as shown in figure 4. this is intended to create a situation as close as possible to the one where people actually pronounce the incorrect command while believing it is correct and proceed with the task. when the error is made clear, they can attempt to find to the command causing the problem. since it is known that verbalisation can reinforce the memorisation of visual stimuli (weatherford et al., 2021), the experiment induces the subjects in verbalising wrong commands to recreate a believable situation, in which the possibility to detect the issue independently from the type of confirmation request is still present. this is especially important to obtain a fair baseline for the presented comparisons. the task ends when people renounce or indicate a slide as the one containing the error. the task ends regardless of whether people are correct or not. furthermore, to better represent a situation in which an error occurs, participating subjects were not informed that there was the possibility that an elicited command could give rise to an error. this way, they would not force themselves into remembering precisely the previous commands they gave. furthermore, to reduce the possibility that the sx was in the subject’s short term memory, sx was introduced at a minimum distance of 5 slides from sk. finally, to better exemplify the experiment, in figure 5 and figure 6, we reported the sequence of slides used for the recipe piadina, highlighting s0 as the source of quantity-related incoherence. the recipe used comprises the following actions: s0 put the flour in the bowl s1 mix salt, lard, baking soda, and water to the flour s2 stir the mixture s3 add water to the mixture s4 stir the mixture s5 add water to the mixture s6 knead the dough s7 put the flour on the counter sq error message [... ] s′0 put part of the flour in the bowl 126 common ground inconsistencies figure 5: the actions sequence in the recipe during the first phase of the experiment, up to the error message slide sq (english translation of the above reported message: should i not have added the flour to the bowl? go back through the recipe until you find the action that prevents the last one from being carried out.), and containing the error slide s0. 127 di maro, origlia, cutugno figure 6: the reversed actions sequence in the recipe, shown during the second part of the experiment, containing the correct slide s′0 3.3 hypothesis as previously mentioned, the goal of the experiment was to check if different system error messages, shown in sq, were more or less efficient in signalling to the user the existence of a conflict arising from sx and its details in a succinct, natural way. our initial hypothesis is that, similar to what has been observed in other production and introspective studies (domaneschi et al., 2017; di maro et al., 2021c), in the conflicting situation presented here, the adoption of an hnpq results perceptually and operationally in a more effective conflict detection. more specifically, we expected that in expressing conflicts between previous beliefs (positive bias) and opposing contextual obser128 common ground inconsistencies 1st recipe 2nd recipe p1 r09 r08 p2 r07 r03 p3 r03 r02 p4 r01 r07 p5 r06 r01 p6 r04 r10 p7 r02 r05 p8 r05 r06 p9 r10 r09 p10 r08 r04 p11 r07 r05 p12 r05 r07 table 3: recipes’ distribution for each experimental session. vations (negative evidence), the use of a negative polar question would: a) improve the capability of the participants to identify the error, suggesting that the use of such questions would improve the observability degree of a dialogue system, i.e., main question, and b) decrease the required effort to intervene on the system to address the inconsistency, i.e., secondary question (section 1). 3.4 methods instructions given to the participants were presented in the first slide of the experiment and contained the following text: in this experiment, we ask you to look at a series of slides describing a recipe and tell the experimenter what to do to make it. each slide is made up of actions and parameters: actions generally describe what to do, while parameters show which objects are involved in the action. in case of problems, you are free to move back and forth in the recipe as you like by asking the experimenter in which direction to move the slides. take your time to think about what to say and give your instructions when you are sure they are correct. we first start with a training recipe, in which you are free to ask any questions you want to the experimenter. once the test starts, the experimenter will no longer be able to answer you, except for carrying out your instructions until the end of the experiment. a total of 36 participants was recruited for the experiment, each of which had to perform 2 tasks (72 tasks were collected in total). participants were divided into three, gender balanced, groups, one for each experimental condition: • control group: the first experimental condition consisted of two tasks combining two different recipes as shown in table 3; this condition was used both as validation for the experimental setup, in order to understand if the slides and the task were understandable for the participants, 129 di maro, origlia, cutugno and as analysis of the general error message which was used to signal the conflict to the participant (i.e., this action is not possible because it clashes with a previous one). the resulting collected values were, therefore, used as a term of comparison for statistical analysis. • ppq group: the second experimental condition differed from the previous one just for the typology of error message presented to the participant, where a positive polar question was instead used (i.e, should i have added the flour to the container?). the use of the most frequent polar question form was useful to test its appropriateness in bias-evidence conflicts in simulated human-machine interactions. • hnpq group: the third experimental condition, similarly as the previous one, made use of a negative polar question, and more specifically of a high negation polar question in the past tense (i.e, should i not have added the flour to the container?), whose appropriateness in the positive bias versus negative evidence scenarios was confirmed in the experiments described in (domaneschi et al., 2017). for the control group, the average age was of 25.5, with an average self-evaluation of their cooking skills equal to 2.33 (on a scale from 1 to 5). for the ppq group, the average age was of 27.25 with an average of 2.67 self-evaluated cooking skills. finally, for the hnpq group, the participants were on average 26.08 years old and their average self-evaluated cooking skills was 2.92 points. no significant differences were found in the comparison of the self-evaluated capabilities among the groups. because of covid-19 restrictions, the experiment was carried out online. before presenting the two recipes used for the experiment, a training recipe was used to let participants familiarize themselves with the setting, to learn how to interpret the slides and to ask questions. this did not contain a conflict, but was just used to explain the structure of the experiment. after this training session, participants did not report significant challenges in interpreting the slides. an interaction example is reported below: recipes: zucchini "alla scapece" u1: cut zucchini; u2: add salt; u3: put the oil in a bowl; --error --u4: add salt and mint to the bowl; -------------correct -u4*: add salt and chili to the bowl; ------------u5: mix the mixture; u6: cut the garlic; u7: add garlic and mint to the mixture; error message: control group: "this action is not possible because it clashes with a previous one"; ppq group: "should i have added mint to the bowl?" 130 common ground inconsistencies hnpq group: "should i not have added mint to the bowl?" in this example4, we reported the sequence of utterances for a recipe. the erroneous action (u4) states to add the mint to the bowl, instead of chili (u4*). at the time u7 is uttered, mint is no longer available, as it was already used. at this point, the error message is prompted and the recipe can not proceed. the correct action, as reported in u4*, suggested, instead, the addition of another ingredient, i.e., chili instead of mint, which, instead, will be required later. the error message is different according to the experimental group, as previously described. 4. results at the end of the data collection phase, about 372 minutes (122 minutes for the control group, 145 minutes for the ppq group, 105 minutes for the hnpq group) of audiovisual recording were collected and annotated using elan (wittenburg et al., 2006). annotations, provided by the authors, were used to compute the results. they marked the slides boundaries appearing in the videos and the speech fragments containing subject-uttered commands. for the rest of the presented analyses, slide change times will be considered to investigate the proposed research questions. first of all, we verified that the error inserted in each recipe was found at least one time and that the error was never systematically found. this indicates that the considered recipes and the kind of error introduced were neither too difficult nor too easy. a summary of the percentage of times the error was correctly identified, in each recipe, is shown in figure 7. next, we consider the number of times in which participants were able to identify the slide containing the action that caused the conflict, according to the type of error message. a summary of the performance of each participant, for each experimental setting (control, ppq, hnpq), is shown in table 4, where the conflicts found are ticked, while aggregate counts are reported in table 5. results for our main research question show that, as expected, the slide causing the conflict was found most often in the hnpq configuration. since both the ppq and the hnpq configurations improved the capability of the participants to identify the error, the significance of the effect was checked using the binomial test, a non-parametric test for binary variables (wagner-menghin, 2014). the test showed that, when using hnpqs, the conflict was found more frequently, with respect to the control group, in a statistically significant way (p = 0.005). this result is also consistent with hnpqs being the preferred form of clarification requests for the considered conflicts in domaneschi et al. (2017) and di maro et al. (2021c). on the other hand, when using ppqs, the difference with the control group resulted not to be statistically significant (p = 0.4). although ppqs do provide more information about the problem than a generic error message does, since the counts of conflicts found is higher in ppq than in control group, its use is not statistically different from that of the generic error message. the interpretation of the results leads us to the following observations: • ppqs and hnpqs are not expected to be mutually exclusive in the considered situation: as a matter of fact, they are both reported to be acceptable but there is a difference in the ability of the subjects to detect errors. this is consistent with what was reported in previous experiments domaneschi et al. (2017); di maro et al. (2021c), where the frequency of occurrence of 4. translated from italian ppq: dovevo aggiungere la menta alla ciotola hnpq: non dovevo aggiungere la menta alla ciotola? 131 di maro, origlia, cutugno figure 7: percentage of times conflicts found per recipe. control ppq hnpq 1st recipe 2nd recipe 1st recipe 2nd recipe 1st recipe 2nd recipe p1 ✓ ✓ ✓ ✓ p2 ✓ p3 ✓ ✓ ✓ p4 ✓ ✓ ✓ ✓ p5 ✓ ✓ ✓ p6 ✓ ✓ ✓ p7 ✓ ✓ p8 ✓ ✓ ✓ ✓ p9 ✓ ✓ ✓ p10 ✓ ✓ p11 ✓ ✓ ✓ ✓ p12 ✓ ✓ ✓ table 4: distribution of conflicts found by each participant: for each experimental group (control, ppq, and hnpq), found conflicts are ticked per each participant for each corresponding recipe (1st and 2nd recipe). 132 common ground inconsistencies 1st recipe 2nd recipe total percentage control 5 4 9 37.5 ppq 7 4 11 45.83 hnpq 10 6 16 66.67 table 5: number of conflicts found in the three experimental setups (per recipe, in total, and in percentage). hnpqs was higher than ppqs in the considered type of conflict. ppqs, however, were not completely unobserved; • since the only difference between the two forms is the negation element non, both convey the same amount of informative content more than the baseline but only the hnpq form outperforms it in a statistically significant way, supporting a view that hnpqs do provide an advantage in functional terms; • the acceptability of the two forms in previous experiments also explains the non statistically significant difference between hnpqs and ppqs: while they can convey the intended message, in general, being used more frequently to communicate other kinds of problem (section 2), also causes them to be sometimes misinterpreted in this pragmatic condition. support to the third point was partially provided by the feedback of one participant, who spontaneously reported that the ppq led him to think that the question was referring to the last presented slide, rather than to a previous one, as if it was a different kind of confirmation. in fact, while ppqs can be considered as a grounding act, which verify the correctness of the previous discourse unit, hnpqs function more as an argumentation act, referring to a global rather than local problem (traum, 1999; di maro, 2021a). this suggests that using the appropriate syntactic form to convey pragmatic meaning when signalling bias/evidence conflicts may not only be a generic preference or a mere matter of appropriateness but it may actually improve the quality of the conveyed message, as people are indeed able to identify problems more easily. from the previous analysis, it was found that people tend to make the right choice, in addressing a p/¬p conflict, more frequently when this is signalled using a hnpq. a performance decrease is observed during the second round of the experiment, in all the conditions: it seems reasonable to attribute this to fatigue effects rather than factors due to the stimuli because it is consistent among the participants and is not caused by the recipes themselves because they all appeared both as the first sequence and as the second one. to answer to our secondary research question we needed to evaluate how efficiently people reach a conclusion after being presented the error prompt. for this reason, the sequence of steps followed when searching for the conflict was considered. as participants may take a quick decision but indicate the wrong slide or they may simply rapidly decide they are not able to find the conflict, it is not possible to simply consider the time spent looking for the conflict. in the indicated cases, for example, a short amount of time to reach a wrong conclusion or to quit the task would be considered equally to cases in which people rapidly got to the right choice. for this reason, the sequence of steps followed by participants during the search phase was compared with the ideal sequence that would lead a person who has understood the problem to the problematic slide. for the comparison, the dynamic time warping (dtw) (müller, 2007) algorithm was used and, specifically, warping 133 di maro, origlia, cutugno control hnpq ppq 0 1 2 3 4 group d t w d is ta n ce figure 8: box plots representing distances distribution in the three experimental conditions. control hnpq hnpq 0.089 ppq 0.75 0.39 table 6: pairwise comparisons using wilcoxon rank sum test with hölm adjustment. distances were considered as an indicator of how efficient the observed sequences of steps were with respect to the ideal one, which directly reaches the problematic slide. an example of a user who goes directly to the conflicting slide is given in in figure 9a, whereas an example of a user who is not sure of which slide caused the problem and goes back and forth in the sequence is shown in figure 9b. in absolute terms, users explored the sequence in a closer way, on average, to the reference curve when the hnpq was used, when compared with a general error message and a ppq. to check whether the average differences were significantly different, the shapirowilk test (shapiro and wilk, 1965) was used to check for normality. since the distributions were not normal, the kruskal-wallis test (kruskal and wallis, 1952) was used to check the differences shown in figure 8. in this figure, the distributions of dynamic time warping distances, for each experimental condition, are shown. the difference was found not to be statistically significant in any case. pairwise comparisons were performed using the wilcoxon rank-sum test with the hölm adjustment to further detail the situation, shown in table 6. therefore, while it is confirmed that, by using hnpqs, participants identify the problem more frequently, when they understand the kind of mistake they should look for, we could not find evidence, in our data, that they find it in statistically different amounts of time. 134 common ground inconsistencies (a) an observed sequence (lower graph) from a user who understood the problem after reading the prompt compared with the optimal sequence (leftmost graph). the number of turns needed to reach the conclusion is exactly 13 as expected in the reference graph. the user directly goes to the wrong slide in the numbered sequence (5), leading to perfect alignment. in this case, the dynamic time warping algorithm reports a warping distance of 0 (no alignment effort needed). (b) an observed sequence (lower graph) from a user who repeatedly moved through the slides sequence looking for the problem after reading the prompt compared with the optimal sequence (leftmost graph). the number of turns needed to reach the conclusions is much higher than expected and the user keeps going back and forth in the numbered slides sequence. in this case, the dynamic time warping algorithm reports a warping distance of 1.93 (significant effort). figure 9: evaluations examples with dynamic time warping. for each figure, the reference graph, on the left, represents the optimal sequence of steps through the recipe to get to the error slide and report the error, thus terminating the exploration. the lower graph represents the observed sequence of steps taken by an example subject before reporting the error. in the first case (a), the user immediately found the error, producing the ideal sequence of steps. in the second case (b), the user went back and forth through the sequence of slides multiple times before identifying the error. the central graph shows the alignment between the two sequences, produced using the dynamic programming algorithm for classic dtw, as reported in müller (2007). the alignment effort represents the dtw distance between the reference sequence and the observed one. in the reference graph, the x axis represents the slide number and the y axis represents the step. the axes are inverted, in the lower graph, with the y axis representing the slide number and the x axis the number of steps. the central graph x and y axes correspond, respectively to the x axis of the observed sequence and to the y axis of the reference sequence. this represents the optimal alignment between the two sequences. 135 di maro, origlia, cutugno 5. discussion this section describes how the application of the results obtained from an experimental methodology on the operation of certain linguistic forms in specific pragmatic contexts can be formalised for the purpose of framing them in a developing theoretical framework for argumentation-based dialogue. our discussion will now cover how these findings contribute to the design of real dialogue management systems. we believe our results have an impact in the following ways: • the kind of conflict we studied has real-world implications for designers of dialogue systems, as the syntactic form in which the error is presented, depending on the nature of the error itself, does have an impact in the users understanding it; • the conflict we analysed can be represented in formal terms so that dialogue systems can do the necessary inference to recognise and respond appropriately. we will show in the rest of this section how the use of graph databases enables this. the presented approach is designed to be compatible with the system architecture assumed by the framework for advanced natural tools and applications with social interactive agents (fantasia) (origlia et al., 2019, 2022). fantasia5 is a plugin for the unreal engine designed to support the development of embodied conversational agents. from a terminological point of view, we will adopt some of the concepts presented in fantasia. first, the type of system in which the pragmatic strategies described here can be applied will be specified. then, details will be provided on the structuring of the knowledge it deals with, with its formalisation and conflict identification strategy. proving the efficiency of such a form in a defined pragmatic situation, like the one tested in the experiment presented in this work, can be used to good advantage in dialogue systems aimed at learning sequences of actions uttered by a human interlocutor. these applications require the user to have a leading role, and therefore a higher knowledge (k+), and the machine to have a subordinate one, corresponding to a lower knowledge (k-). this type of task can be considered as a sub-type of user-initiative tasks. in addition to the characteristics typical of a user-initiative system, in our model, the system checks for consistency based on shared rules. in such situations, the domain information given by the users builds the pcg, that is the set of information collected over time through communicative exchanges with an interlocutor. in other words, it can be considered as a record of shared experiences new to the receiving system, although the general knowledge of the domain are conversely already shared in the ccg, that is the amount of information shared with people that belong to the same community, that is to say, people that share general knowledge, knowledge about social background, education (schools attended, levels of education attained), religion, nationality, and language(s) (clark, 2015) (section 2.1). the user has, therefore, more knowledge (k+ position) of the domain with respect to the system, as the desired goal is known by the user. conversely, the system does not have the same k+ position. nonetheless, the structured pcg is used to build presuppositions with strong confidence, that make the system closer to the k+ position to the point that, in case of inconsistencies, the system can assume the role of questioner. in such a scenario, each set of actions a = {a1, . . . , an} contains both preand post-conditions ai = {pre con, post con}. preand post-conditions-based inconsistencies between two uttered actions occur when the post-condition resulted from a previous action is not compatible with the pre-condition of a new action, based on the rules of the ccg. the system 5. github.com/antori82/fantasia 136 common ground inconsistencies can use a knowledge representation module, in the form of a graph, to verify the compatibility with the pcg, as shown in di maro et al. (2021b). in this case a confirmation request is used. on the one hand, if an inconsistency occurs, the problem is recognised and signalled by using a hnpq, which resulted in an improvement in the efficiency of conflict detection with respect to the baseline. in fact, as already pointed out in domaneschi et al. (2017) for german and english, this polar question form is suitable to the type of conflict arising when a positive bias clashes with a negative contextual evidence. this is because a preceding action, part of the pcg, becomes the system’s presupposition, whereas the new uttered action is made impossible because of an unverified pre-condition, representing the negative evidence. if, on the other hand, the inconsistency is between a pre-condition and the rules of the ccg, the corrective feedback, would be of another type, i.e., explanation. previous experiments showed that ppqs and hnpqs can be both adopted in negative bias/positive evidence conflicts, but with a significant difference in the frequency of occurrence. however, ppqs are more frequent in other pragmatic situations (i.e., positive bias/neutral evidence), with a confirmation function, so that their use in the context studied in this work, while acceptable, may be easily misleading for the participants. while the use of ppqs, therefore, does not implicitly prevent the problem identification, it makes it harder to identify it, not producing a significant improvement with respect to the baseline. furthermore, it has to be remembered that the questions were presented in the written form. the interpretation of bias in polar questions can, conversely, be also affected by intonation (asher and reese, 2007; savino, 2012). participants could, therefore, have interpreted the questions with different intonations. further investigations will also be directed towards this level of analysis. concluding, our results show that, while it is possible for people to identify the problem when a ppq is presented, it is easier for the subjects to find the solution when the expected form, hnpq, is presented. the adoption of hnpq together with its details coming from conflict detection represents the argumentative capabilities needed in this type of dialogue. therefore, conflict detection can constitute a first building block to create a formal theory of argumentation-based dialogue centred on conflict patterns. formally, a pcg is inconsistent if a post-condition of a certain action aj introduces a proposition pj such that pj is in conflict with a subsequent pre-condition for an action ai. this can be formalised as: inconsistent(pcg) =⇒ ∃ai ∈ a,∃pj ∈ p |pre(ai,¬pj) (1) given that the set of propositions p is generated during the dialogue, this corresponds to the pcg. therefore, when a pre-condition pre(aj ,¬p) is not verified because of a post-condition post(ai, p), it is possible to report the cause of the inconsistency, beyond the mere existence of the problem, by reporting ai. in this case, consistently with domaneschi et al. (2017), a hnpq is generated. from the user’s point of view, this model represents our proposal that a human interacting with a dialogue system, when presented a clarification request in the form of a hnpq, is implicitly led to look for specific conflict patterns. this implies a revision of previous beliefs, incrementally accepted in the pcg, to identify conflicts. these would be characterised by the post-conditions of an action ai, in the sequence, causing the violation of the pre-conditions posed by the last requested action an, by establishing the truthfulness of a given proposition p that should not be verified, in order for an to be accepted. 137 di maro, origlia, cutugno figure 10: an example of the graph-based interpretation for the instruction grind the nutmeg. after verifying that the pre-condition of the state of the referent is solid, the consequence of the action is that the powder state is acquired. note that in this way the sequence of transformation is preserved to support clarification requests when necessary. given the previous formal representations of a consistent/¬ consistent pcg, the negative polar question implies a specific error pattern as follows: hnpq(aj) =⇒ pre(ai,¬p) ∧ post(aj , p) (2) this can be represented in the form of a graph, as suggested in di maro et al. (2021b), connecting actions and entities involved in the actions with their specific properties, as preand postconditions. specifically, di maro et al. (2021b) made use of neo4j (webber, 2012). neo4j is an open source graph database manager that has been developed over the last 16 years and applied to a high number of tasks related to data representation (dietze et al., 2016), exploration (drakopoulos et al., 2015) visualisation (jiménez et al., 2016) and dialogue management (di bratto et al., 2024). in neo4j, nodes and relationships may be assigned labels that describe the type of object they are associated with. neo4j is characterised by high scalability, ease of use and its proprietary query language: cypher. cypher is designed to be a declarative language that highlights patterns’ structure using an sql-inspired ascii-art syntax. in figure 10, an example showing how the domain of the interaction is represented using graphbased formalism is presented. in such a graph, using cypher, the conflict can be detected as follows: match (a1:action)-[:refers_to]->(e:entity)-[:assigned_to]->(fe: frame_element {name: 'patient'}), (e)-[:refers_to]->(pe1:perceived_entity) where not (a1)-[:is_followed_by]->() and 'powder' in labels(pe1) listing 1: cypher query checking the pre-conditions of the grinding frame for which a perceived entity cannot have the powder label in order to make the action possible 138 common ground inconsistencies this query lets the system identify a possible inconsistency, as the action refers to an entity, in turn, referring to a specific perceived entity in the interaction. for the action of grinding, reported in the query, the perceived entity must not have the label powder, according to the pre-conditions of the ccg, to not cause the inconsistency to occur. in case of inconsistency, with the use of another query, the action that caused the perceived entity to acquire the label powder, if present as a postcondition, is returned. in this way, the clarification request can be structured with the retrieved information. for more details, the reader is referred to di maro et al. (2021b). the practical consequence of such a model and its future extensions, considering other types of conflict, is a more efficient generation of error messages using the appropriate form of polar questions depending on the reported situation. 5.1 threats to validity the limitations of the present study in the context of dialogue and argumentation can be multifaceted. in this work, we limited the experimental context to a very specific situation to evaluate the impact of the using specific syntactic forms to signal the conflict of interest. specifically, we intended to measure the practical advantage of using patterns observed in linguistic studies to improve the quality of human-machine interaction. nevertheless, we are aware that dialogue and argumentation deal with many other important aspects which are not addressed, such as social dynamics, contextual factors, and user initiated questions. it’s important to consider these limitations to understand the relevance of the study and the possibility to generalise, as well as its implications for different types of dialogue, tasks, and objectives. one significant limitation of the study is its applicability to a specific type of collaborative argumentation-based dialogues. nevertheless, as also reported by other studies (domaneschi et al., 2017; di maro et al., 2021c), the use of hnpq in the described pragmatic scenario should be considered appropriate regardless of the domain and dialogue type. most dialogue systems, including ai-driven ones, can have imperfect processing and interpretation capabilities, which can also lead to inconsistencies. these problems were addressed in a previous study that classified clarification requests based on the type of problem, considering contact, perception, understanding, and intention as different communication levels (di maro, 2021b). common ground inconsistencies are a subset of problems at the understanding level. the exploration of other types of inconsistencies or communicative problems is left to future studies. moreover, since participants were asked to read the error message, one further limitation may lie in the absence of prosodic features which may have disambiguate the interpretation of ppqs. 6. conclusions subtle changes in human communication, also depending on the contextual situation, may lead to different interpretation of the communicative act. recent pragmatic studies have highlighted how the interaction between knowledge that was previously added to the common ground and different kinds of contextual evidence lead to different forms of clarification requests. for the specific case of positive knowledge negated by incoming contextual evidence, the use of high negative polar questions appears to be perceived as more appropriate by human evaluators. in this work, we have shown that this is not simply a matter of preference or naturalness: adopting the most expressive syntactic form to signal specific kinds of conflict between bias and evidence indeed has a potential impact on interaction quality. more specifically, we investigated the following research questions: 139 di maro, origlia, cutugno • main question: does linguistic feedback in the form of different polar questions influence the capability of people to identify the cause of reported common ground inconsistencies? • secondary question: do different polar question forms influence the speed with which people take decisions when searching for the cause of reported common ground inconsistencies? for the case of error signalling in the p/¬p case, we have shown that the use of hnpqs significantly improves the capability of human subjects to understand where the cause of the conflicts considered in this work is, with respect to the baseline. this is coherent with the expectations coming from the reference pragmatic study and it has relevant implications for the design of dialogue systems concerning the observability feature, thus providing a positive answer for our main research question. the use of hnpqs in bias/evidence conflicts is not just a matter of appropriateness but it influences the capability of the subjects to complete the assigned task. concerning the secondary research question, about the efficiency with which a decision is reached by the human participants, results do not show a statistically significant positive effect, in our data. this may suggest that, once the kind of problem is understood by the participant, the time needed to identify the inconsistent action does not vary significantly between hnpq and ppq. the main issue may lie, therefore, in the participants being able to actually find the inconsistent action, at all. however, this claim needs further investigation as it is solely based on the analysis of speech response times. given the time needed, for example, for response planning, it is possible that appreciable differences in reaction times may be observable on other signals, like for example eye tracking and pupil dilation. this is left for future work. the theoretical formalisation presented in section 5 adds a further perspective on possible future applications of our findings. we have shown that a formal representation of conflicts can be related to the superficial form of clarification requests, which are a characteristic element of argumentation-based dialogues. we have also shown that human subjects react positively when presented with a clarification request form coherent with type of conflict that was considered in the experiment. this is consistent with the theory presented in domaneschi et al. (2017) and di maro et al. (2021c) which also describe a number of other types of conflicts. this suggests that a larger number of cases in which argumentation-based capabilities are necessary can be described through the use of conflict patterns, which would become a foundational element of an encompassing theory for argumentation-based dialogue. moreover, to overcome the reported limitations (section 5.1), future plans foresee the implementation of a dialogue system based on these observations to confirm the observed effects in human-machine dialogues. also, the formalisation effort for argumentation-based dialogue will be extended to cover all the cases mentioned literature. references jens allwood, joakim nivre, and elisabeth ahlsén. on the semantics and pragmatics of linguistic feedback. journal of semantics, 9:1–26, 1992. jens allwood, david traum, and kristiina jokinen. cooperation, dialogue and ethics. international journal of human-computer studies, 53(6):871–914, 2000. egbert ammicht, j fosler-lussier, and alexandros potamianos. system and method for representing and resolving ambiguity in spoken dialogue systems, 2003. us patent app. 10/170,510. 140 common ground inconsistencies anne h anderson, miles bader, ellen gurman bard, elizabeth boyle, gwyneth doherty, simon garrod, stephen isard, jacqueline kowtko, jan mcallister, jim miller, et al. the hcrc map task corpus. language and speech, 34(4):351–366, 1991. nicholas asher and brian reese. intonation and discourse: biased questions. interdisciplinary studies on information structure, 8:1–38, 2007. rachel baker and valerie hazan. diapixuk: task materials for the elicitation of multiple spontaneous speech dialogs. behavior research methods, 43(3):761–770, 2011. robbert-jan beun and rogier m van eijk. conceptual discrepancies and feedback in humancomputer interaction. in proceedings of the conference on dutch directions in hci, page 13, 2004. dan bohus. error awareness and recovery in task-oriented spoken dialogue systems. ph.d. thesis, computer science department, carnegie mellon university, 2007. caroline bousquet-vernhettes, régis privat, and nadine vigouroux. error handling in spoken dialogue systems: toward corrective dialogue. in isca tutorial and research workshop on error handling in spoken dialogue systems, 2003. harry bunt and william black. abduction, belief and context in dialogue: studies in computational pragmatics, volume 1. john benjamins publishing, 2000. daniel buring and christine gunlogson. aren’t positive and negative polar questions the same? working paper, 2000. url http://hdl.handle.net/1802/1432. hendrik buschmeier and stefan kopp. communicative listener feedback in human-agent interaction: artificial speakers need to be attentive and adaptive. in proceedings of the 17th international conference on autonomous agents and multiagent systems, aamas ’18, pages 1213–1221, richland, sc, 2018. international foundation for autonomous agents and multiagent systems. url http://dl.acm.org/citation.cfm?id=3237383.3237880. mary sandra carberry. pragmatic modeling in information system interfaces (goals, dialogue, plans, ill-formedness). phd thesis, newark, de, usa, 1985. eve v. clark. common ground. in the handbook of language emergence, page 328–353. wiley, chichester, uk, 2015. doi: 10.1002/9781118346136.ch15. herbert h clark. using language. cambridge university press, 1996. herbert h. clark and susan e. brennan. grounding in communication. in lauren resnick, levine b., m. john, stephanie teasley, and d., editors, perspectives on socially shared cognition, pages 13–1991. american psychological association, 1991. amanda cercas curry, helen hastie, and verena rieser. a review of evaluation techniques for social dialogue systems. in proceedings of the 1st acm sigchi international workshop on investigating social interactions with artificial agents, pages 25–26, 2017. adrian de wynter, xun wang, alex sokolov, qilong gu, and si-qing chen. an evaluation on large language model outputs: discourse and memorization. arxiv preprint arxiv:2304.08637, 2023. 141 di maro, origlia, cutugno martina di bratto, antonio origlia, maria di maro, and sabrina mennella. linguistics-based dialogue simulations to evaluate argumentative conversational recommender systems. user modelling and user adapted interaction, special issue on conversational recommender systems: theory, models, evaluations, and trends, 2024. maria di maro. computational grounding: an overview of common ground applications in conversational agents. ijcol. italian journal of computational linguistics, 7(7-1, 2):133–156, 2021a. maria di maro. “shouldn’t i use a polar question?” proper question forms disentangling inconsistencies in dialogue systems. ph.d. thesis, università degli studi di napoli federico ii, 2021b. maria di maro, mohamed diaoulé diallo, and francesco cutugno. information-processing machines and the access-conscious recognition of common ground inconsistencies: a proposal. in psychobit, 2020. maria di maro, antonio origlia, and francesco cutugno. conflict search graph for common ground consistency checks in dialogue systems. in proceedings of the 25th workshop on the semantics and pragmatics of dialogue, 2021a. maria di maro, antonio origlia, and francesco cutugno. cutting melted butter? common ground inconsistencies management in dialogue systems using graph databases. ijcol. italian journal of computational linguistics, 7(7-1, 2):157–190, 2021b. maria di maro, antonio origlia, and francesco cutugno. polarexpress: polar question forms expressing bias-evidence conflicts in italian. international journal of linguistics, pages 14–35, 2021c. felix dietze, johannes karoff, andré calero valdez, martina ziefle, christoph greven, and ulrik schroeder. an open-source object-graph-mapping framework for neo4j and scala: renesca. in international conference on availability, reliability, and security, pages 204–218. springer, 2016. alan dix, alan john dix, janet finlay, gregory d abowd, and russell beale. human-computer interaction. pearson education, 2003. filippo domaneschi, maribel romero, and bettina braun. bias in polar questions: evidence from english and german production experiments. glossa: a journal of general linguistics, 2(1), 2017. doi: 10.5334/gjgl.27. georgios drakopoulos, andreas kanavos, christos makris, and vasileios megalooikonomou. on converting community detection algorithms for fuzzy graphs in neo4j. in proceedings of the 5th international workshop on combinations of intelligent methods and applications, cima, 2015. raquel fernández, andrea corradini, david schlangen, and manfred stede. towards reducing and managing uncertainty in spoken dialogue systems. in proceedings of the 7th international workshop on computational semantics (iwcs07), 2007. kaoru hayano. 19 question design in conversation. the handbook of conversation analysis, page 395, 2013. 142 common ground inconsistencies john heritage. the limits of questioning: negative interrogatives and hostile question content. journal of pragmatics, 34(10-11):1427–1446, 2002. julian hough and david schlangen. it’s not what you do, it’s how you do it: grounding uncertainty for a simple robot. in 2017 12th acm/ieee international conference on human-robot interaction (hri, pages 274–282. ieee, 2017. julian hough, sina zarrieß, and david schlangen. grounding imperatives to actions is not enough: a challenge for grounded nlu for robots from human-human data. in glu 2017 international workshop on grounding language understanding, pages 88–91, 08 2017. doi: 10.21437/glu.2017-18. yan huang. the oxford handbook of pragmatics. oxford university press, 2017. pablo jiménez, javier villalba diez, and joaquin ordieres-mere. hoshin kanri visualization with neo4j. empowering leaders to operationalize lean structural networks. procedia cirp, 55:284– 289, 2016. irene koshik. a conversation analytic study of yes/no questions which convey reversed polarity assertions. journal of pragmatics, 34(12):1851–1877, 2002. irene koshik. beyond rhetorical questions: assertive questions in everyday interaction, volume 16. john benjamins publishing, 2005. william h kruskal and w allen wallis. use of ranks in one-criterion variance analysis. journal of the american statistical association, 47(260):583–621, 1952. d robert ladd. a first look at the semantics and pragmatics of negative questions and tag questions. in papers from the... regional meeting. chicago ling. soc. chicago, ill, number 17, pages 164– 171, 1981. fabrizio macagno and sarah bigi. analyzing the pragmatic structure of dialogues. discourse studies, 19(2):148–168, 2017. meinard müller. dynamic time warping. information retrieval for music and motion, pages 69–84, 2007. antonio origlia, francesco cutugno, antonio rodà, piero cosi, and claudio zmarich. fantasia: a framework for advanced natural tools and applications in social, interactive approaches. multimedia tools and applications, 78:13613–13648, 2019. antonio origlia, martina di bratto, maria di maro, and sabrina mennella. developing embodied conversational agents in the unreal engine: the fantasia plugin. in proceedings of the 30th acm international conference on multimedia, pages 6950–6951, 2022. volha petukhova, harry bunt, andrei malchanau, and ramkumar aruchamy. experimenting with grounding strategies in dialogue. semdial 2015 godial, page 198, 2015. henry prakken. historical overview of formal argumentation, volume 1. college publications, 2018. 143 di maro, origlia, cutugno matthew purver. clarie: the clarification engine. in proceedings of the 8th workshop on the semantics and pragmatics of dialogue (catalog), pages 77–84. citeseer, 2004. matthew purver. clarie: handling clarification requests in a dialogue system. research on language and computation, 4(2-3):259–288, 2006. antonio roque. dialogue management in spoken dialogue systems with degrees of grounding. university of southern california, 2009. antonio roque and david traum. degrees of grounding based on evidence of understanding. in proceedings of the 9th sigdial workshop on discourse and dialogue, pages 54–63, 2008. michelina savino. the intonation of polar questions in italian: where is the rise? journal of the international phonetic association, 42(1):23–48, 2012. samuel sanford shapiro and martin b wilk. an analysis of variance test for normality (complete samples). biometrika, 52(3/4):591–611, 1965. jack sidnell and tanya stivers. the handbook of conversation analysis, volume 121. john wiley & sons, 2012. gabriel skantze. exploring human error recovery strategies: implications for spoken dialogue systems. speech communication, 45(3):325–341, 2005. jakob steensig and paul drew. questioning and affiliation/disaffiliation in interaction: special issue of discourse processes. sage publications, 2008. tanya stivers and nick j enfield. a coding scheme for question–response sequences in conversation. journal of pragmatics, 42(10):2620–2626, 2010. david r traum. computational models of grounding in collaborative systems. in psychological models of communication in collaborative systems-papers from the aaai fall symposium, pages 124–131, 1999. michaela m wagner-menghin. binomial test. wiley statsref: statistics reference online, 2014. douglas walton and erik cw krabbe. commitment in dialogue: basic concepts of interpersonal reasoning. suny press, 1995. dawn r weatherford, mitchell a meltzer, curt a carlson, and james c bartlett. never forget a face: verbalization facilitates recollection as evidenced by flexible responding to contrasting recognition memory tests. memory & cognition, 49(2):323–339, 2021. jim webber. a programmatic introduction to neo4j. in proceedings of the 3rd annual conference on systems, programming, and applications: software for humanity, pages 217–218, 2012. peter wittenburg, hennie brugman, albert russel, alex klassmann, and han sloetjes. elan: a professional framework for multimodality research. in 5th international conference on language resources and evaluation (lrec 2006), pages 1556–1559, 2006. yiqian zou. an experimental evaluation of grounding strategies for conversational agents. master thesis, department of philosophy, linguistics and theory of science, university of gothenburg, 2020. 144 dialogue & discourse 14 (2) 49-82 doi: 10.5210/dad.2023.202 ©2023 ekaterina tskhovrebova, sandrine zufferey and pascal gygax this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). exploring the sensitivity to non-connective signals of coherence relations: the case of french speaking teenagers ekaterina tskhovrebova ekaterina.tckhovrebova@unibe.ch university of bern, switzerland sandrine zufferey sandrine.zufferey@unibe.ch university of bern, switzerland pascal gygax pascal.gygax@unifr.ch university of fribourg, switzerland editor: manfred stede submitted 01/2023; accepted 09/2023; published online 10/2023 abstract coherence relations between elements of discourse can be signaled by linguistic devices such as connectives and/or non-connective signals. while the use and comprehension of connectives have been studied in different categories of speakers, less is known about the functioning of non-connective signals of coherence relations, especially in younger populations. in the current series of three experiments, we aim to examine the sensitivity of french-speaking teenagers to the non-connective signals of the list relation (adjectives of quantity such as plusieurs ‘several’ and différents ‘various’), combined with connectives varying in frequency and signaling two types of coherence relations (addition: en plus, en outre; consequence: donc, ainsi). our results reveal that, as early as in teenage years, speakers are sensitive (i.e., they produce list continuation sentences) to non-connective signals of list relation (experiments 1, 2, 3). furthermore, the inference of list relation is not significantly changed when a non-connective signal is combined with the more frequent additive connective en plus (experiment 2). however, this inference is inhibited by the less frequent additive connective en outre (experiment 3), and is almost completely hindered by the consequence connectives donc (experiment 2) and ainsi (experiment 3). overall, these results show that non-connective list signals are an important source for the inference of the list relation, even in the presence of more salient and prototypical signals of coherence such as connectives. keywords: discourse connectives, non-connective signals, coherence relations, french, teenagers. 1 introduction coherence is an important property of meaningful discourse (sanders et al., 1992). between discourse segments, coherence hinges on the ability to infer an appropriate coherence relation, such as causality or contrast. there are various linguistic elements that help speakers to infer an appropriate coherence relation. connectives, i.e., words like because or nevertheless, are one of the most studied signals of coherence relations (e.g., bloom et al., 1980; champaud & bassano, 1994; blything, davies & cain, 2015), with studies focusing on speakers from various age groups (see e.g., blything & cain, 2016; evers-vermeul & sanders, 2009; nippold et al., 1992) and linguistic competences (see e.g., crosson et al., 2008; van silfhout et al., 2015; volodina & weinert, 2020). tskhovrebova, zufferey and gygax 50 many coherence relations, however, are not marked by a connective but rather conveyed implicitly. in the penn discourse treebank (pdtb 3), about 41% of the relations are not marked by connectives (webber et al., 2019). liu (2019), examining the distribution of signals across four different text genres, namely academic articles, how-to guides, interviews, and news articles, from the georgetown university multilayer (gum) corpus (zeldes, 2017), also found that connectives signal only 16% of relations as opposed to 84% of relations marked by other signal types. similarly, in the mono-genre rst discourse treebank (carlson et al., 2002), only about 11% of coherence relations are signaled exclusively by connectives, whereas approximatively 75% of relations are marked by other, non-connective types of coherence signals (das & taboada, 2018). in fact, das and taboada (2018) identified at least seven types of non-connective coherence markers in this corpus, such as lexical, semantic, morphological, syntactic, graphical, genre, and numerical features (for other approaches to the annotation of different signal types, see knaebel & stede, 2022; liu & zeldes, 2019; zeldes & liu, 2020). for instance, such syntactic features as the present participial clause in (1) can signal the relation of manner; and the antonyms in (2) are the semantic signals of the contrastive relation (das, 2014). (1) wyse has done well, establishing a distribution business. (2) … higher bidding narrows the investor's return, while lower bidding widens it. although the corpus studies underscore the importance of alternative coherence signaling besides connectives, less is known about the inference of coherence relations from these types of signals (but see brown & fish, 1983; scholman et al., 2020). moreover, very few studies (except for crible, 2021; crible & demberg, 2021; crible & pickering, 2020; grisot & blochowiak, 2021; schwab & liu, 2020) have assessed how different types of signals interact with each other. lexical (schwab & liu, 2020) and syntactic (crible & pickering, 2020) cues, for instance, were found to reinforce inference of a particular coherence relation signaled by polyfunctional connectives, such as but or and. however, it is not clear whether the interaction between non-connective signals and connectives would be the same if the latter were monofunctional, specialized in marking one type of coherence relation. in comparison to nonconnective signals, connectives are more salient markers of coherence, as the signaling of coherence relations is their primary function, and they are often used in a prominent clauseinitial position. in contrast, non-connective signals often occupy less prominent syntactic positions and are not specialized in signaling coherence relations. therefore, an important question is whether the inference from non-connective signals, such as the lexical or semantic features from das and taboada (2018), is still generated on top of the contribution of a stronger cue of coherence relations, such as a connective. as little as we know about the functioning of non-connective signals of coherence relations in adults, even less is known about the sensitivity to these signals in younger populations. in other words, we do not know whether and how young speakers are guided by non-connective signals in their production of coherence relations. to the best of our knowledge, only au (1986) examined the sensitivity of preschool children to the implicit causality verbs and showed that already at the age of 5, speakers could perceive whether it is an agent or a patient who is causing a certain event in a sentence. however, teenage years seem not to have been studied, even though this period is found to be important for the development of an adultlevel linguistic competence (berman, 2004). moreover, even adult speakers show variation in their linguistic competence in general (kidd et al., 2018) and in the sensitivity to non-connective signals of list relations in particular (scholman et al., 2020). thus, scholman et al. (2020) demonstrated that the ability of adult speakers to infer the relation of list from the expressions of quantity like a couple, a few, multiple, and several varied according to the speakers’ degree of exposure to print (as measured by the author recognition test). in order to explore this relation further, we will extend the study of non-connective list signals on a younger population, who has even less linguistic experience than adults and is probably still acquiring a sensitivity to such signals. we therefore expect that, when both a non-connective signal (e.g., expressions of quantity) and a connective are present in a sentence, speakers of all ages, but especially young ones, should be more influenced by exploring the sensitivity to non-connective signals of coherence relations 51 connectives than by the non-connective signals, as connectives represent more salient and specialized cues of coherence. 2 the role of non-connective signals for coherence marking there are various types of non-connective signals that can mark different coherence relations. many have studied the role of lexical cues for the inference of causal relations (e.g., au, 1986; koornneef & van berkum, 2006; pyykkönen & järvikivi, 2010; rohde & horton, 2014). pyykkönen and järvikivi (2010), for instance, showed that in the sentences john feared bill because … and john frightened bill because …, the implicit causality verbs fear and frighten immediately activate verb-based reference toward either the second or the first participant of the action, respectively. kehler (1994) and lascarides & asher (1993) revealed the importance of morphological features, such as the combination of verb tenses, for signaling order of the occurring events. for example, in (3), the usage of the past simple in the first sentence and past perfect in the second one suggests that the event presented in the second clause (swimming in the lake) preceded the one shown in the first clause (illness). (3) jane fell ill. she had swum in the cold lake. there is also evidence about non-connective signals used in other coherence relations. for instance, crible and pickering (2020) found a facilitating effect of syntactic parallelism in combination with the connectives but or and for marking the relation of addition and contrast (4), as sentences with parallel structures were read faster than sentences without parallelism across a series of self-paced reading experiments. schwab and liu (2020) observed in a selfpaced reading task that the lexical cues true and sure, like in the example (5), helped readers to anticipate the upcoming concessive relation, as reflected by shorter reading times at the postcritical region. moreover, crible (2021) demonstrated in a series of four self-paced reading experiments that verbal negation, introduced in the first sentence, facilitates processing of the concessive relation, removing the difference in processing cost between the more complex concessive relation and the less complex result relation. (4) nick always eats in low-budget restaurants and/but grace always eats in fancy places (crible & pickering, 2020, p. 8). (5) james likes to run. true/sure, he has a treadmill in the living room, but he often jogs in parks (schwab & liu, 2020, p. 106). crible and demberg (2021) argued that resultative verbs, as in (6), and antonyms as in (7), respectively generate inferences of consequence and contrast relations. yet, the inference power of these non-connective signals was not as important as that of connectives signaling the same relations. as for temporal relations, grisot and blochowiak (2021) reported in a bilingual french-english corpus study that pluperfects signal backward temporal relations, simple past marks forward temporal relations, and imperfectives convey synchronous temporals. (6) males have been proven to be more skilled at sports. it allows them to win in mixed competitions (crible & demberg, 2021, p. 320). (7) the belgian government decided to create a new tax on solar panels. the french government decided to remove the existing tax (crible & demberg, 2021, p. 321). less is known, however, about the inference generation of an additive relation. still, scholman et al. (2020) examined expressions of quantity such as a couple, a few, multiple, and several, and found that they activate the inference of list relation – a particular type of a more generic additive relation – in adult speakers. in addition, the corpus study by péry-woodley et al. (2017) showed that the relation of list, or enumeration, can be expressed by a variety of enumerative structures of different length and graphical aspect, such as multiparagraph structures and bullet lists. interestingly, it also showed that these structures often have a similar organization. they predominantly start with a trigger, which often includes a lexical cue. the trigger element is followed by a series of items, which in turn can be followed by a closure element. these findings are particularly insightful, because additive relations are one of the relations that are the least signaled by connectives and are conveyed by the greatest variety of non-connective signals (das & taboada, 2018). it even seems that speakers' comprehension of additive tskhovrebova, zufferey and gygax 52 relations is hindered when an additive connective is present between two sentences (kleijn et al., 2019), as in (8). this effect is different from other types of relations such as cause or contrast that elicit better comprehension scores when marked by connectives. a possible reason of this hindering effect, as suggested by kleijn et al. (2019), is that additive connectives draw excessive attention to the coherence relation and elicit an overinterpretation of the intended relation in contrast to a simple juxtaposition. other signals become therefore interesting to investigate, especially to better understand how additive relations work. (8) not everyone can register in the donor register: you must be at least twelve years old and in addition you must be a registered citizen of a dutch municipality (kleijn, 2018, p. 216). another important contribution would be to examine the interaction between nonconnective signals of coherence relations and connectives. only few studies have attempted to explore this interaction, reporting findings for a limited number of coherence relations, namely contrast (crible & demberg, 2021; crible & pickering, 2020), consequence (crible & demberg, 2021), and concession (schwab & liu, 2020). however, more work is needed to describe how this interaction works for other types of coherence relations. in this respect, it would be useful to provide evidence on the interaction between non-connective signals and connectives signaling a less studied additive relation. for instance, assessing the interaction between non-connective list signals, additive and consequence connectives for readers’ propensity to generate inferences of list relations would enable us to evaluate whether these relations are still inferred from non-connective signals. importantly, one could document whether they are inferred even in the presence of stronger coherence signals such as connectives marking the same or a different type of relation. in all, it would constitute an interesting extension to the study by scholman et al. (2020). moreover, examining speakers' sensitivity to non-connective list signals and their interaction with connectives in teenagers would allow us to fill a gap in literature on non-connective signaling in teenage years and to generalize the results of scholman et al. (2020) to other age groups. it is also possible that even connectives conveying the same type of coherence relation but varying in frequency may have a different impact on the generation of inferences. for instance, even adults have difficulties using (tskhovrebova, zufferey & gygax, 2022; zufferey & gygax, 2020b) and identifying correct and incorrect uses (zufferey & gygax, 2020a) of the infrequent connectives aussi ‘therefore’ and en outre ‘in addition’. in consequence, since speakers appear to be less confident about the usage of less frequent connectives, these connectives may also generate weaker inferences of a certain coherence relation, even combined with non-connective signals. an overview of research on the competence with connectives in teenage years will allow us to make predictions on the sensitivity to nonconnective signals in combination with connectives (of different frequency) in this age group. 3 teenagers’ competence with connectives teenage years are an important period of linguistic development between the emergence and mastery of language (berman, 2004). language development in teenagers continues on lexical, semantic, syntactic, and pragmatic levels (see, e.g., nippold, 2008). the mastery of connectives, in turn, is at the interface between lexical, syntactic and pragmatic skills, which are actively developing during this period, and therefore occupy a central role in the development of a full-fledged linguistic competence. previous research has shown that, on average, teenagers' competence with any type of connectives is inferior to that of adult speakers (nippold et al., 1992; tskhovrebova, zufferey & gygax, 2022; zufferey & gygax, 2020b). nippold et al. (1992) assessed the competence of english-speaking teenagers aged 12 to 15 and young adults aged 19 to 23 with connectives mostly used in written language, such as furthermore and nevertheless, in a connective insertion task and a sentence continuation task. the authors found that young adults performed significantly better than teenagers in both tasks. a similar result was obtained in two studies examining the usage of four french connectives bound to the written modality but varying in frequency (tskhovrebova, zufferey & gygax, 2022; zufferey & gygax, 2020b). both studies exploring the sensitivity to non-connective signals of coherence relations 53 demonstrated that even high-school students aged 16 to 18 did not reach the performance of adults in the connective cloze task across all connectives, suggesting that proficiency with connectives continues to develop through the late teenage years. moreover, research on competence with connectives in l2, i.e., for speakers with a lower level of linguistic proficiency and can be in that respect compared to teenagers, shows that language learners also have difficulties detecting incorrect uses, even for very frequent connectives. the study of wetzel et al. (2022) reported, for instance, that german-speaking learners of french did not react to the erroneous uses of the frequent french connective alors ‘so’ in a self-paced reading task. considering the findings on the mastery of connectives by less experienced speakers, we suggest that teenagers may also be less proficient with non-connective signals of coherence relations, and thus less sensitive to them when they are used either alone or combined with connectives. this may be true, but not for all teenagers, as some individual factors may be decisive in determining whether they have a lower sensitivity to non-connective signals or not. for example, in adult populations, exposure to print, as measured by an author recognition test (stanovich & west, 1989), has been shown to modulate reader’s mastery of connectives (zufferey & gygax, 2020a) and sensitivity to non-connective signals (scholman et al., 2020). degree of general exposure to print could therefore be an important factor, modulating the inference generated by non-connective signals in combination with connectives, also in teenage populations. 4 our set of experiments in the current set of experiments, we aim to address the gaps identified in previous research on non-connective signals of coherence relations. our goal is to extend the study by scholman et al. (2020) on a younger cohort of teenagers and to examine their sensitivity to non-connective signals of the list relation (experiment 1) in combination with connectives varying in frequency and signaling two types of coherence relations (experiments 2 & 3). more specifically, we assess french-speaking teenagers' sensitivity to the adjectives of quantity plusieurs ‘several’ and différents ‘various’, and how this sensitivity is modulated by the presence of connectives signaling the relations of addition and consequence. this way, we aim to examine whether a list inference, generated by a non-connective signal, is strong enough to trigger list continuations on top of the inference generated by connectives. the additive connectives were chosen because addition does not compete with the logic of the list relation. in fact, additive relations represent a generic type of relations that include several subtypes, among them the relation of list. in contrast, the consequence connectives were selected because consequence represents a separate class of coherence relations, which is competing with the logic of the list relation (see table 1 for a summary of all the signals used in the set of experiments). we use the following definitions for the three coherence relations included in our experiments: 1. sentences are linked with a list relation when the second sentence enumerates one or several events related to the content of the first sentence; 2. sentences are linked with an additive relation when the second sentence expands and elaborates on the content of the first sentence, except for instances of enumeration that are included in the category of list relations; 3. sentences are linked with a consequence relation when the second sentence describes an event caused by an activity presented in the first sentence. non-connective signals connectives additive consequence experiment 1 plusieurs ‘several’ différents ‘various’ – – experiment 2 en plus donc experiment 3 en outre ainsi table 1. all the connectives and non-connective signals used in our set of experiments. based on the results of scholman et al. (2020), we predict that participants will produce more list continuations after reading items containing adjectives of quantity in all experiments. tskhovrebova, zufferey and gygax 54 however, it is possible that teenagers will be less sensitive than adults to such signals, due to a lower level of linguistic competence. we also expect that after reading sentences including both a list signal and an additive connective (experiments 2 and 3), the proportion of list continuations should not decrease, but rather increase or remain unchanged because an additive connective is not in contradiction with the relation of list. moreover, we predict that the combination of a list signal and a consequence connective will decrease the percentage of list continuations, as this type of connective expresses a non-compatible relation of consequence, and this will override the inference generated by a less salient and more polysemous (in the sense that it is not specialized only in coherence marking) non-connective signal of list (experiments 2 and 3). finally, we expect that the general effect from the less frequent connectives (experiment 3) will be lower than from the more frequent connectives (experiment 2). to identify whether the sensitivity to these signals in young speakers also varies depending on individual differences in linguistic competence, we assessed the participants' degree of exposure to print, as measured by adapted french versions of the author recognition test (tskhovrebova, zufferey & tribushinina, 2022; zufferey & gygax, 2020a). 5 experiment 1 5.1 participants fifty-three teenagers (mage = 14.18, sd = 1.66, range: 12–18) and twenty adults (mage = 31.36, sd = 11.35, range: 21–63) took part in the experiment. all of them were native french speakers. the experiment among teenagers was carried out in secondary and high schools of the french-speaking part of switzerland, and was performed online via a weblink. adults were recruited online on the prolific© platform (prolific, oxford, uk). 5.2 materials and procedure 5.2.1 story-continuation task in this experiment, participants had to write a continuation to a series of pairs of sentences. in each pair, the first sentence introduced an agent and the context it was in, and either included a list signal (the adjectives of quantity plusieurs ‘several’ or différent ‘various’) or not. the second sentence started with a pronoun coreferential with the agent of the first sentence, and developed the situation. example (9) illustrates an experimental item in the list and non-list conditions. (9) list condition: la comédienne a planifié plusieurs rendez-vous pour la journée. elle a prévu d'aller voir son agent. ‘the actress scheduled several appointments for the day. she planned to meet her agent.’ non-list condition: la comédienne se préparait à la maison. elle a prévu d'aller voir son agent. ‘the actress was getting ready at home. she planned to meet her agent.’ the second sentence was identical across both list signal conditions. we did not simply remove the adjectives of quantity from the first sentence of the condition without list signal, but also changed the verb for several reasons. first, we wanted to prevent list and non-list items from being perceived as repetitions of the same sentence after reading multiple task items in a row, which could be the case if we just omitted the list signal. second, we wanted to make sure that participants would perceive list and non-list items as different sentences and treat them as such across the whole task, but without making it obvious that the presence or absence of these nonconnective signals were the focal point of the task. third, we wanted to ensure that list and nonlist items were similar in terms of the expectations that they would create. our objective was to exploring the sensitivity to non-connective signals of coherence relations 55 build neutral sentences in the non-list condition, without any obvious non-connective cues of coherence (such as implicit causality verbs, for instance). the choice of the adjectives plusieurs and différents was based on several criteria. first, they have the same function in french as previously examined english expressions of quantity from the study of scholman et al. (2020). second, both plusieurs and différents belong to the same part of speech (indefinite adjectives) and are used in the same syntactical position before nouns. third, they refer to an indefinite number of things or events in contrast to other indefinite adjectives of quantity like quelques, which normally is used to describe a small number of things, or nombreux and multiple, which refer to large numbers of things of events. finally, both adjectives are frequent in french with 447.03 (for plusieurs) and 144.79 (for différent) occurrences per million words. in total, there were 20 items, with two conditions per item (with and without list signal). each type of list signal was inserted equally frequently across all list conditions. participants were asked to provide a continuation with at least one sentence that had to be complete, grammatically correct, and contain at least three words. all participants saw all the items, both with and without list signal. it was important that participants saw both types of items to examine whether it was the presence of a list signal that affected their inference generation and that the latter was not due to individual bias towards a specific relation. coding procedure continuation sentences were annotated for the analysis as list, additive, consequence or other, depending on their relation to the cue sentence. we defined as list continuations those sentences that contained an enumeration of one or several events related to the content of the first cue sentence. examples (10) – (12) illustrate list continuations that participants wrote in the list condition for item (9). (10) et elle a prévu de passer d'autres castings. ‘and she planned to do more castings.’ (11) puis elle prévu d'aller faire un coucou à ses grand-parent1. ‘then she planned to go and say hello to her grandparents.’ (12) elle doit aussi aller se faire une teinture chez le coiffeur. ‘she also has to get her hair dyed.’ there were several completion sentences that expressed not only list, but also temporal (11) or contrast (13) relations at the same time. in such cases, continuations were labelled as list, as the focal point of this set of experiments was to identify sentences conveying the idea of enumeration in relation to the prompt. (13) task sentences: la journaliste a fait différents commentaires sur le film. elle a apprécié le jeu de l'actrice principale. ‘the journalist made various comments about the film. she appreciated the acting of the lead actress.’ continuation: elle a moins aimé la qualité des dialogues. ‘she liked the quality of the dialogues less.’ we coded as additive those continuations that provided new information or more details about the first or the second task sentences, including exemplification and sub-events, but excluding the instances of enumeration, which were included in the category of list relations. in the continuations (14) and (15), for instance, participants do not list any other activities planned by the actress for the day as in (10) – (12), but rather add a new fact about an actress (14) or her agent (15). for this reason, these continuations were labelled as additive rather than list. we did not make a further distinction about additive and elaboration relations, as it was not relevant for our set of experiments. (14) task sentence (list condition): 1 we kept the faulty original spelling of participants in french examples. tskhovrebova, zufferey and gygax 56 la comédienne a planifié plusieurs rendez-vous pour la journée. elle a prévu d'aller voir son agent. ‘the actress scheduled several appointments for the day. she planned to meet her agent.’ continuation: elle a reçu beaucoup d'argent. ‘she received a lot of money.’ (15) continuation: celui-ci a annulé au dernier moment. ‘the latter cancelled at the last moment.’ when a continuation phrase described an event caused by an activity presented in the task item, it was tagged as a consequence relation. example (16) illustrates a consequence continuation that was written by one of the participants after a task item in the list condition. the fact that the girl’s mother gave her an ice cream is considered here to be a consequence of her good performance at school. (16) task sentence (list condition): la fille a reçu plusieurs bonnes notes à l'école. elle a réussi l'examen d'histoire. ‘the girl received several good marks at school. she passed the history exam.’ continuation: donc sa maman l'a récompensé avec une glace. ‘so her mother rewarded her with an ice cream.’ all the remaining relations were labelled as other. this category included several types of discourse relations, such as temporality, contrast, cause, and goal, which were not further distinguished, as it was not essential for the goals of the present investigation. if a participant provided several continuation sentences, discourse relations between the provided continuations were not labelled, since it was outside of the scope of the present set of experiments. our focus was on discourse relations between the prompt and the completion sentence. out of 8114 continuations2, 10% were annotated together by one of the authors and an independent experienced coder. it is important to mention that, when we deal with the annotation of discourse relations, multiple, non-self-excluding interpretations can be possible and not always all of these interpretations are noticed and taken into account by coders. however, the agreement rate between the two coders on this continuation sample was 95% (κ=.82; gwet's ac1=.92), which granted the remaining of the continuations to be divided in half and annotated independently. note that all instances where one coder had doubts were cross-checked by the other coder. 5.2.2 author recognition tests to assess teenagers’ degree of exposure to print, we used an adapted version of the author recognition test (art) (tskhovrebova, zufferey & tribushinina, 2022), since the art is not only sensitive to cultural differences (e.g., stainthorp, 1997) but also to age (e.g., cunningham & stanovich, 1990). this version of the art (art-f-cl) was based on the names of frenchspeaking authors who are considered to be classics according to three swiss and french national libraries and bookshop chains. the list included 40 author names and 40 names of unknown people, which were randomly mixed. the participants had to select only those names that they knew to be authors. the instruction mentioned that some of the names were not authors, and that one point would be removed if the participants checked a wrong name. for each correct answer, the participants were given 1 point, and for each wrong one -1. the maximum possible score therefore was 40 and the minimum -40, as we computed the general score summing up the points for correct and incorrect answers. for the adult control group, we used a different version of the art, which was developed by zufferey and gygax (2020a, https://osf.io/yxj8q/) and was based on the names of best-selling and prize-winning authors (art-f). the art-f replicated the design of the 2 the details about annotation are reported for the data from all three experiments. exploring the sensitivity to non-connective signals of coherence relations 57 original english art (stanovich & west, 1989). the number of items and the calculation of the final score was the same as for the teenage version of the task, described before. the reliability of the tests was quite high, as indicated by their cronbach's alphas which are close to or greater than .90 (art-f-cl: .88 [.85–.91]; art-f: .92 [.86–.94]). the participants fulfilled the tasks always in the same order, starting with the storycontinuation task and finishing with the art. once the participants gave an answer and proceeded to the next question, they could not go back and correct their initial response. 5.3 analysis continuation sentences were analyzed by fitting generalized mixed-effects logistic regression models on the binary variable (list versus non-list relation), using the r software (rstudio team, 2015). we tested models with the glmer function of the lme4 package (bates et al., 2015) and made model comparison with the anova() function, using a forward-testing approach. we added main and interaction fixed effects one at a time, and each model with an added factor was compared to a previous model that did not have the included factor. p-values of the final model were obtained with the summary() function from the lmertest package (kuznetsova et al., 2014). the statistical significance level was set at <5% and is indicated by bold marking in the corresponding tables. in total, we created three models: the first one only for teenage participants, the second one only for adults, and the third one for both age groups. in addition, for each separate analysis, we performed a pairwise comparison between list signal (absent versus present) and connectives used in the task with the lsmeans() function of the emmeans package in r (lenth, 2020). this analysis at first was performed separately for teenagers and adults, and then together for all participants. age groups were first analyzed separately, given that our primary aim was to shed light on teenage sensitivity to non-connective list signals – and given that art was different across age groups. we also present general analyses considering all groups together, yet without including art. in order to facilitate reading, we report all the details of the model selections in the online appendix3. moreover, separate models for teenagers and adults are also included in the online appendix, since the degree of exposure to print, as measured by artf-cl and art-f, did not predict the variation in the sensitivity to list signals. 5.4 results and discussion our final model included list signal (absent versus present) and group (teenagers versus adults) as fixed factors (both main and interaction effects), and item and participant as random intercepts (see table 2 and figure 1). the results demonstrate that both teenagers and adults were sensitive to list signals, as revealed by an estimated increase of 1.95  se 0.33. however, there was a significant interaction between the factors group and list signal, demonstrating that teenagers were on average less sensitive to list signals than adults. finally, in the condition without list signal, the production of list continuations did not vary between the two groups of participants. the two separate analyses within each age group confirmed the effects found in the general analysis and did not reveal any significant inter-individual variation, related to the degree of exposure to print and age. this result replicates the finding of scholman et al. (2020) on the sensitivity to list signals, applied both to adult and young speakers of french. in the next experiment, we aim to examine further the effect from the adjectives of quantity. more precisely, we assess whether participants are still sensitive to these non-connective signals, even if the task items include both adjectives of quantity and different types of connectives, which are more salient and prototypical signals of coherence relations. 3 the url of the online appendix is provided in the data availability statement. tskhovrebova, zufferey and gygax 58 figure 1. proportions of different types of relations in the continuation sentences in experiment 1 (see table s1 in online appendix for the exact values). estimate se z p all participants (intercept) -1.92 0.36 -5.37 <.001 list signal 1.95 0.33 5.85 <.001 teenagers -0.22 0.36 -0.63 0.529 list signal*teenagers -0.75 0.23 -3.29 0.001 teenagers (intercept) -2.18 0.29 -7.51 <.001 list signal 1.20 0.31 3.92 <.001 adults (intercept) -1.96 0.34 -5.76 <.001 list signal 2.00 0.43 4.64 <.001 table 2. model’s estimates for the best fitting models in experiment 1. note. in the model for all participants, conditional r2δ=.36, marginal r2δ=.08; for teenagers, conditional r2δ=.35, marginal r2δ=.05; for adults, conditional r2δ=.40, marginal r2δ=.13. 5.5 additional analysis of the distribution of connectives in list continuations we noticed that in this experiment, where connectives were not included in the prompt passage, participants added their own connectives in 80% of list continuations. when we fitted the generalized mixed-effects logistic regression model on the binary variable (absence versus presence of the connective in the list continuation), adding the factors of list signal (absence versus presence in the task item) and group (adults versus teenagers) did not improve the model’s fit (list signal: χ2(1) = 0.22, p < .638; group: χ2(1) = 0.22, p < .638). in other words, the teenagers adults present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s relation types in continuations other consequence addition list **** **** **** exploring the sensitivity to non-connective signals of coherence relations 59 insertion of the connective in list continuations was not predicted by the presence or the absence of the adjectives of quantity in the task sentence for both groups of participants. however, it seems that the position in which teenagers and adults used connectives in their productions was different. teenagers inserted connectives in sentence-initial position in 70% of the cases against only 31% for adults. in contrast, sentence-medial or sentence-final position takes up 14% of continuations produced by teenagers against 50% for adults. in the remaining 16% (for teenagers) and 19% (for adults) of continuations, participants did not use any connective. among teenagers, the most popular connective was sentence-initial et ‘and’ (48%), followed by sentence-medial aussi ‘also’ (11%), and sentence-initial mais ‘but’ (6%), and puis ‘then’ (5%). adults used most often sentence-medial connectives aussi ‘also’ (24%) and également ‘also’ (18%), sentence-initial et ‘and’ (10%), sentence-medial ensuite ‘then’ (7%), and sentence-initial puis ‘then’ (5%). examples (17) and (18) illustrate the use of some of these connectives. (17) task item: l'acrobate a signé plusieurs contrats. il va participer au festival du cirque à grenoble. ‘the acrobat has signed several contracts. he will be taking part in the circus festival in grenoble.’ continuation provided by a teenager: et il va y gagner. ‘and he's going to win.’ (18) task item: la fille a reçu plusieurs bonnes notes à l'école. elle a réussi l'examen d'histoire. ‘the girl got several good marks at school. she passed her history exam.’ continuation provided by an adult: elle a aussi réussi l'examen d'anglais. ‘she also passed her english exam.’ 6 experiment 2 6.1 participants fifty-four french native speaking teenagers (mage = 14.44, sd = 1.62, range: 12–17) and twenty-two adults (mage = 26.10, sd = 7.17, range: 18–43) participated in the second experiment. the recruitment modalities of both groups of participants were the same as in experiment 1. 6.2 materials and procedure the art tests were the same as in experiment 1, while the story-continuation task was slightly modified. participants were asked to fulfil an almost identical story-continuation task to the one in the first experiment, with the only difference that the second sentence was this time followed by a connective. the selected connectives en plus and donc respectively encode a relation of addition and consequence and are frequently used in french (respectively, 279.30 and 3'318.41 occurrences per million words 4 ). adding connectives allowed us to examine whether participants' sensitivity to list signals was modulated by the presence of a connective. moreover, by including different types of connectives, we also aimed to study their effect on the generation of inference for the upcoming coherence relation. the additive connective is not in contradiction with the logic of enumeration conveyed by the lexical signal, as this connective encodes a more generic additive relation, and can also introduce a list relation (more specific). the connective en plus was a particularly suitable candidate for this experiment, as it is 4 the connectives' mean frequency was calculated by averaging their frequencies in oral and written language. the frequency in oral speech was calculated based on the oral sub-corpus of orféo (benzitoun et al., 2016). the frequency in writing was obtained based on three different corpora, namely le monde (monde, 1987–2012), the french part of europarl (koehn, 2005), and frantext (atilf, 1998-2022), respectively representing journalistic, argumentative and literary genres. tskhovrebova, zufferey and gygax 60 frequent, monofunctional, and specialized in signaling additive coherence relations (roze et al., 2012). in contrast, we expected that the connective of consequence should decrease the production of list continuations, as this connective cannot be used to introduce a list relation. examples (19) and (20) illustrate the items used in experiment 2. (19) list condition: la comédienne a planifié plusieurs rendez-vous pour la journée. elle a prévu d'aller voir son agent. en plus, … ‘the actress scheduled several appointments for the day. she planned to meet her agent. in addition, …’ non-list condition: la comédienne se préparait à la maison. elle a prévu d'aller voir son agent. en plus, … ‘the actress was getting ready at home. she planned to meet her agent. in addition, …’ (20) list condition: le médecin avait plusieurs lieux de travail. il avait un cabinet à l'hôpital central. donc, … ‘the doctor had several places of work. he had an office at the central hospital. so, ...’ non-list condition: le médecin était spécialisé dans les traitements contre le cancer. il avait un cabinet à l'hôpital central. donc, … ‘the doctor specialized in cancer treatment. he had an office at the central hospital. so, …’ 6.3 analyses we started by making the same statistical analysis as in experiment 1. however, in order to compare the effects from list signals and connectives between the task without connectives (experiment 1) and the one with frequent connectives (experiment 2), we made an additional comparative analysis separately for each connective. 6.4 results 6.4.1 sensitivity to connectives and list signals in experiment 2 our final model included connective (en plus versus donc) as a fixed factor and item and participant as random intercepts (see table 3). this result shows that, in contrast to the connective donc, the additive connective en plus predicted a greater number of list continuations, independently of the presence of the list signal and the age group. figure 2. proportions of different types of relations in the continuation sentences in experiment 2. teenagers adults en plus donc present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s en plus donc present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s *** * relation types in continuations other consequence addition list exploring the sensitivity to non-connective signals of coherence relations 61 estimate se z p all participants (intercept) -6.49 0.52 -12.42 <.001 en plus 6.09 0.54 11.27 <.001 teenagers (intercept) -6.75 0.64 -10.50 <.001 en plus 6.34 0.64 9.88 <.001 adults (intercept) -6.52 0.91 -7.19 <.001 en plus 6.07 0.93 6.52 0.026 table 3. model’s estimates for the best fitting model in experiment 2. note. in the model for all participants, conditional r2δ=.71, marginal r2δ=.59; for teenagers, conditional r2δ=.74, marginal r2δ=.63; for adults, conditional r2δ=.75, marginal r2δ=.56. 6.4.2 additional comparative analysis between experiment 1 and experiment 2 for the connective en plus the final model for the analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after the items with the additive connective en plus (from experiment 2), included list signal (absent versus present), connective (no connective versus en plus), and group (adults versus teenagers) as fixed factors (both main and interaction effects), and item and participant as random intercepts (see table 4 and figure 3). the results from this analysis demonstrate that there was a main effect of list signal and of the connective en plus for the production of list continuations. however, when a list signal was present in the cue sentence, adults were on average more sensitive to it than teenagers. the post-hoc pairwise comparison revealed that there were significantly more list continuations after the sentences with list signals than without list signals, both when the connective en plus was present (log odds ratio=0.57, se=0.28, p=.045) and absent (log odds ratio=1.51, se=0.26, p <.0001). as a result, there was no significant change in the production of lists after the adjectives of quantity between the sentences followed by en plus and the sentences not followed by a connective (log odds ratio=0.13, se=0.25, p <.598). however, when the adjectives of quantity were absent in the task sentences, there was a significant increase in the number of list relations in participants’ responses after the sentences including the connective en plus compared to sentences without this connective (log odds ratio=1.07, se=0.25, p <.0001). the separate models for teenagers and adults had similar effects as the general model for all participants (see table 4). the only difference was that teenagers produced significantly more list continuations after the sentences with list signals than after the sentences without list signals when the connective en plus was not present in the task (log odds ratio=1.14, se=0.26, p <.0001). in contrast, adults wrote significantly more list continuations after the sentences with list signals than after the sentences without list signals both when the connective en plus was present (log odds ratio=0.85, se=0.43, p=0.046) and absent (log odds ratio=1.91, se=0.38, p <.0001). in other words, it seems that teenagers were sensitive to the adjectives of quantity only in the task items that were not followed by a connective, while adults were sensitive to the nonconnective signals in both conditions, independently of the connective en plus. tskhovrebova, zufferey and gygax 62 figure 3. proportions of different types of relations in the continuation after the task items without a connective (from experiment 1) and after those with the connective en plus (from experiment 2) across teenagers and adults. no en plus present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s relation types in continuations other consequence addition list adults no en plus present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s **** **** teenagers **** * ** exploring the sensitivity to non-connective signals of coherence relations 63 estimate se z p all participants (intercept) -1.85 0.34 -5.51 <.001 list signal 1.87 0.30 6.27 <.001 en plus 0.87 0.42 2.08 0.038 teenagers -0.22 0.35 -0.63 0.530 list signal*en plus -1.10 0.30 -3.64 <.001 list signal*teenagers -0.73 0.22 -3.25 0.001 en plus*teenagers 0.40 0.49 0.82 0.413 list signal*en plus*teenagers 0.33 0.35 0.94 0.349 teenagers (intercept) -2.09 0.26 -8.07 <.001 list signal 1.14 0.26 4.40 <.001 en plus 1.35 0.30 4.54 <.001 list signal*en plus -0.82 0.22 -3.70 <.001 adults (intercept) -1.89 0.31 -6.03 <.001 list signal 1.91 0.38 5.02 <.001 en plus 0.81 0.33 2.45 0.015 list signal*en plus -1.06 0.37 -2.83 0.005 table 4. model’s estimates for the best fitting model in the additional analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after those with the connective en plus (from experiment 2). note. in the model for all participants, conditional r2δ=.34, marginal r2δ=.07; for teenagers, conditional r2δ=.35, marginal r2δ=.06; for adults, conditional r2δ=.37, marginal r2δ=.10. 6.4.3 additional comparative analysis between experiment 1 and experiment 2 for the connective donc the final model for the analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after the items with the consequence connective donc (from experiment 2), included list signal (absent versus present), connective (no connective versus donc), and group (adults versus teenagers) as fixed factors (both main and interaction effects), and item and participant as random intercepts (see table 5 and figure 4). this analysis reveals that the presence of the connective donc in experiment 2 significantly decreased the proportion of list continuations in comparison to the task sentences without this connective from experiment 1. in other words, both groups of participants were responsive to list signals only after the sentences without the connective donc, while the presence of the consequence connective almost completely prevented participants from writing list relations in their productions. the analyses within each age group confirmed the overall effects obtained in the general analysis. tskhovrebova, zufferey and gygax 64 figure 4. proportions of different types of relations in the continuation after the task items without a connective (from experiment 1) and after those with the connective donc (from experiment 2) across teenagers and adults. no donc present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s **** **** **** teenagers no donc present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o r ti o n o f c o n ti n u a ti o n s relation types in continuations other consequence addition list **** **** adults exploring the sensitivity to non-connective signals of coherence relations 65 estimate se z p all participants (intercept) -1.92 0.36 -5.31 <.001 list signal 1.95 0.34 5.80 <.001 donc -17.05 7.89 -2.16 0.031 teenagers -0.23 0.36 -0.63 0.528 list signal* donc 11.84 7.88 1.50 0.133 list signal*teenagers -0.75 0.23 -3.24 0.001 donc*teenagers 12.33 7.88 1.57 0.118 list signal* donc *teenagers -12.63 7.89 -1.60 0.109 teenagers (intercept) -2.18 0.29 -7.45 <.001 list signal 1.20 0.30 3.94 <.001 donc -4.88 1.06 -4.62 <.001 list signal* donc -0.70 1.25 -0.56 0.577 adults (intercept) -2.04 0.36 -5.59 <.001 list signal 2.06 0.40 5.21 <.001 donc -16.90 86.55 -0.20 0.845 list signal* donc 11.90 86.54 0.14 0.891 table 5. model’s estimates for the best fitting model in the additional analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after those with the connective donc (from experiment 2). note. in the model for all participants, conditional r2δ=.76, marginal r2δ=.68; for teenagers, conditional r2δ=.56, marginal r2δ=.40; for adults, conditional r2δ=.90, marginal r2δ=.86. 6.4.4 discussion the results of this experiment revealed that the presence of the connectives en plus ‘in addition’ and donc ‘so’ affected the sensitivity to the adjectives of quantity of french speakers. the consequence connective donc completely overrode the inference from the non-connective signals of the list relation in both groups of participants. as for the additive connective en plus, the effects were not the same for the two age groups. teenagers were sensitive to the adjectives of quantity only in the items that were not followed by the connective en plus, while adults remained sensitive to the non-connective signals, independently of the additive connective. however, this finding should be interpreted with caution, as it is based on the comparison between two experiments. in the next experiment, we examine whether the frequency of the connectives following the task items may be an additional factor affecting the sensitivity to the non-connective signals. more precisely, we assess whether the presence of the less frequent additive and consequence connectives en outre ‘in addition’ and ainsi ‘therefore’, respectively, would produce the same effects on the generation of list inferences as the equivalent frequent connectives. it was found tskhovrebova, zufferey and gygax 66 in previous studies, for instance, that certain infrequent connectives, such as en outre ‘in addition’ and aussi ‘therefore’, are particularly challenging both for teenagers and adults (tskhovrebova, zufferey & gygax, 2022; zufferey & gygax, 2020 a,b). this difficulty may stem from the fact that infrequent connectives are mostly used in written modality, and extensive exposure to the written language happens later than that to the oral language, coming with schooling process (nippold, 2004; 2008). it is only starting from secondary school that teenagers become autonomous readers and start to be exposed to written texts of various genres (nippold, 2004; 2008). as a result, connectives that appear mostly in writing, and thus have on average lower frequency, may be mastered less well than those that are often used in oral language. hence, in experiment 3, we included the connectives en outre and ainsi to assess their effect on the generation of list inferences, as these connectives are mostly used in writing and have a lower frequency. the additive connective en outre can be considered as equivalent to en plus, as it signals the same coherence relation, but has a much lower average frequency (46.52 versus 279.30 occurrences per million words, respectively). the consequence connective ainsi can be considered as equivalent to donc, but it is much less frequent (178.61 versus 3'318.41 occurrences per million words, respectively). we include the connective ainsi instead of the previously tested aussi, as the latter is polyfunctional and can convey both relation of addition and that of consequence (roze et al., 2012). including two monofunctional connectives (en outre and ainsi) allowed us to disentangle two coherence relations and avoid possible confusions. 7 experiment 3 7.1 participants in the third experiment, we recruited 50 french native speaking teenagers (mage = 14.34, sd = 1.94, range: 12–19) and 21 adults (mage = 28.64, sd = 10.43, range: 20–57). the recruitment process of both groups of participants were the same as in experiment 1. 7.2 materials and procedure the art tests were again the same as in experiment 1, while the story-continuation task slightly differed. experiment 3 was almost identical to experiment 2, and differed only in the choice of connectives. instead of more frequent connectives, the cue passage included one of the two less frequent connectives, namely en outre ‘in addition’ and ainsi ‘therefore’. 7.3 analyses statistical analyses were the same as in experiment 2. 7.4 results 7.4.1 sensitivity to connectives and list signals in experiment 3 the final model for all participants included list signal, connective (en outre versus ainsi), and group as fixed factors (main and interaction effects), item and participant as random intercepts, and connective as random slope by participant (see table 6 and figure 5). this result shows that, similar to the experiment 2, the additive connective en outre predicted a greater number of list continuations than the consequence connective ainsi. however, in contrast to the experiment 2, teenagers on average wrote fewer list continuations after en outre than adults. the separate analyses for teenagers and adults confirmed the trends from the general analysis and did not reveal variation, predicted by the arts. exploring the sensitivity to non-connective signals of coherence relations 67 figure 5. proportions of different types of relations in the continuation sentences in experiment 3. estimate se z p all participants (intercept) -7.23 1.33 -5.43 <.001 en outre 6.31 1.35 4.66 <.001 list signal 1.45 0.91 1.59 0.112 teenagers 2.40 1.24 1.93 0.054 en outre*list signal -0.44 0.97 -0.45 0.655 en outre*teenagers -3.46 1.27 -2.73 0.006 list signal*teenagers -0.73 0.92 -0.80 0.427 en outre*list signal*teenagers 0.39 0.97 0.40 0.691 teenagers (intercept) -4.45 0.67 -6.61 <.001 en outre 2.50 0.71 3.51 <.001 list signal 0.70 0.38 1.85 0.064 en outre*list signal -0.05 0.49 -0.10 0.923 adults (intercept) -10.43 2.46 -4.24 <.001 en outre 9.37 2.50 3.75 <.001 list signal 1.84 1.15 1.60 0.109 en outre*list signal -0.67 1.27 -0.53 0.597 table 6. model’s estimates for the best fitting model in experiment 3. note. in the model for all participants, conditional r2δ=.66, marginal r2δ=.29; for teenagers, conditional r2δ=.50, marginal r2δ=.13; for adults, conditional r2δ=.35, marginal r2δ=.17. relation types in continuations other consequence addition list teenagers adults en outre ainsi present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s en outre ainsi present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s *** *** ** tskhovrebova, zufferey and gygax 68 7.4.2 additional comparative analysis between experiment 1 and experiment 3 for the connective en outre the final model for the analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after the items with the additive connective en outre (from experiment 3), included list signal (absent versus present), connective (no connective versus en outre), and group (adults versus teenagers) as fixed factors (both main and interaction effects), and item and participant as random intercepts (see table 7 and figure 6). this analysis showed that both groups of participants overall produced more list continuations after the items including the adjectives of quantity than after the items without them. moreover, teenagers were on average less responsive to the presence of list signals than adults across both experiments. the separate analysis within the group of teenagers showed that teenagers produced significantly more list continuations after the sentences with list signals than after the sentences without list signals when the connective en outre was absent in the task (log odds ratio=1.16, se=0.27, p <.0001). in addition, the presence of the non-connective signals and the additive connective en outre significantly decreased the production of lists in comparison to the sentences that included only the non-connective signals (log odds ratio=-0.69, se=0.29, p=0.016). in contrast, the analysis within the group of adults demonstrated that adults wrote significantly more list continuations after the sentences with list signals than after the sentences without list signals both when the connective en outre was present (log odds ratio=0.85, se=0.43, p=0.046) and absent (log odds ratio=1.91, se=0.38, p <.0001). however, the proportion of lists in the adult productions did not significantly change between the sentences with the connective en outre and those without any connective, both when adjectives of quantity were present (log odds ratio=-0.21, se=0.49, p=0.662) and absent (log odds ratio=0.73, se=0.39, p=0.061) in the task items. to summarize, similarly to the experiment 2, teenagers were more sensitive to the adjectives of quantity in the task items that were not followed by a connective, while adults were sensitive to the non-connective signals in both conditions, independently of the connective en outre. moreover, the presence of the connective en outre together with the non-connective signals significantly reduced the proportion of list productions by teenagers, but did not affect the proportion of lists produced by adult speakers. figure 6. proportions of different types of relations in the continuation after the task items without a connective (from experiment 1) and after those with the connective en outre (from experiment 3) across teenagers and adults. no en outre present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s teenagers **** ** no en outre present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s relation types in continuations other consequence addition list adults **** * exploring the sensitivity to non-connective signals of coherence relations 69 estimate se z p all participants (intercept) -1.86 0.33 -5.57 <.001 list signal 1.89 0.30 6.30 <.001 en outre 0.78 0.42 1.86 0.063 teenagers -0.21 0.34 -0.62 0.535 list signal*en outre -1.00 0.31 -3.18 0.001 list signal*teenagers -0.73 0.22 -3.24 0.001 en outre *teenagers -0.85 0.50 -1.69 0.091 list signal*en outre *teenagers 0.35 0.37 0.95 0.342 teenagers (intercept) -2.09 0.26 -8.18 <.001 list signal 1.16 0.27 4.35 <.001 en outre -0.01 0.30 -0.02 0.983 list signal*en outre -0.69 0.25 -2.79 0.005 adults (intercept) -1.93 0.34 -5.61 <.001 list signal 1.97 0.40 4.97 <.001 en outre 0.73 0.39 1.88 0.061 list signal*en outre -0.94 0.40 -2.33 0.020 table 7. model’s estimates for the best fitting model in the additional analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after those with the connective en outre (from experiment 3). note. in the model for all participants, conditional r2δ=.66, marginal r2δ=.29; for teenagers, conditional r2δ=.28, marginal r2δ=.03; for adults, conditional r2δ=.41, marginal r2δ=.10. 7.4.3 additional comparative analysis between experiment 1 and experiment 3 for the connective ainsi the final model for the analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after the items with the additive connective ainsi (from experiment 3), included list signal (absent versus present), connective (no connective versus ainsi), and group (adults versus teenagers) as fixed factors (both main and interaction effects), and item and participant as random intercepts (see table 8 and figure 7). this analysis shows that the presence of the non-connective signals significantly increased the production of lists for all the participants. however, overall, the presence of the consequence connective ainsi almost completely prevented participants from writing list continuations. the two separate within-group analyses confirmed general trends revealed in the analysis for all participants. the only difference was that when the consequence connective ainsi was present in the task sentences, adult speakers were not sensitive to the non-connective list signals (log odds ratio=1.09, se=0.81, p=0.176). in contrast, teenagers responded to the presence of the tskhovrebova, zufferey and gygax 70 adjectives of quantity and produced slightly more lists even in the presence of the consequence connective ainsi (log odds ratio=0.78, se=0.40, p=0.049). we noticed however that not all participants who produced list continuations after the connective ainsi interpreted it as a consequence connective. out of 115 continuations, ainsi was treated as a connective of consequence in only 11 of them. in the other 104 continuations, participants started their sentence with que and, this way, used it as an additive conjunction ainsi que ‘as well as’. in other words, some participants changed the connective intended in the task. as a result, it is complicated to interpret the effects of ainsi as a connective of consequence on list inference generation. figure 7. proportions of different types of relations in the continuation after the task items without a connective (from experiment 1) and after those with the connective ainsi (from experiment 3) across teenagers and adults. no ainsi present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s teenagers **** * **** *** no ainsi present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s relation types in continuations other consequence addition list **** **** *** adults exploring the sensitivity to non-connective signals of coherence relations 71 estimate se z p all participants (intercept) -1.93 0.45 -4.28 <.001 list signal 1.96 0.32 6.11 <.001 ainsi -3.50 0.90 -3.91 <.001 teenagers -0.31 0.49 -0.64 0.525 list signal*ainsi -0.84 0.74 -1.13 0.261 list signal*teenagers -0.74 0.23 -3.18 0.001 ainsi *teenagers 2.09 0.98 2.14 0.032 list signal*ainsi *teenagers 0.37 0.80 0.47 0.641 teenagers (intercept) -2.23 0.32 -6.90 <.001 list signal 1.20 0.28 4.26 <.001 ainsi -1.49 0.47 -3.16 0.002 list signal*ainsi -0.42 0.33 -1.26 0.209 adults (intercept) -2.09 0.48 -4.37 <.001 list signal 2.12 0.39 5.37 <.001 ainsi -3.31 0.95 -3.48 <.001 list signal*ainsi -1.02 0.76 -1.34 0.180 table 8 model’s estimates for the best fitting model in the additional analysis, comparing the production of list continuations after the task items without a connective (from experiment 1) and after those with the connective ainsi (from experiment 3). note. in the model for all participants, conditional r2δ=.48, marginal r2δ=.17; for teenagers, conditional r2δ=.43, marginal r2δ=.09; for adults, conditional r2δ=.62, marginal r2δ=.32. 7.4.4 comparative analysis between experiment 2 and experiment 3 in order to see whether connectives with different frequencies had a different impact on the generation of list inference, we performed an analysis, contrasting the results from experiment 2, which included more frequent connectives, and from experiment 3, which assessed less frequent connectives. however, given the issue in the interpretation of results after the connective ainsi, we excluded all the results for both connectives of consequence (ainsi and donc) from this analysis and focused only on the two connectives of additive relations (en plus and en outre). the statistical procedure remained the same as in previous analyses. we also made three separate models for different groups of participants (for teenagers, adults, and all participants together) and reported the details of model selection for teenagers and adults in the online appendix (see table s9). similar to previous analyses, the measures of exposure to print did not predict the variation in the sensitivity to list signals. finally, treatment contrasts tskhovrebova, zufferey and gygax 72 were applied to the factor of connective, where en plus was set as a reference level for comparison in all three models. the final model for all participants included list signal, connective, and group as fixed factors (both main and interaction effects), and item and participant as random intercepts (see table 9). comparing the results from all participants revealed that teenagers produced significantly fewer list continuations than adults after the prompt including the connective en outre. however, all other interactions were not statistically significant. the separate withingroup analyses demonstrated that connective frequency played a role only for the group of teenagers, as they produced significantly fewer list continuations after the less frequent connective en outre than after the more frequent en plus, both when list signals were absent (log odds ratio=1.30, se=0.28, z=4.67, p=<.0001) or present (log odds ratio=1.13, se=0.27, z=4.16, p=<.0001) in the task. as for the group of adults, the frequency of connectives did not affect the proportion of list continuations. estimate se z p all participants (intercept) -0.82 0.32 -2.56 0.011 list signal 0.88 0.31 2.85 0.004 en outre -0.09 0.41 -0.21 0.830 teenagers 0.19 0.33 0.57 0.569 list signal*en outre 0.11 0.32 0.34 0.735 list signal*teenagers -0.41 0.26 -1.56 0.119 en outre*teenagers -1.24 0.49 -2.53 0.011 list signal*en outre*teenagers 0.06 0.39 0.16 0.870 table 9. model’s estimates of the best fitting models the analysis, comparing the task with the more frequent connective en plus ‘in addition’ (from experiment 2) and the task with less frequent connective en outre ‘in addition’ (from experiment 3). note. in the model for all participants, conditional r2δ= .30, marginal r2δ= .07. 7.4.5 discussion the results of the experiment 3 were similar to those from the experiment 2. it was shown that the presence of the consequence connective ainsi, similar to the more frequent consequence connective donc, almost completely overrode the inference from the non-connective signals of the list relation in both groups of participants. as for the additive connective en outre, the effects again were not the same for the two age groups. teenagers were sensitive to the adjectives of quantity only in the task items that were not followed by the additive connective en outre, while adults remained sensitive to the non-connective signals, independently of the additive connective. in general, the presence of en outre significantly decreased the production of list continuations in teenagers in comparison to the sentences not followed by any connective and to those followed by the more frequent additive connective en plus. however, this finding should be interpreted with caution, as it is based on a comparison between two separate experiments. to sum up, the findings from experiments 2 and 3 demonstrated that the combination of list signals with additive connectives did not significantly increase the production of list continuations, but rather decreased (en outre) or left unchanged (en plus). given that these connectives signal a more generic additive relation, they can be used to express the relation of list, but are not limited to it. as a result, when a connective expressing a more generic additive relation is used together with a non-connective signal of a more specific list relation, it does not significantly improve the inference for a more specific list signal. this effect may stem from exploring the sensitivity to non-connective signals of coherence relations 73 the fact that the inference of a more generic additive relation, coming from a more salient and monofunctional signal such as connective, competes with the inference of the list relation, coming from a less prominent and non-monofunctional non-connective signal. we make in the next section an additional analysis aiming to assess whether participants were more sensitive to the additive connectives and produced significantly more additive continuations in the conditions that included additive connectives en plus and en outre. 8 analysis of additive continuations after the sentences with additive connectives en plus and en outre 8.1 comparative analysis between experiment 1 and experiment 2 for the connective en plus in order to examine whether more additive continuations were produced after the items including the additive connective en plus, we created a statistical model, comparing the proportion of additive continuations after the sentences without any connective (from experiment 1) and those followed by the connective en plus (from experiment 2). the results of both age groups were analyzed together, as we did not need to include the measures of exposure to print in the analysis. the details on model selections can be found in the online appendix also for this analysis. results show that both groups of participants indeed produced significantly more additive continuations after the sentences containing the additive connective en plus than after the sentences without any connective (see table 10 for the model’s estimates and figure 8). the sensitivity to the frequent additive connective en plus was not significantly different between the two age groups. moreover, in the sentences without any connective, participants produced more additive continuations when the adjectives of quantity were absent (log odds ratio=0.73, se=0.25, p=0.003). in the sentences including the additive connective, the presence of adjectives of quantity did not affect the proportion of additive continuations (log odds ratio= 0.47, se=0.28, p=0.086). finally, when both types of signals were absent in the task sentences, adults on average wrote more additive continuations than teenagers (log odds ratio= 0.48, se=0.21, p=0.022). figure 8. proportions of different types of relations in the continuation after the task items without a connective (from experiment 1) and after those with the connective en plus (from experiment 2) across teenagers and adults. no en plus present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s ** *** teenagers no en plus present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s relation types in continuations other consequence addition list adults ** *** * tskhovrebova, zufferey and gygax 74 estimate se z p (intercept) -0.82 0.30 -2.70 0.007 list signal -1.00 0.31 -3.18 0.001 en plus 1.34 0.38 3.55 <0.001 teenagers -0.57 0.29 -1.98 0.048 list signal*en plus 0.27 0.34 0.79 0.427 list signal*teenagers 0.52 0.22 2.35 0.019 en plus*teenagers 0.18 0.42 0.42 0.672 list signal*en plus*teenagers -0.07 0.34 -0.20 0.838 table 10. model’s estimates of the best fitting model in the analysis, comparing additive continuations after the task items without a connective (from experiment 1) and after those with the connective en plus (from experiment 2). note. conditional r2δ =.31, marginal r2δ = .10. 8.2 comparative analysis between experiment 1 and experiment 3 for the connective en outre in order to examine whether more additive continuations were produced after the items including the additive connective en outre, we created a statistical model, comparing the proportion of additive continuations after the sentences without any connective (from experiment 1) and those followed by the connective en outre (from experiment 3). results show that there were also significantly more additive continuations after the sentences containing the additive connectives en outre than after the sentences without connectives (see table 11 for the model’s estimates and figure 9). however, adults were more sensitive to the less frequent additive connective en outre, as they produced significantly more additive sentences than teenagers after the task items including this connective (log odds ratio=1.03, se=0.26, p<0.001). finally, as in the analysis for the connective en plus, after the items without any connective, participants produced more additive continuations when the adjectives of quantity were absent (log odds ratio= 0.76, se=0.26, p=0.004). in contrast, after the sentences including the additive connective en outre, the presence of adjectives of quantity did not affect the proportion of additive continuations (log odds ratio=0.53, se=0.29, p=0.073). figure 9. proportions of different types of relations in the continuation after the task items without a connective (from experiment 1) and after those with the connective en outre (from experiment 3) across teenagers and adults. no en outre present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o rt io n o f co n ti n u a ti o n s teenagers ** *** no en outre present absent present absent 0% 25% 50% 75% 100% list signal in task p ro p o r ti o n o f c o n ti n u a ti o n s relation types in continuations other consequence addition list adults *** *** ** exploring the sensitivity to non-connective signals of coherence relations 75 estimate se z p (intercept) -0.79 0.28 -2.81 0.005 list signal -1.01 0.30 -3.37 <0.001 en outre 1.18 0.34 3.47 <0.001 teenagers -0.56 0.27 -2.10 0.036 list signal*en outre 0.16 0.31 0.53 0.597 list signal*teenagers 0.51 0.22 2.34 0.019 en outre*teenagers -0.79 0.40 -1.99 0.047 list signal*en outre*teenagers 0.13 0.36 0.37 0.709 table 11. model’s estimates of the best fitting model in the analysis, comparing additive continuations after the task items without a connective (from experiment 1) and after those with the connective en outre (from experiment 3). note. conditional r2δ =.25, marginal r2δ =.05. 9 general discussion in the current set of experiments, we examined whether native french-speaking teenagers were sensitive to signals of the list relation, expressed by the adjectives of quantity plusieurs ‘several’ and différents ‘various’ (experiments 1, 2, 3). we also assessed whether this sensitivity was modulated by the presence of another signal of coherence relation, namely connectives of additive and consequence relations, varying in frequency (experiments 2, 3). finally, we systematically contrasted the results obtained by teenagers with those of a control group of adults, and assessed whether their performance in the main task was modulated by their linguistic competence, as measured by the author recognition test. 9.1 sensitivity to non-connective signals of list relation both groups of participants were sensitive to list signals, as they produced more continuations expressing a list relation when one of the adjectives of quantity was present in the first sentence of the task that did not include connectives (see the main analysis of experiment 1). however, teenagers' receptiveness to alternative list signals was still inferior to that of adults. this finding might indicate that sensitivity to alternative signals develops with age and the increasing linguistic experience that is normally associated to it. it is possible that teenagers are less sensitive to alternative signals than adults because they have not yet mastered non-sentenceinitial usage of coherence markers. indeed, when teenagers used connectives in their own productions, they preferred to use them in sentence-initial position and only rarely used them in other positions. in contrast, adults produced connectives in different syntactic positions and even did so more frequently in non-sentence initial positions. in addition, the fact that linguistic experience and level of linguistic proficiency develop with age is reflected in the types of continuations produced by teenagers and adults. we observed that, across all experiments and conditions, teenagers produced more elliptic continuations that lacked subject or verb. out of 5604 continuations written by teenagers, 670 (12 %) were elliptic; whereas only 38 (2%) of the 2509 completions created by adults had an ellipsis. most ellipses were found in list continuations across both age groups (503 (75%) in teenagers and 36 (95%) in adults). note that some participants analyzed ainsi not as a connective of consequence, but as an additive connective ainsi que, by adding que in their continuation sentence (see example 21). since all such instances were elliptic, this accounted for most elliptic sentences produced by adults and an important part of ellipses produced by teenagers (see table 12). however, even when no connective was present in the prompt, the tskhovrebova, zufferey and gygax 76 proportion of elliptic sentences written by teenagers was still greater than that of adults (54% vs. 25%). (21) task sentences: l'enfant a surpris ses parents. il voulait comprendre pourquoi le ciel était bleu. ainsi, … ‘the child surprised his parents. he wanted to understand why the sky was blue. therefore, …’ continuation: que pourquoi la neige est-t-elle blache. ‘as well as why the snow is white.’ no connective ainsi donc en outre en plus total n teenagers 271 (.54) 86 (.17) 0 11 (.02) 135 (.27) 503 adults 9 (.25) 21 (.58) 0 0 6 (.17) 36 table 12. raw number (and proportion) of elliptic sentences in list continuations across all three experiments and all age groups. this finding may of course indicate that teenagers took the task less seriously and paid less attention to it. however, it may also mean that they have not yet mastered all the particularities of written language, which precisely tends to avoid ellipses (see, e.g., menzel, 2016). another indication of the fact that teenagers may not master the written modality is the usage of connective et ‘and’ in sentence-initial position produced in their own sentences. whereas in oral speech it is perfectly normal to use this connective in the beginning of the sentence, in written language it is not stylistically appropriate, as coordinating conjunctions are not possible in sentence-initial position according to reference grammars (see, e.g., riegel, pellat & rioul, 2021). 9.2 sensitivity to list signals combined with connectives when the task combined both non-connective signals and connectives, we found different effects in the production of list continuation sentences. first of all, the difference between the proportion of list continuations after the sentences including and not including the list signal was not the same within three experiments. after the cue sentences with connectives en plus and en outre, teenagers and adults produced more list continuations when a list signal was present than when it was absent. however, the observed effects were significant only for the group of adults, suggesting that teenagers are probably even less sensitive to the non-connective list signals when a more salient signal like connective is also present in the sentence (see comparisons reported in 6.4.2 and 7.4.2). the presence of list signals together with the connectives of consequence donc and ainsi did not have any effect on the generation of list inference. indeed, after the task passages followed by the connectives of consequence, the list relation was almost completely absent in the continuation sentences produced by both teenage and adult participants (see comparisons between the items with consequence connectives and those without connectives reported in 6.4.3 and 7.4.3). presumably, this means that connectives signaling the relation of consequence create a much stronger mental inference of this relation than do the non-connective list signals for the relation of list. however, the results for ainsi should be considered with caution, as in a significant number of cases, it was interpreted as a different type of signal (ainsi que), used for marking addition. secondly, we observed that in the condition without list signal, there were significantly more list continuations after en plus in comparison to the task without connectives in all age groups (see comparisons between the items with the additive connective en plus and those without connectives reported in 6.4.2). in other words, this means that even the additive connective en plus alone can generate inference of the list relation. however, when both en plus exploring the sensitivity to non-connective signals of coherence relations 77 and the list signal were present in the cue sentence, it did not significantly reinforce the inference of a list relation. as demonstrated in the analysis of additive continuations (see 8.1), en plus can generate not only an inference of the list relation, but also that of an additive relation. therefore, when both types of signals are present in the sentence, the additional additive function of the connective may compete with the inference of the list relation from the nonconnective signal. alternatively, and in line with findings of crible and demberg (2021), this effect may be due to a stronger inference power of connectives as a type of coherence signal compared to list signals within the related segments. in the condition without list signal, after the more infrequent additive connective en outre, the proportion of list continuations produced by teenagers was the same as in the task without connectives; while when combined with the list signals, there were even fewer list continuations in comparison to the same condition in the task without connectives (see comparisons between the items with additive connective en outre and those without connectives reported in 7.4.2). as for adults, although they produced slightly more list sentences after the connective en outre, their proportion was not significantly higher than in the task without connectives in both conditions (with and without list signal). as far as the comparison of connectives with different frequencies was concerned, teenagers produced significantly more list continuations after the more frequent additive connective en plus than after the infrequent connective en outre. in contrast, there was no such difference between the effects of the two additive connectives for adults (see 7.4.4 for the comparative analysis between items with the less frequent additive connective and the more frequent one). taken together, these findings suggest that en outre does not facilitate the inference of the list relation and may even hinder this inference, especially in the case of young speakers. indeed, teenagers may be less familiar with the less frequent connective en outre compared to adults. hence, it is more difficult for them to infer a more specific list relation. this finding as well as the fact that teenagers produced some list continuations even in the presence of the consequence connective ainsi also suggest that the mastery of a specific connective may be an additional factor affecting the inference generation. in addition, similar to the connective en plus, a more generic additive meaning triggered by en outre may override the more specific list meaning, as suggested by the analysis of additive continuations (see 8.2). nevertheless, the fact that an important number of list relations was produced even in the presence of more salient, stronger, and prototypical signals of coherence such as the additive connectives en plus and en outre, shows that non-connective list signals are an important source for inferring a list relation. these signals start to be perceived and to affect discourse inferences as early as at the age of 12 and their impact increases with age. it is however important to point out that the presented comparisons should be considered with caution, as they are made between experiments. in contrast to scholman et al. (2020), we did not find an effect of the author recognition test on the sensitivity to non-connective list signals both for teenagers and adults. although the french versions of the art were strong predictors for the use of connectives in other studies (tskhovrebova, zufferey & tribushinina, 2022; zufferey & gygax, 2020b), it probably requires further validation in french. as a matter of fact, the french version of this test included 80 items, while the english art, used by scholman et al. (2020), consisted of 130 items, which might have rendered this version a more sensitive measure. moreover, the performance of both groups of participants was not very high on the measures of exposure to print (teenagers: m=6.44, sd=6.08, observed range: -11 to 28, possible range: -40 to 40; adults: m=8.89, sd=5.43, observed range: -1 to 23, possible range: -40 to 40). this may have created a floor effect that did not allow us to track individual variation. finally, the lack of effect of art scores in the present experiments may also suggest that exposure to print does not necessarily reflect individual differences in the ability to infer an intended coherence relation. it is possible that this ability constitutes a specific type of linguistic competence that should be assessed with a more sensitive measure. tskhovrebova, zufferey and gygax 78 9.3 limitations and future directions the present set of experiments had several limitations that should be taken into account in follow-up research. it is important to point out that we examined the effect of non-connective and connective signals on production data and can only speculate about the comprehension level of the coherence relations included in our study. in other words, the continuation task provided evidence only about one type of coherence relation that a participant chose to write down, while all other relations that might have been inferred as well remain unknown. this suggests that participants might have been more sensitive to alternative list signals, but the task did not always reveal this sensitivity. moreover, one of the most important limitations is related to the design of the experiments, as they involved between-participant design. as a result, the comparisons made between experiments 1, 2, and 3 should be interpreted with caution. future research should therefore address the issue of the design and focus on comprehension measures in order to complement our findings. as for the interpretations of the results on the interaction between adjectives of quantity and additive connectives, it should be noticed that since the relation of list is a subtype of the relation of addition, we cannot exclude that in some continuations both relations simply coexisted, without necessarily competing with each other. furthermore, the analysis of the connective insertions in the participants’ productions have hinted that, perhaps, temporal connectives, such as ensuite and puis, may be even better suited for marking list relations and should be analyzed in future studies. finally, an important contribution to future research would be to unveil other types of non-connective signals that can generate coherence inferences when used alone or together with connectives, and to continue the examination of other linguistic competences that may better explain individual variation in speakers' sensitivity to non-connective signaling. 10 conclusion taken together, the results of the current series of experiments suggest that expressions of quantity are an important source for the inference of the list relation as early as in teenage years, even though the sensitivity to these non-connective signals still develops into adulthood. the fact that the combination of a non-connective signal with the connective en plus did not significantly increase the inference of a list relation in both age groups indicates that a more generic additive relation, signaled by this connective, may compete with a more specific relation of list. furthermore, it seems that the inference of the list relation in teenagers is inhibited by a less frequent additive connective en outre, and is almost completely hindered by both types of consequence connectives. ultimately, the degree of exposure to print, as measured by the art on our data, does not predict the individual differences in the sensitivity to the adjectives of quantity as signals of the list relation. more globally, the presented set of experiments shows that the examination of how different types of coherence signals combine with each other opens many new avenues of enquiry for future research. this type of research sheds light onto the linguistic devices that can reinforce or inhibit the generation of a certain coherence relation, and thus, allows to understand the functioning of this relation better. acknowledgements this work was funded by the swiss national science foundation grant 100012_184882. disclosure statement the authors report there are no competing interests to declare. data availability statement the data that support the findings of this study are openly available in osf repository at https://osf.io/ca7hm/?view_only=b59ac0969a3045f5978074f7015b6291. https://osf.io/ca7hm/?view_only=b59ac0969a3045f5978074f7015b6291 exploring the sensitivity to non-connective signals of coherence relations 79 references atilf. (1998-2022). base textuelle frantext (en ligne) [data set]. atilf-cnrs & université de lorraine. https://www.frantext.fr/ au, t. k. (1986). a verb is worth a thousand words: the causes and consequences of interpersonal events implicit in language. journal of memory and language, 25, 104– 122. https://doi.org/10.1016/0749-596x(86)90024-0 bates, d., maechler, m., bolker, b., & walker, s. (2015). fitting linear mixed-effects models using lme4. journal of statistical software, 67(1), 1–48. https://doi.org/10.18637/jss.v067.i01 benzitoun c., debaisieux, j.-m., & deulofeu, j. (2016). le projet orféo: un corpus détude pour le français contemporain. corpus, 15. https://doi.org/10.4000/corpus.2936 berman, r. (2004). between emergence and mastery: the long developmental route of language acquisition. in r. berman (ed). language development across childhood and adolescence (pp. 9–34). john benjamins publishing company. https://doi.org/10.1075/tilar.3 bloom, l., lahey, m., hood, l., lifter, k., & fiess, k. (1980). complex sentences: acquisition of syntactic connectives and the semantic relations they encode. journal of child language, 7, 235–261. https://doi.org/10.1017/s0305000900002610 blything, l. p., & cain, k. (2016). children’s processing and comprehension of complex sentences containing temporal connectives: the influence of memory on the time course of accurate responses. developmental psychology, 52(10), 1517-1529. http://dx.doi.org/10.1037/dev0000201 blything, l. p., davies, r., & cain, k. (2015). young children's comprehension of temporal relations in complex sentences: the influence of memory on performance. child development, 86(6), 1922–1934. https://doi.org/10.1111/cdev.12412 brown, r., & fish, d. (1983). the psychological causality implicit in language. cognition, 14(3), 237–273. https://doi.org/10.1016/0010-0277(83)90006-9 carlson, l., marcu, d., & okurowski, m. e. (2002). rst discourse treebank, ldc2002t07. https://doi.org/10.35111/4w31-m996 champaud, c. & bassano, d. (1994). french concessive connectives and argumentation: an experimental study in eightto ten-year-old children. journal of child language, 21, 415–438. https://doi.org/10.1017/s0305000900009338 crible, l. (2021). negation cancels discourse-level processing differences: evidence from reading times in concession and result relations. j psycholinguist res, 50, 1283–1308. https://doi.org/10.1007/s10936-021-09802-2 crible, l., & demberg, v. (2021). the role of non-connective discourse cues and their interaction with connectives. pragmatics and cognition, 27(2), 313–338. https://doi.org/10.1075/pc.20003.cri crible, l., & pickering, m. j. (2020). compensating for processing difficulty in discourse: effect of parallelism in contrastive relations. discourse processes, 57, 862–879. https://doi.org/10.1080/0163853x.2020.1813493 crosson, a., lesaux, n., & martiniello, m. (2008). factors that influence comprehension of connectives among language minority children from spanish-speaking backgrounds. applied psycholinguistics, 29, 603–625. https://doi.org/10.1017/s0142716408080260 cunningham, a., & stanovich, k. (1990). assessing print exposure and orthographic processing skill in children: a quick measure of reading experience. journal of educational psychology, 82(4), 733–740. https://doi.org/10.1037/0022-0663.82.4.733 das, d. (2014). signalling of coherence relations in discourse (phd dissertation). simon fraser university, burnaby, canada. das, d., & taboada, m. (2018). rst signalling corpus: a corpus of signals of coherence relations. language resources and evaluation, 52(1), 149–184. https://doi.org/10.1007/s10579-017-9383-x https://www.frantext.fr/ https://doi.org/10.1016/0749-596x(86)90024-0 https://doi.org/10.18637/jss.v067.i01 https://doi.org/10.4000/corpus.2936 https://doi.org/10.1075/tilar.3 https://doi.org/10.1017/s0305000900002610 http://dx.doi.org/10.1037/dev0000201 https://psycnet.apa.org/doi/10.1016/0010-0277(83)90006-9 https://doi.org/10.35111/4w31-m996 https://doi.org/10.1075/pc.20003.cri https://doi.org/10.1017/s0142716408080260 https://doi.org/10.1037/0022-0663.82.4.733 https://doi.org/10.1007/s10579-017-9383-x tskhovrebova, zufferey and gygax 80 evers-vermeul, j., & sanders, t. (2009). the emergence of dutch connectives; how cumulative cognitive complexity explains the order of acquisition. journal of child language, 36, 829–854. https://doi.org/10.1017/s0305000908009227 grisot, c., & blochowiak., j. (2021). temporal relations at the sentence and text genre level: the role of linguistic cueing and non-linguistic biases—an annotation study of a bilingual corpus. corpus pragmatics, 5, 379–419. https://doi.org/10.1007/s41701-02100104-5 kehler. a. (1994). temporal relations: reference or discourse coherence? proceedings of the 32nd annual meeting on association for computational linguistics (acl '94). association for computational linguistics, usa, 319–321. https://doi.org/10.3115/981732.981779 kidd, e., donnelly, s., & christiansen, m. h. (2018). individual differences in language acquisition and processing. trends in cognitive sciences, 22(2), 154–169. https://doi.org/10.1016/j.tics.2017.11.006 kleijn, s. (2018). clozing in on readability: how linguistic features affect and predict text comprehension and on-line processing. [doctoral dissertation, utrecht university]. lot. https://www.lotpublications.nl/documents/493_fulltext.pdf kleijn, s., pander maat, h. l. w., & sanders, t. j. m. (2019). comprehension effects of connectives across texts, readers, and coherence relations. discourse processes, 56, 447–464. https://doi.org/10.1080/0163853x.2019.1605257 knaebel, r., & stede, m. (2022). towards identifying alternative-lexicalization signals of discourse relations. proceedings of the 29th international conference on computational linguistics. international committee on computational linguistics, gyeongju, republic of korea, 837–850. retrieved from https://aclanthology.org/2022.coling-1.70 koehn, p. (2005). europarl: a parallel corpus for statistical machine translation. conference proceedings: the tenth machine translation summit, phuket, thailand, 79–86. retrieved from https://homepages.inf.ed.ac.uk/pkoehn/publications/europarlmtsummit05.pdf koornneef, a. w., & van berkum, j. j. a. (2006). on the use of verb-based implicit causality in sentence comprehension: evidence from selfpaced reading and eye-tracking. journal of memory and language, 54, 445–465. https://doi.org/10.1016/j.jml.2005.12.003 kuznetsova, a., bruun brockhoff, p., & christensen, r. h. b. (2014). lmertest package: tests in linear mixed effects models. journal of statistical software, 82(13), 1–26. https://doi.org/10.18637/jss.v082.i13 lascarides, a., & asher, n. (1993). temporal interpretation, discourse relations and commonsense entailment. linguist philos 16, 437–493 https://doi.org/10.1007/bf00986208 lenth, r. (2020). emmeans: estimated marginal means, aka least-squares means. r package version 1.5.1. https://cran.r-project.org/package=emmeans liu, y. (2019). beyond the wall street journal: anchoring and comparing discourse signals across genres. proceedings of the workshop on discourse relation parsing and treebanking 2019. association for computational linguistics, minneapolis, usa, 72– 81. https://doi.org/10.18653/v1/w19-2710 liu, y., & zeldes, a. (2019). discourse relations and signaling information: anchoring discourse signals in rst-dt. proceedings of the society for computation in linguistics, 2, article 35, 314–317. https://doi.org/10.7275/vh3w-4240 menzel, k. (2016). understanding english-german contrasts: a corpus-based comparative analysis of ellipses as cohesive devices. [doctoral dissertation, saarland university]. der wissenschaftsserver der universität des saarlandes. http://doi.org/10.22028/d291-23663 le monde. (1987–2012). text corpus of ”le monde”. distributed via elra: elra-id w0015, islrn 421-401527-366-2. https://doi.org/10.3115/981732.981779 https://www.lotpublications.nl/documents/493_fulltext.pdf https://doi.org/10.1080/0163853x.2019.1605257 https://aclanthology.org/2022.coling-1.70 https://doi.org/10.1016/j.jml.2005.12.003 https://doi.org/10.18637/jss.v082.i13 https://doi.org/10.1007/bf00986208 https://doi.org/10.18653/v1/w19-2710 http://doi.org/10.22028/d291-23663 exploring the sensitivity to non-connective signals of coherence relations 81 nippold, m., schwartz, i., & undlin, r. (1992). use and understanding of adverbial conjunctions: a developmental study of adolescents and young adults. journal of speech and hearing research, 35, 108–118. https://doi.org/10.1044/jshr.3501.108 nippold, m. (2004). research on later language development international perspectives. in r. berman (ed). language development across childhood and adolescence (pp. 1–8). john benjamins publishing company. https://doi.org/10.1075/tilar.3 nippold, m. (2008). later language development: school-age children, adolescents, and young adults (3rd ed., 2nd printing). pro-ed. péry-woodley, m., ho-dac, l.m., rebeyrolle, j., tanguy, l., & fabre, c. (2017). a corpusdriven approach to discourse organisation: from cues to complex markers. dialogue discourse, 8, 66–105. https://doi.org/10.5087/dad.2017.103 pyykkönen, p., & järvikivi, j. (2010). activation and persistence of implicit causality information in spoken language comprehension. experimental psychology, 57(1), 5– 16. https://doi.org/10.1027/1618-3169/a000002 riegel, m., pellat, j.-c., & rioul, r. (2021). grammaire méthodique du français. (8th ed.). paris: presses universitaires de france. rohde, h., & w. s. horton (2014). anticipatory looks reveal expectations about discourse relations. cognition 133(3), 667–691. https://doi.org/10.1016/j.cognition.2014.08.012 roze, c., danlos, l., & muller, p. (2012). lexconn: a french lexicon of discourse connectives. discours, 10, 1–15. https://doi.org/10.4000/discours.8645 rstudio team. (2015). rstudio: integrated development for r. rstudio. inc., boston, ma. http://www.rstudio.com/ sanders, t. j. m., spooren, w. p. m., & noordman, l. g. m. (1992). toward a taxonomy of coherence relations. discourse processes, 15(1), 1–35. https://doi.org/10.1080/01638539209544800 scholman, m.c.j., demberg, v., & sanders, t. j. m. (2020). individual differences in expecting coherence relations: exploring the variability in sensitivity to contextual signals in discourse. discourse processes, 57(10), 844–861. https://doi.org/10.1080/0163853x.2020.1813492 schwab, j., & liu, m. (2020). lexical and contextual cue effects in discourse expectations: experimenting with german ‘zwar…aber’ and english ‘true/sure…but’. dialogue & discourse, 11(2), 74–109. https://doi.org/10.5087/dad.2020.203 stainthorp, r. (1997). a childrens author recognition test: a useful tool in reading research. journal of research in reading, 20(2), 148–158. https://doi.org/10.1111/14679817.00027 stanovich, k., & west, r. (1989). exposure to print and orthographic processing. reading research quarterly, 24(4), 402–433. https://doi.org/10.2307/747605 tskhovrebova, e., zufferey, s., & gygax, p. (2022). individual variations in the mastery of discourse connectives from teenage years to adulthood. language learning, 72(2): 412–455. https://doi.org/10.1111/lang.12481 tskhovrebova, e., zufferey, s., & tribushinina, e. (2022). french-speaking teenagers’ mastery of connectives: the role of vocabulary size and exposure to print. applied psycholinguistics, 43(5), 1141–1163. http://doi.org/10.1017/s0142716422000303 van silfhout, g., evers-vermeul, j., & sanders, t. j. m. (2015). connectives as processing signals: how students benefit in processing narrative and expository texts. discourse processes, 52, 47–76. https://doi.org/10.1080/0163853x.2014.905237 volodina, a., & weinert, s. (2020). comprehension of connectives: development across primary school age and influencing factors. frontiers in psychology, 11, 814. https://doi.org/10.3389/fpsyg.2020.00814 webber, b., prasad, r., lee, a., & joshi, a. (2019). the penn discourse treebank 3.0 annotation manual. https:// catalog.ldc.upenn.edu/docs/ldc2019t05/ pdtb3annotation-manual.pdf. https://doi.org/10.1075/tilar.3 https://doi.org/10.5087/dad.2017.103 https://doi.org/10.1016/j.cognition.2014.08.012 https://doi.org/10.4000/discours.8645 http://www.rstudio.com/ https://doi.org/10.1080/01638539209544800 https://doi.org/10.1080/0163853x.2020.1813492 https://doi.org/10.5087/dad.2020.203 https://doi.org/10.1111/1467-9817.00027 https://doi.org/10.1111/1467-9817.00027 https://doi.org/10.2307/747605 https://doi.org/10.1111/lang.12481 http://doi.org/10.1017/s0142716422000303 https://doi.org/10.3389/fpsyg.2020.00814 tskhovrebova, zufferey and gygax 82 wetzel, m., crible, l., & zufferey, s. (2022). processing clause-internal discourse relations in a second language: a case study of specifications in german and french. j. second lang. stud. 5 (2), 206e234. https://doi.org/10.1075/jsls.21032.wet. zeldes, a. (2017). the gum corpus: creating multilayer resources in the classroom. language resources and evaluation, 51(3), 581–612. zeldes, a., & liu, y. (2020). a neural approach to discourse relation signal detection. dialogue & discourse, 11(2), 1–33. https://doi.org/10.5087/dad.2020.201 zufferey, s., & gygax, p. (2020a). ‘roger broke his tooth. however, he went to the dentist’: why some readers struggle to evaluate wrong (and right) uses of connectives. discourse processes, 57, 184–200. http://dx.doi.org/10.1080/0163853x.2019.1607446 zufferey, s., & gygax, p. (2020b). do teenagers know how to use connectives from the written mode? lingua, 234, 102779, 1–12. https://doi.org/10.1016/j.lingua.2019.102779 https://doi.org/10.5087/dad.2020.201 http://dx.doi.org/10.1080/0163853x.2019.1607446 https://doi.org/10.1016/j.lingua.2019.102779 1 introduction 2 the role of non-connective signals for coherence marking 3 teenagers’ competence with connectives 4 our set of experiments 5 experiment 1 5.1 participants 5.2 materials and procedure 5.2.1 story-continuation task 5.2.2 author recognition tests 5.3 analysis 5.4 results and discussion 5.5 additional analysis of the distribution of connectives in list continuations 6 experiment 2 6.1 participants 6.2 materials and procedure 6.3 analyses 6.4 results 6.4.1 sensitivity to connectives and list signals in experiment 2 6.4.2 additional comparative analysis between experiment 1 and experiment 2 for the connective en plus 6.4.3 additional comparative analysis between experiment 1 and experiment 2 for the connective donc 6.4.4 discussion 7 experiment 3 7.1 participants 7.2 materials and procedure 7.3 analyses 7.4 results 7.4.1 sensitivity to connectives and list signals in experiment 3 7.4.2 additional comparative analysis between experiment 1 and experiment 3 for the connective en outre 7.4.3 additional comparative analysis between experiment 1 and experiment 3 for the connective ainsi 7.4.4 comparative analysis between experiment 2 and experiment 3 7.4.5 discussion 8 analysis of additive continuations after the sentences with additive connectives en plus and en outre 8.1 comparative analysis between experiment 1 and experiment 2 for the connective en plus 8.2 comparative analysis between experiment 1 and experiment 3 for the connective en outre 9 general discussion 9.1 sensitivity to non-connective signals of list relation 9.2 sensitivity to list signals combined with connectives 9.3 limitations and future directions 10 conclusion dialogue & discourse 15(1) 45-76 doi: 10.5210/dad.2024.102 ©2024 derya çokal and klaus von heusinger this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). 45 german demonstrative pronouns in contrast derya çokal dcokal@uni-koeln.de university of cologne klaus von heusinger klaus.vonheusinger@uni-koeln.de university of cologne editor: manfred stede submitted 12/2023; accepted 07/2024; published online 08/2024 abstract german has two demonstrative pronouns: the der, die, das paradigm and the dieser, diese, dies(es) paradigm. previous studies mainly compared the anaphoric use of der with the personal pronoun er and observed that der refers to less prominent antecedents. however, there are only very few studies that have investigated the differences between these two demonstrative pronouns. we hypothesize that they differ in signaling topic persistence and in accessing contrastive antecedents. we tested these hypotheses in short texts that manipulated the contrast of the antecedent by inducing the expression ‘in contrast to’ vs. ‘together with’ (e.g., the cellist in contrast to the flautist vs. the cellist together with the flautist). results from our eye-tracking reading experiment (experiment 1), in which participants’ eyemovements were monitored while reading sentences, show that (i) readers preferred dieser when referring to the topic of a sentence, and (ii) dieser caused less processing difficulties than der in both contrast and no-contrast contexts. our sentence completion experiment (experiment 2) also confirmed that der and dieser are both used for anaphoric reference to a topical antecedent. collectively, our experiments provide evidence that dieser functions as inducing topic persistence. these results suggest that there is a need for further experimental investigation into the semantic factors and informational structures influencing the usage of demonstrative pronouns in german. keywords: demonstratives, anaphora, online reading, contrast, prominence. 1 introduction discourse management encompasses the organization and flow of information in a conversation or written text, and so plays a crucial role in effective communication. effective discourse management relies on establishing and maintaining clear referential structures to avoid ambiguity. referential structures deal with how referential expressions refer to entities in discourse. we can refer to entities/individuals in discourse by personal or demonstrative pronouns. in particular, german includes various types of demonstratives that can be utilized anaphorically. the most prevalent ones are the demonstratives from the der paradigm (i.e., der, die, das – also known as ‘dpronouns’) and those from the dieser paradigm (i.e., dieser, diese, dieses – ‘dem-pronouns’). for example, in (1) a speaker can say: (1a) ich habe einen perkussionisten und einen gitarristen auf der bühne gesehen. i saw a percussionist and a guitarist on the stage. http://creativecommons.org/licenses/by/3.0/ çokal and von heusinger 46 (1b) der+ perkussionist / er / der / dieser/ hat sehr fleißig gearbeitet. the percussionist / he / that one / this one/ worked very diligently. when referring to the percussionist or guitarist introduced in (1a), one can use a definite noun phrase (der perkussionist), a personal pronoun (er), or one of the two d-pronouns (e.g., der vs. dieser)1. with respect to (1b), scholars have been trying to determine which anaphoric form is selected and what factors play a role in the selection process. the most studied factors are structural elements such as grammatical role (subject vs. object vs. other arguments), linear position of antecedents (first vs. last), topichood (topic vs. non-topic), and thematic role (agent vs. patient) – (see the section on structural factors below.). in addition, der has been extensively investigated and specifically compared to the personal pronoun er using acceptability judgment tasks, sentence completion experiments, forced-choice tasks (e.g., bader & portele, 2019; bouma & hopp, 2007; schumacher et al., 2016), corpus data (bosch et al., 2003), self-paced reading, (bosch & umbach, 2007), visual world paradigm studies (ellert, 2013; schumacher et al., 2017; wilson, 2009), erp studies (schumacher et al., 2015; repp & schumacher, 2023), and an eye-tracking reading study (patterson & schumacher, 2020). a common result across these studies is that while der refers to less prominent antecedents, er refers to more prominent antecedents (e.g., potential antecedents of der: object in kaiser, 2011, or recent/last-mentioned entities, anti-topic bias in bosch & hinterwimmer, 2016; bosch & umbach, 2007; wilson, 2009; and potential antecedents of er: topic, subject, first-mentioned entity in bosch et al., 2003; bouma & hopp, 2007). furthermore, while the antecedent choice of er is flexible, the antecedent of der is not (bader & portele, 2019; kaiser, 2011; schumacher et al., 2015; 2016; 2017). it is worth noting that the comparison of the demonstrative pronouns der and dieser has received far less attention with online processing methods (e.g., eye-tracking reading or erp experiments). dieser has been primarily investigated using offline studies: (i.e., abraham 2002 (centering theory); ahrenholz, 2007; weinert 2007 (corpus studies); fuchs & schumacher, 2020 (sentence completion task); patil et al. 2020; 2023 (forced-choice and acceptability judgement tasks); patterson & schumacher, 2021 (acceptability judgement tasks); patterson & schumacher, 2023 (story completion tasks); zifonun et al., 1997; wöllstein et al., 2022 (in grammar); as well as diessel, 1999; himmelmann, 1997 (more typological research)). when discussed in sentences with two arguments (e.g. subject vs. object), the pronominal use of der and dieser is assumed to signal a topic shift. according to zifonun et al. (1997), abraham (2002); and bosch et al. (2007), der and dieser refer to a non-topic antecedent that introduced a referent that is taken up by the demonstrative as a topic in the current sentence, typically in the first position of the sentence. however, there is no explicit research on the contrastive or topic persistent functions of demonstrative pronouns in german. to fill this gap, we focus on the pronominal use of der and dieser, rather than on their determiner uses (i.e., prenominal use). we specifically focus on the masculine singular forms der/dieser, since these forms are fully specified for number and case, while the feminine forms die/diese are underspecified for case and number, which is not within the scope of the current study. moreover, to the best of our knowledge, there are still no studies that have examined the online reading processing of der and dieser in contrast and no-contrast contexts. by using both online and offline methods (i.e., eye-tracking reading and sentence completion experiments), the current study fills this gap and makes several novel contributions. firstly, we employ eye-movement recording during reading to examine the processing of der and dieser in contrast and no-contrast contexts. secondly, since previous studies have explored the thematic roles of verbs (i.e., agent vs. patient) 1 note that the use of der as in der perkussionist can be either the definite article der, die, das or the demonstrative determiner der, die, das. both paradigms partly overlap, but show different forms for the genitive singular, genitive plural and dative plural (see wöllstein et al., 2022: 744, §1297). german demonstrative pronouns 47 in the antecedent selection of demonstratives and personal pronouns, we contribute new results with respect to der and dieser. our study explores antecedent preferences of der and dieser when all parameters (e.g., thematic roles, recency, and animacy) are kept the same but the contrastiveness of the subject and topic of the previous sentence are controlled. such investigation provides a deeper understanding of the discourse and referential functions of demonstratives and how prominence – boosted by contrast – affects the selection of antecedents of der and dieser. 2 review of literature before delving into our experimental results, we give a brief review of previous studies that investigated the status of noun phrases in anaphora processing and production, including concepts such as the first-mentioned entity and the recency effect. following this, we discuss the outcomes of studies focusing specifically on the d-pronouns der and dieser. 2.1 structural factors in anaphora processing researchers have determined the factors that play a role in anaphora processing, including structural factors (e.g., first-mention advantage, subjecthood, parallelism, last-mentioned entity: arnold et al., 2000; gordon, et al., 1993; stevenson et al., 1994; 1995; 2000) and semantic factors (e.g., implicit causality, animacy: järvikivi et al., 2017; pyykkönen & järvikivi, 2010; van den hoven & ferstl, 2018). it has long been recognized that a recency effect results in ‘more recently introduced’ entities being more likely antecedents of referential expressions (hobbs, 1979). however, choosing the most recent mention is not necessarily the most effective strategy in anaphora resolution. previous studies also show a first-mention advantage, indicating the firstmentioned noun phrase is the preferred antecedent of an ambiguous pronoun (e.g., arnold et al., 2000; gordon, et al., 1993). in addition, while the results of kaiser and trueswell’s (2004) eyetracking study that examined the temporal progression of the first-mention effect initially revealed an inclination towards recency, this preference later shifted to the first-mentioned character. in a subsequent eye-tracking reading study, fukumura and van gompel (2015) demonstrated that the influence of mention order on pronoun comprehension could depend on the type of referential expression. for instance, the personal pronoun er in german has been shown to refer to the subject (e.g., bouma & hopp, 2007; schumacher et al., 2016), whereas der refers to the object (schumacher et al., 2015; 2016). another hypothesis, supported by numerous researchers, suggests that an entity’s prominence is influenced by its syntactic function rather than its sentence position. according to the subject preference account, prominence derives from grammatical role, with entities in a subject position being more salient than those in other grammatical roles (arnold et al., 2000; crawley et al., 1990; gordon et al., 1993; grosz et al., 1995).2 while providing an extensive characterization of topicality or subject matter is not within the scope of this paper, it is important to acknowledge that numerous researchers have pointed out that subject position often hosts topical entities (e.g., chafe, 1976; gundel, 1988; lambrecht, 1994; reinhart, 1981). the topic is what the sentence is about (reinhart, 1981). a relationship between topicality and givenness/pronominalization has been proposed (e.g., ariel, 1990; beaver, 2004). this implies a link between being a suitable antecedent for a subsequent pronoun and being topical. for instance, in the constructions ‘peter in contrast to paul came early’ or ‘peter together with paul came early’, peter is in the subject and topic position whereas the second argument paul is in a prepositional phrase. paul holds a lower position in terms of grammatical role hierarchies compared to peter. additionally, the use of contrastive construction further contributes to a higher prominence level for peter than paul. peter rather than paul is likely to be a more prominent antecedent for referential expressions. as a result, one could deduce that 2 we use the term saliency and prominence in the same way (cf, von heusinger & schumacher 2019 for further discussion). çokal and von heusinger 48 topicality functions as the overarching influence that shapes the prominence of a referent and steers pronoun interpretation by narrowing possible antecedent options.3 therefore, to the best of our knowledge there is still a lack of understanding of how contrastive structure (in contrast to) and topic persistence affect the online processing and production of der and dieser. before discussing the existing theories and accounts concerning the pronominal use of der and dieser (see section 2.3), we will briefly overview eye-tracking studies for demonstratives. 2.2 eye-tracking reading studies on demonstratives to our knowledge, there have been very few eye-tracking reading studies that have examined the online processing of demonstrative pronouns and compared them with the processing of singular pronouns (e.g., it) or personal pronouns (e.g., he) (cf. review article on demonstratives peeters et al., 2021). in this section, we specifically focus on eye-tracking reading studies on demonstratives, excluding visual world paradigm studies due to differences in eye-movement measures and data analysis (e.g., visual world paradigm studies: brown-schmidt et al., 2005; ellert, 2013; wilson, 2009). the eye-tracking reading studies have examined whether different referential expressions encode the distinct cognitive status of the intended referent in the mind of the addressee (çokal et al., 2014; 2018; patterson & schumacher, 2020). for instance, to examine whether singular pronoun (it) and demonstrative (this) signal different procedural instructions, çokal et al. (2018) ran eye-tracking reading experiments. they used three eye-movement measures (i.e., regression path times, second pass reading times, and total time) and four regions of interest (i.e., context, pronoun, disambiguation, and spillover). the interaction between referential expression type (it vs. this) and antecedent types (concrete entity vs. proposition) was observed in late measures such as second pass reading times and total time in the context region. similarly, the same interaction pattern was also evident in the regression path times (i.e., early measures) for the disambiguation region. their results were robust across eye-movement measures (early and late measures) and areas of interests (i.e., context and disambiguation). in another study, çokal et al., (2014) explored whether two english demonstratives (i.e., this and that) direct readers’ attention to different segments of a written discourse (i.e., adjacent frontier vs. distant frontier).4 to address this, they ran two eye-tracking reading experiments, with five regions of interest and three eye-movement measures (i.e., early measures: first pass reading and regression path times; late measure: second pass reading times). again, robust findings (i.e., this and that more readily access the adjacent frontier) were observed across early and late eye-movement measures and areas of interests (i.e., anaphora and disambiguation regions). while the previous two studies focused on english demonstratives and singular pronoun, there is another eye-tracking reading study (patterson & schumacher, 2020) that examined the processing of german personal pronoun (er) and demonstratives (der). patterson and schumacher (2020) predicted that in both accusative and dative items, it would be easier to process er when it refers to np1 (proto-agent) compared to np2 (proto-patient). conversely, it should be easier to process der when it refers to np2 compared to np1. in experiment 1, they had two regions of interest (i.e., pronoun and spill-over regions) and four eye-movement measures (i.e., early 3 it is also worth mentioning that antecedents joined by with showed no plural and a singular entity preference in a sentence completion experiment (hielscher & musseler, 1990). similarly, moxey et. al (2004) observed that the context ‘x with y’ allows individuals to be focused upon entities separately, and thus resulted in few plural continuations in the with condition. 4 the adjacent frontier-only hypothesis uses polanyi’s (1986) left-frontier/right-frontier distinction. polanyi uses “right frontier” to refer to the clause or group of contiguous clauses immediately adjacent to the referential expression, which can thus be argued to be salient; the “left frontier” refers to a clause or clauses separated from the deictic expression by intervening clauses or units (i.e., by the right frontier) and can thus be argued to be less salient. german demonstrative pronouns 49 measures: first-pass times and right-bound reading times and late measures: re-reading times and total times). in experiment 2, they had three regions of interest (i.e., pronoun, buffer, and repeated np regions) but used the same eye-movement measures as in experiment 1. in both experiments, der took longer to read than er in all eye-movement measures in the pronoun region. additionally, irrespective of anaphora type, in both first-pass times and right-bound reading times in the repeated np region, np2 conditions – specifically with the personal pronoun – had longer reading times/fixation times than np1 conditions. the precise timing for when prominence (i.e., np1 reference preference) is used to resolve ambiguity was not determined. overall, these studies employed meticulous experimental designs and analysis methods, incorporating a minimum of three eye-movement measures and examining between two to five regions of interest depending on their hypotheses. the studies examined the effects in both early and late eye-movement measures because each measure signals different linguistic and cognitive processing strategies (cf. rayner & liversedge, 2012 for linguistic and cognitive influences on eye movements in reading). one common challenge in anaphora resolution in eye-tracking reading experiments is discerning a clear and robust interaction pattern across all eye-movement measures and areas of interest. in addition, capturing the precise region where disambiguation occurs can be quite challenging (cf. patterson & schumacher, 2020). however, this does not diminish the insightful information provided by eye-movement studies. in fact, such studies elucidate how linguistic elements interact with discourse structure and cognitive processes in language comprehension, highlighting the complexity of cognitive mechanisms involved. 2.3 existing work on pronominal use of two demonstratives in german in addition to several studies that have focused on english demonstratives (çokal et al., 2014; 2018; fossard et al., 2012; maes et al., 2022a; 2022b), there has been behavioral research on demonstrative determiners (der and dieser +np) in written discourse in german (bader et al., 2022; fuchs & schumacher, 2020; patil et al., 2020). it has been claimed that dieser is conventionally limited to selecting a noun phrase at the end of a clause as its antecedent, since demonstratives generally avoid the most prominent discourse referents as antecedents (hinterwimmer & bosch, 2018; patterson & schumacher 2021; zifonun et al., 1997). specifically, patil et al. (2020) focused on transitive or ditransitive sentences with distinct grammatical and thematic roles and used a forced-choice task with a drop-down menu to select der, dieser, and er. they demonstrated that dieser can refer to nps at the beginning of sentences. however, this occurs particularly when the object comes before the subject in the preceding context sentence. consequently, patil et al. (2020) proposed that dieser exhibits a preference for an object reference. similarly, in a sentence completion experiment, fuchs and schumacher (2020) investigated whether er, der, and dieser bring arguments with different grammatical or thematic roles in a preceding discourse into focus. to test this, they provided a sentence with two arguments (i.e., proto-agents and proto-patient) for two verb types: (1) active accusative verbs (subject/ proto-agent and direct object/proto-patient) and (2) dative experiencer (i.e., indirect object/proto-agent and subject/proto-patient) as in (2a) and (2b). (2a) example for an active accusative verb jeden morgen hat der pfleger den heimbewohner gekämmt. dabei hat er/der/dieser oft... every morning, thenom (male) nurse combed theacc (male) resident. during this process, he/he.dem /he.dem often... (2b) example for a dative experiencer verb im hafen ist dem segler der urlauber aufgefallen. wenig später hat er/der/dieser dann... at the harbour, thedat (male) sailor noticed thenom (male) tourist. shortly afterwards, he/he.dem/he.dem then.... çokal and von heusinger 50 they observed that in 65% of cases, the personal pronoun er refers to the first-mentioned entity, while both demonstrative pronouns tend to refer to the second-mentioned entity (73% for der and 74% for dieser). it should be noted that in both types of sentences, the first-mentioned entity exhibits a high level of agentivity involving either an agent or experiencer as opposed to a protopatient. in addition, in fuchs and schumacher’s (2020) study, the second mentioned entity came immediately before anaphoric expressions. this suggests that the proximity of the second entity (i.e., recency effect) to the demonstrative pronouns could have an impact. however, the preference of a second-mentioned referent for demonstrative pronouns may not always be observed. for instance, in an online electroencephalogram (eeg) experiment, repp and schumacher (2023) investigated the processing of the d-pronoun der and personal pronouns in naturalistic discourse contexts (specifically during story listening). their analysis of the tschick corpus highlighted the unexpected characteristics of the antecedents of der (such as their occurrence as subjects and agents), suggesting the prominent role of the perspectival center in their corpus. consequently, like personal pronoun (e.g., er) der can also refer to subjects/agent antecedents. repp and schumacher (2023) pointed out that their findings on der dramatically deviate from the observations in previous studies (e.g., object antecedent preference in bosch et al., 2003; 2007; proto-patient antecedents in schumacher et al., 2016). moreover, in the eeg experiment, where participants listened to audiobooks, repp and schumacher (2023) also found that a biphasic n400late positivity pattern at posterior electrodes emerged in response to d-pronouns compared to personal pronouns. due to its relative unexpectedness in the context of stories when compared to a personal pronoun, they interpreted that the observed n400 associated with der was indicative of higher processing costs. on the other hand, in a sentence completion experiment, with a subject-experiencer verb (e.g., ‘bernhard fears the real-estate agent because dieser/der has a bad reputation’), bader et al. (2022) extended earlier studies, finding that both dieser and der referred to the object referent. in contrast, with an object-experiencer verb (e.g., ‘the real-estate agent frightens bernhard because dieser/der makes a fraudulent impression’), a majority of references with dieser and der were to the subject referent. additionally, the two demonstratives predominantly referred to the subject in the case of object-experiencer verbs (hinterwimmer & bosch, 2018; patterson & schumacher, 2021; zifonun et al., 1997). given the research indicating that dieser is conventionally associated with selecting a noun phrase positioned at the end of a clause – especially when the object precedes the subject in the preceding context sentence (patil et al., 2020) – dieser was observed to refer to the subject in the case of object-experiencer verbs (bader et al., 2022). to sum up, previous studies have shown that demonstrative pronouns in german are most often associated with a recently mentioned entity/second-mentioned entity as an antecedent compared to the personal pronoun er. however, a recent eeg study has shown that subjects’ preference for the second-mentioned entity is not always observed for der (repp & schumacher, 2023), and thematic roles are also significant in determining references of der/dieser to the object/subject (the subject/first np references in bader et al., 2022). these findings provide a nuanced understanding of how these two demonstratives are used in different situations, such as verb semantics. 2.4 topic persistence with demonstrative pronouns and demonstrative determiners extending this summary of recent research on demonstrative pronouns to demonstrative determiners in complex demonstratives (e.g., that book), complex demonstratives are distinguished from definite noun phrases (e.g., the book) in their anaphoric use by the following properties: first, they express contrast to other referents of the same type. they also express an antiuniqueness/contrast condition, evidenced by the ungrammaticality of this sun, this pope in a neutral context (ahn, 2019; bisle-müller, 1991; bosch & umbach, 2007; king, 2001; wolter, 2006). second, demonstrative determiners are topic shifters (i.e., they refer to a non-topical antecedent and promote it to the topic of the current clause (bosch & umbach, 2007; diessel, 1999; van german demonstrative pronouns 51 kampen, 2007; zifonun et al., 1997). third, they signal topic persistence, which is understood as a primitive form of givón’s (1983) concept of topic continuity. topic persistence means the cataphoric or forward-looking function (fuchs & schumacher, 2020) of demonstratives is best investigated for indefinite demonstratives (deichsel & von heusinger, 2011, gernsbacher & shroyer, 1989, prince, 1981). these patterns are observed more frequently with complex demonstratives with the pronominal dieser than with definite noun phrases with the prenominal der, which is in most cases understood as being the definite article (e.g., the book). from the observations on demonstrative determiners in complex demonstratives, we assume that some functions are similar for demonstrative determiners and demonstrative pronouns – perhaps in a different distribution. therefore, we assume that the pronominal dieser is used for topic persistence and thus signals that in the subsequent discourse the referent is used as topic. 2.5 contrast function with demonstratives the semantics of german demonstratives also involves a contrast function. contrast function, as discussed by diessel (1999), refers to the capability of demonstratives to distinguish a specific object or person from a group of similar entities (cf. further discussion, ahrenholz, 2007). this is described as “…pointing out one member of a group” (anderson & keenan, 1985, p. 289). bislemüller (1991) also emphasized that the contrast and delimitation functions of demonstrative noun phrases with dieser imply a distinction from other referents, suggesting a sense of “only this one matters, forget about the others.” this distinction goes beyond just spatial functions of demonstratives and signifies a fundamental differentiation. bisle-müller’s (1991) perspective on the meaning of dieser is that it serves to address differentiation from other possible referents. all these accounts describe the demonstrative determiner dieser (dieser np), since a complex demonstrative der np cannot easily be distinguished from a definite noun phrase der np. other researchers have posited that the demonstrative pronoun der is used if there is a contrast in alternation with the non-accented personal pronoun er (bosch, 1988; zifonun et al., 1997). additionally, bader et al. (2022) propose that both dieser and der demonstratives are capable of indicating contrast. while these studies focused on german pronouns, van kampen (2007) also examined how referent preferences can also be modulated by contrast in dutch, german and swedish. van kampen (2007) suggests that demonstratives can only shift a topic referent if that referent holds a contrastive position in the previous sentence. while van kampen (2007) investigated antecedents in the focus position, one could also test accented antecedents or antecedents that are contrastively marked by explicit expressions. however, it remains uncertain how sensitive der and dieser are to topic contrast context. consequently, we investigated their usage in both contrast and non-contrast contexts (see current study & design sections below.). 3 the current study examining demonstratives in such a contrast context will provide new results with respect to how der and dieser are processed in context and how they are integrated into a discourse model. overall, there is a lack of comprehensive characterization for dieser, and no systematic distinction has been established between der and dieser in terms of their interpretive preferences. therefore, the current study aims to answer the un-investigated questions: (1) do we observe any distinction between these two demonstrative pronouns uses? and (2) how does contrast affect the function of der and dieser to find an antecedent and their integration into the discourse model? above, we argued that previous studies leave open questions of how contrast influences real-time processing of der and dieser. we predict that the functions are not equally distributed over the two demonstrative pronouns: (1) dieser signals topic persistence (2) der signals that the antecedent argument is contrastive. in order to test these two hypotheses, we constructed short texts that allow us to (i) neutralize the parameter of distinct grammatical and thematic roles; (ii) manipulate the contrast of one argument; (iii) neutralize immediate recency, and (iv) see whether demonstrative pronouns can çokal and von heusinger 52 access topics and whether they also continue the previous topic (i.e., topic persistence). we, therefore, designed short texts of the following type: (3) s1: i listened to a cellist and a flautist at a munich classical music festival. s2: the cellist, in contrast to the flautist /together with the flautist, played a few wrong notes. s3: a few minutes before the opening concert of the event, dieser/der changed the strings skillfully and tuned his instrument. the first sentence (s1) introduces two indefinite noun phrases in direct object position. in the second sentence (s2), the first indefinite noun phrase introduced in s1 is taken up as a definite noun phrase in subject/topic position. the second noun phrase is taken up as a prepositional phrase, either expressing a contrast (np1 in contrast to np2) or as a comitative phrase (np1 together with np2). on the one hand, the syntactic role of np2 is clearly degraded in comparison to the np1, which functions as the subject/topic. on the other hand, np2 has the same thematic role as it refers to an alternative or comitative expression to np1. this design enables us to investigate the subtle difference between der and dieser keeping the thematic role stable. the third sentence (s3) starts with the demonstrative pronoun in the subject (and topic) position and then predicates a property that is only coherent with np1. in other words, we have an ambiguous pronoun that is disambiguated after the verb phrase. to examine subtle differences between the functions of der and dieser, in experiment 1 (reported below), we used eye-movement recording during reading to examine how der/dieser are processed in german, using contrast (e.g., np1 in contrast to np2) and no-contrast contexts (e.g., np1 together with np2). additionally, we ran a sentence-completion experiment (experiment 2 reported below) to examine participants’ antecedent preferences regarding information structure without time pressure (i.e., der/dieser references to np1 or np2 in contrast and no-contrast contexts). 4 experiment 1 experiment 1 was an eye-tracking study using a 2 × 2 within-subject, within-item design, crossing contrast type (contrast and no-contrast) with anaphora type (dieser and der) as seen in example (4) below. (4a) dieser in contrast context bei einem münchener klassikfestival habe ich einem cellisten und einem flötisten zugehört. der cellist im gegensatz zum flötisten hat ein paar schiefe töne gespielt. einige minuten vor dem eröffnungskonzert der veranstaltung hat dieser die saiten routiniert ausgewechselt und sein instrument gestimmt. (4b) der in contrast context bei einem münchener klassikfestival habe ich einem cellisten und einem flötisten zugehört. der cellist im gegensatz zum flötisten hat ein paar schiefe töne gespielt. einige minuten vor dem eröffnungskonzert der veranstaltung hat der die saiten routiniert ausgewechselt und sein instrument gestimmt. (translations of 4a/4b: at a classical music festival in munich, i listened to a cellist and a flautist at a munich classical music festival. the cellist, in contrast to the flautist, played a few wrong notes. a few minutes before the opening concert of the event, dieser/der changed the strings skillfully and tuned his instrument.) german demonstrative pronouns 53 (4c) dieser in no-contrast context bei einem münchener klassikfestival habe ich einem cellisten und einem flötisten zugehört. der cellist zusammen mit dem flötisten hat ein paar schiefe töne gespielt. einige minuten vor dem eröffnungskonzert der veranstaltung hat dieser die saiten routiniert ausgewechselt und sein instrument gestimmt. (4d) der in no-contrast context bei einem münchener klassikfestival habe ich einem cellisten und einem flötisten zugehört. der cellist zusammen mit dem flötisten hat ein paar schiefe töne gespielt. einige minuten vor dem eröffnungskonzert der veranstaltung hat der die saiten routiniert ausgewechselt und sein instrument gestimmt. (translations of 4c/4d: at a classical music festival in munich, i listened to a cellist and a flautist. the cellist, together with the flautist, played a few wrong notes. a few minutes before the opening concert of the event, dieser/der changed the strings skillfully and tuned his instrument.) 4.1 predictions our first prediction is that dieser is a marker for topic persistence (deichsel & von heusinger 2011, gernsbacher & shroyer, 1989, prince, 1981). if this principle does play a role, then dieser in (4a/4c) would likely lead to less processing difficulties, resulting in lower odd ratios of regressions-out as well as shorter total times in the disambiguation5 and/or spill-over regions. both conditions suggest that we talk more about “the cellist” and therefore allow dieser to be processed easily. our second prediction is that der has a congruent feature of contrast and therefore it can be anaphorically linked to a contrasted item in the previous context. if that is the case, in (4b) der would be preferred to der in (4d) because in (4b) a contrast with respect to a competitive referent is made (i.e., the cellist in contrast to the flautist). consequently, we expected to observe processing difficulty (i.e., high proportions of regressions-out and longer total times) in (4d) immediately after der, (the disambiguation region), in comparison to der in (4b). given this experimental design, the presence or absence of topic contrast in the previous context would lead to an interaction (all other things being equal) between contrast type and anaphora type, due to the prediction of a difference between (4b) and (4d) and the absence of a contrast effect in (4a) and (4c). it should be noted that this prediction assumes that der and dieser have a difference in meaning and function. 4.2 method 4.2.1 participants fifty-two paid, native german-speakers aged 22-24 (46 females and 6 males) from the university of cologne participated in the experiment. they were not exposed to another language before they were three years old. all were unaware of the study’s purpose. the experiment involving human participants was reviewed and approved by the ethics board of the german linguistic society. the participants provided written informed consent to participate in this study. 4.2.2 apparatus we used an eyelink 1000 eye-tracker (sr research ltd, ottawa, canada) in tower-mounted mode, with a chin rest to stabilize each participant’s head. 5 in the disambiguation region, we disambiguated the antecedents of der and dieser by employing a referential expression closely related to np1 (strings related to cellist). çokal and von heusinger 54 4.2.3 materials forty items were created based on example 4 above. each item appeared in four conditions, which crossed contrast type (contrast vs. no-contrast) with anaphora type (der vs. dieser). the explicit topic contrast was manipulated using im gegensatz zu (‘in contrast to’) in the contrast condition and zusammen mit (‘together with’) in the no-contrast condition. anaphora type was manipulated by including either der or dieser. we disambiguated the antecedents of der and dieser by using referential expressions/noun phrases (e.g., strings) coming after them, thereby disambiguating the demonstrative pronouns (e.g., der/dieser changed the strings – referring to a cellist in example 4). throughout the experiment, der and dieser were disambiguated to np1, i.e. the subject and topic of the antecedent sentence (i.e. cellist in the example above. while both referents (i.e., cellist and flautist) were congruent in gender and number with the pronoun, reference to the subject/topic was ensured via the plausibility of the critical sentence (e.g., in example 4 it is only plausible that a cellist would check the strings, not the flautist because a flute does not have strings). references to the subject/topic was done to avoid the recency/last-mentioned entity as a confounding factor for demonstrative reference (bosch & hinterwimmer, 2016; bosch & umbach, 2007; kaiser, 2011; wilson, 2009; zifonun et al., 1997). in addition, to reduce the recency effect (i.e., the second np as an alternative antecedent), temporal phrases were also used after the critical sentence (e.g., einige minuten vor dem eröffnungskonzert der veranstaltung ‘a few minutes before the opening concert of the event’) to put a distance between der/dieser and their possible antecedent options. except for our manipulations (the cellist & the flautist), there were no alternative antecedents for der or dieser in the previous context. in the spillover region, the length of adverb (routiniert ‘skillfully’) in each item was matching across four conditions. the forty stimuli were distributed into four lists, following a latin square procedure. in each list, every item appeared in only one condition, and each condition appeared an equal number of times. each list was assigned to 10 participants. additionally, there were 73 fillers with plural pronouns but not personal or demonstrative pronouns in the singular. since our critical items consisted of first np references with der and dieser, the plural pronouns in our filler items referred to the second-mentioned np or its referent, resulting in ambiguity/underspecification (see the filler item below.). in addition, there were three practice items, all of which were similar in length to the experimental sentences. here is an example of a filler item: (5) es war in einem operationssaal eines krankenhauses in köln lindenthal. die ärzte standen am operationstisch. die anästhesisten standen daneben. es war ein typischer arbeitsreicher und produktiver tag. sie checkten sorgfältig ihre aufgaben und erledigten die arbeit. (it was in an operating room of a hospital in cologne lindenthal. the doctors were standing at the operating table. the anesthesiologists stood by. it was a typical busy and productive day. they carefully checked their task and got the job done.) 4.2.4 pre-testing items we pre-tested our stimuli using qualtrics with 120 native speakers of german. we provided the initial text and asked them to choose “who would do the action”. a sample pre-test item is as follows: (6) es gab einen cellisten und einen flötisten. der hat die saiten routiniert ausgewechselt. wer hat die saiten ausgewechselt? (who changed the strings?) (a) cellist (b) flötist german demonstrative pronouns 55 our results showed that 75% of referents were selected as np1, whereas 25% of cases were selected as np2. using qualtrics, we also ran an acceptability judgment task to check that the items and conditions did not differ in their acceptability. the participants did not take part in the previous pre-testing session (n = 46). they rated 40 experimental stimuli and 40 filler items from 1 to 5. we ran a linear mixed-effects model including random intercept for participants and items (intercept: β = 1.028, se = .144, t = 71.084). while the interaction between anaphora type and contrast type was not significant (β = 0.038, se = .054, t = .712), there were main effects of anaphora and contrast types, anaphora: β = 0.064, se = 0.27, t = 2.369; contrast: β = -0.092, se = 0.27, t = -3.337. overall, the use of dieser was significantly preferred to der, der with explicit contrast context with m = 3.22, se = 1.40; dieser with explicit contrast m = 3.36, se = 1.39; der with no-contrast context m = 2.91, se = 1.38; dieser with no-contrast m = 3.33, se = 1.40. in addition, to a significant degree participants had less of a preference for the use of der with a no-contrast context to der with a contrast context: β = -0.325, se = 0.111, t = -2.969. the main effect of contrast indicates that participants less preferred contrast context than no-contrast context. overall, our results indicate that sentences are not considered ungrammatical, receiving an average score of 3 on a scale of 1-5. this was expected given that the demonstrative dieser is often preferred in formal and /or written discourse, while the demonstrative der is preferred in informal or oral discourse (patil et al., 2023, for further discussion). moreover, participants might have learned that the use of der is impolite for human referents. both assumptions would predict a certain preference of dieser over der in our written test items. 4.2.5 procedures we presented 116 texts in times new roman 18 font, in a fixed random order, with no experimental items adjacent to each other. the texts were presented on six or seven written lines, with each line containing between 76 and 85 characters. the critical regions including der/dieser and their referents (i.e., die saiten/the strings) always appeared near the middle of a line and were not the last sentence on the screen. to familiarize participants with the experimental procedure, the experiment began with three fillers. while viewing was binocular, only the right eye was tracked. the items appeared on a 19” monitor, approximately 70 cm away from the participants’ eyes. in order for the experimenter to check the calibration of each participant, the participant fixated on a black square before each item, indicating the position of the first character of the text. the black square was automatically replaced with text once a stable fixation was detected. after reading each item, the participant pressed a button to indicate the end of the sentence. for 37% of the items, a comprehension question then appeared, which the participant answered by pressing a button on the left or right side of the button box. the comprehension questions never probed an anaphora and the antecedent (‘cellist’; or ‘flautist’). the eye-tracking experiment took 50 minutes, and included informing participants about consent forms, instructions, and three short breaks. 4.2.6 data analysis texts were divided into three regions (see table 1). below, we report data for the following regions: anaphora, disambiguation, and spillover. it should be noted there is a length difference in the anaphora region between der (i.e., 3 characters) and dieser (i.e., 6 characters) rendering this main effect in the anaphora region uninterpretable. additionally, due to the skipping rate of functional words in previous studies (rayner, 2011), we have typically not found early effects of syntactic or referential processing difficulty in the anaphora region. we predicted that we would observe an interaction between two factors (i.e., anaphora type and type of contrast) in the disambiguation and spill-over regions, which immediately follow the anaphora region. these regions were matched for length in all conditions. çokal and von heusinger 56 anaphora disambiguation spillover der/dieser der/dieser die saiten the strings routiniert skillfully table 1. regions in experiment 1 fixations of less than 80, or more than 1200 ms, were excluded from analysis. the percentages of data points for total time in each region that were excluded are as follows: 21% in the pronoun region, 8% in the disambiguation region, and 5% in the spillover region. 21% of the excluded data points in the pronoun region involve instances where the pronoun was skipped. all participants correctly answered at least 90% of comprehension questions. von der malsburg and angele’s (2017) study showed that conducting multiple comparisons (i.e., rate of false positives when dependent measures in eye movements with multiple regions are tested) increased the probability of incorrectly rejecting the null hypothesis (type 1 error). since reporting a large number of eyemovement measures might result in false positives due to family-wise errors (von der malsburg & angele, 2017), we selected only two eye-movement measures: regression out and total time. we selected regressions-out, which is the proportion of trials where readers looked back from the region to an earlier piece of the text between the time when the region was first entered from the left to the time when the region was first exited to the right. this measure indicates that lexical semantic information is processed during regressions-out and used in recovery from processing difficulty from the previous text where the current word/text does not meet the expectation of readers (cf. further discussion on regressions in reading sturt & kwon, 2018). for completeness, we also report total time (i.e., the sum of all fixations in the region) as a general measure of processing, even though this does not provide information about initial processing. in cases where the region received no fixations (total time), the trial was treated as missing data and excluded from analysis. in order to keep false positives in check (type i error), we conducted the bonferroni correction (von der malsburg & angele, 2017). there were two measures on three regions, then there were 6 analyses. to remove familywise error, we divided our critical value by 6 (i.e., p < 0.05/6). after bonferroni adjustment, our significance threshold value was p < .008. then, since we assumed |t| had to be greater than 2 for p < .05, we determined the new |t| value 2.65 (based on z distribution). a given co-efficient was judged to be significant at α = .008 if the absolute t-value exceeded 2.65. we used the following packages: lme4 to run logistic mixed effects regressions models for the analysis of regressions-out. we used sjplot to calculate odds ratios, random effects; library (emmeans) to calculate standard errors and confidence intervals; ggplot2 to graph estimate proportions of each condition. the analysis was based on whether or not participants had regressions-out in the region. therefore, trials without regressions-out were retained in the analysis. trials that had no regressions-out were included and coded as 0. we used logistic-mixed effect regression (glmer). we contrast-coded our fixed effects (i.e., anaphora type and contrast type) and centered the coding around zero (i.e., using -0.5 and 0.5). random slope parameters corresponding to the two experimental factors and their interactions were included in the maximal model for both participants and items (barr, levy, scheepers, & tily, 2013; bates et al., 2015). below we report odd ratios for the analysis of regressions-out. to aid convergence, and to avoid spurious overestimates of correlations, random correlation parameters were excluded from the model. in anaphora and disambiguation regions for regressions-out, the model failed to converge, and random slope parameters with the least variance were removed until convergence was achieved. to decide which factors needed to be removed first to achieve convergence, we assessed the variability at german demonstrative pronouns 57 each level in our data. it should be noted that when there is little variability between the levels of a particular random effect, it may not be necessary to include it in our model. the following models converged for the regressions-out: • anaphora region: (anaphora type * contrast type +1||subject) + (anaphora type +1|item) + anaphora type * contrast type • disambiguation region: (anaphora type +1|subject) + (1|item) + anaphora type * contrast type • spillover region: (1|subject) + (1|item) + anaphora type * contrast type for each region and total time, linear mixed effects regression (lmer) models using lme4 r packages were constructed, incorporating all fixed effects and interactions in a single step. an additional package (plyr) was used to compute mean values. factor labels were transformed into numerical values and centered prior to analysis. we performed analyses on log-transformed reading times, and all analyses reported below incorporated crossed random intercepts for participants and items. the model did not converge in the anaphora and disambiguation regions for total time. to achieve convergence, random slope parameters with the lowest variance were eliminated. converged models for each region are as follow: • anaphora region: (1|subject) + (1|item) + anaphora type * contrast type • disambiguation region: (contrast type +1|subject) + (1|item) + anaphora type * contrast type • spillover region: (1|subject) + (1|item) + anaphora type * contrast type after running the above models, for any significant interactions and main effects, we examined the effect of exposure to test items. specifically, we explored the impact of reading first np references with der and dieser during the experiment and whether participants adapted the referential patterns in the experiment (see çokal & ferreira, 2015 for normalization of pronoun errors during an eye-tracking reading experiment; johnson & arnold, 2023 for participants’ adaptation to experimental manipulation/referential patterns within an experiment). to account for this, we added ‘trial order’ as a covariant to our models. considering the possible effects of nesting, we included participants and items as random intercepts. random slope was specified for the variable ‘trial order’ within the grouping variable ‘subject’. this random slope allowed the effect of ‘order’ on the outcome variable to vary across different subjects. each subject may have their own unique relationship between the order of items and their responses. including random slopes allows the model to capture this variability in individual subject’s responses to the variable ‘order’. while regressions-out are reported with odds ratios, standard errors, and p-values, the results for total times also include coefficients, standard errors, and t-values for each fixed effect and interaction (see baayen et al., 2008 for absolute t-value in linear-mixed effect models). data, scripts, and stimuli for all experiments are available at https://osf.io/ubxkr/?view_only=01bc62fb4fd0460f9e7831b0eedfbe4e. 4.3 results and discussion in this section, we present our results region by region. 4.3.1 anaphora region regressions-out did not show main effects of anaphora and contrast types (see tables 2 & 3, figure 1). total time showed a main effect of anaphora type, with longer reading times for dieser than der but this effect was not significant since the absolute t-value did not exceed 2.65 (dieser: m = 297, çokal and von heusinger 58 se = 11, der: m = 272, se = 14). this effect is likely due to the length differences between dieser and der, as discussed above. the interaction between anaphora and context in total times was not significant (see tables 2 & 3 below). regions/parameters regressions-out total time anaphora odds ratio se p-value β se t intercept 0.17 (0.12-0.23) 0.03 .001* 5.70 0.039 142 contrast 0.97 (0.70-1.35) 0.16 .873 -0.023 0.023 -1.019 anaphora 0.77 (0.541.11) 0.14 .167 -0.052 0.023 -2.249 contrast* anaphora 0.101 (0.53-1.92) 0.33 .976 0.011 0.046 .244 disambiguation intercept 0.22 (0.15-0.30) 0.04 .001* 6.065 0.067 89.791 contrast 0.78 (0.62-0.99) 0.09 .043 -0.027 0.026 -1.025 anaphora 0.32 (0.25-0.41) 0.04 .001* -0.054 0.021 -2.548 contrast *anaphora 1.13 (0.171-182) 0.27 .606 0.026 0.042 .617 spillover intercept 0.12 0.01 .001* 5.755 0.041 138 contrast 1.49 (1.12-1.97) 0.21 .006* 0.022 0.027 .818 anaphora 0.81 (0.61-1.07) 0.12 .136 0.001 0.019 .050 contrast* anaphora 1.37 (0.78-2.40) 0.39 .279 -0.056 0.039 -1.441 table 2. results of mixed-effects analysis for regressions-out and total time for experiment 1. regressions-out are reported with odds ratios, standard errors, and p-values, the results for total times also include coefficients, standard errors, and t-values for each fixed effect and interaction. after bonferroni adjustment, statistically significant affects are indicated with asterisk (*). the effects are considered significant when the absolute t-value is greater than 2.65. german demonstrative pronouns 59 pronoun disambiguation spillover regressions-out m (se) m (se) m (se) der, contrast .209 (.034) .288 (.033) .144 (.020) dieser, contrast .177 (.022) .143 (.017) .132 (.018) der, no-contrast .208 (.035) .336 (.033) .113 (.014) dieser, no-contrast .179 (.028) .158 (.024) .080 (.013) total time der, contrast 252 (17) 531 (29) 355 (11) dieser, contrast 293 (15) 511 (25) 340 (12) der, no-contrast 292 (21) 555 (29) 350 (13) dieser, no-contrast 302 (16) 509 (24) 356 (15) table 3. means (and standard errors) for regressions-out and total time for experiment 1 4.3.2 disambiguation region regressions-out showed a main effect of anaphora but not a main effect of contrast types or a twoway interaction between anaphora type and contrast context (see tables 2 & 3, figure 1) 6. the odds ratios with der were higher than those with dieser, which indicates that irrespective of contrast type (i.e., contrast or no-contrast), conditions with der led to more processing difficulties compared to dieser (see figure 1 & table 2). in addition, irrespective of anaphora types, no-contrast context led to more regressions-out than the contrast context. however, after bonferroni correction, contrast-effect was not significant. our results show that references to a topic with dieser led to less processing difficulties than for der, irrespective of contrast type pairwise comparison of dieser and der with contrast context (β = -1.028, se = .173, z = -5.921, pairwise comparisons of dieser and der with no-contrast context: β = -1.284, se = .237, z = -5.408). 6 we ran right-bound reading time (re-reading time as a late measure) and observed the same patterns in regressions-out. (see appendix a). while anaphora and spillover regions did not show any main effects or a significant interaction between the variables, in the disambiguation region, we observed a main effect of anaphora β = -.0046, se = .019, t = 2.326; but no main effect of contrast and an interaction between the two variables (contrast: β = -.018, se = .019, t = .944; .anaphora * contrast: contrast: β = -.001, se = .039, t = -.015). compared to dieser, der, in the disambiguation region, led to longer right bound re-reading times for der irrespective of contrast type contexts. readers’ initial and late processing strategies for der and dieser did not change. çokal and von heusinger 60 figure 1. estimated proportion of regressions-out across regions for experiment 1. error bars show the 95% confidence intervals. german demonstrative pronouns 61 figure 2. estimated total times (ms) across regions for experiment 1. error bars show the 95% confidence intervals. in the disambiguation region, total time did not show a significant interaction between the two variables (see tables 2 & 3 and figure 2). there was a main effect of anaphora, with longer total times with der than dieser but this effect was not significant after the bonferroni and t-value adjustments. there were no significant main effect of trial order and an interaction between the two variables and trial order (regression-out, β = .011, se = .006, z = 1.794; total-time: β = .002, se = .002, z = -.856). 4.3.3 spillover region regressions-out showed a main effect of contrast context but no main effect of anaphora type or an interaction between the factors (see tables 2 & 3, figure 1). compared to the no-contrast context, the contrast context led to more regressions-out, showing that contrast contexts were difficult to process irrespective of anaphora type in the spillover region. regressions-out and total time did not show a two-way interaction (i.e., anaphora * context) or main effects of anaphora type (see tables 2 & 3, figure 2). in addition, trial order as a co-variant did not affect the results of regressions-out and total-time (regression-out, β = .001, se = .007, z = .046; total-time: β = .001, se = .002, z = .381). çokal and von heusinger 62 4.4 overall summary in the disambiguation region, the main effect of anaphora in regressions-out indicates that overall, readers had less processing difficulties with dieser. in the initial processing, regardless of contrast type participants preferred dieser over der. the less processing difficulty observed for dieser might be attributed to its topic persistence function in both contrast and no-contrast contexts. this finding appears to support our initial prediction that dieser serves as a marker for topic persistence. our second prediction was that in the disambiguation region, der in the contrast context would lead to less processing difficulty than der in the no-contrast context, due to the presence of a contrastive antecedent. this assumption is not supported by the findings in the disambiguation region of regressions-out. we observed a pattern where der with the contrast context resulted in lower regressions-out than der with the no-contrast context, although this difference was not statistically significant. regardless of anaphora type, processing a contrast context resulted in higher regressions-out compared to a no-contrast context in the spillover region. initially, due to the presence of two alternatives, readers in the disambiguation region might find no-contrast contexts demanding. while these alternatives are bundled with the use of such phrases as ‘together with’, in the disambiguation region they need to ‘unbundle’ them. after accomplishing this, in the spillover region readers may find no-contrast contexts less taxing than contrast contexts, since contrastive information in online processing, emphasizing a particular element or concept in a sentence or discourse, creates a mental ‘placeholder’ for related alternatives (repp & spalek, 2021). this placeholder can be connected to likely candidates or alternatives presented after the initial focus. our results indicate that these alternatives can be generated or considered as the discourse unfolds. our findings suggest that the effect of contrast contexts in online discourse processing needs to be further tested. in previous studies with offline methods, where participants read two-sentence stories on amazon mechanical turk and answered comprehension questions about the reference of referential expressions, readers adapted referential patterns within the study after exposure to certain structures in the stimuli (johnson & arnold, 2023). similarly, in an online eye-tracking reading study, native speakers of english who encountered pronoun errors (i.e., gender mismatching) normalized the input and resolved the pronoun ambiguity using available relevant information (çokal & ferreira, 2015). however, their sensitivity to errors was weaker when the antecedent of a pronoun error was embedded in the density of information and declined after certain exposure to errors (çokal & ferreira, 2015). in the current study, we did not observe any trial order effect (i.e. the possibility of learning effects) on our results. it should be noted that since the effect of trial order on the processing of der and dieser was not our main goal in the current study, further studies need to be conducted on this topic. while experiment 1 focused on participants’ online processing of der and dieser in contrast and no-contrast contexts, in experiment 2 we explored participants’ offline antecedent preferences for these anaphoric expressions. in this way, we were able to examine which individuals (np1) and (np2) were selected as antecedents using these expressions. experiment 2 also provided further information on whether participants preferred topic references (np1) with der and dieser. 5 experiment 2 experiment 2 tested participants’ antecedent preferences using a sentence completion method approved by the ethics of research committee at university of cologne. 5.1 method in this section, we will describe our participants, materials, and data analysis. german demonstrative pronouns 63 5.1.1 participants experiment participants (n =32) were native german-speakers aged 21-23 (24 females and 8 males) from cologne university, and they were compensated for their participation. all were unaware of the study’s purpose, and none had participated in experiment 1. to recruit participants, we posted fliers to the faculty of philosophy at university of cologne. our main selection criteria were being a native speaker of german and not being bilingual (defined as not growing up in an environment where more than one language was spoken). 5.1.2 materials we used the same sentences as in experiment 1, but we made a slight change to avoid the demonstrative determiner uses of der and dieser in their completions (e.g., der cellist, der flötist). while in experiment 1, the auxiliary hat ‘has’ came before der and dieser (i.e., hat der/dieser ...), in experiment 2, hat (i.e., der/dieser hat) followed der and dieser (please see below). this forced a demonstrative pronoun usage and disallowed a demonstrative determiner usage. each participant was provided with an initial context and asked to provide a completion answer for the sentence fragment ending with der hat or dieser hat in a manner consistent with the previous text. in each condition, participants could refer to either the first np (np1) (e.g., the cellist) or the second np (np2) (e.g., the flautist). one experimental item with its conditions is given below: (7a) /(7b) dieser/der in the contrast context bei einem münchener klassikfestival habe ich einem cellisten und einem flötisten zugehört. der cellist im gegensatz zum flötisten hat ein paar schiefe töne gespielt. es waren einige minuten vor dem eröffnungskonzert. dieser / der hat ... (i listened to a cellist and a flautist at a munich classical music festival. the cellist, in contrast to the flautist, played a few wrong notes. it was a few minutes before the opening concert of the event. dieser / der has …) (7c) /(7d) dieser/der in the no-contrast context bei einem münchener klassikfestival habe ich einem cellisten und einem flötisten zugehört. der cellist zusammen mit dem flötisten hat ein paar schiefe töne gespielt. es waren einige minuten vor dem eröffnungskonzert. dieser / der hat ... (at a classical music festival in munich, i listened to a cellist and a flautist. the cellist, together with the flautist, played a few wrong notes. it was a few minutes before the opening concert of the event. dieser / der has …) there were 40 experimental and 60 filler sentences. there were four experimental conditions as per experiment 1. four lists were constructed, using latin square counterbalancing. in each list, each sentence appeared in only one condition, with an equal number of items from each condition. sentences were presented in a word document in fixed random order. all filler sentences had referential expressions (see a sample filler with a personal pronoun below.). (8) es war in einem operationssaal eines krankenhauses in köln lindenthal. die ärzte standen am operationstisch. die anästhesisten standen daneben. es war ein typischer arbeitsreicher und produktiver tag. sie hatten … (it was in an operating room of a hospital in cologne lindenthal. the doctors were standing at the operating table. the anesthesiologists stood by. it was a typical busy and productive day. they had …) 5.1.3 procedure each participant made an appointment for a sentence completion experiment session. we conducted a two-block data collection process, which has been previously used in studies of plural and singular anaphors (çokal et al., 2023; koh & clifton, 2002). in the first-block, participants çokal and von heusinger 64 finished all sentence completions, and we saved their completions to a different folder. in the second-block, immediately after saving their document, participants were asked to review their completions and underline to what the referential expressions (e.g., der, dieser) referred. since filler sentences included other referential expressions (e.g., sie ‘she’, ‘they’, das ‘that’, er ‘he’), participants also underlined their referents. no feedback was given to participants. the whole experiment took 1 hour 15 minutes. 5.1.4 data analysis we used the following continuation codings and samples for der and dieser. the data analysis below was based on participants’ underlined interpretations. if der or dieser referred to the protagonist introduced first, then its referent was coded as the first np (np1), as in (9) and (10) below: (9) in meiner stammkneipe habe ich mit einem koch und einem barmixer geredet. der koch im gegensatz zu dem barmixer war durchgehend sehr freundlich. es gab ein unterhaltsames gespräch. der hat den anderen viele tipps und ratschläge zum kochen gegeben. (in my regular pub, i had a conversation with a chef and a bartender. the chef, unlike the bartender, was consistently very friendly. it was an entertaining conversation. der gave the bartender many tips and pieces of advice about cooking.) (10) auf einem jazzfestival in dem berühmtesten kölner konzerthaus habe ich einen sänger und einen saxophonisten gemeinsam auftreten sehen. der sänger zusammen mit dem saxophonisten hat stehende ovationen bekommen. es herrschte eine angenehme stimmung in der schönen kölner philharmonie. dieser hat den saxophonisten umarmt, weil er so glücklich über die tolle performance war. (at a jazz festival in the most famous concert hall in cologne, i saw a singer and a saxophonist perform together. the singer, along with the saxophonist, received a standing ovation. there was a pleasant atmosphere in the beautiful cologne philharmonic hall. dieser hugged the saxophonist because he was so delighted with the fantastic performance.) if der or dieser referred to the protagonist introduced as the second noun, then its referent was coded as the second np (np2), as in (11) below: (11) in meiner stammkneipe habe ich mit einem koch und einem barmixer geredet. der koch zusammen mit dem barmixer war durchgehend sehr freundlich. es gab ein unterhaltsames gespräch. dieser hat schon mit 16 seine ersten cocktails gemixt. (in my regular pub, i talked with a chef and a bartender. the chef, along with the bartender, was consistently very friendly. there was an entertaining conversation. dieser started mixing cocktails at the age of 16.) ungrammatical or incoherent sentence completions were excluded from the logistic mixed effects regression analysis. for instance, while der/dieser is used for masculine nouns in nominative singular, they can also be used for feminine nouns in the genitive or dative singular form (see appendix c for other cases). 5.2 results we ran logistic mixed effects regressions, taking pronouns (der vs. dieser) and type of contrast (contrast vs. no-contrast) as the fixed effects, and including crossed random intercepts and slopes for participants and items. when participants had ungrammatical or incoherent sentence completions, these cases were coded as ‘other cases’ using 2 and excluded from the main analysis (as illustrated above). the percentage of other cases was 3.59 (n = 23 cases out of 1280. in the logistic mixed-effects regression, we coded references to the first protagonist (np1/ der cellist) as 1 and to the second protagonist (np2/ der flötist) as 0. we contrast-coded our fixed effects (i.e., german demonstrative pronouns 65 anaphora type and contrast type). however, since the full model did not converge, we reduced the fixed random effect for items and ran the following model instead: response ~ (1 | subject) + (1|item) + contrast type * anaphora type the intercept was significant (β = 2.27, se = .258, z = 8.18) and the analysis revealed main effects of anaphora and contrast type (anaphora: β = -0.366, se = .180, z = -2.039, p = .041; or = 0.67, se = .12, p = .024; contrast type: β = -0.406, se = .180, z = -2.249; p = .024; or = 0.69, se = .14, p = .041). however, the interaction between these two factors was not significant (β = 0.157, se = .355, z = 0.441; or = 1.17, se = .42, p = .659; see appendix b, figure 4). most references were resolved to np1. additionally, there were slightly fewer np1 references in contrast than nocontrast conditions: der-contrast: or = 0.905, se = .025; dieser-contrast: or = 0.873, se = .032; der-no-contrast: or = 0.937, se = .018; dieser-no-contrast: or = 0.902, se = 0.026 (see appendix b and figure 4 below). figure 4. estimated proportions of np1 references out of all np references in topic contrast and no-contrast contexts with der and dieser. note: error bars show the 95% confidence intervals. overall, our results suggest that participants preferred np1 references with der and dieser but this preference slightly changed in the contrast context. 6 general discussion in this paper, we investigated whether a contrast context (i.e., np1 in contrast to np2) or no-contrast context (i.e., np1 together with np2) affects online processing of der and dieser regarding integration to a discourse model and antecedent preference. in addition, we also explored whether der and dieser are used to refer to np1 or np2 in these contexts. we predicted that the functions of der and dieser are not equally distributed. our first assumption was that dieser signals topic persistence. our eye-tracking reading experiment (experiment 1) showed that in both contrast and no-contrast contexts the processing of dieser was easier than that of der, with lower regressions-out in the disambiguation region. this suggests that dieser led to less processing difficulty compared to der when referring to the first np in both contrast and no-contrast contexts. in a sentence completion task (experiment 2), participants mostly 0.00 0.25 0.50 0.75 1.00 der contrast der.no contrast dieser.contrast dieser.no contrast pr op or tio n of n p1 re fe re nc es contrast no contrast çokal and von heusinger 66 preferred continuations in which dieser and der refer to np1 rather than np2 in both contrast and no-contrast contexts (e.g., the cellist (np1) in contrast to/together with the flautist (np2). der/dieser hat). these results support our first prediction on the topic persistence function of dieser. a related question is, ‘why do we not observe such online processing difficulty with dieser?’ one reason may be that dieser serves a ‘singling out’ function, making it a more suitable anaphor for ‘unbundling’ compared to der (ahrenholz, 2007; anderson & keenan, 1985: bislemüller, 1991). our findings from the eye-tracking reading experiment show longer reading times for der compared to dieser (i.e., regressions-out in the disambiguation region). in a sentence completion task (experiment 2), participants mostly preferred continuations in which der and dieser refer to np1 rather than np2 in both contrast and no-contrast contexts (e.g., the cellist (np1) in contrast to/together with the flautist (np2). der/dieser hat). our sentence completion experiment results suggest that these reading difficulties might not be attributed to topic references with der and dieser (i.e., ‘the cellist’ but not ‘the flautist’ and ‘the cellist’ together with ‘the flautist’ played a few wrong notes – ‘der / dieser fixed the strings’). one possible reason behind longer reading times for der might be that readers’ prior expectations in the disambiguation region lean towards the demonstrative determiner use of der (der + cellist) instead of a pronominal use. to investigate this further, we conducted a brief corpus analysis to examine the distribution of pronouns der and dieser as well as determiners der and dieser. to accomplish this, we retrieved instances from the cosmas tagged t2 corpus (https://www2.ids-mannheim.de/cosmas2/projekt/referenz/archive.html). our results indicate that the percentage of determiner use of der is 95%, while the pronominal use is 5% (cases: der-pronoun (n) = 13.352; der-determiner (n) = 253.801). in contrast, determiner use of dieser accounts for 50% of cases, with demonstrative pronoun use also making up 50% (cases: dieser-pronoun (n) = 9.402; der-determiner (n = 9.525)). what is interesting is that, even though the assumption is less usage of der in written discourse (see ahrenholz, 2007; patil et al., 2020; weinert, 2007), our small-scale corpus analysis indicates that people more frequently use der than dieser in written (newspaper) discourse. while further research is needed, it is not the focus of the current study. the combination of our results and the corpus analysis suggests readers tend to prefer the d-pronoun use of der (i.e., determiner), while the d-pronoun use of dieser does not significantly deviate from their expectations. our second assumption was that der would be preferred in the contrast context rather than the no-contrast context because it requires a topic contrast. experiment 1 (an eye-tracking reading study) did not support this prediction. we observed fewer regressions-out in the contrast context with der, compared to the no-contrast context with der or a similar pattern for der in the same region of total times. however, the interaction between anaphora and contrast type was not statistically significant. this finding appears to be inconsistent with the theoretical claims made by bosch (1988) and zifonun et al. (1997) for german, as well as van kampen (2007) for dutch and swedish, who suggested that der could signal a contrastive antecedent. there might be two reasons for the inconsistency. the first reason is methodological differences. the previous studies’ claims were based on observations. therefore, there were no controlled empirical data on der in the contrast context. the second reason is the limitation of the reading study, as the manipulation of the accented antecedent (i.e., the use of prosody) was not possible with our current experimental setting 6.1 implications of topic reference/np1 in the current study our sentence completion experiment results suggest that a topic reference is preferred not only with dieser but also with der. these findings can be aligned with those of other studies that have explored demonstrative determiners cross-linguistically, including works by ahrenholz (2007), anderson and keenan (1985), bisle-müller (1991), and diessel (1999). as mentioned earlier, the literature posits three functions of demonstrative determiners in complex demonstratives: (a) contrast/nonuniqueness, (b) topic shift, and (c) topic persistence. these functions can be partially extended to https://www2.ids-mannheim.de/cosmas2/projekt/referenz/archive.html german demonstrative pronouns 67 the function of demonstrative pronouns. in german, dieser seems to specifically specialize in topic persistence (both online and offline study) (see fuchs & schumacher, 2020, for evidence supporting the topic-persistent function of dieser). in addition, in experiment 2, writers referred to the first mentioned entity, which is in the subject/topic position, but not the second np, which is in a prepositional phrase (e.g., cellist/np1 in contrast to flautist/np2 or cellist/np1 together with flautist /np2). our findings extend previous studies that have shown demonstratives’ preference for a non-topic antecedent, object, lastmentioned entity, or recent entity (bader & portele, 2019; bosch, 2003; fuchs & schumacher, 2020; hinterwimmer & bosch, 2018; kaiser, 2011; patil et al., 2020; zifonun et al., 1997). the recency effect observed in previous studies may be attributed to either (a) the short distance between demonstratives and their recent antecedents/np2, since demonstratives immediately followed np2, or (b) the thematic roles (agent vs. patient) explaining the recent antecedents of these two demonstratives. in our experimental design, we maintained a consistent distance between np1 and np2 in relation to the demonstratives. in addition, the thematic roles of antecedents remained the same across conditions. both demonstratives (particularly dieser in both online and offline processing) can also serve as a topic reference in sentence completion tasks (i.e., np1 references instead of np2 references). these findings extend repp and schumacher’s (2023) corpus analysis of the informal language in the crime novel tschick, where they observed that der can refer to antecedents in the subject position and antecedent with the thematic role of agent. in essence, repp and schumacher (2023) found that both der and the personal pronoun er tend to prefer subject and agent references (see figure 2 in repp & schumacher, 2023). it should be noted that in the eeg study, repp and schumacher (2023) discovered that during an experiment in which participants listened to stories from audiobooks, the demonstrative der elicited a more pronounced n400 response compared to personal pronouns. consequently, repp and schumacher (2023) argue that der, when contrasted with personal pronouns, is perceived as a less expected and more distinct form of reference. in their study, due to the design of the eeg experiment, they were unable to distinguish between the determiner and pronoun uses of der in the stories. however, our study specifically hypothesizes that dieser would cause less processing difficulty in reading compared to der. our findings point to subtle distinctions within the d-pronoun and determiner systems of demonstratives. further research is needed to differentiate demonstrative forms in online processing, encompassing both eeg and eye-tracking studies. a relevant direction for such further research is to investigate the online processing and production of both determiner and pronoun forms of these demonstratives to gain a broader understanding of when and how they are disambiguated. additionally, there is a need to move beyond the recency and spatial approaches to information structures and semantics. a limitation of the current study is the absence of eye movement and reading time data for the determiner use of these two demonstratives when referring to a protagonist in both contrast and nocontrast contexts. our corpus analysis indicates there would be less of a penalty if the determiner der was used (note that this is the definite article in most instances). another limitation is that our design is clearly biased towards np1 to test topic persistency. however, our offline sentence completion experiment results suggest participants predominantly had np1 references with der and dieser, and fewer np2 references occurred in a contrast context compared to a no-contrast context. in conclusion, our results provide novel insights regarding demonstrative pronoun resolution in online reading and offline sentence production tasks. the following findings also contribute to previous research: (i) dieser was preferred when referring to the topic of a sentence in both contrast and no-contrast contexts, (ii) der can be used for np1 references as well. we feel our findings move the state-of-the-art in psychological and theoretical modelling of anaphora resolution forward beyond the recency and/or thematic role approach to demonstrative pronoun in cases requiring such inferences. çokal and von heusinger 68 acknowledgements this research was funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) – project-id 281511265 – sfb 1252 prominence in language – in the project c04 “conceptual and referential activation in discourse” at the university of cologne, department of german language and literature i, linguistics. we would like to express our gratitude to robert voigt and nagihan gökben konuk for their invaluable assistance in crafting the stimuli and conducting piloting sessions with native speakers of german. german demonstrative pronouns 69 appendix a: right-bound rereading across regions for experiment 1 anaphora disambiguation spillover β se t β se t β se t intercept 5.529 .028. 194 5.88 .057. 102 5.642 .034 162 contrast -0.05 .020 -0.267 -0.018 .019 -0.944 0.013 .016 .803 anaphora -0.021 .021. 1.044 -0.046 .019. -2.326 0.013 .016 -.508 contrast * anaphora -0.030 .041. 0.720 -0.001 .039. -0.015 0.008 .033 0.268 table a1. right-bound rereading across regions for experiment 1. pronoun disambiguation spillover right-bound rereading m (se) m (se) m (se) der, contrast 144 (11) 415 (24) 296 (11) dieser, contrast 225 (12) 395 (20) 286 (12) der, no-contrast 155 (11) 421 (23) 295 (12) dieser, no-contrast 218 (11) 387 (18) 286 (12) table a2. means (and standard errors) across regions for right-bound rereading in experiment 1 çokal and von heusinger 70 appendix b: odd ratios in experiment 2 predictors odds ratios std. error p-value (intercept) 9.77 (5.89 – 16.22) 2.53 <0.001* contrast 0.69 (0.49 – 0.99) 0.13 0.041* anaphora 0.67 (0.47 – 0.95) 0.12 0.024* contrast * anaphora 1.17 (0.58 – 2.35) 0.42 0.659 random effects σ2 3.29 τ00 subj 1.55 τ00 item 0.12 icc 0.34 nsubj 32 nitem 40 observations 1257 marginal r2 / conditional r2 0.015/ 0.346 table a3. odd ratios of der and dieser in referring to np1 and np2 in experiment 2. german demonstrative pronouns 71 appendix c: cases coded as ‘other cases’ were excluded in the statistical analysis in experiment 2 (sentence completion experiment) der and dieser are used to refer to other entities, as in (1) below: (1) bei einer verhandlung im kölner landgericht habe ich gestern einen richter und einen anwalt intensiv diskutieren gehört. der richter zusammen mit dem anwalt empörte sich über eine zeugenaussage. es herrschte eine angespannte atmosphäre. der hat scheinbar einfach gelogen. (during a trial at the cologne district court, i heard a judge and a lawyer engaged in an intense discussion yesterday. the judge, along with the lawyer, expressed outrage over a witness statement. there was a tense atmosphere. he...apparently just lied.) der and dieser are used to refer to two nps, as in 2 below: (2) beim morgendlichen einkaufen in der überfüllten markthalle des großen einkaufscenters war ich bei einem metzger und einem gemüsehändler. der metzger zusammen mit dem gemüsehändler hat die leute durchgehend bedient. es gab einige aufmerksame fragen und lustige bemerkungen. der hat den leuten einen guten morgen beschert. (during my morning shopping in the crowded market hall of the large shopping center, i was at a butcher's shop and a vegetable vendor's stall. the butcher, along with the vegetable vendor, served people continuously. there were some attentive questions and amusing remarks. der made people's morning enjoyable.) çokal and von heusinger 72 references werner abraham (2002). pronomina im diskurs: deutsche personalund demonstrativpronomina unter ‘zentrierungsperspektive’. grammatische überlegungen zu einer teiltheorie der textkohärenz. sprachwissenschaft, 27(4):447–491. dorothy ahn (2019). that thesis: a competition mechanism for anaphoric expressions. phd thesis, harvard university, cambridge, massachusetts. bernt ahrenholz (2007). verweise mit demonstrativa im gesprochenen deutsch. grammatik, zweitspracherwerb und deutsch als fremdsprache. de gruyter. doi: https://doi.org/10.1515/9783110894127. stephen r. anderson and edward l. keenan (1985). deixis. in timothy shopen, editor, language, typology, and syntactic description, 2: 259–308. cambridge university press, cambridge. doi: doi.org/10.1017/cbo9780511619434. mira ariel (1990). accessing noun-phrase antecedents. routledge, london. doi: https://doi.org/10.4324/9781315857473. jennifer e. arnold, janet g. eisenband, sarah brown-schmidt, and john trueswell (2000). the rapid use of gender information: evidence of the time course of pronoun resolution from eyetracking. cognition, 76:13–26. doi: https://doi.org/10.1016/s0010-0277(00)00073-1. harald baayen (2008). analyzing linguistic data. a practical introduction to statistics using r. cambridge university press, cambridge. doi: doi.org/10.1017/cbo9780511801686. sarah, brown-schmidt, donna, byron, and michael k. tanenhaus (2005). beyond salience: interpretation of personal and demonstrative pronouns. journal of memory and language, 53(2): 292–313. https://doi.org/10.1016/j.jml.2005.03.003 markus bader and yvonne portele (2019). the interpretation of german personal pronouns and d-pronouns. zeitschrift für sprachwissenschaft, 38(2):155-190. doi: https://doi.org/10.1515/zfs-2019-2002. markus bader, yvonne portele, and alice schäfer. semantic bias in the interpretation of german personal and demonstrative pronouns (2022). in robin hörnig, sophie von wietersheim, andreas konietzko, and sam featherston, editors, proceedings of linguistic evidence: linguistic theory enriched by experimental data, pages 399–419. university of tübingen, tübingen. doi: https://publikationen.unituebingen.de/xmlui/handle/10900/119301. douglas bates, martin mächler, ben bolker, and steve walker (2015). fitting linear mixedeffects models. using lme4. journal of statistical software, 67: 1–48. 10.18637/jss.v067.i01 dale j. barr, roger levy, christoph scheepers, and harry j. tily (2013). random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3): 0.1016/j.jml.2012.11.001. doi: 10.1016/j.jml.2012.11.001 david i. beaver (2004). the optimization of discourse anaphora. linguistics and philosophy, 27:3–56. doi: https://doi.org/10.1023/b:ling.0000010796.76522.7a. hansjörg bisle-müller (1991). artikelwörter im deutschen: semantische und pragmatische aspekte ihrer verwendung. niemeyer, tübingen. peter bosch (1988). representing and accessing focussed referents. language and cognitive processes, 3(3):207–231. doi: https://doi.org/10.1080/01690968808402088. peter bosch and stefan hinterwimmer (2016). anaphoric reference by demonstrative pronouns in german. in search of the relevant parameters. in anke holler and kaja suckow, editors, empirical perspectives on anaphora resolution, pages 193–212. de gruyter, berlin. doi: https://doi.org/10.1515/9783110464108-010. peter bosch and carla umbach. reference determination for demonstrative pronouns (2007). zas papers in linguistics, 48:39-51. doi: https://doi.org/10.21248/zaspil.48.2007.353. peter bosch, tom rozario, and yufan zhao (2003). demonstrative pronouns and personal pronouns. german der vs. er. in proceedings of the eacl, workshop on the computational treatment of anaphora, pages 61–68, budapest. doi: https://aclanthology.org/w03-2609. https://doi.org/10.1515/9783110894127 https://psycnet.apa.org/doi/10.1016/j.jml.2005.03.003 https://doi.org/10.1023/b:ling.0000010796.76522.7a german demonstrative pronouns 73 gerlof bouma and holger hopp (2007). coreference preferences for personal pronouns in german. in dagmar bittner and natalia gargarina, editors, proceedings of the conference on intersentential pronominal reference in child and adult language, (zas papers in linguistics 48), pages 53–74. zas, berlin. doi: https://doi.org/10.21248/zaspil.48.2007.354. wallace chafe (1976). givenness, contrastiveness, definiteness, subjects, topics, and point of view. in charles n. li, editor, subject and topic, pages 25–55. academic press, new york. derya çokal, patrick sturt, and fernanda ferreira (2014). deixis: this and that in written narrative discourse. discourse processes, 51(3):201–229. doi: https://doi.org/10.1080/0163853x.2013.866484. derya çokal, and fernanda ferreira (2015). online processing of noisy input from native and non-native language users. poster presented at the 28th annual conference on human sentence processing (cuny), university of southern california, 19-21 march, losangeles, ca. derya çokal, patrick sturt, and fernanda ferreira (2018). processing of it and this in written narrative discourse. discourse processes, 55(3):272–289. doi: https://doi.org/10.1080/0163853x.2016.1236231. derya çokal, ruth filik, patrick sturt, and massimo poesio (2023). anaphoric reference to mereological entities. discourse processes, 60(3):202–223. doi: https://doi.org/10.1080/0163853x.2023.2197682. rosalind a. crawley, rosemary stevenson, and david kleinman (1990). the use of heuristic strategies in the interpretation of pronouns. journal of psycholinguistic research, 19:245–264. doi: https://doi.org/10.1007/bf01077259. annika deichsel and klaus von heusinger (2011). the cataphoric potential of indefinites in german. in iris hendrickx, sobha lalitha devi, antónio branco, and ruslan mitkov, editors, anaphora processing and applications. daarc 2011, pages 144–156. springer, heidelberg. holger diessel (1999). demonstratives: form, function and grammaticalization. john benjamins, amsterdam. miriam ellert (2013). resolving ambiguous pronouns in a second language: a visual-world eyetracking study with dutch learners of german. international review of applied linguistics in language teaching, 51(2):171–197. doi: https://doi.org/10.1515/iral-2013-0008. marion fossard, alan garnham, and h. wind cowles (2012). between anaphora and deixis... the resolution of the demonstrative noun phrase “that n”. language and cognitive processes, 27(9):1385–1404. doi: https://doi.org/10.1080/01690965.2011.606668. melanie fuchs and petra b. schumacher (2020). referential shift potential of demonstrative pronouns – evidence from text continuation. in ashild næss, anna margetts, and yvonne treis, editors, demonstratives in discourse, pages 185–213. language science press, berlin. doi: https://langscipress.org/catalog/book/282. kumiko fukumura and roger p. g. van gompel (2015). effects of order of mention and grammatical role on anaphor resolution. learning, memory, and cognition, 41(2):501–525. doi: https://doi.org/10.1037/xlm0000041. morton ann gernsbacher and suzanne shroyer (1989). the cataphoric use of the indefinite this in spoken narratives. memory & cognition, 17(5):536–540. talmy givón (1983). topic continuity in discourse: an introduction. in talmy givón, editor, topic continuity in discourse: a quantitative cross-language study, pages 1–42. john benjamins, amsterdam. doi: https://doi.org/10.1075/tsl.3.01giv. peter gordon, barbara grosz, and laura gilliom (1993). pronouns, names, and the centering of attention in discourse. cognitive science, 17(3):311–346. doi: https://doi.org/10.1207/s15516709cog1703. barbara grosz, aravind joshi, and scott weinstein. (1995). centering: a framework for modeling the local coherence of discourse. computational linguistics, 21:203–226. doi: https://doi.org/10.1.1.14.9312. https://doi.org/10.1007/bf01077259 çokal and von heusinger 74 jeanette k. gundel (1988). universals of topic-comment structure. in michael hammond, edith moravcsik, and jessica r. wirth, editors, studies in syntactic typology, pages 209–239. john benjamins, amsterdam. jeanette k. gundel., nancy, hedberg, & ron, zacharski (1993). cognitive status and the form of referring expressions in discourse. language, 69(2), 274–307. doi: https:doi.org/10.2307/416535 elyse, d., johnson, & jennifer, arnold (2023). the frequency of referential patterns guides pronoun comprehension. journal of experimental psychology: learning, memory, and cognition, 49(8), 1325–1344. doi: https://doi.org/10.1037/xlm0001137 martina hielscher and jochen musseler (1990). anaphoric resolution of singular and plural pronouns: the reference to persons being introduced by different co-ordinating structures. journal of semantics, 7:347–364. nikolaus p. himmelmann (1997). deiktikon, artikel, nominalphrase. zur emergenz syntaktischer struktur. max niemeyer verlag, tübingen. stefan hinterwimmer and peter bosch (2018). demonstrative pronouns and propositional attitudes. in pritty patel-grosz, patrick georg grosz, and sarah zobel, editors, pronouns in embedded contexts at the syntax-semantics interface, pages 105–144. springer. jerry r. hobbs (1979). coherence and coreference. cognitive science, 3(1):67–90. doi: https://doi.org/10.1207/s15516709cog0301_4. juhani järvikivi, roger p. g. van gompel, and jukka hyönä (2017). the interplay of implicit causality, structural heuristics, and anaphor type in ambiguous pronoun resolution. journal of psycholinguistic research, 46:525–550. elsi kaiser (2011b). on the relation between coherence relations and anaphoric demonstratives in german. in ingo reich, eva horch, and dennis paul, editors, proceedings of sinn und bedeutung 15, pages 337–351. universaar – saarland university press, saarbrücken. doi: https://ojs.ub.unikonstanz.de/sub/index.php/sub/article/view/385. elsi kaiser and john trueswell (2004). the referential properties of dutch pronouns and demonstratives: is salience enough? in proceedings of sinn und bedeutung 8, pages 137–150. universität konstanz, konstanz. doi: https://doi.org/10.18148/sub/2004.v8i0.754. jeffrey king (2001). complex demonstrativesa quantificational account. mit press, cambridge, massachusetts. doi: doi.org/10.7551/mitpress/1990.003.0011. sungryong koh and charles clifton jr (2002). resolution of the antecedent of a plural pronoun: ontological categories and predicate symmetry. journal of memory and language, 46:830– 844. doi: https://doi.org/10.1006/jmla.2001.2829. knud lambrecht (1994). information structure and sentence form. topic, focus, and the mental representations of discourse referents. cambridge university press, cambridge. alfons maes, emiel krahmer, and david peeters (2022a). explaining variance in writers’ use of demonstratives: a corpus study demonstrating the importance of discourse genre. journal of general linguistics, 7(1):1–36. doi: https://doi.org/10.16995/glossa.5826. alfons maes, emiel krahmer, and david peeters (2022b). understanding demonstrative reference in text: a new taxonomy based on a new corpus. language and cognition, 14(2):185–207. doi: https://doi.org/10.1017/langcog.2021.28. linda m. moxey, anthony j. sanford, and patrick sturt (2004). constraints on the formation of plural reference objects: the influence of role, conjunction, and type of description. journal of memory and language, 51:346–364. livia polanyı (1986). the linguistic discourse model: towards a formal theory of discourse structure. cambridge, ma: laboratories incorp. umesh patil, peter bosch, and stefan hinterwimmer (2020). constraints on german diese demonstratives: language formality and subject-avoidance. glossa: a journal of general linguistics, 5:1–22. doi: https://doi.org/10.5334/gjgl.962. umesh patil, stefan hinterwimmer, and petra b. schumacher (2023). effect of evaluative german demonstrative pronouns 75 expressions on two types of demonstrative pronouns in german. glossa: a journal of general linguistics, 8:1–29. doi: https://doi.org/10.16995/glossa.9577. clare patterson and petra b. schumacher (2020). the timing of prominence information during the resolution of german personal and demonstrative pronouns. dialogue & discourse, 11(1):1–39. doi: https://doi.org/10.5087/dad.2020.101. clare patterson and petra b. schumacher (2021). interpretation preferences in contexts with three antecedents: examining the role of prominence in german pronouns. cambridge university press, 42(6):1427–1461. doi: 10.1017/s0142716421000291. clare patterson and petra b. schumacher (2023). how focus and position affect the interpretation of demonstrative pronouns. collabra: psychology, 9:1–23. ellen f. prince (1981). on the inferencing of indefinite-this nps. in bonnie webber, aravind k. joshi, and ivan sag, editors, elements of discourse understanding, pages 231–250. cambridge university press, cambridge. pirita pyykkönen and juhani järvikivi (2010). activation and persistence of implicit causality information in spoken language comprehension. experimental psychology, 57:5–16. doi: https://doi.org/10.1027/16183169/a000002. keith rayner, timothy j. slattery, denis drieghe, and simon p. liversedge (2011). eye movements and word skipping during reading: effects of word length and predictability. journal of experimental psychology, 37:514–528. keith rayner and simon p. liversedge (2012). linguistic and cognitive influences on eye movements during reading. in simon p. liversedge, iain gilchrist, and stefan everling (eds), the oxford handbook of eye movements, 752-766. oxford university press. doi.org/10.1093/oxfordhb/9780199539789.013.0041. tanya reinhart (1981). pragmatics and linguistics: an analysis of sentence topics. philosophica, 27:53–94. david, peeters, emiel, krahmer, and alfons, maes (2021). a conceptual framework for the study of demonstrative reference. psychonomic bulletin & review, 28(2), 409–433. https://doi.org/10.3758/s13423-020-01822-8 magdalena repp and petra b. schumacher (2023). what naturalistic stimuli tell us about pronoun resolution in real-time processing. frontiers in artificial intelligence, 6:1–15. doi: 10.3389/frai.2023.1058554. repp sophie and spalek katharina (2021). the role of alternatives in language. frontiers in communication, 6. doi: https://doi.org/10.3389/fcomm.2021.682009. petra b. schumacher, jana backhaus, and manuel dangl (2015). backwardand forward-looking potential of anaphors. frontiers in psychology, 6:1–14. doi: https://doi.org/10.3389/fpsyg.2015.01746. petra b. schumacher, manuel dangl, and elyesa uzun (2016). thematic role as prominence cue during pronoun resolution in german. in anke holler and kaja suckow, editors, empirical perspectives on anaphora resolution, pages 213–240. de gruyter, berlin. doi: https://doi.org/10.1515/9783110464108-011. petra b. schumacher, leah roberts, and juhani järvikivi (2017). agentivity drives real-time pronoun resolution: evidence from german er and der. lingua, 185:25–41. doi: https://doi.org/10.1016/j.lingua.2016.07.004. petra b. schumacher, clare patterson, and magdalena repp (2022). die diskursstrukturierende funktion von demonstrativpronomen. in chiara gianollo, lukasz jedrzejowski, and sofiana i. lindemann, editors, paths through meaning and form. festschrift offered to klaus von heusinger on the occasion of his 60th birthday, pages 216–220. universitätsund stadtbibliothek köln, köln. doi: https://doi.org/10.18716/omp.3. rosemary j. stevenson, rosalind crawley, and david kleinman (1994). thematic roles, focus, and the representation of events. language and cognitive processes, 9(4):519–548. doi: https://doi.org/10.1080/01690969408402130. çokal and von heusinger 76 rosemary j. stevenson, alistair knott, jon oberlander, and sharon mcdonald (2000). interpreting pronouns and connectives: interactions among focusing, thematic roles and coherence relations. language and cognitive processes, 15(3):225–262. doi: https://doi.org/10.1080/016909600386048. rosemary j. stevenson, alexander w. r. nelson, and keith stenning (1995). the role of parallelism in strategies of pronoun comprehension. language and speech, 38:393–418. doi: https://doi.org/10.1177/002383099503800404. patrick sturt and nayoung kwon (2018). processing information during regressions: an application of the reverse boundary-change paradigm. frontiers in psychology, 9:1630. doi: https://doi.org/10.3389/fpsyg.2018.01630. emil van den hofen and evelyn c. ferstl (2018). discourse context modulates the effect of implicit causality on rementions. language and cognition, 10(4):561–594. doi: https://doi.org/10.1017/langcog.2018.17. jacqueline van kampen (2007). anaphoric topic-shift devices. paper presented at the ‘workshop on anaphoric uses of demonstrative expressions’, siegen. titus von der malsburg and bernhard angele (2017). false positives and other statistical errors in standard analyses of eye movements in reading. journal of memory and language, 94:119– 133. klaus von heusinger and petra b. schumacher (2019). discourse prominence: definition and application. journal of pragmatics, 154:117–127. doi: https://doi.org/10.1016/j.pragma.2019.07.025. regina weinert (2007). demonstrative and personal pronouns in formal and informal conversations. in regina weinert, editor, spoken language pragmatics: an analysis of formfunction relations, pages 1–28. continuum, london. frances wilson (2009). processing at the syntax-discourse interface in second language acquisition. phd thesis, university of edinburgh, edinburgh. angelika wöllstein (2022). die dudenredaktion. duden. die grammatik. dudenverlag, berlin. lynsey wolter (2006). that’s that: the semantics and pragmatics of demonstrative noun phrases. phd thesis, university of california, santa cruz, ca. gisela zifonun, ludger hoffmann, and bruno strecker (1997). grammatik der deutschen sprache.10., völlig neu verfasste auflage. de gruyter, berlin. dialogue & discourse 12(2) (2021) 81–114 doi: 10.5210/dad.2021.203 user satisfaction reward estimation across domains: domain-independent dialogue policy learning stefan ultes stefan.ultes@daimler.com mercedes-benz ag research & development sindelfingen, germany wolfgang maier wolfgang.mw.maier@daimler.com mercedes-benz ag research & development sindelfingen, germany editor: kallirroi georgila submitted 12/2020; accepted 07/2021; published online 09/2021 abstract learning suitable and well-performing dialogue behaviour in statistical spoken dialogue systems has been in the focus of research for many years. while most work that is based on reinforcement learning employs an objective measure like task success for modelling the reward signal, we propose to use a reward signal based on user satisfaction. we propose a novel estimator and show that it outperforms all previous estimators while learning temporal dependencies implicitly. we show in simulated experiments that a live user satisfaction estimation model may be applied resulting in higher estimated satisfaction whilst achieving similar success rates. moreover, we show that a satisfaction estimation model trained on one domain may be applied in many other domains that cover a similar task. we verify our findings by employing the model to one of the domains for learning a policy from real users and compare its performance to policies using user satisfaction and task success acquired directly from the users as reward. keywords: statistical dialogue management, reinforcement learning, spoken dialogue systems 1. introduction spoken dialogue systems (sdss) enable voice interaction between technical systems and humans. they have been advanced into our everyday lives. prominent examples include apple’s siri, amazon alexa, or google assistant as well as more specialised systems like the in-car voice assistant heymercedes. spoken dialogue systems that target the fulfilment of a certain task are called taskoriented and are usually built using a modular pipeline architecture comprising speech recognition, semantic decoding, dialogue management (consisting of dialogue state tracking and deciding on the next system action), language generation, and speech synthesis (see figure 1). one prominent way of modelling the decision-making component of a spoken dialogue system is to use (partially observable) markov decision processes ((po)mdps) (lemon and pietquin, 2012; young et al., 2013). there, reinforcement learning (rl) (sutton and barto, 1998) is applied to find the optimal system behaviour represented by the policy π. task-oriented dialogue systems model the reward r, which is used to guide the learning process, traditionally with task success as the principal reward component (gašić and young, 2014; lemon and pietquin, 2007; daubigney et al., 2012; levin and pieraccini, 1997; singh et al., 2002; young et al., 2013; su et al., 2015, 2016). c©2021 stefan ultes and wolfgang maier this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). ultes and maier policy state tracking interaction quality reward estimation semantic decoding speech recognition speech synthesis language generation environment at st rt figure 1: the modular pipeline architecture of a spoken dialogue system with the focus of its application in a reinforcement learning setup and the presented extension of integrating an interaction quality reward estimator as originally proposed by ultes et al. (2017a). the policy learns to take action at at time t while being in state st and receiving reward rt. the goal of this article is to demonstrate that user satisfaction (us) may be used as a reward to learn dialogue policies which not only maximise us but also lead to high task success (ts) rates. we apply two user satisfaction reward estimators (ultes et al., 2017a; ultes, 2019) where one uses a support vector machine (vapnik, 1995) and relies on handcrafted temporal features. the other uses a deep learning network that uses long short-term memory (lstm) cells (hochreiter and schmidhuber, 1997) to learn the temporal dependencies of a multi-turn dialogue interaction implicitly. we argue that training a system to maximise us is a good alternative to ts for the following reasons: 1. user satisfaction is favourable over task success as it represents more accurately the user’s view and thus whether the user is likely to use the system again in the future. in fact, task success has only been used as principal reward component as it has been shown to correlate well with user satisfaction (williams and young, 2004). 2. user satisfaction—in contrast to ts—may be linked to interaction phenomena that are independent of the user’s goal (schmitt and ultes, 2015) and hence, no prior knowledge of the goal or any other domain-dependent information is required. 3. as user satisfaction is independent of application domain information, the use of an estimator of user satisfaction has the potential to generalise well across domains. thus, learning dialogue policies for new, previously unseen domains becomes much easier. following up on previous work (ultes et al., 2011, 2012; ultes and minker, 2014; ultes et al., 2015, 2019), interaction quality (iq)—a less subjective version of user satisfaction1—will be used for estimating the reward. we apply a conventional iq estimator (ultes et al., 2017a) using domainindependent, interaction-related features that do not contain any information about the task or the 1. the relation of us and iq has been closely investigated by schmitt and ultes (2015) and ultes et al. (2013), see also section 3.1. 82 user satisfaction reward estimation across domains goal of the dialogue. this allows the reward estimator to be applicable for learning in unseen domains. to circumvent the dependency of the convectional estimator on handcrafted temporal features, we will additionally present a deep learning-based iq estimator (ultes, 2019) that utilises the capabilities of recurrent neural networks to get rid of all handcrafted features that encode temporal effects. by that, these temporal dependencies may be learned instead. the applied rl framework is shown in figure 1. it has previously been applied for in-domain experiments and simulated evaluation (ultes et al., 2019). within this setup, both iq estimators are used for learning dialogue policies in several domains to analyse their impact on general dialogue performance metrics. moreover, one estimator is used in an experiment where the policy is learned through interaction with real humans. the remainder of the paper is organised as follows: in section 2, related work is presented focusing on dialogue learning and the type of reward that is applied. in section 3, the interaction quality is presented and how it is used in the reward model. the deep learning-based interaction quality estimator is then described in detail in section 3.3 followed by the experiments and results both of the estimator itself and the resulting dialogue policies in section 4. the work described in this article builds upon and extends work published by ultes (2019) and ultes et al. (2017a). 2. relevant related work most of previous work on dialogue policy learning focuses on employing task success as the main reward signal (gašić and young, 2014; gašić et al., 2014; lemon and pietquin, 2007; daubigney et al., 2012; levin and pieraccini, 1997; singh et al., 2002; young et al., 2013; su et al., 2015, 2016). however, task success is usually only computable for predefined tasks, e.g., through interactions with simulated or recruited users, where the underlying goal is known in advance. to overcome this, the required information can be requested directly from users at the end of each dialogue (gašić et al., 2013). however, this can be intrusive, and users may not always cooperate. an alternative is to use a task success estimator (el asri et al., 2014b; su et al., 2015, 2016). with the right choice of features, such a task success estimator can also be applied to new and unseen domains (vandyke et al., 2015). however, these models still attempt to estimate completion of the underlying task, whereas our model evaluates the overall user experience. in this paper, we show that an interaction quality reward estimator trained on dialogues from a bus information system will result in well-performing dialogues both in terms of success rate and user satisfaction on five other domains, while only using interaction-related, domain-independent information, i.e., not knowing anything about the task of the domain. others have previously introduced user satisfaction into the reward (walker et al., 1998; walker, 2000; henderson et al., 2008; rieser and lemon, 2008b,a) by using the paradise framework initially proposed by walker et al. (1997). however, paradise relies on the existence of explicit task success information, which is usually hard to obtain. furthermore, to derive user ratings within that framework, users have to answer a questionnaire, which is usually not feasible in real world settings. to overcome this, paradise has been used in conjunction with expert judges instead (el asri et al., 2012, 2013) to enable unintrusive acquisition of dialogues. however, the problem of mapping the results of the questionnaire to a scalar reward value still exists. a similar measure called response quality has been proposed as a measure to capture user satisfaction as an alternative to interaction quality (bodigutla et al., 2019b,a, 2020). in contrast to 83 ultes and maier the interaction quality, the response quality focuses more on the overal performance of a system, e.g., including the functionality of back-end services. thus, it is hard to identify if a negative score should be attributed to the performance of a back-end service or the interaction itself. thus, the response quality is not suitable for learning the behaviour of a dialogue system. other research uses different cues for reward estimation. for instance, misu et al. (2012) investigate domain-independent dialogue policy learning in the context of question-answering dialogues. users are asked to rate dialogues on a likert scale, then regression is used to compute rewards for question-answer pairs. as there is no explicit task success metric (their dialogues are not strictly task-oriented and somewhat similar to chat), their approach is also domain-independent. shi and yu (2018) incorporate user sentiment information obtained from multi-modal cues in both a supervised and a reinforcement learning setup. liu and lane (2018) forego the need for explicit ratings and rely only on adversarial rewards. in the area of non-task-oriented dialogue, cuayáhuitl et al. (2018) rely on deep reinforcement learning and investigates the amount of dialogue history that has to be taken into account for the estimation of reward, also without relying on explicit ratings. in this work, we use interaction quality (section 3) due to the following reasons: interaction quality models user satisfaction instead of task success that targets the measurement of task completion. other than automatically derived measures such as adversarial signals, interaction quality provides expert ratings. these expert ratings reflect the quality of the interaction itself, with no dependency on the connected apis as it is the case for response quality; furthermore, they are easier to obtain than ratings from mixed sources (shi and yu, 2018). interaction quality furthermore uses scalar values applied by experts and only uses task-independent features that are valid across domains and easy to derive, which gives it an edge over the approaches relying on the paradise framework. 3. interaction quality reward estimation in this work, the reward estimator is based on estimating the interaction quality (iq) (schmitt and ultes, 2015) for learning information-seeking dialogue policies. iq represents a less subjective variant of user satisfaction: instead of being acquired from users directly, experts annotate pre-recorded dialogues to avoid the large variance that is often encountered when users rate their dialogues directly (schmitt and ultes, 2015). the iq estimation model will be used as a reward estimator as depicted in figure 1. with parameters that are collected from the dialogue system modules for each time step t, the reward estimator derives the reward rt that is used for learning the dialogue policy π. iq is defined on a five-point scale from five (satisfied) down to one (extremely unsatisfied). to derive a reward from this value, the equation riq = t · (−1) + (iq − 1) · 5 (1) is used where riq describes the final reward. it is applied to the final turn of the dialogue of length t with a final iq value of iq. a per-turn penalty of −1 is added to the dialogue outcome. this results in a reward range of 19 (because there is always at least one turn) down to −t , which is consistent with related work (e.g., gašić and young, 2014; vandyke et al., 2015; su et al., 2016) in which binary task success (ts) was used to define the reward as: rts = t · (−1) + 1ts · 20 , (2) 84 user satisfaction reward estimation across domains s u s u s u s u…s1 u1 s2 u2 s3 u3 sn un … e1 e2 e3 en figure 2: a dialogue may be separated into a sequence of system-user-exchanges where each exchange ei consists of a system turn si followed by a user turn ui. where 1ts = 1 only if the dialogue was successful, 1ts = 0 otherwise. rts will be used as a baseline. 3.1 interaction quality and the lego corpus interaction quality is defined similarly to user satisfaction: while the latter represents the true disposition of the user, iq is the disposition of the user assumed by an expert annotator. here, expert annotators are people who listen to recorded dialogues after the interactions and rate them by assuming the point of view of the actual person performing the dialogue. these experts are supposed to have some experience with dialogue systems. for the lego corpus—the data set used in this work—, expert annotators were “advanced students of computer science and engineering” (schmitt and ultes, 2015; schmitt et al., 2011), i.e., grad students. interaction quality has been shown to be a suitable surrogate for user satisfaction by ultes et al. (2013). comparing user satisfaction ratings and interaction quality ratings for the same data showed a high correlation between the labels. estimation models trained on one (user satisfaction or interaction quality) and evaluated on the other result in estimation performances clearly above chance. furthermore, interaction quality estimation is much more reliable and accurate than user satisfaction estimation. moreover, the interaction quality matches requirements identified by ultes et al. (2012) for an estimation approach to be used for online adaptation: exchange-level ratings, automatically derivable and domain-independent input features, a consistent labelling process and reproducible and unbiased labels. the lego corpus (schmitt et al., 2012b) is based on 200 calls to the “let’s go bus information system” of the carnegie mellon university in pittsburgh (raux et al., 2006) recorded in 2006. labels for iq have been assigned by three expert annotators to 200 calls consisting of 4,885 systemuser-exchanges (see figure 2) in total with an inter-annotator agreement of κ = 0.54 using cohen’s weighted kappa with a linear weighting function 13. this may be considered as a moderate agreement (cf. landis and koch’s kappa benchmark scale (1977)), which is quite good considering the difficulty of the task that required to rate each exchange. for instance, if one annotator reduces the iq value only one exchange earlier than another annotator, both already disagree on two exchanges. the final label of each exchange was derived by using the median of all three individual ratings as the median showed the best overall agreement (schmitt and ultes, 2015). iq was labelled on a scale from 1 (extremely unsatisfied) to 5 (satisfied) considering the complete dialogue up to the current exchange. thus, each exchange has been rated without regarding 85 ultes and maier policy interaction quality reward estimation state tracking semantic decoding speech recognition speech synthesis language generation environment at st rt en… en-1en-2e1 e2 en+1 … dialogue-level window-level exchange-level recognition dialogue figure 3: the overall architecture of deriving the interaction parameters from the spoken dialogue system that model the temporal information explicitly and are used as input to the interaction quality reward estimator. any upcoming user utterance. as the users are expected to be satisfied at the beginning, each dialogue’s initial rating is 5. in order to ensure consistent labelling, the expert annotators had to follow distinct labelling guidelines (schmitt et al., 2012b). an example of an annotated dialogue of the lego corpus is shown in table 1. the dialogue starts off with a good iq until the system first attempts to find results. there, the iq falls drastically as the user input does not correspond to what the system expects and the system remains in a loop until the user reacts appropriately. 3.2 conventional interaction quality estimation the problem of estimating iq has been cast as a classification problem where the target classes are the distinct iq values. the input consists of domain-independent variables called interaction parameters. these parameters incorporate information from the automatic speech recognition (asr) output and the preceding system action (see figure 3): the asr status (one of success, no match, no input), the asr confidence of the highest ranked result, the general type of the system action (one of statement, question), whether the system action is a repetition of the previous system action, and whether the role of the system action is to confirm previous user input. this information, which is available at every turn, is used to compute temporal features by taking sums, means or counts from the turn-based information for a window of the last three system-user-exchanges2 and the complete 2. a system turn followed by a user turn. 86 user satisfaction reward estimation across domains table 1: example of a dialogue annotated with interaction quality for each system-user-exchange. as the system fails to interact properly with the user in the end, the iq value drops down to one. utterance represents either the system utterance or the recognised user utterance. (callid 2061123000 of the lego corpus.) utterance iq s: welcome to the cmu let’s go bus information system. u: 5 s: to get help at any time, just say help or press zero. u: 5 s: what can i do for you? u: penn at butler time is it is the centre at highland from runs the 61a eight thirty 61b can’t oakland 5 s: leaving from oakland . is this correct? u: oakland 4 s: leaving from oakland . is this correct? u: yes 4 s: right u: 4 s: where do you want to go? u: squirrel hill of fifth avenue 4 s: going to fifth avenue . is this correct? u: yes 5 s: alright u: 5 s: when do you wanna travel? u: sunday 4 s: at what time do you wanna travel? u: at five four very about eleven o’clock 4 s: leaving at 11 a.m. . is this correct? u: yes 5 s: right u: 5 s: hold on. let me check that for you. u: 5 s: oakland and fifth avenue are both the same stop. please provide a different start or end point. u: 3 s: to ask about a different trip, you can say, start a new query. if you are finished, you can say goodbye. u: 61a 2 s: you can say start a new query, or goodbye. u: five duquesne 1 s: to ask about a different trip, you can say, start a new query. if you are finished, you can say goodbye. u: last the next 1 s: you can say start a new query, or goodbye. u: start wh 1 s: to ask about a different trip, you can say, start a new query. if you are finished, you can say goodbye. u: the 1 s: you can say start a new query, or goodbye. u: thank you query eleven back 1 s: to ask about a different trip, you can say, start a new query. if you are finished, you can say goodbye. u: goodbye 1 s: thank you for using the cmu let’s go bus information system. goodbye. u: 1 dialogue (see lower part of figure 3). the baseline iq estimation approach as applied by ultes et al. (2017a) (and originating from ultes et al. (2015)) uses a feature set of 16 parameters as shown in table 2 to train a support vector machine (svm) (vapnik, 1995; chang and lin, 2011) with linear kernel. 87 ultes and maier table 2: the interaction parameters extracted from each user input (exchange level) plus counts, sums and rates for the whole dialogue (#, %, mean) and for a window ({·}) of the last 3 turns. parameter description e xc ha ng e le ve l asrrecognitionstatus asr status: success, no match, no input asrconfidence confidence of top asr results reprompt? is the system question the same as in the previous turn? activitytype general type of system action: statement, question confirmation? is system action confirm? d ia lo gu e le ve l meanasrconfidence mean asr confidence if asr is success #exchanges number of exchanges (turns) #asrsuccess count of asr status is success %asrsuccess rate of asr status is success #asrrejections count of asr status is reject %asrrejections rate of asr status is reject w in do w le ve l {mean}asrconfidence mean asr confidence if asr is success {#}asrsuccess count of asr is success {#}asrrejections count of asr status is reject {#}reprompts count of times repromt? is true {#}systemquestions count of activitytype is question the lego corpus (schmitt et al., 2012a) provides data for training and testing and consists of 200 dialogues (4,885 turns) from the let’s go bus information system (raux et al., 2006; eskenazi et al., 2008) of carnegie mellon university in pittsburgh, pa. the system provided information about bus schedules and connections to actual users with real needs and was live from 2006 until 2016. each turn of these 200 dialogues has been annotated with iq (representing the quality of the dialogue up to the current turn) by three experts. the final iq label has been assigned using the median of the three individual labels. previous work has used the lego corpus with a full iq feature set (which includes additional partly domain-related information) achieving an unweighted average recall3 (uar) of 0.55 using ordinal regression (el asri et al., 2014a), 0.53 using a two-level svm approach (ultes and minker, 2013), and 0.51 using a hybrid-hmm (ultes and minker, 2014). human performance on the same task is 0.69 uar (schmitt and ultes, 2015). 3.3 lstm-based interaction quality estimation the architecture of the lstm-based iq estimation model is shown in figure 4. it is based on the idea that the temporal information may be learned by using recurrent neural networks instead of encoding it explicitly with the window and dialogue interaction parameter levels as used by the conventional estimator. thus, only the exchange level parameters et are considered (see table 2). long short-term memory (lstm) cells are at the core of the model and have originally been proposed by hochreiter and schmidhuber (1997) as a recurrent variant that remedies the vanishing gradient problem (bengio et al., 1994). 3. uar is the arithmetic average of all class-wise recalls. 88 user satisfaction reward estimation across domains … … l s t m aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw= l s t m aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw= l s t m aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw= l s t m aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw= l s t m aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw= l s t m aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw=aaahsxiclvxlbtnafl1tig3dq4ufczyraresteoceiwjtsawiao0bawmqryzsa34edlocld6dwzhm/gcpomdqiw493osmjipysueo3fuofc1y5tdxw6jcvnh0nluytx8yupa4dr1gzdvrw/cpgz9xmcpuuu7fnbsgqfybe/vizty1he3uizrouri7ozy+lffbahtewfroktoxapt2s3bmiko6g0njnyp66xytlmuylaoakfe+tr3n3j3qufn8smihrmkykmiskmghbhpqejl6kj3sgl0asrb1hwduwhyhqwulaxoo3i3mtvrwg9z5gwfbcglgycaskhbef4kowlr9qoghxj/4pksuvzmd4kwc4qdjcyy14txdfqrncfiedlvlsnyfim5q4ha9fyysrffvzscpzxi2cnkaf1hvor0qizb4dbl3kcfpix1rmbvhjiujemmrkngjsyezjtaf2dk6nm8s7nrw86a9kz6fnnjwbd0z0ppgqp3oen7zh6g2b8fugldjj1kyblcamhofyubtzka/+nzpmuuml2md3qa62furmsqxj73wkt3rdbde/czuooidig7guxksq/otw9w3r/crs0wfhe3rn3nuv3zkj6bmsylvaluxfbpdh70n2qhj/u5igjv1qlvlxqhzvam9eizp1act0uisrcncc4ynpssbu2izvwddgyrzexm8rin9zhfcdhoj3att/pennh2peo1jp5wmrs2sc63ivmitnhpb9lpgesc2g772zed4kk+tzdvjbiu7lvx+4v83plytkwwerh7la9zlygy5bpbnenpmrfqhld5hi+p5w4ups6h8rsz6otvy4a/tgxyf5ivdqs7fcjvnpzqd/s/zpxu0x16iio8q/9f0t6+gbb8fagv9c3/pf8z/yv/ozvdxtkyozr2ret+ak74uvw= softmax aaahthiclvvlb9naej62klbh1ckba5eitagjwiubcy6rwhaheawanljtidvzpfb8igwnofj9e1zhn3hnf3adjl4zb0it51fs2ts7o983r13b7dp2gjvkp5awv65cza2ureevxb9x89bg5u2j0o8flqpzvumhddmilwn7qhbzkapq3uazrumoy7ozx+vhfrwetu8droouonwntme3bmuioko3qr8vuub8cany2i3jvcgkzs0usv8h/ubkxwpqk3yyqecukfioguyqqshueyptibrqnvicxqdjlnvf55qhtgcrbqsd2g7ebcxotnbdndldqvvw4uajgczqnp6xwmjcmr0qychg33g+i64900mizbzhakmjxnvhfan9rgewwir0teuwlsviziqifj2xbgze1xun52mnepaxekdxkzucvrdlnjhmmfdraq9jdrfwlycmbcm4idgquqmlpxkn8auyufocz+zs2razod2tpsx0wnyt3dlqusbofwjsprofyfzvgvtceemneligkfqyc8vi4jom5v/4homyk0yv6qmdwnoylsuzkrhnprbpfzh18b58ptqgoj3irmc5cuyd8+2dlfcvd2sbdaxcjvhfuw7/pkrnokpxcrwboxdump0h/sfzwen+lqlyw2xcwwtdido8ct3s7kkv59osacxey4izhs2wrfuvabnxaanbjhmrszzu4d3ov0s08ypcwpo+f2e0c6l6jaonvcygtazzbwiwolpc84f0eia5p7bdfnzkh3iszxpmeymujxtt3p4ip3c+llorzv7epstd1kuibpnsdud4muatmuftnsfl47kd+anzotztlpr62xjgd1oe/j9khapkbhnyu6ff6gp9r1mje3sfhqiiz9d/v3salycf5i/0lb7lvud+5n7l/qsmy0sac4fgrtxcx3zkurw=aaahthiclvvlb9naej62klbh1ckba5eitagjwiubcy6rwhaheawanljtidvzpfb8igwnofj9e1zhn3hnf3adjl4zb0it51fs2ts7o983r13b7dp2gjvkp5awv65cza2ureevxb9x89bg5u2j0o8flqpzvumhddmilwn7qhbzkapq3uazrumoy7ozx+vhfrwetu8droouonwntme3bmuioko3qr8vuub8cany2i3jvcgkzs0usv8h/ubkxwpqk3yyqecukfioguyqqshueyptibrqnvicxqdjlnvf55qhtgcrbqsd2g7ebcxotnbdndldqvvw4uajgczqnp6xwmjcmr0qychg33g+i64900mizbzhakmjxnvhfan9rgewwir0teuwlsviziqifj2xbgze1xun52mnepaxekdxkzucvrdlnjhmmfdraq9jdrfwlycmbcm4idgquqmlpxkn8auyufocz+zs2razod2tpsx0wnyt3dlqusbofwjsprofyfzvgvtceemneligkfqyc8vi4jom5v/4homyk0yv6qmdwnoylsuzkrhnprbpfzh18b58ptqgoj3irmc5cuyd8+2dlfcvd2sbdaxcjvhfuw7/pkrnokpxcrwboxdump0h/sfzwen+lqlyw2xcwwtdido8ct3s7kkv59osacxey4izhs2wrfuvabnxaanbjhmrszzu4d3ov0s08ypcwpo+f2e0c6l6jaonvcygtazzbwiwolpc84f0eia5p7bdfnzkh3iszxpmeymujxtt3p4ip3c+llorzv7epstd1kuibpnsdud4muatmuftnsfl47kd+anzotztlpr62xjgd1oe/j9khapkbhnyu6ff6gp9r1mje3sfhqiiz9d/v3salycf5i/0lb7lvud+5n7l/qsmy0sac4fgrtxcx3zkurw=aaahthiclvvlb9naej62klbh1ckba5eitagjwiubcy6rwhaheawanljtidvzpfb8igwnofj9e1zhn3hnf3adjl4zb0it51fs2ts7o983r13b7dp2gjvkp5awv65cza2ureevxb9x89bg5u2j0o8flqpzvumhddmilwn7qhbzkapq3uazrumoy7ozx+vhfrwetu8droouonwntme3bmuioko3qr8vuub8cany2i3jvcgkzs0usv8h/ubkxwpqk3yyqecukfioguyqqshueyptibrqnvicxqdjlnvf55qhtgcrbqsd2g7ebcxotnbdndldqvvw4uajgczqnp6xwmjcmr0qychg33g+i64900mizbzhakmjxnvhfan9rgewwir0teuwlsviziqifj2xbgze1xun52mnepaxekdxkzucvrdlnjhmmfdraq9jdrfwlycmbcm4idgquqmlpxkn8auyufocz+zs2razod2tpsx0wnyt3dlqusbofwjsprofyfzvgvtceemneligkfqyc8vi4jom5v/4homyk0yv6qmdwnoylsuzkrhnprbpfzh18b58ptqgoj3irmc5cuyd8+2dlfcvd2sbdaxcjvhfuw7/pkrnokpxcrwboxdump0h/sfzwen+lqlyw2xcwwtdido8ct3s7kkv59osacxey4izhs2wrfuvabnxaanbjhmrszzu4d3ov0s08ypcwpo+f2e0c6l6jaonvcygtazzbwiwolpc84f0eia5p7bdfnzkh3iszxpmeymujxtt3p4ip3c+llorzv7epstd1kuibpnsdud4muatmuftnsfl47kd+anzotztlpr62xjgd1oe/j9khapkbhnyu6ff6gp9r1mje3sfhqiiz9d/v3salycf5i/0lb7lvud+5n7l/qsmy0sac4fgrtxcx3zkurw=aaahthiclvvlb9naej62klbh1ckba5eitagjwiubcy6rwhaheawanljtidvzpfb8igwnofj9e1zhn3hnf3adjl4zb0it51fs2ts7o983r13b7dp2gjvkp5awv65cza2ureevxb9x89bg5u2j0o8flqpzvumhddmilwn7qhbzkapq3uazrumoy7ozx+vhfrwetu8droouonwntme3bmuioko3qr8vuub8cany2i3jvcgkzs0usv8h/ubkxwpqk3yyqecukfioguyqqshueyptibrqnvicxqdjlnvf55qhtgcrbqsd2g7ebcxotnbdndldqvvw4uajgczqnp6xwmjcmr0qychg33g+i64900mizbzhakmjxnvhfan9rgewwir0teuwlsviziqifj2xbgze1xun52mnepaxekdxkzucvrdlnjhmmfdraq9jdrfwlycmbcm4idgquqmlpxkn8auyufocz+zs2razod2tpsx0wnyt3dlqusbofwjsprofyfzvgvtceemneligkfqyc8vi4jom5v/4homyk0yv6qmdwnoylsuzkrhnprbpfzh18b58ptqgoj3irmc5cuyd8+2dlfcvd2sbdaxcjvhfuw7/pkrnokpxcrwboxdump0h/sfzwen+lqlyw2xcwwtdido8ct3s7kkv59osacxey4izhs2wrfuvabnxaanbjhmrszzu4d3ov0s08ypcwpo+f2e0c6l6jaonvcygtazzbwiwolpc84f0eia5p7bdfnzkh3iszxpmeymujxtt3p4ip3c+llorzv7epstd1kuibpnsdud4muatmuftnsfl47kd+anzotztlpr62xjgd1oe/j9khapkbhnyu6ff6gp9r1mje3sfhqiiz9d/v3salycf5i/0lb7lvud+5n7l/qsmy0sac4fgrtxcx3zkurw= attention aaahpniclvxbbtnaej22qnpya+grl4i0qbk0sgispfzkqtxwkda0fu1fbwetwvelsjchwepf8ar/xb/wwjnxpjrxlq0te2dn55y57dp2x3njxsz+wvhcunhzvm55zfx2nbv37q+tpzimw27kqkotemf0zfux8txavbwrpxxuiztl256q2e0kr9d6kordmdjqg4468a1w4dzdx9jqfa1bwquaxw9rhejoua58vigzoudm2g/xlypupwaf5fcxffiukibskuux7mmquze60j1qal0eyzv1ree0cmwxvgowfrrtvfuyhrttgdlzxoj24mxdewgzp008b4trhjv7vzbjjh/x/bbda6qhrjg5wgfgg4wrwvgeek1nsjih9i3lmjb5sm5ku5nestyu4uuihvn0lnj2sbjb15avpl0wyxy4bjn3uieayxurcjwhdhnjuihrklejs2aylfbfgln6hm/07fqws6a9kz716zmso6azsxsn0xvqph1mp8pspwkpheflpxjyxpcamhpf+sangc31+fjwjh5xqneuxclqr8frye7xhe0dfaedwa+z9awcsux572qz37iepoppltpqqkd2bcsdyr9cxx5y+vzwltgeqx53/wi/ct/+e0jp1q7bjvqa7lxqaxyb9n/lziz6uyxib+uxb03uqj/hteipyu+toj+mromgghucfdhssls7ei3j6mbgi2qmyprhbbxh+qqidnaeg3js9/ymtq9ur1h0piq5sombfbuyxegu93wgnr5izqntsj9t2sgb5pmc8+4fl5k9nmp/mz8735ftlmwexz7nq9zljgz5m9cz4wmst/kyt1ker47ndqxona/lswcxnn/cch+u0vh/kisclndkkd+9kow+mf+wzxpej2klfxmj/r+lfxxzhut6k37r79xw7koumqulposlbvoqrq7c6t+j4wnwaaahpniclvxbbtnaej22qnpya+grl4i0qbk0sgispfzkqtxwkda0fu1fbwetwvelsjchwepf8ar/xb/wwjnxpjrxlq0te2dn55y57dp2x3njxsz+wvhcunhzvm55zfx2nbv37q+tpzimw27kqkotemf0zfux8txavbwrpxxuiztl256q2e0kr9d6kordmdjqg4468a1w4dzdx9jqfa1bwquaxw9rhejoua58vigzoudm2g/xlypupwaf5fcxffiukibskuux7mmquze60j1qal0eyzv1ree0cmwxvgowfrrtvfuyhrttgdlzxoj24mxdewgzp008b4trhjv7vzbjjh/x/bbda6qhrjg5wgfgg4wrwvgeek1nsjih9i3lmjb5sm5ku5nestyu4uuihvn0lnj2sbjb15avpl0wyxy4bjn3uieayxurcjwhdhnjuihrklejs2aylfbfgln6hm/07fqws6a9kz716zmso6azsxsn0xvqph1mp8pspwkpheflpxjyxpcamhpf+sangc31+fjwjh5xqneuxclqr8frye7xhe0dfaedwa+z9awcsux572qz37iepoppltpqqkd2bcsdyr9cxx5y+vzwltgeqx53/wi/ct/+e0jp1q7bjvqa7lxqaxyb9n/lziz6uyxib+uxb03uqj/hteipyu+toj+mromgghucfdhssls7ei3j6mbgi2qmyprhbbxh+qqidnaeg3js9/ymtq9ur1h0piq5sombfbuyxegu93wgnr5izqntsj9t2sgb5pmc8+4fl5k9nmp/mz8735ftlmwexz7nq9zljgz5m9cz4wmst/kyt1ker47ndqxona/lswcxnn/cch+u0vh/kisclndkkd+9kow+mf+wzxpej2klfxmj/r+lfxxzhut6k37r79xw7koumqulposlbvoqrq7c6t+j4wnwaaahpniclvxbbtnaej22qnpya+grl4i0qbk0sgispfzkqtxwkda0fu1fbwetwvelsjchwepf8ar/xb/wwjnxpjrxlq0te2dn55y57dp2x3njxsz+wvhcunhzvm55zfx2nbv37q+tpzimw27kqkotemf0zfux8txavbwrpxxuiztl256q2e0kr9d6kordmdjqg4468a1w4dzdx9jqfa1bwquaxw9rhejoua58vigzoudm2g/xlypupwaf5fcxffiukibskuux7mmquze60j1qal0eyzv1ree0cmwxvgowfrrtvfuyhrttgdlzxoj24mxdewgzp008b4trhjv7vzbjjh/x/bbda6qhrjg5wgfgg4wrwvgeek1nsjih9i3lmjb5sm5ku5nestyu4uuihvn0lnj2sbjb15avpl0wyxy4bjn3uieayxurcjwhdhnjuihrklejs2aylfbfgln6hm/07fqws6a9kz716zmso6azsxsn0xvqph1mp8pspwkpheflpxjyxpcamhpf+sangc31+fjwjh5xqneuxclqr8frye7xhe0dfaedwa+z9awcsux572qz37iepoppltpqqkd2bcsdyr9cxx5y+vzwltgeqx53/wi/ct/+e0jp1q7bjvqa7lxqaxyb9n/lziz6uyxib+uxb03uqj/hteipyu+toj+mromgghucfdhssls7ei3j6mbgi2qmyprhbbxh+qqidnaeg3js9/ymtq9ur1h0piq5sombfbuyxegu93wgnr5izqntsj9t2sgb5pmc8+4fl5k9nmp/mz8735ftlmwexz7nq9zljgz5m9cz4wmst/kyt1ker47ndqxona/lswcxnn/cch+u0vh/kisclndkkd+9kow+mf+wzxpej2klfxmj/r+lfxxzhut6k37r79xw7koumqulposlbvoqrq7c6t+j4wnwaaahpniclvxbbtnaej22qnpya+grl4i0qbk0sgispfzkqtxwkda0fu1fbwetwvelsjchwepf8ar/xb/wwjnxpjrxlq0te2dn55y57dp2x3njxsz+wvhcunhzvm55zfx2nbv37q+tpzimw27kqkotemf0zfux8txavbwrpxxuiztl256q2e0kr9d6kordmdjqg4468a1w4dzdx9jqfa1bwquaxw9rhejoua58vigzoudm2g/xlypupwaf5fcxffiukibskuux7mmquze60j1qal0eyzv1ree0cmwxvgowfrrtvfuyhrttgdlzxoj24mxdewgzp008b4trhjv7vzbjjh/x/bbda6qhrjg5wgfgg4wrwvgeek1nsjih9i3lmjb5sm5ku5nestyu4uuihvn0lnj2sbjb15avpl0wyxy4bjn3uieayxurcjwhdhnjuihrklejs2aylfbfgln6hm/07fqws6a9kz716zmso6azsxsn0xvqph1mp8pspwkpheflpxjyxpcamhpf+sangc31+fjwjh5xqneuxclqr8frye7xhe0dfaedwa+z9awcsux572qz37iepoppltpqqkd2bcsdyr9cxx5y+vzwltgeqx53/wi/ct/+e0jp1q7bjvqa7lxqaxyb9n/lziz6uyxib+uxb03uqj/hteipyu+toj+mromgghucfdhssls7ei3j6mbgi2qmyprhbbxh+qqidnaeg3js9/ymtq9ur1h0piq5sombfbuyxegu93wgnr5izqntsj9t2sgb5pmc8+4fl5k9nmp/mz8735ftlmwexz7nq9zljgz5m9cz4wmst/kyt1ker47ndqxona/lswcxnn/cch+u0vh/kisclndkkd+9kow+mf+wzxpej2klfxmj/r+lfxxzhut6k37r79xw7koumqulposlbvoqrq7c6t+j4wnw ·aaahoxiclvxbbtnaej22qnpwa+grl4i0ehk0sgispezkqscbckvpkzuvsp1naswxyn6ubcvfwct8gv/ca2fgm9leubs27j2dnxpmtmvbpc+ndan0z2v17dbto7n1jfzde/cfpnzcenquh/3iuq0n9mloxlzi5bmbamhxe+qkfynltz11bhdrvh58oalydyndpeypm9/qbg7bdswnvapptel9bbny2ivjvcgkzsmuyvz1cgutrk1quugo9cknrqfpyb5zfom+ptkvqafdgsxqrzbcwvc0ojywfvgpwfjqdvhuyhzqtahmzbkl2oexd08ezif28lwtrhvw7fvbjjh+xfnddj25hhjh5gihgg0wbgjjj+g1ncnigdi3lunylim5k01teipzuiivjxro07nk2cdkbf1xvgr0viw74lblfoekbbgbiicrpgyosmytjjamslgcw2ibl8li1ed45mfxgz0f7bn0auavzn0xny2la4zehybtm/szz/8zscumwjqvwdkg1mackzyapslobsbhsjbxufknkxqlq70zrye7xxe2j/svdme9ztaxciqx572rzx7lejgany211vsd7auwo8i+ui4xyovzwbtgbwwf3m3l/ct9+o8hpvtvg0uocnxiugl2aftf5wrm+rmkym+vkw9t1giq4u3ouwfprtiftktjibobnapybeu0vymwcu0wsewyedhp4y7ek3xfrls4wm086xt5rrvxqtckelbfxngmtl4tzgj0ins+le4pjefudtzpruyqqpj5ixn/kkvjxpu0v8rpnr/iacsyl2of5yhrjua2/e3olfa0y1tlytsij9fhcwfym+djedzzdm03n8kfqzz9n8okr5w9muqvr4rvivmhrdmtekrpujhx6p97qupl6sdft/pfv3pf3idcpxeqmq6ugmxjmrhyp/8a+nnhhg==aaahoxiclvxbbtnaej22qnpwa+grl4i0ehk0sgispezkqscbckvpkzuvsp1naswxyn6ubcvfwct8gv/ca2fgm9leubs27j2dnxpmtmvbpc+ndan0z2v17dbto7n1jfzde/cfpnzcenquh/3iuq0n9mloxlzi5bmbamhxe+qkfynltz11bhdrvh58oalydyndpeypm9/qbg7bdswnvapptel9bbny2ivjvcgkzsmuyvz1cgutrk1quugo9cknrqfpyb5zfom+ptkvqafdgsxqrzbcwvc0ojywfvgpwfjqdvhuyhzqtahmzbkl2oexd08ezif28lwtrhvw7fvbjjh+xfnddj25hhjh5gihgg0wbgjjj+g1ncnigdi3lunylim5k01teipzuiivjxro07nk2cdkbf1xvgr0viw74lblfoekbbgbiicrpgyosmytjjamslgcw2ibl8li1ed45mfxgz0f7bn0auavzn0xny2la4zehybtm/szz/8zscumwjqvwdkg1mackzyapslobsbhsjbxufknkxqlq70zrye7xxe2j/svdme9ztaxciqx572rzx7lejgany211vsd7auwo8i+ui4xyovzwbtgbwwf3m3l/ct9+o8hpvtvg0uocnxiugl2aftf5wrm+rmkym+vkw9t1giq4u3ouwfprtiftktjibobnapybeu0vymwcu0wsewyedhp4y7ek3xfrls4wm086xt5rrvxqtckelbfxngmtl4tzgj0ins+le4pjefudtzpruyqqpj5ixn/kkvjxpu0v8rpnr/iacsyl2of5yhrjua2/e3olfa0y1tlytsij9fhcwfym+djedzzdm03n8kfqzz9n8okr5w9muqvr4rvivmhrdmtekrpujhx6p97qupl6sdft/pfv3pf3idcpxeqmq6ugmxjmrhyp/8a+nnhhg==aaahoxiclvxbbtnaej22qnpwa+grl4i0ehk0sgispezkqscbckvpkzuvsp1naswxyn6ubcvfwct8gv/ca2fgm9leubs27j2dnxpmtmvbpc+ndan0z2v17dbto7n1jfzde/cfpnzcenquh/3iuq0n9mloxlzi5bmbamhxe+qkfynltz11bhdrvh58oalydyndpeypm9/qbg7bdswnvapptel9bbny2ivjvcgkzsmuyvz1cgutrk1quugo9cknrqfpyb5zfom+ptkvqafdgsxqrzbcwvc0ojywfvgpwfjqdvhuyhzqtahmzbkl2oexd08ezif28lwtrhvw7fvbjjh+xfnddj25hhjh5gihgg0wbgjjj+g1ncnigdi3lunylim5k01teipzuiivjxro07nk2cdkbf1xvgr0viw74lblfoekbbgbiicrpgyosmytjjamslgcw2ibl8li1ed45mfxgz0f7bn0auavzn0xny2la4zehybtm/szz/8zscumwjqvwdkg1mackzyapslobsbhsjbxufknkxqlq70zrye7xxe2j/svdme9ztaxciqx572rzx7lejgany211vsd7auwo8i+ui4xyovzwbtgbwwf3m3l/ct9+o8hpvtvg0uocnxiugl2aftf5wrm+rmkym+vkw9t1giq4u3ouwfprtiftktjibobnapybeu0vymwcu0wsewyedhp4y7ek3xfrls4wm086xt5rrvxqtckelbfxngmtl4tzgj0ins+le4pjefudtzpruyqqpj5ixn/kkvjxpu0v8rpnr/iacsyl2of5yhrjua2/e3olfa0y1tlytsij9fhcwfym+djedzzdm03n8kfqzz9n8okr5w9muqvr4rvivmhrdmtekrpujhx6p97qupl6sdft/pfv3pf3idcpxeqmq6ugmxjmrhyp/8a+nnhhg==aaahoxiclvxbbtnaej22qnpwa+grl4i0ehk0sgispezkqscbckvpkzuvsp1naswxyn6ubcvfwct8gv/ca2fgm9leubs27j2dnxpmtmvbpc+ndan0z2v17dbto7n1jfzde/cfpnzcenquh/3iuq0n9mloxlzi5bmbamhxe+qkfynltz11bhdrvh58oalydyndpeypm9/qbg7bdswnvapptel9bbny2ivjvcgkzsmuyvz1cgutrk1quugo9cknrqfpyb5zfom+ptkvqafdgsxqrzbcwvc0ojywfvgpwfjqdvhuyhzqtahmzbkl2oexd08ezif28lwtrhvw7fvbjjh+xfnddj25hhjh5gihgg0wbgjjj+g1ncnigdi3lunylim5k01teipzuiivjxro07nk2cdkbf1xvgr0viw74lblfoekbbgbiicrpgyosmytjjamslgcw2ibl8li1ed45mfxgz0f7bn0auavzn0xny2la4zehybtm/szz/8zscumwjqvwdkg1mackzyapslobsbhsjbxufknkxqlq70zrye7xxe2j/svdme9ztaxciqx572rzx7lejgany211vsd7auwo8i+ui4xyovzwbtgbwwf3m3l/ct9+o8hpvtvg0uocnxiugl2aftf5wrm+rmkym+vkw9t1giq4u3ouwfprtiftktjibobnapybeu0vymwcu0wsewyedhp4y7ek3xfrls4wm086xt5rrvxqtckelbfxngmtl4tzgj0ins+le4pjefudtzpruyqqpj5ixn/kkvjxpu0v8rpnr/iacsyl2of5yhrjua2/e3olfa0y1tlytsij9fhcwfym+djedzzdm03n8kfqzz9n8okr5w9muqvr4rvivmhrdmtekrpujhx6p97qupl6sdft/pfv3pf3idcpxeqmq6ugmxjmrhyp/8a+nnhhg== ~ht aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnjxnug==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnjxnug==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnjxnug==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnjxnug== ~h2 aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwwxnea==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwwxnea==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwwxnea==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwwxnea== ~h1 aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypujnndw==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypujnndw==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypujnndw==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2kl3ljbxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypujnndw== ~h1 aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypug1ndw==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypug1ndw==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypug1ndw==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2qpuleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypug1ndw== ~h2 aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwuznea==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwuznea==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwuznea==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg2q5uleo7hxlymefkhekzk6jchplgcpuo5ac6pbpiglskd2ykmz9qsuquhu6s0qgiyc5sq5oqovadmclygfb28k7gdmf0qaym2csaadepdwrkhnawfnogg1ys1cfocb4f88p0tvmekiemspsy7tbucamh6hx1itfiqrvliexlezyvprq9eaycrffwzscpzpiocrkbf1lvvl0viwb4lbl3kufaowniicrpgtis8y1jjamslgcw2ibl8li1ed4zmfxgj0fbvp61kmxsu6yzsbsnuyfqpp2mf0ms/8epbiglz1kybldqmpofesbn2q0n+njwzt4xkngqlqs1d6m15pd4wvbb/pkx7aezutlbzxy897vzr9lpxwbny211xqa2rcsd4r9cb27yovzwbtgbwx53jxrfuj+/peqnq19g0uoanxaugl2aftf5wsm+7mkym/lcw911kix4u3ouwfprtifuktjibobnd3ybeu0+xit4ypgyitklmkwx128x/kkihz+hnt40vfijhavva9x9lskubdpmxxrmmxofpe8l53us86p7bcfldkhgetzevpoievjxhu3v8rpne/jacsyl2kf5shrjua2/e1oz/e0zvt5wts8j9fhcwfwp86h8rszgjpvboq/v2nyp5uvtst7jcifxxx2n5p/2co9pif0dbv5jf6/pyn8wfm795n+0e9coxeeq+as1hr5ywae0div+/ypwuznea== ~ht aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnhznug==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnhznug==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnhznug==aaahshiclvxbbtnaej22qnpya8sjlxfperk0sgispfzqqtyakncb1fsr7wwse18iexmsrpwkr/bl/aspnblvqhpn0tqyd3z2zpnbrm23ptfwxekfpewvw7fv5fbx1u/eu//g4cbm1mkcdijhntihf0bnthurzw3uixa1p87bkbj821nnduua18+6kordmdjw/ba69k1g4nzdx9jqvte2ko7qjhxf0k27njqhg6qubhske0w58lmhziqcmeso3fw5oarvkcshoustooa0zi8sinffuimk1ibukhloikiurcsa0dqwhvgpwfjqtvbuyhzhtahmzbkl2oexd08ezj528lwtrhvw7fvbjjh+xfnddi2zhhjh5gj7gg0wrgnjr+g1nwgxcokby2esi5gclay6vzfsxmtxfg3n6yx4dresqdeslty9fcsgogyzd1gbaomjiuaqdxnyknenoywjepbamfrgizby9tme2dk1ygdb25q+9eifrdums7f0jdgh0kr9zj/d7d8bqyrbs6cswmaq6phzxxrajxnnzfhy1iyev6oxkk5kttfj9wt3+ml2gb7smayh2fpsqsx2vhe12w9zd1/az0ttnr1a9gxlhwefxmcuwplc8c7yaumed2w0n7gf/z2kz2vf4biqadcqvjpdap13ornjfq6i2ft5wlsdneplebn6bthtk86nlte4imygzw822xltvktluaoy2ckzi5jlcrfvcb4cop0f4tae9l04o91r1wscpa1ilmx6jt8azje6xt3vs6f7knnqo+xns3ziipm8xlwz4lky18btr/jz53ty2rlmi9hnech6izetfxpaczxn81ae8dbp4/xx3ih1qfohpo0shuabg+hpvzr8t2wf0/jecflnv4x9p+yftkqp6qk9q0veo//v6qhfvv7u/arf9dtxzp3nqjkrnv1emphhnhblvv0dnhznug== et aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz/pso1dclcqvzwolixtixafh5kqf6tsgkbzqkk+katkqpbioxn1cjspsb7ptsqclilmyrmha68b2yavgyuhbxruf2ynrbpgzzyxob148pbgqedrg804ybvizvwu5xvgxzw/rtwz6siszi+xjtmg4jowfodd0dotfsn9ydmnzjossndxpjwtjir6oadhpz8szj5uiuras5omtwlbaycv8ahuimfyravd5yjcxjbsylrmvsasg0qjfhjgrz/hmzq4fowvac+ltj17iumm6g0vxgl0ptdpn9jpm/hoqshi0dcqbzqypitlxrad8ktfcj49lbejxprod6upwez1et3apl2wf6csdwnqyrs8vvglpe1eb/zb18av8ttrwuwwyl1jucpvgol6alc8f74jtmorx10f7ifvx30n6tsogl1abuigg0uwc6l/lyrj3cxnf3vymvdvro16gn6hnhj214nyaeo2dagxw9mczjdgwjvrg1chafslcxcypo3ip8xuq7fwit/ck78uz7vypxupoarvzydmz+tywi9ep7nlfot2xnfpbyt/bskmcyecl5t0rl5k9nm5/mz8735ptlmvexd7lq9zljgz5m9cz42mat70jb/m8xh3phvifoh/k085ial65ef5cpcn/vfy42tstqf78qlb+yv5hq/sihtntvoq1+v+edvbldvctn/slfuee5q5ytdy31hr5ywae0tivs/8bqv5kua==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz/pso1dclcqvzwolixtixafh5kqf6tsgkbzqkk+katkqpbioxn1cjspsb7ptsqclilmyrmha68b2yavgyuhbxruf2ynrbpgzzyxob148pbgqedrg804ybvizvwu5xvgxzw/rtwz6siszi+xjtmg4jowfodd0dotfsn9ydmnzjossndxpjwtjir6oadhpz8szj5uiuras5omtwlbaycv8ahuimfyravd5yjcxjbsylrmvsasg0qjfhjgrz/hmzq4fowvac+ltj17iumm6g0vxgl0ptdpn9jpm/hoqshi0dcqbzqypitlxrad8ktfcj49lbejxprod6upwez1et3apl2wf6csdwnqyrs8vvglpe1eb/zb18av8ttrwuwwyl1jucpvgol6alc8f74jtmorx10f7ifvx30n6tsogl1abuigg0uwc6l/lyrj3cxnf3vymvdvro16gn6hnhj214nyaeo2dagxw9mczjdgwjvrg1chafslcxcypo3ip8xuq7fwit/ck78uz7vypxupoarvzydmz+tywi9ep7nlfot2xnfpbyt/bskmcyecl5t0rl5k9nm5/mz8735ptlmvexd7lq9zljgz5m9cz42mat70jb/m8xh3phvifoh/k085ial65ef5cpcn/vfy42tstqf78qlb+yv5hq/sihtntvoq1+v+edvbldvctn/slfuee5q5ytdy31hr5ywae0tivs/8bqv5kua==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz/pso1dclcqvzwolixtixafh5kqf6tsgkbzqkk+katkqpbioxn1cjspsb7ptsqclilmyrmha68b2yavgyuhbxruf2ynrbpgzzyxob148pbgqedrg804ybvizvwu5xvgxzw/rtwz6siszi+xjtmg4jowfodd0dotfsn9ydmnzjossndxpjwtjir6oadhpz8szj5uiuras5omtwlbaycv8ahuimfyravd5yjcxjbsylrmvsasg0qjfhjgrz/hmzq4fowvac+ltj17iumm6g0vxgl0ptdpn9jpm/hoqshi0dcqbzqypitlxrad8ktfcj49lbejxprod6upwez1et3apl2wf6csdwnqyrs8vvglpe1eb/zb18av8ttrwuwwyl1jucpvgol6alc8f74jtmorx10f7ifvx30n6tsogl1abuigg0uwc6l/lyrj3cxnf3vymvdvro16gn6hnhj214nyaeo2dagxw9mczjdgwjvrg1chafslcxcypo3ip8xuq7fwit/ck78uz7vypxupoarvzydmz+tywi9ep7nlfot2xnfpbyt/bskmcyecl5t0rl5k9nm5/mz8735ptlmvexd7lq9zljgz5m9cz42mat70jb/m8xh3phvifoh/k085ial65ef5cpcn/vfy42tstqf78qlb+yv5hq/sihtntvoq1+v+edvbldvctn/slfuee5q5ytdy31hr5ywae0tivs/8bqv5kua==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz/pso1dclcqvzwolixtixafh5kqf6tsgkbzqkk+katkqpbioxn1cjspsb7ptsqclilmyrmha68b2yavgyuhbxruf2ynrbpgzzyxob148pbgqedrg804ybvizvwu5xvgxzw/rtwz6siszi+xjtmg4jowfodd0dotfsn9ydmnzjossndxpjwtjir6oadhpz8szj5uiuras5omtwlbaycv8ahuimfyravd5yjcxjbsylrmvsasg0qjfhjgrz/hmzq4fowvac+ltj17iumm6g0vxgl0ptdpn9jpm/hoqshi0dcqbzqypitlxrad8ktfcj49lbejxprod6upwez1et3apl2wf6csdwnqyrs8vvglpe1eb/zb18av8ttrwuwwyl1jucpvgol6alc8f74jtmorx10f7ifvx30n6tsogl1abuigg0uwc6l/lyrj3cxnf3vymvdvro16gn6hnhj214nyaeo2dagxw9mczjdgwjvrg1chafslcxcypo3ip8xuq7fwit/ck78uz7vypxupoarvzydmz+tywi9ep7nlfot2xnfpbyt/bskmcyecl5t0rl5k9nm5/mz8735ptlmvexd7lq9zljgz5m9cz42mat70jb/m8xh3phvifoh/k085ial65ef5cpcn/vfy42tstqf78qlb+yv5hq/sihtntvoq1+v+edvbldvctn/slfuee5q5ytdy31hr5ywae0tivs/8bqv5kua== e2 aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoaej5vaea8gcjrnuvnvtrnjrpgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzhytdyvmujvta6sa1yew6gktrvnjrprmrybu9v7fy+r1cvvbs7yxck+x115lvnwg24jqwhqtv8s7fsrqig5+xzjujxtyhxpiuujfagcx2gmyv7vkm6herql3xsfjcg7jffme5tklgrotcduqjdbmmvduudwge2cysfcwvant5nze6nnsccownbo/di4ymaznm2nnfcamoavsrimca/eh6irjntqylmhgefow3gnwh8cl2mfiwwix1joyxlmzkz0tsgn5kni/g6oue8nrhpavyi6nqykqe3ytkehy3zc1qgwfhbbfzliunemq5jtgruwhiyrgt8euaupsczo7sm7cxow9knhr2qdcd0npaumfoamrtp7gey/scgltbo6vqcyxhsa3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6gvuww18qqmse9642+y3r4qv4bkmtpn3ivmc5i+yd63gbvj4xvau2wzdhxrvtj+7hfw/p2dozuiqkwa0elwyxqp9dtsa4n8so9lae8nzajxoz3oseg/buivnpsdqoorhb2ypnlks7j9eyrgygtkjmimz53mf7nk+aaodhuiunfs/oaodk9rpht6uyc5ueybeowyxocc/70um+5jzadvvzlh0ssd4vme+oujtstxh7y/zc+z6ctizzivzzhrjeymtl34tohe/tvjunvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhulxbgvz5vwhvifmhrdijekxpuzhx6p97ossx1ufnftiv+p17ljvmvxpfutpljyn5sgnxzv4hzi5kdg==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoaej5vaea8gcjrnuvnvtrnjrpgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzhytdyvmujvta6sa1yew6gktrvnjrprmrybu9v7fy+r1cvvbs7yxck+x115lvnwg24jqwhqtv8s7fsrqig5+xzjujxtyhxpiuujfagcx2gmyv7vkm6herql3xsfjcg7jffme5tklgrotcduqjdbmmvduudwge2cysfcwvant5nze6nnsccownbo/di4ymaznm2nnfcamoavsrimca/eh6irjntqylmhgefow3gnwh8cl2mfiwwix1joyxlmzkz0tsgn5kni/g6oue8nrhpavyi6nqykqe3ytkehy3zc1qgwfhbbfzliunemq5jtgruwhiyrgt8euaupsczo7sm7cxow9knhr2qdcd0npaumfoamrtp7gey/scgltbo6vqcyxhsa3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6gvuww18qqmse9642+y3r4qv4bkmtpn3ivmc5i+yd63gbvj4xvau2wzdhxrvtj+7hfw/p2dozuiqkwa0elwyxqp9dtsa4n8so9lae8nzajxoz3oseg/buivnpsdqoorhb2ypnlks7j9eyrgygtkjmimz53mf7nk+aaodhuiunfs/oaodk9rpht6uyc5ueybeowyxocc/70um+5jzadvvzlh0ssd4vme+oujtstxh7y/zc+z6ctizzivzzhrjeymtl34tohe/tvjunvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhulxbgvz5vwhvifmhrdijekxpuzhx6p97ossx1ufnftiv+p17ljvmvxpfutpljyn5sgnxzv4hzi5kdg==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoaej5vaea8gcjrnuvnvtrnjrpgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzhytdyvmujvta6sa1yew6gktrvnjrprmrybu9v7fy+r1cvvbs7yxck+x115lvnwg24jqwhqtv8s7fsrqig5+xzjujxtyhxpiuujfagcx2gmyv7vkm6herql3xsfjcg7jffme5tklgrotcduqjdbmmvduudwge2cysfcwvant5nze6nnsccownbo/di4ymaznm2nnfcamoavsrimca/eh6irjntqylmhgefow3gnwh8cl2mfiwwix1joyxlmzkz0tsgn5kni/g6oue8nrhpavyi6nqykqe3ytkehy3zc1qgwfhbbfzliunemq5jtgruwhiyrgt8euaupsczo7sm7cxow9knhr2qdcd0npaumfoamrtp7gey/scgltbo6vqcyxhsa3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6gvuww18qqmse9642+y3r4qv4bkmtpn3ivmc5i+yd63gbvj4xvau2wzdhxrvtj+7hfw/p2dozuiqkwa0elwyxqp9dtsa4n8so9lae8nzajxoz3oseg/buivnpsdqoorhb2ypnlks7j9eyrgygtkjmimz53mf7nk+aaodhuiunfs/oaodk9rpht6uyc5ueybeowyxocc/70um+5jzadvvzlh0ssd4vme+oujtstxh7y/zc+z6ctizzivzzhrjeymtl34tohe/tvjunvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhulxbgvz5vwhvifmhrdijekxpuzhx6p97ossx1ufnftiv+p17ljvmvxpfutpljyn5sgnxzv4hzi5kdg==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoaej5vaea8gcjrnuvnvtrnjrpgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzhytdyvmujvta6sa1yew6gktrvnjrprmrybu9v7fy+r1cvvbs7yxck+x115lvnwg24jqwhqtv8s7fsrqig5+xzjujxtyhxpiuujfagcx2gmyv7vkm6herql3xsfjcg7jffme5tklgrotcduqjdbmmvduudwge2cysfcwvant5nze6nnsccownbo/di4ymaznm2nnfcamoavsrimca/eh6irjntqylmhgefow3gnwh8cl2mfiwwix1joyxlmzkz0tsgn5kni/g6oue8nrhpavyi6nqykqe3ytkehy3zc1qgwfhbbfzliunemq5jtgruwhiyrgt8euaupsczo7sm7cxow9knhr2qdcd0npaumfoamrtp7gey/scgltbo6vqcyxhsa3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6gvuww18qqmse9642+y3r4qv4bkmtpn3ivmc5i+yd63gbvj4xvau2wzdhxrvtj+7hfw/p2dozuiqkwa0elwyxqp9dtsa4n8so9lae8nzajxoz3oseg/buivnpsdqoorhb2ypnlks7j9eyrgygtkjmimz53mf7nk+aaodhuiunfs/oaodk9rpht6uyc5ueybeowyxocc/70um+5jzadvvzlh0ssd4vme+oujtstxh7y/zc+z6ctizzivzzhrjeymtl34tohe/tvjunvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhulxbgvz5vwhvifmhrdijekxpuzhx6p97ossx1ufnftiv+p17ljvmvxpfutpljyn5sgnxzv4hzi5kdg== e1 aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz6wzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hxvvkdq==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz6wzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hxvvkdq==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz6wzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hxvvkdq==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzqigz6wzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hxvvkdq== lt aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzuinzvtzrqg4w5qrnxvkriiquq7czzuk1albitnujz8ubaqhe2rrjpueslskdnsnleaxqxjlxdga1ohtwkrbwok2jxclsxojdtbnzljqdrx4ecig87sn550w2rbmrwpyjpevnh+ia830kagzr9jhainxtrg/qq/phbalkl6xhmaygmlzawrsg8ngrxwd0xcezohnhysrdg1zydnbswybw5b5bsoqykwiaq7ykcevgtcwwjiqyqkmoww+ccnxn+oznv0ldha059knhr2qdcd0npaumxofmrtp7gey/scgltbo6vqcyxhse3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6hpuww18qqmse9642+y3r4qv4bkmtpgpkx7dcefbbdbwak58l3gxbymjjro/2e/fjv4f0bjunlqeccanbpdkf0h+xkzhu5zkkve1neguirr0mb0lpdxtqxfk0jroh0djg7mfms6its7smq4oblzk5ifked/ae5ysg2vkrbufj34sz2rlsvcbr0yrmwqzn8m1gfqnt3po+dlovoae2w362zyceks9lzlsjliv7bdz+mj93vienlcu8ih2wh6yxgnnyn6ezx9m0b3st3uz5vdqeo7a+dt6up53f0hxzi/y5spp/qaxwtldbgvz5vah8xpzdvukrpaanqmhr9p89hedl6qamp+kx/c49yx3karlvqenykse8plerz/8d3ajkvw==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzuinzvtzrqg4w5qrnxvkriiquq7czzuk1albitnujz8ubaqhe2rrjpueslskdnsnleaxqxjlxdga1ohtwkrbwok2jxclsxojdtbnzljqdrx4ecig87sn550w2rbmrwpyjpevnh+ia830kagzr9jhainxtrg/qq/phbalkl6xhmaygmlzawrsg8ngrxwd0xcezohnhysrdg1zydnbswybw5b5bsoqykwiaq7ykcevgtcwwjiqyqkmoww+ccnxn+oznv0ldha059knhr2qdcd0npaumxofmrtp7gey/scgltbo6vqcyxhse3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6hpuww18qqmse9642+y3r4qv4bkmtpgpkx7dcefbbdbwak58l3gxbymjjro/2e/fjv4f0bjunlqeccanbpdkf0h+xkzhu5zkkve1neguirr0mb0lpdxtqxfk0jroh0djg7mfms6its7smq4oblzk5ifked/ae5ysg2vkrbufj34sz2rlsvcbr0yrmwqzn8m1gfqnt3po+dlovoae2w362zyceks9lzlsjliv7bdz+mj93vienlcu8ih2wh6yxgnnyn6ezx9m0b3st3uz5vdqeo7a+dt6up53f0hxzi/y5spp/qaxwtldbgvz5vah8xpzdvukrpaanqmhr9p89hedl6qamp+kx/c49yx3karlvqenykse8plerz/8d3ajkvw==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzuinzvtzrqg4w5qrnxvkriiquq7czzuk1albitnujz8ubaqhe2rrjpueslskdnsnleaxqxjlxdga1ohtwkrbwok2jxclsxojdtbnzljqdrx4ecig87sn550w2rbmrwpyjpevnh+ia830kagzr9jhainxtrg/qq/phbalkl6xhmaygmlzawrsg8ngrxwd0xcezohnhysrdg1zydnbswybw5b5bsoqykwiaq7ykcevgtcwwjiqyqkmoww+ccnxn+oznv0ldha059knhr2qdcd0npaumxofmrtp7gey/scgltbo6vqcyxhse3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6hpuww18qqmse9642+y3r4qv4bkmtpgpkx7dcefbbdbwak58l3gxbymjjro/2e/fjv4f0bjunlqeccanbpdkf0h+xkzhu5zkkve1neguirr0mb0lpdxtqxfk0jroh0djg7mfms6its7smq4oblzk5ifked/ae5ysg2vkrbufj34sz2rlsvcbr0yrmwqzn8m1gfqnt3po+dlovoae2w362zyceks9lzlsjliv7bdz+mj93vienlcu8ih2wh6yxgnnyn6ezx9m0b3st3uz5vdqeo7a+dt6up53f0hxzi/y5spp/qaxwtldbgvz5vah8xpzdvukrpaanqmhr9p89hedl6qamp+kx/c49yx3karlvqenykse8plerz/8d3ajkvw==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tzuinzvtzrqg4w5qrnxvkriiquq7czzuk1albitnujz8ubaqhe2rrjpueslskdnsnleaxqxjlxdga1ohtwkrbwok2jxclsxojdtbnzljqdrx4ecig87sn550w2rbmrwpyjpevnh+ia830kagzr9jhainxtrg/qq/phbalkl6xhmaygmlzawrsg8ngrxwd0xcezohnhysrdg1zydnbswybw5b5bsoqykwiaq7ykcevgtcwwjiqyqkmoww+ccnxn+oznv0ldha059knhr2qdcd0npaumxofmrtp7gey/scgltbo6vqcyxhse3ouwa/4jko5hh/l2stjsjug0pws9nq8nuwex9g+0fc6hpuww18qqmse9642+y3r4qv4bkmtpgpkx7dcefbbdbwak58l3gxbymjjro/2e/fjv4f0bjunlqeccanbpdkf0h+xkzhu5zkkve1neguirr0mb0lpdxtqxfk0jroh0djg7mfms6its7smq4oblzk5ifked/ae5ysg2vkrbufj34sz2rlsvcbr0yrmwqzn8m1gfqnt3po+dlovoae2w362zyceks9lzlsjliv7bdz+mj93vienlcu8ih2wh6yxgnnyn6ezx9m0b3st3uz5vdqeo7a+dt6up53f0hxzi/y5spp/qaxwtldbgvz5vah8xpzdvukrpaanqmhr9p89hedl6qamp+kx/c49yx3karlvqenykse8plerz/8d3ajkvw== yt aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tztifnomzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hoh1kza==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tztifnomzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hoh1kza==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tztifnomzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hoh1kza==aaahqhiclvxbbtnaej22qnpya+grl4i0cagtkoiej5fsea8gcjrnuvnvtrnjrfgs2zusyou3eixf4l944mx4e5o4l9awvbozc87cdm2747mxlhb/lc2v3lh5k7e6tn77zt179zc2hxzfytdyvnujvta6tq1yew6gqtrvnjrurmrybu/v7haf12sxkordmdju/y469a1w4dzdx9jq1eu+pc/tztifnomzjujxtyhxpiuujfagcx2emysvqlodqnkosz4pckhd9siigpcjlahihehokyeugutkuqibrqpbhzwchqvtg+8wzidgg2donlgghxjx8era5mkbzzthtghnxhxkgonfpd9e15rpirfmjrcp0qbjmjb+hf7toswwix1joyxlmzkz0tskn5kni/g6oue8nrhpplyi6nqykqe3ytkchy3zc1qgwfhfbfzliunemm5gtgruwhiyrgt8euaupsczo7sw7cxoz6vppxoh647pbcxdy/q+ngmf2c8w+09akmhq0qkeljgkjuzcsr7wsuzzpt6wtynhlwompctz7fv4pdk9vrb9ok90cothtr5uuik9711t9lvwwxfw2vjbtrxivmc5i+yd63gbvj4xvau2wzdhxr/tj+7hfw/p2sobxeif4aacsrmlop8uj2pcz2uue9ub8nzejxoz3oseg/buivnpsjqoorhb2ypnlkrblmgzvwcdwyrzebm87ua9zldatpmj3mktvhdntholeo2jp1xmhu3p5nvaleanuod96xrfck5th/1syw4jjj+xmhdhxer22rj9zx7ufe9ow5z5efssd1kvmbllb0jnjqdp3vymvm3zehu8d2b96nwotzulofnmrvhzlsb/u1nhag+3bpnzq0l5ifmhrdijekxpuzhx6p97oscx1ufnftiv+p17ljvi1xlfutpljyn5sgnxzv4hoh1kza== figure 4: the architecture of the proposed bilstm model with self attention. for each time t, the exchange level parameter of all exchanges ei of the sub-dialogue i ∈ {1 . . . t} are encoded to their respective hidden representation hi and are considered and weighted with the self attention mechanism to finally estimate the iq value yt at time t. as shown in figure 4, the exchange level parameters form the input vector et for each time step or turn t to a bi-directional lstm (graves et al., 2013) layer. the input vector et encodes the nominal parameters asrrecognitionstatus, activitytype, and confirmation? as 1-hot representations. in the bilstm layer, two hidden states are computed: ~ht constitutes the forward pass through the current sub-dialogue and ~ht the backwards pass: ~ht = lstm(et, ~ht−1) (3) ~ht = lstm(et, ~ht+1) (4) the final hidden layer is then computed by concatenating both hidden states: ht = [~ht, ~ht] . (5) even though information from all time steps may contribute to the final iq value, not all time steps may be equally important. thus, an attention mechanism (vaswani et al., 2017) is used that 89 ultes and maier evaluates the importance of each time step t′ for estimating the iq value at time t by calculating a weight vector αt,t′ . gt,t′ = tanh(ht t wt + ht t′wt′ + bt) (6) αt,t′ = softmax(σ(wagt,t′ + ba)) (7) lt = ∑ t′ αt,t′ht′ (8) zheng et al. (2018) describe this as follows: “the attention-focused hidden state representation lt of an [exchange] at time step t is given by the weighted summation of the hidden state representation ht′ of all [exchanges] at time steps t′, and their similarity αt,t′ to the hidden state representation ht of the current [exchange]. essentially, lt dictates how much to attend to an [exchange] at any time step conditioned on their neighbourhood context.” to calculate the final estimate yt of the current iq value at time t, a softmax layer is introduced: yt = softmax(lt) (9) for estimating the interaction quality using a bilstm, the proposed architecture frames the task as a classification problem where each sequence is labelled with one iq value. thus, for each time step t, the iq value needs to be estimated for the corresponding sub-dialogue consisting of all exchanges from the beginning up to t. framing the problem like this is necessary to allow the application of a bilstm-approach while still being able to only use information that would be present at the current time step t in an ongoing dialogue interaction. to analyse the influence of the bilstm, a model with a single forward-lstm layer is also investigated where ht = ~ht . (10) similarly, a model without attention is also analysed where lt = ht . (11) a deep learning approach using only non-temporal features has been previously proposed and achieved an uar on the full feature set of 0.55 (rach et al., 2017). 4. simulation experiments and results the iq estimators are both trained and evaluated on the lego corpus and applied within the iq reward estimation framework (figure 1) on several domains within a simulated environment. 4.1 interaction quality estimation to evaluate the bilstm model with attention (bilstm+att), it is compared with three of its own variants: a bilstm without attention (bilstm) as well as a single forward-lstm layer with attention (lstm+att) and without attention (lstm). an additional baseline is defined by rach et al. (2017) who already proposed an lstm-based architecture that only uses non-temporal features. furthermore, the conventional iq estimator using a linear svm is evaluated as originally used for reward estimation by ultes et al. (2015). 90 user satisfaction reward estimation across domains table 3: performance of the proposed lstm-based variants with the traditional cross-validation setup. due to overlapping sub-dialogues in the train and test sets, the performance of the lstm-based models achieve unrealistically high performance. significant differences are observed between bilstm and all other variants/models (p < 0.05, wilcoxon signedrank test (wilcoxon, 1945)). uar κ ρ ea ep. lstm 0.78 0.85 0.91 0.99 101 bilstm 0.78 0.85 0.92 0.99 100 lstm+att 0.74 0.82 0.91 0.99 101 bilstm+att 0.75 0.83 0.91 0.99 93 rach et al. (2017) 0.55 0.68 0.83 0.94 ultes et al. (2015) 0.55 0.89 the deep neural net models have been implemented with keras4 using the self-attention implementation as provided by zheng et al. (2018)5. the input vector consists of one-hot-encodings of the three nominal features asrrecognitionstatus, activitytype, and confirmation? and the numerical features asrconfidence and reprompt? resulting in a vector of size 11. the target iq labels are also one-hot-encoded. the lstm and bilstm embeddings have a dimension of 64 and 128, respectively. the maximum dialogue length has been set to 60: dialogues with fewer turns are padded wiht zero vectors and dialogues with more turns are truncated. all models were trained against cross-entropy loss using rmsprop (tieleman and hinton, 2012) optimisation with a learning rate of 0.001 and a mini-batch size of 32. interaction quality estimation is evaluated by using three commonly applied evaluation metrics: unweighted average recall (uar), cohen’s kappa (cohen, 1960), and spearman’s rho (spearman, 1904) comparing the estimated iq ratings x with the true iq ratings y. as missing the correct estimated iq value by only one has little impact for modelling the reward, a measure we call the extended accuracy (ea) is used where neighbouring values are taken into account as well. recall in general is defined as the rate of correctly classified samples belonging to one class. the unweighted average recall for multi-class classification problems with c classes is computed by the class-wise recalls recallc for each class c and then averaged over all class-wise recalls: uar = 1 c n∑ c=1 recallc . (12) cohen’s kappa measures the relative agreement between two corresponding sets of ratings, here the estimate x and ground truth y. in our case, cohen’s weighted kappa is applied as ordinal scores are compared (cohen, 1968): κ = 1− ∑c c=1 ∑c k=1wck ·mck∑c c=1 ∑c k=1wck · mc.·m.k n . (13) 4. https://keras.io/ 5. code freely available at https://github.com/cyberzhg/keras-self-attention 91 ultes and maier mc,k is the number of samples where, for a corresponding (xi, yi) pair, the estimator predicted k for the true class c, and n is the total number of samples. mc. represents the sum of all estimated class samples for true class c and m.k the sum of all true class samples for estimated class k. a weighting factor w is introduced reducing the discount of disagreements the smaller the difference is between two ratings: wck = |rc − rk| |rmax − rmin| . (14) here, rc and rk denote the rating pair and rmax and rmin the maximum and minimum ratings possible. the correlation of two variables describes the degree by that one variable can be expressed by the other. spearman’s rank correlation coefficient is a non-parametric method assuming a monotonic function between the two variables (spearman, 1904). it is defined by ρ = ∑ i(xi − x̄)(yi − ȳ)√∑ i(xi − x̄)2 ∑ i(yi − ȳ)2 , (15) where xi and yi are corresponding ranked ratings and x̄ and ȳ the mean ranks. thus, two sets of ratings can have total correlation even if they never agree. this would happen if all ratings are shifted by the same value, for example. the extended accuracy is computed similar to regular accuracy. however, instead of only using the main diagonal of the confusion matrix, the two secondary diagonals are considered additionally. for n classes, the extended accuracy is computed by ea = ∑c c=1mc,c +mc,c−1 +mc,c+1∑n c=1 ∑n k=1mc,k , (16) where mc,k is the number of samples where the estimator predicted k for the true class c. thus, neighbouring values are taken into account as well and being off by one is still considered as correct estimation6. all experiments were conducted with the lego corpus (schmitt et al., 2012a) in a 10-fold cross-validation setup for a total of 100 epochs per fold. the results are presented in table 3. due to the way the task is framed (one label for each sub-dialogue), memorising effects may be observed with the traditional cross-validation setup that has been used in previous work. hence, the results in table 3 show very high performance, which is likely to further increase with ongoing training. however, the corresponding models are likely to generalise poorly. to alleviate this, a dialogue-wise cross-validation setup has been employed also consisting of 10 folds of disjoint sets of dialogues. by that, it can be guaranteed that there are no overlapping subdialogues in the training and test sets. all results of these experiments are presented in table 4 with the absolute improvement of the two main measures uar and ea over the svm-based conventional approach of ultes et al. (2015) visualised in figure 5. the bilstm+att model outperforms the other models and baselines in all four performance measures by achieving an uar of 0.54 and an ea of 0.94 after 40 epochs. furthermore, both the bilstm and the attention mechanism by themselves improve the performance in terms of uar. 92 user satisfaction reward estimation across domains table 4: performance of the proposed lstm-based variants with the dialogue-wise crossvalidation setup. the models by rach et al. (2017) and ultes et al. (2015) have been re-implemented. the bilstm with attention mechanism performs best in all evaluation metrics. all results are significantly different to each other (p < 0.05, wilcoxon signedrank test (wilcoxon, 1945)). uar κ ρ ea ep. lstm 0.51 0.63 0.78 0.93 8 bilstm 0.53 0.63 0.78 0.93 8 lstm+att 0.52 0.63 0.79 0.92 40 bilstm+att 0.54 0.65 0.81 0.94 40 rach et al. (2017) 0.45 0.58 0.79 0.88 82 ultes et al. (2015) 0.44 0.53 0.69 0.86 +0 .0 7 +0 .0 7+0 .0 9 +0 .0 7 +0 .0 8 +0 .0 6 +0 .1 0 +0 .0 8 +0 .0 1 +0 .0 2 0.00 +0.02 +0.04 +0.06 +0.08 +0.10 +0.12 uar ea ab so lu te im pr ov em en t lstm bilstm lstm+att bilstm+att rach et al. (2017) figure 5: absolute improvement of the iq estimation models over the conventional model originally proposed by (ultes et al., 2015) for iq-based reward estimation with the dialoguewise cross-validation setup. uar and ea take values from 0.0 to 1.0. 4.2 dialogue policy learning to analyse the impact of the iq reward estimator on the resulting dialogue policy, experiments are conducted comparing three different reward models. the baseline is in accordance to ultes et al. (2017a): having the objective task success as principal reward component (rts). it will be compared to the conventional estimator having the interaction quality estimated by a support vector machine as principal reward component (rs iq) and to the bilstm+att model to estimate the interaction quality used as principal reward component (rbi iq). ts can be computed by comparing 6. for the bounds, the respective non-existing mc,k is omitted. 93 ultes and maier 0.00 0.20 0.40 0.60 0.80 1.00 0% 15% 30% 0% 15% 30% 0% 15% 30% 0% 15% 30% 0% 15% 30% 0% 15% 30% cr ch sr sh l all ta sk s uc ce ss r at e reward using iq (bilstm) reward using iq (svm) reward using ts figure 6: results using gp-sarsa in task success rate (tsr) of the simulated experiments for all domains and semantic error rates 0%, 15%, and 30%. each value is computed after 100 evaluation / 1,000 training dialogues averaged over three trials. numerical results with significance indicators are shown in table 6. 0.00 0.20 0.40 0.60 0.80 1.00 0% 15% 30% 0% 15% 30% 0% 15% 30% 0% 15% 30% 0% 15% 30% 0% 15% 30% cr sr all cr sr all dqn enac ta sk s uc ce ss r at e reward using iq (bilstm) reward using iq (svm) reward using ts figure 7: results using dqn and enac in task success rate (tsr) of the simulated experiments for camrestaurants and sfrestaurants and semantic error rates 0%, 15%, and 30%. each value is computed after 100 evaluation / 1,000 training dialogues averaged over three trials. numerical results with significance indicators are shown in table 8 and table 7. the outcome of each dialogue with the pre-defined goal. of course, this is only possible in simulation and when evaluating with paid subjects. this goal information is not available to the iq estimators, nor is it required. for learning the dialogue behaviour, three policy models are applied. the first is a policy model based on the gp-sarsa algorithm (gašić and young, 2014), which is a value-based method that 94 user satisfaction reward estimation across domains table 5: statistics of the domains the interaction quality estimators are trained on (letsgo) and applied to (rest). domain code # constraints # db items state size letsgo 4 camrestaurants cr 3 110 268 camhotels ch 5 33 111 sfrestaurants sr 6 271 636 sfhotels sh 6 182 438 laptops l 6 126 204 uses a gaussian process to approximate the state-value function. as it takes into account the uncertainty of the approximation, it is very sample efficient and may even be used to learn a policy directly through real human interaction (gašić et al., 2013). additionally, two deep reinforcement learning algorithms are applied: deep q-network (dqn) (mnih et al., 2015) and episodic natural actor critic (enac) (su et al., 2017). similar to the gpsarsa, the dqn also approximates the state-value function also known as q-function. the enac directly learns the policy using a policy gradient approach in an actor-critic framework. the decisions of the policy are based on a summary space representation of the dialogue state tracker. in this work, the focus tracker (henderson et al., 2014)—an effective rule-based tracker— is used. the state space and summary state space follow previous work (e.g., gašić et al., 2013) and comprise multiple dimensions: for each informable slot one probability distribution over all slot values plus the special values dontcare and none, for each requestable slot a bernoulli distribution indicating whether the slot has been requested by the user, a probability distribution over the query method (e.g., search by constraints, search by name), and a status vector of the search results given the current state. informable slots contain all information that the user has provided to the system as search constraints during the interaction. requestable slots are usually a superset of the informable slots and contain additionally information that is part of the search result (e.g., the phone number). to map this dialogue state to a summary space, each of the informable slot probability distribution vectors are sorted (excluding none). this follows the idea that for making a decision, the actual slot value is not important, only how the probabilities are distributed over all values. for each dialogue decision, the policy chooses exactly one summary action out of a set of summary actions, which are based on general dialogue acts like request, confirm or inform. the exact number of system actions varies for the domains and ranges from 16 to 25. to measure the dialogue performance, the task success rate (tsr) and the average interaction quality (aiq) are measured: the tsr represents the ratio of dialogues for which the system was able to provide the correct result. aiq is calculated based on the estimated iq values of the respective model (aiqbi) for the bilstm and aiqs for the svm) at the end of each dialogue. as there are two iq estimators, a distinction is made between aiqs and aiqbi. additionally, the average dialogue length (adl) is reported. 95 ultes and maier table 6: results of the simulated experiments for all domains using gp-sarsa showing task success rate (tsr), average interaction quality estimated with the svm (aiqs) and the bilstm (aiqbi), and average dialogue length (adl) in number of turns. each value is computed after 100 evaluation / 1,000 training dialogues averaged over three trials with different random seeds. 1,2,3 marks statistically significant difference compared to rts , to rs iq, and to rbi iq, respectively (p < 0.05, t-test for tsr and adl, mann-whitney-u test for aiq). domain ser tsr aiqs aiqbi adl rts rs iq rbi iq rts rs iq rts rbi iq rts rs iq rbi iq cr 0% 1.002,3 0.991 0.991 3.642 3.901 3.683 3.831 4.68 4.88 4.59 15% 0.97 0.94 0.96 3.352 3.651 3.453 3.631 5.853 5.33 5.101 30% 0.94 0.92 0.90 3.152 3.341 3.22 3.30 6.34 6.30 6.25 ch 0% 0.98 0.99 0.99 3.262 3.621 3.33 3.44 5.71 5.61 5.40 15% 0.962 0.891,3 0.932 2.90 2.88 3.14 3.14 6.282 7.261,3 6.312 30% 0.86 0.88 0.87 2.382 2.791 2.793 3.021 7.943 7.31 6.991 sr 0% 0.98 0.97 0.98 3.042 3.531 3.133 3.371 6.26 6.03 5.80 15% 0.903 0.88 0.841 2.402 3.001 2.853 3.011 7.99 7.55 7.33 30% 0.71 0.77 0.78 2.032 2.521 2.463 2.781 9.773 9.41 8.501 sh 0% 0.97 0.99 0.98 3.152 3.521 3.173 3.361 5.992 5.501 5.76 15% 0.88 0.88 0.89 2.632 2.941 2.773 3.171 7.983 7.593 6.631,2 30% 0.832 0.761 0.80 2.50 2.63 2.703 2.871 8.38 9.21 8.37 l 0% 0.98 0.99 0.99 3.262 3.611 3.28 3.41 5.78 5.44 5.60 15% 0.89 0.88 0.92 2.582 2.971 2.923 3.171 7.19 7.34 6.73 30% 0.80 0.74 0.77 2.43 2.57 2.79 2.92 8.222 9.321,3 7.972 all 0% 0.98 0.98 0.98 3.232 3.651 3.31 3.48 5.76 5.50 5.47 15% 0.92 0.89 0.91 2.762 3.101 3.023 3.201 7.13 7.06 6.52 30% 0.83 0.81 0.82 2.49 2.80 2.78 2.97 8.202 8.231,3 7.662 for the simulation experiments with the gp-sarsa, the performance of the trained policies on five different domains was evaluated: cambridge hotels and restaurants, san francisco hotels and restaurants, and laptops. the complexity of each domain is shown in table 5 and compared to the letsgo domain (the domain the estimators have been trained on). the dqn and enac are evaluated only on the two domains cambridge (most simple) and san francisco restaurants (most complex). the dialogues were created using the publicly available spoken dialogue system toolkit pydial (ultes et al., 2017b)7, which contains implementations of all the applied policy models8 and an implementation of the agenda-based user simulator (schatzmann and young, 2009) with an ad7. code freely available at http://www.pydial.org 8. please refer to pydial code for details about the exact model implementation. 96 user satisfaction reward estimation across domains table 7: results of the simulated experiments for cr and sfr using dqn showing task success rate (tsr), average interaction quality estimated with the svm (aiqs) and the bilstm (aiqbi), and average dialogue length (adl) in number of turns. each value is computed after 100 evaluation / 1,000 training dialogues averaged over three trials with different random seeds. 1,2,3 marks statistically significant difference compared to rts , to rs iq, and to rbi iq, respectively (p < 0.05, t-test for tsr and adl, mann-whitney-u test for aiq). domain ser tsr aiqs aiqbi adl rts rs iq rbi iq rts rs iq rts rbi iq rts rs iq rbi iq cr 0% 0.983 0.96 0.951 3.692 3.171 3.52 3.51 4.322 4.681 4.53 15% 0.87 0.87 0.86 3.05 3.11 3.39 3.42 5.37 5.32 5.14 30% 0.76 0.82 0.77 2.652 2.941 3.34 3.28 6.04 5.59 5.74 sr 0% 0.682,3 0.531 0.591 2.572 2.011 3.10 3.07 6.29 6.56 6.56 15% 0.59 0.62 0.57 2.20 2.18 3.10 3.08 6.582 7.101 6.85 30% 0.462 0.381 0.39 1.85 1.71 3.013 3.081 7.552 8.791,3 7.662 all 0% 0.833 0.75 0.771 3.132 2.591 3.31 3.29 5.302 5.621 5.54 15% 0.73 0.75 0.72 2.63 2.65 3.25 3.25 5.98 6.21 5.99 30% 0.61 0.60 0.58 2.252 2.321 3.17 3.18 6.79 7.19 6.70 table 8: results of the simulated experiments for cr and sfr using enac showing task success rate (tsr), average interaction quality estimated with the svm (aiqs) and the bilstm (aiqbi), and average dialogue length (adl) in number of turns. each value is computed after 100 evaluation / 1,000 training dialogues averaged over three trials with different random seeds. 1,2,3 marks statistically significant difference compared to rts , to rs iq, and to rbi iq, respectively (p < 0.05, t-test for tsr and adl, mann-whitney-u test for aiq). domain ser tsr aiqs aiqbi adl rts rs iq rbi iq rts rs iq rts rbi iq rts rs iq rbi iq cr 0% 0.973 0.993 0.741,2 3.312 3.651 3.45 3.41 4.783 4.783 5.821,2 15% 0.912,3 0.741,3 0.341,2 3.25 3.16 3.463 3.181 5.181,2 7.721,3 8.711,2 30% 0.733 0.763 0.391,2 2.78 2.98 3.293 3.141 6.491,2 7.231,3 8.771,2 sr 0% 0.952,3 0.721,3 0.601,2 2.52 2.65 3.213 3.111 5.382,3 7.401,3 8.471,2 15% 0.802,3 0.571,3 0.401,2 2.23 2.35 3.133 3.071 6.672,3 9.391,3 7.901,2 30% 0.632,3 0.511 0.471 1.712 1.941 3.03 3.03 8.602,3 10.011 9.381 all 0% 0.963 0.853 0.671,2 2.922 3.151 3.33 3.26 5.083 6.093 7.151,2 15% 0.862,3 0.651,3 0.371,2 2.74 2.76 3.293 3.131 5.932,3 8.561,3 8.311,2 30% 0.683 0.643 0.431,2 2.25 2.46 3.163 3.091 7.542,3 8.621,3 9.081,2 97 ultes and maier ditional error model. the error model simulates the required semantic error rate (ser) caused in a real system by the noisy speech channel. for each domain, all three reward models are compared on three sers: 0%, 15%, and 30%. more specifically, the applied evaluation environments are based on env. 1, env. 3, and env. 6, respectively, as defined by casanueva et al. (2017). these environments also contain all parameters used for the training of gp-sarsa, dqn and enac implementations of pydial. for each domain and for each ser, policies have been trained using 1,000 dialogues followed by an evaluation step of 100 dialogues. the task success rates for gp-sarsa in figure 6 and dqn and enac in figure 7 with exact numbers shown in table 6, table 7, and table 8, respectively, were computed based on the evaluation step averaged over three train/evaluation cycles with different random seeds. the results of training the gp-sarsa with the svm iq reward estimator are similar in terms of tsr for rs iq and rts in all domains for an ser of 0%. this finding is even stronger when comparing rbi iq and rts . these high tsrs are achieved while having the dialogues of both iqbased models result in higher aiq values compared to rts throughout the experiments. of course, only the iq-based model is aware of the iq concept and indeed is trained to optimise it. for higher sers, the tsrs lightly degrade for the iq-based reward estimators. however, there seems to be a tendency that the tsr for rbi iq is more robust against noise compared to rs iq while still resulting in better aiq values. finally, even though the differences are mostly not significant, there is also a tendency for rbi iq to result in shorter dialogues compared to both rs iq and rts . the results of training a dqn or an enac policy model are different, though. while the dqn shows similar tsrs for all sers and reward models in the cr domain, rts clearly shows best performance in the sfr domain. furthermore, even when using an iq-based reward model, the respective aiq does not result in higher scores than usingrts which shows the stability problems that come with the usage of deep reinforcement learning. the enac policy model is even more prone to these effects having rts always resulting in the highest tsr with rs iq and rbi iq performing rather poorly. 5. analysis of learned behaviour to further analyse the learning behaviour of the different reward estimators and policy models and thus to gain deeper insights, the similarity scoring framework (ultes and maier, 2020) is applied. it uses a standardised setup to feed a fixed set of dialogue states into each trained policy model and compares the resulting system actions and quantifies their similarities. three different metrics are computed: the total match rate (tmr), the dialogue act match rate (dmr) and the concept match rate (cmr). a similarity score is computed for the comparison of two behavioural models π and π′. depending on the nature of the behavioural model, for each context ci ∈ c, each may produce an abstract system response actions ai, and an text response pi. each abstract system action ai = acti(s i 1 = vi1, . . . , s i j = vij) consists of a dialogue act acti, representing the communicative function like inform or request, and a set si of j slot-value-pairs si = {(si1, vi1), . . . , (sij , vij)} 98 user satisfaction reward estimation across domains representing the concepts and their respective values9. to compute each similarity score, |c| action/text response pairs are compared using the following similarity score measures. according to ultes and maier (2020), the metrics are defined as follows: the total match rate (tmr) is based on a binary score that regards two actions a, a′ as equal only if they completely match, i.e., δa,a′ = 1 iff a = a′, else 0. the tmr is then defined by tmr = 1 |c| |c|∑ i=1 δai,a′i . (17) the dialogue act match rate (dmr) is based on a binary score comparing the actions a, a′ where both match if the corresponding dialogue acts are the same: δact ,act ′ = 1 iff act = act ′, else 0. the dmr is defined by dmr = 1 |c| |c|∑ i=1 δacti,act ′i . (18) the symmetric concept match rate (cmr) counts concepts γ that are present in both dialogue actions where m̃(a, a′, γ) defines if a match occurred: m̃(a, a′, γ) = { 1, if γ ∈ s and γ ∈ s′ 0, otherwise . (19) the concept match cm takes into account the dialogue acts and the unified set of concepts s̃ = s1 ∪ s2 of both dialogue actions treating slots s and values v in s̃ as individual γ: c̃m (a, a′) = δact ,act ′ + ∑ ∈s̃ m̃(a, a′, γ) (20) a concept match of two dialogue actions a and a′ is thus defined by cm (a, a′) = c̃m (a, a′) 1 + |s̃| (21) and the concept match rate by cmr = 1 |c| |c|∑ i=1 cm (a, a′) . (22) in short, the tmr only evaluates to true in case of a complete match of both actions. the dmr evaluates to true already if only the dialogue acts of the system actions are equal. the cmr counts concepts that are present in both dialogue actions and normalises this count by the total number of distinct concepts in both actions. using this analysis setup, the following questions are addressed within this section: 1. how dependent is the resulting policy on the chosen random seed? 9. for the abstract system action a = inform(name=’golden house’, area=centre), act = inform and s = {(name,’golden house’),(area,centre)}. 99 ultes and maier 2. how similar are learned policies across different reward models? 3. how similar are learned policies across different policy models? for analysing the learned behaviour and computing similarity scores within and between learning setups, the camrestaurants domain is used. 5.1 consistency of policies as each combination of reward model, policy model and semantic error rate has been trained with three random seeds, this section addresses the question about the consistency of the resulting learned behaviour. thus, for each setup, the policy of each random seed has been evaluated with the dialogue states from 100 evaluation dialogues taken from one of the policies. a different set of states is used for each noise level. each random seed’s policy is then compared with the policy of the two other random seeds resulting in a total of three comparisons. for each policy model and reward model, the similarity scores of these three comparisons are averaged. the full scores are shown in the appendix in table 12. the similarity scores shown in table 9 show the average scores for tmr, dmr, and cmr. in the left table, the scores are averaged over the different reward models to identify which policy model offers the most consistency during training, i.e., ends up in similar behaviour independent of the random seed. overall, enac shows the most consistency in all three metrics. for 0% ser, enac seems to be less dependent on the random seed. the gp-sarsa shows average consistency with the worst cmr for 0% and 15% ser. on the right side of table 9, the scores are averaged over the policy models to identify which reward model offers the most consistency during training. here, rts shows overall good performance but also rs iq and rbi iq are not fare off and in some cases even more consistent than rts , e.g., tmr with 30% ser. 5.2 similarity between reward models to gain deeper understanding of how similar the behaviour of the trained policies are that originate from the different reward models, this section compares policies from different reward models resulting in pair-wise comparisons of the policy of each random seed of one reward model with the policies of each random seed of another reward model within one policy model. with three random seeds each, this results in nine comparisons. as in the previous section, each policy is evaluated with 100 evaluation dialogues taken from one of the policies. for each policy model and reward model comparison, the similarity scores are averaged. the full scores are shown in the appendix in table 13. the similarity scores shown in table 10 show the average scores for tmr, dmr, and cmr. in the left table, the scores are averaged over the different reward model pairs to identify which policy model results in the most similar behaviour during training, i.e., ends up in similar behaviour independent of the reward model. overall, dqn shows the highest similarity in all three metrics. for 0% ser, gp-sarsa seems to result in more similar behaviour independent for the reward model. on the right side of table 10, the scores are averaged over the policy models to identify which reward model pair results in the highest similarity in behaviour independent of the chosen policy model. overall, rts – rbi iq shows the highest similarity. interestingly, the low score of rs iq – rbi iq 100 user satisfaction reward estimation across domains table 9: the similarity scores analysing the consistency of policies. on the left side, the mean match rates for different policy models is shown averaged over the respective reward models. the right side shows the mean match rates for different reward models averaged over the respective policy models. ser type tmr dmr cmr 0% gp 0.46 0.72 0.44 dqn 0.42 0.73 0.50 enac 0.55 0.83 0.64 15% gp 0.52 0.73 0.48 dqn 0.50 0.76 0.42 enac 0.49 0.72 0.62 30% gp 0.45 0.64 0.52 dqn 0.52 0.70 0.45 enac 0.42 0.69 0.57 all gp 0.48 0.70 0.48 dqn 0.48 0.73 0.46 enac 0.49 0.75 0.61 ser type tmr dmr cmr 0% rts 0.51 0.79 0.58 rs iq 0.47 0.75 0.46 rbi iq 0.46 0.73 0.55 15% rts 0.49 0.77 0.57 rs iq 0.48 0.71 0.54 rbi iq 0.54 0.73 0.41 30% rts 0.43 0.68 0.52 rs iq 0.51 0.69 0.53 rbi iq 0.44 0.66 0.48 all rts 0.48 0.74 0.56 rs iq 0.49 0.72 0.51 rbi iq 0.48 0.71 0.48 shows that both reward models result in quite different behaviour even though both models estimate the same quantity, i.e., the interaction quality. 5.3 similarity between policy models to analyse the behaviour of the trained policies that originate from the different policy models, this section compares policies from different policy models resulting in pair-wise comparisons of the policy of each random seed of one policy model with the policies of each random seed of another policy model within one reward model. with three random seeds each, this results in nine comparisons. as in the previous sections, each policy is evaluated with 100 evaluation dialogues taken from one of the policies. for each policy model and reward model comparison, the similarity scores are averaged. the full scores are shown in the appendix in table 14. the similarity scores shown in table 11 show the average scores for tmr, dmr, and cmr. in the left table, the scores are averaged over the different reward models to identify which policy model pair learns the most similar behaviour, i.e., ends up in similar behaviour independent of the reward model. overall, gp – enac shows the highest similarity in all three metrics. for 15% and 30% ser, the two deep reinforcement learning models dqn – enac show rather low similarity in behaviour. on the right side of table 11, the scores are averaged over the policy model pairs to identify which reward model results in the highest similarity among the policy models. overall, rs iq seems to be the strongest learning signal resulting in the most similar behaviour between policy models. the similarity resulting from using rs ts , instead, is rather low. 101 ultes and maier table 10: the similarity scores analysing the similarity between reward models. on the left side, the mean match rates for different policy models is shown averaged over the respective reward model pairs. the right side shows the mean match rates for different reward model pairs averaged over the respective policy models. ser type tmr dmr cmr 0% gp 0.42 0.62 0.47 dqn 0.38 0.59 0.42 enac 0.37 0.53 0.39 15% gp 0.33 0.61 0.38 dqn 0.40 0.64 0.46 enac 0.37 0.58 0.42 30% gp 0.36 0.52 0.42 dqn 0.41 0.60 0.47 enac 0.40 0.56 0.45 all gp 0.37 0.59 0.42 dqn 0.40 0.61 0.45 enac 0.38 0.56 0.42 ser type tmr dmr cmr 0% rts – rs iq 0.34 0.54 0.41 rts – rbi iq 0.53 0.71 0.59 rs iq – rbi iq 0.29 0.48 0.27 15% rts – rs iq 0.37 0.58 0.43 rts – rbi iq 0.47 0.70 0.53 rs iq – rbi iq 0.27 0.55 0.31 30% rts – rs iq 0.34 0.48 0.41 rts – rbi iq 0.52 0.71 0.59 rs iq – rbi iq 0.30 0.49 0.34 all rts – rs iq 0.35 0.53 0.42 rts – rbi iq 0.51 0.71 0.57 rs iq – rbi iq 0.29 0.51 0.31 table 11: the similarity scores analysing the similarity between policy models. on the left side, the mean match rates for different policy model pairs is shown averaged over the respective reward models. the right side shows the mean match rates for different reward models averaged over the respective policy model pairs. ser type tmr dmr cmr 0% gp – dqn 0.23 0.41 0.27 gp – enac 0.35 0.52 0.39 dqn – enac 0.33 0.53 0.41 15% gp – dqn 0.27 0.50 0.35 gp – enac 0.30 0.54 0.38 dqn – enac 0.25 0.53 0.32 30% gp – dqn 0.38 0.54 0.45 gp – enac 0.33 0.51 0.42 dqn – enac 0.34 0.51 0.39 all gp – dqn 0.29 0.48 0.36 gp – enac 0.33 0.52 0.40 dqn – enac 0.31 0.52 0.37 ser type tmr dmr cmr 0% rts 0.29 0.48 0.38 rs iq 0.30 0.50 0.37 rbi iq 0.31 0.48 0.33 15% rts 0.26 0.52 0.35 rs iq 0.31 0.51 0.38 rbi iq 0.25 0.54 0.32 30% rts 0.26 0.43 0.34 rs iq 0.51 0.68 0.62 rbi iq 0.28 0.44 0.31 all rts 0.27 0.48 0.36 rs iq 0.38 0.56 0.45 rbi iq 0.28 0.49 0.32 102 user satisfaction reward estimation across domains 0 0.2 0.4 0.6 0.8 1 0 100 200 300 400 moving task success rate ts iq us 1 2 3 4 5 0 100 200 300 400 moving average interaction quality ts iq us 1 2 3 4 5 6 0 100 200 300 400 moving average user satisfaction ts iq us figure 8: moving tsr (left), moving aiqs (middle) and moving aus (right) for using either rts , rs iq, or rus as reward averaged over two policies respectively, computed on windows consisting of 120 dialogues. 6. learning from real humans for learning a policy directly from the interaction with real humans, the iq-based policy was only learned using rs iq using gp-sarsa for dialogues in the cr domain10. subjects were recruited through amazon mechanical turk to talk to a telephone-based dialogue system. at the end of each dialogue, users were asked two questions. the first was a yes/no question targeting the dialogue success (“have you found all the information you were looking for?”) and has been used as the baseline for rts . as this label is noisy, only the dialogues where this subjective success label matches the objective success were used for policy training (gašić et al., 2013) (obj = subj). a second baseline was also included: directly acquiring a user satisfaction (us) rating from the users after each dialogue. for this, the second question posed was: “how satisfied are you with the interaction?” the users were able to respond on a six-point scale: 6=very satisfied, 5=satisfied, 4=generally ok, 3=unsatisfied, 2=very unsatisfied or 1=extremely unsatisfied. this rating was converted to a reward in correspondence with rs iq: rus = t · (−1) + (us − 1) · 5 . (23) hence, each dialogue was also evaluated using the average user satisfaction (aus). two policies were trained for each reward function. the learning curves show moving tsr, movingaiqs and moving aus and are presented in figure 8. each value in the graphs is calculated using the respective previous 120 training dialogues averaged over the three trials. the results show that rs iq may successfully be used for learning through interactions with real users. it achieves the same tsr as with rts while resulting in better results for aiq. the results for moving aus also indicate a slightly better user satisfaction compared to rts . the second baseline of using rus also resulted in competitive results showing a similar tsr as rs iq and rts whilst resulting in slightly better aus. this indicates that for a problem as defined by the cr domain, using user satisfaction directly as a learning signal is also a viable option for learning policies. 10. these experiments and results have previously been published by ultes et al. (2017a). 103 ultes and maier sys welcome to the cambridge restaurant dialogue system. how may i help you? usr i’m looking for european food in the expensive price range sys welcome to the cambridge restaurant dialogue system. how may i help you? usr i want to european restaurant sys would you like something in the cheap, moderate, or expensive price range? usr i don’t care the area sys would you like something in the cheap, moderate, [. . . ] ? usr european expensive sys would you like something in the cheap, moderate, [. . . ] ? usr european food sys what part of town do you have in mind? usr i don’t care sys michaelhouse cafe is a nice place [. . . ] . usr thank you goodbye figure 9: example dialogue observed during training using rts after 387 training dialogues resulting in iq = 1. sys welcome to the cambridge restaurant dialogue system. how may i help you? usr international food sys you are looking for a international venue right? usr international sys would you like something in the cheap, moderate, or expensive price range? usr don’t care sys you are looking for a international venue right? usr yes sys what part of town do you have in mind? usr center sys bloomsbury restaurant is a nice place [. . . ] . usr thank you goodbye figure 10: example dialogue observed during training using rs iq after 365 training dialogues resulting in iq = 4. figures 9 and 10 show two successful example dialogues for the models trained with rts and rs iq, respectively. one effect of training with rs iq was a reduced number of system repetitions (which may be linked to the reprompt? feature). obviously, the human experiments have only been conducted using rs iq and not with rbi iq. however, the general framework of applying an iq reward estimator for learning a dialogue policy has shown to be applicable in such a setup and it seems rather unlikely that the changes we induce by changing the reward estimator lead to a fundamentally different result. 104 user satisfaction reward estimation across domains 7. discussion one of the major questions of this work addresses the impact of an iq reward estimator on the resulting dialogues as different iq estimators achieves different levels of performance. analysing the results of the dialogue policy learning experiment leads to the conclusion that, for gp-sarsa, the policy learned with rbi iq performs similar or better than rs iq throughout all experiments while still achieving better average user satisfaction compared torts . especially for noisy environments, the improvement is relevant. this finding does not transfer to the other policy models, though. for both, dqn and enac, using rts results in the overall best peformance for both. this can be explained by the nature of the iq reward estimates as they are not as noise-free as rts . here, gaussian processes are known to be able to deal with this type of noisy targets better while deep learning approaches are known to be quite sensitive. the bilstm clearly performs better on the lego corpus while learning the temporal dependencies instead of using handcrafted ones. however, it entails the risk that these learned temporal dependencies are too specific to the original data so that the model does not generalise well anymore. this would mean that it would be less suitable to be applied to dialogue policy learning for different domains. the experiments clearly show that this is not the case. one additional aspect for discussion is the definition of riq, where the interaction quality is paired with a dialogue length penalty to guarantee a better comparison with rts . this length penalty is not strictly necessary as the annotated scores of the lego corpus already contain a notion of dialogue length implicitly as long dialogues usually tend to have lower quality ratings (in contrast to other work, e.g., (foster et al., 2009)). furthermore, the iq estimation approach described in the following uses the dialogue length as an explicit input parameter. however, adding a dialogue length penalty to riq is not regarded as harmfull to the overall approach as dialogue length plays a subordinate role in the learning process. moreover, even though this limits the usage of the lego corpus for its applicability to types of interactions, it does not limit the overall approach presented in this work: for different tasks, new data needs to be collected and annotated. 8. conclusion this article has demonstrated that employing a user satisfaction reward estimator for learning dialogue policies without any knowledge about the domain can yield good performance in terms of both task success rate and (estimated) user satisfaction. this has been demonstrated by training reward estimators on a bus information domain and applying it to learn dialogue policies in five different domains (cambridge restaurants and hotels, san francisco restaurants and hotels, laptops) in a simulated experiment. the reward estimator using bilstms with attention mechanism achieved better estimation performance than a svm-based estimator while learning all temporal dependencies implicitly. moreover, one of the estimators has successfully been applied to learning dialogue policies in the domain of finding a restaurant in cambridge through interaction with real users. for future work, we aim at extending the user satisfaction estimator by incorporating domainindependent linguistic data to further improve the estimation performance. to tackle the problem of degrading performance if the noise level increases, one possible solution would be to have a combination of success and satisfaction as the reward. in addition, active learning will be investigated to mitigate the requirement for iq annotated training data. furthermore, the effects of using a user 105 ultes and maier satisfaction-based reward estimator needs to be applied to more complex tasks, e.g., as defined by the conversational entity dialogue model (ultes et al., 2018). 9. acknowledgements part of this research was funded by the epsrc grant ep/m018946/1 open domain statistical spoken dialogue systems. references yoshua bengio, patrice simard, paolo frasconi, et al. learning long-term dependencies with gradient descent is difficult. ieee transactions on neural networks, 5(2):157–166, 1994. praveen kumar bodigutla, lazaros polymenakos, and spyros matsoukas. multi-domain conversation quality evaluation via user satisfaction estimation. arxiv preprint arxiv:1911.08567, 2019a. praveen kumar bodigutla, longshaokan wang, kate ridgeway, joshua levy, swanand joshi, alborz geramifard, and spyros matsoukas. domain-independent turn-level dialogue quality evaluation via user satisfaction estimation. arxiv preprint arxiv:1908.07064, 2019b. praveen kumar bodigutla, aditya tiwari, spyros matsoukas, josep valls-vargas, and lazaros polymenakos. joint turn and dialogue level user satisfaction estimation on multi-domain conversations. in findings of the association for computational linguistics: emnlp 2020, pages 3897–3909, online, november 2020. association for computational linguistics. doi: 10.18653/v1/2020.findings-emnlp.347. url https://www.aclweb.org/anthology/ 2020.findings-emnlp.347. iñigo casanueva, paweł budzianowski, pei-hao su, nikola mrkšić, tsung-hsien wen, stefan ultes, lina rojas-barahona, steve young, and milica gašić. a benchmarking environment for reinforcement learning based task oriented dialogue management. in deep reinforcement learning symposium, 31st conference on neural information processing systems (nips), 2017. chih-chung chang and chih-jen lin. libsvm: a library for support vector machines. acm transactions on intelligent systems and technology, 2:27:1–27:27, 2011. software available at http://www.csie.ntu.edu.tw/˜cjlin/libsvm. jacob cohen. a coefficient of agreement for nominal scales. in educational and psychological measurement, volume 20, pages 37–46, april 1960. jacob cohen. weighted kappa: nominal scale agreement provision for scaled disagreement or partial credit. psychological bulletin, 70(4):213, 1968. heriberto cuayáhuitl, seonghan ryu, donghyeon lee, and jihie kim. a study on dialogue reward prediction for open-ended conversational agents. in 2018 neurips workshop on conversational ai: “today’s practice and tomorrow’s potential”, montréal, canada., 2018. lucie daubigney, matthieu geist, and olivier pietquin. off-policy learning in large-scale pomdp-based dialogue systems. in proceedings of the 37th ieee international conference 106 user satisfaction reward estimation across domains on acoustics, speech and signal processing (icassp 2012), pages 4989–4992, kyoto (japan), 2012. ieee. doi: 10.1109/icassp.2012.6289040. layla el asri, romain laroche, and olivier pietquin. reward function learning for dialogue management. in proceedings of the 6th starting ai researchers’ symposium (stairs), pages 95–106. ios press, 2012. doi: 10.3233/978-1-61499-096-3-95. layla el asri, romain laroche, and olivier pietquin. reward shaping for statistical optimisation of dialogue management. in adrian-horia dediu, carlos martı́n-vide, ruslan mitkov, and bianca truthe, editors, statistical language and speech processing, pages 93–101, berlin, heidelberg, 2013. springer berlin heidelberg. isbn 978-3-642-39593-2. layla el asri, hatim khouzaimi, romain laroche, and olivier pietquin. ordinal regression for interaction quality prediction. in international conference on acoustics, speech and signal processing (icassp), pages 3245–3249. ieee, may 2014a. doi: 10.1109/icassp.2014.6854195. layla el asri, romain laroche, and olivier pietquin. task completion transfer learning for reward inference. in workshops at the twenty-eighth aaai conference on artificial intelligence, 2014b. maxine eskenazi, alan w black, antoine raux, and brian langner. let’s go lab: a platform for evaluation of spoken dialog systems with real world users. in ninth annual conference of the international speech communication association, 2008. mary ellen foster, manuel giuliani, and alois knoll. comparing objective and subjective measures of usability in a human-robot dialogue system. in proceedings of the joint conference of the 47th annual meeting of the acl and the 4th international joint conference on natural language processing of the afnlp, pages 879–887, suntec, singapore, august 2009. association for computational linguistics. url https://www.aclweb.org/anthology/p09-1099. milica gašić and steve j. young. gaussian processes for pomdp-based dialogue manager optimization. ieeeacm transactions on audio, speech, and language processing, 22(1):28–40, 2014. doi: 10.1109/tasl.2013.2282190. milica gašić, catherine breslin, matthew henderson, dongho kim, martin szummer, blaise thomson, pirros tsiakoulis, and steve j. young. on-line policy optimisation of bayesian spoken dialogue systems via human interaction. in proceedings of the international conference on acoustics, speech and signal processing (icassp), pages 8367–8371. ieee, 2013. doi: 10.1109/icassp.2013.6639297. milica gašić, dongho kim, pirros tsiakoulis, catherine breslin, matthew henderson, martin szummer, blaise thomson, and steve j. young. incremental on-line adaptation of pomdp-based dialogue managers to extended domains. in proceedings of the 15th international conference on spoken language processing (interspeech), pages 140–144. isca, august 2014. alex graves, navdeep jaitly, and abdel-rahman mohamed. hybrid speech recognition with deep bidirectional lstm. in 2013 ieee workshop on automatic speech recognition and understanding, pages 273–278. ieee, 2013. doi: 10.1109/asru.2013.6707742. james henderson, oliver lemon, and kallirroi georgila. hybrid reinforcement/supervised learning of dialogue policies from fixed data sets. computational linguistics, 34(4):487–511, 2008. 107 ultes and maier matthew henderson, blaise thomson, and jason d. williams. the second dialog state tracking challenge. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 263–272, philadelphia, pa, u.s.a., june 2014. association for computational linguistics. doi: 10.3115/v1/w14-4337. url https://www.aclweb.org/ anthology/w14-4337. sepp hochreiter and jürgen schmidhuber. long short-term memory. neural computation, 9(8): 1735–1780, 1997. doi: 10.1162/neco.1997.9.8.1735. url https://doi.org/10.1162/ neco.1997.9.8.1735. j. r. landis and g. g. koch. the measurement of observer agreement for categorical data. biometrics, 33(1):159–174, march 1977. issn 0006-341x. oliver lemon and olivier pietquin. machine learning for spoken dialogue systems. in european conference on speech communication and technologies (interspeech’07), pages 2685–2688, 2007. oliver lemon and olivier pietquin. data-driven methods for adaptive spoken dialogue systems. springer new york, 2012. isbn 978-1-4614-4802-0. doi: 10.1007/978-1-4614-4803-7. esther levin and roberto pieraccini. a stochastic model of computer-human interaction for learning dialogue strategies. in proc. 5th european conference on speech communication and technology (eurospeech 1997), pages 1883–1886, 1997. bing liu and ian lane. adversarial learning of task-oriented neural dialog models. in 19th annual meeting of the special interest group on discourse and dialogue, pages 350–359, 2018. teruhisa misu, kallirroi georgila, anton leuski, and david traum. reinforcement learning of question-answering dialogue policies for virtual museum guides. in proceedings of the 13th annual meeting of the special interest group on discourse and dialogue, pages 84–93, seoul, south korea, july 2012. association for computational linguistics. url http://aclweb. org/anthology-new/w/w12/w12-1611. volodymyr mnih, koray kavukcuoglu, david silver, andrei a rusu, joel veness, marc g bellemare, alex graves, martin riedmiller, andreas k fidjeland, georg ostrovski, stig petersen, charles beattie, amir sadik, ioannis antonoglou, helen king, dharshan kumaran, daan wierstra, shane legg, and demis hassabis. human-level control through deep reinforcement learning. nature, 518(7540):529–533, 2015. doi: 10.1038/nature14236. url https://doi.org/10.1038/nature14236. niklas rach, wolfgang minker, and stefan ultes. interaction quality estimation using long shortterm memories. in proceedings of the 18th annual sigdial meeting on discourse and dialogue, pages 164–169, saarbrücken, germany, august 2017. association for computational linguistics. doi: 10.18653/v1/w17-5520. url https://www.aclweb.org/anthology/ w17-5520. antoine raux, dan bohus, brian langner, alan w. black, and maxine eskenazi. doing research on a deployed spoken dialogue system: one year of let’s go! experience. in proc. of the international conference on speech and language processing (icslp), september 2006. 108 user satisfaction reward estimation across domains verena rieser and oliver lemon. automatic learning and evaluation of user-centered objective functions for dialogue system optimisation. in proceedings of the sixth international conference on language resources and evaluation (lrec’08), marrakech, morocco, may 2008a. european language resources association (elra). url http://www.lrec-conf.org/ proceedings/lrec2008/pdf/592_paper.pdf. verena rieser and oliver lemon. learning effective multimodal dialogue strategies from wizardof-oz data: bootstrapping and evaluation. in proceedings of acl-08: hlt, pages 638–646, columbus, ohio, june 2008b. association for computational linguistics. url https:// www.aclweb.org/anthology/p08-1073. jost schatzmann and steve j. young. the hidden agenda user simulation model. ieee transactions on audio, speech, and language processing, 17(4):733–747, 2009. doi: 10.1109/tasl.2008. 2012071. alexander schmitt and stefan ultes. interaction quality: assessing the quality of ongoing spoken dialog interaction by experts—and how it relates to user satisfaction. speech communication, 74:12–36, november 2015. issn 0167-6393. doi: 10.1016/j.specom.2015.06.003. url http: //www.sciencedirect.com/science/article/pii/s0167639315000679. alexander schmitt, benjamin schatz, and wolfgang minker. modeling and predicting quality in spoken human-computer interaction. in proceedings of the sigdial 2011 conference, pages 173–184, portland, oregon, usa, june 2011. association for computational linguistics. alexander schmitt, stefan ultes, and wolfgang minker. a parameterized and annotated spoken dialog corpus of the cmu let’s go bus information system. in proceedings of the eighth international conference on language resources and evaluation (lrec’12), pages 3369– 3373, istanbul, turkey, may 2012a. european language resources association (elra). url http://www.lrec-conf.org/proceedings/lrec2012/pdf/333_paper.pdf. alexander schmitt, stefan ultes, and wolfgang minker. a parameterized and annotated spoken dialog corpus of the cmu let’s go bus information system. in international conference on language resources and evaluation (lrec), pages 3369–337, may 2012b. weiyan shi and zhou yu. sentiment adaptive end-to-end dialog systems. in proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: long papers), pages 1509–1519, melbourne, australia, july 2018. association for computational linguistics. doi: 10.18653/v1/p18-1140. url https://www.aclweb.org/anthology/p18-1140. satinder singh, diane litman, michael kearns, and marilyn walker. optimizing dialogue management with reinforcement learning: experiments with the njfun system. journal of artificial intelligence research, 16:105–133, 2002. charles edward spearman. the proof and measurement of association between two things. american journal of psychology, 15:88–103, 1904. pei-hao su, david vandyke, milica gašić, dongho kim, nikola mrkšić, tsung-hsien wen, and steve j. young. learning from real users: rating dialogue success with neural networks for 109 ultes and maier reinforcement learning in spoken dialogue systems. in interspeech, pages 2007–2011. isca, september 2015. pei-hao su, milica gašić, nikola mrkšić, lina m. rojas-barahona, stefan ultes, david vandyke, tsung-hsien wen, and steve young. on-line active reward learning for policy optimisation in spoken dialogue systems. in proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), pages 2431–2441, berlin, germany, august 2016. association for computational linguistics. doi: 10.18653/v1/p16-1230. url https: //www.aclweb.org/anthology/p16-1230. pei-hao su, paweł budzianowski, stefan ultes, milica gašić, and steve young. sampleefficient actor-critic reinforcement learning with supervised data for dialogue management. in proceedings of the 18th annual sigdial meeting on discourse and dialogue, pages 147– 157, saarbrücken, germany, august 2017. association for computational linguistics. doi: 10.18653/v1/w17-5518. url https://www.aclweb.org/anthology/w17-5518. richard s. sutton and andrew g. barto. reinforcement learning: an introduction. mit press, cambridge, ma, usa, 1st edition, 1998. isbn 0262193981. url http://portal.acm. org/citation.cfm?id=551283. t. tieleman and g. hinton. lecture 6.5—rmsprop: divide the gradient by a running average of its recent magnitude. coursera: neural networks for machine learning, 2012. stefan ultes. improving interaction quality estimation with bilstms and the impact on dialogue policy learning. in proceedings of the 20th annual sigdial meeting on discourse and dialogue, pages 11–20, stockholm, sweden, september 2019. association for computational linguistics. doi: 10.18653/v1/w19-5902. url https://www.aclweb.org/anthology/ w19-5902. stefan ultes and wolfgang maier. similarity scoring for dialogue behaviour comparison. in proceedings of the 21th annual meeting of the special interest group on discourse and dialogue, pages 311–322, 1st virtual meeting, july 2020. association for computational linguistics. url https://www.aclweb.org/anthology/2020.sigdial-1.38. stefan ultes and wolfgang minker. improving interaction quality recognition using error correction. in proceedings of the sigdial 2013 conference, pages 122–126, metz, france, august 2013. association for computational linguistics. url https://www.aclweb.org/ anthology/w13-4018. stefan ultes and wolfgang minker. interaction quality estimation in spoken dialogue systems using hybrid-hmms. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 208–217, philadelphia, pa, u.s.a., june 2014. association for computational linguistics. doi: 10.3115/v1/w14-4328. url https://www. aclweb.org/anthology/w14-4328. stefan ultes, tobias heinroth, alexander schmitt, and wolfgang minker. a theoretical framework for a user-centered spoken dialog manager. in ramón lópez-cózar and tetsunori kobayashi, editors, proceedings of the paralinguistic information and its integration in spoken dialogue 110 user satisfaction reward estimation across domains systems workshop, pages 241–246, new york, ny, september 2011. springer new york. isbn 978-1-4614-1334-9. doi: 10.1007/978-1-4614-1335-6 24. stefan ultes, alexander schmitt, and wolfgang minker. towards quality-adaptive spoken dialogue management. in naacl-hlt workshop on future directions and needs in the spoken dialog community: tools and data (sdctd 2012), pages 49–52, montréal, canada, june 2012. association for computational linguistics. url https://www.aclweb.org/anthology/ w12-1819. stefan ultes, alexander schmitt, and wolfgang minker. on quality ratings for spoken dialogue systems – experts vs. users. in proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 569–578, atlanta, georgia, june 2013. association for computational linguistics. url https: //www.aclweb.org/anthology/n13-1064. stefan ultes, matthias kraus, alexander schmitt, and wolfgang minker. quality-adaptive spoken dialogue initiative selection and implications on reward modelling. in proceedings of the 16th annual meeting of the special interest group on discourse and dialogue, pages 374– 383, prague, czech republic, september 2015. association for computational linguistics. doi: 10.18653/v1/w15-4649. url https://www.aclweb.org/anthology/w15-4649. stefan ultes, paweł budzianowski, iñigo casanueva, nikola mrkšić, lina rojas-barahona, pei-hao su, tsung-hsien wen, milica gašić, and steve young. domain-independent user satisfaction reward estimation for dialogue policy learning. in proc. interspeech 2017, pages 1721–1725. isca, august 2017a. doi: 10.21437/interspeech.2017-1032. url http://dx.doi.org/ 10.21437/interspeech.2017-1032. stefan ultes, lina m. rojas-barahona, pei-hao su, david vandyke, dongho kim, iñigo casanueva, paweł budzianowski, nikola mrkšić, tsung-hsien wen, milica gašić, and steve young. pydial: a multi-domain statistical dialogue system toolkit. in proceedings of acl 2017, system demonstrations, pages 73–78, vancouver, canada, july 2017b. association for computational linguistics. url https://www.aclweb.org/anthology/p17-4013. stefan ultes, paweł budzianowski, iñigo casanueva, lina m. rojas-barahona, bo-hsiang tseng, yen-chen wu, steve young, and milica gašić. addressing objects and their relations: the conversational entity dialogue model. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 273–283, melbourne, australia, july 2018. association for computational linguistics. doi: 10.18653/v1/w18-5032. url https://www.aclweb.org/ anthology/w18-5032. stefan ultes, juliana miehle, and wolfgang minker. on the applicability of a user satisfactionbased reward for dialogue policy learning, pages 211–217. springer international publishing, cham, 2019. isbn 978-3-319-92108-2. doi: 10.1007/978-3-319-92108-2 22. url https: //doi.org/10.1007/978-3-319-92108-2_22. david vandyke, pei-hao su, milica gašić, nikola mrkšić, tsung-hsien wen, and steve young. multi-domain dialogue success classifiers for policy training. in 2015 ieee workshop on automatic speech recognition and understanding (asru), pages 763–770. ieee, 2015. doi: 10.1109/asru.2015.7404865. 111 ultes and maier vladimir n. vapnik. the nature of statistical learning theory. springer-verlag new york, inc., new york, ny, usa, 1995. isbn 0-387-94559-8. doi: 10.1007/978-1-4757-3264-1. ashish vaswani, noam shazeer, niki parmar, jakob uszkoreit, llion jones, aidan n gomez, ł ukasz kaiser, and illia polosukhin. attention is all you need. in i. guyon, u. v. luxburg, s. bengio, h. wallach, r. fergus, s. vishwanathan, and r. garnett, editors, advances in neural information processing systems, volume 30, pages 5998–6008. curran associates, inc., 2017. url https://proceedings.neurips.cc/paper/2017/file/ 3f5ee243547dee91fbd053c1c4a845aa-paper.pdf. marilyn walker. an application of reinforcement learning to dialogue strategy selection in a spoken dialogue system for email. journal of artificial intelligence research, 12:387–416, 2000. doi: https://doi.org/10.1613/jair.713. marilyn a. walker, diane j. litman, candace a. kamm, and alicia abella. paradise: a framework for evaluating spoken dialogue agents. in 35th annual meeting of the association for computational linguistics and 8th conference of the european chapter of the association for computational linguistics, pages 271–280, madrid, spain, july 1997. association for computational linguistics. doi: 10.3115/976909.979652. url https://www.aclweb.org/ anthology/p97-1035. marilyn a. walker, jeanne c. fromer, and shrikanth narayanan. learning optimal dialogue strategies: a case study of a spoken dialogue agent for email. in coling 1998 volume 2: the 17th international conference on computational linguistics, 1998. url https: //www.aclweb.org/anthology/c98-2214. frank wilcoxon. individual comparisons by ranking methods. biometrics bulletin, 1(6):80–83, 1945. jason d. williams and steve j. young. characterizing task-oriented dialog using a simulated asr chanel. in proceedings of the 8th international conference on spoken language processing (interspeech 2004), pages 185–188, 2004. steve j. young, milica gašić, blaise thomson, and jason d. williams. pomdp-based statistical spoken dialog systems: a review. proceedings of the ieee, 101(5):1160–1179, 2013. doi: 10.1109/jproc.2012.2225812. guineng zheng, subhabrata mukherjee, xin luna dong, and feifei li. opentag: open attribute value extraction from product profiles. in proceedings of the 24th acm sigkdd international conference on knowledge discovery & data mining, pages 1049–1058. acm, 2018. 112 user satisfaction reward estimation across domains appendix a. tables table 12: full results showing the consistency of learned policies in terms of total match rate (tmr), dialogue act match rate (dmr) and concept match rate (cmr) for three different semantic error rates (sers). similarity scores are computed by comparing the learned behaviour of the three policies originating from different random seeds for each policy model and each reward model. ser policy tmr dmr cmr rts rs iq rbi iq rts rs iq rbi iq rts rs iq rbi iq 0% gp 0.48 0.44 0.47 0.72 0.70 0.74 0.55 0.26 0.51 dqn 0.42 0.46 0.39 0.76 0.71 0.70 0.51 0.53 0.48 enac 0.61 0.52 0.52 0.90 0.85 0.76 0.67 0.59 0.67 15% gp 0.51 0.39 0.66 0.74 0.66 0.78 0.58 0.40 0.45 dqn 0.40 0.52 0.57 0.74 0.75 0.78 0.50 0.57 0.18 enac 0.56 0.52 0.39 0.82 0.72 0.63 0.63 0.64 0.60 30% gp 0.36 0.52 0.46 0.62 0.61 0.70 0.45 0.47 0.63 dqn 0.51 0.54 0.50 0.69 0.73 0.66 0.59 0.49 0.27 enac 0.44 0.46 0.37 0.71 0.74 0.61 0.53 0.64 0.55 113 ultes and maier table 13: full results showing the comparison of learned policies using different reward models in terms of total match rate (tmr), dialogue act match rate (dmr) and concept match rate (cmr) for three different semantic error rates (sers). similarity scores are computed by pair-wise comparing the learned behaviour of the three policies of each reward model originating from different random seeds with each trained policy of the other reward model within each policy model. ser type tmr dmr cmr gp dqn enac gp dqn enac gp dqn enac 0% rts – rs iq 0.31 0.33 0.39 0.57 0.55 0.53 0.40 0.40 0.44 rts – rbi iq 0.58 0.58 0.45 0.72 0.75 0.65 0.64 0.64 0.50 rs iq – rbi iq 0.38 0.45 0.27 0.55 0.65 0.42 0.36 0.50 0.23 15% rts – rs iq 0.24 0.41 0.45 0.48 0.64 0.62 0.31 0.48 0.51 rts – rbi iq 0.49 0.53 0.37 0.78 0.74 0.59 0.57 0.60 0.41 rs iq – rbi iq 0.26 0.37 0.27 0.57 0.59 0.53 0.27 0.41 0.34 30% rts – rs iq 0.22 0.36 0.45 0.30 0.55 0.57 0.29 0.43 0.49 rts – rbi iq 0.48 0.57 0.51 0.68 0.76 0.70 0.55 0.63 0.59 rs iq – rbi iq 0.37 0.51 0.24 0.58 0.70 0.41 0.40 0.59 0.26 table 14: full results showing the comparison of learned policies using different policy models in terms of total match rate (tmr), dialogue act match rate (dmr) and concept match rate (cmr) for three different semantic error rates (sers). similarity scores are computed by pair-wise comparing the learned behaviour of the three policies of each policy model originating from different random seeds with each trained policy of the other policy model within each reward model. ser type tmr dmr cmr rts rs iq rbi iq rts rs iq rbi iq rts rs iq rbi iq 0% gp – dqn 0.25 0.20 0.22 0.46 0.40 0.37 0.35 0.27 0.20 gp – enac 0.35 0.37 0.33 0.50 0.56 0.49 0.43 0.43 0.32 dqn – enac 0.28 0.33 0.38 0.48 0.53 0.59 0.36 0.40 0.46 15% gp – dqn 0.24 0.24 0.33 0.45 0.45 0.61 0.34 0.31 0.39 gp – enac 0.30 0.35 0.25 0.57 0.53 0.52 0.39 0.43 0.32 dqn – enac 0.23 0.36 0.18 0.56 0.55 0.47 0.33 0.40 0.23 30% gp – dqn 0.29 0.55 0.31 0.41 0.72 0.49 0.36 0.65 0.35 gp – enac 0.23 0.50 0.26 0.44 0.67 0.41 0.31 0.62 0.33 dqn – enac 0.25 0.49 0.29 0.45 0.66 0.42 0.35 0.58 0.24 114 dialogue & discourse 16(1) (2025) 31–67 doi: 10.5210/dad.2025.102 investigating proactivity in task-oriented dialogues sofia brenna sbrenna@fbk.eu fondazione bruno kessler trento, italy free university of bozen-bolzano bolzano, italy elisabetta jezek elisabetta.jezek@unipv.it university of pavia pavia, italy bernardo magnini magnini@fbk.eu fondazione bruno kessler trento, italy editor: manfred stede submitted 03/2024; accepted 03/2025; published online 03/2025 abstract this paper investigates proactivity, a characteristic phenomenon of collaborative human-human interaction, where a participant in the dialogue offers the addressee some useful and not explicitly requested information. more precisely, a proactive behaviour is: (i) self-prompted and not simply reactive, that is, the speaker does not act merely in response to the requests the other participant has made; (ii) somehow effective for the achievement of the dialogue goal, since the speaker has a long-term, goal-directed behaviour that predicts future states and needs. proactivity has been poorly investigated from a theoretical point of view, and there is a general need of empirical data for both quantitative and qualitative research. the paper provides an extensive analysis of proactivity in several human-human task-oriented dialogic corpora, selected with different characteristics, including chat exchanges and telephone calls, collection modalities such as natural setting and wizard of oz, and two languages, italian and english. the main result is the d-pro corpus, a new resource manually annotated at the utterance level with proactivity and dialogue acts, which allows to investigate proactivity in the context of task-oriented dialogues. there are several findings from our empirical investigation of proactivity: (i) we find that about 20% of turns in our corpus are proactive turns, showing that this is a very diffused and relevant phenomenon; (ii) we confirm the non-reactive nature of proactivity, highlighting the presence of a pattern where a turn in the dialogue triggers a reaction in a following turn and a proactive utterance is then added to the turn; (iii) we show that only a limited number of dialogue acts are actually involved in expressing proactivity, and we discuss the theoretical implications of this finding; (iv) we empirically confirm that proactivity has a crucial role in recovering from goal-failure situations, contributing to the effectiveness of the whole dialogue; (v) we support the intuition of a non-uniform distribution of proactive utterances throughout the dialogue. our empirical findings and the d-pro corpus provide relevant insights for deeper theoretical investigations, as well as crucial resources for improving proactivity in current task-oriented dialogue systems. keywords: proactivity, task-oriented dialogue, annotated resources ©2025 sofia brenna, elisabetta jezek, and bernardo magnini this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). brenna, jezek, and magnini 1. introduction human dialogue is a complex interaction characterised by systematic, coordinated behaviours and a collaborative effort on the part of each participant to communicate (ellefson, 2021). collaborative behaviours in dialogue refer to the various actions and strategies employed by participants to adapt appropriately to each other and work together towards effective communication, shared understanding, and the achievement of conversational goals. there are several collaborative behaviours that have been individuated, which include grounding (clark and schaefer, 1987, 1989; clark and brennan, 1991; clark, 1996), clarification requests (purver et al., 2003a,b), backchanneling (shelley and gonzalez, 2013), proactivity (strauß and minker, 2010; balaraman and magnini, 2020a,b), reformulation (fetzer, 2006), giving examples, and convergence / divergence / maintenance (giles and ogay, 2006). however, although some of such collaborative behaviours have received attention, particularly from the perspective of the development of computational dialogue systems, most of them are still under-investigated in recent data-driven approaches to dialogue models, resulting in a substantial under-representation of collaborative behaviours in human-machine dialogues. this situation is quite evident in the area of task-oriented dialogues (mctear, 2020; louvan and magnini, 2020; balaraman et al., 2021) and conversational search (radlinski and craswell, 2017), where, despite the huge application-oriented interest, there is a gap of empirical studies on collaborative phenomena. the purpose of this study is to investigate proactivity, a collaborative phenomenon representing a fundamental property of human interaction. proactivity can be regarded as the ability to provide the addressee with some useful, yet not explicitly requested information. in example (1), an excerpt from a human-human dialogue between a client (c) and an agent (a) is reported; the agent, in utterance u19, reacts to a question about a point of interest (does it have an entrance fee?) answering the question (that information is not available to me). in the same turn, in utterance u20, the agent takes a non-requested, non-reactive initiative, providing a phone number, which was not explicitly required. we regard utterance u20, marked with pro, as a proactive utterance. example (1) c: u18 does it have an entrance fee? a: u19 that information is not available to me. u20 [pro] the phone number is 00872208000.1 although we have an intuition that proactive behaviours are widespread in human dialogues, to our knowledge there is a lack of quantitative and qualitative analysis supporting this intuition. in our investigation we focus on task-oriented dialogues, because we believe that their inherent collaborative nature (participants jointly aim at and contribute to the achievement of one or more communicative goals) should encourage proactivity. we address the following research questions: (i) what is the amount of proactive utterances in task-oriented dialogues? (ii) is there a relation between proactivity and the dialogue acts employed by the dialogue participants? (iii) are there typical linguistic markers of proactivity? (iv) how is proactivity distributed along the flow of a task-oriented dialogue? to address our research questions, we start by formulating an operative definition of proactivity, that we then use to annotate a selected sample of human-human task-oriented dialogues. in the 1example taken from the multiwoz 2.2 corpus, cfr. zang et al. (2020). a = agent, c = client. 32 investigating proactivity in task-oriented dialogues annotation effort, we focus on a few relevant aspects of proactivity, including the relation between proactivity and dialogue acts, goal-failure situations, and the relation with utterances that typically precede proactivity. as for the data, we exploit five already existing task-oriented dialogue collections, with different characteristics in terms of language, conversational domain, media used for exchanging turns, and collection modalities. the resulting annotated corpus (called d-pro corpus2) is freely distributed for further research, thereby compensating the absence of quantitative studies on the presence of proactivity in task-oriented dialogue corpora, especially with regards to the italian language. there are several findings from our empirical investigation of proactivity: (i) we find that about 20% of turns in our corpus are proactive turns, showing that this is a very diffused and relevant phenomenon. in addition, we find that proactivity is more frequent in spontaneous corpora (e.g., social media chat) than in corpora collected through a guided process (i.e., wizard of oz); (ii) we confirm the non-reactive nature of proactivity, highlighting a pattern in which a turn in the dialogue triggers a reactive utterance in the subsequent turn, followed by the addition of a proactive utterance to the triggered response; (iii) we show that only a limited number of dialogue acts are involved in expressing proactivity, with the vast majority of proactive utterances serving the communicative intent of providing information (60%), suggestions (13.9%), or offers (12.5%); (iv) we empirically confirm that proactivity has a crucial role in recovering from goal-failure situations, contributing to the effectiveness of the whole dialogue; furthermore, that dialogues characterized by higher levels of proactivity experience the fewest instances of failure; (v) we provide evidence supporting the intuition that proactive utterances are non-uniformly distributed throughout the dialogue, with a higher concentration observed in the central segments. our empirical findings and the d-pro corpus provide relevant insights for deeper theoretical investigations, as well as crucial resources for improving proactivity in current task-oriented dialogue systems. the paper is structured as follows. section 2 situates proactivity in the context of linguistics and natural language processing. section 3 introduces the definition of proactivity we adopt in our study and presents our annotation scheme of proactivity, while section 4 provides information about our source corpora and the resulting d-pro corpus, including statistics about the number of dialogues, turns and utterances it contains, and its lexical richness. in sections 5 and 6, we discuss the results of the annotation of proactivity at the utterance-, turn-, and dialogue-act levels. finally, in section 7 we investigate how proactive utterances are positioned within the flow of a task-oriented dialogue. in 8 we report our concluding observations and ongoing work. 2. background and related work collaborative behaviour in dialogue refers to participants’ various actions and strategies to work together towards effective communication, shared understanding, and attainment of conversation goals. these behaviours help maintain the flow, coherence, and relevance of the dialogue while ensuring that all participants have the opportunity to contribute and be heard. as referenced in the introduction, among the most prominent linguistic techniques that participants can use for collaborative purposes in a dialogue, we find proactivity. derived from the definition of proactivity in organisational behaviours (grant and ashford, 2008), the term proactivity has been used in natural language processing since at least li et al. (2016) to refer to conversational agents’ capability to create or control the conversation by taking the initiative and anticipating the 2https://github.com/sofiabrenna/d-pro_corpus 33 brenna, jezek, and magnini impacts on themselves or human users, rather than passively responding to the user’s request (see deng et al. (2023) for an overview). from a theoretical point of view, the concept of collaborative behaviour within which proactivity is couched is not tied to a single theory but instead emerges under different terminologies in different traditions of studies, including philosophy and pragmatics. among the first systematic attempts to specify the rules governing participants’ collaborative behaviour in human interactions, we recall h. paul grice’s cooperative principle and maxims of conversation (grice, 1975, 1989), the latter often interpreted as a way of spelling out what the principle itself articulates. grice’s main focus, however, was not to provide a fully-fledged theory of cooperation in human interactions but rather to account for how participants in a communicative exchange derive the implicated meaning of their interlocutors’ utterances, particularly in cases where there is no apparent relation between the utterances.3 the work of j.l. austin (austin, 1962) and j. searle (searle, 1969, 1975) integrates grice’s contribution by identifying a typology of speech acts and illocutionary forces and by examining their application condition in detail. their proposal has been taken up in computational linguistics under the label of dialogue act (stolcke et al., 2000), conversation act and intent (bunt et al., 2010; bunt and girard, 2005; bunt, 2006; traum and hinkelman, 1992). in other works originating from the social psychology of language, the concept of accommodation has been put forth, which offers a theoretical framework for analysing proactivity in nlp. accommodation is the process of modifying one’s communication style, vocabulary, code, and tone (including politeness, brown and levinson (1987); bargiela-chiappini (2003)) to better align with a conversation partner (cf. speech accommodation theory (giles et al., 1973; giles, 1979; giles et al., 1991; giles and powesland, 1997; burt, 1994; scotton, 1988) and communication accommodation theory (giles and ogay, 2006)). this adaptation facilitates understanding, promotes effective collaboration, and fosters a positive interactional atmosphere. accommodation has already been investigated in the design of spoken dialogue systems (cf. vocal accommodation in raveh (2021) and prosodic accommodation in de looze et al. (2014)). a significant body of work has also been dedicated to the concept of participant initiative in a dialogue. according to traum (1997), a speaker would have the initiative if the speaker had the choice as to the content of the utterance, while the other speaker would have the initiative if the speaker had to frame the utterance in response to speech by the other speaker. this concept is closely connected to proactivity, since being proactive inherently requires taking the initiative to anticipate future needs, rather than responding reactively. initiative has been studied in dialogue and discourse analysis in several contexts, for example in task-oriented and advisory dialogues (whittaker and stenton, 1988; walker and whittaker, 1990), in learning environments (core et al., 2003; kersey et al., 2009), in overlaps of speech (yang and heeman, 2010), in multi-party dialogues (strauß and minker, 2010), and in negotiation dialogues (nouri and traum, 2014). research on mixed-initiative dialogues initially focused mainly on monitoring the control flow in dialogue (whittaker and stenton, 1988) showing that control shifts are predictable based on utterance type. 3see the following example taken from grice (1975), 51: a: smith doesn’t seem to have a girlfriend these days. b: he has been paying a lot of visits to new york lately. in such cases, cooperation is needed as a requirement on the behaviour of speakers to reconstruct the unstated connection between the utterances, to go beyond what is said and to understand what is meant (b implicating that smith has, or may have, a girlfriend in new york). note that what grice actually meant by cooperative is still controversial (see ellefson 2021 for a thorough discussion). 34 investigating proactivity in task-oriented dialogues walker and whittaker (1990) equate initiative to control, associate four utterance types with the allocation of control to either participant, and identify types of control shift4. they suggest that control transfer in mixed-initiative dialogues is often collaborative, even during interruptions when the non-controlling participant takes initiative. in such cases, interruptions help align mutual beliefs needed for the collaborative plan, supporting rather than obstructing the dialogue goal achievement. guinn (1998) explores effective collaboration when participants rely on each other to achieve a common goal, viewing initiative as the decision-making power to manage sub-tasks. he argues that having initiative in task management corresponds to having initiative in dialogue management. challenging this monolithic view, smith (1992) introduces variable initiative, while cohen et al. (1998) propose a non-binary perspective with varying degrees of initiative. chu-carroll and brown (1999), followed by kersey et al. (2009) and others, distinguish between dialogue initiative and task initiative: dialogue initiative is held by the participant guiding the conversation, while task initiative belongs to the one leading goal planning. this distinction separates the two types of initiative, aligning with jordan and di eugenio (1997), who refute walker and whittaker (1990) arguing that control pertains to the dialogue level, whereas initiative relates to the problem-solving level. we propose that proactivity operates at an additional level—the turn or utterance level. this view partly aligns with nouri and traum (2014), who develop an annotation scheme for initiative and response behaviour within dialogue turns. they distinguish between two aspects of initiative: establishing new discourse obligations and providing unsolicited material. on the response side, they examine fulfilling obligations and maintaining relevance to prior turns, reflecting sperber and wilson (1986)’s notion of relevance. we define proactivity as closely related to the second aspect of both initiative and response, involving unsolicited yet relevant contributions to the dialogue’s topics and goals. in section 3.1, we will further elaborate on this by seeking to define proactivity more precisely. several attempts have been made to design dialogue systems that enable the conversational agent to behave proactively, for example, by introducing new topics or useful suggestions during the conversation. particularly, in task-oriented dialogues, proactivity has been addressed primarily in the so-called non-collaborative dialogues, where the system and the user may have divergent objectives or conflicting interests regarding the completion of the task (e.g., the price bargain negotiation), and in enriched task-oriented dialogues (balaraman and magnini, 2020a), where the agent takes the initiative to provide useful supplementary information not explicitly requested by the user (e.g., additional knowledge or chitchats), which can improve the quality and effectiveness of conveying functional service in the conversation. finally, sun et al. (2021) constructed the accentor dataset by adding topical chit-chats into the responses for task-oriented dialogues to make the interactions more engaging and interactive. despite the progress, the nlp community still lacks a comprehensive framework that brings together all the concepts related to proactivity under a unified theoretical perspective. we believe that a deeper understanding of proactivity is essential for improving dialogue system design aimed at simulating natural interactions, and our work is an effort to contribute in this direction. 4walker and whittaker (1990) note that assertions, commands, and questions are typically produced by the controlling participant, whereas prompts leave the control to the hearer. a parallel can be drawn between these controller utterance types and the dialogue acts that we identify as conveying proactivity: assertions with inform, offer, and suggest, commands with request and instruct, questions with requests. yet, we believe that while proactivity always entails some sort of initiative, the same is not necessarily true for control. criticism of the equivalence between initiative and control is addressed further ahead in this section. 35 brenna, jezek, and magnini accordingly, there is a shortage of materials and resources specifically focused on proactivity that we can rely on. however, in recent years, possibly due to the impressive performance of large language models, proactivity has gained significant attention and has been incorporated into data annotation efforts, also in neighbouring fields such as human-computer interaction (hci). a notable work in this area is prodial (kraus et al., 2022), a collection of mixed-initiative humancomputer collaborative interactions including different levels of proactive dialogue actions, meant to create a proactive dialogue model. the dataset contains 3,696 system-user exchanges, collected in a serious game setting based on crowd-sourced interactions with an autonomous agent capable of modelling four actions of proactive behaviour (none, notification, suggestion, intervention). the dialogue actions corpus has been annotated with user information and self-reported assessments of the user’s experience with the dialogue system’s behaviour. while the main prodial focus is on the human-computer trust relationship, our goal is to investigate proactivity in human-human dialogues. 3. annotating proactivity in this section we first present an operative definition of proactive behaviour and then we introduce the schema we developed to annotate proactivity in dialogues. the purpose is to extract useful quantitative and qualitative data about proactivity in human-human task-oriented dialogues. 3.1 defining proactivity we have introduced proactivity (see section 1) as the ability to provide the addressee with some useful and not explicitly requested information. a more operative definition is proposed in balaraman and magnini (2020a), where a proactive behaviour, in the context of a task-oriented dialogue system, is defined as any information that: (i) is introduced by the system; (ii) was not previously introduced in the dialogue by the user; and (iii) is assumed to be relevant to achieve the user needs. while this definition has the merit to relate proactivity with information content, which can be somehow located (i.e., annotated), it requires that proactive units (balaraman and magnini, 2020a) are exactly located within dialogue utterances, making the annotation effort excessively complex. in addition, the definition does not consider the proactive contribution of the user in the dialogue, which is instead a crucial one. in this study, we still base proactivity on information content, although adopting a more comprehensive and usable definition. we say that an utterance in the context of a task-oriented dialogue is considered as proactive when one of the participants, either agent or client: 1. does not act merely in response to the requests the other participant has made, so the behaviour is self-prompted and not simply reactive; 2. has a long-term, goal-directed behaviour that predicts future states and needs, so the behaviour is somehow effective for the achievement of the dialogue goal. when these two conditions are satisfied, the corresponding utterance is marked with the tag pro. in example (2), from the italian corpus jilda (sucameli et al., 2021), utterance u8 (dopo aver fatto la triennale a roma, see translation in footnote) and utterance u11 (però ci sono delle opportunità di lavoro su roma.) are marked as pro, because they are not a direct answer to the interlocutor’s request, and, nonetheless, they provide a piece of useful information for the dialogue. 36 investigating proactivity in task-oriented dialogues example (2) a: u7 hai qualche preferenza riguardo al luogo di lavoro? c: u8 [pro] dopo aver fatto la triennale a roma, u9 mi piacerebbe tornare verso casa, a firenze. a: u10 al momento non abbiamo nessun annuncio che faccia al caso tuo nella u10 zona di firenze, u11 [pro] però ci sono delle opportunità di lavoro su roma.5 there are a few requirements that need to be satisfied for our annotation schema to be applied. we focus on written or transcribed, mixed-initiative task-oriented dialogues that, as mentioned earlier, provide the ideal context for investigating collaborative behaviours. in such dialogues we assume a turn-taking partition of the conversation, where each turn is assigned to a participant: in order to make our annotation homogeneous through different dialogues, participants are generally referred to as client, the participant who provides the initial task to be addressed, and agent, the participant who helps the client to solve the task (see example (1)). finally, we assume that each turn can be segmented into the utterances that compose the turn. the annotation schema is based on four levels: utterance annotation, dialogue act annotation, goal failure annotation, and turn adjacency annotation. we describe them in the following sections. 3.2 utterance annotation the basic units we consider for proactivity annotation in a dialogue are utterances, that is, according with (traum and heeman, 1996), continuous pieces of speech beginning and ending with a clear pause, possibly related to paralinguistic features, including facial expressions, laughter, eye contact, and gestures. more specifically, relating the notion of utterance to dialogue acts, we can state, referencing traum (2004), that an utterance can be defined as a small unit of speech or text within a conversational turn corresponding to a single act that is bordered by the speaker’s silence and/or prosodic boundary tones. thus, as far as written chats are concerned, an utterance is bordered by the writer’s sending of a single message, for instance by pressing ’enter’ on the keyboard or tapping the ’send’ button on a smartphone screen, and/or punctuation. this single-act-matching concept enables us to divide conversational turns into utterances within dialogue corpora that lack an initial segmentation into utterances and align all dialogues to the same splitting criterion (see 4.1 for details). the annotation task includes both agent and client utterances. the task consists in marking as proactive an utterance in its totality: the utterance itself may as well contain some non proactive behaviour, but nonetheless it should be marked as proactive whenever it holds a piece of proactive content. 5example taken from the jilda corpus (sucameli et al., 2020). it may be translated to english as follows: a: u7 do you have any preference about where to work? c: u8 [pro] after completing my bachelor’s degree in rome, u9 i would like to move back towards home, to florence. a: u10 currently we do not have any offer in the florence area that matches your u10 requests, u11 [pro] however, there are job opportunities in rome. 37 brenna, jezek, and magnini 3.3 dialogue act annotation proactive utterances are then further classified according to the dialogue act they convey. dialogue acts refer to the performative dimension of dialogue as originally investigated by austin (1962), and subsequently adapted to the purposes of dialogue systems through the notion of conversation acts (traum and hinkelman, 1992), dialogue acts (stolcke et al., 2000; bunt, 2006; bunt et al., 2010), and, more recently, through the notion of intent (louvan and magnini, 2020). to simplify the manual annotation task, we use a limited number of high-level dialogue acts, selected from bunt et al. (2010)’s iso standard taxonomy developed for annotating dialogue with semantic information. our purpose is to employ the same dialogue act annotation schema for each of the 5 sub-corpora, so we need high-level dialogue act tags. in particular, from the schema in figure 1, we select the following dialogue acts (general-purpose communicative functions in the iso taxonomy), which potentially can express proactive utterances.6 each dialogue act is paired with an operative definition adapted to the proactivity annotation task. • inform = a proactive utterance where the participant provides information; • offer = a proactive utterance where the participant proposes to do something or to provide some further information; • suggest = a proactive utterance where the participant suggests that the addressee should do something; • request = a proactive utterance where the participant demands that the addressee do something or where they demand that the addressee provide some information; • instruct = a proactive utterance where the participant provides the addressee with instructions to follow. • other = a label made available to annotators for tagging utterances that could not be classified within the designated 5 dialogue act labels; however, it was never used in the final set of annotations. given this tagset, the annotator must observe what kind of dialogue act is meant and performed by the participant in utterances that have already been observed to present some proactive behaviour (both the agent’s and the client’s). as an example, consider the following dialogue, where the proactive utterance u20 has been annotated with the dialogue act inform, while utterance u21, in the same turn, has been annotated with request. example (3) p: u15 perfect! u16 can meet there at 8ish? r: u17 sounds good ˆ.ˆ p: u18 i’ll be there at 8.10 is that ok? r: u19 yes perfect! u20 [pro][inform] i’m sitting inside with an italian guy i met at a tandem u20 last week ˆ.ˆ u21 [pro][request] tell me when you arrive! 6the dialogue act selection process was based on a pilot annotation (see also 4.2). 38 investigating proactivity in task-oriented dialogues figure 1: the iso standard taxonomy of general-purpose functions proposed by bunt et al. (2010). dialogue acts that are part of our tag-set are highlighted in blue. p: u22 i’m here!7 3.4 goal-failure annotation we also annotate the presence of situations of failure in utterances where a participant fails in achieving a dialogue goal. such goal failure situations are typically linguistically realised through negative answering to a question or through the impossibility to fulfil a request. goal failure situations are interesting for proactivity because they often require some sort of repair that a proactive behaviour can conveniently bring (balaraman and magnini, 2020a). thus, the annotation of failure situations can give us some insights on the correlation between these two phenomena. we annotate goal failure as follows: • fail = a turn that contains a situation of failure, where the participant can not answer a question or fulfil a request. for instance, in the following dialogue (already presented in example 2), utterance u10 at turn t5 (”currently we do not have any offer in the florence area that matches your requests”) is annotated as fail, and it is connected to the next utterance in the dialogue, u11, where the agent is trying to recover from the goal failure with a practice utterance, informing the client about a job offer of interest in the area of rome. example (4) 7example taken from the italian whatsapp corpus (hewett, 2017). whatsapp chat participants are identified by the first letter of their pseudonym; for example, ”p” represents ”peter” and ”r” represents ”raffaelle”. 39 brenna, jezek, and magnini a: t3 u7 hai qualche preferenza riguardo al luogo di lavoro? c: t4 u8 [pro][inform] dopo aver fatto la triennale a roma, t4 u9 mi piacerebbe tornare verso casa, a firenze. a: t5 u10 [fail] al momento non abbiamo nessun annuncio che faccia al caso t5 u10 tuo nella zona di firenze, t5 u11 [pro][inform] però ci sono delle opportunità di lavoro su roma.8 3.5 turn adjacency annotation in the last annotation level, we consider the relation between a proactive utterance and the utterances of the previous (i.e., adjacent) turn. the intuition is that through turn adjacency annotation it will be possible to better investigate the elements in the dialogue that trigger proactivity. practically, once a certain utterance is marked as pro, the annotator has to look at at the adjacent previous turn in the dialogue. if the utterances of the previous turn provide all required context to motivate the current pro utterance, the adj tag (adjacent) is added to the current proactive utterance. on the other hand, when the adjacent turn does not suffice to provide all required context in order to motivate proactivity, the pro utterance is labelled as na (non adjacent). as an example, the proactive utterances u20 and u21 in the following dialogue (example (5)), are both marked as adj because the origin of the purpose of their proactivity can be found in utterance u18 in the previous turn (i’ll be there at 8.10 is that ok?). example (5) p: t4 u15 perfect! t4 u16 can meet there at 8ish? r: t5 u17 sounds good ˆ.ˆ p: t6 u18 i’ll be there at 8.10 is that ok? r: t7 u19 yes perfect! t7 u20 [pro][inform][adj] i’m sitting inside with an italian guy i met at t7 u20 a tandem last week ˆ.ˆ t7 u21 [pro][request][adj] tell me when you arrive! p: t8 u22 i’m here!9 by contrast, in example (6), the proactive utterances u15 and u16 are annotated as na because their proactivity is motivated by utterance u11 (i also need to take a train on wednesday, leaving after 10:15.), which is not the previous adjacent turn. example (6) c: t7 u11 i also need to take a train on wednesday, leaving after 10:15. a: t8 u12 okay, we have a lot of trains leaving after that time. t8 u13 what is your starting point and destination? c: t9 u14 from leicester to cambridge, please. a: t10 u15 [pro][inform][na] ok, the tr9776 leaves at 11:09 and arrives at t10 u15 12:54, the cost is 37.80 pounds, t10 u16 [pro][offer][na] do you want me to book you? c: t11 u17 yes please book the train for 1 person t11 u18 and make sure you give me the reference number.10 8example taken from the jilda corpus (sucameli et al., 2020). see example (2) for an english translation. 9example taken from the italian whatsapp corpus (hewett, 2017). 10example taken from the multiwoz 2.2 corpus (zang et al., 2020). 40 investigating proactivity in task-oriented dialogues 4. investigating proactivity in a task-oriented dialogic corpus as already mentioned, the focus of this study are task-oriented dialogues, under the assumption that such dialogues are likely to show collaborative phenomena among the interlocutors, including proactivity. in order to provide a representative sample of task-oriented dialogues we considered the following criteria: • language. the main language we use to investigate proactivity in task-oriented dialogues is italian. however, in order to assess potential differences due to language, we include in our sample two english corpora: (i) one corpus (i.e., ubuntu) with comparable (same communicative situation) italian and english dialogues, and (ii) one english corpus (i.e., multiwoz). this choice allows for a comparison of proactivity in the two languages, reported in section 5.3. • medium. we consider the diamesic dimension of dialogue, aiming at including in our sample a sufficient variety of communication media, including telephone call transcriptions (i.e., nespole!), irc chat (i.e., ubuntu), social media chat (i.e., whatsapp), and chat-based platforms (i.e., multiwoz and jilda). • data collection methodology. we try to balance dialogues collected with several methods (e.g., wizard of oz, role-taking, ecological settings), in order to assess how proactivity might be influenced by different degrees of spontaneity and naturalness in the speakers–or writers, as well as dialogues showing different levels of lexical variety and syntactical complexity. • participants. all our dialogues are intended to represent human-human communication. we included both two-party dialogues and multi-party dialogues (for instance, ubuntu and partially whatsapp). although multiwoz is collected through wizard of oz, this corpus is generally considered composed of human-human dialogues (see, for instance, budzianowski et al. (2018); paul et al. (2019); wu et al. (2019)), as the agent-wizard dialogues were generated by humans, as opposed to human-machine dialogues. • domain. our corpus selection covers a variety of interaction domains, including simulated professional support on different topics (e.g., nespole!, multiwoz, and jilda), technical support (e.g., ubuntu), and informal social interactions (e.g., whatsapp). table 1 summarizes the five source corpora we selected, as well as their main characteristics. more details for each dialogue corpus are reported in the next section. 4.1 source dialogic corpora in accordance with the outlined criteria, as sources for our study on task-oriented dialogues we have considered five existing corpora: the italian nespole! corpus, the italian whatsapp corpus, the italian ubuntu chat corpus, multiwoz 2.2, and the jilda corpus (table 1). nespole! (mana et al., 2003) is a voip human-human role-taking (anderson et al., 1991) phone call dialogue collection, part of the multi-language and multi-modal nespole! project. nespole! dialogues span two domains: medicine and tourism. for our analysis, we focus on the 56 italian dialogues from the tourism domain (total recording time: 7h 35’), where a tourist (client) calls a travel operator, and the agent’s goal is to arrange a vacation for them in the trentino region 41 brenna, jezek, and magnini corpus year language medium methodology participants domain nespole! 2003 ita voip call role-taking human-human tourism ubuntu 2013 ita-eng irc chat natural human-humans tech support whatsapp 2017 ita social media chat natural human-human(s) private chat multiwoz 2.2 2020 eng chat wizard of oz human-wizard multi domain jilda 2021 ita chat role-taking human-human job offer table 1: a synoptic view on the dialogue corpora that have been analysed in our study, in chronological order; in those cases where only one specific part of a corpus has been used, such as with the italian nespole! corpus and the italian ubuntu chat corpus, the information is given about that specific part of the corpus. of italy. data preparation: dialogues were already segmented into both turns and utterances, and pauses and prosodic boundaries are transcribed. minor edits have been performed in terms of utterance segmentation according to the single-act-matching criterion. ubuntu chat corpus (uthus and aha, 2013) is a multilingual internet relay chat corpus of multi-party human-humans chats, composed of archived logs from ubuntu’s irc technical support channel for ubuntu users, later assembled into dialogues by lowe et al. (2015). the italian channel chats is rather thin compared to the english chats: it collected 645,375 messages from over 10,300 users, while the english channel reached over 26,360,000 messages from almost 530,000 users. users of the platform, identified by nicknames, ask other users to help them with technical issues related to the linux-based operating system ubuntu. the user initiating the help request is regarded as client, users that help them are regarded as agents. data preparation: the chats were already divided into messages, which roughly correspond to utterances, depending on the participant’s preference for shorter or longer messages. minor edits to utterance division have been made. italian whatsapp corpus (hewett, 2017) is composed of italian two-party and multi-party chat dialogues, consisting in samples of whatsapp private conversations from users based in germany and italy. we manually searched the 6,640 messages composing the corpus for well identifiable tasks and used excerpts from conversations that contained task-oriented dialogues, creating a taskoriented whatsapp sub-corpus. due to participants’ code-mixing and code-switching, minor parts of the corpus are in english and a few utterances are in german. thus our whatsapp sub-corpus is for the most part an italian dialogue corpus, while containing 3 dialogues with one or more german utterances and 4 dialogues with one or more english utterances. data preparation: the same procedure was applied for utterance division as with ubuntu. multiwoz 2.2 (zang et al., 2020) is an updated version of the widely used multiwoz corpus (budzianowski et al., 2018), which gathers 8,438 english multi-domain written short dialogues (average turn per dialogue is 13.68), collected through the wizard-of-oz method (kelley, 1984). scripted conversations take place in various domains between a tourist client (the user) and an information centre agent (the system, namely the wizard pretending to be a conversational machine). notwithstanding the participants’ expectations caused by the wizard of oz framework, multiwoz is regarded as a human-human dialogue dataset in the literature and by its authors. a finer-grained 42 investigating proactivity in task-oriented dialogues ontology than ours is used to annotate dialogue acts for the system turns only: inform, request, offerbook, reqmore, bye, offer, bookinform, welcome, recommend, nooffer, select, greet. data preparation: each dialogue turn corresponds to one single message, so messages were manually split into utterances. jilda (sucameli et al., 2021) is an italian corpus of 525 human-human dialogues in the job search and offer domain, collected through the role-taking method (anderson et al., 1991), in a two-party online chat: a client, who is looking for a job, is assisted by an agent in his goal. the corpus is annotated for the presence of proactive information and with the following dialogue act ontology: greet, inform-basic, inform-proactive, request, select, deny. data preparation: the same as for multiwoz applies. 4.2 the d-pro corpus from the five source corpora presented in table 1, we extract a smaller corpus called dialogue proactivity corpus (d-pro), and manually annotate it in accordance with the schema presented in section 3. as dialogues from different sources may have different length (e.g., dialogues in nespole! are much longer than dialogues in multiwoz), from each source corpus, we sample a sub-corpus containing a total of about 600 turns. the resulting corpus consists in a total amount of 151 dialogues, divided into 2,855 turns and 6,028 utterances.11 the annotation process of the d-pro includes the following steps: • guidelines creation. for the purposes of annotation, a document with precise guidelines and examples for the annotators is realized, suited to clarify any doubts and to give a thread the annotator could follow in places where the parting line between what is considered proactive and what is considered not-proactive becomes blurred, and therefore a more subjective and annotator-dependent decision has to be made. • pilot annotations and guidelines revision. an expert annotator is selected and provided with the guidelines for a pilot annotation of 15 dialogues. after the pilot annotation, the guidelines are slightly modified and the dialogue act annotation schema is consolidated, as a response to the feedback given by the annotator. • inter annotator agreement. after the pilot exercise, a second expert annotator is selected, and a portion of 15% of the d-pro dialogues is annotated by the two annotators, in order to estimate their agreement. a detailed description of the inter annotator agreement is presented in section 4.3. • extensive d-pro annotation. as the inter annotator agreement was high, in the last phase the two annotators are engaged in the extensive annotations of the whole d-pro corpus. this phase lasts for about two months, with the annotation of a single dialogue taking between half an hour to two hours of effort, depending on its length and complexity. table 2 presents the composition of the d-pro corpus. 11the target of 600 turns per sub-corpus is not met for the whatsapp sub-corpus due to the modest size of the italian whatsapp source corpus: 45 dialogues were extracted, yielding 401 turns overall. the whatsapp sub-corpus was added at a later stage, and the annotation of 600 turns had already been completed for the other four sub-corpora by the time the whatsapp sub-corpus was incorporated. 43 brenna, jezek, and magnini d-pro corpus nespole! ubuntu whatsapp multiwoz jilda d-pro tot. d-pro micro avg. # dialogues 14 22 45 41 29 151 30.2 # turns 605 612 401 602 635 2855 571 # agent turns 305 393 / 301 321 1320 330 # client turns 300 219 / 301 318 1138 284.5 # utterances 1722 1181 959 963 1203 6028 1205.6 # tokens 11493 7437 4880 7907 9593 41310 8262 # types 1493 1983 1762 1048 1681 7966 1427.8 # lemmas 1092 1466 1249 688 1127 5622 1124.4 ttr 12.99 26.66 36.11 13.25 17.52 / 15.03 avg. turns per dial. 43.21 27.82 8.91 14.68 21.9 / 18.90 st.dev. turns per dial. 18.51 21.87 4.76 4.81 2.98 / 14.79 avg. utt. per dial. 123 53.68 21.31 23.49 41.48 / 39.92 avg. utt. per turn 2.85 1.93 2.39 1.60 1.89 / 2.11 table 2: the composition of the d-pro corpus. the upper part of the table reports dialogue information and the middle one reports lexical information; the lower part reports statistical measures on dialogues. ttr = type/token ratio. statistics about dialogues. dialogue numbers vary significantly across corpora. for instance, nespole! contains approximately one-third of the dialogues found in whatsapp. this variation is accompanied by notable differences in dialogue length, both between corpora—as reflected in the average turn counts (nespole! 43.21 vs. whatsapp 8.91)—and within individual corpora, as indicated by the standard deviation values. during annotation, it was also observed that each subcorpus contains at least one dialogue that is twice as long as another, with the exception of jilda, where dialogue lengths range more narrowly from 17 to 27 turns. statistics about turns and utterances. shifting to a turn-level perspective, we can determine average turn length by looking at the average number of utterances contained in each turn. in accordance with the preceding discussion, nespole! tops by far the other sub-corpora, followed by whatsapp; both exceed the average threshold of 2.11 utterances per turn. the division of dialogue turns per speaker, identified by their role of either client or agent, is included in table 2. the two roles can not be easily assigned to users in most of whatsapp dialogues, especially in group chats, where tasks are equally shared by the participants and where roles are flexible and can be taken on and left by any user during the course of the same dialogue. a similar situation occurs in multi-party dialogues of the ubuntu sub-corpus, where, however, the chat room’s netiquette regulations demand that the user in need of aid directly announce their technical issue, hence facilitating the identification of the client. the multi-party nature of ubuntu dialogues justify the disproportion of turn allotted to the single client versus the several agents. statistics about lexicon. the counts of tokens, types, and lemmas provide an estimate of the size of each sub-corpus’s vocabulary, while the type/token ratio (ttr) offers insights into lexical richness and variety. ttr is considered an indicator of lexical diversity, with higher values indicating larger variability of the corpus vocabulary. ttr is affected by the length of the corpus, which makes comparisons between larger sub-corpora, like nespole!, and smaller ones, like whatsapp, less 44 investigating proactivity in task-oriented dialogues meaningful. a higher ttr for ubuntu compared to multiwoz and jilda suggests greater lexical diversity, implying that ubuntu has a more varied vocabulary per unit of text. this indicates that ubuntu, despite having longer dialogues, contains a higher proportion of unique words compared to multiwoz and jilda. 4.3 inter annotator agreement as referenced above, as an indication of the quality of the annotation, a portion of 15% of the d-pro corpus is selected to be manually annotated by two annotators, and their annotations are compared to calculate the agreement among them. before entrusting the whole annotation work to the second annotator, some training is done, which implies a pilot annotation and confrontation on a selection of dialogues from each of the five sub-corpora. the annotation includes about 70 turns for each of the five source sub-corpora, with a total of 23 dialogues, 375 turns, and 896 utterances. for the iaa calculation the metric used is cohen’s kappa coefficient, illustrated by landis and koch (1977) and pustejovsky and stubbs (2012).12 the evaluation of the agreement through the resulting kappa is made with reference to landis and koch (1977)’s graduated scale for the k value. table 3 reports the results of the inter-annotator agreement over pro and fail annotation computed at both utterance and turn level and of the dialogue act annotation. each computation is made per single sub-corpus and with all sub-corpora combined together, namely on d-pro corpus as a whole. annotation level nespole! ubuntu whatsapp multiwoz jilda d-pro pro utterance 0.77 0.41 0.63 0.85 0.76 0.77 pro turn 0.81 0.45 0.66 0.84 0.81 0.72 fail utterance 1.0 0.49 0.87 1.0 1.0 0.89 fail turn 1.0 1.0 0.88 1.0 1.0 0.96 dialogue act utterance 0.74 0.62 0.92 1.0 0.72 0.84 table 3: iaa per single sub-corpus and on the whole d-pro (all sub-corpora combined together) computed with cohen’s kappa for both utterance-level and turn-level pro and fail annotation, and for dialogue act annotation. the outcomes for utterance-level pro reveal an almost perfect agreement (0.85) on the annotation of the most structured sub-corpus, namely multiwoz, substantial agreement in nespole!, whatsapp, and jilda, and moderate agreement in the least structured sub-corpus, ubuntu (0.41). combining the five sub-corpora together and computing the iaa over the whole d-pro corpus results in k = 0.77, while the simple and weighted mean of the k values for each sub-corpus is respectively 0.68 and 0.71 (substantial agreement). on the other hand, the overall combined turnlevel pro annotation iaa scored k = 0.72. generally the pro agreement computed at turn level 12cohen’s kappa brings more trustworthiness to the two-annotators agreement calculation since it takes into consideration the likelihood that a particular agreement situation accidentally occurred by chance. it is computed with the following formula: k = pr(a)− pr(e) 1− pr(e) where pr(a) is the actual agreement observed between the two annotators and pr(e) is the expected agreement considering chance. 45 brenna, jezek, and magnini scores lower, as the number of turns is smaller than that of utterances in dialogues. consequently, a disagreement on a single turn carries greater weight in the overall calculation, compared to a disagreement on an individual utterance. concerning the fail annotation, there is no substantial difference in the number of failure turns and the number of failure utterances, as no more than one failure situation occurs per turn. the almost perfect agreement on the annotation of failure situations corresponds to a k of 0.89 at utterance level and 0.96 at turn level: results suggest that identifying the specific utterance that conveys the failure is more challenging than determining where it occurs at a higher level, namely, in which turn. the iaa for dialogue act annotation is calculated specifically on utterances where both annotators agreed on the proactive classification, yielding a kappa value of k = 0.84. 5. turn-level and utterance-level proactivity annotation results this section presents and analyses the results of the proactivity annotation (pro label) described in section 4, both at utterance-level and turn-level. we first discuss proactivity in the whole d-pro corpus, then we provide a detailed analysis of the individual sub-corpora in d-pro, and, finally, we provide a cross-linguistic analysis related to the english and italian portions of the ubuntu corpus. 5.1 proactivity in d-pro sub-corpus-specific and overall results of the annotation of proactivity at both the utterance and the turn level are reported in table 4. d-pro nespole! ubuntu whatsapp multiwoz jilda d-pro micro avg. pro turns % 19.01% 18.95% 35.91% 11.13% 18.27% 19.54% agent pro turns % 76.52% 64.66% / 67.16% 49.14% 64.01% client pro turns % 23.48% 35.34% / 32.84% 50.86% 35.99% pro utterance % 14.59% 14.92% 25.34% 9.35% 13.55% 15.31% agent pro utt. % 83.67% 59.09% / 76.67% 44.79% 67.06% client pro utt. % 16.33% 40.91% / 23.33% 55.21% 33.94% pro utt. per pro turn 2.18 1.52 1.69 1.34 1.41 1.65 avg. turn length 2.84 1.93 2.39 1.60 1.89 2.11 table 4: percentages of sub-corpus-specific and total pro annotation computed at turn and utterance level and divided by speaker. the number of proactive utterances per single proactive turn and average turn length reported for comparison. turn-level proactivity. in quantitative terms, the d-pro corpus annotation, made on a total of 2855 turns and 6028 utterances from 151 dialogues, resulted in the marking of 923 dialogue utterances with the pro label13. perhaps the most significant result of the d-pro annotation is that proactivity is a relevant presence in the task-oriented dialogues that have been investigated, since 13note that for our purposes, a dialogue turn is designated and, therefore, annotated as proactive if, and only if, it contains at least one proactive utterance. a proactive turn can, therefore, either consist of: (i) one single proactive utterance, (ii) only proactive utterances, or (iii) a mix of proactive and non-proactive utterances. 46 investigating proactivity in task-oriented dialogues it can be found within a percentage of 19.54% over the total amount of dialogue turns of the dpro corpus. while this is a significative finding, there are important individual differences (e.g., 11.13% in multiwoz and 35.91% in whatsapp, st.dev. = 9.15), which highlight how proactivity is influenced by different dialogue features. a chi-square test for independence was conducted to determine whether the distribution of pro and non-pro turns varied across the five sub-corpora. the test revealed a statistically significant association between sub-corpus and turn proactivity, χ2(4) = 61.24, p < 0.001. particularly, examination of the residuals shows that the higher proactivity rate is present in the sub-corpus with higher natural setting (whatsapp +5.39), while the lower proactivity is registered in multiwoz (−4.34), whose collection was highly guided through instructions to the wizard. the other sub-corpora show smaller deviations that were not as significant. overall, the high number of proactive turns confirm our initial intuition that proactivity plays a crucial role in human-human task-oriented dialogues. pro turn % in d-pro generally reflects pro utterance % outcomes. as reported in table 4, there are subtle rises in the proportions of proactive turns across all sub-corpora, compared to the proportion of proactive utterances. moreover, the data on average turn length suggests that whatsapp and nespole! in particular exhibit longer turns with more utterances. consequently, within these extended turns filled also with non proactive utterances, proactivity is prone to dispersion: it loses part of its significance when analysed at the utterance-level. on the other hand, the 2% increase for multiwoz 2.2 correlates with the short turn length for this last sub-corpus. these observations are consistent with the average turn length data reported in the lower block of table 4, where the metric is the average number of utterance in each turn. utterance-level proactivity. overall, 15.31% of the d-pro utterances have been annotated as proactive, with a lower standard deviation (st.dev. = 5.91) among sub-corpora than for turns and with χ2(4) = 73.51, p < 0.001 (residuals: whatsapp = +6.60, multiwoz = -4.21). an additional datum reported in table 4 is the #pro-utterance/#pro-turn rate, i.e., the proactive utterances count over proactive turns count, shows the average number of proactive utterances contained within one single proactive turn. this datum offers further details on the distribution of proactive utterances across turns. by comparison with the average turn length, computed by utterances composing each turn and reported in the last line of the table, we can draw two conclusions. first, proactive turns tend to be composed of both proactive and non-proactive utterances, since the average value of proactive utterance within proactive turn is below the average of the number of utterances typically composing one single turn. second, results indicate a positive correlation between percentages of proactivity and turn length. statistical significance tests were conducted to assess this correlation. results indicated a positive, moderate correlation between the percentage of proactive turns and turn length, with pearson’s ρ = 0.491. similarly, the percentage of proactive utterances was also positively correlated with turn length, showing a moderate association, pearson’s ρ = 0.4997. these findings suggest that both proactive turns and utterances are moderately correlated with turn length across sub-corpora, with correlation coefficients falling within the moderate range (0.3 to 0.7). in other words, a competent human speaker or writer will know how abundant or how relevant the information is that he or she can proactively provide in the unfolding of the conversation, or how many times he or she can exploit proactivity within the same dialogue, without violating collaborative rules of dialogue and thus without annoying his or her addressee and without deploying behaviours detrimental to achieving the conversational goal. 47 brenna, jezek, and magnini 5.2 proactivity in d-pro sub-corpora in this section we analyse proactivity in each individual sub-corpora composing d-pro, according to the statistics presented in table 4. analysis of nespole! in nespole!, 19.01% of turns are identified as proactive (pro turns %), which is lower only with respect to whatsapp. along with the pro utterance % (14.59%), this indicates that nespole! is particularly rich in collaboration, a result that reflects the nature of telephonic interaction, where the speakers’ proactive input is essential in managing and supporting the progress of the conversation effectively. the ratio of proactive utterances per proactive turn (2.18) is the highest among the sub-corpora, meaning that proactive turns in nespole! often include multiple proactive utterances, emphasizing a more sustained proactive engagement. additionally, the average turn length in nespole! (2.84) is relatively long, possibly indicating more detailed guidance or suggestions by the agent in this task-oriented setting. in fact, an interesting observation about the nespole! sub-corpus comes from the comparison between the small percentage of proactive utterances offered by the client and the (almost 5 times) larger proportion of proactive utterances provided by the agent. based on the analysed data, two primary reasons explain this pattern: (i) clients’ requests, made while explaining their desired accommodation, are informative and relevant, but are not considered proactive since they serve to set the dialogue goal, a necessary and introductory phase in user-initiated task-oriented dialogues like these, where the tourist (client) calls the travel agency to arrange a vacation; (ii) the agent proactively provides more information than the client initially requests, particularly when describing all-inclusive vacation packages and accommodation options. analysis of ubuntu. the italian ubuntu sub-corpus proactivity content is similar to both nespole! and jilda, showing a consistent presence of proactivity. in terms of proactive utterances annotation, 14.92% of all utterances in ubuntu are proactive, and the average number of proactive utterances per proactive turn is 1.52. this pattern suggests that proactivity in ubuntu is often situational and concise, focusing on specific points of support or troubleshooting advice rather than extensive instructional turns. lastly, the average turn length in ubuntu is 1.93 utterances, indicating short, to-the-point exchanges. the 172 proactive utterances in ubuntu show a distribution of 104 proactive utterances offered by the agents, and 72 provided by the client. that seems fair if one considers that in each ubuntu conversation, one client potentially corresponds to several agents, given that any user happening to be online inside the support chat at the moment when the client poses his request, would be entitled to answer them and thereby become an agent; besides, every dialogue actually unfolds with the intervention of one client and at least 2 or 3 agents. in addition to that, the nature of the ubuntu dialogues itself contributes to these high values of proactivity offered by both the agents and the clients, since these most natural chat interactions regularly contain (i) overlapping conversations related to different tech support topics; (ii) new problems arising during tech assistance processes, resulting in either task-change or task-extension situation14; (iii) many attempts at solving the same problem in different ways or with different tools, or with the help of different agents. 14by task-change we refer to those situations in dialogues where the task in force is essentially replaced by another, because the first one is either achieved, given up, or discarded for some reason; by task-extension we refer to places in dialogues where the main task is temporarily put aside to solve one or some more new side-task problems that have shown up during the interaction, and is eventually resumed. 48 investigating proactivity in task-oriented dialogues analysis of whatsapp. whatsapp dialogues hold the widest proportion of proactivity both at the utterance level (25.34%) and at the turn level (35.91%). average turn length and the number of proactive utterances per proactive turn correlate, settling at 2.39 and 1.69, respectively. on the other hand, the average length of dialogues in the sub-corpus, about 9 turns, 21 utterances per dialogue, is the shortest found. this signals that proactivity may be the key to a fast and effective conclusion of the dialogue. the familiarity of the speakers with each other may also play a role in facilitating task-completion. informants are always friends, sometimes close relatives: they perfectly know each other’s expectations and are capable of anticipating fast and effectively (especially in one-toone conversations) the type of information that the other person needs. as a consequence, there is also little space for greetings, formalities, and courtesies. the exchange of information flows quite smoothly, with straightforward requests and without any particular misunderstanding or need for clarification. all of these factors combined facilitate highly collaborative, effective conversations. analysis of jilda. in the jilda dialogues, proactive behaviours are almost equally distributed among clients and agents turns. a dialogue where both parties can interject proactive guidance reflects a cooperative, balanced approach to the dialogue goal. comparing the data to the pro utterance % divided by speaker, we can argue however that clients on average provide some more proactive utterances within proactive turns. this datum correlates with (i) the observed tendency shown by agents to stick to repeating patterns of questions and answers: this behaviour was particularly emphasised in dialogues produced by specific individuals playing the role of the agent, who created their own routinised conversation-management policy and tended to reproduce it even when their client partner changed; furthermore it correlates with (ii) clients being quite aware of which pieces of information were relevant for the agent to select a suitable job offer so that the clients proactively produced said relevant and concise information during the conversation. analysis of multiwoz. among the five analysed sub-corpora, multiwoz 2.2 stands out for having below-average proactivity proportions: 11.13% of turns and only 9.35% of utterances are proactive. this is possibly due to the wizard of oz methodology employed in the collection, which produces dialogues that follow given scripts and are poor in lexical and syntactical variety. the fact that the few cases of proactivity are mostly offered by the agent (they are twice more frequent than the client’s proactive utterances) agrees with the introduction of mid-talk changes of dialogue goal–a behaviour explicitly encouraged in the informants by the developers of the methodology. overall, the multiwoz sub-corpus embodies an agent-led form of proactivity appropriate for task completion, with minimal deviation from the dialogue main track, highlighting the participants’ role in maintaining the focus on the dialogue task. 5.3 a cross-linguistic analysis: the ubuntu sub-corpus although the analysis of proactivity that we have conducted is mainly based on italian, a relevant question is whether our findings can be extended to other languages. in addition, being the multiwoz corpus in english, it remains open the question whether the low proactivity rate in the sub-corpus (i.e., 11.13%) is due to the specific interaction modality, wizard of oz, or to the different language, english, with respect to the other d-pro sub-corpora. although an extensive cross-language investigation is out of the scope or our work, we took advantage of the fact that the ubuntu chat corpus is available in several languages, including italian and english, which are comparable, as both the dialogue task and the collection methodology is the 49 brenna, jezek, and magnini figure 2: percentage distribution of dialogue acts in proactive utterances divided by sub-corpus (left) and for the totality of d-pro (right). same. to investigate potential cross-linguistic differences in the use of proactivity among dialogue participants, a subsample of 200 turns from the english ubuntu chat corpus was labelled. out of 200 turns and 285 utterances, 20.5% and 17.54% respectively were identified as proactive, showing a statistically insignificant increase compared to the italian ubuntu percentages (18.95% and 14.92%): language does not significantly influence the results at the 0.05 significance level (p-value = 0.1599). these findings suggest that we can reject the hypothesis that the english language is to be accounted for major decreases in proactivity rates in the multiwoz corpus, which, on the contrary, can be explained on the basis of the wizard of oz interaction modality. 6. dialogue acts and proactivity in this section we discuss the results of the annotation of the dialogue act communicative functions performed on proactive utterances. first we provide statistics related to the dialogue acts involved in the annotation, and then we discuss the correlations between the linguistic structure of the utterance and the dialogue act annotation. 6.1 dialogue acts and proactivity in d-pro in the following, we discuss the relation between proactivity and the dialogue acts used to annotate d-pro. statistics are reported in in figure 2. inform. we notice a widespread prevalence of the inform tag: 60.2% of overall proactive utterances display a proactive information-giving attitude. this is somehow expected, considering that our definition of proactive behaviour (see section 3.1) indeed focuses on the ability to add new, relevant, unsolicited information. jilda has the most inform tags in proactive utterances (74.85%), which aligns with section 5, where it is stated that jilda’s proactivity is largely provided by clients, who, as we presume, are aware of the type of information the agent needs to 50 investigating proactivity in task-oriented dialogues achieve the dialogue goal within the well-delimited domain of job search/offer (for instance, see example (4)). suggest. as for the suggest label, which qualifies second for tagging frequency (13.9%), the results show that it appears overall thrice less than the inform tag, still keeping a noticeable advantage on the others. proactive suggestions are more common in spontaneous talks than in structured dialogues, with the only exception of whatsapp. in the first type of dialogues, interactions allow participants to negotiate clients’ preferences, discuss benefits and drawbacks, and make suggestions among many solutions available. as far as ubuntu is concerned, suggestions take the place that offers hold in other sub-corpora, as in example (17) in section 6.2: this makes sense because the ontology from which the agents seek solutions to tech issues is made up of a one’s lifetime experience with the ubuntu operating system, so there is not a set of pre-determined options to chose from. offer. in contrast to the two labels above, the offer label triggers the lowest visible value, that is, a percentage of 1.7%, provided by only three occurrences in ubuntu, suggesting there could be a corpus-specific reason for that. in fact, that datum seems to be due to the peculiar nature of the ubuntu interactions. since the general goal of the dialogues is to resolve a technical problem that has arisen in the client’s operating system, there is no need to select a final item amongst the available solutions as would instead be the case in multiwoz (e.g. selection of a restaurant), jilda (selection of a job offer) or nespole! (selection of an all-inclusive package). consequently, there is no need for the agent to offer such options for selection. on the contrary, the most offer dialogue acts are found in whatsapp and multiwoz, where such acts can be performed, for instance, to meet the needs of the interlocutor or as an offer of action (example (6)). request. the multiwoz sub-corpus, though being the poorest in proactivity content among the five sub-corpora composing d-pro, has the highest absolute and percentage value for the request tag. we argue that also this result is due to the corpus-specific nature of the conversations. as mentioned earlier, the dialogues were gathered through the wizard of oz simulation, which deceived the user participants into believing they were interacting with a dialogue system rather than a human being. the consequences of this collection method entail that users’ expectations about how the interaction will proceed differ from what they would be if the users believed to be interacting with an actual human being. as a consequence, there is an increased tendency to use requests that address the user’s needs straightforwardly, putting restrictions to the naturalness and linguistic richness of the dialogue and reducing collaborative phenomena and politeness mechanisms. the latter, typical of human dialogue, usually make people reluctant to make straightforward requests to mere acquaintances (rather than to close friends and relatives, as in whatsapp: example (5)). also, the generation of some sort of proactive behaviour is encouraged by multiwoz’s researchers thanks to the introduction of mid-conversation task changes: task change brings the need for even new requests made to the system to set the user’s preferences for the new dialogue goal. an exception is the ubuntu sub-corpus, where requests are not so scarce as in nespole! and jilda. it is important to notice that in these latter cases, requests are mostly ”requests of action”, that is, utterances where an agent asks the client to try some operation to narrow down the possible reasons for the issue or to attempt some solutions, with a procedure that advances by trials and errors. instruct. the arguments in the previous paragraph on request are helpful also to interpret the highest value for the instruct tag found in the ubuntu sub-corpus. in proceeding by attempts, it 51 brenna, jezek, and magnini is not uncommon that agents instruct clients on doing a particular operation, often giving step-bystep instructions as those in u43 in example (7) below. example (7) c: u39 mi sono spiegato male: u40 l’icona nella dock unity non compare neanche quando firefox è apperto... u41 *è aperto (correzione) a: u42 rocker: resetta unity allora: u43 [pro][instruct] unity --reset dato dopo aver premuto alt+f2.15 6.2 dialogue act-related linguistic analysis in this section we discuss the correlations that we observed between the utterance linguistic structure and the dialogue act annotation. we performed a qualitative analysis of the proactive utterances in the d-pro corpus to verify if they contain recurrent expressions that function as markers of proactivity. to this end, we uploaded our corpus in the sketch engine online platform and queried it through the functions wordlist (that produces frequency lists of words) and n-gram (that produces frequency lists of sequences of tokens that tend to co-occur). our analysis showed that proactive utterances contain several expressions that act as markers of proactivity and that some lexical-syntactical structures appear to be predominantly tied to one particular dialogue act annotation. in the following, we present some case studies: causal clauses, interrogative clauses, sentences introduced by modal verbs, and a selection of other recurrent patterns. causal clauses. in the d-pro corpus causal connectors are often found, such as perché, visto che, poiché (eng: because, since, as), that introduce proactive causal clauses usually labelled with the inform tag. this occurs, for example, when either (i) the client makes a request and afterwards adds one proactive utterance in order to motivate his or her choice (as in example 8) or (ii) the agent brings an offer or places a suggestion (as in examples 9 and 10) and motivates with the client the choice of that particular offer/suggestion. example (8) c: u6 il mio sogno sarebbe quello di fare l’insegnante u7 [pro][inform] perché mi piace lavorare con i bambini e ragazzi.16 example (9) a: u65 quindi non so se lei ha le catene le consiglierei vivamente di portarleu65 u65 magari se non le vuole montare 15example taken from the ubuntu sub-corpus. eng: c: u39 i didn’t make myself clear: u40 the icon in the unity dock does not appear even when firefox is oppen... u41 *is open (correction) g: u42 rocker: reset unity then: u43 [pro][instruct] unity --reset run after pressing alt+f2. 16example taken from the jilda sub-corpus. eng: c: u6 my dream would be to be a teacher u7 [pro][inform] because i enjoy working with children and teenagers. 52 investigating proactivity in task-oriented dialogues u66 {e} cioè non serve montarle u67 comunque se le tenga nel se le porti u68 [pro][inform] anche perché comunque poi salendo in montagna può sempreu68 u68 capitare una nevicata improvvisa.17 example (10) r: u20 comunque quando vuoi possiamo vederci anche io e te u21 [pro][inform] visto che i tandem "ufficiali" sono solo il martedı̀ eu21 u21 mercoledı̀ :)18 d-pro #? inform suggest offer request instruct tot. #utterances 556 128 115 67 57 923 #? 1 4 31 3 0 39 table 5: interrogative clauses count within proactive utterances, grouped by dialogue act labelling. interrogative clauses. direct interrogative clauses have been collected through automatic search for question marks in the closing part of proactive utterances. results on the distribution of interrogative clauses in co-occurrence with proactivity grouped by dialogue act labelling are presented in table 5. the vast majority (31 out of 39, about 80%) of interrogative utterances has been labelled with the offer tag: about 27% of total offers are brought in an interrogative form. example (11) is typical of proactive offers in the multiwoz corpus. example (11) c: u8 i would like an expensive hotel if you can find one. a: u9 the express by holiday inn cambridge is located in the east and meet u9 your criteria. u10[pro][offer] shall i book you a room?19 as for the other dialogue acts, the interrogative clauses count is virtually negligible. the only interrogative sentence labelled with the inform tag is sai che forse non va fatta riposare?? (eng: maybe you should’nt let the dough rest, you know??), where the speaker is delivering information that she is not entirely confident about. with respect to suggest and request dialogue acts, 17example taken from the nespole! sub-corpus. eng: a: u65 so i don’t know if you have snow chains i would strongly advise you to bring u65 them maybe if you don’t want to put them on u66 {erm} i mean, you don’t need to put them on u67 anyway keep them in the bring them u68 [pro][inform] also because then anyway going up into the mountains a sudden u68 snowfall can always happen. 18example taken from the whatsapp sub-corpus. eng: r: u20 anyway when you want we can also meet just you and me u21 [pro][inform] since the "official" tandems are only on tuesday and wednesday u21 :) 19example taken from the multiwoz sub-corpus. 53 brenna, jezek, and magnini interrogative clauses include sentences starting with maybe...?, may i recommend...?, sicuro che non...? (eng: are you sure you don’t...?), puoi...? (eng: could you...?) and are used as a means of showing courtesy. modal verbs. one set of very frequent expressions is the one containing the modal verbs volere (eng: want), potere (eng: can, may), shall, and should in the following constructions: se vuoi / se vuole (eng: if you want), ti posso / le posso (eng: i can) + infinitive clause, shall / should / may i + infinitive clause. such phrases can be found in the context of use where the speaker, usually the agent, offers either to provide some further information or to do something in order to help the addressee: example (12) c: u81 da kpakager ho soltanto adobe flash plugin a: u82 vai nel sito di adobe flash u83 e vedi se te lo installa da firefox u84 [pro][offer] oppure ti posso dire come installarlo in un altro modo.20 example (13) a: u13 i’ve found several restaurants that are located in the centre with a u13 moderate price range. u14 [pro][offer] may i recommend a british restaurant called the oak bistro?21 as seen in example (11), this construction often occurs in interrogative clauses reporting proactive utterances marked as offer. in fact, even in the case of modals, the linguistic pattern is used as a strategy to soften requests or offers and convey politeness (indirect speech act, davison (1975)), whether or not they take the form of questions, as shown in (11). connectives. another frequent structure is represented by sentences introduced by the connectives ma / però, but, or anzi, invece, instead as in (14) and (15). the connective introduces a logical relation of concession between discourse segments (ferrari, 2010). in the case of example (14), the relation is between u5 and u6, while in example (15) is between u42 and u43. proactivity introduced by this kind of connectives is not particularly tied to specific dialogue acts, but rather follows the general distribution of dialogue acts. example (14) j: u1 ma per la pasta fresca si usa tutto l’uovo? u2 va messa a riposare in frigor? u3 un uovo ogni cento grammi? g: u4 tutto l’uovo u5 sarebbe un uovo ogni cento grammi u6 [pro][instruct] ma io ne metto di meno u7 tipo 3/4 uova per mezzo chilo.22 20example taken from the ubuntu sub-corpus. eng: c: u81 from kpakager i only have adobe flash plugin a: u82 go to the adobe flash website u83 and see if it lets you install it from firefox u84 [pro][offer] or i can tell you how to install it in another way. 21example taken from the multiwoz sub-corpus. 22example taken from the whatsapp sub-corpus. eng: 54 investigating proactivity in task-oriented dialogues example (15) a: u41 trus, ovvero: hai installato ubuntu "dentro" windows? c: u42 mi sa che hai ragione...anzi si... u43 [pro][inform] però mi sembra di ricordare che il disco in qualche modo u43 me lo ha fatto partizionare lo stesso...23 other recurrent patterns. additional recurrent expressions include verbs such as consigliare, suggest, or provare, try, in the constructions (ti) consiglio (di)..., i suggest (that) you..., i recommend (that) you..., and prova (a)..., try (doing)... these expressions are commonly employed to perform suggest dialogue acts (as in 16 and 17) or requests, as for instance in utterances like ”prova e facci sapere”, eng: ”try (doing this) and let us know”. example (16) c: u30 mi potresti fornire informazioni sull’altra proposta di lavoro? a: u31 certo, attendi solo un momento per favore u32 [pro][suggest] ti consiglio di informarti comunque presso la munus s.r.l.24 example (17) c: u13 marcotux, puoi spiegarmi come si fa? a: u14 under flea, provo a vedere se esiste in pacchetto u15 [pro][suggest] prova a vedere nel gestore pacchetti se c’è lastfm.25 7. dialogue structure and proactivity this section analyses how proactive utterances are positioned within the flow of a task-oriented dialogue. we discuss three aspects: (i) the relation between proactive utterances and goal failures; (ii) the relation between proactivity and the dialogue turn that originates a proactive utterance; and (iii) how proactive utterances are actually distributed throughout the whole dialogue. j: u1 do you use the whole egg to make fresh pasta? u2 should it be put to rest in the fridge? u3 one egg for every hundred grams? g: u4 the whole egg. u5 that would be one egg for every hundred grams u6 [pro][instruct] but i put less than that u7 like 3/4 eggs per pound. 23example taken from the ubuntu sub-corpus. eng: a: u41 trus, that is: did you install ubuntu "inside" windows? c: u42 i guess you’re right...actually yes.... u43 [pro][inform] i think i remember, though, that it somehow let me partition u43 the disk anyway... 24example taken from the jilda sub-corpus. eng: c: u30 could you provide me with information about the other job offer? a: u31 sure, just hold on a moment please u32 [pro][suggest] i recommend that you still inquire with munus s.r.l. 25example taken from the ubuntu sub-corpus. eng: c: u13 marcotux, can you explain how to do that? a: u14 under flea, i’ll try to see if a package exists u15 [pro][suggest] try seeing in the package manager if there is lastfm. 55 brenna, jezek, and magnini 7.1 goal-failure situations and proactivity as argued in section 3.4, proactivity may help a participant recover from a failure situation, that is, when a communicative goal can not be satisfied. in task-oriented dialogues, such situations occur quite frequently when some expectations of the client do not match with the knowledge of the agent, most of the time simply because the agent has a partial knowledge of the conversational domain (see example (4) in section 3.4). here we analyse the fail tagging on the d-pro corpus, reported in table 6, and investigate how proactivity is related to goal-failure situations. d-pro #fail nespole! ubuntu whatsapp multiwoz jilda d-pro tot. #fail 13 19 14 27 32 105 #fail+pro 11 12 10 6 21 60 #fail+pro/#fail % 84.62% 63.16% 71.43% 22.22% 65.63% 57.14% #fail/#utterances % 0.75% 1.61% 1.46% 2.80% 2.67% 1.74% #fail+pro/#utterances % 0.64% 1.02% 1.04% 0.62% 1.75% 1.00% #fail+pro/#pro-utterances % 4.38% 6.82% 4.12% 6.67% 12.88% 6.50% table 6: occurrences of failure situations related to proactive behaviour and to total number of utterances within the d-pro corpus (fail = tag for failure situations; fail+pro = fail tag co-occurs with pro tag within one single turn). first of all, we notice that out of a total number of 105 failure situations, 60 of them (57%) originated a proactive initiative, supporting our hypothesis that failure turns are highly productive in terms of proactivity. in such a situation, typically, the agent proactively offers an alternative solution with respect to the client’s goals. in terms of distribution in our corpus, there is high variability. the proportion of proactivity in goal-failure situations is very high in nespole! (85%) and in whatsapp (71%), while it is low in multiwoz (22%). notice that the proportion of fail utterances in multiwoz is the highest among our sub-corpora (2.80%), but, in spite of that, the large majority of them are not recovered through proactivity. a possible explanation may lay in the multiwoz design, which drives the agent’s behaviour towards asking the client to provide an alternative goal instead of offering a proactive solution. generally, our intuition is that proactivity, in addition to promoting recovery from failures, would also help to reduce the occurrence of goal-failure situations, namely, a proactive dialogue is less probable to manifest failures. this intuition is also supported by our data: the percentage of pro utterance tags (see table 4) inversely correlates to the number of fail tags (pearson’s ρ = -0.6). again, the most proactive sub-corpora are whatsapp and nespole!, with 25% and 15% of the utterances being proactive (refer to table 4), while the two sub-corpora also had fewer failure utterances, 1.46% and 0.75%. on the other hand, multiwoz is both the least proactive (9% of proactive utterances) and the most failure-prone (2.80%, or almost four times more than nespole!). the last row in table 6 shows the proportion of failures with proactivity out of the total proactive utterances in a dialogue. these numbers support the intuition that proactivity helps reducing failures in task-oriented dialogues. for instance, in ubuntu only 6.82% of proactive utterances are employed to recover failure situations, indicating that the largely prevalent use of proactivity is outside failure situations: in other words, proactivity is indirectly used to prevent the insurgence of failures. 56 investigating proactivity in task-oriented dialogues 7.2 adjacent and non-adjacent turn proactivity in this section we analyse the relation between a proactive utterance and reactive (i.e., non proactive) utterances in adjacent turns and in the same turn. the aim is to validate the hypothesis that proactivity is not a direct response to a conversational stimulus (such as a question), but instead arises from a more autonomous initiative of a dialogue participant. adjacency nespole! ubuntu whatsapp multiwoz jilda d-pro micro avg. #adj % 74.10% 55.68% 79.01% 95.56% 95.09% 77.68% #non adj % 25.90% 44.32% 20.99% 4.44% 4.91% 22.32% table 7: proportion of pro utterances occurring at a turn ti following a reactive utterance triggered by the previous turn ti−1 . table 7 reports the outcomes of turn adjacency annotation with the adj and na tags on the d-pro corpus. for two of our sub-corpora, namely multiwoz and jilda, almost all proactive utterances are added to a reaction utterance triggered by the previous turn (95.56% and 95.09% of the cases, respectively). for instance, in example (1) in section 1, u20 is the proactive utterance (the phone number is ...), u19 in the same turn is the reactive utterance (that information is not available to me.), and u18 in the previous turn (does it have an entrance fee?) is the utterance triggering u19. this pattern, [trigger utterance + reaction utterance + proactive utterance]26, is by far the most frequent in our d-pro corpus, indicating that adjacent turns alone usually provide sufficient context for the following proactive turn to be understood. by contrast, dialogues from nespole! and, especially, from ubuntu, contain a higher number of na tags, implying longer dependencies between a proactive utterance, the reaction utterance and the triggering turn. a possible explanation is that longer turns in nespole! usually correspond to higher numbers of asynchronous messages: these can be due to the medium (spoken phone conversation) in nespole!, where speakers may miss cues for turn-taking dynamics, resulting in incorrect timing, interruptions, and overlap. in other cases, turn-adjacency is impeded by backchanneling turns, that is, expedients for participants to send feedback to the current speaker often emerging in the form of minimal verbal cues (such as uh-uh, sı̀, yes, mm-hmm, bene, good, capisco, i see) signalling active listening, engagement, and understanding. as for the ubuntu dialogues, the participation of multiple agents at once interrupts the flow of one-to-one conversations: since chat dialogues are necessarily asynchronous, this may happen especially in an uncontrolled environment, where users can easily wait minutes or hours before answering a message; incidentally, a behaviour that was discouraged by design in the creation of the multiwoz and the jilda corpus. 7.3 proactivity distribution in dialogue in this section we analyse proactivity from the point of view of the dialogue structure. in table 4 we reported that, overall, proactivity accounts for almost 20% of the turns in our task-oriented 26according to nouri and traum (2014)’s annotation scheme (r = turn that directly relates to previous turn; f = turn fulfils a pending discourse obligation; i = turn imposes an obligation; n = turn provides new optional material) our turn [reactive utterance + proactive utterance] would be labelled as [r, possibly f and/or i + n, possibly i]. 57 brenna, jezek, and magnini figure 3: 5-segments distribution of proactive turns: each dialogue is divided into five equal parts (turn number is the standard measure for the division), and the percentage of proactive turns is computed within each of the five parts, so that dialogues of different lengths are comparable in a coarse grained analysis. dialogues. however, intuitively, proactivity does not distribute uniformly across all portions of a task-oriented dialogue. for instance, the initial turns in a dialogue are typically introductory and used to reveal the communicative goals of the client, while the last turns serve to finalize the dialogue (e.g., making a reservation) and for final greetings (zhai and williams, 2014). in these portions of the dialogue we expect to have less proactive utterances than in the central part of the dialogue. this intuition is confirmed by our findings on proactivity distribution in the d-pro corpus, as depicted in figure 3. to investigate this aspect, each dialogue in the d-pro corpus was split in five segments containing an equal amount of turns (e.g., given a dialogue with 15 turns, each segment contains exactly three turns). taking advantage of the pro annotations of utterances, we calculated the proportion of proactive utterances for each segment of the dialogue, and for each sub-corpus in d-pro. according to our hypothesis, results show that segment 1 and segment 5 are less proactive then the central segments. on average, about 15% of utterances in segment 1 are proactive, and about 10% in segment 5, while the average proactivity in segment 3 is about 30%, showing that the central turns in a task-oriented dialogue are the most proactive. this distribution holds for all our subcorpora, nespole!, whatsapp, multiwoz and jilda, with the exception of ubuntu, whose proactivity distribution is almost uniform across the dialogue segments. a chi-square test has been conducted to compare the distributions across the five sub-corpora at five observation points. for the full corpus, the chi-square statistic is χ2 = 33.30 with 16 degrees of freedom, and the p-value is p = 0.0067, indicating a significant difference between the groups. after removing the ubuntu sub-corpus, the chi-square statistic is χ2 = 14.92 with 12 degrees of freedom, and the p-value is 58 investigating proactivity in task-oriented dialogues p = 0.2459, indicating there is not any more significant difference in proactivity distribution among the groups. a possible explanation for this, is that ubuntu participants are instructed to join the chat without introducing themselves, thus avoiding initial and final greetings that are instead present in the other four types of dialogues. in addition, ubuntu interactions are multi-party dialogues, where different participants bring their contribution at different points in time (given the asynchronous nature of the chat). 8. conclusions and ongoing work in this research, we focus on proactivity in task-oriented dialogues. we take advantage of investigations of different traditions, including the cooperation principle in language pragmatics, accommodation in social psychology of language, and the notion of initiative in dialogue defined in computational linguistics. we provide an operational definition of proactivity at the utterance level, as a collaborative behaviour occurring when: (i) a participant does not act merely in response to a previous request; and (ii) the participant’s behaviour is somehow effective for the achievement of the dialogue goal. we annotate proactive language behaviours in human-human task-oriented dialogues with the goal of quantifying the extent of the phenomenon, clarifying which dialogue acts are expressed by proactive utterances, and identifying under what conditions proactivity tends to occur. to reach these targets, we gather human-human task-oriented dialogues from five pre-existing corpora with different characteristics in terms of language, conversational domain, media used for exchanging turns, and collection modalities. we develop an annotation scheme and the guidelines to label the presence of proactivity in dialogue utterances and turns, and to tag the dialogue act displayed by proactive utterances. these activities result in the creation of a new corpus (called d-pro) annotated for proactive behaviours. findings. our investigation of proactivity in human-human dialogues enables us to have a much clearer definition of the phenomenon, from both a quantitative and qualitative point of view. first, we are able to quantify that about 20% of turns in the d-pro corpus are proactive turns, showing that this is a pervasive phenomenon. second, we show that only a limited number of dialogue acts are actually involved in expressing proactivity, a result that opens interesting theoretical perspectives. in addition, we confirm the non-reactive nature of proactivity, highlighting the presence of a pattern where a turn ti triggers a reaction in a following turn ti+1 and a proactive utterance is then added to ti+1. moreover, we empirically confirm the hypothesis that proactivity has a crucial role in recovering from goal-failure situations, contributing to the efficacy of the whole dialogue. finally, we demonstrate the non-uniform distribution of proactivity throughout the dialogue. limitations and ongoing work. there are several aspects of proactivity that we could not address in this paper, and that we plan for future research. a first aspect is the impact of proactivity on the efficacy of task-oriented dialogues: this analysis would imply a clear methodology to measure dialogue efficacy (e.g., in terms of goal achievement), which, however, is still a challenging research topic. a second aspect involves analysing proactivity in different kinds of dialogues, including argumentative dialogues. here a potential issue is the lack of clear communicative goals, which helped us to characterize proactivity in task-oriented dialogues. potential impact on computational models of dialogue. finally, in the long term, we are interested in developing computational models of proactive conversational agents, based on large 59 brenna, jezek, and magnini language models (llms). current llms, such as gpt-4 and the open source llama family, are specifically instructed to execute interactive tasks, such as question answering and chat-based exchanges. however, although llms achieve excellent performance in information seeking tasks, their conversational abilities when participants need to collaborate to jointly achieve a communicative goal (e.g., booking a restaurant, fixing an appointment) are still far from those exhibited by humans. in order to model collaborative behaviours, recent approaches investigate how llms can be fine tuned to address dialogue pragmatics. for instance, shaikh et al. (2024) show that grounding acts can be identified and annotated by a large language model and modelled through appropriate fine tuning of the model itself. as a first step in this direction we used the d-pro corpus to exploit a language model to predict whether the last utterance in a task-oriented dialogue is either proactive or not-proactive. in brenna and magnini (2024) we show that a few-shot approach with gpt-4o achieves encouraging performance on a test set composed of dialogue snippets collected from the five d-pro sub-corpora, and that, in particular for the nespole! corpus, the agreement between the model labels and the human-annotated gold labels is nearly equivalent to the agreement between humans. as a following step on proactivity, once we have sufficiently-accurate results, we plan to collect a large number of training dialogue snippets (e.g., 100k) both proactive and not-proactive, and use them to instruction-tune an open-source model like llama. we expect that new llms will be able to manifest much better proactive behaviours, and, in general, better collaborative behaviours, than the current llms. references anne h. anderson, miles bader, ellen gurman bard, elizabeth boyle, gwyneth doherty, simon garrod, stephen isard, jacqueline kowtko, jan mcallister, jim miller, et al. the hcrc map task corpus. language and speech, 34(4):351–366, 1991. doi: 10.1177/002383099103400404. url https://doi.org/10.1177/002383099103400404. john langshaw austin. how to do things with words. william james lectures. oxford university press, 1962. url http://scholar.google.de/scholar.bib?q=info: xi2jvixh8_qj:scholar.google.com/&output=citation&hl=de&as_sdt=0, 5&ct=citation&cd=1. vevake balaraman and bernardo magnini. proactive systems and influenceable users: simulating proactivity in task-oriented dialogues. in proceedings of the 24th workshop on the semantics and pragmatics of dialogue-full papers, virually at brandeis, waltham, new jersey, july. semdial, 2020a. vevake balaraman and bernardo magnini. investigating proactivity in task-oriented dialogues. proceedings of the seventh italian conference on computational linguistics clic-it 2020, 2020b. url https://api.semanticscholar.org/corpusid:229294115. vevake balaraman, seyedmostafa sheikhalishahi, and bernardo magnini. recent neural methods on dialogue state tracking for task-oriented dialogue systems: a survey. in haizhou li, gina-anne levow, zhou yu, chitralekha gupta, berrak sisman, siqi cai, david vandyke, nina dethlefs, yan wu, and junyi jessy li, editors, proceedings of the 22nd annual meeting of the special interest group on discourse and dialogue, sigdial 2021, singapore and on60 investigating proactivity in task-oriented dialogues line, july 29-31, 2021, pages 239–251. association for computational linguistics, 2021. url https://aclanthology.org/2021.sigdial-1.25. francesca bargiela-chiappini. face and politeness: new (insights) for old (concepts). journal of pragmatics, 35(10-11):1453–1469, 2003. sofia brenna and bernardo magnini. last utterance proactivity prediction in task-oriented dialogues. in proceedings of the eighth workshop on natural language for artificial intelligence (nl4ai 2024) co-located with the 23rd international conference of the italian association for artificial intelligence (ai*ia 2024). ceur-ws.org, 2024. penelope brown and stephen c. levinson. politeness: some universals in language usage, volume 4. cambridge university press, 1987. paweł budzianowski, tsung-hsien wen, bo-hsiang tseng, inigo casanueva, stefan ultes, osman ramadan, and milica gašić. multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling. arxiv preprint arxiv:1810.00278, pages 5016–5026, october-november 2018. doi: 10.18653/v1/d18-1547. url https://www.aclweb.org/ anthology/d18-1547. harry bunt. dimensions in dialogue act annotation. in proceedings of the fifth international conference on language resources and evaluation (lrec’06), genoa, italy, may 2006. european language resources association (elra). url http://www.lrec-conf.org/ proceedings/lrec2006/pdf/428_pdf.pdf. harry bunt and yann girard. designing an open, multidimensional dialogue act taxonomy. in dialor’05, proceedings of the ninth workshop on the semantics and pragmatics of dialogue, nancy, pages 37–44, 2005. harry bunt, jan alexandersson, jean carletta, jae-woong choe, alex chengyu fang, koiti hasida, kiyong lee, volha petukhova, andrei popescu-belis, laurent romary, et al. towards an iso standard for dialogue act annotation. in seventh conference on international language resources and evaluation (lrec’10), 2010. susan meredith burt. code choice in intercultural conversation: speech accommodation theory and pragmatics. pragmatics. quarterly publication of the international pragmatics association (ipra), 4(4):535–559, 1994. jennifer chu-carroll and michael k. brown. an evidential model for tracking initiative in collaborative dialogue interactions. in susan haller, alfred kobsa, and susan mcroy, editors, computational models of mixed-initiative interaction, pages 49–87. springer netherlands, dordrecht, 1999. isbn 978-94-017-1118-0. doi: 10.1007/978-94-017-1118-0 2. url https: //doi.org/10.1007/978-94-017-1118-0_2. herbert h. clark. using language. cambridge university press, 1996. herbert h. clark and susan e. brennan. grounding in communication. perspectives on socially shared cognition, 1991. 61 brenna, jezek, and magnini herbert h. clark and edward f. schaefer. collaborating on contributions to conversations. language and cognitive processes, 2(1):19–41, 1987. herbert h. clark and edward f. schaefer. contributing to discourse. cognitive science, 13(2): 259–294, 1989. robin cohen, coralee allaby, christian cumbaa, mark fitzgerald, kinson ho, bowen hui, celine latulipe, fletcher lu, nancy moussa, david pooley, alex qian, and saheem siddiqi. what is initiative? user modeling and user-adapted interaction, 8:171–214, 1998. mark g. core, johanna moore, and claus zinn. the role of initiative in tutorial dialogue. in 10th conference of the european chapter of the association for computational linguistics, pages 67–74. association for computational linguistics, 2003. alice davison. indirect speech acts and what to do with them. in peter cole and jerry l. morgan, editors, speech acts, pages 143–185. brill, 1975. céline de looze, stefan scherer, brian vaughan, and nick campbell. investigating automatic measurements of prosodic accommodation and its dynamics in social interaction. speech communication, 58:11–34, 2014. issn 0167-6393. doi: https://doi.org/10.1016/j.specom. 2013.10.002. url https://www.sciencedirect.com/science/article/pii/ s0167639313001386. yang deng, wenqiang lei, wai lam, and tat-seng chua. a survey on proactive dialogue systems: problems, methods, and prospects. in proceedings of the thirty-second international joint conference on artificial intelligence, ijcai ’23, 2023. isbn 978-1-956792-03-4. doi: 10.24963/ijcai.2023/738. url https://doi.org/10.24963/ijcai.2023/738. gretchen ellefson. conversational cooperation revisited. the southern journal of philosophy, 59 (4):545–571, 2021. angela ferrari. connettivi. in enciclopedia dell’italiano. istituto della enciclopedia italiana, roma, 2010. url https://www.treccani.it/enciclopedia/connettivi_ (enciclopedia-dell%27italiano)/. anita fetzer. reformulation and common grounds. in lexical markers of common grounds, pages 159–181. brill, 2006. howard giles. accommodation theory: optimal levels of convergence. in language and social psychology, pages 45–65. basil blackwell, 1979. howard giles and tania ogay. communication accommodation theory. in bryan b. whaley and wendy samter, editors, explaining communication: contemporary theories and exemplars, pages 293–310. lawrence erlbaum associates publishers, 2006. howard giles and peter powesland. accommodation theory. in nikolas coupland and adam jaworski, editors, sociolinguistics, pages 232–239. macmillan education uk, london, 1997. isbn 978-1-349-25582-5. doi: 10.1007/978-1-349-25582-5 19. url https://doi.org/ 10.1007/978-1-349-25582-5_19. 62 investigating proactivity in task-oriented dialogues howard giles, donald m. taylor, and richard bourhis. towards a theory of interpersonal accommodation through language: some canadian data. language in society, 2(2):177–192, 1973. doi: 10.1017/s0047404500000701. howard giles, nikolas coupland, and justine coupland. accommodation theory: communication, context, and consequence. contexts of accommodation: developments in applied sociolinguistics, 1:1–68, 1991. adam m. grant and susan j. ashford. the dynamics of proactivity at work. research in organizational behavior, 28:3–34, 2008. issn 0191-3085. doi: https://doi.org/10.1016/j.riob. 2008.04.002. url https://www.sciencedirect.com/science/article/pii/ s0191308508000038. paul grice. logic and conversation. in peter cole and jerry l. morgan, editors, speech acts, pages 41–58. brill, 1975. paul grice. studies in the way of words. harvard university press, 1989. curry i. guinn. an analysis of initiative selection in collaborative task-oriented discourse. in judith masthoff, editor, user modeling and user-adapted interaction, volume 8, page 255–314. springer publishing company, usa, 1998. doi: 10.1023/a:1008359330641. url https: //doi.org/10.1023/a:1008359330641. freya hewett. sequential organisation in whatsapp conversations. unpublished bachelor’s thesis, free university of berlin, summer semester, 2017. pamela w. jordan and barbara di eugenio. control and initiative in collaborative problem solving dialogues. in working notes of the aaai spring symposium on computational models for mixedinitiative interaction, pages 81–84, 1997. john f. kelley. an iterative design methodology for user-friendly natural language office information applications. in john o. limb, editor, acm transactions on information systems (tois), volume 2(1), pages 26–41. acm new york, ny, usa, january 1984. doi: 10.1145/357417.357420. url https://doi.org/10.1145/357417.357420. cynthia kersey, barbara di eugenio, pamela jordan, and sandra katz. knowledge co-construction and initiative in peer learning interactions. in proceedings of the 2009 conference on artificial intelligence in education: building learning systems that care: from knowledge representation to affective modelling, page 325–332, nld, 2009. ios press. isbn 9781607500285. matthias kraus, nicolas wagner, and wolfgang minker. prodial – an annotated proactive dialogue act corpus for conversational assistants using crowdsourcing. in nicoletta calzolari, frédéric béchet, philippe blache, khalid choukri, christopher cieri, thierry declerck, sara goggi, hitoshi isahara, bente maegaard, joseph mariani, hélène mazo, jan odijk, and stelios piperidis, editors, proceedings of the thirteenth language resources and evaluation conference, pages 3164–3173, marseille, france, june 2022. european language resources association. url https://aclanthology.org/2022.lrec-1.339/. 63 brenna, jezek, and magnini j. richard landis and gary g. koch. the measurement of observer agreement for categorical data. biometrics, 33(1):159–174, 1977. issn 0006341x, 15410420. url http://www.jstor. org/stable/2529310. xiang li, lili mou, rui yan, and ming zhang. stalematebreaker: a proactive content-introducing approach to automatic human-computer conversation. in subbarao kambhampati, editor, proceedings of the twenty-fifth international joint conference on artificial intelligence, ijcai 2016, new york, ny, usa, 9-15 july 2016, pages 2845–2851. ijcai/aaai press, 2016. url http://www.ijcai.org/abstract/16/404. samuel louvan and bernardo magnini. recent neural methods on slot filling and intent classification for task-oriented dialogue systems: a survey. in proceedings of the 28th international conference on computational linguistics, pages 480–496, barcelona, spain (online), december 2020. international committee on computational linguistics. doi: 10. 18653/v1/2020.coling-main.42. url https://www.aclweb.org/anthology/2020. coling-main.42. ryan lowe, nissan pow, iulian serban, and joelle pineau. the ubuntu dialogue corpus: a large dataset for research in unstructured multi-turn dialogue systems. in alexander koller, gabriel skantze, filip jurcicek, masahiro araki, and carolyn penstein rose, editors, proceedings of the 16th annual meeting of the special interest group on discourse and dialogue, pages 285– 294, prague, czech republic, september 2015. association for computational linguistics. doi: 10.18653/v1/w15-4640. url https://aclanthology.org/w15-4640/. nadia mana, susanne burger, roldano cattoni, laurent besacier, victoria maclaren, john mcdonough, and florian metze. the nespole! voip multilingual corpora in tourism and medical domains. in h. bourlard, editor, proceedings of the 8th european conference on speech communication and technology, eurospeech 2003, geneva, switzerland, 01-04 september 2003, pages 1589–1592. international speech communication association (isca), 2003. doi: 10.21437/eurospeech.2003-464. michael mctear. conversational ai: dialogue systems, conversational agents, and chatbots. synthesis lectures on human language technologies. morgan & claypool publishers, united states, october 2020. isbn 9781636390314. doi: 10.2200/s01060ed1v01y202010hlt048. elnaz nouri and david traum. initiative taking in negotiation. in kallirroi georgila, matthew stone, helen hastie, and ani nenkova, editors, proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 186–193, philadelphia, pa, u.s.a., june 2014. association for computational linguistics. doi: 10.3115/v1/w14-4325. url https://aclanthology.org/w14-4325/. shachi paul, rahul goel, and dilek hakkani-tür. towards universal dialogue act tagging for task-oriented dialogues. in proc. interspeech 2019, pages 1453–1457, 2019. doi: 10.21437/ interspeech.2019-1866. matthew purver, jonathan ginzburg, and patrick healey. on the means for clarification in dialogue. in jan van kuppevelt and ronnie w. smith, editors, current and new directions in 64 investigating proactivity in task-oriented dialogues discourse and dialogue, pages 235–255. springer netherlands, dordrecht, 2003a. isbn 97894-010-0019-2. doi: 10.1007/978-94-010-0019-2 11. url https://doi.org/10.1007/ 978-94-010-0019-2_11. matthew purver, patrick g.t. healey, james king, jonathan ginzburg, and greg j. mills. answering clarification questions. in proceedings of the fourth sigdial workshop of discourse and dialogue, pages 23–33, 2003b. url https://aclanthology.org/w03-2103/. james pustejovsky and amber stubbs. natural language annotation for machine learning: a guide to corpus-building for applications. o’reilly media, inc., 2012. filip radlinski and nick craswell. a theoretical framework for conversational search. in ragnar nordlie, nils pharo, luanne freund, birger larsen, and dan russel, editors, proceedings of the 2017 conference on conference human information interaction and retrieval, chiir 2017, oslo, norway, march 7-11, 2017, pages 117–126. acm, 2017. doi: 10.1145/3020165.3020183. url https://doi.org/10.1145/3020165.3020183. eran raveh. vocal accommodation in human-computer interaction: modeling and integration into spoken dialogue systems. phd thesis, saarländische universitätsund landesbibliothek, 2021. carol myers scotton. odeswitching as indexical of social negotiations. in monica heller, editor, codeswitching: anthropological and sociolinguistic perspectives, pages 151–186. de gruyter mouton, berlin, new york, 1988. isbn 9783110849615. doi: doi:10.1515/9783110849615.151. url https://doi.org/10.1515/9783110849615.151. john r. searle. speech acts: an essay in the philosophy of language. cambridge university press, 1969. doi: 10.1017/cbo9781139173438. john r. searle. indirect speech acts. in peter cole and jerry l. morgan, editors, speech acts, pages 59–82. brill, 1975. omar shaikh, kristina gligoric, ashna khetan, matthias gerstgrasser, diyi yang, and dan jurafsky. grounding gaps in language model generations. in kevin duh, helena gomez, and steven bethard, editors, proceedings of the 2024 conference of the north american chapter of the association for computational linguistics: human language technologies (volume 1: long papers), pages 6279–6296, mexico city, mexico, june 2024. association for computational linguistics. doi: 10.18653/v1/2024.naacl-long.348. url https://aclanthology.org/ 2024.naacl-long.348/. leah shelley and fernando gonzalez. back channeling: function of back channeling and l1 effects on back channeling in l2. linguistic portfolios, 2(1), 2013. url https://repository. stcloudstate.edu/stcloud_ling/vol2/iss1/9. ronnie w. smith. a computational model of expectation-driven mixed-initiative dialog processing. phd thesis, duke university, usa, 1992. dan sperber and deirdre wilson. relevance: communication and cognition, volume 142. harvard university press cambridge, ma, 1986. 65 brenna, jezek, and magnini andreas stolcke, klaus ries, noah coccaro, elizabeth shriberg, rebecca bates, daniel jurafsky, paul taylor, rachel martin, carol van ess-dykema, and marie meteer. dialogue act modeling for automatic tagging and recognition of conversational speech. computational linguistics, 26 (3):339–374, 2000. url https://aclanthology.org/j00-3003/. petra-maria strauß and wolfgang minker. proactive spoken dialogue interaction in multiparty environments. springer new york, ny, 1 edition, 2010. isbn 978-1-4419-59911. doi: https://doi.org/10.1007/978-1-4419-5992-8. url https://doi.org/10.1007/ 978-1-4419-5992-8. ilaria sucameli, alessandro lenci, bernardo magnini, manuela speranza, and maria simi. toward data-driven collaborative dialogue systems: the jilda dataset. italia journal of computational linguistics ijcol (torino), 7(1 — 2):67–90, 2021. doi: 10.4000/ijcol.842. url https:// doi.org/10.4000/ijcol.842. irene sucameli, alessandro lenci, bernardo magnini, maria simi, and manuela speranza. becoming jilda. in johanna monti, felice dell’orletta, and fabio tamburini, editors, proceedings of the seventh italian conference on computational linguistics clic-it 2020, volume 2769 of ceur workshop proceedings, bologna, 2020. ceur-ws. url http://ceur-ws.org/ vol-2769/paper\_69.pdf. kai sun, seungwhan moon, paul crook, stephen roller, becka silvert, bing liu, zhiguang wang, honglei liu, eunjoon cho, and claire cardie. adding chit-chat to enhance task-oriented dialogues. in kristina toutanova, anna rumshisky, luke zettlemoyer, dilek hakkani-tur, iz beltagy, steven bethard, ryan cotterell, tanmoy chakraborty, and yichao zhou, editors, proceedings of the 2021 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 1570–1583, online, june 2021. association for computational linguistics. doi: 10.18653/v1/2021.naacl-main.124. url https://aclanthology.org/2021.naacl-main.124/. david r. traum. views on mixed-initiative interaction. in aaai97 spring symposium on mixedinitiative interaction, pages 169–171, 1997. david r. traum. issues in multiparty dialogues. in frank dignum, editor, advances in agent communication, pages 201–211, berlin, heidelberg, 2004. springer berlin heidelberg. isbn 978-3-540-24608-4. david r. traum and peter a. heeman. utterance units in spoken dialogue. in elisabeth maier, marion mast, and susann luperfoy, editors, dialogue processing in spoken language systems, ecai’96 workshop, budapest, hungary, august 13, 1996, revised papers, volume 1236 of lecture notes in computer science, pages 125–140. springer, 1996. doi: 10.1007/3-540-63175-5\ 42. url https://doi.org/10.1007/3-540-63175-5\_42. david r. traum and elizabeth a. hinkelman. conversation acts in task-oriented spoken dialogue. technical report, university of rochester, usa, 1992. david c. uthus and david w. aha. the ubuntu chat corpus for multiparticipant chat analysis. in analyzing microtext, papers from the 2013 aaai spring symposium, palo alto, california, usa, 66 investigating proactivity in task-oriented dialogues march 25-27, 2013, volume ss-13-01 of aaai technical report. aaai, 2013. url http: //www.aaai.org/ocs/index.php/sss/sss13/paper/view/5706. marilyn walker and steve whittaker. mixed initiative in dialogue: an investigation into discourse segmentation. in 28th annual meeting of the association for computational linguistics, pages 70–78, pittsburgh, pennsylvania, usa, june 1990. association for computational linguistics. doi: 10.3115/981823.981833. url https://aclanthology.org/p90-1010/. steve whittaker and phil stenton. cues and control in expert-client dialogues. in 26th annual meeting of the association for computational linguistics, pages 123–130, buffalo, new york, usa, june 1988. association for computational linguistics. doi: 10.3115/982023.982038. url https://aclanthology.org/p88-1015/. chien-sheng wu, andrea madotto, ehsan hosseini-asl, caiming xiong, richard socher, and pascale fung. transferable multi-domain state generator for task-oriented dialogue systems. in anna korhonen, david traum, and lluı́s màrquez, editors, proceedings of the 57th annual meeting of the association for computational linguistics, pages 808–819, florence, italy, july 2019. association for computational linguistics. doi: 10.18653/v1/p19-1078. url https://aclanthology.org/p19-1078/. fan yang and peter a. heeman. initiative conflicts in task-oriented dialogue. computer speech & language, 24(2):175–189, 2010. url https://api.semanticscholar. org/corpusid:14637469. xiaoxue zang, abhinav rastogi, srinivas sunkara, raghav gupta, jianguo zhang, and jindong chen. multiwoz 2.2 : a dialogue dataset with additional annotation corrections and state tracking baselines. in proceedings of the 2nd workshop on natural language processing for conversational ai, pages 109–117, online, july 2020. association for computational linguistics. doi: 10.18653/v1/2020.nlp4convai-1.13. url https://www.aclweb.org/anthology/ 2020.nlp4convai-1.13. ke zhai and jason d. williams. discovering latent structure in task-oriented dialogues. in kristina toutanova and hua wu, editors, proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: long papers), pages 36–46, baltimore, maryland, june 2014. association for computational linguistics. doi: 10.3115/v1/p14-1004. url https: //aclanthology.org/p14-1004/. 67 dialogue & discourse 13(2) (2022) 1–48 doi: 10.5210/dad.2022.201 studying alignment in a collaborative learning activity via automatic methods: the link between what we say and do utku norman∗, utku.norman@epfl.ch chili lab, epfl, lausanne, switzerland tanvi dinkar∗,† t.dinkar@hw.ac.uk ltci, polytechnique institute of paris, télécom paris, paris, france† barbara bruno barbara.bruno@epfl.ch chili lab, epfl, lausanne, switzerland chloé clavel chloe.clavel@telecom-paris.fr ltci, polytechnique institute of paris, télécom paris, paris, france editor: kallirroi georgila submitted 06/2021; accepted 07/2022; published online 08/2022 abstract a dialogue is successful when there is alignment between the speakers, at different linguistic levels. in this work, we consider the dialogue occurring between interlocutors engaged in a collaborative learning task, where they are evaluated on how well they performed and how much they learnt. our main contribution is to propose new automatic measures to study alignment; focusing on lexical alignment, and a new alignment context that we introduce termed as behavioural alignment (when an instruction given by one interlocutor was followed with concrete actions in a physical environment by another). thus we propose methodologies to create a link between what was said, and what was done as a consequence. to do so, we focus on expressions related to the task in the situated activity. these expressions are minimally required by the interlocutors to make progress in the task. we then observe how these local alignment contexts build to dialogue level phenomena; success in the task. what distinguishes our approach from other works, is the treatment of alignment as a procedure that occurs in stages. since we utilise a dataset of spontaneous speech dialogues elicited from children, a second contribution of our work is to study how spontaneous speech phenomena (such as when interlocutors say “uh”, “oh” . . . ) are used in the process of alignment. lastly, we make public the dataset1 to study alignment in educational dialogues. our results show that all teams lexically and behaviourally align to some degree regardless of their performance and learning, and our measures capture that teams that did not succeed in the task were simply slower to collaborate. thus we find that teams that performed better, were faster to align. furthermore, our methodology captures a productive, collaborative period that includes the time where the interlocutors came up with their best solutions. we also find that well-performing teams verbalise the marker “oh” more when they are behaviourally aligned, compared to other times in the dialogue; showing that this marker is an important cue in alignment. to the best of our knowledge, we are ∗the authors contributed equally to this research. †present address: interaction lab, heriot-watt university, edinburgh, uk 1the anonymised “dataset” (justhink alignment dataset) consisting of transcripts, logs, and responses to the pre-test and the post-test, as well as the description of the network in the activity, and “tools” for our analyses in this paper (justhink alignment analysis), are publicly available online, from the zenodo repositories doi: 10.5281/zenodo.4627104 for the dataset, and doi: 10.5281/zenodo.4675070 for the tools. ©2022 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). https://orcid.org/0000-0002-6802-1444 mailto:utku.norman@epfl.ch mailto:t.dinkar@hw.ac.uk https://orcid.org/0000-0003-0953-7173 mailto:barbara.bruno@epfl.ch https://orcid.org/0000-0003-4850-3398 mailto:chloe.clavel@telecom-paris.fr http://doi.org/10.5281/zenodo.4627104 http://doi.org/10.5281/zenodo.4627104 http://doi.org/10.5281/zenodo.4675070 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel the first to study the role of “oh” as an information management marker in a behavioural context (i.e. in connection to actions taken in a physical environment), compared to only a verbal one. our measures contribute to the research in the field of educational dialogue and the intersection between dialogue and collaborative learning research. keywords: alignment; collaborative learning; educational dialogue; spontaneous speech 1. introduction collaboration occurs in a situation in which individuals work together as members of a group in order to solve a problem (roschelle and teasley, 1995). in a collaborative activity, members build shared, abstract representations of the problem at hand (schwartz, 1995). when a collaborative activity is in an educational setting, it is termed as collaborative learning: where individuals ‘learn’ together. the group interactions in collaborative learning are expected to activate mechanisms that would bring about learning, although there is no guarantee that these beneficial interactions will happen (dillenbourg, 1999). on the other hand, there are studies that do not focus on educational goals, but on how humans understand each other by analysing human-human dialogues. similar to understanding group processes in collaborative learning, studying a dialogue is challenging; as rather than an individual effort, a dialogue is an interaction between two or more people, i.e. interlocutors (clark and wilkes-gibbs, 1986). interlocutors ideally take turns to try to reach a common or mutual understanding (clark and schaefer, 1989). often, works that study mutual understanding largely consider dialogues amongst adult interlocutors solving a problem together e.g. in garrod and anderson (1987); fusaroli and tylén (2016), without the added dimension of a learning goal. however, task performance is not always positively correlated with learning outcomes, as a student can fail in the task but learn from it (even learn from failure, as in “productive failure” (kapur, 2008; kapur and bielaczyc, 2012)) and perform well in the task but not learn from it (i.e. unproductive performers (kuhn, 2015)). thus, works on collaborative learning study behaviours related to learning outcomes (such as gaze and gestures), but can lack in-depth dialogue analysis. on the other hand, several works on dialogue, particularly on mutual understanding, have extensively studied the dialogue, but not always with the added depth of learning outcomes. the objective of this article is to bring together the two perspectives, to contribute to the exploration of the deep, complex and tangled relationship between what we say and what we do, and the outcomes of this. specifically, we are interested in the form this takes when applied to children engaged in a collaborative learning activity, in which what they say and do is not only tied to how they perform in the task, but also what they ultimately learn from the activity. therefore, we consider success in the task, by considering both performance in the task, and learning outcomes. to investigate the above question, we firstly propose novel rule-based algorithms to automatically and empirically measure collaboration, by studying the alignment (or the development of shared representations of interlocutors at different linguistic levels (pickering and garrod, 2004, 2006)) between the children resulting from the activity. in this situated activity, the dialogue has an interdependence on the immediate environment. therefore, to study alignment, we focus on the formation of expressions related to the activity, as focusing on these expressions allows us to target information very specific to making progress in the task. using these expressions we thus study two alignment contexts: i) lexical alignment (what was said), i.e. alignment at a lexical level, and ii) behavioural alignment (what was done), i.e. a new alignment context we propose to mean when 2 studying alignment in a collaborative learning activity via automatic methods instructions provided by one interlocutor are either followed or not followed with physical actions by the other interlocutor. additionally, our research on these two alignment contexts considers the occurrence of spontaneous speech phenomena (e.g. “um”, “uh” . . . ), as our dataset is one of spoken dialogues, where these paralinguistic cues may be an important indicator in alignment; and are often cues that are neglected as noise. we then observe how these local level alignment contexts build into a function of dialogue level phenomena; i.e. success in the task. we make this dataset and tools to study alignment in children’s dialogues publicly available (entitled the “justhink alignment dataset”, containing transcriptions, action logs and alignment tools). aims and research questions following these lines of inquiry, there are hence more immediate, information sharing goals between the dialogue participants (alignment), and broader goals such as success in the task. in this article we thus investigate whether alignment between children in dialogues that emerge from a collaborative learning activity is associated with success in the task. concretely, we investigate the following research questions: • rq1 lexical alignment: how do the interlocutors use expressions related to the task? is this associated with task success? • rq2 behavioural alignment: how do the interlocutors follow up these expressions with actions? is this associated with task success? the rest of the article is organised as follows. sec. 2 discusses related work. sec. 3 describes the dataset we obtained to study alignment, and our dataset contributions. sec. 4 presents our research questions and hypotheses. sec. 5 and sec. 6 give our methodology and results for studying lexical alignment (rq1) and behavioural alignment (rq2), respectively. please note, in all results, we first report empirical results, and then further discussion points – i.e. interesting observations that require additional data to be confirmed. lastly, sec. 7 concludes the findings of the paper. 2. related work we stated in sec. 1 that research on collaborative learning studies behaviours related to learning outcomes but can neglect in-depth dialogue analysis, while works on dialogue do not always consider the depth of learning outcomes. in this section, we highlight relevant research on the intersection between the two in sec. 2.1. since terminology is vast and varied, works on collaborative learning that specifically analyse verbal behaviours are very closely related to works on dialogue focused on educational data. then, we briefly discuss work on alignment outside educational scenarios in sec. 2.2. 2.1 analysis of verbal behaviour in educational dialogues broadly speaking, both verbal and non-verbal behaviours could be indicative of the learning process (trausan-matu and slotta, 2021). non-verbal behaviours (gaze, gestures, laughter . . . ) have been linked with the quality of interaction (see schneider et al. 2021; jermann et al. 2011; jermann and nüssli 2012; bangalore kantharaju et al. 2020). verbal behaviours could consist of textual input such as manual/automatically obtained transcripts, or even acoustic features of the speech (pitch, loudness . . . ). speech duration for instance was found to be longer in participant pairs that collaborated better (jermann and nüssli, 2012). while both acoustic and non-verbal behaviour 3 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel are important to study interaction, analysis can sometimes be limited without further context; for instance it may not always be the case that speech indicates better collaboration, as speech could even consist of negative interactions. 2.1.1 analysing textual input to study lexical and behavioural alignment, in this work we mainly utilise textual input (manually transcribed dialogues). textual inputs can allow for a finer grained analysis of collaborative learning; as using text can give further insight into “the processes of development of collective thinking in and by dialogue” (baker et al., 2021). textual inputs then serve as basis for coding schemes that can be used to represent various types of verbal interactions. the goal is to label each segment in the dialogue, that can range from a word, to a sentence, to the complete dialogue itself (strijbos et al., 2006). for example, to analyse the process of how learners construct arguments through dialogue, weinberger and fischer (2006) proposed a scheme that manually labelled online, written discussions. there were labels along multiple process axes, such as the extent to which learners contributed to the dialogue, and the content of their contribution (then automated in further works of dönmez et al. 2005; rosé et al. 2008; mu et al. 2012). thus, (both manual and automated) representation of dialogue can elucidate the learning process, by showing trends that happen in collaboration, and the effect of particular patterns in dialogue on the learning outcomes (borge and rosé, 2021). the analysis of verbal interactions in educational settings has also been investigated in the domain of intelligent tutoring systems (itss). these systems aim to adaptively facilitate learning by extracting content from the learners’ contributions to dialogue and continuously modelling the evolution of their learning. an important component of dialogue-based itss is the representation of the input text into speech acts (d’mello and graesser, 2013). speech acts serve to classify the discourse into communicative/pragmatic functions, such as backchannels (e.g. “uh-huh”), metacognitive statements (e.g. “i need help”) and so on. this representation is required to generate an appropriate response and model the learner’s progress. for instance, most speech acts used by learners consist of answers to questions asked by the tutor (d’mello and graesser, 2013). following this, many works focused on educational scenarios have developed methodologies to perform dialogue (speech) act classification using both supervised and unsupervised methods (e.g. boyer et al. 2010; ezen-can and boyer 2015b,a, 2014; goodman et al. 2005). in the intersection between collaborative learning and itss, analysing student interactions to see how they communicate and collaborate with each other has served as a basis to give adaptive feedback (tchounikine et al., 2010). these interactions are more complex than conventional oneon-one tutoring, as there are added social dimensions (howley and rosé, 2016). here, itss are used to further the students’ skills in a collaborative activity (scheuer et al., 2010), or even facilitate learning. walker et al. (2011a,b) for instance, have implemented a system in such a scenario, where feedback is adaptively given depending on the interactions students are having in the classroom. here, student collaboration was actively analysed using the system put forth by rosé et al. (2008): an automated text classification software used in educational contexts. thus distinguishing types of interactions based on verbal behaviour has been implemented with success to equip an agent acting as a tutor with better feedback capabilities. in the present work, we utilise a dataset of children engaged in a collaborative learning activity, where a robot uses minimal terminology to explain the task to the children. while the robot intervention is minimal in this dataset (see sec. 3 for further 4 studying alignment in a collaborative learning activity via automatic methods details), a goal of this work is to contribute towards research that could enhance the capabilities of the robot; enabling it to intervene to facilitate task success. 2.1.2 complications of working with speech data since our dataset consists of spoken dialogues among children, in addition to studying the process of collaboration, we would like to observe how the speech modality contributes to collaboration. previous research also indicates the importance of the speech modality in educational scenarios. litman et al. (2004) for instance found that learning gain is positively impacted when the interaction modality includes speech, compared to solely written input for human-human tutoring data. related to collaboration, the co-presence of interlocutors, visibility and audibility in the medium have been found to dramatically affect the process of mutual understanding (dillenbourg and traum, 2006; clark and brennan, 1991). in itss, a first step in the pipeline is the transformation of the input into a parsable format (d’mello and graesser, 2013). this is not straightforward when the input is spoken, as the quality of transcription (i.e. the parsable format) can depend on the performance of the automatic speech recognition (asr) system. furthermore, methodologies could neglect to consider spontaneous speech phenomena. after the input is transcribed (and also for general dialogue systems) it is then collapsed into a semantic frame (see tur and de mori 2011; louvan and magnini 2020 for general task oriented-dialogue), consisting of speech acts and certain keywords/phrases to estimate local and global levels of learning. the local level of learning for instance, may be estimated by comparing certain keywords said by the learner compared to the expected answer keyword (d’mello and graesser, 2013). thus while works on itss and educational scenarios may include the spoken modality, often then the textual input (manually or automatically transcribed) will still remove many of the phenomena arising from speech, to model the learner. however, spoken utterances often contain ‘messy’ (disfluent) speech, for instance, spontaneous speech phenomena such as “uh”, “erm” and so on. they may even be relevant when considering the added dimension of a pedagogical goal, as certain speech phenomena have been linked to signs of hesitation and uncertainty (pickett, 2018; smith and clark, 1993; brennan and williams, 1995) and information management (schiffrin, 1987). thus, since speech has been found to be an important factor in learning outcomes (and indeed, analysis shows the impact of the speech modality over just the written modality on learning (litman et al., 2004), but also, the association of acoustic features with learning such as in ward and litman (2007a)), there should be a focus on developing methodologies to analyse its characteristics. we have previously discussed some methodologies that focus on the acoustic characteristics of speech. however, there is not enough progress on systems capable of analysing spontaneous speech phenomena that can be transcribed in text. 2.1.3 alignment in educational dialogue the alignment of behaviour has been investigated at varying linguistic levels, from acoustic/prosodic (e.g. thomason et al. 2013; ward and litman 2007b,a; lubold and pon-barry 2014), to lexical (e.g. ward and litman 2007b,a), and syntactic (e.g. reitter and moore 2014). research on alignment in educational scenarios can shed light on learning outcomes. for instance, sinclair and schneider (2021) showed that linguistic and gestural alignment in dialogue correlate with learning and are indicative of success in collaborative problem solving. similarly, ward and litman (2007a) found that the convergence of lexical and acoustic behaviours can predict learning outcomes. convergence 5 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel of behaviour can be thought of as a variation of alignment. even in an automated setting, lubold et al. (2018) showed that a robot that aligns with the interlocutor and speaks socially has a positive effect on learning. research on alignment can also give information about the process of communication, the dynamic between interlocutors and so on. for instance, to build rapport (“the development of personal relationships between speakers over time” (sinha and cassell, 2015)), interlocutors become closer to each other in terms of acoustic/prosodic behaviour (lubold and pon-barry, 2014). following this, lubold et al. (2015) found that pitch alignment of a learning companion leads to higher perceptions of rapport. sinha and cassell (2015) found that behavioural convergence and rapport are linked to each other besides being correlated to learning gains; but in this case, they found that rapport leads to convergence (of speech rate) in dyadic peer tutoring conversations. these works serve as a foundation to indicate the positive influence that alignment has on learning outcomes. we propose automatic measures of lexical alignment to observe in our dataset whether local alignment is associated with global task success. we similarly study behavioural alignment (when instructions provided by one interlocutor are either followed or not followed with physical actions by the other interlocutor). as discussed, works have studied how alignment can build to other phenomena such as rapport (lubold and pon-barry, 2014), and vice versa (sinha and cassell, 2015). however, to the best of our knowledge, we are the first to focus on how alignment arises from the timely occurrence of actions in a physical environment, thus concretely forming a link between what was said, and what was done as a result. 2.2 analysis of alignment in other scenarios outside of educational scenarios, several methods have been proposed to automatically compute alignment. this allows for an analysis into the dynamics between in interlocutors; e.g. to study deceptive and truthful speech in interview dialogues (levitan et al., 2018), or study the relationships between levels of alignment in multi-party dialogues for a cooperative board game (rahimi et al., 2017). since the present work is focused on lexical and (novel) behavioural alignment, we discuss other works that have studied lexical alignment and proposed methodologies to automatically extract its occurrences. lexical alignment (also referred to as entrainment) has been studied by focusing on various features such as referring expressions (brennan and clark, 1996), repeated sequences in utterances (dubuisson duplessis et al., 2017, 2021), frequent words in the discourse (levitan et al., 2018), hedge words (levitan et al., 2018) and even expressions related to the task (referred to as “topic words” in rahimi et al. 2017). many of these works build from older linguistic works, to automatically study lexical alignment. we “align” with the literature on lexical entrainment in some aspects, e.g. by focusing on expressions related to the task. however, there are drawbacks with the current methodologies to automatically compute entrainment. firstly, many works approach it as a high-level process (e.g. nenkova et al. 2008; friedberg et al. 2012; rahimi et al. 2017), where the focus is on quantifying the overall entrainment between the interlocutors to see how it can build in the discourse (by looking at proximity: a degree of similarity, and convergence: its evolution). however, while lexical alignment builds in the discourse, there is no consideration for the individual contributions of the interlocutors. hence, we break down lexical alignment into these individual contributions, i.e. we treat it as a process that has different stages. to do so, we consider the first time an interlocutor 6 studying alignment in a collaborative learning activity via automatic methods introduces an expression (priming), and the time when the other interlocutor utilises the same expression (establishment). then, like other works, we see how alignment can build in the discourse. specifically, while we also focus on expressions related to the task (like rahimi et al. 2017), we propose to focus on how such expressions are established: this is distinct from checking overall entrainment on a set of words related to the activity. furthermore, there are works that study how lexical alignment can be correlated to acoustic/prosodic alignment (e.g. rahimi et al. 2017). however, we study the distribution of spontaneous speech phenomena (using transcripts and not acoustic data) in relation to the alignment process, and not the alignment of spontaneous speech characteristics itself. we base this on linguistic research (e.g. brennan and williams 1995) that suggests that there could be a link between spontaneous speech phenomena and alignment, which we discuss as reasoning for our methodology in experiments 2 and 4. dubuisson duplessis et al. (2017, 2021) proposed automatic and generic measures to extract lexical structures of alignment (which they refer to as verbal alignment) in a task-oriented dialogue. the proposed method works on alignment based on surface matching of text and does not focus on other levels of linguistic alignment as envisioned by pickering and garrod (2004). however, it is done with the aim of automatically finding these text patterns in the dialogue, by sequentially processing a transcript in an unsupervised manner. in this article, we provide a new tool/framework for studying situated dialogue building on the automatic and generic methodology by dubuisson duplessis et al. (2017, 2021): this requires modelling how the interlocutors refer to their environment and extending the tool based on this model. we then propose new measures that allow us to study behavioural alignment in situated dialogues; i.e. via automatically inferring instructions given by the interlocutors, and then linking those instructions to actions taken in the physical environment. 3. materials 3.1 the justhink dataset the justhink dataset is based on a collaborative problem solving activity for school children (nasir et al. 2020c; nasir, norman, et al., 2020b). it aims to improve their computational thinking (ct) skills by exercising their abstract reasoning on graphs. recent research on educational curricula stresses the need for learning ct skills in schools, as going beyond simple digital literacy to developing these ct skills becomes crucial (menon et al., 2020). with this in mind, the objective of the activity is to expose school children to minimum-spanning-tree problems. scenario a humanoid robot, acting as the ceo of a gold mining company, presents the activity to the children as a game, asking them to help it collect gold, by connecting gold mines one another with railway tracks. they are told to spend as little money as possible to build the tracks, which change in cost according to how they connect the gold mines. the goal of the activity is to find a solution that minimises the overall cost, i.e. an optimal solution for the given network – see appendix c for details. children participate in teams of two to collaboratively construct a solution, by drawing and erasing tracks. once all gold mines are reachable, i.e. in some way connected to each other, they can submit their solution to the robot for evaluation. they must submit their solution together, and can submit as many times as they want. the robot then reveals whether their solution is an optimal solution or, if not, how far it is from an optimal solution (in terms of its cost). in the latter 7 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel figure 1: the justhink activity setup. case, children are also encouraged by the robot to try again. they can submit a solution as many times as they want until the allotted time for the activity is over. the robot does not initiate any expression that contains a task-specific referent (i.e. the mountain names). note that we treat this triadic activity as a dyadic dialogue. the children are initially prompted by the robot to work with each other, and later simply given the cost difference for their sub-optimal solution(s). however, almost all of the exchanges are between the two interlocutors. after careful observation of the dialogues in the dataset, we observe the tendency to ignore the robot unless submitting a solution. setup two children sit across each other, separated by a barrier. a touch screen is placed horizontally in front of each child. children can see each other, but cannot see the other’s screen, as shown in fig. 1. they are encouraged by the robot to verbally interact with each other, and work together to construct a solution to the activity. the screens display two different views of the current solution to the children. one view is an abstract view, where the gold mines are represented as nodes, and the railway tracks that connect them as edges (see fig. 2a). the other view, or the visual view, represents the gold mines and railway tracks with images (see fig. 2b). a child in the abstract view can see the cost of built edges, but cannot act upon the network. that is, after an edge is built, its cost is shown in this view regardless of whether it was built or removed. conversely, in the visual view, a child can add or delete an edge, which is a railway track, but cannot see its cost. the views of the children are swapped every two edit actions, which is any addition or deletion of an edge. hence, after every two edit actions, the child that was in the abstract view moves to the visual view and vice versa. a turn is thus the time interval between two view swaps, i.e. in which one child is in the visual view and the other child is in the abstract view. a turn lasts for two edit actions. this design aims at encouraging children (interlocutors) to collaborate. participants the dataset consists of 76 children in teams of two (41 females: m = 10.3, sd = 0.75 years old; and 35 males: m = 10.4, sd = 0.61). we have one problem solving session, i.e. task, per team. 8 out of the 38 teams (≈ 21%) found an optimal solution to the activity. the teams were formed randomly, without considering the gender, nationality, or the mother tongue. they include mixed and same gender pairs, and this information is available but not used in our analyses. the study was conducted in multiple international schools in switzerland, where the medium of education is in english, and hence students are proficient in english. content we note that to measure task success, we consider both performance of the teams, and their learning outcomes. from the original justhink dataset (nasir et al., 2020b), we utilise: 8 studying alignment in a collaborative learning activity via automatic methods (a) an example abstract view. (b) corresponding visual view. figure 2: interlocutors’ views during the justhink activity. • the recorded audio files: audio was recorded as two mono audio channels synchronised to each other, with one lavalier microphone per channel. the interlocutors were asked to speak in english. the microphones were clipped onto the interlocutors’ shirts. at a local level, to study lexical and behavioural alignment, we transcribe a representative subset of the audio files. • event log files: event log entries consist of timestamped edit actions and submitted solutions. at local level to study behavioural alignment, we combine the edit actions with the transcripts. at a dialogue level to measure performance in the task, we use the teams’ best score calculated from all the submitted solutions. • pre-test and post-test: at a dialogue level to measure learning outcomes, we use interlocutors’ scores in the pre-test and the post-test2 (see sec. 3.3 for the measure we adopt). 3.2 alignment in the justhink dataset we choose this dataset as it is particularly suited to study alignment, since it is designed in such a way to create interdependence, i.e. a mutual reliance to further the task, between the interlocutors. this interdependence requires interlocutors to align with each other on multiple levels, e.g. how to refer to the environment and how to represent the activity, in order to succeed. concretely, the activity enables: i) swapping and visual view control: since a turn changes every two edit actions, if an interlocutor has a particular action they want to take, they have to either wait for their turn in the visual view to implement the desired change, or instruct the other interlocutor. here, we follow an idealised perspective on the activity: at a given time, the interlocutor in the abstract view is an instruction giver (ig) who describes their instructions for the task by using specific referring expressions, and the other (in visual view) is the instruction follower (if) who executes the action (this is akin to the map task (anderson et al., 1991)). the activity design creates a frequent swapping of views. this aims to discourage interlocutors from working in isolation or in fixed roles of ig and if, which could potentially happen in collaborative tasks. ii) routine expressions and alignment in the task: since the robot uses brief and general instructions to present the activity and its goal, the interlocutors must figure out for themselves the way to approach the 2appendix d gives more details on the pre-test and the post-test. 9 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel activity. the task-specific referents, which are the names of the gold mines (cities in switzerland, a multi-lingual country), are potentially unfamiliar to interlocutors. interlocutors must refer to the task, and then align with the other to form routine expressions. by aligning, they establish a shared lexicon, and thus align their representations of the activity with each other. iii) submission of solutions: since the interlocutors have to submit their solution together by pressing the “submit” button that is present in both views, they have to, at least, align in terms of their intent to submit. alternatively, one interlocutor must convince the other to reach a common intent. 3.3 creating the justhink alignment dataset we put together an anonymised version of the original justhink dataset, called the justhink alignment dataset3, based on the dialogue transcripts, event logs, and test responses. we sampled the data because the state-of-the-art asr tools were not sufficiently good on the dialogues (see appendix a for a comparison). therefore, in order to study lexical and behavioural alignment, we need to manually transcribe the data: for this purpose, we develop a selection process so that the set is as representative as possible of the whole corpus. we firstly select a subset of 10 teams of the dataset. the teams were selected randomly according to the task success distribution (see fig. 3); which we measure through performance in the task and learning outcomes observed from the pre-test and the post-test. we thus choose a representation of the percentage of successful teams (30% compared to 21% of the whole dataset). transcription for our experiments, we focus on gold-standard/manual transcription, due to the poor performance of state-of-the-art automatic speech recognition (asr) systems on this dataset (which consists of children’s speech with music playing in the background). we give the accuracy of asr on the dataset in appendix a. table 1 provides further details about the transcribed subset. as shown, the mean duration of the task is ≈ 23 minutes, and the transcriptions account for ≈ 4 hours of data. the transcripts report which interlocutor is speaking (either a or b) and the start and end timestamps for each utterance, beside the utterance content. utterance segmentation is based on koiso et al. (1998)’s definition of an inter pausal unit (ipu), defined as “a stretch of a single interlocutor’s speech bounded by pauses longer than 100 ms”. we also annotated punctuation markers, such as commas, full stops, exclamation points and question marks. fillers, such as “uh” and “um”, (using meteer et al. 1995 for reference) were transcribed, as well as the discourse marker “oh”. the above 3 spontaneous speech phenomena occur frequently in the dataset {“um”: 236, “uh”: 173, “oh”: 333}. other phenomena, such as “ew” or “oops!” were also transcribed, however, their frequency is too low for analysis. transcription included incomplete elements, such as “mount neuchat-” in “mount neuchatum mount interlaken”. pronunciation differed among and within interlocutors (for example, for the word “montreux”, pronouncing the ending as /ks/ or /ø/), due to the unfamiliarity of the interlocutors with the referents, and individual accents. as our methodology is dependent on matching surface forms (refer to sec. 5), we standardise variations of pronunciation in the transcriptions, and we do not account for e.g. variations in accent. a graduate student completed two passes on each transcript, which were then checked by another native english speaking graduate student with experience in transcription/annotation tasks. 3which we make publicly available online, from the zenodo repository doi: 10.5281/zenodo.4627104. 10 http://doi.org/10.5281/zenodo.4627104 studying alignment in a collaborative learning activity via automatic methods 5 teams learn > 0 5 teams error > mean 5 teams error ≤ mean 5 teams learn ≤ 0 figure 3: scatter plot of the transcribed teams (red dots) and non-transcribed teams (blue dots) in the learning outcome (learn) vs. task performance (error) space. the black lines indicate the criteria for choosing the samples of the dataset to transcribe (i.e. 5 teams with learn > 0 . . . ). the mean of a set (transcribed or other) is shown as a dashed line, with the fit of a univariate kernel density estimate for the corresponding set. numbers denote the id of teams. table 1: descriptive statistics for the transcribed teams (n = 10). sd stands for standard deviation. mean sd min max number of submitted solutions 9.4 4.7 4 19 number of turns in task 50.4 21.4 28 98 total duration (mins) 23.7 7.1 11.2 36.0 time per submission (mins) 2.6 2.3 0.7 12.8 duration of a turn (secs) 25.0 30.3 1.1 240.2 length of utterance (in tokens) 6.6 5.4 1 32 11 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel dialogue level labels: measuring success in the task for task performance, we consider error which is the scaled difference of each submitted solution compared to the optimal solution, i.e. error = (cost− optimal cost)/optimal cost. at the dialogue level, we take the lowest error: this represents the team’s closest solution to an optimal solution. learning measures commonly build upon the difference between the post-test and pre-test results, e.g. in sangin et al. (2011); which indicates how much an interlocutor’s knowledge on the subject has changed due to the activity. we measure the learning outcomes on the basis of the relative learning gain (learnp) of an interlocutor p , which essentially is the difference between pre-test and post-test, normalised by the margin of improvement or decline (sangin et al., 2011). it is computed as: learnp = { post−pre max score−pre , post ≥ pre post−pre pre , post < pre. (1) it indicates how much the interlocutor learnt as a fraction of how much the interlocutor could have learnt. we use learn, the average relative learning gain of both interlocutors, to measure a team’s learning outcomes. 4. research questions the present work focuses on the task-specific referents that interlocutors minimally require to succeed in the task. therefore, we restrict the possible referring expressions to ones that contain taskspecific referents, in particular, to the objects that the interlocutors are explicitly given on the map (see fig. 2). interlocutors only need this terminology with certain function words to progress in the task (e.g. “montreux to basel”). we believe that this design choice is particularly suited to study alignment in this type of activity, as the frequent swapping of views encourages the interlocutors to communicate with the other their intents using these referents. this allows us to focus on verbal contributions that are explicitly linked to the ‘situatedness’ of the task and its association to a final measure of task success. while there are certainly other referring expressions to consider (e.g. “that mountain there”) not containing task-specific referents, it would require some degree of manual annotation. we thus focus on task/domain specific referents that can be automatically extracted. with this in mind, rq1 considers lexical alignment, i.e. the use of task-specific referents, while rq2 considers behavioural alignment, i.e. the follow-up actions taken after these task-specific referents were uttered. thus rq1 focuses on “what did the interlocutors say”, while rq2 builds on this with actions, i.e. “what did the interlocutors do (afterwards)”. we investigate how these local alignment contexts could build to a function of dialogue level task success. please refer to fig. 4 for an overview of the rqs. rq1 lexical alignment: how do the interlocutors use expressions related to the task? is this associated with task success? in rq1, we specifically consider the link between expressions related to the task and task success through the routines’ i) temporality, and ii) surrounding hesitation phenomena. we expand on these in sec. 5. specifically, we hypothesise: • h1.1: task-specific referents become routine early for more successful teams. we expect more successful teams to establish routine expressions earlier in the dialogue. ideally by quicker establishment of routine expressions, teams will understand each other faster and thus have greater task success. 12 studying alignment in a collaborative learning activity via automatic methods • h1.2: hesitation phenomena are more likely to occur in the vicinity of priming and establishment of task-specific referents for more successful teams. we expect that new contributions to the dialogue, via priming (the speaker first introduces the referent) and establishment (the listener utilises this referent for the first time) of routine expressions are associated with hesitation phenomena. a prolonged occurrence of hesitation phenomena not associated with the priming or establishment of routine expressions could highlight greater lack of understanding of the task, and hence be related to lower task success. rq2 behavioural alignment: how do the interlocutors follow up these expressions with actions? is this associated with task success? rq2 investigates how the use of these expressions manifests in the interlocutors’ actions within the task, and whether this is associated with their task success. for this purpose, we consider the instructions of an interlocutor, as the verbalised instructions one interlocutor gives to the other, which we extract through their use of task-specific referents. a physical manifestation of this instruction could result in a corresponding edit action, or a different edit action. we investigate the follow-up actions of the task-specific referents which its effect on task success, through the follow-up actions’ i) temporality, and the ii) surrounding information management phenomena. we expand on these in sec. 6. we hypothesise: • h2.1: instructions are more likely to be followed by a corresponding action early in the dialogue for more successful teams. we expect that the earlier interlocutors align with each other in terms of instructions and follow-up actions, the better they progress in the task, and the greater the chance of success in the task. this idea of verbalised instructions being followed up by corresponding actions is in line with previous research on alignment (i.e. interlocutors being in alignment in a successful dialogue). however, work on collaborative learning suggests that individual cognitive development (in our case, positive learning outcomes) happens via socio-cognitive conflict (mugny and doise, 1978; doise and mugny, 1984), and its regulation (butera et al., 2019). in our task, this means a verbalised instruction could be followed by a corresponding or a different action; as a different action could result in collaboratively resolving conflicts and together building a solution – resulting in task success. • h2.2: when instructions are followed by a corresponding or a different action, the action is more likely to be in the vicinity of information management phenomena for more successful teams. since the task involves the creation of a joint focus of attention between the interlocutors, we expect there to be verbalised information management markers present (such as “oh”). ultimately, an increase in these information management markers associated with the task should lead to an increase in task success. 5. studying lexical alignment (rq1) – how expressions related to the task contribute to lexical alignment a routine is formed when a referring expression is commonly used by both interlocutors. for example, “montreux” is a task-specific referent, and one interlocutor might prime the referring expression “mount montreux”. if the other interlocutor reuses this referring expression, it becomes a routine, i.e. a common part of the dialogue. for our methodology, a routine is specific to the exact matching of token sequences in two utterance strings. we thus formally define a routine expression (adapted from dubuisson duplessis 13 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel figure 4: the research questions investigated in this article. et al. 2017, 2021; pickering and garrod 2004) as a referring expression shared by two interlocutors if i) the referring expression is produced by both interlocutors, and ii) it is produced at least once without being part of a larger routine. in particular, we define the utterance at which a referring expression becomes routine as the establishment of that routine. we extract the utterances at which the routines are primed and established from the transcripts as in dubuisson duplessis et al. (2017, 2021). then, we filter for the routines that contain a task-specific referent. 5.1 experiment 1: studying when interlocutors are lexically aligned and its association with task success (h1.1) 5.1.1 methodology to investigate when the routine expressions become established, we study i) the establishment time of a routine, i.e. the end time of the utterance at which the expression is established, and ii) the collaborative period of a team, i.e. the duration between first quartile (q1) to third quartile (q3), i.e. the interquartile range (iqr) of the establishment, where half of the establishments occur. to then study h1.1, for task performance, we check if the median establishment times are significantly earlier for better-performing teams by spearman’s rank correlation and its statistical significance (between the median and the error). for the learning outcomes, we compare the distribution of the median establishment times of teams with positive learning outcomes (teams 7, 9, 10, 11, 17) and others (teams 8, 18, 20, 28, 47) by kruskal-wallis h test.4 we consider i) the establishment times of all the routines in real time (as teams took different durations to complete the task), ii) the establishment times of the routines that are established in the common time duration (i.e. the first ≈ 11 mins for all the teams, which was the time taken by the quickest team), and iii) normalised establishment times, that are scaled by the duration of each team itself to reflect the ‘progress’ of a team’s interaction, from 0% progress at the beginning of the activity, to 100% when the interaction ends (either by finding an optimal solution, or by being intervened by the experimenters to end the task). to gain further insight, we compare the distribution of the establishment times (from i) to iii)) of the teams, through inspecting the box plots of all teams that are compared side-by-side and sorted by decreasing task success. 4since we have groups of 5 teams for learning, we use kruskal-wallis (scipy’s implementation of kruskal-wallis works with ≥ 5 samples). this can not be used for performance. 14 https://www.scipy.org/ studying alignment in a collaborative learning activity via automatic methods 5.1.2 results and discussion empirical results fig. 5 shows for each team the distribution of the establishment times of the routines, sorted by task performance. the median establishment times for all of the routines is strongly positively correlated with the task performance measure error (spearman’s ρ = 0.69, p < .05), that is, better-performing teams establish routines earlier.5 we see from the figure that the results are influenced by the variation in the duration of the activity: the interaction ends for the well-performing teams when they find a correct solution, whereas the badly-performing teams continue their interaction until the experimenters intervene and stop the activity. fig. 6 shows the normalised establishment times. we see that the establishment occurs around the middle of the dialogue: establishment times have mean of the medians = 63.0% (combined sd = 22.2%). while we hypothesised that establishment will happen early in the dialogue, this is the case for an ideal dialogue; people will ‘share’ expressions earlier. however, we observe that there is an exploratory period (the period before q1), where the interlocutors take the time to understand the task, followed by a collaborative period that corresponds to the establishment period (the period between q1 and q3). we expect that if the interlocutors had to complete the task again, the establishment/collaborative period would be closer to the start of the dialogue. while it is expected that better-performing teams have established routines, from the figures we observe even badly-performing teams still successfully established routines. thus it is to be noted that all teams, regardless of performance, have aligned to some degree. we observe that five teams (7, 8, 10, 18, and 47) that had their error > mean (see fig. 3) started collaborating later in their dialogue, in terms of when they establish most of their routines (median establishment time > 60%). though the measures of task performance would only reflect that these teams performed badly with a final overall score of performance, our alignment measures reflect details of the performance. interlocutors are gently reminded a few minutes before the end of the task the remaining time, but were not rushed to find a solution: which could have changed the way they aligned by forcing an establishment period. thus, we observe through our alignment measures that badly performing teams were simply slower to collaborate and establish routines. the distribution of the median establishment times are not statistically significantly different for teams with positive learning outcomes (teams 7, 9, 10, 11, 17) vs. others (teams 8, 18, 20, 28, 47) by kruskal-wallis h test (h = 0.88, p = .35 for times in absolute values, and h = 0.10, p = .75 for normalised times). further discussion we would like to discuss the notion of collaborative and exploratory periods by examining the error progression of teams’ submitted solutions during the task (fig. 6). during the exploratory period, nine out of ten teams had their highest error. for teams that performed well, their collaborative period had their lowest error before finding an optimal solution (thus ending the task), and for other teams; their closest solution to the optimal solution. there are 8 such teams that exhibit this behaviour. the teams that did not achieve a correct solution, never regressed to their largest error from the exploratory period. looking at this through the number of attempts, for the well-performing teams, their collaborative period was productively used, with their next solution reaching an optimal cost (teams 17 and 28), or one more attempt before their optimal cost (team 20). several teams that did not perform 5the magnitude of spearman’s correlation coefficient (ρ) can be interpreted by using the thresholds e.g. from evans (1996), i.e. 0.00−0.19 “very weak”, 0.20−0.39 “weak”, 0.40−0.59 “moderate”, 0.60−0.79 “strong”, and 0.80−1.00 “very strong” . 15 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel 0 5 10 15 20 25 30 35 establishment time (mins) 28 17 20 11 9 47 10 8 7 18 te am n o learn + + + + +191 19132 36 41 41 41 3627 45 41 23 50 32 23 14 32 14 27 32 14 41 14 27 50 3241 36 32 36 32 27 23 5 18 14 18 9 32 68 68 23 55 18 14 14 18 18 18 23 45 191 127 64 59 45 59 32 9 41 5 14 41 32 23 5 0 45 45 18 23 41 191 109 109 0 36 27 18 5 0 23 73 14 18 27 23 18 14 3214 23 we llpe rfo rm in g (in cr ea sin g er ro r) ba dl ype rfo rm in g (in cr ea sin g er ro r) figure 5: the team’s establishment times for h1.1. the teams are sorted by decreasing task performance (i.e. increasing error). teams that have the same colour (such as team 11 and 9) have the same error, and are sorted by increasing duration for the ties.. the thick, bold boxplots with whiskers with maximum 1.5 iqr show the distributions of the establishments that occurred in the common duration (i.e. as marked by the dashed green line), while the thin boxplots show the distributions through the total duration of interaction. the learning outcome of each team is indicated with a plus (‘+’) for learn > 0, or a minus (‘−’) otherwise. solid lines indicate the end of the interaction, by submitting a correct solution (in green) or timing out (in red). the thin blue lines indicate submission of a solution, with the number showing how far the solution that was submitted is from the optimal solution (e.g. error = 0% means the team has found an optimal solution). the red and blue dots indicate the utterances of the interlocutors, to give an idea of when the interlocutors are speaking versus when are their establishment times. 16 studying alignment in a collaborative learning activity via automatic methods 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% establishment time (%) 28 17 20 11 9 47 10 8 7 18 te am n o 191 191 32 36 41 41 41 36 27 45 41 23 50 32 23 14 32 14 27 32 14 41 14 27 50 32 41 36 32 36 32 27 23 5 18 14 18 9 32 68 68 23 55 18 14 14 18 18 18 23 45 191 127 64 59 45 59 32 9 41 5 14 41 32 23 5 0 45 45 18 23 41 191 109 109 0 36 27 18 5 0 23 73 14 18 27 23 18 14 32 14 23 we llpe rfo rm in g ba dl ype rfo rm in g (in cr ea sin g er ro r) figure 6: the teams’ establishment times as normalised by the duration for each team separately, i.e. 100% indicates the end of that team’s interaction. sorting is as in fig. 5. 17 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel well have a greater number of attempts submitted after their collaborative period (teams 9, 7, 8, 10, 47). their submission pattern of attempts indicate a “trial-and-error” strategy on how to solve the task, for example, teams 8 and 9 increasing their submissions after their collaborative period, or teams 7, 10 and 47 submitting throughout. lastly, we would like to compare the team that had the most task success (team 17, with positive performance and learning gains) with the team that had the lowest task success (team 18). the two teams have established routines later in their dialogue: team 17 that found a correct solution, and team 18 that could not (their median establishment times are around 70%). yet, 17 had a focused establishment period (with a smaller iqr = 18% vs. 30% of the time, respectively, see fig. 6). we interpret that team 17 was able to turn it around and find a correct solution, while team 18 did not; ending up as the worst performing team. 17 has positive learning gain (learn = 25%), while it is negative for team 18 (learn = −56%, the highest decrease among the transcribed teams, see fig. 3). while bad performance could be reflected through a collaborative period starting later in their dialogue (such as team 18; with later establishment, and not finding a solution in time), team 17 shows that there are exceptions to this. team 18 also possibly got confused, reflected in the high and negative learn. synopsis overall, the results support h1.1 for task performance, while they are inconclusive for learning outcomes. surprisingly, we see that all teams, regardless of task success, are (lexically) aligned to some degree. thus, inspecting the distributions of the establishment times together with their proposed solutions gave further insight into the collaboration processes, specifically revealing a “collaborative period” that most of the teams came up with their best solutions. we believe that this captures some local level alignment patterns that might have otherwise been overlooked when only considering dialogue level task success. 5.2 experiment 2: studying hesitation phenomena surrounding lexical alignment and its association with task success (h1.2) 5.2.1 methodology to measure hesitation phenomena, we choose fillers (in particular “uh” and “um”) specifically due to their many links with the uncertainty of the interlocutor, be it by simple hesitation (pickett, 2018), deeper meanings of a speaker’s feeling of how knowledgeable they are (smith and clark, 1993), or even the listener’s impression of how knowledgeable is the speaker (brennan and williams, 1995). fillers can thus be used by the interlocutor to inform the listener about upcoming new information, or even production difficulties that they are facing. particular to the establishment of referring expressions, research has shown that disfluency (studied with the filler “uh”) biases listeners towards new referents (arnold et al., 2004) rather than ones already introduced into the discourse, and helps listeners resolve reference ambiguities (arnold et al., 2007). to investigate h1.2, we inspect the distribution of filler times, in relation to the i) establishment and ii) priming times of routine expressions. this considers when a speaker introduces a new taskspecific referent into the dialogue, and when the listener makes this expression routine for the first time. in particular, we note the order of the tokens in the dialogue for the filler positions and the first token of the priming/establishment instance. then, we check whether the distributions of filler times (by its token position) with establishment times and priming times are significantly different (by its first token’s position), by utilising a mann-whitney u test, and estimate the effect size by 18 studying alignment in a collaborative learning activity via automatic methods table 2: summary statistics for fillers and the established routines of teams, sorted by decreasing task performance. count median (%) team filler routine filler priming establishment 28 38 58 11.6 3.5 15.2 17 27 41 14.0 6.2 17.1 20 59 51 20.7 18.1 22.8 11 35 70 34.5 3.5 24.4 9 81 131 46.3 11.4 37.4 47 56 74 19.5 6.7 21.9 10 20 65 20.1 8.7 20.4 8 26 62 19.8 15.7 36.0 7 27 59 24.9 25.9 37.8 18 52 57 10.8 4.3 11.8 computing cliff’s delta. then, we compare these results for the teams as sorted by increasing task success: for task performance, we compute the spearman’s correlation and its significance between cliff’s delta and error. for learning outcomes, we perform kruskal-wallis h test to compare the distribution of cliff’s delta values for the two groups of learning. 5.2.2 results and discussion we consider hesitation phenomena as cued by the presence of a filler, and thus investigate how are the fillers distributed as compared to the priming and establishment of the routines. table 2 presents the count and median times of fillers and established routines for each team. the results for mann-whitney u test and effect size as estimated by cliff’s delta (δ) are given in table 3. to interpret the results, δ ranges from −1 to 1, where 0 would indicate that the group distributions overlap completely; whereas values of −1 and 1 indicate a complete absence of overlap with the groups. for example, −1 indicates that all fillers occur before priming times, and 1 indicates that all fillers occur after priming. empirical results we observe that filler and priming times differ significantly for most of the teams (for eight out of ten teams by mann-whitney u test p < .05). we see positive δ values, except for one team. thus most teams have fillers that occur after priming times. we see that the magnitude of δ for the significant teams is large except for team 7.6 we interpret that most fillers do occur visibly after priming times (with little overlap from the large effect sizes). we observe that filler and establishment times differ significantly for most of the teams (for seven out of ten teams by mann-whitney u test p < .05). we see δ ranging from−0.67 for team 7, to 0.33 for team 11 (though most are negative). however, we see that the magnitude of δ for most teams is small, with the exception of two teams. this means that the distributions of fillers differ with a small effect size compared to the distributions of establishment times, especially considering 6the magnitude of cliff’s delta (δ) can be interpreted by using the thresholds from romano et al. (2006), i.e. |δ| < 0.147 “negligible”, |δ| < 0.33 “small”, |δ| < 0.474 “medium”, and otherwise “large”. 19 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel table 3: the results for the mann-whitney u test that compares the distribution of filler times (both “uh” and “um”) with establishment and priming times for h1.2. the effect size is estimated by cliff’s delta (δ). teams are sorted by decreasing task performance i.e. increasing error. the horizontal line separates well-performing teams (that found a correct solution) from badly-performing teams. u is the u statistic, and p is the ‘twosided’ p-value of a mann-whitney u test (without continuity correction as there can be no ties, via our unique token number assignment). priming establishment team u p δ u p δ 28 1600.5 < .05 0.45 828.5 < .05 -0.25 17 823.0 < .05 0.49 431.0 .12 -0.22 20 1746.5 .15 0.16 1139.5 < .05 -0.24 11 2229.0 < .05 0.82 1624.0 < .05 0.33 9 8805.5 < .05 0.66 6561.5 < .05 0.24 47 3189.0 < .05 0.54 1627.0 < .05 -0.21 10 982.0 < .05 0.51 660.0 .92 0.02 8 883.0 .48 0.10 461.0 < .05 -0.43 7 562.0 < .05 -0.29 259.0 < .05 -0.67 18 2136.0 < .05 0.44 1326.0 .34 -0.11 that the token number values do not overlap; i.e. we do not expect an effect size of 0. we thus interpret most fillers do occur around establishment times (from positive and negative δ values), with larger overlap given the small effect sizes. the distribution of the cliff’s delta between filler times and both priming or establishment times has a very weak correlation coefficient with error (or task performance) (spearman’s ρ = −0.18, p = .62 for priming, ρ = 0.06, p = .88 for establishment). therefore, we can not conclude that there is a significant relationship between how early are the priming or establishment times, and how well a team performs. similarly, the distribution of the cliff’s delta that quantifies how different are filler times from priming or establishment times are not statistically significantly different for teams with positive learning outcomes (teams 7, 9, 10, 11, 17) vs. others (teams 8, 18, 20, 28, 47) by kruskal-wallis h tests (h = 1.32, p = .25 for both). therefore, we can not conclude that a significant difference exists, between how early fillers occur compared to priming (or establishments) between these two learning groups. further discussion the empirical results regarding the priming and establishment of the filler indicate that in the process of routine formation, in between the priming and the establishment of the expression, there seems to be a period in which the interlocutors use fillers. this placement of fillers is of interest due to the potentially unfamiliar vocabulary of the task-specific referents that the interlocutors had to utilise in the situated environment. the results demonstrate a lack of fillers at the start of the formation of a routine; i.e. they occur visibly after the priming of expressions that contain task-specific referents. the large effect size shows that this is predominantly the case for most teams. in addition to this, fillers were found to occur around establishment times. we suggest 20 studying alignment in a collaborative learning activity via automatic methods that a part of establishment is often a clarification request, as shown by the following example. red indicates the routine expression, while blue indicates the filler. a: mount zurich to mount bern. . . . . . . b: uh isn’t already mount zurich to mount bern isn’t it connected? . . . . . . a: so if we erase mount zurich or mount berto mount bern or mount zurich to mount gallen? b: wait uh zurich to mount . . . ? (2) speakers are able to prime these expressions without using fillers, but this does not guarantee that the primed expression was fully understood by the listener, as shown by the results of the mannwhitney u test. following the idea of if and ig pairs in the dialogue, the ig (speaker) tends to be in the abstract view, and can see the minimally represented names of gold mines. they need only concentrate on the gold mines that have the lowest cost to connect. perhaps the reason why the ig does not use (by our measures) fillers as much is because they can see the task-specific referents available to them written down on the map. expressions that contain task-specific referents could successfully become part of the ig’s expression lexicon when given these new referents. the ig may also feel at ease to read out these expressions. the if (listener) must follow the instructions with actions, and since they can not see the cost of adding and removing edges, they must search for the specific gold mine names given by the ig. this could create an uncertainty in the ig regarding how much they think that the if has understood, and bring about a need to clarify. synopsis regarding h1.2, we can not conclude that the usage of fillers occur in the vicinity of priming or establishments times for more successful teams (for neither better-performing teams, nor higher-learning teams). overall, we see that for most of the teams, the fillers tend to occur visibly after priming times, and around establishment times. this shows that while fillers are associated with local alignment contexts, through our methodology, we cannot conclude about their contribution to overall task success. 6. studying behavioural alignment (rq2) if an instruction is verbalised (e.g. to connect “mount basel to montreux”) by an interlocutor (ig), it could result in an action of connecting the two. we hence say an instruction matches an action when the instruction is executed by the other interlocutor (if) via an action in the situated environment, and within the period of a turn of views before they are swapped. we study the discrepancy created when the if does not follow the ig, which we call a mismatch of instructions-to-actions. in example 3, the instruction (to connect gallen to davos) matches the action (connecting these two). in example 4, we illustrate a dialogue excerpt that results in a mismatch: 21 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel ... tn utterance tn+1 action tn+2 action tn+3 utterance tn+4 utterance ... yes no swap or submit? clear instructions yes utterance action kind of event no recognise instructions match instructions to actions dialogue and actions infer & add instructions no has referent? the action is nonmatched the action is matched event at tk new pending instructions list yes instructions pending? clear instructions clear instructions yes no matches any instruction by the other? the action is mismatched figure 7: representation of the schema in a flowchart (e.g. the parallelogram shows the annotated output), to annotate behavioural alignment between the interlocutors. a: what about mount davos to mount, saint gallen? [instruction to add an edge from davos to gallen.] b: because what if, you did if we could do it? a: what about mount um davos to mount gallen? b: mount . . . b: oh mount davos a: yeah to mount gallen. b: to mount gallen yeah do that. 〈a connects mount gallen to mount davos〉 b: okay my turn. [match!] (3) b: go to mount basel. [instruction to add an edge to basel.] a: that’s, it’s expensive. b: just do it. a: you can’t, you can’t, i can’t because there’s a mountain there . . . a: so i’m going, so i’m going here. 〈a connects mount interlaken to mount bern〉 [mismatch!] (4) recognising instructions we firstly extract instructions from the utterances through the interlocutors’ use of task-specific referents. in the schema as shown in fig. 7, our input consists timestamped dialogue transcripts and action logs. when the input is an utterance, it is processed to infer instructions using entity patterns. to check these entity patterns, we employ named-entity recognition 22 studying alignment in a collaborative learning activity via automatic methods table 4: an example from the output of recognising instructions and detecting follow-up actions (by recognise-instructions (algorithm 1) and match-instructions-toactions (algorithm 3)), from team 10. view denotes which view the interlocutor is in (refer to sec. 3.1), either abstract (ab) or visual (v). annotations denotes the automatically inferred instructions and follow-up actions in the activity. for example, instructa indicates that interlocutor a has given an instruction to add two nodes (inferred from referents), which can be partially recognised (gallen,?). as shown, the algorithm builds up (or “caches”) instructions until an edit action is performed (‘-’ in utt.). note, since b is in the visual view, their inferred instruction is deliberately not matched. utt. view verb utterance annotations a 198 ab says maybe we start from, instructa(add(zermatt,?)) mount zermatt ? b 199 v says no lets do mount davos to, instructa(add(zermatt,?), instructb(add(davos,?)) where do you wanna go? a 200 ab says . . . to mount, st gallen. instructa(add(zermatt,?), instructb(add(davos,?) , instructa(add(gallen,?)) b 201 v says okay. as previous b v adds gallen-davos instructa(add(zermatt,?), matchb (instructa(add(gallen,?))) (ner) feature of the python library spacy that performs this entity recognition. we add the node names of the mountains (e.g. “montreux”), and also verbs; i.e. “add”, “subtract” . . . then, if the input utterance contains our custom entity patterns, then we automatically infer instructions from the utterance by joining these entities together. for example, the result may be add(node1,node2), because the interlocutor explicitly said the verb “add” and also the names of the mountains. we give the complete algorithm in appendix e. matching instructions-to-actions to find (mis)matches of instructions-to-actions, we then follow the logic as given in the schema to determine whether there is an inferred instruction in the instructions list at the time an action was taken. then, we check whether the action matches or mismatches the inferred instruction. in example 3, gallen-davos was the result of a negotiation, rather than a complete given instruction by the ig. b says in one utterance “oh mount davos” and then in another “to mount gallen yeah do that”, resulting in two cached inferred instructions; (davos,?) and (gallen,?) respectively. this accounts for some amount of multiple speaker turn7 negotiations, and possible other ways of referring to task-specific referents (e.g. “now go from there to davos”). though an instruction could be carried out after the views swap again, i.e. in the following turn, in our methodology, the pending instructions are cleared at every swap (if “swap or submit”, then “clear instructions”), resulting sometimes in a nonmatched action, or when an action occurred but there was no inferred 7note, we specifically use “speaker turn” to distinguish from a turn in the collaborative activity; i.e. every swapping of views. 23 https://spacy.io/ utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel instruction. a full and concrete example of the added annotations of our automatically inferred instructions-to-actions is given in table 4. 6.1 experiment 3: studying when interlocutors are behaviourally aligned and its association with task success (h2.1) 6.1.1 methodology to investigate h2.1, we compare the distributions of the match and mismatch times of the teams, by following the same methodology stated for h1.1 (which investigated establishment and priming times). in particular, we check if the median match or mismatch times are significantly earlier for better-performing teams by spearman’s rank correlation, and for the two groups of learning, by kruskal-wallis h test. 6.1.2 results and discussion empirical results the median match times has a moderate positive correlation coefficient with the task performance measure error (spearman’s ρ = 0.59, p = .08). this means that betterperforming teams tend to have matches earlier. while close to significant (however, p > .05), this follows our results regarding lexical alignment; i.e. better performing teams align both verbally and behaviourally earlier (in terms of matched instructions-to-actions) than badly performing teams. h2.1 is weakly supported by these results. better-performing teams tend to have mismatches earlier as well (spearman’s ρ = 0.70, p = .024). fig. 8 shows the distribution of the (mis)match times for each team in a common time frame. when the times are normalised by the duration for each team separately, we still observe a positive, moderate spearman’s correlation coefficient (ρ = 0.54, p = .11), that indicates a general trend that matches happen earlier for the better-performing teams. fig. 9 shows for each team, the distribution of the match and mismatch times, normalised by the duration of each team itself. while it is natural to think that instructions will be followed by a match for more successful teams in a structured and organised manner, we observe that teams that performed badly and teams that did not learn also have a certain period of matches. therefore independent of their task success, all teams have their ig’s instructions matched to the if’s actions to some extent. match times have mean of the medians = 62.6% (combined sd = 21.9%), and mismatches have mean of medians = 67.6% (combined sd = 26.0%). by inspecting the median of the match times with learning outcomes, we observe that they are not significantly different for teams with positive learning outcomes (teams 7, 9, 10, 11, 17) vs. others (teams 8, 18, 20, 28, 47) by kruskal-wallis h tests (h = 0.54, p = .47 for times in absolute values, and h = 0.10, p = .75 for normalised times). further discussion by inspecting the outputs of the automatically inferred instructions in table 4, occasionally, we see that the traditional roles of if and ig are not maintained, even for the brief fixed time of views (i.e. within a turn). the if, having just been switched from ig, may have their own ideas about the next best edit action, as seen in utt. no. 199 – prior to utt. no. 198, the views were swapped, with b being in the abstract view, meaning b was just previously the ig. thus the interlocutors can also conduct a negotiation where they collaboratively decide which action to take. it is therefore not always the case that the ig is in the abstract view deciding and instructing what the if in the visual view should do; and this is supported by the frequent swapping of views. 24 studying alignment in a collaborative learning activity via automatic methods 0 5 10 15 20 25 30 35 time (mins) 28 17 20 11 9 47 10 8 7 18 te am n o learn + + + + +191 19132 36 41 41 41 3627 45 41 23 50 32 23 14 32 14 27 32 14 41 14 27 50 3241 36 32 36 32 27 23 5 18 14 18 9 32 68 68 23 55 18 14 14 18 18 18 23 45 191 127 64 59 45 59 32 9 41 5 14 41 32 23 5 0 45 45 18 23 41 191 109 109 0 36 27 18 5 0 23 73 14 18 27 23 18 14 3214 23 we llpe rfo rm in g (in cr ea sin g er ro r) ba dl ype rfo rm in g (in cr ea sin g er ro r) figure 8: the team’s match and mismatch times in blue and red respectively, to study h2.1. the teams are sorted by decreasing learning outcomes (i.e. decreasing learn). see fig. 5 for a description of the green, red and blue lines. 25 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% time (%) 28 17 20 11 9 47 10 8 7 18 te am n o 191 191 32 36 41 41 41 36 27 45 41 23 50 32 23 14 32 14 27 32 14 41 14 27 50 32 41 36 32 36 32 27 23 5 18 14 18 9 32 68 68 23 55 18 14 14 18 18 18 23 45 191 127 64 59 45 59 32 9 41 5 14 41 32 23 5 0 45 45 18 23 41 191 109 109 0 36 27 18 5 0 23 73 14 18 27 23 18 14 32 14 23 we llpe rfo rm in g ba dl ype rfo rm in g (in cr ea sin g er ro r) figure 9: the teams’ match and mismatch times in blue and red respectively, as normalised by the duration for each team separately. sorting is as in fig. 8. 26 studying alignment in a collaborative learning activity via automatic methods what is interesting, is that from the behavioural alignment algorithm, we see that all teams have more inferred matches than mismatches (see table 5 for reference). teams 18 and 20 that learnt the least, have the highest match-to-mismatch ratio (3.6 and 4.0 matches of instructions-to-actions for every 1 mismatch respectively.), as well as team 7 (3.3). we observe that teams 10, 11 and 17, who had the highest learning gain, comparatively have a lower ratio of matches-to-mismatches of 2.4, 1.8 and 2.3 respectively. while we could not conclude on the statistical significance, we interpret that it is important for learning that interlocutors have a good ratio of matches-to-mismatches, and not just be in total (blind) agreement with the other. this is consistent with previous research, as stated in sec. 4. by just inspecting performance, superficially, team 20 performed well. however, they did not learn (in fact, “unlearnt”, as shown with a negative learning outcome). positive learning teams seem to have conflict(s) that they resolved collaboratively while building and submitting solutions, whereas the others either had less or unresolved conflicts. by inspecting fig. 9 for how the (mis)matches are distributed, and how they are they positioned in relation to the costs of the submitted solutions, we observe nuances of conflicts and their resolutions. for instance, consider a high-learning team 11. the team began with a high cost of 191%, i.e. they basically connected everything to each other, with many redundant connections. they initially misunderstood the goal of the task of connecting the graph minimally. then, they had many mismatches (“a mismatch period”), which is followed by many matches ending with (at q3 of matches) their best solution, getting very close (5%, i.e. they did not notice that they could replace a particular connection with a better/lower-cost one). although the team could not find a correct solution, they learnt a lot from this interaction flow. in comparison, team 20 also started with that high cost of 191%. yet, they could not resolve the conflict, as their (mis)matches did not result in them getting closer to an optimal solution, but rather still keeping to high cost solutions, and repeating these high cost submissions (109%, i.e. more than double the cost they could make with many redundant connections). although they found a correct solution next, since they did not have the conflict resolution period that could be cued by (mis)matches, they did not end up learning. synopsis we see that all teams are behaviourally aligned, irrespective of their task success. in terms of task performance, we observe a general trend that supports the h2.1, i.e. better performing teams tend to follow up the instructions with actions (as matches) early overall (the absolute times), as well as in their dialogues (normalised times). for learning, although the general trend is not statistically confirmed, we gained insight into the nuances of the dynamics of interaction: by inspecting the normalised (mis)match time plots alongside with the submission costs. there seems to be conflicts that may be collaboratively resolved and resulted in learning, or may not be resolved and have adverse effects on the learning outcomes. 6.2 experiment 4: studying information management phenomena surrounding behavioural alignment and its association with task success (h2.2) 6.2.1 methodology we consider the use of “oh” as an information management marker, to mark a focus of a speaker’s attention, which then also becomes a candidate for the listener’s attention. this creation of a joint focus of attention allows for transitions in the information state (schiffrin, 1987). in general, changes in the information state – or what is commonly known about the task – of the participating interlocu27 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel tors should be due to physical actions related to the task, particularly because the situated activity has an interdependence on the physical environment. “oh” as a marker for new information has been studied in various scenarios. aijmer (1987) did a corpus analysis of “oh” to identify the contexts in which the marker is used, starting from a base general description of its usage as a mental reaction to a stimulus; e.g. “oh, flowers!”. among the specific contexts identified, the most relevant to this work is “oh” used as a marker to verbalise reactions to surprising information. fox tree and schrock (1999) also studied the use of “oh” in online comprehension experiments to find that it is used by listener’s to help them integrate information in spontaneous speech. we are interested in the use of “oh” in the context of what is commonly known by the interlocutors regarding the task. thus, we consider the verbalisation of “oh”, as an update in the interlocutor’s knowledge regarding how to solve the situated activity, resulting from actions taken. to the best of our knowledge, the use of “oh” has been studied more in the context of integrating information at the lexical level, but not a behavioural one. thus, an increase in these information management markers associated with the task should lead to an increase in task success. to investigate h2.2, we check if the distributions of (mis)matches and “oh” marker times are significantly different by mann-whitney u test, and estimate the effect size by computing cliff’s delta. this calculation differs from h1.2, as actions are not part of an utterance (compared to establishment time for example); hence we cannot use the order of tokens as was previously used, and instead compare the end times of the utterance that contains the marker with the (mis)matched action times. then, as h1.2, we calculate the relation to performance and learning. 6.2.2 results and discussion empirical results we note that the “oh” marker occurs 333 times in the transcripts (average per transcript = 33.3, sd = 20.3). see table 5 for the number of utterances that contain one or more “oh”s for each team. the results of the tests for each team are given in table 6. here, δ = −1 would mean that all “oh”s occur earlier than (mis)match times, and 1; that all “oh”s occur later.8 the δ values vary from 0 to negative values (see table 6), for all teams except one. this indicates that the distribution of “oh” tends to occur earlier than (mis)match times. observe that “oh” and (mis)match times do not significantly differ for half of the teams (mann-whitney u test p < .05). we see that for the teams that had significantly different distributions, the effect size only ranges from negligible (|δ| < 0.147) to medium (|δ| < 0.474), with the exception of one team. this indicates that overall, “oh” tends to occur close to mismatch times, though earlier, with larger overlap from smaller effect sizes, and half the teams not having distributions significantly different by this test. following this, the δ values between “oh” times and (mis)match times have a medium negative correlation coefficient with performance (spearman’s ρ = −0.53, p = .12). while close to significant (however, p > 0.05) we see the general trend that for better performers there is more the overlap between (mis)match times and “oh”s. thus is a general trend that for well-performing teams the “oh”s occur more in the vicinity of (mis)match times with larger overlap between the two groups, and the badly-performing teams have comparably less overlap. thus the results weakly support h2.2, however this would require more data to be verified. 8as h1.2, δ ranges from −1 to 1, where 0 would mean that the group distributions overlap completely; whereas values of −1 and 1 indicate a absence of overlap between the groups. 28 studying alignment in a collaborative learning activity via automatic methods table 5: summary statistics for the “oh” and (mis)matched instructions-to-actions, sorted by decreasing task performance. count is given as the number of utterances that contained (mis)matched actions. count median (%) team oh match mismatch match/mism. oh match mismatch 28 24 19 13 1.5 61.7 62.7 73.6 17 20 23 10 2.3 61.1 61.1 69.4 20 65 24 6 4.0 35.0 58.0 42.3 11 15 34 18 1.8 57.3 59.5 38.5 9 29 28 9 3.1 42.9 59.9 70.2 47 29 28 21 1.3 53.4 56.8 75.7 10 18 36 15 2.4 47.6 67.3 65.3 8 55 30 11 2.7 36.3 65.3 86.3 7 51 43 13 3.3 49.7 65.5 79.7 18 12 25 7 3.6 69.6 69.5 75.0 the distribution of the cliff’s delta, that quantifies how different “oh” times are from (mis)match times, is not statistically significantly different for teams with positive learning outcomes (teams 7, 9, 10, 11, 17) vs. others (teams 8, 18, 20, 28, 47) by kruskal-wallis h test (h = 0.27, p = .60). further discussion we see from table 6 there is a delicate balance between performance and learning. see table 7 and table 8 for excerpts from teams 20 (performed well but learnt nothing) and 17 (i.e. performed and learnt well), respectively. for both teams, we see the first occurrences of “oh” in the tables (corresponding to the exploratory period, from h1.1) and a period of nonmatched instructions for both teams, as they individually figure out the constraints of the task and navigate the situated environment. in team 20 (utt. no. 56 − 61), we see both interlocutors using “oh” traditionally as an information management marker (e.g. “oh i think we have to connect all of them”); but from the mismatch/nonmatch of instructions-to-actions, we see that this information state of the interlocutor is not transferred to the collective information state of both interlocutors. essentially, the interlocutors are working in isolation, individually gaining (perceived) information about the situated environment (measured by use of “oh”), and then following up with their own intentions in isolation (e.g. mismatch and nonmatch between utt. no. 56 and 57). to contrast, we clearly see that team 17 (utt. no. 61 − 65) has this transition of information from the interlocutor to the listener with their use of “oh”, signalling information in a shared focus of attention (with a saying “oh no that costs more” and b responding “we should erase it”). we see similar patterns of behaviour towards the end of the dialogue (by this point, interlocutors have had opportunity to build a collaboration with each other). the intuition behind h2.2 seems reasonable here; team 20 for example, had 65 occurrences of the information marker “oh”, but they significantly differ from the 30 times the instructions given by the ig was (mis)matched by the if. interestingly, we observe in the last period for team 20 a lack of inferred instructions. inferred instructions are a precursor to an action being (mis)matched or nonmatched. indeed, there is also a lack of nonmatched actions. here, they do not verbalise 29 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel table 6: the results for mann-whitney u tests that compare the distribution of information marker (i.e. “oh”) times, with (mis)match times for h1. the effect size is estimated by cliff’s delta (δ). teams are sorted by decreasing task performance i.e. increasing error. the horizontal line separates well-performing teams (that found a correct solution) from badlyperforming teams. u is the u statistic, and p is the ‘two-sided’ p-value of a mann-whitney u test (without continuity correction as there can be no ties, via our unique token number assignment). (mis)match team u p δ 28 344.0 .51 -0.10 17 318.0 .83 -0.04 20 715.0 < .05 -0.27 11 427.0 .58 0.09 9 429.0 .16 -0.20 47 488.0 < .05 -0.31 10 264.0 < .05 -0.42 8 522.0 < .05 -0.54 7 841.0 < .05 -0.41 18 157.0 .36 -0.18 their own intentions, visible from the lack of inferred instructions and nonmatched actions. this indicates a strong tendency of working in isolation despite the design of the task. synopsis overall, regardless of a team’s task success, we see that “oh” does occur earlier than (mis)matched instructions-to-actions, as evidenced by the negative effect sizes of cliff’s delta. specifically for performance, there is a general trend that for well-performing teams the “oh”s occur more in the vicinity of (mis)match times with larger overlap between the two groups, and the badly-performing teams have comparably less overlap. this result weakly supports h2.2. while we cannot conclude on the results for learning, using our automatically annotated instructions-toactions and occurrences of “oh”, we still gain some insight into the way the teams collaborate. future work could expand more on the inferred instructions and nonmatches and their implications of collaboration, rather than only focusing on (mis)matched actions. 7. conclusion in this article, we are interested in how children collaborate as they solve a problem together, in which what they say and do is strongly tied to how they perform, and subsequently what they will ultimately learn from the situated activity. to investigate this relationship, we consider the corpus of data (dialogue transcripts from audio files and action logs) generated by teams of two children engaged in a collaborative learning activity, which aims at providing an intuitive understanding of graphs and spanning trees. collaborative learning activities are a particularly interesting type of collaborative task, due to their “multi-layered goals”, typically including immediate, performance 30 studying alignment in a collaborative learning activity via automatic methods table 7: excerpts that contain information management marker “oh” from team 20, who performed well but did not learn. annotations could have the pending instructions, and (mis)matched or nonmatched actions. for brevity, nodes can be inferred from the utterance column for (mis)matched or nonmatched actions, unless partial (mis)match. utt. verb utterance annotations . . . . . . . . . . . . . . . b 10 says i’m just gonna . . . a adds luzern-zermatt nonmatcha(doa(add)) a 11 says uh . . . a 12 says uh . . . b 13 says oh there. a 14 says oh two three. b 15 says oh that’s what you’ve been doing this all time. . . . . . . . . . . . . . . . b 56 says oh i think we have to connect all of them. instructa(add(gallen,?)) b adds luzern-interlaken mismatchb (instructa(add(gallen,?))) b adds luzern-zurich nonmatchb(dob(add)) a 57 says oh. b 58 says okay i did some a 59 okay for me. a adds luzern-davos nonmatcha(doa(add)) a 60 says oh no. b 61 says i think we are doing terrible. . . . . . . . . . . . . . . . a 450 says oh. b 451 says 3. b 452 what? a 453 says let me . . . a removes luzern-zermatt nonmatcha(doa(remove)) a 454 says there you go. a 455 what no. b 456 says you are erasing my mistake. b 457 how dare you. a 458 says i know. 459 wait, what? 460 can i get pencil again? 461 oh oh. 462 oh. 463 okay. b 464 says we messed up again didn’t we? . . . . . . . . . . . . . . . 31 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel table 8: excerpts that contain information management marker “oh” from team 17, who had high task success (i.e. performed and learnt well). utt. stands for the utterance number, which is not applicable for the edit actions. see the caption of table 7 for further details. utt. verb utterance annotations . . . . . . . . . . . . . . . i 4 says so you only build from something that is already connected. b 5 says oh. a 6 says oh okay. b adds zermatt-davos nonmatchb(dob(add)) b adds gallen-davos nonmatchb(dob(add)) a adds zurich-davos nonmatcha(doa(add)) . . . . . . . . . . . . . . . b adds basel-bern matchb (instructa(add(basel,?))) a 61 says yeah, and then go to mount zurich. instructa(add(zurich,?)) b adds basel-zurich matchb (instructa(add(zurich,?))) a 62 says yeah. 63 oh no that costs more. 64 uh . . . b 65 says we should erase it. . . . . . . . . . . . . . . . a 266 says then do mount bern to mount zermatt. instructa(add(bern,zermatt)) a 267 says maybe that’s better. b 268 says you can’t do that. a 269 says oh. a 270 says then do . . . b 271 says mount bern to mount interlaken? instructb(add(bern,interlaken)) a 272 says yeah. a 273 says i think that’s 4 though. b adds interlaken-bern mismatchb (instructa(add(bern,zermatt))) a 274 says so don’t do that. b 275 says is that 4? a 276 says oh yeah it’s, it is 4. . . . . . . . . . . . . . . . 32 studying alignment in a collaborative learning activity via automatic methods goals (e.g. finding the solution to a math problem) and deeper, learning goals (e.g. understanding the notion of equation). collaboration often involves a dialogue amongst interlocutors. it has been shown that a dialogue is successful when there is alignment between the interlocutors, at different linguistic levels. we focus on two levels of alignment: i) lexical alignment (what was said), i.e. alignment at a lexical level, and ii) behavioural alignment (what was done), i.e. a new alignment context we propose to mean when instructions provided by one interlocutor are either followed or not followed with physical actions by the other interlocutor. we propose novel rule-based algorithms to automatically and empirically measure these two alignment contexts. what distinguishes our approach from other works, is the treatment of alignment as a procedure that occurs in stages; compared to a holistic approach that has been used in other works. thus our measures on alignment study more in depth how alignment forms (e.g. via priming and establishment), and then builds into a function of the discourse. to study alignment, we focus on the the formation of expressions related to the activity, as focusing on these expressions allows us to target information very specific to making progress in the task. additionally, our research on these two alignment contexts considers the occurrence of spontaneous speech phenomena (e.g. “um”, “uh” . . . ), as our dataset is one of spoken dialogues, where these paralinguistic cues may be an important indicator in alignment; and are often cues that are neglected as noise. we then observe how these local level alignment contexts build into a function of dialogue level phenomena; i.e. success in the task. we make this dataset and tools to study alignment in children’s dialogues publicly available (entitled the “justhink alignment dataset”, containing transcriptions, action logs and alignment tools). a first finding of this work is the discovery that the measures we propose are capable of capturing alignment in such a context. in terms of results, for rq1, we see that all teams establish routines, regardless of task success. an assumption that might be commonly made, is that the more aligned interlocutors are, the more are chances of their task success (usually, measured by performance alone). indeed, our results from rq1 indicate that this is not necessarily the case for lexical alignment: we rather observed that better performing teams were earlier than badly performing teams to align by our measures. in terms of the formation of a routine, we see that hesitation phenomena tend to occur around establishment times, and greatly after priming times, indicating that the if could be using fillers in the role of clarification. for rq2, we observe that all teams, regardless of task success, were behaviourally aligned. similar to rq1, we observed a general trend that better performing teams tend to follow up their instructions with actions earlier in the task rather than later. an interesting finding was that well-performing teams verbalise the marker “oh” more when they are behaviourally aligned, compared to other times in the dialogue. this shows that the information management marker “oh” is an important cue in alignment. to the best of our knowledge, we are the first to study the role of “oh” as an information management marker in a behavioural context (i.e. in connection to actions taken in a physical environment), compared to only a verbal one. while our measures discussed in rq1 and rq2 do not show significant results for learning, we still think considering learning with performance is an important first step in evaluating task success in these activities, as performance does not necessarily bring about learning (nasir, norman, et al., 2020b). our measures still reflect some fine-grained aspects of learning in the dialogue (such as an exploratory and collaborative period the interlocutors go through during the alignment process), even if we cannot conclude that overall they are linked to the final measure of learning. our measures capture more aspects of performance than learning, as our measures focus on what was said specifically about the environment, and the immediate apparent changes to the environment 33 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel – essentially the crux of the task. at a higher level, lack of understanding/learning etc. could be reflected by other phenomena (such as other multimodal features (nasir et al., 2020a, 2022)), other expressions (“i don’t understand”) etc. since there are many ways to learn, and hence different behaviours that could result in learning, it is unsurprising that such patterns are difficult to capture. working with dialogues is complicated, due to i) the nature of spontaneous speech (variable turn taking, disfluencies, etc.) and ii) the lack of automatic evaluation criteria. for example, human annotators might instinctively be able to say that a particular team collaborated well after observing a dialogue, but it is hard to empirically pinpoint the exact reasons such a judgement was taken. our results, albeit limited to a small dataset, highlight that in a situated educational activity, focusing simply on expressions related to the task and certain surrounding spontaneous speech phenomena can still give good insights into the nuances of collaboration between the interlocutors, and its ultimate links to task success (with the awareness that there are several other aspects that remain to be observed in the dialogue such as studying the gaze patterns, prosodic features and so on). it would be very interesting in future work to look at the influence that levels of alignment have on each other, e.g. how the lexical can influence the behavioural and so on. we hope that our findings can inspire further research on the topic and contribute to the design of technologies for the support of learning. 8. acknowledgments this project has received funding from the european union’s horizon 2020 research and innovation programme under grant agreement no 765955 (animatas project). 9. data availability all relevant data (transcripts, logs, and responses to the pre-test and the post-test, as well as the description of the network in the activity) are available from the zenodo repository, doi: 10.5281/zenodo.4627104. the code that reproduces all the results and figures given in this paper are available from the zenodo repository, doi: 10.5281/zenodo.4675070. 10. declaration of interest there is no conflict of interest. references karin aijmer. oh and ah in english conversation. in willem meijs, editor, corpus linguistics and beyond, pages 61–86. brill, 1987. doi: 10.1163/9789004483989 010. anne h. anderson, miles bader, ellen gurman bard, elizabeth boyle, gwyneth doherty, simon garrod, stephen isard, jacqueline kowtko, jan mcallister, jim miller, catherine sotillo, henry s. thompson, and regina weinert. the hcrc map task corpus. language and speech, 34(4):351–366, 1991. doi: 10.1177/002383099103400404. jennifer e arnold, michael k tanenhaus, rebecca j altmann, and maria fagnano. the old and thee, uh, new: disfluency and reference resolution. psychological science, 15(9):578–582, 34 https://www.animatas.eu/ http://doi.org/10.5281/zenodo.4627104 http://doi.org/10.5281/zenodo.4627104 http://doi.org/10.5281/zenodo.4675070 http://doi.org/10.1163/9789004483989_010 http://doi.org/10.1177/002383099103400404 studying alignment in a collaborative learning activity via automatic methods 2004. doi: 10.1111/j.0956-7976.2004.00723.x. jennifer e arnold, carla l hudson kam, and michael k tanenhaus. if you say thee uh you are describing something hard: the on-line attribution of disfluency during reference comprehension. journal of experimental psychology: learning, memory, and cognition, 33(5): 914–930, 2007. doi: 10.1037/0278-7393.33.5.914. michael j. baker, baruch b. schwarz, and sten r. ludvigsen. educational dialogues and computer supported collaborative learning: critical analysis and research perspectives. international journal of computer-supported collaborative learning, 16(4):583–604, 2021. doi: 10 .1007/ s11412-021-09359-1. reshmashree bangalore kantharaju, caroline langlet, mukesh barange, chloé clavel, and catherine pelachaud. multimodal analysis of cohesion in multi-party interactions. in proceedings of the 12th language resources and evaluation conference, pages 498–507. european language resources association, 2020. tim bell, ian h witten, and mike fellows. computer science unplugged: an enrichment and extension programme for primary-aged children. university of canterbury. computer science and software engineering, 2015. url http://hdl.handle.net/10092/247. marcela borge and carolyn p. rosé. quantitative approaches to language in cscl. in ulrike cress, carolyn p. rosé, alyssa friend wise, and jun oshima, editors, international handbook of computer-supported collaborative learning, computer-supported collaborative learning series, pages 585–604. springer international publishing, 2021. doi: 10 . 1007 / 978-3-030-65291-3 32. kristy boyer, eun young ha, robert phillips, michael wallis, mladen vouk, and james lester. dialogue act modeling in a complex task-oriented domain. in proceedings of sigdial 2010: the 11th annual meeting of the special interest group on discourse and dialogue, pages 297– 305. association for computational linguistics, september 2010. susan e brennan and herbert h clark. conceptual pacts and lexical choice in conversation. journal of experimental psychology: learning, memory, and cognition, 22(6):1482–1493, 1996. doi: 10.1037/0278-7393.22.6.1482. susan e brennan and maurice williams. the feeling of another’s knowing: prosody and filled pauses as cues to listeners about the metacognitive states of speakers. journal of memory and language, 34(3):383–398, 1995. doi: 10.1006/jmla.1995.1017. fabrizio butera, nicolas sommet, and céline darnon. sociocognitive conflict regulation: how to make sense of diverging ideas. current directions in psychological science, 28(2):145–151, 2019. doi: 10.1177/0963721418813986. herbert h. clark and susan e. brennan. grounding in communication. in l. b. resnick, j. m. levine, and s. d. teasley, editors, perspectives on socially shared cognition, pages 127–149. american psychological association, 1991. doi: 10.1037/10096-006. herbert h. clark and edward f. schaefer. contributing to discourse. cognitive science, 13(2): 259–294, 1989. doi: 10.1016/0364-0213(89)90008-6. 35 http://doi.org/10.1111/j.0956-7976.2004.00723.x http://doi.org/10.1037/0278-7393.33.5.914 http://doi.org/10.1007/s11412-021-09359-1 http://doi.org/10.1007/s11412-021-09359-1 http://hdl.handle.net/10092/247 http://doi.org/10.1007/978-3-030-65291-3_32 http://doi.org/10.1007/978-3-030-65291-3_32 http://doi.org/10.1037/0278-7393.22.6.1482 http://doi.org/10.1006/jmla.1995.1017 http://doi.org/10.1177/0963721418813986 http://doi.org/10.1037/10096-006 http://doi.org/10.1016/0364-0213(89)90008-6 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel herbert h. clark and deanna wilkes-gibbs. referring as a collaborative process. cognition, 22 (1):1–39, 1986. doi: 10.1016/0010-0277(86)90010-7. pierre dillenbourg. what do you mean by collaborative learning? in pierre dillenbourg, editor, collaborative-learning: cognitive and computational approaches, pages 1–19. oxford: elsevier, 1999. pierre dillenbourg and david traum. sharing solutions: persistence and grounding in multimodal collaborative problem solving. journal of the learning sciences, 15(1):121–151, 2006. doi: 10.1207/s15327809jls1501 9. willem doise and gabriel mugny. the social development of the intellect, volume 10 of international series in experimental social psychology. pergamon press, 1984. original work le développement social de l’intelligence published in 1981. guillaume dubuisson duplessis, chloé clavel, and frédéric landragin. automatic measures to characterise verbal alignment in human-agent interaction. in proceedings of the 18th annual sigdial meeting on discourse and dialogue, pages 71–81. association for computational linguistics, 2017. doi: 10.18653/v1/w17-5510. guillaume dubuisson duplessis, caroline langlet, chloé clavel, and frédéric landragin. towards alignment strategies in human-agent interactions based on measures of lexical repetitions. language resources and evaluation, 55(2):353–388, june 2021. doi: 10 . 1007 / s10579-021-09532-w. pinar dönmez, carolyn rosé, karsten stegmann, armin weinberger, and frank fischer. supporting cscl with automatic corpus analysis technology. in cscl ’05: proceedings of th 2005 conference on computer support for collaborative learning: learning 2005: the next 10 years!, pages 125–134, 2005. doi: 10.3115/1149293.1149310. sidney d’mello and art graesser. design of dialog-based intelligent tutoring systems to simulate human-to-human tutoring. in amy neustein and judith a. markowitz, editors, where humans meet machines: innovative solutions for knotty natural-language problems, pages 233–269. springer, new york, ny, 2013. doi: 10.1007/978-1-4614-6934-6 11. james d evans. straightforward statistics for the behavioral sciences. brooks/cole publishing company, 1996. aysu ezen-can and kristy boyer. combining task and dialogue streams in unsupervised dialogue act models. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 113–122. association for computational linguistics, 2014. doi: 10.3115/v1/w14-4316. aysu ezen-can and kristy elizabeth boyer. a tutorial dialogue system for real-time evaluation of unsupervised dialogue act classifiers: exploring system outcomes. in cristina conati, neil heffernan, antonija mitrovic, and m. felisa verdejo, editors, artificial intelligence in education, lecture notes in computer science, pages 105–114. springer international publishing, 2015a. doi: 10.1007/978-3-319-19773-9 11. 36 http://doi.org/10.1016/0010-0277(86)90010-7 http://doi.org/10.1207/s15327809jls1501_9 http://doi.org/10.18653/v1/w17-5510 http://doi.org/10.1007/s10579-021-09532-w http://doi.org/10.1007/s10579-021-09532-w http://doi.org/10.3115/1149293.1149310 http://doi.org/10.1007/978-1-4614-6934-6_11 http://doi.org/10.3115/v1/w14-4316 http://doi.org/10.1007/978-3-319-19773-9_11 studying alignment in a collaborative learning activity via automatic methods aysu ezen-can and kristy elizabeth boyer. understanding student language: an unsupervised dialogue act classification approach. journal of educational data mining, 7(1):51–78, 2015b. doi: 10.5281/zenodo.3554708. jean e. fox tree and josef c. schrock. discourse markers in spontaneous speech: oh what a difference an oh makes. journal of memory and language, 40(2):280–295, january 1999. doi: 10.1006/jmla.1998.2613. heather friedberg, diane litman, and susannah b. f. paletz. lexical entrainment and success in student engineering groups. in 2012 ieee spoken language technology workshop (slt), pages 404–409, 2012. doi: 10.1109/slt.2012.6424258. riccardo fusaroli and kristian tylén. investigating conversational dynamics: interactive alignment, interpersonal synergy, and collective task performance. cognitive science, 40(1):145– 171, 2016. doi: 10.1111/cogs.12251. simon garrod and anthony anderson. saying what you mean in dialogue: a study in conceptual and semantic co-ordination. cognition, 27:181–218, 1987. doi: 10 . 1016 / 0010-0277(87 ) 90018-7. bradley a. goodman, frank n. linton, robert d. gaimari, janet m. hitzeman, helen j. ross, and guido zarrella. using dialogue features to predict trouble during collaborative learning. user modeling and user-adapted interaction, 15(1):85–134, march 2005. doi: 10.1007/ s11257-004-5269-x. iris k. howley and carolyn p. rosé. towards careful practices for automated linguistic analysis of group learning. journal of learning analytics, 3(3):239–262, 2016. doi: 10.18608/jla.2016. 33.12. patrick jermann and marc-antoine nüssli. effects of sharing text selections on gaze crossrecurrence and interaction quality in a pair programming task. in proceedings of the acm 2012 conference on computer supported cooperative work, cscw ’12, pages 1125–1134. acm, 2012. doi: 10.1145/2145204.2145371. patrick jermann, dejana mullins, marc-antoine nüssli, and pierre dillenbourg. collaborative gaze footprints: correlates of interaction quality. connecting computer-supported collaborative learning to policy and practice: cscl2011 conference proceedings., volume i long papers: 184–191, 2011. manu kapur. productive failure. cognition and instruction, 26(3):379–424, 2008. doi: 10.1080/ 07370000802212669. manu kapur and katerine bielaczyc. designing for productive failure. journal of the learning sciences, 21(1):45–83, 2012. doi: 10.1080/10508406.2011.591717. hanae koiso, yasuo horiuchi, syun tutiya, akira ichikawa, and yasuharu den. an analysis of turn-taking and backchannels based on prosodic and syntactic features in japanese map task dialogs. language and speech, 41(3-4):295–321, 1998. doi: 10.1177/002383099804100404. 37 http://doi.org/10.5281/zenodo.3554708 http://doi.org/10.1006/jmla.1998.2613 http://doi.org/10.1109/slt.2012.6424258 http://doi.org/10.1111/cogs.12251 http://doi.org/10.1016/0010-0277(87)90018-7 http://doi.org/10.1016/0010-0277(87)90018-7 http://doi.org/10.1007/s11257-004-5269-x http://doi.org/10.1007/s11257-004-5269-x http://doi.org/10.18608/jla.2016.33.12 http://doi.org/10.18608/jla.2016.33.12 http://doi.org/10.1145/2145204.2145371 http://doi.org/10.1080/07370000802212669 http://doi.org/10.1080/07370000802212669 http://doi.org/10.1080/10508406.2011.591717 http://doi.org/10.1177/002383099804100404 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel deanna kuhn. thinking together and alone. educational researcher, 44:46–53, january 2015. doi: 10.3102/0013189x15569530. esther le grezause. um and uh, and the expression of stance in conversational speech. phd thesis, university of washington, august 2017. url http://hdl.handle.net/1773/40619. sarah ita levitan, jessica xiang, and julia hirschberg. acoustic-prosodic and lexical entrainment in deceptive dialogue. in proc. speech prosody 2018, pages 532–536, 2018. doi: 10.21437/ speechprosody.2018-108. diane j. litman, carolyn p. rosé, kate forbes-riley, kurt vanlehn, dumisizwe bhembe, and scott silliman. spoken versus typed human and computer dialogue tutoring. in james c. lester, rosa maria vicari, and fábio paraguaçu, editors, intelligent tutoring systems, volume 3220 of lecture notes in computer science, pages 368–379. springer, berlin, heidelberg, 2004. doi: 10.1007/978-3-540-30139-4 35. samuel louvan and bernardo magnini. recent neural methods on slot filling and intent classification for task-oriented dialogue systems: a survey. in proceedings of the 28th international conference on computational linguistics, pages 480–496. international committee on computational linguistics, 2020. doi: 10.18653/v1/2020.coling-main.42. nichola lubold and heather pon-barry. acoustic-prosodic entrainment and rapport in collaborative learning dialogues. in proceedings of the 2014 acm workshop on multimodal learning analytics workshop and grand challenge, mla ’14, pages 5–12. association for computing machinery, 2014. doi: 10.1145/2666633.2666635. nichola lubold, heather pon-barry, and erin walker. naturalness and rapport in a pitch adaptive learning companion. in 2015 ieee workshop on automatic speech recognition and understanding (asru), pages 103–110, 2015. doi: 10.1109/asru.2015.7404781. nichola lubold, erin walker, heather pon-barry, and amy ogan. automated pitch convergence improves learning in a social, teachable robot for middle school mathematics. in carolyn penstein rosé, roberto martı́nez-maldonado, h. ulrich hoppe, rose luckin, manolis mavrikis, kaska porayska-pomsta, bruce mclaren, and benedict du boulay, editors, artificial intelligence in education, lecture notes in computer science, pages 282–296. springer international publishing, 2018. doi: 10.1007/978-3-319-93843-1 21. divya menon, b. p. sowmya, margarida romero, and thierry viéville. going beyond digital literacy to develop computational thinking in k-12 education. in linda daniela, editor, epistemological approaches to digital learning in educational contexts. routledge, 1 edition, may 2020. doi: 10.4324/9780429319501-2. marie w meteer, ann a taylor, robert macintyre, and rukmini iyer. dysfluency annotation stylebook for the switchboard corpus. university of pennsylvania philadelphia, pa, 1995. jin mu, karsten stegmann, elijah mayfield, carolyn rosé, and frank fischer. the acodea framework: developing segmentation and classification schemes for fully automatic analysis of online discussions. international journal of computer-supported collaborative learning, 7: 285–305, june 2012. doi: 10.1007/s11412-012-9147-y. 38 http://doi.org/10.3102/0013189x15569530 http://hdl.handle.net/1773/40619 http://doi.org/10.21437/speechprosody.2018-108 http://doi.org/10.21437/speechprosody.2018-108 http://doi.org/10.1007/978-3-540-30139-4_35 http://doi.org/10.18653/v1/2020.coling-main.42 http://doi.org/10.1145/2666633.2666635 http://doi.org/10.1109/asru.2015.7404781 http://doi.org/10.1007/978-3-319-93843-1_21 http://doi.org/10.4324/9780429319501-2 http://doi.org/10.1007/s11412-012-9147-y studying alignment in a collaborative learning activity via automatic methods gabriel mugny and willem doise. socio-cognitive conflict and structure of individual and collective performances. european journal of social psychology, 8(2):181–192, 1978. doi: 10.1002/ejsp.2420080204. jauwairia nasir, barbara bruno, and pierre dillenbourg. is there ‘one way’ of learning? a data-driven approach. in 22nd acm international conference on multimodal interaction, pages 388–391, october 2020a. doi: 10.1145/3395035.3425200. jauwairia nasir, utku norman, barbara bruno, and pierre dillenbourg. when positive perception of the robot has no effect on learning. in 2020 29th ieee international conference on robot and human interactive communication (ro-man), pages 313–320. ieee, 2020b. doi: 10.1109/ ro-man47096.2020.9223343. jauwairia nasir, utku norman, barbara bruno, and pierre dillenbourg. you tell, i do, and we swap until we connect all the gold mines! ercim news, 2020 (120), 2020c. url https : / / ercim-news . ercim . eu / en120 / special / you-tell-i-do-and-we-swap-until-we-connect-all-the-gold-mines. jauwairia nasir, barbara bruno, mohamed chetouani, and pierre dillenbourg. what if social robots look for productive engagement? international journal of social robotics, 14:55–71, 2022. doi: 10.1007/s12369-021-00766-w. ani nenkova, agustı́n gravano, and julia hirschberg. high frequency word entrainment in spoken dialogue. in proceedings of acl-08: hlt, short papers, pages 169–172. association for computational linguistics, 2008. martin j. pickering and simon garrod. toward a mechanistic psychology of dialogue. behavioral and brain sciences, 27(02):169–225, 2004. doi: 10.1017/s0140525x04000056. martin j. pickering and simon garrod. alignment as the basis for successful communication. research on language and computation, 4(2):203–228, 2006. doi: 10.1007/s11168-006-9004-0. joseph p pickett. the american heritage dictionary of the english language. houghton mifflin harcourt, 2018. zahra rahimi, anish kumar, diane litman, susannah paletz, and mingzhi yu. entrainment in multi-party spoken dialogues at multiple linguistic levels. in proc. interspeech 2017, pages 1696–1700, 2017. doi: 10.21437/interspeech.2017-1568. david reitter and johanna d. moore. alignment and task success in spoken dialogue. journal of memory and language, 76:29–46, 2014. doi: 10.1016/j.jml.2014.05.008. jeanine romano, jeffrey d kromrey, jesse coraggio, and jeff skowronek. appropriate statistics for ordinal level data: should we really be using t-test and cohen’s d for evaluating group differences on the nsse and other surveys? in annual meeting of the florida association of institutional research, volume 177, 2006. jeremy roschelle and stephanie d. teasley. the construction of shared knowledge in collaborative problem solving. computer supported collaborative learning, 128:69–97, 1995. doi: 10.1007/978-3-642-85098-1 5. 39 http://doi.org/10.1002/ejsp.2420080204 http://doi.org/10.1145/3395035.3425200 http://doi.org/10.1109/ro-man47096.2020.9223343 http://doi.org/10.1109/ro-man47096.2020.9223343 https://ercim-news.ercim.eu/en120/special/you-tell-i-do-and-we-swap-until-we-connect-all-the-gold-mines https://ercim-news.ercim.eu/en120/special/you-tell-i-do-and-we-swap-until-we-connect-all-the-gold-mines http://doi.org/10.1007/s12369-021-00766-w http://doi.org/10.1017/s0140525x04000056 http://doi.org/10.1007/s11168-006-9004-0 http://doi.org/10.21437/interspeech.2017-1568 http://doi.org/10.1016/j.jml.2014.05.008 http://doi.org/10.1007/978-3-642-85098-1_5 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel carolyn rosé, yi-chia wang, yue cui, jaime arguello, karsten stegmann, armin weinberger, and frank fischer. analyzing collaborative learning processes automatically: exploiting the advances of computational linguistics in computer-supported collaborative learning. international journal of computer-supported collaborative learning, 3(3):237–271, 2008. doi: 10.1007/s11412-007-9034-0. mirweis sangin, gaëlle molinari, marc-antoine nüssli, and pierre dillenbourg. facilitating peer knowledge modeling: effects of a knowledge awareness tool on collaborative learning outcomes and processes. computers in human behavior, 27(3):1059–1067, 2011. doi: 10.1016/j.chb.2010.05.032. oliver scheuer, frank loll, niels pinkwart, and bruce m. mclaren. computer-supported argumentation: a review of the state of the art. international journal of computer-supported collaborative learning, 5(1):43–102, march 2010. doi: 10.1007/s11412-009-9080-x. deborah schiffrin. discourse markers. number 5 in studies in interactional sociolinguistics. cambridge university press, 1987. doi: 10.1017/cbo9780511611841. bertrand schneider, marcelo worsley, and roberto martinez-maldonado. gesture and gaze: multimodal data in dyadic interactions. in ulrike cress, carolyn rosé, alyssa friend wise, and jun oshima, editors, international handbook of computer-supported collaborative learning, computer-supported collaborative learning series, pages 625–641. springer international publishing, 2021. doi: 10.1007/978-3-030-65291-3 34. daniel l schwartz. the emergence of abstract representations in dyad problem solving. journal of the learning sciences, 4(3):321–354, 1995. doi: 10.1207/s15327809jls0403 3. arabella j. sinclair and bertrand schneider. linguistic and gestural coordination: do learners converge in collaborative dialogue? in proceedings of the 14th international conference on educational data mining (edm 2021), pages 431–437. international educational data mining society, 2021. tanmay sinha and justine cassell. we click, we align, we learn: impact of influence and convergence processes on student learning and rapport building. in proceedings of the 1st workshop on modeling interpersonal synchrony and influence, interpersonal ’15, pages 13–20. association for computing machinery, 2015. doi: 10.1145/2823513.2823516. vicki l smith and herbert h clark. on the course of answering questions. journal of memory and language, 32(1):25–38, 1993. doi: 10.1006/jmla.1993.1002. jan-willem strijbos, rob l. martens, frans j. prins, and wim m. g. jochems. content analysis: what are they talking about? computers & education, 46(1):29–48, 2006. doi: 10.1016/j. compedu.2005.04.002. pierre tchounikine, nikol rummel, and bruce m. mclaren. computer supported collaborative learning and intelligent tutoring systems. in roger nkambou, jacqueline bourdeau, and riichiro mizoguchi, editors, advances in intelligent tutoring systems, studies in computational intelligence, pages 447–463. springer, berlin, heidelberg, 2010. doi: 10 . 1007 / 978-3-642-14363-2 22. 40 http://doi.org/10.1007/s11412-007-9034-0 http://doi.org/10.1016/j.chb.2010.05.032 http://doi.org/10.1007/s11412-009-9080-x http://doi.org/10.1017/cbo9780511611841 http://doi.org/10.1007/978-3-030-65291-3_34 http://doi.org/10.1207/s15327809jls0403_3 http://doi.org/10.1145/2823513.2823516 http://doi.org/10.1006/jmla.1993.1002 http://doi.org/10.1016/j.compedu.2005.04.002 http://doi.org/10.1016/j.compedu.2005.04.002 http://doi.org/10.1007/978-3-642-14363-2_22 http://doi.org/10.1007/978-3-642-14363-2_22 studying alignment in a collaborative learning activity via automatic methods jesse thomason, huy v. nguyen, and diane litman. prosodic entrainment and tutoring dialogue success. in h. chad lane, kalina yacef, jack mostow, and philip pavlik, editors, artificial intelligence in education, lecture notes in computer science, pages 750–753. springer, 2013. doi: 10.1007/978-3-642-39112-5 104. stefan trausan-matu and james d. slotta. artifact analysis. in ulrike cress, carolyn p. rosé, alyssa friend wise, and jun oshima, editors, international handbook of computer-supported collaborative learning, computer-supported collaborative learning series, pages 551–567. springer, 2021. doi: 10.1007/978-3-030-65291-3 30. gokhan tur and renato de mori. spoken language understanding: systems for extracting semantic information from speech. john wiley & sons, 2011. erin walker, nikol rummel, and kenneth r. koedinger. designing automated adaptive support to improve student helping behaviors in a peer tutoring activity. international journal of computer-supported collaborative learning, 6(2):279–306, 2011a. doi: 10 . 1007 / s11412-011-9111-2. erin walker, nikol rummel, and kenneth r. koedinger. using automated dialog analysis to assess peer tutoring and trigger effective support. in gautam biswas, susan bull, judy kay, and antonija mitrovic, editors, artificial intelligence in education, lecture notes in computer science, pages 385–393. springer, 2011b. doi: 10.1007/978-3-642-21869-9 50. arthur ward and diane litman. dialog convergence and learning. in proceedings of the 2007 conference on artificial intelligence in education: building technology rich learning contexts that work, pages 262–269. ios press, june 2007a. arthur ward and diane j. litman. automatically measuring lexical and acoustic/prosodic convergence in tutorial dialog corpora. in slate workshop on speech and language technology in education, 2007b. url http://d-scholarship.pitt.edu/id/eprint/23210. armin weinberger and frank fischer. a framework to analyze argumentative knowledge construction in computer-supported collaborative learning. computers & education, 46(1):71–95, 2006. doi: 10.1016/j.compedu.2005.04.003. vicky zayats, trang tran, richard wright, courtney mansfield, and mari ostendorf. disfluencies and human speech transcription errors. in proc. interspeech 2019, april 2019. doi: 10.21437/ interspeech.2019-3134. appendix a. automatic speech recognition (asr) comparison an accurate transcription of the task-specific referents is crucial for both our lexical and behavioural alignment measures. to evaluate how well automatic speech recognition would perform in obtaining transcripts for our analyses, we used the google cloud speech-to-text services. to configure the recognition system, we extracted a 15 seconds long audio sample that contains task-specific referents and a filler “um”, see the reference transcript at table 9. we need to assist the system towards improving its accuracy for the task-specific referents, as these are context specific 41 http://doi.org/10.1007/978-3-642-39112-5_104 http://doi.org/10.1007/978-3-030-65291-3_30 http://doi.org/10.1007/s11412-011-9111-2 http://doi.org/10.1007/s11412-011-9111-2 http://doi.org/10.1007/978-3-642-21869-9_50 http://d-scholarship.pitt.edu/id/eprint/23210 http://doi.org/10.1016/j.compedu.2005.04.003 http://doi.org/10.21437/interspeech.2019-3134 http://doi.org/10.21437/interspeech.2019-3134 https://cloud.google.com/speech-to-text utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel type start (s) end (s) int. utterance reference 552.5 556.7 a what about mount davos to mount , saint gallen ? 558.6 562.6 b because what if , you did if we could do it ? 562.5 566.1 a what about mount um davos to mount gallen ? default 552.2 556.1 a what about mount davis to mount saint helen ? 558.5 562.3 b cuz what if you did , if we could do it . 558.6 560.9 a what if you did ? 561.7 565.9 a could you it ? what about mount davis to mount gallon ? 566.0 566.7 b mount . adapted 552.6 556.5 a what about mount davos to mount st . gallen ? 558.7 562.8 b cuz what if you did , if we could do it . 562.5 565.8 a what about mount davos to mount gallen ? table 9: automatic transcription results for an audio sample. words that do not occur frequently and thus would be difficult to recognise. to do so, we supplied the task-specific referents as ‘hints’ to the asr system. then, we used the extra boost feature of the system to increase the probability that these specific phrases will be recognised as described in the system’s documentation (boost = 200). furthermore, we set the system to use the enhanced “video” model that is particularly suitable “for audio that was recorded with a high-quality microphone or that has lots of background noise”. this is the case for our data that contains background music of the game and some audio spill where one interlocutor’s microphone could pick up sound from the other interlocutor. for the audio sample, table 9 presents our manually obtained reference transcript, the automatic transcript with the default model, and the transcript with a model adapted to our dataset. the transcription with the adapted model seems promising, as the task-specific referents are correctly transcribed. with the same adapted configuration, we automatically transcribed the complete audio files, for the subset of data. we evaluate using traditional word error rate (wer), but also evaluate the error rate for task-specific referents only (as only task specific referents are required for the methodology for h1.1 and h2.1). fig. 10 presents the error rates for the automatic transcripts. we see that the word error rates are very high (median = 62.6%), and varied between interlocutors of the same team (ranging from = 37.7% to = 253.4%). when we filter for the task-specific referents, we see that the referent error rates are high as well (median = 47.3%). thus it is infeasible to use automatic transcripts for this work, and therefore use gold-standard transcripts. we note that the filler “uh” is not recognised. appendix b. accuracy of algorithms for studying lexical alignment (rq1) the algorithm extracts routine expressions by the exact matching of token sequences: thus, the accuracy of the inference depends only on having accurate transcripts. we have gold-standard transcriptions with standardised variations of pronunciation (the details of the transcription are given sec. 3.3). thus the extraction of routine expressions, and subsequently determining the priming and establishment times are not sources of error. 42 https://cloud.google.com/python/docs/reference/speech/latest/google.cloud.speech_v1p1beta1.types.speechcontext?hl=en https://cloud.google.com/speech-to-text/docs/basics#select-model https://cloud.google.com/speech-to-text/docs/basics#select-model studying alignment in a collaborative learning activity via automatic methods 7 8 9 10 11 17 18 20 28 47 team no 0% 50% 100% 150% 200% 250% w or d er ro r r at e (% ) word error rates interlocutor a b (a) word error rates 7 8 9 10 11 17 18 20 28 47 team no 0% 50% 100% 150% 200% 250% er ro r r at e (% ) task-specific referent error rates interlocutor a b (b) task-specific referent error rates figure 10: (left) word and (right) task-specific referent error rates per interlocutor for the transcripts obtained by automatic speech recognition. however, this exact matching of token sequences is exhaustive. the original work by dubuisson duplessis et al. (2021) is intended to measure alignment by enumerating all existing matches (for example, even if interlocutors primed and established the token “what”, this would be counted towards a routine expression formed). this is why we filter for referents that are specific to the task. therefore, we would like to highlight that the alignment is lexical, based on a surface level matching of token sequences. furthermore, we keep in mind the issues of transcribing disfluent speech (le grezause, 2017; zayats et al., 2019), and thus use transcription software (praat) to ensure no unnecessary insertion, substitution or deletion of disfluencies. for studying behavioural alignment (rq2) there are two possible sources of error: while i) inferring instructions, and then while ii) matching them with actions. to infer instructions, we allow for a build up of instructions over a period of time (caching instructions), and then clear these instructions. this ensures that an instruction given at the start of the interaction is not matched with an action at the end of the interaction, i.e. there is a temporal constraint of when an instruction is a valid instruction. the main source of error in inferring instructions could be in anaphora resolution, e.g. if an interlocutor says “connect that node there”. this is why we allow partially inferred instructions (i.e. only one of the node names is mentioned) and (partial) matching of these instructions with actions. for example, from the utterance “maybe we start from mount zermatt.”, with only one node being explicitly stated, we infer the partial instruction add(zermatt,?). then, we consider the follow-up action as matched, if it is an add action, with ‘zermatt’ as one of the nodes of the added connection. thus, the inferring of instructions could be considered greedy, and suffer from inferences at each iteration without considering broader context. with regards to matching instructions to actions, problems arise when considering the formulaic definition of behavioural alignment. the algorithm is explicit in considering which interlocutor is in which view (i.e. always characterising interlocutors into if and ig). matching instructions-toactions are defined in a computationally strict way, and may not consider matched instructions resulting from negotiations, or consequently the build up of matches over larger periods of time (as the instructions are cleared frequently). furthermore, given this formulaic definition, this may be 43 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel a difficult task to manually annotate by human annotator, i.e. constantly considering who is in the position to give instructions versus follow (which changes throughout the interaction). the annotation of high-level constructs, such as engagement, emotion and in our work, behavioural alignment, all have a perceptual component that is intrinsically hard to define, and thus guide for annotation purposes (e.g. see nasir et al. 2022). thus given the way behavioural alignment is defined, our measures may not capture the whole picture, of what an annotator might perceptually label the interaction to be. appendix c. the activity details the minimum-spanning-tree problem is defined on a graph g = (v,e) with a cost function on its edges e, where v is the set of nodes, e ⊆ v × v is the set of edges that connects pairs of nodes. the goal is to find a subset of edges t ⊆ e such that the total cost of the edges in t is minimised. this corresponds to finding a minimum-cost sub-network of the given network. in the network used in the justhink activity, there exist 10 nodes: { ‘luzern’, ‘interlaken’, ‘montreux’, ‘davos’, ‘zermatt’, ‘neuchatel’, ‘gallen’, ‘bern’, ‘zurich’, ‘basel’ } and 20 edges. a description of the network, with the node labels (e.g. “mount luzern”), x, y position of a node, possible edges between the nodes, and edge costs, is publicly available online with the dataset (justhink alignment dataset), from the zenodo repository doi: 10.5281/zenodo.4627104. appendix d. pre-test and post-test details the pre-test and the post-test are made up of 10 multiple-choice items with a single correct answer. the items of the pre-test and the post-test are defined in a context other than the context of the task (swiss gold mines), and are based on variants of the graphics in the muddy city problem from bell et al. (2015). all items were validated by experts of the domain and experts in education. the score of an item is 1 if it is answered correctly, and 0 otherwise. the pre-test and post-test assess the following concepts: 1. concept-1 (existence): if a spanning tree exists, i.e. the graph is connected. 2. concept-2 (spanningness): if the given subgraph spans the graph. 3. concept-3 (minimumness): if the given subgraph that spans the graph has a minimum cost. there are 3, 3, and 4 items associated with the concepts, respectively. appendix e. algorithms e.1 instruction recognition in rq2 recognise-instructions automatically infers a sequence of instructions for an input utterance via a rule-based algorithm, as described in algorithm 1. note that allows inference of partial instructions i.e. that contain one node name only. it uses recognise-entities in algorithm 2 to find the edit instructions in an utterance in a way. 44 http://doi.org/10.5281/zenodo.4627104 studying alignment in a collaborative learning activity via automatic methods algorithm 1: recognise-instructions finds the edit instructions in an utterance. it is implemented in a script (tools/6 recognise instructions detect follow-up actions.ipynb in the tools) that generates an annotated corpus (processed data/annotated corpus available with the tools), where the column instructions gives the list of instructions that are inferred for that row. input: a sequence of tokens u = ⟨t1, t2, . . . , tn⟩ that make up an utterance u output: a sequence of instructions i = ⟨i1, i2, . . . , ik⟩ 1 e ← recognise-entities(u) 2 i ← an empty sequence // for inferred instructions 3 i← a new instruction object ; i.verb← null ; i.u← null; i.v← null // u is the first and v is the second node by mention 4 foreach entity e ∈ e do 5 if e.label = ‘add’ or e.label = ‘remove’ then 6 if i.verb ̸= null then // already inferring an instruction 7 if i.u ̸= null then // save the partial instruction 8 insert i into i 9 i.u← null; i.v← null // clear node 1 and 2 10 i.verb← e.label 11 else if i.u = null then // that is, e.label = ‘node’ 12 i.u← e.token 13 else if i.v = e.token then 14 if i.u ̸= e.token then // if not repeating node name 15 i.v← e.token 16 if i.verb = null then // default to a verb if not detected 17 if i .length = 0 then // no previous instruction: default to ‘add’ 18 i.verb← ‘add’ 19 else // default to previous instruction’s verb if exists 20 i.verb← i[i.length− 1].verb 21 insert i into i 22 i← a new instruction object (i.verb← null, i.u← null, i.v← null) 23 return i 45 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel algorithm 2: recognise-entities finds the edit entities in an utterance via a simple rule-based named entity recognition procedure. it is implemented in a script (tools/6 recognise instructions detect followup actions.ipynb in the tools) that generates an annotated corpus (processed data/annotated corpus available with the tools, for which we employ named-entity recognition (ner) feature of the python library spacy that performs this entity recognition. input: a sequence of tokens u = ⟨t1, t2, . . . , tn⟩ that make up an utterance u output: a sequence of entities e = ⟨e1, e2, . . . , em⟩ 1 n ← ⟨‘montreux’, ‘bern’, . . . , ‘basel’⟩ // all node names 2 a← ⟨‘add’, ‘build’, ‘connect’, ‘do’, ‘go’, ‘put’⟩ // add verbs 3 r← ⟨‘away’, ‘cut’, ‘delete’, ‘erase’, ‘remove’, ‘rub’⟩ // remove verbs 4 e ← an empty sequence // for inferred entities 5 foreach token t ∈ u do 6 label← null 7 if t ∈ n then label← ‘node’ 8 else if t ∈ a then label← ‘add’ 9 else if t ∈ r then label← ‘remove’ 10 if label ̸= null then 11 e← a new entity object ; e.token← t ; e.label← label 12 insert e into e 13 return e e.2 detecting follow-up actions of the instructions in rq2 match-instructions-to-actions pairs instructions with actions as matches or mismatches, for a verbal and physical actions list a, as described in algorithm 3. preprocessing (to obtain a verbal and physical actions list a): we combine the transcript and edit actions in a subject-verb-object(-turn-attempt) format. • each utterance in the transcript is added as an action with the verb ‘says’. • each edit action from the logs is added with the verb ‘adds’ or ‘removes’, according to whether it is an add action or a remove action, respectively. each action a ∈ a has fields: • a.subject ∈ {‘a’, ‘b’}, the two learners that are collaborating to solve the given problem • a.verb ∈ {‘says’, ‘adds’, ‘removes’}, the edit actions and utterance action for matching instructions with edit actions • a.object ∈ {utterances} ∪ {(u, v) : (u, v) ∈ edges} • a.turn ∈ {1, 2, . . . , n} indicating the turn number of the period the action belongs to (where for utterances, the start time of the utterance belongs to). after every two edits, the turn number incremented by 1 • a.attempt ∈ {1, 2, . . . ,m} indicating the attempt number of the period the action belongs to (where for utterances, the start time of the utterance belongs to). after every submission, the attempt number is incremented by 1 46 https://spacy.io/ studying alignment in a collaborative learning activity via automatic methods algorithm 3: match-instructions-to-actions pairs a list of pending instructions with actions as matches or mismatches. it is implemented in a script (tools/6 recognise instructions detect followup actions.ipynb in the tools) that generates an annotated corpus (processed data/annotated corpus available with the tools, where the column matching gives the the result of the matching for that row). input: a sequence of verbal and physical actions a = ⟨a1, a2, . . . , ak⟩ output: a sequence of m = ⟨m1,m2, . . . ,mk⟩ holding (mis)match info mi for each ai 1 p ← an empty sequence for pending instructions to be matched 2 m ← an empty sequence for match/mismatch for each action in a 3 attempt← 1 // submission no for clearing the pending instructions list 4 turn← 1 // turn no for clearing the pending instructions list 5 foreach action a ∈ a do // clear pending instructions if a new turn (or attempt i.e. submission) 6 if a.turn = turn+ 1 then 7 clear p // remove all items in the sequence p 8 turn← a.turn // update the current episode (i.e. new turn) 9 else if a.attempt = attempt+ 1 then 10 clear p // remove all items in the sequence p 11 attempt← a.attempt // update the current episode (i.e. new attempt) // for say action, recognise instructions and update the pending instructions list 12 if a.verb = ‘says′ then 13 i ← recognise-instructions(a.object) // a.object is the utterance 14 foreach instruction i ∈ i do 15 i.agent← a.subject // set the instructing agent 16 insert i into p // update the pending instructions // for do action, try to match with a pending instruction 17 else if a.verb = ‘does′ then 18 i ′ ← {i : i ∈ i and i.agent ̸= a.subject} // filter for the other interlocutor’s instructions 19 m← a new matching object ; m.match← null 20 if i ′.length > 0 then // there is an instruction that may (mis)match // try to match a pending instruction with the current action 21 foreach instruction i ∈ i ′ do 22 if check-match(i, a) then 23 m.match← true; m.instruction← i; m.action← a 24 if m..match = null then // no matches, hence a mismatch 25 i← i ′[i ′.length− 1] // get the last instruction by the other 26 m.match← false; m.instruction← i; m.action← a // process the match (if matched or mismatched) 27 if m..match ̸= null then // match: true or mismatch: false 28 m [i]← m // add match object to list of matches // remove matching instructions from pending instructions sequence 29 foreach instruction i ∈ p do 30 if check-match(i, a) then remove i from p 31 return m 47 utku norman∗, tanvi dinkar∗, barbara bruno, and chloé clavel algorithm 4: check-match checks if an instruction matches with the action. it allows partial matching for partially inferred instructions (i.e. only one of the node names is mentioned). input: an instruction i and an action a output: true if the intended action in i and action a match, false otherwise 1 if i.action ̸= a.verb then 2 return false 3 u← a.object.u // first node in the edited edge 4 v ← a.object.v // second node in the edited edge, sorted 5 if i.v ̸= null then // instruction is partially inferred i.e. contains i.u only 6 if i.u = u or i.u = v then // if one node matches 7 return true 8 else 9 return false 10 else if (i.u = u or i.u = v) and ( i.v = u or i.v = v) then // both match 11 return true // note that i.v ̸= i.u by its way of inference 12 else 13 return false 48 introduction related work analysis of verbal behaviour in educational dialogues analysing textual input complications of working with speech data alignment in educational dialogue analysis of alignment in other scenarios materials the justhink dataset alignment in the justhink dataset creating the justhink alignment dataset research questions studying lexical alignment (rq1) – how expressions related to the task contribute to lexical alignment experiment 1: studying when interlocutors are lexically aligned and its association with task success (h1.1) methodology results and discussion experiment 2: studying hesitation phenomena surrounding lexical alignment and its association with task success (h1.2) methodology results and discussion studying behavioural alignment (rq2) experiment 3: studying when interlocutors are behaviourally aligned and its association with task success (h2.1) methodology results and discussion experiment 4: studying information management phenomena surrounding behavioural alignment and its association with task success (h2.2) methodology results and discussion conclusion acknowledgments data availability declaration of interest references automatic speech recognition (asr) comparison accuracy of algorithms the activity details pre-test and post-test details algorithms instruction recognition in rq2 detecting follow-up actions of the instructions in rq2 dialogue & discourse 16(2) (2025) 111–124 doi: 10.5210/dad.2025.204 ©2025 holden härtl this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). calling things by their names: towards a unified account for name-informing and mixed quotation holden härtl holden.haertl@uni-kassel.de universität kassel editor: manfred stede submitted 04/2025; accepted 12/2025; published online 12/2025 abstract this paper explores the semantic connection between mixed quotation and name-informing quotation, proposing a unified account for both. while mixed quotation combines direct quotation with indirect reporting, name-informing quotation highlights the linguistic shape of a concept’s conventionalized name. we argue that both types of quotation involve a naming predicate – explicit in name-informing quotation and covert in mixed quotation. a pilot questionnaire study, which presented participants with two-turn dialogue contexts, used the notion of at-issueness to probe the naming component across these types, supporting the hypothesis that both share a similar semantic structure. this unified approach contributes to a broader understanding of quotational constructions and their role in linguistic and discourse representation. keywords: quotation, mixed quotation, pure quotation, name-informing 1 introduction quotation is a metalinguistic device used to talk about certain dimensions of language, see, e.g., cappelen & lepore (1997), davidson (1979), saka (1998). in quotational constructions, expressions are mentioned rather than or in addition to being used denotationally. with an assertion like in (1a), for example, in contrast to (1b), the syllabic setup of the word sofa is described and the quotation marks around sofa indicate this use, which means reference is made to a linguistic dimension of the quoted expression, see, e.g., quine (1981). (1) a. “sofa” has two syllables. b. a sofa is a piece of furniture. the type of quotation displayed in (1a) is an instance of so-called pure quotation. other types involve direct quotation (see 2a), mixed quotation (2b) and scare quotation (2c), see, e.g., cappelen and lepore (1997). (2) a. “something is wrong,” alan whispered softly to his dolls. b. ben declared he’s going to “kick up a huge fuss” today. c. flowers “know” when to bloom. each of these types marks a different metalinguistic function. direct quotation in (2a) reproduces someone’s original wording, typically introduced by a speech-report predicate. mixed quotation in (2b) blends direct quotation and indirect speech, embedding a segment of what someone else has said into the speaker’s own sentence structure. scare quotation in (2c), by contrast, signals that the quoted term is used in a non-standard, ironic, or distanced way, departing from its literal meaning. härtl 112 in what we call name-informing quotation, quotation marks serve a different purpose: they highlight the linguistic shape of a conventionalized term for a given concept. consider the following example. (3) the phenomenon is called “sun halo.” in this example, the quotes highlight the linguistic shape of the expression sun halo, thus referring to the conventionalized name that is used to describe a certain optical phenomenon, i.e., a sun halo. in this paper, we claim that name-informing quotation can be unified with mixed quotation, as illustrated in (2b), under a single semantic account. our main argument is that both types of quotation involve a naming predicate, which is covert in mixed quotation but explicit in name-informing quotation. a pilot questionnaire-based rating study, conducted in german, was designed to determine the naming component of each type by assessing its potential to be treated as at-issue in a comparison between instances of mixed quotation and name-informing quotation. the structure of this paper is as follows. in section 2, the main points of two different approaches towards quotation are surveyed. in section 3, the semantic properties of name-informing constructions are outlined, and we introduce our unified account for the two types of quotation in question. section 4 reports on the experimental study we conducted to test our claim. section 5 concludes our investigation, drawing connections to related analyses and addressing remaining open questions. 2 perspectives on quotation theoretical approaches to quotation differ in whether they assume that a quotational meaning is part of the sentence’s compositional semantic structure or arises through pragmatic inference. this distinction has implications for how different types of quotation are analyzed and whether a unified structural treatment is theoretically plausible. a key point in the debate concerns the optionality of quotation marks and whether they signal a semantic contribution or merely reflect a pragmatic interpretation. proponents of a semantic analysis of quotes often claim the presence of quotes to have truth-conditional effects, and that they are used to produce truth-conditionally relevant content (predelli 2003, simchen 1999). on a semantic account, the apparent optionality of quotes is explained in terms of compositional contextual embedding: when embedded under certain predicates such as speech-report (say) or naming predicates (call), the quoted material semantically contributes its mentioned form to the truth-conditional content. on this view, quotation marks are quasi-obligatory, surfacing either opaquely or at another linguistic level, e.g., the acoustic (cf. cappelen and lepore 1999). the optionality of quotes can also be used as evidence for a pragmatic approach to quotes. under a pragmatic view, the interpretation of a linguistic segment as a quotation does not depend on compositional embedding but on extralinguistic contextual cues in the broader discourse or situation, which are considered sufficient to yield a quotational reading even in the absence of a dedicated quotation marker. in this way, for example, washington (1992) argues that neither graphemic quotes nor their gestural and acoustic equivalents are an essential part of a quotational construction. on his account, quotes are no more than a punctuation device and, as such, “are neither mentioning expressions nor parts of mentioning expressions” (washington 1992: 591). approaches of this sort imply that quotes are not semantic in the sense that their manifestation is not part of the compositional semantic representation of a quotational construction. instead, quotes are considered pragmatic in nature. a pragmatic view is maintained in analyses like de brabanter’s (2013), who, with a focus on mixed quotation, argues that contextual cues alone are sufficient to pragmatically construe a quotational meaning with no need to signal the quotation with a dedicated linguistic marker. schlechtweg & härtl (2020) report on spoken data, which indicate that quotes are acoustically pronounced, primarily triggering a lengthening effect, but that this is independent of the compositional environment they occur in. in quotational contexts in which the target item is not embraced by quotes, an acoustic correlate was not found, which is explained by the fact that the mentioning use of the target expression is resolved pragmatically. semantic and pragmatic approaches differ in how they account for the linguistic representation of quotational expressions. while semantic theories typically posit structural elements – such as an overt or covert quotation operator – that contribute compositionally to the sentence meaning, pragmatic calling things by their names: towards a unified account for name-informing and mixed quotation 113 accounts emphasize contextual inference and interpret quotational meaning as arising from discourse context and speaker intent, even in the absence of explicit markers. approaches that integrate different types of quotation into a unified theoretical framework are scarce (e.g., de brabanter 2023). in the following, we will propose that mixed quotation (see 2c) and name-informing quotation (see 3) can be captured under a single semantic account. our central claim – partly following maier (2014b) – is that both types involve a naming predicate (or operator), which is overt in name-informing quotation and covert in mixed quotation. this account is semantic in that it assumes a compositional contribution of the naming predicate, but it also incorporates a pragmatic dimension, as the realization of quotation marks is treated as optional and context-dependent. 3 towards a unified account quotations as in (2b) and (3) above are typically considered to represent two distinct categories of quotation: mixed quotation (mq) and name-informing quotation (niq). mq involves a blend of direct quotation and indirect speech report, where the speaker quotes a portion of what someone else has said. in contrast, niq is a metalinguistic device used to demonstrate the linguistic shape of a concept’s conventionalized name in a rule-like fashion. 3.1 name-informing quotation naming predicates are highly polysemous, see, among others, anderson (2004) and biro (2012) for analyses. they can be used to express, for instance, an act of baptizing as in (4a), an act of nomination, see (4b), or an evaluative act, as in (4c), conveying a negative attitude of the subject referent toward the target individual. (4) a. they named their son arthur. b. she was named the president of the university. c. max called his neighbor a simpleton. in what follows, we explore a specific variant of constructions involving naming predicates, namely name-informing quotation, and lay out a semantic analysis. constructions involving a niq typically contain predicates like call, name, refer to etc., as embodied in (3) and in (5) below. (5) a. one calls this medical condition “cataract.” b. a function that invokes itself is named “recursive function.” c. the purity of gold is referred to with the word “karat.” name-informing constructions1 (nics) express established naming relations by highlighting the linguistic form of a concept’s conventionalized name. the predicate serves to mark the link between the concept and its linguistic label. this connection arises from an (abstract) act of nomenclature, which assigns a name to a concept, establishing the term for subsequent use. this contrasts with non-nameinforming uses of naming predicates as exemplified in (4c), which does not describe an established naming relation but rather an act of evaluative judgment. we have argued elsewhere that niq represents a special case of pure quotation (pq) (härtl 2018), i.e., a metalinguistic device used to mention a word or phrase as an object of discussion, emphasizing its linguistic form rather than its meaning (e.g., davidson 1979, cappelen & lepore 1997, maier 2014a). similar to pqs, the term quoted in a nic is not used denotationally but rather to cite a name, highlighting a key feature of pq: the expression is referenced as a linguistic entity rather than being used in its typical referential function. as an illustration of their metalinguistic status, pqs in nics can be preceded by appositions such as the word, as exemplified in (5c) above. 1 by “name-informing construction”, we refer to a sentence that contains a name-informing quotation. härtl 114 predicates like call are three-place predicates, which require an argument that can be interpreted metalinguistically.2 in a case like (5a) above, for instance, call is used to assert that a certain occurrence of an eye condition is commonly referred to with the word cataract. thus, call’s verbal root involves three thematic arguments, an agent x (one), a theme y (this medical condition) and a relational argument name3 that introduces the string “n” (“cataract”) of the theme argument’s name. observe that the agent argument is bound generically. (6) a. λy λn λx [call(x, y, name(“n”, y))] b. genx [call(x, this medical condition, name(“cataract”, this medical condition))] we assume the naming predicate contained in a nic to imply an underspecified copular relation p (härtl 2020, hoekstra 2023). (7) λp λy λn λx [call(x, y, name(“n”, y) ˄ p(y, n))] p identifies the relation holding between the denotation of the name n, mentioned as “n” in a nic, and the theme argument y. this relation entails that calling y “n” presupposes that y qualifies as an instance of n, as expressed in (8a). accordingly, negating the copular relation in such contexts typically results in a contradiction, see (8b). (8) a. this medical condition is a cataract. b. one calls this medical condition “cataract,” ↯but this medical condition is not a cataract.4 c. the phenomenon is called a “sun halo.” d. this substance is referred to as a “biomaterial.” a copular account also explains why the quoted noun can in principle occur with a determiner in nics, see (8c & d) above. the determiner accompanying the quoted noun hints to the copular sentence entailed in nics, in which the noun fulfills a referential function as the predicative element and, thus, projects within a determiner phrase. this is indicated by the presence of the accusative suffix -en with weak nouns in german like zeuge (‘witness’) in such niq sentences, e.g., eine solche person nennt man einen zeugen (‘a such person calls one a witness-acc’, one calls such a person a witness) in contrast to niq sentences with no determiner, where the noun is morphologically unmarked, cf. eine solche person nennt man „zeuge“ (‘witness-nom’).5 the type of copula involved in niq is an identificational copula. typically, identificational copular sentences contain a demonstrative or definite nominal expression as subject and are used to teach the names of people or things, introduced in the postcopular phrase, e.g., this fabric is velvet, see, among others, higgins (1979), mikkelsen (2011). the subject in identificational copular sentences has a different semantic type than the subject of predicational copular sentences, see geist (2006) and mikkelsen (2005) for analyses. grammatical evidence for this assumption comes from the fact that in a left-dislocation configuration in german, the subject of a copular sentence like (8a) is referred to using a non 2 we remain agnostic regarding the specifics of the lexical entry for the naming predicate. an underspecified representation λn λy λx [call(x, y, n)] may serve as a plausible base for deriving different interpretive variants contextually. 3 the name predicate is included to capture the relational nature of naming, linking a linguistic expression to a referent within a given linguistic and conceptual system. 4 the contradiction arises if there is formal identity between the first and the second conjunct. examples which target metaphorical naming or misnomer do not yield a contradiction, e.g., this insect is called a butterfly, but it is actually not a fly. the same holds for false labeling (the doctor called the lump a tumor, but it turned out not to be a tumor), which likewise does not produce a contradiction. 5 i wish to thank jan wiślicki for bringing this point to my attention. analogous effects can be observed in polish. calling things by their names: towards a unified account for name-informing and mixed quotation 115 referential pronoun (das ‘that’), cf. diese erkrankung, das / ??die ist ein grauer star (‘this conditionfeminine that-neuter / ??the-feminine is a cataract’), see mikkelsen (2005: 74f). observe that the semantic form in (6) is meant to represent the meaning of call in its name-informing use. we believe an underspecification approach to be desirable, with the different manifestations of the naming predicate to be derived compositionally (cf. footnote 3 above). the verbal event and the agent argument have, for example, a generic meaning in nics like in (5) above but adopt a specific interpretation in the description of a speech act like (4a). our account highlights the argument-structural and semantic properties of the naming predicate explicitly contained in a nic. the analysis entails that both the agent argument and the naming predicate itself exhibit generic properties in niq. 3.2 mixed quotation the established view on mixed quotations (mqs) is that they are blends of direct quotation and indirect report. the speaker quotes a segment of what someone else has said, integrating it into their own sentence structure (cappelen & lepore 1997, tsohatzidis 2003). consider the following example. (9) ben declared he’s going to “kick up a huge fuss” today. in mqs, expressions are both mentioned and used denotationally at the same time. this dual role of mentioning and using distinguishes mqs from purely direct or indirect speech, allowing speakers to preserve fidelity to specific wording while adapting the utterance to suit their own purposes or context. recanati (2001) analyzes mq as a case of so-called open quotation, in which the speaker quotes an exact wording while still integrating it into the sentence semantically. this is done without treating the quoted material as semantically inert, instead allowing it to contribute fully to the sentence’s meaning. thus, the quoted phrase is used for demonstration purposes while preserving its usual semantic function, reflecting the dual nature of mixed quotation as both a linguistic and demonstrative act. mqs are not necessarily introduced by a speech-report predicate but can serve the same function. parts of the expression are quoted (and refer to the exact wording), while other parts are formulated in the speaker’s own words. consider the example in (10). (10) tom’s theory: the “protective state” could morph into a “surveillance regime that exhibits authoritarian features.” here, the speaker who produces (10) quotes the expressions protective state and surveillance regime that exhibits authoritarian features, as used by tom, and merges them into their own utterance, i.e., the rest of the sentence. to explain cases where mixed quotation occurs without quotation marks or overt quotative indicators (e.g., this sentence doesn’t contain a hyperbole, it contains a hyperbole), kirkgiannini (2024) proposes a covert operator that refers to an unpronounced element in the syntactic structure that introduces mq even when there are no explicit quotation marks or visible quotative indicators.6 this covert operator is linked to the mq’s not-at-issue (i.e., secondary) content that encodes the fact that an expression was uttered verbatim by a particular individual at a certain point in time. this contrasts with the at-issue content, which is the main propositional information being asserted or discussed. kirk-giannini’s key insight is that mq arises from the interaction between pq and covert material within the sentence structure. this perspective supports a reductionist analysis, in which different types of quotation – such as mq and pq – are unified under a single theoretical framework. in this vein, the fact that the boundary between mq and scare quotation is not always clear-cut further strengthens the case for such a unified account, since the indeterminacy can give rise to ambiguity in interpretive resolution. consider example (11). (11) kim reported that she had received a “pink slip.” 6 i wish to thank stefan hinterwimmer for his valuable input on this point. härtl 116 in this example, the quotes around pink slip could suggest that kim previously used the expression (indicative of a mq) or that the speaker is using scare quotes to signal that the term pink slip is linguistically noteworthy. this ambiguity is crucial evidence for a pragmatic argumentation towards quotes, such as de brabanter’s (2010), that challenge a strict distinction between mq and scare quotation. 3.3 linking niq and mq the idea that one type of quotation can be at least partially reduced to another has been discussed elsewhere. for example, maier (2007) acknowledges that mq shares features with both indirect discourse and pq, which allows quoted words in mq to retain specific characteristics of pq, such as maintaining the original speaker’s indexical expressions. in our context, the question arises whether niq cases like the phenomenon is called a “sun halo,” in which the quoted noun is accompanied by a determiner, should in fact be analyzed as instances of mq. such cases of niq contain quoted material that are apparently both used and mentioned, the latter indicated by the determiner. with a copular account (see 3.1. above), the denotational use in such examples arises naturally as the corresponding noun functions as the predicative element rooted in the implied copular relation. we will leave this issue open at this time. our main claim is that there is no principled distinction between niq and mq regarding a certain semantic element, which we assume to be involved in both types. specifically, we claim that mq, like niq, involves a naming predicate, which is covert in cases like (9) above. the proposed account aligns with maier’s (2014b) analysis, which emphasizes that mixed-quoted expressions mean what a source speaker referred to with certain words, thereby invoking a presupposition about the original utterance. in maier’s dynamic semantic account, the predicate refer to as in what x referred to as “y” plays a crucial role in the compositional structure of mqs by establishing a link between the original utterance and its reference in the current speaker’s discourse. evidence for the assumption that a naming predicate is implied in mqs comes from examples where a naming predicate is explicated (cf. maier 2014b). consider the following example. (12) bryant said he did what he called more “explosive movements” at practice wednesday […]. cbc news, february 06, 2023 the italics are ours. they mark a parenthetical relative clause that denotes a naming event (called). it introduces a mentioned expression (“explosive movements”7), representing the name argument of call. the name argument denotes the name of the theme argument of call, i.e., of what in this case. the theme-argument position is linked with material from the host clause, that is, with the host clause’s direct object (explosive movements). this is similar in reporting as-clauses as in (13). (13) […] he did, as he called it, “computer stuff in support of space programs” […]. middlebury magazine, september 16, 2022 here, the external argument of as is identified with content expressed in the matrix clause. in their analysis, pittner and frey (2023) state that german wie (‘as’) encodes a two-place relation expressing congruence of its arguments, thus, relating an expression contained in the matrix clause (computer stuff in support of space programs), i.e., wie’s external argument, to its internal argument (it). semantically, wie sets the internal argument equal to the external argument. thus, the paraphrase of (13) entails that he did something, and he referred to it as “computer stuff in support of space program.” 7 note that two constituent structures are feasible regarding the relative clause in (12), see below. we postpone the discussion of this point to a later time. a. … [he did [what he called more “explosive movements” …]] b. … [he did [what he called] more “explosive movements” …] calling things by their names: towards a unified account for name-informing and mixed quotation 117 observe that a naming predicate can also be argued to be covertly implied in examples like (9), repeated in (14) below, that do not overtly contain a naming predicate. the negation test in (14) suggests this, as by negating the calling content a contradiction is produced. (14) ben declared he’s going to “kick up a huge fuss” today (↯but never called it that).8 thus, (9/14) denotes that, at a certain time of speaking, ben produced an utterance about some intended act of complaining or protesting, which he referred to as “kicking up a huge fuss.” (15) call(ben, doing something, name(“kicking up a huge fuss”, doing something) ˄ p(doing something, kicking up a huge fuss)) we assert that the copular relation involved here is, once again, an identificational copula. the intended act of ben’s doing something constitutes an instance of kicking up a huge fuss. or put in slang: ben’s doing something is a kicking up a huge fuss. note that, while the agent argument (ben) is specific, the type of binding for the calling event in example (9/14) is not specified. ben may generally refer to the corresponding act as “kicking up a huge fuss,” or he may have used this term only for this particular act. this contrasts with the example in (13), where the calling event (described in the past tense perfective) is specific and not generic. in summary, we explored the role of the naming predicate in unifying the analyses of niq and mq. while the naming predicate is implicit in mq, covertly integrating quoted material into the speaker’s narrative, it is overt in niq, explicitly highlighting the linguistic form of a concept’s conventionalized name. this dual approach allows both types of quotations to establish a referential link between the linguistic expression and its denoted concept, albeit with differing levels of visibility. 4 experimental study this pilot study aims to test whether the at-issueness methodology previously applied in related work can be effectively used to investigate the case at hand. specifically, we are interested in determining whether mqs contain an implicit call component, compared to niqs, which functions as baseline that lexically contains the predicate call, and compared to pqs, which we do not assume to contain a call component. note that our assumption – that mqs contain an implicit call predicate – does not imply that mqs lack other types of quotative features, for example, that someone has literally uttered the words in quotes. however, our main interest lies in the call component as a reflection of an act of nomenclature. to determine whether a call component is present in mqs and niqs, we make use of the distinction between direct and indirect rejections as a measure of how explicitly the call component is represented in each type. the underlying assumption is that semantically active content is more likely to give rise to direct rejections, while inactive or pragmatically inferred content tends to elicit indirect rejections. the distinction between direct and indirect rejections has typically been associated with the notion of at-issueness – the spectrum between primary and secondary content –, which we adopt in this context to capture gradations in how different parts of an utterance are treated in discourse. on the standard view, at-issue content represents the main assertion and directly addresses the question under discussion. therefore, at-issue content is responsive to a direct negation like no, that is not true. not-at-issue content, in contrast, is linked to secondary aspects of an utterance and does not, or only indirectly, contribute to the question under discussion (e.g., fintel 2004, potts 2015, tonhauser 2012). a typical instance of not-at-issue content is an appositive relative clause as in kim, who lives in berlin, fascinates joan, whose content can only be indirectly rejected by means of a discourse-interrupting protest like wait a minute – kim lives in rome! 8 it remains an empirical question to what extent the effect is actually pronounced. note that the contradiction is not produced if the quotes are not present or if the quotes are interpreted as scare quotes. this could be used as an argument for a semantic analysis of quotes, at least in the case of mq (see section 2). we will refrain from further elaboration at this point. härtl 118 we believe that at-issueness is a useful factor for examining the informational status of content within a linguistic expression. the key idea is that not only different expressions contained in a sentence, like main clause and subordinate clause, can exhibit different levels of at-issueness, but even more fundamentally, whether certain semantic information is present at all in an expression can be systematically investigated using at-issue-sensitive tests. this view implies that even information that is only implicitly conveyed by an utterance may display degrees of at-issueness, as we find it, for instance, in the critical-attitudinal content encoded in an ironic utterance. in this way, and this is the perspective we follow in the subsequent study, at-issueness test also helps uncover whether a particular content, even if it is not inherently at-issue, has the potential to manifest as at-issue in a discourse. if a given content has this potential, it should exhibit a stronger tendency toward at-issue rejections; conversely, if this potential is diminished, the corresponding content should show a stronger tendency toward not-at-issue objections. reasonings of this sort can be found, for example, in the discussions of the different contents of slurring expressions (with respect to their descriptive versus derogatory content; carrus 2017, mcready 2010) and ironic expressions (non-literal versus attitudinal content; härtl and bürger 2021). note that the perspective pursued here suggests that (not-)at-issueness as well as a content’s ability to be rejected is a gradual feature and is therefore present to varying degrees in an utterance (cepollaro 2015, härtl & bürger 2021, gutzmann 2023). our main assumption is that content associated with call is present in mqs. we thus hypothesize that this content has a greater potential to be treated as at-issue in mqs compared to pqs, which we assume to not contain a call component. the latter condition is labelled pq_nai in the experimental design. our comparative measure is niqs, which lexically contain a call predicate. mqs should thus gravitate towards niqs with respect to their call component’s at-issueness potential. our at-issue controls are utterances that involve false statements about pqs, thus giving rise to at-issue rejections (pq_ai).9 (16) degrees of at-issueness ha: pq_ai > niq & mqs hb: niq ≈ mq (no difference predicted) hc: niq & mq > pq_nai when formulating the hypotheses and interpreting the results, it is essential to consider the potential influence of politeness on participants’ responses. this effect, driven by a social preference for rejections perceived as more polite, may cause a systematic shift towards not-at-issue ratings, even for content that is inherently at-issue.10 the connection between politeness and not-at-issue rejections has been noted, for example, by syrett und koev (2015). it necessitates careful consideration when analyzing the results, as the politeness effect introduces a bias that should ideally be accounted for in the interpretation of any observed differences between conditions. recall, however, that our study focuses on the gradient potential of a content to be treated as at-issue, rather than a strict binary categorization of at-issueness versus not-at-issueness. this approach acknowledges that even content primarily perceived as not-atissue can exhibit a stronger potential to be treated as at-issue depending on its contextual embedding. therefore, while the politeness effect may influence absolute ratings (affecting all conditions), it does not undermine the relevance of the observed gradience, which remains central to our investigation. to 9 a reviewer asked why we do not group niq with pq_ai, given that the niq condition targets lexically encoded content that could be considered equally at-issue as pq-related content. from previous work we know that factual inaccuracies (e.g., a whale is a fish, 2+2=5 or the term “white flag” consists of three words (i.e., a pq) tend to elicit more direct (i.e., at-issue) rejections, whereas niq and mq inconsistencies, which are based on conventionalization, discourse knowledge, or lexical information, invite more indirect rejections, presumably because they are more strongly influenced by politeness considerations (see below). 10 based on our previous studies employing the paradigm used in this study, the bias caused by this effect is estimated to result in a distortion of approximately one point on a 5-point likert scale, meaning that genuinely at-issue content may be rated around 2 as at-issue. calling things by their names: towards a unified account for name-informing and mixed quotation 119 test our hypotheses, we devised a rating study, the methodology and results of which we will now present. 4.1 method to clarify whether a call component is involved in the meaning of mqs as we can assume it to be present in niqs, mqs were compared to niqs concerning their compatibility with at-issue vs. not-atissue rejections of a call component. our presumption is that semantic information implicitly conveyed by an expression should exhibit a greater potential for compatibility with an at-issue rejection. conversely, information that is not entailed and must be accommodated pragmatically should lean to be compatible with a not-at-issue rejection. 4.1.1 participants forty german native speakers participated in the online survey (aged < 20 n = 2, 20–25 n = 28, 26–30 n = 7, > 30 n = 3). participants were not paid. 4.1.2 material and design the entire experiment was conducted in german. the experimental items followed a consistent structure in which a situation was presented introducing two individuals. then a statement containing a quotation was made by one of the two interlocutors. all quoted expressions contained conventionalized adjective-noun names like blauer brief (‘blue letter’, pink slip), weiße fahne (white flag), etc., which were controlled for lexical frequency using the wortschatz corpus.11 finally, two options were given that represented the second interlocutor’s rejection of a part conveyed with the preceding statement. across all conditions, rejections referred to a call component and were designed as corrective rejections, i.e., as denials that included a reason for why the corresponding content is rejected. as critical conditions, instances of niq and mq were included in the statements, cf. (17a & b).12 in the mq condition, the quoted material consisted of a vp placed in quotation marks to ensure that the quotation is interpreted as a quoted utterance rather than as an instance of scare quotation. in the corresponding mq rejections, the theme argument of call was realized by a demonstrative pronoun (das ‘that’) referring to the semantic content of that vp, see (17b). as controls, we used pqs. in the first control type, all statements made by the first interlocutor were false, calling for at-issue rejections (pq_ai). in the second control type, pqs were followed by rejections targeting a call component, i.e., a component that can only be interpreted through accommodation and, thus, prompting a not-at-issue rejection (pq_nai). consider the examples of the four conditions below. (17) a. name-informing quotation (niq) ingo erzählt anna von einem schreiben in seinem briefkasten. anna stellt in frage, was ingo sagt. ‘ingo tells anna about a letter in his mailbox. anna questions what ingo says. ingo: ein solches schreiben nennt man „blauer brief.” a such letter calls one “blue letter.” ‘such a letter is called “pink slip”.’ anna: (a) das ist falsch, das nennt man anders. ‘that is wrong, that is called something else.’ 11 wortschatz.uni-leipzig.de 12 a possible criticism of including quotational sentences like those in (17b) in the material could focus on the fact that they involve a subordinate clause, which sets this condition apart from all others. such a criticism would presumably argue that the quoted content is more deeply embedded in this condition, making the content potentially harder to access for direct rejection. however, such an argument would only hold if we predicted mqs to differ from the second critical condition, i.e., niqs. such a difference is neither hypothesized, see (16) above, nor observed in the data. härtl 120 (b) sekunde, das nennt man anders. ‘wait a second, that is called something else.’ b. mixed quotation (mq) finn unterhält sich mit uta über oliver. uta stellt in frage, was finn berichtet. ‘finn is talking to uta about oliver. uta questions what finn is reporting.’ finn: oliver verriet vorhin, man würde „möglicherweise eine oliver revealed earlier one would “possibly a rote karte sehen.” red card see.” ‘oliver revealed earlier one would “possibly see a red card”.’ uta: (a) das ist unwahr, oliver hat das niemals so genannt. ‘that is not true, oliver never called it that.’ (b) sekunde, oliver hat das niemals so genannt. ‘wait a second, oliver never called it that.’ c. pq – at-issue rejection (pq_ai) laura erklärt daniel etwas während der konferenz. daniel bezweifelt lauras aussage. ‘laura explains something to daniel during the conference. daniel doubts laura’s statement.’ laura: der begriff „weiße fahne“ besteht aus drei wörtern. the term “white flag” consists of three words. ‘the term “white flag” consists of three words.’ daniel: (a) das ist falsch, der begriff besteht aus zwei wörtern. ‘that is wrong, the term consists of two words.’ (b) moment mal, der begriff besteht aus zwei wörtern. ‘wait a moment, the term consists of two words.’ d. pq – not-at-issue rejection (pq_nai) lea erklärt tom etwas während der teamsitzung. tom zweifelt an, worauf sie sich bezieht. ‘lena explains something to tom during the team meeting. tom doubts what she is referring to.’ lea: der ausdruck „goldene regel“ beginnt mit dem buchstaben ‚g.‘ the term “golden rule” starts with the letter ‘g’. ‘the term “golden rule” starts with the letter ‘g’. tom: (a) falsch, niemand nennt es „goldene regel.“ ‘that’s wrong, nobody calls it “golden rule”.’ (b) wart mal, niemand nennt es „goldene regel.“ ‘wait, nobody calls it “golden rule”.’ the material was consistently structured but varied regarding the lexical material as well as the proper names used. the rejections contained in all three “naming” conditions (niq, mq, and pq_nai) involved naming predicates (e.g., ‘wait a second, oliver never called it that’.). condition pq_ai contained a correction of the false statement. to avoid repetitiveness, a variety of discourse-interrupting phrases (wart mal ‘wait’, sekunde ‘wait a second’, moment mal ‘one moment’, etc.) was used for the not-at-issue rejections, and varying direct negations (das ist nicht wahr ‘that’s not true’, das stimmt calling things by their names: towards a unified account for name-informing and mixed quotation 121 nicht ‘that’s not correct’, das ist falsch ‘that’s wrong’, etc.) were used for the at-issue rejections. to maintain concentration and warrant lexical processing, six content questions (e.g., kam im letzten dialog der ausdruck „schwarzer humor“ vor? ‘did the expression “black humor” occur in the last dialogue?’) were included in the material. these yes-no questions contained an equal balance of true and false statements and occurred after four to six experimental items. for each condition, ten items were included, with no fillers. the study used a within-subjects design, with a single randomized and manually balanced item order to avoid condition clustering. participants were asked to decide which rejection they perceived as more appropriate. to do this, they rated the rejections on a 5-point likert scale, with value 1 representing the at-issue rejection on the left side of the scale and indicating a clear preference for the direct rejection, and value 5 representing the not-at-issue rejection on the right side, indicating a preference for the indirect rejection. this alignment of the values was kept constant to avoid random guessing. both rejection options were placed below value 1 and value 5, respectively, of the scale. participants were able to rate both rejections as equally (in)appropriate by choosing the mid value. moreover, they could rate one rejection as more appropriate than the other by using the values in between, without being forced to indicate a clear preference towards one option. 4.1.3 procedure in an online questionnaire on the platform sosci survey13, participants were asked to evaluate forty experimental items. every participant received the same questionnaire. prior to the study, participants had to answer questions regarding their native language to be allowed to further partake in the survey, as well as their age. an explanation was given (in german) about the procedure of the study: the participants had to assess how appropriate certain short negative reactions are in a dialogue between two people. further, it was mentioned that in all dialogues, a term or expression appears in quotation marks and that these denote either fixed terms (e.g., “pink slip”) or parts of direct speech illustrated by an example (e.g., max says he enjoyed the day “to the fullest.”). an example item was given in the instruction as well as an explanation of the rating scale. further, participants were informed that they would be asked a brief (yes-no) content-related question at irregular intervals. afterwards, participants had a short training period that included three items exemplifying the range of conditions and a content question. after that, the first test item was presented. the entire experiment lasted about 20 minutes. 4.2 results only questionnaires in which all items were rated were included in the statistical analysis. the accuracy of the responses to the content questions was not considered. a repeated-measures anova was performed with minitab’s14 general linear model procedure by subject for the dependent variable (rating). the independent variable was condition (niq, mq, pq_ai, pq_nai), which was included as a fixed factor. the factor subject was treated as random. table 1 summarizes the mean ratings for each condition. 13 soscisurvey.de 14 minitab.com härtl 122 table 1. rounded mean ratings and sds for the individual conditions the analysis revealed highly significant differences between the individual conditions, f(3,156) = 59.1622, p < .0001. pairwise comparisons between conditions were conducted using dependent samples t-tests. to control for the risk of type i errors, a bonferroni correction was applied. the results revealed significant differences between the pq_ai and mq conditions, t(39) = −13.45, p < 0.001, the pq_ai and niq conditions, t(39) = −11.66, p < 0.001, and the pq_ai and pq_nai conditions, t(39) = −17.78, p < 0.001, indicating consistently lower (i.e., at-issue) ratings for pq_ai compared to all other conditions. additionally, significant differences were found between mq and pq_nai, t(39) = −3.51, p = 0.001, and niq and pq_nai, t(39) = −5.91, p < 0.001, suggesting higher (i.e., not-at-issue) ratings for pq_nai. in contrast, the comparison between mq and niq showed no significant difference after bonferroni correction, t(39) = 2.51, p = 0.012, suggesting no reliable difference between these conditions. 4.3 discussion the results reveal that false statements concerning pqs (see 17c) tend to elicit at-issue rejections more than any other condition, providing clear support for hypothesis ha (see section 4 above). notably, the niq conditions, despite their lexical inclusion of a call predicate, lean towards not-at-issue rejections. this can be explained by the politeness effect discussed above, where at-issue rejections are generally perceived as more face-threatening or less polite than not-at-issue rejections. politeness considerations can influence participants’ judgments, causing them to favor less direct forms of rejection (koev 2018, syrett & koev 2015). this bias introduces a systematic shift towards not-at-issue ratings (affecting all conditions), even for content that is inherently at-issue. moreover, the results indicate that the pq_nai conditions elicit not-at-issue rejections more frequently than any other condition, providing support for hypothesis hc. crucially, no significant difference was observed between the mq and niq conditions, which is consistent with hypothesis hb. taken together, the overall pattern underscores the gradient potential of content to be treated as at-issue in discourse, rather than a binary distinction. even content primarily perceived as not-at-issue can exhibit a stronger potential to be treated as at-issue, depending on its contextual embedding. 5 conclusion this paper has investigated the semantic connection between niq and mq, proposing a unified account. niqs emphasize the linguistic form of a concept’s conventionalized name, while mqs integrate direct and indirect speech. both types of quotation involve a naming predicate, which is explicit in niqs but covert in mqs. in a pilot questionnaire-based rating study, we applied a method used in previous research, which proved effective in this context. by embedding the target expressions in short dialogue sequences, the study provided evidence for similar degrees of the naming component’s potential to be treated as at-issue in both mqs and niqs, supporting the hypothesis that they share a related semantic structure. our account aligns with the views of kirk-giannini (2024) and maier (2014b), both of whom argue for a pure-quotational element in types of quotations in question. additionally, it follows maier’s (2014b) assumption of a free relative clause (i.e., what x referred to as “y”) in mqs. however, we offer a more nuanced analysis by positing that the predicate in question describes a naming relation. furthermore, our analysis parallels kirk-giannini’s (2024) proposition that mqs involve a covert operator, which refers to an unpronounced element in the compositional structure. this operator introduces mq condition mean rating sd pq_ai 2.1 1.4 niq 3.1 1.3 mq 3.3 1.3 pq_nai 3.7 1.2 calling things by their names: towards a unified account for name-informing and mixed quotation 123 even when there are no explicit quotation marks or overt quotative markers. in contrast, our proposal suggests the presence of a covert naming operator in mqs. as concerns the unification aspect of our account, stern’s (2022) proposal that quotations can be understood as pictures may provide a framework for refining the unified semantic approach proposed in this study. stern’s pictorial analogy underscores the dual role of quotations in both representing and exemplifying linguistic content. such a perspective aligns closely with the dual use-and-mention function that is central to both niq and mq. by conceptualizing quotations as simultaneously conveying representational and demonstrative content, stern’s account offers a theoretical foundation that strengthens the case for a shared semantic basis between niq and mq. a couple of open questions remain. first, it is still an open question whether the naming predicate, which we claim to be present in mqs, is semantically entailed or pragmatically accommodated, especially in cases where no parenthetical relative clause makes the naming predicate explicit, as in the example in (11) above. to address this, online processing data could provide further insight. second, a compositional analysis is still pending, showing how the call projection in mqs integrates with the projection of the embedding speech report, on the one hand, and, on the other, how a parenthetical relative clause that explicates the naming predicate integrates with the rest of the clause. we believe the accounts from grosu (2003), tsohatzidis (2005), and pittner & frey (2023) provide good starting points. finally, an explanation still needs to be delivered as to whether our analysis also applies to direct quotation (e.g., ben declared, “i am going to kick up a huge fuss today.”). we do not believe this to be the case, with the main reason being that direct quotations are not compatible with expressions of a call predicate (cf., ben declared, (??what he called) “i am going to kick up a huge fuss today.”). references john m. anderson (2004). on the grammatical status of names. language, 80(3):435–474. john biro (2012). calling names. analysis, 72(2):285–293. herman cappelen and ernie lepore (1997). varieties of quotation. mind, 106(423):429–450. herman cappelen and ernie lepore (1999). using, mentioning and quoting: a reply to saka. mind, 108(432):741–750. simone carrus (2017). slurs. at-issueness and semantic normativity. phenomenology & mind, 12:84–97. bianca cepollaro (2015). in defense of a presuppositional account of slurs. language sciences, 52:36–45. philippe de brabanter (2010). the semantics and pragmatics of hybrid quotations. language and linguistics compass, 4(2):107–120. philippe de brabanter (2013). a pragmaticist feels the tug of semantics: recanati’s “open quotation revisited”. teorema, 32(2):129–147. philippe de brabanter (2023). quotation does not need marks of quotation. linguistics, 61(2):285– 316. donald davidson (1979). quotation. theory and decision, 11(1):27–40. kai von fintel (2004). would you believe it? the king of france is back! presuppositions and truthvalue intuitions. in marga reimer and anne bezuidenhout, editors. descriptions and beyond, pages 315–341. oxford university press, oxford. ljudmila geist (2006). die kopula und ihre komplemente. zur kompositionalität in kopulasätzen. max niemeyer verlag, tübingen. alexander grosu (2003). a unified theory of ‘standard’ and ‘transparent’ free relatives. natural language & linguistic theory, 21(2):247–331. daniel gutzmann (2023). gradient at-issueness, minimum relevance, and propositional prominence. theoretical linguistics, 49(3–4):239–247. holden härtl (2018). name-informing and distancing sogenannt (‘so-called’): name mentioning and the lexicon-pragmatics interface. zeitschrift für sprachwissenschaft, 37(2):139–169. härtl 124 holden härtl (2020). referring nouns in name-informing quotation: a copula-based approach. in michael franke, nikola kompa, mingya liu, jutta l. mueller, and juliane schwab, editors. proceedings of sinn und bedeutung 24 vol. 1, pages 291–304, osnabrück and berlin. holden härtl and tatjana bürger (2021). ‘well, that’s just great!’ – an empirically based analysis of non-literal and attitudinal content of ironic utterances. folia linguistica, 55(2):361–387. jarich hoekstra (2023). from naming verb to copula: the case of wangerooge frisian heit. journal of germanic linguistics, 35:97–147. francis r. higgins (1979). the pseudo-cleft construction in english. london: routledge. cameron d. kirk-giannini (2024). covert mixed quotation. semantics and pragmatics, 17(5):1–67. todor koev (2018). notions of at-issueness. language and linguistics compass, 12(12):1–16. emar maier (2007). mixed quotation: between use and mention. proceedings of lenls 2007, pages 1– 15, miyazaki. emar maier (2014a). pure quotation. philosophy compass, 9(9):615–630. emar maier (2014b). mixed quotation: the grammar of apparently transparent opacity. linguistics and philosophy, 37(3):255–307. elin mccready (2010). varieties of conventional implicature. semantics & pragmatics, 3:1–57. line mikkelsen (2005). copular clauses: specification, predication and equation. john benjamins, amsterdam. line mikkelsen (2011). article 68: copular clauses. in claudia maienborn, klaus von heusinger, and paul portner, editors. semantics. an international handbook of natural language meaning, vol. 1 (= hsk series, 33.1), pages 1805–1829. de gruyter, berlin. karin pittner and werner frey (2023). german wie-comment and reporting clauses: a comparison with so-parentheticals. in łukasz jędrzejowski and carla umbach, editors. non-interrogative subordinate wh-clauses (online ed.), pages 328-364. oxford university press, oxford. christopher potts (2015). presupposition and implicature. in shalom lappin and chris fox, editors. the handbook of contemporary semantic theory, pages 168–202. wiley-blackwell, malden, massachusetts. stefano predelli (2003). scare quotes and their relation to other semantic issues. linguistics and philosophy, 26(1):1–28. willard v. o. quine (1981). mathematical logic (rev. ed.). harvard university press, cambridge, massachusetts and london. françois recanati (2001). open quotation. mind, 110(439):637–687. paul saka (1998). quotation and the use-mention distinction. mind, 107(425):113–135. marcel schlechtweg and holden härtl (2020). do we pronounce quotation? an analysis of name-informing and non-name-informing contexts. language and speech, 63(4):769–798. ori simchen (1999). quotational mixing of use and mention. philosophical quarterly, 49:325–336. josef stern (2022). quotations as pictures. mit press, cambridge, massachusetts. kristen syrett and todor koev (2015). experimental evidence for the truth conditional contribution and shifting information status of appositives. journal of semantics, 32:525–577. judith tonhauser (2012). diagnosing (not-)at-issue content. in h. greene, editor. proceedings of semantics of under-represented languages of the americas (sula), glsa, pages 239–254. amherst, massachusetts. savas l. tsohatzidis (2003). lost hopes and mixed quotes. belgian journal of linguistics, 17(1):213– 229. corey washington (1992). the identity theory of quotation. journal of philosophy, 89(11):582–605. dialogue & discourse 16(2) (2025) 1–34 doi: 10.5210/dad.2025.201 speech acts that support other speech acts katja jasinskaja katja.jasinskaja@uni-koeln.de university of cologne frank zickenheiner frank.zickenheiner@uni-bonn.de university of bonn editor: jonathan ginzburg submitted 06/2024; accepted 11/2025; published online 11/2025 abstract theories of discourse structure and existing discourse structure annotation schemes (e.g. mann and thompson, 1988; sanders, 1997; asher and lascarides, 2003; webber et al., 2019) often make a distinction between propositional, subject-matter, or semantic coherence relations on the one hand, and speech-act-level, presentational, or pragmatic relations, on the other. while there have been several convincing attempts to circumscribe the space of all possible propositional relations and to subdivide it into theoretically motivated subcategories (sanders et al., 1992; kehler, 2002), to date there is no comparable comprehensive taxonomy for speech-act-level relations. this paper develops a fragment of such a taxonomy, which describes what we call support relations—relations that connect two speech acts of the same speaker iff one of them fails to achieve its goal and the other helps achieve that same goal, as for instance, in evidence relations, where one speech act makes the proposition asserted in the other more believable. we provide conceptual motivation for the proposed categories grounded in insights from sociolinguistic, psychological, and philosophical studies of human communication and illustrate the categories with examples from naturally occurring discourse, some of which do not fit easily into any existing classifications. keywords: coherence relations, speech acts, communicative goals, dialogue, monologue, belief 1. introduction the central assumption of a large group of approaches to discourse structure (mann and thompson, 1988; sanders et al., 1992; kehler, 2002; asher and lascarides, 2003) is that a discourse is coherent to the extent that hearers or readers are able to connect the sentences and larger discourse units it consists of with meaningful links—coherence relations (alias rhetorical relations or discourse relations, see jasinskaja and karagjosova, 2020, for an overview). while details might differ between the approaches, many of them make a distinction between propositional, subject-matter, event-level, or semantic relations on the one hand (1-a), and speech-act-level, presentational, or pragmatic relations, on the other (1-b). so in (1-a), a causal coherence relation holds between the described events of pushing and falling. in (1-b), the relation can be characterised in two ways: either as a causal relation that holds between the speech act of asking the question and the fact that there is a good movie on (sweetser, 1990), or as a relation between two speech acts, a question and an assertion, where the purpose of the assertion is to explain why the question is justified or relevant to the speaker. ©2025 katja jasinskaja and frank zickenheiner this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). jasinskaja and zickenheiner (1) a. max fell. john pushed him. b. what are you doing tonight? because there’s a good movie on. the focus of this paper is on cases like (1-b) and on pragmatic relations of the latter kind—those that connect two speech acts in discourse. this paper pursues three goals: a comprehensive taxonomy: while a number of speech-act-level relation types have been identified in previous studies and existing discourse structure annotation schemes (asher and lascarides, 2003; sanders et al., 2009; webber et al., 2019), to date there is no comprehensive taxonomy of such relations. the research reported in this paper seeks to shed light on the question of what kinds of coherence relations between speech acts exist, should exist, and are theoretically possible. the emphasis is here on ‘should’ and on ‘theoretically possible’, as our approach is not empirical or exploratory in the sense of looking for speech-act-level relations in naturally occurring discourse and trying to define fitting categories for the cases we find. our goal is to find out how speech acts should relate to each other based on a theoretical understanding of what a speech act is, what definitional properties it has, and how those properties impose constraints on speech act combinatorics. the question of the existence of an exhaustive list of coherence relations is fundamental in relational theories of discourse coherence. without an exhaustive list, there is no substance in the claim that a discourse is coherent only if we can connect all its parts by some coherence relation. if we fail to connect the discourse units with a relation from a given list, knowing that the list is not exhaustive, we still cannot say whether the discourse is therefore incoherent or we just have not yet encountered a suitable relation. this has been a major point of criticism directed at the empirical, bottom-up approach of rhetorical structure theory (rst, mann and thompson, 1988), which has never claimed to provide a comprehensive taxonomy (see discussion in knott and dale, 1994; kehler, 2002; jasinskaja and karagjosova, 2020). on the other hand, the existing conceptually-driven top-down approaches have not provided a comprehensive taxonomy of pragmatic relations so far. kehler (2002) offers a prime example of the theoretical methodology we adopt, but only covers semantic relations. while sanders et al. (1992) propose a top-down taxonomy of pragmatic relations, in the next section we argue against their specific way of dealing with the problem. some of the same criticism also applies to segmented discourse representation theory (sdrt, asher and lascarides, 2003). in this paper we argue that coherence relations between speech acts are best understood in terms of relations between their communicative goals. at this point, we only develop a fragment of the taxonomy, focusing on what we call support relations, i.e. relations where one speech act is produced as a means to successfully achieve the communicative goal of another speech act of the same speaker. the class of support relations overlaps several prominent relation classes familiar from previous studies, in particular, presentational relations in rst and some kinds of self-repairs (levelt, 1983; clark, 1994). however, it also includes cases that do not seem to fit comfortably into any existing classifications or annotations schemes. the relation between (2-b) and (2-c) is a case in point: (2) a. “pa told peter he wanted him to be chairman.” b. “sure he did. c. if you don’t believe me, ask peter.”1 1. from night over water by ken follett. note that (2-a) is uttered by one speaker, and (2-b) and (2-c) by another, as indicated by the quotation marks. a support relation holds between the utterances of the second speaker. 2 speech acts that support other speech acts the speech act in (2-c) is a directive, which asks the addressee to seek evidence for the proposition asserted in (2-b). its purpose is to make (2-b) more believable, but it does so without directly providing evidence, and therefore does not fall under the standard definitions of evidence relations. the purpose of a comprehensive taxonomy of speech act relations is not only theoretical, but also empirical—to provide a better coverage of previously uncategorised cases, such as (2). coherence relations in dialogue: the focus of relational theories of discourse structure has traditionally been on written monologues that consist predominantly of assertions. even though extensions of the approach to dialogue have been proposed (taboada, 2004; asher and lascarides, 2003; lascarides and asher, 2009), the coverage of non-assertive speech acts remains limited. the framework developed in this paper naturally accommodates speech acts of all types. besides, since the need to start a second attempt to achieve the same communicative goal usually arises when the first attempt was in some way unsuccessful, support relations help describe sequences where things do not go according to the speaker’s plan and cover a range of cases that fall under the notions of disfluency and self-repair—phenomena typical in spontaneous spoken communication. grounding the structure of monologue in interaction: support relations in the definition we propose are relations between speech acts of the same speaker. therefore, they can occur both in conversation, for instance, where one speaker produces a turn at talk that consists of more than one utterance, and in bona fide monologue, like an email, a lecture, or a software user manual. however, by viewing support relations as a communicative practice that is called for when the speaker encounters or anticipates resistance from the addressee in adopting the speaker’s communicative intentions, we look at this coherence relation type as something that exists in interaction, be that with an actual physically present addressee, or a projected hypothetical audience. this is based on the idea that even extended written monologue is embedded in interaction. furthermore, historically, writing and most monological text types are more recent inventions than spoken face-to-face dialogue. they are also acquired later in the course of individual development than face-to-face dialogue communication skills. so it stands to reason that at least some discourse-structural patterns that we find in monologue will have developed from communicative practices used in dialogue. in this paper, we propose a way to look at coherence relations used in monologue through the lens of dialogue. while we are not going to provide proof or empirical evidence for this claim, we would like to put it out there as a hypothesis for future research to engage with. this paper is structured as follows. in section 2, we argue that coherence between speech acts is governed by a different set of principles and should not be treated by analogy with propositionallevel or semantic relations. we discuss previous work on speech-act-level coherence in dialogue and show how those findings can be applied to monologue. finally, the notion of support relation is introduced in that section. in section 3, we define goals of speech acts and identify necessary subgoals that a speech act must achieve in order to be successful. section 4 presents a fine-grained classification of support relations along two main dimensions: (a) the subgoal at which the initial speech act fails, which the supporting speech act tries to repair; (b) the means available for doing so, depending on the type of goal state to be achieved (belief, desire, emotion, action, etc.). finally, in section 5 we outline the general idea for how our approach extends to other kinds of relations between speech acts, going beyond support. 3 jasinskaja and zickenheiner 2. relations between speech acts this section reviews some previous approaches to coherence at the level of speech acts, focusing on the question of how relations between speech acts make a discourse coherent, and, conversely, what it means for a discourse to be incoherent at the speech act level. section 2.1 looks into the matter from the perspective of the by far better understood coherence relations at propositional level and argues against the common assumption that the same coherence principles apply at both levels. section 2.2 summarises some relevant insights concerning coherence relations between speech acts across speakers in dialogue. section 2.3 discusses some previous ideas on how these insights can be made fruitful to understand coherence between utterances of the same speaker, within a single conversational turn or in monologue more generally. finally, building on these ideas, section 2.4 introduces the notion of support relations, a subclass of relations between speech acts of the same speaker whose contribution to discourse coherence is the main subject of this paper. 2.1 semantic relations at speech act level many existing approaches to discourse coherence (mann and thompson, 1988; redeker, 1990; sweetser, 1990; sanders et al., 1992; sanders, 1997; asher and lascarides, 2003) make a distinction between coherence relations at propositional and at speech act level. the exact boundary might be drawn slightly differently depending on the approach, but the general idea is roughly the same. here we will refer to sanders et al.’s (1992, pp. 7–8) definitions of what they call semantic vs. pragmatic relations: ‘a relation is semantic if the discourse segments are related because of their propositional content. in this case the writer refers to the locutionary meaning of the segments. the coherence exists because the world that is described is perceived as coherent.’ ‘a relation is pragmatic if the discourse segments are related because of the illocutionary meaning of one or both of the segments. in pragmatic relations the coherence relation concerns the speech act status of the segments. the coherence exists because of the writer’s goal-oriented communicative acts.’ there are two relevant things to note about these definitions. first, note the asymmetry between them. the definition of semantic relations clearly specifies what it is about the locutionary meaning of the discourse segments that makes the sequence coherent—namely the coherence of the world described. following david hume’s classification of relations between ideas, kehler (2002) proposes that there are exactly three ways in which elements of the described reality can cohere or belong together—by causal relations, by spatio-temporal contiguity, or by resemblance (similarities and differences). if we can recognise such relations between objects, states, and events in the world, we perceive that world as coherent, otherwise we don’t. if we can infer these relations from a description of a complex state of affairs, we perceive that description as a coherent discourse at the semantic level. in contrast, sanders et al.’s definition of pragmatic relations does not specify what it is about the goal-oriented communicative acts that creates coherence. however, in practice, the authors simply transfer the same set of relations that are motivated from the point of view of coherence in the world (e.g. causal relations) to the domain of speech acts. although the issue is not explicitly addressed in that work, in our understanding, this basically amounts to saying that coherence between speech 4 speech acts that support other speech acts acts is governed by the same principles as coherence between propositions and ultimately depends on coherence between states and events in some relevant world. (while this is a sensible null hypothesis, below we argue that it is wrong.) second, note that according to sanders et al.’s definition, in pragmatic relations the discourse segments are related because of the illocutionary meaning of one or both of the segments. that is, we have two kinds of cases. the first one is where the relation holds between the performance of a speech act and the content, or the proposition expressed by another discourse segment. we will refer to this case as a proposition–speech-act relation. the second case is where the relation holds between two performances of two speech acts—a speech-act–speech-act relation. proposition–speech-act relations are the better studied of the two. this is the standard way of looking at examples like (1-b). the connective because encodes a causal relation between the content of its syntactic argument, expressed by the subordinate clause, and another argument recovered from the context. in (1-b), that argument happens to be the performance of the question speech act, so (1-b) can be paraphrased as (3-a). but since the second argument is a proposition rather than a speech act, (1-b) cannot be paraphrased as (3-b). (3) a. i ask you what you are doing tonight because there’s a good movie on. b. #i ask you what you are doing tonight because i inform you that there’s a good movie on. the idea goes back to sweetser (1990), but is also found in metatalk relations in segmented discourse representation theory (sdrt asher and lascarides, 2003), speech act relations in the annotation scheme of the penn discourse tree bank (webber et al., 2019), and in sanders et al.’s own work. as pointed out above, the tacit assumption here is that the same relations that exist at the propositional level will also be possible between a speech act and a proposition. in the more than thirty years since sweetser’s and sanders et al.’s seminal work one would expect to have encountered pragmatic versions of all sorts of semantic relations. however, this only appears to be the case for relations that have a causal or conditional component, primarily explanation, as in (1-b), result (4-a) and conditional (4-b) (result* and consequence*, respectively, in asher and lascarides, 2003, p. 334), negative conditional (4-c) (webber et al., 2019, p. 23) and concession (4-d). (4) a. i’m cold. please close the window. b. if you failed the test, then why should i listen to you? c. unless you’re on a diet, there are some cookies in the cupboard. d. although i hate to say it, please don’t panic purchase or hoard decaf.2 on the other hand, we have not seen any discussion of proposition–speech-act versions of parallel, contrast (other than the denial-of-expectation type of contrast ≈ concession), or narration so far. in fact, we even find it difficult to construct a coherent sequence where the whole point of one sentence is to describe an event that happens before or after the speech act expressed by the other, which would fit the definition of a proposition–speech-act narration. we are also not familiar with any attempts to explain these gaps in the paradigm. sdrt might explain away the absence of metatalk parallel and contrast by pointing out that metatalk counterparts only exist for contentlevel relations in sdrt’s narrow understanding of the term, whereas parallel and contrast are not 2. from https://www.mycuppa.com.au/blogs/news/aug-2022. last checked on november 17, 2025. 5 jasinskaja and zickenheiner content-level, but text-structuring relations (asher and lascarides, 2003, pp. 465–466). however, why isn’t there a metatalk version of narration, which is a content-level relation by all definitions? speech-act–speech-act relations have received less attention, but for instance pdtb treats them in the same way as proposition–speech-act relations. webber et al. (2019, p. 25) characterise (5-a) as a concession that holds between two speech acts and can be paraphrased as (5-b). (5) a. he lived in peking, or should i say beijing, for 20 years. b. while i say he lived in peking, it might be more accurate to say he lived in beijing. the problem is that if we continue to follow the same logic, transferring the same set of relations from the domain of propositions to the domain of speech acts, for speech-act–speech-act relations this makes even less sense. for instance, one could argue that narration holds between the two speech acts in (6-a), because one immediately follows the other in time and the sequence can be paraphrased as (6-b), but the sequence is nevertheless perceived as incoherent. (6) a. # john broke his leg. i like plums. (knott and dale, 1994) b. i say john broke his leg, and then i say i like plums. moreover, any pair of adjacent speech acts in discourse are spatio-temporally contiguous by definition. if spatio-temporal contiguity were enough to make a discourse coherent at speech act level, then all sequences of speech acts would be coherent. clearly, such a notion of discourse coherence would not be very useful. a possible objection to this argument could be that it is not the actual spatio-temporal contiguity or adjacency that is relevant to establish coherence between speech acts, but the expected adjacency (see fetzer, 2013, on a related distinction between adjacency position and adjacency expectation). the sequence in (6-a) is incoherent, because the second speech act presents an unlikely, unexpected follow-up to the first. in contrast, (7) is coherent because the second speech act is a felicitous answer to the question and in that sense is expected in that context. one could even argue that a’s question causes b to answer, so a causal coherence relation holds between these speech acts. (7) a: what’s your favourite fruit? b: i like plums. the problem with this view is that it does not really give an answer to the question what makes (7) coherent at speech act level, but simply restates the question. intuitively, (7) is coherent for a different reason than (1-b), (2), or (5-a) is. we want to understand that difference rather than just saying that all four are coherent because some speech act presents a likely follow-up to another. therefore, we will not attempt to save the idea that coherence at the semantic and the pragmatic level is governed by the same principles and described by the same set of coherence relations. instead, we will approach the issue from a completely different, genuinely speech-act-oriented perspective, focusing on speech-act–speech-act relations. while our approach might ultimately also help explain the gaps in the paradigm of proposition–speech-act relations, this is not our goal in this paper. 2.2 relations between speech acts in dialogue interestingly, the foundation for an answer to the question of what makes a sequence of speech acts coherent was already laid by austin (1962, p. 160), who pointed out that many speech acts 6 speech acts that support other speech acts have illocutionary forces that “invite by convention a response or sequel”. for example, if one asks someone to do something, the hearer is invited to perform a certain action, if someone asks a question the hearer is invited to give an answer, and so on. searle (1969, 1983) labels the relevant property of a speech act as the condition of satisfaction. for example, the condition of satisfaction of an order is that the hearer obeys that order, the condition of satisfaction of a question is that the hearer gives an answer to that question, and so on. obviously, not all speech act types will be “satisfied” by another speech act, but some will (question–answer pairs being the paradigmatic case), and in those cases one could say that the sequence is coherent because the second speech act fulfils the satisfaction condition of the first. in the philosophical tradition of speech act theory, this idea was further developed by sbisà (2002), who argued that conventional effects of illocutionary acts make sequences of speech acts possible, and sequences, in turn, make it possible for individual speech acts to achieve their conventional effects. however, there have been altogether few attempts to apply the notions of traditional speech act theory to explain coherence of speech act sequences (notable exceptions being sbisà (2002); franke (1990), fritz and hundsnurscher (2009) who analysed possible reactions to accusations and hindelang (2010), who classified speech acts in sequences based on the work of franke, fritz and hundschnurscher). the issue received more interest in conversation analysis. sacks and schegloff coined the term adjacency pair, referring to pairs of utterances where the second one constitutes a proper reaction to the first one. typical examples are question – answer, greeting – counter greeting or offer – acceptance/refusal. sacks and schegloff (1973) did not attempt to explain these pairings in terms of the notions of speech act theory, but what is clear from the known listings is that different illocutionary forces are paired with different types of reactions. within this approach one could say that a pair of speech acts is coherent if it constitutes an adjacency pair, and even if we do not fully understand which properties of the speech acts are responsible for the legitimate pairings, there clearly is a connection. finally, an utterance may constitute a coherent follow-up to another one, not only if it satisfies the expectations set up by the first utterance, but also when it frustrates those expectations, but only in a way that ultimately serves the purpose of the conversation. franke (1990), for example, introduced decision-preparing reaction moves, which are performed to help decide between a positive and a negative response to a previous speech act. franke’s decision-preparing moves include clarification requests (which were also studied at length in computationally oriented approaches to dialogue, e.g. ginzburg, 2012) and speech acts that raise other kinds of problems related to a previous speech act. with respect to speech-act-level coherence that means in particular that a clarification request can constitute a coherent follow up to a speech act of any type, as long as it helps solving communicative problems related to it. adjacency pairs and relations between a speech act and a clarification request it triggers are, in our view, genuine speech-act–speech-act coherence relations, because the reason why those speech acts constitute a coherent sequence can only be understood if we take into account their illocutionary forces and/or perlocutionary goals. incidentally, all these relations connect speech acts produced by two different speakers where the second one is a reaction to the first. franke’s (1990) classification of ‘speech acts of the second move’ can be seen as a proposal towards a comprehensive taxonomy of such relations, and ginzburg et al.’s (2022) response spaces is a fragment that describes reactions to questions. these relations will normally occur in interactive discourses between two or more speakers, so one might get the impression that coherence at speech act level is a phenomenon 7 jasinskaja and zickenheiner restricted to dialogue. the next section shows what we can learn from dialogue about coherence relations between speech acts of the same speaker in dialogue or monologue. 2.3 from dialogue to monologue the notions of dialogue and monologue can be understood holistically, as categories of discourse types; for instance, an interview is a dialogue, and a novel is a monologue. in this paper we adopt a more technical definition: a sequence of utterances is a dialogue if it contains an exchange of speaker and hearer roles between at least two participants. a sequence of utterances is a monologue if all those utterances are produced by the same speaker without interruption by another speaker. in this sense, an extended turn at talk is a monologue, so monologues can be part of dialogues. coherence relations between utterances of the same speaker will generally occur in monological sequences, but they can also hold between two non-adjacent utterances of the same speaker separated by an utterance of a different speaker in a dialogical sequence. in a dyadic exchange between speakers s1 and s2, after s2’s reaction to s1’s initial utterance, the turn usually comes back to s1. one might wonder whether the relationship between s1’s initial speech act and her reaction to s2’s reaction can be fruitfully analysed in speech-act-theoretic terms. for instance, s1 could use an ambiguous pronoun in the first move (8-a), s2 could react with a clarification request in the second move (8-b), to which s1 could provide the desired clarification in the third move (8-c). the question is then: how do the first and the third utterance, (8-a) and (8-c), cohere at speech act level? which properties of these utterances as goal-oriented communicative acts are essential for the overall coherence of the sequence? (8) a. s1: at that point, it was over across the road. b. s2: what do you mean ’it’? c. s1: the warehouse. a systematic classification of ‘speech acts of the third move’ has been developed by franke (1990). after a clarification request or some other kind of reactive move that frustrates the goals of the initial utterance, the speech act in the third move may belong to one of three classes according to the relationship between its goals and the goals of the initial speech act: retractive speech acts signal that s1 completely gives up on the goal of her initial speech act; revising speech acts modify the communicative goal of the initial speech act in such a way that it is more likely to be achieved; finally, re-initiative speech acts present a second attempt to achieve the goals of the initial speech act without modification. the third move in (8) is an instance of a re-initiative speech act. in (8-a) s1 is trying to inform s2 that the warehouse was across the road; in (8-c) she is still trying to do so. in this respect, (8-c) presents a coherent follow-up to (8-a) and (8-b) because it continues to pursue the goals of (8-a) after (8-b) indicates that they were not reached in the first attempt. notice that re-initiation is a relation between utterances of the same speaker. but speakers often do not wait for an explicit clarification request from their interlocutor, but detect potential communicative problems by monitoring their own speech, and produce re-initiative speech acts without conversational turn transition, as is the case in most self-repairs (levelt, 1983). therefore re-initiation can function as a speech-act-level coherence relation in a monologue. this idea is developed by ginzburg et al. (2014) as a general approach to self-repair moves and other kinds of speech disfluencies. in this approach, (9) has the same underlying structure as (8), with the 8 speech acts that support other speech acts difference that in (8) the clarification request is made explicit, whereas in (9) it constitutes an implicit question under discussion (qud).3 (9) at that point, it, the warehouse was over across the road. as far as the source of coherence is concerned, we could say that the warehouse constitutes a coherent interjection within the utterance at that point, it was over across the road because it ultimately helps achieve the initial goal of informing the hearer that the warehouse was across the road. coherence is established at the level of the communicative goals. 2.4 support relations adjacency pairs, clarification requests, retractions, revisions, re-initiations, and (self-) repairs are relational categories developed in previous research that, in our view, genuinely pertain to the speech act level of coherence because the source of coherence (the reason why a sequence is coherent) in all these cases has something to do with the communicative goals of the speech acts involved and the relative success or failure in achieving those goals. of course, causality also plays an important role here since one speaker communicating a certain goal can cause the other speaker to want to satisfy that goal, or the failure to achieve a certain goal can cause the speaker to start another attempt, but as we argued in section 2.1, in order to achieve a deeper understanding of coherence at speech act level, we should shift our focus from causality in general to more specific relationships between communicative goals. therefore, the focus of this paper will be on a subclass of relations between speech acts that we will call support relations:4 (10) support relations: speech act s of speaker a supports speech act n of the same speaker iff a. a believes that n has failed or will fail to achieve its goal b. a believes that s will help to achieve that goal. as should be clear from the discussion in previous sections, this definition primarily covers selfrepairs and franke’s re-initiations. it is also closely related to the notion of presentational relations (11) in rst (mann and thompson, 1988): (11) presentational relations: presentational relations are those whose intended effect is to increase some inclination in the reader, such as the desire to act or the degree of positive regard for, belief in, or acceptance of the nucleus. (mann and thompson, 1988) all presentational relations in rst are relations between a nucleus (n), a discourse unit that is more central to the overall purpose of the discourse, and a satellite (s), a discourse unit whose function is defined in relation to the nucleus. the n and s notation in (10) is our adaptation of this distinction to support relations, where s stands for support, and n for the unit being supported. there is a substantial overlap in the range of cases covered by the definitions (10) and (11). both types of relations have in common that s has the function to increase the chances of the success of n. for instance, if the goal of n is that the hearer believe a proposition and s increases the hearer’s 3. (9) is the original example cited by ginzburg et al. (2014). the example in (8) is constructed from it. 4. note that our notion of support is quite different from that of grosz and sidner (1986), who we otherwise have a lot in common with, see section 3 for more discussion. 9 jasinskaja and zickenheiner belief in that proposition, as in the case of evidence relations like that in (12), then s also helps achieve the goal of n. (12) a. he must have been here recently. b. there are his footprints. however, in this paper, we pursue a methodological approach radically different from that of rst. the approach of rst is empirically driven and bottom-up, the proposed set of rhetorical relations is motivated by what can be found in naturally occurring texts, and is potentially open for new additions. this approach has been criticised for its failure to give a principled basis for saying what a theoretically (im)possible coherence relation is (knott and dale, 1994). as a consequence, the resulting list of relations does not provide a basis for making predictions about the coherence of texts. for an incoherent sequence like that in (6-a), we cannot claim that it is incoherent because there is no suitable relation in the list, as nothing prevents us from adding an inform-accident-andmention-fruit relation, as knott and dale (1994) put it. our goal is to ensure that there are no inform-accident-and-mention-fruit relations at the speech act level. in this paper, we address this challenge by taking a strictly top-down approach. we first try to give answers to the following questions: what kinds of goals can speech acts have? what are the possible ways to fail those goals? what are the possible ways to achieve those goals? that provides the basis for our taxonomy of all theoretically possible support relations. as will become clear, this approach does not only reveal systematic similarities between self-repairs, franke’s reinitiations and rst’s presentational relations, but also creates previously unknown categories for speech-act-level coherence relations. one last remark before we embark on this endeavour: presentational relations in rst do not impose any general constraints on the linear order of the nucleus and the satellite, both (n,s) and (s,n) sequences are possible. repair moves can also generally follow the reparandum, or be linearly embedded in it, as in (9) (or even precede it, although in this case the reparandum is rarely fully articulated, cf. forward looking disfluencies in ginzburg et al., 2014). in principle, supporting speech acts can also follow, precede or be linearly embedded in their nuclei. however, this paper will mainly focus on the case of (n,s), where support follows the nucleus.5 3. the goals of a speech act as defined in the previous section, supporting speech acts and support relations are produced in order to help achieve the goal(s) of a speech act that is (anticipated to be) unsuccessful. a goal of an action, and a speech act in particular, is its intended effect—a proposition that describes the way the agent wants the world to become as a result of the action (see zickenheiner, 2020, for a formal implementation of the idea in discourse representation theory). there is a vast theoretical tradition that recognises the central role of speech act goals in the structuring of discourse, most notably grosz and sidner (1986), as well as roberts (1996) and ginzburg (2012) with questions under discussion as a specific conceptualisation of goals. our approach follows the general architecture of grosz and sidner (1986) with respect to relationships between goals (e.g. dominance, or goal–subgoal relationships) and the way they are managed in a stack structure in the flow of communication (again, see zickenheiner, 2020, for formal details). 5. an example of a support relation with a (s,n) order of segments is (25), discussed in section 4.1. 10 speech acts that support other speech acts unlike grosz and sidner, we do not commit to the existence of a single or main goal of a speech act. the way the world should become as a result of an action can be described in many different ways, emphasising different aspects and at different levels of specificity. in principle, all those descriptions can be seen as goal specifications for any given speech act. however, it is important to distinguish what constitutes the goal of a specific speech act, rather than a sequence of multiple speech acts. for instance, in (13) the speaker ultimately wants the hearer (h) to believe that fred ate the beans (f ) and mary ate the eggplant (m), or believe(h, f) ∧ believe(h,m). the first speech act in (13-a) is a step in that direction, but it is not meant to achieve that goal by itself. we would not consider (13-a) to be unsuccessful if the hearer does not also acquire the belief that m in addition to the belief that f . therefore, believe(h, f) ∧ believe(h,m) is not a goal of (13-a), only believe(h, f) is. (13) a. fred ate the beans. b. mary ate the eggplant. the difference is important for distinguishing between grosz and sidner’s dominance and utterances dominated by the same goal on the one hand and our support relations on the other. even though at a certain level of abstraction, (13-a) and (13-b) pursue the same joint goal of believe(h, f) ∧ believe(h,m), they do not pursue that goal each individually. the goals they pursue individually are distinct, and therefore the relation between (13-a) and (13-b) is not support. on the other hand, a whole cascade of goals have to be achieved for an individual speech act to be fully successful. the intermediate steps with their respective goals that we assume to lead to communicative success (inspired by clark’s 1994; 1996, levels of action in communication, going back to allwood et al. 1992) are listed in figure 1. for illustration, consider the speech act off with their heads! performed by the queen of hearts upon seeing that her gardeners had planted roses in the wrong colour in lewis carroll’s alice’s adventures in wonderland: (14) ‘i see!’ said the queen, who had meanwhile been examining the roses. ‘off with their heads!’ and the procession moved on, three of the soldiers remaining behind to execute the unfortunate gardeners, who ran to alice for protection. in spoken communication, the first step is the production of a certain acoustic signal. the goal is to get sound waves of appropriate shape to travel to the right place at the right time (signal transmission). second, the acoustic signal must be processed by the hearer’s auditory system, resulting in a certain pattern of activation in the hearer’s brain and, presumably, a representation in the hearer’s mind (signal processing). the next steps correspond more or less closely to clark’s levels of action in communication.6 after the sounds have been recognised as speech, the hearer must first of all pay attention to the signal (attention), which is a prerequisite for all further processing. while clark treats attention as the first step in the sequence on a par with the subsequent steps, we see it as a state that must hold for the whole duration of communication and in that sense cannot be meaningfully ordered with respect to the other steps. 6. clark emphasises the joint character of communicative action in which both the speaker and the hearer have a role to play. he identifies four levels of action: 1. vocalisation and attention, 2. presentation and identification 3. meaning and understanding and 4. proposal and uptake (the first term in each pair describes the speaker’s part and the second term the hearer’s part in the joint action). in this paper, we think of the hearer’s part as the goal to be achieved by the speaker’s action. 11 jasinskaja and zickenheiner signal transmission signal processing attention, engagement in communication identification understanding uptake execution sequel sound waves travel ear membranes vibrate neurons fire the audience believes that the queen utters “off with their heads” the audience believes that the queen orders execution three soldiers form the intention to execute the gardeners the soldiers execute the gardeners the gardeners’ heads are off figure 1: the workflow of the speech act off with their heads! in (14). provided that the hearer is paying attention and engaging in the communication, they normally need to identify what the speaker said and what the speaker meant. in the next step, the queen’s audience has to map the sounds they hear to the words off, with, their and heads, i.e. create a representation of the linguistic form of the utterance (identification). then, they must come to the belief that the queen produced these words because she wanted to order the execution of the gardeners. that includes recognition of both the semantic content of the utterance and the speaker’s communicative intention (understanding). after the queen’s message was understood, the ultimate success of the speech act will depend on whether or not the hearer reacts in the expected way. we divide the hearer’s reaction into a ‘mental’ part (uptake) and a ‘physical’ or ‘active’ part (execution). the ‘mental’ part of the goal in the present example is that some of the soldiers (three out of ten taking part in the procession) form the intention to execute the gardeners. a few clarifying remarks on our notion of uptake are in order. the term dates back to austin (1962), who used it in a rather broad sense to refer to the whole range of effects a speech act can have upon the hearer, including clark’s identification, understanding, uptake in the narrow sense, and probably our execution as well. clark (1996) draws a line between uptake and understanding: understanding is the correct construal of the speaker’s action, e.g. an order to behead the gardeners is recognised as an order to behead the gardeners. uptake, in turn, is the hearer’s action upon the proposal. however, clark seems to include a number of rather different things in that. 12 speech acts that support other speech acts on the one hand, clark’s notion of uptake includes appropriate reactions in dialogue, i.e. roughly, the second parts of adjacency pairs. for instance, an answer to a question constitutes the hearer’s uptake of the question. to the extent that the notion of adjacency pair is taken to cover cases of non-speech action, the soldiers beheading the gardeners would constitute the uptake of the queen’s order (see also hulstijn and maudet, 2006). on the other hand, clark also counts mere consideration of the speaker’s proposal by the hearer as uptake. in our present example it would mean that uptake already takes place when the soldiers contemplate whether or not they should follow the queen’s order. this weaker notion of uptake has been more widely adopted in subsequent research, to the point of complete replacement of the notion uptake by the notion of consideration (see e.g. rodríguez and schlangen, 2004). in the very narrow sense of uptake that we adopt in this paper, neither counts as uptake. consideration of the speaker’s contribution, in our view, constitutes part of paying attention to the exchange and engagement in communication. as pointed out above, we assume that attention and engagement must be given at all levels of utterance processing, and consideration of the proposal should therefore be seen as engagement in communication at the level of uptake. appropriate reactions in dialogue, such as answering a question or fulfilling an order, on the other hand, are executions of actions that may serve as evidence of the hearer’s uptake, but they do not make up its essence. in this paper, we will use the term ‘uptake’ to refer to the hearer’s mental reaction to an utterance that is in accordance with the speaker’s intentions for that utterance. this notion comes close to what schlöder and fernández (2015) call intention adoption (reaching mutual agreement), which they distinguish from intention recognition (understanding that goes beyond semantic interpretation). along the same lines, we see intention recognition as part of pragmatic understanding, whereas the adoption of the speaker’s intention is the hearer’s mental reaction intended by the speaker. what kind of mental reaction that is will generally depend on the speech act type. uptake of a directive speech act, such as the queen’s order, is the adoption of the intention to fulfil that order. uptake of an assertion is usually the belief of the asserted proposition. uptake of an insult is feeling insulted, and so on. in any case, uptake goes beyond mere understanding of an utterance and includes mental compliance with the speaker’s proposal, but does not go as far as the (physical) action that might result from that mental compliance, such as actually carrying out the order or signalling by means of a nod or a feedback utterance (e.g. yes, mhm) that the hearer accepts the asserted proposition. the latter belong to the level of execution.7 finally, the goal of a speech act may go beyond the hearer’s mental attitudes or physical actions and lie anywhere in the outside world. any direct or indirect consequences of the speech act may constitute its ultimate intended effect. for instance, the state of affairs of the gardeners’ heads being ‘off’ could be seen as the goal of the queen’s order at the level of the sequel. note that the queen does not specify how the gardeners’ heads should get ‘off’, or who should do what to achieve that result. in fact, as we know from the further development of the story, alice ensures that the gardeners’ heads are ‘off’ without being parted from their bodies, and everyone appears happy with that solution. we could say that the literal goal of this speech act is that the gardeners’ heads are 7. admittedly, in some cases it might be difficult to draw a line between uptake and execution, especially in directives that concern mental acts. for instance (i) is literally a request to make the stated assumption, but it is hard to imagine forming an intention to make this assumption without making the assumption yet. it is not theoretically impossible, but something we probably very rarely do with simple mental actions of this kind. (i) suppose x is greater than 0. 13 jasinskaja and zickenheiner ‘off’. this is also the ultimate goal of this speech act. that is, both the literal and the ultimate goal are located at the level of the sequel.8 speech acts of different types expressed by different sentential moods will generally have literal/ultimate goals located at different levels. for instance, a typical directive expressed by an imperative sentence, such as (15), literally expresses the goal that the hearer perform the action of leaving the room. that is, the literal goal of this speech act (and taken at face value, also its ultimate goal) lies at the level of execution. (15) leave the room. one might disagree on whether the goal that the hearer believe that it is raining is literally expressed by (16) or is the result of pragmatic inference. however, that goal clearly lies in the hearer’s mind and therefore at the level of uptake. this is also its ultimate goal if the speech act is taken at face value, but if it is interpreted as an indirect speech act, e.g. as advice to put on appropriate clothing, then the ultimate goal lies at the level of execution. (16) it is raining. in other words, the ultimate goal of a speech act may pertain to the levels of uptake, execution or sequel, but reaching that goal will normally require reaching intermediate goals at the other levels leading up to it, which includes attention, identification and understanding. the linguistic form of the utterance, in particular its sentential mood, gives the hearer a clue of the intended ultimate goal, though it does not necessarily literally express it. things can go wrong at any of these levels. in the next section, we develop a taxonomy of support relations based on the location of the communicative problem as the main criterion for classification. however, first a few words are due on why we choose allwood/clark-style levels of action specifically as basis for our underlying classification of goals. several related concepts that one could use instead or in addition come to mind.9 since the fulfillment of a goal is a criterion for the success of a speech act, there is a certain overlap between our notion of goal and the speech-act-theoretic notion of felicity condition (austin, 1962; searle, 1969); however, the overlap is small. for instance, preparatory conditions of speech acts are not goals, because they need to be given before the speech act is performed, goals are achieved after. besides, classical speech act theory focuses on illocution and has little to say about perlocutionary effects, such as the addressee’s belief in the asserted proposition or emotional state triggered by an expressive speech act, while they are central in our approach. clark’s levels of action are a more useful framework for mapping the entire space of speech act goals because it takes the state of the addressee systematically into account. our notion of intended effect and the extension of clark’s hierarchy based on it is broader than the specific selection of intended effects (positive regard, belief, desire, acceptance of the right to present) that underlies rst’s set of presentational relations, and therefore arguably is more compre8. to put these levels in relation to the standard speech-act-theoretic notions, reaching a goal at the level of identification corresponds roughly to performing the phonetic act in the sense of austin (1962, p. 95). reaching semantic understanding corresponds to the performance of a locutionary act, whereas reaching pragmatic understanding or what schlöder and fernández (2015) call intention recognition corresponds to the performance of an illocutionary act. finally, achieving uptake, execution, or any more far-reaching goals that pertain to the sequel belongs in the domain of perlocution. 9. we thank the anonymous reviewers for bringing up these alternatives. 14 speech acts that support other speech acts hensive. clark’s levels of action have been used as basis for functional classification of clarification requests and self-repairs by rodríguez and schlangen (2004) and by clark (1994) himself. as we will see in the next section, our extended goals ladder provides necessary categories both for selfrepairs and most of rst’s presentational relations, bringing both sets of phenomena into a single conceptual framework. in a taxonomy, there is no such thing as the correct level of granularity. for some applications, it might be enough to distinguish support from other kinds of speech act level relations. in that case, all the different categories presented in the next section are only interesting to the extent that they show what kinds of cases belong to the broad category of support. accordingly, the distinctions between the types of goals underlying those categories would not be relevant for such an application. on the other hand, the categories can be subdivided further if necessary. for instance traum (1994) additionally identifies a level for turn-taking acts (e.g. take turn, keep turn, etc.) and a level for argumentation acts (e.g. summarise, clarify, etc.). the former could be used in an approach like ours to further refine the taxonomy. the latter are relational by their nature. we hope that those or similar categories would follow from our definitions of speech-act-level relations rather than serving as input to them. the reason we believe that the level of granularity provided by clark’s levels of action is at least useful or interesting is because some of the resulting categories roughly match those identified in previous research, such as rst’s evidence, motivation, and enablement relations, as will become clear in the next section. if distinguishing those categories appeared relevant to analysts of naturally occurring discourse, then they might find the underlying categorisation of goals based on clark relevant as well. 4. towards a taxonomy as we promised in section 2, the primary motivation for the taxonomy of support relations to be developed in this paper is theoretical rather than empirical. our goal is to circumscribe all theoretically possible kinds of support relations based on the concept of support as a relation between speech acts. the classification criteria should therefore be based on the understanding of what support is and what a speech act is, on the constitutive parts of these notions and properties that are associated with them by definition. the definition of support relations given in (10) is repeated below. two main classification criteria follow from the two conditions in this definition. one can distinguish between different kinds of support depending on (a) how n failed, and (b) how s solves that problem. (17) support relations: speech act s of speaker a supports speech act n of the same speaker iff a. a believes that n has failed or will fail to achieve its goal, and b. a believes that s will help to achieve that goal. in section 3, we offered our version of the levels of action that expands on and modifies that proposed by clark (1996), cf. figure 1. however, the basic observation concerning these levels remains the same: a speech act can fail at any of the levels that it involves, and therefore a supporting speech act may be called for to repair a problem at any of these levels. so, the first dimension of our taxonomy is the location of the problem targeted by the supporting speech act. section 4.1 gives an overview of the types of support relations according to this criterion. while we will have less to say about the second dimension—the types of support relations according to how they solve the problem at 15 jasinskaja and zickenheiner hand—section 4.2 outlines our general strategy to approach this issue and showcases one group of support types aimed at affecting the hearer’s beliefs. 4.1 locating the problem support for attention / engagement in communication a classical example of a speech act that supports another speech act because the latter failed at the most basic level of securing the addressee’s attention is (18) from clark (1994). here it seems that bob did not even hear ann’s initial utterance (18-a). ann solves the problem by repeating it more loudly (18-c). that is, (18-c) supports (18-a). (18) a. ann: bob b. bob: [3 sec of no response] c. ann: bob [louder] d. bob: what? clark’s original ladder of levels of action creates the impression that attention, as the first step of that ladder, is on a par with the other steps in the sense that success at higher levels presupposes success at the attention level. this, in turn, seems to be based on the implicit assumption that once attention is achieved it cannot be taken away, similar to how understanding, once achieved, cannot be taken away. however, as we pointed out in section 3, while identification, understanding and uptake are events where one is a precondition for the next one, attention is a state that needs to be given throughout the whole process of communication. research on the psychology of attention has found that while attention can be induced as an automatic reaction to abrupt changes in the environment, it is generally given and maintained intentionally and is driven by the agent’s domain-level goals and interests (yantis and jonides, 1984, 1990; theeuwes, 1991; van der lubbe and postma, 2005). thus, the initial ‘grabbing’ of the addressee’s attention that we see in (18) only marks the beginning of the attention state, after which attention needs to be maintained intentionally by the addressee. example (19) shows that attention, or more generally the willingness to engage with the addressee, can be taken away after the utterance is identified and understood: (19-e) makes clear that the first person narrator has perfectly understood the lady’s attempt at contact. nevertheless, he denies her his attention and refuses to engage in the communication. it does not matter whether it is obvious to the speaker that the addressee has understood her utterances. by uttering (19-f) and (19-g) she is trying to (re)gain the addressee’s attention. therefore, (19-f) and (19-g) stand in a support relation to all or any of (19-a), (19-c) and (19-d). (19) a. “sorry,” b. i hear somebody next to me say. c. “aren’t you the man from the television? d. the one who harassed those poor girls?” e. fuck... i don’t answer [...] f. “hey, g. i am talking to you.” h. the lady insists again.10 10. from almost by adriana ls swift. 16 speech acts that support other speech acts support for identification this category of support relations mainly consists of self-repairs that target the hearer’s failure to create the correct representation of the linguistic form of an utterance. in clark’s (1994) example, (20-c) supports (20-a). (20) a. a: yes forty-nine skipton place b. b: forty-one c. a: n i n e . nine d. b: forty-nine, skipton place, support for understanding in section 3, we adopted a broad notion of understanding, which includes both semantic and pragmatic understanding. an utterance counts as fully understood only if the hearer is able to correctly identify the speaker’s meaning behind it (grice, 1957). that includes understanding the conventional meanings of the words and phrases, reference resolution, presupposition resolution, being able to draw the intended implicatures, understanding what kind of speech act the speaker intends to perform by means of the utterance, understanding whether the utterance is meant seriously or jokingly, etc. support relations that target semantic aspects of understanding are otherwise known as selfrepair and reformulation. for instance, the self-repair in (21-c) clarifies the sense in which ken uses the verb evaluate in (21-a). (21) a. ken: k who evaluates the property b. ned: uh whoever you asked, . the surveyor for the building society c. ken: no, i meant who decides what price it’ll go on the market d. ned: ( snorts) . whatever people will pay in (22), another example from clark (1994), sam’s response m in (22-c) to dar’s clarification request supports (22-a) by confirming the reference of this boy: (22) a. sam: well wo uh what shall we do about uh this boy then b. dar: duveen c. sam: m these are instances of more or less spontaneous self-repair. but speech acts whose purpose is to improve the understanding of a previous utterance can also be planned. in (23), the speaker does not only replace the presumably unfamiliar term anacrusis with a more accessible definition, but the reformulation is used as a strategy to establish the equivalence of anacrusis and an unaccented note which is not part of the first full bar (blakemore, 1993), i.e. to define the new term. (23) a. this piece begins with an anacrusis, b. an unaccented note which is not part of the first full bar. mann and thompson’s rst (mann and thompson, 1988, p. 273) includes a presentational relation background, whose definition covers the essence of a support relation that targets a potential problem at the level of understanding (r: reader; n: nucleus; s: satellite): 17 jasinskaja and zickenheiner (24) background: a. r won’t comprehend n sufficiently before reading text of s b. s increases the ability of r to comprehend an element in n however, the cases that this definition is normally applied to are different from those mentioned above. in rst practice, the typical order of discourse units connected by a background relation is (s,n), rather than (n,s). in (25) from mann et al. (1989), the first sentence presents the event of media covering the results of zpg’s 1985 urban stress test, which satisfies the existence presupposition of the definite np this remarkable media coverage in the second sentence. without (25-a) preceding it, (25-b) as it stands would be unacceptable in formally published written discourse, although with slight modifications the reverse order would also be possible, cf. (26). (25) a. the results of zpg’s 1985 urban stress test were reported as a top news story by hundreds of newspapers and tv and radio stations from coast to coast. b. i hope you’ll help us monitor this remarkable media coverage by completing the enclosed reply form. (26) a. i hope you’ll help us monitor the remarkable media coverage of the results of zpg’s 1985 urban stress test by completing the enclosed reply form. b. the results were reported as a top news story by hundreds of newspapers and tv and radio stations from coast to coast. the second sentence in (26) reads like an afterthought that one would like to put in parentheses, and in that sense resembles the more spontaneous self-repairs in (21) and (22), which is probably why backgrounds that follow their nuclei are less acceptable in written texts. however, what is common to (25) and (22) is that the purpose of the supporting speech act is to help establish the reference of a definite description, this remarkable media coverage and this boy, respectively. support relations that target pragmatic aspects of understanding have received less attention in previous research, or might have partly been handled under different unrelated categories. understanding the speaker’s intention behind the utterance involves many different layers, including the understanding of implicatures, illocutionary force, and perlocutionary object, among others. example (27) is an instance of conversational implicature clarification. the first sentence (27-a) has two readings: the ‘normal’ one, without any notable quantity implicature, and the one where the quantity implicature like⇝ not love is drawn in the scope of negation. the second sentence (27-b) makes it clear that the more marked second reading is intended. (27) a. around here, we don’t like coffee, (horn, 1989, p. 382) b. we love it. in (28), an excerpt from the novella the little prince by antoine de saint-exupery, the little prince is in conversation with the king, who he meets on one of the planets he visits on his journey. in (28-d), the king orders the little prince to yawn. just in case the exact illocutionary force of the imperative might be misunderstood, (28-e) supports (28-d) by clarifying that it is an order. 18 speech acts that support other speech acts (28) a. it is years since i have seen anyone yawning. b. yawns, to me, are objects of curiosity. c. come, now! d. yawn again! e. it is an order. a supporting speech act may be called for if the perlocutionary goal of the nucleus needs to be clarified. for instance, (29-a) might be taken as an insult. in (29-b), the speaker makes clear that he does not intend to insult the hearer. (29) a. fuckin’ i hate this guitar! i hate it so much! b. no offence, ed.11 sweetser’s (1990) classical examples of proposition–speech-act causality (30), as well as some instances of the rst relation justify (31) might also belong to this category. (30) a. what are you doing tonight? b. because there’s a good movie on. (31) justify: r’s comprehending s increases r’s readiness to accept w’s right to present n on the face of it, the perlocutionary goal of (30-a) is clear: the addressee should tell the speaker what they are doing tonight, but that might be seen as information the speaker is not entitled to. however, (30-b) explains why the question is relevant. that is, by uttering (30-b) the speaker shows that their underlying goal is to invite the addressee to go to see the film together, something that is meant for the addressee’s benefit as much as their own. after understanding (30-b), the addressee does not even need to give a full answer to the original question. it is enough if they say whether or not they will go to the movies with the speaker as that is actually the ultimate goal. finally, the following example stems from a reference manual for ableton live 5, digital audio workstation software.12 the purpose of user manuals is to give instructions about how to use a product. instruction is a subtype of directive speech act that is appropriate to perform in a situation where the addressee has a certain goal in mind which he or she desires to achieve, but does not know how to do it. therefore, by saying (32-a), the speaker insinuates that the addressee might want, among other things, to get rid of unwelcome house guests or terrifying small pets. in the supporting speech act (32-b), the speaker corrects for a potential misunderstanding, stating that (32-a) was meant as a joke. (32) a. this can be very useful in creating new sounds and textures, as well as getting rid of unwelcome house guests, or terrifying small pets b. (just kidding!). it is important to note that sequences like (28-d)–(28-e), (29-a)–(29-b), (32-a)–(32-b) do not fit comfortably into the existing taxonomies of coherence relations. while they might literally fit the 11. from an interaction between thom yorke and his band during a studio recording: https://www.youtube.com/watch? v=6erp97zrwmk, time stamp 4:45–4:50. last accessed on november 18, 2025. 12. url: http://downloads.ableton.com/manuals/50/ableton_live_5_manual_en.pdf. last accessed on november 18, 2025. 19 jasinskaja and zickenheiner rst definition of background, they do not fit the intuitive notion of background as something that is behind something in some sense (temporally, epistemically, or in the flow of communication) and they all show the non-canonical (n,s) order of segments. none of the pdtb speech act relations seems to apply in these cases. perhaps, a case could be made that these are instances of proposition– speech-act, or metatalk elaboration, but we are not familiar with any proposals that argue that. support for uptake the notion of uptake adopted in section 3 refers to the hearer’s mental reaction that goes beyond mere understanding of the speaker’s intention and encompasses cooperative adoption of that intention. it is the success of the utterance at the perlocutionary level, albeit limited to its mental component. the type of mental reaction in question will depend on the goal. for instance, the perlocutionary goal of an assertion is often to make the speaker believe a proposition. therefore, the uptake of the assertion would consist in the hearer’s adopting that belief. the perlocutionary goal of a directive is to bring the hearer to perform an action. the uptake of the directive would be the hearer’s forming an intention to perform that action. etc. support relations whose purpose is to secure uptake can be distinguished according to the type of mental act or state of the hearer that constitutes the goal of the utterance. if the goal is the addressee’s belief in a certain proposition, then the supporting speech act should make that proposition more believable. the rst relation evidence is geared towards this situation: (33) evidence a. r might not believe n to a degree satisfactory to w b. r’s comprehending s increases r’s belief of n a typical evidence is an assertion that states something more evident, i.e. something more directly observable, than the proposition expressed by the nucleus: (34) a. he must have been here recently. b. there are his footprints. but giving evidence is not the only way to make a proposition more believable. in section 4.2 we will present a number of other support relations that also target the addressee’s belief. if the goal of the utterance is to bring the addressee to perform an action, the mental state that normally precedes that is the addressee’s desire or intention to perform that action. this is typical for directive speech acts. if a directive is anticipated to fail, the speaker might perform a supporting speech act to make the addressee more willing to perform that action. the rst relation motivation (35) covers this case neatly. for instance, (36-c) provides a reason why the google development team might want to fulfil the request expressed in (36-b). (35) motivation a. n presents an action in which r is the actor, unrealized with respect to the context of n b. comprehending s increases r’s desire to perform action presented in n 20 speech acts that support other speech acts (36) a. i have an idea, b. why don’t you guys add the pause and download on the playstore, c. that way it would increase your downloads13 belief and desire/intention are the first kinds of uptake that come to mind as they are directly associated with two kinds of speech acts—assertions and directives.14 speech acts can also pursue the goal of triggering emotions, which may or may not require the intermediate step of imparting a belief. for instance, the use of rude vocabulary can be insulting by itself, regardless of whether the addressee believes the proposition expressed by the utterance, or whether the utterance even expresses a proposition. perlocutionary acts whose purpose is to elicit emotion are altogether less well studied, even less so are the speakers’ discourse strategies when those acts fail. support relations for emotional uptake will have to remain a question for future research. however, it is important to emphasise that they must exist because emotional perlocutionary acts exist, and therefore, a comprehensive taxonomy of support relations must include them. support for execution most speech acts do not require from the addressee the execution of any physical action, but those that do—directive speech acts, first and foremost—can also run into problems at the execution level. in (37), the uptake goes smoothly, speaker b readily agrees to perform the action requested in (37-a). however, in (37-c) speaker a provides information that will make it possible, or easier for b to perform that action. in rst, such support relations are called enablement, (38). (37) a. a: please can you post these letters? b. b: sure. c. a: the stamps are in the drawer. (38) enablement a. n presents an action in which r is the actor, unrealized with respect to the context of n b. r comprehending s increases r’s potential ability to perform the action presented in n concluding remarks we have surveyed different types of support relations according to the location of the communicative problem that the supporting speech act is designed to solve. as it appears, some of them correspond closely to relations previously defined in the literature (evidence, motivation, enablement), while others cross-cut known categories or are difficult to sort under any categories proposed in previous research. 13. a message posted on august 7, 2023 on the google play help forum: https://support.google.com/googleplay/thread/ 229269275/i-have-an-idea-why-don-t-you-guys-add-the-pause-and-download-on-the-playstore-it-will-help-you-guys. last accessed: november 18, 2025. 14. belief can also constitute a perlocutionary goal of a commissive speech act. if commissives were only about regulating the speaker’s own commitments, then what is the point of saying it? just do what you are committed to do! (see gärtner, 2021, for extensive discussion.) by making explicit promises we often try to influence the behaviour of others, for instance, to obtain something in return. that requires that the addressees trust our promises, i.e. believe the propositions they express. 21 jasinskaja and zickenheiner we took a closer look at support relations at the levels of attention, identification, understanding, uptake, and execution, but did not cover signal transmission, signal processing, and sequel in the proposed ladder of levels of action in communication in figure 1. however, support relations should exist at those levels, too. for signal transmission, it is not even that difficult to find natural examples. we are all familiar with the situation in online meetings during the covid pandemic when a participant forgets to switch on their microphone before they start to speak. once they notice their mistake, they usually switch on the microphone and repeat their utterance(s). the combined action of switching on the mic and the repetition stands in a support relation to the initial utterance. this is support for signal transmission.15 we are not going to try to find examples for the last remaining two levels. once again, it follows from the definition of support relations and from the proposed levels of action ladder that such relations must exist. we will either encounter them sooner or later in naturally occurring discourse, or we will have to offer a principled explanation why these levels constitute an exception. both tasks go beyond the scope of this paper. 4.2 solving the problem now we turn to the second criterion for the classification of support relations that follows from the definition in (17)—the way in which the supporting speech act tries to solve the problem caused by the nucleus. it will not be possible to give nearly as broad a survey of possible types as the one we gave in the previous section, but we will present some basic considerations on which, we believe, the taxonomy of possible solutions should be based, and we will develop a fragment of the taxonomy for only one, albeit prominent, subclass of support relations, those whose purpose is to induce a belief. general considerations what means are available for solving a problem will strongly depend on the problem at hand and the nature of the state that the agent desires to achieve. so far we have identified five types of (intermediate or ultimate) goals that directly concern the state of the addressee: (a) attention/engagement 15. note that we have systematically defined the ways in which a speech act can fail in terms of missing a certain effect on the addressee. matej drobňák (p.c.) pointed out that speakers might also perceive their utterances as failing and might feel the need to produce a supporting speech act when they do not comply with a socially defined standard of propriety or correctness, especially in institutional contexts. for instance, a judge can only sentence the accused in a valid and legally binding way in a courtroom. suppose the judge pronounces the sentence while standing in the doorway of the courtroom. they might step back inside the courtroom to pronounce the sentence again to ensure that it has legal consequences. importantly, it does not matter how the audience takes it. they might have completely understood and agreed with the sentence. nevertheless, there is some goal the speech act has obviously failed to achieve if the judge feels the need to take another try at it. another example might be self-corrections for spelling and grammar in instant messaging communication with whatsapp or other applications that allow for later editing of sent messages. again, the addressee might have understood and adopted the intention of the utterance regardless of the incorrect spelling and grammar. nevertheless, the speaker feels the need to repair the utterance because it does not comply with some ideal standard of correctness. we are not sure yet whether these cases could be reduced to failing some less immediate addressee-related goals (e.g. goals in the sequel) or whether they call for a separate category. this is a point where our taxonomy of support relations might still be incomplete. 22 speech acts that support other speech acts in communication; (b) belief; (c) desire/intention; (d) emotion; and (e) action.16 eliciting these reactions or behaviours in another person requires conscious or intuitive domain knowledge about attention, belief, desire and intention, different emotions, and different types of action. for instance, we seem to intuitively know that abruptly increasing the volume or making an abrupt movement is likely to draw the addressee’s attention (theeuwes, 1991; yantis and jonides, 1984, 1990). therefore, if our initial attempt to draw the addressee’s attention did not work, we might want to increase the speech volume, as in example (18) in the previous section, or accompany our speech with a more expressed gesture. in order to keep the addressee engaged for a longer period of time, more sophisticated strategies might be necessary. making people want and intend to perform an action is what persuasion is about. for instance, o’keefe (2006) distinguishes several major strategies of persuasion depending on the addressee’s initial state of mind. one strategy targets the addressee’s positive regard of the action itself and its outcomes for the addressee. as in (36), getting more downloads is presumably something the addressee wants, so the speaker chooses that as an argument to persuade them to fulfil his or her request. other strategies target (a) the influence of the action on the addressee’s perception by others; (b) the addressee’s perceived ability to perform the action when the desire to perform it is a given and the only thing keeping them from turning the desire into a specific intention is the belief that they cannot successfully perform it; and (c) turning a general intention to perform the action into an intention to do it right now. accordingly, one could distinguish further subtypes of the motivation relation depending on which strategy is used by the speaker. if contrary to the speaker’s intention, a speech act fails to insult, frighten, console, or amuse the addressee, one needs specific knowledge in the domain of interpersonal emotional (dis)regulation to be able to produce a supporting speech act that will help induce the desired emotion. some of that knowledge might be intuitively available to the majority of neurotypical population, some may require talent and/or professional training in therapy, advertisement, propaganda, or creative storytelling (thoits, 1996; ochsner and gross, 2005; niven et al., 2009; reddy, 2012; zaki and williams, 2013) to ensure successful execution of an action, one requires domain knowledge about that kind of action. for instance, posting letters requires domain knowledge about posting letters and, in particular, the fact that letters need stamps. that is what makes (37-c) a good supporting speech act for (37-a). in other words, there is no single unified taxonomy of solutions for all kinds of communicative problems. the criteria for more fine-grained classification of supporting speech acts according to the solution strategy they employ will depend on the nature of the domain knowledge required. therefore, it will not be possible to give a comprehensive survey of such criteria and types in this paper. however, to illustrate the basic idea of the approach, we will give a brief overview of strategies for inducing belief and the range of support relations that can be distinguished according to the chosen strategy. 16. this list excludes states that lie outside the addressee’s mind or immediate control, those pertaining to the levels of signal transmission, signal processing and the sequel. for reasons of space, those levels will not be discussed any further in this paper. 23 jasinskaja and zickenheiner support for belief belief plays a role at different levels of a speech act’s workflow, cf. figure 1. the hearer must form the correct belief about the intended linguistic form of the utterance (i.e., identification), she must form the correct belief about the communicative intent of the speech act (i.e., understanding), and so on. in some cases, most typically in assertions, forming a certain belief is the intended ultimate goal of a speech act. since assertions are the most common speech act type in monologues and most previous research on coherence relations has concentrated on assertions, it is useful to look more closely at their typical goal—belief. the most important means of inducing belief is evidence, and several existing frameworks include some version of an evidence relation in their taxonomies (e.g. mann and thompson, 1988; asher and lascarides, 2003). however, since previous research did not consider the full range of possibilities for giving support for a belief, many pragmatic relations conceptually adjacent but not identical to evidence were overlooked. in this section we give several examples of such relations. there are two main approaches to the study of beliefs. the normative approach grounded in philosophy attempts to define the correct, or proper ways of managing our belief states (e.g. feldman, 2000), driven by the goal of holding only true beliefs as well as by other ethical considerations. the empirical approaches in psychology and cognitive science try to answer the question how humans really acquire, store, and change their beliefs (see porot and mandelbaum, 2021, for a recent overview). both trains of research can give us clues as to what speakers might consider as appropriate or effective ways to induce a certain belief in the addressee, and those can be used as basis for classification of discourse patterns of support for belief. evidence-based vs. pragmatic: the philosophical debate has identified two main reasons why an agent may hold a particular belief: evidence-based reasons and pragmatic reasons. according to evidentialism, it is only permissible to believe a proposition if there are a sufficient number of evidence-based reasons for the truth of that proposition. as clifford (1877) put it in his seminal paper on the ethics of belief, “it is always, everywhere, and for everyone, wrong to believe anything on insufficient evidence”. the philosophical counterpoint to this view is that in some cases one might believe a proposition because of the benefits of believing it. for example, one might choose to believe that climate change is not anthropogenic because this belief does not force one to think about the consequences of one’s choices, which is a more comfortable attitude than thinking about one’s own responsibility and possible past mistakes. jordan (1996) argues that sometimes it is morally and rationally permissible to form beliefs on the basis of pragmatic reasons. from a psychological point of view, porot and mandelbaum (2021) point out that belief updating is partly governed by a ‘psychological immune system’ that makes it easier to believe propositions that are consistent with one’s self-image, and to reject beliefs that contradict it. so the psychological immune system helps the development of pragmatic beliefs. we expect this dichotomy to be reflected in the strategies to induce belief employed in supporting speech acts. we have already seen an example of evidence-based belief (34) in section 4.1, repeated in (39). here the speaker wants to induce belief by providing evidence: (39) a. he must have been here recently. b. there are his footprints. 24 speech acts that support other speech acts on the other hand, supporting moves can offer pragmatic reasons for adopting a particular belief. for example, in (40-b) sue urges ann to adopt an evidently false belief because that will help them finish the task they are at before the deadline. (40) a. ann: i’m hungry. b. sue: no, you’re not. c. sue: we need to finish this before seven. note that (40-b)–(40-c) is not an evidence relation, because the fact that sue and ann have urgent work to do has no bearing on the truth of ann being hungry. one might wonder if it could be an instance of rst motivation, as (40-c) attempts to make the belief that ann is not hungry more desirable, cf. the definition in (35). however, (40-b) does not describe a future action of the addressee, so motivation does not fit either. that means that we need a new category to be able to describe this relation, and pragmatic support for belief does the job. providing vs. asking to seek: feldman (2000) points out that our talk about beliefs mirrors our talk about actions and moral judgement. we say that people should do certain things in certain situations, and we blame them when they commit certain actions. on the other hand, we say that, given a certain amount of evidence, people should believe something, and we blame them if they don’t. accordingly, hall and johnson (1998) argue that agents not only have a moral duty to do the right thing, but also an epistemic duty to seek evidence for uncertain propositions—the epistemic duty thesis. applied to a cooperative communicative situation, this thesis implies that communication partners have a duty to seek evidence for the truth of the proposition in question. in the case where one of the interlocutors wants the other to believe something, the joint duty can be fulfilled by providing the missing evidence or by asking the hearer to seek for evidence. the first case is exemplified in (39). the second case can occur, for instance, where the speaker refers the hearer to another, perhaps more trustworthy epistemic authority (zagzebski, 2012), as in (2), repeated in (41). here the speaker urges the hearer to seek evidence from another source. (41) a. “pa told peter he wanted him to be chairman.” b. “sure he did. c. if you don’t believe me, ask peter.” we have already pointed out that this example does not fall under the standard definitions of evidence relations, since (41-c) does not directly provide evidence. this example has some similarity to the sdrt source relation (hunter et al., 2006; hunter, 2016), which relates the content of a message (41-b) to a statement about who said that. however, (41-c) does not directly say that peter said (41-b), maybe he never did before being asked. therefore the source relation does not fit here either. moreover, (43)—another instance of the ’asking to seek’ category that we will discuss presently—is not even similar to source, but is likewise lacking a fitting category in existing classifications. perception vs. testimony vs. inference: epistemologists have argued that we form our beliefs based on sources of evidence that can be broadly categorised as perception (what we directly perceive), testimony (what other people tell us), or inference (what we conclude based on other evidence), as summarised by lesage et al. (2015, see also millar 2011 and davies and matheson 2012). competent speakers are intuitively aware of this difference, as shown by distinct evidential25 jasinskaja and zickenheiner ity marking available in many languages of the world (see e.g. faller, 2002). so it stands to reason that speech acts of support for belief would make recourse to these three sources of evidence and could be categorised along this dimension. the standard evidence relation in (39) involves the speaker telling the addressee the evidence. thus for the addressee, the evidence comes from testimony. similarly, in (41) the speaker asks the addressee to seek further testimony of the proposition expressed by the nucleus. in contrast, the speaker of (42) shows the evidence. by singing the iconic tenor piece from verdi’s opera rigoletto the speaker demonstrates the range of his voice, which the addressee can directly perceive. contrary to what the rst definition of evidence (33) would require, there is no need to comprehend the supporting speech act. in fact, the addressee need not have any knowledge of italian to be convinced by the “argument” in (42-b). (42) a. i’m a tenor. b. [sings:] la donna è mobile qual piuma al vento... in the majority of cases where the target of support is a belief at one of the lower levels in the speech act workflow, cf. figure 1, direct perception will be the most natural way of acquiring evidence by the addressee. for instance, identification of the utterance is essentially the result of processing the auditory and visual input. if the addressee fails to form the belief that the speaker uttered nine as part of (20-a), the most straightforward way to solve the problem is to simply re-enact the same utterance, whose direct perception will help the addressee form the correct belief. showing often requires actions that go beyond speech. for instance, performing a double toe loop is a way to show evidence for the statement i can perform a double toe loop. if the proposition to be shown concerns facts external to the speaker, showing will typically require drawing the addressees attention to events and states that serve as evidence to those facts. pointing gestures or presenting pictures, while not strictly speech acts, are communicative acts that have this function in multimodal communication. rather than providing, the speaker can also ask the addressee to seek direct perceptual evidence. in (43-f), from harry potter and the goblet of fire by j. k. rowling, draco malfoy asks professor snape to look, i.e. to seek perceptual evidence for his assertion that harry potter hit his, malfoy’s, friend goyle in (43-e). and in fact, snape does look, as is made clear in (43-g). so the support relation between (43-f) and (43-e) belongs to the ‘perception’ and ‘asking to seek’ category.17 (43) a. snape pointed a long yellow finger at malfoy b. and said, “explain.” c. “potter attacked me, sir —” d. “we attacked each other at the same time!” harry shouted. e. “— and he hit goyle — f. look —” g. snape examined goyle, whose face now resembled something that would have been at home in a book on poisonous fungi. 17. directive speech acts like (43-f) are often accompanied by pointing gestures, but they are distinct from pointing gestures in our taxonomy. pointing gestures belong to the ‘providing’ category as they facilitate access to perceptual evidence rather than directly tasking the addressee with an action. 26 speech acts that support other speech acts in other words, the distinction between different sources of evidence is orthogonal to the distinction between providing and asking to seek. we should also be able to find instances of support relations that provide and ask to seek inferential evidence, thus amounting to six possible types based on these two features. once again, we would like to stress the fact that (40), (41), (42), and (43-e)–(43-f) do not fit any existing definitions of evidence, or of any coherence relations, to the best of our knowledge. previous studies overlooked such cases probably because, whether by design or not, they focused on common phenomena in the types of discourse they happened to use as their main empirical basis. in contrast, our top-down approach forces us to look at the entire range of possible ways in which a belief can be induced in the addressee. that makes it possible to predict the existence of coherence relations based on strategies that have not been considered before, and once we know what to look for, it turns out that we can also find instances of such relations in naturally occurring discourse, resulting in better empirical coverage. to summarise, we have distinguished between eight possible types of support relations based on the location of the communicative problem within the cascade of goals, or levels of action, of a speech act (cf. figure 1). we have described five of those types in more detail: (a) support for attention, (b) support for identification, (c) support for understanding, (d) support for uptake, and (e) support for execution. we have made a further distinction between five types of goals according to what needs to be achieved, what kind of state or event, giving rise to further types of support relations: (a) support for attention, (b) support for belief, (c) support for desire/intention, (d) support for emotion, and (e) support for action. we have argued that further subdivisions of support relations according to the means the speaker uses to achieve their goal will depend on the type of goal. that is, the means relevant for support for belief and those relevant for support for, let’s say, attention will generally not be the same. finally, we have identified three features that characterise support for belief relations according to the means used to induce belief: (a) evidence-based vs. pragmatic; (b) providing vs. asking to seek; (c) the source of evidence—perception vs. testimony vs. inference. crucially, these are all and only support relations that (should) exist. of course, one can define more features to further subdivide this space, but for the features discussed so far one could already say that if a relation is a support relation, it must belong to one of the proposed categories at each level, otherwise it is not a support relation. for example, the classical evidence relation (39) is support for uptake, support for belief, evidence-based, providing, and based on testimony. the relation in (43-f)–(43-e) is support for uptake, evidence-based, asking to seek, and based on perception. the relation in (20-a)–(20-c) is support for identification, support for belief, providing, and based on perception. finally, to give one more example, the relation in (37-a)–(37-c) is support for execution, support for action, and we have not defined further subtypes based on the solution method for support for action relations. as should have become evident in the meantime, the features this classification is based on are only partly independent. obviously, attention as a stage in the workflow of a speech act and attention as a type of mental state to be achieved inherently go together. desire and emotion as types of goals only seem to be relevant for support for uptake relations. however, belief constitutes the goal at several stages—identification, understanding, and uptake. these asymmetries result from our concept of the levels of action in communication. the fact that there are (probably) no support relations that belong both to the support for understanding and support for desire category is a consequence of the nature of understanding—that the result state of understanding is a belief, 27 jasinskaja and zickenheiner rather than a desire. in other words, the asymmetries are exceptions that prove the rule as they are a logical consequence of our top-down, concept-driven approach. contrary to tradition, we have not given names (e.g. evidence, motivation, enablement) to all possible combinations of the proposed features, and we would indeed need too many to label all the boxes we have opened. if a short label is needed for ease of reference or for the purposes of corpus annotation, one way to go could be to extend the existing labels to broader categories. for instance, one could agree to refer to all evidence-based support for belief relations as evidence. however, one should keep in mind that this is neither rst nor sdrt evidence any more, and that it covers cases like (43-f)–(43-e) and (20-a)–(20-c) that were not thought of as evidence relations previously. users of this taxonomy are welcome to come up with relation names as needed. 5. conclusion and outlook we started out by expressing our general dissatisfaction with previous definitions of speech-actlevel coherence relations which did not specify clearly what it is about speech acts that make some combinations of them coherent while not others. we are now in a position to give a partial answer to that question. in general, coherence between speech acts can be characterised in terms of relations between the goals of those speech acts. there is a limited number of ways in which goals can be coherently related. in this paper, we have investigated one such way—support—where one speech act fails or is anticipated to fail to achieve one of its goals, while the other speech act of the same speaker helps achieve that goal. based on theoretically motivated considerations about levels of action in communication and the nature of different types of goals, we have offered some further subdivisions of the broad category of support relations and have shown how some speech-act-level relations known from previous studies, e.g. self-repairs or rst’s presentational relations, fit into the proposed categories. however, unlike rst, our top-down approach also made it possible to predict what other types of support relations should exist, even if they might not be common in written texts—the type of data that predominantly informed early relational approaches to discourse structure. moreover, for many of the predicted types we were able to find naturally occurring examples, some of which would be really hard to categorise within previously proposed classifications. this is especially true for sequences that include non-assertive speech acts, which have been barely taken into account by approaches to discourse structure based on coherence relations. this is why the proposed taxonomy is particularly relevant for the study of coherence relations in dialogue and will likely turn out useful for the annotation of dialogue corpora. another important advantage of the top-down approach is that it defines the limits of what is theoretically possible. our ultimate goal is to have a comprehensive taxonomy of coherence relations at the speech act level, so that if a sequence of speech acts does not fit one of the categories in that taxonomy, then we should be able to say that the sequence should therefore be incoherent. the taxonomy proposed in this paper is not comprehensive at that level. if a sequence does not fit our definition of a support relation, it only means that it is not a support relation, but it could still be coherent because it features a speech-act-level relation of a different type. however, it is crucial that the range of other possibilities is also theoretically motivated and very limited. apart from support relations, we mentioned adjacency pairs and franke’s (1990) retractions as other ways in which speech acts can be related. yet another possibility, briefly mentioned in section 3, cf. (13), is when two speech acts (usually of the same type, e.g. two assertions) are designed to achieve distinct sub28 speech acts that support other speech acts goals of a bigger communicative goal of the speaker. this case has been studied extensively within the approach to discourse structure based on questions under discussion (roberts, 1996; büring, 2003). as per roberts’ (1996) original intention, questions under discussion are one specific way to operationalise the more general notion of communicative goal. in (13), one could say that the speaker’s goal is to make a ‘big’ assertion in which they inform the hearer about who ate what, but they split it up into two ‘small’ assertions, one to inform the hearer about what fred ate, and the other to inform them about what mary ate. many of the classical semantic or propositional-level coherence relations, e.g. parallel in (13), contrast, and narration, would fall into this category. but then, we believe that that should be all. all coherent sequences of speech acts should be reducible to these few possibilities and combinations thereof, and if they cannot be characterised in these terms, they should be incoherent. it remains a task for future research to review other types of speech-act-level relations along the same methodological guidelines. from the above it should be clear that our taxonomy of support relations, or speech-act-level relations more generally, is not intended to replace existing taxonomies of semantic relations at propositional level. propositional-level taxonomies would still be relevant particularly for sequences of assertions. one could attempt to reformulate them in speech-act-theoretic terms, but it would not necessarily bring new insight. the notions of causality, contiguity, and resemblance have turned out useful to capture generalisations about relations between assertions just by looking at the relations between the communicated propositions. our taxonomy becomes most relevant where propositional level taxonomies hit their limits, that is, especially for heterogeneous speech act sequences, interactive sequences of more than one speaker, and sequences that do not go according to plan. our approach still needs to be tested in an empirical setting. it is true that applying the proposed relation definitions to corpus annotation would require a lot of reasoning with the speakers’ cognitive states, which are accessible to the analyst only to a limited extent. however, in cases where coherence at the speech act level is all we have, i.e. where the more familiar relations at propositional level do not seem to apply, hearers are bound to form at least some hypotheses about the speakers’ communicative intentions to be able to perceive the sequence as coherent. if hearers can do it, then analysts should be able to do it too. how reliable the recognition of support relations is, both for hearers and for corpus annotators, is another open question. finally, it remains to be seen to what extent reasoning about the addressee’s understanding and acceptance of a speech act plays a role in the production of support relations in real time or in the formation of rhetorical patterns in the course of the historical development of specific discourse genres. one of the fundamental ideas underlying our proposal is that even monologue—an uninterrupted sequence of utterances of the same speaker or writer—is in its essence an interactive practice. our notion of support relations substantiates this idea and is a step towards a comprehensive theory of coherence in both dialogue and monologue. acknowledgments the research reported in this paper was funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) – project-id 281511265 – sfb 1252 “prominence in language” in the project c06 at the university of cologne, department of german language und literature i. we would also like to thank the audiences of the workshop on at-issueness, scope, and coherence 29 jasinskaja and zickenheiner (asc) 2019 in cologne, the spagad 2 workshop on speech acts in grammar and discourse 2020 at zas, berlin, and the 24th szklarska poręba workshop on the roots of pragmasemantics, especially matej drobňák, for inspiring discussions of previous versions of this paper. references jens allwood, joakim nivre, and elisabeth ahlsén. on the semantics and pragmatics of linguistic feedback. journal of semantics, 9(1):1–26, 1992. doi: 10.1093/jos/9.1.1. url https://doi.org/ 10.1093/jos/9.1.1. nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. john langshaw austin. how to do things with words. william james lectures. oxford university press, 1962. diane blakemore. the relevance of reformulations. language and literature, 2:101–120, 1993. doi: 10.1177/096394709300200202. url https://doi.org/10.1177/096394709300200202. daniel büring. on d-trees, beans, and b-accents. linguistics and philosophy, 26(5):511–545, 2003. doi: 10.1023/a:1025887707652. url https://doi.org/10.1023/a:1025887707652. herbert h. clark. managing problems in speaking. speech communication, 15(3-4):243–250, 1994. doi: https://doi.org/10.1016/0167-6393(94)90075-2. url https://www.sciencedirect.com/ science/article/pii/0167639394900752. herbert h. clark. using language. cambridge university press, 1996. william k. clifford. the ethics of belief. contemporary review, 1877. jim davies and david matheson. the cognitive importance of testimony. principia: an international journal of epistemology, 16(2):297–318, 2012. doi: 10.5007/1808-1711.2012v16n2p297. url https://doi.org/10.5007/1808-1711.2012v16n2p297. martina faller. remarks on evidential hierarchies. in david i beaver, luis d casillas martínez, brady zack clark, and stefan kaufmann, editors, the construction of meaning, pages 89–111. csli publications, 2002. richard feldman. the ethics of belief. philosophy and phenomenological research, 60(3):667– 695, 2000. doi: 10.2307/2653823. url https://doi.org/10.2307/2653823. anita fetzer. the structuring of discourse. in marina sbisà and ken turner, editors, pragmatics of speech actions, pages 685–711. de gruyter mouton, berlin, boston, 2013. doi: 10.1515/ 9783110214383.685. url https://doi.org/10.1515/9783110214383.685. wilhelm franke. elementare dialogstrukturen: darstellung, analyse, diskussion. germanistische linguistik. niemeyer (de gruyter), 1990. isbn 978-3-484-31101-5. doi: 10.1515/ 9783111631905. url https://doi.org/10.1515/9783111631905. gerd fritz and franz hundsnurscher. sprechaktsequenzen : überlegungen zur vorwurf/rechtfertigungs-interaktion. der deutschunterricht, 27(2):81 – 103, 2009. 30 speech acts that support other speech acts hans-martin gärtner. on the utility of commissive signals and the “promissive gap”. 2021. url http://real.mtak.hu/id/eprint/137527. jonathan ginzburg. the interactive stance: meaning for conversation. oxford university press, 2012. isbn 9780198722991. doi: 10.1093/acprof:oso/9780199697922.001.0001. url https: //doi.org/10.1093/acprof:oso/9780199697922.001.0001. jonathan ginzburg, raquel fernández, and david schlangen. disfluencies as intra-utterance dialogue moves. semantics and pragmatics, 7, 2014. doi: 10.3765/sp.7.9. url https://doi.org/10. 3765/sp.7.9. jonathan ginzburg, zulipiye yusupujiang, chuyuan li, kexin ren, aleksandra kucharska, and pawel lupkowski. characterizing the response space of questions: data and theory. dialogue & discourse, 13(2):79–132, 2022. doi: 10.5210/dad.2022.203. url https://doi.org/10.5210/dad. 2022.203. h. paul grice. meaning. philosophical review, 66(3):377–388, 1957. doi: 10.2307/2182440. url https://doi.org/10.2307/2182440. barbara j. grosz and candace l. sidner. attention, intentions and the structure of discourse. computational linguistics, 12(3):175–204, 1986. richard j. hall and charles r. johnson. the epistemic duty to seek more evidence. american philosophical quarterly, 35:129–139, 1998. url https://www.jstor.org/stable/20009926. götz hindelang. einführung in die sprechakttheorie: sprechakte, äusserungsformen, sprechaktsequenzen. de gruyter, 2010. isbn 9783110231472. laurence r. horn. a natural history of negation. university of chicago press, 1989. joris hulstijn and nicolas maudet. uptake and joint action. cognitive systems research, 7(2):175– 191, 2006. issn 1389-0417. url https://doi.org/10.1016/j.cogsys.2005.11.002. cognition, joint action and collective intentionality. julie hunter. reports in discourse. dialogue & discourse, 7(4):1–35, 2016. doi: 10.5087/dad. 2016.401. url https://doi.org/10.5087/dad.2016.401. julie hunter, nicholas asher, brian reese, and pascal denis. evidentiality and intensionality: two uses of reportative constructions in discourse. in proceedings of the workshop on constraints in discourse 2, maynooth, ireland, 2006. katja jasinskaja and elena karagjosova. rhetorical relations. in daniel gutzmann, lisa matthewson, cécile meier, hotze rullmann, and thomas ede zimmermann, editors, the companion to semantics. wiley, oxford, 2020. doi: 10.1002/9781118788516.sem061. url https: //doi.org/10.1002/9781118788516.sem061. jeff jordan. pragmatic arguments and belief. american philosophical quarterly, 33(4):409–420, 1996. andrew kehler. coherence, reference, and the theory of grammar. csli publications, 2002. 31 jasinskaja and zickenheiner alistair knott and robert dale. using linguistic phenomena to motivate a set of coherence relations. discourse processes, 18:35–62, 1994. doi: 10.1080/01638539409544883. url https://doi.org/ 10.1080/01638539409544883. alex lascarides and nicholas asher. agreement, disputes and commitments in dialogue. journal of semantics, 26(2):109–158, 2009. url https://doi.org/10.1093/jos/ffn013. claire lesage, nalini ramlakhan, ida toivonen, and chris wildman. the reliability of testimony and perception: connecting epistemology and linguistic evidentiality. in cogsci, volume 37, pages 1302–1307, 2015. url https://www.researchgate.net/publication/275582507_ the_reliability_of_testimony_and_perception_connecting_epistemology_and_linguistic_ evidentiality. willem j.m. levelt. monitoring and self-repair in speech. cognition, 14(1):41–104, 1983. issn 0010-0277. doi: 10.1016/0010-0277(83)90026-4. url https://doi.org/10.1016/0010-0277(83) 90026-4. william c. mann and sandra thompson. rhetorical structure theory: toward a functional theory of text organization. text, 8(3):243–281, 1988. doi: 10.1515/text.1.1988.8.3.243. url https: //doi.org/10.1515/text.1.1988.8.3.243. william c. mann, christian m. i. m. matthiessen, and sandra a. thompson. rhetorical structure theory and text analysis. technical report, university of southern california, information sciences institute, 1989. url https://doi.org/10.1075/pbns.16.04man. alan millar. how visual perception yields reasons for belief. philosophical issues, 21:332–351, 2011. doi: 10.1111/j.1533-6077.2011.00207.x. url https://doi.org/10.1111/j.1533-6077.2011. 00207.x. karen niven, peter totterdell, and david holman. a classification of controlled interpersonal affect regulation strategies. emotion, 9(4):498–509, 2009. doi: 10.1037/a0015962. url https://doi. org/10.1037/a0015962. kevin n. ochsner and james j. gross. the cognitive control of emotion. trends in cognitive sciences, 9(5):242–249, 2005. issn 1364-6613. doi: 10.1016/j.tics.2005.03.010. url https: //doi.org/10.1016/j.tics.2005.03.010. daniel j. o’keefe. persuasion. in the handbook of communication skills. 3d edition, pages 323– 341. routledge, 2006. isbn 9780761925392. nicolas porot and eric mandelbaum. the science of belief: a progress report. wiley interdisciplinary reviews: cognitive science, 12, 2021. doi: 10.1002/wcs.1539. vasudevi reddy. moving others matters. in moving ourselves, moving others, consciousness and emotion book series, pages 139–162. philadelphia, 2012. isbn 9789027241566. gisela redeker. ideational and pragmatic markers of discourse structure. journal of pragmatics, 14(3):367–381, 1990. doi: 10.1016/0378-2166(90)90095-u. url https://doi.org/10.1016/ 0378-2166(90)90095-u. 32 speech acts that support other speech acts craige roberts. information structure in discourse: towards an integrated formal theory of pragmatics. osu working papers in linguistics, 49:91–136, 1996. kepa joseba rodríguez and david schlangen. form, intonation and function of clarification requests in german task oriented spoken dialogues. in proceedings of catalog ’04 (the 8th workshop on the semantics and pragmatics of dialogue, semdial04), pages 101–108, 2004. harvey sacks and emanuel a. schegloff. opening up closings. semiotica, 8(4):289–327, 1973. doi: 10.1515/semi.1973.8.4.289. url https://doi.org/10.1515/semi.1973.8.4.289. ted sanders. semantic and pragmatic sources of coherence: on the categorization of coherence relations in context. discourse processes, 24:119–147, 1997. doi: 10.1080/01638539709545009. url https://doi.org/10.1080/01638539709545009. ted sanders, wilbert spooren, and leo noordman. toward a taxonomy of coherence relations. discourse processes, 15(1):1–35, 1992. doi: 10.1080/01638539209544800. url https://doi. org/10.1080/01638539209544800. ted sanders, josé sanders, and eve sweetser. causality, cognition and communication: a mental space analysis of subjectivity in causal connectives. in ted sanders and eve sweetser, editors, causal categories in discourse and cognition, pages 19–59. de gruyter mouton, berlin, new york, 2009. isbn 9783110224429. doi: doi:10.1515/9783110224429.19. url https://doi.org/ 10.1515/9783110224429.19. marina sbisà. cognition and narrativity in speech act sequences. in anita fetzer and c. meierkord, editors, rethinking sequentiality, pages 71–97. john benjamins publishing company, amsterdam, 01 2002. doi: 10.1075/pbns.103.04sbi. url https://doi.org/10.1075/pbns.103.04sbi. julian j. schlöder and raquel fernández. clarifying intentions in dialogue: a corpus study. in matthew purver, mehrnoosh sadrzadeh, and matthew stone, editors, proceedings of the 11th international conference on computational semantics, pages 46–51, london, uk, 2015. association for computational linguistics. url https://aclanthology.org/w15-0106. john r. searle. speech acts: an essay in the philosophy of language. cambridge university press, 1969. doi: 10.1017/cbo9781139173438. url https://doi.org/10.1017/cbo9781139173438. john r. searle. intentionality: an essay in the philosophy of mind. cambridge paperback library. cambridge university press, 1983. doi: 10.1017/cbo9781139173452. url https://doi.org/10. 1017/cbo9781139173452. eve sweetser. from etymology to pragmatics: metaphorical and cultural aspects of semantic structure. cambridge university press, 1990. doi: 10.1017/cbo9780511620904. url https: //doi.org/10.1017/cbo9780511620904. maite taboada. rhetorical relations in dialogue. in discourse across languages and cultures, studies in language companion series, pages 75–97. john benjamins publishing company, 2004. isbn 9789027295262. url https://www.sfu.ca/~mtaboada/docs/publications/taboada_ rst_2004.pdf. 33 jasinskaja and zickenheiner jan theeuwes. exogenous and endogenous control of attention: the effect of visual onsets and offsets. perception & psychophysics, 49(1):83–90, 1991. doi: 10.3758/bf03211619. url https://doi.org/10.3758/bf03211619. peggy a. thoits. managing the emotions of others. symbolic interaction, 19(2):85–109, 1996. url https://doi.org/10.1525/si.1996.19.2.85. david r. traum. a computational theory of grounding in natural language conversation. phd thesis, university of rochester, 1994. rob h. j. van der lubbe and albert postma. interruption from irrelevant auditory and visual onsets even when attention is in a focused state. experimental brain research, 164(4):464–471, 2005. doi: 10.1007/s00221-005-2267-0. url https://pubmed.ncbi.nlm.nih.gov/15785951/. bonnie webber, rashmi prasad, alan lee, and aravind joshi. the penn discourse treebank 3.0 annotation manual, 2019. url https://catalog.ldc.upenn.edu/docs/ldc2019t05/ pdtb3-annotation-manual.pdf. steven yantis and john jonides. abrupt visual onsets and selective attention: evidence from visual search. journal of experimental psychology: human perception and performance, 10(5):601– 621, 1984. doi: 10.1037/0096-1523.10.5.601. url https://doi.org/10.1037/0096-1523.10.5.601. steven yantis and john jonides. abrupt visual onsets and selective attention: voluntary versus automatic allocation. journal of experimental psychology: human perception and performance, 16 (1):121–134, 1990. doi: 10.1037/0096-1523.16.1.121. url https://doi.org/10.1037/0096-1523. 16.1.121. linda trinkaus zagzebski. epistemic authority: a theory of trust, authority, and autonomy in belief. oxford university press, 2012. doi: 10.1093/acprof:oso/9780199936472.001.0001. url https://doi.org/10.1093/acprof:oso/9780199936472.001.0001. jamil zaki and w. craig williams. interpersonal emotion regulation. emotion, 13(5):803–810, 2013. doi: 10.1037/a0033839. url https://doi.org/10.1037/a0033839. frank zickenheiner. goals in discourse: from actions to rhetorical relations. phd thesis, university of cologne, germany, 2020. url http://kups.ub.uni-koeln.de/id/eprint/64204. 34 dialogue and discourse 4(2) (2013) i-x doi: 10.5087/dad.2013.200 beyond semantics: the challenges of annotating pragmatic and discourse phenomena introduction to the special issue stefanie dipper dipper@linguistics.rub.de institute of linguistics ruhr-university bochum heike zinsmeister heike.zinsmeister@uni-hamburg.de institute of german language and literature university of hamburg bonnie webber bonnie@inf.ed.ac.uk institute for language, cognition and computation university of edinburgh 1. introduction manual corpus annotation and automated corpus analysis began with a focus on what was seen as “low-hanging fruit”, such as morphological tags, part-of-speech tags, and syntactic structures. after investment of considerable resources (time, money and effort), development of fairly accurate automated tools, and recognition that other slightly higher-hanging fruit were there for the taking, attention shifted—this time, to characterizing the meaning and intention of language use, including the development of semantic labelling schemes such as word sense tagging (e.g. miller et al., 1993), semantic class labelling such as named-entity typing (muc-6, 1995; brunstein, 2002; aceedt, 2004), semantic role labelling (palmer et al., 2005; fillmore et al., 2004; meyers et al., 2004), coreference linking (passonneau, 1996; hirschman and chinchor, 1998; poesio, 2000) and temporal relations (pustejovsky et al., 2003a,b) and the development of pragmatic labelling schemes such as dialogue act tagging (allen and core, 1997; carletta et al., 1997; jurafsky et al., 1997; alexandersson et al., 1998). these too have seen considerable resources invested in them, as well as the development of fairly accurate automated tools. but horizons change, and attention is now focussed somewhat higher—on problems in understanding and reliably annotating other sorts of semantic and pragmatic phenomena, especially phenomena that involve spans of text larger than a single sentence or clause. the goal of this special issue is to show the challenges faced in reliably annotating abstract semantic and pragmatic information at both the sentence and discourse levels, and how those challenges are being met. such information is frequently not explicitly or unambiguously marked in natural language. it is usually dependent on contextual information, and annotators often have to reconstruct complex relations and situations from the context. annotated data can serve both as the basis of linguistic investigations and as training data for applications developed in the field of natural language processing. most of the papers in this issue c©2013 stefanie dipper, heike zinsmeister and bonnie webber dipper, zinsmeister and webber deal with resources that are intended to ultimately serve automatic applications. still, they have a strong focus on the linguistic foundations underlying the annotations. 2. the challenges linguistically motivated semantic and pragmatic theories are too often based on toy examples, wellcontrolled contexts, and the intuitions of theory’s author. as a consequence, it is often hard to transfer results from theoretical linguistics to naturally-occurring texts. annotation, either by trained experts or untrained workers, is felt to be a way of dealing with semantic and pragmatic phenomena in naturally-occurring text. but as noted, annotating such phenomena often requires annotators to infer from context, complex relations and situations. pertinent examples from the papers in this issue are: • constructing (lowor high-level) questions under discussion to determine the focus of a sentence or the discourse topic (see the articles by riester & baumann, and by versley & gastel) • judging whether a sentence makes broad statements about a topic or provides details (see the article by louis & nenkova) • judging whether some sentence serves to introduce a new entity or to present a situation as a whole (see the article by cook & bildhauer) • disambiguating discourse relations that are either triggered by connectives which are lexically ambiguous, or that are not lexically expressed at all (see the articles by versley & gastel, by cartoni, zufferey & meyer, and by hardt) • deciding whether two sentences are coherently related when there is no overt connective or other surface clue, and coherence is achieved only by inference based on word knowledge (see the article by burstein, tetreault & chodorow) • deciding what speakers “do” with language as they interact across varied communication situations (see the articles by tenbrink, eberhard, shi, kübler & scheutz, and by morgan, oxley, bender, zhu, gracheva & zachry) • discriminating between different lexically triggered inference relations such as presupposition and logical entailment (see the article by tremper & frank) even with explicit annotation guidelines, discourse-related and pragmatic phenomena are often difficult to annotate reliably. for instance, earlier studies (ritz et al., 2008) showed that annotating information-structural features in german often results in inter-annotator agreement scores well below κ = 0.6 (cohen, 1960)—such scores are often assumed to allow for tentative conclusions only (landis and koch, 1977). similarly, the overview on inter-annotator agreement measures by artstein and poesio (2008) showed that annotation of discourse-related features, such as dialogue act tagging, discourse segmentation, or word sense tagging, also achieves low κ scores in many studies. a possible approach to tackling these problems is the use of proxies for more abstract linguistic concepts in terms of surface clues that are more reliably classified by annotators than the original concepts. a prominent example is the practice applied in the penn discourse treebank (prasad ii beyond semantics: introduction et al., 2008), where annotators are asked to generate overt connectives for otherwise unmarked discourse relations. another approach is to use paraphrase tests to elicit interpretations in a way as objective as possible (cf. zhou and xue, 2012). pertinent examples from the current papers are: • inserting overt connectives to determine discourse relations (see the article by versley & gastel) • approximating meaning of ambiguous connectives by means of their translation equivalents (see the article by cartoni, zufferey & meyer) • transforming a sentence into concerning x, . . . to test whether x is an aboutness topic (see the article by cook & bildhauer) • paying attention to morphological endings as indicators of nominal clauses that have predicative potential and can function as arguments of discourse relations (see the article by zeyrek, demirşahin, sevdik callı & cakıcı) • guiding the annotators by a decision-tree like question-based annotation design to elicit complex semantico-pragmatic judgements in a reliable way (see the article by tremper & frank) • taking the reverse route by explicitly annotating the entities that appear to signal coherence in a corpus that already contains coherence annotation (see the article by taboada & das). 3. the case for a special issue the general idea of the special issue was to gather research that reports on the generation (and exploitation) of corpora that are annotated with pragmatic or discourse-related information grounded in linguistic theory. the volume aims at enhancing mutual awareness of people working on different kinds of abstract semantic and pragmatic phenomena from different perspectives. we would like to bring together theoretical linguists who use texts and corpora for pragmatic or discourse-related research questions, and corpus linguists as well as computational linguists who create and annotate relevant corpus resources, or exploit them. the goal of the special issue is to enhance exchange and awareness between researchers of both fields, and to gain insights in the—possibly common—properties and peculiarities of these abstract semantic and pragmatic phenomena. ideally, the volume would allow people to realize that there are problems they share—despite the fact that they are working on quite different tasks—and to recognize (partial) solutions that they too might be able to adopt. we also see it as an important desideratum to promote the application of linguistic theories to naturally-occurring texts. this would enhance the search for operationalizations of theoretical concepts, which can probably then be annotated with higher reliability. it would open up corpusbased development and validation of theoretical hypotheses. at the same time, operationalized theoretical concepts and reliable annotations would facilitate the use of pragmatic and discourserelated knowledge in computational linguistics. this means, on one hand, that we need more theoretical linguists annotating corpora and validating their theories based on corpora, and, on the other hand, more computational linguists drawing from linguistic insights to a greater extent when annotating training data. we hope that this special issue promotes this exchange considerably, by highlighting relevant research. iii dipper, zinsmeister and webber 4. overview of the papers the papers collected in this special volume address topics from different fields: discourse relations and coherence, dialogue acts, text specificity and communicative goals, inference-triggering relations, and information structure. 4.1 discourse relations and coherence the paper by yannick versley & anna gastel: “linguistic tests for discourse relations in the tüba-d/z corpus of written german”, deals with the annotation of discourse relations in a german corpus. annotators first segment the texts in topic segments, which answer a high-level “question under discussion”; annotators are also asked to make the topic explicit. discourse relations usually occur within the boundaries of a topic segment. exceptions are cross-topic relations, which are introduced by the authors to attenuate the effects of topic shifts. relations can be subordinating or coordinating, and fall into five groups: contingency, expansion, temporal, comparison, reporting. they are assigned by means of linguistic tests, such as substitution or insertion of connectives, use of paraphrases or nominalizations, or explicit insertion of the questions under discussion. the article by bruno cartoni, sandrine zufferey & thomas meyer: “annotating the meaning of discourse connectives by looking at their translation: the translation spotting technique” focuses on disambiguating the senses of discourse connectives without annotating their arguments. the paper introduces a method of sense disambiguation based on projecting meaning cross-linguistically in a parallel corpus. they point out that connectives like while are ambiguous and introduce different discourse relations depending on the context (and sometimes even simultaneously). while, for example, has among others a concessive and a contrastive reading, in addition to its temporal reading. the authors highlight that the distinctions of some relations result in very low inter-annotator agreement. the method they suggest instead of sense disambiguation requires the annotators to identify the translation equivalent of the connective in the aligned sentence (particle, paraphrase, or no translation). to cluster monolingual equivalent classes of connectives, the authors conducted a cloze/fill-the-blank test in which annotators had to choose a connective from a list of connectives to fill the blank in a sentence. the authors trained a maximum entropy classifier to determine the senses of the connective while automatically (six senses). finally the authors discuss that the cross-linguistic projection method helps to identify sub-senses in connectives not explicitly distinguished in the source language. deniz zeyrek, işın demirşahin, ayışığı sevdik callı & ruket cakıcı report in their paper “turkish discourse bank: porting a discourse annotation style to a morphologically rich language” on the annotation of explicit discourse connectives in the turkish discourse bank. they detail the annotation cycle and provide an overview of the finalized annotation scheme that has been strongly influenced by the annotation scheme of the english penn discourse treebank (pdtb). extensions and modifications implemented in their scheme concern the tag for shared discourse arguments, and the modifier and the supplementary tags of the pdtb scheme. nominalizations are also annotated as discourse arguments, given their frequency of occurrence in turkish. the evaluation of inter-annotation agreement shows that the task is more feasible for certain discourse relations and more subjective for others. the authors also provide statistical information about the connectives that proved challenging in the annotation process, namely discourse adverbials, subordinators and their polymorphous occurrences. iv beyond semantics: introduction the paper by daniel hardt: “a uniform syntax and discourse structure: the copenhagen dependency treebanks”, introduces the annotation of the danish copenhagen dependency treebank (cdt), which belongs to a word-aligned multi-lingual and multi-level annotated treebank comprising danish texts and their translations into english, german, italian, and spanish subcorpora. even if the cdt discourse annotation is heavily inspired by the english penn discourse treebank (pdtb) annotation guidelines, the annotation procedure of the cdt differs from the pdtb by annotating on top of syntactic (dependency) structure rather than raw text strings. in this paper, the author argues from a semantic perspective for doing so, by discussing cases of contrastive, causal, and conjunctive relations. maite taboada & debopam das in their article “annotation upon annotation: adding signalling information to a corpus” take the reverse way: they start from an already existing discourse treebank, the rst discourse treebank (carlson et al., 2003), and mark linguistic expressions that could serve as indicators of the annotated relations—a case of reverse engineering. in their pilot study, they annotate various kinds of linguistic expressions that can signal discourse relations. the signals include discourse markers, relations between entities (e.g. coreference), semantic relations (e.g. antonyms, hypernyms), etc. their analysis finds 14% of the relations in the rst discourse treebank to lack an overt signal and only 22% of the overtly signalled relations to be signalled by a discourse marker. many relations that are signalled by other means are redundantly marked, i.e. by several signals. jill burstein, joel tetreault & martin chodorow deal in their contribution “holistic annotation of discourse coherence quality in noisy essay writing” with the annotation of discourse coherence in terms of scoring student essays on text-level with a two-level scale (low coherence, high coherence). they started out with a three-level annotation scale that distinguished high coherence proper from essentially coherent in the sense of that “text meaning can basically be constructed, but one or two identifiable points were confusing”—the annotators were also asked to mark these points of coherence breakdown. since inter-annotator agreement for the middle score was low, the authors decided to combine it with the high coherence score in the end. these two-level scores for discourse coherence correlated well with task-independent expert essay ratings on a 5or 6-level scale. finally, the authors used the annotated data to train a decision-tree classifier to distinguish low and high coherence essays in a corpus of 1,555 essays. the best performing system out-performed base-line systems in all but one of ten subcorpora. 4.2 dialogue acts the paper by jonathan morgan, meghan oxley, emily bender, liyi zhu, varya gracheva & mark zachry: “are we there yet?: the development of a corpus annotated for social acts in multilingual online discourse” is concerned with multi-party discourse. the authors created two multi-lingual corpora (english, mandarin, russian) of computer-mediated communication in which they annotated dialogue acts in terms of “social acts”. in particular, they annotated authority claims such as contributors giving reference to their “education, training, or a history of work in an area” and positive/negative interpersonal alignment moves such as explicit agreement when taking the turn by stating “exactly.”. in cases like this, it turned out to be difficult to identify sarcasm reliably. after outlining the corpus sampling from editors’ discussions of wikipedia articles and from (written) chat discussions, the authors describe their iterative annotation process in detail. among others, they had been monitoring annotation quality by performing a longitudinal inter-annotator v dipper, zinsmeister and webber agreement study. finally, the authors discuss some analyses based on the annotated corpora, which identify interactions among social acts, and between participant status and social acts in the data. the article by thora tenbrink, kathleen eberhard, hui shi, sandra kübler & matthias scheutz: “annotation of negotiation processes in joint action dialogues” discuss another type of dialogue act annotation in terms of annotating activity coordination and goal negotiation processes, as well as belief states of the participants in dialogues that are conducted in course of a joint action. the prototypical example of such a joint action dialogue is the edinburgh map task corpus (anderson et al., 1991). one important characteristic of such dialogues is that they are grounded in terms of visual and other non-linguistic cues that also influence the dialogue structure. tenbrink et al. present a kind of reference paper for the annotation of this type of multi-modal dialogue. they give a comprehensive overview of existing schemes and introduce relevant corpora and their annotation schemes, including their own corpora. finally, the authors point to four layers of annotation which are not always captured in existing schemes but which they argue are essential to capture the characteristic properties of such dialogues: intonation, gestures, perception of the talk domain, and task-relevant actions. 4.3 text specificity and communicative goals the article by annie louis & ani nenkova: “a corpus of science journalism for analyzing writing quality”, introduces a new type of corpus consisting of science articles of different writing qualities from the new york times. “great” articles are those that have been chosen for the “best american science writing” annual anthologies. “very good” articles are further articles written by the top authors, and “typical” articles are articles from other authors. the corpus is divided in subcorpora of topically-related articles. the authors hypothesize that text generality/specificity and communicative goals can help distinguishing between top and average writing. in a crowd annotation, five turkers per sentence were asked to mark isolated sentences as “general” or “specific”. 64% of the sentences were assigned the same class by at least four judges. specific sentences predominate, and tend to occur in blocks. the authors then trained a classifier for specificity, based on lexical and non-lexical features, which outperformed the baseline considerably. it turns out that confidence from the classifier is correlated with agreement among the annotators. next, a classifier for distinguishing “great” and “very good” articles was trained. finally, the authors investigate the communicative goals of individual sentences, and exploit syntactic similarity for automatically identifying such goals, and for judging writing quality from them. 4.4 inference-triggering relations the paper galina tremper & anette frank: “a discriminative analysis of fine-grained semantic relations including presupposition: annotation and classification” presents a corpusbased induction study of semantic relations between verbs. the relations under investigation go beyond lexical semantic relations like synonymy and hyperonymy established, for example, in wordnet. the focus of this paper is on relations between verbs that trigger inferences, in particular presupposition, logical entailment, temporal inclusion, antonymy, and synonymy. the authors outline a discriminative analysis of these semantic relations, which draws on the negation test that is well established in semantic and pragmatic literature. in this paper, the authors concentrate on type-level discrimination of the relations, which will be the basis for automatically deriving implicit meaning from text in future work. the authors discuss the guidelines for manual annotation and present the vi beyond semantics: introduction results of manually annotating their gold standard of semantic relations. in addition, they also report results for their automatic classification experiments, which outperfom their baseline by a large margin. 4.5 information structure the paper by philippa cook & felix bildhauer: “identifying ‘aboutness topics’: two annotation experiments” presents two experiments on annotating aboutness topics in naturally-occurring data from german. the first experiment uses götze at al’s (2007) guidelines, and involves annotating sentences with the main verbs geraten, reagieren, profitieren, herrschen ‘get caught, react, profit, reign’. motivated by the (rather poor) results of inter-annotator agreement, a second annotation experiment is carried out using refined guidelines. in contrast to götze et al’s guidelines, they distinguish between entity-central (presentational) and event-central (event-reporting) thetic (i.e. all-rhematic) expressions. the new guidelines result in better (but still low) agreement. the authors hypothesize that high agreement on a sentence indicates a prototypical case, instantiating all typical properties of aboutness-topics. to achieve higher inter-annotator agreement, the authors argue for language-specific rough-and-ready distinctions in the guidelines. the paper by arndt riester & stefan baumann: “focus triggers and focus types from a corpus perspective” presents an integrated analysis of novelty (new information) focus and contrastive focus: both serve to answer (implicit or explicit) questions. with contrastive focus, the hearer is able to name (at least) one contrastive alternative; with novelty focus, the alternatives remain anonymous. the authors present the reflex annotation scheme for information status, marking the given–new distinction at lexical and referential levels. for instance, in a sequence like a man came in. the man coughed., the second occurrence of the word man is lexically given because it is a repetition, and the phrase the man is referentially given because it is coreferent with the phrase a man. novelty focus (information focus) can be identified based on lexical and referential givenness. identifying contrastive focus requires information about contrastive alternatives. the authors propose an annotation scheme for alternative-eliciting features, marking cases involving focus-sensitive particles, overtly contrastive expressions, comparative constructions, etc. finally, the authors address the issue of secondary foci in general, and second occurrence focus in particular. acknowledgments this special issue grew out of the workshop “beyond semantics” at the annual conference of the german linguistic society in february 2011, organized by two of the editors (dipper and zinsmeister, 2011). we received 32 abstracts in response to our open call for statements of intent; seven of them were based on papers presented at the workshop. among these, we invited sixteen to submit full regular papers, and seven to submit short notes. we finally received sixteen submissions. after the reviewing process, three papers were rejected, one withdrawn, which left us with twelve papers (ten regular papers and two short notes). in general, papers were reviewed by three reviewers. we are very grateful to the following reviewers: john bateman, kristy boyer, özlem çetinoğlu, jennifer chu-carroll, micha elsner, david elson, cathrine fabricius-hansen, raquel fernández, dilek hakkani-tür, klaus von heusinger, nancy hedberg, graeme hirst, hans kamp, ruth kempson, ralf klabunde, anke holler, shinichiro ishihara, alistair knott, valia kordoni, alan lee, katja markert, roland meyer, marie-francine moens, malvina nissim, rainer osswald, lilja øvrelid, vii dipper, zinsmeister and webber alexis palmer, paul portner, christopher potts, gisela redeker, craige roberts, josef ruppenhofer, ted sanders, kiril simov, wilbert spooren, caroline sporleder, manfred stede, amanda stent, angelika storrer, elke teich, sara tonelli, carla umbach, yannick versley, thomas weskott, janyce wiebe, nianwen xue. finally, we would like to thank the managing editors of dialogue & discourse, in particular, jonathan ginzburg (editor-in-chief) and raquel fernández for their support throughout the preparation of this issue. references ace-edt. annotation guidelines for entity detection and tracking (edt). linguistic data consortium, 2004. url http://catalog.ldc.upenn.edu/docs/ldc2005t09/ guidelines/englishedtv4-2-6.pdf. version 4.2.6 200400401. jan alexandersson, bianka buschbeck-wolf, tsutomu fujinami, michael kipp, stephan koch, elisabeth maier, norbert reithinger, birte schmitz, and melanie siegel. dialogue acts in verbmobil-2 (second edition). verbmobil report 226, saarland university, saarbrücken, germany, 1998. url http://www.coli.uni-sb.de/publikationen/ softcopies/alexandersson:1998:dav.pdf. james allen and mark core. draft of damsl: dialogue act markup in several layers. technical report, university of rochester, 1997. url http://www.cs.rochester.edu/ research/speech/damsl/revisedmanual/. anne h. anderson, miles bader, ellen gurman bard, elizabeth boyle, gwyneth doherty, simon garrod, stephen isard, jaqueline c. kowtko, jan mcallister, jim miller, cathy sotillo, henry thompson, and regina weinert. the hcrc map task corpus. language and speech, 34(4): 351–366, 1991. ron artstein and massimo poesio. inter-coder agreement for computational linguistics (survey article). computational linguistics, 34(4):555–596, 2008. ada brunstein. annotation guidelines for answer types. technical report, bbn technologies, 2002. url http://catalog.ldc.upenn.edu/docs/ldc2005t33/bbn-typessubtypes.html. jean carletta, amy isard, stephen isard, jacqueline c. kowtko, gwyneth doherty-sneddon, and anne h. anderson. the reliability of a dialogue structure coding scheme. computational linguistics, 23(1):13–31, 1997. lynn carlson, daniel marcu, and mary ellen okurowski. building a discourse-tagged corpus in the framework of rhetorical structure theory. in jan van kuppevelt and ronnie smith, editors, current and new directions in discourse and dialogue, pages 85–112. kluwer academic publishers, 2003. jacob cohen. a coefficient of agreement for nominal scales. educational and psychological measurement, 20(1):37–46, 1960. viii beyond semantics: introduction stefanie dipper and heike zinsmeister, editors. beyond semantics: corpus-based investigations of pragmatic and discourse phenomena. proceedings of the dgfs workshop, göttingen, volume 3 of bla (bochumer linguistische arbeiten). institute of linguistics; ruhr-university bochum, 2011. url http://www.linguistics.ruhr-uni-bochum.de/bla/. charles fillmore, josef ruppenhofer, and collin f. baker. framenet and representing the link between semantic and syntactic relations. in chu-ren huang and winfried lenders, editors, computational linguistics and beyond, language and linguistics monographs series b. frontiers in linguistics i, pages 19–62. institute of linguistics, academia sinica, taipei, 2004. michael götze, thomas weskott, cornelia endriss, ines fiedler, stefan hinterwimmer, svetlana petrova, anne schwarz, stavros skopeteas, and ruben stoel. information structure. in stefanie dipper, michael götze, and stavros skopeteas, editors, information structure in cross-linguistic corpora, number 7 in interdisciplinary studies on information structure (isis), pages 147–187. universitätsverlag potsdam, 2007. lynette hirschman and nancy chinchor. muc-7 coreference task definition—version 3.0. in proceedings of the 7th message understanding conference (muc-7), fairfax, virginia, 1998. dan jurafsky, liz shriberg, and debra biasca. switchboard swbd-damsl shallow-discoursefunction annotation. coders manual, draft 13. technical report rt 97-02, university of colorado at boulder & sri international, 1997. url http://www.stanford.edu/ ˜jurafsky/ws97/manual.august1.html. j. richard landis and gary g. koch. the measurement of observer agreement for categorical data. biometrics, 33(1):159–174, 1977. adam meyers, ruth reeves, catherine macleod, rachel szekely, veronika zielinska, brian young, and ralph grishman. the nombank project: an interim report. in proceedings of the hltnaacl workshop on frontiers in corpus annotation, pages 24–31, boston, massachusetts, 2004. george a. miller, claudia leacock, randee tengi, and ross t. bunker. a semantic concordance. in proceedings of the workshop on human language technology (hlt-93), pages 303–308, stroudsburg, pennsylvania, 1993. muc-6. named entity task definition. in proceedings of the sixth message understanding conference (muc-6), pages 317–322, columbia, maryland, 1995. url http://aclweb.org/ anthology/m/m95/m95-1024.pdf. version 2.1. martha palmer, daniel gildea, and paul kingsbury. the proposition bank: an annotated corpus of semantic roles. computational linguistics, 31(1):71–105, 2005. rebecca j. passonneau. instructions for applying discourse reference annotation for multiple applications (drama). unpublished. department of computer science, columbia university, 1996. massimo poesio. mate dialogue annotation guidelines: coreference. mate deliverable d2.1, pages 134–187, 2000. url http://www.andreasmengel.de/pubs/mdag.pdf. ix dipper, zinsmeister and webber rashmi prasad, nikhil dinesh, alan lee, eleni miltsakaki, livio robaldo, aravind joshi, and bonnie webber. the penn discourse treebank 2.0. in proceedings of the 6th international conference on language resources and evaluation (lrec’08), pages 2961–2968, marrakech, morocco, 2008. james pustejovsky, josé castaño, robert ingria, roser saurı́, robert gaizauskas, andrea setzer, graham katz, and dragomir radev. timeml: robust specification of event and temporal expressions in text. in mark t. maybury, editor, new directions in question answering. papers from the 2003 aaai spring symposium, pages 28–34. aaai press, menlo park, california, 2003a. james pustejovsky, patrick hanks, roser saurı́, andrew see, robert gaizauskas, andrea setzer, dragomir radev, beth sundheim, david day, lisa ferro, and marcia lazo. the timebank corpus. in proceedings of corpus linguistics, pages 647–656, 2003b. julia ritz, stefanie dipper, and michael götze. annotation of information structure: an evaluation across different types of texts. in proceedings of the 6th international conference on language resources and evaluation (lrec’08), pages 2137–2142, marrakech, morocco, 2008. yuping zhou and nianwen xue. pdtb-style discourse annotation of chinese text. in proceedings the 50th annual meeting of the association for computational linguistics (acl’12), pages 69– 77, jeju, south korea, 2012. x dialogue & discourse 14(2) (2023) 1–48 doi: 10.5210/dad.2023.201 scoring coreference chains with split-antecedent anaphors silviu paun∗ spaun3691@gmail.com school of electronic engineering and computer science, queen mary university of london juntao yu∗ juntao.yu@qmul.ac.uk school of electronic engineering and computer science, queen mary university of london nafise sadat moosavi n.s.moosavi@sheffield.ac.uk department of computer science university of sheffield massimo poesio m.poesio@qmul.ac.uk school of electronic engineering and computer science, queen mary university of london department of information and computing science, university of utrecht editor: barbara di eugenio submitted 12/2022; accepted 04/2023; published online 08/2023 abstract anaphoric reference is an aspect of language interpretation covering a variety of types of interpretation beyond the simple case of identity reference to entities introduced via nominal expressions covered by the traditional coreference task in its most recent incarnation in ontonotes and similar datasets. one of these cases that go beyond simple coreference is anaphoric reference to entities that must be added to the discourse model via accommodation, and in particular split-antecedent references to entities constructed out of multiple discourse entities, as in split-antecedent plurals and in some cases of discourse deixis. although this type of anaphoric reference is now annotated in many datasets, systems interpreting such references cannot be evaluated using the reference coreference scorer (pradhan et al., 2014). as part of the work towards a new scorer for anaphoric reference able to evaluate all aspects of anaphoric interpretation in the coverage of the universal anaphora initiative, we propose in this paper a solution to the technical problem of generalizing existing metrics for identity anaphora so that they can also be used to score cases of split-antecedents. this is the first such proposal in the literature on anaphora or coreference, and has been successfully used to score both split-antecedent plural references and discourse deixis in the recent codi/crac anaphora resolution in dialogue shared tasks. keywords: coreference, evaluation, split-antecedent anaphors 1. introduction anaphoric reference is the use of language expressions to refer to entities already introduced in a discourse. the simplest case of anaphoric reference is identity reference, as in (1), where the *. equal contribution. listed by alphabetical order ©2023 silviu paun, juntao yu, nafise sadat moosavi and massimo poesio this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). paun, yu, moosavi and poesio anaphoric expression her is a mention of the same (discourse) entity 1 earlier introduced with the nominal expression mary (the antecedent),1 whereas anaphoric expression it refers to the same discourse entity 2 first introduced with the nominal expression a new dress. notice also that in both cases, the discourse entity is explicitly introduced with a nominal phrase, and that the anaphoric expressions refer to a single entity. we call this type of anaphoric reference single-antecedent identity anaphora. in computational linguistics (cl) / natural language processing (nlp), identity anaphora is better known as coreference (see next section), and the task is equivalently defined as that of clustering the sets of mentions referring to the same entity, or coreference chains. (1) [mary]1 bought [a new dress]2 but [it]2 didn’t fit [her]1. the dataset that currently serves as universally accepted reference testbed for some aspects of anaphoric interpretation including single-antecedent anaphora resolution is ontonotes (pradhan et al., 2012). the performance of models for single-antecedent anaphora resolution on ontonotes has greatly improved in recent years (wiseman et al., 2016; clark and manning, 2016; lee et al., 2017, 2018; kantor and globerson, 2019; joshi et al., 2020). as a consequence, the attention of the community has started to turn to cases of anaphora not annotated or not thoroughly covered in ontonotes. examples of this trend include, for instance, research on the cases of anaphora whose interpretation requires some form of commonsense knowledge tested by benchmarks for the winograd schema challenge (rahman and ng, 2012; liu et al., 2017; sakaguchi et al., 2020), or on pronominal anaphors that cannot be resolved purely using gender, for which benchmarks such as gap have been developed (webster et al., 2018). in addition, more research has been carried out on aspects of anaphoric interpretation that go beyond identity anaphora but are covered by a great many existing datasets, such as ancora for catalan and spanish (recasens and martı́, 2010), arrau (poesio et al., 2018; uryupina et al., 2020) and gum for english (zeldes, 2017), or the prague dependency treebank for czech (nedoluzhko, 2013).2 these aspects include, e.g., bridging reference (clark, 1975; hou et al., 2018; hou, 2020; yu and poesio, 2020), discourse deixis (webber, 1991; marasović et al., 2017; kolhatkar et al., 2018) and, finally, split-antecedent anaphoric reference, the type of anaphoric interpretation on which we are focusing in this paper. the best known example of split-antecedent are cases of plural anaphoric references such as pronoun they in (2) (eschenbach et al., 1989; kamp and reyle, 1993), a plural anaphoric reference to a set composed of two or more discourse entities (john and mary) introduced by separate noun phrases. (2) [john]1 met [mary]2. [he]1 greeted [her]2. [they]1,2 went to the movies. split-antecedent plural references are found in all corpora and in all languages, so that more and more annotation schemes cover them, including ancora (recasens and martı́, 2010), arrau (poesio et al., 2018; uryupina et al., 2020), friends (zhou and choi, 2018), gum (zeldes, 2017), phrase detectives (poesio et al., 2019), the prague dependency treebank (nedoluzhko, 2013), 1. the term ‘antecedent’ is often used to refer to a nominal expression–one of the nominal expressions referring to the same discourse entity as the anaphoric expression (lyons, 1977; kamp and reyle, 1993). in this paper we will however follow the less widely adopted use of cornish (cornish, 2006) and use the term to refer to the discourse entity itself. this is in part because the more traditional use of antecedent suggests that anaphora is a relation between linguistic expressions, instead of between linguistic expressions and a discourse model. in part because the antecedents of the anaphoric expressions studied here–split antecedent plurals and split-antecedent discourse deixis– are not antecedents in this sense. 2. see (poesio et al., 2016a) for a more detailed survey and (nedoluzhko et al., 2021) for a more recent, extensive update. 2 scoring coreference chains with split-antecedent anaphors and the recently created codi/crac 2021 shared task corpus of anaphora resolution in dialogue (khosla et al., 2021; yu et al., 2022a).3 split-antecedent plural references are not always common but e.g., in the friends corpus 9% of the mentions have more than one antecedent (zhou and choi, 2018). a number of computational models for the interpretation of split-antecedent plural anaphoric references have been proposed, but so far, every model has been tested using different evaluation scores (vala et al., 2016; zhou and choi, 2018; yu et al., 2020a, 2021). but other types of anaphoric reference besides plural references can have multiple antecedents introduced (‘evoked’) by separate text segments in exactly the same way. this can happen, for instance, with discourse deixis, illustrated in (3), where the antecedent for discourse-deictic demonstrative that in 4.2 is the plan formed by the actions in 1.5/1.6 and 3.1, evoked by utterances from speaker m separated by utterances from speaker s.4 (3) 1.1 m all right system 1.2 we’ve got a more complicated problem 1.3 uh 1.4 first thing i’d like you to do 1.5 is [send engine e2 off with a boxcar to corning to pick up oranges]1 1.6 uh [as soon as possible]1 2.1 s okay 3.1 m and [while it’s there it should pick up the tanker]2 4.1 s okay 4.2 and [that]1,2 can get 4.3 we can get that done by three split-antecedent plural references and discourse deictic references are important from the point of view of our understanding of anaphora because they illustrate the fact that reference possibilities in discourse models are not limited to entities introduced in the discourse model via nominals–indeed, such cases were one of the reasons for, e.g., webber’s development of the idea of discourse model (webber, 1979). unlike simple cases of identity coreference, such cases of anaphora refer to entities that need to be added to the discourse model at the point in which the anaphoric reference is encountered via some form of inference. this operation on the discourse model is generally called accommodation (lewis, 1979; beaver and zeevat, 2007). assigning an antecedent to splitantecedent anaphors requires a particularly simple form of inference–creating a new plural object out of two atomic objects–but more complex cases are known requiring additional inferences (as in, e.g., context change accommodation (webber and baldwin, 1992; fang et al., 2022)). these types of anaphoric reference thus test the ability of an anaphora resolution system to create antecedents ex novo instead of choosing them from the already introduced mentions, which is what differentiates proper discourse models from simple history lists of referents. assessing this ability, however, requires a scorer that can evaluate the interpretation produced by a system in cases that require accommodation. but even though split-antecedent plurals are cases of identity anaphora, they are not covered by the reference coreference scorer (pradhan et al., 2014). simplified forms of evaluation for this type of references have been developed for discourse deixis and used, e.g., in the 2018 crac shared task (poesio et al., 2018), but have not become part of 3. see the corefud report prepared for the universal anaphora initiative for an extensive discussion of the coverage of split-antecedent and other anaphors in the universal anaphora corpora (nedoluzhko et al., 2021). 4. this example is from the trains subset of the arrau corpus (uryupina et al., 2020). 3 paun, yu, moosavi and poesio a standardized scorer for anaphora, and the existing metrics for discourse deixis do not work for split-antecedent discourse deixis cases such as (3). the previously proposed methods for scoring split-antecedent plural anaphors (vala et al., 2016; zhou and choi, 2018; yu et al., 2021), either work only for split-antecedent plurals in isolation, or require a substantial redefinition of the notion of coreference chain, or only generalize one of the existing metrics for coreference, as discussed in detail in section 6. the objective of this paper is to fill this gap in the literature, as part of the effort to develop the new universal anaphora scorer (yu et al., 2022b),5 an extension of the reference coreference scorer (pradhan et al., 2014; moosavi and strube, 2016; poesio et al., 2018) that can evaluate all aspects of anaphoric interpretation currently covered by the universal anaphora (ua) initiative,6 including bridging reference, discourse deixis, and some cases of accommodation, and was used as the official scorer for the 2021 and 2022 codi/crac shared tasks on anaphoric interpretation in dialogue (khosla et al., 2021; yu et al., 2022a).7 we start in section 2 with a summary of the types of anaphoric interpretation the new scorer is meant to cover and a more detailed discussion of accommodation and split-antecedent plural reference. in section 3 we review the current situation of evaluation for coreference and provide a brief description of the metrics we set to extend. crucially, these include all the most widely used metrics for coreference evaluation (luo and pradhan, 2016): mention and entity-based metrics such as b3 (bagga and baldwin, 1998) and ceaf (luo, 2005) respectively, as well as link-based metrics such as muc (vilain et al., 1995), lea (moosavi and strube, 2016), and blanc (luo et al., 2014; recasens and hovy, 2011). our key contribution in this paper is a method for scoring references to accommodated objects created out of antecedents separately introduced that generalizes all existing metrics and thus can be used to score both single and split antecedent anaphoric reference in exactly the same way. our proposed extensions of these metrics that can score both single and split-antecedent references are presented in section 5 and illustrated in detail with an example in section 4.4. our solution is compared with all alternative proposals regarding the scoring of split antecedent anaphors in section 6. for a more thorough demonstration of how the metrics can be used, we use the new ua scorer incorporating these extended metrics to score both split antecedent plural reference and discourse deixis with multiple antecedents. in this paper, we use this scorer to show how the proposed generalization, unlike the proposal by zhou and choi (2018) and our own proposal in (yu et al., 2021), can be used to score both single and split antecedent anaphora in exactly the same way, as well as to further illustrate and analyze the behavior of our generalized metrics on the data used in previous work on split-antecedent plurals (section 7). finally, in section 8 we discuss how the approach proposed here could be extended to cover other types of anaphoric interpretation requiring accommodation. 2. anaphoric phenomena currently in the scope of the universal anaphora initiative the ultimate objective of the universal anaphora initiative is to develop guidelines for annotating a broad range of anaphoric phenomena–not just identity anaphora–and evaluation methods to score systems carrying out the interpretation of all these types of reference. we briefly summarize here the 5. https://github.com/juntaoy/universal-anaphora-scorer 6. http://www.universalanaphora.org 7. https://competitions.codalab.org/competitions/30312. https://codalab.lisn. upsaclay.fr/competitions/614 4 https://github.com/juntaoy/universal-anaphora-scorer http://www.universalanaphora.org https://competitions.codalab.org/competitions/30312 https://codalab.lisn.upsaclay.fr/competitions/614 https://codalab.lisn.upsaclay.fr/competitions/614 scoring coreference chains with split-antecedent anaphors types of anaphoric reference covered by the current version of the universal anaphora scorer (yu et al., 2022b), and discuss in more detail the type of reference that is the focus of the present paper because of the lack of widely accepted evaluation methods: reference to accommodated entities, as exemplified by split-antecedent anaphora. 2.1 (identity) anaphora and coreference in much cl / nlp literature following the first message understanding conference (muc) shared tasks (chinchor and sundheim, 1995) a distinction is made between anaphora resolution and coreference resolution, and the term ‘anaphora’ is used to indicate pronominal anaphora only, whereas the term coreference is supposed to cover anaphoric reference with all types of nominals as well as, originally at least, predication. in this article, however, the terms ‘anaphora’ and ‘anaphoric reference’ is used in the more general sense of reference to entities in the discourse model adopted in linguistics (see, e.g., lyons (1977); kamp and reyle (1993)) psycholinguistics (see, e.g., garnham (2001)) and the pre-muc computational work such as, e.g., webber (1979); carter (1987); luperfoy (1991) (see (mitkov, 2002; poesio et al., 2016b; gundel and abbott, 2019; poesio et al., 2023) for discussion). according to these theories of anaphoric reference, language is interpreted with respect to a discourse model (or mental model) which, in its simplest form, consists of discourse entities and their properties. interpreting language involves, among other things, updating the discourse model with new entities as they are mentioned, or referring back anaphorically to the entities already introduced in the discourse model. for instance, in possibly the most widely known linguistic theory of interpretation in discourse, discourse representation theory (drt) (kamp and reyle, 1993), the first mention of mary in example (1), proper name mary, introduces a new discourse entity 1 in the discourse model, and all subsequent mentions of mary, whether using pronouns or proper names, are considered (identity) anaphoric references to that entity 1. the term coreference resolution was introduced for the message understanding conference (chinchor and sundheim, 1995) to specify a rather different task covering several interrelated aspects of language interpretation of interest for information extraction including not only identity anaphora resolution as in the example above but also, e.g., the association of properties with entities in cases such as (4), where the np an nlp researcher, normally considered predicative from a linguistic perspective, would be considered ‘co-referent’ with mary: (4) [mary]1 is [an nlp researcher]1. this term ‘coreference’ was maintained in nlp even when criticism of the original definition of the coreference task (see, e.g., van deemter and kibble (2000)) led the field to adopt a revised definition focusing exclusively on (a subset of) identity anaphora in the (psycho-) linguistic sense (passonneau, 1997; poesio et al., 1999) adopted most notably in ontonotes (pradhan et al., 2012). we will use here ‘anaphora’ in the traditional sense from (psycho-)linguistics, but we will adopt the characterization of identity anaphora interpretation and notation that have become standard in the computational literature, in particular in the literature defining metrics for the evaluation of this task. most modern anaphoric annotation projects cover the basic case of identity anaphoric reference to entities introduced via nominals in (1), although substantial differences exist regarding which types of identity anaphora are covered. equally, the several proposed metrics for evaluating ‘coreference’ (reviewed in section 3), and the reference coreference scorer developed for the conll shared task and incorporating many of these metrics (pradhan et al., 2014) all focus exclusively on 5 paun, yu, moosavi and poesio evaluating identity reference / identity anaphora in this sense. these metrics are defined by reference to mentions– the nominal phrases referring to a discourse entity–and a discourse entity i is identified with the set (cluster) ki of mentions referring to that entity: the coreference chain. 8 2.2 beyond identity anaphora: the universal anaphora /corefud specification of anaphoric phenomena as already mentioned, many other types of anaphoric reference exist beyond basic identity anaphora, including, e.g., bridging references or associative anaphora (clark, 1975; hou et al., 2018; hou, 2020; yu and poesio, 2020), and cases of anaphoric reference to entities not introduced using nominals, such as the cases of split-antecedent anaphoric reference to accommodated entities which are the focus of this paper. these other types of anaphoric reference are not covered, or are covered only partially, in ontonotes because of complexity and cost reasons (see (pradhan et al., 2012) as well as the discussion in (zeldes, 2022)), but as discussed above, they are annotated in the majority of the most recent anaphoric corpora, including ancora (recasens and martı́, 2010), arrau (poesio and artstein, 2008; poesio et al., 2018; uryupina et al., 2020), gum (zeldes, 2017), phrase detectives (poesio et al., 2019), the prague dependency treebank (nedoluzhko, 2013), the tüba-dz corpus (versley, 2008) the recently created codi/crac corpus of anaphoric reference in dialogue (khosla et al., 2021; yu et al., 2022a), and many others (for a more complete list, see (poesio et al., 2016a, 2023)). the objectives of the universal anaphora initiative include, first of all, identifying the types of anaphoric reference most frequently annotated in current corpora; second, to define shared guidelines for the annotation of these aspects of anaphoric reference; and third, to develop methods that can be used to evaluate anaphoric resolvers carrying out these types of interpretation. the list of phenomena currently considered includes the distinction between referring and non-referring expressions; identity anaphora also via zeros and split antecedent plurals; bridging reference; and discourse deixis. (details can be found in the draft proposal on the universal anaphora website,9 as well as in the corefud report (nedoluzhko et al., 2021).) the universal anaphora scorer (yu et al., 2022b) can evaluate the interpretation of all of the types of anaphoric reference covered by the current proposal. 2.3 anaphoric references requiring accommodation one of the types of anaphoric reference covered by the current universal anaphora proposal is split-antecedent anaphora: anaphoric reference to a set of entities previously mentioned separately from each other. in fact, split-antecedent anaphora is only one example of a more general class of anaphoric references that require so-called accommodation of a new antecedent (lewis, 1979; van der sandt, 1992; beaver and zeevat, 2007).10 in this paper we propose a general method for scoring split-antecedent anaphora resolution that can be used both for split-antecedent plurals and for split-antecedent discourse deixis. in this section we briefly discuss accommodation in general and the types of split-antecedent anaphoric references covered by the current proposal. 8. one of our reviewers pointed out that this use of sets of mentions as the ’referent’ of anaphoric expressions is only really appropriate for anaphoric reference–a proper notion of entity is required for other types of reference, e.g., to objects in the scene in situated dialogue. 9. https://github.com/universalanaphora/universalanaphora/blob/main/universal_ anaphora_1_0___proposal_for_discussion.pdf 10. note that in fact bridging references discussed earlier require a form of accommodation as well: the entity related to the entity already introduced has to be added to the discourse model. 6 https://github.com/universalanaphora/universalanaphora/blob/main/universal_anaphora_1_0___proposal_for_discussion.pdf https://github.com/universalanaphora/universalanaphora/blob/main/universal_anaphora_1_0___proposal_for_discussion.pdf scoring coreference chains with split-antecedent anaphors accommodation one of the most powerful arguments for the discourse model view of anaphora, as opposed to older history-list approaches11 is the fact that many cases of anaphoric reference cannot be interpreted with respect to the entities already introduced in the discourse model with a nominal, but require new entities to be added to the discourse model, or accommodated (webber, 1979; kamp and reyle, 1993; garnham, 2001). accommodation as conceived by lewis (1979) is a general operation on the discourse model which involves adding new content that is required to process a statement (beaver and zeevat, 2007). for instance, (5) presupposes that it was raining; when processing the statement, that fact has to be added to the discourse model. (5) mary realized it was raining. webber (1979); kamp and reyle (1993) and garnham (2001) discuss a number of cases of anaphoric interpretation requiring new entities to be added to the discourse model to interpret anaphoric reference. we focus on the following two types of split-antecedent reference. split-antecedent plural reference in ontonotes, plural reference is only marked when the antecedent is mentioned by a single noun phrase. however, split-antecedent plural reference is also possible (eschenbach et al., 1989; kamp and reyle, 1993), as in example (2). these are cases of plural identity reference, but where the antecedents are sets whose elements are two or more entities introduced by separate noun phrases, which have to be accommodated in (i.e., added to) the discourse model (kamp and reyle, 1993; beaver and zeevat, 2007). such references are annotated in many modern datasets, including, e.g., arrau (uryupina et al., 2020), gum (zeldes, 2017) and phrase detectives (poesio et al., 2019; yu et al., 2023b). the complexities involved in evaluating systems interpreting split-antecedent plural references are illustrated in our reference (artificial) example (6). in this and other examples in the paper, mentions are delimited by square brackets; the discourse entity realized by a mention is indicated using a subscript; and the mention id is indicated with a superscript. so, for instance, [him5]1 is a mention via noun phrase him; it is the 5th mention in the text; and realizes discourse entity 1. split-antecedent references are indicated by listing all discourse entities that are part of the set, as in [their6]1,2. example (6) was designed to illustrate in a concise way some of the key properties that a scorer for split-antecedent must have. first of all, that it would not be correct to interpret a split-antecedent plural such as [their6]1,2 in the third sentence in terms of links between mentions –say, stipulate that this plural pronoun should be interpreted by every system in terms of links between this mention and mentions [she4]2 and [him5]1. clearly, a scorer must be able to interpret such a reference in terms of the constituent entities 1 and 2, so as to consider as equally correct system interpretations of the pronoun linking it to any of the mentions of these entities. second, the example shows that plurals in a text may refer to sets of more than two entities. third, different split antecedent plurals may refer to different sets. we use example (6) in later sections to provide a detailed illustration of how the extended metrics work. (6) [john1]1 met [mary2]2 after work and proposed to [her3]2 to go see a play. [she4]2 liked the idea but suggested to [him5]1 to have dinner first. on [their6]1,2 way to the restaurant, [they7]1,2 met [bill8]3 and [jane9]4 . [the two10]1,2 were very happy to see [bill11]3, as [they12]1,2,3 go way back. [he13]3 introduced [jane14]4 and [all four15]1,2,3,4 agreed to have dinner together to catch up. 11. see e.g., poesio et al. (2016b) for a review of early approaches to anaphoric interpretation. 7 paun, yu, moosavi and poesio our readers might wonder whether references such as [their6]1,2 should be treated as bridging references–specifically, element-inverse associative references (uryupina et al., 2020) to the entities referred to by the mentions [she4]2 and [him5]1 . this interpretation is not inaccurate,12 but it is incomplete, as it fails to capture the fact that john and mary are the entire set of entities referred to by the plural pronoun [their6]1,2 . split-antecedent plural references are not evaluated either by the standard reference coreference scorer (pradhan et al., 2014) or by our own codi/crac 2018 scorer (poesio et al., 2018), and therefore are not generally attempted by anaphoric resolvers. a few dedicated evaluation methods concerned with evaluating this type of reference were however proposed, as discussed in section 6; but these proposals either work only for split-antecedent plurals in isolation, or require a substantial redefinition of the notion of coreference chain, or only generalize one of the existing metrics. as mentioned earlier, the key contribution of this paper is a method for scoring references to accommodated objects created out of antecedents separately introduced that generalizes all existing metrics and thus can be used to score both single and split antecedent plural anaphoric reference in exactly the same way. this extension also applies to the (much more frequent) phenomenon of split-antecedent discourse deixis, as discussed next. split-antecedent discourse deixis a second case of anaphoric reference that often also involves split antecedents, but that has been the focus of more nlp research than split-antecedent plural reference is discourse deixis, or anaphora with non-nominal antecedents (webber, 1991; byron, 2002; gundel et al., 2003; artstein and poesio, 2006; kolhatkar et al., 2018). discourse deixis, exemplified by this issue in (7), which refers to an abstract object (asher, 1993) whose type would not be easy to specify, is a type of abstract anaphora in which the antecedent is some type of abstract entity ‘evoked’ by the propositional content of a previous sentence. the evidence on discourse deixis interpretation suggests that these antecedents are not introduced in the discourse model immediately, but are accommodated upon encountering the anaphoric reference (kolhatkar et al., 2018). (7) the municipal council had to decide [whether to balance the budget by raising revenue or cutting spending]i. the council had to come to a resolution by the end of the month. [this issue]i was dividing communities across the country. (kolhatkar et al., 2018) discourse deixis was not annotated in many early coreference corpora (artstein and poesio, 2006; kolhatkar et al., 2018), and when it was, the problem was simplified in a number of ways. in ontonotes, for instance, only event anaphora, a subtype of discourse deixis, is marked, as exemplified by that in (8), which refers to the event of a white rabbit with pink ears running past alice.13 (8) ... when suddenly a white rabbit with pink eyes ran close by her. there was nothing so very remarkable in [that]; nor did alice think it so very much out of the way to hear the rabbit say to itself, ’oh dear! oh dear! i shall be late!’ .... however, many of the more recent corpora–such as, e.g., ancora, gum, arrau, and the codi/crac corpus–do annotate the whole range of discourse deixis, covering, in addition to event anaphora, ref12. at least according to the annotation scheme for bridging references in the arrau corpus; [their6]1,2 would not be considered a bridging reference according to other annotation schemes such as, e.g., that for the isnotes corpus (markert et al., 2012). 13. this example is from the annotated version of alice in wonderland in the phrase detectives corpus (poesio et al., 2019; yu et al., 2023b). 8 scoring coreference chains with split-antecedent anaphors erences such as this issue in (7). but even in these corpora the task of annotating discourse deixis is simplified in a number of ways. one such simplification is that the annotators are not asked to mark as antecedent the accommodated abstract entity, but the list of sentences or clausal units that evoke the antecedent (e.g., the clause [whether to balance the budget by raising revenue or cutting spending] in (7)). the result of this simplification is that many if not most discourse deictic references annotated in such corpora are in fact cases of split antecedent discourse deictic reference, like the example in (3). (in fact, whereas split antecedent plurals are relatively rare–see section 7–split antecedent discourse deictic references are very common in, e.g., arrau, from which example (3) was taken, and in the codi/crac corpus.) no standard metric for scoring discourse deixis resolution exists, but typically systems are scored by their ability to identify the sentence evoking the antecedent. the success @ n metric, introduced by kolhatkar et al. (2013), considers a response as correct if the key sentence is among the top n candidate response sentences identified as ‘antecedent’ by the system. this metric was used by marasović et al. (2017) and in the crac 2018 shared task (poesio et al., 2018). the problem with this metric is that only one gold sentence per discourse deictic reference can be specified, meaning that this metric cannot be used to evaluate discourse deictic references with split antecedents, such as the example in (3). in universal anaphora, discourse deixis is treated following the approach that has become standard for event anaphora (lu and ng, 2018). discourse deictic references are marked in a separate discourse deixis layer, and differs from identity reference only in that the antecedents of discourse deixis are utterances rather than nominal phrases (yu et al., 2023b). the interpretation of discourse deictic references can then be scored using the same metrics developed for the interpretation of identity reference (yu et al., 2022b). hence, the proposal in this paper for scoring split antecedent plural references can also be used without any change to handle split antecedent cases of discourse deixis such as the one in (3). (and was in fact used to score such cases in both editions of the codi/crac shared task on anaphora resolution in dialogue.) context change accommodation as a final example of split antecedent anaphoric reference requiring accommodation we will mention the cases of context change accommodation discussed by webber and baldwin (1992) such as (9), where a new entity, the dough, is obtained by mixing together flour and water. (9) add [the water]i to [the flour]j little by little. then work [the dough]i,j . context-change accommodation has recently become again the object of interest in the community and has been annotated in corpora including chemu-ref (fang et al., 2021) and reciperef (fang et al., 2022). although the type of accommodation required by this type of reference is not the focus of this paper as it involves more than combining separate entities in the discourse model in sets, we will argue in section 8 that a version of the proposal in this paper could be used to score this type of anaphoric reference, as well, and could be argued to be more suited than the metrics used e.g., in (fang et al., 2021, 2022). 3. metrics for scoring identity anaphora (coreference) one of the fundamental issues with anaphora resolution / coreference, and a reason of great dissatisfaction among many practitioners, in the fact that although the field has at various time points 9 paun, yu, moosavi and poesio converged on an ‘official’ metric developed for particular shared tasks, and one such metric–the conll score (denis and baldridge, 2007; pradhan et al., 2014)–has become dominant in the last ten years (see below), it is far from clear that this metric captures our intuitions about how anaphoric interpretation should be evaluated–or indeed, what these intuitions are. this is not to say that the field is completely divided. for instance, it is universally accepted that coreference evaluation should be entity-based, in the sense that a system’s interpretation of, e.g., (10) should be evaluated on the extent that system recognizes that 1, 2 and 3 are all mentions of the same entity (alternatively, elements of the same coreference chains), as opposed to merely its ability to identify, say, mention 1 as the previous mention of entity i that is mentioned by mention 2 (‘link 2 to 1’) or mention 3 as the previous mention for mention 2 (‘link 3 to 2’). (10) [mary1]i woke up late that morning, so [she2]i rushed out of bed– [she3]i had an important meeting. however, agreement that evaluation should be entity-centered still leaves many degrees of freedom, with the result that different ways have been proposed of computing precision, recall and f1 by comparing the entities (i.e., the coreference chains) in the gold annotation–in anaphora resolution these gold coreference chains are generally known as the keys–with those produced by a system, or responses (vilain et al., 1995; bagga and baldwin, 1998; luo, 2005) and no consensus has been reached on which metric is most appropriate. this impasse was broken by denis and baldridge (2007), who introduced a measure which was adopted in the conll 2011 and 2012 shared tasks and has since become known as the conll score and adopted as semi-standard (pradhan et al., 2012, 2014), but this hasn’t stopped the development of new metrics (recasens and hovy, 2011; moosavi and strube, 2016). before introducing the proposed extensions to additionally evaluate split-antecedent references, we briefly discuss in this section the metrics standardly used to evaluate single antecedent reference and their existing definitions. please consult, e.g., (luo and pradhan, 2016) for more in-depth discussion. 3.1 notation the standard coreference evaluation metrics are based on the simplification discussed in section 2.1 that discourse entities can be identified with coreference chains of mentions introduced in a text, and the assumption that each mention refers to a single entity. we use the following notation to indicate that an entity ki is identified with a coreference chain of mi mentions mi,j : ki = {mi,1,mi,2, ...,mi,mi} (as we will see in section 5, one of the key proposals in this paper is a more complex representation of entities for entities not introduced via mentions.) following standard convention, we use k and r to refer to entities from the key and from the response sets, respectively. the key entities represent the gold standard, whereas the response entities are the entities proposed by a system to be evaluated. the metrics to be described below adopt different approaches to comparing key and response entities. 10 scoring coreference chains with split-antecedent anaphors 3.2 the standard metrics for anaphora / coreference resolution evaluation 3.2.1 standard muc muc (vilain et al., 1995) is a link-based metric that evaluates response entities on the basis of the number of links they have in common with the entities in the key. it is standard practice to compute this information indirectly, by counting the number of missing links, and discarding them from the maximum number of possible links, as we are about to see.14 for the case of recall, this amounts to: recallmuc = ∑ i |ki| − |p(ki;r)|∑ i |ki| − 1 in the equation above p(ki;r) is a function called the partition function, that returns all the partitions of key entity ki with respect to the response r of a system: p(ki;r) = { ki ∩rj | 1 ≤ j ≤ |r| } ⋃ ki,u∈ki\r { {ki,u} } notice how the number of partitions indicates the number of links found in the key entities but not in the response entities. to compute precision, we simply swap the key and response sets: precisionmuc = ∑ i |ri| − |p(ri;k)|∑ i |ri| − 1 the muc metric reports as a final value an f1-measure, which is the harmonic mean between the precision and recall presented above: f = 2× precision × recall precision + recall 3.2.2 standard b3 one problem with the muc metric is that, by definition, it only scores a system’s ability to identify links between mentions; its ability to recognize that a mention does not belong to any coreference chain–i.e., its ability to classify a mention as a singleton–does not get any reward (bagga and baldwin, 1998; luo and pradhan, 2016). the b3 metric (bagga and baldwin, 1998) was proposed to correct this problem. b3 is a mention-based metric: the evaluation measures the number of mentions common between the entities in the key and in the response. recall is computed by calculating recall for every mention, which is: r(m) = |ki ∩rj | |ki| and then summing all these mention recalls up. this can be done by finding all |ki ∩rj | mentions in the intersection of key entity ki and response entity rj , summing up recall for these: r(i, j) = ∑ m∈ki∩rj r(m) = |ki ∩rj | ∗ |ki ∩rj | |ki| = |ki ∩rj |2 |ki| 14. in muc the maximum number of links in an entity is the minimum number of links needed to connect its mentions. 11 paun, yu, moosavi and poesio and then summing up across all i, j pairs and averaging by the total number of mentions. the result is: recallb3 = ∑ i,j ( |ki∩rj | )2 |ki|∑ i |ki| precision is computed in a similar way, again by swapping the key entities and the response entities. an f1 measure can then be computed from precision and recall in the usual way. 3.2.3 standard ceaf b3 also suffers from a problem–namely, that a single chain in the key or response can be credited several times. this leads to anomalies–e.g., if all coreference chains in the key are merged into one in the response, b3 recall is one (luo, 2005; luo and pradhan, 2016). the solution proposed by luo (2005), ceaf, is an entity-level metric: it aligns one to one the entities in the key with those in the response, and then it computes their similarity following the same function used for the alignment step. the metric comes in two flavors, depending on the similarity function used to align and compare the entities. given a key entity ki and a response entity rj , mention-based ceaf is calculated using a similarity function measuring their mention overlap: ϕm (ki, rj) = |ki ∩rj | whereas in entity-based ceaf, the similarity between the two entities, is computed using the dice coefficient: ϕe(ki, rj) = 2 ( |ki ∩rj | ) |ki|+ |rj | at the alignment step, the kuhn-munkres algorithm (kuhn, 1955; munkres, 1957) is used to find the optimal one-to-one mapping between the entities in the key and the response such that their cumulative similarity is maximal. let k∗ ⊂ k and r∗ ⊂ r be the key and the response entities for which an alignment was established, respectively, and let g : k∗ → r∗ be an alignment function storing the one-to-one mappings. then recall is defined as the sum of the similarities over the maximal possible similarity: recallceaf = ∑ i ϕ(k ∗ i , g(k ∗ i ))∑ i ϕ(ki,ki) and again, precision is computed by swapping the entities from the key with those from the response set, and f1ceaf is computed from recallceaf and precisionceaf as usual. 3.2.4 standard lea lea (moosavi and strube, 2016) is a link-based, but entity-aware metric which measures the resolution score of entities while taking into consideration their importance. so for instance for recall we have: recalllea = ∑ i importance(ki)× resolution-score(ki)∑ i importance(ki) the importance of an entity could be defined in a different ways depending on the task, but for instance, it could be defined as being proportional to the entity’s size so that more weight is given 12 scoring coreference chains with split-antecedent anaphors to more frequently mentioned entities: importance(ki) = |ki| the resolution score, for recall, measures the proportion of links in the key that are recovered in the response: resolution-score(ki) = ∑ j links(ki ∩rj) links(ki) in lea the number of links in an entity is counted as the total number of links that can be formed between their mentions. for example, for an entity ki, we have: links(ki) = ( |ki| 2 ) as with the other metrics seen so far, precision is evaluated by swapping the key with the response entities, and f1 is computed as usual. 3.2.5 standard blanc finally, blanc (recasens and hovy, 2011; luo et al., 2014) is an adaptation for coreference resolution of the rand index used in clustering (rand, 1971). the key feature of the rand index is that it is computed assessing not only a clustering algorithm’s decisions to put two entities in the same cluster, but also its decisions to put them in different clusters. blanc adapts this approach to the case of coreference also taking into account the imbalance between the number of coreference (same cluster, aka coreference chain) vs. non-coreference (different cluster / coreference chain) links. to compute blanc we need to determine the number of common coreference and non-coreference links found in the key and in the response entities. for a key ki the coreference links are computed as follows: ck(i) = {(mi,u,mi,v) | mi,u ∈ ki,mi,v ∈ ki,mi,u ̸= mi,v} whereas the non-coreference links between two key entities ki and kj are computed as follows: nk(i, j) = {(mi,u,mj,v) | mi,u ∈ ki,mj,v ∈ kj} it follows that the set of all coreference and non-coreference links from the keys are: ck = ∪ick(i), nk = ∪i ̸=jnk(i, j) the coreference links cr and the non-coreference links nr for the response entities are computed in a similar way. precision and recall values are then computed for both coreference and noncoreference links. for example, for recall, we have: recallc = |ck ∩ cr| |ck | , recalln = |nk ∩nr| |nk | 13 paun, yu, moosavi and poesio 3.3 the case for diversity in anaphora / coreference resolution evaluation some subfields of nlp appear to have standardized on a single evaluation metric. the stereotypical example of this situation is perhaps machine translation, where the bleu metric (papineni et al., 2002) would appear to have been accepted as the de facto standard for measuring progress in the field. another example is summarization, where rouge (lin, 2004) and its variants are equally dominant. the situation for coreference evaluation is apparently very different, and one might ask whether this should be a worry– shouldn’t we just choose one of these metrics and generalize that? we are not going to argue for choosing a particular metric in this paper. first of all, that would require providing experimental evidence that one metric is ‘best’ in some sense, which would be well beyond the scope of this paper. but we should also keep in mind, secondly, that the existence of multiple evaluation metrics in an nlp subfield is not unusual, and not necessarily problematic; other areas of nlp are also characterized by the existence of multiple metrics. the situation of coreference is very similar to the situation of parsing (kakkonen, 2007), where in addition to the widely used parseval metric, developed in connection with a shared task (black et al., 1991), a great number of other metrics exist, including, e.g., leaf-evaluation, cross-bracketing, minimum tree edit distance, and their variants. in fact, we would argue that the field has converged on a single metrics only for tasks which involve assigning a single label to a linguistic expression (e.g., part of speech tagging, named entity recognition), or for sub-fields in which the progress is driven by shared tasks organized by governments or industry, such as machine translation. multiple metrics are the norm for tasks where the ‘labels’ to be compared are more complex, and it’s far from obvious how these metrics should be compared. and even in fields which have converged on a single metric, such as machine translation, the debate is still raging–other metrics exist, and are often argued to be superior (see, e.g., the case of meteor for mt (banerjee and lavie, 2005)). thirdly, we should point out that a de facto convergence has been reached in anaphora / coreference, on the conll metric proposed by denis and baldridge (2007) as the average among the f1 values obtained using muc, b3 and ceaf. this score was used in the conll shared tasks in 2011 and 2012 (pradhan et al., 2012) and since then has become the standard for the field. but this metric is an average of three of the metrics introduced in this section; so computing it requires generalizing all component metrics. we believe, therefore, that until the field reaches a new convergence, the best approach to promote the development of coreference resolvers able to resolve split antecedents is to extend all the current metrics in a way that is completely transparent to the developers of those systems in particular by extending all three metrics on which the computation of the current reference score, the conll score, is based, as done in this paper. 4. generalizing standard metrics to allow for split-antecedent references: key ideas and terminology the standard metrics for coreference resolution discussed in the previous section all expect mentions to refer to a single entity. in this section we describe an extension of these metrics that also allows for references to multiple entities, as in the cases of split-antecedent anaphors illustrated in (2). the generalization we propose follows the spirit of the existing metrics. the existing metrics assess the proposed single-antecedent references in the response on the basis of whether they refer to the same entity as the key; we propose that split-antecedent references should be evaluated in the same way, i.e., on whether the entities they refer to are the same. this ensures that two systems which propose as antecedents of a split-antecedent anaphor different mentions, but that refer to the 14 scoring coreference chains with split-antecedent anaphors same entities, will be considered equivalent. conversely, two systems will be considered different when a split antecedent anaphor is taken to refer to different entities by the two systems. the key idea on which our generalisation is based on is to compare the entities referred to by a split-antecedent anaphor using the very same metrics we set to extend. so, for example, when scoring a system using muc, we propose to score split-antecedent anaphora resolution according to how well the component entities of an accommodated set in the response match the key (gold) entities according to the same muc metric. similarly, for b3 evaluation we use the b3 metric, and so on for the other metrics. in this way, we can handle split-antecedent anaphora evaluation within coreference evaluation without altering the existing evaluation paradigm. we assess both single and split-antecedent references using the very same metrics, preserving their individual strengths and weaknesses, as evolved over years of research. a second important characteristic of our proposed extension is that when no split-antecedents are present, the scores produced by the extended metrics are identical to the scores obtained using their standard formulation. in this section, we introduce the technical ideas underlying our proposal–accommodated sets, their alignment, and the δ term responsible for comparing them–which will then used in section 5 in which we introduce the proposed generalizations one by one, and demonstrate their computation with reference to our main example (6). 4.1 accommodated sets in order to handle split-antecedent references, we generalize the notion of entity introduced in section 3.1 to also allow entities consisting of the merge of an accommodated object ko i constructed from the discourse model (e.g., a set constructed from the existing entities in the case of split antecedent anaphors) and a traditional coreference chain km i of mentions of that object. we indicate this using the following notation: ki = ko i ⊕km i different types of accommodated objects are involved in anaphoric reference, as discussed earlier in section 2.3 and then later in section 8. in the case of split-antecedents anaphors, we use the notation ks i to refer to the accommodated object–the set that serves as antecedent for the split-antecedent reference. by definition, ks i is a set composed of two or more entities: ks i = {ki,1,ki,2, ...,ki,si} we use the term accommodated set to refer to ks i , i.e., the set of two or more entities in the discourse model which was accommodated in the context to serve as the antecedent of a split antecedent anaphor. the antecedent entities of a split-antecedent pronoun, in our representation above, are atomic entities, –i.e., entities all of whose mentions refer to a single antecedent. it is however possible for a split-antecedent anaphor to refer to an entity which in turn contains a split antecedent anaphor among its mentions. for instance, in the following example, the split antecedent anaphor they all has as split antecedents the entity bill and the set consisting of john and mary, which in turn was accommodated in the discourse as a consequence of the split antecedent anaphor they in the second utterance. (we omit mention indices in this example.) 15 paun, yu, moosavi and poesio (11) [john]1 met [mary]2. [they]1,2 went to the movies, and met [bill]3. afterwards, [they all]1,2,3 went to dinner. in this case, we recursively replace the antecedents which are themselves accommodated sets (i.e., {1,2} = {john,mary}) with their element entities, so that the larger accommodated set (the antecedent of they all) has only ‘atomic’ entities as its elements, i.e., only entities referred to using single-antecedent anaphors: ({1,2,3} = {john,mary,bill}).15 note that this does not lead to any loss of generality; we resolve the split-antecedent references to sets of atomic entities as those are the actual antecedents if you unpack the recursive references. (see section 4.4 for a more detailed discussion of such cases.) we use the following notation to indicate the set of all accommodated sets, and the set of all regular mentions (i.e., the single-antecedent references), respectively: ks = ⋃ i ks i , km = ⋃ i km i the notation above was introduced for the entities in the key, but the corresponding notions will also be used for the entities in the response. also, the formulation of the coreference metrics involves computing the cardinality of an entity. we generalize the notion of cardinality to complex entities ks i ⊕km i in the obvious way as follows: |ks i ⊕km i | = 1 + |km i | finally, notice how an entity without a split-antecedent has the same representation as seen before in section 3.1, where only single-antecedent references were allowed: ki = km i = {mi,1,mi,2, ...,mi,mi} 4.2 aligning accommodated sets all the standard coreference evaluation metrics assume an implicit ‘alignment’ between the mentions in the key and in the response (the single-antecedent references). we say a mention in the key and one in the response are aligned if they share the same boundaries. we need to know which mentions align to compute the metrics: depending on the metric, the aligned mentions or the links between these mentions are used to compare the key and response entities following that metric’s strategy, as discussed in section 3. if only single-antecedent anaphors are present, and if response entities are well-formed, i.e., if mentions are not repeated across entities, each mention from the key is aligned with at most one mention in the response. however, only aligning mentions is not sufficient if we also have split-antecedent anaphors. as discussed in section 4.1, split-antecedent anaphors result in the introduction of entities containing accommodated sets. thus, evaluating split-antecedent anaphora interpretation requires additionally aligning accommodated sets. we align the accommodated sets by aligning their element entities. if we did not specify a one-to-one alignment, the metrics would be ill-defined, i.e., contributions from multiple partially-overlapping accommodated sets may accumulate and inflate the scores. 15. this assumption amounts to adopting what has become known as the ’union’ theory of plurals (link, 1983; schwarzschild, 1996). see (link, 1984; landman, 1989) for alternative views, and (schwarzschild, 1996; winter and scha, 2015) for discussion. 16 scoring coreference chains with split-antecedent anaphors as with the evaluation strategy briefly mentioned in the beginning of this section, we propose to align the accommodated sets using the very same metric that a system is to be evaluated with, for both single and split-antecedent anaphora. i.e., we propose to align the accommodated sets in the key entities and in the response entities for the purpose of computing metric µ using the f1µ scores that the same metric µ returns for the element entities of those accommodated sets. for example, when computing a muc score, we compute the alignment score between a key accommodated set ks i included in an entity i and a response accommodated set rs j from an entity j as follows: ϕ(ks i , r s j) = f1muc(k s i , r s j) the alignment process involves finding the pairs of accommodated sets from the key and the response that lead to the largest cumulative f1 score. since a brute-force approach to this problem can be computationally unfeasible, we use the kuhn-munkres algorithm (kuhn, 1955; munkres, 1957) adopted also in ceaf, which solves the alignment problem in polynomial time. let ks′ ⊂ ks and rs′ ⊂ rs be the subsets of the aligned key and response accommodated sets; then we will use the following function (and its inverse for the reverse mappings) to access the aligned split-antecedent pairs: τ : ks′ → rs′ 4.3 the δ term the generalization of the existing evaluation metrics that can be used to score both single and split-antecedent references proposed in this paper is uniform across all metrics and only requires an additional δ term responsible for the comparison between accommodated sets or between links involving accommodated sets, depending on the metric. as noted earlier in this section, to compute the score between accommodated sets for metric µ we compute the score according to µ between the component entities of these accommodated sets. the accommodated sets will receive scores between 0 and 1 for how well they are resolved by a system. as we will see in section 5, the δ score has different interpretations for different metrics. however, two principles apply to all generalizations. the first is that if entity i contains no accommodated sets, the value of the generalized metric is equivalent to that of the original formulation (see section 3). second, when accommodated sets are perfectly resolved, the contribution from the accommodated sets equals that from regular mentions in the entity; they are after all just another element of the coreference chains. 4.4 the (artificial) illustrative example, revisited we will illustrate the terminology just introduced using example (6), repeated here as example (12). (12) [john1]1 met [mary2]2 after work and proposed to [her3]2 to go see a play. [she4]2 liked the idea but suggested to [him5]1 to have dinner first. on [their6]1,2 way to the restaurant, [they7]1,2 met [bill8]3 and [jane9]4 . [the two10]1,2 were very happy to see [bill11]3, as [they12]1,2,3 go way back. [he13]3 introduced [jane14]4 and [all four15]1,2,3,4 agreed to have dinner together to catch up. only the mentions of entities referred to by split-antecedent anaphors are marked in (12); other mentions that would be interpreted by coreference resolvers in a normal way, such as a a play or 17 paun, yu, moosavi and poesio the restaurant, are left out from our discussion, to keep things simpler, although their interpretation would also be scored by our (generalized) scorer of course. the key entities mentioned in the example above are as follows (again, we use subscripts to indicate entities, and indicate the mention number using superscripts): k1 = {[john1], [him5]} k2 = {[mary2], [her3], [she4]} k3 = {k1,k2} ⊕ {[their6], [they7], [the two10]} k4 = {[bill8], [bill11], [he13]} k5 = {[jane9], [jane14]} k6 = {k1,k2,k4} ⊕ {[they12]} k7 = {k1,k2,k4,k5} ⊕ {[all four15]} the first entities are atomic entities k1 (john) and k2 (mary). the first accommodated set component appears with k3: this is the set with elements entities k1 and k2. notice next the accommodated set component of k6. the anaphor [they12] refers to entity k4 (bill) and to entity k3 referred to by [the two10] but entity k3 in turn contains an accommodated set with elements the entities k1 (john) and k2 (mary). as described in section 4.1, we normalize the representation of accommodated sets so that they only contain atomic entities as constituents: k6 = {k3,k4} ⊕ {[they12]} = {k1,k2,k4} ⊕ {[they12]} in our illustrations of the metrics we will compare the responses produced by two hypothetical coference resolvers on example (12). let us call the first coreference resolver ‘system a’. system a outputted the following response: ra,1 = {[john1], [him5]} ra,2 = {[mary2], [she4]} ra,3 = {ra,1, ra,2} ⊕ {[their6], [they7], [the two10], [they12]} ra,4 = {[bill8], [bill11], } ra,5 = {[jane9], [jane14]} ra,6 = {ra,1, ra,2, ra,5} ⊕ {[all four15]} system a made several mistakes. it did not include mention [her3] in ra,2 and [he13] in ra,4, respectively, and mistakenly interpreted [they12] as a reference to john and mary instead of to john, mary and bill. also, system a only produced a partially-correct interpretation of the split antecedent anaphor [all four15] which in the key refers to john, mary, bill, and jane, whereas a only recovered 3 of the 4 constituent entities. using the method discussed in section 4.2, the accommodated sets in the response from system a align with those in the key as follows:16 16. this alignment is optimal irrespective of the similarity metric used. 18 scoring coreference chains with split-antecedent anaphors τ(ks 3) = rs a,3, τ(ks 6) = ∅, τ(ks 7) = rs a,6, τ(rs a,3) = ks 3 , τ(rs a,6) = ks 7 we will compare the score assigned to system a with those assigned to systems b, c and d whose output is a variation on how the accommodated set in k7 may be resolved that raise interesting questions about the way our proposed generalizations operate: rb,6 = {ra,1, ra,2, ra,4, ra,5} ⊕ {[all four15]} rc,6 = {ra,1, ra,2, ra,4} ⊕ {[all four15]} rd,6 = {ra,2, ra,4, ra,5} ⊕ {[all four15]} compared with ra,6, the accommodated set in rb,6 proposed by system b for mention [all four15] correctly includes all 4 entities: john, mary, bill and jane. the entity rc,6 in c’s response includes an accommodated set consisting of the entities john, mary and bill, that aligns better with the accommodated set in the key interpretation for mention [they12] k6, than with the key interpretation for mention [all four15] k7. system d’s interpretation for [all four15], rd,6, also proposes 3 entities as antecedents for the anaphor, like ra,6, but the interpretation is slightly worse than that proposed in a (ra,1 = k1, but ra,4 = k4 \ {[he13]}). 5. generalized definitions for the coreference metrics we are now in the position to provide generalizations of the existing evaluation metrics that can be used to score both single and split-antecedent references. in this section, we go through the standard coreference metrics one by one, showing how they are generalized using the δ term, and then illustrating their computation using example (12). note that it is not possible to evaluate the new generalized metrics in the traditional way, i.e., by showing that they produce more intuitive outputs than existing metrics on some examples, as done in papers introducing new coreference metrics such as (bagga and baldwin, 1998; luo, 2005; recasens and hovy, 2011; moosavi and strube, 2016), for the simple reason that no generalization of all existing metrics to cover split antecedent references has been proposed before. (we discuss in section 6 the existing proposals regarding scoring such cases, none of which involves generalizing all existing metrics, and all of which have other limitations, as discussed there.) instead, we illustrate our generalizations and make a case for them as follows. in this section, we describe in detail how the scores for each metric are computed under the proposed extension with reference to example (12), with the dual objective of illustrating how our generalization works in practice and showing that the results obtained are sensible in the sense that systems producing intuitively ‘better’ responses get better scores. however, showing a step-bystep computation of all of the metric scores, for both single and split-antecedent references, would be tedious. since the proposed generalization does not modify the computation of the metrics on single-antecedent anaphors, but just adds an additional δ term for the evaluation of split-antecedent references, we only illustrate here the computation of this term. for a step-by-step guide to the computation of the standard version of the metrics, i.e., defined only for single-antecedent references, see (luo and pradhan, 2016). then, in section 6, we compare our generalization metrics in 19 paun, yu, moosavi and poesio detail to the few existing and very partial previous proposals. finally, in section 7, we show that the extended metrics work in practice, in that they are effective at differentiating between systems which are intuitively better at the task and systems which are intuitively worse, and compare to the existing metrics, by using them to evaluate a system carrying out split-antecedent reference resolution on the datasets used in the papers in which the two most recent proposals for scoring split antecedent plural reference were made, yu et al. (2021) and zhou and choi (2018). 5.1 generalized b3 5.1.1 definition the simplest illustration of how the δ term is used to generalize an existing metric is our generalization of b3. we illustrate this with recall∗ b3 . again, the aim is to give a system full credit when the component entities in an accommodated set in the response exactly match those in the key. (e.g., in example (12), a system would get full credit for their interpretation of [their6] if this interpretation refers to an accommodated set r with two constituent entities r1 and r2 which perfectly match the two constituent entities k1 and k2 in the gold interpretation of [their6].) such perfectly resolved split-antecedent references should make the same contribution to recall∗ b3 as correctly identified single-antecedent references. when these conditions are not met, i.e., when the system either does not produce an accommodated set as the interpretation of a split-antecedent anaphor, or this accommodated set is not aligned with that in the key, no credit should be given. (see section 4.2 for why alignment is necessary.) in intermediate cases, when the accommodated set in the response partially matches the accommodated set in the key, a recall score between 0 and 1 should be obtained for that mention. when the document does not contain any accommodated sets, generalized b3 should be equivalent to standard b3. these goals are achieved by adding to the standard formula for recallb3 (discussed in section 3.2.2) a δ term responsible for evaluating the resolution of split-antecedent references, if any. r∗(m) = |km i ∩rm j |+ δi,j |ki| notice that most of the b3 formula stays unchanged (see section 3.2.2 for a direct comparison). what has changed is that the cardinality of an entity, if it contains a split-antecedent reference, is one greater than the cardinality of the entities with single-antecedent references only. δ for b3 is defined as follows. when a key ki and a response rj contain an aligned accommodated set, i.e., τ(ks i ) = rs j , δi,j should specify how well the system resolved the component entities of the accommodated set. we do this using the very same recall metric used for the evaluation of the single antecedents: δi,j = recallb3 ( ks i , r s j ) in this way, the system is given full credit (δi,j = 1) if the component entities in the accommodated set in the response exactly match (recall-wise) those in the key. i.e., perfectly resolved splitantecedent references make a contribution to recall identical to that of correctly identified singleantecedent references. by contrast, when the system does not produce an accommodated set as the interpretation of a split-antecedent anaphor or this accommodated set is not aligned with that in the key, δi,j = 0. when the document does not contain any accommodated sets, δi,j = 0 as well, so 20 scoring coreference chains with split-antecedent anaphors that recall∗ b3 is equivalent to recallb3 . in the intermediate cases–the system response for a splitantecedent anaphor is aligned with an accommodated set in the key, but the match is not perfect– δ gets a intermediate value between 0 and 1 reflecting how well the response matches the key, as desired. (see 5.1.2.) to get overall recall, all the r(m) are summed up as before, giving: r∗(i, j) = ∑ m∈ki∩rj r∗(m) = (|km i ∩rm j |+ δi,j) ∗ |km i ∩rm j |+ δi,j |ki| = (|km i ∩rm j |+ δi,j) 2 |ki| and the following formula for overall recall: recall∗ b3 = ∑ i,j ( |km i ∩rm j |+δi,j )2 |ki|∑ i |ki| to compute precision, we proceed in the usual way and replace the keys with the responses and vice-versa. f1∗ b3 is also computed in the usual way. 5.1.2 computing generalized b3 on the example let us now see in more detail how recall∗ b3 works with reference to example (12). again, we focus on computing recall, as precision proceeds in the same way. the key entities k1,k2,k4 and k5 are only referred to using single-antecedent mentions. this means that there is no link to an accommodated set that the response should recover; thus, w.r.t. these entities, all the δ terms are 0, for all the systems we are considering, a, b, c and d. in these particular cases the value of the metric is not affected by our generalization. moving forward, let us focus first on system a. when a key and a response entity contain aligned accommodated sets, we need to credit the system for how well it resolved the accommodated set in the key. looking at the mappings specified by the alignment function, this only happens in two cases. the first case involves the accommodated sets from the k3 and ra,3 entities: δa,3,3 = recallb3 ( ks 3 , r s a,3 ) = recallb3 ( {k1,k2}, {ra,1, ra,2} ) = |k1∩ra,1|2 k1 + |k2∩ra,2|2 k2 |k1|+ |k2| = 2 3 the second is between the accommodated sets in k7 and ra,6: δa,7,6 = recallb3 ( ks 7 , r s a,6 ) = recallb3 ( {k1,k2,k4,k5}, {ra,1, ra,2, ra,5} ) = 8 15 21 paun, yu, moosavi and poesio in all other cases, i.e., ∀(i, j) ∈ {1, 2, ..., 7} × {1, 2, ..., 6} \ {(3, 3), (7, 6)}, δa,i,j = 0. let us now compare how system a resolved the split-antecedent references with the results of systems b, c, and d, similarly computed. the relevant δ terms are the following: δa,7,6 = 8 15 , δb,7,6 = 2 3 , δc,7,? = 0, δc,6,6 = 7 12 , δd,7,6 = 7 15 among the four systems, system b (correctly) gets the highest credit for identifying all 4 entity elements of the accommodated set in k7. system c gets no credit for k7 as that accommodated set does not align with any coreference chain in c’s response (we indicate this using the notation δc,7,?)–rc,6 aligns best with entity k6. finally, system d gets a slightly lower score compared with system a which makes intuitive sense since entity ra,1 is better resolved compared with ra,4. 5.2 generalized muc 5.2.1 definition the muc metric, as well, can be generalized to score both single and split-antecedent references by including into the original formula an additional δ term responsible for scoring split-antecedent references using the muc metric being generalized, but there is an additional complication. as with b3, we indicate generalized recallmuc as recall∗muc. when computing recall∗muc, the additional δ term again measures how well the system resolves the links in the key involving accommodated sets, just as in the case of generalized b3. however, δ is used in a different way for muc–instead of being added to recall, it is used to compute a penalty term by subtracting it from 1, defined as 1− δ. if an entity ki does not contain an accommodated set, then there is no such link for the response to recover, and therefore we want the penalty to be 0: this is obtained by having δi,j = 1. when this is the case for all entities (i.e., when no split-antecedent anaphor is present in a document) then recall∗muc is equivalent to standard recallmuc. when a key entity does contain an accommodated set, however, we need to score the response for how well it recovers the link to this accommodated set. a response link is credited if one of its nodes matches a split-antecedent anaphor in the key and the other consists of the aligned accommodated set. when the match is perfect, we again want the penalty to be zero, but we want it to grow as the match becomes less perfect. this is done by again computing δi,j using recallmuc, as follows (where i is the index of a gold entity and j the index of a response entity). if ∃(rm j , rs j) s.t. rm j ∈ km i and τ(ks i ) = rs j then: δi,j = recallmuc ( ks i , r s j ) again, notice that we are comparing the accommodated sets in the key and in the response using the very same metric (recallmuc) used to evaluate the system for both single and split-antecedent references. if the key accommodated set ks i and the aligned response accommodated set rs j (τ(ks i )) match perfectly, i.e., if their component entities are the same, we have δi,j = 1, and so the system is fully rewarded for correctly producing the link to a split-antecedent contained by the key. a partial penalty is applied for an accommodated set in the response with recallmuc < 1, i.e., when the component entities of the accommodated set are not perfectly resolved by the system. if there is no link in the response that satisfies the aforementioned conditions, i.e., either its nodes do not contain a regular mention matching with the key, or the aligned accommodated set, then a missing link penalty is applied, and we set δi,j = 0. 22 scoring coreference chains with split-antecedent anaphors recall∗muc is computed by substracting the penalty term from the standard definition, as follows: recall∗muc = ∑ i |ki| − |p(km i ;rm)| − (1− δi,j)∑ i |ki| − 1 the partition function p() takes as arguments the regular mentions portion of the entities, i.e., the single-antecedent references (cfr. section 4.1), so that part of the muc formula stays unchanged. as in the case of b3, what changes is that the cardinality of an entity, if it contains a split-antecedent reference, is one greater than the cardinality of the entities with single-antecedent references only ( section 3.2.1). to generalize muc precision, we simply swap the entities in the key and the response, as per standard practice: precision∗muc = ∑ i |ri| − |p(rm i ;km)| − (1− δi)∑ i |ri| − 1 , δi = precisionmuc ( ks i , r s j ) note also that this time the δ term is computed using the precision of the metric, in accordance with the principle discussed earlier of using the very same metric for both single and split-antecedent references. f1muc is computed as usual. 5.2.2 computing generalized muc on the example we will again focus on recall. and again, we focus on the entities in the key that do contain split-antecedent references, whose interpretation must be found in the responses and assessed. we start with the response provided by system a. the accommodated set in entity k6 cannot be optimally aligned with any accommodated set in ra; therefore, a missing link penalty is applied in this case, i.e., δa,6,? = 0. in the case of the other two entities in the key containing an accommodated set, k3 and k7, links to an accommodated set are recovered in the response (as the conditions to have a matching regular mention and aligned accommodated sets are satisfied) and need to be evaluated. the accommodated set in ra,3, {ra,1, ra,2}, is an example of a system identifying the correct number of antecedents for a split antecedent anaphor, but resolving imperfectly the antecedent entities. its δa,3,3 value is as follows: δa,3,3 = recallmuc ( ks 3 , r s a,3 ) = recallmuc ( {k1,k2}, {ra,1, ra,2} ) = ∑ ki∈{k1,k2} |ki| − |p(ki; {ra,1, ra,2})|∑ ki∈{k1,k2} |ki| − 1 = 2 3 the logic of this is that the entities included in the response accommodated set above have a recall of 2/3, so 1/3 is deducted from system a’s recall for the link to the accommodated set. (note that the ‘mention’ component of ra,3, rm a,3, is also incorrect as it includes an extra mention in comparison with k3, but this aspect of the interpretation is not discussed here as it is not affected by our generalization.) 23 paun, yu, moosavi and poesio in the case of the accommodated set in k7, system a recovers only 2 of the 3 antecedents; the missing link penalty is computed as follows: δa,7,6 = recallmuc ( ks 7 , r s 6 ) = recallmuc({k1,k2,k4,k5}, {ra,1, ra,2, ra,5}) = 1 2 let us now compare the penalty applied to system a for the interpretation of [all four15] with those applied to b,c and d. the relevant δ terms are: δa,7,6 = 1 2 , δb,7,6 = 2 3 , δc,7,? = 0, δc,6,6 = 0, δd,7,6 = 1 2 resulting in the following penalty (1 δ) being deducted from the overall recall: (1− δa,7,6) = 1 2 , (1− δb,7,6) = 1 3 , (1− δc,7,?) = 1, (1− δc,6,6) = 1, (1− δd,7,6) = 1 2 we can see that system b receives a smaller penalty compared to system a, which makes sense considering b recovers all 4 entity references. system c is the one most heavily penalised by the scorer, because, first of all, the optimal alignment for the accommodated set produced by c is not the one in k7, but that in k6, so that c ends up completely missing the accommodated set in k7 that the other systems get some credit for. still, when computing recall with respect to k6, system c continues to get a full penalty even though the accommodated sets align this time around, but a link cannot be determined because no regular mentions match. finally, system d gets the same penalty as system a. this is because ra,1 and ra,4, the entities that are different between the accommodated sets from a and d, both contribute with one link when computing muc recall. (we saw when discussing the b3 metric earlier that system a does a boost in score over system d with that metric–this might perhaps be considered a problem with muc). 5.3 generalized ceaf 5.3.1 definition as discussed in section 3.2.3, central to computing the ceaf metric are the similarity functions used to both align and compare the key and response entities. to extend the metric to evaluate split-antecedent references as well we use, as in the other cases, an additional δ term which, as in the case of b3, is added to both the similarity function used in the computation of mention-based ceaf, and to that used in entity-based ceaf: ϕm (ki, rj) = |km i ∩rm j |+ δmi,j , ϕe(ki, rj) = 2 ( |km i ∩rm j |+ δei,j ) |ki|+ |rj | it is only when a key ki and a response rj contain aligned accommodated sets –i.e., τ(ks i ) = rs j– that we need to evaluate how well their component entities were resolved. δ for recall is computed as follows: δmi,j = recallceafm ( ks i , r s j ) , δei,j = recallceafe ( ks i , r s j ) when there are no accommodated sets or they are not aligned, δi,j = 0. otherwise, recallceaf is computed exactly as discussed in section 3.2.3, and so for precisionceaf and f1ceaf. 24 scoring coreference chains with split-antecedent anaphors 5.3.2 computing generalized ceaf on the example as with b3, it is only when a key and a response entity contain aligned accommodated sets that we need to evaluate how well the two match. again, we start with the computations for system a. we have seen there is an alignment in two cases, one between the accommodated sets in entities k3 and r3: δ m/e a,3,3 = recallceafm/e ( ks 3 , r s a,3 ) = { 4 5 for mention-based ceaf 9 10 for entity-based ceaf the calculation is done for recall, precision is analogous. the other accommodated set alignment is between those in the k7 and ra,6 entities: δ m/e a,7,6 = recallceafm/e ( ks 7 , r s a,6 ) = { 3 5 for mention-based ceaf 7 10 for entity-based ceaf all other cases either involve entities without accommodated sets, or whose accommodated sets do not align, so we have δ m/e i,j = 0. we now introduce, for comparison, the relevant δ terms for the other systems: δma,7,6 = 3 5 , δmb,7,6 = 4 5 , δmc,7,? = 0, δmc,6,6 = 3 4 , δmd,7,6 = 3 5 δea,7,6 = 7 10 , δeb,7,6 = 9 10 , δec,7,? = 0, δec,6,6 = 13 15 , δed,7,6 = 13 20 as with the other metrics presented so far system b gets the highest score for identifying all four entities from the split-antecedent reference in k7. system c is not allocated any credit for the accommodated set in k7, only for the one from k6, because of how the split-antecedents get aligned. and system d is found on par with system a when using mention-based ceaf and slightly worse (as it intuitively should) when the evaluation is conducted using entity-based ceaf. this is another illustration of the strengths and weaknesses of the existing metrics for coreference evaluation which our scorer inherit. for the computation of the rest of the metrics, the observations for systems b, c, and d will be similar to those expressed so far, and will be omitted. from now on, we will only calculate the scores for system a to illustrate the methods, as the calculations for the other metrics are the same. 5.4 generalized lea 5.4.1 definition to extend lea to evaluate both single and split-antecedent references we modify both the importance and the resolution-score functions. starting with the former, we define the importance function to additionally include a β term to further reward entities which contain an accommodated set: importance(ki) = βi|ki| 25 paun, yu, moosavi and poesio the resolution score function is defined as in the standard version of the metric, but computed differently to also consider split-antecedent references: resolution-score(ki) = ∑ j links(ki ∩rj) links(ki) counting the number of links in ki is trivial: links(ki) = (|ki| 2 ) . but special attention needs to be paid when counting the number of links between the set of mentions in common between a key ki and a response rj : links(ki ∩rj) = ( |km i ∩rm j | 2 ) + δi,j × ( |km i ∩rm j | − 1 ) we distinguish two types of links between the mentions common to both a key and a response entity. first, we have links between regular mentions present in both entities, i.e., links between single-antecedent references. the number of these links is expressed in the first term of the equation above. the second type are links involving an accommodated set. when a key ki and a response rj contain aligned accommodated sets, i.e., τ(ks i ) = rs j , we need to assess how well their element entities compare. again, this can be done just using lea’s notion of recall: δi,j = recalllea ( ks i , r s j ) after evaluating how well the key and the response accommodated sets compare, we use this information to weigh the number of links that can have an accommodated set as a node; this is expressed in the second term from the link counting formula presented earlier. when there are no accommodated sets in the entities, or when they are not aligned, we set δi,j = 0. the generalized version of the metric is equivalent to the standard version presented in section 3.2.4 when the entities in the key and response do not contain any accommodated sets (when δi,j = 0, we have link(ki ∩rj) = (|km i ∩rm j | 2 ) ). when accommodated sets do exist, however, and they were perfectly resolved by a system, the scorer allocates full credit to each of these links involving accommodated sets, just as it does for the links between correctly identified single-antecedent references: when δi,j = 1, links(ki ∩ rj) = (|km i ∩rm j |+1 2 ) . for imperfectly resolved accommodated sets the credit allocated to a system lies in-between the two extremes. 5.4.2 computing generalized lea on the example when computing lea, for those key and response sets that contain aligned accommodated sets, we need to evaluate how well these accommodated sets compare. for recall, we have: δi,j =  recalllea ( ks 3 , r s a,3 ) for i = j = 3 recalllea ( ks 7 , r s a,6 ) for i = 7, j = 6 0 otherwise when computing precision, lea precision is used instead to evaluate accommodated sets. 26 scoring coreference chains with split-antecedent anaphors 5.5 generalized blanc 5.5.1 definition we saw back in section 3.2.5 that to compute blanc we need to establish the coreference and the non-coreference links found in the key and in the response entities. in the standard version of the metric these links are only between regular mentions, i.e., between single-antecedent references. in the generalized version, in which entities may also include accommodated sets, we additionally distinguish two types of links: (i) links where both nodes are accommodated sets,17 and (ii) links where one node is an accommodated set, and the other is a single-antecedent reference. blanc evaluates a response by comparing the coreference and non-coreference links in the response set with those in the key. in standard blanc this is done by simply computing the intersection of these sets. in the generalized version of the metric, however, we cannot do this anymore, due to the introduction of accommodated sets and the additional types of links that get created, as discussed above. to help us compare different types of links we introduce a new function δ() that takes as argument two links–one from the key (mk 1,m k 2), the other from the response (mr 1,m r 2)–and specifies how to allocate them credit. in short, this function will allocate full credit (i.e., a value of 1) to links between regular mentions that match, and partial credit (a value between 0 and 1) to those key and response links whose nodes involve aligned accommodated sets. the partial credit in this case will depend on how well the element entities in the accommodated sets compare. the function will not allocate any credit (a value of 0) to all other pairs of links. we can assess how the coreference and the non-coreference links in the response compared to those in the key by evaluating the credit allocated by the function above to all pairs of links found between these sets. if no accommodated sets are present in the entities in the key and the response, using the function as described has the same effect as the set intersection operation mentioned before used in standard blanc. we illustrate below how the δ() function is used to compute recall for the non-coreference links: rn = 1 |nk | ∑ (mk 1 ,m k 2)∈nk (mr 1,m r 2)∈nr δ ( (mk 1,m k 2), (m r 1,m r 2) ) a link from the key and one from the response whose nodes are regular mentions receive a credit of 1 if their mentions match. formally, if mk 1,m k 2 ∈ km,mr 1,m r 2 ∈ rm, and mk 1 = mr 1,m k 2 = mr 2 (or mk 1 = mr 2,m k 2 = mr 1), then: δ ( (mk 1,m k 2), (m r 1,m r 2) ) = 1 two links one of whose nodes is a regular mention while the other is an accommodated set are scored based on how well the accommodated sets compare, assuming they are aligned, and that the regular mentions match. in line with the rest of the metric extensions, we compare accommodated sets by comparing their element entities using the very same metric we are evaluating the system with overall, for both single and split-antecedent references. the current example is for the recall of the non-coreference links, so this is the metric that is used here as well. formally, if ∃mk s ,m k m ∈ {mk 1,m k 2} s.t. mk s ∈ ks,mk m ∈ km, and ∃mr s,m r m ∈ {mr 1,m r 2} s.t. mr s ∈ rs,mr m ∈ rm, and 17. these can occur among the non-coreference links. 27 paun, yu, moosavi and poesio mk m = mr m, τ(mk s) = mr s then: δ ( (mk 1,m k 2), (m r 1,m r 2) ) = recalln (mk s ,m r s) two links whose nodes are aligned accommodated sets receive credit on the basis of how well their element entities compare. formally, if mk 1,m k 2 ∈ ks,mr 1,m r 2 ∈ rs, and τ(mk 1) = mr 1, τ(m k 2) = mr 2 ( or τ(mk 1) = mr 2, τ(m k 2) = mr 1 ) then: δ ( (mk 1,m k 2), (m r 1,m r 2) ) = recalln ( mk 1, τ(m k 1) ) × recalln ( mk 2, τ(m k 2) ) all other links in the key and response are unrelated and receive no credit. for these, we have: δ ( (mk 1,m k 2), (m r 1,m r 2) ) = 0 all the other computations required by blanc, i.e., the precision of the non-coreference links, together with both the precision and the recall of the coreference links, are computed in a similar fashion. blanc then reports as its final value the arithmetic mean of the f1 values for the coreference and the non-coreference links. 5.5.2 computing generalized blanc on the example for this metric we need to determine how the coreference and the non-coreference links in the key and response entities compare. the link space is large, but let us look, for example, at nk(3, 7) and nr(3, 6), the sets of non-coreference links between keys k3 and k7, and response entities ra,3 and ra,6, respectively. starting with the former, we have:18 nk(3, 7) = { (ks 3 ,k s 7), (k s 3 , 15), (6,k s 7), (6, 15), (7,k s 7), (7, 15), (10,k s 7), (10, 15) } the non-coreference links between the response entities ra,3 and ra,6 are: nr(3, 6) = { (rs a,3, r s a,6), (r s a,3, 15), (6, r s a,6), (6, 15), (7, rs a,6), (7, 15), (10, r s a,6), (10, 15), (12, r s a,6), (12, 15) } let us now consider how matching links are determined in a recall-based evaluation. two links, one from the key, and the other from the response, whose nodes are regular mentions, get full credit if their mentions match. in our example that happens in 3 cases: δ ( (6, 15), (6, 15) ) = 1 δ ( (7, 15), (7, 15) ) = 1 δ ( (10, 15), (10, 15) ) = 1 18. we shall use the id of the mentions, single or split-antecedent, for a more concise representation. 28 scoring coreference chains with split-antecedent anaphors two links, where one of the nodes is a regular mention, and the other an accommodated set, are given credit if the regular mentions match and the accommodated sets are aligned. the allocated credit depends on how well the accommodated sets evaluate: δ ( (ks 3 , 15), (r s a,3, 15) ) = blancrn (ks 3 , r s a,3) δ ( (6,ks 7), (6, r s a,6) ) = blancrn (ks 7 , r s a,6) δ ( (7,ks 7), (7, r s a,6) ) = blancrn (ks 7 , r s a,6) δ ( (10,ks 7), (10, r s a,6) ) = blancrn (ks 7 , r s a,6) two links both of whose nodes are accommodated sets receive credit if the accommodated sets are aligned, and the score depends on how well the response evaluates against the key: δ ( (ks 3 ,k s 7), (r s a,3, r s a,6) ) = blancrn ( ks 3 , r s a,3 ) × blancrn ( ks 7 , r s a,6 ) there is no alignment for all other key and response links, so no credit can be allocated in these cases, and δ = 0. finally, notice we used rn to compare the entities included in the accommodated sets, as the computations above were used to determine the credit allocated to non-coreference links in a recall-based evaluation. computing the credit for the coreference links involves the same steps, but using rc instead. and when turning to precision, pn and pc , the precision related metrics from blanc are used. 6. alternative proposals to score split-antecedent anaphora there is limited previous work on split-antecedent anaphora resolution and its evaluation. we are aware of four proposals, two of which we put forward ourselves in previous work. metrics for gold evaluation only vala et al. (2016) and yu et al. (2020a) only evaluate their models on split-antecedent selection task that assumes gold anaphors and mentions were provided. they compute precision, recall, and f1 measures based on the links between split-antecedent anaphors and their antecedent. recall is defined as the percentage of gold split-antecedent anaphoric links that are correctly identified by the system (where each atomic antecedent of a split-antecedent is counted as a separate link). precision is the percentage of split-antecedent anaphoric links proposed by a system that is correct. because they only evaluate on the split-antecedent selection task, the evaluation is much simpler, i.e. it does not require any form of alignment. however, such an evaluation has a number of limitations: it is not entity-based, and is not realistic, as it is based on the assumption that the gold mentions and gold anaphors will be provided. distributing the plural mentions among the coreference chains for singular entities zhou and choi (2018) propose a method to evaluate split-antecedent plural references resolution using the standard conll scorer. this is done by adding the plural mention to each of the clusters for its atomic elements: for example, they would represent the coreference chain {{johni, maryj}, theyt} 29 paun, yu, moosavi and poesio encoded in this paper as the discourse entity kt including an accommodated set : ki = {johni1i , ...johnimi } kj = {maryj1j , ....maryjnj } kt = {ki,kj} ⊕ {theyt1t } by adding [theyt1]t to the coreference chains for john and mary, i.e., as the following two gold coreference chains each consisting of all mentions in one individual entity plus the split-antecedent reference: {[johni1]i, . . . [johnim]i, [theyt1]t} {[maryj1]j , . . . [maryjn]j , [theyt1]t}. first of all, this representation clearly violates the fundamental assumption behind the notion of coreference chain, i.e., that all mentions in the chain refer to the same entity. in the two coreference chains, mention [theyt1]t does not refer to the same entity as mention [johnik ]i of john or mention [maryjl]j of mary: it refers to the set consisting of john and mary. in addition, this approach also causes problems with the evaluation and may produce counterintuitive results. one example of a problem is that the proposed representation cannot be used with the muc metric. muc relies on the partition function to compute the number of links in common between the key and response entities, and the function requires all mentions to participate in a single cluster. as a result, the authors do not use the muc metric in their evaluation. regarding counterintuitive results, the authors themselves point out that the approach may have unpredictable effects on the alignment of the ceafe score. the following example illustrates this. suppose we have a text which first introduces the entity john, mentioned four times; the entity mary, mentioned once; the entity bill, also mentioned once; followed by three mentions of the set {john,mary,bill}. this situation is schematically encoded by the following key entities: k1 = {john1, john2, him3, he4} k2 = {mary5} k3 = {bill6} k4 = {k1,k2,k3} ⊕ {they7, them8, their9} and suppose that the response entities predicted by a system are as follows: r1 = {john1} r2 = {mary5} r3 = {r1, r2} ⊕ {they7, them8, their9} in the approach proposed by zhou and choi (2018), the plural mentions [they7], [them8] and [their9] would be represented by distributing them into the singular clusters (e.g. k1, r1) –i.e., the key and response entities would be represented as follows: 30 scoring coreference chains with split-antecedent anaphors k1 = {john1, john2, him3, he4, they7, them8, their9} k2 = {mary5, they7, them8, their9} k3 = {bill6, they7, them8, their9} r1 = {john1, they7, them8, their9} r2 = {mary5, they7, them8, their9} intuitively, the system’s coreference chain for john, r1, should be aligned with k1; distributing the plurals should not change this alignment. but in reality, r1 will be aligned with k3, because after adding the plurals the similarity score between r1 and k3 is 3/8, which is higher than the score of 4/11 between r1 and k1. by contrast, our approach will correctly align the singular and plural clusters (i.e. k1 and r1; k2 and r2, and k4 with r3). generalizing lea finally, in our own previous work (yu et al., 2021), we proposed an extension of the lea metric to score split-antecedent plural references. a first obvious difference between the proposal in this paper and that earlier proposal is that back then we only proposed a generalization for lea, whereas here we generalize every one of the most used coreference metrics discussed in section 5. but another reason for the current proposal is that after publishing that paper we found issues with the approach to alignment we had used. the generalization proposed in that paper scores split-antecedent references in two steps, each of which requires an alignment step. in the first step, the singular coreference chains in the response are aligned with the gold coreference chains. the atomic antecedents of the split-antecedent references found in the response are then replaced by the aligned gold coreference chains. the alignment between the coreference chains in the key and response in this step is established using ceafe . after that, a second alignment step is used between the converted split-antecedents when computing the link scores in the lea. using the notation introduced in section 4.1, our early approach can be formulated as follows. a gold accommodated set ks i is aligned with an accomodated set rs j in the response, i.e., τ(ks i ) = rs j , if rs j is the accommodated set in the system entities that has the largest number of coreference chains in common with ks i , i.e., rs j = argmaxrs j∗ |ks i ∩ rs j∗|. following our presentation of lea in section 5.4, the aligned accommodated sets, when evaluating recall, are assessed based on the recall of their mentions: δi,j = |ks i ∩rs j | |ks i | however, a one-to-one alignment between the accommodated sets was not imposed, hence there might be multiple accommodated sets in the key being aligned to the same accommodated sets in the response, and vice-versa. this is potentially problematic as the coreference metrics are based on the assumption that a mention can only participate in one coreference chain. the following example shows how the alignment used in our early approach might cause some instability in the metrics. 31 paun, yu, moosavi and poesio suppose we have the following key and response entities: k1 = {john1, john2, him3, he4} k2 = {mary5} k3 = {bill6} k4 = {k1,k2} ⊕ {they7} k5 = {k1,k2,k3} ⊕ {all three8} r1 = {john1, john2, him3, he4} r2 = {mary5} r3 = {bill6} r4 = {r1, r2} ⊕ {they7} r5 = {r1, r3} ⊕ {all three8} (i.e., the system correctly resolved atomic entities 1, 2 and 3 and plural entity 4, but then only identified two of the antecedents for split antecedent reference [all three8].) in our approach in yu et al. (2021), we will first align the singular coreference chains (e.g. k1, r2) without accommodated sets using ceafe , and then replace the ‘atomic’ coreference chains that are part of accommodated sets in the response with the aligned gold atomic coreference chains. as a result, r4 and r5 will become: r4 = {k1,k2} ⊕ {they7} r5 = {k1,k3} ⊕ {all three8} after that, when computing the actual lea score, the accommodated set part of k4, ks 4 ({k1,k2}), will be aligned to rs 4, which is expected as they are identical. ks 5 ({k1,k2,k3}) intuitively should be aligned to rs 5, since rs 4 has already been aligned. but because there is no one-to-one alignment restriction, ks 5 is equally likely to be aligned to either rs 4 or rs 5, as both have the same δi,j of 2/3 (two out of three clusters are overlapping). the actual outcome solely depends on the order in which the response split-antecedents are presented. by contrast, with the new method proposed in this paper, ks 5 will be aligned correctly with rs 5 due to the one-to-one alignment constraint. a second drawback of yu et al. was caused by the pre-alignment step on the atomic entities using the ceafe score. the pre-alignment step is deterministic, and imposes a hard alignment between key and response coreference chains; and the later scoring step does not take into account how good the alignments are. this is problematic as the hard alignment in this stage is not robust enough to align multiple equivalent response clusters, as illustrated by the following example. suppose we have the following key entities: k1 = {john1, john2, him3, he4} k2 = {mary5} k3 = {k1,k2} ⊕ {they6} 32 scoring coreference chains with split-antecedent anaphors and consider the following two responses, which differ on the interpretation of [they6]: r1 = {john1, him3} r2 = {john2, he4} r3 = {mary5} r4 = {r1, r3} ⊕ {they6} r′ 4 = {r2, r3} ⊕ {they6} the two system predictions for [they6], r4 and r′ 4, should be equivalent, since they both get half of k1 and all of k2. however, because the pre-alignment step requires a hard alignment, either r1 or r2 will be aligned with k1. suppose r1 is aligned with k1; the yu et al.’s scorer will then give a score of 100% to {r1, r3} but 50% to {r2, r3}. by contrast, the new scorer proposed in this paper will correctly assign the same scores to both interpretations. in order to provide a more empirical comparison between our new approach to evaluation and the one we proposed in yu et al. (2021), in the next section we use our new generalizations to score the same system tested by yu et al. (2021) on the same corpus. we also test the same systems on the friends corpus used by zhou and choi (2018) to provide a baseline for future work. 7. using the metrics for scoring generalized anaphoric reference because no generalizations of the existing coreference metrics to cover split-antecedent anaphoric reference was previously proposed, it is not possible to compare the proposal in this paper to previous ones except on individual examples, as done in the previous section. however, it is possible to show that, unlike the existing and partial solutions discussed in the previous section, the generalized metrics proposed here can be used to compare anaphoric resolvers in exactly the same way as done using the existing coreference metrics. our generalized metrics were incorporated in the new universal anaphora scorer, a new scorer for anaphora that can also score split antecedent anaphora resolution (as well as non-referring mentions identification, bridging reference resolution and discourse deixis resolution) (yu et al., 2022b, 2023a), and this scorer was then used to score the performance of a state of the art system able to carry out split-antecedent anaphora resolution on real data (yu et al., 2021) and to compare it with simpler baselines, both on the corpus used by yu et al. (2021) and on the corpus used by zhou and choi (2018). in this section we present the results of this evaluation. split-antecedent anaphoric references are rarer compared to single entity anaphoric references, at least for the case of plural split antecedent references (split discourse deictic references are much more common). as a result, they typically do not make a significant contribution to the overall evaluation score. thus, to offer a clear picture of the performance of a system on split-antecedents, our scorer also reports separate scores for the split-antecedent references only. the scores for the separate evaluation of split-antecedents are the micro-average f1 of all the aligned gold and system pairs. those accommodated sets for which an alignment could not be found (e.g., missing or spurious accommodated sets ) were paired with empty sets when computing the scores. 7.1 the models being compared in order to evaluate the generalized coreference metrics on the output of a real system, we obtained the best-performing output and all the baselines from the system by yu et al. (2021), the only mod33 paun, yu, moosavi and poesio muc b3 ceafe conll r p f1 r p f1 r p f1 f1 recent 2 76.6 77.3 76.9 78.8 76.1 77.5 78.0 74.3 76.1 76.8 recent 3 76.7 77.3 77.0 78.9 76.2 77.5 78.0 74.4 76.1 76.9 recent 4 76.8 77.3 77.0 79.0 76.2 77.5 78.0 74.4 76.1 76.9 recent 5 76.7 77.3 77.0 79.0 76.2 77.5 78.0 74.3 76.1 76.9 random 76.5 77.1 76.8 78.8 76.0 77.4 77.9 74.2 76.0 76.7 single ant 76.3 78.4 77.4 78.7 76.9 77.8 78.0 74.4 76.2 77.1 yu et al. 77.1 77.9 77.5 79.1 76.5 77.8 78.1 74.5 76.3 77.2 oracle 77.6 78.4 78.0 79.5 76.8 78.1 78.3 74.6 76.4 77.5 (a) muc, b3, ceafe and conll f1. ceafm blanc lea (β=1) lea(β=10) r p f1 r p f1 r p f1 r p f1 recent 2 77.2 75.4 76.3 75.3 71.5 73.4 70.3 66.9 68.6 59.7 61.2 60.5 recent 3 77.3 75.4 76.3 75.3 71.5 73.4 70.4 66.9 68.6 60.3 61.3 60.8 recent 4 77.3 75.4 76.3 75.3 71.6 73.4 70.4 66.9 68.6 60.7 61.5 61.1 recent 5 77.3 75.4 76.3 75.3 71.6 73.4 70.4 66.9 68.6 60.7 61.1 60.9 random 77.2 75.3 76.2 75.3 71.5 73.3 70.3 66.7 68.4 59.3 59.9 59.6 single ant 77.2 75.8 76.5 75.2 72.3 73.7 70.3 67.4 68.8 59.0 67.4 62.9 yu et al. 77.4 75.7 76.6 75.5 71.9 73.6 70.7 67.2 68.9 62.8 64.7 63.7 oracle 77.8 75.9 76.8 75.9 72.1 74.0 71.2 67.6 69.4 65.9 68.4 67.1 (b) ceafm , blanc and lea with different split-antecedent importance (β). table 1: evaluation of yu et al. (2021)’s model and of the baselines on arrau using the new universal anaphora scorer. muc b3 ceafe ceafm blanc lea conll recent 2 25.6 22.2 15.9 23.6 15.8 21.9 21.2 recent 3 27.1 24.2 20.8 26.3 18.2 23.5 24.0 recent 4 28.0 25.2 21.0 27.5 17.5 24.4 24.7 recent 5 26.6 23.6 19.0 26.0 15.9 22.9 23.1 random 19.6 15.3 7.9 16.8 11.4 14.8 14.3 yu et al. 35.8 31.9 37.0 32.8 18.2 30.9 34.9 oracle 70.1 63.7 62.9 68.1 68.6 61.4 65.6 table 2: split-antecedent f1 scores only for yu et al. (2021)’s system on arrau evaluated using our extension of the coreference scorers. ern system that can process both singleand split-antecedent anaphors, and ran our scorer on all six predictions. the results are illustrated in tables 1 -3. yu et al. (2021)’s model is an extension of the system proposed by yu et al. (2020b) which further interprets split-antecedents. the model shares most of the network architecture proposed by yu et al. (2020b), but in addition, it includes a dedicated feed-forward network for split-antecedents. the baselines are based on heuristic rules or random selection. the same candidate split-antecedent anaphors/singular clusters are used in 34 scoring coreference chains with split-antecedent anaphors yu et al. our approach lenientsplit lea (β=1) lea (β=10) leasplit lea (β=1) lea (β=10) recent 2 12.4 68.7 61.4 21.9 68.6 60.5 recent 3 17.0 68.7 61.4 23.5 68.6 60.8 recent 4 16.0 68.7 61.5 24.4 68.6 61.1 recent 5 13.1 68.7 61.3 22.9 68.6 60.9 random 4.5 68.5 60.4 14.8 68.4 59.6 yu et al. 36.4 69.0 64.1 30.9 68.9 63.7 table 3: comparison between our generalized scores and the score proposed in yu et al. (2021) on the yu et al.’s system output evaluated on the arrau corpus. the baselines as in the yu et al. (2021) model. then these baselines attempt to interpret the mentions belonging to a small list of plural pronouns which could be interpreted as split antecedent anaphor (e.g., they, their, them, we) but were classified as not having an atomic antecedent by the single-antecedent anaphoric resolver, and attempt to resolve them as split-antecedent anaphors. the random baseline randomly assigns two to five antecedents to these candidate split antecedent anaphors. after that, the recent-x baseline uses the x closest singular clusters as antecedents to each chosen anaphor. in addition, we also include scores for the system only resolving single antecedent anaphors (single ant) and system augmented with the gold split-antecedent anaphors (oracle)19. 7.2 evaluation on arrau table 1 shows the overall scores (single and split-antecedent anaphors) on arrau evaluated using the new scorer.20 the general direction of the results doesn’t change from those reported in yu et al. (2021) (see below). what changes is that whereas in yu et al. (2021) the system’s performance on split antecedents could only be evaluated using a single, ad-hoc metric (see, e.g., tables 3 and 5 in that paper), thanks to the extension proposed in this paper it is now possible to score both single-antecedent and split-antecedent anaphors using the same metrics. the results with all the metrics confirm the results obtained by yu et al. with their specialized metric. first of all, as already observed by yu et al., the difference between the baselines and the best model is small when singleand split-antecedent anaphors are evaluated together, because the number of split-antecedents is low (only 0.8% of the clusters containing split-antecedents). however, the difference becomes very clear when considering the performance on split antecedents only (see table 2): up to 20 percentage points in conll score. the oracle setting has again much better split-antecedent f1 scores, and this results in considerable improvements on all the scores when evaluated with singular clusters. confirming our hypothesis on the need for partial credit on split-antecedents, we find that only 16% of the split-antecedents were fully resolved by the yu et al. model. the vast majority of the split-antecedents that were partially resolved rely on the partial credit to get a fair assessment. 19. for oracle setting we allow the system to use the gold split-antecedent anaphors annotations when possible. please note the system is still constrained by the quality of singular clusters. this simulates a better system on resolving the split-antecedent anaphors. 20. following yu et al. (2021), we only report scores for documents in which at least one split-antecedent anaphor is annotated. 35 paun, yu, moosavi and poesio we further compare our new and previous approaches (yu et al., 2021) in table 3. since yu et al. only generalised the lea score, the comparison is primarily based on lea. for each approach, lea scores were computed with two different split-antecedent importance (β ∈ {1, 10}). a large β means more weight is given to the split-antecedents. we additionally include the most relevant split-antecedent-only score (lenientsplit and leasplit for yu et al. and our approach respectively) to assess the correlation between the split-antecedent-only scores and the lea scores. a s we can see from the table, when β = 1 both approaches successfully register a difference between the yu et al. and the baselines, but do not show a visible difference between the baselines apart from the ‘random’ setting that has a lower score overall. this is not surprising given that the number of splitantecedent anaphors is small. when we increase the split-antecedent importance (i.e. β = 10) the score differences become more visible; however, we noticed that the lea score from our previous approach does not correlate well with the split-antecedent only score. even though the ‘recent 3’ baseline has a higher lenientsplit than ‘recent 2’ and ‘recent 4’, it has the same lea score (61.4%) as ‘recent 2’ and lower than ‘recent 4’ ’s 61.5%. on the other hand, the new approach reports lea scores that follow the same trend as its split-antecedent scores (i.e. leasplit), which makes the evaluation more consistent between the split-antecedent only and overall scores. 7.3 evaluation on friends in order to compare our metrics in practice with the evaluation approach proposed by zhou and choi (2018), we further tested our extended metrics on the friends corpus used in their paper, which contains a larger percentage of split-antecedent plural references. we follow zhou and choi (2018) in using episodes 1 19 for training, 20, 21 for development and 22, 23 for testing. the original corpus is annotated for entity linking, so the coreference clusters are created by grouping the mentions that refer to the same entity (a character in the show friends). 14.5% of those clusters contain split-antecedents. in contrast, in the original annotation, split-antecedent anaphors represent 9% of all mentions; this is because in the original version of the friends corpus all subsequent mentions of a set accommodated using a split-antecedent anaphor are also marked as split-antecedents instead of being coreferent with the first mention. (e.g. in our illustrative example (12), [they7] and [the two10] are treated as single-antecedent by putting them in the same cluster as [their6], whereas in the friends corpus would be annotated as split-antecedent references as well.) we transformed all of these cases into single-antecedent anaphors; after this transformation, 4.1% of the mentions remain split-antecedent anaphors. to obtain the system predictions, we trained the yu et al. (2021) system on the friends corpus and computed all the baselines in the same way as for the arrau corpus.21 as shown in table 4, the yu et al. model outperforms the baselines by a large margin according to all the metrics even though the performance improvements on split-antecedent anaphors (see table 5) are smaller than those we observed with the arrau evaluation. this was expected, as the friends corpus contains many more split-antecedent anaphors. when comparing the yu et al. model with the single-antecedent only system (single ant), the model has a better recall but a lower precision, overall having similar f1 scores for most of the matrices. this is because system performance on the split-antecedent part is not good enough to make a clear difference. with the oracle setting, however , the better performance on split-antecedent anaphors contributed to a robust improvement 21. we contacted the authors of zhou and choi (2018) to obtain their system’s outputs but did not get a reply. it should also be noted that they evaluated their system in a gold mention setting which is not realistic. 36 scoring coreference chains with split-antecedent anaphors muc b3 ceafe conll r p f1 r p f1 r p f1 f1 recent 2 80.0 80.4 80.2 71.0 72.0 71.5 59.1 61.7 60.4 70.7 recent 3 80.0 80.6 80.3 71.1 72.1 71.6 59.3 61.6 60.4 70.8 recent 4 80.1 81.0 80.6 71.1 72.5 71.8 59.9 61.9 60.9 71.1 recent 5 79.6 81.3 80.4 70.6 72.9 71.7 60.2 61.8 61.0 71.1 random 79.8 79.9 79.9 71.0 71.4 71.2 60.6 61.3 61.0 70.7 single ant 77.9 85.0 81.3 69.1 77.0 72.8 64.0 63.0 63.5 72.6 yu et al. 80.7 82.2 81.4 71.7 73.5 72.6 63.6 64.8 64.2 72.7 oracle 81.5 83.9 82.7 72.5 75.3 73.9 66.5 65.3 65.9 74.2 (a) muc, b3, ceafe and conll f1. ceafm blanc lea (β=1) lea(β=10) r p f1 r p f1 r p f1 r p f1 recent 2 70.8 71.8 71.3 73.1 75.9 74.5 60.9 62.0 61.4 43.1 42.8 42.9 recent 3 70.8 71.9 71.3 73.3 76.2 74.7 61.1 62.0 61.5 43.4 42.6 43.0 recent 4 70.8 72.2 71.5 73.3 76.5 74.9 61.3 62.3 61.8 43.5 43.8 43.7 recent 5 70.5 72.3 71.4 73.2 76.9 75.0 60.9 62.5 61.7 40.9 43.9 42.3 random 71.1 71.5 71.3 73.5 75.7 74.6 60.8 61.1 61.0 42.5 39.8 41.1 single ant 70.4 74.7 72.5 72.4 80.1 76.0 60.5 65.5 62.9 34.8 65.5 45.5 yu et al. 72.0 73.3 72.6 73.6 76.8 75.2 63.0 64.2 63.6 47.6 50.5 49.0 oracle 73.0 74.1 73.6 74.6 78.1 76.3 64.6 66.2 65.4 53.4 61.2 57.1 (b) ceafm , blanc and lea with different split-antecedent importance (β). table 4: evaluation of yu et al. (2021)’s model and of the baselines on the friends corpus using the new universal anaphora scorer (yu et al., 2022b, 2023a). muc b3 ceafe ceafm blanc lea conll recent 2 43.4 37.5 30.9 42.7 28.3 36.1 37.3 recent 3 45.1 38.9 32.0 43.8 30.7 37.3 38.6 recent 4 44.0 37.7 31.2 42.7 30.2 36.0 37.6 recent 5 42.8 36.3 28.5 41.4 30.6 34.5 35.9 random 43.5 36.9 34.2 42.2 31.1 35.3 38.2 yu et al. 52.3 44.9 43.3 50.8 37.8 43.3 46.8 oracle 65.5 57.1 65.6 65.1 50.9 54.8 62.7 table 5: split-antecedent f1 scores only for yu et al.’s systems evaluated on the friends corpus using our generalization of the coreference scorers. on overall performance on both singleand split-antecedent anaphors. this indicates a better splitantecedent anaphora resolver is needed to achieve a significant improvement when compared with systems that only resolve single-antecedent anaphors. if split-antecedent anaphors are the main focus of the evaluation, one can use the split-antecedent f1 scores to look at the split-antecedent only scores (see table 5) or the lea metric with appropriate split-antecedent importance (e.g. β = 10) to prioritise split-antecedent anaphors. 37 paun, yu, moosavi and poesio zhou&choi our approach b3 ceafe blanc b3 ceafe blanc recent 2 71.2 53.1 76.5 71.5 60.4 74.5 recent 3 71.2 53.1 76.4 71.6 60.4 74.7 recent 4 71.2 52.9 76.3 71.8 60.9 74.9 recent 5 70.8 52.4 76.0 71.7 61.0 75.0 random 70.7 51.8 75.8 71.2 61.0 74.6 sing ant 70.2 54.1 75.0 72.8 63.5 76.0 yu et al. 74.2 59.1 79.0 72.6 64.2 75.2 oracle 75.9 62.6 80.3 73.9 65.9 76.3 gold sing ant 89.3 82.2 89.1 96.1 95.0 97.0 table 6: comparison between our generalization of the coreference scorers and the zhou and choi (2018) method on the yu et al.’s system output evaluated on the friends corpus. in table 6 we compare in detail our scorer with that of zhou and choi (2018). we use the scorer developed by zhou and choi to score the baselines and system output and compare all three metrics proposed by them with our equivalents. as we can see from the table, the main difference is that the zhou and choi approach heavily penalises systems that do not predict split-antecedents: the ‘sing ant’ system that only predicts references to coreference chains without accommodated antecedents scores lower than the baselines in most of the cases. we suggest this is a result of the decision to put the split-antecedent anaphors into the relevant singular clusters. by doing so, the plural mentions are credited multiple times in the evaluation, hence the split-antecedent anaphors gain more weight than normal mentions. to give a clearer picture of the scorer’s behaviour on penalising the missing split-antecedents, we also provide in table 6 the scores on evaluating the gold singular clusters (gold sing ant). you can see this as a system that predicts all singular clusters correctly and only misses the links between the split-antecedent anaphors with their antecedents. given that about 4% of mentions in the corpus are split-antecedent references, we would expect the scores to be close to 96%. the zhou and choi scorer, however, penalises the system by 10% 18%, way beyond the credit that should be assigned to split-antecedent anaphors. in contrast, our scorer gives scores between 95% 97% ,which is close to the expectation. although in some circumstances one might want to give more credit to split-antecedent anaphors if this is the main focus, it might not be a good idea to assume this is the mainstream. instead of permanently boosting the importance of splitantecedent anaphors as done in zhou and choi, we posit that it would be better to allow users (e.g. the shared task organizers) to decide how important should the split-antecedent anaphors be, as we have done in lea, using the split-antecedent importance parameter (β) to configure the importance of split-antecedent plurals flexibly. 8. scoring other types of split-antecedent anaphora and of anaphora involving accommodation as discussed in the introduction and in section 2, split-antecedent plurals are just one example of anaphoric reference referring to an entity which wasn’t previously mentioned, thus requiring accommodation of a new antecedent (beaver and zeevat, 2007; van der sandt, 1992). in all of these 38 scoring coreference chains with split-antecedent anaphors cases, the new entity is composed of a part constructed out of the pre-existing discourse model, together with a ‘coreference chain’ part–i.e., the structure proposed here for split antecedent plurals: ki = ko i ⊕km i the difference is the relation linking the new entity to existing entities in the context. for split antecedents, this relation is set membership: the new set is the set of the entities mentioned by the split antecedents. this type of accommodation is required not just for plurals, but for discourse deixis as well. in the case of bridging references, the relation is associative, not coreference. in the case of context change accommodation, the new entity is typically the result of an action carried out over the entities in the context. so, the proposed notation could potentially also serve as the basis for extensions covering these cases. split-antecedent discourse deixis as discussed in section 2, discourse deixis (webber, 1991; kolhatkar et al., 2018), may also involve split-antecedents, as shown by example (3). we also mentioned there that in universal anaphora, discourse deixis is treated following the approach that has become standard for event anaphora (lu and ng, 2018)–i.e., the interpretation of discourse deictic references is scored using the same metrics developed for the interpretation of identity reference (yu et al., 2022b). (discourse deictic references are marked in a separate discourse deixis layer, but this layer differs from the identity reference only in that the antecedents of discourse deixis are utterances rather than nominal phrases.) the extension of these metrics proposed in this paper can therefore be immediately used for cases of split-antecedent discourse deixis such as (3). and indeed, this approach was used to score discourse deixis, including split-antecedent discourse deixis, in the codi/crac shared tasks (khosla et al., 2021; yu et al., 2022a). context change accomodation a more complex case of split-antecedent anaphoric reference are the cases of context change accommodation discussed by webber and baldwin (1992), where a new entity, the dough, is obtained by mixing together flour and water. (13) add [the water]i to [the flour]j little by little. then work [the dough]i+j context change accommodation was not annotated in any of the best-known datasets for anaphoric reference discussed in section 2, so no evaluation method was proposed to our knowledge. however, as mentioned earlier, the community has started to study context-change accommodation again recenly, and it has been annotated in chemu-ref (fang et al., 2021) and reciperef (fang et al., 2022). in these corpora, context-change anaphoric references are treated as cases of bridging references, and systems interpreting them have been evaluated in this way. this approach suffers however from the same problem as treating split-antecedent plural references as cases of bridging discussed in section 2.3–namely, that it fails to capture the fact that the water and the flour are all of the antecedents of ‘the dough,’ and therefore there is no way of rewarding a system for identifying all antecedents, or penalizing a system for recovering only some of them. using the approach proposed in this paper would better reflect the semantics of these cases, although some work is needed to work out the implications of the fact that the water and the flour are not simply elements of a set, but the two components of a piece of matter. 39 paun, yu, moosavi and poesio 9. conclusions in order to push forward the state of the art in anaphora resolution beyond the simplest form of identity anaphora it is not sufficient to create suitable datasets annotated with the more general cases of anaphora, although that is an important effort. it is also necessary to develop methods for evaluating the performance of anaphoric resolvers on these cases. in this paper we proposed a method for evaluating one of these more general cases–the case of anaphoric reference to entities that need introducing in a discourse model via accomodation, exemplified by split-antecedent anaphors but including other cases as well, such as discourse deixis–that is a straightforward extension of existing proposals for coreference evaluation and thus does not require introducing additional metrics, an issue in a field already over-rich with proposals in this direction. acknowledgments this research was supported in part by the dali project, erc grant 695662, and in part by the arciduca project, epsrc grant number ep/w001632/1. references ron artstein and massimo poesio. identifying reference to abstract objects in dialogue. in brandial 2006: proceedings of the 10th workshop on the semantics and pragmatics of dialogue (semdial-10), potsdam, 2006. nicholas asher. reference to abstract objects in english. d. reidel, dordrecht, 1993. amit bagga and breck baldwin. algorithms for scoring coreference chains. in workshop on linguistics coreference at the first international conference on language resources and evaluation (lrec), volume 1, pages 563–566. association for computational linguistics, 1998. satanjeev banerjee and alon lavie. meteor: an automatic metric for mt evaluation with improved correlation with human judgments. in proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65– 72, ann arbor, michigan, 2005. association for computational linguistics. url https: //aclanthology.org/w05-0909. david beaver and henk zeevat. accommodation. in g. ramchand and c. reiss, editors, the handbook of linguistic interfaces, pages 503–536. oxford, 2007. ezra black, steve abney, dan flickinger, c. gdaniec, ralph grishman, paul harrison, don hindle, rob ingria, fred jelinek, judith klavans, mark liberman, mitch marcus, s. roukos, beatrice santorini, and tomek strzalkowski. a procedure for quantitatively comparing the syntactic coverage of english grammars. in speech and natural language: proceedings of a workshop held at pacific grove, california, february 19-22, 1991. association for computational linguistics, 1991. url https://aclanthology.org/h91-1060. donna k. byron. resolving pronominal reference to abstract entities. in proceedings of the 40th annual meeting of the association for computational linguistics, pages 80–87, philadelphia, 40 https://aclanthology.org/w05-0909 https://aclanthology.org/w05-0909 https://aclanthology.org/h91-1060 scoring coreference chains with split-antecedent anaphors pennsylvania, usa, 2002. association for computational linguistics. doi: 10.3115/1073083. 1073099. url https://aclanthology.org/p02-1011. david m. carter. interpreting anaphors in natural language texts. ellis horwood, chichester, uk, 1987. nancy a. chinchor and beth sundheim. message understanding conference (muc) tests of discourse processing. in proceedings of the aaai spring symposium on empirical methods in discourse interpretation and generation, pages 21–26, stanford, 1995. herbert h. clark. bridging. in theoretical issues in natural language processing, 1975. kevin clark and christopher d. manning. improving coreference resolution by learning entitylevel distributed representations. in proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), pages 643–653, berlin, germany, 2016. association for computational linguistics. doi: 10.18653/v1/p16-1061. url https://www. aclweb.org/anthology/p16-1061. francis cornish. discourse anaphora. in keith brown, editor, encyclopedia of language and linguistics, pages 631–638. oxford university press, 2nd edition, 2006. pascal denis and jason baldridge. joint determination of anaphoricity and coreference resolution using integer programming. in human language technologies 2007: the conference of the north american chapter of the association for computational linguistics; proceedings of the main conference, pages 236–243, rochester, new york, 2007. association for computational linguistics. url https://aclanthology.org/n07-1030. carola eschenbach, christopher habel, michael herweg, and klaus rehkämper. remarks on plural anaphora. in proceedings of the 4th conference of the european chapter of the association for computational linguistics, pages 161–167. association for computational linguistics, 1989. biaoyan fang, christian druckenbrodt, saber a akhondi, jiayuan he, timothy baldwin, and karin verspoor. chemu-ref: a corpus for modeling anaphora resolution in the chemical domain. in proceedings of the 16th conference of the european chapter of the association for computational linguistics (eacl), page 1362–1375. association for computational linguistics, 2021. doi: 10.18653/v1/2021.eacl-main.116. url https://aclanthology.org/ 2021.eacl-main.116. biaoyan fang, timothy baldwin, and karin verspoor. what does it take to bake a cake? the reciperef corpus and anaphora resolution in procedural text. in findings of the association for computational linguistics: acl 2022, page 3481–3495. association for computational linguistics, 2022. doi: 10.18653/v1/2022.findings-acl.275. url https://aclanthology.org/ 2022.findings-acl.275. alan garnham. mental models and the interpretation of anaphora. psychology press, 2001. jeanette k. gundel and barbara abbott, editors. the oxford handbook of reference. oxford university press, 2019. 41 https://aclanthology.org/p02-1011 https://www.aclweb.org/anthology/p16-1061 https://www.aclweb.org/anthology/p16-1061 https://aclanthology.org/n07-1030 https://aclanthology.org/2021.eacl-main.116 https://aclanthology.org/2021.eacl-main.116 https://aclanthology.org/2022.findings-acl.275 https://aclanthology.org/2022.findings-acl.275 paun, yu, moosavi and poesio jeanette k. gundel, michael hegarty, and kaja borthen. cognitive status, information structure, and pronominal reference to clausally introduced entities. journal of logic language and information, 12(3):281–299, 2003. yufang hou. bridging anaphora resolution as question answering. in proceedings of the 58th annual meeting of the association for computational linguistics, pages 1428–1438, online, 2020. association for computational linguistics. doi: 10.18653/v1/2020.acl-main.132. url https://www.aclweb.org/anthology/2020.acl-main.132. yufang hou, katja markert, and michael strube. unrestricted bridging resolution. computational linguistics, 44(2):237–284, 2018. mandar joshi, danqi chen, yinhan liu, daniel s. weld, luke zettlemoyer, and omer levy. spanbert: improving pre-training by representing and predicting spans. transactions of the association for computational linguistics, 8:64–77, 2020. doi: 10.1162/tacl a 00300. url https://aclanthology.org/2020.tacl-1.5. tuomo kakkonen. framework and resources for natural language parsing evaluation. phd thesis, university of joensuu, 2007. hans kamp and uwe reyle. from discourse to logic. d. reidel, dordrecht, 1993. ben kantor and amir globerson. coreference resolution with entity equalization. in proceedings of the 57th annual meeting of the association for computational linguistics, pages 673–677, florence, italy, 2019. association for computational linguistics. doi: 10.18653/v1/p19-1066. url https://www.aclweb.org/anthology/p19-1066. sopan khosla, juntao yu, ramesh manuvinakurike, vincent ng, massimo poesio, michael strube, and carolyn rosé. the codi-crac 2021 shared task on anaphora, bridging, and discourse deixis in dialogue. in proceedings of the codi-crac 2021 shared task on anaphora, bridging, and discourse deixis in dialogue, pages 1–15, punta cana, dominican republic, 2021. association for computational linguistics. doi: 10.18653/v1/2021.codi-sharedtask.1. url https://aclanthology.org/2021.codi-sharedtask.1. varada kolhatkar, heike zinsmeister, and graeme hirst. interpreting anaphoric shell nouns using antecedents of cataphoric shell nouns as training data. in proceedings of the 2013 conference on empirical methods in natural language processing, pages 300–310, seattle, washington, usa, 2013. association for computational linguistics. url https://aclanthology. org/d13-1030. varada kolhatkar, adam roussel, stefanie dipper, and heike zinsmeister. anaphora with non-nominal antecedents in computational linguistics: a survey. computational linguistics, 44(3):547–612, 2018. doi: 10.1162/coli a 00327. url https://www.aclweb.org/ anthology/j18-3007. harold w. kuhn. the hungarian method for the assignment problem. naval research logistics quarterly, 2(1-2):83–97, 1955. fred landman. groups i. linguistics and philosophy, pages 559–605, 1989. 42 https://www.aclweb.org/anthology/2020.acl-main.132 https://aclanthology.org/2020.tacl-1.5 https://www.aclweb.org/anthology/p19-1066 https://aclanthology.org/2021.codi-sharedtask.1 https://aclanthology.org/d13-1030 https://aclanthology.org/d13-1030 https://www.aclweb.org/anthology/j18-3007 https://www.aclweb.org/anthology/j18-3007 scoring coreference chains with split-antecedent anaphors kenton lee, luheng he, mike lewis, and luke zettlemoyer. end-to-end neural coreference resolution. in proceedings of the 2017 conference on empirical methods in natural language processing, pages 188–197, copenhagen, denmark, 2017. association for computational linguistics. doi: 10.18653/v1/d17-1018. url https://www.aclweb.org/anthology/ d17-1018. kenton lee, luheng he, and luke zettlemoyer. higher-order coreference resolution with coarse-tofine inference. in proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 2 (short papers), pages 687–692, new orleans, louisiana, 2018. association for computational linguistics. doi: 10.18653/v1/n18-2108. url https://www.aclweb.org/anthology/n18-2108. david k. lewis. scorekeeping in a language game. journal of philosophical logic, 8:339–359, 1979. chin-yew lin. rouge: a package for automatic evaluation of summaries. in workshop on text summarization branches out, pages 74–81, barcelona, spain, 2004. association for computational linguistics. url https://aclanthology.org/w04-1013. godehard link. the logical analysis of plurals and mass terms: a latticetheoretical approach. in r. bäuerle, c. schwarze, and a. von stechow, editors, meaning, use and interpretation of language, pages 302–323. walter de gruyter, 1983. godehard link. hydras: on the logic of relative clause constructions with multiple heads. in f. landman and f. veltman, editors, varieties of formal semantics. foris, dordrecht, 1984. quan liu, hui jiang, andrew evdokimov, zhen-hua ling, xiaodan zhu, si wei, and yu hu. causeeffect knowledge acquisition and neural association model for solving a set of winograd schema problems. in proceedings of the 26th international joint conference on artificial intelligence, ijcai-17, pages 2344–2350, 2017. doi: 10.24963/ijcai.2017/326. url https://doi.org/ 10.24963/ijcai.2017/326. jing lu and vincent ng. event coreference resolution: a survey of two decades of research. in proceedings of the 27th international joint conference on artificial intelligence, ijcai-18, pages 5479–5486. international joint conferences on artificial intelligence organization, 2018. doi: 10.24963/ijcai.2018/773. url https://doi.org/10.24963/ijcai.2018/773. xiaoqiang luo. on coreference resolution performance metrics. in proceedings of human language technology conference and conference on empirical methods in natural language processing, pages 25–32, vancouver, british columbia, canada, 2005. association for computational linguistics. url https://www.aclweb.org/anthology/h05-1004. xiaoqiang luo and sameer pradhan. evaluation metrics. in massimo poesio, roland stuckardt, and yannick versley, editors, anaphora resolution: algorithms, resources and applications, pages 147–170. springer, 2016. xiaoqiang luo, sameer pradhan, marta recasens, and eduard hovy. an extension of blanc to system mentions. in proceedings of the 52nd annual meeting of the association for computational linguistics (volume 2: short papers), pages 24–29, baltimore, maryland, 2014. as43 https://www.aclweb.org/anthology/d17-1018 https://www.aclweb.org/anthology/d17-1018 https://www.aclweb.org/anthology/n18-2108 https://aclanthology.org/w04-1013 https://doi.org/10.24963/ijcai.2017/326 https://doi.org/10.24963/ijcai.2017/326 https://doi.org/10.24963/ijcai.2018/773 https://www.aclweb.org/anthology/h05-1004 paun, yu, moosavi and poesio sociation for computational linguistics. doi: 10.3115/v1/p14-2005. url https://www. aclweb.org/anthology/p14-2005. susan luperfoy. discourse pegs: a computational analysis of context-dependent referring expressions. phd thesis, the university of texas at austin, dept. of linguistics, austin, tx, 1991. john lyons. semantics. cambridge university press, 1977. ana marasović, leo born, juri opitz, and anette frank. a mention-ranking model for abstract anaphora resolution. in proceedings of the 2017 conference on empirical methods in natural language processing, pages 221–232, copenhagen, denmark, 2017. association for computational linguistics. doi: 10.18653/v1/d17-1021. url https://www.aclweb.org/ anthology/d17-1021. katja markert, yufang hou, and michael strube. collective classification for fine-grained information status. in proceedings of the 50th annual meeting of the association for computational linguistics (volume 1: long papers), juju island, korea, 2012. url http://www.aclweb. org/anthology/p12-1084. ruslan mitkov. anaphora resolution. longman, 2002. nafise sadat moosavi and michael strube. which coreference evaluation metric do you trust? a proposal for a link-based entity aware metric. in proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), pages 632–642, berlin, germany, 2016. association for computational linguistics. doi: 10.18653/v1/p16-1060. url https://www.aclweb.org/anthology/p16-1060. james munkres. algorithms for the assignment and transportation problems. journal of the society for industrial and applied mathematics, 5(1):32–38, 1957. anna nedoluzhko. generic noun phrases and annotation of coreference and bridging relations in the prague dependency treebank. in proceedings of the 7th linguistic annotation workshop and interoperability with discourse, pages 103–111, sofia, bulgaria, 2013. association for computational linguistics. url https://aclanthology.org/w13-2313. anna nedoluzhko, michal novák, martin popel, zdeněk žabokrtský, and daniel zeman. coreference in universal dependencies 0.1 (corefud 0.1), 2021. url http://hdl.handle.net/ 11234/1-3510. lindat/clariah-cz digital library at the institute of formal and applied linguistics (úfal), faculty of mathematics and physics, charles university. kishore papineni, salim roukos, todd ward, and wei-jing zhu. bleu: a method for automatic evaluation of machine translation. in proceedings of the 40th annual meeting of the association for computational linguistics, pages 311–318, philadelphia, pennsylvania, usa, 2002. association for computational linguistics. doi: 10.3115/1073083.1073135. url https: //aclanthology.org/p02-1040. rebecca j. passonneau. instructions for applying discourse reference annotation for multiple applications (drama). unpublished manuscript., december 1997. 44 https://www.aclweb.org/anthology/p14-2005 https://www.aclweb.org/anthology/p14-2005 https://www.aclweb.org/anthology/d17-1021 https://www.aclweb.org/anthology/d17-1021 http://www.aclweb.org/anthology/p12-1084 http://www.aclweb.org/anthology/p12-1084 https://www.aclweb.org/anthology/p16-1060 https://aclanthology.org/w13-2313 http://hdl.handle.net/11234/1-3510 http://hdl.handle.net/11234/1-3510 https://aclanthology.org/p02-1040 https://aclanthology.org/p02-1040 scoring coreference chains with split-antecedent anaphors massimo poesio and ron artstein. anaphoric annotation in the arrau corpus. in proceedings of the sixth international conference on language resources and evaluation (lrec’08), marrakech, morocco, 2008. european language resources association (elra). url http: //www.lrec-conf.org/proceedings/lrec2008/pdf/297_paper.pdf. massimo poesio, florence bruneseaux, and laurent romary. the mate meta-scheme for coreference in dialogues in multiple languages. in workshop towards standards and tools for discourse tagging, 1999. url https://aclanthology.org/w99-0309. massimo poesio, sameer pradhan, marta recasens, kepa rodriguez, and yannick versley. annotated corpora and annotation tools. in m. poesio, r. stuckardt, and y. versley, editors, anaphora resolution: algorithms, resources and applications, chapter 4. springer, 2016a. massimo poesio, roland stuckardt, and yannick versley, editors. anaphora resolution. springer berlin heidelberg, 2016b. doi: 10.1007/978-3-662-47909-4. url https://doi.org/10. 1007%2f978-3-662-47909-4. massimo poesio, yulia grishina, varada kolhatkar, nafise moosavi, ina roesiger, adam roussel, fabian simonjetz, alexandra uma, olga uryupina, juntao yu, and heike zinsmeister. anaphora resolution with the arrau corpus. in proceedings of the first workshop on computational models of reference, anaphora and coreference, pages 11–22, new orleans, louisiana, 2018. association for computational linguistics. doi: 10.18653/v1/w18-0702. url https:// www.aclweb.org/anthology/w18-0702. massimo poesio, jon chamberlain, silviu paun, juntao yu, alexandra uma, and udo kruschwitz. a crowdsourced corpus of multiple judgments and disagreement on anaphoric interpretation. in proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 1778–1789, minneapolis, minnesota, 2019. association for computational linguistics. doi: 10.18653/v1/n19-1176. url https://aclanthology.org/n19-1176. massimo poesio, juntao yu, silviu paun, abdulrahman aloraini, pengcheng lu, janosch haber, and derya cokal. computational models of anaphora. annual review of linguistics, 2023. sameer pradhan, alessandro moschitti, nianwen xue, olga uryupina, and yuchen zhang. conll2012 shared task: modeling multilingual unrestricted coreference in ontonotes. in joint conference on emnlp and conll shared task, pages 1–40, jeju island, korea, 2012. association for computational linguistics. url https://aclanthology.org/w12-4501. sameer pradhan, xiaoqiang luo, marta recasens, eduard hovy, vincent ng, and michael strube. scoring coreference partitions of predicted mentions: a reference implementation. in proceedings of the 52nd annual meeting of the association for computational linguistics (volume 2: short papers), pages 30–35, baltimore, maryland, 2014. association for computational linguistics. doi: 10.3115/v1/p14-2006. url https://www.aclweb.org/anthology/ p14-2006. altaf rahman and vincent ng. resolving complex cases of definite pronouns: the winograd schema challenge. in proceedings of the 2012 joint conference on empirical methods in natural 45 http://www.lrec-conf.org/proceedings/lrec2008/pdf/297_paper.pdf http://www.lrec-conf.org/proceedings/lrec2008/pdf/297_paper.pdf https://aclanthology.org/w99-0309 https://doi.org/10.1007%2f978-3-662-47909-4 https://doi.org/10.1007%2f978-3-662-47909-4 https://www.aclweb.org/anthology/w18-0702 https://www.aclweb.org/anthology/w18-0702 https://aclanthology.org/n19-1176 https://aclanthology.org/w12-4501 https://www.aclweb.org/anthology/p14-2006 https://www.aclweb.org/anthology/p14-2006 paun, yu, moosavi and poesio language processing and computational natural language learning, pages 777–789, jeju island, korea, 2012. association for computational linguistics. url https://www.aclweb. org/anthology/d12-1071. william m. rand. objective criteria for the evaluation of clustering methods. journal of the american statistical association, 66(336):846—-850, 1971. doi: 10.2307/2284239. marta recasens and eduard hovy. blanc: implementing the rand index for coreference evaluation. natural language engineering, 17(4):485–510, 2011. marta recasens and m. antònia martı́. ancora-co: coreferentially annotated corpora for spanish and catalan. language resources and evaluation, 44(4):315–345, 2010. keisuke sakaguchi, ronan le bras, chandra bhagavatula, and yejin choi. winogrande: an adversarial winograd schema challenge at scale. proceedings of the aaai conference on artificial intelligence, 34:8732–8740, 2020. doi: 10.1609/aaai.v34i05.6399. url https: //ojs.aaai.org/index.php/aaai/article/view/6399. roger s. schwarzschild. pluralities. kluwer, 1996. olga uryupina, ron artstein, antonella bristot, federica cavicchio, francesca delogu, kepa j. rodriguez, and massimo poesio. annotating a broad range of anaphoric phenomena, in a variety of genres: the arrau corpus. journal of natural language engineering, 2020. hardik vala, andrew piper, and derek ruths. the more antecedents, the merrier: resolving multi-antecedent anaphors. in proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), pages 2287–2296, berlin, germany, 2016. association for computational linguistics. doi: 10.18653/v1/p16-1216. url https://www.aclweb.org/anthology/p16-1216. kees van deemter and rodger kibble. on coreferring: coreference in muc and related annotation schemes. computational linguistics, 26(4):629–637, 2000. url https://aclanthology. org/j00-4005. rob a. van der sandt. presupposition projection as anaphora resolution. journal of semantics, 9 (4):333–377, 1992. yannick versley. vagueness and referential ambiguity in a large-scale annotated corpus. research on language and computation, 6:333–353, 2008. marc vilain, john burger, john aberdeen, dennis connolly, and lynette hirschman. a modeltheoretic coreference scoring scheme. in proceedings of the sixth message understanding conference (muc-6), 1995. url https://www.aclweb.org/anthology/m95-1005. bonnie l. webber. a formal approach to discourse anaphora. garland, new york, 1979. bonnie l. webber. structure and ostension in the interpretation of discourse deixis. language and cognitive processes, 6(2):107–135, 1991. 46 https://www.aclweb.org/anthology/d12-1071 https://www.aclweb.org/anthology/d12-1071 https://ojs.aaai.org/index.php/aaai/article/view/6399 https://ojs.aaai.org/index.php/aaai/article/view/6399 https://www.aclweb.org/anthology/p16-1216 https://aclanthology.org/j00-4005 https://aclanthology.org/j00-4005 https://www.aclweb.org/anthology/m95-1005 scoring coreference chains with split-antecedent anaphors bonnie lynn webber and breck baldwin. accommodating context change. in 30th annual meeting of the association for computational linguistics, pages 96–103, newark, delaware, usa, 1992. association for computational linguistics. doi: 10.3115/981967.981980. url https:// aclanthology.org/p92-1013. kellie webster, marta recasens, vera axelrod, and jason baldridge. mind the gap: a balanced corpus of gendered ambiguous pronouns. transactions of the association for computational linguistics, 6:605–617, 2018. doi: 10.1162/tacl a 00240. url https://www.aclweb. org/anthology/q18-1042. yoad winter and remko scha. plurals. in shalom lappin and chris fox, editors, the handbook of semantic theory, chapter 3, pages 77–113. wiley blackwell, 2nd edition, 2015. sam wiseman, alexander m. rush, and stuart m. shieber. learning global features for coreference resolution. in proceedings of the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 994–1004, san diego, california, 2016. association for computational linguistics. doi: 10.18653/v1/n16-1114. url https://www.aclweb.org/anthology/n16-1114. juntao yu and massimo poesio. multitask learning based neural bridging reference resolution. in proceedings of the 28th international conference on computational linguistics, pages 3534– 3546, barcelona, spain (online), 2020. international committee on computational linguistics. doi: 10.18653/v1/2020.coling-main.315. url https://www.aclweb.org/anthology/ 2020.coling-main.315. juntao yu, nafise sadat moosavi, silviu paun, and massimo poesio. free the plural: unrestricted split-antecedent anaphora resolution. in proceedings of the 28th international conference on computational linguistics, pages 6113–6125, barcelona, spain (online), 2020a. international committee on computational linguistics. doi: 10.18653/v1/2020.coling-main.538. url https://www.aclweb.org/anthology/2020.coling-main.538. juntao yu, alexandra uma, and massimo poesio. a cluster ranking model for full anaphora resolution. in proceedings of the 12th language resources and evaluation conference, pages 11–20, marseille, france, 2020b. european language resources association. url https: //www.aclweb.org/anthology/2020.lrec-1.2. juntao yu, nafise sadat moosavi, silviu paun, and massimo poesio. stay together: a system for single and split-antecedent anaphora resolution. in proceedings of the 2021 conference of the north american chapter of the association for computational linguistics: human language technologies. association for computational linguistics, 2021. url https://arxiv.org/ abs/2104.05320. juntao yu, sopan khosla, ramesh manuvinakurike, lori levin, vincent ng, massimo poesio, michael strube, and carolyn rosé. the codi-crac 2022 shared task on anaphora, bridging, and discourse deixis in dialogue. in proceedings of the codi-crac 2022 shared task on anaphora, bridging, and discourse deixis in dialogue, pages 1–14, gyeongju, republic of korea, 2022a. association for computational linguistics. url https://aclanthology. org/2022.codi-crac.1. 47 https://aclanthology.org/p92-1013 https://aclanthology.org/p92-1013 https://www.aclweb.org/anthology/q18-1042 https://www.aclweb.org/anthology/q18-1042 https://www.aclweb.org/anthology/n16-1114 https://www.aclweb.org/anthology/2020.coling-main.315 https://www.aclweb.org/anthology/2020.coling-main.315 https://www.aclweb.org/anthology/2020.coling-main.538 https://www.aclweb.org/anthology/2020.lrec-1.2 https://www.aclweb.org/anthology/2020.lrec-1.2 https://arxiv.org/abs/2104.05320 https://arxiv.org/abs/2104.05320 https://aclanthology.org/2022.codi-crac.1 https://aclanthology.org/2022.codi-crac.1 paun, yu, moosavi and poesio juntao yu, sopan khosla, nafise sadat moosavi, silviu paun, sameer pradhan, and massimo poesio. the universal anaphora scorer. in proceedings of the thirteenth language resources and evaluation conference, pages 4873–4883, marseille, france, 2022b. european language resources association. url https://aclanthology.org/2022.lrec-1.521. juntao yu, michal novák, abdulrahman aloraini, nafise sadat moosavi, silviu paun, sameer pradhan, and massimo poesio. the universal anaphora scorer 2.0. in proceedings of the international workshop on computational semantics (iwcs), 2023a. juntao yu, silviu paun, maris camilleri, paloma carretero garcia, jon chamberlain, udo kruschwitz, and massimo poesio. aggregating crowdsourced and automatic judgments to scale up a corpus of anaphoric reference for fiction and wikipedia texts. in proceedings of the 17th conference of the european chapter of the association for computational linguistics (eacl), page 767–781, dubrovnik, croatia, 2023b. association for computational linguistics (acl). url https://aclanthology.org/2023.eacl-main.54. amir zeldes. the gum corpus: creating multilayer resources in the classroom. language resources and evaluation, 51(3):581–612, 2017. doi: http://dx.doi.org/10.1007/ s10579-016-9343-x. amir zeldes. can we fix the scope for coreference? dialogue and discourse, 13(1):41–62, 2022. doi: https://doi.org/10.5210/dad.2022.102. url https://journals.uic.edu/ ojs/index.php/dad/article/view/11706. ethan zhou and jinho d. choi. they exist! introducing plural mentions to coreference resolution and entity linking. in proceedings of the 27th international conference on computational linguistics, pages 24–34, santa fe, new mexico, usa, 2018. association for computational linguistics. url https://www.aclweb.org/anthology/c18-1003. 48 https://aclanthology.org/2022.lrec-1.521 https://aclanthology.org/2023.eacl-main.54 https://journals.uic.edu/ojs/index.php/dad/article/view/11706 https://journals.uic.edu/ojs/index.php/dad/article/view/11706 https://www.aclweb.org/anthology/c18-1003 introduction anaphoric phenomena currently in the scope of the universal anaphora initiative (identity) anaphora and coreference beyond identity anaphora: the universal anaphora /corefud specification of anaphoric phenomena anaphoric references requiring accommodation metrics for scoring identity anaphora (coreference) notation the standard metrics for anaphora / coreference resolution evaluation standard muc standard b3 standard ceaf standard lea standard blanc the case for diversity in anaphora / coreference resolution evaluation generalizing standard metrics to allow for split-antecedent references: key ideas and terminology accommodated sets aligning accommodated sets the term the (artificial) illustrative example, revisited generalized definitions for the coreference metrics generalized b3 definition computing generalized b3 on the example generalized muc definition computing generalized muc on the example generalized ceaf definition computing generalized ceaf on the example generalized lea definition computing generalized lea on the example generalized blanc definition computing generalized blanc on the example alternative proposals to score split-antecedent anaphora using the metrics for scoring generalized anaphoric reference the models being compared evaluation on arrau evaluation on friends scoring other types of split-antecedent anaphora and of anaphora involving accommodation conclusions rieser-dnd24-han dialogue & discourse 15(2) 36-84 doi: 10.5210/dad.2024.202 ©2024 hannes rieser this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). multi-modal anaphora and broadcasting of information by gestural post-holds hannes rieser hannes.rieser@uni-bielefeld.de bielefeld university faculty of linguistics and literary studies germany editor: jonathan ginzburg submitted 04/2022; accepted 08/2024; published online 09/2024 abstract this paper deals with verbal anaphora, multi-modal anaphora and the top-down broadcasting of information using gestural post-holds in multimodal dialogue. first, a new account of definite, pronominal and pro-adverbial anaphora is given. the approach is then extended to multimodal anaphora where part or all of an anaphor’s meaning is contributed by iconic or deictic gestures. as tradition has it, anaphora work “bottom-up”; making use of process algebra, here an inverse relation “broadcasting” is defined, where information is distributed top down and input to receiving ports. anaphora are modelled with this relation; the same is attempted for dialogue coherence relations which generally exhibit anaphora-like behaviour. as video data show, broadcasting is tied to gestural post-holds, the holding of a gesture’s stroke over a stretch of discourse, independently of alignable speech. multi-modal anaphora and broadcasting cross single contributions and turns. this leads to considering post-holds from a new perspective, stressing their speech-independent function and their relevance for indicating topic-continuity. the data come from the saga (the speech and gesture alignment) corpus, a set of route-description dialogues generated in a virtual reality-setting. the calculus used to model the anaphora and broadcasting dynamics is the λψ-calculus, a specifically developed process algebra. it uses the ψ-calculus for input-output, data transport and broadcasting. extending the ψ-calculus, the data transported are typed λ-calculus formulae equipped with compositional links; they can be verbal only, gestural only or multi-modal. following the y-calculus, information chunks are modelled as communicating agents sending or receiving information. the λψ-calculus is also used for the fusion component unifying gestural and verbal information; hence, the paper is also an independent contribution to multi-modal fusion. keywords: anaphora, iconic gestures, gestural post-holds, broadcasting, multimodal fusion, multimodal meaning, ψ-calculus, λψ-calculus 1 introduction multi-modal anaphora is a difficult topic to treat and so is the evaluation of gestural post-holds1 and the modelling of their meaning. both are modelled using information transfer, i.e., broadcasting of information. broadcasting of information in discourse and its seemingly unrestricted distribution top-down might be unfamiliar to researchers in semantics, pragmatics and gesture semiotics. these 1 mcneill (1992, p. 83 and pp. 288-89) also deals with post-stroke holds but tries to harmonise them with his synchrony concept. multi-modal anaphora and broadcasting of information 37 themes are tightly bound up with multi-modal fusion, the idea, especially entertained in ai (see the literature review, sec. 3), that information from different sources, linguistic, gestural, facial or body posture are brought together in one representation. we discuss and model the fusion of speech and aligned iconic gesture using a passage from a route-description dialogue. it is taken from the bielefeld multi-modal saga corpus and provides a wealth of data on (multi-modal) anaphora, twohanded coordinated gestures, gestural post-holds and broadcasting (top down information passing). the data discussion already provides an idea of which properties an algorithm must have to simulate these. these properties are independence of gesture and speech, concurrency, temporal asynchrony of gesture and speech, and composition of gesture semantics and speech semantics. the discussion of the literature, also tailored to the datum, shows that coherence-bound linguistic anaphora theories in the hobbs-hume-kehler-rohde tradition are important as are drt and sdrt accounts of coherence relations. the multimodal fusion mechanism developed is set up against recent multimodal fusion approaches in ai. the afore-mentioned characteristics motivate the calculus we introduce, the ψ-calculus, a powerful extension of milner’s π-calculus (milner, 1999, parrow, 2001). the ψ-calculus has agents as basic entities and works with channels, input-output facilities, information transfer and operators for concurrency, replication of information as well as operators from standard logics. the agents can be seen as monads able to do internal computation and to communicate information. they can embody expressions of higher order logics and transport these via output and input channels. the formulas transported are standard formulas of a typed λcalculus. therefor the two-tiered machinery is called λψ-calculus. the λ-calculus formulas are static, the dynamics rests with the ψ-calculus component. both rely on top-down information processing leading to a new mechanism for anaphora resolution and confirming the independence of gesture information and its role as a topic holder. due to the focus on anaphora and broadcasting there is not much on dialogue structure in this paper. in the appendices it is shown how the mechanism used for anaphora resolution and broadcasting can be used to model rhetorical relations and anaphora resolution in ptt (poesio and traum 1997, poesio and rieser 2010). due to the main topic of the paper several issues are treated with lesser care such as the characteristics of german spoken language syntax, turn-taking regularities, and dialogue update; see however sec. 5.4 on follower’s semantics, follower’s clarification request and follower’s decision for an anaphora antecedent2. the paper extends research on multi-modality elaborated in rieser (2015, 2017), lawler, hahn, and rieser (2017), and, rieser and lawler (2020). it offers a new approach to verbal anaphora taking up a suggestion due to chastain (1975), multi-modal anaphora using a multi-modal fusion account, treats two-handed concurrent gestures and provides a reconstruction of the semantics of gestural holds across contributions and turns. the algorithmic background for that is the λψcalculus. sec. 2 provides the annotated data, a passage from a virtual reality (vr) route description dialogue and an analysis of which algorithmic devices will be needed for its description, an issue taken up in more detail in sec. 4. sec. 3 discusses some of the literature relevant for the anaphora relations in the datum, work in the hobbs-hume-kehler-rohde research line, drt, sdrt, and definite anaphora. a short sub-section deals with speech-gesture alignment. multi-modal fusion approaches integrating speech semantics and information from other modalities are briefly commented upon. against this background, coherence relations, anaphora and multi-modal anaphora in the datum receive a first scrutiny. the literature review section closes with remarks on methodology and the contributions the paper provides. sec. 4 gives the process algebra set-up, i.e. the standard ψcalculus definitions, and introduces the λψ-calculus. this done, we informally map the capabilities of the λψ-calculus onto the structures of the datum outlined in sec. 2. the discussion of the route 2 given the technical possibilities of the ly-calculus it is quite clear how turn-taking and semantic update could be treated, the first by channel communication and the latter by “upwards” recursive processes. rieser 38 description data and their formal rendering in ly comes in sec. 5. anaphora and broadcasting of information are described in some detail, identifying central paradigms of gesture-speech cooperation. finally, sec. 6 offers with conclusions and suggestions for further research. the appendices contain ly-deductions for the central datum and suggest how the λψ-account can be extended to capture rhetorical relations in sdrt and anaphora resolution in ptt. 2. datum and annotation the section of dialogue serving as our main example comes from the saga corpus (see lücking et al., 2013), dialogue v5. the corpus is a collection of 25 vr route-description dialogues in which a route-giver has made a ride on a “bus” through a vr site passing various landmarks. these were a sculpture, a town-hall, a market place with two churches, a park with a pond, a chapel and a fountain, all connected by streets (see figure 1, the route through the scenery, for details). after the ride the experience is reported to a follower expected to do the tour. the route-giver-follower exchange chosen (see figure 1 and the transcript figure 2) captures the ride from a hedge surrounding the park, its doorways and the main entrance towards a pond, and around it up to a place where an ice-cream man offers his services. the “bus”-tour leaves the park at the ice-cream stand and goes on to a chapel (1.1-g to 11-rh below). multi-modal anaphora and broadcasting of information 39 figure 1. route-giver’s vr ride on “bus”, route-giver entering the park with pond in the distance, the route through the scenery, landmarks: sculpture, townhall, two churches, park entry and path around the pond, chapel, fountain; an exchange: route-giver signing the two churches, follower signing left and right church, both concurrently, route-giver and follower interacting. in the annotation (see figure 2 below) we provide the german speech transcription (-g), an english word-by-word translation (-e) and a transcription of the right-hand (rh) and left-hand (lh) gestures. in the main text, the english translation is used except in cases, where german data must be considered. for characterizing the natural hand-gestures of the datum we use approximations to asl-handshapes. in order of appearance these are the following (see figure 3 below): c, small c, q, g, b, loose b (not represented in figure 3), o, and d. next to the aslhandshapes you find an approximation of them in the datum. note the differences in palmand finger positions. complementing the asl-handshapes we work with observational predicates like “small arc-shape”, “large grey arc-shape”, “drawing trajectory right angle”, “palm flat”, “indexing location at embankment of pond”. since we deal with two-handed gestures in the paper, the lhrh-coordination is all important. we notate speech-gesture overlap with [] and {} in the following way: l-handshape is marked with brackets “[….]”, r-handshape with curly brackets “{….}”. in the datum we have a two-handed gesture (or two coordinated gestures) and see that the lhandshape o is held throughout [and at this pond. you {drive towards it and you drive right around it.}] and the r-handshape d underpins {drive towards it and you drive right around it.}. in other words, the pond gets an lh-o shape constantly held and the driving gets a trajectory with d shape extending over most part of 4.1 and 4.2. observe that the gestures cover much more than their natural alignment expression3. this shows a characteristic temporal asymmetry between gesture and speech discussed in (rieser and lawler 2020). in this paper we do not provide the original elan corpus annotation and the map from corpus annotation to semantics but assume it as given. 1.1g route-giver: dann {kommt auf der rechten seite eine hecke}, {eine grüne hecke}. {zum teil auch mit eingängen.} 1.1e then comes at the right side a hedge, a green hedge. partly also with doorways. 1.1lh ……………………………………………………………………………………………………………………… ………………………… 1.1rh {r-handshape c, }, {r-h c, small c }. {r-handshape q }. 1.2g route-giver: die sind dann so bogenförmig, wie so n‘ garteneingang. 1.2e these are then like arch-shaped, like such a garden entrance. 1.2lh ……………………………………………………………………………………………………………………… 1.2rh {r-hands. small-arch-shape}, 1.3g route-giver: 1.3.1 bis irgendwann einmal dann der haupteingang kommt. 1.3.2 das ist ein großer grauer bogen. 1.3e until sometime then the main entrance comes. that is a large grey arch. 3 the standard segmentation of gesture going back to mcneill (1992, pp. 83-85) and kendon (2004, pp. 111-124) is in preparation, stroke and retraction. the cited datum has no observable preparation phases or retractions. the impression is that gesturing proceeds from stroke to stroke. it can well be that post-holds have a dual function; besides preserving the semantics of the stroke they could function like a quasi-retraction. rieser 40 1.3lh …………………………………………………………………………………………………………………………… 1.3rh {r-hands. large arc-shape + small c}. 1.4g route-giver: und da mußt du sofort, scharfer rechter winkel, rechts rein. 1.4e and there must you quickly, sharp right angle, to the right into. 1.4lh …………………………………………………………………………………………………………………………….. 1.4rh {g, drawing trajectory right angle, + modelling bus }. 1.5g follower: in den bogen rein? 1.5e in the arch into? 1.5lh 1.5rg route-giver: { g, drawing trajectory right angle } 2.1g route-giver: 2.1.1 ja, in den bogen rein. 2.1.2 es ist keine wand sondern eine hecke. 2.1e yeah. into the arch. it isn’t a wall, it’s a hedge. 2.1lh [l-handshape palm flat ] 2.1rh { g, drawing trajectory right angle } { r-handshape palm flat } 3.1g follower: ok. 3.1lh ………………………. 3.1rh ………………………. 3.2g route-giver: 3.2.1 wenn du dort eingefahren bist, 3.2.2 fährst du geradeaus auf einen teich zu. einen teich. 3.2e if you have there driven in, you drive [{straight} towards a pond. a pond.] 3.2lh c; [b parallel trajectory ] [b parallel trajectory] [l-handshape o ctnd. ] 3.2rh {b parallel trajectory } {r-handshape c > open o ctnd } 4-g route-giver: [4.1 und an diesem teich. 4.2 du {fährst drauf zu und 4.3 du fährst rechts herum. }] 4-e [and at this pond. you {drive towards it and you drive right around it. }] 4lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.ctnd. ctnd. ] 4rh {r-handshape d ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.ctnd. ctnd. } 5-g route-giver: [die hecke, {die geht noch ungefähr} so 50 m.] 5-e [the hedge,{it runs another roughly} 50 m. ] 5lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ] 5rh {r-handshape loose b ctnd. ctnd. ctnd. ctnd. } 6-g route-giver: [und dann sind dort {auch} hin und {wieder} {sitzbänke}.] 6-e [then there are {also} here and { there } {benches} .] 6lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ] multi-modal anaphora and broadcasting of information 41 6rh {r-handshape loose d} (first two overlaps) gesture expressing doubt: {wiggling of loose d handshape} (third overlap) 7-g route-giver: ab[{er du fährst um den teich herum. rechts herum. }] 7-e [{but you drive around the pond. right around. } ] 7lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ] 7rh {r-handshape d ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. } 8-g route-giver: 8.1 [und manchmal ist da auch {nen eisverkäufer }.8.2 und an dem fährst du rechts ab.] 8-e [and sometimes there is {‘n ice-cream man there.} and at him you drive off to the right.] 8lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.] 8rh {r-handshape g ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.} 9-g follower: was heißt “manchmal”? 9-e what does “sometimes” mean? 10g route-giver: 10.1 [ja, könnte verändert werden. 10.2 auf jeden fall, {auf} meiner tour war dort ein eisverkäufer. ] 10e [well, could be changed. in any case {on} my tour there was an ice-cream man there.] 10lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.] 10rh {r-handshape d, 3 beats/indexing location at embankment of pond} 11g route-giver: ein wagen mit schirm. 11e [{a cart ] with umbrella. } 11lh 11rh [ lh + rh pushing pantomime ] {rh modelling umbrella} figure 2. annotation of the paper’s main datum. information relevant for anaphora is marked yellow; broadcasting cases are marked blue; ”.” represents larger and “,” smaller speech pauses. observe the amount of broadcasted information. c rieser 42 q figure 3. used asl shape names c, small c, q, g, b, o, d (right hand) with natural gesture occurrences for comparison. observe that the natural d gesture in the datum uses the index finger and middle finger. data and algorithmic handling as a preview, we consider what an algorithm modelling some features of the datum must cover: first of all, we need independent renderings of gesture and speech information since they carry independent information: from the annotation of speech and gesture, rh, lh, it can be seen that rh and lh run in parallel with speech and with each other. therefore, we need some mechanism to express parallelism/concurrency. the roundness shape of the arch is initially only given by iconic gesture. in order to get “round arch”, round’4, the gesture semantics, and arch’, the word semantics, have to communicate. specifically, we need mechanisms of information transfer and information composition. we have the relation of 1.1-g to 1.2-g, an antecedent-anaphora relation; to capture this we need a device mapping the antecedent information onto the anaphora. the blue marks indicate persistence of information across contributions and turns, e.g., gesturing a right angle, the 4 information given in red indicates that it was provided by gesture. multi-modal anaphora and broadcasting of information 43 pond or the moving-towards-the-pond trajectory. lines 3.2 and 4 contain conditionals, hence a formal analysis of conditionals is needed. 3. relevant literature background vis-à-vis the datum anaphora accounts differ as to how they conceptualize the relation between antecedent and anaphora and use it in predictions of which antecedent will be selected. this relation can be one of coherence as in coherence theories, identity of discourse referents as in drt and drt-like approaches or a definiteness algorithm in corpus-based research. for a background to speechgesture integration we look into literature on multi-modal fusion. these are broad fields: as general overviews we recommend poesio (2016), poesio et al. (2016) for anaphora, geurts, beaver, maier (2020) for drt and johnston (2019) for finite state models. 3.1 coherence-bound theories the literature on linguistic anaphora mainly focusses on 3rd person pronominal resolution (krahmer and piwek, 2013, king and lewis 2018, muskens 1996, rohde 2018). there, the works of hobbs gave the initial impulse: hobbs (1978) suggested a coherence based semantic system for pronoun interpretation relying on the inference patterns contrast, cause, violated expectation, temporal succession, paraphrase, parallel and example. kehler’s (2002) coherence notion for pronoun resolution drew on hobbs (1978 and 1979 inter alia). it takes up hume’s assumptions on connection among ideas (hume 1748/1955), “resemblance, contiguity in time or place, and cause or effect”. kehler’s class of resemblance relations includes parallel, contrast, exemplification, generalisation, exception, and elaboration. for pronoun resolution he distinguishes coherencedriven approaches (like his own) from attention-driven approaches such as centering theory (brennan et al. 1987). kehler et al. (2007) investigated the coherence-driven approach with respect to four interpretation biases, grammatical role parallelism, thematic role, implicit causality, and subjecthood. recently, rohde (2018), following this line of research, distinguished two types of accounts: a coherence-based one (hobbs 1978, 1979, kehler 2002) and a surface form based one as in, e.g., centering theory. she argues that an efficient model of pronoun reference must cover both, coherenceand surface-relatedness: coherence theory (semantics bound) and centering theory (surface bound) cover distinct factors of her experimental findings. a unified approach subsuming both, coherenceand surface-relatedness, is proposed, formulated in bayesian terms. 3.2 drt and drt-like approaches kamp and reyle (1993) have the widest empirical coverage (singular and plural pronouns, tense and aspect) for anaphora. they treat the anaphora problem as follows: starting from surface syntax, they map syntactic structures onto representations (drss, discourse representation structures) and then provide these with a first order interpretation5. pronominal anaphora are analysed as relations between pronouns and discourse referents introduced in the discourse by an accessible antecedent. they show that numerous grammatical constraints determine accessibility. a further step in understanding the workings of anaphora in discourse was achieved with asher and lascarides (2003), referred to as al below. while building on the insights of kamp and reyle (1993) concerning grammar-based mechanisms of anaphora resolution, they argue that one must consider the rhetorical structure of discourse and the right frontier constraint. both influence anaphora resolution and the regularities of information packaging. in addition, they extend the anaphora notion to sentence pairs linked by cue-words like but, because or and-then and show that rhetorical relations are operative in dialogue. this is similar to the hobbs-hume-kehler coherence relations. von heusinger (2007) treats inter alia alternations of definite nps and pronouns in 5 one of the reviewers noted that muskens (1996) should also be mentioned in this context. rieser 44 anaphoric chains. against standard drt and sdrt approaches he proposes five aspects of accessibility: activation, accessibility relation, accessibility hierarchy, accessibility structure, and salience. one of our papers which touches on many topics relevant in the anaphora context and specific issues of multi-modality is (poesio and rieser, 2011), formulated in ptt. it observes linguistic and psycholinguistic findings which plausibly show that interpretation of reference is incremental and interpreters’ reference hypotheses run in parallel. the focus of the 2011 paper is inter alia on definite and pronominal anaphora. its anaphora theory uses contexts (situations in the sense of situation semantics) represented as drss specifying the antecedents of anaphora. building on poesio and rieser (2011a), this context theory uses visual situations to provide the semantics of indexicals, pointings and focus movement. 3.3 definite anaphora corpus work on definite description anaphora is rare, poesio and vieira (1998) being an exception. they ran two experiments asking subjects to classify the uses of definite descriptions in a corpus of articles from the wall street journal. the classification schemas used were taken from hawkins (1978), prince (1981), and fraurud (1990). in one experiment the classification tags were coreferential, bridging, larger situation, unfamiliar, and doubt; note the difference from the parameters relied on in the hobbs-kehler tradition. remarkably, subjects disagreed about the antecedents of definite descriptions. this is a crucial finding, since it points to the non-determinism of the antecedent-anaphora relation. we now turn to the alignment relation of speech and co-speech gestures. 3.4 speech-gesture alignment6 giorgolo and verstraaten (2008) carried out two perceptual experiments using shifts of gesture with respect to phonological information. their experiments pointed to a correlation between gesture information and phonetic peak. they looked for the anchor at which gesture is coupled to speech. using evidence from motion tracking data they found that gestures’ peak velocity is closely aligned to the pitch (f0) of speech. however, their results also indicated that semantic information played some role. leonard and cummins (2010) investigate asynchrony between beats and speech. subjects’ detection of asynchrony of post-word beats was considerably better than with pre-word beats (first experiment). a second experiment revealed a relatively tight alignment of gestural apex and pitch accent. the authors stress that their findings do not extend to naturally occurring gestures. navarretta (2011) points to eisenstein and davies’ (2006) studies which show that gestural holds can be useful for co-reference resolution. she investigated gestures semantically related to anaphora and co-referring expressions in danish spontaneous two-and three-party conversations. gestures related to anaphora are deictic or iconic; they can occur with speech or on their own. if antecedents and anaphora are affiliated to gestures, the shape descriptions of both show “many common attributes and values” (p. 8). navarretta concludes that “the shape of hand gestures might therefore contribute to the resolution of anaphora and co-referring expressions”. esposito and esposito (2011) conducted experiments with different age groups which showed that speech pauses are highly synchronized with gestures: however, in esposito et al. (2001) also non-synchronisation data were described. navarretta (2021) offers data about corpus investigations and classification experiments concerning the role of silent and filled speech pauses and accents for the identification and classification of individual and abstract anaphora in danish. as far as we are aware, there is virtually no research on gestural post-holds. in contrast to coherence bound theories, drt-like approaches to anaphora resolution or definiteness accounts which all work bottom up, we develop an alternative resting on the idea that antecedents and anaphora 6 this section has been introduced following the suggestion of an anonymous reviewer for d&d. for the perspective on speech-gesture asynchrony maintained here, cf. the section “methodology and contributions of paper” in 3.8. multi-modal anaphora and broadcasting of information 45 communicate information from (in)definites coming top down and that this is taken up by anaphora ports. the topic of multi-modal anaphora is, as far as we can see, modelled here for the first time, although our intuitions resemble those of navarretta (2011). concerning speech-gesture alignment, we start from the assumption of speech-gesture asynchrony mediated by concurrent speech-gesture communication which is orthogonal to the kendon-mcneill tradition. 3.5 multi-modal fusion one of the main aims of this paper is to show how gesture meaning and speech meaning can be integrated. the following aspects of multi-modal integration are important for our concerns: (a) integrated accounts of at least two different modalities, (b) the special mechanism of integration used, (c) composition of information7 , (d) incremental (step-by-step) input, (e) handling of temporal modality co-occurrence. aspect (a) means that modalities occurring together are mapped onto one representation format. aspect (b) looks to the means of integrating the modalities, such as, e.g., frames, types of unification or finite automata. how information from different modalities is integrated is the topic of (c). by (d) we understand the successive modelling of bits of information such as, for example, incoming words and gestures which was the main topic of poesio and rieser (2011). the temporal relation between the modalities is considered in (e). as work from ai and linguistics shows, multimodal integration can start from independent representations for gesture and speech and then provide some fusion machinery; (a) is usually taken for granted. development of fusion machineries is one of the most productive research fields in recent ai. they can be implemented as program-based multi-modal integrators combining uni-modal information as in robotic accounts (bergmann et al. 2011, kopp et al. 2004, pfleger 2007), interrelated frames (koons et al. 1993), semantic networks (mehlmann and andré 2012), as prolog unification (mehlmann and andré 2012), typed feature structure unification (johnston et al. 1997, johnston 1998, alahverdzhieva & lascarides 2010, 2011, lücking 2013, lücking and ginzburg 2021), parallel automata (mehlmann and andré 2012) or finite state automata (bangalore and johnston 2009, johnston 2019). the fusion mechanisms mentioned share some features since they must handle information transfer. lascarides and stone’s (2009) gesture-speech account differs from all of these insofar as they link gesture semantics to speech semantics (both propositional) with rhetorical relations. 3.6 coherence relations in the datum due to the fact that we deal with the description of a vr-tour which goes from turning point to turning point and from landmark to landmark (see figure 1, route), route-giver and follower are focused on end-states of events and the beginning of the next event. this is roughly what is captured by the hobbs-kehler coherence relation of occasion (kehler et al. (2007), p. 24). occasion captures sentence pairs and the end state bias affecting the choice of pronouns in the follow-up sentence, as for example with der haupteingang/das, the main entrance/that in the datum 1.3-g/e. for getting at overall coherence in a dialogue one needs more; on that one can consult the rhetorical relations of al. although al deal with pronoun interpretation biases only marginally, they have a general definition for antecedents to anaphora (definition 15, p. 149) which can be used in tandem with the rhetorical relations listed in their appendix d. these can be extended to dialogue (al, pp. 203273, vide especially “why dialogue and monologue are similar)”) and to multi-modal dialogue (lascarides and stone 2009). in-turn relations of our datum can be modelled using al’s continuation and narration, where narration is roughly equivalent to hobbs-kehler occasion8. for 7 one of the reviewers drew our attention to kuhn (2022) who develops a dynamical semantics for iconic gestures in sign language. although kuhn (2022) and ly are both compositional, they differ in how compositionality is implemented. this is probably due to the fact that sign language is language (one level) and we deal with the speech-co-speech gesture relation (two levels). 8 we take narrations up in the appendix, 2.1 rhetorical relations. rieser 46 an alternative, modelling coherence in terms of contexts/situations, see poesio and rieser (2011). one’s take on the anaphora problem in dialogue must in the end be guided by some dialogue theory, which demands more than a standard version of dynamic semantics can provide. this is clear from an early paper of eckert and strube (1999) hooking up to the then existing dialogue research. they specify the domain of anaphora antecedents using the notion of common ground existing between discourse participants and across turns. these issues are beyond the scope of the present paper. we dealt with some of them in poesio and rieser (2010, 2011). so far, we haven’t considered multi-modal anaphora to which we now turn. 3.7 anaphora and multi-modal anaphora in the datum as could be expected, our take on anaphora in the saga corpus also reveals a tight connection between discourse structures and anaphora occurrence mainly due to the existence of coherence relations between sections of discourse9 but in contrast to the regimented data preferred in cl and experimental linguistics, our datum in fig. 2 shows great variance, cf. (a) to (f): (a) prepphrase/def. article: [mit] eingängen/die, with doorways/these; (b) np/neuter demonstrative pronoun: der haupteingang/das, the main entrance/that; (c) nppl./npsingular (subspecies): mit eingängen/der haupteingang, with doorways/the main entrance; (d) npsing./local adverb: ein großer grauer bogen/da, a large grey arch/there; (e) def. np/def. np: [in]den bogen/[in]den bogen, the arch/the arch; (f) event or situation/local adverb: station der busfahrt/dort, stage of bus tour/there. to provide an impression of the full multi-modal data, we present the anaphora (a) to (f) above in their actual multi-modal occurrence as (a’) to (f’). the multi-modal information is only stated in the german examples: (a’) prepphrase/def. article: mit eingängen + r-handshape q (see figure 4)/die + r-handshape small arc shape (see figure 5); with doorways/these; figure 4. r-handshape q figure 5. r-handshape small arc shape (bottom of arch, representing curb) (b’) np/demonstrative pronoun: der haupteingang + iconic information from above/das + rhandshape large arc-shape + small c (see figures 6, 7); the main entrance/that; 9 see jasinskaja and karagjosova (2020) for a general discussion of rhetorical relations. multi-modal anaphora and broadcasting of information 47 figure 6. r-handshape large arc-shape (by an arc drawing gesture) + small c for curb figure 7. r-handshape large arc-shape + (by a drawing gesture) base of arch (c’) npplural/npsingular(subspecies): mit eingängen + r-handshape small arc shape (see figure 5) /der haupteingang + iconic information from above; with doorways/the main entrance; (d‘) npsing./local adverb: ein großer grauer bogen + r-handshape large arc shape /da + g drawing trajectory right angle + modelling bus plus track of bus (see figures 6 and 8); a large grey arch/there; figure 8. g drawing trajectory right angle (route into main entrance to park) figure 9. modelling (sides of) bus figure 10. stage of bus tour + indexing; observe a pond still signed and held in the left hand (e’) def. np/def. np: den bogen + iconic information from above + g drawing trajectory right angle + modelling bus plus track of bus /den bogen + iconic information from above + g drawing trajectory right angle + modelling bus plus track of bus (see figures 8 and 9); the arch/the arch; (f’) event or situation/local adverb: stage of bus tour + indexing /dort + indexing (see figure 10); stage of bus tour/there. in (a’) the doorways are specified by asl q and the uptake of these by a small arc shape, hence the multi-modal meanings are “arch-shaped doorways” and “doorways with small arches”, respectively. (b') the main entrance being among the doorways is given iconic information from these. (c’) with doorways is affiliated with gesture information for their shape; main entrance is fused with an iconic gesture as before. (d’) the large grey arch gets a large arch shape by drawing using a small c gesture. with the anaphora there a trajectory is gesticulated for the movement of rieser 48 the bus into the arch. (e’) gets information already multi-modally expressed plus trajectory information. (f’) anaphora and antecedent are differently indexed. after these overviews, we give a brief summary of the contributions this paper makes to the topics referred to in the literature section. 3.8 methodology and contributions of paper based on longstanding research on the saga corpus (see lücking et al. 2012), we specified in rieser and lawler (2020) the following desiderata as a methodological guide line for gesturespeech research. concretely, a satisfactory account should (a) asynchrony: accommodate cases where gesture strokes come (substantially) earlier or later than the “fitting” speech part, i.e., cases where gestures introduce new meaning or modify an existing one regardless of their precise temporal occurrence. (b) independence: accommodate the independence of gesture and speech, for instance, it must accommodate cases where gesture strokes are held throughout several utterances. (c) blocking: allow for the blocking (non-communication) or postponing of semantic information. (d) broadcasting: accommodate broadcasting cases, e.g., by allowing for the replication or repetition of meaning pieces across contributions or turns. we also add the desideratum that speech-gesture meaning coordination should be determined algorithmically. there should be perspicuity regarding how the gesture meaning coordinates with speech meaning, and the coordination should not be represented in an ad hoc fashion but rather be the result of (finite) rule-bound procedures. this enables a systematic explanation of speech-gesture meaning coordination and a generalization to a variety of data and contexts: (e) algorithmic determination: a satisfactory account should algorithmically determine a gesture’s speech relatum and its coordination term. these desiderata call for a dynamic machinery, considerably different from dynamic semantics or the information flow idea in barwise and seligman (1997): formulated in terms of processes, we need output processes which can give semantic information a “piggyback ride” and we need input processes which receive this semantic information, get it, and hand it on to the right place, the “right place” being (as a rule) already existing information. in this paper we treat more data-related matters. we suggest solutions to the following problems using the ly-calculus where the focus is restricted to semantic representation: handling of a broad spectrum of anaphora attribution of meaning to iconic and deictic gestures fusion of speech meaning and gesture meaning handling concurrency and temporal asynchrony of speech and gesture treating (chains of) multimodal anaphora investigating broadcasting of post-holds across contributions and turns handling non-deterministic multi-modal antecedents contributing a new field (dialogue, anaphora, multimodal meaning) to the application domain of the y-calculus providing examples for the interaction of l-calculus and y-calculus. this list looks more ambitious than it is. once one has developed a speech-gesture fusion account and a machinery relating antecedents and anaphora, solutions for the other problems follow directly. anaphora intuitions are quite strong as the literature shows. however, there is no general definition of anaphora. instead researchers depend on intuitions tied up with examples. the reason for this seems clear: sec. 3.7 shows that anaphora semantics is denotationally too varied to be subsumed under a common semantic definition. a possibility would be to resort to disjunction. we take a different route: after many dialogue data have been analysed we look into their commonalities in terms of ly-derivations in sec. 5.5 résumé and look-back. multi-modal anaphora and broadcasting of information 49 4. process algebra set up process calculi (or process algebras) belong to an intensively researched field of ai, logics, and philosophy. they have been invented to do away with sequentialism in programming and to establish models for the communication of processes running in parallel (see bergstra et al. (eds.), 2001 for an impressive view of the topics relevant for the field). by now there are many variants, among them csp (communicating sequential processes), ccs (calculus of communicating systems), acp (algebra of communicating processes), ambient calculus, p-calculus, the more recently developed y-calculus, and variants and extensions of these.10 most of them share the following features: processes (agents11, processes, threads) which can run sequentially or concurrently. process interaction is via input and output facilities. process interaction can be regulated to achieve composition, choice, blocking and sequential behaviour. processes are used for data transport where the data can come with their own logics. processes can send information recursively. some of the application domains of process calculi are concurrent programming languages, cryptography, security protocols, business processes, molecular biology or social behaviour (see fokkink 2000, introduction, for examples). we now provide an informal description of the λψcalculus preparing for definitions 1, 2, and 3 below. 4.1 basic intuitions for matters of easier exposition we provide an intuitive description of the λψ-calculus ontology, its objects and relations. its basic dynamic units are so-called agents (processes, threads) interconnected by channels on which communicative processes can run. these agents can receive, internally compute, store and send information. think of these agents as programs equipped with input and output channels and a simple internal computing mechanism. the agents communicate information coded in a formal language. the computing mechanism consists essentially in receiving a value on an agent’s input channel and filling an information slot in the receiving agent. if an agent has completed the internal computations, it can send out the result to arbitrarily many other agents to receive it and work with it. there is a restriction on communication; for communication to succeed between n agents, their output and input channels have to correspond. in effect, there is, hence, a communication pipeline from an outputting agent to a receiving one. each such communication track is typed to handle a special type of informational entity.12 an agent can send information to different agents distributively and can receive information from different agents. see the remarks about broadcasting below. a special type of information propagation is the so-called replication. it happens if an agent iteratively sends information to other agents. the following process algebra set up for the ψand the λ-ψ-calculus is largely adapted from (rieser and lawler 2020) where interested readers can find more detailed information concerning the history of process algebra, the history of the ψ-calculus and some of its formal ai-background. 10 for more information see the wiki entries for “process calculus” and “p-calculus” which provide easily accessible information. abramsky (2008), baeten (2004), and wing (2002) are also useful. johansson (2010) is recommended for the y-calculus, fokkink (2000) for process algebra. 11 the agent notion is widely used in the p-calculus, the y-calculus and other process algebra research traditions. it should not be confused with the “intentional agent” notion used in dialogue theory, action theory and plan-based ai. 12 how the typing is done following a montagovian tradition is explained by rieser and lawler (2020). see also comment under definition 1. rieser 50 4.2 process algebra definitions we do not present and discuss the full ψ-calculus here but only the parts needed to describe multimodal integration and the transfer of linguistic and gestural information13 (for a more elaborate version of the λψ-calculus cf. (rieser and lawler 2020)). what does the ψ-calculus provide to model the intuitions concerning anaphora and multi-modal fusion laid out under sec. three? in principle, we have parameters, i.e., variables or names, (concurrent) operators on agents or processes, and dynamically operating agents at our disposal. the semantic representation of linguistic and gestural information is coded in data terms. these can come from any (higher order) logic. in some process algebras like the π-calculus (cf. parrow 2001) only variables (called names) get associated with input-output channels, but in the y-algebra they can be associated with arbitrary data, e.g., channels or typed λ-terms. channels help to transport data from one increment in a multimodal structure to another, and most importantly, across dialogue contributions and turns; we do not know of any other formalism which has this flexibility while maintaining strict control. hence, they are an obvious means to provide inter alia the semantics of anaphors. the parameters indicating ψ’s data types are given in the following definition 1 (from bengtson et al. 2011, pp. 4-14)14. definition 1: t the (data) terms, ranged over by m, n, n’, …. c the conditions, ranged over by φ a the assertions, ranged over by ψ. in our application data terms t will be standard formulas from a typed λ-calculus such as a property m = λfz (like’(arch-shaped’)(z) ù f(z)); there λ binds higher order information, a complex property variable f and a set variable z; the red coloured arch-shaped’ indicates that the information comes from an annotated arc gesture semantically mapped onto an arch-shape. the types, not given explicitly but presupposed to be intuitively evident, are standard, echoing a montagovian tradition. conditions c are used in the antecedents of case constructs/if-then-elses, see φi below. assertions are l-terms without free variables. they can be combined by ù and are, e.g., properties of agents, embedded in conditions of the case construct of definition 2 or used to specify contexts of derivation. the ψ-calculus agents or processes, indicated by p, q, …, are: definition 2: 0 nil, the 0-agent mn.p output-agent m(λ𝑥$)n.p input-agent case φ1: p1 ;…; φn: pn case construct15 p | q parallel (concurrency) operator between agents p and q; p and q can be executed independently or in parallel !p replication; p can be repeated arbitrarily often δ deadlock the 0-agent is one being inactive. 0 is needed, e.g., if a local computation terminates. 13 for example, we do not integrate mechanisms for hiding (encapsulating) information and for scope extension. we also do not deal with foundational matters in terms of structural operational semantics. 14 here we keep to the wording of bengtson et al. 2011, pp. 4-14, and to their use of “definition”. 15 instead of case constructs we often argue with conventional if-then-elses in the following. multi-modal anaphora and broadcasting of information 51 two agents implement symmetric channel communication: “mn.p” (m overbar, n dot p) puts a data structure n, more precisely, n’s value, onto output channel m% , sends it out and continues with process p, possibly a 0 process, if no “transportable” material remains. “mn.” is informally also called “prefix” and similarly for “m(λ𝑥$)n.” in the next line. “m(λ𝑥$)n.p” indicates that a datum is received on the input channel m and substituted for 𝑥$ in n and p. “m(λ𝑥$)n.p” will be further commented upon below, where we discuss data inputs for an actual example. also, note that the “λ” here is part of the ψ-object language. therefore, it differs from the λ-operator in the transported language. 𝑥$ is a sequence of bound variables (bound by λ). the syntax device “.” in, e.g., “mn.p” separates the prefix “mn” from the subsequent process “p”. “.” also functions as a scope marker. we frequently leave out “.”, if no misunderstanding can arise. in the case construct “case φ1: p1 ;…; φn: pn” alternatives are indicated by “;”; one alternative pi is chosen given that φi is true. the case construct is also used to model the nondeterministic “either or” and the “if then else”. “if then else” is captured by case φ1: p1 ;⌐φ1: pi; a generalized “either or” results, if more than one φi holds. the parallel operator “|” enables p and q to expand independently or to communicate with each other via output and input operators, perhaps after several independent expansions. for the communication case across “|” we have, e.g. m(λ𝑥$)n.p | mn’.q ® n.p[n’/x] | .q, where “®” signifies the transition relation of the operational semantics. here, on the front side, n’ goes on the output channel m (right side of “|”) and enters the input channel m substituting for free x in n and p, hence, n.p[n’/x]. on the right-hand side of “|”.q remains. example: in our application in sec. 516 we have, for example, agent(1): ch1 ent-i λf! ch*****4a ιx(hedge’(x) ù green’(x) ù f(x)).$x(hedge’(x) ù green’(x) ù f(x)).0 (ent-i). this embodies the information hedge’(x) ù green’(x) and can receive information via its input channel ch1. assume, e. g., that some agent(2) provides on his output channel ch***1 the information with-doorways’. since agent(1)s input channel and agent(2)s output channel correspond, indicated by the subindex 1, they communicate. this leads to the substitution of the y-name ent-i with withdoorways’. now, lb-conversion applies and we get the information $x(hedge’(x) ù green’(x) ù with-doorways’ (x)). agent(1) can now send the definite information ιx(hedge’(x) ù green’(x) ù with-doorways’(x))17 in its prefix out via its output channel ! ch*****4a. indeed, it can do so arbitrarily often. on that see the explanation to replication !p below and the sec. 5.5 broadcasting of information. replication !p is understood as equivalent to p|!p, meaning that p can be emitted arbitrarily often, and can for example be broadcast down across subsequent turns in descriptions of multi-modal dialogue. so, !p is a technical means to model trans-turn communication, such as for example, involving multimodal anaphora resolution. we will, for example, use !ch***3 λx(arch’(x)), to indicate the arches of doorways provided by gesture which are referred to several times in the datum in figure 2. 16 the example in sec. 5 has a few downward and upward recursive steps and is therefore slightly more complicated. 17 the definiteness information is based on the revised chastain rule motivated in sec. 5.1. however, a drt-based argument would lead to a similar result. rieser 52 the deadlock δ is taken from fokkink (2000, pp.7, 25) and used here merely terminologically18 instead of the frege ⊥ in the original ψ-calculus papers. in our application, deadlock δ could capture semantic inconsistencies, as among a property ball’ and a property square’, excluding square’ù ball’.19 δ’s difference from 0 is that 0, being an idle agent, marks a non-action but does not indicate a fail, therefore it does not impede the process of semantic computation. in addition to the agents, we also have operators. the equivariant (“equivariance” defined by α-equivalence20) operators are as in definition 3: definition 3: ↔: t ´ t ® c channel equivalence21 ä: a ´ a ® a composition ⊢ a ´ c entailment as remarked, for our descriptive needs we will use a reduced version of the ψ-calculus, focussing on the tuple <0, mn.p, m(λ𝑥$)n.p, case φ1: p1 ;…; φn: pn, p | q, !p, “.”, d> in the following. following bengtson et al. (2011), “ä“ is implemented as “ù“. the transitions of operational semantics used (see “® ” above in the formula m(λ𝑥$)n.p | mn’.q ® n.p[n’/x] | .q) will be captured by informal descriptions dealing with the change of states as in the above example. 0-agents will often be omitted. 4.3 data and algorithmic handling matched we now look back at the description of the datum in sec. two, where we explained some of its features and take up the terminology coined there: our modelling instances are now λψ-agents. agents can be seen as small programs which receive and send information. independence of gesture and speech is captured by establishing different communicating agents for them. concurrency of gesture and speech can be handled with the |-operator ranging over agents. sequentialization is achieved by typing of information transfer. for transfer of information we have the input-outputdevice between agents. compositionality of λ-data transported and the interface between ψvariables and λ-variables such as ent-i and f in the example for agent(1) above, guarantee compositionality among the finally assembled multi-modal meanings. in other words, we work with a two-tiered approach captured by ψ’s transporting agents (the first tier) and the data transported with these (the second tier; see definitions 1 and 2 above).22 persistence or repetition of information across contributions is modelled with the !-operator: information sent via ! can be received by a type-conforming input channel or blocked, i.e. if no fitting input channel is available; this serves as the major means for dealing with temporal gesture-speech (a-)synchrony. finally, natural language conjunction, conditional and ^ have corresponding operators in ψ (see definitions 2 and 3). 18 meaning that we do not use fokkink’s process algebra version here, which we, anyhow, view as an alternative to the ψ-calculus. 19 to really achieve this, we would have to enlarge our λ-component investing words with a fine structure allowing agents to check for compatibility. incompatible readings would then have to be marked with δ. 20 α-equivalence means substitution of bound variables in formulas. 21 taking up a suggestion of cooper and larsson’s (pers. comm.), we notate channel equivalence by identity subindexing, e.g., chi , instead of the axiom ↔: t ́ t ® c. this has repercussions for the use of the calculus in the derivations in sec. 5. the most recent work on channel connectivity we are aware of is pohjola (2020). 22 two-tieredness is, however, not an absolute necessity, as the concurrent λ-calculus of dezani-ciancaglini (1996) shows. multi-modal anaphora and broadcasting of information 53 arguments for the ly-calculus the y-calculus has been chosen both for theoretical and practical reasons, see the main reasons listed and the motivation for them below. it has (a) concurrency operator |, (b) a replication operator ! (c) a case construct, it entertains (d) a channel notion, (e) input-output facilities, (f) it can transport any higher order logic formulas. in more detail: (a) some concurrency operator must be used because of gesture-speech concurrency which cannot be captured using sequential modelling, the deeper reason being that multi-modal communication works in parallel. (b) the replication operator in tandem with (e) cares for the top-down distribution of information. (c) the non-deterministic case construct is useful for expressing natural language conditionals and the context-dependence of gesture (see fn. 24). (d) ys channel notion does service for modelling the visual-auditory channel. (f) y-agents can transport typed l-structures. footing on operational semantics (see, however, the restrictions of this paper mentioned in fn. 28 and fn. 14), the l-formulas care for the composition of verbal meaning on the one hand and co-speech gesture meaning and speech meaning, i.e. multi-modal fusion, on the other hand. this suggests a move towards a unified semantics for multi-modal behaviour. next, we move on to the description of anaphora and broadcasting (sec. 5) and its reconstruction in the λψ-calculus. 5. verbal anaphora, multi-modal anaphora, broadcasting of information 5.1 verbal anaphora we start our investigation of multi-modal anaphora with a closer look at the linguistic anaphora in our datum. we then will explain in greater detail which meanings23 co-verbal gestures contribute to the linguistic anaphora and what the resulting multi-modal meanings will be. anaphora in natural dialogue present problems different from those discussed in the computational linguistics or the philosophy of language literature: they extend over stretches of discourse and non-regimented anaphoric relations are extremely varied as we will see. a remarkably early tool to approach some of these observations is the nearly forgotten chastain (1975), which also relates the anaphora problem to the general discussion of the reference of terms at issue (including russell, strawson, quine, donnellan, and kaplan). we accept the following of chastain’s hypotheses24 with modifications, which we indicate in a. to d. below: (a) anaphora occur not only in pairs but in anaphoric chains as in eingänge/doorways die/these der haupteingang/the main entrance das/it da/there in den bogen/into the arch in den bogen/into the arch dort/there (1.1-g 1.2-g 1.3-g 1.4-g – 1.5-g 2.1g 3.2-g) (p. 204). (b) indefinite descriptions starting an anaphoric chain are not taken as existentially quantified expressions as, for example, the later dpl and drt accounts would essentially suggest, 23 we agree with reviewers a, b, and c that gestures can only be interpreted in context. we dealt with the context-dependence of gesture in lawler, hahn, and rieser (2017) “gesture meaning needs speech meaning to denote. a case of speech-gesture meaning interaction”. there, the idea is, using ys case construct, to make the meaning of a gesture conditional on some vr object(s) introduced by speech. so, for example, the o-gesture as shown in figures 11a, 12a, and 13a would depend for its meaning on pond/teich, the q-shape in figure 13 on its verbal context arch-shaped, like such a garden entrance/bogenförmig wie so’n garteneingang. in this paper we assume that a context-dependent meaning of the gesture has been computed using the case construct device. the gesture meanings supposed in this paper follow the decisions of the saga corpus annotators. 24 unfortunately, we cannot do justice to the philosophical insights of chastain (1975) here. page numbering is given according to chastain (1975). rieser 54 but as singular terms with a definite reference in order to account for definite uptake (pp. 206-215). (c) the descriptive content of singular terms can be enlarged over time in the discourse (p. 232). (d) in chains multiply linked terms can exist, i.e. different anaphoric chains can result (p. 267). now, as to our modifications: (b) is too strong and goes against the tradition in logic as, e.g., originally defended in reichenbach (1947, §§ 42, 47). we will weaken it in the following way, using ψ’s agent notion for dynamic processes and information transfer. a. an indefinite description can send a singular term for further anaphoric use. b. definite terms can always be sent to the outside. c. the terms of an identity can be sent to the outside and computation can continue with either term. d. extending standard ψ-calculus practice, we will allow arrays in the output prefix indicated with “< …>”, the fields of which can be accessed as usual. arrays are a source of non-determinism. they mark a sequence of possible dialogue continuations. the sending and receiving notion is bound to the ψ-calculus definitions (see definition 2, sec. four). it will be demonstrated in more detail in the application sec. 5.4. mainly due to the route description dialogues of the saga corpus, land-marks, objects, definite and indefinite locations, routes, and trajectories are reflected in anaphora formation as shown in (a). (b) states a precondition for explaining the transition from the first indefinite member of an anaphoric chain to a following definite one; clearly, an existentially quantified expression cannot be equivalently substituted by, say, a definite description in general, which has been a main reason for introducing discourse referents in drt. observation (c) allows for dynamically extending the meaning of successive members of an anaphoric chain. as a consequence, a definite description sent in the manner of (b) can contain a slot for further accumulating information; we, however, will only indicate an extension informally and avoid slots to avoid cluttering up the formalism. finally, (d) specifies different antecedents in anaphoric chains. we have to wait for applications of (d) until secs. 5.4 and the appendix 1.2 (see the function of agent(1) there). remark on nl syntax: as the dialogues in saga show, there is an intimate connection between antecedents, anaphora, german word order and information structure management. in order to describe that we informally use the so-called field theory of the german sentence in the following manner25: prefield sentence bracket enclosing the middle field26 postfield aber du fährst um den teich herum. rechts herum. but you drive around the pond. right around. the sentence bracket is given by the finite and the infinitival parts of the verb, fährst/drive and herum/around; the postfield is taken by the local adverb rechts and the factorized herum. the middle field is underlined. so much for the syntactic distinctions which will be addressed. we now successively walk through the antecedent-anaphora relations in the datum. 1.1-g: a hedge, a green hedge. partly also with doorways. here we see an extended indefinite a hedge, a green 25 for more information on the field theory of the german clause, see höhle (1986) and wöllstein (2018). 26 the most extensive study we know of the german middle field is haider (2006). multi-modal anaphora and broadcasting of information 55 hedge. partly also with doorways. we have a rhetorical anaphora a hedge with two elaborations in the repetition. the indefinite a hedge, a green hedge. partly also with doorways is the first member of an anaphoric chain. its descriptive content is hedge, a green hedge. partly also with doorways. remember that the second member in a chain must be marked as a singular term. working with the kamp and reyle (1993 p. 305-482) plural classification, we think that the whole np can be classified as either dependent plural or a case of distribution δi27 over which we can distribute in order to capture the plural eingänge/doorways. this is also taken into account in our λψreconstruction which specifies these two options (see agent(4) in appendix 1.1). in order to shed some light on 5.1 a. above, we turn to eine hecke/a hedge – die hecke/the hedge (1.1g/e-5 g/e). a first order representation would render this pair as $x hedge(x) ιx (hedge(x)), the first being a general and the second a singular referential term. the anaphora intuitions for cases like these are clear: the singular term depends on the general term. hence everything hinges on the explication provided for “depends”. we assume that once an object has been introduced by existential quantification, it can subsequently be taken up by a referential term. the new method used for implementing this intuition is the output channel device of the λψ-calculus (see sec. 4.2, definition 2) working as follows: suppose we have an agent transporting $x hedge(x) …. we use λψ’s output channel ch***i and a definite value ιx hedge(x), in this manner ch***i ιx hedge(x). $x hedge(x) …. it then transports ιx hedge(x) in the prefix to the outside. ιx hedge(x) can then be received by an input channel chi (subscripted “i” to indicate identity) of some follow-up agent invested with chi and enter into compositional relations with it. the anaphorical port, that is the representation of a pronoun, definite description, local adverb or whatever, waits for information from above. 5.2 the contribution of gesture meaning to anaphora meaning in this section we discuss the co-verbal gestures in the datum. remember figure 3 for the aslshapes and their natural data occurrences. the arch information by gesture beginning with a hedge, a green hedge, we have a sequence of rh gestures moving from loose c over small c to a classifier for “thin”. loose c is often used to support figure 11. loose c figure 12. small c going to “thin” figure 13. q-shape a height or width indication for objects but this does not match well with small c and the classifier for “thin” here. it could well be that this sequence is due to the motor planning of a preparatory sequence for the r-handshape q which is basically a rotated version of a sloppy thin classifier. figures 11-13 represent the gesture sequence in the video v5 datum. the handshape q is aligned with partly also with doorways. for all we know from saga, an equally plausible interpretation would be to regard the loose c-shape as an interactional gesture, as an attention seeking device 27 for distribution we use a distribution operator δ and the kamp-reyle *. rieser 56 more often expressed by loose g. however, the q shape can be seen as modelling two vertical trajectories connected by an arc (see figure 13). considered this way, the handshape provides additional information, since then it is about the round shape of the top side of the doorways as against a possibly straight shape. the gesture works in parallel with partly also with [doorways]. note that in order to be semantically appropriate, it has to skip over partly also and combine with doorways achieved in the formal reconstruction via an assumed typing of the output-input facilities. the content of q denotes an arch property. so, we have a concurrent interaction of several agents as detailed in appendix 1.1. 5.3 extracting agents from the video data as already emphasised, the formal machinery is based on the agent idea of the λψ-calculus. the identification of agents proceeds from the multimodal saga datum according to the following guide-lines: (1) gesture and speech yield different agents. gesture information is marked in red for easier identification: λx(arch’(x)), pond’lh, right-angle’ etc. which represent, respectively, gesture meaning arch’, pond’ generated by the left hand and gesture meaning right-angle’28. suprasegmental phonological marking of structures is taken as an indication for identifying agents. these can be a. pauses indicated by “.” and “,” in the datum transcript: eine grüne hecke. zum teil auch mit eingängen/ a green hedge. sometimes with doorways. (1-g/e) this is modelled, e.g., by interacting agents (1), (2), and (4) in appendix 1.1, b. final rises or level tone indicating continuation: ein fluß über den auch brücken führen/a river crossed by bridges. the level tone is on fluß/river. the final tone is on führen/crossed, c. final falls indicating termination of chunk of speech in the audio data: wie so n‘ garteneingang/ like such a garden entrance. see decision for agent(7) in appendix 1.1, d. material in the prefield: und an diesem teich./ and at this pond. observe that there is also a pause, marked by “.”; represented in agent(24) in appendix 1.3, e. material in the postfield: rechts herum/right around it. cf. agent(27), appendix 1.3. and similarly for (2) parentheticals: scharfer rechter winkel/sharp right angle (1.4g/e) with level tone; represented in agent(13), appendix 1.2 . finally, decisions due to the neo-davidsonian verb-frames and considerations for coming as close to a first-order language as possible can also provide motivation for postulating agents: teilweise auch/partly also with and mit eingängen/with doorways (1.1-g/e) yield agent(2) and agent(4), appendix 1.1. 5.4 anaphora resolution in the λψ-calculus we now turn to some of the multi-modal information of the route-giver’s and the follower’s dialogue contributions in the datum fig. 2. we do not think that the semantic values of gesture and speech differ metaphysically. nevertheless, we want to indicate the source of the meaning produced, in order to point out which source contributed which meaning; therefore, as mentioned earlier, we colour gestural meanings. coloured and non-coloured meanings are fused into a multimodal meaning as discussed in sec. 3.5. in the end, colouring will give us a rough idea of which 28 in this paper we do not provide a model-theoretic semantics for gesture and speech but stay at the level of semantic representation to avoid unnecessary length and complexity. if we did, red expressions would be given the same type of semantics as the black ones. this also follows from the idea of a unified semantic representation as developed in sec. 3.5 multi-modal fusion. multi-modal anaphora and broadcasting of information 57 objects were gesticulated, i.e. of the contribution of gesture to ontology. this is perhaps also a proper place to add a caveat as far as the meaning representation for iconic gestures is concerned. what we get through corpus-annotation are essentially constellations of trajectories in r4 (4dimensional reals), non-standard topological entities with their (non-standard) sui generis denotations as geometrical objects. in practice, there is a long way, for example, from two separated orthogonal surfaces in r4 connected by an arc-trajectory, as in 1.2-g/e or in 1.3-g/e, to an architectural arch-concept. here, we fix the relation on an observational basis substantiated by the vr-representations of the scenery got from the saga corpus (see the route-giver’s tour and landmarks in fig.1). a complete representation of the agents and their interaction is provided in appendix 1; the numbering of agents in main text matches the numbering in the appendix. for detailed information it is useful to consult the appendix 1. the arch information in this section the focus is on communication and exchange of information according to y-calculus definitions, the communication of gesture agent and speech agent, and anaphora resolution with the revised chastain anaphora rule plural readings of nps and plural anaphora. anaphora resolution, following the revised chastain rule in sec. 5.1, proceeds in a top-down fashion, originating from a source which is information on some output channel and “looking for” an input channel where it can hook up. this process might be one-to-one, i.e. one output channel linking up with one ready input channel or it might be distributive, i.e. one source information linking up iteratively with several input channels, indicated by “!”. the latter cases are mainly treated in the broadcasting section. examples: agent(8) gets information on input channel ch8 and instantiates the y-name m-entr with it. the value is handed on to the l-variable p by lb-conversion, so the information ch8 m-entr is in place. agent(8) now sends the main-entrance information ιv(main-entrance’(v)) on ch***9 to the corresponding agent(10), agent(8): ch8 m-entr ch***9 ιv(main-entrance’(v)). λp until-sometime-then’(p) (m-entr).0 |, where it is integrated via the y-name m-ent. agent(10): ch9 m-ent λu$y((large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) ù y = u) (m-ent) .0 |. y-names are arbitrary. m-entr and m-ent were chosen to indicate that they belong to different agents. however, it would not do any harm to use either m-entr or m-ent in both places, since they are substituted by the values transported by ch8 and ch9 which are different. agent (22) below, after having got information from channels ch15c, ch14, ch11, ch18, ch19, communicates ιx(pond’(x)), i.e. the description the pond, or the pond gesture semantics ιx(pondlh’(x)) (see revised chastain rule in sec. 5.1) distributively via !ch***17: agent(22): ch15c pa-i ch14 veh-i ch11 the-i ch18 polh ch19 porh λpth2 λveh λthere ! ch***17 <ιx(pond’(x)), ιx(pondlh’(x))>. (drive’(e2, fo, veh, pth2, there) ù straight-forward’(e2, veh, pth2) ù towards’(e2, fo, veh, λrs $x (pond’(x) ù x = r ù x = s)))(pa-i)(veh-i)(the-i)(polh)(porh).0 | similarly, see also the pond-information of agent(24): !ch***20 pond’lh. at’(ιx (pond’(x)) = pond’lh).0 |. this provides also a good example for observing that gesture semantics and speech semantics are treated as being on a par. following the incremental sequence, we start the formal rendering with the sequence of agents agent(1) | ,……, | agent(53), rieser 58 where “|” indicates concurrency according to definition 229. so, a green hedge is the first agent and there was an/the ice-cream man’s cart which had a handle bar (horizontal grasp pantomime) and an umbrella with vertical handle (vertical grasp pantomime) the last one. the numbering of channels is arbitrary. what matters is their identity and the role they fulfil for agents(n). due to channel communication and assumed typing regime, we can sequentialize agents, see for example the interaction of agent(1), agent(2) and agent(4) in appendix 1.1 which prescribes an execution order. initially, the algorithm starts with an agent in need of information to be provided by some other agent further down the line (see agent(1)) or by an agent providing some information for one or several other agents to come. the receiving agents wait for information without which they are blocked. this is similar to an algorithms entrance condition familiar from programming. so, we get in the end a communication net in which the agents depend on each other in terms of information. apart from occasional idle 0-agents, only two agents, 37 and 16, do not communicate. looking at the execution order, agent(1) needs information via the input channel ch1; the ψ-name ent-i receives a value sent from some corresponding output channel ch***1. agent(2) will be able to send on ch***1 but needs first information on ch2. now, information seeking percolates down, agent(2) must have some information on ch2. this is provided by agent(4). here we deal with a parallel agent indicated by | as introduced in sec. 4, definition 2. one side can provide an output via ch***2, once it received a value on its ch3 input channel served by gestural agent(3) modelling the information of q (see figures 3 and 4 for the datum q). the q information for small arch provided by orthogonal sides and the connecting arc is implemented as the arch-information which can be iterated arbitrarily often due to “!” (see sec. 4, definition 2). a closer look at the datum a green hedge. partly also with doorways reveals that we have concurrently a dependent plural reading of doorways, intuitively, the hedge’s doorways, as well as a distributive reading indicated by δi since either could obtain. the information flow branches nondeterministically; only the right-hand side information is considered in the following.30 ch***2 (δi (doorway’(i) ù of’(i,a) ù λx(arch’(x))(a))) finds agent(2) for communication; agent(2) ends up with an output ch***1 ready to interact. communication among agent(2) and agent(1) attributes to any member of a set the property of being a hedge, green and partly also with arched doorways as follows: $x(hedge’(x) ù green’(x) ù (partly-also’(with’(x, δi (doorway’(i) ù of’(i,a) ù arch’(a)))))) |. we let agent(1) send out a definite information according to the revised chastain rule for indefinite introduction and definite resumption (see sec. 5.1, a.). we use an array with two fields (see generalisation in sec. 5.1). the fields read as the green hedge, partly also with arched doorways and the arched doorways, respectively. both have been derived and can be replicated independently to fill anaphora slots to come. the die/these anaphor, the arch gesture and the park entrance in this section we deal with co-operation of handshape semantics and speech semantics, plural descriptions incorporating gesture representations based on the kamp-reyle plural distribution, non-deterministic (“upwards branching”) anaphora, isolated multi-modal agent information, contradiction of gesture meaning and speech meaning, 29obviously, we have, concerning the multi-modal dialogue extract, a sequence (except for speech-gesture mappings) however, remember that sequentialization can be forced by the typing regime. 30 we could also work with the left-hand side. it would not matter much for our example. multi-modal anaphora and broadcasting of information 59 and on the dialogue side, follower’s semantics, dual function of route-giver’s gesture, follower’s clarification requests and route-giver’s answers, follower’s decision for an anaphora antecedent. these in 1.2-g/e takes up the doorways already introduced within the antecedent. on these a bit later. we have a semantic cooperation of arch-shaped and the r-handshape small-arch-shape, where “small” is due to asl q (see figure 4) and the little information it covers in gesture space. we let it denote the constant small arch. as we shall see below, agent(6) embodies the denotation of the r-handshape small-arch-shape information and cooperates with the verbal information like such a garden entrance provided by agent(7). agent(6) can then output something like “small arch like such a garden entrance”. agent(5), representing the meaning of so like arch-shaped fuses first with “small arch like such a garden entrance” and then with the input doorway information. using kamp-reyle style plural distribution31, the result of the computation is informally “the doorways with their arches are arch-shaped, all of them [by distribution] have a small arch, and they [by distribution] are like such garden entrances”. the doorways are taken up by the anaphora these. agent(5) sets the scene for dealing with the anaphora relation, receiving the value δi (doorway’(i) ù of’(i,a) ù arch’(a)) from agent(1)’s output. agent(5) needs input on channel ch5 for the comparison clause like such a garden entrance. agent(6) and (7) cooperate to provide that. agent(7) implements the comparison. 32 the information sitting on ch***6 can be received by agent(6). ch4a second transports the set of arched doorways, δi (doorway’(i) ù of’(i,a) ù arch’(a)), a definite set; its representation can, hence, be used like a singular term. so, the final derivation for agent(5) is: like’(arch-shaped’)(δi (doorway’(i) ù of’(i,a) ù arch’(a))) ù (of’(δi (doorway’(i) ù of’(i,a) ù arch’(a)), small-arch’) ù like’(such’$u *garden-entrance’ (u))( δi (doorway’(i) ù of’(i,a) ù arch’(a)))).0 | this means that each i in the set operated on by δ that is an arched doorway, i.e., is of’(i,a) ù arch’(a), is like’(arch-shaped’) [first conjunct], has a small-arch’ [second conjunct] and is like such a garden entrance [third conjunct]. some of this information will be sent out by ch***7 specified below given that it finds a corresponding input channel. so, we have solved the anaphoric relation among eingängen/doorways and die/these; observe that die/these, the anaphora, is represented as δi(doorway’(i) ù of’(i,a) ù arch’(a)), i.e. as the set of arched doorways. next we consider until sometime then the main entrance comes. main entrance might be an anaphora candidate for doorways. so, the main entrance is in the set of doorways which are archshaped, have a small arch and are like such a garden entrance, as derived. the main entrance has to be tied up with this information in agent(9). that then takes up the definite information “the main entrance is the arch-shaped object, having a small arch like such a garden entrance” as sent by agent(9). agent(10) specifies that it is large, grey and has a curb (see figure 6 repeated). 31 i owe the argument that the example cannot be represented using the kamp-reyle σ-operator (which is an abstraction operator) but must use (an equivalent of) the kamp-reyle * to andy lücking. 32 for the * see kamp and reyle (2003), p. 327, fn 13. we use δ as bare plural operator and * for predicates which can have either singular or plural denotation. rieser 60 figure 6. r-handshape large arc-shape (by an arc drawing gesture) + small c for curb figure 8. g drawing trajectory right angle (route into main entrance to park) figure 9. modelling (sides of) bus the extent information concerning the arch is provided by the gestures q vs an r-handshape archshape and the adjective large. assembling the information, we get a contradiction of small vs large33. the curb-information comes from the r-handshape arch-shape plus the drawn small c information indicating an architectural frame structure (1.3 g/e). it is embodied in agent(10). agent(10) sends the information “the large grey arch with the curb which is the main entrance” to the outside. it is taken up by the anaphora there in and there must you quickly, sharp right angle, to the right into. and modelled in agent(12), see figure 8 repeated for the sharp right angle. by theory, the anaphora there can hook up non-deterministically to either the main entrance or the arch. this is due to the fact that a new definite can be created out of the np in the predication. as the clarification request below shows, the follower takes a large grey arch as the relevant antecedent. as we see in the formal rendering in the appendix, this has major repercussions for the anaphora analysis. in the datum 1.4-rh we also observe a dual function of the route-giver’s rh gesture: it models the moving vehicle and the route to be taken into the arch (see figure 9 repeated). both iconic gestures continue up to the end of the route-giver’s answer to the follower’s clarification request in 2.1-g/e. both, clarification request (agent(17)) and answer (agent(18)) are highly elliptical and get their information from agent(12), multi-modally specifying the way of moving into the arch. now we have to implement the anaphoric relation of the main entrance: the main entrance must be in the set of doorways having a small arch and being like such a garden entrance. agent(8) provides a simplified temporal prefix for until sometime then and will in the end have to send out the main entrance information after input on ch8. agent(9) submits the proposition the main entrance comes and expresses that the main entrance is one of the arched doorways which has a small arch and is like such a garden entrance. agent(5) is active on channel ch7 and provides the necessary input. the information derived is sent to the outside. the main entrance is one of the arched doorways, arch-shaped, has a small arch and is like such a garden entrance. agent(10) cares for that is a large grey arch. (1.3-g/e). the anaphoric indexical that is resolved to ιv(main-entrance’(v)) by information transfer of either 𝑐ℎ***8 or channel 𝑐ℎ***9: we have anaphora non-determinism, an upward branching anaphora as remarked in sec. 5.1 (d). agent(10) embodies the {r-handsh. large arc-shape + small c, figures 6 and 7} information represented as curb’. we are now ready to deal with the move into the park. agent(12) cares for the verbal structure of the contribution (1.4-g/e) and there must you quickly, sharp right angle, to the right into. its multimodal structure is schematically set up as follows: (must-in(current event, follower, vehicle, path taken, destination) ù f(current event)), 33 from the observer’s point of view, contradictions are quite frequent in the saga corpus. they are either overridden or (hardly ever) negotiated by the participants. multi-modal anaphora and broadcasting of information 61 where “must-in” represents informally a modal+verb-combination, “current event” the current event of the route, “follower” the route-giver’s addressee, “vehicle” the bus, “path taken”, the trajectory followed by the bus, and “destination” the current event’s endpoint. the gesture meaning which has to be integrated is {g, drawing trajectory right angle + modelling bus}34, as shown in figures 8 and 9. in more detail, agent(12) gets information from various communicating sources: agent(10) supplies the main entrance information. agent(13) provides the gesture information for right angle, sent on ch***15: ! ch***15 right-angle’.0 |. agent(14) contributes the vehicle gesture on ch***14: ! ch***14 vehicle’.0 |. the right-angle information and the vehicle information, both frequently distributed in the datum, are replicated, “!”. agent (15a) communicates event information and adverbial information on ch***15b.what is represented is roughly, “you must quickly [go] directly into the large grey arch with curb [which is the main entrance]”. we now come to the follower’s clarification request in den bogen rein?/ in the arch into? represented in agent(17). the problem that arises is “what is the semantics of the follower’s notion of arch”? we introduce a test on the follower’s side for a specific interpretation: ιu(arch(u)) = ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) expressing “the arch is the large grey arch with curb”. the question sends the information to the outside and agent(18), encoding the answer, takes it up. so, we have for the follower’s clarification request agent(17) interacting with agents(10), for providing the main entrance, (12), for specifying the event e1, (14), for pinpointing the vehicle, and (13), for giving the route direction into the arch. agent(17)’s output is ch***36 (ιu (arc(u)) = ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’))).0 |. agent(18), encodes the answer to the follower’s clarification request35. the feeding channels are as in agent(17); the output value is the filled verb frame: ch***37 into’(e1, fo, vehicle’, right-angle, in’ (ιu(arc(u)) = ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’))). 0 |. channel ch15, the path information indicated by the route-giver’s gesture, ! ch***15 right-angle’, is active during the clarification request and the answer. therefore, the clarification request is a cooperative undertaking between route-giver and follower: in order to be sure about the direction of the bus route, i.e., the path to be taken, one needs the route-giver’s right-angle’ indication which works like an anaphora, in our account by top-down information transfer. it is not clear why the route-giver remarks it isn’t a wall, it’s a hedge (2.1-g/e), since the bus must go through the arch which was just clarified. from a gestural perspective it is an interesting datum, since the front of the hedge is modelled with a vertical slanted two-handed asl-flat b gesture, taken here as meaning front’ (see figure 16) and modelled with agent(16a). 34 remember that blue marks indicate broadcasted information, the discussion of which is postponed to sec. 5.5. 35 for a kos and ttr perspective on a similar datum see ginzburg and lücking (2021b). rieser 62 figure 16. left: vr-entry into park with large grey arch right: two-handed gesture for hedge figure 11a. pond first gesticulated by co-verbal lh and rh gestures with slightly different o-shapes so we end up with ($x(¬ wall’(x) ù hedge’(x) ù of’(x, front’))). no information is on an out-going channel showing the singularity of the route-giver’s contribution36. the anaphoric chain cited in the context of chastain’s anaphora theory in sec. 5.1 is modelled with agent(1) (doorways) – agent(5) (these) -agent(8) (the main entrance) – agent(10) (it) – agent(10) (there) – agent(17) (into the arch) – agent(15a) (there). 5.5 broadcasting of information we have so far met gestural information about objects, events, routes and directions which have been freely replicated one-to-one across dialogue contributions. the idea is that information of this sort is output, received and integrated somewhere in the dialogue on a single output-target basis. we now turn to an example where information transfer is generated by a continuous post-stroke hold. the existence of these post-holds in saga is one of the main observations dealt with in the present paper. it is relevant likewise for empirical gesture research (such as gesture annotation or experiments concerning the relation of gesture and pitch accent as reported in pow and dixon (2019) or navarretta (2021)) and its algorithmic rendering (e.g. the obvious need for an iterative/recursive information emitting device). the formal device to capture the broadcasting of information is the “!” operator as introduced in definition 2. the !-operator can emit the information in its scope, which is an output pre-fix of a formula, arbitrarily often. it is handed on, if it finds a fitting input channel. however, even after this act, “!” goes on to send the information in its scope; it further emits the information for concurrent or upcoming input channels, if any. the behaviour of “!” can be compared metaphorically to a top-down impulse. a first case of broadcasting is found in 1.4-rh, 1.5 rh, 2.1 rh (blue marks) where the routegiver signs the route through the main entrance into the park. signing continues from the routegiver’s first directive up to the end of the follower’s first clarification request. it lasts as long as the route into the park is at issue. in the formalism this is reflected in the right-angle information of agents(12), (13), and (17). however, the main occurrences of broadcasting in the datum are found in 3.2-g to 10-g, tied up with the pond-information, see the blue marks in the excerpt of the datum below. first, briefly, again about the move into the park. in the conditional’s antecedent in 3.2-g/e if you have driven in there we have there anaphorically related to arch in 2.1-g/e. two parallel trajectories with flat lh and rh asl shapes b indicate again two things, the vehicle’s movement (distance between both hands for the vehicle breadth) and the path into the arch (distance between 36 nevertheless, in sdrt terms it could perhaps be taken as an elaboration. it could therefore be placed on an outgoing channel but perhaps not taken up. multi-modal anaphora and broadcasting of information 63 both hands for the track width, see figure 9, (agent(21)) and towards the pond (agent(22)). this becomes ultimately pooled in agent(20). 3.2g route-giver: 3.2.1 wenn du dort eingefahren bist, 3.2.2 fährst du geradeaus auf einen teich zu. einen teich. 3.2e if you have there driven in, you drive [{straight} towards a pond. a pond.] 3.2lh c; [b parallel trajectory ] [b parallel trajectory] [l-handshape o ctnd. ] 3.2rh {b parallel trajectory } {r-handshape c > open o ctnd } 4-g route-giver: [4.1 und an diesem teich. 4.2 du {fährst drauf zu und 4.3 du fährst rechts herum. }] 4-e [and at this pond. you {drive towards it and you drive right around it. }] 4lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.ctnd. ctnd. ] 4rh {r-handshape d ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.ctnd. ctnd. } 5-g route-giver: [die hecke, {die geht noch ungefähr} so 50 m.] 5-e [the hedge,{it runs another roughly} 50 m. ] 5lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ] 5rh {r-handshape loose b ctnd. ctnd. ctnd. ctnd. } 6-g route-giver: [und dann sind dort {auch} hin und {wieder} {sitzbänke}.] 6-e [then there are {also} here and { there } {benches} .] 6lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ] 6rh {r-handshape loose d} (first two overlaps) gesture expressing doubt: {wiggling of loose d handshape} (third overlap) 7-g route-giver: ab[{er du fährst um den teich herum. rechts herum. }] 7-e [{but you drive around the pond. right around. } ] 7lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ] 7rh {r-handshape d ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. } 8-g route-giver: 8.1 [und manchmal ist da auch {nen eisverkäufer }.8.2 und an dem fährst du rechts ab.] 8-e [and sometimes there is {‘n ice-cream man there.} and at him you drive off to the right.] 8lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.] 8rh {r-handshape g ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.} 9-g follower: was heißt “manchmal”? 9-e what does “sometimes” mean? 10g route-giver: 10.1 [ja, könnte verändert werden. 10.2 auf jeden fall, {auf} meiner tour war dort ein eisverkäufer. ] 10e [well, could be changed. in any case {on} my tour there was an ice-cream man there.] 10lh [l-handshape o ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd. ctnd.] 10rh {r-handshape d, 3 beats/indexing location at embankment of pond} rieser 64 11g route-giver: ein wagen mit schirm. 11e [{a cart ] with umbrella. } 11lh 11rh [ lh + rh pushing pantomime ] {rh modelling umbrella} figure 3: broadcasting fragment of figure 2. as the datum 3.2-g/e shows, the pond (fig. 12a) is first gesticulated by co-verbal lh and rh gestures with slightly different o shapes (figure 11a). since there is only one pond in the vrsetting acting as our model, their denotations must be identical and identical to the pond denotation expressed in speech. for antecedent and consequent of the conditional we have “if you have driven into the arch, you (next) drive towards a pond”. “drive” is underpinned with gestural information for vehicle, path (rh) and the pond goal (lh). this completes the information of agent(22). it distributes the relevant pond information in speech and gesture by “!” replication; ! broadcasts the pond information up to 10-lh and beyond. 4-g/e is concerned with the pond and the path to reach it. instantly, the function of the rh and the lh changes: the lh still broadcasts the pond information as before while the rh now signs the path towards the pond and then around it. we deal with a complicated conditional expressing “as for the pond, if you have driven towards it, you (next) drive right around it”. the route-giver, after again focussing the pond, moves one step back in the route description and repeats you drive towards it and then adds and then right around it. it takes up this pond anaphorically by lh o and continues throughout the contribution 4-g/e. this way it also contributes the pond information for the elliptified towards it and right around it. the l-handshape o is statically held on pond, while the r-handshape d is drawing the route. looking at the vr-rendering of the pond and its environs (see figures 12a and 14), we see that we get the correctly gesticulated information related to the pond and the route around it, in other words, gesture information seems truth functional37. figure 12a. left: the path into the park through the main entrance (arched main entrance from the street into the park at the bottom of slide), around the pond and out of it. right: the rh draws the path towards the pond and then around it. the l-handshape o is statically held while the r-handshape d is drawing the route towards the pond and around it. agents(26) and (27) continue the gestural information for vehicle and path. in 5-g/e, “the hedge, it runs another roughly 50 m”, the route-giver’s left hand still continues to hold the pond 37 see, however, fns 33 and 39 which call truth functionality into question. multi-modal anaphora and broadcasting of information 65 representation while the rh gesticulates, using loose b, the distance covered by the vehicle; the speech-gesture overlap is it runs another roughly. the route-giver continues to hold the pond information with lh in 6-g/e, then, there are also here and there benches. figure 14. the vr-ride towards and around the pond; one of several benches figure 13 a. indexing location of ice-cream man at bank of pond. observe that lh is still holding the pond. the pond gesture continues to be held across the follower’s new clarification request what does “sometimes” mean? (see 9-e). it only stops when the cart of the ice-cream-man is gesticulated by a pantomime using both hands (see fig. 15 below). figure 15. the cart pantomime of the route-giver (left) and the umbrella pantomime of the route-giver with right hand using lax c (right). observe that the left hand is still holding the pond38. the stop makes perfect sense, since the follower has to turn off to the right at the location of the ice-cream man. the route-giver’s gestural message is now that we are done with the route around the pond and the pond-gesture can be dropped for the time being. it is again taken up in a subsequent extensive repeat not discussed in this paper. résumé and look-back: we started with mapping information chunks of the datum onto the formal agents using, e.g., saga pauses and intonation contours. we then began the discussion of the multi-modal anaphora in the datum and explained how the anaphora resolution is working in λψ. multi-modal antecedents consisting of speech and gesture are generated by a fusion device combining speech semantics and gesture semantics in one representation as suggested in the multi 38 observe that the umbrella stick is gesturally larger than the pond and also separated from it. so, if we take that at face value, we have a gestural contradiction: truth functionality of gesture gets lost. rieser 66 modal fusion literature in sec. 3.6. the fused information is sent on ψ’s output channel to the outside and can be received by some anaphoric representation for pronouns, adverbs or definites showing a corresponding input channel. the linking of the antecedent information to the anaphoric site is achieved through ψ’s channel concept and the interface of ψ-names and λ-variables. the outgoing multimodal information can be distributed to several receiving ports broadcasting across stretches of non-alignable speech, thus showing the independence of gesture and speech. due to the communication concept of λψ, the anaphora intuitions laid out are different from what we reported in the literature overview (sec.s 3.1, 3.2, 3.3). they are an inverse to those proposals, which, starting from syntactic anaphora information, look bottom up for exactly one antecedent. essentially, the ly-calculus’ output-input device captures the “familienähnlichkeit” (family resemblance) of multi-modal anaphora of different sorts using top-down communication. at the end we can look back at the sec.s 3.8 “methodology and contributions of paper” and 4.3 “data and algorithmic handling matched” and review what has been cashed in in which ways. as far as we know this is the first paper dealing with multimodal anaphora and the distribution of information in dialogue (broadcasting). these rest on the ly-techniques. for handling independent gesture information co-operating with speech information as in the arch’ and small arch case we need a communication vehicle, served by ly’s output-input agents. for speech-gesture concurrency as with vehicle or right angle one needs a concurrency concept like the |-operator. top-down anaphora resolution also relies on the communication mechanism of ly, more specifically, the distribution of information as in the case of pond’lh, on the !-operator. no existing static grammar-bound speech-gesture account could handle the problems listed in sec. 3.8. therefore, at present, there is no alternative to lytechniques or similar ones which might exist given the broad spectrum of process algebras. 6. conclusions and further research in this paper we dealt with verbal anaphora, multimodal anaphora involving one-handed and twohanded gestures, and the broadcasting of multimodal information across speech contributions and turns. anaphora resolution is based on the λψ-calculus especially on the output-input information transfer by a channel mechanism and the same holds for broadcasting. in contrast to drt-like approaches which treat anaphora by discourse referents (drs) and accessibility relations linking the anaphor dr to the antecedent dr we transfer full terms. here we project the whole antecedent information on an output channel down to its receiving channel changing the type of expression from (existential) quantification to singular term (the “chastainian shift”, if you like), if necessary. note that outputs often get fairly complicated. in effect, if a (multi-modal) singular term has been coined, it can be sent further down. however, it may happen that no corresponding input channel exists. then the information on the firing output channel suggests a possible discourse extension not exploited. simple cases of speech-gesture integration demand that one starts with extracting information from the corpus annotation, map it onto a logical formula (a λ-term) and shift it around using ψ-techniques, always guided by observing speech-gesture asymmetry, i.e. the fact that a gesture can come too early, too late or overlap only partly with the speech semantics it is to be aligned with. in such cases, speech and gesture can be asynchronous but bounded within a construction or at least within one turn. note that broadcasting cases are different: there, gestural information is sent out without meeting a semantically fitting speech or gesture input channel in same contribution. the information held tries to communicate across constructions and turns but is kept from fusing; it is in a wait loop, so to speak. in the cases of the datum discussed broadcasting communication finally succeeds but it could also fail and peter out after some time. from a dialogue perspective, the following problems were treated in the main text: sequential and parallel dialogue contributions using ly’s agent concept, in-turn anaphora and multi-modal anaphora using y’s channel device and its input-output facilities, multi-modal anaphora in clarification requests, in multi-modal anaphora and broadcasting of information 67 turn and trans-turn communication by broadcasting gesture meaning, composition of meanings across (long) distances using ly. in the appendix we provide a trace of the axioms used and show how coherence relations and anaphora in ptt can be modelled using ly-techniques.39 now, questions worthwhile to further investigate are: which other information processes in dialogue could be described using the same ψ-techniques and which cannot, in other words, what are the limits of the λψ-techniques as developed in this paper? we first turn to the cases which concern the dialogue structure and which can possibly be handled: the route descriptions in saga are planbased insofar as the landmarks are clear from the outset and the route-giver instructs the follower how to get from one landmark to the next. selection of objects of orientation is first of all due to the route-giver. in order to proceed on the route to the next landmark/object of orientation certain paths have to be taken. on the whole, the dialogue’s “and next” structure suggests a reconstruction in terms of hobb’s occasion relation, sdrt’s narration relation or ptts context resource situations (see the literature discussion in sec. three). so, as a first starting point, we can resort to a case structure (see definition 2 in sec. four above): if the sculpture has been passed, next comes the round-about and the straight passage out of it etc. however, there seem to be other structures which are decidedly bottom-up, so for example dialogue moves hooking up to the information already given by the route-giver such as acknowledgements, clarification requests, corrections, repeats of part or whole of the bus-tour description by routegiver or follower. repeats are often ad-hoc, following route-giver’s or follower’s sporadic intentions, so the divide seems finally to be between systematic, hence expectable, dialogue moves and merely contingent ones. the systematic ones like acknowledgements could perhaps be rendered as answering to route-giver’s expectations, hence in a top-down way, but we are not sure whether this will work. there is no easy way to deal with the divide and we leave that as our next big question for a follow-up paper. acknowledgements i want to thank andy lücking who read a preliminary version of this paper. thanks go also to the anonymous d&d reviewers and to the editor for their suggestions to improve the paper. appendices 1. trace of ly-rendering of the datum in figure 2 1.1 the arch information agent(1): ® ch1 ent-i λf$x(hedge’(x) ù green’(x) ù f(x)) (ent-i) | ¬ !ch***4a < ιx(hedge’(x) ù green’(x) ù (partly-also’(with’(x, δi (doorway’(i) ù of’(i,a) ù arch’(a)))))), δi (doorway’(i) ù of’(i,a) ù arch’(a))>.0 | agent(2): ® ch2 ent-i ch***1 λyλy(y, partly-also’(with’(y))) (ent-i) | ¬ ch***1 λy (y, partly-also’(with’(δi (doorway’(i) ù of’(i,a) ù arch’(a))))).0 | agent(3): ¬! ch*****3 λfz. f(z) ù of’(z, s1) ù of’(z, s2) ù side’(s1) ù side’(s2) ù connect’(a, s1, s2) ù arch’(a).0 = def. ! ch*****3 λx(arch’(x))(a).0 | agent(4): ® ch3 arc-g ch***4 λfz(doorway’(z) ù of’(z, a) ù f(a)) (arc-g) | ch3 arc-g ch***2 λfδi (doorway’(i) ù of’(i,a) ù f(a)) (arc-g). 0 | ¬ ch***4 λz(doorway’(z) ù of’(z, a) ù λx(arch’(x))(a)).0 | ch***2 (δi (doorway’(i) ù of’(i,a) ù λx(arch’(x))(a))).0 | 39 these issues are still quite low-level in character, however, theories of stepwise turn-production, -reception, in-turn communication, acknowledgement, clarification and mutual belief could build on them mainly due to the dynamics of λψ.. rieser 68 table 1. agents for deriving the arch information tagged for input ® and output ¬.40 agent(1) gets information on ch1 from agent(2) sending on ch***1. agent(2) needs first information on input channel ch2 provided by agent(4). agent(4) is a parallel agent indicated by |, see sec. 4, definition 2. however, agent(4) can only output via ch***2 on the right hand side of | once it received a value on its ch3 input channel served by gestural agent(3) modelling the information of gesture q (see figures 3 and 4 in main text for the datum). the q information is implemented as the archinformation which can be iterated arbitrarily often due to replication “!” (see sec. 4, definition 2). communication among agent(2) and agent(1) attributes to the member of a set the property of being a hedge, green and partly also with arched doorways. so, agent(1) can send out either ιx(hedge’(x) ù green’(x) ù (partly-also’(with’(x, δi (doorway’(i) ù of’(i,a) ù arch’(a)))))) or δi (doorway’(i) ù of’(i,a) ù arch’(a)), see sec. 5.1 (d). the doorways are taken up by the anaphora these. agent(5) (see below) deals with the anaphora relation, receiving the value δi (doorway’(i) ù of’(i,a) ù arch’(a)) from agent(1) via its output-channel. a note on the multi-modal fusion perspective provided by ly as can be seen from the transcript in figure 1, 1.1-g and 1.1 rh, we have speech-gesture overlap of two uni-modal information processes, zum teil auch mit eingängen/partly also with doorways and the gesture {r-handshape c, q}; c indicates width and q the shape built by thumb, index finger and the space these enclose, oriented downwards. the {r-handshape c, q} gesture conveys independent information conceptualized in agent(3) as replication !ch***3 λx(arch’(x))(a).0, see fn. 24 on the context-problem. since partly also with is also in the scope of the gesture stroke, we have speech-gesture asymmetry; this means that the semantics of the q-gesture cannot be directly combined with the semantics of partly also with but has to be shifted until it finds the doorwaysemantics. observe that the definition of !ch***3 λx(arch’(x))(a).0 mimics the handand fingerposition of rh using side’(s1), side’(s2), and connect’. agent(3) distributes its information via !. by channel equivalence, agent(3) can communicate with agent(4) which is a concurrent agent, itself composed of two concurrent agents. λx(arch’(x))(a) fuses with the verbal doorway information by lb-conversion in both agents. the result is a multi-modal expression percolating up to communicate with agent(2), which in turn contributes to agent(1). 1.2 the die/these anaphor, the arch gesture and the park entrance agent(5): ® ch5 comp ch4a second entr-i λf z(like’(arch-shaped’)(z) ù f(z)) (comp) (entr-i).0 | ¬ δi (doorway’(i) ù like’(arch-shaped’)(i) ù of’(i,a) ù arch’(a)) ù of’(i, small-arch’) ù like’(such’$u *garden-entrance’ (u))(i)) . (like’(arch-shaped’)(δi (doorway’(i) ù of’(i,a) ù arch’(a))) ù of’((δi (doorway’(i) ù of’(i,a) ù arch’(a))), small-arch’) ù like’(such’$u *garden-entrance’ (u))(δi (doorway’(i) ù of’(i,a) ù arch’(a)))).0 | agent(6): ® ch6 comp ch***5 λfv(of’(v, small-arch’) ù f(v))(comp).0 | ¬ ch***5 λv(of(v, small-arch’) ù like’(such’$u *garden-entrance’ (u))(v))).0 | agent(7): ¬ ch***6 λs like’(such’$u *garden-entrance’ (u))(s).0 | agent(8): ® ch8 m-entr ch***9 ιv(main-entrance’(v)). λp until-sometime-then’(p) (m-entr).0 | ¬ ch***9 ιv(main-entrance’(v)). until-sometime-then’((ιv(main-entrance’(v)) = j î(δi 40 since channels, except for 0-agents, are always marked for input or output by underscore as, e.g., n or overbar as, e.g., n" (see definition 2) tags are strictly speaking not needed. however, as discussions with anonymous d&d reviewers b and c on the communication behaviour of agents showed, the communication threads among agents are easier to follow with tags. hence, they serve a useful didactic purpose. multi-modal anaphora and broadcasting of information 69 (doorway’(i) ù arch-shaped’(i) ù of’(i,a) ù arch’(a)) ù of’(i, small-arch’) ù like’(such’$u*garden-entrance’ (u)))(i) ù comes’(j))).0 | agent(9): ® ch7 p-ent ch***8 ιv(main-entrance’(v)). λz(ιv(main-entrance’(v)) = j î z ù comes’(j)) (p-ent).0 ¬ ch***8 (ιv(main-entrance’(v)) = j î(δi (doorway’(i) ù arch-shaped’(i) ù of’(i,a) ù arch’(a)) ù of’(i, small-arch’) ù like’(such’$u *garden-entrance’ (u)))(i) ù comes’(j)). 0 |. agent(10): ® ch9 m-ent λu$y((large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) ù y = u) (m-ent) .0 ¬ ch***11 <$y(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v)). 0 | agent(11): .0 | agent(12): ® ch11 garc ch14 veh ch15 quiri ch15b ev λpath λvehi λarc λf (must-in(e1, fo, vehi path, arc) ù f(e1))(garc)(veh)(quiri)(ev).0 | ¬ ! ch*****16 . (must-in(e1, fo, vehicle’, right-angle’, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v))) ù quickly(e1) ù right-into(e1, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v)))). 0 | agent(13): ¬ ! ch***15 right-angle’.0 | agent(14): ¬ ! ch***14 vehicle’.0 | agent(15): ® ch11 arc λarc-i ch***15b λe(quickly(e) ù right-into(e, arc-i)) (arc) ¬ ch***15b λe(quickly(e) ù right-into(e, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)))).0 agent(16): ® ch38 fronth(λfro($x(¬ wall’(x) ù hedge’(x) ù fro(x))))(fronth).0 | agent(16a):¬ ch***38 lz of’(z, front’)) .0 | agent(17): ® ch11 arc-i ch16 second ev ch14 ve ch15 pa λarc λe λvehi λpath ?into’(e, fo, vehi, path, in’(ιu (arc’(u)) = arc))(arc-i)(ev)(ve)(pa).0| ¬ ch***36 (ιu (arc(u)) = ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’))).0 | agent(18): ® ch16 second e ch14 veh ch15 sra ch11 arc λev λvehi λpath λarc(into’(ev, fo, vehi, path, in’ (ιu (arc(u)) = arc)))(e)(veh)(sra)(arc). 0 | ¬ ch***37 into’(e1, fo, vehicle’, right-angle‘, in’ (ιu(arc(u)) = ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’))). 0 | table 2: agents needed to derive the die/these anaphor, the arch gesture and the park entrance gestures with tags for agents’ input ® and agents’ output ¬. agent(5) gets input on channel ch5 for the comparison clause like such a garden entrance. agent(6) and (7) provide that. agent(7) implements the comparison.41 the information on ch***6 is received by agent(6), providing the input for agent(5). the final derivation for agent(5) is: like’(arch-shaped’)(δi (doorway’(i) ù of’(i,a) ù arch’(a))) ù (of’(δi (doorway’(i) ù of’(i,a) ù arch’(a)), small-arch’) ù like’(such’$u *garden-entrance’ (u))( δi (doorway’(i) ù of’(i,a) ù arch’(a)))).0 | this means that each i in the set operated on by δ that is an arched doorway, i.e., is of’(i,a) ù arch’(a), is like’(arch-shaped’) [first conjunct], has a small-arch’ [second conjunct] and is like such a garden entrance [third conjunct]. some of this information will be sent out by ch***7 given that it finds a corresponding input channel: 41 for the * used here, see kamp and reyle (2003), p. 327, fn 13. rieser 70 ch****7 (δi (doorway’(i) ù like’(arch-shaped’)(i) ù of’(i,a) ù arch’(a)) ù of’(i, small-arch’) ù like’(such’$u *garden-entrance’ (u))(i)) . (like’(arch-shaped’)(δi (doorway’(i) ù of’(i,a) ù arch’(a))) ù of’((δi (doorway’(i) ù of’(i,a) ù arch’(a))), small-arch’) ù like’(such’$u *gardenentrance’ (u))(δi (doorway’(i) ù of’(i,a) ù arch’(a)))).0 |. this solves the anaphoric relation among doorways and these; observe that these, the anaphora, is represented as δi(doorway’(i) ù of’(i,a) ù arch’(a)), i.e. as the set of arched doorways. the value in the prefix says that each arched doorway i is arch-shaped and has a small arch and is like such a garden entrance. technically, we assemble properties originally attributed to the δ-set into the δ-set itself following the chastainian observation of referential term extension in sec. 5.1 (b). after action on ch7 agent(9) looks as follows: ch***8 (ιv(main-entrance’(v)) = j î(δi (doorway’(i) ù arch-shaped’(i) ù of’(i,a) ù arch’(a)) ù of’(i, small-arch’) ù like’(such’$u *garden-entrance’ (u)))(i) ù comes’(j)). 0 |. on the convention that an existentially quantified expression can send a singular term, we get for agent(10): = ch***11 <ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v))>. $y(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v)). 0 | remember that an identity can proceed with either term according to 5.1 (b) c. the output derivation for agent(12) is then: !ch***16 <ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v)), e1>. (must-in(e1, fo, vehicle’, right-angle’, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(mainentrance’(v))) ù quickly(e1) ù right-into(e1, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v)))). 0 | agent(13) and agent(14) specify the right-angle and the vehicle information, both given by gesture, respectively. agent(15a) specifies the moving into the arch. agent(17) embodies the follower’s clarification request. 1.3 post-holds and broadcasting agent(12): ® ch11 garc ch14 veh ch15 quiri ch15b ev λpath λvehi λarc λf (must-in(e1, fo, vehi, path, arc) ù f(e1))(garc)(veh)(quiri)(ev).0 | ¬ !ch***16 . (must-in(e1, fo, vehicle’, right-angle’, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v))) ù quickly(e1) ù right-into(e1, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v)))). 0 | agent(13): ¬ !ch***15 right-angle’.0 | agent(14): ¬ !ch***14 vehicle’. 0 | agent(15): ® ch11 arc λarc-i ch***15b λe(quickly(e) ù right-into(e, arc-i)) (a rc) ¬ ch***15b λe(quickly(e) ù right-into(e, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)))).0 agent(16): ® ch***38 fronth(λfro($x(¬ wall’(x) ù hedge’(x) ù fro(x))))(fronth).0 | agent(16a): ¬ ch***38 λz(of’(z, front’)) .0 | agent(17): ® ch11 arc-i ch16 second ev ch14 ve ch15 pa λarc λe λvehi λpath ?into’(e, fo, vehi, path, in’(ιu (arc’(u)) = arc))(arc-i)(ev)(ve)(pa).0| ¬ ch***36 (ιu (arc(u)) = ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’))).0 | multi-modal anaphora and broadcasting of information 71 agent(18): ® ch16 second e ch14 veh ch15 sra ch11 arc λev λvehi λpath λarc(into’(ev, fo, vehi, path, in’(ιu (arc(u)) = arc)))(e)(veh)(sra)(arc). 0 | ¬ ch***37 into’(e1, fo, vehicle’, right-angle’, in’ (ιu(arc(u)) = ιy(large’(y) ù grey’(y) ù (y) ù of’(y, curb’))). 0 | agent(12): ® ch11 garc ch14 veh ch15 quiri ch15b ev λpath λvehi λarc λf (must-in(e1, fo, vehi, path, arc) ù f(e1))(garc)(veh)(quiri)(ev).0 | ¬ !ch***16 . (must-in’(e1, fo, vehicle’, right-angle’, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v))) ù quickly’(e1) ù right-into’(e1, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v)))).0 | agent(13): ¬ ! ch***15 right-angle’ .0 | agent(14), (15), (16): see table 2. agent(17): ® ch11 arc-i ch16 second ev ch14 ve ch15 pa λarc λe λvehi λpath ?into’(e, fo, vehi, path, in’(ιu (arc’(u)) = arc))(arc-i)(ev)(ve)(pa).0 | ¬ ch***36 ? into’(e1, fo, vehicle’, right-angle’, in’(ιu (arc(u)) = ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)))).0 | agent(20) = case agent(21): agent(22).0 | agent(21): ® ch15a pa-i ch14 veh-i ch11 ther-i λpth1 λveh λthere (have-driven-in’(e1, fo, veh, pth1,there)) (pa-i)(veh-i) (there-i) .0 | ¬ ch***16 path1’. (have-driven-in’(e1, fo, vehicle’, path1’, ιx (large’(x) ù grey’(x) ù arch’(x) ù of’(x, curb’)))).0 |. ch***17 agent(22): ® ch15c pa-i ch14 veh-i ch11 the-i ch18 polh ch19 porh λpth2 λveh λthere <ιx(pond’(x)), ιx(pondlh’(x))>. (drive’(e2, fo, veh, pth2, there) ù straight-forward’(e2, veh, pth2) ù towards’(e2, fo, veh, λrs $x (pond’(x) ù x = r ù x = s)))(pa-i)(veh-i)(the i)(polh)(porh).0| ¬ ! ch***17 <ιx(pond’(x)), ιx(pondlh’(x))>. (drive’(e2, fo, vehicle’, path2’, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v))) ù straight-forward’(e2, vehicle’, path2’) ù towards’(e2, fo, vehicle’, $x(pond’(x) = pond’lh = pond’rh))).0| agent(23) = case agent(24): agent(25).0 | agent(24): ¬ ch***20 pond’lh. at’(ιx (pond’(x)) = pond’lh).0 | agent(25) = case agent(26): agent (27).0 | agent(26): ® ch17 first po-i ch14 veh-i ch21 pa-i λpo λvh λpth3 λpolh (drive’(e3, fo, vh, pth3) ù towards’(e3, fo, vh, pth3, po) ù po = pond’lh)(po-i)(veh-i)(pa-i).0 | ¬ (drive’(e3, fo, vehicle’, path’rh) ù towards’(e3, fo, vehicle’, path’rh, pond’lh) ù ιx(pond’(x)) = pond’lh).0 |. agent(27): ® ch14 veh-i ch21 pa-i ch22 pa-ic ch20 po-i λvh λpth3 λpth4 λpolh (drive’(e3, fo, vh, pth3) ù right-around’(e3, fo, vh, pth4, polh))(veh-i)(pa-i)(pa-ic)(po-i) .0 | ¬ ch***22 path-c’rh.(drive’(e3, fo, vehicle’, path’rh) ù right-around’(e3, fo, vehicle’, path-c’rh, pond’lh)).0 | agent(31): ¬ ch***25 (runs’(e4, (ιx(hedge’(x) ù green’(x) ù (partly-also’(with’(x,y) ù doorway’(y) ù of’(y, a) ù arch’(a)))), z) ù z = another‘(roughly‘|50|meter’))). 0 | agent(33): .0 | agent(34): ¬ ch***26 environs’.0 | agent(37): ® ch26 area ch30 ben1 ch31 ben2 ch28 poss-ben λloc λloc-ben1 λloc-ben2 λben(at’(loc, of(ben, here’, loc-ben1) ù of’(ben, there’, loc-ben2))) (area)(ben1)(ben2)(poss ben) .0| (at’(environs’, of’(perhaps’(δi (bench’ (i))), here’, ι loc1(of (δi (bench’(i)), rieser 72 loc1))) ù of(perhaps’(δi (bench’ (i))), there’, ι loc2 (of (δi (bench’(i)), loc2))))) .0 | agent(41): accommodated. agent(42): ® ch14 veh-i ch22 pth-i ch18 polh-i λvh λpth λpolh (drive’(e6, fo, vh, pth) ù round’(ιx (pond’(x)))(e6) ù ιx (pond’(x)) = polh ù right’((around’)(polh))(e6))(veh-i)(pth-i) (polh-i) .0 | agent(43): ¬ !ch***34 ιx(ice-cream-man(x)). $x(ice-cream-man(x)).0 | agent(44): ¬ ch***35 ιz(loc(z)).0 | agent(45): ¬ ch***36 $y sometimes’(at(ιx(ice-cream-man’(x)), ιz(loc(z))) ù boundary’(y) ù of’(y, pond’lh) ù part-of(ιz(loc(z)), y)).0 agent(46): ® ch14 veh-i ch35 loc-icm λveh λloc (drive-off(e7, fo, veh-i, at(loc)) ù to-right(e7)) (veh-i)(loc-icm).0 | drive-off(e7, fo, vehicle’, at(ιz(loc(z))) ù to-right(e7)).0 | agent(47): ¬ ch***37 r. ask(fo, p) ù p = $q imply(sometimes’(q), r).0 | agent(48): ® ch37 chang ch34 icr-m ch35 loc-icm λchang λice-crm λloc could-be(chang = $y(changed’(e8, y, q) ù q = be(ice-crm, at(ice-crm-loc))) (chang)(icr-m)(loc-icm) .0| could-be(r = $y(changed’(e8, y, q) ù q = be’(ιx(ice-cream-man’(x)), at(ιz(loc’(z)))))) agent(49a) = case agent(49): agent(50) | agent(51): .0 | agent(49): ® ch40 tour-icm ch41c-umb λp q"e(case e = e9: p ù q)(tour-icm)(c-umb) .0 | ¬ "e(case e = e9: (was’ (e9, ιx(ice-cream-man’(x)), on-route-giver’stour’at(ιz(loc’(z))))) ù $x was’(cart’(x) ù of’(x, ιx(ice-cream-man’(x))) ù of’(x, y) ù handle-bar’(y) ù horizontally-grasped’(y) ù of’(x, u) ù umbrella’(u) ù of’(u, v) ù vertical-handle’(v))). 0| agent(50): ® ch34 icr-m ch35 loc-icm λicr-m λloc-icm ch***40 on-route-giver’s-tour’(was’ (e9, icr m, at(loc-icm)))(icr-m)(loc-icm) .0 | ¬ ch***40 on-route-giver’s-tour’(was’(e9, ιx(ice-cream-man’(x)), at(ιz(loc’(z))))) .0 | agent(51): ® ch34 icr-m ch38 cart-h ch39 umb-h λicr-m λcart-h λumb-h ch***41 ($x was’ (cart’(x) ù of’(x,icr-m) ù of’(x, y) ù cart-h’(y) ù of’(x, z) ù umbrella’(z) ù of’(z, v) ù umb h’(v)))(icr-m)(cart-h)(umb-h) .0 | ¬ ch***41 ($x was’ (cart’(x) ù of’(x, ιx(ice-cream-man’(x))) ù of’(x, y) ù handle-bar’(y ) ù horizontally-grasped’(y) ù of’(x, u) ù umbrella’(u) ù of’(u, v) ù vertical handle’(v))) .0 | agent(52): ¬ ch***38 λu (handle-bar’(u) ù horizontally-grasped’(u)).0 | agent(53): ¬ ch***39 λz (vertical-handle’(z)).0 | table 3: agents needed to model post-holds and broadcasting tagged for input ® and output ¬. for the route-giver’s 3.2-g/e we use the case construct of ψ: if you have driven in there works as antecedent φ and you drive straight towards a pond. a pond as consequent (see sec. 4, definition 3 for the general format). the anaphora there is resolved to a large grey arch becoming definite. agent(21) amounts in the end to ch****16path1’.(have-driven-in’(e1, fo, vehicle’, path1’, ιx (large’(x) ù grey’(x) ù arch’(x) ù of’(x, curb’)))).0 |. ch15a contributes path1’, ch***14 vehicle’, ch***11 the large grey arch with curb’. the initial section of the path, path1’, is returned on ch***16.the consequent you drive straight towards a pond. a pond. is multi-modal anaphora and broadcasting of information 73 modelled as agent(22). assuming pondlh as value on ch18 and path2’ as value on ch15c 42 agent(22) turns into ! ch***17 <ιx(pond’(x)), ιx(pondlh’(x))>. (drive’(e2, fo, vehicle’, path2’, ιy(large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) = ιv(main-entrance’(v))) ù straight-forward’(e2, vehicle’, path2’) ù towards’(e2, fo, vehicle’, $x(pond’(x) = pond’lh = pond’rh ).0|. we need the path-into, pa-i, its value coming from ch15c, the vehicle, the main entrance, where the bus went in by ch11 and a left hand and a right-hand o-shape gesture, both denoting the pond. given these, we end up with the derivation of the broadcasting agent(22). the route around the pond (see vr figure 12a) is represented as follows: we have an embedded conditional structure: agent(23) = case agent(24): agent(25) and agent(25) = case agent(26): agent(27) with the noticeable occurrence of the linguistic pond information syntactically in the prefield (see sec. 5.1 on this notion). the whole conditional structure is, agent(24) and agent(26) covering respectively the role of φi in the general definition (see sec. four, definition 2): agent(24) represents the antecedent and at this pond, agent(26) the antecedent you drive towards it. in specifying agent(26) we use the linguistic pond information on ch17first (distributively communicated by agent(22)), the vehiclegesture information we already have and an assumed path drawn by the rh gesture on ch21, path’rh, again accommodated, shown in fig. 12a. the lh pond gesture pond’lh is held all the time as indicated by iterative “!”. the central movement information is conveyed by the rh gesture path’rh. we get as final derivation for agent(26), observe the gestural representation of vehicle’, path’rh, and pond’lh: (drive’(e3, fo, vehicle’, path’rh) ù towards’(e3, fo, vehicle’, path’rh, pond’lh) ù ιx(pond’(x)) = pond’lh).0 |. agent(27) represents and you drive right around it. its final derivation is ch***22 path-c’rh.(drive’(e3, fo, vehicle’, path’rh) ù right-around’(e3, fo, vehicle’, path-c’rh, pond’lh)).0 |, where, looking at the λ-variables, the parameters for vehicle, vh, the path towards the pond, pth3, the rh-path around the pond, pth4, and the lh-gestured polh are filled by the values provided by the input channels. note that substantial information is expressed by the lh and rh gestures. the path-information path-c’rh is sent to the outside and will be taken up in agent(42). next, the route-giver refers to the hedge (5-g/e), the hedge, syntactically again situated in the prefield and resumed in the middle field. the long-distance anaphora for the hedge is resolved with the broadcasting hedge-information from above, i.e. by agent(1) via channel !ch***4afirst. this is embodied in agent(31) which is derived as ch***25(runs’(e4, (ιx(hedge’(x) ù green’(x) ù (partly-also’(with’(x,y) ù doorway’(y) ù of’(y, a) ù arch’(a)))), z) ù z = another‘(roughly‘|50|meter’))). 0 |. while the route-giver produces the description the hedge, {it runs another roughly} 50 m., we might expect that the hedge is at issue, nevertheless it is the l-handshape o with denotation pond’ which is broadcasted independently from speech. the route-giver’s next remark is concerned with the environs (6g/e): then there are also here and there benches. again, the l-handshape o is continued. the indexing d-shape shows the temporal asymmetry of gesture and speech characteristic for iconic gesture use. the first indexing is too early, ideally it should be aligned with hin [und wieder]/here [and there]. 42 again, observe that pondlh as value on ch19 and path2’ on ch15c are both accommodated. a more elaborate reconstruction would have to express that path2’ continues path1’ in the direction towards the pond. we forego the more detailed reconstruction here. rieser 74 assume an agent (34) which provides a value environs’ for an indexical gesture and sends it out on ch***26. then agent(37) does the integration work. it gets input from agent(34) which provides the value environs’ on channel ch26. assume agent(41) contributes the “doubtful benches”, possben, on ch28. then the final derivation for agent(37) is (at’(environs’, of’(perhaps’(δi (bench’ (i))), here’, ι loc1(of (δi (bench’(i)), loc1))) ù of(perhaps’(δi (bench’ (i))), there’, ι loc2 (of (δi (bench’(i)), loc2))))) .0 |. there is no outgoing transport channel for this information stressing the singularity of this contribution as in the case of the hedge information. 7-g/e essentially repeats but you drive around the pond. right around, modelled in agent(42). so, almost the same agents for input are active as in agent(22): ch14 veh-i contributes the vehicle information, ch22 pth-i the gestured path around the pond, path-c’rh, and ch18 polh the pond assigned by the left hand. the pond is a non-official landmark, important for turning to the right and continuing the tour to the chapel and on to the fountain. the turn-off to the right is at an ice-cream cart, referred to in 8-g/e: and sometimes there is an ice-cream man there. and at him you drive off to the right. we need new entities for the ice-cream man and his location at the embankment of the pond. the “pond-gesture”, small o, is held as before (see fig. 12a). therefore, we have in agent(45): ch***36 $y sometimes’(at(ιx(ice-cream-man’(x)), ιz(loc(z))) ù boundary’(y) ù of’(y, pond’lh) ù partof(ιz(loc(z)), y)).0 |, where agent(43) and agent(44) have contributed, respectively, the ice-cream man and his location. we hand on the location of the ice-cream man ιz(loc(z)) which we need to resolve the anaphora at him. we have the pond gesture as before and a g-bound gesture indexing the ice-cream man’s location. the datum shows that the r-handshape g being on {nen eisverkäufer}/{‘n ice-cream man} comes too early. ideally, it should align with da/there. we represented “manchmal/sometimes” as propositional operator in agent(45) which is sufficient for our purposes43. using agent(14) and agent(44), agent(46) expands to drive-off(e7, fo, vehicle’, at(ιz(loc(z))) ù to-right(e7)). at this stage of the dialogue we get an interesting clarification request of the follower’s (9g/e): “was heißt ‘manchmal’?/”what does ‘sometimes’ mean”?. due to the polysemy of german “heißen”/”mean” it can be read as “what does ‘sometimes’ imply/come to”? so we render it as agent(47)44. 2. further applications 2.1 rhetorical relations in the literature discussion sec. 3.1 we pointed out the importance of semantic/pragmatic relations among sentence pairs for anaphora resolution as observed in the hobbs-hume-kehler-rohde research line. why does one, in addition, need rhetorical relations? essentially, they cover the intuition that some parts of a discourse are more tightly bound together than others as exemplified 43 we do not want to go into quantifying across time intervals in this paper which an existential quantification for sometimes’ would need. 44 andy lücking attributes the follower’s clarification request to ginzburg’s “intended content clarification”, cf. ginzburg (2012), p. 153. multi-modal anaphora and broadcasting of information 75 in sdrt (al, p. 147). the primary task of rhetorical relations and the right border constraint is to enable anaphora resolution which, as will be clear by now, we do differently, by direct message passing. but we do admit that there is a marked substructure in the route-giver’s directives in the datum: first we have the introduction of the hedge with its doorways and the shape of the doorways, then we get the main-entrance with its properties and so on. we assume that dann + sind/then + are acts as a cue phrase for continuation and dann + finite verb as a cue phrase for narration. we let rhetorical relations supervene on anaphora resolution below. in order to reduce discourse information we use the following simplified version of the given datum: 1-e there are doorways. 1.1-rh {r-handshape q } 1.2-e these are then arch-shaped. 1.2-rh {r-handshape small-arch-shape} 2-e [until sometime] then the main entrance comes. 3-e that is a large grey arch. 3.1-lh {r-handsh. large arc-shape + small c} table2: simplified version of datum. we have by increments a1 < a2 < ,….,< a7 and a1 | a2 | , …., |a7. agent(1) captures the doorways plus an accompanying r-handshape q gesture. agent(2) introduces the continuation relation between this information and the doorways being arch-shaped. the co-verbal gesture signs small arches for the doorways. agent3 shows that these is translated into an input channel plus ψname dw1 receiving the doorway-set δi (doorway’(i) ù of’(i,a) ù arch’(a)) from agent(1). agent(3) hands on the information that each i in the δ-set is arch-shaped and has a small arch. agent(4) encodes the narration relation mapping the generated continuation relation onto the main entrance information, the main entrance being one of the doorways with a small arch. agent(5) outputs the main entrance information and the fact that the main entrance comes in sight on the tour. agent(6) continues the information we have so far with a description of the main entrance being a large grey arch. the gesture depicts a large arc plus the curb-information using small c. by message passing, the main-entrance is identified with the entity being a large grey arch. rhetorical relations are treated like anaphora, agents sending multi-modal propositional information interact with ports receiving this information and accumulating it. this makes rhetorical relations multi-modal anaphoric entities. in sum, our reconstruction of the simplified datum captures the following issues: incrementality, multi-modal fusion of gesture and speech, anaphora resolution, and rhetorical relations continuation and narration45. agent(1) ch***1 (δi (doorway’(i) ù of’(i,a) ù λx(arch’(x))(a))).0 |46 agent(2) ch3 dw ch1 dw2 ch***4 λp1p2 continuation(p1, p2)(dw)(dw2) = ch***4 continuation((arch-shaped’(δi (doorway’(i) ù of’(i,a) ù arch’(a)) ù of’(δi (doorway’(i) ù of’(i,a) ù arch’(a)), small-arch’))), δi (doorway’(i) ù of’(i,a) ù arch’(a))).0 | agent(3) ch1 dw1 ch***3 λ z (arch-shaped’(z) ù of’(z, small-arch’))(dw1) = ch***3 (arch-shaped’(δi (doorway’(i) ù of’(i,a) ù arch’(a)) ù of’(δi (doorway’(i) ù of’(i,a) ù arch’(a)), small-arch’))).0 | 45 we do not use rhetorical relations to tie speech semantics and gesture semantics together as proposed in lascarides and stone (2009). see 3.6 in the literature section on their proposal in the context of multi-modal fusion. 46 agent1 is simplified inasmuch as the whole process of generating the arch-information and attributing it to the doorways is not shown. since the focus here is on rhetorical relations this will not matter. rieser 76 agent(4) ch4 prop1 ch5 second prop2 ch***6 λp1p2 narration(p1, p2)(prop1)(prop2) = ch***6 .0 | agent(5) ! ch***5 < ιv(main-entrance’(v)) = j î(δi (doorway’(i) ù arch-shaped’(i) ù of’(i,a) ù arch’(a)) ù of’(i, small-arch’)), (ιv(main-entrance’(v)) = j î(δi (doorway’(i) ù arch-shaped’(i) ù of’(i,a) ù arch’(a)) ù of’(i, small-arch’)) ù comes’(j)).0> | agent(6) ch6 first prop2 ch7 prop1 λp1p2 continuation(p1, p2)(prop2)(prop1) = continuation(narration(continuation((arch-shaped’(δi (doorway’(i) ù of’(i,a) ù arch’(a)) ù of’(δi (doorway’(i) ù of’(i,a) ù arch’(a)), small-arch’))) , δi (doorway’(i) ù of’(i,a) ù arch’(a))) , $y((large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) ù y = ιv(main-entrance’(v)) = j î(δi (doorway’(i) ù arch-shaped’(i) ù of’(i,a) ù arch’(a) ù of’(i, small-arch’))))) agent(7) ch5 first me ch***7 λu$y((large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) ù y = u)(me) = ch***7 $y((large’(y) ù grey’(y) ù arch’(y) ù of’(y, curb’)) ù y = ιv(main-entrance’(v)) = j î(δi (doorway’(i) ù arch-shaped’(i) ù of’(i,a) ù arch’(a) ù of’(i, small-arch’)))). 1.2 anaphora in ptt poesio and rieser (2011) focussed on the resolution of definite anaphora, nps and pronouns from the perspective of incrementality in ptt. the main insights were47: generalizing löbner’s account of definites (loebner 1987) to instantiate discourse referents. use the resource situations of situation semantics to capture pointings, anaphoric reference and the function of rhetorical relations. capture experimental results of the visual world paradigm. use the ltag formalism for encoding incrementality using prioritized defaults for all levels of the parsing tool developed. developing the concept of micro-conversational events (mces). for illustration purposes we give a full ptt-version of 1-e below48: [k1.1, up1.1, ce 1.1| k1.1 is [x| x is δi (doorway’(i))], up1.1: utter(route-giver, follower, “there are doorways”), sem(up1.1) is k1.1, ce1.1: assert(route-giver, follower, k1.1), generate(up1.1, ce1.1). we want to show how λψ-ideas can help to account for two types of dynamics, the speech-gesture interface and the resolution of multi-modal plural anaphora. as in ptt in general, we also rely on muskens’ compositional drt (cdrt) as a background formalism. the division of labour between cdrt and the λψ-calculus is as follows: box-internal regularities follow the cdrt-regime, communication among box-agents is organized in the λψ-manner. we restrict our focus to semantic representations; up1.1, ce1.1, and generate are not dealt with (see however the remark about further simulations below). agents are now taken as a dynamic version of the mces in poesio (1995). information in red is again marked as contributed by gesture. 47 we do not give detailed explanations and literature references here. see the poesio and rieser (2011) paper on that. 48 k is a discourse referent in cdrss, up1.1 stands for “utterance phrase”, sem for “semantics”, ce1.1 for conversational event, i.e. a discourse act, utterance phrases generate conversational events. multi-modal anaphora and broadcasting of information 77 agent(1) ch***1 x. [x|x is δi (doorway’(i))]; agent(2) ch1 dw λy ch***2 z. [ | of(y,a), arch(a), z is δi (doorway’(i), of’(i,a), λx(arch’(x))(a))] (dw); agent(3) ch2 dwa λu ch***3 v. [ | arch-shaped(u), u is v](dwa); agent(4) ch3 dwas λw ch***4 y. [ | of(w,a), small-arch(a),y is δi (doorway’(i), of’(i,a), λx(arch’(x))(a), small-arch(a))]. the idea is as follows: agent(1) sends the discourse referent for the plural term δi (doorway’(i)) out which agent(2) receives49. agent(2) sends out the multi-modal information tied to doorway. agent(3) attributes arch-shaped to the multi-modal term and sends the input out. agent(4) adds multi-modal information and sends the multi-modal δi (doorway’(i), of’(i,a), λx(arch’(x))(a), small-arch(a)) out. the set variables x, z, v, y do essentially the duty of the resource situation variable in the poesio and rieser (2011) paper. in a similar manner up, ce, and generate can be treated. references the next phase, which syntax is undergoing samson abramsky (2008). information, processes and games. in p. adriaans and j. van benthem, editors. philosophy of information. vol. 8, pages 483 – 549 elsevier, amsterdam katya alahverdzhieva & alex lascarides (2010). analysing language and coverbal gesture in constraint-based grammars. in proceedings of the 17th international conference on head-driven phrase structure grammar, pages 5-25. http://makino.linguist.univ-parisdiderot.fr/files/hpsg2010/file/abstracts/hpsg/alahverdzhieva-lascarides-hpsg2010.pdf katya alahverdzhieva & alex lascarides (2011). an hpsg approach to synchronous speech and deixis. in proceedings of the 18th international conference on head-driven phrase structure grammar, pages 6–24, university of washington, stanford, ca: csli publications. http://cslipublications.stanford.edu/hpsg/2011/alahverdzhieva-lascarides.pdf nicholas asher (1993). reference to abstract objects in discourse. kluwer academic publishers, dordrecht. isbn 0-7923-2242-8 nicholas asher and alex lascarides (2003). logics of conversation. cambridge university press, cambridge. isbn 0521650585 j. c. m. baeten (2004). a brief history of process algebra (pdf). rapport csr 04-02. vakgroupinformatica, technische universiteit eindhoven jon barwise and jerry seligman (1997). information flow. the logic of distributed systems. cambridge university press, cambridge isbn 0 521 58386 1 srinivas bangalore and michael johnston (2009). robust understanding in multimodal interfaces. computational linguistics, 35(3): 345-397. http://dx.doi.org/10.1162/coli.08-022-r2-06-26 jesper bengtson, magnus johansson, joachim parrow, and victor björn (2011). psi-calculi: a framework for mobile processes with nominal data and logic. logical methods in computer science 7 (1,11): 1-44. arxiv:1101.3262 [cs.lo] 49 according to our identity condition 5.1 (b).c we could also send the term to the right of the “is”. rieser 78 kirsten bergmann, hannes rieser and stefan kopp (2011). regulating dialogue with gesture. towards an empirically grounded simulation with conversational agents. in proceedings of sigdial 2011, pages 88-97 j. a. bergstra, a. ponse, s.a. smolka, editors (2001). handbook of process algebra. online resource https://eu04.alma.exlibrisgroup.com/view/uresolver/49hbz_bie/openurl?u.ignore_date_coverag e=true&portfolio_pid=53309823570006442&force_direct=true johannes borgström, shuqin huang, magnus johansson, palle raabjerg, björn victor, johannes ˚aman pohjola, and joachim parrow. broadcast psi-calculi with an application to wireless protocols. software and system modeling, 14(1):201–216, 2015. susan, e. brennan, marilyn w. friedman, carl pollard (1987). a centering approach to pronouns. in proceedings of the 25th acl, pages 155-162 stanford. http://dx.doi.org/10.3115/981175.981197 charles chastain, (1975). reference and context. university of minnesota press, minneapolis. retrieved from the university of minnesota digital conservancy, http://hdl.handle.net/11299/185224 herbert h. clark, (1977). bridging. in p. n. johnson-laird and p.c. wason, editors. thinking: readings in cognitive science, pages 411-420 cambridge university press, cambridge. https://www.aclweb.org/anthology/t75-2034.pdf thomas leonard and fred cummins, (2011). the temporal relation between beat gestures and speech. language and cognitive processes 26: 1457-1471. doi: 10.1080/0169010965.2010.500218 mariangiola dezani-ciancaglini, (1996). logical semantics for concurrent lambda calculus. proefschrift, nijmegen universiteit, nijmegen,the netherlands. miriam eckert and michael strube, (1999). resolving discourse deictic anaphora in dialogues.in proceedings of the eacl’99, pages 37-44. http://dx.doi.org/10.3115/977035.977042 jacob eisenstein, randall davis, (2006). gesture features for coreference resolution. in s. renals, s. bengio, j. fiscus, editors. mlmi 2006, pp. 154-165. doi: 10.1007/11965152_14 jay eisenstein and randall davis, (2006). gesture improves coreference resolution. in proceedings of the human language technology conference of the north american chapter of the acl, pages 37-40, new york, june. https://doi.org/10.3115/1614049.1614059 anna esposito, marc-eric mccullough, francis quek (2001). disfluencies in gesture: gestural correlates to filled and unfilled speech pauses. researchgate anna esposito, antonietta m. esposito (2011). on speech and gestures synchrony. in: esposito, a., vinciarelli, a., vicsi, k., pelachaud, c., nijholt, a., editors. analysis of verbal and nonverbal communication and enactment. the processing issues. lecture notes in computer science, vol 6800. springer, berlin, heidelberg. https://doi.org/10.1007/978-3-642-25775-9_25 multi-modal anaphora and broadcasting of information 79 wan fokkink (2000). introduction to process algebra. springer-verlag berlin, heidelberg isbn:978-3-540-66579-3 kari fraurud (1990). definiteness and the processing of nps in natural discourse. journal of semantics, 7 (issue 4):395–433. https://doi.org/10.1093/jos/7.4.395 bart geurts, david i. beaver, and emar maier, "discourse representation theory", the stanford encyclopedia of philosophy (spring 2020 edition), edward n. zalta, editor, url = . jonathan ginzburg (2012). the interactive stance. oxford university press, oxford jonathan ginzburg and andy lücking (2021a). i thought pointing is rude: a dialogue-semantic analysis of pointing at the addressee. in p. g. grosz, l. marti, h. pearson, y. sudo, and s. zobel, editors, proceedings of sinn and bedeutung 25, pages 276-291. university college london and queen mary university of london, https://ojs.ub.uni-konstanz.de/sub/index.php/sub/article/view/937 jonathan ginzburg and andy lücking (2021b). requesting clarifications with speech and gesture. conference beyond language: multimodal semantic representation. virtual at the university of groningen, held in conjunction with iwcs 2021. jonathan ginzburg and andy lücking (2022). leading voices. dialogue semantics, cognitive science, and its polyphonic structure of multimodal interaction. preprint in language and cognition. november 2022. gianluca giorgolo and frans verstraten (2008). perception of speech-and-gesture integration. in proceedings of the international conference on auditory-visual speech processing 2008, pages 31–36 ramūnas gutkovas (2014). languages, logics, types and tools for concurrent system modelling. dissertation from the faculty of science and technology 1392. upsala universitet, upsala florian hahn, hannes rieser (2011). gestures supporting dialogue structure and interaction in the bielefeld speech and gesture alignment corpus. in semdial 2011: proceedings of the 15th workshop on the semantics and pragmatics of dialogue, pages 182-183, los angeles, california, 21-23 sept. 2011 hubert haider (2006). mittelfeld phenomena (scrambling in germanic). in martin everaert and henk van riemsdijk, editors, the blackwell companion to syntax,vol. iii, pages 204-275, blackwell publishing. http://dx.doi.org/10.1002/9780470996591.ch43 john a. hawkins (1978). definiteness and indefiniteness. lonndon, croom helm. klaus von heusinger (2007) accessibility and definite noun phrases. in m. schwarz-friesel, m. consten, and m. knees, editors. anaphors in text. cognitive, formal and applied approaches to anaphoric reference, pages 123-145, john benjamins, amsterdam. https://doi.org/10.1075/slcs.86.12heu. jerry, r. hobbs, 1978, resolving pronoun references. lingua 44, (1978), 4:311-33 https://doi.org/10.1016/0024-3841(78)90006-28. rieser 80 jerry, r. hobbs, 1979. coherence and coreference. cognitive science 3: 67-90. https://doi.org/10.1207/s15516709cog0301_4 tilman, n. höhle (1986). der begriff “mittelfeld”. anmerkungen über die theorie topologischer felder. in: schöne, albrecht (hrsg.) 1986. kontroversen alte und neue. akten des vii. internationalen germanisten-kongresses, göttingen 1985. bd. 3: weiss, walter e., wiegand, herbert, e. und reis, marga (hrsgg.),textlinguistik kontra stilistik? –wortschatz und wörterbuch – grammatische oder pragmatische organisation von rede?, pages 329-340. https://doi.org/10.5281/zenodo.1169673 david hume (1748). an enquiry concerning human understanding. the liberal arts press, new york, 1955 edition. katja jasinskaja, elena karagjosova (2020). rhetorical relations. in the wiley and blackwell companion to semantics, pages 1-29. doi: 10.1002/9781118788516.sem061 magnus johansson (2010). psi-calculi: a framework for mobile process calculi: cook your own correct process calculus just add data and logic. phd thesis, uppsala university, division of computer systems, 2010. michael johnston (2019). multimodal integration for interactive conversational systems. in sharon oviatt et al., editors. the handbook of multimodal-multisensor interfaces, vol. 3, pages 21-77, acm books. doi:10.1145/3233795.3233798 corpus id: 20028979 michael johnston (1998). unification-based multimodal parsing. in proceedings of the 36th annual meeting on association for computational linguistics – volume i. annual meeting of the acl, pages 624–630, montreal, quebec, canada: association for computational linguistics michael johnston, david mcgee, sharon l. oviatt, james a. pittman, and ira smith (1997). unification-based multimodal integration. in proceedings of the eighth conference on european chapter of the association for computational linguistics. european chapter meeting of the acl, pages 281–288, madrid, spain: association for computational linguistics hans kamp and uwe reyle (1993). from discourse to logic. kluwer, dordrecht. andrew a. kehler (2002). coherence, reference, and the theory of grammar. csli publications. http://www.loc.gov/catdir/description/cam0210/99036170.html andrew a. kehler, laura kertz, hannah rohde, jeffrey, l. elman (2007). coherence and coreference revisited. journal of semantics 25: 1–44. http://dx.doi.org/10.1093/jos/ffm018 andrew a. kehler & hannah rohde (2013). a probabilistic reconciliation of coherence-driven and centering-driven theories of pronoun interpretation. theoretical linguistics 39: 1-37. https://doi.org/10.1515/tl-2013-0001 ruth kempson, eleni gregoromichelaki, ronni cann, stergios chatzikyriakidis (2016). language mechanism for interaction. theoretical linguistics, bd. 42, heft 3-4 (october 2016), pages 203275. doi 10.1515/tl-2016-0011. multi-modal anaphora and broadcasting of information 81 adam kendon (2004). gesture. visible action as utterance. cambridge university press, cambridge jeffrey c. king and karen s. lewis. "anaphora", the stanford encyclopedia of philosophy (fall 2018 edition), edward n. zalta, editor. url=. david b. koons, kristinn, r.thórisson, carlton c. sparrell (1993), integrating simultaneous input from speech, gaze & hand gestures. in m. t. maybury, editor. intelligent multimedia interfaces, pages 257-276, aaai press/mit press, cambridge, ma. stefan kopp, paul tepper & justine cassell (2004). towards integrated microplanning of language and iconic gesture for multimodal output. in proceedings of the 6th international conference on multimodal interfaces, pages 97–104. http://dx.doi.org/10.1145/1027933.1027952 emil krahmer and paul piwek (2013). itri-00-13 introduction: varieties of anaphora. researchgate. https://www.researchgate.net/publication/228712984_itri-003_introduction_varieties_of_anaphora jeremy kuhn (2022). a dynamic semantics for multimodal communication. in 13th international conference dhm 2022, lncs 13319, pages 231-243 alex lascarides & matthew stone (2009). a formal semantic analysis of gesture. journal of semantics 26(4): 1–57. http://dx.doi.org/10.1093/jos/ffp004 peter lasersohn (1999). pragmatic halos. language, vol. 75, no. 3 (sept. 1999) pages 522-551 http://links.jstor.org/sici?sici = 0097-507%28199909%2975%3a%3c522%3aph%3e2.0. insa lawler, florian hahn & hannes rieser (2017). gesture meaning needs speech meaning to denote. a case of speech-gesture meaning interaction. in christine howes & hannes rieser, editors, proceedings of fadli 2017, pages 43-47. https://www.researchgate.net/publication/317339368_gesture_meaning_needs_speech_meaning_ to_denote_-_a_case_of_speech-gesture_meaning_interaction sebastian loebner (1987). definites. journal of semantics 4: 279-326. https://www.researchgate.net/publication/31273905_definites andy lücking, (2013). ikonische gesten. grundzüge einer linguistischen theorie. de gruyter, berlin/boston. https://doi.org/10.1515/9783110301489. andy lücking and jonathan ginzburg (2020). towards the score of communication. in proceedings of semdial 2020/watchdial: proceedings of the 24th workshop on the semantics and pragmatics of dialogue. at: brandeis university, waltham (watch city), ma. andy lücking and jonathan ginzburg (2021). leading voices: the polyphonic structure of multimodal interaction. in language and cognition 15(1), pp. 1-28 andy lücking, kirsten bergmann, florian hahn, stefan kopp and hannes rieser (2012). databased analysis of speech and gesture: the bielefeld speech and gesture alignment corpus rieser 82 (saga) and its applications. journal on multimodal user interfaces, pages 92-98 springer, berlin/heidelberg. http://dx.doi.org/10.1007/s12193-012-0106-8 david mcneill (1992). hand and mind. what gestures reveal about thought. chicago and london, the university of chicago press david mcneill (2005). gesture and thought. university of chicago press, chicago gregor mehlmann and elisabeth andré (2012). modelling multimodal integration with event logic charts. in proceedings of the international conference on multimodal interfaces (icmi), pages125-132. santa monica, ca. https://www.researchgate.net/publication/262410953_modeling_multimodal_integration_with_ev ent_logic_charts robin milner (1999). communicating and mobile systems: the π-calculus. cambridge university press, cambridge. http://dl.acm.org/citation.cfm?id=329902 reinhard muskens (1996). combining montague semantics and discourse representation. linguistics and philosophy 19, pages 143-186 costanza navarretta (2011). anaphora and gestures in multimodal communication. in proceedings of the 8th discourse anaphora and anaphora resolution colloquium (daarc 2011), pages 2-10. https://www.researchgate.net/publication/236649685_anaphora_and_gestures_in_multimodal_c ommunication costanza navarretta (2021). speech pauses and pronominal anaphors. frontiers of computer science 3. https://doi.org/10.3389/fcomp.2021.659539 joachim parrow (2001). an introduction to the π-calculus. in j. a. bergstra, a. ponse and s. a. smolka, editors, handbook of process algebra, pages 479–545. elsevier, amsterdam. pamela, perniss & asli özyürek. 2015. "visible cohesion: a comparison of reference tracking in sign, speech, and co-speech gesture". in: topics in cognitive science 7(1), pages 36–60. pfleger, norbert (2007). context-based multimodal interpretation: an integrated approach to multimodal fusion and discourse processing. dissertation, universität des saarlandes. massimo poesio and david traum, 1997. conversational actions and discourse situations, computational intelligence,v. 13, n.3. massimo poesio (1995). a model of conversational processing based on micro conversational events. in proceedings of the 17th annual conference of the cognitive society, pages 678-703. pittsburgh, july 1995. https://www.researchgate.net/publication/2419229_a_model_of_conversation_processing_base d_on_micro_conversational_events massimo poesio, (2016). linguistic and cognitive evidence about anaphora. in m. poesio, r. stuckardt and y. versley, editors. anaphora resolution. algorithms, resources, and applications, pages 23-55, springer-verlag berlin, heidelberg. multi-modal anaphora and broadcasting of information 83 https://www.researchgate.net/publication/305873349_linguistic_and_cognitive_evidence_abou t_anaphora. massimo poesio and hannes rieser (2010). completions, coordination, and alignment in dialogue. dialogue and discourse 1: 1–89. massimo poesio and hannes rieser (2011). an incremental model of anaphora and reference resolution based on resource situations. dialogue and discourse 2(1): 235-277. doi: 10.5087/dad.2011.110 massimo poesio and hannes rieser (2011a). anaphora and direct reference: empirical evidence from pointing. in proceedings of diaholmia, the 13th workshop on the semantics and pragmatics of dialogue, pages 35-43, stockholm, june 2009. massimo poesio, roland stuckardt, yannik versley and renata vieira (2016). early approaches to anaphora resolution: theoretically inspired and heuristic-based. in m. poesio, r. stuckardt, r. and y. versley, editors. anaphora resolution. algorithms, resources, and applications, pages 55-97. springer-verlag berlin, heidelberg. http://dx.doi.org/10.1007/978-3-662-47909-4_3 massimo poesio and renata vieira (1998). a corpus-based investigation of definite description use. computational linguistics 24(2): 183-216. https://www.researchgate.net/publication/2289214_a_corpusbased_investigation_of_definite_d escription_use johannes åman pohjola (2020). psi-calculi revisited : connectivity and compositionality. in : logical methods in computer science. vol. 16, issue 4, 16 :1 – 16 :28 https : //lmcs.episciences.org/ ellen f. prince (1981). toward a taxonomy of given-new information. in p. cole (editor) radical pragmatics, pages 223-255, academic press, new york.https://www.bibsonomy.org/bibtex/258c8a5f9a8006674c7d0c66a74c86e5f/cbrewster w. pouw, & j. a. dixon (2019). quantifying gesture-speech synchrony. in: grimminger, a., editor. proceedings of the 6th gesture and speech in interaction – gespin 6, pages 75-80. paderborn: universitätsbibliothek paderborn. doi:10.17619/unipb/1-815 matthew purver and jonathan ginzburg, patrick healey (2001). on the means for clarification in dialogue. in proceedings of the second sigdial workshop on discourse and dialogue, pages 235255. doi: 10.1007/978-94-010-0019-2_11 hans reichenbach, (1966). elements of symbolic logic. first free press paperback edition. (feb. 2010). hannes rieser (2010). factoring out a gesture typology from the bielefeld speech-gesturealignment corpus«. in: stefan kopp und ipke wachsmuth, editors. proceedings of gw 2009: gesture in embodied communication and human-computer interaction. pages 47–60, berlin, heidelberg: springer, hannes rieser (2011). gestures indicating dialogue structure. in semdial 2011 proceedings of the 15th workshop on the semantics and pragmatics of dialogue, pages 9-18, los angeles, california, 21-23 sept. 2011 rieser 84 hannes rieser (2015). when hands talk to mouth. gesture and speech as autonomous communicating processes. in proceedings of semdial 2015, gothenburg. http://semdial.org/anthology/z15-rieser_semdial_0017.pdf hannes rieser (2017). a process algebra account of speech-gesture interaction. revised and updated version. in: chr. howes and h. rieser, editors. fadli (esslli 2017 workshop) proceedings 2017, pages 67-71. https://www.researchgate.net/publication/317637747_a_process_algebra_account_of_speechgesture_interaction hannes rieser (2024). small-scale sub-turn communication. a study on the interaction of gaze, speech, gesture, back-channels, interactional anaphora, role changes, and repair within a turn, pages 1-25. under review for multi-modal communication. hannes rieser and insa lawler (2020). multi-modal meaning: an empirically founded process algebra approach. semantics and pragmatics 13 (8). early access version. http://dx.doi.org/10.3765/sp.13.8 hannes rieser and massimo poesio (2009). interactive gesture in dialogue. a ptt model. proceedings of sigdial 2009. the 10th annual meeting of the special interest group in discourse and dialogue, pages 87-96, queen mary university of london. https://aclanthology.org/w093912.pdf hannah rohde (2018). pronoun interpretation and production. in c. cummins & n. katsos, n., editors. oxford handbook of experimental semantics and pragmatics, pages 452-473, oxford university press, oxford. doi:10.1093/oxfordhb/9780198791768.013.21 hannah rohde & andrew kehler (2014). grammatical and information-structural influences on pronoun production. language, cognition, and neuroscience 29:8, pages 912-927. https://doi.org/10.1080/01690965.2013. wikipedia entry process calculus. page last edited on 16th july 2022 wikipedia entry p-calculus. page last edited on 16th june 2022 jeanette m. wing (2002). faq on p-calculus. online resource. columbia university. angelika wöllstein, 2018. topologisches satzmodell. in: jörg hagemann, sven staffeldt, (hrsgg.), syntaxtheorien im vergleich, pages 145-166, stauffenberg 2018, tübingen. dialogue & discourse 16(2) (2025) 35–73 doi: 10.5210/dad.2025.202 a few shades of supervision for discourse segmentation: experiments on a french conversational corpus laurent prévot laurent.prevot@univ-amu.fr cnrs & meae, cefc, taipei, taiwan aix marseille univ & cnrs, lpl, aix-en-provence, france philippe muller philippe.muller@irit.fr irit, université de toulouse & aniti, toulouse, france editor: massimo poesio submitted 03/2025; accepted 11/2025; published online 12/2025 abstract elementary discourse units (edus) constitutes the interface between language grammar and language use. one the one hand, they result from compositional semantic processes that combines individual word meanings into proposition-level representations. on the other hand, edus form the building blocks of most text, discourse, and dialogue frameworks. in written genres, where punctuation is available and reliable, segmenting edus is sometimes seen as a nearly solved problem, as least for high-resource languages. however, this is not the case for spontaneous speech transcripts. in this paper, we use a significant (8-hour) french corpus, manually segmented into edus, to evaluate several large language model (llm)-based approaches for this task. we compare various fine-tuning strategies, including those relying on weakly supervised labels, in relation to the amount of ”gold” manual annotations that can be available. we also experiment with in-context learning, where example instances are provided to condition a generative model (few-shots learning) or in a purely generative approach (zero-shot). our findings indicate that classical fine-tuning is still the most effective approach, requiring only a reasonable amount of gold-annotated data to achieve the best performance in our experiments. beyond traditional quantitative evaluation, we conducted a systematic qualitative analysis, identifying directions for further improvement. these include integrating prosodic considerations while handling pauses when they co-occur with disfluencies or complex discourse markers uses. finally, we argue for the significance of this task and the resulting units, compared to acoustic and syntactic proxies, especially for quantitative linguistics focusing on spontaneous speech. keywords: discourse units, conversation, dialogue, llm, weak-supervision 1. introduction elementary discourse units (edus)1 are both the maximal unit of traditional grammar (commonly referred to as the sentence) and the minimal unit of discourse analysis, that extends beyond the sen1. ”discourse units” (du) are often separated into ”elementary discourse units” (edu) and complex discourse units (cdu) that corresponds to sets of edus related by discourse relations (asher and lascarides, 2003). we focus in this paper on edus. ©2025 laurent prévot, philippe muller this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). prévot & muller tence.2 edus serve as the crucial articulation point between language structure and language use. they address the dual question: how do we form ”ideas”, or more precisely, semantic propositions, using words and grammar, and how do we use these ”ideas” to communicate through discourse and dialogue? (1) [et je cherchais ce cette communication]a [parcequ’elle était de folie]a [il en avait fait des photocopies]a [qu’il avait mises en pile dans un coin]a [quel barjot]b [et puis je les ai jamais retrouvé]a [tu les as jamais trouvé ouais]b [et voilà j’aurai bien aimé avoir ce truc de dingue]a [i was looking for this that talk]a [because it was just insane]a [he’d made some photocopies ]a [and stacked them in a corner somewhere]a [what a weirdo]b [and then i never found them again]a [you never found them yeah]b [and that’s it i wish i still had that crazy thing]a3 given their pivotal role, one might expect edus to be a central focus of many linguistic and natural language processing (nlp) studies. however, the opposite is true. conversation, dialogue or communication studies tend to treat them as pre-existing units, while syntax loosely characterizes them as maximal projections. for written genres, this situation can be explained by the reliability of punctuation as a segmentation cue. however, for spontaneous speech—where no such surface cues exist—the neglect of this segmentation problem is puzzling. one possible explanation for this situation might lie in edus’ position at the intersection of two linguistic domains and research communities. analyzing discourse units requires insights from compositional semantics to understand how they are formed, but also from prosody and higher-level pragmatics related to speakers’ intentions. in natural language processing, early efforts (polanyi and scha, 1983; passonneau and litman, 1997) segmented text and speech flow (hirschberg and grosz, 1992), recognizing edus as potentially useful units for various applications (e.g., discourse summarization, cf marcu, 2000). however, discourse processing and parsing were considered as niche tasks due to their complexity and the low performance of early systems. large language models (llms) have changed this landscape. discourse tasks gained attention in nlp conferences, and are supported by specialized events and workshops. even so, ”segmentation” remains an often-overlooked preliminary step, subordinate to downstream tasks like discourse parsing, argumentation structure analysis, or discourse summarization. see however, (stede, 2012), who provides a detailed overview of segmentation, and the disrpt initiative of zeldes et al. (2019, 2021); braud et al. (2023), reflecting a community interest in this area. finally, most existing work focuses either on written genres or on highly interactional dialogues (e.g., task-oriented or chit-chat), which differ significantly from everyday conversation that combines various kind of conversational activities (e.g alternating storytelling with chit-chat sequences) of the kind we study here. 2. we will later discuss in detail the relationship between the notion of ’sentence’ and discourse units. we use ’sentence’ in this introduction as it is the most familiar term corresponding to both the syntactic maximal and discursive minimal unit. 3. in this example, edus are segmented into brackets and the a/b mentioned correspond to the speaker information. 36 discourse segmentation of conversational transcripts 2. related work work on discourse, including topics such as discourse units, discourse markers, and topic changes, has focused on written genres, both in linguistics and nlp. this tendency reflects the documented ”written bias” in these fields (linell, 2004). in contrast, spontaneous conversational speech, while being a subject of interest within certain linguistic subfields, has been somehow overlooked by mainstream linguistics despite its critical importance for understanding language. discourse segmentation has however been addressed for both written and spoken data, including monologues and dialogues, though with differing lenses. broadly, this research spans at least two levels of segmentation: (1) sentence-, utteranceor clause-like units and (2) paragraphor topic-like units. the latter has been extensively studied in natural language processing for both written texts (hearst, 1994) and spoken data (passonneau and litman, 1997). the former has received more interest from semanticists and discourse analysts as the basic unit of analysis, sometimes referred to explicitly as edu (polanyi and scha, 1983; stede, 2012; asher and lascarides, 2003). relational approaches to discourse have used edus as foundational building blocks for constructing discourse structures. we use example (1) to introduce the terminology for this section. we use utterance as the common sense notion of the result of performing one speech act. speech acts follow austin’s traditional definition and are combination of a propositional content and an illocutionary force (e.g. asking, asserting,...) . they have to been generalized to dialogue acts (synonymous to dialogue moves or communicative acts that both can include various communicative functions such as communicative feedback, potentially related to other level of dialogue management (e.g. turn-taking). 2.1 linguistics there has been significant linguistic interest in discourse units in speech for a long time, beginning with work by conversation analysts (sacks et al., 1974; schegloff and sacks, 1973) and later by interactional linguists (ford and thompson, 1996; selting, 2000). discourse analysts across various frameworks have also contributed, favoring functional over formal approaches and emphasizing corpus-based empirical methods (sinclair and coulthard, 1992; brazil, 1995; roulet et al., 2001). scholars specializing in spoken language syntax have further contributed to the question (pietrandrea and kahane, 2019; haselow, 2017). degand and simon (2005, 2009) already proposed to synthesize these proposals to define basic discourse units. an overview of functional approaches to discourse in romance languages is provided in (pons borderı́a, 2014). finally, researchers exploring the formal semantics and pragmatics of dialogue, such as (asher and lascarides, 2003) and (ginzburg, 2012), have also examined discourse units or dialogue moves in conversation. 2.1.1 interactional units interactional linguistics (couper-kuhlen and selting, 2001), a development of conversation analysis focused on language resources, introduces interactional units (ford and thompson, 1996) as a crucial step for studying interactions. we consider these interactional units to be closely related to our discourse units. in (ford and thompson, 1996), the interactional unit is defined through three key criteria: syntactic completion, intonation completion, and pragmatic completion as illustrated 37 prévot & muller in (2). while prosody is not the focus here, syntactic completion relies on the syntactic structure without being restricted to the traditional notion of the ”sentence”. instead, it helps to make the concept of ‘projectable units of natural language’ concrete (sacks, 1992). this relies on the ’clause’ notion, where a syntactically complete utterance can, ”in its discourse context, be interpreted as a complete clause, that is, with an overt or directly recoverable predicate”. interactional units also depend on identifying a prosodic final intonation contour and require the unit to be interpretable as a ”complete conversational action”. (2) from (ford and thompson, 1996) it was like the other day ! uh > # vera > was talking ! on the phone ! to her mom!>]4 2.1.2 basic discourse units identifying basic discourse units is central to a series of studies by liesbeth degand and colleagues (degand and simon, 2005, 2009; crible and degand, 2019; hu and degand, 2023). more precisely, (hu and degand, 2023) follows the general approach of (degand and simon, 2009) in defining conversational discourse units (cdus).5 they emphasize the importance of treating syntax, prosody, and pragmatics as independent dimensions at the same level of granularity. syntactic units are identified using dependency parsing (nivre, 2010). specifically, a syntactic unit consists of the verb along with its arguments, forming a ‘dependency clause’ that demonstrates maximal syntactic completeness. this approach excludes some linguistic elements (e.g., fragments without verbs, discourse markers), which then need to be promoted to syntactic units, resulting in an exhaustive segmentation of the discourse into syntactic units.6 the pragmatic dimension is inspired by ford and thompson (1996) but operationalized in a way that does not depend on the interlocutor’s response. a pragmatic unit is understood as a step in the implementation of a speaker’s plan. however, cdus should not be limited to the smallest syntactic, prosodic, or pragmatic units. rather, they are the result of the integration of these three levels of analysis. consequently, a cdu is a multidimensional unit, with several subtypes emerging depending on the specific mapping between the three dimensions. 2.1.3 macrosyntax and dualistic approaches building on the macrosyntax movement, which focuses on spoken french syntax (blanchebenveniste et al., 1990; deulofeu, 2003; sabio, 2006; benzitoun and sabio, 2010), pietrandrea and kahane (2019) proposed an operationalization for analyzing spontaneous speech utterances in terms of syntactic, prosodic, and discourse organization. the model suggests that discourse is composed of illocutionary units (ius), based on austin’s concept of illocutionary force (austin, 1975). each unit consists of a nucleus that carries the illocutionary force, along with optional elements, or satellites, which depend on their nucleus. these satellites are studied without directly referring to traditional syntactic categories. additionally, the macrosyntactic framework provides specific notations to transcribe parentheticals, connectives, and other phenomena frequent in spontaneous speech. analyzing a sequence within the macrosyntactic framework requires expertise, as it de4. in this example, ”!” stands for syntactic completion point ; ”>” prosodic completion point and ’]’ correspond to the pragmatic unit. 5. conversational discourse units are conversational counterparts of basic discourse units. 6. by ’exhaustive’, we mean that the segmentation is continuous and no tokens are left behind. 38 discourse segmentation of conversational transcripts mands a fine-grained understanding of the utterance. the framework uses highly specialized terms and concepts. (3) from (pietrandrea and kahane, 2019) [(je suis arrivée euh au kenya) (je voulais travailler d’abord pour le gouvernement)]iu [(i arrived erm in kenya) (i wanted to work first for the government)]]iu as can be seen in example (3), in this framework illocutionary units (here within brackets) can go beyond standard microsyntactic (called here government units here in parentheses). more precisely, the annotation process involves several distinct tasks: (i) identifying illocutionary units; (ii) annotating the internal structure of an iu; and (iii) determining the relations between ius. ius can be defined as all the microsyntactic units (which relate to traditional syntactic relations, e.g., dependencies) that contribute to realizing a single assertion, injunction, interrogation, or exclamation. this is the case when they could be embedded under the scope of a saying verb that makes the illocutionary value of the entire sequence explicit. as mentioned earlier, the nucleus plays a crucial role, as it bears the illocutionary force of the entire sequence. the nucleus could be uttered alone and still retain its illocutionary function. several other attempts to define systems better suited to the peculiarities of spontaneous speech exist. brazil (1995) proposes a functional grammar of speech, aiming not to describe ’sentence forms’ but to describe the successful realization of communicative purposes. he uses the terms ’increment’ and distinguishes between ’telling increments’ and ’asking increments’. when defining the minimum requirements for a telling increment, he mentions the inclusion of nominal and verbal elements, with predication playing again a central role. within his dualistic approach to grammar, haselow (2017) articulates a microand macrogrammar, identifying several ”fields” (e.g., initial, medial, and final) to refine the modeling of spontaneous speech sequences. it should be noted that peripheral utterances have been scrutinized as key locations for analyzing spoken syntax, discourse, or conversational organization (lewis, 2021; herment et al., 2022). 2.1.4 formal models of discourse and dialogue segmented discourse representation theory (sdrt) (asher and lascarides, 2003) is a theory that models the semantic-pragmatic interface by constructing a coherent discourse structure. in sdrt, discourse is divided into elementary and complex discourse units (edus and cdus). despite its compositional semantic foundation, the definition of edus is also top-down. any object in the text that serves as an argument of a discourse relation can be considered a discourse unit. for the standard case, sdrt relies on the semantic notion of a proposition, which is closely tied to the syntactic notion of an independent clause. however, sdrt also accommodates a wide range of forms, including propositional pronouns and non-sentential units, which can be promoted to this level. stent (2000) made also an early proposal to extend rst (mann and thompson, 1987) to dialogue. they proposed a multi-level definition indicating to first segment on ’cue words’ separating ’syntactic phrases’ (involving discourse and syntax factors); then treat ’syntactically complete clause’ (syntax) and then consider ’stretch of continuous speech ended by a pause, a 39 prévot & muller prosodic boundary or a change of speaker’ (prosody, interaction). while framed within a completely different formal apparatus, ginzburg (2012)’s model of dialogue also allows nearly any ’form’ to function as a dialogue move. to account for the broad range of forms—such as elliptical clauses, answers to questions, back-channel responses, as well as laughter or gestures, ginzburg relies on a rich modeling of conversational context. 2.2 natural language and speech processing when considering automatic discourse segmentation for spontaneous speech transcripts, there are two main traditions to consider: textual discourse segmentation and segmentation of speech into sentence-like units. 2.2.1 textual discourse segmentation as early as (polanyi and scha, 1983), a discourse model was proposed, even though it provided little details about its basic building blocks. the primary unit identified was the ’clause’, a semantic object that seemed to require no further specification. much later, passonneau and litman (1997) initiated a new standard of discourse processing studies by conducting an annotation campaign, including inter-annotator agreement evaluation, and developing a system for discourse segmentation of spoken data. this system combined prosodic features (primarily pause duration) and discourse connective information to approximate the human reference. however, their focus was on paragraph or topic segmentation rather than on segmenting elementary discourse units. overall, the most common task in textual discourse segmentation has been the segmentation of long documents based on their thematic and rhetorical organization (hearst, 1994). this was considered an important precursor for other tasks, such as summarization or argumentation structure extraction (marcu, 2000). identifying elementary discourse units is sometimes seen as a necessary step for comprehensive discourse parsing (marcu, 2000). stede (2012) elaborates on this, emphasizing that elementary discourse segmentation is an important step in the discourse processing pipeline. however, in the context of written documents, this task is limited to the identification of sentence-internal discourse units that reside within a sentence. stede offers a semantic perspective on elementary discourse units (stede, 2012, p.89), defining them as: ”a span of text, usually a clause, but in general ranging from minimally a (nominalization) np to maximally a sentence. it denotes a single event or type of event, serving as a complete, distinct unit of information that the subsequent discourse may connect to.” this definition combines both internal criteria (what constitutes a discourse unit) and external criteria (valid attachment points for subsequent discourse). following the general trend in nlp, symbolic rule-based approaches (polanyi et al., 2004) have been largely replaced by statistical methods (soricut and marcu, 2003). as in many other subfields, large language models (llms) have proven highly effective, achieving unprecedented performance on discourse segmentation tasks. however, discourse segmentation within deep learning approaches has been applied to only a limited number of languages, until the recent launch of the disrpt campaigns (zeldes et al., 2019, 2021; braud et al., 2023). the research conducted within the framework of these campaigns has provided the community with powerful tools and frameworks to perform discourse unit (du) segmentation and its evaluation using contemporary 40 discourse segmentation of conversational transcripts methods. however, even for written genres, discourse segmentation performance deteriorates in languages other than english and when gold sentences are not provided, due to the imperfections of sentence splitters (braud et al., 2023). nevertheless, a recent trend involves using sequential models over contextual embeddings for discourse segmentation (pruksachatkun et al., 2020; muller et al., 2019). metheniti et al. (2023) provide an improvement over (muller et al., 2019), achieving new state-of-the-art results for discourse segmentation in various languages.7 several aspects of our study build upon this framework. in addition to fine-tuning approaches, zero-shot and few-shot learning techniques are gaining significant attention across a wide range of tasks. while these methods present challenges related to explainability, control, and prediction-time costs, they constitute a new paradigm for our field. these techniques have been explored for discourse relation identification (metheniti et al., 2024) and discourse segmentation (nayak, 2024). the latter, which directly employed chatgpt8 for english discourse segmentation, found that results were unsatisfactory due to model-generated so-called hallucinations, i.e. generation of spurious elements. 2.2.2 dialogue act identification the development of dialogue systems requires the identification and classification of user responses as communicative acts. early work has primarily focused on dialogue act classification (stolcke et al., 2000), often considering segmentation to be an implicit task, handled by segmenting sentencelike units (see the next section), or implicitly by whatever pre-existing given boundaries (e.g., short speaker turns typically corresponding to a single communicative act). gupta and bangalore (2003) propose a pipeline that starts with the user’s utterance, extracting and removing filled pauses, discourse markers, and explicit edits, followed by the identification of coordinating devices to produce what is termed the ”clausified utterance”. although this work primarily targeted human-computer interaction, which presents distinct challenges compared to segmenting human-human conversational speech, many steps in their pipeline align with the difficulties faced in systematically segmenting spontaneous speech transcripts. the majority of the work in this area targets the classification of dialog acts and accordingly attempted to treat segmentation and classification as joint task (quarteroni et al., 2011). the approaches and the corresponding papers are not very detailed with regard to the segmentation specifically. however, (ang et al., 2005) specifically discusses segmentation evaluation metrics. zhao and kawahara (2018) propose a joint approach and use a bilstm for encoding and a sequence classifier for decoding the segmentation part of the model, evaluated on the switchboard data set (swda, jurafsky et al., 1997). some studies attempted to work directly from the speech signal such as (dang et al., 2020) that integrates the dialogue act segmentation with the asr through the use of a ’dialogue act boundary’ token. finally, dialogue act tagging, due to its specific application context, gives rise to varying interpretations of the base units. for instance, some authors argue that dialogue acts are multi-functional, and thus multiple segmentations may be considered depending on the aspect of the dialogue being 7. code available at https://github.com/phimit/jiant/ 8. https://openai.com/research/gpt-4 41 prévot & muller analyzed at the time of segmentation (bunt, 2011). see example (4) for a case of discontinuous functionality with the specific time-management function interrupting the main communicative function. (4) from (bunt, 2011) [the first train to the airport on sunday is at # [let me see] # 5.32] 2.2.3 sentence segmentation in speech the emergence of large conversational corpora (godfrey et al., 1992) introduced several challenges for the speech processing community, with one of the most significant being the unruly nature of conversational speech flow. to perform any kind of processing, it was necessary to segment this flow into convenient units. the decision was made to segment it into sentence-like units, with various perspectives regarding their relation to written language. the work done in this area, with notable success, can be summarized as addressing how to use language models (initially simple ngrams) for predicting sentence-level breaks in the token flow, how to incorporate acoustic-prosodic information, and how to combine these two sources of information to improve segmentation results. in terms of language modeling, early models based on hidden markov models (hmms) (stolcke and shriberg, 1996) already produced promising results. one difficulty they faced was the strong impact of disfluencies in the natural speech flow, leading to the development of specific models designed to identify and eliminate them (stolcke et al., 1998). over time, these models were improved by adopting conditional random fields (crfs) (liu et al., 2005). these models aimed for efficiency but imposed a strong bias from the written realm on any subsequent data analysis. several studies have shown the benefit of considering acoustic modality characteristics in speech segmentation. building on the pioneering studies of pierrehumbert and hirschberg (1990) and hirschberg and grosz (1992) about the relationship between prosody and discourse structure, shriberg et al. (2000) systematized the use of acoustic-prosodic cues for segmenting speech into sentence-level and topic-level units. furthermore, they discovered that the importance of different cues varied depending on the corpus considered. for instance, pause and pitch features were highly informative for segmenting news speech, while pause duration and language modeling dominated in natural conversation. see also (portes and bertrand, 2011) for an investigation of the issue in french. approaches based on low-level acoustic features, such as pauses, have since been the primary method for segmenting speech into utterances. however, the performance of these simple systems plateaued over time, and with the rise of more advanced machine learning techniques, methods trained on much larger written corpora (allowing for language modeling at a new level) have been explored as effective alternatives. one of the most direct approaches to segmenting speech transcripts into communicative units is to punctuate them as if they were written. the general strategy involves training a punctuation model on a large corpus of written data and then applying this model to speech transcripts lacking punctuation, as done by favre et al. (2008). this method provides an efficient approximation of segmentation into discourse units. it works particularly well for fluent, canonical, and monological sequences, where the method is nearly perfect. however, such models struggle with disfluencies and highly spontaneous sequences absent from the training corpus. more42 discourse segmentation of conversational transcripts over, from a linguistic perspective, it is questionable to impose a notation derived from the written realm onto these spontaneous speech transcripts. despite these limitations, punctuation-based methods remain efficient, with newer models employing multilingual deep language models, e.g., xlm-roberta, which has proven to be the best-performing model (guhr et al., 2021).9 while these models provide impressive results for a relatively simple approach, they note that ”punctuation patterns are domain-specific, and robust punctuation prediction requires training on diverse datasets.” 2.3 take-away from this literature, we retain the following main take-aways. overall, identification of discourse units in spontaneous speech is a complex problem in which syntactic, prosodic and pragmatic / interactional dimensions are deeply intertwined. many proposals (ford and thompson, 1996; degand and simon, 2009; petukhova et al., 2011; hu and degand, 2023) are going as far as treating this issue by considering units at all those different levels and discourse units as products of the relationship between these different sub-level units. a second idea well represented in these works is the need to adapt the grammar itself for the internal structure of the discourse units in spontaneous speech. the main systematic issue is the presence of disfluencies. overall, spoken specific structure of discourse units are built around a more linear organization around a central element and a range of optional initial and final elements. these organizational patterns may offer a way to bridge the ’local’ and ’global’ levels of discourse, with the peripheries naturally serving as predisposed locations for anchoring elements at the global level. the ultimate criterion shared by many seems to be the discourse function / communicative action. this corresponds to the idea that discourse segmentation and discourse relation analysis should not be performed in a strictly sequential way (hoek et al., 2018). roulet et al. (2001) go as far as considering that the minimal discourse unit definition must be a top-down one as there is no criteria to define a syntactic maximal unit. overall however, even though its individuation is recognized to be difficult, the different accounts rely on the core notions of predication, semantic proposition and speech act. in terms of automatic segmentation, it seems that the convergence of nlp (nowadays represented by llms, pretrained on huge amounts of textual data, written or conversation) and speech processing traditions, that also used language modeling coupled with acoustic extraction is not completely achieved. there is currently a convergence in terms of methods but the segmentation of elementary discourse units in spontaneous speech has not yet greatly benefited from this convergence. one explanation is that overall nlp and speech processing are moving toward more end-to-end approaches that avoid processing pipelines in which discourse segmentation would play a role. 9. see also https://github.com/benob/recasepunc for alternative implementations. 43 prévot & muller 3. operationalization: definition and examples based on this review, we adopt a discourse approach that we aim to keep simple enough for seminaive coders after receiving some training. while acknowledging the multidimensional nature of the phenomena discussed, our approach focuses on a single level of discourse/semantic/pragmatic units. even though multiple overlapping layers are analytically justified, we believe that conversation participants and semi-naive coders should identify discourse units as conversational actions without explicitly attending to these complex overlapping layers. our operationalization of elementary discourse units employs both top-down and bottom-up characterizations. this mirrors the interface position of edus as both the maximal unit within the realm of grammar and language and the minimal unit of discourse structure, including in texts, stories, narratives, argumentation, or task-oriented dialogues. from the top-down perspective, any communicative object with, or promoted to have, propositional-level value (including non-assertive values such as questions or requests) that can be related to another discourse object through a discourse relation is considered a discourse unit.10 our approach will group semantic, pragmatic, and discourse aspects under the common-sense ideas of ”what it means” or ”what it does” in context. from the bottom-up perspective, we also rely on syntax, though without limiting ourselves to independent clauses. this flexibility is enabled by the top-down criteria and informed by the variety of phenomena and structures identified in the literature on spoken syntax and spontaneous speech. we adopt a generous approach to what can be included in a discourse unit, drawing on literature on ”satellites,” ”fields,” and ”peripheries” (cf section 2.1.3). this approach consists in allowing for these elements to be included in, potentially large, elementary discourse units as long as they do not clearly convey a full propositional meaning and are not discourse-articulated (e.g. with a discourse marker) to the rest. 3.1 definition the discourse units we aim to annotate are closely related to the ones of (polanyi et al., 2004), where they communicate information about no more than one event, event-type, or state of affairs. these units correspond semantically to independent syntactic clauses. however, the interactional and dialogical nature of discourse requires us to expand this view to include any object that, in context, plays a similar semantic role. our semantic perspective on discourse units leads us to combine criteria to reach the definition proposed further in the definition below. • semantic criteria: following vendler (1957)’s classification of eventualities, we consider a discourse unit as a segment representing a specific event, state, or proposition ; • discourse criteria: we are using discourse markers, whether these markers are connectives, bundles of them or adverbials in a discourse position. 10. we do not elaborate on the notion of discourse relations here as it would take us too far from the segmentation focus. it is enough to mention that discourse units can be related to or attached to other units from the discourse context. 44 discourse segmentation of conversational transcripts • pragmatic criteria: we recognize specific speech acts, such as asking / answering a question or providing conversational feedback, as discourse units. definition in a speech transcript, an elementary discourse unit is a span of contiguous tokens whose combined meaning represents a semantic proposition (describing an event, fact, or state of affairs) or an individual speech act, which can be identified as a basic communicative function, such as asking a question, providing an answer, or offering conversational feedback. all preparatory disfluent material should be included within the elementary discourse unit. a unit is only labeled as an abandoned discourse unit if the introduced material is discarded before the beginning of a different discourse unit. to illustrate, a discourse unit is a segment that describes an eventuality, as in example (5), or carries a clear communicative function, as seen in speech acts such as in (6). in the former case, identifying a predicate (e.g., through the main verb) is a strong clue, as discussed in the related work section. the discourse unit generally includes the arguments of the verb unless strong discourse markers signal a distinction between the arguments and the rest of the clause, as in example (7). (5) eventualities a. [on y va avec des copains]du [we are going there with friends]du [on avait pris le ferry en normandie]du [we took the ferry in normandy]du [puisque j’avais un frère qui était en normandie]du [since i had a brother that was in normandy]du [on traverse]du [we cross]du [on avait passé une nuit épouvantable sur le ferry]du [we spent a terrible night on the ferry]du b. [j’ai eu plusieurs conflits avec des animateurs pas assez sérieux]du [i had several conflicts with group leaders that were not serious enough]du c. [et y en a un qui s’était pris un banc de pierre]du [and there was one who hit a stone bench]du (6) speech act / clear communicative function a. a: a: [tu vois où c’est?]du [you know where it is?]du b: b: [oui]du [yes]du b. a: a: [je ne voulais pas les déranger]du [i did not want to disturb them]du b: b: [oui bien sûr]du [yes of course]du (7) discourse markers inducing segmentation a. [on a appelé euh des les parents d’amis]du [we called um some friend’s parents]du [mais pas d’amis de notre âge d’amis de mes parents]du [but not friends of our age friends or my parents]du b. [donc on était à montréal en fait]du [so we were in montreal in fact]du [et après le congrès on est parti en gaspésie]du [and after the conference we left to gaspésie]du discourse adverbials and conjunctions serve as strong cues for discourse boundaries and are included in the discourse unit they introduce, as illustrated in example (7). final spoken particles like 45 prévot & muller en fait, quoi (in fact, what) also provide useful cues for discourse segmentation and are included in the preceding unit. strictly speaking, attitudinal markers such as tu sais (you know) or je crois (i think) can be seen as discourse units. however, we instructed annotators to include them within the units they qualify. first of all, the discourse attachment of these elements is rather trivial : they qualify the discourse unit within which they occur. this decision also stems from the fact that these markers are often produced as parentheticals, and treating them as standalone discourse units would result in many nested discourse units, which we aim to avoid. furthermore, in terms of methodology, these attitudinal markers are easier to treat specifically at a later stage. 3.2 data the experiments presented in this paper use an existing discourse-segmented corpus, which was segmented according to the approach outlined above. the raw data comes from the corpus of interactional data (cid) (blache et al., 2009, 2017), which consists of 8 dyadic conversations, each lasting approximately 1 hour. the cid contains highly spontaneous data, featuring colloquial sequences, as in example (8), and strong disfluencies, as in example (9), making discourse segmentation much more challenging than for written genres, even for human annotators.11 the full dataset consists of approximately 125,000 tokens and 15,463 discourse units, with 12.4% of the tokens marking edu boundaries. annotations were performed using praat (boersma, 2002) to enable precise word-level alignment of the audio signal when making segmentation decisions. more specifically, during edu segmentation, annotators had access to the time-aligned transcripts of both participants in the conversation. the discourse annotations presented and used in (prévot et al., 2021) were carried out during the otim project (blache et al., 2009, 2017) and are available on the ortolang platform, along with the guidelines used for annotation.12 (8) a:[comme ça # ah ouais non c’était]du [like that # oh yeah no it was]du b:[ah ouais profitez profitez de vos soirées]du [oh yeah enjoy enjoy your evenings]du a:[ouais c’est pour ça]du [yeah like that]du (9) [ou des euh non pas des fpas des frustrations]du [des # espèces de euh # mhm # ouais des des vues différentes sur le boulot quoi]du [or some uh no not some fnot some frustrations]du [some kind of uh # mh # yeah some some different views about work what]du 11. we use the symbol ’#’ to mark pauses in the examples. 12. https://www.ortolang.fr/market/item/ortolang-000918 46 discourse segmentation of conversational transcripts 3.3 manual segmentation evaluation elementary discourse units were obtained through at least two manual annotations, conducted by 4 naive coders and 2 experts. the mean cohen’s κ score across speakers for naive coders is 0.85 (min: 0.83, max: 0.87). for evaluating multiple coders’ agreement we used a standard multi-κ measure discussed in (artstein and poesio, 2008). more specifically, we additionally report, in the appendix, the multi-κ score when available (table 9), the agreement between individual naive coders (table 8), and the agreement between naive coders and experts (table 10).13 while we acknowledge the advantages of using adapted metrics such as windowdiff (wd) (pevzner and hearst, 2002), boundary edit distance (fournier and inkpen, 2012), or γ (mathet et al., 2015) — and have experimented with them in previous work (peshkov and prévot, 2014; prévot et al., 2016) — the present task is a simple binary decision over a skewed, but not highly skewed, distribution. we therefore do not evaluate the task as an interval labeling problem but rather as a token labeling task, following the current trend adopted, for instance, in the disrpt shared tasks (zeldes et al., 2019; braud et al., 2023). 4. automatic discourse segmentation experiments in this section, we present a series of experiments aimed at automatically segmenting the cid corpus into elementary discourse units (edus) as defined earlier. the primary focus of our experiments is on fine-tuning approaches, using large language models (llm), specifically xlm-robertalarge (conneau et al., 2019). this model is on the lower end of current models with respect to size (and thus expressivity), but was chosen as it presents a good compromise between performance and ease of use without large computing resources. it is also one of the rare multilingual models in this capacity range. we compare two distinct fine-tuning strategies: a straightforward fine-tuning approach and one based on weakly supervised annotation, leveraging less annotated data. our models are sequence-to-sequence models, fine-tuned following the methodology proposed by (metheniti et al., 2023) within the jiant framework (pruksachatkun et al., 2020).14 additionally, we explore zero-shot and k-shot in-context learning techniques with larger generative models. the task is defined within the disrpt framework. the data is encoded in the conll format, with a simple binary label indicating whether a token marks the beginning of a discourse unit (see table 1). evaluation is performed using f-score, precision, and recall, calculated specifically for the discourse boundaries (with true negatives excluded from the score). for written data, on french, scores are around 90 of f-score (braud et al., 2023). while there have been few attempts at this task for conversational speech, gravellier et al. (2021) achieved an f-score of 73.7.15 4.1 weak supervision modern nlp techniques require large amounts of annotated data. the general idea behind weak supervision is to avoid the costly process of manually annotating large datasets by using 13. the agreement between experts is not meaningful, as expert double-coding was performed only on a small subset of the data, which was also used for training the naive coders. as a consequence, expert agreement is nearly perfect, as it underwent several rounds of adjudication. however, we obtained an expert agreement of about 0.9 for the same task on a different corpus (prévot et al., 2025). 14. see https://github.com/phimit/jiant-discut 15. at the time, the base llm used was bert. using xlm-roberta would likely result in a higher score. 47 prévot & muller id token connll discourse boundary 1 ouais beginseg=yes 2 # 3 on beginseg=yes 4 dirait 5 des 6 enfants 7 # 8 hein 9 # 10 mais beginseg=yes 11 les 12 enfants table 1: illustration of the data format for the discourse segmentation task. semi-automated methods to generate annotated data. while the accuracy of these annotations is known to be lower than human-generated annotations, the large volume of data is expected to compensate for the noise. more precisely, we adopt here the ”data programming” approach of ratner et al. (2017), using the skweak framework developed by lison et al. (2021).16 this framework provides an api for defining multiple overlapping heuristic classification rules and an aggregation model to combine them. the set of rules, called labeling functions (lfs) in that framework, can be developed and tested with only a small amount of annotated development data. the system builds a profile for each lf, and a model is trained by combining all the lfs, which can be weighted by their estimated reliability. the reliability is estimated without supervision, relying on agreement between rules, and the hypothesis that they are all at least slightly better than pure chance. this ”label model” is then used to label a training set, and finally, a supervised model is trained on the dataset automatically annotated by the label model. a similar approach, based on the snorkel implementation of (ratner et al., 2017), has been used in discourse analysis to enrich a discourse parser (badene et al., 2019), and is also the foundation for the work in (gravellier et al., 2021). in our experiments, we combine three types of rules: (i) base rules, which assign a label (positive or negative) to almost all tokens (acting as a default rule), which can come for instance from a model trained on written text; (ii) recall-oriented rules, which aim to catch more specific discourse boundaries, and (iii) precision-oriented rules, which specify where discourse boundaries should not occur (e.g., by narrowing the application scope of the default rules). for instance, in table 2, the rule lf du non end tok labels tokens that are unlikely to mark the end of a discourse unit, while the rule lf long pause labels tokens following a pause of at least 800 milliseconds.17 we created four sets of rules for weak supervision, based on two key considerations: (1) variations of the base rule, either relying on the discut model (metheniti et al., 2023) or on a sim16. https://github.com/norskregnesentral/skweak 17. a few examples of labeling functions are provided in appendix c, while the complete set is available in the repository at https://github.com/phimit/weakling/tree/main/weaksupervision. 48 discourse segmentation of conversational transcripts # name label conflicts precision recall f-score 8 lf du non end tok ndu 0.017 0.996 0.359 0.528 9 lf du non end pos ndu 0.027 0.990 0.427 0.596 14 lf du very long pause du 0.061 0.948 0.246 0.390 17 lf du long pause du 0.150 0.868 0.378 0.526 19 lf du dm ini forward du 0.347 0.683 0.113 0.194 20 lf du discut du 0.410 0.607 0.876 0.717 table 2: some discourse segmentation labelling function profiles. conflicts : % cases conflicting with at least another labelling function. du : discourse unit start = true ; ndu : discourse unit start = false. as an example rule #14 says that a very long pause implies the start of a new segment, and rule #9 says that at a token from a given list of part-of-speech tag, there is no boundary at that token. precision, recall and f-score are estimated from development data to help rule design, but are not taken into account by the weak supervision process. ple punctuation predictor18; (2) whether to use syntactic chunk information (mostly as precisionoriented rules to prevent splitting chunks into distinct discourse units) or not. these variations led to the following weak supervision models being tested: • mono-discut: simple discourse segmentation fine-tuning (ft) with discut as the base rule for weak supervision. • multi-discut: joint segmentation and chunking with discut as the base rule for discourse noisy labels (and chunks also having noisy labels). • mono-punct: simple discourse segmentation ft with a punctuation predictor as the base rule for weak supervision. • multi-punct: joint segmentation and chunking with a punctuation predictor as the base rule for discourse noisy labels (and chunks also having noisy labels). adding chunk information did not seem to improve discourse segmentation performance, however. as a result, we focused on sets of rules that did not incorporate chunk information when predicting noisy labels. additionally, we conducted experiments both with and without discut, as this represents two realistic scenarios for different languages, depending on whether a high-quality segmentation model exists for the language in other genres. we compared our models with the set of models that are either baselines, toplines or interesting competitors. all the models are fine-tuned models based on xlm-roberta-large: • gold : model fine-tuned on the training set with gold labels for the whole training set (topline); 18. we also experimented with pauses as the base rule, but this resulted in acceptable, yet significantly lower, performance compared to the other approaches. 49 prévot & muller model precision recall f-score nb supervised tokens nb supervised du fra.gold 0.865 0.830 0.847 115850 15630 fra.gold10 0.831 0.839 0.835 11585 1563 fra.mono-discut 0.698 0.850 0.767 11581 1498 fra.multi-discut 0.671 0.876 0.760 11581 1498 fra.mono-punct 0.776 0.684 0.727 11581 1498 fra.multi-punct 0.797 0.662 0.723 11581 1498 fra.discut-w 0.604 0.872 0.714 0 0 fra.pause 0.896 0.423 0.575 0 0 fra.punct 0.776 0.443 0.564 0 0 table 3: results of experiments on the french conversation corpus with the main approaches and a few toplines and baselines. gold is a fully supervised model as a topline comparison, gold10 the same with only 10% of the annotated data. discut w is the discut model trained on written text and taken as it is for conversation segmentation. pause and punct are baselines relying only on pauses or automatically predicted punctuation respectively. in bold are the best weak supervision scores. the number of supervised tokens/du needed for each approach means annotated instances for training a fully supervised model, or in the case of weak supervision, the number of annotated instances used for development/validation of the heuristic rules. this was predetermined to be roughly the same as 10% of the gold annotations, but ideally a weak supervision approach might need less development instances for similar results. • gold 10 : same model fine-tuned on only 10% of the gold-annotated training set (which corresponds roughly to the amount of gold data used to produced the weakly supervised development set); • discut w : discut model directly used to predict on our data set (no fine-tuning) (baseline). pauses longer than 500 ms were sent to the model as commas; • pause : rule-based model segmenting the data set on pause over 500ms (baseline); • punct : rule-based model using strong punctuation (period, question mark) as predicted by the punctuation prediction model (baseline). the main result of this set of experiments, presented in table 3, is that, within this framework and for this language, fine-tuning on a corresponding amount of gold labels is more efficient than adopting a weakly supervised approach where gold annotations are used to develop and evaluate the heuristics rule. this amount has to be predetermined before developing the rules, and it is possible that less annotations are needed, but it is hard to estimate, and the amount of gold annotations used in our experiment is actually not that big (about 11k tokens and 1500 discourse units). this contrasts with previous results from (prevot et al., 2023), where simple fine-tuning required up to 7000 discourse units (dus) to outperform weak supervision. the improvement observed here can be attributed to the use of roberta (compared to bert), the discut framework (compared to tony), or both combined. while it is possible that better labeling function engineering could yield 50 discourse segmentation of conversational transcripts improved results, we have experimented extensively with the available cues and accumulated experience on this issue (prevot et al., 2023). while not completely ruling out this possibility, in terms of efficiency, dedicating time to developing the labeling functions does not seem more efficient than performing annotation on a sample. moreover, we observe that using discut (a discourse segmenter trained on written data) yields better results than a ”simple” punctuation predictor. finally, we do not observe a clear difference between the results obtained in the single-task versus the multi-task approach. one open question is that heuristic rules might be more robust to change in the distribution of data they are applied to, but this goes beyond the present study. 4.2 how much human-labeled data do we actually need? the previous experiment, as well as the work conducted in (prévot and wang, 2024) for another language, led us to explore more systematically and radically how much data is actually needed to create an effective segmentation model. we used again xlm-roberta with a learning rate of 10−5, a batch size of 1, gradient accumulation of 4, and a maximum of 30 epochs with a patience of 10, based on performance on the development set. we performed 8-fold cross-validation based on speaker ids, creating different test, development, and training splits to maximize the distance between training and testing data given our corpus. nb train dus nb train tokens precision recall f-score ≈ 15 100 0.73± 0.29 0.11± 0.06 0.19± 0.09 ≈ 30 200 0.81± 0.11 0.21± 0.08 0.33± 0.10 ≈ 70 500 0.77± 0.03 0.63± 0.04 0.69± 0.03 ≈ 140 1000 0.78± 0.04 0.73± 0.02 0.75± 0.03 ≈ 700 5000 0.80± 0.02 0.80± 0.05 0.80± 0.02 ≈ 1400 10000 0.82± 0.02 0.81± 0.05 0.82± 0.02 ≈ 4200 30000 0.82± 0.02 0.84± 0.04 0.83± 0.02 table 4: global simple fine-tuning results with different amounts of training supervision, averaged with leave-one-speaker-out cross validation (resulting in 8 folds). figure 4 summarizes the results of our experiments. fine-tuning roberta for elementary discourse unit segmentation proves to be highly efficient, even with a small amount of training data. the general pattern observed is that the model starts with a very low recall score but improves quickly as the amount of training data increases. answering the question in the title of this section depends of the precise goal of the discourse segmentation. if one simply wants more accuracy than the usual proxies, then we can seen that even a modest training build within a few days can be enough. however, if one is looking for the best discourse segmenter possible with the goal to approach human level, more data is desirable. the f-score tends to plateau but still goes up when training size grows exponentially.19 19. see however (prévot et al., 2025) for related experiments on a different corpus but with more training data and a more pronounced ”plateau” effect. 51 prévot & muller 4.3 comparison with few-shot generative approaches given the dynamics of the fine-tuning experiments presented above, we decided to explore the potential of a zeroor few-shot approach using a generative model for this task. we initially attempted a pure zero-shot approach using the prompt in example (10), and then adopted a few-shot approach with the prompts provided in appendix d (see example (11) for a short illustration). the few-shot prompt consists of a set of examples derived from our corpus, taken from the training portion. for each example we provide the input data without the segmentation and the expected output using the ”|” pipe symbol as a boundary label. note that we pre-tested both an english prompt and the same prompt in french, and the former seemed to work better. (10) "segment the following dialog in elementary discourse units, where the character # indicates a short pause: only print the original text, indicating segment boundaries with a | character, and do not add anything, do not remove anything. do not present the result with introductory text." we used a version of the freely available llama3-8b base model (dubey et al., 2024), quantized to 4 bits, served via the ollama wrapper and api.20 the model was set with a temperature of 0.3 and topk of 20 with the idea of generating conservative outputs. to test for stability, we ran each experiment 5 times, randomly selecting examples for the prompt from our base of 10 examples. this kind of model is light enough to be run on simple hardware. (11) example 1: input: ah mais c’ est bon je sais plus quoi dire là c’ est bon hein it# enfin je sais pas output: | ah mais c’ est bon | je sais plus quoi dire là | c’ est bon hein | it# enfin je sais pas since the generative model impacts the tokenization (adding tokens in particular), we postprocessed the output with simple rules and excluded pauses from the evaluation, since it is the most added token, but never mark the beginning of a segment. this adjustment helped mitigate biases introduced by the deviant tokenization. note that it is a supplementary step that can increase development time much more than the term ”few-shot” might suggest, and indicates how brittle a generative approach can be for token-level predictions instead of simple classification, at least with moderately sized models. as can be seen in table 5, the 5-shot experiment amounts to about 20 to 30 discourse units that correspond roughly to our smallest training size in our fine-tuning experiment. the first conclusions we can draw from this experiment is that the generative approach works surprisingly well given how little supervision it receives, compared to fine-tuning with a few examples (lines 1-2 of table 4), but it is quite far in performance from more reasonably supervised approaches (supervised or weakly supervised), as the scores plateau quickly after a few examples. 20. https://ollama.com/ 52 discourse segmentation of conversational transcripts figure 1: f-score comparing fine-tuning and few-shots approaches. x-axis: number of tokens, log scale. green lines correspond to 200ms pause baseline and black line to human annotation topline (average across the languages). this approach is prone to some instability by nature and could be investigated more (including with larger models), but it would lose the point of decreasing engineering efforts. 4.4 discussion the key finding of our experiments on french conversations is that simple fine-tuning on gold labels performs well for elementary discourse unit segmentation. the results from the zero/few-shot pilot study, however, yielded contrasting outcomes. few-shot performance was achieved with 20-30 du examples, corresponding to scores obtained with 200-500 training tokens using the fine-tuning approach (roughly 30-70 dus). on one hand, this suggests that for extremely small datasets (or in cases where no gold data is available), prompting could be an option, for instance for scenarios in which pause duration (our baseline) is not available. on the other hand, the scores and error analysis indicate that fine-tuning with a few hundred annotated dus is generally the better strategy, nb of examples precision recall fscore 0 0.47± 0.01 0.58± 0.01 0.52± 0.01 1 0.51± 0.03 0.60± 0.01 0.55± 0.02 5 0.53± 0.02 0.63± 0.03 0.58± 0.02 10 0.52± 0.01 0.63± 0.01 0.57± 0.01 table 5: results for the zero/ in-context k-shot approach. 53 prévot & muller manual predicted mean 7.36 7.45 std 5.46 5.76 median 6 6 mode 1 1 table 6: comparison of du length distributions, manual vs. automated. as it produces higher scores. in contrast, prompting results are unpredictable, making it difficult to plan improvements in a rational, consistent manner. in a broader context, as discussed in section 2, this paper provides a global view of the available strategies for elementary discourse unit segmentation on spontaneous speech transcripts. these strategies include: fine-tuning on human gold annotations (with varying amounts of training data), weak supervision (similar to fine-tuning but using automatic annotations, such as those obtained through data programming), and zero/few-shot approaches. we observe that direct fine-tuning on gold data is the most efficient solution. the amount of data required to reach an 0.8 f-score corresponds to only a couple of days of manual annotation effort, whereas setting up a weakly-supervised approach would require at least the same amount of time with similar or lower results. while the zero/few-shot approach is faster to set up, it yields lower and, more importantly, unpredictable results, which makes it difficult to plan for improvements. as empirical linguists, our primary motivation for building a discourse segmentation tool is to enable access to large amounts of discourse-segmented data, with a comparable quality level, and to draw similar observations as if the dataset had been manually segmented, an option not available for large datasets. the standard metrics presented in table 6 reveal that the distributional shapes of discourse unit lengths are nearly indistinguishable between the automatic and manual segmentation. this is further established by plotting the distribution lengths between the manual and automatic version of the test set, as in figure 2. figure 2: distribution of du lengths in manually and automatically segmented datasets / dispersion of the du lengths in manually and automatically segmented datasets in a violin plot 54 discourse segmentation of conversational transcripts (a) initial (b) final (c) isolated figure 3: frequency distributions from left to right for du-initial, du-final (both for du having at least 3 tokens) and du-isolated (du made of only 1 token). we also plotted the distribution of most frequent lexical items in crucial positions, namely duinitial, du-final and du-isolated as presented in figure 3. these lexical distributions around du boundaries, obtained automatically or manually, are remarkably similar. some divergences occur at lower frequencies, such as the swear word putain, which the model does not treat as a single-token du (despite being relatively frequent as a single du in the manual annotation). 55 prévot & muller type of error disfluencies 34 % discourse markers 27 % non-sentential utterances 8 % other spoken genre structures 8 % table 7: type of errors observed in the systematic error analysis 5. error analysis the scores achieved by our systems are promising but still not at the level obtained on written genres. to better understand the shortcomings of our model and identify what needs to be improved to achieve human-level segmentation performance, we conducted an in-depth error analysis. we reviewed a sub-corpus accounting for about 5000 tokens21 from 10 different speakers in the corpus and systematically labeled the 230 segmentation discrepancies occuring in this sample between the human and automatic segmentations. we categorized errors at multiple levels, including: • the general nature of the discrepancy between gold labels and predicted labels: false positives, false negatives, and gold errors. • a general category of the cause of the errors: disfluencies, discourse markers, pauses, nonsentential units, other spoken structures, dialogical sequences, reported speech, relative clauses. • the potential role of pause presence/absence in errors: pause (for false positives) and no pause (for false negatives). • a finer-grained description of the causes of the errors: abandoned units, discourse marker clusters, discourse marker position inversion (final vs. initial), parentheticals, response insertions, etc. this analysis yielded 230 discrepancies at the token level between gold vs. predicted labels. their systematic characterization is presented in table 7. interestingly, 23% of the errors were found in the gold standard itself.22 additionally, 48% of the discrepancies were false negatives (missed boundaries), and 29% were false positives (spurious boundaries). as seen in table 7, the main cause of errors were predominantly disfluencies and discourse markers. other significant sources of error included spoken genre structures (16%), particularly non-sentential utterances (8%). the remaining errors were attributed to factors such as long pauses, dialogical sequences, relative clauses, or reported speech. these causes of errors are exemplified and further analyzed in the following subsections. we further examined the impact of pauses on 21. the size was arbitrarily chosen as a hopefully representative sample. 22. the manual segmentation was performed by semi-naive coders. the double-checking of these discrepancies was done by a discourse expert, co-author of this paper. this suggests that the overall system scores could be significantly higher than reported if the test set had been segmented and double-checked by experts, rather than relying on a seminaive annotation campaign. see also (nahum et al., 2025) for a discussion on the potential underestimation of model performance due to human errors in evaluation data. 56 discourse segmentation of conversational transcripts these errors. moreover, we found that 63% of the missed boundaries were due to the absence of a pause that the model could use, while 35% of the spurious boundaries were influenced by a pause, which likely contributed to the segmentation decision. 5.1 disfluencies simple disfluencies mostly did not cause issues for the segmentation models. errors appear when the disfluency is combined with a pause at a potential discourse break like (12) in which the sequence of a discourse marker + filler + pause triggered a spurious discourse boundary. (12) disfluencies + pause misused [je venais de temps en temps]du # [et euh # ] !! [ et puis un jour c’ était je sais pas l’ été je crois]du23 [i was coming from time to time]du # [and um # ] !! [ and then one day it was i don’t know summer i think]du complex paradigmatic piles (gerdes and kahane, 2009) generated some errors like in (13), maybe because the material accumulated before the decision point started to accumulate to form an acceptable looking discourse unit while the material after definitely constituted a valid discourse unit in itself. (13) disfluencies + paradigmatic pile [parce qu’ elle était pas très ] !! [ elle était ancienne]du [donc ça pouvait pas le faire]du [because it was not very ] !! [ it was old]du [so it could not work]du disfluencies involving discourse markers were also an issue, like in examples (14): on top of the repetition of the initial discourse marker there is a pause inserted to further mislead the model. (14) disfluencies + discourse markers a. [si euh # ] !! [ si tu veux je vais pas tourner de l’ oeil]du [if uh # ] !! [ if you want i won’t pass out]du b. disfluencies, self-edit with discourse marker [qui vient te euh d# ] !! [ enfin bon te recadrer quoi je veux dire]du [who come to you uh d# ] !! [ well to put you in your place i want to say]du disfluencies sometimes come with truncations that results in infrequent tokens for the model and that seems to have been an issue like in (15) in which in addition to the truncation at the boundary, there is no pause before to help the segmentation decision. (15) disfluency / truncation [mais # mais # si si il y était justement]du [ $$ ] [iil avait obligé à le à le laisser]du 23. in the error analysis examples, a ”]!![” indicates a spurious boundary, while ”[$$]” indicates a missed boundary. 57 prévot & muller [but # but # yes yes he was in finally]du [ $$ ] [ihe forced to let him]du to conclude about disfluencies, our initial discourse segmentation model used the notion of abandoned units (see also pallaud et al., 2013): false starts that are completely abandoned, and impossible to relate to the finally produced utterances. these abandoned units can create issues either because they are coded as separate units that the model would either include within the gold discourse unit at the beginning, as in example (17), or the end, as in example (16) of the discourse, leading to missed boundaries. (16) disfluencies, abandoned units on the right [genre des gens qui étaient au même niveau que moi quoi]du [ $$ ] [qui étaient euh # ] !! [ ou qui euh qui qui l’étaient]du [ $$ ] [mais qui euh de par leur âge # tu vois]du [like some people that were at the same level with me]du [ $$ ] [who were uh # ] !! [ or who uh who who were]du [ $$ ] [but who uh from their age # you see]du (17) disfluencies, abandoned units on the left [enfin si tu veux je normalement je deenfin]du # [ $$ ] [si tout va bien je vais essayer de le faire]du [well if you want i supposedly i dewell]du # [ $$ ] [if everything goes well i will try to do it]du 5.2 discourse markers discourse markers, being key cues for discourse boundaries, are involved in many errors. the main reason for this is that some of the most frequent discourse markers (such as bon / well) can appear both at the beginning and the end of a discourse unit, depending on their function. for instance, in example (18), the discourse marker is considered to be final, while the manual annotation treats it as initial. conversely, in example (19), the composite marker mais bon (but well) is considered to initiate a new discourse unit. however, while mais mostly functions as an initial marker, mais bon is more commonly used as a final attitudinal stance marker (see hancil, 2015, for a similar observation on the english final but) . (18) discourse marker, final initial confusion [et puis en fait non j’ étais]du [ $$ ] [bon ] !! [ j’ avais pas trop le moral]du [and then in fact no i was]du [ $$ ] [well ] !! [ i wasn’t feeling too good]du (19) discourse marker, final initial confusion [ça s’ est jamais bien passé je crois d’ ailleurs ] !! [ mais bon]du # [là vraiment j’ étais verte]du [it never went well i think by the way ] !! [ but well]du # [at that moment i was really disappointed]du 58 discourse segmentation of conversational transcripts conversational spontaneous speech also present large bundles or discourse markers (muller and prévot, 2003). the model struggle to identify where the boundary should be, if any is really present in these sequences like in (20) in which the mais (but) is wrongly considered to segment the unit. (20) discourse marker, cluster [ah ben ouais ] !! [ mais ça]du # [non elles auraient dû au lieu d’ emmener du des croissants]du [oh well yeah] !! [ but this]du # [non they should have instead of bringing some some croissants]du some positive feedback markers, in du-initial position are also used as ’pivot’ (gravano et al., 2011) and in that case tend to be included in their du-host like in (21), that the model failed to adapt to. (21) discourse marker, pivot [tu les as jamais trouvés ouais]du # [ouais # ] !! [ c’ était hallucinant]du [you never found them yeah]du # [yeah # ] !! [ it was amazing]du 5.3 the case of parentheticals we would like to emphasize a last family of errors related to parentheticals (asher, 2000). it was not possible to annotate truly embedded discourse units within another discourse units in our annotation model and framework.24 while deliberate, this choice had some negative consequences for examples like (22) and more crucially (23) for the human coder decided to promote the embedded material as full-fledged discourse units. in a more expressive framework, these two examples would have been handled homogeneously, leading to a clearer situation for these parentheticals (even though they probably would have been harder to learn for the model). (22) parentheticals [elles auraient dû prendre euh ] !! [ je sais pas une ou deux bières]du [they should have brought uh ] !! [ i don’t know one or two beers]du (23) embedded parenthetical [alors que y a]du [ $$ ] [enfin euh à mon sens]du [ $$ ] [y avait pas de norme]du [while there is]du [ $$ ] [well uh in my sense]du [ $$ ] [there was no norm]du 5.4 other sources of errors the majority of the remaining errors relate to various spontaneous speech structures or specific uses, like some relatives that introduce clauses that are missed by the model like the first false negative in (24). 24. this choice was made for efficiency reasons after some pilot annotation studies with and without embedded units. allowing embedded units made the whole process slower and more cumbersome while most of the units obtained this way were rather trivial and could be considered as stance markers for their host, in a medial position. 59 prévot & muller (24) relative [j’ai trouvé derrière le le micro-ondes ] !! [ donc une multiprise]du [ $$ ] [où y avait à la fois le frigo]du [ $$ ] [enfin # tout ce qui pouvait être branché suquelque part]du [i found behind the the microwave ] !! [ so a power strip]du [ $$ ] [where there was at the same time the fridge]du [ $$ ] [well # everything that could be plugged onsomewhere]du as mentioned in the quantitative overview, non-sentential units (fernández and ginzburg, 2002; fernández et al., 2007) were also a source of confusion like for (25) but also some versions of short canonical utterances (for the spontaneous speech genre) could pose problem like (26). (25) non sentential units [des gens qui avaient des expériences]du [ $$ ] [mais dans un domaine différent]du [people who had experiences]du [ $$ ] [but in a different domain]du (26) spoken canonical, very short utterance [bon tu commences]du [ $$ ] [tu en as en tête]du [well you start]du [ $$ ] [you have some in mind]du we can also mention rare but highly problematic cases of reported speech with dialogues present in the reported speech like (27). (27) reported speech, dialogical [euh donc on a tout # effacé]du # [ $$ ] [quoi]du # [mais pourquoi]du [um so we erased everything]du # [ $$ ] [what]du # [but why]du 5.5 error analysis summary to summarize the error analysis, the remaining errors produced by our models correspond to the expected categories. fine-tuning the base model significantly improves performance by adapting it to disfluencies, pauses, and discourse markers. however, when these phenomena diverge further from the written pre-training data, because they involve combination of these phenomena, they are still misclassified by the model. it is important to note that some of these difficult cases are inherently ambiguous and require a deep understanding of discourse sequences to be resolved correctly. as mentioned in the error analysis, even human coders struggled to find a coherent segmentation when disfluencies, pauses, and discourse markers are intertwined into hard-to-interpret sequences. the way pauses are handled, often over-trusted by the model, appears to be the most promising area for further improvement. prosody will play a key role in further characterizing what happens before and after a pause, potentially allowing us to use pause duration in a more nuanced way. finally, the current approach, where the discourse flow is treated on a speaker-by-speaker basis rather than in a truly dialogical context, does not seem to cause significant errors compared to the simplicity gained in terms of pipeline and model. in fact, errors related to dialogical factors 60 discourse segmentation of conversational transcripts account for only a small percentage of the total errors (see appendix 11 for detailed numbers). furthermore, some of these errors were actually related to dialogue occurring within the reported speech of a single speaker. therefore, we do not plan to change the overall pipeline we adopted for this dataset. however, more interactional, heavily dialogical genres might benefit from a truly dialogical discourse segmentation approach. 6. discussion the previous sections have explored methods for performing automatic segmentation of elementary discourse units (edu) in spontaneous speech transcripts, achieving near-human-level results. this task is crucial in the current quantitative linguistic landscape. indeed, without discourse segmentation, the alternatives for obtaining ”sentence” or ”utterance” units are limited to human segmentation, acoustic proxies (such as inter-pausal units), or syntactic proxies (such as re-punctuation methods). our experiments show that, at least for our dataset, a relatively small amount of annotated data is sufficient to approach human performance. in contrast, the two proxies mentioned can only be considered as baselines. some might argue that our dataset is quite specific, combining dialogical and monological dimensions, making the proxies inefficient for part of the data. however, we consider on the contrary that most conversational situations combine or alternate between highly interactive sequences (short turns) and more expository, longer-turn segments (it’s the case in meetings for instance). thus, the proxies mentioned above are not sufficient. we believe that the ability to segment large spontaneous speech corpora with these units opens up various possibilities for revisiting critical research questions, such as turn-taking, discourse prosody, multi-modal interfaces, and, more broadly, language structure as revealed by spontaneous speech. having discourse units that are determined not only from syntax but also from broader semantic and discourse aspects is a step toward more efficient operationalization of other linguistic models. for instance, our discourse boundaries can be used as transition-relevant places (trps) in the conversational analysis framework (sacks et al., 1974), contributing to a more practical definition of conversational units, as suggested by ford and thompson (1996). they can also be systematically explored within dualistic frameworks (haselow, 2017; pietrandrea and kahane, 2019). regarding turn-taking, using inter-pausal units (ipus) directly introduces circularity. elementary discourse unit boundaries, however, provide an operational determination of the trp, opening up the possibility of more accurate (than previous proxies) and large-scale corpus studies on turn-taking and turn transitions. together, with prosody extraction tools, discourse segmentation models of the kind presented here, can label with a near-human level large speech corpora. this is an important step for discourse-prosody interface research as it tend to be a highly variable phenomena and models would benefit from empirical evidence from larger corpora coming from more diverse sub-genres to boost their empirical foundation. 61 prévot & muller concerning language structures, since pauses are not linguistic in nature, they are the source of untypical segmentations, leading to artificially complex structures to describe. the main reason is that pauses can be inserted almost anywhere. this is what happens in many disfluent or interaction sequences in which even rather long pauses (> 500ms) do not interrupt discourse units. ’sentences’, on the other hand, reliably allow for identifying large language processing structures. spontaneous spoken sequences thus differ radically through many different syntactic patterns from canonical traditional grammar as explained in (ginzburg and poesio, 2016). there are several directions for further development of the current approach. although we are near human-level performance, we are not there yet. as seen in the error analysis, many of the errors are tied to issues such as edu boundaries without pauses or over-interpretation of pauses by the model. the next step is to model the prosodic context for each segmentation decision, with the aim of identifying the locations of prosodic breaks without pauses and distinguishing between pauses that do not correspond to real prosodic breaks. another promising direction is to apply our model (perhaps with further fine-tuning on small-scale gold data) to other genres, corpora, and languages. as demonstrated in recent work, fine-tuning roberta for lower-resource languages has proven to be a promising solution. specifically, prévot and wang (2024) showed that fine-tuning roberta yielded good results for taiwan southern min, even though this language is not included in the pretraining corpus. acknowledgments this work was conducted over several years and benefited from multiple sources of support. we gratefully acknowledge funding from the anr project summ-re (anr-20-ce23-0017), the institut convergence ilcb (anr-16-conv-0002), and the institut universitaire de france (iuf). laurent prévot further thanks the institute of linguistics, academia sinica, for hosting him during the final stages of this research. references jeremy ang, yang liu, and elizabeth shriberg. automatic dialog act segmentation and classification in multiparty meetings. in proc. icassp, volume 1, pages i–1061. ieee, 2005. ron artstein and massimo poesio. inter-coder agreement for computational linguistics. computational linguistics, 34(4):555–596, 2008. nicholas asher. truth conditional discourse semantics for parentheticals. journal of semantics, 17 (1):31–50, 2000. nicholas asher and alex lascarides. logics of conversation. cambridge university press, 2003. john langshaw austin. how to do things with words. harvard university press, 1975. sonia badene, kate thompson, jean-pierre lorré, and nicholas asher. weak supervision for learning discourse structure. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), pages 2296–2305, hong kong, china, november 2019. association for 62 discourse segmentation of conversational transcripts computational linguistics. doi: 10.18653/v1/d19-1234. url https://aclanthology. org/d19-1234. christophe benzitoun and frédéric sabio. où finit la phrase? où commence le texte?. l’exemple des regroupements de constructions verbales. discours. revue de linguistique, psycholinguistique et informatique. a journal of linguistics, psycholinguistics and computational linguistics, (7), 2010. philippe blache, roxane bertrand, and gaëlle ferré. creating and exploiting multimodal annotated corpora: the toma project. multimodal corpora, pages 38–53, 2009. philippe blache, roxane bertrand, gaëlle ferré, berthille pallaud, laurent prévot, and stéphane rauzy. the corpus of interactional data: a large multimodal annotated resource. in handbook of linguistic annotation, pages 1323–1356. springer, 2017. claire blanche-benveniste, mireille bilger, christine rouget, karel. van den eynde, piet mertens, and dominique willems. le français parlé (études grammaticales). sciences du langage, 1990. paul boersma. praat, a system for doing phonetics by computer. glot international, 5(9/10):341– 345, 2002. chloé braud, yang janet liu, eleni metheniti, philippe muller, laura rivière, attapol t rutherford, and amir zeldes. the disrpt 2023 shared task on elementary discourse unit segmentation, connective detection, and relation classification. in 3rd shared task on discourse relation parsing and treebanking (disrpt 2023), pages 1–21. acl: association for computational linguistics, 2023. david brazil. a grammar of speech. oxford university press, usa, 1995. harry bunt. multifunctionality in dialogue. computer speech & language, 25(2):222–245, 2011. alexis conneau, kartikay khandelwal, naman goyal, vishrav chaudhary, guillaume wenzek, francisco guzmán, edouard grave, myle ott, luke zettlemoyer, and veselin stoyanov. unsupervised cross-lingual representation learning at scale. corr, abs/1911.02116, 2019. url http://arxiv.org/abs/1911.02116. elizabeth couper-kuhlen and margret selting. introducing interactional linguistics. studies in interactional linguistics. amsterdam: john benjamins, pages 1–22, 2001. ludivine crible and liesbeth degand. domains and functions: a two-dimensional account of discourse markers. discours. revue de linguistique, psycholinguistique et informatique. a journal of linguistics, psycholinguistics and computational linguistics, (24), 2019. viet-trung dang, tianyu zhao, sei ueno, hirofumi inaguma, and tatsuya kawahara. end-to-end speech-to-dialog-act recognition. in helen meng, bo xu, and thomas fang zheng, editors, interspeech 2020, 21st annual conference of the international speech communication association, virtual event, shanghai, china, 25-29 october 2020, pages 3910–3914. isca, 2020. doi: 10.21437/interspeech.2020-1062. url https://doi.org/10.21437/interspeech. 2020-1062. 63 prévot & muller liesbeth degand and anne catherine simon. minimal discourse units: can we define them, and why should we. proceedings of sem-05. connectors, discourse framing and discourse structure: from corpus-based and experimental analyses to discourse theories, biarritz, pages 14–15, 2005. liesbeth degand and anne catherine simon. on identifying basic discourse units in speech: theoretical and empirical issues. discours. revue de linguistique, psycholinguistique et informatique. a journal of linguistics, psycholinguistics and computational linguistics, (4), 2009. henri-josé deulofeu. l’approche macrosyntaxique en syntaxe: un nouveau modèle de rasoir d’occam contre les notions inutiles? scolia: sciences cognitives, linguistiques et intelligence artificielle, 16(1):77–95, 2003. abhimanyu dubey, abhinav jauhri, abhinav pandey, et al. the llama 3 herd of models. arxiv preprint 2407.21783, 2024. url https://arxiv.org/abs/2407.21783. benoit favre, dilek hakkani-tur, slav petrov, and dan klein. efficient sentence segmentation using syntactic features. in 2008 ieee spoken language technology workshop, pages 77–80. ieee, 2008. raquel fernández and jonathan ginzburg. non-sentential utterances: grammar and dialogue dynamics in corpus annotation. in proceedings of the 19th international conference on computational linguistics-volume 1, pages 1–7. association for computational linguistics, 2002. raquel fernández, jonathan ginzburg, and shalom lappin. classifying non-sentential utterances in dialogue: a machine learning approach. computational linguistics, 33(3):397–427, 2007. cecilia e ford and sandra a thompson. interactional units in conversation: syntactic, intonational, and pragmatic resources for the management of turns. studies in interactional sociolinguistics, 13:134–184, 1996. chris fournier and diana inkpen. segmentation similarity and agreement. in proceedings of the conference of the north american chapter of the association for computational linguistics: human language technologies, pages 152–161, montréal, canada, 2012. kim gerdes and sylvain kahane. speaking in piles: paradigmatic annotation of french spoken corpus. in fifth corpus linguistics conference, pages 1–15, 2009. jonathan ginzburg. the interactive stance: meaning for conversation. oxford university press, 2012. jonathan ginzburg and massimo poesio. grammar is a system that characterizes talk in interaction. frontiers in psychology, 7, 2016. john j godfrey, edward c holliman, and jane mcdaniel. switchboard: telephone speech corpus for research and development. in acoustics, speech, and signal processing, 1992. icassp-92., 1992 ieee international conference on, volume 1, pages 517–520. ieee, 1992. agustı́n gravano, julia hirschberg, and štefan beňuš. affirmative cue words in task-oriented dialogue. computational linguistics, 38(1):1–39, 2011. 64 discourse segmentation of conversational transcripts lila gravellier, julie hunter, philippe muller, thomas pellegrini, and isabelle ferrané. weakly supervised discourse segmentation for multiparty oral conversations. in proceedings of emnlp 2021, 2021. oliver guhr, anne-kathrin schumann, frank bahrmann, and hans joachim böhme. fullstop: multilingual deep models for punctuation prediction. june 2021. url http://ceur-ws. org/vol-2957/sepp_paper4.pdf. narendra k gupta and srinivas bangalore. segmenting spoken language utterances into clauses for semantic classification. in 2003 ieee workshop on automatic speech recognition and understanding (ieee cat. no. 03ex721), pages 525–530. ieee, 2003. sylvie hancil. the grammaticalization of final but: from conjunction to final particle. final particles, pages 197–218, 2015. alexander haselow. spontaneous spoken english: an integrated approach to the emergent grammar of speech. cambridge university press, 2017. marti a hearst. multi-paragraph segmentation expository text. in 32nd annual meeting of the association for computational linguistics, pages 9–16, 1994. sophie herment, laetitia leonarduzzi, diana lewis, cristel portes, laurent prévot, frédéric sabio, and gabor turcsan. périphéries gauche et droite. tipa. travaux interdisciplinaires sur la parole et le langage, (38), 2022. julia hirschberg and barbara grosz. intonational features of local and global discourse structure. in proceedings of the darpa workshop on spoken language systems. association for computational linguistics, 1992. jet hoek, jacqueline evers-vermeul, and ted jm sanders. segmenting discourse: incorporating interpretation into segmentation? corpus linguistics and linguistic theory, 14(2):357–386, 2018. junfei hu and liesbeth degand. the conversational discourse unit: identification and its role in conversational turn-taking management. dialogue & discourse, 14(2):83–112, 2023. dan jurafsky, liz shriberg, and debra biasca. switchboard swbd-damsl shallow-discourse-function annotation coders manual. technical report, university of colorado at boulder, 1997. diana m lewis. pragmatic markers at the periphery and discourse prominence. pragmatic markers and peripheries, 325:351, 2021. per linell. the written language bias in linguistics: its nature, origins and transformations. routledge, 2004. pierre lison, jeremy barnes, and aliaksandr hubin. skweak: weak supervision made easy for nlp. in heng ji, jong c. park, and rui xia, editors, proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing: system demonstrations, pages 337–346, online, august 2021. association for computational linguistics. doi: 10.18653/v1/2021.acl-demo.40. url https: //aclanthology.org/2021.acl-demo.40/. 65 prévot & muller yang liu, andreas stolcke, elizabeth shriberg, and mary harper. using conditional random fields for sentence boundary detection in speech. in proceedings of the 43rd annual meeting of the association for computational linguistics (acl’05), pages 451–458, 2005. william c mann and sandra a thompson. rhetorical structure theory: description and construction of text structures. in natural language generation: new results in artificial intelligence, psychology and linguistics, pages 85–95. springer, 1987. daniel marcu. the theory and practice of discourse parsing and summarization, 2000. yann mathet, antoine widlöcher, and jean-philippe métivier. the unified and holistic method gamma (γ) for inter-annotator agreement measure and alignment. computational linguistics, 41 (3):437–479, 2015. eleni metheniti, chloé braud, philippe muller, and laura rivière. discut and discret: melodi at disrpt 2023. in 3rd shared task on discourse relation parsing and treebanking (disrpt 2023), pages 29–42. acl: association for computational linguistics, 2023. eleni metheniti, philippe muller, chloé braud, and margarita hernández casas. zero-shot learning for multilingual discourse relation classification. in nicoletta calzolari, min-yen kan, veronique hoste, alessandro lenci, sakriani sakti, and nianwen xue, editors, proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (lrec-coling 2024), pages 17858–17876, torino, italia, may 2024. elra and iccl. url https://aclanthology.org/2024.lrec-main.1553. philippe muller and laurent prévot. an empirical study of acknowledgment structures. in 7th workshop on the semantics and pragmatics of dialogue, 2003. philippe muller, chloé braud, and mathieu morey. tony: contextual embeddings for accurate multilingual discourse segmentation of full documents. in proceedings of the workshop on discourse relation parsing and treebanking 2019, pages 115–124. association for computational linguistics, 2019. omer nahum, nitay calderon, orgad keller, idan szpektor, and roi reichart. are llms better than reported? detecting label errors and mitigating their effect on model performance. in christos christodoulopoulos, tanmoy chakraborty, carolyn rose, and violet peng, editors, proceedings of the 2025 conference on empirical methods in natural language processing, pages 26770– 26797, suzhou, china, november 2025. association for computational linguistics. isbn 9798-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.1360. url https://aclanthology. org/2025.emnlp-main.1360/. kota shamanth ramanath nayak. does chatgpt measure up to discourse unit segmentation? a comparative analysis utilizing zero-shot custom prompts. 2024. joakim nivre. dependency parsing. language and linguistics compass, 4(3):138–152, 2010. berthille pallaud, stéphane rauzy, and philippe blache. auto-interruptions et disfluences en français parlé dans quatre corpus du cid. tipa. travaux interdisciplinaires sur la parole et le langage, (29), 2013. 66 discourse segmentation of conversational transcripts rebecca j passonneau and diane litman. discourse segmentation by human and automated means. computational linguistics, 23(1):103–139, 1997. klim peshkov and laurent prévot. segmentation evaluation metrics, a comparison grounded on prosodic and discourse units. in lrec, pages 321–325, 2014. volha petukhova, laurent prévot, and harry bunt. multi-level discourse relations between dialogue units. in proceedings 6th joint acl-iso workshop on interoperable semantic annotation (isa-6), oxford, pages 18–27, 2011. lev pevzner and marti a. hearst. a critique and improvement of an evaluation metric for text segmentation. computational linguistics, 28(1):19–36, 2002. janet pierrehumbert and julia bell hirschberg. the meaning of intonational contours in the interpretation of discourse. in intentions in communication. mit press, 1990. paola pietrandrea and sylvain kahane. macrosyntactic annotation. rhapsodie: a prosodic and syntactic treebank for spoken french, john benjamins, amsterdam, pages 97–126, 2019. livia polanyi and remco jh scha. the syntax of discourse. text-interdisciplinary journal for the study of discourse, 3(3):261–270, 1983. livia polanyi, chris culy, martin van den berg, gian lorenzo thione, and david ahn. a rule based approach to discourse parsing. in proceedings of the 5th sigdial workshop on discourse and dialogue at hlt-naacl 2004, pages 108–117, 2004. salvador pons borderı́a. models of discourse segmentation in romance languages: an overview. discourse segmentation in romance languages, pages 1–21, 2014. cristel portes and roxane bertrand. permanence et variation des unités prosodiques dans le discours et l’interaction. journal of french language studies, 21(1), 2011. laurent prévot and sheng-fu wang. experimenting with discourse segmentation of taiwan southern min spontaneous speech. in proceedings of the 5th workshop on computational approaches to discourse (codi 2024), pages 50–63, 2024. laurent prévot, roxane bertrand, klim peshkov, stéphane rauzy, and philippe blache. prosody, discourse and syntax in french conversations: resource creation and evaluation. (submitted), 2016. laurent prévot, roxane bertrand, and stéphane rauzy. investigating disfluencies contribution to discourse-prosody mismatches in french conversations. in the 10th workshop on disfluency in spontaneous speech, 2021. laurent prevot, julie hunter, and philippe muller. comparing methods for segmenting elementary discourse units in a french conversational corpus. in 24th nordic conference on computational linguistics (nodalida 2023). acl: association for computational linguistics; university of tartu library, 2023. 67 prévot & muller laurent prévot, roxane bertrand, and julie hunter. segmenting a large french meeting corpus into elementary discourse units. in proceedings of the 26th annual meeting of the special interest group on discourse and dialogue, pages 183–191, 2025. yada pruksachatkun, phil yeres, haokun liu, jason phang, phu mon htut, alex wang, ian tenney, and samuel r. bowman. jiant: a software toolkit for research on general-purpose text understanding models. in asli celikyilmaz and tsung-hsien wen, editors, proceedings of the 58th annual meeting of the association for computational linguistics: system demonstrations, pages 109–117, online, july 2020. association for computational linguistics. doi: 10.18653/ v1/2020.acl-demos.15. url https://aclanthology.org/2020.acl-demos.15/. silvia quarteroni, alexei v. ivanov, and giuseppe riccardi. simultaneous dialog act segmentation and classification from human-human spoken conversations. in proc. icassp, pages 5596–5599, 2011. doi: 10.1109/icassp.2011.5947628. alexander ratner, stephen h bach, henry ehrenberg, jason fries, sen wu, and christopher ré. snorkel: rapid training data creation with weak supervision. in proceedings of the vldb endowment. international conference on very large data bases, volume 11, page 269. nih public access, 2017. eddy roulet, filliettaz laurent, anne grobet, and marcel burger. un modèle et un instrument d’analyse de l’organisation du discours. bern: lang, 2001. frédéric sabio. phrases et constructions verbales: quelques remarques sur les unités syntaxiques dans le français parlé. in constructions verbales et production de sens, pages 127–139. presses universitaires de franche-comté, 2006. harvey sacks. lecture on conversations. basil blackwell, 1992. harvey sacks, emanuel a schegloff, and gail jefferson. a simplest systematics for the organization of turn-taking for conversation. language, pages 696–735, 1974. emmanuel a. schegloff and harvey sacks. opening up closings. semiotica, 8, 1973. margret selting. the construction of units in conversational talk. language in society, 29(4): 477–517, 2000. elizabeth shriberg, andreas stolcke, dilek hakkani-tür, and gökhan tür. prosody-based automatic segmentation of speech into sentences and topics. speech communication, 32(1):127–154, 2000. john sinclair and malcolm coulthard. toward an analysis of discourse. in advances in spoken discourse analysis. routledge, 1992. radu soricut and daniel marcu. sentence level discourse parsing using syntactic and lexical information. in proceedings of the 2003 human language technology conference of the north american chapter of the association for computational linguistics, pages 228–235, 2003. manfred stede. discourse processing, volume 15. morgan & claypool publishers, 2012. 68 discourse segmentation of conversational transcripts amanda stent. rhetorical structure in dialog. in inlg’2000 proceedings of the first international conference on natural language generation, pages 247–252, 2000. andreas stolcke and elizabeth shriberg. automatic linguistic segmentation of conversational speech. in proceeding of fourth international conference on spoken language processing. icslp’96, volume 2, pages 1005–1008. ieee, 1996. andreas stolcke, elizabeth shriberg, rebecca a bates, mari ostendorf, dilek zeynep hakkani, madelaine plauche, gökhan tür, and yu lu. automatic detection of sentence boundaries and disfluencies based on recognized words. in icslp, volume 2, pages 2247–2250. citeseer, 1998. andreas stolcke, klaus ries, noah coccaro, elizabeth shriberg, rebecca bates, daniel jurafsky, paul taylor, rachel martin, carol van ess-dykema, and marie meteer. dialogue act modeling for automatic tagging and recognition of conversational speech. computational linguistics, 26 (3):339–373, 2000. zeno vendler. verbs and times. the philosophical review, pages 143–160, 1957. amir zeldes, debopam das, erick galani maziero, juliano antonio, and mikel iruskieta. the disrpt 2019 shared task on elementary discourse unit segmentation and connective detection. in proceedings of the workshop on discourse relation parsing and treebanking 2019, pages 97–104, 2019. amir zeldes, yang janet liu, mikel iruskieta, philippe muller, chloé braud, and sonia badene, editors. proceedings of the 2nd shared task on discourse relation parsing and treebanking (disrpt 2021), punta cana, dominican republic, november 2021. association for computational linguistics. url https://aclanthology.org/2021.disrpt-1.0. tianyu zhao and tatsuya kawahara. a unified neural architecture for joint dialog act segmentation and recognition in spoken dialog system. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 201–208, 2018. 69 prévot & muller appendix a. additional inter-coder agreement spk ab ac ag ap bx cm eb im lj ll mb mg ml nh sr ym mean annot a,e b,d d,a c,e a,e a,e c,e d,a c,e d,c b,d a,e d,a d,c c,e d,a κ 0.854 0.857 0.829 0.839 0.868 0.846 0.856 0.825 0.840 0.868 0.853 0.841 0.860 0.856 0.848 0.823 0.848 table 8: κ scores for the discourse boundaries. spk ag ll nh ym mean annot a,d,b a,c,d,e a,c,d,e a,d,b κ 0.837 0.783 0.842 0.853 0.829 table 9: multi-κ scores for the discourse boundaries on 15 minute fragments by 4 speakers, annotated by 3 or 4 naive annotators. annotator expert 1 expert 2 mean annot a 0.753 0.803 0.778 annot c 0.809 0.877 0.843 annot d 0.811 0.885 0.848 annot e 0.805 0.880 0.843 mean 0.828 table 10: κ scores for the discourse boundaries of each naive annotator versus expert on 15 minute fragments by two (30 minutes in total). inter-annotator agreement on discourse units between naive annotators is reported in tables 8 and 9. the whole duration for each speaker was annotated by two naive annotators, the κ scores per speaker are shown in table 8. table 9 contains multi-κ values for 15 minute excerpts of 4 speakers, two of these extracts were annotated by 3 annotators and the other two by 4 annotators. according to these scores, the inter-annotator agreement for discourse boundaries is consistently high across speakers and annotators. 70 discourse segmentation of conversational transcripts appendix b. supplementary figures and tables figure 4: precision / recall comparing fine-tuning and few-shots approaches. x-axis: nb of tokens, log scale. shaded areas are 95% confidence intervals. main apparent error cause count disfluencies 60 discourse marker 48 spoken structure 28 non sentential unit 14 dialogical 7 canonical 7 only pause 5 reported speech / attribution 5 relative clause 2 table 11: error analysis detailed numbers appendix c. examples of labelling functions pseudo-code in python. each document indexes its tokens. tokens have a few attributes, mostly their forms and part-of-speech tags, and a duration in case of a pause. rules output a label, and a token index position, coding if a boundary label must be before or after the token (as this is rule dependent). non_end_tok = pro_suj + pro_oth + pro_rel + dem + det + prep + oth + neg + dm_ini + relatives def non_end_tok(doc): for idx, token in enumerate(doc): 71 prévot & muller if idx > 1: if doc[idx 1].text in non_end_tok: yield idx, idx + 1, nobeg def long_pause(doc): for idx, token in enumerate(doc): if idx > 0: if doc[idx 1].text in [’#’] and doc[idx 1]._.dur > 0.8: if doc[idx].text not in [’#’]: yield idx, idx + 1, beg yield idx 1, idx, nobeg else: if idx < len(doc) 1: yield idx + 1, idx + 2, beg else: yield idx, idx + 1, beg else: if doc[idx].text not in [’#’]: yield idx, idx + 1, beg appendix d. few-shot prompt example your task is to segment a dialog in elementary discourse units, where the character # indicates a short pause here are a few examples of the expected results: example 1: input: ah mais c’ est bon je sais plus quoi dire là c’ est bon hein it# enfin je sais pas output: | ah mais c’ est bon | je sais plus quoi dire là | c’ est bon hein | it# enfin je sais pas example 2: input: ah ouais ah ouais et donc si tu veux il avait acheté tu sais des bouquins de euh # je me demande si c’ était pas le reboul là tu sais le gle le # la bible de la cuisine provençale enfin je sais plus # et bref il voulait faire d# de la crème de marrons # tu vois # donc euh # iil prend la * recette et tout bon il vil dit bon ok # ce c# cet idiot tu sais ce qu’ il fait # il va prendre des marrons tu t’ en rappelles # il va prendre des marrons sur le cours mirabeau ou dans le parc jourdan je sais pas quoi enfin bref il ramasse ses marrons pourris quoi et # il les pèle # enfin je sais pas à quoi il pense il a pris une journée à 72 discourse segmentation of conversational transcripts faire la crème de marrons output: | ah ouais | ah ouais | et donc si tu veux il avait acheté tu sais des bouquins de euh # | je me demande si c’ était pas le reboul là | tu sais le gle le # la bible de la cuisine provençale | enfin je sais plus # | et bref il voulait faire d# de la crème de marrons # tu vois # | donc euh # iil prend la * recette et tout bon | il vil dit | bon ok # | ce c# cet idiot tu sais ce qu’ il fait # | il va prendre des marrons | tu t’ en rappelles # | il va prendre des marrons sur le cours mirabeau ou dans le parc jourdan je sais pas quoi | enfin bref | il ramasse ses marrons pourris quoi | et # il les pèle # | enfin je sais pas à quoi il pense | il a pris une journée à faire la crème de marrons | example 3 input: ça c’ était les # les chocs culturels output: | ça c’ était les # les chocs culturels # | example 4 input: et ouais en plus elle toute seule du coup # bon là c’ est pas rigolo parce qu’ ils vont déménager mais # mais au moins iil sera là-haut quoi tu vois output: | et ouais en plus elle toute seule du coup # | bon là c’ est pas rigolo | parce qu’ ils vont déménager | mais # mais au moins iil sera là-haut quoi tu vois exemple 5 input: ouais ouais # ou en vrai aussi des fois non # ah ouais output: | ouais ouais # | ou en vrai aussi des fois non # | ah ouais | here is the text to segment: # on l’ assimile # après on va la dionly print the original text, indicating segment boundaries with a | character, and do not add anything, do not remove anything. do not present the result with introductory text. 73 journal of machine learning research-microsoft word template dialogue & discourse 16(1) (2025) 68-90 doi: https://doi.org/10.5210/dad.2025.103 ©year nguyen and fox tree this is an open-access article distributed under the terms of a creative commons attribution license (http://creativecommons.org/licenses/by/3.0/). pragmatic uses of i don’t know, boosters, and hedges in text and talk allison nguyen anguye9@ilstu.edu department of psychology illinois state university jean e. fox tree foxtree@ucsc.edu department of psychology university of california santa cruz submitted 10/2024; accepted 03/2025; published online 05/2025 abstract we examined how the phrases i don’t know, i dunno, and idk are used in spontaneously produced speech and writing. we compared functions to the related phrases totally, absolutely, sorta, and kinda. we assessed usage across modalities (face to face, instant messaging, audiovisual), goals (tasks versus casual chat), and relationships (friends versus strangers). we also assessed where the phenomena occurred in a sentence, what words co-occurred with the phenomena, and what functions the phenomena served in the conversations. communicators use phenomena differently depending on modality, goals, and relationships. we found that i don’t know was used more often when people could access cues beyond the voice, and that both i don’t know and i dunno can perform a variety of pragmatic functions. in instant messaging, i don’t know has been lexicalized to idk, but idk does not have as many pragmatic functions as i don’t know and i dunno. keywords: i don’t know, discourse markers, hedges, boosters, communication modality 1 introduction in a commercial from 2007 a mother is talking to her daughter about their high phone bill, asking “who could you be texting?” the daughter responds by saying out loud letters that are normally spoken, “i d k my b f f jill” (solarmax, 2007). in order for the mother to understand what the daughter means, the two have to engage in the process of mutually attempting to understand one another, particularly around what it means when the daughter says idk. while bff is an acronym for best friend forever (which the mother may or may not understand), idk can have multiple meanings, including marking uncertainty, prefacing disagreement, or highlighting commitment to the answer. conversations are processes that participants mutually engage in. a primary goal is to reach mutual understanding (clark, 1996; clark & brennan, 1991; stalnaker, 2002). in the process of reaching mutual understanding, or common ground, speakers and listeners must have access to previously grounded-upon information, as well as knowledge about their interlocutors, such as the type of relationship the two speakers share (metzing and brennan, 2003; myrendal, 2019). interlocutors use this knowledge to engage in small negotiations on things like referential expressions and conceptual pacts (brennan & clark, 1996; soler et al., 2023). this is a delicate process, where negotiators must navigate politeness and social dynamics (brown & levinson, 1987; beltrama & schwarz, 2024; scheibman, 2000; tsui, 1991) while suggesting to their partner that they wish to negotiate on meaning. one way that speakers can navigate the complex social https://doi.org/10.5210/dad.2025.103 i don’t know, boosters, and hedges 69 aspects of conversations is by choosing to use words that cue their conversational partner about upcoming negotiations. negotiation words are a set of words that speakers can choose to deploy to highlight the negotiation aspects of establishing common ground. while negotiation is a key feature of reaching common ground, this process is not always explicitly expressed. when someone says “it’s fun,” the addressee understands that the speaker found something fun. however, saying “it’s kinda fun” requires negotiation about what is meant by “fun.” in spontaneously produced speech, words that highlight negotiation include those commonly identified as hedges (sorta, kinda) and words commonly identified as boosters (absolutely, totally), as well as items like i don’t know, which fall outside the classification of hedge or booster. if conversational partners do not take contextual factors into account as part of the grounding process, we should expect to see no differences across text and talk. that is, speakers should use negotiation words in the same manner, with the same meaning across discourse modalities. additionally, these expressions should not vary across discourse contexts, such as task or non-task conversation, or across friendship status. said another way, an i don’t know should signal a lack of information to the same degree across modality type, task type, and friendship type. in this paper, we assessed how conversation medium (in person face to face, audio-only, text-only), conversation type (on-task or off-task), and relationship type (friends or strangers) affected use of i don’t know, absolutely, totally, kinda, and sorta. our assessment included examination of where i don’t know, boosters, and hedges were used and what they co-occurred with. we included four words of negotiation (two boosters, absolutely and totally, and two hedges, sorta and kinda) to understand how i don’t know contrasts with these similar expressions. the medium, conversation type, and relationship type varied across three corpora. the roommates corpus was of face to face conversation in a laboratory, the artwalk corpus was of on-task and off task audio conversation across friends and strangers in a partially-in-the-lab and partially-in-theworld setting, and the instant messaging corpus was of texted communication in a laboratory. 1.1 boosters and hedges negotiation words occur in both speech and writing, although the set of words found in different communicative settings vary widely. in writing, for example, the top boosters and hedges might be modal verbs like could, would, and will (takimoto, 2015: 98) or emotionally-charged intensifiers such as dishearteningly weak and particularly encouraging (salager-meijer, 1994: 7). in speech, in contrast, the top boosters and hedges might be single words like just and could (nuraniwati et al., 2021: 210) or two word expressions like i think, you know, and sort of (holmes, 1990: 202). defining boosters and hedges is a difficult challenge as the words used in different settings vary greatly; consider, for example, words used in academic writing (e.g., the words thoroughly, undeniably, hypothetically, and reportedly, farrokhi et al., 2008: 97-98), versus prepared speaking like political speeches (e.g., the phrases have to make sure, will never, can probably, and believe that, kashiha, 2020: 92-93), versus spontaneously produced dialogues (e.g., the words seemingly or usually or the phrase in my view, islam et al., 2020: 3111). we now turn to definitions of boosters and hedges, along with information about the subset we explore in this report, which focuses on spontaneously-produced speech and writing. boosters are words that are canonically considered to mark conviction and knowledge on a stance (hyland, 1998a, 2000b) – in other words, boosters are able to maximize commitment to an utterance. in certain scientific disciplines, for example, boosters are strategically deployed in writing in order to increase writer authority and to comment on the strength of arguments (hyland, 1998a; namaziandost, 2017). in planned and spontaneous speech, speakers use boosters to emphasize certainty (holmes, 1990; jalilifar & alavi-nia, 2012) or to make their words credible (kashiha, 2020). speakers and writers can also use boosters to exaggerate, which can sometimes imply sarcasm (d’arcey et al., 2019). nguyen and fox tree 70 hedges, in contrast, are words or short phrases like sort of and kind of that communicate that the speaker may not be sure about the information being presented or how the addressee will receive it (hyland, 1998a). in speech, many think of hedges as undesirable (e.g., berger, 2024). however, there is evidence that hedges provide meaningful information, such as marking information as unsure (jucker & smith 1996) or indicating that speakers want to distance themselves from claims (kashiha, 2020). hedges also affect language processing, such as by increasing memory for items (liu & fox tree, 2012). liu and fox tree (2012) found that in storytelling contexts, hedged information was more likely to be omitted in a retelling, but more likely to be retained in recall tests. there is an extensive list of both boosters and hedges. we selected four of the most common expressions for our study: the boosters absolutely and totally and the hedges kinda and sorta. 1.2 i don’t know, i dunno, and idk in addition to hedges and boosters, there are words that are neither hedges nor boosters, and yet serve some of the same functions. i don’t know is a common expression in this category. on its face, i don’t know means the absence of knowing. the literal meaning is “i don’t have the requested information or knowledge.” but people often use i don’t know for a variety of pragmatic purposes. for example, they can say i don’t know to convey that they both know the information and that they don’t want to say what they know. saying i don’t know throws the ball into the addressees’ court in a way that no other negotiation word does. i don’t know is the opposite of a booster; on its face, it states a lack of knowledge as opposed to a certainty of knowledge. but at the same time, i don’t know is not necessarily a hedge. i don’t know can be said with certainty of lack of knowledge, for example. that is, speakers can say i don’t know to indicate both they don’t know and that they don’t want to say (pichler & hesson, 2016; grant, 2010; brennan & williams, 1995; tsui, 1991). i don’t know can also function as a marker of epistemic certainty and as a discourse marker (grant, 2010; kärkkäinen, 2010; doehler, 2016). in addition, i don’t know can be used to create distance between the speaker and a statement (grant, 2010) or to steer the conversation away from a topic (doehler, 2016). the steering-away use of i don’t know can be considered an ostensible use of i don’t know. ostensible language is language that has a particular use in its literal, on-record form (such as indicating lack of knowledge) but another use in an off-record form. importantly, the off-record interpretation is meant to be recognized by the addressee as an insincere act by the speaker (which differentiates ostensible language from lying). for example, ostensible invitations occur when addressees are invited to an event with the implication that the invitee is expected to turn down the invitation, while maintaining social relationships (isaacs & clark 1990; link & kreuz 2005). by issuing the invitation, the inviter is off-record indicating that they would still like to spend time with the invitee, and by accepting the pretense, the invitee is willing to leave things off the record. ostensibile language, which often occurs with hedges and vague language, requires the two interlocutors to mutually recognize and collude together as to what the off-record reasons are. with an ostensible use of i don’t know, the speaker wants to indicate that they know the information but they don’t want to say it, with the reason remaining off-record. such a reason could be that providing the information would hurt the feelings of the addressee, break someone else’s confidence, or cause the spread of information the speaker wants kept private. unlike lying, ostensible uses of i don’t know can convey subtler information than simple denials of knowledge. indeed, i don’t know in combination with pausing can be indicative of knowledge: when answering trivia questions, people who took time before saying i don’t know were judged as more likely to know the answer than those who said i don’t know more quickly (brennan & williams, 1995). other uses of i don’t know include managing turns, expressing attitudes, and reducing commitment to statements (diani, 2004; grant, 2010; weatherall, 2011). i don’t know, boosters, and hedges 71 i don’t know can also be used in turn organization. it can indicate that the speaker is planning to add additional information or mark a turn as containing non-standard information (doehler 2016). i don’t know also allows the speaker to exit a turn, even when the turn is not complete (doehler 2016). thus, i don’t know has multiple uses in conversation, both literal and pragmatic. what interpretation communicators adopt might depend on many contextual factors, including what the communicators know about each other, what communicative modality is being used, and the purpose of the conversation. there are two other forms of i don’t know that can appear. i dunno is a well attested reduced form of i don’t know that appears in spoken speech and should be considered a bound phrase, just like i don’t know (aijmer, 2009; scheibman, 2000). like i don’t know, i dunno can serve a variety of pragmatic functions. it can be used in a literal sense (“i don’t have the information requested”), but other attested uses include hedging, softening or avoiding disagreement, and turn exchange (schiebman, 2000). likewise, idk is a written form of i don’t know that appears in instant messaging, text messaging, and other forms of technology-mediated communication, and likely serves similar functions, though usage has not been studied. 1.3 differences across settings negotiation words are primarily produced in unprepared, unrehearsed settings. many of them can be categorized as discourse markers, or words that aid the smooth flow of conversation (fox tree, 2010a, 2015b). other spontaneous elements that contribute to conversation management include backchannels (nguyen et a., 2024; tolins & fox tree, 2014; zellers, 2021) and fillers (clark & fox tree, 2002; goodwin, 1981; walker et al., 2014). spontaneous phenomena generally vary across conversational settings. for example, the backchannel yeah, which is used to indicate a speaker should continue telling a story, is more commonly used in face to face communication compared to telepresence communication (fox tree et al., 2021) as another example, discourse markers including well, you know, and like were found more frequently in hyperpartisan communication compared to nonhyperpartisan communication (nguyen et al., 2022). additionally, discourse markers are used more frequently in tasks compared to casual conversations over the phone, but used less frequently in tasks compared to casual conversations over instant messaging (guydish et al., 2024). for these reasons, we believe that i don’t know, absolutely, totally, sorta and kinda may also vary across settings. one setting where a difference has been found is in different types of english for i don’t know. new zealand english speakers tended to use i don’t know to avoid disagreement more often than british english speakers, and british english speakers used i don’t know to avoid making a commitment more often than new zealand speakers (grant, 2010). in our study, we investigated english spoken in the same community (american english), but across different modalities (video chat, phone, text) and different types of conversations (task-related, chit-chat). we also examined potential differences in friendship status. there is also reason to think that friends and strangers may use negotiation words differently. within close social relationships (friends), it might be more acceptable to be vague, which would minimize the use of i don’t know as a marker of lacking knowledge (bristol & rossano, 2020), and increase the number of hedges (sorta, kinda). in contrast, casual acquaintances or strangers may prefer i don’t know over vague statements (bristol & rossano, 2020). friends and strangers might use boosters (absolutely, totally) to emphasize commitment, but it is possible that close social bonds influence use: if emphasis on commitment is considered rude, then friends may be more likely to use them as friends are willing to tolerate more rudeness from friends than strangers (gupta et al., 2007). collocations, or co-occurring words, may also vary across settings. i don’t know often cooccurs with discourse markers such as well, oh, i mean and you know (grant, 2010, see also diani, 2004, and aijmer, 2009). when used together with other markers, negotiation markers may add nguyen and fox tree 72 nuance to conversational negotiation and management. for example, previous researchers suggested that well and oh commonly appear before disagreements or before expressing hesitation and that i mean and you know increase tentativeness and distance the speaker from the opinion (grant, 2010). if a well is used to indicate an upcoming disagreement, a co-occuring i don’t know or hedge might help distance the speaker from the upcoming disagreement in order to soften the impact of the disagreement. 1.4 current study while a lot of work has been done on i don’t know, absolutely, totally, kinda and sorta, their uses have not been studied across conversation modalities, types of conversations, and friendship statuses. word co-occurrences for boosters and i don’t know have also not been explored. co-occurrences have been explored for the hedges sort of and kind of/kinda in the british national corpus (gries & david, 2007). the british national corpus is composed of 90% written language and 10% transcripts of spoken language produced between 1960 and 1994, with material from 1960 to 1975 being written only (burnage & baguley, 1996). some of the conclusions were that kind of is more common in writing and sort of in speaking, that kind of co-occurs with verbs but sort of co-occurs with nouns, that kind of co-occurs with emotional states and sort of with colors, that kind of co-occurs with negative words like depressing more often than sort of but that both occur frequently with positive words like fun, and that sort of modifies quantities and people. in the current report, we tested how i don’t know varied across medium, relationship, and task type, including how it was used in comparison to common boosters and hedges. we also compared uses across our chosen hedges (kinda and sorta) and boosters (totally and absolutely). there are many ways to use these words that go against conventional thought. without examining what people say while communicating, it might seem that the subject and verb phrase i don’t know would either stand alone or occur at the beginning of a sentence (such as before that or whether). but examples of i don’t know being used in other ways can be found in spontaneously produced speech. in the following, i don’t know occurs within a prepositional phrase: (1) “we got home like really really late like at at like i don’t know like 2:30 or 3:00” (spoken example in fox tree, 2006, p. 739) in the following it occurs within a modifying phrase: (2) “he pulls out an ant house um with all this ant furniture in it and stuff, a little ladder for the ant to comclimb up to and ring a bell, and little justi don’t know, kinda like what you’d see in a gerbil cage i guess just for ants” (spoken example in fox tree & schrock, 1999, p. 285; the example is from a corpus collected by herbert clark of speakers retelling a monty python sketch) in the following it occurs in a verb phrase: (3) “i’ve been trying to like i don’t know understand the differences between northern california and southern california” (spoken example from roommates) in the following i don’t know and i dunno occur within adjectival phrases: (4) “the giant sculpture that i’m looking at is not small or gumdrop like i mean it looks like, i don’t know uh oh a pendulum that’s inside a giant grandfather clock” (spoken example from artwalk) (5) “it make a noise exactlyit sounds exactly like i dunno like a dog or a horse” (spoken example adapted from overstreet, 2005, p. 1865) we can also see i dunno open a turn without being followed by a subordinate clause (such as a clause beginning with that or whether): (6) “i dunno it’s very abstract, like the bottom part” (spoken example from artwalk) i don’t know, boosters, and hedges 73 on the flip side, a modifier like kinda might seem like it should only be found preceding what it modifies in the middle of a sentence. that is, a sentence that starts or ends with kinda would not be produced. but examples of such uses can be found in spontaneously produced communication. in the following two examples, kinda appears at the beginning and sort of at the end: (7) “sometimes i have to like, force it out of her. (laughs) so like, are you mad at me!? i know you’re mad at me! (laughs) kinda like that” (spoken example adapted from thorne et al., 2009, p. 640) (8) “it’s kind of walking up stairs to the left sort of “ (spoken example in fox tree & mayer, 2008, p. 167) (9) “okay yeah i think i see, the cut out corner is like wobbly cut kinda?” (texted example from im) the following ungrammatical structure has both i dunno and sort of, where i dunno comes after an adjective in what appears to be a noun phrase, and sort of appears to begin a new phrase: (10) “on the top is kind of this like big i dunno sort of has a rectangle on its left side (spoken example in fox tree & mayer, 2008, p. 174) by not defining phenomena of interest by conventional syntactic form (e.g., i don’t know is a sentence with a noun and a verb and kinda is a modifier), we allow for productions that violate conventional expectations about grammaticality and acceptability. as the examples provided here show, in everyday language use people do violate syntactic expectations. we predicted that uses will vary across task contexts. in the storytelling chit-chat context, we predicted that i don’t know would be more likely to mean “i don’t want to talk about it” than “i don’t have the information” because people are offering their own stories, which, presumably, they could simply not express instead of stating that they have no information about them. in the task contexts, in contrast, we predicted that i don’t know will be more likely to be used to mean “i don’t have information.” this would be observed in the on-task portions of the artwalk and im corpora. in object-identification tasks, people literally need to express that they don’t know what an object is. that is, the tasks afford the “i don’t have access to the information” meaning over the “i don’t want to talk about it” meaning. in both artwalk chit-chat and im chit-chat, we predict that i don’t know will mean “i am marking uncertainty,” where a participant avoids answering or alternatively, knows the answer and gives it but willingly distances themselves from it. in roommates (audiovisual, face to face), we predict that there will be fewer expressions of i don’t know than in artwalk chit-chat. we also predict that i don’t know will serve the same purpose as in the chit-chat — it will mean “i am marking uncertainty” more often than any other meaning, due to the ostensible use of i don’t know. thus, in roommates, which is a conversational task that began with storytelling, it would be conversationally odd for a speaker to mark their own uncertainty about a story that happened to them via i don’t know. we also predict that uses will vary based on communicative modality. we predict fewer numbers of i don’t know/idk in texted communication compared to spoken communication. in the on-task communication, we predicted fewer texted expressions i don’t know/idk compared to spoken i don’t know. there is a start-up cost to sending a message, so people are more likely to send informative messages rather than messages like i don’t know/idk (unless they truly do not know!). so, while we predicted that i don’t know/idk would more often mean lack of knowledge in the texted task-based communication, we predicted fewer numbers of i don’t know/idk overall in text communication. we also predicted that there will be more i don’t knows in task-based artwalk compared to task-based im due to the start-up costs associated with typing compared to speaking. in both corpora, in the task-based portions, we predict that i don’t know will more often mean “i don’t have information.” of the non-texted modalities, we predicted that there will be more i don’t know/i dunnos in face to face communication than phone communication because addressees will be better able to indicate their lack of knowledge or willingness to answer without seeming rude when they have nguyen and fox tree 74 additional facial cues. that is, in a conversation where the speakers have access to non-verbal visual cues, using an expression with multiple meanings like i don’t know can be unambiguous because of the additional cues. alternatively, with only voice cues, i don’t know/i dunno may be interpreted as a rude unwillingness to answer rather than an ostensible distancing from an answer. for artwalk chit-chat and task, where we are able to look at conversations between friends and strangers, we predict that friends will use i don’t know/i dunno to mean “i am marking uncertainty” and strangers will use it to mean “i don’t have information.” see table 1 for a list of predictions. artwalk task-basked more i don’t knows than in artwalk chit-chat more i don’t knows than in im task i don’t know means “i don’t have information” artwalk chit-chat i don’t know means “i am marking uncertainty” im task based more i don’t knows than in im chit-chat i don’t know means “i don’t have information” im chit chat i don’t know means “i am marking uncertainty” roommates more i don’t knows than in artwalk chit-chat i don’t know means “i am marking uncertainty” friends i don’t know means “i don’t want to talk about it” strangers i don’t know means “i don’t know” table 1. predictions. 2 method i don’t know and its variants (i dunno, idk) were analyzed across three corpora for their quantity and function. comparison expressions were analyzed (absolutely, totally, sorta, kinda) for their occurrence across corpora. 2.1 corpora assessed use of i don’t know and the four comparison markers were assessed in three corpora: (1) the roommates corpus of in person face to face communication (bryant, 2010), (2) the artwalk corpus of audio-only communication (liu et al., 2016), and (3) the instant messaging corpus (henceforth the im corpus) of texted communication (guydish & fox tree, 2022). please see table 2 for a breakdown of the corpora. corpus word count dyad type conversation type modality artwalk 236,629 friends / strangers task / chit-chat audio-only im 82,691 strangers task / chit-chat text-only i don’t know, boosters, and hedges 75 roommates 58,441 strangers storytelling chitchat face to face table 2. word count, dyad type, conversation type, and modality type for the three corpora assessed. the roommates corpus consisted of in-person spoken communication where participants were asked to come into the lab and have a conversation (see bryant, 2010, for a full description). the corpus was collected at a west coast research university. there were 22 transcripts of different dyads in the roommates corpus. participants were initially asked to tell a story about a roommate, and the conversation was allowed to continue from there, a task we will refer to as storytelling chitchat. because this was an open-ended task, there is no defined measurement of success, as there is in referential tasks. however, the dialogue was prompted by a task in a laboratory setting. this type of spontaneous conversation resembles both task and non-task related communication with respect to other conversational phenomena, such as backchannels (see nguyen et al., 2024 for more). we predict that storytelling chit-chat will more closely resemble naturalistic conversation (the chat portions of the artwalk and im corpora) than task conversation (the task portions of the artwalk and im corpora) for i don’t know. artwalk is a corpus of conversations between dyads via telephone (liu et al., 2016). there are 69 transcripts total; 48 are part of the original corpus and are publicly available (liu et al., 2016), and 21 are part of an expansion (guydish et al., 2021). one participant was in the lab (the director) and the other was in a west coast downtown location with public art (the follower). the director was responsible for giving the follower directions to different artworks installed in the location. conversational contributions related to finding the artworks comprises the task-oriented portion of the corpus. artwalk also, importantly, has spontaneously generated off-task conversation as well, the chat or chit-chat portion of the corpus (see nguyen et al., 2024, for an argument as to why these sections of conversation should be considered naturalistic). out of the 69 dyads, artwalk has 47 dyads with information recorded about the relationship between the people in the dyad; of these 47, 25 dyads were composed of friends and 22 dyads were composed of strangers. like artwalk, the im corpus has both on-task and off-task sections (see guydish & fox tree, 2022). there are 65 transcripts in the im corpus. this corpus was collected at a west coast research institution. pairs of dyads engaged in a referential task card-matching activity where they had to identify abstract shapes together. participants also had sections where they could have conversations unrelated to the task. these conversations are analogous to the chat portions of artwalk. all conversations in the im corpus took place between strangers. 2.2 coding i don’t know, i dunno, idk, absolutely, totally, sorta, and kinda and their surrounding contexts were extracted using an automated method. extracted information was double checked by a research assistant. the assistant ensured that the correct items of interest were pulled and also ensured that the surrounding context was obtained. occurrence (how often the phrase occurs) and co-occurrence (what it occurs with) rates were counted. functions were coded using an adapted schema based on grant (2010). all coding was done by two research assistants, with a third coder brought in to resolve disagreements. location and co-location information was also coded for these words. utterances were looked at individually, but the phrases of interest were coded in the context of the utterance they were spoken in. research assistants were trained on definitions developed and refined over several iterations of coding (see table 3 for the definitions the ras were provided), and then coded all expressions of nguyen and fox tree 76 interest individually for occurrence, co-occurrence, location, and co-location. irrs were calculated for all coding and interpreted using the ranges proposed by landis and koch (1977). grant’s (2010) coding scheme was adapted to code the expressions of interest. in addition to grant’s original categories, four more were added: prefacing agreements, expressing agreement, highlighting commitment to the answer, and maximizing compliments. because one suggested use of absolutely and totally is that they increase commitment, we added these four categories to capture these uses. there is no “seeking assessment” (the opposite of avoiding assessment) because assessment is focused on evaluating the interlocutor’s contributions, not one’s own. this means that “seeking assessment” would look like agreement or disagreement, both behaviors captured under other codes. there is no “seeking commitment” (the opposite of avoiding commitment) because “seeking commitment” can be interpreted as simply making a statement (e.g., “he’s kinda tall” is avoiding commitment, but “he’s tall” is making a statement). that is, speakers are assumed to commit as much as they can to every utterance, and it is only by marking it in some sense that commitment is lowered. in contrast, “highlighting commitment” can be interpreted as adding extra commitment, which isn’t the counterpart to avoiding commitment (cf., “he’s absolutely tall”). please see table 3 for a synopsis of the coding scheme. coding categories from grant (2010) are listed first, and the newly developed categories are listed after. category definition example grant (2010) inability to provide information / insufficient knowledge the expression is used when the speaker has a lack of knowledge “yeah it woit won’t let me take the picture right now i don’t know what to do” (artwalk corpus) prefacing disagreement the expression is used to manage the social relationship while disagreeing with the speaker “[laugh] it’s kinda not i uh i mean it think it's just supposed to be just suggestive and amorphous i don't i don't know i’m it could be just me” (artwalk corpus) avoiding disagreement the expression is used when the speaker wants to avoid giving a negative response “oh i totally listen . didn’t i remember that you live off campus[?]” (roommates corpus) avoiding assessment the expression is used to avoid judging the truth of their interlocutor’s statements d: yeah they can represent people, i don't know but uh they are very simplified. you only see the oval shape avoiding commitment the expression is stressing the speaker’s lack of confidence of the truth of the utterance “d: uhh it looks kinda like uhm little brown and yellowish? i think” (artwalk corpus) minimizing compliment the expression is downplaying the speaker’s confidence in a [no examples in corpora] i don’t know, boosters, and hedges 77 compliment (from the interlocutor) marking uncertainty the expression is stressing the uncertainty of the utterance “yeah but um the bottom of it its [sic] kinda like the stone texture and everything” (artwalk corpus) unclear / missing data the data can’t be coded due to an inability to make a judgment (lack of context, etc) “i kinda wan” (artwalk corpus) new addition to coding scheme prefacing agreement the expression is used to manage the social relationship while agreeing with the speaker “yeah kinda yeah” (artwalk corpus) expressing agreement the expression is used when the speaker wants to give a response in agreement “yeah, *totally!*” (roommates corpus) highlighting commitment to the answer the expression is emphasizing the speaker’s confidence in an expressed compliment “there was absolutely no drama at all and then to *go from that to*” (roommates corpus) maximizing compliment the expression is emphasizing the speaker’s confidence in an expressed compliment (from the interlocutor) [no examples in corpora] table 3. coding scheme for analysis with the marker of interest bolded. the first 8 categories are from grant (2010). the bottom 4 categories are created for this analysis. because location can be a cue for pragmatic function (aijmer, 2009), location was also coded. start of utterance indicated the word/phrase started the utterance or was a standalone phrase. end of utterance indicated the word/phrase ended the utterance, or appeared within the last sentence of the utterance. middle of utterance was selected for all other locations. co-locators were coded by counting the bigrams immediately to the left and right of the word/phrase. nguyen and fox tree 78 3 results results for each corpus are presented below. location information is presented across corpora in table 4. following that, cross-corpora comparisons are reported. 3.1 roommates all 21 transcripts in the roommates corpus were coded following the procedure laid out above. due to the difficult nature of the task, interrater reliability across the entire corpus was fair (fleiss’s kappa = .32). 3.1.1 hedges and boosters of the four words besides i don’t know, only totally and kinda appeared more than a handful of times. about 40% of the uses of kinda was to mark uncertainty, and about 30% were to avoid commitment (see table 5 in section 3.4 for uses). the most common words to precede kinda were it and i; the most common word to follow was like (see table # for collocations). about 60% of the uses of totally were highlighting commitment to the answer. like was the most common word before totally (“that’s like totally the way [she] is”). about half the words after totally were verbs. absolutely appeared once and sorta appeared twice . 3.1.2 i don’t know i don’t know occurred 219 times in the roommates corpus. i dunno and idk appeared zero times. i don’t know was most commonly coded as marking uncertainty (56 occurrences) followed by indicating insufficient knowledge (33 occurrences). i don’t know was difficult to code because it was often used as a stand-alone particle, resulting in a code of ( missing information (23 occurrences). the rest of the data was split between the rest of the coding scheme. in roommates, people told each other stories and chatted. they may have been unsure of how much their interlocutor might agree, distancing themselves from their statements with i don’t know. all 219 appearances of i don’t know were coded for co-located words. well, oh, and i mean – words that had previously co-occurred with i don’t know in grant (2010) – did appear with i don’t know in the roommates corpus. well appeared twice, once following i don’t know and once before. oh appeared 4 times, twice before i don’t know and twice after i don’t know. i mean appeared 4 times, 3 times before i don’t know and once after. you mean was not present. other common co-appearing words included like, which appeared 32 times before i don’t know (“like i don’t know”) and 22 times after (“i don’t know like this girl [name]”), as well as whwords, which appeared three times before, and 19 times after (“i don’t know why”). phrases like i think, i feel, and i guess also appeared before and after i don’t know. these are phrases that express a thought or opinion, so it makes sense that a speaker using them might qualify them with i don’t know, and also highlight the function i don’t know plays in softening the commitment to an upcoming statement. 3.2 artwalk all 59 transcripts in the artwalk corpus were coded following the procedure laid out above. inter-rater reliability was moderate (fleiss’ kappa = .55). below, we discuss the hedges and boosters (absolutely, totally, sorta, kinda) first, followed by i don’t know. we then discuss marker use by friendship status. 3.2.1 hedges and boosters like in roommates, only kinda and totally appeared in large numbers. kinda was the most frequent marker used in artwalk (306 occurrences). the other three markers were used five times less frequently (61 occurrences). kinda was much more frequent in task-related communication i don’t know, boosters, and hedges 79 (89% of use cases) compared to chit-chat (11% of use cases). totally appeared 41 times in the artwalk corpus, with a roughly even split between task-related (54% of uses) and chit-chat conversation (46% of uses). sorta was rare in the corpus, appearing 13 times. it appeared 11 times in task-related conversation, and twice in chit-chat. absolutely was also rare in the corpus, appearing 7 times in the corpus, with 71% of those uses in task-related conversation and 29% in chit-chat. please see table 5 in section 3.4 for a detailed breakdown. 3.2.2 i don’t know i don’t know appeared 168 times in the artwalk corpus, 138 times in task-related conversation, and 30 times in chit-chat conversation. all instances of i don’t know were coded for co-located words. in task-related conversation, 40 uses of i don’t know were coded as indicating insufficient knowledge, 28 were coded as marking uncertainty, 7 were coded as avoiding commitment or avoiding disagreement, and 8 were coded as missing or insufficient data. i don’t know appeared as a reduplication multiple times, with 7 sequences of an i don’t know followed by another i don’t know. i don’t know was preceded by um or uh a total of 11 times (“um, i don’t know”). a common co-locator with i don’t know was if, following i don’t know 28 times (“i don’t know if”). whwords were also common after i don’t know, appearing 17 times. how also appeared 11 times after i don’t know. i think, i feel, and other expressions of thought, feeling, or belief did occur after i don’t know, which is in line with i don’t know serving as a marker of uncertainty (“i don’t know, i think we’re doing a good job”). in chit-chat conversations, 10 uses were coded as indicating insufficient knowledge, 5 as marking uncertainty, and the rest were split between highlighting commitment, prefacing disagreement, and missing or insufficient data. i don’t know was preceded by but 4 times (“but i don’t know”), yeah 3 times, and um and like 1 time each. i don’t know was followed by if 4 times, whwords 3 times, like 2 times, and how 1 time. i dunno is a spoken, reduced form of i don’t know. because artwalk is a spoken corpus, i dunno was examined for usage. i dunno occurred 136 times in the artwalk corpus, 100 times in task-related conversation and 36 times in chit-chat conversation. all instances of i dunno were coded for co-located words. in the task-related conversation, 27 uses were coded as marking uncertainty, 17 uses were coded as indicating insufficient data, and the rest of the uses were split between avoiding disagreements and avoiding commitment. before i dunno, like was the most common word, with 11 uses (“like i dunno”). other common words include uh/um (9 occurrences), it (3 uses) and well (2 uses; for example, “well i dunno”). you know occurred once. after i dunno, whwords were very common (15 uses), followed by how (8 uses), if (7 uses), and like (6 uses). there were 7 occurrences of expressions like i guess occurring after i dunno, highlighting i dunno’s use as a marker of knowledge state (“i dunno, i guess”). in the chit-chat conversation, i dunno was coded as marking uncertainty 14 times and as indicating insufficient data 13 times. the rest of the data was split between avoiding commitment, avoiding disagreement, and prefacing disagreement. looking at words that co-locate to the left of i dunno, uh preceded i dunno 4 times (“uh i dunno”), but, so, and i/i’m appeared 3 times each, and yeah appeared twice. common words that followed i dunno were i/i’m (8 times as in “i dunno, i have a bunch of loans”), if (4 times), whwords (3 times), and but (2 times). 3.2.3 friends and strangers for the friends and strangers analysis, 12 transcripts were excluded because they did not have information on whether the participants were friends or strangers, resulting in 47 transcripts used in the analysis. of those 47 remaining transcripts, 25 were friend dyads and 22 were stranger dyads. we examined the use of i don’t know across task-related and chit-chat conversation. two nguyen and fox tree 80 different things were analyzed – the number of occurrences of i don’t know / i dunno and the use of i don’t know/i dunno. starting with numbers used, there was no difference in the number of combined i don’t know and i dunnos between friends and strangers across both chit-chat and task related conversation, t(44.96) = .89, p = .37. there was no difference in the number of combined i don’t know and i dunnos in chit-chat related conversations (t(34.40) = -.83, p = .41) or in task related conversation (t(44.57) = 1.47, p = .15). breaking apart i don’t know and i dunno, there were no differences in the number of i don’t knows used by friends and strangers in either chit-chat (t(21.98) = -1.40, p = .17). or task-related conversation (t(41.547) = -.31, p = .76). there was also no difference in the number of i dunnos used in chit-chat by strangers and friends (t(42.62) = .53, p = .60). however, there was a difference in the number of i dunnos used in task related conversation, with more i dunnos in friend dyads working on tasks compared to dyads of strangers (t(25.45) = 2.37, p = .03). to look at uses, we identified the two most common uses of i dunno and i don’t know, which were marking uncertainty and inability to provide requested information. we were particularly interested in how usage frequency (how many times a marker of uncertainty or a literal marker of non-information was used) was affected by conversation type and relationship. the first analysis collapsed across i don’t know and i dunno for the i don’t have the information use. within this use, there was no significant difference in how friends and strangers used these words across conversation types, x2 (1, n = 69) = .357, p = .550. breaking apart i don’t know and i dunno, for the i don’t have the information uses, there was no difference in how friends and strangers used i don’t know across conversation types (x2 (1, n = 45) = 3.794, p = .05). however, for the i don’t have the information uses, i dunno did differ in how it was used across relationship and conversation types, with i dunno being more likely to be used among friends in task-settings (x2 (1, n = 24) = 5.71, p = .017). the second analysis collapsed across i don’t know and i dunno for the marking uncertainty use.there was no difference in usage across relationship or conversation type, (x2 (1, n = 67) = .45, p = .50). breaking apart i don’t know and i dunno, there was also no difference in frequency of marking uncertainty use across relationship or conversation type, (x2 (1, n = 36) = .789, p = .387). there was also no difference in usage frequency for i dunno across relationship or conversation type, (x2 (1, n = 31) = .003, p = .959). looking within chit-chat conversation, there was no difference in how i don’t know/i dunno was used across relationships, (x2 (1, n = 58) = .395, p = .530). looking at the individual forms, there was again no difference in usage frequencies across relationships for i don’t know, (x2 (1, n = 35) = .412, p = .521), and for i dunno (x2 (1, n = 23) = .059, p = .809). looking within task-based conversation, there was no difference in how i don’t know/i dunno was used across relationships, (x2 (1, n = 78) = .423, p = .515). for i don’t know, there was no difference in how it was used across relationships, (x2 (1, n = 46) = 3.05, p = .08). for i dunno, there was a significant difference in how it was used across relationships, with both the i don’t have the information and marking uncertainty uses more likely among friends than strangers. 3.3 im all 65 transcripts in the im corpus were coded following the procedure laid out above. due to the difficult nature of the task, interrater reliability across the entire corpus was fair (fleiss’s kappa = .28). 3.3.1 hedges and boosters kinda was, again, the most frequent of the discourse markers used in the corpus (213 occurrences). taken together, the other markers – absolutely, totally, and sorta – were used six times less frequently (35 occurrences). kinda appeared 213 times in the im corpus, with about i don’t know, boosters, and hedges 81 70% of the uses in task-related conversation and approximately 30% of the uses in chit-chat. totally appeared 17 times in the im corpus, with most of its uses in chit-chat (about 70% of the uses). sorta appeared 14 times in the im corpus, with 70% of the uses in task-related conversation.. absolutely was rare, appearing 4 times in the im corpus, with all uses in the chit-chat portion. see table 5 in section 3.4 for a detailed breakdown. 3.3.2 i don’t know i don’t know appeared 6 times, and all six occurrences were in chit-chat. three were used to indicate insufficient knowledge and three were coded as “i don’t want to say.” i don’t know was followed by how much twice, why twice, too much once, and how many once. there were two occurrences of i dunno, both in chit-chat conversations and both at the start of an utterance (“i dunno”). both of them were coded as missing or insufficient data. i dunno appeared with lol once. because this was a text-based conversation, a written, abbreviated form of i don’t know (idk) was also analyzed. idk appeared 71 times, 7 times in task-related conversation and 64 times in chit-chat related conversation. in the task-related conversation, 3 uses of idk were coded as indicating insufficient knowledge, 3 were coded as marking uncertainty, and 1 was coded as avoiding commitment. in the chit-chat based conversation, there were 64 appearances of idk. the most common use was indicated insufficient knowledge, with 35 occurrences. the next most common was avoiding commitment, with 13 instances. the rest of the data was split between missing or unable to code (8), marking uncertainty uses (7), and avoiding assessment (1). the most common word to precede idk was but, with 6 instances. other words that preceded idk include although, because, oh and yeah. the most common word to follow idk was if, with 13 occurrences (“idk if we can go back”). how also followed idk at a high rate (10) and so did whphrases (7). kinda totally sorta absolutely i don’t know i dunno idk roommates start: 21 middle: 101 end: 34 start: 18 middle: 25 end: 9 start: 2 middle: 0 end: 0 start: 0 middle: 1 end: 0 start: 87 middle: 104 end: 29 none none artwalk task start: 41 middle: 173 end: 57 start: 4 middle: 12 end: 6 start: 0 middle:10 end: 1 start: 1 middle: 1 end: 3 start: 49 middle: 68 end: 10 start: 45 middle: 43 end: 12 none artwalk chat start: 9 middle: 16 end: 10 start: 2 middle: 12 end: 5 start: 2 middle: 0 end: 0 start: 1 middle: 1 end: 0 start: 16 middle: 11 end: 3 start: 14 middle: 15 end: 7 none im task start: 47 middle: 72 end: 27 start: 2 middle: 1 end: 2 start: 3 middle: 4 end: 3 none none none start: 5 middle: 2 end: 0 im chat nguyen and fox tree 82 start: 23 middle: 33 end: 12 start: 6 middle: 5 end: 1 start: 2 middle: 2 end: 0 start: 1 middle: 3 end: 0 start: 5 middle: 1 end: 0 start: 2 middle: 0 end: 0 start: 47 middle: 24 end: 2 table 4. locations of marker 3.4 cross-corpora comparisons in table 5, the uses of hedges and boosters across corpora are presented. results are discussed for each corpus in their respective sections, above. corpus word use type # of uses top use most common preceding word most common following word/part of speech roommates absolutely n/a 1 intensify commitment was no sorta n/a 2 avoid commitment, avoid assessment i’m, was like, a totally n/a 52 highlighting commitment to the answer like verbs kinda n/a 156 marking uncertainty it’s like artwalk absolutely task 5 express agreement yeah no absolutely chit-chat 2 highlight commitment, perform agreement i have, i know no, its sorta task 11 marking uncertainty it’s verb sorta chit-chat 2 marking uncertainty i, it want, weird totally task 22 highlighting commitment i verbs totally chit-chat 19 highlighting commitment i verbs i don’t know, boosters, and hedges 83 kinda task 271 marking uncertainty its like kinda chit-chat 35 highlighting commitment its verbs, adjectives im absolutely task 0 absolutely chit-chat 4 highlighting commitment, expressing agreement yeah, we have, but, breathtaking sorta task 10 marking uncertainty looks like sorta chit-chat 4 marking uncertainty it, well, can, detective verbs totally task 5 highlighting commitment its adjectives totally chit-chat 12 highlighting commitment i verb kinda task 146 marking uncertainty its look/looks/looked kinda chit-chat 68 highlighting commitment its adjectives table 5. hedges and boosters across corpora to do statistical comparisons across corpora, raw counts were converted into percentages, where the number of instances was divided by the number of words in the transcript. we predicted that there would be more i don’t knows in task-based portions of a conversation than in chit-chat portions. in artwalk, there were more i don’t knows in task-based conversation, t(58) = 2.55, p = .01. i dunno was also used significantly more in artwalk task-based conversation than chit-chat, t(58) = 2.70, p < .001. in im, in contrast, we found no evidence of a difference in the rate of i don’t knows in task-based portions and chit-chat portions. however, idk was used more in chit-chat, t(64) = 2.69, p = .009. across artwalk and im task-related conversation, we found no evidence of a difference in the rate of i don’t knows, t(71.1) = 1.77, p = .07. as i don’t know appeared only 138 times across all the task-related conversations, or 0.0006 %, in artwalk, and i don’t know appeared no times in task-related conversations, this result is unsurprising. combining the forms of i don’t know used in artwalk task-based conversations (i don’t know and i dunno) and the forms used in im task-based conversations (i don’t know and idk), there are more overall in im, t(118.77) = 2.89, p = .004. nguyen and fox tree 84 across artwalk chit-chat and roommates conversations, there was a significant difference in the number of i don’t knows, with more in roommates, t(99.7) = 2.34, p = .02. artwalk task-basked more i don’t knows than in artwalk chit-chat. no difference in i don’t knows compared to im task. i don’t know primarily means “i don’t have information.” i dunno primarily means “i don’t have information.” artwalk chit-chat i don’t know primarily means “i don’t have information.” i dunno means “i am marking uncertainty” and “i don’t have information.” im task based no appearances of i don’t know. idk primarily means “i don’t have information.” im chit chat i don’t know means “i am marking uncertainty” and “i don’t have information.” idk primarily means “i don’t have information.” more idk in chat. roommates more i don’t knows than in artwalk chit-chat. i don’t know is used primarily to mark uncertainty. no uses of idk / i dunno. friends no difference in use of i don’t knows. more i dunnos in chit-chat and task-related conversation compared to strangers. strangers no difference in use of i don’t know. no difference in use of i dunno. table 6. summary of results 4 discussion when negotiating common ground, speakers are deliberate about what they want to signal to their interlocutors. saying “it’s totally fun” is different from saying “it’s sorta fun” which is different from saying “it’s fun, i don’t know.” the negotiation words qualify the level of fun and the speaker’s commitment to what they are saying. we found that communicators use negotiation words differently depending on modality, goals, and relationships. 4.1 i don’t know i don’t know (and the spoken and written forms i dunno and idk) were primarily used to indicate a lack of knowledge or inability to provide the requested information. across task-related conversations, we found no evidence of differences in the frequency of i don’t know across audioonly and text-only conversations. however, when looking at all the forms people could use, there were more i don’t knows and associated forms (in this case idk) in the im corpus. there were more idks than i don’t knows, so this is likely due to the low start up cost of typing idk versus the start up cost of saying i don’t know or i dunno. additionally, i dunno and idk have different meanings, with i dunno being used to mark uncertainty more often than idk (see table 6). i don’t know, boosters, and hedges 85 in spoken chit-chat conversations (roommates, artwalk chit-chat), there were more i don’t knows used in roommates, in line with our predictions. one explanation for this is that i don’t know might have been less vague in a situation where people are able to access cues beyond the voice. this is supported by the usage of i don’t know – most of the uses in roommates were marking uncertainty uses, and most of the uses in artwalk chit-chat were i don’t have the information uses. another reason there might be more i don’t knows in roommates compared to artwalk chit-chat is that roommates may function more like a task rather than true spontaneous communication. while it was focused on telling stories and chatting, participants were still brought into the lab and given instructions, whereas with artwalk, participants freely chose to engage in chit-chat without prompting (see nguyen et al., 2024] for similar behavior in backchannels). a third explanation is that the different uses of i don’t know reflect differences between storytelling chit-chat and problem solving in communication, rather than differences between modalities (in person versus audio only) or differences in motivations (asked-to-chat versus chose-to-chat). one interesting finding is that speakers in the im corpus did not use i don’t know or idk very often in the task-related portion of the conversation. when communicating in a medium that is lacking other cues a speaker might use to indicate uncertainty, it seems that speakers prefer to avoid expressing a lack of knowledge. while we did not look at what speakers actually did, it seems likely that rather than saying i don’t know, interlocutors asked questions or expressed that they didn’t understand their partner directly (e.g., “i can’t find what you’re talking about”). relationships between speakers did affect use of i dunno. friends overall used i dunno in tasks more often, and to mean both i don’t have the information and i am marking uncertainty. friends might use i dunno in tasks more often than strangers because of the shared common ground between them; a friend who says i dunno during a difficult task is probably perceived as less rude than a stranger who says i dunno. because speakers must negotiate on meaning while retaining politeness and limiting misunderstandings, it is safer to use i dunno with someone who already understands what you might mean by that. i don’t know, i dunno, and idk show different patterns for where people use them in utterances. i don’t know appeared more in the middle of an utterance, but could also appear at the start of an utterance across the three corpora. in contrast, i dunno appeared at almost equal rates at the start and middle of an utterance across the three corpora. while all three are less likely to be used at the end of an utterance, suggesting they are less likely to be used as a floor-yielding marker (aijmer, 2009), idk is also less likely to be used in the middle of an utterance compared to the start of an utterance across the three corpora. we found that idk rarely performed anything beyond the literal meaning, and tended to appear at the beginning of utterances. kinda, one of the two common hedges, was more likely to appear in the middle of an utterance than the start or end, though it did appear in all three positions across corpora. absolutely (a common booster), and sorta (the other common hedge) were rare, but when they did appear, it was across all three positions. totally (the other common booster) tended to appear in the middle of an utterance rather than at the start or end. it appears that idk is the form of i don’t know that has become lexicalized to mean “i don’t have the information” in instant messaging communication. this process might be similar to what has happened to laughing out loud (lol), which, like idk, started out as a full phrase and shortened to an acronym, and has evolved, like idk, past the original meaning (herring, 2012; tagliamonte & denis, 2008;varnhagen et al., 2010). this may be the result of pressures of technology: while typing on phones has gotten easier with the introduction of the full keyboard, it is still difficult to do. likewise, there are time costs to typing on a keyboard on a computer (people can say more than they can type in a given amount of time), and so people who are using mediated communication might favor shorter messages as a result. this contrasts with i dunno, which acts pragmatically like i don’t know. as we move into a world where communication occurs online more than offline, we should expect to see more phrases lexicalize. providing a preliminary study of the location and functions nguyen and fox tree 86 of idk provides insight into how lexicalized phrases might be used, and when and how they can replace their spoken counterparts in online communication. our work provides a snapshot in time against which future uses can be compared. 4.2 hedges and boosters overall, there were very few boosters in the three corpora. across all three corpora, absolutely appeared 12 times and totally appeared 110 times. when absolutely did appear, it was used to either highlight commitment or to express agreement. thus, absolutely appears to function in the way that boosters are expected to – that is, it maximizes commitment to the utterance. the picture around totally is less clear. most uses of totally were to highlight commitment, and express agreement as expected from a booster. however, totally is able to perform a broader range of functions compared to absolutely. some uses of totally were to avoid disagreement or to preface agreement. this suggests that totally may function in a broader sense compared to canonical boosters — it has an expanded use that can cover more functions. of the hedges, sorta (29 times total) was rare in comparison to kinda (676 times total). looking at the uses of sorta, it was primarily used to mark uncertainty. other uses of sorta were to avoid commitment or assessment. in other words, sorta is a true hedge – it is used in cases where people are interested in distancing themselves from a statement, or when they are unsure of the information they are reporting. on the other hand, kinda has a broader set of uses. the most common uses of kinda were to mark uncertainty and to highlight commitment. kinda could also be used to avoid commitment, or express agreement, but these uses were rare in comparison. kinda and sorta are typically grouped together as a set of hedges that functionally mean the same thing, but these results suggest that kinda and sorta have different pragmatic functions and are not equivalent in use. 4.3 collocations regarding collocations, the common hedges (kinda, sorta) frequently appear before and after like and looks. the hedges can be interpreted as modifying like and looks to suggest “roughly” or “approximately.” in addition, weird and other adjectives (which are easily qualified by approximation) tended to follow common hedges. totally tended to appear after i and before verbs. this supports the idea that one of the main functions of totally is to highlight commitment. absolutely most often appears with agreement, supporting its function of indicating agreement. as for i don’t know, i dunno, and idk, all three appeared before whphrases (which, why), echoing previous findings on i don’t know and i dunno (aijmer, 2009). i dunno and i don’t know, both forms that are spoken aloud, often preceded expressions like i feel, suggesting that i don’t know can be used to mark uncertainty about the listener’s reception, not just the speaker’s beliefs. idk, which is typed, did not appear with expressions like i feel. it was, however, preceded by but and followed by if, which suggests a use in distancing the writer from the intentions conveyed. 5 conclusions negotiation words are often assumed to be equivalent and interchangeable when they are of the same type (the booster totally can substitute for absolutely; the hedge kinda can substitute for sorta). the evidence provided here shows instead that the words have different pragmatic functions and are not as interchangeable as previously thought. one notable finding is that i don’t know appears infrequently in texted conversations, with the shortened form idk more likely to appear, providing evidence of the lexicalization of idk into a carrier of meaning beyond the original i don’t know. another is that within the hedges and boosters, there are significant differences in how frequently kinda and totally are used compared to sorta and absolutely. while previous researchers have shown that these words cluster together on amounts of negotiation and correction i don’t know, boosters, and hedges 87 they imply (nguyen & fox tree, under review), the current evidence highlights key differences: totally and kinda are more general than absolutely and sorta. that is, hedges and boosters are precisely selected for their meanings and cannot be interchanged without altering the meaning the speaker is trying to convey. the findings reported here have broad-ranging applications. for example, there are significant implications for those learning english, whether they be new language learners or artificial agents. while one of the boosters and one of the hedges we looked at (absolutely, sorta) hewed closely to expected patterns, others (totally, kinda) had broader semantic ranges. in addition to needing additional information about these boosters and hedges, communicators would need to be taught or trained how to interpret the intended meaning of i don’t know, or to produce i don’t know in the appropriate way. there is a difference between wanting to convey a lack of knowledge versus wanting to convey a desire to keep information off record, among other uses. contextual information such as the conversational modality, the type of conversation, or the acquaintanceship status of communicators could also be explicitly taught or programmed. another implication applies to the broader understanding of negotiation words. i don’t know is not like boosters or hedges. it wasn’t used to intensify commitment to an answer or to express agreement, like boosters were. i don’t know is also more flexible than a hedge because it can convey a literal lack of information along with marking uncertainty. future work might include assessing the nonverbal cues associated with boosters, hedges, and i don’t know. for example, people might indicate uncertainty by pulling lips back with a grimace, certainty by furrowing brows, or lack of knowledge by raising shoulders and eyebrows. with respect to i don’t know, the bodily movements for indicating a desire to keep information off record may be different (perhaps a mouth movement instead of a shoulder shrug). examining these words in a larger context that not only includes turn by turn analyses but also body gestures would provide insight into how people metalinguistically mark the different uses of i don’t know. researchers might also consider assessing directly how boosters, hedges, and i don’t knows (and its variants) are interpreted across different settings. for example, the uses of i don’t know, boosters, and hedges may vary in legal contexts or among people of different ages. acknowledgements funding for this project was received from the university of california santa cruz social sciences division quarter dissertation fellowship. we thank the research assistants who assisted on this project. references karin aijmer. ”so er i just sort i dunno i think it’s just because. . . ”: a corpus study of i don’t know and dunno in learners’ spoken english. in corpora: pragmatics and discourse, pages 151–168. brill, 2009. andrea beltrama and florian schwarz. (im) precise personae: the effect of socio-indexical information on semantic interpretation. language in society, pages 1–28, 2024. susan e. brennan and maurice williams. the feeling of another’s knowing: prosody and filled pauses as cues to listeners about the metacognitive states of speakers. journal of memory and language, 34(3):383–398, 1995. rachel bristol and federico rossano. epistemic trespassing and disagreement. journal of memory and language, 110:104067, 2020. penelope brown and stephen c. levinson. politeness: some universals in language usage. nguyen and fox tree 88 number 4 in cambridge papers in linguistics. cambridge university press, 1987. gregory a. bryant. prosodic contrasts in ironic speech. discourse processes, 47(7):545–566, 2010. a. gavin burnage and glynis baguley. the british national corpus. library and information briefings, (65), february 1996. herbert h. clark and jean e. fox tree. using uh and um in spontaneous speaking. cognition, 84 (1):73–111, 2002. j. trevor d’arcey, shereen oraby, and jean e. fox tree. wait signals predict sarcasm in online debates. dialogue & discourse, 10(2):56–78, 2019. giuliana diani. the discourse functions of i don’t know in english conversation. in pragmatics and beyond new series, pages 157–172. 2004. simona pekarek doehler. more than an epistemic hedge: french je sais pas ‘i don’t know’as a resource for the sequential organization of turns and actions. journal of pragmatics, 106:148– 162, 2016. farahman farrokhi and safoora emami. hedges and boosters in academic writing: native vs. non native research articles in applied linguistics and engineering. journal of english language pedagogy and practice, 1(2):62–98, 2008. jean e. fox tree. placing like in telling stories. discourse studies, 8(6):723–743, 2006. jean e. fox tree. discourse markers across speakers and settings. language and linguistics compass, 4(5):269–281, 2010. jean e. fox tree. discourse markers in writing. discourse studies, 17(1):64–82, 2015. jean e. fox tree and sarah a. mayer. overhearing single and multiple perspectives. discourse processes, 45(2):160–179, 2008. jean e. fox tree and josef c. schrock. discourse markers in spontaneous speech: oh what a difference an oh makes. journal of memory and language, 40(2):280–295, 1999. jean e. fox tree, steve whittaker, susan c. herring, yasmin chowdhury, allison nguyen, and leila takayama. psychological distance in mobile telepresence. international journal of human computer studies, 151:102629, 2021. charles goodwin. conversational organization: interaction between speakers and hearers. 01 1981. lynn e. grant. a corpus comparison of the use of i don’t know by british and new zealand speakers. journal of pragmatics, 42(8):2282–2296, 2010. stefan th. gries and caroline v. david. this is kind of/sort of interesting: variation in hedging in english. studies in variation, contacts and change in english, 2, 2007. swati gupta, marilyn a. walker, and daniela m. romano. how rude are you? evaluating polite ness and affect in interaction. in proceedings of the 2nd international conference on affective computing and intelligent interaction, 2007. andrew j guydish and jean e fox tree. reciprocity in instant messaging conversations. language and speech, 65(2):404–417, 2022. andrew j guydish, j trevor d’arcey, and jean e fox tree. reciprocity in conversation. language and speech, 64(4):859–872, 2021. andrew j guydish, allison nguyen, and jean e fox tree. discourse markers in small talk and tasks. discourse studies, 26(5):606–620, 2024. janet holmes. hedges and boosters in women’s and men’s speech. language & communication, 10(3):185–205, 1990. ken hyland. persuasion and context: the pragmatics of academic metadiscourse. journal of pragmatics, 30(4):437–455, 1998. i don’t know, boosters, and hedges 89 ken hyland. hedges, boosters and lexical invisibility: noticing modifiers in academic texts. language awareness, 9(4):179–197, 2000. ellen a. isaacs and herbert h. clark. ostensible invitations. language in society, 19(4):493–509, 1990. jumayel islam, lu xiao, and robert e. mercer. a lexicon-based approach for detecting hedges in informal text. in proceedings of the twelfth language resources and evaluation conference, pages 3109–3113, 2020. alireza jalilifar and maryam alavi-nia. we are surprised; wasn’t iran disgraced there? a functional analysis of hedges and boosters in televised iranian and american presidential debates. discourse & communication, 6(2):135–161, 2012. andreas h. jucker and sara w. smith. explicit and implicit ways of enhancing common ground in conversations. pragmatics. quarterly publication of the international pragmatics association (ipra), 6(1):1–18, 1996. elise kärkkäinen. position and scope of epistemic phrases in planned and unplanned american english. in new approaches to hedging, pages 203–236. brill, 2010. hadi kashiha. on persuasive strategies: metadiscourse practices in political speeches. discourse and interaction, 15(1):77–100, 2022. j. richard landis and gary g koch. the measurement of observer agreement for categorical data. biometrics, 33(1): 159-174, 1977. kristen e. link and roger j. kreuz. the comprehension of ostensible speech acts. journal of language and social psychology, 24(3):227–251, 2005. kris liu and jean e. fox tree. hedges enhance memory but inhibit retelling. psychonomic bulletin & review, 19:892–898, 2012. charles metzing and susan e. brennan. when conceptual pacts are broken: partner-specific effects on the comprehension of referring expressions. journal of memory and language, 49(2):201– 213, 2003. jenny myrendal. negotiating meanings online: disagreements about word meaning in discussion forum communication. discourse studies, 21(3):317–339, 2019. islam namaziandost. a comparative study of boosters in the discussion section of medical and applied linguistics articles. international journal of applied linguistics and english literature, 6(7):1–5, 2017. allison nguyen and jean e. fox tree. is kinda less sure than partially?: putting negotiation words on a scale. under review. allison nguyen, tom roberts, pranav anand, and jean e. fox tree. look, dude: how hyperpar tisan and non-hyperpartisan speech differ in online commentary. discourse & society, 33(3): 371–390, 2022. allison nguyen, andrew j guydish, and jean e. fox tree. backchannels in the lab and in the wild. interaction studies, 25(1):70–99, 2024. tri nuraniwati and alfelia nugky permatasari. hedging in ted talks: a corpus-based pragmatic study. jeels (journal of english education and linguistics studies), 8(2):203–226, 2021. maryann overstreet. and stuff und so: investigating pragmatic expressions in english and german. journal of pragmatics, 37(11):1845–1864, 2005. heike pichler and ashley hesson. discourse-pragmatic variation across situations, varieties, ages: i don’t know in sociolinguistic and medical interviews. language & communication, 49:1– 18, 2016. françoise salager-meyer. hedges and textual communicative function in medical english written discourse. english for specific purposes, 13(2):149–170, 1994. nguyen and fox tree 90 joanne scheibman. i dunno: a usage-based account of the phonological reduction of don’t in american english conversation. journal of pragmatics, 32(1):105–124, 2000. solarmax. bff jill. url https://www.youtube.com/watch?v=4niucrjx9-o. aina garí soler, matthieu labeau, and chloe´ clavel. measuring lexico-semantic alignment in debates with contextualized word representations. in proceedings of the first workshop on social influence in conversations (sicon 2023), pages 50–63. association for computational linguistics, 2023. sali a. tagliamonte and derek denis. linguistic ruin? lol! instant messaging and teen language. american speech, 83(1):3–34, 2008. masahiro takimoto. a corpus-based analysis of hedges and boosters in english academic articles. indonesian journal of applied linguistics, 5(1):95–105, 2015. avril thorne, lauren shapiro, kim cardilla, neill korobov, and paul a nelson. caught in the act: how extraverted and introverted friends communally cope with being recorded. journal of research in personality, 43(4):634–642, 2009. jackson tolins and jean e. fox tree. addressee backchannels steer narrative development. journal of pragmatics, 70:152–164, 2014. amy bm tsui. the pragmatic functions of i don’t know. text-interdisciplinary journal for the study of discourse, 11(4):607–622, 1991. connie k. varnhagen, g. peggy mcfall, nicole pugh, lisa routledge, heather sumida macdonald, and trudy e. kwong. lol: new language and spelling in instant messaging. reading and writing, 23:719–733, 2010. esther j. walker, evan f. risko, and alan kingstone. fillers as signals: evidence from a question– answering paradigm. discourse processes, 51(3):264–286, 2014. ann weatherall. i don’t know as a prepositioned epistemic hedge. research on language & social interaction, 44(4):317–337, 2011. margaret zellers. an overview of forms, functions, and configurations of backchannels in ruru uli/lunyala. journal of pragmatics, 175:38–52, 2021. http://www.youtube.com/watch?v=4niucrjx9-o dialogue & discourse 15(1) (2024) 1–44 doi: 10.5210/dad.2024.101 digging communicative intentions: the case of crises events farah benamara farah.benamara@irit.fr irit, université de toulouse, cnrs, toulouse inp, ut3. france ipal, cnrs-nus-astar. singapore alda mari alda.mari@ens.fr ijn cnrs/ens/ehess/psl. france romain meunier romain.meunier@irit.fr irit, université de toulouse, cnrs, toulouse inp, ut3. france véronique moriceau veronique.moriceau@irit.fr irit, université de toulouse, cnrs, toulouse inp, ut3. france leila moudjari leila.moudjari@irit.fr irit, université de toulouse, cnrs, toulouse inp, ut3. france valentin tinarrage v.tinarrage@gmail.com ijn cnrs/ens/ehess/psl. france editor: junyi jessy li submitted 07/2023; accepted 05/2024; published online 05/2024 abstract in emergency situations users of social networks convey all sorts of what have been called communicative intentions, well-known since the work of austin (1962) and searle (1969) as speech acts (sa). while speech acts have been the focus of close scrutiny in the philosophical and linguistic literature (see (portner, 2018) for extended discussion), their role has been only rarely understood and exploited in processing social media content about crisis events, our focus here. current work on communicative intentions in social media are topic-oriented, focusing on the correlation between sa and specific topics such as crisis (e.g., earthquakes) but also politics, celebrities, cooking, travel, etc. it has been observed that people globally tend to react to natural disasters with sa distinct from those used in other contexts (e.g., celebrities, which are essentially made up of comments). here, we explore the further hypothesis of a correlation between different sa types and urgency and propose an in depth linguistic and computational analysis of communicative intentions in tweets from an urgency-oriented perspective. indeed, sa are mostly relevant to identify intentions, desires, plans and preferences towards action and to ultimately produce a system intended to help rescue teams. our contribution is four-fold and consists of: (1) a two-layer annotation scheme of speech acts both at the tweet and sub-tweet levels, (2) a new french dataset of about 13k tweets annotated for both urgency and sa, targeting both expected (e.g., storms) and unexpected or sudden (e.g., building collapse, explosion) events, (3) a thorough analysis of the annotations studying in particular the correlation between sa and the urgency of the message, sa and intentions to act categories (e.g., human damages), and sa and crisis types, finally, (4) a set of deep learning experiments to detect sa in crises related corpora. our results show a strong correlation between sa and urgency annotations at both the tweet and sub-tweet levels with a particular salient correlation in the latter case, which constitutes a first important step towards sa-aware nlp-based crisis management on social media. keywords: speech acts, crisis events, social media ©2024 benamara, mari et al. this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). 1. introduction 1.1 motivation in ordinary interaction as well as in social networks, speakers unveil a variety of communicative intentions among which, make content known, express their own views and opinions or enhance action. since austin (1962) and later and more prominently searle (1975), these communicative intentions are known under the term speech acts. before percolating into the computational literature, speech acts (henceforth sa) have been the object of extensive discussion in the philosophical and the linguistic communities ((hamblin, 1970; brandom, 1994; sadock, 2004; asher and lascarides, 2008; portner, 2018; bach and harnish, 1979) to mention just a few). according to the austinian initial view, sa are to achieve action rather than conveying information. when uttering i now baptize you, the priest accomplishes the action of baptizing rather than just stating a proposition. beyond these prototypical cases, the literature has quickly broadened the understanding of the notion of sa as a special type of linguistic object that encompasses questions, orders and assertions and transcends propositional content revealing communicative intentions on the part of the speaker (bach and harnish, 1979; gunlogson, 2008; asher and lascarides, 2008; giannakidou and mari, 2021c): with an assertion, the speaker intends to present the propositional content and to add it to the common ground (portner, 2018); with a question, the speaker asks the addressee to provide new information; with an order the speaker asks that the content be realized and with exclamatives, a subjective evaluation towards propositional content is conveyed. our study investigates the communicative intentions that sa conveys in urgency situations and more importantly, how intentions vary according to the degree of urgency of the information (urgent vs. not urgent vs. not useful – cf. examples below) when posted in social networks. we focus on messages posted on twitter as tweets are widely used to generate valuable information in crisis situations (reuter et al., 2018). for example, the notre dame fire that occurred in france has been the most used in twitter in 20191 and in the recent earthquake in turkey and syria, some victims trapped in the rubble have been saved thanks to the messages they posted (toraman et al., 2023). sa are particularly helpful in identifying urgent messages. these are messages that raise situational awareness over a crisis situation and some specific aspects that include human/infrastructure damages, security instructions, etc. they provide actionable information that will help human teams to set priorities and decide appropriate actions (vieweg et al., 2014; castillo, 2016; reuter and kaufhold, 2018). therefore, speaking subjects perform qualitatively very different language acts depending on the situation they find themselves in. they mostly aim to make interlocutors react (i.e., perlocutionary level) by different linguistic means (illocutionary level, this is the level at which the speech acts are encoded), in view of achieving a purpose.2 1.2 when communicative intentions reveal urgency by revealing speakers communicative intentions and aiming at triggering the addressee reaction, speech acts become essential in emergency situations where action is to be enhanced. we have thus used two different independent classifications: (i) a new, two level classification of speech acts, and (ii) an independent classification for urgency and actionability elaborated in kozlowski et al. (2020). 1. https://blog.twitter.com/en_us/topics/insights/2019/thishappened-in-2019 2. on perlocutionary / illocutionary, see (austin, 1962; searle, 1975). 2 https://blog.twitter.com/en_us/topics/insights/2019/thishappened-in-2019 digging communicative intentions the following are two examples3 of how these two classifications proceed. we use the → notation, with, at its left, the tweet-level categories, and, at its right, the sub-tweet level categories. a precise definition of the labels will be provided later in the paper (see section 3). (1) a. [ the fire situation in the landiras area is getting worse.]1 [please follow the instructions of the fire brigade and the police.]2 (sa annotation) jussive → 1. proper assertive; 2. open option (urgency annotation) urgent: warning/advice b. [ 5th day of fire fighting, about 6000 hectares of our forest charred here. still the same means at the disposal of our firemen: 2 air-crafts and 1 dash.]1 [ what are you waiting for to give them the means to stop this fire? @emmanuelmacron @gdarmanin #landiras]2 (sa annotation) subjective → 1: proper assertive; 2: evaluation (urgency annotation) urgent: material damage as shown in these examples, a tweet is composed of several parts that contribute to the construction of the communicative intention of the whole message. these parts may convey (and they indeed often do) very different speech acts types. therefore tweets need also to be analyzed at the subtweet level, in order to search for more precise and specific content that provides useful actionable information. for (1-a), the writer publicly expresses an explicit demand (hence a jussive4 speech act at the tweet level) for the population to follow the authorities’ instructions as the wildfires in the landiras region keep spreading. at the sub-tweet level, (1-a) first presents a description of the situation (cf. segment 1 that triggers a speech act of proper assertive) and then provides an advice on how to behave (see segment 2 which qualifies as an open option in our classification, cf. infra). the latter is the most useful piece of content as it provides new and actionable information triggering action expectation. for emergency and actionability, (1-a) qualifies as urgent at the tweet level, specifically providing content that falls in the actionability category advice. as a further example, insofar as the speech act annotation is concerned, (1-b) expresses an intention to complain about the current means at the disposal of the fire brigades. the overall tweet is considered as expressing a subjective stance of the speaker (hence the overall label subjective) in virtue of the question, which reveals a complaint (the part containing the question is labeled as evaluation, cf. infra for details). the first segment is a proper assertive. as for the emergency annotation of the same tweet, (1-b) qualifies as urgent at the tweet level, providing content that is labeled material damages at the actionability level.5 1.3 previous approaches and research questions since the introduction of dialogue acts (see, a.o., the damsl framework (allen and core, 1997; core et al., 1998)), sa have been dedicated an extensive body of work in the computational linguistics literature where various approaches have been proposed to detect them in both synchronous (e.g., meeting, phone) (stolcke et al., 2000; keizer et al., 2002; carvalho and cohen, 2005; joty and mohiuddin, 2018) as well as asynchronous dialogues (e.g., emails, live chats, tweet threads) 3. these are examples taken from our french corpus translated into english. 4. we borrow the latin word for order as standard practice in linguistics, see portner (2018) 5. kozlowski et al. (2020) classification also comprises a prior level of relatedness as we explain later in the paper. 3 (carvalho and cohen, 2005; joty and mohiuddin, 2018; bracewell et al., 2012). sa have shown to be an important step in many downstream nlp applications such as strategic actions prediction (cadilhac et al., 2013), dialogues summarization (goo and chen, 2018) and conversational systems (higashinaka et al., 2014). however, sa for emergency detection has received less attention in the literature and most of related work on communicative intentions in social media are topic-oriented, focusing on the correlation between sa and specific topics such as crisis (e.g., earthquakes, bombing, attacks) but also politics, celebrities, cooking, travel, etc. (zhang et al., 2011; vosoughi, 2015; elmadany et al., 2018a; saha et al., 2020b). these corpus-based studies show that there is a greater similarity of distribution between topics of the same type than between topics of different types. in particular, it has been observed that people globally tend to react to natural disasters with sa distinct from those used in other contexts (e.g., celebrities, which are essentially made up of comments). here, we explore the further hypothesis of a correlation between different sa types and urgency. we thus investigate whether sa can be used to sort urgent from not urgent messages. as far as we know, this is the first study that proposes an in depth linguistic and computational analysis of communicative intentions in tweets from an urgency-oriented perspective: what are the most frequent intentions in urgent vs. not urgent message? are these intentions different from those found in non useful messages? and more importantly, are they particularly correlated with finegrained urgency categories (such as human/infrastructure damages, donations, security instructions etc.)? finally, are the observed sa stable across different types of crisis (flood, hurricane, fire, attack, etc.)? to answer these questions and before moving to real scenarios that rely on sa-aware automatic detection of urgency (this is left for future work), we propose to (1) measure the impact of sa in detecting urgency during crisis events in manually annotated data, and (2) explore the feasibility of sa automatic detection in crisis corpora. 1.4 overview of the main contributions we build on laurenti et al. (2022a) where we performed a preliminary analysis of the role of sa on urgency detection in about 6,6k tweets with of a focus on natural disasters (flood, hurricane, storm, etc.). in laurenti et al. (2022a), we relied on a new annotation scheme of sa that takes into account the variety of linguistic means whereby sa are expressed (including lexical items, punctuation, etc), both at the message and sub-message level. we further extend this initial work by proposing: • the first largest french dataset of about 13,300 tweets annotated for both urgency and sa following the same annotation scheme. in addition, we expend the annotations to 6 new sudden crisis making the dataset spans over 20 crises.6 • a qualitative and quantitative analysis of the annotation campaign intersecting the two-level classification of speech acts with a classification of urgency. in particular, we explore the correlations between sa vs. urgency, sa vs. intention to act categories as well as sa vs. the types of crises for both levels of sa annotations. our results show a strong correlation between sa and urgency annotations at both the tweet and sub-tweet levels with a particular salient correlation in the latter case which constitutes a first important step towards sa-aware nlp-based crisis management on social media. 6. the annotated dataset will be available for research purposes upon request. 4 digging communicative intentions • a set of deep learning experiments to detect speech acts relying on deep learning architectures coupled with relevant linguistic features about how sa are linguistically expressed. we consider several experimental settings ranging from monotask to multitask learning including multi-label classification. our results show that sa detection achieve very encouraging results proposing to the community a novel state of the art of sa detection in french social media. • an error analysis of the automatic detection at both sa levels, highlighting main cases of mis-classification. this paper is organized as follows. section 2 presents related work in sa detection in social media as well as main existing crisis datasets. section 3 provides the classification of sa we propose and the annotation guidelines to annotate them. sections 4 and 5 respectively detail the dataset we relied on and the results of the annotation campaign. section 6 focuses on the experiments we carried out to detect sa automatically. we end by some perspectives for future work. 2. related work speech acts have been extensively studied in the computational linguistics literature since early 2000’s. most studies focus on sa in human-human dialog conversations where several datasets have been annotated relying on various taxonomies of sa (also known as dialogue acts), such as question, acknowledgment and follow-up questions (see serban et al. (2018); gonçalo et al. (2022) for recent surveys in the field). dialogues being out of the scope of this paper, we focus in this section on sa for social media content, a relatively under-explored area of research compared to dialogue. we first provide an overview of sa used to annotate tweets about various events including crises as well as other domains (politics, offensive language, etc.). we then review main approaches for sa automatic detection. as our dataset for the first time combines sa and urgency annotations, we end this section by presenting existing crisis-related datasets highlighting the novelty of this study. 2.1 speech acts in social media 2.1.1 speech acts in the crisis domain the main line of analysis of the role of sa in tweets consists in unveiling how speech acts (as used on twitter) vary qualitatively according to the topic discussed. in this line of questioning, sa have been studied as filters for new topics. zhang et al. (2011) in particular, resorts to a searlian typology of sa that distinguishes between assertive statements (description of the world) and expressive comments (expression of a mental state of the speaker). zhang et al. (2011) also distinguish between interrogative questions and imperative suggestions. finally, a category miscellaneous brings together the searlian declaratives and the commissives, used to make promises. concerning the question of emergency, zhang et al. (2011) showed that the sa’s distribution on twitter in the context of a natural disaster (e.g., earthquake in japan) is distinctive: it is essentially composed by statements, associated to comments and suggestions / orders. in this context new information or ideas on how to (re)act are indeed expected and assertions are the most suitable to this aim. by contrast, discussion over a celebrity will mostly generate comments and 5 almost no order or suggestion. indeed, in this context, subjectivity matters more than immediate action. also inspired by searle’s typology, vosoughi (2015); vosoughi and roy (2016) distinguish six categories: assertions, recommendations, expressions, question requests and miscellaneous. the authors use the definitions of zhang et al. (2011), by distinguishing the topic discussed in the tweets, from the type of topic (entity-oriented, event-oriented topics, or longstanding topics which are topics about subjects that are commonly discussed). six topics were then selected (2 of each type): for entity-oriented, they are interested in ashton kusher and the red sox; for event-oriented, they studied the boston bombings in 2013 and the ferguson demonstrations in 2014; for longstanding topics, they considered cooking and travel. the distribution of speech acts shows a greater similarity of distribution between topics of the same type than between topics of different types. on the other hand, the entity-oriented and event-oriented types are closer to each other, with a majority of assertions and expressions, whereas for the long-standing types, assertions are less abundant and recommendations well represented. in this same perspective of topic identification and relying on the same topic characterization as above, elmadany et al. (2018b) manually annotate 21,000 tweets in arabic according to their topic type and distinguish events like sinai bombings, gulf crisis, arab spring and world cup qualifications, entities (especially people) and various issues such as travel or cooking. each tweet is associated to a pair of speech act/sentiment according to the following classification: assertions, recommendations, expressions and requests, and among sentiments, the standard positive, negative, mixed and neutral categories. their study reveals a salient association between assertions and people/events and neutrality on the one hand and an association between expressivity long-standing topics and negativity on the other. 2.1.2 sa in other domains in a recent and extensive study of sa in social media, bell (2020) takes on a different approach than other studies in the literature on speech act theory and conducts an empirical investigation into the identity of illocutionary force indicating devices, which are the elements responsible for encoding a speaker’s intentions. a corpus of 1,000 twitter threads is collected, manually segmented by an expert and annotated at the sub-tweet level, allowing multiple speech acts per tweet, as opposed to most other studies. they consider the following sa: assertive, directive, interrogative, expressive, commissive, exercitive (with a commissive, the speaker commits themselves, with an exercitive, the speaker requires someone else’s commitment). this study distinguishes direct and indirect illocutionary acts (i.e. acts performed by way of performing another). regarding the direct force, the majority of segments (64.5%) were annotated as assertive. the second most frequent category was expressive, with 16.3%. on the other end, the least frequent category was exercitive, with 0.25%. regarding the indirect force, 83.9% of tweets were determined to perform no indirect act, and those annotated as performing one were about 80% expressives. in plakidis and rehm (2022), an annotation of sa is done using a subset of 600 tweets taken from a german corpora of offensive and non-offensive tweets. mainly inspired by searle (1975), and building upon compagno et al. (2018) and weisser (2018), the tweets are segmented in sentences, which are then annotated on two main levels : the syntactical level (eg. declarative, exclamative, imperative, etc.), which describes the type of sentence, and the speech act level, consisting of a 6 digging communicative intentions coarse-grained and a fine-grained level, which describes the type of speech act. the categories used for the first speech act level are as follows: assertive, expressive, directive, commissive, other and unsure. they are subsequently detailed into 23 sub-classes at the sub-tweet level. for example, the category assertive is further detailed into the following 6 sub-categories: assert (”it costs 200$”), sustain (”i’m going to buy it because it’s very convenient”), guess (”i’m unsure he’s right for her”), predict (”it will be a few hundreds at most”), agree (”you are right”), disagree (”i don’t think so”). the results suggest that offensive language contains more expressives and less assertive than non-offensive language. tweets with implicit offensive language have a lower frequency of expressives and a higher frequency of assertives than tweets with explicit offensive language. in view of the topic – offensive language – the distinction between assertives and expressives is reported as a prominent issue, which does not arise in the context of urgency detection, where the description of facts (assertives) and the evaluation of said facts (expressives) are more clearly distinct. for completeness, we note that sa have also been studied in the context of political campaigns, notably by subramanian et al. (2019), with a corpus of 258 official documents related to the 2016 australian ”federal election cycle”: official statements, tweets, press clippings, etc. from which 7641 utterances are extracted. each utterance is annotated with a sa and a target party (liberal or conservative). the categorization of sa articulates: assertives, commissives-actionspecific, commissive-action-vague, commissives-outcome (about a future reality state), directives, expressives, past-actions and verdictives (an assessment on prospective or retrospective actions). they observe an over-representation of assertives (40%), followed by verdictives (25%) and specific action (12%). the other categories represent less than 10% of the annotations. it is interesting to note that commissives make up a almost a quarter of the assigned speech acts, whereas they are almost absent from our corpus, which is related to emergency. 2.1.3 sa automatic detection sa prediction has been tackled either as a primary task (i.e., multi-class classification problem) or auxiliary task where sa information are used to boost the performances of classification tasks such as sentiment analysis, emotion detection or hate speech detection. some works consider speech acts at the message level while others consider dialogue acts when uttered in conversations. at the message level, most state of the art approaches make use of feature-based machine learning algorithms (svm, naive baise, decision tree) relying on various surface, lexicon and syntactic features such as unigrams, punctuations, pos, emoticons and sentiment words (zhang et al., 2011; rojas-barahona et al., 2012; franovic and šnajder, 2012; vosoughi and roy, 2016; sherkawi et al., 2018; algotiml et al., 2019). deep learning architectures have also been explored. saha et al. (2021) propose a multi-modal approach for detecting sa in arabic tweets relying on a multi-tasking framework based on dyadic attention mechanism (vaswani et al., 2017) and adversarial loss to predict simultaneously sentiment, emotion and speech acts. it employs intra-modal and inter-modal attention to fuse multiple modalities and learn generalized features across all the tasks. subramanian et al. (2019) propose a target based speech act classification on a dataset of political discourse using a semi-supervised learning approach (bigru) by incorporating contextualized word representations 7 (elmo) and a cross-view training framework to augment the initial dataset with in-domain unlabeled text. finally, saha et al. (2020a) combine bert and capsule networks (sabour et al., 2017) to asses the intent of tweets (expression, statement, suggestion, threat, request, question). another line of research focuses on predicting sa in social media conversational thread casting it into a sequence labeling problem. for example, cerisara et al. (2018) use a two-level hierarchical recurrent network (bi-lstm and rnn) to predict dialog acts and sentiments. joty and mohiuddin (2018) experiment with an lstm-rnn architecture to represent sentences of a conversation then crf models to extract the inter-sentence dependencies. the approach has been evaluated on many synchronous and asynchronous corpora, including forum conversations from tripadvisor. other works propose to model sa in dialogues as a multi-label classification problem. for example, xu et al. (2017) rely on a cnn model on top of pre-trained word vectors by utilizing a threshold learning mechanism. the model has been evaluated on the task of dialog state tracking. 2.2 crisis datasets the literature on emergencies detection in social media has been growing fast in the recent years and several datasets (mainly tweets) have been proposed to account for crisis related phenomena such as flood, hurricane, storm and attacks.7 messages are annotated according to relevant categories that are deemed to fit the information needs of various stakeholders like humanitarian organizations, local police and firefighters. annotations are usually done at the text level relying either on crowd-sourced workers, humanitarian volunteers or domain experts.8 relevance criteria found in the literature can be grouped into the following dimensions: • relatedness (also known as usefulness or informativeness) to identify whether the message content is useful provides valuable information that might be relevant to rescue teams. this is generally cats into a binary classification problem: is the message useful vs. non useful. this dimension is used in almost all state of the art annotation guidelines (imran et al., 2016; kaufhold et al., 2020). • urgency (also known as criticality or priority) to filter out on-topic relevant information that can aid people in making decisions, advise others or offer immediate post-impact help, and on-topic irrelevant including offers, supports and solicitations for donations to charities (imran et al., 2013; mccreadie et al., 2019a; sarioglu kayi et al., 2020; kozlowski et al., 2020; kejriwal and zhou, 2020). • intention to act, also know as humanitarian information type (alam et al., 2021). urgency is often associated with a taxonomy of intention to act categories such as: caution or advice, donations, people missing, found, or seen and damage infrastructure (imran et al., 2016; olteanu et al., 2015). • eyewitnesses types. it is used to identify direct (first-hand knowledge and experience of an event), indirect (messages sharing valuable information from direct witnesses) and vulnerable direct eyewitness (users reporting warnings and alerts) (zahra et al., 2020). annotations in most existing datasets are usually carried out at the message level. 7. see https://crisisnlp.qcri.org/ for an overview. 8. some studies propose to additionally annotate images within the tweets (see for example alam et al. (2018)) 8 https://crisisnlp.qcri.org/ digging communicative intentions existing datasets are either annotated according to one of the dimensions above or using several dimensions in cascade like relatedness or urgency first, then information type for messages that have been identified as relevant. most annotated datasets are in english. well known datasets include trec-is 9 (mccreadie et al., 2019b, 2020), a shared task that aims to develop real-time monitoring systems capable of monitoring the development of incidents such as natural disasters, terrorist incidents or public health crises from online text data feeds. we also cite the crisisfacts2022 dataset 10 which aims at generating a summary of crisis. few crisis datasets exist in other languages such as spanish (cobo et al., 2015), arabic (alharbi and lee, 2019), italian (cresci et al., 2015). for french, the only publicly available dataset is the one developed by kozlowski et al. (2020) who propose a three-level classification of tweets : relatedness, urgency and intention to act categories to deal with missing people, human/infrastructure damage, etc. this dataset focuses on several natural disasters (hurricanes, flood, storms, etc.) going beyond the french portion of crisisnlp 11 that only focuses on one type of crisis (landslide). 2.3 contributions as far as we are aware, communicative intentions have been explored in connection with urgency detection in two previous works. first, laurenti et al. (2022a) propose a sa classification for french tweet in the crisis domain. they focus on ecological crises and propose a two-layer annotation scheme to manually annotate a dataset of 6,669 tweets both for urgency (urgent, not urgent and not useful) and sa (tweet level: assertive, subjective, interrogative and jussive). quantitative analysis of the annotations showed a correlation between tweet-level sa and urgency categories. this dataset has been used for supervised sa classification where a set of deep learning experiments have been carried out based on the camembert transformer architecture to classify each tweet into four sa categories at the tweet level. laurenti et al. (2022b) built on this pre-trained classifier and propose sa-aware urgency detection models, showing that injecting sa as external semantic feature is a promising direction to improve urgency detection in social media. in the present paper, we rely on the annotation scheme initially proposed in laurenti et al. (2022a) and advance these previous studies, making six new contributions: 1. we double the dataset in laurenti et al. (2022a) and create the largest french dataset annotated for sa with a total of about 13k at both the tweet and sub-tweet levels. 2. we extend to 6 new sudden crises, making the dataset cover both sudden and expected crises for a total of 20 events. this new dataset is the first that combines both sa and urgency annotations, to the best of our knowledge. 3. we correct the initial dataset addressing shortcomings related to the annotations at the subtweet level. we propose an automatic procedure to check for annotation inconsistencies which yields to a significantly improved version of the annotations. 4. in addition to the quantitative analysis we made on the initial portion of the dataset (6,6k) (i.e., sa vs. (urgent vs. non urgent vs. non useful)), we newly explore: (1) the correlation between sa and 6 intention to act categories, among human/infrastructure damages, warning advice, 9. https://www.dcs.gla.ac.uk/˜richardm/trec_is/ 10. https://crisisfacts.github.io/ 11. https://crisisnlp.qcri.org/ 9 https://www.dcs.gla.ac.uk/~richardm/trec_is/ https://crisisfacts.github.io/ https://crisisnlp.qcri.org/ critics, supports, etc. (2) the distribution of sa across crisis types (sudden vs. expected events), and (3) a study of the sa evolution across time. 5. in addition to sa detection at the tweet-level relying on baseline architectures, we newly: (1) address sub-tweet sa detection as well as joint tweet/sub-tweet predictions relying on monotask and multitask learning approaches while evaluating models performances to classify each message into a single class vs. multi-label. as far as we know, handling sa in social media content as a multi-label problem has not been explored before, (2) experiment model adaptability across crisis types and layers. to the best of our knowledge, this is the first attempt to the automatic sa detection in a french social media dataset. 6. finally, we provide a detailed error analysis of our results at both tweet and sub-tweet levels. overall this paper proposes an in-depth study of speech act in view of their contribution to enhance emergency detection. before moving to real scenarios that rely on sa-aware automatic detection of urgency – which we leave for future work – our aim here is (a) unveil the contribution of speech act to emergency detection on a distributive basis, and (b) explore sa detection in french social media across various crisis types. this is, as far as we know, the first work that addresses the issue in such exhaustive manner. the second step that will consider injecting sa to improving urgency detection is out of the scope of this paper. 3. a two-level annotation of speech acts for urgency detection we deployed two layers of annotation for speech acts: • sa1: at the first level, we use a classification including 5 distinct categories, which we apply to the tweet as an atomic unit. • sa2: at the second level, 8 categories are used to annotate tweets at the sub-tweet level as opposed to the tweet as a whole. the goal of this two-layers annotation is to allow us to dig fine-grained information about speaker’s posture towards the event, to ultimately identify the main communicative intention of the tweet as a whole. in this section, all examples are taken from our corpus and provided in french together with their english translations. urls and private user mentions have been replaced by and respectively. each example comes with its sa1 (cf. sections 3.1 and 3.2) and sa2 (cf. section 3.2). notation-wise, recall from section 1.2 that we use arrows (→) to signify the relation between first-level and second-level sa categories, at the left and right of the arrow respectively. in addition, in order to show the interplay between sa and urgency annotations, all examples come with urgency annotations (urgent vs. not urgent vs. not relevant) as well as six intention to act categories as follows: (1) urgent applies to messages mentioning human, infrastructure damages as well as security instructions to limit these damages during crisis events, (2) not urgent groups support messages to the victims, critics or any other messages that do not have an immediate impact on actionability but contribute in raising situational awareness, and finally (3) not urgent for messages that are not related to the targeted crisis. please note that all urgency annotations have been removed during the sa annotation campaign (cf. section 4.1). 10 digging communicative intentions 3.1 tweet level our classification of sa elaborates on the fundational austinian and later searlian distinction by (i) relying on propositional content and lexical clues such as modals (should, must, can, ...), evaluative adjectives, attitude verbs (think, believe, want, hope ...); (ii) introducing the category subjective, which reshuffles some of the earlier classifications (‘wishes’, for instance are subjective rather than jussive in our classification (e.g.,condoravdi and lauer (2012)); (iii) considering presuppositional content as well (see mari (2016) on french). we distinguish four first-level categories which are mutually exclusive and define tweets as wholes, at a holistic level, as shown in figure 1. figure 1: a classification for tweets that makes use of four illocutionary categories. (1) jussive, as defined by zanuttini et al. (2012), enhance commitment to take action, as in (2). importantly, there is no strict correlation between the imperative form and jussive. as the example shows, the imperative form is not needed to enhance action. in this respect, our classification aligns with accounts that do not ground speech acts in sentence types (see portner (2018) for extended discussion). (2) incendies #feuxdeforêt #gironde 1. ne pas se fier uniquement aux prévisions de météo france 2. si fumée lire le communiqué 3. laisser les #sapeurspompiers effectuer leur rotation de 12 heures au feu 4. 96% des sinistres sont d’origine humaine (source sdis 33) merci (wildfires #forestfires #gironde 1. do not rely solely on météo france forecasts 2. if smoke read the press release 3. let the #firebrigade carry out their 12-hour rotation at the fire 4. 96% of fires are of human origin (source sdis 33) thank you ) sa1: jussive urgency: urgent → security instructions (2) assertive. assertions, like in (3), are considered to convey objective truth (as opposed to subjective truth (giannakidou and mari, 2021c). with assertive, the speaker is committed 11 toward the truthfulness of the proposition that is being uttered ((portner, 2018) a.o.) and require their interlocutor to update the common ground (ginzburg, 2012). (3) direct. deux immeubles s’effondrent à lille: les secours cherchent une victime dans les décombres via @lavoixdunord (direct. two buildings collapse in lille: rescue workers search for a victim in the rubble via @lavoixdunord) sa1: assertive urgency: urgent → human damage at this level of the classification, this is a simplification of what assertions are. when asserting, speakers can lie, or they can use a partial knowledge that undermines the likelihood of the assertion to express true. to nuance this simplification, we elaborate on the notion of assertion at the second level of the annotation, where we introduce some evidentiality-based distinctions. (3) interrogative. this category is dedicated to those questions that require an informative answer, like in (4) the questions that, besides triggering an answer, reveal bias and expectations on the part of the speaker (see ladd (1981)) are classified as subjective (see below). (4) @emmanuelmacron où sont les renforts censés arrivés à saint-martin et que comptez-vous faire. #sxmirma #saintmartin #irma #sxmstrong #sxm (@emmanuelmacron where are the reinforcements supposed to arrive in st. martin and what are you planning on doing. #sxmirma #saintmartin #irma #sxmstrong #sxm) sa1: interrogative urgency: not urgent → critics (4) subjective. finally, with subjective, as in (5) the speaker shares a mental state that can be either a personal evaluation or preference (see among many others (lasersohn, 2005)) or an expressive state (an emotion or a feeling, (giannakidou and mari, 2015)). the interlocutor is asked to update the common ground not just with the content of the evaluation but with the evaluation itself (see simons (2007), and for recent discussion on french: mari and portner (2021)). in our classification, ‘wishes’, for instance, are subjective rather than jussive as they do not trigger any commitment to act so to make the content of the wish true (this is the emotive content of the wish (giannakidou and mari, 2021a)). (5) #incendie l’abbaye de frigolet..la catastrophe... un désastre.. (#fire at frigolet abbey..the catastrophe... a disaster...) sa1: subjective urgency: urgent → infrastucture damage (5) other. additionally, other is added to the classification, for undecidable cases. (6) feu d’artifices du 14 juillet @villedeputeaux (14th of july fireworks @cityofputeaux) sa1: other urgency: not useful one important feature of our classification is that it does not rely on sentence type, but on sentence interpretation. for instance, an imperative is not necessarily classified as a jussive. imperatives that convey wishes, as we noted, are considered to be subjective. likewise, the interrogative 12 digging communicative intentions form, does not necessarily correlate with the interrogative category. an interrogative can express a point of view, or even knowledge, as in the case of rhetorical questions. in (7), the speaker is not really asking a question, but rather wants to express their opinion that the authorities are not doing the right thing, hence expressing a subjective point of view. (7) @prefet974 #berguitta .. alerte orange pour rien hier qui a penalisé l’économie et pas d’alerte rouge pour ne pas pénaliser l’économie quand le danger est réel.. on marche sur la tête? (@prefet974 #berguitta . orange alert for nothing yesterday that penalized the economy and no red alert to not penalize the economy when the danger is real ... are we walking on our heads ?) sa1: subjective urgency: urgent → security instructions 3.2 sub-tweet level we consider each tweet as a discourse unit, composed of one or more statements or sub-segments, so that it can not only be classified at the holistic level but also at the level of its segments (identified in the following examples between ‘[ ... ]’). in order to achieve this, we have elaborated on each of the four categories at the tweet level to annotate the tweets at the segment level relying on eight categories (see figure 2). figure 2: two-layers annotation for tweets and inner segments. for jussive, the annotation distinguishes between (a) open-option – the speaker puts forward a possibility and leaves the addressee free to realize it or not (cf. (8)) – , and (b) utterances that enhance a direct commitment on the part of a discourse participant, ie. commissives, exhortatives, orders and prohibitions, that are called other-jussive (cf. (9)). (8) [cyclone #irma : qu’est ce que le cic?]1 [présentation de cet outil de #gescrise ci-dessous! ]2 ([ cyclone #irma : what is the cic? ]1[ presentation of this #gescrise tool below! ]2) sa2: 1. interrogative→ uninformative, 2. jussive→ open option urgency: not urgent → other messages (9) [un été caniculaire de tous les dangers avec des incendies dans plusieurs régions.]1 [alors redoublons de prudence et de vigilance. pas de barbecues en forêt, de cigarettes allumées...]2 ([ a dangerous hot summer with fires in several regions. ]1[ so let’s be extra cautious. no 13 barbecues in the forest, no lit cigarettes... ]2) sa2: 1. assertive→ proper, 2. jussive→ other jussive urgency: urgent → security instructions for assertive, both second-level categories are determined by the source of knowledge that the speaker relies upon, i.e. the evidentiality condition as defined by saurı́ and pustejovsky (2009). if the speaker grounds their utterance on a third-party source, the assertive utterance is (a) a reported assertive, whereas if there is no such explicit source, it is a (b) proper assertive, see (10) and (11) respectively. (10) [inondations dans l’aude. macron promet 80 millions d’euros:]1 [est ce suffisant ???]2 ([ floods in the aude. macron promises 80 million euros: ]1[ is this enough??? ]2) sa2: 1. assertive→ reported, 2. interrogative→ informative urgency: not urgent → other messages (11) [le feu de landiras (au départ à 40km) s’approche de chez moi. encore 2 villages et c’est à nous d’évacuer. ce soir on sent vachement le brulé.]1 [on ne panique pas mais le stress monte. vais mal dormir.]2 ([ the landiras fire (initially 40km away) is approaching my home. two more villages and it’s up to us to evacuate. tonight it really smells like burnt. ]1[ no panic but the stress is mounting. i’m not going to sleep well. ]2) sa2: 1. assertive→ proper, 2. subjective→ expressive urgency: urgent → human damage it is important to note that the distinction between reported and proper assertive is meant to reveal a difference in degrees of commitment on the part of the speaker. on the assumption that a proper assertive reveals total commitment to the truthfulness of the content of the assertion, by signaling that the content of the assertion is reported, the speaker is considered as willing to distance themselves from the truth of that content (see discussion in aikhenvald (2004); giannakidou and mari (2021a) and subsequent literature). while we are aware that a certain amount of simplification remains (assertions can be lies, for instance (see extended discussion in giannakidou and mari (2021c)), this distinction allows us to introduce a certain degree of complexity in our treatment of the attitudinal domain. for subjective, a distinction is made between (a) expressives/evaluatives whereby the speaker describes a personal evaluation or an expressive state that it is not deemed to become common ground or truth (cf. (12))(lasersohn, 2005; giannakidou and mari, 2021c; mari and portner, 2021)) and (b) other subjective for utterances that do not explicitly fall in the previous category (eg: puns, greetings...), see (13). (12) [@prefet29 enfin !]1 [un grand bravo aux pompiers et aux agriculteurs qui sont venus aider à maı̂triser cet incendie]2 ([ @prefet29 finally! ]1[ congratulations to the firefighters and farmers who came to help control the fire ]2) sa2: 1. subjective→ expressive, 2. subjective→ evaluative urgency: not urgent → support (13) [le paradis en feu.]1 [grosses pensées aux pompiers, policiers, bénévoles qui se battent s’en relâche depuis 1 semaine maintenant, nos cœurs sont serrés et nous prenons notre mal en patience.. #bassindarcachon #feuxdeforet #sud #dunedupilat #incendiesgironde 14 digging communicative intentions #incendies ]2 ([ heaven on fire. ]1[ my thoughts go out to the firemen, policemen, volunteers who have been fighting relentlessly for a week now, our hearts are heavy but we grin and bear it.. #bassindarcachon #forestfires #south #dunedupilat #firesgironde #fires ]2) sa2: 1. subjective→ other subjective, 2. subjective→ expressive urgency: not urgent → support the expressive/evaluative category is a complex one which can be enhanced by a variety of linguistic means, such as evaluative adjectives (including moral adjectives, good, right, epistemic adjective such as clear, evident) modality (must, might, should, would etc), adverbs (obviously, regretfully, etc.), particles (in french, bien (ok), bon (good), ...) (see giannakidou and mari (2021c)). for interrogative, a distinction is made between (a) informative questions to which the speaker cannot answer and which require an answer triggering new information and the ones that are (b) uninformative indicating that the speaker is biased towards an answer, as in (14) and (15) respectively.12 (14) [une semaine après le drame, on continue d’éclairer les zones d’ombre.]1 [pourquoi le médecin qui a trouvé la mort dans l’ #effondrement n’a pas été évacué ? #lille ]2 ([ one week after the tragedy, we continue to shed light on the grey areas. ]1[ why was the doctor who died in the #collapse not evacuated? #lille ]2) sa2: 1. assertive→ proper, 2. interrogative→ informative urgency: not urgent → other messages (15) [etes-vous en zone inondable ?]1 [retrouvez l’actualisation de la carte de prévention du #risque #inondation à #paris sur ]2 ([ are you in a flood zone? ]1[ find the update of the #flood risk prevention map in #paris on ]2) sa2: 1. interrogative→ uninformative, 2. jussive→ other jussive urgency: not urgent → other messages as for sa1, we also add other to the sa2 classification, for undecidable cases. 4. data and annotation in this section, we provide details on the dataset used, the annotation procedure, and the results of the annotation campaign. 4.1 dataset since our focus is on crises that occur in metropolitan france and its overseas departments, we rely on the only available corpus of french tweets by kozlowski et al. (2020)13 and augmented later on by bourgon et al. (2022b) with sudden crises (attacks, explosion, fires, etc.). the collection is composed of 19,595 tweets collected using dedicated keywords about ecological crises that occurred in france from 2016 to 2022 and posted 24h before, during (48h) and up to 72h after the crisis: 2 12. see larrivée and mari (2022) for french and ginzburg (2012); giannakidou and mari (2021b) for a more general discussion and cross-linguistic observations. 13. https://github.com/diegokoz/french_ecological_crisis 15 https://github.com/diegokoz/french_ecological_crisis floods that occurred in aude and corsica regions, 8 storms (béryl, berguitta, fionn, eleanor, bruno, egon, ulrika, susanna), 2 hurricanes (irma and harvey), 2 building collapses (marseille, lille), 2 chemical plants explosions (lubrizol, sanary), 2 fires (notre-dame fire, gironde and landes wildfires) and 1 terrorist attack (trèbes).14 the data comes with additional metadata including: number of likes, retweets, followers and followings of the user. in this dataset, each tweet is annotated following an urgency classification composed of three urgency categories as well as 6 intentions to act categories: (1) urgent that applies to messages mentioning human/infrastructure damages as well as security instructions to limit these damages during crisis events, (2) not urgent that groups support messages to the victims, critics or any other messages that do not have an immediate impact on actionability but contribute in raising situational awareness, and finally (3) not useful for messages that are not related to the targeted crisis or information pertaining to events occurring outside the french territories. this scheme has been used to annotate the dataset by two annotators who achieved a kappa inter-annotator agreement of 0.67 and 0.65 for urgency and intention to act classification respectively (kozlowski et al., 2020). not useful urgent not urgent total security human infra. support other critics instruc. damage damage messages flood aude 1,065 150 34 157 157 184 26 1,773 flood other 993 292 35 111 231 16 19 1,697 flood corse 468 51 58 12 52 66 13 720 storm béryl 612 91 0 2 3 10 2 720 storm bruno 586 107 5 11 2 9 0 720 storm susanna 484 129 11 38 4 54 0 720 storm ulrika 650 47 2 18 0 4 0 721 storm berguitta 587 56 5 9 12 46 5 720 storm fionn 552 138 6 10 0 8 6 720 storm egon 609 66 1 35 0 10 0 721 storm eleanor 590 82 22 19 1 6 0 720 hurricane harvey 628 78 10 2 1 1 0 720 hurricane irma 790 121 47 55 199 199 29 1,440 collapse marseille 627 9 24 11 11 19 19 720 collapse lille 320 2 39 27 12 117 32 549 wildfire gironde landes 1,394 51 23 93 317 380 165 2,423 wildfire notre-dame 86 224 209 519 plant explosion lubrizol 137 583 627 1,347 plant explosion sanary 6 363 164 533 attack trèbes 174 398 810 1,382 total 11,358 3,970 4,257 19,595 table 1: urgency distribution in our dataset per crisis. table 1 presents the distribution by class for all available crises. some crises (plant explosion lubrizol, plant explosion sanary, notre-dame wildfire, attack trèbes) are only annotated for urgency. the ecological crisis (flood, storm, hurricane) are the most represented with 12,112 messages against 7,483 messages for sudden crisis (collapse, wildfire, plant explosion, attack). we also 14. for long-term crises, the end of the crisis has been fixed to the date of resolution of the crisis, e.g. extinction of the first fires in the case of fires land. 16 digging communicative intentions notice that for sudden crises, there are fewer security instruction messages than ecological crisis, explained by the fact that these latter crisis are predictable. the collection is extremely imbalanced with 57.96% not useful and 20.26% for urgent. this is largely due to how tweets are collected. indeed, since tweets posted 24 hours before the crisis have been collected, a large amount of them are not useful. the corpus is also imbalanced regarding the sub-level of urgency categories: 1.93% of the tweet are annotated as human damage with 306 messages while security instruction represents 9.88% of the corpus with 1,470 tweets. these proportions are in line with the ones reported in other crisis corpora (see section 2.1). 4.2 annotation procedure a subset of this dataset composed of 13, 378 tweets has been selected for sa1 annotations, among them 11, 229 have been annotated for both sa1 and sa2. regarding sa1 dataset, it comprises almost all urgent (3,857) and not urgent (4,222) messages. only 5,299 not useful tweets have been selected, in order to reduce the size of that category, but keep it as the majority class. similar urgency annotations split holds for sa2 dataset. note that, during the annotation process, pre-existing urgency tags and metadata information are removed, as to not bias the annotators. the annotators were native french speakers, both master’s degree students in linguistics. the procedure was as follows. first, each segment in a given tweet is annotated at the sub-tweet level (i.e., sa2), then the tweet level annotation (i.e., sa1) is deduced accordingly: • if the tweet is composed of one or several sa2 annotations that subsume the same sa1 category, the final annotation is sa1. for example, for a tweet composed of two segments annotated with sa2=[informative, uninformative], then sa1=interrogative. • in case of several segments annotated with sa2 that do not belong to the same sa1 category, annotators are asked to determine the main communicative purpose of the tweet, and what segment signifies the main communicative intention of the speaker ((simons, 2007; mari and portner, 2021) a.o.). the main criterion to identify the main intention relies on the determination of the background (known) foreground (new) information.for example in (16), a tweet is composed of two segments: a proper assertive, followed by an uninformative question that conveys an evaluation. the annotators have considered the second segment to be dominant, as the fist half is a description of a fact that occurred in the past and that is already part of the common ground. the main point of the tweet is the uninformative question about the present situation, as an expression of a criticism.15 the tweet is thus labeled at the first level as subjective. the sa2 annotation and the background-foreground distinction provides a solid heuristic to identify the main point of the tweet. furthermore, as we shall see in section 5.1, the fist segment is mostly responsible for determining the overall categorization, this providing a reliable criterion to settle undecided cases. finally, as we show in section 5.3, specific subsegments correlate with urgency, thus enhancing emergency detection. (16) [#marseille la mairie de marseille a touché des millions pour la rénovation des immeubles urbains.]1 [qu’en ont-ils fait ?! @joelle dago #ggrmc ]2 ([ the marseille city council has received millions for the renovation of urban buildings. 15. recall that questions can convey a subjective stance rather than a request of information. 17 ]1[ what have they done with it?! ]2) sa2: 1. assertive→ proper, 2. subjective→ evaluative sa1: subjective urgency: not urgent→ critics the annotation has been performed using the brat annotation tool. (stenetorp et al., 2012)16 to ensure consistency between annotations at the sa1 and sa2 levels (i.e., a tweet composed of one segment and annotated with sa1=interrogative and sa2=proper assertive), automatic checks have been conducted and annotators are asked to solve their errors before moving to the next tweet. figure 3 shows an example of the tweet ”a fire is currently in progress in #saintdizier in the city center. avoid the area” annotated in brat, highlighting both the tweet level (in red) and the sub-tweet levels sa annotations (in white). figure 3: example of a tweet annotated in brat. jussif stands for jussive, while propre and autre-jussif for proper assertive and other jussive respectively. the annotators performed a two-step annotation with an intermediate analysis of agreement and disagreement between the annotators. 448 tweets have been annotated in the first step by both annotators to compute the inter-annotator agreement (cohen’s kappa=0.62 for sa1 and 0.48 for sa217). this agreement exhibits a comparatively lower score than what is typically encountered in similar studies involving sa annotations in tweets, with for example 0.78 in vosoughi and roy (2016) and between 0.72 and 0.92 depending on task in subramanian et al. (2019). we found that it is mostly caused by the level of subjectivity involved in this task, in particular, the choice of the dominant segment, as mentioned earlier, has been the source of a lot of discrepancies. to address this issue, we encouraged regular feedback sessions and discussions between the two annotators to address discrepancies, clarify guidelines, and ultimately improve their agreement levels. another cause of disagreement were due to the difficulty of disentangling subjective from assertive, in particular when attitudes and modal expressions are used such as believe, think that, etc. indeed, both the subjective expressions (think, believe, or even more complex modal-tenseaspect combinations as fallait (which translates as ‘should have been’ with an additional implicature of preference in (17))) or its content can be targeted, according to their contextual relevance. (17) et maintenant il n’y a presque plus de fumée... il fallait arrêter le trafic ce matin et pas au milieu de la journée. ( and now there’s hardly any smoke... should have stopped the traffic this morning, not in the middle of the day.) sa1: subjective urgency: not urgent 16. http://brat.nlplab.org 17. we computed sa2 inter-annotation agreements on the basis of the dominant segment. 18 http://brat.nlplab.org digging communicative intentions 5. results of the annotation campaign we provide in this section a detailed analysis of the annotation campaign. we focus in particular on: (a) quantitative results of the sa annotations at both the tweet (sa1) and sub-tweet levels (sa2), (b) an analysis of how sa are expressed across different types of crisis, (c) the correlation between sa and urgency annotations, and finally (d) the evolution of sa over time since the crisis occurs. we end this section highlighting the main findings of this corpus-based study. 5.1 sa annotations: quantitative results table 2 shows the distribution of categories of sa1 annotations (i.e., tweet level). we observe that a majority of the tweets are classified as assertive, with 53.42%. the second-most frequent class is subjective, with 28.18% followed by jussive with 11.72%. interrogative and other are the less frequent with 3.36% and 3.32% respectively. these distributions indicate that in crisis situations, users predominantly tweet to assert their thoughts and views, to express their personal opinions and feelings, and to share information and updates on the given situation. conversely, the low percentage of jussive and interrogative suggests that they are less likely to give advice or ask questions in these circumstances (see also (zhang and liu, 2014)). assertive subjective jussive interrogative other total 7,147 3,770 1,568 449 444 13,378 (53.42%) (28.18%) (11.72%) (3.36%) (3.32%) table 2: frequency of tweet level (sa1) annotations. figure 4 provides the distribution of the sa2 dominant labels (i.e., the ones that drive the sa1 annotations). we observe that proper assertive is the most frequent with 37.19% while the other assertive sub-class, namely reported assertive, was dominant in 14.36% of the tweets. regarding non assertive content, evaluative and expressive sa2 annotations obtained similar frequencies of about 14.94% and 13.19% respectively. figure 5 combines the previous two tables illustrating the distribution of each sa2 sub-categories with their corresponding sa1 annotations. we observe that the pattern (sa1 = assertive, sa2 = proper assertive) is the most frequent with 72.13%. for interrogative, 72.58% of the segments are informative vs. 27.42% for uninformative while for jussive, 63.17% are open option vs. 36.83% other jussive. similar observations hold for the two remaining sa1 categories. finally, the very low percentage of other (i.e., 0.43%) suggests that annotators were able to easily associate a sa2 category to a given segment. this is not the case for sa1 annotations were this frequency increases to 3.32% showing that sub-level sa annotations are important to better capture users’ communicative intentions. the number of other sa2 annotations being relatively low (48 instances), we discard them for the further analysis below. when analyzing tweet segmentation for sa2 annotations (recall that sa2 annotations consist of a sequence of segments [s1, s2, . . . , sn], each with its associated sa2 category), we observe (see table 3) that, among the 11, 229 tweets annotated for sa2, only about 23% are made up of more than one segment. furthermore, 18.01% and 4.12% of tweets contain two and three segments respectively. while all sa2 classes display over 50% of presence in the first position, an interesting observation regarding the distribution of sa2 tags among possible positions is that it differs from class to class. notably, while proper assertive and reported assertive segments are over19 figure 4: distribution of sa2 annotations. figure 5: distribution of sa1 and sa2 annotations in our dataset. 20 digging communicative intentions whelmingly found in the first position (over 93%), all other classes display a much higher rate of non-first position in the sequence, ranging from 24,27% for informative to 44,60% for other jussive. table 3 together with table 4 that shows the distribution of the most frequent sequences within a tweet, suggest that relying only on the first label in the case of multi-label sequences might be a viable approach. however, this approach should consider two potential difficulties. first, it could introduce a bias in favor of the dominant class proper assertive, which tends to appear as the first element in multi-label sequences a lot more than the other classes in our data (about 98% of the time). second, a specific pattern is identified, where a proper assertive is followed by a different type of sa2 that is considered dominant, with the latter being in relation or in reaction to that initial assertion. in these cases, the reaction is the main, new, informative content that the rescue teams might be interested in, whereas the assertive content provides background information. when analyzing the data further, we indeed observe that, for tweets composed of two sequences, the forms [proper assertive, evaluative] and [proper assertive, expressive] are a majority with 414 and 388 tweets respectively, followed by [proper assertive, other jussive] and [proper assertive, open-option] with 189 and 169 tweets respectively. for tweets composed of three sequences, the patterns [proper assertive, evaluative, expressive] and [proper assertive, other jussive, evaluative] have been observed in 88 and 44 tweets respectively. examples (18) and (19) illustrate of the observed patterns. in (18), the interrogative is a direct follow-up to the assertion while in (19), the jussive is a reminder/directives given directly in reaction to the assertion. a final interesting observation concerns the other class, where proper assertive is not over-represented. this is likely due to the fact that this category is used to classify tweets that do not fit any of the other classes, and it should therefore be expected that such tweets follow a different pattern than the other classes. sa2 / position 1st 2nd 3rd 4th 5th total proper assertive 5,195 89 1 0 0 5,285 reported assertive 1,656 85 27 3 2 1,773 evaluative 1,092 774 183 18 4 2,071 expressive 1,272 548 204 42 0 2,066 other subjective 304 208 26 1 0 539 other jussive 422 372 30 8 2 834 open option 766 299 26 3 1 1,095 informative 252 91 26 3 3 375 uninformative 213 89 19 3 0 324 other 57 11 2 0 0 70 total 11,229 2,566 544 81 12 14,432 table 3: distribution of sa2 labels based on their position in the sequence. (18) [antony inondations du 11 juin arrêté de catastrophe naturelle enfin sorti ! les 700 habitations d’antony sinistrées, doivent s’attendre à d’autres désordres vue la situation climatique. le réservoir de fresnes serait sous dimensionné, un autre doit être construit,]1 [mais quand ?!]2 21 sub-tweet sa sequences % proper assertive 30.86 46.13proper assertive + other(s) sa2 15.27 eval./expr. 19.10 21.04eval./expr. + other(s) sa2 1.94 reported assertive 13.37 14.22reported assertive + other(s) sa2 0.85 open-option 6.15 6.84open-option + other(s) sa2 0.69 table 4: distribution of the most frequent sequences in sa2 annotations. ([ antony floods on june 11: natural disaster decree finally issued! the 700 antony homes affected by the floods can expect further disruption, given the weather situation. the fresnes reservoir is said to be undersized, and a new one is due to be built ]1 [ but when?! ]2) sa2: 1. assertive→ proper, 2. interrogative→ uninformative (19) [un été caniculaire de tous les dangers avec des incendies dans plusieurs régions.]1 [alors redoublons de prudence et de vigilance. pas de barbecues en forêt, de cigarettes allumées...]2 ([a dangerous hot summer with fires in several regions.]1[ so let’s be extra cautious. no barbecues in the forest, no lit cigarettes... ]2) sa2: 1. assertive→ proper, 2. jussive→ other jussive 5.2 sa annotations vs. crisis types our dataset is composed of 7 types of crisis, among them five 5 are unexpected or sudden events: floods, storms, hurricanes, building collapses, explosions and fires/ wildfires, and terrorist attacks. in this section, we analyze whether the type of crisis impacts the distribution of sa annotations. table 5 shows the results. overall, the distribution is quite similar across all crises (some more fine-grained observation will be provided in section 5.4), and are inline of those observed in tables 2 and 5. the only exception being the trèbes attack, with 39.36% assertives (the lowest frequency of assertives in the corpus) and 45.88% subjectives (the highest frequency). tweets posted during the sanary explosion displays the polar opposite distribution: 79.92% assertives (the highest frequency of assertives in the corpus) and 9.38% subjectives (the lowest frequency of subjectives in the corpus). those two events, despite having both resulted in several deaths and injuries, have, according to the difference in sa distribution, elicited vastly different reactions on twitter. a possible interpretation is that, in the case of the incident in sanary, users simply shared and discussed facts, as opposed to the terrorist attack in trèbes, where users expressed their emotions and sentiments. the types of crises seem to highlight certain tendencies related to sa1 annotations. for example, the distribution is quite similar between the 3 floods: with 3 of the highest numbers of assertives, averaging to 64.73%, and the 3 lowest numbers (besides sanary) of subjectives, averaging to 17.78%. similarly, the distribution for the 2 fires is such that both sub-corpuses display, by quite a margin (besides trèbes), the lowest numbers for assertives and the highest for subjectives with, respectively, 42.18% and 39.26%. finally, the frequency of others is consis22 digging communicative intentions % ass. % sub. %jus. % int. %oth. floods aude (1,002) 68.66 17.96 8.38 2.00 3.00 autre (1,001) 62.14 17.48 12.39 2.80 5.20 corse (404) 61.39 18.07 11.39 5.94 3.22 total (2,407) 64.73 17.78 10.55 2.99 3.95 storms beryl (320) 55.00 26.88 7.19 3.44 7.50 bruno (350) 57.43 26.29 4.86 4.86 6.57 susanna (391) 59.08 23.53 11.51 1.53 4.34 ulrika (316) 53.80 18.99 13.61 2.22 11.38 berguitta (330) 56.97 22.12 10.61 4.55 5.76 fionn corse (352) 67.33 19.60 7.95 1.70 3.41 egon (332) 55.72 28.31 7.23 3.31 5.42 eleanor (316) 66.14 21.84 8.23 1.90 1.90 total (2,707) 59.03 23.50 8.90 2.92 5.65 hurricane harvey (304) 55.59 19.08 11.84 7.57 5.92 irma (893) 54.88 28.22 10.96 4.03 1.91 total (1,197) 55.02 25.88 11.18 4.92 2.91 collapse marseille (304) 47.37 33.55 8.88 5.59 4.61 lille (549) 45.53 25.86 21.49 4.73 2.38 total (853) 46.14 28.59 16.98 5.04 3.25 incidents lubrizol (1,357) 48.77 23.12 18.05 4.57 0.59 sanary (533) 79.96 9.39 10.51 0.00 0.19 total (1,890) 61.01 19.24 15.92 3.27 0.48 fires landes (2,423) 42.51 39.99 12.55 3.76 1.20 notredame (519) 40.66 35.84 3.08 2.89 17.53 total (2,942) 42.18 39.26 10.88 3.60 4.08 attacks trèbes (1,382) 39.36 45.88 12.52 2.03 0.22 total total (13,378) 53.35 28.14 11.65 3.34 3.52 table 5: sa1 distribution per crisis type. tently very low for the whole corpus (averages 3.32%), with two exceptions: ulrika with 11.39% and notredame with 17.53%. finally, when looking into the distributions of sa2 annotations across crisis types, we observe that proper assertives are the most frequent first segment of the sequence for all the crises except the building collapse and terrorist attack where reported assertives were a majority. we also observe a high proportion of informatives, open options and evaluatives. the distributions for the storms, hurricanes, terrorist attack but also fire crises are different with more expressives than evaluatives. 5.3 sa vs. urgency annotations our dataset is annotated both for urgency and speech acts. all the tweets in our corpus (13, 378) have been annotated for sa1 and urgency (i.e. urgent, not urgent and not useful), whereas 11, 229 have been annotated tweets for sa1, sa2 and intentions to act, namely security instruction, human damage and infrastructure damage for urgent messages, and support, other and critics for not urgent messages. 23 sa1 vs. urgency. table 6 details the frequency of sa1 tags comparatively with the original urgency annotations. regarding the two most frequent sa1 (assertive and subjective), two observations emerge: (1) among 3,857 urgent messages (resp. 4,222 not urgent), 86.13% (resp. 33.82%) are assertive; and (2) only 5.81% of urgent messages are subjective while 44.69% of not urgent messages are. similarly, we observe that 6.82% of jussive are urgent vs. 15.99% not urgent. regarding not urgent messages, assertives mainly occur when messages contain information that is irrelevant to the crisis. it is interesting to note that the proportion of interrogatives are higher for not urgent messages when compared to the urgent ones (3.79% vs. 0.93%). finally, among the 444 messages that have been annotated as other messages, 81.08% are not useful. these frequencies are statistically significant using the χ2 test (χ2 = 2, 831.84, df = 8, p < 0.01). when measuring the dependency strength between urgency and sa1 categories using the cramer’s v test, we get (v = .32, df = 8) which confirms the statistical correlation between these two classifications. these observations indicate a strong correlation between assertivity and urgency when removing the not useful class (v = .54, df = 4). % urgent not urgent not useful total assertive 86.13 33.82 45.23 53.42 subjective 5.81 44.69 31.31 28.18 jussive 6.82 15.99 11.89 11.72 interrogative 0.93 3.79 4.77 3.36 other 0.31 1.71 6.79 3.32 total 100 100 100 100 table 6: urgency vs. sa1 annotation pairs statistics. % inf. hum. sec. sup. cri. oth. not. usf total assertive 86.87 90.52 83.33 21.80 17.30 59.21 46.66 54.56 subjective 7.08 6.21 5.04 70.34 70.75 10.31 31.52 26.95 jussive 4.49 2.61 10.35 7.10 2.52 20.93 11.43 11.22 interrogative 1.21 0.65 0.92 0.51 9.12 5.58 4.69 3.73 other 0.35 0.00 0.35 0.25 0.31 3.97 5.70 3.55 total 100 100 100 100 100 100 100 100 table 7: intention to act categories vs. sa1 annotations pairs statistics. table 7 provides the same analysis, this time with sa1 vs. intention to act annotations. for all urgent subcategories, assertive has the highest frequency with a total of 5,243 tweets, among them 90.52% are human damages, 86.87% infrastructure damages, and 83.33% security instructions. regarding the 2,416 not urgent messages, subjective make up 70.75% of critics and 70.34% of supports vs. 17.30% and 21.80% for assertives respectively. however, for the 1,309 other messages, which are not urgent messages that do not fall in either of the previous two categories, only 10.31% of them are classified as subjective, while 20.93% are jussives, and 59.21% assertives. these frequencies are statistically significant using the χ2 test (χ2 = 2, 502.17, df = 24, p < 0.01). when measuring the dependency strength 24 digging communicative intentions between intention and sa1 categories using the cramer’s v test, we get (v = .25, df = 24) which confirms the statistical correlation between these two classifications. sa2 vs. urgency. table 8 presents the frequency of sub-tweet sa tags (excluding other) when paired with urgency labels. in this table, the frequencies of sa2 are statistically significant (χ2 = 2, 378.84, df = 16, p < 0.01, v = .32), showing that sa2 annotations are of particular importance for urgency detection. % urgent not urgent not useful assertive reported assertive 52.73 24.60 40.21 proper assertive 33.32 9.69 8.53 subjective evaluative/expressive 5.34 42.20 27.75 other subjective 0.69 3.25 4.83 jussive open option 2.08 10.42 8.47 other jussive 4.20 5.44 3.92 interrogative informative 1.22 2.94 3.71 uninformative 0.24 0.92 1.68 table 8: urgency vs. sa2 annotations pairs (other sa2 tags have been removed). when looking into the distributions of sa2 tags against intentions to act categories (cf. table 9), we again observe an over-representation of proper assertive with a total of 3,133 instances, among them 153 are about infrastructure damages whereas 72 human damages. uninformative has the lowest frequency of 112 instances. overall, the relationship between sa2 and the urgency categories suggests that the degree of urgency of a message is correlated to the type of speech act used (χ2 = 2, 928.24, df = 42, p < 0.01, v = .25). the strength of the correlation increases to (v = 0.40, p < 0.01) when excluding the not useful. 5.4 evolution of speech acts over time recall that all the tweets in our dataset has been collected in three periods: 24h before, during (48h) and up to 72h after the crisis. our aim here is to analyze the evolution of speech acts over time focusing on three periods: before, during and after the event happened.18 table 10 shows the distribution of sa1 categories per period in terms of percentage. when looking at tweets over time since the crisis happens, we notice some interesting trends. before a crisis, tweets are a mix of assertions and to a little extent subjective content. during the crisis, tweets become more focused and include a lot of strong statements and questions, showing people intend to provide informative content and express opinions and evaluations. after the crisis, 18. this three time periods have been determined to better meet the french civil security and crisis management department’s specifications who perceives actionability in terms of emergency. 25 % inf. hum. sec. sup. cri. oth. not usf. assertive reported 26.97 40.82 9.81 3.5 3.18 13.84 8.39 proper 57.3 48.98 70.98 18.03 14.33 47.95 41.2 subjective eval/expr. 9.74 7.48 2.71 67.18 63.7 10.38 27.4 other-sub 0.37 0 0.84 3.63 07.01 0.4 5.32 jussive open option 1.87 0.68 1.88 3.5 0.00 14.96 8.13 other-juss 1.12 0.68 11.06 3.63 2.55 6.11 3.53 interrogative infor. 2.62 1.36 1.67 0.13 7.64 3.94 3.5 uninfor. 0.00 0.00 01.04 0.39 1.59 1.93 1.73 table 9: intentions to act categories vs. sa2 annotation pairs (percentage of each sa2 category per intention category). there is still a focus on sharing information, but fewer opinions are shared. after the crisis, assertive language remains substantial, suggesting a continued focus on conveying information. the proportion of subjective and interrogative tweets decreases post-crisis. this nuanced understanding highlights the shifting dynamics in communication styles across different phases of a crisis, with assertiveness and information-seeking becoming more pronounced during heightened situations. jussives are observed as more prominent before the crises happens, which is in line with the interpretation of the jussive: the speakers intend to enhance action most notably when preventing casualties is still possible, that is to say, before the crisis happens. % before during after assertive 10.51 26.05 15.89 subjective 7.68 15.01 6.57 jussive 6.24 2.64 2.8 interrogative 0.71 1.73 1.02 other 0.51 1.45 1.2 table 10: sa1 annotation vs. crisis period, in percentage. we further detail our analysis, this time by studying the distribution of sa1 per crisis type and period (see table 11).19 in the case of storms, floods and hurricanes, there is a notable surge in assertive messages, particularly after the event, indicating a shift toward providing clear information and directives to address the aftermath. concurrently, there is an increase in subjective expressions, possibly reflecting the emotional impact on individuals. likewise, collapse sees a notable increase in assertive messages post-crisis, suggesting a focus on clear statements and instructions once the immediate danger has passed. for explosion/attack and fire-related communication, there is a 19. we removed the other category from table 11. 26 digging communicative intentions significant uptick in assertive messages during the event, possibly aimed at providing immediate guidance. storm flood hurricane collapse explosion attack fire assertive before 2.32 0.87 2.28 0.28 0.00 0.00 0.28 during 4.04 1.33 0.99 0.60 9.47 4.47 9.62 after 6.75 4.33 2.14 2.36 0.01 0.00 0.30 subjective before 0.86 0.21 0.81 0.30 0.00 0.00 0.29 during 1.83 0.54 0.49 0.44 2.99 5.21 8.71 after 2.52 1.06 1.24 1.26 0.00 0.00 0.48 jussive before 0.37 0.11 0.47 0.09 0.00 0.00 0.18 during 0.74 0.27 0.15 0.25 2.47 1.42 2.36 after 0.87 0.51 0.48 0.85 0.00 0.00 0.09 interrogative before 0.11 0.07 0.21 0.07 0.00 0.00 0.02 during 0.23 0.09 0.03 0.09 0.51 0.23 0.78 after 0.30 0.21 0.24 0.20 0.00 0.00 0.07 table 11: sa1 annotation vs. crisis period vs. crisis type, in percentage. the two best scores for each period are in bold font. 5.5 interim conclusions the corpus-based study of speech acts in tweets annotated for urgency allows for multiple statistically relevant observations: • the vast majority of tweets are assertives, seconded by subjectives. more specifically, proper assertives is the dominant class at the sub-tweet level. these results seem to indicate that, in a reaction to a crisis, french twitter users mostly tweet to share information, generally in the form of a single utterance. this corroborates the findings in (zhang and liu, 2014), tending to show that in an emergency, factual information is more relevant than the expression of a personal view-point. • proper assertives are over-represented in the first position in every sa1 category of tweets. in particular, the high frequency of proper assertives in the interrogative, jussive and subjective tweets is explained by the fact that a significant part of those tweets follow a format comprising an assertion, followed by the speaker’s reaction to said assertion, which constitutes the dominant sa, as shown in (16). this reveals an interesting finding: in crisis situations, speakers tend to assert or re-assert a piece of (already known) information, followed by their personal comment in relation to it, thus sharing their perspective. 27 • the distribution of sa1 annotations highlights a general consistency in the data across the different crises, as well as similarities in the sa distribution of similar crises. finally, we found a statistically significant relationship between assertivity and urgency, and between subjectivity and absence of urgency. the picture that emerges, is one on which speakers favor (what they consider) truthful information over orders and commands to enhance action (on the part of the rescuing teams, for instance). indeed, in our classification assertives do not include subjective evaluations, and thus convey content informationally reliable and objectively veridical (i.e. conform to the outer reality and not a mental state) (giannakidou and mari, 2017, 2018, 2021c) and thus ready for uptake and endorsement (e.g. ginzburg (2012), krifka (2019)) on the part of those who will bring help. the fact that speakers favor proper assertives to indicate urgency reveals that they are fully committed to the truthfulness of the message, of which they might present themselves as the primary informational source. on the contrary, we observe that subjectives correlate with absence of urgency. among subjectives evaluatives/expressives are largely used to convey truths that are relativized to a ‘judge’ or an individual (a.o. (lasersohn, 2005; stephenson, 2007)) and are not eligible to function as reliable information for the rescuing services. a minority of subjectives encompass attitudes, whereby truth is also relativized to a particular mental state and cannot (without further negotiation) immediately become common ground (e.g., (gunlogson, 2008; mari and portner, 2021)) and be ready for uptake on the part of the helpers. collapses, while qualifying as sudden crises, behave like non-sudden ones, probably in virtue of the long searches for casualties that make them similar to non-sudden ones. finally, we have discovered that assertives are more prominent after the crises when these are non-sudden, and during the crises when these are sudden. this points to the fact that speakers are active in providing information during the aftermaths of non-sudden crises, which, most of the times, require sustained efforts in view of the intensity of the damages. speakers are keener in using assertives during the crisis with sudden crises, aiming at providing contentful information as the unexpected crisis unfolds and no knowledge had been made previously available by media or other sources. 6. automatic detection of sa 6.1 experimental settings now the dataset has been annotated, the next step is to automatically detect sa. we cast the problem into a classification task, leaving the complex task of discourse-based tweet segmentation into non overlapping units to future work (morabia et al., 2019; aljebreen et al., 2021). we propose the following experimental settings: • sa1 detection: classify each tweet into one of our five sa1 categories, namely assertive, subjective, interrogative, jussive, and other. • sa2 detection: classify each tweet into one of our eight sa2 categories. note that the other instances (48 tweets) have been removed from the dataset for the experiments as they are very less frequent in urgent tweets and have no regular linguistic patterns. we propose two settings: 28 digging communicative intentions – multi-class classification. given a tweet t = [s1, . . . , sn] and its associated sa annotations sa1 = [sa21, . . . , sa2j , . . . , sa2n] where sa2j is the dominant segment, predicts its sa2 category sa2pred. we evaluate the results considering (i) a strict match where sa2pred = sa2j (this is similar to a binary classification), as well as (ii) a partial match such that sa2pred ∈ {sa21, . . . , sa2j , . . . , sa2n}. – multi-label classification. multi-class classification only focuses on the dominant segment ignoring the speech acts conveyed by the other tweet segments (we recall that a tweet can be composed up to 5 segments, see table 3). we believe these informations can be of particular importance for urgency detection. therefore, we aim at capturing label dependencies among the segments by assigning multiple labels for each instance simultaneously (zhang and zhou, 2013; liu et al., 2021). • detecting sa1 and sa2 simultaneously. this is a multi-task learning framework considering there are two classification tasks (sa1 and sa2). the classifiers for both tasks share and update the same low layers except the final task-specific classification layer. 6.2 models we rely on flaubertbase(le et al., 2019) the base cased french transformers models (martin et al., 2020) pre-trained on french texts from various sources from the general domain (e.g., wikipedia and books), as implemented in huggingface. in addition, we use flauberttuned, a flaubert model that was pre-trained on 358,834 unannotated tweets from the crisis domain (kozlowski et al., 2020) achieving better performances compared to flaubert for urgency detection. we also experimented with camembertbase (martin et al., 2019), the other french transformer architecture.20 following laurenti et al. (2022b), we also experiment with two multi-input models that use extra-features added on top of pre-trained contextual word embeddings, among which21: the presence of urls, punctuation (exclamation marks and question marks) and the presence of numbers, as they are often used in tweets to indicates phone numbers of emergency rescue services or weather forecast. we refer to these models as flauberttuned+feat and flaubertbase+feat. in addition to the cross entropy loss (hereafter +c) and to fight class imbalance, we consider the focal loss (hereafter +focal) (lin et al., 2017) or weighted cross entropy (hereafter +w). our aim here is to compare with one of the most effective approach for handling imbalanced data (cui et al., 2019). all our models were trained for four epochs with a learning rate of 2e− 5, on top of which a linear layer for classification was added. for better convergence, we use the adam optimizer during backpropagation. to avoid exploding gradients, we use a gradient clipping of 1.0. for the multi-label task, we adapt the flaubert architecture22 to account for multilabel outputs relying on a sequence classification head on top of the pool layer. the input sequence comprises characters, sub-words, and words, which are processed by the transformer layers. on top of the pooled output, a linear layer is added for the classification task. we then examine each label independently for every message and determine whether the label is predicted by the model. we rely on label-based metrics (f1 macro) following the general trend in multi-label classification (zhang and 20. the majority baseline achieved an accuracy of 0.534 and 0.372 for sa1 and sa2 classifications, respectively. 21. we also tested several other features including tweet meta-features, sentiment and emoticons, number of imperatives verbs, etc., but the results were not conclusive. 22. we used flaubert as the results achieved by camembert on this task were lower. 29 zhou, 2013). this architecture was successfully employed in multilabel classification in other nlp tasks including judicial documents (dai and liu, 2020), sentiment analysis (tang et al., 2020) and diagnoses of patients prediction (hart, 2022). 6.3 evaluation protocol to evaluate sa1 and sa2 models, we designed two evaluation protocols: • random sampling. we mixed the tweets for all the crisis and randomly select 80% for train and 20% for test. for sa1 classification (resp. sa2), the final dataset is composed of 11, 181 tweets (resp. 13, 378) split into 80-20 for train-test while keeping the same distribution as in the train set. • out-of-event. following the general trends in crisis management (kersten et al., 2019; algiriyage et al., 2021; bourgon et al., 2022a), we designed an out-of-type evaluation protocol by training on a pool of events related to different types of crises (e.g., hurricane, storm) and testing on a particular different type (e.g., earthquake). the aim is to evaluate if a model can deal with new types of crisis, which is crucial to ensure the portability of the models to unseen events. to this end, we consider the distinction between expected vs. sudden events and experiment whether the use of speech acts differ according to the type of crises. indeed, compared to ecological disaster like hurricanes and floods, sudden events (like earthquakes, terror attacks, explosions, technological incidents) are difficult to predict (björck, 2016). these events, over which organizations have virtually no control, influence social behavior and the ways the emergency services are organized (james and wooten, 2005; coombs, 2014; quarantelli et al., 2017). we propose two evaluation settings: (a) train on expected events and test on sudden, (b) train on sudden and test on expected. we consider flood, storm and hurricane as expected events while collapse, wildfire, plant explosion and attack as sudden events (see table 1), which corresponds to a total of 6,311 tweets for the former vs. 7,067 for the latter. all the sa1 (resp. sa2) models have been run five times on a randomly selected instances from the test set with a standard deviation of results being 5.3 × 10−6 (resp. 7.8 × 10−4). we therefore report the averaged scores (accuracy, precision, recall, and macro f1). finally, due to the high number of experiments, we only provide those achieved by the best configurations. below we present our results. we end by a qualitative analysis highlighting main causes of misclassification. 6.4 results 6.4.1 random sampling results the experimental results are presented in table 12, showcasing the accuracy (a), precision (p), recall (r), and macro-averaged f1-scores (f1). the results are grouped according to mono vs. multitask learning and whether these models use extra-features. best scores are highlighted in bold font. the results show that flauberttuned has consistently achieved the best scores across all settings and that mono-task learning models outperform its multitask counterpart. however, it is interesting to note that flaubertmultitasktuned+c+feat resulted in the best accuracy of 81.81%. injecting additional features was very helpful when coupling with the focal loss for flauberttuned+focal+feat and 30 digging communicative intentions models a p r f1 camembertbase+w 79.04 69.20 65.67 66.93 flaubertbase+focal 78.29 69.53 61.88 64.96 flauberttuned+c 79.15 68.87 65.19 66.81 flaubertbase+c+feat 78.66 70.95 62.79 65.58 flauberttuned+focal+feat 78.59 71.06 65.02 67.37 flaubertmultitaskbase+c 62.98 53.01 49.33 49.76 flaubertmultitasktuned+c 80.69 60.69 57.57 58.98 flaubertmultitaskbase+c+feat 79.17 59.43 56.23 57.65 flaubertmultitasktuned+c+feat 81.81 61.21 58.28 59.60 table 12: sa1 classification results. cross entropy for flaubertmultitasktuned+c+feat. when looking into the detailed results per class (cf. table 13), we observe that the predictions are closely aligned with the distribution of each class in the dataset. in particular, assertive and subjective were well-predicted with an f-score of 85.32 and 76.14 respectively, whereas jussive and interrogative exhibit lower scores. precision recall f1-score assertive 84.46 86.19 85.32 subjective 75.47 76.82 76.14 jussive 64.26 64.47 64.37 interrogative 63.24 55.84 59.31 other 67.86 41.76 51.70 accuracy = 78.59 table 13: sa1 results per class as given by flauberttuned+focal+feat our best model. level models a p r f1 monotask strict flaubertbase+c 66.58 56.76 52.06 52.79 flauberttuned+w 67.72 59.56 56.98 57.82 flaubertbase+c+feat 65.49 57.55 51.88 51.90 flauberttuned+w+feat 67.95 56.90 55.12 55.44 monotask partial flaubertbase+c 73.71 63.46 60.23 60.42 flauberttuned+w 74.74 67.91 64.74 65.67 multitask flaubertmultitaskbase+c 59.83 48.64 39.69 41.15 flaubertmultitasktuned+c 67.59 59.76 53.09 55.17 flaubertmultitaskbase+c+feat 65.71 52.78 50.53 50.05 flaubertmultitasktuned+c+feat 67.99 60.31 53.51 54.72 multi-label flaubertbase+c 94.68 96.35 82.85 87.80 flauberttuned+c 88.11 78.17 57.50 62.36 table 14: sa2 classification results. 31 the results of the fine-grained sa experiments are presented in table 14. the results indicate that partial evaluation leads to an improvement of approximately 8% in terms of accuracy with flauberttuned+w achieving the best performance, resulting in an f-score of 65.67. this shows that a strict evaluation is not suitable to determine the dominant segment which is predictable given the pragmatic nature of selecting these types of segment. regarding multitask architectures, the results are inline with those observed with sa1, making them less productive for multi-level sa detection. more importantly, the multi-label classification was the best, with flaubertbase+c yielding the highest scores. finally, although features have been very productive for sa1 classification, their injection into the flaubert architecture for sa2 achieved mitigated results, see for example the boost in the f1-scores achieved by flaubertmultitaskbase+c+feat vs. flaubertmultitaskbase+c while we observe a drop when comparing flaubertmultitasktuned+c vs. flaubertmultitasktuned+c+feat. we end this section by detailed results per class, as given by the multi-label model (cf. table 15). overall, the model achieves very good results for all the classes except the less frequent (see for example other subjective and uninformative). precision recall f1-score reported assertive 100 96.78 98.37 proper assertive 91.27 99.81 95.35 expressive/evaluative 98.22 93.96 96.02 other-subjective 100 43.14 60.27 open-option 100 92.36 96.03 other-jussive 98.36 90.91 94.49 uninformative 100 64.29 78.26 informative 81.13 70.49 75.44 accuracy = 94.68 table 15: sa2 results per class as given by the multi-label flaubertbase+c. 6.4.2 out-of-event results table 16 shows the results of our best sa1 (resp. sa2) models when tested following the out-ofevent protocol, i.e., flauberttuned+focal+feat (resp. flaubertbase+c), in terms of precision (p), recall (r) and the averaged f1-score (f1). for sa1, and when compared to random sampling, we observe a small drop in the performances and this is more salient when the model is trained on sudden events. when we look into the results per class, we notice that three out of the five sa1 categories achieved similar results when trained vs. tested on expected events: 82.30 vs. 85.40 f1-score for assertive, 51.58 vs. 54.16 for interrogative, and 59.76 vs. 60.11 for jussive. note that these scores are close to the one reported in table 13 where the sa1 model has been evaluated in a random sampling scenario. other and subjective however exhibits a different behavior: other scores 44.44 vs. 31.58 resulting in an important decrease in performances up to 7.3% in terms of f1-score when compared to random sampling. for subjective, the drop depends on the test set: when trained on expected events and tested on sudden, this category achieved around −6% compared to a random test. on the other hand, when trained on sudden and tested on expected events, the scores were similar (76.85 vs. 76.14 in the random configuration). 32 digging communicative intentions this drop can be explained by the diverse linguistics means by which speech acts are expressed in sudden events (we recall the the distribution of each class in both settings are quite similar (see table 5)). the subjective class encompasses a series of expressions that belong to different grammatical categories (verbs, adjectives, particles, interjections, ...) at different levels of the semantic and pragmatic interpretation (sentence, discourse, ...) and can either have an informational or just an expressive function. a finer grained typology of expressions is to be established to pin down the linguistic differences between the subjective expressions involved in sudden crises and those used in non-sudden crises situations. finally, regarding the sa2 results, we notice that testing on expected (resp. sudden) events does not significantly impact the results when compared to the random sampling. this shows that casting sa2 detection as a multi-labeling problem is quite effective. random expec. → sudden sudden → expec. f1 p r f1 p r f1 sa1: flauberttuned+focal+feat 67.37 68.18 59.81 62.98 65.86 60.88 60.32 sa2: flaubertbase+c 87.80 94.62 76.90 82.35 94.35 83.23 87.47 table 16: sa1 (resp. sa2) best models results when tested in the out-of-event protocol. 6.5 error analysis we end this paper by analyzing most causes of misclassifications as given by our best sa1 (resp. sa2) models, namely flauberttuned+focal+feat, the fine-tuned flaubert with focal loss and feature injection (resp. the multi-label flaubertbase+c trained with a cross entropy loss). figure 6 presents our results. we also provide the confusion matrix for sa1 classification (see table 17).23 it shows that most errors come from the difficult distinction between assertive and subjective, as also observed during the annotation campaign (see section 4.2). in practice, the objectivesubjective distinction may not always be clear-cut, leading to a preference for assertive classification, see the examples (9) and (10) in table 18. we also observe other complex cases like in (2), (3), (7) and (8) that contain declarative statements and for which the model fails to distinguish between assertives and jussives. some other examples lack context, and the gold label may not always be accurate (e.g., (1)), but they can still express assertive sentiments (e.g., (4)). finally, the model struggles with some interrogative texts that are phrased declaratively (e.g., (5)) or affirmatively (e.g., (6)), leading to difficulties in identifying them as interrogative, therefore the model has trouble taking into consideration the interrogative mark present at the end of the text. assertive subjective jussive interrogative other assertive 0 119 66 6 8 subjective 135 0 31 14 2 jussive 59 35 0 6 8 interrogative 13 19 2 0 0 other 18 23 10 2 0 table 17: sa1 misclassification results. 23. we only report the confusion matrix for sa1 classification as the one for sa2 is too sparse. 33 figure 6: distribution of misclassified examples. tweet predicted gold 1 tempête #nekfeu (storm #nekfeu) oth. assert 2 aurore bergé sur les inondations: ”il faut donner des cours de natation aux français” (aurore bergé on the floods: ”we need to give swimming lessons to the french.”) juss assert 3 regarde le prix des billets pour la guadeloupe (look at the price of tickets to guadeloupe) juss assert 4 ouragan #irma : l’inquiétude des antillais de métropole #bourdindirect (hurricane #irma: concerns of antilleans in mainland france #bourdindirect) assert oth. 5 j’suis le seul a quitter une reunion en plein milieu parce qu’elle ne m’intéresse pas? (am i the only one who leaves a meeting in the middle because it doesn’t interest me?) assert inter 6 toulouse prêt pour le déluge ? (is toulouse ready for the deluge?) assert inter 7 #rouen sous la fumée noire de l’entreprise #lubrizol :: url (#rouen under the black smoke of the company #lubrizol :: url ) assert juss 8 les mégapoles face aux risques d’inondation, passé / présent (megacities facing flood risks, past and present.) assert juss 9 journée noire à #rouen après l’incendie de l’usine #lubrizol #pollution url (black day in #rouen after the fire at the #lubrizol factory #pollution url ) assert subj 10 nan mais le temps ici y’a deux secondes c’était inondation et là y’a du soleil (but the weather here, two seconds ago, was flood and now it’s sunny.) assert subj table 18: sa1 misclassified examples. for sa2, more than 84% of the misclassified tweets were due to the proper assertive class, followed by informative. table 19 presents some representative examples of such errors. it appears that the model fails to distinguish between proper assertives and subjective (expressive or evaluative) content, as seen in examples (1) and (2). we observed that the classifier is misguided by the interrogation (e.g., (7)) and exclamation (e.g., (5) and (6)) marks. other examples lack context and present some grammatical errors (e.g., (4)). finally, the model is challenged with some of the interrogative sentences, one possible reason is that these often have a different syntactic structure and may resemble declarative sentences (e.g., (9) and (10)). 7. conclusion in this paper, we presented the first corpus-based study to measure the impact of speech acts in messages posted in social media during various types of crises. we first proposed a new annotation guideline to annotate speech acts both at the tweet and sub-tweet levels, then a new dataset annotated 34 digging communicative intentions text predicted gold 1 ce que l’on sait de l’incendie à rouen qui ravage l’usine lubrizol url (rouen’s lubrizol factory fire: what we know url ) oth-juss proper 2 regarder la nuit tomber c’est dans mon top 3 des activités (watching the night fall is in my top 3 activities.) proper expr./eval. 3 calmez vraiment la tempête (really calm the storm.) expr./eval. oth-subj 4 t’es bon pour intéresser aussi au papi (programmes d’actions de prévention des inondations). (you’re good to also be interested in the papi (programs of action for flood prevention).) info expr. 5 wow! rt @pellepx3: les premières images de l’inondation à montréal le 29 mai #inondation (wow! rt @pellepx3: the first images of the flooding in montreal on may 29 #flooding) info eval 6 @dragonduclos qui sème le vent récolte la tempête ! (@dragonduclos who sows the wind reaps the storm!) info oth-subj 7 user et toi tu t’effondres quand sous le.poids de tes âneries ? (and you, do you collapse under the weight of your nonsense?) info oth-subj 8 une sinistrée des inondations demande que macron apporte ”des solutions” (a flood victim asks macron to bring ”solutions”.) proper rapported 9 hello @bfmtv ça vous intéresse des vidéos de l’incendie de la forêt de brocéliande faites avec un drone ??? (hello @bfmtv, are you interested in videos of the fire in the brocéliande forest made with a drone???) oth-juss info 10 le rétablissement de la continuité écologique des zones humides ne fait-il pas oublier le risque d’inondation ? (does the restoration of ecological continuity in wetlands not forget the risk of flooding?) proper uninfo table 19: sa2 misclassified examples. for both speech acts and urgency categories in french. we conducted a deep corpus-based analysis of the correlation between sa and urgency, sa and intention to act categories, sa and crisis type, and finally, sa evolution over time since the event happens. our results show a strong correlation (i) between assertive messages (particularly those that rely on first hand knowledge, i.e. proper assertives) and urgency, (ii) subjective messages and absence of urgency, with a high frequency of expressives and evaluatives. in addition, we found a strong correlation between assertives and human/infrastructure damages. we finally conducted a set of experiments to detect sa relying on transformer architectures augmented with dedicated features. we propose a set of monotask and multitask learning settings to classify a given tweet in either tweet-level speech act or sub-tweet level speech act categories casting the problem into a multi-label and multi-class task. we also experiment models portability to unseen events to measure sa detection performances in real scenario. our results are encouraging and constitute a new state of the art of speech act detection in french tweets. the next step now is to inject sa information while detecting urgency. our preliminary study on sa-aware urgency detection were very encouraging (laurenti et al., 2022b). these experiments were however conducted on a small set of 6,6k tweets and only focusing on tweet-level speech acts. we plan to extend this work by injecting sub-tweet speech acts as well through new deep learning architectures in a multilingual setting. acknowledgment this work has been supported by the intact project funded by a cnrs pre-maturation grant. alda mari gratefully acknowledges anr-17-eure-0017 frontcog. the research of farah benamara is also partially supported by descartes: the national research foundation, prime minister’s office, singapore under its campus for research excellence and technological enterprise (create) program. 35 authors contributions the order of the authors is alphabetical to reveal contributions of comparable importance. farah benamara and alda mari are the project leaders, and have furthermore coordinated the research and the writing. references alexandra aikhenvald. evidentiality. oxford university press, kettering, northamptonshire, uk, 2004. firoj alam, ferda ofli, and muhammad imran. crisismmd: multimodal twitter datasets from natural disasters. arxiv:1805.00713 [cs], may 2018. url http://arxiv.org/abs/1805. 00713. firoj alam, hassan sajjad, muhammad imran, and ferda ofli. crisisbench: benchmarking crisisrelated social media datasets for humanitarian information processing. in proceedings of the international aaai conference on web and social media, volume 15 of icwsm ’21, pages 923–932, may 2021. url https://ojs.aaai.org/index.php/icwsm/article/ view/18115. nilani algiriyage, rangana sampath, raj prasanna, emma eh doyle, kristin stock, and david johnston. identifying disaster-related tweets: a large-scale detection model comparison. in social media in crises and conflicts, proceedings of the 18th iscram conference, pages 731–743, 2021. url http://idl.iscram.org/files/nilanialgiriyage/2021/2368_ nilanialgiriyage_etal2021.pdf. bushra algotiml, abdelrahim elmadany, and walid magdy. arabic tweet-act: speech act recognition for arabic asynchronous conversations. in proceedings of the fourth arabic natural language processing workshop, pages 183–191, florence, italy, august 2019. association for computational linguistics. doi: 10.18653/v1/w19-4620. url https://aclanthology.org/ w19-4620. alaa alharbi and mark lee. crisis detection from arabic tweets. in proceedings of the 3rd workshop on arabic corpus linguistics, pages 72–79, 2019. url https://www.aclweb. org/anthology/w19-5609. abdullah aljebreen, weiyi meng, and eduard dragut. segmentation of tweets with urls and its applications to sentiment analysis. proceedings of the aaai conference on artificial intelligence, 35(14):12480–12488, may 2021. doi: 10.1609/aaai.v35i14.17480. url https: //ojs.aaai.org/index.php/aaai/article/view/17480. james allen and mark core. draft of damsl: dialog act markup in several layers. 1997. url http://www.fb10.uni-bremen.de/anglistik/ling/ss07/ discourse-materials/damsl97.pdf. nicholas asher and alex lascarides. commitments, beliefs and intentions in dialogue. in proceedings of the 12th workshop on the semantics and pragmatics of dialogue, pages 29–36. citeseer, 2008. 36 http://arxiv.org/abs/1805.00713 http://arxiv.org/abs/1805.00713 https://ojs.aaai.org/index.php/icwsm/article/view/18115 https://ojs.aaai.org/index.php/icwsm/article/view/18115 http://idl.iscram.org/files/nilanialgiriyage/2021/2368_nilanialgiriyage_etal2021.pdf http://idl.iscram.org/files/nilanialgiriyage/2021/2368_nilanialgiriyage_etal2021.pdf https://aclanthology.org/w19-4620 https://aclanthology.org/w19-4620 https://www.aclweb.org/anthology/w19-5609 https://www.aclweb.org/anthology/w19-5609 https://ojs.aaai.org/index.php/aaai/article/view/17480 https://ojs.aaai.org/index.php/aaai/article/view/17480 http://www.fb10.uni-bremen.de/anglistik/ling/ss07/discourse-materials/damsl97.pdf http://www.fb10.uni-bremen.de/anglistik/ling/ss07/discourse-materials/damsl97.pdf digging communicative intentions john langshaw austin. how to do things with words. oxford university press, kettering, northamptonshire, uk, 1962. kent bach and robert m harnish. linguistic communication and speech acts. mit press, 1979. laura beth bell. illocution on twitter: the construction and analysis of a social media speech act corpus. georgetown university, washington, d.c., usa, 2020. albena björck. crisis typologies revisited: an interdisciplinary approach. central european business review, 2016(3):25–37, 2016. doi: 10.18267/j.cebr.156. url https://ideas. repec.org/a/prg/jnlcbr/v2016y2016i3id156p25-37.html. nils bourgon, farah benamara, alda mari, véronique moriceau, gaetan chevalier, and laurent leygue. are sudden crises making me collapse? measuring transfer learning performances on urgency detection. in 19th international conference on information systems for crisis response and management (iscram 2022), 2022a. url https://ut3-toulouseinp. hal.science/hal-03707241/document. nils bourgon, farah benamara, alda mari, véronique moriceau, gaetan chevalier, laurent leygue, and yasmine djadda. are sudden crises making me collapse? measuring transfer learning performances on urgency detection. in rob grace and hossein baharmand, editors, 19th international conference on information systems for crisis response and management, iscram 2022, tarbes, france, may 22-25, 2022, pages 701–709. iscram digital library, 2022b. url https://idl.iscram.org/show.php?record=2449. david bracewell, marc tomlinson, and hui wang. identification of social acts in dialogue. in proceedings of coling 2012, pages 375–390, mumbai, india, december 2012. the coling 2012 organizing committee. url https://aclanthology.org/c12-1024. robert brandom. making it explicit: reasoning, representing, and discursive commitment. harvard university press, cambridge, ma, usa, 1994. anaı̈s cadilhac, nicholas asher, farah benamara, and alex lascarides. grounding strategic conversation: using negotiation dialogues to predict trades in a win-lose game. in proceedings of the 2013 conference on empirical methods in natural language processing, pages 357– 368, seattle, washington, usa, october 2013. association for computational linguistics. url https://aclanthology.org/d13-1035. vitor r. carvalho and william w. cohen. on the collective classification of email ”speech acts”. in proceedings of the 28th annual international acm sigir conference on research and development in information retrieval, sigir ’05, page 345–352, new york, ny, usa, 2005. association for computing machinery. isbn 1595930345. doi: 10.1145/1076034.1076094. url https://doi.org/10.1145/1076034.1076094. carlos castillo. big crisis data: social media in disasters and time-critical situations. cambridge university press, cambridge, england, 2016. isbn 9781316476840. doi: 10.1017/ 9781316476840. 37 https://ideas.repec.org/a/prg/jnlcbr/v2016y2016i3id156p25-37.html https://ideas.repec.org/a/prg/jnlcbr/v2016y2016i3id156p25-37.html https://ut3-toulouseinp.hal.science/hal-03707241/document https://ut3-toulouseinp.hal.science/hal-03707241/document https://idl.iscram.org/show.php?record=2449 https://aclanthology.org/c12-1024 https://aclanthology.org/d13-1035 https://doi.org/10.1145/1076034.1076094 christophe cerisara, somayeh jafaritazehjani, adedayo oluokun, and hoa t. le. multi-task dialog act and sentiment recognition on mastodon. in proceedings of the 27th international conference on computational linguistics, pages 745–754, santa fe, new mexico, usa, august 2018. association for computational linguistics. url https://aclanthology.org/c18-1063. alfredo cobo, denis parra, and jaime navón. identifying relevant messages in a twitter-based citizen channel for natural disaster situations. in proceedings of the 24th international conference on world wide web, www’15, pages 1189–1194, 2015. dario compagno, elena v. epure, rebecca deneckere, and camille salinesi. exploring digital conversation corpora with process mining. corpus pragmatics, 2:193–215, 2018. doi: 10.1007/ s41701-018-0030-6. url https://hal.univ-lorraine.fr/hal-01722928. cleo condoravdi and sven lauer. imperatives: meaning and illocutionary force. empirical issues in syntax and semantics, 9:37–58, 2012. w. t. coombs. ongoing crisis communication: planning, managing, and responding. thousand oaks, ca: sage., 2014. mark core, masato ishizaki, johanna moore, christine nakatani, norbert reithinger, david traum, and syun tutiya. the report of the third workshop of the discourse resource initiative, chiba university and kazusa academia hall. technical report, 1998. stefano cresci, maurizio tesconi, andrea cimino, and felice dell’orletta. a linguistically-driven approach to cross-event damage assessment of natural disasters from social media messages. in proceedings of the 24th international conference on world wide web, www’15, pages 1195– 1200, 2015. yin cui, menglin jia, tsung-yi lin, yang song, and serge belongie. class-balanced loss based on effective number of samples. in proceedings of the ieee/cvf conference on computer vision and pattern recognition (cvpr), june 2019. mian dai and chao-lin liu. multi-label classification of chinese judicial documents based on bert. in 2020 ieee international conference on big data (big data), pages 1866–1867. ieee, 2020. abdelrahim elmadany, hamdy mubarak, and walid magdy. arsas: an arabic speech-act and sentiment corpus of tweets. in hend al-khalifa, king saud university, ksa walid magdy, university of edinburgh, uk kareem darwish, qatar computing research institute, qatar tamer elsayed, qatar university, and qatar, editors, proceedings of the eleventh international conference on language resources and evaluation (lrec 2018), paris, france, may 2018a. european language resources association (elra). isbn 979-10-95546-25-2. abdelrahim elmadany, hamdy mubarak, and walid magdy. arsas: an arabic speech-act and sentiment corpus of tweets. osact, 3:20, 2018b. tin franovic and jan šnajder. speech act based classification of email messages in croatian language. 2012. anastasia giannakidou and alda mari. mixed (non) veridicality and mood choice with emotive verbs. in cls 51, 2015. 38 https://aclanthology.org/c18-1063 https://hal.univ-lorraine.fr/hal-01722928 digging communicative intentions anastasia giannakidou and alda mari. epistemic future and epistemic must: nonveridicality, evidence, and partial knowledge. in mood, aspect, modality revisited, pages 75–118. university of chicago press, chicago, il, usa, 2017. anastasia giannakidou and alda mari. a unified analysis of the future as epistemic modality. natural language & linguistic theory, 36(1):85–129, 2018. anastasia giannakidou and alda mari. a linguistic framework for knowledge, belief, and veridicality judgment. know: a journal on the formation of knowledge, 5(2):255–293, 2021a. anastasia giannakidou and alda mari. modalization and bias in questions. university of chicago and insitut jean nicod, 2021b. anastasia giannakidou and alda mari. truth and veridicality in grammar and thought: mood, modality, and propositional attitudes. university of chicago press, chicago, il, usa, 2021c. jonathan ginzburg. the interactive stance. oxford university press, kettering, northamptonshire, uk, 2012. oliveira hugo gonçalo, patrı́cia ferreira, daniel martins, catarina silva, and ana alves. a brief survey of textual dialogue corpora. in proceedings of the thirteenth language resources and evaluation conference, pages 1264–1274, marseille, france, june 2022. european language resources association. url https://aclanthology.org/2022.lrec-1.135. chih-wen goo and yun-nung chen. abstractive dialogue summarization with sentence-gated modeling optimized by dialogue acts. in proceedings of 7th ieee workshop on spoken language technology, 2018. christine gunlogson. a question of commitment. belgian journal of linguistics, 22(1):101–136, 2008. charles l hamblin. fallacies. tijdschrift voor filosofie, 33(1), 1970. hanna’t hart. predicting diagnoses of patients in the emergency room: a multi-label text classification approach. master’s thesis, 2022. ryuichiro higashinaka, kenji imamura, toyomi meguro, chiaki miyazaki, nozomi kobayashi, hiroaki sugiyama, toru hirano, toshiro makino, and yoshihiro matsuo. towards an open-domain conversational system fully based on natural language processing. in proceedings of coling 2014, the 25th international conference on computational linguistics: technical papers, pages 928–939, dublin, ireland, august 2014. dublin city university and association for computational linguistics. url https://aclanthology.org/c14-1088. muhammad imran, shady mamoon elbassuoni, carlos castillo, fernando diaz, and patrick meier. extracting information nuggets from disaster-related messages in social media. proc. of iscram, baden-baden, germany, 2013. muhammad imran, prasenjit mitra, and carlos castillo. twitter as a lifeline: human-annotated twitter corpora for nlp of crisis-related messages. in proceedings of the tenth international conference on language resources and evaluation (lrec 2016), paris, france, 2016. european language resources association (elra). isbn 978-2-9517408-9-1. 39 https://aclanthology.org/2022.lrec-1.135 https://aclanthology.org/c14-1088 erika james and lynn wooten. leadership as (un)usual: how to display competence in times of crisis. organizational dynamics, 34(2):141–152, 2005. shafiq joty and tasnim mohiuddin. modeling speech acts in asynchronous conversations: a neuralcrf approach. computational linguistics, 44(4):859–894, december 2018. doi: 10.1162/coli a 00339. url https://aclanthology.org/j18-4012. marc-andré kaufhold, markus bayer, and christian reuter. rapid relevance classification of social media posts in disasters and emergencies: a system and evaluation featuring active, incremental and online learning. information processing & management, 57(1):102–132, 2020. simon keizer, rieks op den akker, and anton nijholt. dialogue act recognition with bayesian networks for dutch dialogues. in proceedings of the third sigdial workshop on discourse and dialogue, pages 88–94, philadelphia, pennsylvania, usa, july 2002. association for computational linguistics. doi: 10.3115/1118121.1118134. url https://aclanthology.org/ w02-0213. mayank kejriwal and peilin zhou. on detecting urgency in short crisis messages using minimal supervision and transfer learning. social network analysis and mining, 10(1):58, 2020. jens kersten, anna kruspe, matti wiegmann, and friederike klan. robust filtering of crisisrelated tweets. in iscram 2019 conference proceedings-16th international conference on information systems for crisis response and management, 2019. url http://idl.iscram. org/files/jenskersten/2019/1909_jenskersten_etal2019.pdf. diego kozlowski, elisa lannelongue, frédéric saudemont, farah benamara, alda mari, véronique moriceau, and abdelmoumene boumadane. a three-level classification of french tweets in ecological crises. inf. process. manag., 57(5):102284, 2020. doi: 10.1016/j.ipm.2020.102284. url https://doi.org/10.1016/j.ipm.2020.102284. manfred krifka. commitments and beyond. theoretical linguistics, 45(1-2):73–91, 2019. d robert ladd. a first look at the semantics and pragmatics of negative questions and tag questions. in papers from the... regional meeting. chicago ling. soc. chicago, ill, number 17, pages 164– 171, 1981. pierre larrivée and alda mari. interpreting high negation in negative interrogatives: the role of the other. linguistics vanguard, 8(2):219–226, 2022. peter lasersohn. context dependence, disagreement, and predicates of personal taste. linguistics and philosophy, 28(6):643–686, 2005. enzo laurenti, nils bourgon, farah benamara, alda mari, véronique moriceau, and camille courgeon. give me your intentions, i’ll predict our actions: a two-level classification of speech acts for crisis management in social media. in nicoletta calzolari, frédéric béchet, philippe blache, khalid choukri, christopher cieri, thierry declerck, sara goggi, hitoshi isahara, bente maegaard, joseph mariani, hélène mazo, jan odijk, and stelios piperidis, editors, proceedings of the thirteenth language resources and evaluation conference, lrec 2022, marseille, france, 20-25 june 2022, pages 4333–4343. european language resources association, 2022a. url https://aclanthology.org/2022.lrec-1.462. 40 https://aclanthology.org/j18-4012 https://aclanthology.org/w02-0213 https://aclanthology.org/w02-0213 http://idl.iscram.org/files/jenskersten/2019/1909_jenskersten_etal2019.pdf http://idl.iscram.org/files/jenskersten/2019/1909_jenskersten_etal2019.pdf https://doi.org/10.1016/j.ipm.2020.102284 https://aclanthology.org/2022.lrec-1.462 digging communicative intentions enzo laurenti, nils bourgon, benamara farah, alda mari, moriceau véronique, and camille courgeon. speech acts and communicative intentions for urgency detection. in proceedings of the 11th joint conference on lexical and computational semantics, pages 289–298, seattle, washington, july 2022b. association for computational linguistics. doi: 10.18653/v1/2022.starsem-1.25. url https://aclanthology.org/2022.starsem-1.25. hang le, loı̈c vial, jibril frej, vincent segonne, maximin coavoux, benjamin lecouteux, alexandre allauzen, benoı̂t crabbé, laurent besacier, and didier schwab. flaubert: unsupervised language model pre-training for french. arxiv preprint arxiv:1912.05372, 2019. tsung-yi lin, priya goyal, ross girshick, kaiming he, and piotr dollár. focal loss for dense object detection. in proceedings of the ieee international conference on computer vision, pages 2980–2988, 2017. weiwei liu, haobo wang, xiaobo shen, and ivor w tsang. the emerging trends of multi-label learning. ieee transactions on pattern analysis and machine intelligence, 44(11):7955–7974, 2021. alda mari. assertability conditions of epistemic (and fictional) attitudes and mood variation. in semantics and linguistic theory, volume 26, pages 61–81, 2016. alda mari and paul portner. mood variation with belief predicates: modal comparison and the raisability of questions. glossa: a journal of general linguistics, 40(1), 2021. louis martin, benjamin muller, pedro javier ortiz suárez, yoann dupont, laurent romary, éric villemonte de la clergerie, djamé seddah, and benoı̂t sagot. camembert: a tasty french language model. arxiv e-prints, art. arxiv:1911.03894, nov 2019. louis martin, benjamin muller, pedro javier ortiz suárez, yoann dupont, laurent romary, éric villemonte de la clergerie, djamé seddah, and benoı̂t sagot. camembert: a tasty french language model. in acl 2020 58th annual meeting of the association for computational linguistics, seattle / virtual, united states, july 2020. doi: 10.18653/v1/2020.acl-main.645. url https://hal.inria.fr/hal-02889805. richard mccreadie, cody buntain, and ian soboroff. trec incident streams: finding actionable information on social media. in zeno franco, josé j. gonzález, and josé h. canós, editors, proceedings of the 16th international conference on information systems for crisis response and management, valència, spain, may 19-22, 2019. iscram association, 2019a. url http://idl.iscram.org/files/richardmccreadie/2019/1867_ richardmccreadie_etal2019.pdf. richard mccreadie, cody buntain, and ian soboroff. trec incident streams: finding actionable information on social media. in proceedings of the 16th iscram conference, 2019b. richard mccreadie, cody buntain, and ian soboroff. incident streams 2019: actionable insights and how to find them. in proceedings of the 17th iscram conference, 2020. keval morabia, neti lalita bhanu murthy, aruna malapati, and surender samant. sedtwik: segmentation-based event detection from tweets using wikipedia. in proceedings of the 2019 41 https://aclanthology.org/2022.starsem-1.25 https://hal.inria.fr/hal-02889805 http://idl.iscram.org/files/richardmccreadie/2019/1867_richardmccreadie_etal2019.pdf http://idl.iscram.org/files/richardmccreadie/2019/1867_richardmccreadie_etal2019.pdf conference of the north american chapter of the association for computational linguistics: student research workshop, pages 77–85, minneapolis, minnesota, june 2019. association for computational linguistics. doi: 10.18653/v1/n19-3011. url https://aclanthology. org/n19-3011. alexandra olteanu, sarah vieweg, and carlos castillo. what to expect when the unexpected happens: social media communications across crises. in proceedings of the 18th acm conference on computer supported cooperative work & social computing, cscw ’15, pages 994–1009, 2015. melina plakidis and georg rehm. a dataset of offensive german language tweets annotated for speech acts. in proceedings of the thirteenth language resources and evaluation conference, pages 4799–4807, 2022. paul portner. mood. oxford university press, kettering, northamptonshire, uk, 2018. el quarantelli, a boin, and p lagadec. studying future disasters and crises: a heuristic approach. handbook of disaster research, pages 61–83, 2017. christian reuter and marc-andré kaufhold. fifteen years of social media in emergencies: a retrospective review and future directions for crisis informatics. journal of contingencies and crisis management (jccm), 26(1):41–57, 2018. issn 0966-0879. doi: https://doi.org/10.1111/ 1468-5973.12196. url http://tubiblio.ulb.tu-darmstadt.de/108144/. special issue: human-computer-interaction and social media in safety-critical systems. christian reuter, amanda lee hughes, and marc-andré kaufhold. social media in crisis management: an evaluation and analysis of crisis informatics research. international journal of human–computer interaction, 34(4):280–294, 2018. lina m. rojas-barahona, alejandra lorenzo, and claire gardent. building and exploiting a corpus of dialog interactions between french speaking virtual and human agents. in proceedings of the eighth international conference on language resources and evaluation (lrec’12), pages 1428–1435, istanbul, turkey, may 2012. european language resources association (elra). url http://www.lrec-conf.org/proceedings/lrec2012/pdf/505_ paper.pdf. sara sabour, nicholas frosst, and geoffrey e hinton. dynamic routing between capsules. advances in neural information processing systems, 30, 2017. jerrold sadock. 3 speech acts. the handbook of pragmatics, page 53, 2004. tulika saha, srivatsa ramesh jayashree, sriparna saha, and pushpak bhattacharyya. bert-caps: a transformer-based capsule network for tweet act classification. ieee transactions on computational social systems, 7(5):1168–1179, 2020a. tulika saha, aditya prakash patra, sriparna saha, and pushpak bhattacharyya. a transformer based approach for identification of tweet acts. in 2020 international joint conference on neural networks (ijcnn), pages 1–8. ieee, 2020b. 42 https://aclanthology.org/n19-3011 https://aclanthology.org/n19-3011 http://tubiblio.ulb.tu-darmstadt.de/108144/ http://www.lrec-conf.org/proceedings/lrec2012/pdf/505_paper.pdf http://www.lrec-conf.org/proceedings/lrec2012/pdf/505_paper.pdf digging communicative intentions tulika saha, apoorva upadhyaya, sriparna saha, and pushpak bhattacharyya. towards sentiment and emotion aided multi-modal speech act classification in twitter. in proceedings of the 2021 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 5727–5737, online, june 2021. association for computational linguistics. doi: 10.18653/v1/2021.naacl-main.456. url https: //aclanthology.org/2021.naacl-main.456. efsun sarioglu kayi, linyong nan, bohan qu, mona diab, and kathleen mckeown. detecting urgency status of crisis tweets: a transfer learning approach for low resource languages. in proceedings of the 28th international conference on computational linguistics, pages 4693–4703, barcelona, spain (online), december 2020. international committee on computational linguistics. doi: 10.18653/v1/2020.coling-main.414. url https://aclanthology.org/2020. coling-main.414. roser saurı́ and james pustejovsky. factbank: a corpus annotated with event factuality. language resources and evaluation, 43(3):227, 2009. john r searle. indirect speech acts. in speech acts, pages 59–82. brill, leiden, netherlands, 1975. john rogers searle. speech acts: an essay in the philosophy of language. cambridge university press, cambridge, england, 1969. i. v. serban, r. lowe, p. henderson, l. charlin, and j. pineau. a survey of available corpora for building data-driven dialogue systems: the journal version. dialogue & discourse, 9(1):1–49, 2018. lina sherkawi, nada ghneim, and oumayma al dakkak. arabic speech act recognition techniques. acm trans. asian low-resour. lang. inf. process., 17(3), feb 2018. issn 2375-4699. doi: 10.1145/3170576. url https://doi.org/10.1145/3170576. mandy simons. observations on embedding verbs, evidentiality, and presupposition. lingua, 117 (6):1034–1056, 2007. pontus stenetorp, sampo pyysalo, goran topić, tomoko ohta, sophia ananiadou, and jun’ichi tsujii. brat: a web-based tool for nlp-assisted text annotation. in proceedings of the demonstrations session at eacl 2012, avignon, france, april 2012. association for computational linguistics. tamina stephenson. judge dependence, epistemic modals, and predicates of personal taste. linguistics and philosophy, 30(4):487–525, 2007. andreas stolcke, klaus ries, noah coccaro, elizabeth shriberg, rebecca bates, daniel jurafsky, paul taylor, rachel martin, carol van ess-dykema, and marie meteer. dialogue act modeling for automatic tagging and recognition of conversational speech. computational linguistics, 26 (3):339–374, 2000. url https://aclanthology.org/j00-3003. shivashankar subramanian, trevor cohn, and timothy baldwin. target based speech act classification in political campaign text. in proceedings of the eighth joint conference on lexical and computational semantics (*sem 2019), pages 273–282, minneapolis, minnesota, 43 https://aclanthology.org/2021.naacl-main.456 https://aclanthology.org/2021.naacl-main.456 https://aclanthology.org/2020.coling-main.414 https://aclanthology.org/2020.coling-main.414 https://doi.org/10.1145/3170576 https://aclanthology.org/j00-3003 june 2019. association for computational linguistics. doi: 10.18653/v1/s19-1030. url https://aclanthology.org/s19-1030. tiancheng tang, xinhuai tang, and tianyi yuan. fine-tuning bert for multi-label sentiment analysis in unbalanced code-switching text. ieee access, 8:193248–193256, 2020. cagri toraman, izzet emre kucukkaya, oguzhan ozcelik, and umitcan sahin. tweets under the rubble: detection of messages calling for help in earthquake disaster. arxiv preprint arxiv:2302.13403, 2023. ashish vaswani, noam shazeer, niki parmar, jakob uszkoreit, llion jones, aidan n gomez, łukasz kaiser, and illia polosukhin. attention is all you need. advances in neural information processing systems, 30, 2017. sarah vieweg, carlos castillo, and muhammad imran. integrating social media communications into the rapid assessment of sudden onset disasters. in proceedings of the 6th international conference of social informatics, socinfo’14, pages 444–461, 2014. soroush vosoughi. automatic detection and verification of rumors on twitter. thesis, massachusetts institute of technology, 2015. url https://dspace.mit.edu/handle/ 1721.1/98553. soroush vosoughi and deb roy. tweet acts: a speech act classifier for twitter. proceedings of the tenth international aaai conference on web and social media (icwsm 2016), 2016. m. weisser. how to do corpus pragmatics on pragmatically annotated data: speech acts and beyond. john benjamins publishing company, 2018. guanghao xu, hyunjung lee, myoung-wan koo, and jungyun seo. convolutional neural network using a threshold predictor for multi-label speech act classification. in 2017 ieee international conference on big data and smart computing (bigcomp), pages 126–130, 2017. doi: 10.1109/ bigcomp.2017.7881727. kiran zahra, muhammad imran, and frank o ostermann. automatic identification of eyewitness messages on twitter during disasters. information processing & management, 57(1):102–107, 2020. issn 0306-4573. raffaella zanuttini, miok pak, and paul portner. a syntactic analysis of interpretive restrictions on imperative, promissive, and exhortative subjects. natural language & linguistic theory, 30(4): 1231–1274, 2012. min-ling zhang and zhi-hua zhou. a review on multi-label learning algorithms. ieee transactions on knowledge and data engineering, 26(8):1819–1837, 2013. renxian zhang and naishi liu. recognizing humor on twitter. in proceedings of the 23rd acm international conference on conference on information and knowledge management, pages 889– 898, 2014. renxian zhang, dehong gao, and wenjie li. what are tweeters doing: recognizing speech acts in twitter. in workshops at the twenty-fifth aaai conference on artificial intelligence. citeseer, 2011. 44 https://aclanthology.org/s19-1030 https://dspace.mit.edu/handle/1721.1/98553 https://dspace.mit.edu/handle/1721.1/98553 introduction motivation when communicative intentions reveal urgency previous approaches and research questions overview of the main contributions related work speech acts in social media speech acts in the crisis domain sa in other domains sa automatic detection crisis datasets contributions a two-level annotation of speech acts for urgency detection tweet level sub-tweet level data and annotation dataset annotation procedure results of the annotation campaign sa annotations: quantitative results sa annotations vs. crisis types sa vs. urgency annotations evolution of speech acts over time interim conclusions automatic detection of sa experimental settings models evaluation protocol results random sampling results out-of-event results error analysis conclusion dialogue & discourse 16(3) (2025) 96–130 doi: 10.5210/dad.2025.305 prior lessons of incremental dialogue and robot action management for the age of language models casey kennington caseykennington@boisestate.edu department of computer science boise state university pierre lison plison@nr.no norwegian computing center david schlangen david.schlangen@uni-potsdam.de computational linguistics university of potsdam editor: hendrik buschmeier submitted 10/2023; accepted 11/2025; published online 12/2025 abstract efforts towards endowing robots with the ability to speak have benefited from recent advancements in natural language processing, in particular large language models. however, current language models are not fully incremental, as their processing is inherently monotonic and thus lack the ability to revise their interpretations or output in light of newer observations. this monotonicity has important implications for the development of dialogue systems for human–robot interaction. in this paper, we review the literature on interactive systems that operate incrementally (i.e., at the word level or below it). we motivate the need for incremental systems, survey incremental modeling of important aspects of dialogue like speech recognition and language generation. primary focus is on the part of the system that makes decisions, known as the dialogue manager. we find that there is very little research on incremental dialogue management, offer some requirements for practical incremental dialogue management, and implications of incremental dialogue for embodied, robotic platforms in the age of large language models. keywords: spoken dialogue systems, incremental, human-robot interaction, dialogue management 1. introduction large language models (llms) have become more prominent in robotics, and for good reason. williams et al. (2024) explain that llms can offer “quick-enabling of full-pipeline solutions” for many aspects of robots ranging from enabling robots to engage humans in spoken interaction to generating action plans (singh et al., 2023; cohen et al., 2024; singh et al., 2024; mahadevan et al., 2024) and emotional behaviours (mishra et al., 2023). while promising, some recent work has identified important qualities that llms lack which, if part of the model, would make interaction with robots seem more natural. a recent survey of spoken interaction on robots by reimann et al. (2024) showcases a long history of research that spoken dialogue systems (sdss) are key to endowing robots with handling common artifacts in spoken interaction which are not commonly found in text or written interaction, including (inter alia) turn-taking, requests for clarification and building ©2025 casey kennington, pierre lison, and david schlangen this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). incremental dialogue and robot action management for the age of language models figure 1: example of incremental processing for speech recognition: i, have, and four are recognized, then four is revoked and replaced with forty. diamonds denote the point in time when the information is passed to the next module. figure adapted from kennington et al. (2017). common ground. at the heart of their focus is the dialogue manager because both sdss and robots must make decisions about which actions they will take at any given moment, either by uttering a response, moving a robotic arm, or any other potential action within the capabilities of the robot. the dm (or corresponding robot action manager) not only decides which action to take, but also when to take that action; both are critical for natural interaction between robots and humans (lison and kennington, 2023). both studies look at dm in hri tasks and settings, comparing how different systems divide the decision-making responsibilities, concluding that dm on robots is still rather a new field; more data, tasks, benchmarks, and discussions are needed. fortunately, there exists a body of literature that spans over 25 years (allen et al., 2001) of success in developing and improving systems that enable humans to talk naturally to machines: incremental dialogue, meaning that processing happens at a fine-grained, word-by-word level (see example in figure 1). comparisons between incremental and non-incremental systems have shown that incremental systems significantly improve system performance (ghigi et al., 2014a), are perceived by humans as being more natural (aist et al., 2007; asri et al., 2014) and human-like (edlund et al., 2008), which suggests that the most appropriate systems for robots should be incremental, echoing the requirements of “robot-ready” sds for use in human-robot interaction (hri) settings (kennington et al., 2020). in this paper, we review literature relating to incremental sds and explain how the incremental unit framework has influenced incremental research (section 2.2.1). llm research can greatly benefit from this knowledge on incremental processing. inoue et al. (2024b) points out that, for example, when humans engage in real-time dialogue with robotic agents, humans expect the robot to take seamless turns (i.e., without a gap in conversation, as happens within human-human dialogue) and they expect backchannels (e.g., nodding, or utterances like yeah, or uh huh), neither of which are handled with common llms. their work applied voice activity projection (vap) to enable llms to predict when a person might stop speaking so the llm can respond at an appropriate time. chiba and higashinaka (2025)’s recent vap method also used an llm to predict when to start speaking, and their work directly relied on incremental processing. according to the authors: “[...]even if systems are equipped with a natural turn-taking model, such a model will be ineffective if response generation cannot begin immediately once a turn97 kennington, lison and schlangen shift is detected. incremental response generation is an approach that addresses this issue.” using llms in incremental settings is a positive step, but more work is needed. furthermore, other recent work (hudeček and dusek, 2023) asked if llms are all that is necessary for task-oriented dialogue (albeit outside of dialogue with robots) with some negatives (e.g., “llms underperform[...]” in important aspects of dialogue) and positives (e.g., “llms show the ability to guide the dialogue to a successful ending”). finally, wagner and ultes (2024) investigated the usefulness of llms in dialogue interaction and found that they are effective, but need control and guidance to ensure that dialogue responses are coherent— two critical aspects in scenarios where robots are involved. taken together, while llms can be employed for a wide range of dialogue processing tasks, they still suffer from a number of limitations when it comes to incrementality. in particular, while a llm decoder can process any kind of input including word-level input, they are trained to act upon complete, sentence-level input, rendering them unable to produce behaviour where input and output happen concurrently – although recent work on full-duplex models shows promising results (zhang et al., 2025). in contrast, humans must process individual words while reading text, and psycholinguistic research has shown that speech comprehension happens at a word or even subword level (tanenhaus and spivey-knowlton, 1995). moreover, another requirement of incremental processing is non-monotonicity; i.e., that a model can react to change in input, for example when information coming from a speech recognizer is incorrect, the model needs to be able to revoke the erroneous input and change its internal state–causal language modeling is strictly monotonic, but incremental processing should allow for non-monotonic input. with chatbots, the text-in, text-out nature of the interaction is well-suited for llms, but the expectation of human-like conversation becomes more challenging when people interact with robots due to the anthropomorphic characteristics of many robotic platforms. if, for example, a robot has what appear to be eyes, people expect that the robot can see them, or if the robot has an arm they expect the robot to be able to point or grasp objects. furthermore, it has been shown that people anthropomorphize robots for gender (reich-stiebert and eyssel, 2017; eyssel and hegel, 2012), intelligence (novikova et al., 2015), and even age (plane et al., 2018) depending on the robot’s morphology, size, and movements, which affects the expectations of how robots behave: the more anthropomorphic a robot appears, the more human-like people tend to expect the robot to act. in this paper, we review the literature for incremental sds with a particular focus on the decisionmaking component known as dialogue management (dm; explained further below) for the sake of guiding ongoing and future work related to decision making on robots that interact with humans. we find in our review that that although other elements of sds – such as automatic speech recognition and natural language generation – have seen substantial work on incrementalization, there is a noticeable lack of focus on incremental decision making. we identify some of the challenges and requirements to help guide future research on incremental decision making. the next section begins with background on incremental sdss, focusing first on common modules then fully implemented and evaluated systems. the section that follows then focuses on dm, giving first a brief overview of dm research, then focuses on incremental dm. we then end this review with some concluding remarks and suggested paths for future work. 98 incremental dialogue and robot action management for the age of language models figure 2: traditional architecture for spoken dialogue systems composed of automatic speech recognition (asr), natural language understanding (nlu), dialogue management (dm), natural language generation (nlg), and text-to-speech synthesis (tts). the system is extended to be multimodal, where the added modalities are robot sensors and control. 2. background: incremental spoken dialogue systems in this section, we review literature on common incremental spoken dialogue system modules except dm, which we save for the following section. we explain incremental frameworks that have been adopted, and explain different paradigms of modeling incremental processing. 2.1 spoken dialogue systems: overview equally important to the distinction between incremental (word-level) and non-incremental (utterance or sentence-level) sds is the distinction between end-to-end and modular sdss. an end-to-end system is modeled using a single model that takes in input and produces an expected output directly, such as a question-answering system that produces an answer given a question, or social chatbot that produces responses given text input. end-to-end systems often focus on the capability of producing a written or spoken response no matter what the input is. end-to-end architectures now constitute the dominant approach for developing open-domain dialogue systems where the main focus is the social aspect of the interaction (roller et al., 2020; ni et al., 2023). the social aspects of interaction are, of course, important in a natural dialogue, but in task-oriented dialogue there is often something that is required outside of the dialogue itself for the dialogue to be considered successful; e.g., look up information in a database, perform some kind of robotic action, or complete a payment. modular sdss are often task-based in that they help the user achieve such a goal such as booking a flight; they do not usually focus on social aspects beyond what helps to accomplish the task (budzianowski et al., 2018; zhang et al., 2020b). this traditional distinction between open-ended end-to-end systems and task-based modular architectures is, however, increasingly blurry, as recent years have seen the emergence of end-to-end models specifically designed for task completion (liu et al., 2018; zhang et al., 2020a; hosseini-asl et al., 2020; young et al., 2022) as well as newer agentic architectures. interestingly, end-to-end models for task-oriented systems often operate by augmenting the generative model with implicit “modules” in the form of retrieval mechanisms (qin et al., 2019), knowledge bases (yang et al., 2020) or domain-specific ontologies (chen et al., 2023), or by pre-training the response generation model in a modular fashion (qin et al., 2023). as the name suggests, modular systems are made up of modules that have well-defined roles in the system, and which can be made to communicate with each other. figure 2 depicts visually a modular sds. for example, a prototypical sds is often made up five modules including automatic 99 kennington, lison and schlangen speech recognition (asr) that transcribes speech to a text representation of the human utterances, natural language understanding (nlu) that takes the text and yields a computable semantic abstraction, dialogue management (including dialogue state tracking) that makes a high-level decision about the next action to take (e.g., look up information in a database and respond to the user), natural language generation (nlg) that takes the dialogue manager’s decision and determines which words to use and in what order, and text-to-speech (tts) which actually speaks the words. these modules are further explained below. modular sdss that process incrementally have an added complexity in that all of the modules must operate at granularities that downstream modules can make use of, such as at the word level from asr to nlu. for example, given a system on a robot that is made up of the standard five modules as explained above (along with connections to robotic modules), and someone utters hand me the green book on the left, an incremental sds begins to process as soon as speech is detected. the asr outputs each word, one at a time, and the nlu updates its interpretation each time a word is outputted by the asr and the nlu likewise produces outputs as it gathers information about the utterance, for example, tagging hand as the action as the first word is uttered, and a specific book as the target once the green book it has been uttered. the dm is tasked with querying a module that takes in visual information and instructing an arm to reach for the book in question. an incremental dm might already extend its arm in no particular direction as the first word is uttered to signal understanding, then towards any green book once green book is uttered, then narrow the target down further as on the left is uttered. the nlg could then start uttering a response like green book as it begins to move its arm then ah, here we go once it determines a unique referent. the above example highlights some things that differentiate an incremental dm from a more traditional dm. first, the incremental dm receives installments of information over time, whereas a traditional dm receives all of the information at once after everything has been uttered and the nlu has finished processing the asr’s transcription. the incremental dm has therefore an important role that is lacking in non-incremental dm it not only must decide which action to take, but it also must decide when to take that action given the information that it has so far, and—perhaps a bigger challenge—perform concurrent actions as it is still receiving input (i.e., “full-duplex” models). traditional sds has often relied on endpointing; i.e., waiting for silence after a speaker begins to speak, which burdens the asr with determining when to act. however, pauses in speech are not always signals that someone is done speaking, and incremental sds that relies on a dm to determine when to take an action can potentially use speech, silence, as well as information from the content of the utterance (i.e., via the nlu) to make decisions about when to act. 2.2 frameworks & architectures 2.2.1 the incremental unit framework the incremental unit (iu) framework, a well-established approach to incremental processing, will be a recurring reference in this paper, following the works of schlangen and skantze (2009, 2011). the iu framework views each bit of information created by the modules (e.g., words produced by asr and slots produced by nlu) as part of a global network of interconnected ius no matter which module produced them. the framework defines functions for changing the network including how nodes of the network are added and how the nodes are interconnected. newly created ius by a module (e.g., words by asr ) can be added to the iu network, revoked from the network if the module determines that an iu was erroneously added in light of new information (e.g., the asr first 100 incremental dialogue and robot action management for the age of language models added the word iu four but later revoked and added forty), and ius can be committed, meaning they have already been added to the iu network, and are guaranteed to not be revoked. to be added to the iu network, an iu must be connected to other ius that already exist in the network through two relations: • same-level links which are relations between ius created by the same module – e.g., if the asr recognizes the and dog as two ius, the later word dog has a same level link to the. • grounded-in links where a relation is created between and iu an the iu(s) that gave rise to that iuḟor instance, the ius the and dog from the asr might give rise to a subject tag in the nlu, leading to two grounded-in links, one to each word iu. when ius are operated on (i.e., added, revoked, or committed) the modules that triggered the operation signal downstream modules that consume their output about the change. for example, as the asr module recognizes words from a microphone, it adds each of them to the iu network and signals to the nlu module that a new word has been added. a module’s ability to revoke an iu is important in a natural dialogue interaction. the above asr example of revoking is fairly straight-forward and happens internally between the modules, but there are cases where modules that produce output need to revoke information, for example a robot is planning on uttering something that, given new information, should be changed. in a setting where a robot and a person are working together to move around colored boxes, if the person says ”move the green box to the left” but the asr system mistakenly recognises ”gray” instead of ”green”, the robot will initially approach and attempt to move a gray box. the robot generates a plan to move towards the box and produce an utterance to signal understanding such as okay, i’m on my way to move the gray box. but when the revoke from gray to green happens, the robot must revise its movement and verbal responses accordingly. if the robot hasn’t completed its utterance, it still has a chance to change the utterance to okay, i’m on my way to move the green box and change its direction to the box the person referred to, known in robotics as replanning (cashmore et al., 2019). this is an illustration of non-monotonicity: modules can update the information that they pass to each other, and often updates need to happen after the robot has already taken some kind of action. more on this in section 3. the iu framework has been implemented in several software packages, notably in java as inprotk (baumann and schlangen, 2012) and more recently in python as retico (michael and möller, 2019) and remdis (chiba et al., 2024). other conceptual frameworks such as the information state approach (traum and larsson, 2003) and cohen’s belief-desire-intent model (cohen, 2017) remain valid in incremental sds, including within the iu framework, though they are not strictly incremental dialogue frameworks. later versions of inprotk and, more recently, retico has been extended to include common robot capabilities such as object detection and control of multiple robot platforms (kennington et al., 2020; manaseryan et al., 2025), towards bridging the gap between sds and hri research. while not strictly incremental, opendial (java and python) has been used in spoken hri studies, and can serve as a decision making module in both inprotk and retico (lison and kennington, 2016; jang et al., 2020). 2.2.2 restart vs. update incremental models khouzaimi et al. (2014) points out that not all methods and models are inherently incremental, though many can be made to work incrementally under certain constraints. while their proposed 101 kennington, lison and schlangen method is an important step in improving the incrementality of current systems, it should be highlighted that there are two distinct approaches to modeling incremental systems: restart incremental and update incremental, which we explain below. restart incrementality restart incremental models take in inputs and produce incremental outputs (e.g., at the word level), but the input is repeated as the prefix grows, and models themselves are agnostic to the incremental updates. any model (e.g., a language model using zero-shot classification) could be used restart incrementally. for example, a nlu module that is restart incremental would take in the following input (time moves from top to bottom; each line represents input to a nlu model): the the dog the dog barks update incrementality in contrast to restart incremental models, update incremental models do not need repeated input and the model is designed to maintain a state that updates for each incremental input. an nlu model that works in an update incremental way would not need to repeat a growing prefix from the asr: the dog barks the model explained in kennington and schlangen (2017), for example, is a bayesian updateincremental model that produced a distribution over possible slot values that updated the distribution at each word increment. an open question that we explore below is if a dm model should be either restart or update incremental. 2.3 common modules in incremental, interactive systems 2.3.1 automatic speech recognition current asr systems receive streaming input and produce partial transcriptions, and can often work at word-level increments. early incremental asr were implemented in sphinx (baumann et al., 2009), and newer neural asr systems could be made to operate at the character level (hwang and sung, 2016). the most common evaluation metric for asr is word error rate, and recent neural models have shown very low error rates in common asr benchmark datasets. however, evaluation of incremental asr requires a closer look at how often a model alters its output and latency of results (baumann et al., 2016; whetten et al., 2023). because of the nature of the iu processed by asr, evaluating its performance both independently as well as within larger systems is crucial, as certain errors may affect downstream components. because asr requires streaming input, they are inherently incremental in terms of input, though not all recognizers produce incremental output; they often wait until a pause in the speech (i.e., end-pointing). however, recent asr models have become very effective at accurate transcription in multiple languages and they produce incremental output. the whisper (radford et al., 2022), wav2vec (baevski et al., 2020), and deep speech 2 amodei et al. (2016) all produce incremental output, the former 2 being incorporated into the retico framework. imai et al. (2025) recently eval102 incremental dialogue and robot action management for the age of language models uated conversational speech recognition on several state-of-the-art asr models (including whisper and wacv2vec), taking gender into account in their evaluation. the results were mixed; asr has come a long way in the past two decades, but more work is needed to accommodate different demographics of speakers, and correcting errors. incremental asr is more challenging on robots in spoken hri settings because robot voice can be picked up by the asr, thereby ‘confusing’ the robot. this can be at least partially addressed using diarization (i.e., tracking the voice of particular individuals within a speech signal, including a robot) or by the robot tracking its own speech signal and filtering it out of the asr input. 2.3.2 natural language understanding understanding natural language in sdss also has a long history. in most nlu models, the input corresponds to text. it the case of sds, the input to nlu is transcribed speech. the output of nlu is important to consider here, because it is often what serves as the input to the dm. the output needs to abstracted sufficiently over the input text to form a computable meaning representation that the dm can use for making a decision on how to act. that meaning representation in incremental nlu has been represented in various ways in the literature including tagged words, logical forms, or frames (i.e., a set of key-value pairs known as slots), recent models tend to use tags to produce slots and frames as output, or latent representations (e.g., embeddings). below is an example frame for an utterance made to command a specific robot action: move the red ball into the box on the left made up of four slots: intent command object red ball target left box action move object to target like incremental asr, incremental nlu produces output as early as possible (for example, individual filled slots), but unlike asr the input is discrete words instead of a continuous speech signal, so the intervals of when output is produced can vary depending on the input and the domain. early incremental nlu focused on inferring semantic frames. each input word produced a partially complete frame as output (devault et al., 2011; devault and traum, 2012, 2013; yamauchi et al., 2013; kennington and schlangen, 2014; kennington et al., 2014b, 2015). part of the frame is also the dialogue act; i.e., the overarching type of utterance produced by the interlocutor (e.g., a question or an assertion), which has also a history of incremental models (petukhova and bunt, 2011). early actionable output from nlu is particularly important in hri settings, where robots can already be moving towards referred objects before the person finishes their request, which is an important physical backchannel: a robot beginning to move towards an object is a signal to the user that the robot is understanding the unfolding utterance, as done in hough and schlangen (2016) similar to their non-incremental counterparts, incremental nlu can benefit from syntactic parsing to guide language understanding, but in the case of incremental nlu, the parsers must also work incrementally (i.e., produce a partial syntactic parse such as a tree for each word input). there has been ample research in incremental parsing for different syntactic formalisms, including dependency parsing (nivre, 2008), combinatory categorical grammar parsing (hassan et al., 2008; beuck and menzel, 2013), as well as frameworks with a more semantic focus, such as robust minimal recursion semantics (copestake, 2007; peldszus et al., 2012), dynamic syntax (eshghi et al., 103 kennington, lison and schlangen 2013) (see hough et al. (2015) for a comparison of robust minimal recursion semantics and dynamic syntax for incremental dialogue) and abstract meaning representation (damonte et al., 2017). in multimodal sds and interaction with robots, the nlu component must sometimes resolve references to objects that exist in the shared space with the robotic system and the human interlocutor. incremental reference resolution can be viewed as the ability to narrow down possible referents in a shared visual space to an individual object. an incremental reference resolution model might, for example, understand the word red to refer to objects that have a red color, and book to then further narrow down from all red objects to only red books. incremental reference resolution is sometimes an integral part of nlu (kennington et al., 2014b), but have also been designed for modules that only resolve references (schlangen et al., 2009; paetzel et al., 2015; kennington and schlangen, 2015; schlangen et al., 2016; kennington and schlangen, 2017), information that the dm may need to use for making a decision. other works have explored how deep learning architectures can be used for incremental nlu, including recurrent architectures (shivakumar et al., 2019) and to what degree architectures that are not inherently incremental (e.g., self-attention transformers which are designed to process multiword input in parallel) can be used for incremental nlu (madureira and schlangen, 2020), with mixed results. it is important to explore further how neural models can work incrementally because many dialogue phenomena are incremental in nature. for example, shalyminov et al. (2017) showed that that deep neural dialogue models failed on common spoken phenomena like restarts and selfcorrections. more recent work has explored how transformer lms can be successfully used for incremental nlu. madureira and schlangen (2020) showed that both bidirectional long short-term memory (lstm) models and transformer-based encoders assume that an input sequence to be encoded is available a-prior in its entirety, to be processed either forwards and backwards (in the case of bidirectional lstms) or as a full sequence (in the case of transformer-based encoders). the results of their work support the possibility of using bidirectional encoders in their developed incremental mode while training their non-incremental qualities (i.e., parallel processing). kahardipraja et al. (2021) explored using linear transformers with a recurrence mechanism to examine the feasibility of linear transformers for incremental nlu. they found that linear transformers have better performance and faster inference than standard transformers when used in a restart-incremental fashion. hri settings make the requirements of nlu more challenging due to the fast-paced, multimodal nature of the interaction. llms trained only on text are limited, but the recent proliferation of multimodal llms have real potential for being used on robot platforms. modalities include images, speech, and video, for example the one-peace (wang et al., 2023) and palm-e models (driess et al., 2023). a recent survey explains the modeling trends (e.g., one vs. two-tower; different methods of representing images) for vision lms (fields and kennington, 2023); improvements in visual lms will benefit hri research because robots need to see and talk about objects in a shared space. other recent work focuses on how transformers handle incremental nlu revisions. most transformer models are causal in that they are forced to produce a single output once an ambiguity is resolved, but madureira et al. (2024) proposed an interpretable way to analyse incremental states to show how transformers handle ambiguity. they showed that transformer sequential structures encode information on the garden path effect, as well as the resolution of garden paths. another model, tapir, a two-pass method that modeled the revision process itself showed better performance on incremental metrics compared to transformers used restart-incrementally (kahardipraja et al., 2023). 104 incremental dialogue and robot action management for the age of language models 2.3.3 natural language generation and speech synthesis early work in incremental nlg focused on resolving references in situated dialog. kelleher and kruijff (2006) presented an approach to generating locative expressions using a basic incremental algorithm that considered visual salience as a computation of an object’s perceivable size and centrality relative to the viewer, choosing words that distinguish between the target object and distractor objects. while the algorithm the authors present is “incremental”, it is not evaluated as a word-byword incremental model, but given the co-location and potential of being used at the word level, we include it here. more recent work has shown that incremental installments of words that refer to a visual object using a model trained on visual object/word pairings that uses a beam search to determine the best possible word to utter can use a model of vision/word that isn’t trained specifically for nlg (zarrieß and schlangen, 2016). incremental nlg that builds on the iu framework include work that used a buffer of words to be uttered, and three operations add, revoke, and purge were used for operating on the buffer (dethlefs et al., 2012a). the add operation, of course, means a word is added to the buffer and eventually uttered, unless it was revoked (removed from the buffer) or purged (all words currently in the buffer are removed in favor of a new hypothesis/goal). the nlg module often produced words faster than they could be articulated by a tts, giving an incremental nlg time to determine which words should be uttered, and in which order. the authors also carried out experiments to explore how nlg interacts with output generation of other modalities, such as information on a screen (dethlefs et al., 2012b). in general, the research has shown how incremental generation produces systems that are more reactive and perceived as more natural to human dialogue partners. in hri settings, incremental nlg is important, for example, when referring to real-world objects (zarrieß and schlangen, 2016). others also looked at how incremental multimodal generation affects the interaction qualities when the sds is part of a virtual agent; including not only nlg but also hand gestures and eye gaze by the agent (van welbergen et al., 2012). instead of planning all articulations before they were realized, the model generated behaviours incrementally and linked increments in the multiple output modalities to each other, so what happened corresponded temporally to other modalities (e.g., saying that in conjunction with a pointing gesture). such articulation requires that the speech synthesis also be incremental, as a system utterance currently being generated by the nlg and transferred to the tts might change before the tts actually articulates a word in the utterance, thereby changing prosody or duration; e.g., the system may want to hold the floor longer so will need to take longer to speak or insert artifacts such as ummm (buschmeier et al., 2012; baumann, 2014). improvements in asr and nlu have enabled systems to be far more capable than even a few years ago, but the biggest gains in nlg have been due to lms. generative large lms are flexible in that inputs can be structured text and models can be made to produce useful structured output that is useful for sds and hri. for example, the dm can output a request to look up information in a database, then take that structured information and input it into a lm, which produces a surface utterance that has the necessary information. fortunately, while lm input processing is not incremental as defined, lm output is naturally incremental due to inherent modeling (i.e., autoregression). however, while lm output can be paused (goyal et al., 2023) or interrupted, the output itself cannot be changed given new input. 105 kennington, lison and schlangen 2.3.4 incremental systems & evaluation beyond individual modules, full systems are more complex and difficult to evaluate, though some prior works have shown how incremental systems are better in some domains than their nonincremental counterparts. for example, a virtual in-car dialogue that presented information incrementally was shown to be safer and more effective at helping users remember information (kousidis et al., 2014). the system was able to detect changes in the car’s control (e.g., changing lanes or speed) and if any change was detected, the system would pause its output and resume after the driving was constant. this allowed drivers to focus on driving instead of non co-located interlocutors. in another system, fischer et al. (2021) used incremental speech adaptation to initiate humanrobot interactions in noisy (in-the-wild) scenarios. the robot incrementally adjusted the loudness of its voice depending on the circumstances, and was perceived positively by human users. finally, ghigi et al. (2014b) showed that that an incremental dialogue strategy significantly improved system performance by eliminating long and often off-task utterances that generally produce poor speech recognition results. user behaviour is also affected; the user tends to shorten utterances after being interrupted by the system. in both of these examples of full system evaluation, human participants were recruited to interact with the systems. following common human-agent, human-robot, or human-interface research, human participants are often presented with one of two different versions of the system; a baseline/control version and a test version that focuses on a specific system component. for example, the in-car dialogue system was evaluating incremental information presentation and pausing vs. a system that kept talking; in both cases participants were asked a true/false question about the information they just heard which they answered by pressing a button on the steering wheel. metrics include objective measures and subjective measures. in the case of the in-car system, objective measures included successful lane changes, the true/false answer, and driving speed, all which could be measured at any given time. subjective measures include surveys about the participants’ experience with the system. objective and subjective measures are often combined to give a more holistic picture of the evaluation. for example, participants may have felt that they drove the same in both the control and test versions of the system (subjective), but failed to answer questions correctly, or failed to drive at the prescribed speed. köhn (2018) reviewed incremental processing in the field of natural language processing (including parsing, machine translation, among others which are beyond our scope), pointing out that granularity, grounding, monotonicity, and timeliness are all aspects of incremental processing that play into how incremental systems are perceived. most incremental sds research is performed with the level of granularity set at the word level, but it might be better in certain cases to work at subword or phrase levels, or on speech directly (see kebe et al. (2022) for a non-incremental model grounded in raw speech). grounding, moreover, is how a system aligns its output (in the case of sds, generated speech) to what is happening in the dialogue state including physical context and the conversation up until that point. grounding is particularly important (and challenging) for hri settings and tasks, as the human and the robot need to track objects, dialogue history (including entity tracking; i.e., objects that have been discussed before). task-completion is a common metric (i.e., did the human and robot pair complete the assigned task, like put together a puzzle?), but in many cases the task is more social and more focused on human impressions of their interactions with the robot. 106 incremental dialogue and robot action management for the age of language models another challenge in incremental evaluation that is more specific to incremental processing is monotonicity. monotonicity is an open question in sds research; an ideal incremental asr for example, would only output the correct word as early as possible as they are spoken without the need for revoking and replacing words. thus while monotonicity is an ideal to strive for, system modules make mistakes and need to be able to repair those mistakes (hence the need for the iu framework), but knowing how monotonic a system or an individual module is can be a useful metric for measuring stability. finally, timeliness is important: the system needs to respond quickly, but the system should reach a level of confidence that the response is the proper one. the challenge lies in striking a balance between these two opposing optimisations: ensuring timeliness while maintaining accuracy, which underscores the importance of non-monotonic operations. llms are increasingly being used to evaluate dialogue systems, notably through llm-as-ajudge setups (gu et al., 2025), although trust and reliability remain important concerns compared to actual human evaluation studies (pan et al., 2024). lab settings, however, often do not generalize to real-world settings due to constrained systems and low number of participants. crowd-sourcing is a way to increase the number of humans who can evaluate a full system, but quality control is often difficult. 3. review of incremental dialogue management in this section, we review literature relating to incremental dm. we give an overview of dm, dialogue state tracking, and attempts at modeling incremental dm. 3.1 a brief overview of dialogue management dm lies at the crossroads between nlu and nlg and is responsible for controlling the general flow of the interaction, often in relation with the task(s) that should be fulfilled by the dialogue agent. in their seminal work on the information state approach to dm, traum and larsson (2003) mention four objectives: 1. updating a representation of the dialogue context on the basis of interpreted communication (from all dialogue participants) ; 2. providing context-dependent expectations for interpretation of observed signals as communicative behaviour ; 3. interfacing with task/domain processing (e.g., database, planner, execution module, other back-end system), to coordinate dialogue and non-dialogue behaviour and reasoning ; 4. deciding what content to express next and when to express it. current dm approaches distinguish between two central (and consecutive) components, respectively called dialogue state tracking and action/response selection. 3.1.1 dialogue state tracking the task of maintaining a representation of the current dialogue state over the course of the interaction is called dialogue state tracking (williams et al., 2016; ren et al., 2018; heck et al., 2020). the dialogue state aims to reflect the system knowledge of the current conversational situation, and 107 kennington, lison and schlangen often includes multiple variables related to the dialogue history, common ground, external context (including the physical context, in the case of human–robot interaction), and the task(s) to perform. this update of this dialogue state should occur upon the reception of any new observation that may potentially impact the system’s understanding of the current conversational situation, such as new user utterances, but also changes in the physical context of the interaction (for instance, new entities perceived in the visual scene, or updates on the current location of the robot). for incremental systems, those observations will typically correspond to incremental units produced by the nlu module. in task-oriented systems, the dialogue state is often represented as a list of slots to fill (williams et al., 2016; mrkšić et al., 2017), where a slot typically represents a required or optional attribute whose value should be derived from the user inputs to complete the task. for instance, a restaurant booking system might have slots for the date, time and number of people. although such slot-filling representation can be applied to many domains, it remains restricted to a fixed list of predefined slots, and may therefore be difficult to apply to conversational domains with varying numbers of entities and relations between them. this is notably the case in human–robot interaction, where the number of persons in a room, or the number of objects detected in the current visual scene is not fixed in advance and may change over the course of the interaction. in such settings, representing the dialogue state as a graph of entities connected through various relations is a preferred alternative (ultes et al., 2018; walker et al., 2022). approaches to dialogue state tracking also differ in whether they explicitly represent uncertainty related the current dialogue state using probability distributions. many dm approaches represent the current dialogue state as a mere collection of key-value pairs (slots and their values). although this representation does simplify both dialogue state tracking and action selection (in particular when this selection is optimized using reinforcement learning), it makes it harder to capture uncertain, ambiguous or untrustworthy information, which may arise from e.g. error-prone sensory inputs (e.g., imperfect object recognition or asr) or non-deterministic inference (e.g. linguistic ambiguities). an alternative is to explicitly represent the dialogue state as partially observable and define a probability distribution over possible state values (young et al., 2013; mrkšić et al., 2017), often called the belief state. this belief state can notably be expressed as a bayesian network over state variables (thomson and young, 2010). 3.1.2 action/response selection the second core dm task is action selection, whose role is to determine the next (verbal or nonverbal) action(s) that the system should undertake, based on the dialogue state updated through dialogue state tracking. although those actions frequently correspond to verbal system responses, they may also express other types of actions, such as api calls or high-level physical actions in the case of robotic platforms. a given dialogue state may lead to the selection of several actions to execute in parallel or in sequence (for instance, a robot may simultaneously move to a new location and utter a sentence to describe his movement to the user) or to no action at all. the selection of the next action/response may take several forms, from handcrafted flowcharts and logical rules to data-driven techniques. early work includes larsson (2002), which surveyed existing approaches to dm including logic-based, finite state, form-based, and plan-based approaches, but the author regarded those approaches as limited in their practicality—most were theoretical models without a concrete implementation. to remedy this situation, larsson (2002) introduced 108 incremental dialogue and robot action management for the age of language models issue-based dialogue management. issues are modeled semantically as questions, which can be implemented in multiple theories (e.g., plan-based or form-based). this kind of dialogue is systemdriven in that the system has a specific task that it must perform and it drives the dialogue by asking questions to the user, for example an automated travel agency would ask questions about price ranges, travel dates, origin and destination airports, and airlines if it is going to help a user find an appropriate flight. as is the case with most dialogues, a kind of “information exchange” takes place, the system is not requiring anything of the user beyond responding verbally with requested information. also seminal is the early work of cohen and levesque (1990) on plan-based approaches to dm building on earlier work by allen (1979). more recently, cohen and galescu (2023) showcases a fully working multimodal conversational system that infers users’ intentions and plans to achieve those goals. the system can infer obstacles to goals and actions and find ways to address them collaboratively. the dm is broken down int four parts: plan recognition, obstacle detection and goal adoption, planning, then execution. planning here is an important aspect of the dm; it does not just identify an action to take now, it identifies a plan (i.e., a series of actions) that must be taken to achieve a higher-level user goal, making it potentially more amenable to multimodal (including iva and robotic) control. the mapping from dialogue state to action(s) is called a dialogue policy, and various methods have been developed to automatically learn such policies from real or simulated dialogue data. supervised learning techniques can be employed to imitate the conversational strategies followed by human experts in a corpus of dialogue (griol et al., 2008). however, the behaviour of human experts may be hard to imitate, especially as those experts often base their decisions on a different and richer understanding of the conversational context that what can be captured in a dialogue state. those supervised learning techniques also suffer from data sparsity problems, as only a small fraction of the state space can realistically be covered by the dialogue examples. to this end, a range of reinforcement learning methods have been proposed to automatically optimize dialogue policies based on a reward function (rieser and lemon, 2011; young et al., 2013; williams et al., 2017; peng et al., 2018). although the reward function is often defined manually based on the system objectives, it can also be learned from data (su et al., 2018; takanobu et al., 2019). the dialogues can be generated automatically using a user simulator (schatzmann et al., 2006; chandramohan et al., 2011; ultes et al., 2017) or from actual dialogues with human users (su et al., 2016; shah et al., 2018). the underlying process to optimize may be either framed as a markov decision process (mdp), or, in case the dialogue state itself is consider to be uncertain, a partially observable markov decision process (pomdp). while framing action selection as a pomdp makes it possible to explicitly account for uncertainties about the current dialogue state, it also complicates the dialogue policy optimization, due to the need to derive a policy in a continuous and high-dimensional belief state space. dialogue policies can also be expressed in terms of probabilistic rules with a skeleton provided by the system designer while the rule parameters are estimated from dialogue data, as shown by lison (2015a,b). recent work in reinforcement learning goes well beyond the pomdp model, including reinforcement learning with human feedback (ouyang et al., 2022), and proximal policy optimization (shao et al., 2024), each with potential effectiveness for dm. some recent work shows how lms can be used for dm (niu et al., 2024; zhang et al., 2025). little work has been done, however, on the problem of revising current action plans of the dialogue manager in light of new (incremental) observations. for instance, a robot may start executing 109 kennington, lison and schlangen a particular action plan, but suddenly hear a human user say “stop!”, in which case the robot ought to interrupt its current sequence of actions and devise an alternative response. this ability to revise or regenerate plans is related to the problem of replanning in robotics and automation (garrett et al., 2020; zhou et al., 2023). the output of the action selection should convey what the system should say or do next, and is often structured as a logical form (traum and larsson, 2003; lison, 2015b). in the case of a verbal response, the nlg module is then responsible for converting this representation into an actual utterance. alternatively, the dialogue manager may generate a prompt containing natural language instructions on how to respond, and use this prompt as input to a llm in charge of producing the response. 3.1.3 turn-taking and end-of-turn prediction dialogue management in spoken dialogue systems is not only about what to do, but also about when to do it. this question of timing has, unfortunately, not received as much attention as it should have. a common but sub-optimal approach is to wait until the current speaker has stopped speaking for a given period of time, and seek to predict whether they are likely to continue or not (ferrer et al., 2002). raux and eskenazi (2009) presented a finite-state model for turn-taking in spoken dialogue systems, relying on a cost matrix and a decision-theoretic framework to determine whether to take the dialogue floor, release it, wait or keep the floor. several machine learning models have also been developed to automatically predict when the utterance of the current speak is about to end (de kok and heylen, 2009; maier et al., 2017b). roddy et al. (2018) presented a data-driven approach to predict a range of turn-taking behaviours when encountering pauses or overlaps, based on on speech-related features. skantze (2021) and ohagi et al. (2024) provide a general survey of the various approaches to turn-taking in both embodied and non-embodied speech-based dialogue systems. early deep learning approaches to end-of-turn prediction include maier et al. (2017a), which applied a long short-term memory model to the task using live acoustic features. such recurrent models are inherently incremental. using transformer lms to predict turn-taking is well represented in the recent literature. while not inherently incremental as defined above, turn-taking requires models to handle continuous input. turngpt made early use of lms to detect turn shifts in dialogue (ekstedt and skantze, 2020), with some discussion as to how the model could be used to predict endof-turn. this work was extended in inoue et al. (2024a), which uses vap which includes contrastive predictive coding of a cross-attention transformer (as a plus, the model is effective on a cpu). more recently, chiba and higashinaka (2025) also applied vap within an incremental framework (in their case, remdis (chiba et al., 2024)) and roddy and harte (2020) proposed a model of response timing that is designed for use in incremental systems; human evaluations indicated that they perceived the interaction qualities as more natural when the model was in use. shukuri et al. (2023) were also concerned with timing and turn-taking and proposed a method for using lms as meta-controllers of dialogue (where the dialogue system is made up of lms). 3.2 incrementalizing dialogue management buß et al. (2010) introduced an information state approach to incremental dm using the iu framework where the ius themselves composed the information state. in their method, they focused on the collaborative nature of many dialogues in a micro domain of playing a puzzle game. all mod110 incremental dialogue and robot action management for the age of language models figure 3: from kennington and schlangen (2021), an example of pointer, word, pos, and sem iu annotations for a sample from the localized narrative dataset. solid lines denote slls, dashed denote grins, and the dotted lines denote an alignment between two modalities. image taken from https://google.github.io/localized-narratives. ules, including asr nlu, tts, and a floor tracker were modeled at the incremental word level. the incremental dm reacted to information from the nlu, game board state change (i.e., non-linguistic relevant state actions), and the floor tracker. the central element of information was the iqud (incremental qud, following ginzburg (2012)) and is rule-based. they evaluated using an incremental and a non-incremental version of their system and found that the incremental versions were rated higher human-likeness and reactivity by human observers of recorded dialogue of both incremental and non-incremental interactions. this is promising, but limited as a methodology for incremental dialogue. the same authors followed up this work with buß and schlangen (2011) that introduced dium— dialogue incremental unit manager—that is also rule-based, but builds on their prior work by leveraging edits that can be made in a dialogue system that is built on the iu-framework. one positive aspect of incremental dialogue is that systems can respond appreciably faster than non-incremental counterparts, but a potential drawback of early response is that the response is based on information which has already been, or is currently being, updated in processing modules. for example, an asr recognizes i would like to book a train to hamm passes each word to a nlu module that informs the dm with information about which action to take and which object to take the action on. the dm begins to act by looking up train information in a database and informing the nlg about how to respond, and tts begins to vocalize the response, but at that moment hamm is revoked and replaced with hamburg. this asr update is propagated to the nlu and likewise to the dm. what action should the dm now take given that tts is currently uttering something about the wrong city? this is a shortcoming of incremental systems that needs to be addressed according to the authors. instead of reducing revisions (i.e., waiting for more information) which means waiting longer, and instead of ignoring the problem completely, dium offers a third alternative: acknowledge the problem and repair it explicitly. the iu information state is adaptable to addressing the problem directly because a revoke—an important part of the iu-framework—is an abrupt change to the information state that 111 kennington, lison and schlangen can be addressed by triggering an explicit repair, for example oops, i thought you said hamm, but it was actually hamburg. let me get that information for you. unfortunately, this line of research has not been pursued since the 2011 dium paper. however, recently, kennington and schlangen (2021) proposed a sketch of using the iu-framework as a method of representing a multimodal, fine-grained information state for use in physical, co-located settings such as hri. like buß and schlangen (2011), their sketch explained how the information state can consist of the full iu network including connections between ius, as well as all prior edits. figure 3 shows an example of a fine-grained, incremental information state using an example from the localized narrative dataset (pont-tuset et al., 2020). later work explored incremental dm using a time board where input, output, and decisions made by the dm are posted on the time board (yaghoubzadeh et al., 2015; yaghoubzadeh and kopp, 2016). the time board is an important piece of incremental dm, according to the authors, because not only does it maintain a history of the ongoing dialogue, future events (e.g., decisions) are also posted and coordinated. events that have been initiated, for example the system begins an utterance that the nlg is currently constructing and tts is uttering, can clearly show that they are not yet complete, so a new event that needs to interrupt the ongoing event can produce natural behaviour (e.g., saying um or oops, or sorry). going beyond rule-based incremental dm, selfridge and arizmendi (2012) introduced a first step towards an incremental pomdp-based system. they proposed an incremental interaction manager (iim) to mediate communication between an incremental asr and a partially-observable dm. the iim worked by evaluating potential dm decisions by applying incremental asr output to temporary instances of the dm, allowing the system to maintain multiple dm s across time and prune away dm s that are unlikely to advance the dialogue. this enables the partially observable dm to work with incremental asr n-best lists, but the work demonstrated in selfridge et al. (2012) has regrettably not been pursued further. a barge-in is when person a begins speaking, then person b attempts to take the floor while person a is still speaking. while often rude, this is common in interactive game scenarios, and it is important for a system that needs to have the ability to stop talking when a human barges in because timing is critical. selfridge et al. (2013) modeled a simple method for detecting bargeins, and pincus and traum (2017) brought together multiple aspects of incremental dialogue in a word-game task that required fast-paced dialogue where barge-in was required. see figure 4 for an example. the system had an incremental asr and learned a policy of when it should handle interruptions made while the system was speaking, and learning when to initiate barge ins. though the focus was on barge-ins, there are some useful take-always from this work: first, that people often overlap in speech. second, systems should be ready to yield the floor when they are barged-in on, and they should have the ability to barge in on a human’s ongoing speech if there are appropriate stakes involved (e.g., a system needs to inform a human of an impending problem in a nuclear facility). none of these would be possible without incremental processing, and this work shows that timing is an important aspect of the kind of policy a dm needs to learn about. manuvinakurike et al. (2017) also looked at incremental dialogue policy learning in a fast-paced game scenario where the user was presented with multiple images and needed to produce a referring expression to that object; the system was tasked with identifying which object the user was referring to. the system could highlight the image that it determined was being referred and say got it or it could suggest that the system and user move onto the next set of images (e.g., “let’s move onto the next one”) because it is unlikely to be able to refer to the correct one given the user’s utterance. 112 incremental dialogue and robot action management for the age of language models figure 4: from pincus and traum (2017), an example of human-human and game intelligent update dialogues with barge-in. the learned policy was to either wait (i.e., let the user continue speaking), as-i (i.e., the system selects what it thinks the referred object is), or as-s (i.e., skip to the next set of images). the policy was incremental in that it had to learn at each word increment which action to take. the system and user earned points for identifying images quickly, but it lost points if it referred to an image incorrectly. the system, therefore, had to learn when to wait, select the image, or determine that it was better to move on. the evaluation showed that the learned policy worked better than the hand-coded policy in that it enabled more correctly identified images within a shorter amount of time. like pincus and traum (2017), the focus of this policy revolves around timing of simple actions rather than complex actions, indicating that the purpose of an incremental dm should include handling timing decisions. incremental dm in a multi-party hri setting was reported in kennington et al. (2014a), that used the iu framework, used an independent opendial (lison and kennington, 2016) dm for every human that it detected in a game setting. overall, the dm worked effectively, but it only controlled the dialogue interaction, not robot actions. approaches to incremental dialogue state tracking have also been developed. žilka and jurčı́ček (2015) introduced lectrack, a word-level recurrent neural network state tracker model evaluated on dstc2 data. the recurrent neural network they used was a lstm because it is a kind of neural network that can be modeled to work at the word level and maintain its internal state (i.e., update-incremental) at each word increment (see figure 5). their evaluations on a subset dtsc2 dataset showed as being on-par with state-of-the-art non-incremental state trackers. there also exists various approaches to dialogue state tracking based on autoregressive lms (feng et al., 2023; hudeček and dusek, 2023), which rely on instruction-tuned llms to extract slot-value pairs from the dialogue history. 4. discussion one of the primary challenges of incremental sds in hri settings is handling uncertainty including sensory uncertainty and uncertainty that is inherent when communicating with humans. certainly, all systems are required to handle uncertainty, but the problem is more acute with incremental, multimodal systems in hri settings because they are tasked with acting on incomplete information that could be forthcoming at a later point. 113 kennington, lison and schlangen figure 5: from žilka and jurčı́ček (2015), a schematic of a lstm-based dialogue state tracker. evaluating dm is challenging in general because a proper evaluation usually amounts to a fullyworking sds with human evaluation, but the dm could be working perfectly while the asr or the nlg modules are not working properly for the task, resulting in poor evaluations from the humans. offline evaluation is difficult for two reasons: while other modules like asr, nlu, and even nlg can be evaluated with offline benchmarks, there isn’t a clear offline evaluation for dm, though the dialogue state tracking challenge is one attempt to address this. the second difficulty is that there is not a dataset that is annotated for dm at an incremental level. it is therefore unknown whether a dm should make a decision at a specific point while a user is speaking, or how to handle errors in decisions when they are in the process of being articulated either in speech or a robotic action. the work on incremental dm explained above (buß and schlangen, 2011; yaghoubzadeh et al., 2015) show methods that attempt to address these challenges, but not with properly annotated incremental data. 4.1 desiderata to address these challenges, we offer here desiderata for incremental dm on robotic platforms: • incremental dm is responsible for timing: knowing not just what decision to make, but when to make that decision are both important in fast-paced, incremental settings, particularly on robotic, embodied platforms where additional modalities play a role in understanding and interaction between user and system. devault et al. (2009) explored how learning when to respond to incremental results affects task success, and kennington and schlangen (2016) used a rule-based dm to make timing decisions on when to settle on a final decision on which action to take. recent work relevant to hri tasks in this area include zhang et al. (2025) that used a full-duplex dm (i.e., the asr was always producing input) based on voice activity detection to help determine if the system should wait or act (audiopalm can likewise ”speak and listen”, which is a step in the right direction (rubenstein et al., 2023), and yaghoubzadeh et al. (2015)’s model of dm that used a time board is a likely good place to start. • incremental dm needs to act on incomplete information: when humans interact with each other, there are often backchannels (e.g., nodding) that signal understanding, or as someone is speaking a listener can signal understanding by taking an action. for example, if a speaker 114 incremental dialogue and robot action management for the age of language models makes a request can you hand me the green book on the left? the listener can already be turning and reaching for a green book before the utterance is complete. a robot that interacts with a person where the task involves handling objects should act in a similar way; for example, reaching for an object or driving towards a destination. if indeed the system made the wrong decision about which object to pursue, then the robot can change its course, but it’s important that the robot act as soon as it has enough information to act, even if that act might be incorrect; the movement signals to the user that the robot is in the process of understanding. • incremental dm needs to make fast, small decisions concurrently: traditional dms often take in all information from the user during their turn then make a high-level decision once which can then potentially inform multiple modules like nlg to speak and a robot arm to move. an incremental dm, in contrast, needs to make smaller decisions that may lead to a final outcome, but the outcome may not yet be known. this is similar in principle to acting on incomplete information, but the nature of the actions is more fine-grained. for example, the dm may know that the user wants a robot to fetch an object in the kitchen, and though the robot doesn’t know which object, it makes a smaller decision to move to the kitchen, and by the time it arrives in the kitchen it knows more about the specifics of the object that it is requested to retrieve. • incremental llms: transformer llm architectures generate output incrementally and some aspects of input are incremental, but they are not completely update-incremental. using them in a restart-incremental manner is computationally expensive, but recent work has shown that minor changes to the model can improve incremental metrics and reduce computational overhead (kahardipraja et al., 2021). llms are being used in many ways in robotics (see, for example, singh et al. (2023) that uses llms for robot action planning), but work needs to be done for incremental processing on llms in hri settings. recommendations there is a lack of incremental datasets. most datasets can be used for incremental training and evaluation for some modules (e.g., asr or nlg), but nlu and dm modules that produce incremental output that is on a different level of granularity than the word level, so it is unclear from nlu datasets as to when a slot should be filled or when the dm should make a decision. efforts towards a dataset that has incremental annotations would be very beneficial to research in the setting of dialogue with robots. models may or may not need to be trained incrementally, but evaluation metrics should be on the incremental level. 5. conclusion in this article, we reviewed literature relating to incremental dialogue management motivated by the need for incremental dialogue management in robotic platforms. we showed that there is ample work in incremental processing, but very little in incremental dialogue management itself. the review resulted in several key desiderata for incremental dialogue management, particularly needed in spoken dialogue-enabled human robot interaction. clearly, a decision-making module is a critical component in a robot that can interact with people using spoken dialogue. taken together, this review in conjunction with other recent review work from reimann et al. (2024) and lison and kennington (2023) are useful for robotics researchers who are interested in designing effective dialogue strategies between robots and humans. 115 kennington, lison and schlangen acknowledgments we would like to thank the reviewers and editors for their helpful feedback. references gregory aist, james allen, ellen campana, carlos gomez gallo, scott stoness, and mary swift. incremental understanding in human-computer dialogue and experimental evidence for advantages over nonincremental methods. in pragmatics, volume 1, pages 149–154, trento, italy, 2007. james f allen, d byron, m dzikovska, g ferguson, lucian galescu, and amanda stent. toward conversational human-computer interaction. ai mag., 22(4):27–38, december 2001. james frederick allen. a plan-based approach to speech act recognition. phd thesis, 1979. dario amodei, sundaram ananthanarayanan, rishita anubhai, jingliang bai, eric battenberg, carl case, jared casper, bryan catanzaro, qiang cheng, guoliang chen, jie chen, jingdong chen, zhijie chen, mike chrzanowski, adam coates, greg diamos, ke ding, niandong du, erich elsen, jesse engel, weiwei fang, linxi fan, christopher fougner, liang gao, caixia gong, awni hannun, tony han, lappi vaino johannes, bing jiang, cai ju, billy jun, patrick legresley, libby lin, junjie liu, yang liu, weigao li, xiangang li, dongpeng ma, sharan narang, andrew ng, sherjil ozair, yiping peng, ryan prenger, sheng qian, zongfeng quan, jonathan raiman, vinay rao, sanjeev satheesh, david seetapun, shubho sengupta, kavya srinet, anuroop sriram, haiyuan tang, liliang tang, chong wang, jidong wang, kaifu wang, yi wang, zhijian wang, zhiqian wang, shuang wu, likai wei, bo xiao, wen xie, yan xie, dani yogatama, bin yuan, jun zhan, and zhenyao zhu. deep speech 2: end-to-end speech recognition in english and mandarin. in proceedings of the 33rd international conference on international conference on machine learning volume 48, icml’16, page 173–182, new york, ny, usa, 2016. jmlr.org. layla el asri, romain laroche, olivier pietquin, and hatim khouzaimi. nastia: negotiating appointment setting interface. in proceedings of lrec, pages 266–271, 2014. alexei baevski, henry zhou, abdelrahman mohamed, and michael auli. wav2vec 2.0: a framework for self-supervised learning of speech representations. arxiv [cs.cl], june 2020. timo baumann. partial representations improve the prosody of incremental speech synthesis. in proceedings of the annual conference of the international speech communication association, interspeech, pages 2932–2936, 2014. timo baumann and david schlangen. the inprotk 2012 release. in naacl-hlt workshop on future directions and needs in the spoken dialog community: tools and data (sdctd 2012), pages 29–32, 2012. timo baumann, okko buß, michaela atterer, and david schlangen. evaluating the potential utility of asr n-best lists for incremental spoken dialogue systems. proceedings of the annual conference of the international speech communication association, interspeech, pages 1031– 1034, 2009. 116 incremental dialogue and robot action management for the age of language models timo baumann, casey kennington, julian hough, and david schlangen. recognising conversational speech: what an incremental asr should do for a dialogue system and how to get there. in proceedings of the international workshop series on spoken dialogue systems technology (iwsds) 2016, 2016. niels beuck and wolfgang menzel. structural prediction in incremental dependency parsing. in lecture notes in computer science (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics), volume 7816 lncs, pages 245–257, 2013. paweł budzianowski, tsung-hsien wen, bo-hsiang tseng, inigo casanueva, stefan ultes, osman ramadan, and milica gašić. multiwoz–a large-scale multi-domain wizard-of-oz dataset for taskoriented dialogue modelling. arxiv preprint arxiv:1810.00278, 2018. hendrik buschmeier, timo baumann, benjamin dosch, stefan kopp, and david schlangen. combining incremental language generation and incremental speech synthesis for adaptive information presentation. in proceedings of the 13th annual meeting of the special interest group on discourse and dialogue, pages 295–303, seoul, south korea, july 2012. association for computational linguistics. okko buß and david schlangen. dium an incremental dialogue manager that can produce selfcorrections. in proceedings of semdial 2011 (los angelogue), proceedings of semdial 2011 (los angelogue), 2011. okko buß, timo baumann, and david schlangen. collaborating on utterances with a spoken dialogue system using an isu-based approach to incremental dialogue management. in proceedings of sigdial, pages 233–236, tokyo, japan, september 2010. michael cashmore, andrew coles, bence cserna, erez karpas, daniele magazzeni, and wheeler ruml. replanning for situated robots. proceedings of the international conference on automated planning and scheduling, 29(1):665–673, july 2019. senthilkumar chandramohan, matthieu geist, fabrice lefevre, and olivier pietquin. user simulation in dialogue systems using inverse reinforcement learning. in interspeech 2011, pages 1025–1028, 2011. zhi chen, yuncong liu, lu chen, su zhu, mengyue wu, and kai yu. opal: ontology-aware pretrained language model for end-to-end task-oriented dialogue. transactions of the association for computational linguistics, 11:68–84, 01 2023. issn 2307-387x. doi: 10.1162/tacl a 00534. url https://doi.org/10.1162/tacl_a_00534. yuya chiba and ryuichiro higashinaka. investigating the impact of incremental processing and voice activity projection on spoken dialogue systems. in proceedings of the 31st international conference on computational linguistics, pages 3687–3696, 2025. yuya chiba, koh mitsuda, akinobu lee, and ryuichiro higashinaka. the remdis toolkit: building advanced real-time multimodal dialogue systems with incremental processing and large language models. in proceedings of iwsds, pages 1–6, 2024. 117 https://doi.org/10.1162/tacl_a_00534 kennington, lison and schlangen phil cohen. steps towards collaborative multimodal dialogue (sustained contribution award). in proceedings of the 19th acm international conference on multimodal interaction, icmi ’17, page 4, new york, ny, usa, 2017. association for computing machinery. isbn 9781450355438. doi: 10.1145/3136755.3154480. url https://doi.org/10.1145/ 3136755.3154480. philip r cohen and lucian galescu. a planning-based explainable collaborative dialogue system. arxiv [cs.ai], february 2023. philip r cohen and hector j levesque. intention is choice with commitment. artificial intelligence, 42(2):213–261, 1990. vanya cohen, jason xinyu liu, raymond mooney, stefanie tellex, and david watkins. a survey of robotic language grounding: tradeoffs between symbols and embeddings. arxiv [cs.ro], may 2024. ann copestake. semantic composition with (robust) minimal recursion semantics. in proceedings of the workshop on deep linguistic processing deeplp ’07, page 73, morristown, nj, usa, 2007. association for computational linguistics. marco damonte, shay b cohen, and giorgio satta. an incremental parser for abstract meaning representation. in european chapter of the association for computational linguistics (eacl), volume 1, pages 536–546, 2017. iwan de kok and dirk heylen. multimodal end-of-turn prediction in multi-party meetings. in proceedings of the 2009 international conference on multimodal interfaces, pages 91–98, 2009. nina dethlefs, helen hastie, verena rieser, and oliver lemon. optimising incremental generation for spoken dialogue systems: reducing the need for fillers. in barbara di eugenio and susan mcroy, editors, inlg 2012 proceedings of the seventh international natural language generation conference, pages 49–58, utica, il, may 2012a. association for computational linguistics. nina dethlefs, verena rieser, helen hastie, and oliver lemon. towards optimising modality allocation for multimodal output generation in incremental dialogue. in proceedings of the ecai workshop on machine learning for interactive syustems, pages 31–36, montepellier, france, 2012b. david devault and david traum. a demonstration of incremental speech understanding and confidence estimation in a virtual human dialogue system. in gary geunbae lee, jonathan ginzburg, claire gardent, and amanda stent, editors, proceedings of the 13th annual meeting of the special interest group on discourse and dialogue, pages 131–133, seoul, south korea, july 2012. association for computational linguistics. david devault and david traum. a method for the approximation of incremental understanding of explicit utterance meaning using predictive models in finite domains. in proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 1092–1099, 2013. 118 https://doi.org/10.1145/3136755.3154480 https://doi.org/10.1145/3136755.3154480 incremental dialogue and robot action management for the age of language models david devault, kenji sagae, and david traum. can i finish?: learning when to respond to incremental interpretation results in interactive dialogue. annual meeting of the special interest group on discourse and dialogue (sigdial), (september):11–20, 2009. david devault, kenji sagae, and david traum. incremental interpretation and prediction of utterance meaning for interactive dialogue. dialogue and discourse, 2(1):143–170, 2011. danny driess, fei xia, mehdi s m sajjadi, corey lynch, aakanksha chowdhery, brian ichter, ayzaan wahid, jonathan tompson, quan vuong, tianhe yu, wenlong huang, yevgen chebotar, pierre sermanet, daniel duckworth, sergey levine, vincent vanhoucke, karol hausman, marc toussaint, klaus greff, andy zeng, igor mordatch, and pete florence. palm-e: an embodied multimodal language model. arxiv [cs.lg], march 2023. jens edlund, joakim gustafson, mattias heldner, and anna hjalmarsson. towards human-like spoken dialogue systems. speech communication, 50(8):630–645, 2008. erik ekstedt and gabriel skantze. turngpt: a transformer-based language model for predicting turn-taking in spoken dialog. in findings of the association for computational linguistics: emnlp 2020, pages 2981–2990, stroudsburg, pa, usa, november 2020. association for computational linguistics. arash eshghi, matthew purver, and julian hough. probabilistic induction for an incremental semantic grammar. in proceedings of the 10th international conference on computational semantics (iwcs 2013) – long papers, pages 107–118, potsdam, germany, 2013. friederike eyssel and frank hegel. (s)he’s got the look: gender stereotyping of robots. journal of applied social psychology, 42(9):2213–2230, 2012. yujie feng, zexin lu, bo liu, liming zhan, and xiao-ming wu. towards llm-driven dialogue state tracking. in houda bouamor, juan pino, and kalika bali, editors, proceedings of the 2023 conference on empirical methods in natural language processing, pages 739–755, stroudsburg, pa, usa, 2023. association for computational linguistics. luciana ferrer, elizabeth shriberg, and andreas stolcke. is the speaker done yet? faster and more accurate end-of-utterance detection using prosody. in seventh international conference on spoken language processing, 2002. clayton fields and casey kennington. vision language transformers: a survey. arxiv [cs.cv], july 2023. kerstin fischer, lakshadeep naik, rosalyn m langedijk, timo baumann, matouš jelı́nek, and oskar palinko. initiating human-robot interactions using incremental speech adaptation. in companion of the 2021 acm/ieee international conference on human-robot interaction, hri ’21 companion, pages 421–425, new york, ny, usa, march 2021. association for computing machinery. caelan reed garrett, chris paxton, tomas lozano-perez, leslie pack kaelbling, and dieter fox. online replanning in belief space for partially observable task and motion problems. in 2020 ieee international conference on robotics and automation (icra), pages 5678–5684. ieee, 2020. 119 kennington, lison and schlangen fabrizio ghigi, maxine eskenazi, m ines torres, and sungjin lee. incremental dialog processing in a task-oriented dialog. in fifteenth annual conference of the international speech communication association, 2014a. fabrizio ghigi, maxine eskenazi, m ines torres, and sungjin lee. incremental dialog processing in a task-oriented dialog. in proceedings of the annual conference of the international speech communication association, interspeech, pages 308–312, singapore, 2014b. jonathan ginzburg. the interactive stance. oxford university press, 2012. sachin goyal, ziwei ji, ankit singh rawat, aditya krishna menon, sanjiv kumar, and vaishnavh nagarajan. think before you speak: training language models with pause tokens. arxiv [cs.cl], october 2023. david griol, lluı́s f hurtado, encarna segarra, and emilio sanchis. a statistical approach to spoken dialog systems design and evaluation. speech communication, 50(8-9):666–682, 2008. jiawei gu, xuhui jiang, zhichao shi, hexiang tan, xuehao zhai, chengjin xu, wei li, yinghan shen, shengjie ma, honghao liu, saizhuo wang, kun zhang, yuanzhuo wang, wen gao, lionel ni, and jian guo. a survey on llm-as-a-judge, 2025. url https://arxiv.org/abs/ 2411.15594. hany hassan, khalil simaan, and andy way. a syntactic language model based on incremental ccg parsing. in 2008 ieee workshop on spoken language technology, slt 2008 proceedings, pages 205–208. institute of electrical and electronics engineers, 2008. michael heck, carel van niekerk, nurul lubis, christian geishauser, hsien-chin lin, marco moresi, and milica gašic. trippy: a triple copy strategy for value independent neural dialog state tracking. in 21th annual meeting of the special interest group on discourse and dialogue, page 35, 2020. ehsan hosseini-asl, bryan mccann, chien-sheng wu, semih yavuz, and richard socher. a simple language model for task-oriented dialogue. advances in neural information processing systems, 33:20179–20191, 2020. julian hough and david schlangen. investigating fluidity for human-robot interaction with realtime, real-world grounding strategies. in proceedings of the 17th annual meeting of the special interest group on discourse and dialogue, pages 288–298, los angeles, september 2016. association for computational linguistics. julian hough, casey kennington, david schlangen, and jonathan ginzburg. incremental semantics for dialogue processing : requirements , and a comparison of two approaches. in proceedings of iwcs, pages 206–216. association for computational linguistics, 2015. vojtěch hudeček and ondrej dusek. are large language models all you need for task-oriented dialogue? in proceedings of the 24th meeting of the special interest group on discourse and dialogue, pages 216–228, stroudsburg, pa, usa, 2023. association for computational linguistics. 120 https://arxiv.org/abs/2411.15594 https://arxiv.org/abs/2411.15594 incremental dialogue and robot action management for the age of language models kyuyeon hwang and wonyong sung. character-level incremental speech recognition with recurrent neural networks. in icassp, ieee international conference on acoustics, speech and signal processing proceedings, volume 2016-may, pages 5335–5339, january 2016. saki imai, tahiya chowdhury, and amanda j stent. evaluating open-source asr systems: performance across diverse audio conditions and error correction methods. in owen rambow, leo wanner, marianna apidianaki, hend al-khalifa, barbara di eugenio, and steven schockaert, editors, proceedings of the 31st international conference on computational linguistics, pages 5027–5039, abu dhabi, uae, january 2025. association for computational linguistics. koji inoue, bing’er jiang, erik ekstedt, tatsuya kawahara, and gabriel skantze. real-time and continuous turn-taking prediction using voice activity projection. arxiv [cs.cl], january 2024a. koji inoue, divesh lala, gabriel skantze, and tatsuya kawahara. yeah, un, oh: continuous and real-time backchannel prediction with fine-tuning of voice activity projection. arxiv [cs.cl], october 2024b. youngsoo jang, jongmin lee, jaeyoung park, kyeng hun lee, pierre lison, and kee eung kim. pyopendial: a python-based domain-independent toolkit for developing spoken dialogue systems with probabilistic rules. in emnlp-ijcnlp 2019 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing, proceedings of system demonstrations, pages 187–192, 2020. patrick kahardipraja, brielen madureira, and david schlangen. towards incremental transformers: an empirical analysis of transformer models for incremental nlu. in proceedings of the 2021 conference on empirical methods in natural language processing, pages 1178–1189, online and punta cana, dominican republic, november 2021. association for computational linguistics. patrick kahardipraja, brielen madureira, and david schlangen. tapir: learning adaptive revision for incremental natural language understanding with a two-pass model. in findings of the association for computational linguistics: acl 2023, pages 4173–4197, stroudsburg, pa, usa, july 2023. association for computational linguistics. gaoussou youssouf kebe, luke e richards, edward raff, francis ferraro, and cynthia matuszek. bridging the gap: using deep acoustic representations to learn grounded language from percepts and raw speech. proceedings of the aaai conference on artificial intelligence, 36(10):10884– 10893, june 2022. john d kelleher and g-j geert-jan kruijff. incremental generation of spatial referring expressions in situated dialog. in proceedings of the 21st international conference on computational linguistics and the 44th annual meeting of the association for computational linguistics (colingacl’06), pages 1041–1048, 2006. casey kennington and david schlangen. situated incremental natural language understanding using markov logic networks. computer speech & language, 2014. casey kennington and david schlangen. simple learning and compositional application of perceptually grounded word meanings for incremental reference resolution. in proceedings of the 53rd 121 kennington, lison and schlangen annual meeting of the association for computational linguistics and the 7th international joint conference on natural language processing (volume 1: long papers), pages 292–301, beijing, china, 2015. association for computational linguistics. casey kennington and david schlangen. supporting spoken assistant systems with a graphical user interface that signals incremental understanding and prediction state. in proceedings of the 17th annual meeting of the special interest group on discourse and dialogue, pages 242–251, los angeles, september 2016. association for computational linguistics. casey kennington and david schlangen. a simple generative model of incremental reference resolution in situated dialogue. computer speech & language, 2017. casey kennington and david schlangen. incremental unit networks for multimodal, fine-grained information state representation. in proceedings of the 1st workshop on multimodal semantic representations (mmsr), pages 89–94, groningen, netherlands (online), june 2021. association for computational linguistics. casey kennington, kotaro funakoshi, yuki takahashi, and mikio nakano. probabilistic multiparty dialogue management for a game master robot. in proceedings of the 2014 acm/ieee international conference on human-robot interaction hri ’14, pages 200–201, bielefeld, germany, 2014a. acm. casey kennington, spyros kousidis, and david schlangen. situated incremental natural language understanding using a multimodal, linguistically-driven update model. in proceedings of coling 2014, pages 1803–1812, 2014b. casey kennington, ryu iida, takenobu tokunaga, and david schlangen. incrementally tracking reference in human / human dialogue using linguistic and extra-linguistic information. in hltnaacl 2015 human language technology conference of the north american chapter of the association of computational linguistics, proceedings of the main conference, pages 272–282, denver, u.s.a., 2015. association for computational linguistics. casey kennington, ting han, and david schlangen. temporal alignment using the incremental unit framework. in proceedings of the 19th acm international conference on multimodal interaction, icmi 2017, pages 297–301, new york, ny, usa, 2017. acm. casey kennington, daniele moro, lucas marchand, jake carns, and david mcneill. rrsds: towards a robot-ready spoken dialogue system. in proceedings of the 21th annual meeting of the special interest group on discourse and dialogue, virtual, 2020. association for computational linguistics. hatim khouzaimi, romain laroche, and fabrice lefevre. an easy method to make dialogue systems incremental. in proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (sigdial), pages 98–107, philadelphia, pa, u.s.a., 2014. association for computational linguistics. spyros kousidis, casey kennington, timo baumann, hendrik buschmeier, stefan kopp, and stefan schlangen. situationally aware in-car information presentation using incremental speech generation: safer, and more effective. in proceedings of the workshop on dialogue in motion (dm), eacl 2014, pages 68–72, 2014. 122 incremental dialogue and robot action management for the age of language models arne köhn. incremental natural language processing: challenges, strategies, and evaluation. in emily m bender, leon derczynski, and pierre isabelle, editors, proceedings of the 27th international conference on computational linguistics, pages 2990–3003, santa fe, new mexico, usa, august 2018. association for computational linguistics. staffan larsson. issue-based dialogue management. phd thesis, göteborg university, 2002. pierre lison. a hybrid approach to dialogue management based on probabilistic rules. computer speech & language, 34(1):232–255, 2015a. pierre lison. a hybrid approach to dialogue management based on probabilistic rules. computer speech & language, 34(1):232–255, 2015b. pierre lison and casey kennington. opendial: a toolkit for developing spoken dialogue systems with probabilistic rules. in 54th annual meeting of the association for computational linguistics, acl 2016 system demonstrations, 2016. pierre lison and casey kennington. who’s in charge? roles and responsibilities of decision-making components in conversational robots. in proceedings of the workshop on human-robot conversational interaction, march 2023. bing liu, gökhan tür, dilek hakkani-tur, pararth shah, and larry heck. dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems. in proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long papers), pages 2060–2069, 2018. brielen madureira and david schlangen. incremental processing in the age of non-incremental encoders: an empirical assessment of bidirectional models for incremental nlu. proceedings of the 2020 conference on empirical methods in natural language processing, pages 357–374, 2020. brielen madureira, patrick kahardipraja, and david schlangen. when only time will tell: interpreting how transformers process local ambiguities through the lens of restart-incrementality. in lun-wei ku, andre martins, and vivek srikumar, editors, proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: long papers), pages 4722–4749, bangkok, thailand, august 2024. association for computational linguistics. karthik mahadevan, jonathan chien, noah brown, zhuo xu, carolina parada, fei xia, andy zeng, leila takayama, and dorsa sadigh. generative expressive robot behaviors using large language models. arxiv [cs.ro], january 2024. angelika maier, julian hough, and david schlangen. towards deep end-of-turn prediction for situated spoken dialogue systems. in proceedings interspeech 2017, pages 1676–1680, 2017a. angelika maier, julian hough, and david schlangen. towards deep end-of-turn prediction for situated spoken dialogue systems. proceedings of interspeech 2017, 2017b. anna manaseryan, porter rigby, brooke matthews, catherine henry, josue torres-fonseca, ryan whetten, enoch levandovsky, and casey kennington. rrsds 2.0: incremental, modular, distributed, multimodal spoken dialogue with robotic platforms. in proceedings of the 26th annual 123 kennington, lison and schlangen meeting of the special interest group on discourse and dialogue, avignon, france, 2025. association for computational linguistics. ramesh manuvinakurike, david devault, and kallirroi georgila. using reinforcement learning to model incrementality in a fast-paced dialogue game. in proceedings of the 18th annual sigdial meeting on discourse and dialogue, pages 331–341, saarbrücken, germany, august 2017. association for computational linguistics. thilo michael and sebastian möller. retico: an open-source framework for modeling real-time conversations in spoken dialogue systems. in tagungsband der 30. konferenz elektronische sprachsignalverarbeitung 2019, essv, pages 134–140, dresden, march 2019. tudpress, dresden. chinmaya mishra, rinus verdonschot, peter hagoort, and gabriel skantze. real-time emotion generation in human-robot dialogue using large language models. front robot ai, 10:1271610, december 2023. nikola mrkšić, diarmuid ó séaghdha, tsung-hsien wen, blaise thomson, and steve young. neural belief tracker: data-driven dialogue state tracking. in proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: long papers), pages 1777–1788, 2017. jinjie ni, tom young, vlad pandelea, fuzhao xue, and erik cambria. recent advances in deep learning based dialogue systems: a systematic survey. artificial intelligence review, 56(4):3055– 3155, 2023. cheng niu, xingguang wang, xuxin cheng, juntong song, and tong zhang. enhancing dialogue state tracking models through llm-backed user-agents simulation. in proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: long papers), pages 8724–8741, stroudsburg, pa, usa, 2024. association for computational linguistics. joakim nivre. algorithms for deterministic incremental dependency parsing. computational linguistics, 34(4):513–553, 2008. jekaterina novikova, gang ren, and leon watts. it’s not the way you look, it’s how you move: validating a general scheme for robot affective behaviour. in julio abascal, simone barbosa, mirko fetter, tom gross, philippe palanque, and marco winckler, editors, human-computer interaction – interact 2015, pages 239–258, cham, 2015. springer international publishing. masaya ohagi, tomoya mizumoto, and katsumasa yoshikawa. investigation of look-ahead techniques to improve response time in spoken dialogue system. in interspeech 2024, pages 3580– 3584, 2024. long ouyang, jeff wu, xu jiang, diogo almeida, carroll l wainwright, pamela mishkin, chong zhang, sandhini agarwal, katarina slama, alex ray, john schulman, jacob hilton, fraser kelton, luke miller, maddie simens, amanda askell, peter welinder, paul christiano, jan leike, and ryan lowe. training language models to follow instructions with human feedback. arxiv [cs.cl], pages 27730–27744, 2022. 124 incremental dialogue and robot action management for the age of language models maike paetzel, ramesh manuvinakurike, and david devault. “so, which one is it?” the effect of alternative incremental architectures in a high-performance game-playing agent. in proceedings of the 16th annual meeting of the special interest group on discourse and dialogue, pages 77–86, prague, czech republic, september 2015. association for computational linguistics. qian pan, zahra ashktorab, michael desmond, martı́n santillán cooper, james johnson, rahul nair, elizabeth daly, and werner geyer. human-centered design recommendations for llm-asa-judge. in nikita soni, lucie flek, ashish sharma, diyi yang, sara hooker, and h. andrew schwartz, editors, proceedings of the 1st human-centered large language modeling workshop, pages 16–29, tbd, august 2024. acl. doi: 10.18653/v1/2024.hucllm-1.2. url https: //aclanthology.org/2024.hucllm-1.2/. andreas peldszus, okko buß, timo baumann, and david schlangen. joint satisfaction of syntactic and pragmatic constraints improves incremental spoken language understanding. in proceedings of the 13th eacl, pages 514–523, avignon, france, april 2012. association for computational linguistics. baolin peng, xiujun li, jianfeng gao, jingjing liu, and kam-fai wong. deep dyna-q: integrating planning for task-completion dialogue policy learning. in proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: long papers), pages 2182–2192, 2018. volha petukhova and harry bunt. incremental dialogue act understanding. pragmatics, pages 235–244, 2011. eli pincus and david traum. an incremental response policy in an automatic word-game. in ceur workshop proceedings, volume 1943, pages 1–8, 2017. sarah plane, ariel marvasti, tyler egan, and casey kennington. predicting perceived age: both language ability and appearance are important. in proceedings of sigdial, 2018. jordi pont-tuset, jasper uijlings, soravit changpinyo, radu soricut, and vittorio ferrari. connecting vision and language with localized narratives. in computer vision – eccv 2020, volume 12350 lncs, pages 647–664. springer international publishing, 2020. libo qin, yijia liu, wanxiang che, haoyang wen, yangming li, and ting liu. entity-consistent end-to-end task-oriented dialogue system with kb retriever. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), pages 133–142, 2019. libo qin, xiao xu, lehan wang, yue zhang, and wanxiang che. modularized pre-training for end-to-end task-oriented dialogue. ieee/acm transactions on audio, speech, and language processing, 31:1601–1610, 2023. doi: 10.1109/taslp.2023.3244503. alec radford, jong wook kim, tao xu, greg brockman, christine mcleavey, and ilya sutskever. robust speech recognition via large-scale weak supervision, 2022. antoine raux and maxine eskenazi. a finite-state turn-taking model for spoken dialog systems. in proceedings of human language technologies: the 2009 annual conference of the north american chapter of the association for computational linguistics, pages 629–637, 2009. 125 https://aclanthology.org/2024.hucllm-1.2/ https://aclanthology.org/2024.hucllm-1.2/ kennington, lison and schlangen natalia reich-stiebert and friederike eyssel. (ir)relevance of gender? on the influence of gender stereotypes on learning with a robot. in proceedings of the 2017 acm/ieee international conference on human-robot interaction, hri ’17, page 166–176, new york, ny, usa, 2017. association for computing machinery. merle m reimann, florian a kunneman, catharine oertel, and koen v hindriks. a survey on dialogue management in human-robot interaction. j. hum.-robot interact., march 2024. liliang ren, kaige xie, lu chen, and kai yu. towards universal dialogue state tracking. in proceedings of the 2018 conference on empirical methods in natural language processing, pages 2780–2786, 2018. verena rieser and oliver lemon. reinforcement learning for adaptive dialogue systems: a datadriven methodology for dialogue management and natural language generation. springer science & business media, 2011. m roddy, gabriel skantze, and n harte. investigating speech features for continuous turn-taking prediction using lstms. in 19th annual conference of the international speech communication, interspeech 2018; hyderabad international convention centre (hicc) hyderabad; india; 2 september 2018 through 6 september 2018, pages 586–590. international speech communication association, 2018. matthew roddy and naomi harte. neural generation of dialogue response timings. in proceedings of the 58th annual meeting of the association for computational linguistics, pages 2442–2452, stroudsburg, pa, usa, 2020. association for computational linguistics. stephen roller, y-lan boureau, jason weston, antoine bordes, emily dinan, angela fan, david gunning, da ju, margaret li, spencer poff, et al. open-domain conversational agents: current progress, open problems, and future directions. arxiv preprint arxiv:2006.12442, 2020. paul k rubenstein, chulayuth asawaroengchai, duc dung nguyen, ankur bapna, zalán borsos, félix de chaumont quitry, peter chen, dalia el badawy, wei han, eugene kharitonov, hannah muckenhirn, dirk padfield, james qin, danny rozenberg, tara sainath, johan schalkwyk, matt sharifi, michelle tadmor ramanovich, marco tagliasacchi, alexandru tudor, mihajlo velimirović, damien vincent, jiahui yu, yongqiang wang, vicky zayats, neil zeghidour, yu zhang, zhishuai zhang, lukas zilka, and christian frank. audiopalm: a large language model that can speak and listen. arxiv [cs.cl], june 2023. jost schatzmann, karl weilhammer, matt stuttle, and steve young. a survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies. the knowledge engineering review, 21(2):97–126, 2006. david schlangen and gabriel skantze. a general, abstract model of incremental dialogue processing. in proceedings of eacl, 2009. david schlangen and gabriel skantze. a general, abstract model of incremental dialogue processing. dialogue & discourse, 2(1):83–111, may 2011. 126 incremental dialogue and robot action management for the age of language models david schlangen, timo baumann, and michaela atterer. incremental reference resolution: the task, metrics for evaluation, and a bayesian filtering model that is sensitive to disfluencies. in proceedings of the 10th sigdial, pages 30–37, london, uk, 2009. association for computational linguistics. david schlangen, sina zarriess, and casey kennington. resolving references to objects in photographs using the words-as-classifiers model. in proceedings of the 54th annual meeting of the association for computational linguistics, pages 1213–1223, 2016. ethan selfridge and iker arizmendi. integrating incremental speech recognition and pomdp-based dialogue systems. in in proceedings of dialogue and discourse, page 275–279, seoul, south korea, july 2012. association for computational linguistics. ethan selfridge, iker arizmendi, peter heeman, and jason williams. continuously predicting and processing barge-in during a live spoken dialogue task. in proceedings of the sigdial 2013 conference, pages 384–393, metz, france, august 2013. association for computational linguistics. ethan o selfridge, peter a heeman, iker arizmendi, and jason d williams. demonstrating the incremental interaction manager in an end-toend “lets go!” dialogue system,”. in proc. of ieee workshop on spoken language technology, 2012. pararth shah, dilek hakkani-tur, bing liu, and gökhan tür. bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning. in proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 3 (industry papers), pages 41–51, 2018. igor shalyminov, arash eshghi, and oliver lemon. challenging neural dialogue models with natural data: memory networks fail on incremental phenomena. in semdial 2017 (saardial) workshop on the semantics and pragmatics of dialogue, isca, august 2017. isca. zhihong shao, peiyi wang, qihao zhu, runxin xu, junxiao song, mingchuan zhang, y k li, y wu, and daya guo. deepseekmath: pushing the limits of mathematical reasoning in open language models. arxiv [cs.cl], february 2024. prashanth gurunath shivakumar, naveen kumar, panayiotis georgiou, and shrikanth narayanan. incremental online spoken language understanding. october 2019. kotaro shukuri, ryoma ishigaki, jundai suzuki, tsubasa naganuma, takuma fujimoto, daisuke kawakubo, masaki shuzo, and eisaku maeda. meta-control of dialogue systems using large language models. arxiv [cs.ro], december 2023. ishika singh, valts blukis, arsalan mousavian, ankit goyal, danfei xu, jonathan tremblay, dieter fox, jesse thomason, and animesh garg. progprompt: generating situated robot task plans using large language models. in 2023 ieee international conference on robotics and automation (icra), pages 11523–11530, may 2023. ishika singh, david traum, and jesse thomason. twostep: multi-agent task planning using classical planners and large language models. arxiv [cs.ai], march 2024. 127 kennington, lison and schlangen gabriel skantze. turn-taking in conversational systems and human-robot interaction: a review. computer speech & language, 67:101178, 2021. pei-hao su, milica gasic, nikola mrkšić, lina m rojas barahona, stefan ultes, david vandyke, tsung-hsien wen, and steve young. on-line active reward learning for policy optimisation in spoken dialogue systems. in proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), pages 2431–2441, 2016. pei-hao su, milica gašić, and steve young. reward estimation for dialogue policy optimisation. computer speech & language, 51:24–43, 2018. ryuichi takanobu, hanlin zhu, and minlie huang. guided dialog policy learning: reward estimation for multi-domain task-oriented dialog. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), pages 100–110, 2019. michael k tanenhaus and michael j spivey-knowlton. integration of visual and linguistic information in spoken language comprehension. science, 268(5217):1632, 1995. blaise thomson and steve young. bayesian update of dialogue state: a pomdp framework for spoken dialogue systems. computer speech & language, 24(4):562–588, 2010. david traum and staffan larsson. the information state approach to dialogue management. in current and new directions in discourse and dialogue, pages 325–353. springer, 2003. stefan ultes, lina m rojas barahona, pei-hao su, david vandyke, dongho kim, inigo casanueva, paweł budzianowski, nikola mrkšić, tsung-hsien wen, milica gasic, et al. pydial: a multidomain statistical dialogue system toolkit. in proceedings of acl 2017, system demonstrations, pages 73–78, 2017. stefan ultes, paweł budzianowski, iñigo casanueva, lina m. rojas-barahona, bo-hsiang tseng, yen-chen wu, steve young, and milica gašić. addressing objects and their relations: the conversational entity dialogue model. in proceedings of the 19th annual sigdial meeting on discourse and dialogue, pages 273–283, melbourne, australia, july 2018. association for computational linguistics. doi: 10.18653/v1/w18-5032. url https://aclanthology.org/ w18-5032. herwin van welbergen, dennis reidsma, and stefan kopp. an incremental multimodal realizer for behavior co-articulation and coordination. in lecture notes in computer science (including subseries lecture notes in artificial intelligence and lecture notes in bioinformatics), volume 7502 lnai, pages 175–188, 2012. nicolas wagner and stefan ultes. on the controllability of large language models for dialogue interaction. in proceedings of the 25th annual meeting of the special interest group on discourse and dialogue, pages 216–221, stroudsburg, pa, usa, 2024. association for computational linguistics. nicholas thomas walker, stefan ultes, and pierre lison. graphwoz: dialogue management with conversational knowledge graphs. arxiv preprint arxiv:2211.12852, 2022. 128 https://aclanthology.org/w18-5032 https://aclanthology.org/w18-5032 incremental dialogue and robot action management for the age of language models peng wang, shijie wang, junyang lin, shuai bai, xiaohuan zhou, jingren zhou, xinggang wang, and chang zhou. one-peace: exploring one general representation model toward unlimited modalities. arxiv preprint arxiv:2305. 11172, 2023. ryan whetten, enoch levandovsky, mir tahsin imtiaz, and casey kennington. evaluating automatic speech recognition and natural language understanding in an incremental setting. in proceedings of the 27th workshop on the semantics and pragmatics of dialogue full papers, maribor, slovenia, august 2023. semdial. jason d williams, antoine raux, and matthew henderson. the dialog state tracking challenge series: a review. dialogue & discourse, 7(3):4–33, 2016. jason d williams, kavosh asadi atui, and geoffrey zweig. hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning. in proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: long papers), pages 665–677, 2017. tom williams, cynthia matuszek, ross mead, and nick depalma. scarecrows in oz: the use of large language models in hri. j. hum.-robot interact., 13(1):1–11, january 2024. ramin yaghoubzadeh and stefan kopp. flexdiam – flexible dialogue management for problemaware, incremental spoken interaction for all user groups (demo paper). proceedings of the 7th workshop on speech and language processing for assistive technologies (slpat 2016), 2016. ramin yaghoubzadeh, karola pitsch, and stefan kopp. adaptive grounding and dialogue management for autonomous conversational assistants for elderly users. in intelligent virtual agents, lecture notes in computer science, pages 28–38. springer international publishing, cham, 2015. takashi yamauchi, mikio nakano, and kotaro funakoshi. a robotic agent in a virtual environment that performs situated incremental understanding of navigational utterances. in sigdial 2013, pages 369–371, metz, france, august 2013. association for computational linguistics. shiquan yang, rui zhang, and sarah erfani. graphdialog: integrating graph knowledge into endto-end task-oriented dialogue systems. in proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), pages 1878–1888, 2020. steve young, milica gašić, blaise thomson, and jason d williams. pomdp-based statistical spoken dialog systems: a review. proc. ieee, 101(5):1160–1179, 2013. tom young, frank xing, vlad pandelea, jinjie ni, and erik cambria. fusing task-oriented and open-domain dialogues in conversational agents. in proceedings of the aaai conference on artificial intelligence, volume 36, pages 11622–11629, 2022. sina zarrieß and david schlangen. easy things first: installments improve referring expression generation for objects in photographs. in proceedings of the 54th annual meeting of the association for computational linguistics (acl 2016), 2016. hao zhang, weiwei li, rilin chen, vinay kothapally, meng yu, and dong yu. llm-enhanced dialogue management for full-duplex spoken dialogue systems. arxiv [cs.cl], february 2025. 129 kennington, lison and schlangen yichi zhang, zhijian ou, min hu, and junlan feng. a probabilistic end-to-end task-oriented dialog model with latent belief states towards semi-supervised learning. in proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), pages 9207–9219, 2020a. zheng zhang, ryuichi takanobu, qi zhu, minlie huang, and xiaoyan zhu. recent advances and challenges in task-oriented dialog systems. science china technological sciences, 63(10): 2011–2027, 2020b. siyuan zhou, yilun du, shun zhang, mengdi xu, yikang shen, wei xiao, dit-yan yeung, and chuang gan. adaptive online replanning with diffusion models. in thirty-seventh conference on neural information processing systems, 2023. lukáš žilka and filip jurčı́ček. lectrack: incremental dialog state tracking with long short-term memory networks. in text, speech, and dialogue, pages 174–182. springer international publishing, 2015. 130 introduction background: incremental spoken dialogue systems spoken dialogue systems: overview frameworks & architectures the incremental unit framework restart vs. update incremental models common modules in incremental, interactive systems automatic speech recognition natural language understanding natural language generation and speech synthesis incremental systems & evaluation review of incremental dialogue management a brief overview of dialogue management dialogue state tracking action/response selection turn-taking and end-of-turn prediction incrementalizing dialogue management discussion desiderata conclusion dialogue & discourse 16(3) (2025) 60–95 doi: 10.5210/dad.2025.304 a graph-to-text approach to knowledge-grounded response generation in human–robot interaction nicholas thomas walker nicholas.walker@uni-bamberg.de norwegian computing center stefan ultes stefan.ultes@uni-bamberg.de university of bamberg pierre lison plison@nr.no norwegian computing center editor: hendrik buschmeier submitted 11/2023; accepted 11/2025; published online 12/2025 abstract knowledge graphs are often used to represent structured information in a flexible and efficient manner, but their use in situated dialogue remains under-explored. this paper presents a novel conversational model for human–robot interaction that rests upon a graph-based representation of the dialogue state. the knowledge graph representing the dialogue state is continuously updated with new observations from the robot sensors, including linguistic, situated and multimodal inputs, and is further enriched by other modules, in particular for spatial understanding. the neural conversational model employed to respond to user utterances relies on a simple but effective graph-to-text mechanism that traverses the dialogue state graph and converts the traversals into a natural language form. this conversion of the state graph into text is performed using a set of parameterized functions, and the values for those parameters are optimized based on a small set of wizard-of-oz interactions. after this conversion, the text representation of the dialogue state graph is included as part of the prompt of a large language model used to decode the agent response. the proposed approach is empirically evaluated through a user study with a humanoid robot that acts as conversation partner to evaluate the impact of the graph-to-text mechanism on the response generation. after moving a robot along a tour of an indoor environment, participants interacted with the robot using spoken dialogue and evaluated how well the robot was able to answer questions about what the robot observed during the tour. user scores suggest an improvement in the perceived factuality of the robot responses when the graph-to-text approach is employed compared to a baseline using inputs structured as semantic triples. keywords: dialogue management, human-robot interaction, large language models, graphs, task-oriented dialogue 1. introduction recent advances in nlp have made large strides in improving the conversational abilities of dialogue systems (bommasani et al., 2021). large language models (llms) based on the transformer architecture (vaswani et al., 2017) have become the foundation for a large range of conversational agents, both task-oriented (hosseini-asl et al., 2020; peng et al., 2020; nekvinda and dušek, 2022) and open-domain (zhang et al., 2020; roller et al., 2021; thoppilan et al., 2022). these advances have also raised questions about the capability of these models to perform tasks and reason about ©2025 nicholas thomas walker, stefan ultes, and pierre lison this is an open-access article distributed under the terms of a creative commons attribution license (http ://creativecommons.org/licenses/by/3.0/). a graph-to-text approach to knowledge-grounded response generation in hri information (valmeekam et al., 2022). numerous investigations into the reasoning capabilities of llms have resulted in evidence of their abilities in some areas (bubeck et al., 2023), while indicating shortcomings in others (shi et al., 2023; ullman, 2023). meanwhile, there is a continuing interest and ongoing work in human-robot interaction (hri) using large language models. a central question in this vein of research is how information about a robot’s environment from different modalities can be effectively combined with llms. for instance, huang et al. (2022) demonstrated a system where feedback from the environment is used in the form of a kind of inner monologue. under such a system, multiple sources of information are combined in a form representing the semantics of a situation, in the sense of how objects under discussion relate to each other. these systems thereby make use of multiple modalities obtained from other (particularly pre-trained) models where the combined information enables improved performance even in zero-shot scenarios (zeng et al., 2022). such models suggest that there exists a strong capability for llms to operate over diverse sources of information structured as text input. in this article, we explore the problem of response generation for embodied agents and demonstrate how various sources of information can be collected and acted upon in such a system. our approach rests on the following elements: 1. a dynamic knowledge graph of the dialogue state where locations, physical objects, and dialogue elements such as utterances are represented as nodes and semantic relations between them represented with edges. 2. a simple parameterized function to verbalize the dialogue state graph based on traversals of the graphs, and the optimization of those verbalization parameters from wizard-of-oz data. 3. a large language model that includes the verbalized representation of the graph as part of its prompt and is responsible for generating the robot response. a sketch of the proposed approach is provided in figure 1. as shown in the figure, information detected by the embodied agent including speech, visual, and positional data are brought together into a unified structure in the form of a graph and subsequently transformed into natural language as input for an llm. specifically, this is a dynamic knowledge graph, in which nodes and edges are added incrementally. a dynamic rather than static graph structure allows it to be modified throughout robot exploration and conversation with users. the graph structure thereby enables the continuous inclusion of both situated and dialogue-specific information as it is received by the system. semantic relations between objects, places, and utterances are represented as edges between the nodes of the graph, enabling not only verbal expression of these relations but also of larger traversals of the graph. moreover, probabilities and temporal relations in the graph are also easily exploited in such a transformation. the graph structure is thereby exposed to an llm in a way that enables the strengths of such models to be exploited through a simple, controllable transformation from graph to text. to evaluate the proposed approach, we design an experiment with users to explore its use in a practical setting. our experiment consists of users touring an office space with a robot and subsequently discussing the tour. we evaluate the robot’s ability to converse about the tour when equipped with two llms (specifically gpt-4 and llama), using the graph-to-text approach and a baseline triples conversion. from the resulting dialogues and user feedback, we provide a statistical analysis of the results along with further qualitative analysis of attributes of the resulting dialogues. 61 walker, ultes and lison asr f(g,θ) verbalization prompt “i saw a plant in the hall.” response object recognition (vqa)… input & postprocessing modules figure 1: diagram of the proposed approach. the red path in the graph indicates a traversal which is transformed to a natural language document via a parameterized function. the resulting text description is then inserted into the prompt for the language model. the rest of the paper is structured as follows. in section 2, we provide a general background on embodied conversational agents, knowledge graphs for dialogue, and large language models. section 3 describes how to structure the dialogue state of the agent as a knowledge graph, verbalize its content through graph traversals, and integrate the result into the prompt of the llm employed to generate the robot responses to user utterances. section 4 presents an experimental study that investigates the influence of this graph-to-text approach on the factuality and adequacy of the robot responses. the obtained user ratings show that this graph-to-text approach is able to improve the factuality of the robot responses compared to a baseline relying on inputs where the graph information is provided in the form of semantic triples. finally, in sections 5 and 6 we discuss the experimental results and review the benefits and limitations of the proposed approach, and conclude. 2. background in this section, we discuss the background to our study, starting with contemporary approaches to embodied agents in hri. we then discuss recent work where graphs have been used with dialogue systems to represent situated dialogue information. finally, we review the recent usage of llms in dialogue systems and the variety of approaches to presenting background information to them, including with graphs and for embodied agents. 2.1 embodied conversational agents embodiment refers broadly to the operation of an artificial agent in physical space, though specific definitions vary in precision and strictness (miller and feil-seifer, 2016). whichever definition may be preferred, embodiment has a substantial impact on dialogue and affects how speakers perceive 62 a graph-to-text approach to knowledge-grounded response generation in hri and relate to the dialogue (wainer et al., 2006; pickering and garrod, 2009). in human–robot interaction, grounding the dialogue in the physical environment is often a prerequisite to language understanding. grounded language, in the sense employed by e.g. tellex et al. (2020), refers to situated language, i.e. language with meaning in a physical environment. in a situated scenario, dialogue between a human and a dialogue agent has a direct spatial relation to the surrounding environment. an embodied, robotic agent must take into account numerous input sources from the environment as well as any human users to accomplish complex tasks requiring both natural language or physical outputs, i.e. movement. beyond spatial data and language input, visual perceptions of varying kinds are an important source of information for the dialogue. these kinds of information may also be viewed in relation to time. for instance, the utterances in a dialogue are naturally temporally related. moreover, embodied agents may take different actions depending on what information is most recently observed. linguistic, visual, and spatial information acquired by the robot may also be uncertain, presenting an additional challenge in effectively leveraging world knowledge to accomplish tasks. the agent’s grounding in the physical world is subject to uncertainty arising from both perceptual module errors as well as inherently uncertain aspects of the environment. thus, temporality and probability interact with these kinds of data to present a nuanced picture of the world for the agent to act upon. when quantified and represented by the system, uncertainty about the environment or dialogue state can be communicated to the user (hough and schlangen, 2017). by dialogue state, we refer to a representation of the current dialogue context which contains all information that is relevant for responding to user utterances. within hri, language input is one communication medium among many that can be used to interact with a robot (goodrich et al., 2008). in some cases, it may be used to instruct a robot to perform actions in its physical environment. for example, rankin et al. (2021) demonstrated semantic parsing of instructions to associate the instructions with desired locations for the robot to explore. in a smaller setting, lynch et al. (2023) explored the use of imitation learning through natural language instructions to guide a robotic arm to move objects. the real-time task required an approach that could account for both spatial information and the temporal ordering of actions and instructions given to the robot. another model integrating spatio-temporal grounding was presented by shah et al. (2023). in their model, a robotic navigation system combines a pre-trained visual navigation module with an llm to construct a path following instructions. this approach first identifies landmarks with an llm and subsequently grounds this information in a set of visual inputs from the navigation area using a second module. each of these approaches required alignment of language input to the robot’s physical environment. 2.2 knowledge graphs for dialogue 2.2.1 graph-based representations while dialogue state information may be represented in numerous ways, graphs are a flexible data structure with desirable properties with regard to both natural language and the physical world. a graph g consists of a set of vertices (or nodes) v paired with a set of edges e. each edge describes a pair of vertices, which may be ordered or unordered (describing a directed or undirected graph, respectively). knowledge graphs specifically represent entities, that is to say objects and concepts, as well as relations between them. an example of graph-structured background information for dialogue is the dataset of eric et al. (2017). in this dataset, each dialogue is associated with a knowledge base of entities for the dialogue system to retrieve in response to user queries. the 63 walker, ultes and lison dataset comprises three domains, of which the navigation domain represents entities with a spatial relation (distance) to the speaker and dialogue agent. with this property, the dialogue system can reason over the dialogue context to infer closest locations to recommend to the user. another graphbased model of spatial information for embodied agents is the scene graph (xu et al., 2020), which represents physical objects as graph nodes and spatial relations as edges between them. similarly, in the model presented by li et al. (2022b), an embodied scene graph can also be generated by a model trained with imitation and reinforcement learning. the resulting graphs can then be applied to downstream tasks requiring spatial information. within language data specifically, both utterance and dialogue level structures can be expressively rendered as graphs such as those of the abstract meaning representation paradigm (banarescu et al., 2013). a relevant use of this paradigm was presented by bonial et al. (2020), extending the amr schema for dialogue with a robot performing tasks in search and navigation. their model of situated dialogue constructs graph representations of both user commands and agent responses. these graphs represent a semantic model of both objects in the environment and actions for the system to take. graphs have also been used to represent dialogue structure as a whole and flows between utterances (andreas et al., 2020; gritta et al., 2021). also specifically in hri, wilcock and jokinen (2022) presented a model for robotic conversational agents that used knowledge graphs to present semantic metadata to the system. thus, in the context of previous work it is apparent that combining abstract and dialogue-specific information with situated knowledge provides a fruitful avenue for state representation (papangelis and ultes, 2020). within this line of inquiry, our work builds upon a graph-structured dialogue state model wherein both dialogue elements and physical objects in the world are equally represented as nodes of a graph with edges in the graph representing semantic relations between them (walker et al., 2022). aside from the general usage of graphs representing linguistic structure or purely physical information, an additional kind of dialogue information that can be expressed in a graph is temporal relations between objects. for instance, xing and tsang (2023) introduced the darer model, which used two types of temporal graphs representing semantic relations for dialogue act and sentiment prediction. in their model, user utterances are represented as nodes of a graph with “previous” and “future” relations between them. as demonstrated in a number of models (tuan et al., 2019; he et al., 2023), the dialogue agent’s knowledge state can be captured as a dynamic knowledge graph that is updated throughout the dialogue. dynamic, evolving graphs capture the intuition that the dialogue agent’s knowledge should evolve along with the dialogue, encompassing perception of natural language utterances or physical observations as updates to the graph. a further advantageous property of graphs is the capability of associating probabilities with the constituent nodes or edges of a graph. a common use of probability with graph structures in dialogue systems has been in handling uncertainty in action selection (young et al., 2013). bayesian networks are an additional example of a probabilistic model which has been used for dialogue management, including with multimodal data (prodanov and drygajlo, 2003; thomson et al., 2008). more generally, the problem of epistemic uncertainty (celemin and kober, 2023) about the dialogue state may also be usefully described by probabilities associated with graph elements. 2.2.2 graph processing although the flexibility of graphs allows for a wide diversity of state representations, it also requires adaptation of response generation models to operate on a graph structure. one means of doing so 64 a graph-to-text approach to knowledge-grounded response generation in hri is the use of graph neural networks (gnns). an example of a gnn model used to process a large knowledge base was presented by christmann et al. (2023). their model for conversational question answering used a gnn to construct a reduced representation of a large heterogeneous graph of entities. this model had the particular advantage of enabling explainability in the model along with using numerous combined sources of information. particularly large knowledge graphs may require a filtering method or preprocessing to extract the most salient information (wang et al., 2021). gnns have also been used among other things to verbalize graphs, such as the method presented by yang et al. (2020), wherein a graph attention network is used with a verbalization of a program to verify facts. while gnns have become an area of interest for using graph data in many domains, these models often require large quantities of in-domain training data and may require certain properties in the graph structure (e.g. homogeneity, connectivity). for instance, models such as the influential graphsage of hamilton et al. (2017) have limitations in their ability to distinguish certain non-isomorphic graphs (garg et al., 2020). moreover, gnns also face difficulties when aggregating information across long paths in graphs, an “over-squashing” of graph information explored by alon and yahav (2020). graph verbalization is not limited to gnn-based models. xu et al. (2018) demonstrated a method of transforming structured query language (sql), itself viewable as a graph-structure, to natural language via a graph-to-sequence model. a common thread among approaches to verbalizing graphs is a correspondence of graph structures to natural language surface forms. a basic example is the conversion of a semantic triple (subject, predicate, and object) of two nodes and an edge between them to a sentence: (‘robot’ | ‘in’ | ‘hallway’) → the robot is located in the hallway. such conversion is not limited to small-scale elements of the graph. the entire graph might be labelled or summarized by some natural language description, as in graph classification (zhang et al., 2022). a knowledge graph completion algorithm relying on a transformation of paths in the graph was presented by lin et al. (2023). in their approach, knowledge graph paths are translated to text by concatenating entities and relations separated by special tokens. another approach using traversals knowledge graph traversals is opendialkg (moon et al., 2019), which relates shifts in dialogue context to walks in the graph. other work using this dataset and the dataset of eric et al. (2021) is the diffkg model of tuan et al. (2021). the diffkg model uses a transformer architecture to perform walks over the knowledge graph to select kg knowledge for a dialogue system. this system was evaluated with respect to ground truth paths and entities in fixed knowledge bases which are not available in all dialogue domains. the authors explored a baseline following the model of beygi et al. (2022), which transformed a knowledge graph into a text document of semantic “pseudo-language” statements. in both this model and diffkg, a structured text representation of the graph is a source of information to the dialogue system similar to the semantic triple-based model we employ in our study (see section 4). the attnio model of jung et al. (2020) also makes use of traversals with an attention mechanism to attend to varying sizes of node neighborhoods in the graph. 65 walker, ultes and lison 2.3 large language models (llms) 2.3.1 response generation with llms llms have demonstrated strong capabilities for few-shot learning across tasks (brown et al., 2020), although limitations remain in their use in task-oriented dialogue (hudeček and dušek, 2023). at the outset, these models have a great deal of information available from the large quantities of data consumed in training them (bommasani et al., 2021). despite substantial improvements in llm performance across tasks, limitations remain. dziri et al. (2023) investigated llm performance on reasoning tasks and found that the models often resort to basic pattern recognition rather than generic problem solving. moreover, hallucinations remain a substantial limitation and subject of continuing research (ji et al., 2023). as zhang et al. (2023) describe it, a hallucination exists when an llm generates output that “misaligns with established world knowledge.” for our purposes, we can consider established world knowledge to be defined by the physical environment in which the user and robot operate. thus, the notion of factuality (or, absence of hallucinations) derives from what is true in the physical environment, thus for our study we seek to evaluate factuality of model responses with respect to the environment rather than the graph itself. recent study has also indicated that the linguistic structures within prompts may be a factor in the occurrence of hallucinations, specifically that more formal and concrete prompts may assist in reducing hallucination (rawte et al., 2023). in this context, evaluation of factuality in llm responses remains an important direction for study. 2.3.2 prompting methods in many recent dialogue systems, the model is initialized with a system prompt which consists of an initial input that informs the system what role it should play in the dialogue and how to proceed (liu et al., 2023). besides defining the system’s role in the dialogue, the prompt may include other information that is salient to the system’s tasks. importantly, the input to llms is bound by a context window beyond which input text must be truncated, although some research has investigated means of alleviating this limitation (ratner et al., 2023). when the total scenariospecific information exceeds the limit imposed by the context window, it becomes advantageous to design methods to select relevant information (thulke et al., 2021). complementary to these methods is the possibility of providing further enriched context such as information derived by commonsense or logical reasoning (walker et al., 2023). due to the difficulty of training large lms, the use of in-context learning has also become a heavily investigated method of providing training examples to the largest language models (dong et al., 2022). with in-context learning, a small number of examples can be used to prime a model to generate output of a desired format or style. notably, this technique avoids the need to update model parameters which would otherwise be costly or infeasible given the size of many current llms, instead relying upon the model’s pattern recognition abilities to adapt its output to a desired outcome. combined with the general notion of prompting, this approach enables substantial flexibility when using llms for different tasks and scenarios, including with multimodal tasks such as generalizing models to unseen contexts (tai et al., 2023). in-context learning with llms has also been demonstrated with table-like structures (chen, 2022). 66 a graph-to-text approach to knowledge-grounded response generation in hri 2.3.3 llms in embodied systems as noted previously, llms have been used as a means of processing language input for embodied systems such as commands or instructions from a user. an example of a model that integrated an llm with an embodied agent is the progprompt model presented by gupta and kembhavi (2023). progprompt generated task plans for a robot using an llm by making use of a python-like prompt format which induced the llm to generate actions structured as executable programs. thus, the llm does not operate on strictly natural language, but rather a kind of programming language. a major advantage in the progprompt system was that the set of actions available to the robot could be represented compactly and abstractly as functions in this language structure. a robot equipped with an llm was also presented by billing et al. (2023), who demonstrated a basic model of hri with a pepper robot using the gpt-3 model of openai. in this case, although the dialogue system interacted with the user through an physical agent, the dialogues themselves were not situated in the physical environment. beyond the use of text input (natural language or otherwise), other research has investigated directly incorporating multimodal input to llm architectures. an example of llm capable of generating text based on combined text and visual input is the flamingo model of alayrac et al. (2022), who presented an approach to combine textual and image data with few-shot prompting for visual dialogue. similarly, the palm-e model of driess et al. (2023) addressed multimodality by directly incorporating embeddings of image and other non-language data in the input to an llm architecture. such approaches suggest that models trained or prompted to incorporate mixed forms of information can be effective, however these architectures are often task-specific. moreover, as noted before, it is often cost-prohibitive to train large models. when the system must take into account its physical environment, the challenge of presenting it appropriate information is potentially greater than a non-embodied dialogue system. as direct operation on audio-visual data or haptic feedback is impossible for an llm, a transformation from the raw forms of these inputs into natural language is required. ahn et al. (2022) presented a model called saycan that combined an llm with an “affordance function” describing the probability of success over a set of possible actions for the agent. another example of such a system was presented by zhao et al. (2023), where a robot “chats” with the environment in the sense that multimodal sensory input is iteratively presented to an llm as natural language statements describing the input, with the model outputting actions at each turn. this approach takes advantage of the turn-taking nature of dialogue to inform a system in real-time, but would be difficult to port to an embodied agent intended to explore a wider environment. nonetheless, this work demonstrates the capabilities of an llm in reasoning over diverse sources of input information represented as natural language. a further example of llms integrated in hri systems is the llm-planner of song et al. (2022), a few-shot system using prompting with a gpt-3 model as a planner for an embodied agent. an approach for data augmentation by grounding dialogue responses in text documents was presented by wu et al. (2020). this model used two dialogue agents to simulate conversations across different domains represented by sets of documents. lastly, a question answering model which stands quite close to our approach was presented by lanchantin et al. (2023), whose model made use of a textformatted version of a knowledge graph along with a transformer operating directly on the graph. this model operated in a simulated three-dimensional space, where the queries directed to the agent related to the spatio-temporal information describing the user’s position in the simulated space. this work did not however investigate the use of a natural language transformation of the graph, 67 walker, ultes and lison relying instead on a text document constructed of structured representations of graph information in a similar manner to the semantic triples model employed in our experiments. additionally, this approach relied upon the generation of simulated data in a virtual space, whereas we investigate the real-time creation of graph in a physically situated hri scenario. 3. approach in this section, we describe our proposed approach to llm-based response generation for hri. our proposed model revolves around a dynamic knowledge graph that represents both elements of the environment and dialogue utterances. nodes and edges describing these elements are iteratively added as new data is observed by an agent. 3.1 state representation as explored in section 2, embodied dialogue agents must take into account diverse types of information regarding the dialogue as well as the physical environment. these aspects of embodied dialogue suggest that the flexibility, expressiveness, and semantic interpretability of graphs are effective at representing evolving spatial, temporal, and probabilistic information in a single structure. for these reasons, we use a dynamic knowledge graph to represent the dialogue state. a dynamic, in-memory knowledge graph permits continuous updates as new information is recorded. the nodes of this graph represent entities, in the sense of distinct objects and places in the environments, as well as natural language utterances. likewise, the edges represent semantic relations between these entities. these relations include spatial or temporal information which may be updated in the course of exploring an environment. temporality is also applicable to utterances within the dialogue as a chain of responses. as both dialogue information and situated information from the environment are represented within the graph, a form of “grounding” (in the sense discussed in section 2) of the dialogue in the environment is equally possible through edges between utterances and physical objects. in this model, nodes in the graph are assigned a type label along with a content. the type of the node indicates what kind of conversational entity the node represents, such as an utterance or a physical object. edges in the graph represent semantic relations between the nodes of objects, such as physical location, origination, or in the case of utterances a temporal ordering. to illustrate situated information in a graph, the node representing a laptop in figure 2 has a description of “laptop” along with the type label “entity”. the laptop is likewise related with an “in” relation to an image data node carrying the type label “image”, where binary data of the image itself is stored. in turn, the image node relates to a “location” node named “office”, indicating that the picture was taken in the location labelled “office”. plainly stated, this graph structure describes the fact that a laptop computer was seen in an image taken in an office. nodes and edges are also assigned a probability. the assignment of a probability value to a node allows for the expression of uncertainty arising from sensory data. this is notably the case for visual elements detected in the images captured by the agent, or for alternative transcriptions produced by the asr model. in general, regardless of the source or type of information represented in a node, uncertainty in the data can be expressed as an attribute of the node. likewise, edges can also be assigned to a probability. a further property of nodes is a timestamp. this attribute makes it possible to account for temporality in the data. every node in our model is assigned a timestamp representing its creation 68 a graph-to-text approach to knowledge-grounded response generation in hri figure 2: example of subgraph where an image is assigned a location, and a “laptop” entity detected in the image data is created as a node with an “in” relation to the image. time in the graph. a second timestamp representing the creation time of the information is also assigned to the node metadata. this timestamp is the time the information was recorded by an external module such as robot sensors. the difference between the two timestamps is often small, nevertheless separate representations allow precise measurement of any latency or reference to exact occurrences in physical space as opposed to graph update time. 3.2 state tracking & postprocessing 3.2.1 graph updates dialogue state tracking (williams et al., 2016; ren et al., 2018) is often viewed in the form of a “slot-filling” paradigm, wherein slots of a defined ontology are filled by a predictive model. here, we conceive of state-tracking as the management of the conversational knowledge graph. as the knowledge graph is comparatively small, it is stored in-memory and directly updated through python operations. because the knowledge graph is dynamic, new nodes representing spatial positions, images, and entities seen during the embodied agent’s movement are added to the state. this is a constant stream of information that updates the dialogue state as soon as new information is obtained. thus, spatial coordinates and visual information can be continuously added as nodes to the graph while an embodied agent traverses physical space. equally, language input is added to the graph as it is received, whether concurrently with agent movement or not. this flexibility is an advantage for an embodied agent to allow it to attend to numerous, diverse sources of information simultaneously while also taking into account interactions between them. we rely on a dedicated management module that serves as a central point of access for all updates to the graph. upon each update, the management module triggers all system modules that should be notified of this change. those modules may in turn lead to new graph updates, thereby creating a cascading transformation to the graph, where modules variously react and push changes to the graph manager. an important example is a response generation module, which is triggered upon the addition of a new user utterance node in the graph. using the updated graph, the response 69 walker, ultes and lison robot path end point start point start point end point forward movement rotation figure 3: example path drawn from spatial coordinates, with raw coordinates on the left and the approximated path on the right. the orange points mark coordinates associated with the room label, blue points are those labeled as located in the hallway. approximate forward movements and inplace rotations are highlighted alongside the robot path in red and green, respectively. generation module generates an agent response to the last user utterance and adds the response to the graph with a “follows” relation to the utterance. this graph update in turn is passed to the embodied agent for speech generation. in the same fashion, objects detected in visual sensory input are added to the graph. raw image data collected by a camera can be added as nodes to the graph directly. these additions in turn trigger postprocessing of the image data, whereby objects, people and other entities may be recognized and added to the graph as additional nodes with an “in” relation to the image. we also assign probabilities to these objects derived from the output of the visual model use to identify them. 3.2.2 path generation a particular component integrated in the hri platform employed in our experiments (section 4) is a mechanism for deriving a continuous path from the raw sequence of spatial coordinates recorded by the robot. this mechanism operates in two steps. the spatial coordinates are first converted into a limited number of lines based on the ramer-douglas-peucker algorithm (douglas and peucker, 1973). those lines are in turn converted into an ordered list of movements, divided in two types: straight forward movements and in-place rotations1, the straight movements being specified for a given distance, and the rotations for a given angle. this list of movements is then integrated into the graph. an example path drawn from raw coordinates is given in figure 3. the nodes expressing the spatial coordinates in the state graph are further connected to edges expressing that they are part of a larger location with a given name, such as “hallway” or “break room”. in the experiments, those names are provided by the human participant during the tour, as detailed in section 4.1. 1. this is of course a simplification, based on the fact that the pepper robot uses omni-directional wheels, and that its movements in the indoor environment of the user study were either straight-line motions or in-place turns. 70 a graph-to-text approach to knowledge-grounded response generation in hri turn left 75 degrees creation_time 08:55 bookshelf in office of magna table creation_time 08:39 bookshelf creation_time 08:48 computer mouse creation_time 08:38 door creation_time 08:48 forward 2.8 meters start corridor 3 forward 1.2 meters start corridor for turn left 67 degrees start corridor for person in corridor for turn left 70 degrees start corridor desk creation_time 08:45 chair creation_time 08:48 forward 2.4 meters start storage room forward 9.0 meters start new corridor forward 1.8 meters start survive [...] at 08:38, we were in starting area. in starting area, pepper saw a table, a door, a desk, a chair, a computer mouse, a bookshelf and a tv. pepper may have also seen a computer, a blue curtain, a lamp, a light, paper, a window and a copy machine. we spent 27 seconds in that area and traveled 5.7 meters. in starting area, pepper turned left 24 degrees. at 08:39, we were in corridor. in corridor, pepper saw a table, a fire extinguisher, a door, a desk, a chair, a couch, a bookshelf, a whiteboard and a window. pepper may have also seen a wastebin, a plant, a printer, a pillar, a pen, a blue curtain, a lamp, a light, a tv and a copy machine. we spent 84 seconds in that area and traveled 25.9 meters. [...] triples representation verbalized representation figure 4: comparison of a dialogue graph represented with triples versus verbalization. 3.3 graph-to-text transformation 3.3.1 verbalization function to use the dialogue state graph information with an llm, it is necessary to convert the structured information into natural language. a variety of methods have been applied to generate natural language from structured information such as tables (nema et al., 2018). in contrast to e.g. llmbased approaches to converting structured data to text, our proposed approach the converts dialogue state graph information into text descriptions via a set of templates relying on graph traversals. in this approach, both relations between nodes as well as more complex paths can be rendered with a natural language description. this traversal-based transformation is a simple means to convert structured information to a form amenable to an llm. to traverse the graph, we first order the nodes temporally and iterate over them. for each node, the node’s type determines which paths should be searched for in its neighborhood. the traversal function then collects ingoing and outgoing edges along with the neighboring nodes connected to them. for each node type, the function defines a set of paths to traverse by searching along the ingoing or outgoing edges. as an example, when evaluating an “entity” node, paths containing an “in” relation from the source node to an image node with a subsequent “in location” edge to a “location” node describes the location of an object (as in figure 2). thus, the paths are defined in terms of which edges and node types should be found in a given sequence of hops in the graph. the collection of these paths then constitutes a chronologically ordered description of the graph. node probabilities are also recorded during the traversal. the verbalization process takes the form of a function that converts the graph into a string according to a set of handcrafted parameters. an example parameter is that the system may reference itself within the prompt, such as first or third person pronouns, or a name. depending on this parameter, the description of the graph will then summarize the information from either a personal or 71 walker, ultes and lison parameter values how the agent refers to itself “pepper”, “the robot”, “i” whether to include discourse markers between sentences true — false how to specify distance (in meters) “precise”, “rounding”, “none” whether to specify rotation angle (in degrees) true — false whether to include low-probability entities in description true — false whether to include time in description true — false whether to mention the total number of turns true — false table 1: parameter values for the verbalization function. each parameter specifies an aspect of the verbalization. bold values indicate the parameter values selected by optimization. impersonal perspective with respect to the dialogue agent. as illustrated in table 1, the verbalization function employed in our experiments also included parameters related to the use of discourse markers, the inclusion of low-probability entities in the description, or how to express distances, angles, and time. an abbreviated document from our study data is shown in figure 4, along with the triple-structured document for comparison. 3.3.2 parameter selection as mentioned above, the transformation of the graph to a text document is a parameterized function. to select the best set of parameters for the task, we optimize the parameters with respect to the output of a llama language model (touvron et al., 2023), such that the model assigns the highest probability to the correct response out of a set of candidate responses across dialogues. for this purpose, we collect a small set of wizard-of-oz interactions, each response from the wizard being associated with a corresponding dialogue state graphs (which include sample locations and entities along with the dialogue history). formally, let vp(g) denote a graph-to-text function where g is a knowledge graph and p a set of categorical parameters, and let d = {(gi, ri), 1 ≤ i ≤ n} denote the recorded set of n wizard-of-oz responses ri in a conversational context represented by its dialogue state graph gi. we then search for an optimal set of parameter values p∗ that minimizes the cross-entropy loss over the wizard-of-oz examples, assuming a fixed language model llm: p∗ = argmin p lce(d ;llm,p) (1) where the cross-entropy loss is itself defined as: lce(d ;llm,p) = − 1 n ∑ (gi,ri)∈d logpllm(ri | vp(gi)) (2) the optimization is performed using a tree-structured parzen estimator algorithm (bergstra et al., 2011), as implemented in the optuna package for efficient hyper-parameter optimization (akiba et al., 2019). using the above approach, we are able to optimize the categorical parameters and adjust the verbalization function with only a few example responses from a wizard. 72 a graph-to-text approach to knowledge-grounded response generation in hri 3.4 response generation the final component of the proposed approach is the actual response generation. as already mentioned, this response is generated using a large language model optimized for dialogue use cases. in our experiments, those language models were respectively based on gpt-4 (openai, 2023) and the llama-2 chat model (touvron et al., 2023). the response is generated based on a prompt that includes the verbalized graph along with: • a short, generic instruction text that details how the robot should respond to the user utterances, and stress in particular that the robot responses should all be grounded in the observations expressed in the verbalized graph. • a short list of example responses from the wizard-of-oz interactions to take advantage of incontext learning (brown et al., 2020; wei et al., 2022) and provide cues about how responses are expected to be formulated. once decoding is complete, the response is itself added to the graph as a new node. this node is connected to the user utterance with a responds to relation, thereby capturing the temporal sequence of dialogue turns in the graph. 4. experiments in this section, we describe the design of our experimental study and the results of the corresponding user evaluation. the goal of our study is to explore the capabilities of a robot that relies on a dynamic knowledge graph as the representation of its environment and dialogue state, and converts this graph into natural language to enable the generation of system responses through an llm. to do so, we designed a study wherein users brought the robot to different locations in an office environment. after the completion of this office tour, the participants were then instructed to conduct a short dialogue with the robot to assess its ability to recollect the observations made during the tour. three distinct model setups were investigated, and users conversed with all three in a randomized order immediately after their tour with the robot. after each conversation, the participants were asked to rate the robot responses with respect to their factuality and their adequacy. we recruited a total of 20 participants (for the most part it students and researchers) for our study, the first two participants being asked to test a prototype version of the final system. each participant was given a 200 nok gift card for their participation. for the robot in the experiment, we use the pepper robot of aldebaran2. the dialogues were conducted in english. 4.1 experimental setup the experiment was structured in two parts that respectively correspond to the physical tour of the office and the dialogue between the user and the robot that followed. tour phase during the tour, the robot is moved around an office environment by a participant to record details of its path, locations, and identifiable objects. for the sake of simplicity, we instruct participants to push the robot along its path in the office. before entering a new room or corridor, the participant was instructed to activate the robot microphone by tapping its head and stating the name of the area they were entering. the time spent on each tour ranged from around 5 to 20 minutes. 2. https://www.aldebaran.com/en/pepper 73 https://www.aldebaran.com/en/pepper walker, ultes and lison figure 5: the first author in a dialogue with the robot. a video of a short example tour may be found at https://youtu.be/a52zbcfvgs8 dialogue phase after the tour was completed, the user was then asked to converse with the robot and have it summarize the tour, as illustrated in figure 5. the participant was instructed to ask the robot at least 5 questions relating to the tour, asking it to summarize the tour along with questions such as to describe which locations it visited, what was seen in each location, what distance was traveled, or how much time was spent in an area. participants were encouraged to treat the conversation as a natural, continuous dialogue rather than a set of isolated questions, allowing for follow up questions referencing the previous question (e.g. “what did you see after that?”). this dialogue between robot and user is repeated three times, each time with a distinct configuration for response generation (see below). participants were instructed to treat each dialogue as a completely independent conversation. the model order was randomly reshuffled for each experiment. the participants were also told they did not need to ask the same questions in each dialogue. 4.2 system configuration the software architecture employed for the experiments integrates a number of modules, as illustrated in figure 6. the sensor-related modules ran on the pepper robot itself, while the other modules were located on a separate machine with access to a gpu. 4.2.1 robot perception during the tour, the dialogue state graph is continuously updated with new visual and positional data, along with the timestamp at which these elements are observed by the robot. the user-provided location labels are added to the graph as nodes with the provided name as the node description. as the robot moves along its path, new nodes are added to the state representing spatial positions, images, and entities seen during its movement. the room labels provided by the user are assigned to all subsequent positions recorded by the robot until the next label was provided. the robot takes a picture of what is in front of it approximately every 1.5 seconds, along with its current coordinates relative to an initial starting location. due to odometry drift, we were unable to treat the spatial data as a precise representation of the robot’s path at large distances. nevertheless, these location measurements provided an approximate measure of distance traveled within areas as 74 https://youtu.be/a52zbcfvgs8 a graph-to-text approach to knowledge-grounded response generation in hri sensors camera asr other input object and person detection verbalization response generation (llm) path detection robot figure 6: diagram of the implemented architecture for the experiment. robot sensor modules are colored blue, while post-processing models operating on the graph are colored yellow. well as a means of estimating directional changes by the robot. the image data and x, y coordinates are added to the graph as nodes labeled image and position, respectively, and are each created with an edge to the current location node. as mentioned in section 3.2, a dedicated mechanism is implemented to convert the raw spatial coordinates to a high-level sequence of movements based on the ramer-douglas-peucker algorithm (douglas and peucker, 1973). when an image is added to the graph, a visual question answering (vqa) model processes the image information and returns a score of whether objects were detected in the image. the model we chose for this purpose is the vision transformer model of minderer et al. (2022). we enumerate the possible objects the robot should see including people as well as office objects (see appendix a). when objects are added to the graph with this model, the estimated probability of the object in the image from the vqa model is assigned as the probability of the node. as this is based upon a non-fine-tuned model, there is substantial noise in the observations of the objects. 4.2.2 dialogue system during the dialogue phase, the user activates the microphone by tapping the robot’s head. utterances are transcribed using the google speech api3 and insert to the knowledge graph as nodes. the process of response generation is triggered upon the addition of new utterance nodes in the graph. this process begins with the verbalization of the graph (see section 3.3). the resulting description of the graph is then inserted into a prompt format for the language model. we use gpt-4 (openai, 2023) and llama 2 (touvron et al., 2023) as the llms for our experiment. the gpt-4 model was accessed through the openai api while the llama-2 model, which is specifically optimized for dialogue use cases4, was run on a local gpu, using quantization methods (dettmers et al., 2022b,a) to satisfy memory limitations. to evaluate the impact of the graph verbalization on the dialogues, we test the following setups with these models: 3. https://cloud.google.com/speech-to-text 4. https://huggingface.co/meta-llama/llama-2-13b-chat-hf 75 https://cloud.google.com/speech-to-text https://huggingface.co/meta-llama/llama-2-13b-chat-hf walker, ultes and lison model factuality adequacy gpttriples 3.44 3.69 gptverbal 4.06 3.88 llamaverbal 3.68 3.79 table 2: average participant scores for the model factuality and conversational adequacy. 1. the gpt-4 model provided with a prompt containing the verbalized dialogue state graph. 2. the llama 2 model provided with a prompt containing the verbalized dialogue state graph. 3. the gpt-4 model provided with a prompt containing the dialogue state graph expressed as semantic triples of the state graph facts. the first two models make use of the verbalized graph as produced by the transformation function and selected parameters from section 3.3. the prompt to these models is therefore a concatenation of the general prompt describing the system’s role and the resulting text description. for the third model, the graph information is instead provided as a list of semantic triples directly describing the edges of the graph, as illustrated in figure 4. due to technical issues encountered with a few interactions, there is a slight imbalance in the number of dialogues collected for each model (gpt-triples: 16, gpt-verbal: 17, llama: 19). this, however, does not affect the statistical significance of the experimental results, detailed below. the llm-generated robot responses are sent to the text-to-speech engine embedded on the pepper platform. when speaking, the robot also makes basic hand gestures and movement to accompany its response, using the off-the-shelf library made available on the pepper platform. 4.3 user evaluation right after the conclusion of each dialogue, the participants were asked to rate two statements about the dialogue on a likert scale of 1 to 5 for two statements where 1 is “never”, 2 is “mostly not”, 3 “sometimes”, 4 “mostly”, and 5 “always”: statement 1 (factuality): the robot responded factually according to what we saw on the tour. statement 2 (conversational adequacy): the robot answered my questions and responded in a relevant, natural, and concise manner. table 2 shows the average scores given to each model for both statements, while figure 7 provides a box plot of the user scores for statement 1. the user ratings show that participants generally rated the models using verbalized documents higher in factuality than the model based on semantic triples, while providing similar scores to all models for the second statement regarding the conversational adequacy of the robot responses. to statistically evaluate the difference in scores between the models, we use the wilcoxon signed rank test (wilcoxon, 1945) with the typical threshold for significance α = 0.05 and the required number of samples n = 16. we first evaluate both alternative models pairwise against the baseline model. regarding the first statement, between the gpt-verbal model and the gpt 76 a graph-to-text approach to knowledge-grounded response generation in hri u s e r s c o r e w 1 2 4 5 미 factuality gpt-triples gpt-verbal model llama u s e r s c o r e 3 1 2 4 5 미 adequacy gpt-triples gpt-verbal model llama 미 figure 7: box plots of the results for each statement: the dotted line in each box represents the median, while the white inner square displays the mean user score for each model. model based on semantic triples we find a difference of p = 0.03,w = 1.86. we found no statistically significant difference between the scores of the llama model and the semantic triples model (p = 0.13, w = 1.13) for this statement. likewise, we find no statistically significant difference between either gpt-verbal (p = 0.25, w = 0.67) or llama (p = 0.29,w = 0.54) scores with the semantic triples model for the second statement. while the pairwise difference in distributions of scores for factuality between gpt-verbal and the baseline were below our threshold α = 0.05, the experiment results may also be looked at from the perspective of all models. that is, whether for all models we can say there is a statistically significant difference in outcomes (whatever model is chosen of the three). to correct for multiple comparisons of this sort, we apply bonferroni correction to provide a conservative bound on the resulting p-values amongst all model comparisons. with bonferroni correction, between the gptverbal model and the baseline triples model the adjusted value for factuality is p = 0.1, and for adequacy p = 0.75. for the llama model, the values are p = 0.39 and p = 0.88, respectively. 4.4 analysis 4.4.1 general observations qualitatively, the majority of factuality errors across the three models were due to objects that were not detected by the vqa model employed for object and person detection, and were consequently not mentioned in the provided prompt (either as semantic triples or verbalized graph). although those omissions do indeed result in factually incorrect statements from the robot, there is little that can be done from the dialogue side to correct this shortcoming. the setup based on semantic triples demonstrates a decent performance on both statements rated by the participants. this model was however more prone to asserting it did not have information available, and was generally unable to provide details as to the order of rooms visited. although the times rooms were entered were available to the model based on semantic triples, the structure of the input made it difficult for the llm to provide informed answers over the tour chronology. due to limited gpu resources, the llama model had much higher latency in response time. this condition did not appear to affect the resultant scores, but upon inspection it appears to have 77 walker, ultes and lison resulted in more independent questions rather than ones building upon each other. participants generally appeared to more fluently express themselves to the gpt models, sometimes with offtopic utterances. moreover, the llama model’s responses to questions about what it saw tended to lean to absolute “yes” and “no”. the llama model lastly tended to provide numbered lists of observations within the tour, resulting in more verbose conversational behavior. 4.4.2 negation words to further explore the difference in factuality scores, we also analyzed the occurrence of negation words. the following list was used: “no”, “not”, “can’t”, “don’t”, “unable” and “cannot”. as shown in table 3, the number of negation words was much higher in the model based on semantic triples. the higher occurrence of those words reflects the tendency of this model to refrain from providing certain details, particularly the order of visited locations and the time spent in them. model “not” “cannot” “no” other* total gpttriples 69 13 14 16 112 gptverbal 48 3 2 15 68 llamaverbal 56 0 28 7 91 table 3: count of negation words by model. 4.4.3 expression of uncertainty to investigate the models’ ability to express probability and uncertainty, we analyzed a count of words expressing uncertainty such as “may” and “might”. the frequency of such words reveals that the gpt-4 model with the verbalized graph was substantially more likely to communicate uncertainty to the user, as shown in table 4. model “may” “might” “possibly” “possible” total gpttriples 4 0 2 0 6 gptverbal 25 7 7 1 40 llamaverbal 7 0 0 1 8 table 4: token counts of words expressing uncertainty by model. upon inspecting the dialogues, there appear to have been no instances where the gpt-verbal model expressed uncertainty with the word “may” that was not justified by the observation descriptions. by constrast, the gpt-triples model used “may” in this way twice, with the remaining count arising from utterances where the model asks the user to note that the observations may not describe certain aspects such as order or count of objects. this tendency appears to demonstrate that while probability of observations was presented in semantic triple form, the model made significantly more use of uncertainty in its expression when uncertainty was transformed to natural language. 78 a graph-to-text approach to knowledge-grounded response generation in hri 4.4.4 influence of model order we analyzed the effect of the model order on user scores, shown in table 5. there was not a substantial effect on scores from whether the model was presented first or last to the user. to analyze the interaction of model types with order, we fit a bayesian linear mixed model using the bambi5 library. we treat model order as a fixed effect with individual participant variability as a random effect, as score ∼ model ∗ order + (1|user). the results indicated that order had minimal impact on scores, with (β = −0.10, sd = 0.29, 94% hdi: [−0.64, 0.4]) for factuality and (β = −0.011, sd = 0.247, 94% hdi: [−0.467, 0.459]) for adequacy. model factuality (first) factuality (last) adequacy (first) adequacy (last) gpttriples 3.5 3.5 3.83 3.67 gptverbal 4.4 4 3.8 3.75 llamaverbal 3.71 3.86 3.86 3.71 table 5: average participant scores for model factuality and conversational adequacy when presented to the user first and last. 4.4.5 outliers as shown in the box plot of figure 7, the llama model had the widest range in scores, being the only model to receive scores of “1” to either statement. this score was recorded in two dialogues. in one case, the model frequently answered that it had not seen things in the tour after having enumerated them in an initial summary and demonstrated an unexpected inability to recall facts. this first dialogue is as follows and received a score of 2 for statement 2: user: which room did we spend the most time in robot: based on the observations, we spent the most time in the ”sick room” (205 seconds). user: what was the room we went to after the glass hallway robot: based on the observations, after the glass hallway, we went to the ”printer area” (32 seconds). user: do you see any useful equipment for extinguishing a fire anywhere robot: no, i do not see any useful equipment for extinguishing a fire anywhere in the observations. user: were there any chairs in the glass hallway robot: no, i did not see any chairs in the glass hallway. user: were there any people working in the office robot: yes, i saw a person working in the office. user: what furniture store in the sick room robot: there was no furniture store in the sick room. 5. https://bambinos.github.io/bambi/ 79 walker, ultes and lison while the first two observations were correct according to the recorded time, the model was unable to provide facts about the detected entities, namely a fire extinguisher and a chair in the hallway, furniture in the sick room (chair, sofa) and the absence of people in the office. in the second case, the model failed to describe detected objects from the tour and was given a score of 1 for both statements. however, all model scores given by the participant were quite low. this participant appeared to take a much stricter approach to interpretation of the statements in evaluating all models. the model was nonetheless weak in recalling detected objects in this dialogue. an excerpt from this dialogue is as follows: user: where did we go first robot: we first went to the main hallway. user: and after that robot: after the main hallway, we went to the print room. user: did you see any plants robot: no, i did not see any plants during my tour of the office floor. user: that’s not true robot: i apologize, you are correct. i did see a bookshelf in the first office i visited, but i did not see any plants. overall, the llama model also demonstrated a propensity to answer particularly short introductory greetings with a list of answers to the example questions provided in the prompt, a factor which may have influenced the scores. 4.5 dialogue examples dialogues 1 and 2 provide two additional examples of interactions, respectively based on the gpt-4 model relying on semantic triples and the gpt-4 model using the verbalized graph. as illustrated in dialogue 1, one peculiarity of the responses based on semantic triples was the difficulty of providing information about seen objects in an area even when recalling it in other locations. this kind of error may relate to the model needing to attend to a number of individual facts about the plants in the data at once, thus overlooking facts where the information is available. while the information necessary to determine the chronological ordering of locations visited, time spent in locations, and distances traveled are present in the triple-structured information, the results suggest that it is difficult for the model to fully make use of this information. 5. discussion 5.1 assessment of factuality variability in how participants approached the tour and what was shown to the robot also affected the results. with regard to the first question, participants appeared to differ in their criteria of whether a statement by the robot was “correct” or not. a score of “5” being “always” factual could be interpreted to require that the model hallucinated precisely zero times, while other participants interpreted the statement as that the robot’s responses as a whole were always factual. one alternative evaluation strategy is to evaluate the factuality of the responses against a ground truth. however, defining such a ground truth is particularly difficult in human–robot interaction, as the agent must rely on imperfect and often error-prone sensors to perceive its physical environment. 80 a graph-to-text approach to knowledge-grounded response generation in hri gpt-triples user: can you summarize our tour robot: during our exploration of the office floor, we visited various locations such as corridors, a printer room, storage room, offices, a coffee break area, and a reception area. we found many items such as doors, blue folders, a copy machine, bookshelves, a tv, a light switch, and a wastebin. there were also different movements like turning right or left and proceeding forwards for various meters. i also identified the presence of people, but i cannot confirm their identities. the exploration began and ended at different times as indicated by the creation time of the objects and paths. user: how many people did we see robot: it’s unclear exactly how many people we saw because they were often identified simply as ‘person’, ‘man’, or ‘woman’. these labels could refer to the same or different individuals. therefore, i cannot accurately provide the total number of people we saw without further information. user: did we go up to the starting area robot: yes, there were several mentions of moving to a location labeled as the “starting area” during our exploration. user: how long did we spend in the coffee room robot: the observations do not provide specific information about the amount of time spent in the coffee break area. user: which area did we visit after the coffee room robot: the observations do not provide a sequential order of the locations visited, so it’s unclear which area was visited immediately after the coffee break area. user: when did you see the plants robot: according to the observations, plants were seen at different times. however, specific timestamps were not provided for all of these observations. the time a plant was seen in the “survive” area was at 08:49, one in corridor 2 doesn’t have a time, and the one in the reception area wasn’t time-stamped either. user: did we see people after we left the coffee area robot: the observations do not provide a sequential order of the events, so it’s unclear if we saw people after leaving the coffee break area. dialogue 1: example using the gpt-triples model. this dialogue was given by the participant a score of 3 for both statements. this dialogue illustrates questions to which the triples-based model had difficulty generating responses (highlighted in red), namely concerning the chronology of perceived objects and explored locations. 81 walker, ultes and lison gpt-verbal user: what’s up robot: as a robot, i don’t have personal experiences, but i have recently completed the exploration of an office floor. user: [utterance not recognized] robot: i’m sorry, but your previous message was not understood. could you please provide more details or rephrase the query? user: can you mention all the rooms we went to robot: yes, we visited the starting area, hallway, print the room, lisa’s office and the break room. user: how often do we go in the hallway robot: we visited the hallway once during the exploration. user: did you see any plants robot: yes, i observed a plant in the hallway during the exploration. i may have also seen a plant in lisa’s office, but i’m not certain. user: did you see any fruits robot: i may have seen fruit in the hallway and the break room, but the observations are uncertain. user: [utterance not recognized] robot: i’m sorry, but your previous message was not understood. could you please provide more details or rephrase the query? user: did you see any apples robot: the observations do not provide information on whether i saw any apples during the exploration. user: did you see any paintings robot: i may have seen a painting in the hallway, but the observation is uncertain. user: can you dance for me robot: as a text-based ai, i’m afraid i don’t have the capability to dance. user: did you see any paintings in leases office robot: the observations do not provide information on whether i saw any paintings in lisa’s office during the exploration. user: did we go to lashes office robot: there is no mention of a place called “lashes office” in the observations. we did go to a place called “lisa’s office”. user: did you see a ping pong table robot: the observations do not provide information on whether i saw a ping pong table during the exploration. user: did you see a coffee machine robot: the observations do not provide information on whether i saw a coffee machine during the exploration. dialogue 2: example using the gpt-verbal model. this dialogue was given a score of 4 for statement 1 and 3 for statement 2. this dialogue featured a number of speech recognition errors. additionally, the model responded to several questions with an expression of uncertainty about the observations (highlighted in blue). notably, the final turn was accurate in regard to the observations as no coffee machine was recorded by the object detection model. 82 a graph-to-text approach to knowledge-grounded response generation in hri 5.2 question types in general, the participants tended to focus quite heavily on asking questions related to what was seen in the tour, as opposed to distances traveled or relative locations. this tendency is natural given that it is simpler to evaluate factuality with respect to seen objects as opposed to less easily memorable aspects such as number of robot turns or exact times rooms were entered. this kind of question was handled well by all models. the questions posed by the users can be divided into several categories: • polar questions (e.g. “did you see a chair in the office?”, “did we go to a room called office?”) • wh-questions (“what did you see in the office?”, “what was the second room we visited?”) • indirect / open-ended questions (“how about after that?”, “please summarize the tour.”) moreover, these questions principally centered on 1) visual recognition (as noted above) and 2) routing and path description, along with 3) requests for summary, which came exclusively at the start of dialogues. polar questions requiring a yes or no answer were typically handled well by all models (although as noted above the llama model encountered occasional repeated failures within a dialogue). the advantages of our proposed model are best illustrated in questions which centered on the second category of questions, thereby taking into account temporal ordering of locations or probability of objects. answers to indirect questions were often improved in this scenario, as the semantic triples model was unable to make use of timestamps to discern the order of events. this element of information was also important when the model was asked about distances travelled, given that the information was split between locations. 5.3 comparison between approaches a general benefit of verbalization compared to semantic triples is that the model has access to natural language describing the observations that it can directly use in its answers. this is especially evident in the gpt verbal model’s ability to express the likelihood of having seen objects. while the structured representation presented the same data to the model (and in a more precise form with respect to probabilities), the llm was not able to effectively use this aspect of the data. thus, verbalization avoids relying on the model to make logical inferences beyond its ordinary capabilities. with regards to temporal data, a verbalized form of the information avoids the necessity of the model estimating time spans across disjoint facts from the graph. related to this, when the probability of the observation is presented to the llm in natural language, the model is not required to translate that uncertainty to a specific linguistic expression. this translation does come with some design judgment: specifically, what is considered to be certain, likely, unlikely, or impossible? these judgments do however allow the verbalization to be sensitive to what a human would consider reasonable estimations of likelihood, rather than relying on an a priori estimation by the llm. another benefit of our approach is that the relation of the underlying information to the model prompt is explicit and human comprehensible. common sense intuitions which are not easily understood by llms within structured information can be made explicit in the process of verbalization. the advantage of the models’ strong fluidity in generating responses is preserved while explicitly representing factual information in the state graph and thereby the verbalized graph document. 83 walker, ultes and lison errors arising from failures in the underlying modules are explicit and more easily uncovered compared to dense vector representations which require specialized methods to decipher. 5.4 non-linguistic knowledge the reliance of the proposed approach on verbalized knowledge can make the model oblivious to content that is not easily expressed in a sentence. paralinguistic signals (schuller et al., 2013) such as body language, facial expressions or intonation are all crucial elements of communication which are not readily represented in text. other information like physical feedback in a robotic agent’s components may also not be readily represented. lacking verbalization, integrating these information sources with llms will likely remain challenging in the future. additionally, possible transformations of the dialogue state information might incorporate function calls, logical rule inference, and other methods to allow elements of dialogue state triples to be reformatted to assist the llm. for instance, the chronological elements might be summarized by a function, and then either ordered in a triples model or used as a starting point for our verbalization function. 5.5 limitations of the experimental study regarding the scope of this study, some limitations must be observed. first, the diversity of the study participants was limited to local researchers and students. another limitation with respect to experiment participants is that most participants were not native english speakers. additionally, many local place names were norwegian and thus difficult for an english asr system to accurately recognize. both factors led to various transcription errors that likely affected perceptions of model performance. further study in different populations and specifically with native english speakers would be a warranted follow-up to our experiments to evaluate the approach in broader audiences. a further limitation of the experimental setup is the comparative simplicity of the vqa model. the visual module did not distinguish between distinct entities of a single type in a labeled area, merely whether one was present. this setup avoided complex failures arising from difficulties of segmenting objects such as chairs, yet participants were often interested in the number of one kind of entity in an area. moreover, the vqa model only distinguished entities defined from the outset. while this allowed for a more well-defined set of entities of interest in an office environment, it limited the more creative or unusual possibilities for objects participants might ask about from the tour. some questions from the participants focused on the number of a particular item seen in the tour despite instruction that this information was not available to the robot. lastly, a more sophisticated vqa model may have been able to determine spatial relationships between entities themselves in a given room. doing so would require estimation of the robot’s field of view and position within the room. when an object was observed at the transition point between two rooms, before and after the user provided a label, any objects in the robot’s field of view could be attributed to both locations. in several instances this led to an apparent error of the model stating an object as having been observed in a room when it was seen at the end of its path in that area. our study did not incorporate verbal commands to the robot for movement. this choice was made to facilitate the execution of the office tour. additionally, the robot also suffers from substantial odometry drift, transformations of spatial data such as plotting the robot path in a building map were not possible. future work might use more precise measurements of the robot’s position and 84 a graph-to-text approach to knowledge-grounded response generation in hri integrate verbal commands for movement to the robot. lastly, continuous input of audio data to the model instead of relying on an external cue to begin speech recognition is possible. 6. conclusion in this work, we have presented a model of response generation in human-robot interaction where situated, dialogue-specific data is represented with a graph. as discussed in the preliminaries, graphs present a number of beneficial attributes that are useful for representing these kinds of information. namely, dialogue systems benefit from the ability to represent diverse types of information in the same structure along with semantic relationships between them. this structure also allows for flexible and effective transformation to natural language that expresses common sense intuitions about structured data. importantly for embodied agents, a dynamic knowledge graph can be continuously updated in response to real-time information. using these attributes, we presented a model which represents a dialogue state of diverse and continuously changing information about the physical world as graph of visual, spatial, and natural language data with temporal and probabilistic attributes. the information in this graph is converted to natural language descriptions for use by large-language models. to do so, a simple parameterized function converts paths within the graph to sentences describing the semantics of the relations described by the path. this function is optimized in a few-shot learning approach on manually annotated data. to assess the benefit of this approach, we have also presented a study in human-robot interaction where the robot responses are generated with an llm based on a prompt that includes a verbalized representation of the graph. this study consisted of users taking a robot for a tour of an office environment and subsequently discussing the tour with the robot. the users then rated the robot’s ability to accurately, concisely, and naturally summarize the tour. for each dialogue, the robot made use of two llms, namely gpt-4 and llama. for the llama model and one version of gpt-4, the graph information was presented to the llm as verbalized form of the dialogue state graph. while llms can make use of diversely formatted text, our study indicates that the gpt-4 model likely has improved ability to respond factually when presented data as a natural language document describing the graph information recorded in the graph. while a number of limitations exist within our study, the simplicity of the conversion from graph to natural language allows for the expression of relations between multimodal data in an intuitively useful way for llm-based response generation, perhaps particularly in hri settings where in-domain data is scarce and expensive to acquire. the impact of differing graph verbalization strategies also likely warrants further study, as graph structures and attributes can be verbalized in diverse ways. a promising area for future work is to determine which translations to surface forms produce the best performance improvements in downstream llms. these performance gains may also depend upon the performance of underlying modules such as visual-question answering models and robot sensors. different choices of embodied agents and physical sensors may allow for a more precise and wider variety of information to verbalize. longer explorations of the environment or more finegrained sensory information may require the incorporation of notions of relevance in order to reduce the amount of information presented to the llm. only a small part of the knowledge graph may be necessary to generate a response. retrieval-based generation (lewis et al., 2020; li et al., 2022a; walker et al., 2023) would enable a contextualized verbalization of the graph, where only the most relevant parts of the graph are retrieved and included in the input to the llm response generator. 85 walker, ultes and lison references michael ahn, anthony brohan, noah brown, yevgen chebotar, omar cortes, byron david, chelsea finn, chuyuan fu, keerthana gopalakrishnan, karol hausman, et al. do as i can, not as i say: grounding language in robotic affordances. arxiv preprint arxiv:2204.01691, 2022. url https://arxiv.org/pdf/2204.01691. takuya akiba, shotaro sano, toshihiko yanase, takeru ohta, and masanori koyama. optuna: a next-generation hyperparameter optimization framework. in proceedings of the 25th acm sigkdd international conference on knowledge discovery and data mining, 2019. url https://dl.acm.org/doi/10.1145/3292500.3330701. jean-baptiste alayrac, jeff donahue, pauline luc, antoine miech, iain barr, yana hasson, karel lenc, arthur mensch, katherine millican, malcolm reynolds, et al. flamingo: a visual language model for few-shot learning. advances in neural information processing systems, 35:23716– 23736, 2022. url https://proceedings.neurips.cc/paper_files/paper/ 2022/file/960a172bc7fbf0177ccccbb411a7d800-paper-conference.pdf. uri alon and eran yahav. on the bottleneck of graph neural networks and its practical implications. arxiv preprint arxiv:2006.05205, 2020. url https://arxiv.org/pdf/2006.05205. pdf. jacob andreas, john bufe, david burkett, charles chen, josh clausman, jean crawford, kate crim, jordan deloach, leah dorner, jason eisner, et al. task-oriented dialogue as dataflow synthesis. transactions of the association for computational linguistics, 8:556–571, 2020. url https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_ 00333/96470. laura banarescu, claire bonial, shu cai, madalina georgescu, kira griffitt, ulf hermjakob, kevin knight, philipp koehn, martha palmer, and nathan schneider. abstract meaning representation for sembanking. in proceedings of the 7th linguistic annotation workshop and interoperability with discourse, pages 178–186, 2013. url https://aclanthology.org/w13-2322. pdf. james bergstra, rémi bardenet, yoshua bengio, and balázs kégl. algorithms for hyper-parameter optimization. advances in neural information processing systems, 24, 2011. url https://proceedings.neurips.cc/paper_files/paper/2011/ file/86e8f7ab32cfd12577bc2619bc635690-paper.pdf. sajjad beygi, maryam fazel-zarandi, alessandra cervone, prakash krishnan, and siddhartha reddy jonnalagadda. logical reasoning for task oriented dialogue systems. arxiv preprint arxiv:2202.04161, 2022. url https://arxiv.org/pdf/2202.04161.pdf. erik billing, julia rosén, and maurice lamb. language models for human-robot interaction. in acm/ieee international conference on human-robot interaction, march 13–16, 2023, stockholm, sweden, pages 905–906. acm digital library, 2023. url https://dl.acm.org/ doi/abs/10.1145/3568294.3580040. 86 https://arxiv.org/pdf/2204.01691 https://dl.acm.org/doi/10.1145/3292500.3330701 https://proceedings.neurips.cc/paper_files/paper/2022/file/960a172bc7fbf0177ccccbb411a7d800-paper-conference.pdf https://proceedings.neurips.cc/paper_files/paper/2022/file/960a172bc7fbf0177ccccbb411a7d800-paper-conference.pdf https://arxiv.org/pdf/2006.05205.pdf https://arxiv.org/pdf/2006.05205.pdf https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00333/96470 https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00333/96470 https://aclanthology.org/w13-2322.pdf https://aclanthology.org/w13-2322.pdf https://proceedings.neurips.cc/paper_files/paper/2011/file/86e8f7ab32cfd12577bc2619bc635690-paper.pdf https://proceedings.neurips.cc/paper_files/paper/2011/file/86e8f7ab32cfd12577bc2619bc635690-paper.pdf https://arxiv.org/pdf/2202.04161.pdf https://dl.acm.org/doi/abs/10.1145/3568294.3580040 https://dl.acm.org/doi/abs/10.1145/3568294.3580040 a graph-to-text approach to knowledge-grounded response generation in hri rishi bommasani, drew a hudson, ehsan adeli, russ altman, simran arora, sydney von arx, michael s bernstein, jeannette bohg, antoine bosselut, emma brunskill, et al. on the opportunities and risks of foundation models. arxiv preprint arxiv:2108.07258, 2021. url https://arxiv.org/pdf/2108.07258.pdf?utm_source=morning_brew. claire bonial, lucia donatelli, mitchell abrams, stephanie lukin, stephen tratz, matthew marge, ron artstein, david traum, and clare voss. dialogue-amr: abstract meaning representation for dialogue. in proceedings of the 12th language resources and evaluation conference, pages 684–695, 2020. url https://aclanthology.org/2020.lrec-1.86.pdf. tom brown, benjamin mann, nick ryder, melanie subbiah, jared d kaplan, prafulla dhariwal, arvind neelakantan, pranav shyam, girish sastry, amanda askell, et al. language models are few-shot learners. advances in neural information processing systems, 33:1877–1901, 2020. url https://proceedings.neurips.cc/paper_files/paper/2020/ file/1457c0d6bfcb4967418bfb8ac142f64a-paper.pdf. sébastien bubeck, varun chandrasekaran, ronen eldan, johannes gehrke, eric horvitz, ece kamar, peter lee, yin tat lee, yuanzhi li, scott lundberg, et al. sparks of artificial general intelligence: early experiments with gpt-4. arxiv preprint arxiv:2303.12712, 2023. url https://arxiv.org/pdf/2303.12712.pdf?utm_source=webtekno. carlos celemin and jens kober. knowledge-and ambiguity-aware robot learning from corrective and evaluative feedback. neural computing and applications, pages 1–19, 2023. url https: //link.springer.com/article/10.1007/s00521-022-08118-z. wenhu chen. large language models are few (1)-shot table reasoners. arxiv preprint arxiv:2210.06710, 2022. url https://arxiv.org/pdf/2210.06710. philipp christmann, rishiraj saha roy, and gerhard weikum. explainable conversational question answering over heterogeneous sources via iterative graph neural networks. in proceedings of the 46th international acm sigir conference on research and development in information retrieval, pages 643–653, 2023. url https://dl.acm.org/doi/pdf/10.1145/ 3539618.3591682. tim dettmers, mike lewis, younes belkada, and luke zettlemoyer. llm.int8(): 8-bit matrix multiplication for transformers at scale. arxiv preprint arxiv:2208.07339, 2022a. url https: //arxiv.org/pdf/2208.07339.pdf?trk=public_post_comment-text. tim dettmers, mike lewis, sam shleifer, and luke zettlemoyer. 8-bit optimizers via block-wise quantization. 9th international conference on learning representations, iclr, 2022b. url https://arxiv.org/pdf/2110.02861. qingxiu dong, lei li, damai dai, ce zheng, zhiyong wu, baobao chang, xu sun, jingjing xu, and zhifang sui. a survey for in-context learning. arxiv preprint arxiv:2301.00234, 2022. url https://arxiv.org/pdf/2301.00234.pdf. david h douglas and thomas k peucker. algorithms for the reduction of the number of points required to represent a digitized line or its caricature. cartographica: the international journal 87 https://arxiv.org/pdf/2108.07258.pdf?utm_source=morning_brew https://aclanthology.org/2020.lrec-1.86.pdf https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-paper.pdf https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-paper.pdf https://arxiv.org/pdf/2303.12712.pdf?utm_source=webtekno https://link.springer.com/article/10.1007/s00521-022-08118-z https://link.springer.com/article/10.1007/s00521-022-08118-z https://arxiv.org/pdf/2210.06710 https://dl.acm.org/doi/pdf/10.1145/3539618.3591682 https://dl.acm.org/doi/pdf/10.1145/3539618.3591682 https://arxiv.org/pdf/2208.07339.pdf?trk=public_post_comment-text https://arxiv.org/pdf/2208.07339.pdf?trk=public_post_comment-text https://arxiv.org/pdf/2110.02861 https://arxiv.org/pdf/2301.00234.pdf walker, ultes and lison for geographic information and geovisualization, 10(2):112–122, 1973. url http://www2. ipcku.kansai-u.ac.jp/˜yasumuro/m_infomedia/paper/douglas73.pdf. danny driess, fei xia, mehdi sm sajjadi, corey lynch, aakanksha chowdhery, brian ichter, ayzaan wahid, jonathan tompson, quan vuong, tianhe yu, et al. palm-e: an embodied multimodal language model. arxiv preprint arxiv:2303.03378, 2023. url https://arxiv. org/pdf/2303.03378.pdf?trk=public_post_comment-text. nouha dziri, ximing lu, melanie sclar, xiang lorraine li, liwei jian, bill yuchen lin, peter west, chandra bhagavatula, ronan le bras, jena d hwang, et al. faith and fate: limits of transformers on compositionality. arxiv preprint arxiv:2305.18654, 2023. url https:// arxiv.org/pdf/2305.18654.pdf. mihail eric, lakshmi krishnan, francois charette, and christopher d. manning. key-value retrieval networks for task-oriented dialogue. in proceedings of the 18th annual sigdial meeting on discourse and dialogue, pages 37–49, saarbrücken, germany, august 2017. association for computational linguistics. doi: 10.18653/v1/w17-5506. url https://aclanthology. org/w17-5506. mihail eric, nicole chartier, behnam hedayatnia, karthik gopalakrishnan, pankaj rajan, yang liu, and dilek hakkani-tur. multi-sentence knowledge selection in open-domain dialogue. in proceedings of the 14th international conference on natural language generation, pages 76– 86, aberdeen, scotland, uk, august 2021. association for computational linguistics. url https://aclanthology.org/2021.inlg-1.9. vikas garg, stefanie jegelka, and tommi jaakkola. generalization and representational limits of graph neural networks. in international conference on machine learning, pages 3419–3430. pmlr, 2020. url https://proceedings.mlr.press/v119/garg20c/garg20c. pdf. michael a goodrich, alan c schultz, et al. human–robot interaction: a survey. foundations and trends® in human–computer interaction, 1(3):203–275, 2008. url https://dl.acm. org/doi/abs/10.1561/1100000005. milan gritta, gerasimos lampouras, and ignacio iacobacci. conversation graph: data augmentation, training, and evaluation for non-deterministic dialogue management. transactions of the association for computational linguistics, 9:36–52, 2021. url https://aclanthology. org/2021.tacl-1.3/. tanmay gupta and aniruddha kembhavi. visual programming: compositional visual reasoning without training. in proceedings of the ieee/cvf conference on computer vision and pattern recognition, pages 14953–14962, 2023. url https://openaccess.thecvf.com/ content/cvpr2023/papers/gupta_visual_programming_compositional_ visual_reasoning_without_training_cvpr_2023_paper.pdf. will hamilton, zhitao ying, and jure leskovec. inductive representation learning on large graphs. advances in neural information processing systems, 30, 2017. url https://proceedings.neurips.cc/paper/2017/file/ 5dd9db5e033da9c6fb5ba83c7a7ebea9-paper.pdf. 88 http://www2.ipcku.kansai-u.ac.jp/~yasumuro/m_infomedia/paper/douglas73.pdf http://www2.ipcku.kansai-u.ac.jp/~yasumuro/m_infomedia/paper/douglas73.pdf https://arxiv.org/pdf/2303.03378.pdf?trk=public_post_comment-text https://arxiv.org/pdf/2303.03378.pdf?trk=public_post_comment-text https://arxiv.org/pdf/2305.18654.pdf https://arxiv.org/pdf/2305.18654.pdf https://aclanthology.org/w17-5506 https://aclanthology.org/w17-5506 https://aclanthology.org/2021.inlg-1.9 https://proceedings.mlr.press/v119/garg20c/garg20c.pdf https://proceedings.mlr.press/v119/garg20c/garg20c.pdf https://dl.acm.org/doi/abs/10.1561/1100000005 https://dl.acm.org/doi/abs/10.1561/1100000005 https://aclanthology.org/2021.tacl-1.3/ https://aclanthology.org/2021.tacl-1.3/ https://openaccess.thecvf.com/content/cvpr2023/papers/gupta_visual_programming_compositional_visual_reasoning_without_training_cvpr_2023_paper.pdf https://openaccess.thecvf.com/content/cvpr2023/papers/gupta_visual_programming_compositional_visual_reasoning_without_training_cvpr_2023_paper.pdf https://openaccess.thecvf.com/content/cvpr2023/papers/gupta_visual_programming_compositional_visual_reasoning_without_training_cvpr_2023_paper.pdf https://proceedings.neurips.cc/paper/2017/file/5dd9db5e033da9c6fb5ba83c7a7ebea9-paper.pdf https://proceedings.neurips.cc/paper/2017/file/5dd9db5e033da9c6fb5ba83c7a7ebea9-paper.pdf a graph-to-text approach to knowledge-grounded response generation in hri huang he, hua lu, siqi bao, fan wang, hua wu, zheng-yu niu, and haifeng wang. learning to select external knowledge with multi-scale negative sampling. ieee/acm transactions on audio, speech, and language processing, pages 1–7, 2023. doi: 10.1109/taslp.2023.3301222. url https://ieeexplore.ieee.org/document/10202207. ehsan hosseini-asl, bryan mccann, chien-sheng wu, semih yavuz, and richard socher. a simple language model for task-oriented dialogue. advances in neural information processing systems, 33:20179–20191, 2020. url https://proceedings.neurips.cc/paper/ 2020/file/e946209592563be0f01c844ab2170f0c-paper.pdf. julian hough and david schlangen. it’s not what you do, it’s how you do it: grounding uncertainty for a simple robot. in proceedings of the 2017 acm/ieee international conference on humanrobot interaction, pages 274–282, 2017. url https://clp.ling.uni-potsdam.de/ publications/hough-2017-3.pdf. wenlong huang, fei xia, ted xiao, harris chan, jacky liang, pete florence, andy zeng, jonathan tompson, igor mordatch, yevgen chebotar, et al. inner monologue: embodied reasoning through planning with language models. in 6th annual conference on robot learning, 2022. url https://openreview.net/pdf?id=3r3pz5i0tye. vojtěch hudeček and ondřej dušek. are llms all you need for task-oriented dialogue? arxiv preprint arxiv:2304.06556, 2023. url https://arxiv.org/pdf/2304.06556. ziwei ji, nayeon lee, rita frieske, tiezheng yu, dan su, yan xu, etsuko ishii, ye jin bang, andrea madotto, and pascale fung. survey of hallucination in natural language generation. acm computing surveys, 55(12):1–38, 2023. url https://dl.acm.org/doi/10.1145/ 3571730. jaehun jung, bokyung son, and sungwon lyu. attnio: knowledge graph exploration with inand-out attention flow for knowledge-grounded dialogue. in proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), pages 3484–3497, online, november 2020. association for computational linguistics. doi: 10.18653/v1/2020. emnlp-main.280. url https://aclanthology.org/2020.emnlp-main.280. jack lanchantin, sainbayar sukhbaatar, gabriel synnaeve, yuxuan sun, kavya srinet, and arthur szlam. a data source for reasoning embodied agents. in proceedings of the aaai conference on artificial intelligence, volume 37, pages 8438–8446, 2023. url https://ojs.aaai.org/ index.php/aaai/article/view/26017/25789. patrick lewis, ethan perez, aleksandra piktus, fabio petroni, vladimir karpukhin, naman goyal, heinrich küttler, mike lewis, wen-tau yih, tim rocktäschel, et al. retrieval-augmented generation for knowledge-intensive nlp tasks. advances in neural information processing systems, 33:9459–9474, 2020. huayang li, yixuan su, deng cai, yan wang, and lemao liu. a survey on retrieval-augmented text generation. arxiv preprint arxiv:2202.01110, 2022a. xinghang li, di guo, huaping liu, and fuchun sun. embodied semantic scene graph generation. in conference on robot learning, pages 1585–1594. pmlr, 2022b. url https: //proceedings.mlr.press/v164/li22e/li22e.pdf. 89 https://ieeexplore.ieee.org/document/10202207 https://proceedings.neurips.cc/paper/2020/file/e946209592563be0f01c844ab2170f0c-paper.pdf https://proceedings.neurips.cc/paper/2020/file/e946209592563be0f01c844ab2170f0c-paper.pdf https://clp.ling.uni-potsdam.de/publications/hough-2017-3.pdf https://clp.ling.uni-potsdam.de/publications/hough-2017-3.pdf https://openreview.net/pdf?id=3r3pz5i0tye https://arxiv.org/pdf/2304.06556 https://dl.acm.org/doi/10.1145/3571730 https://dl.acm.org/doi/10.1145/3571730 https://aclanthology.org/2020.emnlp-main.280 https://ojs.aaai.org/index.php/aaai/article/view/26017/25789 https://ojs.aaai.org/index.php/aaai/article/view/26017/25789 https://proceedings.mlr.press/v164/li22e/li22e.pdf https://proceedings.mlr.press/v164/li22e/li22e.pdf walker, ultes and lison qika lin, rui mao, jun liu, fangzhi xu, and erik cambria. fusing topology contexts and logical rules in language models for knowledge graph completion. information fusion, 90:253–264, 2023. url http://ww.sentic.net/knowledge-graph-completion.pdf. pengfei liu, weizhe yuan, jinlan fu, zhengbao jiang, hiroaki hayashi, and graham neubig. pretrain, prompt, and predict: a systematic survey of prompting methods in natural language processing. acm computing surveys, 55(9):1–35, 2023. url https://dl.acm.org/doi/ pdf/10.1145/3560815?trk=public_post_comment-text. corey lynch, ayzaan wahid, jonathan tompson, tianli ding, james betker, robert baruch, travis armstrong, and pete florence. interactive language: talking to robots in real time. ieee robotics and automation letters, 2023. url https://ieeexplore.ieee.org/iel7/ 7083369/7339444/10182264.pdf. blanca miller and david feil-seifer. embodiment, situatedness, and morphology for humanoid robots interacting with people. humanoid robotics: a reference, pages 1–23, 2016. url https://rrl.cse.unr.edu/media/documents/2017/10. 10072f978-94-007-7194-9_130-1.pdf. matthias minderer, alexey gritsenko, austin stone, maxim neumann, dirk weissenborn, alexey dosovitskiy, aravindh mahendran, anurag arnab, mostafa dehghani, zhuoran shen, xiao wang, xiaohua zhai, thomas kipf, and neil houlsby. simple open-vocabulary object detection with vision transformers. eccv, 2022. url https://www.ecva.net/papers/eccv_ 2022/papers_eccv/papers/136700714.pdf. seungwhan moon, pararth shah, anuj kumar, and rajen subba. opendialkg: explainable conversational reasoning with attention-based walks over knowledge graphs. in proceedings of the 57th annual meeting of the association for computational linguistics, pages 845–854, florence, italy, july 2019. association for computational linguistics. doi: 10.18653/v1/p19-1081. url https://aclanthology.org/p19-1081. tomáš nekvinda and ondřej dušek. aargh! end-to-end retrieval-generation for task-oriented dialog. in proceedings of the 23rd annual meeting of the special interest group on discourse and dialogue, pages 283–297, 2022. url https://aclanthology.org/2022. sigdial-1.29.pdf. preksha nema, shreyas shetty, parag jain, anirban laha, karthik sankaranarayanan, and mitesh m. khapra. generating descriptions from structured data using a bifocal attention mechanism and gated orthogonalization. in proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long papers), pages 1539–1550, new orleans, louisiana, june 2018. association for computational linguistics. doi: 10.18653/v1/n18-1139. url https://aclanthology. org/n18-1139. openai. gpt-4 technical report, 2023. url https://arxiv.org/abs/2303.08774. alexandros papangelis and stefan ultes. towards meaningful, grounded conversations with intelligent agents. arxiv preprint arxiv:2006.15768, 2020. url https://arxiv.org/pdf/ 2006.15768. 90 http://ww.sentic.net/knowledge-graph-completion.pdf https://dl.acm.org/doi/pdf/10.1145/3560815?trk=public_post_comment-text https://dl.acm.org/doi/pdf/10.1145/3560815?trk=public_post_comment-text https://ieeexplore.ieee.org/iel7/7083369/7339444/10182264.pdf https://ieeexplore.ieee.org/iel7/7083369/7339444/10182264.pdf https://rrl.cse.unr.edu/media/documents/2017/10.10072f978-94-007-7194-9_130-1.pdf https://rrl.cse.unr.edu/media/documents/2017/10.10072f978-94-007-7194-9_130-1.pdf https://www.ecva.net/papers/eccv_2022/papers_eccv/papers/136700714.pdf https://www.ecva.net/papers/eccv_2022/papers_eccv/papers/136700714.pdf https://aclanthology.org/p19-1081 https://aclanthology.org/2022.sigdial-1.29.pdf https://aclanthology.org/2022.sigdial-1.29.pdf https://aclanthology.org/n18-1139 https://aclanthology.org/n18-1139 https://arxiv.org/abs/2303.08774 https://arxiv.org/pdf/2006.15768 https://arxiv.org/pdf/2006.15768 a graph-to-text approach to knowledge-grounded response generation in hri baolin peng, chenguang zhu, chunyuan li, xiujun li, jinchao li, michael zeng, and jianfeng gao. few-shot natural language generation for task-oriented dialog. in findings of the association for computational linguistics: emnlp 2020, pages 172–182, 2020. url https://aclanthology.org/2020.findings-emnlp.17.pdf. martin j pickering and simon garrod. prediction and embodiment in dialogue. european journal of social psychology, 39(7):1162–1168, 2009. url https://onlinelibrary.wiley. com/doi/abs/10.1002/ejsp.663. plamen prodanov and andrzej drygajlo. bayesian networks for spoken dialogue management in multimodal systems of tour-guide robots. in proceedings of the 8th european conference on speech communication and technology (eurospeech). eth zürich, 2003. url https://www.research-collection.ethz.ch/bitstream/handle/20. 500.11850/82551/1/eth-8214-01.pdf. ian c rankin, seth mccammon, and geoffrey a hollinger. robotic information gathering using semantic language instructions. in 2021 ieee international conference on robotics and automation (icra), pages 4882–4888. ieee, 2021. url https://par.nsf.gov/servlets/ purl/10315084. nir ratner, yoav levine, yonatan belinkov, ori ram, inbal magar, omri abend, ehud karpas, amnon shashua, kevin leyton-brown, and yoav shoham. parallel context windows for large language models. in proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), pages 6383–6402, 2023. url https: //aclanthology.org/2023.acl-long.352.pdf. vipula rawte, prachi priya, sm tonmoy, sm zaman, amit sheth, and amitava das. exploring the relationship between llm hallucinations and prompt linguistic nuances: readability, formality, and concreteness. arxiv preprint arxiv:2309.11064, 2023. url https://arxiv.org/ pdf/2309.11064. liliang ren, kaige xie, lu chen, and kai yu. towards universal dialogue state tracking. in proceedings of the 2018 conference on empirical methods in natural language processing, pages 2780–2786, 2018. url https://aclanthology.org/d18-1299.pdf. stephen roller, emily dinan, naman goyal, da ju, mary williamson, yinhan liu, jing xu, myle ott, eric michael smith, y-lan boureau, et al. recipes for building an open-domain chatbot. in proceedings of the 16th conference of the european chapter of the association for computational linguistics: main volume, pages 300–325, 2021. url https://aclanthology. org/2021.eacl-main.24.pdf. björn schuller, stefan steidl, anton batliner, felix burkhardt, laurence devillers, christian müller, and shrikanth narayanan. paralinguistics in speech and language—state-of-the-art and the challenge. computer speech & language, 27(1):4–39, 2013. dhruv shah, błażej osiński, sergey levine, et al. lm-nav: robotic navigation with large pretrained models of language, vision, and action. in conference on robot learning, pages 492–504. pmlr, 2023. url https://proceedings.mlr.press/v205/shah23b/shah23b. pdf. 91 https://aclanthology.org/2020.findings-emnlp.17.pdf https://onlinelibrary.wiley.com/doi/abs/10.1002/ejsp.663 https://onlinelibrary.wiley.com/doi/abs/10.1002/ejsp.663 https://www.research-collection.ethz.ch/bitstream/handle/20.500.11850/82551/1/eth-8214-01.pdf https://www.research-collection.ethz.ch/bitstream/handle/20.500.11850/82551/1/eth-8214-01.pdf https://par.nsf.gov/servlets/purl/10315084 https://par.nsf.gov/servlets/purl/10315084 https://aclanthology.org/2023.acl-long.352.pdf https://aclanthology.org/2023.acl-long.352.pdf https://arxiv.org/pdf/2309.11064 https://arxiv.org/pdf/2309.11064 https://aclanthology.org/d18-1299.pdf https://aclanthology.org/2021.eacl-main.24.pdf https://aclanthology.org/2021.eacl-main.24.pdf https://proceedings.mlr.press/v205/shah23b/shah23b.pdf https://proceedings.mlr.press/v205/shah23b/shah23b.pdf walker, ultes and lison freda shi, xinyun chen, kanishka misra, nathan scales, david dohan, ed h chi, nathanael schärli, and denny zhou. large language models can be easily distracted by irrelevant context. in international conference on machine learning, pages 31210–31227. pmlr, 2023. url https://proceedings.mlr.press/v202/shi23a/shi23a.pdf. chan hee song, jiaman wu, clayton washington, brian m sadler, wei-lun chao, and yu su. llmplanner: few-shot grounded planning for embodied agents with large language models. arxiv preprint arxiv:2212.04088, 2022. url https://embodied-ai.org/papers/2023/ 7.pdf. yan tai, weichen fan, zhao zhang, feng zhu, rui zhao, and ziwei liu. link-context learning for multimodal llms. arxiv preprint arxiv:2308.07891, 2023. url https://arxiv.org/ pdf/2308.07891. stefanie tellex, nakul gopalan, hadas kress-gazit, and cynthia matuszek. robots that use language. annual review of control, robotics, and autonomous systems, 3:25–55, 2020. url https://h2r.cs.brown.edu/wp-content/uploads/tellex20.pdf. blaise thomson, jost schatzmann, and steve young. bayesian update of dialogue state for robust dialogue systems. in 2008 ieee international conference on acoustics, speech and signal processing, pages 4937–4940. ieee, 2008. url http://svr-www.eng.cam.ac.uk/˜sjy/ papers/thsy08.pdf. romal thoppilan, daniel de freitas, jamie hall, noam shazeer, apoorv kulshreshtha, heng-tze cheng, alicia jin, taylor bos, leslie baker, yu du, et al. lamda: language models for dialog applications. arxiv preprint arxiv:2201.08239, 2022. url https://arxiv.org/pdf/ 2201.08239.pdf?trk=public_post_comment-text. david thulke, nico daheim, christian dugast, and hermann ney. efficient retrieval augmented generation from unstructured knowledge for task-oriented dialog. arxiv preprint arxiv:2102.04643, 2021. url https://arxiv.org/pdf/2102.04643. hugo touvron, thibaut lavril, gautier izacard, xavier martinet, marie-anne lachaux, timothée lacroix, baptiste rozière, naman goyal, eric hambro, faisal azhar, et al. llama: open and efficient foundation language models. arxiv preprint arxiv:2302.13971, 2023. url https: //arxiv.org/pdf/2302.13971. kai-wen tuan, yi-jyun chen, yi-chien lin, chun-ho kwok, hai-lun tu, and jason s. chang. learning to find translation of grammar patterns in parallel corpus. in proceedings of the 33rd conference on computational linguistics and speech processing (rocling 2021), pages 301–309, taoyuan, taiwan, october 2021. the association for computational linguistics and chinese language processing (aclclp). url https://aclanthology.org/2021. rocling-1.39. yi-lin tuan, yun-nung chen, and hung-yi lee. dykgchat: benchmarking dialogue generation grounding on dynamic knowledge graphs. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), pages 1855–1865, hong kong, china, 92 https://proceedings.mlr.press/v202/shi23a/shi23a.pdf https://embodied-ai.org/papers/2023/7.pdf https://embodied-ai.org/papers/2023/7.pdf https://arxiv.org/pdf/2308.07891 https://arxiv.org/pdf/2308.07891 https://h2r.cs.brown.edu/wp-content/uploads/tellex20.pdf http://svr-www.eng.cam.ac.uk/~sjy/papers/thsy08.pdf http://svr-www.eng.cam.ac.uk/~sjy/papers/thsy08.pdf https://arxiv.org/pdf/2201.08239.pdf?trk=public_post_comment-text https://arxiv.org/pdf/2201.08239.pdf?trk=public_post_comment-text https://arxiv.org/pdf/2102.04643 https://arxiv.org/pdf/2302.13971 https://arxiv.org/pdf/2302.13971 https://aclanthology.org/2021.rocling-1.39 https://aclanthology.org/2021.rocling-1.39 a graph-to-text approach to knowledge-grounded response generation in hri november 2019. association for computational linguistics. doi: 10.18653/v1/d19-1194. url https://aclanthology.org/d19-1194. tomer ullman. large language models fail on trivial alterations to theory-of-mind tasks. arxiv preprint arxiv:2302.08399, 2023. url https://arxiv.org/pdf/2302.08399.pdf. karthik valmeekam, alberto olmo, sarath sreedharan, and subbarao kambhampati. large language models still can’t plan (a benchmark for llms on planning and reasoning about change). arxiv preprint arxiv:2206.10498, 2022. url https://arxiv.org/pdf/2206.10498. ashish vaswani, noam shazeer, niki parmar, jakob uszkoreit, llion jones, aidan n gomez, łukasz kaiser, and illia polosukhin. attention is all you need. advances in neural information processing systems, 30, 2017. url https://proceedings.neurips.cc/paper/ 2017/file/3f5ee243547dee91fbd053c1c4a845aa-paper.pdf. joshua wainer, david j feil-seifer, dylan a shell, and maja j mataric. the role of physical embodiment in human-robot interaction. in roman 2006-the 15th ieee international symposium on robot and human interactive communication, pages 117–122. ieee, 2006. url https://rrl.cse.unr.edu/media/documents/2016/498.pdf. nicholas thomas walker, torbjørn dahl, and pierre lison. dialogue management as graph transformations. in conversational ai for natural human-centric interaction, pages 219–227, singapore, 2022. springer nature singapore. isbn 978-981-19-55389. url https://nr.brage.unit.no/nr-xmlui/bitstream/handle/11250/ 3055827/iwsds_2021.pdf?sequence=1. nicholas thomas walker, stefan ultes, and pierre lison. retrieval-augmented neural response generation using logical reasoning and relevance scoring. in proceedings of the 27th workshop on the semantics and pragmatics of dialogue, full papers, 2023. url https://arxiv. org/pdf/2310.13566.pdf. yanmeng wang, ye wang, xingyu lou, wenge rong, zhenghong hao, and shaojun wang. improving dialogue response generation via knowledge graph filter. in icassp 2021-2021 ieee international conference on acoustics, speech and signal processing (icassp), pages 7423– 7427. ieee, 2021. url https://ieeexplore.ieee.org/document/9414324. jason wei, yi tay, rishi bommasani, colin raffel, barret zoph, sebastian borgeaud, dani yogatama, maarten bosma, denny zhou, donald metzler, et al. emergent abilities of large language models. arxiv preprint arxiv:2206.07682, 2022. url https://arxiv.org/pdf/ 2206.07682.pdf?trk=cndc-detail. graham wilcock and kristiina jokinen. conversational ai and knowledge graphs for social robot interaction. in 2022 17th acm/ieee international conference on human-robot interaction (hri), pages 1090–1094. ieee, 2022. url https://helda.helsinki.fi/server/api/ core/bitstreams/60e2a300-e51d-4b86-8558-06e7326e7f3b/content. frank wilcoxon. individual comparisons by ranking methods. biometrics bulletin, 1(6):80–83, 1945. 93 https://aclanthology.org/d19-1194 https://arxiv.org/pdf/2302.08399.pdf https://arxiv.org/pdf/2206.10498 https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-paper.pdf https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-paper.pdf https://rrl.cse.unr.edu/media/documents/2016/498.pdf https://nr.brage.unit.no/nr-xmlui/bitstream/handle/11250/3055827/iwsds_2021.pdf?sequence=1 https://nr.brage.unit.no/nr-xmlui/bitstream/handle/11250/3055827/iwsds_2021.pdf?sequence=1 https://arxiv.org/pdf/2310.13566.pdf https://arxiv.org/pdf/2310.13566.pdf https://ieeexplore.ieee.org/document/9414324 https://arxiv.org/pdf/2206.07682.pdf?trk=cndc-detail https://arxiv.org/pdf/2206.07682.pdf?trk=cndc-detail https://helda.helsinki.fi/server/api/core/bitstreams/60e2a300-e51d-4b86-8558-06e7326e7f3b/content https://helda.helsinki.fi/server/api/core/bitstreams/60e2a300-e51d-4b86-8558-06e7326e7f3b/content walker, ultes and lison jason d williams, antoine raux, and matthew henderson. the dialog state tracking challenge series: a review. dialogue & discourse, 7(3):4–33, 2016. url https://core.ac.uk/ download/pdf/230775555.pdf. peng wu, bowei zou, ridong jiang, and aiti aw. gcdst: a graph-based and copy-augmented multi-domain dialogue state tracking. in findings of the association for computational linguistics: emnlp 2020, pages 1063–1073, online, november 2020. association for computational linguistics. doi: 10.18653/v1/2020.findings-emnlp.95. url https://aclanthology. org/2020.findings-emnlp.95. bowen xing and ivor w tsang. relational temporal graph reasoning for dual-task dialogue language understanding. ieee transactions on pattern analysis and machine intelligence, 2023. url https://arxiv.org/pdf/2306.09114.pdf. kun xu, lingfei wu, zhiguo wang, yansong feng, and vadim sheinin. sql-to-text generation with graph-to-sequence model. arxiv preprint arxiv:1809.05255, 2018. url https://arxiv. org/pdf/1809.05255.pdf. pengfei xu, xiaojun chang, ling guo, po-yao huang, xiaojiang chen, and alexander g hauptmann. a survey of scene graph: generation and application. ieee trans. neural netw. learn. syst, 1:1, 2020. url https://www.researchgate.net/ profile/xiaojun-chang-6/publication/340528049_a_survey_of_scene_ graph_generation_and_application/links/5e8eedb9a6fdcca789020994/ a-survey-of-scene-graph-generation-and-application.pdf. xiaoyu yang, feng nie, yufei feng, quan liu, zhigang chen, and xiaodan zhu. program enhanced fact verification with verbalization and graph attention network. arxiv preprint arxiv:2010.03084, 2020. url https://arxiv.org/pdf/2010.03084.pdf. steve young, milica gašić, blaise thomson, and jason d williams. pomdp-based statistical spoken dialog systems: a review. proceedings of the ieee, 101(5):1160–1179, 2013. url https: //cs.brown.edu/courses/csci2951-k/papers/young13.pdf. andy zeng, maria attarian, brian ichter, krzysztof choromanski, adrian wong, stefan welker, federico tombari, aveek purohit, michael ryoo, vikas sindhwani, et al. socratic models: composing zero-shot multimodal reasoning with language. arxiv preprint arxiv:2204.00598, 2022. url https://arxiv.org/pdf/2204.00598.pdf. min zhang, jingxiang chen, pengfei li, ming jiang, and zhe zhou. topic scene graphs for image captioning. iet computer vision, 16(4):364–375, 2022. url https://ietresearch. onlinelibrary.wiley.com/doi/pdf/10.1049/cvi2.12093. yizhe zhang, siqi sun, michel galley, yen-chun chen, chris brockett, xiang gao, jianfeng gao, jingjing liu, and william b dolan. dialogpt: large-scale generative pre-training for conversational response generation. in proceedings of the 58th annual meeting of the association for computational linguistics: system demonstrations, pages 270–278, 2020. url https://aclanthology.org/2020.acl-demos.30.pdf. 94 https://core.ac.uk/download/pdf/230775555.pdf https://core.ac.uk/download/pdf/230775555.pdf https://aclanthology.org/2020.findings-emnlp.95 https://aclanthology.org/2020.findings-emnlp.95 https://arxiv.org/pdf/2306.09114.pdf https://arxiv.org/pdf/1809.05255.pdf https://arxiv.org/pdf/1809.05255.pdf https://www.researchgate.net/profile/xiaojun-chang-6/publication/340528049_a_survey_of_scene_graph_generation_and_application/links/5e8eedb9a6fdcca789020994/a-survey-of-scene-graph-generation-and-application.pdf https://www.researchgate.net/profile/xiaojun-chang-6/publication/340528049_a_survey_of_scene_graph_generation_and_application/links/5e8eedb9a6fdcca789020994/a-survey-of-scene-graph-generation-and-application.pdf https://www.researchgate.net/profile/xiaojun-chang-6/publication/340528049_a_survey_of_scene_graph_generation_and_application/links/5e8eedb9a6fdcca789020994/a-survey-of-scene-graph-generation-and-application.pdf https://www.researchgate.net/profile/xiaojun-chang-6/publication/340528049_a_survey_of_scene_graph_generation_and_application/links/5e8eedb9a6fdcca789020994/a-survey-of-scene-graph-generation-and-application.pdf https://arxiv.org/pdf/2010.03084.pdf https://cs.brown.edu/courses/csci2951-k/papers/young13.pdf https://cs.brown.edu/courses/csci2951-k/papers/young13.pdf https://arxiv.org/pdf/2204.00598.pdf https://ietresearch.onlinelibrary.wiley.com/doi/pdf/10.1049/cvi2.12093 https://ietresearch.onlinelibrary.wiley.com/doi/pdf/10.1049/cvi2.12093 https://aclanthology.org/2020.acl-demos.30.pdf a graph-to-text approach to knowledge-grounded response generation in hri yue zhang, yafu li, leyang cui, deng cai, lemao liu, tingchen fu, xinting huang, enbo zhao, yu zhang, yulong chen, et al. siren’s song in the ai ocean: a survey on hallucination in large language models. arxiv preprint arxiv:2309.01219, 2023. url https: //arxiv.org/pdf/2309.01219.pdf. xufeng zhao, mengdi li, cornelius weber, muhammad burhan hafez, and stefan wermter. chat with the environment: interactive multimodal perception using large language models. arxiv preprint arxiv:2303.08268, 2023. url https://arxiv.org/pdf/2303.08268.pdf. appendix a. vqa objects “a person”, “a woman”, “a man”, “a computer”, “a desk”, “a table”, “a coffee mug”, “a chair”, “a whiteboard”, “a garbage can”, “a door”, “a window”, “a plant”, “a fire extinguisher”, “a pillar”, “fruit”, “apples”, “bananas”, “flowers”, “magazines”, “books”, “a computer mouse”, “a couch”, “a tv”, “a wastebin”, “a pen”, “a pencil”, “scissors”, “folders”, “a light switch”, ”cables”, ”notebooks”, ”paper”, “a printer”, “a bookshelf”, “bookshelves”, “a painting”, “a camera”, “food”, “cake”, “a light”, “a lamp”, “a stapler”, “a red folder”, “a blue folder”, “a yellow folder”, “a green folder”, “a broom”, “a blue curtain”, “a wooden pallet”, “a marker”, “a copy machine”, ”cardboard boxes” appendix b. prompt format “you are a robot that has just explored an office floor and recorded observations. answer each question in a short, concise sentence, from the robot perspective. example responses are: yes, i saw a woman in the main hallway. i may have seen a fire extinguisher, but i’m not sure. sorry, i didn’t see any plants in the office. we traveled 10 meters in the hallway. i saw a bookshelf in the first office, but not in the second office. we first went to an office, then to the reception area. only use the provided observations to arrive at your answer. if you don’t know the answer, just say that you don’t know. don’t try to make up an answer. observations: [triples or verbalized facts]” 95 https://arxiv.org/pdf/2309.01219.pdf https://arxiv.org/pdf/2309.01219.pdf https://arxiv.org/pdf/2303.08268.pdf introduction background embodied conversational agents knowledge graphs for dialogue graph-based representations graph processing large language models (llms) response generation with llms prompting methods llms in embodied systems approach state representation state tracking & postprocessing graph updates path generation graph-to-text transformation verbalization function parameter selection response generation experiments experimental setup system configuration robot perception dialogue system user evaluation analysis general observations negation words expression of uncertainty influence of model order outliers dialogue examples discussion assessment of factuality question types comparison between approaches non-linguistic knowledge limitations of the experimental study conclusion vqa objects prompt format